@dzhechkov/skills-bto 1.0.1 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -31,6 +31,7 @@ After installation, open Claude Code in your project directory and start using B
31
31
  | **Skill** | 1 | `bto` — core Build-Test-Optimize skill with modules (build, test, optimize) |
32
32
  | **Commands** | 4 | `/bto`, `/bto-build`, `/bto-test`, `/bto-optimize` |
33
33
  | **Rules** | 1 | `bto-quality-gates` — quality gate enforcement |
34
+ | **Shards** | 1 | `bto-evaluation` — context shard for BTO evaluation pipeline |
34
35
  | **Agent Templates** | 2 | `bto-judge-panel`, `bto-optimizer-worker` |
35
36
  | **References** | 4 | Eval patterns, judge rubrics, optimization methods, quality checklist |
36
37
  | **Examples** | 1 | Sample evaluation report |
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@dzhechkov/skills-bto",
3
- "version": "1.0.1",
3
+ "version": "1.1.0",
4
4
  "description": "Build-Test-Optimize skill pack for Claude Code — structured BTO pipeline with quality gates, testing strategies, and optimization workflows",
5
5
  "main": "src/cli.js",
6
6
  "bin": {
@@ -0,0 +1,48 @@
1
+ # BTO Evaluation — Governance Shard
2
+
3
+ ## Skills
4
+ Load: `.claude/skills/bto/SKILL.md` + relevant module
5
+
6
+ ## Layer Architecture
7
+
8
+ | Layer | Model | Gate | Action on Fail |
9
+ |-------|-------|------|----------------|
10
+ | Layer 0 | — (deterministic) | ≥ 80% checks pass | STOP, auto-retry up to 3x |
11
+ | Layer 1 | haiku | avg ≥ 7.0 | NEEDS WORK (flag) |
12
+ | Layer 2 | sonnet × 3 judges | weighted avg ≥ 7.0 | FAIL |
13
+ | Meta | sonnet | disagreement > 3 | Arbitrate |
14
+
15
+ ## Judge Isolation — INVARIANT
16
+ - Each judge reads the SAME artifact
17
+ - Each judge writes to a SEPARATE evaluation
18
+ - Judges do NOT see each other's scores before submitting
19
+ - Generator and judge models MUST differ
20
+
21
+ ## Weights
22
+ - Domain Expert: 0.40
23
+ - Critic: 0.30
24
+ - Completeness Auditor: 0.30
25
+
26
+ ## Optimization Delta Gate
27
+ - Accepted if: new_score - prev_score > 0.5
28
+ - 3 consecutive iterations ≤ 0.5 delta → convergence, stop
29
+ - Score decrease > 1.0 → rollback to previous best
30
+
31
+ ## Model Routing
32
+ - Layer 0: deterministic (no LLM)
33
+ - Layer 1: haiku
34
+ - Layer 2 judges: sonnet
35
+ - Meta-judge: sonnet
36
+ - Crossover synthesis: opus
37
+ - Variant fast-eval: haiku
38
+
39
+ ## Promises
40
+ - `<promise>BTO_LAYER0_PASSED</promise>` — after Layer 0
41
+ - `<promise>BTO_LAYER2_SCORED</promise>` — after Layer 2
42
+ - `<promise>BTO_OPTIMIZED</promise>` — after optimization converges
43
+
44
+ ## Anti-Patterns
45
+ - Score inflation (all > 8.5 first attempt) → add calibration
46
+ - Conformity collapse (identical scores) → enforce isolation
47
+ - Runaway optimization (> 10 iterations) → abort
48
+ - Judge-generator collusion (same model) → BLOCK
@@ -1,5 +1,7 @@
1
1
  # BTO — Build · Test · Optimize
2
2
 
3
+ <!-- Trust Tier: 1 — Structured | Path to Tier 2: Run /bto-test for self-evaluation -->
4
+
3
5
  > Multi-agent evaluation and prompt optimization system for Claude Code artifacts.
4
6
 
5
7
  ## Overview