@dzhechkov/skills-bto 1.0.1 → 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md
CHANGED
|
@@ -31,6 +31,7 @@ After installation, open Claude Code in your project directory and start using B
|
|
|
31
31
|
| **Skill** | 1 | `bto` — core Build-Test-Optimize skill with modules (build, test, optimize) |
|
|
32
32
|
| **Commands** | 4 | `/bto`, `/bto-build`, `/bto-test`, `/bto-optimize` |
|
|
33
33
|
| **Rules** | 1 | `bto-quality-gates` — quality gate enforcement |
|
|
34
|
+
| **Shards** | 1 | `bto-evaluation` — context shard for BTO evaluation pipeline |
|
|
34
35
|
| **Agent Templates** | 2 | `bto-judge-panel`, `bto-optimizer-worker` |
|
|
35
36
|
| **References** | 4 | Eval patterns, judge rubrics, optimization methods, quality checklist |
|
|
36
37
|
| **Examples** | 1 | Sample evaluation report |
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@dzhechkov/skills-bto",
|
|
3
|
-
"version": "1.0
|
|
3
|
+
"version": "1.1.0",
|
|
4
4
|
"description": "Build-Test-Optimize skill pack for Claude Code — structured BTO pipeline with quality gates, testing strategies, and optimization workflows",
|
|
5
5
|
"main": "src/cli.js",
|
|
6
6
|
"bin": {
|
|
@@ -0,0 +1,48 @@
|
|
|
1
|
+
# BTO Evaluation — Governance Shard
|
|
2
|
+
|
|
3
|
+
## Skills
|
|
4
|
+
Load: `.claude/skills/bto/SKILL.md` + relevant module
|
|
5
|
+
|
|
6
|
+
## Layer Architecture
|
|
7
|
+
|
|
8
|
+
| Layer | Model | Gate | Action on Fail |
|
|
9
|
+
|-------|-------|------|----------------|
|
|
10
|
+
| Layer 0 | — (deterministic) | ≥ 80% checks pass | STOP, auto-retry up to 3x |
|
|
11
|
+
| Layer 1 | haiku | avg ≥ 7.0 | NEEDS WORK (flag) |
|
|
12
|
+
| Layer 2 | sonnet × 3 judges | weighted avg ≥ 7.0 | FAIL |
|
|
13
|
+
| Meta | sonnet | disagreement > 3 | Arbitrate |
|
|
14
|
+
|
|
15
|
+
## Judge Isolation — INVARIANT
|
|
16
|
+
- Each judge reads the SAME artifact
|
|
17
|
+
- Each judge writes to a SEPARATE evaluation
|
|
18
|
+
- Judges do NOT see each other's scores before submitting
|
|
19
|
+
- Generator and judge models MUST differ
|
|
20
|
+
|
|
21
|
+
## Weights
|
|
22
|
+
- Domain Expert: 0.40
|
|
23
|
+
- Critic: 0.30
|
|
24
|
+
- Completeness Auditor: 0.30
|
|
25
|
+
|
|
26
|
+
## Optimization Delta Gate
|
|
27
|
+
- Accepted if: new_score - prev_score > 0.5
|
|
28
|
+
- 3 consecutive iterations ≤ 0.5 delta → convergence, stop
|
|
29
|
+
- Score decrease > 1.0 → rollback to previous best
|
|
30
|
+
|
|
31
|
+
## Model Routing
|
|
32
|
+
- Layer 0: deterministic (no LLM)
|
|
33
|
+
- Layer 1: haiku
|
|
34
|
+
- Layer 2 judges: sonnet
|
|
35
|
+
- Meta-judge: sonnet
|
|
36
|
+
- Crossover synthesis: opus
|
|
37
|
+
- Variant fast-eval: haiku
|
|
38
|
+
|
|
39
|
+
## Promises
|
|
40
|
+
- `<promise>BTO_LAYER0_PASSED</promise>` — after Layer 0
|
|
41
|
+
- `<promise>BTO_LAYER2_SCORED</promise>` — after Layer 2
|
|
42
|
+
- `<promise>BTO_OPTIMIZED</promise>` — after optimization converges
|
|
43
|
+
|
|
44
|
+
## Anti-Patterns
|
|
45
|
+
- Score inflation (all > 8.5 first attempt) → add calibration
|
|
46
|
+
- Conformity collapse (identical scores) → enforce isolation
|
|
47
|
+
- Runaway optimization (> 10 iterations) → abort
|
|
48
|
+
- Judge-generator collusion (same model) → BLOCK
|