@dzhechkov/skills-bto 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bin/cli.js +5 -0
- package/package.json +43 -0
- package/src/cli.js +150 -0
- package/src/commands/doctor.js +366 -0
- package/src/commands/init.js +188 -0
- package/src/commands/list.js +161 -0
- package/src/commands/remove.js +211 -0
- package/src/commands/update.js +198 -0
- package/src/utils.js +398 -0
- package/templates/.claude/agents/bto-judge-panel.md +192 -0
- package/templates/.claude/agents/bto-optimizer-worker.md +181 -0
- package/templates/.claude/commands/bto-build.md +169 -0
- package/templates/.claude/commands/bto-optimize.md +208 -0
- package/templates/.claude/commands/bto-test.md +186 -0
- package/templates/.claude/commands/bto.md +171 -0
- package/templates/.claude/rules/bto-quality-gates.md +91 -0
- package/templates/.claude/skills/bto/SKILL.md +266 -0
- package/templates/.claude/skills/bto/examples/sample-eval-report.md +436 -0
- package/templates/.claude/skills/bto/modules/build.md +189 -0
- package/templates/.claude/skills/bto/modules/optimize.md +201 -0
- package/templates/.claude/skills/bto/modules/test.md +348 -0
- package/templates/.claude/skills/bto/references/eval-patterns.md +202 -0
- package/templates/.claude/skills/bto/references/judge-rubrics.md +183 -0
- package/templates/.claude/skills/bto/references/optimization-methods.md +139 -0
- package/templates/.claude/skills/bto/references/quality-checklist.md +220 -0
|
@@ -0,0 +1,208 @@
|
|
|
1
|
+
# /bto-optimize — Evolutionary Prompt Optimization for Skills and Commands
|
|
2
|
+
|
|
3
|
+
## Usage
|
|
4
|
+
```
|
|
5
|
+
/bto-optimize [path to skill or command to optimize]
|
|
6
|
+
```
|
|
7
|
+
|
|
8
|
+
## Parameters
|
|
9
|
+
- $ARGUMENTS — Path to a skill directory (`.claude/skills/<name>/`) or artifact file to optimize. Optionally include a focus dimension (`METHODOLOGY`, `DEPTH`, `CORRECTNESS`, `USABILITY`, `ROBUSTNESS`) to bias mutation strategies toward a specific weakness. Optionally include "rounds N" (e.g., "rounds 5") to override the default 3 rounds.
|
|
10
|
+
|
|
11
|
+
## Protocol
|
|
12
|
+
|
|
13
|
+
### Step 1: Load Skill and Module
|
|
14
|
+
|
|
15
|
+
Read `.claude/skills/bto/SKILL.md`
|
|
16
|
+
Read `.claude/skills/bto/modules/optimize.md`
|
|
17
|
+
|
|
18
|
+
### Step 2: Validate Input
|
|
19
|
+
|
|
20
|
+
If $ARGUMENTS is empty:
|
|
21
|
+
- Ask: "Provide a path to the skill or command you want to optimize."
|
|
22
|
+
- Stop and wait.
|
|
23
|
+
|
|
24
|
+
Resolve the artifact path from $ARGUMENTS:
|
|
25
|
+
- If path does not exist → report "Path not found: [path]" and stop
|
|
26
|
+
- If path is a directory → use `SKILL.md` inside it as the primary optimization target
|
|
27
|
+
- Parse optional focus dimension from $ARGUMENTS (if provided)
|
|
28
|
+
- Parse optional rounds override from $ARGUMENTS (default: 3, max: 5)
|
|
29
|
+
|
|
30
|
+
### Step 3: Baseline Evaluation
|
|
31
|
+
|
|
32
|
+
**Run TEST internally before generating any variants.**
|
|
33
|
+
|
|
34
|
+
Read `.claude/skills/bto/modules/test.md`
|
|
35
|
+
Read `.claude/skills/bto/references/judge-rubrics.md`
|
|
36
|
+
|
|
37
|
+
Execute TEST at Layer 2 (full judge panel, 3 parallel sonnet agents):
|
|
38
|
+
|
|
39
|
+
- Spawn 3 parallel agents (same as `/bto-test` Layer 2)
|
|
40
|
+
- Aggregate scores with weights: Expert 0.40, Critic 0.30, Auditor 0.30
|
|
41
|
+
- Record per-dimension baseline scores
|
|
42
|
+
- Record overall `BASELINE_SCORE`
|
|
43
|
+
- Record key weaknesses from each judge
|
|
44
|
+
|
|
45
|
+
**Early exit condition:**
|
|
46
|
+
If `BASELINE_SCORE` ≥ 8.0 → report:
|
|
47
|
+
```
|
|
48
|
+
Artifact already high quality (BASELINE_SCORE/10).
|
|
49
|
+
Full optimization not recommended. Specific minor suggestions:
|
|
50
|
+
[list top 3 suggestions from baseline judges]
|
|
51
|
+
```
|
|
52
|
+
Stop and wait for user to confirm they still want to proceed.
|
|
53
|
+
|
|
54
|
+
**Identify target dimensions:** all dimensions with score < 7.0.
|
|
55
|
+
If a focus dimension was specified in $ARGUMENTS → add it to target dimensions regardless of score.
|
|
56
|
+
|
|
57
|
+
### Step 4: Round 1 — Generate 5 Variants
|
|
58
|
+
|
|
59
|
+
Generate exactly 5 variants of the artifact. Each applies one mutation strategy from `optimize.md`:
|
|
60
|
+
|
|
61
|
+
| Variant | Mutation Strategy | Prioritize When |
|
|
62
|
+
|---------|-----------------|-----------------|
|
|
63
|
+
| Variant 1 | Rephrase | CLARITY / USABILITY weak |
|
|
64
|
+
| Variant 2 | Restructure | METHODOLOGY / USABILITY weak |
|
|
65
|
+
| Variant 3 | Add Constraints | ROBUSTNESS / CORRECTNESS weak |
|
|
66
|
+
| Variant 4 | Simplify | Artifact verbose or over-engineered |
|
|
67
|
+
| Variant 5 | Specialize | DEPTH weak, needs domain context |
|
|
68
|
+
|
|
69
|
+
If a focus dimension was specified, assign the 3 most relevant strategies to Variants 1-3 and use remaining strategies for Variants 4-5.
|
|
70
|
+
|
|
71
|
+
Generate all 5 variants sequentially (opus-level generation). For each variant, apply ONLY its assigned mutation. Preserve original intent, scope, and artifact type.
|
|
72
|
+
|
|
73
|
+
### Step 5: Evaluate Variants — Round 1
|
|
74
|
+
|
|
75
|
+
**Spawn 5 parallel agents using Agent tool, one per variant:**
|
|
76
|
+
|
|
77
|
+
Each agent runs Layer 1 evaluation (model: haiku) on its assigned variant.
|
|
78
|
+
|
|
79
|
+
Agents:
|
|
80
|
+
- Agent 1: "BTO Eval — Variant 1 (Rephrase)"
|
|
81
|
+
- Agent 2: "BTO Eval — Variant 2 (Restructure)"
|
|
82
|
+
- Agent 3: "BTO Eval — Variant 3 (Add Constraints)"
|
|
83
|
+
- Agent 4: "BTO Eval — Variant 4 (Simplify)"
|
|
84
|
+
- Agent 5: "BTO Eval — Variant 5 (Specialize)"
|
|
85
|
+
|
|
86
|
+
Record scores for all 5. Select top 2 by overall score.
|
|
87
|
+
|
|
88
|
+
### Step 6: Crossover — Generate Round 2 Variants
|
|
89
|
+
|
|
90
|
+
Combine the best elements of the top 2 variants from Round 1:
|
|
91
|
+
|
|
92
|
+
- Identify sections where Variant A scored higher → take from A
|
|
93
|
+
- Identify sections where Variant B scored higher → take from B
|
|
94
|
+
- Generate 3 crossover variants that combine strengths differently
|
|
95
|
+
|
|
96
|
+
Use crossover prompt from `optimize.md`. Output COMPLETE artifacts.
|
|
97
|
+
|
|
98
|
+
**Abort condition:** If any variant fails Layer 0 structural checks during crossover → discard that variant and continue with remaining.
|
|
99
|
+
|
|
100
|
+
### Step 7: Evaluate Variants — Round 2
|
|
101
|
+
|
|
102
|
+
**Spawn 3 parallel agents using Agent tool:**
|
|
103
|
+
|
|
104
|
+
Each agent runs Layer 1 evaluation (model: haiku) on its assigned crossover variant.
|
|
105
|
+
|
|
106
|
+
Agents:
|
|
107
|
+
- Agent 1: "BTO Eval — Round 2 Crossover A"
|
|
108
|
+
- Agent 2: "BTO Eval — Round 2 Crossover B"
|
|
109
|
+
- Agent 3: "BTO Eval — Round 2 Crossover C"
|
|
110
|
+
|
|
111
|
+
Record scores. Select top 2.
|
|
112
|
+
|
|
113
|
+
### Step 8: Crossover — Generate Round 3 Variants
|
|
114
|
+
|
|
115
|
+
Repeat crossover from Step 6 with the new top 2 from Round 2.
|
|
116
|
+
Generate 3 final crossover variants.
|
|
117
|
+
|
|
118
|
+
### Step 9: Final Evaluation — Round 3 (Full Panel)
|
|
119
|
+
|
|
120
|
+
**Round 3 uses Layer 2 — full 3-judge panel per variant.**
|
|
121
|
+
|
|
122
|
+
For each of the 3 final variants, spawn 3 parallel agents (model: sonnet):
|
|
123
|
+
|
|
124
|
+
Total: up to 9 agents running in parallel (or in 3 sequential batches of 3).
|
|
125
|
+
|
|
126
|
+
Select the single best variant as winner.
|
|
127
|
+
|
|
128
|
+
If `max_score(all_variants_round3) < BASELINE_SCORE + 0.3`:
|
|
129
|
+
- Report: "Optimization produced minimal improvement. Consider original."
|
|
130
|
+
- Still present the best variant for review.
|
|
131
|
+
|
|
132
|
+
**Abort conditions (from `optimize.md`):**
|
|
133
|
+
- Any round shows overall regression > 0.5 from baseline → stop immediately
|
|
134
|
+
- Artifact semantics change fundamentally → flag and stop
|
|
135
|
+
- User requests stop
|
|
136
|
+
|
|
137
|
+
### Step 10: Apply Winner
|
|
138
|
+
|
|
139
|
+
1. Write the winning variant to the original artifact path (overwrite in place)
|
|
140
|
+
2. Preserve original as `<filename>.pre-optimize.bak` in the same directory
|
|
141
|
+
|
|
142
|
+
### Step 11: Generate Report
|
|
143
|
+
|
|
144
|
+
Compute before/after delta per dimension.
|
|
145
|
+
|
|
146
|
+
**Winning strategy:** name the mutation strategy (or crossover combination) that produced the winner.
|
|
147
|
+
|
|
148
|
+
**Recommendation logic:**
|
|
149
|
+
- Improvement > 1.0 → "Apply changes — significant improvement"
|
|
150
|
+
- Improvement 0.5-1.0 → "Review changes before applying"
|
|
151
|
+
- Improvement < 0.5 → "Minimal improvement — original may be preferred"
|
|
152
|
+
|
|
153
|
+
---
|
|
154
|
+
|
|
155
|
+
## Checkpoint
|
|
156
|
+
|
|
157
|
+
```
|
|
158
|
+
═══════════════════════════════════════════════════════
|
|
159
|
+
CHECKPOINT: OPTIMIZE Complete
|
|
160
|
+
Artifact: [path]
|
|
161
|
+
Rounds run: [1-3]
|
|
162
|
+
Total evaluations: [N] (baseline + variants)
|
|
163
|
+
|
|
164
|
+
BEFORE → AFTER:
|
|
165
|
+
METHODOLOGY: X.X → X.X ([+/-]X.X)
|
|
166
|
+
DEPTH: X.X → X.X ([+/-]X.X)
|
|
167
|
+
CORRECTNESS: X.X → X.X ([+/-]X.X)
|
|
168
|
+
USABILITY: X.X → X.X ([+/-]X.X)
|
|
169
|
+
ROBUSTNESS: X.X → X.X ([+/-]X.X)
|
|
170
|
+
|
|
171
|
+
OVERALL: X.X → X.X ([+/-]X.X)
|
|
172
|
+
|
|
173
|
+
Winning Strategy: [strategy name or crossover combination]
|
|
174
|
+
|
|
175
|
+
CHANGELOG:
|
|
176
|
+
- [specific change made]
|
|
177
|
+
- [specific change made]
|
|
178
|
+
- [specific change made]
|
|
179
|
+
|
|
180
|
+
Recommendation: [Apply / Review / Original preferred]
|
|
181
|
+
|
|
182
|
+
Backup saved: [path.pre-optimize.bak]
|
|
183
|
+
|
|
184
|
+
• "ок" — done, keep optimized version
|
|
185
|
+
• "ещё раунд" — run one additional optimization round
|
|
186
|
+
• "откат" — restore original from backup
|
|
187
|
+
• "покажи diff" — show before/after diff for review
|
|
188
|
+
═══════════════════════════════════════════════════════
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
Wait for user confirmation.
|
|
192
|
+
|
|
193
|
+
---
|
|
194
|
+
|
|
195
|
+
## Modular Usage
|
|
196
|
+
|
|
197
|
+
This command is also invoked internally by `/bto` as the final step (OPTIMIZE phase) of the full pipeline, receiving `TEST_SCORE` and the artifact path from the preceding TEST step.
|
|
198
|
+
|
|
199
|
+
## Critical Rules
|
|
200
|
+
|
|
201
|
+
- Always establish baseline with Layer 2 before generating variants
|
|
202
|
+
- Never skip baseline — variants cannot be ranked without a reference point
|
|
203
|
+
- Agent tool is REQUIRED for parallel variant evaluation in all rounds
|
|
204
|
+
- Hard cap: 3 rounds by default, 5 maximum — never loop indefinitely
|
|
205
|
+
- Do not optimize if baseline ≥ 8.0 without explicit user confirmation
|
|
206
|
+
- Always save a `.pre-optimize.bak` backup before overwriting original
|
|
207
|
+
- Round 3 evaluation must use Layer 2 (sonnet judges), not Layer 1 (haiku)
|
|
208
|
+
- Preserve original artifact intent — optimization changes HOW, not WHAT
|
|
@@ -0,0 +1,186 @@
|
|
|
1
|
+
# /bto-test — Multi-Agent Evaluation of a Skill or Command
|
|
2
|
+
|
|
3
|
+
## Usage
|
|
4
|
+
```
|
|
5
|
+
/bto-test [path to skill directory or file]
|
|
6
|
+
```
|
|
7
|
+
|
|
8
|
+
## Parameters
|
|
9
|
+
- $ARGUMENTS — Path to a skill directory (`.claude/skills/<name>/`), a command file (`.claude/commands/<name>.md`), a rule file (`.claude/rules/<name>.md`), or any research artifact. Optionally include "full" to force Layer 2 panel regardless of Layer 1 score.
|
|
10
|
+
|
|
11
|
+
## Protocol
|
|
12
|
+
|
|
13
|
+
### Step 1: Load Skill and Module
|
|
14
|
+
|
|
15
|
+
Read `.claude/skills/bto/SKILL.md`
|
|
16
|
+
Read `.claude/skills/bto/modules/test.md`
|
|
17
|
+
Read `.claude/skills/bto/references/judge-rubrics.md`
|
|
18
|
+
|
|
19
|
+
### Step 2: Validate Input
|
|
20
|
+
|
|
21
|
+
If $ARGUMENTS is empty:
|
|
22
|
+
- Ask: "Provide a path to the skill directory or artifact file you want to evaluate."
|
|
23
|
+
- Stop and wait.
|
|
24
|
+
|
|
25
|
+
Resolve the artifact path from $ARGUMENTS:
|
|
26
|
+
- Strip any trailing slash
|
|
27
|
+
- If the path points to a directory → look for `SKILL.md` inside it as the primary file, but include all files in the directory for context
|
|
28
|
+
- If the path points to a file → use that file directly
|
|
29
|
+
- If the path does not exist → report "Path not found: [path]" and stop
|
|
30
|
+
|
|
31
|
+
### Step 3: Detect Artifact Type
|
|
32
|
+
|
|
33
|
+
Auto-detect from path pattern:
|
|
34
|
+
|
|
35
|
+
| Path Pattern | Detected Type |
|
|
36
|
+
|-------------|--------------|
|
|
37
|
+
| `.claude/skills/*/SKILL.md` or `.claude/skills/*/` | skill |
|
|
38
|
+
| `.claude/commands/*.md` | command |
|
|
39
|
+
| `.claude/rules/*.md` | rule |
|
|
40
|
+
| `.claude/agents/*.md` | agent |
|
|
41
|
+
| `researches/**/*.md` | research artifact |
|
|
42
|
+
|
|
43
|
+
If type cannot be determined from path → infer from file content structure.
|
|
44
|
+
|
|
45
|
+
### Step 4: Layer 0 — Deterministic Pre-checks
|
|
46
|
+
|
|
47
|
+
**Run Layer 0 first. Always. No exceptions. Zero cost.**
|
|
48
|
+
|
|
49
|
+
Execute all applicable checks from `test.md` for the detected artifact type:
|
|
50
|
+
|
|
51
|
+
- Universal checks (U1-U5): file exists, UTF-8, has headings, no excessive blanks, size in bounds
|
|
52
|
+
- Type-specific checks: skill (S1-S10), command (C1-C5), rule (R1-R4), agent (A1-A4)
|
|
53
|
+
|
|
54
|
+
Display Layer 0 output as a per-check PASS/FAIL table.
|
|
55
|
+
|
|
56
|
+
**Gate:** If pass rate < 80% → stop here. Do NOT proceed to Layer 1 or Layer 2.
|
|
57
|
+
Report all failures with specific line references where possible.
|
|
58
|
+
|
|
59
|
+
### Step 5: Layer 1 — Single LLM Judge (Quick)
|
|
60
|
+
|
|
61
|
+
Run if and only if Layer 0 gate passes.
|
|
62
|
+
|
|
63
|
+
**Model:** haiku
|
|
64
|
+
|
|
65
|
+
Execute the Layer 1 judge prompt from `test.md` against the artifact content.
|
|
66
|
+
|
|
67
|
+
Rate on 5 dimensions (1-10 each):
|
|
68
|
+
1. CLARITY
|
|
69
|
+
2. COMPLETENESS
|
|
70
|
+
3. ACTIONABILITY
|
|
71
|
+
4. QUALITY
|
|
72
|
+
5. ANTI-PATTERNS
|
|
73
|
+
|
|
74
|
+
Display scores, one-line justification per dimension, top 3 improvement suggestions, and VERDICT.
|
|
75
|
+
|
|
76
|
+
**Gate:**
|
|
77
|
+
- Average ≥ 7.0 → PASS
|
|
78
|
+
- Average 5.0-6.9 → NEEDS WORK (offer Layer 2)
|
|
79
|
+
- Average < 5.0 → FAIL (Layer 2 recommended)
|
|
80
|
+
|
|
81
|
+
Check $ARGUMENTS for "full":
|
|
82
|
+
- If "full" present → proceed directly to Layer 2 without asking
|
|
83
|
+
|
|
84
|
+
Otherwise, if Layer 1 score < 7.0 → automatically offer Layer 2.
|
|
85
|
+
If Layer 1 score ≥ 7.0 → ask: "Run full 3-judge panel for comprehensive evaluation? (yes/no)"
|
|
86
|
+
|
|
87
|
+
### Step 6: Layer 2 — Full Judge Panel (3 Parallel Agents)
|
|
88
|
+
|
|
89
|
+
Run if triggered by score threshold or user request.
|
|
90
|
+
|
|
91
|
+
**Spawn 3 agents in parallel using Agent tool:**
|
|
92
|
+
|
|
93
|
+
- Agent 1: "BTO Judge — Domain Expert" (model: sonnet)
|
|
94
|
+
- Focus: methodology, depth, technical correctness, domain fit
|
|
95
|
+
- Uses Judge 1 prompt from `test.md`
|
|
96
|
+
|
|
97
|
+
- Agent 2: "BTO Judge — Critic" (model: sonnet)
|
|
98
|
+
- Focus: gaps, weaknesses, anti-patterns, failure scenarios
|
|
99
|
+
- Uses Judge 2 prompt from `test.md`
|
|
100
|
+
|
|
101
|
+
- Agent 3: "BTO Judge — Completeness Auditor" (model: sonnet)
|
|
102
|
+
- Focus: structural coverage, cross-references, section completeness
|
|
103
|
+
- Uses Judge 3 prompt from `test.md`
|
|
104
|
+
|
|
105
|
+
**Isolation:** Each agent evaluates independently. No cross-communication.
|
|
106
|
+
|
|
107
|
+
**Aggregate scores after all 3 return:**
|
|
108
|
+
|
|
109
|
+
Weights: Expert 0.40, Critic 0.30, Auditor 0.30
|
|
110
|
+
|
|
111
|
+
Per-dimension weighted score:
|
|
112
|
+
```
|
|
113
|
+
dim_score = expert[dim] * 0.4 + critic[dim] * 0.3 + auditor[dim] * 0.3
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
Overall score = mean of all dimension scores.
|
|
117
|
+
|
|
118
|
+
**Disagreement detection:** For any dimension where `max(scores) - min(scores) > 3` → flag for meta-judge.
|
|
119
|
+
|
|
120
|
+
**Meta-judge (if triggered):**
|
|
121
|
+
- Model: default (opus)
|
|
122
|
+
- Present all 3 evaluations for the flagged dimension
|
|
123
|
+
- Reconcile to a single score with explicit reasoning
|
|
124
|
+
- Flag for human review if still unresolvable
|
|
125
|
+
|
|
126
|
+
### Step 7: Record TEST_SCORE
|
|
127
|
+
|
|
128
|
+
Set `TEST_SCORE` = overall weighted average (Layer 2 if run, else Layer 1 average).
|
|
129
|
+
|
|
130
|
+
This value is consumed by `/bto-optimize` and by the full `/bto` pipeline.
|
|
131
|
+
|
|
132
|
+
---
|
|
133
|
+
|
|
134
|
+
## Checkpoint
|
|
135
|
+
|
|
136
|
+
```
|
|
137
|
+
═══════════════════════════════════════════════════════
|
|
138
|
+
CHECKPOINT: TEST Complete
|
|
139
|
+
Artifact: [path]
|
|
140
|
+
Type: [detected type]
|
|
141
|
+
|
|
142
|
+
Layer 0: X/Y checks passed (Z%) — PASS / FAIL
|
|
143
|
+
Layer 1: X.X/10 — PASS / NEEDS WORK / FAIL
|
|
144
|
+
Layer 2: X.X/10 (weighted) — [if run, else "not run"]
|
|
145
|
+
|
|
146
|
+
OVERALL SCORE: X.X/10
|
|
147
|
+
|
|
148
|
+
Per-Dimension (Layer 2 or Layer 1):
|
|
149
|
+
METHODOLOGY: X.X
|
|
150
|
+
DEPTH: X.X
|
|
151
|
+
CORRECTNESS: X.X
|
|
152
|
+
USABILITY: X.X
|
|
153
|
+
ROBUSTNESS: X.X
|
|
154
|
+
|
|
155
|
+
[Flagged dimensions with judge disagreement, if any]
|
|
156
|
+
|
|
157
|
+
Top Improvements:
|
|
158
|
+
1. [specific suggestion]
|
|
159
|
+
2. [specific suggestion]
|
|
160
|
+
3. [specific suggestion]
|
|
161
|
+
|
|
162
|
+
• "ок" — done
|
|
163
|
+
• "покажи детали [expert/critic/auditor]" — expand judge feedback
|
|
164
|
+
• "оптимизируй" — run /bto-optimize with this artifact
|
|
165
|
+
• "запусти полный" — re-run with Layer 2 if only Layer 1 was run
|
|
166
|
+
═══════════════════════════════════════════════════════
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
Wait for user confirmation.
|
|
170
|
+
|
|
171
|
+
---
|
|
172
|
+
|
|
173
|
+
## Modular Usage
|
|
174
|
+
|
|
175
|
+
This command is also invoked internally by:
|
|
176
|
+
- `/bto` — as Step 4 (TEST phase) of the full pipeline
|
|
177
|
+
- `/bto-optimize` — as baseline evaluation before optimization rounds
|
|
178
|
+
|
|
179
|
+
## Critical Rules
|
|
180
|
+
|
|
181
|
+
- Always run Layer 0 before any LLM evaluation — it is free and fast
|
|
182
|
+
- Never spawn Layer 2 agents if Layer 0 fails (< 80% pass rate)
|
|
183
|
+
- Agent tool is REQUIRED for Layer 2 — do not run judges sequentially
|
|
184
|
+
- Report `TEST_SCORE` explicitly so downstream commands can consume it
|
|
185
|
+
- Rubrics from `judge-rubrics.md` govern scoring anchor points for all judges
|
|
186
|
+
- If "full" is in $ARGUMENTS, skip the confirmation prompt and run Layer 2 automatically
|
|
@@ -0,0 +1,171 @@
|
|
|
1
|
+
# /bto — Full BTO Pipeline (Build · Test · Optimize)
|
|
2
|
+
|
|
3
|
+
## Usage
|
|
4
|
+
```
|
|
5
|
+
/bto [path to existing skill OR description of new skill]
|
|
6
|
+
```
|
|
7
|
+
|
|
8
|
+
## Parameters
|
|
9
|
+
- $ARGUMENTS — Path to an existing skill/command directory or file, OR a natural language description of a new skill to create
|
|
10
|
+
|
|
11
|
+
## Protocol
|
|
12
|
+
|
|
13
|
+
### Step 1: Setup — Load Skill
|
|
14
|
+
|
|
15
|
+
Read `.claude/skills/bto/SKILL.md`
|
|
16
|
+
|
|
17
|
+
### Step 2: Route by Input Type
|
|
18
|
+
|
|
19
|
+
Inspect $ARGUMENTS to determine mode:
|
|
20
|
+
|
|
21
|
+
**If $ARGUMENTS is a path that exists on disk:**
|
|
22
|
+
- Mode: TEST → OPTIMIZE
|
|
23
|
+
- Skip BUILD
|
|
24
|
+
- Proceed to Step 4 (TEST)
|
|
25
|
+
|
|
26
|
+
**If $ARGUMENTS is a natural language description (not a path):**
|
|
27
|
+
- Mode: BUILD → TEST → OPTIMIZE
|
|
28
|
+
- Proceed to Step 3 (BUILD)
|
|
29
|
+
|
|
30
|
+
**If $ARGUMENTS is empty:**
|
|
31
|
+
- Ask the user: "Provide a path to an existing skill, or describe a new skill to build."
|
|
32
|
+
- Stop and wait.
|
|
33
|
+
|
|
34
|
+
---
|
|
35
|
+
|
|
36
|
+
### Step 3: BUILD (only if description provided)
|
|
37
|
+
|
|
38
|
+
Read `.claude/skills/bto/modules/build.md`
|
|
39
|
+
|
|
40
|
+
Execute the BUILD module:
|
|
41
|
+
1. Auto-detect artifact type from description
|
|
42
|
+
2. QUICK mode by default — parse requirements directly from description
|
|
43
|
+
3. If user said "deep" anywhere in $ARGUMENTS — activate DEEP mode (load `explore` skill first)
|
|
44
|
+
4. Generate complete artifact following build templates
|
|
45
|
+
5. Run self-review (Layer 0 structural check)
|
|
46
|
+
6. Create all files in the target directory
|
|
47
|
+
7. Record `BUILD_OUTPUT_PATH` for use in next steps
|
|
48
|
+
|
|
49
|
+
**Checkpoint BUILD:**
|
|
50
|
+
```
|
|
51
|
+
═══════════════════════════════════════════════════════
|
|
52
|
+
CHECKPOINT 1: BUILD Complete
|
|
53
|
+
Artifact generated and self-reviewed.
|
|
54
|
+
Files created: [list generated files]
|
|
55
|
+
Path: [BUILD_OUTPUT_PATH]
|
|
56
|
+
• "ок" — run TEST on the new artifact
|
|
57
|
+
• "переделай [aspect]" — rebuild with changes
|
|
58
|
+
• "углуби [section]" — expand a section
|
|
59
|
+
═══════════════════════════════════════════════════════
|
|
60
|
+
```
|
|
61
|
+
Wait for user confirmation before proceeding.
|
|
62
|
+
|
|
63
|
+
---
|
|
64
|
+
|
|
65
|
+
### Step 4: TEST
|
|
66
|
+
|
|
67
|
+
Read `.claude/skills/bto/modules/test.md`
|
|
68
|
+
Read `.claude/skills/bto/references/judge-rubrics.md`
|
|
69
|
+
|
|
70
|
+
**Resolve artifact path:**
|
|
71
|
+
- If BUILD was run → use `BUILD_OUTPUT_PATH`
|
|
72
|
+
- If path was provided in $ARGUMENTS → use that path
|
|
73
|
+
|
|
74
|
+
Execute TEST module in layers:
|
|
75
|
+
|
|
76
|
+
**Layer 0 — Deterministic pre-checks (always run first):**
|
|
77
|
+
- Run all structural checks for detected artifact type
|
|
78
|
+
- Display Layer 0 output with per-check PASS/FAIL
|
|
79
|
+
- If score < 80% → stop, report failures, do not proceed to LLM layers
|
|
80
|
+
|
|
81
|
+
**Layer 1 — Single LLM judge (haiku):**
|
|
82
|
+
- Run if Layer 0 passes
|
|
83
|
+
- Quick spot-check on all 5 dimensions
|
|
84
|
+
- Display scores and top 3 improvement suggestions
|
|
85
|
+
|
|
86
|
+
**Layer 2 — Full judge panel (user-triggered or if Layer 1 score < 7.0):**
|
|
87
|
+
- Ask: "Run full 3-judge panel for comprehensive evaluation? (yes/no)"
|
|
88
|
+
- If yes → spawn 3 parallel agents using Agent tool:
|
|
89
|
+
- Agent 1: "BTO Judge — Domain Expert" (model: sonnet)
|
|
90
|
+
- Agent 2: "BTO Judge — Critic" (model: sonnet)
|
|
91
|
+
- Agent 3: "BTO Judge — Completeness Auditor" (model: sonnet)
|
|
92
|
+
- Aggregate scores with weights: Expert 0.40, Critic 0.30, Auditor 0.30
|
|
93
|
+
- Flag any dimension where max - min > 3 (trigger meta-judge if needed)
|
|
94
|
+
- Display full evaluation report
|
|
95
|
+
|
|
96
|
+
Record `TEST_SCORE` (overall weighted average) for use in OPTIMIZE step.
|
|
97
|
+
|
|
98
|
+
**Checkpoint TEST:**
|
|
99
|
+
```
|
|
100
|
+
═══════════════════════════════════════════════════════
|
|
101
|
+
CHECKPOINT 2: TEST Complete
|
|
102
|
+
Artifact: [path]
|
|
103
|
+
Layer 0: X/Y checks passed
|
|
104
|
+
Layer 1: X.X/10 — [PASS / NEEDS WORK / FAIL]
|
|
105
|
+
Layer 2: X.X/10 (weighted) — [if run]
|
|
106
|
+
|
|
107
|
+
Overall: X.X/10
|
|
108
|
+
• "ок" — run OPTIMIZE
|
|
109
|
+
• "пропусти" — skip optimize (artifact is good)
|
|
110
|
+
• "покажи детали [judge]" — expand judge feedback
|
|
111
|
+
═══════════════════════════════════════════════════════
|
|
112
|
+
```
|
|
113
|
+
Wait for user confirmation before proceeding.
|
|
114
|
+
|
|
115
|
+
---
|
|
116
|
+
|
|
117
|
+
### Step 5: OPTIMIZE
|
|
118
|
+
|
|
119
|
+
Read `.claude/skills/bto/modules/optimize.md`
|
|
120
|
+
|
|
121
|
+
Execute OPTIMIZE module:
|
|
122
|
+
1. If `TEST_SCORE` ≥ 8.0 → report "Artifact already high quality (X.X/10). Minor tweaks only." Show specific suggestions. Stop.
|
|
123
|
+
2. If `TEST_SCORE` < 8.0 → run full evolutionary optimization:
|
|
124
|
+
- Round 1: Generate 5 variants (one per mutation strategy) → evaluate with Layer 1 (haiku) → select top 2
|
|
125
|
+
- Round 2: Crossover top 2 → generate 3 variants → evaluate with Layer 1 → select top 2
|
|
126
|
+
- Round 3: Crossover → generate 3 variants → evaluate with Layer 2 (3 parallel sonnet judges) → select winner
|
|
127
|
+
3. Apply the winning variant to the artifact file
|
|
128
|
+
4. Display before/after delta report
|
|
129
|
+
|
|
130
|
+
**Checkpoint OPTIMIZE:**
|
|
131
|
+
```
|
|
132
|
+
═══════════════════════════════════════════════════════
|
|
133
|
+
CHECKPOINT 3: OPTIMIZE Complete
|
|
134
|
+
Artifact: [path]
|
|
135
|
+
Rounds run: 3
|
|
136
|
+
Total evaluations: 15
|
|
137
|
+
|
|
138
|
+
BEFORE → AFTER:
|
|
139
|
+
METHODOLOGY: X.X → X.X (+X.X)
|
|
140
|
+
DEPTH: X.X → X.X (+X.X)
|
|
141
|
+
CORRECTNESS: X.X → X.X (+X.X)
|
|
142
|
+
USABILITY: X.X → X.X (+X.X)
|
|
143
|
+
ROBUSTNESS: X.X → X.X (+X.X)
|
|
144
|
+
|
|
145
|
+
OVERALL: X.X → X.X (+X.X)
|
|
146
|
+
|
|
147
|
+
Winning Strategy: [strategy name]
|
|
148
|
+
Recommendation: [Apply / Review / Original preferred]
|
|
149
|
+
• "ок" — done
|
|
150
|
+
• "ещё раунд" — run another optimization round
|
|
151
|
+
• "откат" — restore original artifact
|
|
152
|
+
═══════════════════════════════════════════════════════
|
|
153
|
+
```
|
|
154
|
+
|
|
155
|
+
---
|
|
156
|
+
|
|
157
|
+
## Modular Usage
|
|
158
|
+
|
|
159
|
+
Each module can be run independently:
|
|
160
|
+
- `/bto-build [description]` — BUILD only
|
|
161
|
+
- `/bto-test [path]` — TEST only
|
|
162
|
+
- `/bto-optimize [path]` — OPTIMIZE only
|
|
163
|
+
|
|
164
|
+
## Critical Rules
|
|
165
|
+
|
|
166
|
+
- Always run Layer 0 before any LLM evaluation — it is free and fast
|
|
167
|
+
- Never skip checkpoints — wait for user "ок" between modules
|
|
168
|
+
- Only run OPTIMIZE if TEST score < 8.0
|
|
169
|
+
- Agent tool is REQUIRED for Layer 2 parallel judge panel
|
|
170
|
+
- BUILD mode: QUICK by default, DEEP only if user explicitly requests it
|
|
171
|
+
- Record artifact path from BUILD and pass it through TEST → OPTIMIZE
|
|
@@ -0,0 +1,91 @@
|
|
|
1
|
+
# BTO Quality Gate Rules
|
|
2
|
+
|
|
3
|
+
## What is BTO
|
|
4
|
+
BTO (Build-Test-Optimize) is a pipeline for generating, evaluating, and iteratively
|
|
5
|
+
improving skills, prompts, or any structured artifact via agent-driven evaluation loops.
|
|
6
|
+
These rules apply to ANY evaluation system, not just Keysarium.
|
|
7
|
+
|
|
8
|
+
## Layer Architecture and Model Budget
|
|
9
|
+
|
|
10
|
+
| Layer | Role | Model | Trigger |
|
|
11
|
+
|-------|------|-------|---------|
|
|
12
|
+
| Layer 0 | Structural pre-check (format, completeness) | haiku | Always |
|
|
13
|
+
| Layer 1 | Shallow semantic check (relevance, coherence) | haiku | After Layer 0 passes |
|
|
14
|
+
| Layer 2 | Deep evaluation (quality, domain fit) | sonnet (judge panel) | After Layer 1 passes |
|
|
15
|
+
| Layer 3 | Creative synthesis / optimization crossover | opus | On top-N candidates only |
|
|
16
|
+
|
|
17
|
+
Never promote an artifact to a higher layer if the lower layer gate fails.
|
|
18
|
+
|
|
19
|
+
## Cost Optimization Table
|
|
20
|
+
|
|
21
|
+
| Task | Model | Rationale |
|
|
22
|
+
|------|-------|-----------|
|
|
23
|
+
| Layer 0 structural checks | haiku | High-frequency, pattern-matching only |
|
|
24
|
+
| Layer 1 semantic baseline | haiku | Fast coherence scan, no deep reasoning needed |
|
|
25
|
+
| Judge 1 — Domain Expert | sonnet | Domain knowledge + nuanced scoring |
|
|
26
|
+
| Judge 2 — Critic | sonnet | Adversarial analysis, pattern detection |
|
|
27
|
+
| Judge 3 — Completeness Auditor | sonnet | Structured coverage check |
|
|
28
|
+
| Meta-judge (escalation) | sonnet | Disagreement resolution |
|
|
29
|
+
| Crossover / creative synthesis | opus | Novel combination of best candidates |
|
|
30
|
+
| Mutation workers (standard) | sonnet | Requires reasoning about improvement direction |
|
|
31
|
+
| Variant fast-eval (ranking pass) | haiku | Volume scoring before full panel |
|
|
32
|
+
|
|
33
|
+
Escalate to a higher-cost model only when the lower-cost model has failed or is insufficient.
|
|
34
|
+
|
|
35
|
+
## Layer 0 Mandatory Checks
|
|
36
|
+
Every generated skill or artifact MUST pass ALL of these before entering judge panel:
|
|
37
|
+
|
|
38
|
+
- [ ] Required sections present (structure check)
|
|
39
|
+
- [ ] No empty placeholders (`[TODO]`, `[TBD]`, `<INSERT>`)
|
|
40
|
+
- [ ] Length within bounds (not below minimum, not above maximum)
|
|
41
|
+
- [ ] Encoding valid (no broken unicode, no binary artifacts)
|
|
42
|
+
- [ ] Self-reference loop absent (artifact does not cite itself as source)
|
|
43
|
+
|
|
44
|
+
If any check fails → reject immediately, log reason, do NOT send to judges.
|
|
45
|
+
Layer 0 may auto-retry up to 3 times before escalating to human review.
|
|
46
|
+
|
|
47
|
+
## Judge Panel Rules
|
|
48
|
+
|
|
49
|
+
- Panel MUST have an odd number of judges: 3 (standard) or 5 (high-stakes)
|
|
50
|
+
- Each judge operates in strict isolation: reads the same artifact, writes to a separate evaluation file
|
|
51
|
+
- Judges do NOT see each other's scores before submitting
|
|
52
|
+
- Final score = weighted average (weights defined per panel configuration)
|
|
53
|
+
- Standard weights: Domain Expert 0.4 / Critic 0.3 / Completeness Auditor 0.3
|
|
54
|
+
- Disagreement threshold: if max_score - min_score > 3 points → escalate to meta-judge
|
|
55
|
+
|
|
56
|
+
## Optimization Delta Gate
|
|
57
|
+
|
|
58
|
+
- An optimization iteration is accepted ONLY if: `new_score - prev_score > 0.5`
|
|
59
|
+
- If delta <= 0.5 for 3 consecutive iterations → declare convergence and stop
|
|
60
|
+
- If score DECREASES by > 1.0 → rollback to previous best and log regression
|
|
61
|
+
- Improvement must be measurable on the same rubric used in the previous iteration
|
|
62
|
+
|
|
63
|
+
## Human Checkpoint Rules
|
|
64
|
+
|
|
65
|
+
- NEVER auto-approve any artifact for delivery without a human checkpoint
|
|
66
|
+
- Checkpoint is required after: Layer 2 evaluation, final optimization round, before packaging
|
|
67
|
+
- Checkpoint format follows the standard checkpoint-protocol.md
|
|
68
|
+
- Exception: Layer 0 rejections may be auto-retried up to 3 times before human escalation
|
|
69
|
+
|
|
70
|
+
## BTO-Specific Anti-Patterns
|
|
71
|
+
|
|
72
|
+
| Anti-Pattern | Detection Signal | Required Fix |
|
|
73
|
+
|-------------|-----------------|--------------|
|
|
74
|
+
| Score inflation | All judges score > 8.5 on first attempt | Add calibration prompt to critics |
|
|
75
|
+
| Overfitting to rubric | Artifact optimizes wording to match rubric literally | Blind evaluation: hide rubric from generator |
|
|
76
|
+
| Conformity collapse | Judges converge to identical scores after 1 round | Enforce isolation, re-randomize judge order |
|
|
77
|
+
| Runaway optimization | > 10 iterations without convergence | Abort, log, human review |
|
|
78
|
+
| Phantom improvement | Delta > 0.5 but no substantive content change | Diff-check content, not just score |
|
|
79
|
+
| Judge-generator collusion | Same model used for both generation and evaluation | BLOCK — generator and judge models must differ |
|
|
80
|
+
| Missing rejection log | Failed artifacts silently discarded | Every rejection MUST be logged with reason |
|
|
81
|
+
|
|
82
|
+
## Auto-Detection
|
|
83
|
+
Self-check generated artifacts and evaluation results against the anti-patterns above.
|
|
84
|
+
If detected, flag with a WARNING label and halt the BTO loop pending human review.
|
|
85
|
+
|
|
86
|
+
## Reusability Note
|
|
87
|
+
These rules are artifact-type agnostic. Apply them to:
|
|
88
|
+
- Skill generation pipelines
|
|
89
|
+
- Prompt optimization loops
|
|
90
|
+
- Presentation scoring systems
|
|
91
|
+
- Any multi-judge evaluation workflow
|