@dzhechkov/skills-bto 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,208 @@
1
+ # /bto-optimize — Evolutionary Prompt Optimization for Skills and Commands
2
+
3
+ ## Usage
4
+ ```
5
+ /bto-optimize [path to skill or command to optimize]
6
+ ```
7
+
8
+ ## Parameters
9
+ - $ARGUMENTS — Path to a skill directory (`.claude/skills/<name>/`) or artifact file to optimize. Optionally include a focus dimension (`METHODOLOGY`, `DEPTH`, `CORRECTNESS`, `USABILITY`, `ROBUSTNESS`) to bias mutation strategies toward a specific weakness. Optionally include "rounds N" (e.g., "rounds 5") to override the default 3 rounds.
10
+
11
+ ## Protocol
12
+
13
+ ### Step 1: Load Skill and Module
14
+
15
+ Read `.claude/skills/bto/SKILL.md`
16
+ Read `.claude/skills/bto/modules/optimize.md`
17
+
18
+ ### Step 2: Validate Input
19
+
20
+ If $ARGUMENTS is empty:
21
+ - Ask: "Provide a path to the skill or command you want to optimize."
22
+ - Stop and wait.
23
+
24
+ Resolve the artifact path from $ARGUMENTS:
25
+ - If path does not exist → report "Path not found: [path]" and stop
26
+ - If path is a directory → use `SKILL.md` inside it as the primary optimization target
27
+ - Parse optional focus dimension from $ARGUMENTS (if provided)
28
+ - Parse optional rounds override from $ARGUMENTS (default: 3, max: 5)
29
+
30
+ ### Step 3: Baseline Evaluation
31
+
32
+ **Run TEST internally before generating any variants.**
33
+
34
+ Read `.claude/skills/bto/modules/test.md`
35
+ Read `.claude/skills/bto/references/judge-rubrics.md`
36
+
37
+ Execute TEST at Layer 2 (full judge panel, 3 parallel sonnet agents):
38
+
39
+ - Spawn 3 parallel agents (same as `/bto-test` Layer 2)
40
+ - Aggregate scores with weights: Expert 0.40, Critic 0.30, Auditor 0.30
41
+ - Record per-dimension baseline scores
42
+ - Record overall `BASELINE_SCORE`
43
+ - Record key weaknesses from each judge
44
+
45
+ **Early exit condition:**
46
+ If `BASELINE_SCORE` ≥ 8.0 → report:
47
+ ```
48
+ Artifact already high quality (BASELINE_SCORE/10).
49
+ Full optimization not recommended. Specific minor suggestions:
50
+ [list top 3 suggestions from baseline judges]
51
+ ```
52
+ Stop and wait for user to confirm they still want to proceed.
53
+
54
+ **Identify target dimensions:** all dimensions with score < 7.0.
55
+ If a focus dimension was specified in $ARGUMENTS → add it to target dimensions regardless of score.
56
+
57
+ ### Step 4: Round 1 — Generate 5 Variants
58
+
59
+ Generate exactly 5 variants of the artifact. Each applies one mutation strategy from `optimize.md`:
60
+
61
+ | Variant | Mutation Strategy | Prioritize When |
62
+ |---------|-----------------|-----------------|
63
+ | Variant 1 | Rephrase | CLARITY / USABILITY weak |
64
+ | Variant 2 | Restructure | METHODOLOGY / USABILITY weak |
65
+ | Variant 3 | Add Constraints | ROBUSTNESS / CORRECTNESS weak |
66
+ | Variant 4 | Simplify | Artifact verbose or over-engineered |
67
+ | Variant 5 | Specialize | DEPTH weak, needs domain context |
68
+
69
+ If a focus dimension was specified, assign the 3 most relevant strategies to Variants 1-3 and use remaining strategies for Variants 4-5.
70
+
71
+ Generate all 5 variants sequentially (opus-level generation). For each variant, apply ONLY its assigned mutation. Preserve original intent, scope, and artifact type.
72
+
73
+ ### Step 5: Evaluate Variants — Round 1
74
+
75
+ **Spawn 5 parallel agents using Agent tool, one per variant:**
76
+
77
+ Each agent runs Layer 1 evaluation (model: haiku) on its assigned variant.
78
+
79
+ Agents:
80
+ - Agent 1: "BTO Eval — Variant 1 (Rephrase)"
81
+ - Agent 2: "BTO Eval — Variant 2 (Restructure)"
82
+ - Agent 3: "BTO Eval — Variant 3 (Add Constraints)"
83
+ - Agent 4: "BTO Eval — Variant 4 (Simplify)"
84
+ - Agent 5: "BTO Eval — Variant 5 (Specialize)"
85
+
86
+ Record scores for all 5. Select top 2 by overall score.
87
+
88
+ ### Step 6: Crossover — Generate Round 2 Variants
89
+
90
+ Combine the best elements of the top 2 variants from Round 1:
91
+
92
+ - Identify sections where Variant A scored higher → take from A
93
+ - Identify sections where Variant B scored higher → take from B
94
+ - Generate 3 crossover variants that combine strengths differently
95
+
96
+ Use crossover prompt from `optimize.md`. Output COMPLETE artifacts.
97
+
98
+ **Abort condition:** If any variant fails Layer 0 structural checks during crossover → discard that variant and continue with remaining.
99
+
100
+ ### Step 7: Evaluate Variants — Round 2
101
+
102
+ **Spawn 3 parallel agents using Agent tool:**
103
+
104
+ Each agent runs Layer 1 evaluation (model: haiku) on its assigned crossover variant.
105
+
106
+ Agents:
107
+ - Agent 1: "BTO Eval — Round 2 Crossover A"
108
+ - Agent 2: "BTO Eval — Round 2 Crossover B"
109
+ - Agent 3: "BTO Eval — Round 2 Crossover C"
110
+
111
+ Record scores. Select top 2.
112
+
113
+ ### Step 8: Crossover — Generate Round 3 Variants
114
+
115
+ Repeat crossover from Step 6 with the new top 2 from Round 2.
116
+ Generate 3 final crossover variants.
117
+
118
+ ### Step 9: Final Evaluation — Round 3 (Full Panel)
119
+
120
+ **Round 3 uses Layer 2 — full 3-judge panel per variant.**
121
+
122
+ For each of the 3 final variants, spawn 3 parallel agents (model: sonnet):
123
+
124
+ Total: up to 9 agents running in parallel (or in 3 sequential batches of 3).
125
+
126
+ Select the single best variant as winner.
127
+
128
+ If `max_score(all_variants_round3) < BASELINE_SCORE + 0.3`:
129
+ - Report: "Optimization produced minimal improvement. Consider original."
130
+ - Still present the best variant for review.
131
+
132
+ **Abort conditions (from `optimize.md`):**
133
+ - Any round shows overall regression > 0.5 from baseline → stop immediately
134
+ - Artifact semantics change fundamentally → flag and stop
135
+ - User requests stop
136
+
137
+ ### Step 10: Apply Winner
138
+
139
+ 1. Write the winning variant to the original artifact path (overwrite in place)
140
+ 2. Preserve original as `<filename>.pre-optimize.bak` in the same directory
141
+
142
+ ### Step 11: Generate Report
143
+
144
+ Compute before/after delta per dimension.
145
+
146
+ **Winning strategy:** name the mutation strategy (or crossover combination) that produced the winner.
147
+
148
+ **Recommendation logic:**
149
+ - Improvement > 1.0 → "Apply changes — significant improvement"
150
+ - Improvement 0.5-1.0 → "Review changes before applying"
151
+ - Improvement < 0.5 → "Minimal improvement — original may be preferred"
152
+
153
+ ---
154
+
155
+ ## Checkpoint
156
+
157
+ ```
158
+ ═══════════════════════════════════════════════════════
159
+ CHECKPOINT: OPTIMIZE Complete
160
+ Artifact: [path]
161
+ Rounds run: [1-3]
162
+ Total evaluations: [N] (baseline + variants)
163
+
164
+ BEFORE → AFTER:
165
+ METHODOLOGY: X.X → X.X ([+/-]X.X)
166
+ DEPTH: X.X → X.X ([+/-]X.X)
167
+ CORRECTNESS: X.X → X.X ([+/-]X.X)
168
+ USABILITY: X.X → X.X ([+/-]X.X)
169
+ ROBUSTNESS: X.X → X.X ([+/-]X.X)
170
+
171
+ OVERALL: X.X → X.X ([+/-]X.X)
172
+
173
+ Winning Strategy: [strategy name or crossover combination]
174
+
175
+ CHANGELOG:
176
+ - [specific change made]
177
+ - [specific change made]
178
+ - [specific change made]
179
+
180
+ Recommendation: [Apply / Review / Original preferred]
181
+
182
+ Backup saved: [path.pre-optimize.bak]
183
+
184
+ • "ок" — done, keep optimized version
185
+ • "ещё раунд" — run one additional optimization round
186
+ • "откат" — restore original from backup
187
+ • "покажи diff" — show before/after diff for review
188
+ ═══════════════════════════════════════════════════════
189
+ ```
190
+
191
+ Wait for user confirmation.
192
+
193
+ ---
194
+
195
+ ## Modular Usage
196
+
197
+ This command is also invoked internally by `/bto` as the final step (OPTIMIZE phase) of the full pipeline, receiving `TEST_SCORE` and the artifact path from the preceding TEST step.
198
+
199
+ ## Critical Rules
200
+
201
+ - Always establish baseline with Layer 2 before generating variants
202
+ - Never skip baseline — variants cannot be ranked without a reference point
203
+ - Agent tool is REQUIRED for parallel variant evaluation in all rounds
204
+ - Hard cap: 3 rounds by default, 5 maximum — never loop indefinitely
205
+ - Do not optimize if baseline ≥ 8.0 without explicit user confirmation
206
+ - Always save a `.pre-optimize.bak` backup before overwriting original
207
+ - Round 3 evaluation must use Layer 2 (sonnet judges), not Layer 1 (haiku)
208
+ - Preserve original artifact intent — optimization changes HOW, not WHAT
@@ -0,0 +1,186 @@
1
+ # /bto-test — Multi-Agent Evaluation of a Skill or Command
2
+
3
+ ## Usage
4
+ ```
5
+ /bto-test [path to skill directory or file]
6
+ ```
7
+
8
+ ## Parameters
9
+ - $ARGUMENTS — Path to a skill directory (`.claude/skills/<name>/`), a command file (`.claude/commands/<name>.md`), a rule file (`.claude/rules/<name>.md`), or any research artifact. Optionally include "full" to force Layer 2 panel regardless of Layer 1 score.
10
+
11
+ ## Protocol
12
+
13
+ ### Step 1: Load Skill and Module
14
+
15
+ Read `.claude/skills/bto/SKILL.md`
16
+ Read `.claude/skills/bto/modules/test.md`
17
+ Read `.claude/skills/bto/references/judge-rubrics.md`
18
+
19
+ ### Step 2: Validate Input
20
+
21
+ If $ARGUMENTS is empty:
22
+ - Ask: "Provide a path to the skill directory or artifact file you want to evaluate."
23
+ - Stop and wait.
24
+
25
+ Resolve the artifact path from $ARGUMENTS:
26
+ - Strip any trailing slash
27
+ - If the path points to a directory → look for `SKILL.md` inside it as the primary file, but include all files in the directory for context
28
+ - If the path points to a file → use that file directly
29
+ - If the path does not exist → report "Path not found: [path]" and stop
30
+
31
+ ### Step 3: Detect Artifact Type
32
+
33
+ Auto-detect from path pattern:
34
+
35
+ | Path Pattern | Detected Type |
36
+ |-------------|--------------|
37
+ | `.claude/skills/*/SKILL.md` or `.claude/skills/*/` | skill |
38
+ | `.claude/commands/*.md` | command |
39
+ | `.claude/rules/*.md` | rule |
40
+ | `.claude/agents/*.md` | agent |
41
+ | `researches/**/*.md` | research artifact |
42
+
43
+ If type cannot be determined from path → infer from file content structure.
44
+
45
+ ### Step 4: Layer 0 — Deterministic Pre-checks
46
+
47
+ **Run Layer 0 first. Always. No exceptions. Zero cost.**
48
+
49
+ Execute all applicable checks from `test.md` for the detected artifact type:
50
+
51
+ - Universal checks (U1-U5): file exists, UTF-8, has headings, no excessive blanks, size in bounds
52
+ - Type-specific checks: skill (S1-S10), command (C1-C5), rule (R1-R4), agent (A1-A4)
53
+
54
+ Display Layer 0 output as a per-check PASS/FAIL table.
55
+
56
+ **Gate:** If pass rate < 80% → stop here. Do NOT proceed to Layer 1 or Layer 2.
57
+ Report all failures with specific line references where possible.
58
+
59
+ ### Step 5: Layer 1 — Single LLM Judge (Quick)
60
+
61
+ Run if and only if Layer 0 gate passes.
62
+
63
+ **Model:** haiku
64
+
65
+ Execute the Layer 1 judge prompt from `test.md` against the artifact content.
66
+
67
+ Rate on 5 dimensions (1-10 each):
68
+ 1. CLARITY
69
+ 2. COMPLETENESS
70
+ 3. ACTIONABILITY
71
+ 4. QUALITY
72
+ 5. ANTI-PATTERNS
73
+
74
+ Display scores, one-line justification per dimension, top 3 improvement suggestions, and VERDICT.
75
+
76
+ **Gate:**
77
+ - Average ≥ 7.0 → PASS
78
+ - Average 5.0-6.9 → NEEDS WORK (offer Layer 2)
79
+ - Average < 5.0 → FAIL (Layer 2 recommended)
80
+
81
+ Check $ARGUMENTS for "full":
82
+ - If "full" present → proceed directly to Layer 2 without asking
83
+
84
+ Otherwise, if Layer 1 score < 7.0 → automatically offer Layer 2.
85
+ If Layer 1 score ≥ 7.0 → ask: "Run full 3-judge panel for comprehensive evaluation? (yes/no)"
86
+
87
+ ### Step 6: Layer 2 — Full Judge Panel (3 Parallel Agents)
88
+
89
+ Run if triggered by score threshold or user request.
90
+
91
+ **Spawn 3 agents in parallel using Agent tool:**
92
+
93
+ - Agent 1: "BTO Judge — Domain Expert" (model: sonnet)
94
+ - Focus: methodology, depth, technical correctness, domain fit
95
+ - Uses Judge 1 prompt from `test.md`
96
+
97
+ - Agent 2: "BTO Judge — Critic" (model: sonnet)
98
+ - Focus: gaps, weaknesses, anti-patterns, failure scenarios
99
+ - Uses Judge 2 prompt from `test.md`
100
+
101
+ - Agent 3: "BTO Judge — Completeness Auditor" (model: sonnet)
102
+ - Focus: structural coverage, cross-references, section completeness
103
+ - Uses Judge 3 prompt from `test.md`
104
+
105
+ **Isolation:** Each agent evaluates independently. No cross-communication.
106
+
107
+ **Aggregate scores after all 3 return:**
108
+
109
+ Weights: Expert 0.40, Critic 0.30, Auditor 0.30
110
+
111
+ Per-dimension weighted score:
112
+ ```
113
+ dim_score = expert[dim] * 0.4 + critic[dim] * 0.3 + auditor[dim] * 0.3
114
+ ```
115
+
116
+ Overall score = mean of all dimension scores.
117
+
118
+ **Disagreement detection:** For any dimension where `max(scores) - min(scores) > 3` → flag for meta-judge.
119
+
120
+ **Meta-judge (if triggered):**
121
+ - Model: default (opus)
122
+ - Present all 3 evaluations for the flagged dimension
123
+ - Reconcile to a single score with explicit reasoning
124
+ - Flag for human review if still unresolvable
125
+
126
+ ### Step 7: Record TEST_SCORE
127
+
128
+ Set `TEST_SCORE` = overall weighted average (Layer 2 if run, else Layer 1 average).
129
+
130
+ This value is consumed by `/bto-optimize` and by the full `/bto` pipeline.
131
+
132
+ ---
133
+
134
+ ## Checkpoint
135
+
136
+ ```
137
+ ═══════════════════════════════════════════════════════
138
+ CHECKPOINT: TEST Complete
139
+ Artifact: [path]
140
+ Type: [detected type]
141
+
142
+ Layer 0: X/Y checks passed (Z%) — PASS / FAIL
143
+ Layer 1: X.X/10 — PASS / NEEDS WORK / FAIL
144
+ Layer 2: X.X/10 (weighted) — [if run, else "not run"]
145
+
146
+ OVERALL SCORE: X.X/10
147
+
148
+ Per-Dimension (Layer 2 or Layer 1):
149
+ METHODOLOGY: X.X
150
+ DEPTH: X.X
151
+ CORRECTNESS: X.X
152
+ USABILITY: X.X
153
+ ROBUSTNESS: X.X
154
+
155
+ [Flagged dimensions with judge disagreement, if any]
156
+
157
+ Top Improvements:
158
+ 1. [specific suggestion]
159
+ 2. [specific suggestion]
160
+ 3. [specific suggestion]
161
+
162
+ • "ок" — done
163
+ • "покажи детали [expert/critic/auditor]" — expand judge feedback
164
+ • "оптимизируй" — run /bto-optimize with this artifact
165
+ • "запусти полный" — re-run with Layer 2 if only Layer 1 was run
166
+ ═══════════════════════════════════════════════════════
167
+ ```
168
+
169
+ Wait for user confirmation.
170
+
171
+ ---
172
+
173
+ ## Modular Usage
174
+
175
+ This command is also invoked internally by:
176
+ - `/bto` — as Step 4 (TEST phase) of the full pipeline
177
+ - `/bto-optimize` — as baseline evaluation before optimization rounds
178
+
179
+ ## Critical Rules
180
+
181
+ - Always run Layer 0 before any LLM evaluation — it is free and fast
182
+ - Never spawn Layer 2 agents if Layer 0 fails (< 80% pass rate)
183
+ - Agent tool is REQUIRED for Layer 2 — do not run judges sequentially
184
+ - Report `TEST_SCORE` explicitly so downstream commands can consume it
185
+ - Rubrics from `judge-rubrics.md` govern scoring anchor points for all judges
186
+ - If "full" is in $ARGUMENTS, skip the confirmation prompt and run Layer 2 automatically
@@ -0,0 +1,171 @@
1
+ # /bto — Full BTO Pipeline (Build · Test · Optimize)
2
+
3
+ ## Usage
4
+ ```
5
+ /bto [path to existing skill OR description of new skill]
6
+ ```
7
+
8
+ ## Parameters
9
+ - $ARGUMENTS — Path to an existing skill/command directory or file, OR a natural language description of a new skill to create
10
+
11
+ ## Protocol
12
+
13
+ ### Step 1: Setup — Load Skill
14
+
15
+ Read `.claude/skills/bto/SKILL.md`
16
+
17
+ ### Step 2: Route by Input Type
18
+
19
+ Inspect $ARGUMENTS to determine mode:
20
+
21
+ **If $ARGUMENTS is a path that exists on disk:**
22
+ - Mode: TEST → OPTIMIZE
23
+ - Skip BUILD
24
+ - Proceed to Step 4 (TEST)
25
+
26
+ **If $ARGUMENTS is a natural language description (not a path):**
27
+ - Mode: BUILD → TEST → OPTIMIZE
28
+ - Proceed to Step 3 (BUILD)
29
+
30
+ **If $ARGUMENTS is empty:**
31
+ - Ask the user: "Provide a path to an existing skill, or describe a new skill to build."
32
+ - Stop and wait.
33
+
34
+ ---
35
+
36
+ ### Step 3: BUILD (only if description provided)
37
+
38
+ Read `.claude/skills/bto/modules/build.md`
39
+
40
+ Execute the BUILD module:
41
+ 1. Auto-detect artifact type from description
42
+ 2. QUICK mode by default — parse requirements directly from description
43
+ 3. If user said "deep" anywhere in $ARGUMENTS — activate DEEP mode (load `explore` skill first)
44
+ 4. Generate complete artifact following build templates
45
+ 5. Run self-review (Layer 0 structural check)
46
+ 6. Create all files in the target directory
47
+ 7. Record `BUILD_OUTPUT_PATH` for use in next steps
48
+
49
+ **Checkpoint BUILD:**
50
+ ```
51
+ ═══════════════════════════════════════════════════════
52
+ CHECKPOINT 1: BUILD Complete
53
+ Artifact generated and self-reviewed.
54
+ Files created: [list generated files]
55
+ Path: [BUILD_OUTPUT_PATH]
56
+ • "ок" — run TEST on the new artifact
57
+ • "переделай [aspect]" — rebuild with changes
58
+ • "углуби [section]" — expand a section
59
+ ═══════════════════════════════════════════════════════
60
+ ```
61
+ Wait for user confirmation before proceeding.
62
+
63
+ ---
64
+
65
+ ### Step 4: TEST
66
+
67
+ Read `.claude/skills/bto/modules/test.md`
68
+ Read `.claude/skills/bto/references/judge-rubrics.md`
69
+
70
+ **Resolve artifact path:**
71
+ - If BUILD was run → use `BUILD_OUTPUT_PATH`
72
+ - If path was provided in $ARGUMENTS → use that path
73
+
74
+ Execute TEST module in layers:
75
+
76
+ **Layer 0 — Deterministic pre-checks (always run first):**
77
+ - Run all structural checks for detected artifact type
78
+ - Display Layer 0 output with per-check PASS/FAIL
79
+ - If score < 80% → stop, report failures, do not proceed to LLM layers
80
+
81
+ **Layer 1 — Single LLM judge (haiku):**
82
+ - Run if Layer 0 passes
83
+ - Quick spot-check on all 5 dimensions
84
+ - Display scores and top 3 improvement suggestions
85
+
86
+ **Layer 2 — Full judge panel (user-triggered or if Layer 1 score < 7.0):**
87
+ - Ask: "Run full 3-judge panel for comprehensive evaluation? (yes/no)"
88
+ - If yes → spawn 3 parallel agents using Agent tool:
89
+ - Agent 1: "BTO Judge — Domain Expert" (model: sonnet)
90
+ - Agent 2: "BTO Judge — Critic" (model: sonnet)
91
+ - Agent 3: "BTO Judge — Completeness Auditor" (model: sonnet)
92
+ - Aggregate scores with weights: Expert 0.40, Critic 0.30, Auditor 0.30
93
+ - Flag any dimension where max - min > 3 (trigger meta-judge if needed)
94
+ - Display full evaluation report
95
+
96
+ Record `TEST_SCORE` (overall weighted average) for use in OPTIMIZE step.
97
+
98
+ **Checkpoint TEST:**
99
+ ```
100
+ ═══════════════════════════════════════════════════════
101
+ CHECKPOINT 2: TEST Complete
102
+ Artifact: [path]
103
+ Layer 0: X/Y checks passed
104
+ Layer 1: X.X/10 — [PASS / NEEDS WORK / FAIL]
105
+ Layer 2: X.X/10 (weighted) — [if run]
106
+
107
+ Overall: X.X/10
108
+ • "ок" — run OPTIMIZE
109
+ • "пропусти" — skip optimize (artifact is good)
110
+ • "покажи детали [judge]" — expand judge feedback
111
+ ═══════════════════════════════════════════════════════
112
+ ```
113
+ Wait for user confirmation before proceeding.
114
+
115
+ ---
116
+
117
+ ### Step 5: OPTIMIZE
118
+
119
+ Read `.claude/skills/bto/modules/optimize.md`
120
+
121
+ Execute OPTIMIZE module:
122
+ 1. If `TEST_SCORE` ≥ 8.0 → report "Artifact already high quality (X.X/10). Minor tweaks only." Show specific suggestions. Stop.
123
+ 2. If `TEST_SCORE` < 8.0 → run full evolutionary optimization:
124
+ - Round 1: Generate 5 variants (one per mutation strategy) → evaluate with Layer 1 (haiku) → select top 2
125
+ - Round 2: Crossover top 2 → generate 3 variants → evaluate with Layer 1 → select top 2
126
+ - Round 3: Crossover → generate 3 variants → evaluate with Layer 2 (3 parallel sonnet judges) → select winner
127
+ 3. Apply the winning variant to the artifact file
128
+ 4. Display before/after delta report
129
+
130
+ **Checkpoint OPTIMIZE:**
131
+ ```
132
+ ═══════════════════════════════════════════════════════
133
+ CHECKPOINT 3: OPTIMIZE Complete
134
+ Artifact: [path]
135
+ Rounds run: 3
136
+ Total evaluations: 15
137
+
138
+ BEFORE → AFTER:
139
+ METHODOLOGY: X.X → X.X (+X.X)
140
+ DEPTH: X.X → X.X (+X.X)
141
+ CORRECTNESS: X.X → X.X (+X.X)
142
+ USABILITY: X.X → X.X (+X.X)
143
+ ROBUSTNESS: X.X → X.X (+X.X)
144
+
145
+ OVERALL: X.X → X.X (+X.X)
146
+
147
+ Winning Strategy: [strategy name]
148
+ Recommendation: [Apply / Review / Original preferred]
149
+ • "ок" — done
150
+ • "ещё раунд" — run another optimization round
151
+ • "откат" — restore original artifact
152
+ ═══════════════════════════════════════════════════════
153
+ ```
154
+
155
+ ---
156
+
157
+ ## Modular Usage
158
+
159
+ Each module can be run independently:
160
+ - `/bto-build [description]` — BUILD only
161
+ - `/bto-test [path]` — TEST only
162
+ - `/bto-optimize [path]` — OPTIMIZE only
163
+
164
+ ## Critical Rules
165
+
166
+ - Always run Layer 0 before any LLM evaluation — it is free and fast
167
+ - Never skip checkpoints — wait for user "ок" between modules
168
+ - Only run OPTIMIZE if TEST score < 8.0
169
+ - Agent tool is REQUIRED for Layer 2 parallel judge panel
170
+ - BUILD mode: QUICK by default, DEEP only if user explicitly requests it
171
+ - Record artifact path from BUILD and pass it through TEST → OPTIMIZE
@@ -0,0 +1,91 @@
1
+ # BTO Quality Gate Rules
2
+
3
+ ## What is BTO
4
+ BTO (Build-Test-Optimize) is a pipeline for generating, evaluating, and iteratively
5
+ improving skills, prompts, or any structured artifact via agent-driven evaluation loops.
6
+ These rules apply to ANY evaluation system, not just Keysarium.
7
+
8
+ ## Layer Architecture and Model Budget
9
+
10
+ | Layer | Role | Model | Trigger |
11
+ |-------|------|-------|---------|
12
+ | Layer 0 | Structural pre-check (format, completeness) | haiku | Always |
13
+ | Layer 1 | Shallow semantic check (relevance, coherence) | haiku | After Layer 0 passes |
14
+ | Layer 2 | Deep evaluation (quality, domain fit) | sonnet (judge panel) | After Layer 1 passes |
15
+ | Layer 3 | Creative synthesis / optimization crossover | opus | On top-N candidates only |
16
+
17
+ Never promote an artifact to a higher layer if the lower layer gate fails.
18
+
19
+ ## Cost Optimization Table
20
+
21
+ | Task | Model | Rationale |
22
+ |------|-------|-----------|
23
+ | Layer 0 structural checks | haiku | High-frequency, pattern-matching only |
24
+ | Layer 1 semantic baseline | haiku | Fast coherence scan, no deep reasoning needed |
25
+ | Judge 1 — Domain Expert | sonnet | Domain knowledge + nuanced scoring |
26
+ | Judge 2 — Critic | sonnet | Adversarial analysis, pattern detection |
27
+ | Judge 3 — Completeness Auditor | sonnet | Structured coverage check |
28
+ | Meta-judge (escalation) | sonnet | Disagreement resolution |
29
+ | Crossover / creative synthesis | opus | Novel combination of best candidates |
30
+ | Mutation workers (standard) | sonnet | Requires reasoning about improvement direction |
31
+ | Variant fast-eval (ranking pass) | haiku | Volume scoring before full panel |
32
+
33
+ Escalate to a higher-cost model only when the lower-cost model has failed or is insufficient.
34
+
35
+ ## Layer 0 Mandatory Checks
36
+ Every generated skill or artifact MUST pass ALL of these before entering judge panel:
37
+
38
+ - [ ] Required sections present (structure check)
39
+ - [ ] No empty placeholders (`[TODO]`, `[TBD]`, `<INSERT>`)
40
+ - [ ] Length within bounds (not below minimum, not above maximum)
41
+ - [ ] Encoding valid (no broken unicode, no binary artifacts)
42
+ - [ ] Self-reference loop absent (artifact does not cite itself as source)
43
+
44
+ If any check fails → reject immediately, log reason, do NOT send to judges.
45
+ Layer 0 may auto-retry up to 3 times before escalating to human review.
46
+
47
+ ## Judge Panel Rules
48
+
49
+ - Panel MUST have an odd number of judges: 3 (standard) or 5 (high-stakes)
50
+ - Each judge operates in strict isolation: reads the same artifact, writes to a separate evaluation file
51
+ - Judges do NOT see each other's scores before submitting
52
+ - Final score = weighted average (weights defined per panel configuration)
53
+ - Standard weights: Domain Expert 0.4 / Critic 0.3 / Completeness Auditor 0.3
54
+ - Disagreement threshold: if max_score - min_score > 3 points → escalate to meta-judge
55
+
56
+ ## Optimization Delta Gate
57
+
58
+ - An optimization iteration is accepted ONLY if: `new_score - prev_score > 0.5`
59
+ - If delta <= 0.5 for 3 consecutive iterations → declare convergence and stop
60
+ - If score DECREASES by > 1.0 → rollback to previous best and log regression
61
+ - Improvement must be measurable on the same rubric used in the previous iteration
62
+
63
+ ## Human Checkpoint Rules
64
+
65
+ - NEVER auto-approve any artifact for delivery without a human checkpoint
66
+ - Checkpoint is required after: Layer 2 evaluation, final optimization round, before packaging
67
+ - Checkpoint format follows the standard checkpoint-protocol.md
68
+ - Exception: Layer 0 rejections may be auto-retried up to 3 times before human escalation
69
+
70
+ ## BTO-Specific Anti-Patterns
71
+
72
+ | Anti-Pattern | Detection Signal | Required Fix |
73
+ |-------------|-----------------|--------------|
74
+ | Score inflation | All judges score > 8.5 on first attempt | Add calibration prompt to critics |
75
+ | Overfitting to rubric | Artifact optimizes wording to match rubric literally | Blind evaluation: hide rubric from generator |
76
+ | Conformity collapse | Judges converge to identical scores after 1 round | Enforce isolation, re-randomize judge order |
77
+ | Runaway optimization | > 10 iterations without convergence | Abort, log, human review |
78
+ | Phantom improvement | Delta > 0.5 but no substantive content change | Diff-check content, not just score |
79
+ | Judge-generator collusion | Same model used for both generation and evaluation | BLOCK — generator and judge models must differ |
80
+ | Missing rejection log | Failed artifacts silently discarded | Every rejection MUST be logged with reason |
81
+
82
+ ## Auto-Detection
83
+ Self-check generated artifacts and evaluation results against the anti-patterns above.
84
+ If detected, flag with a WARNING label and halt the BTO loop pending human review.
85
+
86
+ ## Reusability Note
87
+ These rules are artifact-type agnostic. Apply them to:
88
+ - Skill generation pipelines
89
+ - Prompt optimization loops
90
+ - Presentation scoring systems
91
+ - Any multi-judge evaluation workflow