@dzhechkov/skills-bto 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,201 @@
1
+ # OPTIMIZE Module — Evolutionary Prompt Optimization Protocol
2
+
3
+ ## Purpose
4
+
5
+ Improve Claude Code artifacts through evolutionary prompt optimization: generate variants, evaluate, select, mutate, repeat.
6
+
7
+ ## Input
8
+
9
+ - **Path:** Path to artifact to optimize
10
+ - **Rounds:** Number of optimization rounds (default: 3, max: 5)
11
+ - **Budget:** max evaluations (default: 15)
12
+ - **Focus:** Optional dimension to prioritize (METHODOLOGY / DEPTH / CORRECTNESS / USABILITY / ROBUSTNESS)
13
+
14
+ ## Prerequisites
15
+
16
+ - Artifact must pass Layer 0 checks (run TEST first)
17
+ - Baseline score established (Layer 2 evaluation)
18
+ - Only optimize if baseline < 8.0 (otherwise artifact is already good)
19
+
20
+ ---
21
+
22
+ ## Protocol
23
+
24
+ ### Step 1: Baseline Evaluation
25
+
26
+ 1. Run TEST module with level=layer2 on current artifact
27
+ 2. Record:
28
+ - Per-dimension scores
29
+ - Overall score
30
+ - Key weaknesses identified by judges
31
+ 3. If overall ≥ 8.0 → report "Artifact already high quality" and suggest minor tweaks only
32
+ 4. Identify **target dimensions**: dimensions scoring < 7.0
33
+
34
+ ### Step 2: Variant Generation (Round 1)
35
+
36
+ Generate N=5 variants of the artifact. Each variant applies ONE mutation strategy.
37
+
38
+ **Mutation Assignment:**
39
+ - Variant 1: **Rephrase** — reword unclear instructions for precision
40
+ - Variant 2: **Restructure** — reorganize sections for better flow
41
+ - Variant 3: **Add Constraints** — add guardrails, edge cases, boundary conditions
42
+ - Variant 4: **Simplify** — remove redundancy, tighten language, reduce verbosity
43
+ - Variant 5: **Specialize** — add domain-specific context and examples
44
+
45
+ **Strategy-to-Weakness Mapping:**
46
+
47
+ | Weak Dimension | Primary Strategy | Secondary Strategy |
48
+ |---------------|-----------------|-------------------|
49
+ | METHODOLOGY | Restructure | Add Constraints |
50
+ | DEPTH | Specialize | Add Constraints |
51
+ | CORRECTNESS | Add Constraints | Rephrase |
52
+ | USABILITY | Rephrase | Restructure |
53
+ | ROBUSTNESS | Add Constraints | Specialize |
54
+
55
+ If a focus dimension is specified, generate 3 variants targeting that dimension's strategies and 2 variants with other strategies.
56
+
57
+ ### Variant Generation Prompt
58
+
59
+ ```
60
+ You are optimizing a Claude Code {artifact_type}.
61
+
62
+ ## Current Artifact
63
+ {content}
64
+
65
+ ## Baseline Evaluation
66
+ Overall: {score}/10
67
+ Weaknesses: {weaknesses}
68
+ Target dimensions: {target_dimensions}
69
+
70
+ ## Mutation Strategy: {strategy_name}
71
+ {strategy_description}
72
+
73
+ ## Task
74
+ Apply the {strategy_name} mutation to improve this artifact.
75
+ Focus on addressing these specific weaknesses: {target_weaknesses}
76
+
77
+ Rules:
78
+ - Preserve the original intent and scope
79
+ - Maintain all existing sections
80
+ - Do not change the artifact type or structure fundamentally
81
+ - Changes should be targeted and purposeful
82
+ - Output the COMPLETE modified artifact (not a diff)
83
+ ```
84
+
85
+ ### Step 3: Evaluate Variants (Round 1)
86
+
87
+ Run TEST module with level=layer1 (haiku — fast and cheap) on each variant.
88
+
89
+ **Parallel execution:** Spawn 5 agents (model: haiku), one per variant.
90
+
91
+ Record scores for all 5 variants.
92
+
93
+ ### Step 4: Selection + Crossover
94
+
95
+ 1. **Select:** Top 2 variants by overall score
96
+ 2. **Crossover:** Combine best elements:
97
+ - Take sections where Variant A scored higher from A
98
+ - Take sections where Variant B scored higher from B
99
+ - Generate 3 new variants from the crossover
100
+
101
+ **Crossover Prompt:**
102
+ ```
103
+ You are creating an improved Claude Code artifact by combining the best
104
+ elements of two high-scoring variants.
105
+
106
+ ## Variant A (Score: {score_a})
107
+ {variant_a}
108
+ Strengths: {strengths_a}
109
+
110
+ ## Variant B (Score: {score_b})
111
+ {variant_b}
112
+ Strengths: {strengths_b}
113
+
114
+ ## Task
115
+ Create a new variant that combines the strengths of both:
116
+ - From A, take: {specific_sections_a}
117
+ - From B, take: {specific_sections_b}
118
+ - Ensure coherence and consistency
119
+ - Output the COMPLETE artifact
120
+ ```
121
+
122
+ ### Step 5: Evaluate + Select (Rounds 2-3)
123
+
124
+ **Round 2:**
125
+ - Evaluate 3 crossover variants with Layer 1
126
+ - Select top 2
127
+ - Generate 3 new crossover variants
128
+
129
+ **Round 3 (Final):**
130
+ - Evaluate 3 final variants with **Layer 2** (full judge panel — thorough)
131
+ - Select the single best variant
132
+
133
+ ### Step 6: Output
134
+
135
+ 1. **Best variant** — the optimized artifact
136
+ 2. **Before/After comparison:**
137
+ ```
138
+ ═══════════════════════════════════════════════════════
139
+ 🔧 BTO OPTIMIZATION REPORT
140
+ Artifact: <path>
141
+ Rounds: 3
142
+ Total evaluations: 15
143
+
144
+ BEFORE → AFTER:
145
+ METHODOLOGY: 6.2 → 8.1 (+1.9) ⬆️
146
+ DEPTH: 5.8 → 7.5 (+1.7) ⬆️
147
+ CORRECTNESS: 7.0 → 8.3 (+1.3) ⬆️
148
+ USABILITY: 6.5 → 8.0 (+1.5) ⬆️
149
+ ROBUSTNESS: 5.5 → 7.8 (+2.3) ⬆️
150
+
151
+ OVERALL: 6.2 → 7.9 (+1.7) ⬆️
152
+
153
+ Winning Strategy: Restructure + Add Constraints (crossover)
154
+
155
+ CHANGELOG:
156
+ - Restructured protocol into clearer numbered steps
157
+ - Added edge case handling for empty inputs
158
+ - Expanded anti-patterns with 3 new entries
159
+ - Simplified module loading instructions
160
+ - Added concrete examples to each section
161
+ ═══════════════════════════════════════════════════════
162
+ ```
163
+
164
+ 3. **Recommendation:**
165
+ - If improvement > 1.0: "Apply changes"
166
+ - If improvement 0.5-1.0: "Review changes, consider applying"
167
+ - If improvement < 0.5: "Minimal improvement — original may be preferred"
168
+
169
+ ---
170
+
171
+ ## Cost Summary
172
+
173
+ | Operation | Count | Model | Est. Tokens |
174
+ |-----------|-------|-------|-------------|
175
+ | Baseline eval | 1 | sonnet ×3 | ~15K |
176
+ | Variant generation | 5 | opus | ~25K |
177
+ | Round 1 eval | 5 | haiku | ~10K |
178
+ | Crossover generation | 3 | opus | ~15K |
179
+ | Round 2 eval | 3 | haiku | ~6K |
180
+ | Crossover generation | 3 | opus | ~15K |
181
+ | Round 3 eval | 3 | sonnet ×3 | ~45K |
182
+ | **Total** | | | **~131K tokens** |
183
+
184
+ ## Anti-Patterns
185
+
186
+ | Anti-Pattern | Detection | Fix |
187
+ |-------------|-----------|-----|
188
+ | Overfitting to one metric | One dimension +3, others flat or down | Balance mutations across dimensions |
189
+ | Losing generality | Specialized variant breaks other use cases | Test with multiple example inputs |
190
+ | Infinite loop | > 5 rounds without improvement | Hard cap at configured rounds |
191
+ | Semantic drift | Optimized version changes intent | Compare purpose statement before/after |
192
+ | Premature optimization | Baseline ≥ 8.0 | Skip optimization, suggest minor tweaks |
193
+ | Score inflation | Haiku gives high scores to everything | Calibrate with sonnet baseline |
194
+
195
+ ## Abort Conditions
196
+
197
+ Stop optimization immediately if:
198
+ 1. Any round shows overall regression > 0.5 from baseline
199
+ 2. Critical structural checks (Layer 0) fail on any variant
200
+ 3. Artifact semantics change fundamentally
201
+ 4. User requests stop
@@ -0,0 +1,348 @@
1
+ # TEST Module — Multi-Agent Evaluation Protocol
2
+
3
+ ## Purpose
4
+
5
+ Evaluate any Claude Code artifact (skill, command, rule, agent template) using layered evaluation: deterministic pre-checks, single-judge quick eval, and full 3-judge panel.
6
+
7
+ ## Input
8
+
9
+ - **Path:** Path to artifact file or directory
10
+ - **Level:** layer0 | layer1 | layer2 | full (default: full)
11
+ - **Artifact type:** auto-detected from path/content
12
+
13
+ ## Type Detection
14
+
15
+ | Path Pattern | Detected Type |
16
+ |-------------|--------------|
17
+ | `.claude/skills/*/SKILL.md` | skill |
18
+ | `.claude/skills/*/` (directory) | skill |
19
+ | `.claude/commands/*.md` | command |
20
+ | `.claude/rules/*.md` | rule |
21
+ | `.claude/agents/*.md` | agent |
22
+ | `researches/**/*.md` | research artifact |
23
+
24
+ ---
25
+
26
+ ## Layer 0: Deterministic Pre-checks
27
+
28
+ **Cost:** Zero (no LLM calls)
29
+ **Speed:** Instant
30
+ **Purpose:** Catch structural issues before expensive LLM evaluation
31
+
32
+ ### Universal Checks (all types)
33
+
34
+ ```
35
+ CHECK-U1: File exists and is non-empty
36
+ CHECK-U2: File is valid UTF-8 text
37
+ CHECK-U3: Has at least one markdown heading (#)
38
+ CHECK-U4: No consecutive empty lines (> 2)
39
+ CHECK-U5: File size within bounds (see per-type limits)
40
+ ```
41
+
42
+ ### Skill Checks
43
+
44
+ ```
45
+ CHECK-S1: SKILL.md exists in skill directory
46
+ CHECK-S2: Has "# Title" as first heading
47
+ CHECK-S3: Has "## Overview" or "## Purpose" section
48
+ CHECK-S4: Has "## Anti-Patterns" section
49
+ CHECK-S5: All files in modules/ referenced in SKILL.md
50
+ CHECK-S6: All files in references/ referenced in SKILL.md
51
+ CHECK-S7: No empty sections (heading → next heading with no content)
52
+ CHECK-S8: Size: 1KB < SKILL.md < 50KB
53
+ CHECK-S9: Total directory size < 200KB
54
+ CHECK-S10: At least one file in references/ OR examples/
55
+ ```
56
+
57
+ ### Command Checks
58
+
59
+ ```
60
+ CHECK-C1: Contains "$ARGUMENTS" or parameter reference
61
+ CHECK-C2: Has checkpoint banner or protocol
62
+ CHECK-C3: Has skill loading instruction (Read *.SKILL.md)
63
+ CHECK-C4: Size: 500B < file < 20KB
64
+ CHECK-C5: Has "## Usage" or "## Protocol" section
65
+ ```
66
+
67
+ ### Rule Checks
68
+
69
+ ```
70
+ CHECK-R1: Has table or structured list of patterns
71
+ CHECK-R2: Each pattern has detection signal and fix
72
+ CHECK-R3: Size: 200B < file < 10KB
73
+ CHECK-R4: Has "Auto-Detection" or similar section
74
+ ```
75
+
76
+ ### Agent Template Checks
77
+
78
+ ```
79
+ CHECK-A1: Specifies model (haiku/sonnet/opus)
80
+ CHECK-A2: Specifies isolation scope
81
+ CHECK-A3: Has prompt template or instructions
82
+ CHECK-A4: Size: 200B < file < 10KB
83
+ ```
84
+
85
+ ### Layer 0 Scoring
86
+
87
+ - Each check: PASS (1) or FAIL (0)
88
+ - Score = passed / total
89
+ - **Gate: score ≥ 0.80 to proceed to Layer 1+**
90
+ - If score < 0.80: return report with specific failures, skip LLM evaluation
91
+
92
+ ### Layer 0 Output
93
+
94
+ ```
95
+ ═══════════════════════════════════════════════════════
96
+ 📋 LAYER 0: Deterministic Pre-checks
97
+ Artifact: <path>
98
+ Type: <detected type>
99
+
100
+ Results: X/Y passed (Z%)
101
+
102
+ ✅ CHECK-S1: SKILL.md exists
103
+ ✅ CHECK-S2: Has title heading
104
+ ❌ CHECK-S7: Empty section found at line 45
105
+ ...
106
+
107
+ Gate: PASS ✅ / FAIL ❌
108
+ ═══════════════════════════════════════════════════════
109
+ ```
110
+
111
+ ---
112
+
113
+ ## Layer 1: Single LLM Judge (Quick)
114
+
115
+ **Cost:** Low (1 haiku call)
116
+ **Speed:** ~10 seconds
117
+ **Purpose:** Fast quality signal for iteration or batch evaluation
118
+
119
+ ### Model Selection
120
+
121
+ - Default: **haiku** (cost-optimized)
122
+ - For critical artifacts: **sonnet**
123
+
124
+ ### Judge Prompt
125
+
126
+ ```
127
+ You are evaluating a Claude Code {artifact_type}.
128
+
129
+ ## Artifact Content
130
+ {content}
131
+
132
+ ## Evaluation Dimensions
133
+
134
+ Rate each dimension 1-10:
135
+
136
+ 1. CLARITY (Are instructions unambiguous? Can an LLM follow them precisely?)
137
+ 2. COMPLETENESS (Are all necessary sections present? No missing pieces?)
138
+ 3. ACTIONABILITY (Can Claude produce concrete output from these instructions?)
139
+ 4. QUALITY (Well-structured? Professional? Good formatting?)
140
+ 5. ANTI-PATTERNS (Avoids known pitfalls? Has failure mode coverage?)
141
+
142
+ ## Required Output Format
143
+
144
+ SCORES:
145
+ - CLARITY: X/10 — [one-line justification]
146
+ - COMPLETENESS: X/10 — [one-line justification]
147
+ - ACTIONABILITY: X/10 — [one-line justification]
148
+ - QUALITY: X/10 — [one-line justification]
149
+ - ANTI-PATTERNS: X/10 — [one-line justification]
150
+
151
+ AVERAGE: X.X/10
152
+
153
+ TOP 3 IMPROVEMENTS:
154
+ 1. [specific, actionable suggestion]
155
+ 2. [specific, actionable suggestion]
156
+ 3. [specific, actionable suggestion]
157
+
158
+ VERDICT: PASS (≥7.0) / NEEDS WORK (5.0-6.9) / FAIL (<5.0)
159
+ ```
160
+
161
+ ### Layer 1 Gate
162
+
163
+ - Average ≥ 7.0: PASS
164
+ - Average 5.0-6.9: NEEDS WORK (can proceed to Layer 2 for detailed feedback)
165
+ - Average < 5.0: FAIL (fix before proceeding)
166
+
167
+ ---
168
+
169
+ ## Layer 2: Full Judge Panel (3 Agents)
170
+
171
+ **Cost:** Moderate (3 sonnet calls)
172
+ **Speed:** ~30 seconds (parallel)
173
+ **Purpose:** Comprehensive, multi-perspective evaluation
174
+
175
+ ### Architecture
176
+
177
+ Spawn 3 agents in parallel using Agent tool:
178
+
179
+ ```
180
+ Agent 1: "BTO Judge — Domain Expert" model: sonnet
181
+ Agent 2: "BTO Judge — Critic" model: sonnet
182
+ Agent 3: "BTO Judge — Completeness" model: sonnet
183
+ ```
184
+
185
+ **Isolation:** Each agent reads the same artifact independently. No cross-communication.
186
+
187
+ ### Judge 1: Domain Expert
188
+
189
+ Focus: Is the content technically sound and appropriate for the domain?
190
+
191
+ ```
192
+ You are a Domain Expert evaluator for Claude Code artifacts.
193
+
194
+ Evaluate this {artifact_type} for technical quality:
195
+
196
+ {content}
197
+
198
+ Score 1-10 on each dimension:
199
+ 1. METHODOLOGY — Is the approach well-designed? Sound structure?
200
+ 2. DEPTH — Sufficient detail for the task? Not too shallow?
201
+ 3. CORRECTNESS — Are all claims and instructions valid?
202
+ 4. USABILITY — Can a user/agent effectively use this?
203
+ 5. ROBUSTNESS — Handles edge cases and failure modes?
204
+
205
+ For each dimension provide:
206
+ - Score (1-10)
207
+ - 2-3 sentence justification
208
+ - One specific improvement suggestion
209
+
210
+ OVERALL_SCORE: weighted average
211
+ CONFIDENCE: HIGH/MEDIUM/LOW
212
+ KEY_STRENGTHS: [top 2]
213
+ KEY_WEAKNESSES: [top 2]
214
+ ```
215
+
216
+ ### Judge 2: Critic
217
+
218
+ Focus: Find weaknesses, gaps, and anti-patterns.
219
+
220
+ ```
221
+ You are a Critical Evaluator. Your job is to find problems.
222
+
223
+ Evaluate this {artifact_type} adversarially:
224
+
225
+ {content}
226
+
227
+ Score 1-10 on each dimension (be strict — average should be ~5-6):
228
+ 1. METHODOLOGY — Any logical flaws or unjustified assumptions?
229
+ 2. DEPTH — What's missing? What's under-specified?
230
+ 3. CORRECTNESS — Any instructions that could mislead or produce wrong output?
231
+ 4. USABILITY — What would confuse a user? Where would someone get stuck?
232
+ 5. ROBUSTNESS — What breaks this? What edge cases are unhandled?
233
+
234
+ For each dimension provide:
235
+ - Score (1-10) — err on the side of strict
236
+ - 2-3 sentence justification focusing on PROBLEMS
237
+ - One specific failure scenario
238
+
239
+ OVERALL_SCORE: weighted average
240
+ CRITICAL_ISSUES: [list of blocking problems]
241
+ IMPROVEMENT_PRIORITY: [ordered list of what to fix first]
242
+ ```
243
+
244
+ ### Judge 3: Completeness Auditor
245
+
246
+ Focus: Structural completeness and cross-reference integrity.
247
+
248
+ ```
249
+ You are a Completeness Auditor for Claude Code artifacts.
250
+
251
+ Audit this {artifact_type} for structural completeness:
252
+
253
+ {content}
254
+
255
+ Score 1-10 on each dimension:
256
+ 1. METHODOLOGY — Does the structure follow established patterns?
257
+ 2. DEPTH — Is every section adequately populated?
258
+ 3. CORRECTNESS — Do all cross-references resolve? Are all claims supported?
259
+ 4. USABILITY — Is the information well-organized and findable?
260
+ 5. ROBUSTNESS — Are failure modes documented? Are anti-patterns listed?
261
+
262
+ For each dimension provide:
263
+ - Score (1-10)
264
+ - 2-3 sentence justification
265
+ - Missing element (if any)
266
+
267
+ OVERALL_SCORE: weighted average
268
+ MISSING_SECTIONS: [list]
269
+ BROKEN_REFERENCES: [list]
270
+ COVERAGE_GAPS: [list]
271
+ ```
272
+
273
+ ### Score Aggregation
274
+
275
+ After all 3 judges return:
276
+
277
+ **Weights:**
278
+ - Domain Expert: 0.40
279
+ - Critic: 0.30
280
+ - Completeness Auditor: 0.30
281
+
282
+ **Per-dimension:**
283
+ ```
284
+ dimension_score = expert[dim] * 0.4 + critic[dim] * 0.3 + auditor[dim] * 0.3
285
+ ```
286
+
287
+ **Overall:**
288
+ ```
289
+ overall = mean(all dimension_scores)
290
+ ```
291
+
292
+ **Disagreement Detection:**
293
+ For each dimension, if `max(scores) - min(scores) > 3` → FLAG for meta-judge.
294
+
295
+ ---
296
+
297
+ ## Meta-Judge: Disagreement Resolution
298
+
299
+ **Trigger:** Any flagged dimension from Layer 2.
300
+
301
+ **Model:** default (opus)
302
+
303
+ **Prompt:**
304
+ ```
305
+ Three judges evaluated a Claude Code artifact and disagreed significantly
306
+ on the following dimension(s): {flagged_dimensions}
307
+
308
+ Judge 1 (Domain Expert): {score} — {justification}
309
+ Judge 2 (Critic): {score} — {justification}
310
+ Judge 3 (Completeness): {score} — {justification}
311
+
312
+ Reconcile these scores. Provide:
313
+ 1. Your reconciled score (1-10) with reasoning
314
+ 2. Which judge(s) had the most valid perspective and why
315
+ 3. Whether this needs human review (YES/NO)
316
+ ```
317
+
318
+ ---
319
+
320
+ ## Output: Evaluation Report
321
+
322
+ Format defined in `examples/sample-eval-report.md`.
323
+
324
+ Summary structure:
325
+ ```
326
+ ═══════════════════════════════════════════════════════
327
+ 📊 BTO EVALUATION REPORT
328
+ Artifact: <path>
329
+ Type: <type>
330
+ Level: Layer 0 + Layer 1 + Layer 2
331
+
332
+ OVERALL SCORE: X.X / 10 [PASS / NEEDS WORK / FAIL]
333
+
334
+ Per-Dimension:
335
+ METHODOLOGY: X.X ██████████░░
336
+ DEPTH: X.X ████████░░░░
337
+ CORRECTNESS: X.X █████████░░░
338
+ USABILITY: X.X ██████████░░
339
+ ROBUSTNESS: X.X ███████░░░░░
340
+
341
+ Flagged: [dimensions with >3 disagreement]
342
+
343
+ Top Improvements:
344
+ 1. ...
345
+ 2. ...
346
+ 3. ...
347
+ ═══════════════════════════════════════════════════════
348
+ ```