@dzhechkov/skills-bto 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,266 @@
1
+ # BTO — Build · Test · Optimize
2
+
3
+ > Multi-agent evaluation and prompt optimization system for Claude Code artifacts.
4
+
5
+ ## Overview
6
+
7
+ BTO is a 3-module pipeline for creating, evaluating, and optimizing Claude Code skills, commands, rules, and agents. It combines structured generation, multi-agent LLM judging, and evolutionary prompt optimization.
8
+
9
+ ## Modules
10
+
11
+ | Module | Purpose | Command | Agent Pattern |
12
+ |--------|---------|---------|---------------|
13
+ | **BUILD** | Generate skill/command from requirements | `/bto-build` | Single agent + explore |
14
+ | **TEST** | Multi-agent evaluation with judge panel | `/bto-test` | 3 parallel judges |
15
+ | **OPTIMIZE** | Evolutionary prompt optimization | `/bto-optimize` | N parallel workers |
16
+
17
+ ## Quick Start
18
+
19
+ ```
20
+ /bto [skill-path] — Full pipeline (build → test → optimize)
21
+ /bto-build [description] — Generate a new skill or command
22
+ /bto-test [path] — Evaluate an existing artifact
23
+ /bto-optimize [path] — Optimize prompts in a skill
24
+ ```
25
+
26
+ ---
27
+
28
+ ## Module 1: BUILD
29
+
30
+ **Goal:** Generate complete, production-quality Claude Code skills/commands from natural language requirements.
31
+
32
+ ### Protocol
33
+
34
+ 1. **Requirements Clarification** — Load `explore` skill to clarify:
35
+ - What artifact type? (skill / command / rule / agent template)
36
+ - Domain and context
37
+ - Input/output expectations
38
+ - Quality criteria
39
+ - Reference examples (optional)
40
+
41
+ 2. **Template Selection** — Based on artifact type:
42
+ - **Skill:** SKILL.md + modules/ + references/ + examples/
43
+ - **Command:** Single .md file with parameter handling
44
+ - **Rule:** Single .md with detection patterns and fixes
45
+ - **Agent template:** Single .md with agent config
46
+
47
+ 3. **Generation** — Structured generation following conventions:
48
+ - Load `references/quality-checklist.md` for validation criteria
49
+ - Generate with proper sections, headers, formatting
50
+ - Include anti-patterns section
51
+ - Include at least one reference and one example
52
+
53
+ 4. **Self-Review** — Before output:
54
+ - Check against quality checklist (Layer 0)
55
+ - Verify no empty sections
56
+ - Verify all cross-references resolve
57
+ - Verify naming conventions match (`file-conventions` rule)
58
+
59
+ ### BUILD Modes
60
+
61
+ | Mode | When | Agents | Time |
62
+ |------|------|--------|------|
63
+ | QUICK | Simple artifact, clear requirements | 1 | ~2 min |
64
+ | DEEP | Complex skill, unclear requirements | 1 + explore | ~5 min |
65
+
66
+ ### Output Structure (Skill)
67
+
68
+ ```
69
+ .claude/skills/<name>/
70
+ ├── SKILL.md ← Main orchestrator
71
+ ├── modules/ ← Detailed protocols per module
72
+ │ └── <module-name>.md
73
+ ├── references/ ← Supporting materials
74
+ │ └── <reference-name>.md
75
+ └── examples/ ← Few-shot examples
76
+ └── <example-name>.md
77
+ ```
78
+
79
+ ### Anti-Patterns (BUILD)
80
+
81
+ | Anti-Pattern | Fix |
82
+ |-------------|-----|
83
+ | Generic skill without domain context | Add specific domain constraints and examples |
84
+ | Missing references directory | Always include at least one reference |
85
+ | No examples | Add at least one few-shot example |
86
+ | Over-scoped skill | Split into modules, keep SKILL.md as orchestrator |
87
+ | Copy-pasting from other skills | Adapt, don't copy — each skill is unique |
88
+
89
+ ---
90
+
91
+ ## Module 2: TEST
92
+
93
+ **Goal:** Evaluate any Claude Code artifact using deterministic checks + multi-agent LLM judging.
94
+
95
+ ### Evaluation Layers
96
+
97
+ ```
98
+ Layer 0: Deterministic Pre-checks ← Fast, free, catches 60% of issues
99
+ Layer 1: Single LLM Judge ← Quick spot-check (model: haiku)
100
+ Layer 2: Full Judge Panel (3 agents) ← Comprehensive evaluation (model: sonnet)
101
+ Meta-Judge: Disagreement Resolution ← Only if judges disagree >3 points
102
+ ```
103
+
104
+ ### Layer 0: Deterministic Pre-checks
105
+
106
+ Run BEFORE spawning any LLM agents. Fast and free.
107
+
108
+ **For Skills:**
109
+ - [ ] SKILL.md exists and is non-empty
110
+ - [ ] Has `# Title` as first heading
111
+ - [ ] Has `## Overview` section
112
+ - [ ] Has `## Anti-Patterns` section
113
+ - [ ] All referenced modules exist in modules/
114
+ - [ ] All referenced references exist in references/
115
+ - [ ] No empty sections (heading followed immediately by another heading)
116
+ - [ ] File size: 1KB < size < 50KB per file
117
+ - [ ] Total skill directory: < 200KB
118
+
119
+ **For Commands:**
120
+ - [ ] Has `$ARGUMENTS` parameter reference
121
+ - [ ] Has checkpoint protocol
122
+ - [ ] Has skill loading instruction (Read .claude/skills/...)
123
+ - [ ] File size: 500B < size < 20KB
124
+
125
+ **For Rules:**
126
+ - [ ] Has table or list of patterns
127
+ - [ ] Has detection signals and fixes
128
+ - [ ] File size: 200B < size < 10KB
129
+
130
+ **Scoring:** Pass/Fail per check → aggregate pass rate (must be ≥80% to proceed)
131
+
132
+ ### Layer 1: Single LLM Judge (Quick Mode)
133
+
134
+ Use when: fast feedback needed, or evaluating many artifacts in batch.
135
+
136
+ **Model:** haiku (cost-optimized)
137
+
138
+ **Prompt Pattern:**
139
+ ```
140
+ You are a Claude Code artifact evaluator. Rate the following {artifact_type}
141
+ on these dimensions (1-10 each):
142
+
143
+ 1. CLARITY — Are instructions unambiguous?
144
+ 2. COMPLETENESS — Are all necessary sections present?
145
+ 3. ACTIONABILITY — Can Claude follow these instructions to produce output?
146
+ 4. QUALITY — Is the content well-structured and professional?
147
+ 5. ANTI-PATTERNS — Does it avoid known anti-patterns?
148
+
149
+ Provide: score per dimension, brief justification, top 3 improvement suggestions.
150
+ ```
151
+
152
+ **Threshold:** Average ≥ 7.0 to pass
153
+
154
+ ### Layer 2: Full Judge Panel (3 Agents)
155
+
156
+ Use when: thorough evaluation needed, or artifact is critical.
157
+
158
+ **Architecture:** 3 parallel agents, each with a different perspective.
159
+
160
+ Load rubrics from `references/judge-rubrics.md`.
161
+
162
+ | Judge | Role | Focus | Model |
163
+ |-------|------|-------|-------|
164
+ | Agent 1 | Domain Expert | Accuracy, depth, methodology, domain fit | sonnet |
165
+ | Agent 2 | Critic | Gaps, weaknesses, anti-patterns, edge cases | sonnet |
166
+ | Agent 3 | Completeness Auditor | Structure, coverage, cross-references, actionability | sonnet |
167
+
168
+ **Each judge scores independently on 5 dimensions (1-10):**
169
+ 1. METHODOLOGY — Is the approach sound and well-structured?
170
+ 2. DEPTH — Is the content thorough enough for the task?
171
+ 3. CORRECTNESS — Are claims accurate and instructions valid?
172
+ 4. USABILITY — Can a user/agent effectively use this artifact?
173
+ 5. ROBUSTNESS — Does it handle edge cases and failure modes?
174
+
175
+ **Aggregation:**
176
+ - Weighted average: Expert (0.4) + Critic (0.3) + Auditor (0.3)
177
+ - Per-dimension and overall score
178
+ - Flag if any dimension has range > 3 across judges
179
+
180
+ ### Meta-Judge: Disagreement Resolution
181
+
182
+ **Trigger:** Any dimension where max - min > 3 across judges.
183
+
184
+ **Action:**
185
+ 1. Present all three evaluations to a single opus-level agent
186
+ 2. Ask for reconciled score with explicit reasoning
187
+ 3. Flag for human review if still unresolvable
188
+
189
+ ### Output: Evaluation Report
190
+
191
+ See `examples/sample-eval-report.md` for format.
192
+
193
+ ---
194
+
195
+ ## Module 3: OPTIMIZE
196
+
197
+ **Goal:** Improve Claude Code artifacts through evolutionary prompt optimization.
198
+
199
+ ### Protocol
200
+
201
+ 1. **Baseline** — Evaluate current artifact with TEST (Layer 2)
202
+ 2. **Generate Variants** — Create N=5 variants using mutation strategies
203
+ 3. **Evaluate** — Run TEST (Layer 1) on each variant
204
+ 4. **Select + Crossover** — Top 2 → generate 3 new variants
205
+ 5. **Repeat** — 3 rounds total (15 evaluations)
206
+ 6. **Output** — Best variant + before/after delta + changelog
207
+
208
+ ### Mutation Strategies
209
+
210
+ | Strategy | Description | When to Use |
211
+ |----------|-------------|-------------|
212
+ | **Rephrase** | Reword instructions for clarity | Low CLARITY score |
213
+ | **Restructure** | Reorganize sections and flow | Low USABILITY score |
214
+ | **Add Constraints** | Add guardrails and edge cases | Low ROBUSTNESS score |
215
+ | **Simplify** | Remove redundancy, tighten language | Over-engineered artifact |
216
+ | **Specialize** | Add domain-specific context | Low DEPTH score |
217
+
218
+ ### Optimization Loop
219
+
220
+ ```
221
+ Round 1: 5 variants × Layer 1 eval → Select top 2
222
+ Round 2: top 2 crossover → 3 new variants × Layer 1 eval → Select top 2
223
+ Round 3: top 2 crossover → 3 new variants × Layer 2 eval → Final selection
224
+ ```
225
+
226
+ ### Cost Optimization
227
+
228
+ | Operation | Model | Est. Cost |
229
+ |-----------|-------|-----------|
230
+ | Variant generation | default (opus) | Creative work |
231
+ | Layer 1 evaluation | haiku | Cheap, fast |
232
+ | Layer 2 evaluation | sonnet | Moderate, thorough |
233
+ | Meta-judge | default (opus) | Only on disagreement |
234
+
235
+ ### Anti-Patterns (OPTIMIZE)
236
+
237
+ | Anti-Pattern | Fix |
238
+ |-------------|-----|
239
+ | Overfitting to one metric | Balance all 5 dimensions equally |
240
+ | Losing generality | Check original use cases still work |
241
+ | Infinite optimization loop | Hard cap at 3 rounds |
242
+ | Optimizing already-good artifacts | Only optimize if baseline < 8.0 |
243
+ | Changing artifact semantics | Preserve original intent and scope |
244
+
245
+ ---
246
+
247
+ ## Integration with Keysarium Pipeline
248
+
249
+ BTO can be used at any point in the pipeline:
250
+
251
+ | Phase | BTO Usage |
252
+ |-------|-----------|
253
+ | After Phase 0 | `/bto-test researches/<slug>/00_product_discovery.md` |
254
+ | After Phase 2 | `/bto-test researches/<slug>/02_research_findings.md` |
255
+ | After Phase 5 | `/bto-test researches/<slug>/05_presentation_content.md` |
256
+ | New skill creation | `/bto-build "describe your skill"` |
257
+ | Skill improvement | `/bto-optimize .claude/skills/<name>/SKILL.md` |
258
+
259
+ ## Dependencies
260
+
261
+ This skill uses:
262
+ - `explore` skill — for requirements clarification in BUILD module
263
+ - `references/judge-rubrics.md` — evaluation rubrics
264
+ - `references/eval-patterns.md` — multi-agent evaluation patterns
265
+ - `references/optimization-methods.md` — optimization approaches
266
+ - `references/quality-checklist.md` — deterministic pre-checks