@dzhechkov/skills-bto 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bin/cli.js +5 -0
- package/package.json +43 -0
- package/src/cli.js +150 -0
- package/src/commands/doctor.js +366 -0
- package/src/commands/init.js +188 -0
- package/src/commands/list.js +161 -0
- package/src/commands/remove.js +211 -0
- package/src/commands/update.js +198 -0
- package/src/utils.js +398 -0
- package/templates/.claude/agents/bto-judge-panel.md +192 -0
- package/templates/.claude/agents/bto-optimizer-worker.md +181 -0
- package/templates/.claude/commands/bto-build.md +169 -0
- package/templates/.claude/commands/bto-optimize.md +208 -0
- package/templates/.claude/commands/bto-test.md +186 -0
- package/templates/.claude/commands/bto.md +171 -0
- package/templates/.claude/rules/bto-quality-gates.md +91 -0
- package/templates/.claude/skills/bto/SKILL.md +266 -0
- package/templates/.claude/skills/bto/examples/sample-eval-report.md +436 -0
- package/templates/.claude/skills/bto/modules/build.md +189 -0
- package/templates/.claude/skills/bto/modules/optimize.md +201 -0
- package/templates/.claude/skills/bto/modules/test.md +348 -0
- package/templates/.claude/skills/bto/references/eval-patterns.md +202 -0
- package/templates/.claude/skills/bto/references/judge-rubrics.md +183 -0
- package/templates/.claude/skills/bto/references/optimization-methods.md +139 -0
- package/templates/.claude/skills/bto/references/quality-checklist.md +220 -0
|
@@ -0,0 +1,266 @@
|
|
|
1
|
+
# BTO — Build · Test · Optimize
|
|
2
|
+
|
|
3
|
+
> Multi-agent evaluation and prompt optimization system for Claude Code artifacts.
|
|
4
|
+
|
|
5
|
+
## Overview
|
|
6
|
+
|
|
7
|
+
BTO is a 3-module pipeline for creating, evaluating, and optimizing Claude Code skills, commands, rules, and agents. It combines structured generation, multi-agent LLM judging, and evolutionary prompt optimization.
|
|
8
|
+
|
|
9
|
+
## Modules
|
|
10
|
+
|
|
11
|
+
| Module | Purpose | Command | Agent Pattern |
|
|
12
|
+
|--------|---------|---------|---------------|
|
|
13
|
+
| **BUILD** | Generate skill/command from requirements | `/bto-build` | Single agent + explore |
|
|
14
|
+
| **TEST** | Multi-agent evaluation with judge panel | `/bto-test` | 3 parallel judges |
|
|
15
|
+
| **OPTIMIZE** | Evolutionary prompt optimization | `/bto-optimize` | N parallel workers |
|
|
16
|
+
|
|
17
|
+
## Quick Start
|
|
18
|
+
|
|
19
|
+
```
|
|
20
|
+
/bto [skill-path] — Full pipeline (build → test → optimize)
|
|
21
|
+
/bto-build [description] — Generate a new skill or command
|
|
22
|
+
/bto-test [path] — Evaluate an existing artifact
|
|
23
|
+
/bto-optimize [path] — Optimize prompts in a skill
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
## Module 1: BUILD
|
|
29
|
+
|
|
30
|
+
**Goal:** Generate complete, production-quality Claude Code skills/commands from natural language requirements.
|
|
31
|
+
|
|
32
|
+
### Protocol
|
|
33
|
+
|
|
34
|
+
1. **Requirements Clarification** — Load `explore` skill to clarify:
|
|
35
|
+
- What artifact type? (skill / command / rule / agent template)
|
|
36
|
+
- Domain and context
|
|
37
|
+
- Input/output expectations
|
|
38
|
+
- Quality criteria
|
|
39
|
+
- Reference examples (optional)
|
|
40
|
+
|
|
41
|
+
2. **Template Selection** — Based on artifact type:
|
|
42
|
+
- **Skill:** SKILL.md + modules/ + references/ + examples/
|
|
43
|
+
- **Command:** Single .md file with parameter handling
|
|
44
|
+
- **Rule:** Single .md with detection patterns and fixes
|
|
45
|
+
- **Agent template:** Single .md with agent config
|
|
46
|
+
|
|
47
|
+
3. **Generation** — Structured generation following conventions:
|
|
48
|
+
- Load `references/quality-checklist.md` for validation criteria
|
|
49
|
+
- Generate with proper sections, headers, formatting
|
|
50
|
+
- Include anti-patterns section
|
|
51
|
+
- Include at least one reference and one example
|
|
52
|
+
|
|
53
|
+
4. **Self-Review** — Before output:
|
|
54
|
+
- Check against quality checklist (Layer 0)
|
|
55
|
+
- Verify no empty sections
|
|
56
|
+
- Verify all cross-references resolve
|
|
57
|
+
- Verify naming conventions match (`file-conventions` rule)
|
|
58
|
+
|
|
59
|
+
### BUILD Modes
|
|
60
|
+
|
|
61
|
+
| Mode | When | Agents | Time |
|
|
62
|
+
|------|------|--------|------|
|
|
63
|
+
| QUICK | Simple artifact, clear requirements | 1 | ~2 min |
|
|
64
|
+
| DEEP | Complex skill, unclear requirements | 1 + explore | ~5 min |
|
|
65
|
+
|
|
66
|
+
### Output Structure (Skill)
|
|
67
|
+
|
|
68
|
+
```
|
|
69
|
+
.claude/skills/<name>/
|
|
70
|
+
├── SKILL.md ← Main orchestrator
|
|
71
|
+
├── modules/ ← Detailed protocols per module
|
|
72
|
+
│ └── <module-name>.md
|
|
73
|
+
├── references/ ← Supporting materials
|
|
74
|
+
│ └── <reference-name>.md
|
|
75
|
+
└── examples/ ← Few-shot examples
|
|
76
|
+
└── <example-name>.md
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
### Anti-Patterns (BUILD)
|
|
80
|
+
|
|
81
|
+
| Anti-Pattern | Fix |
|
|
82
|
+
|-------------|-----|
|
|
83
|
+
| Generic skill without domain context | Add specific domain constraints and examples |
|
|
84
|
+
| Missing references directory | Always include at least one reference |
|
|
85
|
+
| No examples | Add at least one few-shot example |
|
|
86
|
+
| Over-scoped skill | Split into modules, keep SKILL.md as orchestrator |
|
|
87
|
+
| Copy-pasting from other skills | Adapt, don't copy — each skill is unique |
|
|
88
|
+
|
|
89
|
+
---
|
|
90
|
+
|
|
91
|
+
## Module 2: TEST
|
|
92
|
+
|
|
93
|
+
**Goal:** Evaluate any Claude Code artifact using deterministic checks + multi-agent LLM judging.
|
|
94
|
+
|
|
95
|
+
### Evaluation Layers
|
|
96
|
+
|
|
97
|
+
```
|
|
98
|
+
Layer 0: Deterministic Pre-checks ← Fast, free, catches 60% of issues
|
|
99
|
+
Layer 1: Single LLM Judge ← Quick spot-check (model: haiku)
|
|
100
|
+
Layer 2: Full Judge Panel (3 agents) ← Comprehensive evaluation (model: sonnet)
|
|
101
|
+
Meta-Judge: Disagreement Resolution ← Only if judges disagree >3 points
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
### Layer 0: Deterministic Pre-checks
|
|
105
|
+
|
|
106
|
+
Run BEFORE spawning any LLM agents. Fast and free.
|
|
107
|
+
|
|
108
|
+
**For Skills:**
|
|
109
|
+
- [ ] SKILL.md exists and is non-empty
|
|
110
|
+
- [ ] Has `# Title` as first heading
|
|
111
|
+
- [ ] Has `## Overview` section
|
|
112
|
+
- [ ] Has `## Anti-Patterns` section
|
|
113
|
+
- [ ] All referenced modules exist in modules/
|
|
114
|
+
- [ ] All referenced references exist in references/
|
|
115
|
+
- [ ] No empty sections (heading followed immediately by another heading)
|
|
116
|
+
- [ ] File size: 1KB < size < 50KB per file
|
|
117
|
+
- [ ] Total skill directory: < 200KB
|
|
118
|
+
|
|
119
|
+
**For Commands:**
|
|
120
|
+
- [ ] Has `$ARGUMENTS` parameter reference
|
|
121
|
+
- [ ] Has checkpoint protocol
|
|
122
|
+
- [ ] Has skill loading instruction (Read .claude/skills/...)
|
|
123
|
+
- [ ] File size: 500B < size < 20KB
|
|
124
|
+
|
|
125
|
+
**For Rules:**
|
|
126
|
+
- [ ] Has table or list of patterns
|
|
127
|
+
- [ ] Has detection signals and fixes
|
|
128
|
+
- [ ] File size: 200B < size < 10KB
|
|
129
|
+
|
|
130
|
+
**Scoring:** Pass/Fail per check → aggregate pass rate (must be ≥80% to proceed)
|
|
131
|
+
|
|
132
|
+
### Layer 1: Single LLM Judge (Quick Mode)
|
|
133
|
+
|
|
134
|
+
Use when: fast feedback needed, or evaluating many artifacts in batch.
|
|
135
|
+
|
|
136
|
+
**Model:** haiku (cost-optimized)
|
|
137
|
+
|
|
138
|
+
**Prompt Pattern:**
|
|
139
|
+
```
|
|
140
|
+
You are a Claude Code artifact evaluator. Rate the following {artifact_type}
|
|
141
|
+
on these dimensions (1-10 each):
|
|
142
|
+
|
|
143
|
+
1. CLARITY — Are instructions unambiguous?
|
|
144
|
+
2. COMPLETENESS — Are all necessary sections present?
|
|
145
|
+
3. ACTIONABILITY — Can Claude follow these instructions to produce output?
|
|
146
|
+
4. QUALITY — Is the content well-structured and professional?
|
|
147
|
+
5. ANTI-PATTERNS — Does it avoid known anti-patterns?
|
|
148
|
+
|
|
149
|
+
Provide: score per dimension, brief justification, top 3 improvement suggestions.
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
**Threshold:** Average ≥ 7.0 to pass
|
|
153
|
+
|
|
154
|
+
### Layer 2: Full Judge Panel (3 Agents)
|
|
155
|
+
|
|
156
|
+
Use when: thorough evaluation needed, or artifact is critical.
|
|
157
|
+
|
|
158
|
+
**Architecture:** 3 parallel agents, each with a different perspective.
|
|
159
|
+
|
|
160
|
+
Load rubrics from `references/judge-rubrics.md`.
|
|
161
|
+
|
|
162
|
+
| Judge | Role | Focus | Model |
|
|
163
|
+
|-------|------|-------|-------|
|
|
164
|
+
| Agent 1 | Domain Expert | Accuracy, depth, methodology, domain fit | sonnet |
|
|
165
|
+
| Agent 2 | Critic | Gaps, weaknesses, anti-patterns, edge cases | sonnet |
|
|
166
|
+
| Agent 3 | Completeness Auditor | Structure, coverage, cross-references, actionability | sonnet |
|
|
167
|
+
|
|
168
|
+
**Each judge scores independently on 5 dimensions (1-10):**
|
|
169
|
+
1. METHODOLOGY — Is the approach sound and well-structured?
|
|
170
|
+
2. DEPTH — Is the content thorough enough for the task?
|
|
171
|
+
3. CORRECTNESS — Are claims accurate and instructions valid?
|
|
172
|
+
4. USABILITY — Can a user/agent effectively use this artifact?
|
|
173
|
+
5. ROBUSTNESS — Does it handle edge cases and failure modes?
|
|
174
|
+
|
|
175
|
+
**Aggregation:**
|
|
176
|
+
- Weighted average: Expert (0.4) + Critic (0.3) + Auditor (0.3)
|
|
177
|
+
- Per-dimension and overall score
|
|
178
|
+
- Flag if any dimension has range > 3 across judges
|
|
179
|
+
|
|
180
|
+
### Meta-Judge: Disagreement Resolution
|
|
181
|
+
|
|
182
|
+
**Trigger:** Any dimension where max - min > 3 across judges.
|
|
183
|
+
|
|
184
|
+
**Action:**
|
|
185
|
+
1. Present all three evaluations to a single opus-level agent
|
|
186
|
+
2. Ask for reconciled score with explicit reasoning
|
|
187
|
+
3. Flag for human review if still unresolvable
|
|
188
|
+
|
|
189
|
+
### Output: Evaluation Report
|
|
190
|
+
|
|
191
|
+
See `examples/sample-eval-report.md` for format.
|
|
192
|
+
|
|
193
|
+
---
|
|
194
|
+
|
|
195
|
+
## Module 3: OPTIMIZE
|
|
196
|
+
|
|
197
|
+
**Goal:** Improve Claude Code artifacts through evolutionary prompt optimization.
|
|
198
|
+
|
|
199
|
+
### Protocol
|
|
200
|
+
|
|
201
|
+
1. **Baseline** — Evaluate current artifact with TEST (Layer 2)
|
|
202
|
+
2. **Generate Variants** — Create N=5 variants using mutation strategies
|
|
203
|
+
3. **Evaluate** — Run TEST (Layer 1) on each variant
|
|
204
|
+
4. **Select + Crossover** — Top 2 → generate 3 new variants
|
|
205
|
+
5. **Repeat** — 3 rounds total (15 evaluations)
|
|
206
|
+
6. **Output** — Best variant + before/after delta + changelog
|
|
207
|
+
|
|
208
|
+
### Mutation Strategies
|
|
209
|
+
|
|
210
|
+
| Strategy | Description | When to Use |
|
|
211
|
+
|----------|-------------|-------------|
|
|
212
|
+
| **Rephrase** | Reword instructions for clarity | Low CLARITY score |
|
|
213
|
+
| **Restructure** | Reorganize sections and flow | Low USABILITY score |
|
|
214
|
+
| **Add Constraints** | Add guardrails and edge cases | Low ROBUSTNESS score |
|
|
215
|
+
| **Simplify** | Remove redundancy, tighten language | Over-engineered artifact |
|
|
216
|
+
| **Specialize** | Add domain-specific context | Low DEPTH score |
|
|
217
|
+
|
|
218
|
+
### Optimization Loop
|
|
219
|
+
|
|
220
|
+
```
|
|
221
|
+
Round 1: 5 variants × Layer 1 eval → Select top 2
|
|
222
|
+
Round 2: top 2 crossover → 3 new variants × Layer 1 eval → Select top 2
|
|
223
|
+
Round 3: top 2 crossover → 3 new variants × Layer 2 eval → Final selection
|
|
224
|
+
```
|
|
225
|
+
|
|
226
|
+
### Cost Optimization
|
|
227
|
+
|
|
228
|
+
| Operation | Model | Est. Cost |
|
|
229
|
+
|-----------|-------|-----------|
|
|
230
|
+
| Variant generation | default (opus) | Creative work |
|
|
231
|
+
| Layer 1 evaluation | haiku | Cheap, fast |
|
|
232
|
+
| Layer 2 evaluation | sonnet | Moderate, thorough |
|
|
233
|
+
| Meta-judge | default (opus) | Only on disagreement |
|
|
234
|
+
|
|
235
|
+
### Anti-Patterns (OPTIMIZE)
|
|
236
|
+
|
|
237
|
+
| Anti-Pattern | Fix |
|
|
238
|
+
|-------------|-----|
|
|
239
|
+
| Overfitting to one metric | Balance all 5 dimensions equally |
|
|
240
|
+
| Losing generality | Check original use cases still work |
|
|
241
|
+
| Infinite optimization loop | Hard cap at 3 rounds |
|
|
242
|
+
| Optimizing already-good artifacts | Only optimize if baseline < 8.0 |
|
|
243
|
+
| Changing artifact semantics | Preserve original intent and scope |
|
|
244
|
+
|
|
245
|
+
---
|
|
246
|
+
|
|
247
|
+
## Integration with Keysarium Pipeline
|
|
248
|
+
|
|
249
|
+
BTO can be used at any point in the pipeline:
|
|
250
|
+
|
|
251
|
+
| Phase | BTO Usage |
|
|
252
|
+
|-------|-----------|
|
|
253
|
+
| After Phase 0 | `/bto-test researches/<slug>/00_product_discovery.md` |
|
|
254
|
+
| After Phase 2 | `/bto-test researches/<slug>/02_research_findings.md` |
|
|
255
|
+
| After Phase 5 | `/bto-test researches/<slug>/05_presentation_content.md` |
|
|
256
|
+
| New skill creation | `/bto-build "describe your skill"` |
|
|
257
|
+
| Skill improvement | `/bto-optimize .claude/skills/<name>/SKILL.md` |
|
|
258
|
+
|
|
259
|
+
## Dependencies
|
|
260
|
+
|
|
261
|
+
This skill uses:
|
|
262
|
+
- `explore` skill — for requirements clarification in BUILD module
|
|
263
|
+
- `references/judge-rubrics.md` — evaluation rubrics
|
|
264
|
+
- `references/eval-patterns.md` — multi-agent evaluation patterns
|
|
265
|
+
- `references/optimization-methods.md` — optimization approaches
|
|
266
|
+
- `references/quality-checklist.md` — deterministic pre-checks
|