@dzhechkov/skills-bto 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bin/cli.js +5 -0
- package/package.json +43 -0
- package/src/cli.js +150 -0
- package/src/commands/doctor.js +366 -0
- package/src/commands/init.js +188 -0
- package/src/commands/list.js +161 -0
- package/src/commands/remove.js +211 -0
- package/src/commands/update.js +198 -0
- package/src/utils.js +398 -0
- package/templates/.claude/agents/bto-judge-panel.md +192 -0
- package/templates/.claude/agents/bto-optimizer-worker.md +181 -0
- package/templates/.claude/commands/bto-build.md +169 -0
- package/templates/.claude/commands/bto-optimize.md +208 -0
- package/templates/.claude/commands/bto-test.md +186 -0
- package/templates/.claude/commands/bto.md +171 -0
- package/templates/.claude/rules/bto-quality-gates.md +91 -0
- package/templates/.claude/skills/bto/SKILL.md +266 -0
- package/templates/.claude/skills/bto/examples/sample-eval-report.md +436 -0
- package/templates/.claude/skills/bto/modules/build.md +189 -0
- package/templates/.claude/skills/bto/modules/optimize.md +201 -0
- package/templates/.claude/skills/bto/modules/test.md +348 -0
- package/templates/.claude/skills/bto/references/eval-patterns.md +202 -0
- package/templates/.claude/skills/bto/references/judge-rubrics.md +183 -0
- package/templates/.claude/skills/bto/references/optimization-methods.md +139 -0
- package/templates/.claude/skills/bto/references/quality-checklist.md +220 -0
|
@@ -0,0 +1,202 @@
|
|
|
1
|
+
# Multi-Agent Evaluation Patterns
|
|
2
|
+
|
|
3
|
+
## Overview
|
|
4
|
+
|
|
5
|
+
Research-backed patterns for using multiple LLM agents to evaluate content. Each pattern trades off cost, reliability, and depth.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## Pattern 1: Judge Panel (Used by BTO)
|
|
10
|
+
|
|
11
|
+
**Architecture:** N independent judges evaluate the same artifact in parallel.
|
|
12
|
+
|
|
13
|
+
```
|
|
14
|
+
Artifact ──┬──→ Judge 1 (Expert) ──┐
|
|
15
|
+
├──→ Judge 2 (Critic) ──┼──→ Aggregator → Report
|
|
16
|
+
└──→ Judge 3 (Auditor) ──┘
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
**Strengths:**
|
|
20
|
+
- Multi-perspective coverage
|
|
21
|
+
- Parallel execution (fast)
|
|
22
|
+
- Disagreement detection reveals ambiguous quality
|
|
23
|
+
|
|
24
|
+
**Weaknesses:**
|
|
25
|
+
- Judges may converge to similar opinions (groupthink)
|
|
26
|
+
- More expensive than single judge
|
|
27
|
+
- Requires aggregation logic
|
|
28
|
+
|
|
29
|
+
**When to use:** Standard evaluation of well-defined artifacts. BTO default.
|
|
30
|
+
|
|
31
|
+
**Mitigation for groupthink:**
|
|
32
|
+
- Give each judge a distinct role and perspective
|
|
33
|
+
- Critic judge explicitly instructed to be strict (calibrated lower)
|
|
34
|
+
- Different temperature settings if available
|
|
35
|
+
|
|
36
|
+
---
|
|
37
|
+
|
|
38
|
+
## Pattern 2: Adversarial Red/Blue
|
|
39
|
+
|
|
40
|
+
**Architecture:** One agent creates, another attacks.
|
|
41
|
+
|
|
42
|
+
```
|
|
43
|
+
Artifact ──→ Blue Team (Defender) ──→ Assessment
|
|
44
|
+
──→ Red Team (Attacker) ──→ Vulnerabilities
|
|
45
|
+
──→ Reconciliation
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
**Strengths:**
|
|
49
|
+
- Excellent at finding weaknesses
|
|
50
|
+
- Simulates real-world usage patterns
|
|
51
|
+
- Defender forced to justify every choice
|
|
52
|
+
|
|
53
|
+
**Weaknesses:**
|
|
54
|
+
- Can be overly negative
|
|
55
|
+
- Expensive (multiple rounds)
|
|
56
|
+
- Not suited for overall quality scoring
|
|
57
|
+
|
|
58
|
+
**When to use:** Security-sensitive artifacts, high-stakes decisions, or when robustness is critical.
|
|
59
|
+
|
|
60
|
+
---
|
|
61
|
+
|
|
62
|
+
## Pattern 3: Consensus-Based (Multi-Round Debate)
|
|
63
|
+
|
|
64
|
+
**Architecture:** Agents debate until convergence.
|
|
65
|
+
|
|
66
|
+
```
|
|
67
|
+
Round 1: Each judge evaluates independently
|
|
68
|
+
Round 2: Judges see each other's evaluations, revise
|
|
69
|
+
Round 3: Final convergence (or flag for human)
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
**Strengths:**
|
|
73
|
+
- High-confidence final scores
|
|
74
|
+
- Reduces individual judge bias
|
|
75
|
+
- Self-correcting
|
|
76
|
+
|
|
77
|
+
**Weaknesses:**
|
|
78
|
+
- Very expensive (3x the calls)
|
|
79
|
+
- Risk of conformity pressure
|
|
80
|
+
- Slow (sequential rounds)
|
|
81
|
+
|
|
82
|
+
**When to use:** Critical decisions where confidence matters more than speed.
|
|
83
|
+
|
|
84
|
+
---
|
|
85
|
+
|
|
86
|
+
## Pattern 4: Constitutional Evaluation
|
|
87
|
+
|
|
88
|
+
**Architecture:** Evaluate against a set of explicit principles.
|
|
89
|
+
|
|
90
|
+
```
|
|
91
|
+
Principles ──→ Evaluator ──→ Per-principle score ──→ Report
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
**Principles example for Claude Code artifacts:**
|
|
95
|
+
1. Instructions must be unambiguous
|
|
96
|
+
2. Every section must serve a purpose
|
|
97
|
+
3. Anti-patterns must be documented
|
|
98
|
+
4. Examples must demonstrate the happy path
|
|
99
|
+
5. Failure modes must be handled
|
|
100
|
+
|
|
101
|
+
**Strengths:**
|
|
102
|
+
- Deterministic criteria
|
|
103
|
+
- Easy to explain scores
|
|
104
|
+
- Reproducible
|
|
105
|
+
|
|
106
|
+
**Weaknesses:**
|
|
107
|
+
- Misses emergent quality issues
|
|
108
|
+
- Principles may not cover everything
|
|
109
|
+
- Rigid
|
|
110
|
+
|
|
111
|
+
**When to use:** Compliance checking, standardized quality gates.
|
|
112
|
+
|
|
113
|
+
---
|
|
114
|
+
|
|
115
|
+
## Pattern 5: Hierarchical Evaluation
|
|
116
|
+
|
|
117
|
+
**Architecture:** Fast cheap filter → detailed expensive evaluation.
|
|
118
|
+
|
|
119
|
+
```
|
|
120
|
+
Layer 0: Deterministic ──pass──→ Layer 1: Haiku ──pass──→ Layer 2: Sonnet Panel
|
|
121
|
+
│ │
|
|
122
|
+
fail fail
|
|
123
|
+
↓ ↓
|
|
124
|
+
Quick Fix Detailed Feedback
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
**Strengths:**
|
|
128
|
+
- Cost-efficient (most artifacts filtered early)
|
|
129
|
+
- Progressive detail
|
|
130
|
+
- Fast feedback for obvious issues
|
|
131
|
+
|
|
132
|
+
**Weaknesses:**
|
|
133
|
+
- Layer 0 can miss semantic issues
|
|
134
|
+
- Layer 1 may have different calibration than Layer 2
|
|
135
|
+
|
|
136
|
+
**When to use:** Batch evaluation, CI/CD pipelines, optimization loops. BTO uses this pattern.
|
|
137
|
+
|
|
138
|
+
---
|
|
139
|
+
|
|
140
|
+
## Aggregation Strategies
|
|
141
|
+
|
|
142
|
+
### 1. Weighted Average (BTO Default)
|
|
143
|
+
```
|
|
144
|
+
score = Σ(weight_i × judge_i_score) / Σ(weight_i)
|
|
145
|
+
```
|
|
146
|
+
- Simple, interpretable
|
|
147
|
+
- Weights reflect judge importance
|
|
148
|
+
- BTO weights: Expert 0.4, Critic 0.3, Auditor 0.3
|
|
149
|
+
|
|
150
|
+
### 2. Majority Voting
|
|
151
|
+
```
|
|
152
|
+
verdict = mode(judge_verdicts)
|
|
153
|
+
```
|
|
154
|
+
- Good for pass/fail decisions
|
|
155
|
+
- Requires odd number of judges
|
|
156
|
+
- Ignores score magnitude
|
|
157
|
+
|
|
158
|
+
### 3. Bayesian Aggregation
|
|
159
|
+
```
|
|
160
|
+
P(quality | scores) ∝ Π P(score_i | quality) × P(quality)
|
|
161
|
+
```
|
|
162
|
+
- Accounts for judge reliability
|
|
163
|
+
- Requires calibration data
|
|
164
|
+
- More complex to implement
|
|
165
|
+
|
|
166
|
+
### 4. Min-Score (Conservative)
|
|
167
|
+
```
|
|
168
|
+
score = min(all_judge_scores)
|
|
169
|
+
```
|
|
170
|
+
- Most conservative
|
|
171
|
+
- Good for safety-critical artifacts
|
|
172
|
+
- May be overly pessimistic
|
|
173
|
+
|
|
174
|
+
### 5. Trimmed Mean
|
|
175
|
+
```
|
|
176
|
+
score = mean(scores after removing highest and lowest)
|
|
177
|
+
```
|
|
178
|
+
- Outlier-resistant
|
|
179
|
+
- Requires ≥5 judges
|
|
180
|
+
- Good for large panels
|
|
181
|
+
|
|
182
|
+
---
|
|
183
|
+
|
|
184
|
+
## Anti-Conformity Measures
|
|
185
|
+
|
|
186
|
+
Prevent judges from converging to meaningless agreement:
|
|
187
|
+
|
|
188
|
+
1. **Role Differentiation** — Each judge has unique focus and scoring calibration
|
|
189
|
+
2. **Critic Calibration** — Critic judge instructed to score ~5-6 average (strict)
|
|
190
|
+
3. **Independent Evaluation** — No judge sees other judges' scores
|
|
191
|
+
4. **Disagreement as Signal** — High disagreement flags important quality dimensions
|
|
192
|
+
5. **Diverse Prompts** — Each judge gets a differently-framed evaluation prompt
|
|
193
|
+
|
|
194
|
+
## Cost Optimization
|
|
195
|
+
|
|
196
|
+
| Approach | Judges | Model | Est. Cost | Use When |
|
|
197
|
+
|----------|--------|-------|-----------|----------|
|
|
198
|
+
| Layer 0 only | 0 | none | Free | CI/CD pre-check |
|
|
199
|
+
| Layer 1 | 1 | haiku | ~$0.001 | Quick iteration |
|
|
200
|
+
| Layer 2 | 3 | sonnet | ~$0.01 | Thorough evaluation |
|
|
201
|
+
| Full + Meta | 3+1 | sonnet+opus | ~$0.05 | Critical artifacts |
|
|
202
|
+
| Consensus | 3×3 | sonnet | ~$0.03 | High-stakes decisions |
|
|
@@ -0,0 +1,183 @@
|
|
|
1
|
+
# Judge Rubrics — Evaluation Criteria by Artifact Type
|
|
2
|
+
|
|
3
|
+
## Universal Dimensions (1-10 scale)
|
|
4
|
+
|
|
5
|
+
All artifact types are evaluated on these 5 dimensions:
|
|
6
|
+
|
|
7
|
+
| Dimension | What it Measures | 1-3 (Poor) | 4-6 (Adequate) | 7-8 (Good) | 9-10 (Excellent) |
|
|
8
|
+
|-----------|-----------------|------------|----------------|------------|-------------------|
|
|
9
|
+
| METHODOLOGY | Approach soundness | No clear structure | Basic structure present | Well-designed protocol | Rigorous, innovative approach |
|
|
10
|
+
| DEPTH | Content thoroughness | Superficial | Covers basics | Thorough treatment | Comprehensive with nuances |
|
|
11
|
+
| CORRECTNESS | Accuracy of claims/instructions | Major errors | Minor inaccuracies | Accurate | Precise and verifiable |
|
|
12
|
+
| USABILITY | Ease of use | Confusing | Usable with effort | Clear and navigable | Intuitive, self-documenting |
|
|
13
|
+
| ROBUSTNESS | Edge case handling | No coverage | Some mentioned | Well-covered | Exhaustive with fallbacks |
|
|
14
|
+
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
## Skill Rubric
|
|
18
|
+
|
|
19
|
+
### METHODOLOGY
|
|
20
|
+
- 9-10: Clear multi-step protocol, modular design, explicit decision points
|
|
21
|
+
- 7-8: Good protocol structure, some modularity
|
|
22
|
+
- 4-6: Basic steps listed, no clear decision flow
|
|
23
|
+
- 1-3: No protocol, just description of what skill does
|
|
24
|
+
|
|
25
|
+
### DEPTH
|
|
26
|
+
- 9-10: Detailed modules, rich references, multiple examples, anti-patterns
|
|
27
|
+
- 7-8: Good detail in SKILL.md, has references and examples
|
|
28
|
+
- 4-6: Basic overview, minimal references
|
|
29
|
+
- 1-3: Stub-level content, no supporting files
|
|
30
|
+
|
|
31
|
+
### CORRECTNESS
|
|
32
|
+
- 9-10: All instructions produce expected output, cross-references valid
|
|
33
|
+
- 7-8: Instructions work, minor reference issues
|
|
34
|
+
- 4-6: Some instructions ambiguous, broken references
|
|
35
|
+
- 1-3: Instructions would produce wrong output
|
|
36
|
+
|
|
37
|
+
### USABILITY
|
|
38
|
+
- 9-10: Quick start section, clear navigation, progressive disclosure
|
|
39
|
+
- 7-8: Well-organized, findable information
|
|
40
|
+
- 4-6: Readable but poorly organized
|
|
41
|
+
- 1-3: Wall of text, no structure
|
|
42
|
+
|
|
43
|
+
### ROBUSTNESS
|
|
44
|
+
- 9-10: Anti-patterns documented, failure modes handled, abort conditions defined
|
|
45
|
+
- 7-8: Anti-patterns section present, some edge cases
|
|
46
|
+
- 4-6: Mentions some risks
|
|
47
|
+
- 1-3: No failure mode coverage
|
|
48
|
+
|
|
49
|
+
---
|
|
50
|
+
|
|
51
|
+
## Command Rubric
|
|
52
|
+
|
|
53
|
+
### METHODOLOGY
|
|
54
|
+
- 9-10: Clear step-by-step flow, proper checkpoints, skill loading pattern
|
|
55
|
+
- 7-8: Good flow, checkpoint present
|
|
56
|
+
- 4-6: Basic steps, missing checkpoint
|
|
57
|
+
- 1-3: No clear execution flow
|
|
58
|
+
|
|
59
|
+
### DEPTH
|
|
60
|
+
- 9-10: Handles all parameter combinations, parallel agent usage
|
|
61
|
+
- 7-8: Good parameter handling, some agent usage
|
|
62
|
+
- 4-6: Basic parameter handling
|
|
63
|
+
- 1-3: No parameter handling
|
|
64
|
+
|
|
65
|
+
### CORRECTNESS
|
|
66
|
+
- 9-10: $ARGUMENTS properly used, skill paths correct, agent configs valid
|
|
67
|
+
- 7-8: Mostly correct references
|
|
68
|
+
- 4-6: Some incorrect references
|
|
69
|
+
- 1-3: Wrong skill paths or broken references
|
|
70
|
+
|
|
71
|
+
### USABILITY
|
|
72
|
+
- 9-10: Clear usage line, example invocations, helpful error messages
|
|
73
|
+
- 7-8: Usage documented, some examples
|
|
74
|
+
- 4-6: Basic usage, no examples
|
|
75
|
+
- 1-3: Unclear how to invoke
|
|
76
|
+
|
|
77
|
+
### ROBUSTNESS
|
|
78
|
+
- 9-10: Input validation, fallbacks, timeout handling
|
|
79
|
+
- 7-8: Some input validation
|
|
80
|
+
- 4-6: Minimal validation
|
|
81
|
+
- 1-3: No error handling
|
|
82
|
+
|
|
83
|
+
---
|
|
84
|
+
|
|
85
|
+
## Research Artifact Rubric
|
|
86
|
+
|
|
87
|
+
### METHODOLOGY
|
|
88
|
+
- 9-10: PARANOID mode, systematic research plan, structured sections
|
|
89
|
+
- 7-8: Good research structure, mostly verified
|
|
90
|
+
- 4-6: Basic structure, some unverified claims
|
|
91
|
+
- 1-3: Unstructured, many unverified claims
|
|
92
|
+
|
|
93
|
+
### DEPTH
|
|
94
|
+
- 9-10: Multiple sources per claim, competitive analysis, quantitative data
|
|
95
|
+
- 7-8: Good source coverage, some quantitative data
|
|
96
|
+
- 4-6: Few sources, mostly qualitative
|
|
97
|
+
- 1-3: No sources, opinion-based
|
|
98
|
+
|
|
99
|
+
### CORRECTNESS
|
|
100
|
+
- 9-10: All claims sourced, no [UNVERIFIED] tags, real products/companies
|
|
101
|
+
- 7-8: Mostly sourced, few [UNVERIFIED]
|
|
102
|
+
- 4-6: Some sourced, several [UNVERIFIED]
|
|
103
|
+
- 1-3: Unsourced, potentially hallucinated
|
|
104
|
+
|
|
105
|
+
### USABILITY
|
|
106
|
+
- 9-10: Executive summary, clear sections, actionable takeaways
|
|
107
|
+
- 7-8: Well-organized, some actionable content
|
|
108
|
+
- 4-6: Readable but hard to act on
|
|
109
|
+
- 1-3: Dense, unstructured
|
|
110
|
+
|
|
111
|
+
### ROBUSTNESS
|
|
112
|
+
- 9-10: Limitations documented, alternative views considered, confidence levels
|
|
113
|
+
- 7-8: Some limitations noted
|
|
114
|
+
- 4-6: Minimal caveats
|
|
115
|
+
- 1-3: No limitations or alternative views
|
|
116
|
+
|
|
117
|
+
---
|
|
118
|
+
|
|
119
|
+
## Rule Rubric
|
|
120
|
+
|
|
121
|
+
### METHODOLOGY
|
|
122
|
+
- 9-10: Clear pattern/signal/fix structure, auto-detection instructions
|
|
123
|
+
- 7-8: Good table structure, detection guidance
|
|
124
|
+
- 4-6: List format, basic detection
|
|
125
|
+
- 1-3: Unstructured prose
|
|
126
|
+
|
|
127
|
+
### DEPTH
|
|
128
|
+
- 9-10: 8+ patterns, each with specific signals and actionable fixes
|
|
129
|
+
- 7-8: 5-7 patterns, good specificity
|
|
130
|
+
- 4-6: 3-4 patterns, generic
|
|
131
|
+
- 1-3: 1-2 patterns, vague
|
|
132
|
+
|
|
133
|
+
### CORRECTNESS
|
|
134
|
+
- 9-10: All patterns are real, signals are detectable, fixes work
|
|
135
|
+
- 7-8: Mostly accurate patterns
|
|
136
|
+
- 4-6: Some questionable patterns
|
|
137
|
+
- 1-3: Patterns don't match reality
|
|
138
|
+
|
|
139
|
+
### USABILITY
|
|
140
|
+
- 9-10: Scannable table format, clear action items
|
|
141
|
+
- 7-8: Table format, understandable
|
|
142
|
+
- 4-6: List format, requires interpretation
|
|
143
|
+
- 1-3: Hard to use as reference
|
|
144
|
+
|
|
145
|
+
### ROBUSTNESS
|
|
146
|
+
- 9-10: Covers common AND edge case patterns, priority ordering
|
|
147
|
+
- 7-8: Good pattern coverage
|
|
148
|
+
- 4-6: Only obvious patterns
|
|
149
|
+
- 1-3: Misses critical patterns
|
|
150
|
+
|
|
151
|
+
---
|
|
152
|
+
|
|
153
|
+
## Agent Template Rubric
|
|
154
|
+
|
|
155
|
+
### METHODOLOGY
|
|
156
|
+
- 9-10: Clear isolation, model selection, integration protocol
|
|
157
|
+
- 7-8: Good agent config, clear purpose
|
|
158
|
+
- 4-6: Basic config, unclear integration
|
|
159
|
+
- 1-3: No clear agent architecture
|
|
160
|
+
|
|
161
|
+
### DEPTH
|
|
162
|
+
- 9-10: Detailed prompt, context specification, output format
|
|
163
|
+
- 7-8: Good prompt, some output guidance
|
|
164
|
+
- 4-6: Basic prompt only
|
|
165
|
+
- 1-3: Stub template
|
|
166
|
+
|
|
167
|
+
### CORRECTNESS
|
|
168
|
+
- 9-10: Model choice justified, isolation enforced, no conflict with other agents
|
|
169
|
+
- 7-8: Reasonable model choice, basic isolation
|
|
170
|
+
- 4-6: Questionable model choice
|
|
171
|
+
- 1-3: Wrong model or no isolation
|
|
172
|
+
|
|
173
|
+
### USABILITY
|
|
174
|
+
- 9-10: Ready to use with Agent tool, clear parameters
|
|
175
|
+
- 7-8: Mostly ready, minor adjustments needed
|
|
176
|
+
- 4-6: Needs significant adaptation
|
|
177
|
+
- 1-3: Not usable as-is
|
|
178
|
+
|
|
179
|
+
### ROBUSTNESS
|
|
180
|
+
- 9-10: Timeout handling, failure protocol, cost bounds
|
|
181
|
+
- 7-8: Some failure handling
|
|
182
|
+
- 4-6: No failure handling but stable
|
|
183
|
+
- 1-3: Could cause issues (infinite loops, cost overrun)
|
|
@@ -0,0 +1,139 @@
|
|
|
1
|
+
# Prompt Optimization Methods
|
|
2
|
+
|
|
3
|
+
## Overview
|
|
4
|
+
|
|
5
|
+
Survey of automated prompt optimization approaches. BTO uses an EvoPrompt-inspired evolutionary approach adapted for Claude Code artifacts.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## Method 1: APE (Automatic Prompt Engineer)
|
|
10
|
+
|
|
11
|
+
**Paper:** Zhou et al., 2022 — "Large Language Models Are Human-Level Prompt Engineers"
|
|
12
|
+
|
|
13
|
+
**Approach:**
|
|
14
|
+
1. Generate candidate prompts from task description + examples
|
|
15
|
+
2. Score each candidate on a held-out evaluation set
|
|
16
|
+
3. Select the best-performing prompt
|
|
17
|
+
|
|
18
|
+
**Strengths:** Simple, effective for single-prompt optimization
|
|
19
|
+
**Weaknesses:** Requires evaluation dataset, single-round (no iteration)
|
|
20
|
+
**BTO Relevance:** Inspiration for variant generation step
|
|
21
|
+
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
## Method 2: OPRO (Optimization by Prompting)
|
|
25
|
+
|
|
26
|
+
**Paper:** Yang et al., 2023 — "Large Language Models as Optimizers"
|
|
27
|
+
|
|
28
|
+
**Approach:**
|
|
29
|
+
1. Maintain a trajectory of (prompt, score) pairs
|
|
30
|
+
2. Ask LLM to generate better prompt given the trajectory
|
|
31
|
+
3. Evaluate and add to trajectory
|
|
32
|
+
4. Repeat
|
|
33
|
+
|
|
34
|
+
**Key Insight:** LLMs can optimize when shown optimization trajectory.
|
|
35
|
+
|
|
36
|
+
**Strengths:** Leverages LLM understanding of what makes prompts good
|
|
37
|
+
**Weaknesses:** Can get stuck in local optima, expensive
|
|
38
|
+
**BTO Relevance:** Trajectory concept used in crossover step
|
|
39
|
+
|
|
40
|
+
---
|
|
41
|
+
|
|
42
|
+
## Method 3: EvoPrompt (Evolutionary)
|
|
43
|
+
|
|
44
|
+
**Paper:** Guo et al., 2023 — "Connecting Large Language Models with Evolutionary Algorithms"
|
|
45
|
+
|
|
46
|
+
**Approach:**
|
|
47
|
+
1. Initialize population of N prompts
|
|
48
|
+
2. Apply genetic operators: mutation + crossover
|
|
49
|
+
3. Evaluate fitness (task performance)
|
|
50
|
+
4. Select top-K, repeat
|
|
51
|
+
|
|
52
|
+
**Genetic Operators:**
|
|
53
|
+
- **Mutation:** Rephrase, add/remove constraints, change structure
|
|
54
|
+
- **Crossover:** Combine best parts of two prompts
|
|
55
|
+
|
|
56
|
+
**Strengths:** Explores diverse solutions, avoids local optima
|
|
57
|
+
**Weaknesses:** Expensive (many evaluations), slow convergence
|
|
58
|
+
**BTO Relevance:** Primary inspiration for OPTIMIZE module
|
|
59
|
+
|
|
60
|
+
### BTO Adaptation of EvoPrompt
|
|
61
|
+
|
|
62
|
+
```
|
|
63
|
+
EvoPrompt (Original) BTO OPTIMIZE (Adapted)
|
|
64
|
+
───────────────── ─────────────────────
|
|
65
|
+
Random init population → 5 targeted mutations (strategy-driven)
|
|
66
|
+
Generic mutation → Named strategies (Rephrase/Restructure/etc.)
|
|
67
|
+
Random crossover → Strength-based crossover (section-level)
|
|
68
|
+
Task accuracy metric → Multi-dimensional judge panel scores
|
|
69
|
+
Many generations → 3 rounds (cost-bounded)
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
---
|
|
73
|
+
|
|
74
|
+
## Method 4: TextGrad (Textual Gradients)
|
|
75
|
+
|
|
76
|
+
**Paper:** Yuksekgonul et al., 2024 — "TextGrad: Automatic Differentiation via Text"
|
|
77
|
+
|
|
78
|
+
**Approach:**
|
|
79
|
+
1. Evaluate prompt output quality
|
|
80
|
+
2. Generate "textual gradient" — natural language feedback on what to improve
|
|
81
|
+
3. Apply gradient as edit instructions
|
|
82
|
+
4. Repeat
|
|
83
|
+
|
|
84
|
+
**Strengths:** Targeted improvements, interpretable changes
|
|
85
|
+
**Weaknesses:** Requires clear evaluation criteria, can overfit
|
|
86
|
+
**BTO Relevance:** Judge feedback as textual gradients for mutation targeting
|
|
87
|
+
|
|
88
|
+
---
|
|
89
|
+
|
|
90
|
+
## Method 5: DSPy (Programmatic)
|
|
91
|
+
|
|
92
|
+
**Framework:** Khattab et al., 2023 — "DSPy: Compiling Declarative Language Model Calls"
|
|
93
|
+
|
|
94
|
+
**Approach:**
|
|
95
|
+
1. Define modules with typed signatures
|
|
96
|
+
2. Compile with optimizer (BootstrapFewShot, MIPRO, etc.)
|
|
97
|
+
3. Optimizer tunes prompts + few-shot examples
|
|
98
|
+
4. Evaluate on metric
|
|
99
|
+
|
|
100
|
+
**Strengths:** Systematic, reproducible, handles multi-step pipelines
|
|
101
|
+
**Weaknesses:** Requires code framework, Python-specific
|
|
102
|
+
**BTO Relevance:** Modular skill design parallels DSPy module concept
|
|
103
|
+
|
|
104
|
+
---
|
|
105
|
+
|
|
106
|
+
## Comparison Matrix
|
|
107
|
+
|
|
108
|
+
| Method | Iterations | Cost | Diversity | Interpretability | Best For |
|
|
109
|
+
|--------|-----------|------|-----------|-----------------|----------|
|
|
110
|
+
| APE | 1 | Low | Medium | High | Simple prompts |
|
|
111
|
+
| OPRO | 5-20 | High | Low | High | Single metric |
|
|
112
|
+
| EvoPrompt | 5-50 | Very High | High | Medium | Complex prompts |
|
|
113
|
+
| TextGrad | 3-10 | Medium | Low | Very High | Targeted fixes |
|
|
114
|
+
| DSPy | Auto | Variable | Medium | Low | Pipelines |
|
|
115
|
+
| **BTO** | **3** | **Medium** | **High** | **High** | **Claude Code artifacts** |
|
|
116
|
+
|
|
117
|
+
---
|
|
118
|
+
|
|
119
|
+
## BTO Design Decisions
|
|
120
|
+
|
|
121
|
+
### Why Evolutionary (not gradient)?
|
|
122
|
+
- Claude Code artifacts are structured documents, not single prompts
|
|
123
|
+
- Mutations at section level are more meaningful than token-level edits
|
|
124
|
+
- Diversity prevents converging to one style
|
|
125
|
+
|
|
126
|
+
### Why 3 rounds (not more)?
|
|
127
|
+
- Diminishing returns after round 3
|
|
128
|
+
- Cost budget: ~15 evaluations is practical
|
|
129
|
+
- Layer 2 on final round ensures quality
|
|
130
|
+
|
|
131
|
+
### Why strategy-driven mutations (not random)?
|
|
132
|
+
- Named strategies map to specific weaknesses
|
|
133
|
+
- More interpretable changelog
|
|
134
|
+
- Faster convergence than random exploration
|
|
135
|
+
|
|
136
|
+
### Why section-level crossover?
|
|
137
|
+
- Skills have clear section boundaries
|
|
138
|
+
- Strengths are often section-specific
|
|
139
|
+
- Easier to preserve coherence than token-level mixing
|