@dzhechkov/skills-bto 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,202 @@
1
+ # Multi-Agent Evaluation Patterns
2
+
3
+ ## Overview
4
+
5
+ Research-backed patterns for using multiple LLM agents to evaluate content. Each pattern trades off cost, reliability, and depth.
6
+
7
+ ---
8
+
9
+ ## Pattern 1: Judge Panel (Used by BTO)
10
+
11
+ **Architecture:** N independent judges evaluate the same artifact in parallel.
12
+
13
+ ```
14
+ Artifact ──┬──→ Judge 1 (Expert) ──┐
15
+ ├──→ Judge 2 (Critic) ──┼──→ Aggregator → Report
16
+ └──→ Judge 3 (Auditor) ──┘
17
+ ```
18
+
19
+ **Strengths:**
20
+ - Multi-perspective coverage
21
+ - Parallel execution (fast)
22
+ - Disagreement detection reveals ambiguous quality
23
+
24
+ **Weaknesses:**
25
+ - Judges may converge to similar opinions (groupthink)
26
+ - More expensive than single judge
27
+ - Requires aggregation logic
28
+
29
+ **When to use:** Standard evaluation of well-defined artifacts. BTO default.
30
+
31
+ **Mitigation for groupthink:**
32
+ - Give each judge a distinct role and perspective
33
+ - Critic judge explicitly instructed to be strict (calibrated lower)
34
+ - Different temperature settings if available
35
+
36
+ ---
37
+
38
+ ## Pattern 2: Adversarial Red/Blue
39
+
40
+ **Architecture:** One agent creates, another attacks.
41
+
42
+ ```
43
+ Artifact ──→ Blue Team (Defender) ──→ Assessment
44
+ ──→ Red Team (Attacker) ──→ Vulnerabilities
45
+ ──→ Reconciliation
46
+ ```
47
+
48
+ **Strengths:**
49
+ - Excellent at finding weaknesses
50
+ - Simulates real-world usage patterns
51
+ - Defender forced to justify every choice
52
+
53
+ **Weaknesses:**
54
+ - Can be overly negative
55
+ - Expensive (multiple rounds)
56
+ - Not suited for overall quality scoring
57
+
58
+ **When to use:** Security-sensitive artifacts, high-stakes decisions, or when robustness is critical.
59
+
60
+ ---
61
+
62
+ ## Pattern 3: Consensus-Based (Multi-Round Debate)
63
+
64
+ **Architecture:** Agents debate until convergence.
65
+
66
+ ```
67
+ Round 1: Each judge evaluates independently
68
+ Round 2: Judges see each other's evaluations, revise
69
+ Round 3: Final convergence (or flag for human)
70
+ ```
71
+
72
+ **Strengths:**
73
+ - High-confidence final scores
74
+ - Reduces individual judge bias
75
+ - Self-correcting
76
+
77
+ **Weaknesses:**
78
+ - Very expensive (3x the calls)
79
+ - Risk of conformity pressure
80
+ - Slow (sequential rounds)
81
+
82
+ **When to use:** Critical decisions where confidence matters more than speed.
83
+
84
+ ---
85
+
86
+ ## Pattern 4: Constitutional Evaluation
87
+
88
+ **Architecture:** Evaluate against a set of explicit principles.
89
+
90
+ ```
91
+ Principles ──→ Evaluator ──→ Per-principle score ──→ Report
92
+ ```
93
+
94
+ **Principles example for Claude Code artifacts:**
95
+ 1. Instructions must be unambiguous
96
+ 2. Every section must serve a purpose
97
+ 3. Anti-patterns must be documented
98
+ 4. Examples must demonstrate the happy path
99
+ 5. Failure modes must be handled
100
+
101
+ **Strengths:**
102
+ - Deterministic criteria
103
+ - Easy to explain scores
104
+ - Reproducible
105
+
106
+ **Weaknesses:**
107
+ - Misses emergent quality issues
108
+ - Principles may not cover everything
109
+ - Rigid
110
+
111
+ **When to use:** Compliance checking, standardized quality gates.
112
+
113
+ ---
114
+
115
+ ## Pattern 5: Hierarchical Evaluation
116
+
117
+ **Architecture:** Fast cheap filter → detailed expensive evaluation.
118
+
119
+ ```
120
+ Layer 0: Deterministic ──pass──→ Layer 1: Haiku ──pass──→ Layer 2: Sonnet Panel
121
+ │ │
122
+ fail fail
123
+ ↓ ↓
124
+ Quick Fix Detailed Feedback
125
+ ```
126
+
127
+ **Strengths:**
128
+ - Cost-efficient (most artifacts filtered early)
129
+ - Progressive detail
130
+ - Fast feedback for obvious issues
131
+
132
+ **Weaknesses:**
133
+ - Layer 0 can miss semantic issues
134
+ - Layer 1 may have different calibration than Layer 2
135
+
136
+ **When to use:** Batch evaluation, CI/CD pipelines, optimization loops. BTO uses this pattern.
137
+
138
+ ---
139
+
140
+ ## Aggregation Strategies
141
+
142
+ ### 1. Weighted Average (BTO Default)
143
+ ```
144
+ score = Σ(weight_i × judge_i_score) / Σ(weight_i)
145
+ ```
146
+ - Simple, interpretable
147
+ - Weights reflect judge importance
148
+ - BTO weights: Expert 0.4, Critic 0.3, Auditor 0.3
149
+
150
+ ### 2. Majority Voting
151
+ ```
152
+ verdict = mode(judge_verdicts)
153
+ ```
154
+ - Good for pass/fail decisions
155
+ - Requires odd number of judges
156
+ - Ignores score magnitude
157
+
158
+ ### 3. Bayesian Aggregation
159
+ ```
160
+ P(quality | scores) ∝ Π P(score_i | quality) × P(quality)
161
+ ```
162
+ - Accounts for judge reliability
163
+ - Requires calibration data
164
+ - More complex to implement
165
+
166
+ ### 4. Min-Score (Conservative)
167
+ ```
168
+ score = min(all_judge_scores)
169
+ ```
170
+ - Most conservative
171
+ - Good for safety-critical artifacts
172
+ - May be overly pessimistic
173
+
174
+ ### 5. Trimmed Mean
175
+ ```
176
+ score = mean(scores after removing highest and lowest)
177
+ ```
178
+ - Outlier-resistant
179
+ - Requires ≥5 judges
180
+ - Good for large panels
181
+
182
+ ---
183
+
184
+ ## Anti-Conformity Measures
185
+
186
+ Prevent judges from converging to meaningless agreement:
187
+
188
+ 1. **Role Differentiation** — Each judge has unique focus and scoring calibration
189
+ 2. **Critic Calibration** — Critic judge instructed to score ~5-6 average (strict)
190
+ 3. **Independent Evaluation** — No judge sees other judges' scores
191
+ 4. **Disagreement as Signal** — High disagreement flags important quality dimensions
192
+ 5. **Diverse Prompts** — Each judge gets a differently-framed evaluation prompt
193
+
194
+ ## Cost Optimization
195
+
196
+ | Approach | Judges | Model | Est. Cost | Use When |
197
+ |----------|--------|-------|-----------|----------|
198
+ | Layer 0 only | 0 | none | Free | CI/CD pre-check |
199
+ | Layer 1 | 1 | haiku | ~$0.001 | Quick iteration |
200
+ | Layer 2 | 3 | sonnet | ~$0.01 | Thorough evaluation |
201
+ | Full + Meta | 3+1 | sonnet+opus | ~$0.05 | Critical artifacts |
202
+ | Consensus | 3×3 | sonnet | ~$0.03 | High-stakes decisions |
@@ -0,0 +1,183 @@
1
+ # Judge Rubrics — Evaluation Criteria by Artifact Type
2
+
3
+ ## Universal Dimensions (1-10 scale)
4
+
5
+ All artifact types are evaluated on these 5 dimensions:
6
+
7
+ | Dimension | What it Measures | 1-3 (Poor) | 4-6 (Adequate) | 7-8 (Good) | 9-10 (Excellent) |
8
+ |-----------|-----------------|------------|----------------|------------|-------------------|
9
+ | METHODOLOGY | Approach soundness | No clear structure | Basic structure present | Well-designed protocol | Rigorous, innovative approach |
10
+ | DEPTH | Content thoroughness | Superficial | Covers basics | Thorough treatment | Comprehensive with nuances |
11
+ | CORRECTNESS | Accuracy of claims/instructions | Major errors | Minor inaccuracies | Accurate | Precise and verifiable |
12
+ | USABILITY | Ease of use | Confusing | Usable with effort | Clear and navigable | Intuitive, self-documenting |
13
+ | ROBUSTNESS | Edge case handling | No coverage | Some mentioned | Well-covered | Exhaustive with fallbacks |
14
+
15
+ ---
16
+
17
+ ## Skill Rubric
18
+
19
+ ### METHODOLOGY
20
+ - 9-10: Clear multi-step protocol, modular design, explicit decision points
21
+ - 7-8: Good protocol structure, some modularity
22
+ - 4-6: Basic steps listed, no clear decision flow
23
+ - 1-3: No protocol, just description of what skill does
24
+
25
+ ### DEPTH
26
+ - 9-10: Detailed modules, rich references, multiple examples, anti-patterns
27
+ - 7-8: Good detail in SKILL.md, has references and examples
28
+ - 4-6: Basic overview, minimal references
29
+ - 1-3: Stub-level content, no supporting files
30
+
31
+ ### CORRECTNESS
32
+ - 9-10: All instructions produce expected output, cross-references valid
33
+ - 7-8: Instructions work, minor reference issues
34
+ - 4-6: Some instructions ambiguous, broken references
35
+ - 1-3: Instructions would produce wrong output
36
+
37
+ ### USABILITY
38
+ - 9-10: Quick start section, clear navigation, progressive disclosure
39
+ - 7-8: Well-organized, findable information
40
+ - 4-6: Readable but poorly organized
41
+ - 1-3: Wall of text, no structure
42
+
43
+ ### ROBUSTNESS
44
+ - 9-10: Anti-patterns documented, failure modes handled, abort conditions defined
45
+ - 7-8: Anti-patterns section present, some edge cases
46
+ - 4-6: Mentions some risks
47
+ - 1-3: No failure mode coverage
48
+
49
+ ---
50
+
51
+ ## Command Rubric
52
+
53
+ ### METHODOLOGY
54
+ - 9-10: Clear step-by-step flow, proper checkpoints, skill loading pattern
55
+ - 7-8: Good flow, checkpoint present
56
+ - 4-6: Basic steps, missing checkpoint
57
+ - 1-3: No clear execution flow
58
+
59
+ ### DEPTH
60
+ - 9-10: Handles all parameter combinations, parallel agent usage
61
+ - 7-8: Good parameter handling, some agent usage
62
+ - 4-6: Basic parameter handling
63
+ - 1-3: No parameter handling
64
+
65
+ ### CORRECTNESS
66
+ - 9-10: $ARGUMENTS properly used, skill paths correct, agent configs valid
67
+ - 7-8: Mostly correct references
68
+ - 4-6: Some incorrect references
69
+ - 1-3: Wrong skill paths or broken references
70
+
71
+ ### USABILITY
72
+ - 9-10: Clear usage line, example invocations, helpful error messages
73
+ - 7-8: Usage documented, some examples
74
+ - 4-6: Basic usage, no examples
75
+ - 1-3: Unclear how to invoke
76
+
77
+ ### ROBUSTNESS
78
+ - 9-10: Input validation, fallbacks, timeout handling
79
+ - 7-8: Some input validation
80
+ - 4-6: Minimal validation
81
+ - 1-3: No error handling
82
+
83
+ ---
84
+
85
+ ## Research Artifact Rubric
86
+
87
+ ### METHODOLOGY
88
+ - 9-10: PARANOID mode, systematic research plan, structured sections
89
+ - 7-8: Good research structure, mostly verified
90
+ - 4-6: Basic structure, some unverified claims
91
+ - 1-3: Unstructured, many unverified claims
92
+
93
+ ### DEPTH
94
+ - 9-10: Multiple sources per claim, competitive analysis, quantitative data
95
+ - 7-8: Good source coverage, some quantitative data
96
+ - 4-6: Few sources, mostly qualitative
97
+ - 1-3: No sources, opinion-based
98
+
99
+ ### CORRECTNESS
100
+ - 9-10: All claims sourced, no [UNVERIFIED] tags, real products/companies
101
+ - 7-8: Mostly sourced, few [UNVERIFIED]
102
+ - 4-6: Some sourced, several [UNVERIFIED]
103
+ - 1-3: Unsourced, potentially hallucinated
104
+
105
+ ### USABILITY
106
+ - 9-10: Executive summary, clear sections, actionable takeaways
107
+ - 7-8: Well-organized, some actionable content
108
+ - 4-6: Readable but hard to act on
109
+ - 1-3: Dense, unstructured
110
+
111
+ ### ROBUSTNESS
112
+ - 9-10: Limitations documented, alternative views considered, confidence levels
113
+ - 7-8: Some limitations noted
114
+ - 4-6: Minimal caveats
115
+ - 1-3: No limitations or alternative views
116
+
117
+ ---
118
+
119
+ ## Rule Rubric
120
+
121
+ ### METHODOLOGY
122
+ - 9-10: Clear pattern/signal/fix structure, auto-detection instructions
123
+ - 7-8: Good table structure, detection guidance
124
+ - 4-6: List format, basic detection
125
+ - 1-3: Unstructured prose
126
+
127
+ ### DEPTH
128
+ - 9-10: 8+ patterns, each with specific signals and actionable fixes
129
+ - 7-8: 5-7 patterns, good specificity
130
+ - 4-6: 3-4 patterns, generic
131
+ - 1-3: 1-2 patterns, vague
132
+
133
+ ### CORRECTNESS
134
+ - 9-10: All patterns are real, signals are detectable, fixes work
135
+ - 7-8: Mostly accurate patterns
136
+ - 4-6: Some questionable patterns
137
+ - 1-3: Patterns don't match reality
138
+
139
+ ### USABILITY
140
+ - 9-10: Scannable table format, clear action items
141
+ - 7-8: Table format, understandable
142
+ - 4-6: List format, requires interpretation
143
+ - 1-3: Hard to use as reference
144
+
145
+ ### ROBUSTNESS
146
+ - 9-10: Covers common AND edge case patterns, priority ordering
147
+ - 7-8: Good pattern coverage
148
+ - 4-6: Only obvious patterns
149
+ - 1-3: Misses critical patterns
150
+
151
+ ---
152
+
153
+ ## Agent Template Rubric
154
+
155
+ ### METHODOLOGY
156
+ - 9-10: Clear isolation, model selection, integration protocol
157
+ - 7-8: Good agent config, clear purpose
158
+ - 4-6: Basic config, unclear integration
159
+ - 1-3: No clear agent architecture
160
+
161
+ ### DEPTH
162
+ - 9-10: Detailed prompt, context specification, output format
163
+ - 7-8: Good prompt, some output guidance
164
+ - 4-6: Basic prompt only
165
+ - 1-3: Stub template
166
+
167
+ ### CORRECTNESS
168
+ - 9-10: Model choice justified, isolation enforced, no conflict with other agents
169
+ - 7-8: Reasonable model choice, basic isolation
170
+ - 4-6: Questionable model choice
171
+ - 1-3: Wrong model or no isolation
172
+
173
+ ### USABILITY
174
+ - 9-10: Ready to use with Agent tool, clear parameters
175
+ - 7-8: Mostly ready, minor adjustments needed
176
+ - 4-6: Needs significant adaptation
177
+ - 1-3: Not usable as-is
178
+
179
+ ### ROBUSTNESS
180
+ - 9-10: Timeout handling, failure protocol, cost bounds
181
+ - 7-8: Some failure handling
182
+ - 4-6: No failure handling but stable
183
+ - 1-3: Could cause issues (infinite loops, cost overrun)
@@ -0,0 +1,139 @@
1
+ # Prompt Optimization Methods
2
+
3
+ ## Overview
4
+
5
+ Survey of automated prompt optimization approaches. BTO uses an EvoPrompt-inspired evolutionary approach adapted for Claude Code artifacts.
6
+
7
+ ---
8
+
9
+ ## Method 1: APE (Automatic Prompt Engineer)
10
+
11
+ **Paper:** Zhou et al., 2022 — "Large Language Models Are Human-Level Prompt Engineers"
12
+
13
+ **Approach:**
14
+ 1. Generate candidate prompts from task description + examples
15
+ 2. Score each candidate on a held-out evaluation set
16
+ 3. Select the best-performing prompt
17
+
18
+ **Strengths:** Simple, effective for single-prompt optimization
19
+ **Weaknesses:** Requires evaluation dataset, single-round (no iteration)
20
+ **BTO Relevance:** Inspiration for variant generation step
21
+
22
+ ---
23
+
24
+ ## Method 2: OPRO (Optimization by Prompting)
25
+
26
+ **Paper:** Yang et al., 2023 — "Large Language Models as Optimizers"
27
+
28
+ **Approach:**
29
+ 1. Maintain a trajectory of (prompt, score) pairs
30
+ 2. Ask LLM to generate better prompt given the trajectory
31
+ 3. Evaluate and add to trajectory
32
+ 4. Repeat
33
+
34
+ **Key Insight:** LLMs can optimize when shown optimization trajectory.
35
+
36
+ **Strengths:** Leverages LLM understanding of what makes prompts good
37
+ **Weaknesses:** Can get stuck in local optima, expensive
38
+ **BTO Relevance:** Trajectory concept used in crossover step
39
+
40
+ ---
41
+
42
+ ## Method 3: EvoPrompt (Evolutionary)
43
+
44
+ **Paper:** Guo et al., 2023 — "Connecting Large Language Models with Evolutionary Algorithms"
45
+
46
+ **Approach:**
47
+ 1. Initialize population of N prompts
48
+ 2. Apply genetic operators: mutation + crossover
49
+ 3. Evaluate fitness (task performance)
50
+ 4. Select top-K, repeat
51
+
52
+ **Genetic Operators:**
53
+ - **Mutation:** Rephrase, add/remove constraints, change structure
54
+ - **Crossover:** Combine best parts of two prompts
55
+
56
+ **Strengths:** Explores diverse solutions, avoids local optima
57
+ **Weaknesses:** Expensive (many evaluations), slow convergence
58
+ **BTO Relevance:** Primary inspiration for OPTIMIZE module
59
+
60
+ ### BTO Adaptation of EvoPrompt
61
+
62
+ ```
63
+ EvoPrompt (Original) BTO OPTIMIZE (Adapted)
64
+ ───────────────── ─────────────────────
65
+ Random init population → 5 targeted mutations (strategy-driven)
66
+ Generic mutation → Named strategies (Rephrase/Restructure/etc.)
67
+ Random crossover → Strength-based crossover (section-level)
68
+ Task accuracy metric → Multi-dimensional judge panel scores
69
+ Many generations → 3 rounds (cost-bounded)
70
+ ```
71
+
72
+ ---
73
+
74
+ ## Method 4: TextGrad (Textual Gradients)
75
+
76
+ **Paper:** Yuksekgonul et al., 2024 — "TextGrad: Automatic Differentiation via Text"
77
+
78
+ **Approach:**
79
+ 1. Evaluate prompt output quality
80
+ 2. Generate "textual gradient" — natural language feedback on what to improve
81
+ 3. Apply gradient as edit instructions
82
+ 4. Repeat
83
+
84
+ **Strengths:** Targeted improvements, interpretable changes
85
+ **Weaknesses:** Requires clear evaluation criteria, can overfit
86
+ **BTO Relevance:** Judge feedback as textual gradients for mutation targeting
87
+
88
+ ---
89
+
90
+ ## Method 5: DSPy (Programmatic)
91
+
92
+ **Framework:** Khattab et al., 2023 — "DSPy: Compiling Declarative Language Model Calls"
93
+
94
+ **Approach:**
95
+ 1. Define modules with typed signatures
96
+ 2. Compile with optimizer (BootstrapFewShot, MIPRO, etc.)
97
+ 3. Optimizer tunes prompts + few-shot examples
98
+ 4. Evaluate on metric
99
+
100
+ **Strengths:** Systematic, reproducible, handles multi-step pipelines
101
+ **Weaknesses:** Requires code framework, Python-specific
102
+ **BTO Relevance:** Modular skill design parallels DSPy module concept
103
+
104
+ ---
105
+
106
+ ## Comparison Matrix
107
+
108
+ | Method | Iterations | Cost | Diversity | Interpretability | Best For |
109
+ |--------|-----------|------|-----------|-----------------|----------|
110
+ | APE | 1 | Low | Medium | High | Simple prompts |
111
+ | OPRO | 5-20 | High | Low | High | Single metric |
112
+ | EvoPrompt | 5-50 | Very High | High | Medium | Complex prompts |
113
+ | TextGrad | 3-10 | Medium | Low | Very High | Targeted fixes |
114
+ | DSPy | Auto | Variable | Medium | Low | Pipelines |
115
+ | **BTO** | **3** | **Medium** | **High** | **High** | **Claude Code artifacts** |
116
+
117
+ ---
118
+
119
+ ## BTO Design Decisions
120
+
121
+ ### Why Evolutionary (not gradient)?
122
+ - Claude Code artifacts are structured documents, not single prompts
123
+ - Mutations at section level are more meaningful than token-level edits
124
+ - Diversity prevents converging to one style
125
+
126
+ ### Why 3 rounds (not more)?
127
+ - Diminishing returns after round 3
128
+ - Cost budget: ~15 evaluations is practical
129
+ - Layer 2 on final round ensures quality
130
+
131
+ ### Why strategy-driven mutations (not random)?
132
+ - Named strategies map to specific weaknesses
133
+ - More interpretable changelog
134
+ - Faster convergence than random exploration
135
+
136
+ ### Why section-level crossover?
137
+ - Skills have clear section boundaries
138
+ - Strengths are often section-specific
139
+ - Easier to preserve coherence than token-level mixing