@dzhechkov/skills-bto 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bin/cli.js +5 -0
- package/package.json +43 -0
- package/src/cli.js +150 -0
- package/src/commands/doctor.js +366 -0
- package/src/commands/init.js +188 -0
- package/src/commands/list.js +161 -0
- package/src/commands/remove.js +211 -0
- package/src/commands/update.js +198 -0
- package/src/utils.js +398 -0
- package/templates/.claude/agents/bto-judge-panel.md +192 -0
- package/templates/.claude/agents/bto-optimizer-worker.md +181 -0
- package/templates/.claude/commands/bto-build.md +169 -0
- package/templates/.claude/commands/bto-optimize.md +208 -0
- package/templates/.claude/commands/bto-test.md +186 -0
- package/templates/.claude/commands/bto.md +171 -0
- package/templates/.claude/rules/bto-quality-gates.md +91 -0
- package/templates/.claude/skills/bto/SKILL.md +266 -0
- package/templates/.claude/skills/bto/examples/sample-eval-report.md +436 -0
- package/templates/.claude/skills/bto/modules/build.md +189 -0
- package/templates/.claude/skills/bto/modules/optimize.md +201 -0
- package/templates/.claude/skills/bto/modules/test.md +348 -0
- package/templates/.claude/skills/bto/references/eval-patterns.md +202 -0
- package/templates/.claude/skills/bto/references/judge-rubrics.md +183 -0
- package/templates/.claude/skills/bto/references/optimization-methods.md +139 -0
- package/templates/.claude/skills/bto/references/quality-checklist.md +220 -0
|
@@ -0,0 +1,201 @@
|
|
|
1
|
+
# OPTIMIZE Module — Evolutionary Prompt Optimization Protocol
|
|
2
|
+
|
|
3
|
+
## Purpose
|
|
4
|
+
|
|
5
|
+
Improve Claude Code artifacts through evolutionary prompt optimization: generate variants, evaluate, select, mutate, repeat.
|
|
6
|
+
|
|
7
|
+
## Input
|
|
8
|
+
|
|
9
|
+
- **Path:** Path to artifact to optimize
|
|
10
|
+
- **Rounds:** Number of optimization rounds (default: 3, max: 5)
|
|
11
|
+
- **Budget:** max evaluations (default: 15)
|
|
12
|
+
- **Focus:** Optional dimension to prioritize (METHODOLOGY / DEPTH / CORRECTNESS / USABILITY / ROBUSTNESS)
|
|
13
|
+
|
|
14
|
+
## Prerequisites
|
|
15
|
+
|
|
16
|
+
- Artifact must pass Layer 0 checks (run TEST first)
|
|
17
|
+
- Baseline score established (Layer 2 evaluation)
|
|
18
|
+
- Only optimize if baseline < 8.0 (otherwise artifact is already good)
|
|
19
|
+
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
## Protocol
|
|
23
|
+
|
|
24
|
+
### Step 1: Baseline Evaluation
|
|
25
|
+
|
|
26
|
+
1. Run TEST module with level=layer2 on current artifact
|
|
27
|
+
2. Record:
|
|
28
|
+
- Per-dimension scores
|
|
29
|
+
- Overall score
|
|
30
|
+
- Key weaknesses identified by judges
|
|
31
|
+
3. If overall ≥ 8.0 → report "Artifact already high quality" and suggest minor tweaks only
|
|
32
|
+
4. Identify **target dimensions**: dimensions scoring < 7.0
|
|
33
|
+
|
|
34
|
+
### Step 2: Variant Generation (Round 1)
|
|
35
|
+
|
|
36
|
+
Generate N=5 variants of the artifact. Each variant applies ONE mutation strategy.
|
|
37
|
+
|
|
38
|
+
**Mutation Assignment:**
|
|
39
|
+
- Variant 1: **Rephrase** — reword unclear instructions for precision
|
|
40
|
+
- Variant 2: **Restructure** — reorganize sections for better flow
|
|
41
|
+
- Variant 3: **Add Constraints** — add guardrails, edge cases, boundary conditions
|
|
42
|
+
- Variant 4: **Simplify** — remove redundancy, tighten language, reduce verbosity
|
|
43
|
+
- Variant 5: **Specialize** — add domain-specific context and examples
|
|
44
|
+
|
|
45
|
+
**Strategy-to-Weakness Mapping:**
|
|
46
|
+
|
|
47
|
+
| Weak Dimension | Primary Strategy | Secondary Strategy |
|
|
48
|
+
|---------------|-----------------|-------------------|
|
|
49
|
+
| METHODOLOGY | Restructure | Add Constraints |
|
|
50
|
+
| DEPTH | Specialize | Add Constraints |
|
|
51
|
+
| CORRECTNESS | Add Constraints | Rephrase |
|
|
52
|
+
| USABILITY | Rephrase | Restructure |
|
|
53
|
+
| ROBUSTNESS | Add Constraints | Specialize |
|
|
54
|
+
|
|
55
|
+
If a focus dimension is specified, generate 3 variants targeting that dimension's strategies and 2 variants with other strategies.
|
|
56
|
+
|
|
57
|
+
### Variant Generation Prompt
|
|
58
|
+
|
|
59
|
+
```
|
|
60
|
+
You are optimizing a Claude Code {artifact_type}.
|
|
61
|
+
|
|
62
|
+
## Current Artifact
|
|
63
|
+
{content}
|
|
64
|
+
|
|
65
|
+
## Baseline Evaluation
|
|
66
|
+
Overall: {score}/10
|
|
67
|
+
Weaknesses: {weaknesses}
|
|
68
|
+
Target dimensions: {target_dimensions}
|
|
69
|
+
|
|
70
|
+
## Mutation Strategy: {strategy_name}
|
|
71
|
+
{strategy_description}
|
|
72
|
+
|
|
73
|
+
## Task
|
|
74
|
+
Apply the {strategy_name} mutation to improve this artifact.
|
|
75
|
+
Focus on addressing these specific weaknesses: {target_weaknesses}
|
|
76
|
+
|
|
77
|
+
Rules:
|
|
78
|
+
- Preserve the original intent and scope
|
|
79
|
+
- Maintain all existing sections
|
|
80
|
+
- Do not change the artifact type or structure fundamentally
|
|
81
|
+
- Changes should be targeted and purposeful
|
|
82
|
+
- Output the COMPLETE modified artifact (not a diff)
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
### Step 3: Evaluate Variants (Round 1)
|
|
86
|
+
|
|
87
|
+
Run TEST module with level=layer1 (haiku — fast and cheap) on each variant.
|
|
88
|
+
|
|
89
|
+
**Parallel execution:** Spawn 5 agents (model: haiku), one per variant.
|
|
90
|
+
|
|
91
|
+
Record scores for all 5 variants.
|
|
92
|
+
|
|
93
|
+
### Step 4: Selection + Crossover
|
|
94
|
+
|
|
95
|
+
1. **Select:** Top 2 variants by overall score
|
|
96
|
+
2. **Crossover:** Combine best elements:
|
|
97
|
+
- Take sections where Variant A scored higher from A
|
|
98
|
+
- Take sections where Variant B scored higher from B
|
|
99
|
+
- Generate 3 new variants from the crossover
|
|
100
|
+
|
|
101
|
+
**Crossover Prompt:**
|
|
102
|
+
```
|
|
103
|
+
You are creating an improved Claude Code artifact by combining the best
|
|
104
|
+
elements of two high-scoring variants.
|
|
105
|
+
|
|
106
|
+
## Variant A (Score: {score_a})
|
|
107
|
+
{variant_a}
|
|
108
|
+
Strengths: {strengths_a}
|
|
109
|
+
|
|
110
|
+
## Variant B (Score: {score_b})
|
|
111
|
+
{variant_b}
|
|
112
|
+
Strengths: {strengths_b}
|
|
113
|
+
|
|
114
|
+
## Task
|
|
115
|
+
Create a new variant that combines the strengths of both:
|
|
116
|
+
- From A, take: {specific_sections_a}
|
|
117
|
+
- From B, take: {specific_sections_b}
|
|
118
|
+
- Ensure coherence and consistency
|
|
119
|
+
- Output the COMPLETE artifact
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
### Step 5: Evaluate + Select (Rounds 2-3)
|
|
123
|
+
|
|
124
|
+
**Round 2:**
|
|
125
|
+
- Evaluate 3 crossover variants with Layer 1
|
|
126
|
+
- Select top 2
|
|
127
|
+
- Generate 3 new crossover variants
|
|
128
|
+
|
|
129
|
+
**Round 3 (Final):**
|
|
130
|
+
- Evaluate 3 final variants with **Layer 2** (full judge panel — thorough)
|
|
131
|
+
- Select the single best variant
|
|
132
|
+
|
|
133
|
+
### Step 6: Output
|
|
134
|
+
|
|
135
|
+
1. **Best variant** — the optimized artifact
|
|
136
|
+
2. **Before/After comparison:**
|
|
137
|
+
```
|
|
138
|
+
═══════════════════════════════════════════════════════
|
|
139
|
+
🔧 BTO OPTIMIZATION REPORT
|
|
140
|
+
Artifact: <path>
|
|
141
|
+
Rounds: 3
|
|
142
|
+
Total evaluations: 15
|
|
143
|
+
|
|
144
|
+
BEFORE → AFTER:
|
|
145
|
+
METHODOLOGY: 6.2 → 8.1 (+1.9) ⬆️
|
|
146
|
+
DEPTH: 5.8 → 7.5 (+1.7) ⬆️
|
|
147
|
+
CORRECTNESS: 7.0 → 8.3 (+1.3) ⬆️
|
|
148
|
+
USABILITY: 6.5 → 8.0 (+1.5) ⬆️
|
|
149
|
+
ROBUSTNESS: 5.5 → 7.8 (+2.3) ⬆️
|
|
150
|
+
|
|
151
|
+
OVERALL: 6.2 → 7.9 (+1.7) ⬆️
|
|
152
|
+
|
|
153
|
+
Winning Strategy: Restructure + Add Constraints (crossover)
|
|
154
|
+
|
|
155
|
+
CHANGELOG:
|
|
156
|
+
- Restructured protocol into clearer numbered steps
|
|
157
|
+
- Added edge case handling for empty inputs
|
|
158
|
+
- Expanded anti-patterns with 3 new entries
|
|
159
|
+
- Simplified module loading instructions
|
|
160
|
+
- Added concrete examples to each section
|
|
161
|
+
═══════════════════════════════════════════════════════
|
|
162
|
+
```
|
|
163
|
+
|
|
164
|
+
3. **Recommendation:**
|
|
165
|
+
- If improvement > 1.0: "Apply changes"
|
|
166
|
+
- If improvement 0.5-1.0: "Review changes, consider applying"
|
|
167
|
+
- If improvement < 0.5: "Minimal improvement — original may be preferred"
|
|
168
|
+
|
|
169
|
+
---
|
|
170
|
+
|
|
171
|
+
## Cost Summary
|
|
172
|
+
|
|
173
|
+
| Operation | Count | Model | Est. Tokens |
|
|
174
|
+
|-----------|-------|-------|-------------|
|
|
175
|
+
| Baseline eval | 1 | sonnet ×3 | ~15K |
|
|
176
|
+
| Variant generation | 5 | opus | ~25K |
|
|
177
|
+
| Round 1 eval | 5 | haiku | ~10K |
|
|
178
|
+
| Crossover generation | 3 | opus | ~15K |
|
|
179
|
+
| Round 2 eval | 3 | haiku | ~6K |
|
|
180
|
+
| Crossover generation | 3 | opus | ~15K |
|
|
181
|
+
| Round 3 eval | 3 | sonnet ×3 | ~45K |
|
|
182
|
+
| **Total** | | | **~131K tokens** |
|
|
183
|
+
|
|
184
|
+
## Anti-Patterns
|
|
185
|
+
|
|
186
|
+
| Anti-Pattern | Detection | Fix |
|
|
187
|
+
|-------------|-----------|-----|
|
|
188
|
+
| Overfitting to one metric | One dimension +3, others flat or down | Balance mutations across dimensions |
|
|
189
|
+
| Losing generality | Specialized variant breaks other use cases | Test with multiple example inputs |
|
|
190
|
+
| Infinite loop | > 5 rounds without improvement | Hard cap at configured rounds |
|
|
191
|
+
| Semantic drift | Optimized version changes intent | Compare purpose statement before/after |
|
|
192
|
+
| Premature optimization | Baseline ≥ 8.0 | Skip optimization, suggest minor tweaks |
|
|
193
|
+
| Score inflation | Haiku gives high scores to everything | Calibrate with sonnet baseline |
|
|
194
|
+
|
|
195
|
+
## Abort Conditions
|
|
196
|
+
|
|
197
|
+
Stop optimization immediately if:
|
|
198
|
+
1. Any round shows overall regression > 0.5 from baseline
|
|
199
|
+
2. Critical structural checks (Layer 0) fail on any variant
|
|
200
|
+
3. Artifact semantics change fundamentally
|
|
201
|
+
4. User requests stop
|
|
@@ -0,0 +1,348 @@
|
|
|
1
|
+
# TEST Module — Multi-Agent Evaluation Protocol
|
|
2
|
+
|
|
3
|
+
## Purpose
|
|
4
|
+
|
|
5
|
+
Evaluate any Claude Code artifact (skill, command, rule, agent template) using layered evaluation: deterministic pre-checks, single-judge quick eval, and full 3-judge panel.
|
|
6
|
+
|
|
7
|
+
## Input
|
|
8
|
+
|
|
9
|
+
- **Path:** Path to artifact file or directory
|
|
10
|
+
- **Level:** layer0 | layer1 | layer2 | full (default: full)
|
|
11
|
+
- **Artifact type:** auto-detected from path/content
|
|
12
|
+
|
|
13
|
+
## Type Detection
|
|
14
|
+
|
|
15
|
+
| Path Pattern | Detected Type |
|
|
16
|
+
|-------------|--------------|
|
|
17
|
+
| `.claude/skills/*/SKILL.md` | skill |
|
|
18
|
+
| `.claude/skills/*/` (directory) | skill |
|
|
19
|
+
| `.claude/commands/*.md` | command |
|
|
20
|
+
| `.claude/rules/*.md` | rule |
|
|
21
|
+
| `.claude/agents/*.md` | agent |
|
|
22
|
+
| `researches/**/*.md` | research artifact |
|
|
23
|
+
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## Layer 0: Deterministic Pre-checks
|
|
27
|
+
|
|
28
|
+
**Cost:** Zero (no LLM calls)
|
|
29
|
+
**Speed:** Instant
|
|
30
|
+
**Purpose:** Catch structural issues before expensive LLM evaluation
|
|
31
|
+
|
|
32
|
+
### Universal Checks (all types)
|
|
33
|
+
|
|
34
|
+
```
|
|
35
|
+
CHECK-U1: File exists and is non-empty
|
|
36
|
+
CHECK-U2: File is valid UTF-8 text
|
|
37
|
+
CHECK-U3: Has at least one markdown heading (#)
|
|
38
|
+
CHECK-U4: No consecutive empty lines (> 2)
|
|
39
|
+
CHECK-U5: File size within bounds (see per-type limits)
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
### Skill Checks
|
|
43
|
+
|
|
44
|
+
```
|
|
45
|
+
CHECK-S1: SKILL.md exists in skill directory
|
|
46
|
+
CHECK-S2: Has "# Title" as first heading
|
|
47
|
+
CHECK-S3: Has "## Overview" or "## Purpose" section
|
|
48
|
+
CHECK-S4: Has "## Anti-Patterns" section
|
|
49
|
+
CHECK-S5: All files in modules/ referenced in SKILL.md
|
|
50
|
+
CHECK-S6: All files in references/ referenced in SKILL.md
|
|
51
|
+
CHECK-S7: No empty sections (heading → next heading with no content)
|
|
52
|
+
CHECK-S8: Size: 1KB < SKILL.md < 50KB
|
|
53
|
+
CHECK-S9: Total directory size < 200KB
|
|
54
|
+
CHECK-S10: At least one file in references/ OR examples/
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
### Command Checks
|
|
58
|
+
|
|
59
|
+
```
|
|
60
|
+
CHECK-C1: Contains "$ARGUMENTS" or parameter reference
|
|
61
|
+
CHECK-C2: Has checkpoint banner or protocol
|
|
62
|
+
CHECK-C3: Has skill loading instruction (Read *.SKILL.md)
|
|
63
|
+
CHECK-C4: Size: 500B < file < 20KB
|
|
64
|
+
CHECK-C5: Has "## Usage" or "## Protocol" section
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
### Rule Checks
|
|
68
|
+
|
|
69
|
+
```
|
|
70
|
+
CHECK-R1: Has table or structured list of patterns
|
|
71
|
+
CHECK-R2: Each pattern has detection signal and fix
|
|
72
|
+
CHECK-R3: Size: 200B < file < 10KB
|
|
73
|
+
CHECK-R4: Has "Auto-Detection" or similar section
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
### Agent Template Checks
|
|
77
|
+
|
|
78
|
+
```
|
|
79
|
+
CHECK-A1: Specifies model (haiku/sonnet/opus)
|
|
80
|
+
CHECK-A2: Specifies isolation scope
|
|
81
|
+
CHECK-A3: Has prompt template or instructions
|
|
82
|
+
CHECK-A4: Size: 200B < file < 10KB
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
### Layer 0 Scoring
|
|
86
|
+
|
|
87
|
+
- Each check: PASS (1) or FAIL (0)
|
|
88
|
+
- Score = passed / total
|
|
89
|
+
- **Gate: score ≥ 0.80 to proceed to Layer 1+**
|
|
90
|
+
- If score < 0.80: return report with specific failures, skip LLM evaluation
|
|
91
|
+
|
|
92
|
+
### Layer 0 Output
|
|
93
|
+
|
|
94
|
+
```
|
|
95
|
+
═══════════════════════════════════════════════════════
|
|
96
|
+
📋 LAYER 0: Deterministic Pre-checks
|
|
97
|
+
Artifact: <path>
|
|
98
|
+
Type: <detected type>
|
|
99
|
+
|
|
100
|
+
Results: X/Y passed (Z%)
|
|
101
|
+
|
|
102
|
+
✅ CHECK-S1: SKILL.md exists
|
|
103
|
+
✅ CHECK-S2: Has title heading
|
|
104
|
+
❌ CHECK-S7: Empty section found at line 45
|
|
105
|
+
...
|
|
106
|
+
|
|
107
|
+
Gate: PASS ✅ / FAIL ❌
|
|
108
|
+
═══════════════════════════════════════════════════════
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
---
|
|
112
|
+
|
|
113
|
+
## Layer 1: Single LLM Judge (Quick)
|
|
114
|
+
|
|
115
|
+
**Cost:** Low (1 haiku call)
|
|
116
|
+
**Speed:** ~10 seconds
|
|
117
|
+
**Purpose:** Fast quality signal for iteration or batch evaluation
|
|
118
|
+
|
|
119
|
+
### Model Selection
|
|
120
|
+
|
|
121
|
+
- Default: **haiku** (cost-optimized)
|
|
122
|
+
- For critical artifacts: **sonnet**
|
|
123
|
+
|
|
124
|
+
### Judge Prompt
|
|
125
|
+
|
|
126
|
+
```
|
|
127
|
+
You are evaluating a Claude Code {artifact_type}.
|
|
128
|
+
|
|
129
|
+
## Artifact Content
|
|
130
|
+
{content}
|
|
131
|
+
|
|
132
|
+
## Evaluation Dimensions
|
|
133
|
+
|
|
134
|
+
Rate each dimension 1-10:
|
|
135
|
+
|
|
136
|
+
1. CLARITY (Are instructions unambiguous? Can an LLM follow them precisely?)
|
|
137
|
+
2. COMPLETENESS (Are all necessary sections present? No missing pieces?)
|
|
138
|
+
3. ACTIONABILITY (Can Claude produce concrete output from these instructions?)
|
|
139
|
+
4. QUALITY (Well-structured? Professional? Good formatting?)
|
|
140
|
+
5. ANTI-PATTERNS (Avoids known pitfalls? Has failure mode coverage?)
|
|
141
|
+
|
|
142
|
+
## Required Output Format
|
|
143
|
+
|
|
144
|
+
SCORES:
|
|
145
|
+
- CLARITY: X/10 — [one-line justification]
|
|
146
|
+
- COMPLETENESS: X/10 — [one-line justification]
|
|
147
|
+
- ACTIONABILITY: X/10 — [one-line justification]
|
|
148
|
+
- QUALITY: X/10 — [one-line justification]
|
|
149
|
+
- ANTI-PATTERNS: X/10 — [one-line justification]
|
|
150
|
+
|
|
151
|
+
AVERAGE: X.X/10
|
|
152
|
+
|
|
153
|
+
TOP 3 IMPROVEMENTS:
|
|
154
|
+
1. [specific, actionable suggestion]
|
|
155
|
+
2. [specific, actionable suggestion]
|
|
156
|
+
3. [specific, actionable suggestion]
|
|
157
|
+
|
|
158
|
+
VERDICT: PASS (≥7.0) / NEEDS WORK (5.0-6.9) / FAIL (<5.0)
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
### Layer 1 Gate
|
|
162
|
+
|
|
163
|
+
- Average ≥ 7.0: PASS
|
|
164
|
+
- Average 5.0-6.9: NEEDS WORK (can proceed to Layer 2 for detailed feedback)
|
|
165
|
+
- Average < 5.0: FAIL (fix before proceeding)
|
|
166
|
+
|
|
167
|
+
---
|
|
168
|
+
|
|
169
|
+
## Layer 2: Full Judge Panel (3 Agents)
|
|
170
|
+
|
|
171
|
+
**Cost:** Moderate (3 sonnet calls)
|
|
172
|
+
**Speed:** ~30 seconds (parallel)
|
|
173
|
+
**Purpose:** Comprehensive, multi-perspective evaluation
|
|
174
|
+
|
|
175
|
+
### Architecture
|
|
176
|
+
|
|
177
|
+
Spawn 3 agents in parallel using Agent tool:
|
|
178
|
+
|
|
179
|
+
```
|
|
180
|
+
Agent 1: "BTO Judge — Domain Expert" model: sonnet
|
|
181
|
+
Agent 2: "BTO Judge — Critic" model: sonnet
|
|
182
|
+
Agent 3: "BTO Judge — Completeness" model: sonnet
|
|
183
|
+
```
|
|
184
|
+
|
|
185
|
+
**Isolation:** Each agent reads the same artifact independently. No cross-communication.
|
|
186
|
+
|
|
187
|
+
### Judge 1: Domain Expert
|
|
188
|
+
|
|
189
|
+
Focus: Is the content technically sound and appropriate for the domain?
|
|
190
|
+
|
|
191
|
+
```
|
|
192
|
+
You are a Domain Expert evaluator for Claude Code artifacts.
|
|
193
|
+
|
|
194
|
+
Evaluate this {artifact_type} for technical quality:
|
|
195
|
+
|
|
196
|
+
{content}
|
|
197
|
+
|
|
198
|
+
Score 1-10 on each dimension:
|
|
199
|
+
1. METHODOLOGY — Is the approach well-designed? Sound structure?
|
|
200
|
+
2. DEPTH — Sufficient detail for the task? Not too shallow?
|
|
201
|
+
3. CORRECTNESS — Are all claims and instructions valid?
|
|
202
|
+
4. USABILITY — Can a user/agent effectively use this?
|
|
203
|
+
5. ROBUSTNESS — Handles edge cases and failure modes?
|
|
204
|
+
|
|
205
|
+
For each dimension provide:
|
|
206
|
+
- Score (1-10)
|
|
207
|
+
- 2-3 sentence justification
|
|
208
|
+
- One specific improvement suggestion
|
|
209
|
+
|
|
210
|
+
OVERALL_SCORE: weighted average
|
|
211
|
+
CONFIDENCE: HIGH/MEDIUM/LOW
|
|
212
|
+
KEY_STRENGTHS: [top 2]
|
|
213
|
+
KEY_WEAKNESSES: [top 2]
|
|
214
|
+
```
|
|
215
|
+
|
|
216
|
+
### Judge 2: Critic
|
|
217
|
+
|
|
218
|
+
Focus: Find weaknesses, gaps, and anti-patterns.
|
|
219
|
+
|
|
220
|
+
```
|
|
221
|
+
You are a Critical Evaluator. Your job is to find problems.
|
|
222
|
+
|
|
223
|
+
Evaluate this {artifact_type} adversarially:
|
|
224
|
+
|
|
225
|
+
{content}
|
|
226
|
+
|
|
227
|
+
Score 1-10 on each dimension (be strict — average should be ~5-6):
|
|
228
|
+
1. METHODOLOGY — Any logical flaws or unjustified assumptions?
|
|
229
|
+
2. DEPTH — What's missing? What's under-specified?
|
|
230
|
+
3. CORRECTNESS — Any instructions that could mislead or produce wrong output?
|
|
231
|
+
4. USABILITY — What would confuse a user? Where would someone get stuck?
|
|
232
|
+
5. ROBUSTNESS — What breaks this? What edge cases are unhandled?
|
|
233
|
+
|
|
234
|
+
For each dimension provide:
|
|
235
|
+
- Score (1-10) — err on the side of strict
|
|
236
|
+
- 2-3 sentence justification focusing on PROBLEMS
|
|
237
|
+
- One specific failure scenario
|
|
238
|
+
|
|
239
|
+
OVERALL_SCORE: weighted average
|
|
240
|
+
CRITICAL_ISSUES: [list of blocking problems]
|
|
241
|
+
IMPROVEMENT_PRIORITY: [ordered list of what to fix first]
|
|
242
|
+
```
|
|
243
|
+
|
|
244
|
+
### Judge 3: Completeness Auditor
|
|
245
|
+
|
|
246
|
+
Focus: Structural completeness and cross-reference integrity.
|
|
247
|
+
|
|
248
|
+
```
|
|
249
|
+
You are a Completeness Auditor for Claude Code artifacts.
|
|
250
|
+
|
|
251
|
+
Audit this {artifact_type} for structural completeness:
|
|
252
|
+
|
|
253
|
+
{content}
|
|
254
|
+
|
|
255
|
+
Score 1-10 on each dimension:
|
|
256
|
+
1. METHODOLOGY — Does the structure follow established patterns?
|
|
257
|
+
2. DEPTH — Is every section adequately populated?
|
|
258
|
+
3. CORRECTNESS — Do all cross-references resolve? Are all claims supported?
|
|
259
|
+
4. USABILITY — Is the information well-organized and findable?
|
|
260
|
+
5. ROBUSTNESS — Are failure modes documented? Are anti-patterns listed?
|
|
261
|
+
|
|
262
|
+
For each dimension provide:
|
|
263
|
+
- Score (1-10)
|
|
264
|
+
- 2-3 sentence justification
|
|
265
|
+
- Missing element (if any)
|
|
266
|
+
|
|
267
|
+
OVERALL_SCORE: weighted average
|
|
268
|
+
MISSING_SECTIONS: [list]
|
|
269
|
+
BROKEN_REFERENCES: [list]
|
|
270
|
+
COVERAGE_GAPS: [list]
|
|
271
|
+
```
|
|
272
|
+
|
|
273
|
+
### Score Aggregation
|
|
274
|
+
|
|
275
|
+
After all 3 judges return:
|
|
276
|
+
|
|
277
|
+
**Weights:**
|
|
278
|
+
- Domain Expert: 0.40
|
|
279
|
+
- Critic: 0.30
|
|
280
|
+
- Completeness Auditor: 0.30
|
|
281
|
+
|
|
282
|
+
**Per-dimension:**
|
|
283
|
+
```
|
|
284
|
+
dimension_score = expert[dim] * 0.4 + critic[dim] * 0.3 + auditor[dim] * 0.3
|
|
285
|
+
```
|
|
286
|
+
|
|
287
|
+
**Overall:**
|
|
288
|
+
```
|
|
289
|
+
overall = mean(all dimension_scores)
|
|
290
|
+
```
|
|
291
|
+
|
|
292
|
+
**Disagreement Detection:**
|
|
293
|
+
For each dimension, if `max(scores) - min(scores) > 3` → FLAG for meta-judge.
|
|
294
|
+
|
|
295
|
+
---
|
|
296
|
+
|
|
297
|
+
## Meta-Judge: Disagreement Resolution
|
|
298
|
+
|
|
299
|
+
**Trigger:** Any flagged dimension from Layer 2.
|
|
300
|
+
|
|
301
|
+
**Model:** default (opus)
|
|
302
|
+
|
|
303
|
+
**Prompt:**
|
|
304
|
+
```
|
|
305
|
+
Three judges evaluated a Claude Code artifact and disagreed significantly
|
|
306
|
+
on the following dimension(s): {flagged_dimensions}
|
|
307
|
+
|
|
308
|
+
Judge 1 (Domain Expert): {score} — {justification}
|
|
309
|
+
Judge 2 (Critic): {score} — {justification}
|
|
310
|
+
Judge 3 (Completeness): {score} — {justification}
|
|
311
|
+
|
|
312
|
+
Reconcile these scores. Provide:
|
|
313
|
+
1. Your reconciled score (1-10) with reasoning
|
|
314
|
+
2. Which judge(s) had the most valid perspective and why
|
|
315
|
+
3. Whether this needs human review (YES/NO)
|
|
316
|
+
```
|
|
317
|
+
|
|
318
|
+
---
|
|
319
|
+
|
|
320
|
+
## Output: Evaluation Report
|
|
321
|
+
|
|
322
|
+
Format defined in `examples/sample-eval-report.md`.
|
|
323
|
+
|
|
324
|
+
Summary structure:
|
|
325
|
+
```
|
|
326
|
+
═══════════════════════════════════════════════════════
|
|
327
|
+
📊 BTO EVALUATION REPORT
|
|
328
|
+
Artifact: <path>
|
|
329
|
+
Type: <type>
|
|
330
|
+
Level: Layer 0 + Layer 1 + Layer 2
|
|
331
|
+
|
|
332
|
+
OVERALL SCORE: X.X / 10 [PASS / NEEDS WORK / FAIL]
|
|
333
|
+
|
|
334
|
+
Per-Dimension:
|
|
335
|
+
METHODOLOGY: X.X ██████████░░
|
|
336
|
+
DEPTH: X.X ████████░░░░
|
|
337
|
+
CORRECTNESS: X.X █████████░░░
|
|
338
|
+
USABILITY: X.X ██████████░░
|
|
339
|
+
ROBUSTNESS: X.X ███████░░░░░
|
|
340
|
+
|
|
341
|
+
Flagged: [dimensions with >3 disagreement]
|
|
342
|
+
|
|
343
|
+
Top Improvements:
|
|
344
|
+
1. ...
|
|
345
|
+
2. ...
|
|
346
|
+
3. ...
|
|
347
|
+
═══════════════════════════════════════════════════════
|
|
348
|
+
```
|