@dzhechkov/skills-bto 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bin/cli.js +5 -0
- package/package.json +43 -0
- package/src/cli.js +150 -0
- package/src/commands/doctor.js +366 -0
- package/src/commands/init.js +188 -0
- package/src/commands/list.js +161 -0
- package/src/commands/remove.js +211 -0
- package/src/commands/update.js +198 -0
- package/src/utils.js +398 -0
- package/templates/.claude/agents/bto-judge-panel.md +192 -0
- package/templates/.claude/agents/bto-optimizer-worker.md +181 -0
- package/templates/.claude/commands/bto-build.md +169 -0
- package/templates/.claude/commands/bto-optimize.md +208 -0
- package/templates/.claude/commands/bto-test.md +186 -0
- package/templates/.claude/commands/bto.md +171 -0
- package/templates/.claude/rules/bto-quality-gates.md +91 -0
- package/templates/.claude/skills/bto/SKILL.md +266 -0
- package/templates/.claude/skills/bto/examples/sample-eval-report.md +436 -0
- package/templates/.claude/skills/bto/modules/build.md +189 -0
- package/templates/.claude/skills/bto/modules/optimize.md +201 -0
- package/templates/.claude/skills/bto/modules/test.md +348 -0
- package/templates/.claude/skills/bto/references/eval-patterns.md +202 -0
- package/templates/.claude/skills/bto/references/judge-rubrics.md +183 -0
- package/templates/.claude/skills/bto/references/optimization-methods.md +139 -0
- package/templates/.claude/skills/bto/references/quality-checklist.md +220 -0
|
@@ -0,0 +1,436 @@
|
|
|
1
|
+
# Sample BTO Evaluation Report
|
|
2
|
+
|
|
3
|
+
> This file is a **reference example** showing the complete output format of a BTO TEST evaluation.
|
|
4
|
+
> Use it as the authoritative template when formatting evaluation reports.
|
|
5
|
+
>
|
|
6
|
+
> **Hypothetical subject:** `.claude/skills/data-pipeline-validator/` — a skill for validating ETL pipeline configurations.
|
|
7
|
+
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
## Report Header
|
|
11
|
+
|
|
12
|
+
```
|
|
13
|
+
╔══════════════════════════════════════════════════════════════╗
|
|
14
|
+
║ BTO EVALUATION REPORT — TEST MODULE ║
|
|
15
|
+
╠══════════════════════════════════════════════════════════════╣
|
|
16
|
+
║ Artifact: data-pipeline-validator (Skill) ║
|
|
17
|
+
║ Path: .claude/skills/data-pipeline-validator/ ║
|
|
18
|
+
║ Evaluated: 2026-03-01 14:32 UTC ║
|
|
19
|
+
║ Evaluator: BTO TEST v1.2 ║
|
|
20
|
+
║ Layers run: Layer 0 + Layer 1 + Layer 2 ║
|
|
21
|
+
╚══════════════════════════════════════════════════════════════╝
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## Layer 0: Deterministic Pre-Checks
|
|
27
|
+
|
|
28
|
+
**Artifact type:** Skill
|
|
29
|
+
**Applicable checks:** Universal (12) + Skill-specific (16) = 28 total
|
|
30
|
+
|
|
31
|
+
### Universal Checks (12 applicable)
|
|
32
|
+
|
|
33
|
+
| ID | Check | Result | Note |
|
|
34
|
+
|----|-------|--------|------|
|
|
35
|
+
| U-01 | File exists and is non-empty | PASS | SKILL.md: 8.4 KB |
|
|
36
|
+
| U-02 | UTF-8 encoding | PASS | |
|
|
37
|
+
| U-03 | Starts with level-1 heading | PASS | `# Data Pipeline Validator` |
|
|
38
|
+
| U-04 | No placeholder text | FAIL | Found `[INSERT SCHEMA PATH HERE]` in modules/validate.md |
|
|
39
|
+
| U-05 | No empty sections | PASS | |
|
|
40
|
+
| U-06 | Consistent heading hierarchy | PASS | |
|
|
41
|
+
| U-07 | No broken internal cross-references | FAIL | `references/schema-types.md` referenced but missing |
|
|
42
|
+
| U-08 | File size within bounds | PASS | 8.4 KB (within 200B–100KB) |
|
|
43
|
+
| U-09 | No trailing whitespace | PASS | Auto-fixed during generation |
|
|
44
|
+
| U-10 | Standard Markdown only | PASS | |
|
|
45
|
+
| U-11 | Code blocks properly closed | PASS | |
|
|
46
|
+
| U-12 | No duplicate top-level sections | PASS | |
|
|
47
|
+
|
|
48
|
+
**Universal result: 10 / 12**
|
|
49
|
+
|
|
50
|
+
### Skill-Specific Checks (16 applicable)
|
|
51
|
+
|
|
52
|
+
| ID | Check | Result | Note |
|
|
53
|
+
|----|-------|--------|------|
|
|
54
|
+
| SK-01 | SKILL.md exists at skill root | PASS | |
|
|
55
|
+
| SK-02 | Has `## Overview` section | PASS | |
|
|
56
|
+
| SK-03 | Has `## Anti-Patterns` section | PASS | |
|
|
57
|
+
| SK-04 | Has Quick Start or usage example | PASS | Section present with 2 examples |
|
|
58
|
+
| SK-05 | `modules/` directory exists | PASS | |
|
|
59
|
+
| SK-06 | All modules referenced exist on disk | PASS | 3 modules, all present |
|
|
60
|
+
| SK-07 | `references/` directory exists | PASS | |
|
|
61
|
+
| SK-08 | All references referenced exist on disk | FAIL | `schema-types.md` missing (linked at line 47) |
|
|
62
|
+
| SK-09 | `examples/` directory exists | PASS | |
|
|
63
|
+
| SK-10 | At least one example file present | FAIL | Directory exists but is empty |
|
|
64
|
+
| SK-11 | Skill name matches directory name | PASS | |
|
|
65
|
+
| SK-12 | Has `## Dependencies` section | PASS | |
|
|
66
|
+
| SK-13 | SKILL.md size within skill bounds | PASS | 8.4 KB |
|
|
67
|
+
| SK-14 | Total directory size within bounds | PASS | 22 KB total |
|
|
68
|
+
| SK-15 | Each module has `# Title` heading | PASS | |
|
|
69
|
+
| SK-16 | No circular cross-references | PASS | |
|
|
70
|
+
|
|
71
|
+
**Skill-specific result: 13 / 16**
|
|
72
|
+
|
|
73
|
+
### Layer 0 Summary
|
|
74
|
+
|
|
75
|
+
```
|
|
76
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
77
|
+
LAYER 0 PRE-FLIGHT CHECK
|
|
78
|
+
Artifact: data-pipeline-validator (Skill)
|
|
79
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
80
|
+
Universal checks: 10 / 12
|
|
81
|
+
Skill-specific checks: 13 / 16
|
|
82
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
83
|
+
TOTAL: 23 / 28
|
|
84
|
+
Pass rate: 82.1%
|
|
85
|
+
Status: PASS (threshold: 80%)
|
|
86
|
+
|
|
87
|
+
Failed checks:
|
|
88
|
+
- [U-04] Placeholder text in modules/validate.md line 23 *
|
|
89
|
+
- [U-07] Missing file: references/schema-types.md
|
|
90
|
+
- [SK-08] Broken reference: schema-types.md (line 47 in SKILL.md)
|
|
91
|
+
- [SK-10] Empty examples/ directory
|
|
92
|
+
|
|
93
|
+
(* = auto-fixable)
|
|
94
|
+
|
|
95
|
+
Verdict: PROCEED TO LAYER 1 (with noted issues logged)
|
|
96
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
**Decision:** Layer 0 PASS (82.1% > 80%). Flagged issues logged — will be reflected in Layer 1/2 scoring under CORRECTNESS and USABILITY dimensions.
|
|
100
|
+
|
|
101
|
+
---
|
|
102
|
+
|
|
103
|
+
## Layer 1: Single LLM Judge (Quick Mode)
|
|
104
|
+
|
|
105
|
+
**Model:** claude-haiku (cost-optimized)
|
|
106
|
+
**Purpose:** Rapid spot-check before committing to full 3-judge panel
|
|
107
|
+
|
|
108
|
+
**Judge Prompt Used:**
|
|
109
|
+
```
|
|
110
|
+
You are a Claude Code artifact evaluator. Rate the skill at the provided path
|
|
111
|
+
on 5 dimensions (1-10 each). Provide a score, one-sentence justification,
|
|
112
|
+
and top 3 improvement suggestions. Artifact type: Skill.
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
### Layer 1 Scores
|
|
116
|
+
|
|
117
|
+
| Dimension | Score | Justification |
|
|
118
|
+
|-----------|-------|---------------|
|
|
119
|
+
| CLARITY | 7 | Instructions are mostly clear but the validation protocol section has ambiguous branching logic at step 4 |
|
|
120
|
+
| COMPLETENESS | 6 | Missing example file and one broken reference significantly reduce perceived completeness |
|
|
121
|
+
| ACTIONABILITY | 7 | Core happy path is followable; failure paths are underspecified |
|
|
122
|
+
| QUALITY | 7 | Well-structured with good use of tables; prose in modules could be tighter |
|
|
123
|
+
| ANTI-PATTERNS | 8 | Anti-patterns section is thorough with 6 patterns, each with signals and fixes |
|
|
124
|
+
|
|
125
|
+
**Layer 1 Average: 7.0**
|
|
126
|
+
**Threshold: 7.0**
|
|
127
|
+
**Status: BORDERLINE PASS — proceed to Layer 2 for definitive evaluation**
|
|
128
|
+
|
|
129
|
+
**Layer 1 Top 3 Suggestions:**
|
|
130
|
+
1. Add at least one complete worked example showing a real ETL schema being validated end-to-end
|
|
131
|
+
2. Fix the broken `schema-types.md` reference or remove the link
|
|
132
|
+
3. Clarify the branching logic in step 4 of the validation protocol (what happens if schema is partially valid?)
|
|
133
|
+
|
|
134
|
+
---
|
|
135
|
+
|
|
136
|
+
## Layer 2: Full Judge Panel (3 Agents)
|
|
137
|
+
|
|
138
|
+
**Models:** claude-sonnet (all 3 judges)
|
|
139
|
+
**Rubric source:** `references/judge-rubrics.md`
|
|
140
|
+
**Execution:** Parallel (all 3 judges ran simultaneously)
|
|
141
|
+
**Aggregation:** Weighted average — Expert 0.4 / Critic 0.3 / Auditor 0.3
|
|
142
|
+
|
|
143
|
+
---
|
|
144
|
+
|
|
145
|
+
### Judge 1: Domain Expert
|
|
146
|
+
|
|
147
|
+
**Role:** Evaluates accuracy, depth, methodology, and domain fit.
|
|
148
|
+
**Calibration:** Balanced — scores represent genuine quality assessment.
|
|
149
|
+
|
|
150
|
+
| Dimension | Score | Expert Commentary |
|
|
151
|
+
|-----------|-------|-------------------|
|
|
152
|
+
| METHODOLOGY | 8 | The validation protocol follows a logical sequence: schema → types → constraints → relationships. The hierarchical validation approach is sound for ETL contexts. |
|
|
153
|
+
| DEPTH | 6 | Covers the primary validation scenarios (schema drift, type mismatches, null handling) but misses temporal consistency checks and incremental load edge cases — critical for production pipelines. |
|
|
154
|
+
| CORRECTNESS | 7 | All referenced validation patterns are real and applicable. The one broken reference (`schema-types.md`) is a quality concern but doesn't indicate incorrect claims. |
|
|
155
|
+
| USABILITY | 7 | Quick Start is useful. Navigation between SKILL.md and modules is smooth. The empty examples directory is a significant usability gap — users need a worked example. |
|
|
156
|
+
| ROBUSTNESS | 8 | Anti-patterns section is strong. Handles schema evolution, backwards compatibility, and partial validation failure. Missing: network timeout handling for remote schema fetches. |
|
|
157
|
+
|
|
158
|
+
**Expert subtotal: (8+6+7+7+8)/5 = 7.2**
|
|
159
|
+
|
|
160
|
+
---
|
|
161
|
+
|
|
162
|
+
### Judge 2: Critic
|
|
163
|
+
|
|
164
|
+
**Role:** Identifies gaps, weaknesses, anti-patterns, and edge cases.
|
|
165
|
+
**Calibration:** Strict — intentionally scores conservative to surface issues.
|
|
166
|
+
|
|
167
|
+
| Dimension | Score | Critic Commentary |
|
|
168
|
+
|-----------|-------|-------------------|
|
|
169
|
+
| METHODOLOGY | 6 | The protocol works for simple schemas but collapses under nested/recursive schema structures. Step 4 (constraint validation) assumes a flat schema model. No guidance on when to abort vs. continue validation after first failure. |
|
|
170
|
+
| DEPTH | 5 | No coverage of: streaming data validation, schema registry integration (Confluent, AWS Glue), or multi-source join validation. The "data types" coverage only addresses SQL types — no Avro, Parquet, or Arrow type systems. |
|
|
171
|
+
| CORRECTNESS | 6 | The placeholder `[INSERT SCHEMA PATH HERE]` in modules/validate.md (line 23) is a clear correctness failure — the skill was shipped incomplete. The claim about "99.9% detection rate" on line 31 has no source. |
|
|
172
|
+
| USABILITY | 5 | Zero examples is a dealbreaker for a validation skill. Users cannot learn by example. The broken reference creates a dead end. Anti-patterns section, while good, is buried at the end — should appear earlier as a pre-check. |
|
|
173
|
+
| ROBUSTNESS | 6 | Handles the happy path and common failures. No coverage of: what to do if the target system is unavailable, concurrent validation conflicts, or validation against evolving schemas mid-run. |
|
|
174
|
+
|
|
175
|
+
**Critic subtotal: (6+5+6+5+6)/5 = 5.6**
|
|
176
|
+
|
|
177
|
+
---
|
|
178
|
+
|
|
179
|
+
### Judge 3: Completeness Auditor
|
|
180
|
+
|
|
181
|
+
**Role:** Checks structure, coverage, cross-references, and actionability.
|
|
182
|
+
**Calibration:** Neutral — focuses on structural completeness, not quality of content.
|
|
183
|
+
|
|
184
|
+
| Dimension | Score | Auditor Commentary |
|
|
185
|
+
|-----------|-------|-------------------|
|
|
186
|
+
| METHODOLOGY | 7 | All required sections present in SKILL.md. Module structure follows conventions. Decision flow between modules is documented. Protocol steps are numbered and sequential. |
|
|
187
|
+
| DEPTH | 6 | 3 modules present: validate.md, report.md, fix-suggestions.md — reasonable coverage. However, no module for schema-registry integration or configuration management, which the skill description implies are supported. |
|
|
188
|
+
| CORRECTNESS | 6 | 2 structural correctness failures found: (1) broken reference to schema-types.md, (2) placeholder text in validate.md. These are Layer 0 failures that reached Layer 2 — the skill was not fully completed before evaluation. |
|
|
189
|
+
| USABILITY | 7 | Navigation is clear. Table of contents in SKILL.md maps correctly to modules. Quick Start section provides 2 invocation examples. Dependencies section complete. Loses points for missing examples. |
|
|
190
|
+
| ROBUSTNESS | 7 | Anti-patterns section documents 6 patterns. Failure modes section in validate.md covers 4 scenarios. Abort conditions defined. Minor gap: no documentation of what constitutes a "critical" vs. "warning" validation finding. |
|
|
191
|
+
|
|
192
|
+
**Auditor subtotal: (7+6+6+7+7)/5 = 6.6**
|
|
193
|
+
|
|
194
|
+
---
|
|
195
|
+
|
|
196
|
+
### Aggregated Scores (Weighted)
|
|
197
|
+
|
|
198
|
+
**Formula:** Expert (×0.4) + Critic (×0.3) + Auditor (×0.3)
|
|
199
|
+
|
|
200
|
+
| Dimension | Expert | Critic | Auditor | Weighted Score | Bar |
|
|
201
|
+
|-----------|--------|--------|---------|----------------|-----|
|
|
202
|
+
| METHODOLOGY | 8 | 6 | 7 | 7.1 | `███████░░░` |
|
|
203
|
+
| DEPTH | 6 | 5 | 6 | 5.7 | `█████▒░░░░` |
|
|
204
|
+
| CORRECTNESS | 7 | 6 | 6 | 6.4 | `██████░░░░` |
|
|
205
|
+
| USABILITY | 7 | 5 | 7 | 6.4 | `██████░░░░` |
|
|
206
|
+
| ROBUSTNESS | 8 | 6 | 7 | 7.1 | `███████░░░` |
|
|
207
|
+
| **OVERALL** | **7.2** | **5.6** | **6.6** | **6.5** | `██████▒░░░` |
|
|
208
|
+
|
|
209
|
+
**Bar legend:** `█` = full point, `▒` = half point, `░` = empty
|
|
210
|
+
|
|
211
|
+
**Overall weighted score: 6.5 / 10**
|
|
212
|
+
**Pass threshold: 7.0**
|
|
213
|
+
**Status: BELOW THRESHOLD**
|
|
214
|
+
|
|
215
|
+
---
|
|
216
|
+
|
|
217
|
+
### Disagreement Analysis
|
|
218
|
+
|
|
219
|
+
> Flag trigger: any dimension where `max - min > 3` across judges.
|
|
220
|
+
|
|
221
|
+
| Dimension | Max | Min | Range | Disagreement Flag |
|
|
222
|
+
|-----------|-----|-----|-------|-------------------|
|
|
223
|
+
| METHODOLOGY | 8 | 6 | 2 | None |
|
|
224
|
+
| DEPTH | 6 | 5 | 1 | None |
|
|
225
|
+
| CORRECTNESS | 7 | 6 | 1 | None |
|
|
226
|
+
| USABILITY | 7 | 5 | **2** | None (under threshold of 3) |
|
|
227
|
+
| ROBUSTNESS | 8 | 6 | 2 | None |
|
|
228
|
+
|
|
229
|
+
**No disagreement flags triggered.** Maximum range across all dimensions was 2 (USABILITY), below the trigger threshold of 3. Meta-judge not required.
|
|
230
|
+
|
|
231
|
+
**Interpretation:** The judges are aligned in their assessment. The Critic's USABILITY score of 5 reflects the empty examples directory — a structural deficiency all judges identified, but weighed differently. The Expert and Auditor judged the existing structure more charitably; the Critic penalized the missing examples more severely. This spread is within normal calibration variance.
|
|
232
|
+
|
|
233
|
+
---
|
|
234
|
+
|
|
235
|
+
### Meta-Judge Status
|
|
236
|
+
|
|
237
|
+
**Triggered:** NO
|
|
238
|
+
**Reason:** All dimension ranges within tolerance (max range: 2, threshold: 3)
|
|
239
|
+
**Action:** Proceed with weighted aggregation as final score.
|
|
240
|
+
|
|
241
|
+
---
|
|
242
|
+
|
|
243
|
+
## Top Improvements
|
|
244
|
+
|
|
245
|
+
Priority ordered by weighted impact on overall score.
|
|
246
|
+
|
|
247
|
+
### Priority 1 — Add a worked example (Impact: HIGH)
|
|
248
|
+
**Failing checks:** SK-10, Layer 1 suggestion 1, Critic USABILITY: 5, Auditor USABILITY: 7
|
|
249
|
+
**What to fix:** Create `examples/validate-postgres-to-s3.md` showing a complete validation run of a realistic PostgreSQL-to-S3 ETL pipeline: input schema, validation output, error report, and fix suggestions.
|
|
250
|
+
**Expected score improvement:** USABILITY +1.5 to +2.0 points
|
|
251
|
+
|
|
252
|
+
### Priority 2 — Create missing `references/schema-types.md` (Impact: HIGH)
|
|
253
|
+
**Failing checks:** U-07, SK-08, Critic CORRECTNESS: 6
|
|
254
|
+
**What to fix:** Create the referenced file with a type mapping table covering SQL, Avro, Parquet, and Arrow type systems. Remove the broken link if out of scope.
|
|
255
|
+
**Expected score improvement:** CORRECTNESS +0.8, DEPTH +0.5
|
|
256
|
+
|
|
257
|
+
### Priority 3 — Remove placeholder text (Impact: MEDIUM)
|
|
258
|
+
**Failing checks:** U-04, Critic CORRECTNESS: 6
|
|
259
|
+
**What to fix:** Replace `[INSERT SCHEMA PATH HERE]` in `modules/validate.md` line 23 with a concrete example path or a proper `$ARGUMENTS` reference.
|
|
260
|
+
**Expected score improvement:** CORRECTNESS +0.3
|
|
261
|
+
|
|
262
|
+
### Priority 4 — Expand depth for non-SQL type systems (Impact: MEDIUM)
|
|
263
|
+
**Failing checks:** Critic DEPTH: 5, Expert DEPTH: 6
|
|
264
|
+
**What to fix:** Add a module section or reference covering Avro, Parquet, and Arrow schemas. Add a note on schema registry integration (AWS Glue Data Catalog, Confluent Schema Registry).
|
|
265
|
+
**Expected score improvement:** DEPTH +1.0 to +1.5
|
|
266
|
+
|
|
267
|
+
### Priority 5 — Clarify critical vs. warning severity levels (Impact: LOW)
|
|
268
|
+
**Failing checks:** Auditor ROBUSTNESS partial
|
|
269
|
+
**What to fix:** Add a severity matrix to the validation output format — define which validation failures are CRITICAL (block pipeline) vs. WARNING (log and continue).
|
|
270
|
+
**Expected score improvement:** ROBUSTNESS +0.3
|
|
271
|
+
|
|
272
|
+
---
|
|
273
|
+
|
|
274
|
+
## Overall Verdict
|
|
275
|
+
|
|
276
|
+
```
|
|
277
|
+
╔══════════════════════════════════════════════════════════════╗
|
|
278
|
+
║ FINAL VERDICT ║
|
|
279
|
+
╠══════════════════════════════════════════════════════════════╣
|
|
280
|
+
║ ║
|
|
281
|
+
║ Artifact: data-pipeline-validator (Skill) ║
|
|
282
|
+
║ Score: 6.5 / 10 ║
|
|
283
|
+
║ ║
|
|
284
|
+
║ Score breakdown: ║
|
|
285
|
+
║ METHODOLOGY ███████░░░ 7.1 ║
|
|
286
|
+
║ DEPTH █████▒░░░░ 5.7 ║
|
|
287
|
+
║ CORRECTNESS ██████░░░░ 6.4 ║
|
|
288
|
+
║ USABILITY ██████░░░░ 6.4 ║
|
|
289
|
+
║ ROBUSTNESS ███████░░░ 7.1 ║
|
|
290
|
+
║ ────────── ║
|
|
291
|
+
║ OVERALL ██████▒░░░ 6.5 ║
|
|
292
|
+
║ ║
|
|
293
|
+
║ Verdict: NEEDS IMPROVEMENT ║
|
|
294
|
+
║ Threshold: 7.0 (not met) ║
|
|
295
|
+
║ ║
|
|
296
|
+
║ Recommended action: ║
|
|
297
|
+
║ → Apply Priority 1-3 fixes (est. 30 min) ║
|
|
298
|
+
║ → Re-run /bto-test to verify improvement ║
|
|
299
|
+
║ → Do NOT run /bto-optimize until score ≥ 7.0 ║
|
|
300
|
+
║ ║
|
|
301
|
+
║ Estimated score after Priority 1-3 fixes: 7.5 – 8.0 ║
|
|
302
|
+
╚══════════════════════════════════════════════════════════════╝
|
|
303
|
+
```
|
|
304
|
+
|
|
305
|
+
---
|
|
306
|
+
|
|
307
|
+
## Appendix: Judge Prompts Used
|
|
308
|
+
|
|
309
|
+
### Judge 1 (Domain Expert) Prompt
|
|
310
|
+
|
|
311
|
+
```
|
|
312
|
+
You are a Domain Expert evaluating a Claude Code skill artifact.
|
|
313
|
+
Your role: assess accuracy, depth, methodology, and domain fit.
|
|
314
|
+
Score each dimension 1-10 using the rubric in judge-rubrics.md.
|
|
315
|
+
Be balanced — reward genuine quality, penalize real gaps.
|
|
316
|
+
|
|
317
|
+
Artifact type: Skill
|
|
318
|
+
Artifact path: .claude/skills/data-pipeline-validator/
|
|
319
|
+
|
|
320
|
+
Dimensions to score:
|
|
321
|
+
1. METHODOLOGY — Is the approach sound and well-structured?
|
|
322
|
+
2. DEPTH — Is the content thorough enough for production use?
|
|
323
|
+
3. CORRECTNESS — Are claims accurate and instructions valid?
|
|
324
|
+
4. USABILITY — Can a practitioner effectively use this artifact?
|
|
325
|
+
5. ROBUSTNESS — Does it handle edge cases and failure modes?
|
|
326
|
+
|
|
327
|
+
Output: JSON with keys: methodology, depth, correctness, usability,
|
|
328
|
+
robustness, commentary_per_dimension, top_3_issues
|
|
329
|
+
```
|
|
330
|
+
|
|
331
|
+
### Judge 2 (Critic) Prompt
|
|
332
|
+
|
|
333
|
+
```
|
|
334
|
+
You are a strict Critic evaluating a Claude Code skill artifact.
|
|
335
|
+
Your role: find gaps, weaknesses, anti-patterns, and missing edge cases.
|
|
336
|
+
Score calibration: your scores should average around 5-6 (strict).
|
|
337
|
+
Do not reward potential — only reward what is actually present.
|
|
338
|
+
|
|
339
|
+
Artifact type: Skill
|
|
340
|
+
Artifact path: .claude/skills/data-pipeline-validator/
|
|
341
|
+
|
|
342
|
+
Dimensions to score:
|
|
343
|
+
1. METHODOLOGY — Where does the approach break down?
|
|
344
|
+
2. DEPTH — What important topics are missing?
|
|
345
|
+
3. CORRECTNESS — What is wrong, incomplete, or unverified?
|
|
346
|
+
4. USABILITY — Where will users get stuck?
|
|
347
|
+
5. ROBUSTNESS — What failure modes are unhandled?
|
|
348
|
+
|
|
349
|
+
Output: JSON with keys: methodology, depth, correctness, usability,
|
|
350
|
+
robustness, commentary_per_dimension, top_3_issues
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
### Judge 3 (Completeness Auditor) Prompt
|
|
354
|
+
|
|
355
|
+
```
|
|
356
|
+
You are a Completeness Auditor evaluating a Claude Code skill artifact.
|
|
357
|
+
Your role: check structural completeness, cross-reference integrity,
|
|
358
|
+
section coverage, and actionability.
|
|
359
|
+
Be neutral — focus on what is and is not present, not on quality of content.
|
|
360
|
+
|
|
361
|
+
Artifact type: Skill
|
|
362
|
+
Artifact path: .claude/skills/data-pipeline-validator/
|
|
363
|
+
|
|
364
|
+
Dimensions to score:
|
|
365
|
+
1. METHODOLOGY — Are all required protocol elements present?
|
|
366
|
+
2. DEPTH — Does the number and depth of sections match the scope?
|
|
367
|
+
3. CORRECTNESS — Are all cross-references valid and content accurate?
|
|
368
|
+
4. USABILITY — Is navigation clear and information findable?
|
|
369
|
+
5. ROBUSTNESS — Are failure modes and edge cases documented?
|
|
370
|
+
|
|
371
|
+
Output: JSON with keys: methodology, depth, correctness, usability,
|
|
372
|
+
robustness, commentary_per_dimension, top_3_issues
|
|
373
|
+
```
|
|
374
|
+
|
|
375
|
+
---
|
|
376
|
+
|
|
377
|
+
## Appendix: Raw Judge Output (JSON)
|
|
378
|
+
|
|
379
|
+
```json
|
|
380
|
+
{
|
|
381
|
+
"evaluation_metadata": {
|
|
382
|
+
"artifact": "data-pipeline-validator",
|
|
383
|
+
"type": "skill",
|
|
384
|
+
"timestamp": "2026-03-01T14:32:00Z",
|
|
385
|
+
"bto_version": "1.2"
|
|
386
|
+
},
|
|
387
|
+
"layer_0": {
|
|
388
|
+
"total": 28,
|
|
389
|
+
"passed": 23,
|
|
390
|
+
"pass_rate": 0.821,
|
|
391
|
+
"status": "PASS",
|
|
392
|
+
"failed_ids": ["U-04", "U-07", "SK-08", "SK-10"]
|
|
393
|
+
},
|
|
394
|
+
"layer_1": {
|
|
395
|
+
"model": "claude-haiku",
|
|
396
|
+
"scores": {
|
|
397
|
+
"clarity": 7, "completeness": 6, "actionability": 7,
|
|
398
|
+
"quality": 7, "anti_patterns": 8
|
|
399
|
+
},
|
|
400
|
+
"average": 7.0,
|
|
401
|
+
"status": "BORDERLINE_PASS"
|
|
402
|
+
},
|
|
403
|
+
"layer_2": {
|
|
404
|
+
"judges": {
|
|
405
|
+
"expert": {
|
|
406
|
+
"model": "claude-sonnet",
|
|
407
|
+
"scores": { "methodology": 8, "depth": 6, "correctness": 7, "usability": 7, "robustness": 8 },
|
|
408
|
+
"subtotal": 7.2
|
|
409
|
+
},
|
|
410
|
+
"critic": {
|
|
411
|
+
"model": "claude-sonnet",
|
|
412
|
+
"scores": { "methodology": 6, "depth": 5, "correctness": 6, "usability": 5, "robustness": 6 },
|
|
413
|
+
"subtotal": 5.6
|
|
414
|
+
},
|
|
415
|
+
"auditor": {
|
|
416
|
+
"model": "claude-sonnet",
|
|
417
|
+
"scores": { "methodology": 7, "depth": 6, "correctness": 6, "usability": 7, "robustness": 7 },
|
|
418
|
+
"subtotal": 6.6
|
|
419
|
+
}
|
|
420
|
+
},
|
|
421
|
+
"weighted": {
|
|
422
|
+
"methodology": 7.1, "depth": 5.7, "correctness": 6.4,
|
|
423
|
+
"usability": 6.4, "robustness": 7.1, "overall": 6.5
|
|
424
|
+
},
|
|
425
|
+
"disagreement_flags": [],
|
|
426
|
+
"meta_judge_triggered": false,
|
|
427
|
+
"status": "BELOW_THRESHOLD"
|
|
428
|
+
},
|
|
429
|
+
"verdict": {
|
|
430
|
+
"score": 6.5,
|
|
431
|
+
"threshold": 7.0,
|
|
432
|
+
"result": "NEEDS_IMPROVEMENT",
|
|
433
|
+
"recommended_action": "apply_priority_fixes_then_retest"
|
|
434
|
+
}
|
|
435
|
+
}
|
|
436
|
+
```
|
|
@@ -0,0 +1,189 @@
|
|
|
1
|
+
# BUILD Module — Artifact Generation Protocol
|
|
2
|
+
|
|
3
|
+
## Purpose
|
|
4
|
+
|
|
5
|
+
Generate production-quality Claude Code artifacts (skills, commands, rules, agent templates) from natural language requirements.
|
|
6
|
+
|
|
7
|
+
## Input
|
|
8
|
+
|
|
9
|
+
- **Description:** Natural language description of what the artifact should do
|
|
10
|
+
- **Type:** skill | command | rule | agent (auto-detected or explicit)
|
|
11
|
+
- **Mode:** QUICK | DEEP (default: QUICK)
|
|
12
|
+
- **References:** Optional paths to existing artifacts as examples
|
|
13
|
+
|
|
14
|
+
## Protocol
|
|
15
|
+
|
|
16
|
+
### Step 1: Type Detection
|
|
17
|
+
|
|
18
|
+
If type is not explicitly specified, detect from description:
|
|
19
|
+
|
|
20
|
+
| Signal | Detected Type |
|
|
21
|
+
|--------|--------------|
|
|
22
|
+
| "skill", "module", "capability", "protocol" | skill |
|
|
23
|
+
| "command", "slash command", "/something", "pipeline" | command |
|
|
24
|
+
| "rule", "constraint", "convention", "anti-pattern" | rule |
|
|
25
|
+
| "agent", "worker", "parallel", "swarm" | agent |
|
|
26
|
+
|
|
27
|
+
### Step 2: Requirements Gathering
|
|
28
|
+
|
|
29
|
+
**QUICK mode** — Extract from description directly:
|
|
30
|
+
1. Parse artifact name (kebab-case)
|
|
31
|
+
2. Extract key capabilities
|
|
32
|
+
3. Identify domain constraints
|
|
33
|
+
4. Proceed to generation
|
|
34
|
+
|
|
35
|
+
**DEEP mode** — Use `explore` skill:
|
|
36
|
+
1. Read `.claude/skills/explore/SKILL.md`
|
|
37
|
+
2. Follow explore protocol to clarify:
|
|
38
|
+
- Exact scope and boundaries
|
|
39
|
+
- Target users/consumers
|
|
40
|
+
- Input/output format
|
|
41
|
+
- Quality criteria
|
|
42
|
+
- Edge cases
|
|
43
|
+
3. Produce requirements brief
|
|
44
|
+
4. Confirm with user before generation
|
|
45
|
+
|
|
46
|
+
### Step 3: Generation Templates
|
|
47
|
+
|
|
48
|
+
#### Skill Template
|
|
49
|
+
|
|
50
|
+
```
|
|
51
|
+
.claude/skills/<name>/
|
|
52
|
+
├── SKILL.md
|
|
53
|
+
│ ├── # <Name>
|
|
54
|
+
│ ├── ## Overview
|
|
55
|
+
│ ├── ## Quick Start
|
|
56
|
+
│ ├── ## Protocol
|
|
57
|
+
│ │ ├── Step 1: ...
|
|
58
|
+
│ │ ├── Step 2: ...
|
|
59
|
+
│ │ └── Step N: ...
|
|
60
|
+
│ ├── ## Output Format
|
|
61
|
+
│ ├── ## Anti-Patterns
|
|
62
|
+
│ └── ## Dependencies
|
|
63
|
+
├── modules/ (if multi-module)
|
|
64
|
+
│ └── <module>.md
|
|
65
|
+
├── references/ (always include at least one)
|
|
66
|
+
│ └── <ref>.md
|
|
67
|
+
└── examples/ (always include at least one)
|
|
68
|
+
└── <example>.md
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
#### Command Template
|
|
72
|
+
|
|
73
|
+
```markdown
|
|
74
|
+
# /command-name — Short Description
|
|
75
|
+
|
|
76
|
+
## Usage
|
|
77
|
+
/command-name [arguments]
|
|
78
|
+
|
|
79
|
+
## Parameters
|
|
80
|
+
- $ARGUMENTS — Description
|
|
81
|
+
|
|
82
|
+
## Protocol
|
|
83
|
+
|
|
84
|
+
### Step 1: Setup
|
|
85
|
+
- Validate arguments
|
|
86
|
+
- Load required skills: Read `.claude/skills/<skill>/SKILL.md`
|
|
87
|
+
|
|
88
|
+
### Step 2: Execution
|
|
89
|
+
- Main logic here
|
|
90
|
+
- Use Agent tool for parallelism where applicable
|
|
91
|
+
|
|
92
|
+
### Step 3: Output
|
|
93
|
+
- Create artifact files
|
|
94
|
+
- Display checkpoint
|
|
95
|
+
|
|
96
|
+
## Checkpoint
|
|
97
|
+
═══════════════════════════════════════════════════════
|
|
98
|
+
⏸️ CHECKPOINT: [Command Name] Complete
|
|
99
|
+
...
|
|
100
|
+
═══════════════════════════════════════════════════════
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
#### Rule Template
|
|
104
|
+
|
|
105
|
+
```markdown
|
|
106
|
+
# Rule Name
|
|
107
|
+
|
|
108
|
+
## Patterns
|
|
109
|
+
|
|
110
|
+
| Pattern | Detection Signal | Required Fix |
|
|
111
|
+
|---------|-----------------|-------------|
|
|
112
|
+
| ... | ... | ... |
|
|
113
|
+
|
|
114
|
+
## Auto-Detection
|
|
115
|
+
When generating content, self-check against these patterns.
|
|
116
|
+
If detected, flag with ⚠️ and fix before proceeding.
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
#### Agent Template
|
|
120
|
+
|
|
121
|
+
```markdown
|
|
122
|
+
# Agent Name
|
|
123
|
+
|
|
124
|
+
## Purpose
|
|
125
|
+
What this agent does.
|
|
126
|
+
|
|
127
|
+
## Configuration
|
|
128
|
+
- Model: haiku | sonnet | opus
|
|
129
|
+
- Isolation: reads X, writes Y
|
|
130
|
+
- Max turns: N
|
|
131
|
+
|
|
132
|
+
## Prompt Template
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
### Step 4: Self-Review
|
|
136
|
+
|
|
137
|
+
After generation, validate against quality checklist:
|
|
138
|
+
|
|
139
|
+
1. **Structure check:**
|
|
140
|
+
- Required sections present for artifact type
|
|
141
|
+
- No empty sections
|
|
142
|
+
- Proper markdown formatting
|
|
143
|
+
|
|
144
|
+
2. **Content check:**
|
|
145
|
+
- No generic/placeholder content
|
|
146
|
+
- Specific to the domain
|
|
147
|
+
- Anti-patterns section populated
|
|
148
|
+
- At least one concrete example
|
|
149
|
+
|
|
150
|
+
3. **Convention check:**
|
|
151
|
+
- File naming: kebab-case
|
|
152
|
+
- Directory naming: kebab-case
|
|
153
|
+
- Heading hierarchy: proper nesting
|
|
154
|
+
- References resolve to actual files
|
|
155
|
+
|
|
156
|
+
4. **Size check:**
|
|
157
|
+
- SKILL.md: 2KB-30KB (ideal: 5KB-15KB)
|
|
158
|
+
- Module: 1KB-15KB
|
|
159
|
+
- Reference: 500B-10KB
|
|
160
|
+
- Example: 500B-5KB
|
|
161
|
+
|
|
162
|
+
### Step 5: Output
|
|
163
|
+
|
|
164
|
+
1. Create all files in the target directory
|
|
165
|
+
2. Display summary:
|
|
166
|
+
```
|
|
167
|
+
✅ BUILD Complete
|
|
168
|
+
Artifact: .claude/skills/<name>/
|
|
169
|
+
Files created:
|
|
170
|
+
- SKILL.md (X KB)
|
|
171
|
+
- modules/<m>.md (X KB)
|
|
172
|
+
- references/<r>.md (X KB)
|
|
173
|
+
- examples/<e>.md (X KB)
|
|
174
|
+
|
|
175
|
+
Next: Run /bto-test .claude/skills/<name>/ to evaluate
|
|
176
|
+
```
|
|
177
|
+
|
|
178
|
+
## Anti-Patterns
|
|
179
|
+
|
|
180
|
+
| Anti-Pattern | Detection | Fix |
|
|
181
|
+
|-------------|-----------|-----|
|
|
182
|
+
| Generic skill | No domain-specific terms | Add domain context and constraints |
|
|
183
|
+
| Missing references | references/ empty | Add at least one reference file |
|
|
184
|
+
| No examples | examples/ empty | Add at least one few-shot example |
|
|
185
|
+
| Over-scoped | SKILL.md > 30KB | Split into modules |
|
|
186
|
+
| Under-specified | SKILL.md < 2KB | Expand with more detail |
|
|
187
|
+
| Copy-paste | Identical to another skill | Adapt uniquely |
|
|
188
|
+
| Missing anti-patterns | No anti-patterns section | Add common failure modes |
|
|
189
|
+
| No output format | Doesn't specify expected output | Add explicit output section |
|