@dzhechkov/skills-bto 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bin/cli.js +5 -0
- package/package.json +43 -0
- package/src/cli.js +150 -0
- package/src/commands/doctor.js +366 -0
- package/src/commands/init.js +188 -0
- package/src/commands/list.js +161 -0
- package/src/commands/remove.js +211 -0
- package/src/commands/update.js +198 -0
- package/src/utils.js +398 -0
- package/templates/.claude/agents/bto-judge-panel.md +192 -0
- package/templates/.claude/agents/bto-optimizer-worker.md +181 -0
- package/templates/.claude/commands/bto-build.md +169 -0
- package/templates/.claude/commands/bto-optimize.md +208 -0
- package/templates/.claude/commands/bto-test.md +186 -0
- package/templates/.claude/commands/bto.md +171 -0
- package/templates/.claude/rules/bto-quality-gates.md +91 -0
- package/templates/.claude/skills/bto/SKILL.md +266 -0
- package/templates/.claude/skills/bto/examples/sample-eval-report.md +436 -0
- package/templates/.claude/skills/bto/modules/build.md +189 -0
- package/templates/.claude/skills/bto/modules/optimize.md +201 -0
- package/templates/.claude/skills/bto/modules/test.md +348 -0
- package/templates/.claude/skills/bto/references/eval-patterns.md +202 -0
- package/templates/.claude/skills/bto/references/judge-rubrics.md +183 -0
- package/templates/.claude/skills/bto/references/optimization-methods.md +139 -0
- package/templates/.claude/skills/bto/references/quality-checklist.md +220 -0
|
@@ -0,0 +1,220 @@
|
|
|
1
|
+
# Pre-Flight Quality Checklist — Claude Code Artifacts
|
|
2
|
+
|
|
3
|
+
> Deterministic Layer 0 checks. Run before any LLM evaluation. Free, fast, catches ~60% of issues.
|
|
4
|
+
> Pass rate threshold: **≥ 80%** to proceed to Layer 1/2 evaluation.
|
|
5
|
+
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
## How to Use This Checklist
|
|
9
|
+
|
|
10
|
+
1. Identify artifact type (Skill / Command / Rule / Agent Template / Research Artifact)
|
|
11
|
+
2. Run **Universal Checks** (apply to all types)
|
|
12
|
+
3. Run the **type-specific section**
|
|
13
|
+
4. Count: `pass_rate = passed / total_applicable`
|
|
14
|
+
5. If `pass_rate < 0.80` → return to BUILD phase, fix issues, re-check
|
|
15
|
+
|
|
16
|
+
**Auto-fixable items** (marked `YES`) can be fixed programmatically without human review.
|
|
17
|
+
**Non-auto-fixable items** (marked `NO`) require human or LLM-assisted review.
|
|
18
|
+
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## Section 1 — Universal Checks (All Artifact Types)
|
|
22
|
+
|
|
23
|
+
> These checks apply regardless of artifact type. All are required.
|
|
24
|
+
|
|
25
|
+
| ID | Check | Pass Criteria | Fail Criteria | Auto-fixable |
|
|
26
|
+
|----|-------|---------------|---------------|--------------|
|
|
27
|
+
| U-01 | File exists and is non-empty | File size > 0 bytes | Missing file or 0-byte file | NO |
|
|
28
|
+
| U-02 | UTF-8 encoding | Valid UTF-8, no BOM | Encoding errors, binary content | YES |
|
|
29
|
+
| U-03 | Starts with a level-1 heading (`# Title`) | First non-blank line is `# ...` | Missing or wrong heading level | YES |
|
|
30
|
+
| U-04 | No placeholder text remaining | No occurrences of `TODO`, `FIXME`, `[INSERT`, `<YOUR_`, `...` (as placeholder) | Any placeholder found | NO |
|
|
31
|
+
| U-05 | No empty sections | No heading immediately followed by another heading (without content) | Empty section found | NO |
|
|
32
|
+
| U-06 | Consistent heading hierarchy | Headings increment by 1 level (no jump from `##` to `####`) | Skipped heading level | YES |
|
|
33
|
+
| U-07 | No broken internal cross-references | All referenced filenames exist on disk | Referenced file not found | NO |
|
|
34
|
+
| U-08 | File size within bounds | 200 bytes ≤ size ≤ 100 KB (single file) | Under 200B (stub) or over 100KB (bloated) | NO |
|
|
35
|
+
| U-09 | No trailing whitespace on lines | All lines trimmed | Lines with trailing spaces/tabs | YES |
|
|
36
|
+
| U-10 | Uses standard Markdown (no raw HTML) | No `<div>`, `<span>`, `<table>` tags | HTML tags present | NO |
|
|
37
|
+
| U-11 | Code blocks are properly closed | Every ` ``` ` open has a matching ` ``` ` close | Unclosed code block | YES |
|
|
38
|
+
| U-12 | No duplicate top-level sections | All `## Heading` names are unique within the file | Duplicate heading names | NO |
|
|
39
|
+
|
|
40
|
+
**Universal checks total: 12**
|
|
41
|
+
|
|
42
|
+
---
|
|
43
|
+
|
|
44
|
+
## Section 2 — Skill-Specific Checks
|
|
45
|
+
|
|
46
|
+
> Apply when artifact type is a **Skill** (directory with SKILL.md).
|
|
47
|
+
|
|
48
|
+
| ID | Check | Pass Criteria | Fail Criteria | Auto-fixable |
|
|
49
|
+
|----|-------|---------------|---------------|--------------|
|
|
50
|
+
| SK-01 | `SKILL.md` exists at skill root | File present at `.claude/skills/<name>/SKILL.md` | Missing SKILL.md | NO |
|
|
51
|
+
| SK-02 | Has `## Overview` section | Section `## Overview` present | Missing overview | NO |
|
|
52
|
+
| SK-03 | Has `## Anti-Patterns` section | Section `## Anti-Patterns` present | Anti-patterns section absent | NO |
|
|
53
|
+
| SK-04 | Has at least one `## Quick Start` or usage example | Section `## Quick Start`, `## Usage`, or code block with invocation | No invocation example | NO |
|
|
54
|
+
| SK-05 | `modules/` directory exists | Directory `.claude/skills/<name>/modules/` present | No modules directory | YES |
|
|
55
|
+
| SK-06 | Every module referenced in SKILL.md exists on disk | All filenames in `modules/` references resolve | Broken module reference | NO |
|
|
56
|
+
| SK-07 | `references/` directory exists | Directory `.claude/skills/<name>/references/` present | No references directory | YES |
|
|
57
|
+
| SK-08 | Every reference referenced in SKILL.md exists on disk | All filenames in `references/` references resolve | Broken reference | NO |
|
|
58
|
+
| SK-09 | `examples/` directory exists | Directory `.claude/skills/<name>/examples/` present | No examples directory | YES |
|
|
59
|
+
| SK-10 | At least one example file present | ≥1 file in examples/ | Empty examples directory | NO |
|
|
60
|
+
| SK-11 | Skill name in heading matches directory name | `# <Name>` roughly matches `skills/<name>/` | Name mismatch | YES |
|
|
61
|
+
| SK-12 | Has `## Dependencies` section (or explicitly states none) | Section present or "no dependencies" stated | Undocumented dependencies | NO |
|
|
62
|
+
| SK-13 | SKILL.md size within skill bounds | 2 KB ≤ SKILL.md ≤ 50 KB | Too small (stub) or too large (monolith) | NO |
|
|
63
|
+
| SK-14 | Total skill directory size within bounds | Total ≤ 200 KB | Oversized skill (likely includes binaries) | NO |
|
|
64
|
+
| SK-15 | Each module file has its own `# Title` heading | First non-blank line in every module is `# ...` | Module missing title | YES |
|
|
65
|
+
| SK-16 | No circular cross-references between skill files | No file references itself | Self-reference loop | NO |
|
|
66
|
+
|
|
67
|
+
**Skill-specific checks total: 16**
|
|
68
|
+
|
|
69
|
+
---
|
|
70
|
+
|
|
71
|
+
## Section 3 — Command-Specific Checks
|
|
72
|
+
|
|
73
|
+
> Apply when artifact type is a **Command** (single `.md` file in `.claude/commands/`).
|
|
74
|
+
|
|
75
|
+
| ID | Check | Pass Criteria | Fail Criteria | Auto-fixable |
|
|
76
|
+
|----|-------|---------------|---------------|--------------|
|
|
77
|
+
| CM-01 | File located in `.claude/commands/` | Correct directory | Wrong location | NO |
|
|
78
|
+
| CM-02 | References `$ARGUMENTS` | String `$ARGUMENTS` appears in file | No argument handling | NO |
|
|
79
|
+
| CM-03 | Has checkpoint protocol or defers to checkpoint rule | Section mentioning checkpoint or explicit link to checkpoint-protocol rule | No checkpoint mentioned | NO |
|
|
80
|
+
| CM-04 | Has at least one skill loading instruction | Contains `Read .claude/skills/` or equivalent | No skill loading | NO |
|
|
81
|
+
| CM-05 | Has a usage line (how to invoke) | Contains usage pattern like `` /command-name [args] `` | No usage line | NO |
|
|
82
|
+
| CM-06 | Has at least one invocation example | Code block or `` `example` `` of actual invocation | No example | NO |
|
|
83
|
+
| CM-07 | Phase-specific commands reference correct phase number | Phase N command mentions Phase N consistently | Phase number mismatch | NO |
|
|
84
|
+
| CM-08 | Agent usage (if any) follows agent-swarm rules | Agent spawning uses naming convention from agent-swarm.md | Non-compliant agent naming | NO |
|
|
85
|
+
| CM-09 | File size within command bounds | 500 bytes ≤ size ≤ 20 KB | Too small (stub) or too large | NO |
|
|
86
|
+
| CM-10 | Output artifact paths are inside `researches/<slug>/` | All artifact paths reference `researches/` directory | Paths outside research scope | NO |
|
|
87
|
+
| CM-11 | Has error/fallback handling for missing arguments | Specifies behavior when `$ARGUMENTS` is empty | No empty-argument handling | NO |
|
|
88
|
+
|
|
89
|
+
**Command-specific checks total: 11**
|
|
90
|
+
|
|
91
|
+
---
|
|
92
|
+
|
|
93
|
+
## Section 4 — Rule-Specific Checks
|
|
94
|
+
|
|
95
|
+
> Apply when artifact type is a **Rule** (file in `.claude/rules/`).
|
|
96
|
+
|
|
97
|
+
| ID | Check | Pass Criteria | Fail Criteria | Auto-fixable |
|
|
98
|
+
|----|-------|---------------|---------------|--------------|
|
|
99
|
+
| RL-01 | File located in `.claude/rules/` | Correct directory | Wrong location | NO |
|
|
100
|
+
| RL-02 | Contains a detection signal for each pattern | Each pattern row has a "signal" or "detection" column | Patterns without signals | NO |
|
|
101
|
+
| RL-03 | Contains a fix/remedy for each pattern | Each pattern row has "fix", "remedy", or "action" | Patterns without fixes | NO |
|
|
102
|
+
| RL-04 | Uses table format for patterns | Markdown table with ≥2 columns present | Unstructured list only | YES |
|
|
103
|
+
| RL-05 | Minimum pattern coverage | ≥3 distinct patterns listed | Fewer than 3 patterns | NO |
|
|
104
|
+
| RL-06 | No vague patterns | Each pattern is specific and detectable, no "general sloppiness" | Vague or unmeasurable patterns | NO |
|
|
105
|
+
| RL-07 | Has blocking or non-blocking designation | Each pattern marked as block/warn/flag or severity noted | No severity indication | NO |
|
|
106
|
+
| RL-08 | Auto-detection instructions present | Section or note on how to auto-detect patterns | No auto-detection guidance | NO |
|
|
107
|
+
| RL-09 | File size within rule bounds | 200 bytes ≤ size ≤ 10 KB | Too small or too large | NO |
|
|
108
|
+
|
|
109
|
+
**Rule-specific checks total: 9**
|
|
110
|
+
|
|
111
|
+
---
|
|
112
|
+
|
|
113
|
+
## Section 5 — Agent Template-Specific Checks
|
|
114
|
+
|
|
115
|
+
> Apply when artifact type is an **Agent Template** (file describing an agent configuration).
|
|
116
|
+
|
|
117
|
+
| ID | Check | Pass Criteria | Fail Criteria | Auto-fixable |
|
|
118
|
+
|----|-------|---------------|---------------|--------------|
|
|
119
|
+
| AT-01 | Agent purpose is clearly stated in one sentence | First paragraph or overview states agent goal unambiguously | Vague or missing purpose | NO |
|
|
120
|
+
| AT-02 | Model selection is specified and justified | Explicit model name (haiku/sonnet/default) with reason | No model specified | NO |
|
|
121
|
+
| AT-03 | Agent isolation scope is defined | Specifies which files/directories agent may read and write | No scope defined | NO |
|
|
122
|
+
| AT-04 | Output format is specified | Describes expected output structure, keys, or format | No output format | NO |
|
|
123
|
+
| AT-05 | Timeout or cost bounds defined | Specifies max duration, max tokens, or cost cap | No bounds defined | NO |
|
|
124
|
+
| AT-06 | Failure protocol defined | Specifies behavior on error, timeout, or unexpected output | No failure protocol | NO |
|
|
125
|
+
| AT-07 | No conflicts with other agents in same pipeline | Does not claim exclusive access to resources other agents use | Resource conflicts found | NO |
|
|
126
|
+
| AT-08 | Integration protocol specified | Describes how orchestrator collects and uses agent output | No integration guidance | NO |
|
|
127
|
+
| AT-09 | Naming convention followed | Agent name follows `Phase N [Role]` or descriptive convention | Non-standard name | YES |
|
|
128
|
+
| AT-10 | Has at least one example invocation via Agent tool | Shows how to call this agent using Claude's Agent tool | No invocation example | NO |
|
|
129
|
+
|
|
130
|
+
**Agent template-specific checks total: 10**
|
|
131
|
+
|
|
132
|
+
---
|
|
133
|
+
|
|
134
|
+
## Section 6 — Research Artifact Checks
|
|
135
|
+
|
|
136
|
+
> Apply when artifact type is a **Research Artifact** (files like `00_product_discovery.md`, `02_research_findings.md`, etc.).
|
|
137
|
+
|
|
138
|
+
| ID | Check | Pass Criteria | Fail Criteria | Auto-fixable |
|
|
139
|
+
|----|-------|---------------|---------------|--------------|
|
|
140
|
+
| RA-01 | Located inside `researches/<slug>/` | Correct directory structure | File in project root or wrong location | NO |
|
|
141
|
+
| RA-02 | Numbered prefix matches phase | File name starts with correct phase number (00_, 01_, etc.) | Wrong or missing phase prefix | YES |
|
|
142
|
+
| RA-03 | Has executive summary or TL;DR section | Section `## Summary`, `## TL;DR`, or `## Executive Summary` present | No summary section | NO |
|
|
143
|
+
| RA-04 | All factual claims have source attribution | Every statistic and factual claim ends with `— Source: [...]` | Unsourced claims found | NO |
|
|
144
|
+
| RA-05 | No [UNVERIFIED] tags in final artifact | String `[UNVERIFIED]` absent in completed artifact | Unverified claims remain | NO |
|
|
145
|
+
| RA-06 | Self-generated analysis is marked | LLM analysis sections labeled `[ANALYSIS]` | Unlabeled LLM-generated analysis | NO |
|
|
146
|
+
| RA-07 | No hallucinated company or product names | All named companies/products verified to exist | Fabricated entities | NO |
|
|
147
|
+
| RA-08 | Regulatory references cite actual laws | Law/standard references include number (e.g., ФЗ-152, ISO 27001) | Vague regulatory references | NO |
|
|
148
|
+
| RA-09 | Competitive analysis includes real competitors | Named competitors are real, operating organizations | Fake or outdated competitors | NO |
|
|
149
|
+
| RA-10 | KPIs and metrics are specific and quantified | Metrics include baseline and target numbers | Vague metrics ("improve efficiency") | NO |
|
|
150
|
+
| RA-11 | Has a limitations or confidence section | Section noting research limitations, gaps, or confidence levels | No caveats or limitations | NO |
|
|
151
|
+
| RA-12 | Date of research noted | File contains creation or research date | No date context | YES |
|
|
152
|
+
| RA-13 | Phase artifact created during its phase | Timestamp of file creation matches expected phase window | Deferred creation | NO |
|
|
153
|
+
|
|
154
|
+
**Research artifact checks total: 13**
|
|
155
|
+
|
|
156
|
+
---
|
|
157
|
+
|
|
158
|
+
## Scoring Summary
|
|
159
|
+
|
|
160
|
+
### Calculation
|
|
161
|
+
|
|
162
|
+
```
|
|
163
|
+
total_applicable = Universal (12) + type-specific checks
|
|
164
|
+
passed = count of checks that PASS
|
|
165
|
+
pass_rate = passed / total_applicable
|
|
166
|
+
```
|
|
167
|
+
|
|
168
|
+
### Thresholds
|
|
169
|
+
|
|
170
|
+
| Pass Rate | Status | Action |
|
|
171
|
+
|-----------|--------|--------|
|
|
172
|
+
| ≥ 90% | EXCELLENT | Proceed to Layer 1/2 with high confidence |
|
|
173
|
+
| 80% – 89% | PASS | Proceed to Layer 1/2 evaluation |
|
|
174
|
+
| 60% – 79% | CONDITIONAL | Fix auto-fixable items, re-check before Layer 1 |
|
|
175
|
+
| < 60% | FAIL | Return to BUILD phase — do not evaluate yet |
|
|
176
|
+
|
|
177
|
+
### Score Report Format
|
|
178
|
+
|
|
179
|
+
```
|
|
180
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
181
|
+
LAYER 0 PRE-FLIGHT CHECK
|
|
182
|
+
Artifact: <name> (<type>)
|
|
183
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
184
|
+
Universal checks: X / 12
|
|
185
|
+
Type-specific checks: X / N
|
|
186
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
187
|
+
TOTAL: X / (12+N)
|
|
188
|
+
Pass rate: XX%
|
|
189
|
+
Status: PASS | FAIL | CONDITIONAL
|
|
190
|
+
|
|
191
|
+
Failed checks:
|
|
192
|
+
- [ID] Description of failure
|
|
193
|
+
- [ID] Description of failure
|
|
194
|
+
(auto-fixable marked with *)
|
|
195
|
+
|
|
196
|
+
Verdict: [PROCEED TO LAYER 1] | [FIX AND RE-CHECK] | [RETURN TO BUILD]
|
|
197
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
---
|
|
201
|
+
|
|
202
|
+
## Quick Reference — Check ID Index
|
|
203
|
+
|
|
204
|
+
| Range | Section |
|
|
205
|
+
|-------|---------|
|
|
206
|
+
| U-01 to U-12 | Universal (all types) |
|
|
207
|
+
| SK-01 to SK-16 | Skill-specific |
|
|
208
|
+
| CM-01 to CM-11 | Command-specific |
|
|
209
|
+
| RL-01 to RL-09 | Rule-specific |
|
|
210
|
+
| AT-01 to AT-10 | Agent template-specific |
|
|
211
|
+
| RA-01 to RA-13 | Research artifact-specific |
|
|
212
|
+
|
|
213
|
+
---
|
|
214
|
+
|
|
215
|
+
## Maintenance Notes
|
|
216
|
+
|
|
217
|
+
- This checklist is **domain-agnostic** and applies to any IT project using the Claude Code skill/command/rule/agent architecture.
|
|
218
|
+
- When adding new artifact types, add a new section following the same table format (ID / Check / Pass Criteria / Fail Criteria / Auto-fixable).
|
|
219
|
+
- Review and update thresholds after every 50 evaluations to calibrate against real-world pass rates.
|
|
220
|
+
- Auto-fixable items should be addressed by the BUILD module before outputting an artifact.
|