@uzysjung/agent-harness 26.151.0 → 26.152.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.ko.md +2 -2
- package/README.md +2 -2
- package/dist/{chunk-3QBHZUVB.js → chunk-EQDC2AAU.js} +30 -23
- package/dist/chunk-EQDC2AAU.js.map +1 -0
- package/dist/index.js +133 -263
- package/dist/index.js.map +1 -1
- package/dist/trust-tier-drift.js +1 -1
- package/package.json +1 -1
- package/templates/rules/change-management.md +2 -3
- package/templates/rules/cli-development.md +2 -4
- package/templates/rules/doc-governance.md +1 -1
- package/templates/skills/audit-service-gaps/SKILL.md +0 -2
- package/templates/skills/model-orchestration/SKILL.md +6 -24
- package/templates/track-mcp-map.tsv +1 -1
- package/dist/chunk-3QBHZUVB.js.map +0 -1
- package/templates/agents/build-error-resolver.md +0 -111
- package/templates/agents/plan-checker.md +0 -116
- package/templates/agents/silent-failure-hunter.md +0 -50
- package/templates/skills/agent-introspection-debugging/SKILL.md +0 -153
- package/templates/skills/deep-research/SKILL.md +0 -179
- package/templates/skills/deep-research/agents/openai.yaml +0 -7
- package/templates/skills/eval-harness/SKILL.md +0 -306
- package/templates/skills/eval-harness/agents/openai.yaml +0 -7
- package/templates/skills/verification-loop/SKILL.md +0 -250
- package/templates/skills/verification-loop/agents/openai.yaml +0 -7
- package/templates/skills/verification-loop/references/tracks.md +0 -49
|
@@ -1,306 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: eval-harness
|
|
3
|
-
description: Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles
|
|
4
|
-
origin: ECC
|
|
5
|
-
tools: Read, Write, Edit, Bash, Grep, Glob
|
|
6
|
-
---
|
|
7
|
-
|
|
8
|
-
# Eval Harness Skill
|
|
9
|
-
|
|
10
|
-
A formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles.
|
|
11
|
-
|
|
12
|
-
## When to Activate
|
|
13
|
-
|
|
14
|
-
- Setting up eval-driven development (EDD) for AI-assisted workflows
|
|
15
|
-
- Defining pass/fail criteria for Claude Code task completion
|
|
16
|
-
- Measuring agent reliability with pass@k metrics
|
|
17
|
-
- Creating regression test suites for prompt or agent changes
|
|
18
|
-
- Benchmarking agent performance across model versions
|
|
19
|
-
|
|
20
|
-
## Philosophy
|
|
21
|
-
|
|
22
|
-
Eval-Driven Development treats evals as the "unit tests of AI development":
|
|
23
|
-
- Define expected behavior BEFORE implementation
|
|
24
|
-
- Run evals continuously during development
|
|
25
|
-
- Track regressions with each change
|
|
26
|
-
- Use pass@k metrics for reliability measurement
|
|
27
|
-
|
|
28
|
-
## Eval Types
|
|
29
|
-
|
|
30
|
-
### Capability Evals
|
|
31
|
-
Test if Claude can do something it couldn't before:
|
|
32
|
-
```markdown
|
|
33
|
-
[CAPABILITY EVAL: feature-name]
|
|
34
|
-
Task: Description of what Claude should accomplish
|
|
35
|
-
Success Criteria:
|
|
36
|
-
- [ ] Criterion 1
|
|
37
|
-
- [ ] Criterion 2
|
|
38
|
-
- [ ] Criterion 3
|
|
39
|
-
Expected Output: Description of expected result
|
|
40
|
-
```
|
|
41
|
-
|
|
42
|
-
### Regression Evals
|
|
43
|
-
Ensure changes don't break existing functionality:
|
|
44
|
-
```markdown
|
|
45
|
-
[REGRESSION EVAL: feature-name]
|
|
46
|
-
Baseline: SHA or checkpoint name
|
|
47
|
-
Tests:
|
|
48
|
-
- existing-test-1: PASS/FAIL
|
|
49
|
-
- existing-test-2: PASS/FAIL
|
|
50
|
-
- existing-test-3: PASS/FAIL
|
|
51
|
-
Result: X/Y passed (previously Y/Y)
|
|
52
|
-
```
|
|
53
|
-
|
|
54
|
-
## Grader Types
|
|
55
|
-
|
|
56
|
-
### 1. Code-Based Grader
|
|
57
|
-
Deterministic checks using code:
|
|
58
|
-
```bash
|
|
59
|
-
# Check if file contains expected pattern
|
|
60
|
-
grep -q "export function handleAuth" src/auth.ts && echo "PASS" || echo "FAIL"
|
|
61
|
-
|
|
62
|
-
# Check if tests pass
|
|
63
|
-
npm test -- --testPathPattern="auth" && echo "PASS" || echo "FAIL"
|
|
64
|
-
|
|
65
|
-
# Check if build succeeds
|
|
66
|
-
npm run build && echo "PASS" || echo "FAIL"
|
|
67
|
-
```
|
|
68
|
-
|
|
69
|
-
### 2. Model-Based Grader
|
|
70
|
-
Use Claude to evaluate open-ended outputs:
|
|
71
|
-
```markdown
|
|
72
|
-
[MODEL GRADER PROMPT]
|
|
73
|
-
Evaluate the following code change:
|
|
74
|
-
1. Does it solve the stated problem?
|
|
75
|
-
2. Is it well-structured?
|
|
76
|
-
3. Are edge cases handled?
|
|
77
|
-
4. Is error handling appropriate?
|
|
78
|
-
|
|
79
|
-
Score: 1-5 (1=poor, 5=excellent)
|
|
80
|
-
Reasoning: [explanation]
|
|
81
|
-
```
|
|
82
|
-
|
|
83
|
-
### 3. Human Grader
|
|
84
|
-
Flag for manual review:
|
|
85
|
-
```markdown
|
|
86
|
-
[HUMAN REVIEW REQUIRED]
|
|
87
|
-
Change: Description of what changed
|
|
88
|
-
Reason: Why human review is needed
|
|
89
|
-
Risk Level: LOW/MEDIUM/HIGH
|
|
90
|
-
```
|
|
91
|
-
|
|
92
|
-
## Metrics
|
|
93
|
-
|
|
94
|
-
### pass@k
|
|
95
|
-
"At least one success in k attempts"
|
|
96
|
-
- pass@1: First attempt success rate
|
|
97
|
-
- pass@3: Success within 3 attempts
|
|
98
|
-
- Typical target: pass@3 > 90%
|
|
99
|
-
|
|
100
|
-
### pass^k
|
|
101
|
-
"All k trials succeed"
|
|
102
|
-
- Higher bar for reliability
|
|
103
|
-
- pass^3: 3 consecutive successes
|
|
104
|
-
- Use for critical paths
|
|
105
|
-
|
|
106
|
-
## Eval Workflow
|
|
107
|
-
|
|
108
|
-
### 1. Define (Before Coding)
|
|
109
|
-
|
|
110
|
-
Write the spec to a file (`.claude/evals/<feature>.md`) before implementing. Give every eval a
|
|
111
|
-
stable ID — `C1..Cn` for capability, `R1..Rn` for regression — so the same identifier carries
|
|
112
|
-
from definition to the post-implementation status line, and a reviewer can check them off one by
|
|
113
|
-
one. Prose-only lists ("Can create new user account") can't be referenced or scored.
|
|
114
|
-
|
|
115
|
-
````markdown
|
|
116
|
-
# EVAL: <feature name> (<phase/PR>)
|
|
117
|
-
|
|
118
|
-
**Feature**: <one line>
|
|
119
|
-
**Baseline**: commit <sha> # what "regression" is measured against
|
|
120
|
-
**Target**: pass@1 = 100% for capability evals
|
|
121
|
-
|
|
122
|
-
## Capability Evals
|
|
123
|
-
|
|
124
|
-
### C1: <name>
|
|
125
|
-
- <concrete, checkable expectation — inputs → expected output>
|
|
126
|
-
|
|
127
|
-
### C2: <name>
|
|
128
|
-
- <expectation>
|
|
129
|
-
|
|
130
|
-
## Regression Evals
|
|
131
|
-
|
|
132
|
-
### R1: <existing behavior that must not move>
|
|
133
|
-
- <expectation>
|
|
134
|
-
|
|
135
|
-
## Test Command
|
|
136
|
-
```bash
|
|
137
|
-
pytest tests/test_<area>.py -v -k "<selector>"
|
|
138
|
-
```
|
|
139
|
-
|
|
140
|
-
## Status (after implementation)
|
|
141
|
-
- C1-Cn: PASS via <tests (N개) / route check / manual>
|
|
142
|
-
- R1-Rn: PASS
|
|
143
|
-
|
|
144
|
-
**Overall: pass@1 = <x>%**
|
|
145
|
-
````
|
|
146
|
-
|
|
147
|
-
Two fields carry most of the weight. **Baseline commit** makes "regression" falsifiable — without
|
|
148
|
-
it, R-evals are opinions about the past. **Test Command** makes the spec re-runnable by someone
|
|
149
|
-
who didn't write it; an eval nobody can re-run is documentation, not a gate.
|
|
150
|
-
|
|
151
|
-
Fill the Status section *after* implementing, in the same file. A spec whose status is still empty
|
|
152
|
-
at merge time means the evals were written and never used.
|
|
153
|
-
|
|
154
|
-
### 2. Implement
|
|
155
|
-
Write code to pass the defined evals.
|
|
156
|
-
|
|
157
|
-
### 3. Evaluate
|
|
158
|
-
```bash
|
|
159
|
-
# Run capability evals
|
|
160
|
-
[Run each capability eval, record PASS/FAIL]
|
|
161
|
-
|
|
162
|
-
# Run regression evals
|
|
163
|
-
npm test -- --testPathPattern="existing"
|
|
164
|
-
|
|
165
|
-
# Generate report
|
|
166
|
-
```
|
|
167
|
-
|
|
168
|
-
### 4. Report
|
|
169
|
-
```markdown
|
|
170
|
-
EVAL REPORT: feature-xyz
|
|
171
|
-
========================
|
|
172
|
-
|
|
173
|
-
Capability Evals:
|
|
174
|
-
create-user: PASS (pass@1)
|
|
175
|
-
validate-email: PASS (pass@2)
|
|
176
|
-
hash-password: PASS (pass@1)
|
|
177
|
-
Overall: 3/3 passed
|
|
178
|
-
|
|
179
|
-
Regression Evals:
|
|
180
|
-
login-flow: PASS
|
|
181
|
-
session-mgmt: PASS
|
|
182
|
-
logout-flow: PASS
|
|
183
|
-
Overall: 3/3 passed
|
|
184
|
-
|
|
185
|
-
Metrics:
|
|
186
|
-
pass@1: 67% (2/3)
|
|
187
|
-
pass@3: 100% (3/3)
|
|
188
|
-
|
|
189
|
-
Status: READY FOR REVIEW
|
|
190
|
-
```
|
|
191
|
-
|
|
192
|
-
## Integration Patterns
|
|
193
|
-
|
|
194
|
-
### Pre-Implementation
|
|
195
|
-
```
|
|
196
|
-
/eval define feature-name
|
|
197
|
-
```
|
|
198
|
-
Creates eval definition file at `.claude/evals/feature-name.md`
|
|
199
|
-
|
|
200
|
-
### During Implementation
|
|
201
|
-
```
|
|
202
|
-
/eval check feature-name
|
|
203
|
-
```
|
|
204
|
-
Runs current evals and reports status
|
|
205
|
-
|
|
206
|
-
### Post-Implementation
|
|
207
|
-
```
|
|
208
|
-
/eval report feature-name
|
|
209
|
-
```
|
|
210
|
-
Generates full eval report
|
|
211
|
-
|
|
212
|
-
## Eval Storage (.md + .log Pair Format)
|
|
213
|
-
|
|
214
|
-
각 평가 항목은 **`<topic>.md` (설계) + `<topic>.log` (실행 결과)** 쌍으로 저장. 강제. 단독 .md만 있으면 재현 불가.
|
|
215
|
-
|
|
216
|
-
```
|
|
217
|
-
.claude/
|
|
218
|
-
evals/
|
|
219
|
-
feature-xyz.md # Eval definition (Capability/Regression/Test 3섹션 필수)
|
|
220
|
-
feature-xyz.log # Eval run history (실행 시각, grader, pass/fail)
|
|
221
|
-
session-YYYYMMDD.md # 세션 단위 회고 + 차기 backlog
|
|
222
|
-
session-YYYYMMDD.log # 동일 세션의 grader 출력
|
|
223
|
-
baseline.json # Regression baselines (선택)
|
|
224
|
-
```
|
|
225
|
-
|
|
226
|
-
> eval 산출물은 `docs/evals/*.{md,log}` 로 모은다 — 실행 로그와 판정을 같은 자리에 둔다.
|
|
227
|
-
|
|
228
|
-
### .md 파일 의무 섹션 (3개)
|
|
229
|
-
|
|
230
|
-
```markdown
|
|
231
|
-
# Eval: <topic>
|
|
232
|
-
|
|
233
|
-
## Capability
|
|
234
|
-
[새 능력 — Claude/agent가 무엇을 할 수 있는지]
|
|
235
|
-
- AC: [측정 가능 기준]
|
|
236
|
-
- Grader: code-based / model-based / human
|
|
237
|
-
|
|
238
|
-
## Regression
|
|
239
|
-
[기존 기능 보호 — 변경으로 깨지면 안 되는 baseline]
|
|
240
|
-
- Baseline: <SHA or checkpoint>
|
|
241
|
-
- Tests: [목록]
|
|
242
|
-
|
|
243
|
-
## Test
|
|
244
|
-
[실행 절차 — 누가 다시 돌려도 동일 결과 나와야 함]
|
|
245
|
-
- Setup: [사전 조건]
|
|
246
|
-
- Run: `bash run-eval.sh <topic>` 또는 명시적 명령
|
|
247
|
-
- Expected: [기대 출력]
|
|
248
|
-
```
|
|
249
|
-
|
|
250
|
-
### .log 파일 형식
|
|
251
|
-
|
|
252
|
-
각 실행마다 append. 시간순 누적.
|
|
253
|
-
|
|
254
|
-
```
|
|
255
|
-
=== 2026-04-19 14:32 (run #1) ===
|
|
256
|
-
Capability: 3/3 PASS (pass@1)
|
|
257
|
-
Regression: 5/5 PASS (pass^3)
|
|
258
|
-
Status: SHIP READY
|
|
259
|
-
|
|
260
|
-
=== 2026-04-20 09:15 (run #2 — after refactor) ===
|
|
261
|
-
Capability: 3/3 PASS
|
|
262
|
-
Regression: 4/5 PASS (login-flow regressed at SHA abc123)
|
|
263
|
-
Status: BLOCKED — fix login-flow first
|
|
264
|
-
```
|
|
265
|
-
|
|
266
|
-
## Best Practices
|
|
267
|
-
|
|
268
|
-
1. **Define evals BEFORE coding** - Forces clear thinking about success criteria
|
|
269
|
-
2. **Run evals frequently** - Catch regressions early
|
|
270
|
-
3. **Track pass@k over time** - Monitor reliability trends
|
|
271
|
-
4. **Use code graders when possible** - Deterministic > probabilistic
|
|
272
|
-
5. **Human review for security** - Never fully automate security checks
|
|
273
|
-
6. **Keep evals fast** - Slow evals don't get run
|
|
274
|
-
7. **Version evals with code** - Evals are first-class artifacts
|
|
275
|
-
|
|
276
|
-
## Example: Adding Authentication
|
|
277
|
-
|
|
278
|
-
```markdown
|
|
279
|
-
## EVAL: add-authentication
|
|
280
|
-
|
|
281
|
-
### Phase 1: Define (10 min)
|
|
282
|
-
Capability Evals:
|
|
283
|
-
- [ ] User can register with email/password
|
|
284
|
-
- [ ] User can login with valid credentials
|
|
285
|
-
- [ ] Invalid credentials rejected with proper error
|
|
286
|
-
- [ ] Sessions persist across page reloads
|
|
287
|
-
- [ ] Logout clears session
|
|
288
|
-
|
|
289
|
-
Regression Evals:
|
|
290
|
-
- [ ] Public routes still accessible
|
|
291
|
-
- [ ] API responses unchanged
|
|
292
|
-
- [ ] Database schema compatible
|
|
293
|
-
|
|
294
|
-
### Phase 2: Implement (varies)
|
|
295
|
-
[Write code]
|
|
296
|
-
|
|
297
|
-
### Phase 3: Evaluate
|
|
298
|
-
Run: /eval check add-authentication
|
|
299
|
-
|
|
300
|
-
### Phase 4: Report
|
|
301
|
-
EVAL REPORT: add-authentication
|
|
302
|
-
==============================
|
|
303
|
-
Capability: 5/5 passed (pass@3: 100%)
|
|
304
|
-
Regression: 3/3 passed (pass^3: 100%)
|
|
305
|
-
Status: SHIP IT
|
|
306
|
-
```
|
|
@@ -1,250 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: verification-loop
|
|
3
|
-
description: >-
|
|
4
|
-
A comprehensive verification system. Selects and runs proportional verification tracks for UI,
|
|
5
|
-
API/service, CLI/TUI, library/SDK, documents/configuration, and real user flows, then ends every
|
|
6
|
-
run with a fixed verdict — PASS / PASS_WITH_NITS / FAIL — plus severity-labeled findings
|
|
7
|
-
(CRITICAL/HIGH/MEDIUM/LOW) and the evidence each one rests on. Use after implementation, before
|
|
8
|
-
a PR or handoff, after a refactor, or to verify a claimed fix. Do NOT use a green build, a passing
|
|
9
|
-
type check, or file existence as proof of user-visible completion, and do NOT let the instance
|
|
10
|
-
that wrote the change issue its own verdict.
|
|
11
|
-
origin: ECC
|
|
12
|
-
---
|
|
13
|
-
|
|
14
|
-
# Verification Loop Skill
|
|
15
|
-
|
|
16
|
-
> Derived from the `verification-loop` skill in everything-claude-code (ECC), used under the MIT
|
|
17
|
-
> License.
|
|
18
|
-
|
|
19
|
-
A comprehensive verification system for coding sessions. The job is not "run the gates" — it is to
|
|
20
|
-
**verify the changed behavior through the surface a real user or consumer actually uses**, and to
|
|
21
|
-
end with one verdict that cannot be softened into prose.
|
|
22
|
-
|
|
23
|
-
## When to Use
|
|
24
|
-
|
|
25
|
-
Invoke this skill:
|
|
26
|
-
- After completing a feature or significant code change
|
|
27
|
-
- Before creating a PR
|
|
28
|
-
- When you want to ensure quality gates pass
|
|
29
|
-
- After refactoring
|
|
30
|
-
- To verify that a fix actually closed the reported failure
|
|
31
|
-
|
|
32
|
-
### Positive triggers
|
|
33
|
-
|
|
34
|
-
- "Verify this implementation before handoff."
|
|
35
|
-
- "Confirm the bug is closed in the real CLI."
|
|
36
|
-
- "Run visual and functional QA on the changed flow."
|
|
37
|
-
|
|
38
|
-
### Negative triggers
|
|
39
|
-
|
|
40
|
-
- Pure planning with no artifact to verify.
|
|
41
|
-
- A generic request for more tests with no changed behavior or acceptance criterion in hand.
|
|
42
|
-
- Anything where you would be verifying code you just wrote yourself (see the Verdict Contract).
|
|
43
|
-
|
|
44
|
-
## Pick the real surface first
|
|
45
|
-
|
|
46
|
-
Static gates support verification; they do not constitute it. Before running anything, list the
|
|
47
|
-
acceptance criteria as observable outcomes and select every track the change touches — UI,
|
|
48
|
-
API/service, CLI/TUI, library/SDK, documents/configuration, user flow. Each track has a required
|
|
49
|
-
observation and its own evidence record: read
|
|
50
|
-
[references/tracks.md](references/tracks.md).
|
|
51
|
-
|
|
52
|
-
## Then pick the depth — it scales with the risk
|
|
53
|
-
|
|
54
|
-
The Testing rule says depth follows risk and names what counts as high-risk (authentication,
|
|
55
|
-
authorization, payments and settlement, personal data, data integrity, concurrency, state
|
|
56
|
-
transitions, migrations). It deliberately stops there. **Which** instruments to widen with is a
|
|
57
|
-
per-change judgment, and this is where that menu lives:
|
|
58
|
-
|
|
59
|
-
| Instrument | Reach for it when |
|
|
60
|
-
|---|---|
|
|
61
|
-
| Regression beyond the directly affected scope | the change moves a shared type, a schema, a config default, or anything the affected scope was only *assumed* to bound |
|
|
62
|
-
| Integration / contract tests | it crosses a boundary someone else owns — a service, a queue, a stored format, a published API |
|
|
63
|
-
| Critical-path E2E | a user-visible flow that must not break can only be observed end to end |
|
|
64
|
-
| Mutation testing | **you doubt the tests you already have would catch a defect** |
|
|
65
|
-
|
|
66
|
-
**Mutation testing is an option, not a requirement.** It is expensive, so being labelled
|
|
67
|
-
high-risk is not by itself a reason to run it — the reason is uncertainty about detection power.
|
|
68
|
-
When the existing suite has already been shown to bite (a negative control, a caught regression),
|
|
69
|
-
the doubt it answers is not there and the cost buys nothing.
|
|
70
|
-
|
|
71
|
-
Two things this depth choice is *not*: it is not a coverage target, and it is not the scheduled
|
|
72
|
-
full run. Full regression, full E2E, full mutation, and periodic security scanning belong to the
|
|
73
|
-
CI/CD schedule — do not launch them here for one change.
|
|
74
|
-
|
|
75
|
-
## Verification Phases (static gates)
|
|
76
|
-
|
|
77
|
-
### Phase 1: Build Verification
|
|
78
|
-
```bash
|
|
79
|
-
# Check if project builds
|
|
80
|
-
npm run build 2>&1 | tail -20
|
|
81
|
-
# OR
|
|
82
|
-
pnpm build 2>&1 | tail -20
|
|
83
|
-
```
|
|
84
|
-
|
|
85
|
-
If build fails, STOP and fix — then re-verify from Phase 1 in a fresh instance. Continuing
|
|
86
|
-
through the remaining phases yourself would make you the verifier of code you just wrote,
|
|
87
|
-
which the Verdict Contract below forbids.
|
|
88
|
-
|
|
89
|
-
### Phase 2: Type Check
|
|
90
|
-
```bash
|
|
91
|
-
# TypeScript projects
|
|
92
|
-
npx tsc --noEmit 2>&1 | head -30
|
|
93
|
-
|
|
94
|
-
# Python projects
|
|
95
|
-
pyright . 2>&1 | head -30
|
|
96
|
-
```
|
|
97
|
-
|
|
98
|
-
Report all type errors. Fix critical ones before continuing.
|
|
99
|
-
|
|
100
|
-
### Phase 3: Lint Check
|
|
101
|
-
```bash
|
|
102
|
-
# JavaScript/TypeScript
|
|
103
|
-
npm run lint 2>&1 | head -30
|
|
104
|
-
|
|
105
|
-
# Python
|
|
106
|
-
ruff check . 2>&1 | head -30
|
|
107
|
-
```
|
|
108
|
-
|
|
109
|
-
### Phase 4: Test Suite
|
|
110
|
-
```bash
|
|
111
|
-
# Run tests with coverage
|
|
112
|
-
npm run test -- --coverage 2>&1 | tail -50
|
|
113
|
-
|
|
114
|
-
# Check coverage threshold
|
|
115
|
-
# Target: 80% minimum
|
|
116
|
-
```
|
|
117
|
-
|
|
118
|
-
Report:
|
|
119
|
-
- Total tests: X
|
|
120
|
-
- Passed: X
|
|
121
|
-
- Failed: X
|
|
122
|
-
- Coverage: X%
|
|
123
|
-
|
|
124
|
-
Where the project declares its own threshold, that number wins — 80% is the floor to use when no
|
|
125
|
-
project threshold exists, not a licence to lower one that does.
|
|
126
|
-
|
|
127
|
-
### Phase 5: Security Scan
|
|
128
|
-
```bash
|
|
129
|
-
# Check for secrets
|
|
130
|
-
grep -rn "sk-" --include="*.ts" --include="*.js" . 2>/dev/null | head -10
|
|
131
|
-
grep -rn "api_key" --include="*.ts" --include="*.js" . 2>/dev/null | head -10
|
|
132
|
-
|
|
133
|
-
# Check for console.log
|
|
134
|
-
grep -rn "console.log" --include="*.ts" --include="*.tsx" src/ 2>/dev/null | head -10
|
|
135
|
-
```
|
|
136
|
-
|
|
137
|
-
### Phase 6: Diff Review
|
|
138
|
-
```bash
|
|
139
|
-
# Show what changed
|
|
140
|
-
git diff --stat
|
|
141
|
-
git diff HEAD~1 --name-only
|
|
142
|
-
```
|
|
143
|
-
|
|
144
|
-
Review each changed file for:
|
|
145
|
-
- Unintended changes
|
|
146
|
-
- Missing error handling
|
|
147
|
-
- Potential edge cases
|
|
148
|
-
|
|
149
|
-
Adapt the commands to the repository's stack — the six phases (build, types, lint, tests,
|
|
150
|
-
security, diff) are the contract; `npm`/`pyright`/`ruff` are just this list's defaults.
|
|
151
|
-
|
|
152
|
-
## Then run the live surface
|
|
153
|
-
|
|
154
|
-
Static green with no live run verifies nothing a user can see. For each selected track:
|
|
155
|
-
|
|
156
|
-
- **UI** — browser-driven flow, screenshots, console errors, responsive states, and explicit
|
|
157
|
-
approval before any intentional baseline change.
|
|
158
|
-
- **API/service** — start the service and call the real endpoint; check response *and* side effect.
|
|
159
|
-
- **CLI/TUI** — invoke the built command through its terminal interface; check exit code, stdout,
|
|
160
|
-
and the error path.
|
|
161
|
-
- **Library/SDK** — run a minimal consumer program against the public interface.
|
|
162
|
-
- **Documents/configuration** — parse, resolve references, and exercise the consumer that loads it.
|
|
163
|
-
- **User flow** — complete the representative end-to-end scenario across the affected components.
|
|
164
|
-
|
|
165
|
-
Do not reuse evidence from a different path: one path's green is not another path's evidence.
|
|
166
|
-
|
|
167
|
-
## Output Format
|
|
168
|
-
|
|
169
|
-
After running all phases, produce a verification report:
|
|
170
|
-
|
|
171
|
-
```
|
|
172
|
-
VERIFICATION REPORT
|
|
173
|
-
==================
|
|
174
|
-
|
|
175
|
-
Build: [PASS/FAIL]
|
|
176
|
-
Types: [PASS/FAIL] (X errors)
|
|
177
|
-
Lint: [PASS/FAIL] (X warnings)
|
|
178
|
-
Tests: [PASS/FAIL] (X/Y passed, Z% coverage)
|
|
179
|
-
Security: [PASS/FAIL] (X issues)
|
|
180
|
-
Diff: [X files changed]
|
|
181
|
-
|
|
182
|
-
Live surface: [track] — [command/interaction] → [observed] (artifact: path)
|
|
183
|
-
|
|
184
|
-
Verdict: PASS | PASS_WITH_NITS | FAIL
|
|
185
|
-
|
|
186
|
-
Findings:
|
|
187
|
-
| ID | Severity | Finding | Evidence (file:line / command output) |
|
|
188
|
-
|----|----------|---------|---------------------------------------|
|
|
189
|
-
| F1 | HIGH | ... | ... |
|
|
190
|
-
```
|
|
191
|
-
|
|
192
|
-
Skipped and unverified criteria are listed explicitly, with the reason — silence reads as "passed".
|
|
193
|
-
|
|
194
|
-
## Verdict Contract
|
|
195
|
-
|
|
196
|
-
The report ends with exactly one verdict. Free-prose closings ("looks ready", "should be
|
|
197
|
-
fine") are banned — they leave room to bury defects. A fixed vocabulary makes the report
|
|
198
|
-
honest and machine-checkable.
|
|
199
|
-
|
|
200
|
-
| Verdict | Meaning | Action |
|
|
201
|
-
|---------|---------|--------|
|
|
202
|
-
| **PASS** | All gates green, zero findings at any severity | Ship |
|
|
203
|
-
| **PASS_WITH_NITS** | Ship-safe: only LOW/MEDIUM findings, each recorded with a follow-up | Ship + log follow-ups |
|
|
204
|
-
| **FAIL** | Any gate red, or one or more CRITICAL/HIGH findings | Block → fix → **re-verify** |
|
|
205
|
-
|
|
206
|
-
Every finding gets exactly one severity:
|
|
207
|
-
|
|
208
|
-
- **CRITICAL** — data loss, security hole, or the change misbehaves in real use if shipped
|
|
209
|
-
- **HIGH** — main-path defect or regression; users will hit it
|
|
210
|
-
- **MEDIUM** — edge-case or quality defect; unlikely to block real use
|
|
211
|
-
- **LOW** — nit: style, naming, doc wording
|
|
212
|
-
|
|
213
|
-
Rules:
|
|
214
|
-
- Severity is judged by impact evidence, not by how easy the fix is.
|
|
215
|
-
- FAIL → fix → re-verify is one cycle. A fix alone never upgrades the verdict — the
|
|
216
|
-
re-verification must reproduce green.
|
|
217
|
-
- A run that aborts early (Phase 1 build failure) still emits a report: verdict **FAIL**
|
|
218
|
-
with the failing gate as a CRITICAL finding. Stopping to fix is how you *reach* the next
|
|
219
|
-
verdict, not a reason to skip issuing this one — an unreported run reads as "not run".
|
|
220
|
-
- Missing evidence for a required criterion is a FAIL, not a PASS with a caveat.
|
|
221
|
-
- The instance that wrote the change never issues its own verdict: verification runs in a
|
|
222
|
-
fresh instance (see the model-orchestration skill's V&V separation).
|
|
223
|
-
|
|
224
|
-
## Safety
|
|
225
|
-
|
|
226
|
-
- Never kill broad process patterns, hardcode credentials, or bypass authentication controls.
|
|
227
|
-
- Never update a visual baseline without explicit review of the changed images.
|
|
228
|
-
- Do not claim installed, supported, or tested runtimes that were not actually exercised.
|
|
229
|
-
- Tests, browsers, and services create local artifacts and processes — keep them inside the
|
|
230
|
-
requested workspace and stop the ones you started.
|
|
231
|
-
- Stop before authentication, baseline replacement, or any external mutation the user did not
|
|
232
|
-
authorize.
|
|
233
|
-
|
|
234
|
-
## Continuous Mode
|
|
235
|
-
|
|
236
|
-
For long sessions, run verification every 15 minutes or after major changes:
|
|
237
|
-
|
|
238
|
-
```markdown
|
|
239
|
-
Set a mental checkpoint:
|
|
240
|
-
- After completing each function
|
|
241
|
-
- After finishing a component
|
|
242
|
-
- Before moving to next task
|
|
243
|
-
|
|
244
|
-
Run: /verify
|
|
245
|
-
```
|
|
246
|
-
|
|
247
|
-
## Integration with Hooks
|
|
248
|
-
|
|
249
|
-
This skill complements PostToolUse hooks but provides deeper verification.
|
|
250
|
-
Hooks catch issues immediately; this skill provides comprehensive review.
|
|
@@ -1,49 +0,0 @@
|
|
|
1
|
-
# Verification tracks
|
|
2
|
-
|
|
3
|
-
## Select the real surface
|
|
4
|
-
|
|
5
|
-
| Track | Required observation |
|
|
6
|
-
|---|---|
|
|
7
|
-
| UI | Browser flow, rendered state, console, responsive layout, accessibility-relevant interaction |
|
|
8
|
-
| API or service | Running process, real request, response and side effect |
|
|
9
|
-
| CLI or TUI | Built command, exit code, stdout or screen behavior, error path |
|
|
10
|
-
| Library or SDK | Minimal consumer program using the public interface |
|
|
11
|
-
| Document or configuration | Parser or loader result, resolved references, consumer behavior |
|
|
12
|
-
| User flow | End-to-end outcome across the affected components |
|
|
13
|
-
|
|
14
|
-
Static checks (SKILL.md Phases 1-6) support these tracks but do not replace them. A change that
|
|
15
|
-
touches two tracks needs evidence from both — one track's green is not the other's evidence.
|
|
16
|
-
|
|
17
|
-
## UI baseline policy
|
|
18
|
-
|
|
19
|
-
- Capture deterministic viewports and named states.
|
|
20
|
-
- Compare with the approved baseline using hashes or pixel comparison before semantic review.
|
|
21
|
-
- Treat blank pages, missing core content, and console errors as regressions.
|
|
22
|
-
- Require human approval before adopting an intentional changed baseline.
|
|
23
|
-
- Never launch browsers by killing broad process patterns, and never embed login credentials.
|
|
24
|
-
|
|
25
|
-
## Evidence record
|
|
26
|
-
|
|
27
|
-
For each acceptance criterion record:
|
|
28
|
-
|
|
29
|
-
- exact command or interaction;
|
|
30
|
-
- environment and relevant version;
|
|
31
|
-
- observed outcome;
|
|
32
|
-
- exit status;
|
|
33
|
-
- artifact or screenshot path;
|
|
34
|
-
- `pass` / `fail` / `skipped` / `unverified`;
|
|
35
|
-
- why this evidence covers this criterion.
|
|
36
|
-
|
|
37
|
-
The last field is the one that catches self-deception: evidence that cannot be tied to a criterion
|
|
38
|
-
is output, not proof. `skipped` and `unverified` belong in the report — an omitted criterion reads
|
|
39
|
-
as a passing one.
|
|
40
|
-
|
|
41
|
-
## Verdict
|
|
42
|
-
|
|
43
|
-
- `PASS` — all required gates and user-surface scenarios passed with no findings.
|
|
44
|
-
- `PASS_WITH_NITS` — required behavior passed; only recorded LOW/MEDIUM follow-ups remain.
|
|
45
|
-
- `FAIL` — any required gate failed, a main path is broken, or evidence is missing for a required
|
|
46
|
-
criterion.
|
|
47
|
-
|
|
48
|
-
Severity definitions and the rules that bind a verdict to them are in SKILL.md "Verdict Contract";
|
|
49
|
-
they are not restated here so the two cannot drift apart.
|