@uzysjung/agent-harness 26.151.0 → 26.152.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,306 +0,0 @@
1
- ---
2
- name: eval-harness
3
- description: Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles
4
- origin: ECC
5
- tools: Read, Write, Edit, Bash, Grep, Glob
6
- ---
7
-
8
- # Eval Harness Skill
9
-
10
- A formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles.
11
-
12
- ## When to Activate
13
-
14
- - Setting up eval-driven development (EDD) for AI-assisted workflows
15
- - Defining pass/fail criteria for Claude Code task completion
16
- - Measuring agent reliability with pass@k metrics
17
- - Creating regression test suites for prompt or agent changes
18
- - Benchmarking agent performance across model versions
19
-
20
- ## Philosophy
21
-
22
- Eval-Driven Development treats evals as the "unit tests of AI development":
23
- - Define expected behavior BEFORE implementation
24
- - Run evals continuously during development
25
- - Track regressions with each change
26
- - Use pass@k metrics for reliability measurement
27
-
28
- ## Eval Types
29
-
30
- ### Capability Evals
31
- Test if Claude can do something it couldn't before:
32
- ```markdown
33
- [CAPABILITY EVAL: feature-name]
34
- Task: Description of what Claude should accomplish
35
- Success Criteria:
36
- - [ ] Criterion 1
37
- - [ ] Criterion 2
38
- - [ ] Criterion 3
39
- Expected Output: Description of expected result
40
- ```
41
-
42
- ### Regression Evals
43
- Ensure changes don't break existing functionality:
44
- ```markdown
45
- [REGRESSION EVAL: feature-name]
46
- Baseline: SHA or checkpoint name
47
- Tests:
48
- - existing-test-1: PASS/FAIL
49
- - existing-test-2: PASS/FAIL
50
- - existing-test-3: PASS/FAIL
51
- Result: X/Y passed (previously Y/Y)
52
- ```
53
-
54
- ## Grader Types
55
-
56
- ### 1. Code-Based Grader
57
- Deterministic checks using code:
58
- ```bash
59
- # Check if file contains expected pattern
60
- grep -q "export function handleAuth" src/auth.ts && echo "PASS" || echo "FAIL"
61
-
62
- # Check if tests pass
63
- npm test -- --testPathPattern="auth" && echo "PASS" || echo "FAIL"
64
-
65
- # Check if build succeeds
66
- npm run build && echo "PASS" || echo "FAIL"
67
- ```
68
-
69
- ### 2. Model-Based Grader
70
- Use Claude to evaluate open-ended outputs:
71
- ```markdown
72
- [MODEL GRADER PROMPT]
73
- Evaluate the following code change:
74
- 1. Does it solve the stated problem?
75
- 2. Is it well-structured?
76
- 3. Are edge cases handled?
77
- 4. Is error handling appropriate?
78
-
79
- Score: 1-5 (1=poor, 5=excellent)
80
- Reasoning: [explanation]
81
- ```
82
-
83
- ### 3. Human Grader
84
- Flag for manual review:
85
- ```markdown
86
- [HUMAN REVIEW REQUIRED]
87
- Change: Description of what changed
88
- Reason: Why human review is needed
89
- Risk Level: LOW/MEDIUM/HIGH
90
- ```
91
-
92
- ## Metrics
93
-
94
- ### pass@k
95
- "At least one success in k attempts"
96
- - pass@1: First attempt success rate
97
- - pass@3: Success within 3 attempts
98
- - Typical target: pass@3 > 90%
99
-
100
- ### pass^k
101
- "All k trials succeed"
102
- - Higher bar for reliability
103
- - pass^3: 3 consecutive successes
104
- - Use for critical paths
105
-
106
- ## Eval Workflow
107
-
108
- ### 1. Define (Before Coding)
109
-
110
- Write the spec to a file (`.claude/evals/<feature>.md`) before implementing. Give every eval a
111
- stable ID — `C1..Cn` for capability, `R1..Rn` for regression — so the same identifier carries
112
- from definition to the post-implementation status line, and a reviewer can check them off one by
113
- one. Prose-only lists ("Can create new user account") can't be referenced or scored.
114
-
115
- ````markdown
116
- # EVAL: <feature name> (<phase/PR>)
117
-
118
- **Feature**: <one line>
119
- **Baseline**: commit <sha> # what "regression" is measured against
120
- **Target**: pass@1 = 100% for capability evals
121
-
122
- ## Capability Evals
123
-
124
- ### C1: <name>
125
- - <concrete, checkable expectation — inputs → expected output>
126
-
127
- ### C2: <name>
128
- - <expectation>
129
-
130
- ## Regression Evals
131
-
132
- ### R1: <existing behavior that must not move>
133
- - <expectation>
134
-
135
- ## Test Command
136
- ```bash
137
- pytest tests/test_<area>.py -v -k "<selector>"
138
- ```
139
-
140
- ## Status (after implementation)
141
- - C1-Cn: PASS via <tests (N개) / route check / manual>
142
- - R1-Rn: PASS
143
-
144
- **Overall: pass@1 = <x>%**
145
- ````
146
-
147
- Two fields carry most of the weight. **Baseline commit** makes "regression" falsifiable — without
148
- it, R-evals are opinions about the past. **Test Command** makes the spec re-runnable by someone
149
- who didn't write it; an eval nobody can re-run is documentation, not a gate.
150
-
151
- Fill the Status section *after* implementing, in the same file. A spec whose status is still empty
152
- at merge time means the evals were written and never used.
153
-
154
- ### 2. Implement
155
- Write code to pass the defined evals.
156
-
157
- ### 3. Evaluate
158
- ```bash
159
- # Run capability evals
160
- [Run each capability eval, record PASS/FAIL]
161
-
162
- # Run regression evals
163
- npm test -- --testPathPattern="existing"
164
-
165
- # Generate report
166
- ```
167
-
168
- ### 4. Report
169
- ```markdown
170
- EVAL REPORT: feature-xyz
171
- ========================
172
-
173
- Capability Evals:
174
- create-user: PASS (pass@1)
175
- validate-email: PASS (pass@2)
176
- hash-password: PASS (pass@1)
177
- Overall: 3/3 passed
178
-
179
- Regression Evals:
180
- login-flow: PASS
181
- session-mgmt: PASS
182
- logout-flow: PASS
183
- Overall: 3/3 passed
184
-
185
- Metrics:
186
- pass@1: 67% (2/3)
187
- pass@3: 100% (3/3)
188
-
189
- Status: READY FOR REVIEW
190
- ```
191
-
192
- ## Integration Patterns
193
-
194
- ### Pre-Implementation
195
- ```
196
- /eval define feature-name
197
- ```
198
- Creates eval definition file at `.claude/evals/feature-name.md`
199
-
200
- ### During Implementation
201
- ```
202
- /eval check feature-name
203
- ```
204
- Runs current evals and reports status
205
-
206
- ### Post-Implementation
207
- ```
208
- /eval report feature-name
209
- ```
210
- Generates full eval report
211
-
212
- ## Eval Storage (.md + .log Pair Format)
213
-
214
- 각 평가 항목은 **`<topic>.md` (설계) + `<topic>.log` (실행 결과)** 쌍으로 저장. 강제. 단독 .md만 있으면 재현 불가.
215
-
216
- ```
217
- .claude/
218
- evals/
219
- feature-xyz.md # Eval definition (Capability/Regression/Test 3섹션 필수)
220
- feature-xyz.log # Eval run history (실행 시각, grader, pass/fail)
221
- session-YYYYMMDD.md # 세션 단위 회고 + 차기 backlog
222
- session-YYYYMMDD.log # 동일 세션의 grader 출력
223
- baseline.json # Regression baselines (선택)
224
- ```
225
-
226
- > eval 산출물은 `docs/evals/*.{md,log}` 로 모은다 — 실행 로그와 판정을 같은 자리에 둔다.
227
-
228
- ### .md 파일 의무 섹션 (3개)
229
-
230
- ```markdown
231
- # Eval: <topic>
232
-
233
- ## Capability
234
- [새 능력 — Claude/agent가 무엇을 할 수 있는지]
235
- - AC: [측정 가능 기준]
236
- - Grader: code-based / model-based / human
237
-
238
- ## Regression
239
- [기존 기능 보호 — 변경으로 깨지면 안 되는 baseline]
240
- - Baseline: <SHA or checkpoint>
241
- - Tests: [목록]
242
-
243
- ## Test
244
- [실행 절차 — 누가 다시 돌려도 동일 결과 나와야 함]
245
- - Setup: [사전 조건]
246
- - Run: `bash run-eval.sh <topic>` 또는 명시적 명령
247
- - Expected: [기대 출력]
248
- ```
249
-
250
- ### .log 파일 형식
251
-
252
- 각 실행마다 append. 시간순 누적.
253
-
254
- ```
255
- === 2026-04-19 14:32 (run #1) ===
256
- Capability: 3/3 PASS (pass@1)
257
- Regression: 5/5 PASS (pass^3)
258
- Status: SHIP READY
259
-
260
- === 2026-04-20 09:15 (run #2 — after refactor) ===
261
- Capability: 3/3 PASS
262
- Regression: 4/5 PASS (login-flow regressed at SHA abc123)
263
- Status: BLOCKED — fix login-flow first
264
- ```
265
-
266
- ## Best Practices
267
-
268
- 1. **Define evals BEFORE coding** - Forces clear thinking about success criteria
269
- 2. **Run evals frequently** - Catch regressions early
270
- 3. **Track pass@k over time** - Monitor reliability trends
271
- 4. **Use code graders when possible** - Deterministic > probabilistic
272
- 5. **Human review for security** - Never fully automate security checks
273
- 6. **Keep evals fast** - Slow evals don't get run
274
- 7. **Version evals with code** - Evals are first-class artifacts
275
-
276
- ## Example: Adding Authentication
277
-
278
- ```markdown
279
- ## EVAL: add-authentication
280
-
281
- ### Phase 1: Define (10 min)
282
- Capability Evals:
283
- - [ ] User can register with email/password
284
- - [ ] User can login with valid credentials
285
- - [ ] Invalid credentials rejected with proper error
286
- - [ ] Sessions persist across page reloads
287
- - [ ] Logout clears session
288
-
289
- Regression Evals:
290
- - [ ] Public routes still accessible
291
- - [ ] API responses unchanged
292
- - [ ] Database schema compatible
293
-
294
- ### Phase 2: Implement (varies)
295
- [Write code]
296
-
297
- ### Phase 3: Evaluate
298
- Run: /eval check add-authentication
299
-
300
- ### Phase 4: Report
301
- EVAL REPORT: add-authentication
302
- ==============================
303
- Capability: 5/5 passed (pass@3: 100%)
304
- Regression: 3/3 passed (pass^3: 100%)
305
- Status: SHIP IT
306
- ```
@@ -1,7 +0,0 @@
1
- interface:
2
- display_name: "Eval Harness"
3
- short_description: "Eval-driven development with pass/fail criteria"
4
- brand_color: "#EC4899"
5
- default_prompt: "Set up eval-driven development with pass/fail criteria"
6
- policy:
7
- allow_implicit_invocation: true
@@ -1,250 +0,0 @@
1
- ---
2
- name: verification-loop
3
- description: >-
4
- A comprehensive verification system. Selects and runs proportional verification tracks for UI,
5
- API/service, CLI/TUI, library/SDK, documents/configuration, and real user flows, then ends every
6
- run with a fixed verdict — PASS / PASS_WITH_NITS / FAIL — plus severity-labeled findings
7
- (CRITICAL/HIGH/MEDIUM/LOW) and the evidence each one rests on. Use after implementation, before
8
- a PR or handoff, after a refactor, or to verify a claimed fix. Do NOT use a green build, a passing
9
- type check, or file existence as proof of user-visible completion, and do NOT let the instance
10
- that wrote the change issue its own verdict.
11
- origin: ECC
12
- ---
13
-
14
- # Verification Loop Skill
15
-
16
- > Derived from the `verification-loop` skill in everything-claude-code (ECC), used under the MIT
17
- > License.
18
-
19
- A comprehensive verification system for coding sessions. The job is not "run the gates" — it is to
20
- **verify the changed behavior through the surface a real user or consumer actually uses**, and to
21
- end with one verdict that cannot be softened into prose.
22
-
23
- ## When to Use
24
-
25
- Invoke this skill:
26
- - After completing a feature or significant code change
27
- - Before creating a PR
28
- - When you want to ensure quality gates pass
29
- - After refactoring
30
- - To verify that a fix actually closed the reported failure
31
-
32
- ### Positive triggers
33
-
34
- - "Verify this implementation before handoff."
35
- - "Confirm the bug is closed in the real CLI."
36
- - "Run visual and functional QA on the changed flow."
37
-
38
- ### Negative triggers
39
-
40
- - Pure planning with no artifact to verify.
41
- - A generic request for more tests with no changed behavior or acceptance criterion in hand.
42
- - Anything where you would be verifying code you just wrote yourself (see the Verdict Contract).
43
-
44
- ## Pick the real surface first
45
-
46
- Static gates support verification; they do not constitute it. Before running anything, list the
47
- acceptance criteria as observable outcomes and select every track the change touches — UI,
48
- API/service, CLI/TUI, library/SDK, documents/configuration, user flow. Each track has a required
49
- observation and its own evidence record: read
50
- [references/tracks.md](references/tracks.md).
51
-
52
- ## Then pick the depth — it scales with the risk
53
-
54
- The Testing rule says depth follows risk and names what counts as high-risk (authentication,
55
- authorization, payments and settlement, personal data, data integrity, concurrency, state
56
- transitions, migrations). It deliberately stops there. **Which** instruments to widen with is a
57
- per-change judgment, and this is where that menu lives:
58
-
59
- | Instrument | Reach for it when |
60
- |---|---|
61
- | Regression beyond the directly affected scope | the change moves a shared type, a schema, a config default, or anything the affected scope was only *assumed* to bound |
62
- | Integration / contract tests | it crosses a boundary someone else owns — a service, a queue, a stored format, a published API |
63
- | Critical-path E2E | a user-visible flow that must not break can only be observed end to end |
64
- | Mutation testing | **you doubt the tests you already have would catch a defect** |
65
-
66
- **Mutation testing is an option, not a requirement.** It is expensive, so being labelled
67
- high-risk is not by itself a reason to run it — the reason is uncertainty about detection power.
68
- When the existing suite has already been shown to bite (a negative control, a caught regression),
69
- the doubt it answers is not there and the cost buys nothing.
70
-
71
- Two things this depth choice is *not*: it is not a coverage target, and it is not the scheduled
72
- full run. Full regression, full E2E, full mutation, and periodic security scanning belong to the
73
- CI/CD schedule — do not launch them here for one change.
74
-
75
- ## Verification Phases (static gates)
76
-
77
- ### Phase 1: Build Verification
78
- ```bash
79
- # Check if project builds
80
- npm run build 2>&1 | tail -20
81
- # OR
82
- pnpm build 2>&1 | tail -20
83
- ```
84
-
85
- If build fails, STOP and fix — then re-verify from Phase 1 in a fresh instance. Continuing
86
- through the remaining phases yourself would make you the verifier of code you just wrote,
87
- which the Verdict Contract below forbids.
88
-
89
- ### Phase 2: Type Check
90
- ```bash
91
- # TypeScript projects
92
- npx tsc --noEmit 2>&1 | head -30
93
-
94
- # Python projects
95
- pyright . 2>&1 | head -30
96
- ```
97
-
98
- Report all type errors. Fix critical ones before continuing.
99
-
100
- ### Phase 3: Lint Check
101
- ```bash
102
- # JavaScript/TypeScript
103
- npm run lint 2>&1 | head -30
104
-
105
- # Python
106
- ruff check . 2>&1 | head -30
107
- ```
108
-
109
- ### Phase 4: Test Suite
110
- ```bash
111
- # Run tests with coverage
112
- npm run test -- --coverage 2>&1 | tail -50
113
-
114
- # Check coverage threshold
115
- # Target: 80% minimum
116
- ```
117
-
118
- Report:
119
- - Total tests: X
120
- - Passed: X
121
- - Failed: X
122
- - Coverage: X%
123
-
124
- Where the project declares its own threshold, that number wins — 80% is the floor to use when no
125
- project threshold exists, not a licence to lower one that does.
126
-
127
- ### Phase 5: Security Scan
128
- ```bash
129
- # Check for secrets
130
- grep -rn "sk-" --include="*.ts" --include="*.js" . 2>/dev/null | head -10
131
- grep -rn "api_key" --include="*.ts" --include="*.js" . 2>/dev/null | head -10
132
-
133
- # Check for console.log
134
- grep -rn "console.log" --include="*.ts" --include="*.tsx" src/ 2>/dev/null | head -10
135
- ```
136
-
137
- ### Phase 6: Diff Review
138
- ```bash
139
- # Show what changed
140
- git diff --stat
141
- git diff HEAD~1 --name-only
142
- ```
143
-
144
- Review each changed file for:
145
- - Unintended changes
146
- - Missing error handling
147
- - Potential edge cases
148
-
149
- Adapt the commands to the repository's stack — the six phases (build, types, lint, tests,
150
- security, diff) are the contract; `npm`/`pyright`/`ruff` are just this list's defaults.
151
-
152
- ## Then run the live surface
153
-
154
- Static green with no live run verifies nothing a user can see. For each selected track:
155
-
156
- - **UI** — browser-driven flow, screenshots, console errors, responsive states, and explicit
157
- approval before any intentional baseline change.
158
- - **API/service** — start the service and call the real endpoint; check response *and* side effect.
159
- - **CLI/TUI** — invoke the built command through its terminal interface; check exit code, stdout,
160
- and the error path.
161
- - **Library/SDK** — run a minimal consumer program against the public interface.
162
- - **Documents/configuration** — parse, resolve references, and exercise the consumer that loads it.
163
- - **User flow** — complete the representative end-to-end scenario across the affected components.
164
-
165
- Do not reuse evidence from a different path: one path's green is not another path's evidence.
166
-
167
- ## Output Format
168
-
169
- After running all phases, produce a verification report:
170
-
171
- ```
172
- VERIFICATION REPORT
173
- ==================
174
-
175
- Build: [PASS/FAIL]
176
- Types: [PASS/FAIL] (X errors)
177
- Lint: [PASS/FAIL] (X warnings)
178
- Tests: [PASS/FAIL] (X/Y passed, Z% coverage)
179
- Security: [PASS/FAIL] (X issues)
180
- Diff: [X files changed]
181
-
182
- Live surface: [track] — [command/interaction] → [observed] (artifact: path)
183
-
184
- Verdict: PASS | PASS_WITH_NITS | FAIL
185
-
186
- Findings:
187
- | ID | Severity | Finding | Evidence (file:line / command output) |
188
- |----|----------|---------|---------------------------------------|
189
- | F1 | HIGH | ... | ... |
190
- ```
191
-
192
- Skipped and unverified criteria are listed explicitly, with the reason — silence reads as "passed".
193
-
194
- ## Verdict Contract
195
-
196
- The report ends with exactly one verdict. Free-prose closings ("looks ready", "should be
197
- fine") are banned — they leave room to bury defects. A fixed vocabulary makes the report
198
- honest and machine-checkable.
199
-
200
- | Verdict | Meaning | Action |
201
- |---------|---------|--------|
202
- | **PASS** | All gates green, zero findings at any severity | Ship |
203
- | **PASS_WITH_NITS** | Ship-safe: only LOW/MEDIUM findings, each recorded with a follow-up | Ship + log follow-ups |
204
- | **FAIL** | Any gate red, or one or more CRITICAL/HIGH findings | Block → fix → **re-verify** |
205
-
206
- Every finding gets exactly one severity:
207
-
208
- - **CRITICAL** — data loss, security hole, or the change misbehaves in real use if shipped
209
- - **HIGH** — main-path defect or regression; users will hit it
210
- - **MEDIUM** — edge-case or quality defect; unlikely to block real use
211
- - **LOW** — nit: style, naming, doc wording
212
-
213
- Rules:
214
- - Severity is judged by impact evidence, not by how easy the fix is.
215
- - FAIL → fix → re-verify is one cycle. A fix alone never upgrades the verdict — the
216
- re-verification must reproduce green.
217
- - A run that aborts early (Phase 1 build failure) still emits a report: verdict **FAIL**
218
- with the failing gate as a CRITICAL finding. Stopping to fix is how you *reach* the next
219
- verdict, not a reason to skip issuing this one — an unreported run reads as "not run".
220
- - Missing evidence for a required criterion is a FAIL, not a PASS with a caveat.
221
- - The instance that wrote the change never issues its own verdict: verification runs in a
222
- fresh instance (see the model-orchestration skill's V&V separation).
223
-
224
- ## Safety
225
-
226
- - Never kill broad process patterns, hardcode credentials, or bypass authentication controls.
227
- - Never update a visual baseline without explicit review of the changed images.
228
- - Do not claim installed, supported, or tested runtimes that were not actually exercised.
229
- - Tests, browsers, and services create local artifacts and processes — keep them inside the
230
- requested workspace and stop the ones you started.
231
- - Stop before authentication, baseline replacement, or any external mutation the user did not
232
- authorize.
233
-
234
- ## Continuous Mode
235
-
236
- For long sessions, run verification every 15 minutes or after major changes:
237
-
238
- ```markdown
239
- Set a mental checkpoint:
240
- - After completing each function
241
- - After finishing a component
242
- - Before moving to next task
243
-
244
- Run: /verify
245
- ```
246
-
247
- ## Integration with Hooks
248
-
249
- This skill complements PostToolUse hooks but provides deeper verification.
250
- Hooks catch issues immediately; this skill provides comprehensive review.
@@ -1,7 +0,0 @@
1
- interface:
2
- display_name: "Verification Loop"
3
- short_description: "Build, test, lint, typecheck verification"
4
- brand_color: "#10B981"
5
- default_prompt: "Run verification: build, test, lint, typecheck, security"
6
- policy:
7
- allow_implicit_invocation: true
@@ -1,49 +0,0 @@
1
- # Verification tracks
2
-
3
- ## Select the real surface
4
-
5
- | Track | Required observation |
6
- |---|---|
7
- | UI | Browser flow, rendered state, console, responsive layout, accessibility-relevant interaction |
8
- | API or service | Running process, real request, response and side effect |
9
- | CLI or TUI | Built command, exit code, stdout or screen behavior, error path |
10
- | Library or SDK | Minimal consumer program using the public interface |
11
- | Document or configuration | Parser or loader result, resolved references, consumer behavior |
12
- | User flow | End-to-end outcome across the affected components |
13
-
14
- Static checks (SKILL.md Phases 1-6) support these tracks but do not replace them. A change that
15
- touches two tracks needs evidence from both — one track's green is not the other's evidence.
16
-
17
- ## UI baseline policy
18
-
19
- - Capture deterministic viewports and named states.
20
- - Compare with the approved baseline using hashes or pixel comparison before semantic review.
21
- - Treat blank pages, missing core content, and console errors as regressions.
22
- - Require human approval before adopting an intentional changed baseline.
23
- - Never launch browsers by killing broad process patterns, and never embed login credentials.
24
-
25
- ## Evidence record
26
-
27
- For each acceptance criterion record:
28
-
29
- - exact command or interaction;
30
- - environment and relevant version;
31
- - observed outcome;
32
- - exit status;
33
- - artifact or screenshot path;
34
- - `pass` / `fail` / `skipped` / `unverified`;
35
- - why this evidence covers this criterion.
36
-
37
- The last field is the one that catches self-deception: evidence that cannot be tied to a criterion
38
- is output, not proof. `skipped` and `unverified` belong in the report — an omitted criterion reads
39
- as a passing one.
40
-
41
- ## Verdict
42
-
43
- - `PASS` — all required gates and user-surface scenarios passed with no findings.
44
- - `PASS_WITH_NITS` — required behavior passed; only recorded LOW/MEDIUM follow-ups remain.
45
- - `FAIL` — any required gate failed, a main path is broken, or evidence is missing for a required
46
- criterion.
47
-
48
- Severity definitions and the rules that bind a verdict to them are in SKILL.md "Verdict Contract";
49
- they are not restated here so the two cannot drift apart.