@rune-kit/rune 2.10.0 → 2.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (240) hide show
  1. package/LICENSE +21 -21
  2. package/README.md +65 -6
  3. package/commands/rune.md +168 -168
  4. package/compiler/__tests__/detect-invariants.test.js +136 -0
  5. package/compiler/__tests__/doctor-mesh.test.js +229 -0
  6. package/compiler/__tests__/hook-dispatch.test.js +91 -0
  7. package/compiler/__tests__/hooks-antigravity.test.js +118 -0
  8. package/compiler/__tests__/hooks-cursor.test.js +139 -0
  9. package/compiler/__tests__/hooks-install.test.js +305 -0
  10. package/compiler/__tests__/hooks-merge.test.js +204 -0
  11. package/compiler/__tests__/hooks-tiers.test.js +519 -0
  12. package/compiler/__tests__/hooks-windsurf.test.js +115 -0
  13. package/compiler/__tests__/inject-claude-md.test.js +152 -0
  14. package/compiler/__tests__/load-invariants.test.js +408 -0
  15. package/compiler/__tests__/onboard-invariants.test.js +240 -0
  16. package/compiler/adapters/hooks/antigravity.js +140 -0
  17. package/compiler/adapters/hooks/claude.js +166 -0
  18. package/compiler/adapters/hooks/cursor.js +191 -0
  19. package/compiler/adapters/hooks/index.js +82 -0
  20. package/compiler/adapters/hooks/tier-emitter.js +182 -0
  21. package/compiler/adapters/hooks/windsurf.js +202 -0
  22. package/compiler/bin/rune.js +196 -6
  23. package/compiler/commands/hook-dispatch.js +87 -0
  24. package/compiler/commands/hooks/install.js +120 -0
  25. package/compiler/commands/hooks/merge.js +211 -0
  26. package/compiler/commands/hooks/presets.js +116 -0
  27. package/compiler/commands/hooks/status.js +112 -0
  28. package/compiler/commands/hooks/tiers.js +221 -0
  29. package/compiler/commands/hooks/uninstall.js +94 -0
  30. package/compiler/doctor.js +236 -0
  31. package/contexts/dev.md +34 -34
  32. package/contexts/research.md +43 -43
  33. package/contexts/review.md +55 -55
  34. package/extensions/ai-ml/PACK.md +88 -88
  35. package/extensions/ai-ml/skills/ai-agents.md +172 -172
  36. package/extensions/ai-ml/skills/code-sandbox.md +187 -187
  37. package/extensions/ai-ml/skills/deep-research.md +146 -146
  38. package/extensions/ai-ml/skills/embedding-search.md +66 -66
  39. package/extensions/ai-ml/skills/fine-tuning-guide.md +74 -74
  40. package/extensions/ai-ml/skills/llm-architect.md +125 -125
  41. package/extensions/ai-ml/skills/llm-integration.md +64 -64
  42. package/extensions/ai-ml/skills/prompt-patterns.md +72 -72
  43. package/extensions/ai-ml/skills/rag-patterns.md +66 -66
  44. package/extensions/ai-ml/skills/web-extraction.md +114 -114
  45. package/extensions/analytics/PACK.md +92 -92
  46. package/extensions/analytics/skills/ab-testing.md +72 -72
  47. package/extensions/analytics/skills/dashboard-patterns.md +83 -83
  48. package/extensions/analytics/skills/data-validation.md +68 -68
  49. package/extensions/analytics/skills/funnel-analysis.md +81 -81
  50. package/extensions/analytics/skills/sql-patterns.md +57 -57
  51. package/extensions/analytics/skills/statistical-analysis.md +79 -79
  52. package/extensions/analytics/skills/tracking-setup.md +71 -71
  53. package/extensions/backend/PACK.md +104 -104
  54. package/extensions/backend/skills/api-patterns.md +84 -84
  55. package/extensions/backend/skills/async-pipeline.md +193 -193
  56. package/extensions/backend/skills/auth-patterns.md +97 -97
  57. package/extensions/backend/skills/background-jobs.md +133 -133
  58. package/extensions/backend/skills/caching-patterns.md +108 -108
  59. package/extensions/backend/skills/cli-generation.md +133 -133
  60. package/extensions/backend/skills/database-patterns.md +87 -87
  61. package/extensions/backend/skills/middleware-patterns.md +104 -104
  62. package/extensions/chrome-ext/PACK.md +93 -93
  63. package/extensions/chrome-ext/skills/cws-preflight.md +143 -143
  64. package/extensions/chrome-ext/skills/cws-publish.md +104 -104
  65. package/extensions/chrome-ext/skills/ext-ai-integration.md +251 -251
  66. package/extensions/chrome-ext/skills/ext-messaging.md +139 -139
  67. package/extensions/chrome-ext/skills/ext-storage.md +133 -133
  68. package/extensions/chrome-ext/skills/mv3-scaffold.md +164 -164
  69. package/extensions/content/PACK.md +96 -96
  70. package/extensions/content/skills/blog-patterns.md +88 -88
  71. package/extensions/content/skills/cms-integration.md +131 -131
  72. package/extensions/content/skills/content-scoring.md +107 -107
  73. package/extensions/content/skills/i18n.md +83 -83
  74. package/extensions/content/skills/mdx-authoring.md +137 -137
  75. package/extensions/content/skills/reference.md +1014 -1014
  76. package/extensions/content/skills/seo-patterns.md +67 -67
  77. package/extensions/content/skills/video-repurpose.md +153 -153
  78. package/extensions/devops/PACK.md +101 -101
  79. package/extensions/devops/skills/chaos-testing.md +67 -67
  80. package/extensions/devops/skills/ci-cd.md +75 -75
  81. package/extensions/devops/skills/docker.md +58 -58
  82. package/extensions/devops/skills/edge-serverless.md +163 -163
  83. package/extensions/devops/skills/infra-as-code.md +158 -158
  84. package/extensions/devops/skills/kubernetes.md +110 -110
  85. package/extensions/devops/skills/monitoring.md +57 -57
  86. package/extensions/devops/skills/server-setup.md +64 -64
  87. package/extensions/devops/skills/ssl-domain.md +42 -42
  88. package/extensions/ecommerce/PACK.md +116 -116
  89. package/extensions/ecommerce/skills/cart-system.md +79 -79
  90. package/extensions/ecommerce/skills/inventory-mgmt.md +102 -102
  91. package/extensions/ecommerce/skills/order-management.md +126 -126
  92. package/extensions/ecommerce/skills/payment-integration.md +472 -472
  93. package/extensions/ecommerce/skills/shopify-dev.md +69 -69
  94. package/extensions/ecommerce/skills/subscription-billing.md +93 -93
  95. package/extensions/ecommerce/skills/tax-compliance.md +117 -117
  96. package/extensions/gamedev/PACK.md +142 -142
  97. package/extensions/gamedev/skills/asset-pipeline.md +74 -74
  98. package/extensions/gamedev/skills/audio-system.md +129 -129
  99. package/extensions/gamedev/skills/camera-system.md +87 -87
  100. package/extensions/gamedev/skills/ecs.md +98 -98
  101. package/extensions/gamedev/skills/game-loops.md +72 -72
  102. package/extensions/gamedev/skills/input-system.md +199 -199
  103. package/extensions/gamedev/skills/multiplayer.md +180 -180
  104. package/extensions/gamedev/skills/particles.md +105 -105
  105. package/extensions/gamedev/skills/physics-engine.md +89 -89
  106. package/extensions/gamedev/skills/scene-management.md +146 -146
  107. package/extensions/gamedev/skills/threejs-patterns.md +90 -90
  108. package/extensions/gamedev/skills/webgl.md +71 -71
  109. package/extensions/mobile/PACK.md +106 -106
  110. package/extensions/mobile/skills/app-store-connect.md +152 -152
  111. package/extensions/mobile/skills/app-store-prep.md +66 -66
  112. package/extensions/mobile/skills/deep-linking.md +109 -109
  113. package/extensions/mobile/skills/flutter.md +60 -60
  114. package/extensions/mobile/skills/ios-build-pipeline.md +142 -142
  115. package/extensions/mobile/skills/native-bridge.md +66 -66
  116. package/extensions/mobile/skills/ota-updates.md +97 -97
  117. package/extensions/mobile/skills/push-notifications.md +111 -111
  118. package/extensions/mobile/skills/react-native.md +82 -82
  119. package/extensions/saas/PACK.md +116 -116
  120. package/extensions/saas/skills/billing-integration.md +200 -200
  121. package/extensions/saas/skills/feature-flags.md +130 -130
  122. package/extensions/saas/skills/multi-tenant.md +103 -103
  123. package/extensions/saas/skills/onboarding-flow.md +139 -139
  124. package/extensions/saas/skills/subscription-flow.md +95 -95
  125. package/extensions/saas/skills/team-management.md +144 -144
  126. package/extensions/security/PACK.md +99 -99
  127. package/extensions/security/skills/api-security.md +140 -140
  128. package/extensions/security/skills/compliance.md +68 -68
  129. package/extensions/security/skills/owasp-audit.md +64 -64
  130. package/extensions/security/skills/pentest-patterns.md +77 -77
  131. package/extensions/security/skills/secret-mgmt.md +65 -65
  132. package/extensions/security/skills/supply-chain.md +65 -65
  133. package/extensions/trading/PACK.md +80 -80
  134. package/extensions/trading/skills/chart-components.md +55 -55
  135. package/extensions/trading/skills/experiment-loop.md +125 -125
  136. package/extensions/trading/skills/fintech-patterns.md +47 -47
  137. package/extensions/trading/skills/indicator-library.md +58 -58
  138. package/extensions/trading/skills/quant-analysis.md +111 -111
  139. package/extensions/trading/skills/realtime-data.md +58 -58
  140. package/extensions/trading/skills/trade-logic.md +104 -104
  141. package/extensions/ui/PACK.md +130 -130
  142. package/extensions/ui/skills/a11y-audit.md +91 -91
  143. package/extensions/ui/skills/animation-patterns.md +127 -127
  144. package/extensions/ui/skills/component-patterns.md +100 -100
  145. package/extensions/ui/skills/design-decision.md +108 -108
  146. package/extensions/ui/skills/design-system.md +68 -68
  147. package/extensions/ui/skills/landing-patterns.md +155 -155
  148. package/extensions/ui/skills/palette-picker.md +173 -173
  149. package/extensions/ui/skills/react-health.md +90 -90
  150. package/extensions/ui/skills/type-system.md +125 -125
  151. package/extensions/ui/skills/web-vitals.md +153 -153
  152. package/extensions/zalo/PACK.md +145 -145
  153. package/extensions/zalo/skills/zalo-oa-mcp.md +317 -317
  154. package/extensions/zalo/skills/zalo-oa-messaging.md +429 -429
  155. package/extensions/zalo/skills/zalo-oa-setup.md +236 -236
  156. package/extensions/zalo/skills/zalo-oa-webhook.md +189 -189
  157. package/extensions/zalo/skills/zalo-personal-messaging.md +194 -194
  158. package/extensions/zalo/skills/zalo-personal-setup.md +153 -153
  159. package/extensions/zalo/skills/zalo-rate-guard.md +219 -219
  160. package/hooks/auto-format/index.cjs +48 -48
  161. package/hooks/hooks.json +111 -111
  162. package/hooks/post-session-reflect/index.cjs +189 -189
  163. package/hooks/pre-compact/index.cjs +95 -95
  164. package/hooks/run-hook.cmd +1 -1
  165. package/hooks/secrets-scan/index.cjs +100 -100
  166. package/hooks/session-start/index.cjs +71 -71
  167. package/hooks/typecheck/index.cjs +65 -65
  168. package/package.json +63 -63
  169. package/references/ui-pro-max-data/LICENSE-UI-PRO-MAX +21 -21
  170. package/references/ui-pro-max-data/charts.csv +26 -26
  171. package/references/ui-pro-max-data/colors.csv +161 -161
  172. package/references/ui-pro-max-data/styles.csv +68 -68
  173. package/references/ui-pro-max-data/typography.csv +74 -74
  174. package/references/ui-pro-max-data/ui-reasoning.csv +162 -162
  175. package/references/ui-pro-max-data/ux-guidelines.csv +99 -99
  176. package/skills/adversary/SKILL.md +283 -283
  177. package/skills/asset-creator/SKILL.md +157 -157
  178. package/skills/audit/SKILL.md +147 -2
  179. package/skills/autopsy/SKILL.md +335 -335
  180. package/skills/ba/SKILL.md +85 -1
  181. package/skills/brainstorm/SKILL.md +380 -342
  182. package/skills/browser-pilot/SKILL.md +169 -168
  183. package/skills/constraint-check/SKILL.md +165 -165
  184. package/skills/context-engine/SKILL.md +408 -404
  185. package/skills/cook/SKILL.md +917 -863
  186. package/skills/db/SKILL.md +273 -273
  187. package/skills/debug/SKILL.md +465 -465
  188. package/skills/dependency-doctor/SKILL.md +265 -235
  189. package/skills/deploy/SKILL.md +274 -231
  190. package/skills/design/DESIGN-REFERENCE.md +365 -365
  191. package/skills/design/SKILL.md +590 -589
  192. package/skills/doc-processor/SKILL.md +254 -254
  193. package/skills/docs/SKILL.md +374 -374
  194. package/skills/docs-seeker/SKILL.md +178 -177
  195. package/skills/fix/SKILL.md +332 -330
  196. package/skills/git/SKILL.md +339 -339
  197. package/skills/hallucination-guard/SKILL.md +220 -219
  198. package/skills/incident/SKILL.md +254 -253
  199. package/skills/integrity-check/SKILL.md +169 -169
  200. package/skills/journal/SKILL.md +241 -240
  201. package/skills/launch/SKILL.md +344 -344
  202. package/skills/logic-guardian/SKILL.md +269 -251
  203. package/skills/marketing/SKILL.md +351 -289
  204. package/skills/mcp-builder/SKILL.md +425 -425
  205. package/skills/neural-memory/SKILL.md +359 -362
  206. package/skills/onboard/SKILL.md +432 -403
  207. package/skills/onboard/references/invariants-template.md +76 -0
  208. package/skills/onboard/scripts/detect-invariants.js +439 -0
  209. package/skills/onboard/scripts/inject-claude-md.js +150 -0
  210. package/skills/onboard/scripts/onboard-invariants.js +194 -0
  211. package/skills/perf/SKILL.md +347 -346
  212. package/skills/plan/SKILL.md +435 -428
  213. package/skills/preflight/SKILL.md +415 -415
  214. package/skills/problem-solver/SKILL.md +380 -284
  215. package/skills/rescue/SKILL.md +474 -474
  216. package/skills/research/SKILL.md +4 -0
  217. package/skills/retro/SKILL.md +3 -1
  218. package/skills/review/SKILL.md +614 -588
  219. package/skills/review-intake/SKILL.md +249 -249
  220. package/skills/safeguard/SKILL.md +200 -200
  221. package/skills/sast/SKILL.md +190 -190
  222. package/skills/scaffold/SKILL.md +328 -287
  223. package/skills/scope-guard/SKILL.md +183 -180
  224. package/skills/scout/SKILL.md +269 -263
  225. package/skills/sentinel/SKILL.md +384 -381
  226. package/skills/sentinel-env/SKILL.md +254 -254
  227. package/skills/sequential-thinking/SKILL.md +234 -234
  228. package/skills/session-bridge/SKILL.md +595 -543
  229. package/skills/session-bridge/scripts/load-invariants.js +397 -0
  230. package/skills/skill-forge/SKILL.md +581 -581
  231. package/skills/skill-router/SKILL.md +3 -0
  232. package/skills/slides/SKILL.md +19 -0
  233. package/skills/surgeon/SKILL.md +215 -215
  234. package/skills/team/SKILL.md +557 -537
  235. package/skills/test/SKILL.md +620 -614
  236. package/skills/trend-scout/SKILL.md +145 -145
  237. package/skills/verification/SKILL.md +334 -326
  238. package/skills/video-creator/SKILL.md +201 -201
  239. package/skills/watchdog/SKILL.md +168 -168
  240. package/skills/worktree/SKILL.md +140 -140
@@ -1,326 +1,334 @@
1
- ---
2
- name: verification
3
- description: "Universal verification runner. Runs lint, type-check, tests, and build. Use after any code change to verify nothing is broken."
4
- metadata:
5
- author: runedev
6
- version: "0.5.0"
7
- layer: L3
8
- model: haiku
9
- group: validation
10
- tools: "Read, Bash, Glob, Grep"
11
- listen: code.changed
12
- emit: verification.complete
13
- ---
14
-
15
- # verification
16
-
17
- Runs all automated checks to verify code health. Stateless — runs checks and reports results.
18
-
19
- ## Instructions
20
-
21
- ### Phase 1: Detect Project Type
22
-
23
- Use `Glob` to find project config files:
24
-
25
- 1. Check for `package.json` → Node.js/TypeScript project
26
- 2. Check for `pyproject.toml` or `setup.py` → Python project
27
- 3. Check for `Cargo.toml` → Rust project
28
- 4. Check for `go.mod` → Go project
29
- 5. Check for `pom.xml` or `build.gradle` → Java project
30
-
31
- Use `Read` on the detected config file to find scripts or tool config (e.g., `package.json` scripts block for custom lint/test commands).
32
-
33
- ```
34
- TodoWrite: [
35
- { content: "Detect project type", status: "in_progress" },
36
- { content: "Run lint check", status: "pending" },
37
- { content: "Run type check", status: "pending" },
38
- { content: "Run test suite", status: "pending" },
39
- { content: "Run build", status: "pending" },
40
- { content: "Generate verification report", status: "pending" }
41
- ]
42
- ```
43
-
44
- ### Phase 2: Run Lint
45
-
46
- Use `Bash` to run the appropriate linter. If `package.json` has a `lint` script, prefer that:
47
-
48
- - **Node.js (npm lint script)**: `npm run lint`
49
- - **Node.js (no script)**: `npx eslint . --max-warnings 0`
50
- - **Python**: `ruff check .` (fallback: `flake8 .`)
51
- - **Rust**: `cargo clippy -- -D warnings`
52
- - **Go**: `golangci-lint run` (fallback: `go vet ./...`)
53
-
54
- If lint fails: record the failure output, mark lint as FAIL, continue to next step. Do NOT stop.
55
-
56
- **Verification gate**: Command exits without crashing (even if it reports lint errors — those are FAIL, not errors).
57
-
58
- ### Phase 3: Run Type Check
59
-
60
- Use `Bash`:
61
-
62
- - **TypeScript**: `npx tsc --noEmit`
63
- - **Python**: `mypy .` (fallback: `pyright .`)
64
- - **Rust**: `cargo check`
65
- - **Go**: `go vet ./...`
66
-
67
- If type check fails: record error count and first 10 error lines, mark as FAIL, continue.
68
-
69
- ### Phase 4: Run Tests
70
-
71
- Use `Bash` to run the test suite. Prefer the project script if available:
72
-
73
- - **Node.js (npm test script)**: `npm test`
74
- - **Vitest**: `npx vitest run`
75
- - **Jest**: `npx jest --passWithNoTests`
76
- - **Python**: `pytest -v` (fallback: `python -m unittest discover`)
77
- - **Rust**: `cargo test`
78
- - **Go**: `go test ./...`
79
-
80
- Record: total tests, passed count, failed count, coverage percentage if output includes it.
81
-
82
- If tests fail: record which tests failed (first 20), mark as FAIL, continue to build.
83
-
84
- ### Phase 5: Run Build
85
-
86
- Use `Bash`:
87
-
88
- - **Node.js**: check `package.json` for `build` script → `npm run build` (fallback: `npx tsc`)
89
- - **Python**: check `pyproject.toml` for `[build-system]` section:
90
- - If build backend found (setuptools, poetry-core, hatchling, flit-core): `python -m build --no-isolation 2>&1 | head -20` to verify packaging
91
- - If `setup.py` exists (legacy): `python setup.py check --strict`
92
- - Then always: `pip install -e . --dry-run` to catch broken entry points, missing `__init__.py`, or import path issues
93
- - If no `pyproject.toml` and no `setup.py` (scripts-only project): SKIP
94
- - **Rust**: `cargo build`
95
- - **Go**: `go build ./...`
96
-
97
- If build fails: record first 20 lines of build output, mark as FAIL.
98
-
99
- ### Phase 6: Generate Report
100
-
101
- Compile all results into the structured report. Update all TodoWrite items to completed.
102
-
103
- ### 3-Level Artifact Verification
104
-
105
- Every file created or modified during implementation must pass ALL 3 levels:
106
-
107
- **Level 1 — EXISTS**: File is on disk, non-empty.
108
- ```
109
- Glob("path/to/expected/file") → found
110
- ```
111
-
112
- **Level 2 — SUBSTANTIVE**: Contains real logic, NOT a stub. Scan for these stub patterns:
113
-
114
- | Pattern | Language | Meaning |
115
- |---|---|---|
116
- | Component returns only `<div>Placeholder</div>` or `<div>TODO</div>` | React/Vue | Stub component |
117
- | Route returns `{ message: "Not implemented" }` or `res.status(501)` | API | Stub endpoint |
118
- | Function body is only `return null` / `return {}` / `return []` / `pass` | Any | Stub function |
119
- | Class with all methods throwing `NotImplementedError` | Python/Java | Stub class |
120
- | `useEffect` with empty body / `async function` with no `await` | React/JS | Hollow implementation |
121
- | File has only type/interface exports but no implementation | TypeScript | Stub types-only file |
122
- | `// TODO` or `# TODO` as the only content in a function | Any | Placeholder |
123
-
124
- If ANY stub pattern detected → mark file as STUB, Level 2 FAIL.
125
-
126
- **Level 3 — WIRED**: Actually imported/called/used by the rest of the system.
127
-
128
- | File Type | Wiring Check |
129
- |---|---|
130
- | Component | `Grep("<ComponentName")` in parent files → ≥1 consumer |
131
- | API route | `Grep("fetch\\|axios\\|api.*endpoint")` for this path → ≥1 caller |
132
- | Hook | `Grep("useHookName(")` → ≥1 consumer |
133
- | Utility function | `Grep("import.*from.*this-file")` → ≥1 importer |
134
- | DB model/schema | `Grep("ModelName\\|table_name")` in query files → ≥1 reference |
135
- | CSS/style module | `Grep("import.*from.*this-style")` → ≥1 importer |
136
-
137
- If file has 0 consumers → mark as UNWIRED, Level 3 FAIL.
138
-
139
- **Exception**: Entry-point files (main.ts, index.ts, App.tsx, routes config) are exempt from Level 3 — they ARE the top-level consumers.
140
-
141
- <HARD-GATE name="3-level-verification">
142
- ALL new files must pass Level 1 + Level 2 + Level 3.
143
- EXISTS but STUB = "Existence Theater" — agent created files but didn't implement them.
144
- EXISTS and SUBSTANTIVE but UNWIRED = dead code — created but never connected.
145
- Report which level failed for each file in the Verification Report.
146
- </HARD-GATE>
147
-
148
- ### Artifact Output Verification
149
-
150
- > Inspired by CLI-Anything (HKUDS/CLI-Anything, 14.5k★): "Never trust exit 0."
151
- > Many tools exit 0 even when they fail silently. Always verify ACTUAL output.
152
-
153
- After each phase command, verify that the expected artifact or indicator is present:
154
-
155
- **Test output** — scan stdout for the pass/fail summary line:
156
- - Vitest/Jest: look for `X passed`, `X failed` — if neither appears, output is incomplete
157
- - Pytest: look for `X passed` or `X failed` — exit 0 with no summary = runner crashed silently
158
- - If only exit code available and no summary line found → mark as INCOMPLETE, not PASS
159
-
160
- **Build output** — after `npm run build` / `cargo build` / `go build`:
161
- - Verify the output file exists: `Glob("dist/**/*.js")` or equivalent
162
- - Verify file size > 0 bytes: a zero-byte output = silent truncation failure
163
- - If output directory is missing → FAIL even if command exited 0
164
-
165
- **Lint output** — parse stdout for counts, not just exit code:
166
- - ESLint: look for `X problems (Y errors, Z warnings)` — `0 problems` = PASS
167
- - Ruff/Flake8: zero output lines = PASS; any file:line output = FAIL
168
- - If linter exits 0 but output contains `error` keyword → log as suspicious, mark WARN
169
-
170
- **Generated files** — check magic bytes for binary outputs:
171
- - PDF: first bytes must be `%PDF` — use `Bash("head -c 4 file.pdf")`
172
- - ZIP/XLSX/DOCX: first bytes must be `PK` (ZIP magic) — use `Bash("head -c 2 file.zip")`
173
- - File size must exceed minimum threshold (PDF > 1KB, ZIP > 100 bytes)
174
-
175
- **Type check** — do not trust exit code alone:
176
- - TypeScript `tsc --noEmit`: look for `Found X errors` or absence of error lines
177
- - `Found 0 errors` = PASS; any other count = FAIL
178
- - Empty output from `tsc` = PASS (no errors emitted) — note explicitly
179
-
180
- <HARD-GATE name="artifact-verification">
181
- Verification MUST check actual command output for success indicators, not just exit codes.
182
- Exit 0 without a confirming output artifact or success string = UNVERIFIED.
183
- Report the specific line that confirmed success (e.g., "3 passed, 0 failed").
184
- </HARD-GATE>
185
-
186
- ## Error Recovery
187
-
188
- - If project type cannot be detected: report "Unknown project type" and skip all checks
189
- - If a command is not found (e.g., `ruff` not installed): note "tool not installed", mark check as SKIP
190
- - If a command hangs for more than 60 seconds: kill it, mark check as TIMEOUT, continue
191
-
192
- ## Calls (outbound)
193
-
194
- None — pure runner using Bash for all checks. Does not invoke other skills.
195
-
196
- ## Called By (inbound)
197
-
198
- - `cook` (L1): Phase 6 VERIFY — final check before commit
199
- - `fix` (L2): validate fix doesn't break existing functionality
200
- - `test` (L2): validate test coverage meets threshold
201
- - `deploy` (L2): post-deploy health checks
202
- - `sentinel` (L2): run security audit tools (npm audit, etc.)
203
- - `safeguard` (L2): verify safety net is solid before refactoring
204
- - `db` (L2): run migration in test environment
205
- - `perf` (L2): run benchmark scripts if configured
206
- - `skill-forge` (L2): verify newly created skill passes lint/type/build checks
207
-
208
- ## Output Format
209
-
210
- ```
211
- VERIFICATION REPORT
212
- ===================
213
- Lint: [PASS/FAIL/SKIP] ([details])
214
- Types: [PASS/FAIL/SKIP] ([X errors])
215
- Tests: [PASS/FAIL/SKIP] ([passed]/[total], [coverage]%)
216
- Build: [PASS/FAIL/SKIP]
217
-
218
- ### 3-Level File Verification
219
- | File | L1 Exists | L2 Substantive | L3 Wired | Verdict |
220
- |------|-----------|----------------|----------|---------|
221
- | src/auth/login.ts | ✓ | ✓ | ✓ (imported by routes.ts) | PASS |
222
- | src/auth/reset.ts | ✓ | STUB (returns null) | — | FAIL L2 |
223
- | src/utils/format.ts | ✓ | ✓ | UNWIRED (0 importers) | FAIL L3 |
224
-
225
- Overall: [PASS/FAIL]
226
-
227
- ### Failures (if any)
228
- - Lint: [error details with file:line]
229
- - Types: [first 5 type errors]
230
- - Tests: [first 5 failing test names]
231
- - Build: [first 5 build errors]
232
- - Stubs: [files that failed Level 2 with stub pattern detected]
233
- - Unwired: [files that failed Level 3 with 0 consumers]
234
- ```
235
-
236
- ## Output Completion Enforcement
237
-
238
- > From taste-skill (Leonxlnx/taste-skill, 3.4k★): Truncated code is worse than no code — it passes reviews but breaks at runtime.
239
-
240
- When verifying code files (Level 2 SUBSTANTIVE check), also scan for **truncation patterns** — signs that the agent generated partial output and stopped:
241
-
242
- | Banned Pattern | Language | What It Means |
243
- |---|---|---|
244
- | `// ...` or `/* ... */` as a statement | JS/TS | Agent truncated remaining code |
245
- | `# ...` as a statement (not comment) | Python | Agent truncated |
246
- | `// rest of code` / `// remaining implementation` | Any | Explicit truncation admission |
247
- | `// TODO: implement` as sole function body | Any | Placeholder, not implementation |
248
- | `{ /* same as above */ }` | JS/TS | Copy-paste truncation |
249
- | `...` (bare ellipsis, not spread operator) | JS/TS/Python | Truncation marker |
250
- | `[PAUSED]` / `[CONTINUED]` in source | Any | Agent session marker leaked into code |
251
-
252
- **Action on detection:**
253
- - Mark file as TRUNCATED (distinct from STUB) in Verification Report
254
- - TRUNCATED files are Level 2 FAIL they CANNOT pass verification
255
- - Report the specific line number and pattern detected
256
- - If agent claims "done" with truncated files REJECTED by Evidence-Before-Claims gate
257
-
258
- **Continuation protocol** if the agent hit output limits mid-file:
259
- - Agent MUST log: `[PAUSED — X of Y functions complete]` in its response (NOT in the code file)
260
- - Agent MUST resume and complete the file in the next turn
261
- - Verification re-runs after completion to clear the TRUNCATED flag
262
-
263
- ## Evidence-Before-Claims Gate
264
-
265
- <HARD-GATE>
266
- An agent MUST NOT claim "done", "fixed", "passing", or "verified" without showing the actual command output that proves it.
267
- "I ran the tests and they pass" WITHOUT stdout/stderr = UNVERIFIED CLAIM = REJECTED.
268
- The verification report IS the evidence. No report = no verification happened.
269
- </HARD-GATE>
270
-
271
- ### Claim Validation Protocol
272
-
273
- When any skill calls verification and then reports results upstream:
274
-
275
- 1. **Output capture is mandatory** every Bash command's stdout/stderr must appear in the report
276
- 2. **Pass requires proof** PASS means "tool ran AND output shows zero errors" (not "tool ran without crashing")
277
- 3. **Silence is not success** — if a command produces no output, note it explicitly ("0 errors, 0 warnings")
278
- 4. **Partial runs are labeled** — if only 2 of 4 checks ran, Overall = INCOMPLETE (not PASS)
279
-
280
- ### Red Flags — Agent is Lying
281
-
282
- | Claim | Without | Verdict |
283
- |---|---|---|
284
- | "All tests pass" | Test runner stdout showing pass count | REJECTED re-run and show output |
285
- | "No lint errors" | Linter stdout | REJECTED re-run and show output |
286
- | "Build succeeds" | Build command stdout | REJECTED re-run and show output |
287
- | "I verified it" | Verification Report | REJECTED — run verification skill properly |
288
- | "Fixed and working" | Before/after test output | REJECTED show the diff in results |
289
-
290
- ## Constraints
291
-
292
- 1. MUST run ALL four checks: lint, type-check, tests, buildnot just tests
293
- 2. MUST show actual command outputnever claim "all passed" without evidence
294
- 3. MUST report specific failures with file:line references
295
- 4. MUST NOT skip checks because "changes are small"
296
- 5. MUST include stdout/stderr capture in every check result empty output noted explicitly
297
- 6. MUST mark Overall as INCOMPLETE if any check was skipped without valid reason (tool not installed = valid, "changes are small" = invalid)
298
-
299
- ## Sharp Edges
300
-
301
- Known failure modes for this skill. Check these before declaring done.
302
-
303
- | Failure Mode | Severity | Mitigation |
304
- |---|---|---|
305
- | Claiming "all passed" without showing actual command output | CRITICAL | Evidence-Before-Claims HARD-GATE blocks this stdout/stderr is mandatory |
306
- | Agent says "verified" without producing Verification Report | CRITICAL | No report = no verification. Re-run the skill properly. |
307
- | Skipping build because "changes are small" | HIGH | Constraint 4: all four checks mandatory — size of changes doesn't matter |
308
- | Marking check as PASS when the tool isn't installed | MEDIUM | Mark as SKIP (not PASS) — PASS means the tool ran and reported clean |
309
- | Stopping after first failure instead of running remaining checks | MEDIUM | Run all checks; aggregate all failures so developer can fix everything at once |
310
- | Reporting PASS when output has warnings but zero errors | LOW | PASS is correct but note warning count — caller decides if warnings matter |
311
- | Trusting exit code 0 without output verification | CRITICAL | Artifact Verification HARD-GATE: always confirm success indicator in stdout (pass count, "0 errors", output file exists) |
312
- | Existence Theater — file exists but is a stub | HIGH | 3-Level check: Level 2 scans for stub patterns (`<div>Placeholder</div>`, `return null`, `NotImplementedError`) |
313
- | Dead code file created but never imported/used | MEDIUM | 3-Level check: Level 3 greps for consumers. 0 importers = UNWIRED |
314
- | Truncated code agent hit output limit mid-file | HIGH | Output Completion Enforcement: scan for `// ...`, `// rest of code`, bare ellipsis patterns. TRUNCATED = Level 2 FAIL |
315
-
316
- ## Done When
317
-
318
- - Project type detected from config files
319
- - lint, type-check, tests, and build all executed (or SKIP with reason if tool missing)
320
- - Each check shows actual command output
321
- - Failures include specific file:line references (not just counts)
322
- - Verification Report emitted with Overall PASS/FAIL verdict
323
-
324
- ## Cost Profile
325
-
326
- ~$0.01-0.03 per run. Haiku + Bash commands. Fast and cheap.
1
+ ---
2
+ name: verification
3
+ description: "Universal verification runner. Runs lint, type-check, tests, and build. Use after any code change to verify nothing is broken."
4
+ metadata:
5
+ author: runedev
6
+ version: "0.5.0"
7
+ layer: L3
8
+ model: haiku
9
+ group: validation
10
+ tools: "Read, Bash, Glob, Grep"
11
+ listen: code.changed
12
+ emit: verification.complete
13
+ ---
14
+
15
+ # verification
16
+
17
+ Runs all automated checks to verify code health. Stateless — runs checks and reports results.
18
+
19
+ ## Instructions
20
+
21
+ ### Phase 1: Detect Project Type
22
+
23
+ Use `Glob` to find project config files:
24
+
25
+ 1. Check for `package.json` → Node.js/TypeScript project
26
+ 2. Check for `pyproject.toml` or `setup.py` → Python project
27
+ 3. Check for `Cargo.toml` → Rust project
28
+ 4. Check for `go.mod` → Go project
29
+ 5. Check for `pom.xml` or `build.gradle` → Java project
30
+
31
+ Use `Read` on the detected config file to find scripts or tool config (e.g., `package.json` scripts block for custom lint/test commands).
32
+
33
+ ```
34
+ TodoWrite: [
35
+ { content: "Detect project type", status: "in_progress" },
36
+ { content: "Run lint check", status: "pending" },
37
+ { content: "Run type check", status: "pending" },
38
+ { content: "Run test suite", status: "pending" },
39
+ { content: "Run build", status: "pending" },
40
+ { content: "Generate verification report", status: "pending" }
41
+ ]
42
+ ```
43
+
44
+ ### Phase 2: Run Lint
45
+
46
+ Use `Bash` to run the appropriate linter. If `package.json` has a `lint` script, prefer that:
47
+
48
+ - **Node.js (npm lint script)**: `npm run lint`
49
+ - **Node.js (no script)**: `npx eslint . --max-warnings 0`
50
+ - **Python**: `ruff check .` (fallback: `flake8 .`)
51
+ - **Rust**: `cargo clippy -- -D warnings`
52
+ - **Go**: `golangci-lint run` (fallback: `go vet ./...`)
53
+
54
+ If lint fails: record the failure output, mark lint as FAIL, continue to next step. Do NOT stop.
55
+
56
+ **Verification gate**: Command exits without crashing (even if it reports lint errors — those are FAIL, not errors).
57
+
58
+ ### Phase 3: Run Type Check
59
+
60
+ Use `Bash`:
61
+
62
+ - **TypeScript**: `npx tsc --noEmit`
63
+ - **Python**: `mypy .` (fallback: `pyright .`)
64
+ - **Rust**: `cargo check`
65
+ - **Go**: `go vet ./...`
66
+
67
+ If type check fails: record error count and first 10 error lines, mark as FAIL, continue.
68
+
69
+ ### Phase 4: Run Tests
70
+
71
+ Use `Bash` to run the test suite. Prefer the project script if available:
72
+
73
+ - **Node.js (npm test script)**: `npm test`
74
+ - **Vitest**: `npx vitest run`
75
+ - **Jest**: `npx jest --passWithNoTests`
76
+ - **Python**: `pytest -v` (fallback: `python -m unittest discover`)
77
+ - **Rust**: `cargo test`
78
+ - **Go**: `go test ./...`
79
+
80
+ Record: total tests, passed count, failed count, coverage percentage if output includes it.
81
+
82
+ If tests fail: record which tests failed (first 20), mark as FAIL, continue to build.
83
+
84
+ ### Phase 5: Run Build
85
+
86
+ Use `Bash`:
87
+
88
+ - **Node.js**: check `package.json` for `build` script → `npm run build` (fallback: `npx tsc`)
89
+ - **Python**: check `pyproject.toml` for `[build-system]` section:
90
+ - If build backend found (setuptools, poetry-core, hatchling, flit-core): `python -m build --no-isolation 2>&1 | head -20` to verify packaging
91
+ - If `setup.py` exists (legacy): `python setup.py check --strict`
92
+ - Then always: `pip install -e . --dry-run` to catch broken entry points, missing `__init__.py`, or import path issues
93
+ - If no `pyproject.toml` and no `setup.py` (scripts-only project): SKIP
94
+ - **Rust**: `cargo build`
95
+ - **Go**: `go build ./...`
96
+
97
+ If build fails: record first 20 lines of build output, mark as FAIL.
98
+
99
+ ### Phase 6: Generate Report
100
+
101
+ Compile all results into the structured report. Update all TodoWrite items to completed.
102
+
103
+ ### 3-Level Artifact Verification
104
+
105
+ Every file created or modified during implementation must pass ALL 3 levels:
106
+
107
+ **Level 1 — EXISTS**: File is on disk, non-empty.
108
+ ```
109
+ Glob("path/to/expected/file") → found
110
+ ```
111
+
112
+ **Level 2 — SUBSTANTIVE**: Contains real logic, NOT a stub. Scan for these stub patterns:
113
+
114
+ | Pattern | Language | Meaning |
115
+ |---|---|---|
116
+ | Component returns only `<div>Placeholder</div>` or `<div>TODO</div>` | React/Vue | Stub component |
117
+ | Route returns `{ message: "Not implemented" }` or `res.status(501)` | API | Stub endpoint |
118
+ | Function body is only `return null` / `return {}` / `return []` / `pass` | Any | Stub function |
119
+ | Class with all methods throwing `NotImplementedError` | Python/Java | Stub class |
120
+ | `useEffect` with empty body / `async function` with no `await` | React/JS | Hollow implementation |
121
+ | File has only type/interface exports but no implementation | TypeScript | Stub types-only file |
122
+ | `// TODO` or `# TODO` as the only content in a function | Any | Placeholder |
123
+
124
+ If ANY stub pattern detected → mark file as STUB, Level 2 FAIL.
125
+
126
+ **Level 3 — WIRED**: Actually imported/called/used by the rest of the system.
127
+
128
+ | File Type | Wiring Check |
129
+ |---|---|
130
+ | Component | `Grep("<ComponentName")` in parent files → ≥1 consumer |
131
+ | API route | `Grep("fetch\\|axios\\|api.*endpoint")` for this path → ≥1 caller |
132
+ | Hook | `Grep("useHookName(")` → ≥1 consumer |
133
+ | Utility function | `Grep("import.*from.*this-file")` → ≥1 importer |
134
+ | DB model/schema | `Grep("ModelName\\|table_name")` in query files → ≥1 reference |
135
+ | CSS/style module | `Grep("import.*from.*this-style")` → ≥1 importer |
136
+
137
+ If file has 0 consumers → mark as UNWIRED, Level 3 FAIL.
138
+
139
+ **Exception**: Entry-point files (main.ts, index.ts, App.tsx, routes config) are exempt from Level 3 — they ARE the top-level consumers.
140
+
141
+ <HARD-GATE name="3-level-verification">
142
+ ALL new files must pass Level 1 + Level 2 + Level 3.
143
+ EXISTS but STUB = "Existence Theater" — agent created files but didn't implement them.
144
+ EXISTS and SUBSTANTIVE but UNWIRED = dead code — created but never connected.
145
+ Report which level failed for each file in the Verification Report.
146
+ </HARD-GATE>
147
+
148
+ ### Artifact Output Verification
149
+
150
+ > Inspired by CLI-Anything (HKUDS/CLI-Anything, 14.5k★): "Never trust exit 0."
151
+ > Many tools exit 0 even when they fail silently. Always verify ACTUAL output.
152
+
153
+ After each phase command, verify that the expected artifact or indicator is present:
154
+
155
+ **Test output** — scan stdout for the pass/fail summary line:
156
+ - Vitest/Jest: look for `X passed`, `X failed` — if neither appears, output is incomplete
157
+ - Pytest: look for `X passed` or `X failed` — exit 0 with no summary = runner crashed silently
158
+ - If only exit code available and no summary line found → mark as INCOMPLETE, not PASS
159
+
160
+ **Build output** — after `npm run build` / `cargo build` / `go build`:
161
+ - Verify the output file exists: `Glob("dist/**/*.js")` or equivalent
162
+ - Verify file size > 0 bytes: a zero-byte output = silent truncation failure
163
+ - If output directory is missing → FAIL even if command exited 0
164
+
165
+ **Lint output** — parse stdout for counts, not just exit code:
166
+ - ESLint: look for `X problems (Y errors, Z warnings)` — `0 problems` = PASS
167
+ - Ruff/Flake8: zero output lines = PASS; any file:line output = FAIL
168
+ - If linter exits 0 but output contains `error` keyword → log as suspicious, mark WARN
169
+
170
+ **Generated files** — check magic bytes for binary outputs:
171
+ - PDF: first bytes must be `%PDF` — use `Bash("head -c 4 file.pdf")`
172
+ - ZIP/XLSX/DOCX: first bytes must be `PK` (ZIP magic) — use `Bash("head -c 2 file.zip")`
173
+ - File size must exceed minimum threshold (PDF > 1KB, ZIP > 100 bytes)
174
+
175
+ **Type check** — do not trust exit code alone:
176
+ - TypeScript `tsc --noEmit`: look for `Found X errors` or absence of error lines
177
+ - `Found 0 errors` = PASS; any other count = FAIL
178
+ - Empty output from `tsc` = PASS (no errors emitted) — note explicitly
179
+
180
+ <HARD-GATE name="artifact-verification">
181
+ Verification MUST check actual command output for success indicators, not just exit codes.
182
+ Exit 0 without a confirming output artifact or success string = UNVERIFIED.
183
+ Report the specific line that confirmed success (e.g., "3 passed, 0 failed").
184
+ </HARD-GATE>
185
+
186
+ ## Error Recovery
187
+
188
+ - If project type cannot be detected: report "Unknown project type" and skip all checks
189
+ - If a command is not found (e.g., `ruff` not installed): note "tool not installed", mark check as SKIP
190
+ - If a command hangs for more than 60 seconds: kill it, mark check as TIMEOUT, continue
191
+
192
+ ## Calls (outbound)
193
+
194
+ None — pure runner using Bash for all checks. Does not invoke other skills.
195
+
196
+ ## Called By (inbound)
197
+
198
+ - `cook` (L1): Phase 6 VERIFY — final check before commit
199
+ - `fix` (L2): validate fix doesn't break existing functionality
200
+ - `test` (L2): validate test coverage meets threshold
201
+ - `deploy` (L2): post-deploy health checks
202
+ - `sentinel` (L2): run security audit tools (npm audit, etc.)
203
+ - `safeguard` (L2): verify safety net is solid before refactoring
204
+ - `db` (L2): run migration in test environment
205
+ - `perf` (L2): run benchmark scripts if configured
206
+ - `skill-forge` (L2): verify newly created skill passes lint/type/build checks
207
+ - `team` (L1): verify each parallel workstream before merge
208
+ - `scaffold` (L1): verify scaffolded project builds and passes initial tests
209
+ - `launch` (L1): pre-deploy verification gate
210
+ - `mcp-builder` (L2): verify generated MCP server compiles and starts
211
+ - `preflight` (L2): run verification as part of pre-commit quality gate
212
+ - `logic-guardian` (L2): verify logic invariants hold after changes
213
+ - `dependency-doctor` (L3): verify builds pass after dependency updates
214
+ - `sast` (L3): run verification alongside static analysis
215
+
216
+ ## Output Format
217
+
218
+ ```
219
+ VERIFICATION REPORT
220
+ ===================
221
+ Lint: [PASS/FAIL/SKIP] ([details])
222
+ Types: [PASS/FAIL/SKIP] ([X errors])
223
+ Tests: [PASS/FAIL/SKIP] ([passed]/[total], [coverage]%)
224
+ Build: [PASS/FAIL/SKIP]
225
+
226
+ ### 3-Level File Verification
227
+ | File | L1 Exists | L2 Substantive | L3 Wired | Verdict |
228
+ |------|-----------|----------------|----------|---------|
229
+ | src/auth/login.ts | | ✓ | ✓ (imported by routes.ts) | PASS |
230
+ | src/auth/reset.ts | | STUB (returns null) | — | FAIL L2 |
231
+ | src/utils/format.ts | | ✓ | UNWIRED (0 importers) | FAIL L3 |
232
+
233
+ Overall: [PASS/FAIL]
234
+
235
+ ### Failures (if any)
236
+ - Lint: [error details with file:line]
237
+ - Types: [first 5 type errors]
238
+ - Tests: [first 5 failing test names]
239
+ - Build: [first 5 build errors]
240
+ - Stubs: [files that failed Level 2 with stub pattern detected]
241
+ - Unwired: [files that failed Level 3 with 0 consumers]
242
+ ```
243
+
244
+ ## Output Completion Enforcement
245
+
246
+ > From taste-skill (Leonxlnx/taste-skill, 3.4k★): Truncated code is worse than no code it passes reviews but breaks at runtime.
247
+
248
+ When verifying code files (Level 2 SUBSTANTIVE check), also scan for **truncation patterns** — signs that the agent generated partial output and stopped:
249
+
250
+ | Banned Pattern | Language | What It Means |
251
+ |---|---|---|
252
+ | `// ...` or `/* ... */` as a statement | JS/TS | Agent truncated remaining code |
253
+ | `# ...` as a statement (not comment) | Python | Agent truncated |
254
+ | `// rest of code` / `// remaining implementation` | Any | Explicit truncation admission |
255
+ | `// TODO: implement` as sole function body | Any | Placeholder, not implementation |
256
+ | `{ /* same as above */ }` | JS/TS | Copy-paste truncation |
257
+ | `...` (bare ellipsis, not spread operator) | JS/TS/Python | Truncation marker |
258
+ | `[PAUSED]` / `[CONTINUED]` in source | Any | Agent session marker leaked into code |
259
+
260
+ **Action on detection:**
261
+ - Mark file as TRUNCATED (distinct from STUB) in Verification Report
262
+ - TRUNCATED files are Level 2 FAIL — they CANNOT pass verification
263
+ - Report the specific line number and pattern detected
264
+ - If agent claims "done" with truncated files → REJECTED by Evidence-Before-Claims gate
265
+
266
+ **Continuation protocol** if the agent hit output limits mid-file:
267
+ - Agent MUST log: `[PAUSED X of Y functions complete]` in its response (NOT in the code file)
268
+ - Agent MUST resume and complete the file in the next turn
269
+ - Verification re-runs after completion to clear the TRUNCATED flag
270
+
271
+ ## Evidence-Before-Claims Gate
272
+
273
+ <HARD-GATE>
274
+ An agent MUST NOT claim "done", "fixed", "passing", or "verified" without showing the actual command output that proves it.
275
+ "I ran the tests and they pass" WITHOUT stdout/stderr = UNVERIFIED CLAIM = REJECTED.
276
+ The verification report IS the evidence. No report = no verification happened.
277
+ </HARD-GATE>
278
+
279
+ ### Claim Validation Protocol
280
+
281
+ When any skill calls verification and then reports results upstream:
282
+
283
+ 1. **Output capture is mandatory** — every Bash command's stdout/stderr must appear in the report
284
+ 2. **Pass requires proof** PASS means "tool ran AND output shows zero errors" (not "tool ran without crashing")
285
+ 3. **Silence is not success** if a command produces no output, note it explicitly ("0 errors, 0 warnings")
286
+ 4. **Partial runs are labeled** if only 2 of 4 checks ran, Overall = INCOMPLETE (not PASS)
287
+
288
+ ### Red FlagsAgent is Lying
289
+
290
+ | Claim | Without | Verdict |
291
+ |---|---|---|
292
+ | "All tests pass" | Test runner stdout showing pass count | REJECTED re-run and show output |
293
+ | "No lint errors" | Linter stdout | REJECTED re-run and show output |
294
+ | "Build succeeds" | Build command stdout | REJECTED — re-run and show output |
295
+ | "I verified it" | Verification Report | REJECTED — run verification skill properly |
296
+ | "Fixed and working" | Before/after test output | REJECTEDshow the diff in results |
297
+
298
+ ## Constraints
299
+
300
+ 1. MUST run ALL four checks: lint, type-check, tests, build — not just tests
301
+ 2. MUST show actual command output never claim "all passed" without evidence
302
+ 3. MUST report specific failures with file:line references
303
+ 4. MUST NOT skip checks because "changes are small"
304
+ 5. MUST include stdout/stderr capture in every check result — empty output noted explicitly
305
+ 6. MUST mark Overall as INCOMPLETE if any check was skipped without valid reason (tool not installed = valid, "changes are small" = invalid)
306
+
307
+ ## Sharp Edges
308
+
309
+ Known failure modes for this skill. Check these before declaring done.
310
+
311
+ | Failure Mode | Severity | Mitigation |
312
+ |---|---|---|
313
+ | Claiming "all passed" without showing actual command output | CRITICAL | Evidence-Before-Claims HARD-GATE blocks this stdout/stderr is mandatory |
314
+ | Agent says "verified" without producing Verification Report | CRITICAL | No report = no verification. Re-run the skill properly. |
315
+ | Skipping build because "changes are small" | HIGH | Constraint 4: all four checks mandatory — size of changes doesn't matter |
316
+ | Marking check as PASS when the tool isn't installed | MEDIUM | Mark as SKIP (not PASS) — PASS means the tool ran and reported clean |
317
+ | Stopping after first failure instead of running remaining checks | MEDIUM | Run all checks; aggregate all failures so developer can fix everything at once |
318
+ | Reporting PASS when output has warnings but zero errors | LOW | PASS is correct but note warning count — caller decides if warnings matter |
319
+ | Trusting exit code 0 without output verification | CRITICAL | Artifact Verification HARD-GATE: always confirm success indicator in stdout (pass count, "0 errors", output file exists) |
320
+ | Existence Theater — file exists but is a stub | HIGH | 3-Level check: Level 2 scans for stub patterns (`<div>Placeholder</div>`, `return null`, `NotImplementedError`) |
321
+ | Dead code file created but never imported/used | MEDIUM | 3-Level check: Level 3 greps for consumers. 0 importers = UNWIRED |
322
+ | Truncated code — agent hit output limit mid-file | HIGH | Output Completion Enforcement: scan for `// ...`, `// rest of code`, bare ellipsis patterns. TRUNCATED = Level 2 FAIL |
323
+
324
+ ## Done When
325
+
326
+ - Project type detected from config files
327
+ - lint, type-check, tests, and build all executed (or SKIP with reason if tool missing)
328
+ - Each check shows actual command output
329
+ - Failures include specific file:line references (not just counts)
330
+ - Verification Report emitted with Overall PASS/FAIL verdict
331
+
332
+ ## Cost Profile
333
+
334
+ ~$0.01-0.03 per run. Haiku + Bash commands. Fast and cheap.