@rune-kit/rune 2.8.0 → 2.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (287) hide show
  1. package/LICENSE +21 -21
  2. package/README.md +68 -34
  3. package/agents/adversary.md +27 -0
  4. package/agents/architect.md +19 -29
  5. package/agents/asset-creator.md +18 -4
  6. package/agents/audit.md +25 -4
  7. package/agents/autopsy.md +19 -4
  8. package/agents/ba.md +35 -0
  9. package/agents/brainstorm.md +31 -4
  10. package/agents/browser-pilot.md +21 -4
  11. package/agents/coder.md +21 -29
  12. package/agents/completion-gate.md +20 -4
  13. package/agents/constraint-check.md +18 -4
  14. package/agents/context-engine.md +22 -4
  15. package/agents/context-pack.md +32 -0
  16. package/agents/cook.md +41 -4
  17. package/agents/db.md +19 -4
  18. package/agents/debug.md +33 -4
  19. package/agents/dependency-doctor.md +20 -4
  20. package/agents/deploy.md +27 -4
  21. package/agents/design.md +22 -4
  22. package/agents/doc-processor.md +27 -0
  23. package/agents/docs-seeker.md +19 -4
  24. package/agents/docs.md +31 -0
  25. package/agents/fix.md +37 -4
  26. package/agents/git.md +29 -0
  27. package/agents/hallucination-guard.md +20 -4
  28. package/agents/incident.md +21 -4
  29. package/agents/integrity-check.md +18 -4
  30. package/agents/journal.md +19 -4
  31. package/agents/launch.md +32 -4
  32. package/agents/logic-guardian.md +26 -11
  33. package/agents/marketing.md +23 -4
  34. package/agents/mcp-builder.md +26 -0
  35. package/agents/neural-memory.md +30 -0
  36. package/agents/onboard.md +22 -4
  37. package/agents/perf.md +21 -4
  38. package/agents/plan.md +29 -4
  39. package/agents/preflight.md +22 -4
  40. package/agents/problem-solver.md +20 -4
  41. package/agents/rescue.md +23 -4
  42. package/agents/research.md +19 -4
  43. package/agents/researcher.md +19 -29
  44. package/agents/retro.md +32 -0
  45. package/agents/review-intake.md +20 -4
  46. package/agents/review.md +32 -4
  47. package/agents/reviewer.md +20 -28
  48. package/agents/safeguard.md +19 -4
  49. package/agents/sast.md +18 -4
  50. package/agents/scaffold.md +41 -0
  51. package/agents/scanner.md +19 -28
  52. package/agents/scope-guard.md +18 -4
  53. package/agents/scout.md +23 -4
  54. package/agents/sentinel-env.md +26 -0
  55. package/agents/sentinel.md +33 -4
  56. package/agents/sequential-thinking.md +20 -4
  57. package/agents/session-bridge.md +24 -4
  58. package/agents/skill-forge.md +22 -4
  59. package/agents/skill-router.md +26 -4
  60. package/agents/slides.md +24 -0
  61. package/agents/surgeon.md +19 -4
  62. package/agents/team.md +30 -4
  63. package/agents/test.md +36 -4
  64. package/agents/trend-scout.md +17 -4
  65. package/agents/verification.md +20 -4
  66. package/agents/video-creator.md +20 -4
  67. package/agents/watchdog.md +19 -4
  68. package/agents/worktree.md +17 -4
  69. package/commands/rune.md +168 -168
  70. package/compiler/__tests__/analytics.test.js +370 -0
  71. package/compiler/adapters/openclaw.js +2 -2
  72. package/compiler/analytics.js +385 -0
  73. package/compiler/bin/rune.js +68 -2
  74. package/compiler/dashboard.js +883 -0
  75. package/compiler/transforms/branding.js +1 -1
  76. package/contexts/dev.md +34 -34
  77. package/contexts/research.md +43 -43
  78. package/contexts/review.md +55 -55
  79. package/extensions/ai-ml/PACK.md +88 -88
  80. package/extensions/ai-ml/skills/ai-agents.md +172 -172
  81. package/extensions/ai-ml/skills/code-sandbox.md +187 -187
  82. package/extensions/ai-ml/skills/deep-research.md +146 -146
  83. package/extensions/ai-ml/skills/embedding-search.md +66 -66
  84. package/extensions/ai-ml/skills/fine-tuning-guide.md +74 -74
  85. package/extensions/ai-ml/skills/llm-architect.md +125 -125
  86. package/extensions/ai-ml/skills/llm-integration.md +64 -64
  87. package/extensions/ai-ml/skills/prompt-patterns.md +72 -72
  88. package/extensions/ai-ml/skills/rag-patterns.md +66 -66
  89. package/extensions/ai-ml/skills/web-extraction.md +114 -114
  90. package/extensions/analytics/PACK.md +92 -92
  91. package/extensions/analytics/skills/ab-testing.md +72 -72
  92. package/extensions/analytics/skills/dashboard-patterns.md +83 -83
  93. package/extensions/analytics/skills/data-validation.md +68 -68
  94. package/extensions/analytics/skills/funnel-analysis.md +81 -81
  95. package/extensions/analytics/skills/sql-patterns.md +57 -57
  96. package/extensions/analytics/skills/statistical-analysis.md +79 -79
  97. package/extensions/analytics/skills/tracking-setup.md +71 -71
  98. package/extensions/backend/PACK.md +104 -104
  99. package/extensions/backend/skills/api-patterns.md +84 -84
  100. package/extensions/backend/skills/async-pipeline.md +193 -193
  101. package/extensions/backend/skills/auth-patterns.md +97 -97
  102. package/extensions/backend/skills/background-jobs.md +133 -133
  103. package/extensions/backend/skills/caching-patterns.md +108 -108
  104. package/extensions/backend/skills/cli-generation.md +133 -133
  105. package/extensions/backend/skills/database-patterns.md +87 -87
  106. package/extensions/backend/skills/middleware-patterns.md +104 -104
  107. package/extensions/chrome-ext/PACK.md +93 -93
  108. package/extensions/chrome-ext/skills/cws-preflight.md +143 -143
  109. package/extensions/chrome-ext/skills/cws-publish.md +104 -104
  110. package/extensions/chrome-ext/skills/ext-ai-integration.md +251 -251
  111. package/extensions/chrome-ext/skills/ext-messaging.md +139 -139
  112. package/extensions/chrome-ext/skills/ext-storage.md +133 -133
  113. package/extensions/chrome-ext/skills/mv3-scaffold.md +164 -164
  114. package/extensions/content/PACK.md +96 -96
  115. package/extensions/content/skills/blog-patterns.md +88 -88
  116. package/extensions/content/skills/cms-integration.md +131 -131
  117. package/extensions/content/skills/content-scoring.md +107 -107
  118. package/extensions/content/skills/i18n.md +83 -83
  119. package/extensions/content/skills/mdx-authoring.md +137 -137
  120. package/extensions/content/skills/reference.md +1014 -1014
  121. package/extensions/content/skills/seo-patterns.md +67 -67
  122. package/extensions/content/skills/video-repurpose.md +153 -153
  123. package/extensions/devops/PACK.md +101 -101
  124. package/extensions/devops/skills/chaos-testing.md +67 -67
  125. package/extensions/devops/skills/ci-cd.md +75 -75
  126. package/extensions/devops/skills/docker.md +58 -58
  127. package/extensions/devops/skills/edge-serverless.md +163 -163
  128. package/extensions/devops/skills/infra-as-code.md +158 -158
  129. package/extensions/devops/skills/kubernetes.md +110 -110
  130. package/extensions/devops/skills/monitoring.md +57 -57
  131. package/extensions/devops/skills/server-setup.md +64 -64
  132. package/extensions/devops/skills/ssl-domain.md +42 -42
  133. package/extensions/ecommerce/PACK.md +116 -116
  134. package/extensions/ecommerce/skills/cart-system.md +79 -79
  135. package/extensions/ecommerce/skills/inventory-mgmt.md +102 -102
  136. package/extensions/ecommerce/skills/order-management.md +126 -126
  137. package/extensions/ecommerce/skills/payment-integration.md +472 -472
  138. package/extensions/ecommerce/skills/shopify-dev.md +69 -69
  139. package/extensions/ecommerce/skills/subscription-billing.md +93 -93
  140. package/extensions/ecommerce/skills/tax-compliance.md +117 -117
  141. package/extensions/gamedev/PACK.md +142 -142
  142. package/extensions/gamedev/skills/asset-pipeline.md +74 -74
  143. package/extensions/gamedev/skills/audio-system.md +129 -129
  144. package/extensions/gamedev/skills/camera-system.md +87 -87
  145. package/extensions/gamedev/skills/ecs.md +98 -98
  146. package/extensions/gamedev/skills/game-loops.md +72 -72
  147. package/extensions/gamedev/skills/input-system.md +199 -199
  148. package/extensions/gamedev/skills/multiplayer.md +180 -180
  149. package/extensions/gamedev/skills/particles.md +105 -105
  150. package/extensions/gamedev/skills/physics-engine.md +89 -89
  151. package/extensions/gamedev/skills/scene-management.md +146 -146
  152. package/extensions/gamedev/skills/threejs-patterns.md +90 -90
  153. package/extensions/gamedev/skills/webgl.md +71 -71
  154. package/extensions/mobile/PACK.md +106 -106
  155. package/extensions/mobile/skills/app-store-connect.md +152 -152
  156. package/extensions/mobile/skills/app-store-prep.md +66 -66
  157. package/extensions/mobile/skills/deep-linking.md +109 -109
  158. package/extensions/mobile/skills/flutter.md +60 -60
  159. package/extensions/mobile/skills/ios-build-pipeline.md +142 -142
  160. package/extensions/mobile/skills/native-bridge.md +66 -66
  161. package/extensions/mobile/skills/ota-updates.md +97 -97
  162. package/extensions/mobile/skills/push-notifications.md +111 -111
  163. package/extensions/mobile/skills/react-native.md +82 -82
  164. package/extensions/saas/PACK.md +116 -116
  165. package/extensions/saas/skills/billing-integration.md +200 -200
  166. package/extensions/saas/skills/feature-flags.md +130 -130
  167. package/extensions/saas/skills/multi-tenant.md +103 -103
  168. package/extensions/saas/skills/onboarding-flow.md +139 -139
  169. package/extensions/saas/skills/subscription-flow.md +95 -95
  170. package/extensions/saas/skills/team-management.md +144 -144
  171. package/extensions/security/PACK.md +99 -99
  172. package/extensions/security/skills/api-security.md +140 -140
  173. package/extensions/security/skills/compliance.md +68 -68
  174. package/extensions/security/skills/owasp-audit.md +64 -64
  175. package/extensions/security/skills/pentest-patterns.md +77 -77
  176. package/extensions/security/skills/secret-mgmt.md +65 -65
  177. package/extensions/security/skills/supply-chain.md +65 -65
  178. package/extensions/trading/PACK.md +80 -80
  179. package/extensions/trading/skills/chart-components.md +55 -55
  180. package/extensions/trading/skills/experiment-loop.md +125 -125
  181. package/extensions/trading/skills/fintech-patterns.md +47 -47
  182. package/extensions/trading/skills/indicator-library.md +58 -58
  183. package/extensions/trading/skills/quant-analysis.md +111 -111
  184. package/extensions/trading/skills/realtime-data.md +58 -58
  185. package/extensions/trading/skills/trade-logic.md +104 -104
  186. package/extensions/ui/PACK.md +130 -130
  187. package/extensions/ui/skills/a11y-audit.md +91 -91
  188. package/extensions/ui/skills/animation-patterns.md +127 -106
  189. package/extensions/ui/skills/component-patterns.md +100 -75
  190. package/extensions/ui/skills/design-decision.md +108 -108
  191. package/extensions/ui/skills/design-system.md +68 -68
  192. package/extensions/ui/skills/landing-patterns.md +155 -155
  193. package/extensions/ui/skills/palette-picker.md +173 -173
  194. package/extensions/ui/skills/react-health.md +90 -90
  195. package/extensions/ui/skills/type-system.md +125 -125
  196. package/extensions/ui/skills/web-vitals.md +153 -153
  197. package/extensions/zalo/PACK.md +145 -145
  198. package/extensions/zalo/skills/zalo-oa-mcp.md +317 -317
  199. package/extensions/zalo/skills/zalo-oa-messaging.md +429 -429
  200. package/extensions/zalo/skills/zalo-oa-setup.md +236 -236
  201. package/extensions/zalo/skills/zalo-oa-webhook.md +189 -189
  202. package/extensions/zalo/skills/zalo-personal-messaging.md +194 -194
  203. package/extensions/zalo/skills/zalo-personal-setup.md +153 -153
  204. package/extensions/zalo/skills/zalo-rate-guard.md +219 -219
  205. package/hooks/auto-format/index.cjs +48 -48
  206. package/hooks/context-watch/index.cjs +95 -68
  207. package/hooks/hooks.json +111 -111
  208. package/hooks/metrics-collector/index.cjs +86 -42
  209. package/hooks/post-session-reflect/index.cjs +189 -153
  210. package/hooks/pre-compact/index.cjs +95 -95
  211. package/hooks/run-hook.cmd +1 -1
  212. package/hooks/secrets-scan/index.cjs +100 -100
  213. package/hooks/session-start/index.cjs +71 -65
  214. package/hooks/typecheck/index.cjs +65 -65
  215. package/package.json +63 -63
  216. package/references/ui-pro-max-data/LICENSE-UI-PRO-MAX +21 -21
  217. package/references/ui-pro-max-data/charts.csv +26 -26
  218. package/references/ui-pro-max-data/colors.csv +161 -161
  219. package/references/ui-pro-max-data/styles.csv +68 -68
  220. package/references/ui-pro-max-data/typography.csv +74 -74
  221. package/references/ui-pro-max-data/ui-reasoning.csv +162 -162
  222. package/references/ui-pro-max-data/ux-guidelines.csv +99 -99
  223. package/skills/adversary/SKILL.md +283 -283
  224. package/skills/asset-creator/SKILL.md +157 -157
  225. package/skills/audit/SKILL.md +148 -2
  226. package/skills/autopsy/SKILL.md +335 -259
  227. package/skills/autopsy/references/repo-analysis-patterns.md +113 -0
  228. package/skills/ba/SKILL.md +72 -2
  229. package/skills/brainstorm/SKILL.md +342 -341
  230. package/skills/browser-pilot/SKILL.md +168 -168
  231. package/skills/constraint-check/SKILL.md +165 -165
  232. package/skills/context-engine/SKILL.md +404 -404
  233. package/skills/cook/SKILL.md +917 -834
  234. package/skills/cook/references/output-format.md +33 -0
  235. package/skills/db/SKILL.md +273 -272
  236. package/skills/debug/SKILL.md +465 -443
  237. package/skills/dependency-doctor/SKILL.md +265 -235
  238. package/skills/deploy/SKILL.md +274 -231
  239. package/skills/design/DESIGN-REFERENCE.md +365 -365
  240. package/skills/design/SKILL.md +589 -482
  241. package/skills/doc-processor/SKILL.md +254 -254
  242. package/skills/docs/SKILL.md +374 -373
  243. package/skills/docs-seeker/SKILL.md +177 -177
  244. package/skills/fix/SKILL.md +330 -308
  245. package/skills/git/SKILL.md +339 -339
  246. package/skills/graft/SKILL.md +352 -0
  247. package/skills/graft/references/challenge-framework.md +98 -0
  248. package/skills/graft/references/mode-decision.md +44 -0
  249. package/skills/hallucination-guard/SKILL.md +219 -219
  250. package/skills/incident/SKILL.md +254 -251
  251. package/skills/integrity-check/SKILL.md +169 -169
  252. package/skills/journal/SKILL.md +240 -238
  253. package/skills/launch/SKILL.md +344 -342
  254. package/skills/logic-guardian/SKILL.md +251 -251
  255. package/skills/marketing/SKILL.md +290 -245
  256. package/skills/mcp-builder/SKILL.md +425 -423
  257. package/skills/mcp-builder/references/auto-discovery-pattern.md +169 -0
  258. package/skills/neural-memory/SKILL.md +362 -362
  259. package/skills/onboard/SKILL.md +404 -403
  260. package/skills/perf/SKILL.md +346 -346
  261. package/skills/plan/SKILL.md +433 -370
  262. package/skills/plan/references/feature-map.md +84 -0
  263. package/skills/preflight/SKILL.md +415 -396
  264. package/skills/problem-solver/SKILL.md +380 -284
  265. package/skills/rescue/SKILL.md +474 -450
  266. package/skills/retro/SKILL.md +5 -1
  267. package/skills/review/SKILL.md +612 -535
  268. package/skills/review-intake/SKILL.md +249 -249
  269. package/skills/safeguard/SKILL.md +200 -200
  270. package/skills/sast/SKILL.md +190 -190
  271. package/skills/scaffold/SKILL.md +328 -286
  272. package/skills/scope-guard/SKILL.md +180 -162
  273. package/skills/scout/SKILL.md +263 -263
  274. package/skills/sentinel/SKILL.md +382 -353
  275. package/skills/sentinel-env/SKILL.md +254 -254
  276. package/skills/sequential-thinking/SKILL.md +234 -234
  277. package/skills/session-bridge/SKILL.md +543 -397
  278. package/skills/skill-forge/SKILL.md +581 -539
  279. package/skills/skill-router/{skill.md → SKILL.md} +30 -2
  280. package/skills/surgeon/SKILL.md +215 -215
  281. package/skills/team/SKILL.md +556 -514
  282. package/skills/test/SKILL.md +614 -587
  283. package/skills/trend-scout/SKILL.md +145 -145
  284. package/skills/verification/SKILL.md +326 -325
  285. package/skills/video-creator/SKILL.md +201 -201
  286. package/skills/watchdog/SKILL.md +168 -168
  287. package/skills/worktree/SKILL.md +140 -140
@@ -1,325 +1,326 @@
1
- ---
2
- name: verification
3
- description: "Universal verification runner. Runs lint, type-check, tests, and build. Use after any code change to verify nothing is broken."
4
- metadata:
5
- author: runedev
6
- version: "0.5.0"
7
- layer: L3
8
- model: haiku
9
- group: validation
10
- tools: "Read, Bash, Glob, Grep"
11
- listen: code.changed
12
- ---
13
-
14
- # verification
15
-
16
- Runs all automated checks to verify code health. Stateless — runs checks and reports results.
17
-
18
- ## Instructions
19
-
20
- ### Phase 1: Detect Project Type
21
-
22
- Use `Glob` to find project config files:
23
-
24
- 1. Check for `package.json` → Node.js/TypeScript project
25
- 2. Check for `pyproject.toml` or `setup.py` → Python project
26
- 3. Check for `Cargo.toml` → Rust project
27
- 4. Check for `go.mod` → Go project
28
- 5. Check for `pom.xml` or `build.gradle` Java project
29
-
30
- Use `Read` on the detected config file to find scripts or tool config (e.g., `package.json` scripts block for custom lint/test commands).
31
-
32
- ```
33
- TodoWrite: [
34
- { content: "Detect project type", status: "in_progress" },
35
- { content: "Run lint check", status: "pending" },
36
- { content: "Run type check", status: "pending" },
37
- { content: "Run test suite", status: "pending" },
38
- { content: "Run build", status: "pending" },
39
- { content: "Generate verification report", status: "pending" }
40
- ]
41
- ```
42
-
43
- ### Phase 2: Run Lint
44
-
45
- Use `Bash` to run the appropriate linter. If `package.json` has a `lint` script, prefer that:
46
-
47
- - **Node.js (npm lint script)**: `npm run lint`
48
- - **Node.js (no script)**: `npx eslint . --max-warnings 0`
49
- - **Python**: `ruff check .` (fallback: `flake8 .`)
50
- - **Rust**: `cargo clippy -- -D warnings`
51
- - **Go**: `golangci-lint run` (fallback: `go vet ./...`)
52
-
53
- If lint fails: record the failure output, mark lint as FAIL, continue to next step. Do NOT stop.
54
-
55
- **Verification gate**: Command exits without crashing (even if it reports lint errors — those are FAIL, not errors).
56
-
57
- ### Phase 3: Run Type Check
58
-
59
- Use `Bash`:
60
-
61
- - **TypeScript**: `npx tsc --noEmit`
62
- - **Python**: `mypy .` (fallback: `pyright .`)
63
- - **Rust**: `cargo check`
64
- - **Go**: `go vet ./...`
65
-
66
- If type check fails: record error count and first 10 error lines, mark as FAIL, continue.
67
-
68
- ### Phase 4: Run Tests
69
-
70
- Use `Bash` to run the test suite. Prefer the project script if available:
71
-
72
- - **Node.js (npm test script)**: `npm test`
73
- - **Vitest**: `npx vitest run`
74
- - **Jest**: `npx jest --passWithNoTests`
75
- - **Python**: `pytest -v` (fallback: `python -m unittest discover`)
76
- - **Rust**: `cargo test`
77
- - **Go**: `go test ./...`
78
-
79
- Record: total tests, passed count, failed count, coverage percentage if output includes it.
80
-
81
- If tests fail: record which tests failed (first 20), mark as FAIL, continue to build.
82
-
83
- ### Phase 5: Run Build
84
-
85
- Use `Bash`:
86
-
87
- - **Node.js**: check `package.json` for `build` script → `npm run build` (fallback: `npx tsc`)
88
- - **Python**: check `pyproject.toml` for `[build-system]` section:
89
- - If build backend found (setuptools, poetry-core, hatchling, flit-core): `python -m build --no-isolation 2>&1 | head -20` to verify packaging
90
- - If `setup.py` exists (legacy): `python setup.py check --strict`
91
- - Then always: `pip install -e . --dry-run` to catch broken entry points, missing `__init__.py`, or import path issues
92
- - If no `pyproject.toml` and no `setup.py` (scripts-only project): SKIP
93
- - **Rust**: `cargo build`
94
- - **Go**: `go build ./...`
95
-
96
- If build fails: record first 20 lines of build output, mark as FAIL.
97
-
98
- ### Phase 6: Generate Report
99
-
100
- Compile all results into the structured report. Update all TodoWrite items to completed.
101
-
102
- ### 3-Level Artifact Verification
103
-
104
- Every file created or modified during implementation must pass ALL 3 levels:
105
-
106
- **Level 1 — EXISTS**: File is on disk, non-empty.
107
- ```
108
- Glob("path/to/expected/file") → found
109
- ```
110
-
111
- **Level 2 — SUBSTANTIVE**: Contains real logic, NOT a stub. Scan for these stub patterns:
112
-
113
- | Pattern | Language | Meaning |
114
- |---|---|---|
115
- | Component returns only `<div>Placeholder</div>` or `<div>TODO</div>` | React/Vue | Stub component |
116
- | Route returns `{ message: "Not implemented" }` or `res.status(501)` | API | Stub endpoint |
117
- | Function body is only `return null` / `return {}` / `return []` / `pass` | Any | Stub function |
118
- | Class with all methods throwing `NotImplementedError` | Python/Java | Stub class |
119
- | `useEffect` with empty body / `async function` with no `await` | React/JS | Hollow implementation |
120
- | File has only type/interface exports but no implementation | TypeScript | Stub types-only file |
121
- | `// TODO` or `# TODO` as the only content in a function | Any | Placeholder |
122
-
123
- If ANY stub pattern detected → mark file as STUB, Level 2 FAIL.
124
-
125
- **Level 3 — WIRED**: Actually imported/called/used by the rest of the system.
126
-
127
- | File Type | Wiring Check |
128
- |---|---|
129
- | Component | `Grep("<ComponentName")` in parent files → ≥1 consumer |
130
- | API route | `Grep("fetch\\|axios\\|api.*endpoint")` for this path → ≥1 caller |
131
- | Hook | `Grep("useHookName(")` → ≥1 consumer |
132
- | Utility function | `Grep("import.*from.*this-file")` → ≥1 importer |
133
- | DB model/schema | `Grep("ModelName\\|table_name")` in query files → ≥1 reference |
134
- | CSS/style module | `Grep("import.*from.*this-style")` → ≥1 importer |
135
-
136
- If file has 0 consumers → mark as UNWIRED, Level 3 FAIL.
137
-
138
- **Exception**: Entry-point files (main.ts, index.ts, App.tsx, routes config) are exempt from Level 3 — they ARE the top-level consumers.
139
-
140
- <HARD-GATE name="3-level-verification">
141
- ALL new files must pass Level 1 + Level 2 + Level 3.
142
- EXISTS but STUB = "Existence Theater" agent created files but didn't implement them.
143
- EXISTS and SUBSTANTIVE but UNWIRED = dead code — created but never connected.
144
- Report which level failed for each file in the Verification Report.
145
- </HARD-GATE>
146
-
147
- ### Artifact Output Verification
148
-
149
- > Inspired by CLI-Anything (HKUDS/CLI-Anything, 14.5k★): "Never trust exit 0."
150
- > Many tools exit 0 even when they fail silently. Always verify ACTUAL output.
151
-
152
- After each phase command, verify that the expected artifact or indicator is present:
153
-
154
- **Test output** — scan stdout for the pass/fail summary line:
155
- - Vitest/Jest: look for `X passed`, `X failed` if neither appears, output is incomplete
156
- - Pytest: look for `X passed` or `X failed` — exit 0 with no summary = runner crashed silently
157
- - If only exit code available and no summary line found mark as INCOMPLETE, not PASS
158
-
159
- **Build output** — after `npm run build` / `cargo build` / `go build`:
160
- - Verify the output file exists: `Glob("dist/**/*.js")` or equivalent
161
- - Verify file size > 0 bytes: a zero-byte output = silent truncation failure
162
- - If output directory is missing FAIL even if command exited 0
163
-
164
- **Lint output** — parse stdout for counts, not just exit code:
165
- - ESLint: look for `X problems (Y errors, Z warnings)` `0 problems` = PASS
166
- - Ruff/Flake8: zero output lines = PASS; any file:line output = FAIL
167
- - If linter exits 0 but output contains `error` keyword log as suspicious, mark WARN
168
-
169
- **Generated files** — check magic bytes for binary outputs:
170
- - PDF: first bytes must be `%PDF` use `Bash("head -c 4 file.pdf")`
171
- - ZIP/XLSX/DOCX: first bytes must be `PK` (ZIP magic) — use `Bash("head -c 2 file.zip")`
172
- - File size must exceed minimum threshold (PDF > 1KB, ZIP > 100 bytes)
173
-
174
- **Type check** — do not trust exit code alone:
175
- - TypeScript `tsc --noEmit`: look for `Found X errors` or absence of error lines
176
- - `Found 0 errors` = PASS; any other count = FAIL
177
- - Empty output from `tsc` = PASS (no errors emitted) note explicitly
178
-
179
- <HARD-GATE name="artifact-verification">
180
- Verification MUST check actual command output for success indicators, not just exit codes.
181
- Exit 0 without a confirming output artifact or success string = UNVERIFIED.
182
- Report the specific line that confirmed success (e.g., "3 passed, 0 failed").
183
- </HARD-GATE>
184
-
185
- ## Error Recovery
186
-
187
- - If project type cannot be detected: report "Unknown project type" and skip all checks
188
- - If a command is not found (e.g., `ruff` not installed): note "tool not installed", mark check as SKIP
189
- - If a command hangs for more than 60 seconds: kill it, mark check as TIMEOUT, continue
190
-
191
- ## Calls (outbound)
192
-
193
- None — pure runner using Bash for all checks. Does not invoke other skills.
194
-
195
- ## Called By (inbound)
196
-
197
- - `cook` (L1): Phase 6 VERIFY — final check before commit
198
- - `fix` (L2): validate fix doesn't break existing functionality
199
- - `test` (L2): validate test coverage meets threshold
200
- - `deploy` (L2): post-deploy health checks
201
- - `sentinel` (L2): run security audit tools (npm audit, etc.)
202
- - `safeguard` (L2): verify safety net is solid before refactoring
203
- - `db` (L2): run migration in test environment
204
- - `perf` (L2): run benchmark scripts if configured
205
- - `skill-forge` (L2): verify newly created skill passes lint/type/build checks
206
-
207
- ## Output Format
208
-
209
- ```
210
- VERIFICATION REPORT
211
- ===================
212
- Lint: [PASS/FAIL/SKIP] ([details])
213
- Types: [PASS/FAIL/SKIP] ([X errors])
214
- Tests: [PASS/FAIL/SKIP] ([passed]/[total], [coverage]%)
215
- Build: [PASS/FAIL/SKIP]
216
-
217
- ### 3-Level File Verification
218
- | File | L1 Exists | L2 Substantive | L3 Wired | Verdict |
219
- |------|-----------|----------------|----------|---------|
220
- | src/auth/login.ts | ✓ | ✓ | ✓ (imported by routes.ts) | PASS |
221
- | src/auth/reset.ts | ✓ | STUB (returns null) | | FAIL L2 |
222
- | src/utils/format.ts | ✓ | | UNWIRED (0 importers) | FAIL L3 |
223
-
224
- Overall: [PASS/FAIL]
225
-
226
- ### Failures (if any)
227
- - Lint: [error details with file:line]
228
- - Types: [first 5 type errors]
229
- - Tests: [first 5 failing test names]
230
- - Build: [first 5 build errors]
231
- - Stubs: [files that failed Level 2 with stub pattern detected]
232
- - Unwired: [files that failed Level 3 with 0 consumers]
233
- ```
234
-
235
- ## Output Completion Enforcement
236
-
237
- > From taste-skill (Leonxlnx/taste-skill, 3.4k★): Truncated code is worse than no code — it passes reviews but breaks at runtime.
238
-
239
- When verifying code files (Level 2 SUBSTANTIVE check), also scan for **truncation patterns** — signs that the agent generated partial output and stopped:
240
-
241
- | Banned Pattern | Language | What It Means |
242
- |---|---|---|
243
- | `// ...` or `/* ... */` as a statement | JS/TS | Agent truncated remaining code |
244
- | `# ...` as a statement (not comment) | Python | Agent truncated |
245
- | `// rest of code` / `// remaining implementation` | Any | Explicit truncation admission |
246
- | `// TODO: implement` as sole function body | Any | Placeholder, not implementation |
247
- | `{ /* same as above */ }` | JS/TS | Copy-paste truncation |
248
- | `...` (bare ellipsis, not spread operator) | JS/TS/Python | Truncation marker |
249
- | `[PAUSED]` / `[CONTINUED]` in source | Any | Agent session marker leaked into code |
250
-
251
- **Action on detection:**
252
- - Mark file as TRUNCATED (distinct from STUB) in Verification Report
253
- - TRUNCATED files are Level 2 FAIL they CANNOT pass verification
254
- - Report the specific line number and pattern detected
255
- - If agent claims "done" with truncated files → REJECTED by Evidence-Before-Claims gate
256
-
257
- **Continuation protocol** — if the agent hit output limits mid-file:
258
- - Agent MUST log: `[PAUSED X of Y functions complete]` in its response (NOT in the code file)
259
- - Agent MUST resume and complete the file in the next turn
260
- - Verification re-runs after completion to clear the TRUNCATED flag
261
-
262
- ## Evidence-Before-Claims Gate
263
-
264
- <HARD-GATE>
265
- An agent MUST NOT claim "done", "fixed", "passing", or "verified" without showing the actual command output that proves it.
266
- "I ran the tests and they pass" WITHOUT stdout/stderr = UNVERIFIED CLAIM = REJECTED.
267
- The verification report IS the evidence. No report = no verification happened.
268
- </HARD-GATE>
269
-
270
- ### Claim Validation Protocol
271
-
272
- When any skill calls verification and then reports results upstream:
273
-
274
- 1. **Output capture is mandatory** — every Bash command's stdout/stderr must appear in the report
275
- 2. **Pass requires proof** — PASS means "tool ran AND output shows zero errors" (not "tool ran without crashing")
276
- 3. **Silence is not success** — if a command produces no output, note it explicitly ("0 errors, 0 warnings")
277
- 4. **Partial runs are labeled** — if only 2 of 4 checks ran, Overall = INCOMPLETE (not PASS)
278
-
279
- ### Red Flags — Agent is Lying
280
-
281
- | Claim | Without | Verdict |
282
- |---|---|---|
283
- | "All tests pass" | Test runner stdout showing pass count | REJECTED — re-run and show output |
284
- | "No lint errors" | Linter stdout | REJECTED — re-run and show output |
285
- | "Build succeeds" | Build command stdout | REJECTED — re-run and show output |
286
- | "I verified it" | Verification Report | REJECTED — run verification skill properly |
287
- | "Fixed and working" | Before/after test output | REJECTED — show the diff in results |
288
-
289
- ## Constraints
290
-
291
- 1. MUST run ALL four checks: lint, type-check, tests, build — not just tests
292
- 2. MUST show actual command output never claim "all passed" without evidence
293
- 3. MUST report specific failures with file:line references
294
- 4. MUST NOT skip checks because "changes are small"
295
- 5. MUST include stdout/stderr capture in every check result — empty output noted explicitly
296
- 6. MUST mark Overall as INCOMPLETE if any check was skipped without valid reason (tool not installed = valid, "changes are small" = invalid)
297
-
298
- ## Sharp Edges
299
-
300
- Known failure modes for this skill. Check these before declaring done.
301
-
302
- | Failure Mode | Severity | Mitigation |
303
- |---|---|---|
304
- | Claiming "all passed" without showing actual command output | CRITICAL | Evidence-Before-Claims HARD-GATE blocks this — stdout/stderr is mandatory |
305
- | Agent says "verified" without producing Verification Report | CRITICAL | No report = no verification. Re-run the skill properly. |
306
- | Skipping build because "changes are small" | HIGH | Constraint 4: all four checks mandatory size of changes doesn't matter |
307
- | Marking check as PASS when the tool isn't installed | MEDIUM | Mark as SKIP (not PASS) PASS means the tool ran and reported clean |
308
- | Stopping after first failure instead of running remaining checks | MEDIUM | Run all checks; aggregate all failures so developer can fix everything at once |
309
- | Reporting PASS when output has warnings but zero errors | LOW | PASS is correct but note warning count caller decides if warnings matter |
310
- | Trusting exit code 0 without output verification | CRITICAL | Artifact Verification HARD-GATE: always confirm success indicator in stdout (pass count, "0 errors", output file exists) |
311
- | Existence Theater file exists but is a stub | HIGH | 3-Level check: Level 2 scans for stub patterns (`<div>Placeholder</div>`, `return null`, `NotImplementedError`) |
312
- | Dead code — file created but never imported/used | MEDIUM | 3-Level check: Level 3 greps for consumers. 0 importers = UNWIRED |
313
- | Truncated code — agent hit output limit mid-file | HIGH | Output Completion Enforcement: scan for `// ...`, `// rest of code`, bare ellipsis patterns. TRUNCATED = Level 2 FAIL |
314
-
315
- ## Done When
316
-
317
- - Project type detected from config files
318
- - lint, type-check, tests, and build all executed (or SKIP with reason if tool missing)
319
- - Each check shows actual command output
320
- - Failures include specific file:line references (not just counts)
321
- - Verification Report emitted with Overall PASS/FAIL verdict
322
-
323
- ## Cost Profile
324
-
325
- ~$0.01-0.03 per run. Haiku + Bash commands. Fast and cheap.
1
+ ---
2
+ name: verification
3
+ description: "Universal verification runner. Runs lint, type-check, tests, and build. Use after any code change to verify nothing is broken."
4
+ metadata:
5
+ author: runedev
6
+ version: "0.5.0"
7
+ layer: L3
8
+ model: haiku
9
+ group: validation
10
+ tools: "Read, Bash, Glob, Grep"
11
+ listen: code.changed
12
+ emit: verification.complete
13
+ ---
14
+
15
+ # verification
16
+
17
+ Runs all automated checks to verify code health. Stateless — runs checks and reports results.
18
+
19
+ ## Instructions
20
+
21
+ ### Phase 1: Detect Project Type
22
+
23
+ Use `Glob` to find project config files:
24
+
25
+ 1. Check for `package.json` Node.js/TypeScript project
26
+ 2. Check for `pyproject.toml` or `setup.py` Python project
27
+ 3. Check for `Cargo.toml` → Rust project
28
+ 4. Check for `go.mod` → Go project
29
+ 5. Check for `pom.xml` or `build.gradle` → Java project
30
+
31
+ Use `Read` on the detected config file to find scripts or tool config (e.g., `package.json` scripts block for custom lint/test commands).
32
+
33
+ ```
34
+ TodoWrite: [
35
+ { content: "Detect project type", status: "in_progress" },
36
+ { content: "Run lint check", status: "pending" },
37
+ { content: "Run type check", status: "pending" },
38
+ { content: "Run test suite", status: "pending" },
39
+ { content: "Run build", status: "pending" },
40
+ { content: "Generate verification report", status: "pending" }
41
+ ]
42
+ ```
43
+
44
+ ### Phase 2: Run Lint
45
+
46
+ Use `Bash` to run the appropriate linter. If `package.json` has a `lint` script, prefer that:
47
+
48
+ - **Node.js (npm lint script)**: `npm run lint`
49
+ - **Node.js (no script)**: `npx eslint . --max-warnings 0`
50
+ - **Python**: `ruff check .` (fallback: `flake8 .`)
51
+ - **Rust**: `cargo clippy -- -D warnings`
52
+ - **Go**: `golangci-lint run` (fallback: `go vet ./...`)
53
+
54
+ If lint fails: record the failure output, mark lint as FAIL, continue to next step. Do NOT stop.
55
+
56
+ **Verification gate**: Command exits without crashing (even if it reports lint errors — those are FAIL, not errors).
57
+
58
+ ### Phase 3: Run Type Check
59
+
60
+ Use `Bash`:
61
+
62
+ - **TypeScript**: `npx tsc --noEmit`
63
+ - **Python**: `mypy .` (fallback: `pyright .`)
64
+ - **Rust**: `cargo check`
65
+ - **Go**: `go vet ./...`
66
+
67
+ If type check fails: record error count and first 10 error lines, mark as FAIL, continue.
68
+
69
+ ### Phase 4: Run Tests
70
+
71
+ Use `Bash` to run the test suite. Prefer the project script if available:
72
+
73
+ - **Node.js (npm test script)**: `npm test`
74
+ - **Vitest**: `npx vitest run`
75
+ - **Jest**: `npx jest --passWithNoTests`
76
+ - **Python**: `pytest -v` (fallback: `python -m unittest discover`)
77
+ - **Rust**: `cargo test`
78
+ - **Go**: `go test ./...`
79
+
80
+ Record: total tests, passed count, failed count, coverage percentage if output includes it.
81
+
82
+ If tests fail: record which tests failed (first 20), mark as FAIL, continue to build.
83
+
84
+ ### Phase 5: Run Build
85
+
86
+ Use `Bash`:
87
+
88
+ - **Node.js**: check `package.json` for `build` script → `npm run build` (fallback: `npx tsc`)
89
+ - **Python**: check `pyproject.toml` for `[build-system]` section:
90
+ - If build backend found (setuptools, poetry-core, hatchling, flit-core): `python -m build --no-isolation 2>&1 | head -20` to verify packaging
91
+ - If `setup.py` exists (legacy): `python setup.py check --strict`
92
+ - Then always: `pip install -e . --dry-run` to catch broken entry points, missing `__init__.py`, or import path issues
93
+ - If no `pyproject.toml` and no `setup.py` (scripts-only project): SKIP
94
+ - **Rust**: `cargo build`
95
+ - **Go**: `go build ./...`
96
+
97
+ If build fails: record first 20 lines of build output, mark as FAIL.
98
+
99
+ ### Phase 6: Generate Report
100
+
101
+ Compile all results into the structured report. Update all TodoWrite items to completed.
102
+
103
+ ### 3-Level Artifact Verification
104
+
105
+ Every file created or modified during implementation must pass ALL 3 levels:
106
+
107
+ **Level 1 — EXISTS**: File is on disk, non-empty.
108
+ ```
109
+ Glob("path/to/expected/file") → found
110
+ ```
111
+
112
+ **Level 2 — SUBSTANTIVE**: Contains real logic, NOT a stub. Scan for these stub patterns:
113
+
114
+ | Pattern | Language | Meaning |
115
+ |---|---|---|
116
+ | Component returns only `<div>Placeholder</div>` or `<div>TODO</div>` | React/Vue | Stub component |
117
+ | Route returns `{ message: "Not implemented" }` or `res.status(501)` | API | Stub endpoint |
118
+ | Function body is only `return null` / `return {}` / `return []` / `pass` | Any | Stub function |
119
+ | Class with all methods throwing `NotImplementedError` | Python/Java | Stub class |
120
+ | `useEffect` with empty body / `async function` with no `await` | React/JS | Hollow implementation |
121
+ | File has only type/interface exports but no implementation | TypeScript | Stub types-only file |
122
+ | `// TODO` or `# TODO` as the only content in a function | Any | Placeholder |
123
+
124
+ If ANY stub pattern detected → mark file as STUB, Level 2 FAIL.
125
+
126
+ **Level 3 — WIRED**: Actually imported/called/used by the rest of the system.
127
+
128
+ | File Type | Wiring Check |
129
+ |---|---|
130
+ | Component | `Grep("<ComponentName")` in parent files → ≥1 consumer |
131
+ | API route | `Grep("fetch\\|axios\\|api.*endpoint")` for this path → ≥1 caller |
132
+ | Hook | `Grep("useHookName(")` → ≥1 consumer |
133
+ | Utility function | `Grep("import.*from.*this-file")` → ≥1 importer |
134
+ | DB model/schema | `Grep("ModelName\\|table_name")` in query files → ≥1 reference |
135
+ | CSS/style module | `Grep("import.*from.*this-style")` → ≥1 importer |
136
+
137
+ If file has 0 consumers → mark as UNWIRED, Level 3 FAIL.
138
+
139
+ **Exception**: Entry-point files (main.ts, index.ts, App.tsx, routes config) are exempt from Level 3 — they ARE the top-level consumers.
140
+
141
+ <HARD-GATE name="3-level-verification">
142
+ ALL new files must pass Level 1 + Level 2 + Level 3.
143
+ EXISTS but STUB = "Existence Theater"agent created files but didn't implement them.
144
+ EXISTS and SUBSTANTIVE but UNWIRED = dead code created but never connected.
145
+ Report which level failed for each file in the Verification Report.
146
+ </HARD-GATE>
147
+
148
+ ### Artifact Output Verification
149
+
150
+ > Inspired by CLI-Anything (HKUDS/CLI-Anything, 14.5k★): "Never trust exit 0."
151
+ > Many tools exit 0 even when they fail silently. Always verify ACTUAL output.
152
+
153
+ After each phase command, verify that the expected artifact or indicator is present:
154
+
155
+ **Test output** scan stdout for the pass/fail summary line:
156
+ - Vitest/Jest: look for `X passed`, `X failed` — if neither appears, output is incomplete
157
+ - Pytest: look for `X passed` or `X failed` exit 0 with no summary = runner crashed silently
158
+ - If only exit code available and no summary line found → mark as INCOMPLETE, not PASS
159
+
160
+ **Build output** after `npm run build` / `cargo build` / `go build`:
161
+ - Verify the output file exists: `Glob("dist/**/*.js")` or equivalent
162
+ - Verify file size > 0 bytes: a zero-byte output = silent truncation failure
163
+ - If output directory is missing → FAIL even if command exited 0
164
+
165
+ **Lint output** parse stdout for counts, not just exit code:
166
+ - ESLint: look for `X problems (Y errors, Z warnings)` — `0 problems` = PASS
167
+ - Ruff/Flake8: zero output lines = PASS; any file:line output = FAIL
168
+ - If linter exits 0 but output contains `error` keyword → log as suspicious, mark WARN
169
+
170
+ **Generated files** check magic bytes for binary outputs:
171
+ - PDF: first bytes must be `%PDF` — use `Bash("head -c 4 file.pdf")`
172
+ - ZIP/XLSX/DOCX: first bytes must be `PK` (ZIP magic) use `Bash("head -c 2 file.zip")`
173
+ - File size must exceed minimum threshold (PDF > 1KB, ZIP > 100 bytes)
174
+
175
+ **Type check** do not trust exit code alone:
176
+ - TypeScript `tsc --noEmit`: look for `Found X errors` or absence of error lines
177
+ - `Found 0 errors` = PASS; any other count = FAIL
178
+ - Empty output from `tsc` = PASS (no errors emitted) — note explicitly
179
+
180
+ <HARD-GATE name="artifact-verification">
181
+ Verification MUST check actual command output for success indicators, not just exit codes.
182
+ Exit 0 without a confirming output artifact or success string = UNVERIFIED.
183
+ Report the specific line that confirmed success (e.g., "3 passed, 0 failed").
184
+ </HARD-GATE>
185
+
186
+ ## Error Recovery
187
+
188
+ - If project type cannot be detected: report "Unknown project type" and skip all checks
189
+ - If a command is not found (e.g., `ruff` not installed): note "tool not installed", mark check as SKIP
190
+ - If a command hangs for more than 60 seconds: kill it, mark check as TIMEOUT, continue
191
+
192
+ ## Calls (outbound)
193
+
194
+ None — pure runner using Bash for all checks. Does not invoke other skills.
195
+
196
+ ## Called By (inbound)
197
+
198
+ - `cook` (L1): Phase 6 VERIFY final check before commit
199
+ - `fix` (L2): validate fix doesn't break existing functionality
200
+ - `test` (L2): validate test coverage meets threshold
201
+ - `deploy` (L2): post-deploy health checks
202
+ - `sentinel` (L2): run security audit tools (npm audit, etc.)
203
+ - `safeguard` (L2): verify safety net is solid before refactoring
204
+ - `db` (L2): run migration in test environment
205
+ - `perf` (L2): run benchmark scripts if configured
206
+ - `skill-forge` (L2): verify newly created skill passes lint/type/build checks
207
+
208
+ ## Output Format
209
+
210
+ ```
211
+ VERIFICATION REPORT
212
+ ===================
213
+ Lint: [PASS/FAIL/SKIP] ([details])
214
+ Types: [PASS/FAIL/SKIP] ([X errors])
215
+ Tests: [PASS/FAIL/SKIP] ([passed]/[total], [coverage]%)
216
+ Build: [PASS/FAIL/SKIP]
217
+
218
+ ### 3-Level File Verification
219
+ | File | L1 Exists | L2 Substantive | L3 Wired | Verdict |
220
+ |------|-----------|----------------|----------|---------|
221
+ | src/auth/login.ts | ✓ | | ✓ (imported by routes.ts) | PASS |
222
+ | src/auth/reset.ts | ✓ | STUB (returns null) | — | FAIL L2 |
223
+ | src/utils/format.ts | ✓ | ✓ | UNWIRED (0 importers) | FAIL L3 |
224
+
225
+ Overall: [PASS/FAIL]
226
+
227
+ ### Failures (if any)
228
+ - Lint: [error details with file:line]
229
+ - Types: [first 5 type errors]
230
+ - Tests: [first 5 failing test names]
231
+ - Build: [first 5 build errors]
232
+ - Stubs: [files that failed Level 2 with stub pattern detected]
233
+ - Unwired: [files that failed Level 3 with 0 consumers]
234
+ ```
235
+
236
+ ## Output Completion Enforcement
237
+
238
+ > From taste-skill (Leonxlnx/taste-skill, 3.4k★): Truncated code is worse than no code — it passes reviews but breaks at runtime.
239
+
240
+ When verifying code files (Level 2 SUBSTANTIVE check), also scan for **truncation patterns** — signs that the agent generated partial output and stopped:
241
+
242
+ | Banned Pattern | Language | What It Means |
243
+ |---|---|---|
244
+ | `// ...` or `/* ... */` as a statement | JS/TS | Agent truncated remaining code |
245
+ | `# ...` as a statement (not comment) | Python | Agent truncated |
246
+ | `// rest of code` / `// remaining implementation` | Any | Explicit truncation admission |
247
+ | `// TODO: implement` as sole function body | Any | Placeholder, not implementation |
248
+ | `{ /* same as above */ }` | JS/TS | Copy-paste truncation |
249
+ | `...` (bare ellipsis, not spread operator) | JS/TS/Python | Truncation marker |
250
+ | `[PAUSED]` / `[CONTINUED]` in source | Any | Agent session marker leaked into code |
251
+
252
+ **Action on detection:**
253
+ - Mark file as TRUNCATED (distinct from STUB) in Verification Report
254
+ - TRUNCATED files are Level 2 FAIL they CANNOT pass verification
255
+ - Report the specific line number and pattern detected
256
+ - If agent claims "done" with truncated files → REJECTED by Evidence-Before-Claims gate
257
+
258
+ **Continuation protocol**if the agent hit output limits mid-file:
259
+ - Agent MUST log: `[PAUSED — X of Y functions complete]` in its response (NOT in the code file)
260
+ - Agent MUST resume and complete the file in the next turn
261
+ - Verification re-runs after completion to clear the TRUNCATED flag
262
+
263
+ ## Evidence-Before-Claims Gate
264
+
265
+ <HARD-GATE>
266
+ An agent MUST NOT claim "done", "fixed", "passing", or "verified" without showing the actual command output that proves it.
267
+ "I ran the tests and they pass" WITHOUT stdout/stderr = UNVERIFIED CLAIM = REJECTED.
268
+ The verification report IS the evidence. No report = no verification happened.
269
+ </HARD-GATE>
270
+
271
+ ### Claim Validation Protocol
272
+
273
+ When any skill calls verification and then reports results upstream:
274
+
275
+ 1. **Output capture is mandatory** — every Bash command's stdout/stderr must appear in the report
276
+ 2. **Pass requires proof** — PASS means "tool ran AND output shows zero errors" (not "tool ran without crashing")
277
+ 3. **Silence is not success** — if a command produces no output, note it explicitly ("0 errors, 0 warnings")
278
+ 4. **Partial runs are labeled** — if only 2 of 4 checks ran, Overall = INCOMPLETE (not PASS)
279
+
280
+ ### Red Flags — Agent is Lying
281
+
282
+ | Claim | Without | Verdict |
283
+ |---|---|---|
284
+ | "All tests pass" | Test runner stdout showing pass count | REJECTED — re-run and show output |
285
+ | "No lint errors" | Linter stdout | REJECTED — re-run and show output |
286
+ | "Build succeeds" | Build command stdout | REJECTED — re-run and show output |
287
+ | "I verified it" | Verification Report | REJECTED — run verification skill properly |
288
+ | "Fixed and working" | Before/after test output | REJECTED — show the diff in results |
289
+
290
+ ## Constraints
291
+
292
+ 1. MUST run ALL four checks: lint, type-check, tests, build not just tests
293
+ 2. MUST show actual command output never claim "all passed" without evidence
294
+ 3. MUST report specific failures with file:line references
295
+ 4. MUST NOT skip checks because "changes are small"
296
+ 5. MUST include stdout/stderr capture in every check result empty output noted explicitly
297
+ 6. MUST mark Overall as INCOMPLETE if any check was skipped without valid reason (tool not installed = valid, "changes are small" = invalid)
298
+
299
+ ## Sharp Edges
300
+
301
+ Known failure modes for this skill. Check these before declaring done.
302
+
303
+ | Failure Mode | Severity | Mitigation |
304
+ |---|---|---|
305
+ | Claiming "all passed" without showing actual command output | CRITICAL | Evidence-Before-Claims HARD-GATE blocks this stdout/stderr is mandatory |
306
+ | Agent says "verified" without producing Verification Report | CRITICAL | No report = no verification. Re-run the skill properly. |
307
+ | Skipping build because "changes are small" | HIGH | Constraint 4: all four checks mandatorysize of changes doesn't matter |
308
+ | Marking check as PASS when the tool isn't installed | MEDIUM | Mark as SKIP (not PASS) PASS means the tool ran and reported clean |
309
+ | Stopping after first failure instead of running remaining checks | MEDIUM | Run all checks; aggregate all failures so developer can fix everything at once |
310
+ | Reporting PASS when output has warnings but zero errors | LOW | PASS is correct but note warning count caller decides if warnings matter |
311
+ | Trusting exit code 0 without output verification | CRITICAL | Artifact Verification HARD-GATE: always confirm success indicator in stdout (pass count, "0 errors", output file exists) |
312
+ | Existence Theater — file exists but is a stub | HIGH | 3-Level check: Level 2 scans for stub patterns (`<div>Placeholder</div>`, `return null`, `NotImplementedError`) |
313
+ | Dead code — file created but never imported/used | MEDIUM | 3-Level check: Level 3 greps for consumers. 0 importers = UNWIRED |
314
+ | Truncated code — agent hit output limit mid-file | HIGH | Output Completion Enforcement: scan for `// ...`, `// rest of code`, bare ellipsis patterns. TRUNCATED = Level 2 FAIL |
315
+
316
+ ## Done When
317
+
318
+ - Project type detected from config files
319
+ - lint, type-check, tests, and build all executed (or SKIP with reason if tool missing)
320
+ - Each check shows actual command output
321
+ - Failures include specific file:line references (not just counts)
322
+ - Verification Report emitted with Overall PASS/FAIL verdict
323
+
324
+ ## Cost Profile
325
+
326
+ ~$0.01-0.03 per run. Haiku + Bash commands. Fast and cheap.