@opengsd/gsd-core 1.9.1 → 1.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (426) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/.claude-plugin/plugin.json +2 -3
  3. package/.opencode/plugins/gsd-core.js +8 -1
  4. package/agents/gsd-code-fixer.md +27 -3
  5. package/agents/gsd-debug-session-manager.md +11 -0
  6. package/agents/gsd-debugger.md +12 -246
  7. package/agents/gsd-doc-synthesizer.md +2 -4
  8. package/agents/gsd-executor.md +12 -10
  9. package/agents/gsd-integration-checker.md +3 -0
  10. package/agents/gsd-mempalace-curator.md +5 -2
  11. package/agents/gsd-phase-researcher.md +20 -1
  12. package/agents/gsd-plan-checker.md +46 -0
  13. package/agents/gsd-planner.md +49 -54
  14. package/agents/gsd-roadmapper.md +21 -3
  15. package/agents/gsd-user-profiler.md +3 -0
  16. package/agents/gsd-verifier.md +26 -73
  17. package/bin/install.js +1272 -1238
  18. package/bin/lib/ui-safety-gate.cjs +2 -0
  19. package/commands/gsd/code-review.md +1 -1
  20. package/commands/gsd/execute-phase.md +1 -1
  21. package/commands/gsd/map-codebase.md +1 -1
  22. package/commands/gsd/mempalace-capture.md +2 -2
  23. package/commands/gsd/mempalace-recall.md +1 -1
  24. package/commands/gsd/new-milestone.md +2 -2
  25. package/commands/gsd/plan-phase.md +1 -1
  26. package/commands/gsd/quick.md +1 -1
  27. package/commands/gsd/review-backlog.md +2 -1
  28. package/commands/gsd/verify-work.md +1 -1
  29. package/gsd-core/bin/gsd-tools.cjs +1009 -115
  30. package/gsd-core/bin/lib/active-workstream-store.cjs +153 -12
  31. package/gsd-core/bin/lib/agent-install-check.cjs +268 -38
  32. package/gsd-core/bin/lib/api-coverage.cjs +123 -5
  33. package/gsd-core/bin/lib/artifacts.cjs +3 -0
  34. package/gsd-core/bin/lib/assumption-delta.cjs +2 -4
  35. package/gsd-core/bin/lib/audit-command-router.cjs +9 -2
  36. package/gsd-core/bin/lib/audit.cjs +926 -202
  37. package/gsd-core/bin/lib/broken-windows.cjs +36 -6
  38. package/gsd-core/bin/lib/capability-consent.cjs +149 -15
  39. package/gsd-core/bin/lib/capability-lifecycle.cjs +45 -0
  40. package/gsd-core/bin/lib/capability-registry.cjs +608 -148
  41. package/gsd-core/bin/lib/capability-source.cjs +92 -0
  42. package/gsd-core/bin/lib/capability-trust.cjs +444 -25
  43. package/gsd-core/bin/lib/capability-validator.cjs +507 -24
  44. package/gsd-core/bin/lib/capability-writer.cjs +3 -2
  45. package/gsd-core/bin/lib/check-command-router.cjs +114 -38
  46. package/gsd-core/bin/lib/claude-orchestration.cjs +56 -3
  47. package/gsd-core/bin/lib/codex-agent-toml.cjs +329 -0
  48. package/gsd-core/bin/lib/command-aliases.cjs +94 -0
  49. package/gsd-core/bin/lib/command-roster.cjs +44 -1
  50. package/gsd-core/bin/lib/commands.cjs +665 -99
  51. package/gsd-core/bin/lib/commonjs-marker.cjs +142 -0
  52. package/gsd-core/bin/lib/complexity-trigger.cjs +1172 -0
  53. package/gsd-core/bin/lib/config-loader.cjs +76 -0
  54. package/gsd-core/bin/lib/config.cjs +22 -2
  55. package/gsd-core/bin/lib/context-composer.cjs +278 -0
  56. package/gsd-core/bin/lib/context-predicates.cjs +506 -0
  57. package/gsd-core/bin/lib/core-utils.cjs +217 -40
  58. package/gsd-core/bin/lib/decisions.cjs +23 -0
  59. package/gsd-core/bin/lib/docs.cjs +3 -2
  60. package/gsd-core/bin/lib/external-job.cjs +19 -4
  61. package/gsd-core/bin/lib/fallow-runner.cjs +20 -44
  62. package/gsd-core/bin/lib/frontmatter.cjs +239 -32
  63. package/gsd-core/bin/lib/gap-checker.cjs +68 -7
  64. package/gsd-core/bin/lib/gate-predicate-evaluator.cjs +57 -6
  65. package/gsd-core/bin/lib/git-base-branch.cjs +160 -15
  66. package/gsd-core/bin/lib/graphify.cjs +142 -27
  67. package/gsd-core/bin/lib/gsd2-import.cjs +37 -5
  68. package/gsd-core/bin/lib/health-diagnostic-rules/agent-install.cjs +101 -0
  69. package/gsd-core/bin/lib/health-diagnostic-rules/config-validation.cjs +348 -0
  70. package/gsd-core/bin/lib/health-diagnostic-rules/consistency.cjs +145 -0
  71. package/gsd-core/bin/lib/health-diagnostic-rules/install-surface-shadowing.cjs +98 -0
  72. package/gsd-core/bin/lib/health-diagnostic-rules/milestone-archive-hygiene.cjs +100 -0
  73. package/gsd-core/bin/lib/health-diagnostic-rules/phase-structure.cjs +222 -0
  74. package/gsd-core/bin/lib/health-diagnostic-rules/roadmap-disk-consistency.cjs +265 -0
  75. package/gsd-core/bin/lib/health-diagnostic-rules/root-existence.cjs +161 -0
  76. package/gsd-core/bin/lib/health-diagnostic-rules/state-consistency.cjs +303 -0
  77. package/gsd-core/bin/lib/health-diagnostic-rules/worktree-health.cjs +173 -0
  78. package/gsd-core/bin/lib/health-diagnostic-types.cjs +68 -0
  79. package/gsd-core/bin/lib/health-diagnostic.cjs +431 -0
  80. package/gsd-core/bin/lib/host-integration.cjs +13 -1
  81. package/gsd-core/bin/lib/host-runtime-detection.cjs +134 -0
  82. package/gsd-core/bin/lib/init-command-router.cjs +83 -8
  83. package/gsd-core/bin/lib/init.cjs +1325 -169
  84. package/gsd-core/bin/lib/install-effort-resolver.cjs +73 -30
  85. package/gsd-core/bin/lib/install-engine.cjs +805 -264
  86. package/gsd-core/bin/lib/install-fs-adapter.cjs +262 -0
  87. package/gsd-core/bin/lib/install-model-override-resolver.cjs +203 -0
  88. package/gsd-core/bin/lib/install-profiles.cjs +160 -57
  89. package/gsd-core/bin/lib/install-scope.cjs +270 -0
  90. package/gsd-core/bin/lib/install-shadow-report.cjs +385 -0
  91. package/gsd-core/bin/lib/installed-surface-resolver.cjs +381 -0
  92. package/gsd-core/bin/lib/installer-migration-authoring.cjs +3 -1
  93. package/gsd-core/bin/lib/installer-migration-report.cjs +4 -0
  94. package/gsd-core/bin/lib/installer-migrations/007-retire-config-root-commonjs-marker.cjs +149 -0
  95. package/gsd-core/bin/lib/installer-migrations/008-cursor-retire-commands-surface.cjs +55 -0
  96. package/gsd-core/bin/lib/installer-migrations/009-pi-retire-reserved-hooks-dir.cjs +199 -0
  97. package/gsd-core/bin/lib/installer-migrations.cjs +206 -13
  98. package/gsd-core/bin/lib/io.cjs +38 -3
  99. package/gsd-core/bin/lib/markdown-sectionizer.cjs +8 -1
  100. package/gsd-core/bin/lib/markdown-table.cjs +133 -20
  101. package/gsd-core/bin/lib/mcp-catalog.cjs +518 -0
  102. package/gsd-core/bin/lib/mcp-server.cjs +135 -3
  103. package/gsd-core/bin/lib/milestone-lock.cjs +248 -0
  104. package/gsd-core/bin/lib/milestone.cjs +821 -109
  105. package/gsd-core/bin/lib/model-catalog.cjs +59 -1
  106. package/gsd-core/bin/lib/model-resolver.cjs +183 -40
  107. package/gsd-core/bin/lib/normalize-test-command.cjs +1 -1
  108. package/gsd-core/bin/lib/pattern.cjs +122 -0
  109. package/gsd-core/bin/lib/phase-estimation.cjs +1 -1
  110. package/gsd-core/bin/lib/phase-id.cjs +507 -36
  111. package/gsd-core/bin/lib/phase-lifecycle.cjs +28 -3
  112. package/gsd-core/bin/lib/phase-locator.cjs +258 -58
  113. package/gsd-core/bin/lib/phase.cjs +891 -156
  114. package/gsd-core/bin/lib/plan-dependency-graph.cjs +303 -0
  115. package/gsd-core/bin/lib/plan-drift-guard.cjs +120 -0
  116. package/gsd-core/bin/lib/plan-scan.cjs +86 -2
  117. package/gsd-core/bin/lib/planning-scope.cjs +31 -0
  118. package/gsd-core/bin/lib/planning-snapshot.cjs +890 -0
  119. package/gsd-core/bin/lib/planning-workspace.cjs +60 -6
  120. package/gsd-core/bin/lib/probe-core.cjs +1 -1
  121. package/gsd-core/bin/lib/profile-output.cjs +1 -1
  122. package/gsd-core/bin/lib/prompt-budget.cjs +128 -165
  123. package/gsd-core/bin/lib/refactor-trigger-command-router.cjs +740 -0
  124. package/gsd-core/bin/lib/retired-artifact-cleanup.cjs +85 -0
  125. package/gsd-core/bin/lib/review-lane-descriptor.cjs +108 -0
  126. package/gsd-core/bin/lib/review-lane-invocation.cjs +30 -0
  127. package/gsd-core/bin/lib/review-lane-runner.cjs +447 -68
  128. package/gsd-core/bin/lib/review-reviewer-selection.cjs +13 -18
  129. package/gsd-core/bin/lib/roadmap-command-router.cjs +76 -9
  130. package/gsd-core/bin/lib/roadmap-parser.cjs +1035 -194
  131. package/gsd-core/bin/lib/roadmap-upgrade.cjs +37 -10
  132. package/gsd-core/bin/lib/roadmap.cjs +405 -84
  133. package/gsd-core/bin/lib/runtime-artifact-conversion.cjs +795 -100
  134. package/gsd-core/bin/lib/runtime-artifact-install-plan.cjs +14 -2
  135. package/gsd-core/bin/lib/runtime-artifact-layout.cjs +440 -57
  136. package/gsd-core/bin/lib/runtime-config-adapter-registry.cjs +3 -2
  137. package/gsd-core/bin/lib/runtime-homes.cjs +220 -41
  138. package/gsd-core/bin/lib/runtime-hooks-surface.cjs +220 -44
  139. package/gsd-core/bin/lib/runtime-name-policy.cjs +3 -1
  140. package/gsd-core/bin/lib/runtime-slash.cjs +27 -9
  141. package/gsd-core/bin/lib/section-manifest.cjs +209 -0
  142. package/gsd-core/bin/lib/security.cjs +104 -5
  143. package/gsd-core/bin/lib/shell-command-projection.cjs +388 -30
  144. package/gsd-core/bin/lib/smart-entry.cjs +154 -22
  145. package/gsd-core/bin/lib/state-command-router.cjs +5 -1
  146. package/gsd-core/bin/lib/state-document.cjs +152 -8
  147. package/gsd-core/bin/lib/state-transition.cjs +424 -105
  148. package/gsd-core/bin/lib/state.cjs +1927 -401
  149. package/gsd-core/bin/lib/surface.cjs +35 -10
  150. package/gsd-core/bin/lib/text-lines.cjs +80 -0
  151. package/gsd-core/bin/lib/token-scanner.cjs +76 -0
  152. package/gsd-core/bin/lib/uat-predicate.cjs +20 -4
  153. package/gsd-core/bin/lib/uat.cjs +706 -64
  154. package/gsd-core/bin/lib/ui-frontend-evidence.cjs +157 -0
  155. package/gsd-core/bin/lib/ui-safety-gate.cjs +14 -5
  156. package/gsd-core/bin/lib/unusable-input.cjs +33 -0
  157. package/gsd-core/bin/lib/update-context.cjs +8 -2
  158. package/gsd-core/bin/lib/user-artifact-staging.cjs +705 -0
  159. package/gsd-core/bin/lib/validate.cjs +20 -6
  160. package/gsd-core/bin/lib/vendor/README.md +37 -0
  161. package/gsd-core/bin/lib/vendor/re2js.cjs +6480 -0
  162. package/gsd-core/bin/lib/vendor/re2js.d.cts +938 -0
  163. package/gsd-core/bin/lib/verification-command-router.cjs +2 -1
  164. package/gsd-core/bin/lib/verification.cjs +287 -20
  165. package/gsd-core/bin/lib/verify.cjs +368 -880
  166. package/gsd-core/bin/lib/workflow-fragments.cjs +557 -0
  167. package/gsd-core/bin/lib/workstream-inventory-builder.cjs +203 -19
  168. package/gsd-core/bin/lib/workstream-inventory.cjs +576 -31
  169. package/gsd-core/bin/lib/workstream.cjs +8 -2
  170. package/gsd-core/bin/lib/worktree-base-ref.cjs +50 -6
  171. package/gsd-core/bin/lib/worktree-safety.cjs +450 -125
  172. package/gsd-core/bin/shared/config-defaults.manifest.json +1 -0
  173. package/gsd-core/bin/shared/config-schema.manifest.json +9 -1
  174. package/gsd-core/references/agent-contracts.md +43 -26
  175. package/gsd-core/references/artifact-types.md +10 -3
  176. package/gsd-core/references/autonomous-ui-design-contract.md +42 -0
  177. package/gsd-core/references/checkpoints.md +2 -2
  178. package/gsd-core/references/context-budget.md +1 -1
  179. package/gsd-core/references/debugger-techniques.md +255 -0
  180. package/gsd-core/references/dispatch-isolation-gate.md +138 -0
  181. package/gsd-core/references/doc-conflict-engine.md +1 -1
  182. package/gsd-core/references/execute-mvp-tdd.md +3 -3
  183. package/gsd-core/references/execute-phase-between-wave-reset.md +6 -2
  184. package/gsd-core/references/execute-phase-context-guard.md +1 -1
  185. package/gsd-core/references/execute-phase-response-language.md +1 -1
  186. package/gsd-core/references/execute-phase-wave-guard.md +6 -2
  187. package/gsd-core/references/gate-prompts.md +1 -1
  188. package/gsd-core/references/git-planning-commit.md +2 -1
  189. package/gsd-core/references/loop-hook-dispatch.md +39 -2
  190. package/gsd-core/references/model-profiles.md +12 -4
  191. package/gsd-core/references/mvp-concepts.md +9 -9
  192. package/gsd-core/references/planner-guidance.md +3 -9
  193. package/gsd-core/references/planner-preconditions.md +1 -1
  194. package/gsd-core/references/planner-reviews.md +1 -1
  195. package/gsd-core/references/planning-config.md +8 -6
  196. package/gsd-core/references/research-documentation-lookup.md +5 -3
  197. package/gsd-core/references/revision-loop.md +1 -1
  198. package/gsd-core/references/specless-probe-fallback.md +8 -7
  199. package/gsd-core/references/universal-anti-patterns.md +3 -3
  200. package/gsd-core/references/verifier-phase-gates.md +192 -0
  201. package/gsd-core/references/verifier-wiring-patterns.md +100 -0
  202. package/gsd-core/references/verify-mvp-mode.md +1 -1
  203. package/gsd-core/references/workstream-flag.md +22 -6
  204. package/gsd-core/references/worktree-branch-check.md +2 -2
  205. package/gsd-core/templates/discussion-log.md +1 -1
  206. package/gsd-core/templates/phase-prompt.md +2 -4
  207. package/gsd-core/templates/state.md +4 -4
  208. package/gsd-core/templates/summary-complex.md +2 -0
  209. package/gsd-core/templates/summary-minimal.md +2 -0
  210. package/gsd-core/templates/summary-standard.md +2 -0
  211. package/gsd-core/templates/summary.md +2 -0
  212. package/gsd-core/templates/verification-report.md +9 -1
  213. package/gsd-core/workflows/ai-integration-phase.md +9 -11
  214. package/gsd-core/workflows/audit-milestone.md +3 -0
  215. package/gsd-core/workflows/autonomous/steps/converge-banner.md +1 -0
  216. package/gsd-core/workflows/autonomous/steps/converge-dispatch-bg.md +11 -0
  217. package/gsd-core/workflows/autonomous/steps/converge-dispatch-inline.md +7 -0
  218. package/gsd-core/workflows/autonomous/steps/converge-fail-fast.md +21 -0
  219. package/gsd-core/workflows/autonomous/steps/converge-loop.md +7 -0
  220. package/gsd-core/workflows/autonomous.md +33 -70
  221. package/gsd-core/workflows/cleanup.md +62 -3
  222. package/gsd-core/workflows/code-review/steps/dispatch-fix.md +39 -0
  223. package/gsd-core/workflows/code-review/steps/structural-pre-pass.md +93 -0
  224. package/gsd-core/workflows/code-review-fix.md +37 -10
  225. package/gsd-core/workflows/code-review.md +74 -166
  226. package/gsd-core/workflows/complete-milestone/steps/git-tag.md +29 -0
  227. package/gsd-core/workflows/complete-milestone.md +160 -95
  228. package/gsd-core/workflows/debug.md +16 -17
  229. package/gsd-core/workflows/diagnose-issues.md +56 -8
  230. package/gsd-core/workflows/discuss-phase/modes/chain.md +2 -1
  231. package/gsd-core/workflows/discuss-phase/modes/default.md +1 -1
  232. package/gsd-core/workflows/discuss-phase-assumptions/steps/auto-advance-dispatch.md +15 -0
  233. package/gsd-core/workflows/discuss-phase-assumptions.md +7 -17
  234. package/gsd-core/workflows/docs-update/steps/dispatch-monorepo-packages.md +51 -0
  235. package/gsd-core/workflows/docs-update.md +8 -51
  236. package/gsd-core/workflows/edit-phase.md +26 -1
  237. package/gsd-core/workflows/eval-review.md +3 -5
  238. package/gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md +64 -7
  239. package/gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md +50 -0
  240. package/gsd-core/workflows/execute-phase/steps/partial-wave.md +31 -0
  241. package/gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md +77 -0
  242. package/gsd-core/workflows/execute-phase/steps/per-plan-worktree-gate.md +21 -0
  243. package/gsd-core/workflows/execute-phase/steps/regression-gate-run.md +42 -0
  244. package/gsd-core/workflows/execute-phase/steps/regression-gate.md +43 -37
  245. package/gsd-core/workflows/execute-phase.md +103 -187
  246. package/gsd-core/workflows/execute-plan.md +36 -4
  247. package/gsd-core/workflows/explore.md +131 -4
  248. package/gsd-core/workflows/fast.md +10 -2
  249. package/gsd-core/workflows/health.md +73 -4
  250. package/gsd-core/workflows/help/modes/full.md +6 -1
  251. package/gsd-core/workflows/import.md +4 -4
  252. package/gsd-core/workflows/ingest-docs.md +7 -6
  253. package/gsd-core/workflows/mvp-phase.md +6 -3
  254. package/gsd-core/workflows/new-milestone/steps/project-md-milestone-write.md +16 -0
  255. package/gsd-core/workflows/new-milestone/steps/reset-phase-safety.md +19 -0
  256. package/gsd-core/workflows/new-milestone.md +35 -47
  257. package/gsd-core/workflows/new-project/steps/auto-mode-config.md +176 -0
  258. package/gsd-core/workflows/new-project/steps/auto-mode-detection.md +32 -0
  259. package/gsd-core/workflows/new-project/steps/codebase-map-offer.md +18 -0
  260. package/gsd-core/workflows/new-project.md +27 -240
  261. package/gsd-core/workflows/next.md +12 -0
  262. package/gsd-core/workflows/plan-phase/steps/adr-ingest-express-path.md +15 -0
  263. package/gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md +110 -0
  264. package/gsd-core/workflows/plan-phase/steps/prd-express-gate.md +8 -0
  265. package/gsd-core/workflows/plan-phase/steps/research-only-early-exit.md +17 -0
  266. package/gsd-core/workflows/plan-phase/steps/research-only-modifiers.md +16 -0
  267. package/gsd-core/workflows/plan-phase/steps/reviews-prerequisite.md +17 -0
  268. package/gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md +149 -0
  269. package/gsd-core/workflows/plan-phase.md +89 -209
  270. package/gsd-core/workflows/plan-review-convergence.md +50 -2
  271. package/gsd-core/workflows/progress/steps/forensic-audit.md +125 -0
  272. package/gsd-core/workflows/progress/steps/mvp-display.md +18 -0
  273. package/gsd-core/workflows/progress.md +45 -159
  274. package/gsd-core/workflows/quick/steps/discussion-phase.md +124 -0
  275. package/gsd-core/workflows/quick/steps/plan-checker-loop.md +111 -0
  276. package/gsd-core/workflows/quick/steps/quick-verification.md +67 -0
  277. package/gsd-core/workflows/quick/steps/research-phase.md +72 -0
  278. package/gsd-core/workflows/quick/steps/worktree-pre-dispatch-commit.md +37 -0
  279. package/gsd-core/workflows/quick.md +55 -405
  280. package/gsd-core/workflows/resume-project.md +3 -0
  281. package/gsd-core/workflows/review/steps/reviewer-instances-note-1.md +4 -0
  282. package/gsd-core/workflows/review/steps/reviewer-instances-note-2.md +3 -0
  283. package/gsd-core/workflows/review.md +41 -13
  284. package/gsd-core/workflows/section-manifest.json +219 -0
  285. package/gsd-core/workflows/secure-phase.md +1 -1
  286. package/gsd-core/workflows/session-report.md +2 -1
  287. package/gsd-core/workflows/settings.md +66 -2
  288. package/gsd-core/workflows/ship.md +104 -44
  289. package/gsd-core/workflows/sketch.md +1 -1
  290. package/gsd-core/workflows/spec-phase.md +41 -20
  291. package/gsd-core/workflows/spike-wrap-up.md +20 -5
  292. package/gsd-core/workflows/spike.md +50 -16
  293. package/gsd-core/workflows/sync-skills.md +106 -13
  294. package/gsd-core/workflows/transition/steps/workstream-collision-check.md +17 -0
  295. package/gsd-core/workflows/transition.md +53 -31
  296. package/gsd-core/workflows/ui-phase.md +13 -12
  297. package/gsd-core/workflows/ui-review.md +2 -2
  298. package/gsd-core/workflows/update/steps/channel-banner.md +7 -0
  299. package/gsd-core/workflows/update.md +19 -8
  300. package/gsd-core/workflows/validate-phase.md +1 -1
  301. package/gsd-core/workflows/verify-work/steps/automated-ui-verification.md +36 -0
  302. package/gsd-core/workflows/verify-work/steps/mvp-uat-framing.md +21 -0
  303. package/gsd-core/workflows/verify-work.md +17 -65
  304. package/hooks/dist/gsd-agent-isolation-guard.js +517 -0
  305. package/hooks/dist/gsd-check-update-worker.js +64 -12
  306. package/hooks/dist/gsd-check-update.js +19 -1
  307. package/hooks/dist/gsd-cursor-pre-tool.js +0 -3
  308. package/hooks/dist/gsd-cursor-subagent-start.js +607 -26
  309. package/hooks/dist/gsd-cursor-subagent-stop.js +3 -2
  310. package/hooks/dist/gsd-prompt-guard.js +21 -20
  311. package/hooks/dist/gsd-read-injection-scanner.js +45 -24
  312. package/hooks/dist/gsd-statusline.js +90 -6
  313. package/hooks/dist/gsd-update-banner.js +22 -1
  314. package/hooks/dist/gsd-workflow-guard.js +134 -36
  315. package/hooks/dist/gsd-worktree-path-guard.js +2 -1
  316. package/hooks/dist/gsd-write-guard.js +359 -0
  317. package/hooks/dist/lib/git-cmd.js +92 -59
  318. package/hooks/dist/lib/injection-patterns.js +45 -0
  319. package/hooks/dist/lib/isolation-deny-reason.js +39 -0
  320. package/hooks/dist/lib/isolation-sentinel.js +277 -0
  321. package/hooks/dist/managed-hooks-registry.cjs +2 -0
  322. package/hooks/gsd-agent-isolation-guard.js +517 -0
  323. package/hooks/gsd-check-update-worker.js +64 -12
  324. package/hooks/gsd-check-update.js +19 -1
  325. package/hooks/gsd-cursor-pre-tool.js +0 -3
  326. package/hooks/gsd-cursor-subagent-start.js +607 -26
  327. package/hooks/gsd-cursor-subagent-stop.js +3 -2
  328. package/hooks/gsd-prompt-guard.js +21 -20
  329. package/hooks/gsd-read-injection-scanner.js +45 -24
  330. package/hooks/gsd-statusline.js +90 -6
  331. package/hooks/gsd-update-banner.js +22 -1
  332. package/hooks/gsd-workflow-guard.js +134 -36
  333. package/hooks/gsd-worktree-path-guard.js +2 -1
  334. package/hooks/gsd-write-guard.js +359 -0
  335. package/hooks/hooks.json +12 -0
  336. package/hooks/lib/git-cmd.js +92 -59
  337. package/hooks/lib/injection-patterns.js +45 -0
  338. package/hooks/lib/isolation-deny-reason.js +39 -0
  339. package/hooks/lib/isolation-sentinel.js +277 -0
  340. package/hooks/managed-hooks-registry.cjs +2 -0
  341. package/package.json +31 -10
  342. package/pi/gsd.cjs +71 -12
  343. package/scripts/baselines/planning-prompt-drift-baseline.json +4 -0
  344. package/scripts/baselines/planning-snapshot-bypass-baseline.json +12 -0
  345. package/scripts/baselines/unreachable-guard-drift-baseline.json +4 -0
  346. package/scripts/build-hooks.js +9 -0
  347. package/scripts/changeset/lint.cjs +68 -6
  348. package/scripts/changeset/serialize.cjs +5 -1
  349. package/scripts/check-alias-drift.cjs +7 -43
  350. package/scripts/check-contract-drift.cjs +297 -0
  351. package/scripts/ci-test-scope.cjs +19 -2
  352. package/scripts/command-contract-helpers.cjs +903 -1
  353. package/scripts/gen-adr-index.cjs +728 -38
  354. package/scripts/gen-capability-matrix.cjs +1 -1
  355. package/scripts/gen-capability-registry.cjs +3 -15
  356. package/scripts/gen-context-index.cjs +439 -0
  357. package/scripts/gen-health-docs.cjs +390 -0
  358. package/scripts/gen-inventory-manifest.cjs +150 -4
  359. package/scripts/gen-loop-host-contract.cjs +4 -24
  360. package/scripts/gen-prompt-budget-parity-corpus.cjs +645 -0
  361. package/scripts/gen-registry.cjs +3 -14
  362. package/scripts/gen-section-manifest.cjs +638 -0
  363. package/scripts/generate-package-identity.cjs +4 -2
  364. package/scripts/lib/alias-drift-families.cjs +46 -0
  365. package/scripts/lib/drift-scan.cjs +278 -0
  366. package/scripts/lint-allow-test-rule-refs.allowlist.json +15 -54
  367. package/scripts/lint-allow-test-rule-refs.effective-ceiling.json +4 -0
  368. package/scripts/lint-allow-test-rule-refs.unverified-ceiling.json +3 -0
  369. package/scripts/lint-canary-version-leak.cjs +73 -0
  370. package/scripts/lint-command-contract.cjs +96 -13
  371. package/scripts/lint-compiled-artifact-sync.cjs +6 -1
  372. package/scripts/lint-completion-predicate-drift.cjs +933 -0
  373. package/scripts/lint-completion-ratio-drift.cjs +214 -0
  374. package/scripts/lint-default-flip-documentation.cjs +193 -0
  375. package/scripts/lint-docs-command-form.cjs +195 -0
  376. package/scripts/lint-docs-required.cjs +9 -1
  377. package/scripts/lint-emitted-drift-ack.cjs +215 -20
  378. package/scripts/lint-eslint-glob-coverage.allowlist.json +34 -0
  379. package/scripts/lint-eslint-glob-coverage.cjs +340 -0
  380. package/scripts/lint-example-parser-parity.cjs +395 -0
  381. package/scripts/lint-frontmatter-scalar-broad-grep.cjs +237 -0
  382. package/scripts/lint-health-diagnostic-rule-table.cjs +404 -0
  383. package/scripts/lint-hooks-runtime-build-seam.cjs +262 -0
  384. package/scripts/lint-milestone-window-drift.cjs +468 -0
  385. package/scripts/lint-phase-enumeration-drift.cjs +479 -0
  386. package/scripts/lint-plan-count-drift.cjs +318 -0
  387. package/scripts/lint-planning-artifact-writer-drift.cjs +398 -0
  388. package/scripts/lint-planning-prompt-drift.cjs +434 -0
  389. package/scripts/lint-planning-snapshot-bypass-drift.cjs +544 -0
  390. package/scripts/lint-regression-test-names.cjs +15 -13
  391. package/scripts/lint-removed-but-needed.cjs +320 -0
  392. package/scripts/lint-state-field-drift.cjs +805 -0
  393. package/scripts/lint-state-write-path-drift.cjs +1045 -0
  394. package/scripts/lint-test-file-count.allowlist.json +40 -3
  395. package/scripts/lint-unreachable-guard-drift.cjs +843 -0
  396. package/scripts/lint-vendored-deps.cjs +124 -0
  397. package/scripts/mutation-matrix.cjs +13 -0
  398. package/scripts/pr-changed-files.cjs +63 -0
  399. package/scripts/pr-template-policy.cjs +14 -4
  400. package/scripts/prompt-injection-scan.sh +52 -6
  401. package/scripts/require-issue-link-policy.cjs +192 -0
  402. package/scripts/state-write-path-drift-baseline.json +19 -0
  403. package/scripts/sync-runtime-launcher.cjs +2 -4
  404. package/skills/gsd-autonomous/SKILL.md +0 -1
  405. package/skills/gsd-code-review/SKILL.md +1 -1
  406. package/skills/gsd-execute-phase/SKILL.md +1 -2
  407. package/skills/gsd-map-codebase/SKILL.md +1 -1
  408. package/skills/gsd-mempalace-capture/SKILL.md +2 -2
  409. package/skills/gsd-mempalace-recall/SKILL.md +1 -1
  410. package/skills/gsd-new-milestone/SKILL.md +2 -2
  411. package/skills/gsd-next/SKILL.md +0 -1
  412. package/skills/gsd-plan-phase/SKILL.md +1 -2
  413. package/skills/gsd-progress/SKILL.md +0 -1
  414. package/skills/gsd-quick/SKILL.md +1 -1
  415. package/skills/gsd-review-backlog/SKILL.md +2 -1
  416. package/skills/gsd-stats/SKILL.md +0 -1
  417. package/skills/gsd-verify-work/SKILL.md +1 -1
  418. package/vscode/package.json +1 -1
  419. package/gsd-core/workflows/discovery-phase.md +0 -298
  420. package/gsd-core/workflows/plan-milestone-gaps.md +0 -281
  421. package/gsd-core/workflows/verify-phase.md +0 -577
  422. package/scripts/affected-tests-lib.cjs +0 -554
  423. package/scripts/gen-emitted-baseline.cjs +0 -145
  424. package/scripts/lint-allow-test-rule-refs.cjs +0 -162
  425. package/scripts/run-affected-tests.cjs +0 -7
  426. package/scripts/run-tests.cjs +0 -1050
@@ -1,577 +0,0 @@
1
- <purpose>
2
- Verify phase goal achievement through goal-backward analysis. Check that the codebase delivers what the phase promised, not just that tasks completed.
3
-
4
- Executed by a verification subagent spawned from execute-phase.md.
5
- </purpose>
6
-
7
- <core_principle>
8
- **Task completion ≠ Goal achievement**
9
-
10
- A task "create chat component" can be marked complete when the component is a placeholder. The task was done — but the goal "working chat interface" was not achieved.
11
-
12
- Goal-backward verification:
13
- 1. What must be TRUE for the goal to be achieved?
14
- 2. What must EXIST for those truths to hold?
15
- 3. What must be WIRED for those artifacts to function?
16
- 4. What must TESTS PROVE for those truths to be evidenced?
17
-
18
- Then verify each level against the actual codebase.
19
- </core_principle>
20
-
21
- <required_reading>
22
- @~/.claude/gsd-core/references/verification-patterns.md
23
- @~/.claude/gsd-core/templates/verification-report.md
24
- </required_reading>
25
-
26
- <process>
27
-
28
- <step name="load_context" priority="first">
29
- Load phase operation context:
30
-
31
- ```bash
32
- _GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; if [ -f "$GSD_TOOLS" ]; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif command -v gsd-tools >/dev/null 2>&1; then GSD_TOOLS="$(command -v gsd-tools)"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif [ -f "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd-tools is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi
33
- INIT=$(gsd_run query init.phase-op "${PHASE_ARG}")
34
- if [[ "$INIT" == @file:* ]]; then INIT=$(cat "${INIT#@file:}"); fi
35
- ```
36
-
37
- Extract from init JSON: `phase_dir`, `phase_number`, `phase_name`, `has_plans`, `plan_count`.
38
-
39
- Then load phase details and list plans/summaries:
40
- ```bash
41
- gsd_run query roadmap.get-phase "${phase_number}"
42
- grep -E "^| ${phase_number}" .planning/REQUIREMENTS.md 2>/dev/null || true
43
- ls "$phase_dir"/*-SUMMARY.md "$phase_dir"/*-PLAN.md 2>/dev/null || true
44
- ```
45
-
46
- Load full milestone phases for deferred-item filtering (Step 9b):
47
- ```bash
48
- gsd_run query roadmap.analyze
49
- ```
50
-
51
- Extract **phase goal** from ROADMAP.md (the outcome to verify, not tasks), **requirements** from REQUIREMENTS.md if it exists, and **all milestone phases** from roadmap analyze (for cross-referencing gaps against later phases).
52
- </step>
53
-
54
- <step name="establish_must_haves">
55
- **Option A: Must-haves in PLAN frontmatter**
56
-
57
- Use `gsd-tools.cjs query` verify handlers (or legacy gsd-tools) to extract must_haves from each PLAN:
58
-
59
- ```bash
60
- for plan in "$PHASE_DIR"/*-PLAN.md; do
61
- MUST_HAVES=$(gsd_run query frontmatter.get "$plan" --field must_haves)
62
- echo "=== $plan ===" && echo "$MUST_HAVES"
63
- done
64
- ```
65
-
66
- Returns JSON: `{ truths: [...], artifacts: [...], key_links: [...], prohibitions: [...] }`
67
-
68
- Aggregate all must_haves across plans for phase-level verification.
69
-
70
- **Prohibitions (`must_haves.prohibitions`, ADR-550 D3 — the must-NOT sibling block):** When a plan carries `must_haves.prohibitions`, extract each `{ statement, status, verification }` item and route it by `verification` tier in verdict assembly (ADR-550 D4, "B-with-guard", 2026-06-12 maintainer decision). These are NEGATIVE checks (the must-NOT must NOT have happened), distinct from positive `truths`:
71
-
72
- - **judgment-tier → mode-dependent soft-gate.** Interactive verify defers each item to the end-of-phase human checkpoint (`human_verify_mode: end-of-phase`). Autonomous verify records a NON-AUTHORITATIVE LLM-judge verdict + a prominent `unverified-prohibition — human review recommended` flag (autonomous completion reads "complete with N flagged prohibitions"). NEVER a silent pass; NEVER a hard halt of an AFK run.
73
- - **test-tier → ENFORCED via `check prohibition-enforcement` (green on pass, hard-gate on miss/fail).** Accept the `verification: test` value (the SPEC↔must_haves.prohibitions projection contract holds — no forced schema change later). For each test-tier item, the verifier builds `request.check` **DETERMINISTICALLY from the projected descriptor** — it does NOT invent `{ kind, target, rule }`. Read the flat scalar keys `check_kind` / `check_target` / `check_rule` / `check_violation_fixture` off the `must_haves.prohibitions` item and reconstruct the `CheckDescriptor` via the `descriptorFromProjection` adapter in `prohibition-enforcement` (`descriptorFromProjection(projectedItem)` → `{ kind: check_kind, target: check_target, rule?: check_rule, violationFixture?: check_violation_fixture }`). The `violationFixture` (a path to a KNOWN-BAD subject) is the field that gates **green** and it is **now projected** (`check_violation_fixture`, #1346) — so a prohibition authored with all four scalars greens through the projection alone, **zero hand-authoring at verify time**. Do NOT rely on `failFirst`: it is DEMOTED (#1279) and greens nothing on its own; an item with no projected fixture hard-gates fail-closed. Invoke the producer (CLI surface unchanged):
74
-
75
- ```bash
76
- gsd_run check prohibition-enforcement <request.json>
77
- ```
78
-
79
- where `<request.json>` carries `{ prohibition, check, mode }` — `check` being the wired mechanical-check descriptor `{ kind: 'node-test' | 'lint-rule', target, rule?, violationFixture, cleanFixture?, failFirst? }`, with `kind`/`target`/`rule`/`violationFixture`/`cleanFixture` now sourced from the projected `check_*` scalars (not author/verifier invention — #1278 + #1279 + #1346). For `node-test`, `target` (from `check_target`) is the negative-test file path; for `lint-rule`, `target` is the PATH to lint and `rule` (from `check_rule`) is the eslint rule id (e.g. `local/no-source-grep`) — both required (a lint-rule without `rule` is not a valid wired check). `violationFixture` (from `check_violation_fixture`) is the path to a KNOWN-BAD subject the producer runs the check against to **machine-prove fail-first** (for `node-test`, injected via the `GSD_PROHIB_SUBJECT` env convention — #1279); the optional `cleanFixture` (from `check_clean_fixture`) is a KNOWN-CLEAN control subject the `node-test` prover ALSO requires to stay GREEN, proving the RED is content-caused (#1346); `failFirst` is a DEMOTED, non-authoritative hint kept only for backward route-JSON shape (no path greens on it alone — FF-08). The producer LOCATES the wired check from the projection, **machine-proves it is fail-first** by running it against the violation and confirming it goes RED, RUNS it for a genuine non-vacuous pass, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict (#1259 + #1278 + #1279, ADR-550 D5d). Fail-first is **machine-proven, not caller-attested** — absent a provable violation the producer fails closed, never falling back to attestation. Route the result by its typed fields:
80
- - **`status: 'green'`, `flagged: false`** (a genuinely-passing wired negative test / lint rule, `located: true`, non-empty `evidence`) → the item is satisfiable → it can reach **passed**.
81
- - **missing, non-attested, or genuinely-non-passing check** (`located: false` OR `status: 'unverified'`, `flagged: true`) → **hard-gate**: disposes flagged-unverified, NEVER green, routing to `gaps_found` in BOTH interactive and autonomous modes (a failing mechanical check blocks even AFK; ADR-550 D4 / D3). The deterministic fail-closed default backing every miss/fail is `dispositionForProhibition()` in probe-core (`status: 'unverified'`, `flagged: true` on empty `enforcementEvidence`).
82
-
83
- > **Descriptor source — deterministic locate + machine-proof compose (#1278 + #1346, DELIVERED).** The `check` descriptor's `{ kind, target, rule, violationFixture }` is now sourced **deterministically from the projected `check_kind` / `check_target` / `check_rule` / `check_violation_fixture` scalars** on the `must_haves.prohibitions` item (authored at `/gsd:spec-phase`, projected by `projectProhibitions`, read back via the `descriptorFromProjection` adapter). So both halves close with **zero manual descriptor authoring** — the verifier neither invents the locate (#1278) nor hand-supplies the violation fixture (#1346): a prohibition authored with all four scalars machine-proves fail-first and greens end-to-end through the projection alone (removing the spoofable invent-at-verify-time surface; ADR-857 §147 exogenous grading). **Fail-closed is preserved:** an item with NO projected descriptor, a PARTIAL one (e.g. a `lint-rule` missing `check_rule`), OR a descriptor with **no `check_violation_fixture`** makes `descriptorFromProjection` return `null` / an under-specified or fixture-less descriptor, which falls through to the producer's fail-closed paths (`located: false`, or located-but-unprovable) → flagged-unverified, NEVER green, in BOTH modes. `failFirst` is demoted and greens nothing on its own (#1279, FF-08). Causation (**#1346**): supplying `check_clean_fixture` adds an opt-in control — the `node-test` prover also requires GREEN on a known-clean subject, proving the RED is content-caused; with no clean fixture that one residual case (a deceptive test reding merely because the env var is set) stays a documented constraint, an author opting into the stronger proof by wiring a clean control.
84
-
85
- **Option B: Use Success Criteria from ROADMAP.md**
86
-
87
- If no must_haves in frontmatter (MUST_HAVES returns error or empty), check for Success Criteria:
88
-
89
- ```bash
90
- PHASE_DATA=$(gsd_run query roadmap.get-phase "${phase_number}" --raw)
91
- ```
92
-
93
- Parse the `success_criteria` array from the JSON output. If non-empty:
94
- 1. Use each Success Criterion directly as a **truth** (they are already written as observable, testable behaviors)
95
- 2. Derive **artifacts** (concrete file paths for each truth)
96
- 3. Derive **key links** (critical wiring where stubs hide)
97
- 4. Document the must-haves before proceeding
98
-
99
- Success Criteria from ROADMAP.md are the contract — they override PLAN-level must_haves when both exist.
100
-
101
- **Option C: Derive from phase goal (fallback)**
102
-
103
- If no must_haves in frontmatter AND no Success Criteria in ROADMAP:
104
- 1. State the goal from ROADMAP.md
105
- 2. Derive **truths** (3-7 observable behaviors, each testable)
106
- 3. Derive **artifacts** (concrete file paths for each truth)
107
- 4. Derive **key links** (critical wiring where stubs hide)
108
- 5. Document derived must-haves before proceeding
109
- </step>
110
-
111
- <step name="verify_truths">
112
- For each observable truth, determine if the codebase enables it.
113
-
114
- **Status:** ✓ VERIFIED (all supporting artifacts pass — and, for a behavior-dependent truth, a behavioral test exercises the asserted behavior) | ⚠️ PRESENT_BEHAVIOR_UNVERIFIED (present + wired, but a state transition or cancellation/cleanup/ordering invariant is exercised by no test — routes to human verification, excluded from the score) | ✗ FAILED (artifact missing/stub/unwired) | ? UNCERTAIN (needs human)
115
-
116
- For each truth: identify supporting artifacts → check artifact status → check wiring → determine truth status.
117
-
118
- **Behavior-dependent truths:** when a truth asserts a state transition or a cancellation/cleanup/ordering invariant, symbol presence + wiring is necessary but not sufficient — the code can be present and wired yet still leak state on the path the invariant covers. Mark such a truth ✓ VERIFIED only when a pre-existing test exercises the transition/invariant and passes (one named test, never the full suite); otherwise mark it ⚠️ PRESENT_BEHAVIOR_UNVERIFIED, emit a human-verification item, and exclude it from the verified score.
119
-
120
- **Non-inferable (`backstop`) truths (#1154):** a `must_haves.truths` item in object form `{ statement, verification: backstop }` is non-inferable — the correct behavior is not derivable from the spec alone, so the verifier cannot self-detect the gap and would false-pass it confidently. Branch on the `verification: backstop` field (read via `truthVerification()`, never prose): if confirmable with **explicit evidence** (a passing wired held-out/property test, or a directly-observed behavior) → ✓ VERIFIED; otherwise **abstain** — mark ⚠️ `insufficient_spec`, emit an `unverified — held-out test recommended` human-verification item, exclude from the verified score (routes to `human_needed`). Exogenous only (never a self-judged "abstain if unsure"); an inferable truth is never abstained. See `references/honest-verifier.md`.
121
-
122
- **Example:** Truth "User can see existing messages" depends on Chat.tsx (renders), /api/chat GET (provides), Message model (schema). If Chat.tsx is a stub or API returns hardcoded [] → FAILED. If all exist, are substantive, and connected → VERIFIED.
123
- </step>
124
-
125
- <step name="verify_artifacts">
126
- Use `gsd-tools.cjs query verify.artifacts` (or legacy gsd-tools) for artifact verification against must_haves in each PLAN:
127
-
128
- ```bash
129
- for plan in "$PHASE_DIR"/*-PLAN.md; do
130
- ARTIFACT_RESULT=$(gsd_run query verify.artifacts "$plan")
131
- echo "=== $plan ===" && echo "$ARTIFACT_RESULT"
132
- done
133
- ```
134
-
135
- Parse JSON result: `{ all_passed, passed, total, artifacts: [{path, exists, issues, passed}] }`
136
-
137
- **Artifact status from result:**
138
- - `exists=false` → MISSING
139
- - `issues` not empty → STUB (check issues for "Only N lines" or "Missing pattern")
140
- - `passed=true` → VERIFIED (Levels 1-2 pass)
141
-
142
- **Level 3 — Wired (manual check for artifacts that pass Levels 1-2):**
143
- ```bash
144
- grep -r "import.*$artifact_name" src/ --include="*.ts" --include="*.tsx" # IMPORTED
145
- grep -r "$artifact_name" src/ --include="*.ts" --include="*.tsx" | grep -v "import" # USED
146
- ```
147
- WIRED = imported AND used. ORPHANED = exists but not imported/used.
148
-
149
- | Exists | Substantive | Wired | Status |
150
- |--------|-------------|-------|--------|
151
- | ✓ | ✓ | ✓ | ✓ VERIFIED |
152
- | ✓ | ✓ | ✗ | ⚠️ ORPHANED |
153
- | ✓ | ✗ | - | ✗ STUB |
154
- | ✗ | - | - | ✗ MISSING |
155
-
156
- **Export-level spot check (WARNING severity):**
157
-
158
- For artifacts that pass Level 3, spot-check individual exports:
159
- - Extract key exported symbols (functions, constants, classes — skip types/interfaces)
160
- - For each, grep for usage outside the defining file
161
- - Flag exports with zero external call sites as "exported but unused"
162
-
163
- This catches dead stores like `setPlan()` that exist in a wired file but are
164
- never actually called. Report as WARNING — may indicate incomplete cross-plan
165
- wiring or leftover code from plan revisions.
166
- </step>
167
-
168
- <step name="verify_wiring">
169
- Use `gsd-tools.cjs query verify.key-links` (or legacy gsd-tools) for key link verification against must_haves in each PLAN:
170
-
171
- ```bash
172
- for plan in "$PHASE_DIR"/*-PLAN.md; do
173
- LINKS_RESULT=$(gsd_run query verify.key-links "$plan")
174
- echo "=== $plan ===" && echo "$LINKS_RESULT"
175
- done
176
- ```
177
-
178
- Parse JSON result: `{ all_verified, verified, total, links: [{from, to, via, verified, detail}] }`
179
-
180
- **Link status from result:**
181
- - `verified=true` → WIRED
182
- - `verified=false` with "not found" → NOT_WIRED
183
- - `verified=false` with "Pattern not found" → PARTIAL
184
-
185
- **Fallback patterns (if key_links not in must_haves):**
186
-
187
- | Pattern | Check | Status |
188
- |---------|-------|--------|
189
- | Component → API | fetch/axios call to API path, response used (await/.then/setState) | WIRED / PARTIAL (call but unused response) / NOT_WIRED |
190
- | API → Database | Prisma/DB query on model, result returned via res.json() | WIRED / PARTIAL (query but not returned) / NOT_WIRED |
191
- | Form → Handler | onSubmit with real implementation (fetch/axios/mutate/dispatch), not console.log/empty | WIRED / STUB (log-only/empty) / NOT_WIRED |
192
- | State → Render | useState variable appears in JSX (`{stateVar}` or `{stateVar.property}`) | WIRED / NOT_WIRED |
193
-
194
- Record status and evidence for each key link.
195
- </step>
196
-
197
- <step name="verify_requirements">
198
- If REQUIREMENTS.md exists:
199
- ```bash
200
- grep -E "Phase ${PHASE_NUM}" .planning/REQUIREMENTS.md 2>/dev/null || true
201
- ```
202
-
203
- For each requirement: parse description → identify supporting truths/artifacts → status: ✓ SATISFIED / ✗ BLOCKED / ? NEEDS HUMAN.
204
- </step>
205
-
206
- <step name="verify_decisions">
207
- **Decision coverage validation gate (issue #2492).**
208
-
209
- After requirements coverage, also check that each trackable CONTEXT.md
210
- `<decisions>` entry shows up somewhere in the shipped artifacts (plans,
211
- SUMMARY.md, files modified by the phase, or recent commit subjects on the
212
- phase branch).
213
-
214
- This gate is **non-blocking / warning only** by deliberate asymmetry with
215
- the plan-phase translation gate. The plan-phase gate already blocked at
216
- translation time, so by the time verification runs every decision has
217
- either been translated or explicitly deferred. This gate's job is to
218
- surface decisions that *were* translated but vanished during execution —
219
- that's a soft signal because "honors a decision" is a fuzzy substring
220
- heuristic, and we don't want a paraphrase miss to fail an otherwise good
221
- phase.
222
-
223
- **Skip if** `workflow.context_coverage_gate` is explicitly set to `false`
224
- (absent key = enabled). Also skip cleanly when CONTEXT.md is missing or has
225
- no `<decisions>` block.
226
-
227
- ```bash
228
- GATE_CFG=$(gsd_run query config-get workflow.context_coverage_gate 2>/dev/null || echo "true")
229
- if [ "$GATE_CFG" != "false" ]; then
230
- # Discover the phase CONTEXT.md via glob expansion rather than `ls | head`
231
- # (review F17 / ShellCheck SC2012). Globs preserve filenames containing
232
- # spaces and avoid an extra subprocess.
233
- CONTEXT_PATH=""
234
- for f in "${PHASE_DIR}"/*-CONTEXT.md; do
235
- [ -e "$f" ] && CONTEXT_PATH="$f" && break
236
- done
237
- DECISION_RESULT=$(gsd_run query check.decision-coverage-verify "${PHASE_DIR}" "${CONTEXT_PATH}")
238
- fi
239
- ```
240
-
241
- The handler returns JSON `{ skipped, blocking: false, total, honored,
242
- not_honored: [...], message }`.
243
-
244
- **Reporting:** Append the handler's `message` (a `### Decision Coverage`
245
- section) to VERIFICATION.md regardless of outcome — even when all
246
- decisions are honored, recording the count helps reviewers spot drift over
247
- time. Set `decision_coverage` in the verification result to
248
- `{honored, total, not_honored: [...]}` so downstream tooling can read it.
249
-
250
- **Status impact:** none. The decision gate does NOT influence the
251
- `gaps_found` / `human_needed` / `passed` decision tree in
252
- `determine_status`. Its findings are warnings the user reviews and may act
253
- on by re-opening the phase or by acknowledging the decision was abandoned
254
- intentionally.
255
- </step>
256
-
257
- <step name="behavioral_verification">
258
- **Run the project's test suite and CLI commands to verify behavior, not just structure.**
259
-
260
- Static checks (grep, file existence, wiring) catch structural gaps but miss runtime
261
- failures. This step runs actual tests and project commands to verify the phase goal
262
- is behaviorally achieved.
263
-
264
- This follows Anthropic's harness engineering principle: separating generation from
265
- evaluation, with the evaluator interacting with the running system rather than
266
- inspecting static artifacts.
267
-
268
- **Step 1: Run test suite**
269
-
270
- ```bash
271
- # Resolve test command: project config > Makefile > language sniff
272
- TEST_CMD=$(gsd_run query config-get workflow.test_command --default "" --raw 2>/dev/null || true)
273
- if [ -z "$TEST_CMD" ]; then
274
- if [ -f "Makefile" ] && grep -q "^test:" Makefile; then
275
- TEST_CMD="make test"
276
- elif [ -f "Justfile" ] || [ -f "justfile" ]; then
277
- TEST_CMD="just test"
278
- elif [ -f "package.json" ]; then
279
- TEST_CMD="npm test"
280
- elif [ -f "Cargo.toml" ]; then
281
- TEST_CMD="cargo test"
282
- elif [ -f "go.mod" ]; then
283
- TEST_CMD="go test ./..."
284
- elif [ -f "pyproject.toml" ] || [ -f "requirements.txt" ]; then
285
- TEST_CMD="python -m pytest -q --tb=short 2>&1 || uv run python -m pytest -q --tb=short"
286
- else
287
- TEST_CMD="false"
288
- echo "⚠ No test runner detected — skipping test suite"
289
- fi
290
- fi
291
- # Run all tests (timeout: 5 min). #1857: normalize to one-shot so watch mode exits.
292
- TEST_CMD=$(gsd_run query normalize-test-command "$TEST_CMD" --cwd . 2>/dev/null || echo "$TEST_CMD")
293
- TEST_EXIT=0
294
- gsd_run run-with-timeout 300 -- bash -c "$TEST_CMD" 2>&1
295
- TEST_EXIT=$?
296
- if [ "${TEST_EXIT}" -eq 0 ]; then
297
- echo "✓ Test suite passed"
298
- elif [ "${TEST_EXIT}" -eq 124 ]; then
299
- echo "⚠ Test suite timed out after 5 minutes — likely watch/dev mode"
300
- else
301
- echo "✗ Test suite failed (exit code ${TEST_EXIT})"
302
- fi
303
- ```
304
-
305
- Record: total tests, passed, failed, coverage (if available).
306
-
307
- **If any tests fail:** Mark as `behavioral_failures` — these are BLOCKER severity
308
- regardless of whether static checks passed. A phase cannot be verified if tests fail.
309
-
310
- **Step 2: Run project CLI/commands from success criteria (if testable)**
311
-
312
- For each success criterion that describes a user command (e.g., "User can run
313
- `mixtiq validate`", "User can run `npm start`"):
314
-
315
- 1. Check if the command exists and required inputs are available:
316
- - Look for example files in `templates/`, `fixtures/`, `test/`, `examples/`, or `testdata/`
317
- - Check if the CLI binary/script exists on PATH or in the project
318
- 2. **If no suitable inputs or fixtures exist:** Mark as `? NEEDS HUMAN` with reason
319
- "No test fixtures available — requires manual verification" and move on.
320
- Do NOT invent example inputs.
321
- 3. If inputs are available: run the command and verify it exits successfully.
322
-
323
- ```bash
324
- # Only run if both command and input exist
325
- if command -v {project_cli} &>/dev/null && [ -f "{example_input}" ]; then
326
- {project_cli} {example_input} 2>&1
327
- fi
328
- ```
329
-
330
- Record: command, exit code, output summary, pass/fail (or SKIPPED if no fixtures).
331
-
332
- **Step 3: Report**
333
-
334
- ```
335
- ## Behavioral Verification
336
-
337
- | Check | Result | Detail |
338
- |-------|--------|--------|
339
- | Test suite | {N} passed, {M} failed | {first failure if any} |
340
- | {CLI command 1} | ✓ / ✗ | {output summary} |
341
- | {CLI command 2} | ✓ / ✗ | {output summary} |
342
- ```
343
-
344
- **If all behavioral checks pass:** Continue to scan_antipatterns.
345
- **If any fail:** Add to verification gaps with BLOCKER severity.
346
- </step>
347
-
348
- <step name="scan_antipatterns">
349
- Extract files modified in this phase from SUMMARY.md, scan each:
350
-
351
- | Pattern | Search | Severity |
352
- |---------|--------|----------|
353
- | TBD/FIXME/XXX without same-line `issue #123`, `PR #123`, `#123`, or `DEF-*` reference | `grep -n -e TBD -e FIXME -e XXX` | 🛑 Blocker |
354
- | TODO/HACK | `grep -n -e TODO -e HACK` | ⚠️ Warning |
355
- | Placeholder content | `grep -n -iE "placeholder\|coming soon\|will be here"` | 🛑 Blocker |
356
- | Empty returns | `grep -n -E "return null\|return \{\}\|return \[\]\|=> \{\}"` | ⚠️ Warning |
357
- | Log-only functions | Functions containing only console.log | ⚠️ Warning |
358
-
359
- Categorize: 🛑 Blocker (prevents goal) | ⚠️ Warning (incomplete) | ℹ️ Info (notable).
360
- </step>
361
-
362
- <step name="audit_test_quality">
363
- **Verify that tests PROVE what they claim to prove.**
364
-
365
- This step catches test-level deceptions that pass all prior checks: files exist, are substantive, are wired, and tests pass — but the tests don't actually validate the requirement.
366
-
367
- **1. Identify requirement-linked test files**
368
-
369
- From PLAN and SUMMARY files, map each requirement to the test files that are supposed to prove it.
370
-
371
- **2. Disabled test scan**
372
-
373
- For ALL test files linked to requirements, search for disabled/skipped patterns:
374
-
375
- ```bash
376
- grep -rn -E "it\.skip|describe\.skip|test\.skip|xit\(|xdescribe\(|xtest\(|@pytest\.mark\.skip|@unittest\.skip|#\[ignore\]|\.pending|it\.todo|test\.todo" "$TEST_FILE"
377
- ```
378
-
379
- **Rule:** A disabled test linked to a requirement = requirement NOT tested.
380
- - 🛑 BLOCKER if the disabled test is the only test proving that requirement
381
- - ⚠️ WARNING if other active tests also cover the requirement
382
-
383
- **3. Circular test detection**
384
-
385
- Search for scripts/utilities that generate expected values by running the system under test:
386
-
387
- ```bash
388
- grep -rn -E "writeFileSync|writeFile|fs\.write|open\(.*w\)" "$TEST_DIRS"
389
- ```
390
-
391
- For each match, check if it also imports the system/service/module being tested. If a script both imports the system-under-test AND writes expected output values → CIRCULAR.
392
-
393
- **Circular test indicators:**
394
- - Script imports a service AND writes to fixture files
395
- - Expected values have comments like "computed from engine", "captured from baseline"
396
- - Script filename contains "capture", "baseline", "generate", "snapshot" in test context
397
- - Expected values were added in the same commit as the test assertions
398
-
399
- **Rule:** A test comparing system output against values generated by the same system is circular. It proves consistency, not correctness.
400
-
401
- **4. Expected value provenance** (for comparison/parity/migration requirements)
402
-
403
- When a requirement demands comparison with an external source ("identical to X", "matches Y", "same output as Z"):
404
-
405
- - Is the external source actually invoked or referenced in the test pipeline?
406
- - Do fixture files contain data sourced from the external system?
407
- - Or do all expected values come from the new system itself or from mathematical formulas?
408
-
409
- **Provenance classification:**
410
- - VALID: Expected value from external/legacy system output, manual capture, or independent oracle
411
- - PARTIAL: Expected value from mathematical derivation (proves formula, not system match)
412
- - CIRCULAR: Expected value from the system being tested
413
- - UNKNOWN: No provenance information — treat as SUSPECT
414
-
415
- **5. Assertion strength**
416
-
417
- For each test linked to a requirement, classify the strongest assertion:
418
-
419
- | Level | Examples | Proves |
420
- |-------|---------|--------|
421
- | Existence | `toBeDefined()`, `!= null` | Something returned |
422
- | Type | `typeof x === 'number'` | Correct shape |
423
- | Status | `code === 200` | No error |
424
- | Value | `toEqual(expected)`, `toBeCloseTo(x)` | Specific value |
425
- | Behavioral | Multi-step workflow assertions | End-to-end correctness |
426
-
427
- If a requirement demands value-level or behavioral-level proof and the test only has existence/type/status assertions → INSUFFICIENT.
428
-
429
- **6. Coverage quantity**
430
-
431
- If a requirement specifies a quantity of test cases (e.g., "30 calculations"), check if the actual number of active (non-skipped) test cases meets the requirement.
432
-
433
- **Reporting — add to VERIFICATION.md:**
434
-
435
- ```markdown
436
- ### Test Quality Audit
437
-
438
- | Test File | Linked Req | Active | Skipped | Circular | Assertion Level | Verdict |
439
- |-----------|-----------|--------|---------|----------|----------------|---------|
440
-
441
- **Disabled tests on requirements:** {N} → {BLOCKER if any req has ONLY disabled tests}
442
- **Circular patterns detected:** {N} → {BLOCKER if any}
443
- **Insufficient assertions:** {N} → {WARNING}
444
- ```
445
-
446
- **Impact on status:** Any BLOCKER from test quality audit ��� overall status = `gaps_found`, regardless of other checks passing.
447
- </step>
448
-
449
- <step name="identify_human_verification">
450
- **First: determine if this is an infrastructure/foundation phase.**
451
-
452
- Infrastructure and foundation phases — code foundations, database schema, internal APIs, data models, build tooling, CI/CD, internal service integrations — have no user-facing elements by definition. For these phases:
453
-
454
- - Do NOT invent artificial manual steps (e.g., "manually run git commits", "manually invoke methods", "manually check database state").
455
- - Mark human verification as **N/A** with rationale: "Infrastructure/foundation phase — no user-facing elements to test manually."
456
- - Set `human_verification: []` and do **not** produce a `human_needed` status solely due to lack of user-facing features.
457
- - Only add human verification items if the phase goal or success criteria explicitly describe something a user would interact with (UI, CLI command output visible to end users, external service UX).
458
- - **Exception — behavior-unverified truths still count.** A truth marked ⚠️ PRESENT_BEHAVIOR_UNVERIFIED (a state transition or a cancellation/cleanup/ordering invariant with no test exercising it) is a behavioral-evidence gap, not an artificial user-facing step. Record it in `behavior_unverified_items` and emit a human-verification item for it **even on an infrastructure/foundation phase** — these invariants are exactly where infra phases hide runtime state leaks. Such a truth drives `human_needed`; the auto-pass-UAT shortcut applies only to the absence of user-facing UX, never to a behavior-unverified invariant.
459
-
460
- **How to determine if a phase is infrastructure/foundation:**
461
- - Phase goal or name contains: "foundation", "infrastructure", "schema", "database", "internal API", "data model", "scaffolding", "pipeline", "tooling", "CI", "migrations", "service layer", "backend", "core library"
462
- - Phase success criteria describe only technical artifacts (files exist, tests pass, schema is valid) with no user interaction required
463
- - There is no UI, CLI output visible to end users, or real-time behavior to observe
464
-
465
- **If the phase IS infrastructure/foundation:** auto-pass UAT — skip the human verification items list entirely, **except any ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truth (see exception above), which still emits a human-verification item and drives `human_needed`.** Log:
466
-
467
- ```markdown
468
- ## Human Verification
469
-
470
- N/A — Infrastructure/foundation phase with no user-facing elements.
471
- All acceptance criteria are verifiable programmatically.
472
- ```
473
-
474
- **If the phase IS user-facing:** Only flag items that genuinely require a human. Do not invent steps.
475
-
476
- **Always needs human (user-facing phases only):** Visual appearance, user flow completion, real-time behavior (WebSocket/SSE), external service integration, performance feel, error message clarity.
477
-
478
- **Needs human if uncertain (user-facing phases only):** Complex wiring grep can't trace, dynamic state-dependent behavior, edge cases.
479
-
480
- Format each as: Test Name → What to do → Expected result → Why can't verify programmatically.
481
- </step>
482
-
483
- <step name="determine_status">
484
- Classify status using this decision tree IN ORDER (most restrictive first):
485
-
486
- 1. IF any truth FAILED, artifact MISSING/STUB, key link NOT_WIRED, blocker found, **or test quality audit found blockers (disabled requirement tests, circular tests)**:
487
- → **gaps_found**
488
-
489
- 2. IF any `must_haves.prohibitions` item disposes as flagged-unverified (ADR-550 D4):
490
- - **test-tier, fail-closed when the wired check is MISSING OR FAILS** (now run via `check prohibition-enforcement` — `located: false`, or `dispositionForProhibition()` returns `status: 'unverified'`, `flagged: true`): → **gaps_found** in both interactive and autonomous modes (never green; a missing/failing mechanical check is an unverified gap). A test-tier item whose wired check PASSES disposes `status: 'green'`, `flagged: false` and is NOT a gap — it can reach **passed**.
491
- - **judgment-tier, autonomous run** (non-authoritative LLM-judge verdict): emit the `unverified-prohibition — human review recommended` flag and classify → **human_needed** (autonomous completion reads "complete with N flagged prohibitions"; never a silent pass, never a hard halt).
492
- - **judgment-tier, interactive run**: route to the end-of-phase human checkpoint → **human_needed**.
493
-
494
- 2b. IF any `must_haves.truths` item carries the `verification: backstop` marker (#1154 — the verify-time truth-axis mirror of ADR-550 D4) AND the verifier cannot confirm it with **explicit evidence** (a wired held-out/property-based test that PASSES, or a directly-observed behavior — i.e. `dispositionForUnverifiableTruth()` returns `status: 'unverified'`, `flagged: true`, `reason: 'insufficient_spec'`):
495
- - **abstain → human_needed**, NEVER `passed` and never silently graded green. Emit a prominent `unverified — held-out test recommended` flag carrying the distinguishable `reason: insufficient_spec` (so it is not conflated with ordinary manual-UAT `human_needed`).
496
- - *Autonomous run:* record it and continue — completion reads "complete with N unverified non-inferable checks"; never a hard halt of an AFK run. *Interactive run:* route to the end-of-phase human checkpoint.
497
- - **Exogenous only:** abstention fires SOLELY on the `backstop` tag, never a self-judged "abstain if unsure" (N17). An **inferable** truth is NEVER abstained (over-abstention guard); a `backstop` truth WITH a passing wired held-out test reaches **passed**. Reliable on capable tiers (`sonnet`+); the budget `haiku` tier degrades — see `references/honest-verifier.md`.
498
-
499
- 3. IF the previous step produced ANY human verification items — this includes every ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truth and every abstained `insufficient_spec` backstop truth:
500
- → **human_needed** (even if all other truths VERIFIED)
501
-
502
- 4. IF all checks pass AND no human verification items AND no flagged prohibitions AND no abstained (`insufficient_spec`) truths:
503
- → **passed**
504
-
505
- **passed is ONLY valid when no human verification items, no flagged prohibitions, AND no abstained `insufficient_spec` truths exist.** Neither a prohibition (must-NOT) nor an unconfirmable non-inferable truth can ever be silently absorbed into a `passed` verdict — that is the core failure mode ADR-550 D4 forbids (now closed on both the prohibition and truth axes).
506
-
507
- A ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truth is never FAILED and never VERIFIED: it does not trigger gaps_found (the code is present and wired) and is not counted as verified (its runtime behavior was not exercised). It routes through the existing human_needed sink — no new overall status.
508
-
509
- **Score:** `verified_truths / total_truths` — `verified_truths` counts ✓ VERIFIED truths plus PASSED (override) truths; excluded are ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truths (the `behavior_unverified` count) and abstained ⚠️ `insufficient_spec` backstop truths (#1154) — both are not ✓ VERIFIED and both route to `human_needed`. A headline N/N therefore certifies behavioral evidence for every behavior-dependent truth and explicit evidence for every non-inferable one, not merely symbol presence.
510
- </step>
511
-
512
- <step name="filter_deferred_items">
513
- Before reporting gaps, cross-reference each gap against later phases in the milestone using the full roadmap data loaded in load_context (from `roadmap analyze`).
514
-
515
- For each potential gap identified in determine_status:
516
- 1. Check if the gap's failed truth or missing item is covered by a later phase's goal or success criteria
517
- 2. **Match criteria:** The gap's concern appears in a later phase's goal text, success criteria text, or the later phase's name clearly suggests it covers this area
518
- 3. If a clear match is found → move the gap to a `deferred` list with the matching phase reference and evidence text
519
- 4. If no match in any later phase → keep as a real `gap`
520
-
521
- **Important:** Be conservative. Only defer a gap when there is clear, specific evidence in a later phase. Vague or tangential matches should NOT cause deferral — when in doubt, keep it as a real gap.
522
-
523
- **Deferred items do NOT affect the status determination.** Recalculate after filtering:
524
- - If gaps list is now empty and no human items exist → `passed`
525
- - If gaps list is now empty but human items exist → `human_needed`
526
- - If gaps list still has items → `gaps_found`
527
-
528
- Include deferred items in VERIFICATION.md frontmatter (`deferred:` section) and body (Deferred Items table) for transparency. If no deferred items exist, omit these sections.
529
- </step>
530
-
531
- <step name="generate_fix_plans">
532
- If gaps_found:
533
-
534
- 1. **Cluster related gaps:** API stub + component unwired → "Wire frontend to backend". Multiple missing → "Complete core implementation". Wiring only → "Connect existing components".
535
-
536
- 2. **Generate plan per cluster:** Objective, 2-3 tasks (files/action/verify each), re-verify step. Keep focused: single concern per plan.
537
-
538
- 3. **Order by dependency:** Fix missing → fix stubs → fix wiring → **fix test evidence** → verify.
539
- </step>
540
-
541
- <step name="create_report">
542
- ```bash
543
- REPORT_PATH="$PHASE_DIR/${PHASE_NUM}-VERIFICATION.md"
544
- ```
545
-
546
- Fill template sections: frontmatter (phase/timestamp/status/score), goal achievement, artifact table, wiring table, requirements coverage, anti-patterns, human verification, gaps summary, fix plans (if gaps_found), metadata.
547
-
548
- See ~/.claude/gsd-core/templates/verification-report.md for complete template.
549
- </step>
550
-
551
- <step name="return_to_orchestrator">
552
- Return status (`passed` | `gaps_found` | `human_needed`), score (N/M must-haves), report path.
553
-
554
- If gaps_found: list gaps + recommended fix plan names.
555
- If human_needed: list items requiring human testing.
556
-
557
- Orchestrator routes: `passed` → update_roadmap | `gaps_found` → create/execute fixes, re-verify | `human_needed` → present to user.
558
- </step>
559
-
560
- </process>
561
-
562
- <success_criteria>
563
- - [ ] Must-haves established (from frontmatter or derived)
564
- - [ ] All truths verified with status and evidence
565
- - [ ] All artifacts checked at all three levels
566
- - [ ] All key links verified
567
- - [ ] Requirements coverage assessed (if applicable)
568
- - [ ] CONTEXT.md decisions checked against shipped artifacts (#2492 — non-blocking)
569
- - [ ] Anti-patterns scanned and categorized
570
- - [ ] Test quality audited (disabled tests, circular patterns, assertion strength, provenance)
571
- - [ ] Human verification items identified
572
- - [ ] Overall status determined
573
- - [ ] Deferred items filtered against later milestone phases (if gaps found)
574
- - [ ] Fix plans generated (if gaps_found after filtering)
575
- - [ ] VERIFICATION.md created with complete report
576
- - [ ] Results returned to orchestrator
577
- </success_criteria>