@opengsd/gsd-core 1.7.0 → 1.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (261) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/.claude-plugin/plugin.json +1 -1
  3. package/.opencode/plugins/gsd-core.js +45 -1
  4. package/README.md +2 -0
  5. package/agents/gsd-code-fixer.md +1 -1
  6. package/agents/gsd-codebase-mapper.md +1 -1
  7. package/agents/gsd-debug-session-manager.md +78 -4
  8. package/agents/gsd-debugger.md +87 -29
  9. package/agents/gsd-executor.md +49 -9
  10. package/agents/gsd-intel-updater.md +3 -3
  11. package/agents/gsd-phase-researcher.md +4 -2
  12. package/agents/gsd-plan-checker.md +20 -0
  13. package/agents/gsd-planner.md +44 -59
  14. package/agents/gsd-project-researcher.md +2 -2
  15. package/agents/gsd-ui-auditor.md +0 -40
  16. package/agents/gsd-verifier.md +2 -2
  17. package/bin/install.js +1338 -135
  18. package/commands/gsd/ai-integration-phase.md +1 -1
  19. package/commands/gsd/mempalace-capture.md +9 -5
  20. package/commands/gsd/new-milestone.md +1 -1
  21. package/commands/gsd/plan-phase.md +5 -3
  22. package/commands/gsd/plan-review-convergence.md +7 -2
  23. package/gsd-core/bin/gsd-tools.cjs +2690 -2472
  24. package/gsd-core/bin/lib/adapter-imperative.cjs +8 -1
  25. package/gsd-core/bin/lib/agent-command-router.cjs +20 -5
  26. package/gsd-core/bin/lib/api-coverage.cjs +360 -53
  27. package/gsd-core/bin/lib/audit.cjs +8 -8
  28. package/gsd-core/bin/lib/broken-windows.cjs +716 -0
  29. package/gsd-core/bin/lib/capability-command-router.cjs +733 -0
  30. package/gsd-core/bin/lib/capability-consent.cjs +40 -1
  31. package/gsd-core/bin/lib/capability-lifecycle.cjs +58 -0
  32. package/gsd-core/bin/lib/capability-loader.cjs +23 -1
  33. package/gsd-core/bin/lib/capability-registry.cjs +1450 -160
  34. package/gsd-core/bin/lib/capability-trust.cjs +468 -33
  35. package/gsd-core/bin/lib/capability-validator.cjs +882 -6
  36. package/gsd-core/bin/lib/capability-writer.cjs +6 -1
  37. package/gsd-core/bin/lib/check-command-router.cjs +140 -27
  38. package/gsd-core/bin/lib/cjs-command-router-adapter.cjs +15 -0
  39. package/gsd-core/bin/lib/claude-orchestration-command-router.cjs +209 -31
  40. package/gsd-core/bin/lib/claude-orchestration.cjs +203 -25
  41. package/gsd-core/bin/lib/command-aliases.cjs +14 -0
  42. package/gsd-core/bin/lib/commands.cjs +326 -21
  43. package/gsd-core/bin/lib/config-loader.cjs +214 -30
  44. package/gsd-core/bin/lib/config.cjs +158 -22
  45. package/gsd-core/bin/lib/core-utils.cjs +6 -1
  46. package/gsd-core/bin/lib/decisions.cjs +32 -8
  47. package/gsd-core/bin/lib/docs.cjs +6 -0
  48. package/gsd-core/bin/lib/estimate-cli.cjs +336 -0
  49. package/gsd-core/bin/lib/external-descriptor-trust.cjs +14 -2
  50. package/gsd-core/bin/lib/frontmatter.cjs +125 -15
  51. package/gsd-core/bin/lib/gap-checker.cjs +17 -2
  52. package/gsd-core/bin/lib/host-integration.cjs +215 -8
  53. package/gsd-core/bin/lib/init.cjs +155 -66
  54. package/gsd-core/bin/lib/install-engine.cjs +299 -23
  55. package/gsd-core/bin/lib/install-profiles.cjs +239 -1
  56. package/gsd-core/bin/lib/installer-migrations/005-opencode-baseline-commands-dir.cjs +146 -0
  57. package/gsd-core/bin/lib/installer-migrations/006-pi-extension-cjs-to-js.cjs +91 -0
  58. package/gsd-core/bin/lib/installer-migrations.cjs +44 -5
  59. package/gsd-core/bin/lib/markdown-sectionizer.cjs +107 -0
  60. package/gsd-core/bin/lib/milestone.cjs +248 -14
  61. package/gsd-core/bin/lib/model-catalog.cjs +69 -4
  62. package/gsd-core/bin/lib/model-resolver.cjs +189 -7
  63. package/gsd-core/bin/lib/observability/logger.cjs +7 -2
  64. package/gsd-core/bin/lib/onboard-projection.cjs +11 -8
  65. package/gsd-core/bin/lib/phase-command-router.cjs +10 -1
  66. package/gsd-core/bin/lib/phase-estimation.cjs +398 -0
  67. package/gsd-core/bin/lib/phase-id.cjs +304 -9
  68. package/gsd-core/bin/lib/phase.cjs +258 -17
  69. package/gsd-core/bin/lib/plan-drift-guard.cjs +1 -1
  70. package/gsd-core/bin/lib/plan-scan.cjs +70 -2
  71. package/gsd-core/bin/lib/planning-workspace.cjs +9 -2
  72. package/gsd-core/bin/lib/profile-output.cjs +34 -8
  73. package/gsd-core/bin/lib/review-lane-descriptor.cjs +927 -0
  74. package/gsd-core/bin/lib/review-lane-invocation.cjs +348 -0
  75. package/gsd-core/bin/lib/review-lane-runner.cjs +594 -0
  76. package/gsd-core/bin/lib/review-reviewer-selection.cjs +114 -32
  77. package/gsd-core/bin/lib/roadmap-parser.cjs +61 -10
  78. package/gsd-core/bin/lib/roadmap.cjs +23 -7
  79. package/gsd-core/bin/lib/runtime-artifact-conversion.cjs +38 -5
  80. package/gsd-core/bin/lib/runtime-artifact-layout.cjs +23 -9
  81. package/gsd-core/bin/lib/runtime-hooks-surface.cjs +156 -0
  82. package/gsd-core/bin/lib/runtime-name-policy.cjs +15 -2
  83. package/gsd-core/bin/lib/smart-entry.cjs +70 -5
  84. package/gsd-core/bin/lib/state-document.cjs +171 -24
  85. package/gsd-core/bin/lib/state-transition.cjs +50 -11
  86. package/gsd-core/bin/lib/state.cjs +206 -32
  87. package/gsd-core/bin/lib/surface.cjs +51 -9
  88. package/gsd-core/bin/lib/uat-predicate.cjs +6 -4
  89. package/gsd-core/bin/lib/uat.cjs +428 -11
  90. package/gsd-core/bin/lib/ui-consideration-probe.cjs +2 -2
  91. package/gsd-core/bin/lib/unusable-input.cjs +216 -0
  92. package/gsd-core/bin/lib/validate.cjs +44 -8
  93. package/gsd-core/bin/lib/verification.cjs +163 -31
  94. package/gsd-core/bin/lib/verify.cjs +348 -42
  95. package/gsd-core/bin/lib/worktree-safety.cjs +360 -15
  96. package/gsd-core/bin/shared/config-defaults.manifest.json +1 -0
  97. package/gsd-core/bin/shared/config-schema.manifest.json +4 -15
  98. package/gsd-core/bin/shared/model-catalog.json +5 -0
  99. package/gsd-core/bin/shared/runtime-aliases.manifest.json +5 -0
  100. package/gsd-core/references/api-coverage.md +37 -7
  101. package/gsd-core/references/checkpoints.md +1 -1
  102. package/gsd-core/references/common-bug-patterns.md +13 -0
  103. package/gsd-core/references/context-budget.md +40 -0
  104. package/gsd-core/references/debugger-bug-taxonomy.md +111 -0
  105. package/gsd-core/references/debugger-fix-acceptance.md +157 -0
  106. package/gsd-core/references/debugger-philosophy.md +1 -0
  107. package/gsd-core/references/debugger-prevention.md +98 -0
  108. package/gsd-core/references/debugger-rca-branching.md +98 -0
  109. package/gsd-core/references/debugger-repro-hardening.md +130 -0
  110. package/gsd-core/references/debugger-sbfl.md +110 -0
  111. package/gsd-core/references/debugger-semantic-recall.md +81 -0
  112. package/gsd-core/references/execute-phase-quota-recovery.md +55 -0
  113. package/gsd-core/references/execute-phase-requirement-revert.md +8 -0
  114. package/gsd-core/references/execute-phase-response-language.md +7 -0
  115. package/gsd-core/references/gate-prompts.md +6 -3
  116. package/gsd-core/references/model-profile-resolution.md +64 -13
  117. package/gsd-core/references/offer-next.md +88 -0
  118. package/gsd-core/references/planner-antipatterns.md +6 -0
  119. package/gsd-core/references/planner-mvp-mode.md +12 -13
  120. package/gsd-core/references/planner-preconditions.md +156 -0
  121. package/gsd-core/references/planner-reversibility.md +132 -0
  122. package/gsd-core/references/planning-config.md +2 -1
  123. package/gsd-core/references/reviewer-instances.md +28 -19
  124. package/gsd-core/references/runtime-aware-dispatch.md +42 -0
  125. package/gsd-core/references/skeleton-template.md +1 -1
  126. package/gsd-core/references/thinking-models-planning.md +3 -1
  127. package/gsd-core/references/ui-consideration-probe.md +2 -2
  128. package/gsd-core/references/worktree-branch-check.md +4 -4
  129. package/gsd-core/templates/DEBUG.md +5 -3
  130. package/gsd-core/templates/summary-minimal.md +4 -0
  131. package/gsd-core/templates/summary-standard.md +4 -0
  132. package/gsd-core/templates/summary.md +7 -0
  133. package/gsd-core/workflows/add-phase.md +2 -0
  134. package/gsd-core/workflows/add-tests.md +3 -1
  135. package/gsd-core/workflows/add-todo.md +32 -1
  136. package/gsd-core/workflows/ai-integration-phase.md +8 -6
  137. package/gsd-core/workflows/audit-fix.md +6 -2
  138. package/gsd-core/workflows/audit-milestone.md +8 -0
  139. package/gsd-core/workflows/autonomous.md +19 -15
  140. package/gsd-core/workflows/check-todos.md +5 -3
  141. package/gsd-core/workflows/cleanup.md +7 -1
  142. package/gsd-core/workflows/code-review-fix.md +14 -6
  143. package/gsd-core/workflows/code-review.md +93 -24
  144. package/gsd-core/workflows/complete-milestone.md +3 -0
  145. package/gsd-core/workflows/debug.md +35 -7
  146. package/gsd-core/workflows/diagnose-issues.md +5 -1
  147. package/gsd-core/workflows/discovery-phase.md +7 -0
  148. package/gsd-core/workflows/discuss-phase/modes/advisor.md +2 -4
  149. package/gsd-core/workflows/discuss-phase/modes/auto.md +0 -6
  150. package/gsd-core/workflows/discuss-phase/templates/context.md +16 -2
  151. package/gsd-core/workflows/discuss-phase-assumptions.md +18 -9
  152. package/gsd-core/workflows/discuss-phase.md +2 -2
  153. package/gsd-core/workflows/do.md +7 -1
  154. package/gsd-core/workflows/docs-update.md +9 -0
  155. package/gsd-core/workflows/eval-review.md +4 -1
  156. package/gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md +4 -0
  157. package/gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md +160 -0
  158. package/gsd-core/workflows/execute-phase/steps/post-merge-gate.md +4 -4
  159. package/gsd-core/workflows/execute-phase/steps/regression-gate.md +2 -2
  160. package/gsd-core/workflows/execute-phase.md +110 -149
  161. package/gsd-core/workflows/execute-plan.md +20 -8
  162. package/gsd-core/workflows/explore.md +4 -0
  163. package/gsd-core/workflows/extract-learnings.md +21 -0
  164. package/gsd-core/workflows/graduation.md +3 -0
  165. package/gsd-core/workflows/health.md +7 -1
  166. package/gsd-core/workflows/help/modes/full.md +9 -5
  167. package/gsd-core/workflows/import.md +11 -2
  168. package/gsd-core/workflows/inbox.md +7 -0
  169. package/gsd-core/workflows/ingest-docs.md +19 -10
  170. package/gsd-core/workflows/manager.md +3 -1
  171. package/gsd-core/workflows/map-codebase.md +17 -10
  172. package/gsd-core/workflows/mvp-phase.md +3 -0
  173. package/gsd-core/workflows/new-milestone.md +79 -23
  174. package/gsd-core/workflows/new-project.md +28 -19
  175. package/gsd-core/workflows/new-workspace.md +3 -1
  176. package/gsd-core/workflows/next.md +5 -2
  177. package/gsd-core/workflows/onboard.md +3 -0
  178. package/gsd-core/workflows/plan-phase.md +56 -51
  179. package/gsd-core/workflows/plan-review-convergence.md +61 -12
  180. package/gsd-core/workflows/plant-seed.md +3 -0
  181. package/gsd-core/workflows/profile-user.md +7 -1
  182. package/gsd-core/workflows/progress.md +31 -3
  183. package/gsd-core/workflows/quick.md +33 -10
  184. package/gsd-core/workflows/remove-workspace.md +3 -0
  185. package/gsd-core/workflows/review.md +172 -585
  186. package/gsd-core/workflows/scan.md +10 -2
  187. package/gsd-core/workflows/secure-phase.md +13 -2
  188. package/gsd-core/workflows/settings-integrations.md +3 -0
  189. package/gsd-core/workflows/settings.md +3 -0
  190. package/gsd-core/workflows/ship.md +88 -11
  191. package/gsd-core/workflows/sketch.md +3 -0
  192. package/gsd-core/workflows/smart-entry.md +4 -1
  193. package/gsd-core/workflows/spike.md +7 -1
  194. package/gsd-core/workflows/ui-phase.md +11 -2
  195. package/gsd-core/workflows/ui-review.md +11 -1
  196. package/gsd-core/workflows/undo.md +7 -0
  197. package/gsd-core/workflows/update.md +106 -5
  198. package/gsd-core/workflows/validate-phase.md +13 -2
  199. package/gsd-core/workflows/verify-phase.md +2 -2
  200. package/gsd-core/workflows/verify-work.md +15 -4
  201. package/hooks/dist/gsd-context-monitor.js +27 -9
  202. package/hooks/dist/gsd-cursor-session-start.js +6 -2
  203. package/hooks/dist/gsd-cursor-stop.js +6 -2
  204. package/hooks/dist/gsd-cursor-subagent-start.js +6 -2
  205. package/hooks/dist/gsd-graphify-update.sh +9 -0
  206. package/hooks/dist/gsd-phase-boundary.sh +14 -2
  207. package/hooks/dist/gsd-prompt-guard.js +101 -2
  208. package/hooks/dist/gsd-read-guard.js +100 -2
  209. package/hooks/dist/gsd-read-injection-scanner.js +109 -2
  210. package/hooks/dist/gsd-statusline.js +97 -9
  211. package/hooks/dist/gsd-workflow-guard.js +110 -6
  212. package/hooks/dist/gsd-worktree-path-guard.js +132 -8
  213. package/hooks/dist/lib/cursor-workspace.js +74 -0
  214. package/hooks/gsd-context-monitor.js +27 -9
  215. package/hooks/gsd-cursor-session-start.js +6 -2
  216. package/hooks/gsd-cursor-stop.js +6 -2
  217. package/hooks/gsd-cursor-subagent-start.js +6 -2
  218. package/hooks/gsd-graphify-update.sh +9 -0
  219. package/hooks/gsd-phase-boundary.sh +14 -2
  220. package/hooks/gsd-prompt-guard.js +101 -2
  221. package/hooks/gsd-read-guard.js +100 -2
  222. package/hooks/gsd-read-injection-scanner.js +109 -2
  223. package/hooks/gsd-statusline.js +97 -9
  224. package/hooks/gsd-workflow-guard.js +110 -6
  225. package/hooks/gsd-worktree-path-guard.js +132 -8
  226. package/hooks/lib/cursor-workspace.js +74 -0
  227. package/package.json +10 -8
  228. package/pi/gsd.cjs +34 -3
  229. package/scripts/changeset/lint.cjs +1 -0
  230. package/scripts/changeset/parse.cjs +26 -0
  231. package/scripts/check-coverage-gate.cjs +51 -0
  232. package/scripts/check-glossary-refs.cjs +244 -0
  233. package/scripts/ci-rebase-check.cjs +48 -4
  234. package/scripts/ci-test-scope.cjs +67 -17
  235. package/scripts/gen-adr-index.cjs +528 -0
  236. package/scripts/gen-capability-matrix.cjs +26 -2
  237. package/scripts/gen-capability-registry.cjs +132 -34
  238. package/scripts/gen-emitted-baseline.cjs +145 -0
  239. package/scripts/gen-test-timings.cjs +201 -0
  240. package/scripts/lint-compiled-artifact-sync.cjs +146 -0
  241. package/scripts/lint-emitted-drift-ack.cjs +149 -0
  242. package/scripts/lint-fix-has-regression-test.cjs +131 -0
  243. package/scripts/lint-portable-timeout.cjs +140 -0
  244. package/scripts/lint-resolution-provenance.cjs +9 -0
  245. package/scripts/lint-test-file-count.allowlist.json +1 -0
  246. package/scripts/mutation-matrix.cjs +4 -0
  247. package/scripts/prompt-injection-scan.sh +6 -0
  248. package/scripts/registry-schema.cjs +57 -8
  249. package/scripts/release-notes/conventional-title.cjs +19 -1
  250. package/scripts/release-notes/format-github-release-notes.cjs +7 -3
  251. package/scripts/release-tarball-smoke.cjs +18 -11
  252. package/scripts/run-tests.cjs +420 -58
  253. package/scripts/workflow-size.cjs +16 -8
  254. package/skills/gsd-ai-integration-phase/SKILL.md +1 -1
  255. package/skills/gsd-mempalace-capture/SKILL.md +9 -5
  256. package/skills/gsd-new-milestone/SKILL.md +1 -1
  257. package/skills/gsd-plan-phase/SKILL.md +5 -3
  258. package/skills/gsd-plan-review-convergence/SKILL.md +7 -2
  259. package/vscode/package.json +1 -1
  260. package/scripts/gen-golden-install-parity-zcode.cjs +0 -77
  261. package/scripts/update-size-baseline.cjs +0 -68
@@ -123,7 +123,7 @@ All JSON files include a `_meta` object with `updated_at` (ISO timestamp) and `v
123
123
  }
124
124
  ```
125
125
 
126
- **exports constraint:** Array of ACTUAL exported symbol names extracted from `module.exports` or `export` statements. MUST be real identifiers (e.g., `"configLoad"`, `"stateUpdate"`), NOT descriptions (e.g., `"config operations"`). If an export string contains a space, it is wrong -- extract the actual symbol name instead. Use `gsd-tools intel extract-exports <file>` to get accurate exports.
126
+ **exports constraint:** Array of ACTUAL exported symbol names extracted from `module.exports` or `export` statements. MUST be real identifiers (e.g., `"configLoad"`, `"stateUpdate"`), NOT descriptions (e.g., `"config operations"`). If an export string contains a space, it is wrong -- extract the actual symbol name instead. Use `gsd_run intel extract-exports <file>` to get accurate exports.
127
127
 
128
128
  Types: `entry-point`, `module`, `config`, `test`, `script`, `type-def`, `style`, `template`, `data`.
129
129
 
@@ -255,7 +255,7 @@ gsd_run intel patch-meta .planning/intel/arch-decisions.json
255
255
 
256
256
  ### Step 6.5: Self-Check
257
257
 
258
- Run: `gsd-tools intel validate`
258
+ Run: `gsd_run intel validate`
259
259
 
260
260
  Review the output:
261
261
 
@@ -267,7 +267,7 @@ This step is MANDATORY -- do not skip it.
267
267
 
268
268
  ### Step 7: Snapshot
269
269
 
270
- Run: `gsd-tools intel snapshot`
270
+ Run: `gsd_run intel snapshot`
271
271
 
272
272
  This writes `.last-refresh.json` with accurate timestamps and hashes. Do NOT write `.last-refresh.json` manually.
273
273
  </execution_flow>
@@ -32,6 +32,8 @@ Spawned by `/gsd:plan-phase` (integrated) or `/gsd:plan-phase --research-phase <
32
32
 
33
33
  **Package name provenance rule:** A package name discovered via WebSearch, training data, or any non-authoritative source must be tagged `[ASSUMED]` regardless of whether `npm view` confirms it exists on the registry. Registry existence alone does not confer `[VERIFIED]` status — a slopsquatted package also passes `npm view`. Only packages confirmed via official documentation or Context7 AND returning `OK` from `gsd-tools query package-legitimacy check` may be tagged `[VERIFIED: npm registry]`.
34
34
 
35
+ **In-repo value provenance rule:** A claim about an in-repo *discrete value* — an enum, a schema or type union, an error code, a status constant, or a filesystem path — may be tagged `[VERIFIED: …]` only if you opened the source-of-truth file with `Read` **this session**. A codebase `grep` is not sufficient on its own: it confirms a string occurs, not that you read the definition. Cite the path **and line range** (`[VERIFIED: src/types/order.ts:14-22]`), and quote the values **verbatim** in RESEARCH.md beside the claim — paraphrase is forbidden. The quote is what makes the tag checkable — a citation with no quote beside it does not earn `[VERIFIED]`, however precise the line range looks. Every value appearing in a code example or skeleton must also appear in that verbatim quote; a value that does not is `[ASSUMED]`. For a filesystem path, cite the line in the script that creates it, not the location you expect it to occupy. Training memory and a web search are not substitutes for reading the file — a discrete value that merely looks right fails at the executor's `parse()`/typecheck, the most expensive place to discover it.
36
+
35
37
  Claims tagged `[ASSUMED]` signal to the planner and discuss-phase that the information needs user confirmation before becoming a locked decision. Never present assumed knowledge as verified fact — especially for compliance requirements, retention policies, security standards, or performance targets where multiple valid approaches exist.
36
38
  </role>
37
39
 
@@ -136,7 +138,7 @@ For each item where `fetch` is present, invoke the MCP tool matching `fetch.prov
136
138
  | `exa` | `mcp__exa__web_search_exa` with `fetch.query` |
137
139
  | `tavily` | `mcp__tavily__search` with `fetch.query` |
138
140
  | `perplexity` | `mcp__perplexity__*` (use the appropriate perplexity MCP tool for the query) |
139
- | `brave` | `gsd-tools query websearch "<fetch.query>"` (Brave-backed) or built-in `WebSearch` |
141
+ | `brave` | `gsd_run query websearch "<fetch.query>"` (Brave-backed) or built-in `WebSearch` |
140
142
  | `firecrawl` | `mcp__firecrawl__scrape` with url (scrape kind) or `mcp__firecrawl__search` |
141
143
  | `websearch` | built-in `WebSearch` tool |
142
144
  | `webfetch` | built-in `WebFetch` tool |
@@ -707,7 +709,7 @@ docker info 2>/dev/null | head -3
707
709
 
708
710
  ## Step 3: Execute Research Protocol
709
711
 
710
- For each domain, use the `<tool_strategy>` seam (Steps A–D): build questions JSON, call `gsd-tools query research-plan`, run the indicated provider per item, then cache each digest. Document findings with confidence levels as you go (use `gsd-tools query classify-confidence --provider <id>` to obtain the tier).
712
+ For each domain, use the `<tool_strategy>` seam (Steps A–D): build questions JSON, call `gsd_run query research-plan`, run the indicated provider per item, then cache each digest. Document findings with confidence levels as you go (use `gsd_run query classify-confidence --provider <id>` to obtain the tier).
711
713
 
712
714
  ## Step 4: Validation Architecture Research (if nyquist_validation enabled)
713
715
 
@@ -252,6 +252,19 @@ issue:
252
252
  1. Count tasks per plan
253
253
  2. Estimate files modified per plan
254
254
  3. Check against thresholds
255
+ 4. **Smart-zone estimate check (#2631, ADR-2629).** For each plan carrying an `estimate` block, run the
256
+ `estimate-check --calibrated` verb against its `estimate.tokens` (the `--calibrated` flag is required —
257
+ the plan's figure already has the factor applied, and omitting it would square the correction) (invoked in Step 1 below, after the launcher
258
+ preamble). The verb reads `workflow.smart_zone_tokens` and applies the project's calibration. Report
259
+ one line per plan: plan id, estimated tokens, the budget, and — when `over_budget` is true — the
260
+ returned `recommendation`, which names how many slices the phase should become.
261
+
262
+ **Over budget is a WARNING, never a blocker** (ADR-2629 Decision 5). Recommend re-slicing into a tracer
263
+ plus expansion slices; never fail the check on it. Report `estimate.confidence` alongside: `low` means
264
+ fewer than 3 completed phases carry actuals, so the figure is not yet calibrated for this project — say
265
+ so rather than presenting it as precise, and weigh the task/file thresholds above more heavily.
266
+
267
+ A plan with no `estimate` block is not a defect; the field is optional and additive.
255
268
 
256
269
  **Thresholds:**
257
270
  | Metric | Target | Warning | Blocker |
@@ -706,6 +719,13 @@ gsd_run query phase.list-plans "$phase_number"
706
719
  gsd_run query phase.list-artifacts "$phase_number" --type research
707
720
  gsd_run query roadmap.get-phase "$phase_number"
708
721
  gsd_run query phase.list-artifacts "$phase_number" --type summary
722
+
723
+ # Smart-zone estimate check (#2631) — advisory, never fails the check.
724
+ for plan in "${phase_dir:-$PHASE_DIR}"/*-PLAN.md; do
725
+ [ -f "$plan" ] || continue # unmatched glob leaves the literal pattern — skip it
726
+ EST=$(sed -n '/^estimate:/,/^[a-z_]*:/p' "$plan" | grep -o 'tokens: *[0-9]*' | head -1 | grep -o '[0-9]*')
727
+ [ -n "$EST" ] && gsd_run query estimate-check --tokens "$EST" --calibrated 2>/dev/null || true
728
+ done
709
729
  ```
710
730
 
711
731
  **Extract:** Phase goal, requirements (decompose goal), locked decisions, deferred ideas.
@@ -66,7 +66,7 @@ The orchestrator provides user decisions in `<user_decisions>` tags from `/gsd:d
66
66
  **Self-check before returning:** For each plan, verify:
67
67
  - [ ] Every locked decision (D-01, D-02, etc.) has a task implementing it
68
68
  - [ ] Task actions reference the decision ID they implement (e.g., "per D-03")
69
- (The decision-coverage gate `check.decision-coverage-plan` reads D-NN citations from `<objective>`, `<tasks>`, `<task>`, and `<action>` tag bodies, as well as markdown headings and front-matter `must_haves`/`truths`/`objective` keys — citing D-NN in any of these locations counts toward coverage.)
69
+ (The decision-coverage gate `check.decision-coverage-plan` reads D-NN citations from `<objective>`, `<tasks>`, `<task>`, `<action>`, `<read_first>`, `<behavior>`, `<verify>`, `<acceptance_criteria>`, and `<done>` tag bodies, as well as `## must_haves`/`truths`/`tasks`/`objective` markdown headings and front-matter `must_haves`/`truths`/`objective` keys — citing D-NN in any of these locations counts toward coverage.)
70
70
  - [ ] No task implements a deferred idea
71
71
  - [ ] Discretion areas are handled reasonably
72
72
 
@@ -193,23 +193,21 @@ Every task has four required fields:
193
193
  **Grep gate hygiene:** `grep -c` counts comments, so header prose can be self-invalidating. Use `grep -v '^#' | grep -c token`. Bare `== 0` gates on unfiltered files are forbidden.
194
194
 
195
195
  <comment_text_discipline>
196
- **Comment-text discipline (HARD GATE, #429):** A literal an acceptance criterion negative-greps for (`grep -c 'LIT' file == 0`) must NOT appear verbatim in any `<action>` body — JSDoc samples, head-comment references, or "what NOT to do" snippets echo into the written file and trip the executor's commit-time gate. `validate_plan` (`verify.plan-structure`) fails plan creation on violation. Rephrase the literal by concept, or — when it must legitimately appear — add an allowlist marker on its own line:
197
-
198
- `<!-- planner-discipline-allow: LIT -->`
199
-
200
- Full rules + worked examples: @gsd-core/references/planner-antipatterns.md ("Comment-Text Discipline").
196
+ **Comment-text discipline (HARD GATE, #429):** A literal an acceptance criterion negative-greps for must NOT appear verbatim in any `<action>` body. Full rules + `<!-- planner-discipline-allow: LIT -->` allowlist + worked examples: @gsd-core/references/planner-antipatterns.md ("Comment-Text Discipline").
201
197
  </comment_text_discipline>
202
198
 
203
199
  <region_scoped_negative_gate>
204
- **Region-scoped negative gates (WARN, #968):** Region-scope a file-wide negative grep when a sibling task needs that construct elsewhere in the same file; `validate_plan` WARNS. See: @gsd-core/references/planner-antipatterns.md ("Region-Scoped Negative Gates").
205
-
206
- **Verify-gate hygiene (#1478/#1479):** See @gsd-core/references/planner-antipatterns.md.
200
+ **Region-scoped negative gates (WARN, #968)** and **Verify-gate hygiene (#1478/#1479):** @gsd-core/references/planner-antipatterns.md.
207
201
  </region_scoped_negative_gate>
208
202
 
209
203
  **<done>:** Acceptance criteria - measurable state of completion.
210
204
  - Good: "Valid credentials return 200 + JWT cookie, invalid credentials return 401"
211
205
  - Bad: "Authentication is complete"
212
206
 
207
+ **<precondition>** (optional, one prose line): a runnable/checkable fact the task assumes that plan ordering does not guarantee — external setup (`user_setup`), a prior-phase artifact, or an env var. The executor asserts it before running the task and halts on unmet. Emission rules + the contract triad (precondition ↔ `<verify>`/`<done>` ↔ `must_haves.truths`): @~/.claude/gsd-core/references/planner-preconditions.md.
208
+
209
+ **<reversibility>** (optional): `rating="reversible|costly|one-way"` + one-line rationale for a decision this task implements. `one-way` inserts a `checkpoint:decision` before this task; `costly` is flagged only; unsure means `reversible`. Rules: @~/.claude/gsd-core/references/planner-reversibility.md
210
+
213
211
  See @~/.claude/gsd-core/references/planner-guidance.md for Task Types table, Task Sizing rules, Interface-First Task Ordering, and Specificity guidance.
214
212
 
215
213
  ## TDD Detection
@@ -250,34 +248,33 @@ Exceptions where `tdd="true"` is not needed: `type="checkpoint:*"` tasks, config
250
248
 
251
249
  `workflow.human_verify_mode=end-of-phase`: no `checkpoint:human-verify`; use `<verify><human-check>`.
252
250
 
253
- ## MVP Mode Detection
251
+ ## Tracer-First Decomposition (default)
254
252
 
255
- **When `MVP_MODE` is enabled (passed by the plan-phase orchestrator):** Decompose tasks as **vertical feature slices**, not horizontal layers. Required reading: Read `~/.claude/gsd-core/references/planner-mvp-mode.md` for the vertical-slice rules (lazy — only on MVP runs).
253
+ **Every phase plan LEADS with one `type="tracer"` task** — the thinnest path that touches every layer the phase will modify, wired end-to-end, carrying a real runnable `<verify>`. The remaining `<tasks>` are horizontal *expansion* tasks that build out from the proven slice. This is the default for **every** phase; it is not gated behind a flag. Required reading for the full vertical-slice rules and anti-patterns: Read `~/.claude/gsd-core/references/planner-mvp-mode.md`.
256
254
 
257
- **Core rule:** After each task completes, a real user can do something they could not do after the previous task. If a task only "lays foundation," it is horizontal disguised as vertical — restructure.
255
+ **Why tracer-first:** proving the architecture end-to-end on the agent's best early-context tokens catches an architectural dead-end after one commit instead of after ten already-committed layers.
258
256
 
259
- **Plan structure under MVP_MODE:**
257
+ **A tracer is production-quality, not a prototype.** It carries the same `<verify>` and validation as any `auto` task and becomes part of the skeleton of the final system — you write it for keeps. Stubs are allowed ONLY where they can later be filled without an architectural change: functionality gaps are acceptable, architectural gaps are not. (Glossary: `tracer bullet` vs `prototype` in `CONTEXT.md` — GSD ships tracers, never prototypes.)
260
258
 
261
- 1. Frame the phase goal as a user story at the top of `PLAN.md`. The user story is sourced from the `**Goal:**` line in ROADMAP.md (set by `mvp-phase`). Emit it with bolded keywords:
259
+ **Tracer task shape:**
262
260
 
263
- ```
264
- ## Phase Goal
265
-
266
- **As a** [user role], **I want to** [capability], **so that** [outcome].
267
- ```
261
+ ```xml
262
+ <task type="tracer">
263
+ <name>End-to-end "[capability]" — one path only</name>
264
+ <files>[one file per layer the phase touches]</files>
265
+ <action>Wire ONE entry point through every layer to the far end of the stack. No other call sites, no batching. Real error handling on the single path.</action>
266
+ <verify>[a real, runnable END-TO-END check of the one path — not a per-layer unit test]</verify>
267
+ <done>The single happy path works end-to-end and is committed.</done>
268
+ </task>
269
+ ```
268
270
 
269
- Format rules (Read `~/.claude/gsd-core/references/user-story-template.md`):
270
- - All three slots required. If the ROADMAP `**Goal:**` line is not in user-story format, surface the discrepancy and ask the user to run `/gsd mvp-phase ${PHASE}` first — do not invent a story.
271
- - Bold the three keywords (`**As a**`, `**I want to**`, `**so that**`) when emitting to PLAN.md. The ROADMAP form does not use bolded keywords; the PLAN form does.
272
- 2. First task: failing end-to-end test for the happy path.
273
- 3. Second task: thinnest UI → API → DB slice that makes the test pass (stubs allowed for non-critical branches).
274
- 4. Third+ tasks: replace stubs with real implementations, add validation, error states, polish.
271
+ **Core rule (expansion tasks):** after each task a real user can do something they could not before. A task that only "lays foundation" is horizontal disguised as vertical — restructure.
275
272
 
276
- **Mode is all-or-nothing per phase** (PRD decision Q1). Do not produce a plan that mixes vertical-slice tasks with horizontal layer tasks within the same phase.
273
+ **`--no-tracer` (`TRACER_MODE=false`):** opt out of tracer-first and decompose into horizontal layers (the legacy default). Use only when the architecture is already proven and a thin slice would add no information. Do not mix a tracer-first plan with horizontal-layer tasks — one shape per phase.
277
274
 
278
- **Walking Skeleton mode** (`WALKING_SKELETON=true`, set by orchestrator for Phase 1 + new project under `--mvp`): The first deliverable is a Walking Skeleton — the thinnest possible end-to-end stack. In addition to `PLAN.md`, produce `SKELETON.md` using the template at `~/.claude/gsd-core/references/skeleton-template.md` (Read it now). `SKELETON.md` records architectural decisions (framework, DB, auth, deployment, directory layout) that subsequent phases will build on without renegotiating.
275
+ **MVP enrichment (`MVP_MODE=true`):** layered on top of the tracer-first ordering above (MVP no longer *turns on* vertical slices — that is now the default). It adds: (1) frame the phase goal as a user story at the top of `PLAN.md`, sourced from the ROADMAP `**Goal:**` line, bolding `**As a**` / `**I want to**` / `**so that**` (Read `~/.claude/gsd-core/references/user-story-template.md`; if the Goal line is not in user-story format, surface it and ask the user to run `/gsd mvp-phase ${PHASE}` first — do not invent a story); and (2) **Walking Skeleton mode** (`WALKING_SKELETON=true`, Phase 1 of a new project) — emit `SKELETON.md` from `~/.claude/gsd-core/references/skeleton-template.md` alongside `PLAN.md`. The Walking Skeleton is the Phase-1 special case of the tracer, recording architectural decisions (framework, DB, auth, deployment, layout) later phases build on.
279
276
 
280
- **Compatibility with TDD detection:** When both `MVP_MODE=true` and `workflow.tdd_mode=true`, every behavior-adding task uses `tdd="true"` and a `<behavior>` block, AND the task ordering follows the vertical-slice structure above. The first task is always a failing end-to-end test.
277
+ **TDD composition (`workflow.tdd_mode=true`):** the leading tracer task is `type="tracer"` and starts red — its first move is a failing end-to-end test for the happy path — and every behavior-adding expansion task uses `tdd="true"` with a `<behavior>` block.
281
278
 
282
279
  See @~/.claude/gsd-core/references/planner-guidance.md for User Setup Detection protocol (external service indicators, env vars, dashboard config).
283
280
 
@@ -291,30 +288,15 @@ See @~/.claude/gsd-core/references/planner-guidance.md for dependency graph buil
291
288
 
292
289
  <scope_estimation>
293
290
 
294
- ## Context Budget Rules
295
-
296
- Plans should complete within ~50% context (not 80%). No context anxiety, quality maintained start to finish, room for unexpected complexity.
291
+ ## Sizing and the Estimate Block
297
292
 
298
- **Each plan: 2-3 tasks maximum.**
293
+ Full rules: @~/.claude/gsd-core/references/context-budget.md (Phase Sizing). Read before sizing.
299
294
 
300
- | Context Weight | Tasks/Plan | Context/Task | Total |
301
- |----------------|------------|--------------|-------|
302
- | Light (CRUD, config) | 3 | ~10-15% | ~30-45% |
303
- | Medium (auth, payments) | 2 | ~20-30% | ~40-50% |
304
- | Heavy (migrations, multi-subsystem) | 1-2 | ~30-40% | ~30-50% |
305
-
306
- ## Split Signals
307
-
308
- **ALWAYS split if:**
309
- - More than 3 tasks
310
- - Multiple subsystems (DB + API + UI = separate plans)
311
- - Any task with >5 file modifications
312
- - Checkpoint + implementation in same plan
313
- - Discovery + implementation in same plan
314
-
315
- **CONSIDER splitting:** >5 files total, natural semantic boundaries, context cost estimate exceeds 40% for a single plan. See `<planner_authority_limits>` for prohibited split reasons.
316
-
317
- See @~/.claude/gsd-core/references/planner-guidance.md for Granularity Calibration table (Coarse/Standard/Fine plans-per-phase).
295
+ - **2-3 tasks per plan.** **ALWAYS split if:** >3 tasks, multiple subsystems, or any task touching >5 files.
296
+ - **Emit `estimate`**: run `estimate-calibration`; `tokens` = raw projection x factor, `raw_tokens` = that
297
+ projection before the factor (calibration measures actual/raw), `confidence` verbatim — derived from
298
+ sample count, never self-rated.
299
+ - **Over the smart-zone budget?** Re-slice: tracer + expansion slices. Advisory, never a block.
318
300
 
319
301
  </scope_estimation>
320
302
 
@@ -334,6 +316,12 @@ autonomous: true # false if plan has checkpoints
334
316
  requirements: [] # REQUIRED — Requirement IDs from ROADMAP this plan addresses. MUST NOT be empty.
335
317
  user_setup: [] # Human-required setup (omit if empty)
336
318
 
319
+ estimate: # Projected execution cost (see Estimate Emission)
320
+ tokens: 60000 # calibrated projection
321
+ raw_tokens: 30000 # pre-factor projection
322
+ tasks: 3 # task count the projection assumes
323
+ confidence: low # low | med | high — DERIVED from sample count, never self-rated
324
+
337
325
  must_haves:
338
326
  truths: [] # Observable behaviors
339
327
  artifacts: [] # Files that must exist
@@ -415,6 +403,7 @@ Create `.planning/phases/XX-name/{padded_phase}-{plan}-SUMMARY.md` when done
415
403
  | `autonomous` | Yes | `true` if no checkpoints |
416
404
  | `requirements` | Yes | **MUST** list requirement IDs from ROADMAP. Every roadmap requirement ID MUST appear in at least one plan. |
417
405
  | `user_setup` | No | Human-required setup items |
406
+ | `estimate` | No | Projected cost `{tokens, tasks, confidence}`. See Estimate Emission. |
418
407
  | `must_haves` | Yes | Goal-backward verification criteria |
419
408
 
420
409
  Wave numbers are pre-computed during planning. Execute-phase reads `wave` directly from frontmatter.
@@ -542,15 +531,9 @@ Do NOT use for: Deploying (use CLI), creating webhooks (use API), creating datab
542
531
 
543
532
  When Claude tries CLI/API and gets auth error → creates checkpoint → user authenticates → Claude retries. Auth gates are created dynamically, NOT pre-planned.
544
533
 
545
- ## Writing Guidelines
534
+ ## Writing Guidelines, Anti-Patterns, and Extended Examples
546
535
 
547
- **DO:** Automate everything before checkpoint, be specific ("Visit https://myapp.vercel.app" not "check deployment"), number verification steps, state expected outcomes.
548
-
549
- **DON'T:** Ask human to do work Claude can automate, mix multiple verifications, place checkpoints before automation completes.
550
-
551
- ## Anti-Patterns and Extended Examples
552
-
553
- For checkpoint anti-patterns, specificity comparison tables, context section anti-patterns, and scope reduction patterns:
536
+ For checkpoint writing guidelines (DO/DON'T), anti-patterns, specificity comparison tables, context section anti-patterns, and scope reduction patterns:
554
537
  @~/.claude/gsd-core/references/planner-antipatterns.md
555
538
 
556
539
  </checkpoints>
@@ -743,7 +726,7 @@ Read the most recent milestone retrospective and cross-milestone trends. Extract
743
726
  </step>
744
727
 
745
728
  <step name="inject_global_learnings">
746
- If `features.global_learnings` is `true`: run `gsd-tools query learnings.query --tag <tag> --limit 5` once per tag from PLAN.md frontmatter `tags` (or use the single most specific keyword). The handler matches one `--tag` at a time. Prefix matches with `[Prior learning from <project>]` as weak priors. Project-local decisions take precedence. Skip silently if disabled or no matches.
729
+ If `features.global_learnings` is `true`: run `gsd_run query learnings.query --tag <tag> --limit 5` once per tag from PLAN.md frontmatter `tags` (or use the single most specific keyword). The handler matches one `--tag` at a time. Prefix matches with `[Prior learning from <project>]` as weak priors. Project-local decisions take precedence. Skip silently if disabled or no matches.
747
730
  </step>
748
731
 
749
732
  <step name="gather_phase_context">
@@ -768,6 +751,8 @@ At decision points during plan creation, apply structured reasoning:
768
751
 
769
752
  Decompose phase into tasks. **Think dependencies first, not sequence.**
770
753
 
754
+ **Lead with the tracer.** Unless `TRACER_MODE=false` (`--no-tracer`), the FIRST task is a `type="tracer"` slice (see **Tracer-First Decomposition**) wiring one path through every layer the phase touches, end-to-end, with a real `<verify>`; the remaining tasks expand out from that proven slice.
755
+
771
756
  For each task:
772
757
  1. What does it NEED? (files, types, APIs that must exist)
773
758
  2. What does it CREATE? (files, types, APIs others might need)
@@ -102,7 +102,7 @@ For each item where `fetch` is present, invoke the MCP tool matching `fetch.prov
102
102
  | `exa` | `mcp__exa__web_search_exa` with `fetch.query` |
103
103
  | `tavily` | `mcp__tavily__search` with `fetch.query` |
104
104
  | `perplexity` | `mcp__perplexity__*` (use the appropriate perplexity MCP tool for the query) |
105
- | `brave` | `gsd-tools query websearch "<fetch.query>"` (Brave-backed) or built-in `WebSearch` |
105
+ | `brave` | `gsd_run query websearch "<fetch.query>"` (Brave-backed) or built-in `WebSearch` |
106
106
  | `firecrawl` | `mcp__firecrawl__scrape` with url (scrape kind) or `mcp__firecrawl__search` |
107
107
  | `websearch` | built-in `WebSearch` tool |
108
108
  | `webfetch` | built-in `WebFetch` tool |
@@ -490,7 +490,7 @@ Orchestrator provides: project name/description, research mode, project context,
490
490
 
491
491
  ## Step 3: Execute Research
492
492
 
493
- For each domain, use the `<tool_strategy>` seam (Steps A–D): build questions JSON, call `gsd-tools query research-plan`, run the indicated provider per item, then cache each digest. Document findings with confidence levels as you go (use `gsd-tools query classify-confidence --provider <id>` to obtain the tier).
493
+ For each domain, use the `<tool_strategy>` seam (Steps A–D): build questions JSON, call `gsd_run query research-plan`, run the indicated provider per item, then cache each digest. Document findings with confidence levels as you go (use `gsd_run query classify-confidence --provider <id>` to obtain the tier).
494
494
 
495
495
  ## Step 4: Quality Check
496
496
 
@@ -104,46 +104,6 @@ This gate runs unconditionally on every audit. The .gitignore ensures screenshot
104
104
 
105
105
  </gitignore_gate>
106
106
 
107
- <playwright_mcp_approach>
108
-
109
- ## Automated Screenshot Capture via Playwright-MCP (preferred when available)
110
-
111
- Before attempting the CLI screenshot approach, check whether `mcp__playwright__*`
112
- tools are available in this session. If they are, use them instead of the CLI approach:
113
-
114
- ```
115
- # Preferred: Playwright-MCP automated verification
116
- # 1. Navigate to the component URL
117
- mcp__playwright__navigate(url="http://localhost:3000")
118
-
119
- # 2. Take desktop screenshot
120
- mcp__playwright__screenshot(name="desktop", width=1440, height=900)
121
-
122
- # 3. Take mobile screenshot
123
- mcp__playwright__screenshot(name="mobile", width=375, height=812)
124
-
125
- # 4. For specific components listed in UI-SPEC.md, navigate to each
126
- # component route and capture targeted screenshots for comparison
127
- # against the spec's stated dimensions, colors, and layout.
128
-
129
- # 5. Compare screenshots against UI-SPEC.md requirements:
130
- # - Dimensions: Is component X width 70vw as specified?
131
- # - Color: Is the accent color applied only on declared elements?
132
- # - Layout: Are spacing values within the declared spacing scale?
133
- # Report any visual discrepancies as automated findings.
134
- ```
135
-
136
- **When Playwright-MCP is available:**
137
- - Use it for all screenshot capture (skip the CLI approach below)
138
- - Each UI checkpoint from UI-SPEC.md can be verified automatically
139
- - Discrepancies are reported as pillar findings with screenshot evidence
140
- - Items requiring subjective judgment are flagged as `needs_human_review: true`
141
-
142
- **When Playwright-MCP is NOT available:** fall back to the CLI screenshot approach
143
- below. Behavior is unchanged from the standard code-only audit path.
144
-
145
- </playwright_mcp_approach>
146
-
147
107
  <screenshot_approach>
148
108
 
149
109
  ## Screenshot Capture (CLI only — no MCP, no persistent browser)
@@ -537,11 +537,11 @@ grep -R -n -E 'probe-[^[:space:]]+\.sh|scripts/.*/tests/probe-.*\.sh' "$PHASE_DI
537
537
 
538
538
  1. Build the `PROBES` list from explicit PLAN declarations first; include conventional `scripts/*/tests/probe-*.sh` when the phase is a migration/tooling phase or the success criteria mention probes.
539
539
  2. For every documented probe path, if the file is missing or unreadable, mark `MISSING_PROBE` and set `status: gaps_found`. Do not require the executable bit because probes run through `bash "$probe"`.
540
- 3. Run each probe from the built `PROBES` list (declared + conventional) from the repository root:
540
+ 3. Run each probe from the built `PROBES` list from the repository root:
541
541
 
542
542
  ```bash
543
543
  for probe in "${PROBES[@]}"; do
544
- timeout 30s bash "$probe"
544
+ gsd_run run-with-timeout 30 -- bash "$probe"
545
545
  done
546
546
  ```
547
547