@opengsd/gsd-core 1.7.0-rc.6 → 1.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (195) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/.claude-plugin/plugin.json +1 -1
  3. package/.opencode/plugins/gsd-core.js +14 -0
  4. package/README.md +2 -0
  5. package/agents/gsd-debug-session-manager.md +42 -4
  6. package/agents/gsd-debugger.md +87 -29
  7. package/agents/gsd-executor.md +31 -3
  8. package/agents/gsd-planner.md +29 -36
  9. package/agents/gsd-security-auditor.md +13 -15
  10. package/agents/gsd-verifier.md +2 -2
  11. package/bin/install.js +1157 -84
  12. package/commands/gsd/ai-integration-phase.md +1 -1
  13. package/commands/gsd/mempalace-capture.md +31 -1
  14. package/commands/gsd/new-milestone.md +1 -1
  15. package/commands/gsd/plan-phase.md +5 -3
  16. package/commands/gsd/plan-review-convergence.md +3 -2
  17. package/commands/gsd/surface.md +6 -6
  18. package/gsd-core/bin/gsd-tools.cjs +1866 -2434
  19. package/gsd-core/bin/lib/adapter-imperative.cjs +8 -1
  20. package/gsd-core/bin/lib/agent-command-router.cjs +20 -5
  21. package/gsd-core/bin/lib/api-coverage.cjs +341 -49
  22. package/gsd-core/bin/lib/audit.cjs +7 -6
  23. package/gsd-core/bin/lib/broken-windows.cjs +716 -0
  24. package/gsd-core/bin/lib/capability-command-router.cjs +733 -0
  25. package/gsd-core/bin/lib/capability-registry.cjs +157 -88
  26. package/gsd-core/bin/lib/capability-writer.cjs +6 -1
  27. package/gsd-core/bin/lib/check-command-router.cjs +129 -26
  28. package/gsd-core/bin/lib/claude-orchestration-command-router.cjs +115 -27
  29. package/gsd-core/bin/lib/claude-orchestration.cjs +84 -9
  30. package/gsd-core/bin/lib/clock.cjs +19 -0
  31. package/gsd-core/bin/lib/command-aliases.cjs +14 -0
  32. package/gsd-core/bin/lib/commands.cjs +129 -13
  33. package/gsd-core/bin/lib/config-loader.cjs +20 -4
  34. package/gsd-core/bin/lib/config.cjs +81 -18
  35. package/gsd-core/bin/lib/core-utils.cjs +14 -3
  36. package/gsd-core/bin/lib/decisions.cjs +32 -8
  37. package/gsd-core/bin/lib/docs.cjs +6 -0
  38. package/gsd-core/bin/lib/drift.cjs +4 -4
  39. package/gsd-core/bin/lib/external-descriptor-trust.cjs +14 -2
  40. package/gsd-core/bin/lib/frontmatter.cjs +22 -0
  41. package/gsd-core/bin/lib/gap-checker.cjs +17 -2
  42. package/gsd-core/bin/lib/gsd2-import.cjs +2 -1
  43. package/gsd-core/bin/lib/init.cjs +138 -60
  44. package/gsd-core/bin/lib/install-engine.cjs +301 -25
  45. package/gsd-core/bin/lib/install-profiles.cjs +239 -1
  46. package/gsd-core/bin/lib/installer-migration-authoring.cjs +2 -1
  47. package/gsd-core/bin/lib/installer-migrations/005-opencode-baseline-commands-dir.cjs +146 -0
  48. package/gsd-core/bin/lib/installer-migrations/006-pi-extension-cjs-to-js.cjs +91 -0
  49. package/gsd-core/bin/lib/installer-migrations.cjs +45 -6
  50. package/gsd-core/bin/lib/markdown-sectionizer.cjs +449 -0
  51. package/gsd-core/bin/lib/markdown-table.cjs +698 -0
  52. package/gsd-core/bin/lib/milestone.cjs +463 -43
  53. package/gsd-core/bin/lib/model-catalog.cjs +19 -4
  54. package/gsd-core/bin/lib/model-resolver.cjs +189 -7
  55. package/gsd-core/bin/lib/onboard-projection.cjs +11 -8
  56. package/gsd-core/bin/lib/phase-command-router.cjs +50 -2
  57. package/gsd-core/bin/lib/phase-id.cjs +26 -4
  58. package/gsd-core/bin/lib/phase-lifecycle.cjs +62 -36
  59. package/gsd-core/bin/lib/phase-locator.cjs +23 -2
  60. package/gsd-core/bin/lib/phase.cjs +636 -72
  61. package/gsd-core/bin/lib/plan-scan.cjs +73 -2
  62. package/gsd-core/bin/lib/roadmap-parser.cjs +225 -17
  63. package/gsd-core/bin/lib/roadmap.cjs +113 -52
  64. package/gsd-core/bin/lib/runtime-artifact-conversion.cjs +14 -7
  65. package/gsd-core/bin/lib/runtime-artifact-install-plan.cjs +3 -2
  66. package/gsd-core/bin/lib/runtime-artifact-layout.cjs +24 -9
  67. package/gsd-core/bin/lib/runtime-hooks-surface.cjs +41 -17
  68. package/gsd-core/bin/lib/schema-detect.cjs +2 -1
  69. package/gsd-core/bin/lib/security.cjs +1 -1
  70. package/gsd-core/bin/lib/shell-command-projection.cjs +61 -25
  71. package/gsd-core/bin/lib/smart-entry.cjs +73 -7
  72. package/gsd-core/bin/lib/state-document.cjs +7 -4
  73. package/gsd-core/bin/lib/state-transition.cjs +122 -46
  74. package/gsd-core/bin/lib/state.cjs +456 -137
  75. package/gsd-core/bin/lib/surface.cjs +53 -11
  76. package/gsd-core/bin/lib/template.cjs +2 -1
  77. package/gsd-core/bin/lib/uat.cjs +474 -13
  78. package/gsd-core/bin/lib/ui-safety-gate.cjs +23 -1
  79. package/gsd-core/bin/lib/validate.cjs +12 -8
  80. package/gsd-core/bin/lib/verification.cjs +112 -17
  81. package/gsd-core/bin/lib/verify.cjs +224 -25
  82. package/gsd-core/bin/lib/workstream.cjs +3 -2
  83. package/gsd-core/bin/lib/worktree-safety.cjs +1 -1
  84. package/gsd-core/bin/lib/write-set.cjs +38 -0
  85. package/gsd-core/bin/shared/config-schema.manifest.json +5 -2
  86. package/gsd-core/references/api-coverage.md +37 -7
  87. package/gsd-core/references/checkpoints.md +13 -1
  88. package/gsd-core/references/common-bug-patterns.md +13 -0
  89. package/gsd-core/references/debugger-bug-taxonomy.md +111 -0
  90. package/gsd-core/references/debugger-fix-acceptance.md +157 -0
  91. package/gsd-core/references/debugger-philosophy.md +1 -0
  92. package/gsd-core/references/debugger-prevention.md +98 -0
  93. package/gsd-core/references/debugger-rca-branching.md +98 -0
  94. package/gsd-core/references/debugger-repro-hardening.md +130 -0
  95. package/gsd-core/references/debugger-sbfl.md +110 -0
  96. package/gsd-core/references/debugger-semantic-recall.md +81 -0
  97. package/gsd-core/references/execute-phase-quota-recovery.md +55 -0
  98. package/gsd-core/references/execute-phase-requirement-revert.md +8 -0
  99. package/gsd-core/references/execute-phase-response-language.md +7 -0
  100. package/gsd-core/references/planner-antipatterns.md +6 -0
  101. package/gsd-core/references/planner-mvp-mode.md +12 -13
  102. package/gsd-core/references/planner-preconditions.md +156 -0
  103. package/gsd-core/references/planner-reversibility.md +132 -0
  104. package/gsd-core/references/reviewer-instances.md +9 -7
  105. package/gsd-core/references/skeleton-template.md +1 -1
  106. package/gsd-core/references/thinking-models-planning.md +3 -1
  107. package/gsd-core/templates/DEBUG.md +5 -3
  108. package/gsd-core/workflows/add-phase.md +2 -0
  109. package/gsd-core/workflows/add-tests.md +4 -2
  110. package/gsd-core/workflows/add-todo.md +32 -1
  111. package/gsd-core/workflows/ai-integration-phase.md +4 -2
  112. package/gsd-core/workflows/audit-fix.md +2 -2
  113. package/gsd-core/workflows/check-todos.md +3 -1
  114. package/gsd-core/workflows/cleanup.md +7 -1
  115. package/gsd-core/workflows/code-review.md +17 -5
  116. package/gsd-core/workflows/complete-milestone.md +3 -0
  117. package/gsd-core/workflows/debug.md +27 -5
  118. package/gsd-core/workflows/diagnose-issues.md +1 -1
  119. package/gsd-core/workflows/discovery-phase.md +7 -0
  120. package/gsd-core/workflows/discuss-phase/templates/context.md +16 -2
  121. package/gsd-core/workflows/discuss-phase-assumptions.md +3 -0
  122. package/gsd-core/workflows/do.md +7 -1
  123. package/gsd-core/workflows/docs-update.md +1 -0
  124. package/gsd-core/workflows/eval-review.md +3 -0
  125. package/gsd-core/workflows/execute-phase/steps/post-merge-gate.md +4 -4
  126. package/gsd-core/workflows/execute-phase/steps/regression-gate.md +2 -2
  127. package/gsd-core/workflows/execute-phase.md +30 -37
  128. package/gsd-core/workflows/execute-plan.md +15 -4
  129. package/gsd-core/workflows/fast.md +8 -22
  130. package/gsd-core/workflows/graduation.md +3 -0
  131. package/gsd-core/workflows/health.md +7 -1
  132. package/gsd-core/workflows/help/modes/full.md +6 -2
  133. package/gsd-core/workflows/import.md +8 -2
  134. package/gsd-core/workflows/inbox.md +7 -0
  135. package/gsd-core/workflows/ingest-docs.md +15 -10
  136. package/gsd-core/workflows/manager.md +3 -1
  137. package/gsd-core/workflows/map-codebase.md +4 -4
  138. package/gsd-core/workflows/mvp-phase.md +3 -0
  139. package/gsd-core/workflows/new-milestone.md +69 -21
  140. package/gsd-core/workflows/new-project.md +17 -15
  141. package/gsd-core/workflows/new-workspace.md +3 -1
  142. package/gsd-core/workflows/onboard.md +3 -0
  143. package/gsd-core/workflows/plan-phase.md +14 -5
  144. package/gsd-core/workflows/plan-review-convergence.md +48 -3
  145. package/gsd-core/workflows/plant-seed.md +3 -0
  146. package/gsd-core/workflows/profile-user.md +7 -1
  147. package/gsd-core/workflows/progress.md +33 -5
  148. package/gsd-core/workflows/quick.md +21 -7
  149. package/gsd-core/workflows/remove-workspace.md +3 -0
  150. package/gsd-core/workflows/review.md +123 -68
  151. package/gsd-core/workflows/scan.md +1 -1
  152. package/gsd-core/workflows/secure-phase.md +4 -1
  153. package/gsd-core/workflows/settings-integrations.md +3 -0
  154. package/gsd-core/workflows/settings.md +3 -0
  155. package/gsd-core/workflows/ship.md +58 -5
  156. package/gsd-core/workflows/sketch.md +3 -0
  157. package/gsd-core/workflows/smart-entry.md +3 -0
  158. package/gsd-core/workflows/spec-phase.md +1 -1
  159. package/gsd-core/workflows/spike.md +7 -1
  160. package/gsd-core/workflows/transition.md +1 -1
  161. package/gsd-core/workflows/ui-phase.md +3 -1
  162. package/gsd-core/workflows/ui-review.md +3 -0
  163. package/gsd-core/workflows/undo.md +7 -0
  164. package/gsd-core/workflows/update.md +2 -0
  165. package/gsd-core/workflows/validate-phase.md +3 -0
  166. package/gsd-core/workflows/verify-phase.md +2 -2
  167. package/gsd-core/workflows/verify-work.md +7 -3
  168. package/hooks/dist/gsd-context-monitor.js +27 -9
  169. package/hooks/dist/gsd-statusline.js +252 -17
  170. package/hooks/gsd-context-monitor.js +27 -9
  171. package/hooks/gsd-statusline.js +252 -17
  172. package/package.json +8 -4
  173. package/pi/gsd.cjs +8 -2
  174. package/scripts/changeset/lint.cjs +1 -0
  175. package/scripts/changeset/parse.cjs +26 -0
  176. package/scripts/check-glossary-refs.cjs +220 -0
  177. package/scripts/ci-rebase-check.cjs +48 -4
  178. package/scripts/ci-test-scope.cjs +39 -1
  179. package/scripts/gen-adr-index.cjs +526 -0
  180. package/scripts/gen-golden-install-parity-zcode.cjs +35 -45
  181. package/scripts/gen-install-tree-fixtures.cjs +75 -0
  182. package/scripts/gen-test-timings.cjs +201 -0
  183. package/scripts/lint-allow-test-rule-refs.allowlist.json +0 -1
  184. package/scripts/lint-portable-timeout.cjs +140 -0
  185. package/scripts/lint-table-schema-drift.cjs +157 -0
  186. package/scripts/lint-test-file-count.allowlist.json +1 -0
  187. package/scripts/release-tarball-smoke.cjs +18 -11
  188. package/scripts/run-tests.cjs +420 -58
  189. package/skills/gsd-ai-integration-phase/SKILL.md +1 -1
  190. package/skills/gsd-mempalace-capture/SKILL.md +31 -1
  191. package/skills/gsd-new-milestone/SKILL.md +1 -1
  192. package/skills/gsd-plan-phase/SKILL.md +5 -3
  193. package/skills/gsd-plan-review-convergence/SKILL.md +3 -2
  194. package/skills/gsd-surface/SKILL.md +6 -6
  195. package/vscode/package.json +1 -1
@@ -9,7 +9,7 @@
9
9
  {
10
10
  "name": "gsd-core",
11
11
  "description": "GSD Core is a meta-prompting, context engineering, and spec-driven development system for AI coding agents.",
12
- "version": "1.7.0-rc.6",
12
+ "version": "1.8.0",
13
13
  "source": "./",
14
14
  "author": {
15
15
  "name": "open-gsd",
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "gsd-core",
3
3
  "displayName": "GSD Core",
4
- "version": "1.7.0-rc.6",
4
+ "version": "1.8.0",
5
5
  "description": "GSD Core is a meta-prompting, context engineering, and spec-driven development system for AI coding agents.",
6
6
  "author": {
7
7
  "name": "open-gsd",
@@ -200,9 +200,23 @@ function mapToolInput(args) {
200
200
  * @param {string} [opts.cwd] working directory for the child
201
201
  * @returns {{ stdout: string, exitCode: number, timedOut: boolean }}
202
202
  */
203
+ const warnedMissingHooks = new Set();
204
+
203
205
  function runHook(hookFile, payload, opts = {}) {
204
206
  const hookPath = path.join(HOOKS_DIR, hookFile);
205
207
  if (!fs.existsSync(hookPath)) {
208
+ // A missing guard script means the guard is silently NOT enforced — the
209
+ // exact failure mode of #2305 (plugin staged, hooks bundle not). Never
210
+ // break the tool call (the adapter's design contract), but never be
211
+ // silent about it either: warn loudly, once per hook file.
212
+ if (!warnedMissingHooks.has(hookFile)) {
213
+ warnedMissingHooks.add(hookFile);
214
+ console.error(
215
+ `[gsd-core] hook script missing: ${hookPath} — ${hookFile} is NOT ` +
216
+ "enforced. The GSD install may be incomplete; reinstall (or run " +
217
+ "/gsd-update) to restage the hooks/ bundle.",
218
+ );
219
+ }
206
220
  return { stdout: "", exitCode: 0, timedOut: false };
207
221
  }
208
222
  const timeout = opts.timeout ?? 8000;
package/README.md CHANGED
@@ -60,6 +60,8 @@ New here? Follow [Your first project](docs/tutorials/your-first-project.md) for
60
60
 
61
61
  ## Documentation
62
62
 
63
+ **What's new in 1.7.0** → [docs/whats-new-1.7.0.md](docs/whats-new-1.7.0.md)
64
+
63
65
  **Tutorials** — learning by doing:
64
66
  - [Your first project](docs/tutorials/your-first-project.md)
65
67
  - [Onboarding an existing codebase](docs/tutorials/onboarding-an-existing-codebase.md)
@@ -270,30 +270,67 @@ If user selects 1 or 2: spawn continuation agent (with any additional context pr
270
270
 
271
271
  If user selects 3: proceed to Step 4 with fix = "not applied".
272
272
 
273
+ ### 3f. FIX REJECTED BY GUARDRAIL
274
+
275
+ When agent returns `## FIX REJECTED BY GUARDRAIL`:
276
+
277
+ Present the failing signal and evidence to the user via AskUserQuestion:
278
+ ```
279
+ Fix rejected by the acceptance guardrail.
280
+
281
+ Failing signal: {failing signal}
282
+ Evidence: {why it failed}
283
+
284
+ Options:
285
+ 1. Revise fix — spawn continuation agent to revise the fix so the signal passes
286
+ 2. Accept as technical debt — record the unmet signal + justification (the fix lands without the gate passing; this is never silent)
287
+ 3. Abandon — stop; session stays unresolved
288
+ ```
289
+
290
+ If user selects 1: spawn continuation agent with `goal: find_and_fix` naming the failing signal to revise. Loop back to Step 3.
291
+
292
+ If user selects 2: spawn continuation agent instructed to record `guardrail_verdict: accepted_debt` + the justification in the debug file, then proceed to request_human_verification. Loop back to Step 3.
293
+
294
+ If user selects 3: proceed to Step 4 with fix = "not applied (guardrail rejected)".
295
+
273
296
  ## Step 4: Return Compact Summary
274
297
 
298
+ **Non-terminal early stop — check this FIRST.** Before returning any summary below, ask: is your own turn/context budget exhausted while the debugger (`gsd-debugger`) is still investigating — i.e. you have NOT reached `DEBUG COMPLETE`, a user-chosen `ABANDONED`, or exhausted the `INVESTIGATION INCONCLUSIVE` options? If so, do NOT fabricate a `DEBUG SESSION COMPLETE` or `ABANDONED` summary to fit this shape. Return the non-terminal marker instead:
299
+
300
+ ```markdown
301
+ ## CONTINUE_REQUIRED
302
+
303
+ **Session:** {debug_file_path}
304
+ **Status:** {status from frontmatter, e.g. investigating}
305
+ **Next action:** {next_action from Current Focus}
306
+ **Reason:** session-manager turn/context budget exhausted — investigation still in progress
307
+ ```
308
+
309
+ `CONTINUE_REQUIRED` is distinct from both terminal shapes below AND from `## CHECKPOINT REACHED` (Step 3d): a `CHECKPOINT REACHED` is a genuine user-input/approval checkpoint that already correctly pauses via `AskUserQuestion` before looping back to Step 3 — it is not returned to the orchestrator. `CONTINUE_REQUIRED` is emitted only when no checkpoint is pending and the loop simply cannot proceed further in this turn. The orchestrator resumes by re-spawning this agent with the SAME `slug`/`debug_file_path` — the on-disk checkpoint at `.planning/debug/{slug}.md` (its `status` and `next_action`) is the source of truth for where to pick up. Never return control to the user as if the session were complete when it is not.
310
+
275
311
  Read the resolved (or current) debug file to extract final Resolution values.
276
312
 
277
- Return compact summary:
313
+ Return compact summary (terminal — investigation resolved):
278
314
 
279
315
  ```markdown
280
316
  ## DEBUG SESSION COMPLETE
281
317
 
282
318
  **Session:** {final path — resolved/ if archived, otherwise debug_file_path}
283
- **Root Cause:** {one sentence from Resolution.root_cause, or "not determined"}
319
+ **Root Cause:** {one sentence, or a '; '-joined list when the AND-gate identified multiple contributing causes, from Resolution.root_cause; or "not determined"}
284
320
  **Fix:** {one sentence from Resolution.fix, or "not applied"}
285
321
  **Cycles:** {N} (investigation) + {M} (fix)
286
322
  **TDD:** {yes/no}
287
323
  **Specialist review:** {specialist_hint used, or "none"}
324
+ **Prevention:** {one-line from the blameless postmortem — "why not caught: <gate, or 'none (no gate existed for this class)'>; guard: <artifact>"}
288
325
  ```
289
326
 
290
- If the session was abandoned by user choice, return:
327
+ If the session was abandoned by user choice, return (terminal — user stopped):
291
328
 
292
329
  ```markdown
293
330
  ## DEBUG SESSION COMPLETE
294
331
 
295
332
  **Session:** {debug_file_path}
296
- **Root Cause:** {one sentence if found, or "not determined"}
333
+ **Root Cause:** {one sentence if found (or a '; '-joined list if the AND-gate identified multiple contributing causes), or "not determined"}
297
334
  **Fix:** not applied
298
335
  **Cycles:** {N}
299
336
  **TDD:** {yes/no}
@@ -311,5 +348,6 @@ If the session was abandoned by user choice, return:
311
348
  - [ ] Specialist dispatch executed when specialist_dispatch_enabled and hint maps to a skill
312
349
  - [ ] TDD gate applied when tdd_mode=true and ROOT CAUSE FOUND
313
350
  - [ ] Loop continues until DEBUG COMPLETE, ABANDONED, or user stops
351
+ - [ ] Non-terminal `CONTINUE_REQUIRED` (not a fabricated terminal summary) returned when the manager's own turn/context budget is exhausted mid-investigation
314
352
  - [ ] Compact summary returned (at most 2K tokens)
315
353
  </success_criteria>
@@ -253,6 +253,10 @@ reasoning_checkpoint:
253
253
  falsification_test: "[what specific observation would prove this hypothesis wrong]"
254
254
  fix_rationale: "[why the proposed fix addresses the root cause — not just the symptom]"
255
255
  blind_spots: "[what you haven't tested that could invalidate this hypothesis]"
256
+ candidate_causes:
257
+ - "[cause in category: code|config|environment|data]"
258
+ - "[cause in a DIFFERENT category — single-category is not a branch]"
259
+ and_gate: "[could this failure require >1 contributing condition simultaneously? yes/no + why — see RCA branching]"
256
260
  ```
257
261
 
258
262
  **Check before proceeding:**
@@ -260,8 +264,9 @@ reasoning_checkpoint:
260
264
  - Is the confirming evidence direct observation, not inference?
261
265
  - Does the fix address the root cause or a symptom?
262
266
  - Have you documented your blind spots honestly?
267
+ - **Did you branch across ≥2 categories and answer the AND-gate?** (Single-cause is fine when the AND-gate is no — but you must have checked.)
263
268
 
264
- If you cannot fill all five fields with specific, concrete answers — you do not have a confirmed root cause yet. Return to investigation_loop.
269
+ If you cannot fill all seven fields with specific, concrete answers — you do not have a confirmed root cause yet. Return to investigation_loop.
265
270
 
266
271
  ## Minimal Reproduction
267
272
 
@@ -274,6 +279,7 @@ If you cannot fill all five fields with specific, concrete answers — you do no
274
279
  3. Test: Does it still reproduce? YES = keep removed. NO = put back.
275
280
  4. Repeat until bare minimum
276
281
  5. Bug is now obvious in stripped-down code
282
+ 6. **Shrinking (input-space bugs)** — when the bug triggers on a class of inputs, wrap it in a property (fast-check for JS/TS, Hypothesis for Python) and let the shrinker auto-minimize the counterexample; store the **minimized** input as the regression seed. See `gsd-core/references/debugger-repro-hardening.md`.
277
283
 
278
284
  **Example:**
279
285
  ```jsx
@@ -443,18 +449,21 @@ MISMATCH: Checker looks in wrong directory → hooks "not found" → reported as
443
449
 
444
450
  **The discipline:** Never assume a constructed path is correct. Resolve it to its actual value and verify the other side agrees. When two systems share a resource (file, directory, key), trace the full path in both.
445
451
 
446
- ## Technique Selection
452
+ ## Technique Selection (routed by bug class)
447
453
 
448
- | Situation | Technique |
449
- |-----------|-----------|
450
- | Large codebase, many files | Binary search |
451
- | Confused about what's happening | Rubber duck, Observability first |
452
- | Complex system, many interactions | Minimal reproduction |
453
- | Know the desired output | Working backwards |
454
- | Used to work, now doesn't | Differential debugging, Git bisect |
455
- | Many possible causes | Comment out everything, Binary search |
456
- | Paths, URLs, keys constructed from variables | Follow the indirection |
457
- | Always | Observability first (before making changes) |
454
+ Classify the failure first (Phase 1.75), then route by class — not by ad-hoc
455
+ situation:
456
+
457
+ @~/.claude/gsd-core/references/debugger-bug-taxonomy.md
458
+
459
+ | bug_class | Route to | Revoke if already run |
460
+ |---|---|---|
461
+ | Bohrbug | deterministic reproduction → SBFL (Phase 1.25) → git bisect → binary search | — |
462
+ | Heisenbug / Mandelbug | record-replay (`rr`) → stability-stress → statistical sampling | SBFL — Phase 1.25 runs before classification; if it ran, mark its Evidence entry revoked (flaky spectrum poisons the ranking) |
463
+ | Concurrency | atomicity / order / deadlock checklist (see reference) FIRST | — |
464
+ | General (any class) | Binary search, Working backwards, Differential, Delta debugging, Comment-out-everything, Follow-the-indirection, Rubber duck, Observability first (always, before changes) | — |
465
+
466
+ The class rows pick the first move; the General lane holds situation-cued techniques that apply to any class. When the situation table and the class route disagree, the class route wins.
458
467
 
459
468
  ## Combining Techniques
460
469
 
@@ -590,6 +599,13 @@ function processUserData(user) {
590
599
  // 5. Test is now regression protection forever
591
600
  ```
592
601
 
602
+ **Harden the regression test (so the Phase 1A mutation guardrail bites):**
603
+
604
+ @~/.claude/gsd-core/references/debugger-repro-hardening.md
605
+
606
+ - **Classify the oracle** before writing the assertion — `specified` / `derived` (contract/model) / `metamorphic` / `implicit` (crash, weakest). Record it under `Resolution.oracle_type`. Never default to implicit silently.
607
+ - **Add boundary neighbors** around the fixed defect's equivalence class — off-by-one (N±1), min/max (0/length), empty/singleton — the single reported value misses the adjacent off-by-one.
608
+
593
609
  ## Verification Checklist
594
610
 
595
611
  ```markdown
@@ -788,9 +804,11 @@ Each resolved session appends one entry:
788
804
  ## {slug} — {one-line description}
789
805
  - **Date:** {ISO date}
790
806
  - **Error patterns:** {comma-separated keywords extracted from symptoms.errors and symptoms.actual}
791
- - **Root cause:** {from Resolution.root_cause}
807
+ - **Root cause(s):** {from Resolution.root_cause — one cause, or a '; '-joined list when the AND-gate fired}
792
808
  - **Fix:** {from Resolution.fix}
793
809
  - **Files changed:** {from Resolution.files_changed}
810
+ - **Why not caught:** {which existing gate (test/typecheck/lint/review/verify/build) should have caught it — or "no gate existed for this class"}
811
+ - **Recurrence guard:** {the concrete artifact preventing this class from returning — regression test (path:name) / assertion / lint rule / type refinement / config-default change / KB pattern}
794
812
  ---
795
813
  ```
796
814
 
@@ -804,9 +822,11 @@ At the **end of `archive_session`**, after the session file is moved to `resolve
804
822
 
805
823
  ## Matching Logic
806
824
 
807
- Matching is keyword overlap, not semantic similarity. Extract nouns and error substrings from `Symptoms.errors` and `Symptoms.actual`. Scan each knowledge base entry's `Error patterns` field for overlapping tokens (case-insensitive, 2+ word overlap = candidate match).
825
+ **Semantic-first, keyword-fallback.** Query MemPalace with the current symptoms and surface the top-k meaning-similar prior resolutions — this catches same-root-cause/different-wording cases keyword overlap misses. Fall back to keyword overlap on `knowledge-base.md` when MemPalace is absent. See:
826
+
827
+ @~/.claude/gsd-core/references/debugger-semantic-recall.md
808
828
 
809
- **Important:** A match is a **hypothesis candidate**, not a confirmed diagnosis. Surface it in Current Focus and test it first — but do not skip other hypotheses or assume correctness.
829
+ **Important:** A match is a **hypothesis candidate**, not a confirmed diagnosis — surface it in Current Focus and test it first; do not skip other hypotheses or assume correctness.
810
830
 
811
831
  </knowledge_base_protocol>
812
832
 
@@ -966,12 +986,10 @@ At investigation decision points, apply structured reasoning:
966
986
  **Autonomous investigation. Update file continuously.**
967
987
 
968
988
  **Phase 0: Check knowledge base**
969
- - If `.planning/debug/knowledge-base.md` exists, read it
970
- - Extract keywords from `Symptoms.errors` and `Symptoms.actual` (nouns, error substrings, identifiers)
971
- - Scan knowledge base entries for 2+ keyword overlap (case-insensitive)
989
+ - Query MemPalace semantically with the current symptoms (top-k meaning-similar prior resolutions); fall back to reading `.planning/debug/knowledge-base.md` and keyword overlap when MemPalace is absent
972
990
  - If match found:
973
991
  - Note in Current Focus: `known_pattern_candidate: "{matched slug} — {description}"`
974
- - Add to Evidence: `found: Knowledge base match on [{keywords}] → Root cause was: {root_cause}. Fix was: {fix}.`
992
+ - Add to Evidence: `found: Knowledge base match on [{keywords}] → Root cause was: {root_cause}. Fix was: {fix}. Why not caught: {why_not_caught}. Recurrence guard: {recurrence_guard}.` (the last two are absent on old entries — that's fine; consume them when present)
975
993
  - Test this hypothesis FIRST in Phase 2 — but treat it as one hypothesis, not a certainty
976
994
  - If no match: proceed normally
977
995
 
@@ -983,14 +1001,32 @@ At investigation decision points, apply structured reasoning:
983
1001
  - Run app/tests to observe behavior
984
1002
  - APPEND to Evidence after each finding
985
1003
 
1004
+ **Phase 1.25: Spectrum-based fault localization (optional, coverage-gated)**
1005
+ - When a runnable test suite with per-test coverage exists (≥1 failing AND ≥1 passing test), compute an Ochiai suspiciousness ranking and seed the top-N into Evidence before forming hypotheses — narrows the search space deterministically before LLM reasoning:
1006
+
1007
+ @~/.claude/gsd-core/references/debugger-sbfl.md
1008
+
1009
+ - Skip with a logged note when there is no test suite, no failing tests, or no per-test coverage; investigation proceeds unchanged
1010
+
986
1011
  **Phase 1.5: Check common bug patterns**
987
1012
  - Read @~/.claude/gsd-core/references/common-bug-patterns.md
988
1013
  - Match symptoms to pattern categories using the Symptom-to-Category Quick Map
989
1014
  - Any matching patterns become hypothesis candidates for Phase 2
990
1015
  - If no patterns match, proceed to open-ended hypothesis formation
991
1016
 
1017
+ **Phase 1.75: Classify the failure**
1018
+ - Assign a `bug_class` — Bohrbug (deterministic) / Heisenbug-Mandelbug (transient, non-deterministic) / Concurrency — and record it in Current Focus. The class routes which investigation technique to use:
1019
+
1020
+ @~/.claude/gsd-core/references/debugger-bug-taxonomy.md
1021
+
1022
+ - Bohrbug → reproduction + SBFL + bisect; Heisenbug/Mandelbug → record-replay/stability (skip SBFL — flaky spectra poison it); Concurrency → the atomicity/order/deadlock checklist first
1023
+
992
1024
  **Phase 2: Form hypothesis**
993
1025
  - Based on evidence AND common pattern matches, form SPECIFIC, FALSIFIABLE hypothesis
1026
+ - **Branch, don't chain** — at hypothesis formation (so it's done before the Phase 4 commit), enumerate candidate causes across ≥2 Ishikawa categories (code / config / environment / data) and answer the AND-gate check; `root_cause` may hold a set when the AND-gate fires:
1027
+
1028
+ @~/.claude/gsd-core/references/debugger-rca-branching.md
1029
+
994
1030
  - Update Current Focus with hypothesis, test, expecting, next_action
995
1031
 
996
1032
  **Phase 3: Test hypothesis**
@@ -1043,7 +1079,7 @@ Return structured diagnosis:
1043
1079
 
1044
1080
  **Debug Session:** .planning/debug/{slug}.md
1045
1081
 
1046
- **Root Cause:** {from Resolution.root_cause}
1082
+ **Root Cause:** {from Resolution.root_cause — one cause, or a '; '-joined list when the AND-gate identified multiple contributing causes}
1047
1083
 
1048
1084
  **Evidence Summary:**
1049
1085
  - {key finding 1}
@@ -1083,7 +1119,7 @@ Update status to "fixing".
1083
1119
 
1084
1120
  **0. Structured Reasoning Checkpoint (MANDATORY)**
1085
1121
  - Write the `reasoning_checkpoint` block to Current Focus (see Structured Reasoning Checkpoint in investigation_techniques)
1086
- - Verify all five fields can be filled with specific, concrete answers
1122
+ - Verify every field can be filled with specific, concrete answers — including the RCA `candidate_causes` (≥2 categories) and `and_gate` fields
1087
1123
  - If any field is vague or empty: return to investigation_loop — root cause is not confirmed
1088
1124
 
1089
1125
  **1. Implement minimal fix**
@@ -1091,11 +1127,15 @@ Update status to "fixing".
1091
1127
  - Make SMALLEST change that addresses root cause
1092
1128
  - Update Resolution.fix and Resolution.files_changed
1093
1129
 
1094
- **2. Verify**
1130
+ **2. Verify (Fix-Acceptance Guardrail)**
1095
1131
  - Update status to "verifying"
1096
- - Test against original Symptoms
1097
- - If verification FAILS: status -> "investigating", return to investigation_loop
1098
- - If verification PASSES: Update Resolution.verification, proceed to request_human_verification
1132
+ - Run the multi-signal guardrail before accepting the fix:
1133
+
1134
+ @~/.claude/gsd-core/references/debugger-fix-acceptance.md
1135
+
1136
+ - Record every signal's result under `Resolution.verification` (per-signal schema in the reference)
1137
+ - If ANY applicable signal fails (and no documented technical-debt escape applies): return `## FIX REJECTED BY GUARDRAIL` (see structured_returns) — do NOT request human verification
1138
+ - If all applicable signals pass: set `guardrail_verdict: accepted`, proceed to request_human_verification
1099
1139
  </step>
1100
1140
 
1101
1141
  <step name="request_human_verification">
@@ -1174,9 +1214,13 @@ Then commit planning docs via CLI (respects `commit_docs` config automatically):
1174
1214
  gsd_run query commit "docs: resolve debug {slug}" --files .planning/debug/resolved/{slug}.md
1175
1215
  ```
1176
1216
 
1177
- **Append to knowledge base:**
1217
+ **Append to knowledge base (with the Prevention block):**
1218
+
1219
+ Read `.planning/debug/resolved/{slug}.md` to extract final `Resolution` values. Then produce the **Prevention block** — a blameless postmortem (branching 5-Whys per RCA, "why wasn't this caught?", and a concrete recurrence guard):
1220
+
1221
+ @~/.claude/gsd-core/references/debugger-prevention.md
1178
1222
 
1179
- Read `.planning/debug/resolved/{slug}.md` to extract final `Resolution` values. Then append to `.planning/debug/knowledge-base.md` (create file with header if it doesn't exist):
1223
+ Then append to `.planning/debug/knowledge-base.md` (create file with header if it doesn't exist):
1180
1224
 
1181
1225
  If creating for the first time, write this header first:
1182
1226
  ```markdown
@@ -1193,9 +1237,11 @@ Then append the entry:
1193
1237
  ## {slug} — {one-line description of the bug}
1194
1238
  - **Date:** {ISO date}
1195
1239
  - **Error patterns:** {comma-separated keywords from Symptoms.errors + Symptoms.actual}
1196
- - **Root cause:** {Resolution.root_cause}
1240
+ - **Root cause(s):** {Resolution.root_cause — joined as '; ' when multiple contributing causes were confirmed}
1197
1241
  - **Fix:** {Resolution.fix}
1198
1242
  - **Files changed:** {Resolution.files_changed joined as comma list}
1243
+ - **Why not caught:** {which existing gate (test/typecheck/lint/review/verify/build) should have caught it — or "no gate existed for this class"}
1244
+ - **Recurrence guard:** {concrete artifact preventing this class from returning — regression test (path:name) / assertion / lint rule / KB pattern / type refinement / config-default change}
1199
1245
  ---
1200
1246
 
1201
1247
  ```
@@ -1205,6 +1251,8 @@ Commit the knowledge base update alongside the resolved session:
1205
1251
  gsd_run query commit "docs: update debug knowledge base with {slug}" --files .planning/debug/knowledge-base.md
1206
1252
  ```
1207
1253
 
1254
+ **Index into MemPalace (when available)** per the semantic-recall reference — the Resolution summary (not raw symptoms), redacted — so a future Phase-0 query surfaces it by meaning. Skip with a logged note when MemPalace is absent or the KB write failed; `knowledge-base.md` is the durable fallback.
1255
+
1208
1256
  Report completion and offer next steps.
1209
1257
  </step>
1210
1258
 
@@ -1298,7 +1346,7 @@ Orchestrator presents checkpoint to user, gets response, spawns fresh continuati
1298
1346
 
1299
1347
  **Debug Session:** .planning/debug/{slug}.md
1300
1348
 
1301
- **Root Cause:** {specific cause with evidence}
1349
+ **Root Cause:** {specific cause with evidence — one cause, or a '; '-joined list when the AND-gate identified multiple contributing causes}
1302
1350
 
1303
1351
  **Evidence Summary:**
1304
1352
  - {key finding 1}
@@ -1334,6 +1382,16 @@ Orchestrator presents checkpoint to user, gets response, spawns fresh continuati
1334
1382
 
1335
1383
  Only return this after human verification confirms the fix.
1336
1384
 
1385
+ ## FIX REJECTED BY GUARDRAIL
1386
+
1387
+ Returned when a fix-acceptance guardrail signal fails (see `@~/.claude/gsd-core/references/debugger-fix-acceptance.md`). Do **not** mark the session resolved.
1388
+
1389
+ **Debug Session:** .planning/debug/{slug}.md
1390
+ **Failing signal:** {signal 1–5 name}
1391
+ **Evidence:** {why the signal failed — e.g. "mutant at fix site survived", "deletion-only diff with no RCA justification", "bug did not return on revert"}
1392
+
1393
+ The session-manager continuation surfaces this and offers revise / accept-as-debt / abandon.
1394
+
1337
1395
  ## INVESTIGATION INCONCLUSIVE
1338
1396
 
1339
1397
  ```markdown
@@ -144,6 +144,10 @@ At execution decision points, apply structured reasoning:
144
144
 
145
145
  For each task:
146
146
 
147
+ 0. **Precondition check (before any other task work):** If the task carries a `<precondition>` element, evaluate that single prose line first — it names a runnable/checkable fact the task assumes (env var set, prior-phase artifact present, server responding to `/health`, `user_setup` step done). Verify with **read-only checks only** — file existence, env var presence (no value output), idempotent `GET /health`-style pings. Do NOT run commands with side effects (writes, network POSTs, secret emission) as the check; if a side-effecting check seems required, halt and surface via checkpoint instead.
148
+ - **Met OR absent:** continue with no visible change to execution flow. The precondition is a no-op for the rest of the task loop.
149
+ - **Unmet:** STOP — return a `checkpoint:human-verify` (use `checkpoint_return_format`) with `**Blocked by:** Precondition not met: <precondition text>`. Do NOT partial-commit the task. Unmet preconditions are NEVER auto-approved, even under `AUTO_CFG=true` — a missing prerequisite is not a verification step a human can rubber-stamp; it is a fact the executor cannot establish on its own. The human either satisfies the precondition (sets the env var, completes the `user_setup` step, regenerates the artifact) or reruns `/gsd:plan-phase` to restructure.
150
+
147
151
  1. **If `type="auto"`:**
148
152
  - Check for `tdd="true"` → follow TDD execution flow
149
153
  - Execute task, apply deviation rules as needed
@@ -152,11 +156,17 @@ For each task:
152
156
  - Commit (see task_commit_protocol)
153
157
  - Track completion + commit hash for Summary
154
158
 
155
- 2. **If `type="checkpoint:*"`:**
159
+ 2. **If `type="tracer"`:** (the leading thin end-to-end slice — production-quality, never a throwaway)
160
+ - Execute and commit exactly like `type="auto"` (real implementation, real `<verify>`, atomic commit).
161
+ - **Then run the tracer feedback gate BEFORE any expansion task** — an early integration checkpoint on the proven slice:
162
+ - **Autonomous run (auto mode active — `AUTO_CHAIN` or `AUTO_CFG` is `"true"`, per `<auto_mode_detection>`):** re-run the tracer's `<verify>` end-to-end. If it **fails**, HALT and surface it (deviation Rule 1) — do NOT proceed to expansion tasks. Pouring more layers onto a broken foundation is exactly the failure this gate prevents. If it passes, log `⚡ Tracer verified end-to-end — expanding` and continue.
163
+ - **Interactive run (auto mode not active):** immediately after committing the tracer, STOP and return a `checkpoint:human-verify` for the tracer's `<verify>` (the working slice) using checkpoint_return_format, before any expansion task.
164
+
165
+ 3. **If `type="checkpoint:*"`:**
156
166
  - STOP immediately — return structured checkpoint message
157
167
  - A fresh agent will be spawned to continue
158
168
 
159
- 3. After all tasks: run overall verification, confirm success criteria, document deviations
169
+ 4. After all tasks: run overall verification, confirm success criteria, document deviations
160
170
  </step>
161
171
 
162
172
  </execution_flow>
@@ -310,12 +320,14 @@ For full automation-first patterns, server lifecycle, CLI handling:
310
320
 
311
321
  **Quick reference:** Users NEVER run CLI commands. Users ONLY visit URLs, click UI, evaluate visuals, provide secrets. Claude does all automation.
312
322
 
323
+ **Tracer feedback gate:** a `type="tracer"` task is followed by an early integration checkpoint on the proven slice (see `<execution_flow>` → `execute_tasks`) — in autonomous runs a failing tracer `<verify>` HALTS before any expansion task; in interactive runs the executor emits a `checkpoint:human-verify` for the tracer immediately after committing it.
324
+
313
325
  ---
314
326
 
315
327
  **Auto-mode checkpoint behavior** (when `AUTO_CFG` is `"true"`):
316
328
 
317
329
  - **checkpoint:human-verify** → Auto-approve **except package-legitimacy checkpoints**. If checkpoint has `gate="blocking-human"` OR its purpose indicates package legitimacy verification (`what-built` mentions `Package verification required before install` or `Package install failed — human verification required`), do **not** auto-approve. STOP and return checkpoint_return_format for explicit human confirmation.
318
- - **checkpoint:decision** → Auto-select first option (planners front-load the recommended choice). Log `⚡ Auto-selected: [option name]`. Continue to next task.
330
+ - **checkpoint:decision** → If checkpoint has `gate="blocking-human"`, do **not** auto-select — STOP and return checkpoint_return_format for an explicit human decision (a `blocking-human` decision exists because its default answer would be wrong to assume). Otherwise auto-select first option (planners front-load the recommended choice), log `⚡ Auto-selected: [option name]`, continue to next task.
319
331
  - **checkpoint:human-action** → STOP normally. Auth gates cannot be automated — return structured checkpoint message using checkpoint_return_format.
320
332
 
321
333
  **Standard checkpoint behavior** (when `AUTO_CFG` is not `"true"`):
@@ -340,6 +352,7 @@ When hitting checkpoint or auth gate, return this structure:
340
352
  ## CHECKPOINT REACHED
341
353
 
342
354
  **Type:** [human-verify | decision | human-action]
355
+ **Gate:** [blocking | blocking-human] — copy the task's `gate` attribute verbatim so the orchestrator's carve-out sees it
343
356
  **Plan:** {phase}-{plan}
344
357
  **Progress:** {completed}/{total} tasks complete
345
358
 
@@ -659,6 +672,21 @@ Or: "None - plan executed exactly as written."
659
672
 
660
673
  If any stubs exist, add a `## Known Stubs` section to the SUMMARY listing each stub with its file, line, and reason. These are tracked for the verifier to catch. Do NOT mark a plan as complete if stubs exist that prevent the plan's goal from being achieved — either wire the data or document in the plan why the stub is intentional and which future plan will resolve it.
661
674
 
675
+ **Broken-windows ledger (issue #1950).** For each stub, skipped test, or unrun `<verify>` recorded above, ALSO append it to the cross-phase defect register at `.planning/WINDOWS.md`. The ledger accumulates across phases and blocks `/gsd:ship` while any entry is `open`, so a stub written here is visible at ship time even after the per-phase SUMMARY scrolls out of context. Append one entry per defect:
676
+
677
+ ```bash
678
+ gsd_run windows append \
679
+ --kind stub \
680
+ --phase "${PHASE_NUMBER}" \
681
+ --file "<path-relative-to-repo-root>" \
682
+ --line "<line-number-or-omit>" \
683
+ --description "<one-line description, same wording as the Known Stubs row>"
684
+ ```
685
+
686
+ Use `--kind skipped-test` for a `t.skip(...)` / `test.todo(...)` you left behind, `--kind unrun-verify` for a `<verify>` you could not run, or `--kind deviation` for a documented plan deviation. The full kind vocabulary: `stub | todo | fixme | skipped-test | lint-warning | unmet-truth | unrun-verify | deviation`.
687
+
688
+ The ledger is **optional**: if `gsd_run windows append` returns `windows_ledger_missing` or `windows_ok` without writing, continue without error — population is best-effort and never blocks execution. Recording here is what makes the defect visible to the ship gate later; forgetting to record is the failure mode this ledger exists to prevent.
689
+
662
690
  **Threat surface scan:** Before writing the SUMMARY, check if any files created/modified introduce security-relevant surface NOT in the plan's `<threat_model>` — new network endpoints, auth paths, file access patterns, or schema changes at trust boundaries. If found, add:
663
691
 
664
692
  ```markdown
@@ -66,7 +66,7 @@ The orchestrator provides user decisions in `<user_decisions>` tags from `/gsd:d
66
66
  **Self-check before returning:** For each plan, verify:
67
67
  - [ ] Every locked decision (D-01, D-02, etc.) has a task implementing it
68
68
  - [ ] Task actions reference the decision ID they implement (e.g., "per D-03")
69
- (The decision-coverage gate `check.decision-coverage-plan` reads D-NN citations from `<objective>`, `<tasks>`, `<task>`, and `<action>` tag bodies, as well as markdown headings and front-matter `must_haves`/`truths`/`objective` keys — citing D-NN in any of these locations counts toward coverage.)
69
+ (The decision-coverage gate `check.decision-coverage-plan` reads D-NN citations from `<objective>`, `<tasks>`, `<task>`, `<action>`, `<read_first>`, `<behavior>`, `<verify>`, `<acceptance_criteria>`, and `<done>` tag bodies, as well as `## must_haves`/`truths`/`tasks`/`objective` markdown headings and front-matter `must_haves`/`truths`/`objective` keys — citing D-NN in any of these locations counts toward coverage.)
70
70
  - [ ] No task implements a deferred idea
71
71
  - [ ] Discretion areas are handled reasonably
72
72
 
@@ -193,23 +193,21 @@ Every task has four required fields:
193
193
  **Grep gate hygiene:** `grep -c` counts comments, so header prose can be self-invalidating. Use `grep -v '^#' | grep -c token`. Bare `== 0` gates on unfiltered files are forbidden.
194
194
 
195
195
  <comment_text_discipline>
196
- **Comment-text discipline (HARD GATE, #429):** A literal an acceptance criterion negative-greps for (`grep -c 'LIT' file == 0`) must NOT appear verbatim in any `<action>` body — JSDoc samples, head-comment references, or "what NOT to do" snippets echo into the written file and trip the executor's commit-time gate. `validate_plan` (`verify.plan-structure`) fails plan creation on violation. Rephrase the literal by concept, or — when it must legitimately appear — add an allowlist marker on its own line:
197
-
198
- `<!-- planner-discipline-allow: LIT -->`
199
-
200
- Full rules + worked examples: @gsd-core/references/planner-antipatterns.md ("Comment-Text Discipline").
196
+ **Comment-text discipline (HARD GATE, #429):** A literal an acceptance criterion negative-greps for must NOT appear verbatim in any `<action>` body. Full rules + `<!-- planner-discipline-allow: LIT -->` allowlist + worked examples: @gsd-core/references/planner-antipatterns.md ("Comment-Text Discipline").
201
197
  </comment_text_discipline>
202
198
 
203
199
  <region_scoped_negative_gate>
204
- **Region-scoped negative gates (WARN, #968):** Region-scope a file-wide negative grep when a sibling task needs that construct elsewhere in the same file; `validate_plan` WARNS. See: @gsd-core/references/planner-antipatterns.md ("Region-Scoped Negative Gates").
205
-
206
- **Verify-gate hygiene (#1478/#1479):** See @gsd-core/references/planner-antipatterns.md.
200
+ **Region-scoped negative gates (WARN, #968)** and **Verify-gate hygiene (#1478/#1479):** @gsd-core/references/planner-antipatterns.md.
207
201
  </region_scoped_negative_gate>
208
202
 
209
203
  **<done>:** Acceptance criteria - measurable state of completion.
210
204
  - Good: "Valid credentials return 200 + JWT cookie, invalid credentials return 401"
211
205
  - Bad: "Authentication is complete"
212
206
 
207
+ **<precondition>** (optional, one prose line): a runnable/checkable fact the task assumes that plan ordering does not guarantee — external setup (`user_setup`), a prior-phase artifact, or an env var. The executor asserts it before running the task and halts on unmet. Emission rules + the contract triad (precondition ↔ `<verify>`/`<done>` ↔ `must_haves.truths`): @~/.claude/gsd-core/references/planner-preconditions.md.
208
+
209
+ **<reversibility>** (optional): `rating="reversible|costly|one-way"` + one-line rationale for a decision this task implements. `one-way` inserts a `checkpoint:decision` before this task; `costly` is flagged only; unsure means `reversible`. Rules: @~/.claude/gsd-core/references/planner-reversibility.md
210
+
213
211
  See @~/.claude/gsd-core/references/planner-guidance.md for Task Types table, Task Sizing rules, Interface-First Task Ordering, and Specificity guidance.
214
212
 
215
213
  ## TDD Detection
@@ -250,34 +248,33 @@ Exceptions where `tdd="true"` is not needed: `type="checkpoint:*"` tasks, config
250
248
 
251
249
  `workflow.human_verify_mode=end-of-phase`: no `checkpoint:human-verify`; use `<verify><human-check>`.
252
250
 
253
- ## MVP Mode Detection
254
-
255
- **When `MVP_MODE` is enabled (passed by the plan-phase orchestrator):** Decompose tasks as **vertical feature slices**, not horizontal layers. Required reading: Read `~/.claude/gsd-core/references/planner-mvp-mode.md` for the vertical-slice rules (lazy — only on MVP runs).
251
+ ## Tracer-First Decomposition (default)
256
252
 
257
- **Core rule:** After each task completes, a real user can do something they could not do after the previous task. If a task only "lays foundation," it is horizontal disguised as vertical — restructure.
253
+ **Every phase plan LEADS with one `type="tracer"` task** — the thinnest path that touches every layer the phase will modify, wired end-to-end, carrying a real runnable `<verify>`. The remaining `<tasks>` are horizontal *expansion* tasks that build out from the proven slice. This is the default for **every** phase; it is not gated behind a flag. Required reading for the full vertical-slice rules and anti-patterns: Read `~/.claude/gsd-core/references/planner-mvp-mode.md`.
258
254
 
259
- **Plan structure under MVP_MODE:**
255
+ **Why tracer-first:** proving the architecture end-to-end on the agent's best early-context tokens catches an architectural dead-end after one commit instead of after ten already-committed layers.
260
256
 
261
- 1. Frame the phase goal as a user story at the top of `PLAN.md`. The user story is sourced from the `**Goal:**` line in ROADMAP.md (set by `mvp-phase`). Emit it with bolded keywords:
257
+ **A tracer is production-quality, not a prototype.** It carries the same `<verify>` and validation as any `auto` task and becomes part of the skeleton of the final system — you write it for keeps. Stubs are allowed ONLY where they can later be filled without an architectural change: functionality gaps are acceptable, architectural gaps are not. (Glossary: `tracer bullet` vs `prototype` in `CONTEXT.md` — GSD ships tracers, never prototypes.)
262
258
 
263
- ```
264
- ## Phase Goal
259
+ **Tracer task shape:**
265
260
 
266
- **As a** [user role], **I want to** [capability], **so that** [outcome].
267
- ```
261
+ ```xml
262
+ <task type="tracer">
263
+ <name>End-to-end "[capability]" — one path only</name>
264
+ <files>[one file per layer the phase touches]</files>
265
+ <action>Wire ONE entry point through every layer to the far end of the stack. No other call sites, no batching. Real error handling on the single path.</action>
266
+ <verify>[a real, runnable END-TO-END check of the one path — not a per-layer unit test]</verify>
267
+ <done>The single happy path works end-to-end and is committed.</done>
268
+ </task>
269
+ ```
268
270
 
269
- Format rules (Read `~/.claude/gsd-core/references/user-story-template.md`):
270
- - All three slots required. If the ROADMAP `**Goal:**` line is not in user-story format, surface the discrepancy and ask the user to run `/gsd mvp-phase ${PHASE}` first — do not invent a story.
271
- - Bold the three keywords (`**As a**`, `**I want to**`, `**so that**`) when emitting to PLAN.md. The ROADMAP form does not use bolded keywords; the PLAN form does.
272
- 2. First task: failing end-to-end test for the happy path.
273
- 3. Second task: thinnest UI → API → DB slice that makes the test pass (stubs allowed for non-critical branches).
274
- 4. Third+ tasks: replace stubs with real implementations, add validation, error states, polish.
271
+ **Core rule (expansion tasks):** after each task a real user can do something they could not before. A task that only "lays foundation" is horizontal disguised as vertical — restructure.
275
272
 
276
- **Mode is all-or-nothing per phase** (PRD decision Q1). Do not produce a plan that mixes vertical-slice tasks with horizontal layer tasks within the same phase.
273
+ **`--no-tracer` (`TRACER_MODE=false`):** opt out of tracer-first and decompose into horizontal layers (the legacy default). Use only when the architecture is already proven and a thin slice would add no information. Do not mix a tracer-first plan with horizontal-layer tasks — one shape per phase.
277
274
 
278
- **Walking Skeleton mode** (`WALKING_SKELETON=true`, set by orchestrator for Phase 1 + new project under `--mvp`): The first deliverable is a Walking Skeleton — the thinnest possible end-to-end stack. In addition to `PLAN.md`, produce `SKELETON.md` using the template at `~/.claude/gsd-core/references/skeleton-template.md` (Read it now). `SKELETON.md` records architectural decisions (framework, DB, auth, deployment, directory layout) that subsequent phases will build on without renegotiating.
275
+ **MVP enrichment (`MVP_MODE=true`):** layered on top of the tracer-first ordering above (MVP no longer *turns on* vertical slices — that is now the default). It adds: (1) frame the phase goal as a user story at the top of `PLAN.md`, sourced from the ROADMAP `**Goal:**` line, bolding `**As a**` / `**I want to**` / `**so that**` (Read `~/.claude/gsd-core/references/user-story-template.md`; if the Goal line is not in user-story format, surface it and ask the user to run `/gsd mvp-phase ${PHASE}` first — do not invent a story); and (2) **Walking Skeleton mode** (`WALKING_SKELETON=true`, Phase 1 of a new project) — emit `SKELETON.md` from `~/.claude/gsd-core/references/skeleton-template.md` alongside `PLAN.md`. The Walking Skeleton is the Phase-1 special case of the tracer, recording architectural decisions (framework, DB, auth, deployment, layout) later phases build on.
279
276
 
280
- **Compatibility with TDD detection:** When both `MVP_MODE=true` and `workflow.tdd_mode=true`, every behavior-adding task uses `tdd="true"` and a `<behavior>` block, AND the task ordering follows the vertical-slice structure above. The first task is always a failing end-to-end test.
277
+ **TDD composition (`workflow.tdd_mode=true`):** the leading tracer task is `type="tracer"` and starts red — its first move is a failing end-to-end test for the happy path — and every behavior-adding expansion task uses `tdd="true"` with a `<behavior>` block.
281
278
 
282
279
  See @~/.claude/gsd-core/references/planner-guidance.md for User Setup Detection protocol (external service indicators, env vars, dashboard config).
283
280
 
@@ -542,15 +539,9 @@ Do NOT use for: Deploying (use CLI), creating webhooks (use API), creating datab
542
539
 
543
540
  When Claude tries CLI/API and gets auth error → creates checkpoint → user authenticates → Claude retries. Auth gates are created dynamically, NOT pre-planned.
544
541
 
545
- ## Writing Guidelines
546
-
547
- **DO:** Automate everything before checkpoint, be specific ("Visit https://myapp.vercel.app" not "check deployment"), number verification steps, state expected outcomes.
542
+ ## Writing Guidelines, Anti-Patterns, and Extended Examples
548
543
 
549
- **DON'T:** Ask human to do work Claude can automate, mix multiple verifications, place checkpoints before automation completes.
550
-
551
- ## Anti-Patterns and Extended Examples
552
-
553
- For checkpoint anti-patterns, specificity comparison tables, context section anti-patterns, and scope reduction patterns:
544
+ For checkpoint writing guidelines (DO/DON'T), anti-patterns, specificity comparison tables, context section anti-patterns, and scope reduction patterns:
554
545
  @~/.claude/gsd-core/references/planner-antipatterns.md
555
546
 
556
547
  </checkpoints>
@@ -768,6 +759,8 @@ At decision points during plan creation, apply structured reasoning:
768
759
 
769
760
  Decompose phase into tasks. **Think dependencies first, not sequence.**
770
761
 
762
+ **Lead with the tracer.** Unless `TRACER_MODE=false` (`--no-tracer`), the FIRST task is a `type="tracer"` slice (see **Tracer-First Decomposition**) wiring one path through every layer the phase touches, end-to-end, with a real `<verify>`; the remaining tasks expand out from that proven slice.
763
+
771
764
  For each task:
772
765
  1. What does it NEED? (files, types, APIs that must exist)
773
766
  2. What does it CREATE? (files, types, APIs others might need)