@opengsd/gsd-core 1.7.0 → 1.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (261) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/.claude-plugin/plugin.json +1 -1
  3. package/.opencode/plugins/gsd-core.js +45 -1
  4. package/README.md +2 -0
  5. package/agents/gsd-code-fixer.md +1 -1
  6. package/agents/gsd-codebase-mapper.md +1 -1
  7. package/agents/gsd-debug-session-manager.md +78 -4
  8. package/agents/gsd-debugger.md +87 -29
  9. package/agents/gsd-executor.md +49 -9
  10. package/agents/gsd-intel-updater.md +3 -3
  11. package/agents/gsd-phase-researcher.md +4 -2
  12. package/agents/gsd-plan-checker.md +20 -0
  13. package/agents/gsd-planner.md +44 -59
  14. package/agents/gsd-project-researcher.md +2 -2
  15. package/agents/gsd-ui-auditor.md +0 -40
  16. package/agents/gsd-verifier.md +2 -2
  17. package/bin/install.js +1338 -135
  18. package/commands/gsd/ai-integration-phase.md +1 -1
  19. package/commands/gsd/mempalace-capture.md +9 -5
  20. package/commands/gsd/new-milestone.md +1 -1
  21. package/commands/gsd/plan-phase.md +5 -3
  22. package/commands/gsd/plan-review-convergence.md +7 -2
  23. package/gsd-core/bin/gsd-tools.cjs +2690 -2472
  24. package/gsd-core/bin/lib/adapter-imperative.cjs +8 -1
  25. package/gsd-core/bin/lib/agent-command-router.cjs +20 -5
  26. package/gsd-core/bin/lib/api-coverage.cjs +360 -53
  27. package/gsd-core/bin/lib/audit.cjs +8 -8
  28. package/gsd-core/bin/lib/broken-windows.cjs +716 -0
  29. package/gsd-core/bin/lib/capability-command-router.cjs +733 -0
  30. package/gsd-core/bin/lib/capability-consent.cjs +40 -1
  31. package/gsd-core/bin/lib/capability-lifecycle.cjs +58 -0
  32. package/gsd-core/bin/lib/capability-loader.cjs +23 -1
  33. package/gsd-core/bin/lib/capability-registry.cjs +1450 -160
  34. package/gsd-core/bin/lib/capability-trust.cjs +468 -33
  35. package/gsd-core/bin/lib/capability-validator.cjs +882 -6
  36. package/gsd-core/bin/lib/capability-writer.cjs +6 -1
  37. package/gsd-core/bin/lib/check-command-router.cjs +140 -27
  38. package/gsd-core/bin/lib/cjs-command-router-adapter.cjs +15 -0
  39. package/gsd-core/bin/lib/claude-orchestration-command-router.cjs +209 -31
  40. package/gsd-core/bin/lib/claude-orchestration.cjs +203 -25
  41. package/gsd-core/bin/lib/command-aliases.cjs +14 -0
  42. package/gsd-core/bin/lib/commands.cjs +326 -21
  43. package/gsd-core/bin/lib/config-loader.cjs +214 -30
  44. package/gsd-core/bin/lib/config.cjs +158 -22
  45. package/gsd-core/bin/lib/core-utils.cjs +6 -1
  46. package/gsd-core/bin/lib/decisions.cjs +32 -8
  47. package/gsd-core/bin/lib/docs.cjs +6 -0
  48. package/gsd-core/bin/lib/estimate-cli.cjs +336 -0
  49. package/gsd-core/bin/lib/external-descriptor-trust.cjs +14 -2
  50. package/gsd-core/bin/lib/frontmatter.cjs +125 -15
  51. package/gsd-core/bin/lib/gap-checker.cjs +17 -2
  52. package/gsd-core/bin/lib/host-integration.cjs +215 -8
  53. package/gsd-core/bin/lib/init.cjs +155 -66
  54. package/gsd-core/bin/lib/install-engine.cjs +299 -23
  55. package/gsd-core/bin/lib/install-profiles.cjs +239 -1
  56. package/gsd-core/bin/lib/installer-migrations/005-opencode-baseline-commands-dir.cjs +146 -0
  57. package/gsd-core/bin/lib/installer-migrations/006-pi-extension-cjs-to-js.cjs +91 -0
  58. package/gsd-core/bin/lib/installer-migrations.cjs +44 -5
  59. package/gsd-core/bin/lib/markdown-sectionizer.cjs +107 -0
  60. package/gsd-core/bin/lib/milestone.cjs +248 -14
  61. package/gsd-core/bin/lib/model-catalog.cjs +69 -4
  62. package/gsd-core/bin/lib/model-resolver.cjs +189 -7
  63. package/gsd-core/bin/lib/observability/logger.cjs +7 -2
  64. package/gsd-core/bin/lib/onboard-projection.cjs +11 -8
  65. package/gsd-core/bin/lib/phase-command-router.cjs +10 -1
  66. package/gsd-core/bin/lib/phase-estimation.cjs +398 -0
  67. package/gsd-core/bin/lib/phase-id.cjs +304 -9
  68. package/gsd-core/bin/lib/phase.cjs +258 -17
  69. package/gsd-core/bin/lib/plan-drift-guard.cjs +1 -1
  70. package/gsd-core/bin/lib/plan-scan.cjs +70 -2
  71. package/gsd-core/bin/lib/planning-workspace.cjs +9 -2
  72. package/gsd-core/bin/lib/profile-output.cjs +34 -8
  73. package/gsd-core/bin/lib/review-lane-descriptor.cjs +927 -0
  74. package/gsd-core/bin/lib/review-lane-invocation.cjs +348 -0
  75. package/gsd-core/bin/lib/review-lane-runner.cjs +594 -0
  76. package/gsd-core/bin/lib/review-reviewer-selection.cjs +114 -32
  77. package/gsd-core/bin/lib/roadmap-parser.cjs +61 -10
  78. package/gsd-core/bin/lib/roadmap.cjs +23 -7
  79. package/gsd-core/bin/lib/runtime-artifact-conversion.cjs +38 -5
  80. package/gsd-core/bin/lib/runtime-artifact-layout.cjs +23 -9
  81. package/gsd-core/bin/lib/runtime-hooks-surface.cjs +156 -0
  82. package/gsd-core/bin/lib/runtime-name-policy.cjs +15 -2
  83. package/gsd-core/bin/lib/smart-entry.cjs +70 -5
  84. package/gsd-core/bin/lib/state-document.cjs +171 -24
  85. package/gsd-core/bin/lib/state-transition.cjs +50 -11
  86. package/gsd-core/bin/lib/state.cjs +206 -32
  87. package/gsd-core/bin/lib/surface.cjs +51 -9
  88. package/gsd-core/bin/lib/uat-predicate.cjs +6 -4
  89. package/gsd-core/bin/lib/uat.cjs +428 -11
  90. package/gsd-core/bin/lib/ui-consideration-probe.cjs +2 -2
  91. package/gsd-core/bin/lib/unusable-input.cjs +216 -0
  92. package/gsd-core/bin/lib/validate.cjs +44 -8
  93. package/gsd-core/bin/lib/verification.cjs +163 -31
  94. package/gsd-core/bin/lib/verify.cjs +348 -42
  95. package/gsd-core/bin/lib/worktree-safety.cjs +360 -15
  96. package/gsd-core/bin/shared/config-defaults.manifest.json +1 -0
  97. package/gsd-core/bin/shared/config-schema.manifest.json +4 -15
  98. package/gsd-core/bin/shared/model-catalog.json +5 -0
  99. package/gsd-core/bin/shared/runtime-aliases.manifest.json +5 -0
  100. package/gsd-core/references/api-coverage.md +37 -7
  101. package/gsd-core/references/checkpoints.md +1 -1
  102. package/gsd-core/references/common-bug-patterns.md +13 -0
  103. package/gsd-core/references/context-budget.md +40 -0
  104. package/gsd-core/references/debugger-bug-taxonomy.md +111 -0
  105. package/gsd-core/references/debugger-fix-acceptance.md +157 -0
  106. package/gsd-core/references/debugger-philosophy.md +1 -0
  107. package/gsd-core/references/debugger-prevention.md +98 -0
  108. package/gsd-core/references/debugger-rca-branching.md +98 -0
  109. package/gsd-core/references/debugger-repro-hardening.md +130 -0
  110. package/gsd-core/references/debugger-sbfl.md +110 -0
  111. package/gsd-core/references/debugger-semantic-recall.md +81 -0
  112. package/gsd-core/references/execute-phase-quota-recovery.md +55 -0
  113. package/gsd-core/references/execute-phase-requirement-revert.md +8 -0
  114. package/gsd-core/references/execute-phase-response-language.md +7 -0
  115. package/gsd-core/references/gate-prompts.md +6 -3
  116. package/gsd-core/references/model-profile-resolution.md +64 -13
  117. package/gsd-core/references/offer-next.md +88 -0
  118. package/gsd-core/references/planner-antipatterns.md +6 -0
  119. package/gsd-core/references/planner-mvp-mode.md +12 -13
  120. package/gsd-core/references/planner-preconditions.md +156 -0
  121. package/gsd-core/references/planner-reversibility.md +132 -0
  122. package/gsd-core/references/planning-config.md +2 -1
  123. package/gsd-core/references/reviewer-instances.md +28 -19
  124. package/gsd-core/references/runtime-aware-dispatch.md +42 -0
  125. package/gsd-core/references/skeleton-template.md +1 -1
  126. package/gsd-core/references/thinking-models-planning.md +3 -1
  127. package/gsd-core/references/ui-consideration-probe.md +2 -2
  128. package/gsd-core/references/worktree-branch-check.md +4 -4
  129. package/gsd-core/templates/DEBUG.md +5 -3
  130. package/gsd-core/templates/summary-minimal.md +4 -0
  131. package/gsd-core/templates/summary-standard.md +4 -0
  132. package/gsd-core/templates/summary.md +7 -0
  133. package/gsd-core/workflows/add-phase.md +2 -0
  134. package/gsd-core/workflows/add-tests.md +3 -1
  135. package/gsd-core/workflows/add-todo.md +32 -1
  136. package/gsd-core/workflows/ai-integration-phase.md +8 -6
  137. package/gsd-core/workflows/audit-fix.md +6 -2
  138. package/gsd-core/workflows/audit-milestone.md +8 -0
  139. package/gsd-core/workflows/autonomous.md +19 -15
  140. package/gsd-core/workflows/check-todos.md +5 -3
  141. package/gsd-core/workflows/cleanup.md +7 -1
  142. package/gsd-core/workflows/code-review-fix.md +14 -6
  143. package/gsd-core/workflows/code-review.md +93 -24
  144. package/gsd-core/workflows/complete-milestone.md +3 -0
  145. package/gsd-core/workflows/debug.md +35 -7
  146. package/gsd-core/workflows/diagnose-issues.md +5 -1
  147. package/gsd-core/workflows/discovery-phase.md +7 -0
  148. package/gsd-core/workflows/discuss-phase/modes/advisor.md +2 -4
  149. package/gsd-core/workflows/discuss-phase/modes/auto.md +0 -6
  150. package/gsd-core/workflows/discuss-phase/templates/context.md +16 -2
  151. package/gsd-core/workflows/discuss-phase-assumptions.md +18 -9
  152. package/gsd-core/workflows/discuss-phase.md +2 -2
  153. package/gsd-core/workflows/do.md +7 -1
  154. package/gsd-core/workflows/docs-update.md +9 -0
  155. package/gsd-core/workflows/eval-review.md +4 -1
  156. package/gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md +4 -0
  157. package/gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md +160 -0
  158. package/gsd-core/workflows/execute-phase/steps/post-merge-gate.md +4 -4
  159. package/gsd-core/workflows/execute-phase/steps/regression-gate.md +2 -2
  160. package/gsd-core/workflows/execute-phase.md +110 -149
  161. package/gsd-core/workflows/execute-plan.md +20 -8
  162. package/gsd-core/workflows/explore.md +4 -0
  163. package/gsd-core/workflows/extract-learnings.md +21 -0
  164. package/gsd-core/workflows/graduation.md +3 -0
  165. package/gsd-core/workflows/health.md +7 -1
  166. package/gsd-core/workflows/help/modes/full.md +9 -5
  167. package/gsd-core/workflows/import.md +11 -2
  168. package/gsd-core/workflows/inbox.md +7 -0
  169. package/gsd-core/workflows/ingest-docs.md +19 -10
  170. package/gsd-core/workflows/manager.md +3 -1
  171. package/gsd-core/workflows/map-codebase.md +17 -10
  172. package/gsd-core/workflows/mvp-phase.md +3 -0
  173. package/gsd-core/workflows/new-milestone.md +79 -23
  174. package/gsd-core/workflows/new-project.md +28 -19
  175. package/gsd-core/workflows/new-workspace.md +3 -1
  176. package/gsd-core/workflows/next.md +5 -2
  177. package/gsd-core/workflows/onboard.md +3 -0
  178. package/gsd-core/workflows/plan-phase.md +56 -51
  179. package/gsd-core/workflows/plan-review-convergence.md +61 -12
  180. package/gsd-core/workflows/plant-seed.md +3 -0
  181. package/gsd-core/workflows/profile-user.md +7 -1
  182. package/gsd-core/workflows/progress.md +31 -3
  183. package/gsd-core/workflows/quick.md +33 -10
  184. package/gsd-core/workflows/remove-workspace.md +3 -0
  185. package/gsd-core/workflows/review.md +172 -585
  186. package/gsd-core/workflows/scan.md +10 -2
  187. package/gsd-core/workflows/secure-phase.md +13 -2
  188. package/gsd-core/workflows/settings-integrations.md +3 -0
  189. package/gsd-core/workflows/settings.md +3 -0
  190. package/gsd-core/workflows/ship.md +88 -11
  191. package/gsd-core/workflows/sketch.md +3 -0
  192. package/gsd-core/workflows/smart-entry.md +4 -1
  193. package/gsd-core/workflows/spike.md +7 -1
  194. package/gsd-core/workflows/ui-phase.md +11 -2
  195. package/gsd-core/workflows/ui-review.md +11 -1
  196. package/gsd-core/workflows/undo.md +7 -0
  197. package/gsd-core/workflows/update.md +106 -5
  198. package/gsd-core/workflows/validate-phase.md +13 -2
  199. package/gsd-core/workflows/verify-phase.md +2 -2
  200. package/gsd-core/workflows/verify-work.md +15 -4
  201. package/hooks/dist/gsd-context-monitor.js +27 -9
  202. package/hooks/dist/gsd-cursor-session-start.js +6 -2
  203. package/hooks/dist/gsd-cursor-stop.js +6 -2
  204. package/hooks/dist/gsd-cursor-subagent-start.js +6 -2
  205. package/hooks/dist/gsd-graphify-update.sh +9 -0
  206. package/hooks/dist/gsd-phase-boundary.sh +14 -2
  207. package/hooks/dist/gsd-prompt-guard.js +101 -2
  208. package/hooks/dist/gsd-read-guard.js +100 -2
  209. package/hooks/dist/gsd-read-injection-scanner.js +109 -2
  210. package/hooks/dist/gsd-statusline.js +97 -9
  211. package/hooks/dist/gsd-workflow-guard.js +110 -6
  212. package/hooks/dist/gsd-worktree-path-guard.js +132 -8
  213. package/hooks/dist/lib/cursor-workspace.js +74 -0
  214. package/hooks/gsd-context-monitor.js +27 -9
  215. package/hooks/gsd-cursor-session-start.js +6 -2
  216. package/hooks/gsd-cursor-stop.js +6 -2
  217. package/hooks/gsd-cursor-subagent-start.js +6 -2
  218. package/hooks/gsd-graphify-update.sh +9 -0
  219. package/hooks/gsd-phase-boundary.sh +14 -2
  220. package/hooks/gsd-prompt-guard.js +101 -2
  221. package/hooks/gsd-read-guard.js +100 -2
  222. package/hooks/gsd-read-injection-scanner.js +109 -2
  223. package/hooks/gsd-statusline.js +97 -9
  224. package/hooks/gsd-workflow-guard.js +110 -6
  225. package/hooks/gsd-worktree-path-guard.js +132 -8
  226. package/hooks/lib/cursor-workspace.js +74 -0
  227. package/package.json +10 -8
  228. package/pi/gsd.cjs +34 -3
  229. package/scripts/changeset/lint.cjs +1 -0
  230. package/scripts/changeset/parse.cjs +26 -0
  231. package/scripts/check-coverage-gate.cjs +51 -0
  232. package/scripts/check-glossary-refs.cjs +244 -0
  233. package/scripts/ci-rebase-check.cjs +48 -4
  234. package/scripts/ci-test-scope.cjs +67 -17
  235. package/scripts/gen-adr-index.cjs +528 -0
  236. package/scripts/gen-capability-matrix.cjs +26 -2
  237. package/scripts/gen-capability-registry.cjs +132 -34
  238. package/scripts/gen-emitted-baseline.cjs +145 -0
  239. package/scripts/gen-test-timings.cjs +201 -0
  240. package/scripts/lint-compiled-artifact-sync.cjs +146 -0
  241. package/scripts/lint-emitted-drift-ack.cjs +149 -0
  242. package/scripts/lint-fix-has-regression-test.cjs +131 -0
  243. package/scripts/lint-portable-timeout.cjs +140 -0
  244. package/scripts/lint-resolution-provenance.cjs +9 -0
  245. package/scripts/lint-test-file-count.allowlist.json +1 -0
  246. package/scripts/mutation-matrix.cjs +4 -0
  247. package/scripts/prompt-injection-scan.sh +6 -0
  248. package/scripts/registry-schema.cjs +57 -8
  249. package/scripts/release-notes/conventional-title.cjs +19 -1
  250. package/scripts/release-notes/format-github-release-notes.cjs +7 -3
  251. package/scripts/release-tarball-smoke.cjs +18 -11
  252. package/scripts/run-tests.cjs +420 -58
  253. package/scripts/workflow-size.cjs +16 -8
  254. package/skills/gsd-ai-integration-phase/SKILL.md +1 -1
  255. package/skills/gsd-mempalace-capture/SKILL.md +9 -5
  256. package/skills/gsd-new-milestone/SKILL.md +1 -1
  257. package/skills/gsd-plan-phase/SKILL.md +5 -3
  258. package/skills/gsd-plan-review-convergence/SKILL.md +7 -2
  259. package/vscode/package.json +1 -1
  260. package/scripts/gen-golden-install-parity-zcode.cjs +0 -77
  261. package/scripts/update-size-baseline.cjs +0 -68
@@ -0,0 +1,110 @@
1
+ # Spectrum-Based Fault Localization (SBFL) Pre-filter
2
+
3
+ Loaded by `gsd-debugger` via `@-include`. A deterministic ranked "where to look
4
+ first" list derived from existing test pass/fail coverage, computed before LLM
5
+ reasoning over an unranked search space.
6
+
7
+ ## Why this exists
8
+
9
+ `investigation_loop` Phase 1 searches the codebase and reads files to seed
10
+ hypotheses via (expensive, non-deterministic) LLM reasoning over an unranked
11
+ space. When a runnable test suite with per-test coverage exists, there is a
12
+ cheap, deterministic signal being left on the table: which code is
13
+ disproportionately executed by **failing** vs **passing** tests. SBFL turns
14
+ that coverage spectrum into a suspiciousness ranking. The agent currently has no
15
+ fault-localization step at all.
16
+
17
+ ## When to run it (Phase 1.25 — after initial evidence, before hypothesis formation)
18
+
19
+ Run only when ALL hold:
20
+ - A runnable test suite exists for the failing area.
21
+ - At least one failing test AND at least one passing test exist (a spectrum
22
+ requires both).
23
+ - Per-test coverage is available (which tests executed which code
24
+ element — function, line, or branch).
25
+
26
+ If any precondition fails, **skip** this step with a logged note (see
27
+ Degradation) and proceed with Phase 1's normal evidence gathering unchanged.
28
+
29
+ ## The Ochiai formula
30
+
31
+ For each code element `s` executed by the test suite:
32
+
33
+ ```
34
+ ochiai(s) = failed(s) / sqrt(totalFailed × (failed(s) + passed(s)))
35
+ ```
36
+
37
+ where:
38
+ - `failed(s)` = number of **failing** tests that executed `s`
39
+ - `passed(s)` = number of **passing** tests that executed `s`
40
+ - `totalFailed` = total number of failing tests in the suite
41
+
42
+ The score is in `[0, 1]`. An element executed by every failing test and no
43
+ passing test scores `1.0` (maximum suspiciousness). An element touched only by
44
+ passing tests scores `0`. Ochiai is empirically stronger than Tarantula across
45
+ the SBFL literature; **Tarantula** is the documented fallback formula
46
+ (`tarantula(s) = (failed(s)/totalFailed) / ((failed(s)/totalFailed) + (passed(s)/totalPassed))`)
47
+ if a comparison or secondary signal is wanted.
48
+
49
+ ## Output — top-N shortlist seeded into the hypothesis space
50
+
51
+ Rank all executed elements by descending Ochiai score and take the **top-N**
52
+ (N is judgment — 5–10 is typical; bounded by what narrows the search without
53
+ flooding it). Append each top-N element to the debug file's **Evidence**
54
+ section as a first-class hypothesis candidate:
55
+
56
+ ```
57
+ - timestamp: <now>
58
+ checked: SBFL Ochiai ranking (Phase 1.25)
59
+ found: top-N suspicious locations —
60
+ 1. path/to/file.cts:LINE (score 0.89) — <symbol>
61
+ 2. path/to/other.cts:LINE (score 0.77) — <symbol>
62
+ ...
63
+ implication: investigate these before forming broader hypotheses
64
+ ```
65
+
66
+ This narrows the search space by orders of magnitude before any LLM tokens are
67
+ spent forming hypotheses. Each top-N entry becomes a candidate for Phase 2
68
+ hypothesis formation, ranked ahead of un-evidenced guesses.
69
+
70
+ ## Degradation (Gall's Law — optional step, degrades onto the working agent)
71
+
72
+ This step is **purely additive**. Every miss degrades to today's behavior; a
73
+ skipped step is logged, never a silent pass (Kernighan — the debugger stays
74
+ auditable).
75
+
76
+ | Condition | Behavior |
77
+ |---|---|
78
+ | No test suite for the failing area | **skip** with a logged note in Evidence ("SBFL skipped: no test suite"); Phase 1 proceeds unchanged |
79
+ | Test suite but no failing tests | **skip** with a logged note ("SBFL skipped: no failing tests — no spectrum"); Phase 1 proceeds unchanged |
80
+ | Test suite but no passing tests | **skip** with a logged note ("SBFL skipped: no spectrum — no passing tests"); Phase 1 proceeds unchanged (Tarantula would divide by `totalPassed=0`; do not run it) |
81
+ | Test suite but no per-test coverage | **skip** with a logged note ("SBFL skipped: no per-test coverage available"); Phase 1 proceeds unchanged |
82
+ | Coverage exists but is coarse (file-level, not line/function) | run anyway, rank at the available granularity, and note the granularity in Evidence |
83
+
84
+ ## Bug-class gating (pairs with Phase 2B bug-taxonomy routing)
85
+
86
+ SBFL is the go-to pre-filter for **deterministic failures (Bohrbugs)** — bugs
87
+ that reproduce reliably. It is explicitly **not trusted** on
88
+ **Heisenbug/Mandelbug** spectra (timing, races, environment-dependent failures):
89
+ a flaky suite pollutes the spectrum (a "failing" test that sometimes passes
90
+ poisons `failed(s)`), so the ranking becomes noise. When the failure is
91
+ non-deterministic (Phase 2B classifies it), **skip SBFL** and route to
92
+ record-replay or stability-stress instead. If SBFL has already run before
93
+ classification and the class later resolves to Heisenbug/Mandelbug, mark the
94
+ prior SBFL Evidence entry as revoked (do not delete it — Kernighan
95
+ auditability) and note why in Evidence.
96
+
97
+ ## Scope boundary (Zawinski's Law)
98
+
99
+ This is a deterministic pre-filter that reuses the project's existing
100
+ test/coverage runner — it adds **no new coverage framework** and no new
101
+ subsystem. It narrows the LLM's search space; it does not replace hypothesis
102
+ formation, fix-and-verify, or the knowledge base. Coverage acquisition is the
103
+ agent's adaptive job (use whatever coverage the project produces); the formula
104
+ above is the canonical ranking.
105
+
106
+ **Bound the coverage run** (CLAUDE.md gauntlet — unbounded subprocess): a
107
+ coverage run is often 2–3× slower than a plain test run due to instrumentation,
108
+ so cap it (60s for npm-tier suites; scale with suite size) and **degrade to
109
+ skip with a logged note on timeout** — never let coverage acquisition hang the
110
+ debug session.
@@ -0,0 +1,81 @@
1
+ # Semantic Knowledge-Base Recall via MemPalace
2
+
3
+ Loaded by `gsd-debugger` via `@-include` from the `knowledge_base_protocol`
4
+ Matching Logic. Replaces keyword-overlap matching with **semantic recall** so a
5
+ prior session that resolved "requests hang under load" surfaces for a new "API
6
+ times out when many users connect" — same root cause, no shared keywords.
7
+
8
+ ## Why this exists
9
+
10
+ The knowledge base's self-noted limitation was explicit: *"Matching is keyword
11
+ overlap, not semantic similarity."* Keyword overlap only fires on lexical
12
+ coincidence — the highest-value recalls (same root cause, different wording)
13
+ are exactly the ones it misses, and its value decays as the corpus grows.
14
+
15
+ ## The approach — reuse MemPalace, add no new infrastructure
16
+
17
+ Layer semantic recall on top of the existing knowledge base by **reusing
18
+ MemPalace** (the semantic-memory capability already in this environment) —
19
+ **without adding new embedding or vector infrastructure** (Choose Boring /
20
+ Zawinski: spend no new "innovation token" on a bespoke vector store the
21
+ debugger would own).
22
+
23
+ `.planning/debug/knowledge-base.md` remains the **durable plain-text source of
24
+ truth**; semantic recall is an additive layer over it, not a replacement.
25
+
26
+ ## Write — index resolved sessions at archive
27
+
28
+ At `archive_session`, after appending the entry to `knowledge-base.md` (the KB
29
+ append + commit MUST succeed first — `knowledge-base.md` is the durable source
30
+ of truth; skip indexing on KB-write failure), **index the resolved session into
31
+ MemPalace**.
32
+
33
+ **Index the agent-authored `Resolution` summary — `root_cause(s)` + `fix` + the
34
+ Prevention `recurrence_guard` — NOT the raw user-supplied `Symptoms`.** The
35
+ Resolution is the post-investigation, agent-synthesized signal; indexing it
36
+ (rather than raw symptoms) excludes attacker-controlled prose from the
37
+ cross-session index and reduces secret/PII leakage. Even so, **redact
38
+ secret-shaped values** (API keys, bearer tokens, JWTs, passwords, credentials)
39
+ from the summary before indexing — a bug report's error string can echo a
40
+ secret, and MemPalace is a cross-session, cross-project store.
41
+
42
+ ## Invocation (the agent has no MCP tools — use the CLI)
43
+
44
+ The `gsd-debugger` `tools:` frontmatter grants no MCP tools, so query and index
45
+ via the **Bash CLI** (the headless/autonomous path): `mempalace search
46
+ "<symptoms>" --wing <wing>` to recall, and the matching index command on
47
+ archive. If an `mempalace_search(query, wing)` MCP tool is registered in the
48
+ runtime, prefer it. **Resolve the wing** from `config.mempalace.wing` → else the
49
+ project's `project_code` → else the project directory name (the same precedence
50
+ every other MemPalace integration uses).
51
+
52
+ ## Read — query MemPalace at Phase 0
53
+
54
+ At Phase 0, **query MemPalace semantically with the current symptoms** and
55
+ surface the **top-k meaning-similar prior resolutions** as candidate
56
+ hypotheses. Each surfaced candidate flows into Evidence exactly as a
57
+ keyword-match candidate would — a hypothesis to test first, not a confirmed
58
+ diagnosis.
59
+
60
+ This catches the **same-root-cause / different-wording** case: a prior
61
+ "requests hang under load" resolution surfaces for "API times out when many
62
+ users connect" even though no keywords overlap.
63
+
64
+ ## Graceful degradation — MemPalace absent
65
+
66
+ When MemPalace is unavailable (not installed, not configured, or the query
67
+ errors), **fall back to keyword-overlap matching** against
68
+ `knowledge-base.md`: extract nouns, error substrings, and **identifiers**
69
+ (function/variable names — often the highest-signal token) from
70
+ `Symptoms.errors` and `Symptoms.actual`, and scan each entry's `Error patterns`
71
+ field for **2+ token overlap (case-insensitive)**. The fallback is logged
72
+ (Kernighan — never a silent skip), and `knowledge-base.md` continues to be
73
+ written regardless, so no session is lost to a missing palace.
74
+
75
+ ## Scope boundary (Zawinski's Law)
76
+
77
+ An additive recall layer over the existing knowledge base, reusing an existing
78
+ semantic-memory capability. Not a new command, not a vector database, not an
79
+ embedding pipeline the debugger owns. Where MemPalace is absent the debugger
80
+ behaves exactly as it did before this layer — keyword matching against the
81
+ plain-text knowledge base.
@@ -0,0 +1,55 @@
1
+ **Step 7.1 detail — `class == "quota-exceeded"` recovery.**
2
+
3
+ Do not offer "retry now". Run the step-5 spot-check first; if SUMMARY.md is missing but
4
+ commits exist, route to safe-resume (`state.verify-against-disk`) instead of an immediate
5
+ redispatch.
6
+
7
+ **7.1a — provider escalation (#2296, opt-in).** A heavier tier on the same throttled
8
+ provider is still throttled, so when `dynamic_routing.provider_escalation` is configured
9
+ GSD swaps PROVIDER rather than waiting for a reset. `QUOTA_ATTEMPT` starts at 1 on the
10
+ first quota failure of this phase and increments on each subsequent one.
11
+
12
+ ```bash
13
+ ESC_JSON=$(gsd_run query resolve-execution gsd-executor --attempt "${QUOTA_ATTEMPT:-1}" --failure-class quota-exceeded)
14
+ ESCALATED=$(echo "$ESC_JSON" | jq -r '.escalation.escalated')
15
+ EXHAUSTED=$(echo "$ESC_JSON" | jq -r '.escalation.exhausted')
16
+ ESC_FROM=$(echo "$ESC_JSON" | jq -r '.escalation.from')
17
+ ESC_TO=$(echo "$ESC_JSON" | jq -r '.escalation.to')
18
+ ESC_TRIED=$(echo "$ESC_JSON" | jq -r '.escalation.attempted | join(" -> ")')
19
+ ```
20
+
21
+ - **`ESCALATED == "true"`** — log the switch, honor the provider's own backoff
22
+ (`sleep "$RETRY_AFTER"` when `RETRY_AFTER` is set), then re-dispatch the failed plan with
23
+ `executor_model` overridden to `$ESC_TO` and `QUOTA_ATTEMPT` incremented. Do not prompt —
24
+ this is the configured, opt-in path.
25
+
26
+ ```text
27
+ ⚡ Provider quota hit — escalating model: {ESC_FROM} → {ESC_TO}
28
+ Runtime sentinel: {SENTINEL}
29
+ {RETRY_HINT}
30
+ ```
31
+
32
+ - **`EXHAUSTED == "true"`** — the ladder is spent. Fail loudly naming every model tried,
33
+ then fall through to the manual options below. Never silently retry the last one.
34
+
35
+ ```text
36
+ ⛔ Provider escalation exhausted — tried: {ESC_TRIED}
37
+ ```
38
+
39
+ - **`ESCALATED == "false"` and not exhausted** — escalation is not configured for this
40
+ project; use the manual path below. This is the default.
41
+
42
+ **7.1b — manual recovery (default when escalation is not configured).**
43
+
44
+ ```text
45
+ ⚠ Plan {plan_id} terminated by provider quota / rate limit
46
+ Runtime sentinel: {SENTINEL}
47
+ {RETRY_HINT}
48
+ Partial commits on worktree branch: {N}
49
+ SUMMARY.md present: {yes|no}
50
+ 1. Wait for quota reset, then resume (recommended)
51
+ 2. Switch to a different runtime / model and resume
52
+ 3. Abort phase and report partial state
53
+ ```
54
+
55
+ Re-run `/gsd:execute-phase` after the quota resets for Option 1.
@@ -0,0 +1,8 @@
1
+ **Revert this phase's own requirement IDs out of `Complete` before rendering the gap report (#2388).** A shared requirement ID can already read `Complete` at this point (its first-declaring plan finished before this verification ran) — a `gaps_found` verdict must not leave that premature `Complete` sitting in REQUIREMENTS.md. Scoped strictly to `PHASE_REQ_IDS` (this phase's own citations from `init.execute-phase`), so another phase's `Complete` row is never touched:
2
+
3
+ ```bash
4
+ if [ -n "${PHASE_REQ_IDS}" ]; then
5
+ gsd_run query requirements.revert-phase ${PHASE_REQ_IDS} >/dev/null 2>&1 || true
6
+ gsd_run query commit "docs(phase-{X}): revert premature Complete requirements after gaps found" --files .planning/REQUIREMENTS.md >/dev/null 2>&1 || true
7
+ fi
8
+ ```
@@ -0,0 +1,7 @@
1
+ # Execute-Phase Response-Language Directive (#2402)
2
+
3
+ **If `response_language` is set:** User-facing orchestrator output (questions, narration, report-template prose) in `{response_language}`; technical terms, code, file paths, and subagent prompts stay in English. Pass `response_language: {value}` into every spawned subagent prompt so any user-facing output they produce stays in the configured language.
4
+
5
+ The literal report templates embedded in this workflow (`## Execution Plan`, `## Phase {X}: {Name} Execution Complete`, `## ⚠ Phase {X}: {Name} — Gaps Found`, etc.) are a structural source, not literal output to copy verbatim — render their prose translated into `{response_language}` while keeping headings' structural markers, table columns, IDs, commands, and file paths unchanged.
6
+
7
+ This directive was extracted from `workflows/execute-phase.md` to keep that file under the frozen pre-phase-6 byte ceiling (ADR-857 Phase 6 capstone, `tests/fix-2285-claude-orchestration-wiring.test.cjs`). The `@-reference` is eager, so the runtime still loads this content alongside the workflow — the extraction is purely a file-size discipline, not a lazy-load optimization.
@@ -90,11 +90,14 @@ Up to 4 suggested next actions with selection (status, resume workflows).
90
90
  3-option handler for existing CONTEXT.md in discuss workflow.
91
91
  - question: "Phase {N} already has a CONTEXT.md. How should we handle it?"
92
92
  - header: "Context"
93
- - options: Overwrite | Append | Cancel
93
+ - options: Update it | View it | Skip
94
94
 
95
95
  ## Pattern: gray-area-option
96
96
  Dynamic template for presenting gray area choices in discuss workflow.
97
97
  - question: "{Gray area title}"
98
98
  - header: "Decision"
99
- - options: {Option 1} | {Option 2} | Let Claude decide
100
- - Note: Options generated at runtime. Always include "Let Claude decide" as last option.
99
+ - options: {Option 1} | {Option 2}
100
+ - Note: Options generated at runtime. Present real choices only — do NOT include a
101
+ "skip" or "you decide" option (the user ran discuss-phase to decide). An "Other"
102
+ free-text escape hatch may be offered per modes/default.md, but never a
103
+ "Let Claude decide" cop-out.
@@ -1,38 +1,89 @@
1
1
  # Model Profile Resolution
2
2
 
3
- Resolve model profile once at the start of orchestration, then use it for all Task spawns.
3
+ Resolve each agent's model through `gsd-tools`, then pass it to the `Agent()` spawn — or
4
+ omit the parameter entirely when nothing resolved.
4
5
 
5
6
  ## Resolution Pattern
6
7
 
8
+ Prefer the field your workflow's own `init.*` payload already emits (every orchestrator
9
+ init that spawns subagents carries one — `mapper_model`, `planner_model`,
10
+ `executor_model`, `doc_writer_model`, …). Declare it in the workflow's parse line so the
11
+ binding is stated, not implied:
12
+
13
+ ```
14
+ Parse JSON for: `mapper_model`, …
15
+ ```
16
+
17
+ When the agent type is only known at runtime — a capability hook naming its own agent, for
18
+ example — resolve it directly instead:
19
+
7
20
  ```bash
8
- MODEL_PROFILE=$(cat .planning/config.json 2>/dev/null | grep -o '"model_profile"[[:space:]]*:[[:space:]]*"[^"]*"' | grep -o '"[^"]*"$' | tr -d '"' || echo "balanced")
21
+ AGENT_MODEL=$(gsd_run query resolve-model "<agent-type>" --raw)
9
22
  ```
10
23
 
11
- Default: `balanced` if not set or config missing.
24
+ Both surfaces return the same thing: a model string for the active runtime, or the **empty
25
+ string** when nothing resolved.
12
26
 
13
27
  ## Lookup Table
14
28
 
15
29
  @~/.claude/gsd-core/references/model-profiles.md
16
30
 
17
- Look up the agent in the table for the resolved profile. Pass the model parameter to Task calls:
31
+ ## Passing the model to a spawn
18
32
 
19
33
  ```
20
- Task(
34
+ Agent(
21
35
  prompt="...",
22
36
  subagent_type="gsd-planner",
23
- model="{resolved_model}" # "inherit", "sonnet", or "haiku"
37
+ model="{planner_model}"
38
+ )
39
+ ```
40
+
41
+ Substitute the field **this workflow bound** — never a generic placeholder name.
42
+
43
+ **#2517 — omit, do not emit an empty model.** When the resolved value is `"inherit"` or
44
+ empty, **omit the `model=` parameter entirely**:
45
+
46
+ ```
47
+ Agent(
48
+ prompt="...",
49
+ subagent_type="gsd-planner"
24
50
  )
25
51
  ```
26
52
 
27
- **Note:** Opus-tier agents resolve to `"inherit"` (not `"opus"`). This causes the agent to use the parent session's model, avoiding conflicts with organization policies that may block specific opus versions.
53
+ No model parameter is passed at all — `planner_model` resolved to `"inherit"` or empty, and
54
+ omitting it inherits the orchestrator's model. Passing either value through as an argument
55
+ instead 404s on non-Claude runtimes.
56
+
57
+ This is not cosmetic, and both values really occur. `model_profile: "inherit"` — and any
58
+ opus-tier agent — resolves to the literal string `"inherit"`. `resolve_model_ids: "omit"`
59
+ resolves to the **empty string** whenever the project sets it explicitly or the active
60
+ runtime has no native tier aliases; an agent type absent from the profile table takes that
61
+ same empty-string path, because it has no tier for the earlier steps to resolve. Emitting
62
+ either value verbatim fails the spawn on every runtime without native tier aliases.
63
+
64
+ **#2684 — substitute a field your workflow actually bound.** `model="{…}"` must name a key
65
+ your own `init.*` payload emits, a shell variable you assigned, or a field on your declared
66
+ parse line. A placeholder that resolves to nothing does not fail loudly — the orchestrator
67
+ silently invents a value, which is the invisible partial application ADR-1411 prohibits.
68
+ `tests/model-omit-when-inherit-guard.test.cjs` enforces both rules.
69
+
70
+ ## Profile semantics
71
+
72
+ **Note:** Opus-tier agents resolve to `"inherit"` (not `"opus"`). This causes the agent to
73
+ use the parent session's model, avoiding conflicts with organization policies that may
74
+ block specific opus versions — and, per the rule above, means the `model=` parameter is
75
+ omitted rather than set.
28
76
 
29
- If `model_profile` is `"adaptive"`, agents resolve to role-based assignments (opus/sonnet/haiku based on agent type).
77
+ If `model_profile` is `"adaptive"`, agents resolve to role-based assignments (opus/sonnet/
78
+ haiku based on agent type).
30
79
 
31
- If `model_profile` is `"inherit"`, all agents resolve to `"inherit"` (useful for OpenCode `/model`).
80
+ If `model_profile` is `"inherit"`, all agents resolve to `"inherit"` (useful for OpenCode
81
+ `/model`).
32
82
 
33
83
  ## Usage
34
84
 
35
- 1. Resolve once at orchestration start
36
- 2. Store the profile value
37
- 3. Look up each agent's model from the table when spawning
38
- 4. Pass model parameter to each Task call (values: `"inherit"`, `"sonnet"`, `"haiku"`)
85
+ 1. Bind the model once — from the `init.*` payload field, or `query resolve-model` when the
86
+ agent type is runtime-determined
87
+ 2. Declare the bound field on the workflow's parse line
88
+ 3. Pass `model="{bound_field}"` on each `Agent()` spawn
89
+ 4. Omit `model=` entirely whenever the bound value is `"inherit"` or empty
@@ -0,0 +1,88 @@
1
+ <!--
2
+ offer-next.md — extracted from execute-phase.md step "offer_next" (#2537).
3
+ Eagerly @-referenced from execute-phase.md so runtime behavior is unchanged; the
4
+ extraction restores byte-budget headroom the frozen ceiling exists to provide.
5
+ -->
6
+
7
+
8
+ **Exception:** If `gaps_found`, the `verify_phase_goal` step already presents the gap-closure path (`/gsd:plan-phase {X} --gaps`). No additional routing needed — skip auto-advance.
9
+
10
+ **No-transition check (spawned by auto-advance chain):**
11
+
12
+ Parse `--no-transition` flag from $ARGUMENTS.
13
+
14
+ **If `--no-transition` flag present:**
15
+
16
+ Execute-phase was spawned by plan-phase's auto-advance. Do NOT run transition.md.
17
+ After verification passes and roadmap is updated, return completion status to parent:
18
+
19
+ ```
20
+ ## PHASE COMPLETE
21
+
22
+ Phase: ${PHASE_NUMBER} - ${PHASE_NAME}
23
+ Plans: ${completed_count}/${total_count}
24
+ Verification: {Passed | Gaps Found}
25
+
26
+ [Include aggregate_results output]
27
+ ```
28
+
29
+ STOP. Do not proceed to auto-advance or transition.
30
+
31
+ **If `--no-transition` flag is NOT present:**
32
+
33
+ **Auto-advance detection:**
34
+
35
+ 1. Parse `--auto` flag from $ARGUMENTS
36
+ 2. Read consolidated auto-mode (`active` = chain flag OR user preference; chain flag already synced in init step):
37
+ ```bash
38
+ AUTO_MODE=$(gsd_run query check auto-mode --pick active 2>/dev/null || echo "false")
39
+ ```
40
+
41
+ **If `--auto` flag present OR `AUTO_MODE` is true (AND verification passed with no gaps):**
42
+
43
+ ```
44
+ ╔══════════════════════════════════════════╗
45
+ ║ AUTO-ADVANCING → TRANSITION ║
46
+ ║ Phase {X} verified, continuing chain ║
47
+ ╚══════════════════════════════════════════╝
48
+ ```
49
+
50
+ Execute the transition workflow inline (do NOT use Agent — orchestrator context is ~10-15%, transition needs phase completion data already in context):
51
+
52
+ Read and follow `~/.claude/gsd-core/workflows/transition.md`, passing through the `--auto` flag so it propagates to the next phase invocation.
53
+
54
+ **If neither `--auto` nor `AUTO_MODE` is true:**
55
+
56
+ **STOP. Do not auto-advance. Do not execute transition. Do not plan next phase. Present options to the user and wait.**
57
+
58
+ **IMPORTANT: There is NO `/gsd-transition` command. Never suggest it. The transition workflow is internal only.**
59
+
60
+ Check whether CONTEXT.md already exists for the next phase:
61
+
62
+ ```bash
63
+ ls .planning/phases/*{next}*/{next}-CONTEXT.md 2>/dev/null || echo "no-context"
64
+ ```
65
+
66
+ If CONTEXT.md does **not** exist for the next phase, present:
67
+
68
+ ```
69
+ ## ✓ Phase {X}: {Name} Complete
70
+
71
+ /gsd:progress ${GSD_WS} — see updated roadmap
72
+ /gsd:discuss-phase {next} ${GSD_WS} — start here: discuss next phase before planning ← recommended
73
+ /gsd:plan-phase {next} ${GSD_WS} — plan next phase (skip discuss)
74
+ /gsd:execute-phase {next} ${GSD_WS} — execute next phase (skip discuss and plan)
75
+ ```
76
+
77
+ If CONTEXT.md **exists** for the next phase, present:
78
+
79
+ ```
80
+ ## ✓ Phase {X}: {Name} Complete
81
+
82
+ /gsd:progress ${GSD_WS} — see updated roadmap
83
+ /gsd:plan-phase {next} ${GSD_WS} — start here: plan next phase (CONTEXT.md already present) ← recommended
84
+ /gsd:discuss-phase {next} ${GSD_WS} — re-discuss next phase
85
+ /gsd:execute-phase {next} ${GSD_WS} — execute next phase (skip planning)
86
+ ```
87
+
88
+ Only suggest the commands listed above. Do not invent or hallucinate command names.
@@ -5,6 +5,12 @@
5
5
 
6
6
  ## Checkpoint Anti-Patterns
7
7
 
8
+ ### Writing guidelines
9
+
10
+ **DO:** Automate everything before checkpoint, be specific ("Visit https://myapp.vercel.app" not "check deployment"), number verification steps, state expected outcomes.
11
+
12
+ **DON'T:** Ask human to do work Claude can automate, mix multiple verifications, place checkpoints before automation completes.
13
+
8
14
  ### Bad — Asking human to automate
9
15
 
10
16
  ```xml
@@ -1,32 +1,31 @@
1
- # Planner — MVP Mode (Vertical Slice Strategy)
1
+ # Planner — Tracer-First Decomposition (Vertical Slices)
2
2
 
3
- > Loaded by `gsd-planner` only when `MVP_MODE=true`. Standard horizontal-layer planning rules continue to apply for all other phases.
3
+ > Loaded by `gsd-planner` for the **default** tracer-first decomposition: every phase LEADS with one thin end-to-end `type="tracer"` slice, then expansion tasks. `--no-tracer` (`TRACER_MODE=false`) restores standard horizontal-layer planning. The MVP enrichment (user-story framing) and Walking Skeleton mode apply *on top* when `MVP_MODE=true` / `WALKING_SKELETON=true`.
4
4
 
5
5
  ## Core Rule
6
6
 
7
7
  **Decompose by feature slice, not by technical layer.** Every task must move the user-facing capability forward. After each task, a real user can click through more of the feature than they could before.
8
8
 
9
- **Forbidden** in MVP mode:
9
+ **Forbidden** under tracer-first:
10
10
  - "Create the database schema" as a standalone task
11
11
  - "Build the API layer" as a standalone task
12
12
  - "Wire up the UI" as a final integration task
13
13
 
14
- **Required** in MVP mode:
15
- - The first non-test task produces a working end-to-end path. Stubs are allowed for non-critical branches; the happy path must be real.
16
- - Each subsequent task either adds a new slice OR refines an existing slice (validation, error states, edge cases).
17
- - The phase goal is framed as a user story: "**As a** [user], **I want to** [do X], **so that** [Y]."
14
+ **Required** under tracer-first:
15
+ - The leading `tracer` task produces a working end-to-end path — production-quality, not a prototype. Stubs are allowed ONLY where they can later be filled without an architectural change; the happy path must be real.
16
+ - Each subsequent expansion task either adds a new slice OR refines an existing slice (validation, error states, edge cases).
17
+ - *(MVP enrichment, `MVP_MODE=true`)* The phase goal is framed as a user story: "**As a** [user], **I want to** [do X], **so that** [Y]."
18
18
 
19
19
  ## Task Order Pattern
20
20
 
21
21
  For a feature `F`:
22
22
 
23
- 1. **Failing end-to-end test** for the happy path of `F`.
24
- 2. **Thinnest viable slice** — UI form → API endpoint → DB read/write — that makes the test pass. Hard-coded values, missing validation, no error states are fine here.
25
- 3. **Real data layer** — replace any stubs from Task 2 with real queries.
26
- 4. **Validation + error states** — invalid input, network failure, empty states.
27
- 5. **Production polish** — loading indicators, edge cases, accessibility checks.
23
+ 1. **Tracer slice** — the thinnest end-to-end path (UI form → API endpoint → DB read/write), wired through every layer with a real runnable `<verify>`. This task is always `type="tracer"`; production-quality, not a prototype; stubs only where later-fillable without an architectural change. Under `--tdd` it *also* starts red — its first move is a failing end-to-end test for the happy path of `F`.
24
+ 2. **Real data layer** — replace any stubs from the tracer with real queries.
25
+ 3. **Validation + error states** — invalid input, network failure, empty states.
26
+ 4. **Production polish** — loading indicators, edge cases, accessibility checks.
28
27
 
29
- Tasks 3-5 are not always all needed; gate by the phase's acceptance criteria.
28
+ Tasks 2-4 are not always all needed; gate by the phase's acceptance criteria.
30
29
 
31
30
  ## Walking Skeleton Mode (`WALKING_SKELETON=true`)
32
31