@opengsd/gsd-core 1.7.0 → 1.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (261) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/.claude-plugin/plugin.json +1 -1
  3. package/.opencode/plugins/gsd-core.js +45 -1
  4. package/README.md +2 -0
  5. package/agents/gsd-code-fixer.md +1 -1
  6. package/agents/gsd-codebase-mapper.md +1 -1
  7. package/agents/gsd-debug-session-manager.md +78 -4
  8. package/agents/gsd-debugger.md +87 -29
  9. package/agents/gsd-executor.md +49 -9
  10. package/agents/gsd-intel-updater.md +3 -3
  11. package/agents/gsd-phase-researcher.md +4 -2
  12. package/agents/gsd-plan-checker.md +20 -0
  13. package/agents/gsd-planner.md +44 -59
  14. package/agents/gsd-project-researcher.md +2 -2
  15. package/agents/gsd-ui-auditor.md +0 -40
  16. package/agents/gsd-verifier.md +2 -2
  17. package/bin/install.js +1338 -135
  18. package/commands/gsd/ai-integration-phase.md +1 -1
  19. package/commands/gsd/mempalace-capture.md +9 -5
  20. package/commands/gsd/new-milestone.md +1 -1
  21. package/commands/gsd/plan-phase.md +5 -3
  22. package/commands/gsd/plan-review-convergence.md +7 -2
  23. package/gsd-core/bin/gsd-tools.cjs +2690 -2472
  24. package/gsd-core/bin/lib/adapter-imperative.cjs +8 -1
  25. package/gsd-core/bin/lib/agent-command-router.cjs +20 -5
  26. package/gsd-core/bin/lib/api-coverage.cjs +360 -53
  27. package/gsd-core/bin/lib/audit.cjs +8 -8
  28. package/gsd-core/bin/lib/broken-windows.cjs +716 -0
  29. package/gsd-core/bin/lib/capability-command-router.cjs +733 -0
  30. package/gsd-core/bin/lib/capability-consent.cjs +40 -1
  31. package/gsd-core/bin/lib/capability-lifecycle.cjs +58 -0
  32. package/gsd-core/bin/lib/capability-loader.cjs +23 -1
  33. package/gsd-core/bin/lib/capability-registry.cjs +1450 -160
  34. package/gsd-core/bin/lib/capability-trust.cjs +468 -33
  35. package/gsd-core/bin/lib/capability-validator.cjs +882 -6
  36. package/gsd-core/bin/lib/capability-writer.cjs +6 -1
  37. package/gsd-core/bin/lib/check-command-router.cjs +140 -27
  38. package/gsd-core/bin/lib/cjs-command-router-adapter.cjs +15 -0
  39. package/gsd-core/bin/lib/claude-orchestration-command-router.cjs +209 -31
  40. package/gsd-core/bin/lib/claude-orchestration.cjs +203 -25
  41. package/gsd-core/bin/lib/command-aliases.cjs +14 -0
  42. package/gsd-core/bin/lib/commands.cjs +326 -21
  43. package/gsd-core/bin/lib/config-loader.cjs +214 -30
  44. package/gsd-core/bin/lib/config.cjs +158 -22
  45. package/gsd-core/bin/lib/core-utils.cjs +6 -1
  46. package/gsd-core/bin/lib/decisions.cjs +32 -8
  47. package/gsd-core/bin/lib/docs.cjs +6 -0
  48. package/gsd-core/bin/lib/estimate-cli.cjs +336 -0
  49. package/gsd-core/bin/lib/external-descriptor-trust.cjs +14 -2
  50. package/gsd-core/bin/lib/frontmatter.cjs +125 -15
  51. package/gsd-core/bin/lib/gap-checker.cjs +17 -2
  52. package/gsd-core/bin/lib/host-integration.cjs +215 -8
  53. package/gsd-core/bin/lib/init.cjs +155 -66
  54. package/gsd-core/bin/lib/install-engine.cjs +299 -23
  55. package/gsd-core/bin/lib/install-profiles.cjs +239 -1
  56. package/gsd-core/bin/lib/installer-migrations/005-opencode-baseline-commands-dir.cjs +146 -0
  57. package/gsd-core/bin/lib/installer-migrations/006-pi-extension-cjs-to-js.cjs +91 -0
  58. package/gsd-core/bin/lib/installer-migrations.cjs +44 -5
  59. package/gsd-core/bin/lib/markdown-sectionizer.cjs +107 -0
  60. package/gsd-core/bin/lib/milestone.cjs +248 -14
  61. package/gsd-core/bin/lib/model-catalog.cjs +69 -4
  62. package/gsd-core/bin/lib/model-resolver.cjs +189 -7
  63. package/gsd-core/bin/lib/observability/logger.cjs +7 -2
  64. package/gsd-core/bin/lib/onboard-projection.cjs +11 -8
  65. package/gsd-core/bin/lib/phase-command-router.cjs +10 -1
  66. package/gsd-core/bin/lib/phase-estimation.cjs +398 -0
  67. package/gsd-core/bin/lib/phase-id.cjs +304 -9
  68. package/gsd-core/bin/lib/phase.cjs +258 -17
  69. package/gsd-core/bin/lib/plan-drift-guard.cjs +1 -1
  70. package/gsd-core/bin/lib/plan-scan.cjs +70 -2
  71. package/gsd-core/bin/lib/planning-workspace.cjs +9 -2
  72. package/gsd-core/bin/lib/profile-output.cjs +34 -8
  73. package/gsd-core/bin/lib/review-lane-descriptor.cjs +927 -0
  74. package/gsd-core/bin/lib/review-lane-invocation.cjs +348 -0
  75. package/gsd-core/bin/lib/review-lane-runner.cjs +594 -0
  76. package/gsd-core/bin/lib/review-reviewer-selection.cjs +114 -32
  77. package/gsd-core/bin/lib/roadmap-parser.cjs +61 -10
  78. package/gsd-core/bin/lib/roadmap.cjs +23 -7
  79. package/gsd-core/bin/lib/runtime-artifact-conversion.cjs +38 -5
  80. package/gsd-core/bin/lib/runtime-artifact-layout.cjs +23 -9
  81. package/gsd-core/bin/lib/runtime-hooks-surface.cjs +156 -0
  82. package/gsd-core/bin/lib/runtime-name-policy.cjs +15 -2
  83. package/gsd-core/bin/lib/smart-entry.cjs +70 -5
  84. package/gsd-core/bin/lib/state-document.cjs +171 -24
  85. package/gsd-core/bin/lib/state-transition.cjs +50 -11
  86. package/gsd-core/bin/lib/state.cjs +206 -32
  87. package/gsd-core/bin/lib/surface.cjs +51 -9
  88. package/gsd-core/bin/lib/uat-predicate.cjs +6 -4
  89. package/gsd-core/bin/lib/uat.cjs +428 -11
  90. package/gsd-core/bin/lib/ui-consideration-probe.cjs +2 -2
  91. package/gsd-core/bin/lib/unusable-input.cjs +216 -0
  92. package/gsd-core/bin/lib/validate.cjs +44 -8
  93. package/gsd-core/bin/lib/verification.cjs +163 -31
  94. package/gsd-core/bin/lib/verify.cjs +348 -42
  95. package/gsd-core/bin/lib/worktree-safety.cjs +360 -15
  96. package/gsd-core/bin/shared/config-defaults.manifest.json +1 -0
  97. package/gsd-core/bin/shared/config-schema.manifest.json +4 -15
  98. package/gsd-core/bin/shared/model-catalog.json +5 -0
  99. package/gsd-core/bin/shared/runtime-aliases.manifest.json +5 -0
  100. package/gsd-core/references/api-coverage.md +37 -7
  101. package/gsd-core/references/checkpoints.md +1 -1
  102. package/gsd-core/references/common-bug-patterns.md +13 -0
  103. package/gsd-core/references/context-budget.md +40 -0
  104. package/gsd-core/references/debugger-bug-taxonomy.md +111 -0
  105. package/gsd-core/references/debugger-fix-acceptance.md +157 -0
  106. package/gsd-core/references/debugger-philosophy.md +1 -0
  107. package/gsd-core/references/debugger-prevention.md +98 -0
  108. package/gsd-core/references/debugger-rca-branching.md +98 -0
  109. package/gsd-core/references/debugger-repro-hardening.md +130 -0
  110. package/gsd-core/references/debugger-sbfl.md +110 -0
  111. package/gsd-core/references/debugger-semantic-recall.md +81 -0
  112. package/gsd-core/references/execute-phase-quota-recovery.md +55 -0
  113. package/gsd-core/references/execute-phase-requirement-revert.md +8 -0
  114. package/gsd-core/references/execute-phase-response-language.md +7 -0
  115. package/gsd-core/references/gate-prompts.md +6 -3
  116. package/gsd-core/references/model-profile-resolution.md +64 -13
  117. package/gsd-core/references/offer-next.md +88 -0
  118. package/gsd-core/references/planner-antipatterns.md +6 -0
  119. package/gsd-core/references/planner-mvp-mode.md +12 -13
  120. package/gsd-core/references/planner-preconditions.md +156 -0
  121. package/gsd-core/references/planner-reversibility.md +132 -0
  122. package/gsd-core/references/planning-config.md +2 -1
  123. package/gsd-core/references/reviewer-instances.md +28 -19
  124. package/gsd-core/references/runtime-aware-dispatch.md +42 -0
  125. package/gsd-core/references/skeleton-template.md +1 -1
  126. package/gsd-core/references/thinking-models-planning.md +3 -1
  127. package/gsd-core/references/ui-consideration-probe.md +2 -2
  128. package/gsd-core/references/worktree-branch-check.md +4 -4
  129. package/gsd-core/templates/DEBUG.md +5 -3
  130. package/gsd-core/templates/summary-minimal.md +4 -0
  131. package/gsd-core/templates/summary-standard.md +4 -0
  132. package/gsd-core/templates/summary.md +7 -0
  133. package/gsd-core/workflows/add-phase.md +2 -0
  134. package/gsd-core/workflows/add-tests.md +3 -1
  135. package/gsd-core/workflows/add-todo.md +32 -1
  136. package/gsd-core/workflows/ai-integration-phase.md +8 -6
  137. package/gsd-core/workflows/audit-fix.md +6 -2
  138. package/gsd-core/workflows/audit-milestone.md +8 -0
  139. package/gsd-core/workflows/autonomous.md +19 -15
  140. package/gsd-core/workflows/check-todos.md +5 -3
  141. package/gsd-core/workflows/cleanup.md +7 -1
  142. package/gsd-core/workflows/code-review-fix.md +14 -6
  143. package/gsd-core/workflows/code-review.md +93 -24
  144. package/gsd-core/workflows/complete-milestone.md +3 -0
  145. package/gsd-core/workflows/debug.md +35 -7
  146. package/gsd-core/workflows/diagnose-issues.md +5 -1
  147. package/gsd-core/workflows/discovery-phase.md +7 -0
  148. package/gsd-core/workflows/discuss-phase/modes/advisor.md +2 -4
  149. package/gsd-core/workflows/discuss-phase/modes/auto.md +0 -6
  150. package/gsd-core/workflows/discuss-phase/templates/context.md +16 -2
  151. package/gsd-core/workflows/discuss-phase-assumptions.md +18 -9
  152. package/gsd-core/workflows/discuss-phase.md +2 -2
  153. package/gsd-core/workflows/do.md +7 -1
  154. package/gsd-core/workflows/docs-update.md +9 -0
  155. package/gsd-core/workflows/eval-review.md +4 -1
  156. package/gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md +4 -0
  157. package/gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md +160 -0
  158. package/gsd-core/workflows/execute-phase/steps/post-merge-gate.md +4 -4
  159. package/gsd-core/workflows/execute-phase/steps/regression-gate.md +2 -2
  160. package/gsd-core/workflows/execute-phase.md +110 -149
  161. package/gsd-core/workflows/execute-plan.md +20 -8
  162. package/gsd-core/workflows/explore.md +4 -0
  163. package/gsd-core/workflows/extract-learnings.md +21 -0
  164. package/gsd-core/workflows/graduation.md +3 -0
  165. package/gsd-core/workflows/health.md +7 -1
  166. package/gsd-core/workflows/help/modes/full.md +9 -5
  167. package/gsd-core/workflows/import.md +11 -2
  168. package/gsd-core/workflows/inbox.md +7 -0
  169. package/gsd-core/workflows/ingest-docs.md +19 -10
  170. package/gsd-core/workflows/manager.md +3 -1
  171. package/gsd-core/workflows/map-codebase.md +17 -10
  172. package/gsd-core/workflows/mvp-phase.md +3 -0
  173. package/gsd-core/workflows/new-milestone.md +79 -23
  174. package/gsd-core/workflows/new-project.md +28 -19
  175. package/gsd-core/workflows/new-workspace.md +3 -1
  176. package/gsd-core/workflows/next.md +5 -2
  177. package/gsd-core/workflows/onboard.md +3 -0
  178. package/gsd-core/workflows/plan-phase.md +56 -51
  179. package/gsd-core/workflows/plan-review-convergence.md +61 -12
  180. package/gsd-core/workflows/plant-seed.md +3 -0
  181. package/gsd-core/workflows/profile-user.md +7 -1
  182. package/gsd-core/workflows/progress.md +31 -3
  183. package/gsd-core/workflows/quick.md +33 -10
  184. package/gsd-core/workflows/remove-workspace.md +3 -0
  185. package/gsd-core/workflows/review.md +172 -585
  186. package/gsd-core/workflows/scan.md +10 -2
  187. package/gsd-core/workflows/secure-phase.md +13 -2
  188. package/gsd-core/workflows/settings-integrations.md +3 -0
  189. package/gsd-core/workflows/settings.md +3 -0
  190. package/gsd-core/workflows/ship.md +88 -11
  191. package/gsd-core/workflows/sketch.md +3 -0
  192. package/gsd-core/workflows/smart-entry.md +4 -1
  193. package/gsd-core/workflows/spike.md +7 -1
  194. package/gsd-core/workflows/ui-phase.md +11 -2
  195. package/gsd-core/workflows/ui-review.md +11 -1
  196. package/gsd-core/workflows/undo.md +7 -0
  197. package/gsd-core/workflows/update.md +106 -5
  198. package/gsd-core/workflows/validate-phase.md +13 -2
  199. package/gsd-core/workflows/verify-phase.md +2 -2
  200. package/gsd-core/workflows/verify-work.md +15 -4
  201. package/hooks/dist/gsd-context-monitor.js +27 -9
  202. package/hooks/dist/gsd-cursor-session-start.js +6 -2
  203. package/hooks/dist/gsd-cursor-stop.js +6 -2
  204. package/hooks/dist/gsd-cursor-subagent-start.js +6 -2
  205. package/hooks/dist/gsd-graphify-update.sh +9 -0
  206. package/hooks/dist/gsd-phase-boundary.sh +14 -2
  207. package/hooks/dist/gsd-prompt-guard.js +101 -2
  208. package/hooks/dist/gsd-read-guard.js +100 -2
  209. package/hooks/dist/gsd-read-injection-scanner.js +109 -2
  210. package/hooks/dist/gsd-statusline.js +97 -9
  211. package/hooks/dist/gsd-workflow-guard.js +110 -6
  212. package/hooks/dist/gsd-worktree-path-guard.js +132 -8
  213. package/hooks/dist/lib/cursor-workspace.js +74 -0
  214. package/hooks/gsd-context-monitor.js +27 -9
  215. package/hooks/gsd-cursor-session-start.js +6 -2
  216. package/hooks/gsd-cursor-stop.js +6 -2
  217. package/hooks/gsd-cursor-subagent-start.js +6 -2
  218. package/hooks/gsd-graphify-update.sh +9 -0
  219. package/hooks/gsd-phase-boundary.sh +14 -2
  220. package/hooks/gsd-prompt-guard.js +101 -2
  221. package/hooks/gsd-read-guard.js +100 -2
  222. package/hooks/gsd-read-injection-scanner.js +109 -2
  223. package/hooks/gsd-statusline.js +97 -9
  224. package/hooks/gsd-workflow-guard.js +110 -6
  225. package/hooks/gsd-worktree-path-guard.js +132 -8
  226. package/hooks/lib/cursor-workspace.js +74 -0
  227. package/package.json +10 -8
  228. package/pi/gsd.cjs +34 -3
  229. package/scripts/changeset/lint.cjs +1 -0
  230. package/scripts/changeset/parse.cjs +26 -0
  231. package/scripts/check-coverage-gate.cjs +51 -0
  232. package/scripts/check-glossary-refs.cjs +244 -0
  233. package/scripts/ci-rebase-check.cjs +48 -4
  234. package/scripts/ci-test-scope.cjs +67 -17
  235. package/scripts/gen-adr-index.cjs +528 -0
  236. package/scripts/gen-capability-matrix.cjs +26 -2
  237. package/scripts/gen-capability-registry.cjs +132 -34
  238. package/scripts/gen-emitted-baseline.cjs +145 -0
  239. package/scripts/gen-test-timings.cjs +201 -0
  240. package/scripts/lint-compiled-artifact-sync.cjs +146 -0
  241. package/scripts/lint-emitted-drift-ack.cjs +149 -0
  242. package/scripts/lint-fix-has-regression-test.cjs +131 -0
  243. package/scripts/lint-portable-timeout.cjs +140 -0
  244. package/scripts/lint-resolution-provenance.cjs +9 -0
  245. package/scripts/lint-test-file-count.allowlist.json +1 -0
  246. package/scripts/mutation-matrix.cjs +4 -0
  247. package/scripts/prompt-injection-scan.sh +6 -0
  248. package/scripts/registry-schema.cjs +57 -8
  249. package/scripts/release-notes/conventional-title.cjs +19 -1
  250. package/scripts/release-notes/format-github-release-notes.cjs +7 -3
  251. package/scripts/release-tarball-smoke.cjs +18 -11
  252. package/scripts/run-tests.cjs +420 -58
  253. package/scripts/workflow-size.cjs +16 -8
  254. package/skills/gsd-ai-integration-phase/SKILL.md +1 -1
  255. package/skills/gsd-mempalace-capture/SKILL.md +9 -5
  256. package/skills/gsd-new-milestone/SKILL.md +1 -1
  257. package/skills/gsd-plan-phase/SKILL.md +5 -3
  258. package/skills/gsd-plan-review-convergence/SKILL.md +7 -2
  259. package/vscode/package.json +1 -1
  260. package/scripts/gen-golden-install-parity-zcode.cjs +0 -77
  261. package/scripts/update-size-baseline.cjs +0 -68
@@ -26,19 +26,46 @@ command -v cursor-agent >/dev/null 2>&1 && echo "cursor:available" || echo "curs
26
26
  command -v agy >/dev/null 2>&1 && echo "antigravity:available" || echo "antigravity:missing"
27
27
 
28
28
  # Check local model servers (OpenAI-compatible HTTP API — no CLI binary required)
29
- OLLAMA_HOST=$(gsd_run query config-get review.ollama_host 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
29
+ OLLAMA_HOST=$(gsd_run query config-get review.ollama_host --raw 2>/dev/null || echo "")
30
30
  if [ -z "$OLLAMA_HOST" ] || [ "$OLLAMA_HOST" = "null" ]; then OLLAMA_HOST="http://localhost:11434"; fi
31
31
  curl -s --max-time 2 "${OLLAMA_HOST}/v1/models" >/dev/null 2>&1 && echo "ollama:available" || echo "ollama:missing"
32
32
 
33
- LM_STUDIO_HOST=$(gsd_run query config-get review.lm_studio_host 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
33
+ LM_STUDIO_HOST=$(gsd_run query config-get review.lm_studio_host --raw 2>/dev/null || echo "")
34
34
  if [ -z "$LM_STUDIO_HOST" ] || [ "$LM_STUDIO_HOST" = "null" ]; then LM_STUDIO_HOST="http://localhost:1234"; fi
35
35
  curl -s --max-time 2 "${LM_STUDIO_HOST}/v1/models" >/dev/null 2>&1 && echo "lm_studio:available" || echo "lm_studio:missing"
36
36
 
37
- LLAMA_CPP_HOST=$(gsd_run query config-get review.llama_cpp_host 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
37
+ LLAMA_CPP_HOST=$(gsd_run query config-get review.llama_cpp_host --raw 2>/dev/null || echo "")
38
38
  if [ -z "$LLAMA_CPP_HOST" ] || [ "$LLAMA_CPP_HOST" = "null" ]; then LLAMA_CPP_HOST="http://localhost:8080"; fi
39
39
  curl -s --max-time 2 "${LLAMA_CPP_HOST}/v1/models" >/dev/null 2>&1 && echo "llama_cpp:available" || echo "llama_cpp:missing"
40
+
41
+ # jq prerequisite (#2589). The config/model/budget lookups in this workflow no
42
+ # longer need jq — they use the native --raw/--pick flags. But the lanes listed
43
+ # under "jq-dependent reviewer lanes" below parse structured JSON that gsd-tools
44
+ # does not emit (OpenAI-compatible /v1/chat/completions responses, opencode's
45
+ # JSONL event stream, agy's conversation cache), so they cannot run without jq.
46
+ # Probe it here rather than letting each lane swallow exit 127 into empty output.
47
+ command -v jq >/dev/null 2>&1 && echo "jq:available" || echo "jq:missing"
48
+ ```
49
+
50
+ **jq-dependent reviewer lanes.** `jq` is a production prerequisite for the
51
+ `ollama`, `lm_studio`, `llama_cpp`, `opencode`, and `antigravity` lanes only. If
52
+ `detect_clis` reports `jq:missing`, treat those five as **undetected** — they
53
+ follow the same "known-but-undetected" path as a missing CLI. Which path that is
54
+ depends on how the lane was selected (see the precedence rules below): reached
55
+ through `review.default_reviewers` or `--all` it is an info note and the lane is
56
+ ignored; named by an explicit flag it is an **error**, because the user asserted
57
+ that lane. Tell the user to install jq:
58
+
59
+ ```
60
+ NOTE: jq is not on PATH — the ollama, lm_studio, llama_cpp, opencode, and
61
+ antigravity reviewer lanes are unavailable. Install jq (https://jqlang.org/download/)
62
+ or select a lane that does not require it (--gemini, --claude, --codex,
63
+ --coderabbit, --qwen, --cursor).
40
64
  ```
41
65
 
66
+ The remaining lanes (`gemini`, `claude`, `codex`, `coderabbit`, `qwen`, `cursor`)
67
+ do not require jq and must stay selectable on a jq-less host.
68
+
42
69
  Parse flags from `$ARGUMENTS`:
43
70
  - `--gemini` → include Gemini
44
71
  - `--claude` → include Claude
@@ -60,10 +87,22 @@ Reviewer-selection precedence:
60
87
  3. `review.default_reviewers`
61
88
  4. No key + no flags → all detected reviewers
62
89
 
90
+ **Explicit reviewer flags are an assertion, not a preference (ADR-2782 D4).** A lane the user
91
+ named on the command line and that cannot run is an **error**, surfaced and non-silent — even
92
+ when other named lanes did run. Do not proceed with a thinner reviewer set and report success:
93
+ `--gemini --qwen` on a host without `qwen` fails, it does not quietly become a Gemini-only
94
+ review. This applies however the lane became unavailable — binary missing, prerequisite `jq`
95
+ absent, or a local server not reachable.
96
+
97
+ The asymmetry is deliberate: *not finding a lane nobody asked for is normal; failing to run a
98
+ lane somebody asked for is an error.* A user who wants "whatever is available" has `--all`; a
99
+ user who wants a preferred set has `review.default_reviewers`. Both stay lenient below.
100
+
63
101
  `review.default_reviewers` behavior:
64
102
  - Value must be a non-empty array of slug strings (configured via `gsd config-set review.default_reviewers '["gemini","codex"]'`)
65
103
  - Unknown slugs warn and are ignored
66
- - Known-but-undetected slugs emit an info note and are ignored
104
+ - Known-but-undetected slugs emit an info note and are ignored — a configured default is a
105
+ preference evaluated across many hosts, so a subset being present is expected, not an error
67
106
  - If all configured reviewers are unavailable, fail with an actionable message
68
107
 
69
108
  **Reviewer instances (#1517, optional):** if `review.reviewer_instances` is configured,
@@ -119,10 +158,20 @@ Collect phase artifacts for the review prompt:
119
158
  ```bash
120
159
  INIT=$(gsd_run query init.phase-op "${PHASE_ARG}")
121
160
  if [[ "$INIT" == @file:* ]]; then INIT=$(cat "${INIT#@file:}"); fi
161
+
162
+ # #2358: ONE run-scoped temp dir (portable via ${TMPDIR:-/tmp}) so overlapping
163
+ # runs never collide or read each other's stale files.
164
+ RUN_DIR=$(mktemp -d "${TMPDIR:-/tmp}/gsd-review-XXXXXX")
165
+ echo "RUN_DIR=$RUN_DIR"
122
166
  ```
123
167
 
124
168
  Read from init: `phase_dir`, `phase_number`, `padded_phase`.
125
169
 
170
+ Capture `RUN_DIR` above (created ONCE) and thread it into every `{run_dir}`
171
+ placeholder and `$RUN_DIR`/`${RUN_DIR}` reference within a bash block. Do NOT
172
+ re-run `mktemp -d` later — every block must resolve to this same directory, or
173
+ `build_prompt`'s writes and `invoke_reviewers`' reads split.
174
+
126
175
  Then read:
127
176
  1. `.planning/PROJECT.md` (first 80 lines — project context)
128
177
  2. Phase section from `.planning/ROADMAP.md`
@@ -189,569 +238,160 @@ Focus on:
189
238
  Output your review in markdown format.
190
239
  ```
191
240
 
192
- Write to a temp file: `/tmp/gsd-review-prompt-{phase}.md`
241
+ Write to a temp file: `{run_dir}/gsd-review-prompt.md`
193
242
 
194
243
  Also write individual section files so the budget tool can re-trim per reviewer:
195
244
 
196
245
  ```bash
246
+ RUN_DIR="{run_dir}" # from gather_context
247
+
197
248
  # Write individual section files for per-reviewer budget trimming
198
249
  # These are always written so reviewers with a budget can invoke prompt-budget
199
- cp "$INSTRUCTIONS_BLOCK_FILE" "/tmp/gsd-review-${PHASE}-instructions.md"
200
- cp "$ROADMAP_SECTION_FILE" "/tmp/gsd-review-${PHASE}-roadmap.md"
250
+ cp "$INSTRUCTIONS_BLOCK_FILE" "${RUN_DIR}/gsd-review-instructions.md"
251
+ cp "$ROADMAP_SECTION_FILE" "${RUN_DIR}/gsd-review-roadmap.md"
201
252
 
202
253
  # Plan files: copy each PLAN.md to a predictable numbered path
203
254
  PLAN_INDEX=0
204
255
  for PLAN_FILE in "${PHASE_DIR}"/*-PLAN.md; do
205
256
  PADDED_IDX=$(printf '%02d' "$PLAN_INDEX")
206
- cp "$PLAN_FILE" "/tmp/gsd-review-${PHASE}-plan-${PADDED_IDX}.md"
257
+ cp "$PLAN_FILE" "${RUN_DIR}/gsd-review-plan-${PADDED_IDX}.md"
207
258
  PLAN_INDEX=$((PLAN_INDEX + 1))
208
259
  done
209
260
 
210
261
  # Optional section files (only if content was included in the combined prompt)
211
262
  if [ -f ".planning/PROJECT.md" ]; then
212
- cp .planning/PROJECT.md "/tmp/gsd-review-${PHASE}-project.md"
263
+ cp .planning/PROJECT.md "${RUN_DIR}/gsd-review-project.md"
213
264
  fi
214
265
  if ls "${PHASE_DIR}/"*"-CONTEXT.md" >/dev/null 2>&1; then
215
- cat "${PHASE_DIR}/"*"-CONTEXT.md" > "/tmp/gsd-review-${PHASE}-context.md"
266
+ cat "${PHASE_DIR}/"*"-CONTEXT.md" > "${RUN_DIR}/gsd-review-context.md"
216
267
  fi
217
268
  if ls "${PHASE_DIR}/"*"-RESEARCH.md" >/dev/null 2>&1; then
218
- cat "${PHASE_DIR}/"*"-RESEARCH.md" > "/tmp/gsd-review-${PHASE}-research.md"
269
+ cat "${PHASE_DIR}/"*"-RESEARCH.md" > "${RUN_DIR}/gsd-review-research.md"
219
270
  fi
220
271
  if [ -f ".planning/REQUIREMENTS.md" ]; then
221
- cp .planning/REQUIREMENTS.md "/tmp/gsd-review-${PHASE}-requirements.md"
272
+ cp .planning/REQUIREMENTS.md "${RUN_DIR}/gsd-review-requirements.md"
222
273
  fi
223
274
  ```
224
275
 
225
- Note: The variable names above (`INSTRUCTIONS_BLOCK_FILE`, `ROADMAP_SECTION_FILE`, `PHASE_DIR`, `PHASE`) reference the variables already established during prompt assembly. In practice the AI implementing this step writes the instruction and roadmap blocks to temp files while assembling the combined prompt, then copies those same temp files to the per-reviewer section paths. If the assembled prompt was built inline (string concatenation rather than file-by-file), write each section to the corresponding path after writing the combined file.
276
+ Note: `INSTRUCTIONS_BLOCK_FILE`, `ROADMAP_SECTION_FILE`, and `PHASE_DIR` come from prompt assembly; `RUN_DIR` is the run-scoped dir from `gather_context` (#2358) re-assigned from `{run_dir}` above. Copy the temp files written during prompt assembly to these section paths (or write each section here if the prompt was built inline).
226
277
  </step>
227
278
 
228
279
  <step name="invoke_reviewers">
229
- Read model preferences from planning config. Null/missing values fall back to CLI defaults.
230
-
231
- ```bash
232
- # JSON scalars from gsd-tools.cjs query; use jq -r to strip JSON string quotes (install jq if missing)
233
- GEMINI_MODEL=$(gsd_run query config-get review.models.gemini 2>/dev/null | jq -r '.' 2>/dev/null || true)
234
- CLAUDE_MODEL=$(gsd_run query config-get review.models.claude 2>/dev/null | jq -r '.' 2>/dev/null || true)
235
- CODEX_MODEL=$(gsd_run query config-get review.models.codex 2>/dev/null | jq -r '.' 2>/dev/null || true)
236
- OPENCODE_MODEL=$(gsd_run query config-get review.models.opencode 2>/dev/null | jq -r '.' 2>/dev/null || true)
237
- # review.models.agy, when set, is passed to agy as --model (escape hatch for a
238
- # pinned model that 404s server-side); otherwise agy uses its persisted default.
239
- AGY_MODEL=$(gsd_run query config-get review.models.agy 2>/dev/null | jq -r '.' 2>/dev/null || true)
240
-
241
- # #1115: `--dangerously-bypass-hook-trust` only exists on codex-cli >= 0.137.0.
242
- # Capability-probe it so older installs don't fail with "unexpected argument"
243
- # (which, with stderr suppressed, produced a silent empty review). The codex
244
- # invocation works fine without the flag on older versions.
245
- if codex exec --help 2>/dev/null | grep -q -- '--dangerously-bypass-hook-trust'; then
246
- CODEX_BYPASS_FLAG="--dangerously-bypass-hook-trust"
247
- else
248
- CODEX_BYPASS_FLAG=""
249
- fi
250
- ```
251
-
252
- **Reviewer instances (#1517, optional):** when instances are configured, each selected
253
- instance invokes its base `cli` with its own `model`/`agent` (opaque argv, never
254
- shell-interpolated). Exact invocation in `gsd-core/references/reviewer-instances.md`.
255
-
256
- For each selected CLI, invoke in sequence (not parallel — avoid rate limits):
257
-
258
- **Timeout guidance (#2194):** prompt-fed source-grounded reviews are slow — measured ~570s for Codex at `xhigh` effort and ~525s for headless Claude on a large plan set. Each of the Gemini / Claude / Codex blocks below MUST be invoked with a high Bash `timeout:` — at least `900000` (15 min), and `1200000` (20 min) for Codex `xhigh` or headless Claude — so a lane is not killed mid-review. On Claude Code, raise the host cap via `BASH_MAX_TIMEOUT_MS` if a review can exceed it. A silent empty output after a long run is a **timeout kill, not a crash** — the Codex `0xc0000142` misdiagnosis persisted because the empty-output branches below cannot distinguish the two; treat an empty result on a slow lane as a dropped lane and re-run with more time rather than diagnosing a CLI/sandbox failure. A cross-AI review that silently drops a lane is blind in one eye.
259
-
260
- **Gemini:**
261
- ```bash
262
- if [ -n "$GEMINI_MODEL" ] && [ "$GEMINI_MODEL" != "null" ]; then
263
- cat /tmp/gsd-review-prompt-{phase}.md | gemini -m "$GEMINI_MODEL" -p - 2>/dev/null > /tmp/gsd-review-gemini-{phase}.md
264
- else
265
- cat /tmp/gsd-review-prompt-{phase}.md | gemini -p - 2>/dev/null > /tmp/gsd-review-gemini-{phase}.md
266
- fi
267
- ```
268
-
269
- **Claude (separate session):**
270
- ```bash
271
- if [ -n "$CLAUDE_MODEL" ] && [ "$CLAUDE_MODEL" != "null" ]; then
272
- cat /tmp/gsd-review-prompt-{phase}.md | claude --model "$CLAUDE_MODEL" -p - 2>/dev/null > /tmp/gsd-review-claude-{phase}.md
273
- else
274
- cat /tmp/gsd-review-prompt-{phase}.md | claude -p - 2>/dev/null > /tmp/gsd-review-claude-{phase}.md
275
- fi
276
- ```
277
-
278
- **Codex:**
279
- ```bash
280
- # $CODEX_BYPASS_FLAG is capability-gated above (#1115). Capture stderr to a .err
281
- # file (not /dev/null) so a non-zero exit — e.g. a flag the installed codex-cli
282
- # does not support — is diagnosable instead of a silent empty review.
283
- # Capture the review via codex's own `-o/--output-last-message <FILE>` (only the
284
- # final agent message) and discard stdout (#1698): on some platforms (Windows)
285
- # codex writes process-teardown output to stdout *after* the final message, and a
286
- # stdout redirect would append that noise to a non-empty file — slipping past the
287
- # `[ ! -s … ]` empty-output guard as a silently polluted review.
288
- if [ -n "$CODEX_MODEL" ] && [ "$CODEX_MODEL" != "null" ]; then
289
- cat /tmp/gsd-review-prompt-{phase}.md | codex exec --ephemeral $CODEX_BYPASS_FLAG --model "$CODEX_MODEL" --skip-git-repo-check -o /tmp/gsd-review-codex-{phase}.md - 2>/tmp/gsd-review-codex-{phase}.err >/dev/null
290
- else
291
- cat /tmp/gsd-review-prompt-{phase}.md | codex exec --ephemeral $CODEX_BYPASS_FLAG --skip-git-repo-check -o /tmp/gsd-review-codex-{phase}.md - 2>/tmp/gsd-review-codex-{phase}.err >/dev/null
292
- fi
293
- if [ ! -s /tmp/gsd-review-codex-{phase}.md ]; then
294
- echo "Codex review failed or returned empty output. stderr:" > /tmp/gsd-review-codex-{phase}.md
295
- cat /tmp/gsd-review-codex-{phase}.err >> /tmp/gsd-review-codex-{phase}.md
296
- fi
297
- ```
298
-
299
- **CodeRabbit:**
300
-
301
- Note: CodeRabbit reviews the current git diff/working tree — it does not accept a prompt or model flag. It may take up to 5 minutes. Use `timeout: 360000` on the Bash tool call. The source-grounding requirement in the build_prompt Review Instructions applies only to the prompt-fed reviewers above; CodeRabbit is a diff-only reviewer and never receives it. Treat its output as a diff observation, not a grounded plan-level verdict.
302
-
303
- ```bash
304
- coderabbit review --prompt-only 2>/dev/null > /tmp/gsd-review-coderabbit-{phase}.md
305
- ```
306
-
307
- **OpenCode (via GitHub Copilot):**
308
-
309
- OpenCode's default `build` agent is an agentic coder, not a prompt→completion API.
310
- On a large review prompt it may run a few `read` tool calls and then end its turn
311
- with **zero output tokens** (`reason:"stop"`, `output:0`), so `--format default`
312
- yields empty stdout and the second reviewer is silently lost (#1936). Invoke with
313
- `--format json` and reconstruct the review from the assistant `text` parts; if the
314
- agent emitted none, surface the stop `reason`, output-token count, and captured
315
- stderr so the failure is diagnosable instead of a generic empty stub. Runs are also
316
- nondeterministic in length, so bound this Bash tool call with a wall-clock timeout —
317
- set `timeout: 660000` on the call (same mechanism the CodeRabbit block documents).
318
- That bound is hard: if it fires mid-`opencode run` the tool kills the command and
319
- the jq reconstruction below never runs, so the reviewing agent simply proceeds
320
- without an OpenCode result. The completing zero-output case — the actual #1936 bug —
321
- is fully handled below; the timeout only backstops the rarer nondeterministic hang. A
322
- reviewer instance with `"agent": "review"` (see
323
- `gsd-core/references/reviewer-instances.md`) sidesteps the default `build` agent and
324
- is the durable fix when this recurs.
325
-
326
- ```bash
327
- # stderr → sidecar (never /dev/null) so a real error is diagnosable — mirrors the
328
- # Codex block. --format json is the primary invocation (not a fallback): the review
329
- # text lives in assistant `text` parts, which the default formatter drops when the
330
- # agent stops with no final message (#1936).
331
- if [ -n "$OPENCODE_MODEL" ] && [ "$OPENCODE_MODEL" != "null" ]; then
332
- set -- --model "$OPENCODE_MODEL"
333
- else
334
- set --
335
- fi
336
- cat /tmp/gsd-review-prompt-{phase}.md | opencode run "$@" --format json - 2>/tmp/gsd-review-opencode-{phase}.err > /tmp/gsd-review-opencode-{phase}.json
337
- # Reconstruct the review from the assistant text parts. Capture into a variable and
338
- # test its CONTENT (not the output file's size): an empty extraction still prints a
339
- # trailing newline, which would fool a `[ -s file ]` check into skipping the stub.
340
- OPENCODE_REVIEW=$(jq -rs '[.[] | select(.type=="text") | .part.text // empty] | join("\n")' /tmp/gsd-review-opencode-{phase}.json 2>/dev/null)
341
- if [ -n "$OPENCODE_REVIEW" ]; then
342
- printf '%s\n' "$OPENCODE_REVIEW" > /tmp/gsd-review-opencode-{phase}.md
343
- else
344
- # No assistant text (agent emitted no final message, or stdout was not valid JSON events).
345
- {
346
- echo "OpenCode review returned no assistant text (#1936: agent ended its turn with no final message)."
347
- OPENCODE_DIAG=$(jq -rs '[.[] | select(.type=="step_finish")] | last | "stop reason=\(.part.reason // "?"), output tokens=\(.part.tokens.output // "?")"' /tmp/gsd-review-opencode-{phase}.json 2>/dev/null)
348
- [ -n "$OPENCODE_DIAG" ] && echo "Diagnostic: $OPENCODE_DIAG"
349
- echo "stderr:"
350
- cat /tmp/gsd-review-opencode-{phase}.err
351
- } > /tmp/gsd-review-opencode-{phase}.md
352
- fi
353
- ```
354
-
355
- **Qwen Code:**
356
- ```bash
357
- cat /tmp/gsd-review-prompt-{phase}.md | qwen - 2>/dev/null > /tmp/gsd-review-qwen-{phase}.md
358
- if [ ! -s /tmp/gsd-review-qwen-{phase}.md ]; then
359
- echo "Qwen review failed or returned empty output." > /tmp/gsd-review-qwen-{phase}.md
360
- fi
361
- ```
362
-
363
- **Cursor:**
364
- ```bash
365
- # cursor-agent is a SEPARATE binary from the `cursor` IDE launcher; print mode (-p) takes the
366
- # prompt as an ARGUMENT, not stdin. A full review prompt can exceed the OS argument limit, so
367
- # reference the prompt file by path rather than inlining it. Capture stderr so a failure is
368
- # diagnosable instead of a silent empty result.
369
- # #2176: same absolute-root anchor as the Antigravity block — cursor-agent runs
370
- # in the repo cwd, but repo-relative references in the assembled prompt still
371
- # need an explicit root to resolve against. rev-parse (not bare pwd) so the
372
- # anchor is correct even when /gsd:review is invoked from a repo subdirectory.
373
- _CURSOR_ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
374
- CURSOR_PROMPT_ARG="Read the file at /tmp/gsd-review-prompt-{phase}.md in full and carry out the review request it contains. The repository under review is at $_CURSOR_ROOT — resolve every relative file path in the review request against that absolute root. Output only the resulting markdown review. Do not edit any files."
375
- cursor-agent -p --mode ask --trust --output-format text "$CURSOR_PROMPT_ARG" 2>/tmp/gsd-review-cursor-{phase}.err > /tmp/gsd-review-cursor-{phase}.md
376
- if [ ! -s /tmp/gsd-review-cursor-{phase}.md ]; then
377
- echo "Cursor review failed or returned empty output. stderr:" > /tmp/gsd-review-cursor-{phase}.md
378
- cat /tmp/gsd-review-cursor-{phase}.err >> /tmp/gsd-review-cursor-{phase}.md
379
- fi
380
- ```
381
-
382
- **Antigravity CLI:**
383
-
384
- **Maintainer note — why this block has three layers (last updated against agy 1.0.16):**
385
-
386
- `agy -p` (the `--print` non-interactive flag) works correctly on macOS and Linux: it sends the
387
- prompt, receives the model response, and writes it to stdout. On **native Windows** it silently
388
- produces no stdout output despite the API call succeeding — a bug in `text_drip.go`'s non-TTY
389
- flush path, tracked at https://github.com/google-antigravity/antigravity-cli/issues/27466 and
390
- still open as of agy 1.0.2.
391
-
392
- Regardless of platform, `agy` always persists the full exchange to a transcript file on disk.
393
- The transcript fallback (Step 2 below) reads that file directly, giving Windows users full review
394
- coverage without any extra tooling. This pattern was first documented by the community MCP bridge
395
- at https://github.com/SinanTufekci/Claude-Code-Antigravity-CLI-MCP-Server — we inline the same
396
- logic here in pure bash/jq so no additional dependency is required.
397
-
398
- **Stale-response guard (why the pre-flight watermark matters):**
399
- Without a watermark, the fallback would read the last `PLANNER_RESPONSE` entry in the transcript
400
- regardless of when it was written — including entries from a previous invocation in the same
401
- workspace. To prevent that, we record the transcript's line count *before* calling `agy -p`. In
402
- the fallback, we only read lines appended after that count. If no new lines were written (agy
403
- failed before producing a response), `_AGY_RESULT` is empty and Step 3 fires — never stale. If
404
- the conv-id changed (agy started a fresh session), all lines in the new file are new and we use
405
- skip=0.
406
-
407
- **If the upstream stdout bug is fixed** (check the issue above): Step 2 silently becomes
408
- unreachable; stdout is non-empty and Step 1 handles it. No code change needed.
409
-
410
- **If the transcript paths change** in a future `agy` release: Step 2 silently becomes a no-op
411
- and Step 3 fires with a clear error message in REVIEWS.md. No silent corruption. To debug:
412
- - `~/.gemini/antigravity-cli/cache/last_conversations.json` — workspace → conv-id map
413
- - `~/.gemini/antigravity-cli/brain/<id>/.system_generated/logs/transcript.jsonl`
414
- Filter: `source=="MODEL"`, `status=="DONE"`, `type=="PLANNER_RESPONSE"`, take the last match's `content` field.
415
-
416
- Invocation specifics (verified agy 1.0.0, macOS arm64 and Linux amd64):
417
- - `-p` takes the prompt as a **flag value** — `echo X | agy -p` errors with "flag needs an argument: -p"
418
- - `--print-timeout` defaults to 5m, aligning with this workflow's global timeout
419
- - `--model "<name>"` selects the model (available since agy ~1.0.3; `agy models` lists
420
- them). When `review.models.agy` is set it is passed as `--model`; otherwise agy uses
421
- its persisted default (`agy models`).
422
-
423
- ```bash
424
- # Pre-flight: snapshot the transcript watermark before invoking agy.
425
- # Must run BEFORE agy -p — this is what prevents the fallback from reading a stale prior response.
426
- _AGY_WS=$(git rev-parse --show-toplevel 2>/dev/null || pwd)
427
- _AGY_CACHE="$HOME/.gemini/antigravity-cli/cache/last_conversations.json"
428
- _AGY_MARK_CONV=""
429
- _AGY_MARK_LINES=0
430
- if [ -f "$_AGY_CACHE" ]; then
431
- _AGY_MARK_CONV=$(jq -r --arg ws "$_AGY_WS" '
432
- .[$ws] //
433
- (to_entries
434
- | map(select(.key | ascii_downcase == ($ws | ascii_downcase)))
435
- | first | .value) //
436
- empty
437
- ' "$_AGY_CACHE" 2>/dev/null)
438
- if [ -n "$_AGY_MARK_CONV" ] && [ "$_AGY_MARK_CONV" != "null" ]; then
439
- _AGY_MARK_TX="$HOME/.gemini/antigravity-cli/brain/${_AGY_MARK_CONV}/.system_generated/logs/transcript.jsonl"
440
- [ -f "$_AGY_MARK_TX" ] && _AGY_MARK_LINES=$(wc -l < "$_AGY_MARK_TX" | tr -d ' ')
441
- fi
442
- fi
443
-
444
- # Step 1 — primary invocation: stdout works on macOS, Linux, and WSL.
445
- # Three hardening invariants (#2073), all mirroring the Cursor block's discipline:
446
- # * FILE-REFERENCE prompt (not inline `$(cat …)`) — a large review prompt (≈197 KB
447
- # for 6 plans + CONTEXT + RESEARCH + REQUIREMENTS) overflows the exec arg list
448
- # (`bash: agy: Argument list too long`, rc 126), indistinguishable from a model
449
- # failure when stderr is suppressed.
450
- # * EXTERNAL `timeout` wrapper when available (GNU `timeout` / `gtimeout`) —
451
- # `--print-timeout` is agy's native cap but it CANNOT fire before agy creates a
452
- # session; under concurrent heavy runs one process can stall pre-session (no
453
- # `brain/<conv-id>/` dir, alive at 583 s despite `--print-timeout 300s`). The
454
- # external cap bounds wall-clock regardless. Stock macOS lacks `timeout`, so
455
- # the block probes for it and falls back to --print-timeout alone there.
456
- # * `--model` from `review.models.agy` when set — escape hatch for a pinned model
457
- # that 404s server-side (exits 0 with empty stdout AND empty transcript).
458
- # * stdin tied to /dev/null so agy never blocks on a tty.
459
- # A non-zero exit (external timeout = 124, crash, etc.) discards any partial output
460
- # so the Step 2 transcript fallback / Step 3 diagnostic take over.
461
- if [ -n "$AGY_MODEL" ] && [ "$AGY_MODEL" != "null" ]; then
462
- set -- --model "$AGY_MODEL"
463
- else
464
- set --
465
- fi
466
- # #2176: grant the reviewer the repo under review. Without --add-dir, agy's
467
- # permission context never receives the cwd repo — the agent anchors on its own
468
- # ~/.gemini/antigravity-cli/scratch dir and reviews the plan text in isolation
469
- # (the exact failure the Review Instructions forbid). Capability-probed like the
470
- # Codex bypass flag so an older agy without --add-dir still runs; the prompt
471
- # anchor below keeps absolute-path reads possible on that fallback.
472
- if agy --help 2>/dev/null | grep -q -- '--add-dir'; then
473
- set -- "$@" --add-dir "$_AGY_WS"
474
- fi
475
- # #2176: anchor the prompt to the absolute repo root so repo-relative references
476
- # in the assembled review prompt resolve even on the no---add-dir fallback, and
477
- # require an explicit self-report if the reviewer still cannot read the repo.
478
- _AGY_PROMPT="Read the file at /tmp/gsd-review-prompt-{phase}.md in full and carry out the review request it contains. The repository under review is at $_AGY_WS — resolve every relative file path in the review request against that absolute root and verify claims against those files. If you cannot read files under $_AGY_WS, begin your output with the exact line REVIEWED-WITHOUT-REPO-ACCESS before the review. Output only the resulting markdown review. Do not edit any files."
479
- # Capability-probe an external wall-clock killer (GNU coreutils `timeout` or the
480
- # macOS Homebrew `gtimeout`). Stock macOS ships NEITHER — a bare `timeout …` would
481
- # fail with rc 127 ("command not found") and silently lose the reviewer, so fall
482
- # back to agy's native --print-timeout alone in that case. The external cap, when
483
- # available, is set HIGHER than --print-timeout so it only backstops a pre-session
484
- # stall (which --print-timeout cannot bound — #2073 mode 3) and never pre-empts a
485
- # healthy run. Mirrors the probe in scripts/base64-scan.sh.
486
- _AGY_KILLER="$(command -v timeout 2>/dev/null || command -v gtimeout 2>/dev/null || true)"
487
- if [ -n "$_AGY_KILLER" ]; then
488
- "$_AGY_KILLER" 600 agy --print-timeout 540s "$@" -p "$_AGY_PROMPT" </dev/null 2>/dev/null > /tmp/gsd-review-antigravity-{phase}.md
489
- else
490
- agy --print-timeout 540s "$@" -p "$_AGY_PROMPT" </dev/null 2>/dev/null > /tmp/gsd-review-antigravity-{phase}.md
491
- fi
492
- _AGY_RC=$?
493
- if [ "$_AGY_RC" -ne 0 ]; then
494
- : > /tmp/gsd-review-antigravity-{phase}.md
495
- fi
496
-
497
- # Step 2 — transcript fallback: catches Windows agy -p stdout bug (and any future stdout-silent edge cases).
498
- # Reads only lines appended AFTER the pre-flight watermark. If agy failed before writing a new response,
499
- # _AGY_RESULT is empty and Step 3 fires — no stale content can leak through.
500
- # Undocumented paths, verified agy 1.0.0–1.0.2. See maintainer note above if these break.
501
- if [ ! -s /tmp/gsd-review-antigravity-{phase}.md ]; then
502
- if [ -f "$_AGY_CACHE" ]; then
503
- _AGY_CONV=$(jq -r --arg ws "$_AGY_WS" '
504
- .[$ws] //
505
- (to_entries
506
- | map(select(.key | ascii_downcase == ($ws | ascii_downcase)))
507
- | first | .value) //
508
- empty
509
- ' "$_AGY_CACHE" 2>/dev/null)
510
- if [ -n "$_AGY_CONV" ] && [ "$_AGY_CONV" != "null" ]; then
511
- _AGY_TX="$HOME/.gemini/antigravity-cli/brain/${_AGY_CONV}/.system_generated/logs/transcript.jsonl"
512
- if [ -f "$_AGY_TX" ]; then
513
- # If conv-id changed, agy started a new session — all lines are new, skip 0.
514
- # If same conv-id, only read lines beyond the watermark.
515
- [ "$_AGY_CONV" = "$_AGY_MARK_CONV" ] && _AGY_SKIP=$_AGY_MARK_LINES || _AGY_SKIP=0
516
- _AGY_RESULT=$(tail -n +"$((_AGY_SKIP + 1))" "$_AGY_TX" 2>/dev/null | \
517
- jq -r 'select(.source=="MODEL" and .status=="DONE" and .type=="PLANNER_RESPONSE") | .content' \
518
- 2>/dev/null | tail -1)
519
- [ -n "$_AGY_RESULT" ] && echo "$_AGY_RESULT" > /tmp/gsd-review-antigravity-{phase}.md
520
- fi
521
- fi
522
- fi
523
- fi
524
-
525
- # Step 3 — final guard: both approaches yielded nothing (auth error, first-run setup,
526
- # path schema changed, 404'd pinned model, pre-session stall, etc.)
527
- if [ ! -s /tmp/gsd-review-antigravity-{phase}.md ]; then
528
- {
529
- echo "Antigravity review failed or returned empty output."
530
- # #2073 mode 2: a pinned model that 404s exits 0 with empty stdout AND an empty
531
- # transcript — the only evidence is in agy's own log. Surface it instead of a
532
- # bare generic stub so the failure is diagnosable.
533
- _AGY_LOG="$HOME/.gemini/antigravity-cli/cli.log"
534
- if [ -f "$_AGY_LOG" ]; then
535
- _AGY_ERR=$(grep -iE 'agent executor error|NOT_FOUND|Publisher model' "$_AGY_LOG" | tail -3)
536
- if [ -n "$_AGY_ERR" ]; then
537
- echo "agy log hint (pinned model may be unavailable — run 'agy models' and set review.models.agy):"
538
- echo "$_AGY_ERR"
539
- fi
540
- fi
541
- # #2073 mode 3: pre-session stall tell — no new conversation dir appeared.
542
- echo "If no agy run started, that is the pre-session-stall case: check whether a new ~/.gemini/antigravity-cli/brain/<conv-id>/ dir appeared within ~30s of launch."
543
- } > /tmp/gsd-review-antigravity-{phase}.md
544
- fi
545
-
546
- # #2176: blind-review marker. Two tells that the reviewer ran without repo
547
- # access: the prompt's mandated REVIEWED-WITHOUT-REPO-ACCESS self-report in the
548
- # first lines of output, or the agent DECLARING the scratch dir as its
549
- # workspace. Both patterns are anchored — the self-report to the head of the
550
- # file, the scratch tell to a workspace-declaration phrasing — so a grounded
551
- # review that merely QUOTES these strings (e.g. reviewing this very file) is
552
- # never mis-stamped. Stamp a machine-readable marker so the Consensus Summary
553
- # down-weights the review instead of counting an ungrounded verdict at full
554
- # weight. (Temp file + mv, no in-place sed — BSD/GNU safe.)
555
- if [ -s /tmp/gsd-review-antigravity-{phase}.md ] && \
556
- { head -5 /tmp/gsd-review-antigravity-{phase}.md | grep -q 'REVIEWED-WITHOUT-REPO-ACCESS' || \
557
- grep -qiE '(workspace|working) (directory|dir).{0,40}antigravity-cli/scratch' /tmp/gsd-review-antigravity-{phase}.md; }; then
558
- {
559
- echo "> [reviewed-without-repo-access] This reviewer ran without visibility into the repo under review — down-weight its verdict in the Consensus Summary."
560
- echo ""
561
- cat /tmp/gsd-review-antigravity-{phase}.md
562
- } > /tmp/gsd-review-antigravity-{phase}.md.tmp && \
563
- mv /tmp/gsd-review-antigravity-{phase}.md.tmp /tmp/gsd-review-antigravity-{phase}.md
564
- fi
565
- ```
566
-
567
- **Ollama (local, OpenAI-compatible):**
568
-
569
- Read host and model from config. All three local backends share the same `/v1/chat/completions` endpoint — only host and model differ. Use `jq --rawfile` to safely encode the multi-line prompt as JSON without shell-escaping issues.
280
+ Every reviewer lane is **declared data** (ADR-2782). This step iterates the lanes the selection
281
+ resolved; it does not enumerate them. Adding a reviewer is a capability manifest, not an edit here.
282
+
283
+ **Do not re-add a per-CLI block.** A `<!-- reviewer-lane: … -->` marker anywhere in this step now
284
+ FAILS the parity gate (`checkReviewerLaneParity` → `bespoke_leg_present`). Lane divergence is
285
+ declared in the manifest — timeout floor, probe, prompt/output channel, empty-output policy — and
286
+ behaviour that data genuinely cannot express is a named first-party `handler` (ADR-2782 D6), never
287
+ a bespoke block here.
288
+
289
+ **Timeout guidance (#2194):** prompt-fed source-grounded reviews are slow — measured ~570 s for
290
+ Codex at `xhigh` effort and ~525 s for headless Claude on a large plan set. Each lane declares its
291
+ own `timeoutFloorMs` and the runner enforces it internally, but the **Bash tool call wrapping the
292
+ loop below must still be given a high `timeout:`** — at least `900000`, and `1200000` when Codex or
293
+ headless Claude are in the selection — or the host kills the whole loop mid-lane. On Claude Code,
294
+ raise the host cap via `BASH_MAX_TIMEOUT_MS` if a review can exceed it.
295
+
296
+ A silent empty output after a long run is a **timeout kill, not a crash** — the Codex `0xc0000142`
297
+ misdiagnosis persisted for exactly this reason, because an empty result cannot distinguish the two
298
+ on its own. Treat an empty result on a slow lane as a dropped lane and re-run with more time rather
299
+ than diagnosing a CLI or sandbox failure. A cross-AI review that silently drops a lane is blind in
300
+ one eye.
301
+
302
+ **No hook-trust bypass (#2479):** no lane passes a hook-trust bypass flag and none runs a capability
303
+ probe for one. That flag only bypasses *persisted* hook trust (a first-run condition) and flagless
304
+ invocations work in steady state, while host-harness safety classifiers deny commands carrying it.
305
+ An environment that genuinely hits an untrusted-hook prompt surfaces through the `.err` capture and
306
+ the empty-output stub as a dropped lane with diagnosable stderr, not silent attrition. Do not
307
+ reintroduce the flag (even spelled out in prose — a regression test bans the literal file-wide).
308
+
309
+ **Reviewer instances (#1517, optional):** instances resolve *through* a lane and are not lanes
310
+ themselves (ADR-2782 D8). Each selected instance invokes its base `cli` with its own `model`/`agent`
311
+ as opaque argv. Exact invocation in `gsd-core/references/reviewer-instances.md`.
312
+
313
+ Lanes run **sequentially, not in parallel** — concurrent invocation trips provider rate limits.
570
314
 
571
315
  ```bash
572
- # Shared helper: apply prompt-budget trimming for local reviewers
316
+ RUN_DIR="{run_dir}"
317
+ REPO_ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
318
+ # SELECTED_REVIEWERS is the comma-separated result of reviewer selection (ADR-0011 precedence:
319
+ # explicit flags > --all > review.default_reviewers > all detected). Unchanged by this phase.
320
+
321
+ # Shared budget-trim helper. Was defined inside the Ollama leg; it is lane-agnostic, so it is
322
+ # hoisted here now that any lane may declare a promptBudgetKey. Returns non-zero when the budget
323
+ # is too small for the minimum review set (prompt-budget exit 2 / 11).
573
324
  prepare_trimmed_prompt_for_reviewer() {
574
- REVIEWER_KEY="$1"
575
- REVIEWER_BUDGET="$2"
576
- OUTPUT_PROMPT="$3"
577
- OUTPUT_META="$4"
578
-
579
- [ -z "$REVIEWER_BUDGET" ] && return 0
580
- [ "$REVIEWER_BUDGET" = "null" ] && return 0
581
- [ "$REVIEWER_BUDGET" = "0" ] && return 0
325
+ REVIEWER_KEY="$1"; REVIEWER_BUDGET="$2"; OUTPUT_PROMPT="$3"; OUTPUT_META="$4"
582
326
 
583
327
  PLAN_FILE_ARGS=""
584
- for p in /tmp/gsd-review-{phase}-plan-*.md; do
328
+ for p in "$RUN_DIR"/gsd-review-plan-*.md; do
585
329
  [ -f "$p" ] && PLAN_FILE_ARGS="$PLAN_FILE_ARGS --plan-file $p"
586
330
  done
587
331
  PROJECT_ARG=""
588
- [ -f "/tmp/gsd-review-{phase}-project.md" ] && PROJECT_ARG="--project-file /tmp/gsd-review-{phase}-project.md"
332
+ [ -f "$RUN_DIR/gsd-review-project.md" ] && PROJECT_ARG="--project-file $RUN_DIR/gsd-review-project.md"
589
333
  CONTEXT_ARG=""
590
- [ -f "/tmp/gsd-review-{phase}-context.md" ] && CONTEXT_ARG="--context-file /tmp/gsd-review-{phase}-context.md"
334
+ [ -f "$RUN_DIR/gsd-review-context.md" ] && CONTEXT_ARG="--context-file $RUN_DIR/gsd-review-context.md"
591
335
  RESEARCH_ARG=""
592
- [ -f "/tmp/gsd-review-{phase}-research.md" ] && RESEARCH_ARG="--research-file /tmp/gsd-review-{phase}-research.md"
336
+ [ -f "$RUN_DIR/gsd-review-research.md" ] && RESEARCH_ARG="--research-file $RUN_DIR/gsd-review-research.md"
593
337
  REQUIREMENTS_ARG=""
594
- [ -f "/tmp/gsd-review-{phase}-requirements.md" ] && REQUIREMENTS_ARG="--requirements-file /tmp/gsd-review-{phase}-requirements.md"
338
+ [ -f "$RUN_DIR/gsd-review-requirements.md" ] && REQUIREMENTS_ARG="--requirements-file $RUN_DIR/gsd-review-requirements.md"
595
339
 
596
340
  gsd_run query prompt-budget \
597
341
  --budget "$REVIEWER_BUDGET" \
598
- --instructions-file "/tmp/gsd-review-{phase}-instructions.md" \
599
- --roadmap-file "/tmp/gsd-review-{phase}-roadmap.md" \
342
+ --instructions-file "$RUN_DIR/gsd-review-instructions.md" \
343
+ --roadmap-file "$RUN_DIR/gsd-review-roadmap.md" \
600
344
  $PLAN_FILE_ARGS $PROJECT_ARG $CONTEXT_ARG $RESEARCH_ARG $REQUIREMENTS_ARG \
601
345
  --output-prompt "$OUTPUT_PROMPT" \
602
346
  --output-metadata "$OUTPUT_META"
603
347
  return $?
604
348
  }
605
349
 
606
- # Resolve prompt budget for Ollama: per-reviewer override > global default > null
607
- OLLAMA_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens_per_reviewer.ollama 2>/dev/null | jq -r '.' 2>/dev/null || echo "null")
608
- if [ -z "$OLLAMA_REVIEWER_BUDGET" ] || [ "$OLLAMA_REVIEWER_BUDGET" = "null" ]; then
609
- OLLAMA_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens 2>/dev/null | jq -r '.' 2>/dev/null || echo "null")
610
- fi
611
-
612
- # Apply budget trim for Ollama if a budget is configured
613
- OLLAMA_PROMPT_FILE="/tmp/gsd-review-prompt-{phase}.md"
614
- OLLAMA_SKIP=0
615
- if [ -n "$OLLAMA_REVIEWER_BUDGET" ] && [ "$OLLAMA_REVIEWER_BUDGET" != "null" ] && [ "$OLLAMA_REVIEWER_BUDGET" != "0" ]; then
616
- OLLAMA_TRIMMED_PROMPT="/tmp/gsd-review-prompt-{phase}-ollama.md"
617
- OLLAMA_TRIM_META="/tmp/gsd-review-prompt-{phase}-ollama.metadata.json"
618
- prepare_trimmed_prompt_for_reviewer "ollama" "$OLLAMA_REVIEWER_BUDGET" "$OLLAMA_TRIMMED_PROMPT" "$OLLAMA_TRIM_META"
619
- OLLAMA_EXIT=$?
620
- if [ $OLLAMA_EXIT -ne 0 ]; then
621
- if [ $OLLAMA_EXIT -eq 2 ] || [ $OLLAMA_EXIT -eq 11 ]; then
622
- echo "WARNING: prompt budget for ollama (${OLLAMA_REVIEWER_BUDGET} tokens) is too small for the minimum review set. Skipping Ollama reviewer." >&2
623
- else
624
- echo "WARNING: prompt-budget returned unexpected exit code ${OLLAMA_EXIT} for ollama. Skipping Ollama reviewer." >&2
625
- fi
626
- OLLAMA_SKIP=1
627
- else
628
- OLLAMA_PROMPT_FILE="$OLLAMA_TRIMMED_PROMPT"
629
- fi
630
- fi
631
-
632
- if [ "$OLLAMA_SKIP" != "1" ]; then
633
- OLLAMA_HOST=$(gsd_run query config-get review.ollama_host 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
634
- if [ -z "$OLLAMA_HOST" ] || [ "$OLLAMA_HOST" = "null" ]; then OLLAMA_HOST="http://localhost:11434"; fi
635
- OLLAMA_MODEL=$(gsd_run query config-get review.models.ollama 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
636
- if [ -z "$OLLAMA_MODEL" ] || [ "$OLLAMA_MODEL" = "null" ]; then
637
- OLLAMA_MODEL=$(curl -s --max-time 2 "${OLLAMA_HOST}/v1/models" 2>/dev/null | jq -r '.data[0].id // "llama3"' 2>/dev/null || echo "llama3")
638
- fi
639
- jq -n --rawfile content "$OLLAMA_PROMPT_FILE" \
640
- --arg model "$OLLAMA_MODEL" \
641
- '{model: $model, messages: [{role: "user", content: $content}]}' | \
642
- curl -s --max-time 120 -X POST "${OLLAMA_HOST}/v1/chat/completions" \
643
- -H "Content-Type: application/json" -d @- 2>/dev/null | \
644
- jq -r '.choices[0].message.content // "Ollama review failed or returned empty output."' \
645
- > /tmp/gsd-review-ollama-{phase}.md
646
- if [ ! -s /tmp/gsd-review-ollama-{phase}.md ]; then
647
- echo "Ollama review failed or returned empty output." > /tmp/gsd-review-ollama-{phase}.md
648
- fi
649
- fi
650
- ```
651
-
652
- **LM Studio (local, OpenAI-compatible):**
653
- ```bash
654
- # Resolve prompt budget for LM Studio: per-reviewer override > global default > null
655
- LM_STUDIO_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens_per_reviewer.lm_studio 2>/dev/null | jq -r '.' 2>/dev/null || echo "null")
656
- if [ -z "$LM_STUDIO_REVIEWER_BUDGET" ] || [ "$LM_STUDIO_REVIEWER_BUDGET" = "null" ]; then
657
- LM_STUDIO_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens 2>/dev/null | jq -r '.' 2>/dev/null || echo "null")
658
- fi
659
-
660
- # Apply budget trim for LM Studio if a budget is configured
661
- LM_STUDIO_PROMPT_FILE="/tmp/gsd-review-prompt-{phase}.md"
662
- LM_STUDIO_SKIP=0
663
- if [ -n "$LM_STUDIO_REVIEWER_BUDGET" ] && [ "$LM_STUDIO_REVIEWER_BUDGET" != "null" ] && [ "$LM_STUDIO_REVIEWER_BUDGET" != "0" ]; then
664
- LM_STUDIO_TRIMMED_PROMPT="/tmp/gsd-review-prompt-{phase}-lm_studio.md"
665
- LM_STUDIO_TRIM_META="/tmp/gsd-review-prompt-{phase}-lm_studio.metadata.json"
666
- prepare_trimmed_prompt_for_reviewer "lm_studio" "$LM_STUDIO_REVIEWER_BUDGET" "$LM_STUDIO_TRIMMED_PROMPT" "$LM_STUDIO_TRIM_META"
667
- LM_STUDIO_EXIT=$?
668
- if [ $LM_STUDIO_EXIT -ne 0 ]; then
669
- if [ $LM_STUDIO_EXIT -eq 2 ] || [ $LM_STUDIO_EXIT -eq 11 ]; then
670
- echo "WARNING: prompt budget for lm_studio (${LM_STUDIO_REVIEWER_BUDGET} tokens) is too small for the minimum review set. Skipping LM Studio reviewer." >&2
350
+ gsd_run query review-lane plan \
351
+ --selected "$SELECTED_REVIEWERS" --run-dir "$RUN_DIR" --repo-root "$REPO_ROOT" --json \
352
+ > "$RUN_DIR/gsd-review-lanes.json"
353
+
354
+ for SLUG in $(echo "$SELECTED_REVIEWERS" | tr ',' ' '); do
355
+ # Per-lane prompt budget. The lane declares its own `promptBudgetKey`; `plan` resolved it,
356
+ # applying #2797's sentinel rule (-1 = unset → fall back to the global budget; 0 legitimately
357
+ # means "do not trim this lane"). Trimming itself stays in prompt-budget, which owns it.
358
+ LANE_BUDGET=$(gsd_run query review-lane plan --selected "$SLUG" --run-dir "$RUN_DIR" \
359
+ --repo-root "$REPO_ROOT" --json 2>/dev/null \
360
+ | sed -n 's/.*"promptBudget": *\([0-9-]*\).*/\1/p' | head -1)
361
+ PROMPT_ARG=""
362
+ if [ -n "$LANE_BUDGET" ] && [ "$LANE_BUDGET" != "null" ] && [ "$LANE_BUDGET" -gt 0 ] 2>/dev/null; then
363
+ TRIMMED="$RUN_DIR/gsd-review-prompt-$SLUG.md"
364
+ if prepare_trimmed_prompt_for_reviewer "$SLUG" "$LANE_BUDGET" "$TRIMMED" \
365
+ "$RUN_DIR/gsd-review-prompt-$SLUG.metadata.json"; then
366
+ PROMPT_ARG="--prompt-file $TRIMMED"
671
367
  else
672
- echo "WARNING: prompt-budget returned unexpected exit code ${LM_STUDIO_EXIT} for lm_studio. Skipping LM Studio reviewer." >&2
368
+ # A budget too small for the minimum review set drops the lane just as silently as an empty
369
+ # response used to (#2605), so leave the skip visible in the review output, not only on stderr.
370
+ echo "$SLUG review skipped: prompt budget (${LANE_BUDGET} tokens) too small for the minimum review set." \
371
+ > "$RUN_DIR/gsd-review-$SLUG.md"
372
+ continue
673
373
  fi
674
- LM_STUDIO_SKIP=1
675
- else
676
- LM_STUDIO_PROMPT_FILE="$LM_STUDIO_TRIMMED_PROMPT"
677
374
  fi
678
- fi
679
375
 
680
- if [ "$LM_STUDIO_SKIP" != "1" ]; then
681
- LM_STUDIO_HOST=$(gsd_run query config-get review.lm_studio_host 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
682
- if [ -z "$LM_STUDIO_HOST" ] || [ "$LM_STUDIO_HOST" = "null" ]; then LM_STUDIO_HOST="http://localhost:1234"; fi
683
- LM_STUDIO_MODEL=$(gsd_run query config-get review.models.lm_studio 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
684
- if [ -z "$LM_STUDIO_MODEL" ] || [ "$LM_STUDIO_MODEL" = "null" ]; then
685
- LM_STUDIO_MODEL=$(curl -s --max-time 2 "${LM_STUDIO_HOST}/v1/models" 2>/dev/null | jq -r '.data[0].id // "local-model"' 2>/dev/null || echo "local-model")
686
- fi
687
- LM_STUDIO_RESPONSE=$(jq -n --rawfile content "$LM_STUDIO_PROMPT_FILE" \
688
- --arg model "$LM_STUDIO_MODEL" \
689
- '{model: $model, messages: [{role: "user", content: $content}]}' | \
690
- curl -s --max-time 120 -X POST "${LM_STUDIO_HOST}/v1/chat/completions" \
691
- -H "Content-Type: application/json" -d @- 2>/dev/null)
692
- LM_STUDIO_ACTUAL_MODEL=$(echo "$LM_STUDIO_RESPONSE" | jq -r '.model // ""' 2>/dev/null || echo "")
693
- if [ -n "$LM_STUDIO_ACTUAL_MODEL" ] && [ "$LM_STUDIO_ACTUAL_MODEL" != "null" ] && [ "$LM_STUDIO_ACTUAL_MODEL" != "$LM_STUDIO_MODEL" ]; then
694
- echo "Warning: LM Studio served model '$LM_STUDIO_ACTUAL_MODEL' but '$LM_STUDIO_MODEL' was requested. Review may be from a different model." >&2
695
- fi
696
- LM_STUDIO_CONTENT=$(echo "$LM_STUDIO_RESPONSE" | jq -r '.choices[0].message.content // ""' 2>/dev/null || echo "")
697
- if [ -n "$LM_STUDIO_CONTENT" ]; then
698
- echo "$LM_STUDIO_CONTENT" > /tmp/gsd-review-lm_studio-{phase}.md
699
- else
700
- echo "Warning: LM Studio returned empty content — skipping review." >&2
701
- fi
702
- fi
376
+ # One invocation, whatever the lane's transport, prompt channel, output channel or handler.
377
+ # `--explicit` marks a lane the user NAMED: ADR-2782 D4 — not finding a lane nobody asked for is
378
+ # normal, failing to run one somebody asked for is an error.
379
+ gsd_run query review-lane invoke --slug "$SLUG" \
380
+ --run-dir "$RUN_DIR" --repo-root "$REPO_ROOT" $PROMPT_ARG $EXPLICIT_FLAG --json \
381
+ >> "$RUN_DIR/gsd-review-lane-results.jsonl"
382
+ done
703
383
  ```
704
384
 
705
- **llama.cpp (local, OpenAI-compatible):**
706
- ```bash
707
- # Resolve prompt budget for llama.cpp: per-reviewer override > global default > null
708
- LLAMA_CPP_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens_per_reviewer.llama_cpp 2>/dev/null | jq -r '.' 2>/dev/null || echo "null")
709
- if [ -z "$LLAMA_CPP_REVIEWER_BUDGET" ] || [ "$LLAMA_CPP_REVIEWER_BUDGET" = "null" ]; then
710
- LLAMA_CPP_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens 2>/dev/null | jq -r '.' 2>/dev/null || echo "null")
711
- fi
385
+ Each lane leaves `{run_dir}/gsd-review-<slug>.md` — its review, or a diagnostic stub carrying the
386
+ captured stderr (and, for an OpenAI-compatible lane, the raw response body, where such a server puts
387
+ its error JSON on an HTTP 4xx/5xx while still exiting 0). A stub is never mistaken for a clean
388
+ review: it keeps its "failed or returned empty output" header (#2494/#2605/#2794).
712
389
 
713
- # Apply budget trim for llama.cpp if a budget is configured
714
- LLAMA_CPP_PROMPT_FILE="/tmp/gsd-review-prompt-{phase}.md"
715
- LLAMA_CPP_SKIP=0
716
- if [ -n "$LLAMA_CPP_REVIEWER_BUDGET" ] && [ "$LLAMA_CPP_REVIEWER_BUDGET" != "null" ] && [ "$LLAMA_CPP_REVIEWER_BUDGET" != "0" ]; then
717
- LLAMA_CPP_TRIMMED_PROMPT="/tmp/gsd-review-prompt-{phase}-llama_cpp.md"
718
- LLAMA_CPP_TRIM_META="/tmp/gsd-review-prompt-{phase}-llama_cpp.metadata.json"
719
- prepare_trimmed_prompt_for_reviewer "llama_cpp" "$LLAMA_CPP_REVIEWER_BUDGET" "$LLAMA_CPP_TRIMMED_PROMPT" "$LLAMA_CPP_TRIM_META"
720
- LLAMA_CPP_EXIT=$?
721
- if [ $LLAMA_CPP_EXIT -ne 0 ]; then
722
- if [ $LLAMA_CPP_EXIT -eq 2 ] || [ $LLAMA_CPP_EXIT -eq 11 ]; then
723
- echo "WARNING: prompt budget for llama_cpp (${LLAMA_CPP_REVIEWER_BUDGET} tokens) is too small for the minimum review set. Skipping llama.cpp reviewer." >&2
724
- else
725
- echo "WARNING: prompt-budget returned unexpected exit code ${LLAMA_CPP_EXIT} for llama_cpp. Skipping llama.cpp reviewer." >&2
726
- fi
727
- LLAMA_CPP_SKIP=1
728
- else
729
- LLAMA_CPP_PROMPT_FILE="$LLAMA_CPP_TRIMMED_PROMPT"
730
- fi
731
- fi
732
-
733
- if [ "$LLAMA_CPP_SKIP" != "1" ]; then
734
- LLAMA_CPP_HOST=$(gsd_run query config-get review.llama_cpp_host 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
735
- if [ -z "$LLAMA_CPP_HOST" ] || [ "$LLAMA_CPP_HOST" = "null" ]; then LLAMA_CPP_HOST="http://localhost:8080"; fi
736
- LLAMA_CPP_MODEL=$(gsd_run query config-get review.models.llama_cpp 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
737
- if [ -z "$LLAMA_CPP_MODEL" ] || [ "$LLAMA_CPP_MODEL" = "null" ]; then
738
- LLAMA_CPP_MODEL=$(curl -s --max-time 2 "${LLAMA_CPP_HOST}/v1/models" 2>/dev/null | jq -r '.data[0].id // "local-model"' 2>/dev/null || echo "local-model")
739
- fi
740
- LLAMA_CPP_CONTENT=$(jq -n --rawfile content "$LLAMA_CPP_PROMPT_FILE" \
741
- --arg model "$LLAMA_CPP_MODEL" \
742
- '{model: $model, messages: [{role: "user", content: $content}]}' | \
743
- curl -s --max-time 120 -X POST "${LLAMA_CPP_HOST}/v1/chat/completions" \
744
- -H "Content-Type: application/json" -d @- 2>/dev/null | \
745
- jq -r '.choices[0].message.content // ""' 2>/dev/null || echo "")
746
- if [ -n "$LLAMA_CPP_CONTENT" ]; then
747
- echo "$LLAMA_CPP_CONTENT" > /tmp/gsd-review-llama_cpp-{phase}.md
748
- else
749
- echo "Warning: llama.cpp returned empty content — skipping review." >&2
750
- fi
751
- fi
752
- ```
753
-
754
- If a CLI or local server fails, log the error and continue with remaining reviewers.
390
+ A lane that will not run reports a typed reason rather than an empty file — `missing_binary`,
391
+ `probe_failed`, `probe_timeout`, `missing_required_binary`, `host_unreachable`,
392
+ `egress_host_changed`, `unknown_handler`, `budget_too_small`. **`egress_host_changed` means the lane
393
+ was consented to send plans to one destination and `.planning/config.json` now names another; it is
394
+ blocked, not silently redirected** (ADR-2782 D5).
755
395
 
756
396
  Display progress:
757
397
  ```
@@ -794,83 +434,26 @@ trimmed_reviewers: # only present if at least one reviewer was trimmed
794
434
 
795
435
  # Cross-AI Plan Review — Phase {N}
796
436
 
797
- ## Gemini Review
798
-
799
- {gemini review content}
800
-
801
- ---
802
-
803
- ## Claude Review
804
-
805
- {claude review content}
806
-
807
- ---
808
-
809
- ## Codex Review
810
-
811
- {codex review content}
812
-
813
- ---
814
-
815
- ## CodeRabbit Review
816
-
817
- {coderabbit review content}
818
-
819
- ---
820
-
821
- ## OpenCode Review
822
-
823
- {opencode review content}
824
-
825
- ---
826
-
827
- ## OpenCode Review (opencode-deepseek)
828
-
829
- {opencode-deepseek instance review content — only present when this instance was selected}
830
-
831
- ---
832
-
833
- ## OpenCode Review (opencode-mimo)
834
-
835
- {opencode-mimo instance review content — only present when this instance was selected}
836
-
837
- ---
838
-
839
- ## Qwen Review
437
+ <!-- Sections are RENDERED from each lane's declared `reviewsSection`, in descriptor order.
438
+ There is deliberately no hardcoded per-reviewer heading list here any more: a hand-maintained
439
+ list is exactly the drift #2781 was filed about, and it silently disagreed with the roster.
440
+ `gsd_run query review-lane sections --selected "$SELECTED_REVIEWERS"` emits
441
+ `<slug><TAB><reviewsSection>` in order; for each row, emit:
840
442
 
841
- {qwen review content}
443
+ ## <reviewsSection> Review
842
444
 
843
- ---
844
-
845
- ## Cursor Review
846
-
847
- {cursor review content}
848
-
849
- ---
850
-
851
- ## Antigravity Review
445
+ {contents of {run_dir}/gsd-review-<slug>.md}
852
446
 
853
- {antigravity review content}
447
+ ---
854
448
 
855
- ---
856
-
857
- ## Ollama Review
858
-
859
- {ollama review content}
860
-
861
- ---
449
+ Two headings must NOT be generated from this list, because they are not lanes:
450
+ * `## <Adapter> Review (<instance>)` — an ADR-1517 reviewer INSTANCE resolves THROUGH a lane
451
+ and is rendered from the instance list, not the lane list (ADR-2782 D8).
452
+ * `## Consensus Summary` — not a review section at all.
862
453
 
863
- ## LM Studio Review
864
-
865
- {lm_studio review content}
866
-
867
- ---
868
-
869
- ## llama.cpp Review
870
-
871
- {llama_cpp review content}
872
-
873
- ---
454
+ A lane whose `evidenceClass` is `diff-only` (CodeRabbit) carries its caveat from data: it never
455
+ received the source-grounding prompt, so its verdict is folded in as a diff observation and is
456
+ not weighted as a grounded plan review. -->
874
457
 
875
458
  ## Consensus Summary
876
459
 
@@ -911,7 +494,11 @@ To incorporate feedback into planning:
911
494
  /gsd:plan-phase {N} --reviews
912
495
  ```
913
496
 
914
- Clean up temp files.
497
+ Clean up — remove the run's temp directory now that REVIEWS.md is committed:
498
+
499
+ ```bash
500
+ rm -rf "{run_dir}"
501
+ ```
915
502
  </step>
916
503
 
917
504
  </process>