@opengsd/gsd-core 1.8.0 → 1.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (174) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/.claude-plugin/plugin.json +1 -1
  3. package/.opencode/plugins/gsd-core.js +31 -1
  4. package/agents/gsd-code-fixer.md +1 -1
  5. package/agents/gsd-codebase-mapper.md +1 -1
  6. package/agents/gsd-debug-session-manager.md +36 -0
  7. package/agents/gsd-executor.md +20 -7
  8. package/agents/gsd-intel-updater.md +3 -3
  9. package/agents/gsd-phase-researcher.md +4 -2
  10. package/agents/gsd-plan-checker.md +20 -0
  11. package/agents/gsd-planner.md +15 -23
  12. package/agents/gsd-project-researcher.md +2 -2
  13. package/agents/gsd-ui-auditor.md +0 -40
  14. package/bin/install.js +186 -55
  15. package/commands/gsd/plan-review-convergence.md +5 -1
  16. package/gsd-core/bin/gsd-tools.cjs +849 -2
  17. package/gsd-core/bin/lib/api-coverage.cjs +22 -8
  18. package/gsd-core/bin/lib/audit.cjs +8 -8
  19. package/gsd-core/bin/lib/capability-consent.cjs +40 -1
  20. package/gsd-core/bin/lib/capability-lifecycle.cjs +58 -0
  21. package/gsd-core/bin/lib/capability-loader.cjs +23 -1
  22. package/gsd-core/bin/lib/capability-registry.cjs +1353 -132
  23. package/gsd-core/bin/lib/capability-trust.cjs +468 -33
  24. package/gsd-core/bin/lib/capability-validator.cjs +882 -6
  25. package/gsd-core/bin/lib/check-command-router.cjs +12 -2
  26. package/gsd-core/bin/lib/cjs-command-router-adapter.cjs +15 -0
  27. package/gsd-core/bin/lib/claude-orchestration-command-router.cjs +102 -12
  28. package/gsd-core/bin/lib/claude-orchestration.cjs +125 -22
  29. package/gsd-core/bin/lib/commands.cjs +246 -18
  30. package/gsd-core/bin/lib/config-loader.cjs +200 -28
  31. package/gsd-core/bin/lib/config.cjs +90 -5
  32. package/gsd-core/bin/lib/estimate-cli.cjs +336 -0
  33. package/gsd-core/bin/lib/frontmatter.cjs +125 -15
  34. package/gsd-core/bin/lib/host-integration.cjs +215 -8
  35. package/gsd-core/bin/lib/init.cjs +44 -19
  36. package/gsd-core/bin/lib/install-engine.cjs +1 -0
  37. package/gsd-core/bin/lib/milestone.cjs +5 -5
  38. package/gsd-core/bin/lib/model-catalog.cjs +51 -1
  39. package/gsd-core/bin/lib/observability/logger.cjs +7 -2
  40. package/gsd-core/bin/lib/phase-command-router.cjs +10 -1
  41. package/gsd-core/bin/lib/phase-estimation.cjs +398 -0
  42. package/gsd-core/bin/lib/phase-id.cjs +278 -5
  43. package/gsd-core/bin/lib/phase.cjs +57 -5
  44. package/gsd-core/bin/lib/plan-drift-guard.cjs +1 -1
  45. package/gsd-core/bin/lib/plan-scan.cjs +1 -1
  46. package/gsd-core/bin/lib/planning-workspace.cjs +9 -2
  47. package/gsd-core/bin/lib/profile-output.cjs +34 -8
  48. package/gsd-core/bin/lib/review-lane-descriptor.cjs +927 -0
  49. package/gsd-core/bin/lib/review-lane-invocation.cjs +348 -0
  50. package/gsd-core/bin/lib/review-lane-runner.cjs +594 -0
  51. package/gsd-core/bin/lib/review-reviewer-selection.cjs +114 -32
  52. package/gsd-core/bin/lib/roadmap-parser.cjs +54 -6
  53. package/gsd-core/bin/lib/roadmap.cjs +10 -4
  54. package/gsd-core/bin/lib/runtime-artifact-conversion.cjs +31 -4
  55. package/gsd-core/bin/lib/runtime-artifact-layout.cjs +1 -1
  56. package/gsd-core/bin/lib/runtime-hooks-surface.cjs +140 -0
  57. package/gsd-core/bin/lib/runtime-name-policy.cjs +15 -2
  58. package/gsd-core/bin/lib/smart-entry.cjs +1 -1
  59. package/gsd-core/bin/lib/state-document.cjs +164 -20
  60. package/gsd-core/bin/lib/state-transition.cjs +28 -10
  61. package/gsd-core/bin/lib/state.cjs +141 -21
  62. package/gsd-core/bin/lib/uat-predicate.cjs +6 -4
  63. package/gsd-core/bin/lib/uat.cjs +9 -7
  64. package/gsd-core/bin/lib/ui-consideration-probe.cjs +2 -2
  65. package/gsd-core/bin/lib/unusable-input.cjs +216 -0
  66. package/gsd-core/bin/lib/validate.cjs +32 -0
  67. package/gsd-core/bin/lib/verification.cjs +51 -14
  68. package/gsd-core/bin/lib/verify.cjs +128 -20
  69. package/gsd-core/bin/lib/worktree-safety.cjs +360 -15
  70. package/gsd-core/bin/shared/config-defaults.manifest.json +1 -0
  71. package/gsd-core/bin/shared/config-schema.manifest.json +1 -13
  72. package/gsd-core/bin/shared/model-catalog.json +5 -0
  73. package/gsd-core/bin/shared/runtime-aliases.manifest.json +5 -0
  74. package/gsd-core/references/context-budget.md +40 -0
  75. package/gsd-core/references/gate-prompts.md +6 -3
  76. package/gsd-core/references/model-profile-resolution.md +64 -13
  77. package/gsd-core/references/offer-next.md +88 -0
  78. package/gsd-core/references/planning-config.md +2 -1
  79. package/gsd-core/references/reviewer-instances.md +28 -21
  80. package/gsd-core/references/runtime-aware-dispatch.md +42 -0
  81. package/gsd-core/references/ui-consideration-probe.md +2 -2
  82. package/gsd-core/references/worktree-branch-check.md +4 -4
  83. package/gsd-core/templates/summary-minimal.md +4 -0
  84. package/gsd-core/templates/summary-standard.md +4 -0
  85. package/gsd-core/templates/summary.md +7 -0
  86. package/gsd-core/workflows/ai-integration-phase.md +4 -4
  87. package/gsd-core/workflows/audit-fix.md +4 -0
  88. package/gsd-core/workflows/audit-milestone.md +8 -0
  89. package/gsd-core/workflows/autonomous.md +19 -15
  90. package/gsd-core/workflows/check-todos.md +2 -2
  91. package/gsd-core/workflows/code-review-fix.md +14 -6
  92. package/gsd-core/workflows/code-review.md +76 -19
  93. package/gsd-core/workflows/debug.md +10 -2
  94. package/gsd-core/workflows/diagnose-issues.md +4 -0
  95. package/gsd-core/workflows/discuss-phase/modes/advisor.md +2 -4
  96. package/gsd-core/workflows/discuss-phase/modes/auto.md +0 -6
  97. package/gsd-core/workflows/discuss-phase-assumptions.md +15 -9
  98. package/gsd-core/workflows/discuss-phase.md +2 -2
  99. package/gsd-core/workflows/docs-update.md +8 -0
  100. package/gsd-core/workflows/eval-review.md +1 -1
  101. package/gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md +4 -0
  102. package/gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md +160 -0
  103. package/gsd-core/workflows/execute-phase.md +85 -115
  104. package/gsd-core/workflows/execute-plan.md +5 -4
  105. package/gsd-core/workflows/explore.md +4 -0
  106. package/gsd-core/workflows/extract-learnings.md +21 -0
  107. package/gsd-core/workflows/help/modes/full.md +3 -3
  108. package/gsd-core/workflows/import.md +4 -1
  109. package/gsd-core/workflows/ingest-docs.md +4 -0
  110. package/gsd-core/workflows/map-codebase.md +13 -6
  111. package/gsd-core/workflows/new-milestone.md +10 -2
  112. package/gsd-core/workflows/new-project.md +11 -4
  113. package/gsd-core/workflows/next.md +5 -2
  114. package/gsd-core/workflows/plan-phase.md +42 -46
  115. package/gsd-core/workflows/plan-review-convergence.md +18 -14
  116. package/gsd-core/workflows/progress.md +1 -1
  117. package/gsd-core/workflows/quick.md +14 -3
  118. package/gsd-core/workflows/review.md +146 -575
  119. package/gsd-core/workflows/scan.md +9 -1
  120. package/gsd-core/workflows/secure-phase.md +10 -2
  121. package/gsd-core/workflows/ship.md +41 -11
  122. package/gsd-core/workflows/smart-entry.md +1 -1
  123. package/gsd-core/workflows/ui-phase.md +8 -1
  124. package/gsd-core/workflows/ui-review.md +8 -1
  125. package/gsd-core/workflows/update.md +104 -5
  126. package/gsd-core/workflows/validate-phase.md +10 -2
  127. package/gsd-core/workflows/verify-work.md +8 -1
  128. package/hooks/dist/gsd-cursor-session-start.js +6 -2
  129. package/hooks/dist/gsd-cursor-stop.js +6 -2
  130. package/hooks/dist/gsd-cursor-subagent-start.js +6 -2
  131. package/hooks/dist/gsd-graphify-update.sh +9 -0
  132. package/hooks/dist/gsd-phase-boundary.sh +14 -2
  133. package/hooks/dist/gsd-prompt-guard.js +101 -2
  134. package/hooks/dist/gsd-read-guard.js +100 -2
  135. package/hooks/dist/gsd-read-injection-scanner.js +109 -2
  136. package/hooks/dist/gsd-statusline.js +9 -6
  137. package/hooks/dist/gsd-workflow-guard.js +110 -6
  138. package/hooks/dist/gsd-worktree-path-guard.js +132 -8
  139. package/hooks/dist/lib/cursor-workspace.js +74 -0
  140. package/hooks/gsd-cursor-session-start.js +6 -2
  141. package/hooks/gsd-cursor-stop.js +6 -2
  142. package/hooks/gsd-cursor-subagent-start.js +6 -2
  143. package/hooks/gsd-graphify-update.sh +9 -0
  144. package/hooks/gsd-phase-boundary.sh +14 -2
  145. package/hooks/gsd-prompt-guard.js +101 -2
  146. package/hooks/gsd-read-guard.js +100 -2
  147. package/hooks/gsd-read-injection-scanner.js +109 -2
  148. package/hooks/gsd-statusline.js +9 -6
  149. package/hooks/gsd-workflow-guard.js +110 -6
  150. package/hooks/gsd-worktree-path-guard.js +132 -8
  151. package/hooks/lib/cursor-workspace.js +74 -0
  152. package/package.json +7 -7
  153. package/pi/gsd.cjs +26 -1
  154. package/scripts/check-coverage-gate.cjs +51 -0
  155. package/scripts/check-glossary-refs.cjs +24 -0
  156. package/scripts/ci-test-scope.cjs +67 -17
  157. package/scripts/gen-adr-index.cjs +6 -4
  158. package/scripts/gen-capability-matrix.cjs +26 -2
  159. package/scripts/gen-capability-registry.cjs +132 -34
  160. package/scripts/gen-emitted-baseline.cjs +145 -0
  161. package/scripts/lint-compiled-artifact-sync.cjs +146 -0
  162. package/scripts/lint-emitted-drift-ack.cjs +149 -0
  163. package/scripts/lint-fix-has-regression-test.cjs +131 -0
  164. package/scripts/lint-resolution-provenance.cjs +9 -0
  165. package/scripts/mutation-matrix.cjs +4 -0
  166. package/scripts/prompt-injection-scan.sh +6 -0
  167. package/scripts/registry-schema.cjs +57 -8
  168. package/scripts/release-notes/conventional-title.cjs +19 -1
  169. package/scripts/release-notes/format-github-release-notes.cjs +7 -3
  170. package/scripts/workflow-size.cjs +16 -8
  171. package/skills/gsd-plan-review-convergence/SKILL.md +5 -1
  172. package/vscode/package.json +1 -1
  173. package/scripts/gen-golden-install-parity-zcode.cjs +0 -77
  174. package/scripts/update-size-baseline.cjs +0 -68
@@ -26,19 +26,46 @@ command -v cursor-agent >/dev/null 2>&1 && echo "cursor:available" || echo "curs
26
26
  command -v agy >/dev/null 2>&1 && echo "antigravity:available" || echo "antigravity:missing"
27
27
 
28
28
  # Check local model servers (OpenAI-compatible HTTP API — no CLI binary required)
29
- OLLAMA_HOST=$(gsd_run query config-get review.ollama_host 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
29
+ OLLAMA_HOST=$(gsd_run query config-get review.ollama_host --raw 2>/dev/null || echo "")
30
30
  if [ -z "$OLLAMA_HOST" ] || [ "$OLLAMA_HOST" = "null" ]; then OLLAMA_HOST="http://localhost:11434"; fi
31
31
  curl -s --max-time 2 "${OLLAMA_HOST}/v1/models" >/dev/null 2>&1 && echo "ollama:available" || echo "ollama:missing"
32
32
 
33
- LM_STUDIO_HOST=$(gsd_run query config-get review.lm_studio_host 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
33
+ LM_STUDIO_HOST=$(gsd_run query config-get review.lm_studio_host --raw 2>/dev/null || echo "")
34
34
  if [ -z "$LM_STUDIO_HOST" ] || [ "$LM_STUDIO_HOST" = "null" ]; then LM_STUDIO_HOST="http://localhost:1234"; fi
35
35
  curl -s --max-time 2 "${LM_STUDIO_HOST}/v1/models" >/dev/null 2>&1 && echo "lm_studio:available" || echo "lm_studio:missing"
36
36
 
37
- LLAMA_CPP_HOST=$(gsd_run query config-get review.llama_cpp_host 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
37
+ LLAMA_CPP_HOST=$(gsd_run query config-get review.llama_cpp_host --raw 2>/dev/null || echo "")
38
38
  if [ -z "$LLAMA_CPP_HOST" ] || [ "$LLAMA_CPP_HOST" = "null" ]; then LLAMA_CPP_HOST="http://localhost:8080"; fi
39
39
  curl -s --max-time 2 "${LLAMA_CPP_HOST}/v1/models" >/dev/null 2>&1 && echo "llama_cpp:available" || echo "llama_cpp:missing"
40
+
41
+ # jq prerequisite (#2589). The config/model/budget lookups in this workflow no
42
+ # longer need jq — they use the native --raw/--pick flags. But the lanes listed
43
+ # under "jq-dependent reviewer lanes" below parse structured JSON that gsd-tools
44
+ # does not emit (OpenAI-compatible /v1/chat/completions responses, opencode's
45
+ # JSONL event stream, agy's conversation cache), so they cannot run without jq.
46
+ # Probe it here rather than letting each lane swallow exit 127 into empty output.
47
+ command -v jq >/dev/null 2>&1 && echo "jq:available" || echo "jq:missing"
48
+ ```
49
+
50
+ **jq-dependent reviewer lanes.** `jq` is a production prerequisite for the
51
+ `ollama`, `lm_studio`, `llama_cpp`, `opencode`, and `antigravity` lanes only. If
52
+ `detect_clis` reports `jq:missing`, treat those five as **undetected** — they
53
+ follow the same "known-but-undetected" path as a missing CLI. Which path that is
54
+ depends on how the lane was selected (see the precedence rules below): reached
55
+ through `review.default_reviewers` or `--all` it is an info note and the lane is
56
+ ignored; named by an explicit flag it is an **error**, because the user asserted
57
+ that lane. Tell the user to install jq:
58
+
59
+ ```
60
+ NOTE: jq is not on PATH — the ollama, lm_studio, llama_cpp, opencode, and
61
+ antigravity reviewer lanes are unavailable. Install jq (https://jqlang.org/download/)
62
+ or select a lane that does not require it (--gemini, --claude, --codex,
63
+ --coderabbit, --qwen, --cursor).
40
64
  ```
41
65
 
66
+ The remaining lanes (`gemini`, `claude`, `codex`, `coderabbit`, `qwen`, `cursor`)
67
+ do not require jq and must stay selectable on a jq-less host.
68
+
42
69
  Parse flags from `$ARGUMENTS`:
43
70
  - `--gemini` → include Gemini
44
71
  - `--claude` → include Claude
@@ -60,10 +87,22 @@ Reviewer-selection precedence:
60
87
  3. `review.default_reviewers`
61
88
  4. No key + no flags → all detected reviewers
62
89
 
90
+ **Explicit reviewer flags are an assertion, not a preference (ADR-2782 D4).** A lane the user
91
+ named on the command line and that cannot run is an **error**, surfaced and non-silent — even
92
+ when other named lanes did run. Do not proceed with a thinner reviewer set and report success:
93
+ `--gemini --qwen` on a host without `qwen` fails, it does not quietly become a Gemini-only
94
+ review. This applies however the lane became unavailable — binary missing, prerequisite `jq`
95
+ absent, or a local server not reachable.
96
+
97
+ The asymmetry is deliberate: *not finding a lane nobody asked for is normal; failing to run a
98
+ lane somebody asked for is an error.* A user who wants "whatever is available" has `--all`; a
99
+ user who wants a preferred set has `review.default_reviewers`. Both stay lenient below.
100
+
63
101
  `review.default_reviewers` behavior:
64
102
  - Value must be a non-empty array of slug strings (configured via `gsd config-set review.default_reviewers '["gemini","codex"]'`)
65
103
  - Unknown slugs warn and are ignored
66
- - Known-but-undetected slugs emit an info note and are ignored
104
+ - Known-but-undetected slugs emit an info note and are ignored — a configured default is a
105
+ preference evaluated across many hosts, so a subset being present is expected, not an error
67
106
  - If all configured reviewers are unavailable, fail with an actionable message
68
107
 
69
108
  **Reviewer instances (#1517, optional):** if `review.reviewer_instances` is configured,
@@ -238,532 +277,121 @@ Note: `INSTRUCTIONS_BLOCK_FILE`, `ROADMAP_SECTION_FILE`, and `PHASE_DIR` come fr
238
277
  </step>
239
278
 
240
279
  <step name="invoke_reviewers">
241
- Read model preferences from planning config. Null/missing values fall back to CLI defaults.
242
-
243
- ```bash
244
- # JSON scalars from gsd-tools.cjs query; use jq -r to strip JSON string quotes (install jq if missing)
245
- GEMINI_MODEL=$(gsd_run query config-get review.models.gemini 2>/dev/null | jq -r '.' 2>/dev/null || true)
246
- CLAUDE_MODEL=$(gsd_run query config-get review.models.claude 2>/dev/null | jq -r '.' 2>/dev/null || true)
247
- CODEX_MODEL=$(gsd_run query config-get review.models.codex 2>/dev/null | jq -r '.' 2>/dev/null || true)
248
- OPENCODE_MODEL=$(gsd_run query config-get review.models.opencode 2>/dev/null | jq -r '.' 2>/dev/null || true)
249
- # review.models.agy, when set, is passed to agy as --model (escape hatch for a
250
- # pinned model that 404s server-side); otherwise agy uses its persisted default.
251
- AGY_MODEL=$(gsd_run query config-get review.models.agy 2>/dev/null | jq -r '.' 2>/dev/null || true)
252
-
253
- # #1115: `--dangerously-bypass-hook-trust` only exists on codex-cli >= 0.137.0.
254
- # Capability-probe it so older installs don't fail with "unexpected argument"
255
- # (which, with stderr suppressed, produced a silent empty review). The codex
256
- # invocation works fine without the flag on older versions.
257
- if codex exec --help 2>/dev/null | grep -q -- '--dangerously-bypass-hook-trust'; then
258
- CODEX_BYPASS_FLAG="--dangerously-bypass-hook-trust"
259
- else
260
- CODEX_BYPASS_FLAG=""
261
- fi
262
- ```
263
-
264
- **Reviewer instances (#1517, optional):** when instances are configured, each selected
265
- instance invokes its base `cli` with its own `model`/`agent` (opaque argv, never
266
- shell-interpolated). Exact invocation in `gsd-core/references/reviewer-instances.md`.
267
-
268
- For each selected CLI, invoke in sequence (not parallel — avoid rate limits):
269
-
270
- **Timeout guidance (#2194):** prompt-fed source-grounded reviews are slow — measured ~570s for Codex at `xhigh` effort and ~525s for headless Claude on a large plan set. Each of the Gemini / Claude / Codex blocks below MUST be invoked with a high Bash `timeout:` — at least `900000` (15 min), and `1200000` (20 min) for Codex `xhigh` or headless Claude — so a lane is not killed mid-review. On Claude Code, raise the host cap via `BASH_MAX_TIMEOUT_MS` if a review can exceed it. A silent empty output after a long run is a **timeout kill, not a crash** — the Codex `0xc0000142` misdiagnosis persisted because the empty-output branches below cannot distinguish the two; treat an empty result on a slow lane as a dropped lane and re-run with more time rather than diagnosing a CLI/sandbox failure. A cross-AI review that silently drops a lane is blind in one eye.
271
-
272
- **Gemini:**
273
- ```bash
274
- if [ -n "$GEMINI_MODEL" ] && [ "$GEMINI_MODEL" != "null" ]; then
275
- cat {run_dir}/gsd-review-prompt.md | gemini -m "$GEMINI_MODEL" -p - 2>/dev/null > {run_dir}/gsd-review-gemini.md
276
- else
277
- cat {run_dir}/gsd-review-prompt.md | gemini -p - 2>/dev/null > {run_dir}/gsd-review-gemini.md
278
- fi
279
- ```
280
+ Every reviewer lane is **declared data** (ADR-2782). This step iterates the lanes the selection
281
+ resolved; it does not enumerate them. Adding a reviewer is a capability manifest, not an edit here.
282
+
283
+ **Do not re-add a per-CLI block.** A `<!-- reviewer-lane: … -->` marker anywhere in this step now
284
+ FAILS the parity gate (`checkReviewerLaneParity` → `bespoke_leg_present`). Lane divergence is
285
+ declared in the manifest — timeout floor, probe, prompt/output channel, empty-output policy — and
286
+ behaviour that data genuinely cannot express is a named first-party `handler` (ADR-2782 D6), never
287
+ a bespoke block here.
288
+
289
+ **Timeout guidance (#2194):** prompt-fed source-grounded reviews are slow — measured ~570 s for
290
+ Codex at `xhigh` effort and ~525 s for headless Claude on a large plan set. Each lane declares its
291
+ own `timeoutFloorMs` and the runner enforces it internally, but the **Bash tool call wrapping the
292
+ loop below must still be given a high `timeout:`** — at least `900000`, and `1200000` when Codex or
293
+ headless Claude are in the selection — or the host kills the whole loop mid-lane. On Claude Code,
294
+ raise the host cap via `BASH_MAX_TIMEOUT_MS` if a review can exceed it.
295
+
296
+ A silent empty output after a long run is a **timeout kill, not a crash** — the Codex `0xc0000142`
297
+ misdiagnosis persisted for exactly this reason, because an empty result cannot distinguish the two
298
+ on its own. Treat an empty result on a slow lane as a dropped lane and re-run with more time rather
299
+ than diagnosing a CLI or sandbox failure. A cross-AI review that silently drops a lane is blind in
300
+ one eye.
301
+
302
+ **No hook-trust bypass (#2479):** no lane passes a hook-trust bypass flag and none runs a capability
303
+ probe for one. That flag only bypasses *persisted* hook trust (a first-run condition) and flagless
304
+ invocations work in steady state, while host-harness safety classifiers deny commands carrying it.
305
+ An environment that genuinely hits an untrusted-hook prompt surfaces through the `.err` capture and
306
+ the empty-output stub as a dropped lane with diagnosable stderr, not silent attrition. Do not
307
+ reintroduce the flag (even spelled out in prose — a regression test bans the literal file-wide).
308
+
309
+ **Reviewer instances (#1517, optional):** instances resolve *through* a lane and are not lanes
310
+ themselves (ADR-2782 D8). Each selected instance invokes its base `cli` with its own `model`/`agent`
311
+ as opaque argv. Exact invocation in `gsd-core/references/reviewer-instances.md`.
312
+
313
+ Lanes run **sequentially, not in parallel** — concurrent invocation trips provider rate limits.
280
314
 
281
- **Claude (separate session):**
282
315
  ```bash
283
- if [ -n "$CLAUDE_MODEL" ] && [ "$CLAUDE_MODEL" != "null" ]; then
284
- cat {run_dir}/gsd-review-prompt.md | claude --model "$CLAUDE_MODEL" -p - 2>/dev/null > {run_dir}/gsd-review-claude.md
285
- else
286
- cat {run_dir}/gsd-review-prompt.md | claude -p - 2>/dev/null > {run_dir}/gsd-review-claude.md
287
- fi
288
- ```
289
-
290
- **Codex:**
291
- ```bash
292
- # $CODEX_BYPASS_FLAG is capability-gated above (#1115). Capture stderr to a .err
293
- # file (not /dev/null) so a non-zero exit — e.g. a flag the installed codex-cli
294
- # does not support — is diagnosable instead of a silent empty review.
295
- # Capture the review via codex's own `-o/--output-last-message <FILE>` (only the
296
- # final agent message) and discard stdout (#1698): on some platforms (Windows)
297
- # codex writes process-teardown output to stdout *after* the final message, and a
298
- # stdout redirect would append that noise to a non-empty file — slipping past the
299
- # `[ ! -s … ]` empty-output guard as a silently polluted review.
300
- if [ -n "$CODEX_MODEL" ] && [ "$CODEX_MODEL" != "null" ]; then
301
- cat {run_dir}/gsd-review-prompt.md | codex exec --ephemeral $CODEX_BYPASS_FLAG --model "$CODEX_MODEL" --skip-git-repo-check -o {run_dir}/gsd-review-codex.md - 2>{run_dir}/gsd-review-codex.err >/dev/null
302
- else
303
- cat {run_dir}/gsd-review-prompt.md | codex exec --ephemeral $CODEX_BYPASS_FLAG --skip-git-repo-check -o {run_dir}/gsd-review-codex.md - 2>{run_dir}/gsd-review-codex.err >/dev/null
304
- fi
305
- if [ ! -s {run_dir}/gsd-review-codex.md ]; then
306
- echo "Codex review failed or returned empty output. stderr:" > {run_dir}/gsd-review-codex.md
307
- cat {run_dir}/gsd-review-codex.err >> {run_dir}/gsd-review-codex.md
308
- fi
309
- ```
310
-
311
- **CodeRabbit:**
312
-
313
- Note: CodeRabbit reviews the current git diff/working tree — it does not accept a prompt or model flag. It may take up to 5 minutes. Use `timeout: 360000` on the Bash tool call. The source-grounding requirement in the build_prompt Review Instructions applies only to the prompt-fed reviewers above; CodeRabbit is a diff-only reviewer and never receives it. Treat its output as a diff observation, not a grounded plan-level verdict.
314
-
315
- ```bash
316
- coderabbit review --prompt-only 2>/dev/null > {run_dir}/gsd-review-coderabbit.md
317
- ```
318
-
319
- **OpenCode (via GitHub Copilot):**
320
-
321
- OpenCode's default `build` agent is an agentic coder, not a prompt→completion API.
322
- On a large review prompt it may run a few `read` tool calls and then end its turn
323
- with **zero output tokens** (`reason:"stop"`, `output:0`), so `--format default`
324
- yields empty stdout and the second reviewer is silently lost (#1936). Invoke with
325
- `--format json` and reconstruct the review from the assistant `text` parts; if the
326
- agent emitted none, surface the stop `reason`, output-token count, and captured
327
- stderr so the failure is diagnosable instead of a generic empty stub. Runs are also
328
- nondeterministic in length, so bound this Bash tool call with a wall-clock timeout —
329
- set `timeout: 660000` on the call (same mechanism the CodeRabbit block documents).
330
- That bound is hard: if it fires mid-`opencode run` the tool kills the command and
331
- the jq reconstruction below never runs, so the reviewing agent simply proceeds
332
- without an OpenCode result. The completing zero-output case — the actual #1936 bug —
333
- is fully handled below; the timeout only backstops the rarer nondeterministic hang. A
334
- reviewer instance with `"agent": "review"` (see
335
- `gsd-core/references/reviewer-instances.md`) sidesteps the default `build` agent and
336
- is the durable fix when this recurs.
337
-
338
- ```bash
339
- # stderr → sidecar (never /dev/null) so a real error is diagnosable — mirrors the
340
- # Codex block. --format json is the primary invocation (not a fallback): the review
341
- # text lives in assistant `text` parts, which the default formatter drops when the
342
- # agent stops with no final message (#1936).
343
- if [ -n "$OPENCODE_MODEL" ] && [ "$OPENCODE_MODEL" != "null" ]; then
344
- set -- --model "$OPENCODE_MODEL"
345
- else
346
- set --
347
- fi
348
- cat {run_dir}/gsd-review-prompt.md | opencode run "$@" --format json - 2>{run_dir}/gsd-review-opencode.err > {run_dir}/gsd-review-opencode.json
349
- # Reconstruct the review from the assistant text parts into a variable and test
350
- # its CONTENT, not the file size: an empty extraction still prints a trailing
351
- # newline that would fool a `[ -s file ]` check into skipping the stub.
352
- OPENCODE_REVIEW=$(jq -rs '[.[] | select(.type=="text") | .part.text // empty] | join("\n")' {run_dir}/gsd-review-opencode.json 2>/dev/null)
353
- if [ -n "$OPENCODE_REVIEW" ]; then
354
- printf '%s\n' "$OPENCODE_REVIEW" > {run_dir}/gsd-review-opencode.md
355
- else
356
- # No assistant text (no final message, or stdout was not valid JSON):
357
- {
358
- echo "OpenCode review returned no assistant text (#1936: agent ended its turn with no final message)."
359
- OPENCODE_DIAG=$(jq -rs '[.[] | select(.type=="step_finish")] | last | "stop reason=\(.part.reason // "?"), output tokens=\(.part.tokens.output // "?")"' {run_dir}/gsd-review-opencode.json 2>/dev/null)
360
- [ -n "$OPENCODE_DIAG" ] && echo "Diagnostic: $OPENCODE_DIAG"
361
- echo "stderr:"
362
- cat {run_dir}/gsd-review-opencode.err
363
- } > {run_dir}/gsd-review-opencode.md
364
- fi
365
- ```
366
-
367
- **Qwen Code:**
368
- ```bash
369
- cat {run_dir}/gsd-review-prompt.md | qwen - 2>/dev/null > {run_dir}/gsd-review-qwen.md
370
- if [ ! -s {run_dir}/gsd-review-qwen.md ]; then
371
- echo "Qwen review failed or returned empty output." > {run_dir}/gsd-review-qwen.md
372
- fi
373
- ```
374
-
375
- **Cursor:**
376
- ```bash
377
- # cursor-agent is a SEPARATE binary from the `cursor` IDE launcher; print mode (-p) takes the
378
- # prompt as an ARGUMENT, not stdin. A full review prompt can exceed the OS argument limit, so
379
- # reference the prompt file by path rather than inlining it. Capture stderr so a failure is
380
- # diagnosable instead of a silent empty result.
381
- # #2176: same absolute-root anchor as the Antigravity block — cursor-agent runs
382
- # in the repo cwd, but repo-relative references in the assembled prompt still
383
- # need an explicit root to resolve against. rev-parse (not bare pwd) so the
384
- # anchor is correct even when /gsd:review is invoked from a repo subdirectory.
385
- _CURSOR_ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
386
- CURSOR_PROMPT_ARG="Read the file at {run_dir}/gsd-review-prompt.md in full and carry out the review request it contains. The repository under review is at $_CURSOR_ROOT — resolve every relative file path in the review request against that absolute root. Output only the resulting markdown review. Do not edit any files."
387
- cursor-agent -p --mode ask --trust --output-format text "$CURSOR_PROMPT_ARG" 2>{run_dir}/gsd-review-cursor.err > {run_dir}/gsd-review-cursor.md
388
- if [ ! -s {run_dir}/gsd-review-cursor.md ]; then
389
- echo "Cursor review failed or returned empty output. stderr:" > {run_dir}/gsd-review-cursor.md
390
- cat {run_dir}/gsd-review-cursor.err >> {run_dir}/gsd-review-cursor.md
391
- fi
392
- ```
393
-
394
- **Antigravity CLI:**
395
-
396
- **Maintainer note — why this block has three layers (last updated against agy 1.0.16):**
397
-
398
- `agy -p` (the `--print` non-interactive flag) works correctly on macOS and Linux: it sends the
399
- prompt, receives the model response, and writes it to stdout. On **native Windows** it silently
400
- produces no stdout output despite the API call succeeding — a bug in `text_drip.go`'s non-TTY
401
- flush path, tracked at https://github.com/google-antigravity/antigravity-cli/issues/27466 and
402
- still open as of agy 1.0.2.
403
-
404
- Regardless of platform, `agy` always persists the full exchange to a transcript file on disk.
405
- The transcript fallback (Step 2 below) reads that file directly, giving Windows users full review
406
- coverage without any extra tooling. This pattern was first documented by the community MCP bridge
407
- at https://github.com/SinanTufekci/Claude-Code-Antigravity-CLI-MCP-Server — we inline the same
408
- logic here in pure bash/jq so no additional dependency is required.
409
-
410
- **Stale-response guard (why the pre-flight watermark matters):**
411
- Without a watermark, the fallback would read the last `PLANNER_RESPONSE` entry in the transcript
412
- regardless of when it was written — including entries from a previous invocation in the same
413
- workspace. To prevent that, we record the transcript's line count *before* calling `agy -p`. In
414
- the fallback, we only read lines appended after that count. If no new lines were written (agy
415
- failed before producing a response), `_AGY_RESULT` is empty and Step 3 fires — never stale. If
416
- the conv-id changed (agy started a fresh session), all lines in the new file are new and we use
417
- skip=0.
418
-
419
- **If the upstream stdout bug is fixed** (check the issue above): Step 2 silently becomes
420
- unreachable; stdout is non-empty and Step 1 handles it. No code change needed.
421
-
422
- **If the transcript paths change** in a future `agy` release: Step 2 silently becomes a no-op
423
- and Step 3 fires with a clear error message in REVIEWS.md. No silent corruption. To debug:
424
- - `~/.gemini/antigravity-cli/cache/last_conversations.json` — workspace → conv-id map
425
- - `~/.gemini/antigravity-cli/brain/<id>/.system_generated/logs/transcript.jsonl`
426
- Filter: `source=="MODEL"`, `status=="DONE"`, `type=="PLANNER_RESPONSE"`, take the last match's `content` field.
427
-
428
- Invocation specifics (verified agy 1.0.0, macOS arm64 and Linux amd64):
429
- - `-p` takes the prompt as a **flag value** — `echo X | agy -p` errors with "flag needs an argument: -p"
430
- - `--print-timeout` defaults to 5m, aligning with this workflow's global timeout
431
- - `--model "<name>"` selects the model (available since agy ~1.0.3; `agy models` lists
432
- them). When `review.models.agy` is set it is passed as `--model`; otherwise agy uses
433
- its persisted default (`agy models`).
434
-
435
- ```bash
436
- # Pre-flight: snapshot the transcript watermark before invoking agy.
437
- # Must run BEFORE agy -p — this is what prevents the fallback from reading a stale prior response.
438
- _AGY_WS=$(git rev-parse --show-toplevel 2>/dev/null || pwd)
439
- _AGY_CACHE="$HOME/.gemini/antigravity-cli/cache/last_conversations.json"
440
- _AGY_MARK_CONV=""
441
- _AGY_MARK_LINES=0
442
- if [ -f "$_AGY_CACHE" ]; then
443
- _AGY_MARK_CONV=$(jq -r --arg ws "$_AGY_WS" '
444
- .[$ws] //
445
- (to_entries
446
- | map(select(.key | ascii_downcase == ($ws | ascii_downcase)))
447
- | first | .value) //
448
- empty
449
- ' "$_AGY_CACHE" 2>/dev/null)
450
- if [ -n "$_AGY_MARK_CONV" ] && [ "$_AGY_MARK_CONV" != "null" ]; then
451
- _AGY_MARK_TX="$HOME/.gemini/antigravity-cli/brain/${_AGY_MARK_CONV}/.system_generated/logs/transcript.jsonl"
452
- [ -f "$_AGY_MARK_TX" ] && _AGY_MARK_LINES=$(wc -l < "$_AGY_MARK_TX" | tr -d ' ')
453
- fi
454
- fi
455
-
456
- # Step 1 — primary invocation: stdout works on macOS, Linux, and WSL.
457
- # Three hardening invariants (#2073), all mirroring the Cursor block's discipline:
458
- # * FILE-REFERENCE prompt (not inline `$(cat …)`) — a large review prompt (≈197 KB
459
- # for 6 plans + CONTEXT + RESEARCH + REQUIREMENTS) overflows the exec arg list
460
- # (`bash: agy: Argument list too long`, rc 126), indistinguishable from a model
461
- # failure when stderr is suppressed.
462
- # * EXTERNAL `timeout` wrapper when available (GNU `timeout` / `gtimeout`) —
463
- # `--print-timeout` is agy's native cap but it CANNOT fire before agy creates a
464
- # session; under concurrent heavy runs one process can stall pre-session (no
465
- # `brain/<conv-id>/` dir, alive at 583 s despite `--print-timeout 300s`). The
466
- # external cap bounds wall-clock regardless. Stock macOS lacks `timeout`, so
467
- # the block probes for it and falls back to --print-timeout alone there.
468
- # * `--model` from `review.models.agy` when set — escape hatch for a pinned model
469
- # that 404s server-side (exits 0 with empty stdout AND empty transcript).
470
- # * stdin tied to /dev/null so agy never blocks on a tty.
471
- # A non-zero exit (external timeout = 124, crash, etc.) discards any partial output
472
- # so the Step 2 transcript fallback / Step 3 diagnostic take over.
473
- if [ -n "$AGY_MODEL" ] && [ "$AGY_MODEL" != "null" ]; then
474
- set -- --model "$AGY_MODEL"
475
- else
476
- set --
477
- fi
478
- # #2176: grant the reviewer the repo under review. Without --add-dir, agy's
479
- # permission context never receives the cwd repo — the agent anchors on its own
480
- # ~/.gemini/antigravity-cli/scratch dir and reviews the plan text in isolation
481
- # (the exact failure the Review Instructions forbid). Capability-probed like the
482
- # Codex bypass flag so an older agy without --add-dir still runs; the prompt
483
- # anchor below keeps absolute-path reads possible on that fallback.
484
- if agy --help 2>/dev/null | grep -q -- '--add-dir'; then
485
- set -- "$@" --add-dir "$_AGY_WS"
486
- fi
487
- # #2176: anchor the prompt to the absolute repo root so repo-relative references
488
- # in the assembled review prompt resolve even on the no---add-dir fallback, and
489
- # require an explicit self-report if the reviewer still cannot read the repo.
490
- _AGY_PROMPT="Read the file at {run_dir}/gsd-review-prompt.md in full and carry out the review request it contains. The repository under review is at $_AGY_WS — resolve every relative file path in the review request against that absolute root and verify claims against those files. If you cannot read files under $_AGY_WS, begin your output with the exact line REVIEWED-WITHOUT-REPO-ACCESS before the review. Output only the resulting markdown review. Do not edit any files."
491
- # Capability-probe an external wall-clock killer (GNU coreutils `timeout` or the
492
- # macOS Homebrew `gtimeout`). Stock macOS ships NEITHER — a bare `timeout …` would
493
- # fail with rc 127 ("command not found") and silently lose the reviewer, so fall
494
- # back to agy's native --print-timeout alone in that case. The external cap, when
495
- # available, is set HIGHER than --print-timeout so it only backstops a pre-session
496
- # stall (which --print-timeout cannot bound — #2073 mode 3) and never pre-empts a
497
- # healthy run. Mirrors the probe in scripts/base64-scan.sh.
498
- _AGY_KILLER="$(command -v timeout 2>/dev/null || command -v gtimeout 2>/dev/null || true)"
499
- if [ -n "$_AGY_KILLER" ]; then
500
- "$_AGY_KILLER" 600 agy --print-timeout 540s "$@" -p "$_AGY_PROMPT" </dev/null 2>/dev/null > {run_dir}/gsd-review-antigravity.md
501
- else
502
- agy --print-timeout 540s "$@" -p "$_AGY_PROMPT" </dev/null 2>/dev/null > {run_dir}/gsd-review-antigravity.md
503
- fi
504
- _AGY_RC=$?
505
- if [ "$_AGY_RC" -ne 0 ]; then
506
- : > {run_dir}/gsd-review-antigravity.md
507
- fi
508
-
509
- # Step 2 — transcript fallback: catches Windows agy -p stdout bug (and any future stdout-silent edge cases).
510
- # Reads only lines appended AFTER the pre-flight watermark. If agy failed before writing a new response,
511
- # _AGY_RESULT is empty and Step 3 fires — no stale content can leak through.
512
- # Undocumented paths, verified agy 1.0.0–1.0.2. See maintainer note above if these break.
513
- if [ ! -s {run_dir}/gsd-review-antigravity.md ]; then
514
- if [ -f "$_AGY_CACHE" ]; then
515
- _AGY_CONV=$(jq -r --arg ws "$_AGY_WS" '
516
- .[$ws] //
517
- (to_entries
518
- | map(select(.key | ascii_downcase == ($ws | ascii_downcase)))
519
- | first | .value) //
520
- empty
521
- ' "$_AGY_CACHE" 2>/dev/null)
522
- if [ -n "$_AGY_CONV" ] && [ "$_AGY_CONV" != "null" ]; then
523
- _AGY_TX="$HOME/.gemini/antigravity-cli/brain/${_AGY_CONV}/.system_generated/logs/transcript.jsonl"
524
- if [ -f "$_AGY_TX" ]; then
525
- # If conv-id changed, agy started a new session — all lines are new, skip 0.
526
- # If same conv-id, only read lines beyond the watermark.
527
- [ "$_AGY_CONV" = "$_AGY_MARK_CONV" ] && _AGY_SKIP=$_AGY_MARK_LINES || _AGY_SKIP=0
528
- _AGY_RESULT=$(tail -n +"$((_AGY_SKIP + 1))" "$_AGY_TX" 2>/dev/null | \
529
- jq -r 'select(.source=="MODEL" and .status=="DONE" and .type=="PLANNER_RESPONSE") | .content' \
530
- 2>/dev/null | tail -1)
531
- [ -n "$_AGY_RESULT" ] && echo "$_AGY_RESULT" > {run_dir}/gsd-review-antigravity.md
532
- fi
533
- fi
534
- fi
535
- fi
536
-
537
- # Step 3 — final guard: both approaches yielded nothing (auth error, first-run setup,
538
- # path schema changed, 404'd pinned model, pre-session stall, etc.)
539
- if [ ! -s {run_dir}/gsd-review-antigravity.md ]; then
540
- {
541
- echo "Antigravity review failed or returned empty output."
542
- # #2073 mode 2: a pinned model that 404s exits 0 with empty stdout AND an empty
543
- # transcript — the only evidence is in agy's own log. Surface it instead of a
544
- # bare generic stub so the failure is diagnosable.
545
- _AGY_LOG="$HOME/.gemini/antigravity-cli/cli.log"
546
- if [ -f "$_AGY_LOG" ]; then
547
- _AGY_ERR=$(grep -iE 'agent executor error|NOT_FOUND|Publisher model' "$_AGY_LOG" | tail -3)
548
- if [ -n "$_AGY_ERR" ]; then
549
- echo "agy log hint (pinned model may be unavailable — run 'agy models' and set review.models.agy):"
550
- echo "$_AGY_ERR"
551
- fi
552
- fi
553
- # #2073 mode 3: pre-session stall tell — no new conversation dir appeared.
554
- echo "If no agy run started, that is the pre-session-stall case: check whether a new ~/.gemini/antigravity-cli/brain/<conv-id>/ dir appeared within ~30s of launch."
555
- } > {run_dir}/gsd-review-antigravity.md
556
- fi
557
-
558
- # #2176: blind-review marker. Two tells that the reviewer ran without repo
559
- # access: the prompt's mandated REVIEWED-WITHOUT-REPO-ACCESS self-report in the
560
- # first lines of output, or the agent DECLARING the scratch dir as its
561
- # workspace. Both patterns are anchored — the self-report to the head of the
562
- # file, the scratch tell to a workspace-declaration phrasing — so a grounded
563
- # review that merely QUOTES these strings (e.g. reviewing this very file) is
564
- # never mis-stamped. Stamp a machine-readable marker so the Consensus Summary
565
- # down-weights the review instead of counting an ungrounded verdict at full
566
- # weight. (Temp file + mv, no in-place sed — BSD/GNU safe.)
567
- if [ -s {run_dir}/gsd-review-antigravity.md ] && \
568
- { head -5 {run_dir}/gsd-review-antigravity.md | grep -q 'REVIEWED-WITHOUT-REPO-ACCESS' || \
569
- grep -qiE '(workspace|working) (directory|dir).{0,40}antigravity-cli/scratch' {run_dir}/gsd-review-antigravity.md; }; then
570
- {
571
- echo "> [reviewed-without-repo-access] This reviewer ran without visibility into the repo under review — down-weight its verdict in the Consensus Summary."
572
- echo ""
573
- cat {run_dir}/gsd-review-antigravity.md
574
- } > {run_dir}/gsd-review-antigravity.md.tmp && \
575
- mv {run_dir}/gsd-review-antigravity.md.tmp {run_dir}/gsd-review-antigravity.md
576
- fi
577
- ```
578
-
579
- **Ollama (local, OpenAI-compatible):**
580
-
581
- Read host and model from config. All three local backends share the same `/v1/chat/completions` endpoint — only host and model differ. Use `jq --rawfile` to safely encode the multi-line prompt as JSON without shell-escaping issues.
582
-
583
- ```bash
584
- # Shared helper: apply prompt-budget trimming for local reviewers
316
+ RUN_DIR="{run_dir}"
317
+ REPO_ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
318
+ # SELECTED_REVIEWERS is the comma-separated result of reviewer selection (ADR-0011 precedence:
319
+ # explicit flags > --all > review.default_reviewers > all detected). Unchanged by this phase.
320
+
321
+ # Shared budget-trim helper. Was defined inside the Ollama leg; it is lane-agnostic, so it is
322
+ # hoisted here now that any lane may declare a promptBudgetKey. Returns non-zero when the budget
323
+ # is too small for the minimum review set (prompt-budget exit 2 / 11).
585
324
  prepare_trimmed_prompt_for_reviewer() {
586
- REVIEWER_KEY="$1"
587
- REVIEWER_BUDGET="$2"
588
- OUTPUT_PROMPT="$3"
589
- OUTPUT_META="$4"
590
-
591
- [ -z "$REVIEWER_BUDGET" ] && return 0
592
- [ "$REVIEWER_BUDGET" = "null" ] && return 0
593
- [ "$REVIEWER_BUDGET" = "0" ] && return 0
325
+ REVIEWER_KEY="$1"; REVIEWER_BUDGET="$2"; OUTPUT_PROMPT="$3"; OUTPUT_META="$4"
594
326
 
595
327
  PLAN_FILE_ARGS=""
596
- for p in {run_dir}/gsd-review-plan-*.md; do
328
+ for p in "$RUN_DIR"/gsd-review-plan-*.md; do
597
329
  [ -f "$p" ] && PLAN_FILE_ARGS="$PLAN_FILE_ARGS --plan-file $p"
598
330
  done
599
331
  PROJECT_ARG=""
600
- [ -f "{run_dir}/gsd-review-project.md" ] && PROJECT_ARG="--project-file {run_dir}/gsd-review-project.md"
332
+ [ -f "$RUN_DIR/gsd-review-project.md" ] && PROJECT_ARG="--project-file $RUN_DIR/gsd-review-project.md"
601
333
  CONTEXT_ARG=""
602
- [ -f "{run_dir}/gsd-review-context.md" ] && CONTEXT_ARG="--context-file {run_dir}/gsd-review-context.md"
334
+ [ -f "$RUN_DIR/gsd-review-context.md" ] && CONTEXT_ARG="--context-file $RUN_DIR/gsd-review-context.md"
603
335
  RESEARCH_ARG=""
604
- [ -f "{run_dir}/gsd-review-research.md" ] && RESEARCH_ARG="--research-file {run_dir}/gsd-review-research.md"
336
+ [ -f "$RUN_DIR/gsd-review-research.md" ] && RESEARCH_ARG="--research-file $RUN_DIR/gsd-review-research.md"
605
337
  REQUIREMENTS_ARG=""
606
- [ -f "{run_dir}/gsd-review-requirements.md" ] && REQUIREMENTS_ARG="--requirements-file {run_dir}/gsd-review-requirements.md"
338
+ [ -f "$RUN_DIR/gsd-review-requirements.md" ] && REQUIREMENTS_ARG="--requirements-file $RUN_DIR/gsd-review-requirements.md"
607
339
 
608
340
  gsd_run query prompt-budget \
609
341
  --budget "$REVIEWER_BUDGET" \
610
- --instructions-file "{run_dir}/gsd-review-instructions.md" \
611
- --roadmap-file "{run_dir}/gsd-review-roadmap.md" \
342
+ --instructions-file "$RUN_DIR/gsd-review-instructions.md" \
343
+ --roadmap-file "$RUN_DIR/gsd-review-roadmap.md" \
612
344
  $PLAN_FILE_ARGS $PROJECT_ARG $CONTEXT_ARG $RESEARCH_ARG $REQUIREMENTS_ARG \
613
345
  --output-prompt "$OUTPUT_PROMPT" \
614
346
  --output-metadata "$OUTPUT_META"
615
347
  return $?
616
348
  }
617
349
 
618
- # Resolve prompt budget for Ollama: per-reviewer override > global default > null
619
- OLLAMA_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens_per_reviewer.ollama 2>/dev/null | jq -r '.' 2>/dev/null || echo "null")
620
- if [ -z "$OLLAMA_REVIEWER_BUDGET" ] || [ "$OLLAMA_REVIEWER_BUDGET" = "null" ]; then
621
- OLLAMA_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens 2>/dev/null | jq -r '.' 2>/dev/null || echo "null")
622
- fi
623
-
624
- # Apply budget trim for Ollama if a budget is configured
625
- OLLAMA_PROMPT_FILE="{run_dir}/gsd-review-prompt.md"
626
- OLLAMA_SKIP=0
627
- if [ -n "$OLLAMA_REVIEWER_BUDGET" ] && [ "$OLLAMA_REVIEWER_BUDGET" != "null" ] && [ "$OLLAMA_REVIEWER_BUDGET" != "0" ]; then
628
- OLLAMA_TRIMMED_PROMPT="{run_dir}/gsd-review-prompt-ollama.md"
629
- OLLAMA_TRIM_META="{run_dir}/gsd-review-prompt-ollama.metadata.json"
630
- prepare_trimmed_prompt_for_reviewer "ollama" "$OLLAMA_REVIEWER_BUDGET" "$OLLAMA_TRIMMED_PROMPT" "$OLLAMA_TRIM_META"
631
- OLLAMA_EXIT=$?
632
- if [ $OLLAMA_EXIT -ne 0 ]; then
633
- if [ $OLLAMA_EXIT -eq 2 ] || [ $OLLAMA_EXIT -eq 11 ]; then
634
- echo "WARNING: prompt budget for ollama (${OLLAMA_REVIEWER_BUDGET} tokens) is too small for the minimum review set. Skipping Ollama reviewer." >&2
635
- else
636
- echo "WARNING: prompt-budget returned unexpected exit code ${OLLAMA_EXIT} for ollama. Skipping Ollama reviewer." >&2
637
- fi
638
- OLLAMA_SKIP=1
639
- else
640
- OLLAMA_PROMPT_FILE="$OLLAMA_TRIMMED_PROMPT"
641
- fi
642
- fi
643
-
644
- if [ "$OLLAMA_SKIP" != "1" ]; then
645
- OLLAMA_HOST=$(gsd_run query config-get review.ollama_host 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
646
- if [ -z "$OLLAMA_HOST" ] || [ "$OLLAMA_HOST" = "null" ]; then OLLAMA_HOST="http://localhost:11434"; fi
647
- OLLAMA_MODEL=$(gsd_run query config-get review.models.ollama 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
648
- if [ -z "$OLLAMA_MODEL" ] || [ "$OLLAMA_MODEL" = "null" ]; then
649
- OLLAMA_MODEL=$(curl -s --max-time 2 "${OLLAMA_HOST}/v1/models" 2>/dev/null | jq -r '.data[0].id // "llama3"' 2>/dev/null || echo "llama3")
650
- fi
651
- jq -n --rawfile content "$OLLAMA_PROMPT_FILE" \
652
- --arg model "$OLLAMA_MODEL" \
653
- '{model: $model, messages: [{role: "user", content: $content}]}' | \
654
- curl -s --max-time 120 -X POST "${OLLAMA_HOST}/v1/chat/completions" \
655
- -H "Content-Type: application/json" -d @- 2>/dev/null | \
656
- jq -r '.choices[0].message.content // "Ollama review failed or returned empty output."' \
657
- > {run_dir}/gsd-review-ollama.md
658
- if [ ! -s {run_dir}/gsd-review-ollama.md ]; then
659
- echo "Ollama review failed or returned empty output." > {run_dir}/gsd-review-ollama.md
660
- fi
661
- fi
662
- ```
663
-
664
- **LM Studio (local, OpenAI-compatible):**
665
- ```bash
666
- # Resolve prompt budget for LM Studio: per-reviewer override > global default > null
667
- LM_STUDIO_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens_per_reviewer.lm_studio 2>/dev/null | jq -r '.' 2>/dev/null || echo "null")
668
- if [ -z "$LM_STUDIO_REVIEWER_BUDGET" ] || [ "$LM_STUDIO_REVIEWER_BUDGET" = "null" ]; then
669
- LM_STUDIO_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens 2>/dev/null | jq -r '.' 2>/dev/null || echo "null")
670
- fi
671
-
672
- # Apply budget trim for LM Studio if a budget is configured
673
- LM_STUDIO_PROMPT_FILE="{run_dir}/gsd-review-prompt.md"
674
- LM_STUDIO_SKIP=0
675
- if [ -n "$LM_STUDIO_REVIEWER_BUDGET" ] && [ "$LM_STUDIO_REVIEWER_BUDGET" != "null" ] && [ "$LM_STUDIO_REVIEWER_BUDGET" != "0" ]; then
676
- LM_STUDIO_TRIMMED_PROMPT="{run_dir}/gsd-review-prompt-lm_studio.md"
677
- LM_STUDIO_TRIM_META="{run_dir}/gsd-review-prompt-lm_studio.metadata.json"
678
- prepare_trimmed_prompt_for_reviewer "lm_studio" "$LM_STUDIO_REVIEWER_BUDGET" "$LM_STUDIO_TRIMMED_PROMPT" "$LM_STUDIO_TRIM_META"
679
- LM_STUDIO_EXIT=$?
680
- if [ $LM_STUDIO_EXIT -ne 0 ]; then
681
- if [ $LM_STUDIO_EXIT -eq 2 ] || [ $LM_STUDIO_EXIT -eq 11 ]; then
682
- echo "WARNING: prompt budget for lm_studio (${LM_STUDIO_REVIEWER_BUDGET} tokens) is too small for the minimum review set. Skipping LM Studio reviewer." >&2
350
+ gsd_run query review-lane plan \
351
+ --selected "$SELECTED_REVIEWERS" --run-dir "$RUN_DIR" --repo-root "$REPO_ROOT" --json \
352
+ > "$RUN_DIR/gsd-review-lanes.json"
353
+
354
+ for SLUG in $(echo "$SELECTED_REVIEWERS" | tr ',' ' '); do
355
+ # Per-lane prompt budget. The lane declares its own `promptBudgetKey`; `plan` resolved it,
356
+ # applying #2797's sentinel rule (-1 = unset → fall back to the global budget; 0 legitimately
357
+ # means "do not trim this lane"). Trimming itself stays in prompt-budget, which owns it.
358
+ LANE_BUDGET=$(gsd_run query review-lane plan --selected "$SLUG" --run-dir "$RUN_DIR" \
359
+ --repo-root "$REPO_ROOT" --json 2>/dev/null \
360
+ | sed -n 's/.*"promptBudget": *\([0-9-]*\).*/\1/p' | head -1)
361
+ PROMPT_ARG=""
362
+ if [ -n "$LANE_BUDGET" ] && [ "$LANE_BUDGET" != "null" ] && [ "$LANE_BUDGET" -gt 0 ] 2>/dev/null; then
363
+ TRIMMED="$RUN_DIR/gsd-review-prompt-$SLUG.md"
364
+ if prepare_trimmed_prompt_for_reviewer "$SLUG" "$LANE_BUDGET" "$TRIMMED" \
365
+ "$RUN_DIR/gsd-review-prompt-$SLUG.metadata.json"; then
366
+ PROMPT_ARG="--prompt-file $TRIMMED"
683
367
  else
684
- echo "WARNING: prompt-budget returned unexpected exit code ${LM_STUDIO_EXIT} for lm_studio. Skipping LM Studio reviewer." >&2
368
+ # A budget too small for the minimum review set drops the lane just as silently as an empty
369
+ # response used to (#2605), so leave the skip visible in the review output, not only on stderr.
370
+ echo "$SLUG review skipped: prompt budget (${LANE_BUDGET} tokens) too small for the minimum review set." \
371
+ > "$RUN_DIR/gsd-review-$SLUG.md"
372
+ continue
685
373
  fi
686
- LM_STUDIO_SKIP=1
687
- else
688
- LM_STUDIO_PROMPT_FILE="$LM_STUDIO_TRIMMED_PROMPT"
689
374
  fi
690
- fi
691
375
 
692
- if [ "$LM_STUDIO_SKIP" != "1" ]; then
693
- LM_STUDIO_HOST=$(gsd_run query config-get review.lm_studio_host 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
694
- if [ -z "$LM_STUDIO_HOST" ] || [ "$LM_STUDIO_HOST" = "null" ]; then LM_STUDIO_HOST="http://localhost:1234"; fi
695
- LM_STUDIO_MODEL=$(gsd_run query config-get review.models.lm_studio 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
696
- if [ -z "$LM_STUDIO_MODEL" ] || [ "$LM_STUDIO_MODEL" = "null" ]; then
697
- LM_STUDIO_MODEL=$(curl -s --max-time 2 "${LM_STUDIO_HOST}/v1/models" 2>/dev/null | jq -r '.data[0].id // "local-model"' 2>/dev/null || echo "local-model")
698
- fi
699
- LM_STUDIO_RESPONSE=$(jq -n --rawfile content "$LM_STUDIO_PROMPT_FILE" \
700
- --arg model "$LM_STUDIO_MODEL" \
701
- '{model: $model, messages: [{role: "user", content: $content}]}' | \
702
- curl -s --max-time 120 -X POST "${LM_STUDIO_HOST}/v1/chat/completions" \
703
- -H "Content-Type: application/json" -d @- 2>/dev/null)
704
- LM_STUDIO_ACTUAL_MODEL=$(echo "$LM_STUDIO_RESPONSE" | jq -r '.model // ""' 2>/dev/null || echo "")
705
- if [ -n "$LM_STUDIO_ACTUAL_MODEL" ] && [ "$LM_STUDIO_ACTUAL_MODEL" != "null" ] && [ "$LM_STUDIO_ACTUAL_MODEL" != "$LM_STUDIO_MODEL" ]; then
706
- echo "Warning: LM Studio served model '$LM_STUDIO_ACTUAL_MODEL' but '$LM_STUDIO_MODEL' was requested. Review may be from a different model." >&2
707
- fi
708
- LM_STUDIO_CONTENT=$(echo "$LM_STUDIO_RESPONSE" | jq -r '.choices[0].message.content // ""' 2>/dev/null || echo "")
709
- if [ -n "$LM_STUDIO_CONTENT" ]; then
710
- echo "$LM_STUDIO_CONTENT" > {run_dir}/gsd-review-lm_studio.md
711
- else
712
- echo "Warning: LM Studio returned empty content — skipping review." >&2
713
- fi
714
- fi
376
+ # One invocation, whatever the lane's transport, prompt channel, output channel or handler.
377
+ # `--explicit` marks a lane the user NAMED: ADR-2782 D4 — not finding a lane nobody asked for is
378
+ # normal, failing to run one somebody asked for is an error.
379
+ gsd_run query review-lane invoke --slug "$SLUG" \
380
+ --run-dir "$RUN_DIR" --repo-root "$REPO_ROOT" $PROMPT_ARG $EXPLICIT_FLAG --json \
381
+ >> "$RUN_DIR/gsd-review-lane-results.jsonl"
382
+ done
715
383
  ```
716
384
 
717
- **llama.cpp (local, OpenAI-compatible):**
718
- ```bash
719
- # Resolve prompt budget for llama.cpp: per-reviewer override > global default > null
720
- LLAMA_CPP_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens_per_reviewer.llama_cpp 2>/dev/null | jq -r '.' 2>/dev/null || echo "null")
721
- if [ -z "$LLAMA_CPP_REVIEWER_BUDGET" ] || [ "$LLAMA_CPP_REVIEWER_BUDGET" = "null" ]; then
722
- LLAMA_CPP_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens 2>/dev/null | jq -r '.' 2>/dev/null || echo "null")
723
- fi
724
-
725
- # Apply budget trim for llama.cpp if a budget is configured
726
- LLAMA_CPP_PROMPT_FILE="{run_dir}/gsd-review-prompt.md"
727
- LLAMA_CPP_SKIP=0
728
- if [ -n "$LLAMA_CPP_REVIEWER_BUDGET" ] && [ "$LLAMA_CPP_REVIEWER_BUDGET" != "null" ] && [ "$LLAMA_CPP_REVIEWER_BUDGET" != "0" ]; then
729
- LLAMA_CPP_TRIMMED_PROMPT="{run_dir}/gsd-review-prompt-llama_cpp.md"
730
- LLAMA_CPP_TRIM_META="{run_dir}/gsd-review-prompt-llama_cpp.metadata.json"
731
- prepare_trimmed_prompt_for_reviewer "llama_cpp" "$LLAMA_CPP_REVIEWER_BUDGET" "$LLAMA_CPP_TRIMMED_PROMPT" "$LLAMA_CPP_TRIM_META"
732
- LLAMA_CPP_EXIT=$?
733
- if [ $LLAMA_CPP_EXIT -ne 0 ]; then
734
- if [ $LLAMA_CPP_EXIT -eq 2 ] || [ $LLAMA_CPP_EXIT -eq 11 ]; then
735
- echo "WARNING: prompt budget for llama_cpp (${LLAMA_CPP_REVIEWER_BUDGET} tokens) is too small for the minimum review set. Skipping llama.cpp reviewer." >&2
736
- else
737
- echo "WARNING: prompt-budget returned unexpected exit code ${LLAMA_CPP_EXIT} for llama_cpp. Skipping llama.cpp reviewer." >&2
738
- fi
739
- LLAMA_CPP_SKIP=1
740
- else
741
- LLAMA_CPP_PROMPT_FILE="$LLAMA_CPP_TRIMMED_PROMPT"
742
- fi
743
- fi
744
-
745
- if [ "$LLAMA_CPP_SKIP" != "1" ]; then
746
- LLAMA_CPP_HOST=$(gsd_run query config-get review.llama_cpp_host 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
747
- if [ -z "$LLAMA_CPP_HOST" ] || [ "$LLAMA_CPP_HOST" = "null" ]; then LLAMA_CPP_HOST="http://localhost:8080"; fi
748
- LLAMA_CPP_MODEL=$(gsd_run query config-get review.models.llama_cpp 2>/dev/null | jq -r '.' 2>/dev/null || echo "")
749
- if [ -z "$LLAMA_CPP_MODEL" ] || [ "$LLAMA_CPP_MODEL" = "null" ]; then
750
- LLAMA_CPP_MODEL=$(curl -s --max-time 2 "${LLAMA_CPP_HOST}/v1/models" 2>/dev/null | jq -r '.data[0].id // "local-model"' 2>/dev/null || echo "local-model")
751
- fi
752
- LLAMA_CPP_CONTENT=$(jq -n --rawfile content "$LLAMA_CPP_PROMPT_FILE" \
753
- --arg model "$LLAMA_CPP_MODEL" \
754
- '{model: $model, messages: [{role: "user", content: $content}]}' | \
755
- curl -s --max-time 120 -X POST "${LLAMA_CPP_HOST}/v1/chat/completions" \
756
- -H "Content-Type: application/json" -d @- 2>/dev/null | \
757
- jq -r '.choices[0].message.content // ""' 2>/dev/null || echo "")
758
- if [ -n "$LLAMA_CPP_CONTENT" ]; then
759
- echo "$LLAMA_CPP_CONTENT" > {run_dir}/gsd-review-llama_cpp.md
760
- else
761
- echo "Warning: llama.cpp returned empty content — skipping review." >&2
762
- fi
763
- fi
764
- ```
385
+ Each lane leaves `{run_dir}/gsd-review-<slug>.md` — its review, or a diagnostic stub carrying the
386
+ captured stderr (and, for an OpenAI-compatible lane, the raw response body, where such a server puts
387
+ its error JSON on an HTTP 4xx/5xx while still exiting 0). A stub is never mistaken for a clean
388
+ review: it keeps its "failed or returned empty output" header (#2494/#2605/#2794).
765
389
 
766
- If a CLI or local server fails, log the error and continue with remaining reviewers.
390
+ A lane that will not run reports a typed reason rather than an empty file — `missing_binary`,
391
+ `probe_failed`, `probe_timeout`, `missing_required_binary`, `host_unreachable`,
392
+ `egress_host_changed`, `unknown_handler`, `budget_too_small`. **`egress_host_changed` means the lane
393
+ was consented to send plans to one destination and `.planning/config.json` now names another; it is
394
+ blocked, not silently redirected** (ADR-2782 D5).
767
395
 
768
396
  Display progress:
769
397
  ```
@@ -806,83 +434,26 @@ trimmed_reviewers: # only present if at least one reviewer was trimmed
806
434
 
807
435
  # Cross-AI Plan Review — Phase {N}
808
436
 
809
- ## Gemini Review
810
-
811
- {gemini review content}
437
+ <!-- Sections are RENDERED from each lane's declared `reviewsSection`, in descriptor order.
438
+ There is deliberately no hardcoded per-reviewer heading list here any more: a hand-maintained
439
+ list is exactly the drift #2781 was filed about, and it silently disagreed with the roster.
440
+ `gsd_run query review-lane sections --selected "$SELECTED_REVIEWERS"` emits
441
+ `<slug><TAB><reviewsSection>` in order; for each row, emit:
812
442
 
813
- ---
814
-
815
- ## Claude Review
443
+ ## <reviewsSection> Review
816
444
 
817
- {claude review content}
445
+ {contents of {run_dir}/gsd-review-<slug>.md}
818
446
 
819
- ---
447
+ ---
820
448
 
821
- ## Codex Review
449
+ Two headings must NOT be generated from this list, because they are not lanes:
450
+ * `## <Adapter> Review (<instance>)` — an ADR-1517 reviewer INSTANCE resolves THROUGH a lane
451
+ and is rendered from the instance list, not the lane list (ADR-2782 D8).
452
+ * `## Consensus Summary` — not a review section at all.
822
453
 
823
- {codex review content}
824
-
825
- ---
826
-
827
- ## CodeRabbit Review
828
-
829
- {coderabbit review content}
830
-
831
- ---
832
-
833
- ## OpenCode Review
834
-
835
- {opencode review content}
836
-
837
- ---
838
-
839
- ## OpenCode Review (opencode-deepseek)
840
-
841
- {opencode-deepseek instance review content — only present when this instance was selected}
842
-
843
- ---
844
-
845
- ## OpenCode Review (opencode-mimo)
846
-
847
- {opencode-mimo instance review content — only present when this instance was selected}
848
-
849
- ---
850
-
851
- ## Qwen Review
852
-
853
- {qwen review content}
854
-
855
- ---
856
-
857
- ## Cursor Review
858
-
859
- {cursor review content}
860
-
861
- ---
862
-
863
- ## Antigravity Review
864
-
865
- {antigravity review content}
866
-
867
- ---
868
-
869
- ## Ollama Review
870
-
871
- {ollama review content}
872
-
873
- ---
874
-
875
- ## LM Studio Review
876
-
877
- {lm_studio review content}
878
-
879
- ---
880
-
881
- ## llama.cpp Review
882
-
883
- {llama_cpp review content}
884
-
885
- ---
454
+ A lane whose `evidenceClass` is `diff-only` (CodeRabbit) carries its caveat from data: it never
455
+ received the source-grounding prompt, so its verdict is folded in as a diff observation and is
456
+ not weighted as a grounded plan review. -->
886
457
 
887
458
  ## Consensus Summary
888
459