@opengsd/gsd-core 1.5.0-rc.4 → 1.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (89) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/LICENSE +1 -1
  3. package/agents/gsd-executor.md +22 -0
  4. package/agents/gsd-phase-researcher.md +1 -1
  5. package/agents/gsd-planner.md +4 -0
  6. package/agents/gsd-project-researcher.md +1 -1
  7. package/agents/gsd-verifier.md +45 -13
  8. package/bin/install.js +56 -100
  9. package/commands/gsd/autonomous.md +1 -1
  10. package/commands/gsd/execute-phase.md +1 -1
  11. package/commands/gsd/plan-phase.md +1 -1
  12. package/gemini-extension.json +1 -1
  13. package/gsd-core/bin/gsd-tools.cjs +85 -53
  14. package/gsd-core/bin/lib/agent-command-router.cjs +2 -2
  15. package/gsd-core/bin/lib/agent-install-check.cjs +143 -0
  16. package/gsd-core/bin/lib/audit-command-router.cjs +4 -4
  17. package/gsd-core/bin/lib/capability-activation.cjs +37 -10
  18. package/gsd-core/bin/lib/capability-registry.cjs +2 -0
  19. package/gsd-core/bin/lib/capability-state.cjs +80 -30
  20. package/gsd-core/bin/lib/capability-writer.cjs +5 -4
  21. package/gsd-core/bin/lib/check-command-router.cjs +50 -51
  22. package/gsd-core/bin/lib/commands.cjs +20 -2
  23. package/gsd-core/bin/lib/config-loader.cjs +3 -4
  24. package/gsd-core/bin/lib/config-schema.cjs +1 -1
  25. package/gsd-core/bin/lib/config-types.cjs +2 -1
  26. package/gsd-core/bin/lib/config.cjs +5 -2
  27. package/gsd-core/bin/lib/decisions.cjs +19 -1
  28. package/gsd-core/bin/lib/docs.cjs +14 -2
  29. package/gsd-core/bin/lib/frontmatter.cjs +2 -2
  30. package/gsd-core/bin/lib/gap-checker.cjs +5 -2
  31. package/gsd-core/bin/lib/git-base-branch.cjs +27 -1
  32. package/gsd-core/bin/lib/graphify-command-router.cjs +6 -8
  33. package/gsd-core/bin/lib/graphify.cjs +7 -31
  34. package/gsd-core/bin/lib/gsd2-import.cjs +2 -2
  35. package/gsd-core/bin/lib/init.cjs +30 -4
  36. package/gsd-core/bin/lib/intel-command-router.cjs +6 -3
  37. package/gsd-core/bin/lib/intel.cjs +28 -34
  38. package/gsd-core/bin/lib/io.cjs +2 -4
  39. package/gsd-core/bin/lib/learnings.cjs +2 -2
  40. package/gsd-core/bin/lib/loop-resolver.cjs +45 -167
  41. package/gsd-core/bin/lib/milestone.cjs +13 -5
  42. package/gsd-core/bin/lib/model-resolver.cjs +3 -4
  43. package/gsd-core/bin/lib/phase-id.cjs +3 -5
  44. package/gsd-core/bin/lib/phase-locator.cjs +3 -6
  45. package/gsd-core/bin/lib/phase.cjs +59 -13
  46. package/gsd-core/bin/lib/probe-core.cjs +40 -11
  47. package/gsd-core/bin/lib/profile-output.cjs +5 -2
  48. package/gsd-core/bin/lib/prohibition-enforcement.cjs +660 -0
  49. package/gsd-core/bin/lib/roadmap-command-router.cjs +2 -2
  50. package/gsd-core/bin/lib/roadmap-parser.cjs +28 -25
  51. package/gsd-core/bin/lib/roadmap.cjs +9 -4
  52. package/gsd-core/bin/lib/runtime-artifact-conversion.cjs +4 -1
  53. package/gsd-core/bin/lib/runtime-hooks-surface.cjs +24 -3
  54. package/gsd-core/bin/lib/state.cjs +75 -17
  55. package/gsd-core/bin/lib/task-command-router.cjs +2 -2
  56. package/gsd-core/bin/lib/teams-status.cjs +74 -0
  57. package/gsd-core/bin/lib/template.cjs +11 -2
  58. package/gsd-core/bin/lib/uat.cjs +64 -2
  59. package/gsd-core/bin/lib/verification.cjs +8 -5
  60. package/gsd-core/bin/lib/verify.cjs +311 -4
  61. package/gsd-core/bin/lib/workstream-inventory.cjs +2 -2
  62. package/gsd-core/bin/lib/workstream.cjs +8 -2
  63. package/gsd-core/bin/lib/worktree-safety.cjs +44 -3
  64. package/gsd-core/bin/shared/config-schema.manifest.json +2 -0
  65. package/gsd-core/references/planner-antipatterns.md +46 -0
  66. package/gsd-core/references/planning-config.md +5 -1
  67. package/gsd-core/references/prohibition-probe.md +80 -2
  68. package/gsd-core/references/worktree-branch-check.md +11 -5
  69. package/gsd-core/templates/verification-report.md +16 -3
  70. package/gsd-core/workflows/docs-update.md +23 -31
  71. package/gsd-core/workflows/execute-phase/steps/worktree-recovery-policy.md +9 -0
  72. package/gsd-core/workflows/execute-phase.md +7 -2
  73. package/gsd-core/workflows/map-codebase.md +8 -10
  74. package/gsd-core/workflows/plan-phase.md +6 -0
  75. package/gsd-core/workflows/quick.md +26 -2
  76. package/gsd-core/workflows/settings-advanced.md +5 -5
  77. package/gsd-core/workflows/settings-integrations.md +5 -5
  78. package/gsd-core/workflows/spec-phase.md +30 -2
  79. package/gsd-core/workflows/verify-phase.md +22 -7
  80. package/hooks/dist/gsd-worktree-path-guard.js +34 -18
  81. package/hooks/gsd-worktree-path-guard.js +34 -18
  82. package/package.json +3 -3
  83. package/scripts/ci-prepare-test-scope.cjs +56 -14
  84. package/scripts/diff-touches-shipped-paths.cjs +5 -11
  85. package/scripts/gen-capability-registry.cjs +27 -1
  86. package/scripts/lint-allow-test-rule-refs.allowlist.json +0 -2
  87. package/scripts/research-profiles.cjs +2 -2
  88. package/scripts/run-tests.cjs +1 -0
  89. package/gsd-core/bin/lib/core.cjs +0 -345
@@ -110,20 +110,98 @@ lifecycle is identical to the edge-probe, the verification tiers differ):
110
110
  human/LLM judgment that the framing is not manipulative). It records intent and routes
111
111
  to a judgment-based review rather than a green/red test.
112
112
 
113
+ At verify time these tiers are routed differently (ADR-550 D4):
114
+ - A **test**-tier prohibition is enforced + hard-gates via the deterministic
115
+ `check prohibition-enforcement` sub-command (#1259 + #1279, ADR-550 D5d): it locates the wired
116
+ mechanical check (a `node --test` negative test OR a lint/AST rule run as
117
+ `eslint --format json` and filtered by `ruleId`), **machine-proves it is fail-first** against a
118
+ known violation, runs it for a genuine **non-vacuous** pass, and emits the
119
+ `dispositionForProhibition()` verdict. A passing, fail-first-proven wired check disposes **green**
120
+ (satisfiable → can reach `passed`); a missing, un-provable, or genuinely-non-passing check
121
+ **hard-gates** (flagged, never green → `gaps_found`) in BOTH interactive and autonomous modes —
122
+ never a silent pass.
123
+
124
+ **Machine-proven fail-first (#1279).** `failFirst` is now **machine-proven, not caller-attested**:
125
+ before a clean pass greens, the producer independently runs the wired check against a KNOWN
126
+ VIOLATION and confirms it goes RED (any other outcome — passes-on-violation, can't-prove, throws,
127
+ times out, no violation source — hard-gates). The violation is sourced from a descriptor field:
128
+ - **`violationFixture`** — an author-supplied path to a KNOWN-BAD subject. For `lint-rule`: a file
129
+ whose content violates `rule`; the prover lints it and requires the rule id to appear in the
130
+ JSON report (the rule must have teeth). For `node-test`: a subject the negative test exercises,
131
+ expected to drive it RED.
132
+ - **`GSD_PROHIB_SUBJECT`** — the node-test subject-injection convention: the producer spawns the
133
+ negative test with `GSD_PROHIB_SUBJECT=<violationFixture>` in the child env; the test reads that
134
+ env var to locate its subject-under-test and is expected to go red against the violating subject.
135
+ - **lint-fixture authoring gotcha** — the violating fixture must actually trigger the rule. For the
136
+ `local/no-source-grep` dogfood anchor specifically, use the `path.join('lib','foo.cjs')` form
137
+ (a standalone quoted dir token); a single string literal like `'src/x.cjs'` does NOT trigger the
138
+ rule, so a mis-authored fixture makes the prover report "not proven" and hard-gate a legitimately
139
+ wired check. (`no-source-grep` has no filename guard — any `.cjs` with the pattern fires.)
140
+
141
+ > **PROPOSED, renamable conventions (zero live consumers).** Both `GSD_PROHIB_SUBJECT` and
142
+ > `violationFixture` are net-new surface with **no live in-tree consumer yet** — there is no
143
+ > in-tree `node --test` prohibition; node-test fail-first proof is exercised only by SYNTHETIC
144
+ > temp fixtures in the tests, and the real dogfood remains the LINT-rule `local/no-source-grep`.
145
+ > They are therefore **open to maintainer adjustment (rename, or replacing the env var with an
146
+ > argv) at PR review with zero migration cost.** See the ADR-550 2026-06-15 addendum (#1279).
147
+ - A **judgment**-tier prohibition routes to a never-silent / never-hard-halt soft gate
148
+ (autonomous emits an `unverified-prohibition — human review recommended` flag).
149
+
113
150
  Splitting these axes keeps the lifecycle enum free of a verification fact and lets the
114
151
  prohibition adapter declare `test | judgment` without forking the shared lifecycle enum that
115
152
  the edge-probe's `explicit | backstop` also uses.
116
153
 
154
+ ## Optional wired-check descriptor (deterministic locate + machine-proof, #1278 + #1346)
155
+
156
+ A `resolved`/`test`-tier prohibition MAY carry an **optional `check` descriptor** that names
157
+ the wired mechanical check, so verify-phase locates it deterministically instead of inventing
158
+ `{kind, target, rule}` each run. The descriptor is captured at spec-phase (soft / optional —
159
+ the author wires it when the negative test or lint rule already exists) and is represented as
160
+ **four flat scalar keys** on the `must_haves.prohibitions` item — never a nested `check: {}`
161
+ object:
162
+
163
+ - `check_kind` — `node-test` | `lint-rule` (which producer mechanism runs the check).
164
+ - `check_target` — the test file (`node-test`) or the file the rule runs against (`lint-rule`).
165
+ - `check_rule` — the `ruleId` to filter on, **lint-rule only** (absent for `node-test`).
166
+ - `check_violation_fixture` — path to a KNOWN-BAD subject the #1279 prover runs the check against to
167
+ machine-prove fail-first (rides BOTH kinds; for `node-test` it is injected via `GSD_PROHIB_SUBJECT`).
168
+
169
+ The flat-scalar shape is load-bearing: the shared `parseMustHavesBlock` is a flat parser and a
170
+ nested object would flatten/mangle the round-trip (ADR-550 2026-06-15 addendum; #644 "no parser
171
+ rewrite" precedent). `projectProhibitions` emits these keys **only for a well-formed descriptor**
172
+ (valid `check_kind` + non-empty `check_target`; `check_rule` only on the lint-rule path;
173
+ `check_violation_fixture` only when non-empty), and verify-phase reads them back via
174
+ `descriptorFromProjection` into the `CheckDescriptor` handed to `check prohibition-enforcement`. This
175
+ closes **both** the locate (#1278) and the machine-proof-fixture (#1346) halves with **zero manual
176
+ descriptor authoring**: a prohibition authored with all four scalars greens end-to-end through the
177
+ projection alone.
178
+
179
+ **Fail-closed + backward-compat.** A partial descriptor (`lint-rule` missing `check_rule`), an
180
+ unknown `check_kind`, an **absent** descriptor, OR a descriptor with **no `check_violation_fixture`**
181
+ falls through to the producer's fail-closed paths (`located: false`, or located-but-unprovable) —
182
+ never a silent green. A prohibition with no descriptor parses and disposes byte-identically to today.
183
+ `failFirst` is **not** sourced from the descriptor and is **demoted** (machine-proven fail-first
184
+ DELIVERED in #1279 — no path greens on attestation alone, FF-08); the `dispositionForProhibition`
185
+ policy is unchanged. Residual (tracked **#1346**): the node-test proof confirms the fixture exists and
186
+ the check goes RED, but cannot generically prove the red was *caused by* the subject's content.
187
+
117
188
  ## Output schema
118
189
 
119
190
  The probe emits, per kept prohibition, an item of the form:
120
191
 
121
192
  ```
122
- { requirement_id, category, status, verification, resolution, reason, statement }
193
+ { requirement_id, category, status, verification, resolution, reason, statement,
194
+ check_kind?, check_target?, check_rule? }
123
195
  ```
124
196
 
125
197
  where `statement` is the must-NOT sentence and `category` is the values/safety/ethics class
126
- (`values`, `fairness`, `privacy`, `transparency`, `safety`, …), plus a coverage summary:
198
+ (`values`, `fairness`, `privacy`, `transparency`, `safety`, …). The optional **flat-scalar
199
+ `check_*` descriptor** (#1278) is present only on a resolved `test`-tier prohibition carrying a
200
+ wired check: `check_kind` (`node-test` | `lint-rule`), `check_target`, and `check_rule` (lint-rule
201
+ only). `projectProhibitions` emits these into `must_haves.prohibitions` and `descriptorFromProjection`
202
+ reads them back into a `{ kind, target, rule? }` `CheckDescriptor`; they are flat scalars (never a
203
+ nested `check:{}` object) so they round-trip through the unchanged `parseMustHavesBlock`. Plus a
204
+ coverage summary:
127
205
 
128
206
  ```
129
207
  coverage: { applicable, resolved, unresolved, byVerification: { test, judgment } }
@@ -6,9 +6,13 @@ block — do not inline a copy elsewhere. History of coordinated edits: #2924, #
6
6
 
7
7
  **Contract for orchestrators:** before dispatch, capture `EXPECTED_BASE=$(git rev-parse HEAD)`,
8
8
  then embed the block below into the sub-agent prompt verbatim, substituting `{EXPECTED_BASE}`
9
- with that captured SHA. The sub-agent only *verifies* and fails closed; the orchestrator
10
- (the worktree lifecycle owner) performs any base recovery — the sub-agent never rewrites a
11
- worktree it did not create (#48).
9
+ with that captured SHA. Orchestrators that intentionally create a docs-only pre-dispatch
10
+ plan commit may also substitute `{EXPECTED_BASE_ALTERNATE}` with that commit's immediate
11
+ parent so runtimes that fork from either side of the docs-only commit pass the same
12
+ fail-closed guard (#1265). Otherwise substitute `{EXPECTED_BASE_ALTERNATE}` with an empty
13
+ string. The sub-agent only *verifies* and fails closed; the orchestrator (the worktree
14
+ lifecycle owner) performs any base recovery — the sub-agent never rewrites a worktree it
15
+ did not create (#48).
12
16
 
13
17
  <worktree_branch_check>
14
18
  FIRST ACTION: HEAD assertion MUST run before anything else, and this block is
@@ -30,8 +34,10 @@ if ! echo "$ACTUAL_BRANCH" | grep -Eq '^worktree-agent-[A-Za-z0-9._/-]+$'; then
30
34
  echo "FATAL: worktree HEAD '$ACTUAL_BRANCH' is not in the worktree-agent-* namespace; refusing to commit (#2924)." >&2
31
35
  exit 42
32
36
  fi
33
- if [ "$(git rev-parse HEAD)" != "{EXPECTED_BASE}" ]; then
34
- echo "FATAL: worktree base mismatch — HEAD is $(git rev-parse HEAD), expected {EXPECTED_BASE}. Orchestrator owns recovery; sub-agent refuses to rewrite the worktree (#48)." >&2
37
+ ACTUAL_BASE=$(git rev-parse HEAD)
38
+ EXPECTED_BASE_ALTERNATE="{EXPECTED_BASE_ALTERNATE}"
39
+ if [ "$ACTUAL_BASE" != "{EXPECTED_BASE}" ] && { [ -z "$EXPECTED_BASE_ALTERNATE" ] || [ "$ACTUAL_BASE" != "$EXPECTED_BASE_ALTERNATE" ]; }; then
40
+ echo "FATAL: worktree base mismatch — HEAD is $ACTUAL_BASE, expected {EXPECTED_BASE}${EXPECTED_BASE_ALTERNATE:+ or $EXPECTED_BASE_ALTERNATE}. Orchestrator owns recovery; sub-agent refuses to rewrite the worktree (#48)." >&2
35
41
  exit 42
36
42
  fi
37
43
  ```
@@ -12,6 +12,12 @@ phase: XX-name
12
12
  verified: YYYY-MM-DDTHH:MM:SSZ
13
13
  status: passed | gaps_found | human_needed
14
14
  score: N/M must-haves verified
15
+ behavior_unverified: 0 # Count of ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truths (present + wired, behavior not exercised)
16
+ behavior_unverified_items: # Only if behavior_unverified > 0 — the truths above as structured items; emitted regardless of overall status
17
+ - truth: "Observable truth whose state transition or cancellation/cleanup/ordering invariant no test exercises"
18
+ test: "What to trigger"
19
+ expected: "What state must hold afterward"
20
+ why_human: "Why presence checks can't see it"
15
21
  ---
16
22
 
17
23
  # Phase {X}: {Name} Verification Report
@@ -28,9 +34,10 @@ score: N/M must-haves verified
28
34
  |---|-------|--------|----------|
29
35
  | 1 | {truth from must_haves} | ✓ VERIFIED | {what confirmed it} |
30
36
  | 2 | {truth from must_haves} | ✗ FAILED | {what's wrong} |
31
- | 3 | {truth from must_haves} | ? UNCERTAIN | {why can't verify} |
37
+ | 3 | {truth from must_haves} | ⚠️ PRESENT_BEHAVIOR_UNVERIFIED | {present + wired; transition/invariant not exercised by a test — see Human Verification} |
38
+ | 4 | {truth from must_haves} | ? UNCERTAIN | {why can't verify} |
32
39
 
33
- **Score:** {N}/{M} truths verified
40
+ **Score:** {N}/{M} truths verified ({P} present, behavior-unverified)
34
41
 
35
42
  ### Required Artifacts
36
43
 
@@ -161,11 +168,17 @@ None — all verifiable items checked programmatically.
161
168
 
162
169
  ## Guidelines
163
170
 
164
- **Status values:**
171
+ **Status values (overall, frontmatter `status:`):**
165
172
  - `passed` — All must-haves verified, no blockers
166
173
  - `gaps_found` — One or more critical gaps found
167
174
  - `human_needed` — Automated checks pass but human verification required
168
175
 
176
+ **Per-truth states (Observable Truths `Status` column):**
177
+ - `✓ VERIFIED` — supporting artifacts pass all checks; for a behavior-dependent truth, a behavioral test exercised the asserted behavior
178
+ - `⚠️ PRESENT_BEHAVIOR_UNVERIFIED` — present + wired, but a state transition or cancellation/cleanup/ordering invariant was not exercised by any test. Counts toward `behavior_unverified`, routes to human verification, and is *excluded* from the verified score. Per-truth only — on its own the overall `status:` becomes `human_needed` (unless a higher-precedence `gaps_found` also applies); the item is preserved in `behavior_unverified_items` regardless.
179
+ - `✗ FAILED` — artifact missing, stub, or unwired
180
+ - `? UNCERTAIN` — can't verify programmatically
181
+
169
182
  **Evidence types:**
170
183
  - For EXISTS: "File at path, exports X"
171
184
  - For SUBSTANTIVE: "N lines, has patterns X, Y, Z"
@@ -457,27 +457,23 @@ Continue to collect_wave_1.
457
457
  <step name="collect_wave_1">
458
458
  **Read the work manifest first:** `Read .planning/tmp/docs-work-manifest.json` — update `status` to `"completed"` or `"failed"` for each Wave 1 item after collection. Write the updated manifest back to disk.
459
459
 
460
- Wait for all 3 Wave 1 agents to complete using the TaskOutput tool.
460
+ Wait for all 3 Wave 1 background agents to finish, then read each agent's output file to collect confirmations.
461
461
 
462
- Call TaskOutput for all 3 agents in parallel (single message with 3 TaskOutput calls):
462
+ Each `Agent(...)` call above with `run_in_background=true` returns an `async_launched` result that carries an `outputFile` path (and `canReadOutputFile: true`). Each agent's completion arrives as a message in this conversation when it finishes — do NOT issue a separate blocking call to wait. Once all 3 agents have reported completion, read their output files in parallel (single message with 3 Read calls):
463
463
 
464
464
  ```
465
- TaskOutput tool:
466
- task_id: "{task_id from README agent result}"
467
- block: true
468
- timeout: 300000
465
+ Read tool:
466
+ file_path: "{outputFile from README agent result}"
469
467
 
470
- TaskOutput tool:
471
- task_id: "{task_id from ARCHITECTURE agent result}"
472
- block: true
473
- timeout: 300000
468
+ Read tool:
469
+ file_path: "{outputFile from ARCHITECTURE agent result}"
474
470
 
475
- TaskOutput tool:
476
- task_id: "{task_id from CONFIGURATION agent result}"
477
- block: true
478
- timeout: 300000
471
+ Read tool:
472
+ file_path: "{outputFile from CONFIGURATION agent result}"
479
473
  ```
480
474
 
475
+ > Allow up to 5 minutes (300000 ms) for the slowest agent to finish before treating it as failed.
476
+
481
477
  **Expected confirmation format from each agent:**
482
478
  ```
483
479
  ## Doc Generation Complete
@@ -676,29 +672,25 @@ Continue to collect_wave_2.
676
672
  <step name="collect_wave_2">
677
673
  **Read the work manifest first:** `Read .planning/tmp/docs-work-manifest.json` — update `status` to `"completed"` or `"failed"` for each Wave 2 item after collection. Write the updated manifest back to disk.
678
674
 
679
- Wait for all Wave 2 agents to complete using the TaskOutput tool.
675
+ Wait for all Wave 2 background agents to finish, then read each agent's output file to collect confirmations.
680
676
 
681
- Call TaskOutput for all Wave 2 agents in parallel (single message with N TaskOutput calls — one per spawned Wave 2 agent):
677
+ Each `Agent(...)` call above with `run_in_background=true` returns an `async_launched` result that carries an `outputFile` path (and `canReadOutputFile: true`). Each agent's completion arrives as a message in this conversation when it finishes — do NOT issue a separate blocking call to wait. Once all Wave 2 agents have reported completion, read their output files in parallel (single message with N Read calls — one per spawned Wave 2 agent):
682
678
 
683
679
  ```
684
- TaskOutput tool:
685
- task_id: "{task_id from GETTING-STARTED agent result}"
686
- block: true
687
- timeout: 300000
680
+ Read tool:
681
+ file_path: "{outputFile from GETTING-STARTED agent result}"
688
682
 
689
- TaskOutput tool:
690
- task_id: "{task_id from DEVELOPMENT agent result}"
691
- block: true
692
- timeout: 300000
683
+ Read tool:
684
+ file_path: "{outputFile from DEVELOPMENT agent result}"
693
685
 
694
- TaskOutput tool:
695
- task_id: "{task_id from TESTING agent result}"
696
- block: true
697
- timeout: 300000
686
+ Read tool:
687
+ file_path: "{outputFile from TESTING agent result}"
698
688
 
699
- # Add one TaskOutput call per conditional agent spawned (API, DEPLOYMENT, CONTRIBUTING)
689
+ # Add one Read call per conditional agent spawned (API, DEPLOYMENT, CONTRIBUTING)
700
690
  ```
701
691
 
692
+ > Allow up to 5 minutes (300000 ms) for the slowest agent to finish before treating it as failed.
693
+
702
694
  **After collection, verify all Wave 2 files exist on disk** using the `resolved_path` from each manifest entry:
703
695
  ```bash
704
696
  ls -la {resolved_path for each wave 2 item} 2>/dev/null
@@ -752,9 +744,9 @@ Write {package_dir}/README.md directly. Return confirmation only — do not retu
752
744
  )
753
745
  ```
754
746
 
755
- > **ORCHESTRATOR RULE — CODEX RUNTIME**: After calling all per-package Agent() calls above with `run_in_background=true`, do NOT generate any package READMEs independently while the subagents are active. Wait for all agents to complete via TaskOutput before proceeding. This prevents duplicate work and wasted context.
747
+ > **ORCHESTRATOR RULE — CODEX RUNTIME**: After calling all per-package Agent() calls above with `run_in_background=true`, do NOT generate any package READMEs independently while the subagents are active. Wait for all agents to complete before proceeding. This prevents duplicate work and wasted context.
756
748
 
757
- Collect confirmations via TaskOutput for all package agents. Note failures in the final report.
749
+ Collect confirmations by reading each package agent's `outputFile` once it reports completion — each `run_in_background=true` Agent call returns an `async_launched` result carrying an `outputFile` path (with `canReadOutputFile: true`). Note failures in the final report.
758
750
 
759
751
  **Fallback when Task tool is unavailable:** Generate per-package READMEs sequentially inline after the `sequential_generation` step. For each package directory with a `package.json`, construct the equivalent `doc_assignment` block and generate the README following gsd-doc-writer instructions.
760
752
 
@@ -0,0 +1,9 @@
1
+ # Worktree Recovery Policy
2
+
3
+ ## ORCHESTRATOR FAIL-CLOSED RULE (#48)
4
+
5
+ > **ORCHESTRATOR FAIL-CLOSED RULE (#48):** `worktree_branch_check` is verify-only — an executor that hits a base/HEAD-namespace mismatch prints `FATAL:` and exits **42** instead of self-recovering. If any executor result reports a `FATAL:`/`exit 42` (or its commits never appear because it halted at the check), mark that plan **blocked**: do NOT merge or clean up its worktree (preserve it for inspection), do NOT count the wave as successful, and surface the mismatch with recovery guidance to the user. The orchestrator — the worktree lifecycle owner — performs any base correction (e.g. recreate the worktree on `{EXPECTED_BASE}`); the sub-agent never does. Never proceed past a halted executor on the assumption it succeeded.
6
+
7
+ ## ISOLATED-RUN RECOVERY — FAIL SAFE (#1292)
8
+
9
+ > **ISOLATED-RUN RECOVERY — FAIL SAFE (#1292):** When an isolated (worktree) run is *rejected* — the user declines to merge it, the orchestrator surfaces recovery guidance for a blocked/halted plan, or the run over-reached the requested scope — the worktree-isolation contract MUST hold through recovery. Do **NOT** propose continuing on `main`/the primary checkout as the default or recommended recovery path. Default to a **safe halt** and offer: (a) re-attempt in a **fresh, narrowly-scoped worktree**, or (b) inspect or discard the rejected worktree without merging. Any path that edits the primary checkout requires an **explicit, clearly-labeled confirmation** from the user first — editing `main` directly is never the proposed or default option for a run the user configured to be isolated.
@@ -635,6 +635,7 @@ increases monotonically across waves. `{status}` is `complete` (success),
635
635
  this commit — the orchestrator force-removes the worktree after you return, and
636
636
  any uncommitted SUMMARY.md will be permanently lost (#2070).
637
637
  REQUIRED ORDER: Write SUMMARY.md → commit → only then any narration. No text between Write and commit (truncation risk; #2070 rescue is not primary defense).
638
+
638
639
  </parallel_execution>
639
640
 
640
641
  <execution_context>
@@ -682,9 +683,9 @@ increases monotonically across waves. `{status}` is `complete` (success),
682
683
  )
683
684
  ```
684
685
 
685
- Immediately after each worktree `Agent()` spawn returns metadata, atomically append `{agent_id, worktree_path, branch, expected_base}` to `WAVE_WORKTREE_MANIFEST`. If any field is missing, stop and ask for recovery instead of scanning all agent worktrees.
686
+ After each `Agent()` returns, parse executor-returned worktree metadata (`<worktree_metadata>`) before harness metadata, then atomically append `{agent_id, worktree_path, branch, expected_base}` to `WAVE_WORKTREE_MANIFEST`. Missing: stop and ask for recovery instead of scanning worktrees.
686
687
 
687
- > **ORCHESTRATOR FAIL-CLOSED RULE (#48):** `worktree_branch_check` is verify-only — an executor that hits a base/HEAD-namespace mismatch prints `FATAL:` and exits **42** instead of self-recovering. If any executor result reports a `FATAL:`/`exit 42` (or its commits never appear because it halted at the check), mark that plan **blocked**: do NOT merge or clean up its worktree (preserve it for inspection), do NOT count the wave as successful, and surface the mismatch with recovery guidance to the user. The orchestrator — the worktree lifecycle owner — performs any base correction (e.g. recreate the worktree on `{EXPECTED_BASE}`); the sub-agent never does. Never proceed past a halted executor on the assumption it succeeded.
688
+ > **Worktree recovery policy (#48 + #1292):** See `execute-phase/steps/worktree-recovery-policy.md` — FAIL-CLOSED rule for base/HEAD-namespace mismatches AND isolated-run fail-safe recovery.
688
689
 
689
690
  > **ORCHESTRATOR RULE — CODEX RUNTIME**: After calling Agent() above to spawn executor agent(s), stop working on this task immediately. Do not read more files, edit code, or run tests related to this task while the subagent is active. Wait for the subagent to return its result. This prevents duplicate work, conflicting edits, and wasted context. Only resume when the subagent result is available.
690
691
 
@@ -753,6 +754,8 @@ increases monotonically across waves. `{status}` is `complete` (success),
753
754
  ask for one recovery path: `continue waiting`, `kill and retry`, or
754
755
  `kill and switch to inline execution`.
755
756
 
757
+ If the stalled executor ran in an isolated worktree, `kill and switch to inline execution` edits the primary checkout — see worktree recovery policy (`execute-phase/steps/worktree-recovery-policy.md`). Prefer `kill and retry` in a fresh worktree; inline execution requires explicit confirmation, never the default.
758
+
756
759
  **This fallback applies automatically to all runtimes.** Claude Code's Agent() normally
757
760
  returns synchronously, but the fallback ensures resilience if it doesn't.
758
761
 
@@ -855,6 +858,8 @@ increases monotonically across waves. `{status}` is `complete` (success),
855
858
 
856
859
  **If no worktrees found at runtime:** Skip silently — agents may have been spawned without worktree isolation, or the orchestrator already cleaned them up.
857
860
 
861
+ If the user declines to merge a worktree or a worktree over-reached scope, apply the worktree recovery policy (`execute-phase/steps/worktree-recovery-policy.md`) — never default to editing `main`.
862
+
858
863
  5.6. **Post-merge build & test gate:**
859
864
 
860
865
  After merging all worktrees in a wave (parallel mode), or after the last plan completes
@@ -255,21 +255,19 @@ Continue to collect_confirmations.
255
255
  </step>
256
256
 
257
257
  <step name="collect_confirmations">
258
- Wait for all 4 agents to complete using TaskOutput tool.
258
+ Wait for all 4 background agents to finish, then read each agent's output file to collect confirmations.
259
259
 
260
- **For each agent task_id returned by the Agent tool calls above:**
260
+ Each `Agent(...)` call above with `run_in_background=true` returns an `async_launched` result that carries an `outputFile` path (and `canReadOutputFile: true`). The 4 agents run concurrently and each one's completion arrives as a message in this conversation when it finishes — do NOT issue a separate blocking call to wait for them.
261
+
262
+ **Once all 4 agents have reported completion, read each agent's output file (single message with 4 Read calls):**
261
263
  ```
262
- TaskOutput tool:
263
- task_id: "{task_id from Agent result}"
264
- block: true
265
- timeout: {subagent_timeout from init context, default 300000}
264
+ Read tool:
265
+ file_path: "{outputFile from that agent's async_launched result}"
266
266
  ```
267
267
 
268
- > The timeout is configurable via `workflow.subagent_timeout` in `.planning/config.json` (milliseconds). Default: 300000 (5 minutes). Increase for large codebases or slower models.
269
-
270
- Call TaskOutput for all 4 agents in parallel (single message with 4 TaskOutput calls).
268
+ > Allow up to `workflow.subagent_timeout` for the slowest agent to finish before treating it as failed. The timeout is configurable via `workflow.subagent_timeout` in `.planning/config.json` (milliseconds). Default: 300000 (5 minutes). Increase for large codebases or slower models.
271
269
 
272
- Once all TaskOutput calls return, read each agent's output file to collect confirmations.
270
+ Each output file contains that agent's completion confirmation. Parse the confirmation marker (see below) from the file contents.
273
271
 
274
272
  **Expected confirmation format from each agent:**
275
273
  ```
@@ -499,6 +499,12 @@ Display banner:
499
499
 
500
500
  ### Spawn gsd-phase-researcher
501
501
 
502
+ ```bash
503
+ if gsd_run query teams-status --active >/dev/null 2>&1; then
504
+ echo "⚠️ CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS detected. GSD's multi-agent orchestration is not validated under claude-code agent-teams and may stall (a subagent's completion can fail to route to the orchestrator). Recommend disabling agent-teams for GSD workflows. See https://github.com/open-gsd/gsd-core/issues/1355" >&2
505
+ fi
506
+ ```
507
+
502
508
  ```bash
503
509
  PHASE_DESC=$(gsd_run query roadmap.get-phase "${PHASE}" --pick section)
504
510
  if [ -z "${PLAN_PRE_HOOKS_JSON:-}" ]; then
@@ -632,7 +632,10 @@ When `USE_WORKTREES !== "false"`, commit PLAN.md to the current branch **before*
632
632
  Skip this step entirely if `USE_WORKTREES === "false"` (non-worktree mode: PLAN.md is committed in Step 8 as usual).
633
633
 
634
634
  ```bash
635
+ QUICK_PLAN_PARENT=""
636
+ QUICK_PLAN_COMMIT=""
635
637
  if [ "${USE_WORKTREES}" != "false" ]; then
638
+ QUICK_PLAN_PARENT=$(git rev-parse HEAD)
636
639
  COMMIT_DOCS=$(gsd_run query config-get commit_docs 2>/dev/null || echo "true")
637
640
  if [ "$COMMIT_DOCS" != "false" ]; then
638
641
  git add "${QUICK_DIR}/${quick_id}-PLAN.md"
@@ -650,8 +653,12 @@ if [ "${USE_WORKTREES}" != "false" ]; then
650
653
  git commit -m "docs(${quick_id}): pre-dispatch plan for ${DESCRIPTION}" -- "${QUICK_DIR}/${quick_id}-PLAN.md" \
651
654
  || { echo "ERROR: pre-dispatch PLAN.md commit failed — likely a pre-commit hook failure. Fix the hook output above (or set workflow.worktree_skip_hooks=true to bypass) and re-run." >&2; exit 1; }
652
655
  fi
656
+ QUICK_PLAN_COMMIT=$(git rev-parse HEAD)
653
657
  fi
654
658
  fi
659
+ if [ -z "$QUICK_PLAN_COMMIT" ]; then
660
+ QUICK_PLAN_COMMIT=$(git rev-parse HEAD)
661
+ fi
655
662
  fi
656
663
  ```
657
664
 
@@ -678,8 +685,22 @@ Execute quick task ${quick_id}.
678
685
 
679
686
  ${USE_WORKTREES !== "false" ? `
680
687
  <worktree_branch_check>
681
- ORCHESTRATOR build-time embed (NOT a sub-agent runtime step): before this dispatch, read \`gsd-core/references/worktree-branch-check.md\`, substitute \`{EXPECTED_BASE}\` with the base SHA captured above (${EXPECTED_BASE}), and replace this note with that fragment's \`<worktree_branch_check>\` block so the dispatched prompt carries the runnable guard verbatim — do not pass this instruction through in its place.
688
+ ORCHESTRATOR build-time embed (NOT a sub-agent runtime step): before this dispatch, read \`gsd-core/references/worktree-branch-check.md\`, substitute \`{EXPECTED_BASE}\` with the base SHA captured above (${EXPECTED_BASE}), substitute \`{EXPECTED_BASE_ALTERNATE}\` with \`${QUICK_PLAN_PARENT}\` when it differs from \`${EXPECTED_BASE}\` (otherwise empty), and replace this note with that fragment's \`<worktree_branch_check>\` block so the dispatched prompt carries the runnable guard verbatim — do not pass this instruction through in its place.
682
689
  </worktree_branch_check>
690
+
691
+ FIRST ACTION after the worktree branch check: ensure the quick PLAN.md exists at a worktree-rooted relative path before any Read/Edit/Write path can be primed. If \`${QUICK_DIR}/${quick_id}-PLAN.md\` is absent, materialize it from the shared git object store:
692
+
693
+ \`\`\`bash
694
+ QUICK_PLAN_COMMIT="${QUICK_PLAN_COMMIT}"
695
+ QUICK_PLAN_PATH="${QUICK_DIR}/${quick_id}-PLAN.md"
696
+ if [ ! -f "$QUICK_PLAN_PATH" ]; then
697
+ mkdir -p "$(dirname "$QUICK_PLAN_PATH")"
698
+ git show "${QUICK_PLAN_COMMIT}:${QUICK_PLAN_PATH}" > "$QUICK_PLAN_PATH" || {
699
+ echo "FATAL: unable to materialize quick plan from ${QUICK_PLAN_COMMIT}:${QUICK_PLAN_PATH}; refusing to continue." >&2
700
+ exit 42
701
+ }
702
+ fi
703
+ \`\`\`
683
704
  ` : ''}
684
705
 
685
706
  <files_to_read>
@@ -743,7 +764,7 @@ SUMMARY.md and stop — the user must rerun with worktrees disabled.
743
764
 
744
765
  > **ORCHESTRATOR RULE — CODEX RUNTIME**: After calling Agent() above, stop working on this task immediately. Do not read more files, edit code, or run tests related to this task while the subagent is active. Wait for the subagent to return its result. This prevents duplicate work, conflicting edits, and wasted context. Only resume when the subagent result is available.
745
766
 
746
- If the executor ran with `isolation="worktree"`, append its returned `{agent_id, worktree_path, branch, expected_base}` metadata to `QUICK_WORKTREE_MANIFEST` before cleanup. If any field is unavailable, stop and ask for recovery; do not discover global worktrees.
767
+ If the executor ran with `isolation="worktree"`, append its returned `{agent_id, worktree_path, branch, expected_base, allowed_bases}` metadata to `QUICK_WORKTREE_MANIFEST` before cleanup. Set `expected_base` to `${EXPECTED_BASE}` and `allowed_bases` to `["${EXPECTED_BASE}", "${QUICK_PLAN_PARENT}"]` with duplicates removed. If any required field is unavailable, stop and ask for recovery; do not discover global worktrees.
747
768
 
748
769
  After executor returns:
749
770
  1. **Worktree cleanup:** If the executor ran with `isolation="worktree"`, merge the worktree branch back and clean up:
@@ -761,6 +782,9 @@ After executor returns:
761
782
  gsd_run query worktree.cleanup-wave --manifest "$QUICK_WORKTREE_MANIFEST" || exit 1
762
783
  ```
763
784
  If `workflow.use_worktrees` is `false`, skip this step.
785
+
786
+ > **ISOLATED-RUN RECOVERY — FAIL SAFE (#1292):** When an isolated (worktree) run is *rejected* — the user declines to merge it, the orchestrator surfaces recovery guidance for a blocked/halted plan, or the run over-reached the requested scope — the worktree-isolation contract MUST hold through recovery. Do **NOT** propose continuing on `main`/the primary checkout as the default or recommended recovery path. Default to a **safe halt** and offer: (a) re-attempt in a **fresh, narrowly-scoped worktree**, or (b) inspect or discard the rejected worktree without merging. Any path that edits the primary checkout requires an **explicit, clearly-labeled confirmation** from the user first — editing `main` directly is never the proposed or default option for a run the user configured to be isolated.
787
+
764
788
  2. Verify summary exists at `${QUICK_DIR}/${quick_id}-SUMMARY.md`
765
789
  3. Extract commit hash from executor output
766
790
  4. Report completion status
@@ -50,7 +50,7 @@ Planning Tuning:
50
50
  - `workflow.plan_bounce` (default: `false`)
51
51
  - `workflow.plan_bounce_passes` (default: `2`)
52
52
  - `workflow.plan_bounce_script` (default: `null`)
53
- - `workflow.subagent_timeout` (default: `600`)
53
+ - `workflow.subagent_timeout` (default: `300000`)
54
54
  - `workflow.inline_plan_threshold` (default: `3`)
55
55
 
56
56
  Execution Tuning:
@@ -155,12 +155,12 @@ AskUserQuestion([
155
155
  ]
156
156
  },
157
157
  {
158
- question: "Subagent timeout (seconds)? (current: <value or 600>)",
158
+ question: "Subagent timeout (milliseconds)? (current: <value or 300000>)",
159
159
  header: "Subagent Timeout",
160
160
  multiSelect: false,
161
161
  options: [
162
162
  { label: "Keep current", description: "Leave timeout unchanged." },
163
- { label: "Enter seconds", description: "Integer number of seconds. Non-numeric rejected. Default: 600" }
163
+ { label: "Enter milliseconds", description: "Integer number of milliseconds. Non-numeric rejected. Default: 300000 (5 minutes)." }
164
164
  ]
165
165
  },
166
166
  {
@@ -498,7 +498,7 @@ keys and sibling sub-objects.
498
498
  ```bash
499
499
  # Example — only write keys the user changed. "Keep current" selections are skipped.
500
500
  gsd_run query config-set workflow.plan_bounce_passes 5
501
- gsd_run query config-set workflow.subagent_timeout 900
501
+ gsd_run query config-set workflow.subagent_timeout 300000
502
502
  gsd_run query config-set git.base_branch main
503
503
  gsd_run query config-set context_window 1000000
504
504
  # Runtime model tier examples:
@@ -751,7 +751,7 @@ Display:
751
751
  | workflow.plan_bounce | {on/off} |
752
752
  | workflow.plan_bounce_passes | {n} |
753
753
  | workflow.plan_bounce_script | {path/null} |
754
- | workflow.subagent_timeout | {seconds} |
754
+ | workflow.subagent_timeout | {milliseconds} |
755
755
  | workflow.inline_plan_threshold | {n} |
756
756
  | workflow.node_repair | {on/off} |
757
757
  | workflow.node_repair_budget | {n} |
@@ -173,20 +173,20 @@ AskUserQuestion([
173
173
  multiSelect: false,
174
174
  options: [
175
175
  { label: "Claude", description: "review.models.claude — defaults to session model when unset" },
176
- { label: "Codex", description: "review.models.codex — e.g. 'codex exec --model gpt-5'" },
177
- { label: "Gemini", description: "review.models.gemini — e.g. 'gemini -m gemini-2.5-pro'" },
178
- { label: "OpenCode", description: "review.models.opencode — e.g. 'opencode run --model claude-sonnet-4'" }
176
+ { label: "Codex", description: "review.models.codex — bare model id injected into --model, e.g. 'gpt-5'" },
177
+ { label: "Gemini", description: "review.models.gemini — bare model id injected into -m, e.g. 'gemini-2.5-pro'" },
178
+ { label: "OpenCode", description: "review.models.opencode — bare model id injected into --model, e.g. 'claude-sonnet-4'" }
179
179
  ]
180
180
  }
181
181
  ])
182
182
  ```
183
183
 
184
184
  For the selected CLI, show the current value (or `(unset)`) and offer
185
- Leave / Replace / Clear, followed by a text-input prompt for the new command
185
+ Leave / Replace / Clear, followed by a text-input prompt for the model id
186
186
  string. Write via:
187
187
 
188
188
  ```bash
189
- gsd_run query config-set review.models.<cli> "<command string>"
189
+ gsd_run query config-set review.models.<cli> "<model id>"
190
190
  ```
191
191
 
192
192
  After each update, return to the "Review model CLI mapping — what next?" question.
@@ -356,6 +356,26 @@ For each Requirement gathered so far, run the two-stage recall→precision pass:
356
356
  Criteria AND mark the prohibition `resolved` with a verification tier: `test` (a
357
357
  mechanical negative test/lint/assertion exists) or `judgment` (real but not mechanically
358
358
  checkable — routes to judgment review).
359
+ - **Capture the wired-check descriptor on `test`-tier (#1278, SOFT).** When a prohibition is
360
+ resolved `verification: test`, ALSO capture the descriptor of the wired check so
361
+ `verify-phase` can LOCATE it deterministically (no verifier invention at verify time).
362
+ Capture the flat scalars — persisted into SPEC and projected onto the
363
+ `must_haves.prohibitions` item by `projectProhibitions`:
364
+ - `check_kind` — `node-test` | `lint-rule`.
365
+ - `check_target` — the negative-test file path (for `node-test`), or the path to lint
366
+ (for `lint-rule`).
367
+ - `check_rule` — the eslint rule id (e.g. `local/no-source-grep`); `lint-rule` only.
368
+ - `check_violation_fixture` (#1346) — path to a KNOWN-BAD subject the wired check is run
369
+ against to **machine-prove fail-first**; rides BOTH kinds. Capture it to let the item green
370
+ end-to-end with zero hand-authoring at verify time; for `node-test` the negative test should
371
+ read its subject from the `GSD_PROHIB_SUBJECT` env var so the prover can inject this fixture.
372
+ This is a **SOFT capture (CHK-04): a `test`-tier prohibition WITHOUT a descriptor is still
373
+ allowed** — if the author cannot yet name the wired check, leave the descriptor empty and
374
+ proceed. It is NOT a hard authoring block; the item simply stays fail-closed/flagged
375
+ downstream (an absent/partial descriptor — or one with no `check_violation_fixture` —
376
+ → `descriptorFromProjection` null/under-specified/fixture-less → producer fail-closed
377
+ locate-or-unprovable, never green). Do NOT capture `failFirst` here — it is a
378
+ verify-time caller attestation, not a spec-authored field (#1279).
359
379
  - **Dismiss (reason)** → mark `dismissed` with a REQUIRED non-empty reason (PROB-05). The
360
380
  reason string is the audit trail; silence is not a valid dismissal.
361
381
  - **Defer** → leave `unresolved`.
@@ -374,7 +394,11 @@ For each Requirement gathered so far, run the two-stage recall→precision pass:
374
394
  **`--auto` mode:** auto-`resolved` where a defensible negative acceptance criterion can be
375
395
  written (test or judgment tier); otherwise leave `unresolved`. **`--auto` NEVER auto-dismisses
376
396
  a prohibition** — a wrong dismissal is the exact silent failure this probe eliminates (PROB-06,
377
- the load-bearing safety property). Log: `[auto] prohibitions: R resolved, U unresolved`.
397
+ the load-bearing safety property). On a `test`-tier auto-resolution, capture the `check_kind` /
398
+ `check_target` / `check_rule` / `check_violation_fixture` descriptor **only when a wired check is unambiguous**; otherwise
399
+ leave it empty — `--auto` NEVER fabricates a check path or fixture (a wrong locate is re-validated and
400
+ fails closed at the producer, but a fabricated path is still noise to avoid). Log:
401
+ `[auto] prohibitions: R resolved, U unresolved`.
378
402
 
379
403
  **Text mode (PROB-09):** per Step 5's text-mode rule, replace the AskUserQuestion menus above
380
404
  with plain-text numbered lists — there is NO hard AskUserQuestion dependency, so the probe
@@ -382,7 +406,11 @@ runs identically for non-Claude / text-mode hosts.
382
406
 
383
407
  Populate the `## Prohibitions` section of SPEC.md from the resolved prohibitions (each
384
408
  `resolved`/`test` row is a checkable negative acceptance criterion; `resolved`/`judgment`
385
- rows route to judgment review; `⚠ UNRESOLVED` rows are flagged as assumptions).
409
+ rows route to judgment review; `⚠ UNRESOLVED` rows are flagged as assumptions). A
410
+ `resolved`/`test` row ALSO carries its captured `check_kind` / `check_target` / `check_rule` /
411
+ `check_violation_fixture` descriptor when present (so the projection feeds `verify-phase`'s deterministic locate + machine-proof, #1278 + #1346);
412
+ a `test` row with no captured descriptor is still valid — it stays fail-closed/flagged
413
+ downstream rather than blocking authoring.
386
414
 
387
415
  ## Step 6: Generate SPEC.md
388
416