pi-gauntlet 5.0.1 → 5.0.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,13 @@
1
1
  # Changelog
2
2
 
3
+ ## v5.0.3 - 2026-08-25
4
+
5
+ - gatekeep-pr defaults to green exact-head CI evidence (gh-14): normative six-path "Evidence resolution" table at the top of verification-brief.md Section B (opt-out / failed-CI / CI-sufficient / pending / fallback / stale-head, top-down); the local verification command runs only on fallback/opt-out rows; source-discriminated Verifier output (`source: ci|local`) with the exact CI claim form `verified by CI: <check name(s)> succeeded on <sha> (run <url>)`; any blocking conclusion in the resolved set mints a `P#` with a third disposition `CI-infrastructure-broken` that triggers the fallback run; two new `## PR gate` keys `local verification: always` and `ci checks:`. Spec: `doc/specs/2026-09-06-gh-14-gatekeep-ci-evidence-default.md` (partially supersedes `doc/specs/2026-08-18-gh-9-gatekeep-pr-skill.md`, verification-evidence scope only).
6
+
7
+ ## v5.0.2 - 2026-08-24
8
+
9
+ - Plan fidelity (gh-13): `writing-plans` task template gains a required spec-anchor line (`**Spec:** <path> § "<heading>" L<start>-L<end>`), a verbatim-quote rule for exact-string requirements, an extraction-first `## Spec coverage` table, and four mechanical self-review checks (quote integrity spec->task, anchor resolution, three-leg table closure, paths exist). `subagent-driven-development` spec-reviewer contract becomes spec+task: dispatches pass the spec path + the task's anchors in both modes and the Dispatch sketch, the spec wins every dispute, task-vs-spec divergence is unconditionally flagged with the spec literal, and `spec-reviewer-prompt.md` gains a `## Spec Authority` section plus `plan transcription gap` / `out-of-anchor-slice` finding labels. Spec: `doc/specs/2026-08-23-gh-13-plan-fidelity-anchors.md` (partially supersedes `doc/specs/2026-07-06-parallel-wave-spec-reviewer-dispatch.md`, SR contract scope only).
10
+
3
11
  ## v5.0.1 - 2026-08-23
4
12
 
5
13
  - Council roast hardening: verification scope moves from the `spec-council-member` persona to dispatch task text (ticket roasts are content-only; spec roasts verify bounded - `rg`, explicit paths, `timeout`); persona gains a read-only invariant; explicit silence-kill control blocks (shape-ticket 5 min, spec-roast members 10 min, chair 15 min); shape-ticket mandates two-call dispatch (member fanout, then chair over usable files); mechanical usable-critique probe (`verdict:`/`addresses-problem:` headers, `consensus:` for the chair); targeted single retry of failed members only; quorum salvage (>= 1 usable critique -> chair runs with a `Coverage:` note, rendered to the user at brainstorming's gate and at shape-ticket's confirmation gate when coverage was partial). Spec: `doc/specs/2026-08-23-council-roast-hardening.md`.
package/README.md CHANGED
@@ -69,7 +69,7 @@ Everything between gate 1 and gate 2 - task breakdown, implementation, both revi
69
69
 
70
70
  pi-gauntlet ships three kinds of pieces, layered on top of pi-cohort's dispatch:
71
71
 
72
- - **16 skills** - the workflow logic. Twelve activate automatically when pi sees the matching kind of task, and each one gates the next: `brainstorming`, `writing-plans`, `roasting-the-spec`, `test-driven-development`, `subagent-driven-development`, `dispatching-parallel-agents`, `verification-before-completion`, `requesting-code-review`, `receiving-code-review`, `using-git-worktrees`, `finishing-a-development-branch`, `writing-skills`. Four more are explicit-invocation-only (`disable-model-invocation: true`): `shape-ticket` creates or repairs one tracker issue per run against a Context/Problem/Idea/Acceptance-Criteria template, gated by an AC integrity check, a cheap council roast, and a single human-confirmed write - run it with `/skill:shape-ticket`. `gatekeep-pr` is consent-gated pre-merge verification of a PR against its issue - read-only gathering, running the project's verification command, a rubric-based review, then a deterministic authorship-aware menu with stable finding IDs (P#/L#/C#/F#) and numbered pre-composed courses (fixes execute as a single parallel-safe wave: one gate run, one re-review, one push); nothing mutates (fixes, pushes, reviews, merges) until you pick a row - run it with `/skill:gatekeep-pr <pr>`. `check-delivery` is a post-merge detective control: proves an issue actually shipped (default-branch landing, delivery target, per-AC evidence) before its tracker status advances; it never writes a terminal status - run it with `/skill:check-delivery <ref>`. `chase-bug` is human-only bug triage: read-only root-cause discovery to an evidenced verdict menu (real bug -> ticket/brainstorm/respond; five negative verdicts), then a gated response to the reporter - it never fixes during triage - run it with `/skill:chase-bug`.
72
+ - **16 skills** - the workflow logic. Twelve activate automatically when pi sees the matching kind of task, and each one gates the next: `brainstorming`, `writing-plans`, `roasting-the-spec`, `test-driven-development`, `subagent-driven-development`, `dispatching-parallel-agents`, `verification-before-completion`, `requesting-code-review`, `receiving-code-review`, `using-git-worktrees`, `finishing-a-development-branch`, `writing-skills`. Four more are explicit-invocation-only (`disable-model-invocation: true`): `shape-ticket` creates or repairs one tracker issue per run against a Context/Problem/Idea/Acceptance-Criteria template, gated by an AC integrity check, a cheap council roast, and a single human-confirmed write - run it with `/skill:shape-ticket`. `gatekeep-pr` is consent-gated pre-merge verification of a PR against its issue - read-only gathering, verification evidence resolved CI-first (green checks on the exact assessed head count as evidence; the project's verification command runs only as fallback), a rubric-based review, then a deterministic authorship-aware menu with stable finding IDs (P#/L#/C#/F#) and numbered pre-composed courses (fixes execute as a single parallel-safe wave: one gate run, one re-review, one push); nothing mutates (fixes, pushes, reviews, merges) until you pick a row - run it with `/skill:gatekeep-pr <pr>`. `check-delivery` is a post-merge detective control: proves an issue actually shipped (default-branch landing, delivery target, per-AC evidence) before its tracker status advances; it never writes a terminal status - run it with `/skill:check-delivery <ref>`. `chase-bug` is human-only bug triage: read-only root-cause discovery to an evidenced verdict menu (real bug -> ticket/brainstorm/respond; five negative verdicts), then a gated response to the reporter - it never fixes during triage - run it with `/skill:chase-bug`.
73
73
  - **7 subagent personas** - the specialized child agents the skills dispatch via pi-cohort: `implementer`, `code-reviewer`, `spec-reviewer`, `conformance-reviewer`, `spec-summarizer`, `spec-council-member`, `spec-council-synthesizer`. See [doc/personas.md](./doc/personas.md) for what each one does and why its permissions are scoped the way they are.
74
74
  - **3 runtime extensions** - the enforcement layer. `plan-tracker` and `phase-tracker` are tools skills call to track progress (with a TUI widget); `verify-before-ship` is a hook that warns if you push or open a PR without a passing test run since your last edit; a phase-tracker flow guard reminds on implement-phase commits missing spec/code review. See [doc/configuration.md](./doc/configuration.md) for the settings each one reads.
75
75
 
@@ -355,6 +355,8 @@ customization lives in two places, never in the wrapper itself:
355
355
  - verification command: <command> # required unless documented elsewhere
356
356
  - timeout minutes: 15 # optional; default 15
357
357
  - requires credentials: false # optional; true => skill reports "not run" as missing evidence
358
+ - local verification: always # optional; default (absent) = CI-first; "always" forces the local run even when exact-head CI is green
359
+ - ci checks: <comma-separated check names> # optional; narrows which checks count as evidence; absent = all checks on the assessed head
358
360
  - worktree wrapper: <command> # optional
359
361
  - issue fetch: <command with <ref> placeholder> # optional, replaces gh issue view
360
362
  - merge policy: squash | merge-commit # optional
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-gauntlet",
3
- "version": "5.0.1",
3
+ "version": "5.0.3",
4
4
  "description": "Opinionated, gated workflow skills, subagent personas, and runtime extensions for the pi coding agent.",
5
5
  "author": "Jacek Juraszek",
6
6
  "type": "module",
@@ -9,8 +9,9 @@ argument-hint: "<pr> [issue-ref] (e.g. 123, or 123 gh-45)"
9
9
 
10
10
  Verify, don't trust. A PR description is a claim, not proof: over-claimed coverage,
11
11
  hallucinated references, and "tests pass" that were never rerun are the normal case,
12
- not the exception - especially on generated code. This skill gathers evidence, runs
13
- the project's own verification command, reviews the diff against a rubric, and
12
+ not the exception - especially on generated code. This skill gathers evidence, accepts green CI on the exact assessed head as
13
+ verification evidence (running the project's own verification command only as
14
+ the fallback), reviews the diff against a rubric, and
14
15
  presents a deterministic, authorship-aware menu. Authorship sets which row carries
15
16
  `[recommended]`; it never changes which rows are offered.
16
17
 
@@ -40,7 +41,7 @@ a wrapper skill:
40
41
  baseline and reviewer-persona defaults on any conflict.
41
42
  2. **Gauntlet overrides file** (3-location discovery, first found wins): the
42
43
  `## PR gate` section (verification command, `timeout minutes`, `requires credentials`, issue
43
- fetch, worktree wrapper, merge policy). An existing `## verification-before-completion`
44
+ fetch, worktree wrapper, merge policy, `local verification`, `ci checks`). An existing `## verification-before-completion`
44
45
  section is an accepted equivalent source for the verification command.
45
46
  3. **Repo documentation** - an explicitly documented command or tool (e.g. `AGENTS.md`'s
46
47
  canonical test entrypoint, a documented worktree wrapper, a documented tracker CLI,
@@ -56,6 +57,8 @@ which is required unless documented elsewhere):
56
57
  - verification command: <command> # required unless documented elsewhere
57
58
  - timeout minutes: 15 # optional; default 15
58
59
  - requires credentials: false # optional; true => skill reports "not run" as missing evidence
60
+ - local verification: always # optional; default (absent) = CI-first; "always" forces the local run even when exact-head CI is green
61
+ - ci checks: <comma-separated check names> # optional; narrows which checks count as evidence; absent = all checks on the assessed head
59
62
  - worktree wrapper: <command> # optional
60
63
  - issue fetch: <command with <ref> placeholder> # optional, replaces gh issue view
61
64
  - merge policy: squash | merge-commit # optional
@@ -81,7 +84,7 @@ configuration.
81
84
  ## Progress tracking
82
85
 
83
86
  Use `plan_tracker`, never `phase_tracker`. Init with the stage names: `gather`,
84
- `provision worktree`, `run verification`, `claim-check`, `review`, `consent menu`.
87
+ `provision worktree`, `resolve evidence`, `claim-check`, `review`, `consent menu`.
85
88
  Append one task per material claim as the Verifier enumerates them. A passing stage or
86
89
  a matched claim -> `complete`. A failed stage or a contradicted claim -> `failed`
87
90
  (shown crossed, error color) and stays failed while the skill stops at the menu -
@@ -124,14 +127,15 @@ merge-ready and surfaced (see the merge preconditions below).
124
127
  **Phase 3 - Verify, then Review** (sequential, same worktree - deliberate: the
125
128
  verification command may write to the tree while the Reviewer reads it):
126
129
 
127
- - Run verification-brief.md Section B: the resolved verification command under its
128
- safety contract - self-contained and non-interactive (no prompts; run under a
130
+ - Run verification-brief.md Section B: resolve the verification evidence per its
131
+ Evidence resolution table (green exact-head CI is the default evidence); run
132
+ the resolved verification command only when the table selects a fallback or
133
+ opt-out row, under its safety contract - self-contained and non-interactive (no prompts; run under a
129
134
  non-interactive environment), bounded by a timeout (default 15 minutes, `timeout
130
135
  minutes` override) via the first available mechanism: the harness's own bash
131
136
  timeout parameter, else the `timeout`/`gtimeout` CLI when installed, else a
132
137
  background-and-kill fallback - then material-claim checking against the PR body.
133
- After the run,
134
- the orchestrator asserts tracked-only cleanliness (`git status --porcelain
138
+ After a local run, the orchestrator asserts tracked-only cleanliness (`git status --porcelain
135
139
  --untracked-files=no` empty, equivalently `git diff --quiet && git diff --cached
136
140
  --quiet`; HEAD unmoved) - untracked gate artifacts, including the Verifier's
137
141
  `log_path`, are expected and do not fail this check as long as `log_path` sits
@@ -147,11 +151,16 @@ verification command may write to the tree while the Reviewer reads it):
147
151
  - **Provenance:** `worktree_root` matches the provisioned path, every `run_cwd` is
148
152
  inside it, `head_sha` matches the digest's `headRefOid`. On mismatch, re-fetch the
149
153
  PR head once and re-sync + re-run Phase 3 if it advanced; a second mismatch, or any
150
- path mismatch, is treated as missing evidence - not merge-ready. The claim stated in
151
- output is precisely "reproduced locally under the project's documented verification
152
- command" - nothing stronger; never worded to imply a deployed, staging, or CI
153
- environment.
154
- - **Evidence:** paste each run's `command` and `raw_tail` verbatim, fenced - never
154
+ path mismatch, is treated as missing evidence - not merge-ready. The claim stated in output names its source. Local path: precisely "reproduced
155
+ locally under the project's documented verification command" - nothing
156
+ stronger; never worded to imply a deployed, staging, or CI environment.
157
+ CI path (`source: ci`): precisely
158
+ `verified by CI: <check name(s)> succeeded on <sha> (run <url>)`
159
+ - never phrased as local reproduction, never implying the local command ran;
160
+ `<sha>` is the assessed `headRefOid`, `<url>` degrades to `unavailable` when
161
+ absent. Provenance checks on `worktree_root`/`run_cwd` bind only to the local
162
+ path.
163
+ - **Evidence:** On the CI path, list each satisfying check's name, conclusion, assessed SHA, and run URL - there is no command or raw_tail to paste. On the local path, paste each run's `command` and `raw_tail` verbatim, fenced - never
155
164
  paraphrased. Any authored summary is labeled as a summary and never substitutes for
156
165
  `raw_tail`.
157
166
  - **Severity translation:** Critical -> blocking, Moderate -> blocking, Minor ->
@@ -165,18 +174,25 @@ verification command may write to the tree while the Reviewer reads it):
165
174
  claim is a blocking finding. An `unverifiable-pre-merge` claim used as merge proof
166
175
  (appears in the PR body's evidence/result/test-plan content) is blocking; stated as
167
176
  an explicit post-merge observation instead, it is a non-blocking follow-up.
168
- - **Required CI checks:** a failing **required** status check withholds merge from
169
- every pre-composed course until the user explicitly dispositions it - flaky
170
- (proceed via the custom row) or real (it blocks); it mints a `P#`. A **pending**
171
- required check (still running - the normal case, not a defect) is **wait-until-
172
- green, not dispositionable**: it mints no `P#`, is never flaky/real-dispositioned,
173
- and the withhold auto-lifts the moment it turns green - or, if it instead fails,
174
- converts into an undispositioned failing check with its own `P#` at that point.
175
- While pending, the report notes it under Evidence and every merge course simply
176
- does not render (a pending-only PR is not a blocking verdict - findings groups can
177
+ - **CI checks:** any blocking conclusion in the resolved check set (required or
178
+ not - see the brief's Evidence resolution table) withholds merge from every
179
+ pre-composed course until the user explicitly dispositions it, and mints a
180
+ `P#`. Three dispositions: **flaky** (proceed via the custom row), **real** (it
181
+ blocks until green), **CI-infrastructure-broken** (the checks themselves are
182
+ untrustworthy: triggers the fallback local run, and merge stays withheld until
183
+ that fallback produces green evidence). A **pending** required check (still
184
+ running - the normal case, not a defect) is **wait-until-green, not
185
+ dispositionable**: it mints no `P#`, is never dispositioned, and the withhold
186
+ auto-lifts the moment it turns green - or, if it instead fails, converts into
187
+ an undispositioned failing check with its own `P#` at that point. While
188
+ pending, the report notes it under Evidence and every merge course simply does
189
+ not render (a pending-only PR is not a blocking verdict - findings groups can
177
190
  all read "None" - the recommended course falls to `stop` or `review-comment`,
178
- never a merge course, until it resolves). Non-required checks are informational,
179
- listed in Evidence only.
191
+ never a merge course, until it resolves). The evidence decision is
192
+ independent: a green check elsewhere in the resolved set still satisfies
193
+ verification evidence while a pending required check withholds merge. The
194
+ CI-sufficient path changes no consent surface: still read-only, no auto-merge,
195
+ no posting, no menu change beyond the third disposition.
180
196
  - **Doc drift:** when the review finds committed doc drift as a **blocking** finding,
181
197
  the orchestrator applies the doc fixes itself, in the provisioned worktree (created
182
198
  or reused), as part of assessment - real edits, uncommitted, worktree-local. The
@@ -215,9 +231,9 @@ issue is linked, committed doc drift, anything the merged rubric maps to blockin
215
231
  **follow-ups only** (never gate merge), or **clean**.
216
232
 
217
233
  **Merge preconditions** (all must hold): gate green with every blocking finding fixed,
218
- not deferred; `mergeable == MERGEABLE` (`UNKNOWN` after the one post-provision
219
- re-poll withholds merge, same as `CONFLICTING`); no undispositioned failing or
220
- pending required check; evidence pasted with clean provenance; worktree clean and synced with the remote
234
+ not deferred; verification evidence present per the brief's Evidence resolution table (a CI claim or a green local run - a table-sanctioned CI skip is evidence, not missing "not run" evidence; result: not run blocks only when the table required a fallback run that didn't happen); `mergeable == MERGEABLE` (`UNKNOWN` after the one post-provision
235
+ re-poll withholds merge, same as `CONFLICTING`); no undispositioned failing check in
236
+ the resolved set, no pending required check; evidence pasted with clean provenance; worktree clean and synced with the remote
221
237
  head (fixes pushed first); explicit selection with a head compare-and-swap that
222
238
  passes. A merge selection while any precondition fails is refused, naming the failing
223
239
  precondition, and the menu re-renders - never a dead end, never a silent merge. Merge
@@ -253,7 +269,7 @@ findings, menu. No restating diffs, no narration, no recap prose.
253
269
  <one line + the deciding factor>
254
270
 
255
271
  ## Evidence
256
- <verbatim command + raw_tail per run; claims checked; CI rollup with required-check disposition>
272
+ <CI path: satisfying check name(s)/conclusion/sha/url; local path: verbatim command + raw_tail per run; claims checked; resolved-set check dispositions>
257
273
 
258
274
  ## Findings (blocking)
259
275
  Blocking findings (P#):
@@ -295,8 +311,8 @@ Requirement/doc drift (linked issue, committed doc drift, or spec conflict):
295
311
  nothing. Code-level spec bugs (the diff contradicts the spec) are `P#` `[spec]`;
296
312
  requirement/doc mismatches (the spec or docs are stale relative to intent) are
297
313
  `L#`.
298
- - **Required checks close by disposition, not by fix:** an undispositioned failing
299
- required check is `P#` `[test]` referencing the check name; it is never a target
314
+ - **Failing checks close by disposition, not by fix:** an undispositioned failing check in the resolved set
315
+ is `P#` `[test]` referencing the check name; it is never a target
300
316
  of a worktree `fix`. The user's Phase-4 disposition annotates the same ID rather
301
317
  than closing it outright: dispositioned **flaky** -> annotate
302
318
  `(dispositioned: flaky)`; this annotation excepts the `P#` from the unfixed-blocker
@@ -304,6 +320,9 @@ Requirement/doc drift (linked issue, committed doc drift, or spec conflict):
304
320
  "every blocking finding fixed", and the merge path is Phase 4's explicit flaky
305
321
  disposition via the custom row. Dispositioned **real** -> annotate
306
322
  `(dispositioned: real)` and the `P#` keeps blocking until the check is green.
323
+ Dispositioned **CI-infrastructure-broken** -> annotate
324
+ `(dispositioned: ci-infrastructure-broken)`; the fallback local run executes,
325
+ and the `P#` keeps blocking until that fallback is green.
307
326
  - **Severity is decided at triage, not by the category tag:** a finding lands in
308
327
  `P#` only when it must be fixed before merge (correctness, security, material
309
328
  performance trap, a convention the repo enforces); improvements that don't
@@ -363,7 +382,7 @@ pre-composed or custom, bundles a push-producing action (`fix`, `push-docs`) wit
363
382
  | Author | State | Courses (first = `[recommended]`) |
364
383
  |---|---|---|
365
384
  | you | clean / follow-ups only | 1. merge-squash; 2. merge-commit; 3. stop; 4. review-comment (post no-blockers note) |
366
- | you | blocking | 1. fix (worktree-fixable P#s only - `all` covers only those) [+ push-docs when uncommitted doc edits exist]; 2. push-docs (alone, when doc edits exist); 3. stop; 4. review-comment (post findings). When no P# is worktree-fixable (blocking is required-check-only or L#-only), course 1 (fix) is not rendered: push-docs becomes first when doc edits exist, else stop is first |
385
+ | you | blocking | 1. fix (worktree-fixable P#s only - `all` covers only those) [+ push-docs when uncommitted doc edits exist]; 2. push-docs (alone, when doc edits exist); 3. stop; 4. review-comment (post findings). When no P# is worktree-fixable (blocking is failing-check-only or L#-only), course 1 (fix) is not rendered: push-docs becomes first when doc edits exist, else stop is first |
367
386
  | you | blocking, post-fix re-render (gate green, preconditions hold) | 1. merge-squash; 2. merge-commit; 3. stop; 4. review-comment |
368
387
  | someone else | clean / follow-ups only | 1. approve; 2. merge-squash (offered-unrecommended); 3. review-comment (no-blockers note) |
369
388
  | someone else | blocking | 1. request-changes; 2. fix all (courtesy, their branch - omitted when nothing is worktree-fixable); 3. reply <C#s> (omitted when the `C#` group is None); 4. review-comment |
@@ -382,18 +401,18 @@ also dropped (never offered on your own PR) - fork|you|clean renders
382
401
  branch course is also absent, since it is your own PR). A fork PR authored by someone
383
402
  else uses the someone-else cells above with `fix`/`push-docs`/`merge-*` removed.
384
403
 
385
- **Required-check gate on merge courses:** an undispositioned failing **or pending**
386
- required check withholds every pre-composed course containing `merge-*` (per the
404
+ **CI-check gate on merge courses:** an undispositioned failing check in the resolved
405
+ set, or a pending **required** check, withholds every pre-composed course containing `merge-*` (per the
387
406
  Verdict merge preconditions) - none render, whatever the author/state cell says. A
388
407
  pending check mints no `P#` and is wait-until-green, not dispositionable (see Phase
389
- 4); a failing one mints a `P#` and takes a disposition. A **flaky** disposition does
408
+ 4); a failing one mints a `P#` and takes a disposition (flaky / real / CI-infrastructure-broken). A **flaky** disposition does
390
409
  not restore merge to a pre-composed course; merge proceeds only via the custom row
391
410
  naming the disposition explicitly. A **real** disposition, or an unresolved pending
392
- check, keeps every merge course withheld until the check is green - a pending-only
411
+ check, keeps every merge course withheld until the check is green; a **CI-infrastructure-broken** disposition keeps them withheld until the triggered fallback run is green - a pending-only
393
412
  render is not itself a blocking verdict (findings groups may all read "None"); the
394
413
  recommended course falls to `stop` or `review-comment` in the meantime. This never
395
- falls through to the clean cell's recommended `merge-squash` - a required-check
396
- failure or pend means the PR is not in the clean state to begin with.
414
+ falls through to the clean cell's recommended `merge-squash` - a failing resolved-set
415
+ check or a pending required check means the PR is not in the clean state to begin with.
397
416
 
398
417
  Rows a cell offers but GitHub would refuse (branch protection, missing permission)
399
418
  render listed-but-unavailable with the reason. Zero mutation courses is a legal
@@ -434,7 +453,7 @@ The menu is a state machine, not a one-shot report:
434
453
  next CAS check runs against the new head on the next external write.
435
454
  2. Execute only the selected course. **Fix wave** (`fix <set>`): first filter the
436
455
  selected set to worktree-fixable `P#`s - drop any `P#` closed by disposition
437
- (an undispositioned required-check failure is never a `fix` target; a **flaky**
456
+ (an undispositioned failing-check `P#` is never a `fix` target; a **flaky**
438
457
  disposition already excepts it) - and route file-less `P#`s (a claim or a gate
439
458
  command as `source_ref`, no draft touching a file) to run inline/sequentially,
440
459
  never as part of a parallel file-batch.
@@ -463,7 +482,7 @@ The menu is a state machine, not a one-shot report:
463
482
  Once every dispatched/inline batch returns, the orchestrator commits the golden
464
483
  course as one local commit set - the code fixes plus any already-applied
465
484
  reviewed doc edits selected alongside them (one commit, or one per batch
466
- sequentially; subjects name the fixes) - then re-runs the gate **once**. **On
485
+ sequentially; subjects name the fixes) - then re-resolves the evidence for the new head **once** (the brief's stale-head row: prior evidence is stale; the local command executes only on a fallback/opt-out resolution). **On
467
486
  green**, push **once**; gate and push are per-wave invariants, never per-fix or
468
487
  per-batch. **On red**, do not push: leave the commit(s) local, re-render with
469
488
  the unresolved `P#`s still open, and warn that unpushed fix commits sit in the
@@ -474,9 +493,9 @@ The menu is a state machine, not a one-shot report:
474
493
  with the drafted payload for the selected IDs.
475
494
  3. After any mutation that can change readiness (fix wave pushed, docs pushed, PR
476
495
  head moved), re-run the claim-check and Review on the synced worktree: claims
477
- are re-checked against the new head and findings are re-rendered, but the
478
- verification command itself is **not** re-executed here - step 2's gate run
479
- already was the wave's one and only execution of it. Re-render the report:
496
+ are re-checked against the new head and findings are re-rendered, but
497
+ the verification command itself is **not** re-executed here - step 2's evidence re-resolution
498
+ already was the wave's one and only gate pass. Re-render the report:
480
499
  each selected `P#`/`L#` confirmed resolved is annotated `(fixed in <sha>)`
481
500
  under its original ID; unresolved ones stay open unchanged; new findings
482
501
  continue the sequence. Merge, if now available, renders as row 1.
@@ -512,14 +531,14 @@ overrides file - see Project overrides.
512
531
  - Any mutation (fix, push, review, merge) without an explicit menu selection
513
532
  - Pasting paraphrased evidence instead of verbatim `raw_tail`
514
533
  - A provenance mismatch (worktree, `run_cwd`, or `head_sha`) noticed and ignored
515
- - Merging around an undispositioned blocking finding or required-check failure
534
+ - Merging around an undispositioned blocking finding or failing-check `P#`
516
535
  - Reading configuration (rubric, verification command, or ladder sources) from the
517
536
  PR's head instead of the base branch's merge-base
518
537
  - Renumbering or reusing a finding ID between menu rounds
519
538
  - Presenting findings without IDs, a blocking verdict with no `P#`/`L#`, or a
520
539
  `## Decision` rendered without its action vocabulary
521
540
  - Treating `[quality]` or `[performance]` as a downgrade signal on a `P#` - only
522
- an explicit Phase-4 flaky disposition excepts a required-check `P#` from the
541
+ an explicit Phase-4 flaky disposition excepts a failing-check `P#` from the
523
542
  unfixed-blocker set, never a category tag
524
543
  - A course (pre-composed or custom) bundling a push-producing action with
525
544
  `merge-*`
@@ -57,7 +57,7 @@ invent ACs.
57
57
  provisioning, so it never re-polls; the orchestrator re-polls once after
58
58
  provisioning the worktree (see SKILL.md Phase 2) and treats a still-`UNKNOWN`
59
59
  result as not merge-ready. Bot author noted
60
- (`author_is_bot`). Capture each status check's `isRequired` where exposed.
60
+ (`author_is_bot`). Capture each status check's `isRequired` where exposed (digest field: `required`).
61
61
 
62
62
  **Gather digest output schema (normative):**
63
63
 
@@ -65,7 +65,7 @@ result as not merge-ready. Bot author noted
65
65
  - pr: { number, title, body, author, author_is_bot, state, isDraft, headRefName, baseRefName,
66
66
  isCrossRepository, mergeable, headRefOid, files, additions, deletions, reviewDecision }
67
67
  - viewer: { login, is_author, permission }
68
- - status_checks: [ { name, status, conclusion, required } ] # informational except required-failing
68
+ - status_checks: [ { name, status, conclusion, required, url } ] # evidence semantics: Section B Evidence resolution
69
69
  - comments: { inline[], top_level[], review_threads[]? }
70
70
  - issue: { ref, title, body, acceptance_criteria[], comments[] } | null
71
71
  - worktree_discovery: { expected_path, exists, branch, dirty, ahead, behind }
@@ -73,18 +73,57 @@ result as not merge-ready. Bot author noted
73
73
  ```
74
74
 
75
75
  `viewer_is_author` lives at `viewer.is_author` in the digest, computed as
76
- `viewer.login == pr.author.login`. `status_checks` splits `required` vs
77
- non-required per entry - only a failing or pending required check withholds
78
- merge (see the orchestrator's required-check rule in SKILL.md Phase 4);
79
- non-required checks are informational.
76
+ `viewer.login == pr.author.login`. Each `status_checks` entry's `url` is the CheckRun `detailsUrl` / StatusContext
77
+ `targetUrl` already present in the fetched payload; when the payload omits it,
78
+ downstream CI claims record `url: unavailable` - absence never disqualifies the
79
+ check. GraphQL enums are case-folded; a StatusContext's `state` is its
80
+ conclusion, with `ERROR` blocking and `PENDING` pending. Missing `required` is
81
+ treated as non-required. `ci checks:` matches check name, workflow name, or
82
+ status context, trimmed, case-insensitive. What checks mean for verification evidence is owned by
83
+ Section B's Evidence resolution table; what they mean for merge is owned by the
84
+ orchestrator's required-check rule (SKILL.md Phase 4) - two independent
85
+ consumers of the same data.
80
86
 
81
87
  ## Section B - Verifier
82
88
 
83
- Runs the resolved verification command inside the provisioned worktree, then
84
- claim-checks the PR body against what actually ran. Report only - do not
89
+ Resolves the verification evidence per the table below - green exact-head CI is
90
+ the default evidence; the resolved verification command runs inside the
91
+ provisioned worktree only when the table selects a fallback or opt-out row -
92
+ then claim-checks the PR body. Report only - do not
85
93
  edit, fix, or commit anything; you are running a gate and claim-checking,
86
94
  not implementing.
87
95
 
96
+ **Evidence resolution (normative).** The single rule for whether the local
97
+ command runs. Inputs come from Section A's existing `gh pr view` call - no
98
+ second fetch.
99
+
100
+ - **Resolved check set** = checks named by `ci checks:` if configured, else all
101
+ checks on the assessed `headRefOid`.
102
+ - **Conclusion semantics**: `success` satisfies; `failure`/`timed_out`/
103
+ `action_required`/`error` block; `neutral`/`skipped`/`cancelled`/`stale`/
104
+ `startup_failure` are inert; a check with `status != completed` is pending; a
105
+ completed check with a missing/unreadable conclusion cannot satisfy
106
+ (fail-safe).
107
+
108
+ Row precedence is top-down: the first matching row wins.
109
+
110
+ | Path | Trigger | Action | Evidence recorded |
111
+ |---|---|---|---|
112
+ | Opt-out | `local verification: always` in `## PR gate` | Run local command unconditionally; a Failed-CI block below still applies independently | Local, as today |
113
+ | Failed CI | Any blocking conclusion in resolved set | Blocks: mints a `P#` (any resolved-set failure, required or not). A green local run never overrides it. Only an explicit human CI-infrastructure-broken disposition triggers the fallback run; merge stays withheld until the fallback produces green evidence | The disposition; plus the fallback run's result only when CI-infrastructure-broken triggered one |
114
+ | CI-sufficient | >=1 `success` in resolved set | Skip local run | CI claim: check name(s), conclusion, assessed SHA, run URL |
115
+ | Pending | Zero `success` and >=1 pending check in resolved set | Evidence decision waits until the set reaches a completed conclusion - never a fallback trigger, never an evidence-less merge; merge is withheld as missing evidence until the table re-resolves | n/a (waiting) |
116
+ | Fallback | No checks on assessed head, or zero `success` with none pending (all inert / fail-safe) | Run local command (protocol below, unchanged) | Local command + raw tail, existing provenance rules |
117
+ | Stale head | Head advances since evidence was resolved (e.g. a fix-wave push); fires across runs - a single gather is same-head by construction | All prior evidence (CI or local) is stale; re-resolve this table for the new head before merge is offered | Fresh evidence for the new head |
118
+
119
+ CI-sufficient predicate, stated once (the rows implement exactly this): **>=1
120
+ completed `success`, zero blocking conclusions, no opt-out.** Pending checks are
121
+ excluded from the predicate - they neither satisfy nor veto it. The evidence
122
+ decision ("run local?") and the merge decision ("can this merge?") are separate
123
+ consumers of the same `status_checks` data: a green non-required check satisfies
124
+ evidence even while a pending required check blocks merge under the existing
125
+ wait rule. The only evidence-path wait is the Pending row's zero-success case.
126
+
88
127
  **Safety contract:**
89
128
 
90
129
  - Timeout default 15 minutes, overridable by the resolved `timeout minutes`
@@ -95,7 +134,9 @@ not implementing.
95
134
  non-interactive.
96
135
  - If the resolved config states `requires credentials: true`, do not run
97
136
  the command; report "verification requires credentials, not run" as
98
- missing evidence instead of prompting for secrets.
137
+ missing evidence instead of prompting for secrets; this arises only when
138
+ the table selected a fallback or opt-out row - a CI-sufficient resolution
139
+ needs no credentials.
99
140
  - Capture full output to a `log_path` inside the (disposable) worktree, under a
100
141
  gitignored path (e.g. `.worktrees/pr-<N>/.gatekeep-logs/`) so it never counts as a
101
142
  tracked change; keep only the last ~100 lines verbatim in the digest as `raw_tail`.
@@ -103,19 +144,27 @@ not implementing.
103
144
  **Verifier output schema (normative):**
104
145
 
105
146
  ```text
106
- - worktree_root: <absolute path>
107
- - head_sha: <git rev-parse HEAD at run time>
108
- - runs: [ { run_cwd, command (verbatim), exit_code, result: pass|fail|not run,
109
- raw_tail: <last ~100 lines of combined stdout+stderr, verbatim, fenced>,
110
- log_path: <file inside the worktree holding the full captured output> } ]
111
- - claims: [ { claim, disposition: matched|contradicted|unverifiable-pre-merge, evidence } ]
147
+ - source: ci | local
148
+ - source: ci ->
149
+ head_sha: <the assessed headRefOid>
150
+ checks: [ { name, conclusion, url } ] # url: unavailable when the payload omits it
151
+ (no command, no raw_tail - nothing ran locally)
152
+ - source: local ->
153
+ worktree_root: <absolute path>
154
+ head_sha: <git rev-parse HEAD at run time>
155
+ runs: [ { run_cwd, command (verbatim), exit_code, result: pass|fail|not run,
156
+ raw_tail: <last ~100 lines of combined stdout+stderr, verbatim, fenced>,
157
+ log_path: <file inside the worktree holding the full captured output> } ]
158
+ - claims: [ { claim, disposition: matched|contradicted|unverifiable-pre-merge, evidence } ] # both sources
112
159
  ```
113
160
 
114
- `raw_tail` is captured output, not authored prose; anything written in your
161
+ Local provenance and raw-tail rules bind only to source: local. `raw_tail` is captured output, not authored prose; anything written in your
115
162
  own words is labeled `summary` and must never be pasted in place of
116
163
  `raw_tail`.
117
164
 
118
- **Material-claim check.** After the run, claim-check the PR body -
165
+ **Material-claim check.** Always runs, on both sources - on the CI path,
166
+ dispositions are judged against the recorded CI evidence and the diff; on the
167
+ local path, against the run and the diff. Claim-check the PR body -
119
168
  **material claims only** (test/verification/behavior assertions: "added
120
169
  X", "tests cover Y", "fixed Z"), not qualitative prose. Disposition each
121
170
  claim as one of:
@@ -130,7 +179,7 @@ proof* (it appears in the PR body's evidence/result/test-plan content) is
130
179
  blocking; the same claim stated as an explicit post-merge observation is
131
180
  non-blocking follow-up only.
132
181
 
133
- After the run, the orchestrator asserts tracked-only cleanliness
182
+ After a local run (`source: local` only), the orchestrator asserts tracked-only cleanliness
134
183
  (`git status --porcelain --untracked-files=no` empty, equivalently
135
184
  `git diff --quiet && git diff --cached --quiet`; HEAD unmoved); untracked gate
136
185
  artifacts - including `log_path` itself, provided it sits under a gitignored path
@@ -171,9 +220,12 @@ blocking/follow-up happens later, at integration.
171
220
  judges against stated intent only, never inventing ACs; scope-creep findings
172
221
  do not apply.
173
222
  - No resolvable verification command (ladder exhausted, user asked, user
174
- declines): the gate runs without local verification evidence; record
175
- `result: not run` in the Verifier output. Missing evidence blocks merge
176
- the same as a failed gate - the PR is not merge-ready.
223
+ declines) **when the Evidence resolution table selected a fallback or opt-out
224
+ row**: the gate runs without local verification evidence; record
225
+ `result: not run` in the Verifier output. Missing evidence blocks merge the
226
+ same as a failed gate - the PR is not merge-ready. A table-sanctioned CI skip
227
+ (`source: ci`) is evidence, never missing evidence, and needs no command at
228
+ all.
177
229
  - Linked-issue fetch fails (tracker unreachable, bad ref): proceed judging
178
230
  against the PR's stated intent, mark `issue: null` in the digest plus a
179
231
  truncation/availability note explaining why, and never invent ACs; AC
@@ -55,12 +55,14 @@ For each task in `plan_tracker`:
55
55
 
56
56
  1. **Dispatch implementer.** Pass the full task text + scene-setting context + the task's plan-declared test commands as `SCOPED_TEST_COMMANDS` (or `none`). Don't make the subagent re-read the plan.
57
57
  2. **Handle implementer status** (see below).
58
- 3. **Dispatch spec reviewer.** Verify the diff matches the spec — nothing missing, nothing extra.
58
+ 3. **Dispatch spec reviewer.** Pass the task text, the patch diff, the absolute spec path, and the task's `**Spec:**` anchors — SR reads the anchored ranges from the spec file itself; never inline spec excerpts. The spec wins every dispute; the authority hierarchy and finding labels live in `./spec-reviewer-prompt.md`. Anchor-less tasks (no `**Spec:**` line): task text alone is the contract. Verify the diff matches the anchored spec — nothing missing, nothing extra.
59
59
  4. If spec reviewer finds gaps → re-dispatch implementer to fix → re-review. Loop until ✅, within [Fix-Loop Rounds](#fix-loop-rounds).
60
60
  5. **Dispatch code-quality reviewer.** Only after spec is ✅. Skip for doc-only tasks (every file in the task's `Files:` block documentation-only) — SR-only, same exemption as doc-only waves. Pass `SCOPED_TEST_COMMANDS` = the task's plan-declared commands.
61
61
  6. If quality reviewer finds issues → re-dispatch implementer → re-review. Loop until ✅, within [Fix-Loop Rounds](#fix-loop-rounds).
62
62
  7. Mark task complete in `plan_tracker`.
63
63
 
64
+ The spec is frozen at plan time and the orchestrator is its only writer during execution; if you do edit it mid-run, re-run writing-plans' anchor-resolution check before the next wave. An SR unable to read the spec at a cited anchor (missing file, unresolvable heading/range) returns a blocking finding — the contract is spec+task or stop, never a silent fallback to task-only review.
65
+
64
66
  After all tasks: run the whole-diff code review (`requesting-code-review`). Then [After All Tasks](#after-all-tasks-complete).
65
67
 
66
68
  ## Fix-Loop Rounds
@@ -133,7 +135,7 @@ When in doubt, default. Don't downgrade reviewers — false negatives are expens
133
135
  subagent({ agent: "implementer", task: "<task text + context + SCOPED_TEST_COMMANDS + status protocol>" })
134
136
 
135
137
  // spec compliance
136
- subagent({ agent: "spec-reviewer", task: "<diff range + spec excerpt + ask: does this match?>" })
138
+ subagent({ agent: "spec-reviewer", task: "<task text + patch diff + absolute spec path + task's Spec: anchors>" })
137
139
 
138
140
  // code quality
139
141
  subagent({ agent: "code-reviewer", task: "<diff range + SCOPED_TEST_COMMANDS (task commands; wave: union; whole-diff: none) + ask: production-ready?>" })
@@ -169,7 +171,7 @@ Auto-selected at handoff by `writing-plans` (any wave with ≥2 tasks) when the
169
171
 
170
172
  1. **Independence check.** Parse the wave's tasks' `Files:` blocks; assert pairwise-disjoint (mechanical). Runtime-resource disjointness (DB/schema, port, fixture, external service, shared temp path) is not machine-checkable here — trust the plan's wave grouping, which `writing-plans`' D5 contract guarantees. Either kind of overlap → the wave is mis-grouped; run those tasks as sequential single-task waves and note it.
171
173
  2. **Fan out.** One parallel dispatch (shape below): `implementer` per task, `context: "fresh"`, `worktree: true`. Each returns a status + a patch.
172
- 3. **Status + spec review per task.** Parse each `DONE`/`BLOCKED`/etc. (see [Implementer Status](#implementer-status)) **first**. Then **dispatch a `spec-reviewer` per accepted patch** (`DONE`, or a `DONE_WITH_CONCERNS` you proceeded with) in one parallel fan-out — `context: "fresh"`, `cwd: <this worktree>`, **no `worktree` flag** (read-only) — each passed its task text, the returned **patch diff**, and the absolute spec path. Review is **diff-based**: the diff's hunks carry `file:line`, and test execution is never the reviewer's job - in either mode (persona rule; the wave test gate in step 5 runs the wave's declared test commands). Inline verdicts are fine at normal wave sizes; large waves use `output:` + `outputMode: "file-only"` to keep verdicts out of your context. **Re-dispatch by cause:** `BLOCKED`/`NEEDS_CONTEXT` per the [Implementer Status](#implementer-status) matrix; a **spec gap** re-dispatches the implementer (fresh, `worktree: true`) carrying the prior patch + the reviewer's findings, the new patch superseding the old at step 4. Loop until accepted + spec ✅, within [Fix-Loop Rounds](#fix-loop-rounds), same as sequential.
174
+ 3. **Status + spec review per task.** Parse each `DONE`/`BLOCKED`/etc. (see [Implementer Status](#implementer-status)) **first**. Then **dispatch a `spec-reviewer` per accepted patch** (`DONE`, or a `DONE_WITH_CONCERNS` you proceeded with) in one parallel fan-out — `context: "fresh"`, `cwd: <this worktree>`, **no `worktree` flag** (read-only) — each passed its task text, the returned **patch diff**, the absolute spec path, and the task's `**Spec:**` anchors — SR reads the anchored ranges itself (never inline excerpts; authority hierarchy in `./spec-reviewer-prompt.md`; anchor-less tasks are task-text-only). Review is **diff-based**: the diff's hunks carry `file:line`, and test execution is never the reviewer's job - in either mode (persona rule; the wave test gate in step 5 runs the wave's declared test commands). Inline verdicts are fine at normal wave sizes; large waves use `output:` + `outputMode: "file-only"` to keep verdicts out of your context. **Re-dispatch by cause:** `BLOCKED`/`NEEDS_CONTEXT` per the [Implementer Status](#implementer-status) matrix; a **spec gap** re-dispatches the implementer (fresh, `worktree: true`) carrying the prior patch + the reviewer's findings, the new patch superseding the old at step 4. Loop until accepted + spec ✅, within [Fix-Loop Rounds](#fix-loop-rounds), same as sequential.
173
175
  4. **Integrate.** `git apply` each task's patch sequentially onto HEAD. Apply fails = textual conflict → drop that task, finish the rest, re-run the dropped task sequentially on the updated HEAD.
174
176
  5. **Test gate.** Run the union of the wave's tasks' declared test commands on the integrated tree — the full verification set is the verify phase's job, run once. Failure = semantic conflict or bug → re-run the offending task sequentially, else fix per [When a Subagent Fails](#when-a-subagent-fails).
175
177
  6. **Quality review.** CR binds to the wave: exactly one **initial** code-review dispatch per code-touching wave, over the integrated wave diff - never per task within a wave, never batched across waves. Subsequent dispatches within the wave are re-reviews triggered only by findings, per Fix-Loop Rounds. Pass `SCOPED_TEST_COMMANDS` = the union of the wave's tasks' declared commands. Code-quality review on the integrated wave diff; loop fixes to ✅ within [Fix-Loop Rounds](#fix-loop-rounds), same as sequential. Skip for doc-only waves (SR-only per the commit precondition below).
@@ -14,10 +14,26 @@ Dispatch a subagent with this prompt:
14
14
 
15
15
  [FULL TEXT of task requirements]
16
16
 
17
+ (The task text is a derivative of the spec — a lossy projection into an executable unit. See ## Spec Authority below.)
18
+
17
19
  ## What Implementer Claims They Built
18
20
 
19
21
  [From implementer's report]
20
22
 
23
+ ## Spec Authority
24
+
25
+ Spec: [absolute spec path]
26
+ Anchors: [the task's **Spec:** anchor list, e.g. § "Design" L34-L37 — or "omitted: anchor-less mechanical task"]
27
+
28
+ The spec is the sole authority — human-approved; the task never wins a dispute. Read the anchored ranges from the spec file yourself. Requirements in scope are ONLY the cited anchor ranges; do not extract, review, or flag the rest of the spec file.
29
+
30
+ - **Correctness / wording / completeness:** judged against the anchored spec lines. The spec wins every dispute.
31
+ - **Scope ("nothing more"):** the boundary is the anchor set — the slice of spec this task owns. Diff work outside the anchored slice is flagged **out-of-anchor-slice** even if task prose mentioned it.
32
+ - **Plan transcription gap:** spec-required work inside the anchored slice that is missing from the diff because the task prose omitted it — the requirement still binds; flag it. Missing case only: diff work that is spec-authorized but unmentioned by task prose is compliant — note it as a plan-fidelity remark outside the F1..Fn finding stream, never as a finding.
33
+ - **Task-vs-spec divergence** (task says X, anchored spec says Y): unconditional flag; quote the spec literal with spec file:line so the fix re-dispatch carries authoritative wording. Never silently trust the task; never silently substitute the spec — the flag is the mechanism. Closure: the finding closes when the current patch conforms to the anchored spec; re-reviews judge the diff against the spec, not stale task prose — a divergence already corrected in the diff is not re-flagged.
34
+ - **Anchor-less task** (Anchors: omitted): the task text alone is your contract; no out-of-anchor-slice or transcription-gap flagging — only nothing-extra-vs-the-chore review.
35
+ - **Finding grammar:** divergence findings use the existing F1..Fn finding grammar - a finding kind by prose label, not a new schema; the `Parallel-safe:` and `TRAJECTORY:` grammars are untouched.
36
+
21
37
  ## CRITICAL: Do Not Trust the Report
22
38
 
23
39
  The implementer finished suspiciously quickly. Their report may be incomplete,
@@ -51,11 +67,13 @@ Dispatch a subagent with this prompt:
51
67
  - Did they implement everything that was requested?
52
68
  - Are there requirements they skipped or missed?
53
69
  - Did they claim something works but didn't actually implement it?
70
+ - Anchored spec work absent from the diff because task prose omitted it? Label it "plan transcription gap".
54
71
 
55
72
  **Extra/unneeded work:**
56
73
  - Did they build things that weren't requested?
57
74
  - Did they over-engineer or add unnecessary features?
58
75
  - Did they add "nice to haves" that weren't in spec?
76
+ - Diff work outside the task's anchor slice? Label it "out-of-anchor-slice" (distinct from a spec-declared non-goal).
59
77
 
60
78
  **Misunderstandings:**
61
79
  - Did they interpret requirements differently than intended?
@@ -204,6 +204,8 @@ Each task uses `- [ ]` checkbox steps so execution tools (and humans) can track
204
204
 
205
205
  **TDD scenario:** [New feature — full TDD cycle | Modifying tested code — run existing tests first | Trivial change — use judgment]
206
206
 
207
+ **Spec:** doc/specs/<file>.md § "<heading>" L<start>-L<end>
208
+
207
209
  **Files:**
208
210
  - Create: `exact/path/to/file.py`
209
211
  - Modify: `exact/path/to/existing.py:123-145`
@@ -249,6 +251,27 @@ Each task uses `- [ ]` checkbox steps so execution tools (and humans) can track
249
251
 
250
252
  Every code task carries this step (red -> green -> fmt/lint -> commit). Doc-only tasks omit it unless the project formats Markdown.
251
253
 
254
+ **Anchor rules.** The task's `**Spec:**` line cites the plan header's spec path; multiple anchors sit comma-separated on one line (`§ "A" L10-L18, § "C" L40-L44`). Checks key on the `§` marker, so the header's path-only `**Spec:**` line is never matched. Anchors are captured once against the gated spec at plan-writing time — the spec is frozen once planning starts. A task with no anchorable requirement (pure-mechanics chore) omits the `**Spec:**` line entirely (never `**Spec:** none`) and carries a mechanical-task row in `## Spec coverage` — silence is never valid.
255
+
256
+ ## Spec Coverage Table
257
+
258
+ Every plan ends with a `## Spec coverage` section — authored last, placed after all Task sections (owner IDs do not exist earlier). Build it extraction-first: walk the spec top to bottom and write one row per normative requirement **before** assigning owners — every Design imperative (Add/Remove/Keep/Replace-style directives, not any fixed lexical form), every Edge-cases rule, every Acceptance criterion, every Out-of-scope entry, and every non-none Documentation-impact entry. Then assign owners. Two row kinds:
259
+
260
+ ```markdown
261
+ ## Spec coverage
262
+
263
+ | anchor | requirement (short) | owner |
264
+ |---|---|---|
265
+ | § "Design" L34-L37 | anchor line in task template | Task 2 |
266
+ | § "Edge cases" L120 | stale anchor = blocking SR finding | Task 4, Task 5 |
267
+ | § "Out of scope" L131 | fix-round anchoring | waived: out of scope per spec |
268
+ | - | mechanical: release commit | Task 7 |
269
+ ```
270
+
271
+ - **Requirement rows:** anchor + short requirement + owner = task-ID list, or `waived: <reason>` **only when the spec itself marks the item out of scope**. A waiver on an in-scope normative requirement is a Self-Review failure — there is no human plan-review gate to catch it downstream.
272
+ - **Mechanical-task rows:** anchor `-`, requirement `mechanical: <short>`, owner = the task ID. One such row per anchor-less task.
273
+ - The table is plan-authoring-time only — never passed to implementer or reviewer dispatches.
274
+
252
275
  ## No Placeholders
253
276
 
254
277
  Every plan failure mode:
@@ -257,6 +280,7 @@ Every plan failure mode:
257
280
  - ❌ `# Implement the rest of the function` — incomplete code is invalid code.
258
281
  - ❌ "Add tests for edge cases" — name the edge cases.
259
282
  - ❌ "Wire it up to the existing system" — give file paths and call sites.
283
+ - ❌ "timeout/gtimeout ladder" when the spec fixes the literal `timeout 30` — never paraphrase an exact-string requirement (setting keys, error messages, banner/format strings, command names and invocations, API shapes); transcribe it as a backtick-quoted spec literal: `timeout 30`. Spec-side backtick spans containing `<placeholder>` segments are templates the plan instantiates, not exact-string requirements — exempt from quote integrity.
260
284
  - ❌ "Similar to Task N" — repeat the code. Implementers (and subagents with fresh context) may read tasks out of order; pointing at a sibling task is not a substitute for showing the code.
261
285
  - ❌ References to types, functions, methods, or fields not defined in any task in this plan. If it shows up in Task 5, it must be introduced by Task 1–4 or already exist in the codebase (with a file:line citation).
262
286
  - ❌ `[fill in]`, `<example>`, `xxx` markers anywhere in the doc.
@@ -266,9 +290,12 @@ If a decision is genuinely open, put it in an explicit **Open Questions** sectio
266
290
 
267
291
  ## Self-Review (Before Handoff)
268
292
 
269
- After drafting the plan and before announcing it complete, run three checks yourself. This is a checklist you run yourself — not a subagent dispatch.
293
+ After drafting the plan and before announcing it complete, run these checks yourself. This is a checklist you run yourself — not a subagent dispatch.
270
294
 
271
- - **Spec coverage.** Cross-reference the spec's components/decisions/constraints against the plan. Does every spec section map to one or more tasks? If a spec decision has no implementation task, the plan is missing work or the spec was overspecified. Each Documentation impact entry maps to a plan task (or explicit "none").
295
+ - **Table closure (three legs).** Every `## Spec coverage` row's owner is a task-ID list, a spec-authorized `waived: <reason>`, or a mechanical-task row; every `### Task N` heading appears in >=1 row; every requirement row's anchor is contained in the anchor set of each listed owner task's `**Spec:**` line. Zero orphans, zero waived in-scope normative rows, zero row-vs-owner anchor mismatches. Each Documentation impact entry maps to a plan task (or explicit "none").
296
+ - **Quote integrity (spec -> task).** For every non-waived requirement row, extract each backtick-quoted literal inside the row's anchored spec lines (strip the backticks; skip `<placeholder>` template spans) and `grep -F` it against the owning task's body — zero misses. Planner-authored backticks elsewhere in tasks are never scanned; the input set is spec-side literals only.
297
+ - **Anchor resolution.** For every task-level anchor (a `**Spec:**` line carrying `§`; the plan header's path line is exempt), the quoted heading text matches an ATX heading in the spec file and `L<start>-L<end>` is in-bounds, non-empty, and lies within that heading's section — zero unresolved anchors. Verify with `grep -n '^#'` plus a scoped `sed -n`. Ignore `#`-lines inside fenced code blocks when locating headings and section boundaries - a fenced markdown example is not a heading.
298
+ - **Paths exist.** Every `Modify:` path in `Files:` blocks passes `test -f` after stripping any trailing `:line[-line]` suffix; a `Modify:` glob must expand to >=1 match; `Create:` and `Test:` paths are exempt unless the `Test:` path is also listed under `Modify:`. Zero missing.
272
299
  - **Placeholder scan.** Grep the doc for `TODO`, `TBD`, `xxx`, `[fill in]`, `<example>`, `etc.`, "probably", "something like". Resolve or convert each into an explicit Open Question.
273
300
  - **Type / API consistency.** Function signatures and field names that appear in multiple tasks must match exactly. The plan is its own contract — internal contradictions surface as bugs during execution.
274
301
  - **Wave disjointness.** For every multi-task wave, confirm the tasks' `Files:` sets are pairwise disjoint **and** that no two tasks contend on a shared mutable runtime resource (DB/schema, port, fixture, external service, shared temp path). Either kind of overlap = mis-grouped wave; split or re-order before handoff.