tldr-experts 0.29.0 → 0.31.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (40) hide show
  1. package/CHANGELOG.md +602 -0
  2. package/README.md +2 -0
  3. package/dist/hooks/answer-capture.js +9 -8
  4. package/dist/hooks/budget-gate.js +9 -9
  5. package/dist/hooks/{chunk-erasdab5.js → chunk-04kpbhvk.js} +3 -3
  6. package/dist/hooks/{chunk-pne7afnf.js → chunk-27v2rwa9.js} +1 -1
  7. package/dist/hooks/{chunk-ggjxes7m.js → chunk-3r0hhgcf.js} +40 -7
  8. package/dist/hooks/{chunk-857jhjmb.js → chunk-5kv2vtha.js} +1 -1
  9. package/dist/hooks/{chunk-7xascvee.js → chunk-72et1xzx.js} +1 -1
  10. package/dist/hooks/{chunk-0d4csx4s.js → chunk-8ds2ret8.js} +1 -1
  11. package/dist/hooks/{chunk-qmp5p77c.js → chunk-9bd6b3xf.js} +1 -0
  12. package/dist/hooks/chunk-ae6bkfs5.js +0 -0
  13. package/dist/hooks/{chunk-t6gbx3jg.js → chunk-akemhezh.js} +5 -2
  14. package/dist/hooks/{chunk-gr83pq7x.js → chunk-f9d2efhb.js} +3 -3
  15. package/dist/hooks/{chunk-w3d9kxcp.js → chunk-me0kzc6k.js} +5 -2
  16. package/dist/hooks/{chunk-mg3aj0sa.js → chunk-q1ge3nex.js} +1 -1
  17. package/dist/hooks/{chunk-yanmz0vz.js → chunk-x1f4dxwk.js} +1 -1
  18. package/dist/hooks/{chunk-g4d2v4jc.js → chunk-yyhwf3bp.js} +1 -1
  19. package/dist/hooks/{chunk-z9hv47ps.js → chunk-zyz0fbz5.js} +1 -1
  20. package/dist/hooks/claim-sources.js +5 -5
  21. package/dist/hooks/dod-gate.js +7 -7
  22. package/dist/hooks/no-reask.js +8 -8
  23. package/dist/hooks/session-start.js +17 -12
  24. package/dist/hooks/statusline.js +9 -8
  25. package/dist/tldrx.js +3003 -1582
  26. package/package.json +1 -1
  27. package/plugin/.claude-plugin/plugin.json +1 -1
  28. package/stages/how/stage.md +0 -1
  29. package/stages/how/stage.yml +31 -2
  30. package/templates/experts/stack/dotnet.md +4 -3
  31. package/templates/experts/stack/javascript.md +4 -3
  32. package/templates/experts/stack/overlays/postgres-testcontainers.md +7 -2
  33. package/templates/experts/stack/python.md +4 -3
  34. package/templates/experts/stack/typescript.md +4 -3
  35. package/workflows/bugfix.yml +5 -1
  36. package/workflows/feature.yml +5 -1
  37. package/workflows/integration.yml +5 -1
  38. package/workflows/performance.yml +5 -1
  39. package/workflows/refactor.yml +5 -1
  40. package/workflows/upgrade.yml +5 -0
package/CHANGELOG.md CHANGED
@@ -1,5 +1,607 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.31.0 — 2026-09-16
4
+
5
+ ### Added
6
+
7
+ - **A `[src:]` citation on a Watch card that only has a punctuation slip — the marker missing its
8
+ space, a token sitting mid-sentence, an ASCII `->` in place of the real arrow — is repaired
9
+ locally, for $0.00, before the card is judged, instead of failing the stage over formatting a
10
+ script already knew how to fix (#345, watch-card half). Measured over two field audits: 10 of 54
11
+ failed tasks across 22 runs in one workspace were malformed watch cards — caught only after a
12
+ full paid turn wrote the card, by the same `parseWatcherCard` this repair now runs ahead of.
13
+ `repairSrcSyntax` (`src/core/text/srcToken.ts`) only accepts a fix it can VERIFY
14
+ re-parses clean (mirrors `parseYamlRepairing`'s own rule) and only touches the three rules that
15
+ are pure syntax — `marker-spelling`, `trailing-position`, `cmd-arrow` — never a rule that would
16
+ need an invented line number, path or fact id (AGENTS.md §7). The repair is written back to disk
17
+ whether or not the card ends up valid, and it is never silent: named in the stage's own report
18
+ when it clears the card, and in the failure reason when a genuine defect remains — the file on
19
+ disk already carries the fix, so a retry's `previousCard` prompt (`watchPrompt.ts`) shows a writer
20
+ only the part a script could not repair (owner decision 2026-09-15: auto-repair is visible, never
21
+ silent).
22
+
23
+ - **`how` is skipped when the seed already declares its own technical solution, and `plan` reads
24
+ that section instead (#346).** Owner decision, measured against three samples where `how`
25
+ reworked a seed that already carried the answer: 11.4-45% of a run's spend and 43-53% of its
26
+ own bytes went unread downstream. The trigger is one exact, documented marker — an H2 `##
27
+ Solution` or `## Technical approach` in a seed document — never inferred from prose:
28
+ `src/core/facilitator/seedSolution.ts` reuses `seed/seedClaims.ts`'s fence-aware heading reader
29
+ rather than a second regex, and exposes it to the existing `skip_if` grammar as a new
30
+ `seed_solution` predicate (`skipIf.ts`, `(stories|repos|questions|seed_solution)<op>N`). Since no
31
+ shipped `workflows/*.yml` uses the §2.4 record shape a workflow-level `skip_if` needs, a stage's
32
+ own top-level `skip_if:` in `stage.yml` is now read as its built-in default when the workflow
33
+ names none (`stageSpec.ts`'s `overlay`, a workflow's own `skip_if` for that stage id still wins);
34
+ the shipped `stages/how/stage.yml` sets `skip_if: "seed_solution==1"` this way. The skip is
35
+ recorded exactly like any other `skip_if` (`stage.skipped`, named on the run), and — because this
36
+ is the one skip that would otherwise leave `plan` reading a `02-how/design.md` that silently
37
+ looks empty — `tldrx next` additionally materialises that file from the seed's declared section
38
+ verbatim, cited back to the seed, the moment the skip fires.
39
+
40
+ - **A failed sub-agent turn now carries a named `failure_kind`** — `timeout`, `rate_limit`
41
+ (reusing #298's own quota-frame detection), `process_killed`, `empty_result`, `non_zero_exit`,
42
+ `result_error`, `malformed_result`, or `unclassified` (with the raw signal still in `error`) —
43
+ on the task row and the `agent.result` event, additive and absent on every ordinary/successful
44
+ turn and on a read-cap kill (`stopped_by` already names that one). `result_error` is
45
+ `non_zero_exit`'s exitCode-0 sibling (pre-merge review, 2026-09-15): the result document itself
46
+ reports `is_error: true` with a named reason while the process exited `0`, which used to read
47
+ `non_zero_exit` — a label contradicting its own exit code. Measured, one field audit (workspace
48
+ B, 2026-09-15): 24 of 54 failed tasks in that sample were an undifferentiated "timeout / generic
49
+ error" — the largest single bucket, unnamed until now. Classification only: no new retry
50
+ budgets, and nothing decides a retry off this field (#348, the family of #341/#298).
51
+
52
+ ### Changed
53
+
54
+ - **`watch` now follows `build`'s gate policy in every shipped scope where `build` is already
55
+ `auto` — `feature`, `bugfix`, `integration`, `refactor` and `performance` (#349, owner decision
56
+ 2026-09-16).** Every one of these scopes shipped `build: auto` with `watch: human` regardless,
57
+ so a run that trusted Build's own auto-gate still stopped a person at the very next stage over
58
+ a transcription job with a validator behind it. The reasoning is field evidence, not a guess:
59
+ auto-gating `what`/`how`/`plan` in a real workspace since 2026-09-12 produced zero measured
60
+ auto-gate incidents across 26 runs (77 `gate.rejected`, 100% human, 0 auto-rejections). The
61
+ mechanism is unchanged — `evaluateAutoGate` (`src/core/run/autoGate.ts`) already closes any
62
+ stage's gate the same way once its policy is `auto`, and a self-signed `watch` gate is
63
+ announced through the exact same `gate.requested`/`gate.approved`/`stage.done` events an
64
+ auto-signed `what`/`how`/`plan` gate already is — this does not widen the notify-window gap
65
+ tracked separately in #250, it only extends the existing behavior to one more stage. `upgrade`
66
+ is the one shipped scope excluded: `what`/`plan`/`build` are already all `auto` there, and
67
+ following `build` would leave it with no human gate at all, which spec.md's "no all-auto
68
+ preset" guarantee forbids — its `watch` stays `human`. `hotfix`/`security-patch`/`migration`
69
+ are unaffected: their `build` gate is `human`, so the "build already signed" condition never
70
+ applies. The tutorial's attended chapter (`tldrx learn`, chapter 7) ran its `next --commit` on
71
+ a `feature`-scope run expecting the Watch gate to still stop it (`expectExit: [4]`) — now that
72
+ `watch` self-signs there, the chapter sets it back to `human` explicitly first
73
+ (`run gates set watch:human --note …`), because the lesson is the attended human handoff, not
74
+ the shipped default.
75
+
76
+ - **`how`'s default model is `sonnet`, not `opus`** (#346). Owner decision, measured across a
77
+ single-run analysis and two workspace audits: `how` cost $10.40-$269.90 per sample, 11.4-45% of
78
+ the run's total spend. A workspace or a single invocation that wants the strong model back
79
+ overrides it exactly as before — `.tldrx/stages/how/stage.yml` with `model: opus`, or
80
+ `tldrx next --model opus` / `tldrx run auto --model opus` for one invocation.
81
+
82
+ - **`how` stops generating `risks.md`** (#350). Owner decision 2026-09-16 — "stop generating,"
83
+ not "generate and drop from the bundle" — because nothing downstream ever read it: measured
84
+ on a real run, `risks.md` (8,025 bytes) was excluded by `stages/plan/stage.yml`'s own
85
+ `inputs:` (only `design.md`, `contracts.md`, `test-strategy.md`) and a repo-wide grep found no
86
+ other reader. `handoff.md` and `questions.md` are unchanged — they carry no bundle content
87
+ either, but they are load-bearing elsewhere (the auto gate, `claim-sources`, `boundary`'s
88
+ `SURFACE_HANDOFFS`, the dashboard). `validateOutputs` only ever checks the DECLARED output
89
+ list, so a run directory written before this change that still has `02-how/risks.md` on disk
90
+ is unaffected — the file is simply not asked for anymore, never flagged as missing or
91
+ re-demanded on resume.
92
+
93
+ ### Fixed
94
+
95
+ - **A headless turn whose parent `tldrx` process is SIGKILL-class killed mid-spawn now banks its
96
+ cost as unmetered rather than nowhere at all (#337, the remaining half of #246).** #246's field
97
+ case: a `run auto` process died while blocked in `spawnAgent`'s `await` (`runNext.ts`); the
98
+ orphaned `claude` child kept running to completion, but `run.yml`'s `budget.spent_usd` stayed
99
+ `0.00` forever, because `demoteStaleRunning` — the code that finds the dead `.lock` and puts the
100
+ stage back to `ready` — wrote no task row for that attempt at all, not even an unmetered one. A
101
+ confident `0.00` for money a provider may really have charged is the one answer AGENTS.md §7
102
+ rules out. `demoteStaleRunning` now reads `startedHeadless` (`run/lastStart.ts`, the one
103
+ derivation of which mode opened a stage's last turn — already used by the Ctrl-C path) and, for
104
+ a genuinely orphaned HEADLESS spawn (never a `--prepare` bundle waiting for a human, and never a
105
+ Build executor phase, both of which are `preparedRefusal`'s and the executor's own business),
106
+ records a `status: "failed"` task with `cost_usd: null, metered: false` and
107
+ `error: "turn outlived its parent; no result line was read"` before demoting the stage — so
108
+ `spendBasis`/`tallyOf` (the one derivation, unchanged) count it as a LOWER BOUND, and `run
109
+ status`/`budget show`/the dashboard say so instead of printing a silent `$0.00`.
110
+
111
+ - **`run auto` now names `--wait-gates` as the fix, at the moment an all-`auto` gate's only
112
+ blocking questions are answered but nothing will sign it (#342).** Measured live on a
113
+ `--gates none --questions none` run launched with no `--wait-gates`/`--wait-answers`: the
114
+ loop answers a blocking question itself (`questions_policy: recommended`) or once
115
+ `--wait-answers` polls one in, but the ONLY thing that ever self-closes a parked `auto`
116
+ gate is `selfCloseAutoGate`, reached only from `waitForGate`'s poll — so the next `next`
117
+ call reports the stage as `awaiting_gate` with a generic `gate pending: tldrx approve`,
118
+ never saying the gate would have closed itself with one more flag. `settleDeferredGate`
119
+ (`runAuto.ts`) now says so, by name, at the exact moment the gap is decided. The deeper fix
120
+ — `next` re-evaluating a parked `auto` gate on its own, so this holds even for a `next`
121
+ launched by hand — is recommendation (1) on the issue and stays open there; it touches
122
+ `runNext.ts`, out of scope for this change.
123
+
124
+ - **The base-tree DoD refusal now names the workspace's own declared `install:` as the fix
125
+ (#343).** Measured on a fresh clone where a multi-step installer had only partly run: the
126
+ refusal reported `eslint`/`tsc`/`vitest` all `command not found`, identical in shape to a
127
+ genuinely broken repo, even though the same `WorkspaceContext` the refusal already reads
128
+ carries the workspace's own `install:` line. `baseRefusalLines` (`preflight.ts`) now names
129
+ it, once per repo — never runs it: running an installer against the operator's own
130
+ checkout, unasked, ahead of a refusal that has not yet told them anything, is exactly the
131
+ side effect the base-measurement design already avoids for gate commands. The zero-touch
132
+ first-run guide (EN/ES) now says explicitly to run the declared `install:` in the checkout
133
+ itself before the first `run auto`, not only in a story's worktree.
134
+
135
+ - **`run status` now warns NO-RETRY on a phase that cannot afford a second attempt of its own
136
+ stage, not only `budget show` (#232).** The narrowed #232 asked for this at both screens;
137
+ `budget show` shipped it in `bd19763`/`efb4eee` and this closes the other half. `run status`
138
+ computed no retry verdict at all, so a pre-#170 run (sized for exactly one attempt) read every
139
+ phase as an ordinary progress bar right up to the refusal — the exact "operator discovers it at
140
+ the retry" failure the issue is about, just on the screen that gets opened more often. Shares
141
+ the one `retrySizing` derivation and the one `NO-RETRY` sentence (`noRetryBlock`, now exported
142
+ from `budgetView.ts`) that `budget show` already used, so the two screens cannot disagree about
143
+ the same measured shortfall. `buildStatus` takes an optional `root` (both CLI and `run auto`'s
144
+ heartbeat now pass it) to resolve the next stage's own `attempts:` the same tolerant way
145
+ `budget show` does; omitted, it falls back to the shipped default, so a caller that does not pass
146
+ it behaves exactly as before.
147
+
148
+ - **A fix list dispositioned away from `fix-now` by hand, after the story already parked `review`
149
+ behind it, now settles `done` with no agent spawned instead of buying a fresh developer round and
150
+ a full DoD re-run for nothing (#218).** Measured live: a reviewer signed `fixlist` with an open
151
+ finding, parking the story at `review`; the finding was later re-routed away from `fix-now`
152
+ outside the audited auto-close (which only ever closes what a spawned reviewer's approve was
153
+ shown); nothing else re-read the file, so the next `--prepare`/headless pass read `fixlistFor`'s
154
+ unnamed door returning `null` (0 open — correctly, it is not to be re-rendered) as "no fix list
155
+ ever happened" and dispatched a plain developer bundle for a story nothing faulted — one
156
+ developer task and one full DoD re-run over an unchanged tree. `closedFixlistCandidates` /
157
+ `settleClosedFixlists` (gh #329's pre-pass) now also recognizes a `review` story parked on a spent
158
+ `fixlist` verdict, alongside the `blocked`/`approve` shape it already settled, and asks
159
+ `openFixNow` before either. `--fixlist <path>` naming a file with 0 open findings is now a refusal
160
+ that names the file and the count, not a silent downgrade to a bundle.
161
+
162
+ - **A fix list the parser could not fully read — a typo'd `Disposition:` word, or a round file
163
+ emptied by a crash or a disk-full write — can no longer settle a story `done` (#218, pre-merge
164
+ review).** Reviewer-measured: `**fix-now**` mistyped `**fix-noww**` is silently dropped by
165
+ `parseFixlistFile` (by design — a heading with no readable disposition is not a finding, and
166
+ inventing `fix-now` for it would block a story over a typo), so `openFindings(...).length` read 0
167
+ exactly as it would for a genuinely closed round, and a plain `--prepare` settled the story `done`
168
+ with no spawn and no re-review over a live, undetected defect. `parseFixlistFile`'s parse now also
169
+ reports what it could not read (`unreadableFindings`/`fixlistFullyParsed`, `fixlist.ts`), and
170
+ `openFixNow`/`closedFixlistCandidates` (both the `blocked`/`approve` shape and the `review`/
171
+ `fixlist` shape above share this one check) hold rather than settle when a round parsed
172
+ uncleanly, naming the exact heading and line that could not be read.
173
+
174
+ - **`--fixlist <path>` naming a file this framework cannot fully parse — corrupted, truncated, or
175
+ another story's misnamed file — refuses with exit 1 (usage), not exit 5 (agent failed) (#218,
176
+ pre-merge review).** The refusal used to throw out of `fixlistFor` and land in the generic
177
+ executor `catch`, which reads any thrown `Error` as a STAGE failure: `run.yml`'s stage was stamped
178
+ `status: failed` with `cost_usd: 0.00`, and the next `tldrx next` printed "retrying … (cost already
179
+ spent is not refunded)" — a lie, since the refusal fires before anything is opened or spent. It
180
+ now returns the same `refusedOnSequence` precondition-refusal `offerAtFrontier`'s dependency wait
181
+ already uses: exit 1, `run.yml` untouched, no false retry line. The unnamed door's own corruption
182
+ check is scoped to `--prepare` only (the one caller with a `try`/`catch` around it) — the headless
183
+ in-session fix-round spawn keeps the pre-#218 fallback (a corrupted round is treated as absent
184
+ rather than crashing the run mid-wave), since the correctness half above already closes the
185
+ dangerous direction independently of this door.
186
+
187
+ - **A relaunch of `tldrx run auto <runId>` resuming its OWN already-claimed epic branch now
188
+ verifies it instead of trusting it silently (#347).** `run auto` does not expose `--reuse-epic`
189
+ (there is nobody present to type it), so `foreignEpicRefusal`'s `claimed.has(branch)` case — this
190
+ run's own `build.epic_branch` already naming the branch a killed attempt cut and claimed before
191
+ it died — was a bare `continue`: no check that the branch head still resolved, no gate, and no
192
+ record that anything had been resumed at all. It now confirms the branch is still reachable and,
193
+ when the repo declares a `typecheck` command, runs it once in a throwaway detached worktree
194
+ (removed immediately after) before trusting the claim; a repo with no `typecheck` role is
195
+ adopted with that said explicitly (`typecheck: absent`) rather than assumed green. Recorded
196
+ honestly either way — an `epic.resumed` line and event carrying the sha and the typecheck
197
+ verdict — and a typecheck that FAILS refuses the stage by name instead of building on a possibly
198
+ broken branch (`tldrx next --reuse-epic` still adopts it deliberately, unchanged). A claim
199
+ another run wrote is refused exactly as before — nothing here relaxes that guard.
200
+
201
+ ## 0.30.0 — 2026-09-15
202
+
203
+ ### Fixed
204
+
205
+ - **The plan-over-stage advisory speaks only when the plan asks for more than 2× what the stage
206
+ holds, instead of on every overage (#302).** #281 gave a scaled plan a voice before a spawn —
207
+ right, and the reason a $16.20 stage over a $114.00 plan no longer has to be learned from a dead
208
+ developer. Unconditioned, that voice was loud: MEASURED by the #281 pre-merge reviewer over 30
209
+ priced run dirs on one machine, 8 of them (27%; 54% in one of the two workspaces) had
210
+ `priceScale < 1`, at severities from 0.095 to 0.85 — and at the mild end the plan asks ~18% more
211
+ than the stage and usually finishes without one story reaching its scaled cap. A line on every
212
+ Build entry and every Plan gate for a case that harmless is the wallpaper that teaches an
213
+ operator to skip the channel before the run where it matters, the same shape #285/#290 already
214
+ cost once. So the PROACTIVE channel gets a bar — `PLAN_OVER_STAGE_ADVISORY_SCALE`, strictly below
215
+ 0.5, in `caps.ts` beside the reasoning, one constant that both call sites inherit by calling the
216
+ same function — and the REACTIVE one keeps none: `capDeathReason` still names the formula, the
217
+ scale and the lever at every scale, because a story that actually died on its cap is never
218
+ wallpaper. The severe tail the channel exists for (0.095, >10× the stage) still speaks before the
219
+ spawn.
220
+
221
+ - **The provider's rate-limit warning is read instead of dropped, and a Build parks at a story
222
+ boundary before the wall instead of spending developers into it (#298, warning half).** The signal
223
+ was never missing, and the repo did not have to leave its own tree to find that out: `claude
224
+ --output-format stream-json` emits `{"type":"rate_limit_event","rate_limit_info":{…}}` WHILE a turn
225
+ is still working, `agentEvents.ts`'s own header has documented one since the transcript was
226
+ recorded, and `test/fixtures/agent/stream-json.jsonl:9` carries it in full — while
227
+ `test/agent-stream.test.ts` pinned it as noise, in the same list as `"{broken"`. It fell through
228
+ the dispatch `switch`'s `default:` and vanished, so the only trace a quota wall left was a
229
+ developer dying mid-story with the provider's own `success` on the record. Measured on a live
230
+ field turn (#298's thread): `status: "allowed_warning"` at `utilization: 0.92`, then `0.94`,
231
+ arriving BEFORE the limit bit, on a turn that then finished normally and cost $2.00 — and
232
+ `resetsAt` is an epoch integer, so a reader has a real deadline and never parses "resets 6pm". The
233
+ frame is now a typed `rate-limit` event, `AgentOutcome` carries the last one a turn streamed, and
234
+ a Build stage that sees a status the provider itself does not call `allowed` starts NO further
235
+ story: the stories already running finish and settle, the next one is not started, and the run
236
+ says so with the provider's own words on stdout and an `agent.rate_limited` event carrying the
237
+ status, the window, the utilization and the reset instant — written once per stage whether or not
238
+ a story was left to withhold, because a one-story wave that warns is a warning the operator still
239
+ acts on, and carried as the REASON on every story the handoff lists as not started, where the
240
+ audit record would otherwise say this stage had no reason for it. The park test is the provider's WORD,
241
+ never a utilization threshold this repo picked — the CLI already owns the judgement of when a
242
+ window has surpassed its threshold, and a second opinion here would be a second implementation of
243
+ it. Nothing waits and nothing retries: waiting to a reset the provider stated, and classifying the
244
+ DEATH itself, need a terminal capture nobody has obtained (measured negative, twice) and are filed
245
+ as the other half. Absent stays absent — a frame that states no utilization records
246
+ `utilization_absent: "not recorded — …"` rather than a `0` that would read as an empty window, and
247
+ a Codex turn carries no frame at all rather than a fabricated `allowed`, because nothing here has
248
+ measured what `codex exec --json` says about a quota.
249
+
250
+ - **`tldrx ship` refuses an epic carrying a merged diff NOBODY judged, instead of opening a PR over
251
+ it (#311).** #282 closed the case where a reviewer read a story's diff and said `changes` — a
252
+ rejection that stands, with the rejected code on the epic branch. It left the sibling open: three
253
+ settle paths park a story at `review` with `merged: true` beside a verdict that is not an opinion
254
+ about anything. The reviewer was refused for want of money and never spawned (#289) or the run was
255
+ cancelled before the review (#305), both recorded `n-a`; or the reviewer died mid-read, recorded
256
+ `error`. The run's own ledger already said what that class means — "`n-a` and `error` mean nothing
257
+ did [judge it]" — and `ship` asked it nothing: with one story `done` beside the parked one, #210's
258
+ refusal does not fire, #282's predicate does not match, and the PR opens with an unread diff in it
259
+ while the body lists that story under `## Not done` without one word that its code is in the diff.
260
+ Under `--ship merge` that body is read by nobody before auto-merge arms. It is now a refusal of its
261
+ own (exit 2, the family #282 uses: what is missing is a gate, and one of the three causes is
262
+ literally the money refusal), reading the SAME predicate the as-is path reads — `reviewNeverCompleted`,
263
+ not a second list of verdicts to fall out of step. It is a separate line rather than a widening of
264
+ #282's, because the REMEDY differs and a refusal that names the wrong verb sends the operator to the
265
+ wrong place: a rejection is a story to reopen and build again, while an unjudged merge is a review
266
+ still owed, and `tldrx next` settles a story parked at `review` by re-running the review with no
267
+ developer spawned. The line says which of the two non-verdicts it was, never just "unjudged": never
268
+ spawned and died mid-read are different problems with different next steps.
269
+
270
+ - **`tldrx init` probes a repo's four commands at once, and says which one it is waiting on
271
+ (#180).** MEASURED on `6fd2af2` with four 500 ms probes: they started at 0, 501, 1003 and
272
+ 1504 ms and `probeCommands` returned after 2006 ms — the SUM of the four deadlines, not the
273
+ longest. The same shape at the real `PROBE_TIMEOUT_MS` is 4 × 120 s = eight minutes for ONE
274
+ repo whose toolchain hangs, multiplied by the repo count, and for all of it detection had
275
+ nothing to say: `repoStart` fires once per repo and the next line is the repo's summary.
276
+ Nothing about a probe ever wanted the one before it — each has its own argv, races its OWN
277
+ deadline and writes its OWN key — so the four now run together, which puts the same fixture at
278
+ 502 ms (4.0×) and bounds one repo at one deadline. Bounding the wait was only half of it: two
279
+ minutes of a live view with nothing on it still reads as hung, so a probe that starts a process
280
+ announces itself and its outcome, `init` names the repo and the slots still in flight, and the
281
+ plain-mode heartbeat — the CI-log line, where there is no spinner to prove anything is alive —
282
+ carries that detail instead of only a clock. The rows are still assembled in slot order, not in
283
+ the order the probes happened to finish: `workspace.yml` is a file people diff, and a key order
284
+ that depended on which build was slower today would be a diff that means nothing. A slot decided
285
+ without starting anything — `run`, a command needing a shell, an absent one, a `--no-probe` skip
286
+ — announces nothing, because it costs no wait to report.
287
+ - **A Plan refused for a formatting slip gets ONE bounded repair turn instead of a full re-plan
288
+ (#288).** MEASURED on a live unattended run (0.18.3, `run auto --until-done`): the planner wrote
289
+ `acceptance: [ … ], test_plan: [ … ]` on one line — a flow-mapping comma inside a block mapping —
290
+ in three of three stories. The `plan` check refused it correctly, naming the file, the line and
291
+ the column; the stage then failed, because on a failed stage `tldrx next` IS a fresh stage. A
292
+ $2.19 Plan turn and one of the five `--until-done` relaunches were spent re-deriving a plan whose
293
+ only defect the checker had already localised, and the second planner passed only because it
294
+ happened to put one key per line. Build has always had the shape this needed — a verdict, then a
295
+ second turn on the same branch — and Plan had nothing between "refused" and "plan again". Now a
296
+ `plan` refusal whose every issue names a file that EXISTS re-spawns the planner ONCE over those
297
+ files, with the refusal VERBATIM and the same generated `## Output schemas` contract, and re-runs
298
+ the checks; a second refusal fails the stage exactly as before, in the same exit family, naming
299
+ both. Every bound is deliberate: the class is the CHECK's own answer computed from its issues and
300
+ not from the wording of its message (a refusal with nothing on disk to edit — "the Plan wrote no
301
+ stories" — earns no round, because a repair turn there is a re-plan wearing a fix round's
302
+ clothes); one round per planner turn, read off the `role: plan-fix` row rather than a counter
303
+ that could drift; headless only, since a `--prepare`/`--commit` cycle is a person for whom the
304
+ refusal is already the fix list, and it says so rather than doing nothing silently; and the money
305
+ goes through `wouldExceed`, the same predicate the budget gate decides on, at the same `agentCap`
306
+ the planner's own turn got — a phase that cannot fund the repair is told so and nothing is spent.
307
+ The turn is recorded like any other (a task row, an `agent.result`), with an additive
308
+ `plan.fix_round` event beside them saying why it happened and naming the row that holds the
309
+ dollars — one figure in one place. The Plan prompt also now states the rule the field run broke:
310
+ one key per line, never two joined by a comma.
311
+
312
+ - **Five more stack-pack reviewer Checks asked the read-only reviewer to do something it holds no
313
+ tool for, and one of them had no answer waiting for it anywhere (#195, #182's sibling).** #182
314
+ moved the can-it-fail Check to the developer because the mutation it asked for needs a write and
315
+ a test run, and the reviewer's whole allowance is `Read`, `Grep`, `Glob`, `Bash(git diff *)` — but
316
+ four more Checks, one per language pack, carried the identical contradiction: "run the typecheck
317
+ and lint commands declared in `.tldrx/workspace.yml`" sits sixty-odd lines above the SAME rendered
318
+ prompt's "Do not re-run them. They passed; that is why you are being asked," which follows the
319
+ `## Definition of Done — already re-run by the facilitator` section that already hands the
320
+ reviewer every declared command's exit code. Unlike #182's mutation, this one is not a producer's
321
+ obligation moved to whoever can perform it — the answer was already ON THE PAGE, so the four
322
+ Checks now point the reviewer at the Definition of Done section instead of asking it to run
323
+ anything a second time. The `postgres-testcontainers` overlay's ordering Check ("run the new tests
324
+ alone, then again … and compare") is a different case: nothing re-runs it at the Definition of
325
+ Done, so owner decision (Slack) gives it #182's shape instead — the overlay's own Defaults section
326
+ now asks the developer to run each new database test alone and with its neighbours and record the
327
+ result beside the test, and the Check asks the reviewer only whether that record is there.
328
+
329
+ - **A stage death no longer recommends the retry the budget is about to refuse (#232, remaining
330
+ half).** Every failed stage ended with the same literal — *"cost is recorded, not refunded —
331
+ retry with `tldrx next`"* — printed whether or not the phase could fund that retry. On a phase
332
+ sized to hold exactly one attempt of its stage (what `run new` wrote before #170, and what every
333
+ run created before 2026-09-09 keeps for life) `tldrx next` comes straight back as exit 2, so the
334
+ advice sent the operator to buy the refusal themselves: measured twice in one evening on the run
335
+ #232 was filed from, and once more on a second workspace where the starved phase was the last
336
+ one. The other half of #232 shipped in `efb4eee` — `budget show`'s `next` column says `NO-RETRY`
337
+ before a cent is spent — and this is the same fact said where the operator actually is when it
338
+ bites: the line now names the refusal, both figures, the shortfall and the `budget raise` that
339
+ clears it. It predicts the GATE and not the budget's health, using the gate's OWN inputs — what
340
+ the phase has left, the same remaining-work estimate the brake compares it against, the same
341
+ `planRebalance`, and this invocation's own `--rebalance-finished` state — deliberately: #232's
342
+ own comments measured how badly a derived estimate reads on a partly unmetered run, and a second
343
+ opinion computed here would disagree with the refusal the operator then hits. **That last term is
344
+ the one a first version of this fix got wrong, and a reviewer caught it before it merged**: the
345
+ rebalance is ON by default under `run auto` (#330), so a shortfall that finished phases can cover
346
+ is moved out of them before anything is refused, and the advice was telling the operator to raise
347
+ a ceiling on a run whose very next relaunch would have funded the retry by itself. `tldrx next`
348
+ alone never rebalances, so the same starved phase is refused through one door and funded through
349
+ the other; the advice now names which door it is standing in. Under a rebalancing launch whose
350
+ donors cover the shortfall it says the retry funds itself, names the donor phase and the move,
351
+ and says that a bare `tldrx next` is not that door — it stops short of a promise, because the
352
+ grant check on each move happens after this line is written. When no finished phase can cover it,
353
+ the refusal prediction stands through either door, with how short it still is *with all of it*.
354
+ Where this gate does not decide the retry at all — `on_exceed: warn` refuses nothing,
355
+ a `host-tokens` phase is a category error the dollar brake must never judge, an `attended_by:
356
+ host` run is allowed past it by policy — the plain line is printed unchanged: those are "not this
357
+ gate's call", which is not the same as "affordable", and claiming a refusal there would be the
358
+ invented value the house rules forbid. The failure's `signature` (#297) is untouched and is still
359
+ the failure's own sentence, never this line.
360
+
361
+ - **A Build refusal built from a CACHED base reading says so, and a relaunch re-measures the red
362
+ instead of being served it (#339).** MEASURED on a live `run auto --until-done`, 2026-09-15: two
363
+ Definition-of-Done commands exited 127 on the untouched base tree, the framework refused, the
364
+ operator installed the two missing binaries and confirmed all four commands green by hand — and
365
+ the next attempt printed the SAME refusal, byte for byte, over a base tree that was by then
366
+ repaired. Then the supervisor stopped the loop on it: "a refusal that repeats verbatim is not one
367
+ a relaunch moves", with four relaunches unspent and the run ended early on a workspace where the
368
+ thing complained about was fixed. Two halves compose into that. A cached refusal reproduces
369
+ itself exactly — which is correct for a cache and is the one condition the repeat guard reads —
370
+ and nothing anywhere said the second reading had not been taken. So: a refusal whose evidence was
371
+ re-used now names it as re-used, says when it was measured, against which base sha, and when it
372
+ stops being trusted, and it adds the sentence the generic advice cannot carry — that an edit to
373
+ `.tldrx/workspace.yml` or a base that moves clears a cached reading, but a repair the cache
374
+ cannot see (a binary installed, a service started) is not re-measured until that expiry. The
375
+ supervisor's repeat guard now reads a FRESHNESS field the producer sets rather than diffing text:
376
+ a repeat whose evidence was measured stops the loop exactly as it did, a repeat that was re-used
377
+ is relaunched and says why. And the relaunch is no longer served the same cache — an attempt that
378
+ follows a refusal re-probes a RED base unconditionally, the way `--prepare` already did, because
379
+ the attempt before it asked for exactly this repair. Greens are still re-used, so the cost of
380
+ that is bounded by the commands that are actually broken. The 30-minute red TTL is untouched;
381
+ it was never what saved this run, and the reason it was not is filed separately (#340).
382
+
383
+ - **A red Definition-of-Done command is measured TWICE before it pins a story `blocked`, and the
384
+ record carries both exit codes (#163, sub-fix 1).** MEASURED on a .NET workspace, 2026-09-05: a
385
+ story's DoD gate returned `dotnet test` → exit 2, so the story blocked; the operator then ran the
386
+ identical suite twice — in the story worktree and on the epic branch — and got exit 0 with 2860
387
+ tests and 0 failing. The red was contention over a container runtime, and it had consumed the
388
+ story's last attempt. `blocked` is terminal in-run, so a human had to reopen the story by hand.
389
+ The hole was not the terminality, which is a defensible rule: it was that a story's red was never
390
+ re-measured, so a flake and a defect were written down identically and nobody reading the block
391
+ could tell which one they had. The framework already did this one level up — the base-tree
392
+ pre-flight re-measures a command the cache cannot answer — and the story's own red had no
393
+ equivalent. It does now: when a command goes red on a story, the identical command runs once more
394
+ in the same tree, and the `check.failed` event, the story's `dod` row and the blocked reason all
395
+ carry what the second run said. **Nothing about terminality changes** — a red that reproduces
396
+ blocks exactly as before, and a red that does NOT reproduce also still blocks; it is now blocked
397
+ with a sentence that says the command passed the second time, which is the difference between
398
+ reopening the story and going looking for a bug that is not there. When no second run could be
399
+ taken the record says so with the reason and no number (§7): a command whose binary the tree
400
+ never had would only measure the same absence again, and a refused command never ran, so it has
401
+ no first reading to reproduce. **A red that TIMED OUT is re-run too, so a timing-out
402
+ Definition-of-Done command now costs up to two timeout periods instead of one.** That is stated
403
+ rather than buried: the owner approved one extra run of the suite, and doubling a timeout is a
404
+ fair reading of it but was not in the framing. Timeouts are deliberately not carved out — the
405
+ measured case was contention over a container runtime, and contention is exactly what makes a
406
+ suite hang rather than fail cleanly, so a carve-out would leave the likeliest flake the one thing
407
+ never measured. The blocked reason names both readings and marks either of them `(timed out)`,
408
+ so the record says which kind of red was paid for twice. Owner decision, answered on Slack:
409
+ re-run the red command once and record both exit codes. One reading per spawn, shared with the base-tree path, so the two cannot
410
+ disagree about what a timeout's exit code is or which line of a red suite is the summary.
411
+
412
+ - **A fix-list sha that is too LONG is refused by name, instead of reading as no sha at all
413
+ (#163, sub-fix 3).** The abbreviation edge of `Resolved: yes <sha>` was closed in 0.10.0 — a
414
+ 7-39 hex claim that verifies is rewritten to the 40 `rev-parse` returned — and the same
415
+ grammar, `\b([0-9a-f]{7,40})\b`, was failing silently at the other edge, which is the
416
+ dangerous direction §7 names. No position inside a 41-character hex run is a word boundary, so
417
+ an over-long token produced NO match: the line read as a bare `Resolved: yes` that named
418
+ nothing, and the shape fault was never stated. Measured while fixing it, and worse than the
419
+ issue reported: the scan did not stop there but carried on to the next hex word on the line, so
420
+ `Resolved: yes <41 hex> (see 9f2c1ab)` closed the finding over `9f2c1ab` — a DIFFERENT commit
421
+ from the one the line claims, which is an audit record inventing its own evidence. A token of
422
+ 41+ hex characters is now refused by name and by count, through the `claimed-unverified`
423
+ sentence #130 already writes into the file, so the finding keeps holding the story and the
424
+ record says why; a refusal anywhere on the line refuses the whole read, so the answer cannot
425
+ depend on which side of the bad token a good one happens to sit. 7-39 is still accepted and
426
+ still canonicalised — nothing a person legitimately types is refused. The grammar now lives in
427
+ one leaf, `readResolvedSha`, read by both the parser and the rewriter, which is what makes the
428
+ token that gets verified and the token that gets replaced provably the same span. The refusal is
429
+ written once and read everywhere: `verifyResolutions` is the single site that withdraws a claim,
430
+ so the file, the story's blocked reason and the stdout report all carry the SAME sentence. They
431
+ did not before — the file would have named the 41-character token while the report a human reads
432
+ first said `named no commit to point at`, which is two records of one event disagreeing.
433
+
434
+ - **`budget show` says a phase cannot afford its own retry BEFORE the money is spent, instead of
435
+ labelling it `ok` up to the refusal (closes #232).** A phase whose ceiling equals ONE attempt of
436
+ its stage refuses every retry after the first cent — by arithmetic, not by policy — and the
437
+ framework's own remedy for a failed stage ("cost is recorded, not refunded — retry with `tldrx
438
+ next`") is exactly what it then refuses. Measured twice in one run, and on a second workspace in
439
+ the LAST phase, where every dollar has already gone. `run new` has sized phases at `attempts ×`
440
+ their stage since #170 so it can no longer be created, and #330's rebalance reaches it only when
441
+ some FINISHED phase has slack to give — with none, the run stops at exit 2, the one exit nothing
442
+ unattended can leave. The `next` column now answers the SIZE question in its own word: `NO-RETRY`
443
+ with the exact raise, and `n/e` for a phase there is nothing to size in. The verdict reads the
444
+ stage's DECLARED `budget_usd` and its own `attempts:`, never the derived estimate beside it —
445
+ #232's own comments measured why: on a partly unmetered run the estimate is too small, so
446
+ `ceiling / est.` reports headroom that is not there, and the less a run is metered the safer that
447
+ ratio claims it is. A declaration does not move with metering. `attempts: 1` declares no retry, so
448
+ one attempt's worth is the right size for it and still reads `ok`. Nothing about the refusal moved:
449
+ `blocked`, `remaining` and `est.` are untouched and `tldrx next` still runs. One derivation
450
+ (`retrySizing`, `budget/wouldExceed.ts`), read back additively by `budget show --json`.
451
+
452
+ - **A killed headless turn reads as `interrupted` and is handed a command the next line of code
453
+ accepts, instead of being called a `--prepare` bundle that never existed (closes #246, half
454
+ one).** MEASURED on a 0.16.1 field run: the `tldrx` process driving `run auto` died
455
+ SIGKILL-class (no handler ran), leaving the stage `running`, a dead pid in the `.lock`, and
456
+ `.agent/<stage>/pending.json` on disk. `run status` said "a `--prepare` bundle is waiting — run
457
+ the prompt and `tldrx next --commit <id>`"; that command was then REFUSED ("what is `ready`, not
458
+ `running`"), so the only advice the screen gave was a dead end, and nobody was going to run a
459
+ prompt by hand on an unattended run anyway. The mechanism is that the two states are
460
+ indistinguishable on disk: `runNext` writes the bundle BEFORE it branches on the mode, so every
461
+ headless spawn leaves exactly what a `--prepare` leaves, and `waiting.ts` derived `prepared`
462
+ from stage `running` + dead lock + a bundle without ever reading which mode opened the turn. The
463
+ ledger already knew — `stage.started` has carried `mode` since it was first written — so the
464
+ answer is now read off the newest `stage.started` for that stage
465
+ (`src/core/run/lastStart.ts`, one implementation, imported by both readers that had made this
466
+ identification independently). A tenth waiting kind, `interrupted`, says what happened and
467
+ prescribes `tldrx run auto <id>` — the command that actually resumes the run; `prepared` is
468
+ reserved for a bundle a `--prepare` really wrote. `Ctrl-C` shared the bug and is fixed with it:
469
+ `stopInFlightRun` preserved any `killed === 0` stage with a bundle as `running` forever, and now
470
+ preserves only one whose last start was a `prepare`. A ledger that does not say keeps the old
471
+ reading exactly — an absent record is not evidence of a spawn. The second half of #246, the
472
+ orphaned sub-agent's dollars being recorded nowhere while `spent_usd` reads a confident `0.00`,
473
+ is untouched and stays open.
474
+
475
+ - **A Watch retry no longer re-buys a watcher card that already validated (closes #306).** The
476
+ Watch stage fails on the FIRST card that does not validate, and `tldrx next` / `run auto
477
+ --retry-failed` re-enter the executor from the top — so a run with N features where feature k
478
+ was refused spawned all N writers again, including the k-1 cards that had already validated and
479
+ been stamped. Each of those paid at least the `$0.25` cold-spawn floor for a card nothing had
480
+ refused, and on a two-feature field run (#301, run B) with a different card refused in each
481
+ attempt, every writer was bought at least twice. #301 made the second spend an edit rather than
482
+ a rewrite; this makes it zero. The headless path now takes one snapshot of disk as the stage is
483
+ entered, before anything spawns, and keeps a feature whose card passes the SAME
484
+ `parseWatcherCard` against the SAME source context the stage validates with — so the skip can
485
+ never hide a bad card, only decline to save money, and a card whose citations went stale in the
486
+ meantime is rewritten like any other. It is a snapshot rather than a per-feature read so that
487
+ one writer loose in the run directory can never decide whether the next feature is written at
488
+ all. A kept feature still gets its `tasks[]` row, with `cost_usd: 0.0`, `model: null` and
489
+ `session_id: null` — a measured zero, not the `metered: false` of a turn billed to a host
490
+ session — and the stage's report names every kept card, so a `$0.00` row is never read as a
491
+ writer that silently did nothing. `--prepare`/`--commit` is untouched.
492
+
493
+ - **The epic branch Build cuts is claimed in `run.yml` at the cut, not when the stage returns
494
+ (closes #262).** Reported from a field workspace: a `run auto` loop was killed mid-Build,
495
+ relaunched, and was refused its OWN epic — "`epic/<slug>` already exists … refusing to stack
496
+ this run's commits onto someone else's epic" — with `tldrx next --reuse-epic` the only way back
497
+ in, a flag `run auto` cannot pass. Same sentence as #248, different mechanism, and #248's fix
498
+ cannot reach this one: #248 closed the THROW path, where a process is still alive to run its
499
+ catch. Here the process is GONE. `ExecutorOutcome.epicBranches` merges into `build.epic_branch`
500
+ only after the executor RETURNS, and a Build stage does not return for as long as a developer
501
+ turn takes — so a SIGKILL, an OOM kill or a cut session anywhere in that window left
502
+ `epic/<slug>` in the repo with nothing in the record claiming it. SIGINT and SIGTERM are hooked
503
+ (`src/cli/signals.ts`); SIGKILL cannot be hooked by any process, so no handler was ever going to
504
+ cover this. The cut now records the claim itself, through the `ctx.claimEpicBranch` seam, once
505
+ per branch rather than once per story. It does NOT ride the write serializer and is not claimed
506
+ to: of `openStory`'s eight call sites, four wrap it in `this.writes.run` and four
507
+ (`prepare`, `prepareReview`, `commitReview`, `commit`) call it bare. What makes it safe is that
508
+ there is no `await` between the branch being cut and the claim, and the claim is synchronous —
509
+ so on every one of those paths the cut and its record are one uninterrupted step. The guard is
510
+ NOT widened: a branch this run did not cut is refused exactly as
511
+ before — the claim stays a RECORD, and the fix is to write the record when it is earned rather
512
+ than to teach the guard to accept inferred evidence. A `run.yml` save that cannot take the
513
+ workspace lock is NOT swallowed: it fails the stage naming the branch, the repo, why the write
514
+ failed and both ways back in, because a claim that could not be written is this defect happening
515
+ again in silence. Honest about how far this goes: the save is atomic against a KILLED PROCESS
516
+ (temp file, then `rename`), and it is not `fsync`ed — a power cut can still lose a write the
517
+ kernel had not flushed. It is pinned by a real SIGKILL of a real `tldrx next`, never a simulated
518
+ throw, which would only have re-tested #248.
519
+
520
+ - **A Codex Build review no longer dies before the reviewer runs, and the provider's own refusal
521
+ arrives on one line (#148).** MEASURED on a live Build smoke run, published 0.7.0 +
522
+ codex-cli 0.153.0: the developer stage finished and 4 tests passed, then the review failed at
523
+ the API with `invalid_json_schema: fixlist.items.required must include every property (missing
524
+ n)`. Codex's structured-output API requires every DECLARED property to be listed as required;
525
+ `REVIEW_SCHEMA` declares optional fields at two levels — `fixlist` itself, and `n`, `severity`,
526
+ `where`, `detail`, `do_not` inside each item — and the spawn wrote that object to Codex's
527
+ `--output-schema` file verbatim. Every Codex review was refused before it began, so the story
528
+ paid an attempt for a turn the provider never ran. The translation happens at the Codex spawn
529
+ boundary ALONE (`codexSchema`, one implementation, §7): an optional property becomes
530
+ `anyOf: [<its schema>, {"type": "null"}]` and joins `required`, recursively, and the review
531
+ parser already reads a `null` exactly as it read the absent field — so nothing about
532
+ fail-closed moves, and a `fixlist` verdict with nothing readable in it is still `changes`.
533
+ `REVIEW_SCHEMA` itself is untouched and Claude's `--json-schema` still carries it byte for
534
+ byte; the test asserts both directions, because a fix at a provider seam that quietly edits the
535
+ shared contract is the failure this one was supposed to avoid. Second half, same issue: Codex
536
+ names its refusal in a pretty-printed block, and that sentence is quoted into a handoff where
537
+ every line must carry its own `[src: …]` — a multi-line reason became uncited continuation
538
+ lines that failed the claim-sources check. The Codex reason is now collapsed onto one line
539
+ rather than cut at the first, because the first line of that block is `{` and the WHY is two
540
+ lines down. Claude's failure text is passed through unchanged. The acceptance is MEASURED,
541
+ not argued from the error message: the pre-merge reviewer ran one real
542
+ `codex exec -s read-only --output-schema <the translated schema>` in an isolated directory
543
+ against an authenticated codex-cli 0.153.0 — the issue's own version — and it exited 0 with
544
+ no `invalid_json_schema` and the structured output `{"verdict": "approve", "summary":
545
+ "Approved with no findings.", "findings": [], "fixlist": null}`. That `"fixlist": null` is
546
+ the translation's whole point arriving back off the wire: the field the API refused to leave
547
+ optional comes home as the null this parser already reads as absent. Two readings, taken by
548
+ different hands — the schema bytes the child was handed, asserted in the test on every run,
549
+ and that one live call, which is not re-run by any gate.
550
+
551
+ - **A still-open fix-list finding is re-checked against the EPIC tip before the Build gate, so a
552
+ defect a LATER story closed says so with the closing sha instead of reading `Resolved: no`
553
+ (#163, sub-fix 2).** MEASURED on a Next.js workspace, 2026-09-05: at a Build gate all 6
554
+ `fix-now` findings carried a resolving sha and the 3 `defer-with-log` entries read
555
+ `Resolved: no` — "correct PER STORY", in the host agent's own words — but two of those three
556
+ defects had in fact been closed later in the same run by a DIFFERENT story, which the host
557
+ verified by reading the code. Exactly one finding had genuinely shipped unfixed, the record
558
+ could not tell those apart, and so a human narrated the difference by hand at the gate. The hole
559
+ was structural rather than careless: verification asks whether a claim's sha is reachable from
560
+ the STORY's own branch, and it asks at the moment that story settles — a fix that lands later,
561
+ on somebody else's branch, and reaches the epic through a merge did not exist when the question
562
+ was asked and is not on the branch it was asked about. There was no third answer for it, so it
563
+ was recorded as the second. There is one now. After the last story settles and before the
564
+ handoff is written, every finding not already closed with evidence is re-checked against its
565
+ epic tip, and a sha that is reachable from the epic and NOT from the story's own branch is
566
+ recorded as what it is. **The two kinds of close are spelled apart, never flattened**:
567
+ `Resolved: yes <sha>` stays the story's own close, and a close a later story landed is
568
+ `Resolved: yes-on-epic <sha>` — a different word because it is a different fact about who fixed
569
+ it and where the fix lives. `yes-on-epic` is not `yes`, so nothing the sweep writes changes what
570
+ any gate decides; the finding is exactly as open afterwards as it was before, and a finding that
571
+ genuinely shipped unfixed still reads unfixed. **Nothing is ever re-marked on inference.** A
572
+ reachable commit the record already names is the only thing that closes anything here — never a
573
+ changed file, never a heading that resembles another story's work — so the sweep cannot invent
574
+ the one direction §7 refuses to be wrong in. **And it does not go quiet when it finds nothing.**
575
+ Every finding it examined gets a `Swept:` line saying what was measured and against which tip,
576
+ because a `Resolved: no` that has been re-checked against the whole run and one that only ever
577
+ meant "correct per story" were indistinguishable, and that indistinguishability is the defect.
578
+ A sweep that could NOT be taken — no epic branch recorded, a tip git will not resolve, a file
579
+ that cannot be re-read — names its reason on every finding and in the handoff's `## Unknowns`,
580
+ rather than reporting a clean sweep over a measurement nobody took. **And it asks the weaker
581
+ question too, in words that say it is weaker.** The three entries that Next.js gate got wrong
582
+ named no commit at all — nobody had claimed one — so re-checking shas the record already names
583
+ has nothing to say about exactly the findings a human had to read the code for. For those, the
584
+ sweep asks what that human asked first: did a commit on the EPIC, and not on this story's own
585
+ branch, CHANGE the file this finding cites? Up to three such commits come back named, on the
586
+ `Swept:` line and as their own `## Unknowns` bullet — and they resolve NOTHING. A changed file
587
+ is not a closed defect, and writing `Resolved: yes` off one would be the same lie arrived at
588
+ from the other side; what this replaces is a person grepping the run's commits by hand at the
589
+ gate, never a person's judgement about them. **A probe git REFUSED says that, too.** An
590
+ unresolvable story ref and a file nobody changed both come back as no commits, and writing
591
+ "no later commit changed this" over the first would be a measurement nobody took wearing the
592
+ words of one that was — so the refusal carries git's own reason onto the `Swept:` line and
593
+ into `## Unknowns`, and is never counted as a clean probe. `Swept:` and the
594
+ `yes-on-epic` word are additive: a fix list written before this change reads exactly as it did.
595
+
596
+ ### Changed
597
+
598
+ - **The zero-touch launch recipe (EN and ES) now pairs `run auto` with `--wait-gates` and
599
+ `--wait-answers` everywhere it shows opening a run with `--gates none`, and says why: `--gates
600
+ none` sets the policy, but only `--wait-gates` lets `run auto` re-sign a parked `auto` gate once
601
+ the question that held it is auto-answered — without it the loop exits `awaiting human` on the
602
+ first gate that parks on a question, even though every one of its conditions already holds (see
603
+ #342).**
604
+
3
605
  ## 0.29.0 — 2026-09-15
4
606
 
5
607
  ### Changed