tldr-experts 0.28.1 → 0.30.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (32) hide show
  1. package/CHANGELOG.md +635 -0
  2. package/README.md +2 -0
  3. package/dist/hooks/answer-capture.js +8 -9
  4. package/dist/hooks/budget-gate.js +47 -11
  5. package/dist/hooks/{chunk-hws2gnxj.js → chunk-0d4csx4s.js} +1 -1
  6. package/dist/hooks/{chunk-abswp1vt.js → chunk-77x6v0gg.js} +4 -2
  7. package/dist/hooks/{chunk-zrwq6r7k.js → chunk-7xascvee.js} +1 -1
  8. package/dist/hooks/{chunk-zvxgdssb.js → chunk-857jhjmb.js} +1 -1
  9. package/dist/hooks/{chunk-c62y35wh.js → chunk-fbmj1bk6.js} +3 -3
  10. package/dist/hooks/{chunk-veffvzw5.js → chunk-g4d2v4jc.js} +101 -10
  11. package/dist/hooks/{chunk-48ctk154.js → chunk-gr83pq7x.js} +3 -3
  12. package/dist/hooks/{chunk-q3d4xdhb.js → chunk-mg3aj0sa.js} +1 -1
  13. package/dist/hooks/{chunk-tna1bz8r.js → chunk-pne7afnf.js} +1 -1
  14. package/dist/hooks/{chunk-fjeefnqr.js → chunk-przta0a0.js} +44 -8
  15. package/dist/hooks/{chunk-f6t1xc3f.js → chunk-qmp5p77c.js} +2 -2
  16. package/dist/hooks/{chunk-khy21x6e.js → chunk-yanmz0vz.js} +1 -1
  17. package/dist/hooks/{chunk-khjm356r.js → chunk-z9hv47ps.js} +1 -1
  18. package/dist/hooks/{chunk-e618k6te.js → chunk-zpch8e2j.js} +193 -5
  19. package/dist/hooks/claim-sources.js +5 -5
  20. package/dist/hooks/dod-gate.js +9 -10
  21. package/dist/hooks/no-reask.js +8 -8
  22. package/dist/hooks/session-start.js +16 -13
  23. package/dist/hooks/statusline.js +8 -9
  24. package/dist/tldrx.js +2597 -990
  25. package/package.json +1 -1
  26. package/plugin/.claude-plugin/plugin.json +1 -1
  27. package/templates/experts/stack/dotnet.md +4 -3
  28. package/templates/experts/stack/javascript.md +4 -3
  29. package/templates/experts/stack/overlays/postgres-testcontainers.md +7 -2
  30. package/templates/experts/stack/python.md +4 -3
  31. package/templates/experts/stack/typescript.md +4 -3
  32. package/dist/hooks/chunk-g7wz339n.js +0 -102
package/CHANGELOG.md CHANGED
@@ -1,5 +1,640 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.30.0 — 2026-09-15
4
+
5
+ ### Fixed
6
+
7
+ - **The plan-over-stage advisory speaks only when the plan asks for more than 2× what the stage
8
+ holds, instead of on every overage (#302).** #281 gave a scaled plan a voice before a spawn —
9
+ right, and the reason a $16.20 stage over a $114.00 plan no longer has to be learned from a dead
10
+ developer. Unconditioned, that voice was loud: MEASURED by the #281 pre-merge reviewer over 30
11
+ priced run dirs on one machine, 8 of them (27%; 54% in one of the two workspaces) had
12
+ `priceScale < 1`, at severities from 0.095 to 0.85 — and at the mild end the plan asks ~18% more
13
+ than the stage and usually finishes without one story reaching its scaled cap. A line on every
14
+ Build entry and every Plan gate for a case that harmless is the wallpaper that teaches an
15
+ operator to skip the channel before the run where it matters, the same shape #285/#290 already
16
+ cost once. So the PROACTIVE channel gets a bar — `PLAN_OVER_STAGE_ADVISORY_SCALE`, strictly below
17
+ 0.5, in `caps.ts` beside the reasoning, one constant that both call sites inherit by calling the
18
+ same function — and the REACTIVE one keeps none: `capDeathReason` still names the formula, the
19
+ scale and the lever at every scale, because a story that actually died on its cap is never
20
+ wallpaper. The severe tail the channel exists for (0.095, >10× the stage) still speaks before the
21
+ spawn.
22
+
23
+ - **The provider's rate-limit warning is read instead of dropped, and a Build parks at a story
24
+ boundary before the wall instead of spending developers into it (#298, warning half).** The signal
25
+ was never missing, and the repo did not have to leave its own tree to find that out: `claude
26
+ --output-format stream-json` emits `{"type":"rate_limit_event","rate_limit_info":{…}}` WHILE a turn
27
+ is still working, `agentEvents.ts`'s own header has documented one since the transcript was
28
+ recorded, and `test/fixtures/agent/stream-json.jsonl:9` carries it in full — while
29
+ `test/agent-stream.test.ts` pinned it as noise, in the same list as `"{broken"`. It fell through
30
+ the dispatch `switch`'s `default:` and vanished, so the only trace a quota wall left was a
31
+ developer dying mid-story with the provider's own `success` on the record. Measured on a live
32
+ field turn (#298's thread): `status: "allowed_warning"` at `utilization: 0.92`, then `0.94`,
33
+ arriving BEFORE the limit bit, on a turn that then finished normally and cost $2.00 — and
34
+ `resetsAt` is an epoch integer, so a reader has a real deadline and never parses "resets 6pm". The
35
+ frame is now a typed `rate-limit` event, `AgentOutcome` carries the last one a turn streamed, and
36
+ a Build stage that sees a status the provider itself does not call `allowed` starts NO further
37
+ story: the stories already running finish and settle, the next one is not started, and the run
38
+ says so with the provider's own words on stdout and an `agent.rate_limited` event carrying the
39
+ status, the window, the utilization and the reset instant — written once per stage whether or not
40
+ a story was left to withhold, because a one-story wave that warns is a warning the operator still
41
+ acts on, and carried as the REASON on every story the handoff lists as not started, where the
42
+ audit record would otherwise say this stage had no reason for it. The park test is the provider's WORD,
43
+ never a utilization threshold this repo picked — the CLI already owns the judgement of when a
44
+ window has surpassed its threshold, and a second opinion here would be a second implementation of
45
+ it. Nothing waits and nothing retries: waiting to a reset the provider stated, and classifying the
46
+ DEATH itself, need a terminal capture nobody has obtained (measured negative, twice) and are filed
47
+ as the other half. Absent stays absent — a frame that states no utilization records
48
+ `utilization_absent: "not recorded — …"` rather than a `0` that would read as an empty window, and
49
+ a Codex turn carries no frame at all rather than a fabricated `allowed`, because nothing here has
50
+ measured what `codex exec --json` says about a quota.
51
+
52
+ - **`tldrx ship` refuses an epic carrying a merged diff NOBODY judged, instead of opening a PR over
53
+ it (#311).** #282 closed the case where a reviewer read a story's diff and said `changes` — a
54
+ rejection that stands, with the rejected code on the epic branch. It left the sibling open: three
55
+ settle paths park a story at `review` with `merged: true` beside a verdict that is not an opinion
56
+ about anything. The reviewer was refused for want of money and never spawned (#289) or the run was
57
+ cancelled before the review (#305), both recorded `n-a`; or the reviewer died mid-read, recorded
58
+ `error`. The run's own ledger already said what that class means — "`n-a` and `error` mean nothing
59
+ did [judge it]" — and `ship` asked it nothing: with one story `done` beside the parked one, #210's
60
+ refusal does not fire, #282's predicate does not match, and the PR opens with an unread diff in it
61
+ while the body lists that story under `## Not done` without one word that its code is in the diff.
62
+ Under `--ship merge` that body is read by nobody before auto-merge arms. It is now a refusal of its
63
+ own (exit 2, the family #282 uses: what is missing is a gate, and one of the three causes is
64
+ literally the money refusal), reading the SAME predicate the as-is path reads — `reviewNeverCompleted`,
65
+ not a second list of verdicts to fall out of step. It is a separate line rather than a widening of
66
+ #282's, because the REMEDY differs and a refusal that names the wrong verb sends the operator to the
67
+ wrong place: a rejection is a story to reopen and build again, while an unjudged merge is a review
68
+ still owed, and `tldrx next` settles a story parked at `review` by re-running the review with no
69
+ developer spawned. The line says which of the two non-verdicts it was, never just "unjudged": never
70
+ spawned and died mid-read are different problems with different next steps.
71
+
72
+ - **`tldrx init` probes a repo's four commands at once, and says which one it is waiting on
73
+ (#180).** MEASURED on `6fd2af2` with four 500 ms probes: they started at 0, 501, 1003 and
74
+ 1504 ms and `probeCommands` returned after 2006 ms — the SUM of the four deadlines, not the
75
+ longest. The same shape at the real `PROBE_TIMEOUT_MS` is 4 × 120 s = eight minutes for ONE
76
+ repo whose toolchain hangs, multiplied by the repo count, and for all of it detection had
77
+ nothing to say: `repoStart` fires once per repo and the next line is the repo's summary.
78
+ Nothing about a probe ever wanted the one before it — each has its own argv, races its OWN
79
+ deadline and writes its OWN key — so the four now run together, which puts the same fixture at
80
+ 502 ms (4.0×) and bounds one repo at one deadline. Bounding the wait was only half of it: two
81
+ minutes of a live view with nothing on it still reads as hung, so a probe that starts a process
82
+ announces itself and its outcome, `init` names the repo and the slots still in flight, and the
83
+ plain-mode heartbeat — the CI-log line, where there is no spinner to prove anything is alive —
84
+ carries that detail instead of only a clock. The rows are still assembled in slot order, not in
85
+ the order the probes happened to finish: `workspace.yml` is a file people diff, and a key order
86
+ that depended on which build was slower today would be a diff that means nothing. A slot decided
87
+ without starting anything — `run`, a command needing a shell, an absent one, a `--no-probe` skip
88
+ — announces nothing, because it costs no wait to report.
89
+ - **A Plan refused for a formatting slip gets ONE bounded repair turn instead of a full re-plan
90
+ (#288).** MEASURED on a live unattended run (0.18.3, `run auto --until-done`): the planner wrote
91
+ `acceptance: [ … ], test_plan: [ … ]` on one line — a flow-mapping comma inside a block mapping —
92
+ in three of three stories. The `plan` check refused it correctly, naming the file, the line and
93
+ the column; the stage then failed, because on a failed stage `tldrx next` IS a fresh stage. A
94
+ $2.19 Plan turn and one of the five `--until-done` relaunches were spent re-deriving a plan whose
95
+ only defect the checker had already localised, and the second planner passed only because it
96
+ happened to put one key per line. Build has always had the shape this needed — a verdict, then a
97
+ second turn on the same branch — and Plan had nothing between "refused" and "plan again". Now a
98
+ `plan` refusal whose every issue names a file that EXISTS re-spawns the planner ONCE over those
99
+ files, with the refusal VERBATIM and the same generated `## Output schemas` contract, and re-runs
100
+ the checks; a second refusal fails the stage exactly as before, in the same exit family, naming
101
+ both. Every bound is deliberate: the class is the CHECK's own answer computed from its issues and
102
+ not from the wording of its message (a refusal with nothing on disk to edit — "the Plan wrote no
103
+ stories" — earns no round, because a repair turn there is a re-plan wearing a fix round's
104
+ clothes); one round per planner turn, read off the `role: plan-fix` row rather than a counter
105
+ that could drift; headless only, since a `--prepare`/`--commit` cycle is a person for whom the
106
+ refusal is already the fix list, and it says so rather than doing nothing silently; and the money
107
+ goes through `wouldExceed`, the same predicate the budget gate decides on, at the same `agentCap`
108
+ the planner's own turn got — a phase that cannot fund the repair is told so and nothing is spent.
109
+ The turn is recorded like any other (a task row, an `agent.result`), with an additive
110
+ `plan.fix_round` event beside them saying why it happened and naming the row that holds the
111
+ dollars — one figure in one place. The Plan prompt also now states the rule the field run broke:
112
+ one key per line, never two joined by a comma.
113
+
114
+ - **Five more stack-pack reviewer Checks asked the read-only reviewer to do something it holds no
115
+ tool for, and one of them had no answer waiting for it anywhere (#195, #182's sibling).** #182
116
+ moved the can-it-fail Check to the developer because the mutation it asked for needs a write and
117
+ a test run, and the reviewer's whole allowance is `Read`, `Grep`, `Glob`, `Bash(git diff *)` — but
118
+ four more Checks, one per language pack, carried the identical contradiction: "run the typecheck
119
+ and lint commands declared in `.tldrx/workspace.yml`" sits sixty-odd lines above the SAME rendered
120
+ prompt's "Do not re-run them. They passed; that is why you are being asked," which follows the
121
+ `## Definition of Done — already re-run by the facilitator` section that already hands the
122
+ reviewer every declared command's exit code. Unlike #182's mutation, this one is not a producer's
123
+ obligation moved to whoever can perform it — the answer was already ON THE PAGE, so the four
124
+ Checks now point the reviewer at the Definition of Done section instead of asking it to run
125
+ anything a second time. The `postgres-testcontainers` overlay's ordering Check ("run the new tests
126
+ alone, then again … and compare") is a different case: nothing re-runs it at the Definition of
127
+ Done, so owner decision (Slack) gives it #182's shape instead — the overlay's own Defaults section
128
+ now asks the developer to run each new database test alone and with its neighbours and record the
129
+ result beside the test, and the Check asks the reviewer only whether that record is there.
130
+
131
+ - **A stage death no longer recommends the retry the budget is about to refuse (#232, remaining
132
+ half).** Every failed stage ended with the same literal — *"cost is recorded, not refunded —
133
+ retry with `tldrx next`"* — printed whether or not the phase could fund that retry. On a phase
134
+ sized to hold exactly one attempt of its stage (what `run new` wrote before #170, and what every
135
+ run created before 2026-09-09 keeps for life) `tldrx next` comes straight back as exit 2, so the
136
+ advice sent the operator to buy the refusal themselves: measured twice in one evening on the run
137
+ #232 was filed from, and once more on a second workspace where the starved phase was the last
138
+ one. The other half of #232 shipped in `efb4eee` — `budget show`'s `next` column says `NO-RETRY`
139
+ before a cent is spent — and this is the same fact said where the operator actually is when it
140
+ bites: the line now names the refusal, both figures, the shortfall and the `budget raise` that
141
+ clears it. It predicts the GATE and not the budget's health, using the gate's OWN inputs — what
142
+ the phase has left, the same remaining-work estimate the brake compares it against, the same
143
+ `planRebalance`, and this invocation's own `--rebalance-finished` state — deliberately: #232's
144
+ own comments measured how badly a derived estimate reads on a partly unmetered run, and a second
145
+ opinion computed here would disagree with the refusal the operator then hits. **That last term is
146
+ the one a first version of this fix got wrong, and a reviewer caught it before it merged**: the
147
+ rebalance is ON by default under `run auto` (#330), so a shortfall that finished phases can cover
148
+ is moved out of them before anything is refused, and the advice was telling the operator to raise
149
+ a ceiling on a run whose very next relaunch would have funded the retry by itself. `tldrx next`
150
+ alone never rebalances, so the same starved phase is refused through one door and funded through
151
+ the other; the advice now names which door it is standing in. Under a rebalancing launch whose
152
+ donors cover the shortfall it says the retry funds itself, names the donor phase and the move,
153
+ and says that a bare `tldrx next` is not that door — it stops short of a promise, because the
154
+ grant check on each move happens after this line is written. When no finished phase can cover it,
155
+ the refusal prediction stands through either door, with how short it still is *with all of it*.
156
+ Where this gate does not decide the retry at all — `on_exceed: warn` refuses nothing,
157
+ a `host-tokens` phase is a category error the dollar brake must never judge, an `attended_by:
158
+ host` run is allowed past it by policy — the plain line is printed unchanged: those are "not this
159
+ gate's call", which is not the same as "affordable", and claiming a refusal there would be the
160
+ invented value the house rules forbid. The failure's `signature` (#297) is untouched and is still
161
+ the failure's own sentence, never this line.
162
+
163
+ - **A Build refusal built from a CACHED base reading says so, and a relaunch re-measures the red
164
+ instead of being served it (#339).** MEASURED on a live `run auto --until-done`, 2026-09-15: two
165
+ Definition-of-Done commands exited 127 on the untouched base tree, the framework refused, the
166
+ operator installed the two missing binaries and confirmed all four commands green by hand — and
167
+ the next attempt printed the SAME refusal, byte for byte, over a base tree that was by then
168
+ repaired. Then the supervisor stopped the loop on it: "a refusal that repeats verbatim is not one
169
+ a relaunch moves", with four relaunches unspent and the run ended early on a workspace where the
170
+ thing complained about was fixed. Two halves compose into that. A cached refusal reproduces
171
+ itself exactly — which is correct for a cache and is the one condition the repeat guard reads —
172
+ and nothing anywhere said the second reading had not been taken. So: a refusal whose evidence was
173
+ re-used now names it as re-used, says when it was measured, against which base sha, and when it
174
+ stops being trusted, and it adds the sentence the generic advice cannot carry — that an edit to
175
+ `.tldrx/workspace.yml` or a base that moves clears a cached reading, but a repair the cache
176
+ cannot see (a binary installed, a service started) is not re-measured until that expiry. The
177
+ supervisor's repeat guard now reads a FRESHNESS field the producer sets rather than diffing text:
178
+ a repeat whose evidence was measured stops the loop exactly as it did, a repeat that was re-used
179
+ is relaunched and says why. And the relaunch is no longer served the same cache — an attempt that
180
+ follows a refusal re-probes a RED base unconditionally, the way `--prepare` already did, because
181
+ the attempt before it asked for exactly this repair. Greens are still re-used, so the cost of
182
+ that is bounded by the commands that are actually broken. The 30-minute red TTL is untouched;
183
+ it was never what saved this run, and the reason it was not is filed separately (#340).
184
+
185
+ - **A red Definition-of-Done command is measured TWICE before it pins a story `blocked`, and the
186
+ record carries both exit codes (#163, sub-fix 1).** MEASURED on a .NET workspace, 2026-09-05: a
187
+ story's DoD gate returned `dotnet test` → exit 2, so the story blocked; the operator then ran the
188
+ identical suite twice — in the story worktree and on the epic branch — and got exit 0 with 2860
189
+ tests and 0 failing. The red was contention over a container runtime, and it had consumed the
190
+ story's last attempt. `blocked` is terminal in-run, so a human had to reopen the story by hand.
191
+ The hole was not the terminality, which is a defensible rule: it was that a story's red was never
192
+ re-measured, so a flake and a defect were written down identically and nobody reading the block
193
+ could tell which one they had. The framework already did this one level up — the base-tree
194
+ pre-flight re-measures a command the cache cannot answer — and the story's own red had no
195
+ equivalent. It does now: when a command goes red on a story, the identical command runs once more
196
+ in the same tree, and the `check.failed` event, the story's `dod` row and the blocked reason all
197
+ carry what the second run said. **Nothing about terminality changes** — a red that reproduces
198
+ blocks exactly as before, and a red that does NOT reproduce also still blocks; it is now blocked
199
+ with a sentence that says the command passed the second time, which is the difference between
200
+ reopening the story and going looking for a bug that is not there. When no second run could be
201
+ taken the record says so with the reason and no number (§7): a command whose binary the tree
202
+ never had would only measure the same absence again, and a refused command never ran, so it has
203
+ no first reading to reproduce. **A red that TIMED OUT is re-run too, so a timing-out
204
+ Definition-of-Done command now costs up to two timeout periods instead of one.** That is stated
205
+ rather than buried: the owner approved one extra run of the suite, and doubling a timeout is a
206
+ fair reading of it but was not in the framing. Timeouts are deliberately not carved out — the
207
+ measured case was contention over a container runtime, and contention is exactly what makes a
208
+ suite hang rather than fail cleanly, so a carve-out would leave the likeliest flake the one thing
209
+ never measured. The blocked reason names both readings and marks either of them `(timed out)`,
210
+ so the record says which kind of red was paid for twice. Owner decision, answered on Slack:
211
+ re-run the red command once and record both exit codes. One reading per spawn, shared with the base-tree path, so the two cannot
212
+ disagree about what a timeout's exit code is or which line of a red suite is the summary.
213
+
214
+ - **A fix-list sha that is too LONG is refused by name, instead of reading as no sha at all
215
+ (#163, sub-fix 3).** The abbreviation edge of `Resolved: yes <sha>` was closed in 0.10.0 — a
216
+ 7-39 hex claim that verifies is rewritten to the 40 `rev-parse` returned — and the same
217
+ grammar, `\b([0-9a-f]{7,40})\b`, was failing silently at the other edge, which is the
218
+ dangerous direction §7 names. No position inside a 41-character hex run is a word boundary, so
219
+ an over-long token produced NO match: the line read as a bare `Resolved: yes` that named
220
+ nothing, and the shape fault was never stated. Measured while fixing it, and worse than the
221
+ issue reported: the scan did not stop there but carried on to the next hex word on the line, so
222
+ `Resolved: yes <41 hex> (see 9f2c1ab)` closed the finding over `9f2c1ab` — a DIFFERENT commit
223
+ from the one the line claims, which is an audit record inventing its own evidence. A token of
224
+ 41+ hex characters is now refused by name and by count, through the `claimed-unverified`
225
+ sentence #130 already writes into the file, so the finding keeps holding the story and the
226
+ record says why; a refusal anywhere on the line refuses the whole read, so the answer cannot
227
+ depend on which side of the bad token a good one happens to sit. 7-39 is still accepted and
228
+ still canonicalised — nothing a person legitimately types is refused. The grammar now lives in
229
+ one leaf, `readResolvedSha`, read by both the parser and the rewriter, which is what makes the
230
+ token that gets verified and the token that gets replaced provably the same span. The refusal is
231
+ written once and read everywhere: `verifyResolutions` is the single site that withdraws a claim,
232
+ so the file, the story's blocked reason and the stdout report all carry the SAME sentence. They
233
+ did not before — the file would have named the 41-character token while the report a human reads
234
+ first said `named no commit to point at`, which is two records of one event disagreeing.
235
+
236
+ - **`budget show` says a phase cannot afford its own retry BEFORE the money is spent, instead of
237
+ labelling it `ok` up to the refusal (closes #232).** A phase whose ceiling equals ONE attempt of
238
+ its stage refuses every retry after the first cent — by arithmetic, not by policy — and the
239
+ framework's own remedy for a failed stage ("cost is recorded, not refunded — retry with `tldrx
240
+ next`") is exactly what it then refuses. Measured twice in one run, and on a second workspace in
241
+ the LAST phase, where every dollar has already gone. `run new` has sized phases at `attempts ×`
242
+ their stage since #170 so it can no longer be created, and #330's rebalance reaches it only when
243
+ some FINISHED phase has slack to give — with none, the run stops at exit 2, the one exit nothing
244
+ unattended can leave. The `next` column now answers the SIZE question in its own word: `NO-RETRY`
245
+ with the exact raise, and `n/e` for a phase there is nothing to size in. The verdict reads the
246
+ stage's DECLARED `budget_usd` and its own `attempts:`, never the derived estimate beside it —
247
+ #232's own comments measured why: on a partly unmetered run the estimate is too small, so
248
+ `ceiling / est.` reports headroom that is not there, and the less a run is metered the safer that
249
+ ratio claims it is. A declaration does not move with metering. `attempts: 1` declares no retry, so
250
+ one attempt's worth is the right size for it and still reads `ok`. Nothing about the refusal moved:
251
+ `blocked`, `remaining` and `est.` are untouched and `tldrx next` still runs. One derivation
252
+ (`retrySizing`, `budget/wouldExceed.ts`), read back additively by `budget show --json`.
253
+
254
+ - **A killed headless turn reads as `interrupted` and is handed a command the next line of code
255
+ accepts, instead of being called a `--prepare` bundle that never existed (closes #246, half
256
+ one).** MEASURED on a 0.16.1 field run: the `tldrx` process driving `run auto` died
257
+ SIGKILL-class (no handler ran), leaving the stage `running`, a dead pid in the `.lock`, and
258
+ `.agent/<stage>/pending.json` on disk. `run status` said "a `--prepare` bundle is waiting — run
259
+ the prompt and `tldrx next --commit <id>`"; that command was then REFUSED ("what is `ready`, not
260
+ `running`"), so the only advice the screen gave was a dead end, and nobody was going to run a
261
+ prompt by hand on an unattended run anyway. The mechanism is that the two states are
262
+ indistinguishable on disk: `runNext` writes the bundle BEFORE it branches on the mode, so every
263
+ headless spawn leaves exactly what a `--prepare` leaves, and `waiting.ts` derived `prepared`
264
+ from stage `running` + dead lock + a bundle without ever reading which mode opened the turn. The
265
+ ledger already knew — `stage.started` has carried `mode` since it was first written — so the
266
+ answer is now read off the newest `stage.started` for that stage
267
+ (`src/core/run/lastStart.ts`, one implementation, imported by both readers that had made this
268
+ identification independently). A tenth waiting kind, `interrupted`, says what happened and
269
+ prescribes `tldrx run auto <id>` — the command that actually resumes the run; `prepared` is
270
+ reserved for a bundle a `--prepare` really wrote. `Ctrl-C` shared the bug and is fixed with it:
271
+ `stopInFlightRun` preserved any `killed === 0` stage with a bundle as `running` forever, and now
272
+ preserves only one whose last start was a `prepare`. A ledger that does not say keeps the old
273
+ reading exactly — an absent record is not evidence of a spawn. The second half of #246, the
274
+ orphaned sub-agent's dollars being recorded nowhere while `spent_usd` reads a confident `0.00`,
275
+ is untouched and stays open.
276
+
277
+ - **A Watch retry no longer re-buys a watcher card that already validated (closes #306).** The
278
+ Watch stage fails on the FIRST card that does not validate, and `tldrx next` / `run auto
279
+ --retry-failed` re-enter the executor from the top — so a run with N features where feature k
280
+ was refused spawned all N writers again, including the k-1 cards that had already validated and
281
+ been stamped. Each of those paid at least the `$0.25` cold-spawn floor for a card nothing had
282
+ refused, and on a two-feature field run (#301, run B) with a different card refused in each
283
+ attempt, every writer was bought at least twice. #301 made the second spend an edit rather than
284
+ a rewrite; this makes it zero. The headless path now takes one snapshot of disk as the stage is
285
+ entered, before anything spawns, and keeps a feature whose card passes the SAME
286
+ `parseWatcherCard` against the SAME source context the stage validates with — so the skip can
287
+ never hide a bad card, only decline to save money, and a card whose citations went stale in the
288
+ meantime is rewritten like any other. It is a snapshot rather than a per-feature read so that
289
+ one writer loose in the run directory can never decide whether the next feature is written at
290
+ all. A kept feature still gets its `tasks[]` row, with `cost_usd: 0.0`, `model: null` and
291
+ `session_id: null` — a measured zero, not the `metered: false` of a turn billed to a host
292
+ session — and the stage's report names every kept card, so a `$0.00` row is never read as a
293
+ writer that silently did nothing. `--prepare`/`--commit` is untouched.
294
+
295
+ - **The epic branch Build cuts is claimed in `run.yml` at the cut, not when the stage returns
296
+ (closes #262).** Reported from a field workspace: a `run auto` loop was killed mid-Build,
297
+ relaunched, and was refused its OWN epic — "`epic/<slug>` already exists … refusing to stack
298
+ this run's commits onto someone else's epic" — with `tldrx next --reuse-epic` the only way back
299
+ in, a flag `run auto` cannot pass. Same sentence as #248, different mechanism, and #248's fix
300
+ cannot reach this one: #248 closed the THROW path, where a process is still alive to run its
301
+ catch. Here the process is GONE. `ExecutorOutcome.epicBranches` merges into `build.epic_branch`
302
+ only after the executor RETURNS, and a Build stage does not return for as long as a developer
303
+ turn takes — so a SIGKILL, an OOM kill or a cut session anywhere in that window left
304
+ `epic/<slug>` in the repo with nothing in the record claiming it. SIGINT and SIGTERM are hooked
305
+ (`src/cli/signals.ts`); SIGKILL cannot be hooked by any process, so no handler was ever going to
306
+ cover this. The cut now records the claim itself, through the `ctx.claimEpicBranch` seam, once
307
+ per branch rather than once per story. It does NOT ride the write serializer and is not claimed
308
+ to: of `openStory`'s eight call sites, four wrap it in `this.writes.run` and four
309
+ (`prepare`, `prepareReview`, `commitReview`, `commit`) call it bare. What makes it safe is that
310
+ there is no `await` between the branch being cut and the claim, and the claim is synchronous —
311
+ so on every one of those paths the cut and its record are one uninterrupted step. The guard is
312
+ NOT widened: a branch this run did not cut is refused exactly as
313
+ before — the claim stays a RECORD, and the fix is to write the record when it is earned rather
314
+ than to teach the guard to accept inferred evidence. A `run.yml` save that cannot take the
315
+ workspace lock is NOT swallowed: it fails the stage naming the branch, the repo, why the write
316
+ failed and both ways back in, because a claim that could not be written is this defect happening
317
+ again in silence. Honest about how far this goes: the save is atomic against a KILLED PROCESS
318
+ (temp file, then `rename`), and it is not `fsync`ed — a power cut can still lose a write the
319
+ kernel had not flushed. It is pinned by a real SIGKILL of a real `tldrx next`, never a simulated
320
+ throw, which would only have re-tested #248.
321
+
322
+ - **A Codex Build review no longer dies before the reviewer runs, and the provider's own refusal
323
+ arrives on one line (#148).** MEASURED on a live Build smoke run, published 0.7.0 +
324
+ codex-cli 0.153.0: the developer stage finished and 4 tests passed, then the review failed at
325
+ the API with `invalid_json_schema: fixlist.items.required must include every property (missing
326
+ n)`. Codex's structured-output API requires every DECLARED property to be listed as required;
327
+ `REVIEW_SCHEMA` declares optional fields at two levels — `fixlist` itself, and `n`, `severity`,
328
+ `where`, `detail`, `do_not` inside each item — and the spawn wrote that object to Codex's
329
+ `--output-schema` file verbatim. Every Codex review was refused before it began, so the story
330
+ paid an attempt for a turn the provider never ran. The translation happens at the Codex spawn
331
+ boundary ALONE (`codexSchema`, one implementation, §7): an optional property becomes
332
+ `anyOf: [<its schema>, {"type": "null"}]` and joins `required`, recursively, and the review
333
+ parser already reads a `null` exactly as it read the absent field — so nothing about
334
+ fail-closed moves, and a `fixlist` verdict with nothing readable in it is still `changes`.
335
+ `REVIEW_SCHEMA` itself is untouched and Claude's `--json-schema` still carries it byte for
336
+ byte; the test asserts both directions, because a fix at a provider seam that quietly edits the
337
+ shared contract is the failure this one was supposed to avoid. Second half, same issue: Codex
338
+ names its refusal in a pretty-printed block, and that sentence is quoted into a handoff where
339
+ every line must carry its own `[src: …]` — a multi-line reason became uncited continuation
340
+ lines that failed the claim-sources check. The Codex reason is now collapsed onto one line
341
+ rather than cut at the first, because the first line of that block is `{` and the WHY is two
342
+ lines down. Claude's failure text is passed through unchanged. The acceptance is MEASURED,
343
+ not argued from the error message: the pre-merge reviewer ran one real
344
+ `codex exec -s read-only --output-schema <the translated schema>` in an isolated directory
345
+ against an authenticated codex-cli 0.153.0 — the issue's own version — and it exited 0 with
346
+ no `invalid_json_schema` and the structured output `{"verdict": "approve", "summary":
347
+ "Approved with no findings.", "findings": [], "fixlist": null}`. That `"fixlist": null` is
348
+ the translation's whole point arriving back off the wire: the field the API refused to leave
349
+ optional comes home as the null this parser already reads as absent. Two readings, taken by
350
+ different hands — the schema bytes the child was handed, asserted in the test on every run,
351
+ and that one live call, which is not re-run by any gate.
352
+
353
+ - **A still-open fix-list finding is re-checked against the EPIC tip before the Build gate, so a
354
+ defect a LATER story closed says so with the closing sha instead of reading `Resolved: no`
355
+ (#163, sub-fix 2).** MEASURED on a Next.js workspace, 2026-09-05: at a Build gate all 6
356
+ `fix-now` findings carried a resolving sha and the 3 `defer-with-log` entries read
357
+ `Resolved: no` — "correct PER STORY", in the host agent's own words — but two of those three
358
+ defects had in fact been closed later in the same run by a DIFFERENT story, which the host
359
+ verified by reading the code. Exactly one finding had genuinely shipped unfixed, the record
360
+ could not tell those apart, and so a human narrated the difference by hand at the gate. The hole
361
+ was structural rather than careless: verification asks whether a claim's sha is reachable from
362
+ the STORY's own branch, and it asks at the moment that story settles — a fix that lands later,
363
+ on somebody else's branch, and reaches the epic through a merge did not exist when the question
364
+ was asked and is not on the branch it was asked about. There was no third answer for it, so it
365
+ was recorded as the second. There is one now. After the last story settles and before the
366
+ handoff is written, every finding not already closed with evidence is re-checked against its
367
+ epic tip, and a sha that is reachable from the epic and NOT from the story's own branch is
368
+ recorded as what it is. **The two kinds of close are spelled apart, never flattened**:
369
+ `Resolved: yes <sha>` stays the story's own close, and a close a later story landed is
370
+ `Resolved: yes-on-epic <sha>` — a different word because it is a different fact about who fixed
371
+ it and where the fix lives. `yes-on-epic` is not `yes`, so nothing the sweep writes changes what
372
+ any gate decides; the finding is exactly as open afterwards as it was before, and a finding that
373
+ genuinely shipped unfixed still reads unfixed. **Nothing is ever re-marked on inference.** A
374
+ reachable commit the record already names is the only thing that closes anything here — never a
375
+ changed file, never a heading that resembles another story's work — so the sweep cannot invent
376
+ the one direction §7 refuses to be wrong in. **And it does not go quiet when it finds nothing.**
377
+ Every finding it examined gets a `Swept:` line saying what was measured and against which tip,
378
+ because a `Resolved: no` that has been re-checked against the whole run and one that only ever
379
+ meant "correct per story" were indistinguishable, and that indistinguishability is the defect.
380
+ A sweep that could NOT be taken — no epic branch recorded, a tip git will not resolve, a file
381
+ that cannot be re-read — names its reason on every finding and in the handoff's `## Unknowns`,
382
+ rather than reporting a clean sweep over a measurement nobody took. **And it asks the weaker
383
+ question too, in words that say it is weaker.** The three entries that Next.js gate got wrong
384
+ named no commit at all — nobody had claimed one — so re-checking shas the record already names
385
+ has nothing to say about exactly the findings a human had to read the code for. For those, the
386
+ sweep asks what that human asked first: did a commit on the EPIC, and not on this story's own
387
+ branch, CHANGE the file this finding cites? Up to three such commits come back named, on the
388
+ `Swept:` line and as their own `## Unknowns` bullet — and they resolve NOTHING. A changed file
389
+ is not a closed defect, and writing `Resolved: yes` off one would be the same lie arrived at
390
+ from the other side; what this replaces is a person grepping the run's commits by hand at the
391
+ gate, never a person's judgement about them. **A probe git REFUSED says that, too.** An
392
+ unresolvable story ref and a file nobody changed both come back as no commits, and writing
393
+ "no later commit changed this" over the first would be a measurement nobody took wearing the
394
+ words of one that was — so the refusal carries git's own reason onto the `Swept:` line and
395
+ into `## Unknowns`, and is never counted as a clean probe. `Swept:` and the
396
+ `yes-on-epic` word are additive: a fix list written before this change reads exactly as it did.
397
+
398
+ ### Changed
399
+
400
+ - **The zero-touch launch recipe (EN and ES) now pairs `run auto` with `--wait-gates` and
401
+ `--wait-answers` everywhere it shows opening a run with `--gates none`, and says why: `--gates
402
+ none` sets the policy, but only `--wait-gates` lets `run auto` re-sign a parked `auto` gate once
403
+ the question that held it is auto-answered — without it the loop exits `awaiting human` on the
404
+ first gate that parks on a question, even though every one of its conditions already holds (see
405
+ #342).**
406
+
407
+ ## 0.29.0 — 2026-09-15
408
+
409
+ ### Changed
410
+
411
+ - **A reviewer's `changes` verdict must cite its evidence, or it is refused as a malformed
412
+ envelope (closes #326).** MEASURED on a field run: a reviewer returned `changes` with summary
413
+ "test" and findings `["a","b"]`, and the framework recorded it as a real review — it spent the
414
+ story's attempt, blocked `done`, and would have quoted "test" to the next developer as why the
415
+ story is not done. Nothing checked a `changes` verdict's content; only a `refuted` fix-list
416
+ finding had to carry a citation. Owner decision (2026-09-14): a `changes` now pays the same
417
+ price — its `summary`, or one line of a finding, must END with an `[src: …]` token the §2.8
418
+ parser reads (the diff line that is wrong, the unmet acceptance criterion as
419
+ `03-plan/stories/<id>.md:<line>`, or `absent:<path>` for missing work). One that cites nothing
420
+ is a fault in the REPORT, so it takes the existing bounded format retry — re-prompted with the
421
+ refusal, no attempt spent, two corrections — and the third is recorded as `changes` exactly as
422
+ before. This deliberately reverses the part of #79's pin that called a citation-less `changes` a
423
+ judgement about the work; a `changes` that DOES cite still costs the attempt. The reviewer prompt
424
+ now states the rule under `changes` and carries the `[src: …]` grammar on every review, not only
425
+ when `fixlist` is on the table. The citation is read, not resolved: a token that parses but
426
+ points at nothing still passes, so this refuses the hollow envelope, not a wrong one.
427
+ - **`tldrx run auto` now moves a blocked phase's shortfall out of finished phases by default
428
+ (closes #330).** The rebalance shipped opt-in in #314 because a phase ceiling is a person's
429
+ decision about money. An audit of `run auto`'s intervention episodes across two client
430
+ workspaces ranked stage/phase sizing as the #2 cause of a human stepping in: 22 episodes, 27
431
+ `budget.raised`, 7 `budget.blocked` (counts the audit measured, cited from #330), and not one raise note changed the work
432
+ (the audit's inferred reading of those notes). In a validation run the flag moved money twice
433
+ with nobody present. So a person was signing a move the framework could already prove safe.
434
+ Now a bare `run auto` makes it, and `--no-rebalance-finished` turns it off for a launch whose
435
+ phase ceilings must mean exactly what they say. `--rebalance-finished` is still accepted, and
436
+ passing both flags is exit 1, by name. The rules are unchanged: the run ceiling never grows, no
437
+ move passes a recorded grant, donors must be finished and metered with no unmetered turn, the
438
+ move is the exact shortfall or nothing, and each move is one `budget.raised` with the same
439
+ `source: run auto --rebalance-finished`. `tldrx next` on its own still never rebalances. The
440
+ `tldrx drive` mandate is unchanged on purpose: its unattended mode is refused `run auto`
441
+ outright (`attended_by: host`), and its attended mode drives `tldrx next`. No mandate-driven
442
+ turn ever reaches the move, so a sentence about it would teach something the driver cannot do.
443
+ - **A headless fix list now buys its fix round in the same run, and the approving reviewer's
444
+ signature closes what it was shown (#327, owner decision "A: ronda + auto-cierre").** A reviewer
445
+ that signed with a `fix-now` finding used to park the story at `review`, on the grounds that
446
+ routing a fix list needed a host — measured live, that park stranded the story and the two
447
+ stories of the next wave behind it until a person intervened. Now the story is requeued without
448
+ spending an attempt, the developer is handed the open findings, and the next reviewer is shown
449
+ each one verbatim with the rule "approve only if every one is fixed". Its `approve` rewrites each
450
+ finding it was shown as `Resolved: yes <commit> — auto-closed: …` naming the reviewer session,
451
+ attempt and run — the existing line grammar, so every reader parses it unchanged, and checked
452
+ against git before the story may settle, so a sha that is not on the story branch reopens it. A
453
+ finding that reached the file after the prompt was rendered stays open and still blocks. The
454
+ test that pinned the park ("a headless run … parks the story") is inverted, citing the decision.
455
+ A host review is shown the same findings and closes nothing. No new event, `run.yml` field or
456
+ CLI verb.
457
+ - **An `auto` Build gate no longer stops on paths outside the declared surface — it signs, and
458
+ carries them into the PR (closes #331).** A hold-surface audit of two client workspaces measured
459
+ `boundary` in 48 of 78 Build `held_by` values, the only reason on 6, and 7 intervention episodes
460
+ on it; every one ended in an approval or a `story widen` (the "no real catch" reading is inferred
461
+ from their notes). A condition that stops an unattended run and is always waved through is a
462
+ cost, not a check — but dropping it would lose the one question a reviewer asks first. So the
463
+ fact moves from the gate to the review: the auto note renders the boundary with "does not hold an
464
+ auto gate — it is carried into the PR body for review" and ends in `· warned by: boundary`;
465
+ `gate.requested` gains an additive `warned_by` beside `held_by` (present only when non-empty);
466
+ `tldrx next` prints a `warning: boundary — …` line; `tldrx run status` marks the gate row
467
+ `— warning: boundary`; and `tldrx ship` renders `## Outside declared scope — review these`,
468
+ every outside path, grouped under the story whose own measured `story.touches_widened` row names
469
+ it, the rest under `no story's measured diff names these`. One derivation: the gate and the PR
470
+ body both call `evaluateBoundary`, which now returns the uncapped list beside its detail. An
471
+ `agent` gate still falls to a person on a boundary, and a `human` gate is unchanged. **One
472
+ carve-out keeps the hold where the blast radius is not product code:** an outside path in
473
+ `SENSITIVE_PATH_CLASSES` (one exported constant beside `evaluateBoundary` — `ci`, `secrets`,
474
+ `infra`, `dependencies`) still holds an auto gate, named with its class
475
+ (`app:.github/workflows/deploy.yml [ci]`); declaring it in a story's `touches:` clears it. The
476
+ per-story grouping needs nobody to widen: the executor writes its own measured
477
+ `story.touches_widened` row for each story as it settles, and a real unattended Build then
478
+ `ship` puts the path under the story that wrote it. No exit
479
+ code changed; a check-only refusal with a boundary warning beside it is now eligible for
480
+ `run auto`'s bounded re-run, since the boundary no longer holds it.
481
+
482
+ ### Fixed
483
+
484
+ - **A Build gate now names the story a dependent WAITS on, not only the stories that are
485
+ `blocked` (closes #303).** #280 correctly left a dependent of a `review`/`in_progress` story at
486
+ `todo` instead of recording a terminal `blocked` — but the gate only ever named literal
487
+ `blocked` stories (`storiesView`'s `firstBlocked`), so the sentence saying what S2 waits on
488
+ reached `04-build/handoff.md`'s `## Unknowns` and nowhere an operator watches: `gate.requested`,
489
+ its notification and the terminal `stories:` line said `S2 not started` and named no
490
+ dependency. That is the one thing to act on — `story reopen` refuses a story already `todo`,
491
+ and `run auto --until-done` does not relaunch an awaiting-human exit. MEASURED on the #280
492
+ fixture (S1 parked at `review` by a dying reviewer): the payload carried no key for S2 at all.
493
+ Now `gate.requested` and its notification carry three ADDITIVE keys, `waiting_story`,
494
+ `waiting_on` and `waiting_reason` — absent, never null, when nothing waits — and the summary
495
+ and terminal line say S2 waits on S1 (`review`). The wait is derived with the Build loop's own
496
+ `decidingHold`/`dependencyIsPending`, so a story that will BLOCK on a terminal dependency is
497
+ never reported as merely waiting, and `waiting_reason` is `dependencyWaitReason` — the
498
+ `## Unknowns` sentence, byte for byte. `story reopen`'s `depends_on` reader moved to
499
+ `buildProgress.ts` so both ask one reader. No event version or dashboard model bump.
500
+ - **`run auto --until-done` stops a `host-tokens` refusal in tokens, not in dollar words it
501
+ never measured (closes #270).** Every `budget.blocked` went through one dollar-shaped
502
+ sentence, reading `remaining_usd` / `estimate_usd` through a helper that defaults a missing
503
+ field to `0`. The two `host-tokens` writers carry neither field, so a token refusal stopped
504
+ the supervisor with `remaining_usd $0.00 < estimate_usd $0.00` — a confident figure nothing
505
+ measured — and sent the operator to `tldrx budget raise`, a dollar command for a ceiling that
506
+ is a token allowance (and for the headless refusal, a raise that changes nothing: the way out
507
+ is `--prepare` or re-pricing to `metered-usd`, as the row's own reason already says). The stop
508
+ line now switches on `economy`, as the dashboard model already did: a `host-tokens` row names
509
+ the `host_tokens` / `ceiling_tokens` it carries and quotes its `reason`, and a dollar row
510
+ missing its figures says they are not recorded rather than printing `$0.00`. The verdict is
511
+ unchanged — a `budget.blocked` was never relaunched and still is not.
512
+ - **`budget show`, `run estimate` and the `budget-gate` hook price the remaining work with the
513
+ stage's own `attempts:` (closes #214).** `tldrx next`'s brake has resolved the stage's
514
+ `attempts:` since the key existed; these three readers asked `remainingWork` with the shipped
515
+ 2. On an `attempts: 1` stage that reserved a second developer turn and a second reviewer the
516
+ executor never dispatches, so the page said BLOCKED and the hook DENIED a `tldrx next` the
517
+ brake itself allowed. Reproduced before the fix on the Build fixture with `attempts: 1`, two
518
+ of three stories done and $5.00 left: the brake allowed ($4.40 of work), while `budget show`
519
+ and `run estimate` quoted $6.40 and the hook denied `tldrx next` on "the stage estimate is
520
+ $6.40". The hook and `budget show` now resolve the cursor stage through the same tolerant
521
+ `buildStageDefaults` the `dod-gate` hook already calls on its PreToolUse path (it now takes
522
+ the stage id; an unreadable preset still gives the shipped 2), and `run estimate` passes the
523
+ stage spec it already loads. `reviewer_share`, `story_cap_multiplier` and
524
+ `story_cap_floor_usd` are still asked with their defaults by these three readers — the same
525
+ disagreement on three more knobs, measured and filed as #333; this change moves `attempts` only.
526
+ - **An in-session Watch turn is recorded as UNMETERED, not as a measured `$0.00` (closes #224).**
527
+ The `--commit` path read the result envelope's own `cost_usd` and defaulted it to `0` — and a
528
+ host session has no reason to fill that field in — while `--cost-usd`, the flag the host
529
+ declares what its sub-agent cost with, was never read anywhere in the Watch executor. Measured
530
+ on disk across two live workspaces: 56 watch task rows at `cost_usd: 0.0` with no `metered` key,
531
+ one of them written by a current release, so every zero was counted by `budget.spent_usd` as a
532
+ real measurement and the lower-bound labelling that keys off `metered: false` could not fire for
533
+ the whole of 05-watch. Watch now uses the same three-value contract Build and the single-agent
534
+ stages already use: `--cost-usd` first, then the envelope's figure, and with NEITHER
535
+ `cost_usd: null` + `metered: false` — a named absence instead of a confident zero. A declared
536
+ `--tokens` rides along on the row the same way. A headless watch turn, which this process really
537
+ does meter, is untouched.
538
+ - **A Watch spawn now records the ceiling it was given (closes #190).** The executor computed a
539
+ per-feature ceiling, handed it to the sub-agent and emitted no `agent.spawned` at all, so a
540
+ watcher feature was the one turn in the framework whose measured cost could not be read against
541
+ what it was allowed to cost — every other stage reconciles `agent.spawned.max_budget_usd`
542
+ against the `agent.result` it paired with. Watch now emits one `agent.spawned` per feature
543
+ before the spawn (a turn that dies is still a turn the ceiling was committed to), carrying
544
+ `phase`, `role: developer`, `model`, `effort`, `max_budget_usd` and `key` — the same `key` the
545
+ `agent.result` for that row already carries, which is what joins the two ends of one turn.
546
+ Deliberately not `story:`: that is the field `tldrx cost --stories` counts a row as a Build
547
+ story by, and a Watch feature id is not a story, so that report is unchanged and whether it
548
+ grows a Watch axis stays a separate decision.
549
+ - **The conflict-turn marker guard no longer counts an unreadable path as a clean one, bounds its
550
+ scan, and reads modified files too (closes #324).** #286's guard caught every read failure and
551
+ returned "holds nothing": a guard failing OPEN, and silently, which is the one direction an audit
552
+ record may never fail in. A path that could not be read is now NAMED with its reason — the errno,
553
+ or the cap that skipped it — and only `ENOENT` stays silent, because a file that is gone cannot
554
+ carry a marker into a commit. The scan is bounded for the first time: 2,000 paths and 4 MiB per
555
+ file, neither of which can bite a normal worktree (this whole repo is 1,037 tracked files and its
556
+ largest is a 766 KB `CHANGELOG.md`, measured on 5256fab), and a path past either cap surfaces the
557
+ same way an unreadable one does rather than passing as clean — an un-ignored generated tree used
558
+ to make the guard as slow and as hungry as that tree was big, with nothing afterwards to say it
559
+ had been. The scope now also includes files MODIFIED since the handed sha, closing the hole where
560
+ a developer pastes a conflicted hunk, markers and all, into an already-tracked file; that
561
+ widening was measured before it was made, not assumed — every `M` row of 1,060 commits of this
562
+ repo's history, 8,366 rows, each row's content against the marker regex, 0 false positives (one
563
+ repo is not every repo). What to DO about an unchecked path is one decision at one call site
564
+ (`UNCHECKED_PATH_POLICY` in `src/core/build/git.ts`): it refuses the attempt with the paths
565
+ named, and a `warn` policy that lets the story walk on with them named on the Build lines is
566
+ implemented beside it.
567
+ - **A developer's `notes` written as an ARRAY of strings is now joined with newlines instead of
568
+ read as `""` by both readers (closes #217).** Measured on a 0.14.3 security-patch run: the
569
+ in-session developer wrote one note per array element, and the two readers' `typeof … ===
570
+ "string"` ternaries — `toEnvelope` on the spawned path, `readResult` on the `--commit` path —
571
+ each threw the whole field away. `notes` is the envelope's one free-text channel, so a turn's
572
+ stated caveats (on that run, what a security patch could NOT verify) vanished from `run.yml`,
573
+ the story log and the handoff with nothing on the `--commit` line saying so — a record wrong in
574
+ the dangerous direction. The elements ARE the notes, so they are joined rather than refused:
575
+ refusing would spend the turn again over a file whose content is not in doubt, on the reader
576
+ that is tolerant by design. Silence was the actual defect, so `tldrx next --commit --check` now
577
+ says which of the three readings a file gets — joined (with the count), dropped by index for a
578
+ non-string element the way `outputs[i]` already is, or still `""` for a `notes` that is neither
579
+ a string nor an array. All three paths, the check included, go through ONE function
580
+ (`readNotes`, `envelope.ts`), so the rehearsal cannot disagree with the commit.
581
+ - **The `budget-gate` hook no longer refuses the `run auto` launch that would fix the shortfall
582
+ (closes #321).** It priced a `tldrx run auto` spawn against the cursor phase's own ceiling
583
+ alone. On a phase already short, it denied the launch and wrote `budget.blocked` before the
584
+ loop could run its in-process rebalance, even when finished phases held the money. With the
585
+ rebalance on by default, that turned into the common case. Reproduced on the hook fixture
586
+ before the fix: 02-how was $2.39 short, finished 01-what held $2.86 unspent, and
587
+ `tldrx run auto` was denied. Now, when finished phases cover the WHOLE shortfall under the
588
+ loop's own rules, the gate allows and says why on stderr, writing no event. It checks that
589
+ with the same `planRebalance`, `--take-from` validation and grant verdict the loop uses, read
590
+ against the full `run.yml` so a stale donor still counts as unfinished. The gate itself moves
591
+ nothing; the move and its `budget.raised` are the loop's. Denied exactly as before:
592
+ `tldrx next`, a launch with `--no-rebalance-finished`, a run-ceiling shortfall, a shortfall
593
+ the finished phases cover only in part, a stale donor, and a move past a recorded grant.
594
+ - **A re-reviewed story's `changes` is consumed instead of parking the story (closes #327).** A
595
+ story whose earlier reviewer died or was unfunded is re-reviewed alone, and that path returned
596
+ right after the verdict — so a `changes` with attempts left parked at `review` and its
597
+ dependents never started, on the serial path and in a wave alike. It now enters the attempt loop
598
+ (or the wave's next round) under the same ledger bound every other requeue uses.
599
+ - **A story `blocked` on an approving review whose fix list was closed later now settles `done`
600
+ (closes #329).** Measured live: a fix landed, the reviewer approved, the story blocked on a
601
+ `Resolved: no` nobody had rewritten — and after a person rewrote it, `story reopen` dispatched a
602
+ developer with nothing to change and `--as-is` refused a branch already on its epic. Before the
603
+ frontier walk, `tldrx next` and `--prepare` now settle such a story `done` with no agent spawned,
604
+ after holding every `Resolved: yes` to git, and say why in one line. The refusal that blocks a
605
+ story on its fix list names that remedy instead of `tldrx story reopen`.
606
+ - **A parallel Build wave no longer hands two developers the same unspent money (closes #325).**
607
+ Every per-story cap reads the stage's METERED remainder, which is only right when stories run
608
+ one at a time. Measured live: a $21.60 stage spawned S1 under $21.00 and S2 under $16.50 in one
609
+ wave before either had metered a cent, spend landed at $33.07, and S2's reviewer was refused on
610
+ "$0.00 left" — the build ended 0 of 4. Reproduced on the fake agent before the fix: two lanes on
611
+ a $20 stage were spawned at $9.00 each, $22 claimed with their review floors. Now each lane is
612
+ capped at the stage's remainder less the caps of the lanes still running and a $2.00 reviewer
613
+ floor for every story whose review is still ahead, its own included. A lane that bound would
614
+ put under the least its developer is already allowed (`story_cap_floor_usd` for a priced story,
615
+ its uniform share otherwise) waits for a lane to finish, and the run prints `S2: not dispatched
616
+ beside S1 yet —` with the remainder, the reservation, the floors and the bound; with nothing in
617
+ flight nothing can be freed by waiting, so it runs at that floor instead of stalling. The
618
+ consequence to know: a small stage can now serialise a wide wave — measured, the third of three
619
+ unpriced stories on an $8 stage waits for the other two to meter. `--parallel 1` is
620
+ byte-identical (golden unchanged). No new config key, event or field. Not in this
621
+ change: sizing the stage from the plan once it exists, and `budget show`
622
+ naming what its estimate excludes — both named on #325.
623
+ - **The Plan prompt states the per-item character cap where each capped field is described, and
624
+ an over-cap finding now survives into the retry (closes #328).** Two plan stages failed on the
625
+ same cap in one day: a 683-character `test_plan` item, and a `wave_cap_reason` — held to the
626
+ same cap at the gate, and stated nowhere the Plan agent reads before writing it. The
627
+ `acceptance` and `test_plan` rows now say "one sentence each; split a long … into several
628
+ items", and the Plan-shape wave-cap rule and the Caps list name `wave_cap_reason`'s cap — all
629
+ rendered from `MAX_ITEM_CHARS`, never a typed number. The retry half was measured, not assumed:
630
+ a failed check's reason reaches the next `## Previous attempt` squeezed to one line with its
631
+ middle elided, and on a one-defect plan the old message lost "split it into several items" to
632
+ that ellipsis. The finding is now `<n> characters (cap <cap>) — split it into several items`,
633
+ short enough that the field, the length, the cap and the instruction reach the retry; an
634
+ over-cap `wave_cap_reason` is told its length and the cap too. A prompt byte change by
635
+ design; no Build golden moves. Not in this change: the facilitator's one-line squeeze itself,
636
+ which still elides whatever a longer multi-finding reason puts in the middle.
637
+
3
638
  ## 0.28.1 — 2026-09-15
4
639
 
5
640
  ### Fixed