tldr-experts 0.34.0 → 0.36.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,362 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.36.0 — 2026-09-18
4
+
5
+ ### Added
6
+
7
+ - **`tldrx facts dedupe [--dry-run]` retires the duplicate facts already on a ledger written
8
+ before the append-time check below existed (see #216).** A field audit measured 27 verbatim
9
+ duplicates in one workspace's `facts.yml` (F028-F054 == F001-F027, one import that ran
10
+ twice), inlined into every what/how/build prompt with no filter. The new subcommand groups
11
+ live facts by normalised text, keeps the earliest id per group, and marks every later member
12
+ `superseded_by` — chained rather than fanned onto the earliest id directly for a group of
13
+ three or more, since the reciprocal link `validateFactsFile` enforces is one-to-one; `headOf`
14
+ still resolves every chained member to the one still live. Nothing is deleted, `--dry-run`
15
+ writes nothing, and a ledger with no duplicates is byte-identical after.
16
+ - **The context ledger carries the facts share as its own line, `facts_bytes` in `pending.json`
17
+ and `(facts N B)` in `--prepare`/`--dry-run` (see #216).** A subset of `inputs_bytes`, the same
18
+ shape `questions_bytes` already is of `stage_bytes` — never summed twice into `total_bytes`.
19
+ The field audit that opened #216 found this share reaching 57% of a bundle with nothing
20
+ anywhere naming it; the number used to exist only as something you could `wc -c` a rendered
21
+ prompt to discover by hand.
22
+ - **The `plan` gate refuses a story whose own body names an edit of a path outside its own
23
+ `touches` allowlist (see #376).** Measured on a live autonomous run (#332): a story's step
24
+ asked for a CHANGELOG bullet under a new `## <v> — unreleased` heading while `CHANGELOG.md`
25
+ sat outside its `touches` — the Build developer correctly left the file untouched under the
26
+ write-allowlist rule, and the pre-merge reviewer had to send the branch back for AGENTS.md
27
+ §5's bullet, exactly the human intervention the autonomy north star measures. Same shape as
28
+ the #365 sequencing rule: a short, exported edit-verb set (`TOUCH_EDIT_VERBS`) and a
29
+ backtick-quoted path pattern (`TOUCHED_PATH_PATTERN`, never a `*` glob) in the SAME sentence
30
+ or bullet — a step that only reads or points at a path is unaffected. A `## <v> —
31
+ unreleased` heading named alongside an edit verb counts as an edit of `CHANGELOG.md` even
32
+ when the step never spells the filename. The issue's second half — a `touches_always:`
33
+ workspace-level allowlist addition — is not in scope here; it stays available as a
34
+ follow-up if this half proves insufficient in the field.
35
+ - **An over-cap `acceptance`/`test_plan` item is now split mechanically, for $0.00, before the
36
+ `plan` check ever judges it (see #352, part of #345 family).** `requireStringList`'s own
37
+ refusal already said "split it into several items"; `checkPlan` now tries exactly that,
38
+ reusing `[src:]`'s own trailing-token span finder (`trailingTokenSpan`, `srcToken.ts`,
39
+ exported for this — one derivation, AGENTS.md §7) so a split never cuts through, or drops, a
40
+ citation: the token is set aside before any sentence boundary is looked for, and copied
41
+ VERBATIM onto every resulting piece, not just the last. `splitOverCapItem`
42
+ (`src/core/text/listSplit.ts`) does the mechanical half; `repairOverCapItems`
43
+ (`src/core/plan/repairOverCapItems.ts`) does the file half — locating the field surgically in
44
+ the story's raw front-matter lines (both the flow-list and block-list shapes a story is seen
45
+ in) and rewriting only that one field, never a full YAML round-trip that would reflow
46
+ everything else's quoting. A split is written back to disk ONLY when every piece is under the
47
+ cap AND re-validates through the same checks the original item failed
48
+ (`requireStringList` + the `[src:]` grammar when the item carried a token) — an item with no
49
+ sentence boundary, or whose split still fails re-validation, is left byte-identical and
50
+ refused exactly as before, including by the existing `plan-fix` round (#288) when applicable.
51
+ Because the repair runs inside the SAME check pass, a formatting-only over-cap item never
52
+ even reaches that bounded, paid round. `CheckOutcome` grows an additive `repairs?` field
53
+ (`src/core/run/checks.ts`), surfaced on the stage's own report line and on the
54
+ `check.passed`/`check.failed` event — the same "an auto-repair must be visible" owner
55
+ decision (2026-09-15) #345's Watch repair already follows.
56
+ - **A retry of a failed single-agent stage whose declared outputs are already complete and
57
+ valid on disk settles at $0.00 instead of buying a fresh turn to reproduce them (see #353,
58
+ part of #345 family).** MEASURED first: `runStage`'s headless path called `spawnAgent`
59
+ unconditionally on every re-entry — `validateOutputs` was only ever reached from
60
+ `finishStage`, AFTER a turn was dispatched and paid for, so a stage that failed for a reason
61
+ unrelated to its declared outputs (a `checks:` failure since fixed by hand, a timeout after
62
+ the last file was flushed) bought a full fresh turn to redo work already on disk. Now, on
63
+ re-entry to a stage whose `status` is already `"failed"` (never a first attempt, and headless
64
+ mode only — never `--prepare`/`--commit`, never `--dry-run`), `trySettleFromDisk` runs
65
+ `validateOutputs` against what is already there and, when the stage declares any `checks:`,
66
+ runs them for real too (they are local — free to run, and a checks failure there means the
67
+ "already finished" premise is false and a real turn is still owed, exactly as today). Only
68
+ when BOTH clear does the stage settle: one task row (`status: "done"`, `cost_usd: 0`,
69
+ `model: null`, `session_id: null`) carrying the additive `settled_from: "prior-attempt
70
+ outputs"` (`RunFile.ts`, and `emitRunYaml.ts`'s hand-written task emitter, which needed the
71
+ new key taught to it the same way #375's `exit_disagreement` did), and a matching
72
+ `agent.result` event — named on the stage's own report line too, never silently. Any problem
73
+ falls straight through to the ordinary spawn path, unchanged. **Pre-merge review
74
+ (2026-09-18): structurally valid files are not proof of finished work when the agent that
75
+ wrote them reported the failure itself.** `trySettleFromDisk` now refuses to settle before
76
+ even looking at disk when the prior attempt's own task row names an agent-reported
77
+ `failure_kind` (`AGENT_REPORTED_FAILURE_KINDS`, `spawnAgent.ts`: `result_error`,
78
+ `malformed_result`, `empty_result`, `asked_no_diff`), `unclassified`, an absent kind on a
79
+ `"failed"` row, or `non_zero_exit` whose own `error` text shows a result document that parsed
80
+ and disagreed with itself (`… with is_error=true: …`) — only a `checks:`-only failure or the
81
+ wall dying on it (`timeout`, `process_killed`, `rate_limit`, or a clean `non_zero_exit` with
82
+ nothing parseable) settles. Named on the report line as `prior attempt reported <kind>;
83
+ spawning`. A knob two `run auto` retry-counting test fixtures grew to work around the
84
+ original, too-permissive settle is removed along with it — the failure-kind gate alone keeps
85
+ their canned "fails then succeeds" spawns real, which is the smaller, more honest fix.
86
+
87
+ ### Changed
88
+
89
+ - **A turn killed by the provider's rate limit with no attested `rate_limit_event` frame can
90
+ now be classified `rate_limit` from its raw text (part of #341).** `describeFailure`'s
91
+ residual `unclassified` bucket now checks stdout/stderr against a small exported regex
92
+ (`RATE_LIMIT_TEXT_RE`) for the REPORTED (not measured — no raw capture of a turn that
93
+ actually died against the wall exists in this repo) shape of a session-limit API error
94
+ before giving up; it never overrides a more specific cause (timeout, a kill signal, a
95
+ malformed envelope, a named error or subtype disagreement). The frame-attested capture
96
+ #341 actually asks for still does not exist — this narrows the "no reason named" bucket,
97
+ it does not close the issue.
98
+ - **`renderFacts`'s rendered list — Build's `### Facts already on record` and the `{{facts}}`
99
+ template value alike — sits under a byte ceiling (`DEFAULT_FACTS_MAX_BYTES`, 32 KB) for the
100
+ first time (see #216).** Previously unbounded: one measured section reached 123,938 B, 57% of
101
+ a 218 KB bundle. Over the ceiling the list is cut on a WHOLE fact, never mid-word (#161), and
102
+ the cut is named on the page rather than silently dropped.
103
+ - **What and How stop receiving `.tldrx/memory/facts.yml` raw — an INDEX rides in its place
104
+ (owner decision "Index", 2026-09-18; see #216).** Measured on a 120-fact fixture ledger: a
105
+ 48,150 B `## Inputs` entry became a 17,490 B index (63.7% smaller), one line per live fact —
106
+ `- [F<n>] <area> · <first ~120 chars>…` — with a closing pointer at the file for the full
107
+ text. Build and Watch declare the same input in their `stage.yml` too but read facts through
108
+ their own executors, never `inlineInputs`, so they are unaffected — measured, the declaration
109
+ there was already unused. The `## Inputs` preamble marks the summarised file with its own
110
+ sentence instead of the "nothing else to find" line that would otherwise contradict
111
+ what/how's own prose telling the agent to grep the real file.
112
+
113
+ ### Fixed
114
+
115
+ - **The Build executor's `ExecutorTask` rows never copied `AgentOutcome.failureKind` across, so a
116
+ Build developer or reviewer turn that died with a classifiable cause carried no
117
+ `failure_kind` on its `run.yml` row (see #357, see #348).** `ExecutorTask.failureKind` and
118
+ its `recordExecutorTasks` mapping already existed; the four `this.tasks.push` sites in
119
+ `executors/build.ts` were the only thing not copying it. The spawned-developer and
120
+ spawned-reviewer sites (and the format-retry row beside them) now do; the two HOST-turn
121
+ push sites, which have no `AgentOutcome` to read a kind off, are unchanged and continue to
122
+ leave the key absent rather than guess one. `test/build-golden.test.ts`'s `rounds-run-tasks.txt`
123
+ golden is regenerated ON PURPOSE — a failed developer's row in that fixture now carries
124
+ `failure_kind: "non_zero_exit"`, which it did not before; this is the change this commit
125
+ means, not a drift.
126
+ - **A spawned turn whose result document said `subtype: "success"` but whose process exit
127
+ code disagreed settled as a stage failure — costing a full `run auto` relaunch for work
128
+ that was already done (see #375, see #348).** Measured on a live autonomous run: a watch
129
+ developer wrote its declared output, `claude` exited 1 with `is_error: true`, and the
130
+ result document's own `subtype` said `"success"`; the facilitator classified this as
131
+ `unclassified` ("no reason named (the provider's own subtype said "success")") and `run
132
+ auto` spent one of five relaunches re-running work attempt 1 had already finished. The
133
+ three headless spawn sites in `runNext.ts` now settle the turn as `done` — recording the
134
+ disagreement as an additive `exit_disagreement` note on the row and the `agent.result`
135
+ event — whenever every output the stage declared is on disk and non-empty at the moment
136
+ the verdict is taken (`exitDisagreement.ts`'s `settleExitDisagreement`, one derivation,
137
+ called from all three rather than pasted three times); a turn missing even one declared
138
+ output stays `failed` under its ordinary `failure_kind`, exactly as before. The Build
139
+ executor's `spawnDeveloper` path has the same exit/subtype disagreement shape but no
140
+ per-story declared-outputs list to check existence against, so it is deliberately left
141
+ out of this change — see the follow-up draft at
142
+ `scratchpad/issue-draft-375-build-spawndeveloper-exit-disagreement.md`.
143
+ - **`FactsStore.append` dedupes on normalised text, so re-running a driver's import or
144
+ `tldrx facts add` on the same sentence no longer mints a second row (see #216).** Both
145
+ writers — `captureAnswers` and `tldrx facts add` — now get the existing live fact back
146
+ (`duplicate: true`) instead of a new one when the incoming text, trimmed/whitespace-
147
+ collapsed/case-folded, matches a LIVE fact's; a retired or superseded twin does not block a
148
+ fresh assertion of the same text. `facts add` prints "`<id>` is already on record" and exits
149
+ 0 rather than writing a duplicate.
150
+ - **`merge-wave.sh`'s #336 immutability guard no longer prints `printf: write error: Broken
151
+ pipe` once per dated CHANGELOG section (see #380).** Its `changelog_section()` awk helper
152
+ `exit`ed as soon as it saw the NEXT `## ` heading, while the `printf | changelog_section`
153
+ pipe feeding it the pre-merge CHANGELOG was still writing the rest of the file — so the
154
+ writer could get EPIPE (measured live: 45-46 lines of noise per wave, once per dated
155
+ section that was not the last one in the file). `exit` is now `p = 0`: the awk process
156
+ keeps consuming to EOF instead of closing its end of the pipe early, so the writer always
157
+ finishes; the guard's own verdict (`$WAS`/`$NOW`) is unchanged.
158
+ - **`emitFactsYaml` writes `facts: []` for an empty list instead of a bare `facts:` line (see
159
+ #383).** Both YAML readers behind the runtime seam parse a key with nothing after it as
160
+ `null`, not `[]`, so a facts.yml with zero facts never round-tripped — `FactsStore.save()`
161
+ re-validates its own serialized output before writing and threw on exactly this file,
162
+ `refusing to write an invalid facts.yml: facts expected an array, got null`. Latent until
163
+ now (every shipped writer appends a fact before calling `save()`), but reachable from any
164
+ fresh `FactsStore.loadOrEmpty(...).save()` with nothing appended; non-empty output is
165
+ unchanged.
166
+ - **Watch's kept-card pre-pass (#306) never got the `[src:]` mechanical repair the
167
+ validation loop a few lines below it already had, so a card left on disk with only a
168
+ punctuation slip — the marker missing its space, a token sitting mid-sentence, an ASCII
169
+ `->` in a `cmd` source — re-spawned a writer for text a script already knows how to fix
170
+ (see #351, part of #345).** `keptCard` now runs the same `repairSrcSyntax` before its
171
+ `parseWatcherCard` check: a repair that turns an invalid card valid is written back to
172
+ disk and the feature is kept, at $0.00, with the repair named on the stage's own report
173
+ line the same way the validation loop names its own. A card that stays invalid after
174
+ repair is still not kept (today's behaviour, byte-identical file); a card already valid
175
+ needs no repair and is never rewritten.
176
+
177
+ ## 0.35.0 — 2026-09-17
178
+
179
+ ### Fixed
180
+
181
+ - **The phase-id shape guard (`test/phase-ids-one-derivation.test.ts`, #187) now catches a
182
+ reordered duplicate, not just an identical or drifted one (see #204).** The guard's
183
+ `ORDERED_LIST` pattern was built from `PHASE_IDS` by joining the five ids in the one order
184
+ `PHASE_IDS` itself is written in, so a fourth copy that wrote the same five ids out in a
185
+ DIFFERENT order — `["01-what","02-how","04-build","03-plan","05-watch"]` — matched neither
186
+ guard: not the shape guard (wrong order) nor the equality guard (nothing there compares an
187
+ unrelated inline literal to `QUESTION_PHASES`). Planting exactly that array in an unrelated
188
+ file left the suite green, 5 pass / 0 fail — confirmed here before the fix, reproducing the
189
+ issue's own finding. The pattern is now the alternation of all 5! = 120 orderings of the same
190
+ five exact `PHASE_IDS` (still derived from `PHASE_IDS`, never hand-typed), each id still
191
+ requiring its closing quote immediately after it so ordinary path-segment strings like
192
+ `01-what/intent.md` keep not matching. A real drift (a wrong id) still matches none of the
193
+ 120 orderings, so the drifted case still falls to the equality guard exactly as before —
194
+ only the ORDER stopped being load-bearing for the shape guard.
195
+ - **`validateWorkspace` now enforces spec §2.1's no-shell-metacharacter rule for
196
+ `repos[].commands` on LOAD, not only in `tldrx init`'s own emitted document (see #181).**
197
+ A hand-edited `.tldrx/workspace.yml` carrying a bare `& ; | > \`` in a command — e.g.
198
+ `commands.lint: "npm run test | tee lint.log"` — loaded with no complaint, because
199
+ `validateWorkspace` (`src/core/schemas/workspace.ts`) checked `mode`/`root`/`repos[].name`/
200
+ `.path`/`.command_probes` but never touched `repos[].commands` at all; only
201
+ `validateEmitted.ts`'s own `tldrx init`-time check ever ran the rule. The refusal now names
202
+ the command, its slot and the offending character, reusing the same `isSingleArgvCommand`
203
+ both enforcement points always shared, so the two can never drift on what counts as a
204
+ metacharacter. `null` and `""` both still mean "unavailable" (the shipped
205
+ `templates/workspace.yml` skeleton's own convention) and carry nothing to check.
206
+ - **The rule above is quote-aware, so it never refuses a command the DoD gate would actually
207
+ run (follow-up to #181, same day).** The first cut of the load-time check scanned the whole
208
+ command string, so `sh -c "npm run test | tee out.txt"` — the exact workaround
209
+ `docs/guide/09-troubleshooting.md` documents for a command that needs a shell — failed to
210
+ LOAD even though `src/hooks/lib/story.ts`'s own `splitArgv` already runs it fine (a quoted
211
+ token is one literal argument, never a shell operator). Load-time validation stricter than
212
+ the executor it exists to describe would have broken every workspace that followed that
213
+ documented advice. The tokenizer that already told the two apart lived only in `story.ts`;
214
+ it is now `src/core/detect/argvSplit.ts`, and both `story.ts`'s `splitArgv`/
215
+ `unquotedShellSeparator` and `commands.ts`'s `isSingleArgvCommand`/`singleArgvViolation`
216
+ read a command through it — one derivation of "quoted vs bare" for the whole codebase, each
217
+ caller still applying its own banned-character set (the gate's fifteen, spec §2.1's five)
218
+ unchanged.
219
+ - **`tldrx story reopen` now refuses on a run that is `done` or `cancelled` (`isFinished`),
220
+ or whose Build gate is signed, instead of writing a story back to `todo` that nothing will
221
+ ever dispatch (see #227).** `rollUp` derives a run's `done` from its STAGES being terminal,
222
+ not from every story reaching `done` — a `blocked` story can outlive the run it belongs to,
223
+ so the existing `row.status === "done"` guard never saw this case: measured on a real
224
+ workspace, the reopen wrote the file, appended the event, and `tldrx next` answered "is
225
+ done — nothing to advance" without ever naming the story it had just reopened. Pre-merge
226
+ review caught that the first cut of the fix only checked `=== "done"` and let a `cancelled`
227
+ run through the same hole — `run cancel` never touches a story file either — so the check
228
+ now reads `isFinished` (`RunFile.ts`), the same predicate `runNext.ts`'s own `advance()`
229
+ already refuses to dispatch on. The check runs before either write, for the plain verb and
230
+ `--for-fix`/`--as-is` alike, and refuses in the usage family (exit 1, per the owner's
231
+ decision, distinct from this file's other exit-2 refusals) naming the story and the door —
232
+ `tldrx reject --stage 04-build/build --note "…"`, or a new run — rather than leave a
233
+ signature standing over work it never saw (see #144). `next`'s own "is done — nothing to
234
+ advance" line now also names any `story.reopened` a run already carries that this fix
235
+ cannot retroactively undo, so a run an older binary left in this state is not silent about
236
+ it either.
237
+ - **A `budget.yml` that exists but will not parse no longer resurrects `run.yml`'s frozen
238
+ creation ceiling as a live one (see #245).** Since #236 the dashboard headline and
239
+ `tldrx replay` read the run's ceiling from `budget.yml`, falling back to run.yml's mirror —
240
+ a value `budget raise` never moves — whenever that file could not be read. On a run raised
241
+ after creation and then damaged, that fallback printed #236's own symptom again: `$X spent
242
+ of $Y ceiling` with `Y < X`. Owner decision ("Null con razón"): a DAMAGED `budget.yml` nulls
243
+ the ceiling instead, with an additive `ceilingBasis`/`ceilingReason` on the loaded run and
244
+ dashboard model naming why, and every renderer falls back to the same `$?` it already prints
245
+ for an unreadable spend. A run with no `budget.yml` at all is a different, unchanged case —
246
+ it was never raised through a file it never had, so the mirror is kept exactly as before.
247
+
248
+ ### Added
249
+
250
+ - **A human gate approval whose note names the commit that fixes an open `fix-now` fix-list
251
+ finding now closes that finding — only the finding the note is plainly about (see #344, owner
252
+ decision "Sí cerrarlo").** Before this, `tldrx approve --note "…"` never read the note back
253
+ against a run's fix lists, so a Build gate approved over a note like "S3 fix round landed and
254
+ re-reviewed (`71bfe1e`)" left the fixlist file itself saying `Resolved: no` — the
255
+ honour-system gap #269's on-disk escape hatch left standing — and `run.yml`'s own outcome kept
256
+ reporting the story `blocked` on a defect the gate had already verified fixed. `approve` now
257
+ reads its note for sha-looking tokens and holds each one to the exact evidence rule a
258
+ file-written `Resolved: yes` is held to (extracted to `src/core/build/resolutionVerify.ts` so
259
+ the Build executor's own per-story check and this gate path share one derivation, never two).
260
+ A finding is only ever touched when the note names its own story id (`S3`) or a candidate is
261
+ reachable from its own story branch; a reachable candidate closes it, `Resolved: yes <sha>`,
262
+ and one that does not check out is recorded `claimed-unverified`. Pre-merge review on #344
263
+ caught the first cut trying every candidate against every open finding in the whole run, so a
264
+ note naming only one story's fixing sha stamped an UNMENTIONED finding in a different repo
265
+ `claimed-unverified` over a claim nobody made about it — a finding the note never named, and
266
+ whose branch none of the note's candidates reach, is now not touched at all, not even read. A
267
+ candidate that matches no story anywhere is recorded once on the gate's own `gate.approved`
268
+ event instead of being guessed onto a finding. A note naming no commit touches no fix list at
269
+ all.
270
+ - **A fact can carry a `was_id` alias to the id it was cited under before an id-change (see
271
+ #338).** `id` stays immutable (§7: nothing ever rewrites a fact's own `id`), so an owner
272
+ decision (2026-09-17) settled how a fact's id is ever allowed to change: additively, never by
273
+ rewriting or a bare documentation-only rule. `[src: <was_id>]` now resolves through
274
+ `resolveFact` to the fact that carries the alias — but only when the cited id has no live row
275
+ of its own, so an old citation written before a renumber still resolves once its old id is
276
+ fully vacated. A cited id that is ITSELF a live fact is never silently shadowed by another
277
+ fact's `was_id`: `resolveFact` refuses, naming both the live holder and the `was_id` carrier,
278
+ because `readFacts` never validates and a hand-edited or merged facts.yml can carry exactly
279
+ the collision `validateFactsFile` refuses at write time. `validateFactsFile` itself refuses a
280
+ `was_id` that collides with another live fact id or is claimed by more than one fact. The
281
+ field is optional and additive: a facts.yml with no `was_id` anywhere reads and resolves
282
+ exactly as before.
283
+
284
+ ### Changed
285
+
286
+ - **The fix list's on-disk door is now documented as the human escape hatch it was already
287
+ behaving as (see #269).** `parseFixlistFile` — the read path every settle-time question and a
288
+ hand edit both go through — validates shape only and never charges the `[src: …]` citation
289
+ the reviewer-envelope write path (`parseFixFindings`) requires for `refuted` and for an
290
+ unblocking `docs`/`style` finding (#255). That asymmetry predates #255 and was previously
291
+ unwritten, so the next reader could read it as an oversight and "fix" it; the owner settled it
292
+ (2026-09-17) as deliberate — a person editing the artefact by hand is trusted the way the
293
+ envelope's own reviewer is not — and `docs/spec.md` and a comment at `parseFixlistFile` now say
294
+ so, naming exactly what a hand edit bypasses.
295
+ - **Two illustrative examples stop naming a private workspace, and a guard now catches the
296
+ next one (see #191).** An audit for private-workspace leakage across the public surface found
297
+ two purely illustrative hits with no measurement attached — an area-id string in
298
+ `src/core/dashboard/render.ts`'s dashboard-layout comment, and `docs/spec.md`'s running
299
+ workspace.yml / competencies / questions / handoff example (§2.1, §2.6, §2.13, §2.14) — that
300
+ had copied a real workspace's name instead of inventing one; both now use one neutral,
301
+ internally-consistent placeholder. `test/public-surface-consistency.test.ts` gained a guard
302
+ over `src/**`, `docs/**`, `templates/**`, `docs-site/**` and the top (unreleased) CHANGELOG
303
+ heading: `templates/**` and `docs-site/**` fail on any occurrence, and `src/**`/`docs/**` —
304
+ which still carry real measurement citations naming a real workspace, cited evidence rather
305
+ than decoration — fail on a NEW one, via a per-file hit-count allowlist that can only be
306
+ lowered, never silently raised. Those citations, and the two `test/fixtures/**` families this
307
+ same audit found are real workspace data rather than a synthetic name that merely coincides,
308
+ were #389's call (the owner decided 2026-09-17 to synthesize them, on its own branch) — see
309
+ the next entry for that follow-through.
310
+ - **The two `test/fixtures/**` families a real workspace's own bytes had leaked into are now
311
+ genuinely invented (not just a token renamed inside the same real paths and class names —
312
+ a first pass at this made that mistake, caught in pre-merge review), and #191's guard is a
313
+ blanket ban over `test/**` instead of a per-file allowlist over `src/**`/`docs/**` (see #389).**
314
+ `test/fixtures/competencies/read-only-expert.yml` and its paired knowledge fixture — `git mv`'d
315
+ again, to a name that no longer echoes the real one — were copied verbatim from a real
316
+ workspace, architecture and all; the `aidlc-intent` fixture tree (`ideation/feasibility`,
317
+ `ideation/scope-definition`) carried a real pilot's own product name, team and stack. All three
318
+ are now a different invented stack with invented module, file and class names throughout — a
319
+ token-normalised diff against the pre-#389 bytes is ~100% changed on every data-bearing line,
320
+ and grepping every distinctive identifier from the old files against the new ones returns zero
321
+ — while keeping the same shape and size class: same evidence-row counts, same claim/answer
322
+ counts `test/distill.test.ts` measures, same three execution-claim citations
323
+ `test/knowledge-value.test.ts` measures, so every test that reads them keeps asserting the same
324
+ behaviour against different bytes. A real absolute home-directory path in `docs/plans/v1.md` is
325
+ now home-relative. The 37 `src/**` and 7 `docs/**` citations #191's allowlist exempted, and 50
326
+ more the same audit had missed across 24 `test/**` files (evidence-comment citations and a
327
+ handful of test DATA — a fake root path, a fake GitHub repo slug — fed to the code under test),
328
+ are each rephrased to name the evidence — a workspace id, the run, the date, the figure —
329
+ without the workspace's own name; two of those files (`test/plan-skill.test.ts`,
330
+ `test/maintain-skill.test.ts`) carried their own copy of the private-workspace pattern for an
331
+ unrelated skill-content guard, and now import it from the one file still allowed to hold it.
332
+ With nothing left needing an allowlist, `test/public-surface-consistency.test.ts` drops its
333
+ `SRC_CITATION_COUNTS`/`DOCS_CITATION_COUNTS` maps and now refuses ANY occurrence — in file
334
+ CONTENT and tracked file names — across `src/**`, `docs/**`, `templates/**`, `docs-site/**`,
335
+ `test/**` (itself excepted, since the pattern has to live somewhere to be checked against) and
336
+ the unreleased CHANGELOG heading. Released CHANGELOG sections are history and stay out of the
337
+ guard's reach.
338
+ - **The private-workspace guard matched one literal spelling per name, so the same four
339
+ workspaces kept leaking in every OTHER shape (see #389, #191).** `git grep` measured what the
340
+ regex could not see: a hyphenated repo slug (`<word>-api`, `<word>-platform`,
341
+ `<word>-platform-abstractions`, a knowledge-file name), the bare word used as an English
342
+ project name with no fixed noun after it ("the `<word>` run", "The `<word>` graph", "the whole
343
+ `<word>` sample" — an open-ended list, so disambiguated by the article "the" up to two words
344
+ earlier instead), the same bare word as DATA inside a `repos: [<word>]` list literal, the
345
+ capitalised word as a bare product name mid-sentence, and two OTHER private workspace names
346
+ this guard never scanned for at all — already banned for the skill markdown by
347
+ `test/maintain-skill.test.ts`, but leaking in `src/core/run/ship.ts` and
348
+ `test/ship-policy.test.ts` as a `dev/<name>` branch in a real PR citation, now `dev/W3`. A real
349
+ GitHub `owner/repo` remote (`test/interview.test.ts`, `test/process-answers.test.ts`, 17
350
+ occurrences) is now an invented `acme/billing-api`, keeping every `parseGithubRemote` shape
351
+ (https, ssh, scp-style, `.git` suffix, trailing slash, the negative cases) exercising the same
352
+ parse. Eight sub-patterns now cover all of it, engineered so the bare lowercase word — an
353
+ ordinary Spanish verb ("it appears") used constantly and legitimately in `docs-site/es/**` and
354
+ Spanish audit prose — is never matched on its own; a new false-positive control test asserts
355
+ two genuine Spanish sentences survive. A commit message from the previous
356
+ pass on this branch had leaked one such identifier itself (rewritten, since the sha had not
357
+ reached `main`) and carried its own review record, which is stale the moment a sha changes and
358
+ was dropped rather than rewritten.
359
+
3
360
  ## 0.34.0 — 2026-09-17
4
361
 
5
362
  ### Fixed
package/README.md CHANGED
@@ -335,6 +335,8 @@ back on the registry is 0.3.0.
335
335
 
336
336
  | Version | Date | Status | Contains |
337
337
  |---|---|---|---|
338
+ | 0.36.0 | 2026-09-18 | `beta` | Wave six, the P1 cost and autonomy backlog, every branch reviewed before main: facts.yml stops riding whole into every prompt (a one-line index for what/how, a named byte ceiling for Build, dedupe at append and `tldrx facts dedupe`, the facts share visible in the context ledger); a headless turn whose exit code disagrees with its own `subtype: success` settles as done with the disagreement recorded instead of buying a relaunch; Build rows carry `failure_kind`; a rate-limit death is classified from its raw text; three more savings before a paid turn — kept Watch cards get the `[src:]` repair, over-cap acceptance items are split mechanically, and a retry whose prior attempt died by the wall with complete outputs settles at $0, never when the agent itself reported the failure; the plan gate refuses a step that edits a path outside its `touches`; an empty facts ledger round-trips; merge-wave stops spraying Broken pipe — 10 issues |
339
+ | 0.35.0 | 2026-09-17 | `beta` | Wave five of the audited backlog, every branch reviewed before main: a human gate approval whose note names the fixing commit closes that fix-list finding (reachable from the story's own branch, or `claimed-unverified` — never a silent close), a fact keeps a `was_id` alias across an id change, `validateWorkspace` enforces the quote-aware single-argv rule the DoD gate already applied, `tldrx story reopen` refuses a finished or Build-approved run and names the way back, a damaged `budget.yml` reads as absent-with-reason instead of resurrecting a frozen ceiling, the fix list's on-disk door is documented as the human escape hatch, and the public surface stops naming private workspaces — fixtures are synthetic, citations carry neutral labels, and a blanket guard refuses the next one — 9 issues |
338
340
  | 0.34.0 | 2026-09-17 | `beta` | Four waves of the audited backlog, every branch reviewed before main: run auto no longer stops itself (parked auto gates self-close, Build parks at a story boundary on budget), gates cannot sign stale or backdated evidence, interrupted turns bank their cost as absent instead of a measured zero, the boundary condition diffs the branch the run recorded, mutating generators are refused from dod blocks, merge-wave refuses edits to released CHANGELOG sections and answers `unknown` instead of `idle`, and the test harness refuses a TMPDIR inside a work tree — 14 issues, the first of them shipped end-to-end by a tldrx run against this repo |
339
341
  | 0.33.0 | 2026-09-17 | `beta` | Six follow-ups from the 0.32.0 cycle, each measured on the second unattended run (run #42) or on the review that shipped it. The plan sequencing rule counts a bare camelCase field only when it is unambiguous or the plan spells it elsewhere as snake_case or in backticks, so unrelated stories sharing a word like `toString` no longer refuse each other (#370); `task.done` clamps an oversized `as_is_note` the way it clamps `permission_refused` (#368); the asked-no-diff and red-DoD requeues settle through one helper that takes data (#369); a workspace can opt its base pre-flight probe into a story-shaped worktree with `probe_in_worktree: true`, and the Build-entry probe runs `tool_restore:` first (#371, the remainder of #363); AGENTS §11 names the negated closing verb that still auto-closes and widens the pre-wave grep (#372); and CI keeps merge-wave's sandboxes on failure so the next #237 red carries its `merge.log` (#373). Minor release: a new workspace field and a new plan refusal shape. |
340
342
  | 0.32.0 | 2026-09-16 | `beta` | Seven framework defects measured on the second unattended run (run #42), where every one of nine operator stalls was the framework's, not the code's. A refused compound shell line is a rework note for the developer instead of a human stop (#360); a dependency hold is re-evaluated the moment the dependency lands (#361) and a rescued story is written back to `todo` the instant the release is decided (#366); an event payload over the size cap is clamped to a bounded head instead of failing the whole build stage (#359); a headless developer is told it is unattended, that `## Inputs` is where to start and not an allowlist, and a turn that asks instead of deciding is a named `asked_no_diff` failure (#364); the rate-limit park waits for 90% of the window, not the provider's first warning at half of it (#367); the plan refuses a story that enforces an invariant on a field before the story that populates it, matching the field across spellings (#365); and the preflight restores workspace tools before its first probe and names the tree it ran in (#363, in part). Minor release: a new plan refusal, a changed park threshold and a new failure kind. |
@@ -1,24 +1,24 @@
1
1
  #!/usr/bin/env node
2
2
  import {
3
3
  conflictOf
4
- } from "./chunk-9kzevq39.js";
4
+ } from "./chunk-sqkp3405.js";
5
5
  import {
6
6
  FactsStore,
7
7
  formatJaccard
8
- } from "./chunk-bfvcebn6.js";
8
+ } from "./chunk-fbn9q035.js";
9
9
  import {
10
10
  parseHookInput,
11
11
  readStdin
12
- } from "./chunk-jfaarcb1.js";
12
+ } from "./chunk-73h7hmfx.js";
13
13
  import"./chunk-ae6bkfs5.js";
14
14
  import {
15
15
  EventLog,
16
16
  PHASE_ID_RE
17
- } from "./chunk-4d22xm3c.js";
17
+ } from "./chunk-6cptynke.js";
18
18
  import {
19
19
  PHASE_IDS
20
- } from "./chunk-kcrphpwc.js";
21
- import"./chunk-7ypyp8xk.js";
20
+ } from "./chunk-mx1g05pq.js";
21
+ import"./chunk-gqfbnr1h.js";
22
22
  import {
23
23
  ADVISORY_KEY,
24
24
  MAX_FACT_CHARS,
@@ -28,7 +28,7 @@ import {
28
28
  renderQuestionBlock,
29
29
  replaceBlock,
30
30
  serializeQuestions
31
- } from "./chunk-sjykymar.js";
31
+ } from "./chunk-7p9vw1s2.js";
32
32
  import {
33
33
  ITERATION_ONLY_SLOT,
34
34
  PROJECT_FRAMEWORK_DIR,
@@ -38,7 +38,7 @@ import {
38
38
  isScopedTemplate,
39
39
  parseYaml,
40
40
  scopedSlotOf
41
- } from "./chunk-0a3b8w1c.js";
41
+ } from "./chunk-d2g7h2qb.js";
42
42
 
43
43
  // src/hooks/answer-capture.ts
44
44
  import { existsSync as existsSync5 } from "fs";
@@ -466,7 +466,7 @@ function captureAnswers(questionsPath, ctx) {
466
466
  const prov = answerProvenance(block, ctx.overrides, ctx.repoNames, []);
467
467
  const text = factTextFor(block.title, block.answer);
468
468
  const clash = conflictOf({ match: block.title, area, text }, store.active);
469
- const fact = store.append({
469
+ const { fact, duplicate } = store.append({
470
470
  fact: text,
471
471
  ...truncated ? { truncated: true } : {},
472
472
  ...clash === null ? {} : { conflicts_with: [clash.fact.id] },
@@ -489,15 +489,16 @@ function captureAnswers(questionsPath, ctx) {
489
489
  cost_usd: 0,
490
490
  payload: { q: block.id, answer: block.answer, fact: fact.id }
491
491
  });
492
- log.tryAppend({
493
- ts: ctx.at,
494
- run: ctx.run,
495
- stage: null,
496
- type: "fact.added",
497
- actor: ctx.actor,
498
- cost_usd: 0,
499
- payload: { fact: fact.id, area: fact.area, kind: fact.kind, q: block.id }
500
- });
492
+ if (!duplicate)
493
+ log.tryAppend({
494
+ ts: ctx.at,
495
+ run: ctx.run,
496
+ stage: null,
497
+ type: "fact.added",
498
+ actor: ctx.actor,
499
+ cost_usd: 0,
500
+ payload: { fact: fact.id, area: fact.area, kind: fact.kind, q: block.id }
501
+ });
501
502
  captured.push({
502
503
  q: block.id,
503
504
  fact: fact.id,
@@ -1,15 +1,15 @@
1
1
  #!/usr/bin/env node
2
2
  import {
3
3
  budgetGateDeny
4
- } from "./chunk-vhn2vz83.js";
4
+ } from "./chunk-h543njq3.js";
5
5
  import {
6
6
  allow,
7
7
  deny,
8
8
  readPayload,
9
9
  runHook,
10
10
  toolInput
11
- } from "./chunk-d3jrf6yy.js";
12
- import"./chunk-jfaarcb1.js";
11
+ } from "./chunk-re1efynj.js";
12
+ import"./chunk-73h7hmfx.js";
13
13
  import {
14
14
  REBALANCE_SOURCE,
15
15
  applyRebalance,
@@ -25,12 +25,12 @@ import {
25
25
  validateRunBudget,
26
26
  wouldExceed,
27
27
  wouldExceedHostTokens
28
- } from "./chunk-dw8gggv9.js";
28
+ } from "./chunk-kcwv344p.js";
29
29
  import {
30
30
  EventLog,
31
31
  currentActor,
32
32
  nowRfc3339
33
- } from "./chunk-4d22xm3c.js";
33
+ } from "./chunk-6cptynke.js";
34
34
  import {
35
35
  cursorStage,
36
36
  hostTokensIn,
@@ -39,20 +39,20 @@ import {
39
39
  newestActiveRun,
40
40
  renderRunEconomies,
41
41
  runSpend
42
- } from "./chunk-57s4mdhy.js";
42
+ } from "./chunk-qpy8tgtc.js";
43
43
  import {
44
44
  buildStageDefaults
45
- } from "./chunk-kcrphpwc.js";
45
+ } from "./chunk-mx1g05pq.js";
46
46
  import {
47
47
  noteDeprecations
48
- } from "./chunk-sjykymar.js";
48
+ } from "./chunk-7p9vw1s2.js";
49
49
  import {
50
50
  PROJECT_WORK_DIR,
51
51
  findWorkspaceRoot,
52
52
  locateWork,
53
53
  parseYaml,
54
54
  stageYamlPath
55
- } from "./chunk-0a3b8w1c.js";
55
+ } from "./chunk-d2g7h2qb.js";
56
56
 
57
57
  // src/hooks/budget-gate.ts
58
58
  import { existsSync as existsSync2, readFileSync as readFileSync2, statSync } from "node:fs";
@@ -2,7 +2,7 @@ import {
2
2
  GATE_POLICIES,
3
3
  validateGatesPolicy,
4
4
  validateStagePolicy
5
- } from "./chunk-kcrphpwc.js";
5
+ } from "./chunk-mx1g05pq.js";
6
6
  import {
7
7
  asDocument,
8
8
  isRecord,
@@ -15,7 +15,7 @@ import {
15
15
  requireString,
16
16
  requireVersion,
17
17
  result
18
- } from "./chunk-0a3b8w1c.js";
18
+ } from "./chunk-d2g7h2qb.js";
19
19
 
20
20
  // src/core/events/EventLog.ts
21
21
  import { appendFileSync, existsSync, mkdirSync, readFileSync, statSync, writeFileSync } from "node:fs";
@@ -1,6 +1,6 @@
1
1
  import {
2
2
  runtime
3
- } from "./chunk-0a3b8w1c.js";
3
+ } from "./chunk-d2g7h2qb.js";
4
4
 
5
5
  // src/core/hooks/passthrough.ts
6
6
  async function readStdin() {
@@ -3,7 +3,7 @@ import {
3
3
  parseSrcToken,
4
4
  srcToken,
5
5
  withoutSrcToken
6
- } from "./chunk-0a3b8w1c.js";
6
+ } from "./chunk-d2g7h2qb.js";
7
7
 
8
8
  // src/core/text/questions.ts
9
9
  var REQUIRED_METADATA_KEYS = ["id", "status", "area", "asked_by", "asked_at"];
@@ -10,7 +10,7 @@ import {
10
10
  storyDependsOn,
11
11
  tallyOf,
12
12
  validateRunBudget
13
- } from "./chunk-dw8gggv9.js";
13
+ } from "./chunk-kcwv344p.js";
14
14
  import {
15
15
  EventLog,
16
16
  OUTCOME_NOT_RECORDED,
@@ -24,12 +24,12 @@ import {
24
24
  nowRfc3339,
25
25
  stageAt,
26
26
  validateRunFile
27
- } from "./chunk-4d22xm3c.js";
27
+ } from "./chunk-6cptynke.js";
28
28
  import {
29
29
  cursorStage,
30
30
  isAttendedByHostView,
31
31
  openRunViews
32
- } from "./chunk-57s4mdhy.js";
32
+ } from "./chunk-qpy8tgtc.js";
33
33
  import {
34
34
  backupPathFor,
35
35
  isAlive,
@@ -38,13 +38,13 @@ import {
38
38
  workspaceRootOfRunDir,
39
39
  writeAtomic,
40
40
  yamlScalar
41
- } from "./chunk-7ypyp8xk.js";
41
+ } from "./chunk-gqfbnr1h.js";
42
42
  import {
43
43
  isAdvisory,
44
44
  noteDeprecations,
45
45
  openBlocks,
46
46
  parseQuestions
47
- } from "./chunk-sjykymar.js";
47
+ } from "./chunk-7p9vw1s2.js";
48
48
  import {
49
49
  FRAMEWORK_ROOT,
50
50
  MAX_ITEM_CHARS,
@@ -55,7 +55,7 @@ import {
55
55
  parseYaml,
56
56
  parseYamlRepairing,
57
57
  withoutSrcToken
58
- } from "./chunk-0a3b8w1c.js";
58
+ } from "./chunk-d2g7h2qb.js";
59
59
 
60
60
  // src/core/statusline/runSnapshot.ts
61
61
  import { existsSync as existsSync4, readFileSync as readFileSync5 } from "node:fs";
@@ -140,6 +140,8 @@ function task(t, indent) {
140
140
  `${inner}started_at: ${yamlScalar(t.started_at)}, ended_at: ${yamlScalar(t.ended_at)},`,
141
141
  ...t.stopped_by === undefined || t.stopped_by === null ? [] : [`${inner}stopped_by: ${yamlScalar(t.stopped_by)},`],
142
142
  ...t.failure_kind === undefined || t.failure_kind === null ? [] : [`${inner}failure_kind: ${yamlScalar(t.failure_kind)},`],
143
+ ...t.exit_disagreement === undefined || t.exit_disagreement === null ? [] : [`${inner}exit_disagreement: ${yamlScalar(t.exit_disagreement)},`],
144
+ ...t.settled_from === undefined || t.settled_from === null ? [] : [`${inner}settled_from: ${yamlScalar(t.settled_from)},`],
143
145
  ...t.banked_before_refusal === undefined ? [] : [`${inner}banked_before_refusal: true,`],
144
146
  ...t.dedupe === undefined ? [] : [`${inner}dedupe: ${yamlScalar(t.dedupe)},`],
145
147
  ...t.duration_ms === undefined ? [] : [`${inner}duration_ms: ${String(Math.round(t.duration_ms))}, ` + `duration_basis: ${yamlScalar(t.duration_basis ?? null)},`],
@@ -665,6 +667,10 @@ var PLAN_SHAPE_RULES = [
665
667
  {
666
668
  issue: "#365",
667
669
  text: "**A story that adds an invariant over EXISTING data lands in the same story as, or in a wave after, the " + "story that makes existing rows and fixtures satisfy it.** A check constraint, a `NOT NULL`, a required " + "field or a validation rule written ahead of the story that populates what it enforces makes that earlier " + "story's own dod structurally red — every fixture the enforcing story's own tests run against still lacks " + "the value, and two stories in the SAME wave run in separate worktrees that never see each other's writes, " + "so 'later wave' is the only way one story's data can satisfy another's rule. If the invariant must land " + "first, write it NON-ENFORCING — nullable, consistency-only, no constraint — and let the later story " + "tighten it once the data exists. The `plan` check refuses a story whose acceptance or test plan names an " + "enforcement over a field a later story's acceptance or test plan names populating."
670
+ },
671
+ {
672
+ issue: "#376",
673
+ text: "**A story's own body never names an edit of a file its own `touches` forbids.** A step that tells the " + "developer to edit, add, write, create, rename, delete or otherwise change a backtick-quoted path — or to " + "add a bullet under a `## <v> — unreleased` CHANGELOG heading, which is an edit of `CHANGELOG.md` even when " + "the step never spells the filename — is a plan contradicting its own write allowlist: the developer who " + "obeys `touches` correctly leaves the file untouched, and a reviewer has to send the branch back for it. " + "The `plan` check refuses a story whose body names such an edit outside its own `touches`; a step that " + 'only reads or points at a path ("read `x` for the shape") is unaffected.'
668
674
  }
669
675
  ];
670
676