tldr-experts 0.19.0 → 0.21.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,129 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.21.0 — 2026-09-13
4
+
5
+ ### Fixed
6
+
7
+ - **The raise an operator is told to make now moves a ceiling, and a stage that cannot fund a
8
+ review stops buying one (#244, #289).** One function, one question — how much money is there,
9
+ and who decides when there is not enough — so the two were done together. The model, measured
10
+ case by case by the session running two live unattended runs (their measurement, not this
11
+ file's): the three money figures are three knobs with three different jobs, not three copies of
12
+ one number. The STAGE's `budget_usd` decides the price scale and therefore every per-story
13
+ developer ceiling and the reviewer's; the PHASE ceiling only takes part in the economy refusal
14
+ ("remaining work > what is left") and caps no spawn at all; `per_agent_max_usd` only trims a
15
+ ceiling from above. Their two controls: raising only the stage figure, 16.20 → 60, moved a
16
+ developer ceiling 5.97 → 22.11 on the next spawn, and raising `per_agent_max_usd` and the phase
17
+ ceiling *without* touching the stage moved the ceiling by nothing. So `tldrx budget raise
18
+ <phase> <usd>` — the command every refusal names — moved the one knob that caps nothing a
19
+ sub-agent is dispatched under: the operator raised, re-ran, died on the identical cap, and each
20
+ `reject --and-continue` bought exactly one more turn. It now takes `--stage <id>`, which adds
21
+ the same amount to that stage's own `budget_usd` in `run.yml` (additively, like the phase
22
+ ceiling; recorded on `budget.raised` as `stage`/`stage_budget_before`/`stage_budget_after`, and
23
+ an unknown stage id is a usage refusal that writes neither file), and a raise that names no
24
+ stage now SAYS in its own output that no spawn ceiling moved and names the flag that would
25
+ move one.
26
+ - **A reviewer the stage cannot fund is not spawned, and the refusal is not recorded as a verdict
27
+ (#289).** Measured on a live unattended run (`--budget 60`, `--gates none --questions none
28
+ --ship merge`): story S1's developer went DoD 3/3 green at $10.64, the reviewer right after it
29
+ was handed **$0.43**, died with `Reached maximum budget ($0.43)` before reading a line of the
30
+ diff, and the `verdict: error` that recorded its death parked S1 at `review` with both
31
+ dependent stories blocked behind it — the loop stopped after $17.23 of a $60 run.
32
+ `REVIEWER_FLOOR_USD` ($1.00) already existed and did not save it, because the floor YIELDS to
33
+ the stage remainder (`min(REVIEWER_FLOOR_USD, budget − spent)`): a nearly-exhausted stage buys a
34
+ turn that provably cannot finish. The fix is not a bigger floor. When the stage has less left
35
+ than a review costs the executor refuses BEFORE the spawn — no `agent.spawned`, no task row, not
36
+ a cent — and the story parks at `review` exactly where an unjudged story parks, with its diff
37
+ merged and its one attempt unspent. What it writes down is **not a verdict**: `verdict: "n-a"`
38
+ ("no reviewer ran for this story") beside a `reviewer_unfunded:` reason on the story's
39
+ `task.done`, its review log and the retro — the reviewer-side sibling of the `developerError`
40
+ and `budget_death:` records #271 and #277 shipped, and for the same reason. An agent that never
41
+ started formed no opinion, and `error` already means "it died mid-read"; writing a death nobody
42
+ chose as a review outcome is an audit record lying in the dangerous direction. The reason names
43
+ the lever that actually moves the ceiling — `tldrx budget raise <phase> <usd> --stage <id>`,
44
+ sized to the shortfall — rather than the phase ceiling that caps no spawn. The review stays
45
+ owed, so the next invocation's review-only path picks it up instead of re-running the developer,
46
+ and under an unchanged remainder it refuses again for free. `reviewerCap`'s arithmetic is
47
+ deliberately untouched: it is mirrored on the budget gate's hot path, and the brake's estimate
48
+ must keep reading the schedule the executor would have spent under. A HOST review is never
49
+ refused this way — it costs the stage nothing.
50
+
51
+ ## 0.20.0 — 2026-09-13
52
+
53
+ ### Added
54
+
55
+ - **`tldrx seed check <file|dir>` — a read-only seed validator, and `/tldrx-plan`, a second
56
+ managed skill that plans seeds a run can finish (#291).** Measured this week across four
57
+ unattended runs on two workspaces: the difference between a run that finishes alone and one
58
+ that needs four rescues is decided BEFORE `run new`, in the seed — and nothing in the
59
+ framework helped a person write one. `install --claude` installed exactly one skill, the
60
+ facilitator; the only authoring guidance was the hand-written grammar list in the seeds
61
+ guide, which said nothing about stories, `touches`, `dod` lines, dependencies, size or
62
+ boundaries; the seeds that ran cleanly were written by a subagent that had to build its own
63
+ validator to check them. `seed check` is that validator, in the CLI: the seed goes through the
64
+ SAME importer chain `run new --seed` runs (`collectSeeds` → `seedClaims` →
65
+ `renderSeedHandoff` → `validateHandoff`, into a temp dir that stands in for the run dir), and
66
+ then through the authoring rules each traceable to a run that died without it — bullets
67
+ under 200 characters, `[src:]` last on its line and resolving to a real line, `dod` lines
68
+ byte-equal to a `commands:` value with no shell separator (`unquotedShellSeparator`, the
69
+ gate's own reader), a `Recommended:` on every open question (#251), `touches:` +
70
+ `depends_on:` + a ```dod fence on every story, and no two stories sharing a touched file in
71
+ one wave (#286: one conflicted file became four). One `file:line rule — text` line per
72
+ finding, exit `1`; an `advisory:` line for more than 4 stories or 2 waves — the framework's
73
+ CURRENT limit, named as a patch for #286/#244/#280 so the number is never read as a design
74
+ preference — never moves the exit code. `--budget <usd>` prints the stage split `run new`
75
+ would write and the per-story developer and reviewer caps, off `planBudget` and `caps.ts`
76
+ as they are (#244/#289: a `--budget 60` feature run gives Build $10.80 per attempt, and
77
+ nobody was told). Nothing here is a second reading of a rule: every check calls the function
78
+ that enforces it at run time. The planning skill, `plugin/skills/tldrx-plan/SKILL.md`, is
79
+ installed by `tldrx install --claude` beside the facilitator under the same
80
+ `<!-- tldrx-managed -->` marker, the same foreign-file refusal (a `SKILL.md` without the
81
+ marker at EITHER path refuses the whole install, nothing written) and the same uninstall;
82
+ `--dry-run` lists it as its own row. It reads `workspace.yml` and the repo tree before
83
+ writing, splits work into runs sized to what the framework carries today, chains stories
84
+ that share a counted or snapshot file, names only declared tools (patch for #290 — the
85
+ planner checks by hand what Plan cannot yet detect; the cure names tool + subcommand, and a
86
+ `git` line is never a `commands:` slot), gives every question a `Recommended:`, derives the
87
+ budget with `seed check --budget`, and ends with the exact
88
+ `run new … --gates none --questions none --ship merge` line. Every rule in the skill and on
89
+ the guide page is marked either **craft** (holds for any version) or **patch for #N**
90
+ (delete when the issue closes) — and the skill CITES the guide for the grammar instead of
91
+ restating it, which `test/plan-skill.test.ts` pins the way `test/maintain-skill.test.ts`
92
+ pins the maintain skill. The seed guide's "Writing a seed by hand" gained the story rules
93
+ with the same markers; the unattended recipe, both documentation sites (EN and ES) and the
94
+ README quick start name the validator and the skill.
95
+
96
+ ### Fixed
97
+
98
+ - **A permission refusal the workspace's `commands:` could have granted now names the operator's
99
+ cure, on the FIRST dead developer instead of the fourth (#285).** Measured by a peer session on a
100
+ live unattended run: one story's acceptance criterion required the EF migration to be *generated by
101
+ `dotnet ef`* in a workspace declaring `dotnet build` / `dotnet test`; four developers died and two
102
+ reviewer rounds ran over ~3 h, and the fourth wrote, verbatim, "dotnet ef is not grantable,
103
+ confirmed absent from .tldrx/workspace.yml:1 commands and refused twice already... No code changes
104
+ made this round" — the agent knew why it could not proceed, said so in plain words, and the system
105
+ bought a fifth attempt anyway. The field workaround was three exact-string slots added to
106
+ `workspace.yml` by hand, which a person had to deduce from four failures. That edit is now the
107
+ recorded reason: #278's classifier gains the case it had no name for, so a non-git refusal the
108
+ declared commands do not grant is no longer `unknown` with nothing to say but names the fix an
109
+ OPERATOR can apply — add a `commands:` slot whose value is exactly `dotnet ef`. The slot names the
110
+ tool and its subcommand, never the story's own arguments (`sha256sum s1.txt` asks for `sha256sum`,
111
+ not for a slot with one file baked into it), and the grant it is checked against is the real one:
112
+ `bashGrantsFor` writes `Bash(<declared>)` and `Bash(<declared> *)`, so `dotnet ef …` is ungranted
113
+ where `dotnet build` is declared even though both start with `dotnet`. That one notion lives beside
114
+ `DEVELOPER_GIT_VERBS` in `build/developerGrants.ts` and every reader calls it (§7). A `git` line is
115
+ never blamed on `commands:` — its allowance is that verb list, so "add a slot" would be a false
116
+ cure — and a caller that does not pass the declared commands classifies exactly as before, so every
117
+ sentence #278 shipped stays byte-identical. The story still blocks after one attempt with the same
118
+ exit code: #261's sentence is still the base, byte for byte, and the cure is the half it was
119
+ missing. **A Plan-time reader of acceptance PROSE was built for this too and measured OUT of the
120
+ change** — the signal that caught the field case went blind on the issue's own workspace, and the
121
+ signal that caught that one fired on `user profile` and `dotnet handles retries` — so it is filed
122
+ with both measurements as #290 instead of shipped half-working. A check that sometimes stays quiet
123
+ where it matters and sometimes shouts where it does not is worse than no check. The Definition of
124
+ Done, which is structured and run for real, has refused a command no slot declares since the
125
+ 2026-08-29 audit and is untouched here.
126
+
3
127
  ## 0.19.0 — 2026-09-13
4
128
 
5
129
  ### Added
package/README.md CHANGED
@@ -18,9 +18,13 @@ npm i -g tldr-experts # installs `tldrx` (short) and `tldr-experts` (same bi
18
18
  cd your-project
19
19
  tldrx doctor # check the environment — it is the authority, not a list in a README
20
20
  tldrx init # detect repos, map the code, write .tldrx/, ask only the gaps
21
- tldrx install --claude # write the skill, hooks and status line into ./.claude/
21
+ tldrx install --claude # write the skills, hooks and status line into ./.claude/
22
22
  ```
23
23
 
24
+ Planning a feature for an unattended run? **`/tldrx-plan`** (the second skill `install --claude`
25
+ writes) turns "I want X" into seed files that pass `tldrx seed check`, and ends with the exact
26
+ `run new` line.
27
+
24
28
  Later: **`tldrx update`** pulls the newest published version and prints the CHANGELOG between the
25
29
  one you had and the one you now have. Any command will tell you, in one line, when there is a newer
26
30
  one — off the hot path, cached for a day, silent when it cannot reach the registry, and never in
@@ -331,6 +335,8 @@ back on the registry is 0.3.0.
331
335
 
332
336
  | Version | Date | Status | Contains |
333
337
  |---|---|---|---|
338
+ | 0.21.0 | 2026-09-13 | `beta` | `tldrx budget raise <phase> <usd> --stage <id>` now moves the stage's own `budget_usd` — the figure that actually sets a developer's and a reviewer's spawn ceiling — instead of the phase figure, which caps no spawn at all: measured on a live unattended run where raising the plan's per-story price and the phase ceiling moved a developer's cap by nothing, and only raising the stage figure moved it; a raise naming no `--stage` now says outright that it moved no spawn ceiling. And a reviewer a nearly-exhausted stage cannot fund is refused before it is spawned — the run that surfaced this handed a reviewer **$0.43**, which died before reading a line of diff and recorded `verdict: error`, parking the story with its dependents blocked; the refusal now costs $0 and records `verdict: n-a`, not an error the reviewer never formed (#244, #289). Minor release: a new flag, `--stage`, on `budget raise`. |
339
+ | 0.20.0 | 2026-09-13 | `beta` | `tldrx seed check <file\|dir>` runs a hand-written seed through the same importer chain `run new --seed` runs, plus the authoring rules four unattended runs paid to learn — bullets under 200 characters, `[src:]` citations that resolve, `dod` lines byte-equal to a declared command with no shell separator, `depends_on` wherever two stories touch one file, a `Recommended:` on every open question — and prints the budget split and per-story caps a run would set, creating no run and spending nothing; `tldrx install --claude` now writes a second managed skill, `tldrx-plan`, that reads `workspace.yml` and the repo tree and turns "I want X" into seeds that pass the same check, naming only declared tools (#291). And a permission refusal over a command the workspace's `commands:` could have granted now names the operator's cure on the first dead developer instead of the fourth, measured on a live unattended run where four developers died and two review rounds ran before the missing command turned out to be one `workspace.yml` line away (#285). Minor release: a new CLI verb, `tldrx seed check`. |
334
340
  | 0.19.0 | 2026-09-13 | `beta` | `tldrx story reopen <id> --as-is --note "…"` settles a story a person finished by hand from its branch as it stands, spawning no developer — measured on a live unattended run where a hand-rebased branch had a green DoD and a clean merge-tree but no way to land, blocked by #271's since-the-spawn rule from a developer with nothing to do (#279). Everything after the missing developer is the ordinary path: the epic merges in first if it moved, the same DoD and reviewer judge the diff, the same merge lands it. The record never lies about who did the work — `story.reopened` carries `reason: as_is` with the actor and note, `task.done` carries `as_is`/`as_is_by`/`as_is_note`, and the review log names the branch and the signer. Two refusals are structural with no flag to skip either: a branch with no commit ahead of its epic, and a red DoD. The signature clears the moment a real developer attempt starts, so a requeued story after a `changes` verdict records a developer, not an as-is. Minor release: a new flag, `--as-is`, on `story reopen`. |
335
341
  | 0.18.3 | 2026-09-13 | `beta` | four fixes to the unattended build path, all measured on live runs: a headless permission refusal is now classified — a chained shell command, a git verb outside the developer's allowance, or neither — so a chained line gets one retry with the cure stated first, an ungranted verb names its granted equivalent, and the developer prompt now lists the git verbs it holds from the one constant the grant is built from, after six refusals in one day left a ledger that recorded the symptom and not the cause; a wave no longer merges a story into its epic from the base it was cut from once a sibling has moved that epic — the story is measured against the epic tip immediately before merge, the epic is merged into the story and its Definition of Done re-run only when the epic actually moved, and a conflict blocks the story naming the files rather than merging half of one; a seed bullet over the claim cap no longer has its citation cut in half, with the clip moving back to the citation's start instead; and a failed commit count is no longer read as zero, so `--discard-pending` can no longer delete a plan just because git could not answer. |
336
342
  | 0.18.2 | 2026-09-13 | `beta` | three things measured on the first fully unattended run and on the live runs that followed it, one of them a regression 0.18.1 shipped hours earlier: `--ship merge` armed GitHub's auto-merge over a base that required nothing — the guard counted the checks the PR REPORTED while `--auto` waits on the base's REQUIREMENTS, so a proof run's pull request merged at 06:20:14 while all four of its checks finished between 06:21:38 and 06:24:56, and the record said `queued — GitHub merges when its checks pass`; ship now asks the base what it requires through rulesets and classic protection, names its three absences apart, and every unreadable answer arms nothing; the plan's per-story price had become a hard wall at 0.8x the planner's guess once 0.18.1 made those fields actually read, killing stories mid-flight at $1.28 and $3.25 against real costs an order of magnitude higher, so the ceiling is now `max(price x 3, $4.00)` derived at DISPATCH from the budget as it sits on disk — a run already in flight is covered without migrating anything — and, the half that carries the expensive end, a developer that dies on its cap with work in its tree commits it and lets the Definition of Done decide instead of parking the story, the same authority #271 gave a refused one; and the zero-touch recipe is written down where a person looks for it, seed file to backgrounded loop, on the README quick start and both documentation sites, with the honest boundary that this is the happy path without a person |
@@ -22,7 +22,7 @@ import {
22
22
  validateRunBudget,
23
23
  wouldExceed,
24
24
  wouldExceedHostTokens
25
- } from "./chunk-j5d0nard.js";
25
+ } from "./chunk-f5wkn522.js";
26
26
  import {
27
27
  EventLog
28
28
  } from "./chunk-6z5rmj0b.js";
@@ -8,7 +8,7 @@ import {
8
8
  spentBasis,
9
9
  tallyOf,
10
10
  validateRunBudget
11
- } from "./chunk-j5d0nard.js";
11
+ } from "./chunk-f5wkn522.js";
12
12
  import {
13
13
  EventLog,
14
14
  OUTCOME_NOT_RECORDED,
@@ -20,14 +20,14 @@ import {
20
20
  runSnapshot,
21
21
  statusWithOutcome,
22
22
  whatIsWaiting
23
- } from "./chunk-53wgfn01.js";
23
+ } from "./chunk-6qnsx1tz.js";
24
24
  import {
25
25
  expertsDir,
26
26
  loadExperts,
27
27
  pathsIntersect,
28
28
  readExpertDomain,
29
29
  stackExpertNames
30
- } from "./chunk-j5d0nard.js";
30
+ } from "./chunk-f5wkn522.js";
31
31
  import {
32
32
  isFinished
33
33
  } from "./chunk-6z5rmj0b.js";
@@ -2,8 +2,8 @@
2
2
  import {
3
3
  bar,
4
4
  runSnapshot
5
- } from "./chunk-53wgfn01.js";
6
- import"./chunk-j5d0nard.js";
5
+ } from "./chunk-6qnsx1tz.js";
6
+ import"./chunk-f5wkn522.js";
7
7
  import"./chunk-6z5rmj0b.js";
8
8
  import"./chunk-m1s8a6s0.js";
9
9
  import"./chunk-yre2scxn.js";