@letrquan/book 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (103) hide show
  1. package/CHANGELOG.md +2504 -0
  2. package/LICENSE +123 -0
  3. package/README.md +1383 -0
  4. package/dist/app-MOXXWEOY.js +18945 -0
  5. package/dist/app-MOXXWEOY.js.map +1 -0
  6. package/dist/chunk-24AXE6SP.js +12 -0
  7. package/dist/chunk-24AXE6SP.js.map +1 -0
  8. package/dist/chunk-3U5IBM24.js +213 -0
  9. package/dist/chunk-3U5IBM24.js.map +1 -0
  10. package/dist/chunk-574C6MQR.js +743 -0
  11. package/dist/chunk-574C6MQR.js.map +1 -0
  12. package/dist/chunk-5GDZ22YP.js +449 -0
  13. package/dist/chunk-5GDZ22YP.js.map +1 -0
  14. package/dist/chunk-5KLVH3PY.js +167 -0
  15. package/dist/chunk-5KLVH3PY.js.map +1 -0
  16. package/dist/chunk-5RDRAO4B.js +92 -0
  17. package/dist/chunk-5RDRAO4B.js.map +1 -0
  18. package/dist/chunk-7N4J557C.js +10 -0
  19. package/dist/chunk-7N4J557C.js.map +1 -0
  20. package/dist/chunk-A7JLBK2Y.js +3506 -0
  21. package/dist/chunk-A7JLBK2Y.js.map +1 -0
  22. package/dist/chunk-ANSJPYGI.js +84 -0
  23. package/dist/chunk-ANSJPYGI.js.map +1 -0
  24. package/dist/chunk-BTZZUCLI.js +264 -0
  25. package/dist/chunk-BTZZUCLI.js.map +1 -0
  26. package/dist/chunk-C7D4NJ2D.js +217 -0
  27. package/dist/chunk-C7D4NJ2D.js.map +1 -0
  28. package/dist/chunk-CAGOB2N4.js +208 -0
  29. package/dist/chunk-CAGOB2N4.js.map +1 -0
  30. package/dist/chunk-DGMQEI6Y.js +98 -0
  31. package/dist/chunk-DGMQEI6Y.js.map +1 -0
  32. package/dist/chunk-E75IPST5.js +289 -0
  33. package/dist/chunk-E75IPST5.js.map +1 -0
  34. package/dist/chunk-EYGD6X2T.js +180 -0
  35. package/dist/chunk-EYGD6X2T.js.map +1 -0
  36. package/dist/chunk-GTVB4HEM.js +765 -0
  37. package/dist/chunk-GTVB4HEM.js.map +1 -0
  38. package/dist/chunk-J7MIZ5CA.js +1907 -0
  39. package/dist/chunk-J7MIZ5CA.js.map +1 -0
  40. package/dist/chunk-KWHQ2XXL.js +626 -0
  41. package/dist/chunk-KWHQ2XXL.js.map +1 -0
  42. package/dist/chunk-L6LK2QIQ.js +657 -0
  43. package/dist/chunk-L6LK2QIQ.js.map +1 -0
  44. package/dist/chunk-MAQXCUR4.js +1007 -0
  45. package/dist/chunk-MAQXCUR4.js.map +1 -0
  46. package/dist/chunk-S4XL7HOM.js +402 -0
  47. package/dist/chunk-S4XL7HOM.js.map +1 -0
  48. package/dist/chunk-URLQAXTR.js +60 -0
  49. package/dist/chunk-URLQAXTR.js.map +1 -0
  50. package/dist/chunk-V4BOYR52.js +3504 -0
  51. package/dist/chunk-V4BOYR52.js.map +1 -0
  52. package/dist/chunk-VZQXX3WW.js +210 -0
  53. package/dist/chunk-VZQXX3WW.js.map +1 -0
  54. package/dist/chunk-WKBO5R4O.js +34 -0
  55. package/dist/chunk-WKBO5R4O.js.map +1 -0
  56. package/dist/chunk-XIFFTATU.js +15753 -0
  57. package/dist/chunk-XIFFTATU.js.map +1 -0
  58. package/dist/command-approvals-6Y57VCWZ.js +28 -0
  59. package/dist/command-approvals-6Y57VCWZ.js.map +1 -0
  60. package/dist/context-ZC5IFEAX.js +25 -0
  61. package/dist/context-ZC5IFEAX.js.map +1 -0
  62. package/dist/hook-approvals-CTUFKEJT.js +22 -0
  63. package/dist/hook-approvals-CTUFKEJT.js.map +1 -0
  64. package/dist/index.d.ts +1 -0
  65. package/dist/index.js +2582 -0
  66. package/dist/index.js.map +1 -0
  67. package/dist/interactive-assets-CCAAWVKH.js +31 -0
  68. package/dist/interactive-assets-CCAAWVKH.js.map +1 -0
  69. package/dist/job-runner.d.ts +2 -0
  70. package/dist/job-runner.js +235 -0
  71. package/dist/job-runner.js.map +1 -0
  72. package/dist/loader-4IDS2RIE.js +15 -0
  73. package/dist/loader-4IDS2RIE.js.map +1 -0
  74. package/dist/mcp-approvals-WNPH65BK.js +18 -0
  75. package/dist/mcp-approvals-WNPH65BK.js.map +1 -0
  76. package/dist/mcp-config-YIJIES2C.js +16 -0
  77. package/dist/mcp-config-YIJIES2C.js.map +1 -0
  78. package/dist/model-window-store-YFVNXCIP.js +27 -0
  79. package/dist/model-window-store-YFVNXCIP.js.map +1 -0
  80. package/dist/openai-compatible-RHZUBHV6.js +10 -0
  81. package/dist/openai-compatible-RHZUBHV6.js.map +1 -0
  82. package/dist/permission-approvals-76PVREFD.js +14 -0
  83. package/dist/permission-approvals-76PVREFD.js.map +1 -0
  84. package/dist/permissions-CHCKTFQQ.js +28 -0
  85. package/dist/permissions-CHCKTFQQ.js.map +1 -0
  86. package/dist/sandbox-CNVAQDA7.js +23 -0
  87. package/dist/sandbox-CNVAQDA7.js.map +1 -0
  88. package/dist/scrollback-LUKLIBVI.js +136 -0
  89. package/dist/scrollback-LUKLIBVI.js.map +1 -0
  90. package/dist/sdk.d.ts +4574 -0
  91. package/dist/sdk.js +665 -0
  92. package/dist/sdk.js.map +1 -0
  93. package/dist/settings-loader-M2QP5TX2.js +21 -0
  94. package/dist/settings-loader-M2QP5TX2.js.map +1 -0
  95. package/dist/settings-removed-JWWPT4WA.js +19 -0
  96. package/dist/settings-removed-JWWPT4WA.js.map +1 -0
  97. package/dist/shell-selection-V6QAYXPX.js +17 -0
  98. package/dist/shell-selection-V6QAYXPX.js.map +1 -0
  99. package/package.json +143 -0
  100. package/patches/ink+6.8.0.patch +13 -0
  101. package/scripts/apply-ink-patch.mjs +29 -0
  102. package/scripts/ink-patch.mjs +59 -0
  103. package/scripts/verify-ink-patch.mjs +11 -0
package/README.md ADDED
@@ -0,0 +1,1383 @@
1
+ # Book
2
+
3
+ AI coding agent CLI with rich terminal UI. A provider-agnostic alternative to Claude Code.
4
+
5
+ The implementation-backed status snapshot is [docs/current-state.md](./docs/current-state.md).
6
+ This repository is proprietary and is currently distributed from source/GitHub rather than npm.
7
+
8
+ ## Features
9
+
10
+ - **Interactive TUI** (Ink/React) plus **print mode** (`-p`) with `text` / `json` / `stream-json` output for CI.
11
+ - **Providers**: Anthropic Messages API (prompt caching, adaptive thinking) and any OpenAI-compatible endpoint, auto-detected from `baseUrl` / `--provider`. `--effort` reaches both, as `output_config.effort` and as `reasoning_effort`.
12
+ - **Project context**: walks the tree to load Codex-style `AGENTS.md` and Claude-style `CLAUDE.md` instructions (user-global → broad project → specific project → local/rules) into a fenced, trust-labeled block, alongside platform info and discovered skills, slash commands, and subagents. Content is split by how often it changes: a cached static prefix, an uncached suffix for activation-class policy, and a per-turn `<session-state>` block carrying date, git status, and mode on the newest user turn — so an edit or a mode toggle costs one turn of cache, not the whole conversation.
13
+ - **Auto-memory**: file-based store under `~/.book/projects/<project>/memory/` with a `MEMORY.md` index (first 200 lines auto-loaded). Four memory types (`user` / `feedback` / `project` / `reference`), YAML frontmatter, auto-capture on user corrections/confirmations, and an **approval flow** (`/memory inbox` → `/memory approve|discard`). Secret/unfit text is rejected before writing.
14
+ - **Sessions**: append-only JSONL persistence with automatic titles from the first prompt plus `--resume`, `--continue`, `--session-id`, `--name`, and `--fork-session`; in-TUI `/clear` / `/new` / `/reset`, `/resume`, reference-aware `/compact`, and Claude-style `/rewind` for conversation, code, or both. Compaction reduces provider context without deleting the scrollable transcript: recent turns stay exact, older evidence remains addressable by stable session references, remembered file facts are freshness-checked before reuse, and constraints you stated in your own words are carried verbatim in a host-owned ledger the summarizer can read but never rewrite (see "Carried constraints").
15
+ - **Tools**: a provider-neutral capability catalog keeps a practical core loaded and uses `ToolSearch` to activate up to five authorized git, web, session, skill, agent, notebook, or MCP definitions on the next model turn. File, shell, task, clarification, and plan tools stay immediately available when permitted. Existing names such as `Read`, `Bash`, and `AgentSpawn` remain stable.
16
+ - **Slash commands**: built-ins including `/jobs`, `/agents`, `/agent`, `/init`, `/model`, `/effort`, `/config`, `/permissions`, `/cost`, `/usage`, `/context`, `/memory`, `/diff`, `/export`, `/skills`, `/review`, `/security-review`, `/release-notes`, `/feedback`, `/compact`, `/rewind`, `/clear`, `/resume`, plus custom commands from `.book/commands/*.md`. Print mode resolves commands through the same registries: `/init`, `/security-review`, `/review`, and custom commands run headlessly, and the interactive-only ones fail loudly instead of reaching the model as text.
17
+ - **Permissions**: allow/ask/deny rule matching with six modes — `default`, `acceptEdits` (`accept-edits`), `plan`, `auto`, `dontAsk`, `bypassPermissions` — see `/permissions` or `--permission-mode`. At a tool prompt, `A` arms **Always allow** and presses again to widen the rule it will write (`Bash(npm run check)` → `Bash(npm run *)` → `Bash(npm *)`); the pattern is always shown before Enter commits it. The prompt shows the whole command, wrapped to the terminal, and for `Edit`/`MultiEdit`/`Write`/`ApplyPatch` the diff the call would make, computed against the file on disk before anything is written; `D` opens a cut the card had to make. `/permissions` lists the rules in force and removes the selected one with `x`.
18
+ - **Sandbox & hooks**: optional bubblewrap sandbox for Bash; lifecycle hooks (JSON-over-stdio) for `PreToolUse` / `PostToolUse` / session events. Project-declared hooks require one-time approval per workspace; review provider/MCP settings and custom-command substitutions before opening an untrusted workspace.
19
+ - **Verified managed agents**: adaptive model-directed routing, purpose-named runs, compact parent-facing results, live TUI monitoring, profile model overrides, read-only non-Git exploration, resumable isolated worktrees, strict capabilities, typed evidence, independent validation, and explicit patch application. Built-in `explorer`, `patcher`, and `validator` profiles can be overridden under `.book/agents/`.
20
+ - **MCP**: interoperable MCP tool client with stdio, Streamable HTTP, and legacy SSE transports;
21
+ interactive project-server approval, secret-safe diagnostics, dynamic tool discovery, and
22
+ server-scoped permissions.
23
+ - **CLI helpers**: `book doctor` (diagnose env/config), `book config` (get/set/list settings), `book trust` (approve or reject configuration a repository declared), and `book tool-stats` (measure tool use across sessions — fail counts, rates, durations). None of them require a working credential — they exist to help when the provider is not yet configured, so `book doctor` reports an unresolved key as a finding rather than failing on it. When a settings layer is what is broken, `book doctor` names the layer the failure first appears with, and `book doctor --no-settings` reports everything else with all of them skipped.
24
+
25
+ See [`docs/current-state.md`](./docs/current-state.md) for the verified product snapshot, [`MILESTONES.md`](./MILESTONES.md) for the current roadmap, and [`CHANGELOG.md`](./CHANGELOG.md) for release notes.
26
+
27
+ ## Installation
28
+
29
+ Requires **Node.js 22.13+**.
30
+
31
+ ```bash
32
+ npm install -g @letrquan/book
33
+ book --version
34
+ ```
35
+
36
+ npm 11 blocks install scripts by default, so Book's Ink patch — which fixes a cursor bug in Ink's
37
+ incremental renderer — is not applied on a fresh install. Book detects this and uses the full-frame
38
+ `safe` renderer instead, which is correct but redraws more. To get incremental rendering on macOS
39
+ and Linux, allow the script once:
40
+
41
+ ```bash
42
+ npm approve-scripts @letrquan/book # then reinstall, or run: npm rebuild @letrquan/book
43
+ ```
44
+
45
+ The command is `book`; the package is scoped because the unscoped npm name was
46
+ already taken. To work on Book itself:
47
+
48
+ ```bash
49
+ git clone https://github.com/letrquan/book.git
50
+ cd book
51
+ npm install
52
+ npm run build
53
+ npm link # makes your checkout's `book` the global one
54
+ ```
55
+
56
+ ## Quick Start
57
+
58
+ ```bash
59
+ # Interactive TUI mode
60
+ book
61
+
62
+ # Print mode (non-interactive)
63
+ book -p "What does this codebase do?"
64
+
65
+ # Headless JSON output
66
+ book -p "Refactor auth module" --output-format json
67
+
68
+ # Stream JSON (CI-friendly)
69
+ book -p "Run tests" --output-format stream-json
70
+
71
+ # Resume a previous session
72
+ book --resume <id-or-name>
73
+ book --continue # most recent session in current directory
74
+
75
+ # Diagnose setup / edit settings from the shell (these run without a configured credential)
76
+ book doctor
77
+ book doctor --no-settings # skip every settings layer, when one of them is what is broken
78
+ book config list
79
+ book config get permissions.deny
80
+ book config set permissions.allow '["Read(*)","Glob(*)","Grep(*)"]' # user-global by default
81
+ book config set --local permissions.allow '["Read(*)"]' # just this checkout
82
+ book config list --local # what this checkout overrides
83
+ book config unset --local permissions.allow # drop the override
84
+
85
+ # Manage MCP servers (JSON shape is compatible with the wider MCP ecosystem)
86
+ book mcp list
87
+ book mcp add github npx -- -y @modelcontextprotocol/server-github
88
+ book mcp add remote https://mcp.example.com/mcp --transport http --scope project \
89
+ --header 'Authorization=${GITHUB_TOKEN}'
90
+ book mcp get github
91
+ book mcp remove github
92
+
93
+ # Inspect and measure tool use recorded across sessions
94
+ book tool-stats
95
+ book tool-stats --json # machine-readable aggregate
96
+ book tool-stats --all # ignore the retention window
97
+ book tool-stats --since 7 # only the last 7 days
98
+ ```
99
+
100
+ ### Common flags
101
+
102
+ | Flag | Purpose |
103
+ | ------------------------------------- | ---------------------------------------------------------------------------------- |
104
+ | `-w, --workspace <path>` | Workspace root (default: cwd) |
105
+ | `-m, --model <model>` | Model override |
106
+ | `-p, --print [prompt]` | Non-interactive / CI mode |
107
+ | `--output-format <fmt>` | `text` \| `json` \| `stream-json` |
108
+ | `--input-format <fmt>` | `text` \| `stream-json` (print mode input) |
109
+ | `--permission-mode <mode>` | `default` \| `acceptEdits` \| `plan` \| `auto` \| `dontAsk` \| `bypassPermissions` |
110
+ | `--effort <level>` | Thinking effort: `low` \| `medium` \| `high` \| `xhigh` \| `max`; outranks `BOOK_EFFORT`, `settings.effort`, and model metadata |
111
+ | `--provider <type>` | `anthropic` \| `openai` \| `auto` |
112
+ | `--max-turns <n>` | Cap agent turns (print mode) |
113
+ | `--max-budget-usd <amount>` | Cap spend (print mode) |
114
+ | `--json-schema <schema>` | Structured JSON output (print mode) |
115
+ | `-r, --resume <id\|name>` | Resume a named/id session |
116
+ | `-c, --continue` | Resume most recent session here |
117
+ | `--session-id <uuid>` | Pin a session id |
118
+ | `-n, --name <name>` | Display name for the session |
119
+ | `--fork-session` | On resume, fork to a new session id |
120
+ | `--no-session-persistence` | Do not write the session to disk |
121
+ | `--settings <path>` / `--no-settings` | Ad-hoc settings file, or skip all layers |
122
+ | `--scrollback` | Terminal-native scrollback instead of full-screen TUI |
123
+ | `--agents <mode>` | `adaptive` (default) \| `manual` \| `off` |
124
+ | `--verbose` | Full turn-by-turn output in print mode |
125
+ | `--include-hook-events` | Include hook lifecycle events in stream-JSON output |
126
+ | `--include-partial-messages` | Include partial assistant text deltas in stream-JSON output |
127
+ | `--prompt-suggestions` | Ask for follow-up prompt suggestions after completion |
128
+
129
+ ### Print mode
130
+
131
+ `-p/--print` runs one or more prompts with no terminal attached, for CI and scripting. The prompt
132
+ comes from the flag or from stdin, so a long one need not be interpolated into argv:
133
+
134
+ ```sh
135
+ book -p "explain this repo"
136
+ book -p < prompt.txt
137
+ git diff | book -p # the diff is the prompt
138
+ ```
139
+
140
+ The flag wins when both are given. `--input-format stream-json` reads stdin as newline-delimited
141
+ `{type:'user', content}` records instead, which is how you submit more than one prompt to a single
142
+ process.
143
+
144
+ Three things behave differently in print mode, because there is nobody to ask.
145
+
146
+ **Slash commands.** A prompt beginning with `/name` is resolved through the same command
147
+ registries the TUI uses instead of being sent to the model as literal text. See
148
+ [Slash Commands](#slash-commands) for the supported subset.
149
+
150
+ **Plan mode.** `--permission-mode plan` still refuses mutations until a plan is approved, and print
151
+ mode now has a way to approve one. `bypassPermissions` approves automatically, as before. A host
152
+ that supplied `onUserQuestionRequired` is asked through that same handler: one question with
153
+ `Approve` and `Reject` options, where any other free-text answer is taken as revision feedback and
154
+ the agent submits a new plan. With no handler there is nobody to ask, so the run **stops at the
155
+ first plan** and returns the plan as its deliverable — it no longer auto-rejects and lets the model
156
+ re-plan until `--max-turns` is exhausted. In `text` output the plan is printed followed by a line
157
+ saying nothing was applied; in `json` and `stream-json` the result payload carries:
158
+
159
+ ```json
160
+ {
161
+ "plan": {
162
+ "status": "not_applied",
163
+ "reason": "approval_unavailable",
164
+ "plan": "the plan exactly as ExitPlanMode submitted it",
165
+ "message": "No changes were applied: …"
166
+ }
167
+ }
168
+ ```
169
+
170
+ `reason` is one of `approval_unavailable` (no handler), `approval_declined`, `approval_cancelled`,
171
+ or `invalid_approval_response`. The run's terminal outcome is `completed`/`normal_completion` and
172
+ the process **exits 0** — "finished, and deliberately changed nothing" is expressed by
173
+ `plan.status`, not by an exit code. Under `stream-json` the decision is also announced as
174
+ `{"type":"plan_approval","status":"stop"}`; `status` is one of `approve`, `approve-fresh`,
175
+ `reject`, `revise`, or `stop`. Queued `--input-format stream-json` prompts after a plan stop are
176
+ not run. SDK `query()` callers see the stop through the forwarded `tool_use` event (the full plan)
177
+ and its `tool_result` (`structuredError.code = "plan_approval_unavailable"`); the `plan` object is
178
+ not yet carried on the SDK `result` event.
179
+
180
+ **Exit codes.** Print mode exits 1 when the run throws — a slash command this host cannot perform,
181
+ a command invoked with a bad argument, or a failure inside a host-performed command such as
182
+ `/review`. Everything else exits 0.
183
+
184
+ ## Configuration
185
+
186
+ Settings are loaded in priority order (later wins):
187
+
188
+ 1. `~/.book/settings.json` (user-global)
189
+ 2. `.book/settings.json` (project)
190
+ 3. `.book/settings.local.json` (local, should be gitignored)
191
+ 4. `--settings <path>` CLI flag
192
+
193
+ #### Scopes
194
+
195
+ `book config set` and the TUI's `/config <key>=<value>` write the **user-global** layer
196
+ (`<BOOK_HOME>/settings.json`, normally `~/.book/settings.json`) unless told otherwise, so a
197
+ preference you set once applies in every checkout. Pass `--project` to write the checked-in
198
+ `.book/settings.json`, or `--local` to write the gitignored `.book/settings.local.json`.
199
+ `-g`/`--global` states the default explicitly; passing more than one scope is an error. Both
200
+ surfaces run the same guards through one shared write, so they cannot disagree about which file a
201
+ preference lands in.
202
+
203
+ A write is checked against the *merged* configuration, not just the file it lands in, so a value
204
+ that is valid on its own but would leave a configuration nothing can load is refused before it
205
+ lands rather than bricking every later command. A configuration that is already broken stays
206
+ writable, since repairing one is what the command is for.
207
+
208
+ `book config get` and `book config list` report the *resolved* merge of all layers by default.
209
+ Given a scope they read that one file verbatim instead, which is how you find the stray value
210
+ overriding you — the local layer resolves last, so anything left there outranks a later global
211
+ write. `book config unset <key>` removes a key from a scope (also user-global by default). A
212
+ user-global write that a workspace layer still shadows says so, and names the `unset` that clears
213
+ it.
214
+
215
+ Two groups of keys are refused in every scope. Trust decisions
216
+ (`mcp.projectServers`, `permissions.projectAllowRules`, `hooks.projectEntries`,
217
+ `commands.projectCommands`) live in `<BOOK_HOME>/trust.json` and are recorded with `book trust`.
218
+ The `shell` setting is not writable by `book config` in a workspace scope — it names the program
219
+ every `Bash` command is handed to, so edit the user-global file directly, pass `--settings`, or
220
+ use `BOOK_SHELL`.
221
+
222
+ #### Model ids and provider prefixes
223
+
224
+ A `model` written as `<provider>/<model>` is resolved through the `provider` registry: the prefix
225
+ selects the base URL, credential, and model catalog. The same spelling is also how many endpoints
226
+ name a single model (`meta-llama/llama-3-70b`), so a prefix that matches nothing is not an error —
227
+ Book passes the whole id through to the default endpoint.
228
+
229
+ That fallback is silent by design for the second form, but wrong for a typo. So when you have
230
+ configured providers and the prefix matches none of them, Book says so — on stderr at startup and
231
+ inline in `book doctor` — instead of leaving `Credentials: not resolved` as the only symptom of a
232
+ misspelled provider id.
233
+
234
+ Set `BOOK_HOME` to replace the default `~/.book` user-state root. This relocates user settings,
235
+ sessions, memory, managed-agent state and worktrees, jobs, rewind snapshots, telemetry, tool output,
236
+ learned context windows (`model-windows.json`), MCP configuration, and user-level skills, commands,
237
+ agents, and `AGENTS.md` discovery. Project-local `.book/` directories are unchanged.
238
+
239
+ #### Learned context windows
240
+
241
+ When a provider refuses a request for exceeding its context limit, Book records a ceiling for that
242
+ model in `<BOOK_HOME>/model-windows.json` so the next session sizes compaction against a number the
243
+ provider has actually shown it will not exceed. The value is a fraction of the refused size, never
244
+ the refused size itself, and it only ever ratchets **down** — a later refusal at a smaller size
245
+ lowers it, a larger one does not raise it. It is never read from the workspace, so a repository
246
+ cannot declare a ceiling for a clone.
247
+
248
+ A window you declare yourself always wins: set `contextWindow` on the model in settings and nothing
249
+ is learned or applied over it. `book doctor` lists every learned window with how long ago it was
250
+ learned, and `/context` and the status line mark which source the current window came from
251
+ (`declared`, `learned`, `family`, or `default`). To discard one, delete its entry from
252
+ `model-windows.json`, or delete the file to forget them all — it is rebuilt on demand.
253
+
254
+ ### MCP servers
255
+
256
+ MCP declarations use the interoperable `{"mcpServers": {"name": {...}}}` shape. User-global
257
+ servers live at `~/.book/mcp.json` (or `$BOOK_HOME/mcp.json`) and project declarations live at
258
+ `.mcp.json`. A declaration may use the legacy stdio shape or an explicit transport:
259
+
260
+ ```json
261
+ {
262
+ "mcpServers": {
263
+ "github": {
264
+ "type": "stdio",
265
+ "command": "npx",
266
+ "args": ["-y", "@modelcontextprotocol/server-github"],
267
+ "env": { "GITHUB_PERSONAL_ACCESS_TOKEN": "${GITHUB_TOKEN}" }
268
+ },
269
+ "remote": {
270
+ "type": "http",
271
+ "url": "https://mcp.example.com/mcp",
272
+ "headers": { "Authorization": "Bearer ${MCP_TOKEN}" }
273
+ }
274
+ }
275
+ }
276
+ ```
277
+
278
+ Supported transports are `stdio`, Streamable HTTP (`http`), and legacy SSE (`sse`). Variable
279
+ references support `${NAME}` and `${NAME:-fallback}`; values are never shell-evaluated. URLs must
280
+ be absolute HTTP(S) URLs without embedded credentials. Header names are shown in status output,
281
+ but header values are redacted from logs, prompts, reports, and diagnostics.
282
+
283
+ Project servers are untrusted repository-controlled input. Book displays the exact non-secret
284
+ target and asks for one-time approval before launching or connecting; the decision is stored in
285
+ `~/.book/trust.json` and is invalidated when any command, argument, environment value,
286
+ working directory, URL, transport, or header value changes. Headless and SDK runs skip unapproved
287
+ project servers. `/mcp` shows live status in the TUI; `book mcp list|get|add|remove` manages
288
+ declarations. Permission rules may target one server (`mcp__github`) or one exact tool
289
+ (`mcp__github__create_issue`).
290
+
291
+ Servers may ask the user for input mid-call through MCP form elicitation — a project picker, a
292
+ confirmation, a missing parameter. The interactive TUI answers those requests: the form shows which
293
+ server is asking, offers its fields (text, number, yes/no, and choice lists, which filter as you
294
+ type), and returns the answer inside the still-open tool call. `D` declines, `Esc` cancels, and
295
+ either way the server is told rather than left waiting. URL-mode elicitation is not supported and is
296
+ declined.
297
+
298
+ Only the TUI can prompt. Headless (`--print`) runs, and SDK runs without an `onElicit` callback, do
299
+ not declare the capability at all, so a server fails such a request itself instead of blocking on a
300
+ prompt nobody will see. For unattended runs, pass the value explicitly in the tool call or give the
301
+ server a default — for example the Azure DevOps server reads `ado_mcp_project` from its `env` block
302
+ and skips the project prompt entirely.
303
+
304
+ Legacy `.bookrc.json` is still supported but deprecated. Use `--no-settings` to skip all `settings.json` layers (defaults + legacy only).
305
+
306
+ Scalar values use the highest-priority layer. Permission rules, hook lists, and
307
+ `additionalDirectories` accumulate in layer order; directory entries are normalized and
308
+ deduplicated. Other arrays are replaced by the highest-priority layer that defines them.
309
+
310
+ Two exceptions apply to the **project** layer (`<workspace>/.book/settings.json`), because that
311
+ file is checked in and controlled by whoever wrote the repository:
312
+
313
+ - A `permissions.allow` rule it declares is withheld until you approve it. `ask` and `deny` rules
314
+ apply immediately — they only ever restrict. `book doctor` lists withheld rules and prints the
315
+ `book trust rule ...` command that grants them.
316
+ - Keys recording a trust decision — `mcp.projectServers`, `permissions.projectAllowRules`,
317
+ `hooks.projectEntries`, and `commands.projectCommands` — are ignored from **both** workspace
318
+ layers, so a repository cannot approve itself. They are not settings you write: decisions live
319
+ in `~/.book/trust.json`, keyed by workspace path, and `book trust` is what records them.
320
+ `book config set` refuses these four paths outright rather than writing a value nothing reads.
321
+ Putting them in the gitignored `.book/settings.local.json` was not enough — `.gitignore` does
322
+ not stop a force-added file from reaching a clone, so a repository could ship approvals for the
323
+ hooks, servers, and commands it also shipped. A store outside the workspace is one nothing the
324
+ repository ships can reach. An unreadable store records no decisions, which withholds the gated
325
+ input rather than releasing it.
326
+
327
+ `book config set` and TUI preference changes validate the complete local document before writing.
328
+
329
+ Set `defaultMode` in user-global `~/.book/settings.json` to choose the permission mode used by
330
+ the TUI, print mode, scrollback, and SDK when no invocation-specific mode is supplied. The
331
+ `--permission-mode` CLI option and SDK `permissionMode` option override that default.
332
+ Project and local settings cannot select `bypassPermissions` as the startup default. Setting
333
+ `disableBypassPermissionsMode` to `true` also blocks explicit bypass requests and removes bypass
334
+ from the TUI mode cycle.
335
+ Writes use an atomic sibling-file replacement, and malformed or non-object
336
+ `.book/settings.local.json` files are never overwritten. The reported error includes the file and
337
+ invalid setting path; provider secrets are redacted.
338
+
339
+ Inside the TUI, `/config` opens a visual settings menu. Use it to change the main model, compact
340
+ strategy, compact model, effort, theme, memory auto-capture, startup fire, or the model assigned to
341
+ each managed-agent profile. Choosing a row opens that setting's picker and returns to the menu on
342
+ the same row when it closes, so one `/config` covers as many settings as you want to change.
343
+
344
+ `/config <key>=<value>` is the same command in typed form, and it runs the same guarded write as
345
+ `book config set`: it refuses a key nothing reads, checks the resulting merge before writing, and
346
+ takes the same `--global` / `--project` / `--local` flags (at most one, before or after a
347
+ `compact-model` keyword).
348
+
349
+ Eight settings are held by the running session rather than re-read from `settings` each turn —
350
+ `model`, `compactModel`, `effort`, `theme`, `defaultMode`, `ui.showThinking`,
351
+ `ui.startupAnimation` and `memory.autoSave`. Each is handed to the same code path its menu row
352
+ uses, so `/config model=…` switches the session exactly as `/model …` does rather than writing a
353
+ file this session will not re-read. They land in the layer that setting belongs to: user-global
354
+ for all of them except `theme`, which stays project-local because a theme name can come from a
355
+ project's `.book/themes`. Everything else defaults to the user-global layer.
356
+
357
+ Naming a scope that setting already uses is the same request as naming none. Naming a *different*
358
+ one is a request to write that one file, and the reply then says the change waits for the next
359
+ start — because that is what a file write on its own does — and warns when a later-resolved layer
360
+ still decides the value.
361
+ The startup fire plays only for a new, empty launch session and is skipped automatically for
362
+ screen-reader or reduced-motion mode. Press Esc to skip it.
363
+
364
+ TUI preference changes are saved by whose choice they are. Preferences about how Book behaves for
365
+ *you* — model, effort, compact model, permission default mode, provider registries and API keys,
366
+ thinking display, startup animation, memory auto-capture — are written to the user-global
367
+ `~/.book/settings.json` and follow you across projects. What is genuinely about *this* repository
368
+ stays in `.book/settings.local.json`: skill overrides, approved permission rules, per-profile agent
369
+ models, and the theme (a theme name can come from a project's `.book/themes`, so it would not
370
+ resolve elsewhere). Set `ui.startupAnimation` to `false` in `~/.book/settings.json` to disable the
371
+ effect everywhere.
372
+
373
+ ### File mutations
374
+
375
+ Book exposes the same mutation tools to every model — `ApplyPatch`, `Edit`, `MultiEdit`, and
376
+ `Write` — but the system prompt's recommended tool is **model-conditional**: GPT/Codex-family
377
+ models (trained on the V4A patch envelope) are steered to `ApplyPatch`, and every other model
378
+ (Claude, Qwen, GLM, Gemini, Grok, unknown) is steered to exact-replace `Edit`/`MultiEdit`, the
379
+ format with the best cross-model compliance in published evals. Override per model in settings:
380
+
381
+ ```json
382
+ { "provider": { "myrouter": { "models": { "qc/qwen3.7-max": { "editFormat": "replace" } } } } }
383
+ ```
384
+
385
+ `editFormat` accepts `patch`, `replace`, or `whole` (whole-file `Write`-first guidance).
386
+
387
+ `ApplyPatch` accepts a compact Codex-style envelope:
388
+
389
+ ```text
390
+ *** Begin Patch
391
+ *** Update File: src/example.ts
392
+ @@
393
+ const answer = 41
394
+ -return answer
395
+ +return answer + 1
396
+ *** End Patch
397
+ ```
398
+
399
+ Use `*** Add File: path` with `+`-prefixed lines for new files and `*** Delete File: path` for
400
+ deletions. Update hunks use exact, unique context; reread the affected range and regenerate the
401
+ hunk after a `patch_context_not_found` or `ambiguous_patch_context` error. Patches preserve an
402
+ existing file's LF/CRLF convention and UTF-8 BOM, validate all files before writing, verify the
403
+ post-state, and roll back earlier files if a later commit fails. Binary and mixed-line-ending
404
+ updates are rejected rather than guessed.
405
+
406
+ Mutation reliability guardrails, tuned for heterogeneous models:
407
+
408
+ - **Read-before-edit is enforced.** `Edit`/`MultiEdit` (and `Write` over an existing file) fail
409
+ with `file_not_observed` until the file has been Read or `@`-mentioned this session, and fail
410
+ with `stale_file_observation` when it changed since last observed.
411
+ - **Whitespace-tolerant recovery.** When an exact `oldString` match fails, Book tries two
412
+ deterministic relaxations — trailing-whitespace-insensitive and uniform-indent-shift — and
413
+ applies one only when it matches a single location; the result notes the tolerance used.
414
+ `replaceAll` always requires exact matches.
415
+ - **Cross-harness argument aliases.** Claude Code-style spellings (`file_path`, `old_string`,
416
+ `new_string`, `replace_all`, Grep `glob`/`-A`/`-B`/`-C`, ApplyPatch `input`) are normalized to
417
+ Book's canonical arguments before validation, and `invalid_arguments` errors list the allowed
418
+ argument names.
419
+ - **Retry-loop braking.** Repeating a call that already failed with identical arguments returns
420
+ escalated guidance instead of the same error; structured `Fix:` remediation lines are rendered
421
+ into the model-facing error text.
422
+ - **Reliability visibility.** `/usage` shows per-tool call and failure counters for the session,
423
+ and `npm run eval:edit` runs a deterministic ~25-task edit-reliability eval against the
424
+ configured model, writing a report to `.book/reports/`.
425
+
426
+ `npm run eval:compact` runs a real-provider paired benchmark for `/compact`. Every probe runs
427
+ against the original history and the compacted history, so the report can separate baseline model
428
+ errors from compaction regressions. The default `--suite smoke` keeps the original five low-cost
429
+ static-recall probes. `--suite standard` runs 11 probes spanning static recall, knowledge updates,
430
+ conflict resolution, temporal reasoning, multi-hop synthesis, and abstention. Evidence is placed
431
+ early, late, or across the synthetic history, which remains larger than 6,000 estimated tokens.
432
+
433
+ The V2 report includes per-model and per-category accuracy, retained baseline answers, regressions,
434
+ improvements, JSON/tool protocol failures, compression ratio, prompt and net-token savings,
435
+ measured cost savings when model pricing is known, and token/cost break-even estimates that include
436
+ the one-time compaction call. Reports are written to `.book/reports/compact-eval-v2-*.{json,md}`.
437
+
438
+ ```bash
439
+ npm run eval:compact -- --model 9router/qc/qwen3.7-max --suite smoke
440
+ npm run eval:compact -- --models 9router/ag/gemini-3.6-flash-high,9router/qc/qwen3.7-max --suite standard
441
+ npm run eval:compact -- --model 9router/qc/qwen3.7-max --suite standard --repeat 3 --include-no-history
442
+ ```
443
+
444
+ Use repeated `--model` flags or comma-separated `--models` for cross-model comparisons. `--repeat`
445
+ measures run-to-run stability, `--include-no-history` detects probes that can pass without evidence,
446
+ and `--probes <count>` or `--context-window <tokens>` constrains cost and context. Experiments can
447
+ set `--checkpoint-tokens <tokens>` to override the reducer output cap without changing the
448
+ production default; the JSON report records the requested cap, realized checkpoint size, and full
449
+ checkpoint. `--compact-effort low|medium|high|xhigh|max` overrides reasoning effort only for the
450
+ compaction call, making it possible to test a cheaper reducer while keeping probe models fixed.
451
+ `--compact-model <model>` routes only the compaction call through another configured model, so a
452
+ cheaper reducer or a higher-fidelity reducer can be evaluated independently from the probe model.
453
+ The benchmark requires configured provider credentials and is not part of CI.
454
+
455
+ `Write` remains appropriate for generated or intentional full-file replacement. The
456
+ `apply_patch` provider alias maps to `ApplyPatch`; legacy tools are not silently reinterpreted.
457
+
458
+ ### Example `.book/settings.json`
459
+
460
+ ```json
461
+ {
462
+ "model": "claude-opus-4-6",
463
+ "compactStrategy": "summary",
464
+ "compactModel": "9router/ag/gemini-3.6-flash-high",
465
+ "effort": "high",
466
+ "defaultMode": "default",
467
+ "permissions": {
468
+ "allow": ["Read(*)", "Glob(*)", "Grep(*)"],
469
+ "deny": ["Bash(rm *)", "Write(.env)"]
470
+ },
471
+ "sandbox": {
472
+ "enabled": false
473
+ },
474
+ "hooks": {
475
+ "PreToolUse": [{ "matcher": "Bash(*)", "command": "my-validator" }]
476
+ },
477
+ "memory": {
478
+ "enabled": true,
479
+ "autoSave": true,
480
+ "requireApproval": true
481
+ },
482
+ "ui": {
483
+ "showThinking": true,
484
+ "startupAnimation": true
485
+ },
486
+ "toolDiscovery": {
487
+ "mode": "auto",
488
+ "eagerToolCount": 10,
489
+ "schemaTokenBudget": 8000,
490
+ "maxLoadedTools": 15,
491
+ "searchLimit": 5
492
+ },
493
+ "toolExecution": {
494
+ "maxConcurrent": 4
495
+ },
496
+ "agents": {
497
+ "mode": "adaptive",
498
+ "maxConcurrent": 3,
499
+ "maxSpawned": 8,
500
+ "maxDepth": 1,
501
+ "persist": true,
502
+ "includeUntrackedInSnapshot": true,
503
+ "telemetry": true,
504
+ "retentionDays": 30,
505
+ "checks": {
506
+ "test": "npm test",
507
+ "typecheck": "npm run typecheck"
508
+ },
509
+ "checkTimeoutMs": 120000
510
+ }
511
+ }
512
+ ```
513
+
514
+ `agents.checkTimeoutMs` caps one `Check` run (default 120 s, maximum 2 h). A check that exceeds it
515
+ is killed and reported as `check_timed_out` — explicitly *not* as a failing check, so an agent does
516
+ not "fix" code that was passing. Raise it for repositories whose full suite runs longer than the
517
+ default; `npm test` here builds first and routinely does.
518
+
519
+ ### Unattended runs
520
+
521
+ By default a run ends at the first turn that produces no tool calls: one user message is the whole
522
+ run. For long autonomous work, `continuation` lets the loop keep going while the plan says there is
523
+ work left.
524
+
525
+ ```json
526
+ {
527
+ "continuation": {
528
+ "enabled": true,
529
+ "maxConsecutive": 50,
530
+ "noProgressLimit": 3,
531
+ "blockedToolTurnLimit": 3,
532
+ "planRefreshTurns": 25,
533
+ "maxWallClockMs": 0
534
+ },
535
+ "retry": { "streamReissueAttempts": 3, "outputCapContinuations": 10 }
536
+ }
537
+ ```
538
+
539
+ While a run is going, `<BOOK_HOME>/runs/<session-id>.json` says what it is doing: turn, elapsed,
540
+ spend, the current todo, the last tool, free disk, and — once it stops — the terminal outcome, or a
541
+ `crash` field if the process died without reaching one. It is rewritten at each turn boundary
542
+ (temp-file then rename, so a reader never sees a torn record) and stays a fixed size however long the
543
+ run lasts. This is the signal to watch for "is it stuck": the transcript's mtime advances identically
544
+ whether a run is working or wedged.
545
+
546
+ `--max-budget-usd` bounds the **objective**, not one process and not one prompt: spend is carried
547
+ across restarts and across submitted prompts, and enforced against *inclusive* cost, so work done by
548
+ managed agents and subagents counts against the same ceiling. A cap that cannot be evaluated fails
549
+ closed — a non-finite value is refused at startup rather than silently permitting everything.
550
+
551
+ Continuation never overrides an abort, an approved plan handoff, a spent budget, or a policy
552
+ refusal. `noProgressLimit` is the brake: when the todo list, the observed-file hashes, and the
553
+ tool-call count are all unchanged across that many boundaries, the run ends as `no_progress` instead
554
+ of spinning. Only tool calls that actually *ran* count toward that witness: a refused call is a
555
+ policy decision, not work, and counting it would move the one signal meant to prove nothing moved.
556
+
557
+ `blockedToolTurnLimit` is a second, independent brake, and it is enforced **even when `enabled` is
558
+ false**. It stops a run whose every tool call was refused on that many consecutive turns, ending it
559
+ as `all_tools_blocked` and naming the tools to unblock. It is separate because a refusal spin never
560
+ produces a tool-free turn, so the turn-end gate — and therefore every brake behind it — never fires:
561
+ a headless run in the default permission mode answers each prompt `deny` and would otherwise
562
+ re-issue refused calls until the budget ran out. Set it to `0` to disable. `planRefreshTurns` restates the open plan periodically, which also keeps compaction from
563
+ retaining an empty tail in a run that never stops on its own. Terminal outcomes gain
564
+ `objective_complete`, `continuation_limit`, `blocked_plan`, `no_progress`, and `all_tools_blocked`
565
+ so a supervisor can tell them apart — `blocked_plan` specifically means every remaining task is waiting on unfinished
566
+ work, which is a stall, not a success. A deliberate stop is also no longer indistinguishable from
567
+ success: handing an approved plan back reports `handoff_requested`, and a plan stop reports
568
+ `plan_stop` and carries the approver's own message. Both keep the `completed` status, because
569
+ neither is a failure.
570
+
571
+ Thinking models get their own stall ceiling, on both provider paths. `retry.streamStallTimeoutMs`
572
+ (20 s) is tuned for a chat, where that much silence means something broke; a long quiet stretch
573
+ before the first token from a reasoning model is the model working. When a request enables reasoning
574
+ Book uses `retry.thinkingStallTimeoutMs` instead (default 15 minutes,
575
+ `BOOK_THINKING_STALL_TIMEOUT_MS`). Raise it if a very high-effort run still reports `stream_stall`.
576
+
577
+ What counts as "enables reasoning" differs by path, because the two carry different evidence. On the
578
+ Anthropic path it is adaptive thinking, which is on by default for Opus and Sonnet at `high` effort.
579
+ On an OpenAI-compatible endpoint it is a request that sends `reasoning_effort`, or a model whose
580
+ `provider.<id>.models.<model>.effort` entry declares an effort range — an endpoint that buffers a
581
+ whole thinking block sends nothing at all until it is done, so the declaration is the only signal
582
+ available before the silence starts. A model with `effort: false` stays on the chat ceiling, and so
583
+ does a model with no catalog entry, since an unknown model is more likely a chat model than a
584
+ reasoning one.
585
+
586
+ `retry.streamReissueAttempts` re-sends a turn after a transport fault — a stalled stream, a dropped
587
+ socket — onto the history already committed. Set it to 0 to end the run on any stream error, as
588
+ earlier versions did. `retry.outputCapContinuations` is a separate allowance for continuing after the
589
+ provider's output limit, so a large generated file cannot drain the budget a real socket drop needs.
590
+
591
+ For a supervised loop, use `--session-id` (resume-or-create) rather than `--continue`, which selects
592
+ the most recently touched session in the directory and can be hijacked by an unrelated `book -p`
593
+ invocation:
594
+
595
+ ```sh
596
+ ID=$(uuidgen)
597
+ while ! book -p --session-id "$ID" --output-format stream-json --max-budget-usd 500 "$OBJECTIVE" | jq -e 'select(.type=="result") | .outcome.reason=="objective_complete"'; do sleep 30; done
598
+ ```
599
+
600
+ `--max-budget-usd` accumulates across restarts of the same session, so the cap bounds the objective
601
+ rather than each process.
602
+
603
+ Check on a run at any point, from any shell, with no provider configured:
604
+
605
+ ```sh
606
+ book status # newest session in this workspace
607
+ book status <id|name> --json
608
+ ```
609
+
610
+ It reports the byte-exact original objective, how much history has been compacted away, cumulative
611
+ tokens with an upper-bound USD figure, and the plan as last persisted.
612
+
613
+ It also reports **liveness**, which is the half the session transcript cannot answer. Book writes a
614
+ per-run record to `<BOOK_HOME>/runs/<session-id>.json` at every turn boundary; `book status` reads it
615
+ and leads with it:
616
+
617
+ ```
618
+ Run: finished — timed_out (stream_stall)
619
+ Stream stalled: no data received for 20000ms
620
+ turn 16 • elapsed 3m 44s
621
+ last update: 20m ago
622
+ last tool: Bash
623
+ in progress: run npm check
624
+ open todos: 3
625
+ spend: ~$2.4000 of $10.00 budget
626
+ free disk: 50.0 GiB
627
+ ```
628
+
629
+ The headline is one of four: `running` (the pid answers), `finished` with the terminal status and
630
+ reason, `crashed` when the process died without recording an outcome, or *no longer running, and
631
+ recorded no outcome* — the case that otherwise looks exactly like a run that had done nothing. A
632
+ live process that has not reached a turn boundary in fifteen minutes is called out as possibly
633
+ wedged, since a transcript's mtime advances at the same rate for a productive run and one stuck on a
634
+ permission prompt. `--json` carries the same fields under `run` for a supervisor script.
635
+
636
+ This matters most for `book -p`, which under the default `text` output format writes nothing at all
637
+ until it terminates.
638
+
639
+ To be told when something needs a person, configure a `Notification` hook. Only `severity: "alarm"`
640
+ is meant to wake anyone — everything else is a line to tail:
641
+
642
+ ```json
643
+ {
644
+ "hooks": {
645
+ "Notification": [{ "command": "[ \"$severity\" = alarm ] && ntfy publish my-topic \"$message\"" }]
646
+ }
647
+ }
648
+ ```
649
+
650
+ Managed agents refuse to spawn rather than fill the disk: `agents.maxWorktrees` (default 24) caps
651
+ simultaneous worktrees per repository and `agents.minFreeDiskBytes` (default 2 GB) is the free-space
652
+ floor. Both raise an `alarm` notification when they bite, and 0 disables either check.
653
+
654
+ `compactStrategy` supports only `summary`, the production default; it is the only strategy.
655
+
656
+ `compactModel` is optional. When set, manual and automatic `/compact` calls use that configured
657
+ provider/model only for checkpoint generation while normal agent turns continue on `model`. The
658
+ `BOOK_COMPACT_MODEL` environment variable overrides the setting. This is useful when a cheaper
659
+ reducer preserves the active model's accuracy; validate the pairing with `npm run eval:compact`
660
+ before making it a shared default. You can set it without editing JSON:
661
+ open `/config` and choose **Compact model** (shortcut `C`), or run
662
+ `/config compact-model 9router/ag/gemini-3.6-flash-high` (or `/config compactModel=...`). Both
663
+ reach the same place — the typed form is the menu row, not a separate write.
664
+
665
+ `toolDiscovery.mode` accepts `auto`, `eager`, or `deferred`. Auto mode sends all authorized definitions only when there are at most ten and their schemas fit the configured budget; otherwise the provider receives the practical core plus `ToolSearch`. Search never returns tools outside the current command, skill, agent-role, permission-mode, or runtime-state capability intersection.
666
+
667
+ Tool execution is serial by default. Consecutive calls explicitly reviewed as parallel-safe (`Read`, `Glob`, `Grep`, `GitStatus`, `GitDiff`, `GitLog`, and `GitBranch`) run as bounded ordered waves; every other call is a barrier. Preparation, hooks, mode checks, and permission prompts remain sequential, while wave results are published in provider order without discarding successful siblings when another fails. `toolExecution.maxConcurrent` sets the session-wide limit shared by the root and managed children (default `4`, maximum `8`).
668
+
669
+ ### Carried constraints
670
+
671
+ Compaction replaces older turns with a generated checkpoint, and everything in that checkpoint used
672
+ to be written by the summarizer model and re-fitted under budget pressure at every generation. The
673
+ fitter drops the oldest entries first, which in a coding session means the brief: Book's own
674
+ fidelity harness measured **zero** retention of the constraints a user opened the conversation
675
+ with, one generation in.
676
+
677
+ Book now splits authorship. Directive sentences from your own turns -- "the runtime must remain
678
+ Node 20", "never touch the vendored parser", "only use pnpm" -- are copied verbatim into a
679
+ host-owned **carried ledger** on the checkpoint. The summarizer sees it and is told to honour it,
680
+ but cannot write to it: a `carried` field in a model reply is discarded. The fitter cannot evict
681
+ from it. It accumulates across generations and is never reordered, and the checkpoint states the
682
+ rule for reading it: where two entries conflict, the later one wins.
683
+
684
+ The ledger is bounded rather than unlimited -- at most 32 entries, 1024 tokens, and 35% of the
685
+ checkpoint budget. When the cap binds it evicts restatements first, then softer steers ("prefer",
686
+ "avoid"), then explicit rules, never the newest entry, and the checkpoint discloses how many
687
+ entries were dropped. The exact turns stay retrievable with `SessionHistorySearch` and
688
+ `SessionHistoryRead`.
689
+
690
+ Two things are deliberately excluded. Only the text you typed is read -- never `@file` expansions
691
+ or `!`-shell output -- so a repository cannot plant a rule in a record Book is bound to keep. And
692
+ anything matching the secret detector is refused, because a ledger that never forgets is the last
693
+ place to write a credential.
694
+
695
+ Extraction is a cue-based scan, not a model call: it costs no extra tokens and no extra latency,
696
+ and it will miss a constraint phrased without a directive word. It is a floor under retention, not
697
+ a replacement for the summarizer's own record. The design is documented in
698
+ `plans/carried-ledger-plan.md`.
699
+
700
+ ### Permission rules and modes
701
+
702
+ `permissions.allow`, `permissions.ask`, and `permissions.deny` are matched against every tool call.
703
+ `deny` beats `ask`, and `ask` beats `allow`.
704
+
705
+ Permission *modes* decide whether you are prompted; they never relax a `deny` rule. A rule in
706
+ `permissions.deny` blocks the matching call in every mode, including `auto` and
707
+ `bypassPermissions`, and it is evaluated before the prompt, so a denied call never reaches one.
708
+ Modes differ only in what happens to calls that no `deny` rule matched:
709
+
710
+ | Mode | Unmatched calls |
711
+ | ------------------- | ------------------------------------------------------------------- |
712
+ | `default` | Checked against `allow`/`ask`, then prompted |
713
+ | `acceptEdits` | As `default`, but file mutations are approved without a prompt |
714
+ | `plan` | Read-only tools only; mutations are refused until you approve a plan |
715
+ | `auto` | Run without a prompt |
716
+ | `dontAsk` | Refused — the mode never prompts, and `allow` rules do not exempt a call |
717
+ | `bypassPermissions` | Run without a prompt |
718
+
719
+ `plan` mode needs a host that can approve the plan the agent submits through `ExitPlanMode`. The
720
+ TUI prompts; print/headless and the SDK route the decision through `onUserQuestionRequired`, and a
721
+ host that supplies no handler ends the run with the plan itself rather than rejecting it — see
722
+ [Print mode](#print-mode).
723
+
724
+ ### Hooks
725
+
726
+ `hooks.<event>` takes a list of `{ command, matcher? }` entries run in declaration order over a
727
+ JSON-over-stdio contract. Supported events:
728
+
729
+ | Event | Fires | Awaited | Can change the outcome |
730
+ | ------------------- | ---------------------------------- | ------- | ----------------------- |
731
+ | `SessionStart` | Session opened | yes¹ | no |
732
+ | `UserPromptSubmit` | Before a prompt is sent | yes | block, or rewrite it |
733
+ | `PreToolUse` | Before each tool call | yes | block the call |
734
+ | `PostToolUse` | After each tool call | yes | rewrite the tool output |
735
+ | `PreCompact` | Before compaction | yes | block compaction |
736
+ | `PostCompact` | After compaction | yes | no |
737
+ | `SubagentStart` | Before each managed-agent run | yes | no |
738
+ | `SubagentStop` | After each managed-agent run | yes | no |
739
+ | `Stop` | Once, after the agent stops | no | no |
740
+ | `SessionEnd` | Session left | yes¹ | no |
741
+
742
+ ¹ Awaited by the TUI and other multi-turn hosts; fire-and-forget on the one-shot SDK path.
743
+
744
+ **Awaited is the property that costs you latency**, and it is not the same as being able to veto.
745
+ A slow `PostToolUse` hook cannot block anything, but it still delays *every tool call* by up to its
746
+ runtime — hooks are capped at 10 s each and run sequentially in declaration order. Only
747
+ `UserPromptSubmit`, `PreToolUse`, and `PreCompact` can refuse the operation outright.
748
+
749
+ `matcher` filters `PreToolUse`/`PostToolUse` by tool call (`Bash(*)`) and `PreCompact`/`PostCompact`
750
+ by trigger.
751
+
752
+ Hooks from your own layers (`~/.book/settings.json`, `.book/settings.local.json`, `--settings`)
753
+ run as written. A hook declared in a repository's checked-in `.book/settings.json` is withheld
754
+ until you approve it once per workspace: the decision is recorded in `~/.book/trust.json`, outside
755
+ the workspace, keyed by a fingerprint of the event, matcher, command, and env — edit any of those
756
+ and the hook asks again. Nothing the repository ships can write that store, so a clone cannot
757
+ approve its own hooks. Non-interactive runs skip unapproved hooks with a warning.
758
+
759
+ `book doctor` lists each withheld hook with everything the fingerprint covers — command, matcher,
760
+ and environment — because approval covers all of them: `npm test` carrying
761
+ `NODE_OPTIONS=--require ./payload.js` is not the `npm test` it looks like. Record the decision with
762
+
763
+ ```bash
764
+ book trust hook <fingerprint> # or --all-pending for every withheld hook
765
+ book trust hook <fingerprint> --reject # refuse it, and stop being re-offered it
766
+ book trust rule "Bash(npm run *)" # the same, for a project-declared allow rule
767
+ book trust command deploy # and for a project command that substitutes shell
768
+ ```
769
+
770
+ All three take `--workspace <path>`; `book doctor` prints it for you when it is diagnosing a
771
+ directory other than the one you are in. Each invocation records one decision and leaves every
772
+ other decision — in this workspace and in every other — untouched.
773
+
774
+ `Stop` fires once when the agent stops, not once per provider turn — a task that takes twelve
775
+ tool-call turns still fires it once. It fires on cancellation too, which is usually the point of
776
+ having one. Subagents do not fire it: `Task` and managed agents run the same loop with your hook
777
+ config, and managed agents already report through `SubagentStop`. Like `SessionEnd`, `Stop` is
778
+ skipped when a run ends early through a blocked prompt, context overflow, an exhausted run budget,
779
+ or a provider stream error.
780
+
781
+ ### Which shell `Bash` runs
782
+
783
+ Book picks one shell per session and tells the model which one it got, so the syntax the model
784
+ writes matches the interpreter that will parse it. On macOS and Linux that is the platform default,
785
+ `/bin/sh`. On Windows, Book resolves in this order and stops at the first hit:
786
+
787
+ 1. `BOOK_SHELL`, then the `shell` setting — a name (`bash`, `pwsh`, `powershell`, `cmd`, `sh`) or a
788
+ path to an executable. A request that cannot be found is reported by `book doctor` and the
789
+ automatic order continues, rather than failing every command.
790
+ 2. **Git Bash**, when Book was launched from one (`MSYSTEM` or a POSIX `SHELL` is set) — so the
791
+ shell you see in your own terminal is the shell the model writes for.
792
+ 3. **PowerShell 7** (`pwsh`), then **Windows PowerShell 5.1**.
793
+ 4. Git Bash if it is merely installed.
794
+ 5. `cmd.exe`, only when nothing else exists.
795
+
796
+ `shell` is honoured from `~/.book/settings.json`, an explicit `--settings` document, or the
797
+ environment only. A workspace file cannot set it and `book config set shell` refuses those scopes:
798
+ it names the program every command is handed to, so a repository that could set it would run a
799
+ binary it ships on your first command. `book doctor` prints the resolved shell and why it was
800
+ chosen.
801
+
802
+ PowerShell is driven with `-EncodedCommand`, because 5.1 re-parses a `-Command` argument and
803
+ silently strips embedded double quotes. Under 5.1, Book also merges the error stream and renders
804
+ each record as text: left alone, that shell serializes a redirected stderr as a CLIXML document, so
805
+ a failing `Get-Item` handed the model XML instead of `Cannot find path`. Exit codes follow the last
806
+ statement, as in bash.
807
+
808
+ ### Shell command timeouts
809
+
810
+ A foreground `Bash` command is killed after **300000 ms** (five minutes) by default. The model can
811
+ raise that per call with the `timeout` argument, up to **600000 ms** — reach for it before a full
812
+ build or test suite rather than after the kill. It is validated like any other argument, so a value
813
+ outside the declared range is rejected rather than quietly ignored. Background commands ignore
814
+ `timeout` entirely and take `max_runtime_ms` instead.
815
+
816
+ `BOOK_TOOL_TIMEOUT_MS` overrides the default for every tool, `Bash` included, and where it is set it
817
+ is also the **ceiling** on what a single call may ask for: lowering it to 30000 caps a model that
818
+ asks for ten minutes, and a request above the limit in force is refused rather than quietly shrunk.
819
+ Raising it above 600000 raises the *default* — which needs no argument to reach — but not the
820
+ per-call reach, since the schema publishes and validates 600000 as the maximum. Precedence is the
821
+ call's `timeout` (bounded by that ceiling), then a deliberate per-tool setting such as
822
+ `agents.checkTimeoutMs`, then `BOOK_TOOL_TIMEOUT_MS`, then the tool's default. No source can resolve
823
+ past 2147483647 ms (~24 days), the largest delay a timer can hold; beyond it Node silently fires
824
+ after 1 ms.
825
+
826
+ The same resolution governs every tool that enforces a deadline of its own — `Bash`, `Check`, and
827
+ `WebFetch`. Each declares its budget so the host's backstop outlasts it; when the two are equal the
828
+ backstop fires first and replaces the tool's report, including its output, with a bare timeout. A
829
+ `timeout` argument sets the host budget only for a tool that publishes one, so a stray value cannot
830
+ pull the backstop underneath a tool that times itself.
831
+
832
+ A killed command reports itself as killed rather than failed, and returns whatever it printed on
833
+ stdout and stderr before the kill. The two outcomes call for different next moves — retrying a
834
+ killed command identically is pointless; retrying with a larger `timeout` is not — so the failure
835
+ message names the deadline it hit and the ways past it.
836
+
837
+ ### Bash sandbox
838
+
839
+ `sandbox.enabled` runs `Bash` commands inside [bubblewrap](https://github.com/containers/bubblewrap), which must be installed and is Linux-oriented; the sandbox is unavailable on Windows. When it cannot be created, `sandbox.failIfUnavailable` decides whether the command fails or runs unsandboxed. `sandbox.excludedCommands` skips the sandbox for matching commands, and sandboxed output is prefixed with `[sandboxed]`.
840
+
841
+ The sandbox gives the command fresh PID/IPC/UTS namespaces, a private `/tmp`, read-only system directories, all capabilities dropped, and a lifetime tied to the spawning process. Commands are spawned as a direct argument vector — never as a shell string — so the command text is parsed only by the shell running *inside* the sandbox. Ordinary shell syntax (pipes, `&&`, redirection, substitution) works normally there.
842
+
843
+ **The workspace root is the only directory bound writable by default**, and it is bound regardless of the `workdir` argument. A sandboxed `Bash` call whose `workdir` falls outside the workspace is rejected rather than granted a wider mount; use `sandbox.filesystem.allowWrite` to add directories deliberately.
844
+
845
+ `sandbox.filesystem` adjusts the default mounts, applied after the workspace bind so they take precedence:
846
+
847
+ | Key | Effect |
848
+ | ------------ | ----------------------------------------------------------------------- |
849
+ | `allowWrite` | Bind the path writable inside the sandbox |
850
+ | `denyWrite` | Bind the path read-only |
851
+ | `denyRead` | Mask the path — an empty tmpfs for a directory, `/dev/null` for a file |
852
+
853
+ Entries may start with `~`. Paths that do not exist are skipped — bubblewrap rejects a bind with a missing source — and the skipped entries are reported once at startup and by `book doctor`, since a skipped rule is policy you might otherwise believe is active.
854
+
855
+ `sandbox.network` **fails closed**. Bubblewrap has no DNS or per-domain filtering, so it can only share or unshare the network as a whole. If `allowedDomains` or `deniedDomains` contains anything, Book cannot honour the rule as written and disables network access entirely for sandboxed commands, with a warning. Leave both empty to share the host network.
856
+
857
+ Two further keys decide what happens *around* that boundary. Both are consulted on every `Bash`
858
+ call, and both answer the same question — will this exact command really execute inside a
859
+ namespace? — from one shared predicate, so they can never disagree about a command:
860
+
861
+ | Key | Default | Effect |
862
+ | ---------------------------------- | ------- | --------------------------------------------------------------------------- |
863
+ | `sandbox.allowUnsandboxedCommands` | `true` | Set `false` to refuse any `Bash` command that would run outside the sandbox |
864
+ | `sandbox.autoAllowBashIfSandboxed` | `true` | Run a genuinely sandboxed `Bash` command without a permission prompt |
865
+
866
+ `allowUnsandboxedCommands: false` covers all three ways a command escapes: sandboxing is turned
867
+ off, the command matched an `excludedCommands` pattern, or bubblewrap is missing on this platform.
868
+ The refusal names the setting *and* the specific reason, so it is actionable rather than a bare
869
+ denial. Note that with `sandbox.enabled` at its default `false`, nothing is sandboxed and this key
870
+ therefore refuses **every** `Bash` command — it is meant to be paired with `sandbox.enabled: true`.
871
+ It does not require independent approval for an `excludedCommands` match; it only makes that bypass
872
+ refusable outright.
873
+
874
+ `autoAllowBashIfSandboxed` is deliberately the weakest thing in the permission stack. It replaces
875
+ only the *default* ask — the prompt Book raises when nothing else matched:
876
+
877
+ - `permissions.deny` is evaluated first and is never softened by it.
878
+ - An explicit `permissions.ask` rule still prompts.
879
+ - If you wrote **any** `permissions.deny` or `permissions.ask` rule at all, the default ask stays
880
+ and nothing is auto-allowed. A shell line is not a file path: one command can read a file, write
881
+ another, reach the network, and chain three more behind `&&`, so a glob only matches the shapes
882
+ you thought to write down. Configuring adjudication is read as "ask me about shell commands",
883
+ and being sandboxed does not overrule that.
884
+ - It applies only to `Bash`, and only to the exact command text that will be executed.
885
+ - It grants no ability in `plan` mode: `Bash` is not a read-only plan tool, and plan mode refuses
886
+ it independently of the permission verdict.
887
+ - It does nothing while `sandbox.enabled` is `false` (the default), because then no command is
888
+ sandboxed.
889
+
890
+ `book doctor` prints the number of `excludedCommands` patterns and the **effective** state of both
891
+ keys, reporting `autoAllowBashIfSandboxed: true` as *inert* whenever nothing can actually be
892
+ auto-allowed, so the reported policy never overstates the enforced one.
893
+
894
+ ### Tool-use telemetry
895
+
896
+ When `observability.toolTelemetry` is enabled (default), Book appends one JSON line per finalized tool call to `~/.book/telemetry/tool-use.jsonl` (override the directory with `BOOK_TOOL_TELEMETRY_DIR`). Each record captures the canonical tool, the final status the model saw, a derived `isFailure` flag (`error`/`timed_out` only — permission blocks, plan-mode blocks, user declines, and cancellations are never counted as failures), the error code, duration, retries, model, and subagent attribution. The write is best-effort and off the hot path; it never blocks or fails a session, and the active log is size-rotated into a single `.1` backup.
897
+
898
+ `book tool-stats` reads this log and reports, per tool, calls / failures / fail rate / p50 / p95 duration / retry rate, plus a per-model split and the most frequent error codes:
899
+
900
+ ```
901
+ $ book tool-stats
902
+ Tool use — 1,204 calls across 37 sessions (2026-06-30 → 2026-07-27)
903
+ 18 failed (1.5%)
904
+
905
+ TOOL CALLS FAIL FAIL% P50 P95 RETRY%
906
+ Bash 412 12 2.9% 120ms 2.1s 4.4%
907
+ ApplyPatch 210 4 1.9% 38ms 140ms 5.7%
908
+ Read 402 0 0.0% 15ms 34ms 0.0%
909
+ ```
910
+
911
+ Use `--json` for a machine-readable aggregate, `--since <days>` to change the window, `--all` for full history, and `--prune` to drop records older than the window from disk. `observability.toolTelemetryRetentionDays` sets the default reporting window and the `--prune` target; disk use is otherwise bounded by log rotation. Records store outcomes and hashes only, never prompts or file contents. This is separate from the ephemeral in-session counters shown by `/usage`.
912
+
913
+ ### Managed agents
914
+
915
+ Adaptive mode keeps targeted work inline and nudges the parent toward the read-only `explorer` profile after three successful root `Glob`/`Grep` queries. The reminder is advisory: the fourth lookup is still allowed. Broad exploration receives a purpose name such as `Trace authentication flow`; the reusable profile (`explorer`, `patcher`, or `validator`) remains separate. `--agents manual` keeps the same lifecycle tools but requires explicit user delegation; `--agents off` removes managed-agent tools and routing guidance.
916
+
917
+ Explorer and `reviewer` run in the parent workspace with a hard read-only capability boundary and do not require Git, snapshots, or worktrees. `reviewer` backs `/review`: it is restricted to `Read`, `Glob`, `Grep`, `GitStatus`, `GitLog`, and `GitBranch` — deliberately no diff tool, because the host supplies the review target. Because it is a trust boundary, a project agent definition named `reviewer` cannot replace its role, tools, isolation, or body; tune it through `agents.profiles.reviewer` instead. Such a definition is not silently discarded — `book doctor` reports which layer it came from and what to do about it. Patcher and validator runs retain synthetic Git snapshots and isolated worktrees under `~/.book/worktrees/<repo-hash>/<agent-id>`, with per-record state and transcripts under `~/.book/agents/<repo-hash>/records/`. `agents.maxConcurrent` controls active execution while `agents.maxSpawned` caps outstanding queued/running/waiting children; completed history does not consume the cap. Parent-facing lifecycle results contain compact summaries/evidence IDs, terminal handoffs preserve up to 50 KiB, and `AgentRead` retrieves larger results in bounded chunks. The TUI and SDK host can inspect detailed transcripts separately. A patcher commit cannot be applied until a distinct validator passes the exact candidate commit.
918
+
919
+ Child completion is delivered automatically to the correct parent session as a compact agent-update card and a persisted provider-facing notification; `AgentWait` is only an explicit synchronization barrier. Delivery is split into context-budgeted batches, retried with bounded backoff, and deduplicated by durable delivery ID before acknowledgement. Terminal rows freeze their duration and final preview, and lifecycle tool rows show semantic actions instead of serialized JSON prefixes.
920
+
921
+ Managed-agent state writes are atomic and coordinated per target. If Windows, an antivirus scanner, or another Book process temporarily holds a state file, running agents continue in memory while Book retries the newest pending state in the background. The TUI shows one non-modal storage warning and a short recovery notice. Operations that require durable setup before starting, including plans, snapshots, initial agent records, and evidence publication, fail cleanly with a retryable `agent_store_busy` error instead of leaving a partially started agent. Multiple Book processes may read the same repository store, but only the live owner may mutate an active agent.
922
+
923
+ Recoverable temp files and instance leases live beside managed-agent state under `~/.book/agents/<repo-hash>/`. Startup validates and promotes only the newest logical revision from an abandoned instance. Set `BOOK_DEBUG=1` to record safe `agent-store` lifecycle diagnostics such as degraded storage, retry recovery, stale-lock reclamation, and orphan-temp promotion; JSON payloads, prompts, transcripts, credentials, and environment values are not logged. See `docs/agent-store.md` for the storage and recovery policy.
924
+
925
+ Agent definition tool rules are strict capabilities. Missing or empty `tools` means no tools, while `*` explicitly inherits parent tools except recursive lifecycle, implicit user-question, and implicit MCP access. Argument rules such as `Bash(git status*)` are checked again at execution time. Built-in profiles use file/git tools and `Check`; arbitrary shell access requires an explicit custom-agent rule.
926
+
927
+ Running children and background shells appear in one flat job panel directly below the prompt, with `main` and each executable job listed at the same level. From an empty prompt, press Tab to cycle focus through `main` and the jobs. `/jobs` opens the panel for explicit management; `/tasks` remains an alias. Tab/Up/Down selects a row, Enter opens its transcript or output, `x` stops it, and Esc returns to the main prompt. Finished, failed, timed-out, and user-stopped jobs are removed from the active panel automatically; shell completion is preserved as a local notice and retained briefly for `BashOutput`/`DismissShell` compatibility.
928
+
929
+ `Bash` accepts `run_in_background: true` with an optional `title`, `max_runtime_ms`, `notify`, and `lifetime`. Session jobs are the default and end with Book. `lifetime: "persistent"` is explicit, receives a separate permission decision, and reattaches from repository-scoped state after Book restarts. `notify: "ui"` is the default, `"none"` suppresses completion delivery, and `"agent"` queues one bounded output tail for the parent model when it is idle. Persistent logs are bounded and are removed when the completed job is dismissed.
930
+
931
+ `/agents` explains where subagent definitions are configured rather than opening a second runtime dashboard. Import third-party Claude-style definitions with `/agents import <path>` to preview normalized tools and warnings, then `/agents import --confirm <path>` to install under `.book/agents/`. The lower-level `/agent <id>`, `/agent send <id> <message>`, `/agent stop <id>`, and `/agent apply <id> [evidence-id]` commands remain available for direct scripting and recovery.
932
+
933
+ Profile model precedence is invocation override, `agents.profiles.<name>.model`, definition frontmatter, then the parent model. `inherit` falls through rather than becoming a literal provider model. Stream-json and SDK hosts receive status, activity, question, permission, completion, and evidence events by default; high-volume child text deltas require `forwardSubagentText`.
934
+
935
+ > Snapshot privacy: non-ignored untracked files are written into the local Git object database so managed worktrees can reproduce the parent state. Ignore secrets and other sensitive local files before enabling agents. Dismissing or aging out an agent removes its managed worktree, branch, and orphaned snapshot ref. Agent telemetry stores metrics and hashes only, never prompts or file contents.
936
+
937
+ Book clears sessions and rotated debug-log backups after 30 days. Startup resolves and preserves the active session before cleanup, and the current debug log is never removed by age-based retention.
938
+
939
+ ### Themes
940
+
941
+ Use `/theme` to open the keyboard theme picker, or switch directly with `/theme apple`, `/theme dark`, `/theme light`, `/theme auto`, `/theme catppuccin`, `/theme nord`, `/theme gruvbox`, or `/theme solarized-dark`. The selection is applied immediately and saved to `.book/settings.local.json` for the next launch. A fresh install opens on `apple`; `auto` follows the terminal background and resolves to `apple` on a dark terminal and `light` on a light one. The built-in themes provide thoughtfully tuned palettes:
942
+
943
+ - **apple** (default): Near-black neutral surfaces, bright grey text, and one blue accent for the composer and your own turns. Every other hue is a status colour that appears only when a state needs attention.
944
+ - **dark / light**: Editorial warm charcoal / soft parchment with grounded sage and clay accents.
945
+ - **catppuccin**: Soothing medium-contrast pastel palette based on Catppuccin Mocha for minimal eye fatigue.
946
+ - **nord**: Arctic and glacial slate palette designed to reduce blue-light glare and harsh transitions.
947
+ - **gruvbox**: Warm retro-earthy dark palette with amber and olive tones for evening and low-strain coding.
948
+ - **solarized-dark**: Scientifically engineered Lab color-space palette with tuned luminance contrast.
949
+
950
+ Roles are kept visually distinct on purpose: sage/lavender/frost belongs to the agent, clay/blue/orange to product chrome and user input, teal/cyan to references and usage meters, and distinct status hues carry results. A custom theme that reuses one hue across roles will render those roles identically, which is what the built-ins avoid.
951
+
952
+ Project themes can override any token in `.book/themes/<name>.json`. They appear automatically in the picker and can also be activated with `/theme <name>`. Theme files are partial and inherit unspecified values from the editorial `dark` palette (not from `apple`), so existing custom themes render exactly as before:
953
+
954
+ ```json
955
+ {
956
+ "brand": "#AFC19D",
957
+ "userAccent": "#D3A17E",
958
+ "surface": "#20221D",
959
+ "surfaceActive": "#30362B",
960
+ "border": "#4B4D45",
961
+ "selectionText": "#F3EEE4",
962
+ "assistantAccent": "#AFC19D",
963
+ "toolRail": "#6B7164"
964
+ }
965
+ ```
966
+
967
+ ### Environment variables
968
+
969
+ | Variable | Purpose |
970
+ | ------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------- |
971
+ | `BOOK_API_KEY` | Default API key (or `{env:VAR}` in provider settings) |
972
+ | `BOOK_BASE_URL` | Default OpenAI-compatible base URL |
973
+ | `BOOK_MODEL` | Default model |
974
+ | `BOOK_PROVIDER` | `anthropic` \| `openai` \| `auto` |
975
+ | `BOOK_EFFORT` | Thinking effort level |
976
+ | `BOOK_HOME` | User-state root (default `~/.book`) |
977
+ | `BOOK_SHELL` | Shell for `Bash`: `bash`, `pwsh`, `powershell`, `cmd`, or a path |
978
+ | `BOOK_WORKSPACE` | Default workspace |
979
+ | `BOOK_MAX_TOKENS` / `BOOK_MAX_TURNS` | Generation / turn limits |
980
+ | `BOOK_COMPACT_MODEL` | Model used only for compaction checkpoints |
981
+ | `BOOK_RETRY_*` / `BOOK_REQUEST_TIMEOUT_MS` / `BOOK_STREAM_STALL_TIMEOUT_MS` / `BOOK_TOOL_RETRIES` | Retry and timeout tuning |
982
+ | `BOOK_TOOL_TIMEOUT_MS` / `BOOK_TOOL_TELEMETRY_DIR` | Tool timeout (`Bash` included) and telemetry location |
983
+ | `BOOK_WEB_ALLOW_HTTP` | Opt into plain HTTP for `WebFetch` (disabled by default) |
984
+ | `BOOK_WEB_ALLOW_PRIVATE_NETWORK` | Opt into local/private web destinations (disabled by default) |
985
+ | `BOOK_WEB_MAX_REDIRECTS` | Same-origin redirect limit for `WebFetch` (default 5, maximum 10) |
986
+ | `BOOK_TUI_RENDERER` | `safe`, `incremental`, or experimental scroll renderer |
987
+ | `BOOK_DEBUG` / `BOOK_DEBUG_UI` / `BOOK_DEBUG_RENDER` / `BOOK_DEBUG_FLOW` | Debug logging flags |
988
+ | `BOOK_DEBUG_FILE` / `BOOK_DEBUG_STDERR` / `BOOK_DEBUG_MAX_BYTES` / `BOOK_DEBUG_BACKUPS` | Debug log destination and rotation controls |
989
+
990
+ `WebFetch` requires HTTPS by default, validates DNS results and the address used by the network
991
+ connection, blocks private/special-use destinations, and stops on cross-origin redirects so the
992
+ new origin receives its own permission decision. It returns Markdown by default; `format` can be
993
+ `markdown`, `text`, or sanitized `html`. `WebSearch` works without configuration through the
994
+ built-in Exa MCP provider and accepts optional `limit`, `domains`, `recencyDays`, and `country`
995
+ hints. Its provider endpoint is built in and cannot be overridden through settings or environment
996
+ variables.
997
+
998
+ The TUI defaults to the full-frame `safe` renderer on Windows to avoid ConPTY footer corruption
999
+ during deep transcript scrolling. Other interactive terminals default to `incremental`. Set
1000
+ `BOOK_TUI_RENDERER=incremental` to opt into incremental rendering explicitly on Windows.
1001
+
1002
+ ## Slash Commands
1003
+
1004
+ Create custom slash commands by adding Markdown files to `.book/commands/`:
1005
+
1006
+ ```markdown
1007
+ ---
1008
+ description: Check for spelling errors
1009
+ ---
1010
+
1011
+ Run a spell check on the codebase and fix any issues found.
1012
+ ```
1013
+
1014
+ **Shell substitution needs approval when the command is checked in.** A command body can run
1015
+ shell and paste the output into the prompt — an inline ``!`git log --oneline -5` `` span, or a
1016
+ fenced ` ```! ` block. That happens before the model sees anything, and it runs outside the
1017
+ permission system and outside the sandbox: no rule is consulted, no sandbox applies, and nothing
1018
+ is asked. A `.book/commands/*.md` file is repository-controlled, so cloning a project and typing
1019
+ its command name would otherwise be enough to execute whatever that file says — including under
1020
+ `book -p`, where no terminal is present to notice.
1021
+
1022
+ Book therefore requires a one-time decision per project command that substitutes shell. It is
1023
+ recorded in `~/.book/trust.json`, keyed by workspace path — outside the working tree, like the
1024
+ `.mcp.json`, `permissions.allow`, and hook decisions. Nothing a repository ships can reach it:
1025
+
1026
+ ```bash
1027
+ book trust command deploy # or --all-pending for every command awaiting a decision
1028
+ book trust command deploy --reject # refuse it
1029
+ ```
1030
+
1031
+ `book doctor` lists which project commands are approved, rejected, or still refused, and prints
1032
+ the command that decides the pending ones. Until a decision exists the command is refused — in
1033
+ the TUI and in print mode alike — naming the shell it wanted to run.
1034
+
1035
+ The decision is keyed by command name but *validated* by fingerprint: a name is the handle you
1036
+ already have for `/deploy`, and the fingerprint recorded alongside it is re-checked on every
1037
+ invocation, so editing what the body runs re-asks under the same name. That fingerprint covers
1038
+ the shell the body runs, in order, not the prose around it: rewording the instructions does not
1039
+ ask again. Commands in `~/.book/commands/` are yours and are never gated, and a project command
1040
+ that substitutes no shell has nothing to approve.
1041
+
1042
+ Built-ins include session controls (`/clear`, `/resume`, `/compact`, `/rewind`, `/exit`,
1043
+ `/help`), task and job controls (`/task`, `/jobs`, with `/tasks` as an alias), managed-agent
1044
+ controls (`/agents`, `/agent`), config (`/model`, `/providers`,
1045
+ `/effort [low|medium|high|xhigh|max]`, `/config`, `/permissions`, `/theme`), inspection
1046
+ (`/status`, `/mcp`, `/cost`, `/usage` with `/stats` as an alias, `/context`, `/diff`, `/skills`,
1047
+ `/memory`), local output and reload (`/export`, `/reload-skills`), release/support
1048
+ (`/release-notes`, `/feedback`), agent prompts (`/init`, `/security-review`), and code review
1049
+ (`/review`, see below).
1050
+ `/model` switches models, while `/providers` opens the same picker for provider management. BYOK
1051
+ providers you add - their credentials, model catalog, and active model selection - are saved to the
1052
+ user-global `~/.book/settings.json` so they are shared across projects; such providers are labeled
1053
+ `[BYOK]`, and selecting one of their models and pressing `Alt+D` removes it. `/effort` opens a
1054
+ picker when called without an argument and saves successful selections to
1055
+ `.book/settings.local.json`.
1056
+
1057
+ After the base URL and API key, the add-provider wizard asks where the model list should come
1058
+ from: **discover automatically** (Book calls the endpoint's model-list API and you pick from the
1059
+ result) or **enter model IDs manually** (comma-separate to add several at once). Manual entry is
1060
+ the answer for an endpoint that exposes no model-list API, and it is still offered as a fallback
1061
+ if discovery fails.
1062
+
1063
+ An already-configured provider keeps both routes. With one of its models selected in the picker,
1064
+ `Alt+R` re-reads the catalog from the endpoint and `Alt+M` adds model IDs by hand. A refresh
1065
+ replaces what discovery previously returned, but hand-entered models survive it — they exist
1066
+ precisely because the endpoint does not list them, and are recorded as `"manual": true` in
1067
+ settings. Neither action changes the active model or touches the provider's stored credentials,
1068
+ and the highlighted model stays highlighted when the list re-sorts underneath it. Both are offered
1069
+ only for the `[BYOK]` providers you added, on the same ownership rule as `Alt+D`: catalog edits are
1070
+ written to `~/.book/settings.json`, so applying one to a provider inherited from a project layer
1071
+ would copy that provider's credential into a second file.
1072
+
1073
+ **Slash commands in print mode.** `book -p "/name args"` resolves the command through the same
1074
+ registries, the same `$1..$9` / named-argument / `${BOOK_*}` variable / shell substitution, and the
1075
+ same `allowed-tools` and `model` frontmatter enforcement as the TUI — it is never forwarded to the
1076
+ model as literal text. What differs is only what a host with no interactive surface is allowed to
1077
+ do with the result:
1078
+
1079
+ | Command | In print mode |
1080
+ | ---------------------------------------------------------- | -------------------------------------------------------- |
1081
+ | `/init`, `/security-review`, any `.book/commands/*.md` | Run as the prompt for that turn |
1082
+ | `/review` (and `/review --help`) | Performed by the host itself; see [Code review](#code-review) |
1083
+ | Everything else — session controls, pickers, panels, `/config`, `/export`, `/memory` | Refused with an error listing what *is* supported, and exit code 1 |
1084
+
1085
+ Refusal happens *before* the command's own code runs, so a command with a side effect (`/config`
1086
+ writes `settings.local.json`, `/export` writes a file, `/memory approve` mutates memory) can never
1087
+ half-fire in a host that cannot show its result. A `/name` that is not a command at all is still
1088
+ forwarded to the model verbatim, so an ordinary prompt like `book -p "/etc/hosts is a file"` is
1089
+ unaffected. A `.book/commands/*.md` command whose shell substitution has not been approved is
1090
+ refused the same way and for the same reason: this host cannot ask for the decision.
1091
+
1092
+ A command the host performed itself produces no model turn. Under `text` its output is written to
1093
+ stdout; under `stream-json` it is announced as
1094
+ `{"type":"command_result","command":…,"output":…,"data":…}`; and for `json`, `stream-json`, and the
1095
+ SDK it is also carried on the result payload as `commandResults`. `output` is the human rendering
1096
+ and `data` is the command's machine contract. `--output-format json` therefore remains a single
1097
+ top-level JSON document.
1098
+
1099
+ Expansion covers both print-mode input paths (`--print "…"` and `--input-format stream-json`
1100
+ stdin) and can be turned off with `expandSlashCommands: false` on `HeadlessOptions`, which forwards
1101
+ every prompt verbatim — appropriate for a host relaying untrusted end-user text. `query()` does not
1102
+ surface that option yet, so the SDK always expands.
1103
+
1104
+ `/skills` opens the interactive skill manager. Select a skill with `↑`/`↓`, press `Space` to cycle its visibility (`auto`, `name-only`, `manual`, or `off`), press `E` to cycle execution consent (`inherit`, `ask`, or `deny`), and press `Enter` to prepare an explicit `$skill-name` request. `G` toggles the global emergency switch, `R` reloads the catalog, and `/reload-skills` performs the same reload from the command line. Overrides are saved in `.book/settings.local.json` under `skills.overrides`, `skills.execution`, and `skills.enabled`.
1105
+
1106
+ ### Code review
1107
+
1108
+ `/review` is orchestrated by the host rather than run as an ordinary prompt. Book resolves the
1109
+ change once — base commit, changed files, and a unified diff, including untracked files — and hands
1110
+ that **immutable review target** to read-only `reviewer` agents. Reviewers never choose their own
1111
+ scope, so a review cannot silently widen or drift onto unrelated changes.
1112
+
1113
+ ```text
1114
+ /review Review the working tree (tracked + untracked changes)
1115
+ /review --base main Review against the merge base with a ref
1116
+ /review src/tools Restrict the review to a file or directory
1117
+ /review main...HEAD Review a committed range
1118
+ /review --deep Four parallel lenses + an independent verification pass
1119
+ /review --fix Deep review, then apply verified findings (implies --deep)
1120
+ /review --help Usage
1121
+ ```
1122
+
1123
+ **While it runs (TUI).** A review is minutes of work in background agents, so it reports before it
1124
+ starts: the resolved target — file count, base commit, path scope — and which passes are coming are
1125
+ printed before the first agent is spawned. Every reviewer, lens, verifier and patcher then appears
1126
+ in the job panel below the prompt and in the status line with live activity, so you can watch a
1127
+ pass or open one to read its transcript — Tab from an empty prompt selects a row, or `/jobs` opens
1128
+ the panel for explicit management. Those agents belong to the session for display only; they never
1129
+ deliver a completion notification, so watching a review costs no extra model turn. Press `Esc` to
1130
+ cancel — the in-flight agents are stopped, and a cancelled review reports `inconclusive` with no
1131
+ findings rather than presenting its own stopped passes as a result. `Ctrl+C` cancels the review too;
1132
+ a second press exits, as it does mid-stream. A cancelled `--fix` pass reports what it had already
1133
+ committed before stopping. Progress is a streaming-host feature: a print run has no silence to
1134
+ break, so its stdout stays exactly the report (the same target is on `data.target`).
1135
+
1136
+ A plain `/review` runs one structured pass. `--deep` fans out four specialized reviewers
1137
+ (correctness, security, simplification, efficiency), merges and deduplicates their findings, drops
1138
+ anything below 70% confidence, and then runs a **falsification pass**: an independent verifier tries
1139
+ to disprove each candidate against the real code. Rejected findings are dropped; findings the
1140
+ verifier could not reach stay `inconclusive` rather than being reported as real.
1141
+
1142
+ Coverage is explicit. If a reviewer fails, times out, or does not return the required JSON, the
1143
+ report says so and the verdict is capped at `inconclusive` — a review never reports "clean" from
1144
+ incomplete coverage. Output that fails the JSON contract is preserved verbatim in the report instead
1145
+ of being discarded.
1146
+
1147
+ `--fix` applies only verified findings, one at a time, through the patcher → validator pipeline: a
1148
+ patcher produces a patch candidate as evidence, a separate validator must approve that exact
1149
+ evidence id (agents cannot approve their own work), and only then is it applied.
1150
+
1151
+ Drop a **`REVIEW.md`** at the workspace root to calibrate reviews for your repository — severity
1152
+ conventions, known-noisy areas, project-specific rules. It is read fresh on every run and injected
1153
+ as calibration only: it cannot change the output contract, disable verification, or broaden the
1154
+ tools a reviewer may use.
1155
+
1156
+ **Outside the TUI.** `book -p /review`, `book -p "/review --deep"`, `book -p "/review --base main"`,
1157
+ path scopes and `<base>...<head>` all run headlessly through the identical pipeline — the host still
1158
+ resolves the review target and hands the reviewers an immutable diff. The report is written to
1159
+ stdout as text by default. `--fix` is interactive-only: a non-interactive host cannot approve a
1160
+ patcher's tool calls, so `book -p "/review --fix"` exits 1 with an explanation instead of editing
1161
+ and committing unattended. A review that could not run at all — a bad ref, `agents.mode = off`, an
1162
+ unknown option — also exits 1. An inconclusive *verdict* does not: the review ran, so gate on the
1163
+ verdict field rather than on the exit code.
1164
+
1165
+ Under `--output-format json` and `stream-json` the review is emitted as one record, with `data`
1166
+ holding a stable projection of the pipeline's own types:
1167
+
1168
+ ```json
1169
+ {
1170
+ "type": "command_result",
1171
+ "command": "review",
1172
+ "output": "the text report",
1173
+ "data": {
1174
+ "verdict": "blocking | recommend | clean | inconclusive",
1175
+ "target": {
1176
+ "kind": "working-tree | committed-range",
1177
+ "baseSha": "…",
1178
+ "headSha": "…",
1179
+ "path": "src/tools",
1180
+ "changedFiles": ["src/tools/shell.ts"]
1181
+ },
1182
+ "findings": [
1183
+ {
1184
+ "id": "…",
1185
+ "severity": "critical | major | minor | nit",
1186
+ "category": "correctness | security | simplification | efficiency | conventions | tests",
1187
+ "file": "src/tools/shell.ts",
1188
+ "line": 110,
1189
+ "summary": "…",
1190
+ "evidence": "…",
1191
+ "failure": "…",
1192
+ "suggestedFix": "…",
1193
+ "confidence": 85,
1194
+ "verification": "confirmed | rejected | inconclusive",
1195
+ "verificationReason": "…"
1196
+ }
1197
+ ],
1198
+ "coverage": {
1199
+ "reviewers": [{ "id": "correctness", "status": "completed", "findings": 2 }],
1200
+ "verifier": { "id": "verification", "status": "completed", "findings": 2 }
1201
+ }
1202
+ }
1203
+ }
1204
+ ```
1205
+
1206
+ `findings` are `ReviewFinding` values verbatim and `coverage` is the pipeline's own
1207
+ `ReviewCoverage`, so the report and the JSON can never describe different runs. The unified diff is
1208
+ deliberately omitted — it is the caller's own input and can be megabytes. On the `stream-json` wire
1209
+ the record is emitted as it completes; under `--output-format json` it arrives inside the single
1210
+ result document as `result.commandResults[]`, which keeps that format one top-level JSON object.
1211
+ `stream-json` consumers also see the review's managed-agent events (`agent_start`, `agent_update`,
1212
+ `agent_result`) as they happen; text mode prints nothing until the review finishes, which for
1213
+ `--deep` means two sequential phases under the fixed 10-minute per-pass timeout.
1214
+
1215
+ To score the pipeline against a golden set, pair expectations with reports captured from real runs
1216
+ and run `npm run eval:review -- <fixtures.json>`; it prints precision, recall, F1, usefulness rate,
1217
+ and signal-to-noise ratio. See `evals/review/fixtures.example.json` for the format.
1218
+
1219
+ ### Skills
1220
+
1221
+ Book reads interoperable directory packages whose entrypoint is `SKILL.md`:
1222
+
1223
+ ```text
1224
+ <root>/<skill-name>/
1225
+ SKILL.md
1226
+ references/ optional text references
1227
+ assets/ optional templates or other files
1228
+ scripts/ optional packaged scripts (never auto-executed)
1229
+ ```
1230
+
1231
+ `SKILL.md` must start with YAML frontmatter containing `name` and `description`. The body is loaded
1232
+ only after activation; metadata, validation issues, resource manifests, and digests are available
1233
+ for inspection without putting the body in the initial prompt. Use `references/` and `assets/` for
1234
+ supporting material; Book reads declared resources only through `ReadSkillResource`, as untrusted
1235
+ content. Scripts remain ordinary resources and can run only through Book's existing execution tools
1236
+ and their normal approvals.
1237
+
1238
+ Discovery scans these roots from lowest to highest precedence: user `~/.claude/skills`, user
1239
+ `~/.agents/skills`, user `~/.config/opencode/skills`, user `~/.book/skills`, then the matching
1240
+ `.claude/skills`, `.agents/skills`, `.opencode/skills`, and `.book/skills` directories from the Git
1241
+ root to the current working directory. Deeper project directories and native `.book` roots win;
1242
+ duplicate names are shadowed rather than merged and are shown in `/skills` diagnostics. Skill
1243
+ directories may be symlinked after canonical path and size checks; resource symlinks are rejected.
1244
+
1245
+ Visibility controls determine whether metadata participates in automatic matching: `auto` exposes
1246
+ name and description, `name-only` exposes only the name, `manual` requires explicit `$skill-name`,
1247
+ and `off` disables the skill. Project-sourced implicit activation requires consent. `ask` always
1248
+ requests consent, while `deny` fails closed; no skill can grant tools, bypass permissions, alter the
1249
+ sandbox, or execute a packaged script implicitly. Active instructions are scoped to the current run
1250
+ by default (or the next model step for `lifetime: turn`) and tool declarations are intersections
1251
+ with Book's existing authorized surface.
1252
+
1253
+ Newly discovered skills start in `manual` mode. After evaluating representative positive and
1254
+ negative prompts, enable automatic matching per skill from `/skills` or by setting its override to
1255
+ `auto`; this keeps implicit activation available without treating unmeasured skill descriptions as
1256
+ a safe release default.
1257
+
1258
+ Use `/skills status` for a body-free runtime report containing the catalog digest, active and
1259
+ previous activation frames, effective tool intersection, validation failures, prompt-catalog
1260
+ omissions, and recent lifecycle outcomes. The equivalent settings shape is:
1261
+
1262
+ ```json
1263
+ {
1264
+ "skills": {
1265
+ "enabled": true,
1266
+ "overrides": {
1267
+ "review": "auto",
1268
+ "deploy": "manual"
1269
+ },
1270
+ "execution": {
1271
+ "deploy": "ask"
1272
+ }
1273
+ }
1274
+ }
1275
+ ```
1276
+
1277
+ For portable packages, move an existing `.claude/skills/<name>/` or
1278
+ `.opencode/skills/<name>/` directory to `.agents/skills/<name>/` without changing its `SKILL.md`.
1279
+ Book continues to discover the compatibility locations, so migration can be gradual; use
1280
+ `.book/skills/<name>/` only when the package intentionally depends on Book-specific behavior.
1281
+
1282
+ Book watches skill roots and applies changes at the next safe run boundary. If an editor, network
1283
+ filesystem, or platform watcher misses an update, use `R` in the manager or `/reload-skills` and
1284
+ inspect `/skills status`; watcher errors are also shown in the manager. Reload clears lazy body
1285
+ caches, expires affected frames, refreshes the catalog digest, and invalidates agent context without
1286
+ rewriting an in-flight request.
1287
+
1288
+ Activation quality can be gated with
1289
+ `npm run eval:skills -- observations.jsonl [report.json] [report.md]`. The report measures precision,
1290
+ recall, false-activation cost, prompt/body token cost, activation latency, consent prompts, task
1291
+ completion, corrections, and skill-caused tool failures across direct, indirect, negative,
1292
+ ambiguous, conflicting, disabled, invalid, missing-body, and missing-resource cases. Reports retain
1293
+ prompt hashes and aggregate evidence rather than raw prompts, skill bodies, or resource contents.
1294
+ Run `npm run eval:skills -- --help` to print the command syntax.
1295
+
1296
+ `/rewind` first selects an active user prompt, then restores Conversation, Code, or Both to the state immediately before that prompt. Files are captured into local content-addressed snapshots under `~/.book/rewind/`; `--no-session-persistence` uses temporary storage that is removed on exit. `.git` and workspace-local `.book/` state are never captured by default, Git HEAD and the index are never moved, and Code/Both are disabled when HEAD drifted or a checkpoint exceeded its safety limits. Use `.book/rewindignore` to override the default exclusions for dependency, build, cache, coverage, and virtual-environment directories, or to explicitly opt selected `.book` paths back in. Other hidden, gitignored, and secret-like workspace files remain restorable; their contents stay in local blobs and are never written to session JSON or model context.
1297
+
1298
+ ## SDK Usage
1299
+
1300
+ ```typescript
1301
+ import { query } from 'book';
1302
+
1303
+ for await (const event of query('Explain this code', {
1304
+ workspace: process.cwd(),
1305
+ onUserQuestionRequired: async (request) => ({
1306
+ action: 'answer',
1307
+ answers: Object.fromEntries(
1308
+ request.questions.map((question) => [
1309
+ question.question,
1310
+ question.multiSelect ? [question.options[0].label] : question.options[0].label,
1311
+ ]),
1312
+ ),
1313
+ }),
1314
+ })) {
1315
+ if (event.type === 'text') process.stdout.write(event.content);
1316
+ if (event.type === 'tool_use') console.log('tool:', event.toolCall.name);
1317
+ if (event.type === 'result') console.log('usage:', event.usage);
1318
+ }
1319
+ ```
1320
+
1321
+ `AskUserQuestion` supports 1-4 questions, described single/multi-select choices, and free-text answers in the TUI. Print mode emits `user_question` / `user_question_result` stream events and declines deterministically when no callback is supplied. When a callback is supplied, plan approval is routed through it as an ordinary question and emits the same two events; either way the decision is announced as `plan_approval`, whose `status` is one of `approve`, `approve-fresh`, `reject`, `revise`, or `stop` — see [Print mode](#print-mode). A slash command the host performed itself rather than sending to the model emits `command_result` (`{type, command, output, data}`) and is carried on the `result` event as `commandResults`. Managed workers additionally emit `agent_start`, `agent_update`, `agent_result`, `agent_question`, `evidence_update`, and `agent_apply`. Background shells emit `background_job_start`, `background_job_update`, `background_job_output`, `background_job_result`, and `background_job_dismiss` through stream JSON and the SDK.
1322
+
1323
+ Auth and model selection come from settings / env (`BOOK_API_KEY`, `BOOK_MODEL`, and provider blocks), not from `query()` options. See `src/sdk.ts` for the full `QueryEvent` / `QueryOptions` surface.
1324
+
1325
+ For direct lifecycle control, create a manager with `createAgentManager(loadConfig(workspace))`; its public operations cover planning, spawning, listing, inspection, sending/resuming, waiting, stopping, evidence publishing/review, and validated application.
1326
+
1327
+ ## Development
1328
+
1329
+ ```bash
1330
+ npm run typecheck # TypeScript check
1331
+ npm test # Build, then run unit + contract + integration suites
1332
+ npm run test:unit # Deterministic unit suite
1333
+ npm run test:contract
1334
+ npm run test:integration # Isolated PTY/process suite (one worker)
1335
+ npm run check # Format, lint, types, architecture, unit, and contract checks
1336
+ npm run test:watch # Watch mode
1337
+ npm run test:coverage
1338
+ npm run build # tsup → dist/
1339
+ npm run dev # Run via tsx
1340
+ npm run lint # ESLint
1341
+ npm run format # Prettier
1342
+ npm run format:check
1343
+ npm run bench:ui # TUI micro-benchmarks
1344
+ npm run bench:runtime # Runtime micro-benchmarks
1345
+ npm run deadcode:check # knip dead-code scan (non-zero exit on findings)
1346
+ npm run deadcode:report # knip scan as Markdown (used for the CI job summary)
1347
+ npm run deadcode:json # knip scan as JSON
1348
+ npm run eval:edit # Edit reliability evaluation (configured provider)
1349
+ npm run eval:compact # Compaction paired evaluation (configured provider)
1350
+ npm run eval:skills # Skill activation evaluation
1351
+ npm run verify:ink-patch
1352
+ npm run release:check # Version, audit, and package smoke checks
1353
+ ```
1354
+
1355
+ Main-branch runtime work also follows the [stabilization gate](docs/stabilization.md): three
1356
+ consecutive green full CI runs and no open lifecycle or accounting regression issues.
1357
+
1358
+ ### Maintenance workflow
1359
+
1360
+ `.github/workflows/maintenance.yml` runs the deterministic half of the nightly maintenance work,
1361
+ daily at 01:00 UTC and on every pull request:
1362
+
1363
+ - **Dead-code report** — runs knip against the committed `knip.json` and writes the result to the
1364
+ job summary. Report-only: the repository carries a backlog of a few hundred unused exports and
1365
+ exported types, and deciding which are safe to remove is a judgment call rather than a gate.
1366
+ `knip.json` lists `src/index.ts`, `src/sdk.ts`, and `src/job-runner.ts` as entry points, so the
1367
+ published SDK surface is never flagged.
1368
+ - **Security advisories** — runs `npm audit` on a schedule (not just when someone pushes) and keeps
1369
+ a single rolling `Dependency security advisories` issue in sync, opening it when an advisory at
1370
+ or above `high` appears, rewriting it as the set changes, and closing it once clear. A scan that
1371
+ fails to complete never closes the issue.
1372
+
1373
+ ## License
1374
+
1375
+ Copyright (c) 2026 letrquan.
1376
+
1377
+ [PolyForm Small Business License 1.0.0](LICENSE). Source-available: read it, change it, and use it
1378
+ for your own work or your company's, provided the company has fewer than 100 people and under
1379
+ 1,000,000 USD (2019, inflation-adjusted) revenue in the prior tax year. Larger companies, and anyone
1380
+ wanting terms beyond that, need a separate licence — open an issue.
1381
+
1382
+ This is not an open-source licence: it restricts who may use the software commercially. It does not
1383
+ restrict reading, modifying, or redistributing it under the same terms.