@enderfga/claw-orchestrator 7.5.2 → 7.5.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. package/README.md +26 -27
  2. package/configs/engines/README.md +7 -6
  3. package/dist/bin/cli.js +1 -1
  4. package/dist/bin/cli.js.map +1 -1
  5. package/dist/src/acp-server.d.ts +1 -1
  6. package/dist/src/acp-server.js +7 -5
  7. package/dist/src/acp-server.js.map +1 -1
  8. package/dist/src/autoloop/dispatcher.js +3 -3
  9. package/dist/src/autoloop/dispatcher.js.map +1 -1
  10. package/dist/src/autoloop/notify.d.ts +5 -7
  11. package/dist/src/autoloop/notify.js +21 -20
  12. package/dist/src/autoloop/notify.js.map +1 -1
  13. package/dist/src/base-oneshot-session.js +20 -1
  14. package/dist/src/base-oneshot-session.js.map +1 -1
  15. package/dist/src/dashboard/index.html +94 -15
  16. package/dist/src/embedded-server.js +40 -7
  17. package/dist/src/embedded-server.js.map +1 -1
  18. package/dist/src/fanout.d.ts +6 -0
  19. package/dist/src/fanout.js +1 -0
  20. package/dist/src/fanout.js.map +1 -1
  21. package/dist/src/index.js +19 -11
  22. package/dist/src/index.js.map +1 -1
  23. package/dist/src/kernel/nodes/fanout.js +1 -0
  24. package/dist/src/kernel/nodes/fanout.js.map +1 -1
  25. package/dist/src/kernel/types.d.ts +2 -0
  26. package/dist/src/kernel/types.js.map +1 -1
  27. package/dist/src/models.js +36 -7
  28. package/dist/src/models.js.map +1 -1
  29. package/dist/src/openai-compat.d.ts +2 -2
  30. package/dist/src/openai-compat.js +5 -2
  31. package/dist/src/openai-compat.js.map +1 -1
  32. package/dist/src/persistent-agy-session.js +6 -1
  33. package/dist/src/persistent-agy-session.js.map +1 -1
  34. package/dist/src/session-manager.d.ts +1 -0
  35. package/dist/src/session-manager.js +17 -5
  36. package/dist/src/session-manager.js.map +1 -1
  37. package/dist/src/types.d.ts +2 -0
  38. package/openclaw.plugin.json +1 -1
  39. package/package.json +2 -2
  40. package/skills/SKILL.md +31 -32
  41. package/skills/references/acp.md +19 -36
  42. package/skills/references/autoloop.md +163 -180
  43. package/skills/references/claude-cli-tracking.md +28 -27
  44. package/skills/references/cli.md +62 -79
  45. package/skills/references/council.md +40 -63
  46. package/skills/references/dashboard.md +42 -55
  47. package/skills/references/getting-started.md +21 -15
  48. package/skills/references/inbox.md +6 -4
  49. package/skills/references/mcp.md +29 -24
  50. package/skills/references/multi-engine.md +105 -153
  51. package/skills/references/observability.md +42 -32
  52. package/skills/references/openai-compat.md +169 -303
  53. package/skills/references/sessions.md +20 -29
  54. package/skills/references/tools.md +62 -76
  55. package/skills/references/ultra.md +17 -16
  56. package/skills/references/ultraapp.md +59 -64
  57. package/skills/references/verification.md +29 -52
  58. package/skills/references/workflow.md +37 -104
  59. package/skills/ultraapp/SKILL.md +9 -10
@@ -17,7 +17,7 @@ SessionManager
17
17
  ├── engine: 'grok' → PersistentGrokSession
18
18
  │ └── Wraps: grok -p --output-format json (xAI Grok Build, per-message spawning)
19
19
  ├── engine: 'cursor' → PersistentCursorSession (legacy)
20
- │ └── Wraps: agent -p --force --trust --output-format stream-json (per-message spawning)
20
+ │ └── Wraps: cursor-agent -p --trust --output-format stream-json (per-message spawning)
21
21
  ├── engine: 'opencode' → PersistentOpencodeSession
22
22
  │ └── Wraps: opencode run --format json (per-message spawning)
23
23
  └── engine: 'custom' → PersistentCustomSession
@@ -28,33 +28,28 @@ SessionManager
28
28
 
29
29
  ### Claude Code (`engine: 'claude'`)
30
30
 
31
- Default engine. Long-running subprocess with streaming JSON I/O. Tested with Claude Code CLI **2.1.207**.
31
+ Default engine. Long-running subprocess with streaming JSON I/O. Tested with Claude Code CLI **2.1.280**.
32
32
 
33
33
  - Persistent multi-turn conversations
34
34
  - Real-time streaming (text, tool_use, tool_result, system events)
35
35
  - Session resume via `--resume`
36
36
  - Full cost tracking from API usage data
37
- - Cross-session peer messaging (`crossSessionInbound`): sets this session's policy for messages sent from other Claude Code sessions on the same machine — `accept` delivers straight in, `hold` waits for a human to approve it in that session's terminal, `refuse` rejects it. There is no CLI flag for this; it is a settings key, delivered through the `--settings` merge. Worth setting explicitly for orchestrated sessions: with no value the CLI decides from the two sides' permission modes and holds when they differ, and an orchestrated session (`bypassPermissions` / `acceptEdits`) versus a human terminal (prompting) is exactly that case — so the message parks waiting for approval in a terminal nobody is watching. Sessions started by the orchestrator do register as addressable peers and do receive messages (verified against 2.1.232 by sending to a live one and getting a reply). Note that a user-level `~/.claude/settings.json` value may take precedence over the per-session one; only the `accept` path has been confirmed end-to-end here.
37
+ - Cross-session peer messaging (`crossSessionInbound`): this session's policy for messages from
38
+ other Claude Code sessions on the same machine — `accept` delivers them, `hold` waits for approval
39
+ in that session's terminal, `refuse` rejects them. It is a settings key, passed through
40
+ `--settings`. Set it explicitly on orchestrated sessions: when it is unset, the CLI holds messages
41
+ whenever the two sides' permission modes differ, so a message can wait for approval in a terminal
42
+ nobody is watching. A user-level `~/.claude/settings.json` value may take precedence over the
43
+ per-session one.
38
44
  - Hook lifecycle events (`includeHookEvents`), subagent output forwarding (`forwardSubagentText`), permission delegation (`permissionPromptTool`), prompt cache optimization (`bare` + `excludeDynamicSystemPromptSections` + `enablePromptCaching1H`), debug control, `--from-pr` resume, and MCP channel subscriptions
39
- - `--permission-prompts none` is passed whenever no `permissionPromptTool` is configured (CLI
40
- 2.1.259+). This spawn shape has no TTY and, without a prompt tool, no host to answer a permission
41
- prompt — so before this a tool call the permission mode did not already decide sat waiting for an
42
- answer that could never come, until the turn timeout. `none` denies it instead; the model sees the
43
- denial and can adapt, and the permission mode still decides everything else. With a prompt tool
44
- configured the CLI's default (`host`) is left in place so the tool is asked
45
- - `restricted` → `--restricted` (CLI 2.1.249+): removes the command- and code-running tools and
46
- `WebFetch` from the session, and ignores user/project/local settings files. Not folded into
47
- `sandboxMode: 'read-only'`, which maps to plan mode — measured against 2.1.251, plan mode alone
48
- refused a direct write, a shell write and a delegated subagent write, so this is defence in depth
49
- rather than a fix, and switching it on implicitly would silently drop the caller's CLAUDE.md
50
- - Fork subagent (`forkSubagent`), tool search (`enableToolSearch`), OpenTelemetry logging toggles (`otelLogUserPrompts`, `otelLogRawApiBodies`), `xhigh` effort tier (Opus 4.7), and `stats.pluginErrors` capture — see [CLI 2.1.121 options in SKILL.md](../SKILL.md) and [tools.md](./tools.md)
51
-
52
- > **Behavior changes from upstream Claude CLI 2.1.121** (worth knowing if you set permission rules):
53
- >
54
- > - `--agent` / `--print` now enforce agent frontmatter `permissionMode`, `tools`, `disallowedTools` (was advisory). Affects `council` agent personas.
55
- > - `Bash(find:*)` permission rule no longer auto-approves `find -exec` or `find -delete`. Add explicit rules if you depend on these.
56
- > - `--dangerously-skip-permissions` also skips prompts for `.claude/skills/` directory. Treat with care.
57
- > - Distributed tracing context (`TRACEPARENT` / `TRACESTATE`) is automatically forwarded to the child process — set them in the parent before starting the session.
45
+ - `--permission-prompts none` is passed whenever no `permissionPromptTool` is configured. The
46
+ session has no TTY and no prompt tool, so a tool call the permission mode does not already decide
47
+ is denied (the model sees the denial and can adapt) instead of waiting until the turn timeout.
48
+ With a prompt tool configured, the CLI's default (`host`) is kept so the tool is asked
49
+ - `restricted` → `--restricted`: removes the command- and code-running tools and `WebFetch` from
50
+ the session, and ignores user/project/local settings files (including CLAUDE.md). It is separate
51
+ from `sandboxMode: 'read-only'`, which maps to plan mode
52
+ - Fork subagent (`forkSubagent`), tool search (`enableToolSearch`), OpenTelemetry logging toggles (`otelLogUserPrompts`, `otelLogRawApiBodies`), the `xhigh` effort tier, and `stats.pluginErrors` capture — see [SKILL.md](../SKILL.md#claude-engine-options) and [tools.md](./tools.md)
58
53
 
59
54
  ```typescript
60
55
  await manager.startSession({
@@ -67,26 +62,23 @@ await manager.startSession({
67
62
 
68
63
  ### OpenAI Codex (`engine: 'codex'`)
69
64
 
70
- Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested with `codex` CLI **0.147.0**.
65
+ Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested with `codex` CLI **0.156.1**.
71
66
 
72
- - Non-interactive execution via `codex exec --sandbox workspace-write --json` (replaces the deprecated `--full-auto` flag from earlier Codex versions)
73
- - Real `usage` from the `turn.completed` JSON event (input, output, cached, reasoning tokens). **These are cumulative over the thread, not per turn** — three identical turns on 0.147.0 report `input_tokens` 13,856 → 27,727 → 41,613, each matching `total_token_usage` in that thread's rollout exactly. They are assigned to the session totals, never added; subtracting consecutive values recovers the turn's own prompt
74
- - `contextPercent` is that per-turn prompt (which, for a thread-resuming engine, is the live context occupancy) over **codex's own limit**, harvested from the thread's rollout file (`model_context_window`, 258,400 on 0.147.0). The model registry holds the published window — 1,050,000 for gpt-5.x — which codex does not honour, so measuring against it reads ~4x low. Resuming a thread also seeds the token baseline from the rollout, so the first send does not mistake the whole thread history for one prompt. All of this is best-effort: an unreadable or `--ephemeral` thread falls back to the registry window
67
+ - Non-interactive execution via `codex exec --sandbox workspace-write --skip-git-repo-check --json`
68
+ - Real `usage` from the `turn.completed` JSON event (input, output, cached, reasoning tokens). **These are cumulative over the thread, not per turn**, so they replace the session totals rather than being added to them; subtracting consecutive values gives one turn's prompt
69
+ - `contextPercent` is the per-turn prompt measured against **codex's own context limit** (`model_context_window`, read from the thread's rollout file), not the model's published window in the registry. Resuming a thread also seeds the token baseline from the rollout, so the first send does not count the whole thread history as one prompt. This is best-effort: an unreadable or `--ephemeral` thread falls back to the registry window
75
70
  - `item.completed` parsing distinguishes `reasoning` / `todo_list` (logged, not counted) from real tool items (`command_execution`, `file_change`, `mcp_tool_call`, `web_search`, which increment `toolCalls`; a non-zero `command_execution.exit_code` increments `toolErrors`)
76
- - Reasoning effort: the engine-agnostic `effort` maps to `-c model_reasoning_effort=<level>` and passes straight through. Codex 0.149's ladder runs `low|medium|high|xhigh|max|ultra` — it is the only engine here that reaches `ultra`, and all three top levels were exercised against 0.149.1. `auto` and `ultracode` are omitted. Note `-c` values are not validated at spawn: codex prints `reasoning effort: <whatever>` and sends it, so an unknown level fails at the API rather than at the command line
77
- - `noSessionPersistence` → `--ephemeral` (accepted by `exec` and `exec resume`); `ignoreUserConfig` → `--ignore-user-config`, which stops `$CODEX_HOME/config.toml` from deciding an orchestrated run's model behind the caller's back (auth still resolves from `CODEX_HOME`); `addDir` → `--add-dir` on the first turn only, since `exec resume` rejects it and the resumed thread keeps the roots it opened with
78
- - `codexProfile` → `--profile <name>` (named config profile from `~/.codex/config.toml`)
79
- - `--worktree` (0.154.0, behind `--enable worktrees`) is deliberately not wired. Measured: the turn's
80
- edits land in `~/.codex/worktrees/<hash>/<repo>` on a detached HEAD, not in the session's `cwd`, and
81
- the JSON stream never reports that path except inside individual `file_change` items. Everything
82
- here that checks work — acceptance contracts, evidence diffs, the baseline change set — reads the
83
- session's `cwd`, so passing the flag would have them verify an untouched tree and report on it.
84
- Council's own per-agent git worktrees cover the isolation use case with paths the orchestrator owns
71
+ - Reasoning effort: the engine-agnostic `effort` maps to `-c model_reasoning_effort=<level>` for `low|medium|high|xhigh|max|ultra`; `auto` and `ultracode` pass nothing. `-c` values are not validated at spawn, so an unknown level fails at the API rather than at the command line
72
+ - `appendSystemPrompt`: codex has no system-prompt flag, so it is placed at the top of the first message of each new conversation (a resumed thread already carries it)
73
+ - `jsonSchema` → `--output-schema <file>` (written to a temp file, accepted by `exec` and `exec resume`)
74
+ - `noSessionPersistence` → `--ephemeral` (accepted by `exec` and `exec resume`); `ignoreUserConfig` → `--ignore-user-config`, which stops `$CODEX_HOME/config.toml` from choosing the model for an orchestrated run (auth still resolves from `CODEX_HOME`); `addDir` → `--add-dir` on the first turn only, since `exec resume` rejects it and the resumed thread keeps the roots it opened with
75
+ - `codexProfile` → `--profile <name>` (named config profile from `~/.codex/config.toml`), first turn only
76
+ - `--worktree` is not passed. With it, codex writes the turn's edits to `~/.codex/worktrees/<hash>/<repo>` instead of the session's `cwd`, while acceptance contracts, evidence diffs and the baseline change set all read the session's `cwd`. Council's per-agent git worktrees cover the isolation use case
85
77
  - Per-session continuity: the `thread_id` from the first turn's `thread.started` event is captured and reused via `codex exec resume <id>` for subsequent sends, so the model sees prior turns
86
- - `sandboxMode` maps to `--sandbox <mode>` on the first turn. **A resumed thread does not inherit it**, and `codex exec resume` rejects `--sandbox`, so the policy is restated as `-c sandbox_mode="<mode>"` on every resume. Without that, a `read-only` session goes writable from its second turn onward — verified against 0.146.0, where such a session wrote to disk on turn 2 on every attempt. Re-probed on 0.147.0 (direct write, shell redirect and delegate-to-subagent, each on a resumed turn): no writes
78
+ - `sandboxMode` maps to `--sandbox <mode>` on the first turn. A resumed thread does not keep it, and `codex exec resume` rejects `--sandbox`, so the policy is restated as `-c sandbox_mode="<mode>"` on every resumed turn
87
79
  - One-shot execution per message (no persistent subprocess between sends)
88
80
  - Captures the real Codex thread ID and persists it, so later sends and process-level session resume use `codex exec resume <thread_id>`
89
- - Working directory passed via `-C` flag
81
+ - Working directory passed via `-C` on the first turn
90
82
  - Default model: `gpt-5.5`
91
83
  - Requires `codex` CLI >= 0.119 (for `exec resume`): `npm install -g @openai/codex`
92
84
  - **Does not support `/goal`** — for that, use `engine: 'codex-app'` below
@@ -108,13 +100,14 @@ Wraps `codex app-server --listen stdio:// --enable goals` as a long-running JSON
108
100
  - Long-running subprocess; one `codex app-server` per session
109
101
  - JSON-RPC 2.0 over stdio with v2 protocol method names (`initialize`, `thread/start`, `turn/start`, ...)
110
102
  - Real-time streaming via `item/agentMessage/delta` notifications
111
- - Cumulative token tracking from `thread/tokenUsage/updated` notifications. The same notification's `last` breakdown and `modelContextWindow` drive `contextPercent`, so it reports live occupancy against the window the server actually enforces (258,400 on 0.147.0) rather than a running total over the model's published window
103
+ - Cumulative token tracking from `thread/tokenUsage/updated` notifications. The same notification's `last` breakdown and `modelContextWindow` drive `contextPercent`, so it reports live occupancy against the window the server enforces rather than a running total over the model's published window
104
+ - `appendSystemPrompt` is placed at the top of the first message of a new thread (app-server has no system-prompt flag); a resumed thread already carries it
112
105
  - Goal lifecycle observation via `thread/goal/updated` and `thread/goal/cleared` notifications
113
106
  - Goal control via the `codex_goal_*` tools (which internally send the `/goal` slash command as user text — see [tools.md](./tools.md#codex-13))
114
107
  - v2 RPC tools (Codex 0.137): `codex_interrupt` (`turn/interrupt`), `codex_steer` (`turn/steer`), `codex_fork` (`thread/fork`), `codex_rollback` (`thread/rollback`), `codex_models` (`model/list`), `codex_thread_list` (`thread/list`). A `turn/completed` with `status: 'failed'` rejects the turn and increments `toolErrors`.
115
108
  - Thread resume: starting with `resumeSessionId` loads the existing thread via `thread/resume` instead of `thread/start`.
116
109
 
117
- > **Feature-flag risk.** The `goals` feature is marked "under development" in Codex 0.128.0 and has known bugs (e.g. issue #20591). The session class always passes `--enable goals` so it works the moment upstream stabilizes the feature, but during the transition period some goal commands may fail or be silently dropped on the server side. The wrapper layer is unaffected.
110
+ > **Feature flag.** `goals` is an experimental Codex feature. The session always passes `--enable goals`; some goal commands may still fail or be ignored by the server.
118
111
 
119
112
  ```typescript
120
113
  await manager.startSession({
@@ -133,7 +126,7 @@ await manager.startSession({
133
126
 
134
127
  Wraps Google's **Antigravity CLI** (`agy`) — the successor to Gemini CLI (consumer
135
128
  Gemini CLI tiers stopped serving 2026-06-18). Each `send()` spawns a new process
136
- in print mode. Adapter behavior is covered through `agy` **1.2.2**.
129
+ in print mode. Tested with `agy` **1.2.8**.
137
130
 
138
131
  - One-shot execution per message (no persistent subprocess)
139
132
  - **Structured output and real usage** — `--output-format stream-json` emits an
@@ -145,21 +138,18 @@ in print mode. Adapter behavior is covered through `agy` **1.2.2**.
145
138
  scrape remains as a fallback for turns that die before emitting `init`. Seed it
146
139
  externally via `resumeSessionId` (bare UUID only); read it back from
147
140
  `getStats().agyConversationId`.
148
- - **Empty responses fail adapter-wide and remain recoverable**: an exit-0 result
149
- with a missing, blank, or whitespace-only response rejects instead of becoming
150
- a successful empty reply, whether the caller is Autoloop, MCP, HTTP, or the
151
- library API. agy 1.1.26 may do this after plan mode soft-denies a tool
152
- confirmation. The adapter clears the per-session log before each spawn and
153
- recognizes only the narrow current-turn `tool_confirmation_manager` marker,
154
- returning a fixed sanitized diagnosis without exposing native log content. A
155
- conversation id already emitted by `init` is retained for the caller's next
156
- send; the failed turn is not retried automatically. On agy 1.2.2 the same
157
- denial can accompany `status: SUCCESS` and a non-empty reply; the refused tool
158
- names are emitted as `permission_denials`, which SessionManager exposes as
159
- `SendResult.permissionDenials` without discarding the reply.
141
+ - **Empty responses fail**: an exit-0 result with a missing or blank response
142
+ is treated as a failed turn, not a successful empty reply, and is not retried.
143
+ A conversation id already received from `init` is kept for the next send. When
144
+ the failure follows a tool confirmation refused in plan mode, the error says so
145
+ (a fixed message; native log content is not exposed).
146
+ - **Refused tools**: when agy refuses tools but still replies, the refused tool
147
+ names are exposed as `SendResult.permissionDenials` alongside the reply.
148
+ - `appendSystemPrompt`: agy has no system-prompt flag, so it is placed at the top
149
+ of the first message of each new conversation.
160
150
  - **Reasoning effort**: session `effort` and per-turn `session_send` overrides map
161
151
  to `--effort`. agy accepts `low`, `medium`, and `high`; everything above that
162
- (`xhigh`, `max`, `ultra`) clamps to `high`. agy 1.1.25 requires an effort with unsuffixed base
152
+ (`xhigh`, `max`, `ultra`) clamps to `high`. agy requires an effort with unsuffixed base
163
153
  slugs such as `gemini-3.7-flash`, so `auto` resolves those to `high`; a model
164
154
  already ending in `-low`, `-medium`, or `-high` keeps that qualified effort.
165
155
  Per-turn overrides also work with qualified slugs: the adapter removes a
@@ -173,14 +163,15 @@ in print mode. Adapter behavior is covered through `agy` **1.2.2**.
173
163
  did not ask for; run `agy models` to see the tiers a slug actually exposes.
174
164
 
175
165
  - Permission modes: `bypassPermissions` → `--dangerously-skip-permissions`,
176
- `default` → `--sandbox` (terminal-restricted), and
166
+ `default` and `manual` → `--sandbox` (terminal-restricted), and
177
167
  `sandboxMode: 'read-only'` → `--mode plan` (takes precedence). Other modes
178
168
  run agy's own approval flow, which can block in headless print mode. A caller
179
169
  must explicitly choose `bypassPermissions` for a write-enabled session; it is
180
170
  not a recovery mechanism. In particular, an Autoloop Planner stays on
181
171
  `--mode plan` when its preserved conversation is resumed.
182
- - agy enforces its own print timeout (default 5m); the engine derives
183
- `--print-timeout` from the send timeout so the wrapper timer decides
172
+ - The engine always passes `--print-timeout` (the send timeout plus 5s), so the
173
+ wrapper's timer decides when a turn ends; without it a stuck headless agy turn
174
+ can run indefinitely
184
175
  - Do not rely on an unknown `--model` falling back: current agy versions can
185
176
  report `status: ERROR` with no usable response. The adapter rejects result
186
177
  errors, non-success statuses, and empty responses. `agy-flash` and the engine
@@ -216,49 +207,35 @@ await manager.startSession({
216
207
  ### Grok Build (`engine: 'grok'`)
217
208
 
218
209
  Wraps xAI's **Grok Build** CLI. Each `send()` spawns `grok -p <msg> --output-format json`, which
219
- prints a single JSON object and exits. Verified against `grok` **1.0.5**.
220
-
221
- - **Cost comes from the engine, not from our price table.** The result object carries
222
- `total_cost_usd`, and the wrapper writes it straight into the session's spend. Every other engine
223
- here multiplies tokens by a rate in `models.ts` — the metadata most prone to going stale — so on
224
- this engine the run ledger and the `maxBudgetUsd` gate both read what xAI actually charged.
225
- `grok-4.6` is still registered, for its context window and an indicative breakdown; its two price
226
- tiers ($2/$0.50/$6 under a 200K prompt, $4/$1/$12 at or above, charged across the whole request)
227
- therefore never have to be modelled here.
210
+ prints a single JSON object and exits. Tested with `grok` **1.0.41**.
211
+
212
+ - **Cost comes from the engine, not from the price table.** The result object carries
213
+ `total_cost_usd`, and the wrapper writes it straight into the session's spend, so the run ledger
214
+ and the `maxBudgetUsd` gate both read what xAI charged. Other engines multiply tokens by a rate in
215
+ `models.ts`. `grok-4.6` is still registered, for its context window and an indicative breakdown.
228
216
  - **Real conversation continuity**: the `sessionId` from turn 1 is replayed as `--resume <id>`.
229
- `--continue` is deliberately not used — it means "the most recent session for this cwd", which
230
- collides between concurrent sessions. Confirmed with a two-turn recall test, not inferred.
217
+ `--continue` is not used — it means "the most recent session for this cwd", which collides
218
+ between concurrent sessions.
231
219
  - Real token counts from `usage` (`input_tokens`, `output_tokens`, `cache_read_input_tokens`).
232
- These are **per-turn**, not cumulative over the thread — checked by resuming and reading turn 2,
233
- because the same-looking field on codex is a running total.
220
+ These are **per-turn**, unlike codex, where the same field is cumulative.
234
221
  - Permission modes pass straight through: grok's `--permission-mode` takes the same vocabulary we
235
222
  use. The one exception is our `manual`, which grok spells `default`.
236
- - Reasoning effort maps to `--effort`; grok 1.0.13 accepts `low|medium|high|xhigh` (it names the set in
237
- its own rejection message), so only `max` and `ultra` clamp — to `xhigh`.
238
- - Session options that reach grok since 1.0.13: `appendSystemPrompt` → `--rules` (appends, unlike
223
+ - Reasoning effort maps to `--effort`; grok accepts `low|medium|high|xhigh`, so `max` and `ultra`
224
+ clamp to `xhigh`.
225
+ - Session options that reach grok: `appendSystemPrompt` → `--rules` (appends, unlike
239
226
  `systemPrompt` → `--system-prompt-override`, which replaces), `allowedTools` → `--tools`,
240
227
  `disallowedTools` → `--disallowed-tools`, `jsonSchema` → `--json-schema` (inline, and it implies the
241
228
  JSON output format already asked for), `agent` → `--agent`, `agents` → `--agents`,
242
229
  `dangerouslySkipPermissions` → `--always-approve`, `customSessionId` → `--session-id`, `forkSession`
243
230
  → `--fork-session`. **grok validates neither tool list**: a name that does not exist is ignored
244
231
  rather than rejected, so a typo in a denylist leaves the tool enabled. Prefer an allowlist.
245
- - `-p` is now the short form of `--single`, not `--print`. The short form is what this wrapper passes,
246
- so the 1.0.5 → 1.0.13 rename is invisible here; a call written against the long name is not.
247
- - **`sandboxMode: 'read-only'` is refused, not approximated — and against 1.0.13 that is measured.**
248
- The obvious construction is a `--tools` allowlist of read-only built-ins plus `--permission-mode
249
- plan`. It refuses a direct write and a shell write, then loses to the third prompt in the matrix:
250
- asked to delegate, the session spawns a subagent and the file appears. The subagent does not
251
- inherit the parent's tool restriction — the same load-bearing hole found in OpenCode, where
252
- denying the write tools without denying `task` left the delegation path open. grok also ships
253
- `--no-subagents`, the obvious next probe, deliberately not wired: the run that would have
254
- confirmed it hit the account's free-tier limit, and a probe that fails for lack of quota writes no
255
- file either. A read-only grok session throws rather than running writable under a read-only label.
256
- - **On a spent free tier, `grok -p` can hang silently instead of erroring.** Earlier in the same
257
- session it printed a usage-limit message to stderr and exited 1; later invocations produced nothing
258
- on either stream and never exited — in any directory, with or without `--no-leader`. The session's
259
- turn timeout is what ends such a turn, so a caller on that tier pays the full timeout before seeing
260
- a failure. Nothing in this wrapper can distinguish that hang from a slow turn; check `grok -p` by
261
- hand when a grok session times out with no output.
232
+ - **`sandboxMode: 'read-only'` is refused.** A read-only `--tools` allowlist plus
233
+ `--permission-mode plan` does not stop a delegated subagent from writing, because the subagent
234
+ does not inherit the parent's tool restriction. A read-only grok session therefore throws at start instead of running
235
+ writable.
236
+ - **On an exhausted free tier, `grok -p` may hang with no output instead of exiting with an error.**
237
+ The session's turn timeout is then the only thing that ends the turn. If a grok turn times out with
238
+ no output, run `grok -p` by hand to check your quota.
262
239
  - Binary: `grok` (set `GROK_BIN` to override). Not `agent`: xAI's installer claims that name too,
263
240
  and so did Cursor's.
264
241
  - Requires Grok Build: see `x.ai/cli`.
@@ -278,20 +255,21 @@ await manager.startSession({
278
255
  > The `cursor` engine still exists and still works — existing callers are not broken — but it is
279
256
  > no longer a documented option, is not version-tracked, and gets no new work.
280
257
  >
281
- > Note what this is and is not: Cursor itself is **not** discontinued. Anysphere was acquired by
282
- > SpaceX (closed 2026-08-15) and folded into the SpaceXAI team, and the CLI has shipped since. Two
283
- > practical things pushed it out of the tracked set. Cursor never reports which model actually ran
284
- > — its `system` init event says `"model": "Auto"` — so a router that spans Claude, GPT and Grok
285
- > leaves every cost row attributed to a hardcoded proxy rate. And xAI's Grok installer now claims
286
- > the bare `agent` name, so the binary that name resolves to depends on install order.
258
+ > Cursor itself is still maintained. It left the tracked set because it does not report which
259
+ > model ran (its `system` init event says `"model": "Auto"`), so costs cannot be attributed to a
260
+ > model, and because xAI's Grok installer also claims the bare `agent` binary name.
287
261
 
288
- Wraps the Cursor Agent CLI with `--print --output-format stream-json`. Write-enabled sessions use `--force`. Each `send()` spawns a new process.
262
+ Wraps the Cursor Agent CLI with `-p --output-format stream-json`. Write-enabled sessions use `--force`. Each `send()` spawns a new process.
289
263
 
290
264
  - Conversation continuity: the chat id from the first turn's `system` event is captured and passed back as `--resume <chatId>` on later sends, so the model sees prior turns. `--continue` is deliberately not used: it resumes "the latest chat", which collides between concurrent sessions.
291
265
  - One-shot execution per message (no persistent subprocess)
292
266
  - Working directory via `--workspace` flag
293
267
  - Real token counts from stream-json `result` events (camelCase: `inputTokens`, `outputTokens`, `cacheReadTokens`)
294
- - `--force` enables auto-approval of file changes. `sandboxMode: 'read-only'` does **not** use `--force`; it enforces read-only via a binding `.cursor/cli.json` deny config (`Write`/`Edit`/`Shell` denied) written into an isolated temp dir used as the process cwd, with `--workspace` pointing at the real project (the repo tree is never modified). `--mode plan` is passed too as model steering, but the deny config is the actual boundary — plan mode alone is model-cooperative and was verified to let an adversarial prompt write. Do not add `--sandbox` (it does not restrict in-workspace writes and overrides the mode). Read/grep/search remain available
268
+ - `--force` enables auto-approval of file changes. `sandboxMode: 'read-only'` does **not** use `--force`:
269
+ - read-only is enforced by a `.cursor/cli.json` deny config (`Write`/`Edit`/`Shell` denied), written into an isolated temp dir used as the process cwd, with `--workspace` pointing at the real project (the repo tree is never modified)
270
+ - `--mode plan` is passed as well, but it only steers the model; the deny config is the boundary
271
+ - `--sandbox` is not added, since it does not restrict in-workspace writes and overrides the mode
272
+ - read/grep/search remain available
295
273
  - `--trust` auto-trusts the workspace without prompting
296
274
  - Cursor uses its own model routing (e.g., `sonnet-4`, `gpt-5`, `auto`)
297
275
  - Requires Cursor Agent CLI: `curl https://cursor.com/install -fsSL | bash`
@@ -314,27 +292,22 @@ await manager.startSession({
314
292
  bundled in `configs/engines/`. The preset form exists so a third-party CLI is
315
293
  described once and shipped, rather than retyped by every caller.
316
294
 
317
- Presets are **community** entries, and the tier line is verification rather than
318
- code quality:
295
+ Engines fall into three tiers, decided by how they are verified. Bundled presets
296
+ are always `community`:
319
297
 
320
- | Tier | Maintainer | What this project claims |
298
+ | Tier | Maintainer | What is verified |
321
299
  | ------------- | ----------- | --------------------------------------------------------------------------------------------------- |
322
- | **core** | Maintainers | Wrapped here, exercised live weekly, pinned to a version someone here ran |
300
+ | **core** | Maintainers | Wrapped in code, exercised live weekly, pinned to a tested version |
323
301
  | **community** | Contributor | The schema is validated and the preset ships. Whether it runs is the maintainer's dated attestation |
324
- | **legacy** | — | Still wired, no longer tracked |
302
+ | **legacy** | — | Still wired, no longer tracked |
325
303
 
326
- Core requires that a maintainer can obtain and run the binary, because that tier
327
- promises the weekly live turn. Most CLIs worth supporting are behind credentials
328
- this project does not hold — a bar nobody can clear would keep the community tier
329
- permanently empty, and an empty tier looks the same as never having done the work.
330
- So the bar is an attributable, dated, falsifiable claim: `provenance.maintainer`,
331
- `provenance.verifiedAgainst` (the engine version), `provenance.verifiedOn`, and a
332
- link to the captured smoke transcript.
304
+ A community preset must name its maintainer and the engine version and date it
305
+ was verified against (`provenance.maintainer`, `provenance.verifiedAgainst`,
306
+ `provenance.verifiedOn`), with a link to the captured smoke-test transcript.
333
307
 
334
308
  A preset never carries protocol translation. A CLI speaking its own wire format
335
309
  needs an adapter binary in its author's package, with the preset pointing `bin`
336
- at it — the split `@enderfga/dsh-clawo` already uses. That is what makes a preset
337
- something this project can ship without owning a protocol it cannot test.
310
+ at it — as `@enderfga/dsh-clawo` does.
338
311
 
339
312
  Over HTTP — which includes the `clawo` CLI — a custom engine may be given **only**
340
313
  as a preset id. An inline config names a binary and its arguments and is refused
@@ -355,11 +328,17 @@ Wraps the [sst/opencode](https://github.com/sst/opencode) CLI with `run --format
355
328
  - NDJSON event stream with envelope `{ type, timestamp, sessionID, ... }`
356
329
  - Event types: `text`, `reasoning`, `tool_use`, `step_start`, `step_finish`, `error`
357
330
  - `text` and `tool_use` are **cumulative snapshots** keyed by `part.id` / `part.callID`; the wrapper diffs them to produce streaming deltas for `onText` callbacks and counts each tool invocation once
358
- - Real token counts from `step_finish.part.tokens.{input,output,cache.{read,write}}`. **`input` is the uncached remainder only** — opencode's own `total` is `input + output + cache.read + cache.write` — so the cached part is billed on top of it, not carved out of it, and `contextPercent` is measured against the whole input side. On a resumed turn the remainder is tiny next to the cached part (58 against 26,240), which is what makes the distinction matter
331
+ - Real token counts from `step_finish.part.tokens.{input,output,cache.{read,write}}`. **`input` is the uncached remainder only** — opencode's own `total` is `input + output + cache.read + cache.write` — so the cached part is billed on top of it, not carved out of it, and `contextPercent` is measured against the whole input side. On a resumed turn the uncached remainder is usually a small fraction of the cached part
359
332
  - Reasoning effort maps to `--variant`, opencode's provider-specific effort knob. opencode does not validate the value: a level the provider does not offer runs the turn at its default rather than failing
333
+ - `appendSystemPrompt`: opencode has no system-prompt flag, so it is placed at the top of the first message of each new conversation
360
334
  - The wrapper closes the subprocess's stdin immediately after spawn (opencode otherwise reads stdin and blocks on EOF, hanging the call)
361
335
  - Provider-agnostic: opencode's `--model` expects `provider/model` form (e.g. `anthropic/claude-sonnet-4`). The wrapper passes `--model` through only when the value contains a `/`; otherwise opencode's own default applies
362
- - `sandboxMode: 'read-only'` spawns a generated `clawo-readonly` agent (`--agent clawo-readonly` plus an `OPENCODE_CONFIG_CONTENT` env var defining it) that denies `edit` / `bash` / `external_directory` / `webfetch` / **`task`** at the permission level and additionally removes those tools outright via the agent's `tools` map. It deliberately does **not** use OpenCode's built-in `plan` agent: that is a user-overridable preset whose compiled rules start with `{"permission":"*","action":"allow"}` and deny neither `bash` nor `edit`, so a "read-only" session could still author files through a shell heredoc. **`task` is the load-bearing denial**: denying only the write tools leaves the delegation path open, and the agent will hand the write to a subagent that runs under the default writable agent — asked to delegate, a session denied only `edit`/`bash`/`external_directory` wrote to disk on every attempt. Verify this config only with adversarial writes, and include prompts that ask the agent to delegate; `opencode agent list` renders compiled permission rules that look identical for a safe and an unsafe agent, and a probe that only asks for a direct write passes even when the delegation path is wide open
336
+ - `sandboxMode: 'read-only'` runs a generated `clawo-readonly` agent (`--agent clawo-readonly` plus an `OPENCODE_CONFIG_CONTENT` env var defining it):
337
+ - it denies `edit`, `bash`, `external_directory`, `webfetch` and `task` at the permission level, and also removes those tools via the agent's `tools` map
338
+ - denying `task` matters: otherwise the agent can hand a write to a subagent that runs under the default writable agent
339
+ - OpenCode's built-in `plan` agent is not used, because it permits both `bash` and `edit`
340
+ - if the `clawo-readonly` agent fails to load, the turn is refused rather than run with write access
341
+ - to test a change to this config, use adversarial prompts that include asking the agent to delegate; `opencode agent list` shows compiled rules that look the same for a safe and an unsafe agent
363
342
  - Requires opencode installed: `brew install sst/tap/opencode` or `npm install -g opencode-ai`. Auth via `opencode auth login` **or** any provider env var (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GEMINI_API_KEY`, etc.) — opencode picks up either path
364
343
  - Binary: `opencode` (set `OPENCODE_BIN` env var to override)
365
344
 
@@ -422,12 +401,14 @@ is not the answer on every engine. See "Stats & Monitoring" in `sessions.md`.
422
401
 
423
402
  Team tools (`team_list`, `team_send`) operate on the same virtual-team layer for **every** engine: the "team" is the set of all active sessions managed by SessionManager.
424
403
 
425
- | Engine | `team_list` | `team_send` |
426
- | ----------- | ------------------------------------------ | ------------------------------ |
427
- | Claude | Lists other active SessionManager sessions | Routes via cross-session inbox |
428
- | Codex | Lists other active SessionManager sessions | Routes via cross-session inbox |
429
- | Antigravity | Lists other active SessionManager sessions | Routes via cross-session inbox |
430
- | Cursor | Lists other active SessionManager sessions | Routes via cross-session inbox |
404
+ | Engine | `team_list` | `team_send` |
405
+ | ------------------- | ------------------------------------------ | ------------------------------ |
406
+ | Claude | Lists other active SessionManager sessions | Routes via cross-session inbox |
407
+ | Codex / `codex-app` | Lists other active SessionManager sessions | Routes via cross-session inbox |
408
+ | Antigravity | Lists other active SessionManager sessions | Routes via cross-session inbox |
409
+ | Grok | Lists other active SessionManager sessions | Routes via cross-session inbox |
410
+ | OpenCode | Lists other active SessionManager sessions | Routes via cross-session inbox |
411
+ | Custom | Lists other active SessionManager sessions | Routes via cross-session inbox |
431
412
 
432
413
  Messages are delivered via the inbox system — idle sessions receive immediately, busy sessions queue for later delivery.
433
414
 
@@ -591,37 +572,7 @@ await manager.startSession({
591
572
 
592
573
  ### Example: Google Antigravity CLI (`agy`)
593
574
 
594
- > **Note:** `agy` now has first-class support — use [`engine: 'agy'`](#google-antigravity-engine-agy)
595
- > instead, which adds conversation resume and timeout coherence the recipe below
596
- > lacks. This recipe remains as a reference for driving older agy builds or
597
- > forks with a diverged flag surface:
598
-
599
- ```typescript
600
- await manager.startSession({
601
- name: 'antigravity-task',
602
- engine: 'custom',
603
- cwd: '/project',
604
- dangerouslySkipPermissions: true,
605
- customEngine: {
606
- name: 'antigravity',
607
- bin: 'agy', // install: curl -fsSL https://antigravity.google/cli/install.sh | bash
608
- binEnv: 'AGY_BIN',
609
- persistent: false,
610
- args: {
611
- print: '-p', // single-prompt headless mode
612
- skipPermissions: '--dangerously-skip-permissions',
613
- workspace: '--add-dir',
614
- // NOTE: agy 1.0.2 has NO --output-format flag — output is plain text only.
615
- // Omitting outputFormat makes the wrapper parse plain text and *estimate*
616
- // tokens (no real usage / tool-call events). Watch for a JSON output mode.
617
- },
618
- },
619
- });
620
- ```
621
-
622
- Caveats with `agy` 1.0.2: (1) no structured/stream-json output → token counts are
623
- estimated, not real; (2) requires a one-time `agy` Google OAuth login; (3) resume
624
- by conversation ID isn't wired (no JSON stream to capture the ID from).
575
+ Use the built-in [`engine: 'agy'`](#google-antigravity-engine-agy) instead of a custom config.
625
576
 
626
577
  ### Custom Engine in Council
627
578
 
@@ -649,9 +600,10 @@ manager.councilStart('Build feature X', {
649
600
  To add a built-in engine (for CLIs that need custom protocol handling beyond what `CustomEngineConfig` supports):
650
601
 
651
602
  1. Create `src/persistent-<engine>-session.ts` implementing `ISession`
652
- 2. Add the engine name to `EngineType` in `src/types.ts`
653
- 3. Add a case to `SessionManager._createSession()`
654
- 4. Add model pricing to `MODELS[]` in `src/models.ts`
603
+ 2. Add the engine name to `ENGINE_TYPES` (from which `EngineType` is derived) and its binary to `ENGINE_BINARY_NAMES` in `src/types.ts`
604
+ 3. Add a case to `engineHasNativeConversation()` in `src/types.ts` and, if the engine resumes a conversation by id, to `nativeThreadIsLive()` in `src/openai-compat.ts`
605
+ 4. Add a case to `SessionManager._createSession()`
606
+ 5. Add model pricing to `MODELS[]` in `src/models.ts`
655
607
 
656
608
  The `ISession` interface is deliberately minimal — each engine handles its own subprocess bootstrapping, I/O protocol, and cleanup internally.
657
609
 
@@ -4,17 +4,12 @@ Two related surfaces: a durable record of every turn this runtime executes, and
4
4
  spend cap that is enforced by the runtime rather than by whichever CLI happens to
5
5
  support a budget flag.
6
6
 
7
- ## Why
7
+ ## Purpose
8
8
 
9
- `getStats()` / `getCost()` describe a **live** session. They live in memory, and
10
- per-session history is capped and evicted, so a restart erased everything except
11
- the resume-id registry — there was no way to answer "what did we run today, on
12
- which engine, for how much". The run ledger is that record.
13
-
14
- The same gap made `maxBudgetUsd` a promise the runtime did not keep: it was only
15
- ever translated into Claude Code's `--max-budget-usd` flag, so a council of Codex
16
- agents ran with no cap at all. The cap is now applied in `SessionManager`, which
17
- every engine passes through.
9
+ `getStats()` / `getCost()` describe a **live** session and are held in memory.
10
+ The run ledger is the durable record of every turn across restarts: what ran, on
11
+ which engine, for how much. `maxBudgetUsd` is enforced in `SessionManager`, which
12
+ every engine passes through, not only through Claude Code's `--max-budget-usd`.
18
13
 
19
14
  ## The ledger
20
15
 
@@ -34,7 +29,7 @@ every engine passes through.
34
29
  | ----------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
35
30
  | `ts` | ISO timestamp of turn completion |
36
31
  | `session` | SessionManager session name |
37
- | `engine` | `claude` / `codex` / `codex-app` / `grok` / `opencode` / `agy` / `custom` |
32
+ | `engine` | `claude` / `codex` / `codex-app` / `grok` / `opencode` / `agy` / `cursor` (legacy) / `custom` |
38
33
  | `model` | Configured model, or the engine's own reported model when none was set (Claude Code: the model named in its `init` event). Absent when neither is known |
39
34
  | `cwd` | Working directory the turn ran in |
40
35
  | `turn` | 1-based index of the send within the session, counted by the process that recorded it |
@@ -45,7 +40,7 @@ every engine passes through.
45
40
  | `toolCalls` / `toolErrors` | Per-turn deltas |
46
41
  | `ok` | `false` for a turn that threw, or that the session's own `turnsSucceeded` counter did not count (see `sessions.md`). Falls back to "nothing was thrown" when the counter cannot be read |
47
42
  | `error` | Failure text, truncated to 500 chars. Absent when the turn resolved but the engine did not count it as succeeded (an interrupted or non-SUCCESS turn), so a failed row does not always carry one |
48
- | `parent` | council id / fanout id / autoloop run id, when the turn belongs to one |
43
+ | `parent` | council / fanout / autoloop / workflow run id, when the turn belongs to one |
49
44
 
50
45
  Deltas rather than totals means summing a query window gives that window's spend
51
46
  without double-counting.
@@ -74,7 +69,8 @@ curl "http://127.0.0.1:18796/runs?since=24h&limit=200" -H "Authorization: Bearer
74
69
  ```
75
70
 
76
71
  Returns `{ ok, rows, summary }`, where `summary` carries `rows`, `costUsd`,
77
- `tokensIn`, `tokensOut`, `estimatedRows` and a per-engine breakdown. The dashboard
72
+ `tokensIn`, `tokensOut`, `estimatedRows`, `verifiedRows`, `refutedRows`,
73
+ `unverifiedRows` and a per-engine breakdown `byEngine`. The dashboard
78
74
  header shows the 24-hour figure from the same endpoint.
79
75
 
80
76
  Programmatically: `manager.getRunLedger({ since, session, engine, parent, limit })`.
@@ -106,8 +102,8 @@ Notes:
106
102
 
107
103
  ### Zeroing a model's pricing
108
104
 
109
- `pricingOverrides` zeroes the fiction that a subscription seat gets billed per token, and
110
- the cost figures above stop reporting money nobody paid. It also **disables `maxBudgetUsd`
105
+ `pricingOverrides` sets a model's per-token rate, so the cost figures for a subscription
106
+ seat read 0 instead of an API-rate equivalent. Zeroing a rate also **disables `maxBudgetUsd`
111
107
  for that model**: the cap compares the session's accrued `getCost().totalUsd` against it,
112
108
  and that number is pricing-derived, so at a rate of 0 it never reaches any cap. There is
113
109
  no separate token or turn ceiling behind it.
@@ -138,16 +134,30 @@ Where the engine reports usage, those counts are the engine's own. Where it does
138
134
  not, the wrapper falls back to `estimateTokens()` (characters ÷ 4) and the row is
139
135
  flagged `tokensEstimated: true`; the CLI marks those costs with a trailing `~`.
140
136
 
141
- | Engine | Token counts |
142
- | ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
143
- | `claude` | Engine-reported — and so is the **cost**: the `result` event's `total_cost_usd` is taken as-is, so registry drift cannot affect a Claude row. It is the session running total rather than the turn's, so spend advances by the difference between turns. A process started with `--resume` (a model switch, a session recovered after a restart) inherits the resumed session's total since Claude Code 2.1.277, so its first report is taken as a baseline and that one turn is priced from the registry instead — charging the inherited figure would bill the whole history again, and `maxBudgetUsd` gates against this number. Proxy sessions (`baseUrl` set) are the exception: the CLI is told it is running `opus` while another provider serves the tokens, so its figure is Opus list price for someone else's model and the registry estimate is used instead |
144
- | `codex` | Engine-reported |
145
- | `codex-app` | Engine-reported |
146
- | `grok` | Engine-reported — and so is the **cost**: this engine reports `total_cost_usd`, which the wrapper passes through instead of pricing tokens from the registry, so registry drift cannot affect a grok row |
147
- | `cursor` (legacy) | Engine-reported when the stream carries `usage`, else estimated |
148
- | `opencode` | Engine-reported when the run JSON carries `tokens`, else estimated |
149
- | `agy` | Engine-reported when the result event carries usage, else estimated |
150
- | `custom` | Depends on the CLI; estimated when it emits no usage |
137
+ | Engine | Token counts |
138
+ | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
139
+ | `claude` | Engine-reported, including the **cost** (see below) |
140
+ | `codex` | Engine-reported |
141
+ | `codex-app` | Engine-reported |
142
+ | `grok` | Engine-reported — and so is the **cost**: this engine reports `total_cost_usd`, which the wrapper passes through instead of pricing tokens from the registry, so registry drift cannot affect a grok row |
143
+ | `cursor` (legacy) | Engine-reported when the stream carries `usage`, else estimated |
144
+ | `opencode` | Engine-reported when the run JSON carries `tokens`, else estimated |
145
+ | `agy` | Engine-reported when the result event carries usage, else estimated |
146
+ | `custom` | Depends on the CLI; estimated when it emits no usage |
147
+
148
+ How a `claude` row's cost is taken:
149
+
150
+ - The `result` event's `total_cost_usd` is used as-is, so pricing-table drift
151
+ cannot affect a Claude row.
152
+ - That figure is the session's running total, so a turn's cost is the difference
153
+ from the previous report.
154
+ - A process started with `--resume` (a model switch, a session recovered after a
155
+ restart) reports a total that includes the resumed history. Its first report is
156
+ taken as a baseline, and that one turn is priced from the registry instead, so
157
+ the history is not billed twice (`maxBudgetUsd` gates against this number).
158
+ - Proxy sessions (`baseUrl` set) use the registry estimate instead: the CLI
159
+ believes it is running `opus` while another provider serves the tokens, so its
160
+ own figure would be Opus list price.
151
161
 
152
162
  ### What "input tokens" means is not the same on every engine
153
163
 
@@ -183,14 +193,13 @@ Cost figures are also only as good as the pricing table: a model missing from
183
193
  ChatGPT Pro) bill nothing per token while the ledger still reports the API-rate
184
194
  equivalent. Read `costUsd` as "what this would cost at API rates".
185
195
 
186
- ## `ok` vs `verified` (6.0.0)
196
+ ## `ok` vs `verified`
187
197
 
188
- A row now carries two different judgements, and conflating them is the mistake
189
- this section exists to prevent.
198
+ A row carries two different judgements.
190
199
 
191
200
  - **`ok`** — the engine's own terminal verdict for that turn. Codex fails a turn
192
- that emits `turn.failed` while exiting 0; gemini succeeds on exit 53. It is a
193
- careful signal, but it is the engine talking about itself.
201
+ that emits `turn.failed` while exiting 0. It is a careful signal, but it is the
202
+ engine talking about itself.
194
203
  - **`verified`** — an acceptance contract ran against the work and every required
195
204
  check passed. That is the runtime's own measurement.
196
205
 
@@ -210,7 +219,8 @@ either would make the ledger useless for the thing it is for.
210
219
  `verified` is **not written at turn time**, deliberately. The turns that produce
211
220
  the work all finish before the verifier that judges it, so stamping a verdict on
212
221
  them as they are written would be inventing one. It is joined in at read time
213
- from the run record via the row's `parent`, by `annotateVerdicts()`.
222
+ from the run record via the row's `parent`, by `annotateVerdicts()`. The exception
223
+ is a standalone `verify_run`, which writes its verdict directly.
214
224
 
215
225
  Two consequences worth knowing:
216
226
 
@@ -219,7 +229,7 @@ Two consequences worth knowing:
219
229
  - Filtering on `--verified` happens _after_ the join. Pushing the filter into the
220
230
  ledger read would match on a field no row carries yet and return nothing.
221
231
 
222
- ### Other new row fields
232
+ ### Verification row fields
223
233
 
224
234
  | Field | Source |
225
235
  | -------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |