@enderfga/claw-orchestrator 6.0.3 → 6.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +4 -4
- package/dist/src/base-oneshot-session.d.ts +37 -0
- package/dist/src/base-oneshot-session.js +29 -5
- package/dist/src/base-oneshot-session.js.map +1 -1
- package/dist/src/index.js +3 -3
- package/dist/src/index.js.map +1 -1
- package/dist/src/models.js +39 -2
- package/dist/src/models.js.map +1 -1
- package/dist/src/openai-compat.d.ts +27 -1
- package/dist/src/openai-compat.js +354 -35
- package/dist/src/openai-compat.js.map +1 -1
- package/dist/src/persistent-agy-session.js +5 -4
- package/dist/src/persistent-agy-session.js.map +1 -1
- package/dist/src/persistent-codex-session.d.ts +3 -4
- package/dist/src/persistent-codex-session.js +38 -7
- package/dist/src/persistent-codex-session.js.map +1 -1
- package/dist/src/persistent-grok-session.js +18 -4
- package/dist/src/persistent-grok-session.js.map +1 -1
- package/dist/src/persistent-opencode-session.js +21 -2
- package/dist/src/persistent-opencode-session.js.map +1 -1
- package/dist/src/persistent-session.d.ts +74 -0
- package/dist/src/persistent-session.js +205 -28
- package/dist/src/persistent-session.js.map +1 -1
- package/dist/src/types.d.ts +10 -1
- package/dist/src/types.js.map +1 -1
- package/package.json +1 -1
- package/skills/SKILL.md +1 -1
- package/skills/references/claude-cli-tracking.md +2 -1
- package/skills/references/multi-engine.md +8 -5
- package/skills/references/observability.md +50 -25
- package/skills/references/openai-compat.md +250 -0
- package/skills/references/tools.md +3 -2
|
@@ -2,11 +2,12 @@
|
|
|
2
2
|
|
|
3
3
|
This document tracks which Claude Code CLI version Claw Orchestrator is currently synced to, and which features have been integrated.
|
|
4
4
|
|
|
5
|
-
## Currently tracked: **Claude Code CLI 2.1.
|
|
5
|
+
## Currently tracked: **Claude Code CLI 2.1.246** (as of 2026-08-27, plugin v6.1.0)
|
|
6
6
|
|
|
7
7
|
## Sync history
|
|
8
8
|
|
|
9
9
|
| Plugin Version | Claude CLI Version | Date | Notable integrations |
|
|
10
|
+
| v6.1.0 | 2.1.246 | 2026-08-27 | **Engine sweep that turned into a token-accounting audit.** Claude Code 2.1.237→2.1.246, Codex 0.148.0→0.149.1, agy 1.1.15→1.1.21, OpenCode 1.18.18→1.18.23, Grok unchanged at 1.0.5; each ran a live turn and the ACP stdio entry was smoked. No new Claude Code flag needed integrating in that range — what did was what its `result` event has been reporting all along. Four measurement bugs fixed: a turn's usage was added twice (once on `message_delta`, once on `result` — engine said 2/4/47371, we said 4/8/94742); cache writes were absent from the cost formula (engine `$0.322428` vs our `$0.016` on a 1h-cache turn, and `maxBudgetUsd` gates on ours), so Claude now takes `total_cost_usd` — a session running total, applied as a difference; the `input − cached` subtraction is valid only on codex, whose `input_tokens` includes cached reads, while claude/grok/opencode report them alongside; and `contextPercent` read `input_tokens` alone, so a 47k prompt measured 0% — now the whole prompt over `modelUsage[*].contextWindow`. Also: Codex gained a real `max` and an `ultra` above it (we were folding `max`→`xhigh`), Grok now takes `xhigh` natively (we were folding it to `high`), OpenCode's `--variant` means `effort` finally reaches it, `--ephemeral`/`--ignore-user-config`/`--add-dir` wired up on Codex. Registry: `gemini-3.7-flash`/`gemini-3.6-flash`/`gpt-5.2` registered, `gemini-3.5-flash` repriced $0.5/$3 → $1.50/$9. |
|
|
10
11
|
| v4.8.0 | 2.1.207 | 2026-07-12 | **Autoloop role-level multi-engine support.** Planner, Coder, and Reviewer can select independent engines/models while preserving the Claude defaults. Built-in non-Claude Planners use native read-only/plan modes and receive their role protocol in-band. Spawn selections persist across resume; Codex persists its real thread ID. Runtime and invocation checks used Claude Code 2.1.207 and Codex 0.144.1. |
|
|
11
12
|
| v4.7.0 | 2.1.206 | 2026-07-10 | **Antigravity engine ships + permission-mode sync.** Main feature is the community-contributed first-class `engine: 'agy'` (PR #71, reviewed + hardened: layered resume-ID gating, `agy/` prefix routing, shared `sanitize.ts` across all six engines, ENGINE*TYPES single source). Weekly CLI sync: CC 2.1.200 renamed the `default` permission mode to **`manual`** — verified against 2.1.206 that the choices are now `acceptEdits/auto/bypassPermissions/manual/dontAsk/plan`, that `default` is still accepted (hidden compat), and that **`delegate` is hard-rejected at spawn** — so PermissionMode gains `manual`, drops `delegate`, and agy/gemini map `manual` like `default` (→ `--sandbox`). Codex 0.143.0: empirically re-tested `-c model_reasoning_effort=max` — still 400-rejected for gpt-5.5 (the "first-class max" note is Bedrock GPT-5.6-only), so the `max`→`xhigh` map stays. GPT-5.6 Sol/Terra/Luna registered with official pricing ($5/$30, $2.50/$15, $1/$6; 1M/1M/400K ctx) after the user reported using it — it's a limited preview on API/Codex-auth paths (empirically: ChatGPT-account Codex auth gets a 400, which is why the first probe on this box misread it as Bedrock-only; lesson — an auth-path rejection is not model non-existence). Codex default stays gpt-5.5. Free upside: CC 2.1.205 fixed `--json-schema` invalid-schema silent fallback + `format` keyword rejection; CC 2.1.203 fixed background sessions dropping shell-exported `ANTHROPIC_BASE_URL`. Pins → CC 2.1.206 / Codex 0.143.0 (installed; npm has 0.144.1, exec surface unchanged per release notes). |
|
|
12
13
|
| v4.6.0 | 2.1.199 | 2026-07-03 | **Model registry sync — Claude Fable 5.** Registered `claude-fable-5` (first Claude 5-family model, tier above Opus; standard $10/$50 per Mtok, cache read $1, full 1M context at standard rates per the official pricing page) with new `fable` alias; taught the `isClaudeModel`/`resolveProvider` heuristics to recognize `fable`/`mythos` strings (they only matched claude/opus/sonnet/haiku). Mythos 5 not listed (same price, limited availability). CC 2.1.198–199 are subagent/background-agent reliability fixes — no invocation-surface change; free upside for us: subagent partial output on rate-limit/server error is now returned instead of silently dropped, and API errors in subagents are reported to the parent. Codex unchanged (0.142.5 is a log-scrub patch; pin stays 0.142.4 as installed). |
|
|
@@ -62,7 +62,8 @@ Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested wi
|
|
|
62
62
|
- Real `usage` from the `turn.completed` JSON event (input, output, cached, reasoning tokens). **These are cumulative over the thread, not per turn** — three identical turns on 0.147.0 report `input_tokens` 13,856 → 27,727 → 41,613, each matching `total_token_usage` in that thread's rollout exactly. They are assigned to the session totals, never added; subtracting consecutive values recovers the turn's own prompt
|
|
63
63
|
- `contextPercent` is that per-turn prompt (which, for a thread-resuming engine, is the live context occupancy) over **codex's own limit**, harvested from the thread's rollout file (`model_context_window`, 258,400 on 0.147.0). The model registry holds the published window — 1,050,000 for gpt-5.x — which codex does not honour, so measuring against it reads ~4x low. Resuming a thread also seeds the token baseline from the rollout, so the first send does not mistake the whole thread history for one prompt. All of this is best-effort: an unreadable or `--ephemeral` thread falls back to the registry window
|
|
64
64
|
- `item.completed` parsing distinguishes `reasoning` / `todo_list` (logged, not counted) from real tool items (`command_execution`, `file_change`, `mcp_tool_call`, `web_search`, which increment `toolCalls`; a non-zero `command_execution.exit_code` increments `toolErrors`)
|
|
65
|
-
- Reasoning effort: the engine-agnostic `effort` maps to `-c model_reasoning_effort=<level>`
|
|
65
|
+
- Reasoning effort: the engine-agnostic `effort` maps to `-c model_reasoning_effort=<level>` and passes straight through. Codex 0.149's ladder runs `low|medium|high|xhigh|max|ultra` — it is the only engine here that reaches `ultra`, and all three top levels were exercised against 0.149.1. `auto` and `ultracode` are omitted. Note `-c` values are not validated at spawn: codex prints `reasoning effort: <whatever>` and sends it, so an unknown level fails at the API rather than at the command line
|
|
66
|
+
- `noSessionPersistence` → `--ephemeral` (accepted by `exec` and `exec resume`); `ignoreUserConfig` → `--ignore-user-config`, which stops `$CODEX_HOME/config.toml` from deciding an orchestrated run's model behind the caller's back (auth still resolves from `CODEX_HOME`); `addDir` → `--add-dir` on the first turn only, since `exec resume` rejects it and the resumed thread keeps the roots it opened with
|
|
66
67
|
- `codexProfile` → `--profile <name>` (named config profile from `~/.codex/config.toml`)
|
|
67
68
|
- Per-session continuity: the `thread_id` from the first turn's `thread.started` event is captured and reused via `codex exec resume <id>` for subsequent sends, so the model sees prior turns
|
|
68
69
|
- `sandboxMode` maps to `--sandbox <mode>` on the first turn. **A resumed thread does not inherit it**, and `codex exec resume` rejects `--sandbox`, so the policy is restated as `-c sandbox_mode="<mode>"` on every resume. Without that, a `read-only` session goes writable from its second turn onward — verified against 0.146.0, where such a session wrote to disk on turn 2 on every attempt. Re-probed on 0.147.0 (direct write, shell redirect and delegate-to-subagent, each on a resumed turn): no writes
|
|
@@ -128,8 +129,8 @@ in print mode. Verified against `agy` **1.1.13**.
|
|
|
128
129
|
externally via `resumeSessionId` (bare UUID only); read it back from
|
|
129
130
|
`getStats().agyConversationId`.
|
|
130
131
|
- **Reasoning effort**: session `effort` and per-turn `session_send` overrides map
|
|
131
|
-
to `--effort`. agy accepts `low`, `medium`, and `high`;
|
|
132
|
-
`xhigh`
|
|
132
|
+
to `--effort`. agy accepts `low`, `medium`, and `high`; everything above that
|
|
133
|
+
(`xhigh`, `max`, `ultra`) clamps to `high`. agy 1.1.21 requires an effort with unsuffixed base
|
|
133
134
|
slugs such as `gemini-3.7-flash`, so `auto` resolves those to `high`; a model
|
|
134
135
|
already ending in `-low`, `-medium`, or `-high` keeps that qualified effort.
|
|
135
136
|
Per-turn overrides also work with qualified slugs: the adapter removes a
|
|
@@ -197,7 +198,8 @@ prints a single JSON object and exits. Verified against `grok` **1.0.5**.
|
|
|
197
198
|
because the same-looking field on codex is a running total.
|
|
198
199
|
- Permission modes pass straight through: grok's `--permission-mode` takes the same vocabulary we
|
|
199
200
|
use. The one exception is our `manual`, which grok spells `default`.
|
|
200
|
-
- Reasoning effort maps to `--effort`; grok accepts `low|medium|high
|
|
201
|
+
- Reasoning effort maps to `--effort`; grok 1.0.5 accepts `low|medium|high|xhigh` (it names the set in
|
|
202
|
+
its own rejection message), so only `max` and `ultra` clamp — to `xhigh`.
|
|
201
203
|
- **`sandboxMode: 'read-only'` is refused, not approximated.** grok has `--permission-mode plan` and
|
|
202
204
|
`--deny` rules, but plan mode alone is model-cooperative — the shape that let an adversarial
|
|
203
205
|
prompt write through Cursor's plan mode — and the deny rules have not been through the
|
|
@@ -261,7 +263,8 @@ Wraps the [sst/opencode](https://github.com/sst/opencode) CLI with `run --format
|
|
|
261
263
|
- NDJSON event stream with envelope `{ type, timestamp, sessionID, ... }`
|
|
262
264
|
- Event types: `text`, `reasoning`, `tool_use`, `step_start`, `step_finish`, `error`
|
|
263
265
|
- `text` and `tool_use` are **cumulative snapshots** keyed by `part.id` / `part.callID`; the wrapper diffs them to produce streaming deltas for `onText` callbacks and counts each tool invocation once
|
|
264
|
-
- Real token counts from `step_finish.part.tokens.{input,output,cache.read}`
|
|
266
|
+
- Real token counts from `step_finish.part.tokens.{input,output,cache.{read,write}}`. **`input` is the uncached remainder only** — opencode's own `total` is `input + output + cache.read + cache.write` — so the cached part is billed on top of it, not carved out of it, and `contextPercent` is measured against the whole input side. On a resumed turn the remainder is tiny next to the cached part (58 against 26,240), which is what makes the distinction matter
|
|
267
|
+
- Reasoning effort maps to `--variant`, opencode's provider-specific effort knob. opencode does not validate the value: a level the provider does not offer runs the turn at its default rather than failing
|
|
265
268
|
- The wrapper closes the subprocess's stdin immediately after spawn (opencode otherwise reads stdin and blocks on EOF, hanging the call)
|
|
266
269
|
- Provider-agnostic: opencode's `--model` expects `provider/model` form (e.g. `anthropic/claude-sonnet-4`). The wrapper passes `--model` through only when the value contains a `/`; otherwise opencode's own default applies
|
|
267
270
|
- `sandboxMode: 'read-only'` spawns a generated `clawo-readonly` agent (`--agent clawo-readonly` plus an `OPENCODE_CONFIG_CONTENT` env var defining it) that denies `edit` / `bash` / `external_directory` / `webfetch` / **`task`** at the permission level and additionally removes those tools outright via the agent's `tools` map. It deliberately does **not** use OpenCode's built-in `plan` agent: that is a user-overridable preset whose compiled rules start with `{"permission":"*","action":"allow"}` and deny neither `bash` nor `edit`, so a "read-only" session could still author files through a shell heredoc. **`task` is the load-bearing denial**: denying only the write tools leaves the delegation path open, and the agent will hand the write to a subagent that runs under the default writable agent — asked to delegate, a session denied only `edit`/`bash`/`external_directory` wrote to disk on every attempt. Verify this config only with adversarial writes, and include prompts that ask the agent to delegate; `opencode agent list` renders compiled permission rules that look identical for a safe and an unsafe agent, and a probe that only asks for a direct write passes even when the delegation path is wide open
|
|
@@ -30,22 +30,22 @@ every engine passes through.
|
|
|
30
30
|
|
|
31
31
|
### Row schema
|
|
32
32
|
|
|
33
|
-
| Field | Meaning
|
|
34
|
-
| ----------------------------------------- |
|
|
35
|
-
| `ts` | ISO timestamp of turn completion
|
|
36
|
-
| `session` | SessionManager session name
|
|
37
|
-
| `engine` | `claude` / `codex` / `codex-app` / `grok` / `opencode` / `agy` / `custom`
|
|
38
|
-
| `model` | Configured model, or the engine's own reported model when none was set
|
|
39
|
-
| `cwd` | Working directory the turn ran in
|
|
40
|
-
| `turn` | 1-based turn index within the session
|
|
41
|
-
| `tokensIn` / `tokensOut` / `cachedTokens` | **Per-turn deltas**, not session totals
|
|
33
|
+
| Field | Meaning |
|
|
34
|
+
| ----------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
35
|
+
| `ts` | ISO timestamp of turn completion |
|
|
36
|
+
| `session` | SessionManager session name |
|
|
37
|
+
| `engine` | `claude` / `codex` / `codex-app` / `grok` / `opencode` / `agy` / `custom` |
|
|
38
|
+
| `model` | Configured model, or the engine's own reported model when none was set |
|
|
39
|
+
| `cwd` | Working directory the turn ran in |
|
|
40
|
+
| `turn` | 1-based turn index within the session |
|
|
41
|
+
| `tokensIn` / `tokensOut` / `cachedTokens` | **Per-turn deltas**, not session totals |
|
|
42
42
|
| `costUsd` | Per-turn delta in USD — token count × the registry rate. On a flat-rate plan (a subscription seat) nobody was billed this: set `pricingOverrides` in the plugin config to zero the model out (`{"gpt-5.5": {"input": 0, "output": 0, "cached": 0}}`). Read [Zeroing a model's pricing](#zeroing-a-models-pricing) first — it disables `maxBudgetUsd` for that model, except on `grok`, which prices itself |
|
|
43
|
-
| `tokensEstimated` | `true` when the counts came from `estimateTokens()` (see below)
|
|
44
|
-
| `durationMs` | Wall-clock for the turn
|
|
45
|
-
| `toolCalls` / `toolErrors` | Per-turn deltas
|
|
46
|
-
| `ok` | `false` for a turn that threw, or that the session's own `turnsSucceeded` counter did not count (see `sessions.md`). Falls back to "nothing was thrown" when the counter cannot be read
|
|
47
|
-
| `error` | Failure text, truncated to 500 chars. Absent when the turn resolved but the engine did not count it as succeeded (an interrupted or non-SUCCESS turn), so a failed row does not always carry one
|
|
48
|
-
| `parent` | council id / fanout id / autoloop run id, when the turn belongs to one
|
|
43
|
+
| `tokensEstimated` | `true` when the counts came from `estimateTokens()` (see below) |
|
|
44
|
+
| `durationMs` | Wall-clock for the turn |
|
|
45
|
+
| `toolCalls` / `toolErrors` | Per-turn deltas |
|
|
46
|
+
| `ok` | `false` for a turn that threw, or that the session's own `turnsSucceeded` counter did not count (see `sessions.md`). Falls back to "nothing was thrown" when the counter cannot be read |
|
|
47
|
+
| `error` | Failure text, truncated to 500 chars. Absent when the turn resolved but the engine did not count it as succeeded (an interrupted or non-SUCCESS turn), so a failed row does not always carry one |
|
|
48
|
+
| `parent` | council id / fanout id / autoloop run id, when the turn belongs to one |
|
|
49
49
|
|
|
50
50
|
Deltas rather than totals means summing a query window gives that window's spend
|
|
51
51
|
without double-counting.
|
|
@@ -138,16 +138,41 @@ Where the engine reports usage, those counts are the engine's own. Where it does
|
|
|
138
138
|
not, the wrapper falls back to `estimateTokens()` (characters ÷ 4) and the row is
|
|
139
139
|
flagged `tokensEstimated: true`; the CLI marks those costs with a trailing `~`.
|
|
140
140
|
|
|
141
|
-
| Engine | Token counts
|
|
142
|
-
| ----------------- |
|
|
143
|
-
| `claude` | Engine-reported
|
|
144
|
-
| `codex` | Engine-reported
|
|
145
|
-
| `codex-app` | Engine-reported
|
|
146
|
-
| `grok` | Engine-reported — and so is the **cost**: this engine reports `total_cost_usd`, which the wrapper passes through instead of pricing tokens from the registry, so registry drift cannot affect a grok row
|
|
147
|
-
| `cursor` (legacy) | Engine-reported when the stream carries `usage`, else estimated
|
|
148
|
-
| `opencode` | Engine-reported when the run JSON carries `tokens`, else estimated
|
|
149
|
-
| `agy` | Engine-reported when the result event carries usage, else estimated
|
|
150
|
-
| `custom` | Depends on the CLI; estimated when it emits no usage
|
|
141
|
+
| Engine | Token counts |
|
|
142
|
+
| ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
143
|
+
| `claude` | Engine-reported — and so is the **cost**: the `result` event's `total_cost_usd` is taken as-is, so registry drift cannot affect a Claude row. It is the session running total rather than the turn's, so spend advances by the difference between turns. Proxy sessions (`baseUrl` set) are the exception: the CLI is told it is running `opus` while another provider serves the tokens, so its figure is Opus list price for someone else's model and the registry estimate is used instead |
|
|
144
|
+
| `codex` | Engine-reported |
|
|
145
|
+
| `codex-app` | Engine-reported |
|
|
146
|
+
| `grok` | Engine-reported — and so is the **cost**: this engine reports `total_cost_usd`, which the wrapper passes through instead of pricing tokens from the registry, so registry drift cannot affect a grok row |
|
|
147
|
+
| `cursor` (legacy) | Engine-reported when the stream carries `usage`, else estimated |
|
|
148
|
+
| `opencode` | Engine-reported when the run JSON carries `tokens`, else estimated |
|
|
149
|
+
| `agy` | Engine-reported when the result event carries usage, else estimated |
|
|
150
|
+
| `custom` | Depends on the CLI; estimated when it emits no usage |
|
|
151
|
+
|
|
152
|
+
### What "input tokens" means is not the same on every engine
|
|
153
|
+
|
|
154
|
+
Two engines can report `input` and `cached` and mean different things by them,
|
|
155
|
+
and the difference decides whether the cost math may subtract one from the other:
|
|
156
|
+
|
|
157
|
+
| Engine | Its own arithmetic for one turn | `input` contains cached reads |
|
|
158
|
+
| ---------- | ---------------------------------------------------- | ----------------------------- |
|
|
159
|
+
| `codex` | `total 19704 = input 19699 + output 5` | yes |
|
|
160
|
+
| `grok` | `total 30034 = input 19393 + output 17 + read 10624` | no |
|
|
161
|
+
| `opencode` | `total 26315 = input 58 + output 17 + read 26240` | no |
|
|
162
|
+
| `claude` | `input_tokens 2` against `cache_read 47371` | no |
|
|
163
|
+
|
|
164
|
+
On an engine that excludes them, cached reads and cache writes are billed **on
|
|
165
|
+
top of** `input`, and the prompt the turn actually carried is the sum of all
|
|
166
|
+
three — which is what `contextPercent` measures. Reading `input` alone reports a
|
|
167
|
+
nearly-full context as empty on any resumed conversation, because the history
|
|
168
|
+
arrives as cached reads.
|
|
169
|
+
|
|
170
|
+
Anthropic bills cache _writes_ above the input rate, and the premium depends on
|
|
171
|
+
the TTL (1.25x for the 5-minute cache, 2x for the 1-hour one); the Claude
|
|
172
|
+
estimate prices both tiers from the split the engine reports. Elsewhere cache
|
|
173
|
+
writes are priced at the plain input rate, which is a floor rather than an exact
|
|
174
|
+
figure — one more reason to prefer the engine's own `total_cost_usd` where it
|
|
175
|
+
offers one.
|
|
151
176
|
|
|
152
177
|
So on an estimating engine the cap is best-effort. It will stop a runaway session;
|
|
153
178
|
it is not an accounting guarantee, and it is not a substitute for the spend limits
|
|
@@ -135,6 +135,234 @@ session (in that mode the tool list is deliberately left out of the session hash
|
|
|
135
135
|
It costs the per-turn growth described above — only use it if the tool set really
|
|
136
136
|
does change mid-conversation.
|
|
137
137
|
|
|
138
|
+
## Conversation history on the way in
|
|
139
|
+
|
|
140
|
+
The caller's `messages[]` can carry the whole conversation: earlier `user` turns and the engine's
|
|
141
|
+
own earlier `assistant` replies. Whether those turns need to be sent is the same question the
|
|
142
|
+
section below asks about tool results — does the engine's own conversation already hold them? — and
|
|
143
|
+
it gets the same answer, from the same predicate. The turns that are in scope are serialized into
|
|
144
|
+
one `<conversation_history>` block of `<user>` / `<assistant>` turns and put in front of the
|
|
145
|
+
caller's latest `user` text. The wrapper tag is the one `renderHistory()` in the autoloop dispatcher
|
|
146
|
+
already uses for the same job; the per-turn tags are not — that one labels its two speakers `<user>`
|
|
147
|
+
/ `<agent>`, because its roles are autoloop roles rather than OpenAI wire roles.
|
|
148
|
+
|
|
149
|
+
| On this turn the engine | What is sent |
|
|
150
|
+
| ----------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------ |
|
|
151
|
+
| is **not** resuming a conversation — no session yet, a session that never announced a conversation id, or one being stopped and recreated | every `user`/`assistant` turn except the caller's latest `user` message |
|
|
152
|
+
| **is** resuming a live conversation, but one this bridge never sent these turns to | the same — a live thread is not automatically _this_ thread |
|
|
153
|
+
| **is** resuming the live conversation these turns belong to | nothing — the turns are already in the transcript, and the text goes out alone |
|
|
154
|
+
|
|
155
|
+
The last row is what keeps Anthropic prompt caching warm on `claude` and keeps a resumed `codex`
|
|
156
|
+
thread from being re-fed its own history: on a live thread the message is byte-identical to what it
|
|
157
|
+
was before this block existed. The first row is the one that was losing data. A client that opens a
|
|
158
|
+
new conversation per turn — one whose session key hashes the last message, say — lands in it on
|
|
159
|
+
every turn, and `skipPersistence: true` means an OpenAI-compat session is never auto-resumed from
|
|
160
|
+
disk either, so a follow-up like "yes, go ahead" used to reach the engine with nothing in front of
|
|
161
|
+
it.
|
|
162
|
+
|
|
163
|
+
The middle row is the one that is easy to get wrong. "Is there a live thread under this session
|
|
164
|
+
name?" is not "is that thread holding this conversation?", and the two come apart constantly. A
|
|
165
|
+
caller whose session key hashes its latest message resolves every repeat of a short confirmation —
|
|
166
|
+
"yes, go ahead", typed all day by someone approving invoices — to the session some _earlier_
|
|
167
|
+
invoice opened. A client that sends no session key at all falls back to a hash of
|
|
168
|
+
model+system+tools, which is stable for every turn AND identical between two concurrent chats of
|
|
169
|
+
the same user, so both chats share one session. And on `claude`, the default engine,
|
|
170
|
+
`nativeThreadIsLive()` has no id to check and returns true for anything in the session map, so
|
|
171
|
+
there the question collapses to "does the name exist" with no gate at all.
|
|
172
|
+
|
|
173
|
+
So the bridge records, per session, a fingerprint of the `user` turns it has actually pushed there,
|
|
174
|
+
and replays unless this request continues exactly that. It is the only writer to these sessions, so
|
|
175
|
+
what it sent is what the engine holds. Unknown session, forked conversation, evicted entry: all
|
|
176
|
+
replay. Being wrong in that direction costs a duplicated turn under a framing that says not to act
|
|
177
|
+
on it twice; being wrong in the other direction drops context silently, which is the class of bug
|
|
178
|
+
this exists to fix.
|
|
179
|
+
|
|
180
|
+
What the request is compared against depends on where its array ends, and only an array ending in a
|
|
181
|
+
`user` turn carries a turn the bridge has not sent yet. For that shape the fingerprint of every
|
|
182
|
+
`user` turn _before_ the last one is what the thread should be holding. For every other shape — a
|
|
183
|
+
tool-loop hop ending in `tool`, a prefill/continue ending in `assistant` — the latest `user` turn
|
|
184
|
+
was already pushed on an earlier request, so the whole array is compared. Getting that wrong in the
|
|
185
|
+
other direction is not harmless: comparing a hop against one turn less never matches, and the
|
|
186
|
+
transcript is then replayed into the very session that is already holding it, on every hop, for a
|
|
187
|
+
caller that needed none of this.
|
|
188
|
+
|
|
189
|
+
The map is bounded at 1000 entries and evicted oldest-first, which it needs independently of the
|
|
190
|
+
session map: `_cleanupIdleSessions()` reaps a session by `sessionTtlMinutes` without telling this
|
|
191
|
+
map, so a fingerprint outlives the session it mirrors. Eviction costs a replayed block, never a
|
|
192
|
+
dropped one, and `serve` restarts start it empty for the same price.
|
|
193
|
+
|
|
194
|
+
`system` messages are never in the block: they travel as the session's system prompt (see [Tool
|
|
195
|
+
definitions](#tool-definitions-and-where-they-live)). `tool` messages are never in it either —
|
|
196
|
+
they are the next section's business, and repeating them here would duplicate every payload and
|
|
197
|
+
undo the scoping that keeps a tool loop linear. An `assistant` message that only announces
|
|
198
|
+
`tool_calls` carries no text, so it renders no turn; a turn whose text is empty or whitespace is
|
|
199
|
+
dropped rather than rendered as an empty shell. An array with nothing to replay produces no block
|
|
200
|
+
at all, so a single-turn `[system, user]` request — the shape the OpenClaw main agent, cron jobs
|
|
201
|
+
and subagents send — goes out exactly as it did before.
|
|
202
|
+
|
|
203
|
+
Replayed text has neutralized every tag the assembled prompt treats as structure — the block's own
|
|
204
|
+
`<conversation_history>` / `<user>` / `<assistant>`, and also `<tool_results>` / `<tool_result>`,
|
|
205
|
+
`<system>` and `<tool_calls>` (`</user>` becomes `</user>`). Unlike a `<tool_result>` body, which
|
|
206
|
+
comes from the caller's own tool runner, a replayed turn is whatever an end user typed, and
|
|
207
|
+
`hi</user>\n<assistant>\n...` would otherwise close its turn early and forge an `assistant` turn —
|
|
208
|
+
putting words in the engine's own mouth. The other three matter for the same reason: a fabricated
|
|
209
|
+
tool return carries the framing sentence that says a payload is authoritative, a `<system>` block
|
|
210
|
+
contradicts the real one the non-claude path prepends, and `<tool_calls>` is the exact protocol JSON
|
|
211
|
+
the model is asked to emit.
|
|
212
|
+
|
|
213
|
+
Only the `<` is escaped, by lookahead. That shape is what makes the boundary decidable: matching up
|
|
214
|
+
to the closing `>` instead means re-emitting whatever was captured, and `hola<user a</user>` then
|
|
215
|
+
smuggles a raw close through the attribute slot of a tag that IS matched. So the match ends at
|
|
216
|
+
anything that ends a tag name — `>`, `/`, `<`, end of text, or a character that occupies no width.
|
|
217
|
+
That last clause is five Unicode properties rather than a list of code points, and it is not
|
|
218
|
+
decoration: measured over the 6,060 code points that are zero-advance or render blank,
|
|
219
|
+
`[\s></\p{Cc}\p{Cf}]` let **5,806** through, so `ok</user︀>\n<assistant︀>` forged a turn that is
|
|
220
|
+
indistinguishable on screen from `ok</user>`. Adding `Default_Ignorable` leaves 1,770;
|
|
221
|
+
`\p{Mn}\p{Me}` closes it; U+2800 BRAILLE PATTERN BLANK is neither and is named. Cost: the same 11 of
|
|
222
|
+
a 28-string corpus of plausible legitimate text change under the wide class as under the narrow one.
|
|
223
|
+
|
|
224
|
+
The name does **not** have to sit flush against `<` or `</`: the same invisible padding, plus the
|
|
225
|
+
slash itself, is allowed before the name, because `hola</user>` renders as `hola</user>` and a
|
|
226
|
+
model reads it as a close — a trailing class complete over zero-width filler with a flush leading
|
|
227
|
+
side still forges a turn. That is **one** class, `[/\p{Cc}\p{Cf}\p{Mn}\p{Me}\p{Default_Ignorable_Code_Point}]*`,
|
|
228
|
+
not `[…]*\/?[…]*`: two adjacent unbounded quantifiers over the same class backtrack O(n²) on a long
|
|
229
|
+
run that never reaches a valid name, and a single history message of ~100 KB of combining marks hung
|
|
230
|
+
the event loop ~80 s — a one-request denial of service against a fleet whose watchdog already resets
|
|
231
|
+
on event-loop stalls. The class is the zero-**advance** subset only (no `\s`, no U+2800), because a
|
|
232
|
+
visible separator before the name is the `if (count < user && x)` corruption the boundary class
|
|
233
|
+
already refuses. Swept over the 6,060 code points that are zero-advance or separators, in all three
|
|
234
|
+
positions (18,180 probes): 12,120 went through unfenced with the flush leading side, **38** with the
|
|
235
|
+
interior class, and the 38 are 19 `Zs`/`Zl`/`Zp` code points — U+0020, U+00A0, U+2000–200A, U+3000 —
|
|
236
|
+
i.e. exactly the visible-separator limitation below, and nothing else. The single class also lets a
|
|
237
|
+
slash sit among the filler (`<//user`, `</␀/user`); harmless, since the only action is escaping the
|
|
238
|
+
`<`. Filler INSIDE the name (`</us␀er>`) is still not fenced — the old regex missed it too.
|
|
239
|
+
|
|
240
|
+
**What it does not promise.** A positive class cannot be complete over a _visible_ separator, so
|
|
241
|
+
`</ user>` and `< assistant>` go through raw — a model reads them as a boundary. They are excluded on
|
|
242
|
+
cost: reaching them means corrupting `if (count < user && x)` and `the < user > column`. What already
|
|
243
|
+
gets corrupted for the same reason, since the class contains whitespace: `Promise<User | null>` and
|
|
244
|
+
`count<user && total>limit` come out with `<`. Readable to a model, not byte-identical. And the
|
|
245
|
+
body of a `<tool_result>` is never fenced — its content comes from the caller's own tool runner — so a
|
|
246
|
+
tool return that embeds a whole `<conversation_history>` block is not stopped here.
|
|
247
|
+
|
|
248
|
+
The caller's **latest** `user` turn is fenced too, but only on turns that actually carry a block.
|
|
249
|
+
That turn is the one input an attacker controls end to end, and the block teaches the model in the
|
|
250
|
+
same prompt that `<conversation_history>` holds its own earlier turns — unfenced, it could close the
|
|
251
|
+
real block and open a second one indistinguishable from it. With no block in front of it there is
|
|
252
|
+
nothing to forge, so those turns stay byte-for-byte what they were.
|
|
253
|
+
|
|
254
|
+
A `user` turn carrying only non-text content (an image) renders as `[non-text content]` rather than
|
|
255
|
+
vanishing. Dropping it would leave the `assistant` reply to it standing alone under a framing that
|
|
256
|
+
calls the assistant turns the model's own — a reply to a request the model cannot see, which reads
|
|
257
|
+
as license to act on the reply by itself. "Photo of the invoice", then "yes, go ahead", is an
|
|
258
|
+
everyday shape. A leading `assistant` turn with no `user` turn in front of it at all (content
|
|
259
|
+
`null`, or empty) is dropped instead, since no marker could honestly stand in for it.
|
|
260
|
+
|
|
261
|
+
### What this does not cover
|
|
262
|
+
|
|
263
|
+
- **A turn can be replayed that the engine already had.** The mirror image of the duplicate-once
|
|
264
|
+
trade below, and it comes from the same place: the engine's state is read from the session, not
|
|
265
|
+
from the array. A client whose session looks new to the bridge but whose engine did hold context
|
|
266
|
+
gets those turns a second time. The framing sentences tell the model these are earlier turns and
|
|
267
|
+
not to act on them again, which is what keeps a duplicated turn from becoming a duplicated
|
|
268
|
+
action; nothing enforces it.
|
|
269
|
+
- **The block is not in strict chronological order when the array does not end in the caller's
|
|
270
|
+
latest `user` turn.** Every `user`/`assistant` turn except that one is replayed, including turns
|
|
271
|
+
that come after it — an array ending in `assistant` (prefill, an explicit "continue", a framework
|
|
272
|
+
appending its own reply) keeps that turn, because a transcript presented as complete while
|
|
273
|
+
missing the last thing the model said invites it to redo the work. The cost is that such a turn
|
|
274
|
+
is rendered inside the block, i.e. before the caller's latest text rather than after it.
|
|
275
|
+
- **`X-Session-Reset` is not honored on an array ending in a `tool` result**, so on that one shape a
|
|
276
|
+
reset turn on a live thread gets no history block either. Same pre-existing asymmetry the tool
|
|
277
|
+
results have there, from the same cause — the header is parsed after that branch returns. See the
|
|
278
|
+
note under [Tool definitions](#tool-definitions-and-where-they-live).
|
|
279
|
+
- **`grok` inherits the hole described in the next section**: it resumes by id, but its id is absent
|
|
280
|
+
from `SessionStats`, so `nativeThreadIsLive()` reaches `default: return true` and a `grok` session
|
|
281
|
+
whose first turn died before emitting its id reports a live thread — and so gets no history. Same
|
|
282
|
+
cause, same fix (adding the field), not addressed here.
|
|
283
|
+
- **The two blocks are not interleaved.** When a request carries both, the message is the history
|
|
284
|
+
block, then the tool results, then the caller's new text — chronological between the blocks, but
|
|
285
|
+
a `tool` result that chronologically preceded a replayed `assistant` turn still appears after it.
|
|
286
|
+
- **The block is capped at 24,000 characters — the whole block, not the sum of the turn text.**
|
|
287
|
+
Wrapper tags, per-turn tags, elision markers and the 268 characters of framing are all charged to
|
|
288
|
+
the budget before any turn is. That is the fix for a cap that did not cap: charging only the turn
|
|
289
|
+
text left the retained turn count bounded by nothing but `24,000 / mean-turn-length`, and 8,000
|
|
290
|
+
alternating one-word turns (`ok`, `sí`, `dale`) rendered a **165,008-byte** block — 33 KiB past the
|
|
291
|
+
argv ceiling below, from a request no bigger than a chat backlog. The elision markers were worse:
|
|
292
|
+
32 characters each, added AFTER the arithmetic, once per truncated turn, which is how a "24,000
|
|
293
|
+
cap" emitted more than it said.
|
|
294
|
+
|
|
295
|
+
Oldest turns are dropped first, and the turn the budget runs out inside is truncated (head kept)
|
|
296
|
+
and marked `[… turn truncated for length …]` rather than dropped whole — dropping it whole would
|
|
297
|
+
take the request with it on the pasted-document shape. The marker comes out of that turn's own
|
|
298
|
+
allowance, so a truncated turn cannot push the block over. Below 200 rendered characters of
|
|
299
|
+
remaining room a turn is not started at all: it and the turns behind it are dropped, and the
|
|
300
|
+
remainder goes unspent, because a fragment that short is a sentence with its qualifier cut off
|
|
301
|
+
rather than context.
|
|
302
|
+
|
|
303
|
+
The newest `user` turn in the block — the ask the turns after it answer — is the one exception to
|
|
304
|
+
spending newest-first: the turns after it (all `assistant`, by definition of "newest `user` turn")
|
|
305
|
+
spend against the budget minus its frame plus `min(its length, 200)`. That reserve is the fix for
|
|
306
|
+
a drop, not a refinement. Without it, replies that add past the cap (one 30k pasted listing, or
|
|
307
|
+
two ordinary 12k ones) take the whole budget, the window starts past every `user` turn, the
|
|
308
|
+
leading-`assistant` rule clears what is left, and NO block goes out at all: the caller's latest
|
|
309
|
+
turn reaches the engine alone, which is this block's own failure mode at its worst.
|
|
310
|
+
|
|
311
|
+
The 200-character floor applies **above** the anchor too, and that is the second half of the cap
|
|
312
|
+
fix. Once the post-anchor turns have truncated the budget down near the reserve, an older turn is
|
|
313
|
+
started with a `room` smaller than the 32-character elision marker, and `slice(0, room - 32)` with
|
|
314
|
+
a negative argument slices from the END of the string — emitting nearly the whole turn while
|
|
315
|
+
charging the budget only `room`, which blows the cap. Measured on a narrating tool loop, where each
|
|
316
|
+
hop's `assistant` narration becomes a consecutive post-anchor turn because `tool` messages are
|
|
317
|
+
filtered out: with the floor removed, 30 hops render 27,627 characters and the three-`user` /
|
|
318
|
+
with-reply sweep shapes reach 28,003 / 30,003. With the floor, that run of older turns is dropped
|
|
319
|
+
instead, which the anchor's reserve makes safe: ending the window there would drop the anchor and
|
|
320
|
+
hand the leading-`assistant` rule an all-`assistant` list to clear, i.e. the empty block again.
|
|
321
|
+
|
|
322
|
+
Verification. 44,000 random shapes across three message-count ranges (≤7, ≤60 and ≤400 messages,
|
|
323
|
+
roles `user`/`assistant`/`system`/`tool`, content string/whitespace/null/array/empty-array/no-text,
|
|
324
|
+
lengths straddling 0/1/199/200/201/11,999/12,000/12,001/23,799/23,800/24,000/24,001/30,000/60,000):
|
|
325
|
+
**max block 23,999 characters, zero shapes over 24,000, zero content-free turns, zero blocks with
|
|
326
|
+
no `user` turn in them, and zero shapes that went from a non-empty block to an empty one.** 7,488
|
|
327
|
+
of them went the other way — empty before, non-empty now — which is the anchor reserve doing its
|
|
328
|
+
job. 11,954 non-empty blocks changed, which is the point: the old arithmetic charged less than it
|
|
329
|
+
emitted, so every block near the ceiling gets shorter. Directed shapes: 30/100/2,000 narrating
|
|
330
|
+
hops and 1,200/8,000/24,000 one-word turns are all ≤ 24,000 with no content-free turns.
|
|
331
|
+
|
|
332
|
+
The ceiling is on characters and argv counts bytes, which is the conservative direction only up to
|
|
333
|
+
a point: 24,000 characters of astral-plane text is 48,000 bytes and the worst case (3-byte BMP) is
|
|
334
|
+
72,000, both still inside 128 KiB, but the
|
|
335
|
+
block is not the whole prompt — the tool block, the `<system>` prepend and the caller's own turn
|
|
336
|
+
are added after it. What the cap bounds is the part that scales with the transcript.
|
|
337
|
+
|
|
338
|
+
`MAX_BODY_SIZE` (5 MiB) is **not** a usable bound here: six of the nine `ENGINE_TYPES` pass the
|
|
339
|
+
prompt to the CLI as a single argv element (`codex`, `gemini`, `agy`, `cursor`, `grok`,
|
|
340
|
+
`opencode`), and so does a one-shot `custom` engine; Linux caps one argument at `MAX_ARG_STRLEN` =
|
|
341
|
+
128 KiB whatever `getconf ARG_MAX` reports (measured: 131071 bytes spawns, 131072 throws `E2BIG`).
|
|
342
|
+
Only `claude`, `codex-app` and a persistent `custom` engine write over stdin. Uncapped, ordinary
|
|
343
|
+
traffic reaches that ceiling — 400-character turns with 900-character replies put the message at
|
|
344
|
+
131,063 characters at turn 96 and 132,425 at turn 97 — and the failure is a 500 with the turn
|
|
345
|
+
lost, which is worse than the missing context the block exists to restore. This is the same trade
|
|
346
|
+
`renderHistory()` makes, with the same `REPLAY_CHAR_BUDGET`, feeding the same engines.
|
|
347
|
+
- **A send that threw records nothing.** The fingerprint is written after the send and only when the
|
|
348
|
+
send landed, because the two ways to be wrong are not symmetric: forgetting a turn that landed
|
|
349
|
+
replays it once more, while assuming one landed that did not drops context silently. "Landed" is
|
|
350
|
+
`sendMessage` returning — a returned error is answered with 502 and still records, since the CLI
|
|
351
|
+
received the prompt — so only a throw withholds the record. An earlier revision of this file
|
|
352
|
+
described the opposite placement as deliberate and named its residue: a caller that answered a 5xx
|
|
353
|
+
by appending an `assistant` turn and sending again got the next turn bare. That was the bug, not
|
|
354
|
+
the design.
|
|
355
|
+
- **A second request that arrives while the first is still in flight replays.** It sees no
|
|
356
|
+
fingerprint yet, so the transcript goes out again into the session that already holds it. The safe
|
|
357
|
+
direction — a duplicate rather than a drop — plus a lost cache prefix.
|
|
358
|
+
- **Cost is O(n) per turn for engines that never resume.** `engineHasNativeConversation()` is false
|
|
359
|
+
for `gemini` and one-shot custom engines, so for them the block is re-serialized on every turn,
|
|
360
|
+
capped but never free. Same for any caller that mints a new session per turn.
|
|
361
|
+
- **`X-Session-Reset` now replays the transcript.** A reset turn means the engine holds nothing, so
|
|
362
|
+
the history goes out in full. Under the reading "the caller asked to start clean" that is the
|
|
363
|
+
opposite of what was asked, and a client that sends the header on every request AND re-sends
|
|
364
|
+
`messages[]` pays for the transcript every time.
|
|
365
|
+
|
|
138
366
|
## Tool results on the way back
|
|
139
367
|
|
|
140
368
|
A `tool` role message in the caller's array is the result of a call the model asked for on an
|
|
@@ -199,6 +427,28 @@ relying on it:
|
|
|
199
427
|
Neither applies to a caller that keeps its own transcript and forwards only the latest turn — it
|
|
200
428
|
sends one round at a time.
|
|
201
429
|
|
|
430
|
+
### A live session is not the same thing as this conversation
|
|
431
|
+
|
|
432
|
+
Suppressing the replay needs a stronger fact than "a session under this name is live". That is what
|
|
433
|
+
`nativeThreadIsLive()` reports, and a session name can be live while its transcript belongs to a
|
|
434
|
+
different exchange. Three shapes where the two come apart, all reachable with default settings:
|
|
435
|
+
|
|
436
|
+
| shape | what happens |
|
|
437
|
+
| ---------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
|
|
438
|
+
| a caller whose session key hashes its latest message | every repeat of the same short confirmation resolves to whichever session that phrase opened first — often a different subject entirely |
|
|
439
|
+
| a caller that sends no `X-Session-Id` at all | the key falls back to a hash of model + system prompt + tools, so all of that caller's concurrent chats share one name |
|
|
440
|
+
| `engine: 'claude'` — the default | `nativeThreadIsLive()` has no id to check and returns `true` for anything in the session map, so the name is the only evidence there is |
|
|
441
|
+
|
|
442
|
+
So the bridge tracks what it actually pushed into each session and replays whenever the incoming
|
|
443
|
+
conversation is not the one it remembers seeding. The fingerprint covers the `user` turns only: those
|
|
444
|
+
are the caller's own text echoed back verbatim, while assistant text is what the engine produced and
|
|
445
|
+
a client may normalize it. A mismatch replays, which is the safe direction — the cost is a repeated
|
|
446
|
+
block, never a lost one.
|
|
447
|
+
|
|
448
|
+
The map is bounded and evicted oldest-first, and it has to be independent of the session map:
|
|
449
|
+
`_cleanupIdleSessions()` reaps a session by TTL without telling it, so a fingerprint outlives the
|
|
450
|
+
session it mirrors. Losing an entry costs a replayed block, never a dropped one.
|
|
451
|
+
|
|
202
452
|
## Environment variables
|
|
203
453
|
|
|
204
454
|
| Variable | Default | Purpose |
|
|
@@ -33,7 +33,8 @@ Start a persistent coding session with full CLI flag support.
|
|
|
33
33
|
| `mcpConfig` | string \| string[] | MCP server config file(s) |
|
|
34
34
|
| `settings` | string | Settings.json path or inline JSON |
|
|
35
35
|
| `ultracode` | boolean | Claude only. Enable "ultracode" / dynamic workflows — Claude plans a JS orchestration script per substantive task and fans out to subagents. Injected as the `ultracode:true` settings key (merged into `settings`), **not** a `--effort` value (the CLI rejects `--effort ultracode`). |
|
|
36
|
-
| `noSessionPersistence` | boolean | Do not save session to disk — both the engine's own transcript and this orchestrator's resume registry, so a later start under the same name does not reattach
|
|
36
|
+
| `noSessionPersistence` | boolean | Do not save session to disk — both the engine's own transcript and this orchestrator's resume registry, so a later start under the same name does not reattach |
|
|
37
|
+
| `ignoreUserConfig` | boolean | Codex only. Run without loading `$CODEX_HOME/config.toml`, so an orchestrated run is decided by what the caller passed rather than by the machine's own Codex config — notably a `model = …` line in that file, which otherwise picks the model while the ledger records this engine's default. Auth still resolves from `CODEX_HOME`. |
|
|
37
38
|
| `betas` | string \| string[] | Custom beta headers |
|
|
38
39
|
| `enableAgentTeams` | boolean | Enable experimental agent teams |
|
|
39
40
|
| `enableAutoMode` | boolean | Enable auto permission mode |
|
|
@@ -54,7 +55,7 @@ Start a persistent coding session with full CLI flag support.
|
|
|
54
55
|
| `otelLogUserPrompts` | boolean | OpenTelemetry: log user prompts (sets `OTEL_LOG_USER_PROMPTS=1`) |
|
|
55
56
|
| `otelLogRawApiBodies` | boolean | OpenTelemetry: log raw API request/response bodies (sets `OTEL_LOG_RAW_API_BODIES=1`); debug only |
|
|
56
57
|
| `bedrockServiceTier` | `'default'` \| `'flex'` \| `'priority'` | AWS Bedrock service tier (sets `ANTHROPIC_BEDROCK_SERVICE_TIER`); only effective when routing through Bedrock |
|
|
57
|
-
| `effort` | `'low'` \| `'medium'` \| `'high'` \| `'xhigh'` \| `'max'` \| `'auto'`
|
|
58
|
+
| `effort` | `'low'` \| `'medium'` \| `'high'` \| `'xhigh'` \| `'max'` \| `'ultra'` \| `'auto'` | Reasoning effort level. Engines do not share one ladder: Claude Code takes up to `max`, Codex adds `ultra`, Grok stops at `xhigh`, Antigravity at `high`, and OpenCode forwards the level to its provider unvalidated. A level above an engine's ceiling clamps rather than being dropped. `high` and above trigger the `ultrathink` prefix on Claude user messages. |
|
|
58
59
|
|
|
59
60
|
### `session_send`
|
|
60
61
|
|