@enderfga/claw-orchestrator 6.0.4 → 6.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +4 -4
- package/dist/bin/acp-server.js +0 -0
- package/dist/bin/cli.js +0 -0
- package/dist/bin/mcp-server.js +20 -1
- package/dist/bin/mcp-server.js.map +1 -1
- package/dist/src/base-oneshot-session.d.ts +37 -0
- package/dist/src/base-oneshot-session.js +29 -5
- package/dist/src/base-oneshot-session.js.map +1 -1
- package/dist/src/index.js +3 -3
- package/dist/src/index.js.map +1 -1
- package/dist/src/models.js +39 -2
- package/dist/src/models.js.map +1 -1
- package/dist/src/persistent-agy-session.js +5 -4
- package/dist/src/persistent-agy-session.js.map +1 -1
- package/dist/src/persistent-codex-session.d.ts +3 -4
- package/dist/src/persistent-codex-session.js +38 -7
- package/dist/src/persistent-codex-session.js.map +1 -1
- package/dist/src/persistent-grok-session.js +18 -4
- package/dist/src/persistent-grok-session.js.map +1 -1
- package/dist/src/persistent-opencode-session.js +21 -2
- package/dist/src/persistent-opencode-session.js.map +1 -1
- package/dist/src/persistent-session.d.ts +74 -0
- package/dist/src/persistent-session.js +205 -28
- package/dist/src/persistent-session.js.map +1 -1
- package/dist/src/types.d.ts +10 -1
- package/dist/src/types.js.map +1 -1
- package/package.json +2 -1
- package/skills/SKILL.md +1 -1
- package/skills/references/claude-cli-tracking.md +2 -1
- package/skills/references/multi-engine.md +8 -5
- package/skills/references/observability.md +50 -25
- package/skills/references/tools.md +3 -2
|
@@ -2,11 +2,12 @@
|
|
|
2
2
|
|
|
3
3
|
This document tracks which Claude Code CLI version Claw Orchestrator is currently synced to, and which features have been integrated.
|
|
4
4
|
|
|
5
|
-
## Currently tracked: **Claude Code CLI 2.1.
|
|
5
|
+
## Currently tracked: **Claude Code CLI 2.1.246** (as of 2026-08-27, plugin v6.1.0)
|
|
6
6
|
|
|
7
7
|
## Sync history
|
|
8
8
|
|
|
9
9
|
| Plugin Version | Claude CLI Version | Date | Notable integrations |
|
|
10
|
+
| v6.1.0 | 2.1.246 | 2026-08-27 | **Engine sweep that turned into a token-accounting audit.** Claude Code 2.1.237→2.1.246, Codex 0.148.0→0.149.1, agy 1.1.15→1.1.21, OpenCode 1.18.18→1.18.23, Grok unchanged at 1.0.5; each ran a live turn and the ACP stdio entry was smoked. No new Claude Code flag needed integrating in that range — what did was what its `result` event has been reporting all along. Four measurement bugs fixed: a turn's usage was added twice (once on `message_delta`, once on `result` — engine said 2/4/47371, we said 4/8/94742); cache writes were absent from the cost formula (engine `$0.322428` vs our `$0.016` on a 1h-cache turn, and `maxBudgetUsd` gates on ours), so Claude now takes `total_cost_usd` — a session running total, applied as a difference; the `input − cached` subtraction is valid only on codex, whose `input_tokens` includes cached reads, while claude/grok/opencode report them alongside; and `contextPercent` read `input_tokens` alone, so a 47k prompt measured 0% — now the whole prompt over `modelUsage[*].contextWindow`. Also: Codex gained a real `max` and an `ultra` above it (we were folding `max`→`xhigh`), Grok now takes `xhigh` natively (we were folding it to `high`), OpenCode's `--variant` means `effort` finally reaches it, `--ephemeral`/`--ignore-user-config`/`--add-dir` wired up on Codex. Registry: `gemini-3.7-flash`/`gemini-3.6-flash`/`gpt-5.2` registered, `gemini-3.5-flash` repriced $0.5/$3 → $1.50/$9. |
|
|
10
11
|
| v4.8.0 | 2.1.207 | 2026-07-12 | **Autoloop role-level multi-engine support.** Planner, Coder, and Reviewer can select independent engines/models while preserving the Claude defaults. Built-in non-Claude Planners use native read-only/plan modes and receive their role protocol in-band. Spawn selections persist across resume; Codex persists its real thread ID. Runtime and invocation checks used Claude Code 2.1.207 and Codex 0.144.1. |
|
|
11
12
|
| v4.7.0 | 2.1.206 | 2026-07-10 | **Antigravity engine ships + permission-mode sync.** Main feature is the community-contributed first-class `engine: 'agy'` (PR #71, reviewed + hardened: layered resume-ID gating, `agy/` prefix routing, shared `sanitize.ts` across all six engines, ENGINE*TYPES single source). Weekly CLI sync: CC 2.1.200 renamed the `default` permission mode to **`manual`** — verified against 2.1.206 that the choices are now `acceptEdits/auto/bypassPermissions/manual/dontAsk/plan`, that `default` is still accepted (hidden compat), and that **`delegate` is hard-rejected at spawn** — so PermissionMode gains `manual`, drops `delegate`, and agy/gemini map `manual` like `default` (→ `--sandbox`). Codex 0.143.0: empirically re-tested `-c model_reasoning_effort=max` — still 400-rejected for gpt-5.5 (the "first-class max" note is Bedrock GPT-5.6-only), so the `max`→`xhigh` map stays. GPT-5.6 Sol/Terra/Luna registered with official pricing ($5/$30, $2.50/$15, $1/$6; 1M/1M/400K ctx) after the user reported using it — it's a limited preview on API/Codex-auth paths (empirically: ChatGPT-account Codex auth gets a 400, which is why the first probe on this box misread it as Bedrock-only; lesson — an auth-path rejection is not model non-existence). Codex default stays gpt-5.5. Free upside: CC 2.1.205 fixed `--json-schema` invalid-schema silent fallback + `format` keyword rejection; CC 2.1.203 fixed background sessions dropping shell-exported `ANTHROPIC_BASE_URL`. Pins → CC 2.1.206 / Codex 0.143.0 (installed; npm has 0.144.1, exec surface unchanged per release notes). |
|
|
12
13
|
| v4.6.0 | 2.1.199 | 2026-07-03 | **Model registry sync — Claude Fable 5.** Registered `claude-fable-5` (first Claude 5-family model, tier above Opus; standard $10/$50 per Mtok, cache read $1, full 1M context at standard rates per the official pricing page) with new `fable` alias; taught the `isClaudeModel`/`resolveProvider` heuristics to recognize `fable`/`mythos` strings (they only matched claude/opus/sonnet/haiku). Mythos 5 not listed (same price, limited availability). CC 2.1.198–199 are subagent/background-agent reliability fixes — no invocation-surface change; free upside for us: subagent partial output on rate-limit/server error is now returned instead of silently dropped, and API errors in subagents are reported to the parent. Codex unchanged (0.142.5 is a log-scrub patch; pin stays 0.142.4 as installed). |
|
|
@@ -62,7 +62,8 @@ Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested wi
|
|
|
62
62
|
- Real `usage` from the `turn.completed` JSON event (input, output, cached, reasoning tokens). **These are cumulative over the thread, not per turn** — three identical turns on 0.147.0 report `input_tokens` 13,856 → 27,727 → 41,613, each matching `total_token_usage` in that thread's rollout exactly. They are assigned to the session totals, never added; subtracting consecutive values recovers the turn's own prompt
|
|
63
63
|
- `contextPercent` is that per-turn prompt (which, for a thread-resuming engine, is the live context occupancy) over **codex's own limit**, harvested from the thread's rollout file (`model_context_window`, 258,400 on 0.147.0). The model registry holds the published window — 1,050,000 for gpt-5.x — which codex does not honour, so measuring against it reads ~4x low. Resuming a thread also seeds the token baseline from the rollout, so the first send does not mistake the whole thread history for one prompt. All of this is best-effort: an unreadable or `--ephemeral` thread falls back to the registry window
|
|
64
64
|
- `item.completed` parsing distinguishes `reasoning` / `todo_list` (logged, not counted) from real tool items (`command_execution`, `file_change`, `mcp_tool_call`, `web_search`, which increment `toolCalls`; a non-zero `command_execution.exit_code` increments `toolErrors`)
|
|
65
|
-
- Reasoning effort: the engine-agnostic `effort` maps to `-c model_reasoning_effort=<level>`
|
|
65
|
+
- Reasoning effort: the engine-agnostic `effort` maps to `-c model_reasoning_effort=<level>` and passes straight through. Codex 0.149's ladder runs `low|medium|high|xhigh|max|ultra` — it is the only engine here that reaches `ultra`, and all three top levels were exercised against 0.149.1. `auto` and `ultracode` are omitted. Note `-c` values are not validated at spawn: codex prints `reasoning effort: <whatever>` and sends it, so an unknown level fails at the API rather than at the command line
|
|
66
|
+
- `noSessionPersistence` → `--ephemeral` (accepted by `exec` and `exec resume`); `ignoreUserConfig` → `--ignore-user-config`, which stops `$CODEX_HOME/config.toml` from deciding an orchestrated run's model behind the caller's back (auth still resolves from `CODEX_HOME`); `addDir` → `--add-dir` on the first turn only, since `exec resume` rejects it and the resumed thread keeps the roots it opened with
|
|
66
67
|
- `codexProfile` → `--profile <name>` (named config profile from `~/.codex/config.toml`)
|
|
67
68
|
- Per-session continuity: the `thread_id` from the first turn's `thread.started` event is captured and reused via `codex exec resume <id>` for subsequent sends, so the model sees prior turns
|
|
68
69
|
- `sandboxMode` maps to `--sandbox <mode>` on the first turn. **A resumed thread does not inherit it**, and `codex exec resume` rejects `--sandbox`, so the policy is restated as `-c sandbox_mode="<mode>"` on every resume. Without that, a `read-only` session goes writable from its second turn onward — verified against 0.146.0, where such a session wrote to disk on turn 2 on every attempt. Re-probed on 0.147.0 (direct write, shell redirect and delegate-to-subagent, each on a resumed turn): no writes
|
|
@@ -128,8 +129,8 @@ in print mode. Verified against `agy` **1.1.13**.
|
|
|
128
129
|
externally via `resumeSessionId` (bare UUID only); read it back from
|
|
129
130
|
`getStats().agyConversationId`.
|
|
130
131
|
- **Reasoning effort**: session `effort` and per-turn `session_send` overrides map
|
|
131
|
-
to `--effort`. agy accepts `low`, `medium`, and `high`;
|
|
132
|
-
`xhigh`
|
|
132
|
+
to `--effort`. agy accepts `low`, `medium`, and `high`; everything above that
|
|
133
|
+
(`xhigh`, `max`, `ultra`) clamps to `high`. agy 1.1.21 requires an effort with unsuffixed base
|
|
133
134
|
slugs such as `gemini-3.7-flash`, so `auto` resolves those to `high`; a model
|
|
134
135
|
already ending in `-low`, `-medium`, or `-high` keeps that qualified effort.
|
|
135
136
|
Per-turn overrides also work with qualified slugs: the adapter removes a
|
|
@@ -197,7 +198,8 @@ prints a single JSON object and exits. Verified against `grok` **1.0.5**.
|
|
|
197
198
|
because the same-looking field on codex is a running total.
|
|
198
199
|
- Permission modes pass straight through: grok's `--permission-mode` takes the same vocabulary we
|
|
199
200
|
use. The one exception is our `manual`, which grok spells `default`.
|
|
200
|
-
- Reasoning effort maps to `--effort`; grok accepts `low|medium|high
|
|
201
|
+
- Reasoning effort maps to `--effort`; grok 1.0.5 accepts `low|medium|high|xhigh` (it names the set in
|
|
202
|
+
its own rejection message), so only `max` and `ultra` clamp — to `xhigh`.
|
|
201
203
|
- **`sandboxMode: 'read-only'` is refused, not approximated.** grok has `--permission-mode plan` and
|
|
202
204
|
`--deny` rules, but plan mode alone is model-cooperative — the shape that let an adversarial
|
|
203
205
|
prompt write through Cursor's plan mode — and the deny rules have not been through the
|
|
@@ -261,7 +263,8 @@ Wraps the [sst/opencode](https://github.com/sst/opencode) CLI with `run --format
|
|
|
261
263
|
- NDJSON event stream with envelope `{ type, timestamp, sessionID, ... }`
|
|
262
264
|
- Event types: `text`, `reasoning`, `tool_use`, `step_start`, `step_finish`, `error`
|
|
263
265
|
- `text` and `tool_use` are **cumulative snapshots** keyed by `part.id` / `part.callID`; the wrapper diffs them to produce streaming deltas for `onText` callbacks and counts each tool invocation once
|
|
264
|
-
- Real token counts from `step_finish.part.tokens.{input,output,cache.read}`
|
|
266
|
+
- Real token counts from `step_finish.part.tokens.{input,output,cache.{read,write}}`. **`input` is the uncached remainder only** — opencode's own `total` is `input + output + cache.read + cache.write` — so the cached part is billed on top of it, not carved out of it, and `contextPercent` is measured against the whole input side. On a resumed turn the remainder is tiny next to the cached part (58 against 26,240), which is what makes the distinction matter
|
|
267
|
+
- Reasoning effort maps to `--variant`, opencode's provider-specific effort knob. opencode does not validate the value: a level the provider does not offer runs the turn at its default rather than failing
|
|
265
268
|
- The wrapper closes the subprocess's stdin immediately after spawn (opencode otherwise reads stdin and blocks on EOF, hanging the call)
|
|
266
269
|
- Provider-agnostic: opencode's `--model` expects `provider/model` form (e.g. `anthropic/claude-sonnet-4`). The wrapper passes `--model` through only when the value contains a `/`; otherwise opencode's own default applies
|
|
267
270
|
- `sandboxMode: 'read-only'` spawns a generated `clawo-readonly` agent (`--agent clawo-readonly` plus an `OPENCODE_CONFIG_CONTENT` env var defining it) that denies `edit` / `bash` / `external_directory` / `webfetch` / **`task`** at the permission level and additionally removes those tools outright via the agent's `tools` map. It deliberately does **not** use OpenCode's built-in `plan` agent: that is a user-overridable preset whose compiled rules start with `{"permission":"*","action":"allow"}` and deny neither `bash` nor `edit`, so a "read-only" session could still author files through a shell heredoc. **`task` is the load-bearing denial**: denying only the write tools leaves the delegation path open, and the agent will hand the write to a subagent that runs under the default writable agent — asked to delegate, a session denied only `edit`/`bash`/`external_directory` wrote to disk on every attempt. Verify this config only with adversarial writes, and include prompts that ask the agent to delegate; `opencode agent list` renders compiled permission rules that look identical for a safe and an unsafe agent, and a probe that only asks for a direct write passes even when the delegation path is wide open
|
|
@@ -30,22 +30,22 @@ every engine passes through.
|
|
|
30
30
|
|
|
31
31
|
### Row schema
|
|
32
32
|
|
|
33
|
-
| Field | Meaning
|
|
34
|
-
| ----------------------------------------- |
|
|
35
|
-
| `ts` | ISO timestamp of turn completion
|
|
36
|
-
| `session` | SessionManager session name
|
|
37
|
-
| `engine` | `claude` / `codex` / `codex-app` / `grok` / `opencode` / `agy` / `custom`
|
|
38
|
-
| `model` | Configured model, or the engine's own reported model when none was set
|
|
39
|
-
| `cwd` | Working directory the turn ran in
|
|
40
|
-
| `turn` | 1-based turn index within the session
|
|
41
|
-
| `tokensIn` / `tokensOut` / `cachedTokens` | **Per-turn deltas**, not session totals
|
|
33
|
+
| Field | Meaning |
|
|
34
|
+
| ----------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
35
|
+
| `ts` | ISO timestamp of turn completion |
|
|
36
|
+
| `session` | SessionManager session name |
|
|
37
|
+
| `engine` | `claude` / `codex` / `codex-app` / `grok` / `opencode` / `agy` / `custom` |
|
|
38
|
+
| `model` | Configured model, or the engine's own reported model when none was set |
|
|
39
|
+
| `cwd` | Working directory the turn ran in |
|
|
40
|
+
| `turn` | 1-based turn index within the session |
|
|
41
|
+
| `tokensIn` / `tokensOut` / `cachedTokens` | **Per-turn deltas**, not session totals |
|
|
42
42
|
| `costUsd` | Per-turn delta in USD — token count × the registry rate. On a flat-rate plan (a subscription seat) nobody was billed this: set `pricingOverrides` in the plugin config to zero the model out (`{"gpt-5.5": {"input": 0, "output": 0, "cached": 0}}`). Read [Zeroing a model's pricing](#zeroing-a-models-pricing) first — it disables `maxBudgetUsd` for that model, except on `grok`, which prices itself |
|
|
43
|
-
| `tokensEstimated` | `true` when the counts came from `estimateTokens()` (see below)
|
|
44
|
-
| `durationMs` | Wall-clock for the turn
|
|
45
|
-
| `toolCalls` / `toolErrors` | Per-turn deltas
|
|
46
|
-
| `ok` | `false` for a turn that threw, or that the session's own `turnsSucceeded` counter did not count (see `sessions.md`). Falls back to "nothing was thrown" when the counter cannot be read
|
|
47
|
-
| `error` | Failure text, truncated to 500 chars. Absent when the turn resolved but the engine did not count it as succeeded (an interrupted or non-SUCCESS turn), so a failed row does not always carry one
|
|
48
|
-
| `parent` | council id / fanout id / autoloop run id, when the turn belongs to one
|
|
43
|
+
| `tokensEstimated` | `true` when the counts came from `estimateTokens()` (see below) |
|
|
44
|
+
| `durationMs` | Wall-clock for the turn |
|
|
45
|
+
| `toolCalls` / `toolErrors` | Per-turn deltas |
|
|
46
|
+
| `ok` | `false` for a turn that threw, or that the session's own `turnsSucceeded` counter did not count (see `sessions.md`). Falls back to "nothing was thrown" when the counter cannot be read |
|
|
47
|
+
| `error` | Failure text, truncated to 500 chars. Absent when the turn resolved but the engine did not count it as succeeded (an interrupted or non-SUCCESS turn), so a failed row does not always carry one |
|
|
48
|
+
| `parent` | council id / fanout id / autoloop run id, when the turn belongs to one |
|
|
49
49
|
|
|
50
50
|
Deltas rather than totals means summing a query window gives that window's spend
|
|
51
51
|
without double-counting.
|
|
@@ -138,16 +138,41 @@ Where the engine reports usage, those counts are the engine's own. Where it does
|
|
|
138
138
|
not, the wrapper falls back to `estimateTokens()` (characters ÷ 4) and the row is
|
|
139
139
|
flagged `tokensEstimated: true`; the CLI marks those costs with a trailing `~`.
|
|
140
140
|
|
|
141
|
-
| Engine | Token counts
|
|
142
|
-
| ----------------- |
|
|
143
|
-
| `claude` | Engine-reported
|
|
144
|
-
| `codex` | Engine-reported
|
|
145
|
-
| `codex-app` | Engine-reported
|
|
146
|
-
| `grok` | Engine-reported — and so is the **cost**: this engine reports `total_cost_usd`, which the wrapper passes through instead of pricing tokens from the registry, so registry drift cannot affect a grok row
|
|
147
|
-
| `cursor` (legacy) | Engine-reported when the stream carries `usage`, else estimated
|
|
148
|
-
| `opencode` | Engine-reported when the run JSON carries `tokens`, else estimated
|
|
149
|
-
| `agy` | Engine-reported when the result event carries usage, else estimated
|
|
150
|
-
| `custom` | Depends on the CLI; estimated when it emits no usage
|
|
141
|
+
| Engine | Token counts |
|
|
142
|
+
| ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
143
|
+
| `claude` | Engine-reported — and so is the **cost**: the `result` event's `total_cost_usd` is taken as-is, so registry drift cannot affect a Claude row. It is the session running total rather than the turn's, so spend advances by the difference between turns. Proxy sessions (`baseUrl` set) are the exception: the CLI is told it is running `opus` while another provider serves the tokens, so its figure is Opus list price for someone else's model and the registry estimate is used instead |
|
|
144
|
+
| `codex` | Engine-reported |
|
|
145
|
+
| `codex-app` | Engine-reported |
|
|
146
|
+
| `grok` | Engine-reported — and so is the **cost**: this engine reports `total_cost_usd`, which the wrapper passes through instead of pricing tokens from the registry, so registry drift cannot affect a grok row |
|
|
147
|
+
| `cursor` (legacy) | Engine-reported when the stream carries `usage`, else estimated |
|
|
148
|
+
| `opencode` | Engine-reported when the run JSON carries `tokens`, else estimated |
|
|
149
|
+
| `agy` | Engine-reported when the result event carries usage, else estimated |
|
|
150
|
+
| `custom` | Depends on the CLI; estimated when it emits no usage |
|
|
151
|
+
|
|
152
|
+
### What "input tokens" means is not the same on every engine
|
|
153
|
+
|
|
154
|
+
Two engines can report `input` and `cached` and mean different things by them,
|
|
155
|
+
and the difference decides whether the cost math may subtract one from the other:
|
|
156
|
+
|
|
157
|
+
| Engine | Its own arithmetic for one turn | `input` contains cached reads |
|
|
158
|
+
| ---------- | ---------------------------------------------------- | ----------------------------- |
|
|
159
|
+
| `codex` | `total 19704 = input 19699 + output 5` | yes |
|
|
160
|
+
| `grok` | `total 30034 = input 19393 + output 17 + read 10624` | no |
|
|
161
|
+
| `opencode` | `total 26315 = input 58 + output 17 + read 26240` | no |
|
|
162
|
+
| `claude` | `input_tokens 2` against `cache_read 47371` | no |
|
|
163
|
+
|
|
164
|
+
On an engine that excludes them, cached reads and cache writes are billed **on
|
|
165
|
+
top of** `input`, and the prompt the turn actually carried is the sum of all
|
|
166
|
+
three — which is what `contextPercent` measures. Reading `input` alone reports a
|
|
167
|
+
nearly-full context as empty on any resumed conversation, because the history
|
|
168
|
+
arrives as cached reads.
|
|
169
|
+
|
|
170
|
+
Anthropic bills cache _writes_ above the input rate, and the premium depends on
|
|
171
|
+
the TTL (1.25x for the 5-minute cache, 2x for the 1-hour one); the Claude
|
|
172
|
+
estimate prices both tiers from the split the engine reports. Elsewhere cache
|
|
173
|
+
writes are priced at the plain input rate, which is a floor rather than an exact
|
|
174
|
+
figure — one more reason to prefer the engine's own `total_cost_usd` where it
|
|
175
|
+
offers one.
|
|
151
176
|
|
|
152
177
|
So on an estimating engine the cap is best-effort. It will stop a runaway session;
|
|
153
178
|
it is not an accounting guarantee, and it is not a substitute for the spend limits
|
|
@@ -33,7 +33,8 @@ Start a persistent coding session with full CLI flag support.
|
|
|
33
33
|
| `mcpConfig` | string \| string[] | MCP server config file(s) |
|
|
34
34
|
| `settings` | string | Settings.json path or inline JSON |
|
|
35
35
|
| `ultracode` | boolean | Claude only. Enable "ultracode" / dynamic workflows — Claude plans a JS orchestration script per substantive task and fans out to subagents. Injected as the `ultracode:true` settings key (merged into `settings`), **not** a `--effort` value (the CLI rejects `--effort ultracode`). |
|
|
36
|
-
| `noSessionPersistence` | boolean | Do not save session to disk — both the engine's own transcript and this orchestrator's resume registry, so a later start under the same name does not reattach
|
|
36
|
+
| `noSessionPersistence` | boolean | Do not save session to disk — both the engine's own transcript and this orchestrator's resume registry, so a later start under the same name does not reattach |
|
|
37
|
+
| `ignoreUserConfig` | boolean | Codex only. Run without loading `$CODEX_HOME/config.toml`, so an orchestrated run is decided by what the caller passed rather than by the machine's own Codex config — notably a `model = …` line in that file, which otherwise picks the model while the ledger records this engine's default. Auth still resolves from `CODEX_HOME`. |
|
|
37
38
|
| `betas` | string \| string[] | Custom beta headers |
|
|
38
39
|
| `enableAgentTeams` | boolean | Enable experimental agent teams |
|
|
39
40
|
| `enableAutoMode` | boolean | Enable auto permission mode |
|
|
@@ -54,7 +55,7 @@ Start a persistent coding session with full CLI flag support.
|
|
|
54
55
|
| `otelLogUserPrompts` | boolean | OpenTelemetry: log user prompts (sets `OTEL_LOG_USER_PROMPTS=1`) |
|
|
55
56
|
| `otelLogRawApiBodies` | boolean | OpenTelemetry: log raw API request/response bodies (sets `OTEL_LOG_RAW_API_BODIES=1`); debug only |
|
|
56
57
|
| `bedrockServiceTier` | `'default'` \| `'flex'` \| `'priority'` | AWS Bedrock service tier (sets `ANTHROPIC_BEDROCK_SERVICE_TIER`); only effective when routing through Bedrock |
|
|
57
|
-
| `effort` | `'low'` \| `'medium'` \| `'high'` \| `'xhigh'` \| `'max'` \| `'auto'`
|
|
58
|
+
| `effort` | `'low'` \| `'medium'` \| `'high'` \| `'xhigh'` \| `'max'` \| `'ultra'` \| `'auto'` | Reasoning effort level. Engines do not share one ladder: Claude Code takes up to `max`, Codex adds `ultra`, Grok stops at `xhigh`, Antigravity at `high`, and OpenCode forwards the level to its provider unvalidated. A level above an engine's ceiling clamps rather than being dropped. `high` and above trigger the `ultrathink` prefix on Claude user messages. |
|
|
58
59
|
|
|
59
60
|
### `session_send`
|
|
60
61
|
|