@enderfga/claw-orchestrator 7.5.3 → 7.5.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (53) hide show
  1. package/README.md +21 -22
  2. package/configs/engines/README.md +7 -6
  3. package/dist/bin/cli.js +1 -1
  4. package/dist/bin/cli.js.map +1 -1
  5. package/dist/src/acp-server.d.ts +1 -1
  6. package/dist/src/acp-server.js +7 -5
  7. package/dist/src/acp-server.js.map +1 -1
  8. package/dist/src/autoloop/notify.d.ts +5 -7
  9. package/dist/src/autoloop/notify.js +21 -20
  10. package/dist/src/autoloop/notify.js.map +1 -1
  11. package/dist/src/embedded-server.js +8 -5
  12. package/dist/src/embedded-server.js.map +1 -1
  13. package/dist/src/fanout.d.ts +6 -0
  14. package/dist/src/fanout.js +1 -0
  15. package/dist/src/fanout.js.map +1 -1
  16. package/dist/src/index.js +21 -13
  17. package/dist/src/index.js.map +1 -1
  18. package/dist/src/kernel/engine.d.ts +39 -1
  19. package/dist/src/kernel/engine.js +120 -10
  20. package/dist/src/kernel/engine.js.map +1 -1
  21. package/dist/src/kernel/nodes/fanout.js +1 -0
  22. package/dist/src/kernel/nodes/fanout.js.map +1 -1
  23. package/dist/src/kernel/types.d.ts +9 -0
  24. package/dist/src/kernel/types.js.map +1 -1
  25. package/dist/src/openai-compat.d.ts +2 -2
  26. package/dist/src/openai-compat.js +5 -2
  27. package/dist/src/openai-compat.js.map +1 -1
  28. package/dist/src/session-manager.d.ts +7 -2
  29. package/dist/src/session-manager.js +21 -7
  30. package/dist/src/session-manager.js.map +1 -1
  31. package/dist/src/types.d.ts +2 -0
  32. package/openclaw.plugin.json +1 -1
  33. package/package.json +2 -2
  34. package/skills/SKILL.md +31 -32
  35. package/skills/references/acp.md +19 -36
  36. package/skills/references/autoloop.md +158 -180
  37. package/skills/references/claude-cli-tracking.md +27 -27
  38. package/skills/references/cli.md +62 -79
  39. package/skills/references/council.md +40 -63
  40. package/skills/references/dashboard.md +42 -55
  41. package/skills/references/getting-started.md +20 -14
  42. package/skills/references/inbox.md +6 -4
  43. package/skills/references/mcp.md +29 -24
  44. package/skills/references/multi-engine.md +105 -153
  45. package/skills/references/observability.md +42 -32
  46. package/skills/references/openai-compat.md +169 -303
  47. package/skills/references/sessions.md +20 -29
  48. package/skills/references/tools.md +67 -78
  49. package/skills/references/ultra.md +17 -16
  50. package/skills/references/ultraapp.md +59 -64
  51. package/skills/references/verification.md +29 -52
  52. package/skills/references/workflow.md +49 -107
  53. package/skills/ultraapp/SKILL.md +9 -10
@@ -1,36 +1,36 @@
1
1
  # Claude Code CLI Feature Tracking
2
2
 
3
- This document tracks which Claude Code CLI version Claw Orchestrator is currently synced to, and which features have been integrated.
3
+ This document tracks which Claude Code CLI version Claw Orchestrator is currently synced to, and which features have been integrated. Recent rows also record the other engines' versions checked in the same sync.
4
4
 
5
5
  ## Currently tracked: **Claude Code CLI 2.1.280** (as of 2026-09-23, plugin v7.5.3)
6
6
 
7
7
  ## Sync history
8
8
 
9
- | Plugin Version | Claude CLI Version | Date | Notable integrations |
10
- | v7.5.3 | 2.1.280 | 2026-09-23 | **A new frontier model on both sides, and the `opus` alias moved with it.** CC 2.1.278→2.1.280, Codex 0.155.1→0.156.1, agy 1.2.7→1.2.8, grok 1.0.34→1.0.41, OpenCode 1.18.31→1.18.32. 2.1.280 added Claude Opus 5.5 and made it what `--model opus` resolves to — confirmed against the binary, which reported `claude-opus-5-5` for an `opus` turn. It breaks the flat Opus pricing the registry relied on ($4/$20 against Opus 5's $5/$25) and prices cache reads at 5% of input rather than 10%, so every alias session — the autoloop Planner, the ultraplan default — was costed at the old rate, the number `maxBudgetUsd` gates on. Codex 0.156.1 added GPT-6 Sol and Luna; all three models are registered from the vendors' published price tables, with reverse assertions. The sweep's own missing-model check is what named them. Its Codex upstream lookup then failed reproducibly with HTTP 504 — `releases?per_page=100` carries every release body, and that repo's payload is large enough to time the API out — so it now asks `gh release list` for three fields, and a failed lookup prints why instead of only reporting an empty result. Live turns pass on claude, agy, grok and opencode; Codex could not be exercised (account usage limit until 2026-09-25), which is an account fact rather than a wrapper regression. Read and needing nothing here: agy 1.2.8 is compaction and TUI work; 2.1.280 fixed a symlinked write being judged by its in-tree path, auto-mode retry storms, and a finished subagent's report being lost when its launching conversation compacted. |
11
- | v7.5.1 | 2.1.278 | 2026-09-20 | **A vendor changed what a cost field covers, and the wrapper was reading the old meaning.** CC 2.1.274→2.1.278, Codex 0.154.0→0.155.1, agy 1.2.5→1.2.7; grok 1.0.34 and OpenCode 1.18.31 already latest. Every live turn passed through the real wrapper, the ACP and MCP handshakes are clean, registry 26 models / 0 drift, and no engine's flag surface changed — the whole week's findings came from the changelogs and from measuring what they describe. 2.1.277 made a headless process started with `--resume` restore the totals the resumed session saved at exit, where it used to begin at zero; measured on 2.1.278, a turn reported $0.363044 and the same session resumed in a new process reported $0.386463 for a turn whose own usage was $0.023419. `_applyReportedCost` advances spend by the difference between reports and had no earlier figure to subtract on a fresh process, so it billed the inherited total in full — on every model switch and every session recovered after a restart, against the number `maxBudgetUsd` gates on. A resumed process's first report is now a baseline and that turn keeps its registry estimate; mutation-verified. agy 1.2.6 changed the headless `-p` default timeout from five minutes to unlimited, which the wrapper already covers by deriving `--print-timeout` from the send timeout — the comments claiming agy would have timed out anyway are corrected. Also measured or read and needing nothing here: agy 1.2.6 now exits 3 with a structured `AGY_ERROR: {...}` line on stderr for an agent or model API failure, which the wrapper already fails the turn on (it compares against 0, never 1) and surfaces as the error message; agy 1.2.7 retired `find_by_name`, `grep_search` and `list_dir` from the default toolset, none of which this repo names; 2.1.277 fixed `claude -p` hanging with no result after an internal error, one of the failure modes the send timeout exists for; 2.1.275 fixed `--forward-subagent-text` dropping the messages of subagents spawned by a `context: fork` skill. Codex 0.155.x is TUI and Guardian work. |
12
- | v7.5.0 | 2.1.274 | 2026-09-17 | **Engine updates, and three wrapper fixes measured against 2.1.274.** CC 2.1.271→2.1.274, agy 1.2.2→1.2.5, grok 1.0.30→1.0.34; Codex 0.154.0 and OpenCode 1.18.31 already latest. Every live turn passed through the real wrapper, the ACP and MCP handshakes are clean, registry 26 models / 0 drift, and no engine's flag surface changed. agy 1.2.4 and 1.2.5 re-probed on a refused `RunCommand`: the refusal still arrives as `permission_denials`, and 1.2.5's one new flag (`--remote-control`) is for people, not wrappers. Recorded streams from 2.1.274 settled three things the wrapper had wrong: with `--include-partial-messages` every `tool_use` arrives on `content_block_start` (empty input) and again as an `assistant` event (same id, full input), so tool calls were counted and emitted twice; tool results arrive inside `user` messages, so `toolErrors` never moved; and the `init` event names the model, which the ledger had recorded as `default`. A headless Workflow launch (`--settings '{"ultracode":true}'`) returns a first `result` at launch and a second one, tagged `origin: {kind: "task-notification"}`, when the workflow finishes; every result that answers a written message carries its id in `user_message_uuids` (including `/compact`), so sends now carry an id and a result resolves only its own send. From the changelogs, needing nothing here: stream-json sessions no longer hold the first turn for deferred MCP servers, a backgrounded subagent's final report is no longer dropped from stream-json, and a transcript stuck on "unexpected tool*use_id" now ends with an error; agy no longer ends a turn with `NO_TOOL_CALL` on a schema-invalid tool call. The sweep had gone blind on Codex — 26 prereleases since 0.154.0 filled its 15-release window and the upstream column read "?" without failing — and now fails on an empty upstream lookup and reads grok's upstream from `grok update --check --json`. |
13
- | v7.4.1 | 2.1.271 | 2026-09-15 | **Short sweep: two Claude Code releases and one OpenCode release, nothing broken.** CC 2.1.269→2.1.271, OpenCode 1.18.30→1.18.31; Codex 0.154.0, agy 1.2.2 and grok 1.0.30 already latest by their own updaters. Every live turn passed through the real wrapper, registry 25 models / 0 drift, and no engine's flag surface changed. From the changelogs: 2.1.271 added `omitClaudeMd` to the `--agents` JSON (run a subagent without the user/project/local CLAUDE.md files). It already reached the CLI — the tool schema takes any object and the wrapper stringifies it verbatim, verified accepted by 2.1.271 — but the TypeScript type still named only `description` and `prompt`, while the CLI's schema has long carried `tools`, `model`, `maxTurns`, `background`, `memory`, `isolation` and `effort` too. The type is now `AgentDefinition`: `omitClaudeMd` named, the rest passed through, and a test pins the verbatim hand-over so a future allowlist cannot drop the next field silently. Needing nothing here: MCP-only `-p --resume` sessions now work, `--resume` keeps the `[1m]` tag across model families, Monitor watches get a 10-minute cap in `-p`. OpenCode 1.18.31's fixes are to its own ACP and TUI. Asked whether Opus 5.2 had shipped: it has not — Anthropic's pricing and model pages and the 2.1.271 binary all stop at `claude-opus-5`. That question exposed that the sweep would not have said so either: its missing-model check covered only GPT, and on inspection it had never run at all, because it resolved the engine binary against the working directory rather than `PATH` and an empty result is indistinguishable from "nothing new". Fixed, extended to Claude with retired rows excluded, and made loud when a binary cannot be scanned; mutation-verified by hiding `claude-opus-5` and `gpt-6-astra` from the registry, each of which it then named. Its first real run named `gpt-4.1`, now registered. |
14
- | v7.2.0 | 2.1.269 | 2026-09-13 | **Nine Claude Code releases with no new flag in `--help` — the change that mattered was in the result event.** CC 2.1.260→2.1.269, Codex 0.153.2→0.154.0, agy 1.1.25→1.2.2, grok 1.0.13→1.0.30, OpenCode 1.18.27→1.18.30; every live turn passed through the real wrapper, registry 25 models / 0 drift. The flag diff was empty for every engine but Codex, so the Claude changelog was read for what help cannot show. 2.1.269 made stream-json report \_all* refused tool calls in the result event's `permission_denials`; measured with `--permission-prompts none`, a turn asked to write a file ended `subtype: success`, `is_error: false`, the Bash call listed as denied and no file on disk. `sendMessage` dropped that event, so every caller saw a clean success — it is now surfaced as `SendResult.permissionDenials`. agy 1.2.2 has the same shape: in `--mode plan` a refused `RunCommand` still produced `status: SUCCESS` and a non-empty reply, the refusal visible only as a `soft-denying tool confirmation` line in its log. Codex 0.154.0's `--worktree` (behind `--enable worktrees`) measured and deliberately not wired: edits land in `~/.codex/worktrees/<hash>/<repo>` on a detached HEAD, never in the session cwd, so contracts and evidence would verify an untouched tree. Also in 2.1.261–2.1.269 and needing nothing here: `-p --resume` no longer inserts a spurious "Continue" turn, non-interactive sessions no longer reset cwd per message, a mid-turn model switch no longer loses the reply; `--append-subagent-system-prompt-file` exists but is hidden from help and has no counterpart option in this wrapper. |
15
- | v7.1.0 | 2.1.260 | 2026-09-04 | **Weekly sweep: five engines, zero regressions — the finding was in our own registry, not in a wrapper.** CC 2.1.259→2.1.260, Codex 0.153.0→0.153.2; agy 1.1.25, grok 1.0.13, OpenCode 1.18.27 unchanged. Grok's live turn passed for the first time since its free tier ran out — last week's pin was carried unverified, because an exhausted quota hangs silently rather than erroring. Two registry corrections: all three GPT-5.6 tiers had been repriced by OpenAI after launch while this repo kept the launch rates (Luna over-reported 5x), and `gpt-6-astra` needed registering (1,050,000 window, 10/1/50) or it would have fallen back to Sonnet pricing and a 200K window. Astra is absent from Codex 0.153.0 and present in 0.153.2 — baselining on the installed binary instead of upstream would have hidden it another week, which is the concrete case for that rule. The guard test for bare `gpt-5.6` had pinned launch literals and so failed to notice the repricing; it now asserts equality with `gpt-5.6-sol`, the invariant that actually holds. Still unregistered on purpose: `gpt-5.6-pro` (in the binary, no docs/pricing) and `gpt-5.6-cyber` (documented, not selectable by any engine here). |
16
- | v6.5.0 | 2.1.259 | 2026-09-03 | **First sweep run by script, not by hand — and it caught what by-hand missed.** CC 2.1.258→2.1.259, Codex 0.152.1→0.153.0, agy 1.1.22→1.1.25, OpenCode 1.18.26→1.18.27, grok unverifiable (silent hang on a spent free tier; pin stays 1.0.13). `scripts/sweep.ts` measures versions/pins/upstream, the wrapper's flags against each binary's `--help`, one live turn per engine _through the real wrapper class_, and the ACP/MCP handshakes; `scripts/sweep-workflow.json` wraps it as verifier→router→draft-agent→human-gate. CC delta: `--permission-prompts none` (2.1.259) now passed whenever no prompt tool is configured — a prompt nobody can answer is denied instead of hanging to the turn timeout. agy delta: 1.1.25 dropped `gemini-3.5-flash` (status: ERROR), which was this wrapper's default — every model-less Antigravity session was failing; default and `agy-flash` alias moved to 3.8 (registered, $0.75/$3.75 current rate), verified live. The by-hand sweep and the script's first version both exercised agy with a minimal argv and no `--model`, so both passed on agy's own default and never saw it; the live turn now goes through the wrapper. Two script lessons: help-text diffing flagged four still-accepted Claude flags as removed (advisory now, not a verdict), and `execFile` left a dangling stdin that timed out every one-shot engine (stdin is closed now, as the opencode wrapper always did). |
17
- | v6.2.1 | 2.1.258 | 2026-09-02 | **Fable 5.1 registry sync.** CC 2.1.251→2.1.258 in two days, Codex 0.151.0→0.152.1, OpenCode 1.18.25→1.18.26 (both surface-identical, both re-exercised); exactly one item in the delta is AI-facing. `claude-fable-5-1` is the new default Fable model, and the CLI's own `fable` alias resolves to it — read back from `modelUsage.canonicalModel` on a real turn against 2.1.258, not taken from the release note. Registered with `claude-mythos-5-1`, and the `fable` alias moved off 5, the same alias drift that hit Opus 5 and Sonnet 5 in earlier generations. The subtle part is the cache-read rate: 5.1 reads cache at **0.025x base input** ($0.25/Mtok against $10), where every other Claude model is 0.1x — copying Fable 5's $1 or deriving from the input rate over-reports it 4x. Cache _writes_ keep the usual 1.25x/2x, so only the read is special. Everything else in the range is TUI/settings/auto-mode/gateway and was dropped: `timeFormat`/`timeZone`, `/effort s`, `CLAUDE_CODE_SUBAGENT_MODEL_FORCE`, `permissions.blockReadsOutsideWorkingDirectories`, gateway model discovery, the auto-mode Containment Escape rule. |
18
- | v6.2.0 | 2.1.251 | 2026-08-31 | **Weekly sweep — grok's eight-version jump and a measured read-only failure.** CC 2.1.246→2.1.251, Codex 0.149.1→0.151.0, agy 1.1.21→1.1.22, **grok 1.0.5→1.0.13** (its own updater had reported 1.0.5 as current four days earlier), OpenCode 1.18.23→1.18.25. No registry drift — the first clean week in several; Anthropic also confirmed Sonnet 5's $2/$10 as standard on the last day of the announced window. CC delta: `--restricted` (2.1.249) wired as its own option, NOT folded into read-only — plan mode alone was measured against 2.1.251 and refused direct, shell and delegated-subagent writes, while `--restricted` additionally discards the caller's settings files. grok delta: `-p` is now short for `--single` (short form is what we pass, so invisible); nine existing session options finally reach it (`--rules`, `--tools`, `--disallowed-tools`, `--json-schema`, `--agent`, `--agents`, `--always-approve`, `--session-id`, `--fork-session`); neither tool list is validated, so an unknown name is silently ignored. grok read-only stays refused, now on evidence: a `--tools` allowlist + plan mode held for direct and shell writes and lost to delegation — the OpenCode `task` hole in a second engine. `--no-subagents` is the next probe and was left unwired because the confirming run hit the free-tier limit, and a quota failure writes no file either. |
19
- | v6.1.0 | 2.1.246 | 2026-08-27 | **Engine sweep that turned into a token-accounting audit.** Claude Code 2.1.237→2.1.246, Codex 0.148.0→0.149.1, agy 1.1.15→1.1.21, OpenCode 1.18.18→1.18.23, Grok unchanged at 1.0.5; each ran a live turn and the ACP stdio entry was smoked. No new Claude Code flag needed integrating in that range — what did was what its `result` event has been reporting all along. Four measurement bugs fixed: a turn's usage was added twice (once on `message_delta`, once on `result` — engine said 2/4/47371, we said 4/8/94742); cache writes were absent from the cost formula (engine `$0.322428` vs our `$0.016` on a 1h-cache turn, and `maxBudgetUsd` gates on ours), so Claude now takes `total_cost_usd` — a session running total, applied as a difference; the `input − cached` subtraction is valid only on codex, whose `input_tokens` includes cached reads, while claude/grok/opencode report them alongside; and `contextPercent` read `input_tokens` alone, so a 47k prompt measured 0% — now the whole prompt over `modelUsage[*].contextWindow`. Also: Codex gained a real `max` and an `ultra` above it (we were folding `max`→`xhigh`), Grok now takes `xhigh` natively (we were folding it to `high`), OpenCode's `--variant` means `effort` finally reaches it, `--ephemeral`/`--ignore-user-config`/`--add-dir` wired up on Codex. Registry: `gemini-3.7-flash`/`gemini-3.6-flash`/`gpt-5.2` registered, `gemini-3.5-flash` repriced $0.5/$3 → $1.50/$9. |
20
- | v4.8.0 | 2.1.207 | 2026-07-12 | **Autoloop role-level multi-engine support.** Planner, Coder, and Reviewer can select independent engines/models while preserving the Claude defaults. Built-in non-Claude Planners use native read-only/plan modes and receive their role protocol in-band. Spawn selections persist across resume; Codex persists its real thread ID. Runtime and invocation checks used Claude Code 2.1.207 and Codex 0.144.1. |
21
- | v4.7.0 | 2.1.206 | 2026-07-10 | **Antigravity engine ships + permission-mode sync.** Main feature is the community-contributed first-class `engine: 'agy'` (PR #71, reviewed + hardened: layered resume-ID gating, `agy/` prefix routing, shared `sanitize.ts` across all six engines, ENGINE*TYPES single source). Weekly CLI sync: CC 2.1.200 renamed the `default` permission mode to **`manual`** — verified against 2.1.206 that the choices are now `acceptEdits/auto/bypassPermissions/manual/dontAsk/plan`, that `default` is still accepted (hidden compat), and that **`delegate` is hard-rejected at spawn** — so PermissionMode gains `manual`, drops `delegate`, and agy/gemini map `manual` like `default` (→ `--sandbox`). Codex 0.143.0: empirically re-tested `-c model_reasoning_effort=max` — still 400-rejected for gpt-5.5 (the "first-class max" note is Bedrock GPT-5.6-only), so the `max`→`xhigh` map stays. GPT-5.6 Sol/Terra/Luna registered with official pricing ($5/$30, $2.50/$15, $1/$6; 1M/1M/400K ctx) after the user reported using it — it's a limited preview on API/Codex-auth paths (empirically: ChatGPT-account Codex auth gets a 400, which is why the first probe on this box misread it as Bedrock-only; lesson — an auth-path rejection is not model non-existence). Codex default stays gpt-5.5. Free upside: CC 2.1.205 fixed `--json-schema` invalid-schema silent fallback + `format` keyword rejection; CC 2.1.203 fixed background sessions dropping shell-exported `ANTHROPIC_BASE_URL`. Pins → CC 2.1.206 / Codex 0.143.0 (installed; npm has 0.144.1, exec surface unchanged per release notes). |
22
- | v4.6.0 | 2.1.199 | 2026-07-03 | **Model registry sync — Claude Fable 5.** Registered `claude-fable-5` (first Claude 5-family model, tier above Opus; standard $10/$50 per Mtok, cache read $1, full 1M context at standard rates per the official pricing page) with new `fable` alias; taught the `isClaudeModel`/`resolveProvider` heuristics to recognize `fable`/`mythos` strings (they only matched claude/opus/sonnet/haiku). Mythos 5 not listed (same price, limited availability). CC 2.1.198–199 are subagent/background-agent reliability fixes — no invocation-surface change; free upside for us: subagent partial output on rate-limit/server error is now returned instead of silently dropped, and API errors in subagents are reported to the parent. Codex unchanged (0.142.5 is a log-scrub patch; pin stays 0.142.4 as installed). |
23
- | v4.5.0 | 2.1.197 | 2026-07-01 | **Model registry sync — Claude Sonnet 5 + gpt-5.5 pricing.** CLI 2.1.197 shipped Sonnet 5 as the new default (native 1M-token context; standard $3/$15 per Mtok, launch promo $2/$10 through 2026-08-31 — we price the standard rate). Registered `claude-sonnet-5` in `models.ts` and moved the `sonnet` alias to it (was pinned to `claude-sonnet-4-6`), so `--model sonnet` tracks the CLI's own default and cost/context accounting stays correct; `claude-sonnet-4-6` stays selectable by full id. Also corrected `gpt-5.5` (the default Codex model) from placeholder pricing to OpenAI's published $5/$30 per Mtok + 1M context, and updated docs/examples off `gpt-5.4`. The CC 2.1.179→2.1.197 and Codex 0.138→0.142.x ranges are otherwise bug-fix / TUI / remote-executor / plugin-marketplace work that doesn't touch our invocation flags or the stream-json / codex-exec event schema — no wrapper change. Free upside (no code change): 2.1.181 fixed prompt-caching on custom `ANTHROPIC_BASE_URL` (helps proxy mode), 2.1.187 fixed `--json-schema` StructuredOutput infinite-recall, 2.1.196 turned the 5-min streaming idle watchdog on by default; Codex 0.139 preserves `oneOf`/`allOf` in `--output-schema`. Bumped tested versions Claude 2.1.197 / Codex 0.142.4. |
24
- | v4.3.0 | 2.1.178 | 2026-06-16 | **Parity batch 2 + legacy-subsystem upgrades (local-only, no cloud).** Claude `--fallback-model` array form (CSV, verified via `claude --help`). Codex-app `codex_threads` (`thread/list`) and `thread/resume` on start when `resumeSessionId` is set (param shapes from `generate-json-schema`). Council agents gain per-agent `effort`/`ultracode`. `ultrareview` re-implemented on the new cross-engine `fanout` primitive (opt-in `engines`, default claude-only). Consensus parsing exposes match source for observability. Dropped on purpose: `codex cloud exec`/best-of-N (cloud/managed — loses local control), `--bg` (we own the subprocess), `--output-last-message` (we already capture final text), ultraplan `ultracode` (violates its plan-only contract). Autoloop mid-turn steer deferred — the loop is strictly sequential (Coder fully completes before the Reviewer runs), so steer would always fall back to a fresh turn; a real version needs concurrent review. |
25
- | v4.2.0 | 2.1.178 | 2026-06-16 | **`ultracode` integration + binary-verified parity pass.** Added the `ultracode` option on `session_start` — Claude Code's dynamic-workflow mode, wired as the `ultracode: true` settings key merged into `--settings` (confirmed by spawning: it activates `workflow_agent` events in headless stream-json; `--effort ultracode` is rejected by the CLI). Added `claude_agents_list` (wraps `claude agents --json`). Verified against the binary that `claude continue/respawn/stop/logs` do **not** exist as headless subcommands (session continuation stays on `--resume`). Codex 0.137 side: app-server RPC tools `codex_interrupt`/`codex_steer`/`codex_fork`/`codex_rollback`/`codex_models`, `codex exec` reasoning-effort (`-c model_reasoning_effort`) + `--profile` passthrough, and a cross-engine `fanout**` primitive. Bumped tested versions Claude 2.1.178 / Codex 0.137.0. |
26
- | v4.1.2 | 2.1.161 | 2026-06-03 | **Model registry sync, not a CLI-flag integration.** Registered Opus 4.8 (`claude-opus-4-8`, now the `opus`alias) and 4.7 in`models.ts`— 2.1.154 shipped Opus 4.8 as the new default and our`opus`alias was still pinned to 4.6, mis-attributing cost. Effort ladder`low/medium/high/xhigh/max` was already supported (`index.ts`/`types.ts`). The 2.1.151–2.1.161 range (note: .151/.155 skipped) is otherwise TUI/reliability/telemetry; two fixes silently benefit our spawn path with no code change — 2.1.153 (stream-json stdin-close hang) and 2.1.161 (`-p`stdout corruption from background subagents). **Watch-out documented, not fixed:** 2.1.160 adds permission prompts under`acceptEdits` (our default) for build-tool config files (`.npmrc`/`.bazelrc`/`.pre-commit-config.yaml`/`.devcontainer/`etc.) and shell-startup files — headless flows touching these should set`dangerouslySkipPermissions`or a bypass permission mode. Codex unchanged (0.133.0); separately fixed Codex`turn.failed`/`error`events being swallowed. |
27
- | v4.1.1 | 2.1.150 | 2026-05-24 | **No Claude wrapper change** — 2.1.141–2.1.150 are almost entirely TUI / agent-view / security / visual; 2.1.150 itself is "internal infrastructure only". The one scripting-adjacent addition,`claude agents --json` (2.1.145), lists *CLI-managed\* sessions and is not used by our own session manager. This release's real engine work was on Codex/Gemini: Codex `--output-schema` wired into `jsonSchema` (Codex 0.132+), Gemini `--skip-trust` for the 0.43 trusted-folders gate. Bumped tested versions Claude 2.1.150 / Codex 0.133.0 / Gemini 0.43.0. |
28
- |---|---|---|---|
29
- | v4.1.0 | 2.1.140 | 2026-05-13 | `claude_goal_set` / `claude_goal_clear` / `claude_goal_status` tools (wrap CLI 2.1.139 `/goal` slash command), `plugin_details` tool (wraps `claude plugin details`, 2.1.139), `pluginUrl` config (maps to `--plugin-url`, 2.1.129). Skipped items that are user-controlled via `--settings` (worktree.baseRef, autoMode.hard_deny, skillOverrides, sandbox.bwrapPath / socatPath, parentSettingsBehavior) or auto-set by the CLI (`CLAUDE_CODE_SESSION_ID`, `CLAUDE_CODE_DISABLE_ALTERNATE_SCREEN`, `CLAUDE_CODE_FORCE_SYNC_OUTPUT` — all TTY-only). Hook `args: string[]`, `continueOnBlock`, hook input `effort.level`, subagent `x-claude-code-agent-id` headers are CLI-internal — no wrapper change needed. |
30
- | v2.14.2 | 2.1.126 | 2026-05-04 | `bedrockServiceTier` (Bedrock service-tier env, 2.1.122), `project_purge` tool (wraps `claude project purge`, 2.1.126); skipped passive-only items (OTel numeric attr, `invocation_trigger`, `/v1/models` gateway discovery, PowerShell shell changes) |
31
- | v2.14.0 | 2.1.121 | 2026-04-28 | `forkSubagent` (fork subagent env), `enableToolSearch` (Vertex AI tool search env), `otelLogUserPrompts` / `otelLogRawApiBodies` (OTEL logging toggles), `xhigh` effort level (Opus 4.7), `stats.pluginErrors` capture from `system/init` |
32
- | v2.13.0 | 2.1.111 | 2026-04-16 | Hook events, permission delegation, prompt cache optimization (exclude-dynamic-sections + 1H cache), debug control, `--from-pr`, MCP channels, `system/api_retry` event tracking |
33
- | v2.12.2 and earlier | 2.1.91 | — | Bare mode, worktree, json-schema, mcp-config, betas, fallback-model, effort, agent teams |
9
+ | Plugin Version | Claude CLI Version | Date | Notable integrations |
10
+ | ------------------- | ------------------ | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
11
+ | v7.5.3 | 2.1.280 | 2026-09-23 | CC 2.1.278→2.1.280, Codex 0.155.1→0.156.1, agy 1.2.7→1.2.8, grok 1.0.34→1.0.41, OpenCode 1.18.31→1.18.32. Integrated: Claude Opus 5.5 registered and the `opus` alias moved to `claude-opus-5-5` (the CLI resolves `--model opus` to it); Opus 5.5 is priced at $4/$20 per Mtok with cache reads at 5% of input. GPT-6 Sol and Luna (Codex 0.156.1) registered from the vendors' published price tables. The weekly sweep's Codex upstream lookup now uses `gh release list` and prints the reason when a lookup fails. Codex was not exercised live this cycle. Read, no change needed: agy 1.2.8 (compaction and TUI); CC 2.1.280 fixes for symlinked writes, auto-mode retries, and subagent reports lost on compaction. |
12
+ | v7.5.1 | 2.1.278 | 2026-09-20 | CC 2.1.274→2.1.278, Codex 0.154.0→0.155.1, agy 1.2.5→1.2.7; grok 1.0.34 and OpenCode 1.18.31 unchanged. No flag surface changed. Integrated: since 2.1.277 a process started with `--resume` reports the resumed session's saved cost totals, so the first cost report of a resumed process is now taken as a baseline and that turn keeps its registry estimate, so the inherited total is not counted toward `maxBudgetUsd`. Read, no change needed: agy 1.2.6's unlimited `-p` default timeout (the wrapper passes `--print-timeout`); agy 1.2.6 exit code 3 with an `AGY_ERROR: {...}` line (already treated as a failed turn); agy 1.2.7 removed `find_by_name`, `grep_search` and `list_dir` from the default toolset; CC 2.1.277 fixed `claude -p` hanging after an internal error; CC 2.1.275 fixed `--forward-subagent-text` for `context: fork` skills; Codex 0.155.x is TUI work. |
13
+ | v7.5.0 | 2.1.274 | 2026-09-17 | CC 2.1.271→2.1.274, agy 1.2.2→1.2.5, grok 1.0.30→1.0.34; Codex 0.154.0 and OpenCode 1.18.31 unchanged. No flag surface changed. Integrated, from 2.1.274 stream recordings: tool calls are counted once (with `--include-partial-messages` each `tool_use` appears on `content_block_start` and again in the `assistant` event); `toolErrors` reads tool results from `user` messages; the ledger records the model named in the `init` event. A headless ultracode launch returns one `result` at launch and a second, tagged `origin: {kind: "task-notification"}`, when the workflow finishes; sends now carry an id and resolve only on the `result` whose `user_message_uuids` includes it. The sweep now fails on an empty upstream lookup and reads grok's upstream from `grok update --check --json`. Read, no change needed: agy 1.2.4–1.2.5 still report refused `RunCommand` calls as `permission_denials`; agy's new `--remote-control` flag is interactive; CC stream-json fixes for deferred MCP servers, backgrounded subagent reports and `unexpected tool_use_id`; agy no longer ends a turn with `NO_TOOL_CALL` on a schema-invalid tool call. |
14
+ | v7.4.1 | 2.1.271 | 2026-09-15 | CC 2.1.269→2.1.271, OpenCode 1.18.30→1.18.31; Codex 0.154.0, agy 1.2.2 and grok 1.0.30 unchanged. No flag surface changed. Integrated: `omitClaudeMd` in the `--agents` JSON (2.1.271) — the `agents` value is passed to the CLI verbatim, and its TypeScript type is now `AgentDefinition`, naming `omitClaudeMd` and passing other fields through. The sweep's missing-model check now resolves engine binaries on `PATH`, covers Claude as well as GPT, and reports a binary it cannot scan; `gpt-4.1` registered. Read, no change needed: MCP-only `-p --resume` sessions, `--resume` keeping the `[1m]` tag, a 10-minute Monitor cap in `-p`; OpenCode 1.18.31 ACP and TUI fixes. |
15
+ | v7.2.0 | 2.1.269 | 2026-09-13 | CC 2.1.260→2.1.269, Codex 0.153.2→0.154.0, agy 1.1.25→1.2.2, grok 1.0.13→1.0.30, OpenCode 1.18.27→1.18.30. Integrated: 2.1.269 reports all refused tool calls in the result event's `permission_denials`, even when the turn ends `subtype: success`; these are now surfaced as `SendResult.permissionDenials`. agy 1.2.2 in `--mode plan` also reports a refused `RunCommand` as `status: SUCCESS`. Not wired: Codex 0.154.0 `--worktree` (behind `--enable worktrees`), because edits land in `~/.codex/worktrees/<hash>/<repo>` rather than the session cwd, so contracts and evidence would check an untouched tree. Read, no change needed: `-p --resume` no longer inserts a "Continue" turn, cwd no longer resets per message, a mid-turn model switch keeps the reply; the hidden `--append-subagent-system-prompt-file` has no counterpart option here. |
16
+ | v7.1.0 | 2.1.260 | 2026-09-04 | CC 2.1.259→2.1.260, Codex 0.153.0→0.153.2; agy 1.1.25, grok 1.0.13, OpenCode 1.18.27 unchanged. No wrapper change. Registry: GPT-5.6 tiers repriced to OpenAI's current rates; `gpt-6-astra` registered (1,050,000-token window). Not registered: `gpt-5.6-pro` (no published pricing) and `gpt-5.6-cyber` (not selectable by any engine here). |
17
+ | v6.5.0 | 2.1.259 | 2026-09-03 | CC 2.1.258→2.1.259, Codex 0.152.1→0.153.0, agy 1.1.22→1.1.25, OpenCode 1.18.26→1.18.27; grok not verified live (pin stays 1.0.13). Added `scripts/sweep.ts` (versions, flag surface against each binary's `--help`, one live turn per engine through the wrapper, ACP/MCP handshakes) and `scripts/sweep-workflow.json`. Integrated: `--permission-prompts none` (2.1.259) is passed whenever no permission-prompt tool is configured, so an unanswerable prompt is denied instead of waiting for the turn timeout. agy 1.1.25 dropped `gemini-3.5-flash`, the previous default; the agy default and the `agy-flash` alias moved to `gemini-3.8-flash`. |
18
+ | v6.2.1 | 2.1.258 | 2026-09-02 | CC 2.1.251→2.1.258, Codex 0.151.0→0.152.1, OpenCode 1.18.25→1.18.26. Integrated: `claude-fable-5-1` and `claude-mythos-5-1` registered; the `fable` alias moved to `claude-fable-5-1`. Fable 5.1 cache reads are priced at 0.025× base input ($0.25/Mtok); cache writes keep the usual 1.25×/2×. Read, no change needed: `timeFormat`/`timeZone`, `/effort s`, `CLAUDE_CODE_SUBAGENT_MODEL_FORCE`, `permissions.blockReadsOutsideWorkingDirectories`, gateway model discovery, the auto-mode Containment Escape rule. |
19
+ | v6.2.0 | 2.1.251 | 2026-08-31 | CC 2.1.246→2.1.251, Codex 0.149.1→0.151.0, agy 1.1.21→1.1.22, grok 1.0.5→1.0.13, OpenCode 1.18.23→1.18.25. Integrated: `--restricted` (2.1.249) as the separate `restricted` option, not part of read-only mode. grok options `--rules`, `--tools`, `--disallowed-tools`, `--json-schema`, `--agent`, `--agents`, `--always-approve`, `--session-id` and `--fork-session` wired (grok does not validate tool names). grok read-only sessions stay refused: a `--tools` allowlist plus plan mode blocked direct and shell writes but not writes through a delegated subagent. `--no-subagents` not yet wired. Registry: Sonnet 5's $2/$10 per Mtok is now the standard rate. |
20
+ | v6.1.0 | 2.1.246 | 2026-08-27 | CC 2.1.237→2.1.246, Codex 0.148.0→0.149.1, agy 1.1.15→1.1.21, OpenCode 1.18.18→1.18.23; grok unchanged at 1.0.5. No new Claude flag. Fixed token accounting: usage is counted once per turn; Claude cost uses the engine's `total_cost_usd` (applied as a difference, so cache writes are included); cached reads are subtracted from input only on Codex; `contextPercent` uses the whole prompt over `modelUsage[*].contextWindow`. Effort: Codex `max` and `ultra` passed natively, Grok `xhigh` passed natively, OpenCode receives effort via `--variant`. Codex `--ephemeral`, `--ignore-user-config` and `--add-dir` wired. Registry: `gemini-3.7-flash`, `gemini-3.6-flash` and `gpt-5.2` registered; `gemini-3.5-flash` repriced to $1.50/$9. |
21
+ | v4.8.0 | 2.1.207 | 2026-07-12 | Autoloop Planner, Coder and Reviewer can each use a different engine and model; Claude stays the default. Non-Claude Planners use their engine's read-only/plan mode and receive the role protocol in-band. Engine selections persist across resume; Codex persists its thread ID. Checked against Claude Code 2.1.207 and Codex 0.144.1. |
22
+ | v4.7.0 | 2.1.206 | 2026-07-10 | Added `engine: 'agy'` (Antigravity; PR #71) with resume-ID gating, `agy/` prefix routing, shared `sanitize.ts` across engines and `ENGINE_TYPES` as the single engine list. CC 2.1.200 renamed the `default` permission mode to `manual`: `manual` added, `default` still accepted, `delegate` removed (rejected by the CLI); agy/gemini map `manual` like `default` (→ `--sandbox`). Codex `max` effort still mapped to `xhigh` (rejected for gpt-5.5 on Codex 0.143.0). GPT-5.6 Sol/Terra/Luna registered ($5/$30, $2.50/$15, $1/$6; 1M/1M/400K context); Codex default stays gpt-5.5. Read, no change needed: CC 2.1.205 `--json-schema` fixes, CC 2.1.203 `ANTHROPIC_BASE_URL` fix for background sessions. |
23
+ | v4.6.0 | 2.1.199 | 2026-07-03 | Registered `claude-fable-5` ($10/$50 per Mtok, cache read $1, 1M context) with the `fable` alias; `isClaudeModel` and `resolveProvider` recognise `fable`/`mythos` model names. Read, no change needed: CC 2.1.198–2.1.199 subagent reliability fixes. Codex unchanged (0.142.4). |
24
+ | v4.5.0 | 2.1.197 | 2026-07-01 | Registered `claude-sonnet-5` (1M context, $3/$15 per Mtok standard) and moved the `sonnet` alias to it; `claude-sonnet-4-6` remains selectable by full id. `gpt-5.5` priced at $5/$30 per Mtok with 1M context. Read, no change needed: CC 2.1.179–2.1.197 and Codex 0.138–0.142.x (no invocation or event-schema change), including CC 2.1.181 prompt caching on custom `ANTHROPIC_BASE_URL`, 2.1.187 `--json-schema` fix, 2.1.196 streaming idle watchdog on by default, Codex 0.139 `oneOf`/`allOf` in `--output-schema`. |
25
+ | v4.3.0 | 2.1.178 | 2026-06-16 | Claude `--fallback-model` array form. Codex app-server `thread/list` tool and `thread/resume` on start when `resumeSessionId` is set. Council agents gain per-agent `effort` and `ultracode`. `ultrareview` runs on the cross-engine `fanout` primitive (opt-in `engines`, default Claude only). Consensus parsing reports its match source. Not adopted: `codex cloud exec`, `--bg`, `--output-last-message`, `ultracode` for ultraplan; autoloop mid-turn steer (the loop runs Coder and Reviewer sequentially). |
26
+ | v4.2.0 | 2.1.178 | 2026-06-16 | Added `ultracode` on `session_start` (the `ultracode: true` settings key merged into `--settings`) and `claude_agents_list` (wraps `claude agents --json`). `claude continue/respawn/stop/logs` are not headless subcommands; continuation uses `--resume`. Codex 0.137: app-server tools `codex_interrupt`, `codex_steer`, `codex_fork`, `codex_rollback`, `codex_models`; `codex exec` reasoning effort (`-c model_reasoning_effort`) and `--profile`; cross-engine `fanout` primitive. |
27
+ | v4.1.2 | 2.1.161 | 2026-06-03 | Registered Opus 4.8 (`claude-opus-4-8`, now the `opus` alias) and Opus 4.7. The `low/medium/high/xhigh/max` effort ladder was already supported. Read, no change needed: 2.1.151–2.1.161 TUI/reliability/telemetry work, including 2.1.153 (stream-json stdin-close hang) and 2.1.161 (`-p` stdout corruption from background subagents). Note: since 2.1.160, `acceptEdits` prompts before editing build-tool config files (`.npmrc`, `.bazelrc`, `.pre-commit-config.yaml`, `.devcontainer/`, etc.) and shell startup files; headless flows that touch them need a bypass permission mode. Codex `turn.failed`/`error` events are now surfaced. |
28
+ | v4.1.1 | 2.1.150 | 2026-05-24 | No Claude wrapper change: 2.1.141–2.1.150 is TUI, agent-view, security and visual work; `claude agents --json` (2.1.145) lists _CLI-managed_ sessions and is not used by the session manager. Codex `--output-schema` wired into `jsonSchema` (Codex 0.132+); Gemini `--skip-trust` for the 0.43 trusted-folders check. |
29
+ | v4.1.0 | 2.1.140 | 2026-05-13 | `claude_goal_set` / `claude_goal_clear` / `claude_goal_status` tools (wrap CLI 2.1.139 `/goal` slash command), `plugin_details` tool (wraps `claude plugin details`, 2.1.139), `pluginUrl` config (maps to `--plugin-url`, 2.1.129). Skipped items that are user-controlled via `--settings` (worktree.baseRef, autoMode.hard_deny, skillOverrides, sandbox.bwrapPath / socatPath, parentSettingsBehavior) or auto-set by the CLI (`CLAUDE_CODE_SESSION_ID`, `CLAUDE_CODE_DISABLE_ALTERNATE_SCREEN`, `CLAUDE_CODE_FORCE_SYNC_OUTPUT` — all TTY-only). Hook `args: string[]`, `continueOnBlock`, hook input `effort.level`, subagent `x-claude-code-agent-id` headers are CLI-internal — no wrapper change needed. |
30
+ | v2.14.2 | 2.1.126 | 2026-05-04 | `bedrockServiceTier` (Bedrock service-tier env, 2.1.122), `project_purge` tool (wraps `claude project purge`, 2.1.126); skipped passive-only items (OTel numeric attr, `invocation_trigger`, `/v1/models` gateway discovery, PowerShell shell changes) |
31
+ | v2.14.0 | 2.1.121 | 2026-04-28 | `forkSubagent` (fork subagent env), `enableToolSearch` (Vertex AI tool search env), `otelLogUserPrompts` / `otelLogRawApiBodies` (OTEL logging toggles), `xhigh` effort level (Opus 4.7), `stats.pluginErrors` capture from `system/init` |
32
+ | v2.13.0 | 2.1.111 | 2026-04-16 | Hook events, permission delegation, prompt cache optimization (exclude-dynamic-sections + 1H cache), debug control, `--from-pr`, MCP channels, `system/api_retry` event tracking |
33
+ | v2.12.2 and earlier | 2.1.91 | — | Bare mode, worktree, json-schema, mcp-config, betas, fallback-model, effort, agent teams |
34
34
 
35
35
  ## How to update this
36
36
 
@@ -41,4 +41,4 @@ When syncing to a new Claude Code CLI version:
41
41
  3. Decide which features are valuable for programmatic/agent use (vs human-interactive only)
42
42
  4. Implement worthwhile features (add to `SessionConfig` → wire into `persistent-session.ts` → expose in tool schema → document)
43
43
  5. Update this file with the new version + notable integrations
44
- 6. Update `CLAUDE.md` and `README.md` engine compatibility tables
44
+ 6. Update the engine tables in `AGENTS.md` and `README.md`
@@ -5,14 +5,14 @@ The CLI is an HTTP client that talks to the Claw Orchestrator embedded server. I
5
5
  ## Server
6
6
 
7
7
  ```bash
8
- clawo serve [-p, --port <port>]
8
+ clawo serve [-p, --port <port>] [-H, --host <host>] [--ultraapp-runtime host|docker]
9
9
  ```
10
10
 
11
- Start standalone embedded server (default port 18796). Set `CLAWO_API_URL` to override the base URL.
11
+ Start standalone embedded server (default port 18796, bound to `127.0.0.1`; pass `-H 0.0.0.0` for remote access). `--ultraapp-runtime` selects how ultraapp builds run: `host` (default, spawns Node directly) or `docker`. CLI commands reach the server at `CLAWO_API_URL` (default `http://127.0.0.1:18796`).
12
12
 
13
13
  ### Rate Limiting
14
14
 
15
- The embedded server enforces a sliding-window rate limit of 100 requests per minute per IP address. Requests exceeding the limit receive HTTP 429 (Too Many Requests). This prevents accidental runaway scripts from overwhelming the server.
15
+ The embedded server limits each IP address to 300 requests per minute (sliding window; override with `OPENCLAW_RATE_LIMIT`). Requests over the limit receive HTTP 429 (Too Many Requests).
16
16
 
17
17
  ### OpenAI-Compatible API
18
18
 
@@ -29,7 +29,7 @@ The server exposes an OpenAI-compatible chat completions endpoint, enabling any
29
29
 
30
30
  ```json
31
31
  {
32
- "model": "claude-sonnet-4-6",
32
+ "model": "sonnet",
33
33
  "messages": [{ "role": "user", "content": "Hello!" }],
34
34
  "stream": true
35
35
  }
@@ -39,21 +39,31 @@ The server exposes an OpenAI-compatible chat completions endpoint, enabling any
39
39
 
40
40
  1. `X-Session-Id` header
41
41
  2. `user` field in the request body
42
- 3. Default singleton session
42
+ 3. A hash of the model, system prompt and tool definitions (`sys-<hash>`)
43
+ 4. `default`, when all of those are empty
43
44
 
44
45
  **Model routing:** The `model` field auto-routes to the correct engine:
45
46
 
46
47
  - `claude-*`, `opus`, `sonnet`, `haiku` → Claude engine
47
- - `gpt-*` → Codex engine
48
- - `grok-*` → Grok engine
49
- - `composer-*` → Cursor engine (legacy)
50
- - `gemini-3.5-flash`, `gemini-3.1-pro`, `agy-*`, `agy/*` → Antigravity (`agy`) engine
51
- - other `gemini-*` → the legacy `gemini` engine (Gemini CLI is sunset; prefer `agy`)
48
+ - `gpt-*`, `o3*`, `o4*`, `codex*` → Codex engine
49
+ - `grok*` → Grok engine
50
+ - `composer*`, `cursor*`, `auto` → Cursor engine (legacy)
51
+ - `agy/*`, the `agy-flash` / `agy-pro` aliases, and the registered Antigravity models (e.g. `gemini-3.8-flash`, `gemini-3.1-pro`) → Antigravity (`agy`) engine
52
+ - other `gemini*` → the legacy `gemini` engine (Gemini CLI is sunset; prefer `agy`)
53
+ - anything else → Claude engine
52
54
 
53
55
  **CORS:** `/v1/` paths allow cross-origin requests by default. Set `OPENCLAW_CORS_ORIGINS=*` to allow all origins on all paths.
54
56
 
55
57
  **Auto-compact:** When a session's context utilization exceeds 80%, the endpoint automatically compacts the session before sending the next message.
56
58
 
59
+ ## ACP agent
60
+
61
+ ```bash
62
+ clawo acp
63
+ ```
64
+
65
+ Run as an Agent Client Protocol agent over stdio (same as the `clawo-acp` binary). See [acp.md](./acp.md).
66
+
57
67
  ## Session Management
58
68
 
59
69
  ### session-start
@@ -62,38 +72,37 @@ The server exposes an OpenAI-compatible chat completions endpoint, enabling any
62
72
  clawo session-start [name] [options]
63
73
  ```
64
74
 
65
- | Flag | Description |
66
- | ------------------------------------------------ | -------------------------------------------------------------------------------------------------- |
67
- | `-d, --cwd <dir>` | Working directory |
68
- | `-e, --engine <engine>` | Engine: `claude` (default), `codex`, `codex-app`, `agy`, `grok`, `opencode`, or `custom` |
69
- | `-m, --model <model>` | Model name or alias |
70
- | `--permission-mode <mode>` | `acceptEdits`, `plan`, `auto`, `bypassPermissions`, `manual`, `dontAsk` |
71
- | `--effort <level>` | `low`, `medium`, `high`, `max`, `auto` |
72
- | `--allowed-tools <tools>` | Comma-separated tool whitelist |
73
- | `--max-turns <n>` | Max agent loop turns |
74
- | `--max-budget <usd>` | API cost ceiling |
75
- | `--system-prompt <text>` | Replace system prompt |
76
- | `--append-system-prompt <text>` | Append to system prompt |
77
- | `--agents <json>` | Custom sub-agents JSON |
78
- | `--agent <name>` | Default agent |
79
- | `--bare` | No CLAUDE.md, no git context |
80
- | `-w, --worktree [name]` | Git worktree |
81
- | `--fallback-model <model>` | Fallback model |
82
- | `--json-schema <schema>` | JSON Schema for structured output |
83
- | `--mcp-config <paths>` | MCP config files (comma-separated) |
84
- | `--settings <path>` | Settings.json path |
85
- | `--skip-persistence` | Do not save the session — neither the engine's transcript nor the resume registry |
86
- | `--betas <headers>` | Beta headers (comma-separated) |
87
- | `--enable-agent-teams` | Enable agent teams |
88
- | `--include-hook-events` | Stream hook lifecycle events (PreToolUse/PostToolUse) |
89
- | `--permission-prompt-tool <tool>` | Delegate permission prompts to an MCP tool (non-interactive use) |
90
- | `--exclude-dynamic-system-prompt-sections` | Move cwd/env/git context to user message for better prompt cache hits (auto-enabled with `--bare`) |
91
- | `--debug <categories>` | Enable targeted debug output by category (e.g. `"api,mcp"`) |
92
- | `--debug-file <path>` | Write debug output to file |
93
- | `--from-pr <n>` | Resume a session linked to a GitHub PR number or URL |
94
- | `--channels <spec>` | MCP channel subscription (research preview) |
95
- | `--dangerously-load-development-channels <spec>` | Development MCP channel subscriptions (research preview) |
96
- | `ENABLE_PROMPT_CACHING_1H=1` (env var) | Enable 1-hour prompt cache TTL (auto-set with `--bare`) |
75
+ | Flag | Description |
76
+ | ------------------------------- | ------------------------------------------------------------------------------------------- |
77
+ | `-d, --cwd <dir>` | Working directory |
78
+ | `-e, --engine <engine>` | Engine: `claude` (default), `codex`, `codex-app`, `agy`, `grok`, `opencode`, or `custom` |
79
+ | `-m, --model <model>` | Model name or alias |
80
+ | `--permission-mode <mode>` | `acceptEdits` (default), `plan`, `auto`, `bypassPermissions`, `manual`, `dontAsk` |
81
+ | `--effort <level>` | `low`, `medium`, `high`, `xhigh`, `max`, `ultra`, `auto` |
82
+ | `--allowed-tools <tools>` | Comma-separated tool whitelist |
83
+ | `--disallowed-tools <tools>` | Comma-separated tools to deny |
84
+ | `--max-turns <n>` | Max agent loop turns |
85
+ | `--max-budget <usd>` | API cost ceiling |
86
+ | `--system-prompt <text>` | Replace system prompt |
87
+ | `--append-system-prompt <text>` | Append to system prompt |
88
+ | `--agents <json>` | Custom sub-agents JSON |
89
+ | `--agent <name>` | Default agent |
90
+ | `--bare` | No CLAUDE.md, no git context |
91
+ | `-w, --worktree [name]` | Git worktree |
92
+ | `--fallback-model <model>` | Fallback model |
93
+ | `--json-schema <schema>` | JSON Schema for structured output |
94
+ | `--mcp-config <paths>` | MCP config files (comma-separated) |
95
+ | `--settings <pathOrJson>` | Settings.json path or inline JSON |
96
+ | `--skip-persistence` | Do not save the session — neither the engine's transcript nor the resume registry |
97
+ | `--betas <headers>` | Beta headers (comma-separated) |
98
+ | `--enable-agent-teams` | Enable agent teams |
99
+ | `--enable-auto-mode` | Enable auto permission mode |
100
+ | `--resume-session-id <id>` | Resume an existing session by ID |
101
+ | `--base-url <url>` | Custom API endpoint (for proxy) |
102
+ | `--add-dir <dirs>` | Comma-separated additional working directories |
103
+ | `--custom-engine <preset>` | With `-e custom`: the id of a bundled engine preset (see [`clawo engines`](#clawo-engines)) |
104
+
105
+ The remaining `session_start` options (hook events, permission-prompt tool, debug output, `fromPr`, MCP channels, prompt-cache settings and others) are tool parameters only and have no CLI flag. See [Tools Reference](./tools.md#session_start).
97
106
 
98
107
  ### session-send
99
108
 
@@ -141,15 +150,17 @@ clawo session-compact <name> [--summary <text>]
141
150
  ## Run Ledger
142
151
 
143
152
  ```bash
144
- clawo runs [--since <window>] [-n, --limit <n>] [--session <name>] [--engine <engine>] [--parent <id>] [--json]
153
+ clawo runs [-s, --since <window>] [-n, --limit <n>] [--session <name>] [--engine <engine>] [--parent <id>] [--verified | --refuted] [--json]
145
154
  ```
146
155
 
147
156
  Show the durable per-turn record kept at `~/.claw-orchestrator/runs/`. Unlike
148
157
  `session-status`, this survives restarts and covers sessions this process never
149
158
  owned. `--since` takes `30m` / `24h` / `7d` / `2w` or an ISO timestamp (default
150
- `24h`); `--parent` filters to one council / fanout / autoloop run. Costs marked
151
- with a trailing `~` came from estimated token counts. See
152
- [observability.md](observability.md).
159
+ `24h`); `--parent` filters to one council / fanout / autoloop / workflow run;
160
+ `--verified` / `--refuted` keep only turns whose acceptance contract passed /
161
+ failed. Costs marked with a trailing `~` came from estimated token counts. The
162
+ `VERIFIED` column reads `yes`, `NO`, or `—` ("no contract was declared, so
163
+ nothing checked it"). See [observability.md](observability.md).
153
164
 
154
165
  ## Agent Management
155
166
 
@@ -179,30 +190,13 @@ clawo session-team-list <name>
179
190
  clawo session-team-send <name> <teammate> <message>
180
191
  ```
181
192
 
182
- ## SDK-Only Tools (No CLI Wrapper)
183
-
184
- The following tools are available through the OpenClaw plugin SDK and TypeScript API but do not have CLI commands. Use the SDK directly or call them via OpenClaw's tool system.
193
+ ## Tools Without a CLI Command
185
194
 
186
- | Tool | Description |
187
- | ----------------------- | ------------------------------------------------- |
188
- | `sessions_overview` | Aggregate dashboard of all active sessions |
189
- | `session_update_tools` | Hot-swap allowed/disallowed tools via `--resume` |
190
- | `session_switch_model` | Switch model mid-session via `--resume` |
191
- | `council_start` | Start multi-agent council with worktree isolation |
192
- | `council_status` | Poll council progress and agent responses |
193
- | `council_abort` | Abort a running council |
194
- | `council_inject` | Inject a message into the next council round |
195
- | `session_send_to` | Cross-session messaging (immediate or queued) |
196
- | `session_inbox` | Read inbox messages for a session |
197
- | `session_deliver_inbox` | Deliver queued messages to an idle session |
198
- | `ultraplan_start` | Start background Opus planning session |
199
- | `ultraplan_status` | Poll ultraplan progress |
200
- | `ultrareview_start` | Start fleet of parallel reviewer agents |
201
- | `ultrareview_status` | Poll ultrareview findings |
195
+ Tools without a CLI command are reachable through the OpenClaw plugin, the MCP
196
+ server (`clawo-mcp`, see [mcp.md](./mcp.md)), or the `SessionManager` API. See
197
+ [Tools Reference](./tools.md) for full parameter documentation.
202
198
 
203
- See [Tools Reference](./tools.md) for full parameter documentation.
204
-
205
- ## `clawo workflow` (6.0.0)
199
+ ## `clawo workflow`
206
200
 
207
201
  ```bash
208
202
  clawo workflow list [--state <s>] [--workflow <name>] [--limit N] [--json]
@@ -230,17 +224,6 @@ Prints the evidence bundle: per-check pass/fail with the failing command and its
230
224
  output tail, the base and head commits, and the files the run changed (created
231
225
  files included).
232
226
 
233
- ## `clawo runs` additions
234
-
235
- ```bash
236
- clawo runs --verified # only turns whose acceptance contract passed
237
- clawo runs --refuted # only turns whose acceptance contract failed
238
- ```
239
-
240
- The table gains a `VERIFIED` column with three values: `yes`, `NO`, and `—` for
241
- "no contract was declared, so nothing checked it". See
242
- [`observability.md`](./observability.md).
243
-
244
227
  ## `clawo engines`
245
228
 
246
229
  Lists the community engine presets bundled with the installed package, with each