@enderfga/claw-orchestrator 4.7.0 → 4.8.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (49) hide show
  1. package/README.md +8 -9
  2. package/configs/autoloop-coder-prompt.md +4 -4
  3. package/configs/autoloop-planner-prompt.md +10 -10
  4. package/configs/autoloop-reviewer-prompt.md +3 -4
  5. package/dist/src/autoloop/dispatcher.d.ts +59 -10
  6. package/dist/src/autoloop/dispatcher.js +261 -40
  7. package/dist/src/autoloop/dispatcher.js.map +1 -1
  8. package/dist/src/autoloop/planner-tools.d.ts +4 -3
  9. package/dist/src/autoloop/planner-tools.js +35 -2
  10. package/dist/src/autoloop/planner-tools.js.map +1 -1
  11. package/dist/src/autoloop/types.d.ts +2 -0
  12. package/dist/src/autoloop/types.js.map +1 -1
  13. package/dist/src/dashboard/index.html +185 -94
  14. package/dist/src/embedded-server.js +67 -9
  15. package/dist/src/embedded-server.js.map +1 -1
  16. package/dist/src/index.js +84 -52
  17. package/dist/src/index.js.map +1 -1
  18. package/dist/src/persistent-agy-session.js +4 -1
  19. package/dist/src/persistent-agy-session.js.map +1 -1
  20. package/dist/src/persistent-codex-session.js +3 -0
  21. package/dist/src/persistent-codex-session.js.map +1 -1
  22. package/dist/src/persistent-cursor-session.d.ts +6 -0
  23. package/dist/src/persistent-cursor-session.js +98 -3
  24. package/dist/src/persistent-cursor-session.js.map +1 -1
  25. package/dist/src/persistent-custom-session.js +13 -0
  26. package/dist/src/persistent-custom-session.js.map +1 -1
  27. package/dist/src/persistent-gemini-session.d.ts +4 -0
  28. package/dist/src/persistent-gemini-session.js +55 -2
  29. package/dist/src/persistent-gemini-session.js.map +1 -1
  30. package/dist/src/persistent-opencode-session.d.ts +5 -5
  31. package/dist/src/persistent-opencode-session.js +58 -6
  32. package/dist/src/persistent-opencode-session.js.map +1 -1
  33. package/dist/src/persistent-session.js +6 -1
  34. package/dist/src/persistent-session.js.map +1 -1
  35. package/dist/src/session-manager.d.ts +44 -5
  36. package/dist/src/session-manager.js +190 -18
  37. package/dist/src/session-manager.js.map +1 -1
  38. package/dist/src/types.d.ts +3 -2
  39. package/package.json +1 -1
  40. package/skills/SKILL.md +18 -26
  41. package/skills/references/autoloop.md +49 -6
  42. package/skills/references/claude-cli-tracking.md +2 -1
  43. package/skills/references/cli.md +2 -2
  44. package/skills/references/council.md +1 -1
  45. package/skills/references/getting-started.md +4 -4
  46. package/skills/references/multi-engine.md +21 -34
  47. package/skills/references/openai-compat.md +1 -1
  48. package/skills/references/sessions.md +3 -3
  49. package/skills/references/tools.md +39 -20
@@ -21,12 +21,39 @@ This page is the operator reference.
21
21
 
22
22
  ## Roles
23
23
 
24
- | Agent | Engine (default) | cwd | Owns |
24
+ | Agent | Default | cwd | Owns |
25
25
  |---|---|---|---|
26
26
  | **Planner** | claude / opus | workspace | strategy, `plan.md`, `goal.json`, talking to you |
27
- | **Coder** | claude / sonnet (override per spawn) | workspace | code changes, eval execution |
27
+ | **Coder** | claude / sonnet | workspace | code changes, eval execution |
28
28
  | **Reviewer** | claude / sonnet | `<workspace>/tasks/<run_id>/reviewer_sandbox/` | distrust audit; advance / hold / rollback |
29
29
 
30
+ Each role can use any built-in engine, or a `custom` engine config supplied by a
31
+ local caller (custom engines name an executable, so the HTTP API does not accept
32
+ them — see [tools.md](./tools.md)). If a non-Claude role omits `model`, that CLI
33
+ uses its own default model rather than receiving the Claude `opus` / `sonnet`
34
+ defaults. Role instructions are included in-band for engines that do not expose a
35
+ native system-prompt flag.
36
+
37
+ Engines without native multi-turn conversation (Cursor, OpenCode, one-shot custom
38
+ engines) spawn a fresh process per send, so the dispatcher replays that role's
39
+ transcript in-band as a `<conversation_history>` block, oldest turns dropped past a
40
+ character budget. Claude, Codex and Antigravity keep context themselves and get no
41
+ replay.
42
+
43
+ The Planner runs read-only so strategy cannot turn into source edits, and that is
44
+ enforced by the engine rather than requested politely: Claude uses plan mode,
45
+ Antigravity and Cursor use their plan modes, and OpenCode gets a generated
46
+ `clawo-readonly` agent that denies `edit`/`bash`/`external_directory` (its built-in
47
+ `plan` agent is a user-overridable preset that denies neither, so a "read-only"
48
+ session could otherwise still author files through a shell heredoc). A custom
49
+ Planner receives `permissionMode: 'manual'` and its `CustomEngineConfig` **must**
50
+ map that mode to the CLI's read-only flag — if it cannot, the session refuses to
51
+ start rather than silently running write-enabled.
52
+
53
+ Coder and Reviewer engine/model choices can be overridden by the first successful
54
+ `spawn_subagents`; later attempts to change an already-started role are rejected
55
+ instead of silently diverging from the running session.
56
+
30
57
  Coder and Reviewer **never speak to you directly**. Anything they observe
31
58
  flows through the Planner. The Planner decides what to surface and what to
32
59
  absorb.
@@ -78,7 +105,7 @@ curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_stop \
78
105
 
79
106
  | Tool | Args | What |
80
107
  |---|---|---|
81
- | `autoloop_start` | `run_id`, `workspace`, `planner_model?`, `send_timeout_ms?` | Start a run; launches Planner session. |
108
+ | `autoloop_start` | `run_id`, `workspace`, per-role `*_engine?`, `*_model?`, `*_custom_engine?`, `send_timeout_ms?` | Start a run; launches Planner and stores Coder/Reviewer defaults. Each `custom` role requires its matching config. |
82
109
  | `autoloop_chat` | `run_id`, `text` | Send a chat message to the Planner; returns the Planner's reply. |
83
110
  | `autoloop_status` | `run_id` | Current state (status, iter, push count, subagents_spawned). |
84
111
  | `autoloop_list` | — | All active runs in this manager process. |
@@ -94,7 +121,7 @@ never see the JSON — only the Planner's narrative.
94
121
  | Tool | Args | What |
95
122
  |---|---|---|
96
123
  | `notify_user` | `level` ('info' / 'warn' / 'decision' / 'error'), `summary`, `detail?`, `channel?` ('auto' / 'wechat' / 'webchat' / 'both' / 'email') | Push you out-of-band. |
97
- | `spawn_subagents` | `coder_model?`, `reviewer_model?`, `initial_directive?` | Start Coder + Reviewer. Only after explicit user approval. |
124
+ | `spawn_subagents` | `coder_engine?`, `coder_model?`, `reviewer_engine?`, `reviewer_model?`, `initial_directive?` | Start Coder + Reviewer. Omitted values inherit run defaults. An engine change without a model uses the new engine's default. Once a role session has started, changing its engine/model is rejected. Custom configs cannot be emitted by Planner. Only after explicit user approval. |
98
125
  | `send_directive` | `goal`, `constraints?`, `success_criteria?`, `max_attempts?` | Next iter's instruction to Coder. |
99
126
  | `pause_loop` | `reason` | Halt subloop at next iter boundary; chat keeps working. |
100
127
  | `resume_loop` | — | Resume after pause. |
@@ -103,6 +130,22 @@ never see the JSON — only the Planner's narrative.
103
130
  | `write_plan` | `content` (full plan.md body), `commit_message?` | Write `plan.md` to the workspace and git-commit. The **only** way the Planner can author plan.md — Write/Edit are stripped from the Planner session as a hard role boundary. Re-running replaces the whole file. |
104
131
  | `write_goal` | `content` (full goal.json body), `commit_message?` | Same, for `goal.json`. Content is JSON-validated before write; malformed content errors back to the Planner. |
105
132
 
133
+ ### Custom engines and resume
134
+
135
+ Custom engine configs are accepted only by `autoloop_start` (or the HTTP resume
136
+ body), never through Planner output. This keeps config fields such as `env` and
137
+ static CLI arguments out of the Planner transcript and `decisions.jsonl`.
138
+ The central registry persists only each role's engine and model, including the
139
+ effective Coder/Reviewer selection after a successful spawn. Resume leaves the
140
+ prior append-only row untouched until startup succeeds, so a transient CLI
141
+ failure cannot erase the run. When resuming a run that uses `custom`, provide
142
+ the matching `planner_custom_engine`, `coder_custom_engine`, or
143
+ `reviewer_custom_engine` again; otherwise resume fails with a clear
144
+ configuration error rather than silently switching to Claude. Custom config
145
+ shape is validated at runtime, while its `env` and static CLI arguments remain
146
+ out of registry and audit records. See [`multi-engine.md`](./multi-engine.md)
147
+ for the `CustomEngineConfig` shape.
148
+
106
149
  ## Default push policy
107
150
 
108
151
  | Event | Default |
@@ -211,13 +254,13 @@ Every JSON artifact in the ledger carries a `schema_version` field (currently
211
254
  | Endpoint | Returns |
212
255
  |---|---|
213
256
  | `GET /autoloop/list` | `{ ok, runs: AutoloopState[] }` |
214
- | `POST /autoloop/new` | `{ ok, run_id, planner_session }` — body `{ workspace, run_id?, planner_model?, send_timeout_ms? }` |
257
+ | `POST /autoloop/new` | `{ ok, run_id, planner_session }` — body `{ workspace, run_id?, planner_engine?, planner_model?, planner_custom_engine?, coder_engine?, coder_model?, coder_custom_engine?, reviewer_engine?, reviewer_model?, reviewer_custom_engine?, send_timeout_ms? }` |
215
258
  | `GET /autoloop/<id>/state` | `{ ok, state: AutoloopState }` — also returns a `terminated`-state stub reconstructed from the registry for runs that aren't in this process's memory, so the dashboard can open historical runs without 404'ing. |
216
259
  | `GET /autoloop/<id>/push_log` | `{ ok, entries: PushLogEntry[] }` — served from the ledger via `autoloopStatus`, so historical runs work the same as live ones. |
217
260
  | `GET /autoloop/<id>/chat_history` | `{ ok, entries: ChatEntry[] }` — replays `<ledger>/chat.jsonl`. The dashboard fetches this when opening a run so the Planner-pane conversation survives a page refresh / cross-process / re-opening a terminated run. Returns `[]` when the file doesn't exist (e.g. runs that predate the chat-history feature). |
218
261
  | `GET /autoloop/<id>/events` | SSE: `snapshot` / `message` / `state` / `push` / `iter_done` / `planner_reply` / `planner_error` / `coder_reply` / `reviewer_reply` / `terminated`. For runs that are NOT in this process's memory (terminated, or live in another process), the endpoint emits a single-shot `snapshot` + `terminated` then closes — the dashboard's existing handlers render history without hanging. |
219
262
  | `POST /autoloop/<id>/chat` | **202** `{ ok, queued: true }` — body `{ text }`. Fire-and-forget: the Planner's reply streams back via the `/events` SSE channel as a `planner_reply` event (or `planner_error` on failure); the HTTP response intentionally does NOT wait for it, because first-contact replies routinely exceed reverse-proxy idle limits (e.g. Cloudflare Tunnel cuts at ~100s → 524). 400 on empty text, 404 when the run is not in this process's memory. The MCP `autoloop_chat` tool path keeps the synchronous await-and-return-reply semantics (it runs in-process). |
220
- | `POST /autoloop/<id>/resume` | `{ ok, state }` — bring a terminated run back into this process. Reads the registry entry, re-creates dispatcher + runner; `ensurePlanner` picks up the persisted `claudeSessionId` (kept on disk because autoloop terminate now passes `keepPersisted: true`) so Claude resumes the original conversation. Runs that pre-date this change get a fresh Planner; the dashboard replays `chat.jsonl` visually anyway. 404 when the registry has no record. |
263
+ | `POST /autoloop/<id>/resume` | `{ ok, state }` — restore the role engine/model choices from the registry and re-create dispatcher + runner. Optional body fields `planner_custom_engine`, `coder_custom_engine`, `reviewer_custom_engine` must be supplied again for roles using `custom` because configs are intentionally not persisted. Existing engine-specific conversation resume behavior is reused where supported; `chat.jsonl` remains the visual history fallback. 404 when the registry has no record. |
221
264
  | `POST /autoloop/<id>/delete` | `{ ok }` — stops the runner if still alive, scrubs the row from `~/.claw-orchestrator/autoloop-registry.jsonl`, and purges `persistedSessions` so the run cannot be `/resume`'d back. The ledger directory under `<workspace>/tasks/<run_id>/` is kept on disk. 404 if the run was not present in either memory or the registry. |
222
265
 
223
266
  The 3-pane UI consumes these endpoints:
@@ -2,11 +2,12 @@
2
2
 
3
3
  This document tracks which Claude Code CLI version Claw Orchestrator is currently synced to, and which features have been integrated.
4
4
 
5
- ## Currently tracked: **Claude Code CLI 2.1.206** (as of 2026-07-10, plugin v4.7.0)
5
+ ## Currently tracked: **Claude Code CLI 2.1.207** (as of 2026-07-12, plugin v4.8.0)
6
6
 
7
7
  ## Sync history
8
8
 
9
9
  | Plugin Version | Claude CLI Version | Date | Notable integrations |
10
+ | v4.8.0 | 2.1.207 | 2026-07-12 | **Autoloop role-level multi-engine support.** Planner, Coder, and Reviewer can select independent engines/models while preserving the Claude defaults. Built-in non-Claude Planners use native read-only/plan modes and receive their role protocol in-band. Spawn selections persist across resume; Codex persists its real thread ID. Runtime and invocation checks used Claude Code 2.1.207 and Codex 0.144.1. |
10
11
  | v4.7.0 | 2.1.206 | 2026-07-10 | **Antigravity engine ships + permission-mode sync.** Main feature is the community-contributed first-class `engine: 'agy'` (PR #71, reviewed + hardened: layered resume-ID gating, `agy/` prefix routing, shared `sanitize.ts` across all six engines, ENGINE_TYPES single source). Weekly CLI sync: CC 2.1.200 renamed the `default` permission mode to **`manual`** — verified against 2.1.206 that the choices are now `acceptEdits/auto/bypassPermissions/manual/dontAsk/plan`, that `default` is still accepted (hidden compat), and that **`delegate` is hard-rejected at spawn** — so PermissionMode gains `manual`, drops `delegate`, and agy/gemini map `manual` like `default` (→ `--sandbox`). Codex 0.143.0: empirically re-tested `-c model_reasoning_effort=max` — still 400-rejected for gpt-5.5 (the "first-class max" note is Bedrock GPT-5.6-only), so the `max`→`xhigh` map stays. GPT-5.6 Sol/Terra/Luna registered with official pricing ($5/$30, $2.50/$15, $1/$6; 1M/1M/400K ctx) after the user reported using it — it's a limited preview on API/Codex-auth paths (empirically: ChatGPT-account Codex auth gets a 400, which is why the first probe on this box misread it as Bedrock-only; lesson — an auth-path rejection is not model non-existence). Codex default stays gpt-5.5. Free upside: CC 2.1.205 fixed `--json-schema` invalid-schema silent fallback + `format` keyword rejection; CC 2.1.203 fixed background sessions dropping shell-exported `ANTHROPIC_BASE_URL`. Pins → CC 2.1.206 / Codex 0.143.0 (installed; npm has 0.144.1, exec surface unchanged per release notes). |
11
12
  | v4.6.0 | 2.1.199 | 2026-07-03 | **Model registry sync — Claude Fable 5.** Registered `claude-fable-5` (first Claude 5-family model, tier above Opus; standard $10/$50 per Mtok, cache read $1, full 1M context at standard rates per the official pricing page) with new `fable` alias; taught the `isClaudeModel`/`resolveProvider` heuristics to recognize `fable`/`mythos` strings (they only matched claude/opus/sonnet/haiku). Mythos 5 not listed (same price, limited availability). CC 2.1.198–199 are subagent/background-agent reliability fixes — no invocation-surface change; free upside for us: subagent partial output on rate-limit/server error is now returned instead of silently dropped, and API errors in subagents are reported to the parent. Codex unchanged (0.142.5 is a log-scrub patch; pin stays 0.142.4 as installed). |
12
13
  | v4.5.0 | 2.1.197 | 2026-07-01 | **Model registry sync — Claude Sonnet 5 + gpt-5.5 pricing.** CLI 2.1.197 shipped Sonnet 5 as the new default (native 1M-token context; standard $3/$15 per Mtok, launch promo $2/$10 through 2026-08-31 — we price the standard rate). Registered `claude-sonnet-5` in `models.ts` and moved the `sonnet` alias to it (was pinned to `claude-sonnet-4-6`), so `--model sonnet` tracks the CLI's own default and cost/context accounting stays correct; `claude-sonnet-4-6` stays selectable by full id. Also corrected `gpt-5.5` (the default Codex model) from placeholder pricing to OpenAI's published $5/$30 per Mtok + 1M context, and updated docs/examples off `gpt-5.4`. The CC 2.1.179→2.1.197 and Codex 0.138→0.142.x ranges are otherwise bug-fix / TUI / remote-executor / plugin-marketplace work that doesn't touch our invocation flags or the stream-json / codex-exec event schema — no wrapper change. Free upside (no code change): 2.1.181 fixed prompt-caching on custom `ANTHROPIC_BASE_URL` (helps proxy mode), 2.1.187 fixed `--json-schema` StructuredOutput infinite-recall, 2.1.196 turned the 5-min streaming idle watchdog on by default; Codex 0.139 preserves `oneOf`/`allOf` in `--output-schema`. Bumped tested versions Claude 2.1.197 / Codex 0.142.4. |
@@ -44,7 +44,7 @@ The server exposes an OpenAI-compatible chat completions endpoint, enabling any
44
44
  - `gpt-*` → Codex engine
45
45
  - `composer-*` → Cursor engine
46
46
  - `gemini-3.5-flash`, `gemini-3.1-pro`, `agy-*`, `agy/*` → Antigravity (`agy`) engine
47
- - other `gemini-*` → Gemini engine
47
+ - other `gemini-*` → the legacy `gemini` engine (Gemini CLI is sunset; prefer `agy`)
48
48
 
49
49
  **CORS:** `/v1/` paths allow cross-origin requests by default. Set `OPENCLAW_CORS_ORIGINS=*` to allow all origins on all paths.
50
50
 
@@ -61,7 +61,7 @@ clawo session-start [name] [options]
61
61
  | Flag | Description |
62
62
  |------|-------------|
63
63
  | `-d, --cwd <dir>` | Working directory |
64
- | `-e, --engine <engine>` | Engine: `claude` (default), `codex`, `codex-app`, `gemini`, `agy`, `cursor`, `opencode`, or `custom` |
64
+ | `-e, --engine <engine>` | Engine: `claude` (default), `codex`, `codex-app`, `agy`, `cursor`, `opencode`, or `custom` |
65
65
  | `-m, --model <model>` | Model name or alias |
66
66
  | `--permission-mode <mode>` | `acceptEdits`, `plan`, `auto`, `bypassPermissions`, `manual`, `dontAsk` |
67
67
  | `--effort <level>` | `low`, `medium`, `high`, `max`, `auto` |
@@ -106,7 +106,7 @@ Agents can use different engines and models:
106
106
  "agents": [
107
107
  { "name": "Claude", "emoji": "🎭", "engine": "claude", "model": "opus", "persona": "Deep reasoning" },
108
108
  { "name": "Codex", "emoji": "🧠", "engine": "codex", "model": "gpt-5.5", "persona": "Fast implementation" },
109
- { "name": "Gemini", "emoji": "💎", "engine": "gemini", "model": "gemini-3.1-pro-preview", "persona": "Creative solutions" }
109
+ { "name": "Antigravity", "emoji": "💎", "engine": "agy", "model": "agy-pro", "persona": "Creative solutions" }
110
110
  ]
111
111
  }
112
112
  ```
@@ -23,7 +23,7 @@ openclaw plugins install @enderfga/claw-orchestrator --dangerously-force-unsafe-
23
23
  openclaw gateway restart
24
24
  ```
25
25
 
26
- > **Why `--dangerously-force-unsafe-install`?** Claw Orchestrator spawns Claude Code / Codex / Gemini / Cursor Agent / OpenCode CLI subprocesses via `child_process`, which OpenClaw's security scanner flags by design. The flag is required — there is no way to drive coding CLIs without process spawning.
26
+ > **Why `--dangerously-force-unsafe-install`?** Claw Orchestrator spawns Claude Code / Codex / Antigravity / Cursor Agent / OpenCode CLI subprocesses via `child_process`, which OpenClaw's security scanner flags by design. The flag is required — there is no way to drive coding CLIs without process spawning.
27
27
 
28
28
  Agents automatically get access to all session, council, and management tools.
29
29
 
@@ -52,7 +52,7 @@ await manager.stopSession('backend-fix');
52
52
  - **Claude Code CLI >= 2.1** — `npm install -g @anthropic-ai/claude-code`
53
53
  - **OpenClaw >= 2026.3.0** — for plugin mode (optional)
54
54
  - **OpenAI Codex CLI >= 0.112** — `npm install -g @openai/codex` (optional, for codex engine)
55
- - **Gemini CLI >= 0.35** — `npm install -g @google/gemini-cli` (optional, for gemini engine)
55
+ - **Antigravity CLI** — `curl -fsSL https://antigravity.google/cli/install.sh | bash` (optional, for the `agy` engine — Google's successor to the sunset Gemini CLI)
56
56
 
57
57
  ### Engine Authentication
58
58
 
@@ -60,7 +60,7 @@ Each engine requires its own authentication before use:
60
60
 
61
61
  - **Claude Code** — run `claude /login` or set `ANTHROPIC_API_KEY`
62
62
  - **Codex** — run `codex login` or set `OPENAI_API_KEY`
63
- - **Gemini** — run `gemini login` or set `GEMINI_API_KEY`
63
+ - **Antigravity** — run `agy` once and complete the Google OAuth login
64
64
 
65
65
  The plugin does not manage authentication — it expects each CLI to be ready to run.
66
66
 
@@ -87,7 +87,7 @@ Quick config for any client:
87
87
  |---------|-------|
88
88
  | API Base URL | `http://127.0.0.1:18796/v1` |
89
89
  | API Key | The value of `OPENCLAW_SERVER_TOKEN`, or any string if auth is disabled |
90
- | Model | `claude-fable-5`, `claude-opus-4-8`, `claude-sonnet-5`, `gpt-5.5`, `gemini-3.1-pro-preview`, etc. |
90
+ | Model | `claude-fable-5`, `claude-opus-4-8`, `claude-sonnet-5`, `gpt-5.5`, `agy-pro`, etc. |
91
91
 
92
92
  See [openai-compat.md](./openai-compat.md) for the full session-keying rules, `X-Session-Reset` semantics, the legacy-heuristic env var, and the `/v1/sessions` inspection endpoint.
93
93
 
@@ -12,8 +12,6 @@ SessionManager
12
12
  │ └── Wraps: codex exec --sandbox workspace-write --json (per-message spawning)
13
13
  ├── engine: 'codex-app' → PersistentCodexAppServerSession
14
14
  │ └── Wraps: codex app-server --listen stdio:// (long-running JSON-RPC; required for /goal)
15
- ├── engine: 'gemini' → PersistentGeminiSession
16
- │ └── Wraps: gemini -p --output-format stream-json (per-message spawning)
17
15
  ├── engine: 'agy' → PersistentAgySession
18
16
  │ └── Wraps: agy -p (Google Antigravity CLI, per-message spawning, plain-text output)
19
17
  ├── engine: 'cursor' → PersistentCursorSession
@@ -28,7 +26,7 @@ SessionManager
28
26
 
29
27
  ### Claude Code (`engine: 'claude'`)
30
28
 
31
- Default engine. Long-running subprocess with streaming JSON I/O. Tested with Claude Code CLI **2.1.206**.
29
+ Default engine. Long-running subprocess with streaming JSON I/O. Tested with Claude Code CLI **2.1.207**.
32
30
 
33
31
  - Persistent multi-turn conversations
34
32
  - Real-time streaming (text, tool_use, tool_result, system events)
@@ -54,7 +52,7 @@ await manager.startSession({
54
52
 
55
53
  ### OpenAI Codex (`engine: 'codex'`)
56
54
 
57
- Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested with `codex` CLI **0.143.0**.
55
+ Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested with `codex` CLI **0.144.1**.
58
56
 
59
57
  - Non-interactive execution via `codex exec --sandbox workspace-write --json` (replaces the deprecated `--full-auto` flag from earlier Codex versions)
60
58
  - Real per-turn `usage` from the `turn.completed` JSON event (input, output, cached, reasoning tokens)
@@ -63,6 +61,7 @@ Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested wi
63
61
  - `codexProfile` → `--profile <name>` (named config profile from `~/.codex/config.toml`)
64
62
  - Per-session continuity: the `thread_id` from the first turn's `thread.started` event is captured and reused via `codex exec resume <id>` for subsequent sends, so the model sees prior turns
65
63
  - One-shot execution per message (no persistent subprocess between sends)
64
+ - Captures the real Codex thread ID and persists it, so later sends and process-level session resume use `codex exec resume <thread_id>`
66
65
  - Working directory passed via `-C` flag
67
66
  - Default model: `gpt-5.5`
68
67
  - Requires `codex` CLI >= 0.119 (for `exec resume`): `npm install -g @openai/codex`
@@ -106,33 +105,11 @@ await manager.startSession({
106
105
  // await tool('codex_goal_set', { name: 'codex-goal-task', objective: 'build a tic-tac-toe app' });
107
106
  ```
108
107
 
109
- ### Google Gemini (`engine: 'gemini'`)
110
-
111
- Wraps the `gemini` CLI with `--output-format stream-json`. Each `send()` spawns a new process.
112
-
113
- - One-shot execution per message (no persistent subprocess)
114
- - Working directory carries accumulated changes across sends
115
- - Real token counts from stream-json `result` events (not estimated)
116
- - Permission modes: `bypassPermissions` → `--yolo`, `default` → `--sandbox`
117
- - Always passes `--skip-trust` to bypass the "trusted folders" gate introduced
118
- in Gemini CLI 0.43 (otherwise headless runs in worktrees / arbitrary cwds
119
- abort before producing output)
120
- - Requires `gemini` CLI installed: `npm install -g @google/gemini-cli`
121
-
122
- ```typescript
123
- await manager.startSession({
124
- name: 'gemini-task',
125
- engine: 'gemini',
126
- model: 'gemini-3.1-pro-preview',
127
- cwd: '/project',
128
- });
129
- ```
130
-
131
108
  ### Google Antigravity (`engine: 'agy'`)
132
109
 
133
110
  Wraps Google's **Antigravity CLI** (`agy`) — the successor to Gemini CLI (consumer
134
111
  Gemini CLI tiers stopped serving 2026-06-18). Each `send()` spawns a new process
135
- in print mode. Verified against `agy` **1.0.16**.
112
+ in print mode. Verified against `agy` **1.1.1**.
136
113
 
137
114
  - One-shot execution per message (no persistent subprocess)
138
115
  - **Plain-text output** — agy has no structured/stream-json mode, so stdout is
@@ -143,9 +120,10 @@ in print mode. Verified against `agy` **1.0.16**.
143
120
  externally via `resumeSessionId` (bare UUID only); read it back from
144
121
  `getStats().agyConversationId`
145
122
  - Permission modes: `bypassPermissions` → `--dangerously-skip-permissions`,
146
- `default` → `--sandbox` (terminal-restricted). Other modes run agy's own
147
- approval flow, which blocks in headless print mode — use `bypassPermissions`
148
- for autonomous work
123
+ `default` → `--sandbox` (terminal-restricted), and
124
+ `sandboxMode: 'read-only'` → `--mode plan` (takes precedence). Other modes
125
+ run agy's own approval flow, which blocks in headless print mode — use
126
+ `bypassPermissions` for autonomous write-enabled work
149
127
  - agy enforces its own print timeout (default 5m); the engine derives
150
128
  `--print-timeout` from the send timeout so the wrapper timer decides
151
129
  - Unknown `--model` slugs do **not** error — agy silently falls back to its
@@ -167,14 +145,22 @@ await manager.startSession({
167
145
  });
168
146
  ```
169
147
 
148
+ > **Legacy: `engine: 'gemini'`.** Google sunset the consumer Gemini CLI (tiers stopped
149
+ > serving 2026-06-18) in favour of Antigravity. The `gemini` engine still exists and
150
+ > still works — existing callers are not broken, and `gemini-*` model strings outside
151
+ > agy's registered slugs still route to it — but it is no longer a documented option,
152
+ > is not version-tracked, and gets no new work. Use `agy` for Google. (Unrelated: the
153
+ > multi-model **proxy** still talks to the Gemini **API**; that is a different
154
+ > subsystem and is unaffected.)
155
+
170
156
  ### Cursor Agent (`engine: 'cursor'`)
171
157
 
172
- Wraps the Cursor Agent CLI (`agent`) with `--print --force --output-format stream-json`. Each `send()` spawns a new process.
158
+ Wraps the Cursor Agent CLI (`agent`) with `--print --output-format stream-json`. Write-enabled sessions use `--force`. Each `send()` spawns a new process.
173
159
 
174
160
  - One-shot execution per message (no persistent subprocess)
175
161
  - Working directory via `--workspace` flag
176
162
  - Real token counts from stream-json `result` events (camelCase: `inputTokens`, `outputTokens`, `cacheReadTokens`)
177
- - `--force` enables auto-approval of all file changes
163
+ - `--force` enables auto-approval of file changes. `sandboxMode: 'read-only'` does **not** use `--force`; it enforces read-only via a binding `.cursor/cli.json` deny config (`Write`/`Edit`/`Shell` denied) written into an isolated temp dir used as the process cwd, with `--workspace` pointing at the real project (the repo tree is never modified). `--mode plan` is passed too as model steering, but the deny config is the actual boundary — plan mode alone is model-cooperative and was verified to let an adversarial prompt write. Do not add `--sandbox` (it does not restrict in-workspace writes and overrides the mode). Read/grep/search remain available
178
164
  - `--trust` auto-trusts the workspace without prompting
179
165
  - Cursor uses its own model routing (e.g., `sonnet-4`, `gpt-5`, `auto`)
180
166
  - Requires Cursor Agent CLI: `curl https://cursor.com/install -fsSL | bash`
@@ -200,6 +186,7 @@ Wraps the [sst/opencode](https://github.com/sst/opencode) CLI with `run --format
200
186
  - Real token counts from `step_finish.part.tokens.{input,output,cache.read}`
201
187
  - The wrapper closes the subprocess's stdin immediately after spawn (opencode otherwise reads stdin and blocks on EOF, hanging the call)
202
188
  - Provider-agnostic: opencode's `--model` expects `provider/model` form (e.g. `anthropic/claude-sonnet-4`). The wrapper passes `--model` through only when the value contains a `/`; otherwise opencode's own default applies
189
+ - `sandboxMode: 'read-only'` spawns a generated `clawo-readonly` agent (`--agent clawo-readonly` plus an `OPENCODE_CONFIG_CONTENT` env var defining it) whose permissions deny `edit` / `bash` / `external_directory`. It deliberately does **not** use OpenCode's built-in `plan` agent: that is a user-overridable preset whose compiled rules start with `{"permission":"*","action":"allow"}` and deny neither `bash` nor `edit`, so a "read-only" session could still author files through a shell heredoc. Verified against opencode 1.17.15 by attempting a real write, which the agent cannot perform (its toolset has no write/bash tools). Note that `opencode agent list` renders compiled permission rules and does not show the tool-level restriction — trust an actual write attempt, not that view
203
190
  - Requires opencode installed: `brew install sst/tap/opencode` or `npm install -g opencode-ai`. Auth via `opencode auth login` **or** any provider env var (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GEMINI_API_KEY`, etc.) — opencode picks up either path
204
191
  - Binary: `opencode` (set `OPENCODE_BIN` env var to override)
205
192
 
@@ -262,7 +249,7 @@ Team tools (`team_list`, `team_send`) operate on the same virtual-team layer for
262
249
  |--------|------------|-------------|
263
250
  | Claude | Lists other active SessionManager sessions | Routes via cross-session inbox |
264
251
  | Codex | Lists other active SessionManager sessions | Routes via cross-session inbox |
265
- | Gemini | Lists other active SessionManager sessions | Routes via cross-session inbox |
252
+ | Antigravity | Lists other active SessionManager sessions | Routes via cross-session inbox |
266
253
  | Cursor | Lists other active SessionManager sessions | Routes via cross-session inbox |
267
254
 
268
255
  Messages are delivered via the inbox system — idle sessions receive immediately, busy sessions queue for later delivery.
@@ -318,7 +305,7 @@ Integrate **any** coding agent CLI without writing engine-specific code. You pro
318
305
 
319
306
  Two protocol modes:
320
307
  - **Persistent** (`persistent: true`) — long-running subprocess with stream-json I/O over stdin/stdout (like Claude Code)
321
- - **One-shot** (`persistent: false`, default) — new process spawned per `send()` (like Gemini/Codex)
308
+ - **One-shot** (`persistent: false`, default) — new process spawned per `send()` (like Codex/Antigravity)
322
309
 
323
310
  ### CustomEngineConfig
324
311
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  > **Cost warning**: This bridge routes requests through the Claude Code CLI, which uses your Claude Max subscription's **extra usage** quota. When OpenClaw's agent loop sends its system prompt (with distinctive tool definitions and agent instructions), Anthropic's backend recognizes this as programmatic/agent traffic and bills it against extra usage — **not** the included allowance. This is by design: the bridge does NOT bypass Anthropic's billing or subscription enforcement. Using it as OpenClaw's primary model backend means every agent turn consumes extra usage credits at standard API rates ($15/M input, $75/M output for Opus). Monitor your usage at [claude.ai/settings/usage](https://claude.ai/settings/usage).
4
4
 
5
- The embedded server exposes a drop-in OpenAI-compatible endpoint so any client that speaks `/v1/chat/completions` can talk to a persistent Claude Code (or Codex / Gemini / Cursor) session. The bridge is designed to serve **two kinds of clients as first-class citizens**:
5
+ The embedded server exposes a drop-in OpenAI-compatible endpoint so any client that speaks `/v1/chat/completions` can talk to a persistent Claude Code (or Codex / Antigravity / Cursor) session. The bridge is designed to serve **two kinds of clients as first-class citizens**:
6
6
 
7
7
  1. **Upstream agents** that maintain their own conversation state and forward only the latest user turn — OpenClaw's main agent loop, cron jobs, subagents, programmatic clients.
8
8
  2. **OpenAI-compatible webchat / labeling tools** that re-send the full transcript on every turn — ChatGPT-Next-Web, Open WebUI, LobeChat, data-labeling pipelines.
@@ -29,8 +29,8 @@ Key options:
29
29
 
30
30
  | Option | Description |
31
31
  |--------|-------------|
32
- | `engine` | `'claude'` (default), `'codex'`, `'codex-app'`, `'gemini'`, `'agy'`, `'cursor'`, `'opencode'`, or `'custom'` — see [Multi-Engine](./multi-engine.md) |
33
- | `model` | Model alias (`opus`, `sonnet`, `haiku`, `gemini-pro`) or full name |
32
+ | `engine` | `'claude'` (default), `'codex'`, `'codex-app'`, `'agy'`, `'cursor'`, `'opencode'`, or `'custom'` — see [Multi-Engine](./multi-engine.md) |
33
+ | `model` | Model alias (`fable`, `opus`, `sonnet`, `haiku`, `agy-pro`) or full name |
34
34
  | `permissionMode` | `acceptEdits`, `bypassPermissions`, `plan`, `auto`, `manual`, `dontAsk` (`default` = legacy alias for `manual`) |
35
35
  | `effort` | `low`, `medium`, `high`, `max`, `auto` |
36
36
  | `bare` | Skip hooks, LSP, auto-memory, CLAUDE.md |
@@ -162,7 +162,7 @@ SessionManager tracks consecutive failures per engine type. After 3 consecutive
162
162
 
163
163
  ## Orphaned Process Cleanup
164
164
 
165
- If the plugin crashes without calling `stop()`, child CLI processes (claude, codex, gemini, agent) may become orphans. SessionManager tracks PIDs in `~/.openclaw/session-pids.json` and cleans up stale processes on startup:
165
+ If the plugin crashes without calling `stop()`, child CLI processes (claude, codex, agy, agent, opencode) may become orphans. SessionManager tracks PIDs in `~/.openclaw/session-pids.json` and cleans up stale processes on startup:
166
166
 
167
167
  1. Reads PID file from previous run
168
168
  2. For each PID, checks if process is alive (`kill -0`)
@@ -12,9 +12,10 @@ Start a persistent coding session with full CLI flag support.
12
12
  |-----------|------|-------------|
13
13
  | `name` | string | Session name (auto-generated if omitted) |
14
14
  | `cwd` | string | Working directory |
15
- | `engine` | `'claude'` \| `'codex'` \| `'codex-app'` \| `'gemini'` \| `'agy'` \| `'cursor'` \| `'opencode'` \| `'custom'` | Engine to use (default: `claude`). `agy` wraps Google Antigravity CLI. `opencode` wraps sst/opencode (pass model as `provider/model`). Use `custom` with `customEngine` for any CLI. |
15
+ | `engine` | `'claude'` \| `'codex'` \| `'codex-app'` \| `'agy'` \| `'cursor'` \| `'opencode'` \| `'custom'` | Engine to use (default: `claude`). `agy` wraps Google Antigravity CLI. `opencode` wraps sst/opencode (pass model as `provider/model`). Use `custom` with `customEngine` for any CLI. (`'gemini'` is still accepted for existing callers, but Gemini CLI is sunset — use `agy` for Google.) |
16
16
  | `model` | string | Model alias or full name |
17
17
  | `permissionMode` | string | `acceptEdits`, `bypassPermissions`, `plan`, `auto`, `manual`, `dontAsk` (`default` = legacy alias for `manual`) |
18
+ | `sandboxMode` | `'read-only'` \| `'workspace-write'` \| `'danger-full-access'` | Sandbox policy. Codex supports all values. `read-only` is enforced on every other built-in engine too: Claude → plan mode; Antigravity / Cursor → their plan modes; OpenCode → a generated `clawo-readonly` agent denying `edit`/`bash`. A `custom` engine must map it via `permissionModes`, or the session refuses to start. Persisted across session resume. |
18
19
  | `effort` | string | `low`, `medium`, `high`, `max`, `auto` |
19
20
  | `allowedTools` | string[] | Tools to auto-approve |
20
21
  | `disallowedTools` | string[] | Tools to deny |
@@ -537,57 +538,75 @@ Three-agent autonomous iteration loop (Planner / Coder / Reviewer) over a git wo
537
538
 
538
539
  ### `autoloop_start`
539
540
 
540
- Start an autoloop run. Planner is created persistent; Coder + Reviewer are spawned by the Planner once `plan.md` is ready.
541
+ Start a chat-mode autoloop. Planner starts immediately; Coder + Reviewer start only after the Planner receives plan approval and emits `spawn_subagents`.
541
542
 
542
543
  | Parameter | Type | Required | Description |
543
544
  |-----------|------|----------|-------------|
544
- | `cwd` | string | yes | Workspace (must be a git repo) |
545
- | `goal` | string | yes | High-level user goal in natural language |
546
- | `model` | string | | Planner model (default Opus) |
547
- | `coderModel` | string | | Coder subagent model |
548
- | `reviewerModel` | string | | Reviewer subagent model |
549
- | `maxIters` | number | | Cap on Coder/Reviewer rounds (default 50) |
550
- | `pushChannels` | string[] | | Notification channels (`wechat`, `whatsapp`, `email`) |
545
+ | `run_id` | string | yes | Stable run identifier |
546
+ | `workspace` | string | yes | Git workspace path |
547
+ | `planner_engine` | EngineType | | Planner engine (default `claude`) |
548
+ | `planner_model` | string | | Planner model (Claude default `opus`; other engines use their own default when omitted) |
549
+ | `planner_custom_engine` | object | | Trusted `CustomEngineConfig` when Planner engine is `custom`. **Local callers only** — see below |
550
+ | `coder_engine` | EngineType | | Default Coder engine (default `claude`) |
551
+ | `coder_model` | string | | Default Coder model (Claude default `sonnet`) |
552
+ | `coder_custom_engine` | object | | Trusted config when Coder may use `custom`. **Local callers only** |
553
+ | `reviewer_engine` | EngineType | | Default Reviewer engine (default `claude`) |
554
+ | `reviewer_model` | string | | Default Reviewer model (Claude default `sonnet`) |
555
+ | `reviewer_custom_engine` | object | | Trusted config when Reviewer may use `custom`. **Local callers only** |
556
+ | `send_timeout_ms` | number | | Per-message timeout (default 600000) |
557
+
558
+ > **Custom engines are local-only.** A `CustomEngineConfig` names an executable to
559
+ > spawn plus its argv and env, so it may only be supplied by a caller that already
560
+ > runs on the host (this MCP tool, or the `SessionManager` API). The HTTP API
561
+ > (`POST /autoloop/new`, `POST /autoloop/<id>/resume`) rejects a `*_custom_engine`
562
+ > body field with a 400 — the embedded server is often reverse-tunnelled and its
563
+ > token is a monitoring credential, not permission to choose what binary runs.
564
+ > Built-in engines are fully selectable over HTTP.
565
+
566
+ Custom configs are not persisted or accepted from Planner output. See [`multi-engine.md`](./multi-engine.md) for their shape.
551
567
 
552
568
  ### `autoloop_chat`
553
569
 
554
- Send a message into the Planner conversation (e.g. answer a clarifying question, refine the plan, kick off the subloop).
570
+ Send a message into the Planner conversation.
555
571
 
556
572
  | Parameter | Type | Required |
557
573
  |-----------|------|----------|
558
- | `id` | string | yes |
559
- | `message` | string | yes |
574
+ | `run_id` | string | yes |
575
+ | `text` | string | yes |
560
576
 
561
577
  ### `autoloop_status`
562
578
 
563
- Get current state, phase, recent inbox messages, and ledger summary.
579
+ Get the current state and push log.
564
580
 
565
581
  | Parameter | Type | Required |
566
582
  |-----------|------|----------|
567
- | `id` | string | yes |
583
+ | `run_id` | string | yes |
568
584
 
569
585
  ### `autoloop_list`
570
586
 
571
- List active and recent autoloop runs (in-memory + on-disk registry, deduped by run_id).
587
+ List active Autoloop runs in this manager process.
572
588
 
573
589
  (no params)
574
590
 
575
591
  ### `autoloop_reset_agent`
576
592
 
577
- Reset one of the subagent sessions (Coder or Reviewer) without losing Planner state — useful when a subagent loops on a stale belief.
593
+ Reset one role session while retaining the role's current engine/model selection.
578
594
 
579
595
  | Parameter | Type | Required | Description |
580
596
  |-----------|------|----------|-------------|
581
- | `id` | string | yes | Run id |
582
- | `agent` | `'coder'` \| `'reviewer'` | yes | Which subagent to reset |
597
+ | `run_id` | string | yes | Run id |
598
+ | `agent` | `'planner'` \| `'coder'` \| `'reviewer'` | yes | Role to reset; Planner requires `force: true` |
599
+ | `force` | boolean | | Allow Planner reset |
600
+ | `eager_restart` | boolean | | Start the replacement session immediately |
583
601
 
584
602
  ### `autoloop_stop`
585
603
 
586
- Terminate the run. All sessions are stopped and ledger state is finalised.
604
+ Terminate the run and stop all role sessions.
587
605
 
588
606
  | Parameter | Type | Required |
589
607
  |-----------|------|----------|
590
- | `id` | string | yes |
608
+ | `run_id` | string | yes |
609
+ | `reason` | string | |
591
610
 
592
611
  ---
593
612