@enderfga/claw-orchestrator 4.6.0 → 4.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. package/README.md +6 -5
  2. package/configs/autoloop-coder-prompt.md +4 -4
  3. package/configs/autoloop-planner-prompt.md +10 -10
  4. package/configs/autoloop-reviewer-prompt.md +3 -4
  5. package/dist/bin/cli.js +1 -1
  6. package/dist/bin/cli.js.map +1 -1
  7. package/dist/src/agy-conversation.d.ts +2 -0
  8. package/dist/src/agy-conversation.js +10 -0
  9. package/dist/src/agy-conversation.js.map +1 -0
  10. package/dist/src/autoloop/dispatcher.d.ts +59 -10
  11. package/dist/src/autoloop/dispatcher.js +261 -40
  12. package/dist/src/autoloop/dispatcher.js.map +1 -1
  13. package/dist/src/autoloop/planner-tools.d.ts +4 -3
  14. package/dist/src/autoloop/planner-tools.js +35 -2
  15. package/dist/src/autoloop/planner-tools.js.map +1 -1
  16. package/dist/src/autoloop/types.d.ts +2 -0
  17. package/dist/src/autoloop/types.js.map +1 -1
  18. package/dist/src/dashboard/index.html +185 -94
  19. package/dist/src/embedded-server.js +67 -9
  20. package/dist/src/embedded-server.js.map +1 -1
  21. package/dist/src/index.d.ts +1 -0
  22. package/dist/src/index.js +98 -61
  23. package/dist/src/index.js.map +1 -1
  24. package/dist/src/models.js +68 -3
  25. package/dist/src/models.js.map +1 -1
  26. package/dist/src/persistent-agy-session.d.ts +60 -0
  27. package/dist/src/persistent-agy-session.js +228 -0
  28. package/dist/src/persistent-agy-session.js.map +1 -0
  29. package/dist/src/persistent-codex-session.js +3 -0
  30. package/dist/src/persistent-codex-session.js.map +1 -1
  31. package/dist/src/persistent-cursor-session.js +9 -6
  32. package/dist/src/persistent-cursor-session.js.map +1 -1
  33. package/dist/src/persistent-custom-session.js +18 -32
  34. package/dist/src/persistent-custom-session.js.map +1 -1
  35. package/dist/src/persistent-gemini-session.d.ts +4 -0
  36. package/dist/src/persistent-gemini-session.js +58 -7
  37. package/dist/src/persistent-gemini-session.js.map +1 -1
  38. package/dist/src/persistent-opencode-session.js +37 -7
  39. package/dist/src/persistent-opencode-session.js.map +1 -1
  40. package/dist/src/persistent-session.js +8 -9
  41. package/dist/src/persistent-session.js.map +1 -1
  42. package/dist/src/sanitize.d.ts +23 -0
  43. package/dist/src/sanitize.js +57 -0
  44. package/dist/src/sanitize.js.map +1 -0
  45. package/dist/src/session-manager.d.ts +49 -2
  46. package/dist/src/session-manager.js +254 -35
  47. package/dist/src/session-manager.js.map +1 -1
  48. package/dist/src/types.d.ts +8 -4
  49. package/dist/src/types.js +2 -0
  50. package/dist/src/types.js.map +1 -1
  51. package/package.json +1 -1
  52. package/skills/SKILL.md +17 -14
  53. package/skills/references/autoloop.md +50 -6
  54. package/skills/references/claude-cli-tracking.md +3 -1
  55. package/skills/references/cli.md +4 -3
  56. package/skills/references/mcp.md +1 -1
  57. package/skills/references/multi-engine.md +53 -8
  58. package/skills/references/sessions.md +2 -2
  59. package/skills/references/tools.md +40 -21
@@ -21,12 +21,40 @@ This page is the operator reference.
21
21
 
22
22
  ## Roles
23
23
 
24
- | Agent | Engine (default) | cwd | Owns |
24
+ | Agent | Default | cwd | Owns |
25
25
  |---|---|---|---|
26
26
  | **Planner** | claude / opus | workspace | strategy, `plan.md`, `goal.json`, talking to you |
27
- | **Coder** | claude / sonnet (override per spawn) | workspace | code changes, eval execution |
27
+ | **Coder** | claude / sonnet | workspace | code changes, eval execution |
28
28
  | **Reviewer** | claude / sonnet | `<workspace>/tasks/<run_id>/reviewer_sandbox/` | distrust audit; advance / hold / rollback |
29
29
 
30
+ Each role can use any built-in engine, or a `custom` engine config supplied by a
31
+ local caller (custom engines name an executable, so the HTTP API does not accept
32
+ them — see [tools.md](./tools.md)). If a non-Claude role omits `model`, that CLI
33
+ uses its own default model rather than receiving the Claude `opus` / `sonnet`
34
+ defaults. Role instructions are included in-band for engines that do not expose a
35
+ native system-prompt flag.
36
+
37
+ Engines without native multi-turn conversation (Gemini, Cursor, OpenCode, one-shot
38
+ custom engines) spawn a fresh process per send, so the dispatcher replays that
39
+ role's transcript in-band as a `<conversation_history>` block, oldest turns dropped
40
+ past a character budget. Claude, Codex and Antigravity keep context themselves and
41
+ get no replay.
42
+
43
+ The Planner runs read-only so strategy cannot turn into source edits, and that is
44
+ enforced by the engine rather than requested politely: Claude uses plan mode,
45
+ Gemini gets `--approval-mode plan` **plus an admin policy denying `exit_plan_mode`**
46
+ (plan mode on its own is model-cooperative — the agent can call that tool and walk
47
+ out of it), OpenCode gets a generated `clawo-readonly` agent that denies
48
+ `edit`/`bash`/`external_directory` (its built-in `plan` agent is a user-overridable
49
+ preset that denies neither), and Antigravity/Cursor use their plan modes. A custom
50
+ Planner receives `permissionMode: 'manual'` and its `CustomEngineConfig` **must**
51
+ map that mode to the CLI's read-only flag — if it cannot, the session refuses to
52
+ start rather than silently running write-enabled.
53
+
54
+ Coder and Reviewer engine/model choices can be overridden by the first successful
55
+ `spawn_subagents`; later attempts to change an already-started role are rejected
56
+ instead of silently diverging from the running session.
57
+
30
58
  Coder and Reviewer **never speak to you directly**. Anything they observe
31
59
  flows through the Planner. The Planner decides what to surface and what to
32
60
  absorb.
@@ -78,7 +106,7 @@ curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_stop \
78
106
 
79
107
  | Tool | Args | What |
80
108
  |---|---|---|
81
- | `autoloop_start` | `run_id`, `workspace`, `planner_model?`, `send_timeout_ms?` | Start a run; launches Planner session. |
109
+ | `autoloop_start` | `run_id`, `workspace`, per-role `*_engine?`, `*_model?`, `*_custom_engine?`, `send_timeout_ms?` | Start a run; launches Planner and stores Coder/Reviewer defaults. Each `custom` role requires its matching config. |
82
110
  | `autoloop_chat` | `run_id`, `text` | Send a chat message to the Planner; returns the Planner's reply. |
83
111
  | `autoloop_status` | `run_id` | Current state (status, iter, push count, subagents_spawned). |
84
112
  | `autoloop_list` | — | All active runs in this manager process. |
@@ -94,7 +122,7 @@ never see the JSON — only the Planner's narrative.
94
122
  | Tool | Args | What |
95
123
  |---|---|---|
96
124
  | `notify_user` | `level` ('info' / 'warn' / 'decision' / 'error'), `summary`, `detail?`, `channel?` ('auto' / 'wechat' / 'webchat' / 'both' / 'email') | Push you out-of-band. |
97
- | `spawn_subagents` | `coder_model?`, `reviewer_model?`, `initial_directive?` | Start Coder + Reviewer. Only after explicit user approval. |
125
+ | `spawn_subagents` | `coder_engine?`, `coder_model?`, `reviewer_engine?`, `reviewer_model?`, `initial_directive?` | Start Coder + Reviewer. Omitted values inherit run defaults. An engine change without a model uses the new engine's default. Once a role session has started, changing its engine/model is rejected. Custom configs cannot be emitted by Planner. Only after explicit user approval. |
98
126
  | `send_directive` | `goal`, `constraints?`, `success_criteria?`, `max_attempts?` | Next iter's instruction to Coder. |
99
127
  | `pause_loop` | `reason` | Halt subloop at next iter boundary; chat keeps working. |
100
128
  | `resume_loop` | — | Resume after pause. |
@@ -103,6 +131,22 @@ never see the JSON — only the Planner's narrative.
103
131
  | `write_plan` | `content` (full plan.md body), `commit_message?` | Write `plan.md` to the workspace and git-commit. The **only** way the Planner can author plan.md — Write/Edit are stripped from the Planner session as a hard role boundary. Re-running replaces the whole file. |
104
132
  | `write_goal` | `content` (full goal.json body), `commit_message?` | Same, for `goal.json`. Content is JSON-validated before write; malformed content errors back to the Planner. |
105
133
 
134
+ ### Custom engines and resume
135
+
136
+ Custom engine configs are accepted only by `autoloop_start` (or the HTTP resume
137
+ body), never through Planner output. This keeps config fields such as `env` and
138
+ static CLI arguments out of the Planner transcript and `decisions.jsonl`.
139
+ The central registry persists only each role's engine and model, including the
140
+ effective Coder/Reviewer selection after a successful spawn. Resume leaves the
141
+ prior append-only row untouched until startup succeeds, so a transient CLI
142
+ failure cannot erase the run. When resuming a run that uses `custom`, provide
143
+ the matching `planner_custom_engine`, `coder_custom_engine`, or
144
+ `reviewer_custom_engine` again; otherwise resume fails with a clear
145
+ configuration error rather than silently switching to Claude. Custom config
146
+ shape is validated at runtime, while its `env` and static CLI arguments remain
147
+ out of registry and audit records. See [`multi-engine.md`](./multi-engine.md)
148
+ for the `CustomEngineConfig` shape.
149
+
106
150
  ## Default push policy
107
151
 
108
152
  | Event | Default |
@@ -211,13 +255,13 @@ Every JSON artifact in the ledger carries a `schema_version` field (currently
211
255
  | Endpoint | Returns |
212
256
  |---|---|
213
257
  | `GET /autoloop/list` | `{ ok, runs: AutoloopState[] }` |
214
- | `POST /autoloop/new` | `{ ok, run_id, planner_session }` — body `{ workspace, run_id?, planner_model?, send_timeout_ms? }` |
258
+ | `POST /autoloop/new` | `{ ok, run_id, planner_session }` — body `{ workspace, run_id?, planner_engine?, planner_model?, planner_custom_engine?, coder_engine?, coder_model?, coder_custom_engine?, reviewer_engine?, reviewer_model?, reviewer_custom_engine?, send_timeout_ms? }` |
215
259
  | `GET /autoloop/<id>/state` | `{ ok, state: AutoloopState }` — also returns a `terminated`-state stub reconstructed from the registry for runs that aren't in this process's memory, so the dashboard can open historical runs without 404'ing. |
216
260
  | `GET /autoloop/<id>/push_log` | `{ ok, entries: PushLogEntry[] }` — served from the ledger via `autoloopStatus`, so historical runs work the same as live ones. |
217
261
  | `GET /autoloop/<id>/chat_history` | `{ ok, entries: ChatEntry[] }` — replays `<ledger>/chat.jsonl`. The dashboard fetches this when opening a run so the Planner-pane conversation survives a page refresh / cross-process / re-opening a terminated run. Returns `[]` when the file doesn't exist (e.g. runs that predate the chat-history feature). |
218
262
  | `GET /autoloop/<id>/events` | SSE: `snapshot` / `message` / `state` / `push` / `iter_done` / `planner_reply` / `planner_error` / `coder_reply` / `reviewer_reply` / `terminated`. For runs that are NOT in this process's memory (terminated, or live in another process), the endpoint emits a single-shot `snapshot` + `terminated` then closes — the dashboard's existing handlers render history without hanging. |
219
263
  | `POST /autoloop/<id>/chat` | **202** `{ ok, queued: true }` — body `{ text }`. Fire-and-forget: the Planner's reply streams back via the `/events` SSE channel as a `planner_reply` event (or `planner_error` on failure); the HTTP response intentionally does NOT wait for it, because first-contact replies routinely exceed reverse-proxy idle limits (e.g. Cloudflare Tunnel cuts at ~100s → 524). 400 on empty text, 404 when the run is not in this process's memory. The MCP `autoloop_chat` tool path keeps the synchronous await-and-return-reply semantics (it runs in-process). |
220
- | `POST /autoloop/<id>/resume` | `{ ok, state }` — bring a terminated run back into this process. Reads the registry entry, re-creates dispatcher + runner; `ensurePlanner` picks up the persisted `claudeSessionId` (kept on disk because autoloop terminate now passes `keepPersisted: true`) so Claude resumes the original conversation. Runs that pre-date this change get a fresh Planner; the dashboard replays `chat.jsonl` visually anyway. 404 when the registry has no record. |
264
+ | `POST /autoloop/<id>/resume` | `{ ok, state }` — restore the role engine/model choices from the registry and re-create dispatcher + runner. Optional body fields `planner_custom_engine`, `coder_custom_engine`, `reviewer_custom_engine` must be supplied again for roles using `custom` because configs are intentionally not persisted. Existing engine-specific conversation resume behavior is reused where supported; `chat.jsonl` remains the visual history fallback. 404 when the registry has no record. |
221
265
  | `POST /autoloop/<id>/delete` | `{ ok }` — stops the runner if still alive, scrubs the row from `~/.claw-orchestrator/autoloop-registry.jsonl`, and purges `persistedSessions` so the run cannot be `/resume`'d back. The ledger directory under `<workspace>/tasks/<run_id>/` is kept on disk. 404 if the run was not present in either memory or the registry. |
222
266
 
223
267
  The 3-pane UI consumes these endpoints:
@@ -2,11 +2,13 @@
2
2
 
3
3
  This document tracks which Claude Code CLI version Claw Orchestrator is currently synced to, and which features have been integrated.
4
4
 
5
- ## Currently tracked: **Claude Code CLI 2.1.199** (as of 2026-07-03, plugin v4.6.0)
5
+ ## Currently tracked: **Claude Code CLI 2.1.207** (as of 2026-07-12, plugin v4.8.0)
6
6
 
7
7
  ## Sync history
8
8
 
9
9
  | Plugin Version | Claude CLI Version | Date | Notable integrations |
10
+ | v4.8.0 | 2.1.207 | 2026-07-12 | **Autoloop role-level multi-engine support.** Planner, Coder, and Reviewer can select independent engines/models while preserving the Claude defaults. Built-in non-Claude Planners use native read-only/plan modes and receive their role protocol in-band. Spawn selections persist across resume; Codex persists its real thread ID. Runtime and invocation checks used Claude Code 2.1.207 and Codex 0.144.1. |
11
+ | v4.7.0 | 2.1.206 | 2026-07-10 | **Antigravity engine ships + permission-mode sync.** Main feature is the community-contributed first-class `engine: 'agy'` (PR #71, reviewed + hardened: layered resume-ID gating, `agy/` prefix routing, shared `sanitize.ts` across all six engines, ENGINE_TYPES single source). Weekly CLI sync: CC 2.1.200 renamed the `default` permission mode to **`manual`** — verified against 2.1.206 that the choices are now `acceptEdits/auto/bypassPermissions/manual/dontAsk/plan`, that `default` is still accepted (hidden compat), and that **`delegate` is hard-rejected at spawn** — so PermissionMode gains `manual`, drops `delegate`, and agy/gemini map `manual` like `default` (→ `--sandbox`). Codex 0.143.0: empirically re-tested `-c model_reasoning_effort=max` — still 400-rejected for gpt-5.5 (the "first-class max" note is Bedrock GPT-5.6-only), so the `max`→`xhigh` map stays. GPT-5.6 Sol/Terra/Luna registered with official pricing ($5/$30, $2.50/$15, $1/$6; 1M/1M/400K ctx) after the user reported using it — it's a limited preview on API/Codex-auth paths (empirically: ChatGPT-account Codex auth gets a 400, which is why the first probe on this box misread it as Bedrock-only; lesson — an auth-path rejection is not model non-existence). Codex default stays gpt-5.5. Free upside: CC 2.1.205 fixed `--json-schema` invalid-schema silent fallback + `format` keyword rejection; CC 2.1.203 fixed background sessions dropping shell-exported `ANTHROPIC_BASE_URL`. Pins → CC 2.1.206 / Codex 0.143.0 (installed; npm has 0.144.1, exec surface unchanged per release notes). |
10
12
  | v4.6.0 | 2.1.199 | 2026-07-03 | **Model registry sync — Claude Fable 5.** Registered `claude-fable-5` (first Claude 5-family model, tier above Opus; standard $10/$50 per Mtok, cache read $1, full 1M context at standard rates per the official pricing page) with new `fable` alias; taught the `isClaudeModel`/`resolveProvider` heuristics to recognize `fable`/`mythos` strings (they only matched claude/opus/sonnet/haiku). Mythos 5 not listed (same price, limited availability). CC 2.1.198–199 are subagent/background-agent reliability fixes — no invocation-surface change; free upside for us: subagent partial output on rate-limit/server error is now returned instead of silently dropped, and API errors in subagents are reported to the parent. Codex unchanged (0.142.5 is a log-scrub patch; pin stays 0.142.4 as installed). |
11
13
  | v4.5.0 | 2.1.197 | 2026-07-01 | **Model registry sync — Claude Sonnet 5 + gpt-5.5 pricing.** CLI 2.1.197 shipped Sonnet 5 as the new default (native 1M-token context; standard $3/$15 per Mtok, launch promo $2/$10 through 2026-08-31 — we price the standard rate). Registered `claude-sonnet-5` in `models.ts` and moved the `sonnet` alias to it (was pinned to `claude-sonnet-4-6`), so `--model sonnet` tracks the CLI's own default and cost/context accounting stays correct; `claude-sonnet-4-6` stays selectable by full id. Also corrected `gpt-5.5` (the default Codex model) from placeholder pricing to OpenAI's published $5/$30 per Mtok + 1M context, and updated docs/examples off `gpt-5.4`. The CC 2.1.179→2.1.197 and Codex 0.138→0.142.x ranges are otherwise bug-fix / TUI / remote-executor / plugin-marketplace work that doesn't touch our invocation flags or the stream-json / codex-exec event schema — no wrapper change. Free upside (no code change): 2.1.181 fixed prompt-caching on custom `ANTHROPIC_BASE_URL` (helps proxy mode), 2.1.187 fixed `--json-schema` StructuredOutput infinite-recall, 2.1.196 turned the 5-min streaming idle watchdog on by default; Codex 0.139 preserves `oneOf`/`allOf` in `--output-schema`. Bumped tested versions Claude 2.1.197 / Codex 0.142.4. |
12
14
  | v4.3.0 | 2.1.178 | 2026-06-16 | **Parity batch 2 + legacy-subsystem upgrades (local-only, no cloud).** Claude `--fallback-model` array form (CSV, verified via `claude --help`). Codex-app `codex_threads` (`thread/list`) and `thread/resume` on start when `resumeSessionId` is set (param shapes from `generate-json-schema`). Council agents gain per-agent `effort`/`ultracode`. `ultrareview` re-implemented on the new cross-engine `fanout` primitive (opt-in `engines`, default claude-only). Consensus parsing exposes match source for observability. Dropped on purpose: `codex cloud exec`/best-of-N (cloud/managed — loses local control), `--bg` (we own the subprocess), `--output-last-message` (we already capture final text), ultraplan `ultracode` (violates its plan-only contract). Autoloop mid-turn steer deferred — the loop is strictly sequential (Coder fully completes before the Reviewer runs), so steer would always fall back to a fresh turn; a real version needs concurrent review. |
@@ -43,7 +43,8 @@ The server exposes an OpenAI-compatible chat completions endpoint, enabling any
43
43
  - `claude-*`, `opus`, `sonnet`, `haiku` → Claude engine
44
44
  - `gpt-*` → Codex engine
45
45
  - `composer-*` → Cursor engine
46
- - `gemini-*` → Gemini engine
46
+ - `gemini-3.5-flash`, `gemini-3.1-pro`, `agy-*`, `agy/*` → Antigravity (`agy`) engine
47
+ - other `gemini-*` → Gemini engine
47
48
 
48
49
  **CORS:** `/v1/` paths allow cross-origin requests by default. Set `OPENCLAW_CORS_ORIGINS=*` to allow all origins on all paths.
49
50
 
@@ -60,9 +61,9 @@ clawo session-start [name] [options]
60
61
  | Flag | Description |
61
62
  |------|-------------|
62
63
  | `-d, --cwd <dir>` | Working directory |
63
- | `-e, --engine <engine>` | Engine: `claude` (default), `codex`, or `gemini` |
64
+ | `-e, --engine <engine>` | Engine: `claude` (default), `codex`, `codex-app`, `gemini`, `agy`, `cursor`, `opencode`, or `custom` |
64
65
  | `-m, --model <model>` | Model name or alias |
65
- | `--permission-mode <mode>` | `acceptEdits`, `plan`, `auto`, `bypassPermissions` |
66
+ | `--permission-mode <mode>` | `acceptEdits`, `plan`, `auto`, `bypassPermissions`, `manual`, `dontAsk` |
66
67
  | `--effort <level>` | `low`, `medium`, `high`, `max`, `auto` |
67
68
  | `--allowed-tools <tools>` | Comma-separated tool whitelist |
68
69
  | `--max-turns <n>` | Max agent loop turns |
@@ -223,7 +223,7 @@ Hosts deliberately do not forward your full shell environment to MCP subprocesse
223
223
  | `CLAWO_MCP_TOOLS` | Comma-separated allowlist of tool names; unlisted tools are not advertised |
224
224
  | `CLAWO_NO_EMBEDDED_SERVER` | Suppresses port 18796 binding. `clawo-mcp` sets this automatically |
225
225
 
226
- The engines themselves (`claude`, `codex`, `gemini`, `agent`, `opencode`) must also be installed and authenticated on the host machine — `clawo-mcp` spawns them as subprocesses, it does not bundle them.
226
+ The engines themselves (`claude`, `codex`, `gemini`, `agy`, `agent`, `opencode`) must also be installed and authenticated on the host machine — `clawo-mcp` spawns them as subprocesses, it does not bundle them.
227
227
 
228
228
  ---
229
229
 
@@ -14,6 +14,8 @@ SessionManager
14
14
  │ └── Wraps: codex app-server --listen stdio:// (long-running JSON-RPC; required for /goal)
15
15
  ├── engine: 'gemini' → PersistentGeminiSession
16
16
  │ └── Wraps: gemini -p --output-format stream-json (per-message spawning)
17
+ ├── engine: 'agy' → PersistentAgySession
18
+ │ └── Wraps: agy -p (Google Antigravity CLI, per-message spawning, plain-text output)
17
19
  ├── engine: 'cursor' → PersistentCursorSession
18
20
  │ └── Wraps: agent -p --force --trust --output-format stream-json (per-message spawning)
19
21
  ├── engine: 'opencode' → PersistentOpencodeSession
@@ -26,7 +28,7 @@ SessionManager
26
28
 
27
29
  ### Claude Code (`engine: 'claude'`)
28
30
 
29
- Default engine. Long-running subprocess with streaming JSON I/O. Tested with Claude Code CLI **2.1.199**.
31
+ Default engine. Long-running subprocess with streaming JSON I/O. Tested with Claude Code CLI **2.1.207**.
30
32
 
31
33
  - Persistent multi-turn conversations
32
34
  - Real-time streaming (text, tool_use, tool_result, system events)
@@ -52,7 +54,7 @@ await manager.startSession({
52
54
 
53
55
  ### OpenAI Codex (`engine: 'codex'`)
54
56
 
55
- Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested with `codex` CLI **0.142.4**.
57
+ Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested with `codex` CLI **0.144.1**.
56
58
 
57
59
  - Non-interactive execution via `codex exec --sandbox workspace-write --json` (replaces the deprecated `--full-auto` flag from earlier Codex versions)
58
60
  - Real per-turn `usage` from the `turn.completed` JSON event (input, output, cached, reasoning tokens)
@@ -61,6 +63,7 @@ Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested wi
61
63
  - `codexProfile` → `--profile <name>` (named config profile from `~/.codex/config.toml`)
62
64
  - Per-session continuity: the `thread_id` from the first turn's `thread.started` event is captured and reused via `codex exec resume <id>` for subsequent sends, so the model sees prior turns
63
65
  - One-shot execution per message (no persistent subprocess between sends)
66
+ - Captures the real Codex thread ID and persists it, so later sends and process-level session resume use `codex exec resume <thread_id>`
64
67
  - Working directory passed via `-C` flag
65
68
  - Default model: `gpt-5.5`
66
69
  - Requires `codex` CLI >= 0.119 (for `exec resume`): `npm install -g @openai/codex`
@@ -111,7 +114,7 @@ Wraps the `gemini` CLI with `--output-format stream-json`. Each `send()` spawns
111
114
  - One-shot execution per message (no persistent subprocess)
112
115
  - Working directory carries accumulated changes across sends
113
116
  - Real token counts from stream-json `result` events (not estimated)
114
- - Permission modes: `bypassPermissions` → `--yolo`, `default` → `--sandbox`
117
+ - Permission modes: `bypassPermissions` → `--yolo`, `default` → `--sandbox`; `sandboxMode: 'read-only'` → `--approval-mode plan` (takes precedence)
115
118
  - Always passes `--skip-trust` to bypass the "trusted folders" gate introduced
116
119
  in Gemini CLI 0.43 (otherwise headless runs in worktrees / arbitrary cwds
117
120
  abort before producing output)
@@ -126,14 +129,54 @@ await manager.startSession({
126
129
  });
127
130
  ```
128
131
 
132
+ ### Google Antigravity (`engine: 'agy'`)
133
+
134
+ Wraps Google's **Antigravity CLI** (`agy`) — the successor to Gemini CLI (consumer
135
+ Gemini CLI tiers stopped serving 2026-06-18). Each `send()` spawns a new process
136
+ in print mode. Verified against `agy` **1.1.1**.
137
+
138
+ - One-shot execution per message (no persistent subprocess)
139
+ - **Plain-text output** — agy has no structured/stream-json mode, so stdout is
140
+ forwarded as streaming text and **token counts are estimated** (~4 chars/token)
141
+ - **Real conversation continuity**: agy logs `Created conversation <uuid>` to its
142
+ log file; the engine passes a private `--log-file`, harvests the ID after the
143
+ first turn, and resumes with `--conversation <id>` on subsequent sends. Seed it
144
+ externally via `resumeSessionId` (bare UUID only); read it back from
145
+ `getStats().agyConversationId`
146
+ - Permission modes: `bypassPermissions` → `--dangerously-skip-permissions`,
147
+ `default` → `--sandbox` (terminal-restricted), and
148
+ `sandboxMode: 'read-only'` → `--mode plan` (takes precedence). Other modes
149
+ run agy's own approval flow, which blocks in headless print mode — use
150
+ `bypassPermissions` for autonomous write-enabled work
151
+ - agy enforces its own print timeout (default 5m); the engine derives
152
+ `--print-timeout` from the send timeout so the wrapper timer decides
153
+ - Unknown `--model` slugs do **not** error — agy silently falls back to its
154
+ default model. Registered slugs: `gemini-3.5-flash` (alias `agy-flash`),
155
+ `gemini-3.1-pro` (alias `agy-pro`); agy also proxies Claude and GPT-OSS
156
+ models (`agy models` lists them) which pass through unregistered. The
157
+ `agy/` prefix forces Antigravity routing for provider-like model strings
158
+ - Consumer auth is a one-time `agy` Google OAuth login (subscription quotas, no
159
+ per-token billing — registry pricing mirrors Gemini API rates as a value proxy)
160
+ - Requires `agy` installed: `curl -fsSL https://antigravity.google/cli/install.sh | bash`
161
+ - Binary: `agy` (set `AGY_BIN` env var to override)
162
+
163
+ ```typescript
164
+ await manager.startSession({
165
+ name: 'antigravity-task',
166
+ engine: 'agy',
167
+ model: 'gemini-3.5-flash',
168
+ cwd: '/project',
169
+ });
170
+ ```
171
+
129
172
  ### Cursor Agent (`engine: 'cursor'`)
130
173
 
131
- Wraps the Cursor Agent CLI (`agent`) with `--print --force --output-format stream-json`. Each `send()` spawns a new process.
174
+ Wraps the Cursor Agent CLI (`agent`) with `--print --output-format stream-json`. Write-enabled sessions use `--force`; `sandboxMode: 'read-only'` uses `--mode plan`. Each `send()` spawns a new process.
132
175
 
133
176
  - One-shot execution per message (no persistent subprocess)
134
177
  - Working directory via `--workspace` flag
135
178
  - Real token counts from stream-json `result` events (camelCase: `inputTokens`, `outputTokens`, `cacheReadTokens`)
136
- - `--force` enables auto-approval of all file changes
179
+ - `--force` enables auto-approval of file changes; `sandboxMode: 'read-only'` replaces it with `--mode plan`
137
180
  - `--trust` auto-trusts the workspace without prompting
138
181
  - Cursor uses its own model routing (e.g., `sonnet-4`, `gpt-5`, `auto`)
139
182
  - Requires Cursor Agent CLI: `curl https://cursor.com/install -fsSL | bash`
@@ -159,6 +202,7 @@ Wraps the [sst/opencode](https://github.com/sst/opencode) CLI with `run --format
159
202
  - Real token counts from `step_finish.part.tokens.{input,output,cache.read}`
160
203
  - The wrapper closes the subprocess's stdin immediately after spawn (opencode otherwise reads stdin and blocks on EOF, hanging the call)
161
204
  - Provider-agnostic: opencode's `--model` expects `provider/model` form (e.g. `anthropic/claude-sonnet-4`). The wrapper passes `--model` through only when the value contains a `/`; otherwise opencode's own default applies
205
+ - `sandboxMode: 'read-only'` selects OpenCode's built-in `plan` agent via `--agent plan`
162
206
  - Requires opencode installed: `brew install sst/tap/opencode` or `npm install -g opencode-ai`. Auth via `opencode auth login` **or** any provider env var (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GEMINI_API_KEY`, etc.) — opencode picks up either path
163
207
  - Binary: `opencode` (set `OPENCODE_BIN` env var to override)
164
208
 
@@ -384,9 +428,10 @@ await manager.startSession({
384
428
 
385
429
  ### Example: Google Antigravity CLI (`agy`)
386
430
 
387
- Google is sunsetting Gemini CLI (consumer tiers stop serving **2026-06-18**) in
388
- favour of the Go-based **Antigravity CLI** (`agy`). Until `agy` ships first-class
389
- support here, you can drive it today via a custom engine:
431
+ > **Note:** `agy` now has first-class support — use [`engine: 'agy'`](#google-antigravity-engine-agy)
432
+ > instead, which adds conversation resume and timeout coherence the recipe below
433
+ > lacks. This recipe remains as a reference for driving older agy builds or
434
+ > forks with a diverged flag surface:
390
435
 
391
436
  ```typescript
392
437
  await manager.startSession({
@@ -29,9 +29,9 @@ Key options:
29
29
 
30
30
  | Option | Description |
31
31
  |--------|-------------|
32
- | `engine` | `'claude'` (default), `'codex'`, or `'gemini'` — see [Multi-Engine](./multi-engine.md) |
32
+ | `engine` | `'claude'` (default), `'codex'`, `'codex-app'`, `'gemini'`, `'agy'`, `'cursor'`, `'opencode'`, or `'custom'` — see [Multi-Engine](./multi-engine.md) |
33
33
  | `model` | Model alias (`opus`, `sonnet`, `haiku`, `gemini-pro`) or full name |
34
- | `permissionMode` | `acceptEdits`, `bypassPermissions`, `plan`, `auto`, `default` |
34
+ | `permissionMode` | `acceptEdits`, `bypassPermissions`, `plan`, `auto`, `manual`, `dontAsk` (`default` = legacy alias for `manual`) |
35
35
  | `effort` | `low`, `medium`, `high`, `max`, `auto` |
36
36
  | `bare` | Skip hooks, LSP, auto-memory, CLAUDE.md |
37
37
  | `worktree` | Run in isolated git worktree |
@@ -12,9 +12,10 @@ Start a persistent coding session with full CLI flag support.
12
12
  |-----------|------|-------------|
13
13
  | `name` | string | Session name (auto-generated if omitted) |
14
14
  | `cwd` | string | Working directory |
15
- | `engine` | `'claude'` \| `'codex'` \| `'gemini'` \| `'cursor'` \| `'opencode'` \| `'custom'` | Engine to use (default: `claude`). `opencode` wraps sst/opencode (pass model as `provider/model`). Use `custom` with `customEngine` for any CLI. |
15
+ | `engine` | `'claude'` \| `'codex'` \| `'codex-app'` \| `'gemini'` \| `'agy'` \| `'cursor'` \| `'opencode'` \| `'custom'` | Engine to use (default: `claude`). `agy` wraps Google Antigravity CLI. `opencode` wraps sst/opencode (pass model as `provider/model`). Use `custom` with `customEngine` for any CLI. |
16
16
  | `model` | string | Model alias or full name |
17
- | `permissionMode` | string | `acceptEdits`, `bypassPermissions`, `plan`, `auto`, `default` |
17
+ | `permissionMode` | string | `acceptEdits`, `bypassPermissions`, `plan`, `auto`, `manual`, `dontAsk` (`default` = legacy alias for `manual`) |
18
+ | `sandboxMode` | `'read-only'` \| `'workspace-write'` \| `'danger-full-access'` | Sandbox policy. Codex supports all values. `read-only` is enforced on every other built-in engine too: Claude → plan mode; Gemini → `--approval-mode plan` + an admin policy denying `exit_plan_mode`; OpenCode → a generated `clawo-readonly` agent denying `edit`/`bash`; Antigravity/Cursor → plan mode. A `custom` engine must map it via `permissionModes`, or the session refuses to start. Persisted across session resume. |
18
19
  | `effort` | string | `low`, `medium`, `high`, `max`, `auto` |
19
20
  | `allowedTools` | string[] | Tools to auto-approve |
20
21
  | `disallowedTools` | string[] | Tools to deny |
@@ -537,57 +538,75 @@ Three-agent autonomous iteration loop (Planner / Coder / Reviewer) over a git wo
537
538
 
538
539
  ### `autoloop_start`
539
540
 
540
- Start an autoloop run. Planner is created persistent; Coder + Reviewer are spawned by the Planner once `plan.md` is ready.
541
+ Start a chat-mode autoloop. Planner starts immediately; Coder + Reviewer start only after the Planner receives plan approval and emits `spawn_subagents`.
541
542
 
542
543
  | Parameter | Type | Required | Description |
543
544
  |-----------|------|----------|-------------|
544
- | `cwd` | string | yes | Workspace (must be a git repo) |
545
- | `goal` | string | yes | High-level user goal in natural language |
546
- | `model` | string | | Planner model (default Opus) |
547
- | `coderModel` | string | | Coder subagent model |
548
- | `reviewerModel` | string | | Reviewer subagent model |
549
- | `maxIters` | number | | Cap on Coder/Reviewer rounds (default 50) |
550
- | `pushChannels` | string[] | | Notification channels (`wechat`, `whatsapp`, `email`) |
545
+ | `run_id` | string | yes | Stable run identifier |
546
+ | `workspace` | string | yes | Git workspace path |
547
+ | `planner_engine` | EngineType | | Planner engine (default `claude`) |
548
+ | `planner_model` | string | | Planner model (Claude default `opus`; other engines use their own default when omitted) |
549
+ | `planner_custom_engine` | object | | Trusted `CustomEngineConfig` when Planner engine is `custom`. **Local callers only** — see below |
550
+ | `coder_engine` | EngineType | | Default Coder engine (default `claude`) |
551
+ | `coder_model` | string | | Default Coder model (Claude default `sonnet`) |
552
+ | `coder_custom_engine` | object | | Trusted config when Coder may use `custom`. **Local callers only** |
553
+ | `reviewer_engine` | EngineType | | Default Reviewer engine (default `claude`) |
554
+ | `reviewer_model` | string | | Default Reviewer model (Claude default `sonnet`) |
555
+ | `reviewer_custom_engine` | object | | Trusted config when Reviewer may use `custom`. **Local callers only** |
556
+ | `send_timeout_ms` | number | | Per-message timeout (default 600000) |
557
+
558
+ > **Custom engines are local-only.** A `CustomEngineConfig` names an executable to
559
+ > spawn plus its argv and env, so it may only be supplied by a caller that already
560
+ > runs on the host (this MCP tool, or the `SessionManager` API). The HTTP API
561
+ > (`POST /autoloop/new`, `POST /autoloop/<id>/resume`) rejects a `*_custom_engine`
562
+ > body field with a 400 — the embedded server is often reverse-tunnelled and its
563
+ > token is a monitoring credential, not permission to choose what binary runs.
564
+ > Built-in engines are fully selectable over HTTP.
565
+
566
+ Custom configs are not persisted or accepted from Planner output. See [`multi-engine.md`](./multi-engine.md) for their shape.
551
567
 
552
568
  ### `autoloop_chat`
553
569
 
554
- Send a message into the Planner conversation (e.g. answer a clarifying question, refine the plan, kick off the subloop).
570
+ Send a message into the Planner conversation.
555
571
 
556
572
  | Parameter | Type | Required |
557
573
  |-----------|------|----------|
558
- | `id` | string | yes |
559
- | `message` | string | yes |
574
+ | `run_id` | string | yes |
575
+ | `text` | string | yes |
560
576
 
561
577
  ### `autoloop_status`
562
578
 
563
- Get current state, phase, recent inbox messages, and ledger summary.
579
+ Get the current state and push log.
564
580
 
565
581
  | Parameter | Type | Required |
566
582
  |-----------|------|----------|
567
- | `id` | string | yes |
583
+ | `run_id` | string | yes |
568
584
 
569
585
  ### `autoloop_list`
570
586
 
571
- List active and recent autoloop runs (in-memory + on-disk registry, deduped by run_id).
587
+ List active Autoloop runs in this manager process.
572
588
 
573
589
  (no params)
574
590
 
575
591
  ### `autoloop_reset_agent`
576
592
 
577
- Reset one of the subagent sessions (Coder or Reviewer) without losing Planner state — useful when a subagent loops on a stale belief.
593
+ Reset one role session while retaining the role's current engine/model selection.
578
594
 
579
595
  | Parameter | Type | Required | Description |
580
596
  |-----------|------|----------|-------------|
581
- | `id` | string | yes | Run id |
582
- | `agent` | `'coder'` \| `'reviewer'` | yes | Which subagent to reset |
597
+ | `run_id` | string | yes | Run id |
598
+ | `agent` | `'planner'` \| `'coder'` \| `'reviewer'` | yes | Role to reset; Planner requires `force: true` |
599
+ | `force` | boolean | | Allow Planner reset |
600
+ | `eager_restart` | boolean | | Start the replacement session immediately |
583
601
 
584
602
  ### `autoloop_stop`
585
603
 
586
- Terminate the run. All sessions are stopped and ledger state is finalised.
604
+ Terminate the run and stop all role sessions.
587
605
 
588
606
  | Parameter | Type | Required |
589
607
  |-----------|------|----------|
590
- | `id` | string | yes |
608
+ | `run_id` | string | yes |
609
+ | `reason` | string | |
591
610
 
592
611
  ---
593
612