@enderfga/claw-orchestrator 5.0.0 → 5.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -42,7 +42,8 @@ The server exposes an OpenAI-compatible chat completions endpoint, enabling any
42
42
  **Model routing:** The `model` field auto-routes to the correct engine:
43
43
  - `claude-*`, `opus`, `sonnet`, `haiku` → Claude engine
44
44
  - `gpt-*` → Codex engine
45
- - `composer-*` → Cursor engine
45
+ - `grok-*` → Grok engine
46
+ - `composer-*` → Cursor engine (legacy)
46
47
  - `gemini-3.5-flash`, `gemini-3.1-pro`, `agy-*`, `agy/*` → Antigravity (`agy`) engine
47
48
  - other `gemini-*` → the legacy `gemini` engine (Gemini CLI is sunset; prefer `agy`)
48
49
 
@@ -61,7 +62,7 @@ clawo session-start [name] [options]
61
62
  | Flag | Description |
62
63
  |------|-------------|
63
64
  | `-d, --cwd <dir>` | Working directory |
64
- | `-e, --engine <engine>` | Engine: `claude` (default), `codex`, `codex-app`, `agy`, `cursor`, `opencode`, or `custom` |
65
+ | `-e, --engine <engine>` | Engine: `claude` (default), `codex`, `codex-app`, `agy`, `grok`, `opencode`, or `custom` |
65
66
  | `-m, --model <model>` | Model name or alias |
66
67
  | `--permission-mode <mode>` | `acceptEdits`, `plan`, `auto`, `bypassPermissions`, `manual`, `dontAsk` |
67
68
  | `--effort <level>` | `low`, `medium`, `high`, `max`, `auto` |
@@ -23,7 +23,7 @@ openclaw plugins install @enderfga/claw-orchestrator --dangerously-force-unsafe-
23
23
  openclaw gateway restart
24
24
  ```
25
25
 
26
- > **Why `--dangerously-force-unsafe-install`?** Claw Orchestrator spawns Claude Code / Codex / Antigravity / Cursor Agent / OpenCode CLI subprocesses via `child_process`, which OpenClaw's security scanner flags by design. The flag is required — there is no way to drive coding CLIs without process spawning.
26
+ > **Why `--dangerously-force-unsafe-install`?** Claw Orchestrator spawns Claude Code / Codex / Antigravity / Grok Build / OpenCode CLI subprocesses via `child_process`, which OpenClaw's security scanner flags by design. The flag is required — there is no way to drive coding CLIs without process spawning.
27
27
 
28
28
  Agents automatically get access to all session, council, and management tools.
29
29
 
@@ -14,7 +14,9 @@ SessionManager
14
14
  │ └── Wraps: codex app-server --listen stdio:// (long-running JSON-RPC; required for /goal)
15
15
  ├── engine: 'agy' → PersistentAgySession
16
16
  │ └── Wraps: agy -p (Google Antigravity CLI, per-message spawning, stream-json output)
17
- ├── engine: 'cursor' → PersistentCursorSession
17
+ ├── engine: 'grok' → PersistentGrokSession
18
+ │ └── Wraps: grok -p --output-format json (xAI Grok Build, per-message spawning)
19
+ ├── engine: 'cursor' → PersistentCursorSession (legacy)
18
20
  │ └── Wraps: agent -p --force --trust --output-format stream-json (per-message spawning)
19
21
  ├── engine: 'opencode' → PersistentOpencodeSession
20
22
  │ └── Wraps: opencode run --format json (per-message spawning)
@@ -173,9 +175,59 @@ await manager.startSession({
173
175
  > multi-model **proxy** still talks to the Gemini **API**; that is a different
174
176
  > subsystem and is unaffected.)
175
177
 
176
- ### Cursor Agent (`engine: 'cursor'`)
178
+ ### Grok Build (`engine: 'grok'`)
179
+
180
+ Wraps xAI's **Grok Build** CLI. Each `send()` spawns `grok -p <msg> --output-format json`, which
181
+ prints a single JSON object and exits. Verified against `grok` **1.0.5**.
182
+
183
+ - **Cost comes from the engine, not from our price table.** The result object carries
184
+ `total_cost_usd`, and the wrapper writes it straight into the session's spend. Every other engine
185
+ here multiplies tokens by a rate in `models.ts` — the metadata most prone to going stale — so on
186
+ this engine the run ledger and the `maxBudgetUsd` gate both read what xAI actually charged.
187
+ `grok-4.6` is still registered, for its context window and an indicative breakdown; its two price
188
+ tiers ($2/$0.50/$6 under a 200K prompt, $4/$1/$12 at or above, charged across the whole request)
189
+ therefore never have to be modelled here.
190
+ - **Real conversation continuity**: the `sessionId` from turn 1 is replayed as `--resume <id>`.
191
+ `--continue` is deliberately not used — it means "the most recent session for this cwd", which
192
+ collides between concurrent sessions. Confirmed with a two-turn recall test, not inferred.
193
+ - Real token counts from `usage` (`input_tokens`, `output_tokens`, `cache_read_input_tokens`).
194
+ These are **per-turn**, not cumulative over the thread — checked by resuming and reading turn 2,
195
+ because the same-looking field on codex is a running total.
196
+ - Permission modes pass straight through: grok's `--permission-mode` takes the same vocabulary we
197
+ use. The one exception is our `manual`, which grok spells `default`.
198
+ - Reasoning effort maps to `--effort`; grok accepts `low|medium|high`, so `max` and `xhigh` clamp.
199
+ - **`sandboxMode: 'read-only'` is refused, not approximated.** grok has `--permission-mode plan` and
200
+ `--deny` rules, but plan mode alone is model-cooperative — the shape that let an adversarial
201
+ prompt write through Cursor's plan mode — and the deny rules have not been through the
202
+ write × shell × subagent × resumed-turn matrix this project requires before claiming a boundary.
203
+ A read-only grok session throws rather than running writable under a read-only label.
204
+ - Binary: `grok` (set `GROK_BIN` to override). Not `agent`: xAI's installer claims that name too,
205
+ and so did Cursor's.
206
+ - Requires Grok Build: see `x.ai/cli`.
177
207
 
178
- Wraps the Cursor Agent CLI (`agent`) with `--print --output-format stream-json`. Write-enabled sessions use `--force`. Each `send()` spawns a new process.
208
+ ```typescript
209
+ await manager.startSession({
210
+ name: 'grok-task',
211
+ engine: 'grok',
212
+ model: 'grok-4.6',
213
+ cwd: '/project',
214
+ });
215
+ ```
216
+
217
+ ### Cursor Agent (`engine: 'cursor'`) — legacy
218
+
219
+ > **Legacy: `engine: 'cursor'`.** Superseded in this lineup by Grok Build (`engine: 'grok'`).
220
+ > The `cursor` engine still exists and still works — existing callers are not broken — but it is
221
+ > no longer a documented option, is not version-tracked, and gets no new work.
222
+ >
223
+ > Note what this is and is not: Cursor itself is **not** discontinued. Anysphere was acquired by
224
+ > SpaceX (closed 2026-08-15) and folded into the SpaceXAI team, and the CLI has shipped since. Two
225
+ > practical things pushed it out of the tracked set. Cursor never reports which model actually ran
226
+ > — its `system` init event says `"model": "Auto"` — so a router that spans Claude, GPT and Grok
227
+ > leaves every cost row attributed to a hardcoded proxy rate. And xAI's Grok installer now claims
228
+ > the bare `agent` name, so the binary that name resolves to depends on install order.
229
+
230
+ Wraps the Cursor Agent CLI with `--print --output-format stream-json`. Write-enabled sessions use `--force`. Each `send()` spawns a new process.
179
231
 
180
232
  - Conversation continuity: the chat id from the first turn's `system` event is captured and passed back as `--resume <chatId>` on later sends, so the model sees prior turns. `--continue` is deliberately not used: it resumes "the latest chat", which collides between concurrent sessions.
181
233
  - One-shot execution per message (no persistent subprocess)
@@ -185,7 +237,9 @@ Wraps the Cursor Agent CLI (`agent`) with `--print --output-format stream-json`.
185
237
  - `--trust` auto-trusts the workspace without prompting
186
238
  - Cursor uses its own model routing (e.g., `sonnet-4`, `gpt-5`, `auto`)
187
239
  - Requires Cursor Agent CLI: `curl https://cursor.com/install -fsSL | bash`
188
- - Binary: `agent` (set `CURSOR_BIN` env var to override)
240
+ - Binary: `cursor-agent` (set `CURSOR_BIN` env var to override). The generic `agent` name is
241
+ deliberately not used: xAI's Grok installer symlinks `agent` to its own binary, which rejects
242
+ `--force`/`--trust`/`--workspace` and fails the turn with "unexpected argument"
189
243
 
190
244
  ```typescript
191
245
  await manager.startSession({
@@ -34,7 +34,7 @@ every engine passes through.
34
34
  |---|---|
35
35
  | `ts` | ISO timestamp of turn completion |
36
36
  | `session` | SessionManager session name |
37
- | `engine` | `claude` / `codex` / `codex-app` / `cursor` / `opencode` / `agy` / `custom` |
37
+ | `engine` | `claude` / `codex` / `codex-app` / `grok` / `opencode` / `agy` / `custom` |
38
38
  | `model` | Configured model, or the engine's own reported model when none was set |
39
39
  | `cwd` | Working directory the turn ran in |
40
40
  | `turn` | 1-based turn index within the session |
@@ -116,7 +116,8 @@ flagged `tokensEstimated: true`; the CLI marks those costs with a trailing `~`.
116
116
  | `claude` | Engine-reported |
117
117
  | `codex` | Engine-reported |
118
118
  | `codex-app` | Engine-reported |
119
- | `cursor` | Engine-reported when the stream carries `usage`, else estimated |
119
+ | `grok` | Engine-reported — and so is the **cost**: this engine reports `total_cost_usd`, which the wrapper passes through instead of pricing tokens from the registry, so registry drift cannot affect a grok row |
120
+ | `cursor` (legacy) | Engine-reported when the stream carries `usage`, else estimated |
120
121
  | `opencode` | Engine-reported when the run JSON carries `tokens`, else estimated |
121
122
  | `agy` | Engine-reported when the result event carries usage, else estimated |
122
123
  | `custom` | Depends on the CLI; estimated when it emits no usage |
@@ -2,7 +2,7 @@
2
2
 
3
3
  > **Cost warning**: This bridge routes requests through the Claude Code CLI, which uses your Claude Max subscription's **extra usage** quota. When OpenClaw's agent loop sends its system prompt (with distinctive tool definitions and agent instructions), Anthropic's backend recognizes this as programmatic/agent traffic and bills it against extra usage — **not** the included allowance. This is by design: the bridge does NOT bypass Anthropic's billing or subscription enforcement. Using it as OpenClaw's primary model backend means every agent turn consumes extra usage credits at standard API rates ($15/M input, $75/M output for Opus). Monitor your usage at [claude.ai/settings/usage](https://claude.ai/settings/usage).
4
4
 
5
- The embedded server exposes a drop-in OpenAI-compatible endpoint so any client that speaks `/v1/chat/completions` can talk to a persistent Claude Code (or Codex / Antigravity / Cursor) session. The bridge is designed to serve **two kinds of clients as first-class citizens**:
5
+ The embedded server exposes a drop-in OpenAI-compatible endpoint so any client that speaks `/v1/chat/completions` can talk to a persistent Claude Code (or Codex / Antigravity / Grok) session. The bridge is designed to serve **two kinds of clients as first-class citizens**:
6
6
 
7
7
  1. **Upstream agents** that maintain their own conversation state and forward only the latest user turn — OpenClaw's main agent loop, cron jobs, subagents, programmatic clients.
8
8
  2. **OpenAI-compatible webchat / labeling tools** that re-send the full transcript on every turn — ChatGPT-Next-Web, Open WebUI, LobeChat, data-labeling pipelines.
@@ -81,7 +81,7 @@ mechanism is used depends on whether the engine keeps the conversation itself.
81
81
  | Engine | Turn 1 | Later turns |
82
82
  |---|---|---|
83
83
  | `claude` | Schemas go into the session system prompt (`--system-prompt`) | Nothing injected — the system prompt persists |
84
- | `codex`, `codex-app`, `agy`, `opencode`, `cursor` | Full schema block prepended to the message | A short reminder of the calling convention, no schemas — but only once the conversation id has been captured; until then the full block is sent again |
84
+ | `codex`, `codex-app`, `agy`, `opencode`, `grok` | Full schema block prepended to the message | A short reminder of the calling convention, no schemas — but only once the conversation id has been captured; until then the full block is sent again |
85
85
  | `gemini`, one-shot `custom` | Full schema block prepended to the message | Full schema block again — these have no resume surface, so nothing persists between sends |
86
86
 
87
87
  The middle row is the one worth understanding. Those engines resume a conversation by id, so
@@ -29,7 +29,7 @@ Key options:
29
29
 
30
30
  | Option | Description |
31
31
  |--------|-------------|
32
- | `engine` | `'claude'` (default), `'codex'`, `'codex-app'`, `'agy'`, `'cursor'`, `'opencode'`, or `'custom'` — see [Multi-Engine](./multi-engine.md) |
32
+ | `engine` | `'claude'` (default), `'codex'`, `'codex-app'`, `'agy'`, `'grok'`, `'opencode'`, or `'custom'` — see [Multi-Engine](./multi-engine.md) |
33
33
  | `model` | Model alias (`fable`, `opus`, `sonnet`, `haiku`, `agy-pro`) or full name |
34
34
  | `permissionMode` | `acceptEdits`, `bypassPermissions`, `plan`, `auto`, `manual`, `dontAsk` (`default` = legacy alias for `manual`) |
35
35
  | `effort` | `low`, `medium`, `high`, `max`, `auto` |
@@ -135,7 +135,7 @@ prompt over the window the engine actually enforces — so it rises and falls wi
135
135
  the conversation. `stats.tokensIn` is the different question of how many input
136
136
  tokens the session has been billed for in total, which only ever grows.
137
137
  `compactSession()` is a no-op on engines whose CLI has no compaction command
138
- (`codex`, `agy`, `cursor`, `opencode`); those sessions log a warning the first
138
+ (`codex`, `agy`, `grok`, `opencode`); those sessions log a warning the first
139
139
  time it is called.
140
140
 
141
141
  `stats.turns` and `stats.turnsSucceeded` are the same kind of distinction.
@@ -12,10 +12,10 @@ Start a persistent coding session with full CLI flag support.
12
12
  |-----------|------|-------------|
13
13
  | `name` | string | Session name (auto-generated if omitted) |
14
14
  | `cwd` | string | Working directory |
15
- | `engine` | `'claude'` \| `'codex'` \| `'codex-app'` \| `'agy'` \| `'cursor'` \| `'opencode'` \| `'custom'` | Engine to use (default: `claude`). `agy` wraps Google Antigravity CLI. `opencode` wraps sst/opencode (pass model as `provider/model`). Use `custom` with `customEngine` for any CLI. (`'gemini'` is still accepted for existing callers, but Gemini CLI is sunset — use `agy` for Google.) |
15
+ | `engine` | `'claude'` \| `'codex'` \| `'codex-app'` \| `'agy'` \| `'grok'` \| `'opencode'` \| `'custom'` | Engine to use (default: `claude`). `agy` wraps Google Antigravity CLI. `grok` wraps xAI Grok Build and reports its own per-turn cost. `opencode` wraps sst/opencode (pass model as `provider/model`). Use `custom` with `customEngine` for any CLI. (`'gemini'` and `'cursor'` are still accepted for existing callers but are legacy — not version-tracked; use `agy` for Google and `grok` in place of Cursor.) |
16
16
  | `model` | string | Model alias or full name |
17
17
  | `permissionMode` | string | `acceptEdits`, `bypassPermissions`, `plan`, `auto`, `manual`, `dontAsk` (`default` = legacy alias for `manual`) |
18
- | `sandboxMode` | `'read-only'` \| `'workspace-write'` \| `'danger-full-access'` | Sandbox policy. Codex supports all values. `read-only` is enforced on every other built-in engine too: Claude → plan mode; Antigravity / Cursor → their plan modes; OpenCode → a generated `clawo-readonly` agent denying `edit`/`bash`. A `custom` engine must map it via `permissionModes`, or the session refuses to start. Persisted across session resume. |
18
+ | `sandboxMode` | `'read-only'` \| `'workspace-write'` \| `'danger-full-access'` | Sandbox policy. Codex supports all values. `read-only` is enforced on every other built-in engine too: Claude → plan mode; Antigravity → its plan mode; OpenCode → a generated `clawo-readonly` agent denying `edit`/`bash`. **`grok` refuses a read-only session** rather than approximate one — its enforcement has not been adversarially verified. A `custom` engine must map it via `permissionModes`, or the session refuses to start. Persisted across session resume. |
19
19
  | `effort` | string | `low`, `medium`, `high`, `max`, `auto` |
20
20
  | `allowedTools` | string[] | Tools to auto-approve |
21
21
  | `disallowedTools` | string[] | Tools to deny |