@enderfga/claw-orchestrator 3.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +218 -0
- package/assets/banner.jpg +0 -0
- package/configs/council-reviewer-prompt.md +82 -0
- package/configs/council-system-prompt.md +141 -0
- package/dist/bin/cli.d.ts +13 -0
- package/dist/bin/cli.js +460 -0
- package/dist/bin/cli.js.map +1 -0
- package/dist/src/base-oneshot-session.d.ts +87 -0
- package/dist/src/base-oneshot-session.js +228 -0
- package/dist/src/base-oneshot-session.js.map +1 -0
- package/dist/src/circuit-breaker.d.ts +21 -0
- package/dist/src/circuit-breaker.js +49 -0
- package/dist/src/circuit-breaker.js.map +1 -0
- package/dist/src/consensus.d.ts +20 -0
- package/dist/src/consensus.js +52 -0
- package/dist/src/consensus.js.map +1 -0
- package/dist/src/constants.d.ts +129 -0
- package/dist/src/constants.js +138 -0
- package/dist/src/constants.js.map +1 -0
- package/dist/src/council.d.ts +67 -0
- package/dist/src/council.js +914 -0
- package/dist/src/council.js.map +1 -0
- package/dist/src/embedded-server.d.ts +25 -0
- package/dist/src/embedded-server.js +360 -0
- package/dist/src/embedded-server.js.map +1 -0
- package/dist/src/inbox-manager.d.ts +38 -0
- package/dist/src/inbox-manager.js +111 -0
- package/dist/src/inbox-manager.js.map +1 -0
- package/dist/src/index.d.ts +63 -0
- package/dist/src/index.js +973 -0
- package/dist/src/index.js.map +1 -0
- package/dist/src/logger.d.ts +16 -0
- package/dist/src/logger.js +44 -0
- package/dist/src/logger.js.map +1 -0
- package/dist/src/models.d.ts +69 -0
- package/dist/src/models.js +299 -0
- package/dist/src/models.js.map +1 -0
- package/dist/src/openai-compat.d.ts +224 -0
- package/dist/src/openai-compat.js +756 -0
- package/dist/src/openai-compat.js.map +1 -0
- package/dist/src/persistent-codex-app-session.d.ts +108 -0
- package/dist/src/persistent-codex-app-session.js +465 -0
- package/dist/src/persistent-codex-app-session.js.map +1 -0
- package/dist/src/persistent-codex-session.d.ts +37 -0
- package/dist/src/persistent-codex-session.js +208 -0
- package/dist/src/persistent-codex-session.js.map +1 -0
- package/dist/src/persistent-cursor-session.d.ts +21 -0
- package/dist/src/persistent-cursor-session.js +241 -0
- package/dist/src/persistent-cursor-session.js.map +1 -0
- package/dist/src/persistent-custom-session.d.ts +78 -0
- package/dist/src/persistent-custom-session.js +938 -0
- package/dist/src/persistent-custom-session.js.map +1 -0
- package/dist/src/persistent-gemini-session.d.ts +21 -0
- package/dist/src/persistent-gemini-session.js +216 -0
- package/dist/src/persistent-gemini-session.js.map +1 -0
- package/dist/src/persistent-session.d.ts +80 -0
- package/dist/src/persistent-session.js +745 -0
- package/dist/src/persistent-session.js.map +1 -0
- package/dist/src/proxy/anthropic-adapter.d.ts +136 -0
- package/dist/src/proxy/anthropic-adapter.js +392 -0
- package/dist/src/proxy/anthropic-adapter.js.map +1 -0
- package/dist/src/proxy/handler.d.ts +39 -0
- package/dist/src/proxy/handler.js +365 -0
- package/dist/src/proxy/handler.js.map +1 -0
- package/dist/src/proxy/schema-cleaner.d.ts +11 -0
- package/dist/src/proxy/schema-cleaner.js +34 -0
- package/dist/src/proxy/schema-cleaner.js.map +1 -0
- package/dist/src/proxy/thought-cache.d.ts +19 -0
- package/dist/src/proxy/thought-cache.js +53 -0
- package/dist/src/proxy/thought-cache.js.map +1 -0
- package/dist/src/session-manager.d.ts +317 -0
- package/dist/src/session-manager.js +1528 -0
- package/dist/src/session-manager.js.map +1 -0
- package/dist/src/types.d.ts +513 -0
- package/dist/src/types.js +8 -0
- package/dist/src/types.js.map +1 -0
- package/dist/src/validation.d.ts +31 -0
- package/dist/src/validation.js +104 -0
- package/dist/src/validation.js.map +1 -0
- package/openclaw.plugin.json +122 -0
- package/package.json +84 -0
- package/skills/SKILL.md +184 -0
- package/skills/references/claude-cli-tracking.md +25 -0
- package/skills/references/cli.md +187 -0
- package/skills/references/council.md +210 -0
- package/skills/references/getting-started.md +133 -0
- package/skills/references/inbox.md +81 -0
- package/skills/references/multi-engine.md +382 -0
- package/skills/references/openai-compat.md +203 -0
- package/skills/references/sessions.md +191 -0
- package/skills/references/tools.md +418 -0
- package/skills/references/ultra.md +126 -0
|
@@ -0,0 +1,382 @@
|
|
|
1
|
+
# Multi-Engine
|
|
2
|
+
|
|
3
|
+
Claw Orchestrator supports multiple coding CLI engines behind a unified `ISession` interface. Each engine manages its own subprocess, event stream, and cost tracking independently.
|
|
4
|
+
|
|
5
|
+
## Architecture
|
|
6
|
+
|
|
7
|
+
```
|
|
8
|
+
SessionManager
|
|
9
|
+
├── engine: 'claude' → PersistentClaudeSession
|
|
10
|
+
│ └── Wraps: claude CLI (stream-json protocol, persistent subprocess)
|
|
11
|
+
├── engine: 'codex' → PersistentCodexSession
|
|
12
|
+
│ └── Wraps: codex exec --sandbox workspace-write --json (per-message spawning)
|
|
13
|
+
├── engine: 'codex-app' → PersistentCodexAppServerSession
|
|
14
|
+
│ └── Wraps: codex app-server --listen stdio:// (long-running JSON-RPC; required for /goal)
|
|
15
|
+
├── engine: 'gemini' → PersistentGeminiSession
|
|
16
|
+
│ └── Wraps: gemini -p --output-format stream-json (per-message spawning)
|
|
17
|
+
├── engine: 'cursor' → PersistentCursorSession
|
|
18
|
+
│ └── Wraps: agent -p --force --output-format stream-json (per-message spawning)
|
|
19
|
+
└── engine: 'custom' → PersistentCustomSession
|
|
20
|
+
└── Wraps: any CLI via user-provided CustomEngineConfig
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
## Supported Engines
|
|
24
|
+
|
|
25
|
+
### Claude Code (`engine: 'claude'`)
|
|
26
|
+
|
|
27
|
+
Default engine. Long-running subprocess with streaming JSON I/O. Tested with Claude Code CLI **2.1.126**.
|
|
28
|
+
|
|
29
|
+
- Persistent multi-turn conversations
|
|
30
|
+
- Real-time streaming (text, tool_use, tool_result, system events)
|
|
31
|
+
- Session resume via `--resume`
|
|
32
|
+
- Full cost tracking from API usage data
|
|
33
|
+
- Hook lifecycle events (`includeHookEvents`), permission delegation (`permissionPromptTool`), prompt cache optimization (`bare` + `excludeDynamicSystemPromptSections` + `enablePromptCaching1H`), debug control, `--from-pr` resume, and MCP channel subscriptions
|
|
34
|
+
- Fork subagent (`forkSubagent`), tool search (`enableToolSearch`), OpenTelemetry logging toggles (`otelLogUserPrompts`, `otelLogRawApiBodies`), `xhigh` effort tier (Opus 4.7), and `stats.pluginErrors` capture — see [CLI 2.1.121 options in SKILL.md](../SKILL.md) and [tools.md](./tools.md)
|
|
35
|
+
|
|
36
|
+
> **Behavior changes from upstream Claude CLI 2.1.121** (worth knowing if you set permission rules):
|
|
37
|
+
> - `--agent` / `--print` now enforce agent frontmatter `permissionMode`, `tools`, `disallowedTools` (was advisory). Affects `council` agent personas.
|
|
38
|
+
> - `Bash(find:*)` permission rule no longer auto-approves `find -exec` or `find -delete`. Add explicit rules if you depend on these.
|
|
39
|
+
> - `--dangerously-skip-permissions` also skips prompts for `.claude/skills/` directory. Treat with care.
|
|
40
|
+
> - Distributed tracing context (`TRACEPARENT` / `TRACESTATE`) is automatically forwarded to the child process — set them in the parent before starting the session.
|
|
41
|
+
|
|
42
|
+
```typescript
|
|
43
|
+
await manager.startSession({
|
|
44
|
+
name: 'claude-task',
|
|
45
|
+
engine: 'claude', // default, can omit
|
|
46
|
+
model: 'opus',
|
|
47
|
+
cwd: '/project',
|
|
48
|
+
});
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
### OpenAI Codex (`engine: 'codex'`)
|
|
52
|
+
|
|
53
|
+
Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested with `codex` CLI **0.128.0**.
|
|
54
|
+
|
|
55
|
+
- Non-interactive execution via `codex exec --sandbox workspace-write --json` (replaces the deprecated `--full-auto` flag from earlier Codex versions)
|
|
56
|
+
- Real per-turn `usage` from the `turn.completed` JSON event (input, output, cached, reasoning tokens)
|
|
57
|
+
- Per-session continuity: the `thread_id` from the first turn's `thread.started` event is captured and reused via `codex exec resume <id>` for subsequent sends, so the model sees prior turns
|
|
58
|
+
- One-shot execution per message (no persistent subprocess between sends)
|
|
59
|
+
- Working directory passed via `-C` flag
|
|
60
|
+
- Default model: `gpt-5.5`
|
|
61
|
+
- Requires `codex` CLI >= 0.119 (for `exec resume`): `npm install -g @openai/codex`
|
|
62
|
+
- **Does not support `/goal`** — for that, use `engine: 'codex-app'` below
|
|
63
|
+
|
|
64
|
+
```typescript
|
|
65
|
+
await manager.startSession({
|
|
66
|
+
name: 'codex-task',
|
|
67
|
+
engine: 'codex',
|
|
68
|
+
model: 'gpt-5.5',
|
|
69
|
+
cwd: '/project',
|
|
70
|
+
sandboxMode: 'workspace-write', // optional, this is the default
|
|
71
|
+
});
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
### OpenAI Codex App-Server (`engine: 'codex-app'`)
|
|
75
|
+
|
|
76
|
+
Wraps `codex app-server --listen stdio:// --enable goals` as a long-running JSON-RPC subprocess. **Required for `/goal` long-horizon objective support** — Codex's exec subcommand has no slash-command surface.
|
|
77
|
+
|
|
78
|
+
- Long-running subprocess; one `codex app-server` per session
|
|
79
|
+
- JSON-RPC 2.0 over stdio with v2 protocol method names (`initialize`, `thread/start`, `turn/start`, ...)
|
|
80
|
+
- Real-time streaming via `item/agentMessage/delta` notifications
|
|
81
|
+
- Cumulative token tracking from `thread/tokenUsage/updated` notifications
|
|
82
|
+
- Goal lifecycle observation via `thread/goal/updated` and `thread/goal/cleared` notifications
|
|
83
|
+
- Goal control via the `codex_goal_*` tools (which internally send the `/goal` slash command as user text — see [tools.md](./tools.md#codex-7))
|
|
84
|
+
|
|
85
|
+
> **Feature-flag risk.** The `goals` feature is marked "under development" in Codex 0.128.0 and has known bugs (e.g. issue #20591). The session class always passes `--enable goals` so it works the moment upstream stabilizes the feature, but during the transition period some goal commands may fail or be silently dropped on the server side. The wrapper layer is unaffected.
|
|
86
|
+
|
|
87
|
+
```typescript
|
|
88
|
+
await manager.startSession({
|
|
89
|
+
name: 'codex-goal-task',
|
|
90
|
+
engine: 'codex-app',
|
|
91
|
+
model: 'gpt-5.5',
|
|
92
|
+
cwd: '/project',
|
|
93
|
+
});
|
|
94
|
+
// Then either:
|
|
95
|
+
// await manager.codexGoalCommand('codex-goal-task', 'build a tic-tac-toe app');
|
|
96
|
+
// or via the codex_goal_set tool:
|
|
97
|
+
// await tool('codex_goal_set', { name: 'codex-goal-task', objective: 'build a tic-tac-toe app' });
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
### Google Gemini (`engine: 'gemini'`)
|
|
101
|
+
|
|
102
|
+
Wraps the `gemini` CLI with `--output-format stream-json`. Each `send()` spawns a new process.
|
|
103
|
+
|
|
104
|
+
- One-shot execution per message (no persistent subprocess)
|
|
105
|
+
- Working directory carries accumulated changes across sends
|
|
106
|
+
- Real token counts from stream-json `result` events (not estimated)
|
|
107
|
+
- Permission modes: `bypassPermissions` → `--yolo`, `default` → `--sandbox`
|
|
108
|
+
- Requires `gemini` CLI installed: `npm install -g @google/gemini-cli`
|
|
109
|
+
|
|
110
|
+
```typescript
|
|
111
|
+
await manager.startSession({
|
|
112
|
+
name: 'gemini-task',
|
|
113
|
+
engine: 'gemini',
|
|
114
|
+
model: 'gemini-3.1-pro-preview',
|
|
115
|
+
cwd: '/project',
|
|
116
|
+
});
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
### Cursor Agent (`engine: 'cursor'`)
|
|
120
|
+
|
|
121
|
+
Wraps the Cursor Agent CLI (`agent`) with `--print --force --output-format stream-json`. Each `send()` spawns a new process.
|
|
122
|
+
|
|
123
|
+
- One-shot execution per message (no persistent subprocess)
|
|
124
|
+
- Working directory via `--workspace` flag
|
|
125
|
+
- Real token counts from stream-json `result` events (camelCase: `inputTokens`, `outputTokens`, `cacheReadTokens`)
|
|
126
|
+
- `--force` enables auto-approval of all file changes
|
|
127
|
+
- `--trust` auto-trusts the workspace without prompting
|
|
128
|
+
- Cursor uses its own model routing (e.g., `sonnet-4`, `gpt-5`, `auto`)
|
|
129
|
+
- Requires Cursor Agent CLI: `curl https://cursor.com/install -fsSL | bash`
|
|
130
|
+
- Binary: `agent` (set `CURSOR_BIN` env var to override)
|
|
131
|
+
|
|
132
|
+
```typescript
|
|
133
|
+
await manager.startSession({
|
|
134
|
+
name: 'cursor-task',
|
|
135
|
+
engine: 'cursor',
|
|
136
|
+
model: 'sonnet-4',
|
|
137
|
+
cwd: '/project',
|
|
138
|
+
});
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
## ISession Interface
|
|
142
|
+
|
|
143
|
+
All engines implement `ISession`, making them interchangeable at the `SessionManager` level:
|
|
144
|
+
|
|
145
|
+
```typescript
|
|
146
|
+
interface ISession {
|
|
147
|
+
// State
|
|
148
|
+
sessionId?: string;
|
|
149
|
+
readonly isReady: boolean;
|
|
150
|
+
readonly isPaused: boolean;
|
|
151
|
+
readonly isBusy: boolean;
|
|
152
|
+
|
|
153
|
+
// Lifecycle
|
|
154
|
+
start(): Promise<this>;
|
|
155
|
+
stop(): void;
|
|
156
|
+
pause(): void;
|
|
157
|
+
resume(): void;
|
|
158
|
+
|
|
159
|
+
// Communication
|
|
160
|
+
send(message, options?): Promise<TurnResult | { requestId; sent }>;
|
|
161
|
+
|
|
162
|
+
// Observability
|
|
163
|
+
getStats(): SessionStats & { sessionId?; uptime };
|
|
164
|
+
getHistory(limit?): Array<{ time; type; event }>;
|
|
165
|
+
getCost(): CostBreakdown;
|
|
166
|
+
|
|
167
|
+
// Context
|
|
168
|
+
compact(summary?): Promise<TurnResult | { requestId; sent }>;
|
|
169
|
+
getEffort(): EffortLevel;
|
|
170
|
+
setEffort(level): void;
|
|
171
|
+
|
|
172
|
+
// Model
|
|
173
|
+
resolveModel(alias): string;
|
|
174
|
+
|
|
175
|
+
// Events (EventEmitter)
|
|
176
|
+
on(event, listener): this;
|
|
177
|
+
emit(event, ...args): boolean;
|
|
178
|
+
}
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
## Team Tools Across Engines
|
|
182
|
+
|
|
183
|
+
Team tools (`team_list`, `team_send`) operate on the same virtual-team layer for **every** engine: the "team" is the set of all active sessions managed by SessionManager.
|
|
184
|
+
|
|
185
|
+
| Engine | `team_list` | `team_send` |
|
|
186
|
+
|--------|------------|-------------|
|
|
187
|
+
| Claude | Lists other active SessionManager sessions | Routes via cross-session inbox |
|
|
188
|
+
| Codex | Lists other active SessionManager sessions | Routes via cross-session inbox |
|
|
189
|
+
| Gemini | Lists other active SessionManager sessions | Routes via cross-session inbox |
|
|
190
|
+
| Cursor | Lists other active SessionManager sessions | Routes via cross-session inbox |
|
|
191
|
+
|
|
192
|
+
Messages are delivered via the inbox system — idle sessions receive immediately, busy sessions queue for later delivery.
|
|
193
|
+
|
|
194
|
+
> **Note:** Claude Code does have a native experimental "Agent Teams" feature (v2.1.32+, `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`), but it is an in-process TUI mechanism with no slash command or stdin-driven messaging — a subprocess wrapper cannot access its mailbox. Plugin team tools therefore use the engine-agnostic virtual team across the board.
|
|
195
|
+
|
|
196
|
+
## Proxy: Any Model via OpenClaw Gateway
|
|
197
|
+
|
|
198
|
+
Claude Code CLI only speaks Anthropic protocol. The built-in proxy translates Anthropic ↔ OpenAI format, letting you drive Claude Code with **any model** routed through the OpenClaw gateway.
|
|
199
|
+
|
|
200
|
+
### Zero Config
|
|
201
|
+
|
|
202
|
+
If OpenClaw gateway is running, everything is automatic:
|
|
203
|
+
|
|
204
|
+
```typescript
|
|
205
|
+
// No baseUrl, no env vars, no extra config
|
|
206
|
+
await manager.startSession({
|
|
207
|
+
name: 'task',
|
|
208
|
+
engine: 'claude',
|
|
209
|
+
model: 'openclaw', // gateway routes to your configured model
|
|
210
|
+
cwd: '/project',
|
|
211
|
+
});
|
|
212
|
+
```
|
|
213
|
+
|
|
214
|
+
What happens behind the scenes:
|
|
215
|
+
1. Plugin reads `~/.openclaw/openclaw.json` for gateway port + auth
|
|
216
|
+
2. Starts a local proxy server (random port, auto-managed)
|
|
217
|
+
3. Claude Code CLI sends Anthropic-format requests → proxy converts to OpenAI → gateway → any model
|
|
218
|
+
|
|
219
|
+
### Manual Config (optional)
|
|
220
|
+
|
|
221
|
+
Override with environment variables if needed:
|
|
222
|
+
|
|
223
|
+
| Variable | Default | Description |
|
|
224
|
+
|----------|---------|-------------|
|
|
225
|
+
| `GATEWAY_URL` | Auto-detected from openclaw.json | Gateway endpoint (e.g. `http://127.0.0.1:18789/v1`) |
|
|
226
|
+
| `GATEWAY_KEY` | Auto-detected from openclaw.json | Gateway auth password/token |
|
|
227
|
+
| `GEMINI_API_KEY` | - | Direct Gemini API access (bypasses gateway) |
|
|
228
|
+
| `OPENAI_API_KEY` | - | Direct OpenAI API access (bypasses gateway) |
|
|
229
|
+
|
|
230
|
+
### Architecture
|
|
231
|
+
|
|
232
|
+
```
|
|
233
|
+
Claude Code CLI (Anthropic format)
|
|
234
|
+
→ Auto-proxy (Anthropic → OpenAI conversion)
|
|
235
|
+
→ OpenClaw Gateway (/v1/chat/completions, model="openclaw")
|
|
236
|
+
→ Any model (Gemini, GPT, local, etc.)
|
|
237
|
+
```
|
|
238
|
+
|
|
239
|
+
## Custom Engine (`engine: 'custom'`)
|
|
240
|
+
|
|
241
|
+
Integrate **any** coding agent CLI without writing engine-specific code. You provide a `CustomEngineConfig` that maps your CLI's flags to OpenClaw session concepts.
|
|
242
|
+
|
|
243
|
+
Two protocol modes:
|
|
244
|
+
- **Persistent** (`persistent: true`) — long-running subprocess with stream-json I/O over stdin/stdout (like Claude Code)
|
|
245
|
+
- **One-shot** (`persistent: false`, default) — new process spawned per `send()` (like Gemini/Codex)
|
|
246
|
+
|
|
247
|
+
### CustomEngineConfig
|
|
248
|
+
|
|
249
|
+
| Field | Type | Required | Description |
|
|
250
|
+
|-------|------|----------|-------------|
|
|
251
|
+
| `name` | string | yes | Display name (used in logs, session IDs) |
|
|
252
|
+
| `bin` | string | yes | Binary path or command name |
|
|
253
|
+
| `binEnv` | string | | Env var name that overrides `bin` at runtime |
|
|
254
|
+
| `persistent` | boolean | | `true` = persistent subprocess, `false` = one-shot (default) |
|
|
255
|
+
| `args` | object | yes | CLI flag mappings (see below) |
|
|
256
|
+
| `permissionModes` | object | | Maps OpenClaw mode names to CLI-specific values |
|
|
257
|
+
| `pricing` | object | | `{ input, output, cached? }` per 1M tokens |
|
|
258
|
+
| `contextWindow` | number | | Context window size (default: 200,000) |
|
|
259
|
+
| `env` | object | | Extra environment variables for the CLI process |
|
|
260
|
+
| `sanitizePatterns` | string[] | | Regex patterns to redact from stderr |
|
|
261
|
+
|
|
262
|
+
### args field
|
|
263
|
+
|
|
264
|
+
| Key | Example | Description |
|
|
265
|
+
|-----|---------|-------------|
|
|
266
|
+
| `print` | `"-p"` | Non-interactive/print mode flag |
|
|
267
|
+
| `outputFormat` | `"--output-format"` | Output format flag |
|
|
268
|
+
| `outputFormatValue` | `"stream-json"` | Value for stream-json output |
|
|
269
|
+
| `inputFormat` | `"--input-format"` | Input format flag (persistent only) |
|
|
270
|
+
| `inputFormatValue` | `"stream-json"` | Value for stream-json input |
|
|
271
|
+
| `skipPermissions` | `"-y"` | Skip all permissions flag |
|
|
272
|
+
| `permissionMode` | `"--permission-mode"` | Permission mode flag |
|
|
273
|
+
| `model` | `"--model"` | Model selection flag |
|
|
274
|
+
| `systemPrompt` | `"--system-prompt"` | System prompt override flag |
|
|
275
|
+
| `appendSystemPrompt` | `"--append-system-prompt"` | Append system prompt flag |
|
|
276
|
+
| `maxTurns` | `"--max-turns"` | Max agent turns flag |
|
|
277
|
+
| `resume` | `"--resume"` | Session resume flag (persistent only) |
|
|
278
|
+
| `verbose` | `"--verbose"` | Verbose output flag |
|
|
279
|
+
| `replayUserMessages` | `"--replay-user-messages"` | Replay user messages (persistent only) |
|
|
280
|
+
| `includePartialMessages` | `"--include-partial-messages"` | Include partial messages (persistent only) |
|
|
281
|
+
| `effort` | `"--effort"` | Effort level flag |
|
|
282
|
+
| `workspace` | `"--workspace"` | Workspace/cwd flag (one-shot only) |
|
|
283
|
+
| `extra` | `["--trust"]` | Additional static arguments |
|
|
284
|
+
|
|
285
|
+
### Example: Persistent mode (Claude Code-compatible CLI)
|
|
286
|
+
|
|
287
|
+
```typescript
|
|
288
|
+
await manager.startSession({
|
|
289
|
+
name: 'my-agent-task',
|
|
290
|
+
engine: 'custom',
|
|
291
|
+
cwd: '/project',
|
|
292
|
+
customEngine: {
|
|
293
|
+
name: 'my-agent',
|
|
294
|
+
bin: 'my-agent',
|
|
295
|
+
binEnv: 'MY_AGENT_BIN',
|
|
296
|
+
persistent: true,
|
|
297
|
+
args: {
|
|
298
|
+
print: '-p',
|
|
299
|
+
outputFormat: '--output-format',
|
|
300
|
+
outputFormatValue: 'stream-json',
|
|
301
|
+
inputFormat: '--input-format',
|
|
302
|
+
inputFormatValue: 'stream-json',
|
|
303
|
+
skipPermissions: '-y',
|
|
304
|
+
permissionMode: '--permission-mode',
|
|
305
|
+
model: '--model',
|
|
306
|
+
systemPrompt: '--system-prompt',
|
|
307
|
+
appendSystemPrompt: '--append-system-prompt',
|
|
308
|
+
maxTurns: '--max-turns',
|
|
309
|
+
resume: '--resume',
|
|
310
|
+
verbose: '--verbose',
|
|
311
|
+
replayUserMessages: '--replay-user-messages',
|
|
312
|
+
includePartialMessages: '--include-partial-messages',
|
|
313
|
+
},
|
|
314
|
+
pricing: { input: 3, output: 15, cached: 0.3 },
|
|
315
|
+
contextWindow: 200_000,
|
|
316
|
+
sanitizePatterns: ['MY_API_KEY=[^\\s]+'],
|
|
317
|
+
},
|
|
318
|
+
});
|
|
319
|
+
```
|
|
320
|
+
|
|
321
|
+
### Example: One-shot mode (simple CLI)
|
|
322
|
+
|
|
323
|
+
```typescript
|
|
324
|
+
await manager.startSession({
|
|
325
|
+
name: 'simple-agent-task',
|
|
326
|
+
engine: 'custom',
|
|
327
|
+
cwd: '/project',
|
|
328
|
+
customEngine: {
|
|
329
|
+
name: 'simple-agent',
|
|
330
|
+
bin: '/usr/local/bin/simple-agent',
|
|
331
|
+
persistent: false, // default
|
|
332
|
+
args: {
|
|
333
|
+
print: '-p',
|
|
334
|
+
outputFormat: '--output-format',
|
|
335
|
+
outputFormatValue: 'stream-json',
|
|
336
|
+
skipPermissions: '--yolo',
|
|
337
|
+
model: '--model',
|
|
338
|
+
workspace: '--workspace',
|
|
339
|
+
extra: ['--no-color'],
|
|
340
|
+
},
|
|
341
|
+
permissionModes: {
|
|
342
|
+
bypassPermissions: 'yolo',
|
|
343
|
+
default: 'sandbox',
|
|
344
|
+
},
|
|
345
|
+
pricing: { input: 1, output: 5 },
|
|
346
|
+
},
|
|
347
|
+
});
|
|
348
|
+
```
|
|
349
|
+
|
|
350
|
+
### Custom Engine in Council
|
|
351
|
+
|
|
352
|
+
Custom engines work in council by setting `engine: 'custom'` and `customEngine` on the agent persona:
|
|
353
|
+
|
|
354
|
+
```typescript
|
|
355
|
+
manager.councilStart('Build feature X', {
|
|
356
|
+
agents: [
|
|
357
|
+
{
|
|
358
|
+
name: 'Planner',
|
|
359
|
+
emoji: '🟠',
|
|
360
|
+
persona: 'Architecture expert',
|
|
361
|
+
engine: 'custom',
|
|
362
|
+
customEngine: { name: 'my-agent', bin: 'my-agent', persistent: true, args: { ... } },
|
|
363
|
+
},
|
|
364
|
+
{ name: 'Reviewer', emoji: '🔵', persona: 'Code reviewer', engine: 'claude', model: 'opus' },
|
|
365
|
+
],
|
|
366
|
+
maxRounds: 10,
|
|
367
|
+
projectDir: '/project',
|
|
368
|
+
});
|
|
369
|
+
```
|
|
370
|
+
|
|
371
|
+
## Adding a New Built-in Engine
|
|
372
|
+
|
|
373
|
+
To add a built-in engine (for CLIs that need custom protocol handling beyond what `CustomEngineConfig` supports):
|
|
374
|
+
|
|
375
|
+
1. Create `src/persistent-<engine>-session.ts` implementing `ISession`
|
|
376
|
+
2. Add the engine name to `EngineType` in `src/types.ts`
|
|
377
|
+
3. Add a case to `SessionManager._createSession()`
|
|
378
|
+
4. Add model pricing to `MODELS[]` in `src/models.ts`
|
|
379
|
+
|
|
380
|
+
The `ISession` interface is deliberately minimal — each engine handles its own subprocess bootstrapping, I/O protocol, and cleanup internally.
|
|
381
|
+
|
|
382
|
+
For most third-party CLIs, the `custom` engine with `CustomEngineConfig` is sufficient and requires zero code changes.
|
|
@@ -0,0 +1,203 @@
|
|
|
1
|
+
# OpenAI-Compatible Bridge
|
|
2
|
+
|
|
3
|
+
> **Cost warning**: This bridge routes requests through the Claude Code CLI, which uses your Claude Max subscription's **extra usage** quota. When OpenClaw's agent loop sends its system prompt (with distinctive tool definitions and agent instructions), Anthropic's backend recognizes this as programmatic/agent traffic and bills it against extra usage — **not** the included allowance. This is by design: the bridge does NOT bypass Anthropic's billing or subscription enforcement. Using it as OpenClaw's primary model backend means every agent turn consumes extra usage credits at standard API rates ($15/M input, $75/M output for Opus). Monitor your usage at [claude.ai/settings/usage](https://claude.ai/settings/usage).
|
|
4
|
+
|
|
5
|
+
The embedded server exposes a drop-in OpenAI-compatible endpoint so any client that speaks `/v1/chat/completions` can talk to a persistent Claude Code (or Codex / Gemini / Cursor) session. The bridge is designed to serve **two kinds of clients as first-class citizens**:
|
|
6
|
+
|
|
7
|
+
1. **Upstream agents** that maintain their own conversation state and forward only the latest user turn — OpenClaw's main agent loop, cron jobs, subagents, programmatic clients.
|
|
8
|
+
2. **OpenAI-compatible webchat / labeling tools** that re-send the full transcript on every turn — ChatGPT-Next-Web, Open WebUI, LobeChat, data-labeling pipelines.
|
|
9
|
+
|
|
10
|
+
Both modes share the same wire protocol; the difference is how a "new conversation" is detected. See [Operator Modes](#operator-modes) below.
|
|
11
|
+
|
|
12
|
+
## Endpoint
|
|
13
|
+
|
|
14
|
+
| | |
|
|
15
|
+
|---|---|
|
|
16
|
+
| **URL** | `http://127.0.0.1:18796/v1/chat/completions` |
|
|
17
|
+
| **Models endpoint** | `GET /v1/models` |
|
|
18
|
+
| **Inspection endpoint** | `GET /v1/sessions` (lists active openai-compat sessions with caching stats) |
|
|
19
|
+
| **Auth** | Bearer token via `Authorization: Bearer $OPENCLAW_SERVER_TOKEN` (set the env var to enable; otherwise no auth and the server is loopback-only) |
|
|
20
|
+
| **Wire format** | OpenAI Chat Completions, both streaming (SSE) and non-streaming |
|
|
21
|
+
|
|
22
|
+
## Session keying
|
|
23
|
+
|
|
24
|
+
Each request is mapped to a long-running session. Once a session exists, subsequent requests with the same key reuse the same persistent CLI subprocess — so Anthropic prompt caching warms across turns. The key is derived in priority order:
|
|
25
|
+
|
|
26
|
+
1. **`X-Session-Id` header** — explicit, highest precedence
|
|
27
|
+
2. **`user` field in the request body** — OpenAI standard field, treated as a stable caller identifier
|
|
28
|
+
3. **`sys-<sha1(model + systemPrompt)[0..12]>`** — automatic fallback so unkeyed callers don't all collapse onto a single shared session
|
|
29
|
+
4. **`'default'`** — only when there is no system prompt AND no model (degenerate empty body)
|
|
30
|
+
|
|
31
|
+
The hash fallback exists because the previous behavior collapsed every unkeyed caller onto one `openai-default` session. In multi-caller setups (OpenClaw routing the main agent + cron jobs + subagents through one gateway) that meant requests serialized against each other and frequently picked up the wrong session's `appendSystemPrompt` — also a privacy leak across distinct callers.
|
|
32
|
+
|
|
33
|
+
The model is mixed into the hash so that two callers with the same system prompt but different requested models (e.g. one wants `claude-opus-4-6`, another wants `claude-sonnet-4-6`) don't collide and silently get responses from the wrong model.
|
|
34
|
+
|
|
35
|
+
The full plugin-side session name is `openai-<key>`.
|
|
36
|
+
|
|
37
|
+
## Operator modes
|
|
38
|
+
|
|
39
|
+
### Default mode — agent / programmatic clients
|
|
40
|
+
|
|
41
|
+
When the env var is **not set**, the bridge assumes upstream callers maintain their own conversation transcript and only forward the latest user turn. Sessions are reused indefinitely. The only signal that starts a new conversation is the explicit reset header:
|
|
42
|
+
|
|
43
|
+
```
|
|
44
|
+
X-Session-Reset: 1
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
(also accepted: `true`, case-insensitive, with whitespace)
|
|
48
|
+
|
|
49
|
+
When the header is present, the existing session for this key is stopped and a fresh one is created. Use this from a client that wants "new chat" semantics under your own control — e.g. when your UI's "Clear History" button is pressed.
|
|
50
|
+
|
|
51
|
+
### Webchat mode — `OPENAI_COMPAT_NEW_CONVO_HEURISTIC=1`
|
|
52
|
+
|
|
53
|
+
When the env var is set to `1`, the bridge additionally restores a legacy heuristic: a request whose `messages` array contains exactly one non-system message (i.e. the conversation has no assistant turns yet) is treated as a fresh conversation. This is the only signal that webchat frontends (ChatGPT-Next-Web, Open WebUI, LobeChat) emit when the user clicks "New Chat" — they clear their UI transcript and post `[system, user]`.
|
|
54
|
+
|
|
55
|
+
Without this flag, those frontends would silently continue the previous CLI session and surface stale context the user thought they had cleared.
|
|
56
|
+
|
|
57
|
+
The env var is read on every request, so ops can flip it via `launchctl setenv` (or equivalent) without restarting the server.
|
|
58
|
+
|
|
59
|
+
| Mode | Best for | New-conversation signals |
|
|
60
|
+
|---|---|---|
|
|
61
|
+
| **Default** | OpenClaw main agent, cron jobs, subagents, scripted clients | `X-Session-Reset: 1` only |
|
|
62
|
+
| **`HEURISTIC=1`** | ChatGPT-Next-Web, Open WebUI, LobeChat, data labeling tools | `X-Session-Reset: 1` **and** `[system, user]` shape |
|
|
63
|
+
|
|
64
|
+
## Status webhook
|
|
65
|
+
|
|
66
|
+
When `OPENAI_COMPAT_STATUS_URL` is set (full HTTP URL), each chat completion sends best-effort `POST` requests with `Content-Type: application/json` and body:
|
|
67
|
+
|
|
68
|
+
| Field | Type | Meaning |
|
|
69
|
+
|---|---|---|
|
|
70
|
+
| `state` | string | `thinking` (turn started), `working` (a tool is running), or `idle` (turn finished or stream closed). |
|
|
71
|
+
| `activity` | string | Short human-readable line, e.g. `Processing request...`, `Reading: foo.ts`, `Running: npm test...`. |
|
|
72
|
+
| `tool` | string \| null | Tool name when `state === working`, otherwise `null`. |
|
|
73
|
+
|
|
74
|
+
Failures are ignored (no retries). Use this from a small local HTTP handler that forwards status into your webchat UI.
|
|
75
|
+
|
|
76
|
+
## Environment variables
|
|
77
|
+
|
|
78
|
+
| Variable | Default | Purpose |
|
|
79
|
+
|---|---|---|
|
|
80
|
+
| `OPENCLAW_SERVER_TOKEN` | (unset) | Bearer token for HTTP auth. Set to enable; written to `~/.openclaw/server-token` for the CLI. |
|
|
81
|
+
| `OPENCLAW_RATE_LIMIT` | `300` | Max requests per IP per 60-second sliding window. |
|
|
82
|
+
| `OPENCLAW_CORS_ORIGINS` | (loopback only) | Set to `*` to allow all origins (the `/v1/*` paths already do this). |
|
|
83
|
+
| `OPENAI_COMPAT_NEW_CONVO_HEURISTIC` | (unset) | Set to `1` to enable webchat mode (see above). |
|
|
84
|
+
| `OPENAI_COMPAT_STATUS_URL` | (unset) | If set, the bridge POSTs JSON status updates to this URL (fire-and-forget, 2s timeout). See [Status webhook](#status-webhook). |
|
|
85
|
+
| `OPENCLAW_SERVE_MAX_SESSIONS` | `32` | Max concurrent OpenAI-compat sessions in serve mode. Bumped from the in-plugin default of 5 because each distinct caller now gets its own `sys-<hash>` session. |
|
|
86
|
+
| `OPENCLAW_SERVE_TTL_MINUTES` | `60` | Idle TTL for OpenAI-compat sessions in serve mode. Idle sessions are reaped by a 60s background loop; persisted disk registry is kept for 7 days so a returning caller is auto-resumed. |
|
|
87
|
+
|
|
88
|
+
## Inspection: `GET /v1/sessions`
|
|
89
|
+
|
|
90
|
+
Returns a JSON list of every active OpenAI-compat session and its caching statistics:
|
|
91
|
+
|
|
92
|
+
```bash
|
|
93
|
+
TOKEN=$(cat ~/.openclaw/server-token)
|
|
94
|
+
curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" | jq
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
Sample response:
|
|
98
|
+
|
|
99
|
+
```json
|
|
100
|
+
{
|
|
101
|
+
"object": "list",
|
|
102
|
+
"data": [
|
|
103
|
+
{
|
|
104
|
+
"key": "sys-a3f81c9d0b27",
|
|
105
|
+
"session_name": "openai-sys-a3f81c9d0b27",
|
|
106
|
+
"model": "claude-opus-4-6",
|
|
107
|
+
"cwd": "/home/user/projects",
|
|
108
|
+
"created": "2026-04-09T03:12:18.441Z",
|
|
109
|
+
"turns": 14,
|
|
110
|
+
"tokens_in": 248312,
|
|
111
|
+
"tokens_out": 38201,
|
|
112
|
+
"cached_tokens": 198104,
|
|
113
|
+
"context_percent": 28,
|
|
114
|
+
"cost_usd": 0.4123
|
|
115
|
+
}
|
|
116
|
+
]
|
|
117
|
+
}
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
The single most important field is **`cached_tokens`**. If it grows turn-over-turn, the persistent CLI is being reused and Anthropic prompt caching is warming. If it stays at 0, something is killing the session every turn — check that no client is sending `X-Session-Reset` unintentionally and that `OPENAI_COMPAT_NEW_CONVO_HEURISTIC` is not set when it shouldn't be.
|
|
121
|
+
|
|
122
|
+
## Smoke tests
|
|
123
|
+
|
|
124
|
+
Run after standing up the server. Set `TOKEN=$(cat ~/.openclaw/server-token)` first.
|
|
125
|
+
|
|
126
|
+
**1. Two distinct system prompts produce two distinct sessions.**
|
|
127
|
+
|
|
128
|
+
```bash
|
|
129
|
+
for SYS in 'You are Alice.' 'You are Bob.'; do
|
|
130
|
+
curl -s http://127.0.0.1:18796/v1/chat/completions \
|
|
131
|
+
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
|
|
132
|
+
-d "{\"model\":\"claude-opus-4-6\",\"messages\":[{\"role\":\"system\",\"content\":\"$SYS\"},{\"role\":\"user\",\"content\":\"hi\"}]}" \
|
|
133
|
+
| jq -r '.id'
|
|
134
|
+
done
|
|
135
|
+
curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" \
|
|
136
|
+
| jq '.data[] | {key, model, turns}'
|
|
137
|
+
# Expected: two rows, distinct sys-<hash> keys.
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
**2. Same system prompt + different model produces two sessions.**
|
|
141
|
+
|
|
142
|
+
```bash
|
|
143
|
+
for M in claude-opus-4-6 claude-sonnet-4-6; do
|
|
144
|
+
curl -s http://127.0.0.1:18796/v1/chat/completions \
|
|
145
|
+
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
|
|
146
|
+
-d "{\"model\":\"$M\",\"messages\":[{\"role\":\"system\",\"content\":\"SAME\"},{\"role\":\"user\",\"content\":\"hi\"}]}" > /dev/null
|
|
147
|
+
done
|
|
148
|
+
curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" | jq '.data | length'
|
|
149
|
+
# Expected: 2
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
**3. `X-Session-Reset: 1` resets cleanly.**
|
|
153
|
+
|
|
154
|
+
```bash
|
|
155
|
+
SID=smoke-reset
|
|
156
|
+
curl -s http://127.0.0.1:18796/v1/chat/completions \
|
|
157
|
+
-H "Authorization: Bearer $TOKEN" -H "X-Session-Id: $SID" -H "Content-Type: application/json" \
|
|
158
|
+
-d '{"model":"claude-opus-4-6","messages":[{"role":"user","content":"remember the word banana"}]}' > /dev/null
|
|
159
|
+
curl -s http://127.0.0.1:18796/v1/chat/completions \
|
|
160
|
+
-H "Authorization: Bearer $TOKEN" -H "X-Session-Id: $SID" -H "X-Session-Reset: 1" -H "Content-Type: application/json" \
|
|
161
|
+
-d '{"model":"claude-opus-4-6","messages":[{"role":"user","content":"what word did I just tell you"}]}' \
|
|
162
|
+
| jq -r '.choices[0].message.content'
|
|
163
|
+
# Expected: model says it has no prior context.
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
**4. `cached_tokens` grows turn-over-turn (the success metric).**
|
|
167
|
+
|
|
168
|
+
```bash
|
|
169
|
+
SID=smoke-cache
|
|
170
|
+
PREAMBLE=$(printf 'x%.0s' {1..3000})
|
|
171
|
+
for i in 1 2 3 4; do
|
|
172
|
+
curl -s http://127.0.0.1:18796/v1/chat/completions \
|
|
173
|
+
-H "Authorization: Bearer $TOKEN" -H "X-Session-Id: $SID" -H "Content-Type: application/json" \
|
|
174
|
+
-d "{\"model\":\"claude-opus-4-6\",\"messages\":[{\"role\":\"system\",\"content\":\"long preamble: $PREAMBLE\"},{\"role\":\"user\",\"content\":\"turn $i\"}]}" > /dev/null
|
|
175
|
+
curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" \
|
|
176
|
+
| jq ".data[] | select(.session_name == \"openai-$SID\") | {turn: $i, cached_tokens, tokens_in}"
|
|
177
|
+
done
|
|
178
|
+
# Expected: cached_tokens climbs substantially by turn 3-4. If it stays at 0,
|
|
179
|
+
# the persistent CLI is still being killed every turn — regression.
|
|
180
|
+
```
|
|
181
|
+
|
|
182
|
+
## Error responses
|
|
183
|
+
|
|
184
|
+
Errors use the OpenAI error envelope:
|
|
185
|
+
|
|
186
|
+
```json
|
|
187
|
+
{ "error": { "message": "...", "type": "invalid_request_error" } }
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
| Status | When |
|
|
191
|
+
|---|---|
|
|
192
|
+
| 400 | `messages` empty/missing, no user message, invalid `max_tokens` |
|
|
193
|
+
| 401 | Missing or wrong bearer token (when auth enabled) |
|
|
194
|
+
| 415 | POST without `Content-Type: application/json` |
|
|
195
|
+
| 429 | Rate limited (`OPENCLAW_RATE_LIMIT` exceeded) |
|
|
196
|
+
| 503 | Failed to start a new session (model unavailable, CLI crashed at boot) |
|
|
197
|
+
| 500 | Mid-turn failure |
|
|
198
|
+
|
|
199
|
+
## Related
|
|
200
|
+
|
|
201
|
+
- [getting-started.md](./getting-started.md) — install + auth setup
|
|
202
|
+
- [sessions.md](./sessions.md) — what a session is and how the lifecycle works under the hood
|
|
203
|
+
- [tools.md](./tools.md) — the full plugin tool surface (council, ultraplan, etc.)
|