@enderfga/claw-orchestrator 3.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (93) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +218 -0
  3. package/assets/banner.jpg +0 -0
  4. package/configs/council-reviewer-prompt.md +82 -0
  5. package/configs/council-system-prompt.md +141 -0
  6. package/dist/bin/cli.d.ts +13 -0
  7. package/dist/bin/cli.js +460 -0
  8. package/dist/bin/cli.js.map +1 -0
  9. package/dist/src/base-oneshot-session.d.ts +87 -0
  10. package/dist/src/base-oneshot-session.js +228 -0
  11. package/dist/src/base-oneshot-session.js.map +1 -0
  12. package/dist/src/circuit-breaker.d.ts +21 -0
  13. package/dist/src/circuit-breaker.js +49 -0
  14. package/dist/src/circuit-breaker.js.map +1 -0
  15. package/dist/src/consensus.d.ts +20 -0
  16. package/dist/src/consensus.js +52 -0
  17. package/dist/src/consensus.js.map +1 -0
  18. package/dist/src/constants.d.ts +129 -0
  19. package/dist/src/constants.js +138 -0
  20. package/dist/src/constants.js.map +1 -0
  21. package/dist/src/council.d.ts +67 -0
  22. package/dist/src/council.js +914 -0
  23. package/dist/src/council.js.map +1 -0
  24. package/dist/src/embedded-server.d.ts +25 -0
  25. package/dist/src/embedded-server.js +360 -0
  26. package/dist/src/embedded-server.js.map +1 -0
  27. package/dist/src/inbox-manager.d.ts +38 -0
  28. package/dist/src/inbox-manager.js +111 -0
  29. package/dist/src/inbox-manager.js.map +1 -0
  30. package/dist/src/index.d.ts +63 -0
  31. package/dist/src/index.js +973 -0
  32. package/dist/src/index.js.map +1 -0
  33. package/dist/src/logger.d.ts +16 -0
  34. package/dist/src/logger.js +44 -0
  35. package/dist/src/logger.js.map +1 -0
  36. package/dist/src/models.d.ts +69 -0
  37. package/dist/src/models.js +299 -0
  38. package/dist/src/models.js.map +1 -0
  39. package/dist/src/openai-compat.d.ts +224 -0
  40. package/dist/src/openai-compat.js +756 -0
  41. package/dist/src/openai-compat.js.map +1 -0
  42. package/dist/src/persistent-codex-app-session.d.ts +108 -0
  43. package/dist/src/persistent-codex-app-session.js +465 -0
  44. package/dist/src/persistent-codex-app-session.js.map +1 -0
  45. package/dist/src/persistent-codex-session.d.ts +37 -0
  46. package/dist/src/persistent-codex-session.js +208 -0
  47. package/dist/src/persistent-codex-session.js.map +1 -0
  48. package/dist/src/persistent-cursor-session.d.ts +21 -0
  49. package/dist/src/persistent-cursor-session.js +241 -0
  50. package/dist/src/persistent-cursor-session.js.map +1 -0
  51. package/dist/src/persistent-custom-session.d.ts +78 -0
  52. package/dist/src/persistent-custom-session.js +938 -0
  53. package/dist/src/persistent-custom-session.js.map +1 -0
  54. package/dist/src/persistent-gemini-session.d.ts +21 -0
  55. package/dist/src/persistent-gemini-session.js +216 -0
  56. package/dist/src/persistent-gemini-session.js.map +1 -0
  57. package/dist/src/persistent-session.d.ts +80 -0
  58. package/dist/src/persistent-session.js +745 -0
  59. package/dist/src/persistent-session.js.map +1 -0
  60. package/dist/src/proxy/anthropic-adapter.d.ts +136 -0
  61. package/dist/src/proxy/anthropic-adapter.js +392 -0
  62. package/dist/src/proxy/anthropic-adapter.js.map +1 -0
  63. package/dist/src/proxy/handler.d.ts +39 -0
  64. package/dist/src/proxy/handler.js +365 -0
  65. package/dist/src/proxy/handler.js.map +1 -0
  66. package/dist/src/proxy/schema-cleaner.d.ts +11 -0
  67. package/dist/src/proxy/schema-cleaner.js +34 -0
  68. package/dist/src/proxy/schema-cleaner.js.map +1 -0
  69. package/dist/src/proxy/thought-cache.d.ts +19 -0
  70. package/dist/src/proxy/thought-cache.js +53 -0
  71. package/dist/src/proxy/thought-cache.js.map +1 -0
  72. package/dist/src/session-manager.d.ts +317 -0
  73. package/dist/src/session-manager.js +1528 -0
  74. package/dist/src/session-manager.js.map +1 -0
  75. package/dist/src/types.d.ts +513 -0
  76. package/dist/src/types.js +8 -0
  77. package/dist/src/types.js.map +1 -0
  78. package/dist/src/validation.d.ts +31 -0
  79. package/dist/src/validation.js +104 -0
  80. package/dist/src/validation.js.map +1 -0
  81. package/openclaw.plugin.json +122 -0
  82. package/package.json +84 -0
  83. package/skills/SKILL.md +184 -0
  84. package/skills/references/claude-cli-tracking.md +25 -0
  85. package/skills/references/cli.md +187 -0
  86. package/skills/references/council.md +210 -0
  87. package/skills/references/getting-started.md +133 -0
  88. package/skills/references/inbox.md +81 -0
  89. package/skills/references/multi-engine.md +382 -0
  90. package/skills/references/openai-compat.md +203 -0
  91. package/skills/references/sessions.md +191 -0
  92. package/skills/references/tools.md +418 -0
  93. package/skills/references/ultra.md +126 -0
@@ -0,0 +1,382 @@
1
+ # Multi-Engine
2
+
3
+ Claw Orchestrator supports multiple coding CLI engines behind a unified `ISession` interface. Each engine manages its own subprocess, event stream, and cost tracking independently.
4
+
5
+ ## Architecture
6
+
7
+ ```
8
+ SessionManager
9
+ ├── engine: 'claude' → PersistentClaudeSession
10
+ │ └── Wraps: claude CLI (stream-json protocol, persistent subprocess)
11
+ ├── engine: 'codex' → PersistentCodexSession
12
+ │ └── Wraps: codex exec --sandbox workspace-write --json (per-message spawning)
13
+ ├── engine: 'codex-app' → PersistentCodexAppServerSession
14
+ │ └── Wraps: codex app-server --listen stdio:// (long-running JSON-RPC; required for /goal)
15
+ ├── engine: 'gemini' → PersistentGeminiSession
16
+ │ └── Wraps: gemini -p --output-format stream-json (per-message spawning)
17
+ ├── engine: 'cursor' → PersistentCursorSession
18
+ │ └── Wraps: agent -p --force --output-format stream-json (per-message spawning)
19
+ └── engine: 'custom' → PersistentCustomSession
20
+ └── Wraps: any CLI via user-provided CustomEngineConfig
21
+ ```
22
+
23
+ ## Supported Engines
24
+
25
+ ### Claude Code (`engine: 'claude'`)
26
+
27
+ Default engine. Long-running subprocess with streaming JSON I/O. Tested with Claude Code CLI **2.1.126**.
28
+
29
+ - Persistent multi-turn conversations
30
+ - Real-time streaming (text, tool_use, tool_result, system events)
31
+ - Session resume via `--resume`
32
+ - Full cost tracking from API usage data
33
+ - Hook lifecycle events (`includeHookEvents`), permission delegation (`permissionPromptTool`), prompt cache optimization (`bare` + `excludeDynamicSystemPromptSections` + `enablePromptCaching1H`), debug control, `--from-pr` resume, and MCP channel subscriptions
34
+ - Fork subagent (`forkSubagent`), tool search (`enableToolSearch`), OpenTelemetry logging toggles (`otelLogUserPrompts`, `otelLogRawApiBodies`), `xhigh` effort tier (Opus 4.7), and `stats.pluginErrors` capture — see [CLI 2.1.121 options in SKILL.md](../SKILL.md) and [tools.md](./tools.md)
35
+
36
+ > **Behavior changes from upstream Claude CLI 2.1.121** (worth knowing if you set permission rules):
37
+ > - `--agent` / `--print` now enforce agent frontmatter `permissionMode`, `tools`, `disallowedTools` (was advisory). Affects `council` agent personas.
38
+ > - `Bash(find:*)` permission rule no longer auto-approves `find -exec` or `find -delete`. Add explicit rules if you depend on these.
39
+ > - `--dangerously-skip-permissions` also skips prompts for `.claude/skills/` directory. Treat with care.
40
+ > - Distributed tracing context (`TRACEPARENT` / `TRACESTATE`) is automatically forwarded to the child process — set them in the parent before starting the session.
41
+
42
+ ```typescript
43
+ await manager.startSession({
44
+ name: 'claude-task',
45
+ engine: 'claude', // default, can omit
46
+ model: 'opus',
47
+ cwd: '/project',
48
+ });
49
+ ```
50
+
51
+ ### OpenAI Codex (`engine: 'codex'`)
52
+
53
+ Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested with `codex` CLI **0.128.0**.
54
+
55
+ - Non-interactive execution via `codex exec --sandbox workspace-write --json` (replaces the deprecated `--full-auto` flag from earlier Codex versions)
56
+ - Real per-turn `usage` from the `turn.completed` JSON event (input, output, cached, reasoning tokens)
57
+ - Per-session continuity: the `thread_id` from the first turn's `thread.started` event is captured and reused via `codex exec resume <id>` for subsequent sends, so the model sees prior turns
58
+ - One-shot execution per message (no persistent subprocess between sends)
59
+ - Working directory passed via `-C` flag
60
+ - Default model: `gpt-5.5`
61
+ - Requires `codex` CLI >= 0.119 (for `exec resume`): `npm install -g @openai/codex`
62
+ - **Does not support `/goal`** — for that, use `engine: 'codex-app'` below
63
+
64
+ ```typescript
65
+ await manager.startSession({
66
+ name: 'codex-task',
67
+ engine: 'codex',
68
+ model: 'gpt-5.5',
69
+ cwd: '/project',
70
+ sandboxMode: 'workspace-write', // optional, this is the default
71
+ });
72
+ ```
73
+
74
+ ### OpenAI Codex App-Server (`engine: 'codex-app'`)
75
+
76
+ Wraps `codex app-server --listen stdio:// --enable goals` as a long-running JSON-RPC subprocess. **Required for `/goal` long-horizon objective support** — Codex's exec subcommand has no slash-command surface.
77
+
78
+ - Long-running subprocess; one `codex app-server` per session
79
+ - JSON-RPC 2.0 over stdio with v2 protocol method names (`initialize`, `thread/start`, `turn/start`, ...)
80
+ - Real-time streaming via `item/agentMessage/delta` notifications
81
+ - Cumulative token tracking from `thread/tokenUsage/updated` notifications
82
+ - Goal lifecycle observation via `thread/goal/updated` and `thread/goal/cleared` notifications
83
+ - Goal control via the `codex_goal_*` tools (which internally send the `/goal` slash command as user text — see [tools.md](./tools.md#codex-7))
84
+
85
+ > **Feature-flag risk.** The `goals` feature is marked "under development" in Codex 0.128.0 and has known bugs (e.g. issue #20591). The session class always passes `--enable goals` so it works the moment upstream stabilizes the feature, but during the transition period some goal commands may fail or be silently dropped on the server side. The wrapper layer is unaffected.
86
+
87
+ ```typescript
88
+ await manager.startSession({
89
+ name: 'codex-goal-task',
90
+ engine: 'codex-app',
91
+ model: 'gpt-5.5',
92
+ cwd: '/project',
93
+ });
94
+ // Then either:
95
+ // await manager.codexGoalCommand('codex-goal-task', 'build a tic-tac-toe app');
96
+ // or via the codex_goal_set tool:
97
+ // await tool('codex_goal_set', { name: 'codex-goal-task', objective: 'build a tic-tac-toe app' });
98
+ ```
99
+
100
+ ### Google Gemini (`engine: 'gemini'`)
101
+
102
+ Wraps the `gemini` CLI with `--output-format stream-json`. Each `send()` spawns a new process.
103
+
104
+ - One-shot execution per message (no persistent subprocess)
105
+ - Working directory carries accumulated changes across sends
106
+ - Real token counts from stream-json `result` events (not estimated)
107
+ - Permission modes: `bypassPermissions` → `--yolo`, `default` → `--sandbox`
108
+ - Requires `gemini` CLI installed: `npm install -g @google/gemini-cli`
109
+
110
+ ```typescript
111
+ await manager.startSession({
112
+ name: 'gemini-task',
113
+ engine: 'gemini',
114
+ model: 'gemini-3.1-pro-preview',
115
+ cwd: '/project',
116
+ });
117
+ ```
118
+
119
+ ### Cursor Agent (`engine: 'cursor'`)
120
+
121
+ Wraps the Cursor Agent CLI (`agent`) with `--print --force --output-format stream-json`. Each `send()` spawns a new process.
122
+
123
+ - One-shot execution per message (no persistent subprocess)
124
+ - Working directory via `--workspace` flag
125
+ - Real token counts from stream-json `result` events (camelCase: `inputTokens`, `outputTokens`, `cacheReadTokens`)
126
+ - `--force` enables auto-approval of all file changes
127
+ - `--trust` auto-trusts the workspace without prompting
128
+ - Cursor uses its own model routing (e.g., `sonnet-4`, `gpt-5`, `auto`)
129
+ - Requires Cursor Agent CLI: `curl https://cursor.com/install -fsSL | bash`
130
+ - Binary: `agent` (set `CURSOR_BIN` env var to override)
131
+
132
+ ```typescript
133
+ await manager.startSession({
134
+ name: 'cursor-task',
135
+ engine: 'cursor',
136
+ model: 'sonnet-4',
137
+ cwd: '/project',
138
+ });
139
+ ```
140
+
141
+ ## ISession Interface
142
+
143
+ All engines implement `ISession`, making them interchangeable at the `SessionManager` level:
144
+
145
+ ```typescript
146
+ interface ISession {
147
+ // State
148
+ sessionId?: string;
149
+ readonly isReady: boolean;
150
+ readonly isPaused: boolean;
151
+ readonly isBusy: boolean;
152
+
153
+ // Lifecycle
154
+ start(): Promise<this>;
155
+ stop(): void;
156
+ pause(): void;
157
+ resume(): void;
158
+
159
+ // Communication
160
+ send(message, options?): Promise<TurnResult | { requestId; sent }>;
161
+
162
+ // Observability
163
+ getStats(): SessionStats & { sessionId?; uptime };
164
+ getHistory(limit?): Array<{ time; type; event }>;
165
+ getCost(): CostBreakdown;
166
+
167
+ // Context
168
+ compact(summary?): Promise<TurnResult | { requestId; sent }>;
169
+ getEffort(): EffortLevel;
170
+ setEffort(level): void;
171
+
172
+ // Model
173
+ resolveModel(alias): string;
174
+
175
+ // Events (EventEmitter)
176
+ on(event, listener): this;
177
+ emit(event, ...args): boolean;
178
+ }
179
+ ```
180
+
181
+ ## Team Tools Across Engines
182
+
183
+ Team tools (`team_list`, `team_send`) operate on the same virtual-team layer for **every** engine: the "team" is the set of all active sessions managed by SessionManager.
184
+
185
+ | Engine | `team_list` | `team_send` |
186
+ |--------|------------|-------------|
187
+ | Claude | Lists other active SessionManager sessions | Routes via cross-session inbox |
188
+ | Codex | Lists other active SessionManager sessions | Routes via cross-session inbox |
189
+ | Gemini | Lists other active SessionManager sessions | Routes via cross-session inbox |
190
+ | Cursor | Lists other active SessionManager sessions | Routes via cross-session inbox |
191
+
192
+ Messages are delivered via the inbox system — idle sessions receive immediately, busy sessions queue for later delivery.
193
+
194
+ > **Note:** Claude Code does have a native experimental "Agent Teams" feature (v2.1.32+, `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`), but it is an in-process TUI mechanism with no slash command or stdin-driven messaging — a subprocess wrapper cannot access its mailbox. Plugin team tools therefore use the engine-agnostic virtual team across the board.
195
+
196
+ ## Proxy: Any Model via OpenClaw Gateway
197
+
198
+ Claude Code CLI only speaks Anthropic protocol. The built-in proxy translates Anthropic ↔ OpenAI format, letting you drive Claude Code with **any model** routed through the OpenClaw gateway.
199
+
200
+ ### Zero Config
201
+
202
+ If OpenClaw gateway is running, everything is automatic:
203
+
204
+ ```typescript
205
+ // No baseUrl, no env vars, no extra config
206
+ await manager.startSession({
207
+ name: 'task',
208
+ engine: 'claude',
209
+ model: 'openclaw', // gateway routes to your configured model
210
+ cwd: '/project',
211
+ });
212
+ ```
213
+
214
+ What happens behind the scenes:
215
+ 1. Plugin reads `~/.openclaw/openclaw.json` for gateway port + auth
216
+ 2. Starts a local proxy server (random port, auto-managed)
217
+ 3. Claude Code CLI sends Anthropic-format requests → proxy converts to OpenAI → gateway → any model
218
+
219
+ ### Manual Config (optional)
220
+
221
+ Override with environment variables if needed:
222
+
223
+ | Variable | Default | Description |
224
+ |----------|---------|-------------|
225
+ | `GATEWAY_URL` | Auto-detected from openclaw.json | Gateway endpoint (e.g. `http://127.0.0.1:18789/v1`) |
226
+ | `GATEWAY_KEY` | Auto-detected from openclaw.json | Gateway auth password/token |
227
+ | `GEMINI_API_KEY` | - | Direct Gemini API access (bypasses gateway) |
228
+ | `OPENAI_API_KEY` | - | Direct OpenAI API access (bypasses gateway) |
229
+
230
+ ### Architecture
231
+
232
+ ```
233
+ Claude Code CLI (Anthropic format)
234
+ → Auto-proxy (Anthropic → OpenAI conversion)
235
+ → OpenClaw Gateway (/v1/chat/completions, model="openclaw")
236
+ → Any model (Gemini, GPT, local, etc.)
237
+ ```
238
+
239
+ ## Custom Engine (`engine: 'custom'`)
240
+
241
+ Integrate **any** coding agent CLI without writing engine-specific code. You provide a `CustomEngineConfig` that maps your CLI's flags to OpenClaw session concepts.
242
+
243
+ Two protocol modes:
244
+ - **Persistent** (`persistent: true`) — long-running subprocess with stream-json I/O over stdin/stdout (like Claude Code)
245
+ - **One-shot** (`persistent: false`, default) — new process spawned per `send()` (like Gemini/Codex)
246
+
247
+ ### CustomEngineConfig
248
+
249
+ | Field | Type | Required | Description |
250
+ |-------|------|----------|-------------|
251
+ | `name` | string | yes | Display name (used in logs, session IDs) |
252
+ | `bin` | string | yes | Binary path or command name |
253
+ | `binEnv` | string | | Env var name that overrides `bin` at runtime |
254
+ | `persistent` | boolean | | `true` = persistent subprocess, `false` = one-shot (default) |
255
+ | `args` | object | yes | CLI flag mappings (see below) |
256
+ | `permissionModes` | object | | Maps OpenClaw mode names to CLI-specific values |
257
+ | `pricing` | object | | `{ input, output, cached? }` per 1M tokens |
258
+ | `contextWindow` | number | | Context window size (default: 200,000) |
259
+ | `env` | object | | Extra environment variables for the CLI process |
260
+ | `sanitizePatterns` | string[] | | Regex patterns to redact from stderr |
261
+
262
+ ### args field
263
+
264
+ | Key | Example | Description |
265
+ |-----|---------|-------------|
266
+ | `print` | `"-p"` | Non-interactive/print mode flag |
267
+ | `outputFormat` | `"--output-format"` | Output format flag |
268
+ | `outputFormatValue` | `"stream-json"` | Value for stream-json output |
269
+ | `inputFormat` | `"--input-format"` | Input format flag (persistent only) |
270
+ | `inputFormatValue` | `"stream-json"` | Value for stream-json input |
271
+ | `skipPermissions` | `"-y"` | Skip all permissions flag |
272
+ | `permissionMode` | `"--permission-mode"` | Permission mode flag |
273
+ | `model` | `"--model"` | Model selection flag |
274
+ | `systemPrompt` | `"--system-prompt"` | System prompt override flag |
275
+ | `appendSystemPrompt` | `"--append-system-prompt"` | Append system prompt flag |
276
+ | `maxTurns` | `"--max-turns"` | Max agent turns flag |
277
+ | `resume` | `"--resume"` | Session resume flag (persistent only) |
278
+ | `verbose` | `"--verbose"` | Verbose output flag |
279
+ | `replayUserMessages` | `"--replay-user-messages"` | Replay user messages (persistent only) |
280
+ | `includePartialMessages` | `"--include-partial-messages"` | Include partial messages (persistent only) |
281
+ | `effort` | `"--effort"` | Effort level flag |
282
+ | `workspace` | `"--workspace"` | Workspace/cwd flag (one-shot only) |
283
+ | `extra` | `["--trust"]` | Additional static arguments |
284
+
285
+ ### Example: Persistent mode (Claude Code-compatible CLI)
286
+
287
+ ```typescript
288
+ await manager.startSession({
289
+ name: 'my-agent-task',
290
+ engine: 'custom',
291
+ cwd: '/project',
292
+ customEngine: {
293
+ name: 'my-agent',
294
+ bin: 'my-agent',
295
+ binEnv: 'MY_AGENT_BIN',
296
+ persistent: true,
297
+ args: {
298
+ print: '-p',
299
+ outputFormat: '--output-format',
300
+ outputFormatValue: 'stream-json',
301
+ inputFormat: '--input-format',
302
+ inputFormatValue: 'stream-json',
303
+ skipPermissions: '-y',
304
+ permissionMode: '--permission-mode',
305
+ model: '--model',
306
+ systemPrompt: '--system-prompt',
307
+ appendSystemPrompt: '--append-system-prompt',
308
+ maxTurns: '--max-turns',
309
+ resume: '--resume',
310
+ verbose: '--verbose',
311
+ replayUserMessages: '--replay-user-messages',
312
+ includePartialMessages: '--include-partial-messages',
313
+ },
314
+ pricing: { input: 3, output: 15, cached: 0.3 },
315
+ contextWindow: 200_000,
316
+ sanitizePatterns: ['MY_API_KEY=[^\\s]+'],
317
+ },
318
+ });
319
+ ```
320
+
321
+ ### Example: One-shot mode (simple CLI)
322
+
323
+ ```typescript
324
+ await manager.startSession({
325
+ name: 'simple-agent-task',
326
+ engine: 'custom',
327
+ cwd: '/project',
328
+ customEngine: {
329
+ name: 'simple-agent',
330
+ bin: '/usr/local/bin/simple-agent',
331
+ persistent: false, // default
332
+ args: {
333
+ print: '-p',
334
+ outputFormat: '--output-format',
335
+ outputFormatValue: 'stream-json',
336
+ skipPermissions: '--yolo',
337
+ model: '--model',
338
+ workspace: '--workspace',
339
+ extra: ['--no-color'],
340
+ },
341
+ permissionModes: {
342
+ bypassPermissions: 'yolo',
343
+ default: 'sandbox',
344
+ },
345
+ pricing: { input: 1, output: 5 },
346
+ },
347
+ });
348
+ ```
349
+
350
+ ### Custom Engine in Council
351
+
352
+ Custom engines work in council by setting `engine: 'custom'` and `customEngine` on the agent persona:
353
+
354
+ ```typescript
355
+ manager.councilStart('Build feature X', {
356
+ agents: [
357
+ {
358
+ name: 'Planner',
359
+ emoji: '🟠',
360
+ persona: 'Architecture expert',
361
+ engine: 'custom',
362
+ customEngine: { name: 'my-agent', bin: 'my-agent', persistent: true, args: { ... } },
363
+ },
364
+ { name: 'Reviewer', emoji: '🔵', persona: 'Code reviewer', engine: 'claude', model: 'opus' },
365
+ ],
366
+ maxRounds: 10,
367
+ projectDir: '/project',
368
+ });
369
+ ```
370
+
371
+ ## Adding a New Built-in Engine
372
+
373
+ To add a built-in engine (for CLIs that need custom protocol handling beyond what `CustomEngineConfig` supports):
374
+
375
+ 1. Create `src/persistent-<engine>-session.ts` implementing `ISession`
376
+ 2. Add the engine name to `EngineType` in `src/types.ts`
377
+ 3. Add a case to `SessionManager._createSession()`
378
+ 4. Add model pricing to `MODELS[]` in `src/models.ts`
379
+
380
+ The `ISession` interface is deliberately minimal — each engine handles its own subprocess bootstrapping, I/O protocol, and cleanup internally.
381
+
382
+ For most third-party CLIs, the `custom` engine with `CustomEngineConfig` is sufficient and requires zero code changes.
@@ -0,0 +1,203 @@
1
+ # OpenAI-Compatible Bridge
2
+
3
+ > **Cost warning**: This bridge routes requests through the Claude Code CLI, which uses your Claude Max subscription's **extra usage** quota. When OpenClaw's agent loop sends its system prompt (with distinctive tool definitions and agent instructions), Anthropic's backend recognizes this as programmatic/agent traffic and bills it against extra usage — **not** the included allowance. This is by design: the bridge does NOT bypass Anthropic's billing or subscription enforcement. Using it as OpenClaw's primary model backend means every agent turn consumes extra usage credits at standard API rates ($15/M input, $75/M output for Opus). Monitor your usage at [claude.ai/settings/usage](https://claude.ai/settings/usage).
4
+
5
+ The embedded server exposes a drop-in OpenAI-compatible endpoint so any client that speaks `/v1/chat/completions` can talk to a persistent Claude Code (or Codex / Gemini / Cursor) session. The bridge is designed to serve **two kinds of clients as first-class citizens**:
6
+
7
+ 1. **Upstream agents** that maintain their own conversation state and forward only the latest user turn — OpenClaw's main agent loop, cron jobs, subagents, programmatic clients.
8
+ 2. **OpenAI-compatible webchat / labeling tools** that re-send the full transcript on every turn — ChatGPT-Next-Web, Open WebUI, LobeChat, data-labeling pipelines.
9
+
10
+ Both modes share the same wire protocol; the difference is how a "new conversation" is detected. See [Operator Modes](#operator-modes) below.
11
+
12
+ ## Endpoint
13
+
14
+ | | |
15
+ |---|---|
16
+ | **URL** | `http://127.0.0.1:18796/v1/chat/completions` |
17
+ | **Models endpoint** | `GET /v1/models` |
18
+ | **Inspection endpoint** | `GET /v1/sessions` (lists active openai-compat sessions with caching stats) |
19
+ | **Auth** | Bearer token via `Authorization: Bearer $OPENCLAW_SERVER_TOKEN` (set the env var to enable; otherwise no auth and the server is loopback-only) |
20
+ | **Wire format** | OpenAI Chat Completions, both streaming (SSE) and non-streaming |
21
+
22
+ ## Session keying
23
+
24
+ Each request is mapped to a long-running session. Once a session exists, subsequent requests with the same key reuse the same persistent CLI subprocess — so Anthropic prompt caching warms across turns. The key is derived in priority order:
25
+
26
+ 1. **`X-Session-Id` header** — explicit, highest precedence
27
+ 2. **`user` field in the request body** — OpenAI standard field, treated as a stable caller identifier
28
+ 3. **`sys-<sha1(model + systemPrompt)[0..12]>`** — automatic fallback so unkeyed callers don't all collapse onto a single shared session
29
+ 4. **`'default'`** — only when there is no system prompt AND no model (degenerate empty body)
30
+
31
+ The hash fallback exists because the previous behavior collapsed every unkeyed caller onto one `openai-default` session. In multi-caller setups (OpenClaw routing the main agent + cron jobs + subagents through one gateway) that meant requests serialized against each other and frequently picked up the wrong session's `appendSystemPrompt` — also a privacy leak across distinct callers.
32
+
33
+ The model is mixed into the hash so that two callers with the same system prompt but different requested models (e.g. one wants `claude-opus-4-6`, another wants `claude-sonnet-4-6`) don't collide and silently get responses from the wrong model.
34
+
35
+ The full plugin-side session name is `openai-<key>`.
36
+
37
+ ## Operator modes
38
+
39
+ ### Default mode — agent / programmatic clients
40
+
41
+ When the env var is **not set**, the bridge assumes upstream callers maintain their own conversation transcript and only forward the latest user turn. Sessions are reused indefinitely. The only signal that starts a new conversation is the explicit reset header:
42
+
43
+ ```
44
+ X-Session-Reset: 1
45
+ ```
46
+
47
+ (also accepted: `true`, case-insensitive, with whitespace)
48
+
49
+ When the header is present, the existing session for this key is stopped and a fresh one is created. Use this from a client that wants "new chat" semantics under your own control — e.g. when your UI's "Clear History" button is pressed.
50
+
51
+ ### Webchat mode — `OPENAI_COMPAT_NEW_CONVO_HEURISTIC=1`
52
+
53
+ When the env var is set to `1`, the bridge additionally restores a legacy heuristic: a request whose `messages` array contains exactly one non-system message (i.e. the conversation has no assistant turns yet) is treated as a fresh conversation. This is the only signal that webchat frontends (ChatGPT-Next-Web, Open WebUI, LobeChat) emit when the user clicks "New Chat" — they clear their UI transcript and post `[system, user]`.
54
+
55
+ Without this flag, those frontends would silently continue the previous CLI session and surface stale context the user thought they had cleared.
56
+
57
+ The env var is read on every request, so ops can flip it via `launchctl setenv` (or equivalent) without restarting the server.
58
+
59
+ | Mode | Best for | New-conversation signals |
60
+ |---|---|---|
61
+ | **Default** | OpenClaw main agent, cron jobs, subagents, scripted clients | `X-Session-Reset: 1` only |
62
+ | **`HEURISTIC=1`** | ChatGPT-Next-Web, Open WebUI, LobeChat, data labeling tools | `X-Session-Reset: 1` **and** `[system, user]` shape |
63
+
64
+ ## Status webhook
65
+
66
+ When `OPENAI_COMPAT_STATUS_URL` is set (full HTTP URL), each chat completion sends best-effort `POST` requests with `Content-Type: application/json` and body:
67
+
68
+ | Field | Type | Meaning |
69
+ |---|---|---|
70
+ | `state` | string | `thinking` (turn started), `working` (a tool is running), or `idle` (turn finished or stream closed). |
71
+ | `activity` | string | Short human-readable line, e.g. `Processing request...`, `Reading: foo.ts`, `Running: npm test...`. |
72
+ | `tool` | string \| null | Tool name when `state === working`, otherwise `null`. |
73
+
74
+ Failures are ignored (no retries). Use this from a small local HTTP handler that forwards status into your webchat UI.
75
+
76
+ ## Environment variables
77
+
78
+ | Variable | Default | Purpose |
79
+ |---|---|---|
80
+ | `OPENCLAW_SERVER_TOKEN` | (unset) | Bearer token for HTTP auth. Set to enable; written to `~/.openclaw/server-token` for the CLI. |
81
+ | `OPENCLAW_RATE_LIMIT` | `300` | Max requests per IP per 60-second sliding window. |
82
+ | `OPENCLAW_CORS_ORIGINS` | (loopback only) | Set to `*` to allow all origins (the `/v1/*` paths already do this). |
83
+ | `OPENAI_COMPAT_NEW_CONVO_HEURISTIC` | (unset) | Set to `1` to enable webchat mode (see above). |
84
+ | `OPENAI_COMPAT_STATUS_URL` | (unset) | If set, the bridge POSTs JSON status updates to this URL (fire-and-forget, 2s timeout). See [Status webhook](#status-webhook). |
85
+ | `OPENCLAW_SERVE_MAX_SESSIONS` | `32` | Max concurrent OpenAI-compat sessions in serve mode. Bumped from the in-plugin default of 5 because each distinct caller now gets its own `sys-<hash>` session. |
86
+ | `OPENCLAW_SERVE_TTL_MINUTES` | `60` | Idle TTL for OpenAI-compat sessions in serve mode. Idle sessions are reaped by a 60s background loop; persisted disk registry is kept for 7 days so a returning caller is auto-resumed. |
87
+
88
+ ## Inspection: `GET /v1/sessions`
89
+
90
+ Returns a JSON list of every active OpenAI-compat session and its caching statistics:
91
+
92
+ ```bash
93
+ TOKEN=$(cat ~/.openclaw/server-token)
94
+ curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" | jq
95
+ ```
96
+
97
+ Sample response:
98
+
99
+ ```json
100
+ {
101
+ "object": "list",
102
+ "data": [
103
+ {
104
+ "key": "sys-a3f81c9d0b27",
105
+ "session_name": "openai-sys-a3f81c9d0b27",
106
+ "model": "claude-opus-4-6",
107
+ "cwd": "/home/user/projects",
108
+ "created": "2026-04-09T03:12:18.441Z",
109
+ "turns": 14,
110
+ "tokens_in": 248312,
111
+ "tokens_out": 38201,
112
+ "cached_tokens": 198104,
113
+ "context_percent": 28,
114
+ "cost_usd": 0.4123
115
+ }
116
+ ]
117
+ }
118
+ ```
119
+
120
+ The single most important field is **`cached_tokens`**. If it grows turn-over-turn, the persistent CLI is being reused and Anthropic prompt caching is warming. If it stays at 0, something is killing the session every turn — check that no client is sending `X-Session-Reset` unintentionally and that `OPENAI_COMPAT_NEW_CONVO_HEURISTIC` is not set when it shouldn't be.
121
+
122
+ ## Smoke tests
123
+
124
+ Run after standing up the server. Set `TOKEN=$(cat ~/.openclaw/server-token)` first.
125
+
126
+ **1. Two distinct system prompts produce two distinct sessions.**
127
+
128
+ ```bash
129
+ for SYS in 'You are Alice.' 'You are Bob.'; do
130
+ curl -s http://127.0.0.1:18796/v1/chat/completions \
131
+ -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
132
+ -d "{\"model\":\"claude-opus-4-6\",\"messages\":[{\"role\":\"system\",\"content\":\"$SYS\"},{\"role\":\"user\",\"content\":\"hi\"}]}" \
133
+ | jq -r '.id'
134
+ done
135
+ curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" \
136
+ | jq '.data[] | {key, model, turns}'
137
+ # Expected: two rows, distinct sys-<hash> keys.
138
+ ```
139
+
140
+ **2. Same system prompt + different model produces two sessions.**
141
+
142
+ ```bash
143
+ for M in claude-opus-4-6 claude-sonnet-4-6; do
144
+ curl -s http://127.0.0.1:18796/v1/chat/completions \
145
+ -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
146
+ -d "{\"model\":\"$M\",\"messages\":[{\"role\":\"system\",\"content\":\"SAME\"},{\"role\":\"user\",\"content\":\"hi\"}]}" > /dev/null
147
+ done
148
+ curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" | jq '.data | length'
149
+ # Expected: 2
150
+ ```
151
+
152
+ **3. `X-Session-Reset: 1` resets cleanly.**
153
+
154
+ ```bash
155
+ SID=smoke-reset
156
+ curl -s http://127.0.0.1:18796/v1/chat/completions \
157
+ -H "Authorization: Bearer $TOKEN" -H "X-Session-Id: $SID" -H "Content-Type: application/json" \
158
+ -d '{"model":"claude-opus-4-6","messages":[{"role":"user","content":"remember the word banana"}]}' > /dev/null
159
+ curl -s http://127.0.0.1:18796/v1/chat/completions \
160
+ -H "Authorization: Bearer $TOKEN" -H "X-Session-Id: $SID" -H "X-Session-Reset: 1" -H "Content-Type: application/json" \
161
+ -d '{"model":"claude-opus-4-6","messages":[{"role":"user","content":"what word did I just tell you"}]}' \
162
+ | jq -r '.choices[0].message.content'
163
+ # Expected: model says it has no prior context.
164
+ ```
165
+
166
+ **4. `cached_tokens` grows turn-over-turn (the success metric).**
167
+
168
+ ```bash
169
+ SID=smoke-cache
170
+ PREAMBLE=$(printf 'x%.0s' {1..3000})
171
+ for i in 1 2 3 4; do
172
+ curl -s http://127.0.0.1:18796/v1/chat/completions \
173
+ -H "Authorization: Bearer $TOKEN" -H "X-Session-Id: $SID" -H "Content-Type: application/json" \
174
+ -d "{\"model\":\"claude-opus-4-6\",\"messages\":[{\"role\":\"system\",\"content\":\"long preamble: $PREAMBLE\"},{\"role\":\"user\",\"content\":\"turn $i\"}]}" > /dev/null
175
+ curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" \
176
+ | jq ".data[] | select(.session_name == \"openai-$SID\") | {turn: $i, cached_tokens, tokens_in}"
177
+ done
178
+ # Expected: cached_tokens climbs substantially by turn 3-4. If it stays at 0,
179
+ # the persistent CLI is still being killed every turn — regression.
180
+ ```
181
+
182
+ ## Error responses
183
+
184
+ Errors use the OpenAI error envelope:
185
+
186
+ ```json
187
+ { "error": { "message": "...", "type": "invalid_request_error" } }
188
+ ```
189
+
190
+ | Status | When |
191
+ |---|---|
192
+ | 400 | `messages` empty/missing, no user message, invalid `max_tokens` |
193
+ | 401 | Missing or wrong bearer token (when auth enabled) |
194
+ | 415 | POST without `Content-Type: application/json` |
195
+ | 429 | Rate limited (`OPENCLAW_RATE_LIMIT` exceeded) |
196
+ | 503 | Failed to start a new session (model unavailable, CLI crashed at boot) |
197
+ | 500 | Mid-turn failure |
198
+
199
+ ## Related
200
+
201
+ - [getting-started.md](./getting-started.md) — install + auth setup
202
+ - [sessions.md](./sessions.md) — what a session is and how the lifecycle works under the hood
203
+ - [tools.md](./tools.md) — the full plugin tool surface (council, ultraplan, etc.)