@enderfga/claw-orchestrator 5.0.0 → 6.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (158) hide show
  1. package/README.md +28 -28
  2. package/dist/bin/cli.js +107 -1
  3. package/dist/bin/cli.js.map +1 -1
  4. package/dist/src/acp-server.d.ts +5 -5
  5. package/dist/src/acp-server.js +3 -3
  6. package/dist/src/acp-server.js.map +1 -1
  7. package/dist/src/autoloop/dispatcher.d.ts +22 -0
  8. package/dist/src/autoloop/dispatcher.js +71 -13
  9. package/dist/src/autoloop/dispatcher.js.map +1 -1
  10. package/dist/src/autoloop/messages.d.ts +10 -0
  11. package/dist/src/autoloop/messages.js.map +1 -1
  12. package/dist/src/autoloop/runner.js +6 -0
  13. package/dist/src/autoloop/runner.js.map +1 -1
  14. package/dist/src/constants.d.ts +0 -6
  15. package/dist/src/constants.js +0 -6
  16. package/dist/src/constants.js.map +1 -1
  17. package/dist/src/council.d.ts +15 -0
  18. package/dist/src/council.js +48 -35
  19. package/dist/src/council.js.map +1 -1
  20. package/dist/src/dashboard/index.html +191 -6
  21. package/dist/src/embedded-server.js +132 -9
  22. package/dist/src/embedded-server.js.map +1 -1
  23. package/dist/src/fanout.d.ts +30 -1
  24. package/dist/src/fanout.js +32 -3
  25. package/dist/src/fanout.js.map +1 -1
  26. package/dist/src/index.d.ts +1 -0
  27. package/dist/src/index.js +360 -4
  28. package/dist/src/index.js.map +1 -1
  29. package/dist/src/kernel/agent-step.d.ts +59 -0
  30. package/dist/src/kernel/agent-step.js +100 -0
  31. package/dist/src/kernel/agent-step.js.map +1 -0
  32. package/dist/src/kernel/conditions.d.ts +11 -0
  33. package/dist/src/kernel/conditions.js +24 -0
  34. package/dist/src/kernel/conditions.js.map +1 -0
  35. package/dist/src/kernel/engine.d.ts +319 -0
  36. package/dist/src/kernel/engine.js +1047 -0
  37. package/dist/src/kernel/engine.js.map +1 -0
  38. package/dist/src/kernel/exec.d.ts +43 -0
  39. package/dist/src/kernel/exec.js +112 -0
  40. package/dist/src/kernel/exec.js.map +1 -0
  41. package/dist/src/kernel/file-lock.d.ts +50 -0
  42. package/dist/src/kernel/file-lock.js +135 -0
  43. package/dist/src/kernel/file-lock.js.map +1 -0
  44. package/dist/src/kernel/nodes/agent.d.ts +4 -0
  45. package/dist/src/kernel/nodes/agent.js +35 -0
  46. package/dist/src/kernel/nodes/agent.js.map +1 -0
  47. package/dist/src/kernel/nodes/autoloop.d.ts +78 -0
  48. package/dist/src/kernel/nodes/autoloop.js +75 -0
  49. package/dist/src/kernel/nodes/autoloop.js.map +1 -0
  50. package/dist/src/kernel/nodes/council.d.ts +12 -0
  51. package/dist/src/kernel/nodes/council.js +88 -0
  52. package/dist/src/kernel/nodes/council.js.map +1 -0
  53. package/dist/src/kernel/nodes/fanout.d.ts +11 -0
  54. package/dist/src/kernel/nodes/fanout.js +63 -0
  55. package/dist/src/kernel/nodes/fanout.js.map +1 -0
  56. package/dist/src/kernel/nodes/human-gate.d.ts +4 -0
  57. package/dist/src/kernel/nodes/human-gate.js +7 -0
  58. package/dist/src/kernel/nodes/human-gate.js.map +1 -0
  59. package/dist/src/kernel/nodes/index.d.ts +12 -0
  60. package/dist/src/kernel/nodes/index.js +21 -0
  61. package/dist/src/kernel/nodes/index.js.map +1 -0
  62. package/dist/src/kernel/nodes/router.d.ts +4 -0
  63. package/dist/src/kernel/nodes/router.js +12 -0
  64. package/dist/src/kernel/nodes/router.js.map +1 -0
  65. package/dist/src/kernel/nodes/subflow.d.ts +13 -0
  66. package/dist/src/kernel/nodes/subflow.js +38 -0
  67. package/dist/src/kernel/nodes/subflow.js.map +1 -0
  68. package/dist/src/kernel/nodes/ultraapp.d.ts +60 -0
  69. package/dist/src/kernel/nodes/ultraapp.js +62 -0
  70. package/dist/src/kernel/nodes/ultraapp.js.map +1 -0
  71. package/dist/src/kernel/nodes/verifier.d.ts +14 -0
  72. package/dist/src/kernel/nodes/verifier.js +84 -0
  73. package/dist/src/kernel/nodes/verifier.js.map +1 -0
  74. package/dist/src/kernel/projections.d.ts +42 -0
  75. package/dist/src/kernel/projections.js +133 -0
  76. package/dist/src/kernel/projections.js.map +1 -0
  77. package/dist/src/kernel/repo.d.ts +13 -0
  78. package/dist/src/kernel/repo.js +64 -0
  79. package/dist/src/kernel/repo.js.map +1 -0
  80. package/dist/src/kernel/secrets.d.ts +25 -0
  81. package/dist/src/kernel/secrets.js +48 -0
  82. package/dist/src/kernel/secrets.js.map +1 -0
  83. package/dist/src/kernel/store.d.ts +225 -0
  84. package/dist/src/kernel/store.js +838 -0
  85. package/dist/src/kernel/store.js.map +1 -0
  86. package/dist/src/kernel/templates/index.d.ts +140 -0
  87. package/dist/src/kernel/templates/index.js +266 -0
  88. package/dist/src/kernel/templates/index.js.map +1 -0
  89. package/dist/src/kernel/types.d.ts +326 -0
  90. package/dist/src/kernel/types.js +19 -0
  91. package/dist/src/kernel/types.js.map +1 -0
  92. package/dist/src/models.d.ts +1 -1
  93. package/dist/src/models.js +31 -3
  94. package/dist/src/models.js.map +1 -1
  95. package/dist/src/persistent-cursor-session.js +6 -1
  96. package/dist/src/persistent-cursor-session.js.map +1 -1
  97. package/dist/src/persistent-grok-session.d.ts +40 -0
  98. package/dist/src/persistent-grok-session.js +197 -0
  99. package/dist/src/persistent-grok-session.js.map +1 -0
  100. package/dist/src/run-ledger.d.ts +57 -3
  101. package/dist/src/run-ledger.js +45 -2
  102. package/dist/src/run-ledger.js.map +1 -1
  103. package/dist/src/session-manager.d.ts +176 -129
  104. package/dist/src/session-manager.js +657 -603
  105. package/dist/src/session-manager.js.map +1 -1
  106. package/dist/src/types.d.ts +37 -4
  107. package/dist/src/types.js +15 -1
  108. package/dist/src/types.js.map +1 -1
  109. package/dist/src/ultraapp/build.d.ts +117 -3
  110. package/dist/src/ultraapp/build.js +319 -3
  111. package/dist/src/ultraapp/build.js.map +1 -1
  112. package/dist/src/ultraapp/contract.d.ts +52 -0
  113. package/dist/src/ultraapp/contract.js +83 -0
  114. package/dist/src/ultraapp/contract.js.map +1 -0
  115. package/dist/src/ultraapp/conventions.js +9 -2
  116. package/dist/src/ultraapp/conventions.js.map +1 -1
  117. package/dist/src/ultraapp/fix-on-failure.d.ts +21 -2
  118. package/dist/src/ultraapp/fix-on-failure.js +46 -62
  119. package/dist/src/ultraapp/fix-on-failure.js.map +1 -1
  120. package/dist/src/ultraapp/manager.d.ts +107 -2
  121. package/dist/src/ultraapp/manager.js +305 -86
  122. package/dist/src/ultraapp/manager.js.map +1 -1
  123. package/dist/src/verify/baseline.d.ts +73 -0
  124. package/dist/src/verify/baseline.js +186 -0
  125. package/dist/src/verify/baseline.js.map +1 -0
  126. package/dist/src/verify/contract.d.ts +116 -0
  127. package/dist/src/verify/contract.js +142 -0
  128. package/dist/src/verify/contract.js.map +1 -0
  129. package/dist/src/verify/evidence.d.ts +61 -0
  130. package/dist/src/verify/evidence.js +133 -0
  131. package/dist/src/verify/evidence.js.map +1 -0
  132. package/dist/src/verify/runner.d.ts +63 -0
  133. package/dist/src/verify/runner.js +317 -0
  134. package/dist/src/verify/runner.js.map +1 -0
  135. package/openclaw.plugin.json +8 -0
  136. package/package.json +2 -2
  137. package/skills/SKILL.md +121 -80
  138. package/skills/references/acp.md +18 -18
  139. package/skills/references/autoloop.md +148 -72
  140. package/skills/references/claude-cli-tracking.md +4 -4
  141. package/skills/references/cli.md +103 -60
  142. package/skills/references/council.md +109 -37
  143. package/skills/references/dashboard.md +34 -6
  144. package/skills/references/getting-started.md +14 -14
  145. package/skills/references/inbox.md +4 -4
  146. package/skills/references/mcp.md +39 -34
  147. package/skills/references/multi-engine.md +109 -51
  148. package/skills/references/observability.md +88 -27
  149. package/skills/references/openai-compat.md +40 -40
  150. package/skills/references/sessions.md +44 -26
  151. package/skills/references/tools.md +402 -309
  152. package/skills/references/ultra.md +45 -45
  153. package/skills/references/ultraapp.md +126 -50
  154. package/skills/references/verification.md +187 -0
  155. package/skills/references/workflow.md +362 -0
  156. package/dist/src/ultraapp/fix-on-failure-session.d.ts +0 -23
  157. package/dist/src/ultraapp/fix-on-failure-session.js +0 -51
  158. package/dist/src/ultraapp/fix-on-failure-session.js.map +0 -1
@@ -14,7 +14,9 @@ SessionManager
14
14
  │ └── Wraps: codex app-server --listen stdio:// (long-running JSON-RPC; required for /goal)
15
15
  ├── engine: 'agy' → PersistentAgySession
16
16
  │ └── Wraps: agy -p (Google Antigravity CLI, per-message spawning, stream-json output)
17
- ├── engine: 'cursor' → PersistentCursorSession
17
+ ├── engine: 'grok' → PersistentGrokSession
18
+ │ └── Wraps: grok -p --output-format json (xAI Grok Build, per-message spawning)
19
+ ├── engine: 'cursor' → PersistentCursorSession (legacy)
18
20
  │ └── Wraps: agent -p --force --trust --output-format stream-json (per-message spawning)
19
21
  ├── engine: 'opencode' → PersistentOpencodeSession
20
22
  │ └── Wraps: opencode run --format json (per-message spawning)
@@ -37,6 +39,7 @@ Default engine. Long-running subprocess with streaming JSON I/O. Tested with Cla
37
39
  - Fork subagent (`forkSubagent`), tool search (`enableToolSearch`), OpenTelemetry logging toggles (`otelLogUserPrompts`, `otelLogRawApiBodies`), `xhigh` effort tier (Opus 4.7), and `stats.pluginErrors` capture — see [CLI 2.1.121 options in SKILL.md](../SKILL.md) and [tools.md](./tools.md)
38
40
 
39
41
  > **Behavior changes from upstream Claude CLI 2.1.121** (worth knowing if you set permission rules):
42
+ >
40
43
  > - `--agent` / `--print` now enforce agent frontmatter `permissionMode`, `tools`, `disallowedTools` (was advisory). Affects `council` agent personas.
41
44
  > - `Bash(find:*)` permission rule no longer auto-approves `find -exec` or `find -delete`. Add explicit rules if you depend on these.
42
45
  > - `--dangerously-skip-permissions` also skips prompts for `.claude/skills/` directory. Treat with care.
@@ -45,7 +48,7 @@ Default engine. Long-running subprocess with streaming JSON I/O. Tested with Cla
45
48
  ```typescript
46
49
  await manager.startSession({
47
50
  name: 'claude-task',
48
- engine: 'claude', // default, can omit
51
+ engine: 'claude', // default, can omit
49
52
  model: 'opus',
50
53
  cwd: '/project',
51
54
  });
@@ -138,6 +141,7 @@ in print mode. Verified against `agy` **1.1.13**.
138
141
  `gemini-3.1-pro has no "medium" effort (available: low, high)`. The adapter
139
142
  passes the requested effort through rather than substituting a tier the caller
140
143
  did not ask for; run `agy models` to see the tiers a slug actually exposes.
144
+
141
145
  - Permission modes: `bypassPermissions` → `--dangerously-skip-permissions`,
142
146
  `default` → `--sandbox` (terminal-restricted), and
143
147
  `sandboxMode: 'read-only'` → `--mode plan` (takes precedence). Other modes
@@ -173,9 +177,59 @@ await manager.startSession({
173
177
  > multi-model **proxy** still talks to the Gemini **API**; that is a different
174
178
  > subsystem and is unaffected.)
175
179
 
176
- ### Cursor Agent (`engine: 'cursor'`)
180
+ ### Grok Build (`engine: 'grok'`)
181
+
182
+ Wraps xAI's **Grok Build** CLI. Each `send()` spawns `grok -p <msg> --output-format json`, which
183
+ prints a single JSON object and exits. Verified against `grok` **1.0.5**.
184
+
185
+ - **Cost comes from the engine, not from our price table.** The result object carries
186
+ `total_cost_usd`, and the wrapper writes it straight into the session's spend. Every other engine
187
+ here multiplies tokens by a rate in `models.ts` — the metadata most prone to going stale — so on
188
+ this engine the run ledger and the `maxBudgetUsd` gate both read what xAI actually charged.
189
+ `grok-4.6` is still registered, for its context window and an indicative breakdown; its two price
190
+ tiers ($2/$0.50/$6 under a 200K prompt, $4/$1/$12 at or above, charged across the whole request)
191
+ therefore never have to be modelled here.
192
+ - **Real conversation continuity**: the `sessionId` from turn 1 is replayed as `--resume <id>`.
193
+ `--continue` is deliberately not used — it means "the most recent session for this cwd", which
194
+ collides between concurrent sessions. Confirmed with a two-turn recall test, not inferred.
195
+ - Real token counts from `usage` (`input_tokens`, `output_tokens`, `cache_read_input_tokens`).
196
+ These are **per-turn**, not cumulative over the thread — checked by resuming and reading turn 2,
197
+ because the same-looking field on codex is a running total.
198
+ - Permission modes pass straight through: grok's `--permission-mode` takes the same vocabulary we
199
+ use. The one exception is our `manual`, which grok spells `default`.
200
+ - Reasoning effort maps to `--effort`; grok accepts `low|medium|high`, so `max` and `xhigh` clamp.
201
+ - **`sandboxMode: 'read-only'` is refused, not approximated.** grok has `--permission-mode plan` and
202
+ `--deny` rules, but plan mode alone is model-cooperative — the shape that let an adversarial
203
+ prompt write through Cursor's plan mode — and the deny rules have not been through the
204
+ write × shell × subagent × resumed-turn matrix this project requires before claiming a boundary.
205
+ A read-only grok session throws rather than running writable under a read-only label.
206
+ - Binary: `grok` (set `GROK_BIN` to override). Not `agent`: xAI's installer claims that name too,
207
+ and so did Cursor's.
208
+ - Requires Grok Build: see `x.ai/cli`.
177
209
 
178
- Wraps the Cursor Agent CLI (`agent`) with `--print --output-format stream-json`. Write-enabled sessions use `--force`. Each `send()` spawns a new process.
210
+ ```typescript
211
+ await manager.startSession({
212
+ name: 'grok-task',
213
+ engine: 'grok',
214
+ model: 'grok-4.6',
215
+ cwd: '/project',
216
+ });
217
+ ```
218
+
219
+ ### Cursor Agent (`engine: 'cursor'`) — legacy
220
+
221
+ > **Legacy: `engine: 'cursor'`.** Superseded in this lineup by Grok Build (`engine: 'grok'`).
222
+ > The `cursor` engine still exists and still works — existing callers are not broken — but it is
223
+ > no longer a documented option, is not version-tracked, and gets no new work.
224
+ >
225
+ > Note what this is and is not: Cursor itself is **not** discontinued. Anysphere was acquired by
226
+ > SpaceX (closed 2026-08-15) and folded into the SpaceXAI team, and the CLI has shipped since. Two
227
+ > practical things pushed it out of the tracked set. Cursor never reports which model actually ran
228
+ > — its `system` init event says `"model": "Auto"` — so a router that spans Claude, GPT and Grok
229
+ > leaves every cost row attributed to a hardcoded proxy rate. And xAI's Grok installer now claims
230
+ > the bare `agent` name, so the binary that name resolves to depends on install order.
231
+
232
+ Wraps the Cursor Agent CLI with `--print --output-format stream-json`. Write-enabled sessions use `--force`. Each `send()` spawns a new process.
179
233
 
180
234
  - Conversation continuity: the chat id from the first turn's `system` event is captured and passed back as `--resume <chatId>` on later sends, so the model sees prior turns. `--continue` is deliberately not used: it resumes "the latest chat", which collides between concurrent sessions.
181
235
  - One-shot execution per message (no persistent subprocess)
@@ -185,7 +239,9 @@ Wraps the Cursor Agent CLI (`agent`) with `--print --output-format stream-json`.
185
239
  - `--trust` auto-trusts the workspace without prompting
186
240
  - Cursor uses its own model routing (e.g., `sonnet-4`, `gpt-5`, `auto`)
187
241
  - Requires Cursor Agent CLI: `curl https://cursor.com/install -fsSL | bash`
188
- - Binary: `agent` (set `CURSOR_BIN` env var to override)
242
+ - Binary: `cursor-agent` (set `CURSOR_BIN` env var to override). The generic `agent` name is
243
+ deliberately not used: xAI's Grok installer symlinks `agent` to its own binary, which rejects
244
+ `--force`/`--trust`/`--workspace` and fails the turn with "unexpected argument"
189
245
 
190
246
  ```typescript
191
247
  await manager.startSession({
@@ -271,12 +327,12 @@ is not the answer on every engine. See "Stats & Monitoring" in `sessions.md`.
271
327
 
272
328
  Team tools (`team_list`, `team_send`) operate on the same virtual-team layer for **every** engine: the "team" is the set of all active sessions managed by SessionManager.
273
329
 
274
- | Engine | `team_list` | `team_send` |
275
- |--------|------------|-------------|
276
- | Claude | Lists other active SessionManager sessions | Routes via cross-session inbox |
277
- | Codex | Lists other active SessionManager sessions | Routes via cross-session inbox |
330
+ | Engine | `team_list` | `team_send` |
331
+ | ----------- | ------------------------------------------ | ------------------------------ |
332
+ | Claude | Lists other active SessionManager sessions | Routes via cross-session inbox |
333
+ | Codex | Lists other active SessionManager sessions | Routes via cross-session inbox |
278
334
  | Antigravity | Lists other active SessionManager sessions | Routes via cross-session inbox |
279
- | Cursor | Lists other active SessionManager sessions | Routes via cross-session inbox |
335
+ | Cursor | Lists other active SessionManager sessions | Routes via cross-session inbox |
280
336
 
281
337
  Messages are delivered via the inbox system — idle sessions receive immediately, busy sessions queue for later delivery.
282
338
 
@@ -295,12 +351,13 @@ If OpenClaw gateway is running, everything is automatic:
295
351
  await manager.startSession({
296
352
  name: 'task',
297
353
  engine: 'claude',
298
- model: 'openclaw', // gateway routes to your configured model
354
+ model: 'openclaw', // gateway routes to your configured model
299
355
  cwd: '/project',
300
356
  });
301
357
  ```
302
358
 
303
359
  What happens behind the scenes:
360
+
304
361
  1. Plugin reads `~/.openclaw/openclaw.json` for gateway port + auth
305
362
  2. Starts a local proxy server (random port, auto-managed)
306
363
  3. Claude Code CLI sends Anthropic-format requests → proxy converts to OpenAI → gateway → any model
@@ -309,12 +366,12 @@ What happens behind the scenes:
309
366
 
310
367
  Override with environment variables if needed:
311
368
 
312
- | Variable | Default | Description |
313
- |----------|---------|-------------|
314
- | `GATEWAY_URL` | Auto-detected from openclaw.json | Gateway endpoint (e.g. `http://127.0.0.1:18789/v1`) |
315
- | `GATEWAY_KEY` | Auto-detected from openclaw.json | Gateway auth password/token |
316
- | `GEMINI_API_KEY` | - | Direct Gemini API access (bypasses gateway) |
317
- | `OPENAI_API_KEY` | - | Direct OpenAI API access (bypasses gateway) |
369
+ | Variable | Default | Description |
370
+ | ---------------- | -------------------------------- | --------------------------------------------------- |
371
+ | `GATEWAY_URL` | Auto-detected from openclaw.json | Gateway endpoint (e.g. `http://127.0.0.1:18789/v1`) |
372
+ | `GATEWAY_KEY` | Auto-detected from openclaw.json | Gateway auth password/token |
373
+ | `GEMINI_API_KEY` | - | Direct Gemini API access (bypasses gateway) |
374
+ | `OPENAI_API_KEY` | - | Direct OpenAI API access (bypasses gateway) |
318
375
 
319
376
  ### Architecture
320
377
 
@@ -330,46 +387,47 @@ Claude Code CLI (Anthropic format)
330
387
  Integrate **any** coding agent CLI without writing engine-specific code. You provide a `CustomEngineConfig` that maps your CLI's flags to OpenClaw session concepts.
331
388
 
332
389
  Two protocol modes:
390
+
333
391
  - **Persistent** (`persistent: true`) — long-running subprocess with stream-json I/O over stdin/stdout (like Claude Code)
334
392
  - **One-shot** (`persistent: false`, default) — new process spawned per `send()` (like Codex/Antigravity)
335
393
 
336
394
  ### CustomEngineConfig
337
395
 
338
- | Field | Type | Required | Description |
339
- |-------|------|----------|-------------|
340
- | `name` | string | yes | Display name (used in logs, session IDs) |
341
- | `bin` | string | yes | Binary path or command name |
342
- | `binEnv` | string | | Env var name that overrides `bin` at runtime |
343
- | `persistent` | boolean | | `true` = persistent subprocess, `false` = one-shot (default) |
344
- | `args` | object | yes | CLI flag mappings (see below) |
345
- | `permissionModes` | object | | Maps OpenClaw mode names to CLI-specific values |
346
- | `pricing` | object | | `{ input, output, cached? }` per 1M tokens |
347
- | `contextWindow` | number | | Context window size (default: 200,000) |
348
- | `env` | object | | Extra environment variables for the CLI process |
349
- | `sanitizePatterns` | string[] | | Regex patterns to redact from stderr |
396
+ | Field | Type | Required | Description |
397
+ | ------------------ | -------- | -------- | ------------------------------------------------------------ |
398
+ | `name` | string | yes | Display name (used in logs, session IDs) |
399
+ | `bin` | string | yes | Binary path or command name |
400
+ | `binEnv` | string | | Env var name that overrides `bin` at runtime |
401
+ | `persistent` | boolean | | `true` = persistent subprocess, `false` = one-shot (default) |
402
+ | `args` | object | yes | CLI flag mappings (see below) |
403
+ | `permissionModes` | object | | Maps OpenClaw mode names to CLI-specific values |
404
+ | `pricing` | object | | `{ input, output, cached? }` per 1M tokens |
405
+ | `contextWindow` | number | | Context window size (default: 200,000) |
406
+ | `env` | object | | Extra environment variables for the CLI process |
407
+ | `sanitizePatterns` | string[] | | Regex patterns to redact from stderr |
350
408
 
351
409
  ### args field
352
410
 
353
- | Key | Example | Description |
354
- |-----|---------|-------------|
355
- | `print` | `"-p"` | Non-interactive/print mode flag |
356
- | `outputFormat` | `"--output-format"` | Output format flag |
357
- | `outputFormatValue` | `"stream-json"` | Value for stream-json output |
358
- | `inputFormat` | `"--input-format"` | Input format flag (persistent only) |
359
- | `inputFormatValue` | `"stream-json"` | Value for stream-json input |
360
- | `skipPermissions` | `"-y"` | Skip all permissions flag |
361
- | `permissionMode` | `"--permission-mode"` | Permission mode flag |
362
- | `model` | `"--model"` | Model selection flag |
363
- | `systemPrompt` | `"--system-prompt"` | System prompt override flag |
364
- | `appendSystemPrompt` | `"--append-system-prompt"` | Append system prompt flag |
365
- | `maxTurns` | `"--max-turns"` | Max agent turns flag |
366
- | `resume` | `"--resume"` | Session resume flag (persistent only) |
367
- | `verbose` | `"--verbose"` | Verbose output flag |
368
- | `replayUserMessages` | `"--replay-user-messages"` | Replay user messages (persistent only) |
411
+ | Key | Example | Description |
412
+ | ------------------------ | ------------------------------ | ------------------------------------------ |
413
+ | `print` | `"-p"` | Non-interactive/print mode flag |
414
+ | `outputFormat` | `"--output-format"` | Output format flag |
415
+ | `outputFormatValue` | `"stream-json"` | Value for stream-json output |
416
+ | `inputFormat` | `"--input-format"` | Input format flag (persistent only) |
417
+ | `inputFormatValue` | `"stream-json"` | Value for stream-json input |
418
+ | `skipPermissions` | `"-y"` | Skip all permissions flag |
419
+ | `permissionMode` | `"--permission-mode"` | Permission mode flag |
420
+ | `model` | `"--model"` | Model selection flag |
421
+ | `systemPrompt` | `"--system-prompt"` | System prompt override flag |
422
+ | `appendSystemPrompt` | `"--append-system-prompt"` | Append system prompt flag |
423
+ | `maxTurns` | `"--max-turns"` | Max agent turns flag |
424
+ | `resume` | `"--resume"` | Session resume flag (persistent only) |
425
+ | `verbose` | `"--verbose"` | Verbose output flag |
426
+ | `replayUserMessages` | `"--replay-user-messages"` | Replay user messages (persistent only) |
369
427
  | `includePartialMessages` | `"--include-partial-messages"` | Include partial messages (persistent only) |
370
- | `effort` | `"--effort"` | Effort level flag |
371
- | `workspace` | `"--workspace"` | Workspace/cwd flag (one-shot only) |
372
- | `extra` | `["--trust"]` | Additional static arguments |
428
+ | `effort` | `"--effort"` | Effort level flag |
429
+ | `workspace` | `"--workspace"` | Workspace/cwd flag (one-shot only) |
430
+ | `extra` | `["--trust"]` | Additional static arguments |
373
431
 
374
432
  ### Example: Persistent mode (Claude Code-compatible CLI)
375
433
 
@@ -417,7 +475,7 @@ await manager.startSession({
417
475
  customEngine: {
418
476
  name: 'simple-agent',
419
477
  bin: '/usr/local/bin/simple-agent',
420
- persistent: false, // default
478
+ persistent: false, // default
421
479
  args: {
422
480
  print: '-p',
423
481
  outputFormat: '--output-format',
@@ -451,11 +509,11 @@ await manager.startSession({
451
509
  dangerouslySkipPermissions: true,
452
510
  customEngine: {
453
511
  name: 'antigravity',
454
- bin: 'agy', // install: curl -fsSL https://antigravity.google/cli/install.sh | bash
512
+ bin: 'agy', // install: curl -fsSL https://antigravity.google/cli/install.sh | bash
455
513
  binEnv: 'AGY_BIN',
456
514
  persistent: false,
457
515
  args: {
458
- print: '-p', // single-prompt headless mode
516
+ print: '-p', // single-prompt headless mode
459
517
  skipPermissions: '--dangerously-skip-permissions',
460
518
  workspace: '--add-dir',
461
519
  // NOTE: agy 1.0.2 has NO --output-format flag — output is plain text only.
@@ -30,22 +30,22 @@ every engine passes through.
30
30
 
31
31
  ### Row schema
32
32
 
33
- | Field | Meaning |
34
- |---|---|
35
- | `ts` | ISO timestamp of turn completion |
36
- | `session` | SessionManager session name |
37
- | `engine` | `claude` / `codex` / `codex-app` / `cursor` / `opencode` / `agy` / `custom` |
38
- | `model` | Configured model, or the engine's own reported model when none was set |
39
- | `cwd` | Working directory the turn ran in |
40
- | `turn` | 1-based turn index within the session |
41
- | `tokensIn` / `tokensOut` / `cachedTokens` | **Per-turn deltas**, not session totals |
42
- | `costUsd` | Per-turn delta in USD |
43
- | `tokensEstimated` | `true` when the counts came from `estimateTokens()` (see below) |
44
- | `durationMs` | Wall-clock for the turn |
45
- | `toolCalls` / `toolErrors` | Per-turn deltas |
46
- | `ok` | `false` for a turn that threw, or that the session's own `turnsSucceeded` counter did not count (see `sessions.md`). Falls back to "nothing was thrown" when the counter cannot be read |
47
- | `error` | Failure text, truncated to 500 chars. Absent when the turn resolved but the engine did not count it as succeeded (an interrupted or non-SUCCESS turn), so a failed row does not always carry one |
48
- | `parent` | council id / fanout id / autoloop run id, when the turn belongs to one |
33
+ | Field | Meaning |
34
+ | ----------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
35
+ | `ts` | ISO timestamp of turn completion |
36
+ | `session` | SessionManager session name |
37
+ | `engine` | `claude` / `codex` / `codex-app` / `grok` / `opencode` / `agy` / `custom` |
38
+ | `model` | Configured model, or the engine's own reported model when none was set |
39
+ | `cwd` | Working directory the turn ran in |
40
+ | `turn` | 1-based turn index within the session |
41
+ | `tokensIn` / `tokensOut` / `cachedTokens` | **Per-turn deltas**, not session totals |
42
+ | `costUsd` | Per-turn delta in USD |
43
+ | `tokensEstimated` | `true` when the counts came from `estimateTokens()` (see below) |
44
+ | `durationMs` | Wall-clock for the turn |
45
+ | `toolCalls` / `toolErrors` | Per-turn deltas |
46
+ | `ok` | `false` for a turn that threw, or that the session's own `turnsSucceeded` counter did not count (see `sessions.md`). Falls back to "nothing was thrown" when the counter cannot be read |
47
+ | `error` | Failure text, truncated to 500 chars. Absent when the turn resolved but the engine did not count it as succeeded (an interrupted or non-SUCCESS turn), so a failed row does not always carry one |
48
+ | `parent` | council id / fanout id / autoloop run id, when the turn belongs to one |
49
49
 
50
50
  Deltas rather than totals means summing a query window gives that window's spend
51
51
  without double-counting.
@@ -98,11 +98,11 @@ Notes:
98
98
  - The check is "has the cap been reached", not "would this turn exceed it" — a
99
99
  turn's cost is unknown until it finishes, so the last allowed turn can overshoot.
100
100
  Size the cap accordingly.
101
- - A cap of `0` or a negative number means *unset*, not *refuse everything*.
101
+ - A cap of `0` or a negative number means _unset_, not _refuse everything_.
102
102
  - Claude Code still receives `--max-budget-usd` as well: an in-CLI stop happens
103
103
  earlier and therefore costs less than an after-the-fact refusal.
104
104
  - `session_list` / `GET /session/list` expose `costUsd`, `budgetUsd` and
105
- `budgetExhausted` so a stalled session shows *why* it stopped taking turns.
105
+ `budgetExhausted` so a stalled session shows _why_ it stopped taking turns.
106
106
 
107
107
  ## Accuracy: which engines report real usage
108
108
 
@@ -111,15 +111,16 @@ Where the engine reports usage, those counts are the engine's own. Where it does
111
111
  not, the wrapper falls back to `estimateTokens()` (characters ÷ 4) and the row is
112
112
  flagged `tokensEstimated: true`; the CLI marks those costs with a trailing `~`.
113
113
 
114
- | Engine | Token counts |
115
- |---|---|
116
- | `claude` | Engine-reported |
117
- | `codex` | Engine-reported |
118
- | `codex-app` | Engine-reported |
119
- | `cursor` | Engine-reported when the stream carries `usage`, else estimated |
120
- | `opencode` | Engine-reported when the run JSON carries `tokens`, else estimated |
121
- | `agy` | Engine-reported when the result event carries usage, else estimated |
122
- | `custom` | Depends on the CLI; estimated when it emits no usage |
114
+ | Engine | Token counts |
115
+ | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
116
+ | `claude` | Engine-reported |
117
+ | `codex` | Engine-reported |
118
+ | `codex-app` | Engine-reported |
119
+ | `grok` | Engine-reported — and so is the **cost**: this engine reports `total_cost_usd`, which the wrapper passes through instead of pricing tokens from the registry, so registry drift cannot affect a grok row |
120
+ | `cursor` (legacy) | Engine-reported when the stream carries `usage`, else estimated |
121
+ | `opencode` | Engine-reported when the run JSON carries `tokens`, else estimated |
122
+ | `agy` | Engine-reported when the result event carries usage, else estimated |
123
+ | `custom` | Depends on the CLI; estimated when it emits no usage |
123
124
 
124
125
  So on an estimating engine the cap is best-effort. It will stop a runaway session;
125
126
  it is not an accounting guarantee, and it is not a substitute for the spend limits
@@ -129,3 +130,63 @@ Cost figures are also only as good as the pricing table: a model missing from
129
130
  `models.ts` prices at its family default, and subscription plans (Claude Max,
130
131
  ChatGPT Pro) bill nothing per token while the ledger still reports the API-rate
131
132
  equivalent. Read `costUsd` as "what this would cost at API rates".
133
+
134
+ ## `ok` vs `verified` (6.0.0)
135
+
136
+ A row now carries two different judgements, and conflating them is the mistake
137
+ this section exists to prevent.
138
+
139
+ - **`ok`** — the engine's own terminal verdict for that turn. Codex fails a turn
140
+ that emits `turn.failed` while exiting 0; gemini succeeds on exit 53. It is a
141
+ careful signal, but it is the engine talking about itself.
142
+ - **`verified`** — an acceptance contract ran against the work and every required
143
+ check passed. That is the runtime's own measurement.
144
+
145
+ Three states, not two:
146
+
147
+ | `verified` | Means | CLI column |
148
+ | ---------- | --------------------------------------------------- | ---------- |
149
+ | `true` | A contract ran and passed | `yes` |
150
+ | `false` | A contract ran and a required check failed | `NO` |
151
+ | absent | **No contract was declared. Nothing checked this.** | `—` |
152
+
153
+ Absent is not false. An unchecked run is not a failed one, and reading it as
154
+ either would make the ledger useless for the thing it is for.
155
+
156
+ ### Where the verdict comes from
157
+
158
+ `verified` is **not written at turn time**, deliberately. The turns that produce
159
+ the work all finish before the verifier that judges it, so stamping a verdict on
160
+ them as they are written would be inventing one. It is joined in at read time
161
+ from the run record via the row's `parent`, by `annotateVerdicts()`.
162
+
163
+ Two consequences worth knowing:
164
+
165
+ - A raw `runs/*.jsonl` line usually has no `verified` field. Read through
166
+ `clawo runs` / `GET /runs` / `getRunLedger()` to get the join.
167
+ - Filtering on `--verified` happens _after_ the join. Pushing the filter into the
168
+ ledger read would match on a field no row carries yet and return nothing.
169
+
170
+ ### Other new row fields
171
+
172
+ | Field | Source |
173
+ | -------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
174
+ | `evidenceId`, `contractId` | Joined from the run record |
175
+ | `nodeKind` | The kernel node the turn belonged to (`agent`, `council`, `verifier`, …) |
176
+ | `repoLang` | Detected from a manifest (`package.json`, `pyproject.toml`, `go.mod`, …). Never guessed — a wrong label would corrupt the comparison the field exists to enable |
177
+ | `taskKind` | Caller-declared only. Never inferred from the prompt |
178
+
179
+ All are optional and absent on rows written before 6.0.0. The reader already
180
+ skips unknown and missing keys, so old shards stay readable; nothing backfills
181
+ them, because we cannot know retroactively.
182
+
183
+ ```bash
184
+ clawo runs --since 7d --verified # only turns whose contract passed
185
+ clawo runs --since 7d --refuted # only turns whose contract failed
186
+ clawo runs --parent wf-abc123 # every turn of one workflow run
187
+ ```
188
+
189
+ ## Related
190
+
191
+ - [`verification.md`](./verification.md) — what a contract is and how a verdict is produced
192
+ - [`workflow.md`](./workflow.md) — where run records live
@@ -2,7 +2,7 @@
2
2
 
3
3
  > **Cost warning**: This bridge routes requests through the Claude Code CLI, which uses your Claude Max subscription's **extra usage** quota. When OpenClaw's agent loop sends its system prompt (with distinctive tool definitions and agent instructions), Anthropic's backend recognizes this as programmatic/agent traffic and bills it against extra usage — **not** the included allowance. This is by design: the bridge does NOT bypass Anthropic's billing or subscription enforcement. Using it as OpenClaw's primary model backend means every agent turn consumes extra usage credits at standard API rates ($15/M input, $75/M output for Opus). Monitor your usage at [claude.ai/settings/usage](https://claude.ai/settings/usage).
4
4
 
5
- The embedded server exposes a drop-in OpenAI-compatible endpoint so any client that speaks `/v1/chat/completions` can talk to a persistent Claude Code (or Codex / Antigravity / Cursor) session. The bridge is designed to serve **two kinds of clients as first-class citizens**:
5
+ The embedded server exposes a drop-in OpenAI-compatible endpoint so any client that speaks `/v1/chat/completions` can talk to a persistent Claude Code (or Codex / Antigravity / Grok) session. The bridge is designed to serve **two kinds of clients as first-class citizens**:
6
6
 
7
7
  1. **Upstream agents** that maintain their own conversation state and forward only the latest user turn — OpenClaw's main agent loop, cron jobs, subagents, programmatic clients.
8
8
  2. **OpenAI-compatible webchat / labeling tools** that re-send the full transcript on every turn — ChatGPT-Next-Web, Open WebUI, LobeChat, data-labeling pipelines.
@@ -11,13 +11,13 @@ Both modes share the same wire protocol; the difference is how a "new conversati
11
11
 
12
12
  ## Endpoint
13
13
 
14
- | | |
15
- |---|---|
16
- | **URL** | `http://127.0.0.1:18796/v1/chat/completions` |
17
- | **Models endpoint** | `GET /v1/models` |
18
- | **Inspection endpoint** | `GET /v1/sessions` (lists active openai-compat sessions with caching stats) |
19
- | **Auth** | Bearer token via `Authorization: Bearer $OPENCLAW_SERVER_TOKEN` (set the env var to enable; otherwise no auth and the server is loopback-only) |
20
- | **Wire format** | OpenAI Chat Completions, both streaming (SSE) and non-streaming |
14
+ | | |
15
+ | ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
16
+ | **URL** | `http://127.0.0.1:18796/v1/chat/completions` |
17
+ | **Models endpoint** | `GET /v1/models` |
18
+ | **Inspection endpoint** | `GET /v1/sessions` (lists active openai-compat sessions with caching stats) |
19
+ | **Auth** | Bearer token via `Authorization: Bearer $OPENCLAW_SERVER_TOKEN` (set the env var to enable; otherwise no auth and the server is loopback-only) |
20
+ | **Wire format** | OpenAI Chat Completions, both streaming (SSE) and non-streaming |
21
21
 
22
22
  ## Session keying
23
23
 
@@ -56,20 +56,20 @@ Without this flag, those frontends would silently continue the previous CLI sess
56
56
 
57
57
  The env var is read on every request, so ops can flip it via `launchctl setenv` (or equivalent) without restarting the server.
58
58
 
59
- | Mode | Best for | New-conversation signals |
60
- |---|---|---|
61
- | **Default** | OpenClaw main agent, cron jobs, subagents, scripted clients | `X-Session-Reset: 1` only |
59
+ | Mode | Best for | New-conversation signals |
60
+ | ----------------- | ----------------------------------------------------------- | --------------------------------------------------- |
61
+ | **Default** | OpenClaw main agent, cron jobs, subagents, scripted clients | `X-Session-Reset: 1` only |
62
62
  | **`HEURISTIC=1`** | ChatGPT-Next-Web, Open WebUI, LobeChat, data labeling tools | `X-Session-Reset: 1` **and** `[system, user]` shape |
63
63
 
64
64
  ## Status webhook
65
65
 
66
66
  When `OPENAI_COMPAT_STATUS_URL` is set (full HTTP URL), each chat completion sends best-effort `POST` requests with `Content-Type: application/json` and body:
67
67
 
68
- | Field | Type | Meaning |
69
- |---|---|---|
70
- | `state` | string | `thinking` (turn started), `working` (a tool is running), or `idle` (turn finished or stream closed). |
71
- | `activity` | string | Short human-readable line, e.g. `Processing request...`, `Reading: foo.ts`, `Running: npm test...`. |
72
- | `tool` | string \| null | Tool name when `state === working`, otherwise `null`. |
68
+ | Field | Type | Meaning |
69
+ | ---------- | -------------- | ----------------------------------------------------------------------------------------------------- |
70
+ | `state` | string | `thinking` (turn started), `working` (a tool is running), or `idle` (turn finished or stream closed). |
71
+ | `activity` | string | Short human-readable line, e.g. `Processing request...`, `Reading: foo.ts`, `Running: npm test...`. |
72
+ | `tool` | string \| null | Tool name when `state === working`, otherwise `null`. |
73
73
 
74
74
  Failures are ignored (no retries). Use this from a small local HTTP handler that forwards status into your webchat UI.
75
75
 
@@ -78,17 +78,17 @@ Failures are ignored (no retries). Use this from a small local HTTP handler that
78
78
  When the request carries `tools`, the schemas have to reach the CLI somehow. Which
79
79
  mechanism is used depends on whether the engine keeps the conversation itself.
80
80
 
81
- | Engine | Turn 1 | Later turns |
82
- |---|---|---|
83
- | `claude` | Schemas go into the session system prompt (`--system-prompt`) | Nothing injected — the system prompt persists |
84
- | `codex`, `codex-app`, `agy`, `opencode`, `cursor` | Full schema block prepended to the message | A short reminder of the calling convention, no schemas — but only once the conversation id has been captured; until then the full block is sent again |
85
- | `gemini`, one-shot `custom` | Full schema block prepended to the message | Full schema block again — these have no resume surface, so nothing persists between sends |
81
+ | Engine | Turn 1 | Later turns |
82
+ | ----------------------------------------------- | ------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |
83
+ | `claude` | Schemas go into the session system prompt (`--system-prompt`) | Nothing injected — the system prompt persists |
84
+ | `codex`, `codex-app`, `agy`, `opencode`, `grok` | Full schema block prepended to the message | A short reminder of the calling convention, no schemas — but only once the conversation id has been captured; until then the full block is sent again |
85
+ | `gemini`, one-shot `custom` | Full schema block prepended to the message | Full schema block again — these have no resume surface, so nothing persists between sends |
86
86
 
87
87
  The middle row is the one worth understanding. Those engines resume a conversation by id, so
88
88
  everything injected stays in the transcript. Re-sending the full block each turn
89
89
  grows the prompt without bound — a 54-tool block runs to roughly 17k tokens, so a
90
90
  handful of turns is enough to overflow the context window mid-loop and fail the
91
- run outright. Sending *nothing* on resume turns is not the answer either: the
91
+ run outright. Sending _nothing_ on resume turns is not the answer either: the
92
92
  block also carries the "emit a tool call, do not carry out the work yourself"
93
93
  framing, and without it the CLI starts doing the work directly.
94
94
 
@@ -116,16 +116,16 @@ does change mid-conversation.
116
116
 
117
117
  ## Environment variables
118
118
 
119
- | Variable | Default | Purpose |
120
- |---|---|---|
121
- | `OPENCLAW_SERVER_TOKEN` | (unset) | Bearer token for HTTP auth. Set to enable; written to `~/.openclaw/server-token` for the CLI. |
122
- | `OPENCLAW_RATE_LIMIT` | `300` | Max requests per IP per 60-second sliding window. |
123
- | `OPENCLAW_CORS_ORIGINS` | (loopback only) | Set to `*` to allow all origins (the `/v1/*` paths already do this). |
124
- | `OPENAI_COMPAT_NEW_CONVO_HEURISTIC` | (unset) | Set to `1` to enable webchat mode (see above). |
125
- | `OPENAI_COMPAT_TOOLS_PER_MESSAGE` | (unset) | Set to `1` to re-send the full tool schemas on every turn (see [Tool definitions](#tool-definitions-and-where-they-live)). Needed only when the tool set changes mid-conversation; costs per-turn prompt growth. |
126
- | `OPENAI_COMPAT_STATUS_URL` | (unset) | If set, the bridge POSTs JSON status updates to this URL (fire-and-forget, 2s timeout). See [Status webhook](#status-webhook). |
127
- | `OPENCLAW_SERVE_MAX_SESSIONS` | `32` | Max concurrent OpenAI-compat sessions in serve mode. Bumped from the in-plugin default of 5 because each distinct caller now gets its own `sys-<hash>` session. |
128
- | `OPENCLAW_SERVE_TTL_MINUTES` | `60` | Idle TTL for OpenAI-compat sessions in serve mode. Idle sessions are reaped by a 60s background loop; persisted disk registry is kept for 7 days so a returning caller is auto-resumed. |
119
+ | Variable | Default | Purpose |
120
+ | ----------------------------------- | --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
121
+ | `OPENCLAW_SERVER_TOKEN` | (unset) | Bearer token for HTTP auth. Set to enable; written to `~/.openclaw/server-token` for the CLI. |
122
+ | `OPENCLAW_RATE_LIMIT` | `300` | Max requests per IP per 60-second sliding window. |
123
+ | `OPENCLAW_CORS_ORIGINS` | (loopback only) | Set to `*` to allow all origins (the `/v1/*` paths already do this). |
124
+ | `OPENAI_COMPAT_NEW_CONVO_HEURISTIC` | (unset) | Set to `1` to enable webchat mode (see above). |
125
+ | `OPENAI_COMPAT_TOOLS_PER_MESSAGE` | (unset) | Set to `1` to re-send the full tool schemas on every turn (see [Tool definitions](#tool-definitions-and-where-they-live)). Needed only when the tool set changes mid-conversation; costs per-turn prompt growth. |
126
+ | `OPENAI_COMPAT_STATUS_URL` | (unset) | If set, the bridge POSTs JSON status updates to this URL (fire-and-forget, 2s timeout). See [Status webhook](#status-webhook). |
127
+ | `OPENCLAW_SERVE_MAX_SESSIONS` | `32` | Max concurrent OpenAI-compat sessions in serve mode. Bumped from the in-plugin default of 5 because each distinct caller now gets its own `sys-<hash>` session. |
128
+ | `OPENCLAW_SERVE_TTL_MINUTES` | `60` | Idle TTL for OpenAI-compat sessions in serve mode. Idle sessions are reaped by a 60s background loop; persisted disk registry is kept for 7 days so a returning caller is auto-resumed. |
129
129
 
130
130
  ## Inspection: `GET /v1/sessions`
131
131
 
@@ -235,14 +235,14 @@ Errors use the OpenAI error envelope:
235
235
  { "error": { "message": "...", "type": "invalid_request_error" } }
236
236
  ```
237
237
 
238
- | Status | When |
239
- |---|---|
240
- | 400 | `messages` empty/missing, no user message, invalid `max_tokens` |
241
- | 401 | Missing or wrong bearer token (when auth enabled) |
242
- | 415 | POST without `Content-Type: application/json` |
243
- | 429 | Rate limited (`OPENCLAW_RATE_LIMIT` exceeded) |
244
- | 503 | Failed to start a new session (model unavailable, CLI crashed at boot) |
245
- | 500 | Mid-turn failure |
238
+ | Status | When |
239
+ | ------ | ---------------------------------------------------------------------- |
240
+ | 400 | `messages` empty/missing, no user message, invalid `max_tokens` |
241
+ | 401 | Missing or wrong bearer token (when auth enabled) |
242
+ | 415 | POST without `Content-Type: application/json` |
243
+ | 429 | Rate limited (`OPENCLAW_RATE_LIMIT` exceeded) |
244
+ | 503 | Failed to start a new session (model unavailable, CLI crashed at boot) |
245
+ | 500 | Mid-turn failure |
246
246
 
247
247
  ## Related
248
248