@enderfga/claw-orchestrator 5.1.0 → 6.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (147) hide show
  1. package/README.md +26 -26
  2. package/dist/bin/cli.js +107 -1
  3. package/dist/bin/cli.js.map +1 -1
  4. package/dist/src/acp-server.d.ts +5 -5
  5. package/dist/src/acp-server.js +3 -3
  6. package/dist/src/acp-server.js.map +1 -1
  7. package/dist/src/autoloop/dispatcher.d.ts +22 -0
  8. package/dist/src/autoloop/dispatcher.js +71 -13
  9. package/dist/src/autoloop/dispatcher.js.map +1 -1
  10. package/dist/src/autoloop/messages.d.ts +10 -0
  11. package/dist/src/autoloop/messages.js.map +1 -1
  12. package/dist/src/autoloop/runner.js +6 -0
  13. package/dist/src/autoloop/runner.js.map +1 -1
  14. package/dist/src/constants.d.ts +0 -6
  15. package/dist/src/constants.js +0 -6
  16. package/dist/src/constants.js.map +1 -1
  17. package/dist/src/council.d.ts +15 -0
  18. package/dist/src/council.js +48 -35
  19. package/dist/src/council.js.map +1 -1
  20. package/dist/src/dashboard/index.html +191 -6
  21. package/dist/src/embedded-server.js +132 -9
  22. package/dist/src/embedded-server.js.map +1 -1
  23. package/dist/src/fanout.d.ts +30 -1
  24. package/dist/src/fanout.js +32 -3
  25. package/dist/src/fanout.js.map +1 -1
  26. package/dist/src/index.js +359 -4
  27. package/dist/src/index.js.map +1 -1
  28. package/dist/src/kernel/agent-step.d.ts +59 -0
  29. package/dist/src/kernel/agent-step.js +100 -0
  30. package/dist/src/kernel/agent-step.js.map +1 -0
  31. package/dist/src/kernel/conditions.d.ts +11 -0
  32. package/dist/src/kernel/conditions.js +24 -0
  33. package/dist/src/kernel/conditions.js.map +1 -0
  34. package/dist/src/kernel/engine.d.ts +319 -0
  35. package/dist/src/kernel/engine.js +1047 -0
  36. package/dist/src/kernel/engine.js.map +1 -0
  37. package/dist/src/kernel/exec.d.ts +43 -0
  38. package/dist/src/kernel/exec.js +112 -0
  39. package/dist/src/kernel/exec.js.map +1 -0
  40. package/dist/src/kernel/file-lock.d.ts +50 -0
  41. package/dist/src/kernel/file-lock.js +135 -0
  42. package/dist/src/kernel/file-lock.js.map +1 -0
  43. package/dist/src/kernel/nodes/agent.d.ts +4 -0
  44. package/dist/src/kernel/nodes/agent.js +35 -0
  45. package/dist/src/kernel/nodes/agent.js.map +1 -0
  46. package/dist/src/kernel/nodes/autoloop.d.ts +78 -0
  47. package/dist/src/kernel/nodes/autoloop.js +75 -0
  48. package/dist/src/kernel/nodes/autoloop.js.map +1 -0
  49. package/dist/src/kernel/nodes/council.d.ts +12 -0
  50. package/dist/src/kernel/nodes/council.js +88 -0
  51. package/dist/src/kernel/nodes/council.js.map +1 -0
  52. package/dist/src/kernel/nodes/fanout.d.ts +11 -0
  53. package/dist/src/kernel/nodes/fanout.js +63 -0
  54. package/dist/src/kernel/nodes/fanout.js.map +1 -0
  55. package/dist/src/kernel/nodes/human-gate.d.ts +4 -0
  56. package/dist/src/kernel/nodes/human-gate.js +7 -0
  57. package/dist/src/kernel/nodes/human-gate.js.map +1 -0
  58. package/dist/src/kernel/nodes/index.d.ts +12 -0
  59. package/dist/src/kernel/nodes/index.js +21 -0
  60. package/dist/src/kernel/nodes/index.js.map +1 -0
  61. package/dist/src/kernel/nodes/router.d.ts +4 -0
  62. package/dist/src/kernel/nodes/router.js +12 -0
  63. package/dist/src/kernel/nodes/router.js.map +1 -0
  64. package/dist/src/kernel/nodes/subflow.d.ts +13 -0
  65. package/dist/src/kernel/nodes/subflow.js +38 -0
  66. package/dist/src/kernel/nodes/subflow.js.map +1 -0
  67. package/dist/src/kernel/nodes/ultraapp.d.ts +60 -0
  68. package/dist/src/kernel/nodes/ultraapp.js +62 -0
  69. package/dist/src/kernel/nodes/ultraapp.js.map +1 -0
  70. package/dist/src/kernel/nodes/verifier.d.ts +14 -0
  71. package/dist/src/kernel/nodes/verifier.js +84 -0
  72. package/dist/src/kernel/nodes/verifier.js.map +1 -0
  73. package/dist/src/kernel/projections.d.ts +42 -0
  74. package/dist/src/kernel/projections.js +133 -0
  75. package/dist/src/kernel/projections.js.map +1 -0
  76. package/dist/src/kernel/repo.d.ts +13 -0
  77. package/dist/src/kernel/repo.js +64 -0
  78. package/dist/src/kernel/repo.js.map +1 -0
  79. package/dist/src/kernel/secrets.d.ts +25 -0
  80. package/dist/src/kernel/secrets.js +48 -0
  81. package/dist/src/kernel/secrets.js.map +1 -0
  82. package/dist/src/kernel/store.d.ts +225 -0
  83. package/dist/src/kernel/store.js +838 -0
  84. package/dist/src/kernel/store.js.map +1 -0
  85. package/dist/src/kernel/templates/index.d.ts +140 -0
  86. package/dist/src/kernel/templates/index.js +266 -0
  87. package/dist/src/kernel/templates/index.js.map +1 -0
  88. package/dist/src/kernel/types.d.ts +326 -0
  89. package/dist/src/kernel/types.js +19 -0
  90. package/dist/src/kernel/types.js.map +1 -0
  91. package/dist/src/run-ledger.d.ts +57 -3
  92. package/dist/src/run-ledger.js +45 -2
  93. package/dist/src/run-ledger.js.map +1 -1
  94. package/dist/src/session-manager.d.ts +176 -129
  95. package/dist/src/session-manager.js +652 -603
  96. package/dist/src/session-manager.js.map +1 -1
  97. package/dist/src/types.d.ts +33 -3
  98. package/dist/src/ultraapp/build.d.ts +117 -3
  99. package/dist/src/ultraapp/build.js +319 -3
  100. package/dist/src/ultraapp/build.js.map +1 -1
  101. package/dist/src/ultraapp/contract.d.ts +52 -0
  102. package/dist/src/ultraapp/contract.js +83 -0
  103. package/dist/src/ultraapp/contract.js.map +1 -0
  104. package/dist/src/ultraapp/conventions.js +9 -2
  105. package/dist/src/ultraapp/conventions.js.map +1 -1
  106. package/dist/src/ultraapp/fix-on-failure.d.ts +21 -2
  107. package/dist/src/ultraapp/fix-on-failure.js +46 -62
  108. package/dist/src/ultraapp/fix-on-failure.js.map +1 -1
  109. package/dist/src/ultraapp/manager.d.ts +107 -2
  110. package/dist/src/ultraapp/manager.js +305 -86
  111. package/dist/src/ultraapp/manager.js.map +1 -1
  112. package/dist/src/verify/baseline.d.ts +73 -0
  113. package/dist/src/verify/baseline.js +186 -0
  114. package/dist/src/verify/baseline.js.map +1 -0
  115. package/dist/src/verify/contract.d.ts +116 -0
  116. package/dist/src/verify/contract.js +142 -0
  117. package/dist/src/verify/contract.js.map +1 -0
  118. package/dist/src/verify/evidence.d.ts +61 -0
  119. package/dist/src/verify/evidence.js +133 -0
  120. package/dist/src/verify/evidence.js.map +1 -0
  121. package/dist/src/verify/runner.d.ts +63 -0
  122. package/dist/src/verify/runner.js +317 -0
  123. package/dist/src/verify/runner.js.map +1 -0
  124. package/openclaw.plugin.json +8 -0
  125. package/package.json +2 -2
  126. package/skills/SKILL.md +120 -79
  127. package/skills/references/acp.md +17 -17
  128. package/skills/references/autoloop.md +139 -65
  129. package/skills/references/claude-cli-tracking.md +4 -4
  130. package/skills/references/cli.md +101 -59
  131. package/skills/references/council.md +109 -37
  132. package/skills/references/dashboard.md +34 -6
  133. package/skills/references/getting-started.md +13 -13
  134. package/skills/references/inbox.md +4 -4
  135. package/skills/references/mcp.md +39 -34
  136. package/skills/references/multi-engine.md +51 -47
  137. package/skills/references/observability.md +88 -28
  138. package/skills/references/openai-compat.md +39 -39
  139. package/skills/references/sessions.md +43 -25
  140. package/skills/references/tools.md +402 -309
  141. package/skills/references/ultra.md +45 -45
  142. package/skills/references/ultraapp.md +126 -50
  143. package/skills/references/verification.md +187 -0
  144. package/skills/references/workflow.md +362 -0
  145. package/dist/src/ultraapp/fix-on-failure-session.d.ts +0 -23
  146. package/dist/src/ultraapp/fix-on-failure-session.js +0 -51
  147. package/dist/src/ultraapp/fix-on-failure-session.js.map +0 -1
@@ -21,11 +21,11 @@ This page is the operator reference.
21
21
 
22
22
  ## Roles
23
23
 
24
- | Agent | Default | cwd | Owns |
25
- |---|---|---|---|
26
- | **Planner** | claude / opus | workspace | strategy, `plan.md`, `goal.json`, talking to you |
27
- | **Coder** | claude / sonnet | workspace | code changes, eval execution |
28
- | **Reviewer** | claude / sonnet | `<workspace>/tasks/<run_id>/reviewer_sandbox/` | distrust audit; advance / hold / rollback |
24
+ | Agent | Default | cwd | Owns |
25
+ | ------------ | --------------- | ---------------------------------------------- | ------------------------------------------------ |
26
+ | **Planner** | claude / opus | workspace | strategy, `plan.md`, `goal.json`, talking to you |
27
+ | **Coder** | claude / sonnet | workspace | code changes, eval execution |
28
+ | **Reviewer** | claude / sonnet | `<workspace>/tasks/<run_id>/reviewer_sandbox/` | distrust audit; advance / hold / rollback |
29
29
 
30
30
  Each role can use any built-in engine, or a `custom` engine config supplied by a
31
31
  local caller (custom engines name an executable, so the HTTP API does not accept
@@ -105,14 +105,14 @@ curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_stop \
105
105
 
106
106
  ## Plugin tools
107
107
 
108
- | Tool | Args | What |
109
- |---|---|---|
110
- | `autoloop_start` | `run_id`, `workspace`, per-role `*_engine?`, `*_model?`, `*_custom_engine?`, `send_timeout_ms?` | Start a run; launches Planner and stores Coder/Reviewer defaults. Each `custom` role requires its matching config. |
111
- | `autoloop_chat` | `run_id`, `text` | Send a chat message to the Planner; returns the Planner's reply. |
112
- | `autoloop_status` | `run_id` | Current state (status, iter, push count, subagents_spawned). |
113
- | `autoloop_list` | — | All active runs in this manager process. |
114
- | `autoloop_stop` | `run_id`, `reason?` | Terminate; stops Planner / Coder / Reviewer. |
115
- | `autoloop_reset_agent` | `run_id`, `agent` ('planner' / 'coder' / 'reviewer'), `force?`, `eager_restart?` | Reset one subagent. Planner reset requires `force: true`. |
108
+ | Tool | Args | What |
109
+ | ---------------------- | ----------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
110
+ | `autoloop_start` | `run_id`, `workspace`, per-role `*_engine?`, `*_model?`, `*_custom_engine?`, `send_timeout_ms?` | Start a run; launches Planner and stores Coder/Reviewer defaults. Each `custom` role requires its matching config. |
111
+ | `autoloop_chat` | `run_id`, `text` | Send a chat message to the Planner; returns the Planner's reply. |
112
+ | `autoloop_status` | `run_id` | Current state (status, iter, push count, subagents_spawned). |
113
+ | `autoloop_list` | — | All active runs in this manager process. |
114
+ | `autoloop_stop` | `run_id`, `reason?` | Terminate; stops Planner / Coder / Reviewer. |
115
+ | `autoloop_reset_agent` | `run_id`, `agent` ('planner' / 'coder' / 'reviewer'), `force?`, `eager_restart?` | Reset one subagent. Planner reset requires `force: true`. |
116
116
 
117
117
  ## Planner-emitted control tools
118
118
 
@@ -120,17 +120,17 @@ The Planner controls the run by emitting fenced ` ```autoloop ` JSON blocks
120
120
  inside its replies. The dispatcher parses them out and applies them. You
121
121
  never see the JSON — only the Planner's narrative.
122
122
 
123
- | Tool | Args | What |
124
- |---|---|---|
125
- | `notify_user` | `level` ('info' / 'warn' / 'decision' / 'error'), `summary`, `detail?`, `channel?` ('auto' / 'wechat' / 'webchat' / 'both' / 'email') | Push you out-of-band. |
126
- | `spawn_subagents` | `coder_engine?`, `coder_model?`, `reviewer_engine?`, `reviewer_model?`, `initial_directive?` | Start Coder + Reviewer. Omitted values inherit run defaults. An engine change without a model uses the new engine's default. Once a role session has started, changing its engine/model is rejected. Custom configs cannot be emitted by Planner. Only after explicit user approval. |
127
- | `send_directive` | `goal`, `constraints?`, `success_criteria?`, `max_attempts?` | Next iter's instruction to Coder. |
128
- | `pause_loop` | `reason` | Halt subloop at next iter boundary; chat keeps working. |
129
- | `resume_loop` | — | Resume after pause. |
130
- | `terminate` | `reason` | End run. |
131
- | `update_push_policy` | partial PushPolicy | Mutate notification rules (e.g. when you say "tell me every iter"). |
132
- | `write_plan` | `content` (full plan.md body), `commit_message?` | Write `plan.md` to the workspace and git-commit. The **only** way the Planner can author plan.md — Write/Edit are stripped from the Planner session as a hard role boundary. Re-running replaces the whole file. |
133
- | `write_goal` | `content` (full goal.json body), `commit_message?` | Same, for `goal.json`. Content is JSON-validated before write; malformed content errors back to the Planner. |
123
+ | Tool | Args | What |
124
+ | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
125
+ | `notify_user` | `level` ('info' / 'warn' / 'decision' / 'error'), `summary`, `detail?`, `channel?` ('auto' / 'wechat' / 'webchat' / 'both' / 'email') | Push you out-of-band. |
126
+ | `spawn_subagents` | `coder_engine?`, `coder_model?`, `reviewer_engine?`, `reviewer_model?`, `initial_directive?` | Start Coder + Reviewer. Omitted values inherit run defaults. An engine change without a model uses the new engine's default. Once a role session has started, changing its engine/model is rejected. Custom configs cannot be emitted by Planner. Only after explicit user approval. |
127
+ | `send_directive` | `goal`, `constraints?`, `success_criteria?`, `max_attempts?` | Next iter's instruction to Coder. |
128
+ | `pause_loop` | `reason` | Halt subloop at next iter boundary; chat keeps working. |
129
+ | `resume_loop` | — | Resume after pause. |
130
+ | `terminate` | `reason` | End run. |
131
+ | `update_push_policy` | partial PushPolicy | Mutate notification rules (e.g. when you say "tell me every iter"). |
132
+ | `write_plan` | `content` (full plan.md body), `commit_message?` | Write `plan.md` to the workspace and git-commit. The **only** way the Planner can author plan.md — Write/Edit are stripped from the Planner session as a hard role boundary. Re-running replaces the whole file. |
133
+ | `write_goal` | `content` (full goal.json body), `commit_message?` | Same, for `goal.json`. Content is JSON-validated before write; malformed content errors back to the Planner. |
134
134
 
135
135
  ### Custom engines and resume
136
136
 
@@ -150,24 +150,25 @@ for the `CustomEngineConfig` shape.
150
150
 
151
151
  ## Default push policy
152
152
 
153
- | Event | Default |
154
- |---|---|
155
- | on_start | info / wechat ("loop started, will notify on issues") |
156
- | on_iter_done_ok | silent |
157
- | on_target_hit | info / both (webchat + wechat) |
158
- | on_metric_regression_2 | warn / both |
159
- | on_reviewer_reject_2 | warn / both |
160
- | on_phase_error | error / both |
161
- | on_stall_30min | warn / wechat |
162
- | on_decision_needed | decision / both |
153
+ | Event | Default |
154
+ | ---------------------- | ----------------------------------------------------- |
155
+ | on_start | info / wechat ("loop started, will notify on issues") |
156
+ | on_iter_done_ok | silent |
157
+ | on_target_hit | info / both (webchat + wechat) |
158
+ | on_metric_regression_2 | warn / both |
159
+ | on_reviewer_reject_2 | warn / both |
160
+ | on_phase_error | error / both |
161
+ | on_stall_30min | warn / wechat |
162
+ | on_decision_needed | decision / both |
163
163
 
164
164
  5-minute dedup on (level, summary) prevents duplicate pushes from the same
165
165
  event. Channel chain: `auto` walks wechat → whatsapp → email; `wechat` /
166
166
  `webchat` / `email` route directly; `both` does webchat (if session known)
167
- + wechat fallback chain. **`on_phase_error` and `on_decision_needed` cannot
168
- be set to `silent: true`** by Planner — `update_push_policy` strips the flag
169
- and records the attempt in `decisions.jsonl` (these channels are the
170
- operator's lifeline; they stay loud).
167
+
168
+ - wechat fallback chain. **`on_phase_error` and `on_decision_needed` cannot
169
+ be set to `silent: true`** by Planner — `update_push_policy` strips the flag
170
+ and records the attempt in `decisions.jsonl` (these channels are the
171
+ operator's lifeline; they stay loud).
171
172
 
172
173
  ## Auto-compact
173
174
 
@@ -195,7 +196,7 @@ consecutive `phase_error`s and:
195
196
  1. Fires `on_phase_error` on each one (defaults to error / both channels).
196
197
  2. After `phaseErrorCircuit` consecutive errors (default **3**) emits a
197
198
  `decision`-level push and an automatic `terminate { reason:
198
- 'phase_error_circuit' }`.
199
+ 'phase_error_circuit' }`.
199
200
 
200
201
  A successful (non-error) `iter_done` resets the counter. Override the
201
202
  threshold via `AutoloopConfig.phaseErrorCircuit`.
@@ -214,15 +215,15 @@ with `agent: 'reviewer', eager_restart: true`).
214
215
  `<ledger>/decisions.jsonl` is the auditable trail of runner / dispatcher
215
216
  decisions:
216
217
 
217
- | Kind | When |
218
- |---|---|
219
- | `spawn_subagents` | Planner emits `spawn_subagents` |
220
- | `reset_agent` | Any agent reset (manual or auto-recovery) |
221
- | `compact` | Auto-compact fires |
222
- | `update_push_policy` | Planner mutates the policy |
223
- | `policy_silence_blocked` | Planner tried to silence a critical channel |
224
- | `phase_error` | Surfaced from dispatcher to runner |
225
- | `terminate` | Run ends (planner reason or `phase_error_circuit`) |
218
+ | Kind | When |
219
+ | ------------------------ | -------------------------------------------------- |
220
+ | `spawn_subagents` | Planner emits `spawn_subagents` |
221
+ | `reset_agent` | Any agent reset (manual or auto-recovery) |
222
+ | `compact` | Auto-compact fires |
223
+ | `update_push_policy` | Planner mutates the policy |
224
+ | `policy_silence_blocked` | Planner tried to silence a critical channel |
225
+ | `phase_error` | Surfaced from dispatcher to runner |
226
+ | `terminate` | Run ends (planner reason or `phase_error_circuit`) |
226
227
 
227
228
  JSONL, one entry per line, ts-prefixed.
228
229
 
@@ -260,19 +261,21 @@ Every JSON artifact in the ledger carries a `schema_version` field (currently
260
261
 
261
262
  ## Backend HTTP / SSE
262
263
 
263
- | Endpoint | Returns |
264
- |---|---|
265
- | `GET /autoloop/list` | `{ ok, runs: AutoloopState[] }` |
266
- | `POST /autoloop/new` | `{ ok, run_id, planner_session }` — body `{ workspace, run_id?, planner_engine?, planner_model?, planner_custom_engine?, coder_engine?, coder_model?, coder_custom_engine?, reviewer_engine?, reviewer_model?, reviewer_custom_engine?, send_timeout_ms? }` |
267
- | `GET /autoloop/<id>/state` | `{ ok, state: AutoloopState }` — also returns a `terminated`-state stub reconstructed from the registry for runs that aren't in this process's memory, so the dashboard can open historical runs without 404'ing. |
268
- | `GET /autoloop/<id>/push_log` | `{ ok, entries: PushLogEntry[] }` — served from the ledger via `autoloopStatus`, so historical runs work the same as live ones. |
269
- | `GET /autoloop/<id>/chat_history` | `{ ok, entries: ChatEntry[] }` — replays `<ledger>/chat.jsonl`. The dashboard fetches this when opening a run so the Planner-pane conversation survives a page refresh / cross-process / re-opening a terminated run. Returns `[]` when the file doesn't exist (e.g. runs that predate the chat-history feature). |
270
- | `GET /autoloop/<id>/events` | SSE: `snapshot` / `message` / `state` / `push` / `iter_done` / `planner_reply` / `planner_error` / `coder_reply` / `reviewer_reply` / `terminated`. For runs that are NOT in this process's memory (terminated, or live in another process), the endpoint emits a single-shot `snapshot` + `terminated` then closes — the dashboard's existing handlers render history without hanging. |
271
- | `POST /autoloop/<id>/chat` | **202** `{ ok, queued: true }` — body `{ text }`. Fire-and-forget: the Planner's reply streams back via the `/events` SSE channel as a `planner_reply` event (or `planner_error` on failure); the HTTP response intentionally does NOT wait for it, because first-contact replies routinely exceed reverse-proxy idle limits (e.g. Cloudflare Tunnel cuts at ~100s → 524). 400 on empty text, 404 when the run is not in this process's memory. The MCP `autoloop_chat` tool path keeps the synchronous await-and-return-reply semantics (it runs in-process). |
272
- | `POST /autoloop/<id>/resume` | `{ ok, state }` — restore the role engine/model choices from the registry and re-create dispatcher + runner. Optional body fields `planner_custom_engine`, `coder_custom_engine`, `reviewer_custom_engine` must be supplied again for roles using `custom` because configs are intentionally not persisted. Existing engine-specific conversation resume behavior is reused where supported; `chat.jsonl` remains the visual history fallback. 404 when the registry has no record. |
273
- | `POST /autoloop/<id>/delete` | `{ ok }` — stops the runner if still alive, scrubs the row from `~/.claw-orchestrator/autoloop-registry.jsonl`, and purges `persistedSessions` so the run cannot be `/resume`'d back. The ledger directory under `<workspace>/tasks/<run_id>/` is kept on disk. 404 if the run was not present in either memory or the registry. |
264
+ | Endpoint | Returns |
265
+ | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
266
+ | `GET /autoloop/list` | `{ ok, runs: AutoloopState[] }` |
267
+ | `POST /autoloop/new` | `{ ok, run_id, planner_session }` — body `{ workspace, run_id?, planner_engine?, planner_model?, planner_custom_engine?, coder_engine?, coder_model?, coder_custom_engine?, reviewer_engine?, reviewer_model?, reviewer_custom_engine?, send_timeout_ms? }` |
268
+ | `GET /autoloop/<id>/state` | `{ ok, state: AutoloopState }` — also returns a `terminated`-state stub reconstructed from the registry for runs that aren't in this process's memory, so the dashboard can open historical runs without 404'ing. |
269
+ | `GET /autoloop/<id>/push_log` | `{ ok, entries: PushLogEntry[] }` — served from the ledger via `autoloopStatus`, so historical runs work the same as live ones. |
270
+ | `GET /autoloop/<id>/chat_history` | `{ ok, entries: ChatEntry[] }` — replays `<ledger>/chat.jsonl`. The dashboard fetches this when opening a run so the Planner-pane conversation survives a page refresh / cross-process / re-opening a terminated run. Returns `[]` when the file doesn't exist (e.g. runs that predate the chat-history feature). |
271
+ | `GET /autoloop/<id>/events` | SSE: `snapshot` / `message` / `state` / `push` / `iter_done` / `planner_reply` / `planner_error` / `coder_reply` / `reviewer_reply` / `terminated`. For runs that are NOT in this process's memory (terminated, or live in another process), the endpoint emits a single-shot `snapshot` + `terminated` then closes — the dashboard's existing handlers render history without hanging. |
272
+ | `POST /autoloop/<id>/chat` | **202** `{ ok, queued: true }` — body `{ text }`. Fire-and-forget: the Planner's reply streams back via the `/events` SSE channel as a `planner_reply` event (or `planner_error` on failure); the HTTP response intentionally does NOT wait for it, because first-contact replies routinely exceed reverse-proxy idle limits (e.g. Cloudflare Tunnel cuts at ~100s → 524). 400 on empty text, 404 when the run is not in this process's memory. The MCP `autoloop_chat` tool path keeps the synchronous await-and-return-reply semantics (it runs in-process). |
273
+ | `GET /autoloop/<id>/resume-requirements` | `{ ok, runId, rolesNeedingCustomEngine }` — the roles whose engine was `custom`, so a caller knows which secret references a resume needs. Role names only; nothing sensitive. 404 when there is no such run. |
274
+ | `POST /autoloop/<id>/resume` | `{ ok, state }` — restore the role engine/model choices from the run's spec and re-create dispatcher + runner. A custom-engine config is never persisted and is never accepted over HTTP, so a role using `custom` is re-supplied by **reference**: `plannerCustomEngineRef` / `coderCustomEngineRef` / `reviewerCustomEngineRef` name an environment variable `CLAWO_CUSTOM_ENGINE_<NAME>` on the orchestrator host, which the server reads and resolves. The name is not sensitive, the value never crosses the wire, and an unknown name is an error rather than a silent start without credentials. Existing engine-specific conversation resume behavior is reused where supported; `chat.jsonl` remains the visual history fallback. 404 when there is no such run. |
275
+ | `POST /autoloop/<id>/delete` | `{ ok }` — stops the runner if still alive, scrubs the row from `~/.claw-orchestrator/autoloop-registry.jsonl`, and purges `persistedSessions` so the run cannot be `/resume`'d back. The ledger directory under `<workspace>/tasks/<run_id>/` is kept on disk. 404 if the run was not present in either memory or the registry. |
274
276
 
275
277
  The 3-pane UI consumes these endpoints:
278
+
276
279
  - **Left**: Planner chat (subscribes to `planner_reply`)
277
280
  - **Center**: Coder activity (`coder_reply` + `iter_done`)
278
281
  - **Right**: Reviewer verdicts (`reviewer_reply`)
@@ -293,15 +296,13 @@ wrote down. A typical shape:
293
296
  "name": "test_pass_rate",
294
297
  "direction": "max",
295
298
  "extract_cmd": "bash eval.sh | grep -oE 'metric=[0-9.]+' | cut -d= -f2",
296
- "target": 1.0
299
+ "target": 1.0,
297
300
  },
298
- "gates": [
299
- { "name": "tests_pass", "cmd": "npm test", "must": "exit-0" }
300
- ],
301
+ "gates": [{ "name": "tests_pass", "cmd": "npm test", "must": "exit-0" }],
301
302
  "termination": {
302
303
  "max_iters": 10,
303
- "scalar_target_hit": true
304
- }
304
+ "scalar_target_hit": true,
305
+ },
305
306
  }
306
307
  ```
307
308
 
@@ -347,3 +348,76 @@ iter 0 ledger artifacts (`directive` + `eval_output` + `diff.patch` +
347
348
  map; the on-disk ledger survives but cannot resume a running state.
348
349
  - **Multi-run / same workspace** races on `git index.lock`. Run separate
349
350
  workspaces (or git worktrees) for concurrent runs.
351
+
352
+ ## Acceptance contracts (6.0.0)
353
+
354
+ The Reviewer's `advance` was the only thing standing between an iteration and
355
+ "done", and it could not do the job it was given. Its prompt tells it to
356
+ "re-derive the metric independently if the sandbox has the bits to do so", but
357
+ `stageReviewSandbox` copies in the iteration's artifacts — `directive.json`,
358
+ `diff.patch`, `eval_output.json`, `coder_summary.txt` — plus `plan.md`,
359
+ `goal.json`, and the prior verdict. No code, no evaluator. The verdict was
360
+ therefore a reading of the Coder's own report, and `eval_output` was literally
361
+ whatever the Coder passed to a tool call.
362
+
363
+ Pass a contract at autoloop start and an `advance` is held unless the checks
364
+ pass. The verdict is rewritten to `hold`, the reason is appended to
365
+ `audit_notes`, and the evidence bundle lands at
366
+ `<ledger>/iter/<n>/evidence/`. Without a contract nothing changes.
367
+
368
+ ### `on_target_hit` now fires
369
+
370
+ That push-policy key was declared in `types.ts`, given a default, and whitelisted
371
+ for runtime updates — and had **zero firing sites** anywhere in the codebase.
372
+ Autoloop had four ways to notice it was failing (stall, metric regression,
373
+ reviewer rejections, phase-error circuit) and no way to notice it had succeeded,
374
+ because a Reviewer verdict is not a measurement. An acceptance contract is, so
375
+ `on_target_hit` fires when one passes.
376
+
377
+ ### `diff.patch` sees created files
378
+
379
+ The per-iteration patch was captured with a bare `git diff`, which lists tracked
380
+ modifications only. A file the Coder _created_ appeared in neither the patch nor
381
+ the `--name-only` fallback, while the `git add -A` two lines later committed it —
382
+ so the Reviewer audited a picture that structurally could not show new files.
383
+ The capture now covers tracked changes ∪ untracked files.
384
+
385
+ `files_changed` is also taken from git unconditionally. It previously preferred
386
+ the Coder's own `files_changed` whenever the Coder supplied one, despite the
387
+ comment above it saying the claim was not trusted.
388
+
389
+ ## Related
390
+
391
+ - [`verification.md`](./verification.md) — contracts, checks, evidence
392
+
393
+ ## Lifecycle moved to the kernel (6.0.0)
394
+
395
+ An autoloop is a kernel run whose single node holds the loop for as long as it
396
+ lives. Tool signatures are unchanged.
397
+
398
+ What went away: the `autoloops` map; `~/.claw-orchestrator/autoloop-registry.jsonl`
399
+ with its four bespoke helpers (append, remove-then-append upsert, reverse-scan
400
+ dedup, rewrite-via-tmp-file); and the two `Set`s — `_startingAutoloops` and
401
+ `_deletingAutoloops` — that existed only because a start and a delete could race
402
+ each other over that shared map.
403
+
404
+ Two things get better rather than merely moving:
405
+
406
+ - **`autoloop_status` on a run that is not live in this process** used to return
407
+ an all-zero stub labelled `reconstructed from registry — not in current process
408
+ memory`: iter 0, no metrics, no error history, because the registry only ever
409
+ held identity. The record holds the last state the loop published, so a
410
+ historical run opens with its real iteration count.
411
+ - **The engines `spawn_subagents` actually chose** land on the run record
412
+ alongside the rest of its state, instead of in a parallel file with its own
413
+ lifecycle.
414
+
415
+ `autoloop_resume` restarts a terminated run from the stored spec — the immutable
416
+ record of how it was started — rather than from a registry row whose older
417
+ versions omitted the engine fields entirely. Custom-engine configs are the one
418
+ thing the spec does not carry (they can hold secrets), so a resume must be given
419
+ them again.
420
+
421
+ Cancelling a run now tears the loop down the way a stop does. It previously left
422
+ the three persistent agents running and their session names claimed, which
423
+ surfaced much later as `session name already in use`.
@@ -8,13 +8,13 @@ This document tracks which Claude Code CLI version Claw Orchestrator is currentl
8
8
 
9
9
  | Plugin Version | Claude CLI Version | Date | Notable integrations |
10
10
  | v4.8.0 | 2.1.207 | 2026-07-12 | **Autoloop role-level multi-engine support.** Planner, Coder, and Reviewer can select independent engines/models while preserving the Claude defaults. Built-in non-Claude Planners use native read-only/plan modes and receive their role protocol in-band. Spawn selections persist across resume; Codex persists its real thread ID. Runtime and invocation checks used Claude Code 2.1.207 and Codex 0.144.1. |
11
- | v4.7.0 | 2.1.206 | 2026-07-10 | **Antigravity engine ships + permission-mode sync.** Main feature is the community-contributed first-class `engine: 'agy'` (PR #71, reviewed + hardened: layered resume-ID gating, `agy/` prefix routing, shared `sanitize.ts` across all six engines, ENGINE_TYPES single source). Weekly CLI sync: CC 2.1.200 renamed the `default` permission mode to **`manual`** — verified against 2.1.206 that the choices are now `acceptEdits/auto/bypassPermissions/manual/dontAsk/plan`, that `default` is still accepted (hidden compat), and that **`delegate` is hard-rejected at spawn** — so PermissionMode gains `manual`, drops `delegate`, and agy/gemini map `manual` like `default` (→ `--sandbox`). Codex 0.143.0: empirically re-tested `-c model_reasoning_effort=max` — still 400-rejected for gpt-5.5 (the "first-class max" note is Bedrock GPT-5.6-only), so the `max`→`xhigh` map stays. GPT-5.6 Sol/Terra/Luna registered with official pricing ($5/$30, $2.50/$15, $1/$6; 1M/1M/400K ctx) after the user reported using it — it's a limited preview on API/Codex-auth paths (empirically: ChatGPT-account Codex auth gets a 400, which is why the first probe on this box misread it as Bedrock-only; lesson — an auth-path rejection is not model non-existence). Codex default stays gpt-5.5. Free upside: CC 2.1.205 fixed `--json-schema` invalid-schema silent fallback + `format` keyword rejection; CC 2.1.203 fixed background sessions dropping shell-exported `ANTHROPIC_BASE_URL`. Pins → CC 2.1.206 / Codex 0.143.0 (installed; npm has 0.144.1, exec surface unchanged per release notes). |
11
+ | v4.7.0 | 2.1.206 | 2026-07-10 | **Antigravity engine ships + permission-mode sync.** Main feature is the community-contributed first-class `engine: 'agy'` (PR #71, reviewed + hardened: layered resume-ID gating, `agy/` prefix routing, shared `sanitize.ts` across all six engines, ENGINE*TYPES single source). Weekly CLI sync: CC 2.1.200 renamed the `default` permission mode to **`manual`** — verified against 2.1.206 that the choices are now `acceptEdits/auto/bypassPermissions/manual/dontAsk/plan`, that `default` is still accepted (hidden compat), and that **`delegate` is hard-rejected at spawn** — so PermissionMode gains `manual`, drops `delegate`, and agy/gemini map `manual` like `default` (→ `--sandbox`). Codex 0.143.0: empirically re-tested `-c model_reasoning_effort=max` — still 400-rejected for gpt-5.5 (the "first-class max" note is Bedrock GPT-5.6-only), so the `max`→`xhigh` map stays. GPT-5.6 Sol/Terra/Luna registered with official pricing ($5/$30, $2.50/$15, $1/$6; 1M/1M/400K ctx) after the user reported using it — it's a limited preview on API/Codex-auth paths (empirically: ChatGPT-account Codex auth gets a 400, which is why the first probe on this box misread it as Bedrock-only; lesson — an auth-path rejection is not model non-existence). Codex default stays gpt-5.5. Free upside: CC 2.1.205 fixed `--json-schema` invalid-schema silent fallback + `format` keyword rejection; CC 2.1.203 fixed background sessions dropping shell-exported `ANTHROPIC_BASE_URL`. Pins → CC 2.1.206 / Codex 0.143.0 (installed; npm has 0.144.1, exec surface unchanged per release notes). |
12
12
  | v4.6.0 | 2.1.199 | 2026-07-03 | **Model registry sync — Claude Fable 5.** Registered `claude-fable-5` (first Claude 5-family model, tier above Opus; standard $10/$50 per Mtok, cache read $1, full 1M context at standard rates per the official pricing page) with new `fable` alias; taught the `isClaudeModel`/`resolveProvider` heuristics to recognize `fable`/`mythos` strings (they only matched claude/opus/sonnet/haiku). Mythos 5 not listed (same price, limited availability). CC 2.1.198–199 are subagent/background-agent reliability fixes — no invocation-surface change; free upside for us: subagent partial output on rate-limit/server error is now returned instead of silently dropped, and API errors in subagents are reported to the parent. Codex unchanged (0.142.5 is a log-scrub patch; pin stays 0.142.4 as installed). |
13
13
  | v4.5.0 | 2.1.197 | 2026-07-01 | **Model registry sync — Claude Sonnet 5 + gpt-5.5 pricing.** CLI 2.1.197 shipped Sonnet 5 as the new default (native 1M-token context; standard $3/$15 per Mtok, launch promo $2/$10 through 2026-08-31 — we price the standard rate). Registered `claude-sonnet-5` in `models.ts` and moved the `sonnet` alias to it (was pinned to `claude-sonnet-4-6`), so `--model sonnet` tracks the CLI's own default and cost/context accounting stays correct; `claude-sonnet-4-6` stays selectable by full id. Also corrected `gpt-5.5` (the default Codex model) from placeholder pricing to OpenAI's published $5/$30 per Mtok + 1M context, and updated docs/examples off `gpt-5.4`. The CC 2.1.179→2.1.197 and Codex 0.138→0.142.x ranges are otherwise bug-fix / TUI / remote-executor / plugin-marketplace work that doesn't touch our invocation flags or the stream-json / codex-exec event schema — no wrapper change. Free upside (no code change): 2.1.181 fixed prompt-caching on custom `ANTHROPIC_BASE_URL` (helps proxy mode), 2.1.187 fixed `--json-schema` StructuredOutput infinite-recall, 2.1.196 turned the 5-min streaming idle watchdog on by default; Codex 0.139 preserves `oneOf`/`allOf` in `--output-schema`. Bumped tested versions Claude 2.1.197 / Codex 0.142.4. |
14
14
  | v4.3.0 | 2.1.178 | 2026-06-16 | **Parity batch 2 + legacy-subsystem upgrades (local-only, no cloud).** Claude `--fallback-model` array form (CSV, verified via `claude --help`). Codex-app `codex_threads` (`thread/list`) and `thread/resume` on start when `resumeSessionId` is set (param shapes from `generate-json-schema`). Council agents gain per-agent `effort`/`ultracode`. `ultrareview` re-implemented on the new cross-engine `fanout` primitive (opt-in `engines`, default claude-only). Consensus parsing exposes match source for observability. Dropped on purpose: `codex cloud exec`/best-of-N (cloud/managed — loses local control), `--bg` (we own the subprocess), `--output-last-message` (we already capture final text), ultraplan `ultracode` (violates its plan-only contract). Autoloop mid-turn steer deferred — the loop is strictly sequential (Coder fully completes before the Reviewer runs), so steer would always fall back to a fresh turn; a real version needs concurrent review. |
15
- | v4.2.0 | 2.1.178 | 2026-06-16 | **`ultracode` integration + binary-verified parity pass.** Added the `ultracode` option on `session_start` — Claude Code's dynamic-workflow mode, wired as the `ultracode: true` settings key merged into `--settings` (confirmed by spawning: it activates `workflow_agent` events in headless stream-json; `--effort ultracode` is rejected by the CLI). Added `claude_agents_list` (wraps `claude agents --json`). Verified against the binary that `claude continue/respawn/stop/logs` do **not** exist as headless subcommands (session continuation stays on `--resume`). Codex 0.137 side: app-server RPC tools `codex_interrupt`/`codex_steer`/`codex_fork`/`codex_rollback`/`codex_models`, `codex exec` reasoning-effort (`-c model_reasoning_effort`) + `--profile` passthrough, and a cross-engine `fanout_*` primitive. Bumped tested versions Claude 2.1.178 / Codex 0.137.0. |
16
- | v4.1.2 | 2.1.161 | 2026-06-03 | **Model registry sync, not a CLI-flag integration.** Registered Opus 4.8 (`claude-opus-4-8`, now the `opus` alias) and 4.7 in `models.ts` — 2.1.154 shipped Opus 4.8 as the new default and our `opus` alias was still pinned to 4.6, mis-attributing cost. Effort ladder `low/medium/high/xhigh/max` was already supported (`index.ts`/`types.ts`). The 2.1.151–2.1.161 range (note: .151/.155 skipped) is otherwise TUI/reliability/telemetry; two fixes silently benefit our spawn path with no code change — 2.1.153 (stream-json stdin-close hang) and 2.1.161 (`-p` stdout corruption from background subagents). **Watch-out documented, not fixed:** 2.1.160 adds permission prompts under `acceptEdits` (our default) for build-tool config files (`.npmrc`/`.bazelrc`/`.pre-commit-config.yaml`/`.devcontainer/` etc.) and shell-startup files — headless flows touching these should set `dangerouslySkipPermissions` or a bypass permission mode. Codex unchanged (0.133.0); separately fixed Codex `turn.failed`/`error` events being swallowed. |
17
- | v4.1.1 | 2.1.150 | 2026-05-24 | **No Claude wrapper change** — 2.1.141–2.1.150 are almost entirely TUI / agent-view / security / visual; 2.1.150 itself is "internal infrastructure only". The one scripting-adjacent addition, `claude agents --json` (2.1.145), lists *CLI-managed* sessions and is not used by our own session manager. This release's real engine work was on Codex/Gemini: Codex `--output-schema` wired into `jsonSchema` (Codex 0.132+), Gemini `--skip-trust` for the 0.43 trusted-folders gate. Bumped tested versions Claude 2.1.150 / Codex 0.133.0 / Gemini 0.43.0. |
15
+ | v4.2.0 | 2.1.178 | 2026-06-16 | **`ultracode` integration + binary-verified parity pass.** Added the `ultracode` option on `session_start` — Claude Code's dynamic-workflow mode, wired as the `ultracode: true` settings key merged into `--settings` (confirmed by spawning: it activates `workflow_agent` events in headless stream-json; `--effort ultracode` is rejected by the CLI). Added `claude_agents_list` (wraps `claude agents --json`). Verified against the binary that `claude continue/respawn/stop/logs` do **not** exist as headless subcommands (session continuation stays on `--resume`). Codex 0.137 side: app-server RPC tools `codex_interrupt`/`codex_steer`/`codex_fork`/`codex_rollback`/`codex_models`, `codex exec` reasoning-effort (`-c model_reasoning_effort`) + `--profile` passthrough, and a cross-engine `fanout**` primitive. Bumped tested versions Claude 2.1.178 / Codex 0.137.0. |
16
+ | v4.1.2 | 2.1.161 | 2026-06-03 | **Model registry sync, not a CLI-flag integration.** Registered Opus 4.8 (`claude-opus-4-8`, now the `opus`alias) and 4.7 in`models.ts`— 2.1.154 shipped Opus 4.8 as the new default and our`opus`alias was still pinned to 4.6, mis-attributing cost. Effort ladder`low/medium/high/xhigh/max` was already supported (`index.ts`/`types.ts`). The 2.1.151–2.1.161 range (note: .151/.155 skipped) is otherwise TUI/reliability/telemetry; two fixes silently benefit our spawn path with no code change — 2.1.153 (stream-json stdin-close hang) and 2.1.161 (`-p`stdout corruption from background subagents). **Watch-out documented, not fixed:** 2.1.160 adds permission prompts under`acceptEdits` (our default) for build-tool config files (`.npmrc`/`.bazelrc`/`.pre-commit-config.yaml`/`.devcontainer/`etc.) and shell-startup files — headless flows touching these should set`dangerouslySkipPermissions`or a bypass permission mode. Codex unchanged (0.133.0); separately fixed Codex`turn.failed`/`error`events being swallowed. |
17
+ | v4.1.1 | 2.1.150 | 2026-05-24 | **No Claude wrapper change** — 2.1.141–2.1.150 are almost entirely TUI / agent-view / security / visual; 2.1.150 itself is "internal infrastructure only". The one scripting-adjacent addition,`claude agents --json` (2.1.145), lists *CLI-managed\* sessions and is not used by our own session manager. This release's real engine work was on Codex/Gemini: Codex `--output-schema` wired into `jsonSchema` (Codex 0.132+), Gemini `--skip-trust` for the 0.43 trusted-folders gate. Bumped tested versions Claude 2.1.150 / Codex 0.133.0 / Gemini 0.43.0. |
18
18
  |---|---|---|---|
19
19
  | v4.1.0 | 2.1.140 | 2026-05-13 | `claude_goal_set` / `claude_goal_clear` / `claude_goal_status` tools (wrap CLI 2.1.139 `/goal` slash command), `plugin_details` tool (wraps `claude plugin details`, 2.1.139), `pluginUrl` config (maps to `--plugin-url`, 2.1.129). Skipped items that are user-controlled via `--settings` (worktree.baseRef, autoMode.hard_deny, skillOverrides, sandbox.bwrapPath / socatPath, parentSettingsBehavior) or auto-set by the CLI (`CLAUDE_CODE_SESSION_ID`, `CLAUDE_CODE_DISABLE_ALTERNATE_SCREEN`, `CLAUDE_CODE_FORCE_SYNC_OUTPUT` — all TTY-only). Hook `args: string[]`, `continueOnBlock`, hook input `effort.level`, subagent `x-claude-code-agent-id` headers are CLI-internal — no wrapper change needed. |
20
20
  | v2.14.2 | 2.1.126 | 2026-05-04 | `bedrockServiceTier` (Bedrock service-tier env, 2.1.122), `project_purge` tool (wraps `claude project purge`, 2.1.126); skipped passive-only items (OTel numeric attr, `invocation_trigger`, `/v1/models` gateway discovery, PowerShell shell changes) |
@@ -20,26 +20,29 @@ The server exposes an OpenAI-compatible chat completions endpoint, enabling any
20
20
 
21
21
  **Endpoints:**
22
22
 
23
- | Endpoint | Method | Description |
24
- |----------|--------|-------------|
25
- | `/v1/chat/completions` | POST | Chat completions (streaming + non-streaming) |
26
- | `/v1/models` | GET | List available models |
23
+ | Endpoint | Method | Description |
24
+ | ---------------------- | ------ | -------------------------------------------- |
25
+ | `/v1/chat/completions` | POST | Chat completions (streaming + non-streaming) |
26
+ | `/v1/models` | GET | List available models |
27
27
 
28
28
  **Request format** (same as OpenAI):
29
+
29
30
  ```json
30
31
  {
31
32
  "model": "claude-sonnet-4-6",
32
- "messages": [{"role": "user", "content": "Hello!"}],
33
+ "messages": [{ "role": "user", "content": "Hello!" }],
33
34
  "stream": true
34
35
  }
35
36
  ```
36
37
 
37
38
  **Session routing:** Each conversation maps to a persistent session for prompt cache reuse. Session key resolved from (in priority order):
39
+
38
40
  1. `X-Session-Id` header
39
41
  2. `user` field in the request body
40
42
  3. Default singleton session
41
43
 
42
44
  **Model routing:** The `model` field auto-routes to the correct engine:
45
+
43
46
  - `claude-*`, `opus`, `sonnet`, `haiku` → Claude engine
44
47
  - `gpt-*` → Codex engine
45
48
  - `grok-*` → Grok engine
@@ -59,38 +62,38 @@ The server exposes an OpenAI-compatible chat completions endpoint, enabling any
59
62
  clawo session-start [name] [options]
60
63
  ```
61
64
 
62
- | Flag | Description |
63
- |------|-------------|
64
- | `-d, --cwd <dir>` | Working directory |
65
- | `-e, --engine <engine>` | Engine: `claude` (default), `codex`, `codex-app`, `agy`, `grok`, `opencode`, or `custom` |
66
- | `-m, --model <model>` | Model name or alias |
67
- | `--permission-mode <mode>` | `acceptEdits`, `plan`, `auto`, `bypassPermissions`, `manual`, `dontAsk` |
68
- | `--effort <level>` | `low`, `medium`, `high`, `max`, `auto` |
69
- | `--allowed-tools <tools>` | Comma-separated tool whitelist |
70
- | `--max-turns <n>` | Max agent loop turns |
71
- | `--max-budget <usd>` | API cost ceiling |
72
- | `--system-prompt <text>` | Replace system prompt |
73
- | `--append-system-prompt <text>` | Append to system prompt |
74
- | `--agents <json>` | Custom sub-agents JSON |
75
- | `--agent <name>` | Default agent |
76
- | `--bare` | No CLAUDE.md, no git context |
77
- | `-w, --worktree [name]` | Git worktree |
78
- | `--fallback-model <model>` | Fallback model |
79
- | `--json-schema <schema>` | JSON Schema for structured output |
80
- | `--mcp-config <paths>` | MCP config files (comma-separated) |
81
- | `--settings <path>` | Settings.json path |
82
- | `--skip-persistence` | Disable session persistence |
83
- | `--betas <headers>` | Beta headers (comma-separated) |
84
- | `--enable-agent-teams` | Enable agent teams |
85
- | `--include-hook-events` | Stream hook lifecycle events (PreToolUse/PostToolUse) |
86
- | `--permission-prompt-tool <tool>` | Delegate permission prompts to an MCP tool (non-interactive use) |
87
- | `--exclude-dynamic-system-prompt-sections` | Move cwd/env/git context to user message for better prompt cache hits (auto-enabled with `--bare`) |
88
- | `--debug <categories>` | Enable targeted debug output by category (e.g. `"api,mcp"`) |
89
- | `--debug-file <path>` | Write debug output to file |
90
- | `--from-pr <n>` | Resume a session linked to a GitHub PR number or URL |
91
- | `--channels <spec>` | MCP channel subscription (research preview) |
92
- | `--dangerously-load-development-channels <spec>` | Development MCP channel subscriptions (research preview) |
93
- | `ENABLE_PROMPT_CACHING_1H=1` (env var) | Enable 1-hour prompt cache TTL (auto-set with `--bare`) |
65
+ | Flag | Description |
66
+ | ------------------------------------------------ | -------------------------------------------------------------------------------------------------- |
67
+ | `-d, --cwd <dir>` | Working directory |
68
+ | `-e, --engine <engine>` | Engine: `claude` (default), `codex`, `codex-app`, `agy`, `grok`, `opencode`, or `custom` |
69
+ | `-m, --model <model>` | Model name or alias |
70
+ | `--permission-mode <mode>` | `acceptEdits`, `plan`, `auto`, `bypassPermissions`, `manual`, `dontAsk` |
71
+ | `--effort <level>` | `low`, `medium`, `high`, `max`, `auto` |
72
+ | `--allowed-tools <tools>` | Comma-separated tool whitelist |
73
+ | `--max-turns <n>` | Max agent loop turns |
74
+ | `--max-budget <usd>` | API cost ceiling |
75
+ | `--system-prompt <text>` | Replace system prompt |
76
+ | `--append-system-prompt <text>` | Append to system prompt |
77
+ | `--agents <json>` | Custom sub-agents JSON |
78
+ | `--agent <name>` | Default agent |
79
+ | `--bare` | No CLAUDE.md, no git context |
80
+ | `-w, --worktree [name]` | Git worktree |
81
+ | `--fallback-model <model>` | Fallback model |
82
+ | `--json-schema <schema>` | JSON Schema for structured output |
83
+ | `--mcp-config <paths>` | MCP config files (comma-separated) |
84
+ | `--settings <path>` | Settings.json path |
85
+ | `--skip-persistence` | Disable session persistence |
86
+ | `--betas <headers>` | Beta headers (comma-separated) |
87
+ | `--enable-agent-teams` | Enable agent teams |
88
+ | `--include-hook-events` | Stream hook lifecycle events (PreToolUse/PostToolUse) |
89
+ | `--permission-prompt-tool <tool>` | Delegate permission prompts to an MCP tool (non-interactive use) |
90
+ | `--exclude-dynamic-system-prompt-sections` | Move cwd/env/git context to user message for better prompt cache hits (auto-enabled with `--bare`) |
91
+ | `--debug <categories>` | Enable targeted debug output by category (e.g. `"api,mcp"`) |
92
+ | `--debug-file <path>` | Write debug output to file |
93
+ | `--from-pr <n>` | Resume a session linked to a GitHub PR number or URL |
94
+ | `--channels <spec>` | MCP channel subscription (research preview) |
95
+ | `--dangerously-load-development-channels <spec>` | Development MCP channel subscriptions (research preview) |
96
+ | `ENABLE_PROMPT_CACHING_1H=1` (env var) | Enable 1-hour prompt cache TTL (auto-set with `--bare`) |
94
97
 
95
98
  ### session-send
96
99
 
@@ -98,12 +101,12 @@ clawo session-start [name] [options]
98
101
  clawo session-send <name> <message> [options]
99
102
  ```
100
103
 
101
- | Flag | Description |
102
- |------|-------------|
103
- | `--effort <level>` | Override effort for this message |
104
- | `--plan` | Enable plan mode |
105
- | `-s, --stream` | Collect streaming chunks |
106
- | `-t, --timeout <ms>` | Timeout (default 300000) |
104
+ | Flag | Description |
105
+ | -------------------- | -------------------------------- |
106
+ | `--effort <level>` | Override effort for this message |
107
+ | `--plan` | Enable plan mode |
108
+ | `-s, --stream` | Collect streaming chunks |
109
+ | `-t, --timeout <ms>` | Timeout (default 300000) |
107
110
 
108
111
  ### session-stop
109
112
 
@@ -180,21 +183,60 @@ clawo session-team-send <name> <teammate> <message>
180
183
 
181
184
  The following tools are available through the OpenClaw plugin SDK and TypeScript API but do not have CLI commands. Use the SDK directly or call them via OpenClaw's tool system.
182
185
 
183
- | Tool | Description |
184
- |------|-------------|
185
- | `sessions_overview` | Aggregate dashboard of all active sessions |
186
- | `session_update_tools` | Hot-swap allowed/disallowed tools via `--resume` |
187
- | `session_switch_model` | Switch model mid-session via `--resume` |
188
- | `council_start` | Start multi-agent council with worktree isolation |
189
- | `council_status` | Poll council progress and agent responses |
190
- | `council_abort` | Abort a running council |
191
- | `council_inject` | Inject a message into the next council round |
192
- | `session_send_to` | Cross-session messaging (immediate or queued) |
193
- | `session_inbox` | Read inbox messages for a session |
194
- | `session_deliver_inbox` | Deliver queued messages to an idle session |
195
- | `ultraplan_start` | Start background Opus planning session |
196
- | `ultraplan_status` | Poll ultraplan progress |
197
- | `ultrareview_start` | Start fleet of parallel reviewer agents |
198
- | `ultrareview_status` | Poll ultrareview findings |
186
+ | Tool | Description |
187
+ | ----------------------- | ------------------------------------------------- |
188
+ | `sessions_overview` | Aggregate dashboard of all active sessions |
189
+ | `session_update_tools` | Hot-swap allowed/disallowed tools via `--resume` |
190
+ | `session_switch_model` | Switch model mid-session via `--resume` |
191
+ | `council_start` | Start multi-agent council with worktree isolation |
192
+ | `council_status` | Poll council progress and agent responses |
193
+ | `council_abort` | Abort a running council |
194
+ | `council_inject` | Inject a message into the next council round |
195
+ | `session_send_to` | Cross-session messaging (immediate or queued) |
196
+ | `session_inbox` | Read inbox messages for a session |
197
+ | `session_deliver_inbox` | Deliver queued messages to an idle session |
198
+ | `ultraplan_start` | Start background Opus planning session |
199
+ | `ultraplan_status` | Poll ultraplan progress |
200
+ | `ultrareview_start` | Start fleet of parallel reviewer agents |
201
+ | `ultrareview_status` | Poll ultrareview findings |
199
202
 
200
203
  See [Tools Reference](./tools.md) for full parameter documentation.
204
+
205
+ ## `clawo workflow` (6.0.0)
206
+
207
+ ```bash
208
+ clawo workflow list [--state <s>] [--workflow <name>] [--limit N] [--json]
209
+ clawo workflow show <runId> [--json]
210
+ clawo workflow resume <runId> # re-attach to a run whose process died
211
+ clawo workflow cancel <runId>
212
+ clawo workflow steer <runId> "<text>" # queue a correction for the next agent node
213
+ clawo workflow approve <runId> [reject]
214
+ ```
215
+
216
+ `list` prints one row per run with its outcome as `verified` / `REFUTED` /
217
+ `unchecked`. `unchecked` means no acceptance contract was declared — it is not a
218
+ failure.
219
+
220
+ Runs survive restarts, so `list` sees runs started by other processes and by
221
+ earlier sessions.
222
+
223
+ ## `clawo verify`
224
+
225
+ ```bash
226
+ clawo verify <runId> [--evidence <id>] [--json]
227
+ ```
228
+
229
+ Prints the evidence bundle: per-check pass/fail with the failing command and its
230
+ output tail, the base and head commits, and the files the run changed (created
231
+ files included).
232
+
233
+ ## `clawo runs` additions
234
+
235
+ ```bash
236
+ clawo runs --verified # only turns whose acceptance contract passed
237
+ clawo runs --refuted # only turns whose acceptance contract failed
238
+ ```
239
+
240
+ The table gains a `VERIFIED` column with three values: `yes`, `NO`, and `—` for
241
+ "no contract was declared, so nothing checked it". See
242
+ [`observability.md`](./observability.md).