@enderfga/claw-orchestrator 7.5.3 → 7.5.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (49) hide show
  1. package/README.md +21 -22
  2. package/configs/engines/README.md +7 -6
  3. package/dist/bin/cli.js +1 -1
  4. package/dist/bin/cli.js.map +1 -1
  5. package/dist/src/acp-server.d.ts +1 -1
  6. package/dist/src/acp-server.js +7 -5
  7. package/dist/src/acp-server.js.map +1 -1
  8. package/dist/src/autoloop/notify.d.ts +5 -7
  9. package/dist/src/autoloop/notify.js +21 -20
  10. package/dist/src/autoloop/notify.js.map +1 -1
  11. package/dist/src/embedded-server.js +7 -4
  12. package/dist/src/embedded-server.js.map +1 -1
  13. package/dist/src/fanout.d.ts +6 -0
  14. package/dist/src/fanout.js +1 -0
  15. package/dist/src/fanout.js.map +1 -1
  16. package/dist/src/index.js +19 -11
  17. package/dist/src/index.js.map +1 -1
  18. package/dist/src/kernel/nodes/fanout.js +1 -0
  19. package/dist/src/kernel/nodes/fanout.js.map +1 -1
  20. package/dist/src/kernel/types.d.ts +2 -0
  21. package/dist/src/kernel/types.js.map +1 -1
  22. package/dist/src/openai-compat.d.ts +2 -2
  23. package/dist/src/openai-compat.js +5 -2
  24. package/dist/src/openai-compat.js.map +1 -1
  25. package/dist/src/session-manager.js +14 -5
  26. package/dist/src/session-manager.js.map +1 -1
  27. package/dist/src/types.d.ts +2 -0
  28. package/openclaw.plugin.json +1 -1
  29. package/package.json +2 -2
  30. package/skills/SKILL.md +31 -32
  31. package/skills/references/acp.md +19 -36
  32. package/skills/references/autoloop.md +158 -180
  33. package/skills/references/claude-cli-tracking.md +27 -27
  34. package/skills/references/cli.md +62 -79
  35. package/skills/references/council.md +40 -63
  36. package/skills/references/dashboard.md +42 -55
  37. package/skills/references/getting-started.md +20 -14
  38. package/skills/references/inbox.md +6 -4
  39. package/skills/references/mcp.md +29 -24
  40. package/skills/references/multi-engine.md +105 -153
  41. package/skills/references/observability.md +42 -32
  42. package/skills/references/openai-compat.md +169 -303
  43. package/skills/references/sessions.md +20 -29
  44. package/skills/references/tools.md +62 -76
  45. package/skills/references/ultra.md +17 -16
  46. package/skills/references/ultraapp.md +59 -64
  47. package/skills/references/verification.md +29 -52
  48. package/skills/references/workflow.md +37 -104
  49. package/skills/ultraapp/SKILL.md +9 -10
@@ -2,8 +2,9 @@
2
2
 
3
3
  Three-agent autonomous iteration loop for a git workspace. You converse with
4
4
  the **Planner** to design a plan; on your approval, the Planner spawns the
5
- **Coder** + **Reviewer** subloop, monitors it, and pushes you (wechat →
6
- whatsapp → email fallback chain) only when something needs your attention.
5
+ **Coder** + **Reviewer** subloop, monitors it, and pushes you (WeChat →
6
+ WhatsApp → email fallback chain, see [Notification setup](#notification-setup))
7
+ only when something needs your attention.
7
8
 
8
9
  This page is the operator reference.
9
10
 
@@ -34,10 +35,10 @@ uses its own default model rather than receiving the Claude `opus` / `sonnet`
34
35
  defaults. Role instructions are included in-band for engines that do not expose a
35
36
  native system-prompt flag.
36
37
 
37
- Engines without native multi-turn conversation (Gemini and one-shot custom engines)
38
- spawn a fresh process per send with nothing to resume, so the dispatcher replays that
38
+ Engines without native multi-turn conversation (one-shot custom engines) spawn a
39
+ fresh process per send with nothing to resume, so the dispatcher replays that
39
40
  role's transcript in-band as a `<conversation_history>` block, oldest turns dropped
40
- past a character budget. Claude, Codex, Antigravity, Grok, OpenCode and Cursor each
41
+ past a character budget. Claude, Codex, Antigravity, Grok and OpenCode each
41
42
  resume their own conversation by id and get no replay — see
42
43
  `engineHasNativeConversation` in `types.ts`, which is the single source of truth for
43
44
  this and is checked with a two-turn recall test per engine.
@@ -52,21 +53,16 @@ Planner receives `permissionMode: 'manual'` and its `CustomEngineConfig` **must*
52
53
  map that mode to the CLI's read-only flag — if it cannot, the session refuses to
53
54
  start rather than silently running write-enabled.
54
55
 
55
- Antigravity's read-only boundary remains `--mode plan` on every Planner turn,
56
- including recovery. agy 1.1.26 can soft-deny a tool confirmation, emit a valid
57
- conversation id, leave `result.response` blank, and still exit 0. A missing or
58
- whitespace-only response is therefore a failed turn at the agy adapter boundary
59
- for every caller, not a successful empty Planner reply. If the private log for
60
- that exact turn contains agy's narrow `tool_confirmation_manager` soft-denial
61
- marker, the caller receives a fixed, sanitized diagnosis; native log content is
62
- never returned. The captured conversation id remains usable, so a later
63
- operator-initiated chat resumes with `--conversation`. Recovery neither retries
64
- the failed message automatically nor relaxes permissions. Because the failed
65
- reply never reaches the control parser, it cannot change `plan.md` or
66
- `goal.json`, spawn subagents, or emit an initial directive. agy 1.2.2 can instead
67
- return a non-empty successful reply for the same soft denial; the adapter emits
68
- the refused tool names through `SendResult.permissionDenials` while preserving
69
- that reply.
56
+ Antigravity's read-only boundary is `--mode plan` on every Planner turn,
57
+ including recovery. An empty or whitespace-only agy reply is a failed turn, not
58
+ an empty Planner reply. When agy's log for that turn shows a soft-denied tool
59
+ confirmation, the caller receives a fixed diagnosis (native log content is never
60
+ returned); when agy instead returns a non-empty reply for a soft denial, the
61
+ reply is kept and the refused tool names are reported in
62
+ `SendResult.permissionDenials`. The Planner's conversation id stays resumable,
63
+ the failed message is not retried automatically, and permissions are not
64
+ relaxed. A failed reply never reaches the control parser, so it cannot change
65
+ `plan.md` or `goal.json`, spawn subagents, or emit an initial directive.
70
66
 
71
67
  Coder and Reviewer engine/model choices can be overridden by the first successful
72
68
  `spawn_subagents`; later attempts to change an already-started role are rejected
@@ -91,19 +87,24 @@ failing with "Session not found".
91
87
  3. autoloop_chat { run_id, "go" } → Planner emits spawn_subagents
92
88
  4. Coder + Reviewer self-iterate → ledger writes per iter
93
89
  5. Planner pushes you on target_hit / regression / decision / stall
94
- 6. Run terminates on target hit, plan-defined max_iters, or your terminate.
90
+ 6. Run ends when the Planner emits terminate, the phase-error circuit trips,
91
+ the hard deadline passes, or you stop it. An expired activity lease
92
+ pauses the run instead.
95
93
  ```
96
94
 
95
+ Stopping at a target or at `max_iters` from `goal.json` is the Planner's
96
+ decision: the runtime does not evaluate `goal.json`.
97
+
97
98
  ## Timeout hierarchy and recoverable sends
98
99
 
99
100
  Autoloop has three independent start-time controls. Their bounds are inclusive,
100
101
  and omitting them retains the defaults:
101
102
 
102
- | Wire field | Runtime field | Default | Minimum | Maximum | Meaning |
103
- | ------------------------------ | ------------------------- | -------- | ------- | --------- | ------- |
104
- | `send_timeout_ms` | `sendTimeoutMs` | 600000 | 5000 | 7200000 | Wall-clock cap for one Planner, Coder, or Reviewer delivery |
105
- | `activity_lease_ms` | `activityLeaseMs` | 1800000 | 60000 | 7200000 | Inactivity lease, renewed only by validated user or agent progress |
106
- | `autoloop_hard_timeout_ms` | `autoloopHardTimeoutMs` | 86400000 | 600000 | 259200000 | Absolute run deadline, anchored to start and never renewed |
103
+ | Wire field | Runtime field | Default | Minimum | Maximum | Meaning |
104
+ | -------------------------- | ----------------------- | -------- | ------- | --------- | ------------------------------------------------------------------ |
105
+ | `send_timeout_ms` | `sendTimeoutMs` | 600000 | 5000 | 7200000 | Wall-clock cap for one Planner, Coder, or Reviewer delivery |
106
+ | `activity_lease_ms` | `activityLeaseMs` | 1800000 | 60000 | 7200000 | Inactivity lease, renewed only by validated user or agent progress |
107
+ | `autoloop_hard_timeout_ms` | `autoloopHardTimeoutMs` | 86400000 | 600000 | 259200000 | Absolute run deadline, anchored to start and never renewed |
107
108
 
108
109
  Timer checks and runner-generated bookkeeping do not renew the activity lease.
109
110
  The hard deadline cannot be extended by repeated activity and wins if it fires
@@ -127,44 +128,49 @@ stored spec, chat or iteration evidence, or any earlier audit bytes.
127
128
 
128
129
  ## Quick start
129
130
 
131
+ Over HTTP, against `clawo serve` (default `127.0.0.1:18796`):
132
+
130
133
  ```bash
131
- # Start a run (creates Planner session)
132
- curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_start \
133
- -H 'content-type: application/json' \
134
+ TOKEN=$(cat ~/.openclaw/server-token)
135
+
136
+ # Start a run (creates the Planner session)
137
+ curl -X POST http://127.0.0.1:18796/autoloop/new \
138
+ -H "Authorization: Bearer $TOKEN" -H 'content-type: application/json' \
134
139
  -d '{"run_id":"my-run","workspace":"/abs/path/to/workspace"}'
135
140
 
136
- # Chat with the Planner
137
- curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_chat \
138
- -H 'content-type: application/json' \
139
- -d '{"run_id":"my-run","text":"Read the workspace and design a plan to fix X"}'
141
+ # Chat with the Planner (202; the reply arrives on /events as planner_reply)
142
+ curl -X POST http://127.0.0.1:18796/autoloop/my-run/chat \
143
+ -H "Authorization: Bearer $TOKEN" -H 'content-type: application/json' \
144
+ -d '{"text":"Read the workspace and design a plan to fix X"}'
140
145
 
141
146
  # Inspect state
142
- curl http://127.0.0.1:18789/autoloop/my-run/state
147
+ curl http://127.0.0.1:18796/autoloop/my-run/state -H "Authorization: Bearer $TOKEN"
143
148
 
144
- # Live SSE stream (the 3-pane UI subscribes here)
145
- curl http://127.0.0.1:18789/autoloop/my-run/events
149
+ # Live SSE stream (the dashboard's 3-pane view subscribes here)
150
+ curl -N http://127.0.0.1:18796/autoloop/my-run/events -H "Authorization: Bearer $TOKEN"
151
+ ```
146
152
 
147
- # Reset Coder if it drifts (lazy; eager_restart=true to start a fresh session immediately)
148
- curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_reset_agent \
149
- -H 'content-type: application/json' \
150
- -d '{"run_id":"my-run","agent":"coder","eager_restart":true}'
153
+ Resetting an agent and stopping a run have no HTTP route; call the tools
154
+ (plugin or MCP):
151
155
 
152
- # Stop
153
- curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_stop \
154
- -H 'content-type: application/json' \
155
- -d '{"run_id":"my-run","reason":"done"}'
156
+ ```jsonc
157
+ // Reset the Coder if it drifts (lazy by default; eager_restart starts a fresh session now)
158
+ autoloop_reset_agent({ "run_id": "my-run", "agent": "coder", "eager_restart": true })
159
+
160
+ // Stop
161
+ autoloop_stop({ "run_id": "my-run", "reason": "done" })
156
162
  ```
157
163
 
158
164
  ## Plugin tools
159
165
 
160
- | Tool | Args | What |
161
- | ---------------------- | ----------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
166
+ | Tool | Args | What |
167
+ | ---------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
162
168
  | `autoloop_start` | `run_id`, `workspace`, per-role `*_engine?`, `*_model?`, `*_custom_engine?`, `send_timeout_ms?`, `activity_lease_ms?`, `autoloop_hard_timeout_ms?` | Start a run; launches Planner and stores Coder/Reviewer defaults and timeout controls. Each `custom` role requires its matching config. |
163
- | `autoloop_chat` | `run_id`, `text` | Send a chat message to the Planner; returns the Planner's reply. |
164
- | `autoloop_status` | `run_id` | Current state (status, iter, push count, subagents_spawned). |
165
- | `autoloop_list` | — | All active runs in this manager process. |
166
- | `autoloop_stop` | `run_id`, `reason?` | Terminate; stops Planner / Coder / Reviewer. |
167
- | `autoloop_reset_agent` | `run_id`, `agent` ('planner' / 'coder' / 'reviewer'), `force?`, `eager_restart?` | Reset one subagent. Planner reset requires `force: true`. |
169
+ | `autoloop_chat` | `run_id`, `text` | Send a chat message to the Planner; returns the Planner's reply. |
170
+ | `autoloop_status` | `run_id` | Current state (status, iter, push count, subagents_spawned). |
171
+ | `autoloop_list` | — | All autoloop runs in the run store, live or not. |
172
+ | `autoloop_stop` | `run_id`, `reason?` | Terminate; stops Planner / Coder / Reviewer. |
173
+ | `autoloop_reset_agent` | `run_id`, `agent` ('planner' / 'coder' / 'reviewer'), `force?`, `eager_restart?` | Reset one subagent. Planner reset requires `force: true`. |
168
174
 
169
175
  ## Planner-emitted control tools
170
176
 
@@ -198,18 +204,17 @@ removes partial files, and suppresses both spawn and directives.
198
204
 
199
205
  ### Custom engines and resume
200
206
 
201
- Custom engine configs are accepted only by `autoloop_start` (or the HTTP resume
202
- body), never through Planner output. This keeps config fields such as `env` and
203
- static CLI arguments out of the Planner transcript and `decisions.jsonl`.
204
- The central registry persists only each role's engine and model, including the
205
- effective Coder/Reviewer selection after a successful spawn. Resume leaves the
206
- prior append-only row untouched until startup succeeds, so a transient CLI
207
- failure cannot erase the run. When resuming a run that uses `custom`, provide
208
- the matching `planner_custom_engine`, `coder_custom_engine`, or
209
- `reviewer_custom_engine` again; otherwise resume fails with a clear
207
+ Custom engine configs are accepted only by `autoloop_start` (and, on resume, by
208
+ `SessionManager.autoloopResume()` or by reference in the HTTP resume body), never
209
+ through Planner output. This keeps config fields such as `env` and static CLI
210
+ arguments out of the Planner transcript and `decisions.jsonl`. The run record
211
+ stores only each role's engine and model, including the effective Coder/Reviewer
212
+ selection after a successful spawn. When resuming a run that uses `custom`,
213
+ supply the matching config again (over HTTP, as a `*CustomEngineRef`, see
214
+ [Backend HTTP / SSE](#backend-http--sse)); otherwise resume fails with a clear
210
215
  configuration error rather than silently switching to Claude. Custom config
211
- shape is validated at runtime, while its `env` and static CLI arguments remain
212
- out of registry and audit records. See [`multi-engine.md`](./multi-engine.md)
216
+ shape is validated at runtime, while its `env` and static CLI arguments stay
217
+ out of the run record and audit logs. See [`multi-engine.md`](./multi-engine.md)
213
218
  for the `CustomEngineConfig` shape.
214
219
 
215
220
  ## Default push policy
@@ -218,34 +223,53 @@ for the `CustomEngineConfig` shape.
218
223
  | ---------------------- | ----------------------------------------------------- |
219
224
  | on_start | info / wechat ("loop started, will notify on issues") |
220
225
  | on_iter_done_ok | silent |
221
- | on_target_hit | info / both (webchat + wechat) |
226
+ | on_target_hit | info / both |
222
227
  | on_metric_regression_2 | warn / both |
223
228
  | on_reviewer_reject_2 | warn / both |
224
229
  | on_phase_error | error / both |
225
230
  | on_stall_30min | warn / wechat |
226
231
  | on_decision_needed | decision / both |
227
232
 
228
- 5-minute dedup on (level, summary) prevents duplicate pushes from the same
229
- event. Channel chain: `auto` walks wechat → whatsapp → email; `wechat` /
230
- `webchat` / `email` route directly; `both` does webchat (if session known)
233
+ The runtime fires five of these itself: `on_stall_30min` (30 minutes without
234
+ activity), `on_metric_regression_2`, `on_reviewer_reject_2` (two in a row),
235
+ `on_phase_error`, and `on_target_hit` (only when an acceptance contract passes,
236
+ see [Acceptance contracts](#acceptance-contracts)). `on_start`,
237
+ `on_iter_done_ok` and `on_decision_needed` are policy entries for the Planner's
238
+ own `notify_user` calls.
239
+
240
+ A 5-minute dedup on (level, summary) prevents duplicate pushes from the same
241
+ event. Channels: `auto` and `both` walk WeChat → WhatsApp → email and stop at
242
+ the first that succeeds; `wechat` and `email` go to that channel only;
243
+ `webchat` is a no-op (see [Known limitations](#known-limitations)).
244
+ **`on_phase_error` and `on_decision_needed` cannot be set to `silent: true`**
245
+ by the Planner: `update_push_policy` strips the flag and records the attempt in
246
+ `decisions.jsonl`.
231
247
 
232
- - wechat fallback chain. **`on_phase_error` and `on_decision_needed` cannot
233
- be set to `silent: true`** by Planner — `update_push_policy` strips the flag
234
- and records the attempt in `decisions.jsonl` (these channels are the
235
- operator's lifeline; they stay loud).
248
+ ## Notification setup
249
+
250
+ - **WeChat** needs `AUTOLOOP_WECHAT_RECIPIENT` and `AUTOLOOP_WECHAT_ACCOUNT`.
251
+ - **WhatsApp** needs `AUTOLOOP_WHATSAPP_RECIPIENT`.
252
+ - Both send through the `openclaw` CLI, which must be on `PATH`.
253
+ - **Email** is sent by running `bash "$AUTOLOOP_EMAIL_SCRIPT" -s "<subject>"`
254
+ with the message body on stdin.
255
+
256
+ Any channel whose variables are unset is skipped silently. With none set,
257
+ pushes are recorded in `push_log.jsonl` only.
236
258
 
237
259
  ## Auto-compact
238
260
 
239
261
  Each agent's context is monitored after every turn. When `getStats().contextPercent`
240
262
  crosses the per-agent threshold the dispatcher invokes `/compact` with a
241
263
  role-tuned hint (`compactSummaryFor`). Defaults: Planner 80 %, Coder 70 %,
242
- Reviewer 70 %. Override per run via `compactThresholds`. A 30 s debounce
264
+ Reviewer 70 %. The dispatcher's `compactThresholds` option overrides them; it
265
+ is library-level only and not exposed through `autoloop_start` or the HTTP
266
+ API. A 30 s debounce
243
267
  prevents re-fire while post-compact stats settle. Events: `compact` is
244
268
  emitted on the dispatcher EventEmitter AND appended to `decisions.jsonl`.
245
269
 
246
270
  One-shot engines (`codex`, `agy`, `grok`, `opencode`) cannot compact — their
247
271
  CLIs expose no such command. The threshold is still meaningful there because
248
- `contextPercent` now tracks real occupancy, but crossing it cannot free space:
272
+ `contextPercent` tracks real occupancy, but crossing it cannot free space:
249
273
  the session emits a single warning on its log channel the first time compaction
250
274
  is requested, then the thread keeps growing until the CLI refuses the request.
251
275
  Treat that warning as the signal to start a fresh session.
@@ -262,8 +286,9 @@ consecutive `phase_error`s and:
262
286
  `decision`-level push and an automatic `terminate { reason:
263
287
  'phase_error_circuit' }`.
264
288
 
265
- A successful (non-error) `iter_done` resets the counter. Override the
266
- threshold via `AutoloopConfig.phaseErrorCircuit`.
289
+ A successful (non-error) `iter_done` resets the counter. The threshold is
290
+ `AutoloopConfig.phaseErrorCircuit`, which is library-level only and not exposed
291
+ through `autoloop_start` or the HTTP API.
267
292
 
268
293
  ## Reviewer frozen memory
269
294
 
@@ -299,6 +324,8 @@ JSONL, one entry per line, ts-prefixed.
299
324
  ├── goal.json # Planner-authored, git-committed
300
325
  ├── push_log.jsonl # every notify_user attempt + channel used
301
326
  ├── decisions.jsonl # runner / dispatcher audit trail (see above)
327
+ ├── chat.jsonl # Planner-pane conversation, replayed by /chat_history
328
+ ├── evidence/iter-<n>/ # acceptance-contract bundle, when a contract is configured
302
329
  ├── reviewer_sandbox/ # Reviewer cwd; restaged per iter
303
330
  │ ├── plan.md # copy
304
331
  │ ├── goal.json # copy
@@ -309,7 +336,7 @@ JSONL, one entry per line, ts-prefixed.
309
336
  └── iter/<n>/
310
337
  ├── directive.json # Planner → Coder (schema_version: 1)
311
338
  ├── eval_output.json # what Coder reported (schema_version: 1)
312
- ├── diff.patch # git diff of the iter
339
+ ├── diff.patch # git diff of the iter, created files included
313
340
  ├── verdict.json # Reviewer decision + audit notes (schema_version: 1)
314
341
  └── coder_summary.txt
315
342
  ```
@@ -320,25 +347,28 @@ inside an iter** (pre-commit hook reject, signing key missing, …) the
320
347
  dispatcher emits a `phase_error` instead of writing `iter_artifacts`, so
321
348
  the failure is visible to the runner and counts toward the circuit.
322
349
 
350
+ `files_changed` in the iteration artifacts is taken from git, never from the
351
+ Coder's own report.
352
+
323
353
  Every JSON artifact in the ledger carries a `schema_version` field (currently
324
354
  `1`) to make future migrations explicit.
325
355
 
326
356
  ## Backend HTTP / SSE
327
357
 
328
- | Endpoint | Returns |
329
- | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
330
- | `GET /autoloop/list` | `{ ok, runs: AutoloopState[] }` |
331
- | `POST /autoloop/new` | `{ ok, run_id, planner_session }` — body `{ workspace, run_id?, planner_engine?, planner_model?, planner_custom_engine?, coder_engine?, coder_model?, coder_custom_engine?, reviewer_engine?, reviewer_model?, reviewer_custom_engine?, send_timeout_ms?, activity_lease_ms?, autoloop_hard_timeout_ms? }`. Timeout fields use the defaults and inclusive bounds documented above; malformed or out-of-range values return 400 before a run starts. |
332
- | `GET /autoloop/<id>/state` | `{ ok, state: AutoloopState, live }` — `live` is `true` only when the run is running in this process; also returns a `terminated`-state stub reconstructed from the registry for runs that aren't in this process's memory, so the dashboard can open historical runs without 404'ing. |
333
- | `GET /autoloop/<id>/push_log` | `{ ok, entries: PushLogEntry[] }` — served from the ledger via `autoloopStatus`, so historical runs work the same as live ones. |
334
- | `GET /autoloop/<id>/chat_history` | `{ ok, entries: ChatEntry[] }` — replays `<ledger>/chat.jsonl`. The dashboard fetches this when opening a run so the Planner-pane conversation survives a page refresh / cross-process / re-opening a terminated run. Returns `[]` when the file doesn't exist (e.g. runs that predate the chat-history feature). |
335
- | `GET /autoloop/<id>/events` | SSE: `snapshot` / `message` / `state` / `push` / `iter_done` / `planner_reply` / `planner_error` / `coder_reply` / `reviewer_reply` / `terminated`. For runs that are NOT in this process's memory (terminated, or live in another process), the endpoint emits a single-shot `snapshot` + `terminated` then closes — the dashboard's existing handlers render history without hanging. A run still in memory that has already reached `terminated` or `crashed` gets the same single-shot pair instead of an open stream that would never receive another event. Every such stream sets `retry: 864000000`, so an `EventSource` does not keep reconnecting to a stream that can only end again. |
336
- | `POST /autoloop/<id>/chat` | **202** `{ ok, queued: true }` — body `{ text }`. Fire-and-forget: the Planner's reply streams back via the `/events` SSE channel as a `planner_reply` event (or `planner_error` on failure); the HTTP response intentionally does NOT wait for it, because first-contact replies routinely exceed reverse-proxy idle limits (e.g. Cloudflare Tunnel cuts at ~100s → 524). 400 on empty text. 404 when the run is not in this process's memory: if the store still holds it, the error says so and names `POST /autoloop/<id>/resume`; an unknown or malformed id is plain `not found`. The MCP `autoloop_chat` tool path keeps the synchronous await-and-return-reply semantics (it runs in-process). |
337
- | `GET /autoloop/<id>/resume-requirements` | `{ ok, runId, rolesNeedingCustomEngine }` — the roles whose engine was `custom`, so a caller knows which secret references a resume needs. Role names only; nothing sensitive. 404 when there is no such run. |
358
+ | Endpoint | Returns |
359
+ | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
360
+ | `GET /autoloop/list` | `{ ok, runs: AutoloopState[] }` |
361
+ | `POST /autoloop/new` | `{ ok, run_id, planner_session }` — body `{ workspace, run_id?, planner_engine?, planner_model?, coder_engine?, coder_model?, reviewer_engine?, reviewer_model?, send_timeout_ms?, activity_lease_ms?, autoloop_hard_timeout_ms? }`. Timeout fields use the defaults and inclusive bounds documented above; malformed or out-of-range values return 400 before a run starts. |
362
+ | `GET /autoloop/<id>/state` | `{ ok, state: AutoloopState, live }` — `live` is `true` only when the run is running in this process. For a run that is not live here, `state` is the last state recorded in the run store, so historical runs open with their real iteration count. 404 when there is no such run. |
363
+ | `GET /autoloop/<id>/push_log` | `{ ok, entries: PushLogEntry[] }` — served from the ledger via `autoloopStatus`, so historical runs work the same as live ones. |
364
+ | `GET /autoloop/<id>/chat_history` | `{ ok, entries: ChatEntry[] }` — replays `<ledger>/chat.jsonl`. The dashboard fetches this when opening a run so the Planner-pane conversation survives a page refresh / cross-process / re-opening a terminated run. Returns `[]` when the file doesn't exist. |
365
+ | `GET /autoloop/<id>/events` | SSE: `snapshot` / `message` / `state` / `push` / `iter_done` / `planner_reply` / `planner_error` / `coder_reply` / `reviewer_reply` / `terminated`. For runs that are NOT in this process's memory (terminated, or live in another process), the endpoint emits a single-shot `snapshot` + `terminated` then closes — the dashboard's existing handlers render history without hanging. A run still in memory that has already reached `terminated` or `crashed` gets the same single-shot pair instead of an open stream that would never receive another event. Every such stream sets `retry: 864000000`, so an `EventSource` does not keep reconnecting to a stream that can only end again. |
366
+ | `POST /autoloop/<id>/chat` | **202** `{ ok, queued: true }` — body `{ text }`. Fire-and-forget: the Planner's reply streams back via the `/events` SSE channel as a `planner_reply` event (or `planner_error` on failure); the HTTP response intentionally does NOT wait for it, because first-contact replies routinely exceed reverse-proxy idle limits (e.g. Cloudflare Tunnel cuts at ~100s → 524). 400 on empty text. 404 when the run is not in this process's memory: if the store still holds it, the error says so and names `POST /autoloop/<id>/resume`; an unknown or malformed id is plain `not found`. The MCP `autoloop_chat` tool path keeps the synchronous await-and-return-reply semantics (it runs in-process). |
367
+ | `GET /autoloop/<id>/resume-requirements` | `{ ok, runId, rolesNeedingCustomEngine }` — the roles whose engine was `custom`, so a caller knows which secret references a resume needs. Role names only; nothing sensitive. 404 when there is no such run. |
338
368
  | `POST /autoloop/<id>/resume` | `{ ok, state }` — restore the role engine/model choices from the run's spec and re-create dispatcher + runner. For recoverable send timeouts, body fields `send_timeout_ms` and `pending_dispatch_id` apply the increase-only migration described above; `allow_decrease`, lease overrides, and hard-cap overrides are rejected. A custom-engine config is never persisted and is never accepted over HTTP, so a role using `custom` is re-supplied by **reference**: `plannerCustomEngineRef` / `coderCustomEngineRef` / `reviewerCustomEngineRef` name an environment variable `CLAWO_CUSTOM_ENGINE_<NAME>` on the orchestrator host, which the server reads and resolves. The name is not sensitive, the value never crosses the wire, and an unknown name is an error rather than a silent start without credentials. Existing engine-specific conversation resume behavior is reused where supported; `chat.jsonl` remains the visual history fallback. 404 when there is no such run. |
339
- | `POST /autoloop/<id>/delete` | `{ ok }` — stops the runner if still alive, scrubs the row from `~/.claw-orchestrator/autoloop-registry.jsonl`, and purges `persistedSessions` so the run cannot be `/resume`'d back. The ledger directory under `<workspace>/tasks/<run_id>/` is kept on disk. 404 if the run was not present in either memory or the registry. |
369
+ | `POST /autoloop/<id>/delete` | `{ ok }` — stops the loop if still live, deletes the run record from the run store, and purges the role sessions' persisted resume ids so the run cannot be resumed. The ledger directory under `<workspace>/tasks/<run_id>/` is kept on disk. 404 when there is no such run. |
340
370
 
341
- The 3-pane UI consumes these endpoints:
371
+ The dashboard's 3-pane autoloop view (`/dashboard`) consumes these endpoints:
342
372
 
343
373
  - **Left**: Planner chat (subscribes to `planner_reply`)
344
374
  - **Center**: Coder activity (`coder_reply` + `iter_done`)
@@ -346,8 +376,6 @@ The 3-pane UI consumes these endpoints:
346
376
  - **Top bar**: state (status / iter / metric)
347
377
  - **Bottom**: push_log
348
378
 
349
- The UI itself ships in a separate cross-repo PR.
350
-
351
379
  ## `goal.json` shape
352
380
 
353
381
  The Planner authors goal.json based on your conversation. There is no
@@ -382,106 +410,56 @@ The Planner will riff on this shape during your chat and ask if it's right.
382
410
  - ✅ Reviewer accumulates "fakery patterns I've seen" in `reviewer_memory.md` (persists across iters).
383
411
  - ✅ Reviewer defaults to `hold` under uncertainty; only `advance` after independent verification.
384
412
 
385
- ## Smoke test
386
-
387
- `scripts/smoke-autoloop.ts` runs a buggy `add_two` scenario end-to-end with
388
- Opus Planner + Sonnet × 2. Validates plan.md / goal.json commit, spawn,
389
- iter 0 ledger artifacts (`directive` + `eval_output` + `diff.patch` +
390
- `verdict`), and termination on `target_hit`. Cost ~$1-3, wall-clock
391
- ~5-15 min. Run with `npx tsx scripts/smoke-autoloop.ts` (requires
392
- `~/.claude/settings.json` to have your auth env).
393
-
394
- ## Known limitations
395
-
396
- - **`webchat` channel is a no-op** — `notifyUserFallbackChain` does not yet
397
- carry a webchat session id at the run level, so `channel: 'webchat'`
398
- always returns `channel_used: 'none'`. Use `auto` / `wechat` / `email`
399
- until the inbound route lands.
400
- - **One-way push.** WeChat → Planner inbound replies are not yet wired (would
401
- need an openclaw-gateway tmux-passthrough route). Reply via webchat /
402
- `autoloop_chat`.
403
- - **No webchat UI yet.** Backend SSE is shipped; the UI is a separate
404
- cross-repo PR in ChatGPT-Next-Web.
405
- - **No fork / population mode.** Single linear iter trajectory per run.
406
- - **Cross-run knowledge isolated.** Each run's `reviewer_memory.md` and
407
- `coder_notes.md` live in that run's ledger; no shared meta-store yet.
408
- - **No cost / wall-clock budget cap.** Only `phaseErrorCircuit` + Reviewer
409
- hold/reject streaks bound the run; a steady-but-pointless ratchet could
410
- run for days. Set `max_iters` in `goal.json` to bound iter count.
411
- - **Run state in memory.** SessionManager restart drops the live `autoloops`
412
- map; the on-disk ledger survives but cannot resume a running state.
413
- - **Multi-run / same workspace** races on `git index.lock`. Run separate
414
- workspaces (or git worktrees) for concurrent runs.
413
+ ## Acceptance contracts
415
414
 
416
- ## Acceptance contracts (6.0.0)
415
+ The Reviewer's sandbox holds the iteration's artifacts (`directive.json`,
416
+ `diff.patch`, `eval_output.json`, `coder_summary.txt`, `plan.md`, `goal.json`,
417
+ the prior verdict) but no code or evaluator, so its `advance` is a judgement of
418
+ the Coder's report rather than a measurement.
417
419
 
418
- The Reviewer's `advance` was the only thing standing between an iteration and
419
- "done", and it could not do the job it was given. Its prompt tells it to
420
- "re-derive the metric independently if the sandbox has the bits to do so", but
421
- `stageReviewSandbox` copies in the iteration's artifacts — `directive.json`,
422
- `diff.patch`, `eval_output.json`, `coder_summary.txt` — plus `plan.md`,
423
- `goal.json`, and the prior verdict. No code, no evaluator. The verdict was
424
- therefore a reading of the Coder's own report, and `eval_output` was literally
425
- whatever the Coder passed to a tool call.
420
+ An acceptance contract closes that gap. When one is configured, an `advance`
421
+ stands only if the checks pass against the workspace: otherwise the verdict is
422
+ rewritten to `hold`, the failing checks are appended to `audit_notes`, and the
423
+ bundle is written to `<ledger>/evidence/iter-<n>/`. A passing contract fires
424
+ `on_target_hit`. Without a contract the Reviewer's verdict is used as-is.
426
425
 
427
- Pass a contract at autoloop start and an `advance` is held unless the checks
428
- pass. The verdict is rewritten to `hold`, the reason is appended to
429
- `audit_notes`, and the evidence bundle lands at
430
- `<ledger>/iter/<n>/evidence/`. Without a contract nothing changes.
426
+ **The contract is library-level only.** It is set on the dispatcher config
427
+ (`contract` in `ClaudeAgentDispatcherConfig`) and is not yet exposed through
428
+ `autoloop_start` or the HTTP API, so a run started from the tool or
429
+ `POST /autoloop/new` has no contract.
431
430
 
432
- ### `on_target_hit` now fires
431
+ ## Lifecycle and resume
433
432
 
434
- That push-policy key was declared in `types.ts`, given a default, and whitelisted
435
- for runtime updates — and had **zero firing sites** anywhere in the codebase.
436
- Autoloop had four ways to notice it was failing (stall, metric regression,
437
- reviewer rejections, phase-error circuit) and no way to notice it had succeeded,
438
- because a Reviewer verdict is not a measurement. An acceptance contract is, so
439
- `on_target_hit` fires when one passes.
433
+ An autoloop is a kernel run whose single node holds the loop for as long as it
434
+ lives. The run record holds the last state the loop published and the engines
435
+ `spawn_subagents` chose, so `autoloop_status` and `GET /autoloop/<id>/state`
436
+ show a run's real state after it stops or the process restarts.
440
437
 
441
- ### `diff.patch` sees created files
438
+ Resume is explicit. `POST /autoloop/<id>/resume` (or
439
+ `SessionManager.autoloopResume()`) restarts a run from its stored spec.
440
+ Custom-engine configs are the one thing the spec does not carry (they can hold
441
+ secrets), so a resume must be given them again. Cancelling a run stops all three
442
+ agents, the same as `autoloop_stop`.
442
443
 
443
- The per-iteration patch was captured with a bare `git diff`, which lists tracked
444
- modifications only. A file the Coder _created_ appeared in neither the patch nor
445
- the `--name-only` fallback, while the `git add -A` two lines later committed it —
446
- so the Reviewer audited a picture that structurally could not show new files.
447
- The capture now covers tracked changes ∪ untracked files.
444
+ ## Known limitations
448
445
 
449
- `files_changed` is also taken from git unconditionally. It previously preferred
450
- the Coder's own `files_changed` whenever the Coder supplied one, despite the
451
- comment above it saying the claim was not trusted.
446
+ - **`webchat` channel is a no-op.** No webchat session id is carried at the run
447
+ level, so `channel: 'webchat'` always returns `channel_used: 'none'`. Use
448
+ `auto` / `wechat` / `email`.
449
+ - **One-way push.** Replies to a push are not routed back to the Planner;
450
+ answer with `autoloop_chat`.
451
+ - **No fork / population mode.** Single linear iter trajectory per run.
452
+ - **Cross-run knowledge isolated.** Each run's `reviewer_memory.md` and
453
+ `coder_notes.md` live in that run's ledger; there is no shared store.
454
+ - **No cost cap.** `maxBudgetUsd` is not exposed for autoloop. Wall-clock time
455
+ is bounded by `autoloop_hard_timeout_ms` (default 24 h); the iteration count
456
+ is bounded only by the Planner honouring `max_iters` in `goal.json`.
457
+ - **Resume is not automatic.** After a restart a run is not live in the new
458
+ process until it is resumed, and a send that was in flight is not retried.
459
+ - **Multi-run / same workspace** races on `git index.lock`. Run separate
460
+ workspaces (or git worktrees) for concurrent runs.
452
461
 
453
462
  ## Related
454
463
 
455
464
  - [`verification.md`](./verification.md) — contracts, checks, evidence
456
-
457
- ## Lifecycle moved to the kernel (6.0.0)
458
-
459
- An autoloop is a kernel run whose single node holds the loop for as long as it
460
- lives. Tool signatures are unchanged.
461
-
462
- What went away: the `autoloops` map; `~/.claw-orchestrator/autoloop-registry.jsonl`
463
- with its four bespoke helpers (append, remove-then-append upsert, reverse-scan
464
- dedup, rewrite-via-tmp-file); and the two `Set`s — `_startingAutoloops` and
465
- `_deletingAutoloops` — that existed only because a start and a delete could race
466
- each other over that shared map.
467
-
468
- Two things get better rather than merely moving:
469
-
470
- - **`autoloop_status` on a run that is not live in this process** used to return
471
- an all-zero stub labelled `reconstructed from registry — not in current process
472
- memory`: iter 0, no metrics, no error history, because the registry only ever
473
- held identity. The record holds the last state the loop published, so a
474
- historical run opens with its real iteration count.
475
- - **The engines `spawn_subagents` actually chose** land on the run record
476
- alongside the rest of its state, instead of in a parallel file with its own
477
- lifecycle.
478
-
479
- `autoloop_resume` restarts a terminated run from the stored spec — the immutable
480
- record of how it was started — rather than from a registry row whose older
481
- versions omitted the engine fields entirely. Custom-engine configs are the one
482
- thing the spec does not carry (they can hold secrets), so a resume must be given
483
- them again.
484
-
485
- Cancelling a run now tears the loop down the way a stop does. It previously left
486
- the three persistent agents running and their session names claimed, which
487
- surfaced much later as `session name already in use`.
465
+ - [`workflow.md`](./workflow.md) — the kernel that stores and resumes runs