@enderfga/claw-orchestrator 7.5.2 → 7.5.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. package/README.md +26 -27
  2. package/configs/engines/README.md +7 -6
  3. package/dist/bin/cli.js +1 -1
  4. package/dist/bin/cli.js.map +1 -1
  5. package/dist/src/acp-server.d.ts +1 -1
  6. package/dist/src/acp-server.js +7 -5
  7. package/dist/src/acp-server.js.map +1 -1
  8. package/dist/src/autoloop/dispatcher.js +3 -3
  9. package/dist/src/autoloop/dispatcher.js.map +1 -1
  10. package/dist/src/autoloop/notify.d.ts +5 -7
  11. package/dist/src/autoloop/notify.js +21 -20
  12. package/dist/src/autoloop/notify.js.map +1 -1
  13. package/dist/src/base-oneshot-session.js +20 -1
  14. package/dist/src/base-oneshot-session.js.map +1 -1
  15. package/dist/src/dashboard/index.html +94 -15
  16. package/dist/src/embedded-server.js +40 -7
  17. package/dist/src/embedded-server.js.map +1 -1
  18. package/dist/src/fanout.d.ts +6 -0
  19. package/dist/src/fanout.js +1 -0
  20. package/dist/src/fanout.js.map +1 -1
  21. package/dist/src/index.js +19 -11
  22. package/dist/src/index.js.map +1 -1
  23. package/dist/src/kernel/nodes/fanout.js +1 -0
  24. package/dist/src/kernel/nodes/fanout.js.map +1 -1
  25. package/dist/src/kernel/types.d.ts +2 -0
  26. package/dist/src/kernel/types.js.map +1 -1
  27. package/dist/src/models.js +36 -7
  28. package/dist/src/models.js.map +1 -1
  29. package/dist/src/openai-compat.d.ts +2 -2
  30. package/dist/src/openai-compat.js +5 -2
  31. package/dist/src/openai-compat.js.map +1 -1
  32. package/dist/src/persistent-agy-session.js +6 -1
  33. package/dist/src/persistent-agy-session.js.map +1 -1
  34. package/dist/src/session-manager.d.ts +1 -0
  35. package/dist/src/session-manager.js +17 -5
  36. package/dist/src/session-manager.js.map +1 -1
  37. package/dist/src/types.d.ts +2 -0
  38. package/openclaw.plugin.json +1 -1
  39. package/package.json +2 -2
  40. package/skills/SKILL.md +31 -32
  41. package/skills/references/acp.md +19 -36
  42. package/skills/references/autoloop.md +163 -180
  43. package/skills/references/claude-cli-tracking.md +28 -27
  44. package/skills/references/cli.md +62 -79
  45. package/skills/references/council.md +40 -63
  46. package/skills/references/dashboard.md +42 -55
  47. package/skills/references/getting-started.md +21 -15
  48. package/skills/references/inbox.md +6 -4
  49. package/skills/references/mcp.md +29 -24
  50. package/skills/references/multi-engine.md +105 -153
  51. package/skills/references/observability.md +42 -32
  52. package/skills/references/openai-compat.md +169 -303
  53. package/skills/references/sessions.md +20 -29
  54. package/skills/references/tools.md +62 -76
  55. package/skills/references/ultra.md +17 -16
  56. package/skills/references/ultraapp.md +59 -64
  57. package/skills/references/verification.md +29 -52
  58. package/skills/references/workflow.md +37 -104
  59. package/skills/ultraapp/SKILL.md +9 -10
@@ -2,8 +2,9 @@
2
2
 
3
3
  Three-agent autonomous iteration loop for a git workspace. You converse with
4
4
  the **Planner** to design a plan; on your approval, the Planner spawns the
5
- **Coder** + **Reviewer** subloop, monitors it, and pushes you (wechat →
6
- whatsapp → email fallback chain) only when something needs your attention.
5
+ **Coder** + **Reviewer** subloop, monitors it, and pushes you (WeChat →
6
+ WhatsApp → email fallback chain, see [Notification setup](#notification-setup))
7
+ only when something needs your attention.
7
8
 
8
9
  This page is the operator reference.
9
10
 
@@ -34,10 +35,10 @@ uses its own default model rather than receiving the Claude `opus` / `sonnet`
34
35
  defaults. Role instructions are included in-band for engines that do not expose a
35
36
  native system-prompt flag.
36
37
 
37
- Engines without native multi-turn conversation (Gemini and one-shot custom engines)
38
- spawn a fresh process per send with nothing to resume, so the dispatcher replays that
38
+ Engines without native multi-turn conversation (one-shot custom engines) spawn a
39
+ fresh process per send with nothing to resume, so the dispatcher replays that
39
40
  role's transcript in-band as a `<conversation_history>` block, oldest turns dropped
40
- past a character budget. Claude, Codex, Antigravity, Grok, OpenCode and Cursor each
41
+ past a character budget. Claude, Codex, Antigravity, Grok and OpenCode each
41
42
  resume their own conversation by id and get no replay — see
42
43
  `engineHasNativeConversation` in `types.ts`, which is the single source of truth for
43
44
  this and is checked with a two-turn recall test per engine.
@@ -52,21 +53,16 @@ Planner receives `permissionMode: 'manual'` and its `CustomEngineConfig` **must*
52
53
  map that mode to the CLI's read-only flag — if it cannot, the session refuses to
53
54
  start rather than silently running write-enabled.
54
55
 
55
- Antigravity's read-only boundary remains `--mode plan` on every Planner turn,
56
- including recovery. agy 1.1.26 can soft-deny a tool confirmation, emit a valid
57
- conversation id, leave `result.response` blank, and still exit 0. A missing or
58
- whitespace-only response is therefore a failed turn at the agy adapter boundary
59
- for every caller, not a successful empty Planner reply. If the private log for
60
- that exact turn contains agy's narrow `tool_confirmation_manager` soft-denial
61
- marker, the caller receives a fixed, sanitized diagnosis; native log content is
62
- never returned. The captured conversation id remains usable, so a later
63
- operator-initiated chat resumes with `--conversation`. Recovery neither retries
64
- the failed message automatically nor relaxes permissions. Because the failed
65
- reply never reaches the control parser, it cannot change `plan.md` or
66
- `goal.json`, spawn subagents, or emit an initial directive. agy 1.2.2 can instead
67
- return a non-empty successful reply for the same soft denial; the adapter emits
68
- the refused tool names through `SendResult.permissionDenials` while preserving
69
- that reply.
56
+ Antigravity's read-only boundary is `--mode plan` on every Planner turn,
57
+ including recovery. An empty or whitespace-only agy reply is a failed turn, not
58
+ an empty Planner reply. When agy's log for that turn shows a soft-denied tool
59
+ confirmation, the caller receives a fixed diagnosis (native log content is never
60
+ returned); when agy instead returns a non-empty reply for a soft denial, the
61
+ reply is kept and the refused tool names are reported in
62
+ `SendResult.permissionDenials`. The Planner's conversation id stays resumable,
63
+ the failed message is not retried automatically, and permissions are not
64
+ relaxed. A failed reply never reaches the control parser, so it cannot change
65
+ `plan.md` or `goal.json`, spawn subagents, or emit an initial directive.
70
66
 
71
67
  Coder and Reviewer engine/model choices can be overridden by the first successful
72
68
  `spawn_subagents`; later attempts to change an already-started role are rejected
@@ -76,6 +72,11 @@ Coder and Reviewer **never speak to you directly**. Anything they observe
76
72
  flows through the Planner. The Planner decides what to surface and what to
77
73
  absorb.
78
74
 
75
+ A run left idle past `sessionTtlMinutes` has its role sessions evicted like any
76
+ other session. The next message to a role starts it again under the same name,
77
+ which resumes the persisted conversation where the engine supports it, instead of
78
+ failing with "Session not found".
79
+
79
80
  ## UX flow
80
81
 
81
82
  ```
@@ -86,19 +87,24 @@ absorb.
86
87
  3. autoloop_chat { run_id, "go" } → Planner emits spawn_subagents
87
88
  4. Coder + Reviewer self-iterate → ledger writes per iter
88
89
  5. Planner pushes you on target_hit / regression / decision / stall
89
- 6. Run terminates on target hit, plan-defined max_iters, or your terminate.
90
+ 6. Run ends when the Planner emits terminate, the phase-error circuit trips,
91
+ the hard deadline passes, or you stop it. An expired activity lease
92
+ pauses the run instead.
90
93
  ```
91
94
 
95
+ Stopping at a target or at `max_iters` from `goal.json` is the Planner's
96
+ decision: the runtime does not evaluate `goal.json`.
97
+
92
98
  ## Timeout hierarchy and recoverable sends
93
99
 
94
100
  Autoloop has three independent start-time controls. Their bounds are inclusive,
95
101
  and omitting them retains the defaults:
96
102
 
97
- | Wire field | Runtime field | Default | Minimum | Maximum | Meaning |
98
- | ------------------------------ | ------------------------- | -------- | ------- | --------- | ------- |
99
- | `send_timeout_ms` | `sendTimeoutMs` | 600000 | 5000 | 7200000 | Wall-clock cap for one Planner, Coder, or Reviewer delivery |
100
- | `activity_lease_ms` | `activityLeaseMs` | 1800000 | 60000 | 7200000 | Inactivity lease, renewed only by validated user or agent progress |
101
- | `autoloop_hard_timeout_ms` | `autoloopHardTimeoutMs` | 86400000 | 600000 | 259200000 | Absolute run deadline, anchored to start and never renewed |
103
+ | Wire field | Runtime field | Default | Minimum | Maximum | Meaning |
104
+ | -------------------------- | ----------------------- | -------- | ------- | --------- | ------------------------------------------------------------------ |
105
+ | `send_timeout_ms` | `sendTimeoutMs` | 600000 | 5000 | 7200000 | Wall-clock cap for one Planner, Coder, or Reviewer delivery |
106
+ | `activity_lease_ms` | `activityLeaseMs` | 1800000 | 60000 | 7200000 | Inactivity lease, renewed only by validated user or agent progress |
107
+ | `autoloop_hard_timeout_ms` | `autoloopHardTimeoutMs` | 86400000 | 600000 | 259200000 | Absolute run deadline, anchored to start and never renewed |
102
108
 
103
109
  Timer checks and runner-generated bookkeeping do not renew the activity lease.
104
110
  The hard deadline cannot be extended by repeated activity and wins if it fires
@@ -122,44 +128,49 @@ stored spec, chat or iteration evidence, or any earlier audit bytes.
122
128
 
123
129
  ## Quick start
124
130
 
131
+ Over HTTP, against `clawo serve` (default `127.0.0.1:18796`):
132
+
125
133
  ```bash
126
- # Start a run (creates Planner session)
127
- curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_start \
128
- -H 'content-type: application/json' \
134
+ TOKEN=$(cat ~/.openclaw/server-token)
135
+
136
+ # Start a run (creates the Planner session)
137
+ curl -X POST http://127.0.0.1:18796/autoloop/new \
138
+ -H "Authorization: Bearer $TOKEN" -H 'content-type: application/json' \
129
139
  -d '{"run_id":"my-run","workspace":"/abs/path/to/workspace"}'
130
140
 
131
- # Chat with the Planner
132
- curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_chat \
133
- -H 'content-type: application/json' \
134
- -d '{"run_id":"my-run","text":"Read the workspace and design a plan to fix X"}'
141
+ # Chat with the Planner (202; the reply arrives on /events as planner_reply)
142
+ curl -X POST http://127.0.0.1:18796/autoloop/my-run/chat \
143
+ -H "Authorization: Bearer $TOKEN" -H 'content-type: application/json' \
144
+ -d '{"text":"Read the workspace and design a plan to fix X"}'
135
145
 
136
146
  # Inspect state
137
- curl http://127.0.0.1:18789/autoloop/my-run/state
147
+ curl http://127.0.0.1:18796/autoloop/my-run/state -H "Authorization: Bearer $TOKEN"
138
148
 
139
- # Live SSE stream (the 3-pane UI subscribes here)
140
- curl http://127.0.0.1:18789/autoloop/my-run/events
149
+ # Live SSE stream (the dashboard's 3-pane view subscribes here)
150
+ curl -N http://127.0.0.1:18796/autoloop/my-run/events -H "Authorization: Bearer $TOKEN"
151
+ ```
152
+
153
+ Resetting an agent and stopping a run have no HTTP route; call the tools
154
+ (plugin or MCP):
141
155
 
142
- # Reset Coder if it drifts (lazy; eager_restart=true to start a fresh session immediately)
143
- curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_reset_agent \
144
- -H 'content-type: application/json' \
145
- -d '{"run_id":"my-run","agent":"coder","eager_restart":true}'
156
+ ```jsonc
157
+ // Reset the Coder if it drifts (lazy by default; eager_restart starts a fresh session now)
158
+ autoloop_reset_agent({ "run_id": "my-run", "agent": "coder", "eager_restart": true })
146
159
 
147
- # Stop
148
- curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_stop \
149
- -H 'content-type: application/json' \
150
- -d '{"run_id":"my-run","reason":"done"}'
160
+ // Stop
161
+ autoloop_stop({ "run_id": "my-run", "reason": "done" })
151
162
  ```
152
163
 
153
164
  ## Plugin tools
154
165
 
155
- | Tool | Args | What |
156
- | ---------------------- | ----------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
166
+ | Tool | Args | What |
167
+ | ---------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
157
168
  | `autoloop_start` | `run_id`, `workspace`, per-role `*_engine?`, `*_model?`, `*_custom_engine?`, `send_timeout_ms?`, `activity_lease_ms?`, `autoloop_hard_timeout_ms?` | Start a run; launches Planner and stores Coder/Reviewer defaults and timeout controls. Each `custom` role requires its matching config. |
158
- | `autoloop_chat` | `run_id`, `text` | Send a chat message to the Planner; returns the Planner's reply. |
159
- | `autoloop_status` | `run_id` | Current state (status, iter, push count, subagents_spawned). |
160
- | `autoloop_list` | — | All active runs in this manager process. |
161
- | `autoloop_stop` | `run_id`, `reason?` | Terminate; stops Planner / Coder / Reviewer. |
162
- | `autoloop_reset_agent` | `run_id`, `agent` ('planner' / 'coder' / 'reviewer'), `force?`, `eager_restart?` | Reset one subagent. Planner reset requires `force: true`. |
169
+ | `autoloop_chat` | `run_id`, `text` | Send a chat message to the Planner; returns the Planner's reply. |
170
+ | `autoloop_status` | `run_id` | Current state (status, iter, push count, subagents_spawned). |
171
+ | `autoloop_list` | — | All autoloop runs in the run store, live or not. |
172
+ | `autoloop_stop` | `run_id`, `reason?` | Terminate; stops Planner / Coder / Reviewer. |
173
+ | `autoloop_reset_agent` | `run_id`, `agent` ('planner' / 'coder' / 'reviewer'), `force?`, `eager_restart?` | Reset one subagent. Planner reset requires `force: true`. |
163
174
 
164
175
  ## Planner-emitted control tools
165
176
 
@@ -193,18 +204,17 @@ removes partial files, and suppresses both spawn and directives.
193
204
 
194
205
  ### Custom engines and resume
195
206
 
196
- Custom engine configs are accepted only by `autoloop_start` (or the HTTP resume
197
- body), never through Planner output. This keeps config fields such as `env` and
198
- static CLI arguments out of the Planner transcript and `decisions.jsonl`.
199
- The central registry persists only each role's engine and model, including the
200
- effective Coder/Reviewer selection after a successful spawn. Resume leaves the
201
- prior append-only row untouched until startup succeeds, so a transient CLI
202
- failure cannot erase the run. When resuming a run that uses `custom`, provide
203
- the matching `planner_custom_engine`, `coder_custom_engine`, or
204
- `reviewer_custom_engine` again; otherwise resume fails with a clear
207
+ Custom engine configs are accepted only by `autoloop_start` (and, on resume, by
208
+ `SessionManager.autoloopResume()` or by reference in the HTTP resume body), never
209
+ through Planner output. This keeps config fields such as `env` and static CLI
210
+ arguments out of the Planner transcript and `decisions.jsonl`. The run record
211
+ stores only each role's engine and model, including the effective Coder/Reviewer
212
+ selection after a successful spawn. When resuming a run that uses `custom`,
213
+ supply the matching config again (over HTTP, as a `*CustomEngineRef`, see
214
+ [Backend HTTP / SSE](#backend-http--sse)); otherwise resume fails with a clear
205
215
  configuration error rather than silently switching to Claude. Custom config
206
- shape is validated at runtime, while its `env` and static CLI arguments remain
207
- out of registry and audit records. See [`multi-engine.md`](./multi-engine.md)
216
+ shape is validated at runtime, while its `env` and static CLI arguments stay
217
+ out of the run record and audit logs. See [`multi-engine.md`](./multi-engine.md)
208
218
  for the `CustomEngineConfig` shape.
209
219
 
210
220
  ## Default push policy
@@ -213,34 +223,53 @@ for the `CustomEngineConfig` shape.
213
223
  | ---------------------- | ----------------------------------------------------- |
214
224
  | on_start | info / wechat ("loop started, will notify on issues") |
215
225
  | on_iter_done_ok | silent |
216
- | on_target_hit | info / both (webchat + wechat) |
226
+ | on_target_hit | info / both |
217
227
  | on_metric_regression_2 | warn / both |
218
228
  | on_reviewer_reject_2 | warn / both |
219
229
  | on_phase_error | error / both |
220
230
  | on_stall_30min | warn / wechat |
221
231
  | on_decision_needed | decision / both |
222
232
 
223
- 5-minute dedup on (level, summary) prevents duplicate pushes from the same
224
- event. Channel chain: `auto` walks wechat → whatsapp → email; `wechat` /
225
- `webchat` / `email` route directly; `both` does webchat (if session known)
233
+ The runtime fires five of these itself: `on_stall_30min` (30 minutes without
234
+ activity), `on_metric_regression_2`, `on_reviewer_reject_2` (two in a row),
235
+ `on_phase_error`, and `on_target_hit` (only when an acceptance contract passes,
236
+ see [Acceptance contracts](#acceptance-contracts)). `on_start`,
237
+ `on_iter_done_ok` and `on_decision_needed` are policy entries for the Planner's
238
+ own `notify_user` calls.
239
+
240
+ A 5-minute dedup on (level, summary) prevents duplicate pushes from the same
241
+ event. Channels: `auto` and `both` walk WeChat → WhatsApp → email and stop at
242
+ the first that succeeds; `wechat` and `email` go to that channel only;
243
+ `webchat` is a no-op (see [Known limitations](#known-limitations)).
244
+ **`on_phase_error` and `on_decision_needed` cannot be set to `silent: true`**
245
+ by the Planner: `update_push_policy` strips the flag and records the attempt in
246
+ `decisions.jsonl`.
226
247
 
227
- - wechat fallback chain. **`on_phase_error` and `on_decision_needed` cannot
228
- be set to `silent: true`** by Planner — `update_push_policy` strips the flag
229
- and records the attempt in `decisions.jsonl` (these channels are the
230
- operator's lifeline; they stay loud).
248
+ ## Notification setup
249
+
250
+ - **WeChat** needs `AUTOLOOP_WECHAT_RECIPIENT` and `AUTOLOOP_WECHAT_ACCOUNT`.
251
+ - **WhatsApp** needs `AUTOLOOP_WHATSAPP_RECIPIENT`.
252
+ - Both send through the `openclaw` CLI, which must be on `PATH`.
253
+ - **Email** is sent by running `bash "$AUTOLOOP_EMAIL_SCRIPT" -s "<subject>"`
254
+ with the message body on stdin.
255
+
256
+ Any channel whose variables are unset is skipped silently. With none set,
257
+ pushes are recorded in `push_log.jsonl` only.
231
258
 
232
259
  ## Auto-compact
233
260
 
234
261
  Each agent's context is monitored after every turn. When `getStats().contextPercent`
235
262
  crosses the per-agent threshold the dispatcher invokes `/compact` with a
236
263
  role-tuned hint (`compactSummaryFor`). Defaults: Planner 80 %, Coder 70 %,
237
- Reviewer 70 %. Override per run via `compactThresholds`. A 30 s debounce
264
+ Reviewer 70 %. The dispatcher's `compactThresholds` option overrides them; it
265
+ is library-level only and not exposed through `autoloop_start` or the HTTP
266
+ API. A 30 s debounce
238
267
  prevents re-fire while post-compact stats settle. Events: `compact` is
239
268
  emitted on the dispatcher EventEmitter AND appended to `decisions.jsonl`.
240
269
 
241
270
  One-shot engines (`codex`, `agy`, `grok`, `opencode`) cannot compact — their
242
271
  CLIs expose no such command. The threshold is still meaningful there because
243
- `contextPercent` now tracks real occupancy, but crossing it cannot free space:
272
+ `contextPercent` tracks real occupancy, but crossing it cannot free space:
244
273
  the session emits a single warning on its log channel the first time compaction
245
274
  is requested, then the thread keeps growing until the CLI refuses the request.
246
275
  Treat that warning as the signal to start a fresh session.
@@ -257,8 +286,9 @@ consecutive `phase_error`s and:
257
286
  `decision`-level push and an automatic `terminate { reason:
258
287
  'phase_error_circuit' }`.
259
288
 
260
- A successful (non-error) `iter_done` resets the counter. Override the
261
- threshold via `AutoloopConfig.phaseErrorCircuit`.
289
+ A successful (non-error) `iter_done` resets the counter. The threshold is
290
+ `AutoloopConfig.phaseErrorCircuit`, which is library-level only and not exposed
291
+ through `autoloop_start` or the HTTP API.
262
292
 
263
293
  ## Reviewer frozen memory
264
294
 
@@ -294,6 +324,8 @@ JSONL, one entry per line, ts-prefixed.
294
324
  ├── goal.json # Planner-authored, git-committed
295
325
  ├── push_log.jsonl # every notify_user attempt + channel used
296
326
  ├── decisions.jsonl # runner / dispatcher audit trail (see above)
327
+ ├── chat.jsonl # Planner-pane conversation, replayed by /chat_history
328
+ ├── evidence/iter-<n>/ # acceptance-contract bundle, when a contract is configured
297
329
  ├── reviewer_sandbox/ # Reviewer cwd; restaged per iter
298
330
  │ ├── plan.md # copy
299
331
  │ ├── goal.json # copy
@@ -304,7 +336,7 @@ JSONL, one entry per line, ts-prefixed.
304
336
  └── iter/<n>/
305
337
  ├── directive.json # Planner → Coder (schema_version: 1)
306
338
  ├── eval_output.json # what Coder reported (schema_version: 1)
307
- ├── diff.patch # git diff of the iter
339
+ ├── diff.patch # git diff of the iter, created files included
308
340
  ├── verdict.json # Reviewer decision + audit notes (schema_version: 1)
309
341
  └── coder_summary.txt
310
342
  ```
@@ -315,25 +347,28 @@ inside an iter** (pre-commit hook reject, signing key missing, …) the
315
347
  dispatcher emits a `phase_error` instead of writing `iter_artifacts`, so
316
348
  the failure is visible to the runner and counts toward the circuit.
317
349
 
350
+ `files_changed` in the iteration artifacts is taken from git, never from the
351
+ Coder's own report.
352
+
318
353
  Every JSON artifact in the ledger carries a `schema_version` field (currently
319
354
  `1`) to make future migrations explicit.
320
355
 
321
356
  ## Backend HTTP / SSE
322
357
 
323
- | Endpoint | Returns |
324
- | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
325
- | `GET /autoloop/list` | `{ ok, runs: AutoloopState[] }` |
326
- | `POST /autoloop/new` | `{ ok, run_id, planner_session }` — body `{ workspace, run_id?, planner_engine?, planner_model?, planner_custom_engine?, coder_engine?, coder_model?, coder_custom_engine?, reviewer_engine?, reviewer_model?, reviewer_custom_engine?, send_timeout_ms?, activity_lease_ms?, autoloop_hard_timeout_ms? }`. Timeout fields use the defaults and inclusive bounds documented above; malformed or out-of-range values return 400 before a run starts. |
327
- | `GET /autoloop/<id>/state` | `{ ok, state: AutoloopState }` — also returns a `terminated`-state stub reconstructed from the registry for runs that aren't in this process's memory, so the dashboard can open historical runs without 404'ing. |
328
- | `GET /autoloop/<id>/push_log` | `{ ok, entries: PushLogEntry[] }` — served from the ledger via `autoloopStatus`, so historical runs work the same as live ones. |
329
- | `GET /autoloop/<id>/chat_history` | `{ ok, entries: ChatEntry[] }` — replays `<ledger>/chat.jsonl`. The dashboard fetches this when opening a run so the Planner-pane conversation survives a page refresh / cross-process / re-opening a terminated run. Returns `[]` when the file doesn't exist (e.g. runs that predate the chat-history feature). |
330
- | `GET /autoloop/<id>/events` | SSE: `snapshot` / `message` / `state` / `push` / `iter_done` / `planner_reply` / `planner_error` / `coder_reply` / `reviewer_reply` / `terminated`. For runs that are NOT in this process's memory (terminated, or live in another process), the endpoint emits a single-shot `snapshot` + `terminated` then closes — the dashboard's existing handlers render history without hanging. |
331
- | `POST /autoloop/<id>/chat` | **202** `{ ok, queued: true }` — body `{ text }`. Fire-and-forget: the Planner's reply streams back via the `/events` SSE channel as a `planner_reply` event (or `planner_error` on failure); the HTTP response intentionally does NOT wait for it, because first-contact replies routinely exceed reverse-proxy idle limits (e.g. Cloudflare Tunnel cuts at ~100s → 524). 400 on empty text, 404 when the run is not in this process's memory. The MCP `autoloop_chat` tool path keeps the synchronous await-and-return-reply semantics (it runs in-process). |
332
- | `GET /autoloop/<id>/resume-requirements` | `{ ok, runId, rolesNeedingCustomEngine }` — the roles whose engine was `custom`, so a caller knows which secret references a resume needs. Role names only; nothing sensitive. 404 when there is no such run. |
358
+ | Endpoint | Returns |
359
+ | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
360
+ | `GET /autoloop/list` | `{ ok, runs: AutoloopState[] }` |
361
+ | `POST /autoloop/new` | `{ ok, run_id, planner_session }` — body `{ workspace, run_id?, planner_engine?, planner_model?, coder_engine?, coder_model?, reviewer_engine?, reviewer_model?, send_timeout_ms?, activity_lease_ms?, autoloop_hard_timeout_ms? }`. Timeout fields use the defaults and inclusive bounds documented above; malformed or out-of-range values return 400 before a run starts. |
362
+ | `GET /autoloop/<id>/state` | `{ ok, state: AutoloopState, live }` — `live` is `true` only when the run is running in this process. For a run that is not live here, `state` is the last state recorded in the run store, so historical runs open with their real iteration count. 404 when there is no such run. |
363
+ | `GET /autoloop/<id>/push_log` | `{ ok, entries: PushLogEntry[] }` — served from the ledger via `autoloopStatus`, so historical runs work the same as live ones. |
364
+ | `GET /autoloop/<id>/chat_history` | `{ ok, entries: ChatEntry[] }` — replays `<ledger>/chat.jsonl`. The dashboard fetches this when opening a run so the Planner-pane conversation survives a page refresh / cross-process / re-opening a terminated run. Returns `[]` when the file doesn't exist. |
365
+ | `GET /autoloop/<id>/events` | SSE: `snapshot` / `message` / `state` / `push` / `iter_done` / `planner_reply` / `planner_error` / `coder_reply` / `reviewer_reply` / `terminated`. For runs that are NOT in this process's memory (terminated, or live in another process), the endpoint emits a single-shot `snapshot` + `terminated` then closes — the dashboard's existing handlers render history without hanging. A run still in memory that has already reached `terminated` or `crashed` gets the same single-shot pair instead of an open stream that would never receive another event. Every such stream sets `retry: 864000000`, so an `EventSource` does not keep reconnecting to a stream that can only end again. |
366
+ | `POST /autoloop/<id>/chat` | **202** `{ ok, queued: true }` — body `{ text }`. Fire-and-forget: the Planner's reply streams back via the `/events` SSE channel as a `planner_reply` event (or `planner_error` on failure); the HTTP response intentionally does NOT wait for it, because first-contact replies routinely exceed reverse-proxy idle limits (e.g. Cloudflare Tunnel cuts at ~100s → 524). 400 on empty text. 404 when the run is not in this process's memory: if the store still holds it, the error says so and names `POST /autoloop/<id>/resume`; an unknown or malformed id is plain `not found`. The MCP `autoloop_chat` tool path keeps the synchronous await-and-return-reply semantics (it runs in-process). |
367
+ | `GET /autoloop/<id>/resume-requirements` | `{ ok, runId, rolesNeedingCustomEngine }` — the roles whose engine was `custom`, so a caller knows which secret references a resume needs. Role names only; nothing sensitive. 404 when there is no such run. |
333
368
  | `POST /autoloop/<id>/resume` | `{ ok, state }` — restore the role engine/model choices from the run's spec and re-create dispatcher + runner. For recoverable send timeouts, body fields `send_timeout_ms` and `pending_dispatch_id` apply the increase-only migration described above; `allow_decrease`, lease overrides, and hard-cap overrides are rejected. A custom-engine config is never persisted and is never accepted over HTTP, so a role using `custom` is re-supplied by **reference**: `plannerCustomEngineRef` / `coderCustomEngineRef` / `reviewerCustomEngineRef` name an environment variable `CLAWO_CUSTOM_ENGINE_<NAME>` on the orchestrator host, which the server reads and resolves. The name is not sensitive, the value never crosses the wire, and an unknown name is an error rather than a silent start without credentials. Existing engine-specific conversation resume behavior is reused where supported; `chat.jsonl` remains the visual history fallback. 404 when there is no such run. |
334
- | `POST /autoloop/<id>/delete` | `{ ok }` — stops the runner if still alive, scrubs the row from `~/.claw-orchestrator/autoloop-registry.jsonl`, and purges `persistedSessions` so the run cannot be `/resume`'d back. The ledger directory under `<workspace>/tasks/<run_id>/` is kept on disk. 404 if the run was not present in either memory or the registry. |
369
+ | `POST /autoloop/<id>/delete` | `{ ok }` — stops the loop if still live, deletes the run record from the run store, and purges the role sessions' persisted resume ids so the run cannot be resumed. The ledger directory under `<workspace>/tasks/<run_id>/` is kept on disk. 404 when there is no such run. |
335
370
 
336
- The 3-pane UI consumes these endpoints:
371
+ The dashboard's 3-pane autoloop view (`/dashboard`) consumes these endpoints:
337
372
 
338
373
  - **Left**: Planner chat (subscribes to `planner_reply`)
339
374
  - **Center**: Coder activity (`coder_reply` + `iter_done`)
@@ -341,8 +376,6 @@ The 3-pane UI consumes these endpoints:
341
376
  - **Top bar**: state (status / iter / metric)
342
377
  - **Bottom**: push_log
343
378
 
344
- The UI itself ships in a separate cross-repo PR.
345
-
346
379
  ## `goal.json` shape
347
380
 
348
381
  The Planner authors goal.json based on your conversation. There is no
@@ -377,106 +410,56 @@ The Planner will riff on this shape during your chat and ask if it's right.
377
410
  - ✅ Reviewer accumulates "fakery patterns I've seen" in `reviewer_memory.md` (persists across iters).
378
411
  - ✅ Reviewer defaults to `hold` under uncertainty; only `advance` after independent verification.
379
412
 
380
- ## Smoke test
413
+ ## Acceptance contracts
381
414
 
382
- `scripts/smoke-autoloop.ts` runs a buggy `add_two` scenario end-to-end with
383
- Opus Planner + Sonnet × 2. Validates plan.md / goal.json commit, spawn,
384
- iter 0 ledger artifacts (`directive` + `eval_output` + `diff.patch` +
385
- `verdict`), and termination on `target_hit`. Cost ~$1-3, wall-clock
386
- ~5-15 min. Run with `npx tsx scripts/smoke-autoloop.ts` (requires
387
- `~/.claude/settings.json` to have your auth env).
415
+ The Reviewer's sandbox holds the iteration's artifacts (`directive.json`,
416
+ `diff.patch`, `eval_output.json`, `coder_summary.txt`, `plan.md`, `goal.json`,
417
+ the prior verdict) but no code or evaluator, so its `advance` is a judgement of
418
+ the Coder's report rather than a measurement.
388
419
 
389
- ## Known limitations
420
+ An acceptance contract closes that gap. When one is configured, an `advance`
421
+ stands only if the checks pass against the workspace: otherwise the verdict is
422
+ rewritten to `hold`, the failing checks are appended to `audit_notes`, and the
423
+ bundle is written to `<ledger>/evidence/iter-<n>/`. A passing contract fires
424
+ `on_target_hit`. Without a contract the Reviewer's verdict is used as-is.
390
425
 
391
- - **`webchat` channel is a no-op** — `notifyUserFallbackChain` does not yet
392
- carry a webchat session id at the run level, so `channel: 'webchat'`
393
- always returns `channel_used: 'none'`. Use `auto` / `wechat` / `email`
394
- until the inbound route lands.
395
- - **One-way push.** WeChat → Planner inbound replies are not yet wired (would
396
- need an openclaw-gateway tmux-passthrough route). Reply via webchat /
397
- `autoloop_chat`.
398
- - **No webchat UI yet.** Backend SSE is shipped; the UI is a separate
399
- cross-repo PR in ChatGPT-Next-Web.
400
- - **No fork / population mode.** Single linear iter trajectory per run.
401
- - **Cross-run knowledge isolated.** Each run's `reviewer_memory.md` and
402
- `coder_notes.md` live in that run's ledger; no shared meta-store yet.
403
- - **No cost / wall-clock budget cap.** Only `phaseErrorCircuit` + Reviewer
404
- hold/reject streaks bound the run; a steady-but-pointless ratchet could
405
- run for days. Set `max_iters` in `goal.json` to bound iter count.
406
- - **Run state in memory.** SessionManager restart drops the live `autoloops`
407
- map; the on-disk ledger survives but cannot resume a running state.
408
- - **Multi-run / same workspace** races on `git index.lock`. Run separate
409
- workspaces (or git worktrees) for concurrent runs.
410
-
411
- ## Acceptance contracts (6.0.0)
412
-
413
- The Reviewer's `advance` was the only thing standing between an iteration and
414
- "done", and it could not do the job it was given. Its prompt tells it to
415
- "re-derive the metric independently if the sandbox has the bits to do so", but
416
- `stageReviewSandbox` copies in the iteration's artifacts — `directive.json`,
417
- `diff.patch`, `eval_output.json`, `coder_summary.txt` — plus `plan.md`,
418
- `goal.json`, and the prior verdict. No code, no evaluator. The verdict was
419
- therefore a reading of the Coder's own report, and `eval_output` was literally
420
- whatever the Coder passed to a tool call.
426
+ **The contract is library-level only.** It is set on the dispatcher config
427
+ (`contract` in `ClaudeAgentDispatcherConfig`) and is not yet exposed through
428
+ `autoloop_start` or the HTTP API, so a run started from the tool or
429
+ `POST /autoloop/new` has no contract.
421
430
 
422
- Pass a contract at autoloop start and an `advance` is held unless the checks
423
- pass. The verdict is rewritten to `hold`, the reason is appended to
424
- `audit_notes`, and the evidence bundle lands at
425
- `<ledger>/iter/<n>/evidence/`. Without a contract nothing changes.
431
+ ## Lifecycle and resume
426
432
 
427
- ### `on_target_hit` now fires
428
-
429
- That push-policy key was declared in `types.ts`, given a default, and whitelisted
430
- for runtime updates — and had **zero firing sites** anywhere in the codebase.
431
- Autoloop had four ways to notice it was failing (stall, metric regression,
432
- reviewer rejections, phase-error circuit) and no way to notice it had succeeded,
433
- because a Reviewer verdict is not a measurement. An acceptance contract is, so
434
- `on_target_hit` fires when one passes.
433
+ An autoloop is a kernel run whose single node holds the loop for as long as it
434
+ lives. The run record holds the last state the loop published and the engines
435
+ `spawn_subagents` chose, so `autoloop_status` and `GET /autoloop/<id>/state`
436
+ show a run's real state after it stops or the process restarts.
435
437
 
436
- ### `diff.patch` sees created files
438
+ Resume is explicit. `POST /autoloop/<id>/resume` (or
439
+ `SessionManager.autoloopResume()`) restarts a run from its stored spec.
440
+ Custom-engine configs are the one thing the spec does not carry (they can hold
441
+ secrets), so a resume must be given them again. Cancelling a run stops all three
442
+ agents, the same as `autoloop_stop`.
437
443
 
438
- The per-iteration patch was captured with a bare `git diff`, which lists tracked
439
- modifications only. A file the Coder _created_ appeared in neither the patch nor
440
- the `--name-only` fallback, while the `git add -A` two lines later committed it —
441
- so the Reviewer audited a picture that structurally could not show new files.
442
- The capture now covers tracked changes ∪ untracked files.
444
+ ## Known limitations
443
445
 
444
- `files_changed` is also taken from git unconditionally. It previously preferred
445
- the Coder's own `files_changed` whenever the Coder supplied one, despite the
446
- comment above it saying the claim was not trusted.
446
+ - **`webchat` channel is a no-op.** No webchat session id is carried at the run
447
+ level, so `channel: 'webchat'` always returns `channel_used: 'none'`. Use
448
+ `auto` / `wechat` / `email`.
449
+ - **One-way push.** Replies to a push are not routed back to the Planner;
450
+ answer with `autoloop_chat`.
451
+ - **No fork / population mode.** Single linear iter trajectory per run.
452
+ - **Cross-run knowledge isolated.** Each run's `reviewer_memory.md` and
453
+ `coder_notes.md` live in that run's ledger; there is no shared store.
454
+ - **No cost cap.** `maxBudgetUsd` is not exposed for autoloop. Wall-clock time
455
+ is bounded by `autoloop_hard_timeout_ms` (default 24 h); the iteration count
456
+ is bounded only by the Planner honouring `max_iters` in `goal.json`.
457
+ - **Resume is not automatic.** After a restart a run is not live in the new
458
+ process until it is resumed, and a send that was in flight is not retried.
459
+ - **Multi-run / same workspace** races on `git index.lock`. Run separate
460
+ workspaces (or git worktrees) for concurrent runs.
447
461
 
448
462
  ## Related
449
463
 
450
464
  - [`verification.md`](./verification.md) — contracts, checks, evidence
451
-
452
- ## Lifecycle moved to the kernel (6.0.0)
453
-
454
- An autoloop is a kernel run whose single node holds the loop for as long as it
455
- lives. Tool signatures are unchanged.
456
-
457
- What went away: the `autoloops` map; `~/.claw-orchestrator/autoloop-registry.jsonl`
458
- with its four bespoke helpers (append, remove-then-append upsert, reverse-scan
459
- dedup, rewrite-via-tmp-file); and the two `Set`s — `_startingAutoloops` and
460
- `_deletingAutoloops` — that existed only because a start and a delete could race
461
- each other over that shared map.
462
-
463
- Two things get better rather than merely moving:
464
-
465
- - **`autoloop_status` on a run that is not live in this process** used to return
466
- an all-zero stub labelled `reconstructed from registry — not in current process
467
- memory`: iter 0, no metrics, no error history, because the registry only ever
468
- held identity. The record holds the last state the loop published, so a
469
- historical run opens with its real iteration count.
470
- - **The engines `spawn_subagents` actually chose** land on the run record
471
- alongside the rest of its state, instead of in a parallel file with its own
472
- lifecycle.
473
-
474
- `autoloop_resume` restarts a terminated run from the stored spec — the immutable
475
- record of how it was started — rather than from a registry row whose older
476
- versions omitted the engine fields entirely. Custom-engine configs are the one
477
- thing the spec does not carry (they can hold secrets), so a resume must be given
478
- them again.
479
-
480
- Cancelling a run now tears the loop down the way a stop does. It previously left
481
- the three persistent agents running and their session names claimed, which
482
- surfaced much later as `session name already in use`.
465
+ - [`workflow.md`](./workflow.md) — the kernel that stores and resumes runs