@enderfga/claw-orchestrator 7.5.2 → 7.5.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. package/README.md +26 -27
  2. package/configs/engines/README.md +7 -6
  3. package/dist/bin/cli.js +1 -1
  4. package/dist/bin/cli.js.map +1 -1
  5. package/dist/src/acp-server.d.ts +1 -1
  6. package/dist/src/acp-server.js +7 -5
  7. package/dist/src/acp-server.js.map +1 -1
  8. package/dist/src/autoloop/dispatcher.js +3 -3
  9. package/dist/src/autoloop/dispatcher.js.map +1 -1
  10. package/dist/src/autoloop/notify.d.ts +5 -7
  11. package/dist/src/autoloop/notify.js +21 -20
  12. package/dist/src/autoloop/notify.js.map +1 -1
  13. package/dist/src/base-oneshot-session.js +20 -1
  14. package/dist/src/base-oneshot-session.js.map +1 -1
  15. package/dist/src/dashboard/index.html +94 -15
  16. package/dist/src/embedded-server.js +40 -7
  17. package/dist/src/embedded-server.js.map +1 -1
  18. package/dist/src/fanout.d.ts +6 -0
  19. package/dist/src/fanout.js +1 -0
  20. package/dist/src/fanout.js.map +1 -1
  21. package/dist/src/index.js +19 -11
  22. package/dist/src/index.js.map +1 -1
  23. package/dist/src/kernel/nodes/fanout.js +1 -0
  24. package/dist/src/kernel/nodes/fanout.js.map +1 -1
  25. package/dist/src/kernel/types.d.ts +2 -0
  26. package/dist/src/kernel/types.js.map +1 -1
  27. package/dist/src/models.js +36 -7
  28. package/dist/src/models.js.map +1 -1
  29. package/dist/src/openai-compat.d.ts +2 -2
  30. package/dist/src/openai-compat.js +5 -2
  31. package/dist/src/openai-compat.js.map +1 -1
  32. package/dist/src/persistent-agy-session.js +6 -1
  33. package/dist/src/persistent-agy-session.js.map +1 -1
  34. package/dist/src/session-manager.d.ts +1 -0
  35. package/dist/src/session-manager.js +17 -5
  36. package/dist/src/session-manager.js.map +1 -1
  37. package/dist/src/types.d.ts +2 -0
  38. package/openclaw.plugin.json +1 -1
  39. package/package.json +2 -2
  40. package/skills/SKILL.md +31 -32
  41. package/skills/references/acp.md +19 -36
  42. package/skills/references/autoloop.md +163 -180
  43. package/skills/references/claude-cli-tracking.md +28 -27
  44. package/skills/references/cli.md +62 -79
  45. package/skills/references/council.md +40 -63
  46. package/skills/references/dashboard.md +42 -55
  47. package/skills/references/getting-started.md +21 -15
  48. package/skills/references/inbox.md +6 -4
  49. package/skills/references/mcp.md +29 -24
  50. package/skills/references/multi-engine.md +105 -153
  51. package/skills/references/observability.md +42 -32
  52. package/skills/references/openai-compat.md +169 -303
  53. package/skills/references/sessions.md +20 -29
  54. package/skills/references/tools.md +62 -76
  55. package/skills/references/ultra.md +17 -16
  56. package/skills/references/ultraapp.md +59 -64
  57. package/skills/references/verification.md +29 -52
  58. package/skills/references/workflow.md +37 -104
  59. package/skills/ultraapp/SKILL.md +9 -10
@@ -8,7 +8,7 @@ council → fix-on-failure → deploy → done-mode feedback.
8
8
  This page is the operator reference. The interview behavioural contract
9
9
  lives in [`skills/ultraapp/SKILL.md`](../ultraapp/SKILL.md). The
10
10
  council architectural conventions every generated app must satisfy live
11
- in [`src/ultraapp/conventions.ts`](../../src/ultraapp/conventions.ts).
11
+ in [`src/ultraapp/conventions.ts`](https://github.com/Enderfga/claw-orchestrator/blob/main/src/ultraapp/conventions.ts).
12
12
 
13
13
  ## When to use
14
14
 
@@ -32,15 +32,15 @@ interview ─► queued ─► building ─► build-complete ─► deploying
32
32
  structural)
33
33
  ```
34
34
 
35
- | Mode | Meaning |
36
- | ---------------- | ------------------------------------------------------------------------------- |
37
- | `interview` | AppSpec being filled by Q&A. Chat input goes to the interview Opus. |
38
- | `queued` | Build accepted, waiting for a slot in the FIFO build queue. |
39
- | `building` | Council writing code, fix-on-failure driving install/build/test. |
40
- | `build-complete` | Codebase ready, awaiting `deploy` step. |
41
- | `deploying` | Container/process being started, router map being updated. |
42
- | `done` | App live at `/forge/<slug>/`. Chat input now goes to the done-mode classifier. |
43
- | `failed` | Council didn't reach consensus, or fix-on-failure couldn't get the build green. |
35
+ | Mode | Meaning |
36
+ | ---------------- | ---------------------------------------------------------------------------------- |
37
+ | `interview` | AppSpec being filled by Q&A. Chat input goes to the interview agent (Claude Opus). |
38
+ | `queued` | Build accepted, waiting for a slot in the FIFO build queue. |
39
+ | `building` | Council writing code, fix-on-failure driving install/build/test. |
40
+ | `build-complete` | Codebase ready, awaiting `deploy` step. |
41
+ | `deploying` | Container/process being started, router map being updated. |
42
+ | `done` | App live at `/forge/<slug>/`. Chat input now goes to the done-mode classifier. |
43
+ | `failed` | Council didn't reach consensus, or fix-on-failure couldn't get the build green. |
44
44
 
45
45
  ### The build is a workflow run
46
46
 
@@ -73,17 +73,21 @@ state machine hidden inside it.
73
73
  ## Architectural conventions (§1–§7)
74
74
 
75
75
  Every generated app MUST satisfy these. They're embedded in the council
76
- super-task prompt verbatim from `src/ultraapp/conventions.ts`.
77
-
78
- | § | Topic | Headline rule |
79
- | --- | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
80
- | 1 | Path-based deploy | Mount at `BASE_PATH=/forge/<slug>/`; in-app links MUST be relative. |
81
- | 2 | Async file-queue runtime | Exact endpoints: `GET /`, `POST /run`, `GET /status/:jobId`, `GET /result/:jobId`, `GET /health`. File-based job queue under `$DATA_DIR/jobs/<jobId>/`. NO database. Data path from `process.env.DATA_DIR ?? '/data'`. |
82
- | 3 | BYOK | If `runtime.needsLLM`, API keys live in browser localStorage and are sent direct to the provider. The server MUST NEVER receive the key (enforced by `eslint-plugin-no-server-keys`). |
83
- | 4 | Dockerfile + smoke test | Single multi-stage Dockerfile, `npm run smoke` drives one full job in < 90s using `examples[0].ref`. |
84
- | 5 | Council voting protocol | 3 agents in git worktrees, all-YES vote required, max 8 rounds. |
85
- | 6 | Tech stack | Modern TypeScript / JavaScript framework (Next.js, Vite + Hono, SvelteKit). NO Python, NO pure SSGs. |
86
- | 7 | **Frontend quality** | **Real styling system + real type hierarchy + four-state coverage on every async surface + drag-and-drop forms + appropriate result presentation + one deliberate theme.** §7g requires every agent to capture Chrome-headless screenshots at 1440×900 AND 375×812 and visually inspect the PNGs before voting YES — source-code review is explicitly insufficient evidence. |
76
+ super-task prompt verbatim from [`src/ultraapp/conventions.ts`](https://github.com/Enderfga/claw-orchestrator/blob/main/src/ultraapp/conventions.ts).
77
+
78
+ | § | Topic | Headline rule |
79
+ | --- | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
80
+ | 1 | Path-based deploy | Mount at `BASE_PATH=/forge/<slug>/`; in-app links MUST be relative. |
81
+ | 2 | Async file-queue runtime | Exact endpoints: `GET /`, `POST /run`, `GET /status/:jobId`, `GET /result/:jobId`, `GET /health`. File-based job queue under `$DATA_DIR/jobs/<jobId>/`. NO database. Data path from `process.env.DATA_DIR ?? '/data'`. |
82
+ | 3 | BYOK | If `runtime.needsLLM`, API keys live in browser localStorage and are sent direct to the provider. The server MUST NEVER receive the key (enforced by `eslint-plugin-no-server-keys`). |
83
+ | 4 | Dockerfile + smoke test | Single multi-stage Dockerfile, `npm run smoke` drives one full job in < 90s using `examples[0].ref`. |
84
+ | 5 | Council voting protocol | 3 agents in git worktrees, all-YES vote required, max 8 rounds. |
85
+ | 6 | Tech stack | Modern TypeScript / JavaScript framework (Next.js, Vite + Hono, SvelteKit). NO Python, NO pure SSGs. |
86
+ | 7 | Frontend quality | Real styling + type hierarchy, full async-state coverage, drag-and-drop forms, one deliberate theme; §7g screenshot gate (below). |
87
+
88
+ **§7g:** every agent captures headless-Chrome screenshots at 1440×900 and
89
+ 375×812 and inspects them before voting YES; reading source code is not
90
+ accepted as evidence.
87
91
 
88
92
  ## Runtime modes
89
93
 
@@ -128,29 +132,29 @@ persists to `~/.claw-orchestrator/host-procs.json`.
128
132
  All routes are served by the embedded server (default `:18796`), under
129
133
  `Authorization: Bearer <token>` from `~/.openclaw/server-token`.
130
134
 
131
- | Method + path | Purpose |
132
- | ------------------------------------- | --------------------------------------------------------------------------------- |
133
- | `GET /ultraapp/list` | All runs with mode + createdAt. |
134
- | `POST /ultraapp/new` | Body: `{ firstMessage?: string }`. Returns `{ runId }`. |
135
- | `GET /ultraapp/<id>` | Full snapshot: spec + chat + state. |
136
- | `POST /ultraapp/<id>/answer` | Body: `{ value, freeform? }`. Submit interview answer. |
137
- | `POST /ultraapp/<id>/spec-edit` | Body: RFC 6902 patch ops. Edit the spec mid-interview. |
138
- | `POST /ultraapp/<id>/files` | Multipart upload to `examples/`. |
139
- | `GET /ultraapp/<id>/events` | SSE stream of build/chat events (mode pill, narrator, council activity). |
140
- | `POST /ultraapp/<id>/build` | Validate spec strictly + enqueue. |
141
- | `POST /ultraapp/<id>/build/cancel` | Abort the active build. |
142
- | `GET /ultraapp/<id>/artifacts` | List `versions/vN/`. |
143
- | `POST /ultraapp/<id>/start` | Start the deployed container/process for the active version. |
144
- | `POST /ultraapp/<id>/stop` | Stop without deleting. |
145
- | `POST /ultraapp/<id>/delete` | Stop + remove all per-run state. |
146
- | `POST /ultraapp/<id>/feedback` | Body: `{ text }`. Done-mode classifier routes cosmetic / spec-delta / structural. |
147
- | `POST /ultraapp/<id>/promote-version` | Body: `{ version: "vN" }`. Atomically swap deployed version. |
135
+ | Method + path | Purpose |
136
+ | ------------------------------------- | ------------------------------------------------------------------------------------------------------------- |
137
+ | `GET /ultraapp/list` | All runs with mode + createdAt. |
138
+ | `POST /ultraapp/new` | No body. Returns `{ runId }`. |
139
+ | `GET /ultraapp/<id>` | Full snapshot: spec + chat + state. |
140
+ | `POST /ultraapp/<id>/answer` | Body: `{ value, freeform? }`. Submit interview answer. |
141
+ | `POST /ultraapp/<id>/spec-edit` | Body: `{ patch: [...] }` (RFC 6902 ops). Edit the spec mid-interview. |
142
+ | `POST /ultraapp/<id>/files` | JSON body `{ absolutePath }` (a file on the server host) or `{ filename, dataB64 }`. Copies into `examples/`. |
143
+ | `GET /ultraapp/<id>/events` | SSE stream of build/chat events (mode pill, narrator, council activity). |
144
+ | `POST /ultraapp/<id>/build` | Validate spec strictly + enqueue. |
145
+ | `POST /ultraapp/<id>/build/cancel` | Abort the active build. |
146
+ | `GET /ultraapp/<id>/artifacts` | List `versions/vN/`. |
147
+ | `POST /ultraapp/<id>/start` | Start the deployed container/process for the active version. |
148
+ | `POST /ultraapp/<id>/stop` | Stop without deleting. |
149
+ | `POST /ultraapp/<id>/delete` | Stop + remove all per-run state. |
150
+ | `POST /ultraapp/<id>/feedback` | Body: `{ text }`. Done-mode classifier routes cosmetic / spec-delta / structural. |
151
+ | `POST /ultraapp/<id>/promote-version` | Body: `{ version: "vN" }`. Atomically swap deployed version. |
148
152
 
149
153
  ## MCP tools (14)
150
154
 
151
155
  Same surface as HTTP, callable from any Model Context Protocol host
152
156
  (Claude Desktop, Hermes Agent, Cursor, Cline, Continue, Zed,
153
- Windsurf, Goose). Param schemas in [`tools.md`](./tools.md#ultraapp).
157
+ Windsurf, Goose). Param schemas in [`tools.md`](./tools.md#ultraapp-14).
154
158
 
155
159
  ```text
156
160
  ultraapp_list ultraapp_get ultraapp_status
@@ -165,11 +169,11 @@ ultraapp_start_container ultraapp_stop_container ultraapp_delete
165
169
  After the run reaches `done`, chat input goes to a per-run Haiku
166
170
  classifier. Three classes:
167
171
 
168
- | Class | Routes to | Behaviour |
169
- | ------------ | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
170
- | `cosmetic` | Patcher | Opus generates a unified diff against the deployed worktree → `applyUnifiedDiff` → validate via fix-on-failure → on success snapshot to `versions/vN+1/`, on any failure restore the snapshot atomically and post the reason to chat. |
171
- | `spec-delta` | Focused interview | Flips mode back to `interview` with a bootstrap message that names the field(s) being changed. Completion auto-triggers a fresh `startBuild`. |
172
- | `structural` | Suggestion only | Posts a narrator note: "this sounds like a different app — click + New". |
172
+ | Class | Routes to | Behaviour |
173
+ | ------------ | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
174
+ | `cosmetic` | Patcher | A Claude Opus session generates a unified diff against the deployed worktree → `applyUnifiedDiff` → validate via fix-on-failure → on success snapshot to `versions/vN+1/`, on any failure restore the snapshot atomically and post the reason to chat. |
175
+ | `spec-delta` | Focused interview | Flips mode back to `interview` with a bootstrap message that names the field(s) being changed. Completion auto-triggers a fresh `startBuild`. |
176
+ | `structural` | Suggestion only | Posts a narrator note: "this sounds like a different app — click + New". |
173
177
 
174
178
  To swap which version is live, use `promote-version` (HTTP) or
175
179
  `ultraapp_promote_version` (MCP) — the router map and host-procs map
@@ -177,8 +181,8 @@ update atomically.
177
181
 
178
182
  ## Reference traces + replay
179
183
 
180
- 5 captured JSONL traces of real interviews ground-truth the interview
181
- engine against drift:
184
+ 5 reference interview traces (JSONL) pin the interview engine's
185
+ output against drift:
182
186
 
183
187
  ```text
184
188
  src/__tests__/fixtures/ultraapp-traces/
@@ -192,8 +196,8 @@ src/__tests__/fixtures/ultraapp-traces/
192
196
  ```
193
197
 
194
198
  ```bash
195
- tsx scripts/test-ultraapp-integration.ts --trace=image-batch-resize
196
- tsx scripts/test-ultraapp-integration.ts --trace=all
199
+ npx tsx scripts/test-ultraapp-integration.ts --trace=image-batch-resize
200
+ npx tsx scripts/test-ultraapp-integration.ts --trace=all
197
201
  ```
198
202
 
199
203
  The `spec-extraction-quality.test.ts` test replays each trace through
@@ -218,11 +222,9 @@ open "http://127.0.0.1:18796/dashboard?token=$(cat ~/.openclaw/server-token)"
218
222
  curl http://127.0.0.1:19000/forge/<slug>/health
219
223
  ```
220
224
 
221
- ## Acceptance contract (6.0.0)
225
+ ## Acceptance contract
222
226
 
223
- UltraApp is the one mode whose contract is on by default, because it already ran
224
- most of these commands and because two of its documented gates were not actually
225
- enforced.
227
+ UltraApp is the one mode whose acceptance contract is on by default.
226
228
 
227
229
  **Build stage**, in the council's worktree:
228
230
 
@@ -230,10 +232,7 @@ enforced.
230
232
  npm install → npm run build → npm test → [docker build] → npm run smoke
231
233
  ```
232
234
 
233
- `npm run smoke` is new here. §4 of the architectural conventions has always told
234
- the council that the smoke test gates build success — it was never in the step
235
- list, so the claim was false. A codebase without a working `scripts.smoke` now
236
- fails its build, which is what the brief said all along.
235
+ A codebase without a working `scripts.smoke` fails its build (§4).
237
236
 
238
237
  **Deploy stage**, against the live URL: both §7g viewports (1440×900 and
239
238
  375×812) are captured by the orchestrator with headless Chrome and stored as run
@@ -249,8 +248,7 @@ Evidence lands under the run directory:
249
248
 
250
249
  ### What the visual gate does and does not do
251
250
 
252
- It captures images and stores them, so "did anyone actually look" is now a file
253
- on disk instead of an agent's claim. It does **not** compare pixels — judging
251
+ It captures images and stores them as run evidence. It does **not** compare pixels — judging
254
252
  the rendering is still a reader's job, and the §7g instructions in the council
255
253
  prompt remain the agents' responsibility.
256
254
 
@@ -259,13 +257,10 @@ app to a missing browser. Set `CLAWO_ULTRAAPP_VISUAL_GATE=strict` to make a
259
257
  failed capture block the deploy. Chrome is resolved from `CLAWO_CHROME_BIN`, then
260
258
  the usual macOS app paths, then `PATH`.
261
259
 
262
- ## Durable build queue (6.0.0)
260
+ ## Durable build queue
263
261
 
264
262
  The build queue is persisted to `<store>/build-queue.json` and restored on
265
- startup. Its own comment used to say a restart mid-build meant "the build is
266
- marked failed and the user can rerun"; in practice nothing was marked — the
267
- pending list vanished along with any queued build the user was waiting on, with
268
- no record it had been asked for.
263
+ startup.
269
264
 
270
265
  A build that was in flight when the process died is **re-queued, not resumed**,
271
266
  and goes to the front: each build starts from a fresh council worktree, so
@@ -274,6 +269,6 @@ re-running is safe and continuing a half-built tree is not.
274
269
  ## Known limitations
275
270
 
276
271
  - The done-mode patcher loop occasionally hangs between
277
- feedback-classification and the patcher Opus session creation;
272
+ feedback classification and the start of the patcher's Claude Opus session;
278
273
  cosmetic changes can be applied manually until the underlying race
279
274
  is fixed.
@@ -1,19 +1,10 @@
1
1
  # Verification — acceptance contracts and evidence
2
2
 
3
- Before 6.0.0, nothing in this runtime ever checked an agent's work. Every
4
- "finished" signal was the agent grading itself:
5
-
6
- - **Council** terminated when a regex found `[CONSENSUS: YES]` in agent prose.
7
- - **Autoloop**'s `eval_output` was whatever the Coder passed to a tool call, and
8
- the Reviewer that was supposed to catch fabrication had a sandbox containing
9
- the iteration's artifacts and no code, so it could not re-derive anything.
10
- - **UltraApp**'s frontend gate was a sentence in a persona string telling agents
11
- to capture screenshots. There was no screenshot code anywhere in the project.
12
- - The run ledger's `ok` was the engine's report on its own turn.
13
-
14
- An **acceptance contract** is the opposite of all of that: a list of checks the
15
- runtime executes and whose results it reads. A run that declares one cannot reach
16
- `completed` unless every required check passes.
3
+ An **acceptance contract** is a list of checks the runtime executes itself and
4
+ whose results it reads. A run that declares one cannot reach `completed` unless
5
+ every required check passes. Completion signals that come from an agent or an
6
+ engine (consensus votes, a Coder's reported metric, an engine's `ok`) are
7
+ recorded, but they are never treated as verification.
17
8
 
18
9
  ## The one rule about where contracts come from
19
10
 
@@ -55,24 +46,17 @@ Default true. A failing non-required check is recorded in the evidence bundle bu
55
46
  does not refute the run — use it for signals you want visible without making them
56
47
  blocking.
57
48
 
58
- (The predecessor of this module, `ultraapp/fix-on-failure.ts`, declared a
59
- `required` field on its steps and never read it: every step was fatal. It is
60
- honoured now.)
61
-
62
49
  ### Timeouts
63
50
 
64
- Every check has one, and the default is 10 minutes. This is not cosmetic — the
65
- old pipeline had no timeout at all, so a wedged `npm test` hung the build
66
- forever. A check that overruns is killed (its whole process group, with SIGKILL)
51
+ Every check has one; the default is 10 minutes. A check that overruns is killed (its whole process group, with SIGKILL)
67
52
  and recorded as failed with `timedOut: true`.
68
53
 
69
54
  ### What `screenshot` does and does not claim
70
55
 
71
56
  It captures images and stores them. It does **not** compare pixels, and it is not
72
- visual regression testing. What changed in 6.0.0 is that the capture is performed
73
- by the runtime rather than requested of an agent — so "did anyone actually look"
74
- stops being a claim and starts being a file on disk. Judging the rendering is
75
- still a human's or an agent's job.
57
+ visual regression testing. The runtime performs the capture, so the screenshots
58
+ are files on disk rather than an agent's claim. Judging the rendering is still a
59
+ human's or an agent's job.
76
60
 
77
61
  Chrome is resolved from `CLAWO_CHROME_BIN`, then the usual macOS app paths, then
78
62
  `PATH`. A host with no browser fails the check immediately rather than paying the
@@ -85,8 +69,7 @@ kernel records `git rev-parse HEAD` when a run starts and diffs against that.
85
69
 
86
70
  The change set is **tracked changes ∪ untracked files**. That union matters: a
87
71
  bare `git diff` lists tracked modifications only, so files an agent _created_ are
88
- invisible to it — which is exactly how `autoloop`'s per-iteration `diff.patch`
89
- used to miss every new file while `git add -A` committed them anyway.
72
+ invisible to it.
90
73
 
91
74
  `requirePaths: ["."]` means "the run must have changed something".
92
75
 
@@ -167,22 +150,16 @@ whether it passed or not and runs again before every fixer round.
167
150
 
168
151
  ## Per-mode defaults
169
152
 
170
- | Mode | Contract | Notes |
171
- | ------------------------ | ----------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- |
172
- | **UltraApp** | **On by default** | Build: `npm install` → `npm run build` → `npm test` → (`docker build`) → **`npm run smoke`**. Deploy: **both §7g viewports captured** against the live URL. |
173
- | **Council** | Caller-declared | Consensus votes are recorded on the run as advisory and no longer decide completion. |
174
- | **Autoloop** | Caller-declared | With a contract, a Reviewer `advance` is held unless the checks pass; passing also fires `on_target_hit`. |
175
- | **Fanout / Ultrareview** | Caller-declared | Per-agent `ok` now reads the engine's terminal verdict rather than "the call did not throw". |
176
- | **Plain sessions** | None | Use `verify_run` to check work after the fact. |
153
+ | Mode | Contract | Notes |
154
+ | ------------------------ | ----------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
155
+ | **UltraApp** | **On by default** | Build: `npm install` → `npm run build` → `npm test` → (`docker build`) → **`npm run smoke`**. Deploy: **screenshots at 1440×900 and 375×812** against the live URL. |
156
+ | **Council** | Caller-declared | Consensus votes are recorded on the run as advisory and do not decide completion. |
157
+ | **Autoloop** | Not yet exposed | Library-level only (the dispatcher's `contract` option); `autoloop_start` and the HTTP API do not accept one. See [`autoloop.md`](./autoloop.md#acceptance-contracts). |
158
+ | **Fanout / Ultrareview** | Caller-declared | Per-agent `ok` is the engine's terminal verdict for its turn. |
159
+ | **Plain sessions** | None | Use `verify_run` to check work after the fact. |
177
160
 
178
- Two of UltraApp's defaults close claims the project had been making without
179
- backing them:
180
-
181
- - `conventions.ts` §4 told every council that `npm run smoke` gated build
182
- success. It was not in the step list at all. It is now.
183
- - `ultraapp.md` recorded, as a known limitation since 4.0.0, that the §7g gate
184
- "relies on per-agent honesty about running the screenshot capture", with a
185
- server-side validator promised as a follow-up. That follow-up is this release.
161
+ UltraApp's build gate includes `npm run smoke`, and its deploy gate captures
162
+ screenshots at 1440×900 and 375×812 against the live URL.
186
163
 
187
164
  The visual gate is **advisory by default** so that a host without Chrome does not
188
165
  lose a working app to a missing browser. Set
@@ -203,15 +180,10 @@ older version. The contract is yours; nothing is read from agent output.
203
180
  A run ends `verified`, `refuted`, or `unverified`.
204
181
 
205
182
  `unverified` means **no contract was declared and nothing checked the work**. It
206
- is not a failure and it is not a pass. The CLI prints it as `—` and the summary
207
- line says so in words, because collapsing it into either bucket would let an
208
- unchecked run read as a checked one.
209
-
210
- ## Related
211
-
212
- - [`workflow.md`](./workflow.md) — the kernel that runs verifiers as nodes
213
- - [`observability.md`](./observability.md) — how verdicts reach the run ledger
214
- - [`ultraapp.md`](./ultraapp.md) — the default contract in context
183
+ is not a failure and it is not a pass. `clawo runs` prints it as `—` and
184
+ `clawo workflow list` as `unchecked`, and the `clawo runs` summary line says so in
185
+ words, because collapsing it into either bucket would let an unchecked run read
186
+ as a checked one. A verdict that expired (see below) also reads `unverified`.
215
187
 
216
188
  ## What the guarantee is, precisely
217
189
 
@@ -225,7 +197,6 @@ unchecked run read as a checked one.
225
197
  - Outside a git repository the digest is unavailable. Nothing running after the
226
198
  checks means the verdict stands; something running after it means we decline
227
199
  to vouch, and the run says so.
228
-
229
200
  - If an abandoned attempt (a node past its timeout, which cannot be killed) is
230
201
  still running when the run ends, the outcome is `unverified` with the reason
231
202
  recorded. The runtime will not vouch for a tree something may still be writing
@@ -236,3 +207,9 @@ A node past its timeout keeps running, and if it outlives the short settle
236
207
  window its writes land after the last digest — the run will have said
237
208
  `unverified`, but the file is still changed. Give such nodes a timeout they will
238
209
  not hit, or make their writes safe to arrive late.
210
+
211
+ ## Related
212
+
213
+ - [`workflow.md`](./workflow.md) — the kernel that runs verifiers as nodes
214
+ - [`observability.md`](./observability.md) — how verdicts reach the run ledger
215
+ - [`ultraapp.md`](./ultraapp.md) — the default contract in context
@@ -1,50 +1,18 @@
1
1
  # Workflow kernel — durable runs
2
2
 
3
3
  A durable executor for workflow runs: what is running, what happens when a step
4
- fails, when to stop, and — the part none of the previous state machines had — how
5
- to come back after the process dies.
4
+ fails, when to stop, and how to come back after the process dies.
6
5
 
7
6
  ## Every mode runs on it
8
7
 
9
8
  `council_start`, `fanout_start`, `ultraplan_start`, `ultrareview_start` and
10
- `autoloop_start` all create a kernel run. Their tool signatures are unchanged and
11
- their result shapes are unchanged — `CouncilSession`, `FanoutSession`,
12
- `UltraplanResult`, `UltrareviewResult`, `AutoloopState` are now _projected_ from
13
- the run record rather than held in a map.
14
-
15
- The engines that do the work — `Council`, `Fanout`, the autoloop
16
- planner/coder/reviewer dispatcher — are untouched. What they lost is ownership of
17
- a lifecycle. Deleted along the way:
18
-
19
- | Gone | Was |
20
- | ------------------ | -------------------------------------------------------------------------------- |
21
- | 5 result maps | `councils`, `fanouts`, `ultraplans`, `ultrareviews`, `autoloops` |
22
- | 4 eviction timers | a 30-minute TTL per mode, three of them separate implementations |
23
- | 1 poller | ultrareview asking the fan-out every 5s whether it had finished |
24
- | 2 fences | `_startingAutoloops` / `_deletingAutoloops`, guarding a shared map |
25
- | 2 disk enumerators | a regex over council markdown transcripts; a bespoke JSONL registry for autoloop |
26
-
27
- Concretely, three bugs went with them: a fan-out's results vanished 30 minutes
28
- after it finished; an ultraplan still running when its TTL fired was rewritten as
29
- `error: 'Timed out (TTL expired)'` and deleted, so a long plan could be destroyed
30
- by its own eviction timer; and ultrareview's correctness depended on the
31
- fan-out's TTL — evict first and its poll threw, the interval was cleared, and the
32
- review stayed `running` forever.
33
-
34
- ## Why this exists
35
-
36
- Through 5.1.0 each mode carried its own machinery. The same "start in the
37
- background, poll by id, evict after 30 minutes" was written four separate times
38
- (`council`, `fanout`, `ultraplan`, `ultrareview`), with four timer sites and six
39
- status vocabularies that did not overlap. Cross-process listing was implemented
40
- three incompatible ways — council scraped its own markdown transcripts with a
41
- regex, autoloop read a JSONL registry, ultraapp walked a store directory.
42
-
43
- More to the point, most of it was not durable. A fan-out wrote nothing to disk at
44
- all and its results vanished after 30 minutes. Ultraplan and ultrareview were
45
- entirely in memory. A council that crashed mid-round left worktrees and branches
46
- on disk with no index pointing at them. UltraApp's build queue documented that it
47
- did not persist, so a restart mid-build failed the build.
9
+ `autoloop_start` all create a kernel run. Their result shapes —
10
+ `CouncilSession`, `FanoutSession`, `UltraplanResult`, `UltrareviewResult`,
11
+ `AutoloopState` — are _projected_ from the run record.
12
+
13
+ The mode engines (`Council`, `Fanout`, the autoloop planner/coder/reviewer
14
+ dispatcher) do the work; the kernel owns lifecycle, persistence and listing. A
15
+ run's record stays on disk until you delete it.
48
16
 
49
17
  ## Durability contract
50
18
 
@@ -92,75 +60,46 @@ names four things, and all four are checked on every durable write:
92
60
  - **`commit(guard, batch)` is the only way to change anything durable.**
93
61
  Checkpoints, events and node artifacts all go through it, inside one `O_EXCL`
94
62
  critical section that verifies the guard first. The raw writers are not
95
- exported, so there is no path around it — the previous version stated this rule
96
- in a comment while the engine wrote checkpoints directly from `start`,
97
- `resume`, `publish` and `setChild`, and a rule enforced by a comment is not a
98
- rule.
63
+ exported, so there is no path around it.
99
64
  - **A batch lands whole.** It is staged in a scratch directory and published by a
100
65
  single atomic directory rename; the rename is the commit point, and what
101
66
  follows is replayable application of an already-committed transaction. A reader
102
67
  finishes any transaction a crashed owner left, and applying is idempotent — the
103
68
  manifest records the event log's length from before, so recovery truncates and
104
- re-appends rather than duplicating. Without this, `committed` meant "most of it
105
- was attempted": the event append swallowed its own errors, so a checkpoint
106
- could land with its events silently dropped, and a batch that failed partway
107
- left the artifacts it had already written behind.
69
+ re-appends rather than duplicating.
108
70
  - **Creating a run and claiming it are one step.** The run directory is made with
109
71
  a non-recursive `mkdir`, which _is_ the claim — it fails for everyone but the
110
- first caller. Asking `runExists()` and then creating is a check-then-write
111
- race, and it lost: two processes creating the same id 80 times both "succeeded"
112
- 76 times, leaving one workflow executing under another's `spec.json`.
72
+ first caller, so two processes creating the same id cannot both succeed.
113
73
  - **The lock is exclusive, and release is not "unlink that path".** A vanished
114
74
  lock is retried rather than treated as stale debris; a genuinely stale one is
115
75
  broken by atomic rename; and a holder removes the lock file only if it is still
116
- the one it created. Getting any of those wrong puts two callers in the section
117
- at once, and the symptom is not an error — it is a committed transaction being
118
- emptied by the other caller's cleanup, so writes vanish and the run wedges.
76
+ the one it created.
119
77
  - **A published transaction is authoritative before it is applied.** Readers
120
78
  finish any pending transaction first, and refuse rather than hand back the
121
79
  older checkpoint if it cannot be applied. Applying carries a marker written
122
80
  after the last data step, so a failure during cleanup cannot make a healthy
123
81
  transaction permanently unapplicable.
124
- - **The lock is exclusive, and release is not "unlink that path".** A vanished
125
- lock is retried rather than treated as stale debris; a genuinely stale one is
126
- broken by atomic rename; and a holder removes the lock file only if it is still
127
- the one it created. Getting any of those wrong puts two callers in the section
128
- at once, and the symptom is not an error — it is a committed transaction being
129
- emptied by the other caller's cleanup, so writes vanish and the run wedges.
130
- - **A published transaction is authoritative before it is applied.** Readers
131
- finish any pending transaction first, and refuse rather than hand back the
132
- older checkpoint if it cannot be applied. Applying carries a marker written
133
- after its last data step, so a failure during cleanup cannot make a healthy
134
- transaction permanently unapplicable.
135
- - **`delete` claims before removing.** Releasing the lease first opened a window
136
- in which another process could legally resume the run, only for this one to
137
- remove the directory under its new owner.
82
+ - **`delete` claims before removing**, so no other process can resume the run
83
+ while its directory is being removed.
138
84
  - **Contention is not a takeover.** `commit` reports `committed`, `superseded` or
139
85
  `blocked`, and only `superseded` is permanent. The lock waits briefly rather
140
86
  than failing on sight, and an owner that still cannot write stops _and hands
141
- its claim back_ — because a live local pid is never judged stale, so a lease
142
- left behind by a stopped run can never be taken over and the run is lost for
143
- good. Collapsing the two into one boolean is what made a millisecond of
144
- contention wedge a run permanently.
87
+ its claim back_ — a live local pid is never judged stale, so a lease left
88
+ behind by a stopped run could otherwise never be taken over.
145
89
  - **Copy-on-write.** A change is applied to a clone, committed, and adopted only
146
- if the disk accepted it. So a superseded owner does not merely fail to
147
- persist — the record it hands back to its own caller stops advancing too.
148
- Refusing the write while returning a record that says `completed` is the same
149
- claim one layer up, and callers read the record.
90
+ if the disk accepted it, so the record a superseded owner hands back to its
91
+ caller stops advancing too.
150
92
  - **A deleted run id is a new run.** The fence lives in `incarnation.json`, which
151
93
  survives `releaseLease` (so the counter never restarts while the run exists)
152
94
  and dies with the run directory (so the next run under the same id gets a new
153
- random `incarnationId`). Without that, deleting a run and reusing its id reset
154
- the fence to 1, and an abandoned attempt still holding fence 1 became valid a
155
- second time — a textbook ABA, and not hypothetical, because a timed-out attempt
156
- outlives its run by construction.
95
+ random `incarnationId`). An abandoned attempt from a deleted run therefore
96
+ cannot write into a new run that reuses its id.
157
97
  - **Re-acquiring supersedes.** A second claim, even by the same owner, mints a
158
98
  new `acquisitionId` and kills the previous guard.
159
99
  - **Owner identity is not the pid.** Two kernels in one process share a pid;
160
100
  each has its own owner id, or both would read the other's claim as their own.
161
- - **Atomic acquisition.** The check and the write happen inside the lock.
162
- Read-then-write let two processes both see "free" and both conclude they had
163
- it, which is the failure a lease exists to prevent.
101
+ - **Atomic acquisition.** The check and the write happen inside the lock, so two
102
+ processes cannot both see "free" and both claim the run.
164
103
  - **An independent heartbeat.** Renewed on a timer, not only at checkpoints: a
165
104
  run executing one long node makes no checkpoints, and must not look abandoned
166
105
  for it. On the same host a live pid is the authority and is never judged stale
@@ -173,7 +112,8 @@ A second process trying to resume a run someone else is executing is refused by
173
112
  name, with the owner's pid and host in the message.
174
113
 
175
114
  One thing deliberately sits outside the guard: **evidence bundles**. They are
176
- written by the verifier as its checks run, under `evidence/<node>-<attempt>/`,
115
+ written by the verifier as its checks run, under
116
+ `evidence/<node>-v<visit>-<attempt>/`,
177
117
  and they are append-only artifacts, never read as state. What makes a bundle
178
118
  authoritative is the run record's `evidenceId` pointing at it, and that reference
179
119
  _is_ committed under the guard. So a bundle left behind by an owner that has been
@@ -205,7 +145,7 @@ Every node takes `retry: { max, backoffMs }`, `timeoutMs`, and
205
145
  On `fanout` and `council`, `timeoutMs` bounds the whole node and `agentTimeoutMs` one
206
146
  agent's send. They differ because agents beyond the free session slots wait for one, so
207
147
  the node can run several agents' worth of time. Without `agentTimeoutMs`, `timeoutMs`
208
- serves as both, as it did before the field existed. `fanout_start`, `ultrareview_start`,
148
+ serves as both. `fanout_start`, `ultrareview_start`,
209
149
  `council_start` and the built-in `fanout`, `council` and `solve` workflows set both: the
210
150
  agent's own default, and a node bound for the worst case — one agent at a time, every
211
151
  retry taken.
@@ -235,8 +175,8 @@ next node**, so a gate that failed to match simply hands control to the router
235
175
  after it, which then matches on its own. Chaining reads as AND and behaves as
236
176
  "whatever the last router says".
237
177
 
238
- `maxNodeVisits` (default 50) bounds every loop as a backstop. Use `visits_lt` for
239
- the actual budget — the backstop failing a run is a bug report, not a feature.
178
+ `maxNodeVisits` (default 50) bounds every loop as a safety backstop. Budget loops
179
+ with `visits_lt`.
240
180
 
241
181
  ## Example: repair until green
242
182
 
@@ -288,7 +228,7 @@ test suite would double the most expensive part of the run to learn nothing new.
288
228
  `workflow_start` accepts `template` instead of `spec`:
289
229
 
290
230
  - **`solve`** — the shape above: triage → (optional human gate) → implement →
291
- verify → repair-until-green → optional review.
231
+ (optional review) → verify → repair-until-green.
292
232
  - **`council`** — one council node, plus the implicit verifier when a contract is
293
233
  declared.
294
234
  - **`fanout`** — one fan-out node.
@@ -299,20 +239,13 @@ These are ordinary specs, not privileged paths.
299
239
 
300
240
  A passing verdict only stands while it still describes the tree.
301
241
 
302
- The digest is over **content**, not status: HEAD, the full `git diff HEAD`, and
303
- the bytes of every untracked file. An earlier version hashed
304
- `git status --porcelain`, which reports a file's state rather than its bytes — so
305
- a file already `M` before the checks and rewritten afterwards produced an
306
- identical digest, and the commonest case (an agent editing a file it had already
307
- edited) was the one it could not see.
308
-
309
242
  This cannot be enforced by inspecting the spec — a router can send control
310
243
  anywhere, so which node runs last is not a property of the graph. And "nothing
311
244
  may follow the verifier" would be the wrong rule anyway: what matters is not that
312
245
  a node ran, but that the tree moved. So the kernel measures. Each evidence bundle
313
- records a digest of the working tree (`git rev-parse HEAD` plus
314
- `git status --porcelain`), and when the run ends, if any workspace-touching node
315
- ran after the verdict, the digest is recomputed.
246
+ records a content digest of the working tree (HEAD, the full `git diff HEAD`,
247
+ and the path and bytes of every untracked file), and when the run ends, if any
248
+ workspace-touching node ran after the verdict, the digest is recomputed.
316
249
 
317
250
  If it moved, the outcome drops from `verified` to `unverified` with the reason
318
251
  recorded on the run. Not `refuted` — no check failed; we simply stopped knowing,
@@ -323,9 +256,8 @@ checks means the verdict stands regardless (a contract that passed in a plain
323
256
  directory passed); something running after it means we cannot vouch, and the run
324
257
  says so.
325
258
 
326
- The built-in `solve` template puts its reviewer fan-out **before** the gate for
327
- this reason. It shipped the other way round first, which let reviewers edit a
328
- tree the verifier had already signed off while the run still reported `verified`.
259
+ The built-in `solve` template puts its reviewer fan-out **before** the verifier,
260
+ so anything the reviewers change is covered by the verdict.
329
261
 
330
262
  ## Completion
331
263
 
@@ -373,11 +305,12 @@ belong before the task.
373
305
  - Cancel and a node timeout are different things. A timeout is a node failure
374
306
  (it still gets its retries and still honours `onFailure`); cancelling ends the
375
307
  run.
376
- - Nothing prunes run directories. Delete them yourself, or with
377
- `workflowDelete`.
308
+ - Nothing prunes run directories. Delete them yourself
309
+ (`rm -rf ~/.claw-orchestrator/wf/<runId>`), or programmatically with
310
+ `SessionManager.workflowDelete(runId)`.
378
311
 
379
312
  ## Related
380
313
 
381
314
  - [`verification.md`](./verification.md) — contracts, checks, evidence
382
315
  - [`observability.md`](./observability.md) — how a run's verdict reaches the ledger
383
- - [`council.md`](./council.md), [`autoloop.md`](./autoloop.md), [`ultraapp.md`](./ultraapp.md) — the modes, and what changed for each
316
+ - [`council.md`](./council.md), [`autoloop.md`](./autoloop.md), [`ultraapp.md`](./ultraapp.md) — the modes that run on the kernel