@enderfga/claw-orchestrator 5.1.0 → 6.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (147) hide show
  1. package/README.md +26 -26
  2. package/dist/bin/cli.js +107 -1
  3. package/dist/bin/cli.js.map +1 -1
  4. package/dist/src/acp-server.d.ts +5 -5
  5. package/dist/src/acp-server.js +3 -3
  6. package/dist/src/acp-server.js.map +1 -1
  7. package/dist/src/autoloop/dispatcher.d.ts +22 -0
  8. package/dist/src/autoloop/dispatcher.js +71 -13
  9. package/dist/src/autoloop/dispatcher.js.map +1 -1
  10. package/dist/src/autoloop/messages.d.ts +10 -0
  11. package/dist/src/autoloop/messages.js.map +1 -1
  12. package/dist/src/autoloop/runner.js +6 -0
  13. package/dist/src/autoloop/runner.js.map +1 -1
  14. package/dist/src/constants.d.ts +0 -6
  15. package/dist/src/constants.js +0 -6
  16. package/dist/src/constants.js.map +1 -1
  17. package/dist/src/council.d.ts +15 -0
  18. package/dist/src/council.js +48 -35
  19. package/dist/src/council.js.map +1 -1
  20. package/dist/src/dashboard/index.html +191 -6
  21. package/dist/src/embedded-server.js +132 -9
  22. package/dist/src/embedded-server.js.map +1 -1
  23. package/dist/src/fanout.d.ts +30 -1
  24. package/dist/src/fanout.js +32 -3
  25. package/dist/src/fanout.js.map +1 -1
  26. package/dist/src/index.js +359 -4
  27. package/dist/src/index.js.map +1 -1
  28. package/dist/src/kernel/agent-step.d.ts +59 -0
  29. package/dist/src/kernel/agent-step.js +100 -0
  30. package/dist/src/kernel/agent-step.js.map +1 -0
  31. package/dist/src/kernel/conditions.d.ts +11 -0
  32. package/dist/src/kernel/conditions.js +24 -0
  33. package/dist/src/kernel/conditions.js.map +1 -0
  34. package/dist/src/kernel/engine.d.ts +319 -0
  35. package/dist/src/kernel/engine.js +1047 -0
  36. package/dist/src/kernel/engine.js.map +1 -0
  37. package/dist/src/kernel/exec.d.ts +43 -0
  38. package/dist/src/kernel/exec.js +112 -0
  39. package/dist/src/kernel/exec.js.map +1 -0
  40. package/dist/src/kernel/file-lock.d.ts +50 -0
  41. package/dist/src/kernel/file-lock.js +135 -0
  42. package/dist/src/kernel/file-lock.js.map +1 -0
  43. package/dist/src/kernel/nodes/agent.d.ts +4 -0
  44. package/dist/src/kernel/nodes/agent.js +35 -0
  45. package/dist/src/kernel/nodes/agent.js.map +1 -0
  46. package/dist/src/kernel/nodes/autoloop.d.ts +78 -0
  47. package/dist/src/kernel/nodes/autoloop.js +75 -0
  48. package/dist/src/kernel/nodes/autoloop.js.map +1 -0
  49. package/dist/src/kernel/nodes/council.d.ts +12 -0
  50. package/dist/src/kernel/nodes/council.js +88 -0
  51. package/dist/src/kernel/nodes/council.js.map +1 -0
  52. package/dist/src/kernel/nodes/fanout.d.ts +11 -0
  53. package/dist/src/kernel/nodes/fanout.js +63 -0
  54. package/dist/src/kernel/nodes/fanout.js.map +1 -0
  55. package/dist/src/kernel/nodes/human-gate.d.ts +4 -0
  56. package/dist/src/kernel/nodes/human-gate.js +7 -0
  57. package/dist/src/kernel/nodes/human-gate.js.map +1 -0
  58. package/dist/src/kernel/nodes/index.d.ts +12 -0
  59. package/dist/src/kernel/nodes/index.js +21 -0
  60. package/dist/src/kernel/nodes/index.js.map +1 -0
  61. package/dist/src/kernel/nodes/router.d.ts +4 -0
  62. package/dist/src/kernel/nodes/router.js +12 -0
  63. package/dist/src/kernel/nodes/router.js.map +1 -0
  64. package/dist/src/kernel/nodes/subflow.d.ts +13 -0
  65. package/dist/src/kernel/nodes/subflow.js +38 -0
  66. package/dist/src/kernel/nodes/subflow.js.map +1 -0
  67. package/dist/src/kernel/nodes/ultraapp.d.ts +60 -0
  68. package/dist/src/kernel/nodes/ultraapp.js +62 -0
  69. package/dist/src/kernel/nodes/ultraapp.js.map +1 -0
  70. package/dist/src/kernel/nodes/verifier.d.ts +14 -0
  71. package/dist/src/kernel/nodes/verifier.js +84 -0
  72. package/dist/src/kernel/nodes/verifier.js.map +1 -0
  73. package/dist/src/kernel/projections.d.ts +42 -0
  74. package/dist/src/kernel/projections.js +133 -0
  75. package/dist/src/kernel/projections.js.map +1 -0
  76. package/dist/src/kernel/repo.d.ts +13 -0
  77. package/dist/src/kernel/repo.js +64 -0
  78. package/dist/src/kernel/repo.js.map +1 -0
  79. package/dist/src/kernel/secrets.d.ts +25 -0
  80. package/dist/src/kernel/secrets.js +48 -0
  81. package/dist/src/kernel/secrets.js.map +1 -0
  82. package/dist/src/kernel/store.d.ts +225 -0
  83. package/dist/src/kernel/store.js +838 -0
  84. package/dist/src/kernel/store.js.map +1 -0
  85. package/dist/src/kernel/templates/index.d.ts +140 -0
  86. package/dist/src/kernel/templates/index.js +266 -0
  87. package/dist/src/kernel/templates/index.js.map +1 -0
  88. package/dist/src/kernel/types.d.ts +326 -0
  89. package/dist/src/kernel/types.js +19 -0
  90. package/dist/src/kernel/types.js.map +1 -0
  91. package/dist/src/run-ledger.d.ts +57 -3
  92. package/dist/src/run-ledger.js +45 -2
  93. package/dist/src/run-ledger.js.map +1 -1
  94. package/dist/src/session-manager.d.ts +176 -129
  95. package/dist/src/session-manager.js +652 -603
  96. package/dist/src/session-manager.js.map +1 -1
  97. package/dist/src/types.d.ts +33 -3
  98. package/dist/src/ultraapp/build.d.ts +117 -3
  99. package/dist/src/ultraapp/build.js +319 -3
  100. package/dist/src/ultraapp/build.js.map +1 -1
  101. package/dist/src/ultraapp/contract.d.ts +52 -0
  102. package/dist/src/ultraapp/contract.js +83 -0
  103. package/dist/src/ultraapp/contract.js.map +1 -0
  104. package/dist/src/ultraapp/conventions.js +9 -2
  105. package/dist/src/ultraapp/conventions.js.map +1 -1
  106. package/dist/src/ultraapp/fix-on-failure.d.ts +21 -2
  107. package/dist/src/ultraapp/fix-on-failure.js +46 -62
  108. package/dist/src/ultraapp/fix-on-failure.js.map +1 -1
  109. package/dist/src/ultraapp/manager.d.ts +107 -2
  110. package/dist/src/ultraapp/manager.js +305 -86
  111. package/dist/src/ultraapp/manager.js.map +1 -1
  112. package/dist/src/verify/baseline.d.ts +73 -0
  113. package/dist/src/verify/baseline.js +186 -0
  114. package/dist/src/verify/baseline.js.map +1 -0
  115. package/dist/src/verify/contract.d.ts +116 -0
  116. package/dist/src/verify/contract.js +142 -0
  117. package/dist/src/verify/contract.js.map +1 -0
  118. package/dist/src/verify/evidence.d.ts +61 -0
  119. package/dist/src/verify/evidence.js +133 -0
  120. package/dist/src/verify/evidence.js.map +1 -0
  121. package/dist/src/verify/runner.d.ts +63 -0
  122. package/dist/src/verify/runner.js +317 -0
  123. package/dist/src/verify/runner.js.map +1 -0
  124. package/openclaw.plugin.json +8 -0
  125. package/package.json +2 -2
  126. package/skills/SKILL.md +120 -79
  127. package/skills/references/acp.md +17 -17
  128. package/skills/references/autoloop.md +139 -65
  129. package/skills/references/claude-cli-tracking.md +4 -4
  130. package/skills/references/cli.md +101 -59
  131. package/skills/references/council.md +109 -37
  132. package/skills/references/dashboard.md +34 -6
  133. package/skills/references/getting-started.md +13 -13
  134. package/skills/references/inbox.md +4 -4
  135. package/skills/references/mcp.md +39 -34
  136. package/skills/references/multi-engine.md +51 -47
  137. package/skills/references/observability.md +88 -28
  138. package/skills/references/openai-compat.md +39 -39
  139. package/skills/references/sessions.md +43 -25
  140. package/skills/references/tools.md +402 -309
  141. package/skills/references/ultra.md +45 -45
  142. package/skills/references/ultraapp.md +126 -50
  143. package/skills/references/verification.md +187 -0
  144. package/skills/references/workflow.md +362 -0
  145. package/dist/src/ultraapp/fix-on-failure-session.d.ts +0 -23
  146. package/dist/src/ultraapp/fix-on-failure-session.js +0 -51
  147. package/dist/src/ultraapp/fix-on-failure-session.js.map +0 -1
@@ -20,8 +20,8 @@ Runs in background — poll with `ultraplan_status`.
20
20
  ```typescript
21
21
  const plan = manager.ultraplanStart('Add OAuth2 support with Google and GitHub providers', {
22
22
  cwd: '/path/to/project',
23
- model: 'opus', // default
24
- timeout: 1800000, // 30 min default
23
+ model: 'opus', // default
24
+ timeout: 1800000, // 30 min default
25
25
  });
26
26
 
27
27
  console.log(`Plan ID: ${plan.id}`);
@@ -35,18 +35,18 @@ if (status?.status === 'completed') {
35
35
 
36
36
  ### Tools
37
37
 
38
- | Tool | Description |
39
- |------|-------------|
40
- | `ultraplan_start` | Start planning session (background) |
41
- | `ultraplan_status` | Get status and plan text |
38
+ | Tool | Description |
39
+ | ------------------ | ----------------------------------- |
40
+ | `ultraplan_start` | Start planning session (background) |
41
+ | `ultraplan_status` | Get status and plan text |
42
42
 
43
43
  ### Configuration
44
44
 
45
- | Parameter | Default | Description |
46
- |-----------|---------|-------------|
47
- | `model` | `opus` | Model to use for planning |
48
- | `cwd` | `process.cwd()` | Project directory to explore |
49
- | `timeout` | 1,800,000 ms (30 min) | Maximum planning time |
45
+ | Parameter | Default | Description |
46
+ | --------- | --------------------- | ---------------------------- |
47
+ | `model` | `opus` | Model to use for planning |
48
+ | `cwd` | `process.cwd()` | Project directory to explore |
49
+ | `timeout` | 1,800,000 ms (30 min) | Maximum planning time |
50
50
 
51
51
  Results remain queryable for 30 minutes after completion.
52
52
 
@@ -65,36 +65,36 @@ A fleet of specialized bug-hunting agents that review your codebase in parallel,
65
65
 
66
66
  ### Available Review Angles (20)
67
67
 
68
- | Agent | Focus |
69
- |-------|-------|
70
- | SecurityReviewer | Injection, auth flaws, data exposure, OWASP top 10 |
71
- | LogicReviewer | Off-by-one, race conditions, null handling, edge cases |
68
+ | Agent | Focus |
69
+ | ------------------- | -------------------------------------------------------- |
70
+ | SecurityReviewer | Injection, auth flaws, data exposure, OWASP top 10 |
71
+ | LogicReviewer | Off-by-one, race conditions, null handling, edge cases |
72
72
  | PerformanceReviewer | O(n^2) loops, memory leaks, missing caching, N+1 queries |
73
- | APIReviewer | Inconsistent interfaces, missing validation, error gaps |
74
- | TestReviewer | Untested paths, missing edge case tests, flaky patterns |
75
- | TypeReviewer | `any` casts, unsafe assertions, missing null checks |
76
- | ConcurrencyReviewer | Race conditions, deadlocks, async error handling |
77
- | ErrorReviewer | Swallowed errors, missing try/catch, crash paths |
78
- | DependencyReviewer | Outdated packages, CVEs, unnecessary deps |
79
- | ReadabilityReviewer | Unclear naming, complex functions, dead code |
80
- | DataReviewer | Data validation, schema mismatches, encoding |
81
- | ConfigReviewer | Hardcoded values, missing env vars, insecure defaults |
82
- | ScalabilityReviewer | Single points of failure, unbounded growth |
83
- | DocReviewer | Outdated docs, missing API docs, misleading comments |
84
- | A11yReviewer | ARIA labels, keyboard nav, color contrast |
85
- | I18nReviewer | Hardcoded strings, locale handling, RTL support |
86
- | NetworkReviewer | Missing timeouts, retry logic, connection pooling |
87
- | AuthReviewer | Token handling, CSRF, permission checks |
88
- | CryptoReviewer | Weak algorithms, key management, RNG |
89
- | MemoryReviewer | Memory leaks, circular references, stream handling |
73
+ | APIReviewer | Inconsistent interfaces, missing validation, error gaps |
74
+ | TestReviewer | Untested paths, missing edge case tests, flaky patterns |
75
+ | TypeReviewer | `any` casts, unsafe assertions, missing null checks |
76
+ | ConcurrencyReviewer | Race conditions, deadlocks, async error handling |
77
+ | ErrorReviewer | Swallowed errors, missing try/catch, crash paths |
78
+ | DependencyReviewer | Outdated packages, CVEs, unnecessary deps |
79
+ | ReadabilityReviewer | Unclear naming, complex functions, dead code |
80
+ | DataReviewer | Data validation, schema mismatches, encoding |
81
+ | ConfigReviewer | Hardcoded values, missing env vars, insecure defaults |
82
+ | ScalabilityReviewer | Single points of failure, unbounded growth |
83
+ | DocReviewer | Outdated docs, missing API docs, misleading comments |
84
+ | A11yReviewer | ARIA labels, keyboard nav, color contrast |
85
+ | I18nReviewer | Hardcoded strings, locale handling, RTL support |
86
+ | NetworkReviewer | Missing timeouts, retry logic, connection pooling |
87
+ | AuthReviewer | Token handling, CSRF, permission checks |
88
+ | CryptoReviewer | Weak algorithms, key management, RNG |
89
+ | MemoryReviewer | Memory leaks, circular references, stream handling |
90
90
 
91
91
  ### Usage
92
92
 
93
93
  ```typescript
94
94
  const review = manager.ultrareviewStart('/path/to/project', {
95
- agentCount: 10, // use 10 of the 20 angles
96
- maxDurationMinutes: 15, // 15 min timeout per agent
97
- model: 'sonnet', // model for all reviewers
95
+ agentCount: 10, // use 10 of the 20 angles
96
+ maxDurationMinutes: 15, // 15 min timeout per agent
97
+ model: 'sonnet', // model for all reviewers
98
98
  focus: 'Find security and performance bugs',
99
99
  });
100
100
 
@@ -109,18 +109,18 @@ if (status?.status === 'completed') {
109
109
 
110
110
  ### Tools
111
111
 
112
- | Tool | Description |
113
- |------|-------------|
114
- | `ultrareview_start` | Launch reviewer fleet (background) |
115
- | `ultrareview_status` | Get status and findings |
112
+ | Tool | Description |
113
+ | -------------------- | ---------------------------------- |
114
+ | `ultrareview_start` | Launch reviewer fleet (background) |
115
+ | `ultrareview_status` | Get status and findings |
116
116
 
117
117
  ### Configuration
118
118
 
119
- | Parameter | Default | Range | Description |
120
- |-----------|---------|-------|-------------|
121
- | `agentCount` | 5 | 1-20 | Number of reviewer agents |
122
- | `maxDurationMinutes` | 10 | 5-25 | Per-agent timeout |
123
- | `model` | session default | — | Model for all reviewers |
124
- | `focus` | bugs + security + quality | — | Review focus description |
119
+ | Parameter | Default | Range | Description |
120
+ | -------------------- | ------------------------- | ----- | ------------------------- |
121
+ | `agentCount` | 5 | 1-20 | Number of reviewer agents |
122
+ | `maxDurationMinutes` | 10 | 5-25 | Per-agent timeout |
123
+ | `model` | session default | — | Model for all reviewers |
124
+ | `focus` | bugs + security + quality | — | Review focus description |
125
125
 
126
126
  The council runs with `maxRounds: 2` — one round to find bugs, one to cross-review. Results remain queryable for 30 minutes.
@@ -32,30 +32,58 @@ interview ─► queued ─► building ─► build-complete ─► deploying
32
32
  structural)
33
33
  ```
34
34
 
35
- | Mode | Meaning |
36
- |------|---------|
37
- | `interview` | AppSpec being filled by Q&A. Chat input goes to the interview Opus. |
38
- | `queued` | Build accepted, waiting for a slot in the FIFO build queue. |
39
- | `building` | Council writing code, fix-on-failure driving install/build/test. |
40
- | `build-complete` | Codebase ready, awaiting `deploy` step. |
41
- | `deploying` | Container/process being started, router map being updated. |
42
- | `done` | App live at `/forge/<slug>/`. Chat input now goes to the done-mode classifier. |
43
- | `failed` | Council didn't reach consensus, or fix-on-failure couldn't get the build green. |
35
+ | Mode | Meaning |
36
+ | ---------------- | ------------------------------------------------------------------------------- |
37
+ | `interview` | AppSpec being filled by Q&A. Chat input goes to the interview Opus. |
38
+ | `queued` | Build accepted, waiting for a slot in the FIFO build queue. |
39
+ | `building` | Council writing code, fix-on-failure driving install/build/test. |
40
+ | `build-complete` | Codebase ready, awaiting `deploy` step. |
41
+ | `deploying` | Container/process being started, router map being updated. |
42
+ | `done` | App live at `/forge/<slug>/`. Chat input now goes to the done-mode classifier. |
43
+ | `failed` | Council didn't reach consensus, or fix-on-failure couldn't get the build green. |
44
+
45
+ ### The build is a workflow run
46
+
47
+ Everything from `building` onward is a run on the durable kernel — the same one
48
+ `council_start`, `fanout_start` and `autoloop_start` use. The workflow has three
49
+ nodes:
50
+
51
+ | Node | Kind | Does |
52
+ | -------- | ----------------- | ------------------------------------------------------------ |
53
+ | `synth` | `ultraapp_synth` | The 3-agent council; snapshots consensus into `versions/v1/` |
54
+ | `build` | `verifier` | The build contract, with its fix-on-red loop and evidence |
55
+ | `deploy` | `ultraapp_deploy` | Starts the app, registers the route, runs the §7g gate |
56
+
57
+ Three consequences worth knowing:
58
+
59
+ - The run is listed by `workflow_list` and `clawo workflow list`, and shows up in
60
+ the dashboard's Runs tab, under the id `ultraapp-<run_id>`.
61
+ - **A crash between stages does not redo the stage before it.** A build that died
62
+ after the council resumes into the build stage rather than paying for the
63
+ council twice.
64
+ - The mode column above is a **projection** of that run's state, not a second
65
+ state machine. `interview` and `queued` are the exception — nothing is running
66
+ yet, so there is no run to project from.
67
+
68
+ The interview and the done-mode conversation are deliberately not workflows.
69
+ They are user-driven and open-ended, with no predetermined sequence of steps;
70
+ a workflow expressing them would be a router self-loop or a single node with a
71
+ state machine hidden inside it.
44
72
 
45
73
  ## Architectural conventions (§1–§7)
46
74
 
47
75
  Every generated app MUST satisfy these. They're embedded in the council
48
76
  super-task prompt verbatim from `src/ultraapp/conventions.ts`.
49
77
 
50
- | § | Topic | Headline rule |
51
- |---|-------|---------------|
52
- | 1 | Path-based deploy | Mount at `BASE_PATH=/forge/<slug>/`; in-app links MUST be relative. |
53
- | 2 | Async file-queue runtime | Exact endpoints: `GET /`, `POST /run`, `GET /status/:jobId`, `GET /result/:jobId`, `GET /health`. File-based job queue under `$DATA_DIR/jobs/<jobId>/`. NO database. Data path from `process.env.DATA_DIR ?? '/data'`. |
54
- | 3 | BYOK | If `runtime.needsLLM`, API keys live in browser localStorage and are sent direct to the provider. The server MUST NEVER receive the key (enforced by `eslint-plugin-no-server-keys`). |
55
- | 4 | Dockerfile + smoke test | Single multi-stage Dockerfile, `npm run smoke` drives one full job in < 90s using `examples[0].ref`. |
56
- | 5 | Council voting protocol | 3 agents in git worktrees, all-YES vote required, max 8 rounds. |
57
- | 6 | Tech stack | Modern TypeScript / JavaScript framework (Next.js, Vite + Hono, SvelteKit). NO Python, NO pure SSGs. |
58
- | 7 | **Frontend quality** | **Real styling system + real type hierarchy + four-state coverage on every async surface + drag-and-drop forms + appropriate result presentation + one deliberate theme.** §7g requires every agent to capture Chrome-headless screenshots at 1440×900 AND 375×812 and visually inspect the PNGs before voting YES — source-code review is explicitly insufficient evidence. |
78
+ | § | Topic | Headline rule |
79
+ | --- | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
80
+ | 1 | Path-based deploy | Mount at `BASE_PATH=/forge/<slug>/`; in-app links MUST be relative. |
81
+ | 2 | Async file-queue runtime | Exact endpoints: `GET /`, `POST /run`, `GET /status/:jobId`, `GET /result/:jobId`, `GET /health`. File-based job queue under `$DATA_DIR/jobs/<jobId>/`. NO database. Data path from `process.env.DATA_DIR ?? '/data'`. |
82
+ | 3 | BYOK | If `runtime.needsLLM`, API keys live in browser localStorage and are sent direct to the provider. The server MUST NEVER receive the key (enforced by `eslint-plugin-no-server-keys`). |
83
+ | 4 | Dockerfile + smoke test | Single multi-stage Dockerfile, `npm run smoke` drives one full job in < 90s using `examples[0].ref`. |
84
+ | 5 | Council voting protocol | 3 agents in git worktrees, all-YES vote required, max 8 rounds. |
85
+ | 6 | Tech stack | Modern TypeScript / JavaScript framework (Next.js, Vite + Hono, SvelteKit). NO Python, NO pure SSGs. |
86
+ | 7 | **Frontend quality** | **Real styling system + real type hierarchy + four-state coverage on every async surface + drag-and-drop forms + appropriate result presentation + one deliberate theme.** §7g requires every agent to capture Chrome-headless screenshots at 1440×900 AND 375×812 and visually inspect the PNGs before voting YES — source-code review is explicitly insufficient evidence. |
59
87
 
60
88
  ## Runtime modes
61
89
 
@@ -64,10 +92,10 @@ clawo serve --ultraapp-runtime host # default
64
92
  clawo serve --ultraapp-runtime docker # opt-in
65
93
  ```
66
94
 
67
- | Mode | Build | Run | Pros | Cons |
68
- |------|-------|-----|------|------|
69
- | `host` | `npm install && npm run build` | `npm start` (detached, `setsid`-equivalent) | Zero extra deps; works anywhere Node works; faster start. | No process isolation; deps installed under the user. |
70
- | `docker` | `docker build .` | `docker run -d --restart unless-stopped` | Per-app isolation, restart policy, image is the artefact. | Requires a running Docker daemon. |
95
+ | Mode | Build | Run | Pros | Cons |
96
+ | -------- | ------------------------------ | ------------------------------------------- | --------------------------------------------------------- | ---------------------------------------------------- |
97
+ | `host` | `npm install && npm run build` | `npm start` (detached, `setsid`-equivalent) | Zero extra deps; works anywhere Node works; faster start. | No process isolation; deps installed under the user. |
98
+ | `docker` | `docker build .` | `docker run -d --restart unless-stopped` | Per-app isolation, restart policy, image is the artefact. | Requires a running Docker daemon. |
71
99
 
72
100
  Both modes allocate a backend port in `[19100, 19999]`. The reverse-
73
101
  proxy router runs at port `19000` (auto-fallback up to `19099` if
@@ -100,23 +128,23 @@ persists to `~/.claw-orchestrator/host-procs.json`.
100
128
  All routes are served by the embedded server (default `:18796`), under
101
129
  `Authorization: Bearer <token>` from `~/.openclaw/server-token`.
102
130
 
103
- | Method + path | Purpose |
104
- |----|----|
105
- | `GET /ultraapp/list` | All runs with mode + createdAt. |
106
- | `POST /ultraapp/new` | Body: `{ firstMessage?: string }`. Returns `{ runId }`. |
107
- | `GET /ultraapp/<id>` | Full snapshot: spec + chat + state. |
108
- | `POST /ultraapp/<id>/answer` | Body: `{ value, freeform? }`. Submit interview answer. |
109
- | `POST /ultraapp/<id>/spec-edit` | Body: RFC 6902 patch ops. Edit the spec mid-interview. |
110
- | `POST /ultraapp/<id>/files` | Multipart upload to `examples/`. |
111
- | `GET /ultraapp/<id>/events` | SSE stream of build/chat events (mode pill, narrator, council activity). |
112
- | `POST /ultraapp/<id>/build` | Validate spec strictly + enqueue. |
113
- | `POST /ultraapp/<id>/build/cancel` | Abort the active build. |
114
- | `GET /ultraapp/<id>/artifacts` | List `versions/vN/`. |
115
- | `POST /ultraapp/<id>/start` | Start the deployed container/process for the active version. |
116
- | `POST /ultraapp/<id>/stop` | Stop without deleting. |
117
- | `POST /ultraapp/<id>/delete` | Stop + remove all per-run state. |
118
- | `POST /ultraapp/<id>/feedback` | Body: `{ text }`. Done-mode classifier routes cosmetic / spec-delta / structural. |
119
- | `POST /ultraapp/<id>/promote-version` | Body: `{ version: "vN" }`. Atomically swap deployed version. |
131
+ | Method + path | Purpose |
132
+ | ------------------------------------- | --------------------------------------------------------------------------------- |
133
+ | `GET /ultraapp/list` | All runs with mode + createdAt. |
134
+ | `POST /ultraapp/new` | Body: `{ firstMessage?: string }`. Returns `{ runId }`. |
135
+ | `GET /ultraapp/<id>` | Full snapshot: spec + chat + state. |
136
+ | `POST /ultraapp/<id>/answer` | Body: `{ value, freeform? }`. Submit interview answer. |
137
+ | `POST /ultraapp/<id>/spec-edit` | Body: RFC 6902 patch ops. Edit the spec mid-interview. |
138
+ | `POST /ultraapp/<id>/files` | Multipart upload to `examples/`. |
139
+ | `GET /ultraapp/<id>/events` | SSE stream of build/chat events (mode pill, narrator, council activity). |
140
+ | `POST /ultraapp/<id>/build` | Validate spec strictly + enqueue. |
141
+ | `POST /ultraapp/<id>/build/cancel` | Abort the active build. |
142
+ | `GET /ultraapp/<id>/artifacts` | List `versions/vN/`. |
143
+ | `POST /ultraapp/<id>/start` | Start the deployed container/process for the active version. |
144
+ | `POST /ultraapp/<id>/stop` | Stop without deleting. |
145
+ | `POST /ultraapp/<id>/delete` | Stop + remove all per-run state. |
146
+ | `POST /ultraapp/<id>/feedback` | Body: `{ text }`. Done-mode classifier routes cosmetic / spec-delta / structural. |
147
+ | `POST /ultraapp/<id>/promote-version` | Body: `{ version: "vN" }`. Atomically swap deployed version. |
120
148
 
121
149
  ## MCP tools (14)
122
150
 
@@ -137,11 +165,11 @@ ultraapp_start_container ultraapp_stop_container ultraapp_delete
137
165
  After the run reaches `done`, chat input goes to a per-run Haiku
138
166
  classifier. Three classes:
139
167
 
140
- | Class | Routes to | Behaviour |
141
- |-------|-----------|-----------|
142
- | `cosmetic` | Patcher | Opus generates a unified diff against the deployed worktree → `applyUnifiedDiff` → validate via fix-on-failure → on success snapshot to `versions/vN+1/`, on any failure restore the snapshot atomically and post the reason to chat. |
143
- | `spec-delta` | Focused interview | Flips mode back to `interview` with a bootstrap message that names the field(s) being changed. Completion auto-triggers a fresh `startBuild`. |
144
- | `structural` | Suggestion only | Posts a narrator note: "this sounds like a different app — click + New". |
168
+ | Class | Routes to | Behaviour |
169
+ | ------------ | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
170
+ | `cosmetic` | Patcher | Opus generates a unified diff against the deployed worktree → `applyUnifiedDiff` → validate via fix-on-failure → on success snapshot to `versions/vN+1/`, on any failure restore the snapshot atomically and post the reason to chat. |
171
+ | `spec-delta` | Focused interview | Flips mode back to `interview` with a bootstrap message that names the field(s) being changed. Completion auto-triggers a fresh `startBuild`. |
172
+ | `structural` | Suggestion only | Posts a narrator note: "this sounds like a different app — click + New". |
145
173
 
146
174
  To swap which version is live, use `promote-version` (HTTP) or
147
175
  `ultraapp_promote_version` (MCP) — the router map and host-procs map
@@ -190,14 +218,62 @@ open "http://127.0.0.1:18796/dashboard?token=$(cat ~/.openclaw/server-token)"
190
218
  curl http://127.0.0.1:19000/forge/<slug>/health
191
219
  ```
192
220
 
193
- ## Known limitations (v4.0.0)
221
+ ## Acceptance contract (6.0.0)
222
+
223
+ UltraApp is the one mode whose contract is on by default, because it already ran
224
+ most of these commands and because two of its documented gates were not actually
225
+ enforced.
226
+
227
+ **Build stage**, in the council's worktree:
228
+
229
+ ```
230
+ npm install → npm run build → npm test → [docker build] → npm run smoke
231
+ ```
232
+
233
+ `npm run smoke` is new here. §4 of the architectural conventions has always told
234
+ the council that the smoke test gates build success — it was never in the step
235
+ list, so the claim was false. A codebase without a working `scripts.smoke` now
236
+ fails its build, which is what the brief said all along.
237
+
238
+ **Deploy stage**, against the live URL: both §7g viewports (1440×900 and
239
+ 375×812) are captured by the orchestrator with headless Chrome and stored as run
240
+ evidence.
241
+
242
+ Evidence lands under the run directory:
243
+
244
+ ```
245
+ ~/.claw-orchestrator/ultraapps/<runId>/evidence/
246
+ build-01/ bundle.json, checks/*.log, diff.patch
247
+ deploy-01/ bundle.json, shot-desktop.png, shot-mobile.png
248
+ ```
249
+
250
+ ### What the visual gate does and does not do
251
+
252
+ It captures images and stores them, so "did anyone actually look" is now a file
253
+ on disk instead of an agent's claim. It does **not** compare pixels — judging
254
+ the rendering is still a reader's job, and the §7g instructions in the council
255
+ prompt remain the agents' responsibility.
256
+
257
+ It is **advisory by default**, so a host without Chrome does not lose a working
258
+ app to a missing browser. Set `CLAWO_ULTRAAPP_VISUAL_GATE=strict` to make a
259
+ failed capture block the deploy. Chrome is resolved from `CLAWO_CHROME_BIN`, then
260
+ the usual macOS app paths, then `PATH`.
261
+
262
+ ## Durable build queue (6.0.0)
263
+
264
+ The build queue is persisted to `<store>/build-queue.json` and restored on
265
+ startup. Its own comment used to say a restart mid-build meant "the build is
266
+ marked failed and the user can rerun"; in practice nothing was marked — the
267
+ pending list vanished along with any queued build the user was waiting on, with
268
+ no record it had been asked for.
269
+
270
+ A build that was in flight when the process died is **re-queued, not resumed**,
271
+ and goes to the front: each build starts from a fresh council worktree, so
272
+ re-running is safe and continuing a half-built tree is not.
273
+
274
+ ## Known limitations
194
275
 
195
276
  - The done-mode patcher loop occasionally hangs between
196
277
  feedback-classification and the patcher Opus session creation;
197
278
  cosmetic changes can be applied manually until the underlying race
198
279
  is fixed.
199
- - The §7g frontend gate currently relies on per-agent honesty about
200
- running the screenshot capture; agents that skip the inspection can
201
- still pass the smoke gate. A follow-up will plumb a server-side
202
- screenshot validator into the council verifier so the gate becomes
203
- structurally enforced rather than persona-enforced.
@@ -0,0 +1,187 @@
1
+ # Verification — acceptance contracts and evidence
2
+
3
+ Before 6.0.0, nothing in this runtime ever checked an agent's work. Every
4
+ "finished" signal was the agent grading itself:
5
+
6
+ - **Council** terminated when a regex found `[CONSENSUS: YES]` in agent prose.
7
+ - **Autoloop**'s `eval_output` was whatever the Coder passed to a tool call, and
8
+ the Reviewer that was supposed to catch fabrication had a sandbox containing
9
+ the iteration's artifacts and no code, so it could not re-derive anything.
10
+ - **UltraApp**'s frontend gate was a sentence in a persona string telling agents
11
+ to capture screenshots. There was no screenshot code anywhere in the project.
12
+ - The run ledger's `ok` was the engine's report on its own turn.
13
+
14
+ An **acceptance contract** is the opposite of all of that: a list of checks the
15
+ runtime executes and whose results it reads. A run that declares one cannot reach
16
+ `completed` unless every required check passes.
17
+
18
+ ## The one rule about where contracts come from
19
+
20
+ **A contract comes from the caller or from a mode default. Never from agent
21
+ output.** If an agent could declare its own checks, we would be back to
22
+ self-grading with more steps. Nothing in the kernel reads a contract out of a
23
+ node's result, and `normalizeContract()` drops anything it does not recognise, so
24
+ a contract that arrived through a tool call carries no fields the executor did
25
+ not model.
26
+
27
+ Concretely: there is no shell string anywhere. A `command` check is argv.
28
+
29
+ ## Check types
30
+
31
+ | Type | What it does | Passes when |
32
+ | ------------- | --------------------------------------------------------------------------- | ------------------------------------------------------------------------- |
33
+ | `command` | Runs argv in a directory | Exit code equals `expectExit` (default 0) |
34
+ | `http` | Polls a URL until the deadline | Status equals `expectStatus` (default 200) |
35
+ | `screenshot` | Captures the page at each viewport with headless Chrome and stores the PNGs | Every viewport produced a non-empty image |
36
+ | `diff_policy` | Compares the change set against the recorded baseline | File count, forbidden paths, and required paths all satisfied |
37
+ | `file` | Checks a path | Exists (or is absent when `exists: false`) and matches `matches` if given |
38
+
39
+ ```jsonc
40
+ {
41
+ "id": "ship-it",
42
+ "fixOnFailureRounds": 2,
43
+ "checks": [
44
+ { "type": "command", "cmd": "npm", "args": ["run", "build"], "timeoutMs": 600000 },
45
+ { "type": "command", "cmd": "npm", "args": ["test"] },
46
+ { "type": "command", "cmd": "npm", "args": ["run", "lint"], "required": false },
47
+ { "type": "diff_policy", "forbidPaths": ["configs", ".github"], "maxFiles": 40 },
48
+ ],
49
+ }
50
+ ```
51
+
52
+ ### `required`
53
+
54
+ Default true. A failing non-required check is recorded in the evidence bundle but
55
+ does not refute the run — use it for signals you want visible without making them
56
+ blocking.
57
+
58
+ (The predecessor of this module, `ultraapp/fix-on-failure.ts`, declared a
59
+ `required` field on its steps and never read it: every step was fatal. It is
60
+ honoured now.)
61
+
62
+ ### Timeouts
63
+
64
+ Every check has one, and the default is 10 minutes. This is not cosmetic — the
65
+ old pipeline had no timeout at all, so a wedged `npm test` hung the build
66
+ forever. A check that overruns is killed (its whole process group, with SIGKILL)
67
+ and recorded as failed with `timedOut: true`.
68
+
69
+ ### What `screenshot` does and does not claim
70
+
71
+ It captures images and stores them. It does **not** compare pixels, and it is not
72
+ visual regression testing. What changed in 6.0.0 is that the capture is performed
73
+ by the runtime rather than requested of an agent — so "did anyone actually look"
74
+ stops being a claim and starts being a file on disk. Judging the rendering is
75
+ still a human's or an agent's job.
76
+
77
+ Chrome is resolved from `CLAWO_CHROME_BIN`, then the usual macOS app paths, then
78
+ `PATH`. A host with no browser fails the check immediately rather than paying the
79
+ timeout to find out.
80
+
81
+ ### `diff_policy` and the baseline
82
+
83
+ `diff_policy` needs to know what the run changed, which needs a baseline. The
84
+ kernel records `git rev-parse HEAD` when a run starts and diffs against that.
85
+
86
+ The change set is **tracked changes ∪ untracked files**. That union matters: a
87
+ bare `git diff` lists tracked modifications only, so files an agent _created_ are
88
+ invisible to it — which is exactly how `autoloop`'s per-iteration `diff.patch`
89
+ used to miss every new file while `git add -A` committed them anyway.
90
+
91
+ `requirePaths: ["."]` means "the run must have changed something".
92
+
93
+ ## Evidence bundles
94
+
95
+ Every verifier attempt writes a directory under the run:
96
+
97
+ ```
98
+ ~/.claw-orchestrator/wf/<runId>/evidence/<evidenceId>/
99
+ bundle.json verdict, per-check results, changed files, base/head SHA
100
+ checks/<id>.log output tail for each failed check
101
+ diff.patch the patch the run produced, new files included
102
+ shot-*.png screenshots, when a screenshot check ran
103
+ ```
104
+
105
+ Read one with `clawo verify <runId>`, `GET /workflow/<id>/evidence`, or
106
+ `workflow_status` (which returns the id).
107
+
108
+ Bundle writes are best-effort. The verdict is already decided by the check
109
+ results, so a bundle that fails to land loses the record, never the answer.
110
+
111
+ ## Fix-on-red
112
+
113
+ `fixOnFailureRounds` spawns a repair session against the failing check and then
114
+ **re-runs the whole check list**. The fixer's own claim to have fixed it is
115
+ ignored; only the re-run decides. Set it to 0 (the default) to disable.
116
+
117
+ ## Per-mode defaults
118
+
119
+ | Mode | Contract | Notes |
120
+ | ------------------------ | ----------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- |
121
+ | **UltraApp** | **On by default** | Build: `npm install` → `npm run build` → `npm test` → (`docker build`) → **`npm run smoke`**. Deploy: **both §7g viewports captured** against the live URL. |
122
+ | **Council** | Caller-declared | Consensus votes are recorded on the run as advisory and no longer decide completion. |
123
+ | **Autoloop** | Caller-declared | With a contract, a Reviewer `advance` is held unless the checks pass; passing also fires `on_target_hit`. |
124
+ | **Fanout / Ultrareview** | Caller-declared | Per-agent `ok` now reads the engine's terminal verdict rather than "the call did not throw". |
125
+ | **Plain sessions** | None | Use `verify_run` to check work after the fact. |
126
+
127
+ Two of UltraApp's defaults close claims the project had been making without
128
+ backing them:
129
+
130
+ - `conventions.ts` §4 told every council that `npm run smoke` gated build
131
+ success. It was not in the step list at all. It is now.
132
+ - `ultraapp.md` recorded, as a known limitation since 4.0.0, that the §7g gate
133
+ "relies on per-agent honesty about running the screenshot capture", with a
134
+ server-side validator promised as a follow-up. That follow-up is this release.
135
+
136
+ The visual gate is **advisory by default** so that a host without Chrome does not
137
+ lose a working app to a missing browser. Set
138
+ `CLAWO_ULTRAAPP_VISUAL_GATE=strict` to make a failed capture block the deploy.
139
+
140
+ ## Verifying work that did not come through a workflow
141
+
142
+ ```bash
143
+ # Tool
144
+ verify_run({ cwd: "/repo", contract: { checks: [{ type: "command", cmd: "npm", args: ["test"] }] } })
145
+ ```
146
+
147
+ Use it for a plain `session_send` that edited a repo, or for anything from an
148
+ older version. The contract is yours; nothing is read from agent output.
149
+
150
+ ## Three outcomes, not two
151
+
152
+ A run ends `verified`, `refuted`, or `unverified`.
153
+
154
+ `unverified` means **no contract was declared and nothing checked the work**. It
155
+ is not a failure and it is not a pass. The CLI prints it as `—` and the summary
156
+ line says so in words, because collapsing it into either bucket would let an
157
+ unchecked run read as a checked one.
158
+
159
+ ## Related
160
+
161
+ - [`workflow.md`](./workflow.md) — the kernel that runs verifiers as nodes
162
+ - [`observability.md`](./observability.md) — how verdicts reach the run ledger
163
+ - [`ultraapp.md`](./ultraapp.md) — the default contract in context
164
+
165
+ ## What the guarantee is, precisely
166
+
167
+ - The checks are run by the runtime, and their exit codes are read by the
168
+ runtime. No part of the verdict is an agent's report about itself.
169
+ - A run carrying a contract cannot reach `completed` unless every required check
170
+ passed.
171
+ - If anything that can touch the workspace runs after the checks and the tree's
172
+ **content** changes, the verdict expires: the outcome drops to `unverified`
173
+ with the reason recorded. Not `refuted` — no check failed.
174
+ - Outside a git repository the digest is unavailable. Nothing running after the
175
+ checks means the verdict stands; something running after it means we decline
176
+ to vouch, and the run says so.
177
+
178
+ - If an abandoned attempt (a node past its timeout, which cannot be killed) is
179
+ still running when the run ends, the outcome is `unverified` with the reason
180
+ recorded. The runtime will not vouch for a tree something may still be writing
181
+ to.
182
+
183
+ What it is **not**: a promise that nothing can touch the tree after a run ends.
184
+ A node past its timeout keeps running, and if it outlives the short settle
185
+ window its writes land after the last digest — the run will have said
186
+ `unverified`, but the file is still changed. Give such nodes a timeout they will
187
+ not hit, or make their writes safe to arrive late.