@enderfga/claw-orchestrator 5.1.0 → 6.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +26 -26
- package/dist/bin/cli.js +107 -1
- package/dist/bin/cli.js.map +1 -1
- package/dist/src/acp-server.d.ts +5 -5
- package/dist/src/acp-server.js +3 -3
- package/dist/src/acp-server.js.map +1 -1
- package/dist/src/autoloop/dispatcher.d.ts +22 -0
- package/dist/src/autoloop/dispatcher.js +71 -13
- package/dist/src/autoloop/dispatcher.js.map +1 -1
- package/dist/src/autoloop/messages.d.ts +10 -0
- package/dist/src/autoloop/messages.js.map +1 -1
- package/dist/src/autoloop/runner.js +6 -0
- package/dist/src/autoloop/runner.js.map +1 -1
- package/dist/src/constants.d.ts +0 -6
- package/dist/src/constants.js +0 -6
- package/dist/src/constants.js.map +1 -1
- package/dist/src/council.d.ts +15 -0
- package/dist/src/council.js +48 -35
- package/dist/src/council.js.map +1 -1
- package/dist/src/dashboard/index.html +191 -6
- package/dist/src/embedded-server.js +132 -9
- package/dist/src/embedded-server.js.map +1 -1
- package/dist/src/fanout.d.ts +30 -1
- package/dist/src/fanout.js +32 -3
- package/dist/src/fanout.js.map +1 -1
- package/dist/src/index.js +359 -4
- package/dist/src/index.js.map +1 -1
- package/dist/src/kernel/agent-step.d.ts +59 -0
- package/dist/src/kernel/agent-step.js +100 -0
- package/dist/src/kernel/agent-step.js.map +1 -0
- package/dist/src/kernel/conditions.d.ts +11 -0
- package/dist/src/kernel/conditions.js +24 -0
- package/dist/src/kernel/conditions.js.map +1 -0
- package/dist/src/kernel/engine.d.ts +319 -0
- package/dist/src/kernel/engine.js +1047 -0
- package/dist/src/kernel/engine.js.map +1 -0
- package/dist/src/kernel/exec.d.ts +43 -0
- package/dist/src/kernel/exec.js +112 -0
- package/dist/src/kernel/exec.js.map +1 -0
- package/dist/src/kernel/file-lock.d.ts +50 -0
- package/dist/src/kernel/file-lock.js +135 -0
- package/dist/src/kernel/file-lock.js.map +1 -0
- package/dist/src/kernel/nodes/agent.d.ts +4 -0
- package/dist/src/kernel/nodes/agent.js +35 -0
- package/dist/src/kernel/nodes/agent.js.map +1 -0
- package/dist/src/kernel/nodes/autoloop.d.ts +78 -0
- package/dist/src/kernel/nodes/autoloop.js +75 -0
- package/dist/src/kernel/nodes/autoloop.js.map +1 -0
- package/dist/src/kernel/nodes/council.d.ts +12 -0
- package/dist/src/kernel/nodes/council.js +88 -0
- package/dist/src/kernel/nodes/council.js.map +1 -0
- package/dist/src/kernel/nodes/fanout.d.ts +11 -0
- package/dist/src/kernel/nodes/fanout.js +63 -0
- package/dist/src/kernel/nodes/fanout.js.map +1 -0
- package/dist/src/kernel/nodes/human-gate.d.ts +4 -0
- package/dist/src/kernel/nodes/human-gate.js +7 -0
- package/dist/src/kernel/nodes/human-gate.js.map +1 -0
- package/dist/src/kernel/nodes/index.d.ts +12 -0
- package/dist/src/kernel/nodes/index.js +21 -0
- package/dist/src/kernel/nodes/index.js.map +1 -0
- package/dist/src/kernel/nodes/router.d.ts +4 -0
- package/dist/src/kernel/nodes/router.js +12 -0
- package/dist/src/kernel/nodes/router.js.map +1 -0
- package/dist/src/kernel/nodes/subflow.d.ts +13 -0
- package/dist/src/kernel/nodes/subflow.js +38 -0
- package/dist/src/kernel/nodes/subflow.js.map +1 -0
- package/dist/src/kernel/nodes/ultraapp.d.ts +60 -0
- package/dist/src/kernel/nodes/ultraapp.js +62 -0
- package/dist/src/kernel/nodes/ultraapp.js.map +1 -0
- package/dist/src/kernel/nodes/verifier.d.ts +14 -0
- package/dist/src/kernel/nodes/verifier.js +84 -0
- package/dist/src/kernel/nodes/verifier.js.map +1 -0
- package/dist/src/kernel/projections.d.ts +42 -0
- package/dist/src/kernel/projections.js +133 -0
- package/dist/src/kernel/projections.js.map +1 -0
- package/dist/src/kernel/repo.d.ts +13 -0
- package/dist/src/kernel/repo.js +64 -0
- package/dist/src/kernel/repo.js.map +1 -0
- package/dist/src/kernel/secrets.d.ts +25 -0
- package/dist/src/kernel/secrets.js +48 -0
- package/dist/src/kernel/secrets.js.map +1 -0
- package/dist/src/kernel/store.d.ts +225 -0
- package/dist/src/kernel/store.js +838 -0
- package/dist/src/kernel/store.js.map +1 -0
- package/dist/src/kernel/templates/index.d.ts +140 -0
- package/dist/src/kernel/templates/index.js +266 -0
- package/dist/src/kernel/templates/index.js.map +1 -0
- package/dist/src/kernel/types.d.ts +326 -0
- package/dist/src/kernel/types.js +19 -0
- package/dist/src/kernel/types.js.map +1 -0
- package/dist/src/models.d.ts +7 -0
- package/dist/src/models.js +43 -14
- package/dist/src/models.js.map +1 -1
- package/dist/src/persistent-custom-session.js +8 -3
- package/dist/src/persistent-custom-session.js.map +1 -1
- package/dist/src/run-ledger.d.ts +57 -3
- package/dist/src/run-ledger.js +45 -2
- package/dist/src/run-ledger.js.map +1 -1
- package/dist/src/session-manager.d.ts +176 -129
- package/dist/src/session-manager.js +652 -603
- package/dist/src/session-manager.js.map +1 -1
- package/dist/src/types.d.ts +33 -3
- package/dist/src/ultraapp/build.d.ts +117 -3
- package/dist/src/ultraapp/build.js +319 -3
- package/dist/src/ultraapp/build.js.map +1 -1
- package/dist/src/ultraapp/contract.d.ts +52 -0
- package/dist/src/ultraapp/contract.js +83 -0
- package/dist/src/ultraapp/contract.js.map +1 -0
- package/dist/src/ultraapp/conventions.js +9 -2
- package/dist/src/ultraapp/conventions.js.map +1 -1
- package/dist/src/ultraapp/fix-on-failure.d.ts +21 -2
- package/dist/src/ultraapp/fix-on-failure.js +46 -62
- package/dist/src/ultraapp/fix-on-failure.js.map +1 -1
- package/dist/src/ultraapp/manager.d.ts +107 -2
- package/dist/src/ultraapp/manager.js +305 -86
- package/dist/src/ultraapp/manager.js.map +1 -1
- package/dist/src/verify/baseline.d.ts +73 -0
- package/dist/src/verify/baseline.js +186 -0
- package/dist/src/verify/baseline.js.map +1 -0
- package/dist/src/verify/contract.d.ts +116 -0
- package/dist/src/verify/contract.js +142 -0
- package/dist/src/verify/contract.js.map +1 -0
- package/dist/src/verify/evidence.d.ts +61 -0
- package/dist/src/verify/evidence.js +133 -0
- package/dist/src/verify/evidence.js.map +1 -0
- package/dist/src/verify/runner.d.ts +63 -0
- package/dist/src/verify/runner.js +317 -0
- package/dist/src/verify/runner.js.map +1 -0
- package/openclaw.plugin.json +38 -1
- package/package.json +2 -2
- package/skills/SKILL.md +120 -79
- package/skills/references/acp.md +17 -17
- package/skills/references/autoloop.md +139 -65
- package/skills/references/claude-cli-tracking.md +4 -4
- package/skills/references/cli.md +101 -59
- package/skills/references/council.md +109 -37
- package/skills/references/dashboard.md +34 -6
- package/skills/references/getting-started.md +13 -13
- package/skills/references/inbox.md +4 -4
- package/skills/references/mcp.md +39 -34
- package/skills/references/multi-engine.md +51 -47
- package/skills/references/observability.md +115 -28
- package/skills/references/openai-compat.md +39 -39
- package/skills/references/sessions.md +43 -25
- package/skills/references/tools.md +402 -309
- package/skills/references/ultra.md +45 -45
- package/skills/references/ultraapp.md +126 -50
- package/skills/references/verification.md +187 -0
- package/skills/references/workflow.md +362 -0
- package/dist/src/ultraapp/fix-on-failure-session.d.ts +0 -23
- package/dist/src/ultraapp/fix-on-failure-session.js +0 -51
- package/dist/src/ultraapp/fix-on-failure-session.js.map +0 -1
|
@@ -20,8 +20,8 @@ Runs in background — poll with `ultraplan_status`.
|
|
|
20
20
|
```typescript
|
|
21
21
|
const plan = manager.ultraplanStart('Add OAuth2 support with Google and GitHub providers', {
|
|
22
22
|
cwd: '/path/to/project',
|
|
23
|
-
model: 'opus',
|
|
24
|
-
timeout: 1800000,
|
|
23
|
+
model: 'opus', // default
|
|
24
|
+
timeout: 1800000, // 30 min default
|
|
25
25
|
});
|
|
26
26
|
|
|
27
27
|
console.log(`Plan ID: ${plan.id}`);
|
|
@@ -35,18 +35,18 @@ if (status?.status === 'completed') {
|
|
|
35
35
|
|
|
36
36
|
### Tools
|
|
37
37
|
|
|
38
|
-
| Tool
|
|
39
|
-
|
|
40
|
-
| `ultraplan_start`
|
|
41
|
-
| `ultraplan_status` | Get status and plan text
|
|
38
|
+
| Tool | Description |
|
|
39
|
+
| ------------------ | ----------------------------------- |
|
|
40
|
+
| `ultraplan_start` | Start planning session (background) |
|
|
41
|
+
| `ultraplan_status` | Get status and plan text |
|
|
42
42
|
|
|
43
43
|
### Configuration
|
|
44
44
|
|
|
45
|
-
| Parameter | Default
|
|
46
|
-
|
|
47
|
-
| `model`
|
|
48
|
-
| `cwd`
|
|
49
|
-
| `timeout` | 1,800,000 ms (30 min) | Maximum planning time
|
|
45
|
+
| Parameter | Default | Description |
|
|
46
|
+
| --------- | --------------------- | ---------------------------- |
|
|
47
|
+
| `model` | `opus` | Model to use for planning |
|
|
48
|
+
| `cwd` | `process.cwd()` | Project directory to explore |
|
|
49
|
+
| `timeout` | 1,800,000 ms (30 min) | Maximum planning time |
|
|
50
50
|
|
|
51
51
|
Results remain queryable for 30 minutes after completion.
|
|
52
52
|
|
|
@@ -65,36 +65,36 @@ A fleet of specialized bug-hunting agents that review your codebase in parallel,
|
|
|
65
65
|
|
|
66
66
|
### Available Review Angles (20)
|
|
67
67
|
|
|
68
|
-
| Agent
|
|
69
|
-
|
|
70
|
-
| SecurityReviewer
|
|
71
|
-
| LogicReviewer
|
|
68
|
+
| Agent | Focus |
|
|
69
|
+
| ------------------- | -------------------------------------------------------- |
|
|
70
|
+
| SecurityReviewer | Injection, auth flaws, data exposure, OWASP top 10 |
|
|
71
|
+
| LogicReviewer | Off-by-one, race conditions, null handling, edge cases |
|
|
72
72
|
| PerformanceReviewer | O(n^2) loops, memory leaks, missing caching, N+1 queries |
|
|
73
|
-
| APIReviewer
|
|
74
|
-
| TestReviewer
|
|
75
|
-
| TypeReviewer
|
|
76
|
-
| ConcurrencyReviewer | Race conditions, deadlocks, async error handling
|
|
77
|
-
| ErrorReviewer
|
|
78
|
-
| DependencyReviewer
|
|
79
|
-
| ReadabilityReviewer | Unclear naming, complex functions, dead code
|
|
80
|
-
| DataReviewer
|
|
81
|
-
| ConfigReviewer
|
|
82
|
-
| ScalabilityReviewer | Single points of failure, unbounded growth
|
|
83
|
-
| DocReviewer
|
|
84
|
-
| A11yReviewer
|
|
85
|
-
| I18nReviewer
|
|
86
|
-
| NetworkReviewer
|
|
87
|
-
| AuthReviewer
|
|
88
|
-
| CryptoReviewer
|
|
89
|
-
| MemoryReviewer
|
|
73
|
+
| APIReviewer | Inconsistent interfaces, missing validation, error gaps |
|
|
74
|
+
| TestReviewer | Untested paths, missing edge case tests, flaky patterns |
|
|
75
|
+
| TypeReviewer | `any` casts, unsafe assertions, missing null checks |
|
|
76
|
+
| ConcurrencyReviewer | Race conditions, deadlocks, async error handling |
|
|
77
|
+
| ErrorReviewer | Swallowed errors, missing try/catch, crash paths |
|
|
78
|
+
| DependencyReviewer | Outdated packages, CVEs, unnecessary deps |
|
|
79
|
+
| ReadabilityReviewer | Unclear naming, complex functions, dead code |
|
|
80
|
+
| DataReviewer | Data validation, schema mismatches, encoding |
|
|
81
|
+
| ConfigReviewer | Hardcoded values, missing env vars, insecure defaults |
|
|
82
|
+
| ScalabilityReviewer | Single points of failure, unbounded growth |
|
|
83
|
+
| DocReviewer | Outdated docs, missing API docs, misleading comments |
|
|
84
|
+
| A11yReviewer | ARIA labels, keyboard nav, color contrast |
|
|
85
|
+
| I18nReviewer | Hardcoded strings, locale handling, RTL support |
|
|
86
|
+
| NetworkReviewer | Missing timeouts, retry logic, connection pooling |
|
|
87
|
+
| AuthReviewer | Token handling, CSRF, permission checks |
|
|
88
|
+
| CryptoReviewer | Weak algorithms, key management, RNG |
|
|
89
|
+
| MemoryReviewer | Memory leaks, circular references, stream handling |
|
|
90
90
|
|
|
91
91
|
### Usage
|
|
92
92
|
|
|
93
93
|
```typescript
|
|
94
94
|
const review = manager.ultrareviewStart('/path/to/project', {
|
|
95
|
-
agentCount: 10,
|
|
96
|
-
maxDurationMinutes: 15,
|
|
97
|
-
model: 'sonnet',
|
|
95
|
+
agentCount: 10, // use 10 of the 20 angles
|
|
96
|
+
maxDurationMinutes: 15, // 15 min timeout per agent
|
|
97
|
+
model: 'sonnet', // model for all reviewers
|
|
98
98
|
focus: 'Find security and performance bugs',
|
|
99
99
|
});
|
|
100
100
|
|
|
@@ -109,18 +109,18 @@ if (status?.status === 'completed') {
|
|
|
109
109
|
|
|
110
110
|
### Tools
|
|
111
111
|
|
|
112
|
-
| Tool
|
|
113
|
-
|
|
114
|
-
| `ultrareview_start`
|
|
115
|
-
| `ultrareview_status` | Get status and findings
|
|
112
|
+
| Tool | Description |
|
|
113
|
+
| -------------------- | ---------------------------------- |
|
|
114
|
+
| `ultrareview_start` | Launch reviewer fleet (background) |
|
|
115
|
+
| `ultrareview_status` | Get status and findings |
|
|
116
116
|
|
|
117
117
|
### Configuration
|
|
118
118
|
|
|
119
|
-
| Parameter
|
|
120
|
-
|
|
121
|
-
| `agentCount`
|
|
122
|
-
| `maxDurationMinutes` | 10
|
|
123
|
-
| `model`
|
|
124
|
-
| `focus`
|
|
119
|
+
| Parameter | Default | Range | Description |
|
|
120
|
+
| -------------------- | ------------------------- | ----- | ------------------------- |
|
|
121
|
+
| `agentCount` | 5 | 1-20 | Number of reviewer agents |
|
|
122
|
+
| `maxDurationMinutes` | 10 | 5-25 | Per-agent timeout |
|
|
123
|
+
| `model` | session default | — | Model for all reviewers |
|
|
124
|
+
| `focus` | bugs + security + quality | — | Review focus description |
|
|
125
125
|
|
|
126
126
|
The council runs with `maxRounds: 2` — one round to find bugs, one to cross-review. Results remain queryable for 30 minutes.
|
|
@@ -32,30 +32,58 @@ interview ─► queued ─► building ─► build-complete ─► deploying
|
|
|
32
32
|
structural)
|
|
33
33
|
```
|
|
34
34
|
|
|
35
|
-
| Mode
|
|
36
|
-
|
|
37
|
-
| `interview`
|
|
38
|
-
| `queued`
|
|
39
|
-
| `building`
|
|
40
|
-
| `build-complete` | Codebase ready, awaiting `deploy` step.
|
|
41
|
-
| `deploying`
|
|
42
|
-
| `done`
|
|
43
|
-
| `failed`
|
|
35
|
+
| Mode | Meaning |
|
|
36
|
+
| ---------------- | ------------------------------------------------------------------------------- |
|
|
37
|
+
| `interview` | AppSpec being filled by Q&A. Chat input goes to the interview Opus. |
|
|
38
|
+
| `queued` | Build accepted, waiting for a slot in the FIFO build queue. |
|
|
39
|
+
| `building` | Council writing code, fix-on-failure driving install/build/test. |
|
|
40
|
+
| `build-complete` | Codebase ready, awaiting `deploy` step. |
|
|
41
|
+
| `deploying` | Container/process being started, router map being updated. |
|
|
42
|
+
| `done` | App live at `/forge/<slug>/`. Chat input now goes to the done-mode classifier. |
|
|
43
|
+
| `failed` | Council didn't reach consensus, or fix-on-failure couldn't get the build green. |
|
|
44
|
+
|
|
45
|
+
### The build is a workflow run
|
|
46
|
+
|
|
47
|
+
Everything from `building` onward is a run on the durable kernel — the same one
|
|
48
|
+
`council_start`, `fanout_start` and `autoloop_start` use. The workflow has three
|
|
49
|
+
nodes:
|
|
50
|
+
|
|
51
|
+
| Node | Kind | Does |
|
|
52
|
+
| -------- | ----------------- | ------------------------------------------------------------ |
|
|
53
|
+
| `synth` | `ultraapp_synth` | The 3-agent council; snapshots consensus into `versions/v1/` |
|
|
54
|
+
| `build` | `verifier` | The build contract, with its fix-on-red loop and evidence |
|
|
55
|
+
| `deploy` | `ultraapp_deploy` | Starts the app, registers the route, runs the §7g gate |
|
|
56
|
+
|
|
57
|
+
Three consequences worth knowing:
|
|
58
|
+
|
|
59
|
+
- The run is listed by `workflow_list` and `clawo workflow list`, and shows up in
|
|
60
|
+
the dashboard's Runs tab, under the id `ultraapp-<run_id>`.
|
|
61
|
+
- **A crash between stages does not redo the stage before it.** A build that died
|
|
62
|
+
after the council resumes into the build stage rather than paying for the
|
|
63
|
+
council twice.
|
|
64
|
+
- The mode column above is a **projection** of that run's state, not a second
|
|
65
|
+
state machine. `interview` and `queued` are the exception — nothing is running
|
|
66
|
+
yet, so there is no run to project from.
|
|
67
|
+
|
|
68
|
+
The interview and the done-mode conversation are deliberately not workflows.
|
|
69
|
+
They are user-driven and open-ended, with no predetermined sequence of steps;
|
|
70
|
+
a workflow expressing them would be a router self-loop or a single node with a
|
|
71
|
+
state machine hidden inside it.
|
|
44
72
|
|
|
45
73
|
## Architectural conventions (§1–§7)
|
|
46
74
|
|
|
47
75
|
Every generated app MUST satisfy these. They're embedded in the council
|
|
48
76
|
super-task prompt verbatim from `src/ultraapp/conventions.ts`.
|
|
49
77
|
|
|
50
|
-
| §
|
|
51
|
-
|
|
52
|
-
| 1
|
|
53
|
-
| 2
|
|
54
|
-
| 3
|
|
55
|
-
| 4
|
|
56
|
-
| 5
|
|
57
|
-
| 6
|
|
58
|
-
| 7
|
|
78
|
+
| § | Topic | Headline rule |
|
|
79
|
+
| --- | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
80
|
+
| 1 | Path-based deploy | Mount at `BASE_PATH=/forge/<slug>/`; in-app links MUST be relative. |
|
|
81
|
+
| 2 | Async file-queue runtime | Exact endpoints: `GET /`, `POST /run`, `GET /status/:jobId`, `GET /result/:jobId`, `GET /health`. File-based job queue under `$DATA_DIR/jobs/<jobId>/`. NO database. Data path from `process.env.DATA_DIR ?? '/data'`. |
|
|
82
|
+
| 3 | BYOK | If `runtime.needsLLM`, API keys live in browser localStorage and are sent direct to the provider. The server MUST NEVER receive the key (enforced by `eslint-plugin-no-server-keys`). |
|
|
83
|
+
| 4 | Dockerfile + smoke test | Single multi-stage Dockerfile, `npm run smoke` drives one full job in < 90s using `examples[0].ref`. |
|
|
84
|
+
| 5 | Council voting protocol | 3 agents in git worktrees, all-YES vote required, max 8 rounds. |
|
|
85
|
+
| 6 | Tech stack | Modern TypeScript / JavaScript framework (Next.js, Vite + Hono, SvelteKit). NO Python, NO pure SSGs. |
|
|
86
|
+
| 7 | **Frontend quality** | **Real styling system + real type hierarchy + four-state coverage on every async surface + drag-and-drop forms + appropriate result presentation + one deliberate theme.** §7g requires every agent to capture Chrome-headless screenshots at 1440×900 AND 375×812 and visually inspect the PNGs before voting YES — source-code review is explicitly insufficient evidence. |
|
|
59
87
|
|
|
60
88
|
## Runtime modes
|
|
61
89
|
|
|
@@ -64,10 +92,10 @@ clawo serve --ultraapp-runtime host # default
|
|
|
64
92
|
clawo serve --ultraapp-runtime docker # opt-in
|
|
65
93
|
```
|
|
66
94
|
|
|
67
|
-
| Mode
|
|
68
|
-
|
|
69
|
-
| `host`
|
|
70
|
-
| `docker` | `docker build .`
|
|
95
|
+
| Mode | Build | Run | Pros | Cons |
|
|
96
|
+
| -------- | ------------------------------ | ------------------------------------------- | --------------------------------------------------------- | ---------------------------------------------------- |
|
|
97
|
+
| `host` | `npm install && npm run build` | `npm start` (detached, `setsid`-equivalent) | Zero extra deps; works anywhere Node works; faster start. | No process isolation; deps installed under the user. |
|
|
98
|
+
| `docker` | `docker build .` | `docker run -d --restart unless-stopped` | Per-app isolation, restart policy, image is the artefact. | Requires a running Docker daemon. |
|
|
71
99
|
|
|
72
100
|
Both modes allocate a backend port in `[19100, 19999]`. The reverse-
|
|
73
101
|
proxy router runs at port `19000` (auto-fallback up to `19099` if
|
|
@@ -100,23 +128,23 @@ persists to `~/.claw-orchestrator/host-procs.json`.
|
|
|
100
128
|
All routes are served by the embedded server (default `:18796`), under
|
|
101
129
|
`Authorization: Bearer <token>` from `~/.openclaw/server-token`.
|
|
102
130
|
|
|
103
|
-
| Method + path
|
|
104
|
-
|
|
105
|
-
| `GET /ultraapp/list`
|
|
106
|
-
| `POST /ultraapp/new`
|
|
107
|
-
| `GET /ultraapp/<id>`
|
|
108
|
-
| `POST /ultraapp/<id>/answer`
|
|
109
|
-
| `POST /ultraapp/<id>/spec-edit`
|
|
110
|
-
| `POST /ultraapp/<id>/files`
|
|
111
|
-
| `GET /ultraapp/<id>/events`
|
|
112
|
-
| `POST /ultraapp/<id>/build`
|
|
113
|
-
| `POST /ultraapp/<id>/build/cancel`
|
|
114
|
-
| `GET /ultraapp/<id>/artifacts`
|
|
115
|
-
| `POST /ultraapp/<id>/start`
|
|
116
|
-
| `POST /ultraapp/<id>/stop`
|
|
117
|
-
| `POST /ultraapp/<id>/delete`
|
|
118
|
-
| `POST /ultraapp/<id>/feedback`
|
|
119
|
-
| `POST /ultraapp/<id>/promote-version` | Body: `{ version: "vN" }`. Atomically swap deployed version.
|
|
131
|
+
| Method + path | Purpose |
|
|
132
|
+
| ------------------------------------- | --------------------------------------------------------------------------------- |
|
|
133
|
+
| `GET /ultraapp/list` | All runs with mode + createdAt. |
|
|
134
|
+
| `POST /ultraapp/new` | Body: `{ firstMessage?: string }`. Returns `{ runId }`. |
|
|
135
|
+
| `GET /ultraapp/<id>` | Full snapshot: spec + chat + state. |
|
|
136
|
+
| `POST /ultraapp/<id>/answer` | Body: `{ value, freeform? }`. Submit interview answer. |
|
|
137
|
+
| `POST /ultraapp/<id>/spec-edit` | Body: RFC 6902 patch ops. Edit the spec mid-interview. |
|
|
138
|
+
| `POST /ultraapp/<id>/files` | Multipart upload to `examples/`. |
|
|
139
|
+
| `GET /ultraapp/<id>/events` | SSE stream of build/chat events (mode pill, narrator, council activity). |
|
|
140
|
+
| `POST /ultraapp/<id>/build` | Validate spec strictly + enqueue. |
|
|
141
|
+
| `POST /ultraapp/<id>/build/cancel` | Abort the active build. |
|
|
142
|
+
| `GET /ultraapp/<id>/artifacts` | List `versions/vN/`. |
|
|
143
|
+
| `POST /ultraapp/<id>/start` | Start the deployed container/process for the active version. |
|
|
144
|
+
| `POST /ultraapp/<id>/stop` | Stop without deleting. |
|
|
145
|
+
| `POST /ultraapp/<id>/delete` | Stop + remove all per-run state. |
|
|
146
|
+
| `POST /ultraapp/<id>/feedback` | Body: `{ text }`. Done-mode classifier routes cosmetic / spec-delta / structural. |
|
|
147
|
+
| `POST /ultraapp/<id>/promote-version` | Body: `{ version: "vN" }`. Atomically swap deployed version. |
|
|
120
148
|
|
|
121
149
|
## MCP tools (14)
|
|
122
150
|
|
|
@@ -137,11 +165,11 @@ ultraapp_start_container ultraapp_stop_container ultraapp_delete
|
|
|
137
165
|
After the run reaches `done`, chat input goes to a per-run Haiku
|
|
138
166
|
classifier. Three classes:
|
|
139
167
|
|
|
140
|
-
| Class
|
|
141
|
-
|
|
142
|
-
| `cosmetic`
|
|
143
|
-
| `spec-delta` | Focused interview | Flips mode back to `interview` with a bootstrap message that names the field(s) being changed. Completion auto-triggers a fresh `startBuild`.
|
|
144
|
-
| `structural` | Suggestion only
|
|
168
|
+
| Class | Routes to | Behaviour |
|
|
169
|
+
| ------------ | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
170
|
+
| `cosmetic` | Patcher | Opus generates a unified diff against the deployed worktree → `applyUnifiedDiff` → validate via fix-on-failure → on success snapshot to `versions/vN+1/`, on any failure restore the snapshot atomically and post the reason to chat. |
|
|
171
|
+
| `spec-delta` | Focused interview | Flips mode back to `interview` with a bootstrap message that names the field(s) being changed. Completion auto-triggers a fresh `startBuild`. |
|
|
172
|
+
| `structural` | Suggestion only | Posts a narrator note: "this sounds like a different app — click + New". |
|
|
145
173
|
|
|
146
174
|
To swap which version is live, use `promote-version` (HTTP) or
|
|
147
175
|
`ultraapp_promote_version` (MCP) — the router map and host-procs map
|
|
@@ -190,14 +218,62 @@ open "http://127.0.0.1:18796/dashboard?token=$(cat ~/.openclaw/server-token)"
|
|
|
190
218
|
curl http://127.0.0.1:19000/forge/<slug>/health
|
|
191
219
|
```
|
|
192
220
|
|
|
193
|
-
##
|
|
221
|
+
## Acceptance contract (6.0.0)
|
|
222
|
+
|
|
223
|
+
UltraApp is the one mode whose contract is on by default, because it already ran
|
|
224
|
+
most of these commands and because two of its documented gates were not actually
|
|
225
|
+
enforced.
|
|
226
|
+
|
|
227
|
+
**Build stage**, in the council's worktree:
|
|
228
|
+
|
|
229
|
+
```
|
|
230
|
+
npm install → npm run build → npm test → [docker build] → npm run smoke
|
|
231
|
+
```
|
|
232
|
+
|
|
233
|
+
`npm run smoke` is new here. §4 of the architectural conventions has always told
|
|
234
|
+
the council that the smoke test gates build success — it was never in the step
|
|
235
|
+
list, so the claim was false. A codebase without a working `scripts.smoke` now
|
|
236
|
+
fails its build, which is what the brief said all along.
|
|
237
|
+
|
|
238
|
+
**Deploy stage**, against the live URL: both §7g viewports (1440×900 and
|
|
239
|
+
375×812) are captured by the orchestrator with headless Chrome and stored as run
|
|
240
|
+
evidence.
|
|
241
|
+
|
|
242
|
+
Evidence lands under the run directory:
|
|
243
|
+
|
|
244
|
+
```
|
|
245
|
+
~/.claw-orchestrator/ultraapps/<runId>/evidence/
|
|
246
|
+
build-01/ bundle.json, checks/*.log, diff.patch
|
|
247
|
+
deploy-01/ bundle.json, shot-desktop.png, shot-mobile.png
|
|
248
|
+
```
|
|
249
|
+
|
|
250
|
+
### What the visual gate does and does not do
|
|
251
|
+
|
|
252
|
+
It captures images and stores them, so "did anyone actually look" is now a file
|
|
253
|
+
on disk instead of an agent's claim. It does **not** compare pixels — judging
|
|
254
|
+
the rendering is still a reader's job, and the §7g instructions in the council
|
|
255
|
+
prompt remain the agents' responsibility.
|
|
256
|
+
|
|
257
|
+
It is **advisory by default**, so a host without Chrome does not lose a working
|
|
258
|
+
app to a missing browser. Set `CLAWO_ULTRAAPP_VISUAL_GATE=strict` to make a
|
|
259
|
+
failed capture block the deploy. Chrome is resolved from `CLAWO_CHROME_BIN`, then
|
|
260
|
+
the usual macOS app paths, then `PATH`.
|
|
261
|
+
|
|
262
|
+
## Durable build queue (6.0.0)
|
|
263
|
+
|
|
264
|
+
The build queue is persisted to `<store>/build-queue.json` and restored on
|
|
265
|
+
startup. Its own comment used to say a restart mid-build meant "the build is
|
|
266
|
+
marked failed and the user can rerun"; in practice nothing was marked — the
|
|
267
|
+
pending list vanished along with any queued build the user was waiting on, with
|
|
268
|
+
no record it had been asked for.
|
|
269
|
+
|
|
270
|
+
A build that was in flight when the process died is **re-queued, not resumed**,
|
|
271
|
+
and goes to the front: each build starts from a fresh council worktree, so
|
|
272
|
+
re-running is safe and continuing a half-built tree is not.
|
|
273
|
+
|
|
274
|
+
## Known limitations
|
|
194
275
|
|
|
195
276
|
- The done-mode patcher loop occasionally hangs between
|
|
196
277
|
feedback-classification and the patcher Opus session creation;
|
|
197
278
|
cosmetic changes can be applied manually until the underlying race
|
|
198
279
|
is fixed.
|
|
199
|
-
- The §7g frontend gate currently relies on per-agent honesty about
|
|
200
|
-
running the screenshot capture; agents that skip the inspection can
|
|
201
|
-
still pass the smoke gate. A follow-up will plumb a server-side
|
|
202
|
-
screenshot validator into the council verifier so the gate becomes
|
|
203
|
-
structurally enforced rather than persona-enforced.
|
|
@@ -0,0 +1,187 @@
|
|
|
1
|
+
# Verification — acceptance contracts and evidence
|
|
2
|
+
|
|
3
|
+
Before 6.0.0, nothing in this runtime ever checked an agent's work. Every
|
|
4
|
+
"finished" signal was the agent grading itself:
|
|
5
|
+
|
|
6
|
+
- **Council** terminated when a regex found `[CONSENSUS: YES]` in agent prose.
|
|
7
|
+
- **Autoloop**'s `eval_output` was whatever the Coder passed to a tool call, and
|
|
8
|
+
the Reviewer that was supposed to catch fabrication had a sandbox containing
|
|
9
|
+
the iteration's artifacts and no code, so it could not re-derive anything.
|
|
10
|
+
- **UltraApp**'s frontend gate was a sentence in a persona string telling agents
|
|
11
|
+
to capture screenshots. There was no screenshot code anywhere in the project.
|
|
12
|
+
- The run ledger's `ok` was the engine's report on its own turn.
|
|
13
|
+
|
|
14
|
+
An **acceptance contract** is the opposite of all of that: a list of checks the
|
|
15
|
+
runtime executes and whose results it reads. A run that declares one cannot reach
|
|
16
|
+
`completed` unless every required check passes.
|
|
17
|
+
|
|
18
|
+
## The one rule about where contracts come from
|
|
19
|
+
|
|
20
|
+
**A contract comes from the caller or from a mode default. Never from agent
|
|
21
|
+
output.** If an agent could declare its own checks, we would be back to
|
|
22
|
+
self-grading with more steps. Nothing in the kernel reads a contract out of a
|
|
23
|
+
node's result, and `normalizeContract()` drops anything it does not recognise, so
|
|
24
|
+
a contract that arrived through a tool call carries no fields the executor did
|
|
25
|
+
not model.
|
|
26
|
+
|
|
27
|
+
Concretely: there is no shell string anywhere. A `command` check is argv.
|
|
28
|
+
|
|
29
|
+
## Check types
|
|
30
|
+
|
|
31
|
+
| Type | What it does | Passes when |
|
|
32
|
+
| ------------- | --------------------------------------------------------------------------- | ------------------------------------------------------------------------- |
|
|
33
|
+
| `command` | Runs argv in a directory | Exit code equals `expectExit` (default 0) |
|
|
34
|
+
| `http` | Polls a URL until the deadline | Status equals `expectStatus` (default 200) |
|
|
35
|
+
| `screenshot` | Captures the page at each viewport with headless Chrome and stores the PNGs | Every viewport produced a non-empty image |
|
|
36
|
+
| `diff_policy` | Compares the change set against the recorded baseline | File count, forbidden paths, and required paths all satisfied |
|
|
37
|
+
| `file` | Checks a path | Exists (or is absent when `exists: false`) and matches `matches` if given |
|
|
38
|
+
|
|
39
|
+
```jsonc
|
|
40
|
+
{
|
|
41
|
+
"id": "ship-it",
|
|
42
|
+
"fixOnFailureRounds": 2,
|
|
43
|
+
"checks": [
|
|
44
|
+
{ "type": "command", "cmd": "npm", "args": ["run", "build"], "timeoutMs": 600000 },
|
|
45
|
+
{ "type": "command", "cmd": "npm", "args": ["test"] },
|
|
46
|
+
{ "type": "command", "cmd": "npm", "args": ["run", "lint"], "required": false },
|
|
47
|
+
{ "type": "diff_policy", "forbidPaths": ["configs", ".github"], "maxFiles": 40 },
|
|
48
|
+
],
|
|
49
|
+
}
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
### `required`
|
|
53
|
+
|
|
54
|
+
Default true. A failing non-required check is recorded in the evidence bundle but
|
|
55
|
+
does not refute the run — use it for signals you want visible without making them
|
|
56
|
+
blocking.
|
|
57
|
+
|
|
58
|
+
(The predecessor of this module, `ultraapp/fix-on-failure.ts`, declared a
|
|
59
|
+
`required` field on its steps and never read it: every step was fatal. It is
|
|
60
|
+
honoured now.)
|
|
61
|
+
|
|
62
|
+
### Timeouts
|
|
63
|
+
|
|
64
|
+
Every check has one, and the default is 10 minutes. This is not cosmetic — the
|
|
65
|
+
old pipeline had no timeout at all, so a wedged `npm test` hung the build
|
|
66
|
+
forever. A check that overruns is killed (its whole process group, with SIGKILL)
|
|
67
|
+
and recorded as failed with `timedOut: true`.
|
|
68
|
+
|
|
69
|
+
### What `screenshot` does and does not claim
|
|
70
|
+
|
|
71
|
+
It captures images and stores them. It does **not** compare pixels, and it is not
|
|
72
|
+
visual regression testing. What changed in 6.0.0 is that the capture is performed
|
|
73
|
+
by the runtime rather than requested of an agent — so "did anyone actually look"
|
|
74
|
+
stops being a claim and starts being a file on disk. Judging the rendering is
|
|
75
|
+
still a human's or an agent's job.
|
|
76
|
+
|
|
77
|
+
Chrome is resolved from `CLAWO_CHROME_BIN`, then the usual macOS app paths, then
|
|
78
|
+
`PATH`. A host with no browser fails the check immediately rather than paying the
|
|
79
|
+
timeout to find out.
|
|
80
|
+
|
|
81
|
+
### `diff_policy` and the baseline
|
|
82
|
+
|
|
83
|
+
`diff_policy` needs to know what the run changed, which needs a baseline. The
|
|
84
|
+
kernel records `git rev-parse HEAD` when a run starts and diffs against that.
|
|
85
|
+
|
|
86
|
+
The change set is **tracked changes ∪ untracked files**. That union matters: a
|
|
87
|
+
bare `git diff` lists tracked modifications only, so files an agent _created_ are
|
|
88
|
+
invisible to it — which is exactly how `autoloop`'s per-iteration `diff.patch`
|
|
89
|
+
used to miss every new file while `git add -A` committed them anyway.
|
|
90
|
+
|
|
91
|
+
`requirePaths: ["."]` means "the run must have changed something".
|
|
92
|
+
|
|
93
|
+
## Evidence bundles
|
|
94
|
+
|
|
95
|
+
Every verifier attempt writes a directory under the run:
|
|
96
|
+
|
|
97
|
+
```
|
|
98
|
+
~/.claw-orchestrator/wf/<runId>/evidence/<evidenceId>/
|
|
99
|
+
bundle.json verdict, per-check results, changed files, base/head SHA
|
|
100
|
+
checks/<id>.log output tail for each failed check
|
|
101
|
+
diff.patch the patch the run produced, new files included
|
|
102
|
+
shot-*.png screenshots, when a screenshot check ran
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
Read one with `clawo verify <runId>`, `GET /workflow/<id>/evidence`, or
|
|
106
|
+
`workflow_status` (which returns the id).
|
|
107
|
+
|
|
108
|
+
Bundle writes are best-effort. The verdict is already decided by the check
|
|
109
|
+
results, so a bundle that fails to land loses the record, never the answer.
|
|
110
|
+
|
|
111
|
+
## Fix-on-red
|
|
112
|
+
|
|
113
|
+
`fixOnFailureRounds` spawns a repair session against the failing check and then
|
|
114
|
+
**re-runs the whole check list**. The fixer's own claim to have fixed it is
|
|
115
|
+
ignored; only the re-run decides. Set it to 0 (the default) to disable.
|
|
116
|
+
|
|
117
|
+
## Per-mode defaults
|
|
118
|
+
|
|
119
|
+
| Mode | Contract | Notes |
|
|
120
|
+
| ------------------------ | ----------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
121
|
+
| **UltraApp** | **On by default** | Build: `npm install` → `npm run build` → `npm test` → (`docker build`) → **`npm run smoke`**. Deploy: **both §7g viewports captured** against the live URL. |
|
|
122
|
+
| **Council** | Caller-declared | Consensus votes are recorded on the run as advisory and no longer decide completion. |
|
|
123
|
+
| **Autoloop** | Caller-declared | With a contract, a Reviewer `advance` is held unless the checks pass; passing also fires `on_target_hit`. |
|
|
124
|
+
| **Fanout / Ultrareview** | Caller-declared | Per-agent `ok` now reads the engine's terminal verdict rather than "the call did not throw". |
|
|
125
|
+
| **Plain sessions** | None | Use `verify_run` to check work after the fact. |
|
|
126
|
+
|
|
127
|
+
Two of UltraApp's defaults close claims the project had been making without
|
|
128
|
+
backing them:
|
|
129
|
+
|
|
130
|
+
- `conventions.ts` §4 told every council that `npm run smoke` gated build
|
|
131
|
+
success. It was not in the step list at all. It is now.
|
|
132
|
+
- `ultraapp.md` recorded, as a known limitation since 4.0.0, that the §7g gate
|
|
133
|
+
"relies on per-agent honesty about running the screenshot capture", with a
|
|
134
|
+
server-side validator promised as a follow-up. That follow-up is this release.
|
|
135
|
+
|
|
136
|
+
The visual gate is **advisory by default** so that a host without Chrome does not
|
|
137
|
+
lose a working app to a missing browser. Set
|
|
138
|
+
`CLAWO_ULTRAAPP_VISUAL_GATE=strict` to make a failed capture block the deploy.
|
|
139
|
+
|
|
140
|
+
## Verifying work that did not come through a workflow
|
|
141
|
+
|
|
142
|
+
```bash
|
|
143
|
+
# Tool
|
|
144
|
+
verify_run({ cwd: "/repo", contract: { checks: [{ type: "command", cmd: "npm", args: ["test"] }] } })
|
|
145
|
+
```
|
|
146
|
+
|
|
147
|
+
Use it for a plain `session_send` that edited a repo, or for anything from an
|
|
148
|
+
older version. The contract is yours; nothing is read from agent output.
|
|
149
|
+
|
|
150
|
+
## Three outcomes, not two
|
|
151
|
+
|
|
152
|
+
A run ends `verified`, `refuted`, or `unverified`.
|
|
153
|
+
|
|
154
|
+
`unverified` means **no contract was declared and nothing checked the work**. It
|
|
155
|
+
is not a failure and it is not a pass. The CLI prints it as `—` and the summary
|
|
156
|
+
line says so in words, because collapsing it into either bucket would let an
|
|
157
|
+
unchecked run read as a checked one.
|
|
158
|
+
|
|
159
|
+
## Related
|
|
160
|
+
|
|
161
|
+
- [`workflow.md`](./workflow.md) — the kernel that runs verifiers as nodes
|
|
162
|
+
- [`observability.md`](./observability.md) — how verdicts reach the run ledger
|
|
163
|
+
- [`ultraapp.md`](./ultraapp.md) — the default contract in context
|
|
164
|
+
|
|
165
|
+
## What the guarantee is, precisely
|
|
166
|
+
|
|
167
|
+
- The checks are run by the runtime, and their exit codes are read by the
|
|
168
|
+
runtime. No part of the verdict is an agent's report about itself.
|
|
169
|
+
- A run carrying a contract cannot reach `completed` unless every required check
|
|
170
|
+
passed.
|
|
171
|
+
- If anything that can touch the workspace runs after the checks and the tree's
|
|
172
|
+
**content** changes, the verdict expires: the outcome drops to `unverified`
|
|
173
|
+
with the reason recorded. Not `refuted` — no check failed.
|
|
174
|
+
- Outside a git repository the digest is unavailable. Nothing running after the
|
|
175
|
+
checks means the verdict stands; something running after it means we decline
|
|
176
|
+
to vouch, and the run says so.
|
|
177
|
+
|
|
178
|
+
- If an abandoned attempt (a node past its timeout, which cannot be killed) is
|
|
179
|
+
still running when the run ends, the outcome is `unverified` with the reason
|
|
180
|
+
recorded. The runtime will not vouch for a tree something may still be writing
|
|
181
|
+
to.
|
|
182
|
+
|
|
183
|
+
What it is **not**: a promise that nothing can touch the tree after a run ends.
|
|
184
|
+
A node past its timeout keeps running, and if it outlives the short settle
|
|
185
|
+
window its writes land after the last digest — the run will have said
|
|
186
|
+
`unverified`, but the file is still changed. Give such nodes a timeout they will
|
|
187
|
+
not hit, or make their writes safe to arrive late.
|