@enderfga/claw-orchestrator 3.3.1 → 3.5.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. package/README.md +16 -1
  2. package/configs/autoloop-coder-prompt.md +59 -0
  3. package/configs/autoloop-planner-prompt.md +136 -0
  4. package/configs/autoloop-reviewer-prompt.md +76 -0
  5. package/dist/src/autoloop/agent-tools.d.ts +44 -0
  6. package/dist/src/autoloop/agent-tools.js +71 -0
  7. package/dist/src/autoloop/agent-tools.js.map +1 -0
  8. package/dist/src/autoloop/dispatcher.d.ts +137 -0
  9. package/dist/src/autoloop/dispatcher.js +580 -0
  10. package/dist/src/autoloop/dispatcher.js.map +1 -0
  11. package/dist/src/autoloop/messages.d.ts +114 -0
  12. package/dist/src/autoloop/messages.js +92 -0
  13. package/dist/src/autoloop/messages.js.map +1 -0
  14. package/dist/src/autoloop/notify.d.ts +36 -0
  15. package/dist/src/autoloop/notify.js +166 -0
  16. package/dist/src/autoloop/notify.js.map +1 -0
  17. package/dist/src/autoloop/planner-tools.d.ts +81 -0
  18. package/dist/src/autoloop/planner-tools.js +140 -0
  19. package/dist/src/autoloop/planner-tools.js.map +1 -0
  20. package/dist/src/autoloop/runner.d.ts +48 -0
  21. package/dist/src/autoloop/runner.js +214 -0
  22. package/dist/src/autoloop/runner.js.map +1 -0
  23. package/dist/src/autoloop/types.d.ts +69 -0
  24. package/dist/src/autoloop/types.js +16 -0
  25. package/dist/src/autoloop/types.js.map +1 -0
  26. package/dist/src/dashboard/index.html +910 -0
  27. package/dist/src/embedded-server.js +198 -0
  28. package/dist/src/embedded-server.js.map +1 -1
  29. package/dist/src/index.d.ts +5 -0
  30. package/dist/src/index.js +128 -0
  31. package/dist/src/index.js.map +1 -1
  32. package/dist/src/session-manager.d.ts +51 -0
  33. package/dist/src/session-manager.js +162 -0
  34. package/dist/src/session-manager.js.map +1 -1
  35. package/package.json +2 -2
  36. package/skills/SKILL.md +22 -1
  37. package/skills/references/autoloop.md +223 -0
@@ -0,0 +1,223 @@
1
+ # Autoloop — Reference
2
+
3
+ Three-agent autonomous iteration loop for a git workspace. You converse with
4
+ the **Planner** to design a plan; on your approval, the Planner spawns the
5
+ **Coder** + **Reviewer** subloop, monitors it, and pushes you (wechat →
6
+ whatsapp → email fallback chain) only when something needs your attention.
7
+
8
+ Design rationale: `tasks/autoloop.md`. This page is the operator reference.
9
+
10
+ ## When to use
11
+
12
+ - Goal is exploratory and you want to **design the plan with the agent**
13
+ before running. The Planner will read your workspace, surface ambiguity,
14
+ and write `plan.md` / `goal.json` with you.
15
+ - You want a long-running loop you can walk away from. The Planner pushes
16
+ you on regressions, completion, decisions, or stalls; otherwise it stays
17
+ silent.
18
+ - You can write down what "better" means as a shell command (test pass
19
+ count, latency, loss, gate completion). Without that anchor, Reviewer has
20
+ nothing to ratchet on.
21
+
22
+ ## Roles
23
+
24
+ | Agent | Engine (default) | cwd | Owns |
25
+ |---|---|---|---|
26
+ | **Planner** | claude / opus | workspace | strategy, `plan.md`, `goal.json`, talking to you |
27
+ | **Coder** | claude / sonnet (override per spawn) | workspace | code changes, eval execution |
28
+ | **Reviewer** | claude / sonnet | `<workspace>/tasks/<run_id>/reviewer_sandbox/` | distrust audit; advance / hold / rollback |
29
+
30
+ Coder and Reviewer **never speak to you directly**. Anything they observe
31
+ flows through the Planner. The Planner decides what to surface and what to
32
+ absorb.
33
+
34
+ ## UX flow
35
+
36
+ ```
37
+ 1. autoloop_start { run_id, workspace } → Planner session ready
38
+ 2. autoloop_chat { run_id, "<your goal>" } → Planner reads workspace,
39
+ drafts plan.md + goal.json,
40
+ asks "ready to spawn?"
41
+ 3. autoloop_chat { run_id, "go" } → Planner emits spawn_subagents
42
+ 4. Coder + Reviewer self-iterate → ledger writes per iter
43
+ 5. Planner pushes you on target_hit / regression / decision / stall
44
+ 6. Run terminates on target hit, plan-defined max_iters, or your terminate.
45
+ ```
46
+
47
+ ## Quick start
48
+
49
+ ```bash
50
+ # Start a run (creates Planner session)
51
+ curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_start \
52
+ -H 'content-type: application/json' \
53
+ -d '{"run_id":"my-run","workspace":"/abs/path/to/workspace"}'
54
+
55
+ # Chat with the Planner
56
+ curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_chat \
57
+ -H 'content-type: application/json' \
58
+ -d '{"run_id":"my-run","text":"Read the workspace and design a plan to fix X"}'
59
+
60
+ # Inspect state
61
+ curl http://127.0.0.1:18789/autoloop/my-run/state
62
+
63
+ # Live SSE stream (the 3-pane UI subscribes here)
64
+ curl http://127.0.0.1:18789/autoloop/my-run/events
65
+
66
+ # Reset Coder if it drifts (lazy; eager_restart=true to start a fresh session immediately)
67
+ curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_reset_agent \
68
+ -H 'content-type: application/json' \
69
+ -d '{"run_id":"my-run","agent":"coder","eager_restart":true}'
70
+
71
+ # Stop
72
+ curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_stop \
73
+ -H 'content-type: application/json' \
74
+ -d '{"run_id":"my-run","reason":"done"}'
75
+ ```
76
+
77
+ ## Plugin tools
78
+
79
+ | Tool | Args | What |
80
+ |---|---|---|
81
+ | `autoloop_start` | `run_id`, `workspace`, `planner_model?`, `send_timeout_ms?` | Start a run; launches Planner session. |
82
+ | `autoloop_chat` | `run_id`, `text` | Send a chat message to the Planner; returns the Planner's reply. |
83
+ | `autoloop_status` | `run_id` | Current state (status, iter, push count, subagents_spawned). |
84
+ | `autoloop_list` | — | All active runs in this manager process. |
85
+ | `autoloop_stop` | `run_id`, `reason?` | Terminate; stops Planner / Coder / Reviewer. |
86
+ | `autoloop_reset_agent` | `run_id`, `agent` ('planner' / 'coder' / 'reviewer'), `force?`, `eager_restart?` | Reset one subagent. Planner reset requires `force: true`. |
87
+
88
+ ## Planner-emitted control tools
89
+
90
+ The Planner controls the run by emitting fenced ` ```autoloop ` JSON blocks
91
+ inside its replies. The dispatcher parses them out and applies them. You
92
+ never see the JSON — only the Planner's narrative.
93
+
94
+ | Tool | Args | What |
95
+ |---|---|---|
96
+ | `notify_user` | `level` ('info' / 'warn' / 'decision' / 'error'), `summary`, `detail?`, `channel?` ('auto' / 'wechat' / 'webchat' / 'both' / 'email') | Push you out-of-band. |
97
+ | `spawn_subagents` | `coder_model?`, `reviewer_model?`, `initial_directive?` | Start Coder + Reviewer. Only after explicit user approval. |
98
+ | `send_directive` | `goal`, `constraints?`, `success_criteria?`, `max_attempts?` | Next iter's instruction to Coder. |
99
+ | `pause_loop` | `reason` | Halt subloop at next iter boundary; chat keeps working. |
100
+ | `resume_loop` | — | Resume after pause. |
101
+ | `terminate` | `reason` | End run. |
102
+ | `update_push_policy` | partial PushPolicy | Mutate notification rules (e.g. when you say "tell me every iter"). |
103
+ | `write_plan_committed` | `message?` | git-commit current plan.md. |
104
+ | `write_goal_committed` | `message?` | git-commit current goal.json. |
105
+
106
+ ## Default push policy
107
+
108
+ | Event | Default |
109
+ |---|---|
110
+ | on_start | info / wechat ("loop started, will notify on issues") |
111
+ | on_iter_done_ok | silent |
112
+ | on_target_hit | info / both (webchat + wechat) |
113
+ | on_metric_regression_2 | warn / both |
114
+ | on_reviewer_reject_2 | warn / both |
115
+ | on_phase_error | error / both |
116
+ | on_stall_30min | warn / wechat |
117
+ | on_decision_needed | decision / both |
118
+
119
+ 5-minute dedup on (level, summary) prevents duplicate pushes from the same
120
+ event. Channel chain: `auto` walks wechat → whatsapp → email; `wechat` /
121
+ `webchat` / `email` route directly; `both` does webchat (if session known)
122
+ + wechat fallback chain.
123
+
124
+ ## Ledger layout
125
+
126
+ ```
127
+ <workspace>/tasks/<run_id>/
128
+ ├── plan.md # Planner-authored, git-committed
129
+ ├── goal.json # Planner-authored, git-committed
130
+ ├── push_log.jsonl # every notify_user attempt + channel used
131
+ ├── reviewer_sandbox/ # Reviewer cwd; restaged per iter
132
+ │ ├── plan.md # copy
133
+ │ ├── goal.json # copy
134
+ │ ├── iter-N/ # this iter's directive + diff + eval
135
+ │ ├── prior_verdict.json
136
+ │ └── reviewer_memory.md # persistent across iters
137
+ └── iter/<n>/
138
+ ├── directive.json # Planner → Coder
139
+ ├── eval_output.json # what Coder reported
140
+ ├── diff.patch # git diff of the iter
141
+ ├── verdict.json # Reviewer decision + audit notes
142
+ └── coder_summary.txt
143
+ ```
144
+
145
+ The orchestrator git-commits each iter automatically. Coder must NOT call
146
+ `git commit` itself — that confuses the diff log.
147
+
148
+ ## Backend HTTP / SSE
149
+
150
+ | Endpoint | Returns |
151
+ |---|---|
152
+ | `GET /autoloop/list` | `{ ok, runs: AutoloopState[] }` |
153
+ | `GET /autoloop/<id>/state` | `{ ok, state: AutoloopState }` |
154
+ | `GET /autoloop/<id>/push_log` | `{ ok, entries: PushLogEntry[] }` |
155
+ | `GET /autoloop/<id>/events` | SSE: `snapshot` / `message` / `state` / `push` / `iter_done` / `planner_reply` / `coder_reply` / `reviewer_reply` / `terminated` |
156
+
157
+ The 3-pane UI consumes these endpoints:
158
+ - **Left**: Planner chat (subscribes to `planner_reply`)
159
+ - **Center**: Coder activity (`coder_reply` + `iter_done`)
160
+ - **Right**: Reviewer verdicts (`reviewer_reply`)
161
+ - **Top bar**: state (status / iter / metric)
162
+ - **Bottom**: push_log
163
+
164
+ The UI itself ships in a separate cross-repo PR.
165
+
166
+ ## `goal.json` shape
167
+
168
+ The Planner authors goal.json based on your conversation. There is no
169
+ hard schema — the Coder reads what's there and runs the eval the Planner
170
+ wrote down. A typical shape:
171
+
172
+ ```jsonc
173
+ {
174
+ "scalar": {
175
+ "name": "test_pass_rate",
176
+ "direction": "max",
177
+ "extract_cmd": "bash eval.sh | grep -oE 'metric=[0-9.]+' | cut -d= -f2",
178
+ "target": 1.0
179
+ },
180
+ "gates": [
181
+ { "name": "tests_pass", "cmd": "npm test", "must": "exit-0" }
182
+ ],
183
+ "termination": {
184
+ "max_iters": 10,
185
+ "scalar_target_hit": true
186
+ }
187
+ }
188
+ ```
189
+
190
+ The Planner will riff on this shape during your chat and ask if it's right.
191
+
192
+ ## Hard rules (Coder / Reviewer)
193
+
194
+ - ❌ Coder does NOT modify `plan.md`, `goal.json`, or anything under `tasks/`. Planner owns those.
195
+ - ❌ Coder does NOT manually `git commit` — orchestrator commits per iter.
196
+ - ❌ Reviewer modifies nothing outside its sandbox cwd.
197
+ - ❌ Reviewer never pings Planner / Coder for clarification — operates from artifacts only.
198
+ - ✅ Coder leaves notes in `coder_notes.md` for things future iters need to know.
199
+ - ✅ Reviewer accumulates "fakery patterns I've seen" in `reviewer_memory.md` (persists across iters).
200
+ - ✅ Reviewer defaults to `hold` under uncertainty; only `advance` after independent verification.
201
+
202
+ ## Smoke test
203
+
204
+ `scripts/smoke-autoloop.ts` runs a buggy `add_two` scenario end-to-end with
205
+ Opus Planner + Sonnet × 2. Validates plan.md / goal.json commit, spawn,
206
+ iter 0 ledger artifacts (`directive` + `eval_output` + `diff.patch` +
207
+ `verdict`), and termination on `target_hit`. Cost ~$1-3, wall-clock
208
+ ~5-15 min. Run with `npx tsx scripts/smoke-autoloop.ts` (requires
209
+ `~/.claude/settings.json` to have your auth env).
210
+
211
+ ## Known limitations
212
+
213
+ - **No auto-compact on token budget.** Manual `autoloop_reset_agent` covers
214
+ the same recovery path. Auto-compact is queued for a follow-up once
215
+ `ISession.getStats` exposes token-usage hooks.
216
+ - **One-way push.** WeChat → Planner inbound replies are not yet wired (would
217
+ need an openclaw-gateway tmux-passthrough route). Reply via webchat /
218
+ `autoloop_chat`.
219
+ - **No webchat UI yet.** Backend SSE is shipped; the UI is a separate
220
+ cross-repo PR in ChatGPT-Next-Web.
221
+ - **No fork / population mode.** Single linear iter trajectory per run.
222
+ - **Cross-run knowledge isolated.** Each run's `reviewer_memory.md` and
223
+ `coder_notes.md` live in that run's ledger; no shared meta-store yet.