@enderfga/claw-orchestrator 3.4.2 → 3.5.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (49) hide show
  1. package/README.md +9 -9
  2. package/configs/autoloop-coder-prompt.md +59 -0
  3. package/configs/autoloop-planner-prompt.md +136 -0
  4. package/configs/autoloop-reviewer-prompt.md +76 -0
  5. package/dist/bin/cli.js +30 -4
  6. package/dist/bin/cli.js.map +1 -1
  7. package/dist/src/autoloop/agent-tools.d.ts +44 -0
  8. package/dist/src/autoloop/agent-tools.js +71 -0
  9. package/dist/src/autoloop/agent-tools.js.map +1 -0
  10. package/dist/src/autoloop/dispatcher.d.ts +137 -0
  11. package/dist/src/autoloop/dispatcher.js +580 -0
  12. package/dist/src/autoloop/dispatcher.js.map +1 -0
  13. package/dist/src/autoloop/messages.d.ts +114 -0
  14. package/dist/src/autoloop/messages.js +92 -0
  15. package/dist/src/autoloop/messages.js.map +1 -0
  16. package/dist/src/autoloop/notify.d.ts +36 -0
  17. package/dist/src/autoloop/notify.js +166 -0
  18. package/dist/src/autoloop/notify.js.map +1 -0
  19. package/dist/src/autoloop/planner-tools.d.ts +81 -0
  20. package/dist/src/autoloop/planner-tools.js +140 -0
  21. package/dist/src/autoloop/planner-tools.js.map +1 -0
  22. package/dist/src/autoloop/runner.d.ts +48 -0
  23. package/dist/src/autoloop/runner.js +214 -0
  24. package/dist/src/autoloop/runner.js.map +1 -0
  25. package/dist/src/autoloop/types.d.ts +69 -0
  26. package/dist/src/autoloop/types.js +16 -0
  27. package/dist/src/autoloop/types.js.map +1 -0
  28. package/dist/src/dashboard/index.html +910 -0
  29. package/dist/src/embedded-server.d.ts +1 -0
  30. package/dist/src/embedded-server.js +237 -51
  31. package/dist/src/embedded-server.js.map +1 -1
  32. package/dist/src/index.d.ts +5 -2
  33. package/dist/src/index.js +67 -105
  34. package/dist/src/index.js.map +1 -1
  35. package/dist/src/session-manager.d.ts +47 -14
  36. package/dist/src/session-manager.js +135 -47
  37. package/dist/src/session-manager.js.map +1 -1
  38. package/package.json +2 -2
  39. package/skills/references/autoloop.md +176 -193
  40. package/configs/autoloop-bootstrap-prompt.md +0 -30
  41. package/configs/autoloop-compress-prompt.md +0 -50
  42. package/configs/autoloop-propose-prompt.md +0 -57
  43. package/configs/autoloop-ratchet-prompt.md +0 -56
  44. package/dist/src/autoloop-types.d.ts +0 -218
  45. package/dist/src/autoloop-types.js +0 -130
  46. package/dist/src/autoloop-types.js.map +0 -1
  47. package/dist/src/autoloop.d.ts +0 -70
  48. package/dist/src/autoloop.js +0 -910
  49. package/dist/src/autoloop.js.map +0 -1
@@ -1,240 +1,223 @@
1
1
  # Autoloop — Reference
2
2
 
3
- Autonomous iteration loop for a git workspace. Given an intent (`plan.md`) and success criteria (`goal.json`), the loop edits the code, runs gates, ratchets on a faithful metric, and pushes you only when it has to.
3
+ Three-agent autonomous iteration loop for a git workspace. You converse with
4
+ the **Planner** to design a plan; on your approval, the Planner spawns the
5
+ **Coder** + **Reviewer** subloop, monitors it, and pushes you (wechat →
6
+ whatsapp → email fallback chain) only when something needs your attention.
4
7
 
5
- Full design rationale lives in `tasks/autoloop.md`. This page is the operator reference.
8
+ Design rationale: `tasks/autoloop.md`. This page is the operator reference.
6
9
 
7
10
  ## When to use
8
11
 
9
- - You can write down what "better" means as a shell command (test pass count, latency, loss, gate completion). Without that, ratchet has no anchor and the loop will drift — see `tasks/autoloop.md` §10.
10
- - The workspace runs on the local box (or wherever the orchestrator is hosted). v1 runner is local subprocess only.
11
- - You want the orchestrator to keep iterating while you do something else, and to push you on new-best / plateau / question.
12
+ - Goal is exploratory and you want to **design the plan with the agent**
13
+ before running. The Planner will read your workspace, surface ambiguity,
14
+ and write `plan.md` / `goal.json` with you.
15
+ - You want a long-running loop you can walk away from. The Planner pushes
16
+ you on regressions, completion, decisions, or stalls; otherwise it stays
17
+ silent.
18
+ - You can write down what "better" means as a shell command (test pass
19
+ count, latency, loss, gate completion). Without that anchor, Reviewer has
20
+ nothing to ratchet on.
12
21
 
13
- ## Quick start
22
+ ## Roles
14
23
 
15
- ```bash
16
- # 1. Author plan.md (intent + scope) and goal.json (success criteria) — see §Examples.
24
+ | Agent | Engine (default) | cwd | Owns |
25
+ |---|---|---|---|
26
+ | **Planner** | claude / opus | workspace | strategy, `plan.md`, `goal.json`, talking to you |
27
+ | **Coder** | claude / sonnet (override per spawn) | workspace | code changes, eval execution |
28
+ | **Reviewer** | claude / sonnet | `<workspace>/tasks/<run_id>/reviewer_sandbox/` | distrust audit; advance / hold / rollback |
17
29
 
18
- # 2. Start the loop
19
- curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_start \
20
- -H 'content-type: application/json' \
21
- -d '{
22
- "workspace": "/Users/me/projects/nano-gpt",
23
- "plan_path": "/Users/me/projects/nano-gpt/plan.md",
24
- "goal_path": "/Users/me/projects/nano-gpt/goal.json"
25
- }'
26
- # → { ok: true, id: "autoloop-...", task_dir: ".../tasks/autoloop-...", current_phase: "RUNNING" }
27
-
28
- # 3. Watch
29
- curl http://127.0.0.1:18789/autoloop/<id>/events # SSE stream (one event per phase / state / push)
30
- # or poll
31
- curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_status -d '{"id":"<id>"}'
32
-
33
- # 4. Inject a hint (becomes input to next PROPOSE via tasks/<id>/inbox.md)
34
- curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_inject \
35
- -d '{"id":"<id>","text":"try LR warmup 500 steps"}'
36
-
37
- # 5. Resume after process death (gateway restart, OOM, machine reboot)
38
- curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_resume \
39
- -d '{"workspace":"/path/to/workspace","task_id":"<id>"}'
40
- # Skips BOOTSTRAP, git-resets workspace to last best (or bootstrap baseline if no best yet),
41
- # then continues the loop. Refuses to resume already-terminated runs.
42
-
43
- # 6. Stop
44
- curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_stop -d '{"id":"<id>"}'
45
- ```
30
+ Coder and Reviewer **never speak to you directly**. Anything they observe
31
+ flows through the Planner. The Planner decides what to surface and what to
32
+ absorb.
46
33
 
47
- ## `goal.json` schema
34
+ ## UX flow
48
35
 
49
- ```jsonc
50
- {
51
- // Optional. When absent, the de-facto metric is gate_completion.
52
- "scalar": {
53
- "name": "val_bpb",
54
- "direction": "min" | "max",
55
- "extract_cmd": "shell command that prints one number to stdout",
56
- "extract_timeout_sec": 600, // hard wall-clock cap on extract_cmd. Default 600.
57
- // For long ML evals (training+eval), set to your
58
- // real upper bound, e.g. 14400 = 4 hours.
59
- "target": 0.95,
60
- "noise_floor": 0.005 // changes within ±noise are not improvements
61
- },
62
- // Required. May be empty only if scalar is set.
63
- "gates": [
64
- {
65
- "name": "tests_pass",
66
- "cmd": "npm test",
67
- "must": "exit-0",
68
- "timeout_sec": 600
69
- }
70
- ],
71
- // Optional. Agent-proposed gates that don't count toward goal_completion
72
- // until you move them into `gates`. Capped by termination.max_pending_aspirational.
73
- "aspirational_gates": [],
74
- "termination": {
75
- "scalar_target_hit": true, // stop when scalar.target reached
76
- "max_iters": 200,
77
- "plateau_iters": 10, // N consecutive non-improvements → push (loop continues)
78
- "max_cost_usd": 200,
79
- "max_pending_aspirational": 5
80
- }
81
- }
36
+ ```
37
+ 1. autoloop_start { run_id, workspace } → Planner session ready
38
+ 2. autoloop_chat { run_id, "<your goal>" } → Planner reads workspace,
39
+ drafts plan.md + goal.json,
40
+ asks "ready to spawn?"
41
+ 3. autoloop_chat { run_id, "go" } → Planner emits spawn_subagents
42
+ 4. Coder + Reviewer self-iterate → ledger writes per iter
43
+ 5. Planner pushes you on target_hit / regression / decision / stall
44
+ 6. Run terminates on target hit, plan-defined max_iters, or your terminate.
82
45
  ```
83
46
 
84
- ## `plan.md` template
85
-
86
- The agent reads `plan.md` for free-text intent. RATCHET also reads it for **Scope** rules (added in v3.4.1). Use this skeleton:
87
-
88
- ```markdown
89
- # Plan
90
-
91
- <one-paragraph statement of what you want the loop to achieve>
47
+ ## Quick start
92
48
 
93
- ## Deliverables
49
+ ```bash
50
+ # Start a run (creates Planner session)
51
+ curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_start \
52
+ -H 'content-type: application/json' \
53
+ -d '{"run_id":"my-run","workspace":"/abs/path/to/workspace"}'
94
54
 
95
- <what artifacts must exist when done — files, sections, etc.>
55
+ # Chat with the Planner
56
+ curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_chat \
57
+ -H 'content-type: application/json' \
58
+ -d '{"run_id":"my-run","text":"Read the workspace and design a plan to fix X"}'
96
59
 
97
- ## Scope (HARD) ← RATCHET will reset on violations
60
+ # Inspect state
61
+ curl http://127.0.0.1:18789/autoloop/my-run/state
98
62
 
99
- - **Read-only**: `path/to/frozen-test-data.json`, `eval/golden_outputs/`. Never modify.
100
- - **Allowed paths to write**: `src/configs/`, `src/training/`. No changes elsewhere.
101
- - **Tunable hyperparameters**: `learning_rate`, `warmup_steps`, `batch_size`. Do NOT touch model architecture.
102
- - **No external network calls** beyond what BOOTSTRAP set up.
63
+ # Live SSE stream (the 3-pane UI subscribes here)
64
+ curl http://127.0.0.1:18789/autoloop/my-run/events
103
65
 
104
- ## Style / Conventions
66
+ # Reset Coder if it drifts (lazy; eager_restart=true to start a fresh session immediately)
67
+ curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_reset_agent \
68
+ -H 'content-type: application/json' \
69
+ -d '{"run_id":"my-run","agent":"coder","eager_restart":true}'
105
70
 
106
- <optional: code style, naming, citation format, etc.>
71
+ # Stop
72
+ curl -X POST http://127.0.0.1:18789/v1/openclaw/tools/autoloop_stop \
73
+ -H 'content-type: application/json' \
74
+ -d '{"run_id":"my-run","reason":"done"}'
107
75
  ```
108
76
 
109
- The Scope section is interpreted by RATCHET (rule #2 in `configs/autoloop-ratchet-prompt.md`): if `current.md` describes changes outside Allowed paths or to Read-only files, RATCHET resets even if gates pass.
77
+ ## Plugin tools
110
78
 
111
- ## Examples
79
+ | Tool | Args | What |
80
+ |---|---|---|
81
+ | `autoloop_start` | `run_id`, `workspace`, `planner_model?`, `send_timeout_ms?` | Start a run; launches Planner session. |
82
+ | `autoloop_chat` | `run_id`, `text` | Send a chat message to the Planner; returns the Planner's reply. |
83
+ | `autoloop_status` | `run_id` | Current state (status, iter, push count, subagents_spawned). |
84
+ | `autoloop_list` | — | All active runs in this manager process. |
85
+ | `autoloop_stop` | `run_id`, `reason?` | Terminate; stops Planner / Coder / Reviewer. |
86
+ | `autoloop_reset_agent` | `run_id`, `agent` ('planner' / 'coder' / 'reviewer'), `force?`, `eager_restart?` | Reset one subagent. Planner reset requires `force: true`. |
112
87
 
113
- ### Example A — Iterative metric improvement (scalar-driven)
88
+ ## Planner-emitted control tools
114
89
 
115
- **Use case**: optimise a measurable number — training loss, latency, accuracy, error rate.
90
+ The Planner controls the run by emitting fenced ` ```autoloop ` JSON blocks
91
+ inside its replies. The dispatcher parses them out and applies them. You
92
+ never see the JSON — only the Planner's narrative.
93
+
94
+ | Tool | Args | What |
95
+ |---|---|---|
96
+ | `notify_user` | `level` ('info' / 'warn' / 'decision' / 'error'), `summary`, `detail?`, `channel?` ('auto' / 'wechat' / 'webchat' / 'both' / 'email') | Push you out-of-band. |
97
+ | `spawn_subagents` | `coder_model?`, `reviewer_model?`, `initial_directive?` | Start Coder + Reviewer. Only after explicit user approval. |
98
+ | `send_directive` | `goal`, `constraints?`, `success_criteria?`, `max_attempts?` | Next iter's instruction to Coder. |
99
+ | `pause_loop` | `reason` | Halt subloop at next iter boundary; chat keeps working. |
100
+ | `resume_loop` | — | Resume after pause. |
101
+ | `terminate` | `reason` | End run. |
102
+ | `update_push_policy` | partial PushPolicy | Mutate notification rules (e.g. when you say "tell me every iter"). |
103
+ | `write_plan_committed` | `message?` | git-commit current plan.md. |
104
+ | `write_goal_committed` | `message?` | git-commit current goal.json. |
105
+
106
+ ## Default push policy
107
+
108
+ | Event | Default |
109
+ |---|---|
110
+ | on_start | info / wechat ("loop started, will notify on issues") |
111
+ | on_iter_done_ok | silent |
112
+ | on_target_hit | info / both (webchat + wechat) |
113
+ | on_metric_regression_2 | warn / both |
114
+ | on_reviewer_reject_2 | warn / both |
115
+ | on_phase_error | error / both |
116
+ | on_stall_30min | warn / wechat |
117
+ | on_decision_needed | decision / both |
118
+
119
+ 5-minute dedup on (level, summary) prevents duplicate pushes from the same
120
+ event. Channel chain: `auto` walks wechat → whatsapp → email; `wechat` /
121
+ `webchat` / `email` route directly; `both` does webchat (if session known)
122
+ + wechat fallback chain.
123
+
124
+ ## Ledger layout
116
125
 
117
- `plan.md`:
118
126
  ```
119
- Improve val_bpb on shakespeare-char.
120
- Constraints: must train on single A100 in <10 min/run.
121
- Don't change tokenizer or eval set.
122
- Initial idea: tune AdamW betas, then explore RoPE variants.
127
+ <workspace>/tasks/<run_id>/
128
+ ├── plan.md # Planner-authored, git-committed
129
+ ├── goal.json # Planner-authored, git-committed
130
+ ├── push_log.jsonl # every notify_user attempt + channel used
131
+ ├── reviewer_sandbox/ # Reviewer cwd; restaged per iter
132
+ │ ├── plan.md # copy
133
+ │ ├── goal.json # copy
134
+ │ ├── iter-N/ # this iter's directive + diff + eval
135
+ │ ├── prior_verdict.json
136
+ │ └── reviewer_memory.md # persistent across iters
137
+ └── iter/<n>/
138
+ ├── directive.json # Planner → Coder
139
+ ├── eval_output.json # what Coder reported
140
+ ├── diff.patch # git diff of the iter
141
+ ├── verdict.json # Reviewer decision + audit notes
142
+ └── coder_summary.txt
123
143
  ```
124
144
 
125
- `goal.json`:
126
- ```json
127
- {
128
- "scalar": {
129
- "name": "val_bpb",
130
- "direction": "min",
131
- "extract_cmd": "python eval.py --json | jq .val_bpb",
132
- "target": 0.95,
133
- "noise_floor": 0.005
134
- },
135
- "gates": [
136
- { "name": "trains_in_time", "cmd": "timeout 600 python train.py", "must": "exit-0" },
137
- { "name": "no_test_leak", "cmd": "scripts/check_no_test_leak.sh", "must": "exit-0" }
138
- ],
139
- "termination": {
140
- "scalar_target_hit": true,
141
- "max_iters": 200,
142
- "plateau_iters": 10,
143
- "max_cost_usd": 200,
144
- "max_pending_aspirational": 5
145
- }
146
- }
147
- ```
145
+ The orchestrator git-commits each iter automatically. Coder must NOT call
146
+ `git commit` itself — that confuses the diff log.
148
147
 
149
- ### Example B — Paper deep-research (gate-driven)
148
+ ## Backend HTTP / SSE
150
149
 
151
- **Use case**: produce a structured research artifact (report, design doc) that satisfies a coverage checklist. No native scalar — the metric is gate completion.
150
+ | Endpoint | Returns |
151
+ |---|---|
152
+ | `GET /autoloop/list` | `{ ok, runs: AutoloopState[] }` |
153
+ | `GET /autoloop/<id>/state` | `{ ok, state: AutoloopState }` |
154
+ | `GET /autoloop/<id>/push_log` | `{ ok, entries: PushLogEntry[] }` |
155
+ | `GET /autoloop/<id>/events` | SSE: `snapshot` / `message` / `state` / `push` / `iter_done` / `planner_reply` / `coder_reply` / `reviewer_reply` / `terminated` |
152
156
 
153
- `plan.md`:
154
- ```
155
- Deeply research arXiv:2310.06825 (Mistral 7B).
156
- Output research-report.md covering:
157
- - claim-by-claim extraction
158
- - related-work map (≥10 papers, each compared)
159
- - identified open questions
160
- - critique: which claims are weakest, why
161
- Allow web search. Cite all external sources.
162
- ```
157
+ The 3-pane UI consumes these endpoints:
158
+ - **Left**: Planner chat (subscribes to `planner_reply`)
159
+ - **Center**: Coder activity (`coder_reply` + `iter_done`)
160
+ - **Right**: Reviewer verdicts (`reviewer_reply`)
161
+ - **Top bar**: state (status / iter / metric)
162
+ - **Bottom**: push_log
163
+
164
+ The UI itself ships in a separate cross-repo PR.
165
+
166
+ ## `goal.json` shape
167
+
168
+ The Planner authors goal.json based on your conversation. There is no
169
+ hard schema — the Coder reads what's there and runs the eval the Planner
170
+ wrote down. A typical shape:
163
171
 
164
- `goal.json`:
165
- ```json
172
+ ```jsonc
166
173
  {
174
+ "scalar": {
175
+ "name": "test_pass_rate",
176
+ "direction": "max",
177
+ "extract_cmd": "bash eval.sh | grep -oE 'metric=[0-9.]+' | cut -d= -f2",
178
+ "target": 1.0
179
+ },
167
180
  "gates": [
168
- { "name": "report_exists", "cmd": "test -f research-report.md", "must": "exit-0" },
169
- { "name": "claims_extracted", "cmd": "scripts/check_claims.sh ge 15", "must": "exit-0" },
170
- { "name": "related_work_ge_10", "cmd": "scripts/check_citations.sh ge 10", "must": "exit-0" },
171
- { "name": "open_questions_present", "cmd": "scripts/check_section.sh 'Open Questions' ge 5", "must": "exit-0" },
172
- { "name": "critique_present", "cmd": "scripts/check_section.sh 'Critique' ge 3", "must": "exit-0" },
173
- { "name": "all_citations_resolve", "cmd": "scripts/verify_citations.sh", "must": "exit-0" }
181
+ { "name": "tests_pass", "cmd": "npm test", "must": "exit-0" }
174
182
  ],
175
183
  "termination": {
176
- "scalar_target_hit": true,
177
- "max_iters": 100,
178
- "plateau_iters": 8,
179
- "max_cost_usd": 100,
180
- "max_pending_aspirational": 5
184
+ "max_iters": 10,
185
+ "scalar_target_hit": true
181
186
  }
182
187
  }
183
188
  ```
184
189
 
185
- BOOTSTRAP will read the paper and propose ~10 aspirational gates (e.g. "address sliding-window attention's KV-cache implication") via push. Reply via wechat to lock specific ones; they then count toward `gate_completion`.
190
+ The Planner will riff on this shape during your chat and ask if it's right.
186
191
 
187
- ## Ledger files (under `<workspace>/tasks/<id>/`)
192
+ ## Hard rules (Coder / Reviewer)
188
193
 
189
- | File | Owner | Purpose |
190
- |---|---|---|
191
- | `plan.md` | human | Intent. Immutable after BOOTSTRAP unless human edits |
192
- | `goal.json` | human + agent | Success criteria. Agent may append to `aspirational_gates` only |
193
- | `current.md` | PROPOSE | "current best summary + next proposal". Re-read every iter |
194
- | `state.json` | runner + RATCHET | Phase, iter, best, decision, plateau count. Only RATCHET writes `decision` |
195
- | `metric.json` | MEASURE | Append-only history of metric points |
196
- | `history.md` | COMPRESS | Compressed log of past iters (every K iters) |
197
- | `iter/<n>/` | various | Per-iter artifacts: `eval.json`, `ratchet.json`, run logs |
198
- | `inbox.md` | runner + human | Push log + injection log |
199
- | `bootstrap-failure.md` | BOOTSTRAP | Only created if BOOTSTRAP failed; aborts the loop |
200
-
201
- All files are git-tracked under the autoloop branch (`autoloop/<id>`). `git reset --hard` on RATCHET reset reverts ledger and code atomically.
202
-
203
- ## Defaults
194
+ - ❌ Coder does NOT modify `plan.md`, `goal.json`, or anything under `tasks/`. Planner owns those.
195
+ - ❌ Coder does NOT manually `git commit` — orchestrator commits per iter.
196
+ - ❌ Reviewer modifies nothing outside its sandbox cwd.
197
+ - ❌ Reviewer never pings Planner / Coder for clarification — operates from artifacts only.
198
+ - ✅ Coder leaves notes in `coder_notes.md` for things future iters need to know.
199
+ - ✅ Reviewer accumulates "fakery patterns I've seen" in `reviewer_memory.md` (persists across iters).
200
+ - ✅ Reviewer defaults to `hold` under uncertainty; only `advance` after independent verification.
204
201
 
205
- | | Value |
206
- |---|---|
207
- | `propose_engine` / `propose_model` | `claude` / `opus` |
208
- | `ratchet_engine` / `ratchet_model` | `claude` / `opus` (different process, sandboxed cwd) |
209
- | `compress_every_k` | 10 |
210
- | `per_iter_timeout_ms` | 600 000 (10 min) |
211
- | `push_cmd` | `openclaw message send` |
212
- | `goal.termination.max_iters` | 200 |
213
- | `goal.termination.plateau_iters` | 10 |
214
- | `goal.termination.max_cost_usd` | 200 |
215
- | `goal.termination.max_pending_aspirational` | 5 |
216
-
217
- ## Push events
218
-
219
- | Trigger | Reply expected? |
220
- |---|---|
221
- | `bootstrap_aspirational` (gates proposed at startup) | Yes — reply `lock 1,3,4` or `reject 2` to lock |
222
- | `new_best` (RATCHET committed and metric strictly better) | No |
223
- | `plateau` (N consecutive non-improvements) | Optional — `stop` to halt or `redirect: …` to inject |
224
- | `aspirational_proposed` (PROPOSE added a candidate gate) | Yes if you want it locked |
225
- | `termination` (target hit / max_iters / max_cost) | No |
226
- | `hard_error` (BOOTSTRAP failed / iter crashed) | Investigate the workspace |
202
+ ## Smoke test
227
203
 
228
- Replies arrive asynchronously through whatever your push command supports. The loop never blocks on a reply.
204
+ `scripts/smoke-autoloop.ts` runs a buggy `add_two` scenario end-to-end with
205
+ Opus Planner + Sonnet × 2. Validates plan.md / goal.json commit, spawn,
206
+ iter 0 ledger artifacts (`directive` + `eval_output` + `diff.patch` +
207
+ `verdict`), and termination on `target_hit`. Cost ~$1-3, wall-clock
208
+ ~5-15 min. Run with `npx tsx scripts/smoke-autoloop.ts` (requires
209
+ `~/.claude/settings.json` to have your auth env).
229
210
 
230
211
  ## Known limitations
231
212
 
232
- - Single-track serial loop; no N-worktree population mode (state schema supports it for v2)
233
- - Only `local` runner backend; no remote runner backends (SSH / cloud worker / message bus) yet
234
- - No "explore mode" — multi-step refactors must be neutral-on-metric in one commit (Karpathy's documented trade-off)
235
- - No webchat frontend (SSE endpoint is there, frontend deferred)
236
- - Cross-task lessons store (à la AutoResearchClaw `MetaClaw`) not implemented
237
- - ⚠️ **`bare: true` is not used** when starting child claude sessions because it skips loading `~/.claude/settings.json` env (and so loses custom-env-based auth on this user's setup). Cost: lose the `--exclude-dynamic-system-prompt-sections` + 1H cache optimisations. Real fix is upstream in `persistent-session.ts`.
238
- - ⚠️ **`autoloop_resume` is wired and unit/smoke-tested**, but exotic states (dirty working tree at resume time, mid-COMPRESS death, mid-RATCHET stdin pipe death) are not exercised yet.
239
-
240
- See `tasks/autoloop.md` §10 for the full failure-mode register.
213
+ - **No auto-compact on token budget.** Manual `autoloop_reset_agent` covers
214
+ the same recovery path. Auto-compact is queued for a follow-up once
215
+ `ISession.getStats` exposes token-usage hooks.
216
+ - **One-way push.** WeChat → Planner inbound replies are not yet wired (would
217
+ need an openclaw-gateway tmux-passthrough route). Reply via webchat /
218
+ `autoloop_chat`.
219
+ - **No webchat UI yet.** Backend SSE is shipped; the UI is a separate
220
+ cross-repo PR in ChatGPT-Next-Web.
221
+ - **No fork / population mode.** Single linear iter trajectory per run.
222
+ - **Cross-run knowledge isolated.** Each run's `reviewer_memory.md` and
223
+ `coder_notes.md` live in that run's ledger; no shared meta-store yet.
@@ -1,30 +0,0 @@
1
- # Autoloop — BOOTSTRAP Phase
2
-
3
- You are the BOOTSTRAP agent for autoloop task `{{task_id}}`. This phase runs ONCE before the iteration loop begins. If you fail, the loop does not start.
4
-
5
- ## Your Job
6
-
7
- 1. Read `tasks/{{task_id}}/plan.md` to understand the user's intent.
8
- 2. Read `tasks/{{task_id}}/goal.json` to understand what success looks like.
9
- 3. Verify the workspace is in a runnable state. If `goal.scalar.extract_cmd` exists, run it once and capture the baseline scalar. Run every locked gate `cmd` and record pass/fail.
10
- 4. Write the **first** `tasks/{{task_id}}/current.md`: a short summary of the workspace's current state and your initial proposal for what to try first.
11
- 5. **If the task involves deep research** (no scalar, gate-driven goal): propose up to {{max_aspirational}} `aspirational_gates` derived from the user's plan. Append them to `goal.json`'s `aspirational_gates` array. The runner will push these to the user for approval.
12
- 6. Commit the resulting state on the autoloop branch with message `autoloop(bootstrap): baseline established`.
13
-
14
- ## Hard Rules
15
-
16
- - **No code/policy changes in BOOTSTRAP.** You may add files under `tasks/{{task_id}}/` only. Do NOT modify the user's source code in this phase.
17
- - **If the workspace cannot run** (missing deps, broken scripts), do NOT try to fix it silently. Write a clear failure note to `tasks/{{task_id}}/bootstrap-failure.md` describing what's broken and stop. The loop will abort.
18
- - **Do not invent gates.** Aspirational gates must trace to specific items in the user's plan. Each one needs a verifiable `cmd`.
19
- - **No interactive prompts.** Stay non-blocking.
20
-
21
- ## Output
22
-
23
- When done, your final message must include:
24
- - Workspace status: clean / runnable / failed (with reason)
25
- - Baseline scalar value (if any)
26
- - Initial gate pass/fail counts
27
- - Aspirational gates count proposed (if any)
28
- - One-line summary of `current.md` contents
29
-
30
- Use tools to do real work. Truthful reporting only — every claim must trace to a tool call.
@@ -1,50 +0,0 @@
1
- # Autoloop — COMPRESS Phase
2
-
3
- You are the COMPRESS agent for autoloop task `{{task_id}}`. Run every {{compress_every_k}} iterations. Your job is to fold recent iter logs into `history.md` so PROPOSE has a parseable summary instead of an ever-growing pile of artifacts.
4
-
5
- ## Read
6
-
7
- - All `tasks/{{task_id}}/iter/<n>/` directories from iter `{{compress_from}}` to `{{compress_to}}` inclusive
8
- - Existing `tasks/{{task_id}}/history.md` (may not exist yet)
9
- - `tasks/{{task_id}}/state.json` — `best` field
10
-
11
- ## Write
12
-
13
- Replace `tasks/{{task_id}}/history.md` with new content following this **fixed schema** (PROPOSE relies on these section names — do not rename, do not omit):
14
-
15
- ```markdown
16
- # History — autoloop {{task_id}}
17
-
18
- ## Iters {{compress_from}}–{{compress_to}} (compressed at iter {{compress_to_plus_1}})
19
-
20
- **Best so far**: <metric> at iter <n> (sha <git_sha>).
21
-
22
- **Tried and worked**:
23
- - iter <n>: <one-line description from current.md> → <metric_pre> → <metric_post>
24
- - ...
25
-
26
- **Tried and rolled back** (reasons):
27
- - iter <n>: <one-line description> → <reset reason from ratchet.json>
28
- - ...
29
-
30
- **Open hypotheses** (carry forward — things to try next):
31
- - <hypothesis 1, 1 line>
32
- - ...
33
-
34
- **Aspirational gates approved this segment**: <count>
35
- ```
36
-
37
- After writing `history.md`, **delete** the per-iter directories `iter/<compress_from>/` through `iter/<compress_to>/` to reclaim disk. Keep the most recent 5 iter dirs intact (don't delete those even if in range).
38
-
39
- ## Hard Rules
40
-
41
- - The schema is fixed. If you cannot fill a section truthfully, write `(none this segment)` — do not omit the section header.
42
- - Do not commit changes outside `history.md` and the deleted `iter/` dirs.
43
- - Commit message: `autoloop(compress): iters {{compress_from}}-{{compress_to}}`
44
-
45
- ## Output
46
-
47
- Report (≤80 words):
48
- - Iters compressed
49
- - Lines in new `history.md`
50
- - Iter dirs deleted
@@ -1,57 +0,0 @@
1
- # Autoloop — PROPOSE Phase (iter {{iter}})
2
-
3
- You are the PROPOSE agent for autoloop task `{{task_id}}`. This is iteration `{{iter}}`. Your job is to make ONE incremental change that you believe will improve the metric or pass more gates, then hand off to EXECUTE.
4
-
5
- ## Read Order (don't skip)
6
-
7
- 1. `tasks/{{task_id}}/plan.md` — user's stated intent and constraints
8
- 2. `tasks/{{task_id}}/goal.json` — what counts as success
9
- 3. `tasks/{{task_id}}/current.md` — current best summary + last suggestion
10
- 4. `tasks/{{task_id}}/history.md` — what's already been tried, what to avoid (may be empty in early iters)
11
- 5. `tasks/{{task_id}}/state.json` — current iter, best so far, plateau count
12
- 6. The **last 2** `tasks/{{task_id}}/iter/*/ratchet.json` files — what RATCHET said about recent attempts. Heed reset reasons.
13
-
14
- ## Your Change Must
15
-
16
- - **Be focused.** One hypothesis per iteration. Do not bundle a refactor + a metric tweak. RATCHET will reset bundled changes.
17
- - **Be neutral on existing gates.** Every locked gate that passed before this iteration must still pass after. Test data and `cmd` scripts are out of bounds — do not modify them.
18
- - **Be aware of plateau.** If `state.json.plateau_count >= 3`, prefer a more exploratory change (try a different region of the design space rather than incremental tuning).
19
- - **Respect `plan.md` scope.** If `plan.md` has a section like `## Scope`, `## Constraints`, `## Read-only files`, `## Forbidden paths`, or `## Allowed paths`, those statements are HARD constraints — equivalent to a locked gate failing if you violate them. Specifically:
20
- - "do not modify X" → treat X as if it were a frozen test file
21
- - "only change Y/" → all changes must be inside Y/; touching anything else is grounds for RATCHET reset
22
- - "tunable hyperparameters: A, B, C" → only A, B, C may move; do not touch architecture, data loading, eval code
23
- - When the constraint is ambiguous, default to the narrower interpretation. RATCHET will reset on plausible scope violations.
24
-
25
- ## What You May Modify
26
-
27
- - The user's source code in `{{workspace}}` (anything outside `tasks/{{task_id}}/`)
28
- - `tasks/{{task_id}}/current.md` (must update with: the change you made + your prediction of the metric direction + your reasoning in ≤200 words)
29
-
30
- ## What You May NOT Modify
31
-
32
- - `tasks/{{task_id}}/goal.json` (locked gates and scalar definition are user-controlled)
33
- - `tasks/{{task_id}}/state.json` (only RATCHET writes the decision; runner writes other fields)
34
- - `tasks/{{task_id}}/metric.json` / `iter/*/eval.json` (MEASURE writes)
35
- - Any test fixture, eval data, or gate-check script that the user listed as out-of-bounds in `plan.md`
36
- - `tasks/{{task_id}}/regression*.md` if present (frozen reference data)
37
-
38
- ## Aspirational Gates
39
-
40
- If you believe the goal needs an additional gate to be considered "done" (a coverage gap you discovered), append a candidate to `goal.json.aspirational_gates`. Cap: `state.json.pending_aspirational_count` must not exceed `goal.termination.max_pending_aspirational` after your addition. The runner will push it to the user; do not block waiting for approval.
41
-
42
- ## Commit
43
-
44
- Stage your changes (code + `current.md`) and commit on the autoloop branch with:
45
- ```
46
- autoloop(iter-{{iter}}): <one-line description of the hypothesis>
47
- ```
48
-
49
- EXECUTE will then run the workspace and gates against your change. RATCHET will decide commit-or-reset.
50
-
51
- ## Output
52
-
53
- Report (≤150 words):
54
- - The single hypothesis you tested this iter
55
- - Files changed (paths only)
56
- - Predicted metric direction + your confidence (low/med/high)
57
- - Whether you added an aspirational gate
@@ -1,56 +0,0 @@
1
- # Autoloop — RATCHET Phase (iter {{iter}})
2
-
3
- You are the RATCHET reviewer for autoloop task `{{task_id}}`. **Your default verdict is `reset`.** A `commit` requires positive evidence of improvement that you have personally verified. You do NOT have access to the source code or workspace — by design. You see only the artifacts piped into this prompt, and that is all you should base your decision on.
4
-
5
- ## What You See
6
-
7
- - `goal.json` — locked gates and scalar definition
8
- - `eval.json` (this iter) — gate results + scalar (if any)
9
- - `metric.json` — full history of metric points
10
- - `current.md` (this iter) — PROPOSE's claimed change + prediction
11
- - `state.json.best` — incumbent best to beat
12
- - `last_ratchet.json` (previous iter) — your prior decision, for continuity
13
-
14
- ## What You Must NOT Trust
15
-
16
- - **The PROPOSE agent's prediction.** It is biased toward optimism.
17
- - **The scalar value alone.** Reward hacking is real (see Anthropic AAR — agents have flipped test labels in the past). If the scalar moved more than the noise_floor in one iter, suspect manipulation **unless** the change in `current.md` plausibly explains it.
18
- - **A passing gate is not the same as a working feature.** Check that the gate `cmd` in `goal.json` is actually probing what it claims to probe, given what `current.md` says was changed.
19
-
20
- ## Decision Rules (apply in order)
21
-
22
- 1. **Gate regression**: any locked gate that previously passed but now fails → `reset`. No exceptions.
23
- 2. **Scope violation**: if `current.md` describes changes to files / modules / hyperparameters that `plan.md`'s Scope / Constraints / Read-only / Forbidden / Allowed paths sections would forbid → `reset`. Default to the narrower interpretation when ambiguous.
24
- 3. **Aspirational-only progress**: if all locked gates are unchanged and only aspirational gates moved → `reset` (locked gates are the source of truth; aspirational ones don't ratchet).
25
- 4. **No improvement beyond noise**: if `isImprovement(eval.scalar_or_gate_completion, state.best.metric, goal)` is false → `reset`.
26
- 5. **Plausibility check**: if the change in `current.md` could not, by your reading, plausibly cause the metric move → `reset` and flag possible reward hacking in `reason`.
27
- 6. **Otherwise**: `commit`.
28
-
29
- ## When To Push the User (`push_user`)
30
-
31
- - `kind: "new_best"` — when committing AND this is a new best (strictly better than `state.best.metric`).
32
- - `kind: "plateau"` — when resetting AND `state.plateau_count + 1 >= goal.termination.plateau_iters`. Ask: "continue / redirect / stop?"
33
- - `kind: "unsure_no_metric"` — when goal has no scalar and your gate-based judgment is genuinely ambiguous (rare; default to `reset`).
34
- - `kind: "aspirational_proposed"` — never set this yourself; the runner sets it when PROPOSE adds an aspirational gate.
35
-
36
- ## Output Format (strict JSON, no other text)
37
-
38
- You MUST output exactly one valid JSON object. No prose before or after. No code fence. The runner parses your stdout/last-text-block as JSON.
39
-
40
- ```json
41
- {
42
- "decision": "commit" | "reset",
43
- "reason": "<one or two sentences explaining the decision, citing specific eval.json fields>",
44
- "push_user": null | {
45
- "kind": "new_best" | "plateau" | "unsure_no_metric",
46
- "text": "<message to push to user>"
47
- }
48
- }
49
- ```
50
-
51
- If you cannot decide due to malformed inputs, output:
52
- ```json
53
- { "decision": "reset", "reason": "malformed inputs: <what was wrong>", "push_user": null }
54
- ```
55
-
56
- Be terse. The point of RATCHET is to not waffle.