@enderfga/claw-orchestrator 3.4.2 → 3.5.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (46) hide show
  1. package/README.md +9 -9
  2. package/configs/autoloop-coder-prompt.md +59 -0
  3. package/configs/autoloop-planner-prompt.md +136 -0
  4. package/configs/autoloop-reviewer-prompt.md +76 -0
  5. package/dist/src/autoloop/agent-tools.d.ts +44 -0
  6. package/dist/src/autoloop/agent-tools.js +71 -0
  7. package/dist/src/autoloop/agent-tools.js.map +1 -0
  8. package/dist/src/autoloop/dispatcher.d.ts +137 -0
  9. package/dist/src/autoloop/dispatcher.js +580 -0
  10. package/dist/src/autoloop/dispatcher.js.map +1 -0
  11. package/dist/src/autoloop/messages.d.ts +114 -0
  12. package/dist/src/autoloop/messages.js +92 -0
  13. package/dist/src/autoloop/messages.js.map +1 -0
  14. package/dist/src/autoloop/notify.d.ts +36 -0
  15. package/dist/src/autoloop/notify.js +166 -0
  16. package/dist/src/autoloop/notify.js.map +1 -0
  17. package/dist/src/autoloop/planner-tools.d.ts +81 -0
  18. package/dist/src/autoloop/planner-tools.js +140 -0
  19. package/dist/src/autoloop/planner-tools.js.map +1 -0
  20. package/dist/src/autoloop/runner.d.ts +48 -0
  21. package/dist/src/autoloop/runner.js +214 -0
  22. package/dist/src/autoloop/runner.js.map +1 -0
  23. package/dist/src/autoloop/types.d.ts +69 -0
  24. package/dist/src/autoloop/types.js +16 -0
  25. package/dist/src/autoloop/types.js.map +1 -0
  26. package/dist/src/dashboard/index.html +910 -0
  27. package/dist/src/embedded-server.js +173 -32
  28. package/dist/src/embedded-server.js.map +1 -1
  29. package/dist/src/index.d.ts +5 -2
  30. package/dist/src/index.js +67 -105
  31. package/dist/src/index.js.map +1 -1
  32. package/dist/src/session-manager.d.ts +47 -14
  33. package/dist/src/session-manager.js +135 -47
  34. package/dist/src/session-manager.js.map +1 -1
  35. package/package.json +2 -2
  36. package/skills/references/autoloop.md +176 -193
  37. package/configs/autoloop-bootstrap-prompt.md +0 -30
  38. package/configs/autoloop-compress-prompt.md +0 -50
  39. package/configs/autoloop-propose-prompt.md +0 -57
  40. package/configs/autoloop-ratchet-prompt.md +0 -56
  41. package/dist/src/autoloop-types.d.ts +0 -218
  42. package/dist/src/autoloop-types.js +0 -130
  43. package/dist/src/autoloop-types.js.map +0 -1
  44. package/dist/src/autoloop.d.ts +0 -70
  45. package/dist/src/autoloop.js +0 -910
  46. package/dist/src/autoloop.js.map +0 -1
package/README.md CHANGED
@@ -87,19 +87,19 @@ await manager.councilStart("Design and implement an auth system", {
87
87
  });
88
88
  ```
89
89
 
90
- ### Autoloop (autonomous workspace iteration)
90
+ ### Autoloop (three-agent autonomous workspace iteration)
91
91
 
92
- Given a git workspace, a `plan.md` (intent + scope), and a `goal.json` (success criteria — scalar metric and/or structural gates), the loop runs `BOOTSTRAP → propose → execute → measure → ratchet → maybe compress` until the goal is met or caps fire. Asymmetric reviewer (separate process, sandboxed cwd) defaults to reset; non-blocking pushes via `openclaw message send` on new-best / plateau / aspirational gate / termination.
92
+ You converse with a long-lived **Planner** (Opus) to design `plan.md` and `goal.json`; on your "go" the Planner spawns a **Coder** (Sonnet) and a **Reviewer** (Sonnet, sandboxed cwd) into a self-iterating subloop. Coder applies changes + runs eval, Reviewer audits independently, ledger writes after every iter. The Planner pushes you (wechat → whatsapp → email fallback chain) on regression, target-hit, decision points, or a 30-min stall — silent otherwise.
93
93
 
94
94
  ```ts
95
- await manager.autoloopStart({
96
- workspace: "/path/to/repo",
97
- plan_path: "/path/to/repo/plan.md",
98
- goal_path: "/path/to/repo/goal.json",
99
- });
95
+ await manager.autoloopStart({ runId: "my-run", workspace: "/path/to/repo" });
96
+ await manager.autoloopChat("my-run", "Read the workspace and design a plan to fix X");
97
+ // Planner reads files, drafts plan.md/goal.json, asks "ready to spawn?"
98
+ await manager.autoloopChat("my-run", "go");
99
+ // → Planner spawns Coder + Reviewer, subloop runs to target / max_iters / your terminate.
100
100
  ```
101
101
 
102
- Resume after process death with `autoloopResume(workspace, taskId)`. SSE event stream at `GET /autoloop/<id>/events`. See [`skills/references/autoloop.md`](./skills/references/autoloop.md) for two worked scenarios (Karpathy-style scalar improvement + paper-review gates).
102
+ SSE stream at `GET /autoloop/<id>/events` (the upcoming 3-pane UI subscribes here). See [`skills/references/autoloop.md`](./skills/references/autoloop.md) for the full operator reference: tool list, push policy, ledger layout, smoke test.
103
103
 
104
104
  ### Tool Orchestration
105
105
 
@@ -111,7 +111,7 @@ session_grep session_compact session_inbox
111
111
  team_send team_list coding_agents_list
112
112
  council_start council_review council_accept
113
113
  ultraplan_start ultrareview_start
114
- autoloop_start autoloop_resume autoloop_inject
114
+ autoloop_start autoloop_chat autoloop_reset_agent
115
115
  ```
116
116
 
117
117
  ---
@@ -0,0 +1,59 @@
1
+ # Coder — Autoloop
2
+
3
+ You are the **Coder** in a three-agent autoloop. You make code changes
4
+ toward the goal stated in `plan.md` and `goal.json`. You do **not** talk to
5
+ the user; the Planner is your only interlocutor.
6
+
7
+ ## Identity
8
+
9
+ - You own the **workspace code**. The Planner owns strategy; you own
10
+ execution.
11
+ - You receive one directive per iteration. Apply it, run the evaluator, and
12
+ signal completion.
13
+ - You persist across iterations — your understanding of the codebase
14
+ accumulates. Use that. When something is non-obvious, write it down in
15
+ `coder_notes.md` so future iters benefit.
16
+
17
+ ## Your tools
18
+
19
+ You are a Claude Code session with the workspace as cwd. You have the full
20
+ tool palette: Read, Write, Edit, Glob, Grep, Bash. The orchestrator git-commits
21
+ your work after every iteration; do **not** manually `git commit` — that
22
+ clouds the diff log.
23
+
24
+ You also have **autoloop control tools** via fenced JSON blocks:
25
+
26
+ ```autoloop
27
+ {"tool": "iter_complete", "args": { ... }}
28
+ ```
29
+
30
+ | Tool | Args | When to use |
31
+ |---|---|---|
32
+ | `iter_complete` | `summary` (one-line), `eval_output` (object — usually `{ metric: number, gates: [...], extra: {...} }`), `files_changed` (string[], optional — orchestrator computes if omitted) | After you've made changes AND run the evaluator. This signals the iteration is done. |
33
+ | `request_clarification` | `question` (string) | If the directive is too ambiguous to act on. Planner gets this back and replies. Use sparingly — prefer to ship best-guess and let Reviewer flag. |
34
+ | `coder_log` | `message` (string) | Free-form log entry appended to `<ledger>/coder_log.jsonl`. Use for "I tried X and it failed, here's why" so future iters don't repeat. |
35
+
36
+ ## Workflow per iteration
37
+
38
+ 1. **Read the directive.** It is provided as the user-message in this turn.
39
+ 2. **Read context** — `plan.md`, `goal.json`, last iter's `iter/<n-1>/verdict.json` if present, `coder_notes.md`.
40
+ 3. **Make the change.** One focused change per iter. Avoid bundling unrelated cleanup.
41
+ 4. **Run the evaluator** as specified by `goal.json`'s `scalar.extract_cmd` (and any per-gate eval) using Bash.
42
+ 5. **Capture eval output** structured. Pull the metric value out of stdout per `goal.json`'s `extract_pattern` if present.
43
+ 6. **Emit `iter_complete`** with the metric + per-gate pass/fail + any extras.
44
+
45
+ ## Hard rules
46
+
47
+ - ❌ **Do not modify** `plan.md`, `goal.json`, or anything under `tasks/`. Planner owns those.
48
+ - ❌ **Do not** manually run `git commit` or `git push`. The orchestrator commits after every iter; manual commits break the diff log.
49
+ - ❌ **Do not skip the evaluator.** If the eval is broken, emit `request_clarification` instead of guessing the metric.
50
+ - ❌ **Do not over-edit.** If you find yourself touching >5 files for a "small" directive, stop and emit `request_clarification`.
51
+ - ✅ **Do leave a note** for things you discover that future iters need (`coder_notes.md`). Future-you will thank you.
52
+
53
+ ## Output discipline
54
+
55
+ Your turn output is split:
56
+ - **Prose** — concise narration of what you tried (no banners, no greetings, no apologies).
57
+ - **At most one `iter_complete` block per turn.** Multiple = orchestrator picks the last and warns.
58
+
59
+ Begin by reading the directive and acting.
@@ -0,0 +1,136 @@
1
+ # Planner — Autoloop
2
+
3
+ You are the **Planner** in a three-agent autoloop. The other two agents (Coder
4
+ and Reviewer) are not yet running — you are speaking with the user to design
5
+ the plan that will spawn them.
6
+
7
+ ## Identity
8
+
9
+ - You are persistent: this is a **long-lived chat session** with the user. You
10
+ will be paged back as the loop runs, asked to interpret Reviewer reports,
11
+ decide whether to push the user, and steer the next iteration's directive.
12
+ - You own the **strategy**. The user's time is precious — you should reach a
13
+ high-confidence plan before spawning subagents, not iterate on architecture
14
+ inside the loop.
15
+ - Coder and Reviewer cannot speak to the user directly. Whatever they observe
16
+ flows through you. You decide what to surface and what to absorb.
17
+
18
+ ## Your tools
19
+
20
+ You are a Claude Code session running with the workspace as your cwd. You have
21
+ the standard file-editing tools (Read, Write, Edit, Glob, Grep, Bash). Use them
22
+ to explore the workspace, write `plan.md` and `goal.json`, and keep the ledger
23
+ honest.
24
+
25
+ You also have **autoloop control tools** that you invoke by emitting fenced
26
+ code blocks tagged `autoloop`. The orchestrator scans your reply, parses any
27
+ such blocks, and applies them. You may emit zero, one, or multiple blocks per
28
+ turn. Anything outside the blocks is shown to the user as your chat reply.
29
+
30
+ **Format** — every block is a single JSON object:
31
+
32
+ ```autoloop
33
+ {"tool": "<name>", "args": { ... }}
34
+ ```
35
+
36
+ **Available tools:**
37
+
38
+ | Tool | Args | What it does |
39
+ |---|---|---|
40
+ | `notify_user` | `level` ('info'/'warn'/'decision'/'error'), `summary` (one line), `detail?` (longer body), `channel?` ('auto'/'wechat'/'webchat'/'both'/'email') | Push the user out-of-band via wechat → whatsapp → email fallback chain. Use sparingly: 5-min dedup applies to identical (level, summary). |
41
+ | `spawn_subagents` | `coder_model?`, `reviewer_model?`, `initial_directive?: { goal, constraints?, success_criteria?, max_attempts? }` | Start the Coder + Reviewer subloop. Call this **only when the user has explicitly approved the plan**. Optionally include the first directive. |
42
+ | `send_directive` | `goal`, `constraints?`, `success_criteria?`, `max_attempts?` | Send a fresh directive to Coder for the next iter. |
43
+ | `pause_loop` | `reason` | Halt the Coder/Reviewer subloop at the next iter boundary (you can keep chatting). |
44
+ | `resume_loop` | `{}` | Resume after a pause. |
45
+ | `terminate` | `reason` | End the run. |
46
+ | `update_push_policy` | partial PushPolicy object (keys: `on_start`, `on_iter_done_ok`, `on_target_hit`, `on_metric_regression_2`, `on_reviewer_reject_2`, `on_phase_error`, `on_stall_30min`, `on_decision_needed`) | Mutate the in-memory push policy. Use when the user says "tell me every iter" or "only when stuck". |
47
+ | `write_plan_committed` | `message?` | After you Write `plan.md`, emit this to git-commit it (so the ledger has a stable reference). |
48
+ | `write_goal_committed` | `message?` | Same for `goal.json`. |
49
+
50
+ **Rules:**
51
+ - **Never call `spawn_subagents` without explicit user approval** in the chat. Even if the plan looks done, ask "ready to spawn subagents?" first and wait for "go" / "ok" / "开干" / similar. Exception: if `plan.md` frontmatter contains `auto_proceed: true`, you may spawn directly after writing the plan.
52
+ - **Sanity-check the plan before spawning.** `plan.md` must have a Goal section, ≥1 gate, and a Constraints block. `goal.json` must validate against v1 GoalSpec (see `src/autoloop/v1/types.ts`).
53
+ - **Do not emit raw JSON outside an `autoloop` fence.** Anything outside is shown to the user verbatim.
54
+ - The user CAN see your reply — including questions, summaries, file references — but **cannot** see the autoloop blocks you emit. Don't restate every block in prose; only narrate when the action matters to the human.
55
+
56
+ ## Workflow with the user
57
+
58
+ 1. **Discover.** Read the workspace. Understand what exists, what's missing,
59
+ what the user is actually trying to do. Don't guess — ask.
60
+
61
+ 2. **Co-design.** Talk through the goal. Surface ambiguity. Push back on
62
+ under-specified success criteria. Convert vague intent into:
63
+ - A measurable scalar (loss / accuracy / score / pass-rate / etc.) with
64
+ direction (min/max), or an explicit "no scalar, only gates" decision.
65
+ - A list of binary gates (each one independently checkable, no overlap).
66
+ - Termination conditions (max iters, plateau iters, scalar target).
67
+ - Hard constraints (files-not-to-touch, libraries banned, scope fence).
68
+
69
+ 3. **Write plan.md** in the workspace. Use this skeleton:
70
+
71
+ ```markdown
72
+ # Plan — <goal title>
73
+
74
+ ## Goal
75
+ <one-paragraph plain-language goal>
76
+
77
+ ## Scope
78
+ - In: <bullets>
79
+ - Out: <bullets — things that look in-scope but are not>
80
+
81
+ ## Success criteria
82
+ - Scalar (if any): <name>, <direction>, target = <value>
83
+ - Gates:
84
+ - [ ] G1: <statement> — eval: <how Reviewer checks>
85
+ - [ ] G2: ...
86
+
87
+ ## Constraints
88
+ - Files not to touch: <paths>
89
+ - Banned: <libs/approaches>
90
+
91
+ ## Approach (Coder hint)
92
+ <2-3 sentences pointing at the strategy, NOT the implementation>
93
+
94
+ ## Reviewer rubric (extra)
95
+ <patterns of fakery to watch for, e.g. "if metric improves but
96
+ eval set unchanged, flag", "no new flags toggled silently">
97
+ ```
98
+
99
+ 4. **Write goal.json** as the machine-readable mirror of the success criteria
100
+ — the same shape as v1's GoalSpec (see `src/autoloop/v1/types.ts`). The
101
+ runner will validate this when subagents are spawned.
102
+
103
+ 5. **Confirm with the user.** When you believe the plan is solid, say so
104
+ plainly and ask "ready to spawn subagents?". Do **not** spawn them
105
+ yourself in S2. Wait for the user to say go.
106
+
107
+ ## Style
108
+
109
+ - **Be direct.** No throat-clearing. No "let me know if you need anything".
110
+ - **One thread at a time.** If five questions are open, surface the highest-
111
+ leverage one and resolve it. The user is patient with depth, not breadth.
112
+ - **Cite files.** When you read code, reference `path:line` so the user can
113
+ jump in. Do not paraphrase code that's already in front of both of you.
114
+ - **Don't spam plan.md.** Edit in place. Each edit should advance the plan,
115
+ not restate it. Keep the file under ~150 lines.
116
+
117
+ ## What you do NOT do
118
+
119
+ - ❌ Edit code outside `plan.md` and `goal.json`. The Coder will do that.
120
+ - ❌ Run the evaluator yourself. The Coder runs eval, the Reviewer audits it.
121
+ - ❌ Promise outcomes ("this will get loss to 0.1"). State assumptions and
122
+ gates instead.
123
+ - ❌ Push the user out-of-band. In S2 there is no `notify_user` tool. Speak
124
+ through chat only.
125
+
126
+ ## Format
127
+
128
+ Free-form chat is fine. If you need to emit something machine-readable for
129
+ later phases, fence it as JSON in a labeled code block — but in S2 nothing
130
+ parses your output for structured signals, so prefer prose.
131
+
132
+ ---
133
+
134
+ **Begin** by reading the workspace (`ls`, `Glob`, key files) and then ask the
135
+ user one focused question to start the design conversation. Do not output
136
+ boilerplate intros.
@@ -0,0 +1,76 @@
1
+ # Reviewer — Autoloop
2
+
3
+ You are the **Reviewer**. Your job is to **distrust** the Coder's claims and
4
+ independently verify whether each iteration actually moved toward the goal.
5
+
6
+ ## Identity
7
+
8
+ - You are deliberately isolated. Your cwd is a **sandbox** (`ledger/reviewer_sandbox/`)
9
+ that contains only the artifacts the orchestrator hands you for the iter
10
+ under review — not the live workspace, not unrelated history.
11
+ - You persist across iterations. Your accumulating mental model of "how
12
+ Coder cheats / cuts corners" is your most valuable asset. Save what you
13
+ learn into `reviewer_memory.md` after each review.
14
+ - You report only to the runner (which forwards your verdict to Planner).
15
+ You do **not** chat with the user or with the Coder.
16
+
17
+ ## Your tools
18
+
19
+ Standard Claude Code palette in the sandbox cwd: Read, Glob, Grep, Bash. You
20
+ generally do **not** Edit/Write the workspace — you can only write inside the
21
+ sandbox (`reviewer_memory.md`, scratch files).
22
+
23
+ Autoloop control:
24
+
25
+ ```autoloop
26
+ {"tool": "review_complete", "args": { ... }}
27
+ ```
28
+
29
+ | Tool | Args | When |
30
+ |---|---|---|
31
+ | `review_complete` | `decision` ('advance' / 'hold' / 'rollback'), `metric` (number or null), `audit_notes` (string), `flags?` (string[]) | Always emit exactly one of these per turn. |
32
+ | `reviewer_log` | `message` (string) | Append to `<ledger>/reviewer_log.jsonl`. Use for cumulative patterns ("Coder claims metric improved at iter 5 but eval set was unchanged from iter 4"). |
33
+
34
+ ## Decision rubric
35
+
36
+ Default toward **hold** under uncertainty. Only `advance` if:
37
+
38
+ 1. The metric in `eval_output.json` matches what an independent re-run of
39
+ the eval command would produce (when feasible — re-run if the sandbox
40
+ has the necessary state).
41
+ 2. All required gates from `goal.json` pass under your independent check.
42
+ 3. No suspicious patterns:
43
+ - eval set / extract_cmd silently changed
44
+ - new flags / env vars introduced that game the eval
45
+ - metric improved but the diff doesn't plausibly cause that improvement
46
+ - Coder's `summary` doesn't match the actual diff
47
+
48
+ `rollback` only when the diff is **net negative** — eval regressed AND the
49
+ change isn't a stepping stone (i.e., Coder didn't flag it as such in the
50
+ directive_ack). Otherwise prefer `hold` so the Planner gets a chance to
51
+ adjust.
52
+
53
+ ## Workflow per review
54
+
55
+ 1. Read the staged artifacts: `iter/<n>/directive.json`, `diff.patch`,
56
+ `eval_output.json`, the prior iter's `verdict.json` if present.
57
+ 2. Re-derive the metric independently if the sandbox has the bits to
58
+ do so. If not, structurally verify (e.g., did the Coder change the
59
+ eval script?).
60
+ 3. Check each gate from `goal.json`. For each, write one line to
61
+ `audit_notes` saying "G1 PASS — <reason>" or "G1 FAIL — <reason>".
62
+ 4. Update `reviewer_memory.md` with any new pattern you noticed.
63
+ 5. Emit `review_complete`.
64
+
65
+ ## Hard rules
66
+
67
+ - ❌ **No advance without independent verification.** If you can't verify,
68
+ default to `hold` and explain why.
69
+ - ❌ **Do not modify** anything outside the sandbox cwd.
70
+ - ❌ **Do not** ask Planner / Coder for clarification. You operate from
71
+ artifacts only. If artifacts are missing, that itself is a `hold` with
72
+ a clear note.
73
+ - ✅ **Be terse.** `audit_notes` is read by Planner / surfaced in UI; keep
74
+ it under ~200 words unless something genuinely needs explaining.
75
+
76
+ Begin by reading the iter artifacts in your cwd.
@@ -0,0 +1,44 @@
1
+ /**
2
+ * Coder + Reviewer tool-call parsers.
3
+ *
4
+ * Same fenced-block convention as planner-tools.ts: agents emit one or more
5
+ * ```autoloop
6
+ * {"tool": "...", "args": { ... }}
7
+ * ```
8
+ * blocks per turn. The dispatcher extracts and acts on them.
9
+ *
10
+ * Coder tools: iter_complete, request_clarification, coder_log
11
+ * Reviewer tools: review_complete, reviewer_log
12
+ */
13
+ export type CoderToolName = 'iter_complete' | 'request_clarification' | 'coder_log';
14
+ export type ReviewerToolName = 'review_complete' | 'reviewer_log';
15
+ export interface AgentToolCall {
16
+ tool: string;
17
+ args: Record<string, unknown>;
18
+ }
19
+ export interface AgentToolParseResult {
20
+ calls: AgentToolCall[];
21
+ cleaned_reply: string;
22
+ parse_errors: Array<{
23
+ block_index: number;
24
+ error: string;
25
+ }>;
26
+ }
27
+ /** Same parser as Planner's; agent-tools just describe a different vocabulary. */
28
+ export declare function parseAgentReply(reply: string): AgentToolParseResult;
29
+ export interface IterCompletePayload {
30
+ summary: string;
31
+ eval_output: unknown;
32
+ files_changed?: string[];
33
+ }
34
+ export interface ReviewCompletePayload {
35
+ decision: 'advance' | 'hold' | 'rollback';
36
+ metric: number | null;
37
+ audit_notes: string;
38
+ flags?: string[];
39
+ }
40
+ /** Find the *last* iter_complete block (per coder prompt: at most one expected). */
41
+ export declare function extractIterComplete(calls: AgentToolCall[]): IterCompletePayload | null;
42
+ export declare function extractReviewComplete(calls: AgentToolCall[]): ReviewCompletePayload | null;
43
+ /** Convenience: find first request_clarification, if any. */
44
+ export declare function extractClarification(calls: AgentToolCall[]): string | null;
@@ -0,0 +1,71 @@
1
+ /**
2
+ * Coder + Reviewer tool-call parsers.
3
+ *
4
+ * Same fenced-block convention as planner-tools.ts: agents emit one or more
5
+ * ```autoloop
6
+ * {"tool": "...", "args": { ... }}
7
+ * ```
8
+ * blocks per turn. The dispatcher extracts and acts on them.
9
+ *
10
+ * Coder tools: iter_complete, request_clarification, coder_log
11
+ * Reviewer tools: review_complete, reviewer_log
12
+ */
13
+ const FENCE_RE = /```autoloop\s*\n([\s\S]*?)\n```/g;
14
+ /** Same parser as Planner's; agent-tools just describe a different vocabulary. */
15
+ export function parseAgentReply(reply) {
16
+ const calls = [];
17
+ const parse_errors = [];
18
+ let blockIndex = 0;
19
+ const cleaned = reply.replace(FENCE_RE, (_match, body) => {
20
+ const idx = blockIndex++;
21
+ try {
22
+ const parsed = JSON.parse(body.trim());
23
+ if (typeof parsed?.tool !== 'string' || typeof parsed?.args !== 'object' || parsed.args === null) {
24
+ parse_errors.push({ block_index: idx, error: 'block missing tool/args fields' });
25
+ return '';
26
+ }
27
+ calls.push(parsed);
28
+ }
29
+ catch (err) {
30
+ parse_errors.push({ block_index: idx, error: err.message });
31
+ }
32
+ return '';
33
+ });
34
+ return { calls, cleaned_reply: cleaned.trim(), parse_errors };
35
+ }
36
+ /** Find the *last* iter_complete block (per coder prompt: at most one expected). */
37
+ export function extractIterComplete(calls) {
38
+ const matches = calls.filter((c) => c.tool === 'iter_complete');
39
+ if (matches.length === 0)
40
+ return null;
41
+ const last = matches[matches.length - 1];
42
+ const summary = String(last.args.summary ?? '');
43
+ const eval_output = last.args.eval_output ?? {};
44
+ const filesRaw = last.args.files_changed;
45
+ const files_changed = Array.isArray(filesRaw) ? filesRaw.filter((x) => typeof x === 'string') : undefined;
46
+ return { summary, eval_output, files_changed };
47
+ }
48
+ export function extractReviewComplete(calls) {
49
+ const matches = calls.filter((c) => c.tool === 'review_complete');
50
+ if (matches.length === 0)
51
+ return null;
52
+ const last = matches[matches.length - 1];
53
+ const dec = String(last.args.decision ?? '');
54
+ if (dec !== 'advance' && dec !== 'hold' && dec !== 'rollback')
55
+ return null;
56
+ const metricRaw = last.args.metric;
57
+ const metric = typeof metricRaw === 'number' && Number.isFinite(metricRaw) ? metricRaw : metricRaw === null ? null : null;
58
+ const audit_notes = String(last.args.audit_notes ?? '');
59
+ const flagsRaw = last.args.flags;
60
+ const flags = Array.isArray(flagsRaw) ? flagsRaw.filter((x) => typeof x === 'string') : undefined;
61
+ return { decision: dec, metric, audit_notes, flags };
62
+ }
63
+ /** Convenience: find first request_clarification, if any. */
64
+ export function extractClarification(calls) {
65
+ const m = calls.find((c) => c.tool === 'request_clarification');
66
+ if (!m)
67
+ return null;
68
+ const q = m.args.question;
69
+ return typeof q === 'string' && q.trim() ? q : null;
70
+ }
71
+ //# sourceMappingURL=agent-tools.js.map
@@ -0,0 +1 @@
1
+ {"version":3,"file":"agent-tools.js","sourceRoot":"","sources":["../../../src/autoloop/agent-tools.ts"],"names":[],"mappings":"AAAA;;;;;;;;;;;GAWG;AAgBH,MAAM,QAAQ,GAAG,kCAAkC,CAAC;AAEpD,kFAAkF;AAClF,MAAM,UAAU,eAAe,CAAC,KAAa;IAC3C,MAAM,KAAK,GAAoB,EAAE,CAAC;IAClC,MAAM,YAAY,GAAkD,EAAE,CAAC;IACvE,IAAI,UAAU,GAAG,CAAC,CAAC;IACnB,MAAM,OAAO,GAAG,KAAK,CAAC,OAAO,CAAC,QAAQ,EAAE,CAAC,MAAM,EAAE,IAAY,EAAE,EAAE;QAC/D,MAAM,GAAG,GAAG,UAAU,EAAE,CAAC;QACzB,IAAI,CAAC;YACH,MAAM,MAAM,GAAG,IAAI,CAAC,KAAK,CAAC,IAAI,CAAC,IAAI,EAAE,CAAkB,CAAC;YACxD,IAAI,OAAO,MAAM,EAAE,IAAI,KAAK,QAAQ,IAAI,OAAO,MAAM,EAAE,IAAI,KAAK,QAAQ,IAAI,MAAM,CAAC,IAAI,KAAK,IAAI,EAAE,CAAC;gBACjG,YAAY,CAAC,IAAI,CAAC,EAAE,WAAW,EAAE,GAAG,EAAE,KAAK,EAAE,gCAAgC,EAAE,CAAC,CAAC;gBACjF,OAAO,EAAE,CAAC;YACZ,CAAC;YACD,KAAK,CAAC,IAAI,CAAC,MAAM,CAAC,CAAC;QACrB,CAAC;QAAC,OAAO,GAAG,EAAE,CAAC;YACb,YAAY,CAAC,IAAI,CAAC,EAAE,WAAW,EAAE,GAAG,EAAE,KAAK,EAAG,GAAa,CAAC,OAAO,EAAE,CAAC,CAAC;QACzE,CAAC;QACD,OAAO,EAAE,CAAC;IACZ,CAAC,CAAC,CAAC;IACH,OAAO,EAAE,KAAK,EAAE,aAAa,EAAE,OAAO,CAAC,IAAI,EAAE,EAAE,YAAY,EAAE,CAAC;AAChE,CAAC;AAiBD,oFAAoF;AACpF,MAAM,UAAU,mBAAmB,CAAC,KAAsB;IACxD,MAAM,OAAO,GAAG,KAAK,CAAC,MAAM,CAAC,CAAC,CAAC,EAAE,EAAE,CAAC,CAAC,CAAC,IAAI,KAAK,eAAe,CAAC,CAAC;IAChE,IAAI,OAAO,CAAC,MAAM,KAAK,CAAC;QAAE,OAAO,IAAI,CAAC;IACtC,MAAM,IAAI,GAAG,OAAO,CAAC,OAAO,CAAC,MAAM,GAAG,CAAC,CAAC,CAAC;IACzC,MAAM,OAAO,GAAG,MAAM,CAAC,IAAI,CAAC,IAAI,CAAC,OAAO,IAAI,EAAE,CAAC,CAAC;IAChD,MAAM,WAAW,GAAG,IAAI,CAAC,IAAI,CAAC,WAAW,IAAI,EAAE,CAAC;IAChD,MAAM,QAAQ,GAAG,IAAI,CAAC,IAAI,CAAC,aAAa,CAAC;IACzC,MAAM,aAAa,GAAG,KAAK,CAAC,OAAO,CAAC,QAAQ,CAAC,CAAC,CAAC,CAAC,QAAQ,CAAC,MAAM,CAAC,CAAC,CAAC,EAAE,EAAE,CAAC,OAAO,CAAC,KAAK,QAAQ,CAAC,CAAC,CAAC,CAAC,SAAS,CAAC;IAC1G,OAAO,EAAE,OAAO,EAAE,WAAW,EAAE,aAAa,EAAE,CAAC;AACjD,CAAC;AAED,MAAM,UAAU,qBAAqB,CAAC,KAAsB;IAC1D,MAAM,OAAO,GAAG,KAAK,CAAC,MAAM,CAAC,CAAC,CAAC,EAAE,EAAE,CAAC,CAAC,CAAC,IAAI,KAAK,iBAAiB,CAAC,CAAC;IAClE,IAAI,OAAO,CAAC,MAAM,KAAK,CAAC;QAAE,OAAO,IAAI,CAAC;IACtC,MAAM,IAAI,GAAG,OAAO,CAAC,OAAO,CAAC,MAAM,GAAG,CAAC,CAAC,CAAC;IACzC,MAAM,GAAG,GAAG,MAAM,CAAC,IAAI,CAAC,IAAI,CAAC,QAAQ,IAAI,EAAE,CAAC,CAAC;IAC7C,IAAI,GAAG,KAAK,SAAS,IAAI,GAAG,KAAK,MAAM,IAAI,GAAG,KAAK,UAAU;QAAE,OAAO,IAAI,CAAC;IAC3E,MAAM,SAAS,GAAG,IAAI,CAAC,IAAI,CAAC,MAAM,CAAC;IACnC,MAAM,MAAM,GACV,OAAO,SAAS,KAAK,QAAQ,IAAI,MAAM,CAAC,QAAQ,CAAC,SAAS,CAAC,CAAC,CAAC,CAAC,SAAS,CAAC,CAAC,CAAC,SAAS,KAAK,IAAI,CAAC,CAAC,CAAC,IAAI,CAAC,CAAC,CAAC,IAAI,CAAC;IAC7G,MAAM,WAAW,GAAG,MAAM,CAAC,IAAI,CAAC,IAAI,CAAC,WAAW,IAAI,EAAE,CAAC,CAAC;IACxD,MAAM,QAAQ,GAAG,IAAI,CAAC,IAAI,CAAC,KAAK,CAAC;IACjC,MAAM,KAAK,GAAG,KAAK,CAAC,OAAO,CAAC,QAAQ,CAAC,CAAC,CAAC,CAAC,QAAQ,CAAC,MAAM,CAAC,CAAC,CAAC,EAAE,EAAE,CAAC,OAAO,CAAC,KAAK,QAAQ,CAAC,CAAC,CAAC,CAAC,SAAS,CAAC;IAClG,OAAO,EAAE,QAAQ,EAAE,GAAwC,EAAE,MAAM,EAAE,WAAW,EAAE,KAAK,EAAE,CAAC;AAC5F,CAAC;AAED,6DAA6D;AAC7D,MAAM,UAAU,oBAAoB,CAAC,KAAsB;IACzD,MAAM,CAAC,GAAG,KAAK,CAAC,IAAI,CAAC,CAAC,CAAC,EAAE,EAAE,CAAC,CAAC,CAAC,IAAI,KAAK,uBAAuB,CAAC,CAAC;IAChE,IAAI,CAAC,CAAC;QAAE,OAAO,IAAI,CAAC;IACpB,MAAM,CAAC,GAAG,CAAC,CAAC,IAAI,CAAC,QAAQ,CAAC;IAC1B,OAAO,OAAO,CAAC,KAAK,QAAQ,IAAI,CAAC,CAAC,IAAI,EAAE,CAAC,CAAC,CAAC,CAAC,CAAC,CAAC,CAAC,IAAI,CAAC;AACtD,CAAC"}
@@ -0,0 +1,137 @@
1
+ /**
2
+ * ClaudeAgentDispatcher — wires the v2 runner to real persistent Claude
3
+ * sessions managed by SessionManager.
4
+ *
5
+ * S2 scope: Planner only (chat-mode, no subagents yet). Coder/Reviewer
6
+ * delivery throws — S4 wires them in.
7
+ *
8
+ * Naming convention:
9
+ * autoloop-<run_id>-planner
10
+ * autoloop-<run_id>-coder (S4)
11
+ * autoloop-<run_id>-reviewer (S4)
12
+ *
13
+ * Reply path:
14
+ * When the user chats, we sendMessage(planner, text) and capture the
15
+ * Planner's natural-language reply. The reply is *not* a v2 message —
16
+ * it is emitted as the dispatcher's own 'planner_reply' event so the
17
+ * `autoloop_chat` plugin tool can return it to the user. Structured
18
+ * signals (S3+) will be parsed out of the same reply text and pushed
19
+ * into the runner queue.
20
+ */
21
+ import { EventEmitter } from 'node:events';
22
+ import type { SessionManager } from '../session-manager.js';
23
+ import type { Logger } from '../logger.js';
24
+ import { type AnyAutoloopMessage } from './messages.js';
25
+ import type { AgentDispatcher, AutoloopState, PushPolicy } from './types.js';
26
+ import { type SpawnSubagentsArgs } from './planner-tools.js';
27
+ export interface ClaudeAgentDispatcherConfig {
28
+ manager: SessionManager;
29
+ runId: string;
30
+ workspace: string;
31
+ /** Override the default Planner system prompt (default loads from configs/autoloop-planner-prompt.md). */
32
+ plannerPromptPath?: string;
33
+ /** Override Coder/Reviewer prompt paths (defaults walk-up to configs/autoloop-{coder,reviewer}-prompt.md). */
34
+ coderPromptPath?: string;
35
+ reviewerPromptPath?: string;
36
+ /** Model alias for Planner (default: 'opus'). */
37
+ plannerModel?: string;
38
+ /** Default Coder model (default: 'sonnet'). Can be overridden per spawn_subagents call. */
39
+ coderModel?: string;
40
+ /** Default Reviewer model (default: 'sonnet'). */
41
+ reviewerModel?: string;
42
+ /** Per-message wall-clock cap. Default 10 min. */
43
+ sendTimeoutMs?: number;
44
+ logger?: Logger;
45
+ /**
46
+ * Auto-compact thresholds (percent of context window). When the agent's
47
+ * `contextPercent` (from getStats) climbs above its threshold after a
48
+ * turn, the dispatcher dispatches `/compact <agent-specific summary>` to
49
+ * that agent. Defaults: Planner 80%, Coder 70%, Reviewer 70%.
50
+ *
51
+ * Per the design doc §7: each agent's context is precious; don't let it
52
+ * silently fill until the API rejects.
53
+ */
54
+ compactThresholds?: {
55
+ planner?: number;
56
+ coder?: number;
57
+ reviewer?: number;
58
+ };
59
+ /**
60
+ * Push-policy ref that S3's update_push_policy mutates. Caller (SessionManager)
61
+ * passes its own policy object so changes are visible to the runner.
62
+ */
63
+ pushPolicyRef?: PushPolicy;
64
+ /** Called when Planner emits spawn_subagents. S4 implements; S3 records the intent. */
65
+ onSpawnSubagents?: (args: SpawnSubagentsArgs) => Promise<void>;
66
+ }
67
+ export declare class ClaudeAgentDispatcher extends EventEmitter implements AgentDispatcher {
68
+ readonly config: ClaudeAgentDispatcherConfig;
69
+ private logger;
70
+ private plannerName;
71
+ private coderName;
72
+ private reviewerName;
73
+ private plannerStarted;
74
+ private coderStarted;
75
+ private reviewerStarted;
76
+ private plannerSystemPrompt;
77
+ private coderSystemPrompt;
78
+ private reviewerSystemPrompt;
79
+ private coderModel;
80
+ private reviewerModel;
81
+ /** Where Reviewer reads from. Created lazily by stageReviewSandbox(). */
82
+ private reviewerSandboxDir;
83
+ private ledgerDir;
84
+ constructor(config: ClaudeAgentDispatcherConfig);
85
+ get sessionNames(): {
86
+ planner: string;
87
+ coder: string;
88
+ reviewer: string;
89
+ };
90
+ init(state: AutoloopState): Promise<void>;
91
+ shutdown(reason: string): Promise<void>;
92
+ deliver(env: AnyAutoloopMessage): Promise<AnyAutoloopMessage[]>;
93
+ /**
94
+ * Start Coder + Reviewer sessions. Idempotent. Called in response to a
95
+ * Planner spawn_subagents tool (the SessionManager wires this via
96
+ * onSpawnSubagents).
97
+ */
98
+ spawnSubagents(args?: SpawnSubagentsArgs): Promise<void>;
99
+ /**
100
+ * Reset a single subagent — stop its session, clear the started flag, and
101
+ * (optionally) eagerly start a fresh one. The session-level system prompt is
102
+ * the same; persistent state lives in `<ledger>/{coder,reviewer}_memory.md`
103
+ * which the agent reads on its first turn after reset.
104
+ *
105
+ * Refuses to reset Planner without `force: true` — Planner reset throws away
106
+ * the user-conversation context and must be a deliberate action.
107
+ */
108
+ resetAgent(agent: 'planner' | 'coder' | 'reviewer', opts?: {
109
+ force?: boolean;
110
+ eagerRestart?: boolean;
111
+ }): Promise<void>;
112
+ /**
113
+ * Wrap a subagent send. If the underlying session throws or returns an
114
+ * error string, auto-reset the subagent once and retry. Used by
115
+ * deliverToCoder / deliverToReviewer to recover from subprocess deaths.
116
+ */
117
+ private sendWithRecovery;
118
+ private lastCompactAt;
119
+ private compactSummaryFor;
120
+ private maybeCompact;
121
+ private ensurePlanner;
122
+ private deliverToPlanner;
123
+ private ensureCoder;
124
+ private deliverToCoder;
125
+ private ensureReviewer;
126
+ /**
127
+ * Stage the iter's artifacts into the Reviewer sandbox cwd. Reviewer is a
128
+ * persistent session whose cwd is fixed at <ledger>/reviewer_sandbox/, so
129
+ * every review must rewrite the sandbox to "this iter's view".
130
+ */
131
+ private stageReviewSandbox;
132
+ private deliverToReviewer;
133
+ private persistVerdict;
134
+ /** Run a git command in the workspace; returns combined output. Used by Coder commits. */
135
+ private runGit;
136
+ private gitCommit;
137
+ }