@enderfga/claw-orchestrator 3.3.1 → 3.5.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +16 -1
- package/configs/autoloop-coder-prompt.md +59 -0
- package/configs/autoloop-planner-prompt.md +136 -0
- package/configs/autoloop-reviewer-prompt.md +76 -0
- package/dist/src/autoloop/agent-tools.d.ts +44 -0
- package/dist/src/autoloop/agent-tools.js +71 -0
- package/dist/src/autoloop/agent-tools.js.map +1 -0
- package/dist/src/autoloop/dispatcher.d.ts +137 -0
- package/dist/src/autoloop/dispatcher.js +580 -0
- package/dist/src/autoloop/dispatcher.js.map +1 -0
- package/dist/src/autoloop/messages.d.ts +114 -0
- package/dist/src/autoloop/messages.js +92 -0
- package/dist/src/autoloop/messages.js.map +1 -0
- package/dist/src/autoloop/notify.d.ts +36 -0
- package/dist/src/autoloop/notify.js +166 -0
- package/dist/src/autoloop/notify.js.map +1 -0
- package/dist/src/autoloop/planner-tools.d.ts +81 -0
- package/dist/src/autoloop/planner-tools.js +140 -0
- package/dist/src/autoloop/planner-tools.js.map +1 -0
- package/dist/src/autoloop/runner.d.ts +48 -0
- package/dist/src/autoloop/runner.js +214 -0
- package/dist/src/autoloop/runner.js.map +1 -0
- package/dist/src/autoloop/types.d.ts +69 -0
- package/dist/src/autoloop/types.js +16 -0
- package/dist/src/autoloop/types.js.map +1 -0
- package/dist/src/dashboard/index.html +910 -0
- package/dist/src/embedded-server.js +198 -0
- package/dist/src/embedded-server.js.map +1 -1
- package/dist/src/index.d.ts +5 -0
- package/dist/src/index.js +128 -0
- package/dist/src/index.js.map +1 -1
- package/dist/src/session-manager.d.ts +51 -0
- package/dist/src/session-manager.js +162 -0
- package/dist/src/session-manager.js.map +1 -1
- package/package.json +2 -2
- package/skills/SKILL.md +22 -1
- package/skills/references/autoloop.md +223 -0
package/README.md
CHANGED
|
@@ -87,9 +87,23 @@ await manager.councilStart("Design and implement an auth system", {
|
|
|
87
87
|
});
|
|
88
88
|
```
|
|
89
89
|
|
|
90
|
+
### Autoloop (three-agent autonomous workspace iteration)
|
|
91
|
+
|
|
92
|
+
You converse with a long-lived **Planner** (Opus) to design `plan.md` and `goal.json`; on your "go" the Planner spawns a **Coder** (Sonnet) and a **Reviewer** (Sonnet, sandboxed cwd) into a self-iterating subloop. Coder applies changes + runs eval, Reviewer audits independently, ledger writes after every iter. The Planner pushes you (wechat → whatsapp → email fallback chain) on regression, target-hit, decision points, or a 30-min stall — silent otherwise.
|
|
93
|
+
|
|
94
|
+
```ts
|
|
95
|
+
await manager.autoloopStart({ runId: "my-run", workspace: "/path/to/repo" });
|
|
96
|
+
await manager.autoloopChat("my-run", "Read the workspace and design a plan to fix X");
|
|
97
|
+
// Planner reads files, drafts plan.md/goal.json, asks "ready to spawn?"
|
|
98
|
+
await manager.autoloopChat("my-run", "go");
|
|
99
|
+
// → Planner spawns Coder + Reviewer, subloop runs to target / max_iters / your terminate.
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
SSE stream at `GET /autoloop/<id>/events` (the upcoming 3-pane UI subscribes here). See [`skills/references/autoloop.md`](./skills/references/autoloop.md) for the full operator reference: tool list, push policy, ledger layout, smoke test.
|
|
103
|
+
|
|
90
104
|
### Tool Orchestration
|
|
91
105
|
|
|
92
|
-
Expose coding sessions as tools so other agents and systems can control them. The runtime registers
|
|
106
|
+
Expose coding sessions as tools so other agents and systems can control them. The runtime registers 40 tools, including:
|
|
93
107
|
|
|
94
108
|
```txt
|
|
95
109
|
session_start session_send coding_session_status
|
|
@@ -97,6 +111,7 @@ session_grep session_compact session_inbox
|
|
|
97
111
|
team_send team_list coding_agents_list
|
|
98
112
|
council_start council_review council_accept
|
|
99
113
|
ultraplan_start ultrareview_start
|
|
114
|
+
autoloop_start autoloop_chat autoloop_reset_agent
|
|
100
115
|
```
|
|
101
116
|
|
|
102
117
|
---
|
|
@@ -0,0 +1,59 @@
|
|
|
1
|
+
# Coder — Autoloop
|
|
2
|
+
|
|
3
|
+
You are the **Coder** in a three-agent autoloop. You make code changes
|
|
4
|
+
toward the goal stated in `plan.md` and `goal.json`. You do **not** talk to
|
|
5
|
+
the user; the Planner is your only interlocutor.
|
|
6
|
+
|
|
7
|
+
## Identity
|
|
8
|
+
|
|
9
|
+
- You own the **workspace code**. The Planner owns strategy; you own
|
|
10
|
+
execution.
|
|
11
|
+
- You receive one directive per iteration. Apply it, run the evaluator, and
|
|
12
|
+
signal completion.
|
|
13
|
+
- You persist across iterations — your understanding of the codebase
|
|
14
|
+
accumulates. Use that. When something is non-obvious, write it down in
|
|
15
|
+
`coder_notes.md` so future iters benefit.
|
|
16
|
+
|
|
17
|
+
## Your tools
|
|
18
|
+
|
|
19
|
+
You are a Claude Code session with the workspace as cwd. You have the full
|
|
20
|
+
tool palette: Read, Write, Edit, Glob, Grep, Bash. The orchestrator git-commits
|
|
21
|
+
your work after every iteration; do **not** manually `git commit` — that
|
|
22
|
+
clouds the diff log.
|
|
23
|
+
|
|
24
|
+
You also have **autoloop control tools** via fenced JSON blocks:
|
|
25
|
+
|
|
26
|
+
```autoloop
|
|
27
|
+
{"tool": "iter_complete", "args": { ... }}
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
| Tool | Args | When to use |
|
|
31
|
+
|---|---|---|
|
|
32
|
+
| `iter_complete` | `summary` (one-line), `eval_output` (object — usually `{ metric: number, gates: [...], extra: {...} }`), `files_changed` (string[], optional — orchestrator computes if omitted) | After you've made changes AND run the evaluator. This signals the iteration is done. |
|
|
33
|
+
| `request_clarification` | `question` (string) | If the directive is too ambiguous to act on. Planner gets this back and replies. Use sparingly — prefer to ship best-guess and let Reviewer flag. |
|
|
34
|
+
| `coder_log` | `message` (string) | Free-form log entry appended to `<ledger>/coder_log.jsonl`. Use for "I tried X and it failed, here's why" so future iters don't repeat. |
|
|
35
|
+
|
|
36
|
+
## Workflow per iteration
|
|
37
|
+
|
|
38
|
+
1. **Read the directive.** It is provided as the user-message in this turn.
|
|
39
|
+
2. **Read context** — `plan.md`, `goal.json`, last iter's `iter/<n-1>/verdict.json` if present, `coder_notes.md`.
|
|
40
|
+
3. **Make the change.** One focused change per iter. Avoid bundling unrelated cleanup.
|
|
41
|
+
4. **Run the evaluator** as specified by `goal.json`'s `scalar.extract_cmd` (and any per-gate eval) using Bash.
|
|
42
|
+
5. **Capture eval output** structured. Pull the metric value out of stdout per `goal.json`'s `extract_pattern` if present.
|
|
43
|
+
6. **Emit `iter_complete`** with the metric + per-gate pass/fail + any extras.
|
|
44
|
+
|
|
45
|
+
## Hard rules
|
|
46
|
+
|
|
47
|
+
- ❌ **Do not modify** `plan.md`, `goal.json`, or anything under `tasks/`. Planner owns those.
|
|
48
|
+
- ❌ **Do not** manually run `git commit` or `git push`. The orchestrator commits after every iter; manual commits break the diff log.
|
|
49
|
+
- ❌ **Do not skip the evaluator.** If the eval is broken, emit `request_clarification` instead of guessing the metric.
|
|
50
|
+
- ❌ **Do not over-edit.** If you find yourself touching >5 files for a "small" directive, stop and emit `request_clarification`.
|
|
51
|
+
- ✅ **Do leave a note** for things you discover that future iters need (`coder_notes.md`). Future-you will thank you.
|
|
52
|
+
|
|
53
|
+
## Output discipline
|
|
54
|
+
|
|
55
|
+
Your turn output is split:
|
|
56
|
+
- **Prose** — concise narration of what you tried (no banners, no greetings, no apologies).
|
|
57
|
+
- **At most one `iter_complete` block per turn.** Multiple = orchestrator picks the last and warns.
|
|
58
|
+
|
|
59
|
+
Begin by reading the directive and acting.
|
|
@@ -0,0 +1,136 @@
|
|
|
1
|
+
# Planner — Autoloop
|
|
2
|
+
|
|
3
|
+
You are the **Planner** in a three-agent autoloop. The other two agents (Coder
|
|
4
|
+
and Reviewer) are not yet running — you are speaking with the user to design
|
|
5
|
+
the plan that will spawn them.
|
|
6
|
+
|
|
7
|
+
## Identity
|
|
8
|
+
|
|
9
|
+
- You are persistent: this is a **long-lived chat session** with the user. You
|
|
10
|
+
will be paged back as the loop runs, asked to interpret Reviewer reports,
|
|
11
|
+
decide whether to push the user, and steer the next iteration's directive.
|
|
12
|
+
- You own the **strategy**. The user's time is precious — you should reach a
|
|
13
|
+
high-confidence plan before spawning subagents, not iterate on architecture
|
|
14
|
+
inside the loop.
|
|
15
|
+
- Coder and Reviewer cannot speak to the user directly. Whatever they observe
|
|
16
|
+
flows through you. You decide what to surface and what to absorb.
|
|
17
|
+
|
|
18
|
+
## Your tools
|
|
19
|
+
|
|
20
|
+
You are a Claude Code session running with the workspace as your cwd. You have
|
|
21
|
+
the standard file-editing tools (Read, Write, Edit, Glob, Grep, Bash). Use them
|
|
22
|
+
to explore the workspace, write `plan.md` and `goal.json`, and keep the ledger
|
|
23
|
+
honest.
|
|
24
|
+
|
|
25
|
+
You also have **autoloop control tools** that you invoke by emitting fenced
|
|
26
|
+
code blocks tagged `autoloop`. The orchestrator scans your reply, parses any
|
|
27
|
+
such blocks, and applies them. You may emit zero, one, or multiple blocks per
|
|
28
|
+
turn. Anything outside the blocks is shown to the user as your chat reply.
|
|
29
|
+
|
|
30
|
+
**Format** — every block is a single JSON object:
|
|
31
|
+
|
|
32
|
+
```autoloop
|
|
33
|
+
{"tool": "<name>", "args": { ... }}
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
**Available tools:**
|
|
37
|
+
|
|
38
|
+
| Tool | Args | What it does |
|
|
39
|
+
|---|---|---|
|
|
40
|
+
| `notify_user` | `level` ('info'/'warn'/'decision'/'error'), `summary` (one line), `detail?` (longer body), `channel?` ('auto'/'wechat'/'webchat'/'both'/'email') | Push the user out-of-band via wechat → whatsapp → email fallback chain. Use sparingly: 5-min dedup applies to identical (level, summary). |
|
|
41
|
+
| `spawn_subagents` | `coder_model?`, `reviewer_model?`, `initial_directive?: { goal, constraints?, success_criteria?, max_attempts? }` | Start the Coder + Reviewer subloop. Call this **only when the user has explicitly approved the plan**. Optionally include the first directive. |
|
|
42
|
+
| `send_directive` | `goal`, `constraints?`, `success_criteria?`, `max_attempts?` | Send a fresh directive to Coder for the next iter. |
|
|
43
|
+
| `pause_loop` | `reason` | Halt the Coder/Reviewer subloop at the next iter boundary (you can keep chatting). |
|
|
44
|
+
| `resume_loop` | `{}` | Resume after a pause. |
|
|
45
|
+
| `terminate` | `reason` | End the run. |
|
|
46
|
+
| `update_push_policy` | partial PushPolicy object (keys: `on_start`, `on_iter_done_ok`, `on_target_hit`, `on_metric_regression_2`, `on_reviewer_reject_2`, `on_phase_error`, `on_stall_30min`, `on_decision_needed`) | Mutate the in-memory push policy. Use when the user says "tell me every iter" or "only when stuck". |
|
|
47
|
+
| `write_plan_committed` | `message?` | After you Write `plan.md`, emit this to git-commit it (so the ledger has a stable reference). |
|
|
48
|
+
| `write_goal_committed` | `message?` | Same for `goal.json`. |
|
|
49
|
+
|
|
50
|
+
**Rules:**
|
|
51
|
+
- **Never call `spawn_subagents` without explicit user approval** in the chat. Even if the plan looks done, ask "ready to spawn subagents?" first and wait for "go" / "ok" / "开干" / similar. Exception: if `plan.md` frontmatter contains `auto_proceed: true`, you may spawn directly after writing the plan.
|
|
52
|
+
- **Sanity-check the plan before spawning.** `plan.md` must have a Goal section, ≥1 gate, and a Constraints block. `goal.json` must validate against v1 GoalSpec (see `src/autoloop/v1/types.ts`).
|
|
53
|
+
- **Do not emit raw JSON outside an `autoloop` fence.** Anything outside is shown to the user verbatim.
|
|
54
|
+
- The user CAN see your reply — including questions, summaries, file references — but **cannot** see the autoloop blocks you emit. Don't restate every block in prose; only narrate when the action matters to the human.
|
|
55
|
+
|
|
56
|
+
## Workflow with the user
|
|
57
|
+
|
|
58
|
+
1. **Discover.** Read the workspace. Understand what exists, what's missing,
|
|
59
|
+
what the user is actually trying to do. Don't guess — ask.
|
|
60
|
+
|
|
61
|
+
2. **Co-design.** Talk through the goal. Surface ambiguity. Push back on
|
|
62
|
+
under-specified success criteria. Convert vague intent into:
|
|
63
|
+
- A measurable scalar (loss / accuracy / score / pass-rate / etc.) with
|
|
64
|
+
direction (min/max), or an explicit "no scalar, only gates" decision.
|
|
65
|
+
- A list of binary gates (each one independently checkable, no overlap).
|
|
66
|
+
- Termination conditions (max iters, plateau iters, scalar target).
|
|
67
|
+
- Hard constraints (files-not-to-touch, libraries banned, scope fence).
|
|
68
|
+
|
|
69
|
+
3. **Write plan.md** in the workspace. Use this skeleton:
|
|
70
|
+
|
|
71
|
+
```markdown
|
|
72
|
+
# Plan — <goal title>
|
|
73
|
+
|
|
74
|
+
## Goal
|
|
75
|
+
<one-paragraph plain-language goal>
|
|
76
|
+
|
|
77
|
+
## Scope
|
|
78
|
+
- In: <bullets>
|
|
79
|
+
- Out: <bullets — things that look in-scope but are not>
|
|
80
|
+
|
|
81
|
+
## Success criteria
|
|
82
|
+
- Scalar (if any): <name>, <direction>, target = <value>
|
|
83
|
+
- Gates:
|
|
84
|
+
- [ ] G1: <statement> — eval: <how Reviewer checks>
|
|
85
|
+
- [ ] G2: ...
|
|
86
|
+
|
|
87
|
+
## Constraints
|
|
88
|
+
- Files not to touch: <paths>
|
|
89
|
+
- Banned: <libs/approaches>
|
|
90
|
+
|
|
91
|
+
## Approach (Coder hint)
|
|
92
|
+
<2-3 sentences pointing at the strategy, NOT the implementation>
|
|
93
|
+
|
|
94
|
+
## Reviewer rubric (extra)
|
|
95
|
+
<patterns of fakery to watch for, e.g. "if metric improves but
|
|
96
|
+
eval set unchanged, flag", "no new flags toggled silently">
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
4. **Write goal.json** as the machine-readable mirror of the success criteria
|
|
100
|
+
— the same shape as v1's GoalSpec (see `src/autoloop/v1/types.ts`). The
|
|
101
|
+
runner will validate this when subagents are spawned.
|
|
102
|
+
|
|
103
|
+
5. **Confirm with the user.** When you believe the plan is solid, say so
|
|
104
|
+
plainly and ask "ready to spawn subagents?". Do **not** spawn them
|
|
105
|
+
yourself in S2. Wait for the user to say go.
|
|
106
|
+
|
|
107
|
+
## Style
|
|
108
|
+
|
|
109
|
+
- **Be direct.** No throat-clearing. No "let me know if you need anything".
|
|
110
|
+
- **One thread at a time.** If five questions are open, surface the highest-
|
|
111
|
+
leverage one and resolve it. The user is patient with depth, not breadth.
|
|
112
|
+
- **Cite files.** When you read code, reference `path:line` so the user can
|
|
113
|
+
jump in. Do not paraphrase code that's already in front of both of you.
|
|
114
|
+
- **Don't spam plan.md.** Edit in place. Each edit should advance the plan,
|
|
115
|
+
not restate it. Keep the file under ~150 lines.
|
|
116
|
+
|
|
117
|
+
## What you do NOT do
|
|
118
|
+
|
|
119
|
+
- ❌ Edit code outside `plan.md` and `goal.json`. The Coder will do that.
|
|
120
|
+
- ❌ Run the evaluator yourself. The Coder runs eval, the Reviewer audits it.
|
|
121
|
+
- ❌ Promise outcomes ("this will get loss to 0.1"). State assumptions and
|
|
122
|
+
gates instead.
|
|
123
|
+
- ❌ Push the user out-of-band. In S2 there is no `notify_user` tool. Speak
|
|
124
|
+
through chat only.
|
|
125
|
+
|
|
126
|
+
## Format
|
|
127
|
+
|
|
128
|
+
Free-form chat is fine. If you need to emit something machine-readable for
|
|
129
|
+
later phases, fence it as JSON in a labeled code block — but in S2 nothing
|
|
130
|
+
parses your output for structured signals, so prefer prose.
|
|
131
|
+
|
|
132
|
+
---
|
|
133
|
+
|
|
134
|
+
**Begin** by reading the workspace (`ls`, `Glob`, key files) and then ask the
|
|
135
|
+
user one focused question to start the design conversation. Do not output
|
|
136
|
+
boilerplate intros.
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# Reviewer — Autoloop
|
|
2
|
+
|
|
3
|
+
You are the **Reviewer**. Your job is to **distrust** the Coder's claims and
|
|
4
|
+
independently verify whether each iteration actually moved toward the goal.
|
|
5
|
+
|
|
6
|
+
## Identity
|
|
7
|
+
|
|
8
|
+
- You are deliberately isolated. Your cwd is a **sandbox** (`ledger/reviewer_sandbox/`)
|
|
9
|
+
that contains only the artifacts the orchestrator hands you for the iter
|
|
10
|
+
under review — not the live workspace, not unrelated history.
|
|
11
|
+
- You persist across iterations. Your accumulating mental model of "how
|
|
12
|
+
Coder cheats / cuts corners" is your most valuable asset. Save what you
|
|
13
|
+
learn into `reviewer_memory.md` after each review.
|
|
14
|
+
- You report only to the runner (which forwards your verdict to Planner).
|
|
15
|
+
You do **not** chat with the user or with the Coder.
|
|
16
|
+
|
|
17
|
+
## Your tools
|
|
18
|
+
|
|
19
|
+
Standard Claude Code palette in the sandbox cwd: Read, Glob, Grep, Bash. You
|
|
20
|
+
generally do **not** Edit/Write the workspace — you can only write inside the
|
|
21
|
+
sandbox (`reviewer_memory.md`, scratch files).
|
|
22
|
+
|
|
23
|
+
Autoloop control:
|
|
24
|
+
|
|
25
|
+
```autoloop
|
|
26
|
+
{"tool": "review_complete", "args": { ... }}
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
| Tool | Args | When |
|
|
30
|
+
|---|---|---|
|
|
31
|
+
| `review_complete` | `decision` ('advance' / 'hold' / 'rollback'), `metric` (number or null), `audit_notes` (string), `flags?` (string[]) | Always emit exactly one of these per turn. |
|
|
32
|
+
| `reviewer_log` | `message` (string) | Append to `<ledger>/reviewer_log.jsonl`. Use for cumulative patterns ("Coder claims metric improved at iter 5 but eval set was unchanged from iter 4"). |
|
|
33
|
+
|
|
34
|
+
## Decision rubric
|
|
35
|
+
|
|
36
|
+
Default toward **hold** under uncertainty. Only `advance` if:
|
|
37
|
+
|
|
38
|
+
1. The metric in `eval_output.json` matches what an independent re-run of
|
|
39
|
+
the eval command would produce (when feasible — re-run if the sandbox
|
|
40
|
+
has the necessary state).
|
|
41
|
+
2. All required gates from `goal.json` pass under your independent check.
|
|
42
|
+
3. No suspicious patterns:
|
|
43
|
+
- eval set / extract_cmd silently changed
|
|
44
|
+
- new flags / env vars introduced that game the eval
|
|
45
|
+
- metric improved but the diff doesn't plausibly cause that improvement
|
|
46
|
+
- Coder's `summary` doesn't match the actual diff
|
|
47
|
+
|
|
48
|
+
`rollback` only when the diff is **net negative** — eval regressed AND the
|
|
49
|
+
change isn't a stepping stone (i.e., Coder didn't flag it as such in the
|
|
50
|
+
directive_ack). Otherwise prefer `hold` so the Planner gets a chance to
|
|
51
|
+
adjust.
|
|
52
|
+
|
|
53
|
+
## Workflow per review
|
|
54
|
+
|
|
55
|
+
1. Read the staged artifacts: `iter/<n>/directive.json`, `diff.patch`,
|
|
56
|
+
`eval_output.json`, the prior iter's `verdict.json` if present.
|
|
57
|
+
2. Re-derive the metric independently if the sandbox has the bits to
|
|
58
|
+
do so. If not, structurally verify (e.g., did the Coder change the
|
|
59
|
+
eval script?).
|
|
60
|
+
3. Check each gate from `goal.json`. For each, write one line to
|
|
61
|
+
`audit_notes` saying "G1 PASS — <reason>" or "G1 FAIL — <reason>".
|
|
62
|
+
4. Update `reviewer_memory.md` with any new pattern you noticed.
|
|
63
|
+
5. Emit `review_complete`.
|
|
64
|
+
|
|
65
|
+
## Hard rules
|
|
66
|
+
|
|
67
|
+
- ❌ **No advance without independent verification.** If you can't verify,
|
|
68
|
+
default to `hold` and explain why.
|
|
69
|
+
- ❌ **Do not modify** anything outside the sandbox cwd.
|
|
70
|
+
- ❌ **Do not** ask Planner / Coder for clarification. You operate from
|
|
71
|
+
artifacts only. If artifacts are missing, that itself is a `hold` with
|
|
72
|
+
a clear note.
|
|
73
|
+
- ✅ **Be terse.** `audit_notes` is read by Planner / surfaced in UI; keep
|
|
74
|
+
it under ~200 words unless something genuinely needs explaining.
|
|
75
|
+
|
|
76
|
+
Begin by reading the iter artifacts in your cwd.
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Coder + Reviewer tool-call parsers.
|
|
3
|
+
*
|
|
4
|
+
* Same fenced-block convention as planner-tools.ts: agents emit one or more
|
|
5
|
+
* ```autoloop
|
|
6
|
+
* {"tool": "...", "args": { ... }}
|
|
7
|
+
* ```
|
|
8
|
+
* blocks per turn. The dispatcher extracts and acts on them.
|
|
9
|
+
*
|
|
10
|
+
* Coder tools: iter_complete, request_clarification, coder_log
|
|
11
|
+
* Reviewer tools: review_complete, reviewer_log
|
|
12
|
+
*/
|
|
13
|
+
export type CoderToolName = 'iter_complete' | 'request_clarification' | 'coder_log';
|
|
14
|
+
export type ReviewerToolName = 'review_complete' | 'reviewer_log';
|
|
15
|
+
export interface AgentToolCall {
|
|
16
|
+
tool: string;
|
|
17
|
+
args: Record<string, unknown>;
|
|
18
|
+
}
|
|
19
|
+
export interface AgentToolParseResult {
|
|
20
|
+
calls: AgentToolCall[];
|
|
21
|
+
cleaned_reply: string;
|
|
22
|
+
parse_errors: Array<{
|
|
23
|
+
block_index: number;
|
|
24
|
+
error: string;
|
|
25
|
+
}>;
|
|
26
|
+
}
|
|
27
|
+
/** Same parser as Planner's; agent-tools just describe a different vocabulary. */
|
|
28
|
+
export declare function parseAgentReply(reply: string): AgentToolParseResult;
|
|
29
|
+
export interface IterCompletePayload {
|
|
30
|
+
summary: string;
|
|
31
|
+
eval_output: unknown;
|
|
32
|
+
files_changed?: string[];
|
|
33
|
+
}
|
|
34
|
+
export interface ReviewCompletePayload {
|
|
35
|
+
decision: 'advance' | 'hold' | 'rollback';
|
|
36
|
+
metric: number | null;
|
|
37
|
+
audit_notes: string;
|
|
38
|
+
flags?: string[];
|
|
39
|
+
}
|
|
40
|
+
/** Find the *last* iter_complete block (per coder prompt: at most one expected). */
|
|
41
|
+
export declare function extractIterComplete(calls: AgentToolCall[]): IterCompletePayload | null;
|
|
42
|
+
export declare function extractReviewComplete(calls: AgentToolCall[]): ReviewCompletePayload | null;
|
|
43
|
+
/** Convenience: find first request_clarification, if any. */
|
|
44
|
+
export declare function extractClarification(calls: AgentToolCall[]): string | null;
|
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Coder + Reviewer tool-call parsers.
|
|
3
|
+
*
|
|
4
|
+
* Same fenced-block convention as planner-tools.ts: agents emit one or more
|
|
5
|
+
* ```autoloop
|
|
6
|
+
* {"tool": "...", "args": { ... }}
|
|
7
|
+
* ```
|
|
8
|
+
* blocks per turn. The dispatcher extracts and acts on them.
|
|
9
|
+
*
|
|
10
|
+
* Coder tools: iter_complete, request_clarification, coder_log
|
|
11
|
+
* Reviewer tools: review_complete, reviewer_log
|
|
12
|
+
*/
|
|
13
|
+
const FENCE_RE = /```autoloop\s*\n([\s\S]*?)\n```/g;
|
|
14
|
+
/** Same parser as Planner's; agent-tools just describe a different vocabulary. */
|
|
15
|
+
export function parseAgentReply(reply) {
|
|
16
|
+
const calls = [];
|
|
17
|
+
const parse_errors = [];
|
|
18
|
+
let blockIndex = 0;
|
|
19
|
+
const cleaned = reply.replace(FENCE_RE, (_match, body) => {
|
|
20
|
+
const idx = blockIndex++;
|
|
21
|
+
try {
|
|
22
|
+
const parsed = JSON.parse(body.trim());
|
|
23
|
+
if (typeof parsed?.tool !== 'string' || typeof parsed?.args !== 'object' || parsed.args === null) {
|
|
24
|
+
parse_errors.push({ block_index: idx, error: 'block missing tool/args fields' });
|
|
25
|
+
return '';
|
|
26
|
+
}
|
|
27
|
+
calls.push(parsed);
|
|
28
|
+
}
|
|
29
|
+
catch (err) {
|
|
30
|
+
parse_errors.push({ block_index: idx, error: err.message });
|
|
31
|
+
}
|
|
32
|
+
return '';
|
|
33
|
+
});
|
|
34
|
+
return { calls, cleaned_reply: cleaned.trim(), parse_errors };
|
|
35
|
+
}
|
|
36
|
+
/** Find the *last* iter_complete block (per coder prompt: at most one expected). */
|
|
37
|
+
export function extractIterComplete(calls) {
|
|
38
|
+
const matches = calls.filter((c) => c.tool === 'iter_complete');
|
|
39
|
+
if (matches.length === 0)
|
|
40
|
+
return null;
|
|
41
|
+
const last = matches[matches.length - 1];
|
|
42
|
+
const summary = String(last.args.summary ?? '');
|
|
43
|
+
const eval_output = last.args.eval_output ?? {};
|
|
44
|
+
const filesRaw = last.args.files_changed;
|
|
45
|
+
const files_changed = Array.isArray(filesRaw) ? filesRaw.filter((x) => typeof x === 'string') : undefined;
|
|
46
|
+
return { summary, eval_output, files_changed };
|
|
47
|
+
}
|
|
48
|
+
export function extractReviewComplete(calls) {
|
|
49
|
+
const matches = calls.filter((c) => c.tool === 'review_complete');
|
|
50
|
+
if (matches.length === 0)
|
|
51
|
+
return null;
|
|
52
|
+
const last = matches[matches.length - 1];
|
|
53
|
+
const dec = String(last.args.decision ?? '');
|
|
54
|
+
if (dec !== 'advance' && dec !== 'hold' && dec !== 'rollback')
|
|
55
|
+
return null;
|
|
56
|
+
const metricRaw = last.args.metric;
|
|
57
|
+
const metric = typeof metricRaw === 'number' && Number.isFinite(metricRaw) ? metricRaw : metricRaw === null ? null : null;
|
|
58
|
+
const audit_notes = String(last.args.audit_notes ?? '');
|
|
59
|
+
const flagsRaw = last.args.flags;
|
|
60
|
+
const flags = Array.isArray(flagsRaw) ? flagsRaw.filter((x) => typeof x === 'string') : undefined;
|
|
61
|
+
return { decision: dec, metric, audit_notes, flags };
|
|
62
|
+
}
|
|
63
|
+
/** Convenience: find first request_clarification, if any. */
|
|
64
|
+
export function extractClarification(calls) {
|
|
65
|
+
const m = calls.find((c) => c.tool === 'request_clarification');
|
|
66
|
+
if (!m)
|
|
67
|
+
return null;
|
|
68
|
+
const q = m.args.question;
|
|
69
|
+
return typeof q === 'string' && q.trim() ? q : null;
|
|
70
|
+
}
|
|
71
|
+
//# sourceMappingURL=agent-tools.js.map
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
{"version":3,"file":"agent-tools.js","sourceRoot":"","sources":["../../../src/autoloop/agent-tools.ts"],"names":[],"mappings":"AAAA;;;;;;;;;;;GAWG;AAgBH,MAAM,QAAQ,GAAG,kCAAkC,CAAC;AAEpD,kFAAkF;AAClF,MAAM,UAAU,eAAe,CAAC,KAAa;IAC3C,MAAM,KAAK,GAAoB,EAAE,CAAC;IAClC,MAAM,YAAY,GAAkD,EAAE,CAAC;IACvE,IAAI,UAAU,GAAG,CAAC,CAAC;IACnB,MAAM,OAAO,GAAG,KAAK,CAAC,OAAO,CAAC,QAAQ,EAAE,CAAC,MAAM,EAAE,IAAY,EAAE,EAAE;QAC/D,MAAM,GAAG,GAAG,UAAU,EAAE,CAAC;QACzB,IAAI,CAAC;YACH,MAAM,MAAM,GAAG,IAAI,CAAC,KAAK,CAAC,IAAI,CAAC,IAAI,EAAE,CAAkB,CAAC;YACxD,IAAI,OAAO,MAAM,EAAE,IAAI,KAAK,QAAQ,IAAI,OAAO,MAAM,EAAE,IAAI,KAAK,QAAQ,IAAI,MAAM,CAAC,IAAI,KAAK,IAAI,EAAE,CAAC;gBACjG,YAAY,CAAC,IAAI,CAAC,EAAE,WAAW,EAAE,GAAG,EAAE,KAAK,EAAE,gCAAgC,EAAE,CAAC,CAAC;gBACjF,OAAO,EAAE,CAAC;YACZ,CAAC;YACD,KAAK,CAAC,IAAI,CAAC,MAAM,CAAC,CAAC;QACrB,CAAC;QAAC,OAAO,GAAG,EAAE,CAAC;YACb,YAAY,CAAC,IAAI,CAAC,EAAE,WAAW,EAAE,GAAG,EAAE,KAAK,EAAG,GAAa,CAAC,OAAO,EAAE,CAAC,CAAC;QACzE,CAAC;QACD,OAAO,EAAE,CAAC;IACZ,CAAC,CAAC,CAAC;IACH,OAAO,EAAE,KAAK,EAAE,aAAa,EAAE,OAAO,CAAC,IAAI,EAAE,EAAE,YAAY,EAAE,CAAC;AAChE,CAAC;AAiBD,oFAAoF;AACpF,MAAM,UAAU,mBAAmB,CAAC,KAAsB;IACxD,MAAM,OAAO,GAAG,KAAK,CAAC,MAAM,CAAC,CAAC,CAAC,EAAE,EAAE,CAAC,CAAC,CAAC,IAAI,KAAK,eAAe,CAAC,CAAC;IAChE,IAAI,OAAO,CAAC,MAAM,KAAK,CAAC;QAAE,OAAO,IAAI,CAAC;IACtC,MAAM,IAAI,GAAG,OAAO,CAAC,OAAO,CAAC,MAAM,GAAG,CAAC,CAAC,CAAC;IACzC,MAAM,OAAO,GAAG,MAAM,CAAC,IAAI,CAAC,IAAI,CAAC,OAAO,IAAI,EAAE,CAAC,CAAC;IAChD,MAAM,WAAW,GAAG,IAAI,CAAC,IAAI,CAAC,WAAW,IAAI,EAAE,CAAC;IAChD,MAAM,QAAQ,GAAG,IAAI,CAAC,IAAI,CAAC,aAAa,CAAC;IACzC,MAAM,aAAa,GAAG,KAAK,CAAC,OAAO,CAAC,QAAQ,CAAC,CAAC,CAAC,CAAC,QAAQ,CAAC,MAAM,CAAC,CAAC,CAAC,EAAE,EAAE,CAAC,OAAO,CAAC,KAAK,QAAQ,CAAC,CAAC,CAAC,CAAC,SAAS,CAAC;IAC1G,OAAO,EAAE,OAAO,EAAE,WAAW,EAAE,aAAa,EAAE,CAAC;AACjD,CAAC;AAED,MAAM,UAAU,qBAAqB,CAAC,KAAsB;IAC1D,MAAM,OAAO,GAAG,KAAK,CAAC,MAAM,CAAC,CAAC,CAAC,EAAE,EAAE,CAAC,CAAC,CAAC,IAAI,KAAK,iBAAiB,CAAC,CAAC;IAClE,IAAI,OAAO,CAAC,MAAM,KAAK,CAAC;QAAE,OAAO,IAAI,CAAC;IACtC,MAAM,IAAI,GAAG,OAAO,CAAC,OAAO,CAAC,MAAM,GAAG,CAAC,CAAC,CAAC;IACzC,MAAM,GAAG,GAAG,MAAM,CAAC,IAAI,CAAC,IAAI,CAAC,QAAQ,IAAI,EAAE,CAAC,CAAC;IAC7C,IAAI,GAAG,KAAK,SAAS,IAAI,GAAG,KAAK,MAAM,IAAI,GAAG,KAAK,UAAU;QAAE,OAAO,IAAI,CAAC;IAC3E,MAAM,SAAS,GAAG,IAAI,CAAC,IAAI,CAAC,MAAM,CAAC;IACnC,MAAM,MAAM,GACV,OAAO,SAAS,KAAK,QAAQ,IAAI,MAAM,CAAC,QAAQ,CAAC,SAAS,CAAC,CAAC,CAAC,CAAC,SAAS,CAAC,CAAC,CAAC,SAAS,KAAK,IAAI,CAAC,CAAC,CAAC,IAAI,CAAC,CAAC,CAAC,IAAI,CAAC;IAC7G,MAAM,WAAW,GAAG,MAAM,CAAC,IAAI,CAAC,IAAI,CAAC,WAAW,IAAI,EAAE,CAAC,CAAC;IACxD,MAAM,QAAQ,GAAG,IAAI,CAAC,IAAI,CAAC,KAAK,CAAC;IACjC,MAAM,KAAK,GAAG,KAAK,CAAC,OAAO,CAAC,QAAQ,CAAC,CAAC,CAAC,CAAC,QAAQ,CAAC,MAAM,CAAC,CAAC,CAAC,EAAE,EAAE,CAAC,OAAO,CAAC,KAAK,QAAQ,CAAC,CAAC,CAAC,CAAC,SAAS,CAAC;IAClG,OAAO,EAAE,QAAQ,EAAE,GAAwC,EAAE,MAAM,EAAE,WAAW,EAAE,KAAK,EAAE,CAAC;AAC5F,CAAC;AAED,6DAA6D;AAC7D,MAAM,UAAU,oBAAoB,CAAC,KAAsB;IACzD,MAAM,CAAC,GAAG,KAAK,CAAC,IAAI,CAAC,CAAC,CAAC,EAAE,EAAE,CAAC,CAAC,CAAC,IAAI,KAAK,uBAAuB,CAAC,CAAC;IAChE,IAAI,CAAC,CAAC;QAAE,OAAO,IAAI,CAAC;IACpB,MAAM,CAAC,GAAG,CAAC,CAAC,IAAI,CAAC,QAAQ,CAAC;IAC1B,OAAO,OAAO,CAAC,KAAK,QAAQ,IAAI,CAAC,CAAC,IAAI,EAAE,CAAC,CAAC,CAAC,CAAC,CAAC,CAAC,CAAC,IAAI,CAAC;AACtD,CAAC"}
|
|
@@ -0,0 +1,137 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* ClaudeAgentDispatcher — wires the v2 runner to real persistent Claude
|
|
3
|
+
* sessions managed by SessionManager.
|
|
4
|
+
*
|
|
5
|
+
* S2 scope: Planner only (chat-mode, no subagents yet). Coder/Reviewer
|
|
6
|
+
* delivery throws — S4 wires them in.
|
|
7
|
+
*
|
|
8
|
+
* Naming convention:
|
|
9
|
+
* autoloop-<run_id>-planner
|
|
10
|
+
* autoloop-<run_id>-coder (S4)
|
|
11
|
+
* autoloop-<run_id>-reviewer (S4)
|
|
12
|
+
*
|
|
13
|
+
* Reply path:
|
|
14
|
+
* When the user chats, we sendMessage(planner, text) and capture the
|
|
15
|
+
* Planner's natural-language reply. The reply is *not* a v2 message —
|
|
16
|
+
* it is emitted as the dispatcher's own 'planner_reply' event so the
|
|
17
|
+
* `autoloop_chat` plugin tool can return it to the user. Structured
|
|
18
|
+
* signals (S3+) will be parsed out of the same reply text and pushed
|
|
19
|
+
* into the runner queue.
|
|
20
|
+
*/
|
|
21
|
+
import { EventEmitter } from 'node:events';
|
|
22
|
+
import type { SessionManager } from '../session-manager.js';
|
|
23
|
+
import type { Logger } from '../logger.js';
|
|
24
|
+
import { type AnyAutoloopMessage } from './messages.js';
|
|
25
|
+
import type { AgentDispatcher, AutoloopState, PushPolicy } from './types.js';
|
|
26
|
+
import { type SpawnSubagentsArgs } from './planner-tools.js';
|
|
27
|
+
export interface ClaudeAgentDispatcherConfig {
|
|
28
|
+
manager: SessionManager;
|
|
29
|
+
runId: string;
|
|
30
|
+
workspace: string;
|
|
31
|
+
/** Override the default Planner system prompt (default loads from configs/autoloop-planner-prompt.md). */
|
|
32
|
+
plannerPromptPath?: string;
|
|
33
|
+
/** Override Coder/Reviewer prompt paths (defaults walk-up to configs/autoloop-{coder,reviewer}-prompt.md). */
|
|
34
|
+
coderPromptPath?: string;
|
|
35
|
+
reviewerPromptPath?: string;
|
|
36
|
+
/** Model alias for Planner (default: 'opus'). */
|
|
37
|
+
plannerModel?: string;
|
|
38
|
+
/** Default Coder model (default: 'sonnet'). Can be overridden per spawn_subagents call. */
|
|
39
|
+
coderModel?: string;
|
|
40
|
+
/** Default Reviewer model (default: 'sonnet'). */
|
|
41
|
+
reviewerModel?: string;
|
|
42
|
+
/** Per-message wall-clock cap. Default 10 min. */
|
|
43
|
+
sendTimeoutMs?: number;
|
|
44
|
+
logger?: Logger;
|
|
45
|
+
/**
|
|
46
|
+
* Auto-compact thresholds (percent of context window). When the agent's
|
|
47
|
+
* `contextPercent` (from getStats) climbs above its threshold after a
|
|
48
|
+
* turn, the dispatcher dispatches `/compact <agent-specific summary>` to
|
|
49
|
+
* that agent. Defaults: Planner 80%, Coder 70%, Reviewer 70%.
|
|
50
|
+
*
|
|
51
|
+
* Per the design doc §7: each agent's context is precious; don't let it
|
|
52
|
+
* silently fill until the API rejects.
|
|
53
|
+
*/
|
|
54
|
+
compactThresholds?: {
|
|
55
|
+
planner?: number;
|
|
56
|
+
coder?: number;
|
|
57
|
+
reviewer?: number;
|
|
58
|
+
};
|
|
59
|
+
/**
|
|
60
|
+
* Push-policy ref that S3's update_push_policy mutates. Caller (SessionManager)
|
|
61
|
+
* passes its own policy object so changes are visible to the runner.
|
|
62
|
+
*/
|
|
63
|
+
pushPolicyRef?: PushPolicy;
|
|
64
|
+
/** Called when Planner emits spawn_subagents. S4 implements; S3 records the intent. */
|
|
65
|
+
onSpawnSubagents?: (args: SpawnSubagentsArgs) => Promise<void>;
|
|
66
|
+
}
|
|
67
|
+
export declare class ClaudeAgentDispatcher extends EventEmitter implements AgentDispatcher {
|
|
68
|
+
readonly config: ClaudeAgentDispatcherConfig;
|
|
69
|
+
private logger;
|
|
70
|
+
private plannerName;
|
|
71
|
+
private coderName;
|
|
72
|
+
private reviewerName;
|
|
73
|
+
private plannerStarted;
|
|
74
|
+
private coderStarted;
|
|
75
|
+
private reviewerStarted;
|
|
76
|
+
private plannerSystemPrompt;
|
|
77
|
+
private coderSystemPrompt;
|
|
78
|
+
private reviewerSystemPrompt;
|
|
79
|
+
private coderModel;
|
|
80
|
+
private reviewerModel;
|
|
81
|
+
/** Where Reviewer reads from. Created lazily by stageReviewSandbox(). */
|
|
82
|
+
private reviewerSandboxDir;
|
|
83
|
+
private ledgerDir;
|
|
84
|
+
constructor(config: ClaudeAgentDispatcherConfig);
|
|
85
|
+
get sessionNames(): {
|
|
86
|
+
planner: string;
|
|
87
|
+
coder: string;
|
|
88
|
+
reviewer: string;
|
|
89
|
+
};
|
|
90
|
+
init(state: AutoloopState): Promise<void>;
|
|
91
|
+
shutdown(reason: string): Promise<void>;
|
|
92
|
+
deliver(env: AnyAutoloopMessage): Promise<AnyAutoloopMessage[]>;
|
|
93
|
+
/**
|
|
94
|
+
* Start Coder + Reviewer sessions. Idempotent. Called in response to a
|
|
95
|
+
* Planner spawn_subagents tool (the SessionManager wires this via
|
|
96
|
+
* onSpawnSubagents).
|
|
97
|
+
*/
|
|
98
|
+
spawnSubagents(args?: SpawnSubagentsArgs): Promise<void>;
|
|
99
|
+
/**
|
|
100
|
+
* Reset a single subagent — stop its session, clear the started flag, and
|
|
101
|
+
* (optionally) eagerly start a fresh one. The session-level system prompt is
|
|
102
|
+
* the same; persistent state lives in `<ledger>/{coder,reviewer}_memory.md`
|
|
103
|
+
* which the agent reads on its first turn after reset.
|
|
104
|
+
*
|
|
105
|
+
* Refuses to reset Planner without `force: true` — Planner reset throws away
|
|
106
|
+
* the user-conversation context and must be a deliberate action.
|
|
107
|
+
*/
|
|
108
|
+
resetAgent(agent: 'planner' | 'coder' | 'reviewer', opts?: {
|
|
109
|
+
force?: boolean;
|
|
110
|
+
eagerRestart?: boolean;
|
|
111
|
+
}): Promise<void>;
|
|
112
|
+
/**
|
|
113
|
+
* Wrap a subagent send. If the underlying session throws or returns an
|
|
114
|
+
* error string, auto-reset the subagent once and retry. Used by
|
|
115
|
+
* deliverToCoder / deliverToReviewer to recover from subprocess deaths.
|
|
116
|
+
*/
|
|
117
|
+
private sendWithRecovery;
|
|
118
|
+
private lastCompactAt;
|
|
119
|
+
private compactSummaryFor;
|
|
120
|
+
private maybeCompact;
|
|
121
|
+
private ensurePlanner;
|
|
122
|
+
private deliverToPlanner;
|
|
123
|
+
private ensureCoder;
|
|
124
|
+
private deliverToCoder;
|
|
125
|
+
private ensureReviewer;
|
|
126
|
+
/**
|
|
127
|
+
* Stage the iter's artifacts into the Reviewer sandbox cwd. Reviewer is a
|
|
128
|
+
* persistent session whose cwd is fixed at <ledger>/reviewer_sandbox/, so
|
|
129
|
+
* every review must rewrite the sandbox to "this iter's view".
|
|
130
|
+
*/
|
|
131
|
+
private stageReviewSandbox;
|
|
132
|
+
private deliverToReviewer;
|
|
133
|
+
private persistVerdict;
|
|
134
|
+
/** Run a git command in the workspace; returns combined output. Used by Coder commits. */
|
|
135
|
+
private runGit;
|
|
136
|
+
private gitCommit;
|
|
137
|
+
}
|