@enderfga/claw-orchestrator 4.0.0 → 4.0.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +39 -259
- package/configs/autoloop-coder-prompt.md +65 -28
- package/configs/autoloop-planner-prompt.md +103 -42
- package/configs/autoloop-reviewer-prompt.md +58 -30
- package/dist/src/autoloop/dispatcher.d.ts +10 -1
- package/dist/src/autoloop/dispatcher.js +42 -5
- package/dist/src/autoloop/dispatcher.js.map +1 -1
- package/dist/src/autoloop/planner-tools.d.ts +8 -3
- package/dist/src/autoloop/planner-tools.js +21 -4
- package/dist/src/autoloop/planner-tools.js.map +1 -1
- package/dist/src/dashboard/index.html +410 -6
- package/dist/src/embedded-server.d.ts +7 -0
- package/dist/src/embedded-server.js +282 -26
- package/dist/src/embedded-server.js.map +1 -1
- package/dist/src/session-manager.d.ts +37 -1
- package/dist/src/session-manager.js +179 -6
- package/dist/src/session-manager.js.map +1 -1
- package/dist/src/ultraapp/manager.d.ts +9 -0
- package/dist/src/ultraapp/manager.js +31 -8
- package/dist/src/ultraapp/manager.js.map +1 -1
- package/dist/src/ultraapp/store.js +11 -1
- package/dist/src/ultraapp/store.js.map +1 -1
- package/package.json +1 -1
- package/skills/references/autoloop.md +10 -5
- package/skills/references/dashboard.md +172 -0
- package/skills/references/ultraapp.md +2 -2
package/README.md
CHANGED
|
@@ -4,317 +4,97 @@
|
|
|
4
4
|
|
|
5
5
|
# Claw Orchestrator
|
|
6
6
|
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
Claw Orchestrator turns interactive coding CLIs into programmable, headless agent engines. Start persistent sessions, route tasks across different coding agents, coordinate multi-agent councils, and expose everything through a clean tool-based API.
|
|
10
|
-
|
|
11
|
-
It's a TypeScript runtime for orchestrating Claude Code, OpenAI Codex, Gemini, Cursor Agent, and custom coding CLIs as persistent, programmable coding agents.
|
|
12
|
-
|
|
13
|
-
> Claude Code, Codex, Gemini, Cursor Agent, OpenCode, or your own custom CLI — orchestrated as one runtime.
|
|
14
|
-
>
|
|
15
|
-
> **Runs standalone, as an OpenClaw plugin, or as a Model Context Protocol (MCP) server — drop it into Hermes Agent, Claude Desktop, Cursor, Cline, Continue, Zed, Windsurf, Goose, or any MCP-compatible host.**
|
|
7
|
+
> A runtime for coding agents. Wrap Claude Code, Codex, Gemini, Cursor Agent, OpenCode, or any custom CLI as persistent programmable sessions; coordinate them in multi-agent councils; run autonomous Planner / Coder / Reviewer loops; or hand a five-question interview to an Opus council that ships a deployed web app at `localhost:19000/forge/<slug>/`.
|
|
16
8
|
|
|
17
9
|
[](https://www.npmjs.com/package/@enderfga/claw-orchestrator)
|
|
18
10
|
[](https://github.com/Enderfga/claw-orchestrator/actions/workflows/ci.yml)
|
|
19
11
|
[](https://opensource.org/licenses/MIT)
|
|
20
12
|
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
## Why Claw Orchestrator?
|
|
24
|
-
|
|
25
|
-
Coding agents are powerful, but most are still designed as interactive CLIs.
|
|
26
|
-
|
|
27
|
-
That works well when a human is sitting in front of a terminal. It breaks down when you want agents to:
|
|
28
|
-
|
|
29
|
-
- keep long-running coding sessions alive
|
|
30
|
-
- switch between Claude Code, Codex, Gemini, Cursor Agent, OpenCode, or custom CLIs
|
|
31
|
-
- collaborate as a team on the same codebase
|
|
32
|
-
- integrate coding capabilities into OpenClaw first, and other claw-style agent systems over time
|
|
33
|
-
- manage context, tools, worktrees, and execution state programmatically
|
|
34
|
-
|
|
35
|
-
Claw Orchestrator is the control layer for that.
|
|
13
|
+
Coding CLIs are designed for humans at terminals. Claw Orchestrator turns them into headless engines and stacks an agent platform on top: a 55-tool API that scales from a single session call up to a fully generated, deployed web app — reachable through the CLI, the OpenClaw gateway, the Model Context Protocol, or directly from TypeScript, and visible through an embedded three-tab dashboard.
|
|
36
14
|
|
|
37
15
|
---
|
|
38
16
|
|
|
39
|
-
##
|
|
17
|
+
## Features
|
|
40
18
|
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
-
|
|
45
|
-
-
|
|
46
|
-
-
|
|
47
|
-
|
|
48
|
-
|
|
19
|
+
| Capability | What it does | Reference |
|
|
20
|
+
| --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------- |
|
|
21
|
+
| **Persistent Sessions** | Long-lived coding agents kept alive across requests, with full context, tool, model, and worktree control. | [`sessions.md`](./skills/references/sessions.md) |
|
|
22
|
+
| **Multi-Engine Runtime** | One interface over Claude Code, Codex, Gemini, Cursor Agent, OpenCode, and arbitrary custom CLIs. | [`multi-engine.md`](./skills/references/multi-engine.md) |
|
|
23
|
+
| **Multi-Agent Council** | Parallel agents in isolated git worktrees, voting on consensus until they agree. | [`council.md`](./skills/references/council.md) |
|
|
24
|
+
| **Autoloop** | Three-agent autonomous workspace iteration. Chat with the Planner; it spawns Coder + Reviewer into a self-iterating subloop and pushes you on regression, target-hit, or decision points. | [`autoloop.md`](./skills/references/autoloop.md) |
|
|
25
|
+
| **Ultraapp** | A three-agent Opus council turns a five-question interview into a deployed web app — Tailwind UI, BYOK, file-queue runtime, smoke test, all live at `localhost:19000/forge/<slug>/`. | [`ultraapp.md`](./skills/references/ultraapp.md) |
|
|
26
|
+
| **Embedded Dashboard** | Three-tab UI for Autoloop, Council, and Forge with sidebar lifecycle controls, per-run live event streaming, and cookie-based auth via a `/login` redirect. | [`dashboard.md`](./skills/references/dashboard.md) |
|
|
27
|
+
| **OpenAI-Compatible Proxy** | `POST /v1/chat/completions` translates OpenAI requests into native Anthropic, OpenAI, and Google calls and streams responses back in OpenAI shape. Point any OpenAI-SDK client at the orchestrator without changing call sites. | [`openai-compat.md`](./skills/references/openai-compat.md) |
|
|
49
28
|
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
### Persistent Sessions
|
|
53
|
-
|
|
54
|
-
Keep coding agents alive across requests.
|
|
55
|
-
|
|
56
|
-
```ts
|
|
57
|
-
const session = await manager.startSession({
|
|
58
|
-
name: "fix-tests",
|
|
59
|
-
engine: "claude",
|
|
60
|
-
cwd: "/path/to/project",
|
|
61
|
-
});
|
|
62
|
-
|
|
63
|
-
await manager.sendMessage("fix-tests", "Fix the failing tests");
|
|
64
|
-
```
|
|
65
|
-
|
|
66
|
-
### Multi-Engine Runtime
|
|
67
|
-
|
|
68
|
-
Drive different coding agents through one unified interface.
|
|
69
|
-
|
|
70
|
-
```ts
|
|
71
|
-
await manager.startSession({ name: "claude-task", engine: "claude" });
|
|
72
|
-
await manager.startSession({ name: "codex-task", engine: "codex" });
|
|
73
|
-
await manager.startSession({ name: "gemini-task", engine: "gemini" });
|
|
74
|
-
await manager.startSession({ name: "cursor-task", engine: "cursor" });
|
|
75
|
-
```
|
|
76
|
-
|
|
77
|
-
### Multi-Agent Council
|
|
78
|
-
|
|
79
|
-
Run multiple agents in parallel with isolated git worktrees, independent reasoning, and review-based collaboration.
|
|
80
|
-
|
|
81
|
-
```ts
|
|
82
|
-
await manager.councilStart("Design and implement an auth system", {
|
|
83
|
-
agents: [
|
|
84
|
-
{ name: "Planner", engine: "claude" },
|
|
85
|
-
{ name: "Builder", engine: "codex" },
|
|
86
|
-
{ name: "Reviewer", engine: "claude" },
|
|
87
|
-
],
|
|
88
|
-
});
|
|
89
|
-
```
|
|
90
|
-
|
|
91
|
-
### Autoloop (three-agent autonomous workspace iteration)
|
|
92
|
-
|
|
93
|
-
You converse with a long-lived **Planner** (Opus) to design `plan.md` and `goal.json`; on your "go" the Planner spawns a **Coder** (Sonnet) and a **Reviewer** (Sonnet, sandboxed cwd) into a self-iterating subloop. Coder applies changes + runs eval, Reviewer audits independently, ledger writes after every iter. The Planner pushes you (wechat → whatsapp → email fallback chain) on regression, target-hit, decision points, or a 30-min stall — silent otherwise.
|
|
94
|
-
|
|
95
|
-
```ts
|
|
96
|
-
await manager.autoloopStart({ runId: "my-run", workspace: "/path/to/repo" });
|
|
97
|
-
await manager.autoloopChat("my-run", "Read the workspace and design a plan to fix X");
|
|
98
|
-
// Planner reads files, drafts plan.md/goal.json, asks "ready to spawn?"
|
|
99
|
-
await manager.autoloopChat("my-run", "go");
|
|
100
|
-
// → Planner spawns Coder + Reviewer, subloop runs to target / max_iters / your terminate.
|
|
101
|
-
```
|
|
102
|
-
|
|
103
|
-
SSE stream at `GET /autoloop/<id>/events` (the upcoming 3-pane UI subscribes here). See [`skills/references/autoloop.md`](./skills/references/autoloop.md) for the full operator reference: tool list, push policy, ledger layout, smoke test.
|
|
104
|
-
|
|
105
|
-
### ultraapp (Forge tab)
|
|
106
|
-
|
|
107
|
-
A 3-agent Opus council turns a 5-question interview into a deployed web app — Tailwind UI, BYOK, file-queue runtime, smoke test, all live at `localhost:19000/forge/<slug>/`. Iterate via chat: cosmetic ("make button green") runs an Opus patcher; spec-delta ("also output a thumbnail") flips back to a focused interview and rebuilds. See [`skills/references/ultraapp.md`](./skills/references/ultraapp.md).
|
|
108
|
-
|
|
109
|
-
### Tool Orchestration
|
|
110
|
-
|
|
111
|
-
Expose coding sessions as tools so other agents and systems can control them. The runtime registers 55 tools, including:
|
|
112
|
-
|
|
113
|
-
```txt
|
|
114
|
-
session_start session_send coding_session_status
|
|
115
|
-
session_grep session_compact session_inbox
|
|
116
|
-
team_send team_list coding_agents_list
|
|
117
|
-
council_start council_review council_accept
|
|
118
|
-
ultraplan_start ultrareview_start
|
|
119
|
-
autoloop_start autoloop_chat autoloop_reset_agent
|
|
120
|
-
ultraapp_new ultraapp_answer ultraapp_build_start
|
|
121
|
-
ultraapp_feedback ultraapp_promote_version
|
|
122
|
-
```
|
|
29
|
+
The full 55-tool surface is enumerated in [`tools.md`](./skills/references/tools.md).
|
|
123
30
|
|
|
124
31
|
---
|
|
125
32
|
|
|
126
33
|
## Quick Start
|
|
127
34
|
|
|
128
|
-
### Standalone (no OpenClaw)
|
|
129
|
-
|
|
130
35
|
```bash
|
|
131
36
|
npm install -g @enderfga/claw-orchestrator
|
|
132
|
-
clawo serve
|
|
133
|
-
```
|
|
134
|
-
|
|
135
|
-
```bash
|
|
136
|
-
clawo session-start fix-tests --engine claude --cwd .
|
|
137
|
-
clawo session-send fix-tests "Fix the failing tests"
|
|
37
|
+
clawo serve # dashboard at http://127.0.0.1:18796/dash
|
|
138
38
|
```
|
|
139
39
|
|
|
140
|
-
### Programmatic
|
|
141
|
-
|
|
142
40
|
```ts
|
|
143
|
-
import { SessionManager } from
|
|
41
|
+
import { SessionManager } from '@enderfga/claw-orchestrator';
|
|
144
42
|
|
|
145
43
|
const manager = new SessionManager();
|
|
146
|
-
await manager.startSession({ name:
|
|
147
|
-
const result = await manager.sendMessage(
|
|
44
|
+
await manager.startSession({ name: 'fix-tests', engine: 'claude', cwd: '/project' });
|
|
45
|
+
const result = await manager.sendMessage('fix-tests', 'Fix the failing tests');
|
|
148
46
|
```
|
|
149
47
|
|
|
150
|
-
|
|
48
|
+
---
|
|
151
49
|
|
|
152
|
-
|
|
153
|
-
clawo council start "Refactor the API layer and add tests"
|
|
154
|
-
```
|
|
50
|
+
## Integrations
|
|
155
51
|
|
|
156
|
-
###
|
|
52
|
+
### Standalone CLI
|
|
157
53
|
|
|
158
54
|
```bash
|
|
159
|
-
clawo serve
|
|
55
|
+
clawo serve # dashboard + HTTP server on :18796
|
|
56
|
+
clawo session-start fix-tests --engine claude --cwd . # start a session
|
|
57
|
+
clawo session-send fix-tests "Fix the failing tests" # send into it
|
|
160
58
|
```
|
|
161
59
|
|
|
162
|
-
|
|
60
|
+
Every command is documented in [`cli.md`](./skills/references/cli.md).
|
|
163
61
|
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
### As an OpenClaw plugin
|
|
167
|
-
|
|
168
|
-
If you run OpenClaw, Claw Orchestrator installs as a managed plugin. The same tools (`session_start`, `team_send`, `council_start`, ...) become available to every OpenClaw agent.
|
|
62
|
+
### OpenClaw Plugin
|
|
169
63
|
|
|
170
64
|
```bash
|
|
171
65
|
curl -fsSL https://raw.githubusercontent.com/Enderfga/claw-orchestrator/main/install.sh | bash
|
|
172
66
|
```
|
|
173
67
|
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
### As an MCP server (Hermes Agent, Claude Desktop, Cursor, Cline, Continue, Zed, Windsurf, Goose)
|
|
68
|
+
Installs via npm, registers the plugin in `~/.openclaw/openclaw.json`, restarts the gateway. All 55 tools become available to every OpenClaw agent.
|
|
177
69
|
|
|
178
|
-
|
|
70
|
+
### Model Context Protocol Server
|
|
179
71
|
|
|
180
72
|
```bash
|
|
181
|
-
npm install -g @enderfga/claw-orchestrator
|
|
182
|
-
# now `clawo-mcp` is on PATH
|
|
73
|
+
npm install -g @enderfga/claw-orchestrator # clawo-mcp is now on PATH
|
|
183
74
|
```
|
|
184
75
|
|
|
185
|
-
Register
|
|
186
|
-
|
|
187
|
-
<details>
|
|
188
|
-
<summary><strong>Hermes Agent</strong> — <code>~/.hermes/config.yaml</code></summary>
|
|
189
|
-
|
|
190
|
-
```yaml
|
|
191
|
-
mcp_servers:
|
|
192
|
-
clawo:
|
|
193
|
-
command: clawo-mcp
|
|
194
|
-
env:
|
|
195
|
-
ANTHROPIC_API_KEY: "..."
|
|
196
|
-
OPENAI_API_KEY: "..."
|
|
197
|
-
GEMINI_API_KEY: "..."
|
|
198
|
-
# Optional: keep the model focused — Hermes will surface only these
|
|
199
|
-
tools:
|
|
200
|
-
include: [mcp_clawo_session_start, mcp_clawo_session_send, mcp_clawo_council_start]
|
|
201
|
-
```
|
|
202
|
-
</details>
|
|
203
|
-
|
|
204
|
-
<details>
|
|
205
|
-
<summary><strong>Claude Desktop / Claude Code</strong> — <code>claude_desktop_config.json</code></summary>
|
|
206
|
-
|
|
207
|
-
```json
|
|
208
|
-
{
|
|
209
|
-
"mcpServers": {
|
|
210
|
-
"clawo": {
|
|
211
|
-
"command": "clawo-mcp",
|
|
212
|
-
"env": {
|
|
213
|
-
"ANTHROPIC_API_KEY": "...",
|
|
214
|
-
"OPENAI_API_KEY": "...",
|
|
215
|
-
"CLAWO_MCP_TOOLS": "session_start,session_send,council_start,council_status,ultrareview_start"
|
|
216
|
-
}
|
|
217
|
-
}
|
|
218
|
-
}
|
|
219
|
-
}
|
|
220
|
-
```
|
|
221
|
-
</details>
|
|
222
|
-
|
|
223
|
-
<details>
|
|
224
|
-
<summary><strong>Cursor / Cline / Continue / Zed / Windsurf / Goose</strong></summary>
|
|
225
|
-
|
|
226
|
-
Every one of these hosts speaks the standard MCP stdio config. Add `clawo-mcp` as a server with the same `command + args + env` shape — refer to your host's MCP documentation for the exact file path.
|
|
227
|
-
</details>
|
|
228
|
-
|
|
229
|
-
**Notes**
|
|
230
|
-
|
|
231
|
-
- Hosts prefix MCP tool names with the server slug (e.g. `mcp_clawo_session_start` in Hermes). The model sees the prefixed name; you don't need to call it manually.
|
|
232
|
-
- `CLAWO_MCP_TOOLS` (comma-separated allowlist) keeps the exposed surface tight when the host has a tight tool budget. Without it, all 55 tools are advertised.
|
|
233
|
-
- Hosts do not forward arbitrary shell env vars to MCP servers — list every API key your engines need (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GEMINI_API_KEY`, etc.) explicitly under `env`.
|
|
234
|
-
|
|
235
|
-
Full reference: [`skills/references/mcp.md`](./skills/references/mcp.md).
|
|
76
|
+
Register `clawo-mcp` with any MCP-compatible host: Hermes Agent, Claude Desktop, Cursor, Cline, Continue, Zed, Windsurf, Goose, and others. Per-host stdio-config snippets and the `CLAWO_MCP_TOOLS` allowlist for tight tool budgets are in [`mcp.md`](./skills/references/mcp.md).
|
|
236
77
|
|
|
237
78
|
---
|
|
238
79
|
|
|
239
80
|
## Engine Compatibility
|
|
240
81
|
|
|
241
|
-
| Engine
|
|
242
|
-
|
|
243
|
-
| Claude Code
|
|
244
|
-
| Codex
|
|
245
|
-
| Gemini
|
|
246
|
-
| Cursor Agent
|
|
247
|
-
| OpenCode
|
|
248
|
-
| Custom CLI
|
|
82
|
+
| Engine | CLI | Tested Version |
|
|
83
|
+
| ------------ | ---------- | -------------- |
|
|
84
|
+
| Claude Code | `claude` | 2.1.126 |
|
|
85
|
+
| Codex | `codex` | 0.128.0 |
|
|
86
|
+
| Gemini | `gemini` | 0.36.0 |
|
|
87
|
+
| Cursor Agent | `agent` | 2026.03.30 |
|
|
88
|
+
| OpenCode | `opencode` | 1.1.40 |
|
|
89
|
+
| Custom CLI | any | — |
|
|
249
90
|
|
|
250
|
-
Any coding CLI that
|
|
251
|
-
|
|
252
|
-
---
|
|
253
|
-
|
|
254
|
-
## Architecture
|
|
255
|
-
|
|
256
|
-
```txt
|
|
257
|
-
┌─────────────────────┐
|
|
258
|
-
│ Claw Orchestrator │
|
|
259
|
-
└──────────┬──────────┘
|
|
260
|
-
│
|
|
261
|
-
┌───────────────────┼───────────────────┐
|
|
262
|
-
│ │ │
|
|
263
|
-
┌──────▼──────┐ ┌──────▼──────┐ ┌──────▼──────┐
|
|
264
|
-
│ Claude Code │ │ Codex │ │ Custom CLI │
|
|
265
|
-
└──────┬──────┘ └──────┬──────┘ └──────┬──────┘
|
|
266
|
-
│ │ │
|
|
267
|
-
└───────────┬───────┴───────────┬───────┘
|
|
268
|
-
│ │
|
|
269
|
-
Persistent Sessions Tool API
|
|
270
|
-
│ │
|
|
271
|
-
└──── Multi-Agent Council
|
|
272
|
-
```
|
|
273
|
-
|
|
274
|
-
For source-level architecture, see [`CLAUDE.md`](./CLAUDE.md). For deeper reference docs, see [`skills/references/`](./skills/references/).
|
|
275
|
-
|
|
276
|
-
---
|
|
277
|
-
|
|
278
|
-
## Migrating from v2.x
|
|
279
|
-
|
|
280
|
-
v3.x uses the Claw Orchestrator package, `clawo` CLI, and engine-neutral tool API. The v3.0 compatibility aliases were removed in v3.1.0.
|
|
281
|
-
|
|
282
|
-
| What | v2.x | Current |
|
|
283
|
-
|---|---|---|
|
|
284
|
-
| npm package | `@enderfga/openclaw-claude-code` | `@enderfga/claw-orchestrator` |
|
|
285
|
-
| CLI binary | `claude-code-skill` | `clawo` |
|
|
286
|
-
| Tool names | `claude_session_start`, `claude_session_send`, ... | `session_start`, `session_send`, ... |
|
|
287
|
-
| OpenClaw plugin id | `openclaw-claude-code` | `claw-orchestrator` |
|
|
288
|
-
|
|
289
|
-
To upgrade:
|
|
290
|
-
|
|
291
|
-
```bash
|
|
292
|
-
npm uninstall -g @enderfga/openclaw-claude-code
|
|
293
|
-
npm install -g @enderfga/claw-orchestrator
|
|
294
|
-
curl -fsSL https://raw.githubusercontent.com/Enderfga/claw-orchestrator/main/install.sh | bash
|
|
295
|
-
```
|
|
296
|
-
|
|
297
|
-
If your OpenClaw config still has an old plugin entry, remove it and register `claw-orchestrator`. Update scripts and tool callers before moving to v3.1.0 or newer.
|
|
298
|
-
|
|
299
|
-
---
|
|
300
|
-
|
|
301
|
-
## Project Status
|
|
302
|
-
|
|
303
|
-
Active development. Current focus areas:
|
|
304
|
-
|
|
305
|
-
- stable multi-engine session management
|
|
306
|
-
- richer council workflows
|
|
307
|
-
- custom engine configuration ergonomics
|
|
308
|
-
- runtime control APIs
|
|
309
|
-
- cleaner CLI and OpenClaw integration
|
|
91
|
+
Any coding CLI that runs as a subprocess can be wired up as a custom engine — see [`multi-engine.md`](./skills/references/multi-engine.md#custom-engine-enginecustom).
|
|
310
92
|
|
|
311
93
|
---
|
|
312
94
|
|
|
313
95
|
## Contributing
|
|
314
96
|
|
|
315
|
-
See [`CONTRIBUTING.md`](./CONTRIBUTING.md).
|
|
316
|
-
|
|
317
|
-
---
|
|
97
|
+
See [`CONTRIBUTING.md`](./CONTRIBUTING.md). Run `npm run build && npm run lint && npm run format:check && npm run test` before submitting.
|
|
318
98
|
|
|
319
99
|
## License
|
|
320
100
|
|
|
@@ -4,56 +4,93 @@ You are the **Coder** in a three-agent autoloop. You make code changes
|
|
|
4
4
|
toward the goal stated in `plan.md` and `goal.json`. You do **not** talk to
|
|
5
5
|
the user; the Planner is your only interlocutor.
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
---
|
|
8
8
|
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
9
|
+
## ABSOLUTE RULES (read before doing anything)
|
|
10
|
+
|
|
11
|
+
These rules are non-negotiable. They exist because past iterations of this
|
|
12
|
+
system broke when Coders bent them.
|
|
13
|
+
|
|
14
|
+
### Rule 1 — Stay inside your scope
|
|
15
|
+
|
|
16
|
+
The Planner owns `plan.md`, `goal.json`, and anything under `tasks/`. You
|
|
17
|
+
must never modify those files. They are the contract you work against;
|
|
18
|
+
rewriting them is cheating.
|
|
19
|
+
|
|
20
|
+
If you believe the plan is wrong, do not "fix" it. Emit
|
|
21
|
+
`request_clarification` and let the Planner decide.
|
|
22
|
+
|
|
23
|
+
### Rule 2 — Do not commit, do not push
|
|
24
|
+
|
|
25
|
+
The orchestrator git-commits your work after every iteration. Manual
|
|
26
|
+
`git commit` or `git push` from your end pollutes the diff log and
|
|
27
|
+
breaks the Reviewer's ability to see exactly what changed this iter.
|
|
28
|
+
|
|
29
|
+
If you find yourself reaching for `git commit`, stop. Your work is
|
|
30
|
+
captured by the orchestrator.
|
|
31
|
+
|
|
32
|
+
### Rule 3 — Never skip the evaluator
|
|
33
|
+
|
|
34
|
+
If the eval is broken or you can't reach it, emit `request_clarification`
|
|
35
|
+
with a precise description. Do **not** invent a metric value. Do **not**
|
|
36
|
+
report `iter_complete` without a real `eval_output`. The Reviewer will
|
|
37
|
+
catch fabrication, and a `rollback` verdict erases the iter.
|
|
38
|
+
|
|
39
|
+
### Rule 4 — One focused change per iter
|
|
40
|
+
|
|
41
|
+
If a directive seems to need touching more than ~5 files or unrelated
|
|
42
|
+
subsystems, stop and emit `request_clarification`. Multi-concern iters
|
|
43
|
+
make Reviewer audits unreliable.
|
|
44
|
+
|
|
45
|
+
---
|
|
16
46
|
|
|
17
47
|
## Your tools
|
|
18
48
|
|
|
19
49
|
You are a Claude Code session with the workspace as cwd. You have the full
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
50
|
+
file-editing palette: Read, Write, Edit, Glob, Grep, Bash. Use them freely
|
|
51
|
+
on workspace code — that's your job. The role boundary is Rule 1: do not
|
|
52
|
+
touch `plan.md`, `goal.json`, or `tasks/`.
|
|
23
53
|
|
|
24
54
|
You also have **autoloop control tools** via fenced JSON blocks:
|
|
25
55
|
|
|
56
|
+
````
|
|
26
57
|
```autoloop
|
|
27
58
|
{"tool": "iter_complete", "args": { ... }}
|
|
28
59
|
```
|
|
60
|
+
````
|
|
29
61
|
|
|
30
62
|
| Tool | Args | When to use |
|
|
31
63
|
|---|---|---|
|
|
32
|
-
| `iter_complete` | `summary` (one-line), `eval_output` (object — usually `{ metric: number, gates: [...], extra: {...} }`), `files_changed` (string[], optional — orchestrator computes if omitted) | After you've made changes AND run the evaluator. This signals the iteration is done. |
|
|
33
|
-
| `request_clarification` | `question` (string) | If the directive is too ambiguous to act on. Planner gets this back and replies. Use sparingly — prefer to ship best-guess and let Reviewer flag. |
|
|
64
|
+
| `iter_complete` | `summary` (one-line), `eval_output` (object — usually `{ metric: number, gates: [...], extra: {...} }`), `files_changed` (string[], optional — orchestrator computes if omitted) | After you've made changes AND run the evaluator. This signals the iteration is done. **At most one per turn.** |
|
|
65
|
+
| `request_clarification` | `question` (string) | If the directive is too ambiguous to act on, or if Rule 1/3/4 trip. Planner gets this back and replies. Use sparingly — prefer to ship best-guess and let Reviewer flag, unless ambiguity is load-bearing. |
|
|
34
66
|
| `coder_log` | `message` (string) | Free-form log entry appended to `<ledger>/coder_log.jsonl`. Use for "I tried X and it failed, here's why" so future iters don't repeat. |
|
|
35
67
|
|
|
68
|
+
---
|
|
69
|
+
|
|
36
70
|
## Workflow per iteration
|
|
37
71
|
|
|
38
72
|
1. **Read the directive.** It is provided as the user-message in this turn.
|
|
39
|
-
2. **Read context** — `plan.md`, `goal.json`, last iter's
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
73
|
+
2. **Read context** — `plan.md`, `goal.json`, last iter's
|
|
74
|
+
`iter/<n-1>/verdict.json` if present, `coder_notes.md`.
|
|
75
|
+
3. **Make the change.** One focused change per iter (Rule 4). Avoid
|
|
76
|
+
bundling unrelated cleanup.
|
|
77
|
+
4. **Run the evaluator** as specified by `goal.json`'s `scalar.extract_cmd`
|
|
78
|
+
(and any per-gate eval) using Bash.
|
|
79
|
+
5. **Capture eval output** structured. Pull the metric value out of stdout
|
|
80
|
+
per `goal.json`'s `extract_pattern` if present.
|
|
43
81
|
6. **Emit `iter_complete`** with the metric + per-gate pass/fail + any extras.
|
|
44
82
|
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
- ❌ **Do not modify** `plan.md`, `goal.json`, or anything under `tasks/`. Planner owns those.
|
|
48
|
-
- ❌ **Do not** manually run `git commit` or `git push`. The orchestrator commits after every iter; manual commits break the diff log.
|
|
49
|
-
- ❌ **Do not skip the evaluator.** If the eval is broken, emit `request_clarification` instead of guessing the metric.
|
|
50
|
-
- ❌ **Do not over-edit.** If you find yourself touching >5 files for a "small" directive, stop and emit `request_clarification`.
|
|
51
|
-
- ✅ **Do leave a note** for things you discover that future iters need (`coder_notes.md`). Future-you will thank you.
|
|
83
|
+
---
|
|
52
84
|
|
|
53
|
-
##
|
|
85
|
+
## Style
|
|
54
86
|
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
- **
|
|
87
|
+
- **No banners, no greetings, no apologies.** Concise narration of what
|
|
88
|
+
you tried.
|
|
89
|
+
- **Cite files at `path:line`** so Reviewer / Planner can verify.
|
|
90
|
+
- **Leave notes for your future self.** Append to `coder_notes.md` when
|
|
91
|
+
you discover something non-obvious (Rule 1 does not apply to
|
|
92
|
+
`coder_notes.md` — it's yours).
|
|
93
|
+
- **Output discipline.** Prose outside the autoloop fence is shown
|
|
94
|
+
upstream verbatim. At most one `iter_complete` block per turn.
|
|
58
95
|
|
|
59
96
|
Begin by reading the directive and acting.
|