headlesscode 1.2.0 → 1.2.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -5,520 +5,306 @@
5
5
  [![Node.js >=18](https://img.shields.io/badge/node-%3E%3D18-339933?logo=node.js&logoColor=white)](https://nodejs.org/)
6
6
  [![License: Apache-2.0](https://img.shields.io/badge/License-Apache--2.0-blue.svg)](./LICENSE)
7
7
 
8
- Run a coding agent where the work actually happens: in a terminal, a CI job,
9
- or an isolated git worktree. `headlesscode` brings the useful parts of
10
- [Zoo Code](https://github.com/Zoo-Code-Org/Zoo-Code) to a plain Node.js process,
11
- without requiring a VS Code window.
12
-
13
- It is made for real repositories and real engineering loops. The harness reads
14
- the target project's modes and rules, gives the model a controlled tool set,
15
- keeps work separated in git worktrees, and leaves behind logs and state that a
16
- person can inspect.
17
-
18
- ## Why use it?
19
-
20
- - **Run repeatable repository tasks.** Give it a task or GitHub issue and let
21
- it work through files, commands, and tests from a non-interactive process.
22
- - **Keep parallel work organized.** Split issues into worker groups, run them
23
- in separate worktrees, review the results, and optionally run QA before a
24
- human-approved deploy.
25
- - **Remember the project.** Opt-in local memory stores project facts and
26
- rolling session summaries, with a pluggable storage boundary for a future
27
- remote backend.
28
- - **Bound the expensive parts.** Per-session cost, duration, iteration, and
29
- fleet-concurrency limits are built into the orchestration and watcher paths.
30
- - **Experiment with improvement.** The `improve` command runs a bounded
31
- recursive self-improvement loop against an external evaluator, with the
32
- default worker using a local Qwen 3.5 9B model through Ollama.
33
-
34
- The project also includes a GitHub issue watcher and an evaluation-only cloud
35
- provider interface. No live cloud resources are launched by the current
36
- implementation.
37
-
38
- > **Security warning: default-allow arbitrary command execution.**
39
- > By default, `headlesscode` runs arbitrary shell commands with the invoking
40
- > user's privileges. It can read and modify files, including credentials such
41
- > as `~/.ssh` and `~/.aws`, without approval prompts. Use it only with trusted
42
- > tasks and isolate it with a container, VM, or dedicated user when untrusted
43
- > content is involved. The optional permissions layer is defense in depth, not
44
- > a security boundary. See [`SECURITY.md`](./SECURITY.md).
8
+ `headlesscode` runs the Zoo Code agent loop from a Node.js process, without a
9
+ VS Code window. Run one task directly, launch parallel workers in Git
10
+ worktrees, or use the NVIDIA OpenShell provider for sandboxed sessions.
11
+
12
+ | Use it for | What it does |
13
+ | --- | --- |
14
+ | One repository task | Reads project modes and instructions, edits files, runs commands, and reports a result. |
15
+ | Parallel issue work | Splits GitHub issues into Git worktrees, runs workers, monitors completion, and reviews changes. |
16
+ | Intake automation | Watches a GitHub label and starts bounded batches of work. |
17
+ | OpenShell sessions | Runs worker and review sessions in disposable clones, then validates and imports results after sandbox deletion. |
18
+ | Harness experiments | Runs a bounded recursive self-improvement loop against an external evaluator. |
19
+
20
+ ## Security
21
+
22
+ > **Default-allow arbitrary command execution.** By default, `headlesscode`
23
+ > runs shell commands with the invoking user's privileges. It can read and
24
+ > modify files, including credentials such as `~/.ssh` and `~/.aws`, without
25
+ > approval prompts. Use it only with trusted tasks and isolate it with a
26
+ > container, VM, or dedicated user when untrusted content is involved. The
27
+ > optional permissions layer is defense in depth, not a security boundary.
28
+ > See [`SECURITY.md`](./SECURITY.md).
29
+
30
+ OpenShell runs each session in a disposable Git clone mounted at `/workspace`;
31
+ the host worktree and harness control files are not mounted. After the sandbox
32
+ is deleted, HeadlessCode validates a Git bundle and imports the result into the
33
+ host worktree. The gateway must support Docker bind mounts and must see the
34
+ configured scratch root at the same absolute path as the host. See the
35
+ [OpenShell guide](./docs/openshell-integration.md) for setup and security
36
+ details. This integration does not configure NVIDIA Sentry or hardware
37
+ monitoring.
45
38
 
46
39
  ## Quick start
47
40
 
48
- Install from npm:
41
+ Requires Node.js 18 or newer. Install globally or run with `npx`:
49
42
 
50
43
  ```bash
51
44
  npm install -g headlesscode
45
+ # or
46
+ npx headlesscode --task "Fix the bug in src/index.ts" --workspace /path/to/repo
52
47
  ```
53
48
 
54
- Or run it without installing, via npx:
49
+ For a real model call, set the OpenRouter key. `--dry-run` does not need a key.
55
50
 
56
- ```bash
57
- npx headlesscode --task "Fix the bug in src/index.ts" --workspace /path/to/target/repo
58
- ```
51
+ | Variable | Required | Purpose |
52
+ | --- | --- | --- |
53
+ | `HEADLESSCODE_OPENROUTER_API_KEY` | For OpenRouter calls | OpenRouter API key. |
54
+ | `OPENROUTER_MODEL` | No | Default model; `deepseek/deepseek-v4-flash-0731`. `--model` overrides it. |
55
+ | `OPENROUTER_HTTP_REFERER` | No | OpenRouter application referer header. |
56
+ | `OPENROUTER_APP_TITLE` | No | OpenRouter application title header. |
57
+ | `HEADLESSCODE_WORKSPACE_ROOT` | No | Default workspace; otherwise the current directory. |
59
58
 
60
59
  ```bash
61
- # Required (except for --dry-run):
62
60
  export HEADLESSCODE_OPENROUTER_API_KEY=sk-or-...
63
-
64
- # Optional:
65
- export OPENROUTER_MODEL=deepseek/deepseek-v4-flash-0731 # default model
66
- export OPENROUTER_HTTP_REFERER=https://example.com # OpenRouter app header
67
- export OPENROUTER_APP_TITLE="headlesscode" # OpenRouter X-Title header
68
- export HEADLESSCODE_WORKSPACE_ROOT=/path/to/target/repo # default workspace root
69
-
70
- headlesscode --task "Fix the bug in src/index.ts" --workspace /path/to/target/repo
61
+ headlesscode --task "Fix the bug in src/index.ts" --workspace /path/to/repo
71
62
  ```
72
63
 
73
- To work from a source checkout instead (for contributing):
64
+ For a source checkout:
74
65
 
75
66
  ```bash
76
67
  git clone https://github.com/Capsize-Games/headlesscode.git
77
68
  cd headlesscode
78
69
  npm install
79
- node bin/headlesscode.mjs --task "Fix the bug in src/index.ts" --workspace /path/to/target/repo
70
+ node bin/headlesscode.mjs --task "Fix the bug in src/index.ts" --workspace /path/to/repo
80
71
  ```
81
72
 
82
- ### Systemwide `headlesscode` command
73
+ To install a `headlesscode` wrapper on your `PATH`, run `scripts/install-cli.sh`.
74
+ It writes to `~/.local/bin` by default (override with `HEADLESSCODE_BIN_DIR`)
75
+ and runs this checkout's local `tsx`. Re-run it if you move the checkout.
83
76
 
84
- To avoid re-typing `node <checkout>/bin/headlesscode.mjs` from every project,
85
- install a `headlesscode` command onto your `PATH` once:
77
+ ### Check project instructions without calling a model
86
78
 
87
- ```bash
88
- scripts/install-cli.sh
89
- ```
90
-
91
- This writes a wrapper to `~/.local/bin/headlesscode` (override with
92
- `HEADLESSCODE_BIN_DIR`) that runs this checkout's `src/cli.ts` via its local
93
- `tsx`, without `cd`-ing — so `--repo`/`--workspace` still default to whatever
94
- directory you're standing in when you invoke it. Re-run the script any time
95
- after `git pull` to point it at a moved checkout; the wrapper itself doesn't
96
- need updating for ordinary code changes.
79
+ `--dry-run` builds the system prompt and validates project configuration. It
80
+ prints the prompt and a summary of loaded modes, tools, and prompt size. A
81
+ successful exit means prompt building and mode/rules loading succeeded.
97
82
 
98
83
  ```bash
99
- # from any repo, no HEADLESSCODE_ROOT plumbing needed:
100
- headlesscode orchestrate --repo . --issue 42
84
+ headlesscode --dry-run --mode code --workspace /path/to/repo
101
85
  ```
102
86
 
103
- The rest of this README uses the plain `headlesscode` form for brevity —
104
- substitute `node <checkout>/bin/headlesscode.mjs` if you haven't run
105
- `scripts/install-cli.sh` yet.
87
+ Project `.roomodes`, `.roo/rules-<slug>/`, and `AGENTS.md` instructions are
88
+ loaded by the prompt builder when applicable. The selected mode controls which
89
+ tools are exposed. See [Architecture](#architecture) for the relevant code.
106
90
 
107
- ## The `--dry-run` flow (no API key needed)
91
+ ## OpenShell
108
92
 
109
- `--dry-run` builds the full system prompt and validates configuration loading
110
- without calling the LLM — useful for CI and for checking that a target repo's
111
- `.roomodes` / `.roo/rules-<slug>/` / `AGENTS.md` are picked up:
93
+ Build the worker image and import the OpenRouter provider profile once. Configure
94
+ a Docker-backed OpenShell gateway with driver-config and bind-mount support;
95
+ keep that gateway restricted to trusted operators.
112
96
 
113
97
  ```bash
114
- headlesscode --dry-run --mode code --workspace /path/to/target/repo
115
- ```
116
-
117
- It prints the assembled system prompt plus a summary line (mode, custom modes
118
- loaded, exposed tools, prompt size). Exit code 0 means prompt building + mode /
119
- rules loading succeeded; non-zero means a config error.
98
+ # Run these commands from a HeadlessCode source checkout.
99
+ docker build -f docker/OpenShell.Dockerfile -t headlesscode-openshell:local .
100
+ openshell profile import -f shared/openshell/headlesscode-openrouter.yaml
120
101
 
121
- ## Real-usage example against a target repo
122
-
123
- ```bash
124
102
  export HEADLESSCODE_OPENROUTER_API_KEY=sk-or-...
125
- export HEADLESSCODE_WORKSPACE_ROOT=~/Projects/some-target-repo
126
-
127
- # Point the agent at an already-scoped issue, in an isolated worktree:
128
- headlesscode \
129
- --mode code \
130
- --task "Implement issue #29: add retry logic to the HTTP client (see .roo/rules for project conventions)." \
131
- --workspace ~/Projects/some-target-repo \
132
- --max-iterations 50 \
133
- --log-file ./headlesscode-session.log
103
+ export HEADLESSCODE_OPENSHELL_POLICY="$PWD/shared/openshell/headlesscode-policy.yaml"
104
+ headlesscode orchestrate --repo /path/to/repo --issue 42 \
105
+ --execution-provider openshell
134
106
  ```
135
107
 
136
- The agent reads files, runs commands, writes code, and finishes by calling
137
- `attempt_completion` (or by giving a final text answer). The final result is
138
- printed to stdout. Exit code 0 = success; 1 = task failed (max iterations or
139
- consecutive-mistake limit); 2 = usage/config error (e.g. missing
140
- `HEADLESSCODE_OPENROUTER_API_KEY`).
141
-
142
- ## Recursive self-improvement
143
-
144
- `headlesscode improve` is a bounded research loop for improving the harness
145
- itself. A supervisor creates isolated candidate worktrees, asks the local Qwen
146
- 3.5 9B worker to make focused changes, runs regression plus visible and hidden
147
- evaluations, and keeps the archive, score, and selection decision outside the
148
- candidate worktree. Each generation also produces a report for human review.
108
+ Each worker gets an independent clone and a temporary result directory. Neither
109
+ bind source is the host worktree. After the sandbox is deleted, HeadlessCode
110
+ imports a verified Git bundle into the host branch. The default scratch root is
111
+ `~/.local/share/headlesscode/openshell-sessions`; create it before starting the
112
+ gateway and bind-mount it into the gateway container at the same absolute path.
113
+ If you choose another root, set `HEADLESSCODE_OPENSHELL_SCRATCH_ROOT` and mount
114
+ that host directory into the gateway container at the same path. Project data and shared instructions
115
+ are copied into the scratch tree and mounted separately. The writable scratch
116
+ checkpoint store supports checkpoints during a session and is discarded at
117
+ teardown. Opt-in memory uses a scratch snapshot and merges validated JSONL
118
+ records after sandbox deletion. No OpenRouter key is passed as a normal sandbox
119
+ environment argument; OpenShell injects the imported provider credential.
120
+
121
+ `headlesscode watch` accepts the same `--execution-provider openshell` option.
122
+ OpenShell applies to worker, plan-first, continuation, reviewer, QA, and rework
123
+ sessions. The one-off `headlesscode openshell-session` command runs a single
124
+ task and keeps its host worktree for review. Set up a gateway callback address
125
+ that the sandbox's Docker bridge can reach; keep the client endpoint loopback
126
+ only when possible. Read the [OpenShell guide](./docs/openshell-integration.md)
127
+ before adapting the sample gateway or filesystem policy.
128
+
129
+ | Setting | Default | Meaning |
130
+ | --- | --- | --- |
131
+ | `--execution-provider` | `local` | `orchestrate`/`watch` runtime: `local` or `openshell`. Also settable with `HEADLESSCODE_EXECUTION_PROVIDER`. |
132
+ | `HEADLESSCODE_OPENSHELL_POLICY` | No default | Base OpenShell policy YAML; required for OpenShell sessions. |
133
+ | `HEADLESSCODE_OPENSHELL_IMAGE` | `headlesscode-openshell:local` | Sandbox image. |
134
+ | `HEADLESSCODE_OPENSHELL_PROVIDERS` | `headlesscode-openrouter` | Comma-separated credential profile names. |
135
+ | `HEADLESSCODE_OPENSHELL_CPU` / `HEADLESSCODE_OPENSHELL_MEMORY` | `2` / `4Gi` | Per-sandbox resource requests. |
136
+ | `HEADLESSCODE_OPENSHELL_NAME_PREFIX` | `hcls` | Prefix for deterministic sandbox names. |
137
+ | `HEADLESSCODE_OPENSHELL_SCRATCH_ROOT` | `~/.local/share/headlesscode/openshell-sessions` | Host scratch directory that the gateway container must see at the same path. |
138
+
139
+ The `headlesscode openshell-session` command provides a single-session live
140
+ entry point; run `headlesscode openshell-session --help` for options.
141
+
142
+ ## Parallel work and GitHub intake
143
+
144
+ `orchestrate` runs a parallel round from GitHub issue numbers or a JSON file.
145
+ It splits work, creates a worktree per group, starts workers, watches for
146
+ completion, and reviews results. QA and deployment approval are optional.
149
147
 
150
148
  ```bash
151
- # Inspect the planned experiment without creating worktrees or calling a model:
152
- npx tsx src/cli.ts improve --repo . --dry-run
153
-
154
- # Run one small generation with the local Ollama model:
155
- npx tsx src/cli.ts improve --repo . --population 2 --generations 1
149
+ headlesscode orchestrate --repo /path/to/repo --issue 42 --issue 43 --qa
150
+ headlesscode orchestrate status --repo /path/to/repo --wait --timeout-ms 7200000
156
151
  ```
157
152
 
158
- This is an experiment, not unattended model training and not a security
159
- boundary. The current loop records model-candidate and training interfaces,
160
- trajectory datasets, parent-selection policies, and resource-tagged jobs, but
161
- does not yet train adapters, schedule multiple trajectories, or provide
162
- OS-level candidate isolation. Those limits are tracked in [the RSI roadmap](#rsi-roadmap).
153
+ `watch` polls GitHub for issues with a chosen label and starts bounded batches.
154
+ It requires `GH_TOKEN` or `GITHUB_TOKEN`; it performs intake only. Use a later
155
+ `orchestrate` run to monitor, review, QA, or deploy a batch.
163
156
 
164
- ## RSI roadmap
165
-
166
- The implemented loop is intentionally honest about what remains. Follow-up
167
- work is tracked in GitHub:
168
-
169
- - [OS-level candidate sandbox](https://github.com/Capsize-Games/headlesscode/issues/3)
170
- - [Cryptographically verifiable evaluator and artifacts](https://github.com/Capsize-Games/headlesscode/issues/4)
171
- - [Resource-aware resumable scheduler](https://github.com/Capsize-Games/headlesscode/issues/5)
172
- - [Adaptive multi-trajectory search](https://github.com/Capsize-Games/headlesscode/issues/6)
173
- - [Validated curriculum fixtures](https://github.com/Capsize-Games/headlesscode/issues/7)
174
- - [Adversarial evaluation and cross-model supervision](https://github.com/Capsize-Games/headlesscode/issues/8)
175
- - [Real LoRA or QLoRA backend](https://github.com/Capsize-Games/headlesscode/issues/9)
176
-
177
- See [`docs/recursive-self-improvement.md`](./docs/recursive-self-improvement.md)
178
- for the design and [`docs/rsi-progress.md`](./docs/rsi-progress.md) for the
179
- record of the first bounded runs.
180
-
181
- ## Registering a new project
157
+ ```bash
158
+ GH_TOKEN=... headlesscode watch \
159
+ --owner my-org --repo /path/to/repo --label needs-agent --run-once
160
+ ```
182
161
 
183
- To use headlesscode against a project that isn't set up as a headlesscode
184
- project yet, register it in one step:
162
+ | Command | Important options | Behavior and limits |
163
+ | --- | --- | --- |
164
+ | `headlesscode orchestrate --repo <path>` | Repeat `--issue <n>` or use `--issues-json <file>`; `--execution-provider local|openshell`, `--qa`, `--deploy`, `--no-review`, `--dry-run`, `--max-iterations <n>`, `--max-rework-cycles <n>`, `--max-continuations <n>` | Worktree per group; default review is enabled. OpenShell applies to worker, plan-first, continuation, reviewer, QA, and rework sessions. `--deploy` uses the human-approval gate. |
165
+ | `headlesscode orchestrate status --repo <path>` | `--wait`, `--timeout-ms <n>`, `--json` | Reads durable round state; `--wait` blocks until terminal state or timeout. |
166
+ | `headlesscode orchestrate stop --repo <path>` | Repeat `--group <name>` | Stops selected groups' worker process trees. |
167
+ | `headlesscode watch --owner <o> --repo <path> --label <name>` | `--execution-provider local|openshell`, `--run-once`, `--dry-run`, `--max-per-sweep <n>`, `--max-concurrent-sessions <n>`, `--retry-failed` | Label-based intake. `--dry-run` still needs a GitHub token to list issues, but writes no state and starts no workers. |
168
+
169
+ Both orchestration and watcher enforce a shared concurrency cap, defaulting to
170
+ 3 (`HEADLESSCODE_MAX_CONCURRENT_SESSIONS`). Orchestration aborts when the cap
171
+ is full; the watcher leaves excess issues pending. The watcher records
172
+ idempotency state under `<repo>/.worktrees/.watcher-state.json` by default.
173
+ Use `--help` on a command for its complete options. Detailed guides:
174
+ [orchestration](./docs/phase2-orchestration.md),
175
+ [issue watcher](./docs/phase5-issue-watcher.md),
176
+ [QA implementation](./src/qa/qa.ts), and
177
+ [deployment gate](./docs/phase4-deploy-gate.md).
178
+
179
+ ## Project registration and search
180
+
181
+ Registering a project detects its stack, ensures `.gitignore` excludes
182
+ `.headlesscode/`, and builds a code-search index and codemap. Indexing calls an
183
+ embedding API and can incur charges unless skipped or configured for Ollama.
185
184
 
186
185
  ```bash
187
186
  headlesscode init --workspace ~/Projects/your-project
187
+ headlesscode init --workspace ~/Projects/your-project --skip-index
188
188
  ```
189
189
 
190
- This detects the project's stack(s) (which drives per-session instruction
191
- selection), makes sure the project's `.gitignore` excludes `.headlesscode/`
192
- session artifacts, and builds the codebase-search index and the codemap — no
193
- manual `index`/`codemap`/`.gitignore` steps needed. The index step calls the
194
- embedding API and costs real money unless you pass `--skip-index` or
195
- `--embedding-backend ollama`. See `headlesscode init --help` for the full
196
- options.
190
+ See `headlesscode init --help` for `--skip-codemap` and embedding-backend
191
+ options. The CLI also provides `index`, `codemap`, project-store, GitHub App,
192
+ and decision-proxy commands; run `headlesscode --help` for the command list
193
+ and each subcommand's `--help` for complete usage.
197
194
 
198
- ## CLI reference
195
+ ## Recursive self-improvement
199
196
 
200
- ```
201
- headlesscode --task "<task text>" [options]
202
- headlesscode --task-file <path> [options]
203
- headlesscode --dry-run [options]
204
- headlesscode orchestrate --repo <path> --issue <n> [--issue <n> ...] [options]
205
- headlesscode watch --owner <o> --repo <path> --label <name> [options]
206
-
207
- --mode <slug> Mode to run in (built-in or from .roomodes). Default: code
208
- --task <text> The task description for the agent
209
- --task-file <path> Read the task from a file (relative to workspace)
210
- --workspace <root> Workspace root (default: $HEADLESSCODE_WORKSPACE_ROOT or cwd)
211
- --model <id> OpenRouter model id (default: $OPENROUTER_MODEL or deepseek/deepseek-v4-flash-0731)
212
- --max-iterations <n> Loop iteration cap (default: 50)
213
- --consecutive-error-limit <n> Consecutive mistakes before giving up (default: 3)
214
- --max-cost-usd <n> Phase 6: per-session cost cap in USD (decimal). Default
215
- $HEADLESSCODE_MAX_COST_USD; off when neither is set
216
- --max-duration-ms <n> Phase 6: per-session wall-clock cap in ms. Default
217
- $HEADLESSCODE_MAX_DURATION_MS; off when neither is set.
218
- A tripped cap aborts with reason "budget"
219
- --log-file <path> Also append structured logs to this file
220
- --memory-dir <path> Phase 3 memory: store facts + session summaries under <path>
221
- (enabled; default $HEADLESSCODE_MEMORY_DIR or
222
- <workspace>/.headlesscode/memory). Memory is OFF unless set.
223
- --no-memory Explicitly disable memory even if HEADLESSCODE_MEMORY_DIR is set
224
- --allowed-commands <list> Comma-separated command prefixes the agent may run.
225
- Default: $HEADLESSCODE_ALLOWED_COMMANDS, else
226
- .headlesscode/permissions.json, else empty (=
227
- allow everything except --denied-commands; see
228
- SECURITY.md)
229
- --denied-commands <list> Comma-separated command prefixes that are ALWAYS
230
- refused (deny wins over allow; dangerous shell
231
- substitutions are always blocked regardless).
232
- Default: $HEADLESSCODE_DENIED_COMMANDS, else
233
- .headlesscode/permissions.json, else empty
234
- --protected-files <list> Comma-separated glob patterns of files the agent may
235
- not write. Default: $HEADLESSCODE_PROTECTED_FILES,
236
- else .headlesscode/permissions.json, else
237
- ".env,.env.*,*.pem,*.key,id_rsa*"
238
- --allow-protected-writes Escape hatch: permit writes to protected files
239
- (default: OFF). Also settable via
240
- "allowProtectedWrites": true in
241
- .headlesscode/permissions.json
242
- --dry-run Build the system prompt + validate config, then exit (no API key)
243
- --version / --help
244
-
245
- orchestrate subcommand (Phase 2 — parallel worktrees, headless workers):
246
- --repo <path> Target repo root (required)
247
- --issue <n> Issue number to include (repeatable)
248
- --issues-json <file> Read issues from a JSON array of {number,title,body}
249
- (used when gh is unavailable, or for tests)
250
- --file-issues With --issues-json: file a REAL GitHub issue for
251
- each synthetic entry (gh issue create --repo
252
- <origin-owner>/<origin-repo>), swap in the real
253
- number returned, and print one confirmation line
254
- per created issue. A real, visible write to
255
- GitHub — opt-in, never automatic. Requires
256
- --issues-json
257
- --batch <name> Batch id in the state file (default round-<date>)
258
- --review-mode <slug> Mode slug for review sessions (default deepseek-reviewer)
259
- --no-review Spawn + watch only; skip the reviewer
260
- --qa Phase 4: run a headless QA session (--mode qa-agent)
261
- on each group after its review passes; record
262
- qa {status,verdict,evidence} in the state file
263
- --qa-mode <slug> Mode slug for QA sessions (default qa-agent; the
264
- target repo's .roomodes + .roo/rules-<slug>/ are
265
- spliced automatically)
266
- --deploy Phase 4: after all groups done + reviewed + QA passed,
267
- run the human-approval deploy gate
268
- (scripts/deploy-gate.sh) — a hard stop that never runs
269
- the repo's deploy-production.sh without explicit human
270
- approval (interactive on a TTY, token/file otherwise)
271
- --deploy-args <str> Deploy args forwarded to the deploy script after the
272
- gate approves (space-separated flags; also DEPLOY_ARGS env)
273
- --poll-interval-ms <n> Watcher poll interval (default 5000)
274
- --max-concurrent-sessions <n> Phase 6 global cap on concurrent sessions across
275
- processes (default $HEADLESSCODE_MAX_CONCURRENT_
276
- SESSIONS or 3). At/over the cap this run ABORTS
277
- with a clear message, exit 1
278
- --dry-run Print the split plan + spawn commands, spawn nothing
279
-
280
- watch subcommand (Phase 5 — GitHub issue watcher, poll-based intake):
281
- --owner <o> GitHub owner (required)
282
- --repo <path> Local clone of the target repo (required; worktrees
283
- are spawned under <path>/.worktrees/). The GitHub repo
284
- name defaults to the directory basename (--gh-repo overrides)
285
- --label <name> The label that triggers processing, e.g. needs-agent
286
- --poll-interval-ms <n> Sweep interval in continuous mode (default 60000)
287
- --run-once One sweep then exit 0 (or 1 if any spawn failed)
288
- --max-per-sweep <n> Max NEW issues spawned per sweep (default 5); the rest
289
- stay 'pending' in state and are picked up next sweep
290
- --max-concurrent-sessions <n> Phase 6 GLOBAL cap on concurrent sessions across
291
- processes (default $HEADLESSCODE_MAX_CONCURRENT_
292
- SESSIONS or 3). Interplay: maxPerSweep bounds one
293
- sweep's burst; this bounds the total fleet — issues
294
- beyond it stay 'pending' until slots free up
295
- --state-file <path> Durable idempotency state (default
296
- <repo>/.worktrees/.watcher-state.json)
297
- --mode <slug> / --memory-dir <path>
298
- Forwarded to the spawner (ORCHESTRATOR_MODE /
299
- HEADLESSCODE_MEMORY_DIR)
300
- --qa / --deploy Pass-through: recorded per batch for the follow-up
301
- orchestrate completion run
302
- --dry-run Sweep + print the spawn plan, spawn nothing, write no
303
- state (still needs a GitHub token for listIssues)
304
- --retry-failed Retry previously-failed spawns next sweep
197
+ `headlesscode improve` runs a bounded experiment against an external evaluator.
198
+ The supervisor creates candidate worktrees, asks the local Ollama worker to
199
+ make focused changes, runs regression and visible/hidden evaluations, and
200
+ keeps selection state outside candidate worktrees. Each generation produces a
201
+ report for human review.
202
+
203
+ ```bash
204
+ npx tsx src/cli.ts improve --repo . --dry-run
205
+ npx tsx src/cli.ts improve --repo . --population 2 --generations 1
305
206
  ```
306
207
 
208
+ This loop does not yet train adapters, schedule multiple trajectories, or
209
+ provide OS-level candidate isolation. Follow-up work is tracked as:
210
+
211
+ | Issue | Work |
212
+ | --- | --- |
213
+ | [#3](https://github.com/Capsize-Games/headlesscode/issues/3) | OS-level candidate sandbox |
214
+ | [#4](https://github.com/Capsize-Games/headlesscode/issues/4) | Cryptographically verifiable evaluator and artifacts |
215
+ | [#5](https://github.com/Capsize-Games/headlesscode/issues/5) | Resource-aware resumable scheduler |
216
+ | [#6](https://github.com/Capsize-Games/headlesscode/issues/6) | Adaptive multi-trajectory search |
217
+ | [#7](https://github.com/Capsize-Games/headlesscode/issues/7) | Validated curriculum fixtures |
218
+ | [#8](https://github.com/Capsize-Games/headlesscode/issues/8) | Adversarial evaluation and cross-model supervision |
219
+ | [#9](https://github.com/Capsize-Games/headlesscode/issues/9) | Real LoRA or QLoRA backend |
220
+
221
+ See the [RSI design](./docs/recursive-self-improvement.md) and
222
+ [run progress](./docs/rsi-progress.md).
223
+
224
+ ## Core CLI options
225
+
226
+ | Option | Default or source | Purpose |
227
+ | --- | --- | --- |
228
+ | `--task <text>` / `--task-file <path>` | One is required unless `--dry-run` | Task prompt; task file is resolved relative to the workspace. |
229
+ | `--workspace <path>` | `HEADLESSCODE_WORKSPACE_ROOT` or current directory | Repository directory exposed to the worker. |
230
+ | `--mode <slug>` | `code` | Built-in or project `.roomodes` mode. |
231
+ | `--model <id>` | `OPENROUTER_MODEL` or `deepseek/deepseek-v4-flash-0731` | Model ID. |
232
+ | `--max-iterations <n>` | 250 for direct CLI sessions; orchestrator workers default to 50 | Per-session loop bound. |
233
+ | `--consecutive-error-limit <n>` | 3 (local code backend may use 6) | Consecutive tool/model errors before stopping. |
234
+ | `--max-cost-usd <n>` / `--max-duration-ms <n>` | `HEADLESSCODE_MAX_COST_USD` / `HEADLESSCODE_MAX_DURATION_MS`; disabled if unset | Per-session budget caps. A tripped cap aborts the session. |
235
+ | `--memory-dir <path>` / `--no-memory` | Off unless `HEADLESSCODE_MEMORY_DIR` or option is set | Enable project facts and rolling session summaries, or explicitly disable. |
236
+ | `--allowed-commands <list>` / `--denied-commands <list>` | CLI, environment, `.headlesscode/permissions.json`; allow list is empty by default | Command-prefix policy. Denials win; dangerous shell substitutions are blocked. See [`SECURITY.md`](./SECURITY.md). |
237
+ | `--protected-files <globs>` / `--allow-protected-writes` | `.env,.env.*,*.pem,*.key,id_rsa*` | Block writes matching protected files; explicit escape hatch. |
238
+ | `--dry-run` | Off | Build prompt and validate configuration without a model call or API key. |
239
+ | `--log-file <path>` | No file | Also append structured session logs to a file. |
240
+
241
+ Run `headlesscode --help` for advanced controls such as context condensation,
242
+ checkpoints, recursive delegation, session pause, and LLM timeout. Some options
243
+ are backend-specific; help output and source defaults are authoritative.
244
+
307
245
  ## Architecture
308
246
 
309
- - [`src/llm/openrouter.ts`](./src/llm/openrouter.ts) — OpenRouter
310
- chat-completions client using native `fetch` (no axios/node-fetch/openai
311
- SDK). Reads `HEADLESSCODE_OPENROUTER_API_KEY`, optional `OPENROUTER_HTTP_REFERER` /
312
- `OPENROUTER_APP_TITLE`; model default `deepseek/deepseek-v4-flash-0731` (env
313
- `OPENROUTER_MODEL` overrides). Non-2xx → typed `OpenRouterError` with status +
314
- body excerpt; supports `AbortSignal` timeouts.
315
- - [`src/tools/executor.ts`](./src/tools/executor.ts) — headless executor for
316
- `read_file`, `write_to_file`, `execute_command`, `list_files` (plain
317
- `fs`/`child_process`), plus `attempt_completion` / `ask_followup_question`
318
- handlers and "not implemented" stubs for every other vendored tool schema.
319
- All file operations are resolved relative to the workspace root and rejected
320
- if they escape it (path-traversal guard via `path.resolve` + containment
321
- check). Command results are truncated (~30k chars) to keep context bounded.
322
- - [`src/tools/output-summarizer.ts`](./src/tools/output-summarizer.ts) —
323
- OPT-IN local summarization of oversized `execute_command` output: when
324
- `HEADLESSCODE_LOCAL_SUMMARIZATION=1`, a result that would exceed the 30k
325
- char cap is compressed by a local Ollama chat model before reaching the
326
- cloud model, with a `[Output summarized by local model…]` transparency
327
- header. OFF by default; on any failure it falls back to today's exact blunt
328
- truncation (never an error, never a hang). Endpoint/model configurable via
329
- `HEADLESSCODE_OLLAMA_URL` (default `http://localhost:11434`) and
330
- `HEADLESSCODE_SUMMARIZATION_MODEL` (default `qwen3:8b`). Deliberately
331
- limited to command output — file/diff content always stays verbatim.
332
- - [`src/engine/parser.ts`](./src/engine/parser.ts) — OpenAI function-calling
333
- parser: JSON.parses `tool_calls[].function.arguments` with a best-effort
334
- partial-JSON fallback; parse failures are marked and fed back as errors.
335
- - [`src/engine/prompt.ts`](./src/engine/prompt.ts) — wraps the vendored
336
- `SYSTEM_PROMPT` builder: loads project `.roomodes` (same zod schema as Zoo
337
- Code's `CustomModesManager`), passes the workspace as `cwd` so `.roo/rules-*`
338
- / `AGENTS.md` splice in, and selects the mode's exposed tools.
339
- - [`src/engine/loop.ts`](./src/engine/loop.ts) — `HeadlessSession`, the
340
- orchestration loop. System + user → LLM → assistant (with tool_calls) → parse
341
- → execute → `tool` role message → repeat. Terminates on `attempt_completion`
342
- (its `args.result` is the final answer) or a text-only reply; fails bounded
343
- on max iterations or `consecutiveErrorLimit` consecutive mistakes (tool
344
- errors / parse errors / identical repeated calls). History truncation is a
345
- Phase 1 placeholder: system + first user always kept, sliding window of the
346
- last ~40 messages. The loop accepts an injected `llmClient` (DI) so tests use
347
- a fake; the CLI wires `OpenRouterClient`.
348
- - [`src/engine/logger.ts`](./src/engine/logger.ts) — structured logger
349
- (timestamped lines to stdout/stderr, optional file).
350
- - [`src/memory/`](./src/memory/index.ts) — memory subsystem:
351
- - [`src/memory/types.ts`](./src/memory/types.ts) — the `MemoryFact` /
352
- `SessionSummary` schema (per-project scoped, kinds
353
- `convention|decision|failure|knowledge` where "things that didn't work"
354
- are `failure`) and the two contracts: `MemoryStore` (the pluggable
355
- storage boundary) and `Embedder`. Hard data-isolation requirement
356
- documented: the harness knowledge is a dedicated schema, never reachable
357
- through any customer tenant route.
358
- - [`src/memory/local.ts`](./src/memory/local.ts) — `LocalMemoryStore`: the
359
- fully working file backend (`facts/<project>.jsonl` +
360
- `sessions/<project>.jsonl`, append-only, idempotent `addFact` by content
361
- hash, `queryRecall` = keyword matches (high weight) + local-embedder
362
- cosine similarity, deterministic ordering).
363
- - [`src/memory/embed.ts`](./src/memory/embed.ts) — `createLocalEmbedder()`:
364
- zero-dependency, deterministic lexical-hash embedder (lowercase word +
365
- char-bigram tokens → fixed-dim L2-normalized vector). Placeholder for a
366
- real local embedding model behind the same `Embedder` interface.
367
- - [`src/memory/summarizer.ts`](./src/memory/summarizer.ts) —
368
- `extractSessionSummary` (deterministic; files/commands derived from the
369
- tool history, facts via keyword heuristics) + `buildRollingSummary`
370
- (compact markdown recap of the last N sessions, so a session never needs
371
- the infinite raw history).
372
- - [`src/memory/uwuchat.ts`](./src/memory/uwuchat.ts) — a remote
373
- `MemoryStore` implementation stub for a future hosted memory API. Throws
374
- "not implemented" until its base-URL/token env vars are set.
375
- - Memory is wired into `HeadlessSession` as an **opt-in** config (`memory`
376
- + `project`); when unset the loop behaves exactly as before. When set, the
377
- loop injects a `## PROJECT MEMORY` section (recalled facts + rolling
378
- recap) into the first user message and records the session + extracted
379
- facts afterwards — and memory failures are always non-fatal.
380
- - [`src/orchestrator/`](./src/orchestrator/index.ts) — Phase 2 orchestration
381
- layer: `split.ts` (issue-splitting heuristics, deterministic +
382
- unit-tested), `state.ts` (`.worktrees/.orchestrator-state.json` read/write),
383
- `reviewer.ts` (adversarial fresh-context review run with a read-only
384
- executor), `watch.ts` (completion polling of `.harness.done` markers + stall
385
- guard), and `cli.ts` (the `orchestrate` subcommand — split → spawn via
386
- `scripts/spawn-parallel-worktrees.sh` → watch → review → QA → deploy gate).
387
- - [`src/qa/qa.ts`](./src/qa/qa.ts) — Phase 4 headless QA: `runQa()` runs a
388
- second harness session against a worktree in the target repo's `qa-agent`
389
- mode (auto-spliced from `.roomodes` + `.roo/rules-qa-agent/`), with a
390
- generic checklist fallback when the repo has no such mode. Read + command
391
- tools only (no `write_to_file`) — QA verifies and reports, it never edits.
392
- Verdict parsing is fail-closed (`pass`/`fail`/`error`, default `fail`).
393
- - [`src/deploy/gate.ts`](./src/deploy/gate.ts) + [`src/deploy/gate-cli.ts`](./src/deploy/gate-cli.ts) —
394
- Phase 4 human-approval deploy gate: the pure, unit-tested decision function
395
- `decideApproval` (interactive y/N, one-time approval file, or
396
- `DEPLOY_APPROVAL_TOKEN` matching `<repo>/.deploy-approval`; never
397
- auto-approves) plus a thin CLI the bash wrapper calls.
398
- - [`src/watcher/`](./src/watcher/index.ts) — Phase 5 GitHub issue watcher:
399
- [`github.ts`](./src/watcher/github.ts) (native-fetch GitHub REST client with
400
- label filter, PR filtering, pagination, `GITHUB_API_BASE_URL` override for
401
- tests/mocks), [`state.ts`](./src/watcher/state.ts) (durable idempotency
402
- state file — write-ahead `spawned` → `done`/`failed`, `pending` for capped
403
- issues, restart-safe), [`watch.ts`](./src/watcher/watch.ts) (the poll loop:
404
- list by label → split → spawn via the existing bash spawner, bounded by
405
- `maxPerSweep` and the Phase 6 global cap), and [`cli.ts`](./src/watcher/cli.ts)
406
- (the `watch` subcommand — continuous or `--run-once`, `--dry-run`,
407
- `--retry-failed`).
408
- - [`src/budget/`](./src/budget/index.ts) — Phase 6 guardrails:
409
- [`cost.ts`](./src/budget/cost.ts) (model pricing table + `estimateCost`,
410
- `HEADLESSCODE_PRICING_JSON` override, conservative fallback for unlisted
411
- models), [`budget.ts`](./src/budget/budget.ts) (`SessionBudget` +
412
- `BudgetTracker`: `tick()` before each LLM call, `record()` after with usage
413
- tokens, `check()` snapshot, `BudgetExceededError`), and
414
- [`concurrency.ts`](./src/budget/concurrency.ts) (`ConcurrencyLimiter` —
415
- fail-fast in-process semaphore — plus `activeSessionCount` reading the
416
- durable orchestrator/watcher state files for a cross-process view). Wired
417
- into `HeadlessSession` (`budget` config, `budgetUsage` on results), the base
418
- CLI (`--max-cost-usd` / `--max-duration-ms`), `orchestrate` (aborts at the
419
- cap), the watcher (defers cap-exceeding issues to `pending`), and
420
- `run-worker.sh`/`run-qa.sh` (env forwarding).
421
- - [`src/cloud/`](./src/cloud/provider.ts) — Phase 6 ephemeral compute
422
- abstraction: the `CloudProvider` lifecycle interface
423
- (`spawnWorktreeSession` → `waitReady` → `runHarness` → `collectResults` →
424
- `teardown`) with `LocalProcessProvider` as the current local behavior behind
425
- it (reuses `spawn-parallel-worktrees.sh` + `run-worker.sh`), so a
426
- container/VM-per-issue backend slots in without touching the orchestration
427
- layer. A cloud-provider sketch is documented (evaluation only — no
428
- live setup; see [`docs/phase6-cloud.md`](./docs/phase6-cloud.md)).
429
- - [`src/cli.ts`](./src/cli.ts) — the `headlesscode` bin entry (+ `orchestrate`
430
- and `watch` subcommand dispatch).
431
- - [`scripts/run-worker.sh`](./scripts/run-worker.sh) — launches one headless
432
- harness worker per worktree (pid, log, exit code, `.harness.done` marker).
433
- - [`scripts/run-qa.sh`](./scripts/run-qa.sh) — Phase 4 QA wrapper mirroring
434
- run-worker.sh: launches one harness QA session per worktree (`.qa-task.md`,
435
- `.qa.pid`, `qa.log`, `.qa.exit`, `.qa.done/`).
436
- - [`scripts/deploy-gate.sh`](./scripts/deploy-gate.sh) — Phase 4 gate wrapper:
437
- path safety, deployment summary, interactive + token/file approval, then
438
- (and only then) invokes the repo's `scripts/deploy-production.sh` with
439
- forwarded deploy args. Exit 3 = human DENIED (hard stop).
440
- - [`scripts/spawn-parallel-worktrees.sh`](./scripts/spawn-parallel-worktrees.sh)
441
- — spawns one git worktree + harness worker per group (worktree/.env/branch
442
- conventions, `run-worker.sh` + state-file writes, no GUI involved).
443
-
444
- ### Phase 1 tool filtering decision
445
-
446
- The loop exposes to the model exactly the tools the executor can actually run
447
- for the selected mode: the intersection of the vendored mode tool groups
448
- (`getToolsForMode`) with the Phase 1 executable set
449
- (`read_file`, `write_to_file`, `execute_command`, `list_files`,
450
- `attempt_completion`, `ask_followup_question`). Stub-only tools (`apply_diff`,
451
- `search_files`, …) stay registered in the executor purely as a safety net
452
- (clear "not implemented" error) but are NOT advertised to the model, so it
453
- doesn't waste turns calling them.
247
+ The core loop builds the project-specific prompt, calls the configured model,
248
+ executes the selected mode's available tools, and repeats until completion or a
249
+ configured limit. File tools are confined to the workspace path; shell commands
250
+ run with the process user's privileges unless an external boundary such as
251
+ OpenShell is configured. Mode instructions and safety controls are distinct:
252
+ prompt instructions guide the model, while filesystem or command restrictions
253
+ must be enforced by the runtime or sandbox.
254
+
255
+ | Area | Entry points | Responsibility |
256
+ | --- | --- | --- |
257
+ | CLI and session loop | [`src/cli.ts`](./src/cli.ts), [`src/engine/`](./src/engine/) | Command dispatch, prompt construction, LLM/tool loop, iteration and context limits. |
258
+ | Model and tools | [`src/llm/`](./src/llm/), [`src/tools/`](./src/tools/) | OpenRouter client, tool execution, output handling, and workspace path checks. |
259
+ | Parallel orchestration | [`src/orchestrator/`](./src/orchestrator/), [`scripts/spawn-parallel-worktrees.sh`](./scripts/spawn-parallel-worktrees.sh) | Issue grouping, worktree workers, completion state, review, rework, QA, and deploy gating. |
260
+ | GitHub intake | [`src/watcher/`](./src/watcher/) | Label polling, durable idempotency, pending work, and bounded spawning. |
261
+ | Runtime providers | [`src/cloud/`](./src/cloud/) | Local process, Docker, and OpenShell implementations of the session lifecycle. `orchestrate` and `watch` can select OpenShell for worker sessions. |
262
+ | Memory and budgets | [`src/memory/`](./src/memory/), [`src/budget/`](./src/budget/) | Opt-in project memory, cost/time budgets, and session concurrency accounting. |
263
+ | QA and deploy gate | [`src/qa/`](./src/qa/), [`src/deploy/`](./src/deploy/) | Read-only QA verdicts and explicit human approval before deployment. |
264
+ | Project map and search | [`src/codemap/`](./src/codemap/), [`src/codesearch/`](./src/codesearch/) | Deterministic module maps and repository search indexes. |
265
+
266
+ The model only receives tools supported by the selected mode and executable
267
+ runtime. Unsupported vendored tool schemas are not advertised. This keeps the
268
+ model's action set aligned with the current executor; it does not make arbitrary
269
+ shell commands safe by itself.
454
270
 
455
271
  ## Development
456
272
 
457
- ```bash
458
- npm run typecheck # npx tsc --noEmit (whole repo incl. vendored core)
459
- npm run smoke # vendored prompt builder smoke test (no network)
460
- npm test # unit tests with fake LLM clients (no network/key)
461
- bash scripts/e2e/run.sh # Phase 1 integration (mock OpenRouter)
462
- bash scripts/e2e-phase2/run.sh # Phase 2 integration (spawn + watch + review)
463
- bash scripts/e2e-phase4/run.sh # Phase 4 integration (QA + deploy gate, fake deploy)
464
- bash scripts/e2e-phase5/run.sh # Phase 5 integration (watcher vs fake GitHub server,
465
- # stubbed spawner: state transitions + no double-spawn)
466
- bash scripts/e2e-phase6/run.sh # Phase 6 integration (budget abort via mock OpenRouter
467
- # + concurrency-cap sweep via fake GitHub + stubbed spawner)
468
- npm run cli -- --dry-run --workspace . # build this repo's system prompt
469
- ```
470
-
471
- The engine tests (`src/engine/__tests__/loop.test.ts`) run the full loop with a
472
- fake `LlmClient` injected via the `HeadlessSession` constructor — no network,
473
- no API key required. Phase 2 adds `src/orchestrator/__tests__/` (split
474
- heuristics, state round-trip, reviewer verdict parsing), and the e2e scripts
475
- drive the real CLI through a local mock OpenRouter server, including a
476
- 2-worktree parallel spawn, completion-marker polling, and a read-only review
477
- invocation. Phase 5 adds `src/watcher/__tests__/` (github client with an
478
- injected fetch, watcher-state idempotency semantics, and the watch loop with
479
- an injected gh client + spawner covering spawn/idempotency/cap/failure/
480
- dry-run/abort) and `scripts/e2e-phase5/run.sh` (watcher against a fake GitHub
481
- server with a stubbed spawner).
482
-
483
- ## Phase status
484
-
485
- - ✅ Phase 1 Subtask 1 — vendored portable Zoo Code core
486
- ([`src/vendor/zoo-code/`](./src/vendor/zoo-code/), read-only dependency).
487
- - ✅ Phase 1 Subtask 2 — runtime engine (OpenRouter client, tool executor,
488
- parser, orchestration loop, CLI, tests).
489
- - ✅ Phase 2 — headless orchestration layer: drop-in
490
- `spawn-parallel-worktrees.sh` + `run-worker.sh` (harness subprocess per
491
- worktree, `.harness.done` completion markers), issue-splitting heuristics
492
- port, `.orchestrator-state.json` state management, headless reviewer, and
493
- the `orchestrate` CLI subcommand (see
494
- [`docs/phase2-orchestration.md`](./docs/phase2-orchestration.md)).
495
- - ✅ Phase 3 — memory subsystem: per-project knowledge facts + rolling session
496
- summaries (`src/memory/`), local deterministic embedder for semantic recall,
497
- opt-in `HeadlessSession`/CLI wiring (`--memory-dir`, `--no-memory`), and a
498
- pluggable `MemoryStore` contract with a remote-backend client stub.
499
- - ✅ Phase 4 — QA + deploy gate: headless QA runs the target repo's `qa-agent`
500
- mode (`--qa` / `--qa-mode`), verdict parsing is fail-closed, results land in
501
- the state file's per-group `qa` field; the human-approval deploy gate
502
- (`--deploy`, `scripts/deploy-gate.sh` + `src/deploy/gate.ts`) is a hard stop
503
- in front of `deploy-production.sh` that never auto-approves (see
504
- [`docs/phase4-qa.md`](./docs/phase4-qa.md) and
505
- [`docs/phase4-deploy-gate.md`](./docs/phase4-deploy-gate.md)).
506
- - ✅ Phase 5 — GitHub issue watcher: poll-based intake (`watch` subcommand) —
507
- detect issues by label via the GitHub REST API (`GH_TOKEN`), fan each out
508
- through `splitIssues` + the existing `spawn-parallel-worktrees.sh`, track
509
- idempotency durably (`.worktrees/.watcher-state.json`, write-ahead
510
- ordering, `pending` cap deferral, restart-safe), optional `--dry-run` /
511
- `--run-once` / `--retry-failed`; webhook upgrade designed but not built as
512
- a server (see [`docs/phase5-issue-watcher.md`](./docs/phase5-issue-watcher.md)).
513
- - ✅ Phase 6 — cloud scaling, guardrails-first: per-session cost/time/iteration
514
- budget (`src/budget/`, `--max-cost-usd` / `--max-duration-ms`, budgetUsage on
515
- results, worker env forwarding) + a hard concurrent-session cap
516
- (`HEADLESSCODE_MAX_CONCURRENT_SESSIONS`, orchestrate aborts / watcher defers
517
- to `pending`) + the `CloudProvider` abstraction with `LocalProcessProvider`
518
- and an evaluation-only cloud-provider sketch (see
519
- [`docs/phase6-cloud.md`](./docs/phase6-cloud.md)). No live cloud launched;
520
- a container/VM-per-issue backend slots in behind the same interface.
521
- - ⏳ Phase 3 (remaining) — token-based condensation.
273
+ | Command | Purpose |
274
+ | --- | --- |
275
+ | `npm install` | Install dependencies. |
276
+ | `npm run typecheck` | TypeScript type check. |
277
+ | `npm run smoke` | Vendored prompt-builder smoke test; no network. |
278
+ | `npm test` | Unit tests using local/fake dependencies where provided. |
279
+ | `npm run cli -- --dry-run --workspace .` | Build this repository's system prompt. |
280
+ | `bash scripts/e2e/run.sh` | Phase 1 CLI integration against mock OpenRouter. |
281
+ | `bash scripts/e2e-phase2/run.sh` | Orchestrator spawn/watch/review integration. |
282
+ | `bash scripts/e2e-phase4/run.sh` | QA and deploy-gate integration with fake deploy. |
283
+ | `bash scripts/e2e-phase5/run.sh` | Watcher integration against fake GitHub and stubbed spawner. |
284
+ | `bash scripts/e2e-phase6/run.sh` | Budget-abort and concurrency-cap integration. |
285
+
286
+ Unit tests can inject a fake LLM client into `HeadlessSession`; the listed
287
+ end-to-end scripts use local mock services. OpenShell provider tests use a
288
+ fake CLI. A live OpenShell 0.1.2 run called DeepSeek V4 Flash through OpenRouter,
289
+ imported its Git result, and denied a host-worktree canary write. The earlier
290
+ shared-worktree mount was removed after a live probe showed it exposed host Git
291
+ and control files. See the [OpenShell guide](./docs/openshell-integration.md)
292
+ for the current clone-and-import boundary and its validation.
293
+
294
+ ## Implementation status
295
+
296
+ | Area | Status | Details |
297
+ | --- | --- | --- |
298
+ | Runtime engine | Implemented | Portable Zoo Code core, model client, tool executor, parser, CLI, and tests. |
299
+ | Orchestration | Implemented | Parallel worktrees, issue splitting, status tracking, review, and recovery. See [guide](./docs/phase2-orchestration.md). |
300
+ | Memory | Implemented | Opt-in per-project facts and session summaries with a local store; remote store remains a stub. |
301
+ | QA and deploy approval | Implemented | Read-only QA plus a fail-closed human-approval deploy gate. See [orchestration guide](./docs/phase2-orchestration.md) and [deploy gate](./docs/phase4-deploy-gate.md). |
302
+ | GitHub watcher | Implemented | Poll-based label intake; webhook server is not implemented. See [guide](./docs/phase5-issue-watcher.md). |
303
+ | Budget and concurrency controls | Implemented | Per-session limits and a cross-process concurrency cap. |
304
+ | Local process provider | Implemented | Existing host-process worker path behind the provider lifecycle interface. |
305
+ | Docker provider | Implemented | Container-backed provider implementation; `orchestrate` does not currently select it. |
306
+ | OpenShell provider | Implemented | Disposable clone, separate scratch data, validated Git bundle import, and sandbox teardown. See [guide](./docs/openshell-integration.md). |
307
+ | Context management | Implemented | Token-aware condensation summarizes older turns; message-count truncation remains the non-fatal fallback. See [`src/engine/condense.ts`](./src/engine/condense.ts). |
522
308
 
523
309
  ## Attribution
524
310