headlesscode 1.2.1 → 1.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (51) hide show
  1. package/README.md +272 -456
  2. package/package.json +8 -3
  3. package/shared/openshell/headlesscode-openrouter.yaml +21 -0
  4. package/shared/openshell/headlesscode-policy.yaml +21 -0
  5. package/src/cli.ts +77 -18
  6. package/src/cloud/openshell-preflight.ts +16 -0
  7. package/src/cloud/openshell-provider.ts +642 -0
  8. package/src/cloud/openshell-session.ts +119 -0
  9. package/src/cloud/openshell-subsession.ts +63 -0
  10. package/src/cloud/openshell-worker.ts +124 -0
  11. package/src/engine/events.ts +3 -0
  12. package/src/engine/loop.ts +143 -0
  13. package/src/engine/types.ts +42 -0
  14. package/src/llm/ollama.ts +25 -19
  15. package/src/monitoring/controller.ts +73 -0
  16. package/src/monitoring/features.ts +103 -0
  17. package/src/monitoring/index.ts +4 -0
  18. package/src/monitoring/predictor.ts +246 -0
  19. package/src/monitoring/types.ts +192 -0
  20. package/src/orchestrator/cli.ts +24 -0
  21. package/src/orchestrator/reviewer.ts +11 -0
  22. package/src/project-store.ts +14 -0
  23. package/src/qa/qa.ts +12 -0
  24. package/src/rsi/adaptive.ts +49 -0
  25. package/src/rsi/adversarial.ts +106 -0
  26. package/src/rsi/archive.ts +6 -5
  27. package/src/rsi/artifact-store.ts +158 -0
  28. package/src/rsi/config.ts +50 -2
  29. package/src/rsi/controller.ts +561 -42
  30. package/src/rsi/curriculum.ts +135 -16
  31. package/src/rsi/evaluator.ts +6 -37
  32. package/src/rsi/fitness.ts +39 -4
  33. package/src/rsi/index.ts +1 -0
  34. package/src/rsi/migrations/001_postgres_fleet_queue.sql +65 -0
  35. package/src/rsi/migrations/002_external_artifacts_and_job_leases.sql +39 -0
  36. package/src/rsi/migrations/003_model_training_jobs.sql +6 -0
  37. package/src/rsi/model-training.ts +256 -0
  38. package/src/rsi/mutation.ts +1 -77
  39. package/src/rsi/openshell.ts +639 -0
  40. package/src/rsi/postgres-queue.ts +424 -0
  41. package/src/rsi/promote-curriculum.ts +21 -0
  42. package/src/rsi/reports.ts +29 -2
  43. package/src/rsi/roles.ts +13 -3
  44. package/src/rsi/selection.ts +7 -1
  45. package/src/rsi/training-data.ts +103 -0
  46. package/src/rsi/trajectory.ts +1 -1
  47. package/src/rsi/types.ts +115 -2
  48. package/src/rsi/worker.ts +264 -0
  49. package/src/rsi/workspace.ts +14 -3
  50. package/src/watcher/cli.ts +18 -0
  51. package/src/watcher/watch.ts +3 -1
package/README.md CHANGED
@@ -2,523 +2,339 @@
2
2
 
3
3
  [![CI](https://github.com/Capsize-Games/headlesscode/actions/workflows/ci.yml/badge.svg)](https://github.com/Capsize-Games/headlesscode/actions/workflows/ci.yml)
4
4
  [![npm](https://img.shields.io/npm/v/headlesscode?logo=npm)](https://www.npmjs.com/package/headlesscode)
5
- [![Node.js >=18](https://img.shields.io/badge/node-%3E%3D18-339933?logo=node.js&logoColor=white)](https://nodejs.org/)
5
+ [![Node.js >=22.19](https://img.shields.io/badge/node-%3E%3D22.19-339933?logo=node.js&logoColor=white)](https://nodejs.org/)
6
6
  [![License: Apache-2.0](https://img.shields.io/badge/License-Apache--2.0-blue.svg)](./LICENSE)
7
7
 
8
- Run a coding agent where the work actually happens: in a terminal, a CI job,
9
- or an isolated git worktree. `headlesscode` brings the useful parts of
10
- [Zoo Code](https://github.com/Zoo-Code-Org/Zoo-Code) to a plain Node.js process,
11
- without requiring a VS Code window.
12
-
13
- It is made for real repositories and real engineering loops. The harness reads
14
- the target project's modes and rules, gives the model a controlled tool set,
15
- keeps work separated in git worktrees, and leaves behind logs and state that a
16
- person can inspect.
17
-
18
- ## Why use it?
19
-
20
- - **Run repeatable repository tasks.** Give it a task or GitHub issue and let
21
- it work through files, commands, and tests from a non-interactive process.
22
- - **Keep parallel work organized.** Split issues into worker groups, run them
23
- in separate worktrees, review the results, and optionally run QA before a
24
- human-approved deploy.
25
- - **Remember the project.** Opt-in local memory stores project facts and
26
- rolling session summaries, with a pluggable storage boundary for a future
27
- remote backend.
28
- - **Bound the expensive parts.** Per-session cost, duration, iteration, and
29
- fleet-concurrency limits are built into the orchestration and watcher paths.
30
- - **Experiment with improvement.** The `improve` command runs a bounded
31
- recursive self-improvement loop against an external evaluator, with the
32
- default worker using a local Qwen 3.5 9B model through Ollama.
33
-
34
- The project also includes a GitHub issue watcher and an evaluation-only cloud
35
- provider interface. No live cloud resources are launched by the current
36
- implementation.
37
-
38
- > **Security warning: default-allow arbitrary command execution.**
39
- > By default, `headlesscode` runs arbitrary shell commands with the invoking
40
- > user's privileges. It can read and modify files, including credentials such
41
- > as `~/.ssh` and `~/.aws`, without approval prompts. Use it only with trusted
42
- > tasks and isolate it with a container, VM, or dedicated user when untrusted
43
- > content is involved. The optional permissions layer is defense in depth, not
44
- > a security boundary. See [`SECURITY.md`](./SECURITY.md).
8
+ `headlesscode` runs the Zoo Code agent loop from a Node.js process, without a
9
+ VS Code window. Run one task directly, launch parallel workers in Git
10
+ worktrees, or use the NVIDIA OpenShell provider for sandboxed sessions.
11
+
12
+ | Use it for | What it does |
13
+ | --- | --- |
14
+ | One repository task | Reads project modes and instructions, edits files, runs commands, and reports a result. |
15
+ | Parallel issue work | Splits GitHub issues into Git worktrees, runs workers, monitors completion, and reviews changes. |
16
+ | Intake automation | Watches a GitHub label and starts bounded batches of work. |
17
+ | OpenShell sessions | Runs worker and review sessions in disposable clones, then validates and imports results after sandbox deletion. |
18
+ | Harness experiments | Runs a bounded recursive self-improvement loop against an external evaluator. |
19
+
20
+ ## Security
21
+
22
+ > **Default-allow arbitrary command execution.** By default, `headlesscode`
23
+ > runs shell commands with the invoking user's privileges. It can read and
24
+ > modify files, including credentials such as `~/.ssh` and `~/.aws`, without
25
+ > approval prompts. Use it only with trusted tasks and isolate it with a
26
+ > container, VM, or dedicated user when untrusted content is involved. The
27
+ > optional permissions layer is defense in depth, not a security boundary.
28
+ > See [`SECURITY.md`](./SECURITY.md).
29
+
30
+ OpenShell runs each session in a disposable Git clone mounted at `/workspace`;
31
+ the host worktree and harness control files are not mounted. After the sandbox
32
+ is deleted, HeadlessCode validates a Git bundle and imports the result into the
33
+ host worktree. The gateway must support Docker bind mounts and must see the
34
+ configured scratch root at the same absolute path as the host. See the
35
+ [OpenShell guide](./docs/openshell-integration.md) for setup and security
36
+ details. This integration does not configure NVIDIA Sentry or hardware
37
+ monitoring.
45
38
 
46
39
  ## Quick start
47
40
 
48
- Install from npm:
41
+ Requires Node.js 22.19 or newer. Install globally or run with `npx`:
49
42
 
50
43
  ```bash
51
44
  npm install -g headlesscode
45
+ # or
46
+ npx headlesscode --task "Fix the bug in src/index.ts" --workspace /path/to/repo
52
47
  ```
53
48
 
54
- Or run it without installing, via npx:
49
+ For a real model call, set the OpenRouter key. `--dry-run` does not need a key.
55
50
 
56
- ```bash
57
- npx headlesscode --task "Fix the bug in src/index.ts" --workspace /path/to/target/repo
58
- ```
51
+ | Variable | Required | Purpose |
52
+ | --- | --- | --- |
53
+ | `HEADLESSCODE_OPENROUTER_API_KEY` | For OpenRouter calls | OpenRouter API key. |
54
+ | `OPENROUTER_MODEL` | No | Default model; `deepseek/deepseek-v4-flash-0731`. `--model` overrides it. |
55
+ | `OPENROUTER_HTTP_REFERER` | No | OpenRouter application referer header. |
56
+ | `OPENROUTER_APP_TITLE` | No | OpenRouter application title header. |
57
+ | `HEADLESSCODE_WORKSPACE_ROOT` | No | Default workspace; otherwise the current directory. |
59
58
 
60
59
  ```bash
61
- # Required (except for --dry-run):
62
60
  export HEADLESSCODE_OPENROUTER_API_KEY=sk-or-...
63
-
64
- # Optional:
65
- export OPENROUTER_MODEL=deepseek/deepseek-v4-flash-0731 # default model
66
- export OPENROUTER_HTTP_REFERER=https://example.com # OpenRouter app header
67
- export OPENROUTER_APP_TITLE="headlesscode" # OpenRouter X-Title header
68
- export HEADLESSCODE_WORKSPACE_ROOT=/path/to/target/repo # default workspace root
69
-
70
- headlesscode --task "Fix the bug in src/index.ts" --workspace /path/to/target/repo
61
+ headlesscode --task "Fix the bug in src/index.ts" --workspace /path/to/repo
71
62
  ```
72
63
 
73
- To work from a source checkout instead (for contributing):
64
+ For a source checkout:
74
65
 
75
66
  ```bash
76
67
  git clone https://github.com/Capsize-Games/headlesscode.git
77
68
  cd headlesscode
78
69
  npm install
79
- node bin/headlesscode.mjs --task "Fix the bug in src/index.ts" --workspace /path/to/target/repo
70
+ node bin/headlesscode.mjs --task "Fix the bug in src/index.ts" --workspace /path/to/repo
80
71
  ```
81
72
 
82
- ### Systemwide `headlesscode` command
73
+ To install a `headlesscode` wrapper on your `PATH`, run `scripts/install-cli.sh`.
74
+ It writes to `~/.local/bin` by default (override with `HEADLESSCODE_BIN_DIR`)
75
+ and runs this checkout's local `tsx`. Re-run it if you move the checkout.
83
76
 
84
- To avoid re-typing `node <checkout>/bin/headlesscode.mjs` from every project,
85
- install a `headlesscode` command onto your `PATH` once:
77
+ ### Check project instructions without calling a model
78
+
79
+ `--dry-run` builds the system prompt and validates project configuration. It
80
+ prints the prompt and a summary of loaded modes, tools, and prompt size. A
81
+ successful exit means prompt building and mode/rules loading succeeded.
86
82
 
87
83
  ```bash
88
- scripts/install-cli.sh
84
+ headlesscode --dry-run --mode code --workspace /path/to/repo
89
85
  ```
90
86
 
91
- This writes a wrapper to `~/.local/bin/headlesscode` (override with
92
- `HEADLESSCODE_BIN_DIR`) that runs this checkout's `src/cli.ts` via its local
93
- `tsx`, without `cd`-ing — so `--repo`/`--workspace` still default to whatever
94
- directory you're standing in when you invoke it. Re-run the script any time
95
- after `git pull` to point it at a moved checkout; the wrapper itself doesn't
96
- need updating for ordinary code changes.
87
+ Project `.roomodes`, `.roo/rules-<slug>/`, and `AGENTS.md` instructions are
88
+ loaded by the prompt builder when applicable. The selected mode controls which
89
+ tools are exposed. See [Architecture](#architecture) for the relevant code.
90
+
91
+ ## OpenShell
92
+
93
+ Build the worker image and import the OpenRouter provider profile once. Configure
94
+ a Docker-backed OpenShell gateway with driver-config and bind-mount support;
95
+ keep that gateway restricted to trusted operators.
97
96
 
98
97
  ```bash
99
- # from any repo, no HEADLESSCODE_ROOT plumbing needed:
100
- headlesscode orchestrate --repo . --issue 42
98
+ # Run these commands from a HeadlessCode source checkout.
99
+ docker build -f docker/OpenShell.Dockerfile -t headlesscode-openshell:local .
100
+ openshell profile import -f shared/openshell/headlesscode-openrouter.yaml
101
+
102
+ export HEADLESSCODE_OPENROUTER_API_KEY=sk-or-...
103
+ export HEADLESSCODE_OPENSHELL_POLICY="$PWD/shared/openshell/headlesscode-policy.yaml"
104
+ headlesscode orchestrate --repo /path/to/repo --issue 42 \
105
+ --execution-provider openshell
101
106
  ```
102
107
 
103
- The rest of this README uses the plain `headlesscode` form for brevity —
104
- substitute `node <checkout>/bin/headlesscode.mjs` if you haven't run
105
- `scripts/install-cli.sh` yet.
108
+ Each worker gets an independent clone and a temporary result directory. Neither
109
+ bind source is the host worktree. After the sandbox is deleted, HeadlessCode
110
+ imports a verified Git bundle into the host branch. The default scratch root is
111
+ `~/.local/share/headlesscode/openshell-sessions`; create it before starting the
112
+ gateway and bind-mount it into the gateway container at the same absolute path.
113
+ If you choose another root, set `HEADLESSCODE_OPENSHELL_SCRATCH_ROOT` and mount
114
+ that host directory into the gateway container at the same path. Project data and shared instructions
115
+ are copied into the scratch tree and mounted separately. The writable scratch
116
+ checkpoint store supports checkpoints during a session and is discarded at
117
+ teardown. Opt-in memory uses a scratch snapshot and merges validated JSONL
118
+ records after sandbox deletion. No OpenRouter key is passed as a normal sandbox
119
+ environment argument; OpenShell injects the imported provider credential.
120
+
121
+ `headlesscode watch` accepts the same `--execution-provider openshell` option.
122
+ OpenShell applies to worker, plan-first, continuation, reviewer, QA, and rework
123
+ sessions. The one-off `headlesscode openshell-session` command runs a single
124
+ task and keeps its host worktree for review. Set up a gateway callback address
125
+ that the sandbox's Docker bridge can reach; keep the client endpoint loopback
126
+ only when possible. Read the [OpenShell guide](./docs/openshell-integration.md)
127
+ before adapting the sample gateway or filesystem policy.
128
+
129
+ | Setting | Default | Meaning |
130
+ | --- | --- | --- |
131
+ | `--execution-provider` | `local` | `orchestrate`/`watch` runtime: `local` or `openshell`. Also settable with `HEADLESSCODE_EXECUTION_PROVIDER`. |
132
+ | `HEADLESSCODE_OPENSHELL_POLICY` | No default | Base OpenShell policy YAML; required for OpenShell sessions. |
133
+ | `HEADLESSCODE_OPENSHELL_IMAGE` | `headlesscode-openshell:local` | Sandbox image. |
134
+ | `HEADLESSCODE_OPENSHELL_PROVIDERS` | `headlesscode-openrouter` | Comma-separated credential profile names. |
135
+ | `HEADLESSCODE_OPENSHELL_CPU` / `HEADLESSCODE_OPENSHELL_MEMORY` | `2` / `4Gi` | Per-sandbox resource requests. |
136
+ | `HEADLESSCODE_OPENSHELL_NAME_PREFIX` | `hcls` | Prefix for deterministic sandbox names. |
137
+ | `HEADLESSCODE_OPENSHELL_SCRATCH_ROOT` | `~/.local/share/headlesscode/openshell-sessions` | Host scratch directory that the gateway container must see at the same path. |
138
+
139
+ The `headlesscode openshell-session` command provides a single-session live
140
+ entry point; run `headlesscode openshell-session --help` for options.
141
+
142
+ ## Parallel work and GitHub intake
143
+
144
+ `orchestrate` runs a parallel round from GitHub issue numbers or a JSON file.
145
+ It splits work, creates a worktree per group, starts workers, watches for
146
+ completion, and reviews results. QA and deployment approval are optional.
106
147
 
107
- ## The `--dry-run` flow (no API key needed)
148
+ ```bash
149
+ headlesscode orchestrate --repo /path/to/repo --issue 42 --issue 43 --qa
150
+ headlesscode orchestrate status --repo /path/to/repo --wait --timeout-ms 7200000
151
+ ```
108
152
 
109
- `--dry-run` builds the full system prompt and validates configuration loading
110
- without calling the LLM — useful for CI and for checking that a target repo's
111
- `.roomodes` / `.roo/rules-<slug>/` / `AGENTS.md` are picked up:
153
+ `watch` polls GitHub for issues with a chosen label and starts bounded batches.
154
+ It requires `GH_TOKEN` or `GITHUB_TOKEN`; it performs intake only. Use a later
155
+ `orchestrate` run to monitor, review, QA, or deploy a batch.
112
156
 
113
157
  ```bash
114
- headlesscode --dry-run --mode code --workspace /path/to/target/repo
158
+ GH_TOKEN=... headlesscode watch \
159
+ --owner my-org --repo /path/to/repo --label needs-agent --run-once
115
160
  ```
116
161
 
117
- It prints the assembled system prompt plus a summary line (mode, custom modes
118
- loaded, exposed tools, prompt size). Exit code 0 means prompt building + mode /
119
- rules loading succeeded; non-zero means a config error.
120
-
121
- ## Real-usage example against a target repo
162
+ | Command | Important options | Behavior and limits |
163
+ | --- | --- | --- |
164
+ | `headlesscode orchestrate --repo <path>` | Repeat `--issue <n>` or use `--issues-json <file>`; `--execution-provider local|openshell`, `--qa`, `--deploy`, `--no-review`, `--dry-run`, `--max-iterations <n>`, `--max-rework-cycles <n>`, `--max-continuations <n>` | Worktree per group; default review is enabled. OpenShell applies to worker, plan-first, continuation, reviewer, QA, and rework sessions. `--deploy` uses the human-approval gate. |
165
+ | `headlesscode orchestrate status --repo <path>` | `--wait`, `--timeout-ms <n>`, `--json` | Reads durable round state; `--wait` blocks until terminal state or timeout. |
166
+ | `headlesscode orchestrate stop --repo <path>` | Repeat `--group <name>` | Stops selected groups' worker process trees. |
167
+ | `headlesscode watch --owner <o> --repo <path> --label <name>` | `--execution-provider local|openshell`, `--run-once`, `--dry-run`, `--max-per-sweep <n>`, `--max-concurrent-sessions <n>`, `--retry-failed` | Label-based intake. `--dry-run` still needs a GitHub token to list issues, but writes no state and starts no workers. |
168
+
169
+ Both orchestration and watcher enforce a shared concurrency cap, defaulting to
170
+ 3 (`HEADLESSCODE_MAX_CONCURRENT_SESSIONS`). Orchestration aborts when the cap
171
+ is full; the watcher leaves excess issues pending. The watcher records
172
+ idempotency state under `<repo>/.worktrees/.watcher-state.json` by default.
173
+ Use `--help` on a command for its complete options. Detailed guides:
174
+ [orchestration](./docs/phase2-orchestration.md),
175
+ [issue watcher](./docs/phase5-issue-watcher.md),
176
+ [QA implementation](./src/qa/qa.ts), and
177
+ [deployment gate](./docs/phase4-deploy-gate.md).
178
+
179
+ ## Project registration and search
180
+
181
+ Registering a project detects its stack, ensures `.gitignore` excludes
182
+ `.headlesscode/`, and builds a code-search index and codemap. Indexing calls an
183
+ embedding API and can incur charges unless skipped or configured for Ollama.
122
184
 
123
185
  ```bash
124
- export HEADLESSCODE_OPENROUTER_API_KEY=sk-or-...
125
- export HEADLESSCODE_WORKSPACE_ROOT=~/Projects/some-target-repo
126
-
127
- # Point the agent at an already-scoped issue, in an isolated worktree:
128
- headlesscode \
129
- --mode code \
130
- --task "Implement issue #29: add retry logic to the HTTP client (see .roo/rules for project conventions)." \
131
- --workspace ~/Projects/some-target-repo \
132
- --max-iterations 50 \
133
- --log-file ./headlesscode-session.log
186
+ headlesscode init --workspace ~/Projects/your-project
187
+ headlesscode init --workspace ~/Projects/your-project --skip-index
134
188
  ```
135
189
 
136
- The agent reads files, runs commands, writes code, and finishes by calling
137
- `attempt_completion` (or by giving a final text answer). The final result is
138
- printed to stdout. Exit code 0 = success; 1 = task failed (max iterations or
139
- consecutive-mistake limit); 2 = usage/config error (e.g. missing
140
- `HEADLESSCODE_OPENROUTER_API_KEY`).
190
+ See `headlesscode init --help` for `--skip-codemap` and embedding-backend
191
+ options. The CLI also provides `index`, `codemap`, project-store, GitHub App,
192
+ and decision-proxy commands; run `headlesscode --help` for the command list
193
+ and each subcommand's `--help` for complete usage.
141
194
 
142
195
  ## Recursive self-improvement
143
196
 
144
- `headlesscode improve` is a bounded research loop for improving the harness
145
- itself. A supervisor creates isolated candidate worktrees, asks the local Qwen
146
- 3.5 9B worker to make focused changes, runs regression plus visible and hidden
147
- evaluations, and keeps the archive, score, and selection decision outside the
148
- candidate worktree. Each generation also produces a report for human review.
197
+ `headlesscode improve --dry-run` resolves the base commit and prints planned
198
+ candidate worktrees. A real run requires PostgreSQL, a shared S3-compatible
199
+ artifact store, and at least one registered `headlesscode rsi-worker`. The
200
+ coordinator queues sanitized mutation snapshots and visible evaluation jobs;
201
+ workers run each in a fresh OpenShell guest. Fitness, paired baseline
202
+ comparison, archives, and selection stay in the coordinator. Hidden evaluation
203
+ remains disabled. A 2026-09-29 bounded run used two workers on the same local
204
+ OpenShell gateway; both candidate evaluation jobs were leased concurrently.
205
+ The run used a temporary `clamp.js` fixture and does not establish general
206
+ headlesscode improvement or multi-gateway operation. An earlier bounded run
207
+ passed visible checks but made no source change and was rejected.
149
208
 
150
209
  ```bash
151
- # Inspect the planned experiment without creating worktrees or calling a model:
152
210
  npx tsx src/cli.ts improve --repo . --dry-run
153
-
154
- # Run one small generation with the local Ollama model:
155
211
  npx tsx src/cli.ts improve --repo . --population 2 --generations 1
156
212
  ```
157
213
 
158
- This is an experiment, not unattended model training and not a security
159
- boundary. The current loop records model-candidate and training interfaces,
160
- trajectory datasets, parent-selection policies, and resource-tagged jobs, but
161
- does not yet train adapters, schedule multiple trajectories, or provide
162
- OS-level candidate isolation. Those limits are tracked in [the RSI roadmap](#rsi-roadmap).
163
-
164
- ## RSI roadmap
165
-
166
- The implemented loop is intentionally honest about what remains. Follow-up
167
- work is tracked in GitHub:
168
-
169
- - [OS-level candidate sandbox](https://github.com/Capsize-Games/headlesscode/issues/3)
170
- - [Cryptographically verifiable evaluator and artifacts](https://github.com/Capsize-Games/headlesscode/issues/4)
171
- - [Resource-aware resumable scheduler](https://github.com/Capsize-Games/headlesscode/issues/5)
172
- - [Adaptive multi-trajectory search](https://github.com/Capsize-Games/headlesscode/issues/6)
173
- - [Validated curriculum fixtures](https://github.com/Capsize-Games/headlesscode/issues/7)
174
- - [Adversarial evaluation and cross-model supervision](https://github.com/Capsize-Games/headlesscode/issues/8)
175
- - [Real LoRA or QLoRA backend](https://github.com/Capsize-Games/headlesscode/issues/9)
176
-
177
- See [`docs/recursive-self-improvement.md`](./docs/recursive-self-improvement.md)
178
- for the design and [`docs/rsi-progress.md`](./docs/rsi-progress.md) for the
179
- record of the first bounded runs.
180
-
181
- ## Registering a new project
182
-
183
- To use headlesscode against a project that isn't set up as a headlesscode
184
- project yet, register it in one step:
214
+ Before starting workers, build the standard and RSI guest images on each
215
+ OpenShell gateway host that will run RSI jobs. Build the larger training image
216
+ only on gateways whose workers will accept model-training or paired
217
+ model-evaluation jobs:
185
218
 
186
219
  ```bash
187
- headlesscode init --workspace ~/Projects/your-project
220
+ docker build -f docker/OpenShell.Dockerfile -t headlesscode-openshell:local .
221
+ docker build -f docker/OpenShell-RSI.Dockerfile -t headlesscode-openshell-rsi:local .
222
+ docker build -f docker/OpenShell-RSI-Training.Dockerfile -t headlesscode-openshell-rsi-training:local .
188
223
  ```
189
224
 
190
- This detects the project's stack(s) (which drives per-session instruction
191
- selection), makes sure the project's `.gitignore` excludes `.headlesscode/`
192
- session artifacts, and builds the codebase-search index and the codemap — no
193
- manual `index`/`codemap`/`.gitignore` steps needed. The index step calls the
194
- embedding API and costs real money unless you pass `--skip-index` or
195
- `--embedding-backend ollama`. See `headlesscode init --help` for the full
196
- options.
197
-
198
- ## CLI reference
199
-
200
- ```
201
- headlesscode --task "<task text>" [options]
202
- headlesscode --task-file <path> [options]
203
- headlesscode --dry-run [options]
204
- headlesscode orchestrate --repo <path> --issue <n> [--issue <n> ...] [options]
205
- headlesscode watch --owner <o> --repo <path> --label <name> [options]
206
-
207
- --mode <slug> Mode to run in (built-in or from .roomodes). Default: code
208
- --task <text> The task description for the agent
209
- --task-file <path> Read the task from a file (relative to workspace)
210
- --workspace <root> Workspace root (default: $HEADLESSCODE_WORKSPACE_ROOT or cwd)
211
- --model <id> OpenRouter model id (default: $OPENROUTER_MODEL or deepseek/deepseek-v4-flash-0731)
212
- --max-iterations <n> Loop iteration cap (default: 50)
213
- --consecutive-error-limit <n> Consecutive mistakes before giving up (default: 3)
214
- --max-cost-usd <n> Phase 6: per-session cost cap in USD (decimal). Default
215
- $HEADLESSCODE_MAX_COST_USD; off when neither is set
216
- --max-duration-ms <n> Phase 6: per-session wall-clock cap in ms. Default
217
- $HEADLESSCODE_MAX_DURATION_MS; off when neither is set.
218
- A tripped cap aborts with reason "budget"
219
- --log-file <path> Also append structured logs to this file
220
- --memory-dir <path> Phase 3 memory: store facts + session summaries under <path>
221
- (enabled; default $HEADLESSCODE_MEMORY_DIR or
222
- <workspace>/.headlesscode/memory). Memory is OFF unless set.
223
- --no-memory Explicitly disable memory even if HEADLESSCODE_MEMORY_DIR is set
224
- --allowed-commands <list> Comma-separated command prefixes the agent may run.
225
- Default: $HEADLESSCODE_ALLOWED_COMMANDS, else
226
- .headlesscode/permissions.json, else empty (=
227
- allow everything except --denied-commands; see
228
- SECURITY.md)
229
- --denied-commands <list> Comma-separated command prefixes that are ALWAYS
230
- refused (deny wins over allow; dangerous shell
231
- substitutions are always blocked regardless).
232
- Default: $HEADLESSCODE_DENIED_COMMANDS, else
233
- .headlesscode/permissions.json, else empty
234
- --protected-files <list> Comma-separated glob patterns of files the agent may
235
- not write. Default: $HEADLESSCODE_PROTECTED_FILES,
236
- else .headlesscode/permissions.json, else
237
- ".env,.env.*,*.pem,*.key,id_rsa*"
238
- --allow-protected-writes Escape hatch: permit writes to protected files
239
- (default: OFF). Also settable via
240
- "allowProtectedWrites": true in
241
- .headlesscode/permissions.json
242
- --dry-run Build the system prompt + validate config, then exit (no API key)
243
- --version / --help
244
-
245
- orchestrate subcommand (Phase 2 — parallel worktrees, headless workers):
246
- --repo <path> Target repo root (required)
247
- --issue <n> Issue number to include (repeatable)
248
- --issues-json <file> Read issues from a JSON array of {number,title,body}
249
- (used when gh is unavailable, or for tests)
250
- --file-issues With --issues-json: file a REAL GitHub issue for
251
- each synthetic entry (gh issue create --repo
252
- <origin-owner>/<origin-repo>), swap in the real
253
- number returned, and print one confirmation line
254
- per created issue. A real, visible write to
255
- GitHub — opt-in, never automatic. Requires
256
- --issues-json
257
- --batch <name> Batch id in the state file (default round-<date>)
258
- --review-mode <slug> Mode slug for review sessions (default deepseek-reviewer)
259
- --no-review Spawn + watch only; skip the reviewer
260
- --qa Phase 4: run a headless QA session (--mode qa-agent)
261
- on each group after its review passes; record
262
- qa {status,verdict,evidence} in the state file
263
- --qa-mode <slug> Mode slug for QA sessions (default qa-agent; the
264
- target repo's .roomodes + .roo/rules-<slug>/ are
265
- spliced automatically)
266
- --deploy Phase 4: after all groups done + reviewed + QA passed,
267
- run the human-approval deploy gate
268
- (scripts/deploy-gate.sh) — a hard stop that never runs
269
- the repo's deploy-production.sh without explicit human
270
- approval (interactive on a TTY, token/file otherwise)
271
- --deploy-args <str> Deploy args forwarded to the deploy script after the
272
- gate approves (space-separated flags; also DEPLOY_ARGS env)
273
- --poll-interval-ms <n> Watcher poll interval (default 5000)
274
- --max-concurrent-sessions <n> Phase 6 global cap on concurrent sessions across
275
- processes (default $HEADLESSCODE_MAX_CONCURRENT_
276
- SESSIONS or 3). At/over the cap this run ABORTS
277
- with a clear message, exit 1
278
- --dry-run Print the split plan + spawn commands, spawn nothing
279
-
280
- watch subcommand (Phase 5 — GitHub issue watcher, poll-based intake):
281
- --owner <o> GitHub owner (required)
282
- --repo <path> Local clone of the target repo (required; worktrees
283
- are spawned under <path>/.worktrees/). The GitHub repo
284
- name defaults to the directory basename (--gh-repo overrides)
285
- --label <name> The label that triggers processing, e.g. needs-agent
286
- --poll-interval-ms <n> Sweep interval in continuous mode (default 60000)
287
- --run-once One sweep then exit 0 (or 1 if any spawn failed)
288
- --max-per-sweep <n> Max NEW issues spawned per sweep (default 5); the rest
289
- stay 'pending' in state and are picked up next sweep
290
- --max-concurrent-sessions <n> Phase 6 GLOBAL cap on concurrent sessions across
291
- processes (default $HEADLESSCODE_MAX_CONCURRENT_
292
- SESSIONS or 3). Interplay: maxPerSweep bounds one
293
- sweep's burst; this bounds the total fleet — issues
294
- beyond it stay 'pending' until slots free up
295
- --state-file <path> Durable idempotency state (default
296
- <repo>/.worktrees/.watcher-state.json)
297
- --mode <slug> / --memory-dir <path>
298
- Forwarded to the spawner (ORCHESTRATOR_MODE /
299
- HEADLESSCODE_MEMORY_DIR)
300
- --qa / --deploy Pass-through: recorded per batch for the follow-up
301
- orchestrate completion run
302
- --dry-run Sweep + print the spawn plan, spawn nothing, write no
303
- state (still needs a GitHub token for listIssues)
304
- --retry-failed Retry previously-failed spawns next sweep
305
- ```
225
+ The RSI Dockerfile extends `headlesscode-openshell:local` and removes the RSI
226
+ source, test files, and hidden evaluation suite from the guest image. See the [RSI design](./docs/recursive-self-improvement.md)
227
+ for worker and queue configuration.
228
+
229
+ The queue stores job state, worker registrations, leases, retries, and admission
230
+ policy in PostgreSQL; artifact bytes remain in object storage. Local integration
231
+ tests used 12 simulated workers and 100 queued jobs, including a 20 MiB artifact. Each
232
+ worker process currently runs one guest at a time and must be configured with a
233
+ unique worker ID and OpenShell gateway ID. The two-worker run observed 4.008
234
+ seconds of concurrent evaluation leases on one gateway. A local PostgreSQL/
235
+ RustFS test drained the remaining 96 jobs with 12 simulated workers in 446 ms
236
+ (215.2 jobs/s), after the initial four jobs were claimed and completed to
237
+ verify admission limits;
238
+ claims still serialize on a queue-policy row, and neither result qualifies a
239
+ multi-host deployment. Remaining work is:
240
+
241
+ | Issue | Work |
242
+ | --- | --- |
243
+ | [#3](https://github.com/Capsize-Games/headlesscode/issues/3) | OpenShell candidate isolation is implemented; live multi-gateway qualification remains. |
244
+ | [#4](https://github.com/Capsize-Games/headlesscode/issues/4) | Content-addressed artifacts are implemented; signed evaluator provenance and hidden evaluation remain. |
245
+ | [#5](https://github.com/Capsize-Games/headlesscode/issues/5) | PostgreSQL queue, leases, retries, admission, and configured workers are implemented; live fleet qualification remains. |
246
+ | [#6](https://github.com/Capsize-Games/headlesscode/issues/6) | Bounded adaptive independent search supports an initial population plus one evidence-driven follow-up; repeated stages remain out of scope. |
247
+ | [#7](https://github.com/Capsize-Games/headlesscode/issues/7) | Fixture-backed curriculum replay and explicit promotion are implemented; wording-to-capability measurement remains unproven. |
248
+ | [#8](https://github.com/Capsize-Games/headlesscode/issues/8) | Adversarial review and bounded OpenShell break tests are implemented; live provider-backed review is unvalidated. |
249
+ | [#9](https://github.com/Capsize-Games/headlesscode/issues/9) | An OpenShell QLoRA prototype is covered by mocked job tests; the supplied GGUF is rejected before enqueue, and no live training is validated. |
250
+
251
+ See the [RSI design](./docs/recursive-self-improvement.md) and
252
+ [run progress](./docs/rsi-progress.md).
253
+
254
+ ## Core CLI options
255
+
256
+ | Option | Default or source | Purpose |
257
+ | --- | --- | --- |
258
+ | `--task <text>` / `--task-file <path>` | One is required unless `--dry-run` | Task prompt; task file is resolved relative to the workspace. |
259
+ | `--workspace <path>` | `HEADLESSCODE_WORKSPACE_ROOT` or current directory | Repository directory exposed to the worker. |
260
+ | `--mode <slug>` | `code` | Built-in or project `.roomodes` mode. |
261
+ | `--model <id>` | `OPENROUTER_MODEL` or `deepseek/deepseek-v4-flash-0731` | Model ID. |
262
+ | `--max-iterations <n>` | 250 for direct CLI sessions; orchestrator workers default to 50 | Per-session loop bound. |
263
+ | `--consecutive-error-limit <n>` | 3 (local code backend may use 6) | Consecutive tool/model errors before stopping. |
264
+ | `--max-cost-usd <n>` / `--max-duration-ms <n>` | `HEADLESSCODE_MAX_COST_USD` / `HEADLESSCODE_MAX_DURATION_MS`; disabled if unset | Per-session budget caps. A tripped cap aborts the session. |
265
+ | `--memory-dir <path>` / `--no-memory` | Off unless `HEADLESSCODE_MEMORY_DIR` or option is set | Enable project facts and rolling session summaries, or explicitly disable. |
266
+ | `--allowed-commands <list>` / `--denied-commands <list>` | CLI, environment, `.headlesscode/permissions.json`; allow list is empty by default | Command-prefix policy. Denials win; dangerous shell substitutions are blocked. See [`SECURITY.md`](./SECURITY.md). |
267
+ | `--protected-files <globs>` / `--allow-protected-writes` | `.env,.env.*,*.pem,*.key,id_rsa*` | Block writes matching protected files; explicit escape hatch. |
268
+ | `--dry-run` | Off | Build prompt and validate configuration without a model call or API key. |
269
+ | `--log-file <path>` | No file | Also append structured session logs to a file. |
270
+
271
+ Run `headlesscode --help` for advanced controls such as context condensation,
272
+ checkpoints, recursive delegation, session pause, and LLM timeout. Some options
273
+ are backend-specific; help output and source defaults are authoritative.
306
274
 
307
275
  ## Architecture
308
276
 
309
- - [`src/llm/openrouter.ts`](./src/llm/openrouter.ts) — OpenRouter
310
- chat-completions client using native `fetch` (no axios/node-fetch/openai
311
- SDK). Reads `HEADLESSCODE_OPENROUTER_API_KEY`, optional `OPENROUTER_HTTP_REFERER` /
312
- `OPENROUTER_APP_TITLE`; model default `deepseek/deepseek-v4-flash-0731` (env
313
- `OPENROUTER_MODEL` overrides). Non-2xx → typed `OpenRouterError` with status +
314
- body excerpt; supports `AbortSignal` timeouts.
315
- - [`src/tools/executor.ts`](./src/tools/executor.ts) — headless executor for
316
- `read_file`, `write_to_file`, `execute_command`, `list_files` (plain
317
- `fs`/`child_process`), plus `attempt_completion` / `ask_followup_question`
318
- handlers and "not implemented" stubs for every other vendored tool schema.
319
- All file operations are resolved relative to the workspace root and rejected
320
- if they escape it (path-traversal guard via `path.resolve` + containment
321
- check). Command results are truncated (~30k chars) to keep context bounded.
322
- - [`src/tools/output-summarizer.ts`](./src/tools/output-summarizer.ts) —
323
- OPT-IN local summarization of oversized `execute_command` output: when
324
- `HEADLESSCODE_LOCAL_SUMMARIZATION=1`, a result that would exceed the 30k
325
- char cap is compressed by a local Ollama chat model before reaching the
326
- cloud model, with a `[Output summarized by local model…]` transparency
327
- header. OFF by default; on any failure it falls back to today's exact blunt
328
- truncation (never an error, never a hang). Endpoint/model configurable via
329
- `HEADLESSCODE_OLLAMA_URL` (default `http://localhost:11434`) and
330
- `HEADLESSCODE_SUMMARIZATION_MODEL` (default `qwen3:8b`). Deliberately
331
- limited to command output — file/diff content always stays verbatim.
332
- - [`src/engine/parser.ts`](./src/engine/parser.ts) — OpenAI function-calling
333
- parser: JSON.parses `tool_calls[].function.arguments` with a best-effort
334
- partial-JSON fallback; parse failures are marked and fed back as errors.
335
- - [`src/engine/prompt.ts`](./src/engine/prompt.ts) — wraps the vendored
336
- `SYSTEM_PROMPT` builder: loads project `.roomodes` (same zod schema as Zoo
337
- Code's `CustomModesManager`), passes the workspace as `cwd` so `.roo/rules-*`
338
- / `AGENTS.md` splice in, and selects the mode's exposed tools.
339
- - [`src/engine/loop.ts`](./src/engine/loop.ts) — `HeadlessSession`, the
340
- orchestration loop. System + user → LLM → assistant (with tool_calls) → parse
341
- → execute → `tool` role message → repeat. Terminates on `attempt_completion`
342
- (its `args.result` is the final answer) or a text-only reply; fails bounded
343
- on max iterations or `consecutiveErrorLimit` consecutive mistakes (tool
344
- errors / parse errors / identical repeated calls). History truncation is a
345
- Phase 1 placeholder: system + first user always kept, sliding window of the
346
- last ~40 messages. The loop accepts an injected `llmClient` (DI) so tests use
347
- a fake; the CLI wires `OpenRouterClient`.
348
- - [`src/engine/logger.ts`](./src/engine/logger.ts) — structured logger
349
- (timestamped lines to stdout/stderr, optional file).
350
- - [`src/memory/`](./src/memory/index.ts) — memory subsystem:
351
- - [`src/memory/types.ts`](./src/memory/types.ts) — the `MemoryFact` /
352
- `SessionSummary` schema (per-project scoped, kinds
353
- `convention|decision|failure|knowledge` where "things that didn't work"
354
- are `failure`) and the two contracts: `MemoryStore` (the pluggable
355
- storage boundary) and `Embedder`. Hard data-isolation requirement
356
- documented: the harness knowledge is a dedicated schema, never reachable
357
- through any customer tenant route.
358
- - [`src/memory/local.ts`](./src/memory/local.ts) — `LocalMemoryStore`: the
359
- fully working file backend (`facts/<project>.jsonl` +
360
- `sessions/<project>.jsonl`, append-only, idempotent `addFact` by content
361
- hash, `queryRecall` = keyword matches (high weight) + local-embedder
362
- cosine similarity, deterministic ordering).
363
- - [`src/memory/embed.ts`](./src/memory/embed.ts) — `createLocalEmbedder()`:
364
- zero-dependency, deterministic lexical-hash embedder (lowercase word +
365
- char-bigram tokens → fixed-dim L2-normalized vector). Placeholder for a
366
- real local embedding model behind the same `Embedder` interface.
367
- - [`src/memory/summarizer.ts`](./src/memory/summarizer.ts) —
368
- `extractSessionSummary` (deterministic; files/commands derived from the
369
- tool history, facts via keyword heuristics) + `buildRollingSummary`
370
- (compact markdown recap of the last N sessions, so a session never needs
371
- the infinite raw history).
372
- - [`src/memory/uwuchat.ts`](./src/memory/uwuchat.ts) — a remote
373
- `MemoryStore` implementation stub for a future hosted memory API. Throws
374
- "not implemented" until its base-URL/token env vars are set.
375
- - Memory is wired into `HeadlessSession` as an **opt-in** config (`memory`
376
- + `project`); when unset the loop behaves exactly as before. When set, the
377
- loop injects a `## PROJECT MEMORY` section (recalled facts + rolling
378
- recap) into the first user message and records the session + extracted
379
- facts afterwards — and memory failures are always non-fatal.
380
- - [`src/orchestrator/`](./src/orchestrator/index.ts) — Phase 2 orchestration
381
- layer: `split.ts` (issue-splitting heuristics, deterministic +
382
- unit-tested), `state.ts` (`.worktrees/.orchestrator-state.json` read/write),
383
- `reviewer.ts` (adversarial fresh-context review run with a read-only
384
- executor), `watch.ts` (completion polling of `.harness.done` markers + stall
385
- guard), and `cli.ts` (the `orchestrate` subcommand — split → spawn via
386
- `scripts/spawn-parallel-worktrees.sh` → watch → review → QA → deploy gate).
387
- - [`src/qa/qa.ts`](./src/qa/qa.ts) — Phase 4 headless QA: `runQa()` runs a
388
- second harness session against a worktree in the target repo's `qa-agent`
389
- mode (auto-spliced from `.roomodes` + `.roo/rules-qa-agent/`), with a
390
- generic checklist fallback when the repo has no such mode. Read + command
391
- tools only (no `write_to_file`) — QA verifies and reports, it never edits.
392
- Verdict parsing is fail-closed (`pass`/`fail`/`error`, default `fail`).
393
- - [`src/deploy/gate.ts`](./src/deploy/gate.ts) + [`src/deploy/gate-cli.ts`](./src/deploy/gate-cli.ts) —
394
- Phase 4 human-approval deploy gate: the pure, unit-tested decision function
395
- `decideApproval` (interactive y/N, one-time approval file, or
396
- `DEPLOY_APPROVAL_TOKEN` matching `<repo>/.deploy-approval`; never
397
- auto-approves) plus a thin CLI the bash wrapper calls.
398
- - [`src/watcher/`](./src/watcher/index.ts) — Phase 5 GitHub issue watcher:
399
- [`github.ts`](./src/watcher/github.ts) (native-fetch GitHub REST client with
400
- label filter, PR filtering, pagination, `GITHUB_API_BASE_URL` override for
401
- tests/mocks), [`state.ts`](./src/watcher/state.ts) (durable idempotency
402
- state file — write-ahead `spawned` → `done`/`failed`, `pending` for capped
403
- issues, restart-safe), [`watch.ts`](./src/watcher/watch.ts) (the poll loop:
404
- list by label → split → spawn via the existing bash spawner, bounded by
405
- `maxPerSweep` and the Phase 6 global cap), and [`cli.ts`](./src/watcher/cli.ts)
406
- (the `watch` subcommand — continuous or `--run-once`, `--dry-run`,
407
- `--retry-failed`).
408
- - [`src/budget/`](./src/budget/index.ts) — Phase 6 guardrails:
409
- [`cost.ts`](./src/budget/cost.ts) (model pricing table + `estimateCost`,
410
- `HEADLESSCODE_PRICING_JSON` override, conservative fallback for unlisted
411
- models), [`budget.ts`](./src/budget/budget.ts) (`SessionBudget` +
412
- `BudgetTracker`: `tick()` before each LLM call, `record()` after with usage
413
- tokens, `check()` snapshot, `BudgetExceededError`), and
414
- [`concurrency.ts`](./src/budget/concurrency.ts) (`ConcurrencyLimiter` —
415
- fail-fast in-process semaphore — plus `activeSessionCount` reading the
416
- durable orchestrator/watcher state files for a cross-process view). Wired
417
- into `HeadlessSession` (`budget` config, `budgetUsage` on results), the base
418
- CLI (`--max-cost-usd` / `--max-duration-ms`), `orchestrate` (aborts at the
419
- cap), the watcher (defers cap-exceeding issues to `pending`), and
420
- `run-worker.sh`/`run-qa.sh` (env forwarding).
421
- - [`src/cloud/`](./src/cloud/provider.ts) — Phase 6 ephemeral compute
422
- abstraction: the `CloudProvider` lifecycle interface
423
- (`spawnWorktreeSession` → `waitReady` → `runHarness` → `collectResults` →
424
- `teardown`) with `LocalProcessProvider` as the current local behavior behind
425
- it (reuses `spawn-parallel-worktrees.sh` + `run-worker.sh`), so a
426
- container/VM-per-issue backend slots in without touching the orchestration
427
- layer. A cloud-provider sketch is documented (evaluation only — no
428
- live setup; see [`docs/phase6-cloud.md`](./docs/phase6-cloud.md)).
429
- - [`src/cli.ts`](./src/cli.ts) — the `headlesscode` bin entry (+ `orchestrate`
430
- and `watch` subcommand dispatch).
431
- - [`scripts/run-worker.sh`](./scripts/run-worker.sh) — launches one headless
432
- harness worker per worktree (pid, log, exit code, `.harness.done` marker).
433
- - [`scripts/run-qa.sh`](./scripts/run-qa.sh) — Phase 4 QA wrapper mirroring
434
- run-worker.sh: launches one harness QA session per worktree (`.qa-task.md`,
435
- `.qa.pid`, `qa.log`, `.qa.exit`, `.qa.done/`).
436
- - [`scripts/deploy-gate.sh`](./scripts/deploy-gate.sh) — Phase 4 gate wrapper:
437
- path safety, deployment summary, interactive + token/file approval, then
438
- (and only then) invokes the repo's `scripts/deploy-production.sh` with
439
- forwarded deploy args. Exit 3 = human DENIED (hard stop).
440
- - [`scripts/spawn-parallel-worktrees.sh`](./scripts/spawn-parallel-worktrees.sh)
441
- — spawns one git worktree + harness worker per group (worktree/.env/branch
442
- conventions, `run-worker.sh` + state-file writes, no GUI involved).
443
-
444
- ### Phase 1 tool filtering decision
445
-
446
- The loop exposes to the model exactly the tools the executor can actually run
447
- for the selected mode: the intersection of the vendored mode tool groups
448
- (`getToolsForMode`) with the Phase 1 executable set
449
- (`read_file`, `write_to_file`, `execute_command`, `list_files`,
450
- `attempt_completion`, `ask_followup_question`). Stub-only tools (`apply_diff`,
451
- `search_files`, …) stay registered in the executor purely as a safety net
452
- (clear "not implemented" error) but are NOT advertised to the model, so it
453
- doesn't waste turns calling them.
277
+ The core loop builds the project-specific prompt, calls the configured model,
278
+ executes the selected mode's available tools, and repeats until completion or a
279
+ configured limit. File tools are confined to the workspace path; shell commands
280
+ run with the process user's privileges unless an external boundary such as
281
+ OpenShell is configured. Mode instructions and safety controls are distinct:
282
+ prompt instructions guide the model, while filesystem or command restrictions
283
+ must be enforced by the runtime or sandbox.
284
+
285
+ | Area | Entry points | Responsibility |
286
+ | --- | --- | --- |
287
+ | CLI and session loop | [`src/cli.ts`](./src/cli.ts), [`src/engine/`](./src/engine/) | Command dispatch, prompt construction, LLM/tool loop, iteration and context limits. |
288
+ | Model and tools | [`src/llm/`](./src/llm/), [`src/tools/`](./src/tools/) | OpenRouter client, tool execution, output handling, and workspace path checks. |
289
+ | Parallel orchestration | [`src/orchestrator/`](./src/orchestrator/), [`scripts/spawn-parallel-worktrees.sh`](./scripts/spawn-parallel-worktrees.sh) | Issue grouping, worktree workers, completion state, review, rework, QA, and deploy gating. |
290
+ | GitHub intake | [`src/watcher/`](./src/watcher/) | Label polling, durable idempotency, pending work, and bounded spawning. |
291
+ | Runtime providers | [`src/cloud/`](./src/cloud/) | Local process, Docker, and OpenShell implementations of the session lifecycle. `orchestrate` and `watch` can select OpenShell for worker sessions. |
292
+ | Memory and budgets | [`src/memory/`](./src/memory/), [`src/budget/`](./src/budget/) | Opt-in project memory, cost/time budgets, and session concurrency accounting. |
293
+ | QA and deploy gate | [`src/qa/`](./src/qa/), [`src/deploy/`](./src/deploy/) | Read-only QA verdicts and explicit human approval before deployment. |
294
+ | Project map and search | [`src/codemap/`](./src/codemap/), [`src/codesearch/`](./src/codesearch/) | Deterministic module maps and repository search indexes. |
295
+
296
+ The model only receives tools supported by the selected mode and executable
297
+ runtime. Unsupported vendored tool schemas are not advertised. This keeps the
298
+ model's action set aligned with the current executor; it does not make arbitrary
299
+ shell commands safe by itself.
454
300
 
455
301
  ## Development
456
302
 
457
- ```bash
458
- npm run typecheck # npx tsc --noEmit (whole repo incl. vendored core)
459
- npm run smoke # vendored prompt builder smoke test (no network)
460
- npm test # unit tests with fake LLM clients (no network/key)
461
- bash scripts/e2e/run.sh # Phase 1 integration (mock OpenRouter)
462
- bash scripts/e2e-phase2/run.sh # Phase 2 integration (spawn + watch + review)
463
- bash scripts/e2e-phase4/run.sh # Phase 4 integration (QA + deploy gate, fake deploy)
464
- bash scripts/e2e-phase5/run.sh # Phase 5 integration (watcher vs fake GitHub server,
465
- # stubbed spawner: state transitions + no double-spawn)
466
- bash scripts/e2e-phase6/run.sh # Phase 6 integration (budget abort via mock OpenRouter
467
- # + concurrency-cap sweep via fake GitHub + stubbed spawner)
468
- npm run cli -- --dry-run --workspace . # build this repo's system prompt
469
- ```
470
-
471
- The engine tests (`src/engine/__tests__/loop.test.ts`) run the full loop with a
472
- fake `LlmClient` injected via the `HeadlessSession` constructor — no network,
473
- no API key required. Phase 2 adds `src/orchestrator/__tests__/` (split
474
- heuristics, state round-trip, reviewer verdict parsing), and the e2e scripts
475
- drive the real CLI through a local mock OpenRouter server, including a
476
- 2-worktree parallel spawn, completion-marker polling, and a read-only review
477
- invocation. Phase 5 adds `src/watcher/__tests__/` (github client with an
478
- injected fetch, watcher-state idempotency semantics, and the watch loop with
479
- an injected gh client + spawner covering spawn/idempotency/cap/failure/
480
- dry-run/abort) and `scripts/e2e-phase5/run.sh` (watcher against a fake GitHub
481
- server with a stubbed spawner).
482
-
483
- ## Phase status
484
-
485
- - ✅ Phase 1 Subtask 1 — vendored portable Zoo Code core
486
- ([`src/vendor/zoo-code/`](./src/vendor/zoo-code/), read-only dependency).
487
- - ✅ Phase 1 Subtask 2 — runtime engine (OpenRouter client, tool executor,
488
- parser, orchestration loop, CLI, tests).
489
- - ✅ Phase 2 — headless orchestration layer: drop-in
490
- `spawn-parallel-worktrees.sh` + `run-worker.sh` (harness subprocess per
491
- worktree, `.harness.done` completion markers), issue-splitting heuristics
492
- port, `.orchestrator-state.json` state management, headless reviewer, and
493
- the `orchestrate` CLI subcommand (see
494
- [`docs/phase2-orchestration.md`](./docs/phase2-orchestration.md)).
495
- - ✅ Phase 3 — memory subsystem: per-project knowledge facts + rolling session
496
- summaries (`src/memory/`), local deterministic embedder for semantic recall,
497
- opt-in `HeadlessSession`/CLI wiring (`--memory-dir`, `--no-memory`), and a
498
- pluggable `MemoryStore` contract with a remote-backend client stub.
499
- - ✅ Phase 4 — QA + deploy gate: headless QA runs the target repo's `qa-agent`
500
- mode (`--qa` / `--qa-mode`), verdict parsing is fail-closed, results land in
501
- the state file's per-group `qa` field; the human-approval deploy gate
502
- (`--deploy`, `scripts/deploy-gate.sh` + `src/deploy/gate.ts`) is a hard stop
503
- in front of `deploy-production.sh` that never auto-approves (see
504
- [`docs/phase4-qa.md`](./docs/phase4-qa.md) and
505
- [`docs/phase4-deploy-gate.md`](./docs/phase4-deploy-gate.md)).
506
- - ✅ Phase 5 — GitHub issue watcher: poll-based intake (`watch` subcommand) —
507
- detect issues by label via the GitHub REST API (`GH_TOKEN`), fan each out
508
- through `splitIssues` + the existing `spawn-parallel-worktrees.sh`, track
509
- idempotency durably (`.worktrees/.watcher-state.json`, write-ahead
510
- ordering, `pending` cap deferral, restart-safe), optional `--dry-run` /
511
- `--run-once` / `--retry-failed`; webhook upgrade designed but not built as
512
- a server (see [`docs/phase5-issue-watcher.md`](./docs/phase5-issue-watcher.md)).
513
- - ✅ Phase 6 — cloud scaling, guardrails-first: per-session cost/time/iteration
514
- budget (`src/budget/`, `--max-cost-usd` / `--max-duration-ms`, budgetUsage on
515
- results, worker env forwarding) + a hard concurrent-session cap
516
- (`HEADLESSCODE_MAX_CONCURRENT_SESSIONS`, orchestrate aborts / watcher defers
517
- to `pending`) + the `CloudProvider` abstraction with `LocalProcessProvider`
518
- and an evaluation-only cloud-provider sketch (see
519
- [`docs/phase6-cloud.md`](./docs/phase6-cloud.md)). No live cloud launched;
520
- a container/VM-per-issue backend slots in behind the same interface.
521
- - ⏳ Phase 3 (remaining) — token-based condensation.
303
+ | Command | Purpose |
304
+ | --- | --- |
305
+ | `npm install` | Install dependencies. |
306
+ | `npm run typecheck` | TypeScript type check. |
307
+ | `npm run smoke` | Vendored prompt-builder smoke test; no network. |
308
+ | `npm test` | Unit tests using local/fake dependencies where provided. |
309
+ | `npm run cli -- --dry-run --workspace .` | Build this repository's system prompt. |
310
+ | `bash scripts/e2e/run.sh` | Phase 1 CLI integration against mock OpenRouter. |
311
+ | `bash scripts/e2e-phase2/run.sh` | Orchestrator spawn/watch/review integration. |
312
+ | `bash scripts/e2e-phase4/run.sh` | QA and deploy-gate integration with fake deploy. |
313
+ | `bash scripts/e2e-phase5/run.sh` | Watcher integration against fake GitHub and stubbed spawner. |
314
+ | `bash scripts/e2e-phase6/run.sh` | Budget-abort and concurrency-cap integration. |
315
+
316
+ Unit tests can inject a fake LLM client into `HeadlessSession`; the listed
317
+ end-to-end scripts use local mock services. OpenShell provider tests use a
318
+ fake CLI. A live OpenShell 0.1.2 run called DeepSeek V4 Flash through OpenRouter,
319
+ imported its Git result, and denied a host-worktree canary write. The earlier
320
+ shared-worktree mount was removed after a live probe showed it exposed host Git
321
+ and control files. See the [OpenShell guide](./docs/openshell-integration.md)
322
+ for the current clone-and-import boundary and its validation.
323
+
324
+ ## Implementation status
325
+
326
+ | Area | Status | Details |
327
+ | --- | --- | --- |
328
+ | Runtime engine | Implemented | Portable Zoo Code core, model client, tool executor, parser, CLI, and tests. |
329
+ | Orchestration | Implemented | Parallel worktrees, issue splitting, status tracking, review, and recovery. See [guide](./docs/phase2-orchestration.md). |
330
+ | Memory | Implemented | Opt-in per-project facts and session summaries with a local store; remote store remains a stub. |
331
+ | QA and deploy approval | Implemented | Read-only QA plus a fail-closed human-approval deploy gate. See [orchestration guide](./docs/phase2-orchestration.md) and [deploy gate](./docs/phase4-deploy-gate.md). |
332
+ | GitHub watcher | Implemented | Poll-based label intake; webhook server is not implemented. See [guide](./docs/phase5-issue-watcher.md). |
333
+ | Budget and concurrency controls | Implemented | Per-session limits and a cross-process concurrency cap. |
334
+ | Local process provider | Implemented | Existing host-process worker path behind the provider lifecycle interface. |
335
+ | Docker provider | Implemented | Container-backed provider implementation; `orchestrate` does not currently select it. |
336
+ | OpenShell provider | Implemented | Disposable clone, separate scratch data, validated Git bundle import, and sandbox teardown. See [guide](./docs/openshell-integration.md). |
337
+ | Context management | Implemented | Token-aware condensation summarizes older turns; message-count truncation remains the non-fatal fallback. See [`src/engine/condense.ts`](./src/engine/condense.ts). |
522
338
 
523
339
  ## Attribution
524
340