headlesscode 1.2.0 → 1.2.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +245 -459
- package/package.json +2 -1
- package/shared/openshell/headlesscode-openrouter.yaml +21 -0
- package/shared/openshell/headlesscode-policy.yaml +21 -0
- package/src/cli.ts +30 -0
- package/src/cloud/openshell-preflight.ts +16 -0
- package/src/cloud/openshell-provider.ts +582 -0
- package/src/cloud/openshell-session.ts +119 -0
- package/src/cloud/openshell-subsession.ts +63 -0
- package/src/cloud/openshell-worker.ts +124 -0
- package/src/engine/events.ts +3 -0
- package/src/engine/loop.ts +154 -4
- package/src/engine/types.ts +42 -0
- package/src/monitoring/controller.ts +73 -0
- package/src/monitoring/features.ts +103 -0
- package/src/monitoring/index.ts +4 -0
- package/src/monitoring/predictor.ts +246 -0
- package/src/monitoring/types.ts +192 -0
- package/src/orchestrator/cli.ts +24 -0
- package/src/orchestrator/reviewer.ts +11 -0
- package/src/project-store.ts +11 -0
- package/src/qa/qa.ts +12 -0
- package/src/watcher/cli.ts +18 -0
- package/src/watcher/watch.ts +3 -1
package/README.md
CHANGED
|
@@ -5,520 +5,306 @@
|
|
|
5
5
|
[](https://nodejs.org/)
|
|
6
6
|
[](./LICENSE)
|
|
7
7
|
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
> **Security warning: default-allow arbitrary command execution.**
|
|
39
|
-
> By default, `headlesscode` runs arbitrary shell commands with the invoking
|
|
40
|
-
> user's privileges. It can read and modify files, including credentials such
|
|
41
|
-
> as `~/.ssh` and `~/.aws`, without approval prompts. Use it only with trusted
|
|
42
|
-
> tasks and isolate it with a container, VM, or dedicated user when untrusted
|
|
43
|
-
> content is involved. The optional permissions layer is defense in depth, not
|
|
44
|
-
> a security boundary. See [`SECURITY.md`](./SECURITY.md).
|
|
8
|
+
`headlesscode` runs the Zoo Code agent loop from a Node.js process, without a
|
|
9
|
+
VS Code window. Run one task directly, launch parallel workers in Git
|
|
10
|
+
worktrees, or use the NVIDIA OpenShell provider for sandboxed sessions.
|
|
11
|
+
|
|
12
|
+
| Use it for | What it does |
|
|
13
|
+
| --- | --- |
|
|
14
|
+
| One repository task | Reads project modes and instructions, edits files, runs commands, and reports a result. |
|
|
15
|
+
| Parallel issue work | Splits GitHub issues into Git worktrees, runs workers, monitors completion, and reviews changes. |
|
|
16
|
+
| Intake automation | Watches a GitHub label and starts bounded batches of work. |
|
|
17
|
+
| OpenShell sessions | Runs worker and review sessions in disposable clones, then validates and imports results after sandbox deletion. |
|
|
18
|
+
| Harness experiments | Runs a bounded recursive self-improvement loop against an external evaluator. |
|
|
19
|
+
|
|
20
|
+
## Security
|
|
21
|
+
|
|
22
|
+
> **Default-allow arbitrary command execution.** By default, `headlesscode`
|
|
23
|
+
> runs shell commands with the invoking user's privileges. It can read and
|
|
24
|
+
> modify files, including credentials such as `~/.ssh` and `~/.aws`, without
|
|
25
|
+
> approval prompts. Use it only with trusted tasks and isolate it with a
|
|
26
|
+
> container, VM, or dedicated user when untrusted content is involved. The
|
|
27
|
+
> optional permissions layer is defense in depth, not a security boundary.
|
|
28
|
+
> See [`SECURITY.md`](./SECURITY.md).
|
|
29
|
+
|
|
30
|
+
OpenShell runs each session in a disposable Git clone mounted at `/workspace`;
|
|
31
|
+
the host worktree and harness control files are not mounted. After the sandbox
|
|
32
|
+
is deleted, HeadlessCode validates a Git bundle and imports the result into the
|
|
33
|
+
host worktree. The gateway must support Docker bind mounts and must see the
|
|
34
|
+
configured scratch root at the same absolute path as the host. See the
|
|
35
|
+
[OpenShell guide](./docs/openshell-integration.md) for setup and security
|
|
36
|
+
details. This integration does not configure NVIDIA Sentry or hardware
|
|
37
|
+
monitoring.
|
|
45
38
|
|
|
46
39
|
## Quick start
|
|
47
40
|
|
|
48
|
-
Install
|
|
41
|
+
Requires Node.js 18 or newer. Install globally or run with `npx`:
|
|
49
42
|
|
|
50
43
|
```bash
|
|
51
44
|
npm install -g headlesscode
|
|
45
|
+
# or
|
|
46
|
+
npx headlesscode --task "Fix the bug in src/index.ts" --workspace /path/to/repo
|
|
52
47
|
```
|
|
53
48
|
|
|
54
|
-
|
|
49
|
+
For a real model call, set the OpenRouter key. `--dry-run` does not need a key.
|
|
55
50
|
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
51
|
+
| Variable | Required | Purpose |
|
|
52
|
+
| --- | --- | --- |
|
|
53
|
+
| `HEADLESSCODE_OPENROUTER_API_KEY` | For OpenRouter calls | OpenRouter API key. |
|
|
54
|
+
| `OPENROUTER_MODEL` | No | Default model; `deepseek/deepseek-v4-flash-0731`. `--model` overrides it. |
|
|
55
|
+
| `OPENROUTER_HTTP_REFERER` | No | OpenRouter application referer header. |
|
|
56
|
+
| `OPENROUTER_APP_TITLE` | No | OpenRouter application title header. |
|
|
57
|
+
| `HEADLESSCODE_WORKSPACE_ROOT` | No | Default workspace; otherwise the current directory. |
|
|
59
58
|
|
|
60
59
|
```bash
|
|
61
|
-
# Required (except for --dry-run):
|
|
62
60
|
export HEADLESSCODE_OPENROUTER_API_KEY=sk-or-...
|
|
63
|
-
|
|
64
|
-
# Optional:
|
|
65
|
-
export OPENROUTER_MODEL=deepseek/deepseek-v4-flash-0731 # default model
|
|
66
|
-
export OPENROUTER_HTTP_REFERER=https://example.com # OpenRouter app header
|
|
67
|
-
export OPENROUTER_APP_TITLE="headlesscode" # OpenRouter X-Title header
|
|
68
|
-
export HEADLESSCODE_WORKSPACE_ROOT=/path/to/target/repo # default workspace root
|
|
69
|
-
|
|
70
|
-
headlesscode --task "Fix the bug in src/index.ts" --workspace /path/to/target/repo
|
|
61
|
+
headlesscode --task "Fix the bug in src/index.ts" --workspace /path/to/repo
|
|
71
62
|
```
|
|
72
63
|
|
|
73
|
-
|
|
64
|
+
For a source checkout:
|
|
74
65
|
|
|
75
66
|
```bash
|
|
76
67
|
git clone https://github.com/Capsize-Games/headlesscode.git
|
|
77
68
|
cd headlesscode
|
|
78
69
|
npm install
|
|
79
|
-
node bin/headlesscode.mjs --task "Fix the bug in src/index.ts" --workspace /path/to/
|
|
70
|
+
node bin/headlesscode.mjs --task "Fix the bug in src/index.ts" --workspace /path/to/repo
|
|
80
71
|
```
|
|
81
72
|
|
|
82
|
-
|
|
73
|
+
To install a `headlesscode` wrapper on your `PATH`, run `scripts/install-cli.sh`.
|
|
74
|
+
It writes to `~/.local/bin` by default (override with `HEADLESSCODE_BIN_DIR`)
|
|
75
|
+
and runs this checkout's local `tsx`. Re-run it if you move the checkout.
|
|
83
76
|
|
|
84
|
-
|
|
85
|
-
install a `headlesscode` command onto your `PATH` once:
|
|
77
|
+
### Check project instructions without calling a model
|
|
86
78
|
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
This writes a wrapper to `~/.local/bin/headlesscode` (override with
|
|
92
|
-
`HEADLESSCODE_BIN_DIR`) that runs this checkout's `src/cli.ts` via its local
|
|
93
|
-
`tsx`, without `cd`-ing — so `--repo`/`--workspace` still default to whatever
|
|
94
|
-
directory you're standing in when you invoke it. Re-run the script any time
|
|
95
|
-
after `git pull` to point it at a moved checkout; the wrapper itself doesn't
|
|
96
|
-
need updating for ordinary code changes.
|
|
79
|
+
`--dry-run` builds the system prompt and validates project configuration. It
|
|
80
|
+
prints the prompt and a summary of loaded modes, tools, and prompt size. A
|
|
81
|
+
successful exit means prompt building and mode/rules loading succeeded.
|
|
97
82
|
|
|
98
83
|
```bash
|
|
99
|
-
|
|
100
|
-
headlesscode orchestrate --repo . --issue 42
|
|
84
|
+
headlesscode --dry-run --mode code --workspace /path/to/repo
|
|
101
85
|
```
|
|
102
86
|
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
87
|
+
Project `.roomodes`, `.roo/rules-<slug>/`, and `AGENTS.md` instructions are
|
|
88
|
+
loaded by the prompt builder when applicable. The selected mode controls which
|
|
89
|
+
tools are exposed. See [Architecture](#architecture) for the relevant code.
|
|
106
90
|
|
|
107
|
-
##
|
|
91
|
+
## OpenShell
|
|
108
92
|
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
93
|
+
Build the worker image and import the OpenRouter provider profile once. Configure
|
|
94
|
+
a Docker-backed OpenShell gateway with driver-config and bind-mount support;
|
|
95
|
+
keep that gateway restricted to trusted operators.
|
|
112
96
|
|
|
113
97
|
```bash
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
It prints the assembled system prompt plus a summary line (mode, custom modes
|
|
118
|
-
loaded, exposed tools, prompt size). Exit code 0 means prompt building + mode /
|
|
119
|
-
rules loading succeeded; non-zero means a config error.
|
|
98
|
+
# Run these commands from a HeadlessCode source checkout.
|
|
99
|
+
docker build -f docker/OpenShell.Dockerfile -t headlesscode-openshell:local .
|
|
100
|
+
openshell profile import -f shared/openshell/headlesscode-openrouter.yaml
|
|
120
101
|
|
|
121
|
-
## Real-usage example against a target repo
|
|
122
|
-
|
|
123
|
-
```bash
|
|
124
102
|
export HEADLESSCODE_OPENROUTER_API_KEY=sk-or-...
|
|
125
|
-
export
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
headlesscode \
|
|
129
|
-
--mode code \
|
|
130
|
-
--task "Implement issue #29: add retry logic to the HTTP client (see .roo/rules for project conventions)." \
|
|
131
|
-
--workspace ~/Projects/some-target-repo \
|
|
132
|
-
--max-iterations 50 \
|
|
133
|
-
--log-file ./headlesscode-session.log
|
|
103
|
+
export HEADLESSCODE_OPENSHELL_POLICY="$PWD/shared/openshell/headlesscode-policy.yaml"
|
|
104
|
+
headlesscode orchestrate --repo /path/to/repo --issue 42 \
|
|
105
|
+
--execution-provider openshell
|
|
134
106
|
```
|
|
135
107
|
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
108
|
+
Each worker gets an independent clone and a temporary result directory. Neither
|
|
109
|
+
bind source is the host worktree. After the sandbox is deleted, HeadlessCode
|
|
110
|
+
imports a verified Git bundle into the host branch. The default scratch root is
|
|
111
|
+
`~/.local/share/headlesscode/openshell-sessions`; create it before starting the
|
|
112
|
+
gateway and bind-mount it into the gateway container at the same absolute path.
|
|
113
|
+
If you choose another root, set `HEADLESSCODE_OPENSHELL_SCRATCH_ROOT` and mount
|
|
114
|
+
that host directory into the gateway container at the same path. Project data and shared instructions
|
|
115
|
+
are copied into the scratch tree and mounted separately. The writable scratch
|
|
116
|
+
checkpoint store supports checkpoints during a session and is discarded at
|
|
117
|
+
teardown. Opt-in memory uses a scratch snapshot and merges validated JSONL
|
|
118
|
+
records after sandbox deletion. No OpenRouter key is passed as a normal sandbox
|
|
119
|
+
environment argument; OpenShell injects the imported provider credential.
|
|
120
|
+
|
|
121
|
+
`headlesscode watch` accepts the same `--execution-provider openshell` option.
|
|
122
|
+
OpenShell applies to worker, plan-first, continuation, reviewer, QA, and rework
|
|
123
|
+
sessions. The one-off `headlesscode openshell-session` command runs a single
|
|
124
|
+
task and keeps its host worktree for review. Set up a gateway callback address
|
|
125
|
+
that the sandbox's Docker bridge can reach; keep the client endpoint loopback
|
|
126
|
+
only when possible. Read the [OpenShell guide](./docs/openshell-integration.md)
|
|
127
|
+
before adapting the sample gateway or filesystem policy.
|
|
128
|
+
|
|
129
|
+
| Setting | Default | Meaning |
|
|
130
|
+
| --- | --- | --- |
|
|
131
|
+
| `--execution-provider` | `local` | `orchestrate`/`watch` runtime: `local` or `openshell`. Also settable with `HEADLESSCODE_EXECUTION_PROVIDER`. |
|
|
132
|
+
| `HEADLESSCODE_OPENSHELL_POLICY` | No default | Base OpenShell policy YAML; required for OpenShell sessions. |
|
|
133
|
+
| `HEADLESSCODE_OPENSHELL_IMAGE` | `headlesscode-openshell:local` | Sandbox image. |
|
|
134
|
+
| `HEADLESSCODE_OPENSHELL_PROVIDERS` | `headlesscode-openrouter` | Comma-separated credential profile names. |
|
|
135
|
+
| `HEADLESSCODE_OPENSHELL_CPU` / `HEADLESSCODE_OPENSHELL_MEMORY` | `2` / `4Gi` | Per-sandbox resource requests. |
|
|
136
|
+
| `HEADLESSCODE_OPENSHELL_NAME_PREFIX` | `hcls` | Prefix for deterministic sandbox names. |
|
|
137
|
+
| `HEADLESSCODE_OPENSHELL_SCRATCH_ROOT` | `~/.local/share/headlesscode/openshell-sessions` | Host scratch directory that the gateway container must see at the same path. |
|
|
138
|
+
|
|
139
|
+
The `headlesscode openshell-session` command provides a single-session live
|
|
140
|
+
entry point; run `headlesscode openshell-session --help` for options.
|
|
141
|
+
|
|
142
|
+
## Parallel work and GitHub intake
|
|
143
|
+
|
|
144
|
+
`orchestrate` runs a parallel round from GitHub issue numbers or a JSON file.
|
|
145
|
+
It splits work, creates a worktree per group, starts workers, watches for
|
|
146
|
+
completion, and reviews results. QA and deployment approval are optional.
|
|
149
147
|
|
|
150
148
|
```bash
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
# Run one small generation with the local Ollama model:
|
|
155
|
-
npx tsx src/cli.ts improve --repo . --population 2 --generations 1
|
|
149
|
+
headlesscode orchestrate --repo /path/to/repo --issue 42 --issue 43 --qa
|
|
150
|
+
headlesscode orchestrate status --repo /path/to/repo --wait --timeout-ms 7200000
|
|
156
151
|
```
|
|
157
152
|
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
does not yet train adapters, schedule multiple trajectories, or provide
|
|
162
|
-
OS-level candidate isolation. Those limits are tracked in [the RSI roadmap](#rsi-roadmap).
|
|
153
|
+
`watch` polls GitHub for issues with a chosen label and starts bounded batches.
|
|
154
|
+
It requires `GH_TOKEN` or `GITHUB_TOKEN`; it performs intake only. Use a later
|
|
155
|
+
`orchestrate` run to monitor, review, QA, or deploy a batch.
|
|
163
156
|
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
- [OS-level candidate sandbox](https://github.com/Capsize-Games/headlesscode/issues/3)
|
|
170
|
-
- [Cryptographically verifiable evaluator and artifacts](https://github.com/Capsize-Games/headlesscode/issues/4)
|
|
171
|
-
- [Resource-aware resumable scheduler](https://github.com/Capsize-Games/headlesscode/issues/5)
|
|
172
|
-
- [Adaptive multi-trajectory search](https://github.com/Capsize-Games/headlesscode/issues/6)
|
|
173
|
-
- [Validated curriculum fixtures](https://github.com/Capsize-Games/headlesscode/issues/7)
|
|
174
|
-
- [Adversarial evaluation and cross-model supervision](https://github.com/Capsize-Games/headlesscode/issues/8)
|
|
175
|
-
- [Real LoRA or QLoRA backend](https://github.com/Capsize-Games/headlesscode/issues/9)
|
|
176
|
-
|
|
177
|
-
See [`docs/recursive-self-improvement.md`](./docs/recursive-self-improvement.md)
|
|
178
|
-
for the design and [`docs/rsi-progress.md`](./docs/rsi-progress.md) for the
|
|
179
|
-
record of the first bounded runs.
|
|
180
|
-
|
|
181
|
-
## Registering a new project
|
|
157
|
+
```bash
|
|
158
|
+
GH_TOKEN=... headlesscode watch \
|
|
159
|
+
--owner my-org --repo /path/to/repo --label needs-agent --run-once
|
|
160
|
+
```
|
|
182
161
|
|
|
183
|
-
|
|
184
|
-
|
|
162
|
+
| Command | Important options | Behavior and limits |
|
|
163
|
+
| --- | --- | --- |
|
|
164
|
+
| `headlesscode orchestrate --repo <path>` | Repeat `--issue <n>` or use `--issues-json <file>`; `--execution-provider local|openshell`, `--qa`, `--deploy`, `--no-review`, `--dry-run`, `--max-iterations <n>`, `--max-rework-cycles <n>`, `--max-continuations <n>` | Worktree per group; default review is enabled. OpenShell applies to worker, plan-first, continuation, reviewer, QA, and rework sessions. `--deploy` uses the human-approval gate. |
|
|
165
|
+
| `headlesscode orchestrate status --repo <path>` | `--wait`, `--timeout-ms <n>`, `--json` | Reads durable round state; `--wait` blocks until terminal state or timeout. |
|
|
166
|
+
| `headlesscode orchestrate stop --repo <path>` | Repeat `--group <name>` | Stops selected groups' worker process trees. |
|
|
167
|
+
| `headlesscode watch --owner <o> --repo <path> --label <name>` | `--execution-provider local|openshell`, `--run-once`, `--dry-run`, `--max-per-sweep <n>`, `--max-concurrent-sessions <n>`, `--retry-failed` | Label-based intake. `--dry-run` still needs a GitHub token to list issues, but writes no state and starts no workers. |
|
|
168
|
+
|
|
169
|
+
Both orchestration and watcher enforce a shared concurrency cap, defaulting to
|
|
170
|
+
3 (`HEADLESSCODE_MAX_CONCURRENT_SESSIONS`). Orchestration aborts when the cap
|
|
171
|
+
is full; the watcher leaves excess issues pending. The watcher records
|
|
172
|
+
idempotency state under `<repo>/.worktrees/.watcher-state.json` by default.
|
|
173
|
+
Use `--help` on a command for its complete options. Detailed guides:
|
|
174
|
+
[orchestration](./docs/phase2-orchestration.md),
|
|
175
|
+
[issue watcher](./docs/phase5-issue-watcher.md),
|
|
176
|
+
[QA implementation](./src/qa/qa.ts), and
|
|
177
|
+
[deployment gate](./docs/phase4-deploy-gate.md).
|
|
178
|
+
|
|
179
|
+
## Project registration and search
|
|
180
|
+
|
|
181
|
+
Registering a project detects its stack, ensures `.gitignore` excludes
|
|
182
|
+
`.headlesscode/`, and builds a code-search index and codemap. Indexing calls an
|
|
183
|
+
embedding API and can incur charges unless skipped or configured for Ollama.
|
|
185
184
|
|
|
186
185
|
```bash
|
|
187
186
|
headlesscode init --workspace ~/Projects/your-project
|
|
187
|
+
headlesscode init --workspace ~/Projects/your-project --skip-index
|
|
188
188
|
```
|
|
189
189
|
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
embedding API and costs real money unless you pass `--skip-index` or
|
|
195
|
-
`--embedding-backend ollama`. See `headlesscode init --help` for the full
|
|
196
|
-
options.
|
|
190
|
+
See `headlesscode init --help` for `--skip-codemap` and embedding-backend
|
|
191
|
+
options. The CLI also provides `index`, `codemap`, project-store, GitHub App,
|
|
192
|
+
and decision-proxy commands; run `headlesscode --help` for the command list
|
|
193
|
+
and each subcommand's `--help` for complete usage.
|
|
197
194
|
|
|
198
|
-
##
|
|
195
|
+
## Recursive self-improvement
|
|
199
196
|
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
--task-file <path> Read the task from a file (relative to workspace)
|
|
210
|
-
--workspace <root> Workspace root (default: $HEADLESSCODE_WORKSPACE_ROOT or cwd)
|
|
211
|
-
--model <id> OpenRouter model id (default: $OPENROUTER_MODEL or deepseek/deepseek-v4-flash-0731)
|
|
212
|
-
--max-iterations <n> Loop iteration cap (default: 50)
|
|
213
|
-
--consecutive-error-limit <n> Consecutive mistakes before giving up (default: 3)
|
|
214
|
-
--max-cost-usd <n> Phase 6: per-session cost cap in USD (decimal). Default
|
|
215
|
-
$HEADLESSCODE_MAX_COST_USD; off when neither is set
|
|
216
|
-
--max-duration-ms <n> Phase 6: per-session wall-clock cap in ms. Default
|
|
217
|
-
$HEADLESSCODE_MAX_DURATION_MS; off when neither is set.
|
|
218
|
-
A tripped cap aborts with reason "budget"
|
|
219
|
-
--log-file <path> Also append structured logs to this file
|
|
220
|
-
--memory-dir <path> Phase 3 memory: store facts + session summaries under <path>
|
|
221
|
-
(enabled; default $HEADLESSCODE_MEMORY_DIR or
|
|
222
|
-
<workspace>/.headlesscode/memory). Memory is OFF unless set.
|
|
223
|
-
--no-memory Explicitly disable memory even if HEADLESSCODE_MEMORY_DIR is set
|
|
224
|
-
--allowed-commands <list> Comma-separated command prefixes the agent may run.
|
|
225
|
-
Default: $HEADLESSCODE_ALLOWED_COMMANDS, else
|
|
226
|
-
.headlesscode/permissions.json, else empty (=
|
|
227
|
-
allow everything except --denied-commands; see
|
|
228
|
-
SECURITY.md)
|
|
229
|
-
--denied-commands <list> Comma-separated command prefixes that are ALWAYS
|
|
230
|
-
refused (deny wins over allow; dangerous shell
|
|
231
|
-
substitutions are always blocked regardless).
|
|
232
|
-
Default: $HEADLESSCODE_DENIED_COMMANDS, else
|
|
233
|
-
.headlesscode/permissions.json, else empty
|
|
234
|
-
--protected-files <list> Comma-separated glob patterns of files the agent may
|
|
235
|
-
not write. Default: $HEADLESSCODE_PROTECTED_FILES,
|
|
236
|
-
else .headlesscode/permissions.json, else
|
|
237
|
-
".env,.env.*,*.pem,*.key,id_rsa*"
|
|
238
|
-
--allow-protected-writes Escape hatch: permit writes to protected files
|
|
239
|
-
(default: OFF). Also settable via
|
|
240
|
-
"allowProtectedWrites": true in
|
|
241
|
-
.headlesscode/permissions.json
|
|
242
|
-
--dry-run Build the system prompt + validate config, then exit (no API key)
|
|
243
|
-
--version / --help
|
|
244
|
-
|
|
245
|
-
orchestrate subcommand (Phase 2 — parallel worktrees, headless workers):
|
|
246
|
-
--repo <path> Target repo root (required)
|
|
247
|
-
--issue <n> Issue number to include (repeatable)
|
|
248
|
-
--issues-json <file> Read issues from a JSON array of {number,title,body}
|
|
249
|
-
(used when gh is unavailable, or for tests)
|
|
250
|
-
--file-issues With --issues-json: file a REAL GitHub issue for
|
|
251
|
-
each synthetic entry (gh issue create --repo
|
|
252
|
-
<origin-owner>/<origin-repo>), swap in the real
|
|
253
|
-
number returned, and print one confirmation line
|
|
254
|
-
per created issue. A real, visible write to
|
|
255
|
-
GitHub — opt-in, never automatic. Requires
|
|
256
|
-
--issues-json
|
|
257
|
-
--batch <name> Batch id in the state file (default round-<date>)
|
|
258
|
-
--review-mode <slug> Mode slug for review sessions (default deepseek-reviewer)
|
|
259
|
-
--no-review Spawn + watch only; skip the reviewer
|
|
260
|
-
--qa Phase 4: run a headless QA session (--mode qa-agent)
|
|
261
|
-
on each group after its review passes; record
|
|
262
|
-
qa {status,verdict,evidence} in the state file
|
|
263
|
-
--qa-mode <slug> Mode slug for QA sessions (default qa-agent; the
|
|
264
|
-
target repo's .roomodes + .roo/rules-<slug>/ are
|
|
265
|
-
spliced automatically)
|
|
266
|
-
--deploy Phase 4: after all groups done + reviewed + QA passed,
|
|
267
|
-
run the human-approval deploy gate
|
|
268
|
-
(scripts/deploy-gate.sh) — a hard stop that never runs
|
|
269
|
-
the repo's deploy-production.sh without explicit human
|
|
270
|
-
approval (interactive on a TTY, token/file otherwise)
|
|
271
|
-
--deploy-args <str> Deploy args forwarded to the deploy script after the
|
|
272
|
-
gate approves (space-separated flags; also DEPLOY_ARGS env)
|
|
273
|
-
--poll-interval-ms <n> Watcher poll interval (default 5000)
|
|
274
|
-
--max-concurrent-sessions <n> Phase 6 global cap on concurrent sessions across
|
|
275
|
-
processes (default $HEADLESSCODE_MAX_CONCURRENT_
|
|
276
|
-
SESSIONS or 3). At/over the cap this run ABORTS
|
|
277
|
-
with a clear message, exit 1
|
|
278
|
-
--dry-run Print the split plan + spawn commands, spawn nothing
|
|
279
|
-
|
|
280
|
-
watch subcommand (Phase 5 — GitHub issue watcher, poll-based intake):
|
|
281
|
-
--owner <o> GitHub owner (required)
|
|
282
|
-
--repo <path> Local clone of the target repo (required; worktrees
|
|
283
|
-
are spawned under <path>/.worktrees/). The GitHub repo
|
|
284
|
-
name defaults to the directory basename (--gh-repo overrides)
|
|
285
|
-
--label <name> The label that triggers processing, e.g. needs-agent
|
|
286
|
-
--poll-interval-ms <n> Sweep interval in continuous mode (default 60000)
|
|
287
|
-
--run-once One sweep then exit 0 (or 1 if any spawn failed)
|
|
288
|
-
--max-per-sweep <n> Max NEW issues spawned per sweep (default 5); the rest
|
|
289
|
-
stay 'pending' in state and are picked up next sweep
|
|
290
|
-
--max-concurrent-sessions <n> Phase 6 GLOBAL cap on concurrent sessions across
|
|
291
|
-
processes (default $HEADLESSCODE_MAX_CONCURRENT_
|
|
292
|
-
SESSIONS or 3). Interplay: maxPerSweep bounds one
|
|
293
|
-
sweep's burst; this bounds the total fleet — issues
|
|
294
|
-
beyond it stay 'pending' until slots free up
|
|
295
|
-
--state-file <path> Durable idempotency state (default
|
|
296
|
-
<repo>/.worktrees/.watcher-state.json)
|
|
297
|
-
--mode <slug> / --memory-dir <path>
|
|
298
|
-
Forwarded to the spawner (ORCHESTRATOR_MODE /
|
|
299
|
-
HEADLESSCODE_MEMORY_DIR)
|
|
300
|
-
--qa / --deploy Pass-through: recorded per batch for the follow-up
|
|
301
|
-
orchestrate completion run
|
|
302
|
-
--dry-run Sweep + print the spawn plan, spawn nothing, write no
|
|
303
|
-
state (still needs a GitHub token for listIssues)
|
|
304
|
-
--retry-failed Retry previously-failed spawns next sweep
|
|
197
|
+
`headlesscode improve` runs a bounded experiment against an external evaluator.
|
|
198
|
+
The supervisor creates candidate worktrees, asks the local Ollama worker to
|
|
199
|
+
make focused changes, runs regression and visible/hidden evaluations, and
|
|
200
|
+
keeps selection state outside candidate worktrees. Each generation produces a
|
|
201
|
+
report for human review.
|
|
202
|
+
|
|
203
|
+
```bash
|
|
204
|
+
npx tsx src/cli.ts improve --repo . --dry-run
|
|
205
|
+
npx tsx src/cli.ts improve --repo . --population 2 --generations 1
|
|
305
206
|
```
|
|
306
207
|
|
|
208
|
+
This loop does not yet train adapters, schedule multiple trajectories, or
|
|
209
|
+
provide OS-level candidate isolation. Follow-up work is tracked as:
|
|
210
|
+
|
|
211
|
+
| Issue | Work |
|
|
212
|
+
| --- | --- |
|
|
213
|
+
| [#3](https://github.com/Capsize-Games/headlesscode/issues/3) | OS-level candidate sandbox |
|
|
214
|
+
| [#4](https://github.com/Capsize-Games/headlesscode/issues/4) | Cryptographically verifiable evaluator and artifacts |
|
|
215
|
+
| [#5](https://github.com/Capsize-Games/headlesscode/issues/5) | Resource-aware resumable scheduler |
|
|
216
|
+
| [#6](https://github.com/Capsize-Games/headlesscode/issues/6) | Adaptive multi-trajectory search |
|
|
217
|
+
| [#7](https://github.com/Capsize-Games/headlesscode/issues/7) | Validated curriculum fixtures |
|
|
218
|
+
| [#8](https://github.com/Capsize-Games/headlesscode/issues/8) | Adversarial evaluation and cross-model supervision |
|
|
219
|
+
| [#9](https://github.com/Capsize-Games/headlesscode/issues/9) | Real LoRA or QLoRA backend |
|
|
220
|
+
|
|
221
|
+
See the [RSI design](./docs/recursive-self-improvement.md) and
|
|
222
|
+
[run progress](./docs/rsi-progress.md).
|
|
223
|
+
|
|
224
|
+
## Core CLI options
|
|
225
|
+
|
|
226
|
+
| Option | Default or source | Purpose |
|
|
227
|
+
| --- | --- | --- |
|
|
228
|
+
| `--task <text>` / `--task-file <path>` | One is required unless `--dry-run` | Task prompt; task file is resolved relative to the workspace. |
|
|
229
|
+
| `--workspace <path>` | `HEADLESSCODE_WORKSPACE_ROOT` or current directory | Repository directory exposed to the worker. |
|
|
230
|
+
| `--mode <slug>` | `code` | Built-in or project `.roomodes` mode. |
|
|
231
|
+
| `--model <id>` | `OPENROUTER_MODEL` or `deepseek/deepseek-v4-flash-0731` | Model ID. |
|
|
232
|
+
| `--max-iterations <n>` | 250 for direct CLI sessions; orchestrator workers default to 50 | Per-session loop bound. |
|
|
233
|
+
| `--consecutive-error-limit <n>` | 3 (local code backend may use 6) | Consecutive tool/model errors before stopping. |
|
|
234
|
+
| `--max-cost-usd <n>` / `--max-duration-ms <n>` | `HEADLESSCODE_MAX_COST_USD` / `HEADLESSCODE_MAX_DURATION_MS`; disabled if unset | Per-session budget caps. A tripped cap aborts the session. |
|
|
235
|
+
| `--memory-dir <path>` / `--no-memory` | Off unless `HEADLESSCODE_MEMORY_DIR` or option is set | Enable project facts and rolling session summaries, or explicitly disable. |
|
|
236
|
+
| `--allowed-commands <list>` / `--denied-commands <list>` | CLI, environment, `.headlesscode/permissions.json`; allow list is empty by default | Command-prefix policy. Denials win; dangerous shell substitutions are blocked. See [`SECURITY.md`](./SECURITY.md). |
|
|
237
|
+
| `--protected-files <globs>` / `--allow-protected-writes` | `.env,.env.*,*.pem,*.key,id_rsa*` | Block writes matching protected files; explicit escape hatch. |
|
|
238
|
+
| `--dry-run` | Off | Build prompt and validate configuration without a model call or API key. |
|
|
239
|
+
| `--log-file <path>` | No file | Also append structured session logs to a file. |
|
|
240
|
+
|
|
241
|
+
Run `headlesscode --help` for advanced controls such as context condensation,
|
|
242
|
+
checkpoints, recursive delegation, session pause, and LLM timeout. Some options
|
|
243
|
+
are backend-specific; help output and source defaults are authoritative.
|
|
244
|
+
|
|
307
245
|
## Architecture
|
|
308
246
|
|
|
309
|
-
-
|
|
310
|
-
|
|
311
|
-
|
|
312
|
-
|
|
313
|
-
|
|
314
|
-
|
|
315
|
-
|
|
316
|
-
|
|
317
|
-
|
|
318
|
-
|
|
319
|
-
|
|
320
|
-
|
|
321
|
-
|
|
322
|
-
|
|
323
|
-
|
|
324
|
-
|
|
325
|
-
|
|
326
|
-
|
|
327
|
-
|
|
328
|
-
|
|
329
|
-
|
|
330
|
-
|
|
331
|
-
|
|
332
|
-
- [`src/engine/parser.ts`](./src/engine/parser.ts) — OpenAI function-calling
|
|
333
|
-
parser: JSON.parses `tool_calls[].function.arguments` with a best-effort
|
|
334
|
-
partial-JSON fallback; parse failures are marked and fed back as errors.
|
|
335
|
-
- [`src/engine/prompt.ts`](./src/engine/prompt.ts) — wraps the vendored
|
|
336
|
-
`SYSTEM_PROMPT` builder: loads project `.roomodes` (same zod schema as Zoo
|
|
337
|
-
Code's `CustomModesManager`), passes the workspace as `cwd` so `.roo/rules-*`
|
|
338
|
-
/ `AGENTS.md` splice in, and selects the mode's exposed tools.
|
|
339
|
-
- [`src/engine/loop.ts`](./src/engine/loop.ts) — `HeadlessSession`, the
|
|
340
|
-
orchestration loop. System + user → LLM → assistant (with tool_calls) → parse
|
|
341
|
-
→ execute → `tool` role message → repeat. Terminates on `attempt_completion`
|
|
342
|
-
(its `args.result` is the final answer) or a text-only reply; fails bounded
|
|
343
|
-
on max iterations or `consecutiveErrorLimit` consecutive mistakes (tool
|
|
344
|
-
errors / parse errors / identical repeated calls). History truncation is a
|
|
345
|
-
Phase 1 placeholder: system + first user always kept, sliding window of the
|
|
346
|
-
last ~40 messages. The loop accepts an injected `llmClient` (DI) so tests use
|
|
347
|
-
a fake; the CLI wires `OpenRouterClient`.
|
|
348
|
-
- [`src/engine/logger.ts`](./src/engine/logger.ts) — structured logger
|
|
349
|
-
(timestamped lines to stdout/stderr, optional file).
|
|
350
|
-
- [`src/memory/`](./src/memory/index.ts) — memory subsystem:
|
|
351
|
-
- [`src/memory/types.ts`](./src/memory/types.ts) — the `MemoryFact` /
|
|
352
|
-
`SessionSummary` schema (per-project scoped, kinds
|
|
353
|
-
`convention|decision|failure|knowledge` where "things that didn't work"
|
|
354
|
-
are `failure`) and the two contracts: `MemoryStore` (the pluggable
|
|
355
|
-
storage boundary) and `Embedder`. Hard data-isolation requirement
|
|
356
|
-
documented: the harness knowledge is a dedicated schema, never reachable
|
|
357
|
-
through any customer tenant route.
|
|
358
|
-
- [`src/memory/local.ts`](./src/memory/local.ts) — `LocalMemoryStore`: the
|
|
359
|
-
fully working file backend (`facts/<project>.jsonl` +
|
|
360
|
-
`sessions/<project>.jsonl`, append-only, idempotent `addFact` by content
|
|
361
|
-
hash, `queryRecall` = keyword matches (high weight) + local-embedder
|
|
362
|
-
cosine similarity, deterministic ordering).
|
|
363
|
-
- [`src/memory/embed.ts`](./src/memory/embed.ts) — `createLocalEmbedder()`:
|
|
364
|
-
zero-dependency, deterministic lexical-hash embedder (lowercase word +
|
|
365
|
-
char-bigram tokens → fixed-dim L2-normalized vector). Placeholder for a
|
|
366
|
-
real local embedding model behind the same `Embedder` interface.
|
|
367
|
-
- [`src/memory/summarizer.ts`](./src/memory/summarizer.ts) —
|
|
368
|
-
`extractSessionSummary` (deterministic; files/commands derived from the
|
|
369
|
-
tool history, facts via keyword heuristics) + `buildRollingSummary`
|
|
370
|
-
(compact markdown recap of the last N sessions, so a session never needs
|
|
371
|
-
the infinite raw history).
|
|
372
|
-
- [`src/memory/uwuchat.ts`](./src/memory/uwuchat.ts) — a remote
|
|
373
|
-
`MemoryStore` implementation stub for a future hosted memory API. Throws
|
|
374
|
-
"not implemented" until its base-URL/token env vars are set.
|
|
375
|
-
- Memory is wired into `HeadlessSession` as an **opt-in** config (`memory`
|
|
376
|
-
+ `project`); when unset the loop behaves exactly as before. When set, the
|
|
377
|
-
loop injects a `## PROJECT MEMORY` section (recalled facts + rolling
|
|
378
|
-
recap) into the first user message and records the session + extracted
|
|
379
|
-
facts afterwards — and memory failures are always non-fatal.
|
|
380
|
-
- [`src/orchestrator/`](./src/orchestrator/index.ts) — Phase 2 orchestration
|
|
381
|
-
layer: `split.ts` (issue-splitting heuristics, deterministic +
|
|
382
|
-
unit-tested), `state.ts` (`.worktrees/.orchestrator-state.json` read/write),
|
|
383
|
-
`reviewer.ts` (adversarial fresh-context review run with a read-only
|
|
384
|
-
executor), `watch.ts` (completion polling of `.harness.done` markers + stall
|
|
385
|
-
guard), and `cli.ts` (the `orchestrate` subcommand — split → spawn via
|
|
386
|
-
`scripts/spawn-parallel-worktrees.sh` → watch → review → QA → deploy gate).
|
|
387
|
-
- [`src/qa/qa.ts`](./src/qa/qa.ts) — Phase 4 headless QA: `runQa()` runs a
|
|
388
|
-
second harness session against a worktree in the target repo's `qa-agent`
|
|
389
|
-
mode (auto-spliced from `.roomodes` + `.roo/rules-qa-agent/`), with a
|
|
390
|
-
generic checklist fallback when the repo has no such mode. Read + command
|
|
391
|
-
tools only (no `write_to_file`) — QA verifies and reports, it never edits.
|
|
392
|
-
Verdict parsing is fail-closed (`pass`/`fail`/`error`, default `fail`).
|
|
393
|
-
- [`src/deploy/gate.ts`](./src/deploy/gate.ts) + [`src/deploy/gate-cli.ts`](./src/deploy/gate-cli.ts) —
|
|
394
|
-
Phase 4 human-approval deploy gate: the pure, unit-tested decision function
|
|
395
|
-
`decideApproval` (interactive y/N, one-time approval file, or
|
|
396
|
-
`DEPLOY_APPROVAL_TOKEN` matching `<repo>/.deploy-approval`; never
|
|
397
|
-
auto-approves) plus a thin CLI the bash wrapper calls.
|
|
398
|
-
- [`src/watcher/`](./src/watcher/index.ts) — Phase 5 GitHub issue watcher:
|
|
399
|
-
[`github.ts`](./src/watcher/github.ts) (native-fetch GitHub REST client with
|
|
400
|
-
label filter, PR filtering, pagination, `GITHUB_API_BASE_URL` override for
|
|
401
|
-
tests/mocks), [`state.ts`](./src/watcher/state.ts) (durable idempotency
|
|
402
|
-
state file — write-ahead `spawned` → `done`/`failed`, `pending` for capped
|
|
403
|
-
issues, restart-safe), [`watch.ts`](./src/watcher/watch.ts) (the poll loop:
|
|
404
|
-
list by label → split → spawn via the existing bash spawner, bounded by
|
|
405
|
-
`maxPerSweep` and the Phase 6 global cap), and [`cli.ts`](./src/watcher/cli.ts)
|
|
406
|
-
(the `watch` subcommand — continuous or `--run-once`, `--dry-run`,
|
|
407
|
-
`--retry-failed`).
|
|
408
|
-
- [`src/budget/`](./src/budget/index.ts) — Phase 6 guardrails:
|
|
409
|
-
[`cost.ts`](./src/budget/cost.ts) (model pricing table + `estimateCost`,
|
|
410
|
-
`HEADLESSCODE_PRICING_JSON` override, conservative fallback for unlisted
|
|
411
|
-
models), [`budget.ts`](./src/budget/budget.ts) (`SessionBudget` +
|
|
412
|
-
`BudgetTracker`: `tick()` before each LLM call, `record()` after with usage
|
|
413
|
-
tokens, `check()` snapshot, `BudgetExceededError`), and
|
|
414
|
-
[`concurrency.ts`](./src/budget/concurrency.ts) (`ConcurrencyLimiter` —
|
|
415
|
-
fail-fast in-process semaphore — plus `activeSessionCount` reading the
|
|
416
|
-
durable orchestrator/watcher state files for a cross-process view). Wired
|
|
417
|
-
into `HeadlessSession` (`budget` config, `budgetUsage` on results), the base
|
|
418
|
-
CLI (`--max-cost-usd` / `--max-duration-ms`), `orchestrate` (aborts at the
|
|
419
|
-
cap), the watcher (defers cap-exceeding issues to `pending`), and
|
|
420
|
-
`run-worker.sh`/`run-qa.sh` (env forwarding).
|
|
421
|
-
- [`src/cloud/`](./src/cloud/provider.ts) — Phase 6 ephemeral compute
|
|
422
|
-
abstraction: the `CloudProvider` lifecycle interface
|
|
423
|
-
(`spawnWorktreeSession` → `waitReady` → `runHarness` → `collectResults` →
|
|
424
|
-
`teardown`) with `LocalProcessProvider` as the current local behavior behind
|
|
425
|
-
it (reuses `spawn-parallel-worktrees.sh` + `run-worker.sh`), so a
|
|
426
|
-
container/VM-per-issue backend slots in without touching the orchestration
|
|
427
|
-
layer. A cloud-provider sketch is documented (evaluation only — no
|
|
428
|
-
live setup; see [`docs/phase6-cloud.md`](./docs/phase6-cloud.md)).
|
|
429
|
-
- [`src/cli.ts`](./src/cli.ts) — the `headlesscode` bin entry (+ `orchestrate`
|
|
430
|
-
and `watch` subcommand dispatch).
|
|
431
|
-
- [`scripts/run-worker.sh`](./scripts/run-worker.sh) — launches one headless
|
|
432
|
-
harness worker per worktree (pid, log, exit code, `.harness.done` marker).
|
|
433
|
-
- [`scripts/run-qa.sh`](./scripts/run-qa.sh) — Phase 4 QA wrapper mirroring
|
|
434
|
-
run-worker.sh: launches one harness QA session per worktree (`.qa-task.md`,
|
|
435
|
-
`.qa.pid`, `qa.log`, `.qa.exit`, `.qa.done/`).
|
|
436
|
-
- [`scripts/deploy-gate.sh`](./scripts/deploy-gate.sh) — Phase 4 gate wrapper:
|
|
437
|
-
path safety, deployment summary, interactive + token/file approval, then
|
|
438
|
-
(and only then) invokes the repo's `scripts/deploy-production.sh` with
|
|
439
|
-
forwarded deploy args. Exit 3 = human DENIED (hard stop).
|
|
440
|
-
- [`scripts/spawn-parallel-worktrees.sh`](./scripts/spawn-parallel-worktrees.sh)
|
|
441
|
-
— spawns one git worktree + harness worker per group (worktree/.env/branch
|
|
442
|
-
conventions, `run-worker.sh` + state-file writes, no GUI involved).
|
|
443
|
-
|
|
444
|
-
### Phase 1 tool filtering decision
|
|
445
|
-
|
|
446
|
-
The loop exposes to the model exactly the tools the executor can actually run
|
|
447
|
-
for the selected mode: the intersection of the vendored mode tool groups
|
|
448
|
-
(`getToolsForMode`) with the Phase 1 executable set
|
|
449
|
-
(`read_file`, `write_to_file`, `execute_command`, `list_files`,
|
|
450
|
-
`attempt_completion`, `ask_followup_question`). Stub-only tools (`apply_diff`,
|
|
451
|
-
`search_files`, …) stay registered in the executor purely as a safety net
|
|
452
|
-
(clear "not implemented" error) but are NOT advertised to the model, so it
|
|
453
|
-
doesn't waste turns calling them.
|
|
247
|
+
The core loop builds the project-specific prompt, calls the configured model,
|
|
248
|
+
executes the selected mode's available tools, and repeats until completion or a
|
|
249
|
+
configured limit. File tools are confined to the workspace path; shell commands
|
|
250
|
+
run with the process user's privileges unless an external boundary such as
|
|
251
|
+
OpenShell is configured. Mode instructions and safety controls are distinct:
|
|
252
|
+
prompt instructions guide the model, while filesystem or command restrictions
|
|
253
|
+
must be enforced by the runtime or sandbox.
|
|
254
|
+
|
|
255
|
+
| Area | Entry points | Responsibility |
|
|
256
|
+
| --- | --- | --- |
|
|
257
|
+
| CLI and session loop | [`src/cli.ts`](./src/cli.ts), [`src/engine/`](./src/engine/) | Command dispatch, prompt construction, LLM/tool loop, iteration and context limits. |
|
|
258
|
+
| Model and tools | [`src/llm/`](./src/llm/), [`src/tools/`](./src/tools/) | OpenRouter client, tool execution, output handling, and workspace path checks. |
|
|
259
|
+
| Parallel orchestration | [`src/orchestrator/`](./src/orchestrator/), [`scripts/spawn-parallel-worktrees.sh`](./scripts/spawn-parallel-worktrees.sh) | Issue grouping, worktree workers, completion state, review, rework, QA, and deploy gating. |
|
|
260
|
+
| GitHub intake | [`src/watcher/`](./src/watcher/) | Label polling, durable idempotency, pending work, and bounded spawning. |
|
|
261
|
+
| Runtime providers | [`src/cloud/`](./src/cloud/) | Local process, Docker, and OpenShell implementations of the session lifecycle. `orchestrate` and `watch` can select OpenShell for worker sessions. |
|
|
262
|
+
| Memory and budgets | [`src/memory/`](./src/memory/), [`src/budget/`](./src/budget/) | Opt-in project memory, cost/time budgets, and session concurrency accounting. |
|
|
263
|
+
| QA and deploy gate | [`src/qa/`](./src/qa/), [`src/deploy/`](./src/deploy/) | Read-only QA verdicts and explicit human approval before deployment. |
|
|
264
|
+
| Project map and search | [`src/codemap/`](./src/codemap/), [`src/codesearch/`](./src/codesearch/) | Deterministic module maps and repository search indexes. |
|
|
265
|
+
|
|
266
|
+
The model only receives tools supported by the selected mode and executable
|
|
267
|
+
runtime. Unsupported vendored tool schemas are not advertised. This keeps the
|
|
268
|
+
model's action set aligned with the current executor; it does not make arbitrary
|
|
269
|
+
shell commands safe by itself.
|
|
454
270
|
|
|
455
271
|
## Development
|
|
456
272
|
|
|
457
|
-
|
|
458
|
-
|
|
459
|
-
npm
|
|
460
|
-
npm
|
|
461
|
-
|
|
462
|
-
|
|
463
|
-
|
|
464
|
-
bash scripts/e2e
|
|
465
|
-
|
|
466
|
-
bash scripts/e2e-
|
|
467
|
-
|
|
468
|
-
|
|
469
|
-
|
|
470
|
-
|
|
471
|
-
|
|
472
|
-
fake
|
|
473
|
-
|
|
474
|
-
|
|
475
|
-
|
|
476
|
-
|
|
477
|
-
|
|
478
|
-
|
|
479
|
-
|
|
480
|
-
|
|
481
|
-
|
|
482
|
-
|
|
483
|
-
|
|
484
|
-
|
|
485
|
-
-
|
|
486
|
-
|
|
487
|
-
|
|
488
|
-
|
|
489
|
-
-
|
|
490
|
-
|
|
491
|
-
|
|
492
|
-
port, `.orchestrator-state.json` state management, headless reviewer, and
|
|
493
|
-
the `orchestrate` CLI subcommand (see
|
|
494
|
-
[`docs/phase2-orchestration.md`](./docs/phase2-orchestration.md)).
|
|
495
|
-
- ✅ Phase 3 — memory subsystem: per-project knowledge facts + rolling session
|
|
496
|
-
summaries (`src/memory/`), local deterministic embedder for semantic recall,
|
|
497
|
-
opt-in `HeadlessSession`/CLI wiring (`--memory-dir`, `--no-memory`), and a
|
|
498
|
-
pluggable `MemoryStore` contract with a remote-backend client stub.
|
|
499
|
-
- ✅ Phase 4 — QA + deploy gate: headless QA runs the target repo's `qa-agent`
|
|
500
|
-
mode (`--qa` / `--qa-mode`), verdict parsing is fail-closed, results land in
|
|
501
|
-
the state file's per-group `qa` field; the human-approval deploy gate
|
|
502
|
-
(`--deploy`, `scripts/deploy-gate.sh` + `src/deploy/gate.ts`) is a hard stop
|
|
503
|
-
in front of `deploy-production.sh` that never auto-approves (see
|
|
504
|
-
[`docs/phase4-qa.md`](./docs/phase4-qa.md) and
|
|
505
|
-
[`docs/phase4-deploy-gate.md`](./docs/phase4-deploy-gate.md)).
|
|
506
|
-
- ✅ Phase 5 — GitHub issue watcher: poll-based intake (`watch` subcommand) —
|
|
507
|
-
detect issues by label via the GitHub REST API (`GH_TOKEN`), fan each out
|
|
508
|
-
through `splitIssues` + the existing `spawn-parallel-worktrees.sh`, track
|
|
509
|
-
idempotency durably (`.worktrees/.watcher-state.json`, write-ahead
|
|
510
|
-
ordering, `pending` cap deferral, restart-safe), optional `--dry-run` /
|
|
511
|
-
`--run-once` / `--retry-failed`; webhook upgrade designed but not built as
|
|
512
|
-
a server (see [`docs/phase5-issue-watcher.md`](./docs/phase5-issue-watcher.md)).
|
|
513
|
-
- ✅ Phase 6 — cloud scaling, guardrails-first: per-session cost/time/iteration
|
|
514
|
-
budget (`src/budget/`, `--max-cost-usd` / `--max-duration-ms`, budgetUsage on
|
|
515
|
-
results, worker env forwarding) + a hard concurrent-session cap
|
|
516
|
-
(`HEADLESSCODE_MAX_CONCURRENT_SESSIONS`, orchestrate aborts / watcher defers
|
|
517
|
-
to `pending`) + the `CloudProvider` abstraction with `LocalProcessProvider`
|
|
518
|
-
and an evaluation-only cloud-provider sketch (see
|
|
519
|
-
[`docs/phase6-cloud.md`](./docs/phase6-cloud.md)). No live cloud launched;
|
|
520
|
-
a container/VM-per-issue backend slots in behind the same interface.
|
|
521
|
-
- ⏳ Phase 3 (remaining) — token-based condensation.
|
|
273
|
+
| Command | Purpose |
|
|
274
|
+
| --- | --- |
|
|
275
|
+
| `npm install` | Install dependencies. |
|
|
276
|
+
| `npm run typecheck` | TypeScript type check. |
|
|
277
|
+
| `npm run smoke` | Vendored prompt-builder smoke test; no network. |
|
|
278
|
+
| `npm test` | Unit tests using local/fake dependencies where provided. |
|
|
279
|
+
| `npm run cli -- --dry-run --workspace .` | Build this repository's system prompt. |
|
|
280
|
+
| `bash scripts/e2e/run.sh` | Phase 1 CLI integration against mock OpenRouter. |
|
|
281
|
+
| `bash scripts/e2e-phase2/run.sh` | Orchestrator spawn/watch/review integration. |
|
|
282
|
+
| `bash scripts/e2e-phase4/run.sh` | QA and deploy-gate integration with fake deploy. |
|
|
283
|
+
| `bash scripts/e2e-phase5/run.sh` | Watcher integration against fake GitHub and stubbed spawner. |
|
|
284
|
+
| `bash scripts/e2e-phase6/run.sh` | Budget-abort and concurrency-cap integration. |
|
|
285
|
+
|
|
286
|
+
Unit tests can inject a fake LLM client into `HeadlessSession`; the listed
|
|
287
|
+
end-to-end scripts use local mock services. OpenShell provider tests use a
|
|
288
|
+
fake CLI. A live OpenShell 0.1.2 run called DeepSeek V4 Flash through OpenRouter,
|
|
289
|
+
imported its Git result, and denied a host-worktree canary write. The earlier
|
|
290
|
+
shared-worktree mount was removed after a live probe showed it exposed host Git
|
|
291
|
+
and control files. See the [OpenShell guide](./docs/openshell-integration.md)
|
|
292
|
+
for the current clone-and-import boundary and its validation.
|
|
293
|
+
|
|
294
|
+
## Implementation status
|
|
295
|
+
|
|
296
|
+
| Area | Status | Details |
|
|
297
|
+
| --- | --- | --- |
|
|
298
|
+
| Runtime engine | Implemented | Portable Zoo Code core, model client, tool executor, parser, CLI, and tests. |
|
|
299
|
+
| Orchestration | Implemented | Parallel worktrees, issue splitting, status tracking, review, and recovery. See [guide](./docs/phase2-orchestration.md). |
|
|
300
|
+
| Memory | Implemented | Opt-in per-project facts and session summaries with a local store; remote store remains a stub. |
|
|
301
|
+
| QA and deploy approval | Implemented | Read-only QA plus a fail-closed human-approval deploy gate. See [orchestration guide](./docs/phase2-orchestration.md) and [deploy gate](./docs/phase4-deploy-gate.md). |
|
|
302
|
+
| GitHub watcher | Implemented | Poll-based label intake; webhook server is not implemented. See [guide](./docs/phase5-issue-watcher.md). |
|
|
303
|
+
| Budget and concurrency controls | Implemented | Per-session limits and a cross-process concurrency cap. |
|
|
304
|
+
| Local process provider | Implemented | Existing host-process worker path behind the provider lifecycle interface. |
|
|
305
|
+
| Docker provider | Implemented | Container-backed provider implementation; `orchestrate` does not currently select it. |
|
|
306
|
+
| OpenShell provider | Implemented | Disposable clone, separate scratch data, validated Git bundle import, and sandbox teardown. See [guide](./docs/openshell-integration.md). |
|
|
307
|
+
| Context management | Implemented | Token-aware condensation summarizes older turns; message-count truncation remains the non-fatal fallback. See [`src/engine/condense.ts`](./src/engine/condense.ts). |
|
|
522
308
|
|
|
523
309
|
## Attribution
|
|
524
310
|
|