headlesscode 1.2.2 → 1.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  [![CI](https://github.com/Capsize-Games/headlesscode/actions/workflows/ci.yml/badge.svg)](https://github.com/Capsize-Games/headlesscode/actions/workflows/ci.yml)
4
4
  [![npm](https://img.shields.io/npm/v/headlesscode?logo=npm)](https://www.npmjs.com/package/headlesscode)
5
- [![Node.js >=18](https://img.shields.io/badge/node-%3E%3D18-339933?logo=node.js&logoColor=white)](https://nodejs.org/)
5
+ [![Node.js >=22.19](https://img.shields.io/badge/node-%3E%3D22.19-339933?logo=node.js&logoColor=white)](https://nodejs.org/)
6
6
  [![License: Apache-2.0](https://img.shields.io/badge/License-Apache--2.0-blue.svg)](./LICENSE)
7
7
 
8
8
  `headlesscode` runs the Zoo Code agent loop from a Node.js process, without a
@@ -38,7 +38,7 @@ monitoring.
38
38
 
39
39
  ## Quick start
40
40
 
41
- Requires Node.js 18 or newer. Install globally or run with `npx`:
41
+ Requires Node.js 22.19 or newer. Install globally or run with `npx`:
42
42
 
43
43
  ```bash
44
44
  npm install -g headlesscode
@@ -194,29 +194,59 @@ and each subcommand's `--help` for complete usage.
194
194
 
195
195
  ## Recursive self-improvement
196
196
 
197
- `headlesscode improve` runs a bounded experiment against an external evaluator.
198
- The supervisor creates candidate worktrees, asks the local Ollama worker to
199
- make focused changes, runs regression and visible/hidden evaluations, and
200
- keeps selection state outside candidate worktrees. Each generation produces a
201
- report for human review.
197
+ `headlesscode improve --dry-run` resolves the base commit and prints planned
198
+ candidate worktrees. A real run requires PostgreSQL, a shared S3-compatible
199
+ artifact store, and at least one registered `headlesscode rsi-worker`. The
200
+ coordinator queues sanitized mutation snapshots and visible evaluation jobs;
201
+ workers run each in a fresh OpenShell guest. Fitness, paired baseline
202
+ comparison, archives, and selection stay in the coordinator. Hidden evaluation
203
+ remains disabled. A 2026-09-29 bounded run used two workers on the same local
204
+ OpenShell gateway; both candidate evaluation jobs were leased concurrently.
205
+ The run used a temporary `clamp.js` fixture and does not establish general
206
+ headlesscode improvement or multi-gateway operation. An earlier bounded run
207
+ passed visible checks but made no source change and was rejected.
202
208
 
203
209
  ```bash
204
210
  npx tsx src/cli.ts improve --repo . --dry-run
205
211
  npx tsx src/cli.ts improve --repo . --population 2 --generations 1
206
212
  ```
207
213
 
208
- This loop does not yet train adapters, schedule multiple trajectories, or
209
- provide OS-level candidate isolation. Follow-up work is tracked as:
214
+ Before starting workers, build the standard and RSI guest images on each
215
+ OpenShell gateway host that will run RSI jobs. Build the larger training image
216
+ only on gateways whose workers will accept model-training or paired
217
+ model-evaluation jobs:
218
+
219
+ ```bash
220
+ docker build -f docker/OpenShell.Dockerfile -t headlesscode-openshell:local .
221
+ docker build -f docker/OpenShell-RSI.Dockerfile -t headlesscode-openshell-rsi:local .
222
+ docker build -f docker/OpenShell-RSI-Training.Dockerfile -t headlesscode-openshell-rsi-training:local .
223
+ ```
224
+
225
+ The RSI Dockerfile extends `headlesscode-openshell:local` and removes the RSI
226
+ source, test files, and hidden evaluation suite from the guest image. See the [RSI design](./docs/recursive-self-improvement.md)
227
+ for worker and queue configuration.
228
+
229
+ The queue stores job state, worker registrations, leases, retries, and admission
230
+ policy in PostgreSQL; artifact bytes remain in object storage. Local integration
231
+ tests used 12 simulated workers and 100 queued jobs, including a 20 MiB artifact. Each
232
+ worker process currently runs one guest at a time and must be configured with a
233
+ unique worker ID and OpenShell gateway ID. The two-worker run observed 4.008
234
+ seconds of concurrent evaluation leases on one gateway. A local PostgreSQL/
235
+ RustFS test drained the remaining 96 jobs with 12 simulated workers in 446 ms
236
+ (215.2 jobs/s), after the initial four jobs were claimed and completed to
237
+ verify admission limits;
238
+ claims still serialize on a queue-policy row, and neither result qualifies a
239
+ multi-host deployment. Remaining work is:
210
240
 
211
241
  | Issue | Work |
212
242
  | --- | --- |
213
- | [#3](https://github.com/Capsize-Games/headlesscode/issues/3) | OS-level candidate sandbox |
214
- | [#4](https://github.com/Capsize-Games/headlesscode/issues/4) | Cryptographically verifiable evaluator and artifacts |
215
- | [#5](https://github.com/Capsize-Games/headlesscode/issues/5) | Resource-aware resumable scheduler |
216
- | [#6](https://github.com/Capsize-Games/headlesscode/issues/6) | Adaptive multi-trajectory search |
217
- | [#7](https://github.com/Capsize-Games/headlesscode/issues/7) | Validated curriculum fixtures |
218
- | [#8](https://github.com/Capsize-Games/headlesscode/issues/8) | Adversarial evaluation and cross-model supervision |
219
- | [#9](https://github.com/Capsize-Games/headlesscode/issues/9) | Real LoRA or QLoRA backend |
243
+ | [#3](https://github.com/Capsize-Games/headlesscode/issues/3) | OpenShell candidate isolation is implemented; live multi-gateway qualification remains. |
244
+ | [#4](https://github.com/Capsize-Games/headlesscode/issues/4) | Content-addressed artifacts are implemented; signed evaluator provenance and hidden evaluation remain. |
245
+ | [#5](https://github.com/Capsize-Games/headlesscode/issues/5) | PostgreSQL queue, leases, retries, admission, and configured workers are implemented; live fleet qualification remains. |
246
+ | [#6](https://github.com/Capsize-Games/headlesscode/issues/6) | Bounded adaptive independent search supports an initial population plus one evidence-driven follow-up; repeated stages remain out of scope. |
247
+ | [#7](https://github.com/Capsize-Games/headlesscode/issues/7) | Fixture-backed curriculum replay and explicit promotion are implemented; wording-to-capability measurement remains unproven. |
248
+ | [#8](https://github.com/Capsize-Games/headlesscode/issues/8) | Adversarial review and bounded OpenShell break tests are implemented; live provider-backed review is unvalidated. |
249
+ | [#9](https://github.com/Capsize-Games/headlesscode/issues/9) | An OpenShell QLoRA prototype is covered by mocked job tests; the supplied GGUF is rejected before enqueue, and no live training is validated. |
220
250
 
221
251
  See the [RSI design](./docs/recursive-self-improvement.md) and
222
252
  [run progress](./docs/rsi-progress.md).
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "headlesscode",
3
- "version": "1.2.2",
3
+ "version": "1.3.0",
4
4
  "description": "Standalone headless coding-agent harness: runs the Zoo Code agent loop (prompts, tools, modes) without a VS Code UI, driven by CLI, HTTP, and parallel worktree orchestration.",
5
5
  "license": "Apache-2.0",
6
6
  "repository": {
@@ -17,7 +17,7 @@
17
17
  ],
18
18
  "type": "module",
19
19
  "engines": {
20
- "node": ">=18"
20
+ "node": ">=22.19"
21
21
  },
22
22
  "bin": {
23
23
  "headlesscode": "bin/headlesscode.mjs"
@@ -42,14 +42,17 @@
42
42
  "start": "tsx src/cli.ts",
43
43
  "cli": "tsx src/cli.ts",
44
44
  "monitor:pilot": "tsx scripts/monitor-pilot/run.ts",
45
+ "rsi:promote-curriculum": "tsx src/rsi/promote-curriculum.ts",
45
46
  "test": "node scripts/run-tests.mjs",
46
47
  "prepublishOnly": "npm ci && npm test"
47
48
  },
48
49
  "dependencies": {
50
+ "@aws-sdk/client-s3": "^3.1142.0",
49
51
  "@modelcontextprotocol/sdk": "^1.30.0",
50
52
  "@octokit/auth-app": "^8.2.0",
51
53
  "fastest-levenshtein": "^1.0.16",
52
54
  "p-wait-for": "^5.0.2",
55
+ "pg": "^8.23.0",
53
56
  "playwright": "^1.62.1",
54
57
  "simple-git": "^3.36.0",
55
58
  "tsx": "^4.19.0",
@@ -59,6 +62,7 @@
59
62
  "zod": "^3.25.76"
60
63
  },
61
64
  "devDependencies": {
62
- "@types/node": "^22.10.0"
65
+ "@types/node": "^22.10.0",
66
+ "@types/pg": "^8.23.1"
63
67
  }
64
68
  }
package/src/cli.ts CHANGED
@@ -43,7 +43,6 @@ import { analyzeCliMain } from "./orchestrator/analyze-cli.js"
43
43
  import { costHistoryCliMain } from "./orchestrator/cost-history-cli.js"
44
44
  import { resolvePermissions, type PermissionsConfig } from "./permissions/config.js"
45
45
  import { resolveModelForMode, resolveReasoningEffortForMode } from "./config/mode-models.js"
46
- import { improveMain } from "./rsi/controller.js"
47
46
  import { openShellSessionMain } from "./cloud/openshell-session.js"
48
47
 
49
48
  const VERSION = "0.1.0"
@@ -751,10 +750,47 @@ export async function main(argv: string[] = process.argv.slice(2)): Promise<numb
751
750
  if (argv[0] === "openshell-session") {
752
751
  return openShellSessionMain(argv.slice(1))
753
752
  }
753
+ if (argv[0] === "rsi-worker") {
754
+ const { rsiWorkerMain } = await import("./rsi/worker.js")
755
+ return rsiWorkerMain()
756
+ }
757
+ if (argv[0] === "rsi-train-model") {
758
+ const { parseRsiArgs } = await import("./rsi/config.js")
759
+ const { readArchive } = await import("./rsi/archive.js")
760
+ const { PostgresRsiJobQueue, fleetQueuePolicy } = await import("./rsi/postgres-queue.js")
761
+ const { runBoundedModelCandidate } = await import("./rsi/model-training.js")
762
+ const parsed = parseRsiArgs(argv.slice(1))
763
+ if (parsed.help) {
764
+ process.stdout.write("Usage: headlesscode rsi-train-model --repo <path> --archive-dir <path> --resume <run-id> --model-candidate-id <id>\nThe registered worker must use an OpenShell training image and a local Hugging Face checkpoint.\n")
765
+ return 0
766
+ }
767
+ if (parsed.error || !parsed.config || !parsed.config.resumeRunId || !parsed.config.modelCandidateId || parsed.config.dryRun) {
768
+ process.stderr.write(`headlesscode rsi-train-model: ${parsed.error ?? "--resume and --model-candidate-id are required; dry-run is unsupported"}\n`)
769
+ return 2
770
+ }
771
+ const config = parsed.config
772
+ const archive = await readArchive(config.archiveDir)
773
+ const run = archive.activeRuns.find((entry) => entry.runId === config.resumeRunId) ?? archive.runs.find((entry) => entry.runId === config.resumeRunId)
774
+ if (!run) {
775
+ process.stderr.write(`headlesscode rsi-train-model: RSI run not found in ${config.archiveDir}\n`)
776
+ return 2
777
+ }
778
+ const policy = fleetQueuePolicy(process.env.HEADLESSCODE_RSI_QUEUE_NAME?.trim() || "rsi-default", config.maxConcurrent, process.env)
779
+ const queue = PostgresRsiJobQueue.fromEnvironment(policy, process.env)
780
+ try {
781
+ await queue.migrate()
782
+ const model = await runBoundedModelCandidate(config, run, queue, config.modelCandidateId!)
783
+ process.stdout.write(`[rsi-model] ${model.id}: ${model.status}; eligible=${model.eligibleForSelection === true}\n`)
784
+ return model.eligibleForSelection ? 0 : 1
785
+ } finally {
786
+ await queue.close()
787
+ }
788
+ }
754
789
  // Bounded recursive self-improvement: the supervisor owns the evaluator,
755
790
  // archive, and selection logic while each candidate runs in its own
756
791
  // worktree. See docs/recursive-self-improvement.md.
757
792
  if (argv[0] === "improve") {
793
+ const { improveMain } = await import("./rsi/controller.js")
758
794
  return improveMain(argv.slice(1))
759
795
  }
760
796
 
@@ -1049,9 +1085,17 @@ export async function main(argv: string[] = process.argv.slice(2)): Promise<numb
1049
1085
  }
1050
1086
  }
1051
1087
 
1052
- // ── real run: API key required ────────────────────────────────────────────
1088
+ // ── real run: OpenRouter key required unless this mode is local ────────────
1089
+ const codeModeBackend = process.env.HEADLESSCODE_CODE_MODE_BACKEND ?? "openrouter"
1090
+ const localBackendModes = new Set(
1091
+ (process.env.HEADLESSCODE_LOCAL_BACKEND_MODES ?? "code")
1092
+ .split(",")
1093
+ .map((s) => s.trim())
1094
+ .filter(Boolean),
1095
+ )
1096
+ const useLocalCodeBackend = localBackendModes.has(options.mode) && codeModeBackend === "ollama"
1053
1097
  const apiKey = process.env.HEADLESSCODE_OPENROUTER_API_KEY
1054
- if (!apiKey) {
1098
+ if (!apiKey && !useLocalCodeBackend) {
1055
1099
  process.stderr.write(
1056
1100
  "headlesscode: HEADLESSCODE_OPENROUTER_API_KEY is not set.\n" +
1057
1101
  " Export it (e.g. export HEADLESSCODE_OPENROUTER_API_KEY=sk-or-...) or use --dry-run to\n" +
@@ -1098,21 +1142,6 @@ export async function main(argv: string[] = process.argv.slice(2)): Promise<numb
1098
1142
  env: process.env,
1099
1143
  })
1100
1144
 
1101
- // Generalized 2026-08-21 (issue #142) beyond the original `code`-only
1102
- // scope (plans/local-dual-model-code-agent.md, D2) — real need: two
1103
- // separate local daemons on two different GPUs (a coder model and a
1104
- // review model), each needing its own mode -> URL/model mapping.
1105
- // HEADLESSCODE_LOCAL_BACKEND_MODES defaults to "code" alone, so
1106
- // nobody's existing setup changes behavior unless they opt in.
1107
- const codeModeBackend = process.env.HEADLESSCODE_CODE_MODE_BACKEND ?? "openrouter"
1108
- const localBackendModes = new Set(
1109
- (process.env.HEADLESSCODE_LOCAL_BACKEND_MODES ?? "code")
1110
- .split(",")
1111
- .map((s) => s.trim())
1112
- .filter(Boolean),
1113
- )
1114
- const useLocalCodeBackend = localBackendModes.has(options.mode) && codeModeBackend === "ollama"
1115
-
1116
1145
  // The local daemon's own proxy (ollama_shim.py) deliberately runs a 600s
1117
1146
  // upstream request timeout — its own comment documents why: a shorter
1118
1147
  // shim timeout was once found to cut off calls before the harness's own
@@ -1,10 +1,10 @@
1
1
  import { spawnSync } from "node:child_process"
2
2
  import * as fs from "node:fs"
3
3
 
4
- export function openshellPreflight(): string | undefined {
4
+ export function openshellPreflight(options: { requireOpenRouterCredential?: boolean } = {}): string | undefined {
5
5
  const policy = process.env.HEADLESSCODE_OPENSHELL_POLICY
6
6
  if (!policy || !fs.existsSync(policy)) return "set HEADLESSCODE_OPENSHELL_POLICY to an existing base policy YAML"
7
- if (!process.env.HEADLESSCODE_OPENROUTER_API_KEY) return "set HEADLESSCODE_OPENROUTER_API_KEY so OpenShell can provide the imported credential profile"
7
+ if (options.requireOpenRouterCredential !== false && !process.env.HEADLESSCODE_OPENROUTER_API_KEY) return "set HEADLESSCODE_OPENROUTER_API_KEY so OpenShell can provide the imported credential profile"
8
8
  const version = spawnSync("openshell", ["--version"], { encoding: "utf8", timeout: 5000 })
9
9
  if (version.error || version.status !== 0) return "OpenShell CLI is unavailable"
10
10
  const gateway = spawnSync("openshell", ["gateway", "info"], { encoding: "utf8", timeout: 10000 })
@@ -59,12 +59,24 @@ export interface OpenShellSessionProviderOptions {
59
59
  image?: string
60
60
  policy?: string
61
61
  providers?: string[]
62
+ /** Attach automatically discovered host provider profiles to the sandbox. */
63
+ autoProviders?: boolean
62
64
  cpu?: string
63
65
  memory?: string
64
66
  namePrefix?: string
65
67
  readOnlyMounts?: Array<{ source: string; target: string }>
66
68
  importResults?: boolean
67
69
  scratchRoot?: string
70
+ /** Omit project/shared data snapshots for experiments that must not inherit operator context. */
71
+ includeProjectData?: boolean
72
+ /** Use a guest-safe identity path instead of the supervisor checkout path. */
73
+ projectIdentityRoot?: string
74
+ /** Additional exact OpenShell network rules needed by the workload. */
75
+ networkPolicy?: Record<string, unknown>
76
+ /** Timeout for sandbox commands, including agent mutation. */
77
+ commandTimeoutMs?: number
78
+ /** Number of GPU devices requested for a training-only sandbox. */
79
+ gpu?: number
68
80
  run?: (args: string[], timeoutMs?: number) => CommandResult
69
81
  listSandboxes?: () => string[]
70
82
  }
@@ -74,6 +86,7 @@ export class OpenShellSessionProvider implements CloudProvider {
74
86
  private readonly image: string
75
87
  private readonly policy?: string
76
88
  private readonly providers: string[]
89
+ private readonly autoProviders: boolean
77
90
  private readonly cpu: string
78
91
  private readonly memory: string
79
92
  private readonly namePrefix: string
@@ -82,18 +95,30 @@ export class OpenShellSessionProvider implements CloudProvider {
82
95
  private readonly listSandboxes: () => string[]
83
96
  private readonly importResults: boolean
84
97
  private readonly scratchRoot: string
98
+ private readonly includeProjectData: boolean
99
+ private readonly projectIdentityRoot?: string
100
+ private readonly networkPolicy: Record<string, unknown>
101
+ private readonly commandTimeoutMs: number
102
+ private readonly gpu?: number
85
103
  private readonly workspaces = new Map<string, OpenShellWorkspaceState>()
86
104
 
87
105
  constructor(options: OpenShellSessionProviderOptions = {}) {
88
106
  this.image = options.image ?? process.env.HEADLESSCODE_OPENSHELL_IMAGE ?? "headlesscode-openshell:local"
89
107
  this.policy = options.policy ?? process.env.HEADLESSCODE_OPENSHELL_POLICY
90
108
  this.providers = options.providers ?? (process.env.HEADLESSCODE_OPENSHELL_PROVIDERS ?? "headlesscode-openrouter").split(",").map((name) => name.trim()).filter(Boolean)
109
+ this.autoProviders = options.autoProviders ?? true
91
110
  this.cpu = options.cpu ?? process.env.HEADLESSCODE_OPENSHELL_CPU ?? "2"
92
111
  this.memory = options.memory ?? process.env.HEADLESSCODE_OPENSHELL_MEMORY ?? "4Gi"
93
112
  this.namePrefix = options.namePrefix ?? process.env.HEADLESSCODE_OPENSHELL_NAME_PREFIX ?? "hcls"
94
113
  this.readOnlyMounts = options.readOnlyMounts ?? []
95
114
  this.importResults = options.importResults ?? true
96
115
  this.scratchRoot = path.resolve(options.scratchRoot ?? process.env.HEADLESSCODE_OPENSHELL_SCRATCH_ROOT ?? path.join(os.homedir(), ".local", "share", "headlesscode", "openshell-sessions"))
116
+ this.includeProjectData = options.includeProjectData ?? true
117
+ this.projectIdentityRoot = options.projectIdentityRoot
118
+ this.networkPolicy = options.networkPolicy ?? {}
119
+ this.commandTimeoutMs = options.commandTimeoutMs ?? 15 * 60_000
120
+ if (options.gpu !== undefined && (!Number.isInteger(options.gpu) || options.gpu < 1 || options.gpu > 8)) throw new Error("OpenShell GPU count must be an integer from 1 to 8")
121
+ this.gpu = options.gpu
97
122
  this.run = options.run ?? ((args, timeoutMs) => {
98
123
  const result = spawnSync("openshell", args, { encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: timeoutMs })
99
124
  return { exitCode: result.status ?? 1, output: `${result.stdout ?? ""}${result.stderr ?? ""}`.trim() }
@@ -258,7 +283,8 @@ export class OpenShellSessionProvider implements CloudProvider {
258
283
  ): Promise<SessionHandle> {
259
284
  const { sandboxWorkspace } = isolated
260
285
  if (!this.policy || !fs.existsSync(this.policy)) throw new Error("OpenShell requires a base policy file")
261
- const args = ["sandbox", "create", "--detach", "--auto-providers", "--cpu", this.cpu, "--memory", this.memory, "--name", name, "--from", this.image]
286
+ const args = ["sandbox", "create", "--detach", this.autoProviders ? "--auto-providers" : "--no-auto-providers", "--cpu", this.cpu, "--memory", this.memory, "--name", name, "--from", this.image]
287
+ if (this.gpu !== undefined) args.push("--gpu", String(this.gpu))
262
288
  if (this.policy) args.push("--policy", path.resolve(this.policy))
263
289
  for (const provider of this.providers) args.push("--provider", provider)
264
290
  // Only the disposable clone is writable inside the sandbox. Host Git
@@ -268,13 +294,7 @@ export class OpenShellSessionProvider implements CloudProvider {
268
294
  { type: "bind", source: sandboxWorkspace, target: OPENSHELL_WORKSPACE_TARGET, read_only: false },
269
295
  { type: "bind", source: isolated.exportDir, target: "/result", read_only: false },
270
296
  ]
271
- const projectData = resolveProjectDataDir(repo)
272
- const projectDataTarget = "/opt/headlesscode-data/projects/" + path.basename(projectData)
273
297
  const dataSnapshotRoot = path.join(isolated.tempRoot, "data")
274
- const projectDataSnapshot = path.join(dataSnapshotRoot, "projects", path.basename(projectData))
275
- fs.mkdirSync(path.dirname(projectDataSnapshot), { recursive: true })
276
- fs.cpSync(projectData, projectDataSnapshot, { recursive: true, dereference: true })
277
- mounts.push({ type: "bind", source: projectDataSnapshot, target: projectDataTarget, read_only: true })
278
298
  const policy = this.policy ? parse(fs.readFileSync(this.policy, "utf8")) as Record<string, any> : undefined
279
299
  const filesystemPolicy = policy?.filesystem_policy
280
300
  if (!filesystemPolicy || !Array.isArray(filesystemPolicy.read_only) || !Array.isArray(filesystemPolicy.read_write)) {
@@ -282,18 +302,33 @@ export class OpenShellSessionProvider implements CloudProvider {
282
302
  }
283
303
  filesystemPolicy.read_write.push(OPENSHELL_WORKSPACE_TARGET)
284
304
  filesystemPolicy.read_write.push("/result")
285
- filesystemPolicy.read_only.push(projectDataTarget, OPENSHELL_HARNESS_ROOT)
286
- const sharedData = path.join(projectStoreRoot(), "shared")
287
- if (fs.existsSync(sharedData)) {
288
- const sharedDataSnapshot = path.join(dataSnapshotRoot, "shared")
289
- fs.cpSync(sharedData, sharedDataSnapshot, { recursive: true, dereference: true })
290
- mounts.push({ type: "bind", source: sharedDataSnapshot, target: "/opt/headlesscode-data/shared", read_only: true })
291
- filesystemPolicy.read_only.push("/opt/headlesscode-data/shared")
305
+ policy.network_policies = { ...(policy.network_policies ?? {}), ...this.networkPolicy }
306
+ filesystemPolicy.read_only.push(OPENSHELL_HARNESS_ROOT)
307
+ let projectDataMounted = false
308
+ if (this.includeProjectData) {
309
+ const projectData = resolveProjectDataDir(repo)
310
+ if (fs.existsSync(projectData)) {
311
+ const projectDataTarget = "/opt/headlesscode-data/projects/" + path.basename(projectData)
312
+ const projectDataSnapshot = path.join(dataSnapshotRoot, "projects", path.basename(projectData))
313
+ fs.mkdirSync(path.dirname(projectDataSnapshot), { recursive: true })
314
+ fs.cpSync(projectData, projectDataSnapshot, { recursive: true, dereference: true })
315
+ mounts.push({ type: "bind", source: projectDataSnapshot, target: projectDataTarget, read_only: true })
316
+ filesystemPolicy.read_only.push(projectDataTarget)
317
+ projectDataMounted = true
318
+ }
319
+ const sharedData = path.join(projectStoreRoot(), "shared")
320
+ if (fs.existsSync(sharedData)) {
321
+ const sharedDataSnapshot = path.join(dataSnapshotRoot, "shared")
322
+ fs.cpSync(sharedData, sharedDataSnapshot, { recursive: true, dereference: true })
323
+ mounts.push({ type: "bind", source: sharedDataSnapshot, target: "/opt/headlesscode-data/shared", read_only: true })
324
+ filesystemPolicy.read_only.push("/opt/headlesscode-data/shared")
325
+ }
292
326
  }
293
327
  const checkpointsSnapshot = path.join(dataSnapshotRoot, "checkpoints")
294
328
  fs.mkdirSync(checkpointsSnapshot, { recursive: true })
295
329
  mounts.push({ type: "bind", source: checkpointsSnapshot, target: "/opt/headlesscode-data/checkpoints", read_only: false })
296
- filesystemPolicy.read_only.push("/opt/headlesscode-data")
330
+ if (projectDataMounted) filesystemPolicy.read_only.push("/opt/headlesscode-data")
331
+ else filesystemPolicy.read_write.push("/opt/headlesscode-data")
297
332
  filesystemPolicy.read_write.push("/opt/headlesscode-data/checkpoints")
298
333
  const memoryDir = request.env?.HEADLESSCODE_MEMORY_DIR
299
334
  if (memoryDir) {
@@ -312,12 +347,13 @@ export class OpenShellSessionProvider implements CloudProvider {
312
347
  mounts.push({ type: "bind", source: memorySnapshot, target: "/opt/headlesscode-memory", read_only: false })
313
348
  filesystemPolicy.read_write.push("/opt/headlesscode-memory")
314
349
  }
315
- for (const mount of this.readOnlyMounts) {
316
- const relativeTarget = path.relative(OPENSHELL_WORKSPACE_TARGET, path.resolve(mount.target))
317
- if (relativeTarget.startsWith("..") || path.isAbsolute(relativeTarget)) throw new Error(`OpenShell read-only mount target must be under ${OPENSHELL_WORKSPACE_TARGET}: ${mount.target}`)
318
- const target = path.join(sandboxWorkspace, relativeTarget)
319
- fs.mkdirSync(path.dirname(target), { recursive: true })
320
- fs.cpSync(path.resolve(mount.source), target, { recursive: true, dereference: true })
350
+ const mountTargets = validateReadOnlyMountTargets(this.readOnlyMounts)
351
+ for (const [index, mount] of this.readOnlyMounts.entries()) {
352
+ const sourceSnapshot = path.join(isolated.tempRoot, "read-only-mounts", String(index))
353
+ fs.mkdirSync(path.dirname(sourceSnapshot), { recursive: true, mode: 0o700 })
354
+ fs.cpSync(path.resolve(mount.source), sourceSnapshot, { recursive: true, dereference: true })
355
+ mounts.push({ type: "bind", source: sourceSnapshot, target: mountTargets[index]!, read_only: true })
356
+ filesystemPolicy.read_only.push(mountTargets[index]!)
321
357
  }
322
358
  const generatedPolicy = path.join(os.tmpdir(), `headlesscode-openshell-policy-${process.pid}-${Date.now()}.yaml`)
323
359
  fs.writeFileSync(generatedPolicy, stringify(policy), { mode: 0o600 })
@@ -327,13 +363,13 @@ export class OpenShellSessionProvider implements CloudProvider {
327
363
  if (key === "HEADLESSCODE_MEMORY_DIR" && isolated.memorySnapshot) args.push("--env", "HEADLESSCODE_MEMORY_DIR=/opt/headlesscode-memory")
328
364
  else if (key !== "TARGET_REPO") args.push("--env", `${key}=${value}`)
329
365
  }
330
- args.push("--env", `TARGET_REPO=${OPENSHELL_WORKSPACE_TARGET}`, "--env", `HEADLESSCODE_ROOT=${OPENSHELL_HARNESS_ROOT}`, "--env", "HEADLESSCODE_DATA_DIR=/opt/headlesscode-data", "--env", `HEADLESSCODE_PROJECT_IDENTITY_ROOT=${repo}`, "--env", "GIT_CONFIG_NOSYSTEM=1", "--env", "GIT_CONFIG_GLOBAL=/dev/null")
366
+ args.push("--env", `TARGET_REPO=${OPENSHELL_WORKSPACE_TARGET}`, "--env", `HEADLESSCODE_ROOT=${OPENSHELL_HARNESS_ROOT}`, "--env", "HEADLESSCODE_DATA_DIR=/opt/headlesscode-data", "--env", `HEADLESSCODE_PROJECT_IDENTITY_ROOT=${this.projectIdentityRoot ?? repo}`, "--env", "GIT_CONFIG_NOSYSTEM=1", "--env", "GIT_CONFIG_GLOBAL=/dev/null")
331
367
  args.push("--", "sleep", "infinity")
332
368
  this.workspaces.set(name, isolated)
333
369
  writeOpenShellSandboxMarker(hostWorkspace, name)
334
370
  let result: CommandResult
335
371
  try {
336
- result = this.run(args)
372
+ result = this.run(args, this.commandTimeoutMs)
337
373
  } finally {
338
374
  fs.rmSync(generatedPolicy, { force: true })
339
375
  }
@@ -403,7 +439,7 @@ export class OpenShellSessionProvider implements CloudProvider {
403
439
  "for f in /workspace/.harness.* /workspace/.qa.* /workspace/.review.*; do [ -e \"$f\" ] || continue; rel=\"${f#/workspace/}\"; [ -z \"$(git ls-files -- \"$rel\")\" ] && rm -f -- \"$f\"; done",
404
440
  "git ls-files -z -- .env '.env.*' | while IFS= read -r -d '' rel; do git show \"HEAD:$rel\" > \"/workspace/$rel\"; done",
405
441
  "for f in /workspace/.env /workspace/.env.*; do [ -e \"$f\" ] || continue; rel=\"${f#/workspace/}\"; git ls-files --error-unmatch -- \"$rel\" >/dev/null 2>&1 || rm -f -- \"$f\"; done",
406
- "git add -A",
442
+ "git rm -r --cached --ignore-unmatch -- .rsi-base-model && git add -A -- . ':!.rsi-base-model'",
407
443
  "if ! git diff --cached --quiet HEAD; then git -c core.hooksPath=/dev/null commit -m 'HeadlessCode sandbox result'; fi",
408
444
  `git bundle create /result/result.bundle ${branchRef}`,
409
445
  ].join(" && ")
@@ -559,6 +595,30 @@ function pathsOverlap(left: string, right: string): boolean {
559
595
  return isSameOrInside(left, right) || isSameOrInside(right, left)
560
596
  }
561
597
 
598
+ /** Validate nested read-only bind targets before taking filesystem snapshots. */
599
+ export function validateReadOnlyMountTargets(readOnlyMounts: Array<{ source: string; target: string }>): string[] {
600
+ const targets = readOnlyMounts.map(({ target }) => {
601
+ if (typeof target !== "string" || target.includes("\0") || target.includes("\\") || /[\r\n]/.test(target)) {
602
+ throw new Error(`OpenShell read-only mount target is unsafe: ${String(target)}`)
603
+ }
604
+ const normalized = path.posix.normalize(target)
605
+ if (target === "/" || !target.startsWith(`${OPENSHELL_WORKSPACE_TARGET}/`) || normalized !== target || target.slice(1).split("/").some((part) => part === ".." || part === "." || part === "")) {
606
+ throw new Error(`OpenShell read-only mount target must be a normalized non-root path beneath ${OPENSHELL_WORKSPACE_TARGET}: ${target}`)
607
+ }
608
+ return normalized
609
+ })
610
+ for (let left = 0; left < targets.length; left += 1) {
611
+ for (let right = left + 1; right < targets.length; right += 1) {
612
+ const a = targets[left]!
613
+ const b = targets[right]!
614
+ if (a === b || a.startsWith(`${b}/`) || b.startsWith(`${a}/`)) {
615
+ throw new Error(`OpenShell read-only mount targets overlap: ${a} and ${b}`)
616
+ }
617
+ }
618
+ }
619
+ return targets
620
+ }
621
+
562
622
  function isSameOrInside(parent: string, child: string): boolean {
563
623
  const relative = path.relative(path.resolve(parent), path.resolve(child))
564
624
  return relative === "" || (!relative.startsWith(`..${path.sep}`) && relative !== ".." && !path.isAbsolute(relative))
package/src/llm/ollama.ts CHANGED
@@ -87,6 +87,14 @@ function envThinkEnabled(env: NodeJS.ProcessEnv = process.env): boolean {
87
87
  return v !== undefined && v !== "" && v !== "0" && v.toLowerCase() !== "false"
88
88
  }
89
89
 
90
+ function isOpenShellHost(baseUrl: string): boolean {
91
+ try {
92
+ return new URL(baseUrl).hostname === "host.openshell.internal"
93
+ } catch {
94
+ return false
95
+ }
96
+ }
97
+
90
98
  export class OllamaClient implements LlmClient {
91
99
  private readonly baseUrl: string
92
100
  private readonly defaultModel: string
@@ -108,26 +116,21 @@ export class OllamaClient implements LlmClient {
108
116
  * real prefill+generation call against a 9B local model at deep prompt
109
117
  * lengths can legitimately take longer than undici's default assumes.
110
118
  *
111
- * Overriding this requires a per-instance `Agent` with both timeouts
112
- * derived from `this.timeoutMs` (comfortable margin above it) — but
113
- * passing an `Agent` built from the standalone `undici` npm package as
114
- * Node's GLOBAL `fetch`'s `dispatcher` does not reliably work: verified
115
- * live, this threw `InvalidArgumentError: invalid onRequestStart
116
- * method` (code UND_ERR_INVALID_ARG) — Node's global fetch runs on its
117
- * OWN internal, bundled copy of undici (`node:internal/deps/undici`),
118
- * whose Dispatcher/Handler interface isn't guaranteed to match whatever
119
- * version the standalone package resolves to. The fix is to use
120
- * undici's own `fetch` (imported below as `undiciFetch`) together with
121
- * its own `Agent`, so both come from the SAME package/version and the
122
- * interface always matches — never Node's global `fetch` plus an
123
- * externally-built dispatcher.
119
+ * For direct localhost and remote daemon endpoints, use the standalone
120
+ * undici fetch and a matching Agent so these timeouts are configurable.
121
+ * OpenShell's host relay must use Node's global fetch: a standalone undici
122
+ * Agent bypasses OpenShell's transparent network path and the relay closes
123
+ * the connection for larger chat requests. The OpenShell worker has an
124
+ * outer command timeout, while global fetch retains its conservative
125
+ * built-in header/body timeouts.
124
126
  */
125
- private readonly dispatcher: Agent
127
+ private readonly dispatcher?: Agent
126
128
 
127
129
  constructor(options: OllamaClientOptions = {}) {
128
130
  this.baseUrl = (options.baseUrl ?? process.env[OLLAMA_URL_ENV] ?? DEFAULT_OLLAMA_URL).replace(/\/+$/, "")
129
131
  this.defaultModel = options.defaultModel ?? ""
130
132
  this.timeoutMs = options.timeoutMs ?? DEFAULT_OLLAMA_TIMEOUT_MS
133
+ const usesOpenShellHost = isOpenShellHost(this.baseUrl)
131
134
  // Default to undici's OWN fetch, not Node's global one — see
132
135
  // `dispatcher`'s doc comment for why: only undici's own fetch is
133
136
  // guaranteed interface-compatible with an Agent built from the same
@@ -135,13 +138,16 @@ export class OllamaClient implements LlmClient {
135
138
  // (structurally close enough for test fakes to satisfy); the cast
136
139
  // here just reconciles undici's own Response/RequestInit types
137
140
  // against the DOM-lib ones the public field declares.
138
- const defaultFetchImpl: unknown = (url: unknown, init: unknown) =>
139
- undiciFetch(url as Parameters<typeof undiciFetch>[0], init as Parameters<typeof undiciFetch>[1])
141
+ const defaultFetchImpl: unknown = usesOpenShellHost
142
+ ? globalThis.fetch
143
+ : (url: unknown, init: unknown) => undiciFetch(url as Parameters<typeof undiciFetch>[0], init as Parameters<typeof undiciFetch>[1])
140
144
  this.fetchImpl = options.fetchImpl ?? (defaultFetchImpl as typeof fetch)
141
145
  this.think = options.think ?? envThinkEnabled()
142
146
  this.nodeId = options.nodeId ?? randomUUID()
143
- const undiciTimeoutMs = this.timeoutMs + 60_000
144
- this.dispatcher = new Agent({ headersTimeout: undiciTimeoutMs, bodyTimeout: undiciTimeoutMs })
147
+ if (!usesOpenShellHost) {
148
+ const undiciTimeoutMs = this.timeoutMs + 60_000
149
+ this.dispatcher = new Agent({ headersTimeout: undiciTimeoutMs, bodyTimeout: undiciTimeoutMs })
150
+ }
145
151
  }
146
152
 
147
153
  resolveModel(requestModel?: string): string {
@@ -173,7 +179,7 @@ export class OllamaClient implements LlmClient {
173
179
  // against the DOM-lib `RequestInit` type this call site is
174
180
  // statically typed against. A fake fetchImpl injected in
175
181
  // tests simply ignores this extra field.
176
- dispatcher: this.dispatcher as unknown as RequestInit["dispatcher"],
182
+ ...(this.dispatcher ? { dispatcher: this.dispatcher as unknown as RequestInit["dispatcher"] } : {}),
177
183
  body: JSON.stringify({
178
184
  model,
179
185
  node_id: this.nodeId,
@@ -99,7 +99,10 @@ export function resolveProjectIdentity(workspaceRoot: string): ProjectIdentity {
99
99
  // expose host repository metadata. Preserve the original project's store key
100
100
  // without mounting or reading the original repository path in the sandbox.
101
101
  const identityRoot = process.env.HEADLESSCODE_PROJECT_IDENTITY_ROOT?.trim()
102
- if (identityRoot && path.isAbsolute(identityRoot)) return { keySource: path.resolve(identityRoot), kind: "git" }
102
+ const targetRepo = process.env.TARGET_REPO?.trim()
103
+ if (identityRoot && path.isAbsolute(identityRoot) && targetRepo && root === path.resolve(targetRepo)) {
104
+ return { keySource: path.resolve(identityRoot), kind: "git" }
105
+ }
103
106
  // Sandboxed workers may set these to a private Git directory so ordinary
104
107
  // commits cannot touch shared host metadata. Project identity must still
105
108
  // resolve through the workspace's real .git pointer and shared common dir.
@@ -0,0 +1,49 @@
1
+ import type { AdaptiveSearchDecision, CandidateRecord, RsiConfig } from "./types.js"
2
+
3
+ /** Decide whether a bounded adaptive run should allocate one follow-up trajectory. */
4
+ export function decideAdaptiveContinuation(input: {
5
+ generation: number
6
+ generationCandidates: CandidateRecord[]
7
+ allCandidates: CandidateRecord[]
8
+ config: RsiConfig
9
+ elapsedMs: number
10
+ decidedAt: string
11
+ }): AdaptiveSearchDecision {
12
+ const { generation, generationCandidates, allCandidates, config, elapsedMs, decidedAt } = input
13
+ const maxTrajectories = config.maxTrajectories ?? config.population + 1
14
+ const maxIterations = config.maxTotalIterations ?? maxTrajectories * config.maxIterations
15
+ const allocatedTrajectories = allCandidates.length
16
+ const allocatedIterations = allocatedTrajectories * config.maxIterations
17
+ const evidence = generationCandidates.map((candidate) => ({
18
+ candidateId: candidate.id,
19
+ status: candidate.status,
20
+ ...(candidate.fitness
21
+ ? { visiblePassRate: candidate.fitness.metrics.generalization }
22
+ : {}),
23
+ }))
24
+ const distinctResults = new Set(evidence.map((entry) => `${entry.status}:${entry.visiblePassRate ?? "unknown"}`))
25
+ const mixed = evidence.length >= 2 && distinctResults.size > 1
26
+ let reason: AdaptiveSearchDecision["reason"]
27
+ if (allocatedTrajectories >= maxTrajectories) reason = "trajectory-cap"
28
+ else if (allocatedIterations + config.maxIterations > maxIterations) reason = "iteration-cap"
29
+ else if (elapsedMs >= (config.maxRuntimeMs ?? 60 * 60_000)) reason = "runtime-cap"
30
+ else if (generation >= config.generations) reason = "generation-cap"
31
+ else if (evidence.length < 2) reason = "insufficient-results"
32
+ else if (!mixed) reason = "consistent-evidence"
33
+ else reason = "mixed-evidence"
34
+ const remainingTrajectories = Math.max(0, maxTrajectories - allocatedTrajectories)
35
+ const remainingIterations = Math.max(0, maxIterations - allocatedIterations)
36
+ return {
37
+ generation,
38
+ decision: reason === "mixed-evidence" ? "continue" : "stop",
39
+ reason,
40
+ candidateIds: generationCandidates.map((candidate) => candidate.id),
41
+ evidence,
42
+ allocatedTrajectories,
43
+ remainingTrajectories,
44
+ allocatedIterations,
45
+ remainingIterations,
46
+ elapsedMs,
47
+ decidedAt,
48
+ }
49
+ }