headlesscode 1.2.2 → 1.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +46 -16
- package/package.json +7 -3
- package/src/cli.ts +47 -18
- package/src/cloud/openshell-preflight.ts +2 -2
- package/src/cloud/openshell-provider.ts +84 -24
- package/src/llm/ollama.ts +25 -19
- package/src/project-store.ts +4 -1
- package/src/rsi/adaptive.ts +49 -0
- package/src/rsi/adversarial.ts +106 -0
- package/src/rsi/archive.ts +6 -5
- package/src/rsi/artifact-store.ts +158 -0
- package/src/rsi/config.ts +50 -2
- package/src/rsi/controller.ts +561 -42
- package/src/rsi/curriculum.ts +135 -16
- package/src/rsi/evaluator.ts +6 -37
- package/src/rsi/fitness.ts +39 -4
- package/src/rsi/index.ts +1 -0
- package/src/rsi/migrations/001_postgres_fleet_queue.sql +65 -0
- package/src/rsi/migrations/002_external_artifacts_and_job_leases.sql +39 -0
- package/src/rsi/migrations/003_model_training_jobs.sql +6 -0
- package/src/rsi/model-training.ts +256 -0
- package/src/rsi/mutation.ts +1 -77
- package/src/rsi/openshell.ts +639 -0
- package/src/rsi/postgres-queue.ts +424 -0
- package/src/rsi/promote-curriculum.ts +21 -0
- package/src/rsi/reports.ts +29 -2
- package/src/rsi/roles.ts +13 -3
- package/src/rsi/selection.ts +7 -1
- package/src/rsi/training-data.ts +103 -0
- package/src/rsi/trajectory.ts +1 -1
- package/src/rsi/types.ts +115 -2
- package/src/rsi/worker.ts +264 -0
- package/src/rsi/workspace.ts +14 -3
package/README.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
[](https://github.com/Capsize-Games/headlesscode/actions/workflows/ci.yml)
|
|
4
4
|
[](https://www.npmjs.com/package/headlesscode)
|
|
5
|
-
[](https://nodejs.org/)
|
|
6
6
|
[](./LICENSE)
|
|
7
7
|
|
|
8
8
|
`headlesscode` runs the Zoo Code agent loop from a Node.js process, without a
|
|
@@ -38,7 +38,7 @@ monitoring.
|
|
|
38
38
|
|
|
39
39
|
## Quick start
|
|
40
40
|
|
|
41
|
-
Requires Node.js
|
|
41
|
+
Requires Node.js 22.19 or newer. Install globally or run with `npx`:
|
|
42
42
|
|
|
43
43
|
```bash
|
|
44
44
|
npm install -g headlesscode
|
|
@@ -194,29 +194,59 @@ and each subcommand's `--help` for complete usage.
|
|
|
194
194
|
|
|
195
195
|
## Recursive self-improvement
|
|
196
196
|
|
|
197
|
-
`headlesscode improve`
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
|
|
201
|
-
|
|
197
|
+
`headlesscode improve --dry-run` resolves the base commit and prints planned
|
|
198
|
+
candidate worktrees. A real run requires PostgreSQL, a shared S3-compatible
|
|
199
|
+
artifact store, and at least one registered `headlesscode rsi-worker`. The
|
|
200
|
+
coordinator queues sanitized mutation snapshots and visible evaluation jobs;
|
|
201
|
+
workers run each in a fresh OpenShell guest. Fitness, paired baseline
|
|
202
|
+
comparison, archives, and selection stay in the coordinator. Hidden evaluation
|
|
203
|
+
remains disabled. A 2026-09-29 bounded run used two workers on the same local
|
|
204
|
+
OpenShell gateway; both candidate evaluation jobs were leased concurrently.
|
|
205
|
+
The run used a temporary `clamp.js` fixture and does not establish general
|
|
206
|
+
headlesscode improvement or multi-gateway operation. An earlier bounded run
|
|
207
|
+
passed visible checks but made no source change and was rejected.
|
|
202
208
|
|
|
203
209
|
```bash
|
|
204
210
|
npx tsx src/cli.ts improve --repo . --dry-run
|
|
205
211
|
npx tsx src/cli.ts improve --repo . --population 2 --generations 1
|
|
206
212
|
```
|
|
207
213
|
|
|
208
|
-
|
|
209
|
-
|
|
214
|
+
Before starting workers, build the standard and RSI guest images on each
|
|
215
|
+
OpenShell gateway host that will run RSI jobs. Build the larger training image
|
|
216
|
+
only on gateways whose workers will accept model-training or paired
|
|
217
|
+
model-evaluation jobs:
|
|
218
|
+
|
|
219
|
+
```bash
|
|
220
|
+
docker build -f docker/OpenShell.Dockerfile -t headlesscode-openshell:local .
|
|
221
|
+
docker build -f docker/OpenShell-RSI.Dockerfile -t headlesscode-openshell-rsi:local .
|
|
222
|
+
docker build -f docker/OpenShell-RSI-Training.Dockerfile -t headlesscode-openshell-rsi-training:local .
|
|
223
|
+
```
|
|
224
|
+
|
|
225
|
+
The RSI Dockerfile extends `headlesscode-openshell:local` and removes the RSI
|
|
226
|
+
source, test files, and hidden evaluation suite from the guest image. See the [RSI design](./docs/recursive-self-improvement.md)
|
|
227
|
+
for worker and queue configuration.
|
|
228
|
+
|
|
229
|
+
The queue stores job state, worker registrations, leases, retries, and admission
|
|
230
|
+
policy in PostgreSQL; artifact bytes remain in object storage. Local integration
|
|
231
|
+
tests used 12 simulated workers and 100 queued jobs, including a 20 MiB artifact. Each
|
|
232
|
+
worker process currently runs one guest at a time and must be configured with a
|
|
233
|
+
unique worker ID and OpenShell gateway ID. The two-worker run observed 4.008
|
|
234
|
+
seconds of concurrent evaluation leases on one gateway. A local PostgreSQL/
|
|
235
|
+
RustFS test drained the remaining 96 jobs with 12 simulated workers in 446 ms
|
|
236
|
+
(215.2 jobs/s), after the initial four jobs were claimed and completed to
|
|
237
|
+
verify admission limits;
|
|
238
|
+
claims still serialize on a queue-policy row, and neither result qualifies a
|
|
239
|
+
multi-host deployment. Remaining work is:
|
|
210
240
|
|
|
211
241
|
| Issue | Work |
|
|
212
242
|
| --- | --- |
|
|
213
|
-
| [#3](https://github.com/Capsize-Games/headlesscode/issues/3) |
|
|
214
|
-
| [#4](https://github.com/Capsize-Games/headlesscode/issues/4) |
|
|
215
|
-
| [#5](https://github.com/Capsize-Games/headlesscode/issues/5) |
|
|
216
|
-
| [#6](https://github.com/Capsize-Games/headlesscode/issues/6) |
|
|
217
|
-
| [#7](https://github.com/Capsize-Games/headlesscode/issues/7) |
|
|
218
|
-
| [#8](https://github.com/Capsize-Games/headlesscode/issues/8) | Adversarial
|
|
219
|
-
| [#9](https://github.com/Capsize-Games/headlesscode/issues/9) |
|
|
243
|
+
| [#3](https://github.com/Capsize-Games/headlesscode/issues/3) | OpenShell candidate isolation is implemented; live multi-gateway qualification remains. |
|
|
244
|
+
| [#4](https://github.com/Capsize-Games/headlesscode/issues/4) | Content-addressed artifacts are implemented; signed evaluator provenance and hidden evaluation remain. |
|
|
245
|
+
| [#5](https://github.com/Capsize-Games/headlesscode/issues/5) | PostgreSQL queue, leases, retries, admission, and configured workers are implemented; live fleet qualification remains. |
|
|
246
|
+
| [#6](https://github.com/Capsize-Games/headlesscode/issues/6) | Bounded adaptive independent search supports an initial population plus one evidence-driven follow-up; repeated stages remain out of scope. |
|
|
247
|
+
| [#7](https://github.com/Capsize-Games/headlesscode/issues/7) | Fixture-backed curriculum replay and explicit promotion are implemented; wording-to-capability measurement remains unproven. |
|
|
248
|
+
| [#8](https://github.com/Capsize-Games/headlesscode/issues/8) | Adversarial review and bounded OpenShell break tests are implemented; live provider-backed review is unvalidated. |
|
|
249
|
+
| [#9](https://github.com/Capsize-Games/headlesscode/issues/9) | An OpenShell QLoRA prototype is covered by mocked job tests; the supplied GGUF is rejected before enqueue, and no live training is validated. |
|
|
220
250
|
|
|
221
251
|
See the [RSI design](./docs/recursive-self-improvement.md) and
|
|
222
252
|
[run progress](./docs/rsi-progress.md).
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "headlesscode",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.3.0",
|
|
4
4
|
"description": "Standalone headless coding-agent harness: runs the Zoo Code agent loop (prompts, tools, modes) without a VS Code UI, driven by CLI, HTTP, and parallel worktree orchestration.",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
6
|
"repository": {
|
|
@@ -17,7 +17,7 @@
|
|
|
17
17
|
],
|
|
18
18
|
"type": "module",
|
|
19
19
|
"engines": {
|
|
20
|
-
"node": ">=
|
|
20
|
+
"node": ">=22.19"
|
|
21
21
|
},
|
|
22
22
|
"bin": {
|
|
23
23
|
"headlesscode": "bin/headlesscode.mjs"
|
|
@@ -42,14 +42,17 @@
|
|
|
42
42
|
"start": "tsx src/cli.ts",
|
|
43
43
|
"cli": "tsx src/cli.ts",
|
|
44
44
|
"monitor:pilot": "tsx scripts/monitor-pilot/run.ts",
|
|
45
|
+
"rsi:promote-curriculum": "tsx src/rsi/promote-curriculum.ts",
|
|
45
46
|
"test": "node scripts/run-tests.mjs",
|
|
46
47
|
"prepublishOnly": "npm ci && npm test"
|
|
47
48
|
},
|
|
48
49
|
"dependencies": {
|
|
50
|
+
"@aws-sdk/client-s3": "^3.1142.0",
|
|
49
51
|
"@modelcontextprotocol/sdk": "^1.30.0",
|
|
50
52
|
"@octokit/auth-app": "^8.2.0",
|
|
51
53
|
"fastest-levenshtein": "^1.0.16",
|
|
52
54
|
"p-wait-for": "^5.0.2",
|
|
55
|
+
"pg": "^8.23.0",
|
|
53
56
|
"playwright": "^1.62.1",
|
|
54
57
|
"simple-git": "^3.36.0",
|
|
55
58
|
"tsx": "^4.19.0",
|
|
@@ -59,6 +62,7 @@
|
|
|
59
62
|
"zod": "^3.25.76"
|
|
60
63
|
},
|
|
61
64
|
"devDependencies": {
|
|
62
|
-
"@types/node": "^22.10.0"
|
|
65
|
+
"@types/node": "^22.10.0",
|
|
66
|
+
"@types/pg": "^8.23.1"
|
|
63
67
|
}
|
|
64
68
|
}
|
package/src/cli.ts
CHANGED
|
@@ -43,7 +43,6 @@ import { analyzeCliMain } from "./orchestrator/analyze-cli.js"
|
|
|
43
43
|
import { costHistoryCliMain } from "./orchestrator/cost-history-cli.js"
|
|
44
44
|
import { resolvePermissions, type PermissionsConfig } from "./permissions/config.js"
|
|
45
45
|
import { resolveModelForMode, resolveReasoningEffortForMode } from "./config/mode-models.js"
|
|
46
|
-
import { improveMain } from "./rsi/controller.js"
|
|
47
46
|
import { openShellSessionMain } from "./cloud/openshell-session.js"
|
|
48
47
|
|
|
49
48
|
const VERSION = "0.1.0"
|
|
@@ -751,10 +750,47 @@ export async function main(argv: string[] = process.argv.slice(2)): Promise<numb
|
|
|
751
750
|
if (argv[0] === "openshell-session") {
|
|
752
751
|
return openShellSessionMain(argv.slice(1))
|
|
753
752
|
}
|
|
753
|
+
if (argv[0] === "rsi-worker") {
|
|
754
|
+
const { rsiWorkerMain } = await import("./rsi/worker.js")
|
|
755
|
+
return rsiWorkerMain()
|
|
756
|
+
}
|
|
757
|
+
if (argv[0] === "rsi-train-model") {
|
|
758
|
+
const { parseRsiArgs } = await import("./rsi/config.js")
|
|
759
|
+
const { readArchive } = await import("./rsi/archive.js")
|
|
760
|
+
const { PostgresRsiJobQueue, fleetQueuePolicy } = await import("./rsi/postgres-queue.js")
|
|
761
|
+
const { runBoundedModelCandidate } = await import("./rsi/model-training.js")
|
|
762
|
+
const parsed = parseRsiArgs(argv.slice(1))
|
|
763
|
+
if (parsed.help) {
|
|
764
|
+
process.stdout.write("Usage: headlesscode rsi-train-model --repo <path> --archive-dir <path> --resume <run-id> --model-candidate-id <id>\nThe registered worker must use an OpenShell training image and a local Hugging Face checkpoint.\n")
|
|
765
|
+
return 0
|
|
766
|
+
}
|
|
767
|
+
if (parsed.error || !parsed.config || !parsed.config.resumeRunId || !parsed.config.modelCandidateId || parsed.config.dryRun) {
|
|
768
|
+
process.stderr.write(`headlesscode rsi-train-model: ${parsed.error ?? "--resume and --model-candidate-id are required; dry-run is unsupported"}\n`)
|
|
769
|
+
return 2
|
|
770
|
+
}
|
|
771
|
+
const config = parsed.config
|
|
772
|
+
const archive = await readArchive(config.archiveDir)
|
|
773
|
+
const run = archive.activeRuns.find((entry) => entry.runId === config.resumeRunId) ?? archive.runs.find((entry) => entry.runId === config.resumeRunId)
|
|
774
|
+
if (!run) {
|
|
775
|
+
process.stderr.write(`headlesscode rsi-train-model: RSI run not found in ${config.archiveDir}\n`)
|
|
776
|
+
return 2
|
|
777
|
+
}
|
|
778
|
+
const policy = fleetQueuePolicy(process.env.HEADLESSCODE_RSI_QUEUE_NAME?.trim() || "rsi-default", config.maxConcurrent, process.env)
|
|
779
|
+
const queue = PostgresRsiJobQueue.fromEnvironment(policy, process.env)
|
|
780
|
+
try {
|
|
781
|
+
await queue.migrate()
|
|
782
|
+
const model = await runBoundedModelCandidate(config, run, queue, config.modelCandidateId!)
|
|
783
|
+
process.stdout.write(`[rsi-model] ${model.id}: ${model.status}; eligible=${model.eligibleForSelection === true}\n`)
|
|
784
|
+
return model.eligibleForSelection ? 0 : 1
|
|
785
|
+
} finally {
|
|
786
|
+
await queue.close()
|
|
787
|
+
}
|
|
788
|
+
}
|
|
754
789
|
// Bounded recursive self-improvement: the supervisor owns the evaluator,
|
|
755
790
|
// archive, and selection logic while each candidate runs in its own
|
|
756
791
|
// worktree. See docs/recursive-self-improvement.md.
|
|
757
792
|
if (argv[0] === "improve") {
|
|
793
|
+
const { improveMain } = await import("./rsi/controller.js")
|
|
758
794
|
return improveMain(argv.slice(1))
|
|
759
795
|
}
|
|
760
796
|
|
|
@@ -1049,9 +1085,17 @@ export async function main(argv: string[] = process.argv.slice(2)): Promise<numb
|
|
|
1049
1085
|
}
|
|
1050
1086
|
}
|
|
1051
1087
|
|
|
1052
|
-
// ── real run:
|
|
1088
|
+
// ── real run: OpenRouter key required unless this mode is local ────────────
|
|
1089
|
+
const codeModeBackend = process.env.HEADLESSCODE_CODE_MODE_BACKEND ?? "openrouter"
|
|
1090
|
+
const localBackendModes = new Set(
|
|
1091
|
+
(process.env.HEADLESSCODE_LOCAL_BACKEND_MODES ?? "code")
|
|
1092
|
+
.split(",")
|
|
1093
|
+
.map((s) => s.trim())
|
|
1094
|
+
.filter(Boolean),
|
|
1095
|
+
)
|
|
1096
|
+
const useLocalCodeBackend = localBackendModes.has(options.mode) && codeModeBackend === "ollama"
|
|
1053
1097
|
const apiKey = process.env.HEADLESSCODE_OPENROUTER_API_KEY
|
|
1054
|
-
if (!apiKey) {
|
|
1098
|
+
if (!apiKey && !useLocalCodeBackend) {
|
|
1055
1099
|
process.stderr.write(
|
|
1056
1100
|
"headlesscode: HEADLESSCODE_OPENROUTER_API_KEY is not set.\n" +
|
|
1057
1101
|
" Export it (e.g. export HEADLESSCODE_OPENROUTER_API_KEY=sk-or-...) or use --dry-run to\n" +
|
|
@@ -1098,21 +1142,6 @@ export async function main(argv: string[] = process.argv.slice(2)): Promise<numb
|
|
|
1098
1142
|
env: process.env,
|
|
1099
1143
|
})
|
|
1100
1144
|
|
|
1101
|
-
// Generalized 2026-08-21 (issue #142) beyond the original `code`-only
|
|
1102
|
-
// scope (plans/local-dual-model-code-agent.md, D2) — real need: two
|
|
1103
|
-
// separate local daemons on two different GPUs (a coder model and a
|
|
1104
|
-
// review model), each needing its own mode -> URL/model mapping.
|
|
1105
|
-
// HEADLESSCODE_LOCAL_BACKEND_MODES defaults to "code" alone, so
|
|
1106
|
-
// nobody's existing setup changes behavior unless they opt in.
|
|
1107
|
-
const codeModeBackend = process.env.HEADLESSCODE_CODE_MODE_BACKEND ?? "openrouter"
|
|
1108
|
-
const localBackendModes = new Set(
|
|
1109
|
-
(process.env.HEADLESSCODE_LOCAL_BACKEND_MODES ?? "code")
|
|
1110
|
-
.split(",")
|
|
1111
|
-
.map((s) => s.trim())
|
|
1112
|
-
.filter(Boolean),
|
|
1113
|
-
)
|
|
1114
|
-
const useLocalCodeBackend = localBackendModes.has(options.mode) && codeModeBackend === "ollama"
|
|
1115
|
-
|
|
1116
1145
|
// The local daemon's own proxy (ollama_shim.py) deliberately runs a 600s
|
|
1117
1146
|
// upstream request timeout — its own comment documents why: a shorter
|
|
1118
1147
|
// shim timeout was once found to cut off calls before the harness's own
|
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
import { spawnSync } from "node:child_process"
|
|
2
2
|
import * as fs from "node:fs"
|
|
3
3
|
|
|
4
|
-
export function openshellPreflight(): string | undefined {
|
|
4
|
+
export function openshellPreflight(options: { requireOpenRouterCredential?: boolean } = {}): string | undefined {
|
|
5
5
|
const policy = process.env.HEADLESSCODE_OPENSHELL_POLICY
|
|
6
6
|
if (!policy || !fs.existsSync(policy)) return "set HEADLESSCODE_OPENSHELL_POLICY to an existing base policy YAML"
|
|
7
|
-
if (!process.env.HEADLESSCODE_OPENROUTER_API_KEY) return "set HEADLESSCODE_OPENROUTER_API_KEY so OpenShell can provide the imported credential profile"
|
|
7
|
+
if (options.requireOpenRouterCredential !== false && !process.env.HEADLESSCODE_OPENROUTER_API_KEY) return "set HEADLESSCODE_OPENROUTER_API_KEY so OpenShell can provide the imported credential profile"
|
|
8
8
|
const version = spawnSync("openshell", ["--version"], { encoding: "utf8", timeout: 5000 })
|
|
9
9
|
if (version.error || version.status !== 0) return "OpenShell CLI is unavailable"
|
|
10
10
|
const gateway = spawnSync("openshell", ["gateway", "info"], { encoding: "utf8", timeout: 10000 })
|
|
@@ -59,12 +59,24 @@ export interface OpenShellSessionProviderOptions {
|
|
|
59
59
|
image?: string
|
|
60
60
|
policy?: string
|
|
61
61
|
providers?: string[]
|
|
62
|
+
/** Attach automatically discovered host provider profiles to the sandbox. */
|
|
63
|
+
autoProviders?: boolean
|
|
62
64
|
cpu?: string
|
|
63
65
|
memory?: string
|
|
64
66
|
namePrefix?: string
|
|
65
67
|
readOnlyMounts?: Array<{ source: string; target: string }>
|
|
66
68
|
importResults?: boolean
|
|
67
69
|
scratchRoot?: string
|
|
70
|
+
/** Omit project/shared data snapshots for experiments that must not inherit operator context. */
|
|
71
|
+
includeProjectData?: boolean
|
|
72
|
+
/** Use a guest-safe identity path instead of the supervisor checkout path. */
|
|
73
|
+
projectIdentityRoot?: string
|
|
74
|
+
/** Additional exact OpenShell network rules needed by the workload. */
|
|
75
|
+
networkPolicy?: Record<string, unknown>
|
|
76
|
+
/** Timeout for sandbox commands, including agent mutation. */
|
|
77
|
+
commandTimeoutMs?: number
|
|
78
|
+
/** Number of GPU devices requested for a training-only sandbox. */
|
|
79
|
+
gpu?: number
|
|
68
80
|
run?: (args: string[], timeoutMs?: number) => CommandResult
|
|
69
81
|
listSandboxes?: () => string[]
|
|
70
82
|
}
|
|
@@ -74,6 +86,7 @@ export class OpenShellSessionProvider implements CloudProvider {
|
|
|
74
86
|
private readonly image: string
|
|
75
87
|
private readonly policy?: string
|
|
76
88
|
private readonly providers: string[]
|
|
89
|
+
private readonly autoProviders: boolean
|
|
77
90
|
private readonly cpu: string
|
|
78
91
|
private readonly memory: string
|
|
79
92
|
private readonly namePrefix: string
|
|
@@ -82,18 +95,30 @@ export class OpenShellSessionProvider implements CloudProvider {
|
|
|
82
95
|
private readonly listSandboxes: () => string[]
|
|
83
96
|
private readonly importResults: boolean
|
|
84
97
|
private readonly scratchRoot: string
|
|
98
|
+
private readonly includeProjectData: boolean
|
|
99
|
+
private readonly projectIdentityRoot?: string
|
|
100
|
+
private readonly networkPolicy: Record<string, unknown>
|
|
101
|
+
private readonly commandTimeoutMs: number
|
|
102
|
+
private readonly gpu?: number
|
|
85
103
|
private readonly workspaces = new Map<string, OpenShellWorkspaceState>()
|
|
86
104
|
|
|
87
105
|
constructor(options: OpenShellSessionProviderOptions = {}) {
|
|
88
106
|
this.image = options.image ?? process.env.HEADLESSCODE_OPENSHELL_IMAGE ?? "headlesscode-openshell:local"
|
|
89
107
|
this.policy = options.policy ?? process.env.HEADLESSCODE_OPENSHELL_POLICY
|
|
90
108
|
this.providers = options.providers ?? (process.env.HEADLESSCODE_OPENSHELL_PROVIDERS ?? "headlesscode-openrouter").split(",").map((name) => name.trim()).filter(Boolean)
|
|
109
|
+
this.autoProviders = options.autoProviders ?? true
|
|
91
110
|
this.cpu = options.cpu ?? process.env.HEADLESSCODE_OPENSHELL_CPU ?? "2"
|
|
92
111
|
this.memory = options.memory ?? process.env.HEADLESSCODE_OPENSHELL_MEMORY ?? "4Gi"
|
|
93
112
|
this.namePrefix = options.namePrefix ?? process.env.HEADLESSCODE_OPENSHELL_NAME_PREFIX ?? "hcls"
|
|
94
113
|
this.readOnlyMounts = options.readOnlyMounts ?? []
|
|
95
114
|
this.importResults = options.importResults ?? true
|
|
96
115
|
this.scratchRoot = path.resolve(options.scratchRoot ?? process.env.HEADLESSCODE_OPENSHELL_SCRATCH_ROOT ?? path.join(os.homedir(), ".local", "share", "headlesscode", "openshell-sessions"))
|
|
116
|
+
this.includeProjectData = options.includeProjectData ?? true
|
|
117
|
+
this.projectIdentityRoot = options.projectIdentityRoot
|
|
118
|
+
this.networkPolicy = options.networkPolicy ?? {}
|
|
119
|
+
this.commandTimeoutMs = options.commandTimeoutMs ?? 15 * 60_000
|
|
120
|
+
if (options.gpu !== undefined && (!Number.isInteger(options.gpu) || options.gpu < 1 || options.gpu > 8)) throw new Error("OpenShell GPU count must be an integer from 1 to 8")
|
|
121
|
+
this.gpu = options.gpu
|
|
97
122
|
this.run = options.run ?? ((args, timeoutMs) => {
|
|
98
123
|
const result = spawnSync("openshell", args, { encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: timeoutMs })
|
|
99
124
|
return { exitCode: result.status ?? 1, output: `${result.stdout ?? ""}${result.stderr ?? ""}`.trim() }
|
|
@@ -258,7 +283,8 @@ export class OpenShellSessionProvider implements CloudProvider {
|
|
|
258
283
|
): Promise<SessionHandle> {
|
|
259
284
|
const { sandboxWorkspace } = isolated
|
|
260
285
|
if (!this.policy || !fs.existsSync(this.policy)) throw new Error("OpenShell requires a base policy file")
|
|
261
|
-
const args = ["sandbox", "create", "--detach", "--auto-providers", "--cpu", this.cpu, "--memory", this.memory, "--name", name, "--from", this.image]
|
|
286
|
+
const args = ["sandbox", "create", "--detach", this.autoProviders ? "--auto-providers" : "--no-auto-providers", "--cpu", this.cpu, "--memory", this.memory, "--name", name, "--from", this.image]
|
|
287
|
+
if (this.gpu !== undefined) args.push("--gpu", String(this.gpu))
|
|
262
288
|
if (this.policy) args.push("--policy", path.resolve(this.policy))
|
|
263
289
|
for (const provider of this.providers) args.push("--provider", provider)
|
|
264
290
|
// Only the disposable clone is writable inside the sandbox. Host Git
|
|
@@ -268,13 +294,7 @@ export class OpenShellSessionProvider implements CloudProvider {
|
|
|
268
294
|
{ type: "bind", source: sandboxWorkspace, target: OPENSHELL_WORKSPACE_TARGET, read_only: false },
|
|
269
295
|
{ type: "bind", source: isolated.exportDir, target: "/result", read_only: false },
|
|
270
296
|
]
|
|
271
|
-
const projectData = resolveProjectDataDir(repo)
|
|
272
|
-
const projectDataTarget = "/opt/headlesscode-data/projects/" + path.basename(projectData)
|
|
273
297
|
const dataSnapshotRoot = path.join(isolated.tempRoot, "data")
|
|
274
|
-
const projectDataSnapshot = path.join(dataSnapshotRoot, "projects", path.basename(projectData))
|
|
275
|
-
fs.mkdirSync(path.dirname(projectDataSnapshot), { recursive: true })
|
|
276
|
-
fs.cpSync(projectData, projectDataSnapshot, { recursive: true, dereference: true })
|
|
277
|
-
mounts.push({ type: "bind", source: projectDataSnapshot, target: projectDataTarget, read_only: true })
|
|
278
298
|
const policy = this.policy ? parse(fs.readFileSync(this.policy, "utf8")) as Record<string, any> : undefined
|
|
279
299
|
const filesystemPolicy = policy?.filesystem_policy
|
|
280
300
|
if (!filesystemPolicy || !Array.isArray(filesystemPolicy.read_only) || !Array.isArray(filesystemPolicy.read_write)) {
|
|
@@ -282,18 +302,33 @@ export class OpenShellSessionProvider implements CloudProvider {
|
|
|
282
302
|
}
|
|
283
303
|
filesystemPolicy.read_write.push(OPENSHELL_WORKSPACE_TARGET)
|
|
284
304
|
filesystemPolicy.read_write.push("/result")
|
|
285
|
-
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
|
|
290
|
-
|
|
291
|
-
|
|
305
|
+
policy.network_policies = { ...(policy.network_policies ?? {}), ...this.networkPolicy }
|
|
306
|
+
filesystemPolicy.read_only.push(OPENSHELL_HARNESS_ROOT)
|
|
307
|
+
let projectDataMounted = false
|
|
308
|
+
if (this.includeProjectData) {
|
|
309
|
+
const projectData = resolveProjectDataDir(repo)
|
|
310
|
+
if (fs.existsSync(projectData)) {
|
|
311
|
+
const projectDataTarget = "/opt/headlesscode-data/projects/" + path.basename(projectData)
|
|
312
|
+
const projectDataSnapshot = path.join(dataSnapshotRoot, "projects", path.basename(projectData))
|
|
313
|
+
fs.mkdirSync(path.dirname(projectDataSnapshot), { recursive: true })
|
|
314
|
+
fs.cpSync(projectData, projectDataSnapshot, { recursive: true, dereference: true })
|
|
315
|
+
mounts.push({ type: "bind", source: projectDataSnapshot, target: projectDataTarget, read_only: true })
|
|
316
|
+
filesystemPolicy.read_only.push(projectDataTarget)
|
|
317
|
+
projectDataMounted = true
|
|
318
|
+
}
|
|
319
|
+
const sharedData = path.join(projectStoreRoot(), "shared")
|
|
320
|
+
if (fs.existsSync(sharedData)) {
|
|
321
|
+
const sharedDataSnapshot = path.join(dataSnapshotRoot, "shared")
|
|
322
|
+
fs.cpSync(sharedData, sharedDataSnapshot, { recursive: true, dereference: true })
|
|
323
|
+
mounts.push({ type: "bind", source: sharedDataSnapshot, target: "/opt/headlesscode-data/shared", read_only: true })
|
|
324
|
+
filesystemPolicy.read_only.push("/opt/headlesscode-data/shared")
|
|
325
|
+
}
|
|
292
326
|
}
|
|
293
327
|
const checkpointsSnapshot = path.join(dataSnapshotRoot, "checkpoints")
|
|
294
328
|
fs.mkdirSync(checkpointsSnapshot, { recursive: true })
|
|
295
329
|
mounts.push({ type: "bind", source: checkpointsSnapshot, target: "/opt/headlesscode-data/checkpoints", read_only: false })
|
|
296
|
-
filesystemPolicy.read_only.push("/opt/headlesscode-data")
|
|
330
|
+
if (projectDataMounted) filesystemPolicy.read_only.push("/opt/headlesscode-data")
|
|
331
|
+
else filesystemPolicy.read_write.push("/opt/headlesscode-data")
|
|
297
332
|
filesystemPolicy.read_write.push("/opt/headlesscode-data/checkpoints")
|
|
298
333
|
const memoryDir = request.env?.HEADLESSCODE_MEMORY_DIR
|
|
299
334
|
if (memoryDir) {
|
|
@@ -312,12 +347,13 @@ export class OpenShellSessionProvider implements CloudProvider {
|
|
|
312
347
|
mounts.push({ type: "bind", source: memorySnapshot, target: "/opt/headlesscode-memory", read_only: false })
|
|
313
348
|
filesystemPolicy.read_write.push("/opt/headlesscode-memory")
|
|
314
349
|
}
|
|
315
|
-
|
|
316
|
-
|
|
317
|
-
|
|
318
|
-
|
|
319
|
-
fs.
|
|
320
|
-
|
|
350
|
+
const mountTargets = validateReadOnlyMountTargets(this.readOnlyMounts)
|
|
351
|
+
for (const [index, mount] of this.readOnlyMounts.entries()) {
|
|
352
|
+
const sourceSnapshot = path.join(isolated.tempRoot, "read-only-mounts", String(index))
|
|
353
|
+
fs.mkdirSync(path.dirname(sourceSnapshot), { recursive: true, mode: 0o700 })
|
|
354
|
+
fs.cpSync(path.resolve(mount.source), sourceSnapshot, { recursive: true, dereference: true })
|
|
355
|
+
mounts.push({ type: "bind", source: sourceSnapshot, target: mountTargets[index]!, read_only: true })
|
|
356
|
+
filesystemPolicy.read_only.push(mountTargets[index]!)
|
|
321
357
|
}
|
|
322
358
|
const generatedPolicy = path.join(os.tmpdir(), `headlesscode-openshell-policy-${process.pid}-${Date.now()}.yaml`)
|
|
323
359
|
fs.writeFileSync(generatedPolicy, stringify(policy), { mode: 0o600 })
|
|
@@ -327,13 +363,13 @@ export class OpenShellSessionProvider implements CloudProvider {
|
|
|
327
363
|
if (key === "HEADLESSCODE_MEMORY_DIR" && isolated.memorySnapshot) args.push("--env", "HEADLESSCODE_MEMORY_DIR=/opt/headlesscode-memory")
|
|
328
364
|
else if (key !== "TARGET_REPO") args.push("--env", `${key}=${value}`)
|
|
329
365
|
}
|
|
330
|
-
args.push("--env", `TARGET_REPO=${OPENSHELL_WORKSPACE_TARGET}`, "--env", `HEADLESSCODE_ROOT=${OPENSHELL_HARNESS_ROOT}`, "--env", "HEADLESSCODE_DATA_DIR=/opt/headlesscode-data", "--env", `HEADLESSCODE_PROJECT_IDENTITY_ROOT=${repo}`, "--env", "GIT_CONFIG_NOSYSTEM=1", "--env", "GIT_CONFIG_GLOBAL=/dev/null")
|
|
366
|
+
args.push("--env", `TARGET_REPO=${OPENSHELL_WORKSPACE_TARGET}`, "--env", `HEADLESSCODE_ROOT=${OPENSHELL_HARNESS_ROOT}`, "--env", "HEADLESSCODE_DATA_DIR=/opt/headlesscode-data", "--env", `HEADLESSCODE_PROJECT_IDENTITY_ROOT=${this.projectIdentityRoot ?? repo}`, "--env", "GIT_CONFIG_NOSYSTEM=1", "--env", "GIT_CONFIG_GLOBAL=/dev/null")
|
|
331
367
|
args.push("--", "sleep", "infinity")
|
|
332
368
|
this.workspaces.set(name, isolated)
|
|
333
369
|
writeOpenShellSandboxMarker(hostWorkspace, name)
|
|
334
370
|
let result: CommandResult
|
|
335
371
|
try {
|
|
336
|
-
result = this.run(args)
|
|
372
|
+
result = this.run(args, this.commandTimeoutMs)
|
|
337
373
|
} finally {
|
|
338
374
|
fs.rmSync(generatedPolicy, { force: true })
|
|
339
375
|
}
|
|
@@ -403,7 +439,7 @@ export class OpenShellSessionProvider implements CloudProvider {
|
|
|
403
439
|
"for f in /workspace/.harness.* /workspace/.qa.* /workspace/.review.*; do [ -e \"$f\" ] || continue; rel=\"${f#/workspace/}\"; [ -z \"$(git ls-files -- \"$rel\")\" ] && rm -f -- \"$f\"; done",
|
|
404
440
|
"git ls-files -z -- .env '.env.*' | while IFS= read -r -d '' rel; do git show \"HEAD:$rel\" > \"/workspace/$rel\"; done",
|
|
405
441
|
"for f in /workspace/.env /workspace/.env.*; do [ -e \"$f\" ] || continue; rel=\"${f#/workspace/}\"; git ls-files --error-unmatch -- \"$rel\" >/dev/null 2>&1 || rm -f -- \"$f\"; done",
|
|
406
|
-
"git add -A",
|
|
442
|
+
"git rm -r --cached --ignore-unmatch -- .rsi-base-model && git add -A -- . ':!.rsi-base-model'",
|
|
407
443
|
"if ! git diff --cached --quiet HEAD; then git -c core.hooksPath=/dev/null commit -m 'HeadlessCode sandbox result'; fi",
|
|
408
444
|
`git bundle create /result/result.bundle ${branchRef}`,
|
|
409
445
|
].join(" && ")
|
|
@@ -559,6 +595,30 @@ function pathsOverlap(left: string, right: string): boolean {
|
|
|
559
595
|
return isSameOrInside(left, right) || isSameOrInside(right, left)
|
|
560
596
|
}
|
|
561
597
|
|
|
598
|
+
/** Validate nested read-only bind targets before taking filesystem snapshots. */
|
|
599
|
+
export function validateReadOnlyMountTargets(readOnlyMounts: Array<{ source: string; target: string }>): string[] {
|
|
600
|
+
const targets = readOnlyMounts.map(({ target }) => {
|
|
601
|
+
if (typeof target !== "string" || target.includes("\0") || target.includes("\\") || /[\r\n]/.test(target)) {
|
|
602
|
+
throw new Error(`OpenShell read-only mount target is unsafe: ${String(target)}`)
|
|
603
|
+
}
|
|
604
|
+
const normalized = path.posix.normalize(target)
|
|
605
|
+
if (target === "/" || !target.startsWith(`${OPENSHELL_WORKSPACE_TARGET}/`) || normalized !== target || target.slice(1).split("/").some((part) => part === ".." || part === "." || part === "")) {
|
|
606
|
+
throw new Error(`OpenShell read-only mount target must be a normalized non-root path beneath ${OPENSHELL_WORKSPACE_TARGET}: ${target}`)
|
|
607
|
+
}
|
|
608
|
+
return normalized
|
|
609
|
+
})
|
|
610
|
+
for (let left = 0; left < targets.length; left += 1) {
|
|
611
|
+
for (let right = left + 1; right < targets.length; right += 1) {
|
|
612
|
+
const a = targets[left]!
|
|
613
|
+
const b = targets[right]!
|
|
614
|
+
if (a === b || a.startsWith(`${b}/`) || b.startsWith(`${a}/`)) {
|
|
615
|
+
throw new Error(`OpenShell read-only mount targets overlap: ${a} and ${b}`)
|
|
616
|
+
}
|
|
617
|
+
}
|
|
618
|
+
}
|
|
619
|
+
return targets
|
|
620
|
+
}
|
|
621
|
+
|
|
562
622
|
function isSameOrInside(parent: string, child: string): boolean {
|
|
563
623
|
const relative = path.relative(path.resolve(parent), path.resolve(child))
|
|
564
624
|
return relative === "" || (!relative.startsWith(`..${path.sep}`) && relative !== ".." && !path.isAbsolute(relative))
|
package/src/llm/ollama.ts
CHANGED
|
@@ -87,6 +87,14 @@ function envThinkEnabled(env: NodeJS.ProcessEnv = process.env): boolean {
|
|
|
87
87
|
return v !== undefined && v !== "" && v !== "0" && v.toLowerCase() !== "false"
|
|
88
88
|
}
|
|
89
89
|
|
|
90
|
+
function isOpenShellHost(baseUrl: string): boolean {
|
|
91
|
+
try {
|
|
92
|
+
return new URL(baseUrl).hostname === "host.openshell.internal"
|
|
93
|
+
} catch {
|
|
94
|
+
return false
|
|
95
|
+
}
|
|
96
|
+
}
|
|
97
|
+
|
|
90
98
|
export class OllamaClient implements LlmClient {
|
|
91
99
|
private readonly baseUrl: string
|
|
92
100
|
private readonly defaultModel: string
|
|
@@ -108,26 +116,21 @@ export class OllamaClient implements LlmClient {
|
|
|
108
116
|
* real prefill+generation call against a 9B local model at deep prompt
|
|
109
117
|
* lengths can legitimately take longer than undici's default assumes.
|
|
110
118
|
*
|
|
111
|
-
*
|
|
112
|
-
*
|
|
113
|
-
*
|
|
114
|
-
*
|
|
115
|
-
*
|
|
116
|
-
*
|
|
117
|
-
*
|
|
118
|
-
* whose Dispatcher/Handler interface isn't guaranteed to match whatever
|
|
119
|
-
* version the standalone package resolves to. The fix is to use
|
|
120
|
-
* undici's own `fetch` (imported below as `undiciFetch`) together with
|
|
121
|
-
* its own `Agent`, so both come from the SAME package/version and the
|
|
122
|
-
* interface always matches — never Node's global `fetch` plus an
|
|
123
|
-
* externally-built dispatcher.
|
|
119
|
+
* For direct localhost and remote daemon endpoints, use the standalone
|
|
120
|
+
* undici fetch and a matching Agent so these timeouts are configurable.
|
|
121
|
+
* OpenShell's host relay must use Node's global fetch: a standalone undici
|
|
122
|
+
* Agent bypasses OpenShell's transparent network path and the relay closes
|
|
123
|
+
* the connection for larger chat requests. The OpenShell worker has an
|
|
124
|
+
* outer command timeout, while global fetch retains its conservative
|
|
125
|
+
* built-in header/body timeouts.
|
|
124
126
|
*/
|
|
125
|
-
private readonly dispatcher
|
|
127
|
+
private readonly dispatcher?: Agent
|
|
126
128
|
|
|
127
129
|
constructor(options: OllamaClientOptions = {}) {
|
|
128
130
|
this.baseUrl = (options.baseUrl ?? process.env[OLLAMA_URL_ENV] ?? DEFAULT_OLLAMA_URL).replace(/\/+$/, "")
|
|
129
131
|
this.defaultModel = options.defaultModel ?? ""
|
|
130
132
|
this.timeoutMs = options.timeoutMs ?? DEFAULT_OLLAMA_TIMEOUT_MS
|
|
133
|
+
const usesOpenShellHost = isOpenShellHost(this.baseUrl)
|
|
131
134
|
// Default to undici's OWN fetch, not Node's global one — see
|
|
132
135
|
// `dispatcher`'s doc comment for why: only undici's own fetch is
|
|
133
136
|
// guaranteed interface-compatible with an Agent built from the same
|
|
@@ -135,13 +138,16 @@ export class OllamaClient implements LlmClient {
|
|
|
135
138
|
// (structurally close enough for test fakes to satisfy); the cast
|
|
136
139
|
// here just reconciles undici's own Response/RequestInit types
|
|
137
140
|
// against the DOM-lib ones the public field declares.
|
|
138
|
-
const defaultFetchImpl: unknown =
|
|
139
|
-
|
|
141
|
+
const defaultFetchImpl: unknown = usesOpenShellHost
|
|
142
|
+
? globalThis.fetch
|
|
143
|
+
: (url: unknown, init: unknown) => undiciFetch(url as Parameters<typeof undiciFetch>[0], init as Parameters<typeof undiciFetch>[1])
|
|
140
144
|
this.fetchImpl = options.fetchImpl ?? (defaultFetchImpl as typeof fetch)
|
|
141
145
|
this.think = options.think ?? envThinkEnabled()
|
|
142
146
|
this.nodeId = options.nodeId ?? randomUUID()
|
|
143
|
-
|
|
144
|
-
|
|
147
|
+
if (!usesOpenShellHost) {
|
|
148
|
+
const undiciTimeoutMs = this.timeoutMs + 60_000
|
|
149
|
+
this.dispatcher = new Agent({ headersTimeout: undiciTimeoutMs, bodyTimeout: undiciTimeoutMs })
|
|
150
|
+
}
|
|
145
151
|
}
|
|
146
152
|
|
|
147
153
|
resolveModel(requestModel?: string): string {
|
|
@@ -173,7 +179,7 @@ export class OllamaClient implements LlmClient {
|
|
|
173
179
|
// against the DOM-lib `RequestInit` type this call site is
|
|
174
180
|
// statically typed against. A fake fetchImpl injected in
|
|
175
181
|
// tests simply ignores this extra field.
|
|
176
|
-
dispatcher: this.dispatcher as unknown as RequestInit["dispatcher"],
|
|
182
|
+
...(this.dispatcher ? { dispatcher: this.dispatcher as unknown as RequestInit["dispatcher"] } : {}),
|
|
177
183
|
body: JSON.stringify({
|
|
178
184
|
model,
|
|
179
185
|
node_id: this.nodeId,
|
package/src/project-store.ts
CHANGED
|
@@ -99,7 +99,10 @@ export function resolveProjectIdentity(workspaceRoot: string): ProjectIdentity {
|
|
|
99
99
|
// expose host repository metadata. Preserve the original project's store key
|
|
100
100
|
// without mounting or reading the original repository path in the sandbox.
|
|
101
101
|
const identityRoot = process.env.HEADLESSCODE_PROJECT_IDENTITY_ROOT?.trim()
|
|
102
|
-
|
|
102
|
+
const targetRepo = process.env.TARGET_REPO?.trim()
|
|
103
|
+
if (identityRoot && path.isAbsolute(identityRoot) && targetRepo && root === path.resolve(targetRepo)) {
|
|
104
|
+
return { keySource: path.resolve(identityRoot), kind: "git" }
|
|
105
|
+
}
|
|
103
106
|
// Sandboxed workers may set these to a private Git directory so ordinary
|
|
104
107
|
// commits cannot touch shared host metadata. Project identity must still
|
|
105
108
|
// resolve through the workspace's real .git pointer and shared common dir.
|
|
@@ -0,0 +1,49 @@
|
|
|
1
|
+
import type { AdaptiveSearchDecision, CandidateRecord, RsiConfig } from "./types.js"
|
|
2
|
+
|
|
3
|
+
/** Decide whether a bounded adaptive run should allocate one follow-up trajectory. */
|
|
4
|
+
export function decideAdaptiveContinuation(input: {
|
|
5
|
+
generation: number
|
|
6
|
+
generationCandidates: CandidateRecord[]
|
|
7
|
+
allCandidates: CandidateRecord[]
|
|
8
|
+
config: RsiConfig
|
|
9
|
+
elapsedMs: number
|
|
10
|
+
decidedAt: string
|
|
11
|
+
}): AdaptiveSearchDecision {
|
|
12
|
+
const { generation, generationCandidates, allCandidates, config, elapsedMs, decidedAt } = input
|
|
13
|
+
const maxTrajectories = config.maxTrajectories ?? config.population + 1
|
|
14
|
+
const maxIterations = config.maxTotalIterations ?? maxTrajectories * config.maxIterations
|
|
15
|
+
const allocatedTrajectories = allCandidates.length
|
|
16
|
+
const allocatedIterations = allocatedTrajectories * config.maxIterations
|
|
17
|
+
const evidence = generationCandidates.map((candidate) => ({
|
|
18
|
+
candidateId: candidate.id,
|
|
19
|
+
status: candidate.status,
|
|
20
|
+
...(candidate.fitness
|
|
21
|
+
? { visiblePassRate: candidate.fitness.metrics.generalization }
|
|
22
|
+
: {}),
|
|
23
|
+
}))
|
|
24
|
+
const distinctResults = new Set(evidence.map((entry) => `${entry.status}:${entry.visiblePassRate ?? "unknown"}`))
|
|
25
|
+
const mixed = evidence.length >= 2 && distinctResults.size > 1
|
|
26
|
+
let reason: AdaptiveSearchDecision["reason"]
|
|
27
|
+
if (allocatedTrajectories >= maxTrajectories) reason = "trajectory-cap"
|
|
28
|
+
else if (allocatedIterations + config.maxIterations > maxIterations) reason = "iteration-cap"
|
|
29
|
+
else if (elapsedMs >= (config.maxRuntimeMs ?? 60 * 60_000)) reason = "runtime-cap"
|
|
30
|
+
else if (generation >= config.generations) reason = "generation-cap"
|
|
31
|
+
else if (evidence.length < 2) reason = "insufficient-results"
|
|
32
|
+
else if (!mixed) reason = "consistent-evidence"
|
|
33
|
+
else reason = "mixed-evidence"
|
|
34
|
+
const remainingTrajectories = Math.max(0, maxTrajectories - allocatedTrajectories)
|
|
35
|
+
const remainingIterations = Math.max(0, maxIterations - allocatedIterations)
|
|
36
|
+
return {
|
|
37
|
+
generation,
|
|
38
|
+
decision: reason === "mixed-evidence" ? "continue" : "stop",
|
|
39
|
+
reason,
|
|
40
|
+
candidateIds: generationCandidates.map((candidate) => candidate.id),
|
|
41
|
+
evidence,
|
|
42
|
+
allocatedTrajectories,
|
|
43
|
+
remainingTrajectories,
|
|
44
|
+
allocatedIterations,
|
|
45
|
+
remainingIterations,
|
|
46
|
+
elapsedMs,
|
|
47
|
+
decidedAt,
|
|
48
|
+
}
|
|
49
|
+
}
|