mcp-castor 2026.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. package/README.md +487 -0
  2. package/bin/castor.js +706 -0
  3. package/index.js +206 -0
  4. package/package.json +97 -0
  5. package/skills/canary-test-staging/SKILL.md +24 -0
  6. package/skills/evo-mutation-rollback/SKILL.md +29 -0
  7. package/skills/hypothesis-generation/SKILL.md +26 -0
  8. package/skills/traceback-condensing/SKILL.md +26 -0
  9. package/src/castor_runner.js +469 -0
  10. package/src/config.js +1204 -0
  11. package/src/env.js +10 -0
  12. package/src/evo_engine.js +214 -0
  13. package/src/harness/core/events.js +75 -0
  14. package/src/harness/core/kernel.js +209 -0
  15. package/src/harness/evo/evaluator.js +156 -0
  16. package/src/harness/evo/evo_operator.js +550 -0
  17. package/src/harness/evo/lineage_dag.js +383 -0
  18. package/src/harness/evo/trace_repair.js +173 -0
  19. package/src/harness/evo/watchdog.js +72 -0
  20. package/src/harness/loop_detector.js +135 -0
  21. package/src/harness/runner.js +1216 -0
  22. package/src/harness/services/ast_service.js +1813 -0
  23. package/src/harness/services/event_logger.js +275 -0
  24. package/src/harness/services/mcp_bridge.js +408 -0
  25. package/src/harness/services/provider_vllm.js +728 -0
  26. package/src/harness/services/sandbox_fs.js +1238 -0
  27. package/src/harness/services/searxng_lifecycle.js +254 -0
  28. package/src/harness/services/shell_executor.js +264 -0
  29. package/src/harness/services/shell_validator.js +506 -0
  30. package/src/harness/services/web_service.js +828 -0
  31. package/src/platform.js +344 -0
  32. package/src/repetition_detector.js +139 -0
  33. package/src/semaphore.js +373 -0
  34. package/src/server_lifecycle.js +781 -0
  35. package/src/skills.js +400 -0
  36. package/src/state_pruner.js +392 -0
  37. package/src/task_registry.js +1357 -0
  38. package/src/telemetry.js +638 -0
  39. package/src/tools.js +997 -0
  40. package/src/wsl_bridge.js +629 -0
  41. package/src/wsl_env.js +171 -0
  42. package/stream_proxy.js +453 -0
package/README.md ADDED
@@ -0,0 +1,487 @@
1
+ # mcp-castor
2
+
3
+ Unified local Castor agent harness & MCP server. Exposes the
4
+ coworker (supporting Qwen3.8-27B default with vLLM + DFlash2 + KVarN, 245K context, or any configured model/provider) to two runtimes:
5
+ **Claude Code CLI** and **Google Antigravity IDE**.
6
+
7
+ - In-process Castor microkernel with reversible plugin mount/unmount
8
+ - Structural AST surgery via `@ast-grep/napi` (in-process) + CLI fallback
9
+ - 137-vector zero-trust containment (123 attack vectors blocked, 14 allow vectors)
10
+ - Closed-loop evolutionary optimization (`.evo/lineage.json`)
11
+ - Zero-turn OS-level wait (`curl` long-poll on `:18021` eliminating polling tax across 1.31B+ tokens)
12
+ - Engine wedge detection + auto-heal
13
+ - Full 66-suite test gate (`npm run test:all`) validated live with zero skips
14
+
15
+ ## Quickstart
16
+
17
+ ### Claude Code CLI
18
+
19
+ ```bash
20
+ claude mcp add --scope user castor node <repo-root>/mcp-castor/index.js
21
+ ```
22
+
23
+ The server registers three tools: `castor_coworker`, `castor_task`, `castor_server`
24
+ (customizable prefix via `CASTOR_TOOL_PREFIX`).
25
+ On first dispatch the serving engine boots automatically (if configured) and the
26
+ stream proxy starts on `:18022`.
27
+
28
+ ### Google Antigravity IDE
29
+
30
+ ```bash
31
+ npx -y mcp-castor init
32
+ ```
33
+
34
+ This writes `~/.castor/config.json` (baseURL, model, max_context,
35
+ tool_prefix) and then runs `castor install`, which writes the tool JSON
36
+ schemas (from `getToolManifest()`) into
37
+ `~/.gemini/antigravity-ide/mcp/castor/` and merges the `mcpServers.castor`
38
+ entry into `~/.gemini/config/mcp_config.json`. From a local checkout use
39
+ `node bin/castor.js init` (or `node bin/castor.js install --antigravity`
40
+ to skip the config prompt). Re-run after any schema change in `src/tools.js`.
41
+
42
+ ## Module Map
43
+
44
+ | File | Purpose |
45
+ |------|---------|
46
+ | `index.js` | MCP server entry: registers 3 tools, lifecycle handlers, `isMain` guard |
47
+ | `src/config.js` | All constants + env-var parsing (ports, timeouts, budgets, profiles) |
48
+ | `src/platform.js` | Platform abstraction: WSL/Windows path translation, shell resolution, spawn-profile builder |
49
+ | `src/wsl_bridge.js` | WSL bridge: path translators, `canonicalizePath`, `killProcessTree` (anchored sweep + verify) |
50
+ | `src/semaphore.js` | Cross-process task slot semaphore (disk-lease, O_EXCL claim, heartbeat, reclaim) |
51
+ | `src/task_registry.js` | Task registry + HTTP status server (`:18021`): long-poll wait, cancel, orphan detection |
52
+ | `src/server_lifecycle.js` | Serving engine lifecycle: boot, wedge detection (stats silence + canary), auto-heal, stream-proxy ensure |
53
+ | `src/tools.js` | MCP tool registration: `castor_coworker`, `castor_task`, `castor_server` |
54
+ | `src/castor_runner.js` | Task dispatch: watchdog, Evo lineage |
55
+ | `src/skills.js` | Skills library: frontmatter parsing, keyword matching, budget-capped injection |
56
+ | `src/evo_engine.js` | Evo lineage engine (git-commit-based, `getLineageContext`, `extractMetric`, `recordCandidate`) |
57
+ | `src/repetition_detector.js` | Stateful SSE repetition detector (tiered char limits, block-level pattern) |
58
+ | `src/harness/runner.js` | Castor runner: agent loop, continuation, empty-stream guard, reasoning ceiling |
59
+ | `src/harness/core/kernel.js` | Castor microkernel: Context, plugin mount/unmount, tool registry, EventBus |
60
+ | `src/harness/core/events.js` | EventBus (typed event emission) |
61
+ | `src/harness/services/sandbox_fs.js` | Sandboxed FS: `read_file`, `write_file`, `edit_file`, `list_dir`, `search_code` |
62
+ | `src/harness/services/shell_executor.js` | Shell executor: `bash` / `exec_command` (POSIX routing, WSL, dry-run) |
63
+ | `src/harness/services/shell_validator.js` | Pure in-memory shell security validator (zero-exec enclave) |
64
+ | `src/harness/services/ast_service.js` | AST service: `ast_search`, `ast_replace`, `ast_replace_batch` (napi + CLI) |
65
+ | `src/harness/services/provider_vllm.js` | OpenAI-compatible SSE streaming client (reasoning accounting, idle watchdog, ceiling) |
66
+ | `src/harness/services/mcp_bridge.js` | MCP extension bridge: spawns remote MCP servers, registers `ext_<server>_<tool>` |
67
+ | `src/harness/services/event_logger.js` | Append-only JSONL session event logger |
68
+ | `src/harness/evo/evo_operator.js` | Evo operator: `evo_propose/evaluate/select/revert_candidate`, `evo_status` |
69
+ | `src/harness/evo/lineage_dag.js` | Lineage DAG (`.evo/lineage.json`): candidate tree, fitness tracking |
70
+ | `src/harness/evo/evaluator.js` | Closed-loop metric evaluator (fitness score, failure digest) |
71
+ | `src/harness/evo/trace_repair.js` | Traceback condenser (≤100-token failure digest) |
72
+ | `src/harness/evo/watchdog.js` | Evo watchdog: stagnation breaker, token-velocity decay |
73
+ | `stream_proxy.js` | Universal stream proxy (`:18022`): UTF-8 reassembly, SSE keep-alive, multimodal guard, repetition breaker |
74
+ | `bin/castor.js` | Castor CLI: `init` (config), `install` (Antigravity schemas + mcp_config merge, Claude Code registration), `config`, `status` |
75
+
76
+ ## Dual-Runtime Quickstart
77
+
78
+ ### Claude Code (stdio MCP)
79
+
80
+ ```bash
81
+ # Register (one-time)
82
+ claude mcp add --scope user castor node <repo-root>/mcp-castor/index.js
83
+
84
+ # In a Claude Code session, dispatch:
85
+ # castor_coworker(prompt="...", cwd="<repo-root>/my-project", session_id="task1")
86
+ #
87
+ # Long tasks yield a taskId + wait_command:
88
+ # curl -s http://127.0.0.1:18021/task/<id>/wait
89
+ #
90
+ # Check status:
91
+ # castor_task(action="status", task_id="<id>")
92
+ #
93
+ # Engine lifecycle:
94
+ # castor_server(action="status") # gauges, wedge state, canary
95
+ # castor_server(action="start") # boot serving engine + stream proxy
96
+ # castor_server(action="stop") # stop (refuses if tasks in flight)
97
+ ```
98
+
99
+ ### Antigravity IDE
100
+
101
+ ```bash
102
+ # Install (one-time, re-run after schema changes)
103
+ npx -y mcp-castor init
104
+ # or from a local checkout:
105
+ node bin/castor.js install --antigravity
106
+
107
+ # Antigravity auto-discovers the MCP server from:
108
+ # ~/.gemini/antigravity-ide/mcp/castor/
109
+ # (mcpServers.castor is merged into ~/.gemini/config/mcp_config.json)
110
+ #
111
+ # Call via:
112
+ # call_mcp_tool(serverName="castor", toolName="qwen_coworker", arguments={...})
113
+ ```
114
+
115
+ ## Environment Variables
116
+
117
+ All variables are read at process start (module-level) unless noted.
118
+
119
+ | Variable | Default | What it does |
120
+ |----------|---------|--------------|
121
+ | `VLLM_PORT` | `18020` | vLLM engine port |
122
+ | `STATUS_PORT` | `18021` | Task status HTTP server port (long-poll wait, cancel) |
123
+ | `STREAM_PROXY_PORT` | `18022` | Stream proxy port (used by `src/config.js` for the provider) |
124
+ | `VLLM_PROXY_PORT` | `18022` | Stream proxy listen port (used by `stream_proxy.js`) |
125
+ | `VLLM_PROXY_HOST` | `127.0.0.1` | Stream proxy bind address (loopback only; WSL2 forwards it to the Windows host) |
126
+ | `QWEN_STATE_DIR` | `~/.castor` (or WSL-mapped Windows home; auto-migrates from legacy `~/.qwen`) | Root for task JSON, slot leases, session logs, wedge counter |
127
+ | `QWEN_MAX_TOKENS` | `49152` | Per-turn output token budget |
128
+ | `QWEN_MAX_REASONING_TOKENS` | `32768` | Per-turn reasoning (thinking) token ceiling; hit → `finish_reason: "length"` |
129
+ | `QWEN_STREAM_IDLE_TIMEOUT_MS` | `1200000` (20 min) | SSE stream idle watchdog (first-byte + inter-chunk) |
130
+ | `QWEN_REASONING_EFFORT` | `medium` | Fallback effort when a dispatch sends none; per-dispatch `reasoning_effort` overrides. Engine accepts exactly {xhigh, medium, low}. `xhigh` is explicit-only. |
131
+ | `QWEN_RACE_MS` | `15000` (15 s) | Client-side race deadline before yielding `taskId` + `wait_command` |
132
+ | `QWEN_MIN_TIMEOUT_MS` | `600000` (10 min) | Floor for task timeout |
133
+ | `QWEN_INACTIVITY_TIMEOUT_MS` | `1800000` (30 min) | Task execution inactivity watchdog |
134
+ | `QWEN_FIRST_TOKEN_TIMEOUT_MS` | `240000` (4 min) | **Reserved, no consumer yet** (retained as near-term knob per honesty-drift decision). Live zero-output protection is `QWEN_STREAM_IDLE_TIMEOUT_MS`, whose first-byte watchdog already covers this case |
135
+ | `QWEN_TASK_RETENTION_MS` | `604800000` (7 days) | Task-telemetry retention window; floored at `DEFAULT_TIMEOUT_MS + 30min` so a live task's JSON is never unlinked mid-run |
136
+ | `QWEN_MAX_CONCURRENT` | `1` | Global task slot count (cross-process, disk-lease); kept 1:1 with the engine's `MAX_SEQS` |
137
+ | `QWEN_MAX_TURNS` | *(null = unbounded)* | Max agent turns per dispatch |
138
+ | `QWEN_MAX_CONTINUATION_TURNS` | `8` | Max re-prompts after `finish_reason: "length"` |
139
+ | `QWEN_EMPTY_STREAM_RETRIES` | `2` | Retries for empty/zero-byte generations before honest failure |
140
+ | `QWEN_WEDGE_SILENCE_S` | `120` | Engine stats silence threshold (seconds) for wedge declaration |
141
+ | `QWEN_AUTO_HEAL` | `true` (disable with `0`) | Auto-reboot wedged engine |
142
+ | `QWEN_LOG_PATH` | `/tmp/mcp_launch_huge.log` | vLLM engine log path (for stats-line wedge detection) |
143
+ | `QWEN_WSL_DISTRO` | `Ubuntu` | WSL distro name |
144
+ | `QWEN_WSL_USER` | *(probed via `whoami`)* | WSL user |
145
+ | `QWEN_WSL_HOME` | `/home/<user>` | WSL home directory |
146
+ | `QWEN_WIN_HOME` | `os.homedir()` | Windows host home |
147
+ | `QWEN_WIN_HOME_WSL` | *(derived from `QWEN_WIN_HOME`)* | Windows home as seen from WSL (`/mnt/c/Users/...`) |
148
+ | `QWEN_POSIX_SHELL` | *(probed: Git Bash, then PATH)* | POSIX shell for the `bash` tool |
149
+ | `QWEN_SHELL_MODE` | *(unset = POSIX)* | Set to `cmd` to force `cmd.exe` instead of bash |
150
+ | `QWEN_SHELL_DRY_RUN` | *(unset)* | Set to `1` to simulate shell commands (no spawn) |
151
+ | `QWEN_FORCE_WSL` | `true` on Windows (disable with `0`) | Force WSL execution for non-WSL cwds |
152
+ | `QWEN_STREAM_PROXY_PATH` | *(derived from `import.meta.url`)* | WSL-side path to `stream_proxy.js` |
153
+
154
+ ## Extensions Guide (P8 Bridge)
155
+
156
+ The `qwen_coworker` tool accepts an optional `extensions` array. Each entry is
157
+ a string (e.g. `"npx -y @upstash/context7-mcp"`) or a structured object
158
+ (`{ command, args, name? }`).
159
+
160
+ **How it works:**
161
+
162
+ 1. **Normalization** — `normalizeExtensionSpec()` in `src/harness/services/mcp_bridge.js`
163
+ tokenizes the string into an argv array (no shell), derives a server name
164
+ from the last non-flag token, and sanitizes it to `[a-z0-9_]`.
165
+ 2. **Spawn** — Each extension is spawned as a piped stdio child via
166
+ `buildSpawnProfile()` (WSL or Windows mode). On Windows, bare package
167
+ runners (`npx`, `uvx`) are resolved to absolute paths via `where.exe`
168
+ first (Node's `spawn` cannot resolve `.cmd` shims without a shell).
169
+ 3. **Handshake** — JSON-RPC over stdio: `initialize` →
170
+ `notifications/initialized` → `tools/list`. Per-step timeout: 60s.
171
+ 4. **Registration** — Each remote tool is registered on the Castor Context as
172
+ `ext_<server>_<tool>` (e.g. `ext_context7_mcp_get_library_docs`).
173
+ Per-server tool cap: 64.
174
+ 5. **Lifecycle** — The bridge is booted **before** the agent loop (so tools
175
+ are visible on the first turn) and disposed in the runner's `finally`
176
+ block (process-tree kill + tool unregistration). A module-level
177
+ `LIVE_BRIDGES` registry ensures `disposeAllBridges()` in `index.js`
178
+ reaps any surviving children on process shutdown.
179
+
180
+ **Never throws:** bad specs or failed handshakes are logged to stderr and
181
+ skipped; they cannot take down the dispatch.
182
+
183
+ ## Skills Guide (P9)
184
+
185
+ Skills are reusable workflow recipes at `skills/<name>/SKILL.md`.
186
+
187
+ **Frontmatter format:**
188
+
189
+ ```markdown
190
+ ---
191
+ name: my-skill
192
+ description: One-line summary
193
+ keywords: [keyword1, keyword2, keyword3]
194
+ ---
195
+ <workflow body in markdown>
196
+ ```
197
+
198
+ **Matching:** `matchSkills({ prompt, cwd })` in `src/skills.js` performs a
199
+ case-insensitive substring match of each keyword against both the prompt text
200
+ and the cwd path string. A repo named `evo` triggers the `evo` skill.
201
+
202
+ **Budgets** (enforced in `matchSkills`):
203
+ - `MAX_SKILLS = 3` — at most 3 skills injected per dispatch
204
+ - `MAX_BODY_CHARS = 2000` — per-skill body truncation
205
+ - `MAX_TOTAL_CHARS = 6000` — total injected budget
206
+
207
+ **Injection:** `injectSkills(prompt, cwd)` appends a delimited block
208
+ (`--- Matching skills (auto-injected from skills/) ---`) to the prompt only
209
+ when matches exist. Never throws; malformed skills are skipped with a stderr
210
+ note.
211
+
212
+ **Shipped skills (4):**
213
+
214
+ | Skill | Keywords |
215
+ |-------|----------|
216
+ | `evo-mutation-rollback` | rollback, revert, mutation, regression, snapshot, candidate |
217
+ | `canary-test-staging` | canary, smoke test, staging suite, test gate, narrow test, targeted test |
218
+ | `hypothesis-generation` | hypothesis, evo, fitness, benchmark, optimization, candidate, lineage |
219
+ | `traceback-condensing` | traceback, stack trace, error digest, trace_repair, assertion failure, pytest, failed test |
220
+
221
+ ## Security Model
222
+
223
+ Five-layer containment, all fail-closed:
224
+
225
+ 1. **Path/symlink containment** (`src/harness/services/sandbox_fs.js`,
226
+ `src/harness/services/ast_service.js`): every path is canonicalized
227
+ through `fs.realpathSync` (resolves NTFS junctions + POSIX symlinks),
228
+ then compared against the canonical root. `PathEscapeError` on `..`
229
+ traversal; `SymlinkEscapeError` when a symlink inside the workspace
230
+ resolves outside.
231
+
232
+ 2. **Null-byte rejection** (`NullByteError`): any `\0` in a path is
233
+ rejected before normalization.
234
+
235
+ 3. **Device-name rejection** (`DeviceNameError`): Windows reserved names
236
+ (`CON`, `PRN`, `AUX`, `NUL`, `COM1-9`, `LPT1-9`) are blocked.
237
+
238
+ 4. **Shell validator & isolation** (`src/harness/services/shell_validator.js`,
239
+ `src/harness/services/shell_executor.js`):
240
+ pure in-memory, zero-execution enclave protecting against 15 evasion classes:
241
+ - **Pattern-level blocks**: `mkfs`, `fdisk`, `parted`, `format X:`,
242
+ `dd of=/dev/...`, fork bomb.
243
+ - **Chain/delimiter splitting**: `;`, `&&`, `||`, `|`, newlines — every
244
+ segment is recursively segmented and analyzed, not just the first.
245
+ - **Wrapper unwrapping**: `sudo`, `doas`, `env`, `nice`, `nohup`,
246
+ `xargs`, `bash -c`, `cmd /c`, `powershell -Command` (bounded depth 4).
247
+ - **Unexpanded-reference rejection**: `$VAR`, `${VAR}`, `$(cmd)`,
248
+ backticks in destructive operands → `CommandSecurityError`.
249
+ - **Tilde expansion**: `~`, `~/`, `~user`, `~root` resolved via platform resolvers.
250
+ - **Homoglyph defense**: NFKC normalization folds fullwidth/compatibility
251
+ forms to ASCII before protected-root comparison (`\uFF37indows` -> `Windows`).
252
+ - **Protected roots**: `/`, `/root`, `/home`, `/home/<user>`, all
253
+ `/mnt/a-z`, `/mnt/c/windows`, `/mnt/c/users`, `/mnt/c/program files`,
254
+ WSL home, Windows home (as `/mnt/c/Users/...`), raw device names.
255
+ - **Flag disambiguation**: Windows `/s /q /f` recognized as flags,
256
+ not path operands. `--name=value` flag values treated as operands.
257
+ - **ANSI color isolation** (P14, vector `b4`): `shell_executor.js` strips
258
+ color-forcing environment variables (`FORCE_COLOR`, `CLICOLOR`, `CLICOLOR_FORCE`)
259
+ and injects `NO_COLOR=1` into child processes, preventing terminal escapes from
260
+ polluting piped output under Claude Code. WSL-bound commands are additionally
261
+ sanitized inside the bash payload (P15), since Windows-side env does not cross
262
+ into WSL without `WSLENV`.
263
+
264
+ 5. **Dead-man fuse** (`assertDeadManFuse`): a final pre-spawn barrier that
265
+ re-checks pattern-level blocks and a synthetic canary token.
266
+
267
+ **`edit_file` guards & Rule 7** (`src/harness/services/sandbox_fs.js`):
268
+ - **Zero-occurrence**: target not found → explicit error, no write.
269
+ - **`AmbiguousTargetError`**: >1 occurrence without `replace_all` → refusal
270
+ with count, no write (prevents code block duplication).
271
+ - **`LineEndingMismatchError` & Universal UNIX LF Invariant (Rule 7)**:
272
+ exact match fails but CRLF-normalized match succeeds → explicit error, no write.
273
+ All files across this workspace strictly use UNIX LF (`\n`). BPE tokenizers
274
+ split `\r\n` into multiple tokens, degrading DFlash2 speculative decoding;
275
+ `LineEndingMismatchError` halts mutations before disk contamination occurs.
276
+
277
+ **Syntax gates + rollback-on-invalid** (`src/harness/services/ast_service.js`):
278
+ every AST rewrite is written to disk, then validated by a language-specific
279
+ checker (TypeScript `transpileModule`, `node --check`, `python -c
280
+ "import ast; ast.parse(...)"`, `gofmt`, `rustfmt`, JSON parse). If the
281
+ checker reports invalid syntax, the file is **rolled back to its pre-rewrite
282
+ content**. When no checker is available, the gate degrades honestly
283
+ (`checked: false`) and logs a stderr note — it never reports an unchecked
284
+ rewrite as valid.
285
+
286
+ ## Troubleshooting
287
+
288
+ ### Engine wedge auto-heal + busy-gate
289
+
290
+ The engine runs `MAX_SEQS=1` (one generation at a time; the 2026-09-12
291
+ two-seat raise was reverted on 2026-09-25). While a task is
292
+ executing, a canary probe would queue behind the active generation and time
293
+ out — measuring queue depth, not health. The **busy-gate** in
294
+ `src/server_lifecycle.js` (`engineWedgeState`) reads `/metrics` gauges first:
295
+ when `running_requests > 0` or `waiting_requests > 0`, the canary is
296
+ **skipped** (`canary.skipped = "engine_busy"`) and wedge is declared from
297
+ stats-line silence alone (`>120s` without a new `Engine 000: ... Running:`
298
+ line in the engine log). When idle, the canary is authoritative.
299
+
300
+ On wedge detection, `healWedgedEngine` auto-reboots the engine (stop +
301
+ start), guarded by a cross-instance lock file (`.engine_heal.lock`, 5-min
302
+ TTL) and a **heal gatekeeper** that refuses to reboot while live work is in
303
+ flight (`hasLiveWork()` in `src/task_registry.js`).
304
+
305
+ ### Orphaned tasks + `qwen_task cancel_all`
306
+
307
+ If the MCP server process dies, its tasks become orphans. On the next
308
+ `readTaskFromDisk`, `isTaskOrphaned` checks the owner PID liveness and
309
+ heartbeat age; orphans are marked `FAILED` with a `WORKER_PROCESS_TERMINATED`
310
+ error. `qwen_task(action="cancel_all")` sweeps session processes via an
311
+ anchored per-session sweep (pgrep → `/proc` cmdline boundary verify → kill),
312
+ cancels all in-memory and disk tasks, and clears slot leases.
313
+
314
+ ### Stream-proxy lifecycle
315
+
316
+ `ensureStreamProxyRunning` in `src/server_lifecycle.js`:
317
+ 1. Probes `/health` (2 quick attempts).
318
+ 2. **Targeted pre-spawn cleanup**: kills the current listener on the port by
319
+ its specific PID (probed from `/health` or `ss -ltnp`), never a broad
320
+ `pkill -f`.
321
+ 3. Spawns the proxy (injected spawner for tests; real spawner uses
322
+ `setsid` in WSL or detached `node` on Linux).
323
+ 4. **15s health window** (75 × 200ms polls) with early-exit success.
324
+
325
+ ### Reasoning loops
326
+
327
+ Two layers of defense:
328
+ - **Stream-proxy repetition breaker** (`src/repetition_detector.js`):
329
+ detects literal character/phrase repetition in SSE deltas (both
330
+ `delta.content` and `delta.reasoning`/`delta.reasoning_content`) and
331
+ truncates the stream with a guard message.
332
+ - **Reasoning token ceiling** (`src/harness/services/provider_vllm.js`):
333
+ when estimated reasoning tokens exceed `QWEN_MAX_REASONING_TOKENS`
334
+ (default 32768), the stream is ended locally with `finish_reason: "length"`
335
+ + `hadReasoning: true`, routing the runner into the reasoning-cutoff
336
+ continuation directive ("stop deliberating, emit edits now").
337
+
338
+ ### Kill certainty
339
+
340
+ `killProcessTree` in `src/wsl_bridge.js`:
341
+ 1. **Anchored session-id sweep**: `pgrep -f` lists candidates, then each
342
+ candidate's `/proc/<pid>/cmdline` is verified to have the session ID at
343
+ an exact boundary (space or end-of-line). A decoy that merely *contains*
344
+ the ID as a substring is never killed.
345
+ 2. **Direct child kill**: `taskkill /T /F` (Windows) or `SIGKILL` to the
346
+ process group (Linux).
347
+ 3. **Post-kill verification**: two liveness probes 500ms apart. If still
348
+ alive, the kill is re-issued (escalation).
349
+ 4. Returns `{ killed, escalations }` — a pid-less target is a no-op success
350
+ (`killed: false, escalations: 0`), not a failure.
351
+
352
+ ## Test Suite
353
+
354
+ `npm test` runs 9 critical suites (offline, zero engine interruption). `npm run test:all` runs all 66 suites (all must exit 0). The table below is a representative listing of the core suites; the full 66-suite list is the authoritative `test:all` invocation in `package.json`.
355
+
356
+ | # | Suite | Command | Type | Purpose |
357
+ |---|-------|---------|------|---------|
358
+ | 1 | `security.test.js` | `npm run test:security` | Offline | 137-vector containment (123 attack vectors blocked, 14 allow vectors; path, symlink, null-byte, device, shell chains/homoglyphs) |
359
+ | 2 | `canary.test.js` | `npm run test:canary` | Offline | Fast canary pilot: AST search/rewrite, syntax gate, traceback condenser, Evo eval |
360
+ | 3 | `ast_engine.test.js` | `npm test` | Offline | napi-vs-CLI equivalence (search + replace, byte-identical) |
361
+ | 4 | `ast_batch.test.js` | `npm run test:batch` | Offline | Batch replace (directory/glob target, dry_run preview) |
362
+ | 5 | `edit_file_guard.test.js` | `npm test` | Offline | edit_file guards: zero-occurrence, AmbiguousTargetError, LineEndingMismatchError |
363
+ | 6 | `evo.test.js` | `npm run test:evo` | Live / Skip | Full Evo system (kernel, sandbox, AST, shell, Evo, live vLLM; honest-skip offline) |
364
+ | 7 | `semaphore.test.js` | `npm run test:semaphore` | Offline | Cross-process task slot semaphore (lease files in private temp dir) |
365
+ | 8 | `runner_continuation.test.js` | `npm run test:continuation` | Offline | Continuation-on-cutoff: length, empty-generation, reasoning-landing |
366
+ | 9 | `wedge_guard.test.js` | `npm run test:wedge` | Offline | Busy-gate + heal backstop (fully offline) |
367
+ | 10 | `platform.test.js` | `npm test` | Offline | Platform abstraction: WSL/Windows, spawn profiles, resolvers |
368
+ | 11 | `posix_routing.test.js` | `npm test` | Offline | POSIX shell routing: Git Bash resolution, array argv, QWEN_POSIX_SHELL |
369
+ | 12 | `path_canonicalization.test.js` | `npm run test:canonical` | Offline | Path canonicalization: NTFS junction (`D:\mnt\d`), symlink containment |
370
+ | 13 | `syntax_gates.test.js` | `npm run test:syntax` | Offline | Universal syntax gates: all languages checked or honestly degraded |
371
+ | 14 | `syntax_integrity.test.js` | `npm run test:syntax_integrity` | Offline | Syntax-integrity scan: `node --check` on every `.js` file in `src/` + `tests/` |
372
+ | 15 | `provider_reasoning.test.js` | `npm run test:reasoning` | Offline | Provider reasoning-token accounting & ceiling (offline SSE) |
373
+ | 16 | `mcp_bridge.test.js` | `npm run test:bridge` | Offline | MCP extension bridge (offline echo fixture, tool registration) |
374
+ | 17 | `skills.test.js` | `npm run test:skills` | Offline | Skills library: frontmatter, matching, budgets, injection |
375
+ | 18 | `reaping.test.js` | `npm run test:reaping` | Offline | Process-reaping: kill certainty, anchored sweep, decoy survival, proxy lifecycle |
376
+ | 19 | `mcp_client.test.js` | `npm test` | Live / Skip | MCP client integration (live engine; honest-skip offline) |
377
+ | 20 | `fifo_queue.test.js` | `npm test` | Live / Skip | FIFO queue semantics (live engine; honest-skip offline) |
378
+ | 21 | `stdio_purity.test.js` | `npm run test:stdio` | Offline | Zero-stdout-write invariant lock (stdio JSON-RPC purity) |
379
+ | 22 | `tool_errors.test.js` | `npm run test:tool_errors` | Offline | MCP-conformant tool-error envelopes (isError, no thrown exceptions) |
380
+ | 23 | `schema_parity.test.js` | `npm run test:schema_parity` | Offline | Schema drift lock: live-served zod schemas vs `castor install` JSON + `getToolManifest()` |
381
+ | 24 | `stream_proxy.test.js` | `npm run test:proxy` | Live / Mock | SSE stream proxy: repetition tiering, UTF-8 reassembly, real proxy regression |
382
+ | 25 | `utf8_proxy.test.js` | `npm run test:proxy` | Live / Mock | Multi-byte UTF-8 split across chunks reassembly verification |
383
+ | 26 | `benchmark.test.js` | `npm run test:benchmark` | Live / Skip | Head-to-head Evo benchmark |
384
+ | 27 | `runner_mapping.test.js` | `npm test` | Offline | status→isError mapping: `completed`/`completed_ceiling` = success, fail-closed on unknown/null |
385
+ | 28 | `lifecycle_locks.test.js` | `npm run test:all` | Offline | Exclusive lock acquire/reject/stale-recovery, PID-ownership release, atomic wedge counter |
386
+ | 29 | `signal_hardening.test.js` | `npm run test:all` | Offline | vLLM headroom clamping, `ContextExhaustedError`, `readTaskFromDisk` transientLock/ENOENT |
387
+ | 30 | `honesty_drift.test.js` | `npm test` | Offline | `listModels` honest error, corrupt-task quarantine, dead-constant culling |
388
+ | 31 | `lineage_integrity.test.js` | `npm test` | Offline | Per-workspace DAG singleton, version gate, honest legacy status mapping |
389
+ | 32 | `status_lifecycle.test.js` | `npm test` | Offline | Status-server keeper re-election, elapsed fix, cancel slot release, honest `stopServer` |
390
+ | 33 | `shell_hardening.test.js` | `npm test` | Offline | Shell-injection hardening (session-id charset, pgrep escape), credential honesty |
391
+
392
+ **Live-engine test gating**: Suites 6, 19, 20, and 26 use `tests/helpers/engine_probe.js` (`isEngineAvailable` / `requireEngineOrSkip`) to probe `/v1/models` with a 3s timeout. When vLLM is running, all 66 suites execute live; when offline, those four print `[SKIP]` and exit 0 (the remaining 62 run offline or against a mock upstream). Under active engine operation, `npm run test:all` runs all 66 suites with **zero skips and zero failures**.
393
+
394
+ ## Production Verification & Telemetry Ledger
395
+
396
+ Across continuous production pair-programming on consumer 24 GB hardware (RTX 3090), Castor tracks all token consumption, cache hits, tool calls, and financial savings in `~/.castor/telemetry/stats.json`.
397
+
398
+ ### Cumulative Lifetime Production Telemetry (1.31B+ Tokens @ $0 Cost)
399
+
400
+ | Metric | Measured Value | Operational Value & Grounding |
401
+ |--------|----------------|-----------------------|
402
+ | **Cumulative Prompt / Ingest Volume** | **1,309,571,077 tokens (1.31 Billion)** | Absorbed full codebase ASTs, git diffs & build telemetry at **$0 local cost** |
403
+ | **Cumulative Completion Output** | **22,045,497 tokens (22.05 Million)** | Generated structural AST edits, code implementations & tests |
404
+ | **Test-Time Deliberation (Reasoning)** | **23,832,473 tokens (23.83 Million)** | Deep chain-of-thought at zero marginal API billing |
405
+ | **Conversational Turns** | **17,616 turns** across **453 sessions** | Continuous multi-project pair-programming |
406
+ | **Sandboxed Tool Operations** | **23,965 total calls** | 11,809 bash, 5,173 read, 2,898 edit, 1,266 search, 1,161 write, 557 web research |
407
+ | **Warm Prefix Cache Hit Rate** | **92.6%** | Sustained prefix reuse yielding 8,000–15,000+ tok/s prefill speeds |
408
+ | **Speculative Drafter (DFlash2)** | **61.9% acceptance (5.33 tok/step)** | Mean draft acceptance yielding instantaneous decode $\approx$ 61.1 tok/s |
409
+ | **Zero-Turn OS Wait Savings** | **~1.3B tokens eliminated** | Zero-turn HTTP long-poll (`:18021`) eradicated supervisor polling loops |
410
+ | **Net Financial Savings vs Claude Sonnet 5** | **$2,839.60 USD saved** | At $2.00/M prompt, $10.00/M completion ($0 local execution cost) |
411
+ | **Net Financial Savings vs Frontier Tier** | **$14,198.00+ USD saved** | vs Opus 5.5 / GPT-6 Astra ($10.00/M prompt, $50.00/M completion) |
412
+
413
+ ### Milestone Marathon Telemetry (11-Hour & 13-Hour Overhauls)
414
+
415
+ The architecture has been stress-tested across extended multi-agent production marathons between Gemini 3.8 Flash (Meta-Supervisor in Antigravity), Claude Code (Lead Architect), and local Qwen3.8-27B (Execution Coworker).
416
+
417
+ #### Empirical Engine & Hardware Telemetry (Session Snapshot)
418
+
419
+ | Metric | Measured Value | Operational Rationale |
420
+ |--------|----------------|-----------------------|
421
+ | **vLLM Prefill / Prompt Tokens** | **56,312,194 tokens** | Ingested locally on RTX 3090 at $0 token cost |
422
+ | **vLLM Generation Tokens** | **1,749,192 tokens** | Codebase exploration, test runs, structural AST surgery |
423
+ | **Empirical Prefill Throughput** | **9,454.3 tok/s** | Average across 56.3M prompt tokens (prefix-cache hit rates reaching 8,000–9,500+ tok/s) |
424
+ | **Empirical Generation (Decode) Speed** | **58.2 tok/s** | Sustained pure decode throughput (mean TPOT 15.42 ms $\to$ **64.9 tok/s** instantaneous) |
425
+ | **Effective End-to-End Turn Speed** | **48.5 tok/s** | Round-trip throughput across conversation turns including prefill & tool handling |
426
+ | **Speculative Accepted Tokens** | **1,379,412 tokens** | **78.86% acceptance rate** via DFlash2 1.92B drafter |
427
+ | **Active Qwen Sessions** | **134 sessions** | Micro-session roll cadence preventing KV cache decay |
428
+ | **Logged Microkernel Events** | **6,140+ events** | Append-only session telemetry (`~/.qwen/sessions/`) |
429
+ | **Total Tool Executions** | **1,939+ calls** | Autonomous execution across Windows and WSL environments |
430
+ | **VRAM Footprint** | **24,136 MiB / 24,576 MiB** | Universal 245K context + KVarN k4v2 KV cache |
431
+ | **GPU Operating Temp** | **31°C - 58°C** | Steady thermal curve under 250W power limit |
432
+ | **Zero-Turn Wait Savings** | **~590M tokens** | Zero-turn HTTP long-poll (`:18021`) eliminated polling tax |
433
+
434
+ ### 13-Hour Autonomous Production Marathon Telemetry (v5.2.0 End-to-End Overhaul — Sept 7, 2026)
435
+
436
+ In an unbroken 13.1-hour autonomous pairing session driving the full 6-phase frontend UI overhaul of an enterprise web application across Claude Code (GLM-5.3 / GLM-5.3-Flash) and local Qwen3.8-27B (Castor Coworker on RTX 3090), the stack delivered the following production metrics:
437
+
438
+ | Production Telemetry Dimension | Empirical Measurement | Operational Value |
439
+ |---|---|---|
440
+ | **Cumulative Prefill Volume** | **73,317,306 tokens** (~73.3M) | Absorbed large multi-file ASTs & git diffs locally at **$0 token cost** |
441
+ | **Prefix Cache Hit Rate** | **93.80% (68,842,624 tokens)** | Sustained warm prefix cache throughput (~8,000–9,500 tok/s) |
442
+ | **Cumulative Generation Volume** | **1,490,578 tokens** (~1.49M) | Full test-time reasoning and code generation delivered at $0 |
443
+ | **DFlash2 Speculative Decoding** | **53.45% draft acceptance** | 1,176,978 accepted / 2,202,011 drafted across 7 draft positions |
444
+ | **Total Engine Requests** | **883 requests** | 882 `stop`, 1 `length`, **0 error, 0 abort, 0 repetition** |
445
+ | **Orchestrator Token Volume** | **53,081,643 tokens** (~53.1M) | 11.08M input, 341.8K output, 41.66M cache read |
446
+ | **Orchestrator Tool Calls** | **245 calls** | 38 Qwen coworker dispatches, 38 zero-turn curl waits, 21 PowerShell |
447
+ | **MCP Process Stability (PID 16912)** | **80.59 MB RSS, 0 crashes** | Zero memory growth or socket leaks over 13 hours continuous uptime |
448
+ | **Zero-Turn OS Wait Savings** | **~650M tokens saved** | 38 background blocking curl tasks eliminated supervisor polling tax |
449
+ | **Shipped Production Deliverables** | **6 UI Overhaul Phases** | Commits `d2383c6` $\to$ `9869945`, paying down 10 lint errors (76 baseline) |
450
+
451
+ #### Empirical Operational Friction Analysis & Resolution (4+ Incidents Audited)
452
+ Detailed audit of the transcripts reveals **6 critical operational friction modes** encountered across the marathon:
453
+ 1. **`engine_empty_response` Stream Cutoff** (`06:19 UTC`): Phase 0 review final response cut off after 32 tool calls; recovered by resuming the warm session with a compact verdict-only directive. Fixed in `runner.js` via honest retry classification.
454
+ 2. **`curl (56) Connection reset by peer`** (`06:53 UTC`): Wait endpoint dropped connection under concurrent SSE load; recovered by verifying task liveness and adding `--retry-all-errors`. Hardened in `tools.js` and `task_registry.js`.
455
+ 3. **Universal 245K Context Ceiling Overflow** (`09:46 UTC`): Multi-turn accumulation of 900+ LOC files filled the 245K context; recovered by rolling session ID to `ui_ovh_p2_b`. Codified in micro-session roll protocol.
456
+ 4. **vLLM Stream Idle Watchdog (900s) & 4 Stalled Intervals** (`17:11 UTC`): Monolithic review prompt reading 1,000+ LOC and multiple diffs caused extended prefill/deliberation that tripped the 15-min idle watchdog after Claude waited through 3 consecutive 10-minute task timeouts (30 min total); recovered by compacting prompt to targeted greps which passed in 197.9s.
457
+ 5. **vLLM JSON Serialization Glitch / Malformed Wake Payload** (`13:15 UTC`): `Unterminated string` masked by stream proxy with HTTP 200 and exit code 0 (`isError=false`); resolved by enforcing Rule 8 (Fail-Fast, zero error masking).
458
+ 6. **Report Generation Stream Cutoff** (`14:46 UTC`): Output stream truncated mid-sentence due to output token ceiling exhaustion; resolved by decoupling reasoning tokens via `QWEN_MAX_REASONING_TOKENS=32768`.
459
+
460
+ *Complete raw logs, Prometheus dumps, GPU telemetry, and parsed metrics are preserved in [`benchmarks/sessions/session_20260907/`](benchmarks/sessions/session_20260907/).*
461
+
462
+ ### The 14 Engineering Passes (P1–P14 Complete Implementation Map)
463
+
464
+ | Pass | Commit | Scope & Subsystem | Core Resolution & Verification |
465
+ |------|--------|-------------------|--------------------------------|
466
+ | **P1** | `8a05ccf` | Manifest & Registration Hygiene | Manifest-driven server identity, engine pin, README env reference |
467
+ | **P2** | `d241ff3`<br>`27727a7`<br>`bbc17c8`<br>`bb0f6d8` | Engine Liveness & Stream Resilience | Reasoning token accounting (`QWEN_MAX_REASONING_TOKENS=32768`), empty-stream retries, busy-gate wedge detection (`running_requests > 0`), cross-instance heal lock |
468
+ | **P3** | `fc55827` | Platform Abstraction | Single platform resolver (`src/platform.js`) + spawn-profile builder for Windows DrvFs and WSL2 |
469
+ | **P4** | `dcbfbe7`<br>`fb89fc3`<br>`3248c1f`<br>`5208f0c`<br>`86c1fbf`<br>`dc06624` | Shell Safety, POSIX Routing & Path Canonicalization | Path-aware protected roots, 15 shell evasion vectors closed (123 attack vectors blocked in `security.test.js`), `LineEndingMismatchError`, POSIX bash routing via Git Bash, universal LF `.gitattributes`, junction canonicalization via `realpathSync` |
470
+ | **P5** | `954443e` | In-Process AST Surgery | `@ast-grep/napi` native module integration with CLI fallback, byte-identical equivalence verified in `ast_engine.test.js` |
471
+ | **P6** | `a9d5174` | Universal Syntax Gates | Pre-commit syntax validation for JS, TS, Python, Go, Rust, JSON with rollback-on-invalid-syntax |
472
+ | **P7** | `cb0559f`<br>`5a68cdf` | AST Surface & Stream Hardening | `ast_replace_batch` tool, dry-run safety preview, repetition breaker extended to `delta.reasoning` (closing reasoning loop death class) |
473
+ | **P8** | `d94d4d8` | Generic MCP Extension Bridge | Dynamic stdio MCP extension spawning, argument tokenization, `ext_<server>_<tool>` registration, verified live with `@upstash/context7-mcp` |
474
+ | **P9** | `3e52bce` | Packaged Skills Library | Reusable workflow recipes (`skills/<name>/SKILL.md`), keyword matching against prompt and cwd, budget capping (3 skills / 2000 chars / 6000 total) |
475
+ | **P10** | `88efb97`<br>`b3c151e` | Process-Reaping Hardening | Anchored session-id sweep (`/proc/<pid>/cmdline`), decoy survival, double liveness probe with escalation, 15s stream proxy health window, test state dir isolation |
476
+ | **P11** | `8bde746` | Offline Test Resilience & Stdio Purity | Syntax integrity scan (`node --check` across all files), honest engine-down skips, stdio zero-stdout-write lock frame verification |
477
+ | **P12** | `9df8a73` | Dual-Runtime Parity Audit | MCP-conformant tool-error envelopes (`isError: true`), schema drift lock test (`tests/schema_parity.test.js`) asserting Antigravity JSON vs live zod schemas |
478
+ | **P13** | `6f217dd` | Architecture Documentation | Comprehensive production README rewrite and `docs/DESIGN.md` incident-wisdom distillation |
479
+ | **P14** | `a1dabdc` | Final E2E Verification & Release | Stripping color-forcing env vars (`FORCE_COLOR`, `CLICOLOR`) and setting `NO_COLOR=1` in `shell_executor.js` (vector `b4`), tag `v5.1.0`, full 26/26 test suites green |
480
+
481
+ ## Version
482
+
483
+ **2026.3.0** (tracked in `package.json` and git tag `v2026.3.0` — the only current-version literal; the MCP server serves it from there).
484
+ - **2026.1.0**: Established the Universal 245K context baseline (vLLM + DFlash2 + KVarN), the 14 engineering passes (P1–P14), 137-vector zero-trust containment, and zero-turn reactive OS wait (`:18021/task/<id>/wait`).
485
+ - **2026.2.0**: Established native multi-provider live web & framework documentation research (`web_search` with Brave, Tavily, Context7 docs, SearXNG, DuckDuckGo + `web_fetch`), zero-progress-loss Cooperative Landing at turn ceilings, Automatic Prefix Caching (APC) stability (~8,000–9,500 tok/s), and Socratic collaborative pair-programming where local coworker pairs with the Lead Architect as an autonomous Staff Software Engineer peer. Retires the legacy 5.x versioning sequence.
486
+ - **2026.3.0**: Castor major release — Universal model- and provider-agnosticism across all OpenAI-compatible endpoints (Ollama, LM Studio, vLLM, SGLang, LiteLLM), namespace-protected stdio tools (`castor_coworker`, `castor_task`, `castor_server`), multimodal vision authority, unified CLI (`castor`), and `mcp-castor` packaging.
487
+