model-orchestrator 0.1.16 → 0.1.18

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (34) hide show
  1. package/CHANGELOG.md +29 -1
  2. package/README.md +14 -10
  3. package/bin/cli-run.mjs +123 -4
  4. package/bin/cli.js +10 -1
  5. package/docs/README.md +1 -1
  6. package/docs/audit-brief.md +13 -0
  7. package/docs/catalog.md +1 -1
  8. package/docs/part-1-beginner.md +4 -4
  9. package/docs/part-2-intermediate.md +4 -4
  10. package/llms.txt +1 -1
  11. package/package.json +1 -1
  12. package/src/catalog.js +1 -1
  13. package/src/install.js +41 -14
  14. package/templates/advanced/vm/jobs/weekly-audit.sh +11 -1
  15. package/templates/agents/agy/README.md +1 -1
  16. package/templates/agents/agy/builder.md +1 -1
  17. package/templates/agents/agy/finding-verifier.md +1 -1
  18. package/templates/agents/claude-code/finding-verifier.md +1 -1
  19. package/templates/agents/snippets/chat.md +2 -2
  20. package/templates/agents/snippets/claude-code.md +1 -1
  21. package/templates/agents/snippets/generic.md +1 -1
  22. package/templates/agents/snippets/route-gate.mjs +1 -1
  23. package/templates/beginner/ORCHESTRATOR.md +2 -2
  24. package/templates/common/README.md +1 -1
  25. package/templates/common/protocols/build-protocol.md +10 -10
  26. package/templates/common/protocols/deep-research.md +2 -2
  27. package/templates/common/protocols/gap-analysis.md +1 -1
  28. package/templates/common/protocols/propagate.md +2 -2
  29. package/templates/intermediate/CLI-RUN.md +3 -3
  30. package/templates/intermediate/DELEGATION_MATRIX.md +1 -1
  31. package/templates/intermediate/ROUTING.md +4 -4
  32. package/templates/intermediate/TIERS.md +12 -8
  33. package/templates/tools/codecalc/CODECALC.md +1 -1
  34. package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +3 -3
package/CHANGELOG.md CHANGED
@@ -4,6 +4,32 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.18] - 2026-09-11
8
+
9
+ ### Fixed
10
+
11
+ - **The level 3 box's weekly audit could record a hung `--version` probe as a version string instead of "UNVERIFIED: timed out".** `weekly-audit.sh`'s `killtree` killed a hung process's children before the process itself, so in that gap the probe's own shell could print and exit 0, and `bounded()` reported success. Seen once on macOS CI ("codex never" in the report); 0 of 20 local runs reproduced it, so it is rare but real on the box. `killtree` now freezes each process (SIGSTOP) before walking its children, and the watchdog leaves a marker when it fires so `bounded()` returns 124 whatever order the kills land in. A new test forces the bad ordering and checks the timeout still reads as one.
12
+ - **On Windows, `bin/cli-run.mjs` could not run any lane at all: `spawn()` threw `EINVAL` for every `.cmd` binary, which is how npm installs every agent CLI there.** Since Node's fix for CVE-2024-27980 (18.20.2, 20.12.2, and every 22.x), spawning a `.bat`/`.cmd` target without `shell: true` throws instead of silently running it through an unsafely-escaped `cmd.exe`. This was a real, shipped defect, not a test gap: a Windows user following this README could not have run a single lane before this release. `bin/cli-run.mjs` now resolves the `.cmd` shim to the Node script npm's own `cmd-shim` tool wrote underneath it and spawns Node directly on that script (`resolveCmdShim`, `windowsSpawnPlan`), so a prompt (untrusted text this tool does not control) never passes through a shell at all in the common case. A lane whose `.cmd`/`.bat` cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, never run through `cmd.exe`: a batch file re-reads its arguments through `%*` after `cmd.exe` has parsed them once, and no escaping fully contains a prompt through both passes. The escaped `cmd.exe` path (the algorithm documented at [qntm.org/cmd](https://qntm.org/cmd) and used by `cross-spawn`, with `windowsVerbatimArguments`) survives only as an explicit opt-in for the installer's own `npm install -g <pinned spec>`, whose arguments never include user text. Verified against the real, byte-for-byte output of `cmd-shim@9.0.2` (the package npm itself uses), not a guessed shape; the escaping is pinned to exact expected strings for `&`, `|`, `^`, `%`, `"`, a trailing backslash and a literal newline in `test/judges.test.js`.
13
+ - **A reconfigure's terminal report (`runtime upgraded:`, `runtime CONFLICT, kept:`, `documents kept:`, and the rest) printed a Windows install's file paths with backslashes**, while every other path this tool prints in generated text uses forward slashes; `src/install.js`'s manifest key was already posix-normalized, but the label built alongside it for the human-readable report was not. `writeFiles()` now builds both from the same posix-normalized value.
14
+ - **A `--dir` outside the project root, or the `vm/` level-3 templates' `INSTALL_DIR`/`INSTALL_DIR_SH`/`INSTALL_DIR_SYSTEMD`, ran through this host's own `path.resolve()`, which reads a leading `/` as drive-relative on win32**, wrong for both: the `vm/` templates describe a REMOTE Linux box (`weekly-audit.sh` is bash, `weekly-audit.service` is a systemd unit, neither of which can run anywhere but Linux), and the outside-project case is documentation prose, not a local filesystem path. An absolute `--dir` given as a bare POSIX path now renders unchanged on every host for both; a real local Windows path (one naming a drive) is untouched, since that case never took this branch.
15
+ - **`killTree`'s win32 branch spawned a bare `taskkill`, which depends on PATH containing `System32`; when it does not, the spawn's `ENOENT` reaches this tool as an unheard `error` event on the returned process and crashes the whole run over what should be a best-effort cleanup step.** Found on `windows-latest` CI the first time a lane actually ran end to end there, once the `EINVAL` fix above stopped hiding it. `taskkillPath()` now resolves the executable under `%SystemRoot%` (falling back through `%windir%` to a fixed path), independent of PATH, and `killTree` attaches an error listener so a spawn failure can never crash the wrapper.
16
+ - **`bin/cli.js`'s own opt-in `npm install -g <ai>` prompt had the identical `EINVAL`-shaped defect as the lane spawn above, in a different file**, since it called `spawnSync('npm', ...)` directly with no shell. It now resolves `npm` with `which()` and runs it through the same `windowsSpawnPlan()` `bin/cli-run.mjs` exports, instead of a second, duplicated fix.
17
+
18
+ ### Changed
19
+
20
+ - **The three Windows test-skip groups tracked in [#30](https://github.com/aunysillyme/model-orchestrator/issues/30) are gone, replaced by two narrower, individually-justified skips found by actually running the unskipped suite on `windows-latest` CI, plus the one pre-existing skip this pass never touched (`statSync().mode`'s executable bit; NTFS has nothing equivalent).** `test/cli.test.js`'s fake lane binaries now install as an npm-style `.cmd` shim (verified against the real `cmd-shim@9.0.2` output) pointing at a small Node script, the same shape a real vendor CLI's install takes through `windowsSpawnPlan()` above, instead of a bespoke `sh.exe` bridge that never exercised cli-run.mjs's own spawn path at all; every test in that group but one, and the `#12` upgrade-path pair (fixed by the report-label change above), now runs unconditionally. `test/install.test.js`'s 7 formerly-skipped tests split on what their path assertion actually describes: the ones naming a `vm/` remote-box path stay literal POSIX (now true on every host, per the `INSTALL_DIR*` fix above); the ones naming a LOCAL path (where this run wrote files on this host) now build their expectation with `resolve()`/`join()` instead of a hardcoded POSIX literal, so they assert the same real, platform-native value the product renders rather than a string that only happened to match on POSIX. Two of those seven also had their own, separate bug once actually run on Windows: their fake `curl`/`jq`/`node`/`codex` binaries were placed on a hand-built PATH containing the literal strings `/usr/bin` and `/bin`, which name nothing on that OS; they now prepend the stub directory to the REAL `process.env.PATH` (this job already runs under Git Bash, so that PATH already carries what `bash` itself needs) instead of replacing it with a POSIX-only guess.
21
+ - **New skip: a lane dying mid-run from a real POSIX signal genuinely cannot be reproduced on win32.** cli-run.mjs's `r.signal || r.status === null` branch exists for a real lane crashing or being sent a signal, but a real Windows lane is a plain `node <script>` process (via the resolved cmd-shim), so it cannot die "by signal" any more than the product it is testing can; Windows has no OS-level POSIX signals. The only way this file's fixture can even simulate a signal death is a nested `sh -c "...; kill -TERM $$"` (writeShellStub's win32 branch has to bridge through `sh` for the shell body to run at all), which puts an extra node process between cli-run.mjs and the dying shell; measured on `windows-latest` CI, MSYS bash's own self-kill status leaks through as a plain nonzero exit code (3840), which this tool already handles correctly, just under a different, honest verdict (`exit_nonzero`, not `killed`).
22
+ - **`#13` (SIGTERM/SIGINT to the wrapper) is Windows-aware now, not skipped, for both signals: Windows has no OS-level signals at all**, proven on `windows-latest` CI (the wrapper died as `{code: null, signal: sig}` for SIGTERM AND SIGINT alike; a hypothesis that SIGINT gets a real, catchable console-control event there was tried first and measured false in this exact scenario, not assumed). The test now expects an unhandled termination for both signals on win32; the graceful exit-143/130-and-kill-the-lane-first behavior stays a POSIX guarantee, asserted as before on every other OS.
23
+ - **New skip: `#10`'s watchdog-kill test, and the new `#10b` `bounded()` timeout check, for the same reason.** `weekly-audit.sh`'s `bounded()`/`killtree()` rely on `pgrep -P` and killing a backgrounded subshell's process tree, real bash job control this script only ever runs under on the box it targets (a systemd-scheduled job on Ubuntu, never something a Windows user runs locally). Actually executing that watchdog against a genuinely hanging stub under `windows-latest` CI's Git Bash, rather than just rendering and syntax-checking the script (which the rest of this test group does, and which passes), hung past a 20s outer timeout: MSYS's job-control emulation does not reliably propagate a `kill -KILL` to the underlying Windows process tree of a backgrounded `( subshell ) &`, a known class of MSYS/Cygwin limitation, not a defect in the generated script.
24
+ - One test-report-label assertion in `test/install.test.js` and one in `test/cli.test.js` still hardcoded `path.join()`'s native separator for what is now posix-normalized generated text (the report-label fix above); both now match the posix form.
25
+ - `docs/audit-brief.md` gained a section on the Windows spawn path: what runs, why no shell in the common case, and why a lane is refused rather than run through `cmd.exe`, and how the installer's one opt-in `cmd.exe` call escapes its arguments.
26
+
27
+ ## [0.1.17] - 2026-09-11
28
+
29
+ ### Changed
30
+
31
+ - **User-facing text now uses plain language instead of security-audit jargon.** Words like "risk", "attack lane", "adversarial", "blast radius", "fail closed" and "threat model" read as alarming to someone deciding whether to try the tool, so they scared off exactly the readers this project needs. No rule any of them described changed, only the wording: "risk" is now "stakes" everywhere it names a routing input (with a one-line definition added to `README.md` and `TIERS.md`), "attack lane" / "Stage 5 Attack" / "attack pass" are now "challenge lane" / "Stage 5 Challenge" / "challenge pass", "adversarial" (auditor, read, critique, turn, pass) is now "second-opinion", "blast radius" is now "everything it touches", "fail(s) closed" is now "refuses by default", and "threat model" is now "security notes" in the files that link to it. `test/prose.test.js` gained a permanent check (`no alarming security wording in user-facing text`) over the purely-prose, user-facing surface (`docs/`, `templates/`, `README.md`, `llms.txt`, `CONTRIBUTING.md`, the PR template) so the old wording cannot silently creep back in. `docs/audit-brief.md`, `SECURITY.md`, `CODE_OF_CONDUCT.md` and code identifiers/comments (for example the `ATTACK_LANE` render var) are unchanged, since these words are expected or load-bearing there.
32
+
7
33
  ## [0.1.16] - 2026-09-11
8
34
 
9
35
  ### Added
@@ -256,7 +282,9 @@ First release.
256
282
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
257
283
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
258
284
 
259
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.16...HEAD
285
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.18...HEAD
286
+ [0.1.18]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.17...v0.1.18
287
+ [0.1.17]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.16...v0.1.17
260
288
  [0.1.16]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.15...v0.1.16
261
289
  [0.1.15]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...v0.1.15
262
290
  [0.1.14]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...v0.1.14
package/README.md CHANGED
@@ -44,7 +44,7 @@ Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/pa
44
44
  | Id | What | Level |
45
45
  |---|---|---|
46
46
  | `claude-code` | Claude Code CLI, the default orchestrator | 1+ |
47
- | `codex` | Codex CLI on a ChatGPT plan: second coder, adversarial auditor | 1+ |
47
+ | `codex` | Codex CLI on a ChatGPT plan: second coder and second-opinion reviewer (a different model family reading your diff) | 1+ |
48
48
  | `agy` | Antigravity CLI on a Google AI plan: research sweeps, concurrent fan-out | 1+ |
49
49
  | `grok` | Grok CLI on X Premium: live X and web reads at $0 | 1+ |
50
50
  | `hermes` | Hermes Agent: the free tier | 2+ |
@@ -173,7 +173,7 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
173
173
  6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, only the brief carries what this task needs. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
174
174
  7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
175
175
 
176
- ## Routing by role, complexity and risk
176
+ ## Routing by role, complexity and stakes
177
177
 
178
178
  Role picks the agent. Two more inputs move the choice, and they move it in
179
179
  different directions, so `TIERS.md` states them separately rather than folding
@@ -182,10 +182,14 @@ them into the role:
182
182
  - **Complexity moves the effort.** A worker executing a finished plan needs less
183
183
  reasoning than the reviewer judging its output. When the plan is airtight the
184
184
  spec is carrying the thinking.
185
- - **Risk moves the tier and the reader.** Security, privacy, data loss and
186
- irreversible changes buy the attack lane, a named check, a rollback path or a
187
- human yes. A one-line change to an auth check is simple and high-risk at the
188
- same time, and it is the risk that decides.
185
+ - **Stakes move the tier and the reader.** Security, privacy, data loss and
186
+ irreversible changes buy the challenge lane, a named check, a rollback path or
187
+ a human yes. A one-line change to an auth check is simple and high-stakes at
188
+ the same time, and it is the stakes that decide.
189
+
190
+ Stakes means what a mistake would cost: a security hole, leaked personal data,
191
+ lost data, or something you can't undo. Most tasks are low-stakes and route
192
+ normally.
189
193
 
190
194
  The top of the ladder is bought with evidence: a reproduced failure, an
191
195
  unresolved checkpoint, an irreversible change. A task that merely feels hard is
@@ -221,7 +225,7 @@ The report prints turns, **route-marker coverage** (the percentage of turns whos
221
225
  A lane with no `--model`, no `--effort` and no `defaults` entry in
222
226
  `bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
223
227
  CLI configured months ago at a low reasoning effort keeps auditing at that
224
- effort while your routing docs describe an adversarial pass.
228
+ effort while your routing docs describe a second-opinion pass.
225
229
 
226
230
  ```bash
227
231
  node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
@@ -244,7 +248,7 @@ Install for the tools you have, then let the generated `ROUTING.md` decide the t
244
248
 
245
249
  ### How do I route tasks to cheaper models?
246
250
 
247
- The rules route by role, complexity and risk (see [Routing by role, complexity and risk](#routing-by-role-complexity-and-risk)). Role picks the agent, complexity moves the effort, risk moves the tier. A task a cheap tier finishes correctly never gets a frontier token.
251
+ The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly never gets a frontier token.
248
252
 
249
253
  ### Is this an LLM router or an AI gateway?
250
254
 
@@ -256,7 +260,7 @@ Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run`
256
260
 
257
261
  ## Requirements
258
262
 
259
- Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22). Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested there; a named set of tests is skipped on Windows, each with its reason in the test file, mainly running a lane end to end through `cli-run`, so treat lane execution on Windows as unproven until someone reports otherwise.
263
+ Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22), including lane execution end to end through `cli-run` against a fake CLI installed the same way npm installs a real one (a `.cmd` shim). `cli-run` never runs a lane through `cmd.exe` when it can avoid it: it resolves the shim to the Node script underneath and spawns Node directly, so a prompt reaching a real lane never passes through a Windows shell. A `.cmd` or `.bat` lane that cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, rather than run through `cmd.exe`: a batch file re-reads its arguments after `cmd.exe` has parsed them once, and no escaping fully contains a prompt through both passes. Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested on Windows too, including SIGTERM/SIGINT to the wrapper (Windows has no OS-level signals: both terminate it unconditionally, verified there rather than treated the same as POSIX). A few narrow skips remain on Windows, each for a POSIX behavior the OS or the CI shell genuinely does not have: `statSync().mode`'s executable bit (NTFS has none), a lane dying mid-run from a real POSIX signal (a real Windows lane cannot die "by signal"), and running `weekly-audit.sh`'s watchdog functions for real under Git Bash's job control, both the end-to-end run and the `bounded()` timeout check (the script itself only ever runs on the Ubuntu box it targets).
260
264
 
261
265
  **Privacy.** The installer sends no telemetry and makes no network call of its own once it is running. Two things around that are worth being exact about:
262
266
 
@@ -272,7 +276,7 @@ Add an AI to `src/catalog.js` and every prompt, table, config and doc picks it u
272
276
  ## Credits
273
277
 
274
278
  - [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for
275
- separating role, complexity and risk instead of compressing them into one
279
+ separating role, complexity and stakes instead of compressing them into one
276
280
  scale, for recording the model and effort a lane was actually asked for, and
277
281
  for verifying findings before they trigger repairs. All three shipped in
278
282
  0.1.14.
package/bin/cli-run.mjs CHANGED
@@ -322,12 +322,129 @@ export function unfence(text) {
322
322
  return m ? m[1].trim() : t;
323
323
  }
324
324
 
325
+ // --- Windows: spawning a lane without a shell ------------------------------
326
+ // Node's fix for CVE-2024-27980 makes spawn() throw EINVAL for a .bat/.cmd
327
+ // target unless shell:true is set: launching a batch file always goes
328
+ // through cmd.exe, and cmd.exe reads metacharacters (& | ^ < > ( ) % " and
329
+ // space) directly off the command line before the target program's own argv
330
+ // is parsed, even inside quotes. A lane's argv[1] here is a user PROMPT, text
331
+ // this wrapper does not control the contents of, so that is a real injection
332
+ // surface, not a theoretical one.
333
+ //
334
+ // npm installs every CLI on Windows as a "cmd-shim": a short .cmd launcher
335
+ // that hands off to node with a script path (see npm's own `cmd-shim`
336
+ // package). Reading that path out and spawning node directly sidesteps
337
+ // cmd.exe, and the injection surface it carries, entirely: this is the
338
+ // preferred path, used whenever the shim matches the shape cmd-shim writes.
339
+ //
340
+ // A LANE whose .cmd/.bat does not match (hand-written, or an older cmd-shim
341
+ // layout) is refused, not run through cmd.exe: a batch file re-reads its
342
+ // arguments through %* after cmd.exe has already parsed them once, which is
343
+ // the case CVE-2024-27980 is about, and no escaping fully contains user text
344
+ // through both passes. Removing that path beats guarding it.
345
+ //
346
+ // The cmd.exe path survives only for a caller that opts in with
347
+ // { allowCmdFallback: true } and passes arguments it fully controls (the
348
+ // installer's own `npm install -g <pinned spec>`, whose npm.cmd is not a
349
+ // cmd-shim). It uses the caret-escaping algorithm documented at
350
+ // https://qntm.org/cmd and used by `cross-spawn`: quote each argument for
351
+ // CommandLineToArgvW, THEN caret-escape cmd.exe's own metacharacters, THEN
352
+ // pass the whole line with windowsVerbatimArguments so Node does not
353
+ // re-quote it a second, conflicting way.
354
+ const NPM_CMD_SHIM = /"%_prog%"\s+"([^"]+)"\s*%\*/;
355
+ export function resolveCmdShim(cmdPath) {
356
+ let text;
357
+ try {
358
+ text = readFileSync(cmdPath, 'utf8');
359
+ } catch {
360
+ return null;
361
+ }
362
+ const m = NPM_CMD_SHIM.exec(text);
363
+ if (!m) return null;
364
+ const dp0 = /^%~?dp0%?[\\/]?/i;
365
+ if (!dp0.test(m[1])) return null; // only the %dp0%-relative shape cmd-shim writes
366
+ const rel = m[1].replace(dp0, '').replace(/\\/g, '/');
367
+ let script;
368
+ try {
369
+ script = resolve(dirname(cmdPath), rel);
370
+ if (!statSync(script).isFile()) return null;
371
+ } catch {
372
+ return null;
373
+ }
374
+ // Only ever hand off to node for a real JS entry point; anything else (a
375
+ // shim generated for a non-node binary, or a hand-edited file) falls
376
+ // through to the cmd.exe fallback instead of being executed as a script.
377
+ return /\.(m?js|cjs)$/i.test(script) ? script : null;
378
+ }
379
+
380
+ function escapeCmdArg(arg) {
381
+ let s = String(arg);
382
+ // A run of backslashes immediately before a quote (or at the very end of
383
+ // the argument) must be doubled, or CommandLineToArgvW on the receiving
384
+ // end eats one; this is the standard Windows argv-quoting rule, not a
385
+ // cmd.exe-specific one.
386
+ s = s.replace(/(\\*)"/g, '$1$1\\"');
387
+ s = s.replace(/(\\*)$/, '$1$1');
388
+ s = `"${s}"`;
389
+ // cmd.exe reads these characters off the raw command line and acts on
390
+ // them (pipe, redirect, chain, subshell, percent-expand, the caret escape
391
+ // itself) whether or not they sit inside a quoted argument.
392
+ return s.replace(/[()%!^"<>&|;, ]/g, '^$&');
393
+ }
394
+
395
+ function buildCmdExeCommand(cmdPath, args) {
396
+ return [escapeCmdArg(cmdPath), ...args.map(escapeCmdArg)].join(' ');
397
+ }
398
+
399
+ // Decides what spawn() actually receives. POSIX and a plain .exe/extensionless
400
+ // binary on win32 are unchanged: no shell, argv passed straight through.
401
+ export function windowsSpawnPlan(argv, platform = process.platform, { allowCmdFallback = false } = {}) {
402
+ const [bin, ...args] = argv;
403
+ if (platform !== 'win32' || !/\.(cmd|bat)$/i.test(bin)) {
404
+ return { command: bin, args, options: {} };
405
+ }
406
+ const script = resolveCmdShim(bin);
407
+ if (script) return { command: process.execPath, args: [script, ...args], options: {} };
408
+ if (!allowCmdFallback) {
409
+ return {
410
+ refuse: `${bin} is a batch file that is not a standard npm shim, and cli-run never passes a prompt through cmd.exe. Reinstall the CLI with npm (npm install -g <package>) so npm writes a standard shim, or put the CLI's .exe first on PATH.`
411
+ };
412
+ }
413
+ const comspec = process.env.ComSpec || process.env.COMSPEC || 'C:\\Windows\\System32\\cmd.exe';
414
+ return { command: comspec, args: ['/d', '/s', '/c', buildCmdExeCommand(bin, args)], options: { windowsVerbatimArguments: true } };
415
+ }
416
+
325
417
  // Kill a lane and everything it spawned. POSIX: the detached process group.
326
- // Windows has no process groups a signal can reach, so taskkill walks the tree (#18).
327
- // Windows is not exercised by CI; this branch is unit-tested by argv capture only.
418
+ // Windows has no process groups a signal can reach, so taskkill walks the
419
+ // tree (#18): whether the direct child is node (the resolved-shim path) or
420
+ // cmd.exe (the fallback), taskkill /T reaches every descendant either way.
421
+ // taskkill is resolved by an absolute path under SystemRoot rather than a
422
+ // bare command name: this call must not depend on PATH containing
423
+ // System32, which real callers cannot guarantee (this project's own test
424
+ // harness deliberately narrows PATH to isolate a fake lane, and hit
425
+ // exactly this on windows-latest CI: `spawn taskkill ENOENT`) and a
426
+ // sandboxed or otherwise stripped-down environment might not either.
427
+ // %SystemRoot% is the documented, always-set location; %windir% is the
428
+ // older equivalent kept as a fallback; C:\Windows is the last resort.
429
+ export function taskkillPath(env = process.env) {
430
+ // Always a Windows path, built with a literal backslash rather than
431
+ // node:path's join(): join() picks its separator from the HOST running
432
+ // this code, not from the OS the path describes, so on a POSIX host (this
433
+ // test suite runs on all three) it would join with "/" and silently
434
+ // produce a path Windows itself would not recognize as one.
435
+ const root = String(env.SystemRoot || env.windir || 'C:\\Windows').replace(/[\\/]+$/, '');
436
+ return `${root}\\System32\\taskkill.exe`;
437
+ }
438
+
328
439
  export function killTree(pid, platform = process.platform, deps = { kill: (p, sig) => process.kill(p, sig), spawn }) {
329
440
  if (platform === 'win32') {
330
- deps.spawn('taskkill', ['/pid', String(pid), '/T', '/F'], { stdio: 'ignore', windowsHide: true });
441
+ const child = deps.spawn(taskkillPath(), ['/pid', String(pid), '/T', '/F'], { stdio: 'ignore', windowsHide: true });
442
+ // Fire-and-forget: nothing here awaits taskkill's own exit. But a spawn
443
+ // failure (ENOENT if this host's layout is unusual, EPERM, ...) still
444
+ // emits an async 'error' event on the returned ChildProcess, and Node
445
+ // treats an EventEmitter's unheard 'error' as fatal, crashing the whole
446
+ // wrapper mid-run over what should be a best-effort cleanup step.
447
+ if (child && typeof child.on === 'function') child.on('error', () => {});
331
448
  return 'taskkill';
332
449
  }
333
450
  deps.kill(-pid, 'SIGKILL');
@@ -367,7 +484,9 @@ export function runBounded(argv, timeoutSec, maxBuffer = 16 * 1024 * 1024) {
367
484
  process.on('SIGINT', onSignal);
368
485
  process.on('SIGTERM', onSignal);
369
486
  try {
370
- child = spawn(argv[0], argv.slice(1), { stdio: ['ignore', 'pipe', 'pipe'], detached: process.platform !== 'win32' });
487
+ const plan = windowsSpawnPlan(argv);
488
+ if (plan.refuse) throw new Error(plan.refuse); // reported as lane unavailable, exit 13
489
+ child = spawn(plan.command, plan.args, { stdio: ['ignore', 'pipe', 'pipe'], detached: process.platform !== 'win32', ...plan.options });
371
490
  } catch (e) {
372
491
  process.off('SIGINT', onSignal);
373
492
  process.off('SIGTERM', onSignal);
package/bin/cli.js CHANGED
@@ -11,6 +11,12 @@ import { resolve, join } from 'node:path';
11
11
  import { which } from '../src/detect.js';
12
12
  import { AIS, LEVELS, TOOLS, PROVIDERS, aisForLevel, agentCandidates, byId, npmSpec } from '../src/catalog.js';
13
13
  import { planFiles, writeFiles, resolveSelection, resolveTools, resolveApis, dirProblems, readManifest, activationSteps, MACHINE_OWNED, RUNTIME, toPosixRel, GENERATOR_VERSION } from '../src/install.js';
14
+ // npm resolves to npm.cmd on Windows; spawning that bare name with no shell
15
+ // hits the same EINVAL bin/cli-run.mjs's lanes did (Node's fix for
16
+ // CVE-2024-27980). windowsSpawnPlan is the same fix reused here rather than
17
+ // duplicated: resolve npm's own cmd-shim and run node on it directly, no
18
+ // shell, or fall back to the escaped cmd.exe path it also provides.
19
+ import { windowsSpawnPlan } from './cli-run.mjs';
14
20
 
15
21
  // One strict parse. Unknown flags, missing values and duplicates are usage
16
22
  // errors (exit 2) before anything is planned, so a typo like --dryy can never
@@ -354,7 +360,10 @@ async function main() {
354
360
  const spec = npmSpec(a); // the same pinned spec the table and the box script use
355
361
  const run = flag('no-install') || yes ? 'n' : await ask(` ${a.name}: run \`npm install -g ${spec}\` now? [y/N]: `, 'n');
356
362
  if (/^y/i.test(run)) {
357
- const r = spawnSync('npm', ['install', '-g', spec], { stdio: 'inherit' });
363
+ // Opt-in cmd.exe fallback: npm.cmd is not a cmd-shim, and every argument here
364
+ // is the catalog's pinned spec, never user text (lanes refuse this path).
365
+ const plan = windowsSpawnPlan([which('npm') || 'npm', 'install', '-g', spec], process.platform, { allowCmdFallback: true });
366
+ const r = spawnSync(plan.command, plan.args, { stdio: 'inherit', ...plan.options });
358
367
  console.log(r.status === 0 ? ` installed ${spec}` : ` npm exited ${r.status}; install it by hand`);
359
368
  } else {
360
369
  console.log(` ${a.name}: npm install -g ${spec} (pinned to the version this installer was released with)`);
package/docs/README.md CHANGED
@@ -7,7 +7,7 @@ The three parts, as reading. The installer writes the working files; these expla
7
7
  | 1 Beginner | you use one LLM or one agent and want it to route well | [part-1-beginner.md](part-1-beginner.md) |
8
8
  | 2 Intermediate | you have several AIs and want to call them through their CLIs from one orchestrator | [part-2-intermediate.md](part-2-intermediate.md) |
9
9
  | 3 Advanced | you want the whole thing running unattended on a virtual machine | [part-3-advanced.md](part-3-advanced.md) |
10
- | Audit brief | the threat model and the two adversarial audit rounds this shipped with | [audit-brief.md](audit-brief.md) |
10
+ | Audit brief | the security notes and the two second-opinion audit rounds this shipped with | [audit-brief.md](audit-brief.md) |
11
11
  | Catalog | what each AI and companion tool in the installer is for, how it installs, how it signs in | [catalog.md](catalog.md) |
12
12
 
13
13
  Each part ends with "what the installer gives you at this level" so the doc and the files agree.
@@ -113,3 +113,16 @@ Three findings reproduced against the 0.1.15 branch before it shipped, none of t
113
113
  - **Fail-open, on purpose.** Every code path that can fail (a malformed state file, a full disk, a rotation race, invalid JSON on stdin, an unrecognized event) is caught and produces no record rather than a thrown error or a non-zero exit; the process always exits 0. A miss here is a missing line in a telemetry log, never a blocked turn, so there is nothing to gate.
114
114
  - **Bounded.** Stdin is drained asynchronously against a combined 1s time cap and 8 MB size cap; a payload that exceeds either is treated as truncated and parsed as nothing, never partially. `--summary` reads the log directly (never spawns anything, never executes a line in it).
115
115
  - **Not yet attacked.** Untested here: two processes racing the same rotation at once (a rename plus an append landing on the same file); a state directory with thousands of leaked files from a long-lived session with a crashed hook (pruning runs, but only on `SubagentStart`, so an install that never starts a subagent again would never prune); behavior if `agent_id` collides across two concurrent subagents (sha256 makes this astronomically unlikely, not impossible).
116
+
117
+ ## New in 0.1.18: the Windows spawn path
118
+
119
+ `bin/cli-run.mjs` runs a lane's binary through `windowsSpawnPlan()` before every `spawn()` call. On POSIX, and for a plain `.exe` or extensionless binary on Windows, this is a no-op: the same argv reaches `spawn()` with no shell, exactly as before. What changed is the two shapes Windows can hand it that used to reach `spawn()` unchanged and throw `EINVAL` (Node's fix for CVE-2024-27980: a `.bat`/`.cmd` target without `shell: true` is refused rather than run through an unsafely-escaped `cmd.exe`).
120
+
121
+ - **What runs, in order.** `resolveCmdShim(cmdPath)` reads the `.cmd` file and looks for the exact line npm's `cmd-shim` package writes: `"%_prog%" ... "<path>" %*`, where `<path>` is `%dp0%`-relative (verified against the real, byte-for-byte output of `cmd-shim@9.0.2`, the package npm itself uses to write a shim from a package.json `bin` entry with a `#!/usr/bin/env node` shebang; `test/judges.test.js` pins that exact fixture). If it matches, the `%dp0%`-relative path is resolved against the `.cmd` file's own directory and checked with `statSync` (must exist, must be a file, must end in `.js`/`.mjs`/`.cjs`); on success, `windowsSpawnPlan()` returns `{ command: process.execPath, args: [scriptPath, ...args] }`, and `spawn()` runs `node <script> <args>` directly. A lane's prompt (argv[1] and on) is text this tool does not control the contents of; this path never puts it anywhere a shell parses it.
122
+ - **Why no shell, ever, for a lane.** Every real lane (grok, codex, agy, hermes, qwen) is an npm-installed Node CLI, so on a real Windows install the resolved-shim branch is the one every run takes. If `resolveCmdShim` returns nothing for a lane (an old cmd-shim layout, a hand-written `.cmd`, or a `.bat`), `windowsSpawnPlan()` returns `{ refuse }` and `cli-run` reports the lane unavailable (exit 13) with a message saying how to fix it. It does not fall back to `cmd.exe`: a batch file re-reads its arguments through `%*` after `cmd.exe` has parsed them once, which is the case CVE-2024-27980 is about, and no escaping fully contains user text through both passes. Removing that path was chosen over guarding it.
123
+ - **The one opt-in `cmd.exe` path, and how its arguments are escaped.** Only a caller passing `{ allowCmdFallback: true }` with arguments it fully controls gets the `cmd.exe` path: today that is the installer's own `npm install -g <pinned spec>` (`npm.cmd` is not a cmd-shim, and every argument comes from the catalog, none from a user). For that caller, `windowsSpawnPlan()` builds one command-line string with `escapeCmdArg`/`buildCmdExeCommand` and returns `{ command: <ComSpec>, args: ['/d', '/s', '/c', <built string>], options: { windowsVerbatimArguments: true } }`. The algorithm is the one documented at [qntm.org/cmd](https://qntm.org/cmd) (the reference writeup of `cmd.exe`'s quoting behavior) and used by the widely-deployed `cross-spawn` package: each argument is quoted the way `CommandLineToArgvW` expects (backslash-doubling before an embedded quote or at the end of the string, then wrapped in `"`), and THEN every `cmd.exe` metacharacter in that quoted text (`( ) % ! ^ " < > & | ; ,` and space) is caret-escaped, because `cmd.exe`'s own line scanner reads those characters off the raw command line before the quoting is honored, quote or no quote. `windowsVerbatimArguments: true` tells Node not to re-quote the string a second, conflicting way. `test/judges.test.js` pins exact expected output for `&`, `|`, `^`, `%`, a literal `"`, a trailing backslash and a literal newline (the last one deliberately unescaped: it is not a `cmd.exe` metacharacter).
124
+ - **`killTree` needs no change for which process it targets, either path.** `taskkill /pid <pid> /T /F` walks the whole descendant tree regardless of whether the direct child is `node` (the resolved-shim path) or `cmd.exe` (the fallback); there is no intermediate shell layer to lose track of in the common case, since there is no shell there at all. It DID need a change for how `taskkill` itself is found: `windows-latest` CI caught a bare `spawn('taskkill', ...)` failing `ENOENT` the first time a lane actually ran end to end there (this project's own test harness deliberately narrows PATH to isolate a fake lane, and that narrowed PATH does not include `System32`; a sandboxed or otherwise stripped-down real environment might not either), and the resulting unheard `error` event on the returned process crashed the whole run over what should be a best-effort cleanup step. `taskkillPath()` resolves the executable under `%SystemRoot%` (falling back through `%windir%` to a fixed path) instead of relying on PATH, built with a literal backslash rather than `node:path`'s `join()`, which picks its separator from the HOST running the code, not the OS the path describes; `killTree` now attaches an `error` listener so any future spawn failure stays a missed cleanup, never a crash.
125
+ - **`bin/cli.js`'s own `npm install -g` prompt reuses this, rather than duplicating it.** The installer's opt-in "run `npm install -g <ai>` now?" prompt had the identical `EINVAL`-shaped defect (`spawnSync('npm', ...)` with no shell), found the same way: it failed the moment its own test actually ran on `windows-latest`. It now resolves `npm` with `which()` and calls `windowsSpawnPlan()`, the same function above, instead of a second copy of the fix.
126
+ - **Not yet attacked for real.** `windowsSpawnPlan`, `resolveCmdShim` and the escaping functions are unit-tested (pure string logic, runs on every CI host) and the resolved-shim path is exercised end to end on `windows-latest` through the fake-lane fixtures in `test/cli.test.js` (installed as a real npm-style `.cmd` shim). A lane can no longer reach `cmd.exe` at all (a unit test pins the refusal). The opt-in `cmd.exe` branch, used only by the installer's own `npm install -g`, is not exercised end to end through a live Windows process in this suite; its escaping is proven by exact-string unit tests only, and its arguments never include user text.
127
+ - **A platform limit found the same way, unrelated to the spawn path itself: Windows has no OS-level signals at all.** `cli-run.mjs`'s graceful shutdown (`process.on('SIGTERM', ...)`, kill the lane's process group, then exit 143/130) is a POSIX guarantee only: `ChildProcess.kill(sig)` on Windows calls `TerminateProcess()` unconditionally for SIGTERM AND SIGINT alike, giving the target process no chance to run any handler at all, proven on `windows-latest` CI (the wrapper died as `{code: null, signal: sig}` for both; a hypothesis that SIGINT gets a real, catchable console-control event on Windows was tried first and measured false in this exact scenario, not assumed). `test/cli.test.js`'s `#13` now expects an unhandled termination for either signal on win32, and the original graceful-exit assertion elsewhere.
128
+ - **Two narrow, individually-verified Windows skips remain, neither in the spawn path itself.** (1) A lane dying mid-run from a real POSIX signal cannot be reproduced on win32: a real Windows lane is a plain `node <script>` process, so it cannot die "by signal" any more than the product being tested can, and the only way a test fixture can even simulate one (a nested `sh -c "...; kill -TERM $$"`) puts an extra node process between cli-run.mjs and the dying shell, so cli-run.mjs observes only that node's translated exit code (measured: MSYS bash's self-kill status leaks through as a plain nonzero exit code, 3840, which this tool already handles honestly via `exit_nonzero`). (2) `weekly-audit.sh`'s watchdog (`bounded()`/`killtree()`, `pgrep -P` plus killing a backgrounded subshell's tree) relies on real bash job control this script only ever runs under on the Ubuntu box it targets; actually executing it against a genuinely hanging stub under Git Bash's job-control emulation hung past a 20s outer timeout on `windows-latest` CI, a known class of MSYS/Cygwin limitation (a `kill -KILL` not reliably reaching the underlying Windows process tree of a backgrounded subshell), not a defect in the generated script, which still renders and syntax-checks correctly.
package/docs/catalog.md CHANGED
@@ -24,7 +24,7 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
24
24
  ### `codex` · Codex CLI (OpenAI, ChatGPT plan)
25
25
 
26
26
  - **Kind:** agent-cli · **Access:** subscription · **Lane:** A · **Level:** 1+
27
- - **Wins at:** second coder and adversarial auditor (a different model family reading your diff)
27
+ - **Wins at:** second coder and second-opinion reviewer (a different model family reading your diff)
28
28
  - **Install:** `npm install -g @openai/codex@0.153.4`
29
29
  - **Sign in:** `codex login` (add `--device-auth` on a machine with no browser)
30
30
  - **Reads rules from:** `AGENTS.md`
@@ -14,7 +14,7 @@ If your agent exposes model choice (Claude Code, Codex, Antigravity), map the ti
14
14
 
15
15
  Three cost levers, always together: tier (price per token), token discipline (how many tokens: read only what you will touch, never re-read, deliverables not narration), effort (how hard each call thinks).
16
16
 
17
- And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **Risk** moves the tier and who reads the result: security, privacy, data loss and irreversible changes are the four worth naming, because none of their failures can be fixed by editing the code afterwards. A one-line change to an auth check is simple and high-risk at once, and it is the risk that decides.
17
+ And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **Stakes** move the tier and who reads the result: security, privacy, data loss and irreversible changes are the four worth naming, because none of their failures can be fixed by editing the code afterwards. A one-line change to an auth check is simple and high-stakes at once, and it is the stakes that decide.
18
18
 
19
19
  Robustness first, cost second. You split tiers because the split produces better work.
20
20
 
@@ -32,9 +32,9 @@ Modifiers: plan big, execute small · never silently retry a failed attempt at t
32
32
 
33
33
  > A gate you cannot fail is not a gate.
34
34
 
35
- "Does this look good?" passes every time. "Name the single biggest risk and the flaw in the request as filed" can come back empty, which is how you know it worked.
35
+ "Does this look good?" passes every time. "Name what is most likely to go wrong, and what the request as filed missed" can come back empty, which is how you know it worked.
36
36
 
37
- Every build gets two checkpoints. **Before writing:** you map what it touches and what could break, then ask the deep tier on the finished map for a named risk and a named flaw. **After it is green:** a fresh context attacks it (bad input, failing dependency, drift from the plan), allowed to answer CLEAN, every finding reproduced before it reaches you. Then the ship, with a rollback named and an explicit yes. Then the loud negative: re-check the old name everywhere and expect zero.
37
+ Every build gets two checkpoints. **Before writing:** you map what it touches and what could break, then ask the deep tier on the finished map for one named weak spot and one gap in the request. **After it is green:** a fresh context challenges it (bad input, failing dependency, drift from the plan), allowed to answer CLEAN, every finding reproduced before it reaches you. Then the ship, with a rollback named and an explicit yes. Then the loud negative: re-check the old name everywhere and expect zero.
38
38
 
39
39
  ## 4. Every hand-off carries a brief
40
40
 
@@ -46,7 +46,7 @@ After anything comprehensive, a fresh turn that hunts for what is **missing**, n
46
46
 
47
47
  ## 6. Deep research, single agent
48
48
 
49
- Plan the sub-questions as their own turn and inspect them before spending anything. Sweep. Then a fresh adversarial turn told to attack the premise. Plant one deliberately wrong figure and see whether it corrects it. Mark every claim CONFIRMED / DISAGREEMENT / REPORTED / UNVERIFIED. Agreement is weak evidence; disagreement is the signal.
49
+ Plan the sub-questions as their own turn and inspect them before spending anything. Sweep. Then a fresh second-opinion turn told to question the premise. Plant one deliberately wrong figure and see whether it corrects it. Mark every claim CONFIRMED / DISAGREEMENT / REPORTED / UNVERIFIED. Agreement is weak evidence; disagreement is the signal.
50
50
 
51
51
  ## 7. Numbers and logic are computed, never guessed
52
52
 
@@ -13,7 +13,7 @@ Rule: never spend a frontier token on a task a cheap tier finishes correctly. Es
13
13
  | Lane | Wins at |
14
14
  |---|---|
15
15
  | the orchestrator (Claude Code, or whichever you chose) | routes, maps, builds, verifies, records; drives the others as CLIs |
16
- | Codex | second coder and adversarial auditor: a different model family reading your diff |
16
+ | Codex | second coder and second-opinion reviewer: a different model family reading your diff |
17
17
  | Antigravity `agy` | deep research sweeps; concurrent fan-out (its subagent call takes an array) |
18
18
  | Grok CLI | X and live web reads at $0 (the same search on the API bills per call) |
19
19
  | Hermes | the free tier: rough drafts, first-pass summaries, divergent reads, cron jobs |
@@ -28,7 +28,7 @@ Every agent CLI can report success and deliver nothing. `bin/cli-run.mjs` builds
28
28
 
29
29
  Every call goes through it. "This lane is flaky" becomes a query over its log instead of an argument. `node bin/cli-run.mjs --doctor` is the first thing to run after install: enabled lanes, binaries on PATH, the route each lane is pinned to, and with `--run` a one-word canary per lane.
30
30
 
31
- There is a second thing a lane can be quietly wrong about. Left unpinned, it runs on **its own config file**, which the runner cannot see: a CLI set up months ago at a low reasoning effort keeps auditing at that effort while your routing docs describe an adversarial pass, and no error is ever raised. `--model` and `--effort` pin it per call, `defaults` in `bin/lanes.json` pins it per lane, and every run records the value requested and where it came from (`flag`, `lanes.json`, `lane_default`). The log claims no actual: grok reports a model id in its output, the other four lanes report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string the durable log never holds.
31
+ There is a second thing a lane can be quietly wrong about. Left unpinned, it runs on **its own config file**, which the runner cannot see: a CLI set up months ago at a low reasoning effort keeps auditing at that effort while your routing docs describe a second-opinion pass, and no error is ever raised. `--model` and `--effort` pin it per call, `defaults` in `bin/lanes.json` pins it per lane, and every run records the value requested and where it came from (`flag`, `lanes.json`, `lane_default`). The log claims no actual: grok reports a model id in its output, the other four lanes report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string the durable log never holds.
32
32
 
33
33
  ## 4. Every delegation carries a task bundle, on both surfaces
34
34
 
@@ -36,7 +36,7 @@ Subagents and CLI lanes are close to the same problem: something that may hold n
36
36
 
37
37
  ## 5. Research: three engines, one triager
38
38
 
39
- Fan the same plan to three model families (web sweep, adversarial read, live data), each as one `cli-run` call. The orchestrator opens the primary sources itself, marks every claim, and writes the only durable record. Expect one engine to return confident unsourced numerics; downgrade it. Weight the engines that report their own gaps. Count dispositions, not briefs.
39
+ Fan the same plan to three model families (web sweep, second-opinion read, live data), each as one `cli-run` call. The orchestrator opens the primary sources itself, marks every claim, and writes the only durable record. Expect one engine to return confident unsourced numerics; downgrade it. Weight the engines that report their own gaps. Count dispositions, not briefs.
40
40
 
41
41
  ## 5a. A finding is a claim, not a fact
42
42
 
@@ -48,7 +48,7 @@ The second pass is now a different model reading the same artifact, in read-only
48
48
 
49
49
  ## 7. The build protocol, bound to lanes
50
50
 
51
- Stage 1 Map: the orchestrator sweeps; CLI lanes critique the map at $0. Stage 2: deep tier, a named risk and a named flaw. Stage 4: scanners on the added lines, fail closed. Stage 5: security-shaped diff → the second coder in read-only audit mode; architecture-shaped → deep tier reviewing build against plan; never both. Two deep checkpoints per build; CLI lanes are uncapped.
51
+ Stage 1 Map: the orchestrator sweeps; CLI lanes critique the map at $0. Stage 2: deep tier, one named weak spot and one gap in the request. Stage 4: scanners on the added lines, refuses by default. Stage 5: security-shaped diff → the second coder in read-only audit mode; architecture-shaped → deep tier reviewing build against plan; never both. Two deep checkpoints per build; CLI lanes are uncapped.
52
52
 
53
53
  ## 8. Privacy gate
54
54
 
package/llms.txt CHANGED
@@ -24,4 +24,4 @@ Levels: 1 beginner (one agent or chat app), 2 intermediate (several agent CLIs,
24
24
  ## Optional
25
25
 
26
26
  - [Security policy](https://github.com/aunysillyme/model-orchestrator/blob/main/SECURITY.md)
27
- - [Audit brief](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/audit-brief.md): the threat model and what has already been attacked
27
+ - [Audit brief](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/audit-brief.md): the security notes and what has already been security-reviewed
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.16",
3
+ "version": "0.1.18",
4
4
  "description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
5
5
  "type": "module",
6
6
  "bin": {
package/src/catalog.js CHANGED
@@ -87,7 +87,7 @@ export const AIS = [
87
87
  bin: 'codex',
88
88
  access: 'subscription',
89
89
  lane: 'A',
90
- role: 'second coder and adversarial auditor (a different model family reading your diff)',
90
+ role: 'second coder and second-opinion reviewer (a different model family reading your diff)',
91
91
  minLevel: 1,
92
92
  install: { npm: '@openai/codex', pin: '0.153.4' },
93
93
  builtAgainst: '0.153.4',
package/src/install.js CHANGED
@@ -155,21 +155,21 @@ export function laneVars(selected) {
155
155
  if (has('hermes')) step0.push(`${cr('hermes')} (the free tier) for rough drafts and divergent reads`);
156
156
  if (has('qwen')) step0.push(`${cr('qwen')} (the cheapest metered lane) for structured bulk, never for anything citing a line, number or source`);
157
157
  if (has('grok')) step0.push(`${cr('grok')} for X and live web reads at $0`);
158
- if (has('codex')) step0.push(`${cr('codex --audit')} for an adversarial read by a second model family`);
158
+ if (has('codex')) step0.push(`${cr('codex --audit')} for a second-opinion read by a second model family`);
159
159
  if (has('agy')) step0.push(`${cr('agy')} for research sweeps and concurrent fan-out`);
160
160
  const stage1 = [];
161
- if (has('codex')) stage1.push(`${cr('codex')} for adversarial critique of the map`);
161
+ if (has('codex')) stage1.push(`${cr('codex')} for a second-opinion critique of the map`);
162
162
  if (has('grok')) stage1.push(`${cr('grok')} to verify current API behaviour instead of trusting recall`);
163
163
  if (has('hermes')) stage1.push(`${cr('hermes')} for a divergent read`);
164
164
  if (has('agy')) stage1.push(`${cr('agy')} for a wide sweep of prior art`);
165
165
  const examples = [];
166
166
  examples.push(has('grok') ? `| "What is trending on X today" | ${cr('grok')} |` : '| "What is trending on X today" | live-researcher (standard tier with web tools) |');
167
- examples.push(has('codex') ? `| "Audit this auth diff" | ${cr('codex --audit')} |` : '| "Audit this auth diff" | code-reviewer at deep tier, in a fresh context told to attack |');
167
+ examples.push(has('codex') ? `| "Audit this auth diff" | ${cr('codex --audit')} |` : '| "Audit this auth diff" | code-reviewer at deep tier, in a fresh context told to challenge |');
168
168
  examples.push(has('qwen') ? `| "Classify these 200 items" | bulk-worker, or ${cr('qwen')} if the items may leave the machine |` : '| "Classify these 200 items" | bulk-worker |');
169
- examples.push(enabled.length >= 2 ? '| "Research this topic properly" | several engines in parallel, see `RESEARCH_TRIAGE.md` |' : '| "Research this topic properly" | deep tier plans, standard tier sweeps, a fresh context attacks; see `RESEARCH_TRIAGE.md` |');
169
+ examples.push(enabled.length >= 2 ? '| "Research this topic properly" | several engines in parallel, see `RESEARCH_TRIAGE.md` |' : '| "Research this topic properly" | deep tier plans, standard tier sweeps, a fresh context challenges; see `RESEARCH_TRIAGE.md` |');
170
170
  const roles = [];
171
171
  if (has('agy')) roles.push('| Web sweep | `cli-run agy` | widest landscape pass |');
172
- if (has('codex')) roles.push('| Adversarial read | `cli-run codex --audit` | attack the premise, hunt for what the others would get wrong |');
172
+ if (has('codex')) roles.push('| Second-opinion read | `cli-run codex --audit` | question the premise, hunt for what the others would get wrong |');
173
173
  if (has('grok')) roles.push('| Live data | `cli-run grok` | dated primary sources, real-time reads |');
174
174
  if (has('hermes')) roles.push('| Cheap divergent read | `cli-run hermes` | another opinion at $0 |');
175
175
  if (has('qwen')) roles.push('| Structured extraction | `cli-run qwen` | pull the facts into a table; never trust its citations without a check |');
@@ -183,12 +183,12 @@ export function laneVars(selected) {
183
183
  return {
184
184
  LANE_STEP0: step0.length ? step0.map((l) => ' - ' + l).join('\n') : ' - none selected yet: every task stays on your primary agent\'s tiers until you add a lane (re-run the installer with more AIs)',
185
185
  STAGE1_LANES: stage1.length ? '; ' + stage1.join(', ') : '',
186
- ATTACK_LANE: has('codex') ? '`cli-run codex --audit` (a second model family in a read-only sandbox)' : 'code-reviewer at deep tier, in a fresh context told to attack and allowed to answer CLEAN',
186
+ ATTACK_LANE: has('codex') ? '`cli-run codex --audit` (a second model family in a read-only sandbox)' : 'code-reviewer at deep tier, in a fresh context told to challenge and allowed to answer CLEAN',
187
187
  LIVE_LANE: has('grok') ? '`cli-run grok` first ($0), then' : '',
188
188
  BULK_LANE: has('qwen') ? ', or `cli-run qwen` if the data may leave your machine' : has('hermes') ? ', or `cli-run hermes` for a free rough pass' : '',
189
189
  LANE_EXAMPLES: examples.join('\n'),
190
190
  RESEARCH_ROLES: roles.join('\n'),
191
- RESEARCH_RUN: run.length ? run.join('\n') : '# no cli-run lane selected: run the sweep on your primary agent, then a fresh adversarial turn (protocols/deep-research.md, level 1 shape)',
191
+ RESEARCH_RUN: run.length ? run.join('\n') : '# no cli-run lane selected: run the sweep on your primary agent, then a fresh second-opinion turn (protocols/deep-research.md, level 1 shape)',
192
192
  RESEARCH_ENGINES: String(run.length)
193
193
  };
194
194
  }
@@ -405,9 +405,30 @@ function vars(opts) {
405
405
  const codecalc = tools.some((t) => t.id === 'codecalc');
406
406
  const dirAbs = resolve(opts.dir || 'ai-orchestrator');
407
407
  const projectAbs = resolve(opts.project || process.cwd());
408
+ // dirPosix backs two things that must read the same on every host:
409
+ // 1. INSTALL_DIR / INSTALL_DIR_SH / INSTALL_DIR_SYSTEMD (below), rendered
410
+ // into vm/jobs/weekly-audit.sh (bash) and vm/jobs/weekly-audit.service
411
+ // (a systemd unit) for the REMOTE Linux box, neither of which can run
412
+ // anywhere but Linux;
413
+ // 2. rulesPath below when --dir falls outside --project, which is
414
+ // prose in generated markdown ("a dir outside the project renders an
415
+ // absolute path"), not a filesystem call.
416
+ // An absolute --dir given as a bare POSIX path ("/opt/x") is never
417
+ // re-resolved through this host's own path semantics for either: a real
418
+ // Windows path always names a drive ("C:\...", caught by the `startsWith`
419
+ // check below falling through to dirAbs), so a bare "/opt/x" only ever
420
+ // means "a Linux path, or documentation text, verbatim" - resolving it
421
+ // with plain path.resolve() reads that leading "/" as drive-relative on
422
+ // win32 and silently turns it into a local path that does not exist,
423
+ // on the box or in the doc. A relative --dir resolves against this
424
+ // host's cwd exactly as before, which is already correct in the common
425
+ // case: level 3 is normally installed by running this CLI ON the box,
426
+ // where "this host" and "the box" are the same filesystem.
427
+ const rawDir = opts.dir || 'ai-orchestrator';
428
+ const dirPosix = rawDir.startsWith('/') ? posix.normalize(rawDir) : dirAbs;
408
429
  let rulesPath = relative(projectAbs, dirAbs).split(sep).join(posix.sep);
409
430
  if (rulesPath === '') rulesPath = '.';
410
- else if (rulesPath.startsWith('..')) rulesPath = dirAbs; // outside the project: absolute is the only honest path
431
+ else if (rulesPath.startsWith('..')) rulesPath = dirPosix; // outside the project: absolute is the only honest path
411
432
  const pinOf = (id) => (toolById[id] && toolById[id].pin) || 'latest';
412
433
  const snippet = snippetFor(primary);
413
434
  const steps = activationSteps({ level, selected, primary, tools, dir: opts.dir, project: opts.project });
@@ -416,7 +437,7 @@ function vars(opts) {
416
437
  // The path route-gate.mjs and subagent-context.mjs resolve at runtime,
417
438
  // relative to CLAUDE_PROJECT_DIR. Mirrors the RULES_PATH fallback below:
418
439
  // outside the project, the honest path is absolute, never a hardcoded one.
419
- const relJoin = (name) => (rulesPath === dirAbs ? join(dirAbs, name) : rulesPath === '.' ? name : rulesPath + '/' + name);
440
+ const relJoin = (name) => (rulesPath === dirPosix ? posix.join(dirPosix, name) : rulesPath === '.' ? name : rulesPath + '/' + name);
420
441
  const rulesFileRel = relJoin(routingFile);
421
442
  const taskBundleRel = relJoin('TASK_BUNDLE.md');
422
443
  // Only claude-code and agy put files under the project root. A chat primary
@@ -449,9 +470,9 @@ function vars(opts) {
449
470
  CODECALC_PIN: pinOf('codecalc'),
450
471
  OBSIDIAN_TC_PIN: pinOf('obsidian-tc'),
451
472
  APIS_LIST: apis.length ? apis.map((prov) => '- ' + prov.name + ' (`' + prov.envName + '`)').join('\n') : '- none: no metered provider key was selected, so the gateway serves only a local lane if you picked one',
452
- INSTALL_DIR: dirAbs,
453
- INSTALL_DIR_SH: shellQuote(dirAbs),
454
- INSTALL_DIR_SYSTEMD: systemdEscape(dirAbs),
473
+ INSTALL_DIR: dirPosix,
474
+ INSTALL_DIR_SH: shellQuote(dirPosix),
475
+ INSTALL_DIR_SYSTEMD: systemdEscape(dirPosix),
455
476
  // vm/README.md step 3 named `grok login` and `agy` whatever you picked (#26).
456
477
  VM_SIGNIN: (() => {
457
478
  const lines = selected.filter((a) => a.bin && a.kind === 'agent-cli').map((a) => ` - ${a.name}: ${a.auth}`);
@@ -784,8 +805,14 @@ export function writeFiles(files, opts) {
784
805
  for (const f of groups[k]) {
785
806
  const abs = resolve(root, f.rel);
786
807
  const exists = existsSync(abs);
787
- const label = k === 'project' ? '[project] ' + f.rel : f.rel;
788
- const key = (k === 'project' ? '[project] ' : '') + f.rel.split(sep).join('/');
808
+ // label is what reaches the terminal report (bin/cli.js's "runtime
809
+ // upgraded:", "runtime CONFLICT, kept:", etc lines): posix-normalized
810
+ // like key, below, so the report reads the same on every host. Before
811
+ // this it carried f.rel verbatim, which is native-separated (join()),
812
+ // so on win32 the report named "bin\cli-run.mjs" while everything
813
+ // else in this tool (docs, other path prose) uses forward slashes.
814
+ const label = (k === 'project' ? '[project] ' : '') + f.rel.split(sep).join('/');
815
+ const key = label;
789
816
  const cls = k === 'dir' ? fileClass(f.rel) : 'document';
790
817
  if (exists && !force) {
791
818
  if (cls === 'document') {
@@ -31,17 +31,27 @@ find reports -maxdepth 1 -name '.audit-*' -type f -mtime +0 -delete 2>/dev/null
31
31
  # pipe open). Process groups do not help here: bash disables job control inside
32
32
  # pipeline subshells, so `kill -- -pid` would kill nothing. A recursive tree
33
33
  # kill via pgrep works on macOS and Linux alike; `timeout(1)` is not on macOS.
34
+ #
35
+ # Two things make a timeout always read as a timeout. killtree freezes each
36
+ # process (SIGSTOP) before walking its children, so a parent cannot run on,
37
+ # print, and exit 0 in the gap after its child dies. And the watchdog leaves a
38
+ # marker when it fires, so bounded returns 124 whatever order the kills land
39
+ # in. Without both, a hung `--version` probe could be recorded as a version
40
+ # string instead of "UNVERIFIED: timed out" (seen once on macOS CI).
34
41
  killtree() {
35
42
  local p="$1" c
43
+ kill -STOP "$p" 2>/dev/null
36
44
  for c in $(pgrep -P "$p" 2>/dev/null); do killtree "$c"; done
37
45
  kill -KILL "$p" 2>/dev/null
38
46
  }
39
47
  bounded() {
40
48
  local secs="$1"; shift
49
+ local fired; fired="$(mktemp "${TMPDIR:-/tmp}/wa-fired.XXXXXX" 2>/dev/null)" && rm -f "$fired"
41
50
  ( "$@" ) & local pid=$!
42
- ( sleep "$secs"; killtree "$pid" ) >/dev/null 2>&1 & local wd=$!
51
+ ( sleep "$secs"; [ -n "$fired" ] && : > "$fired"; killtree "$pid" ) >/dev/null 2>&1 & local wd=$!
43
52
  wait "$pid" 2>/dev/null; local rc=$?
44
53
  killtree "$wd" >/dev/null 2>&1; wait "$wd" 2>/dev/null
54
+ if [ -n "$fired" ] && [ -e "$fired" ]; then rm -f "$fired"; return 124; fi
45
55
  return $rc
46
56
  }
47
57
 
@@ -2,4 +2,4 @@
2
2
 
3
3
  Antigravity CLI custom agents, one per tier plus three checks (`finding-verifier`, `done-verifier`, `reader`), in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
4
4
 
5
- `commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; `auto` keeps deletes and other high-risk commands gated) and `off` for the read-only agents: `code-reviewer`, `finding-verifier`, `live-researcher`, `done-verifier`, `reader`. `model` is a tier: `pro` for deep-planner, `flash` for the rest. `done-verifier` and `reader` never write and, with `commandExecutionPolicy: off`, cannot execute any command at all here, mutating or not: unlike its claude-code counterpart, which does carry an unrestricted `Bash` and stays read-only by its prompt rather than by the tool grant, agy's `done-verifier` is mechanically blocked from shelling out and probes artifacts through whatever read or fetch capability it has instead. Neither is `bulk-worker`, which classifies, tags and transforms items and does write.
5
+ `commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; deletes and other destructive commands still ask before running) and `off` for the read-only agents: `code-reviewer`, `finding-verifier`, `live-researcher`, `done-verifier`, `reader`. `model` is a tier: `pro` for deep-planner, `flash` for the rest. `done-verifier` and `reader` never write and, with `commandExecutionPolicy: off`, cannot execute any command at all here, mutating or not: unlike its claude-code counterpart, which does carry an unrestricted `Bash` and stays read-only by its prompt rather than by the tool grant, agy's `done-verifier` is mechanically blocked from shelling out and probes artifacts through whatever read or fetch capability it has instead. Neither is `bulk-worker`, which classifies, tags and transforms items and does write.
@@ -4,7 +4,7 @@ description: Well-specified execution of a bounded sub-part of a build.
4
4
  model: flash
5
5
  subagent: true
6
6
  mainAgent: true
7
- commandExecutionPolicy: auto # standard build/test commands run unattended; high-risk commands stay gated
7
+ commandExecutionPolicy: auto # standard build/test commands run unattended; destructive commands, like deletes, still ask before running
8
8
  ---
9
9
 
10
10
  # builder
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: finding-verifier
3
- description: Adversarial verification of review findings; tries to disprove each one and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE. Read-only, never repairs.
3
+ description: Second-opinion verification of review findings; tries to disprove each one and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE. Read-only, never repairs.
4
4
  model: flash
5
5
  subagent: true
6
6
  mainAgent: true
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: finding-verifier
3
- description: Adversarial verification of review findings. Use after a review or audit returns findings and before any of them trigger a repair. No file-editing tools; Bash is for read-only checks, bound by the prompt below, not by the tool grant. Tries to DISPROVE each finding and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Do not use to find new problems, and do not use to fix anything.
3
+ description: Second-opinion verification of review findings. Use after a review or audit returns findings and before any of them trigger a repair. No file-editing tools; Bash is for read-only checks, bound by the prompt below, not by the tool grant. Tries to DISPROVE each finding and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Do not use to find new problems, and do not use to fix anything.
4
4
  tools: Read, Glob, Grep, Bash
5
5
  model: sonnet
6
6
  effort: high
@@ -4,8 +4,8 @@
4
4
 
5
5
  ```
6
6
  You follow a model-orchestrator workflow inside this chat. Tiers describe effort, not automatic model switching or cost savings.
7
- Route first: bulk/formatting -> fast; live data -> standard with tools; review -> standard, read-only; ambiguous/high-risk -> deep; otherwise standard. State the tier. Escalate on failure instead of silently retrying.
8
- For builds: map affected parts; identify the biggest risk and any flaw in the request (none is valid with reasons); build and verify; use a fresh turn to challenge the result. Before irreversible actions, explain rollback and ask for approval.
7
+ Route first: bulk/formatting -> fast; live data -> standard with tools; review -> standard, read-only; ambiguous/high-stakes -> deep; otherwise standard. State the tier. Escalate on failure instead of silently retrying.
8
+ For builds: map affected parts; identify what is most likely to go wrong and any gap in the request (none is valid with reasons); build and verify; use a fresh turn to challenge the result. Before irreversible actions, explain rollback and ask for approval.
9
9
  For hand-offs: include purpose, scope, allowed and denied actions, required output, and stopping conditions. A fresh context has none of these instructions.
10
10
  After comprehensive work, check for omissions. Compute consequential numbers and comparisons with a tool; report what was checked and what remains unverified.
11
11
  Before durable writes, search existing records, update their index, use one writer, and label inferences.
@@ -17,7 +17,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
17
17
 
18
18
  A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Measure yours once: spawn a subagent with a one-line task and read its token count. Work smaller than that stays inline.
19
19
 
20
- Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one adversarial pass, an explicit human yes before anything irreversible, then the loud negative.
20
+ Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one challenge pass, an explicit human yes before anything irreversible, then the loud negative.
21
21
 
22
22
  Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A Claude Code subagent loads this CLAUDE.md hierarchy, so it holds the standing rules already, just not this task's scope; a second CLI or a fresh chat window may hold none of them. Absence is denial either way.
23
23
 
@@ -9,7 +9,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
9
9
 
10
10
  Route by capability tier, first match wins: bulk and mechanical -> fast tier · needs live data -> standard tier with tools · review without changing -> standard, read-only · ambiguous or expensive to get wrong -> deep tier, then hand the plan down · everything else -> build it directly at standard tier.
11
11
 
12
- Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: map the blast radius yourself, ask the deep tier for a named risk and a named flaw, build green, scan the added lines, one adversarial pass with every finding reproduced, an explicit human yes before anything irreversible, then re-grep the old identifier and expect zero.
12
+ Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: map everything it touches yourself, ask the deep tier for one named weak spot and one gap in the request, build green, scan the added lines, one challenge pass with every finding reproduced, an explicit human yes before anything irreversible, then re-grep the old identifier and expect zero.
13
13
 
14
14
  Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A fresh context holds none of these rules; absence is denial.
15
15
 
@@ -10,7 +10,7 @@
10
10
  // This script always exits 0, never blocks on stdin past a short bound,
11
11
  // reads at most 64 KB of the rules file through a fixed-size buffer (never
12
12
  // a full read of an arbitrarily large or non-regular file), and never
13
- // executes anything it reads. See docs/audit-brief.md for the threat model.
13
+ // executes anything it reads. See docs/audit-brief.md for the security notes.
14
14
  import { statSync, openSync, readSync, closeSync, realpathSync } from 'node:fs';
15
15
  import { join, isAbsolute } from 'node:path';
16
16
 
@@ -29,8 +29,8 @@ Modifiers:
29
29
 
30
30
  ## The two checkpoints (every build)
31
31
 
32
- - **Checkpoint 1, before writing anything.** You map the blast radius yourself (files, systems, docs, tickets). Then ask the deep tier, on the finished map: *is this the simplest way, what is the single biggest risk, where is the request as filed wrong?* It must return a named risk and a named flaw. Approval alone is not an answer.
33
- - **Checkpoint 2, after the build is green.** Security-shaped diffs get an adversarial read (in a fresh context, told to attack, allowed to answer CLEAN). Architecture-shaped diffs get the deep tier reviewing build against plan. Never both on one diff. Every finding reproduced before it reaches a human.
32
+ - **Checkpoint 1, before writing anything.** You map everything it touches yourself (files, systems, docs, tickets). Then ask the deep tier, on the finished map: *is this the simplest way, what is most likely to go wrong, what did the request miss?* It must return one named weak spot and one gap in the request. Approval alone is not an answer.
33
+ - **Checkpoint 2, after the build is green.** Security-shaped diffs get a second-opinion read (in a fresh context, told to challenge, allowed to answer CLEAN). Architecture-shaped diffs get the deep tier reviewing build against plan. Never both on one diff. Every finding reproduced before it reaches a human.
34
34
 
35
35
  Cap: two deep-tier consults per build. The full procedure is `protocols/build-protocol.md`.
36
36
 
@@ -47,7 +47,7 @@ Level 2 adds `ROUTING.md`, `TIERS.md`, `DELEGATION_MATRIX.md`, `RESEARCH_TRIAGE.
47
47
  ## The three rules that carry everything
48
48
 
49
49
  1. **Route by capability tier, not by model name.** deep = ambiguous or expensive to get wrong · standard = well-specified execution and review · fast = bulk and mechanical. Default down, escalate on evidence.
50
- 2. **A gate you cannot fail is not a gate.** "Does it look good?" passes every time. "Name the single biggest risk and the flaw in the request" can come back empty, which is how you know it worked.
50
+ 2. **A gate you cannot fail is not a gate.** "Does it look good?" passes every time. "Name what is most likely to go wrong, and what the request missed" can come back empty, which is how you know it worked.
51
51
  3. **Exit 0 is not a deliverable.** Any tool, CLI or subagent can report success and hand back nothing. Check for the artifact, not the status line.
52
52
 
53
53
  ## Where things went
@@ -2,7 +2,7 @@
2
2
 
3
3
  **Three phases, eight stages, and every gate is a question that can be answered wrong.**
4
4
 
5
- Fires on any task that builds, codes, implements, migrates or deploys. Rough test: if it would earn an adversarial audit or a tracker issue, it runs this.
5
+ Fires on any task that builds, codes, implements, migrates or deploys. Rough test: if it would earn a second-opinion audit or a tracker issue, it runs this.
6
6
 
7
7
  > **The one rule underneath:** a gate you cannot fail is not a gate. If a stage's exit reads like "confirm it looks good", it is written wrong and it will pass every time, including the times it should not.
8
8
 
@@ -14,7 +14,7 @@ Three corollaries:
14
14
  | Phase | Master question | Stages |
15
15
  |---|---|---|
16
16
  | 1 Pre-build | What exactly are we building, what do we need first, and what does this touch or break? | 0 Route · 1 Map · 2 Judge |
17
- | 2 Build | Is it secure, built on current code, and correct without hidden flaws? | 3 Build · 4 Scan · 5 Attack · 5b Ship gate |
17
+ | 2 Build | Is it secure, built on current code, and correct without hidden flaws? | 3 Build · 4 Scan · 5 Challenge · 5b Ship gate |
18
18
  | 3 Post-build | Did it land everywhere, is it proven against the real thing, and is it recorded? | 6 Verify · 7 Record |
19
19
 
20
20
  The two seams are the point. Pre-build to Build: nothing is written yet, changing your mind costs a conversation. Build to Post-build: the ship, the only irreversible step, the only one that needs an explicit human yes.
@@ -38,9 +38,9 @@ Four bounded questions, not four exhaustive scans. **The builder maps; the judgm
38
38
  ### Stage 2 · Judge (Checkpoint 1)
39
39
  Ask the judgment tier, on the finished map:
40
40
  1. Is this the simplest way to build it, or are we overcomplicating?
41
- 2. What is the single biggest risk, and where is the request as filed wrong?
41
+ 2. What is most likely to go wrong, and what did the request miss?
42
42
 
43
- **Gate:** a **named risk** and a **named flaw in the request**. Approval alone is not an exit; an advisor asked only to approve will approve. If a consult comes back mostly restating the map, the brief asked it to retrieve when it should have asked it to decide.
43
+ **Gate:** **one named weak spot** and **one gap in the request**. Approval alone is not an exit; an advisor asked only to approve will approve. If a consult comes back mostly restating the map, the brief asked it to retrieve when it should have asked it to decide.
44
44
 
45
45
  ## Phase 2 · Build
46
46
 
@@ -56,15 +56,15 @@ Ask the judgment tier, on the finished map:
56
56
  1. Any secret, key or token in the new code?
57
57
  2. Any vulnerability or vulnerable dependency in the lines we added?
58
58
 
59
- Secret detection, static analysis and dependency scanning, filtered to lines this diff added. Fail closed: a missing or erroring scanner exits non-zero, never a silent green.
59
+ Secret detection, static analysis and dependency scanning, filtered to lines this diff added. Refuses by default: a missing or erroring scanner exits non-zero, never a silent green.
60
60
 
61
61
  **Gate:** zero flags on added lines. Pre-existing flags are reported, never inherited as blockers, and never waved through unread. A scanner finding is a claim; read the code before calling it anything.
62
62
 
63
- ### Stage 5 · Attack (Checkpoint 2, one pass, never two)
63
+ ### Stage 5 · Challenge (Checkpoint 2, one pass, never two)
64
64
  1. Can bad input or a bad actor break it, and what happens when a dependency fails?
65
65
  2. Did the build stick to the approved plan, or did unintended changes sneak in?
66
66
 
67
- Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to an adversarial auditor, ideally a **different model family**. Architecture-shaped diffs go to the judgment tier reviewing build against plan. Never both on one diff.
67
+ Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to a second-opinion reviewer, ideally a **different model family**. Architecture-shaped diffs go to the judgment tier reviewing build against plan. Never both on one diff.
68
68
 
69
69
  Findings do not go straight to a repair. Hand them to the **finding-verifier**, whose job is to DISPROVE each one: read the cited line, state what would trigger it, then hunt for the guard, caller or test that makes it impossible. It returns one of three verdicts per finding, and rounding between them is the failure mode to watch for.
70
70
 
@@ -106,7 +106,7 @@ Use a different model family from the one that produced the finding where you ha
106
106
  |---|---|---|
107
107
  {{ROLES_BUILDER_ROW}}
108
108
  | Judgment tier | Stage 2 and the architectural arm of Stage 5. Argues with a finished map | Perform the retrieval |
109
- | Adversarial auditor | The security arm of Stage 5. Attacks the diff | Fix anything |
109
+ | Second-opinion reviewer | The security arm of Stage 5. Reviews the diff | Fix anything |
110
110
  | Mechanical gates | Stage 4 and any always-on guard | Be overridden without reading |
111
111
  | Cheap workers | Bounded sub-parts: bulk passes, wide searches, long loops | Own a stage |
112
112
  | Human | Stage 5b, and any irreversible or architectural call | Be the first line of review |
@@ -119,9 +119,9 @@ Use a different model family from the one that produced the finding where you ha
119
119
  PRE-BUILD
120
120
  [ ] 0 Inputs and access verified by live probe, not assumed
121
121
  [ ] 0 Confirmed this is a build and not a quick fix
122
- [ ] 1 Blast radius written: files, systems, issues
122
+ [ ] 1 Everything it touches written: files, systems, issues
123
123
  [ ] 1 Asked what could break, and whether this already exists
124
- [ ] 2 Judgment tier named a risk AND a flaw in the request
124
+ [ ] 2 Judgment tier named a weak spot AND a gap in the request
125
125
 
126
126
  BUILD
127
127
  [ ] 3 Repo clean, on a branch, base ref recorded
@@ -31,11 +31,11 @@ Agreement is weak evidence. Disagreement is the signal.
31
31
 
32
32
  ## Level 1: one agent
33
33
 
34
- You still get the shape. Run PLAN as its own turn and inspect it before spending anything. Run the sweep. Then run a **fresh-context adversarial turn** with a brief that says "attack the premise; list what this report would get wrong if its sources were stale". Plant one deliberately wrong figure in the brief and see whether it corrects it: if it does not, its confirmations are worth less than they look. Mark every claim.
34
+ You still get the shape. Run PLAN as its own turn and inspect it before spending anything. Run the sweep. Then run a **fresh-context second-opinion turn** with a brief that says "question the premise; list what this report would get wrong if its sources were stale". Plant one deliberately wrong figure in the brief and see whether it corrects it: if it does not, its confirmations are worth less than they look. Mark every claim.
35
35
 
36
36
  ## Level 2 and up: three engines, one triager
37
37
 
38
- Fan out the same PLAN to three different model families through their CLIs (a web-sweep lane, an adversarial-read lane, a live-data lane). Run them through `cli-run` so a run that produced nothing is caught as `rc=10` rather than read as an empty finding. The orchestrator triages: it opens the primary sources itself, marks each claim, and writes the brief. Only the orchestrator writes the durable record; every other engine proposes.
38
+ Fan out the same PLAN to three different model families through their CLIs (a web-sweep lane, a second-opinion-read lane, a live-data lane). Run them through `cli-run` so a run that produced nothing is caught as `rc=10` rather than read as an empty finding. The orchestrator triages: it opens the primary sources itself, marks each claim, and writes the brief. Only the orchestrator writes the durable record; every other engine proposes.
39
39
 
40
40
  Known failure shape: one engine will return confident unsourced numerics and claim full coverage. Downgrade those to hypothesis. The engines that report their own gaps honestly are the ones to weight.
41
41
 
@@ -14,7 +14,7 @@ Verification asks "is what I did correct?". Gap analysis asks "what did I not do
14
14
  ## Who runs it
15
15
 
16
16
  - **Level 1 (one agent):** the same agent, in a fresh turn, with a brief that says "you are looking for what is missing; do not re-verify what is present". Fresh context matters more than a different model.
17
- - **Level 2 and up:** a **different model family** reading the same artifact. Disagreement between two families is the cheapest available signal that something is soft. The adversarial coder lane (a second-opinion CLI in read-only mode) is the natural fit.
17
+ - **Level 2 and up:** a **different model family** reading the same artifact. Disagreement between two families is the cheapest available signal that something is soft. The second-opinion coder lane (read-only mode) is the natural fit.
18
18
  - **Level 3:** make it recurring. A weekly audit job enumerates live state (lanes, jobs, services, model lists), diffs it against the plan, and files a report. It catches the dead lane and the silently renamed model nobody noticed.
19
19
 
20
20
  ## The second half: analyze, compare, suggest
@@ -1,10 +1,10 @@
1
1
  # Propagate: change completeness
2
2
 
3
- **A rename is a refactor, not a single-file edit.** Any change to a name, term, path, slug, schema field, routing rule or shared convention has a blast radius, and the goal is zero silent strays.
3
+ **A rename is a refactor, not a single-file edit.** Any change to a name, term, path, slug, schema field, routing rule or shared convention reaches everything that uses it, and the goal is zero silent strays.
4
4
 
5
5
  This is retrieval work. It stays with the orchestrator (or a cheap worker for the grep sweep). It never goes to the deep tier: a judgment model re-deriving a file list is the most expensive routing mistake there is.
6
6
 
7
- ## 1. Map the blast radius (before editing anything)
7
+ ## 1. Map everything it touches (before editing anything)
8
8
 
9
9
  - **Docs and notes:** backlinks to the thing being renamed; literal search for the old term and its link forms. With obsidian-tc: `get_backlinks`, `search_text`, then `find_unresolved_links` after the change (`protocols/memory-and-record.md`).
10
10
  - **Memory / instructions:** grep every instructions file your agents read (`CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, `QWEN.md`, custom instructions) and any memory store.
@@ -68,7 +68,7 @@ qwen is the lane whose own success flags lie: an upstream 400 comes back as exit
68
68
 
69
69
  ## The route: which model, and how hard it thinks
70
70
 
71
- A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe an adversarial pass, and nothing anywhere says so.
71
+ A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe a second-opinion pass, and nothing anywhere says so.
72
72
 
73
73
  Pin it per call, or per lane:
74
74
 
@@ -114,9 +114,9 @@ The log records what was **requested**, on every record including a run refused
114
114
 
115
115
  That is each vendor's documented headless shape (`-p`, `exec`). Two consequences: argv is visible to other processes on the machine, so a prompt is never the place for a key; and argv is bounded by the OS (`ARG_MAX`), so a very large brief should be referenced by path inside the prompt rather than pasted whole.
116
116
 
117
- ## lanes.json fails closed
117
+ ## lanes.json refuses by default
118
118
 
119
- Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all fail the whole file closed rather than being skipped quietly.
119
+ Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all make the whole file refuse by default rather than being skipped quietly.
120
120
 
121
121
  ## A killed lane is not a deliverable
122
122
 
@@ -15,7 +15,7 @@ Generated {{DATE}} from the AIs you said you have: `{{AI_IDS}}`.
15
15
  | Many independent items each needing its own agent turn | a concurrent fan-out lane | one call, N children, on a subscription |
16
16
  | Live web or social reads | the live-data CLI | subscription-covered; the same search on the API bills per call |
17
17
  | Code review, no changes | standard tier, or the second-coder CLI | a different model family catches what one misses |
18
- | Adversarial audit of a security-shaped diff | the second-coder CLI in read-only audit mode | Claude writes, a second family attacks, the orchestrator reproduces |
18
+ | Second-opinion audit of a security-shaped diff | the second-coder CLI in read-only audit mode | Claude writes, a second family challenges, the orchestrator reproduces |
19
19
  | Deep architecture / planning | deep tier | expensive to get wrong |
20
20
  | Well-specified execution | the orchestrator | execution does not need the top tier |
21
21
  | Long-document analysis | the largest-context lane, or caching on the primary | window size vs re-query cost |
@@ -31,10 +31,10 @@ Rule of thumb: never spend a frontier token on a task a cheap tier finishes corr
31
31
  |---|---|
32
32
  | 0 Route | live probe for access; `cli-run` lanes are $0 and uncapped |
33
33
  | 1 Map | the orchestrator sweeps{{STAGE1_LANES}} |
34
- | 2 Judge | deep tier, on the finished map: a named risk and a named flaw |
34
+ | 2 Judge | deep tier, on the finished map: one named weak spot and one gap in the request |
35
35
  | 3 Build | the orchestrator, against the installed dependency's source |
36
- | 4 Scan | secret + static + dependency scanners, diff-scoped, fail closed |
37
- | 5 Attack | security-shaped diff → {{ATTACK_LANE}}. Architecture-shaped → deep tier, build against plan. Never both |
36
+ | 4 Scan | secret + static + dependency scanners, diff-scoped, refuses by default |
37
+ | 5 Challenge | security-shaped diff → {{ATTACK_LANE}}. Architecture-shaped → deep tier, build against plan. Never both |
38
38
  | 5a Verify findings | finding-verifier, a different model family where you have one: CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a repair |
39
39
  | 5b Ship | rollback id recorded, explicit human yes |
40
40
  | 6 Verify | real test, negative test seen red, old identifier re-grepped to zero |
@@ -58,7 +58,7 @@ One writer per run; every other lane proposes. Search before writing, index in t
58
58
  - **Long context:** mechanical digestion → fast tier in chunks; judgment over a long input → standard tier.
59
59
  - **Token discipline on every delegation:** pass only the context the delegate needs, never the conversation.
60
60
  - **Effort per agent:** deep xhigh, review, verification and build high, live research medium, bulk low.
61
- - **Three inputs, not one:** role picks the agent, complexity moves the effort, risk moves the tier and who reads it. A one-line auth change is simple and high-risk at once, and the risk decides. See `TIERS.md`.
61
+ - **Three inputs, not one:** role picks the agent, complexity moves the effort, stakes move the tier and who reads it. A one-line auth change is simple and high-stakes at once, and the stakes decide. See `TIERS.md`.
62
62
  - **Pin the route when it matters:** a lane with no `--model`/`--effort` and no `defaults` entry in `bin/lanes.json` runs on its own config, which may be nothing like what this file describes. `cli-run --doctor` prints what each lane is pinned to, and every run logs the value requested and where it came from.
63
63
 
64
64
  ## Example routings
@@ -47,21 +47,25 @@ reasoning than the reviewer judging its output.** When the plan is airtight the
47
47
  spec is carrying the thinking, so builder drops to medium. When the plan is
48
48
  vague, fix the plan; do not buy reasoning to paper over it.
49
49
 
50
- **Risk moves the tier and the reader, never just the effort.** These four are
50
+ **Stakes move the tier and the reader, never just the effort.** These four are
51
51
  the ones worth naming, because their failures are not recoverable by editing the
52
52
  code afterwards.
53
53
 
54
- | Risk | Present when the change touches | What it buys |
54
+ Stakes means what a mistake would cost: a security hole, leaked personal data,
55
+ lost data, or something you can't undo. Most tasks are low-stakes and route
56
+ normally.
57
+
58
+ | Stakes | Present when the change touches | What it buys |
55
59
  |---|---|---|
56
- | security | auth, tokens, sessions, routes, untrusted input | the attack pass, ideally a different model family |
60
+ | security | auth, tokens, sessions, routes, untrusted input | the challenge pass, ideally a different model family |
57
61
  | privacy | personal data, anything leaving the machine | the local lane, and a named check on what is sent |
58
62
  | data loss | deletion, bulk mutation, migrations, overwrites | a reviewed rollback path before the change is written |
59
63
  | irreversible | publishing, sending, rotating, anything with an audience | a human yes at Stage 5b, never an agent's |
60
64
 
61
- A risk raises code-reviewer to xhigh, and a security-shaped diff goes to the
62
- attack lane rather than to a second read by the same family. Risk is not a
63
- synonym for difficulty: a one-line change to an auth check is simple and
64
- high-risk at the same time, and it is the risk that decides the route.
65
+ High stakes raise code-reviewer to xhigh, and a security-shaped diff goes to
66
+ the challenge lane rather than to a second read by the same family. Stakes are
67
+ not a synonym for difficulty: a one-line change to an auth check is simple and
68
+ high-stakes at the same time, and it is the stakes that decide the route.
65
69
 
66
70
  **Reserve the top of the ladder for evidence.** xhigh and the escalation tier are
67
71
  bought with a named reason: a reproduced failure, a checkpoint that came back
@@ -70,7 +74,7 @@ task, not an escalation.
70
74
 
71
75
  ## Why split tiers: robustness first, cost second
72
76
 
73
- The split produces better work. The deep tier steers every build twice, and what it steers is **judgment, never retrieval**: the orchestrator sweeps the blast radius itself and hands the deep tier a finished map. Paying deep-tier rates for a file list is the most expensive routing mistake available.
77
+ The split produces better work. The deep tier steers every build twice, and what it steers is **judgment, never retrieval**: the orchestrator sweeps everything it touches itself and hands the deep tier a finished map. Paying deep-tier rates for a file list is the most expensive routing mistake available.
74
78
 
75
79
  Against a baseline of "standard tier with no consults", default checkpoints are a spend increase. That is the accepted trade, not a saving to claim.
76
80
 
@@ -36,7 +36,7 @@ The tools cannot help a model that never reaches for them. codecalc ships `SKILL
36
36
 
37
37
  ## What it is not
38
38
 
39
- Not a cloud sandbox for multi-tenant loads, not a replacement for a vendor's built-in interpreter when zero setup matters more than measurement. Its threat model is single-operator, local, stdio. It earns its keep when the correctness of a claim, not "it ran", is the point.
39
+ Not a cloud sandbox for multi-tenant loads, not a replacement for a vendor's built-in interpreter when zero setup matters more than measurement. It assumes a single-operator, local, stdio setup. It earns its keep when the correctness of a claim, not "it ran", is the point.
40
40
 
41
41
  ## On a box (level 3)
42
42
 
@@ -11,7 +11,7 @@ A durable, searchable, governed store that the protocols can call by name:
11
11
  | Need in the protocols | obsidian-tc tool |
12
12
  |---|---|
13
13
  | find what exists before writing (deep research dedupe, gap analysis) | `semantic_search`, `search_text`, `search_regex` |
14
- | map a rename's blast radius (propagate) | `get_backlinks`, `find_unresolved_links`, `rewrite_link` |
14
+ | map everything a rename touches (propagate) | `get_backlinks`, `find_unresolved_links`, `rewrite_link` |
15
15
  | record the end-to-end doc (build Stage 7) | `write_note` (compare-and-swap, confirmation on overwrite), `patch_note`, `append_note` |
16
16
  | keep inferred content honest | `write_note` with `provenance: "agent_synthesis"` runs a poison scan before the write lands |
17
17
  | keep a shared vault safe for several agents | JWT scopes, per-vault folder ACLs, a read-only kill switch, human-in-the-loop tokens |
@@ -56,9 +56,9 @@ Merge the block; do not replace the file.
56
56
 
57
57
  ## Security posture, read before a second agent touches it
58
58
 
59
- Zero-config mode boots with **auth off and no folder ACL**: anything that can reach the server has the same authority as raw filesystem access to the vault. That is acceptable only because the surface is local-only (the config fail-closes if you enable HTTP on a non-loopback host with auth off, and a DNS-rebinding guard protects loopback). Before exposing it to partially-trusted, remote or multi-agent callers, turn on `auth.mode: "jwt"` and set `acl.readPaths` / `writePaths` / `deletePaths` in the config file. Upstream `SECURITY.md` has the threat model and a private disclosure path.
59
+ Zero-config mode boots with **auth off and no folder ACL**: anything that can reach the server has the same authority as raw filesystem access to the vault. That is acceptable only because the surface is local-only (the config refuses by default if you enable HTTP on a non-loopback host with auth off, and a DNS-rebinding guard protects loopback). Before exposing it to partially-trusted, remote or multi-agent callers, turn on `auth.mode: "jwt"` and set `acl.readPaths` / `writePaths` / `deletePaths` in the config file. Upstream `SECURITY.md` has the security notes and a private disclosure path.
60
60
 
61
- Track record worth knowing: an independent code audit of v1.8.1 (July 2026) found three security-relevant gaps (an ACL fail-closed bypass in enumeration tools, a compare-and-swap bypass through `upsert`, a poison-eligibility gap in preference extraction). All three were fixed upstream before they were filed; verified against the v1.25.0 source on 2026-09-03.
61
+ Track record worth knowing: an independent code audit of v1.8.1 (July 2026) found three security-relevant gaps (an ACL bypass that let enumeration tools skip its refuse-by-default rule, a compare-and-swap bypass through `upsert`, a poison-eligibility gap in preference extraction). All three were fixed upstream before they were filed; verified against the v1.25.0 source on 2026-09-03.
62
62
 
63
63
  ## Level 3
64
64