model-orchestrator 0.1.16 → 0.1.18
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +29 -1
- package/README.md +14 -10
- package/bin/cli-run.mjs +123 -4
- package/bin/cli.js +10 -1
- package/docs/README.md +1 -1
- package/docs/audit-brief.md +13 -0
- package/docs/catalog.md +1 -1
- package/docs/part-1-beginner.md +4 -4
- package/docs/part-2-intermediate.md +4 -4
- package/llms.txt +1 -1
- package/package.json +1 -1
- package/src/catalog.js +1 -1
- package/src/install.js +41 -14
- package/templates/advanced/vm/jobs/weekly-audit.sh +11 -1
- package/templates/agents/agy/README.md +1 -1
- package/templates/agents/agy/builder.md +1 -1
- package/templates/agents/agy/finding-verifier.md +1 -1
- package/templates/agents/claude-code/finding-verifier.md +1 -1
- package/templates/agents/snippets/chat.md +2 -2
- package/templates/agents/snippets/claude-code.md +1 -1
- package/templates/agents/snippets/generic.md +1 -1
- package/templates/agents/snippets/route-gate.mjs +1 -1
- package/templates/beginner/ORCHESTRATOR.md +2 -2
- package/templates/common/README.md +1 -1
- package/templates/common/protocols/build-protocol.md +10 -10
- package/templates/common/protocols/deep-research.md +2 -2
- package/templates/common/protocols/gap-analysis.md +1 -1
- package/templates/common/protocols/propagate.md +2 -2
- package/templates/intermediate/CLI-RUN.md +3 -3
- package/templates/intermediate/DELEGATION_MATRIX.md +1 -1
- package/templates/intermediate/ROUTING.md +4 -4
- package/templates/intermediate/TIERS.md +12 -8
- package/templates/tools/codecalc/CODECALC.md +1 -1
- package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +3 -3
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,32 @@ All notable changes to this project are documented here. The format follows [Kee
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [0.1.18] - 2026-09-11
|
|
8
|
+
|
|
9
|
+
### Fixed
|
|
10
|
+
|
|
11
|
+
- **The level 3 box's weekly audit could record a hung `--version` probe as a version string instead of "UNVERIFIED: timed out".** `weekly-audit.sh`'s `killtree` killed a hung process's children before the process itself, so in that gap the probe's own shell could print and exit 0, and `bounded()` reported success. Seen once on macOS CI ("codex never" in the report); 0 of 20 local runs reproduced it, so it is rare but real on the box. `killtree` now freezes each process (SIGSTOP) before walking its children, and the watchdog leaves a marker when it fires so `bounded()` returns 124 whatever order the kills land in. A new test forces the bad ordering and checks the timeout still reads as one.
|
|
12
|
+
- **On Windows, `bin/cli-run.mjs` could not run any lane at all: `spawn()` threw `EINVAL` for every `.cmd` binary, which is how npm installs every agent CLI there.** Since Node's fix for CVE-2024-27980 (18.20.2, 20.12.2, and every 22.x), spawning a `.bat`/`.cmd` target without `shell: true` throws instead of silently running it through an unsafely-escaped `cmd.exe`. This was a real, shipped defect, not a test gap: a Windows user following this README could not have run a single lane before this release. `bin/cli-run.mjs` now resolves the `.cmd` shim to the Node script npm's own `cmd-shim` tool wrote underneath it and spawns Node directly on that script (`resolveCmdShim`, `windowsSpawnPlan`), so a prompt (untrusted text this tool does not control) never passes through a shell at all in the common case. A lane whose `.cmd`/`.bat` cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, never run through `cmd.exe`: a batch file re-reads its arguments through `%*` after `cmd.exe` has parsed them once, and no escaping fully contains a prompt through both passes. The escaped `cmd.exe` path (the algorithm documented at [qntm.org/cmd](https://qntm.org/cmd) and used by `cross-spawn`, with `windowsVerbatimArguments`) survives only as an explicit opt-in for the installer's own `npm install -g <pinned spec>`, whose arguments never include user text. Verified against the real, byte-for-byte output of `cmd-shim@9.0.2` (the package npm itself uses), not a guessed shape; the escaping is pinned to exact expected strings for `&`, `|`, `^`, `%`, `"`, a trailing backslash and a literal newline in `test/judges.test.js`.
|
|
13
|
+
- **A reconfigure's terminal report (`runtime upgraded:`, `runtime CONFLICT, kept:`, `documents kept:`, and the rest) printed a Windows install's file paths with backslashes**, while every other path this tool prints in generated text uses forward slashes; `src/install.js`'s manifest key was already posix-normalized, but the label built alongside it for the human-readable report was not. `writeFiles()` now builds both from the same posix-normalized value.
|
|
14
|
+
- **A `--dir` outside the project root, or the `vm/` level-3 templates' `INSTALL_DIR`/`INSTALL_DIR_SH`/`INSTALL_DIR_SYSTEMD`, ran through this host's own `path.resolve()`, which reads a leading `/` as drive-relative on win32**, wrong for both: the `vm/` templates describe a REMOTE Linux box (`weekly-audit.sh` is bash, `weekly-audit.service` is a systemd unit, neither of which can run anywhere but Linux), and the outside-project case is documentation prose, not a local filesystem path. An absolute `--dir` given as a bare POSIX path now renders unchanged on every host for both; a real local Windows path (one naming a drive) is untouched, since that case never took this branch.
|
|
15
|
+
- **`killTree`'s win32 branch spawned a bare `taskkill`, which depends on PATH containing `System32`; when it does not, the spawn's `ENOENT` reaches this tool as an unheard `error` event on the returned process and crashes the whole run over what should be a best-effort cleanup step.** Found on `windows-latest` CI the first time a lane actually ran end to end there, once the `EINVAL` fix above stopped hiding it. `taskkillPath()` now resolves the executable under `%SystemRoot%` (falling back through `%windir%` to a fixed path), independent of PATH, and `killTree` attaches an error listener so a spawn failure can never crash the wrapper.
|
|
16
|
+
- **`bin/cli.js`'s own opt-in `npm install -g <ai>` prompt had the identical `EINVAL`-shaped defect as the lane spawn above, in a different file**, since it called `spawnSync('npm', ...)` directly with no shell. It now resolves `npm` with `which()` and runs it through the same `windowsSpawnPlan()` `bin/cli-run.mjs` exports, instead of a second, duplicated fix.
|
|
17
|
+
|
|
18
|
+
### Changed
|
|
19
|
+
|
|
20
|
+
- **The three Windows test-skip groups tracked in [#30](https://github.com/aunysillyme/model-orchestrator/issues/30) are gone, replaced by two narrower, individually-justified skips found by actually running the unskipped suite on `windows-latest` CI, plus the one pre-existing skip this pass never touched (`statSync().mode`'s executable bit; NTFS has nothing equivalent).** `test/cli.test.js`'s fake lane binaries now install as an npm-style `.cmd` shim (verified against the real `cmd-shim@9.0.2` output) pointing at a small Node script, the same shape a real vendor CLI's install takes through `windowsSpawnPlan()` above, instead of a bespoke `sh.exe` bridge that never exercised cli-run.mjs's own spawn path at all; every test in that group but one, and the `#12` upgrade-path pair (fixed by the report-label change above), now runs unconditionally. `test/install.test.js`'s 7 formerly-skipped tests split on what their path assertion actually describes: the ones naming a `vm/` remote-box path stay literal POSIX (now true on every host, per the `INSTALL_DIR*` fix above); the ones naming a LOCAL path (where this run wrote files on this host) now build their expectation with `resolve()`/`join()` instead of a hardcoded POSIX literal, so they assert the same real, platform-native value the product renders rather than a string that only happened to match on POSIX. Two of those seven also had their own, separate bug once actually run on Windows: their fake `curl`/`jq`/`node`/`codex` binaries were placed on a hand-built PATH containing the literal strings `/usr/bin` and `/bin`, which name nothing on that OS; they now prepend the stub directory to the REAL `process.env.PATH` (this job already runs under Git Bash, so that PATH already carries what `bash` itself needs) instead of replacing it with a POSIX-only guess.
|
|
21
|
+
- **New skip: a lane dying mid-run from a real POSIX signal genuinely cannot be reproduced on win32.** cli-run.mjs's `r.signal || r.status === null` branch exists for a real lane crashing or being sent a signal, but a real Windows lane is a plain `node <script>` process (via the resolved cmd-shim), so it cannot die "by signal" any more than the product it is testing can; Windows has no OS-level POSIX signals. The only way this file's fixture can even simulate a signal death is a nested `sh -c "...; kill -TERM $$"` (writeShellStub's win32 branch has to bridge through `sh` for the shell body to run at all), which puts an extra node process between cli-run.mjs and the dying shell; measured on `windows-latest` CI, MSYS bash's own self-kill status leaks through as a plain nonzero exit code (3840), which this tool already handles correctly, just under a different, honest verdict (`exit_nonzero`, not `killed`).
|
|
22
|
+
- **`#13` (SIGTERM/SIGINT to the wrapper) is Windows-aware now, not skipped, for both signals: Windows has no OS-level signals at all**, proven on `windows-latest` CI (the wrapper died as `{code: null, signal: sig}` for SIGTERM AND SIGINT alike; a hypothesis that SIGINT gets a real, catchable console-control event there was tried first and measured false in this exact scenario, not assumed). The test now expects an unhandled termination for both signals on win32; the graceful exit-143/130-and-kill-the-lane-first behavior stays a POSIX guarantee, asserted as before on every other OS.
|
|
23
|
+
- **New skip: `#10`'s watchdog-kill test, and the new `#10b` `bounded()` timeout check, for the same reason.** `weekly-audit.sh`'s `bounded()`/`killtree()` rely on `pgrep -P` and killing a backgrounded subshell's process tree, real bash job control this script only ever runs under on the box it targets (a systemd-scheduled job on Ubuntu, never something a Windows user runs locally). Actually executing that watchdog against a genuinely hanging stub under `windows-latest` CI's Git Bash, rather than just rendering and syntax-checking the script (which the rest of this test group does, and which passes), hung past a 20s outer timeout: MSYS's job-control emulation does not reliably propagate a `kill -KILL` to the underlying Windows process tree of a backgrounded `( subshell ) &`, a known class of MSYS/Cygwin limitation, not a defect in the generated script.
|
|
24
|
+
- One test-report-label assertion in `test/install.test.js` and one in `test/cli.test.js` still hardcoded `path.join()`'s native separator for what is now posix-normalized generated text (the report-label fix above); both now match the posix form.
|
|
25
|
+
- `docs/audit-brief.md` gained a section on the Windows spawn path: what runs, why no shell in the common case, and why a lane is refused rather than run through `cmd.exe`, and how the installer's one opt-in `cmd.exe` call escapes its arguments.
|
|
26
|
+
|
|
27
|
+
## [0.1.17] - 2026-09-11
|
|
28
|
+
|
|
29
|
+
### Changed
|
|
30
|
+
|
|
31
|
+
- **User-facing text now uses plain language instead of security-audit jargon.** Words like "risk", "attack lane", "adversarial", "blast radius", "fail closed" and "threat model" read as alarming to someone deciding whether to try the tool, so they scared off exactly the readers this project needs. No rule any of them described changed, only the wording: "risk" is now "stakes" everywhere it names a routing input (with a one-line definition added to `README.md` and `TIERS.md`), "attack lane" / "Stage 5 Attack" / "attack pass" are now "challenge lane" / "Stage 5 Challenge" / "challenge pass", "adversarial" (auditor, read, critique, turn, pass) is now "second-opinion", "blast radius" is now "everything it touches", "fail(s) closed" is now "refuses by default", and "threat model" is now "security notes" in the files that link to it. `test/prose.test.js` gained a permanent check (`no alarming security wording in user-facing text`) over the purely-prose, user-facing surface (`docs/`, `templates/`, `README.md`, `llms.txt`, `CONTRIBUTING.md`, the PR template) so the old wording cannot silently creep back in. `docs/audit-brief.md`, `SECURITY.md`, `CODE_OF_CONDUCT.md` and code identifiers/comments (for example the `ATTACK_LANE` render var) are unchanged, since these words are expected or load-bearing there.
|
|
32
|
+
|
|
7
33
|
## [0.1.16] - 2026-09-11
|
|
8
34
|
|
|
9
35
|
### Added
|
|
@@ -256,7 +282,9 @@ First release.
|
|
|
256
282
|
- Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
|
|
257
283
|
- Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
|
|
258
284
|
|
|
259
|
-
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.
|
|
285
|
+
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.18...HEAD
|
|
286
|
+
[0.1.18]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.17...v0.1.18
|
|
287
|
+
[0.1.17]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.16...v0.1.17
|
|
260
288
|
[0.1.16]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.15...v0.1.16
|
|
261
289
|
[0.1.15]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...v0.1.15
|
|
262
290
|
[0.1.14]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...v0.1.14
|
package/README.md
CHANGED
|
@@ -44,7 +44,7 @@ Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/pa
|
|
|
44
44
|
| Id | What | Level |
|
|
45
45
|
|---|---|---|
|
|
46
46
|
| `claude-code` | Claude Code CLI, the default orchestrator | 1+ |
|
|
47
|
-
| `codex` | Codex CLI on a ChatGPT plan: second coder
|
|
47
|
+
| `codex` | Codex CLI on a ChatGPT plan: second coder and second-opinion reviewer (a different model family reading your diff) | 1+ |
|
|
48
48
|
| `agy` | Antigravity CLI on a Google AI plan: research sweeps, concurrent fan-out | 1+ |
|
|
49
49
|
| `grok` | Grok CLI on X Premium: live X and web reads at $0 | 1+ |
|
|
50
50
|
| `hermes` | Hermes Agent: the free tier | 2+ |
|
|
@@ -173,7 +173,7 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
|
|
|
173
173
|
6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, only the brief carries what this task needs. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
|
|
174
174
|
7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
|
|
175
175
|
|
|
176
|
-
## Routing by role, complexity and
|
|
176
|
+
## Routing by role, complexity and stakes
|
|
177
177
|
|
|
178
178
|
Role picks the agent. Two more inputs move the choice, and they move it in
|
|
179
179
|
different directions, so `TIERS.md` states them separately rather than folding
|
|
@@ -182,10 +182,14 @@ them into the role:
|
|
|
182
182
|
- **Complexity moves the effort.** A worker executing a finished plan needs less
|
|
183
183
|
reasoning than the reviewer judging its output. When the plan is airtight the
|
|
184
184
|
spec is carrying the thinking.
|
|
185
|
-
- **
|
|
186
|
-
irreversible changes buy the
|
|
187
|
-
human yes. A one-line change to an auth check is simple and high-
|
|
188
|
-
same time, and it is the
|
|
185
|
+
- **Stakes move the tier and the reader.** Security, privacy, data loss and
|
|
186
|
+
irreversible changes buy the challenge lane, a named check, a rollback path or
|
|
187
|
+
a human yes. A one-line change to an auth check is simple and high-stakes at
|
|
188
|
+
the same time, and it is the stakes that decide.
|
|
189
|
+
|
|
190
|
+
Stakes means what a mistake would cost: a security hole, leaked personal data,
|
|
191
|
+
lost data, or something you can't undo. Most tasks are low-stakes and route
|
|
192
|
+
normally.
|
|
189
193
|
|
|
190
194
|
The top of the ladder is bought with evidence: a reproduced failure, an
|
|
191
195
|
unresolved checkpoint, an irreversible change. A task that merely feels hard is
|
|
@@ -221,7 +225,7 @@ The report prints turns, **route-marker coverage** (the percentage of turns whos
|
|
|
221
225
|
A lane with no `--model`, no `--effort` and no `defaults` entry in
|
|
222
226
|
`bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
|
|
223
227
|
CLI configured months ago at a low reasoning effort keeps auditing at that
|
|
224
|
-
effort while your routing docs describe
|
|
228
|
+
effort while your routing docs describe a second-opinion pass.
|
|
225
229
|
|
|
226
230
|
```bash
|
|
227
231
|
node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
|
|
@@ -244,7 +248,7 @@ Install for the tools you have, then let the generated `ROUTING.md` decide the t
|
|
|
244
248
|
|
|
245
249
|
### How do I route tasks to cheaper models?
|
|
246
250
|
|
|
247
|
-
The rules route by role, complexity and
|
|
251
|
+
The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly never gets a frontier token.
|
|
248
252
|
|
|
249
253
|
### Is this an LLM router or an AI gateway?
|
|
250
254
|
|
|
@@ -256,7 +260,7 @@ Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run`
|
|
|
256
260
|
|
|
257
261
|
## Requirements
|
|
258
262
|
|
|
259
|
-
Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22). Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested there
|
|
263
|
+
Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22), including lane execution end to end through `cli-run` against a fake CLI installed the same way npm installs a real one (a `.cmd` shim). `cli-run` never runs a lane through `cmd.exe` when it can avoid it: it resolves the shim to the Node script underneath and spawns Node directly, so a prompt reaching a real lane never passes through a Windows shell. A `.cmd` or `.bat` lane that cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, rather than run through `cmd.exe`: a batch file re-reads its arguments after `cmd.exe` has parsed them once, and no escaping fully contains a prompt through both passes. Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested on Windows too, including SIGTERM/SIGINT to the wrapper (Windows has no OS-level signals: both terminate it unconditionally, verified there rather than treated the same as POSIX). A few narrow skips remain on Windows, each for a POSIX behavior the OS or the CI shell genuinely does not have: `statSync().mode`'s executable bit (NTFS has none), a lane dying mid-run from a real POSIX signal (a real Windows lane cannot die "by signal"), and running `weekly-audit.sh`'s watchdog functions for real under Git Bash's job control, both the end-to-end run and the `bounded()` timeout check (the script itself only ever runs on the Ubuntu box it targets).
|
|
260
264
|
|
|
261
265
|
**Privacy.** The installer sends no telemetry and makes no network call of its own once it is running. Two things around that are worth being exact about:
|
|
262
266
|
|
|
@@ -272,7 +276,7 @@ Add an AI to `src/catalog.js` and every prompt, table, config and doc picks it u
|
|
|
272
276
|
## Credits
|
|
273
277
|
|
|
274
278
|
- [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for
|
|
275
|
-
separating role, complexity and
|
|
279
|
+
separating role, complexity and stakes instead of compressing them into one
|
|
276
280
|
scale, for recording the model and effort a lane was actually asked for, and
|
|
277
281
|
for verifying findings before they trigger repairs. All three shipped in
|
|
278
282
|
0.1.14.
|
package/bin/cli-run.mjs
CHANGED
|
@@ -322,12 +322,129 @@ export function unfence(text) {
|
|
|
322
322
|
return m ? m[1].trim() : t;
|
|
323
323
|
}
|
|
324
324
|
|
|
325
|
+
// --- Windows: spawning a lane without a shell ------------------------------
|
|
326
|
+
// Node's fix for CVE-2024-27980 makes spawn() throw EINVAL for a .bat/.cmd
|
|
327
|
+
// target unless shell:true is set: launching a batch file always goes
|
|
328
|
+
// through cmd.exe, and cmd.exe reads metacharacters (& | ^ < > ( ) % " and
|
|
329
|
+
// space) directly off the command line before the target program's own argv
|
|
330
|
+
// is parsed, even inside quotes. A lane's argv[1] here is a user PROMPT, text
|
|
331
|
+
// this wrapper does not control the contents of, so that is a real injection
|
|
332
|
+
// surface, not a theoretical one.
|
|
333
|
+
//
|
|
334
|
+
// npm installs every CLI on Windows as a "cmd-shim": a short .cmd launcher
|
|
335
|
+
// that hands off to node with a script path (see npm's own `cmd-shim`
|
|
336
|
+
// package). Reading that path out and spawning node directly sidesteps
|
|
337
|
+
// cmd.exe, and the injection surface it carries, entirely: this is the
|
|
338
|
+
// preferred path, used whenever the shim matches the shape cmd-shim writes.
|
|
339
|
+
//
|
|
340
|
+
// A LANE whose .cmd/.bat does not match (hand-written, or an older cmd-shim
|
|
341
|
+
// layout) is refused, not run through cmd.exe: a batch file re-reads its
|
|
342
|
+
// arguments through %* after cmd.exe has already parsed them once, which is
|
|
343
|
+
// the case CVE-2024-27980 is about, and no escaping fully contains user text
|
|
344
|
+
// through both passes. Removing that path beats guarding it.
|
|
345
|
+
//
|
|
346
|
+
// The cmd.exe path survives only for a caller that opts in with
|
|
347
|
+
// { allowCmdFallback: true } and passes arguments it fully controls (the
|
|
348
|
+
// installer's own `npm install -g <pinned spec>`, whose npm.cmd is not a
|
|
349
|
+
// cmd-shim). It uses the caret-escaping algorithm documented at
|
|
350
|
+
// https://qntm.org/cmd and used by `cross-spawn`: quote each argument for
|
|
351
|
+
// CommandLineToArgvW, THEN caret-escape cmd.exe's own metacharacters, THEN
|
|
352
|
+
// pass the whole line with windowsVerbatimArguments so Node does not
|
|
353
|
+
// re-quote it a second, conflicting way.
|
|
354
|
+
const NPM_CMD_SHIM = /"%_prog%"\s+"([^"]+)"\s*%\*/;
|
|
355
|
+
export function resolveCmdShim(cmdPath) {
|
|
356
|
+
let text;
|
|
357
|
+
try {
|
|
358
|
+
text = readFileSync(cmdPath, 'utf8');
|
|
359
|
+
} catch {
|
|
360
|
+
return null;
|
|
361
|
+
}
|
|
362
|
+
const m = NPM_CMD_SHIM.exec(text);
|
|
363
|
+
if (!m) return null;
|
|
364
|
+
const dp0 = /^%~?dp0%?[\\/]?/i;
|
|
365
|
+
if (!dp0.test(m[1])) return null; // only the %dp0%-relative shape cmd-shim writes
|
|
366
|
+
const rel = m[1].replace(dp0, '').replace(/\\/g, '/');
|
|
367
|
+
let script;
|
|
368
|
+
try {
|
|
369
|
+
script = resolve(dirname(cmdPath), rel);
|
|
370
|
+
if (!statSync(script).isFile()) return null;
|
|
371
|
+
} catch {
|
|
372
|
+
return null;
|
|
373
|
+
}
|
|
374
|
+
// Only ever hand off to node for a real JS entry point; anything else (a
|
|
375
|
+
// shim generated for a non-node binary, or a hand-edited file) falls
|
|
376
|
+
// through to the cmd.exe fallback instead of being executed as a script.
|
|
377
|
+
return /\.(m?js|cjs)$/i.test(script) ? script : null;
|
|
378
|
+
}
|
|
379
|
+
|
|
380
|
+
function escapeCmdArg(arg) {
|
|
381
|
+
let s = String(arg);
|
|
382
|
+
// A run of backslashes immediately before a quote (or at the very end of
|
|
383
|
+
// the argument) must be doubled, or CommandLineToArgvW on the receiving
|
|
384
|
+
// end eats one; this is the standard Windows argv-quoting rule, not a
|
|
385
|
+
// cmd.exe-specific one.
|
|
386
|
+
s = s.replace(/(\\*)"/g, '$1$1\\"');
|
|
387
|
+
s = s.replace(/(\\*)$/, '$1$1');
|
|
388
|
+
s = `"${s}"`;
|
|
389
|
+
// cmd.exe reads these characters off the raw command line and acts on
|
|
390
|
+
// them (pipe, redirect, chain, subshell, percent-expand, the caret escape
|
|
391
|
+
// itself) whether or not they sit inside a quoted argument.
|
|
392
|
+
return s.replace(/[()%!^"<>&|;, ]/g, '^$&');
|
|
393
|
+
}
|
|
394
|
+
|
|
395
|
+
function buildCmdExeCommand(cmdPath, args) {
|
|
396
|
+
return [escapeCmdArg(cmdPath), ...args.map(escapeCmdArg)].join(' ');
|
|
397
|
+
}
|
|
398
|
+
|
|
399
|
+
// Decides what spawn() actually receives. POSIX and a plain .exe/extensionless
|
|
400
|
+
// binary on win32 are unchanged: no shell, argv passed straight through.
|
|
401
|
+
export function windowsSpawnPlan(argv, platform = process.platform, { allowCmdFallback = false } = {}) {
|
|
402
|
+
const [bin, ...args] = argv;
|
|
403
|
+
if (platform !== 'win32' || !/\.(cmd|bat)$/i.test(bin)) {
|
|
404
|
+
return { command: bin, args, options: {} };
|
|
405
|
+
}
|
|
406
|
+
const script = resolveCmdShim(bin);
|
|
407
|
+
if (script) return { command: process.execPath, args: [script, ...args], options: {} };
|
|
408
|
+
if (!allowCmdFallback) {
|
|
409
|
+
return {
|
|
410
|
+
refuse: `${bin} is a batch file that is not a standard npm shim, and cli-run never passes a prompt through cmd.exe. Reinstall the CLI with npm (npm install -g <package>) so npm writes a standard shim, or put the CLI's .exe first on PATH.`
|
|
411
|
+
};
|
|
412
|
+
}
|
|
413
|
+
const comspec = process.env.ComSpec || process.env.COMSPEC || 'C:\\Windows\\System32\\cmd.exe';
|
|
414
|
+
return { command: comspec, args: ['/d', '/s', '/c', buildCmdExeCommand(bin, args)], options: { windowsVerbatimArguments: true } };
|
|
415
|
+
}
|
|
416
|
+
|
|
325
417
|
// Kill a lane and everything it spawned. POSIX: the detached process group.
|
|
326
|
-
// Windows has no process groups a signal can reach, so taskkill walks the
|
|
327
|
-
//
|
|
418
|
+
// Windows has no process groups a signal can reach, so taskkill walks the
|
|
419
|
+
// tree (#18): whether the direct child is node (the resolved-shim path) or
|
|
420
|
+
// cmd.exe (the fallback), taskkill /T reaches every descendant either way.
|
|
421
|
+
// taskkill is resolved by an absolute path under SystemRoot rather than a
|
|
422
|
+
// bare command name: this call must not depend on PATH containing
|
|
423
|
+
// System32, which real callers cannot guarantee (this project's own test
|
|
424
|
+
// harness deliberately narrows PATH to isolate a fake lane, and hit
|
|
425
|
+
// exactly this on windows-latest CI: `spawn taskkill ENOENT`) and a
|
|
426
|
+
// sandboxed or otherwise stripped-down environment might not either.
|
|
427
|
+
// %SystemRoot% is the documented, always-set location; %windir% is the
|
|
428
|
+
// older equivalent kept as a fallback; C:\Windows is the last resort.
|
|
429
|
+
export function taskkillPath(env = process.env) {
|
|
430
|
+
// Always a Windows path, built with a literal backslash rather than
|
|
431
|
+
// node:path's join(): join() picks its separator from the HOST running
|
|
432
|
+
// this code, not from the OS the path describes, so on a POSIX host (this
|
|
433
|
+
// test suite runs on all three) it would join with "/" and silently
|
|
434
|
+
// produce a path Windows itself would not recognize as one.
|
|
435
|
+
const root = String(env.SystemRoot || env.windir || 'C:\\Windows').replace(/[\\/]+$/, '');
|
|
436
|
+
return `${root}\\System32\\taskkill.exe`;
|
|
437
|
+
}
|
|
438
|
+
|
|
328
439
|
export function killTree(pid, platform = process.platform, deps = { kill: (p, sig) => process.kill(p, sig), spawn }) {
|
|
329
440
|
if (platform === 'win32') {
|
|
330
|
-
deps.spawn(
|
|
441
|
+
const child = deps.spawn(taskkillPath(), ['/pid', String(pid), '/T', '/F'], { stdio: 'ignore', windowsHide: true });
|
|
442
|
+
// Fire-and-forget: nothing here awaits taskkill's own exit. But a spawn
|
|
443
|
+
// failure (ENOENT if this host's layout is unusual, EPERM, ...) still
|
|
444
|
+
// emits an async 'error' event on the returned ChildProcess, and Node
|
|
445
|
+
// treats an EventEmitter's unheard 'error' as fatal, crashing the whole
|
|
446
|
+
// wrapper mid-run over what should be a best-effort cleanup step.
|
|
447
|
+
if (child && typeof child.on === 'function') child.on('error', () => {});
|
|
331
448
|
return 'taskkill';
|
|
332
449
|
}
|
|
333
450
|
deps.kill(-pid, 'SIGKILL');
|
|
@@ -367,7 +484,9 @@ export function runBounded(argv, timeoutSec, maxBuffer = 16 * 1024 * 1024) {
|
|
|
367
484
|
process.on('SIGINT', onSignal);
|
|
368
485
|
process.on('SIGTERM', onSignal);
|
|
369
486
|
try {
|
|
370
|
-
|
|
487
|
+
const plan = windowsSpawnPlan(argv);
|
|
488
|
+
if (plan.refuse) throw new Error(plan.refuse); // reported as lane unavailable, exit 13
|
|
489
|
+
child = spawn(plan.command, plan.args, { stdio: ['ignore', 'pipe', 'pipe'], detached: process.platform !== 'win32', ...plan.options });
|
|
371
490
|
} catch (e) {
|
|
372
491
|
process.off('SIGINT', onSignal);
|
|
373
492
|
process.off('SIGTERM', onSignal);
|
package/bin/cli.js
CHANGED
|
@@ -11,6 +11,12 @@ import { resolve, join } from 'node:path';
|
|
|
11
11
|
import { which } from '../src/detect.js';
|
|
12
12
|
import { AIS, LEVELS, TOOLS, PROVIDERS, aisForLevel, agentCandidates, byId, npmSpec } from '../src/catalog.js';
|
|
13
13
|
import { planFiles, writeFiles, resolveSelection, resolveTools, resolveApis, dirProblems, readManifest, activationSteps, MACHINE_OWNED, RUNTIME, toPosixRel, GENERATOR_VERSION } from '../src/install.js';
|
|
14
|
+
// npm resolves to npm.cmd on Windows; spawning that bare name with no shell
|
|
15
|
+
// hits the same EINVAL bin/cli-run.mjs's lanes did (Node's fix for
|
|
16
|
+
// CVE-2024-27980). windowsSpawnPlan is the same fix reused here rather than
|
|
17
|
+
// duplicated: resolve npm's own cmd-shim and run node on it directly, no
|
|
18
|
+
// shell, or fall back to the escaped cmd.exe path it also provides.
|
|
19
|
+
import { windowsSpawnPlan } from './cli-run.mjs';
|
|
14
20
|
|
|
15
21
|
// One strict parse. Unknown flags, missing values and duplicates are usage
|
|
16
22
|
// errors (exit 2) before anything is planned, so a typo like --dryy can never
|
|
@@ -354,7 +360,10 @@ async function main() {
|
|
|
354
360
|
const spec = npmSpec(a); // the same pinned spec the table and the box script use
|
|
355
361
|
const run = flag('no-install') || yes ? 'n' : await ask(` ${a.name}: run \`npm install -g ${spec}\` now? [y/N]: `, 'n');
|
|
356
362
|
if (/^y/i.test(run)) {
|
|
357
|
-
|
|
363
|
+
// Opt-in cmd.exe fallback: npm.cmd is not a cmd-shim, and every argument here
|
|
364
|
+
// is the catalog's pinned spec, never user text (lanes refuse this path).
|
|
365
|
+
const plan = windowsSpawnPlan([which('npm') || 'npm', 'install', '-g', spec], process.platform, { allowCmdFallback: true });
|
|
366
|
+
const r = spawnSync(plan.command, plan.args, { stdio: 'inherit', ...plan.options });
|
|
358
367
|
console.log(r.status === 0 ? ` installed ${spec}` : ` npm exited ${r.status}; install it by hand`);
|
|
359
368
|
} else {
|
|
360
369
|
console.log(` ${a.name}: npm install -g ${spec} (pinned to the version this installer was released with)`);
|
package/docs/README.md
CHANGED
|
@@ -7,7 +7,7 @@ The three parts, as reading. The installer writes the working files; these expla
|
|
|
7
7
|
| 1 Beginner | you use one LLM or one agent and want it to route well | [part-1-beginner.md](part-1-beginner.md) |
|
|
8
8
|
| 2 Intermediate | you have several AIs and want to call them through their CLIs from one orchestrator | [part-2-intermediate.md](part-2-intermediate.md) |
|
|
9
9
|
| 3 Advanced | you want the whole thing running unattended on a virtual machine | [part-3-advanced.md](part-3-advanced.md) |
|
|
10
|
-
| Audit brief | the
|
|
10
|
+
| Audit brief | the security notes and the two second-opinion audit rounds this shipped with | [audit-brief.md](audit-brief.md) |
|
|
11
11
|
| Catalog | what each AI and companion tool in the installer is for, how it installs, how it signs in | [catalog.md](catalog.md) |
|
|
12
12
|
|
|
13
13
|
Each part ends with "what the installer gives you at this level" so the doc and the files agree.
|
package/docs/audit-brief.md
CHANGED
|
@@ -113,3 +113,16 @@ Three findings reproduced against the 0.1.15 branch before it shipped, none of t
|
|
|
113
113
|
- **Fail-open, on purpose.** Every code path that can fail (a malformed state file, a full disk, a rotation race, invalid JSON on stdin, an unrecognized event) is caught and produces no record rather than a thrown error or a non-zero exit; the process always exits 0. A miss here is a missing line in a telemetry log, never a blocked turn, so there is nothing to gate.
|
|
114
114
|
- **Bounded.** Stdin is drained asynchronously against a combined 1s time cap and 8 MB size cap; a payload that exceeds either is treated as truncated and parsed as nothing, never partially. `--summary` reads the log directly (never spawns anything, never executes a line in it).
|
|
115
115
|
- **Not yet attacked.** Untested here: two processes racing the same rotation at once (a rename plus an append landing on the same file); a state directory with thousands of leaked files from a long-lived session with a crashed hook (pruning runs, but only on `SubagentStart`, so an install that never starts a subagent again would never prune); behavior if `agent_id` collides across two concurrent subagents (sha256 makes this astronomically unlikely, not impossible).
|
|
116
|
+
|
|
117
|
+
## New in 0.1.18: the Windows spawn path
|
|
118
|
+
|
|
119
|
+
`bin/cli-run.mjs` runs a lane's binary through `windowsSpawnPlan()` before every `spawn()` call. On POSIX, and for a plain `.exe` or extensionless binary on Windows, this is a no-op: the same argv reaches `spawn()` with no shell, exactly as before. What changed is the two shapes Windows can hand it that used to reach `spawn()` unchanged and throw `EINVAL` (Node's fix for CVE-2024-27980: a `.bat`/`.cmd` target without `shell: true` is refused rather than run through an unsafely-escaped `cmd.exe`).
|
|
120
|
+
|
|
121
|
+
- **What runs, in order.** `resolveCmdShim(cmdPath)` reads the `.cmd` file and looks for the exact line npm's `cmd-shim` package writes: `"%_prog%" ... "<path>" %*`, where `<path>` is `%dp0%`-relative (verified against the real, byte-for-byte output of `cmd-shim@9.0.2`, the package npm itself uses to write a shim from a package.json `bin` entry with a `#!/usr/bin/env node` shebang; `test/judges.test.js` pins that exact fixture). If it matches, the `%dp0%`-relative path is resolved against the `.cmd` file's own directory and checked with `statSync` (must exist, must be a file, must end in `.js`/`.mjs`/`.cjs`); on success, `windowsSpawnPlan()` returns `{ command: process.execPath, args: [scriptPath, ...args] }`, and `spawn()` runs `node <script> <args>` directly. A lane's prompt (argv[1] and on) is text this tool does not control the contents of; this path never puts it anywhere a shell parses it.
|
|
122
|
+
- **Why no shell, ever, for a lane.** Every real lane (grok, codex, agy, hermes, qwen) is an npm-installed Node CLI, so on a real Windows install the resolved-shim branch is the one every run takes. If `resolveCmdShim` returns nothing for a lane (an old cmd-shim layout, a hand-written `.cmd`, or a `.bat`), `windowsSpawnPlan()` returns `{ refuse }` and `cli-run` reports the lane unavailable (exit 13) with a message saying how to fix it. It does not fall back to `cmd.exe`: a batch file re-reads its arguments through `%*` after `cmd.exe` has parsed them once, which is the case CVE-2024-27980 is about, and no escaping fully contains user text through both passes. Removing that path was chosen over guarding it.
|
|
123
|
+
- **The one opt-in `cmd.exe` path, and how its arguments are escaped.** Only a caller passing `{ allowCmdFallback: true }` with arguments it fully controls gets the `cmd.exe` path: today that is the installer's own `npm install -g <pinned spec>` (`npm.cmd` is not a cmd-shim, and every argument comes from the catalog, none from a user). For that caller, `windowsSpawnPlan()` builds one command-line string with `escapeCmdArg`/`buildCmdExeCommand` and returns `{ command: <ComSpec>, args: ['/d', '/s', '/c', <built string>], options: { windowsVerbatimArguments: true } }`. The algorithm is the one documented at [qntm.org/cmd](https://qntm.org/cmd) (the reference writeup of `cmd.exe`'s quoting behavior) and used by the widely-deployed `cross-spawn` package: each argument is quoted the way `CommandLineToArgvW` expects (backslash-doubling before an embedded quote or at the end of the string, then wrapped in `"`), and THEN every `cmd.exe` metacharacter in that quoted text (`( ) % ! ^ " < > & | ; ,` and space) is caret-escaped, because `cmd.exe`'s own line scanner reads those characters off the raw command line before the quoting is honored, quote or no quote. `windowsVerbatimArguments: true` tells Node not to re-quote the string a second, conflicting way. `test/judges.test.js` pins exact expected output for `&`, `|`, `^`, `%`, a literal `"`, a trailing backslash and a literal newline (the last one deliberately unescaped: it is not a `cmd.exe` metacharacter).
|
|
124
|
+
- **`killTree` needs no change for which process it targets, either path.** `taskkill /pid <pid> /T /F` walks the whole descendant tree regardless of whether the direct child is `node` (the resolved-shim path) or `cmd.exe` (the fallback); there is no intermediate shell layer to lose track of in the common case, since there is no shell there at all. It DID need a change for how `taskkill` itself is found: `windows-latest` CI caught a bare `spawn('taskkill', ...)` failing `ENOENT` the first time a lane actually ran end to end there (this project's own test harness deliberately narrows PATH to isolate a fake lane, and that narrowed PATH does not include `System32`; a sandboxed or otherwise stripped-down real environment might not either), and the resulting unheard `error` event on the returned process crashed the whole run over what should be a best-effort cleanup step. `taskkillPath()` resolves the executable under `%SystemRoot%` (falling back through `%windir%` to a fixed path) instead of relying on PATH, built with a literal backslash rather than `node:path`'s `join()`, which picks its separator from the HOST running the code, not the OS the path describes; `killTree` now attaches an `error` listener so any future spawn failure stays a missed cleanup, never a crash.
|
|
125
|
+
- **`bin/cli.js`'s own `npm install -g` prompt reuses this, rather than duplicating it.** The installer's opt-in "run `npm install -g <ai>` now?" prompt had the identical `EINVAL`-shaped defect (`spawnSync('npm', ...)` with no shell), found the same way: it failed the moment its own test actually ran on `windows-latest`. It now resolves `npm` with `which()` and calls `windowsSpawnPlan()`, the same function above, instead of a second copy of the fix.
|
|
126
|
+
- **Not yet attacked for real.** `windowsSpawnPlan`, `resolveCmdShim` and the escaping functions are unit-tested (pure string logic, runs on every CI host) and the resolved-shim path is exercised end to end on `windows-latest` through the fake-lane fixtures in `test/cli.test.js` (installed as a real npm-style `.cmd` shim). A lane can no longer reach `cmd.exe` at all (a unit test pins the refusal). The opt-in `cmd.exe` branch, used only by the installer's own `npm install -g`, is not exercised end to end through a live Windows process in this suite; its escaping is proven by exact-string unit tests only, and its arguments never include user text.
|
|
127
|
+
- **A platform limit found the same way, unrelated to the spawn path itself: Windows has no OS-level signals at all.** `cli-run.mjs`'s graceful shutdown (`process.on('SIGTERM', ...)`, kill the lane's process group, then exit 143/130) is a POSIX guarantee only: `ChildProcess.kill(sig)` on Windows calls `TerminateProcess()` unconditionally for SIGTERM AND SIGINT alike, giving the target process no chance to run any handler at all, proven on `windows-latest` CI (the wrapper died as `{code: null, signal: sig}` for both; a hypothesis that SIGINT gets a real, catchable console-control event on Windows was tried first and measured false in this exact scenario, not assumed). `test/cli.test.js`'s `#13` now expects an unhandled termination for either signal on win32, and the original graceful-exit assertion elsewhere.
|
|
128
|
+
- **Two narrow, individually-verified Windows skips remain, neither in the spawn path itself.** (1) A lane dying mid-run from a real POSIX signal cannot be reproduced on win32: a real Windows lane is a plain `node <script>` process, so it cannot die "by signal" any more than the product being tested can, and the only way a test fixture can even simulate one (a nested `sh -c "...; kill -TERM $$"`) puts an extra node process between cli-run.mjs and the dying shell, so cli-run.mjs observes only that node's translated exit code (measured: MSYS bash's self-kill status leaks through as a plain nonzero exit code, 3840, which this tool already handles honestly via `exit_nonzero`). (2) `weekly-audit.sh`'s watchdog (`bounded()`/`killtree()`, `pgrep -P` plus killing a backgrounded subshell's tree) relies on real bash job control this script only ever runs under on the Ubuntu box it targets; actually executing it against a genuinely hanging stub under Git Bash's job-control emulation hung past a 20s outer timeout on `windows-latest` CI, a known class of MSYS/Cygwin limitation (a `kill -KILL` not reliably reaching the underlying Windows process tree of a backgrounded subshell), not a defect in the generated script, which still renders and syntax-checks correctly.
|
package/docs/catalog.md
CHANGED
|
@@ -24,7 +24,7 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
|
|
|
24
24
|
### `codex` · Codex CLI (OpenAI, ChatGPT plan)
|
|
25
25
|
|
|
26
26
|
- **Kind:** agent-cli · **Access:** subscription · **Lane:** A · **Level:** 1+
|
|
27
|
-
- **Wins at:** second coder and
|
|
27
|
+
- **Wins at:** second coder and second-opinion reviewer (a different model family reading your diff)
|
|
28
28
|
- **Install:** `npm install -g @openai/codex@0.153.4`
|
|
29
29
|
- **Sign in:** `codex login` (add `--device-auth` on a machine with no browser)
|
|
30
30
|
- **Reads rules from:** `AGENTS.md`
|
package/docs/part-1-beginner.md
CHANGED
|
@@ -14,7 +14,7 @@ If your agent exposes model choice (Claude Code, Codex, Antigravity), map the ti
|
|
|
14
14
|
|
|
15
15
|
Three cost levers, always together: tier (price per token), token discipline (how many tokens: read only what you will touch, never re-read, deliverables not narration), effort (how hard each call thinks).
|
|
16
16
|
|
|
17
|
-
And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **
|
|
17
|
+
And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **Stakes** move the tier and who reads the result: security, privacy, data loss and irreversible changes are the four worth naming, because none of their failures can be fixed by editing the code afterwards. A one-line change to an auth check is simple and high-stakes at once, and it is the stakes that decide.
|
|
18
18
|
|
|
19
19
|
Robustness first, cost second. You split tiers because the split produces better work.
|
|
20
20
|
|
|
@@ -32,9 +32,9 @@ Modifiers: plan big, execute small · never silently retry a failed attempt at t
|
|
|
32
32
|
|
|
33
33
|
> A gate you cannot fail is not a gate.
|
|
34
34
|
|
|
35
|
-
"Does this look good?" passes every time. "Name
|
|
35
|
+
"Does this look good?" passes every time. "Name what is most likely to go wrong, and what the request as filed missed" can come back empty, which is how you know it worked.
|
|
36
36
|
|
|
37
|
-
Every build gets two checkpoints. **Before writing:** you map what it touches and what could break, then ask the deep tier on the finished map for
|
|
37
|
+
Every build gets two checkpoints. **Before writing:** you map what it touches and what could break, then ask the deep tier on the finished map for one named weak spot and one gap in the request. **After it is green:** a fresh context challenges it (bad input, failing dependency, drift from the plan), allowed to answer CLEAN, every finding reproduced before it reaches you. Then the ship, with a rollback named and an explicit yes. Then the loud negative: re-check the old name everywhere and expect zero.
|
|
38
38
|
|
|
39
39
|
## 4. Every hand-off carries a brief
|
|
40
40
|
|
|
@@ -46,7 +46,7 @@ After anything comprehensive, a fresh turn that hunts for what is **missing**, n
|
|
|
46
46
|
|
|
47
47
|
## 6. Deep research, single agent
|
|
48
48
|
|
|
49
|
-
Plan the sub-questions as their own turn and inspect them before spending anything. Sweep. Then a fresh
|
|
49
|
+
Plan the sub-questions as their own turn and inspect them before spending anything. Sweep. Then a fresh second-opinion turn told to question the premise. Plant one deliberately wrong figure and see whether it corrects it. Mark every claim CONFIRMED / DISAGREEMENT / REPORTED / UNVERIFIED. Agreement is weak evidence; disagreement is the signal.
|
|
50
50
|
|
|
51
51
|
## 7. Numbers and logic are computed, never guessed
|
|
52
52
|
|
|
@@ -13,7 +13,7 @@ Rule: never spend a frontier token on a task a cheap tier finishes correctly. Es
|
|
|
13
13
|
| Lane | Wins at |
|
|
14
14
|
|---|---|
|
|
15
15
|
| the orchestrator (Claude Code, or whichever you chose) | routes, maps, builds, verifies, records; drives the others as CLIs |
|
|
16
|
-
| Codex | second coder and
|
|
16
|
+
| Codex | second coder and second-opinion reviewer: a different model family reading your diff |
|
|
17
17
|
| Antigravity `agy` | deep research sweeps; concurrent fan-out (its subagent call takes an array) |
|
|
18
18
|
| Grok CLI | X and live web reads at $0 (the same search on the API bills per call) |
|
|
19
19
|
| Hermes | the free tier: rough drafts, first-pass summaries, divergent reads, cron jobs |
|
|
@@ -28,7 +28,7 @@ Every agent CLI can report success and deliver nothing. `bin/cli-run.mjs` builds
|
|
|
28
28
|
|
|
29
29
|
Every call goes through it. "This lane is flaky" becomes a query over its log instead of an argument. `node bin/cli-run.mjs --doctor` is the first thing to run after install: enabled lanes, binaries on PATH, the route each lane is pinned to, and with `--run` a one-word canary per lane.
|
|
30
30
|
|
|
31
|
-
There is a second thing a lane can be quietly wrong about. Left unpinned, it runs on **its own config file**, which the runner cannot see: a CLI set up months ago at a low reasoning effort keeps auditing at that effort while your routing docs describe
|
|
31
|
+
There is a second thing a lane can be quietly wrong about. Left unpinned, it runs on **its own config file**, which the runner cannot see: a CLI set up months ago at a low reasoning effort keeps auditing at that effort while your routing docs describe a second-opinion pass, and no error is ever raised. `--model` and `--effort` pin it per call, `defaults` in `bin/lanes.json` pins it per lane, and every run records the value requested and where it came from (`flag`, `lanes.json`, `lane_default`). The log claims no actual: grok reports a model id in its output, the other four lanes report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string the durable log never holds.
|
|
32
32
|
|
|
33
33
|
## 4. Every delegation carries a task bundle, on both surfaces
|
|
34
34
|
|
|
@@ -36,7 +36,7 @@ Subagents and CLI lanes are close to the same problem: something that may hold n
|
|
|
36
36
|
|
|
37
37
|
## 5. Research: three engines, one triager
|
|
38
38
|
|
|
39
|
-
Fan the same plan to three model families (web sweep,
|
|
39
|
+
Fan the same plan to three model families (web sweep, second-opinion read, live data), each as one `cli-run` call. The orchestrator opens the primary sources itself, marks every claim, and writes the only durable record. Expect one engine to return confident unsourced numerics; downgrade it. Weight the engines that report their own gaps. Count dispositions, not briefs.
|
|
40
40
|
|
|
41
41
|
## 5a. A finding is a claim, not a fact
|
|
42
42
|
|
|
@@ -48,7 +48,7 @@ The second pass is now a different model reading the same artifact, in read-only
|
|
|
48
48
|
|
|
49
49
|
## 7. The build protocol, bound to lanes
|
|
50
50
|
|
|
51
|
-
Stage 1 Map: the orchestrator sweeps; CLI lanes critique the map at $0. Stage 2: deep tier,
|
|
51
|
+
Stage 1 Map: the orchestrator sweeps; CLI lanes critique the map at $0. Stage 2: deep tier, one named weak spot and one gap in the request. Stage 4: scanners on the added lines, refuses by default. Stage 5: security-shaped diff → the second coder in read-only audit mode; architecture-shaped → deep tier reviewing build against plan; never both. Two deep checkpoints per build; CLI lanes are uncapped.
|
|
52
52
|
|
|
53
53
|
## 8. Privacy gate
|
|
54
54
|
|
package/llms.txt
CHANGED
|
@@ -24,4 +24,4 @@ Levels: 1 beginner (one agent or chat app), 2 intermediate (several agent CLIs,
|
|
|
24
24
|
## Optional
|
|
25
25
|
|
|
26
26
|
- [Security policy](https://github.com/aunysillyme/model-orchestrator/blob/main/SECURITY.md)
|
|
27
|
-
- [Audit brief](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/audit-brief.md): the
|
|
27
|
+
- [Audit brief](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/audit-brief.md): the security notes and what has already been security-reviewed
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "model-orchestrator",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.18",
|
|
4
4
|
"description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
package/src/catalog.js
CHANGED
|
@@ -87,7 +87,7 @@ export const AIS = [
|
|
|
87
87
|
bin: 'codex',
|
|
88
88
|
access: 'subscription',
|
|
89
89
|
lane: 'A',
|
|
90
|
-
role: 'second coder and
|
|
90
|
+
role: 'second coder and second-opinion reviewer (a different model family reading your diff)',
|
|
91
91
|
minLevel: 1,
|
|
92
92
|
install: { npm: '@openai/codex', pin: '0.153.4' },
|
|
93
93
|
builtAgainst: '0.153.4',
|
package/src/install.js
CHANGED
|
@@ -155,21 +155,21 @@ export function laneVars(selected) {
|
|
|
155
155
|
if (has('hermes')) step0.push(`${cr('hermes')} (the free tier) for rough drafts and divergent reads`);
|
|
156
156
|
if (has('qwen')) step0.push(`${cr('qwen')} (the cheapest metered lane) for structured bulk, never for anything citing a line, number or source`);
|
|
157
157
|
if (has('grok')) step0.push(`${cr('grok')} for X and live web reads at $0`);
|
|
158
|
-
if (has('codex')) step0.push(`${cr('codex --audit')} for
|
|
158
|
+
if (has('codex')) step0.push(`${cr('codex --audit')} for a second-opinion read by a second model family`);
|
|
159
159
|
if (has('agy')) step0.push(`${cr('agy')} for research sweeps and concurrent fan-out`);
|
|
160
160
|
const stage1 = [];
|
|
161
|
-
if (has('codex')) stage1.push(`${cr('codex')} for
|
|
161
|
+
if (has('codex')) stage1.push(`${cr('codex')} for a second-opinion critique of the map`);
|
|
162
162
|
if (has('grok')) stage1.push(`${cr('grok')} to verify current API behaviour instead of trusting recall`);
|
|
163
163
|
if (has('hermes')) stage1.push(`${cr('hermes')} for a divergent read`);
|
|
164
164
|
if (has('agy')) stage1.push(`${cr('agy')} for a wide sweep of prior art`);
|
|
165
165
|
const examples = [];
|
|
166
166
|
examples.push(has('grok') ? `| "What is trending on X today" | ${cr('grok')} |` : '| "What is trending on X today" | live-researcher (standard tier with web tools) |');
|
|
167
|
-
examples.push(has('codex') ? `| "Audit this auth diff" | ${cr('codex --audit')} |` : '| "Audit this auth diff" | code-reviewer at deep tier, in a fresh context told to
|
|
167
|
+
examples.push(has('codex') ? `| "Audit this auth diff" | ${cr('codex --audit')} |` : '| "Audit this auth diff" | code-reviewer at deep tier, in a fresh context told to challenge |');
|
|
168
168
|
examples.push(has('qwen') ? `| "Classify these 200 items" | bulk-worker, or ${cr('qwen')} if the items may leave the machine |` : '| "Classify these 200 items" | bulk-worker |');
|
|
169
|
-
examples.push(enabled.length >= 2 ? '| "Research this topic properly" | several engines in parallel, see `RESEARCH_TRIAGE.md` |' : '| "Research this topic properly" | deep tier plans, standard tier sweeps, a fresh context
|
|
169
|
+
examples.push(enabled.length >= 2 ? '| "Research this topic properly" | several engines in parallel, see `RESEARCH_TRIAGE.md` |' : '| "Research this topic properly" | deep tier plans, standard tier sweeps, a fresh context challenges; see `RESEARCH_TRIAGE.md` |');
|
|
170
170
|
const roles = [];
|
|
171
171
|
if (has('agy')) roles.push('| Web sweep | `cli-run agy` | widest landscape pass |');
|
|
172
|
-
if (has('codex')) roles.push('|
|
|
172
|
+
if (has('codex')) roles.push('| Second-opinion read | `cli-run codex --audit` | question the premise, hunt for what the others would get wrong |');
|
|
173
173
|
if (has('grok')) roles.push('| Live data | `cli-run grok` | dated primary sources, real-time reads |');
|
|
174
174
|
if (has('hermes')) roles.push('| Cheap divergent read | `cli-run hermes` | another opinion at $0 |');
|
|
175
175
|
if (has('qwen')) roles.push('| Structured extraction | `cli-run qwen` | pull the facts into a table; never trust its citations without a check |');
|
|
@@ -183,12 +183,12 @@ export function laneVars(selected) {
|
|
|
183
183
|
return {
|
|
184
184
|
LANE_STEP0: step0.length ? step0.map((l) => ' - ' + l).join('\n') : ' - none selected yet: every task stays on your primary agent\'s tiers until you add a lane (re-run the installer with more AIs)',
|
|
185
185
|
STAGE1_LANES: stage1.length ? '; ' + stage1.join(', ') : '',
|
|
186
|
-
ATTACK_LANE: has('codex') ? '`cli-run codex --audit` (a second model family in a read-only sandbox)' : 'code-reviewer at deep tier, in a fresh context told to
|
|
186
|
+
ATTACK_LANE: has('codex') ? '`cli-run codex --audit` (a second model family in a read-only sandbox)' : 'code-reviewer at deep tier, in a fresh context told to challenge and allowed to answer CLEAN',
|
|
187
187
|
LIVE_LANE: has('grok') ? '`cli-run grok` first ($0), then' : '',
|
|
188
188
|
BULK_LANE: has('qwen') ? ', or `cli-run qwen` if the data may leave your machine' : has('hermes') ? ', or `cli-run hermes` for a free rough pass' : '',
|
|
189
189
|
LANE_EXAMPLES: examples.join('\n'),
|
|
190
190
|
RESEARCH_ROLES: roles.join('\n'),
|
|
191
|
-
RESEARCH_RUN: run.length ? run.join('\n') : '# no cli-run lane selected: run the sweep on your primary agent, then a fresh
|
|
191
|
+
RESEARCH_RUN: run.length ? run.join('\n') : '# no cli-run lane selected: run the sweep on your primary agent, then a fresh second-opinion turn (protocols/deep-research.md, level 1 shape)',
|
|
192
192
|
RESEARCH_ENGINES: String(run.length)
|
|
193
193
|
};
|
|
194
194
|
}
|
|
@@ -405,9 +405,30 @@ function vars(opts) {
|
|
|
405
405
|
const codecalc = tools.some((t) => t.id === 'codecalc');
|
|
406
406
|
const dirAbs = resolve(opts.dir || 'ai-orchestrator');
|
|
407
407
|
const projectAbs = resolve(opts.project || process.cwd());
|
|
408
|
+
// dirPosix backs two things that must read the same on every host:
|
|
409
|
+
// 1. INSTALL_DIR / INSTALL_DIR_SH / INSTALL_DIR_SYSTEMD (below), rendered
|
|
410
|
+
// into vm/jobs/weekly-audit.sh (bash) and vm/jobs/weekly-audit.service
|
|
411
|
+
// (a systemd unit) for the REMOTE Linux box, neither of which can run
|
|
412
|
+
// anywhere but Linux;
|
|
413
|
+
// 2. rulesPath below when --dir falls outside --project, which is
|
|
414
|
+
// prose in generated markdown ("a dir outside the project renders an
|
|
415
|
+
// absolute path"), not a filesystem call.
|
|
416
|
+
// An absolute --dir given as a bare POSIX path ("/opt/x") is never
|
|
417
|
+
// re-resolved through this host's own path semantics for either: a real
|
|
418
|
+
// Windows path always names a drive ("C:\...", caught by the `startsWith`
|
|
419
|
+
// check below falling through to dirAbs), so a bare "/opt/x" only ever
|
|
420
|
+
// means "a Linux path, or documentation text, verbatim" - resolving it
|
|
421
|
+
// with plain path.resolve() reads that leading "/" as drive-relative on
|
|
422
|
+
// win32 and silently turns it into a local path that does not exist,
|
|
423
|
+
// on the box or in the doc. A relative --dir resolves against this
|
|
424
|
+
// host's cwd exactly as before, which is already correct in the common
|
|
425
|
+
// case: level 3 is normally installed by running this CLI ON the box,
|
|
426
|
+
// where "this host" and "the box" are the same filesystem.
|
|
427
|
+
const rawDir = opts.dir || 'ai-orchestrator';
|
|
428
|
+
const dirPosix = rawDir.startsWith('/') ? posix.normalize(rawDir) : dirAbs;
|
|
408
429
|
let rulesPath = relative(projectAbs, dirAbs).split(sep).join(posix.sep);
|
|
409
430
|
if (rulesPath === '') rulesPath = '.';
|
|
410
|
-
else if (rulesPath.startsWith('..')) rulesPath =
|
|
431
|
+
else if (rulesPath.startsWith('..')) rulesPath = dirPosix; // outside the project: absolute is the only honest path
|
|
411
432
|
const pinOf = (id) => (toolById[id] && toolById[id].pin) || 'latest';
|
|
412
433
|
const snippet = snippetFor(primary);
|
|
413
434
|
const steps = activationSteps({ level, selected, primary, tools, dir: opts.dir, project: opts.project });
|
|
@@ -416,7 +437,7 @@ function vars(opts) {
|
|
|
416
437
|
// The path route-gate.mjs and subagent-context.mjs resolve at runtime,
|
|
417
438
|
// relative to CLAUDE_PROJECT_DIR. Mirrors the RULES_PATH fallback below:
|
|
418
439
|
// outside the project, the honest path is absolute, never a hardcoded one.
|
|
419
|
-
const relJoin = (name) => (rulesPath ===
|
|
440
|
+
const relJoin = (name) => (rulesPath === dirPosix ? posix.join(dirPosix, name) : rulesPath === '.' ? name : rulesPath + '/' + name);
|
|
420
441
|
const rulesFileRel = relJoin(routingFile);
|
|
421
442
|
const taskBundleRel = relJoin('TASK_BUNDLE.md');
|
|
422
443
|
// Only claude-code and agy put files under the project root. A chat primary
|
|
@@ -449,9 +470,9 @@ function vars(opts) {
|
|
|
449
470
|
CODECALC_PIN: pinOf('codecalc'),
|
|
450
471
|
OBSIDIAN_TC_PIN: pinOf('obsidian-tc'),
|
|
451
472
|
APIS_LIST: apis.length ? apis.map((prov) => '- ' + prov.name + ' (`' + prov.envName + '`)').join('\n') : '- none: no metered provider key was selected, so the gateway serves only a local lane if you picked one',
|
|
452
|
-
INSTALL_DIR:
|
|
453
|
-
INSTALL_DIR_SH: shellQuote(
|
|
454
|
-
INSTALL_DIR_SYSTEMD: systemdEscape(
|
|
473
|
+
INSTALL_DIR: dirPosix,
|
|
474
|
+
INSTALL_DIR_SH: shellQuote(dirPosix),
|
|
475
|
+
INSTALL_DIR_SYSTEMD: systemdEscape(dirPosix),
|
|
455
476
|
// vm/README.md step 3 named `grok login` and `agy` whatever you picked (#26).
|
|
456
477
|
VM_SIGNIN: (() => {
|
|
457
478
|
const lines = selected.filter((a) => a.bin && a.kind === 'agent-cli').map((a) => ` - ${a.name}: ${a.auth}`);
|
|
@@ -784,8 +805,14 @@ export function writeFiles(files, opts) {
|
|
|
784
805
|
for (const f of groups[k]) {
|
|
785
806
|
const abs = resolve(root, f.rel);
|
|
786
807
|
const exists = existsSync(abs);
|
|
787
|
-
|
|
788
|
-
|
|
808
|
+
// label is what reaches the terminal report (bin/cli.js's "runtime
|
|
809
|
+
// upgraded:", "runtime CONFLICT, kept:", etc lines): posix-normalized
|
|
810
|
+
// like key, below, so the report reads the same on every host. Before
|
|
811
|
+
// this it carried f.rel verbatim, which is native-separated (join()),
|
|
812
|
+
// so on win32 the report named "bin\cli-run.mjs" while everything
|
|
813
|
+
// else in this tool (docs, other path prose) uses forward slashes.
|
|
814
|
+
const label = (k === 'project' ? '[project] ' : '') + f.rel.split(sep).join('/');
|
|
815
|
+
const key = label;
|
|
789
816
|
const cls = k === 'dir' ? fileClass(f.rel) : 'document';
|
|
790
817
|
if (exists && !force) {
|
|
791
818
|
if (cls === 'document') {
|
|
@@ -31,17 +31,27 @@ find reports -maxdepth 1 -name '.audit-*' -type f -mtime +0 -delete 2>/dev/null
|
|
|
31
31
|
# pipe open). Process groups do not help here: bash disables job control inside
|
|
32
32
|
# pipeline subshells, so `kill -- -pid` would kill nothing. A recursive tree
|
|
33
33
|
# kill via pgrep works on macOS and Linux alike; `timeout(1)` is not on macOS.
|
|
34
|
+
#
|
|
35
|
+
# Two things make a timeout always read as a timeout. killtree freezes each
|
|
36
|
+
# process (SIGSTOP) before walking its children, so a parent cannot run on,
|
|
37
|
+
# print, and exit 0 in the gap after its child dies. And the watchdog leaves a
|
|
38
|
+
# marker when it fires, so bounded returns 124 whatever order the kills land
|
|
39
|
+
# in. Without both, a hung `--version` probe could be recorded as a version
|
|
40
|
+
# string instead of "UNVERIFIED: timed out" (seen once on macOS CI).
|
|
34
41
|
killtree() {
|
|
35
42
|
local p="$1" c
|
|
43
|
+
kill -STOP "$p" 2>/dev/null
|
|
36
44
|
for c in $(pgrep -P "$p" 2>/dev/null); do killtree "$c"; done
|
|
37
45
|
kill -KILL "$p" 2>/dev/null
|
|
38
46
|
}
|
|
39
47
|
bounded() {
|
|
40
48
|
local secs="$1"; shift
|
|
49
|
+
local fired; fired="$(mktemp "${TMPDIR:-/tmp}/wa-fired.XXXXXX" 2>/dev/null)" && rm -f "$fired"
|
|
41
50
|
( "$@" ) & local pid=$!
|
|
42
|
-
( sleep "$secs"; killtree "$pid" ) >/dev/null 2>&1 & local wd=$!
|
|
51
|
+
( sleep "$secs"; [ -n "$fired" ] && : > "$fired"; killtree "$pid" ) >/dev/null 2>&1 & local wd=$!
|
|
43
52
|
wait "$pid" 2>/dev/null; local rc=$?
|
|
44
53
|
killtree "$wd" >/dev/null 2>&1; wait "$wd" 2>/dev/null
|
|
54
|
+
if [ -n "$fired" ] && [ -e "$fired" ]; then rm -f "$fired"; return 124; fi
|
|
45
55
|
return $rc
|
|
46
56
|
}
|
|
47
57
|
|
|
@@ -2,4 +2,4 @@
|
|
|
2
2
|
|
|
3
3
|
Antigravity CLI custom agents, one per tier plus three checks (`finding-verifier`, `done-verifier`, `reader`), in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
|
|
4
4
|
|
|
5
|
-
`commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests;
|
|
5
|
+
`commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; deletes and other destructive commands still ask before running) and `off` for the read-only agents: `code-reviewer`, `finding-verifier`, `live-researcher`, `done-verifier`, `reader`. `model` is a tier: `pro` for deep-planner, `flash` for the rest. `done-verifier` and `reader` never write and, with `commandExecutionPolicy: off`, cannot execute any command at all here, mutating or not: unlike its claude-code counterpart, which does carry an unrestricted `Bash` and stays read-only by its prompt rather than by the tool grant, agy's `done-verifier` is mechanically blocked from shelling out and probes artifacts through whatever read or fetch capability it has instead. Neither is `bulk-worker`, which classifies, tags and transforms items and does write.
|
|
@@ -4,7 +4,7 @@ description: Well-specified execution of a bounded sub-part of a build.
|
|
|
4
4
|
model: flash
|
|
5
5
|
subagent: true
|
|
6
6
|
mainAgent: true
|
|
7
|
-
commandExecutionPolicy: auto # standard build/test commands run unattended;
|
|
7
|
+
commandExecutionPolicy: auto # standard build/test commands run unattended; destructive commands, like deletes, still ask before running
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# builder
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: finding-verifier
|
|
3
|
-
description:
|
|
3
|
+
description: Second-opinion verification of review findings; tries to disprove each one and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE. Read-only, never repairs.
|
|
4
4
|
model: flash
|
|
5
5
|
subagent: true
|
|
6
6
|
mainAgent: true
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: finding-verifier
|
|
3
|
-
description:
|
|
3
|
+
description: Second-opinion verification of review findings. Use after a review or audit returns findings and before any of them trigger a repair. No file-editing tools; Bash is for read-only checks, bound by the prompt below, not by the tool grant. Tries to DISPROVE each finding and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Do not use to find new problems, and do not use to fix anything.
|
|
4
4
|
tools: Read, Glob, Grep, Bash
|
|
5
5
|
model: sonnet
|
|
6
6
|
effort: high
|
|
@@ -4,8 +4,8 @@
|
|
|
4
4
|
|
|
5
5
|
```
|
|
6
6
|
You follow a model-orchestrator workflow inside this chat. Tiers describe effort, not automatic model switching or cost savings.
|
|
7
|
-
Route first: bulk/formatting -> fast; live data -> standard with tools; review -> standard, read-only; ambiguous/high-
|
|
8
|
-
For builds: map affected parts; identify
|
|
7
|
+
Route first: bulk/formatting -> fast; live data -> standard with tools; review -> standard, read-only; ambiguous/high-stakes -> deep; otherwise standard. State the tier. Escalate on failure instead of silently retrying.
|
|
8
|
+
For builds: map affected parts; identify what is most likely to go wrong and any gap in the request (none is valid with reasons); build and verify; use a fresh turn to challenge the result. Before irreversible actions, explain rollback and ask for approval.
|
|
9
9
|
For hand-offs: include purpose, scope, allowed and denied actions, required output, and stopping conditions. A fresh context has none of these instructions.
|
|
10
10
|
After comprehensive work, check for omissions. Compute consequential numbers and comparisons with a tool; report what was checked and what remains unverified.
|
|
11
11
|
Before durable writes, search existing records, update their index, use one writer, and label inferences.
|
|
@@ -17,7 +17,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
|
|
|
17
17
|
|
|
18
18
|
A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Measure yours once: spawn a subagent with a one-line task and read its token count. Work smaller than that stays inline.
|
|
19
19
|
|
|
20
|
-
Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one
|
|
20
|
+
Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one challenge pass, an explicit human yes before anything irreversible, then the loud negative.
|
|
21
21
|
|
|
22
22
|
Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A Claude Code subagent loads this CLAUDE.md hierarchy, so it holds the standing rules already, just not this task's scope; a second CLI or a fresh chat window may hold none of them. Absence is denial either way.
|
|
23
23
|
|
|
@@ -9,7 +9,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
|
|
|
9
9
|
|
|
10
10
|
Route by capability tier, first match wins: bulk and mechanical -> fast tier · needs live data -> standard tier with tools · review without changing -> standard, read-only · ambiguous or expensive to get wrong -> deep tier, then hand the plan down · everything else -> build it directly at standard tier.
|
|
11
11
|
|
|
12
|
-
Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: map
|
|
12
|
+
Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: map everything it touches yourself, ask the deep tier for one named weak spot and one gap in the request, build green, scan the added lines, one challenge pass with every finding reproduced, an explicit human yes before anything irreversible, then re-grep the old identifier and expect zero.
|
|
13
13
|
|
|
14
14
|
Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A fresh context holds none of these rules; absence is denial.
|
|
15
15
|
|
|
@@ -10,7 +10,7 @@
|
|
|
10
10
|
// This script always exits 0, never blocks on stdin past a short bound,
|
|
11
11
|
// reads at most 64 KB of the rules file through a fixed-size buffer (never
|
|
12
12
|
// a full read of an arbitrarily large or non-regular file), and never
|
|
13
|
-
// executes anything it reads. See docs/audit-brief.md for the
|
|
13
|
+
// executes anything it reads. See docs/audit-brief.md for the security notes.
|
|
14
14
|
import { statSync, openSync, readSync, closeSync, realpathSync } from 'node:fs';
|
|
15
15
|
import { join, isAbsolute } from 'node:path';
|
|
16
16
|
|
|
@@ -29,8 +29,8 @@ Modifiers:
|
|
|
29
29
|
|
|
30
30
|
## The two checkpoints (every build)
|
|
31
31
|
|
|
32
|
-
- **Checkpoint 1, before writing anything.** You map
|
|
33
|
-
- **Checkpoint 2, after the build is green.** Security-shaped diffs get
|
|
32
|
+
- **Checkpoint 1, before writing anything.** You map everything it touches yourself (files, systems, docs, tickets). Then ask the deep tier, on the finished map: *is this the simplest way, what is most likely to go wrong, what did the request miss?* It must return one named weak spot and one gap in the request. Approval alone is not an answer.
|
|
33
|
+
- **Checkpoint 2, after the build is green.** Security-shaped diffs get a second-opinion read (in a fresh context, told to challenge, allowed to answer CLEAN). Architecture-shaped diffs get the deep tier reviewing build against plan. Never both on one diff. Every finding reproduced before it reaches a human.
|
|
34
34
|
|
|
35
35
|
Cap: two deep-tier consults per build. The full procedure is `protocols/build-protocol.md`.
|
|
36
36
|
|
|
@@ -47,7 +47,7 @@ Level 2 adds `ROUTING.md`, `TIERS.md`, `DELEGATION_MATRIX.md`, `RESEARCH_TRIAGE.
|
|
|
47
47
|
## The three rules that carry everything
|
|
48
48
|
|
|
49
49
|
1. **Route by capability tier, not by model name.** deep = ambiguous or expensive to get wrong · standard = well-specified execution and review · fast = bulk and mechanical. Default down, escalate on evidence.
|
|
50
|
-
2. **A gate you cannot fail is not a gate.** "Does it look good?" passes every time. "Name
|
|
50
|
+
2. **A gate you cannot fail is not a gate.** "Does it look good?" passes every time. "Name what is most likely to go wrong, and what the request missed" can come back empty, which is how you know it worked.
|
|
51
51
|
3. **Exit 0 is not a deliverable.** Any tool, CLI or subagent can report success and hand back nothing. Check for the artifact, not the status line.
|
|
52
52
|
|
|
53
53
|
## Where things went
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
**Three phases, eight stages, and every gate is a question that can be answered wrong.**
|
|
4
4
|
|
|
5
|
-
Fires on any task that builds, codes, implements, migrates or deploys. Rough test: if it would earn
|
|
5
|
+
Fires on any task that builds, codes, implements, migrates or deploys. Rough test: if it would earn a second-opinion audit or a tracker issue, it runs this.
|
|
6
6
|
|
|
7
7
|
> **The one rule underneath:** a gate you cannot fail is not a gate. If a stage's exit reads like "confirm it looks good", it is written wrong and it will pass every time, including the times it should not.
|
|
8
8
|
|
|
@@ -14,7 +14,7 @@ Three corollaries:
|
|
|
14
14
|
| Phase | Master question | Stages |
|
|
15
15
|
|---|---|---|
|
|
16
16
|
| 1 Pre-build | What exactly are we building, what do we need first, and what does this touch or break? | 0 Route · 1 Map · 2 Judge |
|
|
17
|
-
| 2 Build | Is it secure, built on current code, and correct without hidden flaws? | 3 Build · 4 Scan · 5
|
|
17
|
+
| 2 Build | Is it secure, built on current code, and correct without hidden flaws? | 3 Build · 4 Scan · 5 Challenge · 5b Ship gate |
|
|
18
18
|
| 3 Post-build | Did it land everywhere, is it proven against the real thing, and is it recorded? | 6 Verify · 7 Record |
|
|
19
19
|
|
|
20
20
|
The two seams are the point. Pre-build to Build: nothing is written yet, changing your mind costs a conversation. Build to Post-build: the ship, the only irreversible step, the only one that needs an explicit human yes.
|
|
@@ -38,9 +38,9 @@ Four bounded questions, not four exhaustive scans. **The builder maps; the judgm
|
|
|
38
38
|
### Stage 2 · Judge (Checkpoint 1)
|
|
39
39
|
Ask the judgment tier, on the finished map:
|
|
40
40
|
1. Is this the simplest way to build it, or are we overcomplicating?
|
|
41
|
-
2. What is
|
|
41
|
+
2. What is most likely to go wrong, and what did the request miss?
|
|
42
42
|
|
|
43
|
-
**Gate:**
|
|
43
|
+
**Gate:** **one named weak spot** and **one gap in the request**. Approval alone is not an exit; an advisor asked only to approve will approve. If a consult comes back mostly restating the map, the brief asked it to retrieve when it should have asked it to decide.
|
|
44
44
|
|
|
45
45
|
## Phase 2 · Build
|
|
46
46
|
|
|
@@ -56,15 +56,15 @@ Ask the judgment tier, on the finished map:
|
|
|
56
56
|
1. Any secret, key or token in the new code?
|
|
57
57
|
2. Any vulnerability or vulnerable dependency in the lines we added?
|
|
58
58
|
|
|
59
|
-
Secret detection, static analysis and dependency scanning, filtered to lines this diff added.
|
|
59
|
+
Secret detection, static analysis and dependency scanning, filtered to lines this diff added. Refuses by default: a missing or erroring scanner exits non-zero, never a silent green.
|
|
60
60
|
|
|
61
61
|
**Gate:** zero flags on added lines. Pre-existing flags are reported, never inherited as blockers, and never waved through unread. A scanner finding is a claim; read the code before calling it anything.
|
|
62
62
|
|
|
63
|
-
### Stage 5 ·
|
|
63
|
+
### Stage 5 · Challenge (Checkpoint 2, one pass, never two)
|
|
64
64
|
1. Can bad input or a bad actor break it, and what happens when a dependency fails?
|
|
65
65
|
2. Did the build stick to the approved plan, or did unintended changes sneak in?
|
|
66
66
|
|
|
67
|
-
Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to
|
|
67
|
+
Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to a second-opinion reviewer, ideally a **different model family**. Architecture-shaped diffs go to the judgment tier reviewing build against plan. Never both on one diff.
|
|
68
68
|
|
|
69
69
|
Findings do not go straight to a repair. Hand them to the **finding-verifier**, whose job is to DISPROVE each one: read the cited line, state what would trigger it, then hunt for the guard, caller or test that makes it impossible. It returns one of three verdicts per finding, and rounding between them is the failure mode to watch for.
|
|
70
70
|
|
|
@@ -106,7 +106,7 @@ Use a different model family from the one that produced the finding where you ha
|
|
|
106
106
|
|---|---|---|
|
|
107
107
|
{{ROLES_BUILDER_ROW}}
|
|
108
108
|
| Judgment tier | Stage 2 and the architectural arm of Stage 5. Argues with a finished map | Perform the retrieval |
|
|
109
|
-
|
|
|
109
|
+
| Second-opinion reviewer | The security arm of Stage 5. Reviews the diff | Fix anything |
|
|
110
110
|
| Mechanical gates | Stage 4 and any always-on guard | Be overridden without reading |
|
|
111
111
|
| Cheap workers | Bounded sub-parts: bulk passes, wide searches, long loops | Own a stage |
|
|
112
112
|
| Human | Stage 5b, and any irreversible or architectural call | Be the first line of review |
|
|
@@ -119,9 +119,9 @@ Use a different model family from the one that produced the finding where you ha
|
|
|
119
119
|
PRE-BUILD
|
|
120
120
|
[ ] 0 Inputs and access verified by live probe, not assumed
|
|
121
121
|
[ ] 0 Confirmed this is a build and not a quick fix
|
|
122
|
-
[ ] 1
|
|
122
|
+
[ ] 1 Everything it touches written: files, systems, issues
|
|
123
123
|
[ ] 1 Asked what could break, and whether this already exists
|
|
124
|
-
[ ] 2 Judgment tier named a
|
|
124
|
+
[ ] 2 Judgment tier named a weak spot AND a gap in the request
|
|
125
125
|
|
|
126
126
|
BUILD
|
|
127
127
|
[ ] 3 Repo clean, on a branch, base ref recorded
|
|
@@ -31,11 +31,11 @@ Agreement is weak evidence. Disagreement is the signal.
|
|
|
31
31
|
|
|
32
32
|
## Level 1: one agent
|
|
33
33
|
|
|
34
|
-
You still get the shape. Run PLAN as its own turn and inspect it before spending anything. Run the sweep. Then run a **fresh-context
|
|
34
|
+
You still get the shape. Run PLAN as its own turn and inspect it before spending anything. Run the sweep. Then run a **fresh-context second-opinion turn** with a brief that says "question the premise; list what this report would get wrong if its sources were stale". Plant one deliberately wrong figure in the brief and see whether it corrects it: if it does not, its confirmations are worth less than they look. Mark every claim.
|
|
35
35
|
|
|
36
36
|
## Level 2 and up: three engines, one triager
|
|
37
37
|
|
|
38
|
-
Fan out the same PLAN to three different model families through their CLIs (a web-sweep lane,
|
|
38
|
+
Fan out the same PLAN to three different model families through their CLIs (a web-sweep lane, a second-opinion-read lane, a live-data lane). Run them through `cli-run` so a run that produced nothing is caught as `rc=10` rather than read as an empty finding. The orchestrator triages: it opens the primary sources itself, marks each claim, and writes the brief. Only the orchestrator writes the durable record; every other engine proposes.
|
|
39
39
|
|
|
40
40
|
Known failure shape: one engine will return confident unsourced numerics and claim full coverage. Downgrade those to hypothesis. The engines that report their own gaps honestly are the ones to weight.
|
|
41
41
|
|
|
@@ -14,7 +14,7 @@ Verification asks "is what I did correct?". Gap analysis asks "what did I not do
|
|
|
14
14
|
## Who runs it
|
|
15
15
|
|
|
16
16
|
- **Level 1 (one agent):** the same agent, in a fresh turn, with a brief that says "you are looking for what is missing; do not re-verify what is present". Fresh context matters more than a different model.
|
|
17
|
-
- **Level 2 and up:** a **different model family** reading the same artifact. Disagreement between two families is the cheapest available signal that something is soft. The
|
|
17
|
+
- **Level 2 and up:** a **different model family** reading the same artifact. Disagreement between two families is the cheapest available signal that something is soft. The second-opinion coder lane (read-only mode) is the natural fit.
|
|
18
18
|
- **Level 3:** make it recurring. A weekly audit job enumerates live state (lanes, jobs, services, model lists), diffs it against the plan, and files a report. It catches the dead lane and the silently renamed model nobody noticed.
|
|
19
19
|
|
|
20
20
|
## The second half: analyze, compare, suggest
|
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
# Propagate: change completeness
|
|
2
2
|
|
|
3
|
-
**A rename is a refactor, not a single-file edit.** Any change to a name, term, path, slug, schema field, routing rule or shared convention
|
|
3
|
+
**A rename is a refactor, not a single-file edit.** Any change to a name, term, path, slug, schema field, routing rule or shared convention reaches everything that uses it, and the goal is zero silent strays.
|
|
4
4
|
|
|
5
5
|
This is retrieval work. It stays with the orchestrator (or a cheap worker for the grep sweep). It never goes to the deep tier: a judgment model re-deriving a file list is the most expensive routing mistake there is.
|
|
6
6
|
|
|
7
|
-
## 1. Map
|
|
7
|
+
## 1. Map everything it touches (before editing anything)
|
|
8
8
|
|
|
9
9
|
- **Docs and notes:** backlinks to the thing being renamed; literal search for the old term and its link forms. With obsidian-tc: `get_backlinks`, `search_text`, then `find_unresolved_links` after the change (`protocols/memory-and-record.md`).
|
|
10
10
|
- **Memory / instructions:** grep every instructions file your agents read (`CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, `QWEN.md`, custom instructions) and any memory store.
|
|
@@ -68,7 +68,7 @@ qwen is the lane whose own success flags lie: an upstream 400 comes back as exit
|
|
|
68
68
|
|
|
69
69
|
## The route: which model, and how hard it thinks
|
|
70
70
|
|
|
71
|
-
A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe
|
|
71
|
+
A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe a second-opinion pass, and nothing anywhere says so.
|
|
72
72
|
|
|
73
73
|
Pin it per call, or per lane:
|
|
74
74
|
|
|
@@ -114,9 +114,9 @@ The log records what was **requested**, on every record including a run refused
|
|
|
114
114
|
|
|
115
115
|
That is each vendor's documented headless shape (`-p`, `exec`). Two consequences: argv is visible to other processes on the machine, so a prompt is never the place for a key; and argv is bounded by the OS (`ARG_MAX`), so a very large brief should be referenced by path inside the prompt rather than pasted whole.
|
|
116
116
|
|
|
117
|
-
## lanes.json
|
|
117
|
+
## lanes.json refuses by default
|
|
118
118
|
|
|
119
|
-
Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all
|
|
119
|
+
Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all make the whole file refuse by default rather than being skipped quietly.
|
|
120
120
|
|
|
121
121
|
## A killed lane is not a deliverable
|
|
122
122
|
|
|
@@ -15,7 +15,7 @@ Generated {{DATE}} from the AIs you said you have: `{{AI_IDS}}`.
|
|
|
15
15
|
| Many independent items each needing its own agent turn | a concurrent fan-out lane | one call, N children, on a subscription |
|
|
16
16
|
| Live web or social reads | the live-data CLI | subscription-covered; the same search on the API bills per call |
|
|
17
17
|
| Code review, no changes | standard tier, or the second-coder CLI | a different model family catches what one misses |
|
|
18
|
-
|
|
|
18
|
+
| Second-opinion audit of a security-shaped diff | the second-coder CLI in read-only audit mode | Claude writes, a second family challenges, the orchestrator reproduces |
|
|
19
19
|
| Deep architecture / planning | deep tier | expensive to get wrong |
|
|
20
20
|
| Well-specified execution | the orchestrator | execution does not need the top tier |
|
|
21
21
|
| Long-document analysis | the largest-context lane, or caching on the primary | window size vs re-query cost |
|
|
@@ -31,10 +31,10 @@ Rule of thumb: never spend a frontier token on a task a cheap tier finishes corr
|
|
|
31
31
|
|---|---|
|
|
32
32
|
| 0 Route | live probe for access; `cli-run` lanes are $0 and uncapped |
|
|
33
33
|
| 1 Map | the orchestrator sweeps{{STAGE1_LANES}} |
|
|
34
|
-
| 2 Judge | deep tier, on the finished map:
|
|
34
|
+
| 2 Judge | deep tier, on the finished map: one named weak spot and one gap in the request |
|
|
35
35
|
| 3 Build | the orchestrator, against the installed dependency's source |
|
|
36
|
-
| 4 Scan | secret + static + dependency scanners, diff-scoped,
|
|
37
|
-
| 5
|
|
36
|
+
| 4 Scan | secret + static + dependency scanners, diff-scoped, refuses by default |
|
|
37
|
+
| 5 Challenge | security-shaped diff → {{ATTACK_LANE}}. Architecture-shaped → deep tier, build against plan. Never both |
|
|
38
38
|
| 5a Verify findings | finding-verifier, a different model family where you have one: CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a repair |
|
|
39
39
|
| 5b Ship | rollback id recorded, explicit human yes |
|
|
40
40
|
| 6 Verify | real test, negative test seen red, old identifier re-grepped to zero |
|
|
@@ -58,7 +58,7 @@ One writer per run; every other lane proposes. Search before writing, index in t
|
|
|
58
58
|
- **Long context:** mechanical digestion → fast tier in chunks; judgment over a long input → standard tier.
|
|
59
59
|
- **Token discipline on every delegation:** pass only the context the delegate needs, never the conversation.
|
|
60
60
|
- **Effort per agent:** deep xhigh, review, verification and build high, live research medium, bulk low.
|
|
61
|
-
- **Three inputs, not one:** role picks the agent, complexity moves the effort,
|
|
61
|
+
- **Three inputs, not one:** role picks the agent, complexity moves the effort, stakes move the tier and who reads it. A one-line auth change is simple and high-stakes at once, and the stakes decide. See `TIERS.md`.
|
|
62
62
|
- **Pin the route when it matters:** a lane with no `--model`/`--effort` and no `defaults` entry in `bin/lanes.json` runs on its own config, which may be nothing like what this file describes. `cli-run --doctor` prints what each lane is pinned to, and every run logs the value requested and where it came from.
|
|
63
63
|
|
|
64
64
|
## Example routings
|
|
@@ -47,21 +47,25 @@ reasoning than the reviewer judging its output.** When the plan is airtight the
|
|
|
47
47
|
spec is carrying the thinking, so builder drops to medium. When the plan is
|
|
48
48
|
vague, fix the plan; do not buy reasoning to paper over it.
|
|
49
49
|
|
|
50
|
-
**
|
|
50
|
+
**Stakes move the tier and the reader, never just the effort.** These four are
|
|
51
51
|
the ones worth naming, because their failures are not recoverable by editing the
|
|
52
52
|
code afterwards.
|
|
53
53
|
|
|
54
|
-
|
|
54
|
+
Stakes means what a mistake would cost: a security hole, leaked personal data,
|
|
55
|
+
lost data, or something you can't undo. Most tasks are low-stakes and route
|
|
56
|
+
normally.
|
|
57
|
+
|
|
58
|
+
| Stakes | Present when the change touches | What it buys |
|
|
55
59
|
|---|---|---|
|
|
56
|
-
| security | auth, tokens, sessions, routes, untrusted input | the
|
|
60
|
+
| security | auth, tokens, sessions, routes, untrusted input | the challenge pass, ideally a different model family |
|
|
57
61
|
| privacy | personal data, anything leaving the machine | the local lane, and a named check on what is sent |
|
|
58
62
|
| data loss | deletion, bulk mutation, migrations, overwrites | a reviewed rollback path before the change is written |
|
|
59
63
|
| irreversible | publishing, sending, rotating, anything with an audience | a human yes at Stage 5b, never an agent's |
|
|
60
64
|
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
synonym for difficulty: a one-line change to an auth check is simple and
|
|
64
|
-
high-
|
|
65
|
+
High stakes raise code-reviewer to xhigh, and a security-shaped diff goes to
|
|
66
|
+
the challenge lane rather than to a second read by the same family. Stakes are
|
|
67
|
+
not a synonym for difficulty: a one-line change to an auth check is simple and
|
|
68
|
+
high-stakes at the same time, and it is the stakes that decide the route.
|
|
65
69
|
|
|
66
70
|
**Reserve the top of the ladder for evidence.** xhigh and the escalation tier are
|
|
67
71
|
bought with a named reason: a reproduced failure, a checkpoint that came back
|
|
@@ -70,7 +74,7 @@ task, not an escalation.
|
|
|
70
74
|
|
|
71
75
|
## Why split tiers: robustness first, cost second
|
|
72
76
|
|
|
73
|
-
The split produces better work. The deep tier steers every build twice, and what it steers is **judgment, never retrieval**: the orchestrator sweeps
|
|
77
|
+
The split produces better work. The deep tier steers every build twice, and what it steers is **judgment, never retrieval**: the orchestrator sweeps everything it touches itself and hands the deep tier a finished map. Paying deep-tier rates for a file list is the most expensive routing mistake available.
|
|
74
78
|
|
|
75
79
|
Against a baseline of "standard tier with no consults", default checkpoints are a spend increase. That is the accepted trade, not a saving to claim.
|
|
76
80
|
|
|
@@ -36,7 +36,7 @@ The tools cannot help a model that never reaches for them. codecalc ships `SKILL
|
|
|
36
36
|
|
|
37
37
|
## What it is not
|
|
38
38
|
|
|
39
|
-
Not a cloud sandbox for multi-tenant loads, not a replacement for a vendor's built-in interpreter when zero setup matters more than measurement.
|
|
39
|
+
Not a cloud sandbox for multi-tenant loads, not a replacement for a vendor's built-in interpreter when zero setup matters more than measurement. It assumes a single-operator, local, stdio setup. It earns its keep when the correctness of a claim, not "it ran", is the point.
|
|
40
40
|
|
|
41
41
|
## On a box (level 3)
|
|
42
42
|
|
|
@@ -11,7 +11,7 @@ A durable, searchable, governed store that the protocols can call by name:
|
|
|
11
11
|
| Need in the protocols | obsidian-tc tool |
|
|
12
12
|
|---|---|
|
|
13
13
|
| find what exists before writing (deep research dedupe, gap analysis) | `semantic_search`, `search_text`, `search_regex` |
|
|
14
|
-
| map a rename
|
|
14
|
+
| map everything a rename touches (propagate) | `get_backlinks`, `find_unresolved_links`, `rewrite_link` |
|
|
15
15
|
| record the end-to-end doc (build Stage 7) | `write_note` (compare-and-swap, confirmation on overwrite), `patch_note`, `append_note` |
|
|
16
16
|
| keep inferred content honest | `write_note` with `provenance: "agent_synthesis"` runs a poison scan before the write lands |
|
|
17
17
|
| keep a shared vault safe for several agents | JWT scopes, per-vault folder ACLs, a read-only kill switch, human-in-the-loop tokens |
|
|
@@ -56,9 +56,9 @@ Merge the block; do not replace the file.
|
|
|
56
56
|
|
|
57
57
|
## Security posture, read before a second agent touches it
|
|
58
58
|
|
|
59
|
-
Zero-config mode boots with **auth off and no folder ACL**: anything that can reach the server has the same authority as raw filesystem access to the vault. That is acceptable only because the surface is local-only (the config
|
|
59
|
+
Zero-config mode boots with **auth off and no folder ACL**: anything that can reach the server has the same authority as raw filesystem access to the vault. That is acceptable only because the surface is local-only (the config refuses by default if you enable HTTP on a non-loopback host with auth off, and a DNS-rebinding guard protects loopback). Before exposing it to partially-trusted, remote or multi-agent callers, turn on `auth.mode: "jwt"` and set `acl.readPaths` / `writePaths` / `deletePaths` in the config file. Upstream `SECURITY.md` has the security notes and a private disclosure path.
|
|
60
60
|
|
|
61
|
-
Track record worth knowing: an independent code audit of v1.8.1 (July 2026) found three security-relevant gaps (an ACL
|
|
61
|
+
Track record worth knowing: an independent code audit of v1.8.1 (July 2026) found three security-relevant gaps (an ACL bypass that let enumeration tools skip its refuse-by-default rule, a compare-and-swap bypass through `upsert`, a poison-eligibility gap in preference extraction). All three were fixed upstream before they were filed; verified against the v1.25.0 source on 2026-09-03.
|
|
62
62
|
|
|
63
63
|
## Level 3
|
|
64
64
|
|