model-orchestrator 0.1.14 → 0.1.16

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (36) hide show
  1. package/AGENTS.md +26 -0
  2. package/CHANGELOG.md +43 -2
  3. package/README.md +52 -8
  4. package/bin/cli-run.mjs +22 -7
  5. package/bin/cli.js +4 -4
  6. package/docs/audit-brief.md +32 -0
  7. package/docs/part-1-beginner.md +1 -1
  8. package/docs/part-2-intermediate.md +1 -1
  9. package/llms.txt +27 -0
  10. package/package.json +28 -4
  11. package/src/README.md +1 -1
  12. package/src/catalog.js +9 -1
  13. package/src/detect.js +28 -9
  14. package/src/install.js +206 -5
  15. package/templates/README.md +1 -1
  16. package/templates/agents/README.md +3 -1
  17. package/templates/agents/agy/README.md +2 -2
  18. package/templates/agents/agy/done-verifier.md +35 -0
  19. package/templates/agents/agy/finding-verifier.md +6 -0
  20. package/templates/agents/agy/reader.md +22 -0
  21. package/templates/agents/claude-code/README.md +7 -5
  22. package/templates/agents/claude-code/builder.md +6 -1
  23. package/templates/agents/claude-code/code-reviewer.md +9 -2
  24. package/templates/agents/claude-code/done-verifier.md +44 -0
  25. package/templates/agents/claude-code/finding-verifier.md +9 -2
  26. package/templates/agents/claude-code/reader.md +26 -0
  27. package/templates/agents/snippets/claude-code.md +9 -4
  28. package/templates/agents/snippets/route-gate.mjs +151 -0
  29. package/templates/agents/snippets/route-metrics.mjs +356 -0
  30. package/templates/agents/snippets/settings.hooks.snippet.json +70 -0
  31. package/templates/agents/snippets/subagent-context.mjs +76 -0
  32. package/templates/beginner/ORCHESTRATOR.md +4 -3
  33. package/templates/common/TASK_BUNDLE.md +2 -2
  34. package/templates/common/protocols/build-protocol.md +2 -2
  35. package/templates/intermediate/ROUTING.md +9 -10
  36. package/templates/intermediate/TIERS.md +2 -0
package/AGENTS.md ADDED
@@ -0,0 +1,26 @@
1
+ # AGENTS.md
2
+
3
+ Two audiences: an agent that wants to USE this package for a project, and an agent that is working ON this repository.
4
+
5
+ ## Using this package from an agent
6
+
7
+ model-orchestrator writes routing rules, subagents and a CLI lane runner so an agent sends each task to the right model, subagent or CLI and spends fewer frontier tokens. It is not a proxy or gateway. Headless use:
8
+
9
+ - `npx model-orchestrator --list` prints every supported AI id.
10
+ - `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project <repo> --dir <repo>/ai-orchestrator --dry-run` prints the plan and writes nothing.
11
+ - Drop `--dry-run` to write it. Existing files are never overwritten without `--force`; activation snippets (for example `CLAUDE.snippet.md`, `settings.hooks.snippet.json`) are written for a person or agent to merge.
12
+ - The generated `README.md` in `--dir` lists what to copy where and one smoke command to prove the rules took.
13
+ - A summary for LLMs, with links to every doc: [`llms.txt`](llms.txt).
14
+
15
+ ## Working on this repository
16
+
17
+ The files the installer writes for end users live under `templates/`.
18
+
19
+ - Read `CONTRIBUTING.md` first, then `src/README.md` (the catalog drives everything) and `docs/audit-brief.md` (the threat model and what has already been attacked).
20
+ - Run `npm test` before proposing a change and quote the count and the exit code; the suite prints the current number.
21
+ - Everything renders from `src/catalog.js`. Add an AI or a tool there, not in a template. Templates carry no logic.
22
+ - Never put a value that looks like a credential anywhere in this repo, including tests and examples. Environment variable names only.
23
+ - `bin/cli.js` writes only inside `--dir` and `--project`, never over a document without `--force`, and never runs a vendor script. A change that weakens any of those will be refused in review; the tests that hold them are in `test/install.test.js` and `test/cli.test.js`.
24
+ - `bin/cli-run.mjs` must exit non-zero when a lane produced nothing. Every judge has a red case in `test/judges.test.js`; add one before you change a judge.
25
+ - Prose in this repo uses no em dashes (`test/prose.test.js` enforces it).
26
+ - Why these rules exist: each one is the fix for a failure that reached an audit or CI. `CHANGELOG.md` names the issue behind each.
package/CHANGELOG.md CHANGED
@@ -4,7 +4,46 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
- ## [0.1.14] - 2026-09-09
7
+ ## [0.1.16] - 2026-09-11
8
+
9
+ ### Added
10
+
11
+ - **A third claude-code-only hook, `route-metrics.mjs`, answers "is my agent actually routing and delegating?"** A routing rule nobody measures is a rule nobody knows is followed. Wired to five events (`UserPromptSubmit`, `PreToolUse` on `Agent`/`Task`, `SubagentStart`, `SubagentStop`, `Stop`), it appends one JSON line per event to `~/.ai-orchestrator/route-metrics.jsonl` (the same directory and home resolution `bin/cli-run.mjs` already logs to): a turn, a subagent dispatch (`subagent_type`, background flag), a subagent start and stop (so a duration can be computed from a small state file keyed by `sha256(agent_id)`), and the lane parsed from a new hidden marker, `<!-- route: <lane> | <why> -->`, that the route-gate block now asks every reply to end with. Only named, charset-bounded fields ever reach the log; prompt text, tool descriptions, the raw assistant message, and the marker's "why" half never do. `node .claude/hooks/route-metrics.mjs --summary [--since <ISO date>]` reports turns, route-marker coverage, lanes by count, dispatches by `subagent_type`, dispatches with no matching start, and mean/max duration per agent type. Plain Node, zero deps, prints nothing to stdout on any event, fail-open (a miss is a missing log line, never a blocked turn). Installed and wired only when claude-code is the primary, same no-overwrite rules as the other two hooks. See `docs/audit-brief.md` for the full threat-model writeup.
12
+ - **CI now runs on `windows-latest` too, node 18/20/22, alongside Ubuntu and macOS.** `defaults.run.shell: bash` makes every workflow step Git Bash on the Windows runner instead of the default `pwsh`, so the same script runs on all three OSes with no parallel Windows rewrite.
13
+
14
+ ### Fixed
15
+
16
+ - **`finding-verifier` and `code-reviewer` (claude-code) called themselves unqualified "Read-only" in their descriptions while carrying an unrestricted `Bash` grant**, the same overclaim `done-verifier` shipped with in 0.1.15 and was fixed there; nothing in that grant stops either from running a mutating command. Both descriptions and bodies now say plainly that they carry no file-editing tools and that Bash is bound by the prompt, not the tool grant. `templates/agents/claude-code/README.md` and `templates/agents/snippets/claude-code.md` are corrected the same way. The done-verifier-only test is replaced with one that walks every claude-code agent file: any agent whose `tools:` line includes `Bash` must qualify any "Read-only" claim, checked against its own file and against every generated doc surface.
17
+ - **`CLAUDE.md` and `CONTRIBUTING.md` each pinned a fixed test count that drifted the moment a test was added or removed**, the same class of drift `AGENTS.md` already avoided by saying the suite prints the current number instead. Both now say the same thing `AGENTS.md` does. A new test in `test/prose.test.js` fails if any top-level `.md` states a fixed count of test cases.
18
+ - **`which()` (`src/detect.js`, and its standalone copy in `bin/cli-run.mjs`) never found a Windows lane binary, because a PATH entry there never holds a bare `grok`: npm and vendor installers drop `grok.cmd` (or `.exe`/`.bat`), the same way any Windows shell resolves a bare command through `%PATHEXT%`.** `which()` now tries the bare name first (a no-op on POSIX, and still matching an already-extensioned name on Windows), then each `%PATHEXT%` suffix. `platform` is a parameter (default `process.platform`), the same pattern `killTree(pid, platform, deps)` already used, so the win32 branch has a real test (`test/detect.test.js`, new) on every OS this suite runs on. `home` is a parameter too, for the same testability reason: a dev machine with a real vendor CLI already on `~/.local/bin` made the new win32 tests false-negative until it was injectable.
19
+ - **Committed text files were checked out as CRLF on `windows-latest`, breaking every test that compares a file's exact bytes against a string built in memory with `\n`** (`docs/catalog.md` vs. `catalogMarkdown()`, the README's generated vendor-table section, `llms.txt`'s first line, and a backslash-continued shell example's line-rejoin regex in a test). New `.gitattributes` (`* text=auto eol=lf`) forces LF on checkout regardless of a contributor's or a runner's `core.autocrlf`; `test/fixtures/` (real captured vendor output, byte-exact on purpose) is marked `-text` so line-ending normalization never touches it.
20
+ - **A large block of `test/install.test.js` and `test/cli.test.js` compared a planned file's `f.rel` (built with `path.join`, so backslash-separated on win32) against a hardcoded forward-slash literal** (`'protocols/build-protocol.md'`, `'vm/README.md'`, `'.claude/agents/deep-planner.md'`, and about a dozen more), which is never equal on Windows; several other assertions hardcoded a POSIX `:` `PATH` delimiter and `/usr/bin`, `/bin` absolute paths, or replaced a spawned child's `env` outright and dropped `PATH`/`USERPROFILE`/the rest of the parent environment Windows itself needs. Path literals now go through `join(...)`; PATH construction goes through `node:path`'s `delimiter`; every replaced `env` object spreads `process.env` first and sets `USERPROFILE` alongside `HOME` (`os.homedir()` does not consult `HOME` on win32).
21
+ - **On a real Windows install, a reconfigure's "applied:" line always said "nothing" and the existing-runtime upgrade/conflict checks never fired**, because `bin/cli.js` checked a written file's `f.rel` (backslash-separated on win32, built by `path.join`) directly against `MACHINE_OWNED`/`RUNTIME` (`src/install.js`, hand-written with forward slashes), which never match on that OS. `fileClass()` already normalized before checking; it now exports that normalizer (`toPosixRel`) for `bin/cli.js`'s two direct checks to use too. Both take an optional `separator` parameter (default the real `path.sep`) so the win32 case has a test (`test/install.test.js`) provable from any host. Found via `test/cli.test.js`'s "#6: rerunning with an added lane..." on windows-latest CI.
22
+ - **A replaced child `env` object could carry both `PATH` and the host's own differently-cased `Path` key at once** (`{ ...process.env, PATH: x }` adds a new key next to whichever case the real environment block used; Windows env vars are case-insensitive, plain JS object keys are not), so which one a spawned child actually saw was implementation-defined, not last-key-wins. `test/cli.test.js`'s `mergeEnv()` replaces any existing case-variant of an overridden key instead of adding a second one.
23
+ - **`test/cli.test.js`'s fake lane binaries are `#!/bin/sh` scripts, which Windows cannot execute as `argv[0]`** (no shebang interpretation in `CreateProcess`). `writeShellStub()`/`writeNodeStub()` write a `.cmd` launcher beside the POSIX file on win32 that hands off to Git Bash's `sh.exe` or `node`; `which()`'s detection of the resulting binary works, but running one all the way through `cli-run.mjs`'s own spawn path is not yet reliable on Windows CI for a reason this pass did not fully root-cause. Those specific tests (and the `#12` upgrade-path pair, which diverges in its own, separately unclear way) are skipped on win32 with a stated reason rather than shipped flaky or silently broken; `killTree`'s win32 branch and `which()`'s `%PATHEXT%` resolution each keep their own direct, passing test. A handful of other tests assert an exact POSIX `--dir`/`--project` path string verbatim in output, which `path.resolve()` reinterprets as drive-relative on win32 (a real design question - should a level-3 `--dir` describing a remote Linux box's path ever go through the local host's path semantics at all? - out of scope to decide here) and are skipped the same way. `statSync().mode`'s executable bit is a POSIX-only assertion, dropped on win32 rather than asserted against a filesystem that has no equivalent concept.
24
+
25
+ ## [0.1.15] - 2026-09-10
26
+
27
+ The portable parts of a live routing revision, delegate by default, gated on one verified fact rather than a guess: [code.claude.com/docs/en/sub-agents](https://code.claude.com/docs/en/sub-agents) states that a non-fork Claude Code subagent's initial context includes "every level of the CLAUDE.md hierarchy the main conversation loads", and that the built-in Explore and Plan agents skip it. No other lane in this catalog has that documented, so everything below is gated on `subagentsLoadRules(primary)`, currently true for claude-code alone; every other primary keeps its original wording unchanged.
28
+
29
+ ### Added
30
+
31
+ - **`subagentsLoadRules: true` on the claude-code catalog entry**, with the doc quote as its comment. Drives every new render var below through `src/install.js`; nothing here is a template branch, per the house rule that templates carry no logic.
32
+ - **Builder executes by default, on claude-code.** `ROUTING.md` rule 5, its "Who builds" section, the "Add an endpoint" example, and the claude-code `CLAUDE.snippet.md` now say: the orchestrator plans, briefs, verifies and talks to the human; it stays inline only when (a) the brief would cost as much as the work, (b) the task needs this conversation's own context, or (c) it is the human's decision or the final verification of delegated work. Every other primary keeps "the orchestrator builds it directly."
33
+ - **Two hooks, claude-code only: `route-gate.mjs` (`UserPromptSubmit`) and `subagent-context.mjs` (`SubagentStart`)**, written to `.claude/hooks/`, wired by a `settings.hooks.snippet.json` the user merges into `.claude/settings.json` themselves, never written over one they have. `route-gate.mjs` reads the `<!-- route-gate:start -->...<!-- route-gate:end -->` block out of the rendered routing rules file and injects it every turn, so the table is read from the one place it is generated, not recited from memory; a missing file or block still exits 0 with a one-line fallback naming the path it looked for. `subagent-context.mjs` injects a static reminder of where the rules and `TASK_BUNDLE.md` live and that a delegate does not route further or verify its own work as final. Both are plain Node, zero deps, bounded reads, fail-open by design (a miss is a stray context string, not a gate): see `docs/audit-brief.md`.
34
+ - **`done-verifier` and `reader`, two new fast-tier agents with no file-editing tools, in both the claude-code and agy formats.** `done-verifier` probes a tracker item's stated done-signal (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything. `reader` reads and digests many files or notes and returns exactly what the brief asks (facts, quotes cited `path:line`, an index, a digest); unlike `bulk-worker`, it never classifies, tags, transforms or writes. `reader` is read-only by tool grant on both formats (no `Bash`); `done-verifier` on claude-code carries `Bash` for its probes, bound only by its prompt, not by the grant, and its description says so; on agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `ROUTING.md`'s decision tree, `TIERS.md`'s effort table, and both agent-folder READMEs now name them.
35
+ - **The inline threshold gets a measurement instead of a guessed figure.** "A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Measure yours once: spawn a subagent with a one-line task and read its token count. Work smaller than that stays inline." (claude-code `ROUTING.md` and `CLAUDE.snippet.md` only; no private number shipped.)
36
+
37
+ ### Changed
38
+
39
+ - **The package now says what it is for in the first line people and agents read.** npm search, GitHub search and the installer banner showed "Routing instructions and a CLI runner", which named the parts and not the purpose. The description, the README opening and the banner now lead with the goal (each task to the right model, agent or LLM, fewer frontier tokens) while keeping the 0.1.11 correction intact: routing is an instruction your agent follows, and the README still states it does not automatically compare prices or select models. A test holds all three surfaces to that. New: an "At a glance" block and question-shaped "Common questions" in the README, request-level alternatives named for readers who want a proxy, `llms.txt` at the root and a headless-use section in `AGENTS.md` (both now ship in the package), and search keywords matching what comparable routers use.
40
+ - **Corrected the unqualified premise "a subagent holds none of these rules" everywhere it appeared** (`TASK_BUNDLE.md`, `ORCHESTRATOR.md`, the claude-code snippet, `docs/part-1-beginner.md`, `docs/part-2-intermediate.md`, README principle 6, and `ROUTING.md`'s "Who builds"). The corrected fact: a Claude Code subagent loads CLAUDE.md and so keeps the standing rules, just not this task's scope; a second CLI or a fresh chat window may still hold none of it. "Absence is denial" is unchanged; only the premise about who is absent what was wrong.
41
+ - **The claude-code snippet's closing "available as ..." agent list is generated from the files actually shipped in `templates/agents/claude-code/`, never hand-typed.** It had drifted once already: `finding-verifier` shipped in 0.1.14 and was missing from this sentence until now. `claudeAgentIds()` in `src/install.js` reads the folder; a test ties the rendered list to it.
42
+ - **The delegate-by-default gate now reaches every generated surface it should, not just three of them.** `build-protocol.md`'s "Roles, as capabilities" table and its "Why the builder does not hand off the main build" line, `builder.md`'s description, and `ROUTING.md`'s "Plan big, execute small" modifier still said, on a claude-code install, that the orchestrator writes the main build itself, never hands it off whole, and that a delegate inherits none of the session's rules: the exact premise the rest of this release corrects. All four now render through `subagentsLoadRules(primary)` the same way the decision tree and "Who builds" already did; every other primary is unchanged. A semantic-regression test asserts a claude-code install contains none of the old phrasing and a codex install still does.
43
+
44
+ ### Fixed
45
+
46
+ - **Both new hooks could hang, and `route-gate.mjs` could read an unbounded or blocking file (pre-release audit finding, never shipped).** `readFileSync(0)` in both `route-gate.mjs` and `subagent-context.mjs` blocked until stdin reached EOF, so a caller that piped input in without closing its end (or ran the hook from a bare TTY) left the process running indefinitely; reproduced with `sleep 3 | CLAUDE_PROJECT_DIR=... node route-gate.mjs` still running past 1.5s. Separately, `route-gate.mjs` read the whole rules file into memory before bounding it (`readFileSync(path).slice(0, MAX_READ)`), so a FIFO planted at the rules path blocked forever on open, and a very large file was read in full before being truncated. Fixed in both hooks: stdin is now drained asynchronously against a 250ms hard cap, never blocking past it. `route-gate.mjs` additionally `statSync`s the resolved path and refuses anything that is not `isFile()` (a FIFO, socket, device or directory, symlink target included) before ever calling open, then reads through a single fixed 64 KB buffer via `openSync`/`readSync`, closed in a `finally`, so neither the read time nor the memory used depends on the file's on-disk size. Tests: an open, never-closed stdin pipe now exits within 1s for both hooks; a FIFO at the rules path returns the fallback instead of hanging; a 200 MB sparse rules file completes in well under a second with output still capped.
8
47
 
9
48
  Three refinements to the routing model, from a review by [@shawnwows](https://x.com/shawnwows). The theme is the same in all three: a routing decision that was implied, inherited or asserted is now stated, pinned or checked.
10
49
 
@@ -217,7 +256,9 @@ First release.
217
256
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
218
257
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
219
258
 
220
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...HEAD
259
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.16...HEAD
260
+ [0.1.16]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.15...v0.1.16
261
+ [0.1.15]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...v0.1.15
221
262
  [0.1.14]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...v0.1.14
222
263
  [0.1.13]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.12...v0.1.13
223
264
  [0.1.12]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.11...v0.1.12
package/README.md CHANGED
@@ -2,7 +2,16 @@
2
2
 
3
3
  [![npm](https://img.shields.io/npm/v/model-orchestrator.svg)](https://www.npmjs.com/package/model-orchestrator) [![test](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml/badge.svg)](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml) [![license: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![node >=18](https://img.shields.io/badge/node-%3E%3D18-brightgreen.svg)](package.json)
4
4
 
5
- **Routing instructions and a CLI runner for your AI tools.** One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine. Your primary agent follows the instructions to choose a tier or lane; the runner executes the lane it is given. It does not automatically compare prices or select models.
5
+ **Route every task to the right model, agent or LLM, and spend fewer tokens.** A model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. One installer asks what you have access to and writes routing rules, subagents and a CLI runner for exactly that setup, from one chat app to several agent CLIs or a virtual machine. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models; the runner executes the lane it is given. On Claude Code it also delegates execution to subagents by default, with two hooks that inject the routing table every turn.
6
+
7
+ ## At a glance
8
+
9
+ - **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
10
+ - **What it is not:** a proxy, a gateway or an API router. It does not automatically compare prices or select models; your agent follows the rules and chooses.
11
+ - **Install:** `npx model-orchestrator` (interactive), or headless from a script or an agent: `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator`.
12
+ - **Use it when:** you run more than one model or agent and want each task sent to the smallest one that can do it well.
13
+ - **What it saves:** frontier-model tokens. Bulk work, reading and checks go to fast tiers; the expensive tier is kept for planning and judgment.
14
+ - **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
6
15
 
7
16
  Built from a working system, not a diagram: the routing rules, the protocols and the lane runner here run in production, generalized so they transfer to any stack.
8
17
 
@@ -18,7 +27,7 @@ The installer asks a few things, then writes a folder:
18
27
  2. **Which AIs do you have access to?** (it marks the ones already on your PATH)
19
28
  3. **Which one is your primary agent?** (the one that runs the system)
20
29
 
21
- It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: follow the generated README. Inspect the manifest and remove only the individual managed subagent files you no longer need, preserve edited or pre-existing files, and remove your manually pasted activation block. Never delete a shared subagent folder.
30
+ It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions (and, on Claude Code, two hook scripts) go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: follow the generated README. Inspect the manifest and remove only the individual managed subagent files you no longer need, preserve edited or pre-existing files, and remove your manually pasted activation block. Never delete a shared subagent folder.
22
31
 
23
32
  ## The three levels
24
33
 
@@ -63,7 +72,7 @@ An install has two targets, and a scripted run should set both.
63
72
  | Flag | Default | What lands there |
64
73
  |---|---|---|
65
74
  | `--dir` | `./ai-orchestrator` | the docs, protocols and (level 2+) `bin/cli-run.mjs`. Named after what it contains, not after this package, so a project can hold one without looking like a checkout of it. Pass `--dir ./model-orchestrator` if you prefer the package name. |
66
- | `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. |
75
+ | `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets three hook scripts in `.claude/hooks/`, wired by a settings snippet you merge yourself. |
67
76
 
68
77
  `--project` defaulting to the current directory is the one that surprises people: run the command from your home folder with Claude Code as the primary and five agent files land in your home folder. The installer prints the resolved project path in the plan and says when you left it at the default. Set it.
69
78
 
@@ -90,8 +99,10 @@ ai-orchestrator/
90
99
  TASK_BUNDLE.md the brief every delegation carries
91
100
  protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record
92
101
  CODECALC.md OBSIDIAN-TC.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
93
- <project>/.claude/agents/ six subagents, one per tier plus finding-verifier, at the PROJECT root (if Claude Code is primary)
102
+ <project>/.claude/agents/ one per tier plus finding-verifier, done-verifier, reader, at the PROJECT root (if Claude Code is primary)
103
+ <project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart) + route-metrics.mjs (all five: see "Measuring routing" below), Claude Code only
94
104
  CLAUDE.snippet.md the block to paste into your CLAUDE.md
105
+ settings.hooks.snippet.json the hooks block to merge into .claude/settings.json (Claude Code only)
95
106
  ROUTING.md multi-lane decision tree (level 2+)
96
107
  TIERS.md DELEGATION_MATRIX.md RESEARCH_TRIAGE.md CLI-RUN.md
97
108
  bin/cli-run.mjs bin/lanes.json (node bin/cli-run.mjs --doctor is the smoke test)
@@ -106,7 +117,7 @@ ai-orchestrator/
106
117
  | [`src/`](src/README.md) | the catalog, the pure planner, detection, rendering |
107
118
  | [`templates/`](templates/README.md) | everything the installer can write, by level, plus `tools/` for companions |
108
119
  | [`docs/`](docs/README.md) | the three parts and the catalog |
109
- | [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu and macOS, Node 18/20/22 |
120
+ | [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu, macOS and Windows, Node 18/20/22 |
110
121
 
111
122
  ## What is enforced, what is delegated, what is an instruction
112
123
 
@@ -150,7 +161,7 @@ One number per lane, and it is the same number the installer pins: where a lane
150
161
 
151
162
  **The live canary runs on your machine, with your credentials.** That is what `node bin/cli-run.mjs --doctor --run` is: it sends every enabled lane one tiny prompt through your own sign-ins and reports `canary ok` or `canary FAILED rc=` per lane. Run it after install, and again after any vendor upgrade.
152
163
 
153
- It deliberately does not run in this repository's CI. A canary is only meaningful against real credentials, and there are no credentials a maintainer could supply that would tell **you** anything about **your** lanes: your sign-ins, your quota, your vendor versions. A maintainer-credential canary in CI would prove one machine works and bill someone per run to do it. So CI runs the full suite against stub lanes on Ubuntu and macOS, Node 18/20/22, plus a packaged install into a clean consumer, and the live check ships to you instead.
164
+ It deliberately does not run in this repository's CI. A canary is only meaningful against real credentials, and there are no credentials a maintainer could supply that would tell **you** anything about **your** lanes: your sign-ins, your quota, your vendor versions. A maintainer-credential canary in CI would prove one machine works and bill someone per run to do it. So CI runs the full suite against stub lanes on Ubuntu, macOS and Windows, Node 18/20/22, plus a packaged install into a clean consumer, and the live check ships to you instead.
154
165
 
155
166
  ## Principles the whole thing rests on
156
167
 
@@ -159,7 +170,7 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
159
170
  3. **Exit 0 is not a deliverable.** Check for the artifact, not the status line. `cli-run` checks the response is structurally there; `--expect-file` checks the artifact.
160
171
  4. **Numbers are computed, never guessed.** A tool that calculates beats a model that feels finished.
161
172
  5. **A write nobody can find again did not happen.** Search first, keep the index true, one writer.
162
- 6. **The orchestrator owns the main build.** Delegates hold none of your rules; they get bounded sub-parts and a brief.
173
+ 6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, only the brief carries what this task needs. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
163
174
  7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
164
175
 
165
176
  ## Routing by role, complexity and risk
@@ -190,6 +201,21 @@ change. Use a different model family from the one that produced the finding
190
201
  where you have one: a family asked to check its own claim tends to agree with
191
202
  itself.
192
203
 
204
+ ## Two more fast-tier checks
205
+
206
+ `done-verifier` probes the artifact a tracker item's done-signal names (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything itself. It carries no file-editing tools, but on claude-code it does carry `Bash` for those probes (`git log`, `grep`, `wc -l`, `test -f`); staying to read-only commands there is a rule in its prompt, not a restriction on the tool grant, and its own description says so. On agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `reader` is the one that is read-only by tool grant on both: no `Write`, `Edit`, or `Bash`. It reads and digests many files or notes and hands back exactly what the brief asked for, cited by `path:line`; it never classifies, tags or writes, which is what separates it from `bulk-worker`. Both ship in the claude-code and agy agent sets, at the fast tier.
207
+
208
+ ## Measuring routing
209
+
210
+ A routing rule nobody measures is a rule nobody knows is followed. On a claude-code install, `route-metrics.mjs` (the third hook, wired to `UserPromptSubmit`, `PreToolUse` on `Agent`/`Task`, `SubagentStart`, `SubagentStop` and `Stop`) turns each of those into one JSON line under `~/.ai-orchestrator/route-metrics.jsonl`: a turn started, a subagent was dispatched (and with what, and in the background or not), a subagent started and stopped (so a duration can be computed), and the lane your agent named in its own hidden `<!-- route: <lane> | <why> -->` marker, which the route-gate block now asks for on every reply. It never logs prompt text, tool descriptions, or the "why" half of the marker: only the named fields above, charset-bounded, same principle as `cli-run.mjs`'s log.
211
+
212
+ ```bash
213
+ node .claude/hooks/route-metrics.mjs --summary # since the log began
214
+ node .claude/hooks/route-metrics.mjs --summary --since 2026-09-01 # since a date
215
+ ```
216
+
217
+ The report prints turns, **route-marker coverage** (the percentage of turns whose `Stop` event carried a real lane, not `missing`, which is the number that answers "is the agent actually tagging its routing decisions?"), lanes by count, dispatches by `subagent_type`, dispatches with no matching start (a hook or guard blocked the subagent before it launched), and mean/max duration per agent type. Fail-open by design, like the other two hooks: a miss here is a missing log line, never a blocked turn, and it prints nothing to stdout on any event since stdout on `UserPromptSubmit`/`SubagentStart` becomes model context.
218
+
193
219
  ## Pin the route, or know that you did not
194
220
 
195
221
  A lane with no `--model`, no `--effort` and no `defaults` entry in
@@ -210,9 +236,27 @@ an actual field would be present for one lane and missing for four, and it
210
236
  would be a provider-supplied string, which the durable log deliberately never
211
237
  holds.
212
238
 
239
+ ## Common questions
240
+
241
+ ### How do I cut token usage across Claude Code, Codex and Gemini?
242
+
243
+ Install for the tools you have, then let the generated `ROUTING.md` decide the tier per task: bulk, reading and verification go to the fast tier or a cheaper CLI lane, and the deep tier only plans and judges. On Claude Code, execution goes to the `builder` subagent by default and the main session plans and verifies. Every lane call through `cli-run` logs the model and effort it ran with, so you can check where the tokens went.
244
+
245
+ ### How do I route tasks to cheaper models?
246
+
247
+ The rules route by role, complexity and risk (see [Routing by role, complexity and risk](#routing-by-role-complexity-and-risk)). Role picks the agent, complexity moves the effort, risk moves the tier. A task a cheap tier finishes correctly never gets a frontier token.
248
+
249
+ ### Is this an LLM router or an AI gateway?
250
+
251
+ No. It routes at the task level, through instructions your agent follows and a runner for agent CLIs. If you want a service or proxy that picks or forwards the model on every API request, look at request-level routers and gateways such as RouteLLM, LiteLLM, OpenRouter or claude-code-router. They solve a different problem and can sit underneath this.
252
+
253
+ ### Can an agent install and run it without a person?
254
+
255
+ Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run` previews the plan, and `--list` prints every supported AI. Nothing is appended to a file you already have; activation snippets are written next to your files for you to merge.
256
+
213
257
  ## Requirements
214
258
 
215
- Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows is untested: `cli-run` ends a lane's process tree there with `taskkill`, but nothing in CI runs on Windows, so treat it as unsupported until someone reports otherwise.
259
+ Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22). Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested there; a named set of tests is skipped on Windows, each with its reason in the test file, mainly running a lane end to end through `cli-run`, so treat lane execution on Windows as unproven until someone reports otherwise.
216
260
 
217
261
  **Privacy.** The installer sends no telemetry and makes no network call of its own once it is running. Two things around that are worth being exact about:
218
262
 
package/bin/cli-run.mjs CHANGED
@@ -75,18 +75,33 @@ export const REASONS = new Set([
75
75
 
76
76
  const LOG = join(homedir(), '.ai-orchestrator', 'cli-run.log.jsonl');
77
77
 
78
+ // On win32, a PATH entry never holds a bare "grok": npm and vendor installers
79
+ // drop "grok.cmd" (or .exe/.bat/.ps1), the same way any Windows shell resolves
80
+ // a bare command through %PATHEXT%. Trying the bare name first keeps this a
81
+ // no-op on POSIX and matches an already-extensioned name (a .exe someone put
82
+ // on PATH directly) on Windows too. Kept in sync with src/detect.js's which(),
83
+ // which this file cannot import: it ships standalone into a user's install.
84
+ function candidateExtensions() {
85
+ if (process.platform !== 'win32') return [''];
86
+ const pathext = process.env.PATHEXT || '.COM;.EXE;.BAT;.CMD';
87
+ return ['', ...pathext.split(';').filter(Boolean)];
88
+ }
89
+
78
90
  function which(bin) {
79
91
  const dirs = (process.env['PATH'] || '').split(delimiter).filter(Boolean);
80
92
  const home = homedir();
81
93
  dirs.push(join(home, '.local', 'bin'), join(home, '.grok', 'bin'), join(home, '.npm-global', 'bin'));
94
+ const exts = candidateExtensions();
82
95
  for (const d of dirs) {
83
- const p = join(d, bin);
84
- try {
85
- if (!statSync(p).isFile()) continue; // a directory named like the binary is not the binary
86
- accessSync(p, constants.X_OK);
87
- return p;
88
- } catch {
89
- /* next */
96
+ for (const ext of exts) {
97
+ const p = join(d, bin + ext);
98
+ try {
99
+ if (!statSync(p).isFile()) continue; // a directory named like the binary is not the binary
100
+ accessSync(p, constants.X_OK);
101
+ return p;
102
+ } catch {
103
+ /* next */
104
+ }
90
105
  }
91
106
  }
92
107
  return null;
package/bin/cli.js CHANGED
@@ -10,7 +10,7 @@ import { spawnSync } from 'node:child_process';
10
10
  import { resolve, join } from 'node:path';
11
11
  import { which } from '../src/detect.js';
12
12
  import { AIS, LEVELS, TOOLS, PROVIDERS, aisForLevel, agentCandidates, byId, npmSpec } from '../src/catalog.js';
13
- import { planFiles, writeFiles, resolveSelection, resolveTools, resolveApis, dirProblems, readManifest, activationSteps, MACHINE_OWNED, RUNTIME, GENERATOR_VERSION } from '../src/install.js';
13
+ import { planFiles, writeFiles, resolveSelection, resolveTools, resolveApis, dirProblems, readManifest, activationSteps, MACHINE_OWNED, RUNTIME, toPosixRel, GENERATOR_VERSION } from '../src/install.js';
14
14
 
15
15
  // One strict parse. Unknown flags, missing values and duplicates are usage
16
16
  // errors (exit 2) before anything is planned, so a typo like --dryy can never
@@ -144,7 +144,7 @@ async function main() {
144
144
  // Says what this generates, not what it guarantees. The old line promised
145
145
  // routing this package does not perform: lane choice is an instruction an
146
146
  // agent follows, never something enforced here (#11).
147
- console.log('\nmodel-orchestrator\nRouting instructions and a CLI runner for the AIs you actually have.\n');
147
+ console.log('\nmodel-orchestrator\nA model orchestrator: routing rules and a CLI runner for the AIs you actually have.\n');
148
148
 
149
149
  // 1. Level
150
150
  let level = Number(opt('level'));
@@ -310,10 +310,10 @@ async function main() {
310
310
  if (e && e.code === 'PREFLIGHT') bad(e.message);
311
311
  throw e;
312
312
  }
313
- const ownedWritten = written.filter((w) => MACHINE_OWNED.has(w));
313
+ const ownedWritten = written.filter((w) => MACHINE_OWNED.has(toPosixRel(w)));
314
314
  console.log(`\nWrote ${written.length} file(s)` + (skipped.length ? `, kept ${skipped.length} existing:` : '.'));
315
315
  for (const s of skipped) console.log(' kept ' + s);
316
- const existingRuntime = files.filter((f) => f.root !== 'project' && RUNTIME.has(f.rel)).length;
316
+ const existingRuntime = files.filter((f) => f.root !== 'project' && RUNTIME.has(toPosixRel(f.rel))).length;
317
317
  if (prev || existingRuntime && (upgraded.length || conflicts.length || unverifiable.length) || docsUnverifiable.length) {
318
318
  console.log(`\nExisting installation found${prev ? ` (MANIFEST.json from generator ${prev.generatorVersion || 'pre-0.1.1'}, ${prev.generatedAt || 'undated'}; this run is ${GENERATOR_VERSION})` : ' (no MANIFEST.json: it predates 0.1.1)'}.`);
319
319
  if (prev) {
@@ -81,3 +81,35 @@ Re-audit the same scope. Every round-1 finding was reproduced before it was touc
81
81
  Also new since round 1: the companion-tool path (`--tools codecalc`, `--no-tools`, `templates/tools/codecalc/`, `protocols/numbers-and-logic.md`, `resolveTools`). Attack it the same way: unknown tool ids, interaction with `--yes`, the extra interactive question, and whether any written snippet could be confused for a file the installer should not touch.
82
82
 
83
83
  Report only what reproduces on the current tree. `CLEAN` per area is expected where the fix holds.
84
+
85
+ ## New in 0.1.15: two claude-code-only hooks
86
+
87
+ `route-gate.mjs` (`UserPromptSubmit`) and `subagent-context.mjs` (`SubagentStart`) ship to `.claude/hooks/` only when claude-code is the primary. Both are plain Node, zero deps, and installed with mode `0o755`.
88
+
89
+ - **Reads.** `route-gate.mjs` reads at most 64 KB from one file: the routing rules file (`ROUTING.md` or `ORCHESTRATOR.md`) at a path rendered in at install time relative to `CLAUDE_PROJECT_DIR`, never a hardcoded absolute path. Before opening it, it `statSync`s the resolved path (following a symlink to its target) and refuses anything that is not `isFile()`, a FIFO, socket, device or directory included, so the read never touches a path that could block on open. The read itself is one `openSync` + one bounded `readSync` into a fixed 64 KB buffer, closed in a `finally`, so neither the time nor the memory this hook uses depends on how large the file on disk actually is. It then extracts the text between `<!-- route-gate:start -->` and `<!-- route-gate:end -->` and nothing else. `subagent-context.mjs` reads nothing from disk; its context is static text plus the same two rendered paths. Neither parses or executes anything it reads; the extracted block is passed through as a string.
90
+ - **Stdin.** Neither hook uses a field from the JSON input Claude Code sends on stdin, but both must still consume the pipe rather than ignore it. Both drain stdin asynchronously against a 250ms hard cap: whichever comes first, the real `end` event or the timeout, the hook proceeds. Neither ever calls a blocking, synchronous read of stdin.
91
+ - **Writes.** Neither writes a file. Both write one JSON object to stdout: `{"hookSpecificOutput":{"hookEventName":"...","additionalContext":"..."}}`, and both exit only after that write's callback fires, so a buffered write to a pipe is not truncated by an exit racing ahead of it.
92
+ - **Fail-open, on purpose.** A missing `CLAUDE_PROJECT_DIR`, a missing rules file, a non-regular file at the rules path, or a missing block each produce a one-line fallback `additionalContext` naming what was found, and the script still exits 0. This is acceptable because a miss here is a stray context string reaching the model, not a security gate: nothing downstream trusts the hook's output for anything but a routing suggestion, and the settings snippet that wires it in is a document the user merges by hand, never written automatically over an existing `settings.json`.
93
+ - **Bounded.** `route-gate.mjs` caps the read at 64 KB regardless of the file's reported size and the injected string at 4000 characters, so a rules file bloated by a bad edit, or truncated to an arbitrary length, cannot balloon the context or the read time on every turn. Both hooks cap stdin drain at 250ms.
94
+ - **Not yet attacked.** Untested here: a rules file with a `route-gate:start` marker but no matching end marker very far into the file (bounded by `MAX_READ`, so the end marker past that point is treated as absent, which is the intended fail-open path, but worth a deliberate case); a `CLAUDE_PROJECT_DIR` pointing at a path with no read permission; behavior under the Windows exec-form `node` + `args` invocation named in the settings snippet; a `statSync` that itself hangs (a stalled network filesystem, for instance) rather than the FIFO-at-open case this round fixed.
95
+
96
+ ### Round 1 (pre-release), fixed before shipping
97
+
98
+ Three findings reproduced against the 0.1.15 branch before it shipped, none of them ever released:
99
+
100
+ | # | Finding | Fix |
101
+ |---|---|---|
102
+ | 1 | HIGH. `readFileSync(0)` in both hooks blocked until stdin reached EOF (`sleep 3 \| ... node route-gate.mjs` still running past 1.5s); `route-gate.mjs` also read the whole rules file into memory before bounding it, so a FIFO planted at the rules path blocked forever on open. | Stdin is drained asynchronously against a 250ms hard cap in both hooks. `route-gate.mjs` refuses anything that is not `isFile()` via `statSync` before ever calling open, then reads through one fixed 64 KB buffer via `openSync`/`readSync`. Tests: an open, never-closed stdin pipe exits within 1s for both hooks; a FIFO at the rules path returns the fallback instead of hanging; a 200 MB sparse rules file completes in well under a second with output still capped. |
103
+ | 2 | MEDIUM. `done-verifier`'s description, both agent-folder READMEs, and the root README called it "read-only" without qualification, while its claude-code file carries an unrestricted `Bash` grant; nothing in that grant stops it from running a mutating command. | Every one of those surfaces now says plainly that `done-verifier` carries no file-editing tools and that its Bash use is bound by its own prompt, not by the tool grant; `reader` is named as the one that is read-only by tool grant (no Bash) on both formats. |
104
+ | 3 | MEDIUM. Three generated surfaces still stated the pre-0.1.15 premise on a claude-code install: `builder.md`'s description ("... or the main build itself"), `build-protocol.md`'s roles table and its "why the builder does not hand off" note, and `ROUTING.md`'s "Plan big, execute small" line ("the orchestrator executes"). | All three now render through `subagentsLoadRules(primary)`, the same gate the decision tree and "Who builds" already used; every other primary is unchanged. A semantic-regression test asserts a claude-code install contains none of the old phrasing and a codex install still does. |
105
+
106
+ ## New in 0.1.16: a third hook, route-metrics.mjs
107
+
108
+ `route-metrics.mjs` ships to `.claude/hooks/` alongside `route-gate.mjs` and `subagent-context.mjs`, only when claude-code is the primary. Same shape as the other two: plain Node, zero deps, mode `0o755`. Unlike them, it is wired to five events at once (`UserPromptSubmit`, `PreToolUse` matched to `Agent|Task`, `SubagentStart`, `SubagentStop`, `Stop`), and it does write, deliberately: one JSON line per event, appended to `~/.ai-orchestrator/route-metrics.jsonl`.
109
+
110
+ - **Reads.** Only its own stdin (the JSON Claude Code sends per event) and, for `--summary`, its own log file. It never reads `transcript_path` even though that field is present on every event: the documented source for the route marker is `last_assistant_message`, and the docs say the transcript can lag, so a hook that read it instead could log a stale or absent marker as if it were current. It never reads the rules file, the task bundle, or any other project file.
111
+ - **Writes.** `~/.ai-orchestrator/route-metrics.jsonl` (append-only, rotated to `.jsonl.1` above 5 MB) and `~/.ai-orchestrator/route-metrics.state/<sha256(agent_id)>.json`, a small file recording a subagent's start time and type so `SubagentStop` can compute a duration; it is deleted on stop, and anything older than 24h is pruned on the next `SubagentStart`. Nothing outside `~/.ai-orchestrator/`. It never writes to stdout: on `UserPromptSubmit` and `SubagentStart`, stdout becomes model context, and this hook has nothing to say there, so it stays silent on every event, not just those two.
112
+ - **What it never logs.** Prompt text, tool descriptions, the full `tool_input`, `last_assistant_message` itself, or the "why" half of a route marker. Only six named fields ever reach a record: `session_id`, `subagent_type`, `agent_type`, and the parsed `lane`, each stripped to `[A-Za-z0-9_.+-]` and capped at 64 characters (128 for `session_id`) before being written, plus the event name and a `duration_s` number it computed itself. This mirrors `bin/cli-run.mjs`'s own log, which stores a fixed reason code and never a provider-supplied string.
113
+ - **Fail-open, on purpose.** Every code path that can fail (a malformed state file, a full disk, a rotation race, invalid JSON on stdin, an unrecognized event) is caught and produces no record rather than a thrown error or a non-zero exit; the process always exits 0. A miss here is a missing line in a telemetry log, never a blocked turn, so there is nothing to gate.
114
+ - **Bounded.** Stdin is drained asynchronously against a combined 1s time cap and 8 MB size cap; a payload that exceeds either is treated as truncated and parsed as nothing, never partially. `--summary` reads the log directly (never spawns anything, never executes a line in it).
115
+ - **Not yet attacked.** Untested here: two processes racing the same rotation at once (a rename plus an append landing on the same file); a state directory with thousands of leaked files from a long-lived session with a crashed hook (pruning runs, but only on `SubagentStart`, so an install that never starts a subagent again would never prune); behavior if `agent_id` collides across two concurrent subagents (sha256 makes this astronomically unlikely, not impossible).
@@ -38,7 +38,7 @@ Every build gets two checkpoints. **Before writing:** you map what it touches an
38
38
 
39
39
  ## 4. Every hand-off carries a brief
40
40
 
41
- A subagent, a fresh chat, a second window holds none of your rules and reads an unspecified edge as an open one. The brief: purpose, task class, granted scope, capabilities, denied actions, conventions it does not have, report contract (what was not done, what is unverified), exit parameters (when to stop, and what to return if it hits the bound). Absence is denial.
41
+ A subagent, a fresh chat, a second window may hold none of your rules, and that is the default to assume. The one documented exception is a Claude Code subagent: it loads the project's CLAUDE.md hierarchy at start, so it keeps the standing rules, just not this task's scope. Either way, it reads an unspecified edge as an open one. The brief: purpose, task class, granted scope, capabilities, denied actions, conventions it does not have, report contract (what was not done, what is unverified), exit parameters (when to stop, and what to return if it hits the bound). Absence is denial.
42
42
 
43
43
  ## 5. The second pass
44
44
 
@@ -32,7 +32,7 @@ There is a second thing a lane can be quietly wrong about. Left unpinned, it run
32
32
 
33
33
  ## 4. Every delegation carries a task bundle, on both surfaces
34
34
 
35
- Subagents and CLI lanes are the same problem: something with none of your rules and broad tool access. The brief (purpose, task class, scope, capabilities, denied actions, conventions, report contract, exit parameters) goes in the prompt or in the file passed to `--brief`. If you can, gate it mechanically: a pre-dispatch hook that refuses a brief missing purpose, denied actions or a report contract.
35
+ Subagents and CLI lanes are close to the same problem: something that may hold none of your rules, and broad tool access. A Claude Code subagent is the one documented exception, loading the project's CLAUDE.md hierarchy at start, so it keeps the standing rules but not this task's scope; a CLI lane and a fresh chat window get no such credit. The brief (purpose, task class, scope, capabilities, denied actions, conventions, report contract, exit parameters) goes in the prompt or in the file passed to `--brief` either way. If you can, gate it mechanically: a pre-dispatch hook that refuses a brief missing purpose, denied actions or a report contract. On claude-code, a `SubagentStart` hook can inject the essentials (where the rules and the brief format live) automatically; `.claude/hooks/subagent-context.mjs` is the generated example. A third hook, `.claude/hooks/route-metrics.mjs`, turns that same delegation into a measurement instead of an assumption: it logs every turn, dispatch, subagent start/stop and the lane named in the reply's hidden route marker, and `--summary` reports route-marker coverage, dispatches with no matching start, and duration per agent type.
36
36
 
37
37
  ## 5. Research: three engines, one triager
38
38
 
package/llms.txt ADDED
@@ -0,0 +1,27 @@
1
+ # model-orchestrator
2
+
3
+ > Model orchestrator for AI coding agents and LLMs (Claude Code, Codex, Gemini, Grok, Qwen, Ollama). One installer writes routing rules, subagent definitions and a CLI lane runner for the AI tools you already have. The rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. It is not a proxy or gateway: it does not automatically compare prices or select models; your agent follows the rules and chooses.
4
+
5
+ Install and run: `npx model-orchestrator` (interactive), or headless: `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator`. Preview without writing: add `--dry-run`. List every supported AI: `npx model-orchestrator --list`. Node 18 or newer, zero runtime dependencies, MIT licence.
6
+
7
+ Levels: 1 beginner (one agent or chat app), 2 intermediate (several agent CLIs, each called through `cli-run`), 3 advanced (adds a virtual machine with a gateway and a scheduled audit job). On Claude Code the install also delegates execution to subagents by default and ships two hooks (UserPromptSubmit, SubagentStart) that inject the routing table every turn.
8
+
9
+ ## Docs
10
+
11
+ - [README](https://github.com/aunysillyme/model-orchestrator/blob/main/README.md): what it writes, flags, principles, what is enforced versus instructed
12
+ - [Part 1: beginner](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-1-beginner.md): one agent, tiers, the task bundle every delegation carries
13
+ - [Part 2: intermediate](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-2-intermediate.md): several agent CLIs, the lane runner, pinning model and effort per lane
14
+ - [Part 3: advanced](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-3-advanced.md): the VM, the gateway, scheduled jobs, privacy gates
15
+ - [AI catalog](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/catalog.md): every supported AI, how to sign in, how each is detected
16
+
17
+ ## Reference
18
+
19
+ - [CLI runner](https://github.com/aunysillyme/model-orchestrator/blob/main/bin/README.md): `cli-run` lanes, exit codes, the route logged per run
20
+ - [Templates](https://github.com/aunysillyme/model-orchestrator/blob/main/templates/README.md): the routing, tiers, task bundle and protocol files the installer renders
21
+ - [Changelog](https://github.com/aunysillyme/model-orchestrator/blob/main/CHANGELOG.md): every release and the issue behind each fix
22
+ - [Agent instructions](https://github.com/aunysillyme/model-orchestrator/blob/main/AGENTS.md): running the installer from an agent, and contributing
23
+
24
+ ## Optional
25
+
26
+ - [Security policy](https://github.com/aunysillyme/model-orchestrator/blob/main/SECURITY.md)
27
+ - [Audit brief](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/audit-brief.md): the threat model and what has already been attacked
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.14",
4
- "description": "Routing instructions and a CLI runner for your AI tools. One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine. Routes by role, complexity and risk; pins and logs the model and reasoning effort each lane runs with.",
3
+ "version": "0.1.16",
4
+ "description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
5
5
  "type": "module",
6
6
  "bin": {
7
7
  "model-orchestrator": "bin/cli.js"
@@ -15,7 +15,9 @@
15
15
  "LICENSE",
16
16
  "CHANGELOG.md",
17
17
  "SECURITY.md",
18
- "scripts"
18
+ "scripts",
19
+ "llms.txt",
20
+ "AGENTS.md"
19
21
  ],
20
22
  "scripts": {
21
23
  "start": "node bin/cli.js",
@@ -55,7 +57,29 @@
55
57
  "reasoning-effort",
56
58
  "model-routing",
57
59
  "code-review",
58
- "ai-agents"
60
+ "ai-agents",
61
+ "hooks",
62
+ "claude-code-hooks",
63
+ "delegation",
64
+ "llm-router",
65
+ "ai-router",
66
+ "model-switching",
67
+ "multi-model",
68
+ "token-optimization",
69
+ "token-efficiency",
70
+ "cost-optimization",
71
+ "llm-orchestration",
72
+ "agent-orchestration",
73
+ "ai-orchestration",
74
+ "claude",
75
+ "anthropic",
76
+ "openai",
77
+ "gemini",
78
+ "claude-code-router",
79
+ "agent-routing",
80
+ "prompt-routing",
81
+ "coding-agent",
82
+ "ai-coding"
59
83
  ],
60
84
  "author": "aunysillyme (https://github.com/aunysillyme)",
61
85
  "license": "MIT"
package/src/README.md CHANGED
@@ -4,6 +4,6 @@
4
4
  |---|---|
5
5
  | `catalog.js` | the single list of levels and AIs. Add an AI here and the prompts, docs tables, delegation matrix, gateway config and installer all pick it up. Nothing else lists AIs. |
6
6
  | `detect.js` | PATH lookup for a binary, plus the few places vendor installers drop binaries without touching PATH. No shell-outs. |
7
- | `install.js` | pure planner: turns (level, selection, primary) into a list of files to write, rendering templates and computing every generated table. `writeFiles` is the only thing that touches disk. `activationSteps()` and `snippetFor()` live here so the terminal summary and the generated README render the same list. |
7
+ | `install.js` | pure planner: turns (level, selection, primary) into a list of files to write, rendering templates and computing every generated table. `writeFiles` is the only thing that touches disk. `activationSteps()` and `snippetFor()` live here so the terminal summary and the generated README render the same list. `subagentsLoadRules(primary)` gates every delegate-by-default render var (builder-by-default wording, the route-gate table, the inline-threshold note) on the one verified premise: a claude-code subagent loads CLAUDE.md. |
8
8
  | `prompt.js` | line-buffered questions for the interactive path; piped answers are queued, EOF mid-prompt aborts instead of confirming a write. |
9
9
  | `render.js` | `{{KEY}}` substitution. Throws on an unknown key, so a template typo fails the test suite instead of shipping a literal placeholder. |
package/src/catalog.js CHANGED
@@ -69,7 +69,15 @@ export const AIS = [
69
69
  rulesFile: 'CLAUDE.md',
70
70
  agentsDir: '.claude/agents',
71
71
  cliRun: false,
72
- models: { deep: 'opus', standard: 'sonnet', fast: 'haiku' }
72
+ models: { deep: 'opus', standard: 'sonnet', fast: 'haiku' },
73
+ // Verified at code.claude.com/docs/en/sub-agents (fetched 2026-09-10): "A
74
+ // non-fork subagent's initial context contains: CLAUDE.md files: every
75
+ // level of the CLAUDE.md hierarchy the main conversation loads ... The
76
+ // built-in Explore and Plan agents skip this." No other lane in this
77
+ // catalog has that documented, so the builder-by-default routing, the
78
+ // route-gate hook and the inline-threshold note are gated on this field
79
+ // and stay claude-code only.
80
+ subagentsLoadRules: true
73
81
  },
74
82
  {
75
83
  id: 'codex',
package/src/detect.js CHANGED
@@ -2,24 +2,43 @@ import { accessSync, statSync, constants } from 'node:fs';
2
2
  import { delimiter, join } from 'node:path';
3
3
  import { homedir } from 'node:os';
4
4
 
5
+ // On win32, a PATH entry never holds a bare "grok": npm and vendor installers
6
+ // drop "grok.cmd" (or .exe/.bat/.ps1), the same way any Windows shell resolves
7
+ // a bare command through %PATHEXT%. Trying the bare name first keeps this a
8
+ // no-op on POSIX and matches an already-extensioned name (a .exe someone put
9
+ // on PATH directly) on Windows too.
10
+ // platform is a parameter (default process.platform), not a hardcoded read,
11
+ // so the win32 branch has a test on every OS this suite runs on: the same
12
+ // pattern bin/cli-run.mjs's killTree(pid, platform, deps) already uses.
13
+ export function candidateExtensions(platform = process.platform, pathext = process.env.PATHEXT) {
14
+ if (platform !== 'win32') return [''];
15
+ return ['', ...(pathext || '.COM;.EXE;.BAT;.CMD').split(';').filter(Boolean)];
16
+ }
17
+
5
18
  // PATH lookup plus the handful of places vendor installers drop binaries
6
19
  // without touching PATH. Never a shell function, never a shell out.
7
20
  // A directory with the binary's name is not a binary (X_OK passes on
8
21
  // searchable directories), so the candidate must be a regular file.
9
- export function which(bin) {
22
+ // home is a parameter too (default homedir()), for the same reason platform
23
+ // is: a dev machine with a real ~/.local/bin/grok on it made the win32 tests
24
+ // here false-negative until this was injectable, since PATH alone was never
25
+ // the whole search.
26
+ export function which(bin, platform = process.platform, home = homedir()) {
10
27
  if (!bin) return null;
11
- const home = homedir();
12
28
  const searchPath = process.env['PATH'] || '';
13
29
  const dirs = searchPath.split(delimiter).filter(Boolean);
14
30
  dirs.push(join(home, '.local', 'bin'), join(home, '.grok', 'bin'), join(home, '.npm-global', 'bin'));
31
+ const exts = candidateExtensions(platform);
15
32
  for (const d of dirs) {
16
- const p = join(d, bin);
17
- try {
18
- if (!statSync(p).isFile()) continue;
19
- accessSync(p, constants.X_OK);
20
- return p;
21
- } catch {
22
- /* keep looking */
33
+ for (const ext of exts) {
34
+ const p = join(d, bin + ext);
35
+ try {
36
+ if (!statSync(p).isFile()) continue;
37
+ accessSync(p, constants.X_OK);
38
+ return p;
39
+ } catch {
40
+ /* keep looking */
41
+ }
23
42
  }
24
43
  }
25
44
  return null;