model-orchestrator 0.1.13 → 0.1.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (34) hide show
  1. package/AGENTS.md +26 -0
  2. package/CHANGELOG.md +41 -1
  3. package/README.md +95 -6
  4. package/bin/cli-run.mjs +144 -19
  5. package/bin/cli.js +1 -1
  6. package/docs/audit-brief.md +21 -0
  7. package/docs/part-1-beginner.md +3 -1
  8. package/docs/part-2-intermediate.md +8 -2
  9. package/llms.txt +27 -0
  10. package/package.json +32 -4
  11. package/src/README.md +1 -1
  12. package/src/catalog.js +9 -1
  13. package/src/install.js +188 -5
  14. package/templates/README.md +1 -1
  15. package/templates/agents/README.md +3 -1
  16. package/templates/agents/agy/README.md +2 -2
  17. package/templates/agents/agy/done-verifier.md +35 -0
  18. package/templates/agents/agy/finding-verifier.md +29 -0
  19. package/templates/agents/agy/reader.md +22 -0
  20. package/templates/agents/claude-code/README.md +7 -4
  21. package/templates/agents/claude-code/builder.md +6 -1
  22. package/templates/agents/claude-code/done-verifier.md +44 -0
  23. package/templates/agents/claude-code/finding-verifier.md +43 -0
  24. package/templates/agents/claude-code/reader.md +26 -0
  25. package/templates/agents/snippets/claude-code.md +9 -3
  26. package/templates/agents/snippets/route-gate.mjs +151 -0
  27. package/templates/agents/snippets/settings.hooks.snippet.json +26 -0
  28. package/templates/agents/snippets/subagent-context.mjs +76 -0
  29. package/templates/beginner/ORCHESTRATOR.md +4 -3
  30. package/templates/common/TASK_BUNDLE.md +2 -2
  31. package/templates/common/protocols/build-protocol.md +13 -3
  32. package/templates/intermediate/CLI-RUN.md +40 -3
  33. package/templates/intermediate/ROUTING.md +15 -11
  34. package/templates/intermediate/TIERS.md +44 -1
package/AGENTS.md ADDED
@@ -0,0 +1,26 @@
1
+ # AGENTS.md
2
+
3
+ Two audiences: an agent that wants to USE this package for a project, and an agent that is working ON this repository.
4
+
5
+ ## Using this package from an agent
6
+
7
+ model-orchestrator writes routing rules, subagents and a CLI lane runner so an agent sends each task to the right model, subagent or CLI and spends fewer frontier tokens. It is not a proxy or gateway. Headless use:
8
+
9
+ - `npx model-orchestrator --list` prints every supported AI id.
10
+ - `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project <repo> --dir <repo>/ai-orchestrator --dry-run` prints the plan and writes nothing.
11
+ - Drop `--dry-run` to write it. Existing files are never overwritten without `--force`; activation snippets (for example `CLAUDE.snippet.md`, `settings.hooks.snippet.json`) are written for a person or agent to merge.
12
+ - The generated `README.md` in `--dir` lists what to copy where and one smoke command to prove the rules took.
13
+ - A summary for LLMs, with links to every doc: [`llms.txt`](llms.txt).
14
+
15
+ ## Working on this repository
16
+
17
+ The files the installer writes for end users live under `templates/`.
18
+
19
+ - Read `CONTRIBUTING.md` first, then `src/README.md` (the catalog drives everything) and `docs/audit-brief.md` (the threat model and what has already been attacked).
20
+ - Run `npm test` before proposing a change and quote the count and the exit code; the suite prints the current number.
21
+ - Everything renders from `src/catalog.js`. Add an AI or a tool there, not in a template. Templates carry no logic.
22
+ - Never put a value that looks like a credential anywhere in this repo, including tests and examples. Environment variable names only.
23
+ - `bin/cli.js` writes only inside `--dir` and `--project`, never over a document without `--force`, and never runs a vendor script. A change that weakens any of those will be refused in review; the tests that hold them are in `test/install.test.js` and `test/cli.test.js`.
24
+ - `bin/cli-run.mjs` must exit non-zero when a lane produced nothing. Every judge has a red case in `test/judges.test.js`; add one before you change a judge.
25
+ - Prose in this repo uses no em dashes (`test/prose.test.js` enforces it).
26
+ - Why these rules exist: each one is the fix for a failure that reached an audit or CI. `CHANGELOG.md` names the issue behind each.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,44 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.15] - 2026-09-10
8
+
9
+ The portable parts of a live routing revision, delegate by default, gated on one verified fact rather than a guess: [code.claude.com/docs/en/sub-agents](https://code.claude.com/docs/en/sub-agents) states that a non-fork Claude Code subagent's initial context includes "every level of the CLAUDE.md hierarchy the main conversation loads", and that the built-in Explore and Plan agents skip it. No other lane in this catalog has that documented, so everything below is gated on `subagentsLoadRules(primary)`, currently true for claude-code alone; every other primary keeps its original wording unchanged.
10
+
11
+ ### Added
12
+
13
+ - **`subagentsLoadRules: true` on the claude-code catalog entry**, with the doc quote as its comment. Drives every new render var below through `src/install.js`; nothing here is a template branch, per the house rule that templates carry no logic.
14
+ - **Builder executes by default, on claude-code.** `ROUTING.md` rule 5, its "Who builds" section, the "Add an endpoint" example, and the claude-code `CLAUDE.snippet.md` now say: the orchestrator plans, briefs, verifies and talks to the human; it stays inline only when (a) the brief would cost as much as the work, (b) the task needs this conversation's own context, or (c) it is the human's decision or the final verification of delegated work. Every other primary keeps "the orchestrator builds it directly."
15
+ - **Two hooks, claude-code only: `route-gate.mjs` (`UserPromptSubmit`) and `subagent-context.mjs` (`SubagentStart`)**, written to `.claude/hooks/`, wired by a `settings.hooks.snippet.json` the user merges into `.claude/settings.json` themselves, never written over one they have. `route-gate.mjs` reads the `<!-- route-gate:start -->...<!-- route-gate:end -->` block out of the rendered routing rules file and injects it every turn, so the table is read from the one place it is generated, not recited from memory; a missing file or block still exits 0 with a one-line fallback naming the path it looked for. `subagent-context.mjs` injects a static reminder of where the rules and `TASK_BUNDLE.md` live and that a delegate does not route further or verify its own work as final. Both are plain Node, zero deps, bounded reads, fail-open by design (a miss is a stray context string, not a gate): see `docs/audit-brief.md`.
16
+ - **`done-verifier` and `reader`, two new fast-tier agents with no file-editing tools, in both the claude-code and agy formats.** `done-verifier` probes a tracker item's stated done-signal (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything. `reader` reads and digests many files or notes and returns exactly what the brief asks (facts, quotes cited `path:line`, an index, a digest); unlike `bulk-worker`, it never classifies, tags, transforms or writes. `reader` is read-only by tool grant on both formats (no `Bash`); `done-verifier` on claude-code carries `Bash` for its probes, bound only by its prompt, not by the grant, and its description says so; on agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `ROUTING.md`'s decision tree, `TIERS.md`'s effort table, and both agent-folder READMEs now name them.
17
+ - **The inline threshold gets a measurement instead of a guessed figure.** "A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Measure yours once: spawn a subagent with a one-line task and read its token count. Work smaller than that stays inline." (claude-code `ROUTING.md` and `CLAUDE.snippet.md` only; no private number shipped.)
18
+
19
+ ### Changed
20
+
21
+ - **The package now says what it is for in the first line people and agents read.** npm search, GitHub search and the installer banner showed "Routing instructions and a CLI runner", which named the parts and not the purpose. The description, the README opening and the banner now lead with the goal (each task to the right model, agent or LLM, fewer frontier tokens) while keeping the 0.1.11 correction intact: routing is an instruction your agent follows, and the README still states it does not automatically compare prices or select models. A test holds all three surfaces to that. New: an "At a glance" block and question-shaped "Common questions" in the README, request-level alternatives named for readers who want a proxy, `llms.txt` at the root and a headless-use section in `AGENTS.md` (both now ship in the package), and search keywords matching what comparable routers use.
22
+ - **Corrected the unqualified premise "a subagent holds none of these rules" everywhere it appeared** (`TASK_BUNDLE.md`, `ORCHESTRATOR.md`, the claude-code snippet, `docs/part-1-beginner.md`, `docs/part-2-intermediate.md`, README principle 6, and `ROUTING.md`'s "Who builds"). The corrected fact: a Claude Code subagent loads CLAUDE.md and so keeps the standing rules, just not this task's scope; a second CLI or a fresh chat window may still hold none of it. "Absence is denial" is unchanged; only the premise about who is absent what was wrong.
23
+ - **The claude-code snippet's closing "available as ..." agent list is generated from the files actually shipped in `templates/agents/claude-code/`, never hand-typed.** It had drifted once already: `finding-verifier` shipped in 0.1.14 and was missing from this sentence until now. `claudeAgentIds()` in `src/install.js` reads the folder; a test ties the rendered list to it.
24
+ - **The delegate-by-default gate now reaches every generated surface it should, not just three of them.** `build-protocol.md`'s "Roles, as capabilities" table and its "Why the builder does not hand off the main build" line, `builder.md`'s description, and `ROUTING.md`'s "Plan big, execute small" modifier still said, on a claude-code install, that the orchestrator writes the main build itself, never hands it off whole, and that a delegate inherits none of the session's rules: the exact premise the rest of this release corrects. All four now render through `subagentsLoadRules(primary)` the same way the decision tree and "Who builds" already did; every other primary is unchanged. A semantic-regression test asserts a claude-code install contains none of the old phrasing and a codex install still does.
25
+
26
+ ### Fixed
27
+
28
+ - **Both new hooks could hang, and `route-gate.mjs` could read an unbounded or blocking file (pre-release audit finding, never shipped).** `readFileSync(0)` in both `route-gate.mjs` and `subagent-context.mjs` blocked until stdin reached EOF, so a caller that piped input in without closing its end (or ran the hook from a bare TTY) left the process running indefinitely; reproduced with `sleep 3 | CLAUDE_PROJECT_DIR=... node route-gate.mjs` still running past 1.5s. Separately, `route-gate.mjs` read the whole rules file into memory before bounding it (`readFileSync(path).slice(0, MAX_READ)`), so a FIFO planted at the rules path blocked forever on open, and a very large file was read in full before being truncated. Fixed in both hooks: stdin is now drained asynchronously against a 250ms hard cap, never blocking past it. `route-gate.mjs` additionally `statSync`s the resolved path and refuses anything that is not `isFile()` (a FIFO, socket, device or directory, symlink target included) before ever calling open, then reads through a single fixed 64 KB buffer via `openSync`/`readSync`, closed in a `finally`, so neither the read time nor the memory used depends on the file's on-disk size. Tests: an open, never-closed stdin pipe now exits within 1s for both hooks; a FIFO at the rules path returns the fallback instead of hanging; a 200 MB sparse rules file completes in well under a second with output still capped.
29
+
30
+ Three refinements to the routing model, from a review by [@shawnwows](https://x.com/shawnwows). The theme is the same in all three: a routing decision that was implied, inherited or asserted is now stated, pinned or checked.
31
+
32
+ ### Added
33
+
34
+ - **`--model` and `--effort` on every lane, and a route recorded per run.** A lane with no flag and no `defaults` entry in `bin/lanes.json` runs on its own config file, which `cli-run` cannot see: a CLI configured months ago at a low reasoning effort keeps auditing at that effort while the routing docs describe an adversarial pass, and nothing raises an error. Each vendor spells the flags differently and `cli-run` translates (`grok -m/--reasoning-effort`, `codex -m/-c model_reasoning_effort="X"`, `agy --model/--effort`, `hermes -m/--reasoning`, `qwen -m` and no reasoning flag), each one read from that CLI's own `--help`. Flags beat `defaults`, `defaults` beats nothing, `--doctor` prints what each lane is pinned to, and the log carries `model_requested`, `effort_requested`, `model_source` and `effort_source` on every record, including runs refused before the lane started. It records no "actual": one lane of five (grok) reports a model id in its own output and the other four report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string, which the durable log never holds. `--effort` on qwen is a usage error rather than a silent drop, and route values are charset-bounded because a model id becomes an argv element and, on codex, part of a TOML value.
35
+ - **`finding-verifier`, a sixth subagent, in both agent formats.** Review and scanner findings no longer go straight to a repair. It reads the cited line, states what would trigger the problem, hunts for the guard, caller or test that makes it impossible, and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a change; INCONCLUSIVE is never rounded up to be safe or down to be tidy. Bound into the build protocol as Stage 5a, into `ROUTING.md`, and into the Claude Code activation snippet. The reproduction rule already existed in Stage 5; it had no owner, no separate model family and no way to say "I could not settle this".
36
+ - **Complexity and risk as inputs, alongside role** (`TIERS.md`). Complexity moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output. Risk (security, privacy, data loss, irreversible) moves the tier and who reads the result, because none of those failures is fixable by editing the code afterwards. A one-line change to an auth check is simple and high-risk at once, and the risk decides. Deliberately two rules and two small tables rather than a role by complexity by risk matrix: an 80-cell table is not maintained, and an unmaintained routing table is worse than none because it is believed.
37
+
38
+ ### Changed
39
+
40
+ - `--model` is no longer qwen-only. `--safe-mode` still is.
41
+ - The route is resolved before the "lane disabled" and "binary missing" refusals, so those records carry it too. Found by the pre-release audit: a run refused for a missing binary is still a run that requested a route, and a failure record without one is the gap this release exists to close.
42
+ - `bin/lanes.json` gains an optional `defaults` block. It fails closed with the rest of the file: an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane with no reasoning flag refuses every lane until it is fixed, rather than being skipped quietly.
43
+ - The generated activation list gains a step about pinning the route, and `--doctor` output gains a route column with a plain sentence about what "not pinned" means.
44
+
7
45
  ## [0.1.13] - 2026-09-08
8
46
 
9
47
  Three issues from a fresh first-run walkthrough of 0.1.12 (#26, #27, #28). Same class as 0.1.12's five: a surface describing an install that did not happen. A fourth, #25, was filed and closed as a mistake on the reporter's side, not a defect: the warning it said was missing has been printed since 0.1.12 and the repro had been read through a truncated pipe.
@@ -200,7 +238,9 @@ First release.
200
238
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
201
239
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
202
240
 
203
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...HEAD
241
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.15...HEAD
242
+ [0.1.15]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...v0.1.15
243
+ [0.1.14]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...v0.1.14
204
244
  [0.1.13]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.12...v0.1.13
205
245
  [0.1.12]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.11...v0.1.12
206
246
  [0.1.11]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.10...v0.1.11
package/README.md CHANGED
@@ -2,7 +2,16 @@
2
2
 
3
3
  [![npm](https://img.shields.io/npm/v/model-orchestrator.svg)](https://www.npmjs.com/package/model-orchestrator) [![test](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml/badge.svg)](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml) [![license: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![node >=18](https://img.shields.io/badge/node-%3E%3D18-brightgreen.svg)](package.json)
4
4
 
5
- **Routing instructions and a CLI runner for your AI tools.** One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine. Your primary agent follows the instructions to choose a tier or lane; the runner executes the lane it is given. It does not automatically compare prices or select models.
5
+ **Route every task to the right model, agent or LLM, and spend fewer tokens.** A model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. One installer asks what you have access to and writes routing rules, subagents and a CLI runner for exactly that setup, from one chat app to several agent CLIs or a virtual machine. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models; the runner executes the lane it is given. On Claude Code it also delegates execution to subagents by default, with two hooks that inject the routing table every turn.
6
+
7
+ ## At a glance
8
+
9
+ - **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
10
+ - **What it is not:** a proxy, a gateway or an API router. It does not automatically compare prices or select models; your agent follows the rules and chooses.
11
+ - **Install:** `npx model-orchestrator` (interactive), or headless from a script or an agent: `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator`.
12
+ - **Use it when:** you run more than one model or agent and want each task sent to the smallest one that can do it well.
13
+ - **What it saves:** frontier-model tokens. Bulk work, reading and checks go to fast tiers; the expensive tier is kept for planning and judgment.
14
+ - **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
6
15
 
7
16
  Built from a working system, not a diagram: the routing rules, the protocols and the lane runner here run in production, generalized so they transfer to any stack.
8
17
 
@@ -18,14 +27,14 @@ The installer asks a few things, then writes a folder:
18
27
  2. **Which AIs do you have access to?** (it marks the ones already on your PATH)
19
28
  3. **Which one is your primary agent?** (the one that runs the system)
20
29
 
21
- It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: follow the generated README. Inspect the manifest and remove only the individual managed subagent files you no longer need, preserve edited or pre-existing files, and remove your manually pasted activation block. Never delete a shared subagent folder.
30
+ It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions (and, on Claude Code, two hook scripts) go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: follow the generated README. Inspect the manifest and remove only the individual managed subagent files you no longer need, preserve edited or pre-existing files, and remove your manually pasted activation block. Never delete a shared subagent folder.
22
31
 
23
32
  ## The three levels
24
33
 
25
34
  | Level | You have | You get |
26
35
  |---|---|---|
27
36
  | **1 · Beginner** | one LLM or one agent | tiers, task classification, the two build checkpoints, the protocols (build, propagate, gap analysis, deep research, numbers and logic, memory and record), a task-bundle template, and your agent set up to follow them |
28
- | **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts), a delegation matrix generated from your selection, research triage across the lanes you have |
37
+ | **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts; `--model` / `--effort` to pin the route and log it), a delegation matrix generated from your selection, research triage across the lanes you have |
29
38
  | **3 · Advanced** | a virtual machine | everything above, plus a gateway config rendered from the API keys you hold (asked separately from your CLIs), pinned images, box rules, privacy gates, and a weekly gap-analysis job with "what watches it" written down |
30
39
 
31
40
  Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md).
@@ -63,7 +72,7 @@ An install has two targets, and a scripted run should set both.
63
72
  | Flag | Default | What lands there |
64
73
  |---|---|---|
65
74
  | `--dir` | `./ai-orchestrator` | the docs, protocols and (level 2+) `bin/cli-run.mjs`. Named after what it contains, not after this package, so a project can hold one without looking like a checkout of it. Pass `--dir ./model-orchestrator` if you prefer the package name. |
66
- | `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. |
75
+ | `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets two hook scripts in `.claude/hooks/`, wired by a settings snippet you merge yourself. |
67
76
 
68
77
  `--project` defaulting to the current directory is the one that surprises people: run the command from your home folder with Claude Code as the primary and five agent files land in your home folder. The installer prints the resolved project path in the plan and says when you left it at the default. Set it.
69
78
 
@@ -90,8 +99,10 @@ ai-orchestrator/
90
99
  TASK_BUNDLE.md the brief every delegation carries
91
100
  protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record
92
101
  CODECALC.md OBSIDIAN-TC.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
93
- <project>/.claude/agents/ five subagents, one per tier, at the PROJECT root (if Claude Code is primary)
102
+ <project>/.claude/agents/ one per tier plus finding-verifier, done-verifier, reader, at the PROJECT root (if Claude Code is primary)
103
+ <project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart), Claude Code only
94
104
  CLAUDE.snippet.md the block to paste into your CLAUDE.md
105
+ settings.hooks.snippet.json the hooks block to merge into .claude/settings.json (Claude Code only)
95
106
  ROUTING.md multi-lane decision tree (level 2+)
96
107
  TIERS.md DELEGATION_MATRIX.md RESEARCH_TRIAGE.md CLI-RUN.md
97
108
  bin/cli-run.mjs bin/lanes.json (node bin/cli-run.mjs --doctor is the smoke test)
@@ -159,9 +170,79 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
159
170
  3. **Exit 0 is not a deliverable.** Check for the artifact, not the status line. `cli-run` checks the response is structurally there; `--expect-file` checks the artifact.
160
171
  4. **Numbers are computed, never guessed.** A tool that calculates beats a model that feels finished.
161
172
  5. **A write nobody can find again did not happen.** Search first, keep the index true, one writer.
162
- 6. **The orchestrator owns the main build.** Delegates hold none of your rules; they get bounded sub-parts and a brief.
173
+ 6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, only the brief carries what this task needs. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
163
174
  7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
164
175
 
176
+ ## Routing by role, complexity and risk
177
+
178
+ Role picks the agent. Two more inputs move the choice, and they move it in
179
+ different directions, so `TIERS.md` states them separately rather than folding
180
+ them into the role:
181
+
182
+ - **Complexity moves the effort.** A worker executing a finished plan needs less
183
+ reasoning than the reviewer judging its output. When the plan is airtight the
184
+ spec is carrying the thinking.
185
+ - **Risk moves the tier and the reader.** Security, privacy, data loss and
186
+ irreversible changes buy the attack lane, a named check, a rollback path or a
187
+ human yes. A one-line change to an auth check is simple and high-risk at the
188
+ same time, and it is the risk that decides.
189
+
190
+ The top of the ladder is bought with evidence: a reproduced failure, an
191
+ unresolved checkpoint, an irreversible change. A task that merely feels hard is
192
+ a deep-tier task, not an escalation.
193
+
194
+ ## A finding is a claim, not a fact
195
+
196
+ Review findings do not go straight to a repair. `finding-verifier` reads the
197
+ cited line, states what would trigger the problem, then hunts for the guard,
198
+ caller or test that makes it impossible, and returns **CONFIRMED**,
199
+ **NOT_REPRODUCED** or **INCONCLUSIVE** per finding. Only CONFIRMED earns a
200
+ change. Use a different model family from the one that produced the finding
201
+ where you have one: a family asked to check its own claim tends to agree with
202
+ itself.
203
+
204
+ ## Two more fast-tier checks
205
+
206
+ `done-verifier` probes the artifact a tracker item's done-signal names (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything itself. It carries no file-editing tools, but on claude-code it does carry `Bash` for those probes (`git log`, `grep`, `wc -l`, `test -f`); staying to read-only commands there is a rule in its prompt, not a restriction on the tool grant, and its own description says so. On agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `reader` is the one that is read-only by tool grant on both: no `Write`, `Edit`, or `Bash`. It reads and digests many files or notes and hands back exactly what the brief asked for, cited by `path:line`; it never classifies, tags or writes, which is what separates it from `bulk-worker`. Both ship in the claude-code and agy agent sets, at the fast tier.
207
+
208
+ ## Pin the route, or know that you did not
209
+
210
+ A lane with no `--model`, no `--effort` and no `defaults` entry in
211
+ `bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
212
+ CLI configured months ago at a low reasoning effort keeps auditing at that
213
+ effort while your routing docs describe an adversarial pass.
214
+
215
+ ```bash
216
+ node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
217
+ node bin/cli-run.mjs --doctor # prints what each lane is pinned to, and what is not pinned
218
+ ```
219
+
220
+ Every run logs the model and effort **requested** and where the request came
221
+ from: `flag`, `lanes.json`, or `lane_default`, on every record including the
222
+ runs that never reached a lane. It does not log an actual. One lane of five
223
+ (grok) reports a model id in its own output and the other four report none, so
224
+ an actual field would be present for one lane and missing for four, and it
225
+ would be a provider-supplied string, which the durable log deliberately never
226
+ holds.
227
+
228
+ ## Common questions
229
+
230
+ ### How do I cut token usage across Claude Code, Codex and Gemini?
231
+
232
+ Install for the tools you have, then let the generated `ROUTING.md` decide the tier per task: bulk, reading and verification go to the fast tier or a cheaper CLI lane, and the deep tier only plans and judges. On Claude Code, execution goes to the `builder` subagent by default and the main session plans and verifies. Every lane call through `cli-run` logs the model and effort it ran with, so you can check where the tokens went.
233
+
234
+ ### How do I route tasks to cheaper models?
235
+
236
+ The rules route by role, complexity and risk (see [Routing by role, complexity and risk](#routing-by-role-complexity-and-risk)). Role picks the agent, complexity moves the effort, risk moves the tier. A task a cheap tier finishes correctly never gets a frontier token.
237
+
238
+ ### Is this an LLM router or an AI gateway?
239
+
240
+ No. It routes at the task level, through instructions your agent follows and a runner for agent CLIs. If you want a service or proxy that picks or forwards the model on every API request, look at request-level routers and gateways such as RouteLLM, LiteLLM, OpenRouter or claude-code-router. They solve a different problem and can sit underneath this.
241
+
242
+ ### Can an agent install and run it without a person?
243
+
244
+ Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run` previews the plan, and `--list` prints every supported AI. Nothing is appended to a file you already have; activation snippets are written next to your files for you to merge.
245
+
165
246
  ## Requirements
166
247
 
167
248
  Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows is untested: `cli-run` ends a lane's process tree there with `taskkill`, but nothing in CI runs on Windows, so treat it as unsupported until someone reports otherwise.
@@ -177,6 +258,14 @@ Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box tem
177
258
 
178
259
  Add an AI to `src/catalog.js` and every prompt, table, config and doc picks it up. Run `npm test`. Keep templates free of logic and free of anything that looks like a credential. The rest is in [CONTRIBUTING.md](CONTRIBUTING.md); releases in [RELEASING.md](RELEASING.md); security reports in [SECURITY.md](SECURITY.md).
179
260
 
261
+ ## Credits
262
+
263
+ - [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for
264
+ separating role, complexity and risk instead of compressing them into one
265
+ scale, for recording the model and effort a lane was actually asked for, and
266
+ for verifying findings before they trigger repairs. All three shipped in
267
+ 0.1.14.
268
+
180
269
  ## License
181
270
 
182
271
  [MIT](LICENSE)
package/bin/cli-run.mjs CHANGED
@@ -39,6 +39,17 @@
39
39
  //
40
40
  // The durable log stores a FIXED reason code per run (see REASONS), never a
41
41
  // provider-supplied string. Bounded vendor stderr goes to your terminal only.
42
+ //
43
+ // ROUTE: which model and reasoning effort a lane ran with.
44
+ // A lane with no --model and no lanes.json default inherits whatever its own
45
+ // config file says, which is invisible from here and is how a documented route
46
+ // silently stops being the route that runs. --model / --effort pin it per call,
47
+ // `defaults` in lanes.json pins it per lane, and every run logs the value that
48
+ // was REQUESTED plus where the request came from (flag, lanes.json, or nothing
49
+ // at all). It does not log an "actual". One lane of five (grok) does report a
50
+ // model id in its own output; the other four report none, and a field present
51
+ // for one lane and absent for four is worse than no field. It would also be a
52
+ // provider-supplied string, which this log deliberately never holds.
42
53
 
43
54
  import { spawn } from 'node:child_process';
44
55
  import { StringDecoder } from 'node:string_decoder';
@@ -190,28 +201,71 @@ export function judgeQwen(rc, out) {
190
201
  return pass(text, `subtype=success, totalErrors=0 across ${Object.keys(models).length} model(s)`);
191
202
  }
192
203
 
204
+ // --- route: model and effort per lane --------------------------------------
205
+ // Each vendor spells these differently, and the spelling was read from each
206
+ // CLI's own --help, not remembered. A lane with `effort: null` has no reasoning
207
+ // flag at all; asking for one there is a usage error, never a silent drop.
208
+ // grok -m MODEL --reasoning-effort EFFORT
209
+ // codex -m MODEL -c model_reasoning_effort="EFFORT" (a TOML override, hence the quotes)
210
+ // agy --model M --effort EFFORT (low|medium|high)
211
+ // hermes -m MODEL --reasoning LEVEL (none|minimal|...)
212
+ // qwen -m MODEL no reasoning flag
213
+ export const LANE_FLAGS = {
214
+ grok: { model: (v) => ['-m', v], effort: (v) => ['--reasoning-effort', v] },
215
+ codex: { model: (v) => ['-m', v], effort: (v) => ['-c', `model_reasoning_effort="${v}"`] },
216
+ agy: { model: (v) => ['--model', v], effort: (v) => ['--effort', v] },
217
+ hermes: { model: (v) => ['-m', v], effort: (v) => ['--reasoning', v] },
218
+ qwen: { model: (v) => ['-m', v], effort: null }
219
+ };
220
+
221
+ // A model id or effort level becomes an argv element and, for codex, part of a
222
+ // TOML value. Bounding the charset is what makes both safe: no leading dash (a
223
+ // value cannot become a flag), no quote, space or control character (a value
224
+ // cannot break out of the TOML string), and a length cap so a config file
225
+ // cannot push an unbounded string into the durable log.
226
+ export const ROUTE_VALUE = /^[A-Za-z0-9][A-Za-z0-9._:@/+-]{0,63}$/;
227
+ export function badRouteValue(kind, v) {
228
+ if (typeof v !== 'string' || !ROUTE_VALUE.test(v)) {
229
+ return `--${kind} must be 1 to 64 characters of letters, digits, dot, underscore, colon, at, slash, plus or dash, and may not start with a dash: ${JSON.stringify(v)}`;
230
+ }
231
+ return null;
232
+ }
233
+
193
234
  // --- adapters: build argv for a lane -------------------------------------
235
+ // Route flags go in front of the prompt for every lane, because two lanes
236
+ // (hermes, codex) take the prompt as a positional argument and a flag after it
237
+ // is either ignored or read as part of it.
238
+ function routeFlags(lane, opts) {
239
+ const spec = LANE_FLAGS[lane];
240
+ const out = [];
241
+ if (!spec) return out;
242
+ if (opts.model) out.push(...spec.model(opts.model));
243
+ if (opts.effort && spec.effort) out.push(...spec.effort(opts.effort));
244
+ return out;
245
+ }
246
+
194
247
  export function buildArgv(lane, binary, prompt, opts, tmp) {
195
248
  const timeout = opts.timeout;
249
+ const route = routeFlags(lane, opts);
196
250
  switch (lane) {
197
251
  case 'grok':
198
- return { argv: [binary, '--output-format', 'json', '-p', prompt] };
252
+ return { argv: [binary, '--output-format', 'json', ...route, '-p', prompt] };
199
253
  case 'codex': {
200
254
  const last = join(tmp, 'last.txt');
201
255
  const argv = [binary, 'exec', '--json', '--color', 'never', '--skip-git-repo-check', '-o', last];
202
256
  if (opts.audit) argv.push('--sandbox', 'read-only'); // an audit lane that can write is a bug
257
+ argv.push(...route);
203
258
  argv.push(prompt);
204
259
  return { argv, outFile: last };
205
260
  }
206
261
  case 'agy': {
207
262
  const mins = Math.max(1, Math.round(timeout / 60));
208
- return { argv: [binary, '--print-timeout', `${mins}m`, '--output-format', 'stream-json', '-p', prompt] };
263
+ return { argv: [binary, '--print-timeout', `${mins}m`, '--output-format', 'stream-json', ...route, '-p', prompt] };
209
264
  }
210
265
  case 'hermes':
211
- return { argv: [binary, '-z', prompt, '--usage-file', join(tmp, 'usage.json')] };
266
+ return { argv: [binary, '-z', ...route, prompt, '--usage-file', join(tmp, 'usage.json')] };
212
267
  case 'qwen': {
213
- const argv = [binary, '-o', 'json'];
214
- if (opts.model) argv.push('-m', opts.model);
268
+ const argv = [binary, '-o', 'json', ...route];
215
269
  if (opts.safeMode) argv.push('--safe-mode');
216
270
  argv.push('-p', prompt);
217
271
  return { argv };
@@ -370,26 +424,67 @@ function log(rec) {
370
424
  // documented default). PRESENT BUT UNREADABLE OR MALFORMED = no lane enabled:
371
425
  // a half-written config must fail closed, never re-enable what the installer
372
426
  // disabled. Returns null when the file is bad so the caller can say so.
373
- export function enabledLanes(here = dirname(fileURLToPath(import.meta.url))) {
427
+ export function laneConfig(here = dirname(fileURLToPath(import.meta.url))) {
374
428
  const p = join(here, 'lanes.json');
375
- if (!existsSync(p)) return LANES;
429
+ if (!existsSync(p)) return { enabled: LANES, defaults: {} };
376
430
  try {
377
431
  const j = JSON.parse(readFileSync(p, 'utf8'));
378
432
  if (!j || typeof j !== 'object' || !Array.isArray(j.enabled)) return null;
379
433
  if (!j.enabled.every((l) => typeof l === 'string' && LANES.includes(l))) return null;
380
- return j.enabled;
434
+ // `defaults` pins a model and effort per lane. It is optional; present and
435
+ // malformed fails closed with the rest of the file, because a half-written
436
+ // route is exactly the silent-inheritance problem this field exists to fix.
437
+ const defaults = {};
438
+ if (j.defaults !== undefined) {
439
+ if (!j.defaults || typeof j.defaults !== 'object' || Array.isArray(j.defaults)) return null;
440
+ for (const [lane, d] of Object.entries(j.defaults)) {
441
+ if (!LANES.includes(lane)) return null;
442
+ if (!d || typeof d !== 'object' || Array.isArray(d)) return null;
443
+ const { model, effort, ...rest } = d;
444
+ if (Object.keys(rest).length) return null;
445
+ if (model !== undefined && badRouteValue('model', model)) return null;
446
+ if (effort !== undefined) {
447
+ if (badRouteValue('effort', effort)) return null;
448
+ if (!LANE_FLAGS[lane] || !LANE_FLAGS[lane].effort) return null; // a lane with no reasoning flag cannot have one pinned
449
+ }
450
+ defaults[lane] = { model: model ?? null, effort: effort ?? null };
451
+ }
452
+ }
453
+ return { enabled: j.enabled, defaults };
381
454
  } catch {
382
455
  return null;
383
456
  }
384
457
  }
385
458
 
459
+ // Kept as the narrow question most callers ask. null still means malformed.
460
+ export function enabledLanes(here = dirname(fileURLToPath(import.meta.url))) {
461
+ const c = laneConfig(here);
462
+ return c === null ? null : c.enabled;
463
+ }
464
+
465
+ // Flag beats lanes.json beats nothing. `source` is what makes the log audit-worthy:
466
+ // 'lane_default' means this run inherited the vendor CLI's own config, unseen from here.
467
+ export function resolveRoute(lane, opts, defaults) {
468
+ const d = (defaults && defaults[lane]) || {};
469
+ const model = opts.model ?? d.model ?? null;
470
+ const effort = opts.effort ?? d.effort ?? null;
471
+ const src = (flag, def) => (flag != null ? 'flag' : def != null ? 'lanes.json' : 'lane_default');
472
+ return { model, effort, model_source: src(opts.model, d.model), effort_source: src(opts.effort, d.effort) };
473
+ }
474
+
386
475
  function usage(msg) {
387
476
  if (msg) console.error('cli-run: ' + msg);
388
477
  console.error(`usage: cli-run <${LANES.join('|')}> "<prompt>" [--brief FILE] [--timeout SECS] [--quiet]
389
- [--expect-file PATH] [--expect-json]
478
+ [--model ID] [--effort LEVEL] [--expect-file PATH] [--expect-json]
390
479
  cli-run codex --audit "<prompt>" read-only sandbox (audit shape)
391
- cli-run qwen [--model ID] [--safe-mode] "<prompt>"
392
- cli-run --doctor [--run] enabled lanes, binaries on PATH; --run sends each a tiny prompt`);
480
+ cli-run qwen [--safe-mode] "<prompt>" qwen-only flag
481
+ cli-run --doctor [--run] enabled lanes, binaries, and the route each one is pinned to
482
+
483
+ --model / --effort pin what a lane runs with, instead of letting it inherit its
484
+ own config. Every lane takes --model; every lane except qwen takes --effort.
485
+ Levels are the vendor's own (agy low|medium|high, hermes none|minimal|...): an
486
+ unknown level is rejected by the lane, and reported as that lane's exit code.
487
+ Pin them per lane instead of per call with "defaults" in bin/lanes.json.`);
393
488
  return USAGE;
394
489
  }
395
490
 
@@ -416,11 +511,12 @@ function installedPrimary(here = dirname(fileURLToPath(import.meta.url))) {
416
511
 
417
512
  // --doctor: the first thing to run after install.
418
513
  export async function doctor(run) {
419
- const enabled = enabledLanes();
420
- if (enabled === null) {
514
+ const cfg = laneConfig();
515
+ if (cfg === null) {
421
516
  console.error('doctor: lanes.json exists but is malformed; fix it first');
422
517
  return USAGE;
423
518
  }
519
+ const { enabled, defaults } = cfg;
424
520
  let bad = 0;
425
521
  console.log(`doctor: ${enabled.length} enabled lane(s): ${enabled.join(', ') || 'none'}`);
426
522
  const primary = installedPrimary();
@@ -432,7 +528,11 @@ export async function doctor(run) {
432
528
  for (const lane of LANES) {
433
529
  const on = enabled.includes(lane);
434
530
  const bin = which(lane);
435
- let line = ` ${lane.padEnd(7)} ${on ? 'enabled ' : 'disabled'} ${bin ? 'binary ok' : 'binary MISSING'}`;
531
+ const d = defaults[lane] || {};
532
+ // A disabled lane has no route worth reporting; saying "not pinned" there
533
+ // reads as a finding about a lane that is not going to run.
534
+ const route = !on ? '' : d.model || d.effort ? `route ${d.model || 'lane default'}/${d.effort || 'lane default'}` : 'route not pinned (inherits the lane\'s own config)';
535
+ let line = ` ${lane.padEnd(7)} ${on ? 'enabled ' : 'disabled'} ${bin ? 'binary ok' : 'binary MISSING'}${route ? ' ' + route : ''}`;
436
536
  if (on && !bin) bad++;
437
537
  if (on && bin && run) {
438
538
  const rc = await main([lane, 'Reply with exactly the word OK and nothing else.', '--timeout', '120', '--quiet']);
@@ -443,6 +543,7 @@ export async function doctor(run) {
443
543
  }
444
544
  console.log(bad ? `doctor: ${bad} problem(s)` : 'doctor: all enabled lanes ' + (run ? 'answered' : 'present'));
445
545
  console.log('doctor checks presence and, with --run, a one-word canary. It does not check vendor versions.');
546
+ console.log('"route not pinned" means that lane runs on whatever its own config file says, which this tool cannot see. Pin it in lanes.json "defaults" if the route matters.');
446
547
  return bad ? NO_DELIVERABLE : OK;
447
548
  }
448
549
 
@@ -485,10 +586,10 @@ export function checkContracts(opts, text, before) {
485
586
  }
486
587
 
487
588
  export async function main(argv) {
488
- const VALUE = new Set(['--brief', '--timeout', '--model', '--expect-file']);
589
+ const VALUE = new Set(['--brief', '--timeout', '--model', '--effort', '--expect-file']);
489
590
  const BOOL = new Set(['--quiet', '--audit', '--safe-mode', '--doctor', '--run', '--expect-json']);
490
591
  const args = [...argv];
491
- const opts = { timeout: 900, quiet: false, audit: false, model: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
592
+ const opts = { timeout: 900, quiet: false, audit: false, model: null, effort: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
492
593
  const positional = [];
493
594
  while (args.length) {
494
595
  const a = args.shift();
@@ -498,6 +599,7 @@ export async function main(argv) {
498
599
  if (a === '--brief') opts.brief = v;
499
600
  else if (a === '--timeout') opts.timeout = Number(v);
500
601
  else if (a === '--expect-file') opts.expectFile = v;
602
+ else if (a === '--effort') opts.effort = v;
501
603
  else opts.model = v;
502
604
  } else if (BOOL.has(a)) {
503
605
  if (a === '--quiet') opts.quiet = true;
@@ -530,11 +632,33 @@ export async function main(argv) {
530
632
  if (!prompt) return usage('give a prompt or --brief FILE');
531
633
  if (!Number.isFinite(opts.timeout) || opts.timeout <= 0) return usage('--timeout must be a positive number of seconds');
532
634
  if (opts.audit && lane !== 'codex') return usage('--audit is codex-only');
533
- if ((opts.model || opts.safeMode) && lane !== 'qwen') return usage('--model and --safe-mode are qwen-only');
635
+ if (opts.safeMode && lane !== 'qwen') return usage('--safe-mode is qwen-only');
636
+ for (const [kind, v] of [['model', opts.model], ['effort', opts.effort]]) {
637
+ if (v == null) continue;
638
+ const bad = badRouteValue(kind, v);
639
+ if (bad) return usage(bad);
640
+ }
641
+ // qwen has no reasoning flag. Dropping --effort silently would leave the caller
642
+ // believing a route that never happened, which is the defect this feature fixes.
643
+ if (opts.effort && !(LANE_FLAGS[lane] && LANE_FLAGS[lane].effort)) return usage(`${lane} has no reasoning-effort flag; --effort is not available on this lane`);
534
644
 
535
645
  const digest = createHash('sha256').update(prompt).digest('hex').slice(0, 12);
536
646
  const base = { lane, prompt_sha256_12: digest, prompt_chars: prompt.length };
537
- const enabled = enabledLanes();
647
+ const cfg = laneConfig();
648
+ // Resolve the route BEFORE the refusals below. A run that never reached a lane
649
+ // was still a request for one, and a failure record with no route is the exact
650
+ // gap this feature exists to close. A malformed lanes.json has no usable
651
+ // defaults, so the flags stand alone and say so.
652
+ const route = resolveRoute(lane, opts, cfg === null ? {} : cfg.defaults);
653
+ opts.model = route.model;
654
+ opts.effort = route.effort;
655
+ Object.assign(base, {
656
+ model_requested: route.model,
657
+ effort_requested: route.effort,
658
+ model_source: route.model_source,
659
+ effort_source: route.effort_source
660
+ });
661
+ const enabled = cfg === null ? null : cfg.enabled;
538
662
  if (enabled === null) {
539
663
  console.error('cli-run: lanes.json exists but is not a valid {"enabled": [...]} file; refusing every lane until it is fixed');
540
664
  log({ ...base, verdict: 'unavailable', rc: UNAVAILABLE, reason: 'lanes_json_malformed' });
@@ -601,7 +725,8 @@ export async function main(argv) {
601
725
  }
602
726
  }
603
727
  if (text && code === OK) process.stdout.write(text + '\n');
604
- if (!opts.quiet) console.error(`cli-run[${lane}] ${verdict} rc=${code} ${r.seconds.toFixed(1)}s raw=${r.outBytes || 0}B :: ${detail}`);
728
+ const routeNote = route.model || route.effort ? `${route.model || 'lane default'}/${route.effort || 'lane default'}` : 'lane default';
729
+ if (!opts.quiet) console.error(`cli-run[${lane}] ${verdict} rc=${code} ${r.seconds.toFixed(1)}s raw=${r.outBytes || 0}B route=${routeNote} :: ${detail}`);
605
730
  // Durable log: fixed reason code and structural numbers only.
606
731
  log({ ...base, verdict, rc: code, cli_rc: r.status, signal: r.signal || null, seconds: Math.round(r.seconds * 100) / 100, raw_bytes: r.outBytes || 0, deliverable_bytes: Buffer.byteLength(text), reason: REASONS.has(reason) ? reason : 'unknown' });
607
732
  return code;
package/bin/cli.js CHANGED
@@ -144,7 +144,7 @@ async function main() {
144
144
  // Says what this generates, not what it guarantees. The old line promised
145
145
  // routing this package does not perform: lane choice is an instruction an
146
146
  // agent follows, never something enforced here (#11).
147
- console.log('\nmodel-orchestrator\nRouting instructions and a CLI runner for the AIs you actually have.\n');
147
+ console.log('\nmodel-orchestrator\nA model orchestrator: routing rules and a CLI runner for the AIs you actually have.\n');
148
148
 
149
149
  // 1. Level
150
150
  let level = Number(opt('level'));
@@ -81,3 +81,24 @@ Re-audit the same scope. Every round-1 finding was reproduced before it was touc
81
81
  Also new since round 1: the companion-tool path (`--tools codecalc`, `--no-tools`, `templates/tools/codecalc/`, `protocols/numbers-and-logic.md`, `resolveTools`). Attack it the same way: unknown tool ids, interaction with `--yes`, the extra interactive question, and whether any written snippet could be confused for a file the installer should not touch.
82
82
 
83
83
  Report only what reproduces on the current tree. `CLEAN` per area is expected where the fix holds.
84
+
85
+ ## New in 0.1.15: two claude-code-only hooks
86
+
87
+ `route-gate.mjs` (`UserPromptSubmit`) and `subagent-context.mjs` (`SubagentStart`) ship to `.claude/hooks/` only when claude-code is the primary. Both are plain Node, zero deps, and installed with mode `0o755`.
88
+
89
+ - **Reads.** `route-gate.mjs` reads at most 64 KB from one file: the routing rules file (`ROUTING.md` or `ORCHESTRATOR.md`) at a path rendered in at install time relative to `CLAUDE_PROJECT_DIR`, never a hardcoded absolute path. Before opening it, it `statSync`s the resolved path (following a symlink to its target) and refuses anything that is not `isFile()`, a FIFO, socket, device or directory included, so the read never touches a path that could block on open. The read itself is one `openSync` + one bounded `readSync` into a fixed 64 KB buffer, closed in a `finally`, so neither the time nor the memory this hook uses depends on how large the file on disk actually is. It then extracts the text between `<!-- route-gate:start -->` and `<!-- route-gate:end -->` and nothing else. `subagent-context.mjs` reads nothing from disk; its context is static text plus the same two rendered paths. Neither parses or executes anything it reads; the extracted block is passed through as a string.
90
+ - **Stdin.** Neither hook uses a field from the JSON input Claude Code sends on stdin, but both must still consume the pipe rather than ignore it. Both drain stdin asynchronously against a 250ms hard cap: whichever comes first, the real `end` event or the timeout, the hook proceeds. Neither ever calls a blocking, synchronous read of stdin.
91
+ - **Writes.** Neither writes a file. Both write one JSON object to stdout: `{"hookSpecificOutput":{"hookEventName":"...","additionalContext":"..."}}`, and both exit only after that write's callback fires, so a buffered write to a pipe is not truncated by an exit racing ahead of it.
92
+ - **Fail-open, on purpose.** A missing `CLAUDE_PROJECT_DIR`, a missing rules file, a non-regular file at the rules path, or a missing block each produce a one-line fallback `additionalContext` naming what was found, and the script still exits 0. This is acceptable because a miss here is a stray context string reaching the model, not a security gate: nothing downstream trusts the hook's output for anything but a routing suggestion, and the settings snippet that wires it in is a document the user merges by hand, never written automatically over an existing `settings.json`.
93
+ - **Bounded.** `route-gate.mjs` caps the read at 64 KB regardless of the file's reported size and the injected string at 4000 characters, so a rules file bloated by a bad edit, or truncated to an arbitrary length, cannot balloon the context or the read time on every turn. Both hooks cap stdin drain at 250ms.
94
+ - **Not yet attacked.** Untested here: a rules file with a `route-gate:start` marker but no matching end marker very far into the file (bounded by `MAX_READ`, so the end marker past that point is treated as absent, which is the intended fail-open path, but worth a deliberate case); a `CLAUDE_PROJECT_DIR` pointing at a path with no read permission; behavior under the Windows exec-form `node` + `args` invocation named in the settings snippet; a `statSync` that itself hangs (a stalled network filesystem, for instance) rather than the FIFO-at-open case this round fixed.
95
+
96
+ ### Round 1 (pre-release), fixed before shipping
97
+
98
+ Three findings reproduced against the 0.1.15 branch before it shipped, none of them ever released:
99
+
100
+ | # | Finding | Fix |
101
+ |---|---|---|
102
+ | 1 | HIGH. `readFileSync(0)` in both hooks blocked until stdin reached EOF (`sleep 3 \| ... node route-gate.mjs` still running past 1.5s); `route-gate.mjs` also read the whole rules file into memory before bounding it, so a FIFO planted at the rules path blocked forever on open. | Stdin is drained asynchronously against a 250ms hard cap in both hooks. `route-gate.mjs` refuses anything that is not `isFile()` via `statSync` before ever calling open, then reads through one fixed 64 KB buffer via `openSync`/`readSync`. Tests: an open, never-closed stdin pipe exits within 1s for both hooks; a FIFO at the rules path returns the fallback instead of hanging; a 200 MB sparse rules file completes in well under a second with output still capped. |
103
+ | 2 | MEDIUM. `done-verifier`'s description, both agent-folder READMEs, and the root README called it "read-only" without qualification, while its claude-code file carries an unrestricted `Bash` grant; nothing in that grant stops it from running a mutating command. | Every one of those surfaces now says plainly that `done-verifier` carries no file-editing tools and that its Bash use is bound by its own prompt, not by the tool grant; `reader` is named as the one that is read-only by tool grant (no Bash) on both formats. |
104
+ | 3 | MEDIUM. Three generated surfaces still stated the pre-0.1.15 premise on a claude-code install: `builder.md`'s description ("... or the main build itself"), `build-protocol.md`'s roles table and its "why the builder does not hand off" note, and `ROUTING.md`'s "Plan big, execute small" line ("the orchestrator executes"). | All three now render through `subagentsLoadRules(primary)`, the same gate the decision tree and "Who builds" already used; every other primary is unchanged. A semantic-regression test asserts a claude-code install contains none of the old phrasing and a codex install still does. |
@@ -14,6 +14,8 @@ If your agent exposes model choice (Claude Code, Codex, Antigravity), map the ti
14
14
 
15
15
  Three cost levers, always together: tier (price per token), token discipline (how many tokens: read only what you will touch, never re-read, deliverables not narration), effort (how hard each call thinks).
16
16
 
17
+ And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **Risk** moves the tier and who reads the result: security, privacy, data loss and irreversible changes are the four worth naming, because none of their failures can be fixed by editing the code afterwards. A one-line change to an auth check is simple and high-risk at once, and it is the risk that decides.
18
+
17
19
  Robustness first, cost second. You split tiers because the split produces better work.
18
20
 
19
21
  ## 2. Classify every task, first match wins
@@ -36,7 +38,7 @@ Every build gets two checkpoints. **Before writing:** you map what it touches an
36
38
 
37
39
  ## 4. Every hand-off carries a brief
38
40
 
39
- A subagent, a fresh chat, a second window holds none of your rules and reads an unspecified edge as an open one. The brief: purpose, task class, granted scope, capabilities, denied actions, conventions it does not have, report contract (what was not done, what is unverified), exit parameters (when to stop, and what to return if it hits the bound). Absence is denial.
41
+ A subagent, a fresh chat, a second window may hold none of your rules, and that is the default to assume. The one documented exception is a Claude Code subagent: it loads the project's CLAUDE.md hierarchy at start, so it keeps the standing rules, just not this task's scope. Either way, it reads an unspecified edge as an open one. The brief: purpose, task class, granted scope, capabilities, denied actions, conventions it does not have, report contract (what was not done, what is unverified), exit parameters (when to stop, and what to return if it hits the bound). Absence is denial.
40
42
 
41
43
  ## 5. The second pass
42
44