model-orchestrator 0.1.14 → 0.1.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md ADDED
@@ -0,0 +1,26 @@
1
+ # AGENTS.md
2
+
3
+ Two audiences: an agent that wants to USE this package for a project, and an agent that is working ON this repository.
4
+
5
+ ## Using this package from an agent
6
+
7
+ model-orchestrator writes routing rules, subagents and a CLI lane runner so an agent sends each task to the right model, subagent or CLI and spends fewer frontier tokens. It is not a proxy or gateway. Headless use:
8
+
9
+ - `npx model-orchestrator --list` prints every supported AI id.
10
+ - `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project <repo> --dir <repo>/ai-orchestrator --dry-run` prints the plan and writes nothing.
11
+ - Drop `--dry-run` to write it. Existing files are never overwritten without `--force`; activation snippets (for example `CLAUDE.snippet.md`, `settings.hooks.snippet.json`) are written for a person or agent to merge.
12
+ - The generated `README.md` in `--dir` lists what to copy where and one smoke command to prove the rules took.
13
+ - A summary for LLMs, with links to every doc: [`llms.txt`](llms.txt).
14
+
15
+ ## Working on this repository
16
+
17
+ The files the installer writes for end users live under `templates/`.
18
+
19
+ - Read `CONTRIBUTING.md` first, then `src/README.md` (the catalog drives everything) and `docs/audit-brief.md` (the threat model and what has already been attacked).
20
+ - Run `npm test` before proposing a change and quote the count and the exit code; the suite prints the current number.
21
+ - Everything renders from `src/catalog.js`. Add an AI or a tool there, not in a template. Templates carry no logic.
22
+ - Never put a value that looks like a credential anywhere in this repo, including tests and examples. Environment variable names only.
23
+ - `bin/cli.js` writes only inside `--dir` and `--project`, never over a document without `--force`, and never runs a vendor script. A change that weakens any of those will be refused in review; the tests that hold them are in `test/install.test.js` and `test/cli.test.js`.
24
+ - `bin/cli-run.mjs` must exit non-zero when a lane produced nothing. Every judge has a red case in `test/judges.test.js`; add one before you change a judge.
25
+ - Prose in this repo uses no em dashes (`test/prose.test.js` enforces it).
26
+ - Why these rules exist: each one is the fix for a failure that reached an audit or CI. `CHANGELOG.md` names the issue behind each.
package/CHANGELOG.md CHANGED
@@ -4,7 +4,28 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
- ## [0.1.14] - 2026-09-09
7
+ ## [0.1.15] - 2026-09-10
8
+
9
+ The portable parts of a live routing revision, delegate by default, gated on one verified fact rather than a guess: [code.claude.com/docs/en/sub-agents](https://code.claude.com/docs/en/sub-agents) states that a non-fork Claude Code subagent's initial context includes "every level of the CLAUDE.md hierarchy the main conversation loads", and that the built-in Explore and Plan agents skip it. No other lane in this catalog has that documented, so everything below is gated on `subagentsLoadRules(primary)`, currently true for claude-code alone; every other primary keeps its original wording unchanged.
10
+
11
+ ### Added
12
+
13
+ - **`subagentsLoadRules: true` on the claude-code catalog entry**, with the doc quote as its comment. Drives every new render var below through `src/install.js`; nothing here is a template branch, per the house rule that templates carry no logic.
14
+ - **Builder executes by default, on claude-code.** `ROUTING.md` rule 5, its "Who builds" section, the "Add an endpoint" example, and the claude-code `CLAUDE.snippet.md` now say: the orchestrator plans, briefs, verifies and talks to the human; it stays inline only when (a) the brief would cost as much as the work, (b) the task needs this conversation's own context, or (c) it is the human's decision or the final verification of delegated work. Every other primary keeps "the orchestrator builds it directly."
15
+ - **Two hooks, claude-code only: `route-gate.mjs` (`UserPromptSubmit`) and `subagent-context.mjs` (`SubagentStart`)**, written to `.claude/hooks/`, wired by a `settings.hooks.snippet.json` the user merges into `.claude/settings.json` themselves, never written over one they have. `route-gate.mjs` reads the `<!-- route-gate:start -->...<!-- route-gate:end -->` block out of the rendered routing rules file and injects it every turn, so the table is read from the one place it is generated, not recited from memory; a missing file or block still exits 0 with a one-line fallback naming the path it looked for. `subagent-context.mjs` injects a static reminder of where the rules and `TASK_BUNDLE.md` live and that a delegate does not route further or verify its own work as final. Both are plain Node, zero deps, bounded reads, fail-open by design (a miss is a stray context string, not a gate): see `docs/audit-brief.md`.
16
+ - **`done-verifier` and `reader`, two new fast-tier agents with no file-editing tools, in both the claude-code and agy formats.** `done-verifier` probes a tracker item's stated done-signal (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything. `reader` reads and digests many files or notes and returns exactly what the brief asks (facts, quotes cited `path:line`, an index, a digest); unlike `bulk-worker`, it never classifies, tags, transforms or writes. `reader` is read-only by tool grant on both formats (no `Bash`); `done-verifier` on claude-code carries `Bash` for its probes, bound only by its prompt, not by the grant, and its description says so; on agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `ROUTING.md`'s decision tree, `TIERS.md`'s effort table, and both agent-folder READMEs now name them.
17
+ - **The inline threshold gets a measurement instead of a guessed figure.** "A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Measure yours once: spawn a subagent with a one-line task and read its token count. Work smaller than that stays inline." (claude-code `ROUTING.md` and `CLAUDE.snippet.md` only; no private number shipped.)
18
+
19
+ ### Changed
20
+
21
+ - **The package now says what it is for in the first line people and agents read.** npm search, GitHub search and the installer banner showed "Routing instructions and a CLI runner", which named the parts and not the purpose. The description, the README opening and the banner now lead with the goal (each task to the right model, agent or LLM, fewer frontier tokens) while keeping the 0.1.11 correction intact: routing is an instruction your agent follows, and the README still states it does not automatically compare prices or select models. A test holds all three surfaces to that. New: an "At a glance" block and question-shaped "Common questions" in the README, request-level alternatives named for readers who want a proxy, `llms.txt` at the root and a headless-use section in `AGENTS.md` (both now ship in the package), and search keywords matching what comparable routers use.
22
+ - **Corrected the unqualified premise "a subagent holds none of these rules" everywhere it appeared** (`TASK_BUNDLE.md`, `ORCHESTRATOR.md`, the claude-code snippet, `docs/part-1-beginner.md`, `docs/part-2-intermediate.md`, README principle 6, and `ROUTING.md`'s "Who builds"). The corrected fact: a Claude Code subagent loads CLAUDE.md and so keeps the standing rules, just not this task's scope; a second CLI or a fresh chat window may still hold none of it. "Absence is denial" is unchanged; only the premise about who is absent what was wrong.
23
+ - **The claude-code snippet's closing "available as ..." agent list is generated from the files actually shipped in `templates/agents/claude-code/`, never hand-typed.** It had drifted once already: `finding-verifier` shipped in 0.1.14 and was missing from this sentence until now. `claudeAgentIds()` in `src/install.js` reads the folder; a test ties the rendered list to it.
24
+ - **The delegate-by-default gate now reaches every generated surface it should, not just three of them.** `build-protocol.md`'s "Roles, as capabilities" table and its "Why the builder does not hand off the main build" line, `builder.md`'s description, and `ROUTING.md`'s "Plan big, execute small" modifier still said, on a claude-code install, that the orchestrator writes the main build itself, never hands it off whole, and that a delegate inherits none of the session's rules: the exact premise the rest of this release corrects. All four now render through `subagentsLoadRules(primary)` the same way the decision tree and "Who builds" already did; every other primary is unchanged. A semantic-regression test asserts a claude-code install contains none of the old phrasing and a codex install still does.
25
+
26
+ ### Fixed
27
+
28
+ - **Both new hooks could hang, and `route-gate.mjs` could read an unbounded or blocking file (pre-release audit finding, never shipped).** `readFileSync(0)` in both `route-gate.mjs` and `subagent-context.mjs` blocked until stdin reached EOF, so a caller that piped input in without closing its end (or ran the hook from a bare TTY) left the process running indefinitely; reproduced with `sleep 3 | CLAUDE_PROJECT_DIR=... node route-gate.mjs` still running past 1.5s. Separately, `route-gate.mjs` read the whole rules file into memory before bounding it (`readFileSync(path).slice(0, MAX_READ)`), so a FIFO planted at the rules path blocked forever on open, and a very large file was read in full before being truncated. Fixed in both hooks: stdin is now drained asynchronously against a 250ms hard cap, never blocking past it. `route-gate.mjs` additionally `statSync`s the resolved path and refuses anything that is not `isFile()` (a FIFO, socket, device or directory, symlink target included) before ever calling open, then reads through a single fixed 64 KB buffer via `openSync`/`readSync`, closed in a `finally`, so neither the read time nor the memory used depends on the file's on-disk size. Tests: an open, never-closed stdin pipe now exits within 1s for both hooks; a FIFO at the rules path returns the fallback instead of hanging; a 200 MB sparse rules file completes in well under a second with output still capped.
8
29
 
9
30
  Three refinements to the routing model, from a review by [@shawnwows](https://x.com/shawnwows). The theme is the same in all three: a routing decision that was implied, inherited or asserted is now stated, pinned or checked.
10
31
 
@@ -217,7 +238,8 @@ First release.
217
238
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
218
239
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
219
240
 
220
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...HEAD
241
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.15...HEAD
242
+ [0.1.15]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...v0.1.15
221
243
  [0.1.14]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...v0.1.14
222
244
  [0.1.13]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.12...v0.1.13
223
245
  [0.1.12]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.11...v0.1.12
package/README.md CHANGED
@@ -2,7 +2,16 @@
2
2
 
3
3
  [![npm](https://img.shields.io/npm/v/model-orchestrator.svg)](https://www.npmjs.com/package/model-orchestrator) [![test](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml/badge.svg)](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml) [![license: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![node >=18](https://img.shields.io/badge/node-%3E%3D18-brightgreen.svg)](package.json)
4
4
 
5
- **Routing instructions and a CLI runner for your AI tools.** One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine. Your primary agent follows the instructions to choose a tier or lane; the runner executes the lane it is given. It does not automatically compare prices or select models.
5
+ **Route every task to the right model, agent or LLM, and spend fewer tokens.** A model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. One installer asks what you have access to and writes routing rules, subagents and a CLI runner for exactly that setup, from one chat app to several agent CLIs or a virtual machine. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models; the runner executes the lane it is given. On Claude Code it also delegates execution to subagents by default, with two hooks that inject the routing table every turn.
6
+
7
+ ## At a glance
8
+
9
+ - **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
10
+ - **What it is not:** a proxy, a gateway or an API router. It does not automatically compare prices or select models; your agent follows the rules and chooses.
11
+ - **Install:** `npx model-orchestrator` (interactive), or headless from a script or an agent: `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator`.
12
+ - **Use it when:** you run more than one model or agent and want each task sent to the smallest one that can do it well.
13
+ - **What it saves:** frontier-model tokens. Bulk work, reading and checks go to fast tiers; the expensive tier is kept for planning and judgment.
14
+ - **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
6
15
 
7
16
  Built from a working system, not a diagram: the routing rules, the protocols and the lane runner here run in production, generalized so they transfer to any stack.
8
17
 
@@ -18,7 +27,7 @@ The installer asks a few things, then writes a folder:
18
27
  2. **Which AIs do you have access to?** (it marks the ones already on your PATH)
19
28
  3. **Which one is your primary agent?** (the one that runs the system)
20
29
 
21
- It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: follow the generated README. Inspect the manifest and remove only the individual managed subagent files you no longer need, preserve edited or pre-existing files, and remove your manually pasted activation block. Never delete a shared subagent folder.
30
+ It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions (and, on Claude Code, two hook scripts) go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: follow the generated README. Inspect the manifest and remove only the individual managed subagent files you no longer need, preserve edited or pre-existing files, and remove your manually pasted activation block. Never delete a shared subagent folder.
22
31
 
23
32
  ## The three levels
24
33
 
@@ -63,7 +72,7 @@ An install has two targets, and a scripted run should set both.
63
72
  | Flag | Default | What lands there |
64
73
  |---|---|---|
65
74
  | `--dir` | `./ai-orchestrator` | the docs, protocols and (level 2+) `bin/cli-run.mjs`. Named after what it contains, not after this package, so a project can hold one without looking like a checkout of it. Pass `--dir ./model-orchestrator` if you prefer the package name. |
66
- | `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. |
75
+ | `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets two hook scripts in `.claude/hooks/`, wired by a settings snippet you merge yourself. |
67
76
 
68
77
  `--project` defaulting to the current directory is the one that surprises people: run the command from your home folder with Claude Code as the primary and five agent files land in your home folder. The installer prints the resolved project path in the plan and says when you left it at the default. Set it.
69
78
 
@@ -90,8 +99,10 @@ ai-orchestrator/
90
99
  TASK_BUNDLE.md the brief every delegation carries
91
100
  protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record
92
101
  CODECALC.md OBSIDIAN-TC.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
93
- <project>/.claude/agents/ six subagents, one per tier plus finding-verifier, at the PROJECT root (if Claude Code is primary)
102
+ <project>/.claude/agents/ one per tier plus finding-verifier, done-verifier, reader, at the PROJECT root (if Claude Code is primary)
103
+ <project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart), Claude Code only
94
104
  CLAUDE.snippet.md the block to paste into your CLAUDE.md
105
+ settings.hooks.snippet.json the hooks block to merge into .claude/settings.json (Claude Code only)
95
106
  ROUTING.md multi-lane decision tree (level 2+)
96
107
  TIERS.md DELEGATION_MATRIX.md RESEARCH_TRIAGE.md CLI-RUN.md
97
108
  bin/cli-run.mjs bin/lanes.json (node bin/cli-run.mjs --doctor is the smoke test)
@@ -159,7 +170,7 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
159
170
  3. **Exit 0 is not a deliverable.** Check for the artifact, not the status line. `cli-run` checks the response is structurally there; `--expect-file` checks the artifact.
160
171
  4. **Numbers are computed, never guessed.** A tool that calculates beats a model that feels finished.
161
172
  5. **A write nobody can find again did not happen.** Search first, keep the index true, one writer.
162
- 6. **The orchestrator owns the main build.** Delegates hold none of your rules; they get bounded sub-parts and a brief.
173
+ 6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, only the brief carries what this task needs. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
163
174
  7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
164
175
 
165
176
  ## Routing by role, complexity and risk
@@ -190,6 +201,10 @@ change. Use a different model family from the one that produced the finding
190
201
  where you have one: a family asked to check its own claim tends to agree with
191
202
  itself.
192
203
 
204
+ ## Two more fast-tier checks
205
+
206
+ `done-verifier` probes the artifact a tracker item's done-signal names (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything itself. It carries no file-editing tools, but on claude-code it does carry `Bash` for those probes (`git log`, `grep`, `wc -l`, `test -f`); staying to read-only commands there is a rule in its prompt, not a restriction on the tool grant, and its own description says so. On agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `reader` is the one that is read-only by tool grant on both: no `Write`, `Edit`, or `Bash`. It reads and digests many files or notes and hands back exactly what the brief asked for, cited by `path:line`; it never classifies, tags or writes, which is what separates it from `bulk-worker`. Both ship in the claude-code and agy agent sets, at the fast tier.
207
+
193
208
  ## Pin the route, or know that you did not
194
209
 
195
210
  A lane with no `--model`, no `--effort` and no `defaults` entry in
@@ -210,6 +225,24 @@ an actual field would be present for one lane and missing for four, and it
210
225
  would be a provider-supplied string, which the durable log deliberately never
211
226
  holds.
212
227
 
228
+ ## Common questions
229
+
230
+ ### How do I cut token usage across Claude Code, Codex and Gemini?
231
+
232
+ Install for the tools you have, then let the generated `ROUTING.md` decide the tier per task: bulk, reading and verification go to the fast tier or a cheaper CLI lane, and the deep tier only plans and judges. On Claude Code, execution goes to the `builder` subagent by default and the main session plans and verifies. Every lane call through `cli-run` logs the model and effort it ran with, so you can check where the tokens went.
233
+
234
+ ### How do I route tasks to cheaper models?
235
+
236
+ The rules route by role, complexity and risk (see [Routing by role, complexity and risk](#routing-by-role-complexity-and-risk)). Role picks the agent, complexity moves the effort, risk moves the tier. A task a cheap tier finishes correctly never gets a frontier token.
237
+
238
+ ### Is this an LLM router or an AI gateway?
239
+
240
+ No. It routes at the task level, through instructions your agent follows and a runner for agent CLIs. If you want a service or proxy that picks or forwards the model on every API request, look at request-level routers and gateways such as RouteLLM, LiteLLM, OpenRouter or claude-code-router. They solve a different problem and can sit underneath this.
241
+
242
+ ### Can an agent install and run it without a person?
243
+
244
+ Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run` previews the plan, and `--list` prints every supported AI. Nothing is appended to a file you already have; activation snippets are written next to your files for you to merge.
245
+
213
246
  ## Requirements
214
247
 
215
248
  Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows is untested: `cli-run` ends a lane's process tree there with `taskkill`, but nothing in CI runs on Windows, so treat it as unsupported until someone reports otherwise.
package/bin/cli.js CHANGED
@@ -144,7 +144,7 @@ async function main() {
144
144
  // Says what this generates, not what it guarantees. The old line promised
145
145
  // routing this package does not perform: lane choice is an instruction an
146
146
  // agent follows, never something enforced here (#11).
147
- console.log('\nmodel-orchestrator\nRouting instructions and a CLI runner for the AIs you actually have.\n');
147
+ console.log('\nmodel-orchestrator\nA model orchestrator: routing rules and a CLI runner for the AIs you actually have.\n');
148
148
 
149
149
  // 1. Level
150
150
  let level = Number(opt('level'));
@@ -81,3 +81,24 @@ Re-audit the same scope. Every round-1 finding was reproduced before it was touc
81
81
  Also new since round 1: the companion-tool path (`--tools codecalc`, `--no-tools`, `templates/tools/codecalc/`, `protocols/numbers-and-logic.md`, `resolveTools`). Attack it the same way: unknown tool ids, interaction with `--yes`, the extra interactive question, and whether any written snippet could be confused for a file the installer should not touch.
82
82
 
83
83
  Report only what reproduces on the current tree. `CLEAN` per area is expected where the fix holds.
84
+
85
+ ## New in 0.1.15: two claude-code-only hooks
86
+
87
+ `route-gate.mjs` (`UserPromptSubmit`) and `subagent-context.mjs` (`SubagentStart`) ship to `.claude/hooks/` only when claude-code is the primary. Both are plain Node, zero deps, and installed with mode `0o755`.
88
+
89
+ - **Reads.** `route-gate.mjs` reads at most 64 KB from one file: the routing rules file (`ROUTING.md` or `ORCHESTRATOR.md`) at a path rendered in at install time relative to `CLAUDE_PROJECT_DIR`, never a hardcoded absolute path. Before opening it, it `statSync`s the resolved path (following a symlink to its target) and refuses anything that is not `isFile()`, a FIFO, socket, device or directory included, so the read never touches a path that could block on open. The read itself is one `openSync` + one bounded `readSync` into a fixed 64 KB buffer, closed in a `finally`, so neither the time nor the memory this hook uses depends on how large the file on disk actually is. It then extracts the text between `<!-- route-gate:start -->` and `<!-- route-gate:end -->` and nothing else. `subagent-context.mjs` reads nothing from disk; its context is static text plus the same two rendered paths. Neither parses or executes anything it reads; the extracted block is passed through as a string.
90
+ - **Stdin.** Neither hook uses a field from the JSON input Claude Code sends on stdin, but both must still consume the pipe rather than ignore it. Both drain stdin asynchronously against a 250ms hard cap: whichever comes first, the real `end` event or the timeout, the hook proceeds. Neither ever calls a blocking, synchronous read of stdin.
91
+ - **Writes.** Neither writes a file. Both write one JSON object to stdout: `{"hookSpecificOutput":{"hookEventName":"...","additionalContext":"..."}}`, and both exit only after that write's callback fires, so a buffered write to a pipe is not truncated by an exit racing ahead of it.
92
+ - **Fail-open, on purpose.** A missing `CLAUDE_PROJECT_DIR`, a missing rules file, a non-regular file at the rules path, or a missing block each produce a one-line fallback `additionalContext` naming what was found, and the script still exits 0. This is acceptable because a miss here is a stray context string reaching the model, not a security gate: nothing downstream trusts the hook's output for anything but a routing suggestion, and the settings snippet that wires it in is a document the user merges by hand, never written automatically over an existing `settings.json`.
93
+ - **Bounded.** `route-gate.mjs` caps the read at 64 KB regardless of the file's reported size and the injected string at 4000 characters, so a rules file bloated by a bad edit, or truncated to an arbitrary length, cannot balloon the context or the read time on every turn. Both hooks cap stdin drain at 250ms.
94
+ - **Not yet attacked.** Untested here: a rules file with a `route-gate:start` marker but no matching end marker very far into the file (bounded by `MAX_READ`, so the end marker past that point is treated as absent, which is the intended fail-open path, but worth a deliberate case); a `CLAUDE_PROJECT_DIR` pointing at a path with no read permission; behavior under the Windows exec-form `node` + `args` invocation named in the settings snippet; a `statSync` that itself hangs (a stalled network filesystem, for instance) rather than the FIFO-at-open case this round fixed.
95
+
96
+ ### Round 1 (pre-release), fixed before shipping
97
+
98
+ Three findings reproduced against the 0.1.15 branch before it shipped, none of them ever released:
99
+
100
+ | # | Finding | Fix |
101
+ |---|---|---|
102
+ | 1 | HIGH. `readFileSync(0)` in both hooks blocked until stdin reached EOF (`sleep 3 \| ... node route-gate.mjs` still running past 1.5s); `route-gate.mjs` also read the whole rules file into memory before bounding it, so a FIFO planted at the rules path blocked forever on open. | Stdin is drained asynchronously against a 250ms hard cap in both hooks. `route-gate.mjs` refuses anything that is not `isFile()` via `statSync` before ever calling open, then reads through one fixed 64 KB buffer via `openSync`/`readSync`. Tests: an open, never-closed stdin pipe exits within 1s for both hooks; a FIFO at the rules path returns the fallback instead of hanging; a 200 MB sparse rules file completes in well under a second with output still capped. |
103
+ | 2 | MEDIUM. `done-verifier`'s description, both agent-folder READMEs, and the root README called it "read-only" without qualification, while its claude-code file carries an unrestricted `Bash` grant; nothing in that grant stops it from running a mutating command. | Every one of those surfaces now says plainly that `done-verifier` carries no file-editing tools and that its Bash use is bound by its own prompt, not by the tool grant; `reader` is named as the one that is read-only by tool grant (no Bash) on both formats. |
104
+ | 3 | MEDIUM. Three generated surfaces still stated the pre-0.1.15 premise on a claude-code install: `builder.md`'s description ("... or the main build itself"), `build-protocol.md`'s roles table and its "why the builder does not hand off" note, and `ROUTING.md`'s "Plan big, execute small" line ("the orchestrator executes"). | All three now render through `subagentsLoadRules(primary)`, the same gate the decision tree and "Who builds" already used; every other primary is unchanged. A semantic-regression test asserts a claude-code install contains none of the old phrasing and a codex install still does. |
@@ -38,7 +38,7 @@ Every build gets two checkpoints. **Before writing:** you map what it touches an
38
38
 
39
39
  ## 4. Every hand-off carries a brief
40
40
 
41
- A subagent, a fresh chat, a second window holds none of your rules and reads an unspecified edge as an open one. The brief: purpose, task class, granted scope, capabilities, denied actions, conventions it does not have, report contract (what was not done, what is unverified), exit parameters (when to stop, and what to return if it hits the bound). Absence is denial.
41
+ A subagent, a fresh chat, a second window may hold none of your rules, and that is the default to assume. The one documented exception is a Claude Code subagent: it loads the project's CLAUDE.md hierarchy at start, so it keeps the standing rules, just not this task's scope. Either way, it reads an unspecified edge as an open one. The brief: purpose, task class, granted scope, capabilities, denied actions, conventions it does not have, report contract (what was not done, what is unverified), exit parameters (when to stop, and what to return if it hits the bound). Absence is denial.
42
42
 
43
43
  ## 5. The second pass
44
44
 
@@ -32,7 +32,7 @@ There is a second thing a lane can be quietly wrong about. Left unpinned, it run
32
32
 
33
33
  ## 4. Every delegation carries a task bundle, on both surfaces
34
34
 
35
- Subagents and CLI lanes are the same problem: something with none of your rules and broad tool access. The brief (purpose, task class, scope, capabilities, denied actions, conventions, report contract, exit parameters) goes in the prompt or in the file passed to `--brief`. If you can, gate it mechanically: a pre-dispatch hook that refuses a brief missing purpose, denied actions or a report contract.
35
+ Subagents and CLI lanes are close to the same problem: something that may hold none of your rules, and broad tool access. A Claude Code subagent is the one documented exception, loading the project's CLAUDE.md hierarchy at start, so it keeps the standing rules but not this task's scope; a CLI lane and a fresh chat window get no such credit. The brief (purpose, task class, scope, capabilities, denied actions, conventions, report contract, exit parameters) goes in the prompt or in the file passed to `--brief` either way. If you can, gate it mechanically: a pre-dispatch hook that refuses a brief missing purpose, denied actions or a report contract. On claude-code, a `SubagentStart` hook can inject the essentials (where the rules and the brief format live) automatically; `.claude/hooks/subagent-context.mjs` is the generated example.
36
36
 
37
37
  ## 5. Research: three engines, one triager
38
38
 
package/llms.txt ADDED
@@ -0,0 +1,27 @@
1
+ # model-orchestrator
2
+
3
+ > Model orchestrator for AI coding agents and LLMs (Claude Code, Codex, Gemini, Grok, Qwen, Ollama). One installer writes routing rules, subagent definitions and a CLI lane runner for the AI tools you already have. The rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. It is not a proxy or gateway: it does not automatically compare prices or select models; your agent follows the rules and chooses.
4
+
5
+ Install and run: `npx model-orchestrator` (interactive), or headless: `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator`. Preview without writing: add `--dry-run`. List every supported AI: `npx model-orchestrator --list`. Node 18 or newer, zero runtime dependencies, MIT licence.
6
+
7
+ Levels: 1 beginner (one agent or chat app), 2 intermediate (several agent CLIs, each called through `cli-run`), 3 advanced (adds a virtual machine with a gateway and a scheduled audit job). On Claude Code the install also delegates execution to subagents by default and ships two hooks (UserPromptSubmit, SubagentStart) that inject the routing table every turn.
8
+
9
+ ## Docs
10
+
11
+ - [README](https://github.com/aunysillyme/model-orchestrator/blob/main/README.md): what it writes, flags, principles, what is enforced versus instructed
12
+ - [Part 1: beginner](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-1-beginner.md): one agent, tiers, the task bundle every delegation carries
13
+ - [Part 2: intermediate](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-2-intermediate.md): several agent CLIs, the lane runner, pinning model and effort per lane
14
+ - [Part 3: advanced](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-3-advanced.md): the VM, the gateway, scheduled jobs, privacy gates
15
+ - [AI catalog](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/catalog.md): every supported AI, how to sign in, how each is detected
16
+
17
+ ## Reference
18
+
19
+ - [CLI runner](https://github.com/aunysillyme/model-orchestrator/blob/main/bin/README.md): `cli-run` lanes, exit codes, the route logged per run
20
+ - [Templates](https://github.com/aunysillyme/model-orchestrator/blob/main/templates/README.md): the routing, tiers, task bundle and protocol files the installer renders
21
+ - [Changelog](https://github.com/aunysillyme/model-orchestrator/blob/main/CHANGELOG.md): every release and the issue behind each fix
22
+ - [Agent instructions](https://github.com/aunysillyme/model-orchestrator/blob/main/AGENTS.md): running the installer from an agent, and contributing
23
+
24
+ ## Optional
25
+
26
+ - [Security policy](https://github.com/aunysillyme/model-orchestrator/blob/main/SECURITY.md)
27
+ - [Audit brief](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/audit-brief.md): the threat model and what has already been attacked
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.14",
4
- "description": "Routing instructions and a CLI runner for your AI tools. One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine. Routes by role, complexity and risk; pins and logs the model and reasoning effort each lane runs with.",
3
+ "version": "0.1.15",
4
+ "description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
5
5
  "type": "module",
6
6
  "bin": {
7
7
  "model-orchestrator": "bin/cli.js"
@@ -15,7 +15,9 @@
15
15
  "LICENSE",
16
16
  "CHANGELOG.md",
17
17
  "SECURITY.md",
18
- "scripts"
18
+ "scripts",
19
+ "llms.txt",
20
+ "AGENTS.md"
19
21
  ],
20
22
  "scripts": {
21
23
  "start": "node bin/cli.js",
@@ -55,7 +57,29 @@
55
57
  "reasoning-effort",
56
58
  "model-routing",
57
59
  "code-review",
58
- "ai-agents"
60
+ "ai-agents",
61
+ "hooks",
62
+ "claude-code-hooks",
63
+ "delegation",
64
+ "llm-router",
65
+ "ai-router",
66
+ "model-switching",
67
+ "multi-model",
68
+ "token-optimization",
69
+ "token-efficiency",
70
+ "cost-optimization",
71
+ "llm-orchestration",
72
+ "agent-orchestration",
73
+ "ai-orchestration",
74
+ "claude",
75
+ "anthropic",
76
+ "openai",
77
+ "gemini",
78
+ "claude-code-router",
79
+ "agent-routing",
80
+ "prompt-routing",
81
+ "coding-agent",
82
+ "ai-coding"
59
83
  ],
60
84
  "author": "aunysillyme (https://github.com/aunysillyme)",
61
85
  "license": "MIT"
package/src/README.md CHANGED
@@ -4,6 +4,6 @@
4
4
  |---|---|
5
5
  | `catalog.js` | the single list of levels and AIs. Add an AI here and the prompts, docs tables, delegation matrix, gateway config and installer all pick it up. Nothing else lists AIs. |
6
6
  | `detect.js` | PATH lookup for a binary, plus the few places vendor installers drop binaries without touching PATH. No shell-outs. |
7
- | `install.js` | pure planner: turns (level, selection, primary) into a list of files to write, rendering templates and computing every generated table. `writeFiles` is the only thing that touches disk. `activationSteps()` and `snippetFor()` live here so the terminal summary and the generated README render the same list. |
7
+ | `install.js` | pure planner: turns (level, selection, primary) into a list of files to write, rendering templates and computing every generated table. `writeFiles` is the only thing that touches disk. `activationSteps()` and `snippetFor()` live here so the terminal summary and the generated README render the same list. `subagentsLoadRules(primary)` gates every delegate-by-default render var (builder-by-default wording, the route-gate table, the inline-threshold note) on the one verified premise: a claude-code subagent loads CLAUDE.md. |
8
8
  | `prompt.js` | line-buffered questions for the interactive path; piped answers are queued, EOF mid-prompt aborts instead of confirming a write. |
9
9
  | `render.js` | `{{KEY}}` substitution. Throws on an unknown key, so a template typo fails the test suite instead of shipping a literal placeholder. |
package/src/catalog.js CHANGED
@@ -69,7 +69,15 @@ export const AIS = [
69
69
  rulesFile: 'CLAUDE.md',
70
70
  agentsDir: '.claude/agents',
71
71
  cliRun: false,
72
- models: { deep: 'opus', standard: 'sonnet', fast: 'haiku' }
72
+ models: { deep: 'opus', standard: 'sonnet', fast: 'haiku' },
73
+ // Verified at code.claude.com/docs/en/sub-agents (fetched 2026-09-10): "A
74
+ // non-fork subagent's initial context contains: CLAUDE.md files: every
75
+ // level of the CLAUDE.md hierarchy the main conversation loads ... The
76
+ // built-in Explore and Plan agents skip this." No other lane in this
77
+ // catalog has that documented, so the builder-by-default routing, the
78
+ // route-gate hook and the inline-threshold note are gated on this field
79
+ // and stay claude-code only.
80
+ subagentsLoadRules: true
73
81
  },
74
82
  {
75
83
  id: 'codex',
package/src/install.js CHANGED
@@ -193,6 +193,146 @@ export function laneVars(selected) {
193
193
  };
194
194
  }
195
195
 
196
+ // Only claude-code has a verified sub-agents doc quote saying its subagents
197
+ // load the project's CLAUDE.md hierarchy (code.claude.com/docs/en/sub-agents,
198
+ // see the comment on the catalog entry). Every delegate-by-default surface below
199
+ // (builder-by-default wording, the route-gate hook, the inline-threshold
200
+ // note) is gated on this so a primary with no verified premise keeps the
201
+ // original, more conservative wording.
202
+ export function subagentsLoadRules(primary) {
203
+ return !!(primary && primary.subagentsLoadRules);
204
+ }
205
+
206
+ // Canonical agent order, tier-first. Used to render a stable, non-hardcoded
207
+ // "available as" list for the claude-code snippet from the files actually
208
+ // shipped, so a future agent addition or removal cannot leave the sentence
209
+ // stale the way the finding-verifier omission did.
210
+ const AGENT_ORDER = ['deep-planner', 'builder', 'code-reviewer', 'finding-verifier', 'live-researcher', 'bulk-worker', 'done-verifier', 'reader'];
211
+ export function claudeAgentIds() {
212
+ const dir = join(TEMPLATES, 'agents', 'claude-code');
213
+ if (!existsSync(dir)) return [];
214
+ const files = readdirSync(dir).filter((f) => f.endsWith('.md') && f !== 'README.md').map((f) => f.replace(/\.md$/, ''));
215
+ const set = new Set(files);
216
+ const ordered = AGENT_ORDER.filter((id) => set.has(id));
217
+ const extra = files.filter((id) => !AGENT_ORDER.includes(id)).sort();
218
+ return [...ordered, ...extra];
219
+ }
220
+
221
+ // The compact "pick the lane before acting" table, rendered from the AIs the
222
+ // user actually selected and the agents actually installed, never a second
223
+ // hand-typed copy of ROUTING.md's decision tree.
224
+ export function routeGateTable(selected) {
225
+ const rows = [
226
+ ['Bulk or mechanical, many similar items', 'bulk-worker'],
227
+ ['Needs live data', 'live-researcher'],
228
+ ['Review without changing', 'code-reviewer'],
229
+ ['Findings from a review or a scanner', 'finding-verifier, before any repair'],
230
+ ['Reading or digesting many files or notes', 'reader'],
231
+ ['Checking a tracker item against its stated done-signal', 'done-verifier'],
232
+ ['Ambiguous, architectural, expensive to get wrong', 'deep-planner'],
233
+ ['Everything else that changes files', 'builder, by default']
234
+ ];
235
+ for (const a of selected.filter((x) => x.cliRun)) rows.push([a.role, '`cli-run ' + a.id + '`']);
236
+ return table(rows, ['Task', 'Lane']);
237
+ }
238
+
239
+ // The marked block route-gate.mjs extracts at runtime. Installed only for
240
+ // claude-code so the hook always finds a block to read; other primaries get
241
+ // no hook and so get no block.
242
+ export function routeGateSection(selected) {
243
+ return [
244
+ '<!-- route-gate:start -->',
245
+ '## Route gate: pick the lane before acting',
246
+ '',
247
+ 'Injected on every turn by the `route-gate` hook, so this table is read at runtime rather than recalled from memory.',
248
+ '',
249
+ routeGateTable(selected),
250
+ '',
251
+ "Stay inline only when: (a) the brief would cost as much as the work itself, (b) the task needs this conversation's own context, (c) it is the human's decision or the final verification of delegated work (a delegate never verifies itself).",
252
+ '',
253
+ 'Never: the built-in Explore or Plan agents for rule-bound work (they skip CLAUDE.md). general-purpose taking work a named agent already owns.',
254
+ '<!-- route-gate:end -->'
255
+ ].join('\n');
256
+ }
257
+
258
+ // ROUTING.md / ORCHESTRATOR.md decision-tree rule 5 and the "Who builds"
259
+ // section read differently for claude-code, because only claude-code has the
260
+ // verified premise that its subagents load CLAUDE.md. Every other primary
261
+ // keeps the original wording: the orchestrator builds the main line directly
262
+ // and a subagent or second CLI is assumed to hold none of these rules.
263
+ export function decisionRule5(primary) {
264
+ return subagentsLoadRules(primary)
265
+ ? `5. **Everything else that changes files** → builder executes by default. The orchestrator plans, briefs, verifies and talks to the human; it stays inline only when (a) the brief would cost as much as the work, (b) the task needs this conversation's own context, or (c) it is the human's decision, or the final verification of delegated work (a delegate never verifies itself). Never route rule-bound work to the built-in Explore or Plan agents: both skip CLAUDE.md. general-purpose should not take work a named agent already owns.`
266
+ : `5. **Everything else that changes files** → the orchestrator builds it directly. Bounded sub-parts go to cheaper tiers; the main build is never handed off whole.`;
267
+ }
268
+ export function decisionRule5Beginner(primary) {
269
+ return subagentsLoadRules(primary)
270
+ ? `5. **Everything else that changes files or executes a known plan** → builder executes by default, at standard tier. The orchestrator plans, briefs, verifies and talks to you; it stays inline only when (a) the brief would cost as much as the work, (b) the task needs this conversation's own context, or (c) it is your decision, or the final verification of delegated work. Never route rule-bound work to the built-in Explore or Plan agents: both skip CLAUDE.md.`
271
+ : `5. **Everything else that changes files or executes a known plan** → you build it directly, at standard tier. The main build is never handed off whole; bounded sub-parts (a bulk pass, a wide search, a long audit loop) can go to cheaper tiers.`;
272
+ }
273
+ export function whoBuildsSection(primary) {
274
+ if (subagentsLoadRules(primary)) {
275
+ return [
276
+ '## Who builds',
277
+ '',
278
+ `**Builder executes by default.** A Claude Code subagent loads this project's CLAUDE.md hierarchy at start (verified: code.claude.com/docs/en/sub-agents), so it already carries the standing rules; the orchestrator's job is to plan, brief, verify and talk to the human, not to hold work a delegate can do. Stay inline only when: (a) the brief would cost as much as the work itself, (b) the task needs this conversation's own context, or (c) it is the human's decision to make, or the final verification of delegated work (a delegate never verifies its own output as final). Never route rule-bound work to the built-in Explore or Plan agents: both skip CLAUDE.md and the git status the router depends on. general-purpose should not take work a named agent already owns.`,
279
+ '',
280
+ `Delegate: the main build, background and long-running tasks, small tasks, scoping, verification, research, bounded sub-parts. Never delegate: the human's own decision, or the final sign-off on a delegate's work.`,
281
+ '',
282
+ `Every delegation carries \`TASK_BUNDLE.md\`. Its brief must restate this task's scope: a Claude Code subagent already has the standing rules, just not that.`
283
+ ].join('\n');
284
+ }
285
+ return [
286
+ '## Who builds',
287
+ '',
288
+ '**The orchestrator owns the main build.** It is the only surface that holds these rules: a subagent or a second CLI starts with none of them and cannot route. Handing the main build to one hands it to something the router cannot reach.',
289
+ '',
290
+ 'Delegate: background and long-running tasks, small tasks, scoping, verification, research, bounded sub-parts. Never delegate: the main build, or any step that must carry a house rule (secrets handling, the loud-negative verification, the durable record).',
291
+ '',
292
+ 'Every delegation carries `TASK_BUNDLE.md`. Its brief must restate every convention the delegate needs.'
293
+ ].join('\n');
294
+ }
295
+ export function addEndpointRow(primary) {
296
+ return subagentsLoadRules(primary)
297
+ ? '| "Add an endpoint" | builder, briefed and verified by the orchestrator |'
298
+ : '| "Add an endpoint" | the orchestrator builds it |';
299
+ }
300
+ export function inlineThresholdNote(primary) {
301
+ return subagentsLoadRules(primary)
302
+ ? '\n- **Measure your inline threshold once.** A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Spawn one with a one-line task and read its token count. Work smaller than that stays inline.'
303
+ : '';
304
+ }
305
+ export function delegateRulesNote(primary) {
306
+ return subagentsLoadRules(primary)
307
+ ? `Subagents, a fresh chat, a second window: a Claude Code subagent loads this project's CLAUDE.md hierarchy, so it holds the standing rules already, just not this task's scope; a second CLI or a fresh chat window may hold none of them.`
308
+ : 'Subagents, a fresh chat, a second window: each one holds none of these rules.';
309
+ }
310
+
311
+ // Pre-release audit finding 3: the delegate-by-default gate reached the
312
+ // decision tree and "Who builds" but missed three other generated surfaces
313
+ // stating the same old premise (the orchestrator writes the main build
314
+ // itself; a delegate inherits none of the session's rules). These three
315
+ // close that gap the same way: gated on subagentsLoadRules(primary), every
316
+ // other primary keeps the original wording unchanged.
317
+ export function planBigExecuteSmallLine(primary) {
318
+ return subagentsLoadRules(primary)
319
+ ? `- **Plan big, execute small**, within a build: deep tier plans at Checkpoint 1, builder executes from the orchestrator's brief, bulk and wide searches go down.`
320
+ : '- **Plan big, execute small**, within a build: deep tier plans at Checkpoint 1, the orchestrator executes, bulk and wide searches go down.';
321
+ }
322
+ export function rolesBuilderRow(primary) {
323
+ return subagentsLoadRules(primary)
324
+ ? [
325
+ '| Orchestrator | Routes, maps, briefs, verifies, records. Stages 0, 1, 2, 4, 5b, 6, 7 | Write the build |',
326
+ "| Builder | Executes Stage 3 from the orchestrator's brief | Route further, or verify its own work as final |"
327
+ ].join('\n')
328
+ : '| Builder / orchestrator | Routes, maps, writes, verifies, records. Stages 0, 1, 3, 6, 7 | Hand off the main build |';
329
+ }
330
+ export function builderHandoffNote(primary) {
331
+ return subagentsLoadRules(primary)
332
+ ? `**Why Stage 3 goes to builder by default:** a Claude Code subagent loads this project's CLAUDE.md hierarchy at start, so it already carries the standing rules; the orchestrator's brief only has to restate this task's scope (see \`TASK_BUNDLE.md\`). The orchestrator keeps Stage 3 for itself only when the brief would cost as much as the work, the task needs this conversation's own context, or it is the human's decision or the final verification of delegated work.`
333
+ : `**Why the builder does not hand off the main build:** a delegated agent does not inherit the session's standing rules and usually cannot delegate further. Any brief must restate every convention it needs (see \`TASK_BUNDLE.md\`), and that cost is itself a reason to build directly when the work fits.`;
334
+ }
335
+
196
336
  // Which activation file this primary gets. ONE decision, read by three
197
337
  // surfaces: planFiles writes the file, vars() names it in the generated README,
198
338
  // and bin/cli.js prints it in the terminal. Before 0.1.12 the README hardcoded
@@ -218,6 +358,9 @@ export function activationSteps(opts) {
218
358
  // replaces (#22).
219
359
  else if (snippet) steps.push(`open ${primary.chatName || primary.name} and paste the block in ${join(dirAbs, snippet)} into its ${primary.chatSurface || 'custom instructions'}`);
220
360
  if (primary && primary.agentsDir) steps.push(`subagents are in ${join(projectAbs, primary.agentsDir)}; run ${primary.bin} from ${projectAbs} to pick them up`);
361
+ // Only claude-code ships hooks (route-gate, subagent-context): the wiring
362
+ // lives in a snippet, never written into a settings.json the user already has.
363
+ if (subagentsLoadRules(primary)) steps.push(`merge the hooks in ${join(dirAbs, 'settings.hooks.snippet.json')} into ${join(projectAbs, '.claude', 'settings.json')} (create it if missing) to wire the route-gate and subagent-context hooks`);
221
364
  for (const a of selected.filter((a) => a.bin && a.kind === 'agent-cli')) steps.push(`sign in to ${a.name}: ${a.auth}`);
222
365
  // A local runtime has a bin but no sign-in, so the agent-cli loop above skips it
223
366
  // and before this it appeared in no ordered list at any level (#26).
@@ -232,7 +375,7 @@ export function activationSteps(opts) {
232
375
  // activationSteps is: level 1 writes no bin/, so a step naming cli-run.mjs or
233
376
  // lanes.json there described an install that did not happen (#27).
234
377
  export function proofSteps(opts) {
235
- const { level } = opts;
378
+ const { level, primary } = opts;
236
379
  const steps = [
237
380
  'Start a fresh agent session and ask: "Read the orchestrator instructions. Quote the routing rule you will use, then sort pear, apple, banana alphabetically. Name the tier and whether you delegated."',
238
381
  'Expect the fast tier and `apple, banana, pear`. If the agent cannot quote the routing rule, check the snippet location or chat instructions before continuing. This is a manual activation check, not proof that every future task follows the rules.'
@@ -242,6 +385,12 @@ export function proofSteps(opts) {
242
385
  steps.push('Decide whether the route matters to you. Every lane starts unpinned, which means it runs on whatever its own config file says: a CLI configured months ago at a low reasoning effort will keep auditing at that effort while your docs describe something stronger. Pin it in `bin/lanes.json` under `defaults`, or per call with `--model` and `--effort`. Either way the run is recorded in the log with the value requested and where it came from.');
243
386
  steps.push('To test a real output contract, choose an enabled lane from `bin/lanes.json` and run `node bin/cli-run.mjs <lane> \'Return only {"sorted":["apple","banana","pear"]}\' --expect-json`. This uses quota. Expect JSON and exit 0; inspect the array yourself. A non-JSON response exits 10, a missing binary exits 13, and an authentication failure reports the vendor error. The explicit lane tests execution; your primary agent still makes delegation decisions.');
244
387
  }
388
+ // Only claude-code ships the route-gate hook, so only claude-code gets a
389
+ // proof step that checks it fired: the table must come from the hook's
390
+ // injected context, not from the agent reciting ROUTING.md from memory.
391
+ if (subagentsLoadRules(primary)) {
392
+ steps.push('Ask the agent: "Quote the route-gate table you were given this turn." It should quote the injected table verbatim, not recite it from memory. If it cannot, the hooks snippet was not merged into `.claude/settings.json`, or the hook found no rules file: check both before trusting the routing docs are actually reaching the agent.');
393
+ }
245
394
  return steps;
246
395
  }
247
396
 
@@ -260,7 +409,14 @@ function vars(opts) {
260
409
  const pinOf = (id) => (toolById[id] && toolById[id].pin) || 'latest';
261
410
  const snippet = snippetFor(primary);
262
411
  const steps = activationSteps({ level, selected, primary, tools, dir: opts.dir, project: opts.project });
263
- const proofs = proofSteps({ level });
412
+ const proofs = proofSteps({ level, primary });
413
+ const routingFile = level >= 2 ? 'ROUTING.md' : 'ORCHESTRATOR.md';
414
+ // The path route-gate.mjs and subagent-context.mjs resolve at runtime,
415
+ // relative to CLAUDE_PROJECT_DIR. Mirrors the RULES_PATH fallback below:
416
+ // outside the project, the honest path is absolute, never a hardcoded one.
417
+ const relJoin = (name) => (rulesPath === dirAbs ? join(dirAbs, name) : rulesPath === '.' ? name : rulesPath + '/' + name);
418
+ const rulesFileRel = relJoin(routingFile);
419
+ const taskBundleRel = relJoin('TASK_BUNDLE.md');
264
420
  // Only claude-code and agy put files under the project root. A chat primary
265
421
  // puts nothing there, so naming a project root would name a folder this run
266
422
  // never created (#21).
@@ -336,7 +492,24 @@ function vars(opts) {
336
492
  NPM_PACKAGES: selected.map(npmSpec).filter(Boolean).join(' ') || '""',
337
493
  SCRIPT_INSTALLERS: scriptInstallers(selected),
338
494
  COMPOSE_ENV: composeEnv(selected, apis),
339
- COMPOSE_OLLAMA: composeOllama(selected)
495
+ COMPOSE_OLLAMA: composeOllama(selected),
496
+ // Delegate by default (0.1.15): gated on subagentsLoadRules(primary), currently
497
+ // claude-code only. Every other primary keeps the original, more
498
+ // conservative wording these replace.
499
+ DECISION_RULE5: decisionRule5(primary),
500
+ DECISION_RULE5_L1: decisionRule5Beginner(primary),
501
+ WHO_BUILDS: whoBuildsSection(primary),
502
+ ADD_ENDPOINT_ROW: addEndpointRow(primary),
503
+ INLINE_THRESHOLD_NOTE: inlineThresholdNote(primary),
504
+ DELEGATE_RULES_NOTE: delegateRulesNote(primary),
505
+ PLAN_BIG_LINE: planBigExecuteSmallLine(primary),
506
+ ROLES_BUILDER_ROW: rolesBuilderRow(primary),
507
+ BUILDER_HANDOFF_NOTE: builderHandoffNote(primary),
508
+ ROUTE_GATE_SECTION: subagentsLoadRules(primary) ? '\n' + routeGateSection(selected) + '\n' : '',
509
+ AGENTS_LIST_LINE: claudeAgentIds().map((id) => '`' + id + '`').join(', '),
510
+ RULES_FILE_REL: rulesFileRel,
511
+ RULES_FILE_REL_JSON: JSON.stringify(rulesFileRel),
512
+ TASK_BUNDLE_REL_JSON: JSON.stringify(taskBundleRel)
340
513
  };
341
514
  }
342
515
 
@@ -366,6 +539,13 @@ export function planFiles(opts) {
366
539
  add(join('.claude', 'agents', f.rel), render(readFileSync(f.abs, 'utf8'), v), 0o644, 'project');
367
540
  }
368
541
  add('CLAUDE.snippet.md', render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'claude-code.md'), 'utf8'), v));
542
+ // Delegate-by-default hooks (0.1.15), claude-code only: route-gate.mjs (UserPromptSubmit)
543
+ // and subagent-context.mjs (SubagentStart) live where Claude Code looks for
544
+ // project hooks; the wiring snippet is a document the user merges in, never
545
+ // written into a settings.json they already have.
546
+ add(join('.claude', 'hooks', 'route-gate.mjs'), render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'route-gate.mjs'), 'utf8'), v), 0o755, 'project');
547
+ add(join('.claude', 'hooks', 'subagent-context.mjs'), render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'subagent-context.mjs'), 'utf8'), v), 0o755, 'project');
548
+ add('settings.hooks.snippet.json', render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'settings.hooks.snippet.json'), 'utf8'), v));
369
549
  } else if (primary && primary.id === 'agy') {
370
550
  for (const f of walk(join(TEMPLATES, 'agents', 'agy'))) {
371
551
  if (!installable('agents', f.rel)) continue;
@@ -6,7 +6,7 @@ Everything the installer can write, organized by the level that adds it. Files a
6
6
  |---|---|---|
7
7
  | `common/` | every level | the start-here README, `TASK_BUNDLE.md`, `protocols/` (build, propagate, gap analysis, deep research, numbers and logic, memory and record) |
8
8
  | `beginner/` | every level | `ORCHESTRATOR.md`, the single-agent routing rules |
9
- | `agents/` | every level, one variant | the primary agent's loading surface: Claude Code subagents, Antigravity custom agents, or a paste snippet |
9
+ | `agents/` | every level, one variant | the primary agent's loading surface: Claude Code subagents (plus `.claude/hooks/route-gate.mjs` and `subagent-context.mjs`, and `settings.hooks.snippet.json` to wire them in), Antigravity custom agents, or a paste snippet |
10
10
  | `intermediate/` | level 2+ | `ROUTING.md`, `TIERS.md`, `DELEGATION_MATRIX.md`, `RESEARCH_TRIAGE.md`, `CLI-RUN.md` |
11
11
  | `advanced/` | level 3 | `vm/`: gateway config, compose file, box rules, privacy gates, scheduled jobs |
12
12
  | `tools/` | when selected | companion tools the AIs call: `codecalc/` and `obsidian-tc/` (install doc + MCP snippets each). See `tools/README.md` |