model-orchestrator 0.1.19 → 0.1.21

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md CHANGED
@@ -10,6 +10,7 @@ model-orchestrator writes routing rules, subagents and a CLI lane runner so an a
10
10
  - `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project <repo> --dir <repo>/ai-orchestrator --dry-run` prints the plan and writes nothing.
11
11
  - Drop `--dry-run` to write it. Existing files are never overwritten without `--force`; activation snippets (for example `CLAUDE.snippet.md`, `settings.hooks.snippet.json`) are written for a person or agent to merge.
12
12
  - The generated `README.md` in `--dir` lists what to copy where and one smoke command to prove the rules took.
13
+ - Claude Code users can install the hooks and subagents as a plugin instead of merging snippets: `claude plugin marketplace add aunysillyme/model-orchestrator`, then `claude plugin install model-orchestrator@model-orchestrator`. The rules still come from the installer above; see [`plugin/README.md`](plugin/README.md).
13
14
  - A summary for LLMs, with links to every doc: [`llms.txt`](llms.txt).
14
15
 
15
16
  ## Working on this repository
@@ -19,6 +20,7 @@ The files the installer writes for end users live under `templates/`.
19
20
  - Read `CONTRIBUTING.md` first, then `src/README.md` (the catalog drives everything) and `docs/audit-brief.md` (the threat model and what has already been attacked).
20
21
  - Run `npm test` before proposing a change and quote the count and the exit code; the suite prints the current number.
21
22
  - Everything renders from `src/catalog.js`. Add an AI or a tool there, not in a template. Templates carry no logic.
23
+ - `plugin/` is generated. Edit the agent or hook in `templates/`, then `npm run gen:plugin`; `test/plugin.test.js` fails when the committed bundle drifts. A plugin hook may only read: no network, no file writes, no subprocess.
22
24
  - Never put a value that looks like a credential anywhere in this repo, including tests and examples. Environment variable names only.
23
25
  - `bin/cli.js` writes only inside `--dir` and `--project`, never over a document without `--force`, and never runs a vendor script. A change that weakens any of those will be refused in review; the tests that hold them are in `test/install.test.js` and `test/cli.test.js`.
24
26
  - `bin/cli-run.mjs` must exit non-zero when a lane produced nothing. Every judge has a red case in `test/judges.test.js`; add one before you change a judge.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,42 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.21] - 2026-09-12
8
+
9
+ ### Added
10
+
11
+ - **Subscription plans, stated by you and never guessed.** The installer asks which plan you hold for Claude Code, Codex, `agy` and `grok` (interactive, or `--plans codex=pro-20x,agy=ultra-5x`; `--plans none` clears). Each plan row in `src/catalog.js` carries only a name, a headroom level (`base`, `high`, `max`), its official source page and the date it was checked; no prices and no model ids, because both change faster than releases. `--list` prints them.
12
+ - **Plan guidance in the generated docs.** The lanes table gains a `Plan` column, and `ROUTING.md` and `DELEGATION_MATRIX.md` gain a plan guidance block: high and max headroom lanes take volume (scoped builds, pre-ship second-family audits, first-pass research), base headroom lanes keep short second opinions. Headroom moves volume only; who reviews what does not change.
13
+ - **`--effort auto` in `cli-run`.** Accepted as a flag or as a `lanes.json` default on every lane with an effort flag (qwen still refuses it). A call sizes from prompt length: under 4,000 characters is `medium`, otherwise `high`. A codex `--audit` always runs at `high` and logs the larger of prompt length and changed lines as evidence. Auto never resolves above `high`: `xhigh`, `max` and `ultra` are sent only when named. The durable log gains `effort_resolved`, `effort_basis` (`explicit`, `prompt_chars`, `audit_floor`, `none`), `effort_scope` and `effort_truncated`.
14
+ - **Bounded change counting for audits.** Untracked files are listed NUL-delimited, only regular files are opened (FIFOs, devices and symlinks are skipped by `lstat`), and the pass stops at 200 files, 256 KiB per file, 2 MiB total or 2 seconds, marking the evidence truncated. An unborn `HEAD` or any git failure falls back to prompt length.
15
+
16
+ ### Changed
17
+
18
+ - **Automatic effort is opt-in.** A high or max headroom plan does not change `bin/lanes.json` by itself; `--effort-auto` (or yes to the question) writes `"effort": "auto"` for exactly those cli-run lanes. Claude Code is never written there, since it is not a cli-run lane and an entry for it would fail the whole file closed. With no plan stated, `bin/lanes.json` is byte-identical to 0.1.20 and `MANIFEST.json` records no plan keys. A re-run that omits `--plans` keeps the previous plans.
19
+ - **Five documented Windows skips, not four.** The FIFO half of the new untracked-file test needs `mkfifo`; its symlink half runs everywhere.
20
+
21
+ ## [0.1.20] - 2026-09-12
22
+
23
+ ### Added
24
+
25
+ - **A Claude Code plugin.** `/plugin marketplace add aunysillyme/model-orchestrator`, then `/plugin install model-orchestrator@model-orchestrator`, installs `route-gate.mjs`, `subagent-context.mjs` and the eight subagents without merging a settings snippet by hand. The bundle lives in `plugin/`, listed by `.claude-plugin/marketplace.json` at the repo root. It passes `claude plugin validate --strict`, the check Anthropic's community marketplace review runs, and all eight checks of Sigistry's public plugin verification methodology (1.2), run standalone before release as a quality bar; the plugin is not listed there.
26
+ - **The plugin is generated, never a second copy.** `npm run gen:plugin` renders `plugin/` from the same `templates/` the installer uses, and `test/plugin.test.js` fails when the committed bundle drifts from that, when `plugin.json`'s version is not `package.json`'s, when `hooks/hooks.json` references a hook that is not shipped, when a plugin hook gains a network call, a file write, credential access, dynamic evaluation or a subprocess, when an agent has no `tools:` line or a review-type agent carries Write or Edit, and when the plugin README loses its install commands. Each check was proved red against the real files before it was trusted.
27
+ - **The plugin's route gate works without an install step.** A plugin cannot be rendered per project, so its `route-gate.mjs` reads the installer's default locations, `ai-orchestrator/ROUTING.md` then `ai-orchestrator/ORCHESTRATOR.md`, and takes the first that exists. Something at the first path that is not a readable file (a directory, a FIFO) is reported, never skipped for the second. With neither present it tells Claude on every prompt, and the user once at session start, to run `npx model-orchestrator`, so a project with no rules is never a silent no-op.
28
+
29
+ ### Changed
30
+
31
+ - **`builder`, `deep-planner` and `live-researcher` now declare their tools, for npm installs too.** Until now they carried no `tools:` line and inherited every tool the session had, MCP tools included. `builder` gets `Read, Write, Edit, Glob, Grep, Bash`; `deep-planner` gets `Read, Glob, Grep` (its prompt already says it never edits); `live-researcher` gets `WebSearch, WebFetch`. This narrows what those three agents can do in an existing install once regenerated: if you relied on `builder` calling an MCP tool, or `deep-planner` running a command, add the tool to that agent's `tools:` line or delete the line.
32
+ - **`route-gate.mjs` takes a list of rules paths instead of one.** An installer render is a one-element list with no setup hint, so an npm install behaves exactly as before; a test pins that render.
33
+
34
+ ### Fixed
35
+
36
+ - **`route-gate.mjs` could emit more than Claude Code's 10,000-character hook output cap, on npm installs too.** The routing table was capped at 4,000 characters, but a fallback message embeds the resolved project path and the error text, so a 12,000-character `CLAUDE_PROJECT_DIR` produced 12,146 characters from an installer render. Every string the hook emits is now capped at 8,000 characters, with a test on both renders. Found by the pre-release audit round and reproduced before the fix.
37
+ - **The plugin's hook-safety test could not see an async write or subprocess.** `writeFileSync?` matches `writeFileSyn` and `writeFileSync`, never `writeFile`, and every `process.env` read was exempt. The check now matches the Sync and async form of every file write and subprocess call, refuses dynamic `import(` and `require(`, allows static imports of `node:fs` and `node:path` only, requires `openSync` to open read-only, and allows no environment variable but `CLAUDE_PROJECT_DIR`, with a red case for each. Found by the same audit round.
38
+
39
+ ### Not changed
40
+
41
+ - `route-metrics.mjs` still installs with `npx model-orchestrator`, unchanged. It is left out of the plugin only, because it writes a log to disk and the plugin ships only hooks that read.
42
+
7
43
  ## [0.1.19] - 2026-09-12
8
44
 
9
45
  ### Added
@@ -293,7 +329,9 @@ First release.
293
329
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
294
330
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
295
331
 
296
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.19...HEAD
332
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.21...HEAD
333
+ [0.1.21]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.20...v0.1.21
334
+ [0.1.20]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.19...v0.1.20
297
335
  [0.1.19]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.18...v0.1.19
298
336
  [0.1.18]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.17...v0.1.18
299
337
  [0.1.17]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.16...v0.1.17
package/README.md CHANGED
@@ -2,13 +2,14 @@
2
2
 
3
3
  [![npm](https://img.shields.io/npm/v/model-orchestrator.svg)](https://www.npmjs.com/package/model-orchestrator) [![test](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml/badge.svg)](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml) [![license: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![node >=18](https://img.shields.io/badge/node-%3E%3D18-brightgreen.svg)](package.json)
4
4
 
5
- **Route every task to the right model, agent or LLM, and spend fewer tokens.** A model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. One installer asks what you have access to and writes routing rules, subagents and a CLI runner for exactly that setup, from one chat app to several agent CLIs or a virtual machine. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models; the runner executes the lane it is given. On Claude Code it also delegates execution to subagents by default, with two hooks that inject the routing table every turn.
5
+ **Route every task to the right model, agent or LLM, and spend fewer tokens.** A model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. One installer asks what you have access to and writes routing rules, subagents and a CLI runner for exactly that setup, from one chat app to several agent CLIs or a virtual machine. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models; the runner executes the lane it is given. On Claude Code it also delegates execution to subagents by default, with two hooks that inject the routing table every turn, and the hooks and subagents also install as a [Claude Code plugin](#claude-code-plugin).
6
6
 
7
7
  ## At a glance
8
8
 
9
9
  - **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
10
10
  - **What it is not:** a proxy, a gateway or an API router. It does not automatically compare prices or select models; your agent follows the rules and chooses.
11
11
  - **Install:** `npx model-orchestrator` (interactive), or headless from a script or an agent: `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator`.
12
+ - **Claude Code plugin:** `/plugin marketplace add aunysillyme/model-orchestrator`, then `/plugin install model-orchestrator@model-orchestrator`. The routing rules still come from the installer; see [Claude Code plugin](#claude-code-plugin).
12
13
  - **Use it when:** you run more than one model or agent and want each task sent to the smallest one that can do it well.
13
14
  - **What it saves:** frontier-model tokens. Bulk work, reading and checks go to fast tiers; the expensive tier is kept for planning and judgment.
14
15
  - **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
@@ -29,6 +30,10 @@ The installer asks a few things, then writes a folder:
29
30
 
30
31
  It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions (and, on Claude Code, two hook scripts) go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: follow the generated README. Inspect the manifest and remove only the individual managed subagent files you no longer need, preserve edited or pre-existing files, and remove your manually pasted activation block. Never delete a shared subagent folder.
31
32
 
33
+ ## Plans and automatic effort
34
+
35
+ State known subscription plans with `--plans codex=pro-20x,agy=ultra-5x`. The generated guidance uses plan headroom to allocate volume only. It never changes capability or independent-review rules. `--effort-auto` is explicit consent to set `auto` only for selected high or max headroom CLI lanes. Auto chooses medium below 4,000 prompt characters and high otherwise, never higher. A codex audit is always high. Name `xhigh` explicitly for security-critical or irreversible work.
36
+
32
37
  ## The three levels
33
38
 
34
39
  | Level | You have | You get |
@@ -90,6 +95,22 @@ npx model-orchestrator --yes --level 2 --ais claude-code,codex --project ~/my-ap
90
95
  npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --update-docs # added a lane: regenerate the docs you never edited
91
96
  ```
92
97
 
98
+ ## Claude Code plugin
99
+
100
+ The Claude Code hooks and subagents also ship as a plugin, so they install and update through Claude Code itself instead of a settings snippet you merge by hand:
101
+
102
+ ```
103
+ /plugin marketplace add aunysillyme/model-orchestrator
104
+ /plugin install model-orchestrator@model-orchestrator
105
+ ```
106
+
107
+ - **What it ships:** `route-gate.mjs` (UserPromptSubmit, plus a one-line notice at session start when the project has no rules), `subagent-context.mjs` (SubagentStart), and the eight subagents, each with an explicit tool list. Agents load namespaced, as `model-orchestrator:builder`.
108
+ - **What it still needs from the installer:** the routing rules. A plugin runs no install step, so it reads the installer's default locations, `ai-orchestrator/ROUTING.md` then `ai-orchestrator/ORCHESTRATOR.md`, and names `npx model-orchestrator` when neither exists. A project installed with a different `--dir` should wire the installer's rendered hooks instead.
109
+ - **What it leaves out:** `route-metrics.mjs`, the routing log. It writes to disk and the plugin ships only hooks that read, so `npx model-orchestrator` is how you get it.
110
+ - **How it is kept honest:** `plugin/` is generated from `templates/` by `npm run gen:plugin`, and `test/plugin.test.js` fails when the committed bundle drifts, when a hook gains a network call, a write or a subprocess, or when an agent loses its tool list. The bundle passes `claude plugin validate --strict`, the check Anthropic's community marketplace review runs on every submission.
111
+
112
+ Details, including running the plugin next to an npm install: [plugin/README.md](plugin/README.md).
113
+
93
114
  ## What gets written (level 3, everything)
94
115
 
95
116
  ```
@@ -117,6 +138,7 @@ ai-orchestrator/
117
138
  | [`src/`](src/README.md) | the catalog, the pure planner, detection, rendering |
118
139
  | [`templates/`](templates/README.md) | everything the installer can write, by level, plus `tools/` for companions |
119
140
  | [`docs/`](docs/README.md) | the three parts and the catalog |
141
+ | [`plugin/`](plugin/README.md) | the Claude Code plugin, generated from `templates/` by `npm run gen:plugin`; `.claude-plugin/marketplace.json` at the root lists it |
120
142
  | [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu, macOS and Windows, Node 18/20/22 |
121
143
 
122
144
  ## What is enforced, what is delegated, what is an instruction
@@ -260,7 +282,7 @@ Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run`
260
282
 
261
283
  ## Requirements
262
284
 
263
- Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22), including lane execution end to end through `cli-run` against a fake CLI installed the same way npm installs a real one (a `.cmd` shim). `cli-run` never runs a lane through `cmd.exe` when it can avoid it: it resolves the shim to the Node script underneath and spawns Node directly, so a prompt reaching a real lane never passes through a Windows shell. A `.cmd` or `.bat` lane that cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, rather than run through `cmd.exe`: a batch file re-reads its arguments after `cmd.exe` has parsed them once, and no escaping fully contains a prompt through both passes. Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested on Windows too, including SIGTERM/SIGINT to the wrapper (Windows has no OS-level signals: both terminate it unconditionally, verified there rather than treated the same as POSIX). Four narrow skips remain on Windows, each for a POSIX behavior the OS or the CI shell genuinely does not have, and each named here because a test that is quietly skipped reads as a test that passed: `statSync().mode`'s executable bit (NTFS has none, so that one assertion is conditional inside a test that otherwise runs everywhere); a lane dying mid-run from a real POSIX signal (a real Windows lane cannot die "by signal"); running `weekly-audit.sh`'s watchdog functions for real under Git Bash's job control, both the end-to-end run and the `bounded()` timeout check (the script itself only ever runs on the Ubuntu box it targets); and a `mkfifo` FIFO at the rules path, the one case that proves `route-gate.mjs` cannot HANG on a non-regular file, since Windows has no `mkfifo` to build one (the guard behind it is covered on every OS by a directory at the same path). The list is not prose on trust: `test/prose.test.js` counts every `skip:` in the suite and fails if one of them is not documented here.
285
+ Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22), including lane execution end to end through `cli-run` against a fake CLI installed the same way npm installs a real one (a `.cmd` shim). `cli-run` never runs a lane through `cmd.exe` when it can avoid it: it resolves the shim to the Node script underneath and spawns Node directly, so a prompt reaching a real lane never passes through a Windows shell. A `.cmd` or `.bat` lane that cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, rather than run through `cmd.exe`: a batch file re-reads its arguments after `cmd.exe` has parsed them once, and no escaping fully contains a prompt through both passes. Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested on Windows too, including SIGTERM/SIGINT to the wrapper (Windows has no OS-level signals: both terminate it unconditionally, verified there rather than treated the same as POSIX). Five narrow skips remain on Windows, each for a POSIX behavior the OS or the CI shell genuinely does not have, and each named here because a test that is quietly skipped reads as a test that passed: `statSync().mode`'s executable bit (NTFS has none, so that one assertion is conditional inside a test that otherwise runs everywhere); a lane dying mid-run from a real POSIX signal (a real Windows lane cannot die "by signal"); running `weekly-audit.sh`'s watchdog functions for real under Git Bash's job control, both the end-to-end run and the `bounded()` timeout check (the script itself only ever runs on the Ubuntu box it targets); and a `mkfifo` FIFO at the rules path, the one case that proves `route-gate.mjs` cannot HANG on a non-regular file, since Windows has no `mkfifo` to build one (the guard behind it is covered on every OS by a directory at the same path); and an untracked `mkfifo` FIFO in the repository `cli-run --audit` sizes, the case that proves `--effort auto` never opens a non-regular file (the symlink half of that test runs on every OS). The list is not prose on trust: `test/prose.test.js` counts every `skip:` in the suite and fails if one of them is not documented here.
264
286
 
265
287
  **Privacy.** The installer sends no telemetry and makes no network call of its own once it is running. Two things around that are worth being exact about:
266
288
 
package/bin/cli-run.mjs CHANGED
@@ -51,10 +51,10 @@
51
51
  // for one lane and absent for four is worse than no field. It would also be a
52
52
  // provider-supplied string, which this log deliberately never holds.
53
53
 
54
- import { spawn } from 'node:child_process';
54
+ import { spawn, spawnSync } from 'node:child_process';
55
55
  import { StringDecoder } from 'node:string_decoder';
56
56
  import { createHash } from 'node:crypto';
57
- import { readFileSync, existsSync, mkdirSync, appendFileSync, mkdtempSync, rmSync, accessSync, constants, realpathSync, statSync } from 'node:fs';
57
+ import { readFileSync, existsSync, mkdirSync, appendFileSync, mkdtempSync, rmSync, accessSync, constants, realpathSync, statSync, lstatSync, openSync, readSync, closeSync } from 'node:fs';
58
58
  import { join, dirname, delimiter, resolve } from 'node:path';
59
59
  import { tmpdir, homedir } from 'node:os';
60
60
  import { fileURLToPath, pathToFileURL } from 'node:url';
@@ -233,6 +233,82 @@ export const LANE_FLAGS = {
233
233
  qwen: { model: (v) => ['-m', v], effort: null }
234
234
  };
235
235
 
236
+ // Auto is deliberately a small, static ladder. It is not a vendor capability
237
+ // probe and it never chooses the top of a vendor's effort range.
238
+ export const AUTO_EFFORT = {
239
+ codex: { small: 'medium', large: 'high' },
240
+ grok: { small: 'medium', large: 'high' },
241
+ agy: { small: 'medium', large: 'high' },
242
+ hermes: { small: 'medium', large: 'high' }
243
+ };
244
+
245
+ export function gitChangedLines(cwd) {
246
+ try {
247
+ const r = spawnSync('git', ['diff', '--numstat', '-z', 'HEAD'], { cwd, timeout: 5000, maxBuffer: 1024 * 1024 });
248
+ if (r.error || r.status !== 0) return null;
249
+ let total = 0;
250
+ for (const row of String(r.stdout).split('\0')) {
251
+ if (!row) continue;
252
+ const [added, removed] = row.split('\t');
253
+ total += (Number.isFinite(Number(added)) ? Number(added) : 0) + (Number.isFinite(Number(removed)) ? Number(removed) : 0);
254
+ }
255
+ const listed = spawnSync('git', ['ls-files', '--others', '--exclude-standard', '-z'], { cwd, timeout: 5000, maxBuffer: 1024 * 1024 });
256
+ if (listed.error || listed.status !== 0) return null;
257
+ const root = realpathSync(cwd);
258
+ const deadline = Date.now() + 2000;
259
+ let files = 0, bytes = 0, truncated = false;
260
+ for (const rel of String(listed.stdout).split('\0')) {
261
+ if (!rel) continue;
262
+ if (files >= 200 || Date.now() > deadline) { truncated = true; break; }
263
+ const p = resolve(root, rel);
264
+ if (p !== root && !p.startsWith(root + '/') && !p.startsWith(root + '\\')) continue;
265
+ let st;
266
+ try { st = lstatSync(p); } catch { continue; }
267
+ if (!st.isFile()) continue;
268
+ const limit = Math.min(st.size, 256 * 1024, 2 * 1024 * 1024 - bytes);
269
+ if (limit <= 0) { truncated = true; break; }
270
+ let fd;
271
+ try {
272
+ fd = openSync(p, 'r');
273
+ const buf = Buffer.alloc(Math.min(limit, 64 * 1024));
274
+ let read = 0, last = -1;
275
+ while (read < limit && Date.now() <= deadline) {
276
+ const n = readSync(fd, buf, 0, Math.min(buf.length, limit - read), read);
277
+ if (!n) break;
278
+ for (let i = 0; i < n; i++) if (buf[i] === 10) total++;
279
+ last = buf[n - 1];
280
+ read += n;
281
+ }
282
+ if (read && last !== 10) total++;
283
+ bytes += read;
284
+ if (read < st.size || Date.now() > deadline) truncated = true;
285
+ } catch {
286
+ // A changed file disappearing is ordinary repository churn. The caller
287
+ // falls back to prompt sizing only when the git probes themselves fail.
288
+ } finally { if (fd !== undefined) try { closeSync(fd); } catch {} }
289
+ files++;
290
+ if (bytes >= 2 * 1024 * 1024) { truncated = true; break; }
291
+ }
292
+ return { lines: total, truncated };
293
+ } catch {
294
+ return null;
295
+ }
296
+ }
297
+
298
+ export function resolveAutoEffort(lane, requested, prompt, audit, cwd = process.cwd()) {
299
+ if (requested !== 'auto') return { resolved: requested || null, basis: requested ? 'explicit' : 'none', scope: null, truncated: false };
300
+ const promptScope = prompt.length;
301
+ if (!audit) {
302
+ const bucket = promptScope < 4000 ? 'small' : 'large';
303
+ return { resolved: AUTO_EFFORT[lane][bucket], basis: 'prompt_chars', scope: promptScope, truncated: false };
304
+ }
305
+ const git = gitChangedLines(cwd);
306
+ const scope = Math.max(promptScope, git ? git.lines : 0);
307
+ // Audit is a stakes floor. Scope records the larger independently observed
308
+ // input, but never moves an audit above or below high.
309
+ return { resolved: 'high', basis: 'audit_floor', scope, truncated: !!(git && git.truncated) };
310
+ }
311
+
236
312
  // A model id or effort level becomes an argv element and, for codex, part of a
237
313
  // TOML value. Bounding the charset is what makes both safe: no leading dash (a
238
314
  // value cannot become a flag), no quote, space or control character (a value
@@ -579,6 +655,7 @@ export function laneConfig(here = dirname(fileURLToPath(import.meta.url))) {
579
655
  if (model !== undefined && badRouteValue('model', model)) return null;
580
656
  if (effort !== undefined) {
581
657
  if (badRouteValue('effort', effort)) return null;
658
+ if (effort.startsWith('auto') && effort !== 'auto') return null;
582
659
  if (!LANE_FLAGS[lane] || !LANE_FLAGS[lane].effort) return null; // a lane with no reasoning flag cannot have one pinned
583
660
  }
584
661
  defaults[lane] = { model: model ?? null, effort: effort ?? null };
@@ -618,6 +695,7 @@ function usage(msg) {
618
695
  own config. Every lane takes --model; every lane except qwen takes --effort.
619
696
  Levels are the vendor's own (agy low|medium|high, hermes none|minimal|...): an
620
697
  unknown level is rejected by the lane, and reported as that lane's exit code.
698
+ auto sizes per call: below 4,000 prompt characters is medium, otherwise high; a codex audit is always high. Auto never resolves above high.
621
699
  Pin them per lane instead of per call with "defaults" in bin/lanes.json.`);
622
700
  return USAGE;
623
701
  }
@@ -665,7 +743,7 @@ export async function doctor(run) {
665
743
  const d = defaults[lane] || {};
666
744
  // A disabled lane has no route worth reporting; saying "not pinned" there
667
745
  // reads as a finding about a lane that is not going to run.
668
- const route = !on ? '' : d.model || d.effort ? `route ${d.model || 'lane default'}/${d.effort || 'lane default'}` : 'route not pinned (inherits the lane\'s own config)';
746
+ const route = !on ? '' : d.effort === 'auto' ? `route ${d.model || 'lane default'}/auto (sized per call)` : d.model || d.effort ? `route ${d.model || 'lane default'}/${d.effort || 'lane default'}` : 'route not pinned (inherits the lane\'s own config)';
669
747
  let line = ` ${lane.padEnd(7)} ${on ? 'enabled ' : 'disabled'} ${bin ? 'binary ok' : 'binary MISSING'}${route ? ' ' + route : ''}`;
670
748
  if (on && !bin) bad++;
671
749
  if (on && bin && run) {
@@ -771,6 +849,7 @@ export async function main(argv) {
771
849
  if (v == null) continue;
772
850
  const bad = badRouteValue(kind, v);
773
851
  if (bad) return usage(bad);
852
+ if (kind === 'effort' && v.startsWith('auto') && v !== 'auto') return usage('--effort auto must be spelled exactly');
774
853
  }
775
854
  // qwen has no reasoning flag. Dropping --effort silently would leave the caller
776
855
  // believing a route that never happened, which is the defect this feature fixes.
@@ -784,13 +863,18 @@ export async function main(argv) {
784
863
  // gap this feature exists to close. A malformed lanes.json has no usable
785
864
  // defaults, so the flags stand alone and say so.
786
865
  const route = resolveRoute(lane, opts, cfg === null ? {} : cfg.defaults);
866
+ const sizing = resolveAutoEffort(lane, route.effort, prompt, opts.audit);
787
867
  opts.model = route.model;
788
- opts.effort = route.effort;
868
+ opts.effort = sizing.resolved;
789
869
  Object.assign(base, {
790
870
  model_requested: route.model,
791
871
  effort_requested: route.effort,
792
872
  model_source: route.model_source,
793
- effort_source: route.effort_source
873
+ effort_source: route.effort_source,
874
+ effort_resolved: sizing.resolved,
875
+ effort_basis: sizing.basis,
876
+ effort_scope: sizing.scope,
877
+ effort_truncated: sizing.truncated
794
878
  });
795
879
  const enabled = cfg === null ? null : cfg.enabled;
796
880
  if (enabled === null) {
@@ -810,6 +894,8 @@ export async function main(argv) {
810
894
  return UNAVAILABLE;
811
895
  }
812
896
 
897
+ if (route.effort === 'auto' && !opts.quiet) console.error(`cli-run: effort auto -> ${sizing.resolved} (${sizing.basis})`);
898
+
813
899
  const tmp = mkdtempSync(join(tmpdir(), 'cli-run-'));
814
900
  const before = opts.expectFile ? snapshotFile(resolve(opts.expectFile)) : null;
815
901
  try {
package/bin/cli.js CHANGED
@@ -22,8 +22,8 @@ import { windowsSpawnPlan } from './cli-run.mjs';
22
22
  // errors (exit 2) before anything is planned, so a typo like --dryy can never
23
23
  // turn a dry run into a real one.
24
24
  const SPEC = {
25
- level: 'value', ais: 'value', primary: 'value', dir: 'value', project: 'value', tools: 'value', apis: 'value',
26
- yes: 'bool', force: 'bool', dry: 'bool', 'dry-run': 'bool', 'no-install': 'bool', 'no-tools': 'bool', 'no-apis': 'bool', 'upgrade-runtime': 'bool', 'update-docs': 'bool', list: 'bool', help: 'bool', h: 'bool', version: 'bool', v: 'bool'
25
+ level: 'value', ais: 'value', primary: 'value', dir: 'value', project: 'value', tools: 'value', apis: 'value', plans: 'value',
26
+ yes: 'bool', force: 'bool', dry: 'bool', 'dry-run': 'bool', 'no-install': 'bool', 'no-tools': 'bool', 'no-apis': 'bool', 'effort-auto': 'bool', 'upgrade-runtime': 'bool', 'update-docs': 'bool', list: 'bool', help: 'bool', h: 'bool', version: 'bool', v: 'bool'
27
27
  };
28
28
  export function parseArgs(argv) {
29
29
  const out = {};
@@ -104,6 +104,8 @@ Flags
104
104
  --tools a,b companion tools to set up, all optional (default with --yes: codecalc only); --no-tools for none
105
105
  --apis a,b level 3 only: metered API keys you HOLD (anthropic,openai,google,xai,openrouter); --no-apis for none.
106
106
  Asked separately from the CLIs because a subscription is not an API key.
107
+ --plans a=plan,b=plan stated subscription plans for guidance; --plans none clears prior stated plans
108
+ --effort-auto consent to write auto effort defaults for selected high or max plan cli-run lanes
107
109
  --dir path where to write the docs and protocols (default ./ai-orchestrator)
108
110
  --project path the project root your agent runs from; subagent definitions go here (default: current directory,
109
111
  so set it: a run from your home folder otherwise drops the subagent files there)
@@ -129,6 +131,7 @@ if (flag('list')) {
129
131
  for (const a of AIS) {
130
132
  const here = a.bin ? (which(a.bin) ? 'installed' : 'not on PATH') : 'app';
131
133
  console.log(`${a.id.padEnd(13)} ${a.name}\n${''.padEnd(13)} level ${a.minLevel}+ · ${a.access} · ${here}\n${''.padEnd(13)} ${a.role}`);
134
+ if (a.plans) for (const p of a.plans) console.log(`${''.padEnd(13)} plan ${p.id}: ${p.name} (${p.headroom} headroom, checked ${p.checked}, ${p.source})`);
132
135
  }
133
136
  console.log('\nmetered API providers (--apis a,b, level 3 gateway only):');
134
137
  for (const prov of PROVIDERS) console.log(`${prov.id.padEnd(13)} ${prov.name} (variable name: ${prov.envName})`);
@@ -146,6 +149,29 @@ function bad(msg) {
146
149
  process.exit(2);
147
150
  }
148
151
 
152
+ function plansFromIds(raw, selected) {
153
+ if (raw === 'none') return {};
154
+ const out = {};
155
+ for (const part of raw.split(',')) {
156
+ const [id, plan, ...extra] = part.split('=');
157
+ if (!id || !plan || extra.length) bad(`--plans entry must be AI=plan: ${part}`);
158
+ if (Object.hasOwn(out, id)) bad(`--plans names ${id} more than once`);
159
+ const ai = byId[id];
160
+ if (!ai) bad(`--plans names unknown AI id: ${id}`);
161
+ if (!selected.includes(ai)) bad(`--plans names ${id}, which is not selected`);
162
+ const found = (ai.plans || []).find((p) => p.id === plan);
163
+ if (!found) bad(`--plans names unknown plan ${plan} for ${id}`);
164
+ out[id] = found;
165
+ }
166
+ return Object.fromEntries(Object.entries(out).sort(([a], [b]) => a.localeCompare(b)));
167
+ }
168
+
169
+ function plansFromManifest(manifest, selected) {
170
+ const raw = manifest && manifest.plans;
171
+ if (!raw || typeof raw !== 'object' || Array.isArray(raw)) return {};
172
+ return plansFromIds(Object.entries(raw).map(([id, plan]) => `${id}=${plan}`).join(','), selected);
173
+ }
174
+
149
175
  async function main() {
150
176
  // Says what this generates, not what it guarantees. The old line promised
151
177
  // routing this package does not perform: lane choice is an instruction an
@@ -278,11 +304,51 @@ async function main() {
278
304
  const projectBad = dirProblems(project);
279
305
  if (projectBad.length) bad('--project: ' + projectBad.join('; '));
280
306
 
307
+ // A prior manifest is read before planning so an omitted reconfiguration
308
+ // preserves consent and stated plans instead of silently clearing either.
309
+ const prev = readManifest(dir);
310
+ let plans;
311
+ let plansKept = false;
312
+ if (opt('plans') !== null) {
313
+ plans = plansFromIds(opt('plans'), selected);
314
+ } else if (prev) {
315
+ plans = plansFromManifest(prev, selected);
316
+ plansKept = Object.keys(plans).length > 0;
317
+ } else if (!yes) {
318
+ plans = {};
319
+ for (const ai of selected.filter((a) => a.plans)) {
320
+ const first = ai.plans[0];
321
+ console.log(`\nWhich ${ai.vendor} plan? (checked ${first.checked}, source ${first.source})`);
322
+ ai.plans.forEach((p, i) => console.log(` ${i + 1} ${p.name} (${p.headroom} headroom)`));
323
+ console.log(` ${ai.plans.length + 1} not sure`);
324
+ const answer = Number(await ask(`Plan [${ai.plans.length + 1}]: `, String(ai.plans.length + 1)));
325
+ if (!Number.isInteger(answer) || answer < 1 || answer > ai.plans.length + 1) bad('pick a listed plan number');
326
+ if (answer <= ai.plans.length) plans[ai.id] = ai.plans[answer - 1];
327
+ }
328
+ } else {
329
+ plans = {};
330
+ }
331
+
332
+ const eligibleAuto = selected.filter((a) => a.cliRun && plans[a.id] && ['high', 'max'].includes(plans[a.id].headroom)).map((a) => a.id).sort();
333
+ let effortAuto;
334
+ if (flag('effort-auto')) {
335
+ effortAuto = eligibleAuto;
336
+ } else if (prev && Array.isArray(prev.effortAuto)) {
337
+ effortAuto = prev.effortAuto.filter((id) => selected.some((a) => a.id === id && a.cliRun));
338
+ } else if (!yes && eligibleAuto.length) {
339
+ const answer = await ask(`\nSize reasoning effort per task automatically on ${eligibleAuto.join(', ')}? [y/N] `, 'n');
340
+ effortAuto = /^y/i.test(answer) ? eligibleAuto : [];
341
+ } else {
342
+ effortAuto = [];
343
+ }
344
+
281
345
  // 5. Plan
282
- const files = planFiles({ level, selected, primary, dir, project, tools, apis });
346
+ const files = planFiles({ level, selected, primary, dir, project, tools, apis, plans, effortAuto });
283
347
  const lvl = LEVELS.find((l) => l.id === level);
284
348
  const agentFiles = files.filter((f) => f.root === 'project');
285
- console.log(`\nPlan\n level ${lvl.id} ${lvl.name}\n access ${selected.map((a) => a.id).join(', ')}\n primary ${primary ? primary.id : 'none'}${primaryAutoPicked ? ` (chosen for you from ${candidates.map((a) => a.id).join(', ')}; pass --primary to decide it yourself)` : ''}\n tools ${tools.map((t) => t.id).join(', ') || 'none'}` + (level >= 3 ? `\n api keys ${apis.map((p) => p.id).join(', ') || 'none'}` : '') + `\n folder ${dir}\n project ${project}${agentFiles.length ? ' (' + agentFiles.length + ' subagent files go here)' : ''}\n files ${files.length}`);
349
+ const statedPlans = Object.entries(plans).map(([id, p]) => `${id}=${p.id}`).join(', ');
350
+ console.log(`\nPlan\n level ${lvl.id} ${lvl.name}\n access ${selected.map((a) => a.id).join(', ')}\n primary ${primary ? primary.id : 'none'}${primaryAutoPicked ? ` (chosen for you from ${candidates.map((a) => a.id).join(', ')}; pass --primary to decide it yourself)` : ''}\n tools ${tools.map((t) => t.id).join(', ') || 'none'}\n plans ${statedPlans || 'none stated'}` + (level >= 3 ? `\n api keys ${apis.map((p) => p.id).join(', ') || 'none'}` : '') + `\n folder ${dir}\n project ${project}${agentFiles.length ? ' (' + agentFiles.length + ' subagent files go here)' : ''}\n files ${files.length}`);
351
+ if (plansKept) console.log(' plans kept from the previous run');
286
352
  if (agentFiles.length && !opt('project')) {
287
353
  console.log(`\nNote: --project was not given, so the ${agentFiles.length} subagent file(s) go to the current directory (${project}). Pass --project to put them somewhere else.`);
288
354
  }
@@ -303,10 +369,11 @@ async function main() {
303
369
  }
304
370
 
305
371
  // Reconfiguration: compare what a previous run recorded with what was asked now.
306
- const prev = readManifest(dir);
307
372
  const changed = prev
308
373
  ? ['level', 'primary'].filter((k) => String(prev[k]) !== String(k === 'level' ? level : primary ? primary.id : null))
309
374
  .concat(['ais', 'tools', 'apis'].filter((k) => JSON.stringify(prev[k] || []) !== JSON.stringify({ ais: selected, tools, apis }[k].map((x) => x.id))))
375
+ .concat(JSON.stringify(prev.plans || {}) !== JSON.stringify(Object.fromEntries(Object.entries(plans).map(([id, p]) => [id, p.id]))) ? ['plans'] : [])
376
+ .concat(JSON.stringify((prev.effortAuto || []).slice().sort()) !== JSON.stringify(effortAuto.slice().sort()) ? ['effortAuto'] : [])
310
377
  : [];
311
378
 
312
379
  let written, skipped, upgraded, conflicts, unverifiable, docsUpdated, docsConflict, docsUnverifiable;
package/docs/catalog.md CHANGED
@@ -19,6 +19,10 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
19
19
  - **Install:** `npm install -g @anthropic-ai/claude-code@2.1.226`
20
20
  - **Sign in:** run `claude` once and sign in with your Anthropic account
21
21
  - **Reads rules from:** `CLAUDE.md` · subagents in `.claude/agents/`
22
+ - **Plans:**
23
+ - Claude Pro (base headroom, checked 2026-09-12): https://support.claude.com/en/articles/11049762-choose-a-claude-plan
24
+ - Claude Max 5x (high headroom, checked 2026-09-12): https://support.claude.com/en/articles/11049762-choose-a-claude-plan
25
+ - Claude Max 20x (max headroom, checked 2026-09-12): https://support.claude.com/en/articles/11049762-choose-a-claude-plan
22
26
  - **Built against:** 2.1.226 (the same number the npm pin uses)
23
27
 
24
28
  ### `codex` · Codex CLI (OpenAI, ChatGPT plan)
@@ -29,6 +33,10 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
29
33
  - **Sign in:** `codex login` (add `--device-auth` on a machine with no browser)
30
34
  - **Reads rules from:** `AGENTS.md`
31
35
  - **cli-run lane:** yes
36
+ - **Plans:**
37
+ - ChatGPT Plus (base headroom, checked 2026-09-12): https://learn.chatgpt.com/codex/pricing.md
38
+ - ChatGPT Pro 5x (high headroom, checked 2026-09-12): https://learn.chatgpt.com/codex/pricing.md
39
+ - ChatGPT Pro 20x (max headroom, checked 2026-09-12): https://learn.chatgpt.com/codex/pricing.md
32
40
  - **Built against:** 0.153.4 (the same number the npm pin uses)
33
41
 
34
42
  ### `agy` · Antigravity CLI `agy` (Google AI plan)
@@ -39,6 +47,10 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
39
47
  - **Sign in:** first run opens a device-code sign-in with your Google account
40
48
  - **Reads rules from:** `GEMINI.md` · subagents in `.agents/agents/`
41
49
  - **cli-run lane:** yes
50
+ - **Plans:**
51
+ - Google AI Pro (base headroom, checked 2026-09-12): https://gemini.google/subscriptions/
52
+ - Google AI Ultra 5x (high headroom, checked 2026-09-12): https://gemini.google/subscriptions/
53
+ - Google AI Ultra 20x (max headroom, checked 2026-09-12): https://gemini.google/subscriptions/
42
54
  - **Built against:** 1.1.27
43
55
  - **Note:** Gemini CLI was retired by Google in June 2026. agy is the successor. Do not install `gemini`.
44
56
 
@@ -49,6 +61,9 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
49
61
  - **Install:** vendor script (read it first): `https://x.ai/cli/install.sh`
50
62
  - **Sign in:** `grok login` (add `--device-auth` on a headless machine)
51
63
  - **cli-run lane:** yes
64
+ - **Plans:**
65
+ - SuperGrok (base headroom, checked 2026-09-12): https://x.ai/news/grok-build-cli
66
+ - X Premium Plus (base headroom, checked 2026-09-12): https://x.ai/news/grok-build-cli
52
67
  - **Built against:** 1.0.5
53
68
 
54
69
  ### `hermes` · Hermes Agent (Nous Research)
@@ -2,6 +2,10 @@
2
2
 
3
3
  Everything in Part 1, plus lanes. One agent stays the orchestrator; every other AI becomes a lane it calls from the terminal.
4
4
 
5
+ ## Plans and automatic effort
6
+
7
+ The installer can record plans with `--plans codex=pro-20x,agy=ultra-5x`. Plan headroom changes volume allocation only, never capability or the independent-review rule. `--effort-auto` is explicit consent to write `auto` for eligible high or max headroom CLI lanes. Auto resolves to medium or high from prompt size, and a codex audit is always high. It is a heuristic, not a measurement: name xhigh explicitly for security-critical or irreversible work.
8
+
5
9
  ## 1. Two kinds of lane
6
10
 
7
11
  **Lane A, subscription CLIs.** Claude Code, Codex, Antigravity, Grok, Hermes. Already paid for, $0 per call, used for interactive and agentic work. **Lane B, metered APIs.** Per token, used for programmatic bulk where a subscription CLI cannot serve. **Local.** A privacy lane, never a cost lane.
package/llms.txt CHANGED
@@ -6,6 +6,8 @@ Install and run: `npx model-orchestrator` (interactive), or headless: `npx model
6
6
 
7
7
  Levels: 1 beginner (one agent or chat app), 2 intermediate (several agent CLIs, each called through `cli-run`), 3 advanced (adds a virtual machine with a gateway and a scheduled audit job). On Claude Code the install also delegates execution to subagents by default and ships two hooks (UserPromptSubmit, SubagentStart) that inject the routing table every turn.
8
8
 
9
+ Claude Code plugin: `/plugin marketplace add aunysillyme/model-orchestrator`, then `/plugin install model-orchestrator@model-orchestrator`. It installs the two hooks and eight subagents; the routing rules still come from `npx model-orchestrator`, and the routing log (`route-metrics.mjs`) is npm-only.
10
+
9
11
  ## Docs
10
12
 
11
13
  - [README](https://github.com/aunysillyme/model-orchestrator/blob/main/README.md): what it writes, flags, principles, what is enforced versus instructed
@@ -16,7 +18,12 @@ Levels: 1 beginner (one agent or chat app), 2 intermediate (several agent CLIs,
16
18
 
17
19
  ## Reference
18
20
 
21
+ - [Claude Code plugin](https://github.com/aunysillyme/model-orchestrator/blob/main/plugin/README.md): installing the hooks and subagents with `/plugin install`, what the plugin reads, and what it leaves to the installer
19
22
  - [CLI runner](https://github.com/aunysillyme/model-orchestrator/blob/main/bin/README.md): `cli-run` lanes, exit codes, the route logged per run
23
+
24
+ ## Plans and automatic effort
25
+
26
+ Use `--plans AI=plan` to state a subscription plan and receive volume-allocation guidance. `--effort-auto` is opt-in and writes bounded `auto` effort only for selected high or max headroom CLI lanes. Auto is medium or high, never above high; name xhigh explicitly for security-critical or irreversible work.
20
27
  - [Templates](https://github.com/aunysillyme/model-orchestrator/blob/main/templates/README.md): the routing, tiers, task bundle and protocol files the installer renders
21
28
  - [Changelog](https://github.com/aunysillyme/model-orchestrator/blob/main/CHANGELOG.md): every release and the issue behind each fix
22
29
  - [Agent instructions](https://github.com/aunysillyme/model-orchestrator/blob/main/AGENTS.md): running the installer from an agent, and contributing
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.19",
3
+ "version": "0.1.21",
4
4
  "description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
5
5
  "type": "module",
6
6
  "bin": {
@@ -24,7 +24,8 @@
24
24
  "test": "node --test",
25
25
  "prepublishOnly": "npm test",
26
26
  "dry-run": "node bin/cli.js --yes --level 2 --ais claude-code,codex,grok --dir ./tmp-dry-run --dry",
27
- "gen:catalog": "node scripts/gen-catalog.js"
27
+ "gen:catalog": "node scripts/gen-catalog.js",
28
+ "gen:plugin": "node scripts/gen-plugin.js"
28
29
  },
29
30
  "engines": {
30
31
  "node": ">=18"
package/scripts/README.md CHANGED
@@ -3,3 +3,4 @@
3
3
  | File | Job |
4
4
  |---|---|
5
5
  | `gen-catalog.js` | regenerates `docs/catalog.md` AND the vendor compatibility table in `README.md` (between the `vendor-table` markers) from `src/catalog.js`; `npm run gen:catalog`. `test/catalog.test.js` fails if either generated surface disagrees with the catalog. |
6
+ | `gen-plugin.js` | regenerates the Claude Code plugin bundle in `plugin/` (agents, the two read-only hooks, `plugin.json`, `LICENSE`) from `templates/`, using the plan in `src/plugin.js`; `npm run gen:plugin`. `test/plugin.test.js` fails if the committed bundle disagrees. `plugin/README.md` and `plugin/hooks/hooks.json` are hand-owned. |
@@ -22,6 +22,10 @@ export function catalogMarkdown() {
22
22
  md += `### \`${a.id}\` · ${a.name}\n\n- **Kind:** ${a.kind} · **Access:** ${a.access} · **Lane:** ${a.lane} · **Level:** ${a.minLevel}+\n- **Wins at:** ${a.role}\n- **Install:** ${how}\n- **Sign in:** ${a.auth}\n`;
23
23
  if (a.rulesFile) md += `- **Reads rules from:** \`${a.rulesFile}\`` + (a.agentsDir ? ` · subagents in \`${a.agentsDir}/\`` : '') + '\n';
24
24
  if (a.cliRun) md += '- **cli-run lane:** yes\n';
25
+ if (a.plans) {
26
+ md += '- **Plans:**\n';
27
+ for (const p of a.plans) md += ` - ${p.name} (${p.headroom} headroom, checked ${p.checked}): ${p.source}\n`;
28
+ }
25
29
  if (a.builtAgainst) md += `- **Built against:** ${a.builtAgainst}` + (a.install.npm ? ' (the same number the npm pin uses)' : '') + '\n';
26
30
  if (a.note) md += `- **Note:** ${a.note}\n`;
27
31
  md += '\n';
@@ -0,0 +1,16 @@
1
+ #!/usr/bin/env node
2
+ // Regenerates the Claude Code plugin bundle under plugin/ from the installer's
3
+ // templates (src/plugin.js has the plan). The test suite checks the committed
4
+ // bundle matches, so an agent or hook edited in templates/ cannot ship to npm
5
+ // users and silently not to plugin users.
6
+ import { writeFileSync, mkdirSync } from 'node:fs';
7
+ import { join, dirname } from 'node:path';
8
+ import { planPluginFiles, PLUGIN_DIR } from '../src/plugin.js';
9
+
10
+ const files = planPluginFiles();
11
+ for (const f of files) {
12
+ const abs = join(PLUGIN_DIR, ...f.rel.split('/'));
13
+ mkdirSync(dirname(abs), { recursive: true });
14
+ writeFileSync(abs, f.content);
15
+ }
16
+ console.log('plugin/ regenerated: ' + files.length + ' files');
package/src/README.md CHANGED
@@ -5,5 +5,6 @@
5
5
  | `catalog.js` | the single list of levels and AIs. Add an AI here and the prompts, docs tables, delegation matrix, gateway config and installer all pick it up. Nothing else lists AIs. |
6
6
  | `detect.js` | PATH lookup for a binary, plus the few places vendor installers drop binaries without touching PATH. No shell-outs. |
7
7
  | `install.js` | pure planner: turns (level, selection, primary) into a list of files to write, rendering templates and computing every generated table. `writeFiles` is the only thing that touches disk. `activationSteps()` and `snippetFor()` live here so the terminal summary and the generated README render the same list. `subagentsLoadRules(primary)` gates every delegate-by-default render var (builder-by-default wording, the route-gate table, the inline-threshold note) on the one verified premise: a claude-code subagent loads CLAUDE.md. |
8
+ | `plugin.js` | the plan for the Claude Code plugin bundle in `plugin/`: `planPluginFiles()` renders the claude-code agents and the two read-only hooks from the same templates `install.js` uses, with plugin render vars (the installer's default rules paths, the setup hint for a project with no rules) instead of per-install ones. Pure; `scripts/gen-plugin.js` writes it and `test/plugin.test.js` checks the committed copy. |
8
9
  | `prompt.js` | line-buffered questions for the interactive path; piped answers are queued, EOF mid-prompt aborts instead of confirming a write. |
9
10
  | `render.js` | `{{KEY}}` substitution. Throws on an unknown key, so a template typo fails the test suite instead of shipping a literal placeholder. |
package/src/catalog.js CHANGED
@@ -27,6 +27,8 @@
27
27
  // "(chat only, no CLI)" catalog note, so the activation line stays
28
28
  // a sentence you can read once (#22)
29
29
  // chatSurface chat apps only: where the pasted block goes in that app
30
+ // plans optional known subscription plans: { id, name, headroom,
31
+ // source, checked }. Guidance uses headroom only, never prices.
30
32
 
31
33
  export const LEVELS = [
32
34
  {
@@ -70,6 +72,11 @@ export const AIS = [
70
72
  agentsDir: '.claude/agents',
71
73
  cliRun: false,
72
74
  models: { deep: 'opus', standard: 'sonnet', fast: 'haiku' },
75
+ plans: [
76
+ { id: 'pro', name: 'Claude Pro', headroom: 'base', source: 'https://support.claude.com/en/articles/11049762-choose-a-claude-plan', checked: '2026-09-12' },
77
+ { id: 'max-5x', name: 'Claude Max 5x', headroom: 'high', source: 'https://support.claude.com/en/articles/11049762-choose-a-claude-plan', checked: '2026-09-12' },
78
+ { id: 'max-20x', name: 'Claude Max 20x', headroom: 'max', source: 'https://support.claude.com/en/articles/11049762-choose-a-claude-plan', checked: '2026-09-12' }
79
+ ],
73
80
  // Verified at code.claude.com/docs/en/sub-agents (fetched 2026-09-10): "A
74
81
  // non-fork subagent's initial context contains: CLAUDE.md files: every
75
82
  // level of the CLAUDE.md hierarchy the main conversation loads ... The
@@ -95,6 +102,11 @@ export const AIS = [
95
102
  rulesFile: 'AGENTS.md',
96
103
  agentsDir: null,
97
104
  cliRun: true
105
+ , plans: [
106
+ { id: 'plus', name: 'ChatGPT Plus', headroom: 'base', source: 'https://learn.chatgpt.com/codex/pricing.md', checked: '2026-09-12' },
107
+ { id: 'pro-5x', name: 'ChatGPT Pro 5x', headroom: 'high', source: 'https://learn.chatgpt.com/codex/pricing.md', checked: '2026-09-12' },
108
+ { id: 'pro-20x', name: 'ChatGPT Pro 20x', headroom: 'max', source: 'https://learn.chatgpt.com/codex/pricing.md', checked: '2026-09-12' }
109
+ ]
98
110
  },
99
111
  {
100
112
  id: 'agy',
@@ -113,6 +125,11 @@ export const AIS = [
113
125
  agentsDir: '.agents/agents',
114
126
  cliRun: true,
115
127
  models: { deep: 'pro', standard: 'flash', fast: 'flash' },
128
+ plans: [
129
+ { id: 'ai-pro', name: 'Google AI Pro', headroom: 'base', source: 'https://gemini.google/subscriptions/', checked: '2026-09-12' },
130
+ { id: 'ultra-5x', name: 'Google AI Ultra 5x', headroom: 'high', source: 'https://gemini.google/subscriptions/', checked: '2026-09-12' },
131
+ { id: 'ultra-20x', name: 'Google AI Ultra 20x', headroom: 'max', source: 'https://gemini.google/subscriptions/', checked: '2026-09-12' }
132
+ ],
116
133
  note: 'Gemini CLI was retired by Google in June 2026. agy is the successor. Do not install `gemini`.'
117
134
  },
118
135
  {
@@ -131,6 +148,10 @@ export const AIS = [
131
148
  rulesFile: null,
132
149
  agentsDir: null,
133
150
  cliRun: true
151
+ , plans: [
152
+ { id: 'supergrok', name: 'SuperGrok', headroom: 'base', source: 'https://x.ai/news/grok-build-cli', checked: '2026-09-12' },
153
+ { id: 'x-premium-plus', name: 'X Premium Plus', headroom: 'base', source: 'https://x.ai/news/grok-build-cli', checked: '2026-09-12' }
154
+ ]
134
155
  },
135
156
  {
136
157
  id: 'hermes',
package/src/install.js CHANGED
@@ -37,14 +37,26 @@ function table(rows, header) {
37
37
  return [line(header), line(header.map(() => '---')), ...rows.map(line)].join('\n');
38
38
  }
39
39
 
40
- export function lanesTable(selected) {
40
+ export function lanesTable(selected, plans = {}) {
41
41
  const rows = selected.map((a) => [
42
42
  a.name,
43
43
  a.lane === 'A' ? 'A (subscription, $0 per call)' : a.lane === 'B' ? 'B (metered)' : a.lane === 'local' ? 'local' : 'chat',
44
44
  a.role,
45
- a.cliRun ? '`cli-run ' + a.id + '`' : a.bin ? '`' + a.bin + '`' : 'the app'
45
+ a.cliRun ? '`cli-run ' + a.id + '`' : a.bin ? '`' + a.bin + '`' : 'the app',
46
+ plans[a.id] ? `${plans[a.id].name} (${plans[a.id].headroom} headroom)` : 'not stated'
46
47
  ]);
47
- return table(rows, ['AI', 'Lane', 'Wins at', 'Call it with']);
48
+ return table(rows, ['AI', 'Lane', 'Wins at', 'Call it with', 'Plan']);
49
+ }
50
+
51
+ function planGuidance(selected, plans = {}) {
52
+ const lines = selected.filter((a) => plans[a.id]).map((a) => {
53
+ const p = plans[a.id];
54
+ const volume = p.headroom === 'base'
55
+ ? 'Keep this base-headroom lane for short second opinions. If it is primary, delegate volume to high or max headroom lanes.'
56
+ : 'Use this high or max headroom lane for volume: scoped well-specified builds, pre-ship second-family checks through cli-run, and first-pass research.';
57
+ return `- **${a.name}: ${p.name} (${p.headroom} headroom).** ${volume} Capability and independent-review rules are unchanged. Checked ${p.checked}.`;
58
+ });
59
+ return lines.length ? lines.join('\n') : 'State subscription plans with `--plans` to receive volume-allocation guidance. Capability and independent-review rules stay unchanged.';
48
60
  }
49
61
 
50
62
  export function installTable(selected) {
@@ -398,6 +410,7 @@ export function proofSteps(opts) {
398
410
 
399
411
  function vars(opts) {
400
412
  const { level, selected, primary } = opts;
413
+ const plans = opts.plans || {};
401
414
  const tools = opts.tools || [];
402
415
  const apis = opts.apis || [];
403
416
  const lvl = LEVELS.find((l) => l.id === level);
@@ -506,7 +519,8 @@ function vars(opts) {
506
519
  PRIMARY_FAST: primary && primary.models ? primary.models.fast : 'your cheapest model',
507
520
  AIS_LIST: selected.map((a) => '- ' + a.name + ': ' + a.role).join('\n'),
508
521
  AI_IDS: selected.map((a) => a.id).join(','),
509
- LANES_TABLE: lanesTable(selected),
522
+ LANES_TABLE: lanesTable(selected, plans),
523
+ PLAN_GUIDANCE: planGuidance(selected, plans),
510
524
  INSTALL_TABLE: installTable(selected),
511
525
  CLI_RUN_LANES: selected.filter((a) => a.cliRun).map((a) => a.id).join(', ') || 'none selected',
512
526
  GATEWAY_MODELS: gatewayModels(selected, apis),
@@ -532,7 +546,14 @@ function vars(opts) {
532
546
  AGENTS_LIST_LINE: claudeAgentIds().map((id) => '`' + id + '`').join(', '),
533
547
  RULES_FILE_REL: rulesFileRel,
534
548
  RULES_FILE_REL_JSON: JSON.stringify(rulesFileRel),
535
- TASK_BUNDLE_REL_JSON: JSON.stringify(taskBundleRel)
549
+ TASK_BUNDLE_REL_JSON: JSON.stringify(taskBundleRel),
550
+ // route-gate.mjs takes a candidate list so the plugin bundle (src/plugin.js)
551
+ // can render the installer's default locations from the same template. An
552
+ // install knows its one rules file, and wrote it, so it needs no hint.
553
+ RULES_CANDIDATES_JSON: JSON.stringify([rulesFileRel]),
554
+ SETUP_HINT_JSON: JSON.stringify(''),
555
+ SETUP_NOTICE_JSON: JSON.stringify(''),
556
+ CONTEXT_SUFFIX_JSON: JSON.stringify('')
536
557
  };
537
558
  }
538
559
 
@@ -598,7 +619,7 @@ export function planFiles(opts) {
598
619
  JSON.stringify(
599
620
  {
600
621
  enabled: selected.filter((a) => a.cliRun).map((a) => a.id),
601
- defaults: {},
622
+ defaults: Object.fromEntries((opts.effortAuto || []).map((lane) => [lane, { effort: 'auto' }])),
602
623
  note: 'Lanes cli-run may call. Edit to enable or disable a lane. A lane not listed here exits 13 (unavailable).',
603
624
  defaultsNote: 'Pin what a lane runs with, so the route in your docs is the route that runs: "defaults": {"codex": {"model": "gpt-6-astra", "effort": "high"}}. Left empty, a lane inherits its own config file, which cli-run cannot see and does not guess. `--model` and `--effort` override this per call, and `--doctor` prints what each lane is pinned to. Every lane takes a model; every lane except qwen takes an effort.'
604
625
  },
@@ -633,6 +654,8 @@ export function planFiles(opts) {
633
654
  primary: primary ? primary.id : null,
634
655
  tools: (opts.tools || []).map((t) => t.id),
635
656
  apis: (opts.apis || []).map((p) => p.id),
657
+ ...(Object.keys(opts.plans || {}).length ? { plans: Object.fromEntries(Object.entries(opts.plans).sort(([a], [b]) => a.localeCompare(b)).map(([id, p]) => [id, p.id])) } : {}),
658
+ ...((opts.effortAuto || []).length ? { effortAuto: [...opts.effortAuto].sort() } : {}),
636
659
  dir: resolve(opts.dir || 'ai-orchestrator'),
637
660
  project: resolve(opts.project || process.cwd()),
638
661
  files: fileHashes,
package/src/plugin.js ADDED
@@ -0,0 +1,83 @@
1
+ // The Claude Code plugin bundle under plugin/. Every file this plans is
2
+ // generated from the templates the installer renders, so the plugin carries
3
+ // no second copy of an agent or a hook: `npm run gen:plugin` writes them and
4
+ // test/plugin.test.js fails when the committed bundle drifts from this plan.
5
+ // plugin/README.md and plugin/hooks/hooks.json are plugin-only and hand-owned.
6
+ import { readFileSync } from 'node:fs';
7
+ import { join, dirname } from 'node:path';
8
+ import { fileURLToPath } from 'node:url';
9
+ import { render } from './render.js';
10
+ import { TEMPLATES, GENERATOR_VERSION, claudeAgentIds } from './install.js';
11
+
12
+ const HERE = dirname(fileURLToPath(import.meta.url));
13
+ export const ROOT = join(HERE, '..');
14
+ export const PLUGIN_DIR = join(ROOT, 'plugin');
15
+ export const PLUGIN_NAME = 'model-orchestrator';
16
+ export const REPO_URL = 'https://github.com/aunysillyme/model-orchestrator';
17
+
18
+ // The installer's default --dir, and the rules file each level writes there:
19
+ // ROUTING.md at level 2 and 3, ORCHESTRATOR.md at level 1. Level 2 first, so
20
+ // a project that moved up a level reads the newer file.
21
+ export const DEFAULT_RULES = ['ai-orchestrator/ROUTING.md', 'ai-orchestrator/ORCHESTRATOR.md'];
22
+ export const DEFAULT_TASK_BUNDLE = 'ai-orchestrator/TASK_BUNDLE.md';
23
+
24
+ // route-metrics.mjs is not here on purpose: it appends a routing log to disk,
25
+ // and the plugin ships only hooks that read. `npx model-orchestrator` still
26
+ // installs it.
27
+ export const PLUGIN_HOOKS = ['route-gate.mjs', 'subagent-context.mjs'];
28
+
29
+ // Files in plugin/ that are written by hand, not by planPluginFiles().
30
+ export const HAND_OWNED = ['README.md', 'hooks/hooks.json'];
31
+
32
+ export function pluginVars() {
33
+ return {
34
+ PRIMARY_NAME: 'Claude Code',
35
+ RULES_FILE_REL: DEFAULT_RULES.join(' or '),
36
+ RULES_FILE_REL_JSON: JSON.stringify(DEFAULT_RULES[0] + ' (' + DEFAULT_RULES[1] + ' on a level 1 install)'),
37
+ TASK_BUNDLE_REL_JSON: JSON.stringify(DEFAULT_TASK_BUNDLE),
38
+ RULES_CANDIDATES_JSON: JSON.stringify(DEFAULT_RULES),
39
+ SETUP_HINT_JSON: JSON.stringify(
40
+ 'This project has no model-orchestrator routing rules yet. Running `npx model-orchestrator` in the project root writes them (' +
41
+ DEFAULT_RULES[0] +
42
+ '), and this hook reads them from the next prompt on. Until then, pick the lane yourself, and name that command if the user asks about routing.'
43
+ ),
44
+ SETUP_NOTICE_JSON: JSON.stringify(
45
+ 'model-orchestrator: this project has no routing rules yet, so the plugin has no routing table to inject. Run `npx model-orchestrator` in the project root to write them; the plugin reads them from your next prompt.'
46
+ ),
47
+ CONTEXT_SUFFIX_JSON: JSON.stringify(
48
+ '\nThe model-orchestrator plugin installs the agents named above as model-orchestrator:<name>, for example model-orchestrator:builder.'
49
+ )
50
+ };
51
+ }
52
+
53
+ export function pluginManifest() {
54
+ const n = claudeAgentIds().length;
55
+ return {
56
+ name: PLUGIN_NAME,
57
+ version: GENERATOR_VERSION,
58
+ description:
59
+ 'Routing for Claude Code: a hook injects your project\'s routing table on every prompt, so each task goes to the right subagent tier and fewer tokens go to the most expensive model. Ships ' +
60
+ n +
61
+ ' subagents across three model tiers.',
62
+ author: { name: 'model-orchestrator maintainers', url: REPO_URL },
63
+ homepage: REPO_URL + '#readme',
64
+ repository: REPO_URL,
65
+ license: 'MIT',
66
+ keywords: ['routing', 'subagents', 'hooks', 'model-router', 'token-optimization', 'delegation']
67
+ };
68
+ }
69
+
70
+ // Pure: reads templates, writes nothing. Paths are posix, relative to plugin/.
71
+ export function planPluginFiles() {
72
+ const v = pluginVars();
73
+ const files = [];
74
+ files.push({ rel: '.claude-plugin/plugin.json', content: JSON.stringify(pluginManifest(), null, 2) + '\n' });
75
+ for (const id of claudeAgentIds()) {
76
+ files.push({ rel: 'agents/' + id + '.md', content: render(readFileSync(join(TEMPLATES, 'agents', 'claude-code', id + '.md'), 'utf8'), v) });
77
+ }
78
+ for (const hook of PLUGIN_HOOKS) {
79
+ files.push({ rel: 'hooks/' + hook, content: render(readFileSync(join(TEMPLATES, 'agents', 'snippets', hook), 'utf8'), v) });
80
+ }
81
+ files.push({ rel: 'LICENSE', content: readFileSync(join(ROOT, 'LICENSE'), 'utf8') });
82
+ return files;
83
+ }
@@ -14,3 +14,5 @@ One per tier, plus two checks and two agents with no file-editing tools: `findin
14
14
  | reader | fast | haiku | low | reads and digests many files or notes; read-only |
15
15
 
16
16
  Aliases resolve to the newest model in each family, so a version bump needs no edit here. Each agent carries its own token-discipline rule; the `effort` field is the third cost lever. None of `done-verifier`, `finding-verifier`, `code-reviewer` or `reader` carries `Write` or `Edit` in its `tools:` line. `reader` is read-only by tool grant as well: it carries no `Bash`. `done-verifier`, `finding-verifier` and `code-reviewer` do carry `Bash`, for their probes and checks (`git log`, `grep`, `wc -l`, `test -f`); nothing in that grant stops any of them from running a command that changes state, so staying read-only there is a rule in each one's prompt, not a restriction on the tool, and each file says so.
17
+
18
+ Every agent names its tools explicitly, so none inherits every tool the session has: `builder` carries `Read, Write, Edit, Glob, Grep, Bash` (it changes files and runs checks), `deep-planner` carries `Read, Glob, Grep` (it plans and never edits), and `live-researcher` carries `WebSearch, WebFetch` (it answers from the web, not local files). The same files ship in the Claude Code plugin under `plugin/agents/`, generated from this folder.
@@ -1,6 +1,7 @@
1
1
  ---
2
2
  name: builder
3
3
  description: Executes builds by default on this router, including the main build, from a brief the orchestrator wrote. Use for writing code, editing files, wiring configs, running commands, and implementing a plan the orchestrator briefed. Do not use for open-ended architecture questions or bulk classification; those still go to deep-planner or bulk-worker.
4
+ tools: Read, Write, Edit, Glob, Grep, Bash
4
5
  model: sonnet
5
6
  effort: high
6
7
  ---
@@ -1,6 +1,7 @@
1
1
  ---
2
2
  name: deep-planner
3
3
  description: Ambiguous or high-stakes thinking. Use for architecture design, strategy, planning multi-step projects, hard debugging where the cause is unknown, and any "figure out what to even do" request. Do not use for well-specified execution or bulk work.
4
+ tools: Read, Glob, Grep
4
5
  model: opus
5
6
  effort: xhigh
6
7
  ---
@@ -1,6 +1,7 @@
1
1
  ---
2
2
  name: live-researcher
3
3
  description: Real-time information. Use for anything that needs current data such as latest news, current API docs or pricing, or recent events. Do not use for questions answerable from local files or general knowledge.
4
+ tools: WebSearch, WebFetch
4
5
  model: sonnet
5
6
  effort: medium
6
7
  ---
@@ -14,36 +14,93 @@
14
14
  import { statSync, openSync, readSync, closeSync, realpathSync } from 'node:fs';
15
15
  import { join, isAbsolute } from 'node:path';
16
16
 
17
- // Rendered at install time from the level and directory the user chose.
18
- // Never hardcoded: a level 1 install points this at ORCHESTRATOR.md, level
19
- // 2+ at ROUTING.md, and a --dir outside the project resolves to an absolute
20
- // path instead of a relative one.
21
- const RULES_FILE_REL = {{RULES_FILE_REL_JSON}};
17
+ // One template, two renders. The installer renders a one-element list from
18
+ // the level and directory the user chose: a level 1 install points it at
19
+ // ORCHESTRATOR.md, level 2+ at ROUTING.md, and a --dir outside the project
20
+ // resolves to an absolute path instead of a relative one. The Claude Code
21
+ // plugin (plugin/, generated by scripts/gen-plugin.js) has no install step to
22
+ // render from, so it lists the installer's default locations and reads the
23
+ // first one that exists.
24
+ const RULES_CANDIDATES = {{RULES_CANDIDATES_JSON}};
25
+ const RULES_FILE_REL = RULES_CANDIDATES.join(' or ');
26
+ // Empty in an installer render, whose rules file was written by the same run.
27
+ // The plugin renders a next step for a project that has no rules yet, so a
28
+ // missing file is never a silent no-op: SETUP_HINT goes to the model on every
29
+ // prompt, SETUP_NOTICE to the user once, at session start.
30
+ const SETUP_HINT = {{SETUP_HINT_JSON}};
31
+ const SETUP_NOTICE = {{SETUP_NOTICE_JSON}};
32
+ // Appended after the table. Empty in an installer render.
33
+ const CONTEXT_SUFFIX = {{CONTEXT_SUFFIX_JSON}};
34
+ // Passed only by the plugin's SessionStart entry in hooks/hooks.json.
35
+ const SESSION_START = process.argv.includes('--session-start');
22
36
 
23
37
  const MAX_READ = 64 * 1024; // bounded read: this is a rules file, not a log
24
38
  const MAX_CONTEXT = 4000; // bounded injection: a table, not the whole file
39
+ // Bounded output, whatever produced it. The table is capped above, but a
40
+ // fallback or missing-rules message embeds resolved paths and error text
41
+ // whose length the project directory controls, and Claude Code caps hook
42
+ // output strings at 10,000 characters. Every string this hook emits passes
43
+ // through bound().
44
+ const MAX_OUTPUT = 8000;
25
45
  const STDIN_DRAIN_MS = 250; // hard cap: never let an open, never-closed stdin pipe hold this hook open
26
46
  const START = '<!-- route-gate:start -->';
27
47
  const END = '<!-- route-gate:end -->';
28
48
 
49
+ function bound(text) {
50
+ return text.length > MAX_OUTPUT ? text.slice(0, MAX_OUTPUT - 3) + '...' : text;
51
+ }
52
+
29
53
  function fallback(reason) {
30
54
  return 'route-gate: ' + reason + '. Pick the lane before acting: read ' + RULES_FILE_REL + ' yourself.';
31
55
  }
32
56
 
33
- function resolveRulesPath() {
34
- if (isAbsolute(RULES_FILE_REL)) return RULES_FILE_REL;
57
+ function projectRoot() {
35
58
  const projectDir = process.env.CLAUDE_PROJECT_DIR;
36
59
  if (!projectDir) return null;
37
60
  // Resolve through whatever part of the project dir already exists, so a
38
61
  // symlinked project folder still resolves to the real path the rules file
39
62
  // was written under.
40
- let root = projectDir;
41
63
  try {
42
- root = realpathSync(projectDir);
64
+ return realpathSync(projectDir);
43
65
  } catch {
44
- /* keep the unresolved value; the read below reports the real failure */
66
+ return projectDir; // keep the unresolved value; the read below reports the real failure
45
67
  }
46
- return join(root, RULES_FILE_REL);
68
+ }
69
+
70
+ function resolveRulesPath(rel, root) {
71
+ if (isAbsolute(rel)) return rel;
72
+ return root ? join(root, rel) : null;
73
+ }
74
+
75
+ // ENOENT and ENOTDIR mean "nothing here, try the next candidate". Anything
76
+ // else (a directory, a FIFO, a permission error) means something IS at that
77
+ // path, so it is chosen and readBounded reports the real problem, rather than
78
+ // this hook quietly reading a different file than the one the project has.
79
+ function present(path) {
80
+ try {
81
+ statSync(path);
82
+ return true;
83
+ } catch (e) {
84
+ const code = e && e.code;
85
+ return !(code === 'ENOENT' || code === 'ENOTDIR');
86
+ }
87
+ }
88
+
89
+ // { path } to read, { error } when the rules cannot be located at all, or
90
+ // { missing } (every path looked at) when a multi-candidate render finds
91
+ // nothing. A one-candidate render never returns { missing }: it keeps the
92
+ // installer's original behaviour exactly, and the read reports what is wrong.
93
+ function locateRules() {
94
+ const root = projectRoot();
95
+ const paths = [];
96
+ for (const rel of RULES_CANDIDATES) {
97
+ const p = resolveRulesPath(rel, root);
98
+ if (!p) return { error: 'CLAUDE_PROJECT_DIR is not set, so ' + RULES_FILE_REL + ' could not be located' };
99
+ paths.push(p);
100
+ }
101
+ if (paths.length === 1) return { path: paths[0] };
102
+ const found = paths.find(present);
103
+ return found ? { path: found } : { missing: paths };
47
104
  }
48
105
 
49
106
  // Bounded, regular-file-only read. statSync (not lstatSync) follows a
@@ -76,21 +133,31 @@ function readBounded(path) {
76
133
  }
77
134
 
78
135
  function computeContext() {
79
- const path = resolveRulesPath();
80
- if (!path) return fallback('CLAUDE_PROJECT_DIR is not set, so ' + RULES_FILE_REL + ' could not be located');
136
+ const loc = locateRules();
137
+ if (loc.error) return fallback(loc.error);
138
+ if (loc.missing) return 'route-gate: no routing rules file at ' + loc.missing.join(' or ') + '. ' + SETUP_HINT;
81
139
 
82
140
  let text;
83
141
  try {
84
- text = readBounded(path);
142
+ text = readBounded(loc.path);
85
143
  } catch (e) {
86
144
  return fallback((e && e.message) || String(e));
87
145
  }
88
146
 
89
147
  const s = text.indexOf(START);
90
148
  const e = s === -1 ? -1 : text.indexOf(END, s);
91
- if (s === -1 || e === -1) return fallback(path + ' has no route-gate block');
149
+ if (s === -1 || e === -1) return fallback(loc.path + ' has no route-gate block');
150
+
151
+ return text.slice(s, e + END.length).slice(0, MAX_CONTEXT) + CONTEXT_SUFFIX;
152
+ }
92
153
 
93
- return text.slice(s, e + END.length).slice(0, MAX_CONTEXT);
154
+ // The message a person sees at session start, or null when there is nothing
155
+ // to say: an installer render (no notice), rules that exist, or a project
156
+ // root that cannot be located.
157
+ function sessionStartNotice() {
158
+ if (!SETUP_NOTICE) return null;
159
+ const loc = locateRules();
160
+ return loc.missing ? SETUP_NOTICE : null;
94
161
  }
95
162
 
96
163
  // Drain stdin without ever blocking on it. A bare `readFileSync(0)` waits
@@ -130,20 +197,35 @@ function drainStdin(timeoutMs) {
130
197
  });
131
198
  }
132
199
 
133
- let additionalContext;
134
- try {
135
- additionalContext = computeContext();
136
- } catch (err) {
137
- additionalContext = fallback('route-gate.mjs failed unexpectedly (' + ((err && err.message) || err) + ')');
138
- }
139
-
140
- drainStdin(STDIN_DRAIN_MS).then(() => {
141
- const payload = JSON.stringify({
200
+ let payload = null;
201
+ if (SESSION_START) {
202
+ let notice = null;
203
+ try {
204
+ notice = sessionStartNotice();
205
+ } catch {
206
+ notice = null; // a session-start notice is a courtesy; never fail a session over it
207
+ }
208
+ if (notice) payload = JSON.stringify({ systemMessage: bound(notice) });
209
+ } else {
210
+ let additionalContext;
211
+ try {
212
+ additionalContext = computeContext();
213
+ } catch (err) {
214
+ additionalContext = fallback('route-gate.mjs failed unexpectedly (' + ((err && err.message) || err) + ')');
215
+ }
216
+ payload = JSON.stringify({
142
217
  hookSpecificOutput: {
143
218
  hookEventName: 'UserPromptSubmit',
144
- additionalContext
219
+ additionalContext: bound(additionalContext)
145
220
  }
146
221
  });
222
+ }
223
+
224
+ drainStdin(STDIN_DRAIN_MS).then(() => {
225
+ if (payload === null) {
226
+ process.exit(0);
227
+ return;
228
+ }
147
229
  // Exit only after the write's callback fires, so a buffered write to a
148
230
  // pipe (the common case on Windows, and possible anywhere output exceeds
149
231
  // one write's worth) is not truncated by an exit that races ahead of it.
@@ -19,6 +19,7 @@ Enabled lanes (edit `bin/lanes.json`): {{CLI_RUN_LANES}}
19
19
  node bin/cli-run.mjs <grok|codex|agy|hermes|qwen> "<prompt>" [--brief FILE] [--timeout SECS] [--quiet]
20
20
  node bin/cli-run.mjs codex --audit "<prompt>" # read-only sandbox, the audit shape
21
21
  node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
22
+ node bin/cli-run.mjs codex "<prompt>" --effort auto # medium or high, bounded heuristic
22
23
  node bin/cli-run.mjs qwen [--safe-mode] "<prompt>" # qwen-only flag
23
24
  ```
24
25
 
@@ -98,6 +99,7 @@ Three rules that keep this honest:
98
99
 
99
100
  - **A level `cli-run` does not recognise is not rejected here.** Levels are the vendor's, they change, and guessing the valid set would date this tool. An unknown level is refused by the lane and surfaces as that lane's own exit code and stderr.
100
101
  - **`--effort` on qwen is a usage error, not a silent drop.** A flag that vanishes leaves you believing a route that never ran.
102
+ - **`--effort auto` is bounded.** It uses prompt size, or an audit's changed-file evidence, and resolves only medium or high. An audit is always high. Auto is a heuristic, not a measurement: name `xhigh` explicitly for security-critical or irreversible work.
101
103
  - **Values are charset-bounded** (letters, digits, and `. _ : @ / + -`, no leading dash, 64 characters). A model id becomes an argv element and, on codex, part of a TOML value; bounding it is what stops either from being escaped.
102
104
 
103
105
  ## Permissions are a separate layer
@@ -106,7 +108,7 @@ Three rules that keep this honest:
106
108
 
107
109
  ## Log
108
110
 
109
- `~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, rc, the lane's own exit code, signal, seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, the route (`model_requested`, `effort_requested`, and `model_source` / `effort_source`, each one of `flag`, `lanes.json` or `lane_default`), and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument, and so does "we route audits at high effort".
111
+ `~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, rc, the lane's own exit code, signal, seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, the route (`model_requested`, `effort_requested`, and `model_source` / `effort_source`, each one of `flag`, `lanes.json` or `lane_default`), auto-sizing evidence (`effort_resolved`, `effort_basis`, `effort_scope`), and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). `effort_basis` is exactly one of `explicit`, `prompt_chars`, `audit_floor`, or `none`. Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument, and so does "we route audits at high effort".
110
112
 
111
113
  The log records what was **requested**, on every record including a run refused before the lane started. It does not record an actual. Reporting is inconsistent: grok returns a `modelUsage` block naming a model, the other four lanes return nothing of the kind, so an `actual` field would be populated for one lane and empty for four. It would also be a provider-supplied string, and this log holds fixed codes and bounded caller-supplied values only. `model_source: "lane_default"` is the honest way to say this run inherited something invisible from here.
112
114
 
@@ -6,6 +6,10 @@ Generated {{DATE}} from the AIs you said you have: `{{AI_IDS}}`.
6
6
 
7
7
  {{LANES_TABLE}}
8
8
 
9
+ ## Plan guidance
10
+
11
+ {{PLAN_GUIDANCE}}
12
+
9
13
  ## Task → lane
10
14
 
11
15
  | Task type | Pick | Why |
@@ -6,6 +6,10 @@ Your lanes:
6
6
 
7
7
  {{LANES_TABLE}}
8
8
 
9
+ ## Plan guidance
10
+
11
+ {{PLAN_GUIDANCE}}
12
+
9
13
  Two kinds of lane. **Lane A** = subscription CLIs: $0 marginal, already paid for, used for interactive and agentic work. **Lane B** = metered APIs: per token, used for programmatic bulk where a subscription CLI cannot serve. **Local** = stays on the machine; a privacy lane, never a cost lane.
10
14
 
11
15
  Rule of thumb: never spend a frontier token on a task a cheap tier finishes correctly. Escalate on signal (low confidence, explicit complexity, a failed verification), not by default. And an external lane must earn the hop with a real strength; when in doubt, stay in-house.
@@ -15,6 +15,8 @@ Non-primary lanes are owned by `DELEGATION_MATRIX.md`.
15
15
 
16
16
  Tier sets the price per token. Token discipline sets how many tokens. **Effort sets how hard each call thinks.**
17
17
 
18
+ `cli-run --effort auto` is a heuristic, not a measurement: prompts below 4,000 characters resolve to medium and longer prompts resolve to high. A codex `--audit` always resolves to high because stakes set the floor. Auto never resolves above high. Name `xhigh` explicitly for a security-critical or irreversible audit.
19
+
18
20
  | Agent | Tier | Effort | Why |
19
21
  |---|---|---|---|
20
22
  | deep-planner | deep | xhigh | judges every build twice; expensive to get wrong |