model-orchestrator 0.1.13 → 0.1.15
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +26 -0
- package/CHANGELOG.md +41 -1
- package/README.md +95 -6
- package/bin/cli-run.mjs +144 -19
- package/bin/cli.js +1 -1
- package/docs/audit-brief.md +21 -0
- package/docs/part-1-beginner.md +3 -1
- package/docs/part-2-intermediate.md +8 -2
- package/llms.txt +27 -0
- package/package.json +32 -4
- package/src/README.md +1 -1
- package/src/catalog.js +9 -1
- package/src/install.js +188 -5
- package/templates/README.md +1 -1
- package/templates/agents/README.md +3 -1
- package/templates/agents/agy/README.md +2 -2
- package/templates/agents/agy/done-verifier.md +35 -0
- package/templates/agents/agy/finding-verifier.md +29 -0
- package/templates/agents/agy/reader.md +22 -0
- package/templates/agents/claude-code/README.md +7 -4
- package/templates/agents/claude-code/builder.md +6 -1
- package/templates/agents/claude-code/done-verifier.md +44 -0
- package/templates/agents/claude-code/finding-verifier.md +43 -0
- package/templates/agents/claude-code/reader.md +26 -0
- package/templates/agents/snippets/claude-code.md +9 -3
- package/templates/agents/snippets/route-gate.mjs +151 -0
- package/templates/agents/snippets/settings.hooks.snippet.json +26 -0
- package/templates/agents/snippets/subagent-context.mjs +76 -0
- package/templates/beginner/ORCHESTRATOR.md +4 -3
- package/templates/common/TASK_BUNDLE.md +2 -2
- package/templates/common/protocols/build-protocol.md +13 -3
- package/templates/intermediate/CLI-RUN.md +40 -3
- package/templates/intermediate/ROUTING.md +15 -11
- package/templates/intermediate/TIERS.md +44 -1
package/AGENTS.md
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# AGENTS.md
|
|
2
|
+
|
|
3
|
+
Two audiences: an agent that wants to USE this package for a project, and an agent that is working ON this repository.
|
|
4
|
+
|
|
5
|
+
## Using this package from an agent
|
|
6
|
+
|
|
7
|
+
model-orchestrator writes routing rules, subagents and a CLI lane runner so an agent sends each task to the right model, subagent or CLI and spends fewer frontier tokens. It is not a proxy or gateway. Headless use:
|
|
8
|
+
|
|
9
|
+
- `npx model-orchestrator --list` prints every supported AI id.
|
|
10
|
+
- `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project <repo> --dir <repo>/ai-orchestrator --dry-run` prints the plan and writes nothing.
|
|
11
|
+
- Drop `--dry-run` to write it. Existing files are never overwritten without `--force`; activation snippets (for example `CLAUDE.snippet.md`, `settings.hooks.snippet.json`) are written for a person or agent to merge.
|
|
12
|
+
- The generated `README.md` in `--dir` lists what to copy where and one smoke command to prove the rules took.
|
|
13
|
+
- A summary for LLMs, with links to every doc: [`llms.txt`](llms.txt).
|
|
14
|
+
|
|
15
|
+
## Working on this repository
|
|
16
|
+
|
|
17
|
+
The files the installer writes for end users live under `templates/`.
|
|
18
|
+
|
|
19
|
+
- Read `CONTRIBUTING.md` first, then `src/README.md` (the catalog drives everything) and `docs/audit-brief.md` (the threat model and what has already been attacked).
|
|
20
|
+
- Run `npm test` before proposing a change and quote the count and the exit code; the suite prints the current number.
|
|
21
|
+
- Everything renders from `src/catalog.js`. Add an AI or a tool there, not in a template. Templates carry no logic.
|
|
22
|
+
- Never put a value that looks like a credential anywhere in this repo, including tests and examples. Environment variable names only.
|
|
23
|
+
- `bin/cli.js` writes only inside `--dir` and `--project`, never over a document without `--force`, and never runs a vendor script. A change that weakens any of those will be refused in review; the tests that hold them are in `test/install.test.js` and `test/cli.test.js`.
|
|
24
|
+
- `bin/cli-run.mjs` must exit non-zero when a lane produced nothing. Every judge has a red case in `test/judges.test.js`; add one before you change a judge.
|
|
25
|
+
- Prose in this repo uses no em dashes (`test/prose.test.js` enforces it).
|
|
26
|
+
- Why these rules exist: each one is the fix for a failure that reached an audit or CI. `CHANGELOG.md` names the issue behind each.
|
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,44 @@ All notable changes to this project are documented here. The format follows [Kee
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [0.1.15] - 2026-09-10
|
|
8
|
+
|
|
9
|
+
The portable parts of a live routing revision, delegate by default, gated on one verified fact rather than a guess: [code.claude.com/docs/en/sub-agents](https://code.claude.com/docs/en/sub-agents) states that a non-fork Claude Code subagent's initial context includes "every level of the CLAUDE.md hierarchy the main conversation loads", and that the built-in Explore and Plan agents skip it. No other lane in this catalog has that documented, so everything below is gated on `subagentsLoadRules(primary)`, currently true for claude-code alone; every other primary keeps its original wording unchanged.
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- **`subagentsLoadRules: true` on the claude-code catalog entry**, with the doc quote as its comment. Drives every new render var below through `src/install.js`; nothing here is a template branch, per the house rule that templates carry no logic.
|
|
14
|
+
- **Builder executes by default, on claude-code.** `ROUTING.md` rule 5, its "Who builds" section, the "Add an endpoint" example, and the claude-code `CLAUDE.snippet.md` now say: the orchestrator plans, briefs, verifies and talks to the human; it stays inline only when (a) the brief would cost as much as the work, (b) the task needs this conversation's own context, or (c) it is the human's decision or the final verification of delegated work. Every other primary keeps "the orchestrator builds it directly."
|
|
15
|
+
- **Two hooks, claude-code only: `route-gate.mjs` (`UserPromptSubmit`) and `subagent-context.mjs` (`SubagentStart`)**, written to `.claude/hooks/`, wired by a `settings.hooks.snippet.json` the user merges into `.claude/settings.json` themselves, never written over one they have. `route-gate.mjs` reads the `<!-- route-gate:start -->...<!-- route-gate:end -->` block out of the rendered routing rules file and injects it every turn, so the table is read from the one place it is generated, not recited from memory; a missing file or block still exits 0 with a one-line fallback naming the path it looked for. `subagent-context.mjs` injects a static reminder of where the rules and `TASK_BUNDLE.md` live and that a delegate does not route further or verify its own work as final. Both are plain Node, zero deps, bounded reads, fail-open by design (a miss is a stray context string, not a gate): see `docs/audit-brief.md`.
|
|
16
|
+
- **`done-verifier` and `reader`, two new fast-tier agents with no file-editing tools, in both the claude-code and agy formats.** `done-verifier` probes a tracker item's stated done-signal (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything. `reader` reads and digests many files or notes and returns exactly what the brief asks (facts, quotes cited `path:line`, an index, a digest); unlike `bulk-worker`, it never classifies, tags, transforms or writes. `reader` is read-only by tool grant on both formats (no `Bash`); `done-verifier` on claude-code carries `Bash` for its probes, bound only by its prompt, not by the grant, and its description says so; on agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `ROUTING.md`'s decision tree, `TIERS.md`'s effort table, and both agent-folder READMEs now name them.
|
|
17
|
+
- **The inline threshold gets a measurement instead of a guessed figure.** "A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Measure yours once: spawn a subagent with a one-line task and read its token count. Work smaller than that stays inline." (claude-code `ROUTING.md` and `CLAUDE.snippet.md` only; no private number shipped.)
|
|
18
|
+
|
|
19
|
+
### Changed
|
|
20
|
+
|
|
21
|
+
- **The package now says what it is for in the first line people and agents read.** npm search, GitHub search and the installer banner showed "Routing instructions and a CLI runner", which named the parts and not the purpose. The description, the README opening and the banner now lead with the goal (each task to the right model, agent or LLM, fewer frontier tokens) while keeping the 0.1.11 correction intact: routing is an instruction your agent follows, and the README still states it does not automatically compare prices or select models. A test holds all three surfaces to that. New: an "At a glance" block and question-shaped "Common questions" in the README, request-level alternatives named for readers who want a proxy, `llms.txt` at the root and a headless-use section in `AGENTS.md` (both now ship in the package), and search keywords matching what comparable routers use.
|
|
22
|
+
- **Corrected the unqualified premise "a subagent holds none of these rules" everywhere it appeared** (`TASK_BUNDLE.md`, `ORCHESTRATOR.md`, the claude-code snippet, `docs/part-1-beginner.md`, `docs/part-2-intermediate.md`, README principle 6, and `ROUTING.md`'s "Who builds"). The corrected fact: a Claude Code subagent loads CLAUDE.md and so keeps the standing rules, just not this task's scope; a second CLI or a fresh chat window may still hold none of it. "Absence is denial" is unchanged; only the premise about who is absent what was wrong.
|
|
23
|
+
- **The claude-code snippet's closing "available as ..." agent list is generated from the files actually shipped in `templates/agents/claude-code/`, never hand-typed.** It had drifted once already: `finding-verifier` shipped in 0.1.14 and was missing from this sentence until now. `claudeAgentIds()` in `src/install.js` reads the folder; a test ties the rendered list to it.
|
|
24
|
+
- **The delegate-by-default gate now reaches every generated surface it should, not just three of them.** `build-protocol.md`'s "Roles, as capabilities" table and its "Why the builder does not hand off the main build" line, `builder.md`'s description, and `ROUTING.md`'s "Plan big, execute small" modifier still said, on a claude-code install, that the orchestrator writes the main build itself, never hands it off whole, and that a delegate inherits none of the session's rules: the exact premise the rest of this release corrects. All four now render through `subagentsLoadRules(primary)` the same way the decision tree and "Who builds" already did; every other primary is unchanged. A semantic-regression test asserts a claude-code install contains none of the old phrasing and a codex install still does.
|
|
25
|
+
|
|
26
|
+
### Fixed
|
|
27
|
+
|
|
28
|
+
- **Both new hooks could hang, and `route-gate.mjs` could read an unbounded or blocking file (pre-release audit finding, never shipped).** `readFileSync(0)` in both `route-gate.mjs` and `subagent-context.mjs` blocked until stdin reached EOF, so a caller that piped input in without closing its end (or ran the hook from a bare TTY) left the process running indefinitely; reproduced with `sleep 3 | CLAUDE_PROJECT_DIR=... node route-gate.mjs` still running past 1.5s. Separately, `route-gate.mjs` read the whole rules file into memory before bounding it (`readFileSync(path).slice(0, MAX_READ)`), so a FIFO planted at the rules path blocked forever on open, and a very large file was read in full before being truncated. Fixed in both hooks: stdin is now drained asynchronously against a 250ms hard cap, never blocking past it. `route-gate.mjs` additionally `statSync`s the resolved path and refuses anything that is not `isFile()` (a FIFO, socket, device or directory, symlink target included) before ever calling open, then reads through a single fixed 64 KB buffer via `openSync`/`readSync`, closed in a `finally`, so neither the read time nor the memory used depends on the file's on-disk size. Tests: an open, never-closed stdin pipe now exits within 1s for both hooks; a FIFO at the rules path returns the fallback instead of hanging; a 200 MB sparse rules file completes in well under a second with output still capped.
|
|
29
|
+
|
|
30
|
+
Three refinements to the routing model, from a review by [@shawnwows](https://x.com/shawnwows). The theme is the same in all three: a routing decision that was implied, inherited or asserted is now stated, pinned or checked.
|
|
31
|
+
|
|
32
|
+
### Added
|
|
33
|
+
|
|
34
|
+
- **`--model` and `--effort` on every lane, and a route recorded per run.** A lane with no flag and no `defaults` entry in `bin/lanes.json` runs on its own config file, which `cli-run` cannot see: a CLI configured months ago at a low reasoning effort keeps auditing at that effort while the routing docs describe an adversarial pass, and nothing raises an error. Each vendor spells the flags differently and `cli-run` translates (`grok -m/--reasoning-effort`, `codex -m/-c model_reasoning_effort="X"`, `agy --model/--effort`, `hermes -m/--reasoning`, `qwen -m` and no reasoning flag), each one read from that CLI's own `--help`. Flags beat `defaults`, `defaults` beats nothing, `--doctor` prints what each lane is pinned to, and the log carries `model_requested`, `effort_requested`, `model_source` and `effort_source` on every record, including runs refused before the lane started. It records no "actual": one lane of five (grok) reports a model id in its own output and the other four report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string, which the durable log never holds. `--effort` on qwen is a usage error rather than a silent drop, and route values are charset-bounded because a model id becomes an argv element and, on codex, part of a TOML value.
|
|
35
|
+
- **`finding-verifier`, a sixth subagent, in both agent formats.** Review and scanner findings no longer go straight to a repair. It reads the cited line, states what would trigger the problem, hunts for the guard, caller or test that makes it impossible, and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a change; INCONCLUSIVE is never rounded up to be safe or down to be tidy. Bound into the build protocol as Stage 5a, into `ROUTING.md`, and into the Claude Code activation snippet. The reproduction rule already existed in Stage 5; it had no owner, no separate model family and no way to say "I could not settle this".
|
|
36
|
+
- **Complexity and risk as inputs, alongside role** (`TIERS.md`). Complexity moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output. Risk (security, privacy, data loss, irreversible) moves the tier and who reads the result, because none of those failures is fixable by editing the code afterwards. A one-line change to an auth check is simple and high-risk at once, and the risk decides. Deliberately two rules and two small tables rather than a role by complexity by risk matrix: an 80-cell table is not maintained, and an unmaintained routing table is worse than none because it is believed.
|
|
37
|
+
|
|
38
|
+
### Changed
|
|
39
|
+
|
|
40
|
+
- `--model` is no longer qwen-only. `--safe-mode` still is.
|
|
41
|
+
- The route is resolved before the "lane disabled" and "binary missing" refusals, so those records carry it too. Found by the pre-release audit: a run refused for a missing binary is still a run that requested a route, and a failure record without one is the gap this release exists to close.
|
|
42
|
+
- `bin/lanes.json` gains an optional `defaults` block. It fails closed with the rest of the file: an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane with no reasoning flag refuses every lane until it is fixed, rather than being skipped quietly.
|
|
43
|
+
- The generated activation list gains a step about pinning the route, and `--doctor` output gains a route column with a plain sentence about what "not pinned" means.
|
|
44
|
+
|
|
7
45
|
## [0.1.13] - 2026-09-08
|
|
8
46
|
|
|
9
47
|
Three issues from a fresh first-run walkthrough of 0.1.12 (#26, #27, #28). Same class as 0.1.12's five: a surface describing an install that did not happen. A fourth, #25, was filed and closed as a mistake on the reporter's side, not a defect: the warning it said was missing has been printed since 0.1.12 and the repro had been read through a truncated pipe.
|
|
@@ -200,7 +238,9 @@ First release.
|
|
|
200
238
|
- Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
|
|
201
239
|
- Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
|
|
202
240
|
|
|
203
|
-
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.
|
|
241
|
+
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.15...HEAD
|
|
242
|
+
[0.1.15]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...v0.1.15
|
|
243
|
+
[0.1.14]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...v0.1.14
|
|
204
244
|
[0.1.13]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.12...v0.1.13
|
|
205
245
|
[0.1.12]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.11...v0.1.12
|
|
206
246
|
[0.1.11]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.10...v0.1.11
|
package/README.md
CHANGED
|
@@ -2,7 +2,16 @@
|
|
|
2
2
|
|
|
3
3
|
[](https://www.npmjs.com/package/model-orchestrator) [](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml) [](LICENSE) [](package.json)
|
|
4
4
|
|
|
5
|
-
**
|
|
5
|
+
**Route every task to the right model, agent or LLM, and spend fewer tokens.** A model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. One installer asks what you have access to and writes routing rules, subagents and a CLI runner for exactly that setup, from one chat app to several agent CLIs or a virtual machine. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models; the runner executes the lane it is given. On Claude Code it also delegates execution to subagents by default, with two hooks that inject the routing table every turn.
|
|
6
|
+
|
|
7
|
+
## At a glance
|
|
8
|
+
|
|
9
|
+
- **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
|
|
10
|
+
- **What it is not:** a proxy, a gateway or an API router. It does not automatically compare prices or select models; your agent follows the rules and chooses.
|
|
11
|
+
- **Install:** `npx model-orchestrator` (interactive), or headless from a script or an agent: `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator`.
|
|
12
|
+
- **Use it when:** you run more than one model or agent and want each task sent to the smallest one that can do it well.
|
|
13
|
+
- **What it saves:** frontier-model tokens. Bulk work, reading and checks go to fast tiers; the expensive tier is kept for planning and judgment.
|
|
14
|
+
- **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
|
|
6
15
|
|
|
7
16
|
Built from a working system, not a diagram: the routing rules, the protocols and the lane runner here run in production, generalized so they transfer to any stack.
|
|
8
17
|
|
|
@@ -18,14 +27,14 @@ The installer asks a few things, then writes a folder:
|
|
|
18
27
|
2. **Which AIs do you have access to?** (it marks the ones already on your PATH)
|
|
19
28
|
3. **Which one is your primary agent?** (the one that runs the system)
|
|
20
29
|
|
|
21
|
-
It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: follow the generated README. Inspect the manifest and remove only the individual managed subagent files you no longer need, preserve edited or pre-existing files, and remove your manually pasted activation block. Never delete a shared subagent folder.
|
|
30
|
+
It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions (and, on Claude Code, two hook scripts) go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: follow the generated README. Inspect the manifest and remove only the individual managed subagent files you no longer need, preserve edited or pre-existing files, and remove your manually pasted activation block. Never delete a shared subagent folder.
|
|
22
31
|
|
|
23
32
|
## The three levels
|
|
24
33
|
|
|
25
34
|
| Level | You have | You get |
|
|
26
35
|
|---|---|---|
|
|
27
36
|
| **1 · Beginner** | one LLM or one agent | tiers, task classification, the two build checkpoints, the protocols (build, propagate, gap analysis, deep research, numbers and logic, memory and record), a task-bundle template, and your agent set up to follow them |
|
|
28
|
-
| **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts), a delegation matrix generated from your selection, research triage across the lanes you have |
|
|
37
|
+
| **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts; `--model` / `--effort` to pin the route and log it), a delegation matrix generated from your selection, research triage across the lanes you have |
|
|
29
38
|
| **3 · Advanced** | a virtual machine | everything above, plus a gateway config rendered from the API keys you hold (asked separately from your CLIs), pinned images, box rules, privacy gates, and a weekly gap-analysis job with "what watches it" written down |
|
|
30
39
|
|
|
31
40
|
Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md).
|
|
@@ -63,7 +72,7 @@ An install has two targets, and a scripted run should set both.
|
|
|
63
72
|
| Flag | Default | What lands there |
|
|
64
73
|
|---|---|---|
|
|
65
74
|
| `--dir` | `./ai-orchestrator` | the docs, protocols and (level 2+) `bin/cli-run.mjs`. Named after what it contains, not after this package, so a project can hold one without looking like a checkout of it. Pass `--dir ./model-orchestrator` if you prefer the package name. |
|
|
66
|
-
| `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. |
|
|
75
|
+
| `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets two hook scripts in `.claude/hooks/`, wired by a settings snippet you merge yourself. |
|
|
67
76
|
|
|
68
77
|
`--project` defaulting to the current directory is the one that surprises people: run the command from your home folder with Claude Code as the primary and five agent files land in your home folder. The installer prints the resolved project path in the plan and says when you left it at the default. Set it.
|
|
69
78
|
|
|
@@ -90,8 +99,10 @@ ai-orchestrator/
|
|
|
90
99
|
TASK_BUNDLE.md the brief every delegation carries
|
|
91
100
|
protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record
|
|
92
101
|
CODECALC.md OBSIDIAN-TC.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
|
|
93
|
-
<project>/.claude/agents/
|
|
102
|
+
<project>/.claude/agents/ one per tier plus finding-verifier, done-verifier, reader, at the PROJECT root (if Claude Code is primary)
|
|
103
|
+
<project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart), Claude Code only
|
|
94
104
|
CLAUDE.snippet.md the block to paste into your CLAUDE.md
|
|
105
|
+
settings.hooks.snippet.json the hooks block to merge into .claude/settings.json (Claude Code only)
|
|
95
106
|
ROUTING.md multi-lane decision tree (level 2+)
|
|
96
107
|
TIERS.md DELEGATION_MATRIX.md RESEARCH_TRIAGE.md CLI-RUN.md
|
|
97
108
|
bin/cli-run.mjs bin/lanes.json (node bin/cli-run.mjs --doctor is the smoke test)
|
|
@@ -159,9 +170,79 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
|
|
|
159
170
|
3. **Exit 0 is not a deliverable.** Check for the artifact, not the status line. `cli-run` checks the response is structurally there; `--expect-file` checks the artifact.
|
|
160
171
|
4. **Numbers are computed, never guessed.** A tool that calculates beats a model that feels finished.
|
|
161
172
|
5. **A write nobody can find again did not happen.** Search first, keep the index true, one writer.
|
|
162
|
-
6. **
|
|
173
|
+
6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, only the brief carries what this task needs. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
|
|
163
174
|
7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
|
|
164
175
|
|
|
176
|
+
## Routing by role, complexity and risk
|
|
177
|
+
|
|
178
|
+
Role picks the agent. Two more inputs move the choice, and they move it in
|
|
179
|
+
different directions, so `TIERS.md` states them separately rather than folding
|
|
180
|
+
them into the role:
|
|
181
|
+
|
|
182
|
+
- **Complexity moves the effort.** A worker executing a finished plan needs less
|
|
183
|
+
reasoning than the reviewer judging its output. When the plan is airtight the
|
|
184
|
+
spec is carrying the thinking.
|
|
185
|
+
- **Risk moves the tier and the reader.** Security, privacy, data loss and
|
|
186
|
+
irreversible changes buy the attack lane, a named check, a rollback path or a
|
|
187
|
+
human yes. A one-line change to an auth check is simple and high-risk at the
|
|
188
|
+
same time, and it is the risk that decides.
|
|
189
|
+
|
|
190
|
+
The top of the ladder is bought with evidence: a reproduced failure, an
|
|
191
|
+
unresolved checkpoint, an irreversible change. A task that merely feels hard is
|
|
192
|
+
a deep-tier task, not an escalation.
|
|
193
|
+
|
|
194
|
+
## A finding is a claim, not a fact
|
|
195
|
+
|
|
196
|
+
Review findings do not go straight to a repair. `finding-verifier` reads the
|
|
197
|
+
cited line, states what would trigger the problem, then hunts for the guard,
|
|
198
|
+
caller or test that makes it impossible, and returns **CONFIRMED**,
|
|
199
|
+
**NOT_REPRODUCED** or **INCONCLUSIVE** per finding. Only CONFIRMED earns a
|
|
200
|
+
change. Use a different model family from the one that produced the finding
|
|
201
|
+
where you have one: a family asked to check its own claim tends to agree with
|
|
202
|
+
itself.
|
|
203
|
+
|
|
204
|
+
## Two more fast-tier checks
|
|
205
|
+
|
|
206
|
+
`done-verifier` probes the artifact a tracker item's done-signal names (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything itself. It carries no file-editing tools, but on claude-code it does carry `Bash` for those probes (`git log`, `grep`, `wc -l`, `test -f`); staying to read-only commands there is a rule in its prompt, not a restriction on the tool grant, and its own description says so. On agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `reader` is the one that is read-only by tool grant on both: no `Write`, `Edit`, or `Bash`. It reads and digests many files or notes and hands back exactly what the brief asked for, cited by `path:line`; it never classifies, tags or writes, which is what separates it from `bulk-worker`. Both ship in the claude-code and agy agent sets, at the fast tier.
|
|
207
|
+
|
|
208
|
+
## Pin the route, or know that you did not
|
|
209
|
+
|
|
210
|
+
A lane with no `--model`, no `--effort` and no `defaults` entry in
|
|
211
|
+
`bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
|
|
212
|
+
CLI configured months ago at a low reasoning effort keeps auditing at that
|
|
213
|
+
effort while your routing docs describe an adversarial pass.
|
|
214
|
+
|
|
215
|
+
```bash
|
|
216
|
+
node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
|
|
217
|
+
node bin/cli-run.mjs --doctor # prints what each lane is pinned to, and what is not pinned
|
|
218
|
+
```
|
|
219
|
+
|
|
220
|
+
Every run logs the model and effort **requested** and where the request came
|
|
221
|
+
from: `flag`, `lanes.json`, or `lane_default`, on every record including the
|
|
222
|
+
runs that never reached a lane. It does not log an actual. One lane of five
|
|
223
|
+
(grok) reports a model id in its own output and the other four report none, so
|
|
224
|
+
an actual field would be present for one lane and missing for four, and it
|
|
225
|
+
would be a provider-supplied string, which the durable log deliberately never
|
|
226
|
+
holds.
|
|
227
|
+
|
|
228
|
+
## Common questions
|
|
229
|
+
|
|
230
|
+
### How do I cut token usage across Claude Code, Codex and Gemini?
|
|
231
|
+
|
|
232
|
+
Install for the tools you have, then let the generated `ROUTING.md` decide the tier per task: bulk, reading and verification go to the fast tier or a cheaper CLI lane, and the deep tier only plans and judges. On Claude Code, execution goes to the `builder` subagent by default and the main session plans and verifies. Every lane call through `cli-run` logs the model and effort it ran with, so you can check where the tokens went.
|
|
233
|
+
|
|
234
|
+
### How do I route tasks to cheaper models?
|
|
235
|
+
|
|
236
|
+
The rules route by role, complexity and risk (see [Routing by role, complexity and risk](#routing-by-role-complexity-and-risk)). Role picks the agent, complexity moves the effort, risk moves the tier. A task a cheap tier finishes correctly never gets a frontier token.
|
|
237
|
+
|
|
238
|
+
### Is this an LLM router or an AI gateway?
|
|
239
|
+
|
|
240
|
+
No. It routes at the task level, through instructions your agent follows and a runner for agent CLIs. If you want a service or proxy that picks or forwards the model on every API request, look at request-level routers and gateways such as RouteLLM, LiteLLM, OpenRouter or claude-code-router. They solve a different problem and can sit underneath this.
|
|
241
|
+
|
|
242
|
+
### Can an agent install and run it without a person?
|
|
243
|
+
|
|
244
|
+
Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run` previews the plan, and `--list` prints every supported AI. Nothing is appended to a file you already have; activation snippets are written next to your files for you to merge.
|
|
245
|
+
|
|
165
246
|
## Requirements
|
|
166
247
|
|
|
167
248
|
Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows is untested: `cli-run` ends a lane's process tree there with `taskkill`, but nothing in CI runs on Windows, so treat it as unsupported until someone reports otherwise.
|
|
@@ -177,6 +258,14 @@ Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box tem
|
|
|
177
258
|
|
|
178
259
|
Add an AI to `src/catalog.js` and every prompt, table, config and doc picks it up. Run `npm test`. Keep templates free of logic and free of anything that looks like a credential. The rest is in [CONTRIBUTING.md](CONTRIBUTING.md); releases in [RELEASING.md](RELEASING.md); security reports in [SECURITY.md](SECURITY.md).
|
|
179
260
|
|
|
261
|
+
## Credits
|
|
262
|
+
|
|
263
|
+
- [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for
|
|
264
|
+
separating role, complexity and risk instead of compressing them into one
|
|
265
|
+
scale, for recording the model and effort a lane was actually asked for, and
|
|
266
|
+
for verifying findings before they trigger repairs. All three shipped in
|
|
267
|
+
0.1.14.
|
|
268
|
+
|
|
180
269
|
## License
|
|
181
270
|
|
|
182
271
|
[MIT](LICENSE)
|
package/bin/cli-run.mjs
CHANGED
|
@@ -39,6 +39,17 @@
|
|
|
39
39
|
//
|
|
40
40
|
// The durable log stores a FIXED reason code per run (see REASONS), never a
|
|
41
41
|
// provider-supplied string. Bounded vendor stderr goes to your terminal only.
|
|
42
|
+
//
|
|
43
|
+
// ROUTE: which model and reasoning effort a lane ran with.
|
|
44
|
+
// A lane with no --model and no lanes.json default inherits whatever its own
|
|
45
|
+
// config file says, which is invisible from here and is how a documented route
|
|
46
|
+
// silently stops being the route that runs. --model / --effort pin it per call,
|
|
47
|
+
// `defaults` in lanes.json pins it per lane, and every run logs the value that
|
|
48
|
+
// was REQUESTED plus where the request came from (flag, lanes.json, or nothing
|
|
49
|
+
// at all). It does not log an "actual". One lane of five (grok) does report a
|
|
50
|
+
// model id in its own output; the other four report none, and a field present
|
|
51
|
+
// for one lane and absent for four is worse than no field. It would also be a
|
|
52
|
+
// provider-supplied string, which this log deliberately never holds.
|
|
42
53
|
|
|
43
54
|
import { spawn } from 'node:child_process';
|
|
44
55
|
import { StringDecoder } from 'node:string_decoder';
|
|
@@ -190,28 +201,71 @@ export function judgeQwen(rc, out) {
|
|
|
190
201
|
return pass(text, `subtype=success, totalErrors=0 across ${Object.keys(models).length} model(s)`);
|
|
191
202
|
}
|
|
192
203
|
|
|
204
|
+
// --- route: model and effort per lane --------------------------------------
|
|
205
|
+
// Each vendor spells these differently, and the spelling was read from each
|
|
206
|
+
// CLI's own --help, not remembered. A lane with `effort: null` has no reasoning
|
|
207
|
+
// flag at all; asking for one there is a usage error, never a silent drop.
|
|
208
|
+
// grok -m MODEL --reasoning-effort EFFORT
|
|
209
|
+
// codex -m MODEL -c model_reasoning_effort="EFFORT" (a TOML override, hence the quotes)
|
|
210
|
+
// agy --model M --effort EFFORT (low|medium|high)
|
|
211
|
+
// hermes -m MODEL --reasoning LEVEL (none|minimal|...)
|
|
212
|
+
// qwen -m MODEL no reasoning flag
|
|
213
|
+
export const LANE_FLAGS = {
|
|
214
|
+
grok: { model: (v) => ['-m', v], effort: (v) => ['--reasoning-effort', v] },
|
|
215
|
+
codex: { model: (v) => ['-m', v], effort: (v) => ['-c', `model_reasoning_effort="${v}"`] },
|
|
216
|
+
agy: { model: (v) => ['--model', v], effort: (v) => ['--effort', v] },
|
|
217
|
+
hermes: { model: (v) => ['-m', v], effort: (v) => ['--reasoning', v] },
|
|
218
|
+
qwen: { model: (v) => ['-m', v], effort: null }
|
|
219
|
+
};
|
|
220
|
+
|
|
221
|
+
// A model id or effort level becomes an argv element and, for codex, part of a
|
|
222
|
+
// TOML value. Bounding the charset is what makes both safe: no leading dash (a
|
|
223
|
+
// value cannot become a flag), no quote, space or control character (a value
|
|
224
|
+
// cannot break out of the TOML string), and a length cap so a config file
|
|
225
|
+
// cannot push an unbounded string into the durable log.
|
|
226
|
+
export const ROUTE_VALUE = /^[A-Za-z0-9][A-Za-z0-9._:@/+-]{0,63}$/;
|
|
227
|
+
export function badRouteValue(kind, v) {
|
|
228
|
+
if (typeof v !== 'string' || !ROUTE_VALUE.test(v)) {
|
|
229
|
+
return `--${kind} must be 1 to 64 characters of letters, digits, dot, underscore, colon, at, slash, plus or dash, and may not start with a dash: ${JSON.stringify(v)}`;
|
|
230
|
+
}
|
|
231
|
+
return null;
|
|
232
|
+
}
|
|
233
|
+
|
|
193
234
|
// --- adapters: build argv for a lane -------------------------------------
|
|
235
|
+
// Route flags go in front of the prompt for every lane, because two lanes
|
|
236
|
+
// (hermes, codex) take the prompt as a positional argument and a flag after it
|
|
237
|
+
// is either ignored or read as part of it.
|
|
238
|
+
function routeFlags(lane, opts) {
|
|
239
|
+
const spec = LANE_FLAGS[lane];
|
|
240
|
+
const out = [];
|
|
241
|
+
if (!spec) return out;
|
|
242
|
+
if (opts.model) out.push(...spec.model(opts.model));
|
|
243
|
+
if (opts.effort && spec.effort) out.push(...spec.effort(opts.effort));
|
|
244
|
+
return out;
|
|
245
|
+
}
|
|
246
|
+
|
|
194
247
|
export function buildArgv(lane, binary, prompt, opts, tmp) {
|
|
195
248
|
const timeout = opts.timeout;
|
|
249
|
+
const route = routeFlags(lane, opts);
|
|
196
250
|
switch (lane) {
|
|
197
251
|
case 'grok':
|
|
198
|
-
return { argv: [binary, '--output-format', 'json', '-p', prompt] };
|
|
252
|
+
return { argv: [binary, '--output-format', 'json', ...route, '-p', prompt] };
|
|
199
253
|
case 'codex': {
|
|
200
254
|
const last = join(tmp, 'last.txt');
|
|
201
255
|
const argv = [binary, 'exec', '--json', '--color', 'never', '--skip-git-repo-check', '-o', last];
|
|
202
256
|
if (opts.audit) argv.push('--sandbox', 'read-only'); // an audit lane that can write is a bug
|
|
257
|
+
argv.push(...route);
|
|
203
258
|
argv.push(prompt);
|
|
204
259
|
return { argv, outFile: last };
|
|
205
260
|
}
|
|
206
261
|
case 'agy': {
|
|
207
262
|
const mins = Math.max(1, Math.round(timeout / 60));
|
|
208
|
-
return { argv: [binary, '--print-timeout', `${mins}m`, '--output-format', 'stream-json', '-p', prompt] };
|
|
263
|
+
return { argv: [binary, '--print-timeout', `${mins}m`, '--output-format', 'stream-json', ...route, '-p', prompt] };
|
|
209
264
|
}
|
|
210
265
|
case 'hermes':
|
|
211
|
-
return { argv: [binary, '-z', prompt, '--usage-file', join(tmp, 'usage.json')] };
|
|
266
|
+
return { argv: [binary, '-z', ...route, prompt, '--usage-file', join(tmp, 'usage.json')] };
|
|
212
267
|
case 'qwen': {
|
|
213
|
-
const argv = [binary, '-o', 'json'];
|
|
214
|
-
if (opts.model) argv.push('-m', opts.model);
|
|
268
|
+
const argv = [binary, '-o', 'json', ...route];
|
|
215
269
|
if (opts.safeMode) argv.push('--safe-mode');
|
|
216
270
|
argv.push('-p', prompt);
|
|
217
271
|
return { argv };
|
|
@@ -370,26 +424,67 @@ function log(rec) {
|
|
|
370
424
|
// documented default). PRESENT BUT UNREADABLE OR MALFORMED = no lane enabled:
|
|
371
425
|
// a half-written config must fail closed, never re-enable what the installer
|
|
372
426
|
// disabled. Returns null when the file is bad so the caller can say so.
|
|
373
|
-
export function
|
|
427
|
+
export function laneConfig(here = dirname(fileURLToPath(import.meta.url))) {
|
|
374
428
|
const p = join(here, 'lanes.json');
|
|
375
|
-
if (!existsSync(p)) return LANES;
|
|
429
|
+
if (!existsSync(p)) return { enabled: LANES, defaults: {} };
|
|
376
430
|
try {
|
|
377
431
|
const j = JSON.parse(readFileSync(p, 'utf8'));
|
|
378
432
|
if (!j || typeof j !== 'object' || !Array.isArray(j.enabled)) return null;
|
|
379
433
|
if (!j.enabled.every((l) => typeof l === 'string' && LANES.includes(l))) return null;
|
|
380
|
-
|
|
434
|
+
// `defaults` pins a model and effort per lane. It is optional; present and
|
|
435
|
+
// malformed fails closed with the rest of the file, because a half-written
|
|
436
|
+
// route is exactly the silent-inheritance problem this field exists to fix.
|
|
437
|
+
const defaults = {};
|
|
438
|
+
if (j.defaults !== undefined) {
|
|
439
|
+
if (!j.defaults || typeof j.defaults !== 'object' || Array.isArray(j.defaults)) return null;
|
|
440
|
+
for (const [lane, d] of Object.entries(j.defaults)) {
|
|
441
|
+
if (!LANES.includes(lane)) return null;
|
|
442
|
+
if (!d || typeof d !== 'object' || Array.isArray(d)) return null;
|
|
443
|
+
const { model, effort, ...rest } = d;
|
|
444
|
+
if (Object.keys(rest).length) return null;
|
|
445
|
+
if (model !== undefined && badRouteValue('model', model)) return null;
|
|
446
|
+
if (effort !== undefined) {
|
|
447
|
+
if (badRouteValue('effort', effort)) return null;
|
|
448
|
+
if (!LANE_FLAGS[lane] || !LANE_FLAGS[lane].effort) return null; // a lane with no reasoning flag cannot have one pinned
|
|
449
|
+
}
|
|
450
|
+
defaults[lane] = { model: model ?? null, effort: effort ?? null };
|
|
451
|
+
}
|
|
452
|
+
}
|
|
453
|
+
return { enabled: j.enabled, defaults };
|
|
381
454
|
} catch {
|
|
382
455
|
return null;
|
|
383
456
|
}
|
|
384
457
|
}
|
|
385
458
|
|
|
459
|
+
// Kept as the narrow question most callers ask. null still means malformed.
|
|
460
|
+
export function enabledLanes(here = dirname(fileURLToPath(import.meta.url))) {
|
|
461
|
+
const c = laneConfig(here);
|
|
462
|
+
return c === null ? null : c.enabled;
|
|
463
|
+
}
|
|
464
|
+
|
|
465
|
+
// Flag beats lanes.json beats nothing. `source` is what makes the log audit-worthy:
|
|
466
|
+
// 'lane_default' means this run inherited the vendor CLI's own config, unseen from here.
|
|
467
|
+
export function resolveRoute(lane, opts, defaults) {
|
|
468
|
+
const d = (defaults && defaults[lane]) || {};
|
|
469
|
+
const model = opts.model ?? d.model ?? null;
|
|
470
|
+
const effort = opts.effort ?? d.effort ?? null;
|
|
471
|
+
const src = (flag, def) => (flag != null ? 'flag' : def != null ? 'lanes.json' : 'lane_default');
|
|
472
|
+
return { model, effort, model_source: src(opts.model, d.model), effort_source: src(opts.effort, d.effort) };
|
|
473
|
+
}
|
|
474
|
+
|
|
386
475
|
function usage(msg) {
|
|
387
476
|
if (msg) console.error('cli-run: ' + msg);
|
|
388
477
|
console.error(`usage: cli-run <${LANES.join('|')}> "<prompt>" [--brief FILE] [--timeout SECS] [--quiet]
|
|
389
|
-
[--expect-file PATH] [--expect-json]
|
|
478
|
+
[--model ID] [--effort LEVEL] [--expect-file PATH] [--expect-json]
|
|
390
479
|
cli-run codex --audit "<prompt>" read-only sandbox (audit shape)
|
|
391
|
-
cli-run qwen [--
|
|
392
|
-
cli-run --doctor [--run] enabled lanes, binaries
|
|
480
|
+
cli-run qwen [--safe-mode] "<prompt>" qwen-only flag
|
|
481
|
+
cli-run --doctor [--run] enabled lanes, binaries, and the route each one is pinned to
|
|
482
|
+
|
|
483
|
+
--model / --effort pin what a lane runs with, instead of letting it inherit its
|
|
484
|
+
own config. Every lane takes --model; every lane except qwen takes --effort.
|
|
485
|
+
Levels are the vendor's own (agy low|medium|high, hermes none|minimal|...): an
|
|
486
|
+
unknown level is rejected by the lane, and reported as that lane's exit code.
|
|
487
|
+
Pin them per lane instead of per call with "defaults" in bin/lanes.json.`);
|
|
393
488
|
return USAGE;
|
|
394
489
|
}
|
|
395
490
|
|
|
@@ -416,11 +511,12 @@ function installedPrimary(here = dirname(fileURLToPath(import.meta.url))) {
|
|
|
416
511
|
|
|
417
512
|
// --doctor: the first thing to run after install.
|
|
418
513
|
export async function doctor(run) {
|
|
419
|
-
const
|
|
420
|
-
if (
|
|
514
|
+
const cfg = laneConfig();
|
|
515
|
+
if (cfg === null) {
|
|
421
516
|
console.error('doctor: lanes.json exists but is malformed; fix it first');
|
|
422
517
|
return USAGE;
|
|
423
518
|
}
|
|
519
|
+
const { enabled, defaults } = cfg;
|
|
424
520
|
let bad = 0;
|
|
425
521
|
console.log(`doctor: ${enabled.length} enabled lane(s): ${enabled.join(', ') || 'none'}`);
|
|
426
522
|
const primary = installedPrimary();
|
|
@@ -432,7 +528,11 @@ export async function doctor(run) {
|
|
|
432
528
|
for (const lane of LANES) {
|
|
433
529
|
const on = enabled.includes(lane);
|
|
434
530
|
const bin = which(lane);
|
|
435
|
-
|
|
531
|
+
const d = defaults[lane] || {};
|
|
532
|
+
// A disabled lane has no route worth reporting; saying "not pinned" there
|
|
533
|
+
// reads as a finding about a lane that is not going to run.
|
|
534
|
+
const route = !on ? '' : d.model || d.effort ? `route ${d.model || 'lane default'}/${d.effort || 'lane default'}` : 'route not pinned (inherits the lane\'s own config)';
|
|
535
|
+
let line = ` ${lane.padEnd(7)} ${on ? 'enabled ' : 'disabled'} ${bin ? 'binary ok' : 'binary MISSING'}${route ? ' ' + route : ''}`;
|
|
436
536
|
if (on && !bin) bad++;
|
|
437
537
|
if (on && bin && run) {
|
|
438
538
|
const rc = await main([lane, 'Reply with exactly the word OK and nothing else.', '--timeout', '120', '--quiet']);
|
|
@@ -443,6 +543,7 @@ export async function doctor(run) {
|
|
|
443
543
|
}
|
|
444
544
|
console.log(bad ? `doctor: ${bad} problem(s)` : 'doctor: all enabled lanes ' + (run ? 'answered' : 'present'));
|
|
445
545
|
console.log('doctor checks presence and, with --run, a one-word canary. It does not check vendor versions.');
|
|
546
|
+
console.log('"route not pinned" means that lane runs on whatever its own config file says, which this tool cannot see. Pin it in lanes.json "defaults" if the route matters.');
|
|
446
547
|
return bad ? NO_DELIVERABLE : OK;
|
|
447
548
|
}
|
|
448
549
|
|
|
@@ -485,10 +586,10 @@ export function checkContracts(opts, text, before) {
|
|
|
485
586
|
}
|
|
486
587
|
|
|
487
588
|
export async function main(argv) {
|
|
488
|
-
const VALUE = new Set(['--brief', '--timeout', '--model', '--expect-file']);
|
|
589
|
+
const VALUE = new Set(['--brief', '--timeout', '--model', '--effort', '--expect-file']);
|
|
489
590
|
const BOOL = new Set(['--quiet', '--audit', '--safe-mode', '--doctor', '--run', '--expect-json']);
|
|
490
591
|
const args = [...argv];
|
|
491
|
-
const opts = { timeout: 900, quiet: false, audit: false, model: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
|
|
592
|
+
const opts = { timeout: 900, quiet: false, audit: false, model: null, effort: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
|
|
492
593
|
const positional = [];
|
|
493
594
|
while (args.length) {
|
|
494
595
|
const a = args.shift();
|
|
@@ -498,6 +599,7 @@ export async function main(argv) {
|
|
|
498
599
|
if (a === '--brief') opts.brief = v;
|
|
499
600
|
else if (a === '--timeout') opts.timeout = Number(v);
|
|
500
601
|
else if (a === '--expect-file') opts.expectFile = v;
|
|
602
|
+
else if (a === '--effort') opts.effort = v;
|
|
501
603
|
else opts.model = v;
|
|
502
604
|
} else if (BOOL.has(a)) {
|
|
503
605
|
if (a === '--quiet') opts.quiet = true;
|
|
@@ -530,11 +632,33 @@ export async function main(argv) {
|
|
|
530
632
|
if (!prompt) return usage('give a prompt or --brief FILE');
|
|
531
633
|
if (!Number.isFinite(opts.timeout) || opts.timeout <= 0) return usage('--timeout must be a positive number of seconds');
|
|
532
634
|
if (opts.audit && lane !== 'codex') return usage('--audit is codex-only');
|
|
533
|
-
if (
|
|
635
|
+
if (opts.safeMode && lane !== 'qwen') return usage('--safe-mode is qwen-only');
|
|
636
|
+
for (const [kind, v] of [['model', opts.model], ['effort', opts.effort]]) {
|
|
637
|
+
if (v == null) continue;
|
|
638
|
+
const bad = badRouteValue(kind, v);
|
|
639
|
+
if (bad) return usage(bad);
|
|
640
|
+
}
|
|
641
|
+
// qwen has no reasoning flag. Dropping --effort silently would leave the caller
|
|
642
|
+
// believing a route that never happened, which is the defect this feature fixes.
|
|
643
|
+
if (opts.effort && !(LANE_FLAGS[lane] && LANE_FLAGS[lane].effort)) return usage(`${lane} has no reasoning-effort flag; --effort is not available on this lane`);
|
|
534
644
|
|
|
535
645
|
const digest = createHash('sha256').update(prompt).digest('hex').slice(0, 12);
|
|
536
646
|
const base = { lane, prompt_sha256_12: digest, prompt_chars: prompt.length };
|
|
537
|
-
const
|
|
647
|
+
const cfg = laneConfig();
|
|
648
|
+
// Resolve the route BEFORE the refusals below. A run that never reached a lane
|
|
649
|
+
// was still a request for one, and a failure record with no route is the exact
|
|
650
|
+
// gap this feature exists to close. A malformed lanes.json has no usable
|
|
651
|
+
// defaults, so the flags stand alone and say so.
|
|
652
|
+
const route = resolveRoute(lane, opts, cfg === null ? {} : cfg.defaults);
|
|
653
|
+
opts.model = route.model;
|
|
654
|
+
opts.effort = route.effort;
|
|
655
|
+
Object.assign(base, {
|
|
656
|
+
model_requested: route.model,
|
|
657
|
+
effort_requested: route.effort,
|
|
658
|
+
model_source: route.model_source,
|
|
659
|
+
effort_source: route.effort_source
|
|
660
|
+
});
|
|
661
|
+
const enabled = cfg === null ? null : cfg.enabled;
|
|
538
662
|
if (enabled === null) {
|
|
539
663
|
console.error('cli-run: lanes.json exists but is not a valid {"enabled": [...]} file; refusing every lane until it is fixed');
|
|
540
664
|
log({ ...base, verdict: 'unavailable', rc: UNAVAILABLE, reason: 'lanes_json_malformed' });
|
|
@@ -601,7 +725,8 @@ export async function main(argv) {
|
|
|
601
725
|
}
|
|
602
726
|
}
|
|
603
727
|
if (text && code === OK) process.stdout.write(text + '\n');
|
|
604
|
-
|
|
728
|
+
const routeNote = route.model || route.effort ? `${route.model || 'lane default'}/${route.effort || 'lane default'}` : 'lane default';
|
|
729
|
+
if (!opts.quiet) console.error(`cli-run[${lane}] ${verdict} rc=${code} ${r.seconds.toFixed(1)}s raw=${r.outBytes || 0}B route=${routeNote} :: ${detail}`);
|
|
605
730
|
// Durable log: fixed reason code and structural numbers only.
|
|
606
731
|
log({ ...base, verdict, rc: code, cli_rc: r.status, signal: r.signal || null, seconds: Math.round(r.seconds * 100) / 100, raw_bytes: r.outBytes || 0, deliverable_bytes: Buffer.byteLength(text), reason: REASONS.has(reason) ? reason : 'unknown' });
|
|
607
732
|
return code;
|
package/bin/cli.js
CHANGED
|
@@ -144,7 +144,7 @@ async function main() {
|
|
|
144
144
|
// Says what this generates, not what it guarantees. The old line promised
|
|
145
145
|
// routing this package does not perform: lane choice is an instruction an
|
|
146
146
|
// agent follows, never something enforced here (#11).
|
|
147
|
-
console.log('\nmodel-orchestrator\
|
|
147
|
+
console.log('\nmodel-orchestrator\nA model orchestrator: routing rules and a CLI runner for the AIs you actually have.\n');
|
|
148
148
|
|
|
149
149
|
// 1. Level
|
|
150
150
|
let level = Number(opt('level'));
|
package/docs/audit-brief.md
CHANGED
|
@@ -81,3 +81,24 @@ Re-audit the same scope. Every round-1 finding was reproduced before it was touc
|
|
|
81
81
|
Also new since round 1: the companion-tool path (`--tools codecalc`, `--no-tools`, `templates/tools/codecalc/`, `protocols/numbers-and-logic.md`, `resolveTools`). Attack it the same way: unknown tool ids, interaction with `--yes`, the extra interactive question, and whether any written snippet could be confused for a file the installer should not touch.
|
|
82
82
|
|
|
83
83
|
Report only what reproduces on the current tree. `CLEAN` per area is expected where the fix holds.
|
|
84
|
+
|
|
85
|
+
## New in 0.1.15: two claude-code-only hooks
|
|
86
|
+
|
|
87
|
+
`route-gate.mjs` (`UserPromptSubmit`) and `subagent-context.mjs` (`SubagentStart`) ship to `.claude/hooks/` only when claude-code is the primary. Both are plain Node, zero deps, and installed with mode `0o755`.
|
|
88
|
+
|
|
89
|
+
- **Reads.** `route-gate.mjs` reads at most 64 KB from one file: the routing rules file (`ROUTING.md` or `ORCHESTRATOR.md`) at a path rendered in at install time relative to `CLAUDE_PROJECT_DIR`, never a hardcoded absolute path. Before opening it, it `statSync`s the resolved path (following a symlink to its target) and refuses anything that is not `isFile()`, a FIFO, socket, device or directory included, so the read never touches a path that could block on open. The read itself is one `openSync` + one bounded `readSync` into a fixed 64 KB buffer, closed in a `finally`, so neither the time nor the memory this hook uses depends on how large the file on disk actually is. It then extracts the text between `<!-- route-gate:start -->` and `<!-- route-gate:end -->` and nothing else. `subagent-context.mjs` reads nothing from disk; its context is static text plus the same two rendered paths. Neither parses or executes anything it reads; the extracted block is passed through as a string.
|
|
90
|
+
- **Stdin.** Neither hook uses a field from the JSON input Claude Code sends on stdin, but both must still consume the pipe rather than ignore it. Both drain stdin asynchronously against a 250ms hard cap: whichever comes first, the real `end` event or the timeout, the hook proceeds. Neither ever calls a blocking, synchronous read of stdin.
|
|
91
|
+
- **Writes.** Neither writes a file. Both write one JSON object to stdout: `{"hookSpecificOutput":{"hookEventName":"...","additionalContext":"..."}}`, and both exit only after that write's callback fires, so a buffered write to a pipe is not truncated by an exit racing ahead of it.
|
|
92
|
+
- **Fail-open, on purpose.** A missing `CLAUDE_PROJECT_DIR`, a missing rules file, a non-regular file at the rules path, or a missing block each produce a one-line fallback `additionalContext` naming what was found, and the script still exits 0. This is acceptable because a miss here is a stray context string reaching the model, not a security gate: nothing downstream trusts the hook's output for anything but a routing suggestion, and the settings snippet that wires it in is a document the user merges by hand, never written automatically over an existing `settings.json`.
|
|
93
|
+
- **Bounded.** `route-gate.mjs` caps the read at 64 KB regardless of the file's reported size and the injected string at 4000 characters, so a rules file bloated by a bad edit, or truncated to an arbitrary length, cannot balloon the context or the read time on every turn. Both hooks cap stdin drain at 250ms.
|
|
94
|
+
- **Not yet attacked.** Untested here: a rules file with a `route-gate:start` marker but no matching end marker very far into the file (bounded by `MAX_READ`, so the end marker past that point is treated as absent, which is the intended fail-open path, but worth a deliberate case); a `CLAUDE_PROJECT_DIR` pointing at a path with no read permission; behavior under the Windows exec-form `node` + `args` invocation named in the settings snippet; a `statSync` that itself hangs (a stalled network filesystem, for instance) rather than the FIFO-at-open case this round fixed.
|
|
95
|
+
|
|
96
|
+
### Round 1 (pre-release), fixed before shipping
|
|
97
|
+
|
|
98
|
+
Three findings reproduced against the 0.1.15 branch before it shipped, none of them ever released:
|
|
99
|
+
|
|
100
|
+
| # | Finding | Fix |
|
|
101
|
+
|---|---|---|
|
|
102
|
+
| 1 | HIGH. `readFileSync(0)` in both hooks blocked until stdin reached EOF (`sleep 3 \| ... node route-gate.mjs` still running past 1.5s); `route-gate.mjs` also read the whole rules file into memory before bounding it, so a FIFO planted at the rules path blocked forever on open. | Stdin is drained asynchronously against a 250ms hard cap in both hooks. `route-gate.mjs` refuses anything that is not `isFile()` via `statSync` before ever calling open, then reads through one fixed 64 KB buffer via `openSync`/`readSync`. Tests: an open, never-closed stdin pipe exits within 1s for both hooks; a FIFO at the rules path returns the fallback instead of hanging; a 200 MB sparse rules file completes in well under a second with output still capped. |
|
|
103
|
+
| 2 | MEDIUM. `done-verifier`'s description, both agent-folder READMEs, and the root README called it "read-only" without qualification, while its claude-code file carries an unrestricted `Bash` grant; nothing in that grant stops it from running a mutating command. | Every one of those surfaces now says plainly that `done-verifier` carries no file-editing tools and that its Bash use is bound by its own prompt, not by the tool grant; `reader` is named as the one that is read-only by tool grant (no Bash) on both formats. |
|
|
104
|
+
| 3 | MEDIUM. Three generated surfaces still stated the pre-0.1.15 premise on a claude-code install: `builder.md`'s description ("... or the main build itself"), `build-protocol.md`'s roles table and its "why the builder does not hand off" note, and `ROUTING.md`'s "Plan big, execute small" line ("the orchestrator executes"). | All three now render through `subagentsLoadRules(primary)`, the same gate the decision tree and "Who builds" already used; every other primary is unchanged. A semantic-regression test asserts a claude-code install contains none of the old phrasing and a codex install still does. |
|
package/docs/part-1-beginner.md
CHANGED
|
@@ -14,6 +14,8 @@ If your agent exposes model choice (Claude Code, Codex, Antigravity), map the ti
|
|
|
14
14
|
|
|
15
15
|
Three cost levers, always together: tier (price per token), token discipline (how many tokens: read only what you will touch, never re-read, deliverables not narration), effort (how hard each call thinks).
|
|
16
16
|
|
|
17
|
+
And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **Risk** moves the tier and who reads the result: security, privacy, data loss and irreversible changes are the four worth naming, because none of their failures can be fixed by editing the code afterwards. A one-line change to an auth check is simple and high-risk at once, and it is the risk that decides.
|
|
18
|
+
|
|
17
19
|
Robustness first, cost second. You split tiers because the split produces better work.
|
|
18
20
|
|
|
19
21
|
## 2. Classify every task, first match wins
|
|
@@ -36,7 +38,7 @@ Every build gets two checkpoints. **Before writing:** you map what it touches an
|
|
|
36
38
|
|
|
37
39
|
## 4. Every hand-off carries a brief
|
|
38
40
|
|
|
39
|
-
A subagent, a fresh chat, a second window
|
|
41
|
+
A subagent, a fresh chat, a second window may hold none of your rules, and that is the default to assume. The one documented exception is a Claude Code subagent: it loads the project's CLAUDE.md hierarchy at start, so it keeps the standing rules, just not this task's scope. Either way, it reads an unspecified edge as an open one. The brief: purpose, task class, granted scope, capabilities, denied actions, conventions it does not have, report contract (what was not done, what is unverified), exit parameters (when to stop, and what to return if it hits the bound). Absence is denial.
|
|
40
42
|
|
|
41
43
|
## 5. The second pass
|
|
42
44
|
|