model-orchestrator 0.1.26 → 0.1.27
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +8 -1
- package/README.md +98 -205
- package/docs/README.md +8 -1
- package/docs/companions.md +17 -0
- package/docs/guarantees.md +18 -0
- package/docs/how-it-routes.md +60 -0
- package/docs/install.md +73 -0
- package/llms.txt +5 -1
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,12 @@ All notable changes to this project are documented here. The format follows [Kee
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [0.1.27] - 2026-09-21
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
- **README restructured for a first-time reader: 32,083 bytes to 17,586, with the first screen now a claim, the install command and a real `--dry` plan.** The old opening was a 200-word paragraph followed by "What it is not", the flag-conflict rules and the uninstall procedure, all before the reader had seen the tool do anything. The detail moved rather than went away: `docs/install.md` (every flag, the two target folders, headless examples, the full file list), `docs/how-it-routes.md` (role, complexity and stakes; the three verifier agents; pinning model and effort), `docs/guarantees.md` (enforced by code, delegated to a vendor flag, or only an instruction), `docs/companions.md` (the full companion-tool table). `docs/README.md` and `llms.txt` index all four. Platform support and the vendor compatibility table stay in the README inside collapsed `<details>` blocks, because `scripts/gen-catalog.js` writes the vendor table between markers there and `test/prose.test.js` requires the skipped-test explanations to live in the README; collapsing them keeps both mechanisms pointed at the same file. The opening line still satisfies `test/copy.test.js`: it names the package, carries the shared purpose clause, and keeps the 0.1.11 correction that it does not select models itself.
|
|
12
|
+
|
|
7
13
|
## [0.1.26] - 2026-09-19
|
|
8
14
|
|
|
9
15
|
### Added
|
|
@@ -375,7 +381,8 @@ First release.
|
|
|
375
381
|
- Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
|
|
376
382
|
- Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
|
|
377
383
|
|
|
378
|
-
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.
|
|
384
|
+
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.27...HEAD
|
|
385
|
+
[0.1.27]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.26...v0.1.27
|
|
379
386
|
[0.1.26]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.25...v0.1.26
|
|
380
387
|
[0.1.25]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.24...v0.1.25
|
|
381
388
|
[0.1.24]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.23...v0.1.24
|
package/README.md
CHANGED
|
@@ -2,48 +2,58 @@
|
|
|
2
2
|
|
|
3
3
|
[](https://www.npmjs.com/package/model-orchestrator) [](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml) [](LICENSE) [](package.json)
|
|
4
4
|
|
|
5
|
-
**
|
|
6
|
-
|
|
7
|
-
## At a glance
|
|
8
|
-
|
|
9
|
-
- **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
|
|
10
|
-
- **What it is not:** a proxy, a gateway or an API router. It does not automatically compare prices or select models; your agent follows the rules and chooses.
|
|
11
|
-
- **Install:** `npx model-orchestrator` (interactive), or headless from a script or an agent: `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator`.
|
|
12
|
-
- **Claude Code plugin:** `/plugin marketplace add aunysillyme/model-orchestrator`, then `/plugin install model-orchestrator@model-orchestrator`. The routing rules still come from the installer; see [Claude Code plugin](#claude-code-plugin).
|
|
13
|
-
- **Use it when:** you run more than one model or agent and want each task sent to the smallest one that can do it well.
|
|
14
|
-
- **What it saves:** frontier-model tokens. Bulk work, reading and checks go to fast tiers; the expensive tier is kept for planning and judgment.
|
|
15
|
-
- **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
|
|
16
|
-
|
|
17
|
-
Built from a working system, not a diagram: the routing rules, the protocols and the lane runner here run in production, generalized so they transfer to any stack.
|
|
5
|
+
**A model orchestrator that sends every task to the smallest model that can do it,** so small work goes to cheap tiers and fewer tokens go to frontier models. One command reads which AIs you actually have, then writes the routing rules, the subagents and the lane runner for exactly that set: Claude Code, Codex, Gemini, Grok, Qwen, Ollama.
|
|
18
6
|
|
|
19
7
|
```bash
|
|
20
8
|
npx model-orchestrator
|
|
21
9
|
```
|
|
22
10
|
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
11
|
+
Three questions, then 38 files. Here is a real `--dry` run, which prints the plan and writes nothing:
|
|
12
|
+
|
|
13
|
+
```text
|
|
14
|
+
Plan
|
|
15
|
+
level 2 Intermediate
|
|
16
|
+
access claude-code, codex, grok
|
|
17
|
+
primary claude-code
|
|
18
|
+
tools codecalc
|
|
19
|
+
folder ./ai-orchestrator
|
|
20
|
+
project . (11 subagent files go here)
|
|
21
|
+
files 38
|
|
22
|
+
- ROUTING.md multi-lane decision tree
|
|
23
|
+
- TIERS.md DELEGATION_MATRIX.md which lane, at what effort
|
|
24
|
+
- TASK_BUNDLE.md the brief every delegation carries
|
|
25
|
+
- protocols/ build, propagate, gap-analysis, deep-research, and three more
|
|
26
|
+
- [project] .claude/agents/ builder, deep-planner, code-reviewer, bulk-worker,
|
|
27
|
+
live-researcher, reader, finding-verifier, done-verifier
|
|
28
|
+
- [project] .claude/hooks/ route-gate, subagent-context, route-metrics
|
|
29
|
+
- bin/cli-run.mjs bin/lanes.json the lane runner
|
|
30
|
+
|
|
31
|
+
--dry: nothing written.
|
|
32
|
+
```
|
|
26
33
|
|
|
27
|
-
|
|
28
|
-
2. **Which AIs do you have access to?** (it marks the ones already on your PATH)
|
|
29
|
-
3. **Which one is your primary agent?** (the one that runs the system)
|
|
34
|
+
Reproduce it:
|
|
30
35
|
|
|
31
|
-
|
|
36
|
+
```bash
|
|
37
|
+
npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --dry
|
|
38
|
+
```
|
|
32
39
|
|
|
33
|
-
|
|
40
|
+
- **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
|
|
41
|
+
- **What it is not:** a proxy, a gateway or an API router. It does not automatically compare prices or select models; your agent follows the rules and chooses.
|
|
42
|
+
- **Use it when:** you run more than one model or agent and want the expensive tier kept for planning and judgment.
|
|
43
|
+
- **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
|
|
34
44
|
|
|
35
|
-
|
|
45
|
+
Built from a working system, not a diagram: the routing rules, the protocols and the lane runner here run in production, generalized so they transfer to any stack.
|
|
36
46
|
|
|
37
47
|
## The three levels
|
|
38
48
|
|
|
49
|
+
|
|
39
50
|
| Level | You have | You get |
|
|
40
51
|
|---|---|---|
|
|
41
52
|
| **1 · Beginner** | one LLM or one agent | tiers, task classification, the two build checkpoints, the protocols (build, propagate, gap analysis, deep research, numbers and logic, memory and record, docs and proof), a task-bundle template, and your agent set up to follow them |
|
|
42
53
|
| **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts; `--model` / `--effort` to pin the route and log it), a delegation matrix generated from your selection, research triage across the lanes you have |
|
|
43
54
|
| **3 · Advanced** | a virtual machine | everything above, plus a gateway config rendered from the API keys you hold (asked separately from your CLIs), pinned images, box rules, privacy gates, and a weekly gap-analysis job with "what watches it" written down |
|
|
44
55
|
|
|
45
|
-
|
|
46
|
-
|
|
56
|
+
Levels explained: [Part 1](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md).
|
|
47
57
|
## The AIs it knows about
|
|
48
58
|
|
|
49
59
|
| Id | What | Level |
|
|
@@ -59,134 +69,39 @@ Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/pa
|
|
|
59
69
|
|
|
60
70
|
`npx model-orchestrator --list` prints the catalog with install and sign-in notes. Details: [docs/catalog.md](docs/catalog.md).
|
|
61
71
|
|
|
62
|
-
##
|
|
63
|
-
|
|
64
|
-
An orchestrator routes work. It does not make a model stop guessing numbers, it does not give it a memory, and it does not make it check a library's current docs before writing against it. Three tools close those gaps: codecalc and obsidian-tc are from the same maintainer, Context7 is from Upstash. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc,context7` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
|
|
65
|
-
|
|
66
|
-
| Tool | Closes | Default | You need first |
|
|
67
|
-
|---|---|---|---|
|
|
68
|
-
| [codecalc](https://github.com/The-40-Thieves/codecalc) | guessed numbers, comparisons, complexity and equivalence claims: exact arithmetic, code execution in 31 languages, SMT logic checks, `verify_translation` / `verify_optimization`; offline, no key | yes | Python 3.10+ and `uv`. `uvx 'codecalc[full]' setup --write` registers it with Claude Code, Claude Desktop, Cursor, VS Code, Zed; snippets for Codex, Antigravity, Qwen Code are written for you |
|
|
69
|
-
| [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes with a confirmation gate, folder ACLs, a poison scan on inferred writes; 163 tools, local by default; AGPL-3.0 | no | an Obsidian vault folder; Node 24+ or Bun 1.1+ (stricter than this installer); Ollama with `nomic-embed-text` or a cloud embeddings key; the Obsidian app and its Local REST API plugin only for live bridge tools. Skip it if you do not keep notes in Obsidian |
|
|
70
|
-
| [Context7](https://github.com/upstash/context7) | stale library recall: current, version-specific docs and code examples pulled into the prompt for any library, SDK, API or CLI; hosted, or `npx` locally; MIT | no | nothing to install for the hosted endpoint; Node.js 18+ for the local alternative; a free API key is optional, for a higher rate limit. Always makes a network call, unlike the other two: skip it offline |
|
|
71
|
-
|
|
72
|
-
Context7 pairs with codecalc rather than duplicating it: Context7 tells the agent what a library is documented to do on this version, codecalc runs the code and proves what it actually does. Docs never stand as proof on their own, and where the two disagree the run wins.
|
|
73
|
-
|
|
74
|
-
Whether or not you select them, every level carries the three rules they serve: `protocols/numbers-and-logic.md` (when calling a calculator is mandatory, how to report a computed figure, why a thought log is not evidence), `protocols/memory-and-record.md` (search before writing, the folder index is part of the change, one writer, inferred content marked as inferred), and `protocols/docs-then-prove.md` (current docs before writing a call, then a run proves it, the run wins on disagreement).
|
|
75
|
-
|
|
76
|
-
## The two folders every run writes to
|
|
77
|
-
|
|
78
|
-
An install has two targets, and a scripted run should set both.
|
|
79
|
-
|
|
80
|
-
| Flag | Default | What lands there |
|
|
81
|
-
|---|---|---|
|
|
82
|
-
| `--dir` | `./ai-orchestrator` | the docs, protocols and (level 2+) `bin/cli-run.mjs`. Named after what it contains, not after this package, so a project can hold one without looking like a checkout of it. Pass `--dir ./model-orchestrator` if you prefer the package name. |
|
|
83
|
-
| `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets three hook scripts in `.claude/hooks/`, wired by a settings snippet you merge yourself. |
|
|
84
|
-
|
|
85
|
-
`--project` defaulting to the current directory is the one that surprises people: run the command from your home folder with Claude Code as the primary and five agent files land in your home folder. The installer prints the resolved project path in the plan and says when you left it at the default. Set it.
|
|
72
|
+
## Measuring routing
|
|
86
73
|
|
|
87
|
-
|
|
74
|
+
A routing rule nobody measures is a rule nobody knows is followed. On a claude-code install, `route-metrics.mjs` turns every turn, dispatch and subagent start/stop into one JSON line under `~/.ai-orchestrator/route-metrics.jsonl`, including the lane your agent named in its own `<!-- route: <lane> | <why> -->` marker.
|
|
88
75
|
|
|
89
76
|
```bash
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
--dir ./ai-orchestrator --project ./my-app
|
|
93
|
-
# --yes selects the recommended companion tool (codecalc), which writes CODECALC.md and mcp/ snippets.
|
|
94
|
-
# Add --no-tools for none, or --tools codecalc,obsidian-tc,context7 to choose.
|
|
95
|
-
|
|
96
|
-
npx model-orchestrator --yes --level 3 --ais claude-code,codex,agy,grok,hermes,qwen,ollama --apis anthropic,openrouter --dry # print the plan, write nothing
|
|
97
|
-
npx model-orchestrator --yes --level 2 --ais claude-code,codex --project ~/my-app --dir ~/my-app/ai-orchestrator --no-tools # subagents into ~/my-app/.claude/agents
|
|
98
|
-
npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --update-docs # added a lane: regenerate the docs you never edited
|
|
77
|
+
node .claude/hooks/route-metrics.mjs --summary # since the log began
|
|
78
|
+
node .claude/hooks/route-metrics.mjs --summary --since 2026-09-01 # since a date
|
|
99
79
|
```
|
|
100
80
|
|
|
81
|
+
The report prints turns, **route-marker coverage** (the share of turns that carried a real lane, which answers "is the agent actually tagging its routing decisions?"), lanes by count, dispatches by `subagent_type`, dispatches with no matching start, and mean/max duration per agent type. It never logs prompt text, tool descriptions, or the "why" half of the marker. Fail-open by design: a miss is a missing log line, never a blocked turn.
|
|
82
|
+
|
|
101
83
|
## Claude Code plugin
|
|
102
84
|
|
|
103
|
-
The
|
|
85
|
+
The hooks and subagents also ship as a plugin, so they install and update through Claude Code itself:
|
|
104
86
|
|
|
105
87
|
```
|
|
106
88
|
/plugin marketplace add aunysillyme/model-orchestrator
|
|
107
89
|
/plugin install model-orchestrator@model-orchestrator
|
|
108
90
|
```
|
|
109
91
|
|
|
110
|
-
|
|
111
|
-
- **What it still needs from the installer:** the routing rules. A plugin runs no install step, so it reads the installer's default locations, `ai-orchestrator/ROUTING.md` then `ai-orchestrator/ORCHESTRATOR.md`, and names `npx model-orchestrator` when neither exists. A project installed with a different `--dir` should wire the installer's rendered hooks instead.
|
|
112
|
-
- **What it leaves out:** `route-metrics.mjs`, the routing log. It writes to disk and the plugin ships only hooks that read, so `npx model-orchestrator` is how you get it.
|
|
113
|
-
- **How it is kept honest:** `plugin/` is generated from `templates/` by `npm run gen:plugin`, and `test/plugin.test.js` fails when the committed bundle drifts, when a hook gains a network call, a write or a subprocess, or when an agent loses its tool list. The bundle passes `claude plugin validate --strict`, the check Anthropic's community marketplace review runs on every submission.
|
|
92
|
+
It ships the three hooks and the eight subagents, each with an explicit tool list, and loads them namespaced as `model-orchestrator:builder`. It does not ship the routing rules, because a plugin runs no install step: `npx model-orchestrator` still writes those. `plugin/` is generated from `templates/` and `test/plugin.test.js` fails when the committed bundle drifts, when a hook gains a network call, a write or a subprocess, or when an agent loses its tool list. Details: [plugin/README.md](plugin/README.md).
|
|
114
93
|
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
## What gets written (level 3, everything)
|
|
118
|
-
|
|
119
|
-
```
|
|
120
|
-
ai-orchestrator/
|
|
121
|
-
README.md start here, written for your level and your AIs
|
|
122
|
-
ORCHESTRATOR.md single-agent routing rules (level 1)
|
|
123
|
-
TASK_BUNDLE.md the brief every delegation carries
|
|
124
|
-
protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record · docs-then-prove
|
|
125
|
-
CODECALC.md OBSIDIAN-TC.md CONTEXT7.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
|
|
126
|
-
<project>/.claude/agents/ one per tier plus finding-verifier, done-verifier, reader, at the PROJECT root (if Claude Code is primary)
|
|
127
|
-
<project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart) + route-metrics.mjs (all five: see "Measuring routing" below), Claude Code only
|
|
128
|
-
CLAUDE.snippet.md the block to paste into your CLAUDE.md
|
|
129
|
-
settings.hooks.snippet.json the hooks block to merge into .claude/settings.json (Claude Code only)
|
|
130
|
-
ROUTING.md multi-lane decision tree (level 2+)
|
|
131
|
-
TIERS.md DELEGATION_MATRIX.md RESEARCH_TRIAGE.md CLI-RUN.md
|
|
132
|
-
bin/cli-run.mjs bin/lanes.json (node bin/cli-run.mjs --doctor is the smoke test)
|
|
133
|
-
vm/ gateway config, compose, box rules, privacy gates, jobs/ (level 3)
|
|
134
|
-
```
|
|
135
|
-
|
|
136
|
-
## Repo layout
|
|
137
|
-
|
|
138
|
-
| Folder | What |
|
|
139
|
-
|---|---|
|
|
140
|
-
| [`bin/`](bin/README.md) | `cli.js` (the installer) and `cli-run.mjs` (the lane runner) |
|
|
141
|
-
| [`src/`](src/README.md) | the catalog, the pure planner, detection, rendering |
|
|
142
|
-
| [`templates/`](templates/README.md) | everything the installer can write, by level, plus `tools/` for companions |
|
|
143
|
-
| [`docs/`](docs/README.md) | the three parts and the catalog |
|
|
144
|
-
| [`plugin/`](plugin/README.md) | the Claude Code plugin, generated from `templates/` by `npm run gen:plugin`; `.claude-plugin/marketplace.json` at the root lists it |
|
|
145
|
-
| [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu, macOS and Windows, Node 18/20/22 |
|
|
146
|
-
|
|
147
|
-
## What is enforced, what is delegated, what is an instruction
|
|
148
|
-
|
|
149
|
-
Most of what this package ships is text an agent is asked to follow. Be clear about which is which before relying on it unattended.
|
|
150
|
-
|
|
151
|
-
| Property | How it holds |
|
|
152
|
-
|---|---|
|
|
153
|
-
| Installer writes only inside `--dir` and `--project`, never a secret, never over a document without `--force` (or `--update-docs`, which touches only documents provably untouched since a previous run); machine-owned config always, runtime files only when provably untouched or with `--upgrade-runtime` | **enforced by code** (preflight, exclusive create, rollback, manifest hashes; tested) |
|
|
154
|
-
| `cli-run` exit codes, process-group kill on timeout and on SIGINT/SIGTERM, UTF-8-safe streaming, fixed-code durable log, `--expect-*` contracts with a pre-run snapshot | **enforced by code** (tested with stub lanes) |
|
|
155
|
-
| Codex audit lane runs read-only | **delegated to the vendor flag** (`--audit` → `--sandbox read-only`); commands and network still follow your codex config |
|
|
156
|
-
| Other lanes' permissions, sign-in state, model versions | **delegated to each vendor's own config**; `--doctor` checks presence, not versions |
|
|
157
|
-
| Gateway binds to loopback, keys by name only | **enforced in the generated files**; whether the gateway authenticates is your environment |
|
|
158
|
-
| Lane selection, tiers, privacy classes, one-writer, escalation, the protocols | **agent instructions**. Nothing here stops an agent that ignores its rules; the task bundle and the protocols make ignoring them visible, not impossible |
|
|
159
|
-
| Weekly audit bounded, previous report preserved | **enforced in the generated script and unit** (watchdog, temp-and-rename, `TimeoutStartSec`) |
|
|
160
|
-
|
|
161
|
-
If you need a property in the third row to be enforced, that is a router, a policy engine or a sandbox, and this package does not claim to be one.
|
|
162
|
-
|
|
163
|
-
### Vendor version compatibility
|
|
164
|
-
|
|
165
|
-
**This package detects that a binary exists. It does not check its version, and a present binary is not a working lane.** `--doctor` reports presence, and with `--run` sends one lane a one-word canary; neither validates that the vendor's flags, output shape or auth still match what the generated files assume.
|
|
166
|
-
|
|
167
|
-
The lane wiring and the output judges were written against these versions, which are the ones this release was exercised on:
|
|
168
|
-
|
|
169
|
-
<!-- vendor-table:start -->
|
|
170
|
-
|
|
171
|
-
| Lane | Vendor | Version this release was built against | Where that number is proved |
|
|
172
|
-
|---|---|---|---|
|
|
173
|
-
| `claude` | Anthropic | 2.1.226 | the npm pin the installer writes, `@anthropic-ai/claude-code@2.1.226` |
|
|
174
|
-
| `codex` | OpenAI | 0.153.4 | `test/fixtures/codex-0.153.4.jsonl`, a recorded run |
|
|
175
|
-
| `agy` | Google | 1.1.27 | `test/fixtures/agy-1.1.27.jsonl`, a recorded run |
|
|
176
|
-
| `grok` | xAI | 1.0.5 | `test/fixtures/grok-1.0.5.json`, a recorded run |
|
|
177
|
-
| `hermes` | Nous Research | 0.20.0 | `test/fixtures/hermes-0.20.0.txt`, a recorded run |
|
|
178
|
-
| `qwen` | Alibaba | 0.22.3 | `test/fixtures/qwen-0.22.3-nokey.json`, a recorded run |
|
|
179
|
-
| `ollama` | Ollama | 0.33.3 | the pinned image the level 3 box runs, `ollama/ollama:0.33.3` |
|
|
180
|
-
|
|
181
|
-
Generated from `src/catalog.js` by `npm run gen:catalog`; `npm test` fails if this table and the catalog disagree. Fixtures were captured 2026-09-06.
|
|
182
|
-
|
|
183
|
-
<!-- vendor-table:end -->
|
|
94
|
+
## Companion tools (all optional)
|
|
184
95
|
|
|
185
|
-
|
|
96
|
+
An orchestrator routes work. It does not make a model stop guessing numbers, give it a memory, or make it check a library's current docs before writing against it. Three tools close those gaps:
|
|
186
97
|
|
|
187
|
-
|
|
98
|
+
| Tool | Closes | Default |
|
|
99
|
+
|---|---|---|
|
|
100
|
+
| [codecalc](https://github.com/The-40-Thieves/codecalc) | guessed numbers: exact arithmetic, code execution in 31 languages, logic checks; offline, no key | yes |
|
|
101
|
+
| [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes; local by default | no |
|
|
102
|
+
| [Context7](https://github.com/upstash/context7) | stale library recall: current, version-specific docs pulled into the prompt | no |
|
|
188
103
|
|
|
189
|
-
|
|
104
|
+
Selecting one writes a doc and config snippets; it installs nothing. What each needs first, and why Context7 pairs with codecalc rather than duplicating it: [docs/companions.md](docs/companions.md). Whether or not you select them, every level carries the three rules they serve: `protocols/numbers-and-logic.md`, `protocols/memory-and-record.md` and `protocols/docs-then-prove.md`.
|
|
190
105
|
|
|
191
106
|
## Principles the whole thing rests on
|
|
192
107
|
|
|
@@ -198,101 +113,83 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
|
|
|
198
113
|
6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, only the brief carries what this task needs. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
|
|
199
114
|
7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
|
|
200
115
|
|
|
201
|
-
##
|
|
116
|
+
## Common questions
|
|
202
117
|
|
|
203
|
-
|
|
204
|
-
different directions, so `TIERS.md` states them separately rather than folding
|
|
205
|
-
them into the role:
|
|
118
|
+
### How do I cut token usage across Claude Code, Codex and Gemini?
|
|
206
119
|
|
|
207
|
-
|
|
208
|
-
reasoning than the reviewer judging its output. When the plan is airtight the
|
|
209
|
-
spec is carrying the thinking.
|
|
210
|
-
- **Stakes move the tier and the reader.** Security, privacy, data loss and
|
|
211
|
-
irreversible changes buy the challenge lane, a named check, a rollback path or
|
|
212
|
-
a human yes. A one-line change to an auth check is simple and high-stakes at
|
|
213
|
-
the same time, and it is the stakes that decide.
|
|
120
|
+
Install for the tools you have, then let the generated `ROUTING.md` decide the tier per task: bulk, reading and verification go to the fast tier or a cheaper CLI lane, and the deep tier only plans and judges. On Claude Code, execution goes to the `builder` subagent by default and the main session plans and verifies. Every lane call through `cli-run` logs the model and effort it ran with, so you can check where the tokens went.
|
|
214
121
|
|
|
215
|
-
|
|
216
|
-
lost data, or something you can't undo. Most tasks are low-stakes and route
|
|
217
|
-
normally.
|
|
122
|
+
### How do I route tasks to cheaper models?
|
|
218
123
|
|
|
219
|
-
The
|
|
220
|
-
unresolved checkpoint, an irreversible change. A task that merely feels hard is
|
|
221
|
-
a deep-tier task, not an escalation.
|
|
124
|
+
The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly never gets a frontier token.
|
|
222
125
|
|
|
223
|
-
|
|
126
|
+
### Is this an LLM router or an AI gateway?
|
|
224
127
|
|
|
225
|
-
|
|
226
|
-
cited line, states what would trigger the problem, then hunts for the guard,
|
|
227
|
-
caller or test that makes it impossible, and returns **CONFIRMED**,
|
|
228
|
-
**NOT_REPRODUCED** or **INCONCLUSIVE** per finding. Only CONFIRMED earns a
|
|
229
|
-
change. Use a different model family from the one that produced the finding
|
|
230
|
-
where you have one: a family asked to check its own claim tends to agree with
|
|
231
|
-
itself.
|
|
128
|
+
No. It routes at the task level, through instructions your agent follows and a runner for agent CLIs. If you want a service or proxy that picks or forwards the model on every API request, look at request-level routers and gateways such as RouteLLM, LiteLLM, OpenRouter or claude-code-router. They solve a different problem and can sit underneath this.
|
|
232
129
|
|
|
233
|
-
|
|
130
|
+
### Can an agent install and run it without a person?
|
|
234
131
|
|
|
235
|
-
|
|
132
|
+
Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run` previews the plan, and `--list` prints every supported AI. Nothing is appended to a file you already have; activation snippets are written next to your files for you to merge.
|
|
236
133
|
|
|
237
|
-
## Measuring routing
|
|
238
134
|
|
|
239
|
-
|
|
135
|
+
## Read next
|
|
240
136
|
|
|
241
|
-
|
|
242
|
-
|
|
243
|
-
|
|
244
|
-
|
|
137
|
+
| Doc | What is in it |
|
|
138
|
+
|---|---|
|
|
139
|
+
| [docs/install.md](docs/install.md) | every flag, the two folders a run writes to, headless examples, the full file list |
|
|
140
|
+
| [docs/how-it-routes.md](docs/how-it-routes.md) | role, complexity and stakes; the three verifier agents; pinning a lane's model and effort |
|
|
141
|
+
| [docs/guarantees.md](docs/guarantees.md) | what is enforced by code, what is delegated to a vendor flag, and what is only an instruction |
|
|
142
|
+
| [docs/part-1-beginner.md](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md) | the thinking behind each level |
|
|
143
|
+
| [docs/catalog.md](docs/catalog.md) | every supported AI with install and sign-in notes |
|
|
245
144
|
|
|
246
|
-
|
|
145
|
+
<details>
|
|
146
|
+
<summary><strong>Platform support, and every test this suite skips</strong></summary>
|
|
247
147
|
|
|
248
|
-
## Pin the route, or know that you did not
|
|
249
148
|
|
|
250
|
-
A lane with no
|
|
251
|
-
`bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
|
|
252
|
-
CLI configured months ago at a low reasoning effort keeps auditing at that
|
|
253
|
-
effort while your routing docs describe a second-opinion pass.
|
|
149
|
+
Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22), including lane execution end to end through `cli-run` against a fake CLI installed the same way npm installs a real one (a `.cmd` shim). `cli-run` never runs a lane through `cmd.exe` when it can avoid it: it resolves the shim to the Node script underneath and spawns Node directly, so a prompt reaching a real lane never passes through a Windows shell. A `.cmd` or `.bat` lane that cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, rather than run through `cmd.exe`: a batch file re-reads its arguments after `cmd.exe` has parsed them once, and no escaping fully contains a prompt through both passes. Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested on Windows too, including SIGTERM/SIGINT to the wrapper (Windows has no OS-level signals: both terminate it unconditionally, verified there rather than treated the same as POSIX). Five narrow skips remain on Windows, each for a POSIX behavior the OS or the CI shell genuinely does not have, and each named here because a test that is quietly skipped reads as a test that passed: `statSync().mode`'s executable bit (NTFS has none, so that one assertion is conditional inside a test that otherwise runs everywhere); a lane dying mid-run from a real POSIX signal (a real Windows lane cannot die "by signal"); running `weekly-audit.sh`'s watchdog functions for real under Git Bash's job control, both the end-to-end run and the `bounded()` timeout check (the script itself only ever runs on the Ubuntu box it targets); and a `mkfifo` FIFO at the rules path, the one case that proves `route-gate.mjs` cannot HANG on a non-regular file, since Windows has no `mkfifo` to build one (the guard behind it is covered on every OS by a directory at the same path); and an untracked `mkfifo` FIFO in the repository `cli-run --audit` sizes, the case that proves `--effort auto` never opens a non-regular file (the symlink half of that test runs on every OS). The list is not prose on trust: `test/prose.test.js` counts every `skip:` in the suite and fails if one of them is not documented here.
|
|
254
150
|
|
|
255
|
-
|
|
256
|
-
node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
|
|
257
|
-
node bin/cli-run.mjs --doctor # prints what each lane is pinned to, and what is not pinned
|
|
258
|
-
```
|
|
151
|
+
**Privacy.** The installer sends no telemetry and makes no network call of its own once it is running. Two things around that are worth being exact about:
|
|
259
152
|
|
|
260
|
-
|
|
261
|
-
|
|
262
|
-
runs that never reached a lane. It does not log an actual. One lane of five
|
|
263
|
-
(grok) reports a model id in its own output and the other four report none, so
|
|
264
|
-
an actual field would be present for one lane and missing for four, and it
|
|
265
|
-
would be a provider-supplied string, which the durable log deliberately never
|
|
266
|
-
holds.
|
|
153
|
+
- `npx model-orchestrator` is itself a download: npm fetches this package from the registry before any of it runs. `npm install -g model-orchestrator` once, then run `model-orchestrator`, if you would rather that happen exactly one time.
|
|
154
|
+
- A missing vendor CLI is *printed*, not installed. In an interactive run the installer offers to run one pinned `npm install -g` per package and only runs the ones you answer yes to; with `--yes` or `--no-install` it answers no for you and prints the command instead. Vendor shell installers (Antigravity, Grok) are only ever printed, alongside the `curl … | less` you would use to read one before running it.
|
|
267
155
|
|
|
268
|
-
|
|
156
|
+
`cli-run` talks to nothing but the vendor CLI you name.
|
|
269
157
|
|
|
270
|
-
### How do I cut token usage across Claude Code, Codex and Gemini?
|
|
271
158
|
|
|
272
|
-
|
|
159
|
+
</details>
|
|
273
160
|
|
|
274
|
-
|
|
161
|
+
<details>
|
|
162
|
+
<summary><strong>Vendor version compatibility</strong></summary>
|
|
275
163
|
|
|
276
|
-
The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly never gets a frontier token.
|
|
277
164
|
|
|
278
|
-
|
|
165
|
+
**This package detects that a binary exists. It does not check its version, and a present binary is not a working lane.** `--doctor` reports presence, and with `--run` sends one lane a one-word canary; neither validates that the vendor's flags, output shape or auth still match what the generated files assume.
|
|
279
166
|
|
|
280
|
-
|
|
167
|
+
The lane wiring and the output judges were written against these versions, which are the ones this release was exercised on:
|
|
281
168
|
|
|
282
|
-
|
|
169
|
+
<!-- vendor-table:start -->
|
|
283
170
|
|
|
284
|
-
|
|
171
|
+
| Lane | Vendor | Version this release was built against | Where that number is proved |
|
|
172
|
+
|---|---|---|---|
|
|
173
|
+
| `claude` | Anthropic | 2.1.226 | the npm pin the installer writes, `@anthropic-ai/claude-code@2.1.226` |
|
|
174
|
+
| `codex` | OpenAI | 0.153.4 | `test/fixtures/codex-0.153.4.jsonl`, a recorded run |
|
|
175
|
+
| `agy` | Google | 1.1.27 | `test/fixtures/agy-1.1.27.jsonl`, a recorded run |
|
|
176
|
+
| `grok` | xAI | 1.0.5 | `test/fixtures/grok-1.0.5.json`, a recorded run |
|
|
177
|
+
| `hermes` | Nous Research | 0.20.0 | `test/fixtures/hermes-0.20.0.txt`, a recorded run |
|
|
178
|
+
| `qwen` | Alibaba | 0.22.3 | `test/fixtures/qwen-0.22.3-nokey.json`, a recorded run |
|
|
179
|
+
| `ollama` | Ollama | 0.33.3 | the pinned image the level 3 box runs, `ollama/ollama:0.33.3` |
|
|
285
180
|
|
|
286
|
-
|
|
181
|
+
Generated from `src/catalog.js` by `npm run gen:catalog`; `npm test` fails if this table and the catalog disagree. Fixtures were captured 2026-09-06.
|
|
287
182
|
|
|
288
|
-
|
|
183
|
+
<!-- vendor-table:end -->
|
|
289
184
|
|
|
290
|
-
|
|
185
|
+
One number per lane, and it is the same number the installer pins: where a lane installs from npm, `builtAgainst` in the catalog *is* the pin, so "built against" and "pinned to" can never be two answers. That pin is a floor, not a ceiling: these CLIs ship breaking flag changes on their own schedules, so a newer version may work perfectly, or may change a flag the generated wiring passes. When a lane starts failing after a vendor upgrade, compare against this table first.
|
|
291
186
|
|
|
292
|
-
|
|
293
|
-
- A missing vendor CLI is *printed*, not installed. In an interactive run the installer offers to run one pinned `npm install -g` per package and only runs the ones you answer yes to; with `--yes` or `--no-install` it answers no for you and prints the command instead. Vendor shell installers (Antigravity, Grok) are only ever printed, alongside the `curl … | less` you would use to read one before running it.
|
|
187
|
+
**The live canary runs on your machine, with your credentials.** That is what `node bin/cli-run.mjs --doctor --run` is: it sends every enabled lane one tiny prompt through your own sign-ins and reports `canary ok` or `canary FAILED rc=` per lane. Run it after install, and again after any vendor upgrade.
|
|
294
188
|
|
|
295
|
-
|
|
189
|
+
It deliberately does not run in this repository's CI. A canary is only meaningful against real credentials, and there are no credentials a maintainer could supply that would tell **you** anything about **your** lanes: your sign-ins, your quota, your vendor versions. A maintainer-credential canary in CI would prove one machine works and bill someone per run to do it. So CI runs the full suite against stub lanes on Ubuntu, macOS and Windows, Node 18/20/22, plus a packaged install into a clean consumer, and the live check ships to you instead.
|
|
190
|
+
|
|
191
|
+
|
|
192
|
+
</details>
|
|
296
193
|
|
|
297
194
|
## Contributing
|
|
298
195
|
|
|
@@ -300,11 +197,7 @@ Add an AI to `src/catalog.js` and every prompt, table, config and doc picks it u
|
|
|
300
197
|
|
|
301
198
|
## Credits
|
|
302
199
|
|
|
303
|
-
- [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for
|
|
304
|
-
separating role, complexity and stakes instead of compressing them into one
|
|
305
|
-
scale, for recording the model and effort a lane was actually asked for, and
|
|
306
|
-
for verifying findings before they trigger repairs. All three shipped in
|
|
307
|
-
0.1.14.
|
|
200
|
+
- [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for separating role, complexity and stakes instead of compressing them into one scale, for recording the model and effort a lane was actually asked for, and for verifying findings before they trigger repairs. All three shipped in 0.1.14.
|
|
308
201
|
|
|
309
202
|
## License
|
|
310
203
|
|
package/docs/README.md
CHANGED
|
@@ -1,6 +1,13 @@
|
|
|
1
1
|
# docs/
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Reference, then reading. The reference pages carry the detail the README links out to; the three parts explain the thinking behind each level and how to grow from one to the next.
|
|
4
|
+
|
|
5
|
+
| Reference | What is in it | File |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| Installing | every flag, the two folders a run writes to, headless examples, the full file list | [install.md](install.md) |
|
|
8
|
+
| How it routes | role, complexity and stakes; the three verifier agents; pinning model and effort | [how-it-routes.md](how-it-routes.md) |
|
|
9
|
+
| Guarantees | what is enforced by code, delegated to a vendor flag, or only an instruction | [guarantees.md](guarantees.md) |
|
|
10
|
+
| Companion tools | codecalc, obsidian-tc and Context7: what each closes and what it needs first | [companions.md](companions.md) |
|
|
4
11
|
|
|
5
12
|
| Part | Read if | File |
|
|
6
13
|
|---|---|---|
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
# Companion tools
|
|
2
|
+
|
|
3
|
+
All optional. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc,context7` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
|
|
4
|
+
|
|
5
|
+
## Companion tools (all optional)
|
|
6
|
+
|
|
7
|
+
An orchestrator routes work. It does not make a model stop guessing numbers, it does not give it a memory, and it does not make it check a library's current docs before writing against it. Three tools close those gaps: codecalc and obsidian-tc are from the same maintainer, Context7 is from Upstash. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc,context7` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
|
|
8
|
+
|
|
9
|
+
| Tool | Closes | Default | You need first |
|
|
10
|
+
|---|---|---|---|
|
|
11
|
+
| [codecalc](https://github.com/The-40-Thieves/codecalc) | guessed numbers, comparisons, complexity and equivalence claims: exact arithmetic, code execution in 31 languages, SMT logic checks, `verify_translation` / `verify_optimization`; offline, no key | yes | Python 3.10+ and `uv`. `uvx 'codecalc[full]' setup --write` registers it with Claude Code, Claude Desktop, Cursor, VS Code, Zed; snippets for Codex, Antigravity, Qwen Code are written for you |
|
|
12
|
+
| [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes with a confirmation gate, folder ACLs, a poison scan on inferred writes; 163 tools, local by default; AGPL-3.0 | no | an Obsidian vault folder; Node 24+ or Bun 1.1+ (stricter than this installer); Ollama with `nomic-embed-text` or a cloud embeddings key; the Obsidian app and its Local REST API plugin only for live bridge tools. Skip it if you do not keep notes in Obsidian |
|
|
13
|
+
| [Context7](https://github.com/upstash/context7) | stale library recall: current, version-specific docs and code examples pulled into the prompt for any library, SDK, API or CLI; hosted, or `npx` locally; MIT | no | nothing to install for the hosted endpoint; Node.js 18+ for the local alternative; a free API key is optional, for a higher rate limit. Always makes a network call, unlike the other two: skip it offline |
|
|
14
|
+
|
|
15
|
+
Context7 pairs with codecalc rather than duplicating it: Context7 tells the agent what a library is documented to do on this version, codecalc runs the code and proves what it actually does. Docs never stand as proof on their own, and where the two disagree the run wins.
|
|
16
|
+
|
|
17
|
+
Whether or not you select them, every level carries the three rules they serve: `protocols/numbers-and-logic.md` (when calling a calculator is mandatory, how to report a computed figure, why a thought log is not evidence), `protocols/memory-and-record.md` (search before writing, the folder index is part of the change, one writer, inferred content marked as inferred), and `protocols/docs-then-prove.md` (current docs before writing a call, then a run proves it, the run wins on disagreement).
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
# What holds, and what is only asked for
|
|
2
|
+
|
|
3
|
+
Read this before running any of it unattended.
|
|
4
|
+
|
|
5
|
+
Most of what this package ships is text an agent is asked to follow. Be clear about which is which before relying on it unattended.
|
|
6
|
+
|
|
7
|
+
| Property | How it holds |
|
|
8
|
+
|---|---|
|
|
9
|
+
| Installer writes only inside `--dir` and `--project`, never a secret, never over a document without `--force` (or `--update-docs`, which touches only documents provably untouched since a previous run); machine-owned config always, runtime files only when provably untouched or with `--upgrade-runtime` | **enforced by code** (preflight, exclusive create, rollback, manifest hashes; tested) |
|
|
10
|
+
| `cli-run` exit codes, process-group kill on timeout and on SIGINT/SIGTERM, UTF-8-safe streaming, fixed-code durable log, `--expect-*` contracts with a pre-run snapshot | **enforced by code** (tested with stub lanes) |
|
|
11
|
+
| Codex audit lane runs read-only | **delegated to the vendor flag** (`--audit` → `--sandbox read-only`); commands and network still follow your codex config |
|
|
12
|
+
| Other lanes' permissions, sign-in state, model versions | **delegated to each vendor's own config**; `--doctor` checks presence, not versions |
|
|
13
|
+
| Gateway binds to loopback, keys by name only | **enforced in the generated files**; whether the gateway authenticates is your environment |
|
|
14
|
+
| Lane selection, tiers, privacy classes, one-writer, escalation, the protocols | **agent instructions**. Nothing here stops an agent that ignores its rules; the task bundle and the protocols make ignoring them visible, not impossible |
|
|
15
|
+
| Weekly audit bounded, previous report preserved | **enforced in the generated script and unit** (watchdog, temp-and-rename, `TimeoutStartSec`) |
|
|
16
|
+
|
|
17
|
+
If you need a property in the third row to be enforced, that is a router, a policy engine or a sandbox, and this package does not claim to be one.
|
|
18
|
+
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# How it routes
|
|
2
|
+
|
|
3
|
+
The reasoning behind the lane choice, and the three verifier agents that keep a routed answer honest. Install mechanics are in [install.md](install.md).
|
|
4
|
+
|
|
5
|
+
## Routing by role, complexity and stakes
|
|
6
|
+
|
|
7
|
+
Role picks the agent. Two more inputs move the choice, and they move it in
|
|
8
|
+
different directions, so `TIERS.md` states them separately rather than folding
|
|
9
|
+
them into the role:
|
|
10
|
+
|
|
11
|
+
- **Complexity moves the effort.** A worker executing a finished plan needs less
|
|
12
|
+
reasoning than the reviewer judging its output. When the plan is airtight the
|
|
13
|
+
spec is carrying the thinking.
|
|
14
|
+
- **Stakes move the tier and the reader.** Security, privacy, data loss and
|
|
15
|
+
irreversible changes buy the challenge lane, a named check, a rollback path or
|
|
16
|
+
a human yes. A one-line change to an auth check is simple and high-stakes at
|
|
17
|
+
the same time, and it is the stakes that decide.
|
|
18
|
+
|
|
19
|
+
Stakes means what a mistake would cost: a security hole, leaked personal data,
|
|
20
|
+
lost data, or something you can't undo. Most tasks are low-stakes and route
|
|
21
|
+
normally.
|
|
22
|
+
|
|
23
|
+
The top of the ladder is bought with evidence: a reproduced failure, an
|
|
24
|
+
unresolved checkpoint, an irreversible change. A task that merely feels hard is
|
|
25
|
+
a deep-tier task, not an escalation.
|
|
26
|
+
|
|
27
|
+
## A finding is a claim, not a fact
|
|
28
|
+
|
|
29
|
+
Review findings do not go straight to a repair. `finding-verifier` reads the
|
|
30
|
+
cited line, states what would trigger the problem, then hunts for the guard,
|
|
31
|
+
caller or test that makes it impossible, and returns **CONFIRMED**,
|
|
32
|
+
**NOT_REPRODUCED** or **INCONCLUSIVE** per finding. Only CONFIRMED earns a
|
|
33
|
+
change. Use a different model family from the one that produced the finding
|
|
34
|
+
where you have one: a family asked to check its own claim tends to agree with
|
|
35
|
+
itself.
|
|
36
|
+
|
|
37
|
+
## Two more fast-tier checks
|
|
38
|
+
|
|
39
|
+
`done-verifier` probes the artifact a tracker item's done-signal names (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything itself. It carries no file-editing tools, but on claude-code it does carry `Bash` for those probes (`git log`, `grep`, `wc -l`, `test -f`); staying to read-only commands there is a rule in its prompt, not a restriction on the tool grant, and its own description says so. On agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `reader` is the one that is read-only by tool grant on both: no `Write`, `Edit`, or `Bash`. It reads and digests many files or notes and hands back exactly what the brief asked for, cited by `path:line`; it never classifies, tags or writes, which is what separates it from `bulk-worker`. Both ship in the claude-code and agy agent sets, at the fast tier.
|
|
40
|
+
|
|
41
|
+
## Pin the route, or know that you did not
|
|
42
|
+
|
|
43
|
+
A lane with no `--model`, no `--effort` and no `defaults` entry in
|
|
44
|
+
`bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
|
|
45
|
+
CLI configured months ago at a low reasoning effort keeps auditing at that
|
|
46
|
+
effort while your routing docs describe a second-opinion pass.
|
|
47
|
+
|
|
48
|
+
```bash
|
|
49
|
+
node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
|
|
50
|
+
node bin/cli-run.mjs --doctor # prints what each lane is pinned to, and what is not pinned
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Every run logs the model and effort **requested** and where the request came
|
|
54
|
+
from: `flag`, `lanes.json`, or `lane_default`, on every record including the
|
|
55
|
+
runs that never reached a lane. It does not log an actual. One lane of five
|
|
56
|
+
(grok) reports a model id in its own output and the other four report none, so
|
|
57
|
+
an actual field would be present for one lane and missing for four, and it
|
|
58
|
+
would be a provider-supplied string, which the durable log deliberately never
|
|
59
|
+
holds.
|
|
60
|
+
|
package/docs/install.md
ADDED
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
# Installing
|
|
2
|
+
|
|
3
|
+
Everything the installer asks, writes and accepts as a flag. The short version is in the [README](../README.md).
|
|
4
|
+
|
|
5
|
+
## What a run does
|
|
6
|
+
|
|
7
|
+
|
|
8
|
+
The installer asks a few things, then writes a folder:
|
|
9
|
+
|
|
10
|
+
1. **Which level?** 1 beginner · 2 intermediate · 3 advanced
|
|
11
|
+
2. **Which AIs do you have access to?** (it marks the ones already on your PATH)
|
|
12
|
+
3. **Which one is your primary agent?** (the one that runs the system)
|
|
13
|
+
|
|
14
|
+
It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions (and, on Claude Code, two hook scripts) go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: follow the generated README. Inspect the manifest and remove only the individual managed subagent files you no longer need, preserve edited or pre-existing files, and remove your manually pasted activation block. Never delete a shared subagent folder.
|
|
15
|
+
|
|
16
|
+
## Plans and automatic effort
|
|
17
|
+
|
|
18
|
+
State known subscription plans with `--plans codex=pro-20x,agy=ultra-5x`. The generated guidance uses plan headroom to allocate volume only. It never changes capability or independent-review rules. `--effort-auto` is explicit consent to set `auto` only for selected high or max headroom CLI lanes. Auto chooses medium below 4,000 prompt characters and high otherwise, never higher. A codex audit is always high. Name `xhigh` explicitly for security-critical or irreversible work.
|
|
19
|
+
## The two folders every run writes to
|
|
20
|
+
|
|
21
|
+
An install has two targets, and a scripted run should set both.
|
|
22
|
+
|
|
23
|
+
| Flag | Default | What lands there |
|
|
24
|
+
|---|---|---|
|
|
25
|
+
| `--dir` | `./ai-orchestrator` | the docs, protocols and (level 2+) `bin/cli-run.mjs`. Named after what it contains, not after this package, so a project can hold one without looking like a checkout of it. Pass `--dir ./model-orchestrator` if you prefer the package name. |
|
|
26
|
+
| `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets three hook scripts in `.claude/hooks/`, wired by a settings snippet you merge yourself. |
|
|
27
|
+
|
|
28
|
+
`--project` defaulting to the current directory is the one that surprises people: run the command from your home folder with Claude Code as the primary and five agent files land in your home folder. The installer prints the resolved project path in the plan and says when you left it at the default. Set it.
|
|
29
|
+
|
|
30
|
+
## Non-interactive
|
|
31
|
+
|
|
32
|
+
```bash
|
|
33
|
+
# both targets set: docs in ./ai-orchestrator, subagents into ./my-app/.claude/agents
|
|
34
|
+
npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code \
|
|
35
|
+
--dir ./ai-orchestrator --project ./my-app
|
|
36
|
+
# --yes selects the recommended companion tool (codecalc), which writes CODECALC.md and mcp/ snippets.
|
|
37
|
+
# Add --no-tools for none, or --tools codecalc,obsidian-tc,context7 to choose.
|
|
38
|
+
|
|
39
|
+
npx model-orchestrator --yes --level 3 --ais claude-code,codex,agy,grok,hermes,qwen,ollama --apis anthropic,openrouter --dry # print the plan, write nothing
|
|
40
|
+
npx model-orchestrator --yes --level 2 --ais claude-code,codex --project ~/my-app --dir ~/my-app/ai-orchestrator --no-tools # subagents into ~/my-app/.claude/agents
|
|
41
|
+
npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --update-docs # added a lane: regenerate the docs you never edited
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
## What gets written (level 3, everything)
|
|
45
|
+
|
|
46
|
+
```
|
|
47
|
+
ai-orchestrator/
|
|
48
|
+
README.md start here, written for your level and your AIs
|
|
49
|
+
ORCHESTRATOR.md single-agent routing rules (level 1)
|
|
50
|
+
TASK_BUNDLE.md the brief every delegation carries
|
|
51
|
+
protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record · docs-then-prove
|
|
52
|
+
CODECALC.md OBSIDIAN-TC.md CONTEXT7.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
|
|
53
|
+
<project>/.claude/agents/ one per tier plus finding-verifier, done-verifier, reader, at the PROJECT root (if Claude Code is primary)
|
|
54
|
+
<project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart) + route-metrics.mjs (all five: see "Measuring routing" below), Claude Code only
|
|
55
|
+
CLAUDE.snippet.md the block to paste into your CLAUDE.md
|
|
56
|
+
settings.hooks.snippet.json the hooks block to merge into .claude/settings.json (Claude Code only)
|
|
57
|
+
ROUTING.md multi-lane decision tree (level 2+)
|
|
58
|
+
TIERS.md DELEGATION_MATRIX.md RESEARCH_TRIAGE.md CLI-RUN.md
|
|
59
|
+
bin/cli-run.mjs bin/lanes.json (node bin/cli-run.mjs --doctor is the smoke test)
|
|
60
|
+
vm/ gateway config, compose, box rules, privacy gates, jobs/ (level 3)
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
## Repo layout
|
|
64
|
+
|
|
65
|
+
| Folder | What |
|
|
66
|
+
|---|---|
|
|
67
|
+
| [`bin/`](bin/README.md) | `cli.js` (the installer) and `cli-run.mjs` (the lane runner) |
|
|
68
|
+
| [`src/`](src/README.md) | the catalog, the pure planner, detection, rendering |
|
|
69
|
+
| [`templates/`](templates/README.md) | everything the installer can write, by level, plus `tools/` for companions |
|
|
70
|
+
| [`docs/`](docs/README.md) | the three parts and the catalog |
|
|
71
|
+
| [`plugin/`](plugin/README.md) | the Claude Code plugin, generated from `templates/` by `npm run gen:plugin`; `.claude-plugin/marketplace.json` at the root lists it |
|
|
72
|
+
| [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu, macOS and Windows, Node 18/20/22 |
|
|
73
|
+
|
package/llms.txt
CHANGED
|
@@ -10,7 +10,11 @@ Claude Code plugin: `/plugin marketplace add aunysillyme/model-orchestrator`, th
|
|
|
10
10
|
|
|
11
11
|
## Docs
|
|
12
12
|
|
|
13
|
-
- [README](https://github.com/aunysillyme/model-orchestrator/blob/main/README.md): what it
|
|
13
|
+
- [README](https://github.com/aunysillyme/model-orchestrator/blob/main/README.md): what it is, a real dry-run plan, the levels, the AIs, the principles
|
|
14
|
+
- [Installing](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/install.md): every flag, the two folders a run writes to, headless examples, the full file list
|
|
15
|
+
- [How it routes](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/how-it-routes.md): role, complexity and stakes; the three verifier agents; pinning model and effort per lane
|
|
16
|
+
- [Guarantees](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/guarantees.md): what is enforced by code, what is delegated to a vendor flag, what is only an instruction
|
|
17
|
+
- [Companion tools](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/companions.md): codecalc, obsidian-tc and Context7, what each closes and what it needs first
|
|
14
18
|
- [Part 1: beginner](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-1-beginner.md): one agent, tiers, the task bundle every delegation carries
|
|
15
19
|
- [Part 2: intermediate](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-2-intermediate.md): several agent CLIs, the lane runner, pinning model and effort per lane
|
|
16
20
|
- [Part 3: advanced](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-3-advanced.md): the VM, the gateway, scheduled jobs, privacy gates
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "model-orchestrator",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.27",
|
|
4
4
|
"description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|