model-orchestrator 0.1.27 → 0.1.29

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,25 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.29] - 2026-09-21
8
+
9
+ ### Changed
10
+
11
+ - **The seven principles state what to do, rather than what fails.** "A gate you cannot fail is not a gate" became "Every gate can come back wrong. A checkpoint earns its place by being answerable both ways"; "Exit 0 is not a deliverable" became "Check for the artifact. A deliverable is a file, a commit or a line you can point at"; "A write nobody can find again did not happen" became "A write stays findable". Same rules, same gates, stated forward. This finishes the copy pass 0.1.28 started, which left the principle list untouched.
12
+
13
+
14
+ ## [0.1.28] - 2026-09-21
15
+
16
+ ### Added
17
+
18
+ - **A recording of the command doing its job, at the top of the README.** `docs/demo.gif` (60 KB) shows `npx model-orchestrator ... --dry` typed and run: the level, the AIs, both target folders, all 38 files it would write, and the closing line that nothing was written. `scripts/record-demo.sh` re-records it from the PUBLISHED package inside a temp folder, so the frames stay the program's own output rather than a staged screen, and anyone can reproduce them with `brew install asciinema agg`. The text plan stays in the README under the image, and the image carries alt text describing what it prints, so the page still reads with images off.
19
+ - **`--summary` now answers the question the package exists for: how much work left the main session.** `work sent off the main session` counts covered turns whose route marker named any lane that is not an inline name, over covered turns. It is derived only from lane names the log already holds, with no price table and no token estimate, because the log holds neither. `--inline a,b` renames what counts as inline, since the lane vocabulary belongs to your own `ROUTING.md`. A log with no lane markers yet says so instead of printing a number.
20
+
21
+ ### Changed
22
+
23
+ - **Every line of user-facing copy states what the package is and does.** The opening bullet was "What it is not: a proxy, a gateway or an API router", which told a new reader what to stop expecting before they knew what they were looking at. It now reads "Where it sits: above the request layer. Your agent reads the rules and picks the lane", and request-level routers are described as composing underneath rather than as the thing this is not. Same for the plugin section, the companion-tool intro and the gateway question in Common Questions. `test/copy.test.js` keeps the boundary it was written to protect: the OVERCLAIM guard is unchanged, and two positive phrases are now required in both README and `llms.txt`, so the claim cannot quietly widen and the copy cannot quietly lose it.
24
+
25
+
7
26
  ## [0.1.27] - 2026-09-21
8
27
 
9
28
  ### Changed
@@ -381,7 +400,9 @@ First release.
381
400
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
382
401
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
383
402
 
384
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.27...HEAD
403
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.29...HEAD
404
+ [0.1.29]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.28...v0.1.29
405
+ [0.1.28]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.27...v0.1.28
385
406
  [0.1.27]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.26...v0.1.27
386
407
  [0.1.26]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.25...v0.1.26
387
408
  [0.1.25]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.24...v0.1.25
package/README.md CHANGED
@@ -8,7 +8,9 @@
8
8
  npx model-orchestrator
9
9
  ```
10
10
 
11
- Three questions, then 38 files. Here is a real `--dry` run, which prints the plan and writes nothing:
11
+ <img src="docs/demo.gif" alt="A terminal running npx model-orchestrator with --dry: it prints the level, the AIs detected, both target folders and all 38 files it would write, then says nothing was written." width="100%" />
12
+
13
+ Three questions, then 38 files. The same plan as text:
12
14
 
13
15
  ```text
14
16
  Plan
@@ -31,18 +33,20 @@ Plan
31
33
  --dry: nothing written.
32
34
  ```
33
35
 
34
- Reproduce it:
36
+ Run it yourself, in any folder:
35
37
 
36
38
  ```bash
37
39
  npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --dry
38
40
  ```
39
41
 
42
+ The recording above comes from the published package under `asciinema`, rendered with `agg`: `bash scripts/record-demo.sh`.
43
+
40
44
  - **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
41
- - **What it is not:** a proxy, a gateway or an API router. It does not automatically compare prices or select models; your agent follows the rules and chooses.
45
+ - **Where it sits:** above the request layer. Your agent reads the rules and picks the lane, so the decision stays somewhere you can read, version and edit. Request-level routers and gateways sit underneath it.
42
46
  - **Use it when:** you run more than one model or agent and want the expensive tier kept for planning and judgment.
43
47
  - **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
44
48
 
45
- Built from a working system, not a diagram: the routing rules, the protocols and the lane runner here run in production, generalized so they transfer to any stack.
49
+ Built from a working system: the routing rules, the protocols and the lane runner here run in production every day, generalized so they transfer to any stack.
46
50
 
47
51
  ## The three levels
48
52
 
@@ -78,7 +82,7 @@ node .claude/hooks/route-metrics.mjs --summary # since th
78
82
  node .claude/hooks/route-metrics.mjs --summary --since 2026-09-01 # since a date
79
83
  ```
80
84
 
81
- The report prints turns, **route-marker coverage** (the share of turns that carried a real lane, which answers "is the agent actually tagging its routing decisions?"), lanes by count, dispatches by `subagent_type`, dispatches with no matching start, and mean/max duration per agent type. It never logs prompt text, tool descriptions, or the "why" half of the marker. Fail-open by design: a miss is a missing log line, never a blocked turn.
85
+ The report prints turns, **route-marker coverage** (the share of turns that carried a real lane, which answers "is the agent actually tagging its routing decisions?"), **work sent off the main session**, lanes by count, dispatches by `subagent_type`, and mean/max duration per agent type. The log holds exactly five things: a timestamp, the event, the session id, the lane name and the agent type. Every turn keeps running whatever the hook does, so the worst case is a quieter report.
82
86
 
83
87
  ## Claude Code plugin
84
88
 
@@ -89,11 +93,11 @@ The hooks and subagents also ship as a plugin, so they install and update throug
89
93
  /plugin install model-orchestrator@model-orchestrator
90
94
  ```
91
95
 
92
- It ships the three hooks and the eight subagents, each with an explicit tool list, and loads them namespaced as `model-orchestrator:builder`. It does not ship the routing rules, because a plugin runs no install step: `npx model-orchestrator` still writes those. `plugin/` is generated from `templates/` and `test/plugin.test.js` fails when the committed bundle drifts, when a hook gains a network call, a write or a subprocess, or when an agent loses its tool list. Details: [plugin/README.md](plugin/README.md).
96
+ It ships the three hooks and the eight subagents, each with an explicit tool list, and loads them namespaced as `model-orchestrator:builder`. The routing rules come from `npx model-orchestrator`, which is the step that reads your setup and writes rules to match it. `plugin/` is generated from `templates/`, and `test/plugin.test.js` holds the bundle to that shape: committed output matches the generator, hooks stay read-only, every agent keeps its tool list. Details: [plugin/README.md](plugin/README.md).
93
97
 
94
98
  ## Companion tools (all optional)
95
99
 
96
- An orchestrator routes work. It does not make a model stop guessing numbers, give it a memory, or make it check a library's current docs before writing against it. Three tools close those gaps:
100
+ An orchestrator routes work. Three companion tools cover the rest of what a working agent needs, exact numbers, a memory, and current library docs:
97
101
 
98
102
  | Tool | Closes | Default |
99
103
  |---|---|---|
@@ -101,17 +105,17 @@ An orchestrator routes work. It does not make a model stop guessing numbers, giv
101
105
  | [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes; local by default | no |
102
106
  | [Context7](https://github.com/upstash/context7) | stale library recall: current, version-specific docs pulled into the prompt | no |
103
107
 
104
- Selecting one writes a doc and config snippets; it installs nothing. What each needs first, and why Context7 pairs with codecalc rather than duplicating it: [docs/companions.md](docs/companions.md). Whether or not you select them, every level carries the three rules they serve: `protocols/numbers-and-logic.md`, `protocols/memory-and-record.md` and `protocols/docs-then-prove.md`.
108
+ Selecting one writes a doc and the config snippets for your agents, so the install stays yours to run. What each needs first, and how Context7 and codecalc pair up (docs say what an API should do, a run proves what it does): [docs/companions.md](docs/companions.md). Every level carries the three rules they serve either way: `protocols/numbers-and-logic.md`, `protocols/memory-and-record.md` and `protocols/docs-then-prove.md`.
105
109
 
106
110
  ## Principles the whole thing rests on
107
111
 
108
- 1. **Route by capability tier, not model name.** Default down, escalate on evidence.
109
- 2. **A gate you cannot fail is not a gate.** Every checkpoint is a question that can come back wrong.
110
- 3. **Exit 0 is not a deliverable.** Check for the artifact, not the status line. `cli-run` checks the response is structurally there; `--expect-file` checks the artifact.
111
- 4. **Numbers are computed, never guessed.** A tool that calculates beats a model that feels finished.
112
- 5. **A write nobody can find again did not happen.** Search first, keep the index true, one writer.
113
- 6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, only the brief carries what this task needs. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
114
- 7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
112
+ 1. **Route by capability tier.** Start at the smallest tier that fits, and let evidence move it up.
113
+ 2. **Every gate can come back wrong.** A checkpoint earns its place by being answerable both ways.
114
+ 3. **Check for the artifact.** A deliverable is a file, a commit or a line you can point at. `cli-run` checks the response is structurally there; `--expect-file` checks the artifact.
115
+ 4. **Numbers are computed.** A tool that calculates beats a model that feels finished.
116
+ 5. **A write stays findable.** Search first, keep the index true, one writer.
117
+ 6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, the brief is what carries this task. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
118
+ 7. **Only one process holds keys.** Names in the environment, values in a secrets manager.
115
119
 
116
120
  ## Common questions
117
121
 
@@ -121,15 +125,15 @@ Install for the tools you have, then let the generated `ROUTING.md` decide the t
121
125
 
122
126
  ### How do I route tasks to cheaper models?
123
127
 
124
- The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly never gets a frontier token.
128
+ The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly stays on the cheap tier, and the frontier tokens go to the work that earns them.
125
129
 
126
- ### Is this an LLM router or an AI gateway?
130
+ ### Where does this sit next to an LLM router or an AI gateway?
127
131
 
128
- No. It routes at the task level, through instructions your agent follows and a runner for agent CLIs. If you want a service or proxy that picks or forwards the model on every API request, look at request-level routers and gateways such as RouteLLM, LiteLLM, OpenRouter or claude-code-router. They solve a different problem and can sit underneath this.
132
+ One layer up, and they compose. This routes at the task level, through instructions your agent follows and a runner for agent CLIs. Request-level routers and gateways (RouteLLM, LiteLLM, OpenRouter, claude-code-router) forward the model on every API call, and they sit underneath this happily: pick the lane here, let the gateway carry the call.
129
133
 
130
134
  ### Can an agent install and run it without a person?
131
135
 
132
- Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run` previews the plan, and `--list` prints every supported AI. Nothing is appended to a file you already have; activation snippets are written next to your files for you to merge.
136
+ Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run` previews the plan, and `--list` prints every supported AI. Your existing files stay as they are: activation snippets land beside them, ready to merge when you choose.
133
137
 
134
138
 
135
139
  ## Read next
package/docs/demo.gif ADDED
Binary file
package/llms.txt CHANGED
@@ -1,6 +1,6 @@
1
1
  # model-orchestrator
2
2
 
3
- > Model orchestrator for AI coding agents and LLMs (Claude Code, Codex, Gemini, Grok, Qwen, Ollama). One installer writes routing rules, subagent definitions and a CLI lane runner for the AI tools you already have. The rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. It is not a proxy or gateway: it does not automatically compare prices or select models; your agent follows the rules and chooses.
3
+ > Model orchestrator for AI coding agents and LLMs (Claude Code, Codex, Gemini, Grok, Qwen, Ollama). One installer writes routing rules, subagent definitions and a CLI lane runner for the AI tools you already have. The rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. It sits above the request layer: your agent reads the rules and picks the lane, so the decision stays readable, versioned and editable. Request-level routers and gateways compose underneath it.
4
4
 
5
5
  Install and run: `npx model-orchestrator` (interactive), or headless: `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator`. Preview without writing: add `--dry-run`. List every supported AI: `npx model-orchestrator --list`. Node 18 or newer, zero runtime dependencies, MIT licence.
6
6
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.27",
3
+ "version": "0.1.29",
4
4
  "description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
5
5
  "type": "module",
6
6
  "bin": {
package/scripts/README.md CHANGED
@@ -3,4 +3,5 @@
3
3
  | File | Job |
4
4
  |---|---|
5
5
  | `gen-catalog.js` | regenerates `docs/catalog.md` AND the vendor compatibility table in `README.md` (between the `vendor-table` markers) from `src/catalog.js`; `npm run gen:catalog`. `test/catalog.test.js` fails if either generated surface disagrees with the catalog. |
6
+ | `record-demo.sh` | re-records `docs/demo.gif` by installing the published package into a temp folder and running it under `asciinema`, then rendering the cast with `agg`. Pass a version to pin one: `bash scripts/record-demo.sh 0.1.27`. Needs `brew install asciinema agg`. The frames are the installer's own output, so the GIF stays true to what the command prints. |
6
7
  | `gen-plugin.js` | regenerates the Claude Code plugin bundle in `plugin/` (agents, the two read-only hooks, `plugin.json`, `LICENSE`) from `templates/`, using the plan in `src/plugin.js`; `npm run gen:plugin`. `test/plugin.test.js` fails if the committed bundle disagrees. `plugin/README.md` and `plugin/hooks/hooks.json` are hand-owned. |
@@ -0,0 +1,45 @@
1
+ #!/usr/bin/env bash
2
+ # Re-records docs/demo.gif from the PUBLISHED package, so the frames are the
3
+ # program's own output and anyone can reproduce them.
4
+ #
5
+ # bash scripts/record-demo.sh # records the current published version
6
+ # bash scripts/record-demo.sh 0.1.27 # pins a version
7
+ #
8
+ # Needs asciinema (the recorder) and agg (cast to GIF), both from the asciinema
9
+ # project: brew install asciinema agg
10
+ set -euo pipefail
11
+
12
+ VERSION="${1:-latest}"
13
+ REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
14
+ OUT="$REPO_ROOT/docs/demo.gif"
15
+ WORK="$(mktemp -d)"
16
+ trap 'rm -rf "$WORK"' EXIT
17
+
18
+ for tool in asciinema agg npm; do
19
+ command -v "$tool" >/dev/null || { echo "missing: $tool (brew install asciinema agg)" >&2; exit 1; }
20
+ done
21
+
22
+ # Installed locally first, so the recorded npx call resolves from node_modules
23
+ # and never stops on npx's own "Ok to proceed?" prompt mid-take.
24
+ cd "$WORK"
25
+ npm install --silent "model-orchestrator@$VERSION" >/dev/null
26
+
27
+ cat > "$WORK/take.sh" <<'TAKE'
28
+ #!/bin/bash
29
+ cd "$(dirname "$0")"
30
+ printf '~/my-app $ '
31
+ sleep 0.8
32
+ cmd='npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dry'
33
+ for (( i=0; i<${#cmd}; i++ )); do printf '%s' "${cmd:$i:1}"; sleep 0.022; done
34
+ sleep 0.6; printf '\n'
35
+ npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dry
36
+ printf '~/my-app $ '
37
+ sleep 2.5
38
+ TAKE
39
+ chmod +x "$WORK/take.sh"
40
+
41
+ asciinema rec --cols 108 --rows 34 --overwrite --command "$WORK/take.sh" "$WORK/demo.cast"
42
+ agg --theme monokai --font-size 16 --speed 1.3 --idle-time-limit 1.2 --last-frame-duration 3 \
43
+ "$WORK/demo.cast" "$OUT"
44
+
45
+ echo "wrote $OUT ($(wc -c < "$OUT") bytes)"
@@ -273,6 +273,10 @@ async function runHook() {
273
273
  process.exit(0); // fail-open, always: a miss here is a missing log line, never a blocked turn
274
274
  }
275
275
 
276
+ // Lane names that mean "the main session did it itself". Defaults only: the lane
277
+ // vocabulary belongs to the user's own ROUTING.md, and --inline overrides this.
278
+ const INLINE_LANE_NAMES = ['inline', 'main', 'main inline'];
279
+
276
280
  // ---- --summary: a plain-text report, no stdin involved ----
277
281
 
278
282
  function parseLines(text) {
@@ -313,6 +317,22 @@ function runSummary(args) {
313
317
  for (const lane of Array.isArray(r.lane) ? r.lane : []) laneCounts.set(lane, (laneCounts.get(lane) || 0) + 1);
314
318
  }
315
319
 
320
+ // The one number that answers "did the routing actually move work off the main
321
+ // session?". Derived only from lane names the marker already carries: no price
322
+ // table, no token count, nothing this log does not hold. A covered turn counts as
323
+ // sent off when its marker names any lane that is not one of the inline names.
324
+ // The inline names are overridable because the lane vocabulary is the user's own
325
+ // ROUTING.md, not a list this package gets to fix.
326
+ const inlineIdx = args.indexOf('--inline');
327
+ const inlineNames = new Set(
328
+ (inlineIdx !== -1 && args[inlineIdx + 1] ? args[inlineIdx + 1].split(',') : INLINE_LANE_NAMES)
329
+ .map((n) => n.trim().toLowerCase())
330
+ .filter(Boolean)
331
+ );
332
+ const coveredRoutes = routes.filter((r) => Array.isArray(r.lane) && r.lane.some((l) => l !== 'missing'));
333
+ const sentOff = coveredRoutes.filter((r) => r.lane.some((l) => l !== 'missing' && !inlineNames.has(String(l).trim().toLowerCase()))).length;
334
+ const sentOffPct = coveredRoutes.length > 0 ? (sentOff / coveredRoutes.length) * 100 : null;
335
+
316
336
  const dispatches = records.filter((r) => r.event === 'dispatch');
317
337
  const dispatchCounts = new Map();
318
338
  for (const r of dispatches) dispatchCounts.set(r.subagent_type, (dispatchCounts.get(r.subagent_type) || 0) + 1);
@@ -331,6 +351,12 @@ function runSummary(args) {
331
351
  lines.push('route-metrics summary' + (Number.isNaN(since) ? '' : ' since ' + args[sinceIdx + 1]));
332
352
  lines.push('turns: ' + turns);
333
353
  lines.push('route-marker coverage: ' + (coveragePct === null ? 'no turns yet' : formatNumber(coveragePct) + '%') + ' (' + covered + '/' + turns + ')');
354
+ lines.push(
355
+ 'work sent off the main session: ' +
356
+ (sentOffPct === null ? 'no covered turns yet' : formatNumber(sentOffPct) + '%') +
357
+ ' (' + sentOff + '/' + coveredRoutes.length + ' covered turns)'
358
+ );
359
+ lines.push(' inline names for this report: ' + [...inlineNames].join(', ') + ' (override with --inline a,b)');
334
360
  lines.push('lanes by count:');
335
361
  if (laneCounts.size === 0) lines.push(' (none)');
336
362
  for (const [lane, count] of [...laneCounts.entries()].sort((a, b) => b[1] - a[1])) lines.push(' ' + lane + ': ' + count);