model-orchestrator 0.1.27 → 0.1.29
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +22 -1
- package/README.md +23 -19
- package/docs/demo.gif +0 -0
- package/llms.txt +1 -1
- package/package.json +1 -1
- package/scripts/README.md +1 -0
- package/scripts/record-demo.sh +45 -0
- package/templates/agents/snippets/route-metrics.mjs +26 -0
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,25 @@ All notable changes to this project are documented here. The format follows [Kee
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [0.1.29] - 2026-09-21
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
- **The seven principles state what to do, rather than what fails.** "A gate you cannot fail is not a gate" became "Every gate can come back wrong. A checkpoint earns its place by being answerable both ways"; "Exit 0 is not a deliverable" became "Check for the artifact. A deliverable is a file, a commit or a line you can point at"; "A write nobody can find again did not happen" became "A write stays findable". Same rules, same gates, stated forward. This finishes the copy pass 0.1.28 started, which left the principle list untouched.
|
|
12
|
+
|
|
13
|
+
|
|
14
|
+
## [0.1.28] - 2026-09-21
|
|
15
|
+
|
|
16
|
+
### Added
|
|
17
|
+
|
|
18
|
+
- **A recording of the command doing its job, at the top of the README.** `docs/demo.gif` (60 KB) shows `npx model-orchestrator ... --dry` typed and run: the level, the AIs, both target folders, all 38 files it would write, and the closing line that nothing was written. `scripts/record-demo.sh` re-records it from the PUBLISHED package inside a temp folder, so the frames stay the program's own output rather than a staged screen, and anyone can reproduce them with `brew install asciinema agg`. The text plan stays in the README under the image, and the image carries alt text describing what it prints, so the page still reads with images off.
|
|
19
|
+
- **`--summary` now answers the question the package exists for: how much work left the main session.** `work sent off the main session` counts covered turns whose route marker named any lane that is not an inline name, over covered turns. It is derived only from lane names the log already holds, with no price table and no token estimate, because the log holds neither. `--inline a,b` renames what counts as inline, since the lane vocabulary belongs to your own `ROUTING.md`. A log with no lane markers yet says so instead of printing a number.
|
|
20
|
+
|
|
21
|
+
### Changed
|
|
22
|
+
|
|
23
|
+
- **Every line of user-facing copy states what the package is and does.** The opening bullet was "What it is not: a proxy, a gateway or an API router", which told a new reader what to stop expecting before they knew what they were looking at. It now reads "Where it sits: above the request layer. Your agent reads the rules and picks the lane", and request-level routers are described as composing underneath rather than as the thing this is not. Same for the plugin section, the companion-tool intro and the gateway question in Common Questions. `test/copy.test.js` keeps the boundary it was written to protect: the OVERCLAIM guard is unchanged, and two positive phrases are now required in both README and `llms.txt`, so the claim cannot quietly widen and the copy cannot quietly lose it.
|
|
24
|
+
|
|
25
|
+
|
|
7
26
|
## [0.1.27] - 2026-09-21
|
|
8
27
|
|
|
9
28
|
### Changed
|
|
@@ -381,7 +400,9 @@ First release.
|
|
|
381
400
|
- Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
|
|
382
401
|
- Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
|
|
383
402
|
|
|
384
|
-
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.
|
|
403
|
+
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.29...HEAD
|
|
404
|
+
[0.1.29]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.28...v0.1.29
|
|
405
|
+
[0.1.28]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.27...v0.1.28
|
|
385
406
|
[0.1.27]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.26...v0.1.27
|
|
386
407
|
[0.1.26]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.25...v0.1.26
|
|
387
408
|
[0.1.25]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.24...v0.1.25
|
package/README.md
CHANGED
|
@@ -8,7 +8,9 @@
|
|
|
8
8
|
npx model-orchestrator
|
|
9
9
|
```
|
|
10
10
|
|
|
11
|
-
|
|
11
|
+
<img src="docs/demo.gif" alt="A terminal running npx model-orchestrator with --dry: it prints the level, the AIs detected, both target folders and all 38 files it would write, then says nothing was written." width="100%" />
|
|
12
|
+
|
|
13
|
+
Three questions, then 38 files. The same plan as text:
|
|
12
14
|
|
|
13
15
|
```text
|
|
14
16
|
Plan
|
|
@@ -31,18 +33,20 @@ Plan
|
|
|
31
33
|
--dry: nothing written.
|
|
32
34
|
```
|
|
33
35
|
|
|
34
|
-
|
|
36
|
+
Run it yourself, in any folder:
|
|
35
37
|
|
|
36
38
|
```bash
|
|
37
39
|
npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --dry
|
|
38
40
|
```
|
|
39
41
|
|
|
42
|
+
The recording above comes from the published package under `asciinema`, rendered with `agg`: `bash scripts/record-demo.sh`.
|
|
43
|
+
|
|
40
44
|
- **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
|
|
41
|
-
- **
|
|
45
|
+
- **Where it sits:** above the request layer. Your agent reads the rules and picks the lane, so the decision stays somewhere you can read, version and edit. Request-level routers and gateways sit underneath it.
|
|
42
46
|
- **Use it when:** you run more than one model or agent and want the expensive tier kept for planning and judgment.
|
|
43
47
|
- **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
|
|
44
48
|
|
|
45
|
-
Built from a working system
|
|
49
|
+
Built from a working system: the routing rules, the protocols and the lane runner here run in production every day, generalized so they transfer to any stack.
|
|
46
50
|
|
|
47
51
|
## The three levels
|
|
48
52
|
|
|
@@ -78,7 +82,7 @@ node .claude/hooks/route-metrics.mjs --summary # since th
|
|
|
78
82
|
node .claude/hooks/route-metrics.mjs --summary --since 2026-09-01 # since a date
|
|
79
83
|
```
|
|
80
84
|
|
|
81
|
-
The report prints turns, **route-marker coverage** (the share of turns that carried a real lane, which answers "is the agent actually tagging its routing decisions?"), lanes by count, dispatches by `subagent_type`,
|
|
85
|
+
The report prints turns, **route-marker coverage** (the share of turns that carried a real lane, which answers "is the agent actually tagging its routing decisions?"), **work sent off the main session**, lanes by count, dispatches by `subagent_type`, and mean/max duration per agent type. The log holds exactly five things: a timestamp, the event, the session id, the lane name and the agent type. Every turn keeps running whatever the hook does, so the worst case is a quieter report.
|
|
82
86
|
|
|
83
87
|
## Claude Code plugin
|
|
84
88
|
|
|
@@ -89,11 +93,11 @@ The hooks and subagents also ship as a plugin, so they install and update throug
|
|
|
89
93
|
/plugin install model-orchestrator@model-orchestrator
|
|
90
94
|
```
|
|
91
95
|
|
|
92
|
-
It ships the three hooks and the eight subagents, each with an explicit tool list, and loads them namespaced as `model-orchestrator:builder`.
|
|
96
|
+
It ships the three hooks and the eight subagents, each with an explicit tool list, and loads them namespaced as `model-orchestrator:builder`. The routing rules come from `npx model-orchestrator`, which is the step that reads your setup and writes rules to match it. `plugin/` is generated from `templates/`, and `test/plugin.test.js` holds the bundle to that shape: committed output matches the generator, hooks stay read-only, every agent keeps its tool list. Details: [plugin/README.md](plugin/README.md).
|
|
93
97
|
|
|
94
98
|
## Companion tools (all optional)
|
|
95
99
|
|
|
96
|
-
An orchestrator routes work.
|
|
100
|
+
An orchestrator routes work. Three companion tools cover the rest of what a working agent needs, exact numbers, a memory, and current library docs:
|
|
97
101
|
|
|
98
102
|
| Tool | Closes | Default |
|
|
99
103
|
|---|---|---|
|
|
@@ -101,17 +105,17 @@ An orchestrator routes work. It does not make a model stop guessing numbers, giv
|
|
|
101
105
|
| [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes; local by default | no |
|
|
102
106
|
| [Context7](https://github.com/upstash/context7) | stale library recall: current, version-specific docs pulled into the prompt | no |
|
|
103
107
|
|
|
104
|
-
Selecting one writes a doc and config snippets
|
|
108
|
+
Selecting one writes a doc and the config snippets for your agents, so the install stays yours to run. What each needs first, and how Context7 and codecalc pair up (docs say what an API should do, a run proves what it does): [docs/companions.md](docs/companions.md). Every level carries the three rules they serve either way: `protocols/numbers-and-logic.md`, `protocols/memory-and-record.md` and `protocols/docs-then-prove.md`.
|
|
105
109
|
|
|
106
110
|
## Principles the whole thing rests on
|
|
107
111
|
|
|
108
|
-
1. **Route by capability tier
|
|
109
|
-
2. **
|
|
110
|
-
3. **
|
|
111
|
-
4. **Numbers are computed
|
|
112
|
-
5. **A write
|
|
113
|
-
6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way,
|
|
114
|
-
7. **Only one process holds keys.** Names in the environment, values in a secrets manager
|
|
112
|
+
1. **Route by capability tier.** Start at the smallest tier that fits, and let evidence move it up.
|
|
113
|
+
2. **Every gate can come back wrong.** A checkpoint earns its place by being answerable both ways.
|
|
114
|
+
3. **Check for the artifact.** A deliverable is a file, a commit or a line you can point at. `cli-run` checks the response is structurally there; `--expect-file` checks the artifact.
|
|
115
|
+
4. **Numbers are computed.** A tool that calculates beats a model that feels finished.
|
|
116
|
+
5. **A write stays findable.** Search first, keep the index true, one writer.
|
|
117
|
+
6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, the brief is what carries this task. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
|
|
118
|
+
7. **Only one process holds keys.** Names in the environment, values in a secrets manager.
|
|
115
119
|
|
|
116
120
|
## Common questions
|
|
117
121
|
|
|
@@ -121,15 +125,15 @@ Install for the tools you have, then let the generated `ROUTING.md` decide the t
|
|
|
121
125
|
|
|
122
126
|
### How do I route tasks to cheaper models?
|
|
123
127
|
|
|
124
|
-
The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly
|
|
128
|
+
The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly stays on the cheap tier, and the frontier tokens go to the work that earns them.
|
|
125
129
|
|
|
126
|
-
###
|
|
130
|
+
### Where does this sit next to an LLM router or an AI gateway?
|
|
127
131
|
|
|
128
|
-
|
|
132
|
+
One layer up, and they compose. This routes at the task level, through instructions your agent follows and a runner for agent CLIs. Request-level routers and gateways (RouteLLM, LiteLLM, OpenRouter, claude-code-router) forward the model on every API call, and they sit underneath this happily: pick the lane here, let the gateway carry the call.
|
|
129
133
|
|
|
130
134
|
### Can an agent install and run it without a person?
|
|
131
135
|
|
|
132
|
-
Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run` previews the plan, and `--list` prints every supported AI.
|
|
136
|
+
Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run` previews the plan, and `--list` prints every supported AI. Your existing files stay as they are: activation snippets land beside them, ready to merge when you choose.
|
|
133
137
|
|
|
134
138
|
|
|
135
139
|
## Read next
|
package/docs/demo.gif
ADDED
|
Binary file
|
package/llms.txt
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# model-orchestrator
|
|
2
2
|
|
|
3
|
-
> Model orchestrator for AI coding agents and LLMs (Claude Code, Codex, Gemini, Grok, Qwen, Ollama). One installer writes routing rules, subagent definitions and a CLI lane runner for the AI tools you already have. The rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. It
|
|
3
|
+
> Model orchestrator for AI coding agents and LLMs (Claude Code, Codex, Gemini, Grok, Qwen, Ollama). One installer writes routing rules, subagent definitions and a CLI lane runner for the AI tools you already have. The rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. It sits above the request layer: your agent reads the rules and picks the lane, so the decision stays readable, versioned and editable. Request-level routers and gateways compose underneath it.
|
|
4
4
|
|
|
5
5
|
Install and run: `npx model-orchestrator` (interactive), or headless: `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator`. Preview without writing: add `--dry-run`. List every supported AI: `npx model-orchestrator --list`. Node 18 or newer, zero runtime dependencies, MIT licence.
|
|
6
6
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "model-orchestrator",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.29",
|
|
4
4
|
"description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
package/scripts/README.md
CHANGED
|
@@ -3,4 +3,5 @@
|
|
|
3
3
|
| File | Job |
|
|
4
4
|
|---|---|
|
|
5
5
|
| `gen-catalog.js` | regenerates `docs/catalog.md` AND the vendor compatibility table in `README.md` (between the `vendor-table` markers) from `src/catalog.js`; `npm run gen:catalog`. `test/catalog.test.js` fails if either generated surface disagrees with the catalog. |
|
|
6
|
+
| `record-demo.sh` | re-records `docs/demo.gif` by installing the published package into a temp folder and running it under `asciinema`, then rendering the cast with `agg`. Pass a version to pin one: `bash scripts/record-demo.sh 0.1.27`. Needs `brew install asciinema agg`. The frames are the installer's own output, so the GIF stays true to what the command prints. |
|
|
6
7
|
| `gen-plugin.js` | regenerates the Claude Code plugin bundle in `plugin/` (agents, the two read-only hooks, `plugin.json`, `LICENSE`) from `templates/`, using the plan in `src/plugin.js`; `npm run gen:plugin`. `test/plugin.test.js` fails if the committed bundle disagrees. `plugin/README.md` and `plugin/hooks/hooks.json` are hand-owned. |
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# Re-records docs/demo.gif from the PUBLISHED package, so the frames are the
|
|
3
|
+
# program's own output and anyone can reproduce them.
|
|
4
|
+
#
|
|
5
|
+
# bash scripts/record-demo.sh # records the current published version
|
|
6
|
+
# bash scripts/record-demo.sh 0.1.27 # pins a version
|
|
7
|
+
#
|
|
8
|
+
# Needs asciinema (the recorder) and agg (cast to GIF), both from the asciinema
|
|
9
|
+
# project: brew install asciinema agg
|
|
10
|
+
set -euo pipefail
|
|
11
|
+
|
|
12
|
+
VERSION="${1:-latest}"
|
|
13
|
+
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
|
14
|
+
OUT="$REPO_ROOT/docs/demo.gif"
|
|
15
|
+
WORK="$(mktemp -d)"
|
|
16
|
+
trap 'rm -rf "$WORK"' EXIT
|
|
17
|
+
|
|
18
|
+
for tool in asciinema agg npm; do
|
|
19
|
+
command -v "$tool" >/dev/null || { echo "missing: $tool (brew install asciinema agg)" >&2; exit 1; }
|
|
20
|
+
done
|
|
21
|
+
|
|
22
|
+
# Installed locally first, so the recorded npx call resolves from node_modules
|
|
23
|
+
# and never stops on npx's own "Ok to proceed?" prompt mid-take.
|
|
24
|
+
cd "$WORK"
|
|
25
|
+
npm install --silent "model-orchestrator@$VERSION" >/dev/null
|
|
26
|
+
|
|
27
|
+
cat > "$WORK/take.sh" <<'TAKE'
|
|
28
|
+
#!/bin/bash
|
|
29
|
+
cd "$(dirname "$0")"
|
|
30
|
+
printf '~/my-app $ '
|
|
31
|
+
sleep 0.8
|
|
32
|
+
cmd='npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dry'
|
|
33
|
+
for (( i=0; i<${#cmd}; i++ )); do printf '%s' "${cmd:$i:1}"; sleep 0.022; done
|
|
34
|
+
sleep 0.6; printf '\n'
|
|
35
|
+
npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dry
|
|
36
|
+
printf '~/my-app $ '
|
|
37
|
+
sleep 2.5
|
|
38
|
+
TAKE
|
|
39
|
+
chmod +x "$WORK/take.sh"
|
|
40
|
+
|
|
41
|
+
asciinema rec --cols 108 --rows 34 --overwrite --command "$WORK/take.sh" "$WORK/demo.cast"
|
|
42
|
+
agg --theme monokai --font-size 16 --speed 1.3 --idle-time-limit 1.2 --last-frame-duration 3 \
|
|
43
|
+
"$WORK/demo.cast" "$OUT"
|
|
44
|
+
|
|
45
|
+
echo "wrote $OUT ($(wc -c < "$OUT") bytes)"
|
|
@@ -273,6 +273,10 @@ async function runHook() {
|
|
|
273
273
|
process.exit(0); // fail-open, always: a miss here is a missing log line, never a blocked turn
|
|
274
274
|
}
|
|
275
275
|
|
|
276
|
+
// Lane names that mean "the main session did it itself". Defaults only: the lane
|
|
277
|
+
// vocabulary belongs to the user's own ROUTING.md, and --inline overrides this.
|
|
278
|
+
const INLINE_LANE_NAMES = ['inline', 'main', 'main inline'];
|
|
279
|
+
|
|
276
280
|
// ---- --summary: a plain-text report, no stdin involved ----
|
|
277
281
|
|
|
278
282
|
function parseLines(text) {
|
|
@@ -313,6 +317,22 @@ function runSummary(args) {
|
|
|
313
317
|
for (const lane of Array.isArray(r.lane) ? r.lane : []) laneCounts.set(lane, (laneCounts.get(lane) || 0) + 1);
|
|
314
318
|
}
|
|
315
319
|
|
|
320
|
+
// The one number that answers "did the routing actually move work off the main
|
|
321
|
+
// session?". Derived only from lane names the marker already carries: no price
|
|
322
|
+
// table, no token count, nothing this log does not hold. A covered turn counts as
|
|
323
|
+
// sent off when its marker names any lane that is not one of the inline names.
|
|
324
|
+
// The inline names are overridable because the lane vocabulary is the user's own
|
|
325
|
+
// ROUTING.md, not a list this package gets to fix.
|
|
326
|
+
const inlineIdx = args.indexOf('--inline');
|
|
327
|
+
const inlineNames = new Set(
|
|
328
|
+
(inlineIdx !== -1 && args[inlineIdx + 1] ? args[inlineIdx + 1].split(',') : INLINE_LANE_NAMES)
|
|
329
|
+
.map((n) => n.trim().toLowerCase())
|
|
330
|
+
.filter(Boolean)
|
|
331
|
+
);
|
|
332
|
+
const coveredRoutes = routes.filter((r) => Array.isArray(r.lane) && r.lane.some((l) => l !== 'missing'));
|
|
333
|
+
const sentOff = coveredRoutes.filter((r) => r.lane.some((l) => l !== 'missing' && !inlineNames.has(String(l).trim().toLowerCase()))).length;
|
|
334
|
+
const sentOffPct = coveredRoutes.length > 0 ? (sentOff / coveredRoutes.length) * 100 : null;
|
|
335
|
+
|
|
316
336
|
const dispatches = records.filter((r) => r.event === 'dispatch');
|
|
317
337
|
const dispatchCounts = new Map();
|
|
318
338
|
for (const r of dispatches) dispatchCounts.set(r.subagent_type, (dispatchCounts.get(r.subagent_type) || 0) + 1);
|
|
@@ -331,6 +351,12 @@ function runSummary(args) {
|
|
|
331
351
|
lines.push('route-metrics summary' + (Number.isNaN(since) ? '' : ' since ' + args[sinceIdx + 1]));
|
|
332
352
|
lines.push('turns: ' + turns);
|
|
333
353
|
lines.push('route-marker coverage: ' + (coveragePct === null ? 'no turns yet' : formatNumber(coveragePct) + '%') + ' (' + covered + '/' + turns + ')');
|
|
354
|
+
lines.push(
|
|
355
|
+
'work sent off the main session: ' +
|
|
356
|
+
(sentOffPct === null ? 'no covered turns yet' : formatNumber(sentOffPct) + '%') +
|
|
357
|
+
' (' + sentOff + '/' + coveredRoutes.length + ' covered turns)'
|
|
358
|
+
);
|
|
359
|
+
lines.push(' inline names for this report: ' + [...inlineNames].join(', ') + ' (override with --inline a,b)');
|
|
334
360
|
lines.push('lanes by count:');
|
|
335
361
|
if (laneCounts.size === 0) lines.push(' (none)');
|
|
336
362
|
for (const [lane, count] of [...laneCounts.entries()].sort((a, b) => b[1] - a[1])) lines.push(' ' + lane + ': ' + count);
|