model-orchestrator 0.1.11 → 0.1.13

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,40 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.13] - 2026-09-08
8
+
9
+ Three issues from a fresh first-run walkthrough of 0.1.12 (#26, #27, #28). Same class as 0.1.12's five: a surface describing an install that did not happen. A fourth, #25, was filed and closed as a mistake on the reporter's side, not a defect: the warning it said was missing has been printed since 0.1.12 and the repro had been read through a truncated pipe.
10
+
11
+ ### Fixed
12
+
13
+ - **The "Then prove it took" list no longer sends a level 1 reader to a file level 1 never wrote** (#27). Step 4 told every reader, at every level, to pick a lane out of `bin/lanes.json` and run `node bin/cli-run.mjs`. Level 1 writes no `bin/` at all, and step 3 immediately above it hedged correctly with "At level 2+" while step 4 did not. The list is now `proofSteps()` in `src/install.js`, gated on level the same way `activationSteps()` is, and the template renders it. Two tests: the README's section must equal the array exactly for every level and primary, and no `bin/` path may appear in it that the plan did not write.
14
+ - **The box setup no longer tells you to sign in to CLIs you did not pick** (#26). `templates/advanced/vm/README.md` step 3 was a fixed sentence naming `codex login --device-auth`, `grok login --device-auth` and `agy`. A level 3 install of claude-code, codex, qwen and ollama was told to sign in to two CLIs it does not have and never told about the one it does. The step now renders each selected CLI's own `auth` string from the catalog. Everything else in that file was already computed from the selection, which is what made the one hardcoded line easy to miss.
15
+ - **A selected local runtime is finally told to install itself** (#26). `activationSteps()` filtered on `kind === 'agent-cli'`, so Ollama, which has a binary and a download page, appeared in no ordered list at any level. Its only mention was one row of a URL table in `DELEGATION_MATRIX.md`. It now gets a step naming the download page and the `ollama pull <model>` that has to follow it.
16
+ - **The tool block stopped saying the same word twice** (#28). Every run that selected a tool printed `optional: Optional. Needs Python 3.10+ and uv.`, because the label repeated the note's own first word. The label is `note:` now. The note keeps the word, because `--list` and the interactive picker print it bare with no label.
17
+
18
+ ### Changed
19
+
20
+ - `--primary` is documented as what it is. `--help` called it "required when several qualify", and then a `--yes` run with several candidates silently picked one in catalog order. The run now names the choice in the plan (`primary claude-code (chosen for you from claude-code, codex; pass --primary to decide it yourself)`) and the help says the same thing. Behaviour is unchanged: the default was sensible, only the promise was wrong.
21
+
22
+ ## [0.1.12] - 2026-09-08
23
+
24
+ Five issues from one first-run walkthrough of 0.1.11 (#20 to #24). Every one of them is the same failure: a page describing an install that did not happen. Each fix removes the second copy of a fact rather than correcting it.
25
+
26
+ ### Fixed
27
+
28
+ - **The generated README no longer describes a different install from the one the terminal just printed** (#20). The activation list existed twice: once as an array built in `bin/cli.js`, once as prose in `templates/common/README.md` that assumed a chat app. A level 2 Claude Code install was told, on the page it was pointed at, to paste `PASTE-INTO-YOUR-AGENT.md`, a file that run never wrote, and a level 1 chat install was told its rules file was `your agent's instructions file`, a leftover placeholder. `activationSteps()` and `snippetFor()` now live in `src/install.js` and both surfaces render the same array, so the page can only ever name the file that was written. A test renders every level against every possible primary and fails if the README omits a printed step or names any other agent's snippet.
29
+ - **A chat install no longer claims a project root it never created** (#21). Level 1 with a chat app writes no project files, and the README still printed `--project` as "where your agent reads rules and subagents" next to "subagent definitions: none". It now says there is no project root and why. A CLI primary that reads a rules file but gets no subagent folder (codex, qwen) keeps its project path and gains the missing half: whether this run created that folder.
30
+ - **The chat activation line is a sentence again** (#22). It read `paste ... into Claude app or claude.ai (chat only, no CLI)'s custom instructions or Project`: the catalog's disambiguating note sat inside a possessive. Chat entries in the catalog now carry `chatName` and `chatSurface`, and the line reads `open the Claude app or claude.ai and paste the block in <path> into its custom instructions or a Project`. The catalog note stays where it is useful, in the picker list.
31
+ - **"Built against" and "pinned to" are one number per lane, by construction** (#23). The README's compatibility table was hand-written and the installer's npm pins were edited separately, so a user comparing them found claude at 2.1.226 and 2.1.260, codex at 0.153.4 and 0.153.2, with no rule for which to trust. `builtAgainst` in `src/catalog.js` is now the single source: `npm run gen:catalog` renders the README table from it, the npm pin **is** that value wherever a lane installs from npm, and the tests fail if the table drifts, if a pin disagrees with its `builtAgainst`, or if a recorded fixture's vendor version disagrees with either. The pins moved to the exercised versions rather than the table moving to the pins, because the exercised version is the one with evidence behind it.
32
+ - **The README stops pinning a release tag the registry has moved past** (#23). The GitHub one-liner still said `#v0.1.7` while npm served 0.1.11. It now points at main, says where the tags are, and a test fails on any `#vX.Y.Z` in the README that is not this package's own version.
33
+ - **The documented example sets both write targets** (#24). The first non-interactive example set `--dir` and left `--project` at the current directory, so a copied command run from a home folder dropped five agent files into it. Both flags are now set in the example, a table explains what lands where and why `--dir` defaults to a folder named for its contents, and the installer prints a line when `--project` was left at the default and subagent files are going there. A test fails on any documented `--yes` example that sets one target and not the other.
34
+ - **Privacy describes what actually runs** (#24). It claimed the only network step was an `npm install -g` you approve, when the thing the user runs is `npx` (a download in itself) and the installer under `--yes` prints vendor install commands without running them. Both are now stated, along with `--yes` selecting the recommended companion tool unless `--no-tools` is passed.
35
+
36
+ ### Changed
37
+
38
+ - Vendor CLI pins move to the versions this release was exercised against: `@anthropic-ai/claude-code@2.1.226`, `@openai/codex@0.153.4`, `@qwen-code/qwen-code@0.22.3`. A pin is a floor, not a ceiling: newer versions may work, and the table exists so a lane that breaks after a vendor upgrade has something to compare against.
39
+ - `npm run gen:catalog` now regenerates two surfaces, `docs/catalog.md` and the README vendor table between its `vendor-table` markers.
40
+
7
41
  ## [0.1.11] - 2026-09-07
8
42
 
9
43
  ### Fixed
@@ -166,7 +200,9 @@ First release.
166
200
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
167
201
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
168
202
 
169
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.7...HEAD
203
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...HEAD
204
+ [0.1.13]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.12...v0.1.13
205
+ [0.1.12]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.11...v0.1.12
170
206
  [0.1.11]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.10...v0.1.11
171
207
  [0.1.10]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.9...v0.1.10
172
208
  [0.1.9]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.8...v0.1.9
package/README.md CHANGED
@@ -10,7 +10,7 @@ Built from a working system, not a diagram: the routing rules, the protocols and
10
10
  npx model-orchestrator
11
11
  ```
12
12
 
13
- That runs the latest published release from the npm registry. To run a specific release or the current main straight from GitHub: `npx github:aunysillyme/model-orchestrator#v0.1.7` (drop `#v0.1.7` for main).
13
+ That runs the latest published release from the npm registry, and `npx model-orchestrator --version` prints which one you got. To run the current main straight from GitHub instead: `npx github:aunysillyme/model-orchestrator`. Add `#vX.Y.Z` for one specific release; the tags are on the [releases page](https://github.com/aunysillyme/model-orchestrator/releases), so this page never pins a number the registry has moved past.
14
14
 
15
15
  The installer asks a few things, then writes a folder:
16
16
 
@@ -56,13 +56,29 @@ An orchestrator routes work. It does not make a model stop guessing numbers, and
56
56
 
57
57
  Whether or not you select them, every level carries the two rules they serve: `protocols/numbers-and-logic.md` (when calling a calculator is mandatory, how to report a computed figure, why a thought log is not evidence) and `protocols/memory-and-record.md` (search before writing, the folder index is part of the change, one writer, inferred content marked as inferred).
58
58
 
59
+ ## The two folders every run writes to
60
+
61
+ An install has two targets, and a scripted run should set both.
62
+
63
+ | Flag | Default | What lands there |
64
+ |---|---|---|
65
+ | `--dir` | `./ai-orchestrator` | the docs, protocols and (level 2+) `bin/cli-run.mjs`. Named after what it contains, not after this package, so a project can hold one without looking like a checkout of it. Pass `--dir ./model-orchestrator` if you prefer the package name. |
66
+ | `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. |
67
+
68
+ `--project` defaulting to the current directory is the one that surprises people: run the command from your home folder with Claude Code as the primary and five agent files land in your home folder. The installer prints the resolved project path in the plan and says when you left it at the default. Set it.
69
+
59
70
  ## Non-interactive
60
71
 
61
72
  ```bash
62
- npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator
73
+ # both targets set: docs in ./ai-orchestrator, subagents into ./my-app/.claude/agents
74
+ npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code \
75
+ --dir ./ai-orchestrator --project ./my-app
76
+ # --yes selects the recommended companion tool (codecalc), which writes CODECALC.md and mcp/ snippets.
77
+ # Add --no-tools for none, or --tools codecalc,obsidian-tc to choose.
78
+
63
79
  npx model-orchestrator --yes --level 3 --ais claude-code,codex,agy,grok,hermes,qwen,ollama --apis anthropic,openrouter --dry # print the plan, write nothing
64
- npx model-orchestrator --yes --level 2 --ais claude-code,codex --project ~/my-app --dir ~/my-app/ai-orchestrator # subagents into ~/my-app/.claude/agents
65
- npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --update-docs # added a lane: regenerate the docs you never edited
80
+ npx model-orchestrator --yes --level 2 --ais claude-code,codex --project ~/my-app --dir ~/my-app/ai-orchestrator --no-tools # subagents into ~/my-app/.claude/agents
81
+ npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --update-docs # added a lane: regenerate the docs you never edited
66
82
  ```
67
83
 
68
84
  ## What gets written (level 3, everything)
@@ -114,17 +130,23 @@ If you need a property in the third row to be enforced, that is a router, a poli
114
130
 
115
131
  The lane wiring and the output judges were written against these versions, which are the ones this release was exercised on:
116
132
 
117
- | Lane | Vendor | Version this release was built against |
118
- |---|---|---|
119
- | `claude` | Anthropic | 2.1.226 |
120
- | `codex` | OpenAI | codex-cli 0.153.4 |
121
- | `agy` | Google | 1.1.27 |
122
- | `grok` | xAI | 1.0.5 |
123
- | `hermes` | Nous Research | 0.20.0 |
124
- | `qwen` | Alibaba | 0.22.3 |
125
- | `ollama` | local | 0.33.3 |
133
+ <!-- vendor-table:start -->
126
134
 
127
- Recorded on 2026-09-06 from the maintainer's own installs, by running each CLI's own version flag. Nothing here is pinned: these CLIs ship breaking flag changes on their own schedules, so a newer version may work perfectly, or may change a flag the generated wiring passes. When a lane starts failing after a vendor upgrade, compare against this table first.
135
+ | Lane | Vendor | Version this release was built against | Where that number is proved |
136
+ |---|---|---|---|
137
+ | `claude` | Anthropic | 2.1.226 | the npm pin the installer writes, `@anthropic-ai/claude-code@2.1.226` |
138
+ | `codex` | OpenAI | 0.153.4 | `test/fixtures/codex-0.153.4.jsonl`, a recorded run |
139
+ | `agy` | Google | 1.1.27 | `test/fixtures/agy-1.1.27.jsonl`, a recorded run |
140
+ | `grok` | xAI | 1.0.5 | `test/fixtures/grok-1.0.5.json`, a recorded run |
141
+ | `hermes` | Nous Research | 0.20.0 | `test/fixtures/hermes-0.20.0.txt`, a recorded run |
142
+ | `qwen` | Alibaba | 0.22.3 | `test/fixtures/qwen-0.22.3-nokey.json`, a recorded run |
143
+ | `ollama` | Ollama | 0.33.3 | the pinned image the level 3 box runs, `ollama/ollama:0.33.3` |
144
+
145
+ Generated from `src/catalog.js` by `npm run gen:catalog`; `npm test` fails if this table and the catalog disagree. Fixtures were captured 2026-09-06.
146
+
147
+ <!-- vendor-table:end -->
148
+
149
+ One number per lane, and it is the same number the installer pins: where a lane installs from npm, `builtAgainst` in the catalog *is* the pin, so "built against" and "pinned to" can never be two answers. That pin is a floor, not a ceiling: these CLIs ship breaking flag changes on their own schedules, so a newer version may work perfectly, or may change a flag the generated wiring passes. When a lane starts failing after a vendor upgrade, compare against this table first.
128
150
 
129
151
  **The live canary runs on your machine, with your credentials.** That is what `node bin/cli-run.mjs --doctor --run` is: it sends every enabled lane one tiny prompt through your own sign-ins and reports `canary ok` or `canary FAILED rc=` per lane. Run it after install, and again after any vendor upgrade.
130
152
 
@@ -144,7 +166,12 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
144
166
 
145
167
  Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows is untested: `cli-run` ends a lane's process tree there with `taskkill`, but nothing in CI runs on Windows, so treat it as unsupported until someone reports otherwise.
146
168
 
147
- **Privacy.** The installer makes no network call of its own and sends no telemetry; the only network activity is the `npm install -g` you approve per package. `cli-run` talks to nothing but the vendor CLI you name.
169
+ **Privacy.** The installer sends no telemetry and makes no network call of its own once it is running. Two things around that are worth being exact about:
170
+
171
+ - `npx model-orchestrator` is itself a download: npm fetches this package from the registry before any of it runs. `npm install -g model-orchestrator` once, then run `model-orchestrator`, if you would rather that happen exactly one time.
172
+ - A missing vendor CLI is *printed*, not installed. In an interactive run the installer offers to run one pinned `npm install -g` per package and only runs the ones you answer yes to; with `--yes` or `--no-install` it answers no for you and prints the command instead. Vendor shell installers (Antigravity, Grok) are only ever printed, alongside the `curl … | less` you would use to read one before running it.
173
+
174
+ `cli-run` talks to nothing but the vendor CLI you name.
148
175
 
149
176
  ## Contributing
150
177
 
package/bin/cli.js CHANGED
@@ -10,7 +10,7 @@ import { spawnSync } from 'node:child_process';
10
10
  import { resolve, join } from 'node:path';
11
11
  import { which } from '../src/detect.js';
12
12
  import { AIS, LEVELS, TOOLS, PROVIDERS, aisForLevel, agentCandidates, byId, npmSpec } from '../src/catalog.js';
13
- import { planFiles, writeFiles, resolveSelection, resolveTools, resolveApis, dirProblems, readManifest, MACHINE_OWNED, RUNTIME, GENERATOR_VERSION } from '../src/install.js';
13
+ import { planFiles, writeFiles, resolveSelection, resolveTools, resolveApis, dirProblems, readManifest, activationSteps, MACHINE_OWNED, RUNTIME, GENERATOR_VERSION } from '../src/install.js';
14
14
 
15
15
  // One strict parse. Unknown flags, missing values and duplicates are usage
16
16
  // errors (exit 2) before anything is planned, so a typo like --dryy can never
@@ -88,17 +88,19 @@ if (flag('help') || flag('h')) {
88
88
  Usage
89
89
  npx model-orchestrator interactive
90
90
  npx model-orchestrator --list show the AI catalog and exit
91
- npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok [--primary claude-code] [--dir ./ai-orchestrator]
91
+ npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok [--primary claude-code] [--dir ./ai-orchestrator] [--project .]
92
92
 
93
93
  Flags
94
94
  --level 1|2|3 1 beginner (one agent), 2 intermediate (many CLIs), 3 advanced (plus a VM)
95
95
  --ais a,b,c catalog ids you have access to (see --list)
96
- --primary id the agent that runs the system and receives the subagents (any level; required when several qualify)
96
+ --primary id the agent that runs the system and receives the subagents (any level). When several qualify
97
+ and --yes is set, the run picks one and says so in the plan; pass this to decide it yourself.
97
98
  --tools a,b companion tools to set up, all optional (default with --yes: codecalc only); --no-tools for none
98
99
  --apis a,b level 3 only: metered API keys you HOLD (anthropic,openai,google,xai,openrouter); --no-apis for none.
99
100
  Asked separately from the CLIs because a subscription is not an API key.
100
101
  --dir path where to write the docs and protocols (default ./ai-orchestrator)
101
- --project path the project root your agent runs from; subagent definitions go here (default: current directory)
102
+ --project path the project root your agent runs from; subagent definitions go here (default: current directory,
103
+ so set it: a run from your home folder otherwise drops the subagent files there)
102
104
  --yes skip confirmations
103
105
  --force overwrite every file that already exists, documents included
104
106
  --upgrade-runtime replace the runtime files (cli-run, the audit job, compose, gateway config, setup script) even
@@ -188,6 +190,7 @@ async function main() {
188
190
  // 3. Primary agent (the one that runs the system)
189
191
  const candidates = agentCandidates(selected);
190
192
  let primary = null;
193
+ let primaryAutoPicked = false;
191
194
  if (opt('primary')) {
192
195
  primary = byId[opt('primary')];
193
196
  if (!primary || !candidates.includes(primary)) bad('--primary must be one of: ' + candidates.map((a) => a.id).join(', '));
@@ -199,7 +202,10 @@ async function main() {
199
202
  // --yes picks for the user: claude-code if present, else the first agent that can load subagent
200
203
  // definitions (it gets five files written for it), else the first candidate. #19: codex listed
201
204
  // before agy used to win and nothing was written to the project root.
202
- if (yes) primary = candidates.find((a) => a.id === 'claude-code') || candidates.find((a) => a.agentsDir) || candidates[0];
205
+ if (yes) {
206
+ primary = candidates.find((a) => a.id === 'claude-code') || candidates.find((a) => a.agentsDir) || candidates[0];
207
+ primaryAutoPicked = true;
208
+ }
203
209
  else {
204
210
  console.log('\nWhich one is your primary agent (the one that runs the system)?');
205
211
  candidates.forEach((a, i) => console.log(` ${i + 1} ${a.name}`));
@@ -270,7 +276,10 @@ async function main() {
270
276
  const files = planFiles({ level, selected, primary, dir, project, tools, apis });
271
277
  const lvl = LEVELS.find((l) => l.id === level);
272
278
  const agentFiles = files.filter((f) => f.root === 'project');
273
- console.log(`\nPlan\n level ${lvl.id} ${lvl.name}\n access ${selected.map((a) => a.id).join(', ')}\n primary ${primary ? primary.id : 'none'}\n tools ${tools.map((t) => t.id).join(', ') || 'none'}` + (level >= 3 ? `\n api keys ${apis.map((p) => p.id).join(', ') || 'none'}` : '') + `\n folder ${dir}\n project ${project}${agentFiles.length ? ' (' + agentFiles.length + ' subagent files go here)' : ''}\n files ${files.length}`);
279
+ console.log(`\nPlan\n level ${lvl.id} ${lvl.name}\n access ${selected.map((a) => a.id).join(', ')}\n primary ${primary ? primary.id : 'none'}${primaryAutoPicked ? ` (chosen for you from ${candidates.map((a) => a.id).join(', ')}; pass --primary to decide it yourself)` : ''}\n tools ${tools.map((t) => t.id).join(', ') || 'none'}` + (level >= 3 ? `\n api keys ${apis.map((p) => p.id).join(', ') || 'none'}` : '') + `\n folder ${dir}\n project ${project}${agentFiles.length ? ' (' + agentFiles.length + ' subagent files go here)' : ''}\n files ${files.length}`);
280
+ if (agentFiles.length && !opt('project')) {
281
+ console.log(`\nNote: --project was not given, so the ${agentFiles.length} subagent file(s) go to the current directory (${project}). Pass --project to put them somewhere else.`);
282
+ }
274
283
  if (level >= 2 && !selected.some((a) => a.cliRun)) {
275
284
  console.log('\nWarning: no executable lanes selected; delegation is inactive. Use level 1 for a single-agent setup, or add a supported CLI. Doctor will exit 13 until a lane is enabled.');
276
285
  }
@@ -361,20 +370,13 @@ async function main() {
361
370
 
362
371
  for (const t of tools) {
363
372
  const doc = t.id.toUpperCase() + '.md';
364
- console.log(`\n${t.name}\n optional: ${t.optionalNote}\n needs: ${t.requires}\n run: ${t.install}\n one-click or self-registering for: ${t.autoClients.join(', ')}. Other agents and the details: ${dir}/${doc}`);
373
+ console.log(`\n${t.name}\n note: ${t.optionalNote}\n needs: ${t.requires}\n run: ${t.install}\n one-click or self-registering for: ${t.autoClients.join(', ')}. Other agents and the details: ${dir}/${doc}`);
365
374
  }
366
375
 
367
376
  // 7. Activation summary: writing the folder is half the job. Say exactly what
368
- // turns it on, in order, with one command that proves it.
369
- const snippet = files.find((f) => /\.snippet\.md$|^PASTE-INTO-YOUR-AGENT\.md$/.test(f.rel));
370
- const steps = [];
371
- if (snippet && primary && primary.rulesFile) steps.push(`copy the block in ${join(dir, snippet.rel)} into ${join(project, primary.rulesFile)} (create it if missing)`);
372
- else if (snippet) steps.push(`paste ${join(dir, snippet.rel)} into ${primary.name}'s custom instructions or Project`);
373
- if (primary && primary.agentsDir) steps.push(`subagents are in ${join(project, primary.agentsDir)}; run ${primary.bin} from ${project} to pick them up`);
374
- for (const a of selected.filter((a) => a.bin && a.kind === 'agent-cli')) steps.push(`sign in to ${a.name}: ${a.auth}`);
375
- for (const t of tools) steps.push(`${t.id}: ${t.install}`);
376
- if (level >= 2) steps.push(`smoke test: node ${join(dir, 'bin', 'cli-run.mjs')} --doctor (add --run to send each lane one tiny prompt)`);
377
- if (level >= 3) steps.push(`box: read ${join(dir, 'vm', 'README.md')}; keys named in vm/ENVIRONMENT.md go in your secrets manager, never a file`);
377
+ // turns it on, in order, with one command that proves it. The generated
378
+ // README renders this same array, so the two surfaces cannot disagree (#20).
379
+ const steps = activationSteps({ level, selected, primary, tools, dir, project });
378
380
  console.log('\nTo activate, in order:');
379
381
  steps.forEach((st, i) => console.log(` ${i + 1}. ${st}`));
380
382
  console.log(`\nStart here: ${join(dir, 'README.md')} (written for level ${level} and the AIs you picked).`);
package/docs/catalog.md CHANGED
@@ -16,18 +16,20 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
16
16
 
17
17
  - **Kind:** agent-cli · **Access:** subscription · **Lane:** A · **Level:** 1+
18
18
  - **Wins at:** orchestrator: routes, maps, builds, verifies, records
19
- - **Install:** `npm install -g @anthropic-ai/claude-code@2.1.260`
19
+ - **Install:** `npm install -g @anthropic-ai/claude-code@2.1.226`
20
20
  - **Sign in:** run `claude` once and sign in with your Anthropic account
21
21
  - **Reads rules from:** `CLAUDE.md` · subagents in `.claude/agents/`
22
+ - **Built against:** 2.1.226 (the same number the npm pin uses)
22
23
 
23
24
  ### `codex` · Codex CLI (OpenAI, ChatGPT plan)
24
25
 
25
26
  - **Kind:** agent-cli · **Access:** subscription · **Lane:** A · **Level:** 1+
26
27
  - **Wins at:** second coder and adversarial auditor (a different model family reading your diff)
27
- - **Install:** `npm install -g @openai/codex@0.153.2`
28
+ - **Install:** `npm install -g @openai/codex@0.153.4`
28
29
  - **Sign in:** `codex login` (add `--device-auth` on a machine with no browser)
29
30
  - **Reads rules from:** `AGENTS.md`
30
31
  - **cli-run lane:** yes
32
+ - **Built against:** 0.153.4 (the same number the npm pin uses)
31
33
 
32
34
  ### `agy` · Antigravity CLI `agy` (Google AI plan)
33
35
 
@@ -37,6 +39,7 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
37
39
  - **Sign in:** first run opens a device-code sign-in with your Google account
38
40
  - **Reads rules from:** `GEMINI.md` · subagents in `.agents/agents/`
39
41
  - **cli-run lane:** yes
42
+ - **Built against:** 1.1.27
40
43
  - **Note:** Gemini CLI was retired by Google in June 2026. agy is the successor. Do not install `gemini`.
41
44
 
42
45
  ### `grok` · Grok CLI (xAI, X Premium)
@@ -46,6 +49,7 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
46
49
  - **Install:** vendor script (read it first): `https://x.ai/cli/install.sh`
47
50
  - **Sign in:** `grok login` (add `--device-auth` on a headless machine)
48
51
  - **cli-run lane:** yes
52
+ - **Built against:** 1.0.5
49
53
 
50
54
  ### `hermes` · Hermes Agent (Nous Research)
51
55
 
@@ -54,15 +58,17 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
54
58
  - **Install:** https://github.com/NousResearch/hermes-agent
55
59
  - **Sign in:** `hermes auth add <provider>` per provider; its own fallback chain handles outages
56
60
  - **cli-run lane:** yes
61
+ - **Built against:** 0.20.0
57
62
 
58
63
  ### `qwen` · Qwen Code CLI (Alibaba, provider-agnostic)
59
64
 
60
65
  - **Kind:** agent-cli · **Access:** metered · **Lane:** B · **Level:** 2+
61
66
  - **Wins at:** cheapest metered bulk lane for structured output; never for anything that cites a line, a number or a source
62
- - **Install:** `npm install -g @qwen-code/qwen-code@0.23.0`
67
+ - **Install:** `npm install -g @qwen-code/qwen-code@0.22.3`
63
68
  - **Sign in:** a provider key in an environment variable, named (not stored) in ~/.qwen/settings.json. There is no free Qwen cloud tier any more.
64
69
  - **Reads rules from:** `QWEN.md`
65
70
  - **cli-run lane:** yes
71
+ - **Built against:** 0.22.3 (the same number the npm pin uses)
66
72
  - **Note:** Its own success flags lie on API failures. cli-run checks the two honest signals for you.
67
73
 
68
74
  ### `ollama` · Ollama (local models)
@@ -71,6 +77,7 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
71
77
  - **Wins at:** the privacy lane: anything that must never leave the machine. Not a cost lane.
72
78
  - **Install:** https://ollama.com/download (or `brew install ollama`)
73
79
  - **Sign in:** none
80
+ - **Built against:** 0.33.3
74
81
 
75
82
  ### `claude-app` · Claude app or claude.ai (chat only, no CLI)
76
83
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.11",
3
+ "version": "0.1.13",
4
4
  "description": "Routing instructions and a CLI runner for your AI tools. One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine.",
5
5
  "type": "module",
6
6
  "bin": {
@@ -46,7 +46,12 @@
46
46
  "ollama",
47
47
  "litellm",
48
48
  "multi-agent",
49
- "cli"
49
+ "cli",
50
+ "llm",
51
+ "agents",
52
+ "subagents",
53
+ "qwen",
54
+ "mcp"
50
55
  ],
51
56
  "author": "aunysillyme (https://github.com/aunysillyme)",
52
57
  "license": "MIT"
package/scripts/README.md CHANGED
@@ -2,4 +2,4 @@
2
2
 
3
3
  | File | Job |
4
4
  |---|---|
5
- | `gen-catalog.js` | regenerates `docs/catalog.md` from `src/catalog.js`; `npm run gen:catalog`. `test/catalog.test.js` fails if the two disagree. |
5
+ | `gen-catalog.js` | regenerates `docs/catalog.md` AND the vendor compatibility table in `README.md` (between the `vendor-table` markers) from `src/catalog.js`; `npm run gen:catalog`. `test/catalog.test.js` fails if either generated surface disagrees with the catalog. |
@@ -1,7 +1,8 @@
1
1
  #!/usr/bin/env node
2
- // Regenerates docs/catalog.md from src/catalog.js. The test suite checks they agree.
3
- import { writeFileSync } from 'node:fs';
4
- import { AIS, LEVELS, TOOLS, npmSpec } from '../src/catalog.js';
2
+ // Regenerates docs/catalog.md and the README's vendor compatibility table from
3
+ // src/catalog.js. The test suite checks they agree, so neither can drift (#23).
4
+ import { writeFileSync, readFileSync } from 'node:fs';
5
+ import { AIS, LEVELS, TOOLS, IMAGES, npmSpec } from '../src/catalog.js';
5
6
  import { readdirSync } from 'node:fs';
6
7
 
7
8
  export function protocolCount() {
@@ -21,6 +22,7 @@ export function catalogMarkdown() {
21
22
  md += `### \`${a.id}\` · ${a.name}\n\n- **Kind:** ${a.kind} · **Access:** ${a.access} · **Lane:** ${a.lane} · **Level:** ${a.minLevel}+\n- **Wins at:** ${a.role}\n- **Install:** ${how}\n- **Sign in:** ${a.auth}\n`;
22
23
  if (a.rulesFile) md += `- **Reads rules from:** \`${a.rulesFile}\`` + (a.agentsDir ? ` · subagents in \`${a.agentsDir}/\`` : '') + '\n';
23
24
  if (a.cliRun) md += '- **cli-run lane:** yes\n';
25
+ if (a.builtAgainst) md += `- **Built against:** ${a.builtAgainst}` + (a.install.npm ? ' (the same number the npm pin uses)' : '') + '\n';
24
26
  if (a.note) md += `- **Note:** ${a.note}\n`;
25
27
  md += '\n';
26
28
  }
@@ -31,7 +33,45 @@ export function catalogMarkdown() {
31
33
  return md;
32
34
  }
33
35
 
36
+ export const VENDOR_TABLE_START = '<!-- vendor-table:start -->';
37
+ export const VENDOR_TABLE_END = '<!-- vendor-table:end -->';
38
+
39
+ export function fixtureManifest() {
40
+ return JSON.parse(readFileSync(new URL('../test/fixtures/manifest.json', import.meta.url), 'utf8'));
41
+ }
42
+
43
+ // The compatibility table in README.md. Every number comes from `builtAgainst`
44
+ // in src/catalog.js, which is also the npm pin where there is one, so "built
45
+ // against" and "pinned to" are the same number by construction (#23).
46
+ export function vendorTableMarkdown() {
47
+ const fx = fixtureManifest();
48
+ const byLane = Object.fromEntries(fx.fixtures.map((f) => [f.lane, f]));
49
+ let md = VENDOR_TABLE_START + '\n\n| Lane | Vendor | Version this release was built against | Where that number is proved |\n|---|---|---|---|\n';
50
+ for (const a of AIS) {
51
+ if (!a.bin || !a.builtAgainst) continue;
52
+ const proof = byLane[a.id]
53
+ ? '`test/fixtures/' + byLane[a.id].file + '`, a recorded run'
54
+ : a.install.npm
55
+ ? 'the npm pin the installer writes, `' + npmSpec(a) + '`'
56
+ : a.id === 'ollama'
57
+ ? 'the pinned image the level 3 box runs, `' + IMAGES.ollama + '`'
58
+ : "the maintainer's own install";
59
+ md += `| \`${a.bin}\` | ${a.vendor} | ${a.builtAgainst} | ${proof} |\n`;
60
+ }
61
+ md += `\nGenerated from \`src/catalog.js\` by \`npm run gen:catalog\`; \`npm test\` fails if this table and the catalog disagree. Fixtures were captured ${fx.capturedAt}.\n\n` + VENDOR_TABLE_END;
62
+ return md;
63
+ }
64
+
65
+ export function readmeWithVendorTable(current) {
66
+ const a = current.indexOf(VENDOR_TABLE_START);
67
+ const b = current.indexOf(VENDOR_TABLE_END);
68
+ if (a === -1 || b === -1) throw new Error('README.md has no vendor-table markers');
69
+ return current.slice(0, a) + vendorTableMarkdown() + current.slice(b + VENDOR_TABLE_END.length);
70
+ }
71
+
34
72
  if (process.argv[1] && process.argv[1].endsWith('gen-catalog.js')) {
35
73
  writeFileSync(new URL('../docs/catalog.md', import.meta.url), catalogMarkdown());
36
- console.log('docs/catalog.md regenerated');
74
+ const readmeUrl = new URL('../README.md', import.meta.url);
75
+ writeFileSync(readmeUrl, readmeWithVendorTable(readFileSync(readmeUrl, 'utf8')));
76
+ console.log('docs/catalog.md and the README vendor table regenerated');
37
77
  }
package/src/README.md CHANGED
@@ -4,6 +4,6 @@
4
4
  |---|---|
5
5
  | `catalog.js` | the single list of levels and AIs. Add an AI here and the prompts, docs tables, delegation matrix, gateway config and installer all pick it up. Nothing else lists AIs. |
6
6
  | `detect.js` | PATH lookup for a binary, plus the few places vendor installers drop binaries without touching PATH. No shell-outs. |
7
- | `install.js` | pure planner: turns (level, selection, primary) into a list of files to write, rendering templates and computing every generated table. `writeFiles` is the only thing that touches disk. |
7
+ | `install.js` | pure planner: turns (level, selection, primary) into a list of files to write, rendering templates and computing every generated table. `writeFiles` is the only thing that touches disk. `activationSteps()` and `snippetFor()` live here so the terminal summary and the generated README render the same list. |
8
8
  | `prompt.js` | line-buffered questions for the interactive path; piped answers are queued, EOF mid-prompt aborts instead of confirming a write. |
9
9
  | `render.js` | `{{KEY}}` substitution. Throws on an unknown key, so a template typo fails the test suite instead of shipping a literal placeholder. |
package/src/catalog.js CHANGED
@@ -17,6 +17,16 @@
17
17
  // rulesFile the instructions file that agent reads from a project root, if any
18
18
  // agentsDir where that agent keeps project-level subagent definitions, if any
19
19
  // cliRun true when bin/cli-run.mjs has a judge for this lane
20
+ // builtAgainst the vendor version this release's lane wiring and judges were
21
+ // exercised against. ONE number per lane: the README compatibility
22
+ // table is generated from it, and where install.npm exists the pin
23
+ // IS this number, so "built against" and "pinned to" cannot drift
24
+ // into two answers (#23). Lanes with a recorded fixture are
25
+ // cross-checked against test/fixtures/manifest.json by the tests.
26
+ // chatName chat apps only: the app's name in a sentence, without the
27
+ // "(chat only, no CLI)" catalog note, so the activation line stays
28
+ // a sentence you can read once (#22)
29
+ // chatSurface chat apps only: where the pasted block goes in that app
20
30
 
21
31
  export const LEVELS = [
22
32
  {
@@ -53,7 +63,8 @@ export const AIS = [
53
63
  lane: 'A',
54
64
  role: 'orchestrator: routes, maps, builds, verifies, records',
55
65
  minLevel: 1,
56
- install: { npm: '@anthropic-ai/claude-code', pin: '2.1.260' },
66
+ install: { npm: '@anthropic-ai/claude-code', pin: '2.1.226' },
67
+ builtAgainst: '2.1.226',
57
68
  auth: 'run `claude` once and sign in with your Anthropic account',
58
69
  rulesFile: 'CLAUDE.md',
59
70
  agentsDir: '.claude/agents',
@@ -70,7 +81,8 @@ export const AIS = [
70
81
  lane: 'A',
71
82
  role: 'second coder and adversarial auditor (a different model family reading your diff)',
72
83
  minLevel: 1,
73
- install: { npm: '@openai/codex', pin: '0.153.2' },
84
+ install: { npm: '@openai/codex', pin: '0.153.4' },
85
+ builtAgainst: '0.153.4',
74
86
  auth: '`codex login` (add `--device-auth` on a machine with no browser)',
75
87
  rulesFile: 'AGENTS.md',
76
88
  agentsDir: null,
@@ -87,6 +99,7 @@ export const AIS = [
87
99
  role: 'deep research sweeps and concurrent fan-out (its subagent call takes an array)',
88
100
  minLevel: 1,
89
101
  install: { script: 'https://antigravity.google/cli/install.sh' },
102
+ builtAgainst: '1.1.27',
90
103
  auth: 'first run opens a device-code sign-in with your Google account',
91
104
  rulesFile: 'GEMINI.md',
92
105
  agentsDir: '.agents/agents',
@@ -105,6 +118,7 @@ export const AIS = [
105
118
  role: 'X and live web reads at no per-call cost (its search tools bill on the API, not on the CLI)',
106
119
  minLevel: 1,
107
120
  install: { script: 'https://x.ai/cli/install.sh' },
121
+ builtAgainst: '1.0.5',
108
122
  auth: '`grok login` (add `--device-auth` on a headless machine)',
109
123
  rulesFile: null,
110
124
  agentsDir: null,
@@ -121,6 +135,7 @@ export const AIS = [
121
135
  role: 'the free tier: rough drafts, first-pass summaries, cheap divergent reads, cron jobs on a box',
122
136
  minLevel: 2,
123
137
  install: { url: 'https://github.com/NousResearch/hermes-agent' },
138
+ builtAgainst: '0.20.0',
124
139
  auth: '`hermes auth add <provider>` per provider; its own fallback chain handles outages',
125
140
  rulesFile: null,
126
141
  agentsDir: null,
@@ -136,7 +151,8 @@ export const AIS = [
136
151
  lane: 'B',
137
152
  role: 'cheapest metered bulk lane for structured output; never for anything that cites a line, a number or a source',
138
153
  minLevel: 2,
139
- install: { npm: '@qwen-code/qwen-code', pin: '0.23.0' },
154
+ install: { npm: '@qwen-code/qwen-code', pin: '0.22.3' },
155
+ builtAgainst: '0.22.3',
140
156
  auth: 'a provider key in an environment variable, named (not stored) in ~/.qwen/settings.json. There is no free Qwen cloud tier any more.',
141
157
  rulesFile: 'QWEN.md',
142
158
  agentsDir: null,
@@ -154,6 +170,7 @@ export const AIS = [
154
170
  role: 'the privacy lane: anything that must never leave the machine. Not a cost lane.',
155
171
  minLevel: 2,
156
172
  install: { url: 'https://ollama.com/download', brew: 'ollama' },
173
+ builtAgainst: '0.33.3',
157
174
  auth: 'none',
158
175
  rulesFile: null,
159
176
  agentsDir: null,
@@ -171,6 +188,8 @@ export const AIS = [
171
188
  minLevel: 1,
172
189
  install: { url: 'https://claude.ai' },
173
190
  auth: 'sign in',
191
+ chatName: 'the Claude app or claude.ai',
192
+ chatSurface: 'custom instructions or a Project',
174
193
  rulesFile: null,
175
194
  agentsDir: null,
176
195
  cliRun: false
@@ -187,6 +206,8 @@ export const AIS = [
187
206
  minLevel: 1,
188
207
  install: { url: 'https://chatgpt.com' },
189
208
  auth: 'sign in',
209
+ chatName: 'ChatGPT',
210
+ chatSurface: 'custom instructions or a Project',
190
211
  rulesFile: null,
191
212
  agentsDir: null,
192
213
  cliRun: false
@@ -203,6 +224,8 @@ export const AIS = [
203
224
  minLevel: 1,
204
225
  install: { url: 'https://gemini.google.com' },
205
226
  auth: 'sign in',
227
+ chatName: 'the Gemini app',
228
+ chatSurface: 'saved instructions or a Gem',
206
229
  rulesFile: null,
207
230
  agentsDir: null,
208
231
  cliRun: false
package/src/install.js CHANGED
@@ -193,6 +193,57 @@ export function laneVars(selected) {
193
193
  };
194
194
  }
195
195
 
196
+ // Which activation file this primary gets. ONE decision, read by three
197
+ // surfaces: planFiles writes the file, vars() names it in the generated README,
198
+ // and bin/cli.js prints it in the terminal. Before 0.1.12 the README hardcoded
199
+ // PASTE-INTO-YOUR-AGENT.md and named it for CLI installs that never got one (#20).
200
+ export function snippetFor(primary) {
201
+ if (!primary) return null;
202
+ return primary.rulesFile ? primary.rulesFile.replace(/\.md$/, '.snippet.md') : 'PASTE-INTO-YOUR-AGENT.md';
203
+ }
204
+
205
+ // The activation list, in order. The terminal prints this array at the end of a
206
+ // run and the generated README renders the same array, so the page cannot
207
+ // describe a different first step from the one the user just read (#20).
208
+ export function activationSteps(opts) {
209
+ const { level, selected = [], primary } = opts;
210
+ const tools = opts.tools || [];
211
+ const dirAbs = resolve(opts.dir || 'ai-orchestrator');
212
+ const projectAbs = resolve(opts.project || process.cwd());
213
+ const snippet = snippetFor(primary);
214
+ const steps = [];
215
+ if (snippet && primary.rulesFile) steps.push(`copy the block in ${join(dirAbs, snippet)} into ${join(projectAbs, primary.rulesFile)} (create it if missing)`);
216
+ // A chat app has no possessive that survives its catalog note: "Claude app or
217
+ // claude.ai (chat only, no CLI)'s custom instructions" was the sentence this
218
+ // replaces (#22).
219
+ else if (snippet) steps.push(`open ${primary.chatName || primary.name} and paste the block in ${join(dirAbs, snippet)} into its ${primary.chatSurface || 'custom instructions'}`);
220
+ if (primary && primary.agentsDir) steps.push(`subagents are in ${join(projectAbs, primary.agentsDir)}; run ${primary.bin} from ${projectAbs} to pick them up`);
221
+ for (const a of selected.filter((a) => a.bin && a.kind === 'agent-cli')) steps.push(`sign in to ${a.name}: ${a.auth}`);
222
+ // A local runtime has a bin but no sign-in, so the agent-cli loop above skips it
223
+ // and before this it appeared in no ordered list at any level (#26).
224
+ for (const a of selected.filter((a) => a.bin && a.kind === 'local')) steps.push(`install ${a.name}: ${a.install.url}, then \`${a.bin} pull <model>\` before the local lane can answer`);
225
+ for (const t of tools) steps.push(`${t.id}: ${t.install}`);
226
+ if (level >= 2) steps.push(`smoke test: node ${join(dirAbs, 'bin', 'cli-run.mjs')} --doctor (add --run to send each lane one tiny prompt)`);
227
+ if (level >= 3) steps.push(`box: read ${join(dirAbs, 'vm', 'README.md')}; keys named in vm/ENVIRONMENT.md go in your secrets manager, never a file`);
228
+ return steps;
229
+ }
230
+
231
+ // The verification list, in order. Gated on level for the same reason
232
+ // activationSteps is: level 1 writes no bin/, so a step naming cli-run.mjs or
233
+ // lanes.json there described an install that did not happen (#27).
234
+ export function proofSteps(opts) {
235
+ const { level } = opts;
236
+ const steps = [
237
+ 'Start a fresh agent session and ask: "Read the orchestrator instructions. Quote the routing rule you will use, then sort pear, apple, banana alphabetically. Name the tier and whether you delegated."',
238
+ 'Expect the fast tier and `apple, banana, pear`. If the agent cannot quote the routing rule, check the snippet location or chat instructions before continuing. This is a manual activation check, not proof that every future task follows the rules.'
239
+ ];
240
+ if (level >= 2) {
241
+ steps.push('Run `node bin/cli-run.mjs --doctor` from this folder. It checks binary presence, not authentication or loaded instructions. `--doctor --run` additionally uses a little quota to test live responses. No enabled lanes means delegation is inactive.');
242
+ steps.push('To test a real output contract, choose an enabled lane from `bin/lanes.json` and run `node bin/cli-run.mjs <lane> \'Return only {"sorted":["apple","banana","pear"]}\' --expect-json`. This uses quota. Expect JSON and exit 0; inspect the array yourself. A non-JSON response exits 10, a missing binary exits 13, and an authentication failure reports the vendor error. The explicit lane tests execution; your primary agent still makes delegation decisions.');
243
+ }
244
+ return steps;
245
+ }
246
+
196
247
  function vars(opts) {
197
248
  const { level, selected, primary } = opts;
198
249
  const tools = opts.tools || [];
@@ -206,8 +257,30 @@ function vars(opts) {
206
257
  if (rulesPath === '') rulesPath = '.';
207
258
  else if (rulesPath.startsWith('..')) rulesPath = dirAbs; // outside the project: absolute is the only honest path
208
259
  const pinOf = (id) => (toolById[id] && toolById[id].pin) || 'latest';
260
+ const snippet = snippetFor(primary);
261
+ const steps = activationSteps({ level, selected, primary, tools, dir: opts.dir, project: opts.project });
262
+ const proofs = proofSteps({ level });
263
+ // Only claude-code and agy put files under the project root. A chat primary
264
+ // puts nothing there, so naming a project root would name a folder this run
265
+ // never created (#21).
266
+ const writesProject = !!(primary && primary.agentsDir);
267
+ const readsProjectRules = !!(primary && primary.rulesFile);
268
+ const whereThingsWent = [`- This folder: \`${dirAbs}\``];
269
+ if (writesProject) whereThingsWent.push(`- Project root (where your agent reads rules and subagents): \`${projectAbs}\``, `- Subagent definitions: \`${join(projectAbs, primary.agentsDir)}\``);
270
+ else if (readsProjectRules) whereThingsWent.push(`- Project root (where ${primary.name} reads \`${primary.rulesFile}\`): \`${projectAbs}\`` + (existsSync(projectAbs) ? '' : ' (this run wrote nothing there; create the folder before you copy the snippet in)'), '- Subagent definitions: none, this agent has no subagent folder');
271
+ else whereThingsWent.push('- Project root: none. A chat app reads pasted instructions, not files, so this install wrote nothing to a project folder.', '- Subagent definitions: none');
272
+ whereThingsWent.push(`- The rules path your snippets use: \`${rulesPath}\``);
209
273
  return {
210
274
  ...laneVars(selected),
275
+ ACTIVATION_STEPS: steps.map((st, i) => `${i + 1}. ${st}`).join('\n'),
276
+ PROOF_STEPS: proofs.map((st, i) => `${i + 1}. ${st}`).join('\n'),
277
+ LOAD_IT: readsProjectRules
278
+ ? `${primary.name} reads its rules from \`${primary.rulesFile}\` in the project root. The installer wrote \`${snippet}\` next to this README; copy its contents into \`${join(projectAbs, primary.rulesFile)}\`, creating that file if it does not exist. Nothing was appended to a file you already had.`
279
+ : snippet
280
+ ? `${primary.name} has no project rules file, so the rules travel by paste. The installer wrote \`${snippet}\` next to this README; open ${primary.chatName || primary.name} and paste its contents into ${primary.chatSurface || 'custom instructions'}. Nothing was appended to a file you already had.`
281
+ : 'No primary agent was selected, so no activation file was written. Re-run the installer and pick one.',
282
+ CHAT_UPLOAD_NOTE: primary && primary.kind === 'chat' ? ' A chat app cannot open a local path: upload or paste any protocol file you want it to read.' : '',
283
+ WHERE_THINGS_WENT: whereThingsWent.join('\n'),
211
284
  RULES_PATH: rulesPath,
212
285
  ROUTING_FILE: level >= 2 ? 'ROUTING.md' : 'ORCHESTRATOR.md',
213
286
  PROJECT_DIR: projectAbs,
@@ -220,6 +293,12 @@ function vars(opts) {
220
293
  INSTALL_DIR: dirAbs,
221
294
  INSTALL_DIR_SH: shellQuote(dirAbs),
222
295
  INSTALL_DIR_SYSTEMD: systemdEscape(dirAbs),
296
+ // vm/README.md step 3 named `grok login` and `agy` whatever you picked (#26).
297
+ VM_SIGNIN: (() => {
298
+ const lines = selected.filter((a) => a.bin && a.kind === 'agent-cli').map((a) => ` - ${a.name}: ${a.auth}`);
299
+ for (const a of selected.filter((a) => a.bin && a.kind === 'local')) lines.push(` - ${a.name}: no sign-in. Install it from ${a.install.url}, then \`${a.bin} pull <model>\`.`);
300
+ return lines.length ? lines.join('\n') : ' - none: no CLI you selected needs a sign-in on the box.';
301
+ })(),
223
302
  AUDIT_LANE: lane || 'none',
224
303
  // Enforced boundary per lane: codex has a read-only sandbox flag; the others
225
304
  // run with whatever their own config allows, and the script says so.
@@ -23,7 +23,8 @@ Generated {{DATE}} for: `{{AI_IDS}}`. Installed at `{{INSTALL_DIR}}`; the system
23
23
 
24
24
  1. Provision a box. Ubuntu, 2+ vCPU, 8 GB is comfortable. Put it on a private mesh network if you can; do not open ports to the internet.
25
25
  2. `bash setup-vm.sh`. It installs system deps and the npm-installable CLIs, then **prints** the vendor shell installers for the rest. Read those scripts before running them.
26
- 3. Sign each CLI in with its device-code flow (`codex login --device-auth`, `grok login --device-auth`, `agy` on first run). Run these inside `tmux` so a dropped SSH session does not kill the prompt. Headless Linux has no keyring by default; `setup-vm.sh` installs one so the CLIs stop re-prompting.
26
+ 3. Sign each CLI in, using the flow its vendor gives you. Run these inside `tmux` so a dropped SSH session does not kill the prompt. Headless Linux has no keyring by default; `setup-vm.sh` installs one so the CLIs stop re-prompting.
27
+ {{VM_SIGNIN}}
27
28
  4. Put provider keys in your secrets manager and export the names listed in `ENVIRONMENT.md` into the gateway's environment at start time. Never write a value into a file in this folder. The gateway config was rendered from the API keys you said you hold, not from your CLI subscriptions: those are different entitlements.
28
29
  5. `docker compose up -d`, then list the lanes without putting the key in argv (the key must be a single token, `^[A-Za-z0-9._-]+$`, because it is interpolated into curl's config grammar):
29
30
  ```bash
@@ -13,13 +13,15 @@ Companion tools:
13
13
 
14
14
  This folder gives your agent routing instructions and, at level 2+, a runner for explicitly selected CLI lanes. The agent chooses the tier or lane; the runner does not automatically compare prices or choose a model.
15
15
 
16
- ## Your first task
16
+ ## Activate it
17
17
 
18
- 1. Activate the snippet using the instructions below. For a chat app, paste the block in `PASTE-INTO-YOUR-AGENT.md`; upload any full protocols you want it to read because local paths alone do not share files.
19
- 2. Start a fresh agent session and ask: "Read the orchestrator instructions. Quote the routing rule you will use, then sort pear, apple, banana alphabetically. Name the tier and whether you delegated."
20
- 3. Expect the fast tier and `apple, banana, pear`. If the agent cannot quote the routing rule, check the snippet location or chat instructions before continuing. This is a manual activation check, not proof that every future task follows the rules.
21
- 4. At level 2+, run `node bin/cli-run.mjs --doctor` from this folder. It checks binary presence, not authentication or loaded instructions. `--doctor --run` additionally uses a little quota to test live responses. No enabled lanes means delegation is inactive.
22
- 5. To test a real output contract, choose an enabled lane from `bin/lanes.json` and run `node bin/cli-run.mjs <lane> 'Return only {"sorted":["apple","banana","pear"]}' --expect-json`. This uses quota. Expect JSON and exit 0; inspect the array yourself. A non-JSON response exits 10, a missing binary exits 13, and an authentication failure reports the vendor error. The explicit lane tests execution; your primary agent still makes delegation decisions.
18
+ These are the same steps, in the same order, that the installer printed in your terminal.{{CHAT_UPLOAD_NOTE}}
19
+
20
+ {{ACTIVATION_STEPS}}
21
+
22
+ ## Then prove it took
23
+
24
+ {{PROOF_STEPS}}
23
25
 
24
26
  ## What is in this folder
25
27
 
@@ -38,9 +40,9 @@ This folder gives your agent routing instructions and, at level 2+, a runner for
38
40
 
39
41
  Level 2 adds `ROUTING.md`, `TIERS.md`, `DELEGATION_MATRIX.md`, `RESEARCH_TRIAGE.md`, `CLI-RUN.md` and `bin/cli-run.mjs`. Level 3 adds `vm/`. If those files are here, read `ROUTING.md` instead of `ORCHESTRATOR.md`: it is the multi-lane version, and the snippet your agent loads already points at it. `ORCHESTRATOR.md` stays as the single-agent fallback for a session where only one AI is available.
40
42
 
41
- ## Load it into your agent
43
+ ## Where the rules load from
42
44
 
43
- For a CLI primary, the project rules file is `{{PRIMARY_RULES_FILE}}`. Chat apps use pasted instructions instead. The installer wrote a snippet file next to this README (`*.snippet.md`, or `PASTE-INTO-YOUR-AGENT.md` for a chat app). Copy its contents into that file, or paste it into the agent's custom instructions. Nothing was appended to a file you already had.
45
+ {{LOAD_IT}}
44
46
 
45
47
  ## The three rules that carry everything
46
48
 
@@ -50,10 +52,7 @@ For a CLI primary, the project rules file is `{{PRIMARY_RULES_FILE}}`. Chat apps
50
52
 
51
53
  ## Where things went
52
54
 
53
- - This folder: `{{INSTALL_DIR}}`
54
- - Project root (where your agent reads rules and subagents): `{{PROJECT_DIR}}`
55
- - Subagent definitions: `{{AGENTS_DIR}}`
56
- - The rules path your snippets use: `{{RULES_PATH}}`
55
+ {{WHERE_THINGS_WENT}}
57
56
 
58
57
  ## Uninstall
59
58