model-orchestrator 0.1.12 → 0.1.13

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,21 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.13] - 2026-09-08
8
+
9
+ Three issues from a fresh first-run walkthrough of 0.1.12 (#26, #27, #28). Same class as 0.1.12's five: a surface describing an install that did not happen. A fourth, #25, was filed and closed as a mistake on the reporter's side, not a defect: the warning it said was missing has been printed since 0.1.12 and the repro had been read through a truncated pipe.
10
+
11
+ ### Fixed
12
+
13
+ - **The "Then prove it took" list no longer sends a level 1 reader to a file level 1 never wrote** (#27). Step 4 told every reader, at every level, to pick a lane out of `bin/lanes.json` and run `node bin/cli-run.mjs`. Level 1 writes no `bin/` at all, and step 3 immediately above it hedged correctly with "At level 2+" while step 4 did not. The list is now `proofSteps()` in `src/install.js`, gated on level the same way `activationSteps()` is, and the template renders it. Two tests: the README's section must equal the array exactly for every level and primary, and no `bin/` path may appear in it that the plan did not write.
14
+ - **The box setup no longer tells you to sign in to CLIs you did not pick** (#26). `templates/advanced/vm/README.md` step 3 was a fixed sentence naming `codex login --device-auth`, `grok login --device-auth` and `agy`. A level 3 install of claude-code, codex, qwen and ollama was told to sign in to two CLIs it does not have and never told about the one it does. The step now renders each selected CLI's own `auth` string from the catalog. Everything else in that file was already computed from the selection, which is what made the one hardcoded line easy to miss.
15
+ - **A selected local runtime is finally told to install itself** (#26). `activationSteps()` filtered on `kind === 'agent-cli'`, so Ollama, which has a binary and a download page, appeared in no ordered list at any level. Its only mention was one row of a URL table in `DELEGATION_MATRIX.md`. It now gets a step naming the download page and the `ollama pull <model>` that has to follow it.
16
+ - **The tool block stopped saying the same word twice** (#28). Every run that selected a tool printed `optional: Optional. Needs Python 3.10+ and uv.`, because the label repeated the note's own first word. The label is `note:` now. The note keeps the word, because `--list` and the interactive picker print it bare with no label.
17
+
18
+ ### Changed
19
+
20
+ - `--primary` is documented as what it is. `--help` called it "required when several qualify", and then a `--yes` run with several candidates silently picked one in catalog order. The run now names the choice in the plan (`primary claude-code (chosen for you from claude-code, codex; pass --primary to decide it yourself)`) and the help says the same thing. Behaviour is unchanged: the default was sensible, only the promise was wrong.
21
+
7
22
  ## [0.1.12] - 2026-09-08
8
23
 
9
24
  Five issues from one first-run walkthrough of 0.1.11 (#20 to #24). Every one of them is the same failure: a page describing an install that did not happen. Each fix removes the second copy of a fact rather than correcting it.
@@ -185,7 +200,8 @@ First release.
185
200
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
186
201
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
187
202
 
188
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.12...HEAD
203
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...HEAD
204
+ [0.1.13]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.12...v0.1.13
189
205
  [0.1.12]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.11...v0.1.12
190
206
  [0.1.11]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.10...v0.1.11
191
207
  [0.1.10]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.9...v0.1.10
package/bin/cli.js CHANGED
@@ -93,7 +93,8 @@ Usage
93
93
  Flags
94
94
  --level 1|2|3 1 beginner (one agent), 2 intermediate (many CLIs), 3 advanced (plus a VM)
95
95
  --ais a,b,c catalog ids you have access to (see --list)
96
- --primary id the agent that runs the system and receives the subagents (any level; required when several qualify)
96
+ --primary id the agent that runs the system and receives the subagents (any level). When several qualify
97
+ and --yes is set, the run picks one and says so in the plan; pass this to decide it yourself.
97
98
  --tools a,b companion tools to set up, all optional (default with --yes: codecalc only); --no-tools for none
98
99
  --apis a,b level 3 only: metered API keys you HOLD (anthropic,openai,google,xai,openrouter); --no-apis for none.
99
100
  Asked separately from the CLIs because a subscription is not an API key.
@@ -189,6 +190,7 @@ async function main() {
189
190
  // 3. Primary agent (the one that runs the system)
190
191
  const candidates = agentCandidates(selected);
191
192
  let primary = null;
193
+ let primaryAutoPicked = false;
192
194
  if (opt('primary')) {
193
195
  primary = byId[opt('primary')];
194
196
  if (!primary || !candidates.includes(primary)) bad('--primary must be one of: ' + candidates.map((a) => a.id).join(', '));
@@ -200,7 +202,10 @@ async function main() {
200
202
  // --yes picks for the user: claude-code if present, else the first agent that can load subagent
201
203
  // definitions (it gets five files written for it), else the first candidate. #19: codex listed
202
204
  // before agy used to win and nothing was written to the project root.
203
- if (yes) primary = candidates.find((a) => a.id === 'claude-code') || candidates.find((a) => a.agentsDir) || candidates[0];
205
+ if (yes) {
206
+ primary = candidates.find((a) => a.id === 'claude-code') || candidates.find((a) => a.agentsDir) || candidates[0];
207
+ primaryAutoPicked = true;
208
+ }
204
209
  else {
205
210
  console.log('\nWhich one is your primary agent (the one that runs the system)?');
206
211
  candidates.forEach((a, i) => console.log(` ${i + 1} ${a.name}`));
@@ -271,7 +276,7 @@ async function main() {
271
276
  const files = planFiles({ level, selected, primary, dir, project, tools, apis });
272
277
  const lvl = LEVELS.find((l) => l.id === level);
273
278
  const agentFiles = files.filter((f) => f.root === 'project');
274
- console.log(`\nPlan\n level ${lvl.id} ${lvl.name}\n access ${selected.map((a) => a.id).join(', ')}\n primary ${primary ? primary.id : 'none'}\n tools ${tools.map((t) => t.id).join(', ') || 'none'}` + (level >= 3 ? `\n api keys ${apis.map((p) => p.id).join(', ') || 'none'}` : '') + `\n folder ${dir}\n project ${project}${agentFiles.length ? ' (' + agentFiles.length + ' subagent files go here)' : ''}\n files ${files.length}`);
279
+ console.log(`\nPlan\n level ${lvl.id} ${lvl.name}\n access ${selected.map((a) => a.id).join(', ')}\n primary ${primary ? primary.id : 'none'}${primaryAutoPicked ? ` (chosen for you from ${candidates.map((a) => a.id).join(', ')}; pass --primary to decide it yourself)` : ''}\n tools ${tools.map((t) => t.id).join(', ') || 'none'}` + (level >= 3 ? `\n api keys ${apis.map((p) => p.id).join(', ') || 'none'}` : '') + `\n folder ${dir}\n project ${project}${agentFiles.length ? ' (' + agentFiles.length + ' subagent files go here)' : ''}\n files ${files.length}`);
275
280
  if (agentFiles.length && !opt('project')) {
276
281
  console.log(`\nNote: --project was not given, so the ${agentFiles.length} subagent file(s) go to the current directory (${project}). Pass --project to put them somewhere else.`);
277
282
  }
@@ -365,7 +370,7 @@ async function main() {
365
370
 
366
371
  for (const t of tools) {
367
372
  const doc = t.id.toUpperCase() + '.md';
368
- console.log(`\n${t.name}\n optional: ${t.optionalNote}\n needs: ${t.requires}\n run: ${t.install}\n one-click or self-registering for: ${t.autoClients.join(', ')}. Other agents and the details: ${dir}/${doc}`);
373
+ console.log(`\n${t.name}\n note: ${t.optionalNote}\n needs: ${t.requires}\n run: ${t.install}\n one-click or self-registering for: ${t.autoClients.join(', ')}. Other agents and the details: ${dir}/${doc}`);
369
374
  }
370
375
 
371
376
  // 7. Activation summary: writing the folder is half the job. Say exactly what
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.12",
3
+ "version": "0.1.13",
4
4
  "description": "Routing instructions and a CLI runner for your AI tools. One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine.",
5
5
  "type": "module",
6
6
  "bin": {
package/src/install.js CHANGED
@@ -219,12 +219,31 @@ export function activationSteps(opts) {
219
219
  else if (snippet) steps.push(`open ${primary.chatName || primary.name} and paste the block in ${join(dirAbs, snippet)} into its ${primary.chatSurface || 'custom instructions'}`);
220
220
  if (primary && primary.agentsDir) steps.push(`subagents are in ${join(projectAbs, primary.agentsDir)}; run ${primary.bin} from ${projectAbs} to pick them up`);
221
221
  for (const a of selected.filter((a) => a.bin && a.kind === 'agent-cli')) steps.push(`sign in to ${a.name}: ${a.auth}`);
222
+ // A local runtime has a bin but no sign-in, so the agent-cli loop above skips it
223
+ // and before this it appeared in no ordered list at any level (#26).
224
+ for (const a of selected.filter((a) => a.bin && a.kind === 'local')) steps.push(`install ${a.name}: ${a.install.url}, then \`${a.bin} pull <model>\` before the local lane can answer`);
222
225
  for (const t of tools) steps.push(`${t.id}: ${t.install}`);
223
226
  if (level >= 2) steps.push(`smoke test: node ${join(dirAbs, 'bin', 'cli-run.mjs')} --doctor (add --run to send each lane one tiny prompt)`);
224
227
  if (level >= 3) steps.push(`box: read ${join(dirAbs, 'vm', 'README.md')}; keys named in vm/ENVIRONMENT.md go in your secrets manager, never a file`);
225
228
  return steps;
226
229
  }
227
230
 
231
+ // The verification list, in order. Gated on level for the same reason
232
+ // activationSteps is: level 1 writes no bin/, so a step naming cli-run.mjs or
233
+ // lanes.json there described an install that did not happen (#27).
234
+ export function proofSteps(opts) {
235
+ const { level } = opts;
236
+ const steps = [
237
+ 'Start a fresh agent session and ask: "Read the orchestrator instructions. Quote the routing rule you will use, then sort pear, apple, banana alphabetically. Name the tier and whether you delegated."',
238
+ 'Expect the fast tier and `apple, banana, pear`. If the agent cannot quote the routing rule, check the snippet location or chat instructions before continuing. This is a manual activation check, not proof that every future task follows the rules.'
239
+ ];
240
+ if (level >= 2) {
241
+ steps.push('Run `node bin/cli-run.mjs --doctor` from this folder. It checks binary presence, not authentication or loaded instructions. `--doctor --run` additionally uses a little quota to test live responses. No enabled lanes means delegation is inactive.');
242
+ steps.push('To test a real output contract, choose an enabled lane from `bin/lanes.json` and run `node bin/cli-run.mjs <lane> \'Return only {"sorted":["apple","banana","pear"]}\' --expect-json`. This uses quota. Expect JSON and exit 0; inspect the array yourself. A non-JSON response exits 10, a missing binary exits 13, and an authentication failure reports the vendor error. The explicit lane tests execution; your primary agent still makes delegation decisions.');
243
+ }
244
+ return steps;
245
+ }
246
+
228
247
  function vars(opts) {
229
248
  const { level, selected, primary } = opts;
230
249
  const tools = opts.tools || [];
@@ -240,6 +259,7 @@ function vars(opts) {
240
259
  const pinOf = (id) => (toolById[id] && toolById[id].pin) || 'latest';
241
260
  const snippet = snippetFor(primary);
242
261
  const steps = activationSteps({ level, selected, primary, tools, dir: opts.dir, project: opts.project });
262
+ const proofs = proofSteps({ level });
243
263
  // Only claude-code and agy put files under the project root. A chat primary
244
264
  // puts nothing there, so naming a project root would name a folder this run
245
265
  // never created (#21).
@@ -253,6 +273,7 @@ function vars(opts) {
253
273
  return {
254
274
  ...laneVars(selected),
255
275
  ACTIVATION_STEPS: steps.map((st, i) => `${i + 1}. ${st}`).join('\n'),
276
+ PROOF_STEPS: proofs.map((st, i) => `${i + 1}. ${st}`).join('\n'),
256
277
  LOAD_IT: readsProjectRules
257
278
  ? `${primary.name} reads its rules from \`${primary.rulesFile}\` in the project root. The installer wrote \`${snippet}\` next to this README; copy its contents into \`${join(projectAbs, primary.rulesFile)}\`, creating that file if it does not exist. Nothing was appended to a file you already had.`
258
279
  : snippet
@@ -272,6 +293,12 @@ function vars(opts) {
272
293
  INSTALL_DIR: dirAbs,
273
294
  INSTALL_DIR_SH: shellQuote(dirAbs),
274
295
  INSTALL_DIR_SYSTEMD: systemdEscape(dirAbs),
296
+ // vm/README.md step 3 named `grok login` and `agy` whatever you picked (#26).
297
+ VM_SIGNIN: (() => {
298
+ const lines = selected.filter((a) => a.bin && a.kind === 'agent-cli').map((a) => ` - ${a.name}: ${a.auth}`);
299
+ for (const a of selected.filter((a) => a.bin && a.kind === 'local')) lines.push(` - ${a.name}: no sign-in. Install it from ${a.install.url}, then \`${a.bin} pull <model>\`.`);
300
+ return lines.length ? lines.join('\n') : ' - none: no CLI you selected needs a sign-in on the box.';
301
+ })(),
275
302
  AUDIT_LANE: lane || 'none',
276
303
  // Enforced boundary per lane: codex has a read-only sandbox flag; the others
277
304
  // run with whatever their own config allows, and the script says so.
@@ -23,7 +23,8 @@ Generated {{DATE}} for: `{{AI_IDS}}`. Installed at `{{INSTALL_DIR}}`; the system
23
23
 
24
24
  1. Provision a box. Ubuntu, 2+ vCPU, 8 GB is comfortable. Put it on a private mesh network if you can; do not open ports to the internet.
25
25
  2. `bash setup-vm.sh`. It installs system deps and the npm-installable CLIs, then **prints** the vendor shell installers for the rest. Read those scripts before running them.
26
- 3. Sign each CLI in with its device-code flow (`codex login --device-auth`, `grok login --device-auth`, `agy` on first run). Run these inside `tmux` so a dropped SSH session does not kill the prompt. Headless Linux has no keyring by default; `setup-vm.sh` installs one so the CLIs stop re-prompting.
26
+ 3. Sign each CLI in, using the flow its vendor gives you. Run these inside `tmux` so a dropped SSH session does not kill the prompt. Headless Linux has no keyring by default; `setup-vm.sh` installs one so the CLIs stop re-prompting.
27
+ {{VM_SIGNIN}}
27
28
  4. Put provider keys in your secrets manager and export the names listed in `ENVIRONMENT.md` into the gateway's environment at start time. Never write a value into a file in this folder. The gateway config was rendered from the API keys you said you hold, not from your CLI subscriptions: those are different entitlements.
28
29
  5. `docker compose up -d`, then list the lanes without putting the key in argv (the key must be a single token, `^[A-Za-z0-9._-]+$`, because it is interpolated into curl's config grammar):
29
30
  ```bash
@@ -21,10 +21,7 @@ These are the same steps, in the same order, that the installer printed in your
21
21
 
22
22
  ## Then prove it took
23
23
 
24
- 1. Start a fresh agent session and ask: "Read the orchestrator instructions. Quote the routing rule you will use, then sort pear, apple, banana alphabetically. Name the tier and whether you delegated."
25
- 2. Expect the fast tier and `apple, banana, pear`. If the agent cannot quote the routing rule, check the snippet location or chat instructions before continuing. This is a manual activation check, not proof that every future task follows the rules.
26
- 3. At level 2+, run `node bin/cli-run.mjs --doctor` from this folder. It checks binary presence, not authentication or loaded instructions. `--doctor --run` additionally uses a little quota to test live responses. No enabled lanes means delegation is inactive.
27
- 4. To test a real output contract, choose an enabled lane from `bin/lanes.json` and run `node bin/cli-run.mjs <lane> 'Return only {"sorted":["apple","banana","pear"]}' --expect-json`. This uses quota. Expect JSON and exit 0; inspect the array yourself. A non-JSON response exits 10, a missing binary exits 13, and an authentication failure reports the vendor error. The explicit lane tests execution; your primary agent still makes delegation decisions.
24
+ {{PROOF_STEPS}}
28
25
 
29
26
  ## What is in this folder
30
27