model-orchestrator 0.1.34 → 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (108) hide show
  1. package/AGENTS.md +31 -21
  2. package/CHANGELOG.md +51 -1
  3. package/README.md +127 -110
  4. package/bin/README.md +57 -6
  5. package/bin/aunx.js +7 -0
  6. package/bin/cli-run.mjs +21 -15
  7. package/bin/cli.js +376 -257
  8. package/docs/README.md +15 -18
  9. package/docs/catalog.md +228 -38
  10. package/docs/companions.md +28 -10
  11. package/docs/guarantees.md +21 -12
  12. package/docs/how-it-routes.md +49 -42
  13. package/docs/install.md +135 -33
  14. package/docs/part-1-beginner.md +37 -45
  15. package/docs/part-2-intermediate.md +34 -52
  16. package/docs/part-3-advanced.md +36 -26
  17. package/docs/security-review-history.md +38 -0
  18. package/llms.txt +24 -25
  19. package/package.json +16 -8
  20. package/proof/README.md +100 -0
  21. package/proof/gate-demo.cast +9 -0
  22. package/proof/gate-demo.gif +0 -0
  23. package/proof/results.json +198 -0
  24. package/proof/scripts/check-gate.js +26 -0
  25. package/proof/scripts/install-time.js +16 -0
  26. package/proof/scripts/lib.js +73 -0
  27. package/proof/scripts/measure.js +15 -0
  28. package/proof/scripts/missing-results.js +30 -0
  29. package/proof/scripts/record-gate.js +38 -0
  30. package/proof/scripts/render.js +18 -0
  31. package/proof/scripts/runner-overhead.js +21 -0
  32. package/src/README.md +9 -3
  33. package/src/activation-ownership.js +19 -0
  34. package/src/apply-companions.js +104 -0
  35. package/src/apply-snippets.js +60 -28
  36. package/src/aunx.js +262 -0
  37. package/src/catalog.js +253 -117
  38. package/src/install.js +478 -209
  39. package/src/plugin.js +13 -4
  40. package/src/postinstall.js +57 -0
  41. package/src/roles.js +184 -0
  42. package/src/uninstall.js +125 -8
  43. package/templates/README.md +19 -2
  44. package/templates/advanced/README.md +2 -2
  45. package/templates/advanced/vm/PRIVACY_GATES.md +17 -19
  46. package/templates/advanced/vm/README.md +25 -20
  47. package/templates/advanced/vm/box-CLAUDE.md +19 -18
  48. package/templates/advanced/vm/jobs/README.md +3 -1
  49. package/templates/advanced/vm/jobs/weekly-audit.service +3 -0
  50. package/templates/advanced/vm/jobs/weekly-audit.sh +2 -2
  51. package/templates/advanced/vm/setup-vm.sh +49 -2
  52. package/templates/agents/README.md +2 -2
  53. package/templates/agents/agy/README.md +20 -3
  54. package/templates/agents/agy/builder.md +11 -7
  55. package/templates/agents/agy/bulk-worker.md +9 -7
  56. package/templates/agents/agy/code-reviewer.md +13 -7
  57. package/templates/agents/agy/deep-planner.md +10 -7
  58. package/templates/agents/agy/done-verifier.md +13 -22
  59. package/templates/agents/agy/finding-verifier.md +14 -22
  60. package/templates/agents/agy/live-researcher.md +10 -7
  61. package/templates/agents/agy/reader.md +10 -12
  62. package/templates/agents/claude-code/README.md +18 -14
  63. package/templates/agents/claude-code/builder.md +10 -15
  64. package/templates/agents/claude-code/bulk-worker.md +8 -10
  65. package/templates/agents/claude-code/code-reviewer.md +11 -17
  66. package/templates/agents/claude-code/deep-planner.md +9 -11
  67. package/templates/agents/claude-code/done-verifier.md +12 -33
  68. package/templates/agents/claude-code/finding-verifier.md +13 -39
  69. package/templates/agents/claude-code/live-researcher.md +9 -11
  70. package/templates/agents/claude-code/reader.md +9 -18
  71. package/templates/agents/snippets/chat.md +9 -10
  72. package/templates/agents/snippets/claude-code.md +17 -18
  73. package/templates/agents/snippets/generic.md +9 -11
  74. package/templates/agents/snippets/route-gate.mjs +2 -2
  75. package/templates/agents/snippets/route-metrics.mjs +1 -1
  76. package/templates/agents/snippets/subagent-context.mjs +4 -4
  77. package/templates/beginner/ORCHESTRATOR.md +31 -36
  78. package/templates/beginner/README.md +1 -1
  79. package/templates/common/ACCEPTANCE_CHECKS.json +12 -0
  80. package/templates/common/CONTEXT.md +37 -0
  81. package/templates/common/DECISIONS.md +11 -0
  82. package/templates/common/README.md +24 -11
  83. package/templates/common/TASK_BRIEF.md +84 -0
  84. package/templates/common/protocols/README.md +14 -11
  85. package/templates/common/protocols/acceptance-checks.md +14 -0
  86. package/templates/common/protocols/build-protocol.md +91 -106
  87. package/templates/common/protocols/context-file.md +10 -0
  88. package/templates/common/protocols/decision-log.md +9 -0
  89. package/templates/common/protocols/deep-research.md +20 -34
  90. package/templates/common/protocols/docs-then-prove.md +13 -18
  91. package/templates/common/protocols/gap-analysis.md +15 -21
  92. package/templates/common/protocols/memory-and-record.md +21 -20
  93. package/templates/common/protocols/numbers-and-logic.md +20 -26
  94. package/templates/common/protocols/propagate.md +18 -27
  95. package/templates/intermediate/CLI-RUN.md +83 -113
  96. package/templates/intermediate/DELEGATION_MATRIX.md +9 -3
  97. package/templates/intermediate/README.md +3 -3
  98. package/templates/intermediate/RESEARCH_TRIAGE.md +23 -15
  99. package/templates/intermediate/ROUTING.md +54 -51
  100. package/templates/intermediate/TIERS.md +37 -76
  101. package/templates/tools/README.md +1 -1
  102. package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +1 -1
  103. package/docs/audit-brief.md +0 -148
  104. package/scripts/README.md +0 -7
  105. package/scripts/gen-catalog.js +0 -81
  106. package/scripts/gen-plugin.js +0 -16
  107. package/scripts/record-demo.sh +0 -45
  108. package/templates/common/TASK_BUNDLE.md +0 -56
package/AGENTS.md CHANGED
@@ -1,28 +1,38 @@
1
- # AGENTS.md
1
+ # Agent instructions
2
2
 
3
- Two audiences: an agent that wants to USE this package for a project, and an agent that is working ON this repository.
3
+ ## Use the model router in a project
4
4
 
5
- ## Using this package from an agent
5
+ When configuring a project, run `npx model-orchestrator --list` to inspect supported IDs. Preview a scoped setup with:
6
6
 
7
- model-orchestrator writes routing rules, subagents and a CLI lane runner so an agent sends each task to the right model, subagent or CLI and spends fewer frontier tokens. It is not a proxy or gateway. Headless use:
7
+ ```bash
8
+ npx model-orchestrator --yes --level 2 --ais claude-code,codex --project <repo> --dir <repo>/ai-orchestrator --dry-run
9
+ ```
8
10
 
9
- - `npx model-orchestrator --list` prints every supported AI id.
10
- - `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project <repo> --dir <repo>/ai-orchestrator --dry-run` prints the plan and writes nothing.
11
- - Drop `--dry-run` to write it. Existing files are never overwritten without `--force`; activation snippets (for example `CLAUDE.snippet.md`, `settings.hooks.snippet.json`) are written for a person or agent to merge.
12
- - The generated `README.md` in `--dir` lists what to copy where and one smoke command to prove the rules took.
13
- - Claude Code users can install the hooks and subagents as a plugin instead of merging snippets: `claude plugin marketplace add aunysillyme/model-orchestrator`, then `claude plugin install model-orchestrator@model-orchestrator`. The rules still come from the installer above; see [`plugin/README.md`](plugin/README.md).
14
- - A summary for LLMs, with links to every doc: [`llms.txt`](llms.txt).
11
+ When the selection and target paths are correct, remove `--dry-run`. Interactive installs apply the main agent's catalog-supported project rules and settings with backups under the single confirmation; use `--no-apply` or the edit screen to keep activation manual. Headless `--yes` applies activation only with `--apply-snippets`. Every non-dry install runs the local presence health check automatically; live `--run` canaries stay opt-in. Read the generated `README.md` and the final "What's left for you" list for any remaining sign-in or paste steps. When updating an existing install, use `--update-docs` to regenerate unchanged managed documents; edited files stay and are named.
15
12
 
16
- ## Working on this repository
13
+ When choosing tools, select companions explicitly with `--tools`. Defaults select none, including with `--yes`. With activation enabled, supported project MCP configuration is merged with backups; global configuration remains a manual step. The installer runs no third-party installs or login flows. Uninstall removes unchanged recorded activation blocks and the hook or MCP entries it added, preserving surrounding content.
17
14
 
18
- The files the installer writes for end users live under `templates/`.
15
+ When `aunx` is installed, use:
19
16
 
20
- - Read `CONTRIBUTING.md` first, then `src/README.md` (the catalog drives everything) and `docs/audit-brief.md` (the threat model and what has already been attacked).
21
- - Run `npm test` before proposing a change and quote the count and the exit code; the suite prints the current number.
22
- - Everything renders from `src/catalog.js`. Add an AI or a tool there, not in a template. Templates carry no logic.
23
- - `plugin/` is generated. Edit the agent or hook in `templates/`, then `npm run gen:plugin`; `test/plugin.test.js` fails when the committed bundle drifts. A plugin hook may only read: no network, no file writes, no subprocess.
24
- - Never put a value that looks like a credential anywhere in this repo, including tests and examples. Environment variable names only.
25
- - `bin/cli.js` writes only inside `--dir` and `--project`, never over a document without `--force`, and never runs a vendor script. A change that weakens any of those will be refused in review; the tests that hold them are in `test/install.test.js` and `test/cli.test.js`.
26
- - `bin/cli-run.mjs` must exit non-zero when a lane produced nothing. Every judge has a red case in `test/judges.test.js`; add one before you change a judge.
27
- - Prose in this repo uses no em dashes (`test/prose.test.js` enforces it).
28
- - Why these rules exist: each one is the fix for a failure that reached an audit or CI. `CHANGELOG.md` names the issue behind each.
17
+ - **Worker call:** `aunx cli-run codex --brief TASK_BRIEF.md --effort high`.
18
+ - **Shared facts:** `aunx context CONTEXT.md`.
19
+ - **Task scope:** `aunx brief new TASK_BRIEF.md`.
20
+ - **Verification:** `aunx checks ACCEPTANCE_CHECKS.json`, then fill and run trusted commands with `aunx checks run ACCEPTANCE_CHECKS.json`.
21
+ - **Routing suggestion:** `aunx route "rename this file"`, then apply the installed `ROUTING.md` to the actual context.
22
+ - **Own measurements:** `aunx route-metrics --summary`; use [proof/README.md](proof/README.md) for package measurements and scripts.
23
+
24
+ When using the Claude Code plugin, follow [plugin/README.md](plugin/README.md). The installer supplies project-specific routing rules. Use [llms.txt](llms.txt) for the documentation index.
25
+
26
+ ## Change this repository
27
+
28
+ - **Read first:** `CONTRIBUTING.md`, `src/README.md` and `docs/security-review-history.md`.
29
+ - **Catalog:** when adding an AI or companion, edit `src/catalog.js`; keep templates free of logic.
30
+ - **Templates:** when changing installed instructions, edit `templates/`. Use public terms: task brief, context file, acceptance checks, definition of done and result.
31
+ - **Generated files:** run `npm run gen:catalog` after catalog changes and `npm run gen:plugin` after plugin-template changes. `test/plugin.test.js` checks the committed bundle.
32
+ - **Plugin safety:** hooks may only read and emit context. No network, file writes, subprocesses or credential access.
33
+ - **Installer safety:** writes remain inside `--dir` and `--project`; preserve user edits according to manifest hashes and explicit flags. Run no third-party installer.
34
+ - **Runner safety:** preserve exit codes and the log schema. Before changing an output judge, add a failing case in `test/judges.test.js`.
35
+ - **Secrets:** use environment-variable names only. Never add a credential value to code, examples or tests.
36
+ - **Proof:** measure through `proof/scripts/`, store results in `proof/results.json` and regenerate the proof page. The suite rejects expired entries.
37
+ - **Verify:** run `npm test` and report tests, pass, fail, skipped and exit code. The suite prints current counts. Use `npm pack --dry-run` to inspect publication contents.
38
+ - **Style:** use short condition-to-action instructions and no em dashes. `test/prose.test.js` checks public vocabulary and examples.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,54 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [1.0.0] - 2026-09-27
8
+
9
+ ### Added
10
+
11
+ - **The `aunx` command.** The existing package now ships a short command alongside `model-orchestrator`. Run the installer with `aunx`, call an AI through `aunx cli-run`, scaffold a task brief and context file, execute acceptance checks, get a routing suggestion or summarize your local routing activity.
12
+ - **A build process from requirements to verified use.** Shared context, executable acceptance checks, live capability probes, scoped assignments, background heartbeat guidance, split-build ownership and one independent audit step with a companion reviewer.
13
+ - **Reproducible proof.** Measurement scripts, dated results with method and sample size, a generated proof page and a weekly refresh workflow. The test suite rejects expired figures.
14
+
15
+ ### Changed
16
+
17
+ - **Automatic project activation:** interactive installs apply the main agent's catalog-supported rules block and Claude Code hooks under the existing single confirmation, with backups. The edit screen and `--no-apply` keep activation manual; headless `--yes` requires `--apply-snippets` as before.
18
+ - **Remaining actions:** every non-dry install runs the local CLI presence check automatically. "What's left for you" omits completed activation and informational items, checks Codex's reliable sign-in status, and makes other sign-in instructions conditional. Live canaries remain opt-in.
19
+ - **Companion registration and cleanup:** selected companions register in supported project MCP config when activation is enabled; global configuration stays manual. Activation ownership lets uninstall remove unchanged applied blocks and added hooks or MCP entries while preserving unrelated content and backups.
20
+ - **Stack-dependent assignment:** planning, building, review, verification, research, bulk work, reading and private work are assigned from selected capability facts. Generated stack tables explain each choice; the manifest stores the current assignment and `aunx route` reads it. Independent review requires a known different model family, and private work requires local execution.
21
+ - **Model tiers:** generated agent definitions use planning, working and cheap model tiers. Current plan-to-model mappings are unverified, so definitions leave the model to the user's tool configuration while preserving effort guidance.
22
+ - **One-confirm installation:** interactive setup detects available AI tools and shows a proposed level, main agent, paths and role table before one confirmation. The edit menu changes individual settings; an empty detection asks for the user's AIs first. Headless `--yes` still requires explicit level and AI selection.
23
+ - **Capability catalog:** AI entries carry `summary` and `facts`; role claims and duplicated capability fields are retired. The generated catalog marks unknown facts and inherited vendor claims as unverified.
24
+ - **Companions are opt-in.** Default and `--yes` installs select none. Missing tools appear together under "Install these yourself", with their official commands and links.
25
+ - **Public terms match the work.** Task brief, context file, acceptance checks, subscription lanes, pay-per-token lanes, planning model, working model and cheap model.
26
+ - **Documentation starts with the model router.** New command examples, search-phrased questions, a trailer hero, an install walkthrough and a concise security-review history with regression evidence.
27
+ - **A smaller publication surface.** Development-only `scripts/` and the old raw audit brief are excluded; reproducible proof scripts remain available.
28
+
29
+ ### Fixed
30
+
31
+ - **Coherent level installs:** routing uses installed agents or available role fallbacks, review guidance accounts for the main agent's model family, smoke paths are quoted, and README first reads and output-contract checks match the level and selected workers.
32
+ - **Level 3 setup:** initialize the configured local model inside Compose, verify `local-small`, and configure the scheduled service's executable search path.
33
+ - **Accurate level descriptions:** qualify activation and agent sets by main agent, keep companions opt-in, and describe the weekly job's fixed lane, observed data and configurable budget and privacy policies.
34
+ - **Post-install review round:** oversized settings and MCP JSON files are refused before any write, the rules block keeps a CRLF file's line endings, Claude Code sign-in is hidden only when its status reports signed in, Ollama install text appears only when it is absent and names the configured model, CLI main agents without a rules file get accurate wording, and a new rules file is labelled "create".
35
+
36
+ ### Removed
37
+
38
+ - **Automatic vendor installation.** The installer never runs third-party installs. `--no-install` remains accepted for existing scripts.
39
+
40
+ ### Upgrading from 0.1.x
41
+
42
+ - Keep your existing `--dir` and `--project`. Run with `--update-docs --dry` first, then remove `--dry` to apply.
43
+ - Unedited `TASK_BUNDLE.md` migrates to `TASK_BRIEF.md`; an edited legacy brief is kept and named for manual migration.
44
+ - Unchanged runtime files upgrade automatically. `--upgrade-runtime` replaces runtime files only and preserves document edits.
45
+ - Existing codecalc guides and snippets stay managed; uninstall removes them only when unedited. New installs select companions explicitly with `--tools`.
46
+ - Every previous installer flag remains accepted. Runner exit codes, logs, hook contracts and safe uninstall behavior remain compatible.
47
+
48
+ ## [0.1.35] - 2026-09-25
49
+
50
+ ### Changed
51
+
52
+ - **The description names all three things the installer writes.** npm, the GitHub About box and `package.json` now read: rules, subagents and a lane runner tell agents which model to use. The README opening already listed the lane runner; the description stopped at rules and subagents.
53
+ - **Every GitHub topic is an npm keyword.** `coding-agents` was a topic with no matching keyword.
54
+
7
55
  ## [0.1.34] - 2026-09-24
8
56
 
9
57
  ### Changed
@@ -451,7 +499,9 @@ First release.
451
499
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
452
500
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
453
501
 
454
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.34...HEAD
502
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.0...HEAD
503
+ [1.0.0]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.35...v1.0.0
504
+ [0.1.35]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.34...v0.1.35
455
505
  [0.1.34]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.33...v0.1.34
456
506
  [0.1.33]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.32...v0.1.33
457
507
  [0.1.32]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.31...v0.1.32
package/README.md CHANGED
@@ -2,200 +2,217 @@
2
2
 
3
3
  [![npm](https://img.shields.io/npm/v/model-orchestrator.svg)](https://www.npmjs.com/package/model-orchestrator) [![test](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml/badge.svg)](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml) [![license: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![node >=18](https://img.shields.io/badge/node-%3E%3D18-brightgreen.svg)](package.json)
4
4
 
5
- **A model orchestrator for AI coding agents: routing rules, subagents and a lane runner that tell your agent which model handles each task,** so small work goes to cheap tiers and fewer tokens go to frontier models. Answer a few questions and it writes the setup for exactly the AIs you have: Claude Code, Codex, Antigravity (Google), Grok, Qwen, Ollama.
5
+ **Model router for AI coding agents: installs routing rules, 8 subagents, hooks and a CLI runner so your AI picks model and effort per task and saves tokens.**
6
6
 
7
- **The problem:** one agent does every task on its biggest model, so renaming a file costs the same as designing a system.
7
+ Routine work can use a cheap model. Planning and difficult decisions can use a stronger one. Your agent gets editable rules for making that choice and a runner that checks whether delegated work returned a result.
8
8
 
9
- **What you get:** rules your agent follows to keep planning on the frontier model and hand routine work to cheaper tiers and the other AIs you already pay for, plus a log that shows where the work went.
9
+ A **lane** is one AI tool or model your agent can hand work to. A **tier** describes a model's strength and cost: planning model, working model or cheap model.
10
10
 
11
11
  ```bash
12
12
  npx model-orchestrator
13
13
  ```
14
14
 
15
- <img src="docs/demo.gif" alt="A terminal running npx model-orchestrator with --dry: it prints the level, the AIs detected, both target folders and all 38 files it would write, then says nothing was written." width="100%" />
15
+ ![Tasks routed to the right model and effort](docs/router-trailer.gif)
16
16
 
17
- A few questions, then it writes the setup for the AIs you picked. For Claude Code, Codex and Grok that is 38 files:
17
+ ## What the model router gives you
18
18
 
19
- | Part | What it does for you |
19
+ | Part | What you get |
20
20
  |---|---|
21
- | `ROUTING.md`, `TIERS.md`, `DELEGATION_MATRIX.md` | tell your agent which model or CLI handles each kind of task, and at what effort |
22
- | `TASK_BUNDLE.md` and `protocols/` | the brief every hand-off carries, plus build, research, audit and record-keeping steps |
23
- | 8 subagents in `.claude/agents/` | builder, planner, reviewer, researcher, bulk worker, reader and two verifiers, each with its own tools |
24
- | 3 hooks in `.claude/hooks/` | put the routing table in front of your agent on every prompt and every subagent start, and log where work went |
25
- | `bin/cli-run.mjs` | calls Codex, Grok and the other CLIs, and counts a run as done only when it returns a result |
26
- | `mcp/` and `CODECALC.md` | ready-to-paste configs for the companion tools you chose |
21
+ | Routing rules | A decision tree for bulk work, reading, live data, review, verification, planning and builds |
22
+ | Subagents and hooks | Named jobs with explicit tools, plus routing reminders inside Claude Code |
23
+ | Step-by-step playbooks (protocols) | A build process from acceptance checks through independent review and verified use |
24
+ | Task brief and context file | One shared set of facts, plus each worker's scope, permissions and checks |
25
+ | Lane runner | Honest exit codes, optional file or JSON checks, requested model and effort recorded per run |
26
+ | Routing metrics | Your local routing split and subagent activity |
27
27
 
28
- Preview it in any folder, nothing is written:
28
+ The package sits **above the request layer**: your agent reads the rules and picks the lane. Request-level proxies and gateways can carry the API calls underneath it. `aunx route` prints a deterministic suggestion for your agent to consider.
29
29
 
30
- ```bash
31
- npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --dry
32
- ```
33
-
34
- Every file, folder by folder: [docs/install.md](docs/install.md#what-gets-written-level-3-everything).
30
+ **For agents:** [llms.txt](llms.txt) links the reference docs; [AGENTS.md](AGENTS.md) gives headless commands.
35
31
 
36
- The recording above comes from the published package under `asciinema`, rendered with `agg`: `bash scripts/record-demo.sh`.
32
+ Preview a setup:
37
33
 
38
- - **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
39
- - **Where it sits:** above the request layer. Your agent reads the rules and picks the lane, so the decision stays somewhere you can read, version and edit. Request-level routers and gateways sit underneath it.
40
- - **Use it when:** you run more than one model or agent and want the expensive tier kept for planning and judgment.
41
- - **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
42
-
43
- ## Model orchestrator or a model proxy
44
-
45
- An HTTP proxy or gateway such as LiteLLM, Portkey, OpenRouter or claude-code-router
46
- swaps the model per request underneath the agent.
47
- model-orchestrator is an installer that writes routing rules, subagents, hooks
48
- and a lane runner above the request layer for the coding agents and subscription CLIs you already pay for.
49
- Pick a proxy for request-level model routing and a shared API entry point.
50
- Pick model-orchestrator for task delegation across your agents, tiers and CLIs.
51
- They compose: your agent follows the installed rules, and a proxy can route its API requests underneath.
34
+ ```bash
35
+ npx model-orchestrator --yes --level 2 --ais claude-code,codex --primary claude-code --project . --dir ./ai-orchestrator --dry
36
+ ```
52
37
 
53
38
  ## After you install
54
39
 
55
- For the Claude Code setup above, follow the activation summary from the project folder:
56
-
57
- 1. **Rules:** copy the block in `ai-orchestrator/CLAUDE.snippet.md` into `CLAUDE.md` (create it if missing).
58
- 2. **Hooks:** merge `ai-orchestrator/settings.hooks.snippet.json` into `.claude/settings.json` (create it if missing).
59
- 3. **Smoke test:** run `node ./ai-orchestrator/bin/cli-run.mjs --doctor` to check the enabled lanes. Add `--run` to send each lane one tiny prompt.
40
+ The interactive installer applies the main agent's project rules and, for Claude Code, merges its hooks. The summary names these changes before the single confirmation; existing files get timestamped backups. A local health check runs automatically after installation.
60
41
 
61
- Add `--apply-snippets` to apply the rules and hooks steps for a Claude Code primary. It is off by default. The installer replaces one marked block in `CLAUDE.md`, preserves the surrounding bytes, and merges hooks while keeping existing settings and avoiding duplicate commands with the same arguments. Existing files changed by the run get timestamped backups beside them (`<name>.bak-YYYYMMDDTHHMMSS`); every backup path is printed. Preview with `--apply-snippets --dry-run`. Invalid settings JSON stops the run before any writes. Other primaries get the snippet name to paste by hand.
42
+ - **What's left for you:** follow the remaining sign-in or setup steps printed at the end. Codex sign-in is listed only when its reliable status check cannot confirm it; other CLIs get an "if you have not signed in yet" instruction. A chat-app main agent keeps one paste step. A CLI main agent with no cataloged project rules file (for example Grok or Hermes) gets a "load the block" instruction instead of a file to copy.
43
+ - **Your agent:** start a fresh session in the project. `ai-orchestrator/README.md` (or the README in your chosen `--dir`) explains the installed rules and activation check.
44
+ - **Control:** use `--no-apply` or the edit screen to keep activation manual. Headless `--yes` keeps project rules and settings untouched unless you add `--apply-snippets`. Use `--dry` to preview.
62
45
 
63
- Run `claude` from the project folder to load the subagents. Follow the sign-in and companion-tool steps printed for your selection; the same steps are saved in `ai-orchestrator/README.md`.
46
+ Selected companions get project-scoped registration where the host supports it; global configuration remains a printed step. The health check checks CLI presence and sends no prompt. Live checks remain opt-in with `aunx cli-run --doctor --run`. [Installation and upgrades](docs/install.md) cover every flag.
64
47
 
65
- A real call through the lane runner, captured from a fresh install on 2026-09-23 (Codex CLI 0.154.0). Your agent sends a small read to Codex at low effort, and `cli-run` prints one status line with the route it used:
48
+ Install the command once to use the shorter forms below:
66
49
 
67
- ```text
68
- $ node ai-orchestrator/bin/cli-run.mjs codex "In one sentence, what is ai-orchestrator/ROUTING.md for?" --effort low
69
- cli-run[codex] ok rc=0 class=ok refused=null 18.9s raw=20418B route=lane default/low :: turn.completed
50
+ ```bash
51
+ npm install -g model-orchestrator
52
+ aunx --help
70
53
  ```
71
54
 
72
- Each call also appends one line to `~/.ai-orchestrator/cli-run.log.jsonl` with the lane, the model and effort requested and resolved, the verdict, the exit code, the seconds and the deliverable size, so you can see where the work went. A run that produces no deliverable exits non-zero: on the same install, `--expect-file summary.md` for a file the lane never wrote printed `no_deliverable rc=10 class=empty` and a fix line.
55
+ `aunx` without a subcommand runs the same installer as `model-orchestrator`. The installer writes its files and the project activation changes shown in the summary. Every companion is opt-in, including with `--yes`; missing tools appear together under **Install these yourself**, with official commands and links.
73
56
 
74
57
  ## Part of a set
75
58
 
76
- Three open-source tools that work on their own and fit together:
77
-
78
59
  | Repo | What it gives you |
79
60
  |---|---|
80
61
  | [agent-personalizer](https://github.com/aunysillyme/agent-personalizer) | One interview writes the profile and rules every AI you use reads, kept in sync from one source. |
81
- | **model-orchestrator** | Routing rules that tell your agent which model handles each task, so frontier models do the hard work and cheaper tiers do the rest. |
62
+ | **model-orchestrator** | Model router for AI coding agents: installs routing rules, 8 subagents, hooks and a CLI runner so your AI picks model and effort per task and saves tokens |
82
63
  | [website-build-skill](https://github.com/aunysillyme/website-build-skill) | A skill pack that teaches your AI current website-building expertise: research, design, code, accessibility, performance, search and security. |
83
64
 
84
- ## The three levels
65
+ ## Model routing and request-level proxies
66
+
67
+ An HTTP proxy or gateway such as LiteLLM, Portkey, OpenRouter or claude-code-router swaps the model per request underneath the agent. model-orchestrator is an installer that writes routing rules, subagents, hooks and a lane runner above the request layer.
68
+
69
+ - **Pick a proxy** for request-level model routing and a shared API entry point.
70
+ - **Pick model-orchestrator** for task delegation across your agents, model tiers and CLIs.
71
+
72
+ They compose: your agent follows the installed rules, and a proxy can route its API requests underneath.
85
73
 
74
+ ## Choose the setup that fits your tools
86
75
 
87
- | Level | You have | You get |
76
+ | Level | Your setup | What it adds |
88
77
  |---|---|---|
89
- | **1 · Beginner** | one LLM or one agent | tiers, task classification, the two build checkpoints, the protocols (build, propagate, gap analysis, deep research, numbers and logic, memory and record, docs and proof), a task-bundle template, and your agent set up to follow them |
90
- | **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts; `--model` / `--effort` to pin the route and log it), a delegation matrix generated from your selection, research triage across the lanes you have |
91
- | **3 · Advanced** | a virtual machine | everything above, plus a gateway config rendered from the API keys you hold (asked separately from your CLIs), pinned images, box rules, privacy gates, and a weekly gap-analysis job with "what watches it" written down |
78
+ | Beginner | One agent or chat app | Task classification, model tiers, a task brief, acceptance checks and build protocols |
79
+ | Intermediate | Several AI CLIs | A lane runner, delegation matrix, research triage and model/effort flags |
80
+ | Advanced | An always-on Linux machine | Gateway templates, privacy rules and a scheduled review job |
92
81
 
93
- Levels explained: [Part 1](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md).
82
+ Read [beginner](docs/part-1-beginner.md), [intermediate](docs/part-2-intermediate.md) or [advanced](docs/part-3-advanced.md).
94
83
 
95
- ## The AIs it knows about
84
+ ## Works with the AIs you already pay for
96
85
 
97
- | Id | What | Level |
86
+ The installer detects your AI tools, shows the proposed setup and asks **Write these files?** with `[Y/n/e]`. Confirm once, or enter `e` to change a setting. With no detected tools, select the AIs you have first. The main agent receives its supported agent set; at level 2 and up, selected CLIs supported by the runner become worker lanes.
87
+
88
+ | AI | Installer ID | What it is and gives |
98
89
  |---|---|---|
99
- | `claude-code` | Claude Code CLI, the default orchestrator | 1+ |
100
- | `codex` | Codex CLI on a ChatGPT plan: second coder and second-opinion reviewer (a different model family reading your diff) | 1+ |
101
- | `agy` | Antigravity CLI on a Google AI plan: research sweeps, concurrent fan-out | 1+ |
102
- | `grok` | Grok CLI on X Premium: live X and web reads at $0 | 1+ |
103
- | `hermes` | Hermes Agent: the free tier | 2+ |
104
- | `qwen` | Qwen Code CLI with a cheap metered model: structured bulk | 2+ |
105
- | `ollama` | local models: the privacy lane | 2+ |
106
- | `claude-app`, `chatgpt-app`, `gemini-app` | chat apps: level 1 via a paste block | 1 |
90
+ | Claude Code | `claude-code` | Anthropic's terminal agent; project rules, subagents and routing hooks when main |
91
+ | Codex | `codex` | OpenAI's terminal agent; read-only filesystem sandbox for `--audit`, project rules when main |
92
+ | Antigravity | `agy` | Google's terminal agent; project rules and custom agents when main |
93
+ | Grok | `grok` | xAI's terminal agent on a subscription |
94
+ | Hermes | `hermes` | A free terminal agent using the providers you authenticate |
95
+ | Qwen Code | `qwen` | A terminal agent using your provider key; project rules when main |
96
+ | Ollama | `ollama` | A model runtime that runs on your machine |
97
+ | Claude, ChatGPT and Gemini apps | `claude-app`, `chatgpt-app`, `gemini-app` | `PASTE-INTO-YOUR-AGENT.md`: a routing block for your chat app |
98
+
99
+ Every install includes **Your stack: who does what**: planning, building, independent review, verification, research, bulk work, reading and private work, plus fan-out and long-context roles when supported. One deterministic assignment uses the selected tools' capabilities, billing and selection order. Independent review requires a known different model family; private work requires a local runtime. Unavailable roles are stated explicitly.
100
+
101
+ Subagent definitions name planning, working or cheap model tiers. Your plan and tool configuration select the actual models. `MANIFEST.json` stores the assignment, and `aunx route` reads it even when you keep edited routing documents. [How assignment works](docs/how-it-routes.md).
107
102
 
108
- `npx model-orchestrator --list` prints the catalog with install and sign-in notes. Details: [docs/catalog.md](docs/catalog.md).
103
+ `npx model-orchestrator --list` prints supported IDs and setup notes. [The catalog](docs/catalog.md) lists installation, sign-in and detection details.
109
104
 
110
- ## Measuring routing
105
+ ## Lane runner (`aunx cli-run`)
111
106
 
112
- See where your agent sends the work. On a claude-code install, `route-metrics.mjs` turns every turn, dispatch and subagent start/stop into one JSON line under `~/.ai-orchestrator/route-metrics.jsonl`, including the lane your agent named in its own `<!-- route: <lane> | <why> -->` marker.
107
+ Run from your project. `aunx cli-run` uses the package runner; pass `--dir` to use a project's installed runner. Replace `<lane>` below with a CLI lane from your generated stack table.
113
108
 
114
109
  ```bash
115
- node .claude/hooks/route-metrics.mjs --summary # since the log began
116
- node .claude/hooks/route-metrics.mjs --summary --since 2026-09-01 # since a date
110
+ aunx cli-run --doctor
111
+ aunx cli-run '<lane>' --brief TASK_BRIEF.md
112
+ # Direct form from the installed rules folder:
113
+ node bin/cli-run.mjs '<lane>' --brief TASK_BRIEF.md
117
114
  ```
118
115
 
119
- The report prints turns, **route-marker coverage** (the share of turns that carried a real lane, which answers "is the agent actually tagging its routing decisions?"), **work sent off the main session**, lanes by count, dispatches by `subagent_type`, and mean/max duration per agent type. The log holds exactly five things: a timestamp, the event, the session id, the lane name and the agent type. Every turn keeps running whatever the hook does, so the worst case is a quieter report.
116
+ `--doctor --run` sends a small live check through your own vendor sign-ins. Each run records the requested model and effort and a fixed result class in a local log. A missing result returns nonzero. Add `--expect-file` or `--expect-json` when success needs a concrete output contract. [Runner reference](bin/README.md).
120
117
 
121
- ## Claude Code plugin
118
+ ## Task brief (`aunx brief`) and acceptance checks (`aunx checks`)
122
119
 
123
- The hooks and subagents also ship as a plugin, so they install and update through Claude Code itself:
120
+ ```bash
121
+ aunx context CONTEXT.md
122
+ aunx brief new TASK_BRIEF.md
123
+ aunx checks ACCEPTANCE_CHECKS.json
124
+ aunx checks run ACCEPTANCE_CHECKS.json
125
+ ```
126
+
127
+ Fill the context file with verified facts, quote the user's ask in the brief, and give every requirement a check command. The check runner reports PASS or FAIL and exits 1 when a check fails. Run only check files you trust: their commands execute with your shell's permissions. See [the acceptance-check protocol](templates/common/protocols/acceptance-checks.md).
128
+
129
+ ## Routing suggestions (`aunx route`)
130
+
131
+ ```bash
132
+ aunx route "rename this file"
133
+ aunx route "design the auth system"
134
+ ```
135
+
136
+ The first suggests cheap bulk work; the second suggests planning. Each prints the AI assigned to that role in your installation. Use `--dir PATH` for a custom rules folder; otherwise it reads `./ai-orchestrator/MANIFEST.json`, then `./MANIFEST.json`. With no install, it prints a generic suggestion and an install notice. The classifier uses keywords and points unmatched requests to `ROUTING.md`. It reads JSON and launches no worker.
137
+
138
+ ## See where your agent sends work (`aunx route-metrics`)
124
139
 
140
+ ```bash
141
+ aunx route-metrics --summary
142
+ aunx route-metrics --summary --since 2026-09-01
125
143
  ```
144
+
145
+ On Claude Code installs, the routing hook records turns, route markers and subagent activity locally. The summary reports your routing split, route-marker coverage and agent durations. Prompt text and the explanation inside a route marker never enter that log. [Measured results and reproduction scripts](proof/README.md) show dated figures with methods and sample sizes; expired entries fail the test suite.
146
+
147
+ ## Claude Code plugin
148
+
149
+ ```text
126
150
  /plugin marketplace add aunysillyme/model-orchestrator
127
151
  /plugin install model-orchestrator@model-orchestrator
128
152
  ```
129
153
 
130
- It ships the two read-only hooks (`route-gate`, `subagent-context`) and the eight subagents, each with an explicit tool list, and loads them namespaced as `model-orchestrator:builder`. The routing rules come from `npx model-orchestrator`, which is the step that reads your setup and writes rules to match it. `plugin/` is generated from `templates/`, and `test/plugin.test.js` holds the bundle to that shape: committed output matches the generator, hooks stay read-only, every agent keeps its tool list. The third hook, `route-metrics`, writes a log, so it comes only with the npm install. Details: [plugin/README.md](plugin/README.md).
154
+ The plugin carries read-only routing hooks and the subagents. Generate your project's routing rules with `npx model-orchestrator`. The npm installer also supplies the local metrics hook. [Plugin setup](plugin/README.md).
131
155
 
132
- ## Companion tools (all optional)
156
+ ## Works well with
133
157
 
134
- An orchestrator routes work. Three companion tools cover the rest of what a working agent needs, exact numbers, a memory, and current library docs:
158
+ These are other authors' projects, maintained in their own repositories. All companions start unselected. Choosing one writes guidance and configuration snippets. With activation enabled, supported project configuration is merged automatically; you install the tool and complete any printed setup steps yourself.
135
159
 
136
- | Tool | What it gives you | Set up |
160
+ | Project | Author | What it adds |
137
161
  |---|---|---|
138
- | [codecalc](https://github.com/The-40-Thieves/codecalc) | exact arithmetic, code execution in 31 languages and logic checks, running offline | by default |
139
- | [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | durable memory with hybrid search, backlinks and compare-and-swap writes; local by default | when you pick it |
140
- | [Context7](https://github.com/upstash/context7) | current, version-specific library docs pulled into the prompt | when you pick it |
162
+ | [codecalc](https://github.com/The-40-Thieves/codecalc) | The-40-Thieves | Local calculation, code execution and logic checks |
163
+ | [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | The-40-Thieves | Searchable notes and controlled writes over an Obsidian vault |
164
+ | [Context7](https://github.com/upstash/context7) | Upstash | Current, version-specific library documentation |
165
+
166
+ Use `--tools codecalc,obsidian-tc,context7` to select them. Without a companion, use your available calculator or runtime, a searchable notes folder and official library docs. [Companion setup and upstream support](docs/companions.md).
141
167
 
142
- Selecting one writes a doc and the config snippets for your agents, so the install stays yours to run. What each needs first, and how Context7 and codecalc pair up (docs say what an API should do, a run proves what it does): [docs/companions.md](docs/companions.md). Every level carries the three rules they serve either way: `protocols/numbers-and-logic.md`, `protocols/memory-and-record.md` and `protocols/docs-then-prove.md`.
168
+ ## Common questions
143
169
 
144
- ## Principles the whole thing rests on
170
+ ### What is a model router for coding agents?
145
171
 
146
- 1. **Route by capability tier.** Start at the smallest tier that fits, and let evidence move it up.
147
- 2. **Every gate can come back wrong.** A checkpoint earns its place by being answerable both ways.
148
- 3. **Check for the artifact.** A deliverable is a file, a commit or a line you can point at. `cli-run` checks the response is structurally there; `--expect-file` checks the artifact.
149
- 4. **Numbers are computed.** A tool that calculates beats a model that feels finished.
150
- 5. **A write stays findable.** Search first, keep the index true, one writer.
151
- 6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, the brief is what carries this task. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
152
- 7. **Only one process holds keys.** Names in the environment, values in a secrets manager.
172
+ It helps an agent match a task to a model, effort level and toolset. Run `npx model-orchestrator` to install editable rules, subagents and a CLI runner for your setup; your agent makes the routing decision.
153
173
 
154
- ## Common questions
174
+ ### How do I use Claude Code and Codex together?
175
+
176
+ Run `npx model-orchestrator --yes --level 2 --ais claude-code,codex --primary claude-code --project . --dir ./ai-orchestrator`, then follow the activation summary. Claude Code can dispatch scoped work through `aunx cli-run codex --brief TASK_BRIEF.md` and use a different model family for review.
155
177
 
156
- ### How do I cut token usage across Claude Code, Codex and Antigravity (Google)?
178
+ ### How do I reduce Claude Code token usage?
157
179
 
158
- Install for the tools you have, then let the generated `ROUTING.md` decide the tier per task: bulk, reading and verification go to the fast tier or a cheaper CLI lane, and the deep tier only plans and judges. On Claude Code, execution goes to the `builder` subagent by default and the main session plans and verifies. Every lane call through `cli-run` logs the model and effort it ran with, so you can check where the tokens went.
180
+ Install routing rules with `npx model-orchestrator`, so your agent has guidance for sending routine work to cheaper models and keeping reads scoped. Use `aunx route-metrics --summary` to measure where your work goes; savings depend on your tasks and model choices.
159
181
 
160
182
  ### How do I route tasks to cheaper models?
161
183
 
162
- The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](docs/how-it-routes.md#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly stays on the cheap tier, and the frontier tokens go to the work that earns them.
184
+ Use `aunx route "rename this file"` for a keyword-based suggestion, then apply your installed `ORCHESTRATOR.md` at level 1, or `ROUTING.md` and `TIERS.md` at level 2 and up, to the actual task. Role selects the job, complexity sets effort, and the consequences of a mistake affect the model and reviewer.
163
185
 
164
- ### Where does this sit next to an LLM router or an AI gateway?
186
+ ### How does this work with an AI gateway or LLM router?
165
187
 
166
- One layer up, and they compose. This routes at the task level, through instructions your agent follows and a runner for agent CLIs. Request-level routers and gateways (RouteLLM, LiteLLM, OpenRouter, claude-code-router) forward the model on every API call, and they sit underneath this happily: pick the lane here, let the gateway carry the call.
188
+ It coordinates multi-agent work at task level; a proxy such as LiteLLM or OpenRouter can route the API requests underneath it. `npx model-orchestrator --level 3` includes gateway templates when you want that setup.
167
189
 
168
190
  ### How does an agent install and run it headlessly?
169
191
 
170
- Use the CLI flags. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run` previews the plan, and `--list` prints every supported AI. Your own documents are kept unless you pass `--force`; activation snippets are ready to merge. `MANIFEST.json` and `bin/lanes.json` are rewritten each run. Runtime files upgrade when they match the recorded hash; edited copies are kept and named. See [re-running an install](docs/install.md#what-a-run-does) for `--update-docs` and `--upgrade-runtime`.
171
-
192
+ Pass `--yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator` to `npx model-orchestrator`; add `--dry-run` to preview. Existing edits are preserved by default, and `--update-docs` refreshes files whose installed hashes still match.
172
193
 
173
194
  ## Uninstall
174
195
 
175
- Remove unedited files recorded by the installer; edited files stay and are listed.
176
- Run `npx model-orchestrator --uninstall --dir ./ai-orchestrator --project .` (add `--dry` to preview).
177
- Remove the pasted rules block and merged hooks entry by hand. [Removal details](docs/install.md#uninstall).
196
+ Run `npx model-orchestrator --uninstall --dir ./ai-orchestrator --project .` (add `--dry` to preview). The installer removes unedited managed files, its recorded activation block and the hook entries it added, preserving surrounding rules and settings. It names edited or manually pasted entries that need your attention; backups stay. [Removal details](docs/install.md#uninstall).
178
197
 
179
198
  ## Read next
180
199
 
181
- | Doc | What is in it |
182
- |---|---|
183
- | [docs/install.md](docs/install.md) | every flag, the two folders a run writes to, headless examples, the full file list |
184
- | [docs/how-it-routes.md](docs/how-it-routes.md) | role, complexity and stakes; the three verifier agents; pinning a lane's model and effort |
185
- | [docs/guarantees.md](docs/guarantees.md) | what is enforced by code, what is delegated to a vendor flag, and what is only an instruction |
186
- | [docs/part-1-beginner.md](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md) | the thinking behind each level |
187
- | [docs/catalog.md](docs/catalog.md) | every supported AI with install and sign-in notes |
200
+ - [Install and upgrade](docs/install.md): flags, folders, walkthrough and safe reruns.
201
+ - [How routing works](docs/how-it-routes.md): model, effort and independent verification.
202
+ - [Guarantees](docs/guarantees.md): executable checks and agent instructions.
203
+ - [Proof](proof/README.md): dated measurements and scripts you can rerun.
204
+ - [Security review history](docs/security-review-history.md): findings, fixes and regression evidence.
188
205
 
189
206
  <details>
190
207
  <summary><strong>Platform support, and every test this suite skips</strong></summary>
191
208
 
192
209
 
193
- Node 18 or newer, with zero runtime dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22), including lane execution end to end through `cli-run` against a fake CLI installed the same way npm installs a real one (a `.cmd` shim). `cli-run` never runs a lane through `cmd.exe` when it can avoid it: it resolves the shim to the Node script underneath and spawns Node directly, so a prompt reaching a real lane never passes through a Windows shell. A `.cmd` or `.bat` lane that cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, rather than run through `cmd.exe`: a batch file re-reads its arguments after `cmd.exe` has parsed them once, and no escaping fully contains a prompt through both passes. Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested on Windows too, including SIGTERM/SIGINT to the wrapper (Windows has no OS-level signals: both terminate it unconditionally, verified there rather than treated the same as POSIX). The Windows skip list covers POSIX behavior, with each skip pinned by `test/prose.test.js`: `statSync().mode`'s executable bit (NTFS has none, so that one assertion is conditional inside a test that otherwise runs everywhere); a lane dying mid-run from a real POSIX signal (a real Windows lane cannot die "by signal"); running `weekly-audit.sh`'s watchdog functions for real under Git Bash's job control, both the end-to-end run and the `bounded()` timeout check (the script itself only ever runs on the Ubuntu box it targets); and a `mkfifo` FIFO at the rules path, the one case that proves `route-gate.mjs` cannot HANG on a non-regular file, since Windows has no `mkfifo` to build one (the guard behind it is covered on every OS by a directory at the same path); and an untracked `mkfifo` FIFO in the repository `cli-run --audit` sizes, the case that proves `--effort auto` never opens a non-regular file (the symlink half of that test runs on every OS). `test/prose.test.js` counts every `skip:` in the suite and requires this list to document each one.
210
+ Node 18 or newer, with zero runtime dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22), including lane execution end to end through `cli-run` against a fake CLI installed the same way npm installs a real one (a `.cmd` shim). `cli-run` never runs a lane through `cmd.exe`: it resolves the shim to the Node script underneath and spawns Node directly, so a prompt reaching a real lane never passes through a Windows shell. A `.cmd` or `.bat` lane that cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, rather than run through `cmd.exe`: a batch file re-reads its arguments after `cmd.exe` has parsed them once, and no escaping fully contains a prompt through both passes. Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested on Windows too, including SIGTERM/SIGINT to the wrapper (Windows has no OS-level signals: both terminate it unconditionally, verified there rather than treated the same as POSIX). The Windows skip list covers POSIX behavior, with each skip pinned by `test/prose.test.js`: `statSync().mode`'s executable bit (NTFS has none, so that one assertion is conditional inside a test that otherwise runs everywhere); a lane dying mid-run from a real POSIX signal (a real Windows lane cannot die "by signal"); running `weekly-audit.sh`'s watchdog functions for real under Git Bash's job control, both the end-to-end run and the `bounded()` timeout check (the script itself only ever runs on the Ubuntu box it targets); and a `mkfifo` FIFO at the rules path, the one case that proves `route-gate.mjs` cannot HANG on a non-regular file, since Windows has no `mkfifo` to build one (the guard behind it is covered on every OS by a directory at the same path); and an untracked `mkfifo` FIFO in the repository `cli-run --audit` sizes, the case that proves `--effort auto` never opens a non-regular file (the symlink half of that test runs on every OS). `test/prose.test.js` counts every `skip:` in the suite and requires this list to document each one.
194
211
 
195
212
  **Privacy.** The installer sends no telemetry and makes no network call of its own once it is running. Two things around that are worth being exact about:
196
213
 
197
214
  - `npx model-orchestrator` is itself a download: npm fetches this package from the registry before any of it runs. `npm install -g model-orchestrator` once, then run `model-orchestrator`, if you would rather that happen exactly one time.
198
- - For a missing vendor CLI, the installer prints the install command. An interactive run offers to run one pinned `npm install -g` per package with your confirmation; `--yes` and `--no-install` keep installation in your hands. Vendor shell installers (Antigravity, Grok) are only ever printed, alongside the `curl … | less` you would use to read one before running it.
215
+ - Every missing vendor CLI or selected companion is listed under **Install these yourself**, with an official command and link. The installer runs no third-party installs. `--no-install` remains accepted for existing scripts.
199
216
 
200
217
  `cli-run` calls the vendor CLI you name.
201
218
 
@@ -206,7 +223,7 @@ Node 18 or newer, with zero runtime dependencies. Works on macOS and Linux; the
206
223
  <summary><strong>Vendor version compatibility</strong></summary>
207
224
 
208
225
 
209
- **Detection checks whether a binary is present.** Use the compatibility table below to compare vendor versions. `--doctor` reports presence; add `--run` to send every enabled lane a small canary and check its sign-in and output against the runner's success criteria.
226
+ **Detection checks whether a binary is present.** Use the compatibility table below to compare vendor versions. `--doctor` reports presence; add `--run` to send a small canary to every enabled worker and check its sign-in and output against the runner's success criteria.
210
227
 
211
228
  The lane wiring and the output judges were written against these versions, which are the ones this release was exercised on:
212
229
 
@@ -228,7 +245,7 @@ Generated from `src/catalog.js` by `npm run gen:catalog`; `npm test` fails if th
228
245
 
229
246
  For npm-installed lanes, `builtAgainst` in the catalog supplies both the compatibility table and the install pin. Newer vendor versions may work or may change a flag the generated wiring uses. When a lane starts failing after a vendor upgrade, compare against this table first.
230
247
 
231
- **The live canary runs on your machine, with your credentials.** That is what `node bin/cli-run.mjs --doctor --run` is: it sends every enabled lane one tiny prompt through your own sign-ins and reports `canary ok` or `canary FAILED rc=` per lane. Run it after install, and again after any vendor upgrade.
248
+ **The live canary runs on your machine, with your credentials.** That is what `aunx cli-run --doctor --run` (direct form: `node bin/cli-run.mjs --doctor --run`) is: it sends every enabled lane one tiny prompt through your own sign-ins and reports `canary ok` or `canary FAILED rc=` per lane. Choose this optional live check after setup or a vendor upgrade when you want to verify actual responses.
232
249
 
233
250
  CI runs the full suite against stub lanes on Ubuntu, macOS and Windows, Node 18/20/22, plus a packaged install into a clean consumer. Run the live check locally to verify your own sign-ins, quota and vendor versions.
234
251