model-orchestrator 0.1.25 → 0.1.27

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,22 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.27] - 2026-09-21
8
+
9
+ ### Changed
10
+
11
+ - **README restructured for a first-time reader: 32,083 bytes to 17,586, with the first screen now a claim, the install command and a real `--dry` plan.** The old opening was a 200-word paragraph followed by "What it is not", the flag-conflict rules and the uninstall procedure, all before the reader had seen the tool do anything. The detail moved rather than went away: `docs/install.md` (every flag, the two target folders, headless examples, the full file list), `docs/how-it-routes.md` (role, complexity and stakes; the three verifier agents; pinning model and effort), `docs/guarantees.md` (enforced by code, delegated to a vendor flag, or only an instruction), `docs/companions.md` (the full companion-tool table). `docs/README.md` and `llms.txt` index all four. Platform support and the vendor compatibility table stay in the README inside collapsed `<details>` blocks, because `scripts/gen-catalog.js` writes the vendor table between markers there and `test/prose.test.js` requires the skipped-test explanations to live in the README; collapsing them keeps both mechanisms pointed at the same file. The opening line still satisfies `test/copy.test.js`: it names the package, carries the shared purpose clause, and keeps the 0.1.11 correction that it does not select models itself.
12
+
13
+ ## [0.1.26] - 2026-09-19
14
+
15
+ ### Added
16
+
17
+ - **obsidian-tc on Windows: the answer is that no client needs a `cmd /c` snippet, and `OBSIDIAN-TC.md` now says so with the code it was read from.** All five obsidian-tc snippets spawn `"command": "npx"`, and on Windows `npx` is `npx.cmd`, which a bare `CreateProcess` will not start. Each client was read instead of assumed, and all five resolve it themselves: Zed hands the command to the system shell (`crates/context_server/src/transport/stdio_transport.rs`, `ShellBuilder::new(&Shell::System, ..)`, from its PR #42382, 2025-12-10), VS Code's `formatSubprocessArguments` in `src/vs/workbench/api/node/extHostMcpNode.ts` resolves the executable and re-spawns with `shell: true` for a `.bat` or `.cmd`, the Codex CLI calls `which::which_in` in `codex-rs/rmcp-client/src/program_resolver.rs`, and Cursor and Antigravity's `agy` both reach `npx` through `@modelcontextprotocol/sdk`, whose `client/stdio.js` imports `cross-spawn` and re-invokes a non-`.exe` as `%COMSPEC% /d /s /c`. That last point is wider than those two clients: every published SDK from 1.23.0 through 1.30.0 depends on `cross-spawn ^7.0.5`, so `shell: false` in that transport is not the whole story. The doc also names the one Windows case that can still fail and is not about `.cmd`: Zed prefers PowerShell, which resolves a bare `npx` to npm's `npx.ps1` shim, and a `.ps1` will not run under the `Restricted` execution policy that is Windows' client default. A test pins the snippet set, fails if any `.windows` file appears, and fails if the doc drops a citation. Not run on a Windows machine.
18
+
19
+ ### Fixed
20
+
21
+ - **Context7's Windows guidance said something untrue about Zed, and its `cmd` wrapper was missing `/d`** ([#34](https://github.com/aunysillyme/model-orchestrator/issues/34) follow-up). 0.1.25 shipped `mcp/context7.zed.windows.settings.json` on the premise that Zed spawns `"command": "npx"` directly and so cannot start it. Reading Zed's own source for the obsidian-tc work above showed it has launched MCP stdio servers through the system shell since PR #42382 (2025-12-10), so a current Zed starts the plain snippet. The file stays, because it is still the fix for an older Zed and for the PowerShell execution-policy case, but `CONTEXT7.md` no longer sends every Windows user to it and now carries both reasons. The snippet's args change from `["/c", "npx", ...]` to `["/d", "/c", "npx", ...]`: without `/d`, `cmd` first runs whatever sits in the Command Processor `AutoRun` registry value, which can print non-JSON into the protocol stream. `cross-spawn` passes `/d /s /c` for the same reason. The `#34` test follows the new args and now also fails if the doc stops naming why a current Zed does not need the file.
22
+
7
23
  ## [0.1.25] - 2026-09-19
8
24
 
9
25
  ### Fixed
@@ -365,7 +381,9 @@ First release.
365
381
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
366
382
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
367
383
 
368
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.25...HEAD
384
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.27...HEAD
385
+ [0.1.27]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.26...v0.1.27
386
+ [0.1.26]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.25...v0.1.26
369
387
  [0.1.25]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.24...v0.1.25
370
388
  [0.1.24]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.23...v0.1.24
371
389
  [0.1.23]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.22...v0.1.23
package/README.md CHANGED
@@ -2,48 +2,58 @@
2
2
 
3
3
  [![npm](https://img.shields.io/npm/v/model-orchestrator.svg)](https://www.npmjs.com/package/model-orchestrator) [![test](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml/badge.svg)](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml) [![license: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![node >=18](https://img.shields.io/badge/node-%3E%3D18-brightgreen.svg)](package.json)
4
4
 
5
- **Route every task to the right model, agent or LLM, and spend fewer tokens.** A model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. One installer asks what you have access to and writes routing rules, subagents and a CLI runner for exactly that setup, from one chat app to several agent CLIs or a virtual machine. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models; the runner executes the lane it is given. On Claude Code it also delegates execution to subagents by default, with two hooks that inject the routing table every turn, and the hooks and subagents also install as a [Claude Code plugin](#claude-code-plugin).
6
-
7
- ## At a glance
8
-
9
- - **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
10
- - **What it is not:** a proxy, a gateway or an API router. It does not automatically compare prices or select models; your agent follows the rules and chooses.
11
- - **Install:** `npx model-orchestrator` (interactive), or headless from a script or an agent: `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator`.
12
- - **Claude Code plugin:** `/plugin marketplace add aunysillyme/model-orchestrator`, then `/plugin install model-orchestrator@model-orchestrator`. The routing rules still come from the installer; see [Claude Code plugin](#claude-code-plugin).
13
- - **Use it when:** you run more than one model or agent and want each task sent to the smallest one that can do it well.
14
- - **What it saves:** frontier-model tokens. Bulk work, reading and checks go to fast tiers; the expensive tier is kept for planning and judgment.
15
- - **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
16
-
17
- Built from a working system, not a diagram: the routing rules, the protocols and the lane runner here run in production, generalized so they transfer to any stack.
5
+ **A model orchestrator that sends every task to the smallest model that can do it,** so small work goes to cheap tiers and fewer tokens go to frontier models. One command reads which AIs you actually have, then writes the routing rules, the subagents and the lane runner for exactly that set: Claude Code, Codex, Gemini, Grok, Qwen, Ollama.
18
6
 
19
7
  ```bash
20
8
  npx model-orchestrator
21
9
  ```
22
10
 
23
- That runs the latest published release from the npm registry, and `npx model-orchestrator --version` prints which one you got. To run the current main straight from GitHub instead: `npx github:aunysillyme/model-orchestrator`. Add `#vX.Y.Z` for one specific release; the tags are on the [releases page](https://github.com/aunysillyme/model-orchestrator/releases), so this page never pins a number the registry has moved past.
24
-
25
- The installer asks a few things, then writes a folder:
11
+ Three questions, then 38 files. Here is a real `--dry` run, which prints the plan and writes nothing:
12
+
13
+ ```text
14
+ Plan
15
+ level 2 Intermediate
16
+ access claude-code, codex, grok
17
+ primary claude-code
18
+ tools codecalc
19
+ folder ./ai-orchestrator
20
+ project . (11 subagent files go here)
21
+ files 38
22
+ - ROUTING.md multi-lane decision tree
23
+ - TIERS.md DELEGATION_MATRIX.md which lane, at what effort
24
+ - TASK_BUNDLE.md the brief every delegation carries
25
+ - protocols/ build, propagate, gap-analysis, deep-research, and three more
26
+ - [project] .claude/agents/ builder, deep-planner, code-reviewer, bulk-worker,
27
+ live-researcher, reader, finding-verifier, done-verifier
28
+ - [project] .claude/hooks/ route-gate, subagent-context, route-metrics
29
+ - bin/cli-run.mjs bin/lanes.json the lane runner
30
+
31
+ --dry: nothing written.
32
+ ```
26
33
 
27
- 1. **Which level?** 1 beginner · 2 intermediate · 3 advanced
28
- 2. **Which AIs do you have access to?** (it marks the ones already on your PATH)
29
- 3. **Which one is your primary agent?** (the one that runs the system)
34
+ Reproduce it:
30
35
 
31
- It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions (and, on Claude Code, two hook scripts) go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: follow the generated README. Inspect the manifest and remove only the individual managed subagent files you no longer need, preserve edited or pre-existing files, and remove your manually pasted activation block. Never delete a shared subagent folder.
36
+ ```bash
37
+ npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --dry
38
+ ```
32
39
 
33
- ## Plans and automatic effort
40
+ - **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
41
+ - **What it is not:** a proxy, a gateway or an API router. It does not automatically compare prices or select models; your agent follows the rules and chooses.
42
+ - **Use it when:** you run more than one model or agent and want the expensive tier kept for planning and judgment.
43
+ - **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
34
44
 
35
- State known subscription plans with `--plans codex=pro-20x,agy=ultra-5x`. The generated guidance uses plan headroom to allocate volume only. It never changes capability or independent-review rules. `--effort-auto` is explicit consent to set `auto` only for selected high or max headroom CLI lanes. Auto chooses medium below 4,000 prompt characters and high otherwise, never higher. A codex audit is always high. Name `xhigh` explicitly for security-critical or irreversible work.
45
+ Built from a working system, not a diagram: the routing rules, the protocols and the lane runner here run in production, generalized so they transfer to any stack.
36
46
 
37
47
  ## The three levels
38
48
 
49
+
39
50
  | Level | You have | You get |
40
51
  |---|---|---|
41
52
  | **1 · Beginner** | one LLM or one agent | tiers, task classification, the two build checkpoints, the protocols (build, propagate, gap analysis, deep research, numbers and logic, memory and record, docs and proof), a task-bundle template, and your agent set up to follow them |
42
53
  | **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts; `--model` / `--effort` to pin the route and log it), a delegation matrix generated from your selection, research triage across the lanes you have |
43
54
  | **3 · Advanced** | a virtual machine | everything above, plus a gateway config rendered from the API keys you hold (asked separately from your CLIs), pinned images, box rules, privacy gates, and a weekly gap-analysis job with "what watches it" written down |
44
55
 
45
- Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md).
46
-
56
+ Levels explained: [Part 1](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md).
47
57
  ## The AIs it knows about
48
58
 
49
59
  | Id | What | Level |
@@ -59,134 +69,39 @@ Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/pa
59
69
 
60
70
  `npx model-orchestrator --list` prints the catalog with install and sign-in notes. Details: [docs/catalog.md](docs/catalog.md).
61
71
 
62
- ## Companion tools (all optional)
63
-
64
- An orchestrator routes work. It does not make a model stop guessing numbers, it does not give it a memory, and it does not make it check a library's current docs before writing against it. Three tools close those gaps: codecalc and obsidian-tc are from the same maintainer, Context7 is from Upstash. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc,context7` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
65
-
66
- | Tool | Closes | Default | You need first |
67
- |---|---|---|---|
68
- | [codecalc](https://github.com/The-40-Thieves/codecalc) | guessed numbers, comparisons, complexity and equivalence claims: exact arithmetic, code execution in 31 languages, SMT logic checks, `verify_translation` / `verify_optimization`; offline, no key | yes | Python 3.10+ and `uv`. `uvx 'codecalc[full]' setup --write` registers it with Claude Code, Claude Desktop, Cursor, VS Code, Zed; snippets for Codex, Antigravity, Qwen Code are written for you |
69
- | [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes with a confirmation gate, folder ACLs, a poison scan on inferred writes; 163 tools, local by default; AGPL-3.0 | no | an Obsidian vault folder; Node 24+ or Bun 1.1+ (stricter than this installer); Ollama with `nomic-embed-text` or a cloud embeddings key; the Obsidian app and its Local REST API plugin only for live bridge tools. Skip it if you do not keep notes in Obsidian |
70
- | [Context7](https://github.com/upstash/context7) | stale library recall: current, version-specific docs and code examples pulled into the prompt for any library, SDK, API or CLI; hosted, or `npx` locally; MIT | no | nothing to install for the hosted endpoint; Node.js 18+ for the local alternative; a free API key is optional, for a higher rate limit. Always makes a network call, unlike the other two: skip it offline |
71
-
72
- Context7 pairs with codecalc rather than duplicating it: Context7 tells the agent what a library is documented to do on this version, codecalc runs the code and proves what it actually does. Docs never stand as proof on their own, and where the two disagree the run wins.
73
-
74
- Whether or not you select them, every level carries the three rules they serve: `protocols/numbers-and-logic.md` (when calling a calculator is mandatory, how to report a computed figure, why a thought log is not evidence), `protocols/memory-and-record.md` (search before writing, the folder index is part of the change, one writer, inferred content marked as inferred), and `protocols/docs-then-prove.md` (current docs before writing a call, then a run proves it, the run wins on disagreement).
75
-
76
- ## The two folders every run writes to
77
-
78
- An install has two targets, and a scripted run should set both.
79
-
80
- | Flag | Default | What lands there |
81
- |---|---|---|
82
- | `--dir` | `./ai-orchestrator` | the docs, protocols and (level 2+) `bin/cli-run.mjs`. Named after what it contains, not after this package, so a project can hold one without looking like a checkout of it. Pass `--dir ./model-orchestrator` if you prefer the package name. |
83
- | `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets three hook scripts in `.claude/hooks/`, wired by a settings snippet you merge yourself. |
84
-
85
- `--project` defaulting to the current directory is the one that surprises people: run the command from your home folder with Claude Code as the primary and five agent files land in your home folder. The installer prints the resolved project path in the plan and says when you left it at the default. Set it.
72
+ ## Measuring routing
86
73
 
87
- ## Non-interactive
74
+ A routing rule nobody measures is a rule nobody knows is followed. On a claude-code install, `route-metrics.mjs` turns every turn, dispatch and subagent start/stop into one JSON line under `~/.ai-orchestrator/route-metrics.jsonl`, including the lane your agent named in its own `<!-- route: <lane> | <why> -->` marker.
88
75
 
89
76
  ```bash
90
- # both targets set: docs in ./ai-orchestrator, subagents into ./my-app/.claude/agents
91
- npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code \
92
- --dir ./ai-orchestrator --project ./my-app
93
- # --yes selects the recommended companion tool (codecalc), which writes CODECALC.md and mcp/ snippets.
94
- # Add --no-tools for none, or --tools codecalc,obsidian-tc,context7 to choose.
95
-
96
- npx model-orchestrator --yes --level 3 --ais claude-code,codex,agy,grok,hermes,qwen,ollama --apis anthropic,openrouter --dry # print the plan, write nothing
97
- npx model-orchestrator --yes --level 2 --ais claude-code,codex --project ~/my-app --dir ~/my-app/ai-orchestrator --no-tools # subagents into ~/my-app/.claude/agents
98
- npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --update-docs # added a lane: regenerate the docs you never edited
77
+ node .claude/hooks/route-metrics.mjs --summary # since the log began
78
+ node .claude/hooks/route-metrics.mjs --summary --since 2026-09-01 # since a date
99
79
  ```
100
80
 
81
+ The report prints turns, **route-marker coverage** (the share of turns that carried a real lane, which answers "is the agent actually tagging its routing decisions?"), lanes by count, dispatches by `subagent_type`, dispatches with no matching start, and mean/max duration per agent type. It never logs prompt text, tool descriptions, or the "why" half of the marker. Fail-open by design: a miss is a missing log line, never a blocked turn.
82
+
101
83
  ## Claude Code plugin
102
84
 
103
- The Claude Code hooks and subagents also ship as a plugin, so they install and update through Claude Code itself instead of a settings snippet you merge by hand:
85
+ The hooks and subagents also ship as a plugin, so they install and update through Claude Code itself:
104
86
 
105
87
  ```
106
88
  /plugin marketplace add aunysillyme/model-orchestrator
107
89
  /plugin install model-orchestrator@model-orchestrator
108
90
  ```
109
91
 
110
- - **What it ships:** `route-gate.mjs` (UserPromptSubmit, plus a one-line notice at session start when the project has no rules), `subagent-context.mjs` (SubagentStart), and the eight subagents, each with an explicit tool list. Agents load namespaced, as `model-orchestrator:builder`.
111
- - **What it still needs from the installer:** the routing rules. A plugin runs no install step, so it reads the installer's default locations, `ai-orchestrator/ROUTING.md` then `ai-orchestrator/ORCHESTRATOR.md`, and names `npx model-orchestrator` when neither exists. A project installed with a different `--dir` should wire the installer's rendered hooks instead.
112
- - **What it leaves out:** `route-metrics.mjs`, the routing log. It writes to disk and the plugin ships only hooks that read, so `npx model-orchestrator` is how you get it.
113
- - **How it is kept honest:** `plugin/` is generated from `templates/` by `npm run gen:plugin`, and `test/plugin.test.js` fails when the committed bundle drifts, when a hook gains a network call, a write or a subprocess, or when an agent loses its tool list. The bundle passes `claude plugin validate --strict`, the check Anthropic's community marketplace review runs on every submission.
92
+ It ships the three hooks and the eight subagents, each with an explicit tool list, and loads them namespaced as `model-orchestrator:builder`. It does not ship the routing rules, because a plugin runs no install step: `npx model-orchestrator` still writes those. `plugin/` is generated from `templates/` and `test/plugin.test.js` fails when the committed bundle drifts, when a hook gains a network call, a write or a subprocess, or when an agent loses its tool list. Details: [plugin/README.md](plugin/README.md).
114
93
 
115
- Details, including running the plugin next to an npm install: [plugin/README.md](plugin/README.md).
116
-
117
- ## What gets written (level 3, everything)
118
-
119
- ```
120
- ai-orchestrator/
121
- README.md start here, written for your level and your AIs
122
- ORCHESTRATOR.md single-agent routing rules (level 1)
123
- TASK_BUNDLE.md the brief every delegation carries
124
- protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record · docs-then-prove
125
- CODECALC.md OBSIDIAN-TC.md CONTEXT7.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
126
- <project>/.claude/agents/ one per tier plus finding-verifier, done-verifier, reader, at the PROJECT root (if Claude Code is primary)
127
- <project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart) + route-metrics.mjs (all five: see "Measuring routing" below), Claude Code only
128
- CLAUDE.snippet.md the block to paste into your CLAUDE.md
129
- settings.hooks.snippet.json the hooks block to merge into .claude/settings.json (Claude Code only)
130
- ROUTING.md multi-lane decision tree (level 2+)
131
- TIERS.md DELEGATION_MATRIX.md RESEARCH_TRIAGE.md CLI-RUN.md
132
- bin/cli-run.mjs bin/lanes.json (node bin/cli-run.mjs --doctor is the smoke test)
133
- vm/ gateway config, compose, box rules, privacy gates, jobs/ (level 3)
134
- ```
135
-
136
- ## Repo layout
137
-
138
- | Folder | What |
139
- |---|---|
140
- | [`bin/`](bin/README.md) | `cli.js` (the installer) and `cli-run.mjs` (the lane runner) |
141
- | [`src/`](src/README.md) | the catalog, the pure planner, detection, rendering |
142
- | [`templates/`](templates/README.md) | everything the installer can write, by level, plus `tools/` for companions |
143
- | [`docs/`](docs/README.md) | the three parts and the catalog |
144
- | [`plugin/`](plugin/README.md) | the Claude Code plugin, generated from `templates/` by `npm run gen:plugin`; `.claude-plugin/marketplace.json` at the root lists it |
145
- | [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu, macOS and Windows, Node 18/20/22 |
146
-
147
- ## What is enforced, what is delegated, what is an instruction
148
-
149
- Most of what this package ships is text an agent is asked to follow. Be clear about which is which before relying on it unattended.
150
-
151
- | Property | How it holds |
152
- |---|---|
153
- | Installer writes only inside `--dir` and `--project`, never a secret, never over a document without `--force` (or `--update-docs`, which touches only documents provably untouched since a previous run); machine-owned config always, runtime files only when provably untouched or with `--upgrade-runtime` | **enforced by code** (preflight, exclusive create, rollback, manifest hashes; tested) |
154
- | `cli-run` exit codes, process-group kill on timeout and on SIGINT/SIGTERM, UTF-8-safe streaming, fixed-code durable log, `--expect-*` contracts with a pre-run snapshot | **enforced by code** (tested with stub lanes) |
155
- | Codex audit lane runs read-only | **delegated to the vendor flag** (`--audit` → `--sandbox read-only`); commands and network still follow your codex config |
156
- | Other lanes' permissions, sign-in state, model versions | **delegated to each vendor's own config**; `--doctor` checks presence, not versions |
157
- | Gateway binds to loopback, keys by name only | **enforced in the generated files**; whether the gateway authenticates is your environment |
158
- | Lane selection, tiers, privacy classes, one-writer, escalation, the protocols | **agent instructions**. Nothing here stops an agent that ignores its rules; the task bundle and the protocols make ignoring them visible, not impossible |
159
- | Weekly audit bounded, previous report preserved | **enforced in the generated script and unit** (watchdog, temp-and-rename, `TimeoutStartSec`) |
160
-
161
- If you need a property in the third row to be enforced, that is a router, a policy engine or a sandbox, and this package does not claim to be one.
162
-
163
- ### Vendor version compatibility
164
-
165
- **This package detects that a binary exists. It does not check its version, and a present binary is not a working lane.** `--doctor` reports presence, and with `--run` sends one lane a one-word canary; neither validates that the vendor's flags, output shape or auth still match what the generated files assume.
166
-
167
- The lane wiring and the output judges were written against these versions, which are the ones this release was exercised on:
168
-
169
- <!-- vendor-table:start -->
170
-
171
- | Lane | Vendor | Version this release was built against | Where that number is proved |
172
- |---|---|---|---|
173
- | `claude` | Anthropic | 2.1.226 | the npm pin the installer writes, `@anthropic-ai/claude-code@2.1.226` |
174
- | `codex` | OpenAI | 0.153.4 | `test/fixtures/codex-0.153.4.jsonl`, a recorded run |
175
- | `agy` | Google | 1.1.27 | `test/fixtures/agy-1.1.27.jsonl`, a recorded run |
176
- | `grok` | xAI | 1.0.5 | `test/fixtures/grok-1.0.5.json`, a recorded run |
177
- | `hermes` | Nous Research | 0.20.0 | `test/fixtures/hermes-0.20.0.txt`, a recorded run |
178
- | `qwen` | Alibaba | 0.22.3 | `test/fixtures/qwen-0.22.3-nokey.json`, a recorded run |
179
- | `ollama` | Ollama | 0.33.3 | the pinned image the level 3 box runs, `ollama/ollama:0.33.3` |
180
-
181
- Generated from `src/catalog.js` by `npm run gen:catalog`; `npm test` fails if this table and the catalog disagree. Fixtures were captured 2026-09-06.
182
-
183
- <!-- vendor-table:end -->
94
+ ## Companion tools (all optional)
184
95
 
185
- One number per lane, and it is the same number the installer pins: where a lane installs from npm, `builtAgainst` in the catalog *is* the pin, so "built against" and "pinned to" can never be two answers. That pin is a floor, not a ceiling: these CLIs ship breaking flag changes on their own schedules, so a newer version may work perfectly, or may change a flag the generated wiring passes. When a lane starts failing after a vendor upgrade, compare against this table first.
96
+ An orchestrator routes work. It does not make a model stop guessing numbers, give it a memory, or make it check a library's current docs before writing against it. Three tools close those gaps:
186
97
 
187
- **The live canary runs on your machine, with your credentials.** That is what `node bin/cli-run.mjs --doctor --run` is: it sends every enabled lane one tiny prompt through your own sign-ins and reports `canary ok` or `canary FAILED rc=` per lane. Run it after install, and again after any vendor upgrade.
98
+ | Tool | Closes | Default |
99
+ |---|---|---|
100
+ | [codecalc](https://github.com/The-40-Thieves/codecalc) | guessed numbers: exact arithmetic, code execution in 31 languages, logic checks; offline, no key | yes |
101
+ | [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes; local by default | no |
102
+ | [Context7](https://github.com/upstash/context7) | stale library recall: current, version-specific docs pulled into the prompt | no |
188
103
 
189
- It deliberately does not run in this repository's CI. A canary is only meaningful against real credentials, and there are no credentials a maintainer could supply that would tell **you** anything about **your** lanes: your sign-ins, your quota, your vendor versions. A maintainer-credential canary in CI would prove one machine works and bill someone per run to do it. So CI runs the full suite against stub lanes on Ubuntu, macOS and Windows, Node 18/20/22, plus a packaged install into a clean consumer, and the live check ships to you instead.
104
+ Selecting one writes a doc and config snippets; it installs nothing. What each needs first, and why Context7 pairs with codecalc rather than duplicating it: [docs/companions.md](docs/companions.md). Whether or not you select them, every level carries the three rules they serve: `protocols/numbers-and-logic.md`, `protocols/memory-and-record.md` and `protocols/docs-then-prove.md`.
190
105
 
191
106
  ## Principles the whole thing rests on
192
107
 
@@ -198,101 +113,83 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
198
113
  6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, only the brief carries what this task needs. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
199
114
  7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
200
115
 
201
- ## Routing by role, complexity and stakes
116
+ ## Common questions
202
117
 
203
- Role picks the agent. Two more inputs move the choice, and they move it in
204
- different directions, so `TIERS.md` states them separately rather than folding
205
- them into the role:
118
+ ### How do I cut token usage across Claude Code, Codex and Gemini?
206
119
 
207
- - **Complexity moves the effort.** A worker executing a finished plan needs less
208
- reasoning than the reviewer judging its output. When the plan is airtight the
209
- spec is carrying the thinking.
210
- - **Stakes move the tier and the reader.** Security, privacy, data loss and
211
- irreversible changes buy the challenge lane, a named check, a rollback path or
212
- a human yes. A one-line change to an auth check is simple and high-stakes at
213
- the same time, and it is the stakes that decide.
120
+ Install for the tools you have, then let the generated `ROUTING.md` decide the tier per task: bulk, reading and verification go to the fast tier or a cheaper CLI lane, and the deep tier only plans and judges. On Claude Code, execution goes to the `builder` subagent by default and the main session plans and verifies. Every lane call through `cli-run` logs the model and effort it ran with, so you can check where the tokens went.
214
121
 
215
- Stakes means what a mistake would cost: a security hole, leaked personal data,
216
- lost data, or something you can't undo. Most tasks are low-stakes and route
217
- normally.
122
+ ### How do I route tasks to cheaper models?
218
123
 
219
- The top of the ladder is bought with evidence: a reproduced failure, an
220
- unresolved checkpoint, an irreversible change. A task that merely feels hard is
221
- a deep-tier task, not an escalation.
124
+ The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly never gets a frontier token.
222
125
 
223
- ## A finding is a claim, not a fact
126
+ ### Is this an LLM router or an AI gateway?
224
127
 
225
- Review findings do not go straight to a repair. `finding-verifier` reads the
226
- cited line, states what would trigger the problem, then hunts for the guard,
227
- caller or test that makes it impossible, and returns **CONFIRMED**,
228
- **NOT_REPRODUCED** or **INCONCLUSIVE** per finding. Only CONFIRMED earns a
229
- change. Use a different model family from the one that produced the finding
230
- where you have one: a family asked to check its own claim tends to agree with
231
- itself.
128
+ No. It routes at the task level, through instructions your agent follows and a runner for agent CLIs. If you want a service or proxy that picks or forwards the model on every API request, look at request-level routers and gateways such as RouteLLM, LiteLLM, OpenRouter or claude-code-router. They solve a different problem and can sit underneath this.
232
129
 
233
- ## Two more fast-tier checks
130
+ ### Can an agent install and run it without a person?
234
131
 
235
- `done-verifier` probes the artifact a tracker item's done-signal names (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything itself. It carries no file-editing tools, but on claude-code it does carry `Bash` for those probes (`git log`, `grep`, `wc -l`, `test -f`); staying to read-only commands there is a rule in its prompt, not a restriction on the tool grant, and its own description says so. On agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `reader` is the one that is read-only by tool grant on both: no `Write`, `Edit`, or `Bash`. It reads and digests many files or notes and hands back exactly what the brief asked for, cited by `path:line`; it never classifies, tags or writes, which is what separates it from `bulk-worker`. Both ship in the claude-code and agy agent sets, at the fast tier.
132
+ Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run` previews the plan, and `--list` prints every supported AI. Nothing is appended to a file you already have; activation snippets are written next to your files for you to merge.
236
133
 
237
- ## Measuring routing
238
134
 
239
- A routing rule nobody measures is a rule nobody knows is followed. On a claude-code install, `route-metrics.mjs` (the third hook, wired to `UserPromptSubmit`, `PreToolUse` on `Agent`/`Task`, `SubagentStart`, `SubagentStop` and `Stop`) turns each of those into one JSON line under `~/.ai-orchestrator/route-metrics.jsonl`: a turn started, a subagent was dispatched (and with what, and in the background or not), a subagent started and stopped (so a duration can be computed), and the lane your agent named in its own hidden `<!-- route: <lane> | <why> -->` marker, which the route-gate block now asks for on every reply. It never logs prompt text, tool descriptions, or the "why" half of the marker: only the named fields above, charset-bounded, same principle as `cli-run.mjs`'s log.
135
+ ## Read next
240
136
 
241
- ```bash
242
- node .claude/hooks/route-metrics.mjs --summary # since the log began
243
- node .claude/hooks/route-metrics.mjs --summary --since 2026-09-01 # since a date
244
- ```
137
+ | Doc | What is in it |
138
+ |---|---|
139
+ | [docs/install.md](docs/install.md) | every flag, the two folders a run writes to, headless examples, the full file list |
140
+ | [docs/how-it-routes.md](docs/how-it-routes.md) | role, complexity and stakes; the three verifier agents; pinning a lane's model and effort |
141
+ | [docs/guarantees.md](docs/guarantees.md) | what is enforced by code, what is delegated to a vendor flag, and what is only an instruction |
142
+ | [docs/part-1-beginner.md](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md) | the thinking behind each level |
143
+ | [docs/catalog.md](docs/catalog.md) | every supported AI with install and sign-in notes |
245
144
 
246
- The report prints turns, **route-marker coverage** (the percentage of turns whose `Stop` event carried a real lane, not `missing`, which is the number that answers "is the agent actually tagging its routing decisions?"), lanes by count, dispatches by `subagent_type`, dispatches with no matching start (a hook or guard blocked the subagent before it launched), and mean/max duration per agent type. Fail-open by design, like the other two hooks: a miss here is a missing log line, never a blocked turn, and it prints nothing to stdout on any event since stdout on `UserPromptSubmit`/`SubagentStart` becomes model context.
145
+ <details>
146
+ <summary><strong>Platform support, and every test this suite skips</strong></summary>
247
147
 
248
- ## Pin the route, or know that you did not
249
148
 
250
- A lane with no `--model`, no `--effort` and no `defaults` entry in
251
- `bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
252
- CLI configured months ago at a low reasoning effort keeps auditing at that
253
- effort while your routing docs describe a second-opinion pass.
149
+ Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22), including lane execution end to end through `cli-run` against a fake CLI installed the same way npm installs a real one (a `.cmd` shim). `cli-run` never runs a lane through `cmd.exe` when it can avoid it: it resolves the shim to the Node script underneath and spawns Node directly, so a prompt reaching a real lane never passes through a Windows shell. A `.cmd` or `.bat` lane that cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, rather than run through `cmd.exe`: a batch file re-reads its arguments after `cmd.exe` has parsed them once, and no escaping fully contains a prompt through both passes. Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested on Windows too, including SIGTERM/SIGINT to the wrapper (Windows has no OS-level signals: both terminate it unconditionally, verified there rather than treated the same as POSIX). Five narrow skips remain on Windows, each for a POSIX behavior the OS or the CI shell genuinely does not have, and each named here because a test that is quietly skipped reads as a test that passed: `statSync().mode`'s executable bit (NTFS has none, so that one assertion is conditional inside a test that otherwise runs everywhere); a lane dying mid-run from a real POSIX signal (a real Windows lane cannot die "by signal"); running `weekly-audit.sh`'s watchdog functions for real under Git Bash's job control, both the end-to-end run and the `bounded()` timeout check (the script itself only ever runs on the Ubuntu box it targets); and a `mkfifo` FIFO at the rules path, the one case that proves `route-gate.mjs` cannot HANG on a non-regular file, since Windows has no `mkfifo` to build one (the guard behind it is covered on every OS by a directory at the same path); and an untracked `mkfifo` FIFO in the repository `cli-run --audit` sizes, the case that proves `--effort auto` never opens a non-regular file (the symlink half of that test runs on every OS). The list is not prose on trust: `test/prose.test.js` counts every `skip:` in the suite and fails if one of them is not documented here.
254
150
 
255
- ```bash
256
- node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
257
- node bin/cli-run.mjs --doctor # prints what each lane is pinned to, and what is not pinned
258
- ```
151
+ **Privacy.** The installer sends no telemetry and makes no network call of its own once it is running. Two things around that are worth being exact about:
259
152
 
260
- Every run logs the model and effort **requested** and where the request came
261
- from: `flag`, `lanes.json`, or `lane_default`, on every record including the
262
- runs that never reached a lane. It does not log an actual. One lane of five
263
- (grok) reports a model id in its own output and the other four report none, so
264
- an actual field would be present for one lane and missing for four, and it
265
- would be a provider-supplied string, which the durable log deliberately never
266
- holds.
153
+ - `npx model-orchestrator` is itself a download: npm fetches this package from the registry before any of it runs. `npm install -g model-orchestrator` once, then run `model-orchestrator`, if you would rather that happen exactly one time.
154
+ - A missing vendor CLI is *printed*, not installed. In an interactive run the installer offers to run one pinned `npm install -g` per package and only runs the ones you answer yes to; with `--yes` or `--no-install` it answers no for you and prints the command instead. Vendor shell installers (Antigravity, Grok) are only ever printed, alongside the `curl … | less` you would use to read one before running it.
267
155
 
268
- ## Common questions
156
+ `cli-run` talks to nothing but the vendor CLI you name.
269
157
 
270
- ### How do I cut token usage across Claude Code, Codex and Gemini?
271
158
 
272
- Install for the tools you have, then let the generated `ROUTING.md` decide the tier per task: bulk, reading and verification go to the fast tier or a cheaper CLI lane, and the deep tier only plans and judges. On Claude Code, execution goes to the `builder` subagent by default and the main session plans and verifies. Every lane call through `cli-run` logs the model and effort it ran with, so you can check where the tokens went.
159
+ </details>
273
160
 
274
- ### How do I route tasks to cheaper models?
161
+ <details>
162
+ <summary><strong>Vendor version compatibility</strong></summary>
275
163
 
276
- The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly never gets a frontier token.
277
164
 
278
- ### Is this an LLM router or an AI gateway?
165
+ **This package detects that a binary exists. It does not check its version, and a present binary is not a working lane.** `--doctor` reports presence, and with `--run` sends one lane a one-word canary; neither validates that the vendor's flags, output shape or auth still match what the generated files assume.
279
166
 
280
- No. It routes at the task level, through instructions your agent follows and a runner for agent CLIs. If you want a service or proxy that picks or forwards the model on every API request, look at request-level routers and gateways such as RouteLLM, LiteLLM, OpenRouter or claude-code-router. They solve a different problem and can sit underneath this.
167
+ The lane wiring and the output judges were written against these versions, which are the ones this release was exercised on:
281
168
 
282
- ### Can an agent install and run it without a person?
169
+ <!-- vendor-table:start -->
283
170
 
284
- Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run` previews the plan, and `--list` prints every supported AI. Nothing is appended to a file you already have; activation snippets are written next to your files for you to merge.
171
+ | Lane | Vendor | Version this release was built against | Where that number is proved |
172
+ |---|---|---|---|
173
+ | `claude` | Anthropic | 2.1.226 | the npm pin the installer writes, `@anthropic-ai/claude-code@2.1.226` |
174
+ | `codex` | OpenAI | 0.153.4 | `test/fixtures/codex-0.153.4.jsonl`, a recorded run |
175
+ | `agy` | Google | 1.1.27 | `test/fixtures/agy-1.1.27.jsonl`, a recorded run |
176
+ | `grok` | xAI | 1.0.5 | `test/fixtures/grok-1.0.5.json`, a recorded run |
177
+ | `hermes` | Nous Research | 0.20.0 | `test/fixtures/hermes-0.20.0.txt`, a recorded run |
178
+ | `qwen` | Alibaba | 0.22.3 | `test/fixtures/qwen-0.22.3-nokey.json`, a recorded run |
179
+ | `ollama` | Ollama | 0.33.3 | the pinned image the level 3 box runs, `ollama/ollama:0.33.3` |
285
180
 
286
- ## Requirements
181
+ Generated from `src/catalog.js` by `npm run gen:catalog`; `npm test` fails if this table and the catalog disagree. Fixtures were captured 2026-09-06.
287
182
 
288
- Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22), including lane execution end to end through `cli-run` against a fake CLI installed the same way npm installs a real one (a `.cmd` shim). `cli-run` never runs a lane through `cmd.exe` when it can avoid it: it resolves the shim to the Node script underneath and spawns Node directly, so a prompt reaching a real lane never passes through a Windows shell. A `.cmd` or `.bat` lane that cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, rather than run through `cmd.exe`: a batch file re-reads its arguments after `cmd.exe` has parsed them once, and no escaping fully contains a prompt through both passes. Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested on Windows too, including SIGTERM/SIGINT to the wrapper (Windows has no OS-level signals: both terminate it unconditionally, verified there rather than treated the same as POSIX). Five narrow skips remain on Windows, each for a POSIX behavior the OS or the CI shell genuinely does not have, and each named here because a test that is quietly skipped reads as a test that passed: `statSync().mode`'s executable bit (NTFS has none, so that one assertion is conditional inside a test that otherwise runs everywhere); a lane dying mid-run from a real POSIX signal (a real Windows lane cannot die "by signal"); running `weekly-audit.sh`'s watchdog functions for real under Git Bash's job control, both the end-to-end run and the `bounded()` timeout check (the script itself only ever runs on the Ubuntu box it targets); and a `mkfifo` FIFO at the rules path, the one case that proves `route-gate.mjs` cannot HANG on a non-regular file, since Windows has no `mkfifo` to build one (the guard behind it is covered on every OS by a directory at the same path); and an untracked `mkfifo` FIFO in the repository `cli-run --audit` sizes, the case that proves `--effort auto` never opens a non-regular file (the symlink half of that test runs on every OS). The list is not prose on trust: `test/prose.test.js` counts every `skip:` in the suite and fails if one of them is not documented here.
183
+ <!-- vendor-table:end -->
289
184
 
290
- **Privacy.** The installer sends no telemetry and makes no network call of its own once it is running. Two things around that are worth being exact about:
185
+ One number per lane, and it is the same number the installer pins: where a lane installs from npm, `builtAgainst` in the catalog *is* the pin, so "built against" and "pinned to" can never be two answers. That pin is a floor, not a ceiling: these CLIs ship breaking flag changes on their own schedules, so a newer version may work perfectly, or may change a flag the generated wiring passes. When a lane starts failing after a vendor upgrade, compare against this table first.
291
186
 
292
- - `npx model-orchestrator` is itself a download: npm fetches this package from the registry before any of it runs. `npm install -g model-orchestrator` once, then run `model-orchestrator`, if you would rather that happen exactly one time.
293
- - A missing vendor CLI is *printed*, not installed. In an interactive run the installer offers to run one pinned `npm install -g` per package and only runs the ones you answer yes to; with `--yes` or `--no-install` it answers no for you and prints the command instead. Vendor shell installers (Antigravity, Grok) are only ever printed, alongside the `curl … | less` you would use to read one before running it.
187
+ **The live canary runs on your machine, with your credentials.** That is what `node bin/cli-run.mjs --doctor --run` is: it sends every enabled lane one tiny prompt through your own sign-ins and reports `canary ok` or `canary FAILED rc=` per lane. Run it after install, and again after any vendor upgrade.
294
188
 
295
- `cli-run` talks to nothing but the vendor CLI you name.
189
+ It deliberately does not run in this repository's CI. A canary is only meaningful against real credentials, and there are no credentials a maintainer could supply that would tell **you** anything about **your** lanes: your sign-ins, your quota, your vendor versions. A maintainer-credential canary in CI would prove one machine works and bill someone per run to do it. So CI runs the full suite against stub lanes on Ubuntu, macOS and Windows, Node 18/20/22, plus a packaged install into a clean consumer, and the live check ships to you instead.
190
+
191
+
192
+ </details>
296
193
 
297
194
  ## Contributing
298
195
 
@@ -300,11 +197,7 @@ Add an AI to `src/catalog.js` and every prompt, table, config and doc picks it u
300
197
 
301
198
  ## Credits
302
199
 
303
- - [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for
304
- separating role, complexity and stakes instead of compressing them into one
305
- scale, for recording the model and effort a lane was actually asked for, and
306
- for verifying findings before they trigger repairs. All three shipped in
307
- 0.1.14.
200
+ - [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for separating role, complexity and stakes instead of compressing them into one scale, for recording the model and effort a lane was actually asked for, and for verifying findings before they trigger repairs. All three shipped in 0.1.14.
308
201
 
309
202
  ## License
310
203
 
package/docs/README.md CHANGED
@@ -1,6 +1,13 @@
1
1
  # docs/
2
2
 
3
- The three parts, as reading. The installer writes the working files; these explain the thinking behind them and how to grow from one level to the next.
3
+ Reference, then reading. The reference pages carry the detail the README links out to; the three parts explain the thinking behind each level and how to grow from one to the next.
4
+
5
+ | Reference | What is in it | File |
6
+ |---|---|---|
7
+ | Installing | every flag, the two folders a run writes to, headless examples, the full file list | [install.md](install.md) |
8
+ | How it routes | role, complexity and stakes; the three verifier agents; pinning model and effort | [how-it-routes.md](how-it-routes.md) |
9
+ | Guarantees | what is enforced by code, delegated to a vendor flag, or only an instruction | [guarantees.md](guarantees.md) |
10
+ | Companion tools | codecalc, obsidian-tc and Context7: what each closes and what it needs first | [companions.md](companions.md) |
4
11
 
5
12
  | Part | Read if | File |
6
13
  |---|---|---|
@@ -0,0 +1,17 @@
1
+ # Companion tools
2
+
3
+ All optional. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc,context7` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
4
+
5
+ ## Companion tools (all optional)
6
+
7
+ An orchestrator routes work. It does not make a model stop guessing numbers, it does not give it a memory, and it does not make it check a library's current docs before writing against it. Three tools close those gaps: codecalc and obsidian-tc are from the same maintainer, Context7 is from Upstash. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc,context7` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
8
+
9
+ | Tool | Closes | Default | You need first |
10
+ |---|---|---|---|
11
+ | [codecalc](https://github.com/The-40-Thieves/codecalc) | guessed numbers, comparisons, complexity and equivalence claims: exact arithmetic, code execution in 31 languages, SMT logic checks, `verify_translation` / `verify_optimization`; offline, no key | yes | Python 3.10+ and `uv`. `uvx 'codecalc[full]' setup --write` registers it with Claude Code, Claude Desktop, Cursor, VS Code, Zed; snippets for Codex, Antigravity, Qwen Code are written for you |
12
+ | [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes with a confirmation gate, folder ACLs, a poison scan on inferred writes; 163 tools, local by default; AGPL-3.0 | no | an Obsidian vault folder; Node 24+ or Bun 1.1+ (stricter than this installer); Ollama with `nomic-embed-text` or a cloud embeddings key; the Obsidian app and its Local REST API plugin only for live bridge tools. Skip it if you do not keep notes in Obsidian |
13
+ | [Context7](https://github.com/upstash/context7) | stale library recall: current, version-specific docs and code examples pulled into the prompt for any library, SDK, API or CLI; hosted, or `npx` locally; MIT | no | nothing to install for the hosted endpoint; Node.js 18+ for the local alternative; a free API key is optional, for a higher rate limit. Always makes a network call, unlike the other two: skip it offline |
14
+
15
+ Context7 pairs with codecalc rather than duplicating it: Context7 tells the agent what a library is documented to do on this version, codecalc runs the code and proves what it actually does. Docs never stand as proof on their own, and where the two disagree the run wins.
16
+
17
+ Whether or not you select them, every level carries the three rules they serve: `protocols/numbers-and-logic.md` (when calling a calculator is mandatory, how to report a computed figure, why a thought log is not evidence), `protocols/memory-and-record.md` (search before writing, the folder index is part of the change, one writer, inferred content marked as inferred), and `protocols/docs-then-prove.md` (current docs before writing a call, then a run proves it, the run wins on disagreement).
@@ -0,0 +1,18 @@
1
+ # What holds, and what is only asked for
2
+
3
+ Read this before running any of it unattended.
4
+
5
+ Most of what this package ships is text an agent is asked to follow. Be clear about which is which before relying on it unattended.
6
+
7
+ | Property | How it holds |
8
+ |---|---|
9
+ | Installer writes only inside `--dir` and `--project`, never a secret, never over a document without `--force` (or `--update-docs`, which touches only documents provably untouched since a previous run); machine-owned config always, runtime files only when provably untouched or with `--upgrade-runtime` | **enforced by code** (preflight, exclusive create, rollback, manifest hashes; tested) |
10
+ | `cli-run` exit codes, process-group kill on timeout and on SIGINT/SIGTERM, UTF-8-safe streaming, fixed-code durable log, `--expect-*` contracts with a pre-run snapshot | **enforced by code** (tested with stub lanes) |
11
+ | Codex audit lane runs read-only | **delegated to the vendor flag** (`--audit` → `--sandbox read-only`); commands and network still follow your codex config |
12
+ | Other lanes' permissions, sign-in state, model versions | **delegated to each vendor's own config**; `--doctor` checks presence, not versions |
13
+ | Gateway binds to loopback, keys by name only | **enforced in the generated files**; whether the gateway authenticates is your environment |
14
+ | Lane selection, tiers, privacy classes, one-writer, escalation, the protocols | **agent instructions**. Nothing here stops an agent that ignores its rules; the task bundle and the protocols make ignoring them visible, not impossible |
15
+ | Weekly audit bounded, previous report preserved | **enforced in the generated script and unit** (watchdog, temp-and-rename, `TimeoutStartSec`) |
16
+
17
+ If you need a property in the third row to be enforced, that is a router, a policy engine or a sandbox, and this package does not claim to be one.
18
+
@@ -0,0 +1,60 @@
1
+ # How it routes
2
+
3
+ The reasoning behind the lane choice, and the three verifier agents that keep a routed answer honest. Install mechanics are in [install.md](install.md).
4
+
5
+ ## Routing by role, complexity and stakes
6
+
7
+ Role picks the agent. Two more inputs move the choice, and they move it in
8
+ different directions, so `TIERS.md` states them separately rather than folding
9
+ them into the role:
10
+
11
+ - **Complexity moves the effort.** A worker executing a finished plan needs less
12
+ reasoning than the reviewer judging its output. When the plan is airtight the
13
+ spec is carrying the thinking.
14
+ - **Stakes move the tier and the reader.** Security, privacy, data loss and
15
+ irreversible changes buy the challenge lane, a named check, a rollback path or
16
+ a human yes. A one-line change to an auth check is simple and high-stakes at
17
+ the same time, and it is the stakes that decide.
18
+
19
+ Stakes means what a mistake would cost: a security hole, leaked personal data,
20
+ lost data, or something you can't undo. Most tasks are low-stakes and route
21
+ normally.
22
+
23
+ The top of the ladder is bought with evidence: a reproduced failure, an
24
+ unresolved checkpoint, an irreversible change. A task that merely feels hard is
25
+ a deep-tier task, not an escalation.
26
+
27
+ ## A finding is a claim, not a fact
28
+
29
+ Review findings do not go straight to a repair. `finding-verifier` reads the
30
+ cited line, states what would trigger the problem, then hunts for the guard,
31
+ caller or test that makes it impossible, and returns **CONFIRMED**,
32
+ **NOT_REPRODUCED** or **INCONCLUSIVE** per finding. Only CONFIRMED earns a
33
+ change. Use a different model family from the one that produced the finding
34
+ where you have one: a family asked to check its own claim tends to agree with
35
+ itself.
36
+
37
+ ## Two more fast-tier checks
38
+
39
+ `done-verifier` probes the artifact a tracker item's done-signal names (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything itself. It carries no file-editing tools, but on claude-code it does carry `Bash` for those probes (`git log`, `grep`, `wc -l`, `test -f`); staying to read-only commands there is a rule in its prompt, not a restriction on the tool grant, and its own description says so. On agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `reader` is the one that is read-only by tool grant on both: no `Write`, `Edit`, or `Bash`. It reads and digests many files or notes and hands back exactly what the brief asked for, cited by `path:line`; it never classifies, tags or writes, which is what separates it from `bulk-worker`. Both ship in the claude-code and agy agent sets, at the fast tier.
40
+
41
+ ## Pin the route, or know that you did not
42
+
43
+ A lane with no `--model`, no `--effort` and no `defaults` entry in
44
+ `bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
45
+ CLI configured months ago at a low reasoning effort keeps auditing at that
46
+ effort while your routing docs describe a second-opinion pass.
47
+
48
+ ```bash
49
+ node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
50
+ node bin/cli-run.mjs --doctor # prints what each lane is pinned to, and what is not pinned
51
+ ```
52
+
53
+ Every run logs the model and effort **requested** and where the request came
54
+ from: `flag`, `lanes.json`, or `lane_default`, on every record including the
55
+ runs that never reached a lane. It does not log an actual. One lane of five
56
+ (grok) reports a model id in its own output and the other four report none, so
57
+ an actual field would be present for one lane and missing for four, and it
58
+ would be a provider-supplied string, which the durable log deliberately never
59
+ holds.
60
+
@@ -0,0 +1,73 @@
1
+ # Installing
2
+
3
+ Everything the installer asks, writes and accepts as a flag. The short version is in the [README](../README.md).
4
+
5
+ ## What a run does
6
+
7
+
8
+ The installer asks a few things, then writes a folder:
9
+
10
+ 1. **Which level?** 1 beginner · 2 intermediate · 3 advanced
11
+ 2. **Which AIs do you have access to?** (it marks the ones already on your PATH)
12
+ 3. **Which one is your primary agent?** (the one that runs the system)
13
+
14
+ It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions (and, on Claude Code, two hook scripts) go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: follow the generated README. Inspect the manifest and remove only the individual managed subagent files you no longer need, preserve edited or pre-existing files, and remove your manually pasted activation block. Never delete a shared subagent folder.
15
+
16
+ ## Plans and automatic effort
17
+
18
+ State known subscription plans with `--plans codex=pro-20x,agy=ultra-5x`. The generated guidance uses plan headroom to allocate volume only. It never changes capability or independent-review rules. `--effort-auto` is explicit consent to set `auto` only for selected high or max headroom CLI lanes. Auto chooses medium below 4,000 prompt characters and high otherwise, never higher. A codex audit is always high. Name `xhigh` explicitly for security-critical or irreversible work.
19
+ ## The two folders every run writes to
20
+
21
+ An install has two targets, and a scripted run should set both.
22
+
23
+ | Flag | Default | What lands there |
24
+ |---|---|---|
25
+ | `--dir` | `./ai-orchestrator` | the docs, protocols and (level 2+) `bin/cli-run.mjs`. Named after what it contains, not after this package, so a project can hold one without looking like a checkout of it. Pass `--dir ./model-orchestrator` if you prefer the package name. |
26
+ | `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets three hook scripts in `.claude/hooks/`, wired by a settings snippet you merge yourself. |
27
+
28
+ `--project` defaulting to the current directory is the one that surprises people: run the command from your home folder with Claude Code as the primary and five agent files land in your home folder. The installer prints the resolved project path in the plan and says when you left it at the default. Set it.
29
+
30
+ ## Non-interactive
31
+
32
+ ```bash
33
+ # both targets set: docs in ./ai-orchestrator, subagents into ./my-app/.claude/agents
34
+ npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code \
35
+ --dir ./ai-orchestrator --project ./my-app
36
+ # --yes selects the recommended companion tool (codecalc), which writes CODECALC.md and mcp/ snippets.
37
+ # Add --no-tools for none, or --tools codecalc,obsidian-tc,context7 to choose.
38
+
39
+ npx model-orchestrator --yes --level 3 --ais claude-code,codex,agy,grok,hermes,qwen,ollama --apis anthropic,openrouter --dry # print the plan, write nothing
40
+ npx model-orchestrator --yes --level 2 --ais claude-code,codex --project ~/my-app --dir ~/my-app/ai-orchestrator --no-tools # subagents into ~/my-app/.claude/agents
41
+ npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --update-docs # added a lane: regenerate the docs you never edited
42
+ ```
43
+
44
+ ## What gets written (level 3, everything)
45
+
46
+ ```
47
+ ai-orchestrator/
48
+ README.md start here, written for your level and your AIs
49
+ ORCHESTRATOR.md single-agent routing rules (level 1)
50
+ TASK_BUNDLE.md the brief every delegation carries
51
+ protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record · docs-then-prove
52
+ CODECALC.md OBSIDIAN-TC.md CONTEXT7.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
53
+ <project>/.claude/agents/ one per tier plus finding-verifier, done-verifier, reader, at the PROJECT root (if Claude Code is primary)
54
+ <project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart) + route-metrics.mjs (all five: see "Measuring routing" below), Claude Code only
55
+ CLAUDE.snippet.md the block to paste into your CLAUDE.md
56
+ settings.hooks.snippet.json the hooks block to merge into .claude/settings.json (Claude Code only)
57
+ ROUTING.md multi-lane decision tree (level 2+)
58
+ TIERS.md DELEGATION_MATRIX.md RESEARCH_TRIAGE.md CLI-RUN.md
59
+ bin/cli-run.mjs bin/lanes.json (node bin/cli-run.mjs --doctor is the smoke test)
60
+ vm/ gateway config, compose, box rules, privacy gates, jobs/ (level 3)
61
+ ```
62
+
63
+ ## Repo layout
64
+
65
+ | Folder | What |
66
+ |---|---|
67
+ | [`bin/`](bin/README.md) | `cli.js` (the installer) and `cli-run.mjs` (the lane runner) |
68
+ | [`src/`](src/README.md) | the catalog, the pure planner, detection, rendering |
69
+ | [`templates/`](templates/README.md) | everything the installer can write, by level, plus `tools/` for companions |
70
+ | [`docs/`](docs/README.md) | the three parts and the catalog |
71
+ | [`plugin/`](plugin/README.md) | the Claude Code plugin, generated from `templates/` by `npm run gen:plugin`; `.claude-plugin/marketplace.json` at the root lists it |
72
+ | [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu, macOS and Windows, Node 18/20/22 |
73
+
package/llms.txt CHANGED
@@ -10,7 +10,11 @@ Claude Code plugin: `/plugin marketplace add aunysillyme/model-orchestrator`, th
10
10
 
11
11
  ## Docs
12
12
 
13
- - [README](https://github.com/aunysillyme/model-orchestrator/blob/main/README.md): what it writes, flags, principles, what is enforced versus instructed
13
+ - [README](https://github.com/aunysillyme/model-orchestrator/blob/main/README.md): what it is, a real dry-run plan, the levels, the AIs, the principles
14
+ - [Installing](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/install.md): every flag, the two folders a run writes to, headless examples, the full file list
15
+ - [How it routes](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/how-it-routes.md): role, complexity and stakes; the three verifier agents; pinning model and effort per lane
16
+ - [Guarantees](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/guarantees.md): what is enforced by code, what is delegated to a vendor flag, what is only an instruction
17
+ - [Companion tools](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/companions.md): codecalc, obsidian-tc and Context7, what each closes and what it needs first
14
18
  - [Part 1: beginner](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-1-beginner.md): one agent, tiers, the task bundle every delegation carries
15
19
  - [Part 2: intermediate](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-2-intermediate.md): several agent CLIs, the lane runner, pinning model and effort per lane
16
20
  - [Part 3: advanced](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-3-advanced.md): the VM, the gateway, scheduled jobs, privacy gates
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.25",
3
+ "version": "0.1.27",
4
4
  "description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
5
5
  "type": "module",
6
6
  "bin": {
@@ -48,7 +48,7 @@ Pinned form, if you want the version this installer was released against: `npx -
48
48
 
49
49
  ## Register it with your agent (snippets in `mcp/`)
50
50
 
51
- Every snippet below ships **keyless**: the remote ones point at the hosted endpoint with no `Authorization` header at all (Qwen Code's snippet keeps the non-credential `Accept` header upstream itself ships), and the two Zed snippets (one per OS) run the local, version-pinned `npx` server with no key in its `env` block. That is deliberate, not an oversight: see "Higher rate limits" next for why a header is not shipped by default.
51
+ Every snippet below ships **keyless**: the remote ones point at the hosted endpoint with no `Authorization` header at all (Qwen Code's snippet keeps the non-credential `Accept` header upstream itself ships), and both Zed snippets run the local, version-pinned `npx` server with no key in its `env` block. That is deliberate, not an oversight: see "Higher rate limits" next for why a header is not shipped by default.
52
52
 
53
53
  | Agent | File to edit | Snippet |
54
54
  |---|---|---|
@@ -56,22 +56,27 @@ Every snippet below ships **keyless**: the remote ones point at the hosted endpo
56
56
  | Claude Desktop | no config file: `Settings > Connectors > Add Custom Connector`, name `Context7`, URL `https://mcp.context7.com/mcp` | none, it is a UI step |
57
57
  | Cursor | one-click install in the upstream README, or `~/.cursor/mcp.json` | `mcp/context7.mcpServers.json` |
58
58
  | VS Code | `.vscode/mcp.json` (key is `servers`, remote type is `http`) | `mcp/context7.vscode.mcp.json` |
59
- | Zed | `~/.config/zed/settings.json` (key is `context_servers`); upstream ships only a local (`npx`) config for Zed, so this snippet runs the server locally rather than remote. **Windows: use the `.windows` snippet**, see "Local `npx` on Windows" below | macOS/Linux: `mcp/context7.zed.settings.json` · Windows: `mcp/context7.zed.windows.settings.json` |
59
+ | Zed | `~/.config/zed/settings.json` (key is `context_servers`); upstream ships only a local (`npx`) config for Zed, so this snippet runs the server locally rather than remote. On Windows a current Zed starts it as written; see "Local `npx` on Windows" below for the two cases that need the `.windows` file | `mcp/context7.zed.settings.json` · fallback: `mcp/context7.zed.windows.settings.json` |
60
60
  | Codex CLI | `~/.codex/config.toml` | `mcp/context7.codex.config.toml` |
61
61
  | Antigravity `agy` | its MCP config file | `mcp/context7.agy.mcp_config.json` |
62
62
  | Qwen Code | `~/.qwen/settings.json` under `mcpServers`; note the field is `httpUrl`, not `url`, a different shape from every other client here | `mcp/context7.qwen.settings.json` |
63
63
 
64
64
  ### Local `npx` on Windows
65
65
 
66
- On Windows `npx` is a batch file (`npx.cmd`), and a client that spawns `"command": "npx"` directly fails to start it. Context7's own client guide says to wrap it in `cmd` on Windows ([all clients, Windows section](https://context7.com/docs/resources/all-clients)). This only affects the local form: every remote snippet above connects over HTTPS and spawns nothing.
66
+ On Windows `npx` is a batch file (`npx.cmd`) that a bare `CreateProcess` will not start, and Context7's own client guide says to wrap it in `cmd` on Windows ([all clients, Windows section](https://context7.com/docs/resources/all-clients)). That guide is about clients in general. **Zed is not one that needs it:** since its PR #42382, "Use shell to launch MCP and ACP servers" (2025-12-10), `crates/context_server/src/transport/stdio_transport.rs` builds every MCP launch through `ShellBuilder::new(&Shell::System, ..)`, so the shell resolves `npx` and a plain `"command": "npx"` starts. The `.windows` snippet is kept for two cases it does fix: a Zed older than that build, and the PowerShell one below. This only affects the local form either way: every remote snippet above connects over HTTPS and spawns nothing.
67
67
 
68
68
  | Where | Use |
69
69
  |---|---|
70
70
  | Zed on macOS or Linux | `mcp/context7.zed.settings.json` (`"command": "npx"`) |
71
- | Zed on Windows | `mcp/context7.zed.windows.settings.json` (`"command": "cmd"`, `"args": ["/c", "npx", ...]`) |
72
- | Any other client, local alternative, on Windows | the same change by hand: `"command": "cmd"` and put `"/c", "npx"` in front of the existing args, e.g. `"args": ["/c", "npx", "-y", "@upstash/context7-mcp@{{CONTEXT7_PIN}}"]` |
71
+ | Zed on Windows, current build | `mcp/context7.zed.settings.json` as well; the shell Zed opens resolves `npx` |
72
+ | Zed on Windows, older build or a blocked `npx.ps1` | `mcp/context7.zed.windows.settings.json` (`"command": "cmd"`, `"args": ["/d", "/c", "npx", ...]`) |
73
+ | Any other client, local alternative, on Windows | usually nothing: Cursor, VS Code, the Codex CLI and every client on the stock `@modelcontextprotocol/sdk` resolve `PATHEXT` themselves. If yours genuinely does not, the change is `"command": "cmd"` with `"/d", "/c", "npx"` in front of the existing args |
73
74
 
74
- Not yet run on a Windows machine by this project; the shape is the vendor's own.
75
+ The one Windows failure that is not about `.cmd`: Zed prefers PowerShell for the system shell (`get_windows_system_shell` in `crates/gpui_util/src/lib.rs` falls back to `cmd.exe` only when PowerShell is missing), and PowerShell resolves a bare `npx` to npm's `npx.ps1` shim when one is installed. Under the `Restricted` execution policy that is Windows' client default, a `.ps1` will not run. Check `Get-ExecutionPolicy`, and if that is the cause the `.windows` snippet is the fix.
76
+
77
+ The `/d` in those args is deliberate: without it `cmd` first runs whatever is in the Command Processor `AutoRun` registry value, which can print non-JSON into the protocol stream. `cross-spawn`, the library the MCP TypeScript SDK uses for exactly this job, passes `/d /s /c` for the same reason.
78
+
79
+ Not yet run on a Windows machine by this project.
75
80
 
76
81
  Merge the block; do not replace the file. `mcp/context7.mcpServers.json` (Cursor) and `mcp/context7.claude-code.mcp.json` look alike but are not interchangeable: Claude Code requires the `"type": "http"` field and Cursor's own docs show plain `{"url": ...}` with no `type` at all.
77
82
 
@@ -3,7 +3,7 @@
3
3
  "Context7": {
4
4
  "source": "custom",
5
5
  "command": "cmd",
6
- "args": ["/c", "npx", "-y", "@upstash/context7-mcp@{{CONTEXT7_PIN}}"]
6
+ "args": ["/d", "/c", "npx", "-y", "@upstash/context7-mcp@{{CONTEXT7_PIN}}"]
7
7
  }
8
8
  }
9
9
  }
@@ -54,6 +54,24 @@ Every snippet points `OBSIDIAN_TC_CONFIG` at your config file. Replace `/ABSOLUT
54
54
 
55
55
  Merge the block; do not replace the file.
56
56
 
57
+ ### Local `npx` on Windows
58
+
59
+ Every snippet here spawns `npx`, and on Windows `npx` is a batch file (`npx.cmd`) that a bare `CreateProcess` will not start. **All five clients handle that themselves, so there is no Windows snippet here and the files above are the ones to use on every OS.** Each was read rather than assumed, because the usual advice ("on Windows, wrap it in `cmd /c`") is about clients in general and is wrong for all five of these:
60
+
61
+ | Client | How it starts `"command": "npx"` on Windows | Checked against |
62
+ |---|---|---|
63
+ | Zed | hands the whole command to the system shell | `crates/context_server/src/transport/stdio_transport.rs` builds through `ShellBuilder::new(&Shell::System, ..)`, from PR #42382 "Use shell to launch MCP and ACP servers" (2025-12-10) |
64
+ | VS Code | resolves the executable, then re-spawns it through a shell | `src/vs/workbench/api/node/extHostMcpNode.ts`, `formatSubprocessArguments` resolves the extension and sets `shell: true` when it ends in `.bat` or `.cmd` |
65
+ | Codex CLI | resolves `PATHEXT` itself before spawning | `codex-rs/rmcp-client/src/program_resolver.rs`, whose Windows arm calls `which::which_in` so that "tools like `npx`, `pnpm`, and `yarn` work correctly on Windows" |
66
+ | Cursor | the MCP SDK it bundles spawns through `cross-spawn` | Cursor 3.18.9 ships `@modelcontextprotocol/sdk` 1.25.1, whose `dist/esm/client/stdio.js` opens with `import spawn from 'cross-spawn'`; `cross-spawn/lib/parse.js` resolves the command and, when it is not an `.exe`, re-invokes it as `%COMSPEC% /d /s /c "<escaped>"` |
67
+ | Antigravity `agy` | same SDK path, through the Gemini CLI config namespace it reads | Gemini CLI pins `@modelcontextprotocol/sdk` 1.23.0 in `packages/core/package.json` and hands `mcpServerConfig.command` to that transport; 1.23.0's `client/stdio.js` imports `cross-spawn` too |
68
+
69
+ The SDK point covers more than these two clients: every published `@modelcontextprotocol/sdk` from 1.23.0 through 1.30.0 depends on `cross-spawn ^7.0.5` and uses it in the stdio client, so any client that connects through the stock TypeScript SDK inherits the same `PATHEXT` resolution. `shell: false` in that transport is not the whole story, and reading only that line is how a client gets mistaken for one that cannot start `npx`.
70
+
71
+ One Windows case can still fail, and it is not about `.cmd`. Zed prefers PowerShell for the system shell (`get_windows_system_shell` in `crates/gpui_util/src/lib.rs` falls back to `cmd.exe` only when PowerShell is missing), and PowerShell resolves a bare `npx` to npm's `npx.ps1` shim when one is installed. Under the `Restricted` execution policy that is Windows' client default, running a `.ps1` is blocked. If Zed reports that the server would not start, check `Get-ExecutionPolicy` first, and if that is the cause, change the Zed entry by hand to `"command": "cmd"` with `"args": ["/d", "/c", "npx", "-y", "obsidian-tc"]`, keeping the rest of the block. `/d` is there on purpose: it skips any Command Processor `AutoRun` command, which would otherwise run first and can print non-JSON into the protocol stream.
72
+
73
+ None of this was run on a Windows machine by this project. The five verdicts are from each client's own shipped code; the PowerShell case is from Zed's shell choice plus documented `Restricted` behaviour, and is the one worth reporting back if you hit it.
74
+
57
75
  ## Security posture, read before a second agent touches it
58
76
 
59
77
  Zero-config mode boots with **auth off and no folder ACL**: anything that can reach the server has the same authority as raw filesystem access to the vault. That is acceptable only because the surface is local-only (the config refuses by default if you enable HTTP on a non-loopback host with auth off, and a DNS-rebinding guard protects loopback). Before exposing it to partially-trusted, remote or multi-agent callers, turn on `auth.mode: "jwt"` and set `acl.readPaths` / `writePaths` / `deletePaths` in the config file. Upstream `SECURITY.md` has the security notes and a private disclosure path.