model-orchestrator 0.1.25 → 0.1.27
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +19 -1
- package/README.md +98 -205
- package/docs/README.md +8 -1
- package/docs/companions.md +17 -0
- package/docs/guarantees.md +18 -0
- package/docs/how-it-routes.md +60 -0
- package/docs/install.md +73 -0
- package/llms.txt +5 -1
- package/package.json +1 -1
- package/templates/tools/context7/CONTEXT7.md +11 -6
- package/templates/tools/context7/mcp/context7.zed.windows.settings.json +1 -1
- package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +18 -0
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,22 @@ All notable changes to this project are documented here. The format follows [Kee
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [0.1.27] - 2026-09-21
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
- **README restructured for a first-time reader: 32,083 bytes to 17,586, with the first screen now a claim, the install command and a real `--dry` plan.** The old opening was a 200-word paragraph followed by "What it is not", the flag-conflict rules and the uninstall procedure, all before the reader had seen the tool do anything. The detail moved rather than went away: `docs/install.md` (every flag, the two target folders, headless examples, the full file list), `docs/how-it-routes.md` (role, complexity and stakes; the three verifier agents; pinning model and effort), `docs/guarantees.md` (enforced by code, delegated to a vendor flag, or only an instruction), `docs/companions.md` (the full companion-tool table). `docs/README.md` and `llms.txt` index all four. Platform support and the vendor compatibility table stay in the README inside collapsed `<details>` blocks, because `scripts/gen-catalog.js` writes the vendor table between markers there and `test/prose.test.js` requires the skipped-test explanations to live in the README; collapsing them keeps both mechanisms pointed at the same file. The opening line still satisfies `test/copy.test.js`: it names the package, carries the shared purpose clause, and keeps the 0.1.11 correction that it does not select models itself.
|
|
12
|
+
|
|
13
|
+
## [0.1.26] - 2026-09-19
|
|
14
|
+
|
|
15
|
+
### Added
|
|
16
|
+
|
|
17
|
+
- **obsidian-tc on Windows: the answer is that no client needs a `cmd /c` snippet, and `OBSIDIAN-TC.md` now says so with the code it was read from.** All five obsidian-tc snippets spawn `"command": "npx"`, and on Windows `npx` is `npx.cmd`, which a bare `CreateProcess` will not start. Each client was read instead of assumed, and all five resolve it themselves: Zed hands the command to the system shell (`crates/context_server/src/transport/stdio_transport.rs`, `ShellBuilder::new(&Shell::System, ..)`, from its PR #42382, 2025-12-10), VS Code's `formatSubprocessArguments` in `src/vs/workbench/api/node/extHostMcpNode.ts` resolves the executable and re-spawns with `shell: true` for a `.bat` or `.cmd`, the Codex CLI calls `which::which_in` in `codex-rs/rmcp-client/src/program_resolver.rs`, and Cursor and Antigravity's `agy` both reach `npx` through `@modelcontextprotocol/sdk`, whose `client/stdio.js` imports `cross-spawn` and re-invokes a non-`.exe` as `%COMSPEC% /d /s /c`. That last point is wider than those two clients: every published SDK from 1.23.0 through 1.30.0 depends on `cross-spawn ^7.0.5`, so `shell: false` in that transport is not the whole story. The doc also names the one Windows case that can still fail and is not about `.cmd`: Zed prefers PowerShell, which resolves a bare `npx` to npm's `npx.ps1` shim, and a `.ps1` will not run under the `Restricted` execution policy that is Windows' client default. A test pins the snippet set, fails if any `.windows` file appears, and fails if the doc drops a citation. Not run on a Windows machine.
|
|
18
|
+
|
|
19
|
+
### Fixed
|
|
20
|
+
|
|
21
|
+
- **Context7's Windows guidance said something untrue about Zed, and its `cmd` wrapper was missing `/d`** ([#34](https://github.com/aunysillyme/model-orchestrator/issues/34) follow-up). 0.1.25 shipped `mcp/context7.zed.windows.settings.json` on the premise that Zed spawns `"command": "npx"` directly and so cannot start it. Reading Zed's own source for the obsidian-tc work above showed it has launched MCP stdio servers through the system shell since PR #42382 (2025-12-10), so a current Zed starts the plain snippet. The file stays, because it is still the fix for an older Zed and for the PowerShell execution-policy case, but `CONTEXT7.md` no longer sends every Windows user to it and now carries both reasons. The snippet's args change from `["/c", "npx", ...]` to `["/d", "/c", "npx", ...]`: without `/d`, `cmd` first runs whatever sits in the Command Processor `AutoRun` registry value, which can print non-JSON into the protocol stream. `cross-spawn` passes `/d /s /c` for the same reason. The `#34` test follows the new args and now also fails if the doc stops naming why a current Zed does not need the file.
|
|
22
|
+
|
|
7
23
|
## [0.1.25] - 2026-09-19
|
|
8
24
|
|
|
9
25
|
### Fixed
|
|
@@ -365,7 +381,9 @@ First release.
|
|
|
365
381
|
- Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
|
|
366
382
|
- Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
|
|
367
383
|
|
|
368
|
-
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.
|
|
384
|
+
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.27...HEAD
|
|
385
|
+
[0.1.27]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.26...v0.1.27
|
|
386
|
+
[0.1.26]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.25...v0.1.26
|
|
369
387
|
[0.1.25]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.24...v0.1.25
|
|
370
388
|
[0.1.24]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.23...v0.1.24
|
|
371
389
|
[0.1.23]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.22...v0.1.23
|
package/README.md
CHANGED
|
@@ -2,48 +2,58 @@
|
|
|
2
2
|
|
|
3
3
|
[](https://www.npmjs.com/package/model-orchestrator) [](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml) [](LICENSE) [](package.json)
|
|
4
4
|
|
|
5
|
-
**
|
|
6
|
-
|
|
7
|
-
## At a glance
|
|
8
|
-
|
|
9
|
-
- **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
|
|
10
|
-
- **What it is not:** a proxy, a gateway or an API router. It does not automatically compare prices or select models; your agent follows the rules and chooses.
|
|
11
|
-
- **Install:** `npx model-orchestrator` (interactive), or headless from a script or an agent: `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator`.
|
|
12
|
-
- **Claude Code plugin:** `/plugin marketplace add aunysillyme/model-orchestrator`, then `/plugin install model-orchestrator@model-orchestrator`. The routing rules still come from the installer; see [Claude Code plugin](#claude-code-plugin).
|
|
13
|
-
- **Use it when:** you run more than one model or agent and want each task sent to the smallest one that can do it well.
|
|
14
|
-
- **What it saves:** frontier-model tokens. Bulk work, reading and checks go to fast tiers; the expensive tier is kept for planning and judgment.
|
|
15
|
-
- **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
|
|
16
|
-
|
|
17
|
-
Built from a working system, not a diagram: the routing rules, the protocols and the lane runner here run in production, generalized so they transfer to any stack.
|
|
5
|
+
**A model orchestrator that sends every task to the smallest model that can do it,** so small work goes to cheap tiers and fewer tokens go to frontier models. One command reads which AIs you actually have, then writes the routing rules, the subagents and the lane runner for exactly that set: Claude Code, Codex, Gemini, Grok, Qwen, Ollama.
|
|
18
6
|
|
|
19
7
|
```bash
|
|
20
8
|
npx model-orchestrator
|
|
21
9
|
```
|
|
22
10
|
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
11
|
+
Three questions, then 38 files. Here is a real `--dry` run, which prints the plan and writes nothing:
|
|
12
|
+
|
|
13
|
+
```text
|
|
14
|
+
Plan
|
|
15
|
+
level 2 Intermediate
|
|
16
|
+
access claude-code, codex, grok
|
|
17
|
+
primary claude-code
|
|
18
|
+
tools codecalc
|
|
19
|
+
folder ./ai-orchestrator
|
|
20
|
+
project . (11 subagent files go here)
|
|
21
|
+
files 38
|
|
22
|
+
- ROUTING.md multi-lane decision tree
|
|
23
|
+
- TIERS.md DELEGATION_MATRIX.md which lane, at what effort
|
|
24
|
+
- TASK_BUNDLE.md the brief every delegation carries
|
|
25
|
+
- protocols/ build, propagate, gap-analysis, deep-research, and three more
|
|
26
|
+
- [project] .claude/agents/ builder, deep-planner, code-reviewer, bulk-worker,
|
|
27
|
+
live-researcher, reader, finding-verifier, done-verifier
|
|
28
|
+
- [project] .claude/hooks/ route-gate, subagent-context, route-metrics
|
|
29
|
+
- bin/cli-run.mjs bin/lanes.json the lane runner
|
|
30
|
+
|
|
31
|
+
--dry: nothing written.
|
|
32
|
+
```
|
|
26
33
|
|
|
27
|
-
|
|
28
|
-
2. **Which AIs do you have access to?** (it marks the ones already on your PATH)
|
|
29
|
-
3. **Which one is your primary agent?** (the one that runs the system)
|
|
34
|
+
Reproduce it:
|
|
30
35
|
|
|
31
|
-
|
|
36
|
+
```bash
|
|
37
|
+
npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --dry
|
|
38
|
+
```
|
|
32
39
|
|
|
33
|
-
|
|
40
|
+
- **What it is:** routing rules, subagent definitions and a CLI lane runner (`cli-run`) for the AI tools you already pay for.
|
|
41
|
+
- **What it is not:** a proxy, a gateway or an API router. It does not automatically compare prices or select models; your agent follows the rules and chooses.
|
|
42
|
+
- **Use it when:** you run more than one model or agent and want the expensive tier kept for planning and judgment.
|
|
43
|
+
- **For agents:** [`llms.txt`](llms.txt) summarizes the package and links every doc; [`AGENTS.md`](AGENTS.md) has the headless commands.
|
|
34
44
|
|
|
35
|
-
|
|
45
|
+
Built from a working system, not a diagram: the routing rules, the protocols and the lane runner here run in production, generalized so they transfer to any stack.
|
|
36
46
|
|
|
37
47
|
## The three levels
|
|
38
48
|
|
|
49
|
+
|
|
39
50
|
| Level | You have | You get |
|
|
40
51
|
|---|---|---|
|
|
41
52
|
| **1 · Beginner** | one LLM or one agent | tiers, task classification, the two build checkpoints, the protocols (build, propagate, gap analysis, deep research, numbers and logic, memory and record, docs and proof), a task-bundle template, and your agent set up to follow them |
|
|
42
53
|
| **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts; `--model` / `--effort` to pin the route and log it), a delegation matrix generated from your selection, research triage across the lanes you have |
|
|
43
54
|
| **3 · Advanced** | a virtual machine | everything above, plus a gateway config rendered from the API keys you hold (asked separately from your CLIs), pinned images, box rules, privacy gates, and a weekly gap-analysis job with "what watches it" written down |
|
|
44
55
|
|
|
45
|
-
|
|
46
|
-
|
|
56
|
+
Levels explained: [Part 1](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md).
|
|
47
57
|
## The AIs it knows about
|
|
48
58
|
|
|
49
59
|
| Id | What | Level |
|
|
@@ -59,134 +69,39 @@ Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/pa
|
|
|
59
69
|
|
|
60
70
|
`npx model-orchestrator --list` prints the catalog with install and sign-in notes. Details: [docs/catalog.md](docs/catalog.md).
|
|
61
71
|
|
|
62
|
-
##
|
|
63
|
-
|
|
64
|
-
An orchestrator routes work. It does not make a model stop guessing numbers, it does not give it a memory, and it does not make it check a library's current docs before writing against it. Three tools close those gaps: codecalc and obsidian-tc are from the same maintainer, Context7 is from Upstash. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc,context7` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
|
|
65
|
-
|
|
66
|
-
| Tool | Closes | Default | You need first |
|
|
67
|
-
|---|---|---|---|
|
|
68
|
-
| [codecalc](https://github.com/The-40-Thieves/codecalc) | guessed numbers, comparisons, complexity and equivalence claims: exact arithmetic, code execution in 31 languages, SMT logic checks, `verify_translation` / `verify_optimization`; offline, no key | yes | Python 3.10+ and `uv`. `uvx 'codecalc[full]' setup --write` registers it with Claude Code, Claude Desktop, Cursor, VS Code, Zed; snippets for Codex, Antigravity, Qwen Code are written for you |
|
|
69
|
-
| [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes with a confirmation gate, folder ACLs, a poison scan on inferred writes; 163 tools, local by default; AGPL-3.0 | no | an Obsidian vault folder; Node 24+ or Bun 1.1+ (stricter than this installer); Ollama with `nomic-embed-text` or a cloud embeddings key; the Obsidian app and its Local REST API plugin only for live bridge tools. Skip it if you do not keep notes in Obsidian |
|
|
70
|
-
| [Context7](https://github.com/upstash/context7) | stale library recall: current, version-specific docs and code examples pulled into the prompt for any library, SDK, API or CLI; hosted, or `npx` locally; MIT | no | nothing to install for the hosted endpoint; Node.js 18+ for the local alternative; a free API key is optional, for a higher rate limit. Always makes a network call, unlike the other two: skip it offline |
|
|
71
|
-
|
|
72
|
-
Context7 pairs with codecalc rather than duplicating it: Context7 tells the agent what a library is documented to do on this version, codecalc runs the code and proves what it actually does. Docs never stand as proof on their own, and where the two disagree the run wins.
|
|
73
|
-
|
|
74
|
-
Whether or not you select them, every level carries the three rules they serve: `protocols/numbers-and-logic.md` (when calling a calculator is mandatory, how to report a computed figure, why a thought log is not evidence), `protocols/memory-and-record.md` (search before writing, the folder index is part of the change, one writer, inferred content marked as inferred), and `protocols/docs-then-prove.md` (current docs before writing a call, then a run proves it, the run wins on disagreement).
|
|
75
|
-
|
|
76
|
-
## The two folders every run writes to
|
|
77
|
-
|
|
78
|
-
An install has two targets, and a scripted run should set both.
|
|
79
|
-
|
|
80
|
-
| Flag | Default | What lands there |
|
|
81
|
-
|---|---|---|
|
|
82
|
-
| `--dir` | `./ai-orchestrator` | the docs, protocols and (level 2+) `bin/cli-run.mjs`. Named after what it contains, not after this package, so a project can hold one without looking like a checkout of it. Pass `--dir ./model-orchestrator` if you prefer the package name. |
|
|
83
|
-
| `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets three hook scripts in `.claude/hooks/`, wired by a settings snippet you merge yourself. |
|
|
84
|
-
|
|
85
|
-
`--project` defaulting to the current directory is the one that surprises people: run the command from your home folder with Claude Code as the primary and five agent files land in your home folder. The installer prints the resolved project path in the plan and says when you left it at the default. Set it.
|
|
72
|
+
## Measuring routing
|
|
86
73
|
|
|
87
|
-
|
|
74
|
+
A routing rule nobody measures is a rule nobody knows is followed. On a claude-code install, `route-metrics.mjs` turns every turn, dispatch and subagent start/stop into one JSON line under `~/.ai-orchestrator/route-metrics.jsonl`, including the lane your agent named in its own `<!-- route: <lane> | <why> -->` marker.
|
|
88
75
|
|
|
89
76
|
```bash
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
--dir ./ai-orchestrator --project ./my-app
|
|
93
|
-
# --yes selects the recommended companion tool (codecalc), which writes CODECALC.md and mcp/ snippets.
|
|
94
|
-
# Add --no-tools for none, or --tools codecalc,obsidian-tc,context7 to choose.
|
|
95
|
-
|
|
96
|
-
npx model-orchestrator --yes --level 3 --ais claude-code,codex,agy,grok,hermes,qwen,ollama --apis anthropic,openrouter --dry # print the plan, write nothing
|
|
97
|
-
npx model-orchestrator --yes --level 2 --ais claude-code,codex --project ~/my-app --dir ~/my-app/ai-orchestrator --no-tools # subagents into ~/my-app/.claude/agents
|
|
98
|
-
npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --update-docs # added a lane: regenerate the docs you never edited
|
|
77
|
+
node .claude/hooks/route-metrics.mjs --summary # since the log began
|
|
78
|
+
node .claude/hooks/route-metrics.mjs --summary --since 2026-09-01 # since a date
|
|
99
79
|
```
|
|
100
80
|
|
|
81
|
+
The report prints turns, **route-marker coverage** (the share of turns that carried a real lane, which answers "is the agent actually tagging its routing decisions?"), lanes by count, dispatches by `subagent_type`, dispatches with no matching start, and mean/max duration per agent type. It never logs prompt text, tool descriptions, or the "why" half of the marker. Fail-open by design: a miss is a missing log line, never a blocked turn.
|
|
82
|
+
|
|
101
83
|
## Claude Code plugin
|
|
102
84
|
|
|
103
|
-
The
|
|
85
|
+
The hooks and subagents also ship as a plugin, so they install and update through Claude Code itself:
|
|
104
86
|
|
|
105
87
|
```
|
|
106
88
|
/plugin marketplace add aunysillyme/model-orchestrator
|
|
107
89
|
/plugin install model-orchestrator@model-orchestrator
|
|
108
90
|
```
|
|
109
91
|
|
|
110
|
-
|
|
111
|
-
- **What it still needs from the installer:** the routing rules. A plugin runs no install step, so it reads the installer's default locations, `ai-orchestrator/ROUTING.md` then `ai-orchestrator/ORCHESTRATOR.md`, and names `npx model-orchestrator` when neither exists. A project installed with a different `--dir` should wire the installer's rendered hooks instead.
|
|
112
|
-
- **What it leaves out:** `route-metrics.mjs`, the routing log. It writes to disk and the plugin ships only hooks that read, so `npx model-orchestrator` is how you get it.
|
|
113
|
-
- **How it is kept honest:** `plugin/` is generated from `templates/` by `npm run gen:plugin`, and `test/plugin.test.js` fails when the committed bundle drifts, when a hook gains a network call, a write or a subprocess, or when an agent loses its tool list. The bundle passes `claude plugin validate --strict`, the check Anthropic's community marketplace review runs on every submission.
|
|
92
|
+
It ships the three hooks and the eight subagents, each with an explicit tool list, and loads them namespaced as `model-orchestrator:builder`. It does not ship the routing rules, because a plugin runs no install step: `npx model-orchestrator` still writes those. `plugin/` is generated from `templates/` and `test/plugin.test.js` fails when the committed bundle drifts, when a hook gains a network call, a write or a subprocess, or when an agent loses its tool list. Details: [plugin/README.md](plugin/README.md).
|
|
114
93
|
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
## What gets written (level 3, everything)
|
|
118
|
-
|
|
119
|
-
```
|
|
120
|
-
ai-orchestrator/
|
|
121
|
-
README.md start here, written for your level and your AIs
|
|
122
|
-
ORCHESTRATOR.md single-agent routing rules (level 1)
|
|
123
|
-
TASK_BUNDLE.md the brief every delegation carries
|
|
124
|
-
protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record · docs-then-prove
|
|
125
|
-
CODECALC.md OBSIDIAN-TC.md CONTEXT7.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
|
|
126
|
-
<project>/.claude/agents/ one per tier plus finding-verifier, done-verifier, reader, at the PROJECT root (if Claude Code is primary)
|
|
127
|
-
<project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart) + route-metrics.mjs (all five: see "Measuring routing" below), Claude Code only
|
|
128
|
-
CLAUDE.snippet.md the block to paste into your CLAUDE.md
|
|
129
|
-
settings.hooks.snippet.json the hooks block to merge into .claude/settings.json (Claude Code only)
|
|
130
|
-
ROUTING.md multi-lane decision tree (level 2+)
|
|
131
|
-
TIERS.md DELEGATION_MATRIX.md RESEARCH_TRIAGE.md CLI-RUN.md
|
|
132
|
-
bin/cli-run.mjs bin/lanes.json (node bin/cli-run.mjs --doctor is the smoke test)
|
|
133
|
-
vm/ gateway config, compose, box rules, privacy gates, jobs/ (level 3)
|
|
134
|
-
```
|
|
135
|
-
|
|
136
|
-
## Repo layout
|
|
137
|
-
|
|
138
|
-
| Folder | What |
|
|
139
|
-
|---|---|
|
|
140
|
-
| [`bin/`](bin/README.md) | `cli.js` (the installer) and `cli-run.mjs` (the lane runner) |
|
|
141
|
-
| [`src/`](src/README.md) | the catalog, the pure planner, detection, rendering |
|
|
142
|
-
| [`templates/`](templates/README.md) | everything the installer can write, by level, plus `tools/` for companions |
|
|
143
|
-
| [`docs/`](docs/README.md) | the three parts and the catalog |
|
|
144
|
-
| [`plugin/`](plugin/README.md) | the Claude Code plugin, generated from `templates/` by `npm run gen:plugin`; `.claude-plugin/marketplace.json` at the root lists it |
|
|
145
|
-
| [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu, macOS and Windows, Node 18/20/22 |
|
|
146
|
-
|
|
147
|
-
## What is enforced, what is delegated, what is an instruction
|
|
148
|
-
|
|
149
|
-
Most of what this package ships is text an agent is asked to follow. Be clear about which is which before relying on it unattended.
|
|
150
|
-
|
|
151
|
-
| Property | How it holds |
|
|
152
|
-
|---|---|
|
|
153
|
-
| Installer writes only inside `--dir` and `--project`, never a secret, never over a document without `--force` (or `--update-docs`, which touches only documents provably untouched since a previous run); machine-owned config always, runtime files only when provably untouched or with `--upgrade-runtime` | **enforced by code** (preflight, exclusive create, rollback, manifest hashes; tested) |
|
|
154
|
-
| `cli-run` exit codes, process-group kill on timeout and on SIGINT/SIGTERM, UTF-8-safe streaming, fixed-code durable log, `--expect-*` contracts with a pre-run snapshot | **enforced by code** (tested with stub lanes) |
|
|
155
|
-
| Codex audit lane runs read-only | **delegated to the vendor flag** (`--audit` → `--sandbox read-only`); commands and network still follow your codex config |
|
|
156
|
-
| Other lanes' permissions, sign-in state, model versions | **delegated to each vendor's own config**; `--doctor` checks presence, not versions |
|
|
157
|
-
| Gateway binds to loopback, keys by name only | **enforced in the generated files**; whether the gateway authenticates is your environment |
|
|
158
|
-
| Lane selection, tiers, privacy classes, one-writer, escalation, the protocols | **agent instructions**. Nothing here stops an agent that ignores its rules; the task bundle and the protocols make ignoring them visible, not impossible |
|
|
159
|
-
| Weekly audit bounded, previous report preserved | **enforced in the generated script and unit** (watchdog, temp-and-rename, `TimeoutStartSec`) |
|
|
160
|
-
|
|
161
|
-
If you need a property in the third row to be enforced, that is a router, a policy engine or a sandbox, and this package does not claim to be one.
|
|
162
|
-
|
|
163
|
-
### Vendor version compatibility
|
|
164
|
-
|
|
165
|
-
**This package detects that a binary exists. It does not check its version, and a present binary is not a working lane.** `--doctor` reports presence, and with `--run` sends one lane a one-word canary; neither validates that the vendor's flags, output shape or auth still match what the generated files assume.
|
|
166
|
-
|
|
167
|
-
The lane wiring and the output judges were written against these versions, which are the ones this release was exercised on:
|
|
168
|
-
|
|
169
|
-
<!-- vendor-table:start -->
|
|
170
|
-
|
|
171
|
-
| Lane | Vendor | Version this release was built against | Where that number is proved |
|
|
172
|
-
|---|---|---|---|
|
|
173
|
-
| `claude` | Anthropic | 2.1.226 | the npm pin the installer writes, `@anthropic-ai/claude-code@2.1.226` |
|
|
174
|
-
| `codex` | OpenAI | 0.153.4 | `test/fixtures/codex-0.153.4.jsonl`, a recorded run |
|
|
175
|
-
| `agy` | Google | 1.1.27 | `test/fixtures/agy-1.1.27.jsonl`, a recorded run |
|
|
176
|
-
| `grok` | xAI | 1.0.5 | `test/fixtures/grok-1.0.5.json`, a recorded run |
|
|
177
|
-
| `hermes` | Nous Research | 0.20.0 | `test/fixtures/hermes-0.20.0.txt`, a recorded run |
|
|
178
|
-
| `qwen` | Alibaba | 0.22.3 | `test/fixtures/qwen-0.22.3-nokey.json`, a recorded run |
|
|
179
|
-
| `ollama` | Ollama | 0.33.3 | the pinned image the level 3 box runs, `ollama/ollama:0.33.3` |
|
|
180
|
-
|
|
181
|
-
Generated from `src/catalog.js` by `npm run gen:catalog`; `npm test` fails if this table and the catalog disagree. Fixtures were captured 2026-09-06.
|
|
182
|
-
|
|
183
|
-
<!-- vendor-table:end -->
|
|
94
|
+
## Companion tools (all optional)
|
|
184
95
|
|
|
185
|
-
|
|
96
|
+
An orchestrator routes work. It does not make a model stop guessing numbers, give it a memory, or make it check a library's current docs before writing against it. Three tools close those gaps:
|
|
186
97
|
|
|
187
|
-
|
|
98
|
+
| Tool | Closes | Default |
|
|
99
|
+
|---|---|---|
|
|
100
|
+
| [codecalc](https://github.com/The-40-Thieves/codecalc) | guessed numbers: exact arithmetic, code execution in 31 languages, logic checks; offline, no key | yes |
|
|
101
|
+
| [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes; local by default | no |
|
|
102
|
+
| [Context7](https://github.com/upstash/context7) | stale library recall: current, version-specific docs pulled into the prompt | no |
|
|
188
103
|
|
|
189
|
-
|
|
104
|
+
Selecting one writes a doc and config snippets; it installs nothing. What each needs first, and why Context7 pairs with codecalc rather than duplicating it: [docs/companions.md](docs/companions.md). Whether or not you select them, every level carries the three rules they serve: `protocols/numbers-and-logic.md`, `protocols/memory-and-record.md` and `protocols/docs-then-prove.md`.
|
|
190
105
|
|
|
191
106
|
## Principles the whole thing rests on
|
|
192
107
|
|
|
@@ -198,101 +113,83 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
|
|
|
198
113
|
6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, only the brief carries what this task needs. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
|
|
199
114
|
7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
|
|
200
115
|
|
|
201
|
-
##
|
|
116
|
+
## Common questions
|
|
202
117
|
|
|
203
|
-
|
|
204
|
-
different directions, so `TIERS.md` states them separately rather than folding
|
|
205
|
-
them into the role:
|
|
118
|
+
### How do I cut token usage across Claude Code, Codex and Gemini?
|
|
206
119
|
|
|
207
|
-
|
|
208
|
-
reasoning than the reviewer judging its output. When the plan is airtight the
|
|
209
|
-
spec is carrying the thinking.
|
|
210
|
-
- **Stakes move the tier and the reader.** Security, privacy, data loss and
|
|
211
|
-
irreversible changes buy the challenge lane, a named check, a rollback path or
|
|
212
|
-
a human yes. A one-line change to an auth check is simple and high-stakes at
|
|
213
|
-
the same time, and it is the stakes that decide.
|
|
120
|
+
Install for the tools you have, then let the generated `ROUTING.md` decide the tier per task: bulk, reading and verification go to the fast tier or a cheaper CLI lane, and the deep tier only plans and judges. On Claude Code, execution goes to the `builder` subagent by default and the main session plans and verifies. Every lane call through `cli-run` logs the model and effort it ran with, so you can check where the tokens went.
|
|
214
121
|
|
|
215
|
-
|
|
216
|
-
lost data, or something you can't undo. Most tasks are low-stakes and route
|
|
217
|
-
normally.
|
|
122
|
+
### How do I route tasks to cheaper models?
|
|
218
123
|
|
|
219
|
-
The
|
|
220
|
-
unresolved checkpoint, an irreversible change. A task that merely feels hard is
|
|
221
|
-
a deep-tier task, not an escalation.
|
|
124
|
+
The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly never gets a frontier token.
|
|
222
125
|
|
|
223
|
-
|
|
126
|
+
### Is this an LLM router or an AI gateway?
|
|
224
127
|
|
|
225
|
-
|
|
226
|
-
cited line, states what would trigger the problem, then hunts for the guard,
|
|
227
|
-
caller or test that makes it impossible, and returns **CONFIRMED**,
|
|
228
|
-
**NOT_REPRODUCED** or **INCONCLUSIVE** per finding. Only CONFIRMED earns a
|
|
229
|
-
change. Use a different model family from the one that produced the finding
|
|
230
|
-
where you have one: a family asked to check its own claim tends to agree with
|
|
231
|
-
itself.
|
|
128
|
+
No. It routes at the task level, through instructions your agent follows and a runner for agent CLIs. If you want a service or proxy that picks or forwards the model on every API request, look at request-level routers and gateways such as RouteLLM, LiteLLM, OpenRouter or claude-code-router. They solve a different problem and can sit underneath this.
|
|
232
129
|
|
|
233
|
-
|
|
130
|
+
### Can an agent install and run it without a person?
|
|
234
131
|
|
|
235
|
-
|
|
132
|
+
Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run` previews the plan, and `--list` prints every supported AI. Nothing is appended to a file you already have; activation snippets are written next to your files for you to merge.
|
|
236
133
|
|
|
237
|
-
## Measuring routing
|
|
238
134
|
|
|
239
|
-
|
|
135
|
+
## Read next
|
|
240
136
|
|
|
241
|
-
|
|
242
|
-
|
|
243
|
-
|
|
244
|
-
|
|
137
|
+
| Doc | What is in it |
|
|
138
|
+
|---|---|
|
|
139
|
+
| [docs/install.md](docs/install.md) | every flag, the two folders a run writes to, headless examples, the full file list |
|
|
140
|
+
| [docs/how-it-routes.md](docs/how-it-routes.md) | role, complexity and stakes; the three verifier agents; pinning a lane's model and effort |
|
|
141
|
+
| [docs/guarantees.md](docs/guarantees.md) | what is enforced by code, what is delegated to a vendor flag, and what is only an instruction |
|
|
142
|
+
| [docs/part-1-beginner.md](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md) | the thinking behind each level |
|
|
143
|
+
| [docs/catalog.md](docs/catalog.md) | every supported AI with install and sign-in notes |
|
|
245
144
|
|
|
246
|
-
|
|
145
|
+
<details>
|
|
146
|
+
<summary><strong>Platform support, and every test this suite skips</strong></summary>
|
|
247
147
|
|
|
248
|
-
## Pin the route, or know that you did not
|
|
249
148
|
|
|
250
|
-
A lane with no
|
|
251
|
-
`bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
|
|
252
|
-
CLI configured months ago at a low reasoning effort keeps auditing at that
|
|
253
|
-
effort while your routing docs describe a second-opinion pass.
|
|
149
|
+
Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22), including lane execution end to end through `cli-run` against a fake CLI installed the same way npm installs a real one (a `.cmd` shim). `cli-run` never runs a lane through `cmd.exe` when it can avoid it: it resolves the shim to the Node script underneath and spawns Node directly, so a prompt reaching a real lane never passes through a Windows shell. A `.cmd` or `.bat` lane that cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, rather than run through `cmd.exe`: a batch file re-reads its arguments after `cmd.exe` has parsed them once, and no escaping fully contains a prompt through both passes. Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested on Windows too, including SIGTERM/SIGINT to the wrapper (Windows has no OS-level signals: both terminate it unconditionally, verified there rather than treated the same as POSIX). Five narrow skips remain on Windows, each for a POSIX behavior the OS or the CI shell genuinely does not have, and each named here because a test that is quietly skipped reads as a test that passed: `statSync().mode`'s executable bit (NTFS has none, so that one assertion is conditional inside a test that otherwise runs everywhere); a lane dying mid-run from a real POSIX signal (a real Windows lane cannot die "by signal"); running `weekly-audit.sh`'s watchdog functions for real under Git Bash's job control, both the end-to-end run and the `bounded()` timeout check (the script itself only ever runs on the Ubuntu box it targets); and a `mkfifo` FIFO at the rules path, the one case that proves `route-gate.mjs` cannot HANG on a non-regular file, since Windows has no `mkfifo` to build one (the guard behind it is covered on every OS by a directory at the same path); and an untracked `mkfifo` FIFO in the repository `cli-run --audit` sizes, the case that proves `--effort auto` never opens a non-regular file (the symlink half of that test runs on every OS). The list is not prose on trust: `test/prose.test.js` counts every `skip:` in the suite and fails if one of them is not documented here.
|
|
254
150
|
|
|
255
|
-
|
|
256
|
-
node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
|
|
257
|
-
node bin/cli-run.mjs --doctor # prints what each lane is pinned to, and what is not pinned
|
|
258
|
-
```
|
|
151
|
+
**Privacy.** The installer sends no telemetry and makes no network call of its own once it is running. Two things around that are worth being exact about:
|
|
259
152
|
|
|
260
|
-
|
|
261
|
-
|
|
262
|
-
runs that never reached a lane. It does not log an actual. One lane of five
|
|
263
|
-
(grok) reports a model id in its own output and the other four report none, so
|
|
264
|
-
an actual field would be present for one lane and missing for four, and it
|
|
265
|
-
would be a provider-supplied string, which the durable log deliberately never
|
|
266
|
-
holds.
|
|
153
|
+
- `npx model-orchestrator` is itself a download: npm fetches this package from the registry before any of it runs. `npm install -g model-orchestrator` once, then run `model-orchestrator`, if you would rather that happen exactly one time.
|
|
154
|
+
- A missing vendor CLI is *printed*, not installed. In an interactive run the installer offers to run one pinned `npm install -g` per package and only runs the ones you answer yes to; with `--yes` or `--no-install` it answers no for you and prints the command instead. Vendor shell installers (Antigravity, Grok) are only ever printed, alongside the `curl … | less` you would use to read one before running it.
|
|
267
155
|
|
|
268
|
-
|
|
156
|
+
`cli-run` talks to nothing but the vendor CLI you name.
|
|
269
157
|
|
|
270
|
-
### How do I cut token usage across Claude Code, Codex and Gemini?
|
|
271
158
|
|
|
272
|
-
|
|
159
|
+
</details>
|
|
273
160
|
|
|
274
|
-
|
|
161
|
+
<details>
|
|
162
|
+
<summary><strong>Vendor version compatibility</strong></summary>
|
|
275
163
|
|
|
276
|
-
The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly never gets a frontier token.
|
|
277
164
|
|
|
278
|
-
|
|
165
|
+
**This package detects that a binary exists. It does not check its version, and a present binary is not a working lane.** `--doctor` reports presence, and with `--run` sends one lane a one-word canary; neither validates that the vendor's flags, output shape or auth still match what the generated files assume.
|
|
279
166
|
|
|
280
|
-
|
|
167
|
+
The lane wiring and the output judges were written against these versions, which are the ones this release was exercised on:
|
|
281
168
|
|
|
282
|
-
|
|
169
|
+
<!-- vendor-table:start -->
|
|
283
170
|
|
|
284
|
-
|
|
171
|
+
| Lane | Vendor | Version this release was built against | Where that number is proved |
|
|
172
|
+
|---|---|---|---|
|
|
173
|
+
| `claude` | Anthropic | 2.1.226 | the npm pin the installer writes, `@anthropic-ai/claude-code@2.1.226` |
|
|
174
|
+
| `codex` | OpenAI | 0.153.4 | `test/fixtures/codex-0.153.4.jsonl`, a recorded run |
|
|
175
|
+
| `agy` | Google | 1.1.27 | `test/fixtures/agy-1.1.27.jsonl`, a recorded run |
|
|
176
|
+
| `grok` | xAI | 1.0.5 | `test/fixtures/grok-1.0.5.json`, a recorded run |
|
|
177
|
+
| `hermes` | Nous Research | 0.20.0 | `test/fixtures/hermes-0.20.0.txt`, a recorded run |
|
|
178
|
+
| `qwen` | Alibaba | 0.22.3 | `test/fixtures/qwen-0.22.3-nokey.json`, a recorded run |
|
|
179
|
+
| `ollama` | Ollama | 0.33.3 | the pinned image the level 3 box runs, `ollama/ollama:0.33.3` |
|
|
285
180
|
|
|
286
|
-
|
|
181
|
+
Generated from `src/catalog.js` by `npm run gen:catalog`; `npm test` fails if this table and the catalog disagree. Fixtures were captured 2026-09-06.
|
|
287
182
|
|
|
288
|
-
|
|
183
|
+
<!-- vendor-table:end -->
|
|
289
184
|
|
|
290
|
-
|
|
185
|
+
One number per lane, and it is the same number the installer pins: where a lane installs from npm, `builtAgainst` in the catalog *is* the pin, so "built against" and "pinned to" can never be two answers. That pin is a floor, not a ceiling: these CLIs ship breaking flag changes on their own schedules, so a newer version may work perfectly, or may change a flag the generated wiring passes. When a lane starts failing after a vendor upgrade, compare against this table first.
|
|
291
186
|
|
|
292
|
-
|
|
293
|
-
- A missing vendor CLI is *printed*, not installed. In an interactive run the installer offers to run one pinned `npm install -g` per package and only runs the ones you answer yes to; with `--yes` or `--no-install` it answers no for you and prints the command instead. Vendor shell installers (Antigravity, Grok) are only ever printed, alongside the `curl … | less` you would use to read one before running it.
|
|
187
|
+
**The live canary runs on your machine, with your credentials.** That is what `node bin/cli-run.mjs --doctor --run` is: it sends every enabled lane one tiny prompt through your own sign-ins and reports `canary ok` or `canary FAILED rc=` per lane. Run it after install, and again after any vendor upgrade.
|
|
294
188
|
|
|
295
|
-
|
|
189
|
+
It deliberately does not run in this repository's CI. A canary is only meaningful against real credentials, and there are no credentials a maintainer could supply that would tell **you** anything about **your** lanes: your sign-ins, your quota, your vendor versions. A maintainer-credential canary in CI would prove one machine works and bill someone per run to do it. So CI runs the full suite against stub lanes on Ubuntu, macOS and Windows, Node 18/20/22, plus a packaged install into a clean consumer, and the live check ships to you instead.
|
|
190
|
+
|
|
191
|
+
|
|
192
|
+
</details>
|
|
296
193
|
|
|
297
194
|
## Contributing
|
|
298
195
|
|
|
@@ -300,11 +197,7 @@ Add an AI to `src/catalog.js` and every prompt, table, config and doc picks it u
|
|
|
300
197
|
|
|
301
198
|
## Credits
|
|
302
199
|
|
|
303
|
-
- [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for
|
|
304
|
-
separating role, complexity and stakes instead of compressing them into one
|
|
305
|
-
scale, for recording the model and effort a lane was actually asked for, and
|
|
306
|
-
for verifying findings before they trigger repairs. All three shipped in
|
|
307
|
-
0.1.14.
|
|
200
|
+
- [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for separating role, complexity and stakes instead of compressing them into one scale, for recording the model and effort a lane was actually asked for, and for verifying findings before they trigger repairs. All three shipped in 0.1.14.
|
|
308
201
|
|
|
309
202
|
## License
|
|
310
203
|
|
package/docs/README.md
CHANGED
|
@@ -1,6 +1,13 @@
|
|
|
1
1
|
# docs/
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Reference, then reading. The reference pages carry the detail the README links out to; the three parts explain the thinking behind each level and how to grow from one to the next.
|
|
4
|
+
|
|
5
|
+
| Reference | What is in it | File |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| Installing | every flag, the two folders a run writes to, headless examples, the full file list | [install.md](install.md) |
|
|
8
|
+
| How it routes | role, complexity and stakes; the three verifier agents; pinning model and effort | [how-it-routes.md](how-it-routes.md) |
|
|
9
|
+
| Guarantees | what is enforced by code, delegated to a vendor flag, or only an instruction | [guarantees.md](guarantees.md) |
|
|
10
|
+
| Companion tools | codecalc, obsidian-tc and Context7: what each closes and what it needs first | [companions.md](companions.md) |
|
|
4
11
|
|
|
5
12
|
| Part | Read if | File |
|
|
6
13
|
|---|---|---|
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
# Companion tools
|
|
2
|
+
|
|
3
|
+
All optional. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc,context7` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
|
|
4
|
+
|
|
5
|
+
## Companion tools (all optional)
|
|
6
|
+
|
|
7
|
+
An orchestrator routes work. It does not make a model stop guessing numbers, it does not give it a memory, and it does not make it check a library's current docs before writing against it. Three tools close those gaps: codecalc and obsidian-tc are from the same maintainer, Context7 is from Upstash. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc,context7` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
|
|
8
|
+
|
|
9
|
+
| Tool | Closes | Default | You need first |
|
|
10
|
+
|---|---|---|---|
|
|
11
|
+
| [codecalc](https://github.com/The-40-Thieves/codecalc) | guessed numbers, comparisons, complexity and equivalence claims: exact arithmetic, code execution in 31 languages, SMT logic checks, `verify_translation` / `verify_optimization`; offline, no key | yes | Python 3.10+ and `uv`. `uvx 'codecalc[full]' setup --write` registers it with Claude Code, Claude Desktop, Cursor, VS Code, Zed; snippets for Codex, Antigravity, Qwen Code are written for you |
|
|
12
|
+
| [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes with a confirmation gate, folder ACLs, a poison scan on inferred writes; 163 tools, local by default; AGPL-3.0 | no | an Obsidian vault folder; Node 24+ or Bun 1.1+ (stricter than this installer); Ollama with `nomic-embed-text` or a cloud embeddings key; the Obsidian app and its Local REST API plugin only for live bridge tools. Skip it if you do not keep notes in Obsidian |
|
|
13
|
+
| [Context7](https://github.com/upstash/context7) | stale library recall: current, version-specific docs and code examples pulled into the prompt for any library, SDK, API or CLI; hosted, or `npx` locally; MIT | no | nothing to install for the hosted endpoint; Node.js 18+ for the local alternative; a free API key is optional, for a higher rate limit. Always makes a network call, unlike the other two: skip it offline |
|
|
14
|
+
|
|
15
|
+
Context7 pairs with codecalc rather than duplicating it: Context7 tells the agent what a library is documented to do on this version, codecalc runs the code and proves what it actually does. Docs never stand as proof on their own, and where the two disagree the run wins.
|
|
16
|
+
|
|
17
|
+
Whether or not you select them, every level carries the three rules they serve: `protocols/numbers-and-logic.md` (when calling a calculator is mandatory, how to report a computed figure, why a thought log is not evidence), `protocols/memory-and-record.md` (search before writing, the folder index is part of the change, one writer, inferred content marked as inferred), and `protocols/docs-then-prove.md` (current docs before writing a call, then a run proves it, the run wins on disagreement).
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
# What holds, and what is only asked for
|
|
2
|
+
|
|
3
|
+
Read this before running any of it unattended.
|
|
4
|
+
|
|
5
|
+
Most of what this package ships is text an agent is asked to follow. Be clear about which is which before relying on it unattended.
|
|
6
|
+
|
|
7
|
+
| Property | How it holds |
|
|
8
|
+
|---|---|
|
|
9
|
+
| Installer writes only inside `--dir` and `--project`, never a secret, never over a document without `--force` (or `--update-docs`, which touches only documents provably untouched since a previous run); machine-owned config always, runtime files only when provably untouched or with `--upgrade-runtime` | **enforced by code** (preflight, exclusive create, rollback, manifest hashes; tested) |
|
|
10
|
+
| `cli-run` exit codes, process-group kill on timeout and on SIGINT/SIGTERM, UTF-8-safe streaming, fixed-code durable log, `--expect-*` contracts with a pre-run snapshot | **enforced by code** (tested with stub lanes) |
|
|
11
|
+
| Codex audit lane runs read-only | **delegated to the vendor flag** (`--audit` → `--sandbox read-only`); commands and network still follow your codex config |
|
|
12
|
+
| Other lanes' permissions, sign-in state, model versions | **delegated to each vendor's own config**; `--doctor` checks presence, not versions |
|
|
13
|
+
| Gateway binds to loopback, keys by name only | **enforced in the generated files**; whether the gateway authenticates is your environment |
|
|
14
|
+
| Lane selection, tiers, privacy classes, one-writer, escalation, the protocols | **agent instructions**. Nothing here stops an agent that ignores its rules; the task bundle and the protocols make ignoring them visible, not impossible |
|
|
15
|
+
| Weekly audit bounded, previous report preserved | **enforced in the generated script and unit** (watchdog, temp-and-rename, `TimeoutStartSec`) |
|
|
16
|
+
|
|
17
|
+
If you need a property in the third row to be enforced, that is a router, a policy engine or a sandbox, and this package does not claim to be one.
|
|
18
|
+
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# How it routes
|
|
2
|
+
|
|
3
|
+
The reasoning behind the lane choice, and the three verifier agents that keep a routed answer honest. Install mechanics are in [install.md](install.md).
|
|
4
|
+
|
|
5
|
+
## Routing by role, complexity and stakes
|
|
6
|
+
|
|
7
|
+
Role picks the agent. Two more inputs move the choice, and they move it in
|
|
8
|
+
different directions, so `TIERS.md` states them separately rather than folding
|
|
9
|
+
them into the role:
|
|
10
|
+
|
|
11
|
+
- **Complexity moves the effort.** A worker executing a finished plan needs less
|
|
12
|
+
reasoning than the reviewer judging its output. When the plan is airtight the
|
|
13
|
+
spec is carrying the thinking.
|
|
14
|
+
- **Stakes move the tier and the reader.** Security, privacy, data loss and
|
|
15
|
+
irreversible changes buy the challenge lane, a named check, a rollback path or
|
|
16
|
+
a human yes. A one-line change to an auth check is simple and high-stakes at
|
|
17
|
+
the same time, and it is the stakes that decide.
|
|
18
|
+
|
|
19
|
+
Stakes means what a mistake would cost: a security hole, leaked personal data,
|
|
20
|
+
lost data, or something you can't undo. Most tasks are low-stakes and route
|
|
21
|
+
normally.
|
|
22
|
+
|
|
23
|
+
The top of the ladder is bought with evidence: a reproduced failure, an
|
|
24
|
+
unresolved checkpoint, an irreversible change. A task that merely feels hard is
|
|
25
|
+
a deep-tier task, not an escalation.
|
|
26
|
+
|
|
27
|
+
## A finding is a claim, not a fact
|
|
28
|
+
|
|
29
|
+
Review findings do not go straight to a repair. `finding-verifier` reads the
|
|
30
|
+
cited line, states what would trigger the problem, then hunts for the guard,
|
|
31
|
+
caller or test that makes it impossible, and returns **CONFIRMED**,
|
|
32
|
+
**NOT_REPRODUCED** or **INCONCLUSIVE** per finding. Only CONFIRMED earns a
|
|
33
|
+
change. Use a different model family from the one that produced the finding
|
|
34
|
+
where you have one: a family asked to check its own claim tends to agree with
|
|
35
|
+
itself.
|
|
36
|
+
|
|
37
|
+
## Two more fast-tier checks
|
|
38
|
+
|
|
39
|
+
`done-verifier` probes the artifact a tracker item's done-signal names (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything itself. It carries no file-editing tools, but on claude-code it does carry `Bash` for those probes (`git log`, `grep`, `wc -l`, `test -f`); staying to read-only commands there is a rule in its prompt, not a restriction on the tool grant, and its own description says so. On agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `reader` is the one that is read-only by tool grant on both: no `Write`, `Edit`, or `Bash`. It reads and digests many files or notes and hands back exactly what the brief asked for, cited by `path:line`; it never classifies, tags or writes, which is what separates it from `bulk-worker`. Both ship in the claude-code and agy agent sets, at the fast tier.
|
|
40
|
+
|
|
41
|
+
## Pin the route, or know that you did not
|
|
42
|
+
|
|
43
|
+
A lane with no `--model`, no `--effort` and no `defaults` entry in
|
|
44
|
+
`bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
|
|
45
|
+
CLI configured months ago at a low reasoning effort keeps auditing at that
|
|
46
|
+
effort while your routing docs describe a second-opinion pass.
|
|
47
|
+
|
|
48
|
+
```bash
|
|
49
|
+
node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
|
|
50
|
+
node bin/cli-run.mjs --doctor # prints what each lane is pinned to, and what is not pinned
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Every run logs the model and effort **requested** and where the request came
|
|
54
|
+
from: `flag`, `lanes.json`, or `lane_default`, on every record including the
|
|
55
|
+
runs that never reached a lane. It does not log an actual. One lane of five
|
|
56
|
+
(grok) reports a model id in its own output and the other four report none, so
|
|
57
|
+
an actual field would be present for one lane and missing for four, and it
|
|
58
|
+
would be a provider-supplied string, which the durable log deliberately never
|
|
59
|
+
holds.
|
|
60
|
+
|
package/docs/install.md
ADDED
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
# Installing
|
|
2
|
+
|
|
3
|
+
Everything the installer asks, writes and accepts as a flag. The short version is in the [README](../README.md).
|
|
4
|
+
|
|
5
|
+
## What a run does
|
|
6
|
+
|
|
7
|
+
|
|
8
|
+
The installer asks a few things, then writes a folder:
|
|
9
|
+
|
|
10
|
+
1. **Which level?** 1 beginner · 2 intermediate · 3 advanced
|
|
11
|
+
2. **Which AIs do you have access to?** (it marks the ones already on your PATH)
|
|
12
|
+
3. **Which one is your primary agent?** (the one that runs the system)
|
|
13
|
+
|
|
14
|
+
It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions (and, on Claude Code, two hook scripts) go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: follow the generated README. Inspect the manifest and remove only the individual managed subagent files you no longer need, preserve edited or pre-existing files, and remove your manually pasted activation block. Never delete a shared subagent folder.
|
|
15
|
+
|
|
16
|
+
## Plans and automatic effort
|
|
17
|
+
|
|
18
|
+
State known subscription plans with `--plans codex=pro-20x,agy=ultra-5x`. The generated guidance uses plan headroom to allocate volume only. It never changes capability or independent-review rules. `--effort-auto` is explicit consent to set `auto` only for selected high or max headroom CLI lanes. Auto chooses medium below 4,000 prompt characters and high otherwise, never higher. A codex audit is always high. Name `xhigh` explicitly for security-critical or irreversible work.
|
|
19
|
+
## The two folders every run writes to
|
|
20
|
+
|
|
21
|
+
An install has two targets, and a scripted run should set both.
|
|
22
|
+
|
|
23
|
+
| Flag | Default | What lands there |
|
|
24
|
+
|---|---|---|
|
|
25
|
+
| `--dir` | `./ai-orchestrator` | the docs, protocols and (level 2+) `bin/cli-run.mjs`. Named after what it contains, not after this package, so a project can hold one without looking like a checkout of it. Pass `--dir ./model-orchestrator` if you prefer the package name. |
|
|
26
|
+
| `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets three hook scripts in `.claude/hooks/`, wired by a settings snippet you merge yourself. |
|
|
27
|
+
|
|
28
|
+
`--project` defaulting to the current directory is the one that surprises people: run the command from your home folder with Claude Code as the primary and five agent files land in your home folder. The installer prints the resolved project path in the plan and says when you left it at the default. Set it.
|
|
29
|
+
|
|
30
|
+
## Non-interactive
|
|
31
|
+
|
|
32
|
+
```bash
|
|
33
|
+
# both targets set: docs in ./ai-orchestrator, subagents into ./my-app/.claude/agents
|
|
34
|
+
npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code \
|
|
35
|
+
--dir ./ai-orchestrator --project ./my-app
|
|
36
|
+
# --yes selects the recommended companion tool (codecalc), which writes CODECALC.md and mcp/ snippets.
|
|
37
|
+
# Add --no-tools for none, or --tools codecalc,obsidian-tc,context7 to choose.
|
|
38
|
+
|
|
39
|
+
npx model-orchestrator --yes --level 3 --ais claude-code,codex,agy,grok,hermes,qwen,ollama --apis anthropic,openrouter --dry # print the plan, write nothing
|
|
40
|
+
npx model-orchestrator --yes --level 2 --ais claude-code,codex --project ~/my-app --dir ~/my-app/ai-orchestrator --no-tools # subagents into ~/my-app/.claude/agents
|
|
41
|
+
npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator --project . --update-docs # added a lane: regenerate the docs you never edited
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
## What gets written (level 3, everything)
|
|
45
|
+
|
|
46
|
+
```
|
|
47
|
+
ai-orchestrator/
|
|
48
|
+
README.md start here, written for your level and your AIs
|
|
49
|
+
ORCHESTRATOR.md single-agent routing rules (level 1)
|
|
50
|
+
TASK_BUNDLE.md the brief every delegation carries
|
|
51
|
+
protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record · docs-then-prove
|
|
52
|
+
CODECALC.md OBSIDIAN-TC.md CONTEXT7.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
|
|
53
|
+
<project>/.claude/agents/ one per tier plus finding-verifier, done-verifier, reader, at the PROJECT root (if Claude Code is primary)
|
|
54
|
+
<project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart) + route-metrics.mjs (all five: see "Measuring routing" below), Claude Code only
|
|
55
|
+
CLAUDE.snippet.md the block to paste into your CLAUDE.md
|
|
56
|
+
settings.hooks.snippet.json the hooks block to merge into .claude/settings.json (Claude Code only)
|
|
57
|
+
ROUTING.md multi-lane decision tree (level 2+)
|
|
58
|
+
TIERS.md DELEGATION_MATRIX.md RESEARCH_TRIAGE.md CLI-RUN.md
|
|
59
|
+
bin/cli-run.mjs bin/lanes.json (node bin/cli-run.mjs --doctor is the smoke test)
|
|
60
|
+
vm/ gateway config, compose, box rules, privacy gates, jobs/ (level 3)
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
## Repo layout
|
|
64
|
+
|
|
65
|
+
| Folder | What |
|
|
66
|
+
|---|---|
|
|
67
|
+
| [`bin/`](bin/README.md) | `cli.js` (the installer) and `cli-run.mjs` (the lane runner) |
|
|
68
|
+
| [`src/`](src/README.md) | the catalog, the pure planner, detection, rendering |
|
|
69
|
+
| [`templates/`](templates/README.md) | everything the installer can write, by level, plus `tools/` for companions |
|
|
70
|
+
| [`docs/`](docs/README.md) | the three parts and the catalog |
|
|
71
|
+
| [`plugin/`](plugin/README.md) | the Claude Code plugin, generated from `templates/` by `npm run gen:plugin`; `.claude-plugin/marketplace.json` at the root lists it |
|
|
72
|
+
| [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu, macOS and Windows, Node 18/20/22 |
|
|
73
|
+
|
package/llms.txt
CHANGED
|
@@ -10,7 +10,11 @@ Claude Code plugin: `/plugin marketplace add aunysillyme/model-orchestrator`, th
|
|
|
10
10
|
|
|
11
11
|
## Docs
|
|
12
12
|
|
|
13
|
-
- [README](https://github.com/aunysillyme/model-orchestrator/blob/main/README.md): what it
|
|
13
|
+
- [README](https://github.com/aunysillyme/model-orchestrator/blob/main/README.md): what it is, a real dry-run plan, the levels, the AIs, the principles
|
|
14
|
+
- [Installing](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/install.md): every flag, the two folders a run writes to, headless examples, the full file list
|
|
15
|
+
- [How it routes](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/how-it-routes.md): role, complexity and stakes; the three verifier agents; pinning model and effort per lane
|
|
16
|
+
- [Guarantees](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/guarantees.md): what is enforced by code, what is delegated to a vendor flag, what is only an instruction
|
|
17
|
+
- [Companion tools](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/companions.md): codecalc, obsidian-tc and Context7, what each closes and what it needs first
|
|
14
18
|
- [Part 1: beginner](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-1-beginner.md): one agent, tiers, the task bundle every delegation carries
|
|
15
19
|
- [Part 2: intermediate](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-2-intermediate.md): several agent CLIs, the lane runner, pinning model and effort per lane
|
|
16
20
|
- [Part 3: advanced](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-3-advanced.md): the VM, the gateway, scheduled jobs, privacy gates
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "model-orchestrator",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.27",
|
|
4
4
|
"description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
@@ -48,7 +48,7 @@ Pinned form, if you want the version this installer was released against: `npx -
|
|
|
48
48
|
|
|
49
49
|
## Register it with your agent (snippets in `mcp/`)
|
|
50
50
|
|
|
51
|
-
Every snippet below ships **keyless**: the remote ones point at the hosted endpoint with no `Authorization` header at all (Qwen Code's snippet keeps the non-credential `Accept` header upstream itself ships), and
|
|
51
|
+
Every snippet below ships **keyless**: the remote ones point at the hosted endpoint with no `Authorization` header at all (Qwen Code's snippet keeps the non-credential `Accept` header upstream itself ships), and both Zed snippets run the local, version-pinned `npx` server with no key in its `env` block. That is deliberate, not an oversight: see "Higher rate limits" next for why a header is not shipped by default.
|
|
52
52
|
|
|
53
53
|
| Agent | File to edit | Snippet |
|
|
54
54
|
|---|---|---|
|
|
@@ -56,22 +56,27 @@ Every snippet below ships **keyless**: the remote ones point at the hosted endpo
|
|
|
56
56
|
| Claude Desktop | no config file: `Settings > Connectors > Add Custom Connector`, name `Context7`, URL `https://mcp.context7.com/mcp` | none, it is a UI step |
|
|
57
57
|
| Cursor | one-click install in the upstream README, or `~/.cursor/mcp.json` | `mcp/context7.mcpServers.json` |
|
|
58
58
|
| VS Code | `.vscode/mcp.json` (key is `servers`, remote type is `http`) | `mcp/context7.vscode.mcp.json` |
|
|
59
|
-
| Zed | `~/.config/zed/settings.json` (key is `context_servers`); upstream ships only a local (`npx`) config for Zed, so this snippet runs the server locally rather than remote.
|
|
59
|
+
| Zed | `~/.config/zed/settings.json` (key is `context_servers`); upstream ships only a local (`npx`) config for Zed, so this snippet runs the server locally rather than remote. On Windows a current Zed starts it as written; see "Local `npx` on Windows" below for the two cases that need the `.windows` file | `mcp/context7.zed.settings.json` · fallback: `mcp/context7.zed.windows.settings.json` |
|
|
60
60
|
| Codex CLI | `~/.codex/config.toml` | `mcp/context7.codex.config.toml` |
|
|
61
61
|
| Antigravity `agy` | its MCP config file | `mcp/context7.agy.mcp_config.json` |
|
|
62
62
|
| Qwen Code | `~/.qwen/settings.json` under `mcpServers`; note the field is `httpUrl`, not `url`, a different shape from every other client here | `mcp/context7.qwen.settings.json` |
|
|
63
63
|
|
|
64
64
|
### Local `npx` on Windows
|
|
65
65
|
|
|
66
|
-
On Windows `npx` is a batch file (`npx.cmd`)
|
|
66
|
+
On Windows `npx` is a batch file (`npx.cmd`) that a bare `CreateProcess` will not start, and Context7's own client guide says to wrap it in `cmd` on Windows ([all clients, Windows section](https://context7.com/docs/resources/all-clients)). That guide is about clients in general. **Zed is not one that needs it:** since its PR #42382, "Use shell to launch MCP and ACP servers" (2025-12-10), `crates/context_server/src/transport/stdio_transport.rs` builds every MCP launch through `ShellBuilder::new(&Shell::System, ..)`, so the shell resolves `npx` and a plain `"command": "npx"` starts. The `.windows` snippet is kept for two cases it does fix: a Zed older than that build, and the PowerShell one below. This only affects the local form either way: every remote snippet above connects over HTTPS and spawns nothing.
|
|
67
67
|
|
|
68
68
|
| Where | Use |
|
|
69
69
|
|---|---|
|
|
70
70
|
| Zed on macOS or Linux | `mcp/context7.zed.settings.json` (`"command": "npx"`) |
|
|
71
|
-
| Zed on Windows | `mcp/context7.zed.
|
|
72
|
-
|
|
|
71
|
+
| Zed on Windows, current build | `mcp/context7.zed.settings.json` as well; the shell Zed opens resolves `npx` |
|
|
72
|
+
| Zed on Windows, older build or a blocked `npx.ps1` | `mcp/context7.zed.windows.settings.json` (`"command": "cmd"`, `"args": ["/d", "/c", "npx", ...]`) |
|
|
73
|
+
| Any other client, local alternative, on Windows | usually nothing: Cursor, VS Code, the Codex CLI and every client on the stock `@modelcontextprotocol/sdk` resolve `PATHEXT` themselves. If yours genuinely does not, the change is `"command": "cmd"` with `"/d", "/c", "npx"` in front of the existing args |
|
|
73
74
|
|
|
74
|
-
|
|
75
|
+
The one Windows failure that is not about `.cmd`: Zed prefers PowerShell for the system shell (`get_windows_system_shell` in `crates/gpui_util/src/lib.rs` falls back to `cmd.exe` only when PowerShell is missing), and PowerShell resolves a bare `npx` to npm's `npx.ps1` shim when one is installed. Under the `Restricted` execution policy that is Windows' client default, a `.ps1` will not run. Check `Get-ExecutionPolicy`, and if that is the cause the `.windows` snippet is the fix.
|
|
76
|
+
|
|
77
|
+
The `/d` in those args is deliberate: without it `cmd` first runs whatever is in the Command Processor `AutoRun` registry value, which can print non-JSON into the protocol stream. `cross-spawn`, the library the MCP TypeScript SDK uses for exactly this job, passes `/d /s /c` for the same reason.
|
|
78
|
+
|
|
79
|
+
Not yet run on a Windows machine by this project.
|
|
75
80
|
|
|
76
81
|
Merge the block; do not replace the file. `mcp/context7.mcpServers.json` (Cursor) and `mcp/context7.claude-code.mcp.json` look alike but are not interchangeable: Claude Code requires the `"type": "http"` field and Cursor's own docs show plain `{"url": ...}` with no `type` at all.
|
|
77
82
|
|
|
@@ -54,6 +54,24 @@ Every snippet points `OBSIDIAN_TC_CONFIG` at your config file. Replace `/ABSOLUT
|
|
|
54
54
|
|
|
55
55
|
Merge the block; do not replace the file.
|
|
56
56
|
|
|
57
|
+
### Local `npx` on Windows
|
|
58
|
+
|
|
59
|
+
Every snippet here spawns `npx`, and on Windows `npx` is a batch file (`npx.cmd`) that a bare `CreateProcess` will not start. **All five clients handle that themselves, so there is no Windows snippet here and the files above are the ones to use on every OS.** Each was read rather than assumed, because the usual advice ("on Windows, wrap it in `cmd /c`") is about clients in general and is wrong for all five of these:
|
|
60
|
+
|
|
61
|
+
| Client | How it starts `"command": "npx"` on Windows | Checked against |
|
|
62
|
+
|---|---|---|
|
|
63
|
+
| Zed | hands the whole command to the system shell | `crates/context_server/src/transport/stdio_transport.rs` builds through `ShellBuilder::new(&Shell::System, ..)`, from PR #42382 "Use shell to launch MCP and ACP servers" (2025-12-10) |
|
|
64
|
+
| VS Code | resolves the executable, then re-spawns it through a shell | `src/vs/workbench/api/node/extHostMcpNode.ts`, `formatSubprocessArguments` resolves the extension and sets `shell: true` when it ends in `.bat` or `.cmd` |
|
|
65
|
+
| Codex CLI | resolves `PATHEXT` itself before spawning | `codex-rs/rmcp-client/src/program_resolver.rs`, whose Windows arm calls `which::which_in` so that "tools like `npx`, `pnpm`, and `yarn` work correctly on Windows" |
|
|
66
|
+
| Cursor | the MCP SDK it bundles spawns through `cross-spawn` | Cursor 3.18.9 ships `@modelcontextprotocol/sdk` 1.25.1, whose `dist/esm/client/stdio.js` opens with `import spawn from 'cross-spawn'`; `cross-spawn/lib/parse.js` resolves the command and, when it is not an `.exe`, re-invokes it as `%COMSPEC% /d /s /c "<escaped>"` |
|
|
67
|
+
| Antigravity `agy` | same SDK path, through the Gemini CLI config namespace it reads | Gemini CLI pins `@modelcontextprotocol/sdk` 1.23.0 in `packages/core/package.json` and hands `mcpServerConfig.command` to that transport; 1.23.0's `client/stdio.js` imports `cross-spawn` too |
|
|
68
|
+
|
|
69
|
+
The SDK point covers more than these two clients: every published `@modelcontextprotocol/sdk` from 1.23.0 through 1.30.0 depends on `cross-spawn ^7.0.5` and uses it in the stdio client, so any client that connects through the stock TypeScript SDK inherits the same `PATHEXT` resolution. `shell: false` in that transport is not the whole story, and reading only that line is how a client gets mistaken for one that cannot start `npx`.
|
|
70
|
+
|
|
71
|
+
One Windows case can still fail, and it is not about `.cmd`. Zed prefers PowerShell for the system shell (`get_windows_system_shell` in `crates/gpui_util/src/lib.rs` falls back to `cmd.exe` only when PowerShell is missing), and PowerShell resolves a bare `npx` to npm's `npx.ps1` shim when one is installed. Under the `Restricted` execution policy that is Windows' client default, running a `.ps1` is blocked. If Zed reports that the server would not start, check `Get-ExecutionPolicy` first, and if that is the cause, change the Zed entry by hand to `"command": "cmd"` with `"args": ["/d", "/c", "npx", "-y", "obsidian-tc"]`, keeping the rest of the block. `/d` is there on purpose: it skips any Command Processor `AutoRun` command, which would otherwise run first and can print non-JSON into the protocol stream.
|
|
72
|
+
|
|
73
|
+
None of this was run on a Windows machine by this project. The five verdicts are from each client's own shipped code; the PowerShell case is from Zed's shell choice plus documented `Restricted` behaviour, and is the one worth reporting back if you hit it.
|
|
74
|
+
|
|
57
75
|
## Security posture, read before a second agent touches it
|
|
58
76
|
|
|
59
77
|
Zero-config mode boots with **auth off and no folder ACL**: anything that can reach the server has the same authority as raw filesystem access to the vault. That is acceptable only because the surface is local-only (the config refuses by default if you enable HTTP on a non-loopback host with auth off, and a DNS-rebinding guard protects loopback). Before exposing it to partially-trusted, remote or multi-agent callers, turn on `auth.mode: "jwt"` and set `acl.readPaths` / `writePaths` / `deletePaths` in the config file. Upstream `SECURITY.md` has the security notes and a private disclosure path.
|