model-orchestrator 0.1.35 → 1.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (124) hide show
  1. package/AGENTS.md +31 -21
  2. package/CHANGELOG.md +58 -1
  3. package/README.md +129 -110
  4. package/SECURITY.md +7 -3
  5. package/bin/README.md +57 -6
  6. package/bin/aunx.js +7 -0
  7. package/bin/cli-run.mjs +21 -15
  8. package/bin/cli.js +376 -257
  9. package/docs/README.md +15 -18
  10. package/docs/catalog.md +236 -44
  11. package/docs/companions.md +28 -10
  12. package/docs/guarantees.md +21 -12
  13. package/docs/how-it-routes.md +49 -42
  14. package/docs/install.md +141 -33
  15. package/docs/part-1-beginner.md +37 -45
  16. package/docs/part-2-intermediate.md +34 -52
  17. package/docs/part-3-advanced.md +36 -26
  18. package/docs/security-review-history.md +39 -0
  19. package/llms.txt +24 -25
  20. package/package.json +15 -8
  21. package/proof/README.md +100 -0
  22. package/proof/gate-demo.cast +9 -0
  23. package/proof/gate-demo.gif +0 -0
  24. package/proof/results.json +198 -0
  25. package/proof/scripts/check-gate.js +26 -0
  26. package/proof/scripts/install-time.js +16 -0
  27. package/proof/scripts/lib.js +73 -0
  28. package/proof/scripts/measure.js +15 -0
  29. package/proof/scripts/missing-results.js +30 -0
  30. package/proof/scripts/record-gate.js +38 -0
  31. package/proof/scripts/render.js +18 -0
  32. package/proof/scripts/runner-overhead.js +21 -0
  33. package/src/README.md +10 -3
  34. package/src/activation-ownership.js +19 -0
  35. package/src/apply-companions.js +104 -0
  36. package/src/apply-snippets.js +60 -28
  37. package/src/aunx.js +272 -0
  38. package/src/bounded-file.js +31 -0
  39. package/src/catalog.js +257 -121
  40. package/src/install.js +483 -212
  41. package/src/plugin.js +13 -4
  42. package/src/postinstall.js +57 -0
  43. package/src/roles.js +184 -0
  44. package/src/uninstall.js +128 -10
  45. package/templates/README.md +19 -2
  46. package/templates/advanced/README.md +2 -2
  47. package/templates/advanced/vm/ENVIRONMENT.md +8 -0
  48. package/templates/advanced/vm/PRIVACY_GATES.md +17 -19
  49. package/templates/advanced/vm/README.md +25 -20
  50. package/templates/advanced/vm/box-CLAUDE.md +19 -18
  51. package/templates/advanced/vm/docker-compose.yml +2 -1
  52. package/templates/advanced/vm/jobs/README.md +31 -2
  53. package/templates/advanced/vm/jobs/weekly-audit.service +7 -2
  54. package/templates/advanced/vm/jobs/weekly-audit.sh +24 -17
  55. package/templates/advanced/vm/setup-vm.sh +49 -2
  56. package/templates/agents/README.md +2 -2
  57. package/templates/agents/agy/README.md +20 -3
  58. package/templates/agents/agy/builder.md +11 -7
  59. package/templates/agents/agy/bulk-worker.md +9 -7
  60. package/templates/agents/agy/code-reviewer.md +13 -7
  61. package/templates/agents/agy/deep-planner.md +10 -7
  62. package/templates/agents/agy/done-verifier.md +13 -22
  63. package/templates/agents/agy/finding-verifier.md +14 -22
  64. package/templates/agents/agy/live-researcher.md +10 -7
  65. package/templates/agents/agy/reader.md +10 -12
  66. package/templates/agents/claude-code/README.md +18 -14
  67. package/templates/agents/claude-code/builder.md +10 -15
  68. package/templates/agents/claude-code/bulk-worker.md +8 -10
  69. package/templates/agents/claude-code/code-reviewer.md +11 -17
  70. package/templates/agents/claude-code/deep-planner.md +9 -11
  71. package/templates/agents/claude-code/done-verifier.md +12 -33
  72. package/templates/agents/claude-code/finding-verifier.md +13 -39
  73. package/templates/agents/claude-code/live-researcher.md +9 -11
  74. package/templates/agents/claude-code/reader.md +9 -18
  75. package/templates/agents/snippets/chat.md +9 -10
  76. package/templates/agents/snippets/claude-code.md +17 -18
  77. package/templates/agents/snippets/generic.md +9 -11
  78. package/templates/agents/snippets/route-gate.mjs +2 -2
  79. package/templates/agents/snippets/route-metrics.mjs +1 -1
  80. package/templates/agents/snippets/subagent-context.mjs +4 -4
  81. package/templates/beginner/ORCHESTRATOR.md +31 -36
  82. package/templates/beginner/README.md +1 -1
  83. package/templates/common/ACCEPTANCE_CHECKS.json +12 -0
  84. package/templates/common/CONTEXT.md +37 -0
  85. package/templates/common/DECISIONS.md +11 -0
  86. package/templates/common/README.md +24 -11
  87. package/templates/common/TASK_BRIEF.md +84 -0
  88. package/templates/common/protocols/README.md +14 -11
  89. package/templates/common/protocols/acceptance-checks.md +15 -0
  90. package/templates/common/protocols/build-protocol.md +91 -106
  91. package/templates/common/protocols/context-file.md +10 -0
  92. package/templates/common/protocols/decision-log.md +9 -0
  93. package/templates/common/protocols/deep-research.md +20 -34
  94. package/templates/common/protocols/docs-then-prove.md +13 -18
  95. package/templates/common/protocols/gap-analysis.md +15 -21
  96. package/templates/common/protocols/memory-and-record.md +21 -20
  97. package/templates/common/protocols/numbers-and-logic.md +20 -26
  98. package/templates/common/protocols/propagate.md +18 -27
  99. package/templates/intermediate/CLI-RUN.md +83 -113
  100. package/templates/intermediate/DELEGATION_MATRIX.md +9 -3
  101. package/templates/intermediate/README.md +3 -3
  102. package/templates/intermediate/RESEARCH_TRIAGE.md +23 -15
  103. package/templates/intermediate/ROUTING.md +54 -51
  104. package/templates/intermediate/TIERS.md +37 -76
  105. package/templates/tools/README.md +1 -1
  106. package/templates/tools/codecalc/CODECALC.md +4 -4
  107. package/templates/tools/codecalc/mcp/agy.mcp_config.json +1 -1
  108. package/templates/tools/codecalc/mcp/codex.config.toml +1 -1
  109. package/templates/tools/codecalc/mcp/mcpServers.json +1 -1
  110. package/templates/tools/codecalc/mcp/vscode.mcp.json +1 -1
  111. package/templates/tools/codecalc/mcp/zed.settings.json +1 -1
  112. package/templates/tools/context7/CONTEXT7.md +6 -10
  113. package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +3 -3
  114. package/templates/tools/obsidian-tc/mcp/obsidian-tc.agy.mcp_config.json +1 -1
  115. package/templates/tools/obsidian-tc/mcp/obsidian-tc.codex.config.toml +1 -1
  116. package/templates/tools/obsidian-tc/mcp/obsidian-tc.mcpServers.json +1 -1
  117. package/templates/tools/obsidian-tc/mcp/obsidian-tc.vscode.mcp.json +1 -1
  118. package/templates/tools/obsidian-tc/mcp/obsidian-tc.zed.settings.json +1 -1
  119. package/docs/audit-brief.md +0 -148
  120. package/scripts/README.md +0 -7
  121. package/scripts/gen-catalog.js +0 -81
  122. package/scripts/gen-plugin.js +0 -16
  123. package/scripts/record-demo.sh +0 -45
  124. package/templates/common/TASK_BUNDLE.md +0 -56
package/docs/README.md CHANGED
@@ -1,20 +1,17 @@
1
- # docs/
1
+ # Documentation
2
2
 
3
- Reference, then reading. The reference pages carry the detail the README links out to; the three parts explain the thinking behind each level and how to grow from one to the next.
3
+ Use these pages for setup, operation and evidence. The [front page](../README.md) gives the short introduction.
4
4
 
5
- | Reference | What is in it | File |
6
- |---|---|---|
7
- | Installing | every flag, the two folders a run writes to, headless examples, the full file list | [install.md](install.md) |
8
- | How it routes | role, complexity and stakes; the three verifier agents; pinning model and effort | [how-it-routes.md](how-it-routes.md) |
9
- | Guarantees | what is enforced by code, delegated to a vendor flag, or only an instruction | [guarantees.md](guarantees.md) |
10
- | Companion tools | codecalc, obsidian-tc and Context7: what each closes and what it needs first | [companions.md](companions.md) |
11
-
12
- | Part | Read if | File |
13
- |---|---|---|
14
- | 1 Beginner | you use one LLM or one agent and want it to route well | [part-1-beginner.md](part-1-beginner.md) |
15
- | 2 Intermediate | you have several AIs and want to call them through their CLIs from one orchestrator | [part-2-intermediate.md](part-2-intermediate.md) |
16
- | 3 Advanced | you want the whole thing running unattended on a virtual machine | [part-3-advanced.md](part-3-advanced.md) |
17
- | Audit brief | the security notes and the two second-opinion audit rounds this shipped with | [audit-brief.md](audit-brief.md) |
18
- | Catalog | what each AI and companion tool in the installer is for, how it installs, how it signs in | [catalog.md](catalog.md) |
19
-
20
- Each part ends with "what the installer gives you at this level" so the doc and the files agree.
5
+ | Page | What it gives you |
6
+ |---|---|
7
+ | [Install and upgrade](install.md) | One-confirm setup, edit menu, flags, project commands, safe updates and uninstall |
8
+ | [How routing works](how-it-routes.md) | Stack-dependent role assignment, model tiers and independent verification |
9
+ | [Guarantees](guarantees.md) | What code enforces and what the agent follows as instructions |
10
+ | [Companions](companions.md) | Optional projects, setup links and alternatives using existing tools |
11
+ | [Beginner](part-1-beginner.md) | Routing inside one agent or chat app |
12
+ | [Intermediate](part-2-intermediate.md) | Delegation across several AI CLIs |
13
+ | [Advanced](part-3-advanced.md) | Gateway templates and scheduled work on a Linux host |
14
+ | [Catalog](catalog.md) | Supported tools, capability facts, unverified values, installation and sign-in notes |
15
+ | [Security review history](security-review-history.md) | Review rounds, reproduced findings, fixes and regression tests |
16
+ | [Proof](../proof/README.md) | Dated measurements, methods, sample sizes and reproduction scripts |
17
+ | [Commands](../bin/README.md) | `aunx` subcommands and lane-runner exit codes |
package/docs/catalog.md CHANGED
@@ -1,144 +1,336 @@
1
1
  # Catalog
2
2
 
3
- Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrites it. Protocols shipped at every level: 7 (counted from `templates/common/protocols/`).
3
+ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrites it. Protocols shipped at every level: 10 (counted from `templates/common/protocols/`).
4
4
 
5
5
  ## Levels
6
6
 
7
7
  | Level | Name | Tagline | Gives |
8
8
  |---|---|---|---|
9
9
  | 1 | Beginner | one LLM or agent, routed well | tiers, task classification, every protocol, one agent set up to follow them |
10
- | 2 | Intermediate | several LLMs and agents, called through their CLIs | everything in Beginner plus cli-run, a delegation matrix, task bundles and three-engine research triage |
11
- | 3 | Advanced | everything above, plus a virtual machine that runs it unattended | everything in Intermediate plus a gateway config, scheduled jobs, a dispatch layer and privacy gates for a box |
10
+ | 2 | Intermediate | several LLMs and agents, called through their CLIs | everything in Beginner plus cli-run, a delegation matrix and multi-engine research triage |
11
+ | 3 | Advanced | everything above, plus templates for your always-on Linux machine | everything in Intermediate plus a gateway config, a scheduled review job, dispatch guidance and configurable privacy gates |
12
12
 
13
13
  ## AIs
14
14
 
15
15
  ### `claude-code` · Claude Code (Anthropic)
16
16
 
17
- - **Kind:** agent-cli · **Access:** subscription · **Lane:** A · **Level:** 1+
18
- - **Wins at:** orchestrator: routes, maps, builds, verifies, records
17
+ - **Kind:** agent-cli · **Billing:** subscription · **Level:** 1+
18
+ - **What it is:** Anthropic's terminal coding agent; its subagents load the project rules file
19
19
  - **Install:** `npm install -g @anthropic-ai/claude-code@2.1.226`
20
20
  - **Sign in:** run `claude` once and sign in with your Anthropic account
21
21
  - **Reads rules from:** `CLAUDE.md` · subagents in `.claude/agents/`
22
22
  - **Plans:**
23
- - Claude Pro (base headroom, checked 2026-09-12): https://support.claude.com/en/articles/11049762-choose-a-claude-plan
24
- - Claude Max 5x (high headroom, checked 2026-09-12): https://support.claude.com/en/articles/11049762-choose-a-claude-plan
25
- - Claude Max 20x (max headroom, checked 2026-09-12): https://support.claude.com/en/articles/11049762-choose-a-claude-plan
23
+ - Claude Pro (base headroom, checked 2026-09-12; tier models: unverified): https://claude.com/pricing
24
+ - Claude Max 5x (high headroom, checked 2026-09-12; tier models: unverified): https://claude.com/pricing
25
+ - Claude Max 20x (max headroom, checked 2026-09-12; tier models: unverified): https://claude.com/pricing
26
26
  - **Built against:** 2.1.226 (the same number the npm pin uses)
27
27
 
28
+ **Capability facts.** An unverified value needs a current capability check before use. Role assignments come from the selected stack, using these facts.
29
+
30
+ | Fact | Value |
31
+ |---|---|
32
+ | `modelFamily` | Anthropic |
33
+ | `kind` | agent-cli |
34
+ | `billing` | subscription |
35
+ | `pricing` | unverified |
36
+ | `headless` | yes |
37
+ | `cliRun` | no |
38
+ | `writesFiles` | yes |
39
+ | `readOnlyMode` | no |
40
+ | `liveWeb` | yes |
41
+ | `runsLocally` | no |
42
+ | `fanOut` | unverified |
43
+ | `contextWindow` | unverified |
44
+ | `loadsProjectRules` | yes |
45
+ | `agentDefinitions` | .claude/agents |
46
+
28
47
  ### `codex` · Codex CLI (OpenAI, ChatGPT plan)
29
48
 
30
- - **Kind:** agent-cli · **Access:** subscription · **Lane:** A · **Level:** 1+
31
- - **Wins at:** second coder and second-opinion reviewer (a different model family reading your diff)
49
+ - **Kind:** agent-cli · **Billing:** subscription · **Level:** 1+
50
+ - **What it is:** OpenAI's terminal coding agent on a ChatGPT plan; `--audit` runs it in a read-only filesystem sandbox
32
51
  - **Install:** `npm install -g @openai/codex@0.153.4`
33
52
  - **Sign in:** `codex login` (add `--device-auth` on a machine with no browser)
34
53
  - **Reads rules from:** `AGENTS.md`
35
54
  - **cli-run lane:** yes
36
55
  - **Plans:**
37
- - ChatGPT Plus (base headroom, checked 2026-09-12): https://learn.chatgpt.com/codex/pricing.md
38
- - ChatGPT Pro 5x (high headroom, checked 2026-09-12): https://learn.chatgpt.com/codex/pricing.md
39
- - ChatGPT Pro 20x (max headroom, checked 2026-09-12): https://learn.chatgpt.com/codex/pricing.md
56
+ - ChatGPT Plus (base headroom, checked 2026-09-12; tier models: unverified): https://learn.chatgpt.com/codex/pricing.md
57
+ - ChatGPT Pro 5x (high headroom, checked 2026-09-12; tier models: unverified): https://learn.chatgpt.com/codex/pricing.md
58
+ - ChatGPT Pro 20x (max headroom, checked 2026-09-12; tier models: unverified): https://learn.chatgpt.com/codex/pricing.md
40
59
  - **Built against:** 0.153.4 (the same number the npm pin uses)
41
60
 
61
+ **Capability facts.** An unverified value needs a current capability check before use. Role assignments come from the selected stack, using these facts.
62
+
63
+ | Fact | Value |
64
+ |---|---|
65
+ | `modelFamily` | OpenAI |
66
+ | `kind` | agent-cli |
67
+ | `billing` | subscription |
68
+ | `pricing` | unverified |
69
+ | `headless` | yes |
70
+ | `cliRun` | yes |
71
+ | `writesFiles` | yes |
72
+ | `readOnlyMode` | yes |
73
+ | `liveWeb` | unverified |
74
+ | `runsLocally` | no |
75
+ | `fanOut` | no |
76
+ | `contextWindow` | unverified |
77
+ | `loadsProjectRules` | no |
78
+ | `agentDefinitions` | unverified |
79
+
42
80
  ### `agy` · Antigravity CLI `agy` (Google AI plan)
43
81
 
44
- - **Kind:** agent-cli · **Access:** subscription · **Lane:** A · **Level:** 1+
45
- - **Wins at:** deep research sweeps and concurrent fan-out (its subagent call takes an array)
82
+ - **Kind:** agent-cli · **Billing:** subscription · **Level:** 1+
83
+ - **What it is:** Google's Antigravity terminal agent; one subagent call starts several children. fanOut: UNVERIFIED against a vendor doc; inherited from the 0.1.x catalog.
46
84
  - **Install:** vendor script (read it first): `https://antigravity.google/cli/install.sh`
47
- - **Sign in:** first run opens a device-code sign-in with your Google account
85
+ - **Sign in:** run `agy`; the first run opens a device-code sign-in with your Google account
48
86
  - **Reads rules from:** `GEMINI.md` · subagents in `.agents/agents/`
49
87
  - **cli-run lane:** yes
50
88
  - **Plans:**
51
- - Google AI Pro (base headroom, checked 2026-09-12): https://gemini.google/subscriptions/
52
- - Google AI Ultra 5x (high headroom, checked 2026-09-12): https://gemini.google/subscriptions/
53
- - Google AI Ultra 20x (max headroom, checked 2026-09-12): https://gemini.google/subscriptions/
89
+ - Google AI Pro (base headroom, checked 2026-09-12; tier models: unverified): https://gemini.google/subscriptions/
90
+ - Google AI Ultra 5x (high headroom, checked 2026-09-12; tier models: unverified): https://gemini.google/subscriptions/
91
+ - Google AI Ultra 20x (max headroom, checked 2026-09-12; tier models: unverified): https://gemini.google/subscriptions/
54
92
  - **Built against:** 1.1.27
55
93
  - **Note:** Gemini CLI was retired by Google in June 2026. agy is the successor. Do not install `gemini`.
56
94
 
95
+ **Capability facts.** An unverified value needs a current capability check before use. Role assignments come from the selected stack, using these facts.
96
+
97
+ | Fact | Value |
98
+ |---|---|
99
+ | `modelFamily` | Google |
100
+ | `kind` | agent-cli |
101
+ | `billing` | subscription |
102
+ | `pricing` | unverified |
103
+ | `headless` | yes |
104
+ | `cliRun` | yes |
105
+ | `writesFiles` | yes |
106
+ | `readOnlyMode` | no |
107
+ | `liveWeb` | unverified |
108
+ | `runsLocally` | no |
109
+ | `fanOut` | yes (UNVERIFIED against a vendor doc; inherited from the 0.1.x catalog.) |
110
+ | `contextWindow` | unverified |
111
+ | `loadsProjectRules` | unverified |
112
+ | `agentDefinitions` | .agents/agents |
113
+
57
114
  ### `grok` · Grok CLI (xAI, X Premium)
58
115
 
59
- - **Kind:** agent-cli · **Access:** subscription · **Lane:** A · **Level:** 1+
60
- - **Wins at:** X and live web reads at no per-call cost (its search tools bill on the API, not on the CLI)
116
+ - **Kind:** agent-cli · **Billing:** subscription · **Level:** 1+
117
+ - **What it is:** xAI's terminal agent with first-party X and web search tools; searches are covered by the subscription rather than billed per call. liveWeb: UNVERIFIED against a vendor doc; inherited from the 0.1.x catalog.
61
118
  - **Install:** vendor script (read it first): `https://x.ai/cli/install.sh`
62
119
  - **Sign in:** `grok login` (add `--device-auth` on a headless machine)
63
120
  - **cli-run lane:** yes
64
121
  - **Plans:**
65
- - SuperGrok (base headroom, checked 2026-09-12): https://x.ai/news/grok-build-cli
66
- - SuperGrok Plus (high headroom, checked 2026-09-12): https://x.ai/pricing
67
- - X Premium Plus (base headroom, checked 2026-09-12): https://x.ai/news/grok-build-cli
122
+ - SuperGrok (base headroom, checked 2026-09-12; tier models: unverified): https://x.ai/news/grok-build-cli
123
+ - SuperGrok Plus (high headroom, checked 2026-09-12; tier models: unverified): https://x.ai/pricing
124
+ - X Premium Plus (base headroom, checked 2026-09-12; tier models: unverified): https://x.ai/news/grok-build-cli
68
125
  - **Built against:** 1.0.5
69
126
 
127
+ **Capability facts.** An unverified value needs a current capability check before use. Role assignments come from the selected stack, using these facts.
128
+
129
+ | Fact | Value |
130
+ |---|---|
131
+ | `modelFamily` | xAI |
132
+ | `kind` | agent-cli |
133
+ | `billing` | subscription |
134
+ | `pricing` | unverified |
135
+ | `headless` | yes |
136
+ | `cliRun` | yes |
137
+ | `writesFiles` | yes |
138
+ | `readOnlyMode` | no |
139
+ | `liveWeb` | yes (UNVERIFIED against a vendor doc; inherited from the 0.1.x catalog.) |
140
+ | `runsLocally` | no |
141
+ | `fanOut` | no |
142
+ | `contextWindow` | unverified |
143
+ | `loadsProjectRules` | no |
144
+ | `agentDefinitions` | unverified |
145
+
70
146
  ### `hermes` · Hermes Agent (Nous Research)
71
147
 
72
- - **Kind:** agent-cli · **Access:** free · **Lane:** A · **Level:** 2+
73
- - **Wins at:** the free tier: rough drafts, first-pass summaries, cheap divergent reads, cron jobs on a box
148
+ - **Kind:** agent-cli · **Billing:** free · **Level:** 2+
149
+ - **What it is:** A free terminal agent that chains whichever providers you authenticate
74
150
  - **Install:** https://github.com/NousResearch/hermes-agent
75
151
  - **Sign in:** `hermes auth add <provider>` per provider; its own fallback chain handles outages
76
152
  - **cli-run lane:** yes
77
153
  - **Built against:** 0.20.0
78
154
 
155
+ **Capability facts.** An unverified value needs a current capability check before use. Role assignments come from the selected stack, using these facts.
156
+
157
+ | Fact | Value |
158
+ |---|---|
159
+ | `modelFamily` | unverified |
160
+ | `kind` | agent-cli |
161
+ | `billing` | free |
162
+ | `pricing` | unverified |
163
+ | `headless` | yes |
164
+ | `cliRun` | yes |
165
+ | `writesFiles` | yes |
166
+ | `readOnlyMode` | no |
167
+ | `liveWeb` | unverified |
168
+ | `runsLocally` | no |
169
+ | `fanOut` | no |
170
+ | `contextWindow` | unverified |
171
+ | `loadsProjectRules` | no |
172
+ | `agentDefinitions` | unverified |
173
+
79
174
  ### `qwen` · Qwen Code CLI (Alibaba, provider-agnostic)
80
175
 
81
- - **Kind:** agent-cli · **Access:** metered · **Lane:** B · **Level:** 2+
82
- - **Wins at:** cheapest metered bulk lane for structured output; never for anything that cites a line, a number or a source
176
+ - **Kind:** agent-cli · **Billing:** pay-per-token · **Level:** 2+
177
+ - **What it is:** A provider-agnostic terminal agent; you supply the API key, so its rate is your provider's rate
83
178
  - **Install:** `npm install -g @qwen-code/qwen-code@0.22.3`
84
- - **Sign in:** a provider key in an environment variable, named (not stored) in ~/.qwen/settings.json. There is no free Qwen cloud tier any more.
179
+ - **Sign in:** run `qwen` and use `/auth` to configure your provider
85
180
  - **Reads rules from:** `QWEN.md`
86
181
  - **cli-run lane:** yes
87
182
  - **Built against:** 0.22.3 (the same number the npm pin uses)
88
183
  - **Note:** Its own success flags lie on API failures. cli-run checks the two honest signals for you.
89
184
 
185
+ **Capability facts.** An unverified value needs a current capability check before use. Role assignments come from the selected stack, using these facts.
186
+
187
+ | Fact | Value |
188
+ |---|---|
189
+ | `modelFamily` | unverified |
190
+ | `kind` | agent-cli |
191
+ | `billing` | pay-per-token |
192
+ | `pricing` | unverified |
193
+ | `headless` | yes |
194
+ | `cliRun` | yes |
195
+ | `writesFiles` | yes |
196
+ | `readOnlyMode` | no |
197
+ | `liveWeb` | unverified |
198
+ | `runsLocally` | no |
199
+ | `fanOut` | no |
200
+ | `contextWindow` | unverified |
201
+ | `loadsProjectRules` | no |
202
+ | `agentDefinitions` | unverified |
203
+
90
204
  ### `ollama` · Ollama (local models)
91
205
 
92
- - **Kind:** local · **Access:** local · **Lane:** local · **Level:** 2+
93
- - **Wins at:** the privacy lane: anything that must never leave the machine. Not a cost lane.
206
+ - **Kind:** local-runtime · **Billing:** local · **Level:** 2+
207
+ - **What it is:** A local model runtime; work sent here stays on the machine
94
208
  - **Install:** https://ollama.com/download (or `brew install ollama`)
95
209
  - **Sign in:** none
96
210
  - **Built against:** 0.33.3
97
211
 
212
+ **Capability facts.** An unverified value needs a current capability check before use. Role assignments come from the selected stack, using these facts.
213
+
214
+ | Fact | Value |
215
+ |---|---|
216
+ | `modelFamily` | unverified |
217
+ | `kind` | local-runtime |
218
+ | `billing` | local |
219
+ | `pricing` | unverified |
220
+ | `headless` | yes |
221
+ | `cliRun` | no |
222
+ | `writesFiles` | no |
223
+ | `readOnlyMode` | no |
224
+ | `liveWeb` | no |
225
+ | `runsLocally` | yes |
226
+ | `fanOut` | no |
227
+ | `contextWindow` | unverified |
228
+ | `loadsProjectRules` | no |
229
+ | `agentDefinitions` | unverified |
230
+
98
231
  ### `claude-app` · Claude app or claude.ai (chat only, no CLI)
99
232
 
100
- - **Kind:** chat · **Access:** subscription · **Lane:** chat · **Level:** 1+
101
- - **Wins at:** single-agent use through Projects and custom instructions
233
+ - **Kind:** chat · **Billing:** subscription · **Level:** 1+
234
+ - **What it is:** A chat app; it reads pasted instructions, not files
102
235
  - **Install:** https://claude.ai
103
236
  - **Sign in:** sign in
104
237
 
238
+ **Capability facts.** An unverified value needs a current capability check before use. Role assignments come from the selected stack, using these facts.
239
+
240
+ | Fact | Value |
241
+ |---|---|
242
+ | `modelFamily` | Anthropic |
243
+ | `kind` | chat |
244
+ | `billing` | subscription |
245
+ | `pricing` | unverified |
246
+ | `headless` | no |
247
+ | `cliRun` | no |
248
+ | `writesFiles` | no |
249
+ | `readOnlyMode` | no |
250
+ | `liveWeb` | unverified |
251
+ | `runsLocally` | no |
252
+ | `fanOut` | no |
253
+ | `contextWindow` | unverified |
254
+ | `loadsProjectRules` | no |
255
+ | `agentDefinitions` | unverified |
256
+
105
257
  ### `chatgpt-app` · ChatGPT (chat only, no CLI)
106
258
 
107
- - **Kind:** chat · **Access:** subscription · **Lane:** chat · **Level:** 1+
108
- - **Wins at:** single-agent use through custom instructions and Projects
259
+ - **Kind:** chat · **Billing:** subscription · **Level:** 1+
260
+ - **What it is:** A chat app; it reads pasted instructions, not files
109
261
  - **Install:** https://chatgpt.com
110
262
  - **Sign in:** sign in
111
263
 
264
+ **Capability facts.** An unverified value needs a current capability check before use. Role assignments come from the selected stack, using these facts.
265
+
266
+ | Fact | Value |
267
+ |---|---|
268
+ | `modelFamily` | OpenAI |
269
+ | `kind` | chat |
270
+ | `billing` | subscription |
271
+ | `pricing` | unverified |
272
+ | `headless` | no |
273
+ | `cliRun` | no |
274
+ | `writesFiles` | no |
275
+ | `readOnlyMode` | no |
276
+ | `liveWeb` | unverified |
277
+ | `runsLocally` | no |
278
+ | `fanOut` | no |
279
+ | `contextWindow` | unverified |
280
+ | `loadsProjectRules` | no |
281
+ | `agentDefinitions` | unverified |
282
+
112
283
  ### `gemini-app` · Gemini app (chat only, no CLI)
113
284
 
114
- - **Kind:** chat · **Access:** subscription · **Lane:** chat · **Level:** 1+
115
- - **Wins at:** single-agent use through Gems and saved instructions
285
+ - **Kind:** chat · **Billing:** subscription · **Level:** 1+
286
+ - **What it is:** A chat app; it reads pasted instructions, not files
116
287
  - **Install:** https://gemini.google.com
117
288
  - **Sign in:** sign in
118
289
 
290
+ **Capability facts.** An unverified value needs a current capability check before use. Role assignments come from the selected stack, using these facts.
291
+
292
+ | Fact | Value |
293
+ |---|---|
294
+ | `modelFamily` | Google |
295
+ | `kind` | chat |
296
+ | `billing` | subscription |
297
+ | `pricing` | unverified |
298
+ | `headless` | no |
299
+ | `cliRun` | no |
300
+ | `writesFiles` | no |
301
+ | `readOnlyMode` | no |
302
+ | `liveWeb` | unverified |
303
+ | `runsLocally` | no |
304
+ | `fanOut` | no |
305
+ | `contextWindow` | unverified |
306
+ | `loadsProjectRules` | no |
307
+ | `agentDefinitions` | unverified |
308
+
119
309
  ## Companion tools
120
310
 
311
+ Model-orchestrator project activation merges supported MCP entries for selected companions when activation is enabled. The vendor command auto-registration field below describes what the listed vendor command does when you run it yourself. Global client configuration remains manual.
312
+
121
313
  ### `codecalc` · codecalc (calculator, code runner, logic checker for your agent)
122
314
 
123
315
  - **Repo:** https://github.com/The-40-Thieves/codecalc
124
316
  - **Gives:** exact arithmetic, code execution in 31 languages, SMT logic checks, complexity and equivalence proofs; offline, no key, no telemetry
125
- - **Install:** `uvx 'codecalc[full]' setup --write` (needs uv (https://docs.astral.sh/uv/) and Python 3.10+)
126
- - **Registers itself with:** Claude Code, Claude Desktop, Cursor, VS Code, Zed; snippets for the rest are written to `mcp/`
127
- - **Default:** selected
317
+ - **Install:** `uvx 'codecalc[full]==0.5.0' setup --write` (needs uv (https://docs.astral.sh/uv/) and Python 3.10+)
318
+ - **Vendor command auto-registration:** Claude Code, Claude Desktop, Cursor, VS Code, Zed; snippets for the rest are written to `mcp/`
319
+ - **Default:** not selected
128
320
 
129
321
  ### `obsidian-tc` · obsidian-tc (governed memory: an agent-ready MCP server over an Obsidian vault)
130
322
 
131
323
  - **Repo:** https://github.com/The-40-Thieves/obsidian-tc
132
324
  - **Gives:** durable memory and record for your agents: hybrid retrieval (BM25 + dense + link graph), backlinks, compare-and-swap writes with a confirmation gate, folder ACLs, a poison scan on inferred writes; 163 tools, local by default
133
- - **Install:** `npm install -g obsidian-tc && obsidian-tc /path/to/your/vault` (needs an Obsidian vault folder (the Obsidian app itself is only needed for live plugin bridges); Node 24+ or Bun 1.1+ (stricter than this installer); Ollama with `nomic-embed-text` for local embeddings, or a cloud embeddings key; the Local REST API plugin only for bridge tools)
134
- - **Registers itself with:** Cursor, VS Code; snippets for the rest are written to `mcp/`
325
+ - **Install:** `npm install -g obsidian-tc@1.26.0 && obsidian-tc /path/to/your/vault` (needs an Obsidian vault folder (the Obsidian app itself is only needed for live plugin bridges); Node 24+ or Bun 1.1+ (stricter than this installer); Ollama with `nomic-embed-text` for local embeddings, or a cloud embeddings key; the Local REST API plugin only for bridge tools)
326
+ - **Vendor command auto-registration:** Cursor, VS Code; snippets for the rest are written to `mcp/`
135
327
  - **Default:** not selected
136
328
 
137
329
  ### `context7` · Context7 (Upstash: version-aware docs for the libraries your agent calls)
138
330
 
139
331
  - **Repo:** https://github.com/upstash/context7
140
332
  - **Gives:** up-to-date, version-specific documentation and code examples for libraries, SDKs, APIs and CLIs, pulled into the prompt; tells the agent what the code is SUPPOSED to do. Paired with codecalc, which runs the code and proves what it actually does: docs never stand as proof, and where they disagree the run wins
141
- - **Install:** `npx ctx7 setup` (needs Node.js 18+ for the local server or the ctx7 CLI; a free CONTEXT7_API_KEY is optional, for higher rate limits (it works anonymously at the base rate))
142
- - **Registers itself with:** Claude Code, Cursor, Codex CLI, Qwen Code; snippets for the rest are written to `mcp/`
333
+ - **Install:** `npx -y @upstash/context7-mcp@4.1.1` (needs Node.js 18+ for the local server or the ctx7 CLI; a free CONTEXT7_API_KEY is optional, for higher rate limits (it works anonymously at the base rate))
334
+ - **Vendor command auto-registration:** none; use project activation or merge the supplied snippets in `mcp/` using the tool guide
143
335
  - **Default:** not selected
144
336
 
@@ -1,17 +1,35 @@
1
- # Companion tools
1
+ # Works well with
2
2
 
3
- All optional. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc,context7` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
3
+ Companions add calculation, searchable notes or current library documentation. They are other authors' projects, with their own releases and support channels. All are opt-in; a default install and `--yes` both select none.
4
4
 
5
- ## Companion tools (all optional)
5
+ ```bash
6
+ npx model-orchestrator --tools codecalc,obsidian-tc,context7
7
+ ```
6
8
 
7
- An orchestrator routes work. It does not make a model stop guessing numbers, it does not give it a memory, and it does not make it check a library's current docs before writing against it. Three tools close those gaps: codecalc and obsidian-tc are from the same maintainer, Context7 is from Upstash. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc,context7` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
9
+ Selecting a companion writes model-orchestrator's setup guide and configuration snippets. With automatic activation enabled, a Claude Code main agent also gets the selected servers merged into the project's `.mcp.json`. The summary names this file before the single confirmation; an existing file gets a backup, and existing server entries stay intact. Other hosts keep a manual registration step when the catalog has no verified project-scoped config. `--no-apply` leaves registration manual; `--yes` applies it only with `--apply-snippets`.
8
10
 
9
- | Tool | Closes | Default | You need first |
11
+ The installer writes its managed files and the project configuration changes named in the summary, prints missing tools together under **Install these yourself**, and leaves third-party commands for you to run. It never edits global user config. codecalc still needs uv and Python 3.10+; obsidian-tc still needs your vault configuration path. Context7's hosted connection needs no local package install. `--no-tools` remains accepted for existing scripts.
12
+
13
+ ## Choose a companion
14
+
15
+ | Project | Maintainer | What it adds | Setup and support |
10
16
  |---|---|---|---|
11
- | [codecalc](https://github.com/The-40-Thieves/codecalc) | guessed numbers, comparisons, complexity and equivalence claims: exact arithmetic, code execution in 31 languages, SMT logic checks, `verify_translation` / `verify_optimization`; offline, no key | yes | Python 3.10+ and `uv`. `uvx 'codecalc[full]' setup --write` registers it with Claude Code, Claude Desktop, Cursor, VS Code, Zed; snippets for Codex, Antigravity, Qwen Code are written for you |
12
- | [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes with a confirmation gate, folder ACLs, a poison scan on inferred writes; 163 tools, local by default; AGPL-3.0 | no | an Obsidian vault folder; Node 24+ or Bun 1.1+ (stricter than this installer); Ollama with `nomic-embed-text` or a cloud embeddings key; the Obsidian app and its Local REST API plugin only for live bridge tools. Skip it if you do not keep notes in Obsidian |
13
- | [Context7](https://github.com/upstash/context7) | stale library recall: current, version-specific docs and code examples pulled into the prompt for any library, SDK, API or CLI; hosted, or `npx` locally; MIT | no | nothing to install for the hosted endpoint; Node.js 18+ for the local alternative; a free API key is optional, for a higher rate limit. Always makes a network call, unlike the other two: skip it offline |
17
+ | [codecalc](https://github.com/The-40-Thieves/codecalc) | The-40-Thieves | Local arithmetic, code execution, logic and equivalence checks | Follow the generated `CODECALC.md` and the [upstream instructions](https://github.com/The-40-Thieves/codecalc#readme); report tool issues [upstream](https://github.com/The-40-Thieves/codecalc/issues) |
18
+ | [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | The-40-Thieves | Search, backlinks and controlled writes for an Obsidian notes folder | Follow `OBSIDIAN-TC.md` and the [upstream instructions](https://github.com/The-40-Thieves/obsidian-tc#readme); report tool issues [upstream](https://github.com/The-40-Thieves/obsidian-tc/issues) |
19
+ | [Context7](https://github.com/upstash/context7) | Upstash | Current documentation and examples for a specific library version | Follow `CONTEXT7.md` and the [upstream instructions](https://github.com/upstash/context7#readme); report tool issues [upstream](https://github.com/upstash/context7/issues) |
20
+
21
+ Check each project's current runtime and sign-in requirements before installing it. Context7 sends documentation queries over the network, including when its MCP server runs locally. Keep private source and secrets out of query text.
22
+
23
+ ## Use the protocols with your existing tools
24
+
25
+ | Need | With a companion | With your existing tools |
26
+ |---|---|---|
27
+ | Compute and verify | codecalc | A calculator, a local runtime or your project's test command |
28
+ | Search and record | obsidian-tc | A notes folder, file search and version control |
29
+ | Check an API | Context7 | Current official documentation and upstream source, then a local execution check |
30
+
31
+ Every install includes `protocols/numbers-and-logic.md`, `protocols/memory-and-record.md` and `protocols/docs-then-prove.md`. The protocols describe the job; companion selection changes which tool can do it.
14
32
 
15
- Context7 pairs with codecalc rather than duplicating it: Context7 tells the agent what a library is documented to do on this version, codecalc runs the code and proves what it actually does. Docs never stand as proof on their own, and where the two disagree the run wins.
33
+ ## Keep an existing setup
16
34
 
17
- Whether or not you select them, every level carries the three rules they serve: `protocols/numbers-and-logic.md` (when calling a calculator is mandatory, how to report a computed figure, why a thought log is not evidence), `protocols/memory-and-record.md` (search before writing, the folder index is part of the change, one writer, inferred content marked as inferred), and `protocols/docs-then-prove.md` (current docs before writing a call, then a run proves it, the run wins on disagreement).
35
+ Upgrading from 0.1.x preserves existing codecalc files. `--uninstall` removes a managed companion guide or snippet only when its content still matches the recorded hash; edited files stay and are listed. Automatically merged project MCP entries are recorded separately: uninstall removes only entries the installer added that still match their recorded configuration. Other servers, user edits and backups stay. See [Upgrading from 0.1.x](install.md#upgrading-from-01x).
@@ -1,18 +1,27 @@
1
- # What holds, and what is only asked for
1
+ # Guarantees and verification
2
2
 
3
- Read this before running any of it unattended.
3
+ model-orchestrator writes its managed files and the project activation changes shown in the summary. It runs no third-party installs, edits no global user config and sends no telemetry. `npx` itself downloads this package through npm before the installer starts.
4
4
 
5
- Most of what this package ships is text an agent is asked to follow. Be clear about which is which before relying on it unattended.
5
+ Use the distinction below when choosing which parts to run unattended.
6
6
 
7
- | Property | How it holds |
7
+ | Property | What enforces it |
8
8
  |---|---|
9
- | Installer writes only inside `--dir` and `--project`, never a secret, never over a document without `--force` (or `--update-docs`, which touches only documents provably untouched since a previous run); machine-owned config always, runtime files only when provably untouched or with `--upgrade-runtime` | **enforced by code** (preflight, exclusive create, rollback, manifest hashes; tested) |
10
- | `cli-run` exit codes, process-group kill on timeout and on SIGINT/SIGTERM, UTF-8-safe streaming, fixed-code durable log, `--expect-*` contracts with a pre-run snapshot | **enforced by code** (tested with stub lanes) |
11
- | Codex audit lane runs read-only | **delegated to the vendor flag** (`--audit` → `--sandbox read-only`); commands and network still follow your codex config |
12
- | Other lanes' permissions, sign-in state, model versions | **delegated to each vendor's own config**; `--doctor` checks presence, not versions |
13
- | Gateway binds to loopback, keys by name only | **enforced in the generated files**; whether the gateway authenticates is your environment |
14
- | Lane selection, tiers, privacy classes, one-writer, escalation, the protocols | **agent instructions**. Nothing here stops an agent that ignores its rules; the task bundle and the protocols make ignoring them visible, not impossible |
15
- | Weekly audit bounded, previous report preserved | **enforced in the generated script and unit** (watchdog, temp-and-rename, `TimeoutStartSec`) |
9
+ | Installer writes stay inside `--dir` and `--project`; existing edits are preserved by default | Code: path preflight, exclusive creation, rollback and manifest hashes |
10
+ | Companion tools start unselected, including with `--yes` | Code: explicit tool selection; missing tools get printed setup commands |
11
+ | `--update-docs` refreshes unedited documents; `--upgrade-runtime` replaces runtime files only | Code: recorded hashes, file classes and named conflicts |
12
+ | Interactive activation applies the main agent's supported rules and settings with backups; `--yes` requires `--apply-snippets` | Code: one confirmation, the `--no-apply` flag and edit toggle, marked rules blocks and validated settings merges |
13
+ | Selected companions register automatically only in the main agent's catalog-supported project config | Code: Claude Code's project `.mcp.json` target; other hosts retain manual setup steps |
14
+ | Uninstall removes unchanged managed files and unchanged recorded activation entries | Code: complete manifest validation before removal, containment checks, hash comparison and activation ownership records |
15
+ | Every non-dry install checks CLI presence; live prompts stay opt-in | Code: automatic `--doctor`; `--doctor --run` explicitly enables vendor canaries |
16
+ | `cli-run` returns nonzero for missing results, timeouts and rejected output contracts | Code: vendor-specific output checks, process cleanup and `--expect-*` checks against a pre-run snapshot |
17
+ | Acceptance checks report PASS or FAIL and return exit 1 for any failure | Code: `aunx checks run ACCEPTANCE_CHECKS.json`; run it before the action you want to gate |
18
+ | Proof entries carry measurement dates, methods and expiry dates | Code: data validation and an expiry check in `npm test` |
19
+ | Codex audits request read-only filesystem access | Vendor flag: `--audit` selects `--sandbox read-only`; command and network permissions still follow the vendor configuration |
20
+ | Other workers' permissions, sign-in state and model availability | Vendor configuration and your live verification with `aunx cli-run --doctor --run` |
21
+ | Gateway listens on loopback and refers to secrets by environment-variable name | Generated configuration; authentication and deployment remain your responsibility |
22
+ | Model choice, effort guidance, privacy rules, one writer and independent review | Agent instructions in the routing rules, task brief and protocols |
23
+ | Weekly review has bounded execution and preserves the previous report on failure | Generated script and service: watchdog, temporary output and rename |
16
24
 
17
- If you need a property in the third row to be enforced, that is a router, a policy engine or a sandbox, and this package does not claim to be one.
25
+ The acceptance-check runner executes commands you supply, with your shell's permissions. Review a check file before running it. Wire its exit code into your own release command or CI when you want it to block that action.
18
26
 
27
+ The plugin's routing hooks only read files and emit context; they run no subprocess, perform no network access and write no files. The separate installed metrics hook writes a local routing log without prompt text. [Security review history](security-review-history.md) links the regression evidence for these boundaries.