model-orchestrator 0.1.22 → 0.1.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,30 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.24] - 2026-09-18
8
+
9
+ ### Added
10
+
11
+ - **Context7 as a third companion tool, paired with codecalc.** [Context7](https://github.com/upstash/context7) (Upstash) hands the agent current, version-specific documentation and code examples for any library, SDK, API or CLI, hosted or run locally with `npx`. It answers what a library is documented to do; codecalc still answers what the code actually does by running it. A new protocol, `protocols/docs-then-prove.md`, states the rule the pairing serves: pull current docs before writing against anything unconfirmed this session, then a run proves it, and where a doc and a run disagree the run wins. Optional and off by default, like obsidian-tc (`--tools context7`, or the interactive question); unlike the other two companions it always makes a network call, so it is the one to skip offline. `src/catalog.js` carries the entry (repo, role, install, requirements, the clients it self-registers with); `CONTEXT7_STATUS` renders both selected and not-selected wording the same way `CODECALC_STATUS` and `OBSIDIAN_TC_STATUS` already do, at every level and in `ROUTING.md`.
12
+ - **`templates/tools/context7/CONTEXT7.md` plus seven per-client `mcp/` snippets**, each read from that client's own docs rather than shared across clients that do not actually share a config shape: `context7.claude-code.mcp.json` (`"type": "http"` next to `url`, required or Claude Code skips the server), `context7.mcpServers.json` (Cursor's own documented shape, `url` with no `type`), `context7.vscode.mcp.json`, `context7.qwen.settings.json` (`httpUrl`, not `url`, plus the non-credential `Accept` header upstream ships), `context7.zed.settings.json` (local `npx`, pinned to the catalog's `context7` version), `context7.codex.config.toml`, `context7.agy.mcp_config.json`. Claude Desktop has no file: its remote connection is a UI step (`Settings > Connectors > Add Custom Connector`), documented in `CONTEXT7.md` instead of a snippet nothing there reads.
13
+ - **Every context7 snippet ships keyless by default.** The anonymous tier works with no header, while an unexpanded or empty `Bearer ${CONTEXT7_API_KEY}` makes every call return "Invalid API key" (probed live against `https://mcp.context7.com/mcp`), and clients like Codex `http_headers` never expand it. `CONTEXT7.md` has a "Higher rate limits (optional key)" section with one mechanism per client: Codex `bearer_token_env_var`, Claude Code `${VAR}` expansion in `.mcp.json` headers, Claude Desktop's own Connectors key field, an exported shell variable for local `npx`, and "check your client's docs" where expansion is not confirmed; the fallback for a client that does not pass its environment to a spawned child is to stay anonymous, never to paste the key into a snippet's args. A test fails if a shipped snippet carries `Authorization` or `CONTEXT7_API_KEY`.
14
+
15
+ ## [0.1.23] - 2026-09-15
16
+
17
+ ### Added
18
+
19
+ - **`cli-run` says WHY a lane failed.** Every run lands in one class with its own exit code: `auth` 14, `quota` 15, `rejected` 16, `refused` 17, `cut_short` 18, next to the existing `empty` 10, `no_output` 11, `timeout` 12 and `unavailable` 13. A missing API key, a spent quota and an unknown model id used to share exit 10 or the vendor's own code, and each needs a different response: a missing key is not a model fault, and retrying a spent quota cannot help. Signals are read from each lane's authoritative error fields only, never the model's prose, with precedence auth, quota, rejected, refused, cut_short, empty.
20
+ - **`refused=N` on every run.** Tool calls a hook or deny rule blocked, counted from qwen's `permission_denials`, agy's deny-rule steps, codex's router `Rejected(` lines, and grok's session transcript (its `sessionId` is charset-checked, the path is contained to the sessions root, lines must name the same session, and the read is capped at 5 MiB). `null` when a lane gives no signal. A deliverable with refused calls is still exit 0.
21
+ - **A problem line and a fix line** on the terminal for every failure, and for an `ok` run with refused calls, so a calling agent can relay "this is what went wrong, this is the fix" and ask.
22
+ - **Terminal output is redacted** before it prints: JSON credential keys, `Authorization:` values, bearer values, URL query credentials, and common key prefixes, redacted before any clipping.
23
+ - **The durable log gains `class` and `refused`.** Both are fixed values; the log still never holds provider text, the problem line or stderr.
24
+
25
+ ### Changed
26
+
27
+ - **A nonzero vendor exit is no longer passed through as cli-run's exit code.** The class owns the code, and the vendor's own code stays in the log as `cli_rc`. A nonzero exit nothing else explains is `cut_short` (18), and it is still never `ok`. If a script compared `cli-run`'s exit code with a specific vendor code, compare `cli_rc` in the log instead; `!= 0` checks are unaffected.
28
+ - **A lane killed by a signal, or output past the 16 MiB buffer, is `cut_short` (18)**, not 10. Exit 10 now means only an empty run or an unmet `--expect-*` contract.
29
+ - **agy's judge refuses a non-object terminal `result` as `bad_last_event`** (was `bad_status`), and **qwen's judge refuses a non-string `error.message` as `error_message_not_string`**, so both classify as `cut_short`. hermes' stderr cause (degraded free tier, bad `--toolsets`) is now named in its detail line.
30
+
7
31
  ## [0.1.22] - 2026-09-12
8
32
 
9
33
  ### Added
@@ -335,7 +359,9 @@ First release.
335
359
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
336
360
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
337
361
 
338
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.22...HEAD
362
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.24...HEAD
363
+ [0.1.24]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.23...v0.1.24
364
+ [0.1.23]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.22...v0.1.23
339
365
  [0.1.22]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.21...v0.1.22
340
366
  [0.1.21]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.20...v0.1.21
341
367
  [0.1.20]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.19...v0.1.20
package/README.md CHANGED
@@ -38,7 +38,7 @@ State known subscription plans with `--plans codex=pro-20x,agy=ultra-5x`. The ge
38
38
 
39
39
  | Level | You have | You get |
40
40
  |---|---|---|
41
- | **1 · Beginner** | one LLM or one agent | tiers, task classification, the two build checkpoints, the protocols (build, propagate, gap analysis, deep research, numbers and logic, memory and record), a task-bundle template, and your agent set up to follow them |
41
+ | **1 · Beginner** | one LLM or one agent | tiers, task classification, the two build checkpoints, the protocols (build, propagate, gap analysis, deep research, numbers and logic, memory and record, docs and proof), a task-bundle template, and your agent set up to follow them |
42
42
  | **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts; `--model` / `--effort` to pin the route and log it), a delegation matrix generated from your selection, research triage across the lanes you have |
43
43
  | **3 · Advanced** | a virtual machine | everything above, plus a gateway config rendered from the API keys you hold (asked separately from your CLIs), pinned images, box rules, privacy gates, and a weekly gap-analysis job with "what watches it" written down |
44
44
 
@@ -59,16 +59,19 @@ Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/pa
59
59
 
60
60
  `npx model-orchestrator --list` prints the catalog with install and sign-in notes. Details: [docs/catalog.md](docs/catalog.md).
61
61
 
62
- ## Companion tools (both optional)
62
+ ## Companion tools (all optional)
63
63
 
64
- An orchestrator routes work. It does not make a model stop guessing numbers, and it does not give it a memory. Two tools from the same maintainer close those gaps. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
64
+ An orchestrator routes work. It does not make a model stop guessing numbers, it does not give it a memory, and it does not make it check a library's current docs before writing against it. Three tools close those gaps: codecalc and obsidian-tc are from the same maintainer, Context7 is from Upstash. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc,context7` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
65
65
 
66
66
  | Tool | Closes | Default | You need first |
67
67
  |---|---|---|---|
68
68
  | [codecalc](https://github.com/The-40-Thieves/codecalc) | guessed numbers, comparisons, complexity and equivalence claims: exact arithmetic, code execution in 31 languages, SMT logic checks, `verify_translation` / `verify_optimization`; offline, no key | yes | Python 3.10+ and `uv`. `uvx 'codecalc[full]' setup --write` registers it with Claude Code, Claude Desktop, Cursor, VS Code, Zed; snippets for Codex, Antigravity, Qwen Code are written for you |
69
69
  | [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes with a confirmation gate, folder ACLs, a poison scan on inferred writes; 163 tools, local by default; AGPL-3.0 | no | an Obsidian vault folder; Node 24+ or Bun 1.1+ (stricter than this installer); Ollama with `nomic-embed-text` or a cloud embeddings key; the Obsidian app and its Local REST API plugin only for live bridge tools. Skip it if you do not keep notes in Obsidian |
70
+ | [Context7](https://github.com/upstash/context7) | stale library recall: current, version-specific docs and code examples pulled into the prompt for any library, SDK, API or CLI; hosted, or `npx` locally; MIT | no | nothing to install for the hosted endpoint; Node.js 18+ for the local alternative; a free API key is optional, for a higher rate limit. Always makes a network call, unlike the other two: skip it offline |
70
71
 
71
- Whether or not you select them, every level carries the two rules they serve: `protocols/numbers-and-logic.md` (when calling a calculator is mandatory, how to report a computed figure, why a thought log is not evidence) and `protocols/memory-and-record.md` (search before writing, the folder index is part of the change, one writer, inferred content marked as inferred).
72
+ Context7 pairs with codecalc rather than duplicating it: Context7 tells the agent what a library is documented to do on this version, codecalc runs the code and proves what it actually does. Docs never stand as proof on their own, and where the two disagree the run wins.
73
+
74
+ Whether or not you select them, every level carries the three rules they serve: `protocols/numbers-and-logic.md` (when calling a calculator is mandatory, how to report a computed figure, why a thought log is not evidence), `protocols/memory-and-record.md` (search before writing, the folder index is part of the change, one writer, inferred content marked as inferred), and `protocols/docs-then-prove.md` (current docs before writing a call, then a run proves it, the run wins on disagreement).
72
75
 
73
76
  ## The two folders every run writes to
74
77
 
@@ -88,7 +91,7 @@ An install has two targets, and a scripted run should set both.
88
91
  npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code \
89
92
  --dir ./ai-orchestrator --project ./my-app
90
93
  # --yes selects the recommended companion tool (codecalc), which writes CODECALC.md and mcp/ snippets.
91
- # Add --no-tools for none, or --tools codecalc,obsidian-tc to choose.
94
+ # Add --no-tools for none, or --tools codecalc,obsidian-tc,context7 to choose.
92
95
 
93
96
  npx model-orchestrator --yes --level 3 --ais claude-code,codex,agy,grok,hermes,qwen,ollama --apis anthropic,openrouter --dry # print the plan, write nothing
94
97
  npx model-orchestrator --yes --level 2 --ais claude-code,codex --project ~/my-app --dir ~/my-app/ai-orchestrator --no-tools # subagents into ~/my-app/.claude/agents
@@ -118,8 +121,8 @@ ai-orchestrator/
118
121
  README.md start here, written for your level and your AIs
119
122
  ORCHESTRATOR.md single-agent routing rules (level 1)
120
123
  TASK_BUNDLE.md the brief every delegation carries
121
- protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record
122
- CODECALC.md OBSIDIAN-TC.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
124
+ protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record · docs-then-prove
125
+ CODECALC.md OBSIDIAN-TC.md CONTEXT7.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
123
126
  <project>/.claude/agents/ one per tier plus finding-verifier, done-verifier, reader, at the PROJECT root (if Claude Code is primary)
124
127
  <project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart) + route-metrics.mjs (all five: see "Measuring routing" below), Claude Code only
125
128
  CLAUDE.snippet.md the block to paste into your CLAUDE.md