model-orchestrator 0.1.22 → 0.1.24
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +27 -1
- package/README.md +10 -7
- package/bin/cli-run.mjs +488 -48
- package/docs/audit-brief.md +20 -0
- package/docs/catalog.md +9 -1
- package/docs/part-1-beginner.md +5 -1
- package/docs/part-2-intermediate.md +5 -1
- package/docs/part-3-advanced.md +4 -0
- package/package.json +1 -1
- package/src/catalog.js +12 -0
- package/src/install.js +2 -0
- package/templates/README.md +2 -2
- package/templates/advanced/vm/README.md +1 -0
- package/templates/advanced/vm/jobs/weekly-audit.sh +1 -1
- package/templates/beginner/ORCHESTRATOR.md +6 -2
- package/templates/common/README.md +2 -0
- package/templates/common/protocols/README.md +2 -1
- package/templates/common/protocols/docs-then-prove.md +26 -0
- package/templates/intermediate/CLI-RUN.md +37 -13
- package/templates/intermediate/ROUTING.md +4 -0
- package/templates/tools/README.md +2 -1
- package/templates/tools/context7/CONTEXT7.md +82 -0
- package/templates/tools/context7/mcp/context7.agy.mcp_config.json +7 -0
- package/templates/tools/context7/mcp/context7.claude-code.mcp.json +8 -0
- package/templates/tools/context7/mcp/context7.codex.config.toml +3 -0
- package/templates/tools/context7/mcp/context7.mcpServers.json +7 -0
- package/templates/tools/context7/mcp/context7.qwen.settings.json +8 -0
- package/templates/tools/context7/mcp/context7.vscode.mcp.json +8 -0
- package/templates/tools/context7/mcp/context7.zed.settings.json +9 -0
package/docs/audit-brief.md
CHANGED
|
@@ -126,3 +126,23 @@ Three findings reproduced against the 0.1.15 branch before it shipped, none of t
|
|
|
126
126
|
- **Not yet attacked for real.** `windowsSpawnPlan`, `resolveCmdShim` and the escaping functions are unit-tested (pure string logic, runs on every CI host) and the resolved-shim path is exercised end to end on `windows-latest` through the fake-lane fixtures in `test/cli.test.js` (installed as a real npm-style `.cmd` shim). A lane can no longer reach `cmd.exe` at all (a unit test pins the refusal). The opt-in `cmd.exe` branch, used only by the installer's own `npm install -g`, is not exercised end to end through a live Windows process in this suite; its escaping is proven by exact-string unit tests only, and its arguments never include user text.
|
|
127
127
|
- **A platform limit found the same way, unrelated to the spawn path itself: Windows has no OS-level signals at all.** `cli-run.mjs`'s graceful shutdown (`process.on('SIGTERM', ...)`, kill the lane's process group, then exit 143/130) is a POSIX guarantee only: `ChildProcess.kill(sig)` on Windows calls `TerminateProcess()` unconditionally for SIGTERM AND SIGINT alike, giving the target process no chance to run any handler at all, proven on `windows-latest` CI (the wrapper died as `{code: null, signal: sig}` for both; a hypothesis that SIGINT gets a real, catchable console-control event on Windows was tried first and measured false in this exact scenario, not assumed). `test/cli.test.js`'s `#13` now expects an unhandled termination for either signal on win32, and the original graceful-exit assertion elsewhere.
|
|
128
128
|
- **Two narrow, individually-verified Windows skips remain, neither in the spawn path itself.** (1) A lane dying mid-run from a real POSIX signal cannot be reproduced on win32: a real Windows lane is a plain `node <script>` process, so it cannot die "by signal" any more than the product being tested can, and the only way a test fixture can even simulate one (a nested `sh -c "...; kill -TERM $$"`) puts an extra node process between cli-run.mjs and the dying shell, so cli-run.mjs observes only that node's translated exit code (measured: MSYS bash's self-kill status leaks through as a plain nonzero exit code, 3840, which this tool already handles honestly via `exit_nonzero`). (2) `weekly-audit.sh`'s watchdog (`bounded()`/`killtree()`, `pgrep -P` plus killing a backgrounded subshell's tree) relies on real bash job control this script only ever runs under on the Ubuntu box it targets; actually executing it against a genuinely hanging stub under Git Bash's job-control emulation hung past a 20s outer timeout on `windows-latest` CI, a known class of MSYS/Cygwin limitation (a `kill -KILL` not reliably reaching the underlying Windows process tree of a backgrounded subshell), not a defect in the generated script, which still renders and syntax-checks correctly.
|
|
129
|
+
|
|
130
|
+
## New in 0.1.23: failure classes in cli-run
|
|
131
|
+
|
|
132
|
+
`bin/cli-run.mjs` now puts every run in a closed failure class that owns the exit code (`auth` 14, `quota` 15, `rejected` 16, `refused` 17, `cut_short` 18, beside `ok` 0, `empty` 10, `no_output` 11, `timeout` 12, `unavailable` 13), counts refused tool calls per lane (including a read of grok's session transcript, whose `sessionId` comes from lane stdout), and prints redacted problem and fix lines. The durable log gains only `class` and `refused`. Threat model: a false exit 0, provider text or a secret reaching the log or the terminal, a transcript read escaping the sessions root, misclassification that sends a user the wrong way, and anything that throws or stalls.
|
|
133
|
+
|
|
134
|
+
### Round 1 (pre-release, GPT-6 Astra at xhigh), fixed before shipping
|
|
135
|
+
|
|
136
|
+
Every finding was reproduced as a failing case in `test/classify.test.js` against the unfixed code (all seven red), then fixed; the suite is green after. One round, by rule; the regression cases verify the fixes.
|
|
137
|
+
|
|
138
|
+
| # | Sev | Finding | Fix | Test |
|
|
139
|
+
|---|---|---|---|---|
|
|
140
|
+
| 1 | HIGH | qwen's display detail clipped an error message at 120 characters before redaction, so a long JSON password lost its closing quote and printed its prefix | redact before every clip: judge detail, API-error text, denial text, deliverable snippet, problem cause | R1 |
|
|
141
|
+
| 2 | HIGH | the JSON credential pattern ended at an escaped quote and missed unterminated values | match JSON string escapes, run an unterminated value to the end, and accept a JSON body escaped inside a string | R2 |
|
|
142
|
+
| 3 | HIGH | `firstDenialText` scanned 16 MiB with an unbounded `[^)]*`: 1 MiB of unclosed `Permission denied for command(` took 15.6 s after the lane had exited | scan at most 256 KiB of stdout and of stderr, with bounded repeats | R3 |
|
|
143
|
+
| 4 | MEDIUM | transcript containment was a string prefix with both separators, so on POSIX a directory literally named `sessions\outside` passed | `isInsideRoot()` via `path.relative`, rejecting `..`, absolute and cross-drive results; tested with `path.posix` and `path.win32` on every OS | R4 |
|
|
144
|
+
| 5 | MEDIUM | qwen signals were searched in the clipped display detail, so a quota error after 120 characters read as `empty`, and a model name in the detail could fake a signal | classify on `qwenErrorText()`: the terminal event's full error and an API-error result only | R5 |
|
|
145
|
+
| 6 | MEDIUM | hermes fell back to stdout when stderr was empty, so an answer saying "no rate limit" read as `quota` | stderr only | R6 |
|
|
146
|
+
| 7 | MEDIUM | a nonzero vendor exit with a recognised empty shape (agy exit 7, `SUCCESS`, empty response) returned `empty`, contradicting the documented `cut_short` fallback | an unexplained nonzero exit is `cut_short`, except hermes' own exit 2 | R7 |
|
|
147
|
+
|
|
148
|
+
Reported clean by the same audit: no false exit 0 across 675 malformed probes and five lanes; no provider text or secret fragment in probed log records; the real codex 0.153.4 fixture's non-fatal error item still succeeds; grok traversal, symlink, FIFO, session-match and read-cap guards; timeout, interrupt, overrun, killed lane, JSON contract, `--quiet` and the doctor canary. Not verified by it: native Windows execution, live vendors, filesystem races.
|
package/docs/catalog.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Catalog
|
|
2
2
|
|
|
3
|
-
Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrites it. Protocols shipped at every level:
|
|
3
|
+
Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrites it. Protocols shipped at every level: 7 (counted from `templates/common/protocols/`).
|
|
4
4
|
|
|
5
5
|
## Levels
|
|
6
6
|
|
|
@@ -134,3 +134,11 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
|
|
|
134
134
|
- **Registers itself with:** Cursor, VS Code; snippets for the rest are written to `mcp/`
|
|
135
135
|
- **Default:** not selected
|
|
136
136
|
|
|
137
|
+
### `context7` · Context7 (Upstash: version-aware docs for the libraries your agent calls)
|
|
138
|
+
|
|
139
|
+
- **Repo:** https://github.com/upstash/context7
|
|
140
|
+
- **Gives:** up-to-date, version-specific documentation and code examples for libraries, SDKs, APIs and CLIs, pulled into the prompt; tells the agent what the code is SUPPOSED to do. Paired with codecalc, which runs the code and proves what it actually does: docs never stand as proof, and where they disagree the run wins
|
|
141
|
+
- **Install:** `npx ctx7 setup` (needs Node.js 18+ for the local server or the ctx7 CLI; a free CONTEXT7_API_KEY is optional, for higher rate limits (it works anonymously at the base rate))
|
|
142
|
+
- **Registers itself with:** Claude Code, Cursor, Codex CLI, Qwen Code; snippets for the rest are written to `mcp/`
|
|
143
|
+
- **Default:** not selected
|
|
144
|
+
|
package/docs/part-1-beginner.md
CHANGED
|
@@ -58,9 +58,13 @@ Logic flow follows the same rule. Reasoning scaffolds help models with no native
|
|
|
58
58
|
|
|
59
59
|
Every protocol ends in a write. Search before you write (duplicates are how a store starts lying), correct the folder index in the same pass, one writer per session, mark inferred content as inferred. The optional companion for that is [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc), a governed MCP server over an Obsidian vault: hybrid search, backlinks, compare-and-swap writes, folder ACLs. It needs an Obsidian vault, Node 24+ or Bun, and Ollama or a cloud embeddings key, so it is off by default; without it the rule still binds against a notes folder and `grep`.
|
|
60
60
|
|
|
61
|
+
## 9. Docs, then prove
|
|
62
|
+
|
|
63
|
+
A model's recall of a library's API is training data, not a live source; it goes stale the moment the vendor ships a release it never saw. Before writing code against a library, SDK, API or CLI you have not confirmed this session, pull current, version-specific docs; then a run, not the doc, is what proves the code behaves that way. The optional companion is [Context7](https://github.com/upstash/context7) (Upstash): it hands the agent current, version-aware documentation and code examples on request, hosted or run locally with `npx`. It pairs with codecalc rather than replacing it: Context7 says what the code is supposed to do, codecalc's run says what it actually does, and the run wins where they disagree. It needs a network call (there is no offline mode), so it is off by default; without it the rule still binds, read the vendor's own docs or source by hand.
|
|
64
|
+
|
|
61
65
|
## What the installer gives you at this level
|
|
62
66
|
|
|
63
|
-
`README.md` (start here) · `ORCHESTRATOR.md` · `TASK_BUNDLE.md` · `protocols/{build-protocol, propagate, gap-analysis, deep-research, numbers-and-logic, memory-and-record}.md` · `CODECALC.md
|
|
67
|
+
`README.md` (start here) · `ORCHESTRATOR.md` · `TASK_BUNDLE.md` · `protocols/{build-protocol, propagate, gap-analysis, deep-research, numbers-and-logic, memory-and-record, docs-then-prove}.md` · `CODECALC.md`, `OBSIDIAN-TC.md` and `CONTEXT7.md` with `mcp/` snippets for the companion tools you selected · the loading surface for your primary agent (Claude Code subagents, Antigravity custom agents, a rules-file snippet, or a paste block for a chat app).
|
|
64
68
|
|
|
65
69
|
## When you have outgrown it
|
|
66
70
|
|
|
@@ -28,7 +28,7 @@ One driver, no second AI in the mix: the orchestrator invokes the CLIs; it never
|
|
|
28
28
|
|
|
29
29
|
## 3. Exit 0 is a lie on every lane
|
|
30
30
|
|
|
31
|
-
Every agent CLI can report success and deliver nothing. `bin/cli-run.mjs` builds the right invocation per lane, reads that lane's **native** terminal event, and exits `10` when a run produced no deliverable, `12` on timeout, `13` when the lane is missing. Byte count is not a check either; a run can emit hundreds of kilobytes and no conclusion. One lane's own success flags lie outright (an upstream 400 reported as success), so its judge reads the two honest signals instead.
|
|
31
|
+
Every agent CLI can report success and deliver nothing. `bin/cli-run.mjs` builds the right invocation per lane, reads that lane's **native** terminal event, and exits `10` when a run produced no deliverable, `12` on timeout, `13` when the lane is missing, and `14` to `18` when the lane's own error says why (auth, quota, rejected, refused, cut short), so a missing API key is never blamed on the model. Byte count is not a check either; a run can emit hundreds of kilobytes and no conclusion. One lane's own success flags lie outright (an upstream 400 reported as success), so its judge reads the two honest signals instead.
|
|
32
32
|
|
|
33
33
|
Every call goes through it. "This lane is flaky" becomes a query over its log instead of an argument. `node bin/cli-run.mjs --doctor` is the first thing to run after install: enabled lanes, binaries on PATH, the route each lane is pinned to, and with `--run` a one-word canary per lane.
|
|
34
34
|
|
|
@@ -66,6 +66,10 @@ Measured on the cheapest metered lane: conclusions right, 0 of 11 line citations
|
|
|
66
66
|
|
|
67
67
|
With several lanes proposing, the store is where they meet. obsidian-tc (optional) gives every CLI the same `semantic_search`, `get_backlinks` and compare-and-swap `write_note`, with folder ACLs so a research lane can read what it needs and write nothing. The orchestrator stays the one writer.
|
|
68
68
|
|
|
69
|
+
## 11. Docs, then prove, across lanes
|
|
70
|
+
|
|
71
|
+
Every lane's recall of a library's API is a lead, the same as its arithmetic (see item 9 above). Context7 (optional) gives every CLI the same current, version-aware docs lookup, registered for Claude Code, Cursor, Codex and Qwen Code with the snippets in `CONTEXT7.md`. It pairs with codecalc: a lane's claim about what a library does, cited from memory or from a doc, is confirmed by a run before code ships on it.
|
|
72
|
+
|
|
69
73
|
## What the installer gives you at this level
|
|
70
74
|
|
|
71
75
|
Everything from Part 1, plus `ROUTING.md` · `TIERS.md` · `DELEGATION_MATRIX.md` (generated from your selection) · `RESEARCH_TRIAGE.md` · `CLI-RUN.md` · `bin/cli-run.mjs` · `bin/lanes.json`.
|
package/docs/part-3-advanced.md
CHANGED
|
@@ -60,6 +60,10 @@ Runs as a stdio MCP server next to the orchestrator CLI: offline, no key, nothin
|
|
|
60
60
|
|
|
61
61
|
Stdio next to the orchestrator, or the upstream Docker service against a bind-mounted vault. Embeddings on the box's Ollama, so nothing leaves the machine. HTTP transport stays off unless every caller is on the private mesh and auth is on.
|
|
62
62
|
|
|
63
|
+
## 10. context7 on the box
|
|
64
|
+
|
|
65
|
+
Unlike the other two companions, it is never fully local: the hosted endpoint is a network call over HTTPS from the box, or a local `npx` server over stdio still needs no cloud account to run anonymously. Either way, only the library name and the query text leave the box, never source code. Scheduled jobs that write code against a vendored dependency pull its current docs through Context7 first, then prove the shape with codecalc before it ships.
|
|
66
|
+
|
|
63
67
|
## What the installer gives you at this level
|
|
64
68
|
|
|
65
69
|
Everything from Parts 1 and 2, plus `vm/README.md` · `vm/setup-vm.sh` · `vm/docker-compose.yml` · `vm/gateway.config.yaml` (one lane per provider you selected, keys by name only) · `vm/ENVIRONMENT.md` · `vm/box-CLAUDE.md` · `vm/PRIVACY_GATES.md` · `vm/jobs/` (a weekly audit timer + service, and an index that names what watches each job).
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "model-orchestrator",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.24",
|
|
4
4
|
"description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
package/src/catalog.js
CHANGED
|
@@ -287,6 +287,18 @@ export const TOOLS = [
|
|
|
287
287
|
autoClients: ['Cursor', 'VS Code'],
|
|
288
288
|
recommended: false,
|
|
289
289
|
optionalNote: 'Optional and heavier than codecalc. Skip it if you do not keep notes in Obsidian. AGPL-3.0.'
|
|
290
|
+
},
|
|
291
|
+
{
|
|
292
|
+
id: 'context7',
|
|
293
|
+
name: 'Context7 (Upstash: version-aware docs for the libraries your agent calls)',
|
|
294
|
+
repo: 'https://github.com/upstash/context7',
|
|
295
|
+
role: 'up-to-date, version-specific documentation and code examples for libraries, SDKs, APIs and CLIs, pulled into the prompt; tells the agent what the code is SUPPOSED to do. Paired with codecalc, which runs the code and proves what it actually does: docs never stand as proof, and where they disagree the run wins',
|
|
296
|
+
install: 'npx ctx7 setup',
|
|
297
|
+
pin: '4.1.1',
|
|
298
|
+
requires: 'Node.js 18+ for the local server or the ctx7 CLI; a free CONTEXT7_API_KEY is optional, for higher rate limits (it works anonymously at the base rate)',
|
|
299
|
+
autoClients: ['Claude Code', 'Cursor', 'Codex CLI', 'Qwen Code'],
|
|
300
|
+
recommended: false,
|
|
301
|
+
optionalNote: 'Optional, and from a different maintainer than codecalc and obsidian-tc (Upstash, not The-40-Thieves). Needs a network call even at the anonymous rate; skip it offline. MIT.'
|
|
290
302
|
}
|
|
291
303
|
];
|
|
292
304
|
export const toolById = Object.fromEntries(TOOLS.map((t) => [t.id, t]));
|
package/src/install.js
CHANGED
|
@@ -482,6 +482,7 @@ function vars(opts) {
|
|
|
482
482
|
OLLAMA_IMAGE: IMAGES.ollama,
|
|
483
483
|
CODECALC_PIN: pinOf('codecalc'),
|
|
484
484
|
OBSIDIAN_TC_PIN: pinOf('obsidian-tc'),
|
|
485
|
+
CONTEXT7_PIN: pinOf('context7'),
|
|
485
486
|
APIS_LIST: apis.length ? apis.map((prov) => '- ' + prov.name + ' (`' + prov.envName + '`)').join('\n') : '- none: no metered provider key was selected, so the gateway serves only a local lane if you picked one',
|
|
486
487
|
INSTALL_DIR: dirPosix,
|
|
487
488
|
INSTALL_DIR_SH: shellQuote(dirPosix),
|
|
@@ -507,6 +508,7 @@ function vars(opts) {
|
|
|
507
508
|
TOOLS_LIST: tools.length ? tools.map((t) => '- ' + t.name + ': ' + t.role).join('\n') : '- none selected (re-run the installer with --tools codecalc to add the calculator and code runner)',
|
|
508
509
|
CODECALC_STATUS: codecalc ? 'installed alongside this folder (see `CODECALC.md`)' : 'not selected; the rule below still binds, do the arithmetic with any tool that computes rather than guesses',
|
|
509
510
|
OBSIDIAN_TC_STATUS: tools.some((t) => t.id === 'obsidian-tc') ? 'selected (see `OBSIDIAN-TC.md`); the tool names below are live calls' : 'not selected; the rule below still binds against whatever store you keep (a notes folder, a wiki, a repo of markdown), the tool names are what obsidian-tc would give you',
|
|
511
|
+
CONTEXT7_STATUS: tools.some((t) => t.id === 'context7') ? 'selected (see `CONTEXT7.md`); the tool names below are live calls' : 'not selected; the rule below still binds, read the vendor docs or source by hand before trusting them',
|
|
510
512
|
DATE: new Date().toISOString().slice(0, 10),
|
|
511
513
|
LEVEL_ID: String(level),
|
|
512
514
|
LEVEL_NAME: lvl.name,
|
package/templates/README.md
CHANGED
|
@@ -4,11 +4,11 @@ Everything the installer can write, organized by the level that adds it. Files a
|
|
|
4
4
|
|
|
5
5
|
| Folder | Written at | Contents |
|
|
6
6
|
|---|---|---|
|
|
7
|
-
| `common/` | every level | the start-here README, `TASK_BUNDLE.md`, `protocols/` (build, propagate, gap analysis, deep research, numbers and logic, memory and record) |
|
|
7
|
+
| `common/` | every level | the start-here README, `TASK_BUNDLE.md`, `protocols/` (build, propagate, gap analysis, deep research, numbers and logic, memory and record, docs then prove) |
|
|
8
8
|
| `beginner/` | every level | `ORCHESTRATOR.md`, the single-agent routing rules |
|
|
9
9
|
| `agents/` | every level, one variant | the primary agent's loading surface: Claude Code subagents (plus `.claude/hooks/route-gate.mjs` and `subagent-context.mjs`, and `settings.hooks.snippet.json` to wire them in), Antigravity custom agents, or a paste snippet |
|
|
10
10
|
| `intermediate/` | level 2+ | `ROUTING.md`, `TIERS.md`, `DELEGATION_MATRIX.md`, `RESEARCH_TRIAGE.md`, `CLI-RUN.md` |
|
|
11
11
|
| `advanced/` | level 3 | `vm/`: gateway config, compose file, box rules, privacy gates, scheduled jobs |
|
|
12
|
-
| `tools/` | when selected | companion tools the AIs call: `codecalc
|
|
12
|
+
| `tools/` | when selected | companion tools the AIs call: `codecalc/`, `obsidian-tc/` and `context7/` (install doc + MCP snippets each). See `tools/README.md` |
|
|
13
13
|
|
|
14
14
|
Agent definitions under `agents/claude-code/` and `agents/agy/` are written to the PROJECT root (`--project`), not `--dir`, because that is where those CLIs read them. A `README.md` at the root of a tier folder (like this one) documents the repo and is not installed. `common/README.md` is the exception: it is the user's start-here file. READMEs deeper in (`protocols/`, `vm/`, `vm/jobs/`) are installed as folder indexes.
|
|
@@ -18,6 +18,7 @@ Generated {{DATE}} for: `{{AI_IDS}}`. Installed at `{{INSTALL_DIR}}`; the system
|
|
|
18
18
|
| A local model runtime (if selected) | the privacy lane | nothing leaves the box |
|
|
19
19
|
| Scheduled jobs (`jobs/`) | the weekly gap-analysis audit (lane: `{{AUDIT_LANE}}`), and anything else recurring | the gateway, or `cli-run` |
|
|
20
20
|
| codecalc (if selected) | the calculator, code runner and logic checker every agent here calls; stdio, offline, no key | nothing; it computes locally |
|
|
21
|
+
| context7 (if selected) | version-aware docs for the libraries the box's agents build against; paired with codecalc, docs then a run | the hosted endpoint over HTTPS (or a local `npx` server over stdio, still no cloud account required) |
|
|
21
22
|
|
|
22
23
|
## Setup, in order
|
|
23
24
|
|
|
@@ -7,7 +7,7 @@
|
|
|
7
7
|
# - the previous successful report is NEVER truncated: output goes to a temp file and is
|
|
8
8
|
# renamed into place only on a clean exit; failed output is kept beside it for diagnosis
|
|
9
9
|
# - the lane runs with the strongest boundary it offers ({{AUDIT_LANE_BOUNDARY_NOTE}})
|
|
10
|
-
# - rc
|
|
10
|
+
# - any nonzero rc from cli-run (10 to 18) means no report was produced; the timer's journal shows it
|
|
11
11
|
set -uo pipefail
|
|
12
12
|
INSTALL_DIR={{INSTALL_DIR_SH}}
|
|
13
13
|
AUDIT_LANE="{{AUDIT_LANE}}"
|
|
@@ -46,9 +46,13 @@ Any figure someone will act on, any comparison you state, any complexity, equiva
|
|
|
46
46
|
|
|
47
47
|
Anything durable is searched for before it is written, its folder index is corrected in the same pass, and one writer records. `protocols/memory-and-record.md`. Companion tool (optional, needs an Obsidian vault): obsidian-tc, {{OBSIDIAN_TC_STATUS}}.
|
|
48
48
|
|
|
49
|
-
##
|
|
49
|
+
## Docs, then prove
|
|
50
50
|
|
|
51
|
-
|
|
51
|
+
Before writing code against a library, SDK, API or CLI you have not confirmed the current shape of, pull current docs; then a run, not the doc, is what proves it behaves that way. `protocols/docs-then-prove.md`. Companion tool (optional, needs a network call): Context7, {{CONTEXT7_STATUS}}. Paired with codecalc, {{CODECALC_STATUS}}: docs say what it is supposed to do, codecalc's run says what it actually does.
|
|
52
|
+
|
|
53
|
+
## The seven protocols
|
|
54
|
+
|
|
55
|
+
`protocols/build-protocol.md` · `protocols/propagate.md` · `protocols/gap-analysis.md` · `protocols/deep-research.md` · `protocols/numbers-and-logic.md` · `protocols/memory-and-record.md` · `protocols/docs-then-prove.md`. Each is a set of questions that can be answered wrong. That is the design, not a flaw.
|
|
52
56
|
|
|
53
57
|
## When you outgrow this
|
|
54
58
|
|
|
@@ -35,8 +35,10 @@ These are the same steps, in the same order, that the installer printed in your
|
|
|
35
35
|
| `protocols/deep-research.md` | The source set is unknown, several sources must be reconciled, and the answer will be cited later. |
|
|
36
36
|
| `protocols/numbers-and-logic.md` | You are about to state a number, a comparison, a complexity or an equivalence. Compute it. |
|
|
37
37
|
| `protocols/memory-and-record.md` | You are about to write anything durable. Search first, keep the index true, one writer. |
|
|
38
|
+
| `protocols/docs-then-prove.md` | You are about to write code against a library, SDK, API or CLI. Current docs first, then a run proves it. |
|
|
38
39
|
| `CODECALC.md` | Present when you selected codecalc: install, per-agent registration, the skill. |
|
|
39
40
|
| `OBSIDIAN-TC.md` | Present when you selected obsidian-tc: what you need first, install, per-agent registration, the security posture. |
|
|
41
|
+
| `CONTEXT7.md` | Present when you selected context7: what you need first, install, per-agent registration, the security posture. |
|
|
40
42
|
|
|
41
43
|
Level 2 adds `ROUTING.md`, `TIERS.md`, `DELEGATION_MATRIX.md`, `RESEARCH_TRIAGE.md`, `CLI-RUN.md` and `bin/cli-run.mjs`. Level 3 adds `vm/`. If those files are here, read `ROUTING.md` instead of `ORCHESTRATOR.md`: it is the multi-lane version, and the snippet your agent loads already points at it. `ORCHESTRATOR.md` stays as the single-agent fallback for a session where only one AI is available.
|
|
42
44
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# protocols/
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Seven procedures. Each one is a list of questions whose answers can be wrong.
|
|
4
4
|
|
|
5
5
|
| File | Fires when | The gate |
|
|
6
6
|
|---|---|---|
|
|
@@ -10,5 +10,6 @@ Six procedures. Each one is a list of questions whose answers can be wrong.
|
|
|
10
10
|
| `deep-research.md` | The source set is unknown and the answer will be cited later | Parallel engines, then triage; disagreement is the signal |
|
|
11
11
|
| `numbers-and-logic.md` | You are about to state a number, a comparison, a complexity, an equivalence | Computed by a tool (codecalc) or not stated |
|
|
12
12
|
| `memory-and-record.md` | You are about to write anything durable | Searched first, indexed in the same pass, one writer (obsidian-tc when selected) |
|
|
13
|
+
| `docs-then-prove.md` | You are about to write code against a library, SDK, API or CLI | Current docs first (Context7 when selected), then a run proves it; the run wins on disagreement |
|
|
13
14
|
|
|
14
15
|
Not for lookups, prose edits, bulk classification or one-line config. Those get none of this.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# Docs, then prove: documentation is a lead, never a verdict
|
|
2
|
+
|
|
3
|
+
**Why this is not `numbers-and-logic.md`.** That protocol is scoped to arithmetic, comparisons and complexity claims. This one is scoped to a different failure: trusting what a library, SDK, API or CLI is documented to do instead of checking what it actually does on this version, in this codebase. Two different mistakes, two different tools, one rule underneath both: a model that feels finished is not the same thing as a model that checked.
|
|
4
|
+
|
|
5
|
+
Companion tool for this rule: **Context7** (Upstash), {{CONTEXT7_STATUS}}. It pairs with **codecalc**, {{CODECALC_STATUS}}: Context7 tells the agent what the library is SUPPOSED to do (current, version-aware docs); codecalc runs the code and proves what it actually does. Docs never stand as proof on their own, and where a doc and a run disagree, the run wins and the source settles it.
|
|
6
|
+
|
|
7
|
+
## When calling is mandatory
|
|
8
|
+
|
|
9
|
+
| You are about to | Use |
|
|
10
|
+
|---|---|
|
|
11
|
+
| write code against a library, SDK, API or CLI you have not confirmed the current signature for | Context7 (`resolve-library-id` then `query-docs`, or `ctx7 library` / `ctx7 docs`), then write the call |
|
|
12
|
+
| claim a documented behaviour is what the code actually does | run it (`execute_code`, the project's own test suite, or a REPL), never the doc alone |
|
|
13
|
+
| a doc and a run disagree | the run wins. Say so, and say what the doc got wrong (a stale version, a changed default, a deprecated flag) |
|
|
14
|
+
| training-data recall of a library's API surface, with no source open | treat it as a guess until a doc or a run confirms it; version drift and renamed APIs are the common failure, not a rare one |
|
|
15
|
+
|
|
16
|
+
Trivial, unversioned standard-library calls you would bet the build on are exempt. Anything with a version number attached to its behaviour is not.
|
|
17
|
+
|
|
18
|
+
## How to report a documentation-derived claim
|
|
19
|
+
|
|
20
|
+
Name the library, the version if Context7 returned one, and that it came from docs, not a run: "per Context7, `/vercel/next.js@14.0.0`'s middleware API takes X" reads differently from "Next.js middleware takes X", because the first can be checked against a version and the second cannot. If the claim was then verified by running it, say that too, and which one actually settled the question.
|
|
21
|
+
|
|
22
|
+
## Without Context7
|
|
23
|
+
|
|
24
|
+
The rule still binds. Read the vendor's own README, changelog or source before writing code against it, the same way this repository's own `AGENTS.md` asks: read the actual API, not a recollection of it. What is not allowed is code written against a remembered shape of a library that was never opened this session.
|
|
25
|
+
|
|
26
|
+
Source: https://github.com/upstash/context7
|
|
@@ -56,16 +56,40 @@ qwen is the lane whose own success flags lie: an upstream 400 comes back as exit
|
|
|
56
56
|
|
|
57
57
|
## Exit codes
|
|
58
58
|
|
|
59
|
-
| Code | Meaning |
|
|
60
|
-
|
|
61
|
-
| 0 | structurally accepted non-empty response, every `--expect-*` contract met |
|
|
62
|
-
| 10 | ran and
|
|
63
|
-
| 11 | no output at all |
|
|
64
|
-
| 12 | timed out; the lane and every descendant in its process group were killed |
|
|
65
|
-
| 13 |
|
|
66
|
-
|
|
|
67
|
-
|
|
|
68
|
-
|
|
|
59
|
+
| Code | Class | Meaning | What to do |
|
|
60
|
+
|---|---|---|---|
|
|
61
|
+
| 0 | `ok` | structurally accepted non-empty response, every `--expect-*` contract met | use it; if `refused=N` is above 0, read the problem line |
|
|
62
|
+
| 10 | `empty` | ran and delivered nothing, or a contract was unmet | rerun once, or use another lane |
|
|
63
|
+
| 11 | `no_output` | no output at all | rerun once; check the lane runs on its own |
|
|
64
|
+
| 12 | `timeout` | timed out; the lane and every descendant in its process group were killed | raise `--timeout` or split the brief |
|
|
65
|
+
| 13 | `unavailable` | binary missing, disabled in `lanes.json`, or `lanes.json` malformed | install or enable the lane |
|
|
66
|
+
| 14 | `auth` | the lane's own error says a credential is missing or it is not logged in | set the credential; retrying cannot help |
|
|
67
|
+
| 15 | `quota` | the lane's own error says usage limit, credits or rate limit | switch lanes or wait for the reset |
|
|
68
|
+
| 16 | `rejected` | the upstream rejected the request: an unknown model id, a malformed request | fix the id, flag or request it names |
|
|
69
|
+
| 17 | `refused` | no deliverable, and the lane reports tool calls a hook or deny rule blocked | adjust the rule, or give the lane the tool |
|
|
70
|
+
| 18 | `cut_short` | no trustworthy finish: a missing or non-success terminal event, a lane killed by a signal, output past the 16 MiB buffer, or a nonzero vendor exit nothing above explains (hermes' own exit 2, bad args or an empty response, stays `empty`) | rerun once, then read the lane's stderr |
|
|
71
|
+
| 130 / 143 | `interrupted` | cli-run itself received SIGINT / SIGTERM; the lane's process group was killed first, then the temp dir removed | |
|
|
72
|
+
| 2 | | usage error in cli-run itself | |
|
|
73
|
+
|
|
74
|
+
A nonzero vendor exit is never `ok`, even when parseable text came back; the vendor's own code is kept in the log as `cli_rc`, and the bounded head of its stderr is shown on your terminal.
|
|
75
|
+
|
|
76
|
+
## Why a run failed: the class
|
|
77
|
+
|
|
78
|
+
Exit 10 used to cover causes that need opposite responses. A missing API key is not a model fault, and retrying a spent quota cannot help. So every run lands in exactly one class, and on the terminal every failure (and every `ok` with refused calls) gets two more lines:
|
|
79
|
+
|
|
80
|
+
```
|
|
81
|
+
cli-run[qwen] exit_nonzero rc=14 class=auth refused=0 0.8s raw=212B route=lane default :: lane exited 1; subtype="error_during_execution": Missing API key ...
|
|
82
|
+
cli-run problem: cli-run[qwen] auth: Missing API key ...
|
|
83
|
+
cli-run fix: set the credential the message above names (its environment variable, or the lane's own login command), then rerun
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
A calling agent relays both lines and asks before fixing anything.
|
|
87
|
+
|
|
88
|
+
- **Signals come from each lane's authoritative error fields only.** codex: its own `error`, `turn.failed` and error-item events plus stderr. agy: the terminal result's status and error plus stderr. qwen: the terminal event's full error text (never the clipped display line) plus stderr. hermes: stderr on a nonzero exit, never its stdout. Never the model's prose, so an answer that explains what "rate limit" means is not a quota failure. grok and hermes have no native auth signal, and none is invented.
|
|
89
|
+
- **Precedence** when several are present: `auth`, `quota`, `rejected`, `refused`, `cut_short`, `empty`. A missing credential explains everything downstream of it.
|
|
90
|
+
- **`refused`** counts tool calls a hook or deny rule blocked: qwen's `permission_denials`, agy's deny-rule `TOOL_ERROR` steps, one per codex router `Rejected(` line on stderr, and grok's session transcript (`~/.grok/sessions/<cwd>/<sessionId>/updates.jsonl`: hook runs with `blocked`, and "Denied by permission policy"). grok's `sessionId` comes from lane output, so it must match `^[A-Za-z0-9-]{8,64}$`, the resolved path must stay inside the sessions root, every counted line must name that session, and the read stops at 5 MiB. `null` means the lane gave no readable signal (hermes always), which is an unknown, never 0.
|
|
91
|
+
- **Refused with a deliverable is still exit 0.** The deliverable exists; `refused=N` and the problem and fix lines say what was blocked.
|
|
92
|
+
- **Redacted.** The status, problem and stderr lines pass through one redaction pass before they are printed: JSON credential keys (escaped quotes, unterminated values and JSON escaped inside a string included), `Authorization:` values of any scheme, bearer values, URL query credentials, and the `sk-`, `xai-`, `ghp_` and `AIza` key prefixes. Redaction runs before any clipping, so a long token is never cut into an unrecognisable fragment. Denial text is searched in at most the first 256 KiB of stdout and of stderr, with bounded patterns, so hostile output cannot stall the wrapper after the lane has exited. None of these lines is ever logged.
|
|
69
93
|
|
|
70
94
|
## The route: which model, and how hard it thinks
|
|
71
95
|
|
|
@@ -97,7 +121,7 @@ A flag beats `defaults`; `defaults` beats nothing. Each vendor spells these diff
|
|
|
97
121
|
|
|
98
122
|
Three rules that keep this honest:
|
|
99
123
|
|
|
100
|
-
- **A level `cli-run` does not recognise is not rejected here.** Levels are the vendor's, they change, and guessing the valid set would date this tool. An unknown level is refused by the lane and surfaces as
|
|
124
|
+
- **A level `cli-run` does not recognise is not rejected here.** Levels are the vendor's, they change, and guessing the valid set would date this tool. An unknown level is refused by the lane and surfaces as a class (codex reports it as `rejected`, exit 16; a lane with no rejection signal as `cut_short`, exit 18), with the lane's own exit code in the log as `cli_rc` and its stderr on your terminal.
|
|
101
125
|
- **`--effort` on qwen is a usage error, not a silent drop.** A flag that vanishes leaves you believing a route that never ran.
|
|
102
126
|
- **`--effort auto` is bounded.** It uses prompt size, or an audit's changed-file evidence, and resolves only medium or high. An audit is always high. Auto is a heuristic, not a measurement: name `xhigh` explicitly for security-critical or irreversible work.
|
|
103
127
|
- **Values are charset-bounded** (letters, digits, and `. _ : @ / + -`, no leading dash, 64 characters). A model id becomes an argv element and, on codex, part of a TOML value; bounding it is what stops either from being escaped.
|
|
@@ -108,7 +132,7 @@ Three rules that keep this honest:
|
|
|
108
132
|
|
|
109
133
|
## Log
|
|
110
134
|
|
|
111
|
-
`~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, rc, the lane's own exit code, signal, seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, the route (`model_requested`, `effort_requested`, and `model_source` / `effort_source`, each one of `flag`, `lanes.json` or `lane_default`), auto-sizing evidence (`effort_resolved`, `effort_basis`, `effort_scope`), and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). `effort_basis` is exactly one of `explicit`, `prompt_chars`, `audit_floor`, or `none`. Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument, and so does "we route audits at high effort".
|
|
135
|
+
`~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, `class` (one of the classes above), rc, the lane's own exit code (`cli_rc`), signal, `refused` (an integer, or `null` when the lane gives no signal), seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, the route (`model_requested`, `effort_requested`, and `model_source` / `effort_source`, each one of `flag`, `lanes.json` or `lane_default`), auto-sizing evidence (`effort_resolved`, `effort_basis`, `effort_scope`), and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). `effort_basis` is exactly one of `explicit`, `prompt_chars`, `audit_floor`, or `none`. Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument, and so does "we route audits at high effort".
|
|
112
136
|
|
|
113
137
|
The log records what was **requested**, on every record including a run refused before the lane started. It does not record an actual. Reporting is inconsistent: grok returns a `modelUsage` block naming a model, the other four lanes return nothing of the kind, so an `actual` field would be populated for one lane and empty for four. It would also be a provider-supplied string, and this log holds fixed codes and bounded caller-supplied values only. `model_source: "lane_default"` is the honest way to say this run inherited something invisible from here.
|
|
114
138
|
|
|
@@ -122,7 +146,7 @@ Absent: every lane enabled, nothing pinned. Present but malformed or unreadable:
|
|
|
122
146
|
|
|
123
147
|
## A killed lane is not a deliverable
|
|
124
148
|
|
|
125
|
-
A lane that dies by signal has no honest exit status. Whatever it printed first is discarded; the run reports `killed`
|
|
149
|
+
A lane that dies by signal has no honest exit status. Whatever it printed first is discarded; the run reports `killed`, class `cut_short`, exit 18.
|
|
126
150
|
|
|
127
151
|
## Interrupting cli-run kills the lane too
|
|
128
152
|
|
|
@@ -54,6 +54,10 @@ Every number, comparison, complexity or equivalence claim goes through a tool th
|
|
|
54
54
|
|
|
55
55
|
One writer per run; every other lane proposes. Search before writing, index in the same pass (`protocols/memory-and-record.md`; companion, optional: obsidian-tc, {{OBSIDIAN_TC_STATUS}}).
|
|
56
56
|
|
|
57
|
+
## Docs, then prove
|
|
58
|
+
|
|
59
|
+
A lane's recall of a library's API is a lead, not a verdict, the same as its arithmetic. Pull current, version-specific docs before writing a call against anything you have not confirmed this session (`protocols/docs-then-prove.md`; companion, optional: Context7, {{CONTEXT7_STATUS}}). Then prove the doc was right by running it, the same tool that already owns numbers: codecalc, {{CODECALC_STATUS}}. Where the two disagree, the run wins.
|
|
60
|
+
|
|
57
61
|
## Modifier rules
|
|
58
62
|
|
|
59
63
|
{{PLAN_BIG_LINE}}{{INLINE_THRESHOLD_NOTE}}
|
|
@@ -6,5 +6,6 @@ Companion tools: not AIs, but things the AIs call. Each subfolder is written onl
|
|
|
6
6
|
|---|---|---|
|
|
7
7
|
| `codecalc/` | when codecalc is selected (recommended, default yes) | `CODECALC.md` (install, per-client registration, the skill) and `mcp/` snippets for the agents its own `setup --write` does not cover |
|
|
8
8
|
| `obsidian-tc/` | when obsidian-tc is selected (optional, default no; needs an Obsidian vault, Node 24+, Ollama or a cloud embeddings key) | `OBSIDIAN-TC.md` (what you need first, install, per-agent registration, security posture) and `mcp/` snippets |
|
|
9
|
+
| `context7/` | when context7 is selected (optional, default no; needs a network call, a Node 18+ local alternative, an optional API key) | `CONTEXT7.md` (what you need first, install, per-agent registration, security posture) and `mcp/` snippets |
|
|
9
10
|
|
|
10
|
-
The rules the tools serve, `protocols/numbers-and-logic.md
|
|
11
|
+
The rules the tools serve, `protocols/numbers-and-logic.md`, `protocols/memory-and-record.md` and `protocols/docs-then-prove.md`, are in `common/` and are written at every level whether or not a tool was selected: the rule binds, the tool makes it cheap to follow. `src/install.js` writes `templates/tools/<id>/` for every selected tool that has a folder here.
|
|
@@ -0,0 +1,82 @@
|
|
|
1
|
+
# Context7: version-aware docs for your agent (optional)
|
|
2
|
+
|
|
3
|
+
Repo: https://github.com/upstash/context7 · Upstash · MIT · hosted MCP server, or run it yourself with `npx`
|
|
4
|
+
|
|
5
|
+
**Optional.** Skip this if your agent already opens the real source or docs of anything it calls before writing code against it. The rule it serves, `protocols/docs-then-prove.md`, binds either way.
|
|
6
|
+
|
|
7
|
+
**Paired with codecalc, not a replacement for it.** Context7 answers "what is this library documented to do, on this version": current docs and code examples, pulled straight into the prompt. codecalc answers "what does this code actually do": it runs the thing. A doc can be stale, a version can drift, a default can change between releases; only a run proves current behaviour. Where the two disagree, the run wins and the source settles it.
|
|
8
|
+
|
|
9
|
+
## What it gives every agent in this folder
|
|
10
|
+
|
|
11
|
+
| Need in the protocol | Context7 tool |
|
|
12
|
+
|---|---|
|
|
13
|
+
| find the Context7 id for a library, SDK, API or CLI by name | `resolve-library-id` (MCP) or `ctx7 library <name> <query>` (CLI) |
|
|
14
|
+
| pull current, version-specific docs and code examples before writing a call | `query-docs` (MCP, needs a library id) or `ctx7 docs <libraryId> <query>` (CLI) |
|
|
15
|
+
| skip the library-matching step when you already know the exact package | address it directly, `/org/project` (e.g. `/vercel/next.js`), optionally with `@version` |
|
|
16
|
+
|
|
17
|
+
Context7 indexes documentation from GitHub, GitLab, Bitbucket, websites, `llms.txt` files, OpenAPI specs and Confluence spaces. It is community-contributed: the upstream disclaimer says it cannot guarantee every project's docs are accurate, complete or current, which is exactly why `protocols/docs-then-prove.md` treats a Context7 answer as a lead codecalc (or your own test run) still has to confirm, never as a verdict.
|
|
18
|
+
|
|
19
|
+
## What you need first
|
|
20
|
+
|
|
21
|
+
| Requirement | Why | Notes |
|
|
22
|
+
|---|---|---|
|
|
23
|
+
| **Node.js 18+** | runs the local server or the `ctx7` CLI | only needed for the local (`npx`) connection; the hosted endpoint needs nothing local at all |
|
|
24
|
+
| **A network call to `mcp.context7.com`** | Context7 is a hosted service; there is no fully offline mode | the anonymous rate limit works with no account. Unlike codecalc (offline) and obsidian-tc (local), this tool always leaves the machine |
|
|
25
|
+
| A free **`CONTEXT7_API_KEY`** (optional) | raises the anonymous rate limit | get one at [context7.com/dashboard](https://context7.com/dashboard); it is never pasted into a snippet in `mcp/` (see "Higher rate limits" below) |
|
|
26
|
+
|
|
27
|
+
## Install
|
|
28
|
+
|
|
29
|
+
Two ways to connect, remote first (Context7's own documented default, and the one that needs nothing installed). Both connect keyless, at the anonymous rate limit, which is enough to try it:
|
|
30
|
+
|
|
31
|
+
```bash
|
|
32
|
+
# Remote (recommended): no install. Point your MCP client at the hosted endpoint,
|
|
33
|
+
# https://mcp.context7.com/mcp. No key needed: anonymous requests work at a
|
|
34
|
+
# lower rate limit. OAuth is available where your client supports it
|
|
35
|
+
# (mcp.context7.com/mcp/oauth), as an alternative to an API key, not required either.
|
|
36
|
+
|
|
37
|
+
# Local alternative: runs the MCP server on your machine over stdio.
|
|
38
|
+
npx -y @upstash/context7-mcp
|
|
39
|
+
|
|
40
|
+
# Or the one-command setup Context7 itself ships, which authenticates via OAuth,
|
|
41
|
+
# writes an API key, and can install a CLI-based skill instead of MCP:
|
|
42
|
+
npx ctx7 setup
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
Pinned form, if you want the version this installer was released against: `npx -y @upstash/context7-mcp@{{CONTEXT7_PIN}}`. Drop the pin for latest; `npx` always resolves fresh, so pinning only matters when you want a reproducible version rather than whatever shipped this week.
|
|
46
|
+
|
|
47
|
+
`npx ctx7 setup` is upstream's own guided installer; it is not run by this installer, only documented here, the same way this project never runs a vendor script for you.
|
|
48
|
+
|
|
49
|
+
## Register it with your agent (snippets in `mcp/`)
|
|
50
|
+
|
|
51
|
+
Every snippet below ships **keyless**: the remote ones point at the hosted endpoint with no `Authorization` header at all (Qwen Code's snippet keeps the non-credential `Accept` header upstream itself ships), and the Zed snippet runs the local, version-pinned `npx` server with no key in its `env` block. That is deliberate, not an oversight: see "Higher rate limits" next for why a header is not shipped by default.
|
|
52
|
+
|
|
53
|
+
| Agent | File to edit | Snippet |
|
|
54
|
+
|---|---|---|
|
|
55
|
+
| Claude Code | its MCP config (`.mcp.json`); needs `"type": "http"` next to `url` or Claude Code skips the server as misconfigured | `mcp/context7.claude-code.mcp.json` |
|
|
56
|
+
| Claude Desktop | no config file: `Settings > Connectors > Add Custom Connector`, name `Context7`, URL `https://mcp.context7.com/mcp` | none, it is a UI step |
|
|
57
|
+
| Cursor | one-click install in the upstream README, or `~/.cursor/mcp.json` | `mcp/context7.mcpServers.json` |
|
|
58
|
+
| VS Code | `.vscode/mcp.json` (key is `servers`, remote type is `http`) | `mcp/context7.vscode.mcp.json` |
|
|
59
|
+
| Zed | `~/.config/zed/settings.json` (key is `context_servers`); upstream ships only a local (`npx`) config for Zed, so this snippet runs the server locally rather than remote | `mcp/context7.zed.settings.json` |
|
|
60
|
+
| Codex CLI | `~/.codex/config.toml` | `mcp/context7.codex.config.toml` |
|
|
61
|
+
| Antigravity `agy` | its MCP config file | `mcp/context7.agy.mcp_config.json` |
|
|
62
|
+
| Qwen Code | `~/.qwen/settings.json` under `mcpServers`; note the field is `httpUrl`, not `url`, a different shape from every other client here | `mcp/context7.qwen.settings.json` |
|
|
63
|
+
|
|
64
|
+
Merge the block; do not replace the file. `mcp/context7.mcpServers.json` (Cursor) and `mcp/context7.claude-code.mcp.json` look alike but are not interchangeable: Claude Code requires the `"type": "http"` field and Cursor's own docs show plain `{"url": ...}` with no `type` at all.
|
|
65
|
+
|
|
66
|
+
## Higher rate limits (optional key)
|
|
67
|
+
|
|
68
|
+
Anonymous works. If you hit the rate limit and want a key, add it the correct way for your client, and never as a literal value pasted into any of the snippets above:
|
|
69
|
+
|
|
70
|
+
- **Codex CLI**: under the `[mcp_servers.context7]` table in `~/.codex/config.toml`, add a `bearer_token_env_var` entry naming the environment variable `CONTEXT7_API_KEY`. Codex reads the token from that variable at connect time and sends it as the `Authorization` header itself; the config file never holds the value ([Codex MCP docs](https://developers.openai.com/codex/mcp)).
|
|
71
|
+
- **Claude Code** (`mcp/context7.claude-code.mcp.json`): add a `headers` object to the `context7` entry with an `Authorization` field whose value is `Bearer` followed by a `${CONTEXT7_API_KEY}` reference. Claude Code expands `${VAR}` references in a remote server's `headers` at load time, and `CONTEXT7_API_KEY` is not one of the credential names it deliberately reads as empty (those are Claude/Anthropic-specific). Set the variable in your environment before launching; an unset variable still loads with the literal, unexpanded reference sent as the header, and Context7 answers every call with "Invalid API key" instead of running anonymously ([Claude Code MCP docs](https://code.claude.com/docs/en/mcp)). That failure, not a missing feature, is why this snippet ships with no header at all.
|
|
72
|
+
- **Claude Desktop**: the Connectors UI has its own key field; use it there rather than editing a file.
|
|
73
|
+
- **Any local `npx` connection** (Zed, or the local alternative for any other client): export `CONTEXT7_API_KEY` in the shell that launches your editor or agent. A spawned stdio child process inherits its parent's environment by default, so `npx -y @upstash/context7-mcp` picks it up with no config edit; this is the same environment variable name Context7's own Docker MCP Toolkit config and its GitHub Copilot integration use to feed the server a key. If your client does not pass its environment through to the child (uncommon), either stay anonymous, or check whether that client's own config format has an `env` block that itself supports an environment-variable reference (Claude Code's does, described above; not every client's does) rather than typing the key in.
|
|
74
|
+
- **Cursor, VS Code, Qwen Code, Antigravity `agy`, or any other client using the `mcpServers.json`/`vscode.mcp.json`/`qwen.settings.json`/`agy.mcp_config.json` snippet**: check that client's own docs for whether it expands an environment-variable reference inside a remote server's `headers` before adding one. This is not confirmed for any of them here. If it does not expand, the literal, unexpanded text becomes the header value and every call fails with "Invalid API key" instead of running anonymously, which is worse than shipping no header at all.
|
|
75
|
+
|
|
76
|
+
## Security posture, read before you send anything through it
|
|
77
|
+
|
|
78
|
+
Only the library name and your query text reach Context7's API; your source code is never uploaded. It is still a third-party network call on every lookup, unlike codecalc (offline) and obsidian-tc (local by default): do not route a query that would leak a private project name, an internal library name, or anything else you would not put in a public search box. The docs it indexes are community-contributed, not vetted by Context7 or by this installer; a suspicious or malicious-looking result is reportable upstream from the project's page. An API key raises your rate limit; it is not a secret worth protecting the way a database credential is, but it still never belongs in a committed file, only in your environment.
|
|
79
|
+
|
|
80
|
+
## Level 3
|
|
81
|
+
|
|
82
|
+
Runs the same way at any level: the remote endpoint over HTTPS, or the local `npx` server over stdio next to the orchestrator CLI on the box. Nothing about it changes on a box except that the box, not your laptop, is the machine making the network call.
|