model-orchestrator 0.1.35 → 1.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +31 -21
- package/CHANGELOG.md +58 -1
- package/README.md +129 -110
- package/SECURITY.md +7 -3
- package/bin/README.md +57 -6
- package/bin/aunx.js +7 -0
- package/bin/cli-run.mjs +21 -15
- package/bin/cli.js +376 -257
- package/docs/README.md +15 -18
- package/docs/catalog.md +236 -44
- package/docs/companions.md +28 -10
- package/docs/guarantees.md +21 -12
- package/docs/how-it-routes.md +49 -42
- package/docs/install.md +141 -33
- package/docs/part-1-beginner.md +37 -45
- package/docs/part-2-intermediate.md +34 -52
- package/docs/part-3-advanced.md +36 -26
- package/docs/security-review-history.md +39 -0
- package/llms.txt +24 -25
- package/package.json +15 -8
- package/proof/README.md +100 -0
- package/proof/gate-demo.cast +9 -0
- package/proof/gate-demo.gif +0 -0
- package/proof/results.json +198 -0
- package/proof/scripts/check-gate.js +26 -0
- package/proof/scripts/install-time.js +16 -0
- package/proof/scripts/lib.js +73 -0
- package/proof/scripts/measure.js +15 -0
- package/proof/scripts/missing-results.js +30 -0
- package/proof/scripts/record-gate.js +38 -0
- package/proof/scripts/render.js +18 -0
- package/proof/scripts/runner-overhead.js +21 -0
- package/src/README.md +10 -3
- package/src/activation-ownership.js +19 -0
- package/src/apply-companions.js +104 -0
- package/src/apply-snippets.js +60 -28
- package/src/aunx.js +272 -0
- package/src/bounded-file.js +31 -0
- package/src/catalog.js +257 -121
- package/src/install.js +483 -212
- package/src/plugin.js +13 -4
- package/src/postinstall.js +57 -0
- package/src/roles.js +184 -0
- package/src/uninstall.js +128 -10
- package/templates/README.md +19 -2
- package/templates/advanced/README.md +2 -2
- package/templates/advanced/vm/ENVIRONMENT.md +8 -0
- package/templates/advanced/vm/PRIVACY_GATES.md +17 -19
- package/templates/advanced/vm/README.md +25 -20
- package/templates/advanced/vm/box-CLAUDE.md +19 -18
- package/templates/advanced/vm/docker-compose.yml +2 -1
- package/templates/advanced/vm/jobs/README.md +31 -2
- package/templates/advanced/vm/jobs/weekly-audit.service +7 -2
- package/templates/advanced/vm/jobs/weekly-audit.sh +24 -17
- package/templates/advanced/vm/setup-vm.sh +49 -2
- package/templates/agents/README.md +2 -2
- package/templates/agents/agy/README.md +20 -3
- package/templates/agents/agy/builder.md +11 -7
- package/templates/agents/agy/bulk-worker.md +9 -7
- package/templates/agents/agy/code-reviewer.md +13 -7
- package/templates/agents/agy/deep-planner.md +10 -7
- package/templates/agents/agy/done-verifier.md +13 -22
- package/templates/agents/agy/finding-verifier.md +14 -22
- package/templates/agents/agy/live-researcher.md +10 -7
- package/templates/agents/agy/reader.md +10 -12
- package/templates/agents/claude-code/README.md +18 -14
- package/templates/agents/claude-code/builder.md +10 -15
- package/templates/agents/claude-code/bulk-worker.md +8 -10
- package/templates/agents/claude-code/code-reviewer.md +11 -17
- package/templates/agents/claude-code/deep-planner.md +9 -11
- package/templates/agents/claude-code/done-verifier.md +12 -33
- package/templates/agents/claude-code/finding-verifier.md +13 -39
- package/templates/agents/claude-code/live-researcher.md +9 -11
- package/templates/agents/claude-code/reader.md +9 -18
- package/templates/agents/snippets/chat.md +9 -10
- package/templates/agents/snippets/claude-code.md +17 -18
- package/templates/agents/snippets/generic.md +9 -11
- package/templates/agents/snippets/route-gate.mjs +2 -2
- package/templates/agents/snippets/route-metrics.mjs +1 -1
- package/templates/agents/snippets/subagent-context.mjs +4 -4
- package/templates/beginner/ORCHESTRATOR.md +31 -36
- package/templates/beginner/README.md +1 -1
- package/templates/common/ACCEPTANCE_CHECKS.json +12 -0
- package/templates/common/CONTEXT.md +37 -0
- package/templates/common/DECISIONS.md +11 -0
- package/templates/common/README.md +24 -11
- package/templates/common/TASK_BRIEF.md +84 -0
- package/templates/common/protocols/README.md +14 -11
- package/templates/common/protocols/acceptance-checks.md +15 -0
- package/templates/common/protocols/build-protocol.md +91 -106
- package/templates/common/protocols/context-file.md +10 -0
- package/templates/common/protocols/decision-log.md +9 -0
- package/templates/common/protocols/deep-research.md +20 -34
- package/templates/common/protocols/docs-then-prove.md +13 -18
- package/templates/common/protocols/gap-analysis.md +15 -21
- package/templates/common/protocols/memory-and-record.md +21 -20
- package/templates/common/protocols/numbers-and-logic.md +20 -26
- package/templates/common/protocols/propagate.md +18 -27
- package/templates/intermediate/CLI-RUN.md +83 -113
- package/templates/intermediate/DELEGATION_MATRIX.md +9 -3
- package/templates/intermediate/README.md +3 -3
- package/templates/intermediate/RESEARCH_TRIAGE.md +23 -15
- package/templates/intermediate/ROUTING.md +54 -51
- package/templates/intermediate/TIERS.md +37 -76
- package/templates/tools/README.md +1 -1
- package/templates/tools/codecalc/CODECALC.md +4 -4
- package/templates/tools/codecalc/mcp/agy.mcp_config.json +1 -1
- package/templates/tools/codecalc/mcp/codex.config.toml +1 -1
- package/templates/tools/codecalc/mcp/mcpServers.json +1 -1
- package/templates/tools/codecalc/mcp/vscode.mcp.json +1 -1
- package/templates/tools/codecalc/mcp/zed.settings.json +1 -1
- package/templates/tools/context7/CONTEXT7.md +6 -10
- package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +3 -3
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.agy.mcp_config.json +1 -1
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.codex.config.toml +1 -1
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.mcpServers.json +1 -1
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.vscode.mcp.json +1 -1
- package/templates/tools/obsidian-tc/mcp/obsidian-tc.zed.settings.json +1 -1
- package/docs/audit-brief.md +0 -148
- package/scripts/README.md +0 -7
- package/scripts/gen-catalog.js +0 -81
- package/scripts/gen-plugin.js +0 -16
- package/scripts/record-demo.sh +0 -45
- package/templates/common/TASK_BUNDLE.md +0 -56
|
@@ -1,93 +1,54 @@
|
|
|
1
|
-
# TIERS.md: capability
|
|
1
|
+
# TIERS.md: choose capability, effort and tool reach
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
{{STACK_SUMMARY}}
|
|
4
4
|
|
|
5
|
-
|
|
6
|
-
|---|---|---|
|
|
7
|
-
| deep | ambiguous planning, architecture, strategy, hard debugging | {{PRIMARY_DEEP}} |
|
|
8
|
-
| standard | code writing, code review, execution, live research synthesis | {{PRIMARY_STANDARD}} |
|
|
9
|
-
| fast | classification, extraction, formatting, bulk summarization | {{PRIMARY_FAST}} |
|
|
10
|
-
| escalation | above deep: only when the human asks, or when deep has already run, the call is still unresolved, and the change is irreversible | your vendor's strongest model, if you have one |
|
|
11
|
-
|
|
12
|
-
Non-primary lanes are owned by `DELEGATION_MATRIX.md`.
|
|
13
|
-
|
|
14
|
-
## Effort per agent (the third lever)
|
|
15
|
-
|
|
16
|
-
Tier sets the price per token. Token discipline sets how many tokens. **Effort sets how hard each call thinks.**
|
|
17
|
-
|
|
18
|
-
`cli-run --effort auto` is a heuristic, not a measurement: prompts below 4,000 characters resolve to medium and longer prompts resolve to high. A codex `--audit` always resolves to high because stakes set the floor. Auto never resolves above high. Name `xhigh` explicitly for a security-critical or irreversible audit.
|
|
19
|
-
|
|
20
|
-
| Agent | Tier | Effort | Why |
|
|
21
|
-
|---|---|---|---|
|
|
22
|
-
| deep-planner | deep | xhigh | judges every build twice; expensive to get wrong |
|
|
23
|
-
| code-reviewer | standard | high | every endpoint is internet-facing |
|
|
24
|
-
| finding-verifier | standard | high | judging a claim is harder than producing it |
|
|
25
|
-
| builder | standard | high | a botched deploy is the costly failure |
|
|
26
|
-
| live-researcher | standard | medium | tools do the retrieval |
|
|
27
|
-
| bulk-worker | fast | low | the biggest cost win |
|
|
28
|
-
| done-verifier | fast | low | a done-signal check is a lookup, not a judgment call |
|
|
29
|
-
| reader | fast | low | digestion, not judgment |
|
|
30
|
-
|
|
31
|
-
## Three inputs, not one
|
|
32
|
-
|
|
33
|
-
Role alone does not decide a route. Two more inputs move it, and they move it in
|
|
34
|
-
opposite directions, so state them separately instead of folding them into the
|
|
35
|
-
role.
|
|
5
|
+
When choosing a model, use the job's required capability and the lane's live roster. A tier names a job; a current family alias or explicit model choice implements it.
|
|
36
6
|
|
|
37
|
-
|
|
38
|
-
on every task.
|
|
39
|
-
|
|
40
|
-
| Complexity | What it looks like | What moves |
|
|
7
|
+
| Tier | Use for | On {{PRIMARY_NAME}} |
|
|
41
8
|
|---|---|---|
|
|
42
|
-
|
|
|
43
|
-
|
|
|
44
|
-
|
|
|
45
|
-
| critical | irreversible, or it rewrites a standing rule | the escalation rule below applies |
|
|
9
|
+
| planning model | ambiguous planning, architecture, strategy, unknown causes | {{PRIMARY_DEEP}} |
|
|
10
|
+
| working model | code writing, review, execution and research synthesis | {{PRIMARY_STANDARD}} |
|
|
11
|
+
| cheap model | classification, extraction, formatting and bulk summaries | {{PRIMARY_FAST}} |
|
|
46
12
|
|
|
47
|
-
|
|
48
|
-
reasoning than the reviewer judging its output.** When the plan is airtight the
|
|
49
|
-
spec is carrying the thinking, so builder drops to medium. When the plan is
|
|
50
|
-
vague, fix the plan; do not buy reasoning to paper over it.
|
|
13
|
+
When selecting another AI tool, use `DELEGATION_MATRIX.md` and verify its current availability. When data must remain local, choose a local runtime with the required privacy boundary.
|
|
51
14
|
|
|
52
|
-
|
|
53
|
-
the ones worth naming, because their failures are not recoverable by editing the
|
|
54
|
-
code afterwards.
|
|
15
|
+
## Model and effort per job
|
|
55
16
|
|
|
56
|
-
|
|
57
|
-
lost data, or something you can't undo. Most tasks are low-stakes and route
|
|
58
|
-
normally.
|
|
17
|
+
When routing, consider three levers together: tier sets capability and price; scoped context limits token use; effort sets how much reasoning the call applies.
|
|
59
18
|
|
|
60
|
-
|
|
|
61
|
-
|
|
62
|
-
|
|
|
63
|
-
|
|
|
64
|
-
|
|
|
65
|
-
|
|
|
19
|
+
| Role | Starting tier | Starting effort | Reassess when |
|
|
20
|
+
|---|---|---|---|
|
|
21
|
+
| {{TIER_PLANNER_ROLE}} | planning model | xhigh where supported | The decision can be resolved from a known plan or needs a new capability |
|
|
22
|
+
| {{TIER_REVIEW_ROLE}} | working model | high | Security, privacy or irreversible effects raise the review scope |
|
|
23
|
+
| {{TIER_FINDING_ROLE}} | working model | high | Reproduction needs another runtime or access path |
|
|
24
|
+
| {{TIER_BUILDER_ROLE}} | working model | high | Architecture, security or irreversible work needs xhigh and suitable model capability |
|
|
25
|
+
| {{TIER_LIVE_ROLE}} | working model | medium | Synthesis becomes complex or sources disagree |
|
|
26
|
+
| {{TIER_BULK_ROLE}} | cheap model | low | The input stops fitting the given categories |
|
|
27
|
+
| {{TIER_DONE_ROLE}} | cheap model | low | The definition of done requires interpretation or unavailable tools |
|
|
28
|
+
| {{TIER_READER_ROLE}} | cheap model | low | The requested result needs judgment across sources |
|
|
66
29
|
|
|
67
|
-
|
|
68
|
-
the challenge lane rather than to a second read by the same family. Stakes are
|
|
69
|
-
not a synonym for difficulty: a one-line change to an auth check is simple and
|
|
70
|
-
high-stakes at the same time, and it is the stakes that decide the route.
|
|
30
|
+
When the vendor uses different effort names, choose its equivalent after reading its current capabilities. Before each build, probe the live model roster, compare the configured pin with the lane's default, and record the selected model and effort with a reason. A pin below the current default calls for review; a newer model still needs to fit the job.
|
|
71
31
|
|
|
72
|
-
|
|
73
|
-
bought with a named reason: a reproduced failure, a checkpoint that came back
|
|
74
|
-
unresolved, an irreversible change. A task that merely feels hard is a deep-tier
|
|
75
|
-
task, not an escalation.
|
|
32
|
+
## Role, complexity and stakes
|
|
76
33
|
|
|
77
|
-
|
|
34
|
+
- **Role:** choose the agent whose tools and task match the work.
|
|
35
|
+
- **Complexity:** increase reasoning for ambiguity, interacting systems or an unknown cause; lower it for bounded retrieval and mechanical work.
|
|
36
|
+
- **Stakes:** choose the reviewer and verification needed for the consequence of a mistake.
|
|
37
|
+
- **Reach:** choose a lane that can read the required sources, run the relevant checks and hold enough context.
|
|
38
|
+
- **Capacity:** account for context headroom, concurrency and the lane's current usage limits.
|
|
78
39
|
|
|
79
|
-
|
|
40
|
+
When security, personal data, deletion, bulk mutation or irreversible actions are involved, name the boundary and test it. Use a different model family for the build's single audit pass, with a companion reviewer asking scope versus ask in the same step. Reserve authorization for actions outside the user's existing mandate.
|
|
80
41
|
|
|
81
|
-
|
|
42
|
+
## Effort through the lane runner
|
|
82
43
|
|
|
83
|
-
|
|
44
|
+
When using `aunx cli-run --effort auto`, treat its result as a prompt-size or audit-scope heuristic. It resolves to medium or high, and an audit lane may impose its own effort floor; check the lane's configuration. For a build, select high explicitly; for security-critical or irreversible work, select xhigh explicitly where the lane supports it.
|
|
84
45
|
|
|
85
|
-
|
|
46
|
+
When a lane lacks an effort flag, choose its model and task scope directly. The runner reports an unsupported effort request as a usage error.
|
|
86
47
|
|
|
87
|
-
##
|
|
48
|
+
## Reassess a route
|
|
88
49
|
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
50
|
+
- When a check fails, diagnose its failure before another attempt and state any route change.
|
|
51
|
+
- When the lane lacks required access, hand that part to a lane with authorized reach and continue independent work.
|
|
52
|
+
- When a vendor deprecates a model, verify the replacement from the current roster and update the affected configuration.
|
|
53
|
+
- When changing a model would change cost, privacy or authority beyond the approved scope, present the choice to the user.
|
|
54
|
+
- When the strongest eligible route still cannot resolve the task, return the evidence, partial result and needed decision.
|
|
@@ -4,7 +4,7 @@ Companion tools: not AIs, but things the AIs call. Each subfolder is written onl
|
|
|
4
4
|
|
|
5
5
|
| Folder | Written | Contents |
|
|
6
6
|
|---|---|---|
|
|
7
|
-
| `codecalc/` | when codecalc is selected (
|
|
7
|
+
| `codecalc/` | when codecalc is selected (optional, default no) | `CODECALC.md` (install, per-client registration, the skill) and `mcp/` snippets for the agents its own `setup --write` does not cover |
|
|
8
8
|
| `obsidian-tc/` | when obsidian-tc is selected (optional, default no; needs an Obsidian vault, Node 24+, Ollama or a cloud embeddings key) | `OBSIDIAN-TC.md` (what you need first, install, per-agent registration, security posture) and `mcp/` snippets |
|
|
9
9
|
| `context7/` | when context7 is selected (optional, default no; needs a network call, a Node 18+ local alternative, an optional API key) | `CONTEXT7.md` (what you need first, install, per-agent registration, security posture) and `mcp/` snippets |
|
|
10
10
|
|
|
@@ -9,13 +9,13 @@ What it gives every agent in this folder: exact arithmetic (`evaluate_expression
|
|
|
9
9
|
Needs `uv` (https://docs.astral.sh/uv/) and Python 3.10+.
|
|
10
10
|
|
|
11
11
|
```bash
|
|
12
|
-
uvx 'codecalc[full]' setup # prints what it would do, changes nothing
|
|
13
|
-
uvx 'codecalc[full]' setup --write # merges the codecalc entry into your client's config, copies the skill
|
|
12
|
+
uvx 'codecalc[full]=={{CODECALC_PIN}}' setup # prints what it would do, changes nothing
|
|
13
|
+
uvx 'codecalc[full]=={{CODECALC_PIN}}' setup --write # merges the codecalc entry into your client's config, copies the skill
|
|
14
14
|
```
|
|
15
15
|
|
|
16
|
-
`setup` detects Claude Desktop, Claude Code, Cursor, VS Code and Zed (`--client=NAME` if several), runs two real canaries and ends in one verdict: `ready` / `degraded` / `not-ready`. It backs up the client config it touches. `uvx 'codecalc[full]' doctor` prints a config block with your absolute paths.
|
|
16
|
+
`setup` detects Claude Desktop, Claude Code, Cursor, VS Code and Zed (`--client=NAME` if several), runs two real canaries and ends in one verdict: `ready` / `degraded` / `not-ready`. It backs up the client config it touches. `uvx 'codecalc[full]=={{CODECALC_PIN}}' doctor` prints a config block with your absolute paths.
|
|
17
17
|
|
|
18
|
-
|
|
18
|
+
These commands and the shipped launch snippets use the catalog version this installer was released with. Review a newer release before changing that pin. `[full]` is the edition that actually runs everything documented (about 120 MB). Base `codecalc` is execution only; symbolic tools then return a `dependency_missing` error naming the extra, never a silent failure.
|
|
19
19
|
|
|
20
20
|
## Register more agents with `mcp/` snippets
|
|
21
21
|
|
|
@@ -35,19 +35,15 @@ Two ways to connect, remote first (Context7's own documented default, and the on
|
|
|
35
35
|
# (mcp.context7.com/mcp/oauth), as an alternative to an API key, not required either.
|
|
36
36
|
|
|
37
37
|
# Local alternative: runs the MCP server on your machine over stdio.
|
|
38
|
-
npx -y @upstash/context7-mcp
|
|
39
|
-
|
|
40
|
-
# Or the one-command setup Context7 itself ships, which authenticates via OAuth,
|
|
41
|
-
# writes an API key, and can install a CLI-based skill instead of MCP:
|
|
42
|
-
npx ctx7 setup
|
|
38
|
+
npx -y @upstash/context7-mcp@{{CONTEXT7_PIN}}
|
|
43
39
|
```
|
|
44
40
|
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
`npx ctx7 setup` is upstream's own guided installer; it is not run by this installer, only documented here, the same way this project never runs a vendor script for you.
|
|
41
|
+
The local command and shipped launch snippets use the catalog version this installer was released against. Review a newer release before changing that pin. The remote endpoint is operated by Upstash and cannot be version-pinned by this package. This installer runs neither connection nor a third-party setup command for you.
|
|
48
42
|
|
|
49
43
|
## Register it with your agent (snippets in `mcp/`)
|
|
50
44
|
|
|
45
|
+
When Context7 is selected and project activation is enabled, model-orchestrator merges its entry into a supported project MCP configuration, including Claude Code's `.mcp.json`. Existing conflicting entries are kept for review. With activation disabled or a global client configuration, merge the relevant snippet manually.
|
|
46
|
+
|
|
51
47
|
Every snippet below ships **keyless**: the remote ones point at the hosted endpoint with no `Authorization` header at all (Qwen Code's snippet keeps the non-credential `Accept` header upstream itself ships), and both Zed snippets run the local, version-pinned `npx` server with no key in its `env` block. That is deliberate, not an oversight: see "Higher rate limits" next for why a header is not shipped by default.
|
|
52
48
|
|
|
53
49
|
| Agent | File to edit | Snippet |
|
|
@@ -87,12 +83,12 @@ Anonymous works. If you hit the rate limit and want a key, add it the correct wa
|
|
|
87
83
|
- **Codex CLI**: under the `[mcp_servers.context7]` table in `~/.codex/config.toml`, add a `bearer_token_env_var` entry naming the environment variable `CONTEXT7_API_KEY`. Codex reads the token from that variable at connect time and sends it as the `Authorization` header itself; the config file never holds the value ([Codex MCP docs](https://developers.openai.com/codex/mcp)).
|
|
88
84
|
- **Claude Code** (`mcp/context7.claude-code.mcp.json`): add a `headers` object to the `context7` entry with an `Authorization` field whose value is `Bearer` followed by a `${CONTEXT7_API_KEY}` reference. Claude Code expands `${VAR}` references in a remote server's `headers` at load time, and `CONTEXT7_API_KEY` is not one of the credential names it deliberately reads as empty (those are Claude/Anthropic-specific). Set the variable in your environment before launching; an unset variable still loads with the literal, unexpanded reference sent as the header, and Context7 answers every call with "Invalid API key" instead of running anonymously ([Claude Code MCP docs](https://code.claude.com/docs/en/mcp)). That failure, not a missing feature, is why this snippet ships with no header at all.
|
|
89
85
|
- **Claude Desktop**: the Connectors UI has its own key field; use it there rather than editing a file.
|
|
90
|
-
- **Any local `npx` connection** (Zed, or the local alternative for any other client): export `CONTEXT7_API_KEY` in the shell that launches your editor or agent. A spawned stdio child process inherits its parent's environment by default, so `npx -y @upstash/context7-mcp` picks it up with no config edit; this is the same environment variable name Context7's own Docker MCP Toolkit config and its GitHub Copilot integration use to feed the server a key. If your client does not pass its environment through to the child (uncommon), either stay anonymous, or check whether that client's own config format has an `env` block that itself supports an environment-variable reference (Claude Code's does, described above; not every client's does) rather than typing the key in.
|
|
86
|
+
- **Any local `npx` connection** (Zed, or the local alternative for any other client): export `CONTEXT7_API_KEY` in the shell that launches your editor or agent. A spawned stdio child process inherits its parent's environment by default, so `npx -y @upstash/context7-mcp@{{CONTEXT7_PIN}}` picks it up with no config edit; this is the same environment variable name Context7's own Docker MCP Toolkit config and its GitHub Copilot integration use to feed the server a key. If your client does not pass its environment through to the child (uncommon), either stay anonymous, or check whether that client's own config format has an `env` block that itself supports an environment-variable reference (Claude Code's does, described above; not every client's does) rather than typing the key in.
|
|
91
87
|
- **Cursor, VS Code, Qwen Code, Antigravity `agy`, or any other client using the `mcpServers.json`/`vscode.mcp.json`/`qwen.settings.json`/`agy.mcp_config.json` snippet**: check that client's own docs for whether it expands an environment-variable reference inside a remote server's `headers` before adding one. This is not confirmed for any of them here. If it does not expand, the literal, unexpanded text becomes the header value and every call fails with "Invalid API key" instead of running anonymously, which is worse than shipping no header at all.
|
|
92
88
|
|
|
93
89
|
## Security posture, read before you send anything through it
|
|
94
90
|
|
|
95
|
-
|
|
91
|
+
Library names and query text are sent to Context7's API. If an agent includes source code, credentials or private context in a query, that content leaves the machine too. Review what your agent sends and use public library names and non-sensitive queries. The local stdio server still calls the remote service; it does not make lookups private or offline. The docs it indexes are community-contributed, not vetted by this installer; report suspicious results upstream. Keep API keys in your environment or your client's supported credential store, never in committed files.
|
|
96
92
|
|
|
97
93
|
## Level 3
|
|
98
94
|
|
|
@@ -12,7 +12,7 @@ A durable, searchable, governed store that the protocols can call by name:
|
|
|
12
12
|
|---|---|
|
|
13
13
|
| find what exists before writing (deep research dedupe, gap analysis) | `semantic_search`, `search_text`, `search_regex` |
|
|
14
14
|
| map everything a rename touches (propagate) | `get_backlinks`, `find_unresolved_links`, `rewrite_link` |
|
|
15
|
-
| record the end-to-end doc (build
|
|
15
|
+
| record the end-to-end doc (build Record step) | `write_note` (compare-and-swap, confirmation on overwrite), `patch_note`, `append_note` |
|
|
16
16
|
| keep inferred content honest | `write_note` with `provenance: "agent_synthesis"` runs a poison scan before the write lands |
|
|
17
17
|
| keep a shared vault safe for several agents | JWT scopes, per-vault folder ACLs, a read-only kill switch, human-in-the-loop tokens |
|
|
18
18
|
|
|
@@ -29,7 +29,7 @@ A durable, searchable, governed store that the protocols can call by name:
|
|
|
29
29
|
## Install
|
|
30
30
|
|
|
31
31
|
```bash
|
|
32
|
-
npm install -g obsidian-tc@{{OBSIDIAN_TC_PIN}} #
|
|
32
|
+
npm install -g obsidian-tc@{{OBSIDIAN_TC_PIN}} # reviewed catalog version; review an upgrade before changing the pin
|
|
33
33
|
ollama pull nomic-embed-text
|
|
34
34
|
obsidian-tc /path/to/your/vault # zero-config: one vault named "main", local only
|
|
35
35
|
obsidian-tc plugin install --vault /path/to/your/vault # optional companion plugin, then enable it in Obsidian
|
|
@@ -68,7 +68,7 @@ Every snippet here spawns `npx`, and on Windows `npx` is a batch file (`npx.cmd`
|
|
|
68
68
|
|
|
69
69
|
The SDK point covers more than these two clients: every published `@modelcontextprotocol/sdk` from 1.23.0 through 1.30.0 depends on `cross-spawn ^7.0.5` and uses it in the stdio client, so any client that connects through the stock TypeScript SDK inherits the same `PATHEXT` resolution. `shell: false` in that transport is not the whole story, and reading only that line is how a client gets mistaken for one that cannot start `npx`.
|
|
70
70
|
|
|
71
|
-
One Windows case can still fail, and it is not about `.cmd`. Zed prefers PowerShell for the system shell (`get_windows_system_shell` in `crates/gpui_util/src/lib.rs` falls back to `cmd.exe` only when PowerShell is missing), and PowerShell resolves a bare `npx` to npm's `npx.ps1` shim when one is installed. Under the `Restricted` execution policy that is Windows' client default, running a `.ps1` is blocked. If Zed reports that the server would not start, check `Get-ExecutionPolicy` first, and if that is the cause, change the Zed entry by hand to `"command": "cmd"` with `"args": ["/d", "/c", "npx", "-y", "obsidian-tc"]`, keeping the rest of the block. `/d` is there on purpose: it skips any Command Processor `AutoRun` command, which would otherwise run first and can print non-JSON into the protocol stream.
|
|
71
|
+
One Windows case can still fail, and it is not about `.cmd`. Zed prefers PowerShell for the system shell (`get_windows_system_shell` in `crates/gpui_util/src/lib.rs` falls back to `cmd.exe` only when PowerShell is missing), and PowerShell resolves a bare `npx` to npm's `npx.ps1` shim when one is installed. Under the `Restricted` execution policy that is Windows' client default, running a `.ps1` is blocked. If Zed reports that the server would not start, check `Get-ExecutionPolicy` first, and if that is the cause, change the Zed entry by hand to `"command": "cmd"` with `"args": ["/d", "/c", "npx", "-y", "obsidian-tc@{{OBSIDIAN_TC_PIN}}"]`, keeping the rest of the block. `/d` is there on purpose: it skips any Command Processor `AutoRun` command, which would otherwise run first and can print non-JSON into the protocol stream.
|
|
72
72
|
|
|
73
73
|
None of this was run on a Windows machine by this project. The five verdicts are from each client's own shipped code; the PowerShell case is from Zed's shell choice plus documented `Restricted` behaviour, and is the one worth reporting back if you hit it.
|
|
74
74
|
|
package/docs/audit-brief.md
DELETED
|
@@ -1,148 +0,0 @@
|
|
|
1
|
-
# AUDIT_BRIEF.md: adversarial audit of model-orchestrator (round 1)
|
|
2
|
-
|
|
3
|
-
> Historical document. Counts in it (tests, files) are as of the round they describe; `npm test` prints the current number.
|
|
4
|
-
|
|
5
|
-
Read-only audit. Report findings only; do not modify files. Rank by severity. For each finding: file, line, what breaks, a concrete reproduction. `CLEAN` is a valid answer for any area with no reproducible finding. Skip style.
|
|
6
|
-
|
|
7
|
-
## Task bundle
|
|
8
|
-
|
|
9
|
-
**Purpose.** Adversarial, read-only audit of this npm package before it is published, so that reproducible defects are fixed before strangers run it.
|
|
10
|
-
**Task class.** read_only
|
|
11
|
-
|
|
12
|
-
**Granted scope.**
|
|
13
|
-
- Every file under this repository root: `bin/`, `src/`, `templates/`, `test/`, `scripts/`, `docs/`, `package.json`, `README.md`.
|
|
14
|
-
- Anything outside this repository is out of scope. Do not widen it on your own judgment.
|
|
15
|
-
|
|
16
|
-
**Capabilities.** read files, run `npm test`, run `node bin/cli.js` with `--dry`, run `node bin/cli-run.mjs` with bad arguments, run scripts against a temp directory under /tmp.
|
|
17
|
-
|
|
18
|
-
**Denied actions.** Do not modify, create or delete any file in this repository. Do not run `npm install -g`. Do not run any vendor installer. Do not commit, push, or publish. Do not read files outside this repository except /tmp scratch you created. Do not call any network service.
|
|
19
|
-
- Anything absent from Capabilities is denied. Absence is not permission.
|
|
20
|
-
|
|
21
|
-
**Conventions you do not have.** Report in plain prose with a findings list. No style nitpicks. `CLEAN` is a valid verdict per area. Every finding needs a concrete reproduction (command + observed vs expected). Never print a value that looks like a credential.
|
|
22
|
-
|
|
23
|
-
**Report contract.** Return: a severity-ranked list of findings (file, line, what breaks, reproduction, suggested fix in one or two sentences), then a `CLEAN` line for each area in "Attack these" that had no reproducible finding, then a short "not covered" list naming anything you did not check or could not verify.
|
|
24
|
-
|
|
25
|
-
**Exit parameters.** Stop after 12 minutes of wall clock or after reading every file once and running at most 30 commands, whichever comes first. If you hit a bound, report what you have and name what you did not cover. Never return nothing.
|
|
26
|
-
|
|
27
|
-
## What this is
|
|
28
|
-
An npm package (`npx model-orchestrator`) that asks a user which level (1/2/3) and which AIs they have access to, then writes markdown + config templates into a folder, and optionally runs `npm install -g <pkg>` for known packages after an explicit per-package yes. It also ships `bin/cli-run.mjs`, a wrapper that runs one of five agent CLIs (grok, codex, agy, hermes, qwen) and exits non-zero unless the lane produced a deliverable.
|
|
29
|
-
|
|
30
|
-
## Runtime
|
|
31
|
-
Node >= 18, ESM, zero dependencies. Runs on a stranger's laptop (macOS/Linux) with their PATH and HOME. Level 3 writes shell/systemd/compose templates the user will run on a Linux box.
|
|
32
|
-
|
|
33
|
-
## Threat model
|
|
34
|
-
- The user is not an adversary but is careless: runs it in the wrong directory, passes odd flags, has files with the same names.
|
|
35
|
-
- Untrusted input reaches `cli-run.mjs` through CLI stdout (JSON from third-party binaries) and through `lanes.json` on disk.
|
|
36
|
-
- The installer must never: write a secret value anywhere, overwrite a user file without --force, run a remote shell script, escape the target dir (path traversal via template rel paths or --dir), or leave a placeholder unrendered.
|
|
37
|
-
- `cli-run.mjs` must never: throw on malformed CLI output (a throw is misreported as a usage error), pass a secret in argv, leave temp files, hang on stdin, or report success without a deliverable.
|
|
38
|
-
- Generated templates (`vm/setup-vm.sh`, `vm/jobs/weekly-audit.sh`, `docker-compose.yml`, `gateway.config.yaml`) must not put a key in argv, bind to 0.0.0.0, or pipe a remote script into bash.
|
|
39
|
-
|
|
40
|
-
## Already verified (do not repeat)
|
|
41
|
-
- 60 node --test cases pass, including every judge's failure shapes and a mutation check that turns one case red.
|
|
42
|
-
- Placeholders: every template renders for every level and primary without a leftover `{{KEY}}`.
|
|
43
|
-
- README inside `.claude/agents/` is not installed.
|
|
44
|
-
|
|
45
|
-
## Attack these
|
|
46
|
-
1. `bin/cli.js` argument parsing: `opt()` takes the next argv token; what happens with `--dir --force`, `--ais ""`, duplicate flags, `--level 2.5`, unicode, a `--dir` that is a file, a `--dir` of `/`?
|
|
47
|
-
2. `src/install.js` `writeFiles`: path traversal if a template rel path or the --dir resolves outside; symlink in the target dir; mode handling on Windows; partial writes.
|
|
48
|
-
3. `src/detect.js` `which`: PATH entries that are files, empty PATH, relative PATH entries, a directory named like the binary.
|
|
49
|
-
4. `bin/cli-run.mjs`: `jsonLines` on huge output; `maxBuffer`; `spawnSync` with `timeout` and `killSignal` behaviour; `enabledLanes()` with a malicious lanes.json; `--brief` pointing at a directory or a huge file; prompt containing newlines; the codex `-o` temp file when the CLI writes elsewhere; the `import.meta.url === pathToFileURL(argv[1])` main guard when invoked via a symlink; rc pass-through logic (`if (r.status !== 0 && code === OK) code = r.status`).
|
|
50
|
-
5. Templates: `vm/setup-vm.sh` (set -euo pipefail, the for loop over {{NPM_PACKAGES}} when empty), `vm/jobs/weekly-audit.sh` (curl --config - header injection if GATEWAY_MASTER_KEY contains a quote or newline), `docker-compose.yml` env pass-through, systemd unit paths.
|
|
51
|
-
6. Anything that could make the installer write outside `--dir` or read a file it should not.
|
|
52
|
-
|
|
53
|
-
## Design decisions to challenge, with reasoning
|
|
54
|
-
- Zero dependencies (no inquirer): smaller audit surface, but the prompt code is hand-rolled. Is the readline path safe with piped stdin and EOF?
|
|
55
|
-
- Vendor scripts are printed, never run: correct? Or does printing `curl | bash` still encourage the unsafe pattern?
|
|
56
|
-
- `npm install -g` is run after a per-package yes, with the package name from the catalog (never user input). Confirm user input cannot reach that argv.
|
|
57
|
-
- `writeFiles` refuses existing files unless --force but does not check that the target is inside cwd. Deliberate (users may want `~/project`). Is there a traversal risk from template names?
|
|
58
|
-
|
|
59
|
-
## How to run
|
|
60
|
-
`npm test` · `node bin/cli.js --help` · `node bin/cli.js --yes --level 3 --ais claude-code,codex,agy,grok,hermes,qwen,ollama --dir /tmp/x --dry`
|
|
61
|
-
|
|
62
|
-
## ROUND 2 (after round-1 fixes)
|
|
63
|
-
|
|
64
|
-
Re-audit the same scope. Every round-1 finding was reproduced before it was touched. What changed:
|
|
65
|
-
|
|
66
|
-
| # | Finding | Change |
|
|
67
|
-
|---|---|---|
|
|
68
|
-
| 1 | writeFiles escape / symlink follow | `preflight()` in `src/install.js`: containment under the resolved root, lstat every existing component (symlink or non-directory parent refused), exclusive `wx` create unless `--force`, rollback of files this run created if a later write fails. Tests: escape, symlinked component, conflicting parent leaves nothing behind. Mutation-checked. |
|
|
69
|
-
| 2 | prompt in argv + logged head | argv kept (each vendor's documented headless shape); DEFERRED as a vendor constraint, documented in `CLI-RUN.md` and the file header (no secrets in prompts, ARG_MAX, reference big briefs by path). Log now stores a 12-hex sha256 prefix and length, never text. Test: marker absent from log. |
|
|
70
|
-
| 3 | unknown flags / missing values | strict `parseArgs` in `bin/cli.js`: unknown flag, missing value, empty value, duplicate, positional → exit 2 before planning. Tests for each. |
|
|
71
|
-
| 4 | audit job hardcoded hermes | `auditLane()` picks the first ENABLED cli-run lane (hermes, qwen, codex, agy, grok); none → rendered guard exits 13. Tests. |
|
|
72
|
-
| 5 | audit never received live state | script composes `reports/audit-brief-<date>.md` = task bundle + protocol + DELEGATION_MATRIX + live-state, and passes THAT as `--brief`. Test. |
|
|
73
|
-
| 6 | `--dir` ignored by systemd paths | `INSTALL_DIR` rendered into the service and the script from the resolved `--dir`. Test. |
|
|
74
|
-
| 7 | curl config injection | script refuses a key not matching `^[A-Za-z0-9._-]+$` (exit 2) before any curl; documented in ENVIRONMENT.md and vm/README. Test executes the rendered script with an injecting key. |
|
|
75
|
-
| 8 | malformed lanes.json fail-open | `enabledLanes()` returns null on present-but-invalid; main refuses every lane (13) and says so. Test proves no spawn happens. |
|
|
76
|
-
| 9 | signal → exit 0 | `r.signal || r.status === null` → verdict killed, exit 10, partial output discarded. Test. |
|
|
77
|
-
| 10 | partial install | covered by preflight + rollback (finding 1). Test. |
|
|
78
|
-
| 11 | directory detected as binary | `isFile()` check in both `which()` implementations. Test. |
|
|
79
|
-
| 12 | `curl \| bash` printed | download / read / run form printed instead. |
|
|
80
|
-
|
|
81
|
-
Also new since round 1: the companion-tool path (`--tools codecalc`, `--no-tools`, `templates/tools/codecalc/`, `protocols/numbers-and-logic.md`, `resolveTools`). Attack it the same way: unknown tool ids, interaction with `--yes`, the extra interactive question, and whether any written snippet could be confused for a file the installer should not touch.
|
|
82
|
-
|
|
83
|
-
Report only what reproduces on the current tree. `CLEAN` per area is expected where the fix holds.
|
|
84
|
-
|
|
85
|
-
## New in 0.1.15: two claude-code-only hooks
|
|
86
|
-
|
|
87
|
-
`route-gate.mjs` (`UserPromptSubmit`) and `subagent-context.mjs` (`SubagentStart`) ship to `.claude/hooks/` only when claude-code is the primary. Both are plain Node, zero deps, and installed with mode `0o755`.
|
|
88
|
-
|
|
89
|
-
- **Reads.** `route-gate.mjs` reads at most 64 KB from one file: the routing rules file (`ROUTING.md` or `ORCHESTRATOR.md`) at a path rendered in at install time relative to `CLAUDE_PROJECT_DIR`, never a hardcoded absolute path. Before opening it, it `statSync`s the resolved path (following a symlink to its target) and refuses anything that is not `isFile()`, a FIFO, socket, device or directory included, so the read never touches a path that could block on open. The read itself is one `openSync` + one bounded `readSync` into a fixed 64 KB buffer, closed in a `finally`, so neither the time nor the memory this hook uses depends on how large the file on disk actually is. It then extracts the text between `<!-- route-gate:start -->` and `<!-- route-gate:end -->` and nothing else. `subagent-context.mjs` reads nothing from disk; its context is static text plus the same two rendered paths. Neither parses or executes anything it reads; the extracted block is passed through as a string.
|
|
90
|
-
- **Stdin.** Neither hook uses a field from the JSON input Claude Code sends on stdin, but both must still consume the pipe rather than ignore it. Both drain stdin asynchronously against a 250ms hard cap: whichever comes first, the real `end` event or the timeout, the hook proceeds. Neither ever calls a blocking, synchronous read of stdin.
|
|
91
|
-
- **Writes.** Neither writes a file. Both write one JSON object to stdout: `{"hookSpecificOutput":{"hookEventName":"...","additionalContext":"..."}}`, and both exit only after that write's callback fires, so a buffered write to a pipe is not truncated by an exit racing ahead of it.
|
|
92
|
-
- **Fail-open, on purpose.** A missing `CLAUDE_PROJECT_DIR`, a missing rules file, a non-regular file at the rules path, or a missing block each produce a one-line fallback `additionalContext` naming what was found, and the script still exits 0. This is acceptable because a miss here is a stray context string reaching the model, not a security gate: nothing downstream trusts the hook's output for anything but a routing suggestion, and the settings snippet that wires it in is a document the user merges by hand, never written automatically over an existing `settings.json`.
|
|
93
|
-
- **Bounded.** `route-gate.mjs` caps the read at 64 KB regardless of the file's reported size and the injected string at 4000 characters, so a rules file bloated by a bad edit, or truncated to an arbitrary length, cannot balloon the context or the read time on every turn. Both hooks cap stdin drain at 250ms.
|
|
94
|
-
- **Not yet attacked.** Untested here: a rules file with a `route-gate:start` marker but no matching end marker very far into the file (bounded by `MAX_READ`, so the end marker past that point is treated as absent, which is the intended fail-open path, but worth a deliberate case); a `CLAUDE_PROJECT_DIR` pointing at a path with no read permission; behavior under the Windows exec-form `node` + `args` invocation named in the settings snippet; a `statSync` that itself hangs (a stalled network filesystem, for instance) rather than the FIFO-at-open case this round fixed.
|
|
95
|
-
|
|
96
|
-
### Round 1 (pre-release), fixed before shipping
|
|
97
|
-
|
|
98
|
-
Three findings reproduced against the 0.1.15 branch before it shipped, none of them ever released:
|
|
99
|
-
|
|
100
|
-
| # | Finding | Fix |
|
|
101
|
-
|---|---|---|
|
|
102
|
-
| 1 | HIGH. `readFileSync(0)` in both hooks blocked until stdin reached EOF (`sleep 3 \| ... node route-gate.mjs` still running past 1.5s); `route-gate.mjs` also read the whole rules file into memory before bounding it, so a FIFO planted at the rules path blocked forever on open. | Stdin is drained asynchronously against a 250ms hard cap in both hooks. `route-gate.mjs` refuses anything that is not `isFile()` via `statSync` before ever calling open, then reads through one fixed 64 KB buffer via `openSync`/`readSync`. Tests: an open, never-closed stdin pipe exits within 1s for both hooks; a FIFO at the rules path returns the fallback instead of hanging; a 200 MB sparse rules file completes in well under a second with output still capped. |
|
|
103
|
-
| 2 | MEDIUM. `done-verifier`'s description, both agent-folder READMEs, and the root README called it "read-only" without qualification, while its claude-code file carries an unrestricted `Bash` grant; nothing in that grant stops it from running a mutating command. | Every one of those surfaces now says plainly that `done-verifier` carries no file-editing tools and that its Bash use is bound by its own prompt, not by the tool grant; `reader` is named as the one that is read-only by tool grant (no Bash) on both formats. |
|
|
104
|
-
| 3 | MEDIUM. Three generated surfaces still stated the pre-0.1.15 premise on a claude-code install: `builder.md`'s description ("... or the main build itself"), `build-protocol.md`'s roles table and its "why the builder does not hand off" note, and `ROUTING.md`'s "Plan big, execute small" line ("the orchestrator executes"). | All three now render through `subagentsLoadRules(primary)`, the same gate the decision tree and "Who builds" already used; every other primary is unchanged. A semantic-regression test asserts a claude-code install contains none of the old phrasing and a codex install still does. |
|
|
105
|
-
|
|
106
|
-
## New in 0.1.16: a third hook, route-metrics.mjs
|
|
107
|
-
|
|
108
|
-
`route-metrics.mjs` ships to `.claude/hooks/` alongside `route-gate.mjs` and `subagent-context.mjs`, only when claude-code is the primary. Same shape as the other two: plain Node, zero deps, mode `0o755`. Unlike them, it is wired to five events at once (`UserPromptSubmit`, `PreToolUse` matched to `Agent|Task`, `SubagentStart`, `SubagentStop`, `Stop`), and it does write, deliberately: one JSON line per event, appended to `~/.ai-orchestrator/route-metrics.jsonl`.
|
|
109
|
-
|
|
110
|
-
- **Reads.** Only its own stdin (the JSON Claude Code sends per event) and, for `--summary`, its own log file. It never reads `transcript_path` even though that field is present on every event: the documented source for the route marker is `last_assistant_message`, and the docs say the transcript can lag, so a hook that read it instead could log a stale or absent marker as if it were current. It never reads the rules file, the task bundle, or any other project file.
|
|
111
|
-
- **Writes.** `~/.ai-orchestrator/route-metrics.jsonl` (append-only, rotated to `.jsonl.1` above 5 MB) and `~/.ai-orchestrator/route-metrics.state/<sha256(agent_id)>.json`, a small file recording a subagent's start time and type so `SubagentStop` can compute a duration; it is deleted on stop, and anything older than 24h is pruned on the next `SubagentStart`. Nothing outside `~/.ai-orchestrator/`. It never writes to stdout: on `UserPromptSubmit` and `SubagentStart`, stdout becomes model context, and this hook has nothing to say there, so it stays silent on every event, not just those two.
|
|
112
|
-
- **What it never logs.** Prompt text, tool descriptions, the full `tool_input`, `last_assistant_message` itself, or the "why" half of a route marker. Only six named fields ever reach a record: `session_id`, `subagent_type`, `agent_type`, and the parsed `lane`, each stripped to `[A-Za-z0-9_.+-]` and capped at 64 characters (128 for `session_id`) before being written, plus the event name and a `duration_s` number it computed itself. This mirrors `bin/cli-run.mjs`'s own log, which stores a fixed reason code and never a provider-supplied string.
|
|
113
|
-
- **Fail-open, on purpose.** Every code path that can fail (a malformed state file, a full disk, a rotation race, invalid JSON on stdin, an unrecognized event) is caught and produces no record rather than a thrown error or a non-zero exit; the process always exits 0. A miss here is a missing line in a telemetry log, never a blocked turn, so there is nothing to gate.
|
|
114
|
-
- **Bounded.** Stdin is drained asynchronously against a combined 1s time cap and 8 MB size cap; a payload that exceeds either is treated as truncated and parsed as nothing, never partially. `--summary` reads the log directly (never spawns anything, never executes a line in it).
|
|
115
|
-
- **Not yet attacked.** Untested here: two processes racing the same rotation at once (a rename plus an append landing on the same file); a state directory with thousands of leaked files from a long-lived session with a crashed hook (pruning runs, but only on `SubagentStart`, so an install that never starts a subagent again would never prune); behavior if `agent_id` collides across two concurrent subagents (sha256 makes this astronomically unlikely, not impossible).
|
|
116
|
-
|
|
117
|
-
## New in 0.1.18: the Windows spawn path
|
|
118
|
-
|
|
119
|
-
`bin/cli-run.mjs` runs a lane's binary through `windowsSpawnPlan()` before every `spawn()` call. On POSIX, and for a plain `.exe` or extensionless binary on Windows, this is a no-op: the same argv reaches `spawn()` with no shell, exactly as before. What changed is the two shapes Windows can hand it that used to reach `spawn()` unchanged and throw `EINVAL` (Node's fix for CVE-2024-27980: a `.bat`/`.cmd` target without `shell: true` is refused rather than run through an unsafely-escaped `cmd.exe`).
|
|
120
|
-
|
|
121
|
-
- **What runs, in order.** `resolveCmdShim(cmdPath)` reads the `.cmd` file and looks for the exact line npm's `cmd-shim` package writes: `"%_prog%" ... "<path>" %*`, where `<path>` is `%dp0%`-relative (verified against the real, byte-for-byte output of `cmd-shim@9.0.2`, the package npm itself uses to write a shim from a package.json `bin` entry with a `#!/usr/bin/env node` shebang; `test/judges.test.js` pins that exact fixture). If it matches, the `%dp0%`-relative path is resolved against the `.cmd` file's own directory and checked with `statSync` (must exist, must be a file, must end in `.js`/`.mjs`/`.cjs`); on success, `windowsSpawnPlan()` returns `{ command: process.execPath, args: [scriptPath, ...args] }`, and `spawn()` runs `node <script> <args>` directly. A lane's prompt (argv[1] and on) is text this tool does not control the contents of; this path never puts it anywhere a shell parses it.
|
|
122
|
-
- **Why no shell, ever, for a lane.** Every real lane (grok, codex, agy, hermes, qwen) is an npm-installed Node CLI, so on a real Windows install the resolved-shim branch is the one every run takes. If `resolveCmdShim` returns nothing for a lane (an old cmd-shim layout, a hand-written `.cmd`, or a `.bat`), `windowsSpawnPlan()` returns `{ refuse }` and `cli-run` reports the lane unavailable (exit 13) with a message saying how to fix it. It does not fall back to `cmd.exe`: a batch file re-reads its arguments through `%*` after `cmd.exe` has parsed them once, which is the case CVE-2024-27980 is about, and no escaping fully contains user text through both passes. Removing that path was chosen over guarding it.
|
|
123
|
-
- **The one opt-in `cmd.exe` path, and how its arguments are escaped.** Only a caller passing `{ allowCmdFallback: true }` with arguments it fully controls gets the `cmd.exe` path: today that is the installer's own `npm install -g <pinned spec>` (`npm.cmd` is not a cmd-shim, and every argument comes from the catalog, none from a user). For that caller, `windowsSpawnPlan()` builds one command-line string with `escapeCmdArg`/`buildCmdExeCommand` and returns `{ command: <ComSpec>, args: ['/d', '/s', '/c', <built string>], options: { windowsVerbatimArguments: true } }`. The algorithm is the one documented at [qntm.org/cmd](https://qntm.org/cmd) (the reference writeup of `cmd.exe`'s quoting behavior) and used by the widely-deployed `cross-spawn` package: each argument is quoted the way `CommandLineToArgvW` expects (backslash-doubling before an embedded quote or at the end of the string, then wrapped in `"`), and THEN every `cmd.exe` metacharacter in that quoted text (`( ) % ! ^ " < > & | ; ,` and space) is caret-escaped, because `cmd.exe`'s own line scanner reads those characters off the raw command line before the quoting is honored, quote or no quote. `windowsVerbatimArguments: true` tells Node not to re-quote the string a second, conflicting way. `test/judges.test.js` pins exact expected output for `&`, `|`, `^`, `%`, a literal `"`, a trailing backslash and a literal newline (the last one deliberately unescaped: it is not a `cmd.exe` metacharacter).
|
|
124
|
-
- **`killTree` needs no change for which process it targets, either path.** `taskkill /pid <pid> /T /F` walks the whole descendant tree regardless of whether the direct child is `node` (the resolved-shim path) or `cmd.exe` (the fallback); there is no intermediate shell layer to lose track of in the common case, since there is no shell there at all. It DID need a change for how `taskkill` itself is found: `windows-latest` CI caught a bare `spawn('taskkill', ...)` failing `ENOENT` the first time a lane actually ran end to end there (this project's own test harness deliberately narrows PATH to isolate a fake lane, and that narrowed PATH does not include `System32`; a sandboxed or otherwise stripped-down real environment might not either), and the resulting unheard `error` event on the returned process crashed the whole run over what should be a best-effort cleanup step. `taskkillPath()` resolves the executable under `%SystemRoot%` (falling back through `%windir%` to a fixed path) instead of relying on PATH, built with a literal backslash rather than `node:path`'s `join()`, which picks its separator from the HOST running the code, not the OS the path describes; `killTree` now attaches an `error` listener so any future spawn failure stays a missed cleanup, never a crash.
|
|
125
|
-
- **`bin/cli.js`'s own `npm install -g` prompt reuses this, rather than duplicating it.** The installer's opt-in "run `npm install -g <ai>` now?" prompt had the identical `EINVAL`-shaped defect (`spawnSync('npm', ...)` with no shell), found the same way: it failed the moment its own test actually ran on `windows-latest`. It now resolves `npm` with `which()` and calls `windowsSpawnPlan()`, the same function above, instead of a second copy of the fix.
|
|
126
|
-
- **Not yet attacked for real.** `windowsSpawnPlan`, `resolveCmdShim` and the escaping functions are unit-tested (pure string logic, runs on every CI host) and the resolved-shim path is exercised end to end on `windows-latest` through the fake-lane fixtures in `test/cli.test.js` (installed as a real npm-style `.cmd` shim). A lane can no longer reach `cmd.exe` at all (a unit test pins the refusal). The opt-in `cmd.exe` branch, used only by the installer's own `npm install -g`, is not exercised end to end through a live Windows process in this suite; its escaping is proven by exact-string unit tests only, and its arguments never include user text.
|
|
127
|
-
- **A platform limit found the same way, unrelated to the spawn path itself: Windows has no OS-level signals at all.** `cli-run.mjs`'s graceful shutdown (`process.on('SIGTERM', ...)`, kill the lane's process group, then exit 143/130) is a POSIX guarantee only: `ChildProcess.kill(sig)` on Windows calls `TerminateProcess()` unconditionally for SIGTERM AND SIGINT alike, giving the target process no chance to run any handler at all, proven on `windows-latest` CI (the wrapper died as `{code: null, signal: sig}` for both; a hypothesis that SIGINT gets a real, catchable console-control event on Windows was tried first and measured false in this exact scenario, not assumed). `test/cli.test.js`'s `#13` now expects an unhandled termination for either signal on win32, and the original graceful-exit assertion elsewhere.
|
|
128
|
-
- **Two narrow, individually-verified Windows skips remain, neither in the spawn path itself.** (1) A lane dying mid-run from a real POSIX signal cannot be reproduced on win32: a real Windows lane is a plain `node <script>` process, so it cannot die "by signal" any more than the product being tested can, and the only way a test fixture can even simulate one (a nested `sh -c "...; kill -TERM $$"`) puts an extra node process between cli-run.mjs and the dying shell, so cli-run.mjs observes only that node's translated exit code (measured: MSYS bash's self-kill status leaks through as a plain nonzero exit code, 3840, which this tool already handles honestly via `exit_nonzero`). (2) `weekly-audit.sh`'s watchdog (`bounded()`/`killtree()`, `pgrep -P` plus killing a backgrounded subshell's tree) relies on real bash job control this script only ever runs under on the Ubuntu box it targets; actually executing it against a genuinely hanging stub under Git Bash's job-control emulation hung past a 20s outer timeout on `windows-latest` CI, a known class of MSYS/Cygwin limitation (a `kill -KILL` not reliably reaching the underlying Windows process tree of a backgrounded subshell), not a defect in the generated script, which still renders and syntax-checks correctly.
|
|
129
|
-
|
|
130
|
-
## New in 0.1.23: failure classes in cli-run
|
|
131
|
-
|
|
132
|
-
`bin/cli-run.mjs` now puts every run in a closed failure class that owns the exit code (`auth` 14, `quota` 15, `rejected` 16, `refused` 17, `cut_short` 18, beside `ok` 0, `empty` 10, `no_output` 11, `timeout` 12, `unavailable` 13), counts refused tool calls per lane (including a read of grok's session transcript, whose `sessionId` comes from lane stdout), and prints redacted problem and fix lines. The durable log gains only `class` and `refused`. Threat model: a false exit 0, provider text or a secret reaching the log or the terminal, a transcript read escaping the sessions root, misclassification that sends a user the wrong way, and anything that throws or stalls.
|
|
133
|
-
|
|
134
|
-
### Round 1 (pre-release, GPT-6 Astra at xhigh), fixed before shipping
|
|
135
|
-
|
|
136
|
-
Every finding was reproduced as a failing case in `test/classify.test.js` against the unfixed code (all seven red), then fixed; the suite is green after. One round, by rule; the regression cases verify the fixes.
|
|
137
|
-
|
|
138
|
-
| # | Sev | Finding | Fix | Test |
|
|
139
|
-
|---|---|---|---|---|
|
|
140
|
-
| 1 | HIGH | qwen's display detail clipped an error message at 120 characters before redaction, so a long JSON password lost its closing quote and printed its prefix | redact before every clip: judge detail, API-error text, denial text, deliverable snippet, problem cause | R1 |
|
|
141
|
-
| 2 | HIGH | the JSON credential pattern ended at an escaped quote and missed unterminated values | match JSON string escapes, run an unterminated value to the end, and accept a JSON body escaped inside a string | R2 |
|
|
142
|
-
| 3 | HIGH | `firstDenialText` scanned 16 MiB with an unbounded `[^)]*`: 1 MiB of unclosed `Permission denied for command(` took 15.6 s after the lane had exited | scan at most 256 KiB of stdout and of stderr, with bounded repeats | R3 |
|
|
143
|
-
| 4 | MEDIUM | transcript containment was a string prefix with both separators, so on POSIX a directory literally named `sessions\outside` passed | `isInsideRoot()` via `path.relative`, rejecting `..`, absolute and cross-drive results; tested with `path.posix` and `path.win32` on every OS | R4 |
|
|
144
|
-
| 5 | MEDIUM | qwen signals were searched in the clipped display detail, so a quota error after 120 characters read as `empty`, and a model name in the detail could fake a signal | classify on `qwenErrorText()`: the terminal event's full error and an API-error result only | R5 |
|
|
145
|
-
| 6 | MEDIUM | hermes fell back to stdout when stderr was empty, so an answer saying "no rate limit" read as `quota` | stderr only | R6 |
|
|
146
|
-
| 7 | MEDIUM | a nonzero vendor exit with a recognised empty shape (agy exit 7, `SUCCESS`, empty response) returned `empty`, contradicting the documented `cut_short` fallback | an unexplained nonzero exit is `cut_short`, except hermes' own exit 2 | R7 |
|
|
147
|
-
|
|
148
|
-
Reported clean by the same audit: no false exit 0 across 675 malformed probes and five lanes; no provider text or secret fragment in probed log records; the real codex 0.153.4 fixture's non-fatal error item still succeeds; grok traversal, symlink, FIFO, session-match and read-cap guards; timeout, interrupt, overrun, killed lane, JSON contract, `--quiet` and the doctor canary. Not verified by it: native Windows execution, live vendors, filesystem races.
|
package/scripts/README.md
DELETED
|
@@ -1,7 +0,0 @@
|
|
|
1
|
-
# scripts/
|
|
2
|
-
|
|
3
|
-
| File | Job |
|
|
4
|
-
|---|---|
|
|
5
|
-
| `gen-catalog.js` | regenerates `docs/catalog.md` AND the vendor compatibility table in `README.md` (between the `vendor-table` markers) from `src/catalog.js`; `npm run gen:catalog`. `test/catalog.test.js` fails if either generated surface disagrees with the catalog. |
|
|
6
|
-
| `record-demo.sh` | re-records `docs/demo.gif` by installing the published package into a temp folder and running it under `asciinema`, then rendering the cast with `agg`. Pass a version to pin one: `bash scripts/record-demo.sh 0.1.27`. Needs `brew install asciinema agg`. The frames are the installer's own output, so the GIF stays true to what the command prints. |
|
|
7
|
-
| `gen-plugin.js` | regenerates the Claude Code plugin bundle in `plugin/` (agents, the two read-only hooks, `plugin.json`, `LICENSE`) from `templates/`, using the plan in `src/plugin.js`; `npm run gen:plugin`. `test/plugin.test.js` fails if the committed bundle disagrees. `plugin/README.md` and `plugin/hooks/hooks.json` are hand-owned. |
|