model-orchestrator 0.1.35 → 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +31 -21
- package/CHANGELOG.md +43 -1
- package/README.md +127 -110
- package/bin/README.md +57 -6
- package/bin/aunx.js +7 -0
- package/bin/cli-run.mjs +21 -15
- package/bin/cli.js +376 -257
- package/docs/README.md +15 -18
- package/docs/catalog.md +228 -38
- package/docs/companions.md +28 -10
- package/docs/guarantees.md +21 -12
- package/docs/how-it-routes.md +49 -42
- package/docs/install.md +135 -33
- package/docs/part-1-beginner.md +37 -45
- package/docs/part-2-intermediate.md +34 -52
- package/docs/part-3-advanced.md +36 -26
- package/docs/security-review-history.md +38 -0
- package/llms.txt +24 -25
- package/package.json +15 -8
- package/proof/README.md +100 -0
- package/proof/gate-demo.cast +9 -0
- package/proof/gate-demo.gif +0 -0
- package/proof/results.json +198 -0
- package/proof/scripts/check-gate.js +26 -0
- package/proof/scripts/install-time.js +16 -0
- package/proof/scripts/lib.js +73 -0
- package/proof/scripts/measure.js +15 -0
- package/proof/scripts/missing-results.js +30 -0
- package/proof/scripts/record-gate.js +38 -0
- package/proof/scripts/render.js +18 -0
- package/proof/scripts/runner-overhead.js +21 -0
- package/src/README.md +9 -3
- package/src/activation-ownership.js +19 -0
- package/src/apply-companions.js +104 -0
- package/src/apply-snippets.js +60 -28
- package/src/aunx.js +262 -0
- package/src/catalog.js +253 -117
- package/src/install.js +478 -209
- package/src/plugin.js +13 -4
- package/src/postinstall.js +57 -0
- package/src/roles.js +184 -0
- package/src/uninstall.js +125 -8
- package/templates/README.md +19 -2
- package/templates/advanced/README.md +2 -2
- package/templates/advanced/vm/PRIVACY_GATES.md +17 -19
- package/templates/advanced/vm/README.md +25 -20
- package/templates/advanced/vm/box-CLAUDE.md +19 -18
- package/templates/advanced/vm/jobs/README.md +3 -1
- package/templates/advanced/vm/jobs/weekly-audit.service +3 -0
- package/templates/advanced/vm/jobs/weekly-audit.sh +2 -2
- package/templates/advanced/vm/setup-vm.sh +49 -2
- package/templates/agents/README.md +2 -2
- package/templates/agents/agy/README.md +20 -3
- package/templates/agents/agy/builder.md +11 -7
- package/templates/agents/agy/bulk-worker.md +9 -7
- package/templates/agents/agy/code-reviewer.md +13 -7
- package/templates/agents/agy/deep-planner.md +10 -7
- package/templates/agents/agy/done-verifier.md +13 -22
- package/templates/agents/agy/finding-verifier.md +14 -22
- package/templates/agents/agy/live-researcher.md +10 -7
- package/templates/agents/agy/reader.md +10 -12
- package/templates/agents/claude-code/README.md +18 -14
- package/templates/agents/claude-code/builder.md +10 -15
- package/templates/agents/claude-code/bulk-worker.md +8 -10
- package/templates/agents/claude-code/code-reviewer.md +11 -17
- package/templates/agents/claude-code/deep-planner.md +9 -11
- package/templates/agents/claude-code/done-verifier.md +12 -33
- package/templates/agents/claude-code/finding-verifier.md +13 -39
- package/templates/agents/claude-code/live-researcher.md +9 -11
- package/templates/agents/claude-code/reader.md +9 -18
- package/templates/agents/snippets/chat.md +9 -10
- package/templates/agents/snippets/claude-code.md +17 -18
- package/templates/agents/snippets/generic.md +9 -11
- package/templates/agents/snippets/route-gate.mjs +2 -2
- package/templates/agents/snippets/route-metrics.mjs +1 -1
- package/templates/agents/snippets/subagent-context.mjs +4 -4
- package/templates/beginner/ORCHESTRATOR.md +31 -36
- package/templates/beginner/README.md +1 -1
- package/templates/common/ACCEPTANCE_CHECKS.json +12 -0
- package/templates/common/CONTEXT.md +37 -0
- package/templates/common/DECISIONS.md +11 -0
- package/templates/common/README.md +24 -11
- package/templates/common/TASK_BRIEF.md +84 -0
- package/templates/common/protocols/README.md +14 -11
- package/templates/common/protocols/acceptance-checks.md +14 -0
- package/templates/common/protocols/build-protocol.md +91 -106
- package/templates/common/protocols/context-file.md +10 -0
- package/templates/common/protocols/decision-log.md +9 -0
- package/templates/common/protocols/deep-research.md +20 -34
- package/templates/common/protocols/docs-then-prove.md +13 -18
- package/templates/common/protocols/gap-analysis.md +15 -21
- package/templates/common/protocols/memory-and-record.md +21 -20
- package/templates/common/protocols/numbers-and-logic.md +20 -26
- package/templates/common/protocols/propagate.md +18 -27
- package/templates/intermediate/CLI-RUN.md +83 -113
- package/templates/intermediate/DELEGATION_MATRIX.md +9 -3
- package/templates/intermediate/README.md +3 -3
- package/templates/intermediate/RESEARCH_TRIAGE.md +23 -15
- package/templates/intermediate/ROUTING.md +54 -51
- package/templates/intermediate/TIERS.md +37 -76
- package/templates/tools/README.md +1 -1
- package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +1 -1
- package/docs/audit-brief.md +0 -148
- package/scripts/README.md +0 -7
- package/scripts/gen-catalog.js +0 -81
- package/scripts/gen-plugin.js +0 -16
- package/scripts/record-demo.sh +0 -45
- package/templates/common/TASK_BUNDLE.md +0 -56
|
@@ -1,93 +1,54 @@
|
|
|
1
|
-
# TIERS.md: capability
|
|
1
|
+
# TIERS.md: choose capability, effort and tool reach
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
{{STACK_SUMMARY}}
|
|
4
4
|
|
|
5
|
-
|
|
6
|
-
|---|---|---|
|
|
7
|
-
| deep | ambiguous planning, architecture, strategy, hard debugging | {{PRIMARY_DEEP}} |
|
|
8
|
-
| standard | code writing, code review, execution, live research synthesis | {{PRIMARY_STANDARD}} |
|
|
9
|
-
| fast | classification, extraction, formatting, bulk summarization | {{PRIMARY_FAST}} |
|
|
10
|
-
| escalation | above deep: only when the human asks, or when deep has already run, the call is still unresolved, and the change is irreversible | your vendor's strongest model, if you have one |
|
|
11
|
-
|
|
12
|
-
Non-primary lanes are owned by `DELEGATION_MATRIX.md`.
|
|
13
|
-
|
|
14
|
-
## Effort per agent (the third lever)
|
|
15
|
-
|
|
16
|
-
Tier sets the price per token. Token discipline sets how many tokens. **Effort sets how hard each call thinks.**
|
|
17
|
-
|
|
18
|
-
`cli-run --effort auto` is a heuristic, not a measurement: prompts below 4,000 characters resolve to medium and longer prompts resolve to high. A codex `--audit` always resolves to high because stakes set the floor. Auto never resolves above high. Name `xhigh` explicitly for a security-critical or irreversible audit.
|
|
19
|
-
|
|
20
|
-
| Agent | Tier | Effort | Why |
|
|
21
|
-
|---|---|---|---|
|
|
22
|
-
| deep-planner | deep | xhigh | judges every build twice; expensive to get wrong |
|
|
23
|
-
| code-reviewer | standard | high | every endpoint is internet-facing |
|
|
24
|
-
| finding-verifier | standard | high | judging a claim is harder than producing it |
|
|
25
|
-
| builder | standard | high | a botched deploy is the costly failure |
|
|
26
|
-
| live-researcher | standard | medium | tools do the retrieval |
|
|
27
|
-
| bulk-worker | fast | low | the biggest cost win |
|
|
28
|
-
| done-verifier | fast | low | a done-signal check is a lookup, not a judgment call |
|
|
29
|
-
| reader | fast | low | digestion, not judgment |
|
|
30
|
-
|
|
31
|
-
## Three inputs, not one
|
|
32
|
-
|
|
33
|
-
Role alone does not decide a route. Two more inputs move it, and they move it in
|
|
34
|
-
opposite directions, so state them separately instead of folding them into the
|
|
35
|
-
role.
|
|
5
|
+
When choosing a model, use the job's required capability and the lane's live roster. A tier names a job; a current family alias or explicit model choice implements it.
|
|
36
6
|
|
|
37
|
-
|
|
38
|
-
on every task.
|
|
39
|
-
|
|
40
|
-
| Complexity | What it looks like | What moves |
|
|
7
|
+
| Tier | Use for | On {{PRIMARY_NAME}} |
|
|
41
8
|
|---|---|---|
|
|
42
|
-
|
|
|
43
|
-
|
|
|
44
|
-
|
|
|
45
|
-
| critical | irreversible, or it rewrites a standing rule | the escalation rule below applies |
|
|
9
|
+
| planning model | ambiguous planning, architecture, strategy, unknown causes | {{PRIMARY_DEEP}} |
|
|
10
|
+
| working model | code writing, review, execution and research synthesis | {{PRIMARY_STANDARD}} |
|
|
11
|
+
| cheap model | classification, extraction, formatting and bulk summaries | {{PRIMARY_FAST}} |
|
|
46
12
|
|
|
47
|
-
|
|
48
|
-
reasoning than the reviewer judging its output.** When the plan is airtight the
|
|
49
|
-
spec is carrying the thinking, so builder drops to medium. When the plan is
|
|
50
|
-
vague, fix the plan; do not buy reasoning to paper over it.
|
|
13
|
+
When selecting another AI tool, use `DELEGATION_MATRIX.md` and verify its current availability. When data must remain local, choose a local runtime with the required privacy boundary.
|
|
51
14
|
|
|
52
|
-
|
|
53
|
-
the ones worth naming, because their failures are not recoverable by editing the
|
|
54
|
-
code afterwards.
|
|
15
|
+
## Model and effort per job
|
|
55
16
|
|
|
56
|
-
|
|
57
|
-
lost data, or something you can't undo. Most tasks are low-stakes and route
|
|
58
|
-
normally.
|
|
17
|
+
When routing, consider three levers together: tier sets capability and price; scoped context limits token use; effort sets how much reasoning the call applies.
|
|
59
18
|
|
|
60
|
-
|
|
|
61
|
-
|
|
62
|
-
|
|
|
63
|
-
|
|
|
64
|
-
|
|
|
65
|
-
|
|
|
19
|
+
| Role | Starting tier | Starting effort | Reassess when |
|
|
20
|
+
|---|---|---|---|
|
|
21
|
+
| {{TIER_PLANNER_ROLE}} | planning model | xhigh where supported | The decision can be resolved from a known plan or needs a new capability |
|
|
22
|
+
| {{TIER_REVIEW_ROLE}} | working model | high | Security, privacy or irreversible effects raise the review scope |
|
|
23
|
+
| {{TIER_FINDING_ROLE}} | working model | high | Reproduction needs another runtime or access path |
|
|
24
|
+
| {{TIER_BUILDER_ROLE}} | working model | high | Architecture, security or irreversible work needs xhigh and suitable model capability |
|
|
25
|
+
| {{TIER_LIVE_ROLE}} | working model | medium | Synthesis becomes complex or sources disagree |
|
|
26
|
+
| {{TIER_BULK_ROLE}} | cheap model | low | The input stops fitting the given categories |
|
|
27
|
+
| {{TIER_DONE_ROLE}} | cheap model | low | The definition of done requires interpretation or unavailable tools |
|
|
28
|
+
| {{TIER_READER_ROLE}} | cheap model | low | The requested result needs judgment across sources |
|
|
66
29
|
|
|
67
|
-
|
|
68
|
-
the challenge lane rather than to a second read by the same family. Stakes are
|
|
69
|
-
not a synonym for difficulty: a one-line change to an auth check is simple and
|
|
70
|
-
high-stakes at the same time, and it is the stakes that decide the route.
|
|
30
|
+
When the vendor uses different effort names, choose its equivalent after reading its current capabilities. Before each build, probe the live model roster, compare the configured pin with the lane's default, and record the selected model and effort with a reason. A pin below the current default calls for review; a newer model still needs to fit the job.
|
|
71
31
|
|
|
72
|
-
|
|
73
|
-
bought with a named reason: a reproduced failure, a checkpoint that came back
|
|
74
|
-
unresolved, an irreversible change. A task that merely feels hard is a deep-tier
|
|
75
|
-
task, not an escalation.
|
|
32
|
+
## Role, complexity and stakes
|
|
76
33
|
|
|
77
|
-
|
|
34
|
+
- **Role:** choose the agent whose tools and task match the work.
|
|
35
|
+
- **Complexity:** increase reasoning for ambiguity, interacting systems or an unknown cause; lower it for bounded retrieval and mechanical work.
|
|
36
|
+
- **Stakes:** choose the reviewer and verification needed for the consequence of a mistake.
|
|
37
|
+
- **Reach:** choose a lane that can read the required sources, run the relevant checks and hold enough context.
|
|
38
|
+
- **Capacity:** account for context headroom, concurrency and the lane's current usage limits.
|
|
78
39
|
|
|
79
|
-
|
|
40
|
+
When security, personal data, deletion, bulk mutation or irreversible actions are involved, name the boundary and test it. Use a different model family for the build's single audit pass, with a companion reviewer asking scope versus ask in the same step. Reserve authorization for actions outside the user's existing mandate.
|
|
80
41
|
|
|
81
|
-
|
|
42
|
+
## Effort through the lane runner
|
|
82
43
|
|
|
83
|
-
|
|
44
|
+
When using `aunx cli-run --effort auto`, treat its result as a prompt-size or audit-scope heuristic. It resolves to medium or high, and an audit lane may impose its own effort floor; check the lane's configuration. For a build, select high explicitly; for security-critical or irreversible work, select xhigh explicitly where the lane supports it.
|
|
84
45
|
|
|
85
|
-
|
|
46
|
+
When a lane lacks an effort flag, choose its model and task scope directly. The runner reports an unsupported effort request as a usage error.
|
|
86
47
|
|
|
87
|
-
##
|
|
48
|
+
## Reassess a route
|
|
88
49
|
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
50
|
+
- When a check fails, diagnose its failure before another attempt and state any route change.
|
|
51
|
+
- When the lane lacks required access, hand that part to a lane with authorized reach and continue independent work.
|
|
52
|
+
- When a vendor deprecates a model, verify the replacement from the current roster and update the affected configuration.
|
|
53
|
+
- When changing a model would change cost, privacy or authority beyond the approved scope, present the choice to the user.
|
|
54
|
+
- When the strongest eligible route still cannot resolve the task, return the evidence, partial result and needed decision.
|
|
@@ -4,7 +4,7 @@ Companion tools: not AIs, but things the AIs call. Each subfolder is written onl
|
|
|
4
4
|
|
|
5
5
|
| Folder | Written | Contents |
|
|
6
6
|
|---|---|---|
|
|
7
|
-
| `codecalc/` | when codecalc is selected (
|
|
7
|
+
| `codecalc/` | when codecalc is selected (optional, default no) | `CODECALC.md` (install, per-client registration, the skill) and `mcp/` snippets for the agents its own `setup --write` does not cover |
|
|
8
8
|
| `obsidian-tc/` | when obsidian-tc is selected (optional, default no; needs an Obsidian vault, Node 24+, Ollama or a cloud embeddings key) | `OBSIDIAN-TC.md` (what you need first, install, per-agent registration, security posture) and `mcp/` snippets |
|
|
9
9
|
| `context7/` | when context7 is selected (optional, default no; needs a network call, a Node 18+ local alternative, an optional API key) | `CONTEXT7.md` (what you need first, install, per-agent registration, security posture) and `mcp/` snippets |
|
|
10
10
|
|
|
@@ -12,7 +12,7 @@ A durable, searchable, governed store that the protocols can call by name:
|
|
|
12
12
|
|---|---|
|
|
13
13
|
| find what exists before writing (deep research dedupe, gap analysis) | `semantic_search`, `search_text`, `search_regex` |
|
|
14
14
|
| map everything a rename touches (propagate) | `get_backlinks`, `find_unresolved_links`, `rewrite_link` |
|
|
15
|
-
| record the end-to-end doc (build
|
|
15
|
+
| record the end-to-end doc (build Record step) | `write_note` (compare-and-swap, confirmation on overwrite), `patch_note`, `append_note` |
|
|
16
16
|
| keep inferred content honest | `write_note` with `provenance: "agent_synthesis"` runs a poison scan before the write lands |
|
|
17
17
|
| keep a shared vault safe for several agents | JWT scopes, per-vault folder ACLs, a read-only kill switch, human-in-the-loop tokens |
|
|
18
18
|
|
package/docs/audit-brief.md
DELETED
|
@@ -1,148 +0,0 @@
|
|
|
1
|
-
# AUDIT_BRIEF.md: adversarial audit of model-orchestrator (round 1)
|
|
2
|
-
|
|
3
|
-
> Historical document. Counts in it (tests, files) are as of the round they describe; `npm test` prints the current number.
|
|
4
|
-
|
|
5
|
-
Read-only audit. Report findings only; do not modify files. Rank by severity. For each finding: file, line, what breaks, a concrete reproduction. `CLEAN` is a valid answer for any area with no reproducible finding. Skip style.
|
|
6
|
-
|
|
7
|
-
## Task bundle
|
|
8
|
-
|
|
9
|
-
**Purpose.** Adversarial, read-only audit of this npm package before it is published, so that reproducible defects are fixed before strangers run it.
|
|
10
|
-
**Task class.** read_only
|
|
11
|
-
|
|
12
|
-
**Granted scope.**
|
|
13
|
-
- Every file under this repository root: `bin/`, `src/`, `templates/`, `test/`, `scripts/`, `docs/`, `package.json`, `README.md`.
|
|
14
|
-
- Anything outside this repository is out of scope. Do not widen it on your own judgment.
|
|
15
|
-
|
|
16
|
-
**Capabilities.** read files, run `npm test`, run `node bin/cli.js` with `--dry`, run `node bin/cli-run.mjs` with bad arguments, run scripts against a temp directory under /tmp.
|
|
17
|
-
|
|
18
|
-
**Denied actions.** Do not modify, create or delete any file in this repository. Do not run `npm install -g`. Do not run any vendor installer. Do not commit, push, or publish. Do not read files outside this repository except /tmp scratch you created. Do not call any network service.
|
|
19
|
-
- Anything absent from Capabilities is denied. Absence is not permission.
|
|
20
|
-
|
|
21
|
-
**Conventions you do not have.** Report in plain prose with a findings list. No style nitpicks. `CLEAN` is a valid verdict per area. Every finding needs a concrete reproduction (command + observed vs expected). Never print a value that looks like a credential.
|
|
22
|
-
|
|
23
|
-
**Report contract.** Return: a severity-ranked list of findings (file, line, what breaks, reproduction, suggested fix in one or two sentences), then a `CLEAN` line for each area in "Attack these" that had no reproducible finding, then a short "not covered" list naming anything you did not check or could not verify.
|
|
24
|
-
|
|
25
|
-
**Exit parameters.** Stop after 12 minutes of wall clock or after reading every file once and running at most 30 commands, whichever comes first. If you hit a bound, report what you have and name what you did not cover. Never return nothing.
|
|
26
|
-
|
|
27
|
-
## What this is
|
|
28
|
-
An npm package (`npx model-orchestrator`) that asks a user which level (1/2/3) and which AIs they have access to, then writes markdown + config templates into a folder, and optionally runs `npm install -g <pkg>` for known packages after an explicit per-package yes. It also ships `bin/cli-run.mjs`, a wrapper that runs one of five agent CLIs (grok, codex, agy, hermes, qwen) and exits non-zero unless the lane produced a deliverable.
|
|
29
|
-
|
|
30
|
-
## Runtime
|
|
31
|
-
Node >= 18, ESM, zero dependencies. Runs on a stranger's laptop (macOS/Linux) with their PATH and HOME. Level 3 writes shell/systemd/compose templates the user will run on a Linux box.
|
|
32
|
-
|
|
33
|
-
## Threat model
|
|
34
|
-
- The user is not an adversary but is careless: runs it in the wrong directory, passes odd flags, has files with the same names.
|
|
35
|
-
- Untrusted input reaches `cli-run.mjs` through CLI stdout (JSON from third-party binaries) and through `lanes.json` on disk.
|
|
36
|
-
- The installer must never: write a secret value anywhere, overwrite a user file without --force, run a remote shell script, escape the target dir (path traversal via template rel paths or --dir), or leave a placeholder unrendered.
|
|
37
|
-
- `cli-run.mjs` must never: throw on malformed CLI output (a throw is misreported as a usage error), pass a secret in argv, leave temp files, hang on stdin, or report success without a deliverable.
|
|
38
|
-
- Generated templates (`vm/setup-vm.sh`, `vm/jobs/weekly-audit.sh`, `docker-compose.yml`, `gateway.config.yaml`) must not put a key in argv, bind to 0.0.0.0, or pipe a remote script into bash.
|
|
39
|
-
|
|
40
|
-
## Already verified (do not repeat)
|
|
41
|
-
- 60 node --test cases pass, including every judge's failure shapes and a mutation check that turns one case red.
|
|
42
|
-
- Placeholders: every template renders for every level and primary without a leftover `{{KEY}}`.
|
|
43
|
-
- README inside `.claude/agents/` is not installed.
|
|
44
|
-
|
|
45
|
-
## Attack these
|
|
46
|
-
1. `bin/cli.js` argument parsing: `opt()` takes the next argv token; what happens with `--dir --force`, `--ais ""`, duplicate flags, `--level 2.5`, unicode, a `--dir` that is a file, a `--dir` of `/`?
|
|
47
|
-
2. `src/install.js` `writeFiles`: path traversal if a template rel path or the --dir resolves outside; symlink in the target dir; mode handling on Windows; partial writes.
|
|
48
|
-
3. `src/detect.js` `which`: PATH entries that are files, empty PATH, relative PATH entries, a directory named like the binary.
|
|
49
|
-
4. `bin/cli-run.mjs`: `jsonLines` on huge output; `maxBuffer`; `spawnSync` with `timeout` and `killSignal` behaviour; `enabledLanes()` with a malicious lanes.json; `--brief` pointing at a directory or a huge file; prompt containing newlines; the codex `-o` temp file when the CLI writes elsewhere; the `import.meta.url === pathToFileURL(argv[1])` main guard when invoked via a symlink; rc pass-through logic (`if (r.status !== 0 && code === OK) code = r.status`).
|
|
50
|
-
5. Templates: `vm/setup-vm.sh` (set -euo pipefail, the for loop over {{NPM_PACKAGES}} when empty), `vm/jobs/weekly-audit.sh` (curl --config - header injection if GATEWAY_MASTER_KEY contains a quote or newline), `docker-compose.yml` env pass-through, systemd unit paths.
|
|
51
|
-
6. Anything that could make the installer write outside `--dir` or read a file it should not.
|
|
52
|
-
|
|
53
|
-
## Design decisions to challenge, with reasoning
|
|
54
|
-
- Zero dependencies (no inquirer): smaller audit surface, but the prompt code is hand-rolled. Is the readline path safe with piped stdin and EOF?
|
|
55
|
-
- Vendor scripts are printed, never run: correct? Or does printing `curl | bash` still encourage the unsafe pattern?
|
|
56
|
-
- `npm install -g` is run after a per-package yes, with the package name from the catalog (never user input). Confirm user input cannot reach that argv.
|
|
57
|
-
- `writeFiles` refuses existing files unless --force but does not check that the target is inside cwd. Deliberate (users may want `~/project`). Is there a traversal risk from template names?
|
|
58
|
-
|
|
59
|
-
## How to run
|
|
60
|
-
`npm test` · `node bin/cli.js --help` · `node bin/cli.js --yes --level 3 --ais claude-code,codex,agy,grok,hermes,qwen,ollama --dir /tmp/x --dry`
|
|
61
|
-
|
|
62
|
-
## ROUND 2 (after round-1 fixes)
|
|
63
|
-
|
|
64
|
-
Re-audit the same scope. Every round-1 finding was reproduced before it was touched. What changed:
|
|
65
|
-
|
|
66
|
-
| # | Finding | Change |
|
|
67
|
-
|---|---|---|
|
|
68
|
-
| 1 | writeFiles escape / symlink follow | `preflight()` in `src/install.js`: containment under the resolved root, lstat every existing component (symlink or non-directory parent refused), exclusive `wx` create unless `--force`, rollback of files this run created if a later write fails. Tests: escape, symlinked component, conflicting parent leaves nothing behind. Mutation-checked. |
|
|
69
|
-
| 2 | prompt in argv + logged head | argv kept (each vendor's documented headless shape); DEFERRED as a vendor constraint, documented in `CLI-RUN.md` and the file header (no secrets in prompts, ARG_MAX, reference big briefs by path). Log now stores a 12-hex sha256 prefix and length, never text. Test: marker absent from log. |
|
|
70
|
-
| 3 | unknown flags / missing values | strict `parseArgs` in `bin/cli.js`: unknown flag, missing value, empty value, duplicate, positional → exit 2 before planning. Tests for each. |
|
|
71
|
-
| 4 | audit job hardcoded hermes | `auditLane()` picks the first ENABLED cli-run lane (hermes, qwen, codex, agy, grok); none → rendered guard exits 13. Tests. |
|
|
72
|
-
| 5 | audit never received live state | script composes `reports/audit-brief-<date>.md` = task bundle + protocol + DELEGATION_MATRIX + live-state, and passes THAT as `--brief`. Test. |
|
|
73
|
-
| 6 | `--dir` ignored by systemd paths | `INSTALL_DIR` rendered into the service and the script from the resolved `--dir`. Test. |
|
|
74
|
-
| 7 | curl config injection | script refuses a key not matching `^[A-Za-z0-9._-]+$` (exit 2) before any curl; documented in ENVIRONMENT.md and vm/README. Test executes the rendered script with an injecting key. |
|
|
75
|
-
| 8 | malformed lanes.json fail-open | `enabledLanes()` returns null on present-but-invalid; main refuses every lane (13) and says so. Test proves no spawn happens. |
|
|
76
|
-
| 9 | signal → exit 0 | `r.signal || r.status === null` → verdict killed, exit 10, partial output discarded. Test. |
|
|
77
|
-
| 10 | partial install | covered by preflight + rollback (finding 1). Test. |
|
|
78
|
-
| 11 | directory detected as binary | `isFile()` check in both `which()` implementations. Test. |
|
|
79
|
-
| 12 | `curl \| bash` printed | download / read / run form printed instead. |
|
|
80
|
-
|
|
81
|
-
Also new since round 1: the companion-tool path (`--tools codecalc`, `--no-tools`, `templates/tools/codecalc/`, `protocols/numbers-and-logic.md`, `resolveTools`). Attack it the same way: unknown tool ids, interaction with `--yes`, the extra interactive question, and whether any written snippet could be confused for a file the installer should not touch.
|
|
82
|
-
|
|
83
|
-
Report only what reproduces on the current tree. `CLEAN` per area is expected where the fix holds.
|
|
84
|
-
|
|
85
|
-
## New in 0.1.15: two claude-code-only hooks
|
|
86
|
-
|
|
87
|
-
`route-gate.mjs` (`UserPromptSubmit`) and `subagent-context.mjs` (`SubagentStart`) ship to `.claude/hooks/` only when claude-code is the primary. Both are plain Node, zero deps, and installed with mode `0o755`.
|
|
88
|
-
|
|
89
|
-
- **Reads.** `route-gate.mjs` reads at most 64 KB from one file: the routing rules file (`ROUTING.md` or `ORCHESTRATOR.md`) at a path rendered in at install time relative to `CLAUDE_PROJECT_DIR`, never a hardcoded absolute path. Before opening it, it `statSync`s the resolved path (following a symlink to its target) and refuses anything that is not `isFile()`, a FIFO, socket, device or directory included, so the read never touches a path that could block on open. The read itself is one `openSync` + one bounded `readSync` into a fixed 64 KB buffer, closed in a `finally`, so neither the time nor the memory this hook uses depends on how large the file on disk actually is. It then extracts the text between `<!-- route-gate:start -->` and `<!-- route-gate:end -->` and nothing else. `subagent-context.mjs` reads nothing from disk; its context is static text plus the same two rendered paths. Neither parses or executes anything it reads; the extracted block is passed through as a string.
|
|
90
|
-
- **Stdin.** Neither hook uses a field from the JSON input Claude Code sends on stdin, but both must still consume the pipe rather than ignore it. Both drain stdin asynchronously against a 250ms hard cap: whichever comes first, the real `end` event or the timeout, the hook proceeds. Neither ever calls a blocking, synchronous read of stdin.
|
|
91
|
-
- **Writes.** Neither writes a file. Both write one JSON object to stdout: `{"hookSpecificOutput":{"hookEventName":"...","additionalContext":"..."}}`, and both exit only after that write's callback fires, so a buffered write to a pipe is not truncated by an exit racing ahead of it.
|
|
92
|
-
- **Fail-open, on purpose.** A missing `CLAUDE_PROJECT_DIR`, a missing rules file, a non-regular file at the rules path, or a missing block each produce a one-line fallback `additionalContext` naming what was found, and the script still exits 0. This is acceptable because a miss here is a stray context string reaching the model, not a security gate: nothing downstream trusts the hook's output for anything but a routing suggestion, and the settings snippet that wires it in is a document the user merges by hand, never written automatically over an existing `settings.json`.
|
|
93
|
-
- **Bounded.** `route-gate.mjs` caps the read at 64 KB regardless of the file's reported size and the injected string at 4000 characters, so a rules file bloated by a bad edit, or truncated to an arbitrary length, cannot balloon the context or the read time on every turn. Both hooks cap stdin drain at 250ms.
|
|
94
|
-
- **Not yet attacked.** Untested here: a rules file with a `route-gate:start` marker but no matching end marker very far into the file (bounded by `MAX_READ`, so the end marker past that point is treated as absent, which is the intended fail-open path, but worth a deliberate case); a `CLAUDE_PROJECT_DIR` pointing at a path with no read permission; behavior under the Windows exec-form `node` + `args` invocation named in the settings snippet; a `statSync` that itself hangs (a stalled network filesystem, for instance) rather than the FIFO-at-open case this round fixed.
|
|
95
|
-
|
|
96
|
-
### Round 1 (pre-release), fixed before shipping
|
|
97
|
-
|
|
98
|
-
Three findings reproduced against the 0.1.15 branch before it shipped, none of them ever released:
|
|
99
|
-
|
|
100
|
-
| # | Finding | Fix |
|
|
101
|
-
|---|---|---|
|
|
102
|
-
| 1 | HIGH. `readFileSync(0)` in both hooks blocked until stdin reached EOF (`sleep 3 \| ... node route-gate.mjs` still running past 1.5s); `route-gate.mjs` also read the whole rules file into memory before bounding it, so a FIFO planted at the rules path blocked forever on open. | Stdin is drained asynchronously against a 250ms hard cap in both hooks. `route-gate.mjs` refuses anything that is not `isFile()` via `statSync` before ever calling open, then reads through one fixed 64 KB buffer via `openSync`/`readSync`. Tests: an open, never-closed stdin pipe exits within 1s for both hooks; a FIFO at the rules path returns the fallback instead of hanging; a 200 MB sparse rules file completes in well under a second with output still capped. |
|
|
103
|
-
| 2 | MEDIUM. `done-verifier`'s description, both agent-folder READMEs, and the root README called it "read-only" without qualification, while its claude-code file carries an unrestricted `Bash` grant; nothing in that grant stops it from running a mutating command. | Every one of those surfaces now says plainly that `done-verifier` carries no file-editing tools and that its Bash use is bound by its own prompt, not by the tool grant; `reader` is named as the one that is read-only by tool grant (no Bash) on both formats. |
|
|
104
|
-
| 3 | MEDIUM. Three generated surfaces still stated the pre-0.1.15 premise on a claude-code install: `builder.md`'s description ("... or the main build itself"), `build-protocol.md`'s roles table and its "why the builder does not hand off" note, and `ROUTING.md`'s "Plan big, execute small" line ("the orchestrator executes"). | All three now render through `subagentsLoadRules(primary)`, the same gate the decision tree and "Who builds" already used; every other primary is unchanged. A semantic-regression test asserts a claude-code install contains none of the old phrasing and a codex install still does. |
|
|
105
|
-
|
|
106
|
-
## New in 0.1.16: a third hook, route-metrics.mjs
|
|
107
|
-
|
|
108
|
-
`route-metrics.mjs` ships to `.claude/hooks/` alongside `route-gate.mjs` and `subagent-context.mjs`, only when claude-code is the primary. Same shape as the other two: plain Node, zero deps, mode `0o755`. Unlike them, it is wired to five events at once (`UserPromptSubmit`, `PreToolUse` matched to `Agent|Task`, `SubagentStart`, `SubagentStop`, `Stop`), and it does write, deliberately: one JSON line per event, appended to `~/.ai-orchestrator/route-metrics.jsonl`.
|
|
109
|
-
|
|
110
|
-
- **Reads.** Only its own stdin (the JSON Claude Code sends per event) and, for `--summary`, its own log file. It never reads `transcript_path` even though that field is present on every event: the documented source for the route marker is `last_assistant_message`, and the docs say the transcript can lag, so a hook that read it instead could log a stale or absent marker as if it were current. It never reads the rules file, the task bundle, or any other project file.
|
|
111
|
-
- **Writes.** `~/.ai-orchestrator/route-metrics.jsonl` (append-only, rotated to `.jsonl.1` above 5 MB) and `~/.ai-orchestrator/route-metrics.state/<sha256(agent_id)>.json`, a small file recording a subagent's start time and type so `SubagentStop` can compute a duration; it is deleted on stop, and anything older than 24h is pruned on the next `SubagentStart`. Nothing outside `~/.ai-orchestrator/`. It never writes to stdout: on `UserPromptSubmit` and `SubagentStart`, stdout becomes model context, and this hook has nothing to say there, so it stays silent on every event, not just those two.
|
|
112
|
-
- **What it never logs.** Prompt text, tool descriptions, the full `tool_input`, `last_assistant_message` itself, or the "why" half of a route marker. Only six named fields ever reach a record: `session_id`, `subagent_type`, `agent_type`, and the parsed `lane`, each stripped to `[A-Za-z0-9_.+-]` and capped at 64 characters (128 for `session_id`) before being written, plus the event name and a `duration_s` number it computed itself. This mirrors `bin/cli-run.mjs`'s own log, which stores a fixed reason code and never a provider-supplied string.
|
|
113
|
-
- **Fail-open, on purpose.** Every code path that can fail (a malformed state file, a full disk, a rotation race, invalid JSON on stdin, an unrecognized event) is caught and produces no record rather than a thrown error or a non-zero exit; the process always exits 0. A miss here is a missing line in a telemetry log, never a blocked turn, so there is nothing to gate.
|
|
114
|
-
- **Bounded.** Stdin is drained asynchronously against a combined 1s time cap and 8 MB size cap; a payload that exceeds either is treated as truncated and parsed as nothing, never partially. `--summary` reads the log directly (never spawns anything, never executes a line in it).
|
|
115
|
-
- **Not yet attacked.** Untested here: two processes racing the same rotation at once (a rename plus an append landing on the same file); a state directory with thousands of leaked files from a long-lived session with a crashed hook (pruning runs, but only on `SubagentStart`, so an install that never starts a subagent again would never prune); behavior if `agent_id` collides across two concurrent subagents (sha256 makes this astronomically unlikely, not impossible).
|
|
116
|
-
|
|
117
|
-
## New in 0.1.18: the Windows spawn path
|
|
118
|
-
|
|
119
|
-
`bin/cli-run.mjs` runs a lane's binary through `windowsSpawnPlan()` before every `spawn()` call. On POSIX, and for a plain `.exe` or extensionless binary on Windows, this is a no-op: the same argv reaches `spawn()` with no shell, exactly as before. What changed is the two shapes Windows can hand it that used to reach `spawn()` unchanged and throw `EINVAL` (Node's fix for CVE-2024-27980: a `.bat`/`.cmd` target without `shell: true` is refused rather than run through an unsafely-escaped `cmd.exe`).
|
|
120
|
-
|
|
121
|
-
- **What runs, in order.** `resolveCmdShim(cmdPath)` reads the `.cmd` file and looks for the exact line npm's `cmd-shim` package writes: `"%_prog%" ... "<path>" %*`, where `<path>` is `%dp0%`-relative (verified against the real, byte-for-byte output of `cmd-shim@9.0.2`, the package npm itself uses to write a shim from a package.json `bin` entry with a `#!/usr/bin/env node` shebang; `test/judges.test.js` pins that exact fixture). If it matches, the `%dp0%`-relative path is resolved against the `.cmd` file's own directory and checked with `statSync` (must exist, must be a file, must end in `.js`/`.mjs`/`.cjs`); on success, `windowsSpawnPlan()` returns `{ command: process.execPath, args: [scriptPath, ...args] }`, and `spawn()` runs `node <script> <args>` directly. A lane's prompt (argv[1] and on) is text this tool does not control the contents of; this path never puts it anywhere a shell parses it.
|
|
122
|
-
- **Why no shell, ever, for a lane.** Every real lane (grok, codex, agy, hermes, qwen) is an npm-installed Node CLI, so on a real Windows install the resolved-shim branch is the one every run takes. If `resolveCmdShim` returns nothing for a lane (an old cmd-shim layout, a hand-written `.cmd`, or a `.bat`), `windowsSpawnPlan()` returns `{ refuse }` and `cli-run` reports the lane unavailable (exit 13) with a message saying how to fix it. It does not fall back to `cmd.exe`: a batch file re-reads its arguments through `%*` after `cmd.exe` has parsed them once, which is the case CVE-2024-27980 is about, and no escaping fully contains user text through both passes. Removing that path was chosen over guarding it.
|
|
123
|
-
- **The one opt-in `cmd.exe` path, and how its arguments are escaped.** Only a caller passing `{ allowCmdFallback: true }` with arguments it fully controls gets the `cmd.exe` path: today that is the installer's own `npm install -g <pinned spec>` (`npm.cmd` is not a cmd-shim, and every argument comes from the catalog, none from a user). For that caller, `windowsSpawnPlan()` builds one command-line string with `escapeCmdArg`/`buildCmdExeCommand` and returns `{ command: <ComSpec>, args: ['/d', '/s', '/c', <built string>], options: { windowsVerbatimArguments: true } }`. The algorithm is the one documented at [qntm.org/cmd](https://qntm.org/cmd) (the reference writeup of `cmd.exe`'s quoting behavior) and used by the widely-deployed `cross-spawn` package: each argument is quoted the way `CommandLineToArgvW` expects (backslash-doubling before an embedded quote or at the end of the string, then wrapped in `"`), and THEN every `cmd.exe` metacharacter in that quoted text (`( ) % ! ^ " < > & | ; ,` and space) is caret-escaped, because `cmd.exe`'s own line scanner reads those characters off the raw command line before the quoting is honored, quote or no quote. `windowsVerbatimArguments: true` tells Node not to re-quote the string a second, conflicting way. `test/judges.test.js` pins exact expected output for `&`, `|`, `^`, `%`, a literal `"`, a trailing backslash and a literal newline (the last one deliberately unescaped: it is not a `cmd.exe` metacharacter).
|
|
124
|
-
- **`killTree` needs no change for which process it targets, either path.** `taskkill /pid <pid> /T /F` walks the whole descendant tree regardless of whether the direct child is `node` (the resolved-shim path) or `cmd.exe` (the fallback); there is no intermediate shell layer to lose track of in the common case, since there is no shell there at all. It DID need a change for how `taskkill` itself is found: `windows-latest` CI caught a bare `spawn('taskkill', ...)` failing `ENOENT` the first time a lane actually ran end to end there (this project's own test harness deliberately narrows PATH to isolate a fake lane, and that narrowed PATH does not include `System32`; a sandboxed or otherwise stripped-down real environment might not either), and the resulting unheard `error` event on the returned process crashed the whole run over what should be a best-effort cleanup step. `taskkillPath()` resolves the executable under `%SystemRoot%` (falling back through `%windir%` to a fixed path) instead of relying on PATH, built with a literal backslash rather than `node:path`'s `join()`, which picks its separator from the HOST running the code, not the OS the path describes; `killTree` now attaches an `error` listener so any future spawn failure stays a missed cleanup, never a crash.
|
|
125
|
-
- **`bin/cli.js`'s own `npm install -g` prompt reuses this, rather than duplicating it.** The installer's opt-in "run `npm install -g <ai>` now?" prompt had the identical `EINVAL`-shaped defect (`spawnSync('npm', ...)` with no shell), found the same way: it failed the moment its own test actually ran on `windows-latest`. It now resolves `npm` with `which()` and calls `windowsSpawnPlan()`, the same function above, instead of a second copy of the fix.
|
|
126
|
-
- **Not yet attacked for real.** `windowsSpawnPlan`, `resolveCmdShim` and the escaping functions are unit-tested (pure string logic, runs on every CI host) and the resolved-shim path is exercised end to end on `windows-latest` through the fake-lane fixtures in `test/cli.test.js` (installed as a real npm-style `.cmd` shim). A lane can no longer reach `cmd.exe` at all (a unit test pins the refusal). The opt-in `cmd.exe` branch, used only by the installer's own `npm install -g`, is not exercised end to end through a live Windows process in this suite; its escaping is proven by exact-string unit tests only, and its arguments never include user text.
|
|
127
|
-
- **A platform limit found the same way, unrelated to the spawn path itself: Windows has no OS-level signals at all.** `cli-run.mjs`'s graceful shutdown (`process.on('SIGTERM', ...)`, kill the lane's process group, then exit 143/130) is a POSIX guarantee only: `ChildProcess.kill(sig)` on Windows calls `TerminateProcess()` unconditionally for SIGTERM AND SIGINT alike, giving the target process no chance to run any handler at all, proven on `windows-latest` CI (the wrapper died as `{code: null, signal: sig}` for both; a hypothesis that SIGINT gets a real, catchable console-control event on Windows was tried first and measured false in this exact scenario, not assumed). `test/cli.test.js`'s `#13` now expects an unhandled termination for either signal on win32, and the original graceful-exit assertion elsewhere.
|
|
128
|
-
- **Two narrow, individually-verified Windows skips remain, neither in the spawn path itself.** (1) A lane dying mid-run from a real POSIX signal cannot be reproduced on win32: a real Windows lane is a plain `node <script>` process, so it cannot die "by signal" any more than the product being tested can, and the only way a test fixture can even simulate one (a nested `sh -c "...; kill -TERM $$"`) puts an extra node process between cli-run.mjs and the dying shell, so cli-run.mjs observes only that node's translated exit code (measured: MSYS bash's self-kill status leaks through as a plain nonzero exit code, 3840, which this tool already handles honestly via `exit_nonzero`). (2) `weekly-audit.sh`'s watchdog (`bounded()`/`killtree()`, `pgrep -P` plus killing a backgrounded subshell's tree) relies on real bash job control this script only ever runs under on the Ubuntu box it targets; actually executing it against a genuinely hanging stub under Git Bash's job-control emulation hung past a 20s outer timeout on `windows-latest` CI, a known class of MSYS/Cygwin limitation (a `kill -KILL` not reliably reaching the underlying Windows process tree of a backgrounded subshell), not a defect in the generated script, which still renders and syntax-checks correctly.
|
|
129
|
-
|
|
130
|
-
## New in 0.1.23: failure classes in cli-run
|
|
131
|
-
|
|
132
|
-
`bin/cli-run.mjs` now puts every run in a closed failure class that owns the exit code (`auth` 14, `quota` 15, `rejected` 16, `refused` 17, `cut_short` 18, beside `ok` 0, `empty` 10, `no_output` 11, `timeout` 12, `unavailable` 13), counts refused tool calls per lane (including a read of grok's session transcript, whose `sessionId` comes from lane stdout), and prints redacted problem and fix lines. The durable log gains only `class` and `refused`. Threat model: a false exit 0, provider text or a secret reaching the log or the terminal, a transcript read escaping the sessions root, misclassification that sends a user the wrong way, and anything that throws or stalls.
|
|
133
|
-
|
|
134
|
-
### Round 1 (pre-release, GPT-6 Astra at xhigh), fixed before shipping
|
|
135
|
-
|
|
136
|
-
Every finding was reproduced as a failing case in `test/classify.test.js` against the unfixed code (all seven red), then fixed; the suite is green after. One round, by rule; the regression cases verify the fixes.
|
|
137
|
-
|
|
138
|
-
| # | Sev | Finding | Fix | Test |
|
|
139
|
-
|---|---|---|---|---|
|
|
140
|
-
| 1 | HIGH | qwen's display detail clipped an error message at 120 characters before redaction, so a long JSON password lost its closing quote and printed its prefix | redact before every clip: judge detail, API-error text, denial text, deliverable snippet, problem cause | R1 |
|
|
141
|
-
| 2 | HIGH | the JSON credential pattern ended at an escaped quote and missed unterminated values | match JSON string escapes, run an unterminated value to the end, and accept a JSON body escaped inside a string | R2 |
|
|
142
|
-
| 3 | HIGH | `firstDenialText` scanned 16 MiB with an unbounded `[^)]*`: 1 MiB of unclosed `Permission denied for command(` took 15.6 s after the lane had exited | scan at most 256 KiB of stdout and of stderr, with bounded repeats | R3 |
|
|
143
|
-
| 4 | MEDIUM | transcript containment was a string prefix with both separators, so on POSIX a directory literally named `sessions\outside` passed | `isInsideRoot()` via `path.relative`, rejecting `..`, absolute and cross-drive results; tested with `path.posix` and `path.win32` on every OS | R4 |
|
|
144
|
-
| 5 | MEDIUM | qwen signals were searched in the clipped display detail, so a quota error after 120 characters read as `empty`, and a model name in the detail could fake a signal | classify on `qwenErrorText()`: the terminal event's full error and an API-error result only | R5 |
|
|
145
|
-
| 6 | MEDIUM | hermes fell back to stdout when stderr was empty, so an answer saying "no rate limit" read as `quota` | stderr only | R6 |
|
|
146
|
-
| 7 | MEDIUM | a nonzero vendor exit with a recognised empty shape (agy exit 7, `SUCCESS`, empty response) returned `empty`, contradicting the documented `cut_short` fallback | an unexplained nonzero exit is `cut_short`, except hermes' own exit 2 | R7 |
|
|
147
|
-
|
|
148
|
-
Reported clean by the same audit: no false exit 0 across 675 malformed probes and five lanes; no provider text or secret fragment in probed log records; the real codex 0.153.4 fixture's non-fatal error item still succeeds; grok traversal, symlink, FIFO, session-match and read-cap guards; timeout, interrupt, overrun, killed lane, JSON contract, `--quiet` and the doctor canary. Not verified by it: native Windows execution, live vendors, filesystem races.
|
package/scripts/README.md
DELETED
|
@@ -1,7 +0,0 @@
|
|
|
1
|
-
# scripts/
|
|
2
|
-
|
|
3
|
-
| File | Job |
|
|
4
|
-
|---|---|
|
|
5
|
-
| `gen-catalog.js` | regenerates `docs/catalog.md` AND the vendor compatibility table in `README.md` (between the `vendor-table` markers) from `src/catalog.js`; `npm run gen:catalog`. `test/catalog.test.js` fails if either generated surface disagrees with the catalog. |
|
|
6
|
-
| `record-demo.sh` | re-records `docs/demo.gif` by installing the published package into a temp folder and running it under `asciinema`, then rendering the cast with `agg`. Pass a version to pin one: `bash scripts/record-demo.sh 0.1.27`. Needs `brew install asciinema agg`. The frames are the installer's own output, so the GIF stays true to what the command prints. |
|
|
7
|
-
| `gen-plugin.js` | regenerates the Claude Code plugin bundle in `plugin/` (agents, the two read-only hooks, `plugin.json`, `LICENSE`) from `templates/`, using the plan in `src/plugin.js`; `npm run gen:plugin`. `test/plugin.test.js` fails if the committed bundle disagrees. `plugin/README.md` and `plugin/hooks/hooks.json` are hand-owned. |
|
package/scripts/gen-catalog.js
DELETED
|
@@ -1,81 +0,0 @@
|
|
|
1
|
-
#!/usr/bin/env node
|
|
2
|
-
// Regenerates docs/catalog.md and the README's vendor compatibility table from
|
|
3
|
-
// src/catalog.js. The test suite checks they agree, so neither can drift (#23).
|
|
4
|
-
import { writeFileSync, readFileSync } from 'node:fs';
|
|
5
|
-
import { AIS, LEVELS, TOOLS, IMAGES, npmSpec } from '../src/catalog.js';
|
|
6
|
-
import { readdirSync } from 'node:fs';
|
|
7
|
-
|
|
8
|
-
export function protocolCount() {
|
|
9
|
-
return readdirSync(new URL('../templates/common/protocols/', import.meta.url)).filter((f) => f.endsWith('.md') && f !== 'README.md').length;
|
|
10
|
-
}
|
|
11
|
-
|
|
12
|
-
export function catalogMarkdown() {
|
|
13
|
-
let md = '# Catalog\n\nGenerated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrites it. Protocols shipped at every level: ' + protocolCount() + ' (counted from `templates/common/protocols/`).\n\n## Levels\n\n| Level | Name | Tagline | Gives |\n|---|---|---|---|\n';
|
|
14
|
-
for (const l of LEVELS) md += `| ${l.id} | ${l.name} | ${l.tagline} | ${l.gives} |\n`;
|
|
15
|
-
md += '\n## AIs\n\n';
|
|
16
|
-
for (const a of AIS) {
|
|
17
|
-
const how = a.install.npm
|
|
18
|
-
? '`npm install -g ' + npmSpec(a) + '`'
|
|
19
|
-
: a.install.script
|
|
20
|
-
? 'vendor script (read it first): `' + a.install.script + '`'
|
|
21
|
-
: a.install.url + (a.install.brew ? ' (or `brew install ' + a.install.brew + '`)' : '');
|
|
22
|
-
md += `### \`${a.id}\` · ${a.name}\n\n- **Kind:** ${a.kind} · **Access:** ${a.access} · **Lane:** ${a.lane} · **Level:** ${a.minLevel}+\n- **Wins at:** ${a.role}\n- **Install:** ${how}\n- **Sign in:** ${a.auth}\n`;
|
|
23
|
-
if (a.rulesFile) md += `- **Reads rules from:** \`${a.rulesFile}\`` + (a.agentsDir ? ` · subagents in \`${a.agentsDir}/\`` : '') + '\n';
|
|
24
|
-
if (a.cliRun) md += '- **cli-run lane:** yes\n';
|
|
25
|
-
if (a.plans) {
|
|
26
|
-
md += '- **Plans:**\n';
|
|
27
|
-
for (const p of a.plans) md += ` - ${p.name} (${p.headroom} headroom, checked ${p.checked}): ${p.source}\n`;
|
|
28
|
-
}
|
|
29
|
-
if (a.builtAgainst) md += `- **Built against:** ${a.builtAgainst}` + (a.install.npm ? ' (the same number the npm pin uses)' : '') + '\n';
|
|
30
|
-
if (a.note) md += `- **Note:** ${a.note}\n`;
|
|
31
|
-
md += '\n';
|
|
32
|
-
}
|
|
33
|
-
md += '## Companion tools\n\n';
|
|
34
|
-
for (const t of TOOLS) {
|
|
35
|
-
md += `### \`${t.id}\` · ${t.name}\n\n- **Repo:** ${t.repo}\n- **Gives:** ${t.role}\n- **Install:** \`${t.install}\` (needs ${t.requires})\n- **Registers itself with:** ${t.autoClients.join(', ')}; snippets for the rest are written to \`mcp/\`\n- **Default:** ${t.recommended ? 'selected' : 'not selected'}\n\n`;
|
|
36
|
-
}
|
|
37
|
-
return md;
|
|
38
|
-
}
|
|
39
|
-
|
|
40
|
-
export const VENDOR_TABLE_START = '<!-- vendor-table:start -->';
|
|
41
|
-
export const VENDOR_TABLE_END = '<!-- vendor-table:end -->';
|
|
42
|
-
|
|
43
|
-
export function fixtureManifest() {
|
|
44
|
-
return JSON.parse(readFileSync(new URL('../test/fixtures/manifest.json', import.meta.url), 'utf8'));
|
|
45
|
-
}
|
|
46
|
-
|
|
47
|
-
// The compatibility table in README.md. Every number comes from `builtAgainst`
|
|
48
|
-
// in src/catalog.js, which is also the npm pin where there is one, so "built
|
|
49
|
-
// against" and "pinned to" are the same number by construction (#23).
|
|
50
|
-
export function vendorTableMarkdown() {
|
|
51
|
-
const fx = fixtureManifest();
|
|
52
|
-
const byLane = Object.fromEntries(fx.fixtures.map((f) => [f.lane, f]));
|
|
53
|
-
let md = VENDOR_TABLE_START + '\n\n| Lane | Vendor | Version this release was built against | Where that number is proved |\n|---|---|---|---|\n';
|
|
54
|
-
for (const a of AIS) {
|
|
55
|
-
if (!a.bin || !a.builtAgainst) continue;
|
|
56
|
-
const proof = byLane[a.id]
|
|
57
|
-
? '`test/fixtures/' + byLane[a.id].file + '`, a recorded run'
|
|
58
|
-
: a.install.npm
|
|
59
|
-
? 'the npm pin the installer writes, `' + npmSpec(a) + '`'
|
|
60
|
-
: a.id === 'ollama'
|
|
61
|
-
? 'the pinned image the level 3 box runs, `' + IMAGES.ollama + '`'
|
|
62
|
-
: "the maintainer's own install";
|
|
63
|
-
md += `| \`${a.bin}\` | ${a.vendor} | ${a.builtAgainst} | ${proof} |\n`;
|
|
64
|
-
}
|
|
65
|
-
md += `\nGenerated from \`src/catalog.js\` by \`npm run gen:catalog\`; \`npm test\` fails if this table and the catalog disagree. Fixtures were captured ${fx.capturedAt}.\n\n` + VENDOR_TABLE_END;
|
|
66
|
-
return md;
|
|
67
|
-
}
|
|
68
|
-
|
|
69
|
-
export function readmeWithVendorTable(current) {
|
|
70
|
-
const a = current.indexOf(VENDOR_TABLE_START);
|
|
71
|
-
const b = current.indexOf(VENDOR_TABLE_END);
|
|
72
|
-
if (a === -1 || b === -1) throw new Error('README.md has no vendor-table markers');
|
|
73
|
-
return current.slice(0, a) + vendorTableMarkdown() + current.slice(b + VENDOR_TABLE_END.length);
|
|
74
|
-
}
|
|
75
|
-
|
|
76
|
-
if (process.argv[1] && process.argv[1].endsWith('gen-catalog.js')) {
|
|
77
|
-
writeFileSync(new URL('../docs/catalog.md', import.meta.url), catalogMarkdown());
|
|
78
|
-
const readmeUrl = new URL('../README.md', import.meta.url);
|
|
79
|
-
writeFileSync(readmeUrl, readmeWithVendorTable(readFileSync(readmeUrl, 'utf8')));
|
|
80
|
-
console.log('docs/catalog.md and the README vendor table regenerated');
|
|
81
|
-
}
|
package/scripts/gen-plugin.js
DELETED
|
@@ -1,16 +0,0 @@
|
|
|
1
|
-
#!/usr/bin/env node
|
|
2
|
-
// Regenerates the Claude Code plugin bundle under plugin/ from the installer's
|
|
3
|
-
// templates (src/plugin.js has the plan). The test suite checks the committed
|
|
4
|
-
// bundle matches, so an agent or hook edited in templates/ cannot ship to npm
|
|
5
|
-
// users and silently not to plugin users.
|
|
6
|
-
import { writeFileSync, mkdirSync } from 'node:fs';
|
|
7
|
-
import { join, dirname } from 'node:path';
|
|
8
|
-
import { planPluginFiles, PLUGIN_DIR } from '../src/plugin.js';
|
|
9
|
-
|
|
10
|
-
const files = planPluginFiles();
|
|
11
|
-
for (const f of files) {
|
|
12
|
-
const abs = join(PLUGIN_DIR, ...f.rel.split('/'));
|
|
13
|
-
mkdirSync(dirname(abs), { recursive: true });
|
|
14
|
-
writeFileSync(abs, f.content);
|
|
15
|
-
}
|
|
16
|
-
console.log('plugin/ regenerated: ' + files.length + ' files');
|
package/scripts/record-demo.sh
DELETED
|
@@ -1,45 +0,0 @@
|
|
|
1
|
-
#!/usr/bin/env bash
|
|
2
|
-
# Re-records docs/demo.gif from the PUBLISHED package, so the frames are the
|
|
3
|
-
# program's own output and anyone can reproduce them.
|
|
4
|
-
#
|
|
5
|
-
# bash scripts/record-demo.sh # records the current published version
|
|
6
|
-
# bash scripts/record-demo.sh 0.1.27 # pins a version
|
|
7
|
-
#
|
|
8
|
-
# Needs asciinema (the recorder) and agg (cast to GIF), both from the asciinema
|
|
9
|
-
# project: brew install asciinema agg
|
|
10
|
-
set -euo pipefail
|
|
11
|
-
|
|
12
|
-
VERSION="${1:-latest}"
|
|
13
|
-
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
|
14
|
-
OUT="$REPO_ROOT/docs/demo.gif"
|
|
15
|
-
WORK="$(mktemp -d)"
|
|
16
|
-
trap 'rm -rf "$WORK"' EXIT
|
|
17
|
-
|
|
18
|
-
for tool in asciinema agg npm; do
|
|
19
|
-
command -v "$tool" >/dev/null || { echo "missing: $tool (brew install asciinema agg)" >&2; exit 1; }
|
|
20
|
-
done
|
|
21
|
-
|
|
22
|
-
# Installed locally first, so the recorded npx call resolves from node_modules
|
|
23
|
-
# and never stops on npx's own "Ok to proceed?" prompt mid-take.
|
|
24
|
-
cd "$WORK"
|
|
25
|
-
npm install --silent "model-orchestrator@$VERSION" >/dev/null
|
|
26
|
-
|
|
27
|
-
cat > "$WORK/take.sh" <<'TAKE'
|
|
28
|
-
#!/bin/bash
|
|
29
|
-
cd "$(dirname "$0")"
|
|
30
|
-
printf '~/my-app $ '
|
|
31
|
-
sleep 0.8
|
|
32
|
-
cmd='npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dry'
|
|
33
|
-
for (( i=0; i<${#cmd}; i++ )); do printf '%s' "${cmd:$i:1}"; sleep 0.022; done
|
|
34
|
-
sleep 0.6; printf '\n'
|
|
35
|
-
npx model-orchestrator --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dry
|
|
36
|
-
printf '~/my-app $ '
|
|
37
|
-
sleep 2.5
|
|
38
|
-
TAKE
|
|
39
|
-
chmod +x "$WORK/take.sh"
|
|
40
|
-
|
|
41
|
-
asciinema rec --cols 108 --rows 34 --overwrite --command "$WORK/take.sh" "$WORK/demo.cast"
|
|
42
|
-
agg --theme monokai --font-size 16 --speed 1.3 --idle-time-limit 1.2 --last-frame-duration 3 \
|
|
43
|
-
"$WORK/demo.cast" "$OUT"
|
|
44
|
-
|
|
45
|
-
echo "wrote $OUT ($(wc -c < "$OUT") bytes)"
|
|
@@ -1,56 +0,0 @@
|
|
|
1
|
-
# Task Bundle: the brief every delegation carries
|
|
2
|
-
|
|
3
|
-
A subagent, a second CLI, or a fresh chat window may hold none of the rules your main session is holding, and that is the default to assume. One exception: a Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already carries the standing rules, just not this task's scope. Either way, it cannot see this task's conventions and will read an unspecified edge as an open one.
|
|
4
|
-
|
|
5
|
-
> A delegate gets an approved, bounded brief. Absence is not permission.
|
|
6
|
-
|
|
7
|
-
A brief is under-specified if it is missing **purpose**, **denied actions**, **report contract**, or **exit parameters**.
|
|
8
|
-
|
|
9
|
-
## Template
|
|
10
|
-
|
|
11
|
-
Copy this into the delegate's prompt. Delete nothing; write `none` where a field is genuinely empty, so a reader can tell "nothing denied" from "nobody thought about it".
|
|
12
|
-
|
|
13
|
-
```markdown
|
|
14
|
-
## Task bundle
|
|
15
|
-
|
|
16
|
-
**Purpose.** <one sentence: what this task is for, and why>
|
|
17
|
-
**Task class.** <read_only | draft_only | mutating> (draft_only = produce, do not apply)
|
|
18
|
-
|
|
19
|
-
**Granted scope.**
|
|
20
|
-
- <paths, globs, topics, or record sets this brief covers>
|
|
21
|
-
- Anything outside this list is out of scope. Do not widen it on your own judgment.
|
|
22
|
-
|
|
23
|
-
**Capabilities.** <the actions you MAY take: read, search, write to <path>, run <cmd>>
|
|
24
|
-
|
|
25
|
-
**Denied actions.** <explicit list: do not commit, push, deploy, delete, send, publish, close a ticket...>
|
|
26
|
-
- Anything absent from Capabilities is denied. Absence is not permission.
|
|
27
|
-
|
|
28
|
-
**Conventions you do not have.** <restate every house rule this task needs; even a delegate that loaded the standing rules still needs this task's scope, and a second CLI or a fresh chat window may hold none of it>
|
|
29
|
-
|
|
30
|
-
**Report contract.** Return: <exactly what to hand back>. State plainly what you did NOT do
|
|
31
|
-
and anything you could not verify. "Unverified" is an acceptable answer; a confident guess is not.
|
|
32
|
-
|
|
33
|
-
**Exit parameters.** <at least one bound: a wall-clock ceiling, a work ceiling ("at most 20 files"),
|
|
34
|
-
or a stop condition. Plus the partial-result clause: if you hit a bound, report what you have and
|
|
35
|
-
name what you did not cover. Never keep going past a bound, never return nothing.>
|
|
36
|
-
```
|
|
37
|
-
|
|
38
|
-
## When to skip it
|
|
39
|
-
|
|
40
|
-
A one-line read-only lookup can say so outright:
|
|
41
|
-
|
|
42
|
-
```
|
|
43
|
-
Task bundle: none (one-line lookup, read-only, no artifact)
|
|
44
|
-
```
|
|
45
|
-
|
|
46
|
-
The reason is required. A bare `none` is indistinguishable from "forgot".
|
|
47
|
-
|
|
48
|
-
Never skip it for anything touching secrets, deletion, bulk mutation, deploys, or someone else's data. Better: do not delegate those at all.
|
|
49
|
-
|
|
50
|
-
## Field notes
|
|
51
|
-
|
|
52
|
-
- **Purpose** is what makes the rest checkable. "Fix the thing" has no edge to exceed.
|
|
53
|
-
- **Task class** is the cheapest safety win. Most delegation wants `read_only` or `draft_only`.
|
|
54
|
-
- **Denied actions** must be written even when they feel obvious. Nothing is obvious to a delegate with no context.
|
|
55
|
-
- **Report contract** is what turns a result into evidence rather than a claim.
|
|
56
|
-
- **Exit parameters** exist because a delegate that never returns is more expensive than one that returns wrong. A wrong answer is corrected next turn; a hang burns the session while looking like progress. Bound your own shell calls the same way (pass a timeout; scope recursive searches away from `.git`, `node_modules`, build output).
|