model-orchestrator 0.1.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (81) hide show
  1. package/CHANGELOG.md +96 -0
  2. package/LICENSE +21 -0
  3. package/README.md +133 -0
  4. package/SECURITY.md +17 -0
  5. package/bin/README.md +10 -0
  6. package/bin/cli-run.mjs +599 -0
  7. package/bin/cli.js +372 -0
  8. package/docs/README.md +13 -0
  9. package/docs/audit-brief.md +83 -0
  10. package/docs/catalog.md +113 -0
  11. package/docs/part-1-beginner.md +65 -0
  12. package/docs/part-2-intermediate.md +65 -0
  13. package/docs/part-3-advanced.md +65 -0
  14. package/package.json +52 -0
  15. package/scripts/README.md +5 -0
  16. package/scripts/gen-catalog.js +37 -0
  17. package/src/README.md +9 -0
  18. package/src/catalog.js +277 -0
  19. package/src/detect.js +26 -0
  20. package/src/install.js +628 -0
  21. package/src/prompt.js +34 -0
  22. package/src/render.js +8 -0
  23. package/templates/README.md +14 -0
  24. package/templates/advanced/README.md +14 -0
  25. package/templates/advanced/vm/ENVIRONMENT.md +18 -0
  26. package/templates/advanced/vm/PRIVACY_GATES.md +33 -0
  27. package/templates/advanced/vm/README.md +60 -0
  28. package/templates/advanced/vm/box-CLAUDE.md +28 -0
  29. package/templates/advanced/vm/docker-compose.yml +19 -0
  30. package/templates/advanced/vm/gateway.config.yaml +12 -0
  31. package/templates/advanced/vm/jobs/README.md +39 -0
  32. package/templates/advanced/vm/jobs/weekly-audit.service +17 -0
  33. package/templates/advanced/vm/jobs/weekly-audit.sh +107 -0
  34. package/templates/advanced/vm/jobs/weekly-audit.timer +10 -0
  35. package/templates/advanced/vm/setup-vm.sh +46 -0
  36. package/templates/agents/README.md +13 -0
  37. package/templates/agents/agy/README.md +5 -0
  38. package/templates/agents/agy/builder.md +17 -0
  39. package/templates/agents/agy/bulk-worker.md +17 -0
  40. package/templates/agents/agy/code-reviewer.md +17 -0
  41. package/templates/agents/agy/deep-planner.md +17 -0
  42. package/templates/agents/agy/live-researcher.md +17 -0
  43. package/templates/agents/claude-code/README.md +13 -0
  44. package/templates/agents/claude-code/builder.md +17 -0
  45. package/templates/agents/claude-code/bulk-worker.md +18 -0
  46. package/templates/agents/claude-code/code-reviewer.md +19 -0
  47. package/templates/agents/claude-code/deep-planner.md +18 -0
  48. package/templates/agents/claude-code/live-researcher.md +18 -0
  49. package/templates/agents/snippets/chat.md +25 -0
  50. package/templates/agents/snippets/claude-code.md +27 -0
  51. package/templates/agents/snippets/generic.md +21 -0
  52. package/templates/beginner/ORCHESTRATOR.md +55 -0
  53. package/templates/beginner/README.md +3 -0
  54. package/templates/common/README.md +52 -0
  55. package/templates/common/TASK_BUNDLE.md +56 -0
  56. package/templates/common/protocols/README.md +14 -0
  57. package/templates/common/protocols/build-protocol.md +133 -0
  58. package/templates/common/protocols/deep-research.md +44 -0
  59. package/templates/common/protocols/gap-analysis.md +28 -0
  60. package/templates/common/protocols/memory-and-record.md +30 -0
  61. package/templates/common/protocols/numbers-and-logic.md +35 -0
  62. package/templates/common/protocols/propagate.md +34 -0
  63. package/templates/intermediate/CLI-RUN.md +100 -0
  64. package/templates/intermediate/DELEGATION_MATRIX.md +41 -0
  65. package/templates/intermediate/README.md +13 -0
  66. package/templates/intermediate/RESEARCH_TRIAGE.md +30 -0
  67. package/templates/intermediate/ROUTING.md +73 -0
  68. package/templates/intermediate/TIERS.md +44 -0
  69. package/templates/tools/README.md +10 -0
  70. package/templates/tools/codecalc/CODECALC.md +43 -0
  71. package/templates/tools/codecalc/mcp/agy.mcp_config.json +8 -0
  72. package/templates/tools/codecalc/mcp/codex.config.toml +4 -0
  73. package/templates/tools/codecalc/mcp/mcpServers.json +8 -0
  74. package/templates/tools/codecalc/mcp/vscode.mcp.json +8 -0
  75. package/templates/tools/codecalc/mcp/zed.settings.json +9 -0
  76. package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +65 -0
  77. package/templates/tools/obsidian-tc/mcp/obsidian-tc.agy.mcp_config.json +9 -0
  78. package/templates/tools/obsidian-tc/mcp/obsidian-tc.codex.config.toml +7 -0
  79. package/templates/tools/obsidian-tc/mcp/obsidian-tc.mcpServers.json +9 -0
  80. package/templates/tools/obsidian-tc/mcp/obsidian-tc.vscode.mcp.json +10 -0
  81. package/templates/tools/obsidian-tc/mcp/obsidian-tc.zed.settings.json +9 -0
package/CHANGELOG.md ADDED
@@ -0,0 +1,96 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented here. The format follows [Keep a Changelog 2.0.0](https://keepachangelog.com/en/2.0.0/) and the project uses [Semantic Versioning](https://semver.org/spec/v2.0.0.html); on `0.y.z` anything may change.
4
+
5
+ ## [Unreleased]
6
+
7
+ ### Changed
8
+
9
+ - `RELEASING.md`: the first npm publish is manual and must not pass `--provenance` (npm only generates provenance inside a supported CI runner); the workflow route comes after the package exists.
10
+
11
+ ## [0.1.4] - 2026-09-05
12
+
13
+ ### Added
14
+
15
+ - `--update-docs`: after a selection change, regenerate the documents a previous run wrote and nobody edited since. The check is the same hash rule the runtime class uses: an installed copy that matches the hash `MANIFEST.json` recorded is regenerated and named under "documents updated"; one that differs is kept and named under "document CONFLICT, kept"; without a manifest every changed document is kept as UNVERIFIABLE. `--force` still replaces everything; `--dry` reports and writes nothing. The reconfiguration hint names the flag.
16
+
17
+ ### Fixed
18
+
19
+ - `MANIFEST.json` recorded the hash of content a run planned for a document it then kept, so the next hash check read every kept document (and, on a reconfiguration, every kept runtime file) as edited. The manifest now records the previous run's hash for kept files and no entry when there was no previous manifest, so the file classes tell the truth about what is on disk. Project-root agent definitions are keyed under `[project]`.
20
+
21
+ ## [0.1.3] - 2026-09-05
22
+
23
+ Professional-repo pass and the eight findings from the agy scored audit (overall 9.4/10; the findings are in the issue tracker's audit record).
24
+
25
+ ### Added
26
+
27
+ - Community files: `CONTRIBUTING.md`, `CODE_OF_CONDUCT.md` (Contributor Covenant 2.1), `MAINTAINERS.md`, `RELEASING.md`, `AGENTS.md` and `CLAUDE.md` for contributors' agents, yml issue forms with blank issues disabled, a pull request template, `CODEOWNERS`, `.editorconfig`, Dependabot for the workflow actions.
28
+ - `test/prose.test.js`: fails on an em dash anywhere in the repo's text files, so the house rule is checked rather than requested.
29
+ - README badges (CI, licence, Node) and a one-line privacy statement.
30
+
31
+ ### Fixed
32
+
33
+ - `--help`: `--primary` was described as "level 1 only"; it applies at every level and is required when several agents qualify.
34
+ - `docs/catalog.md`: install lines now carry the same npm pin the installer uses (`scripts/gen-catalog.js` went through `npmSpec`, closing the last gap #8 left open).
35
+ - `--upgrade-runtime`: the report names the runtime files it replaced; before, the files were replaced and the "runtime upgraded:" line never printed.
36
+ - `cli-run --doctor`: prints a note when the primary agent is absent from the lane list, so "1 enabled lane(s): codex" after a Claude Code + Codex install no longer reads as a missed install.
37
+ - `templates/README.md`: listed four protocols and one companion tool; there are six and two. A test now checks that table against the tree.
38
+ - Agent snippets name the routing file for the level (`ORCHESTRATOR.md` at 1, `ROUTING.md` at 2 and 3) instead of a conditional clause; `common/README.md` says why both files exist at level 2.
39
+
40
+ ### Changed
41
+
42
+ - CI: `cli-run --doctor` is no longer masked with `|| true`; exit 0 or 10 (an enabled lane's binary absent on the runner) passes, anything else fails the job.
43
+ - The tarball ships `CHANGELOG.md` and `SECURITY.md` (added to `files`).
44
+ - CI: actions pinned to full commit SHAs with a version comment, `permissions: contents: read`, one run per branch with `cancel-in-progress`, `fail-fast: false` so one leg's failure does not hide another's result.
45
+ - `SECURITY.md` states a response window (7 days to acknowledge, 30 to fix or decline), that only the latest release receives fixes, and the scope.
46
+ - This changelog reshaped to Keep a Changelog 2.0.0 with dated releases and compare links.
47
+
48
+ ## [0.1.2] - 2026-09-05
49
+
50
+ Follow-up audit of 0.1.1 (issues #12 to #15).
51
+
52
+ ### Fixed
53
+
54
+ - `cli-run`: SIGINT/SIGTERM to the wrapper kill the lane's process group before exiting 130/143, with the handlers registered before the spawn so a slow runner cannot signal between the two (#13).
55
+ - `cli-run`: stdout and stderr go through streaming UTF-8 decoders and limits are counted in bytes, so a multibyte character split across chunks survives (#14).
56
+ - `cli-run`: `--expect-file` snapshots the target before the run and requires it to be new or changed, so a pre-existing artifact fails however recent (#15).
57
+
58
+ ### Changed
59
+
60
+ - Installer: three file classes. `MANIFEST.json` records the generator version and a hash per generated file; runtime files (`cli-run.mjs`, the audit job, unit and timer, compose, gateway config, setup script) are upgraded when the installed copy is provably untouched, kept and reported as a conflict when edited, kept and reported as unverifiable when no manifest exists; `--upgrade-runtime` replaces runtime files only (#12). The README discloses the machine-owned and runtime exceptions to "never overwrite".
61
+
62
+ ## [0.1.1] - 2026-09-05
63
+
64
+ Audit follow-up (issues #1 to #10 on the repo).
65
+
66
+ ### Fixed
67
+
68
+ - `cli-run`: lanes run in their own process group and the group is killed on timeout or buffer overrun (#1); the durable log stores only a fixed reason code (#4); nonzero vendor exits pass through with an `exit_nonzero` verdict and a bounded stderr head on the terminal (#9).
69
+ - Weekly audit: temp-and-rename so a failed rerun never truncates the last good report, failed output kept beside it (#3); every probe under a watchdog, `TimeoutStartSec=900`, `UNVERIFIED` lines for timed-out probes (#10).
70
+ - Installer: one `npmSpec` helper so the interactive install, the printed command, the table and the box script use the same pinned version (#8).
71
+
72
+ ### Added
73
+
74
+ - `cli-run`: `--expect-file` and `--expect-json` opt-in contracts, with the guarantee of a bare run stated exactly (#5).
75
+ - Weekly audit: `--audit` for codex and an explicit boundary note for other lanes (#2).
76
+ - Installer: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run, with a requested-vs-applied report on reconfiguration (#6).
77
+ - README leads with the GitHub route pinned to the release until the npm publish, and a table stating which properties are enforced, delegated or instructions (#7, #11). CI installs the packed tarball into a clean consumer and runs it.
78
+
79
+ ## [0.1.0] - 2026-09-04
80
+
81
+ First release.
82
+
83
+ ### Added
84
+
85
+ - Installer: three levels, access-aware AI selection, primary-agent loading surface, companion-tool questions (codecalc recommended, obsidian-tc optional), strict flags, containment preflight, exclusive create with rollback, no vendor scripts run.
86
+ - `bin/cli-run.mjs`: one entrypoint for grok, codex, agy, hermes and qwen with each lane's native success signal; exit 10 on a run that produced nothing, fail-closed `lanes.json`, signal handling, digest-only log, `--doctor`.
87
+ - Templates: six protocols (build, propagate, gap analysis, deep research, numbers and logic, memory and record), task bundle, single-agent and multi-lane routing, tiers, generated delegation matrix, research triage, VM tier (gateway config by env-var name, compose on loopback, box rules, privacy gates, weekly audit timer).
88
+ - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
89
+ - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
90
+
91
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.4...HEAD
92
+ [0.1.4]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.3...v0.1.4
93
+ [0.1.3]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.2...v0.1.3
94
+ [0.1.2]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.1...v0.1.2
95
+ [0.1.1]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.0...v0.1.1
96
+ [0.1.0]: https://github.com/aunysillyme/model-orchestrator/releases/tag/v0.1.0
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 model-orchestrator contributors
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,133 @@
1
+ # model-orchestrator
2
+
3
+ [![test](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml/badge.svg)](https://github.com/aunysillyme/model-orchestrator/actions/workflows/test.yml) [![license: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![node >=18](https://img.shields.io/badge/node-%3E%3D18-brightgreen.svg)](package.json)
4
+
5
+ **A model orchestrator you can `npm run`.** Route every task to the cheapest AI that does it well, whether you have one chat app, five agent CLIs, or a virtual machine running them unattended. One installer asks what you have access to and writes only what fits.
6
+
7
+ Built from a working system, not a diagram: the routing rules, the protocols and the lane runner here run in production, generalized so they transfer to any stack.
8
+
9
+ ```bash
10
+ npx github:aunysillyme/model-orchestrator#v0.1.4
11
+ ```
12
+
13
+ That runs the reviewed release straight from GitHub (drop `#v0.1.4` for the current main). `npx model-orchestrator` will work once the package is on the npm registry; until then it is not a command you can run.
14
+
15
+ The installer asks a few things, then writes a folder:
16
+
17
+ 1. **Which level?** 1 beginner · 2 intermediate · 3 advanced
18
+ 2. **Which AIs do you have access to?** (it marks the ones already on your PATH)
19
+ 3. **Which one is your primary agent?** (the one that runs the system)
20
+
21
+ It never writes a secret, never runs a vendor shell script for you, and never overwrites a document you already have unless you pass `--force`. Two exceptions, both stated when they happen: `MANIFEST.json` and `bin/lanes.json` are machine-owned and rewritten on every run so a changed selection applies; runtime files (`cli-run`, the audit job, compose, gateway config, setup script) are upgraded when the installed copy matches the hash a previous run recorded, kept and reported as a conflict when you edited them, and kept as unverifiable when no manifest exists (`--upgrade-runtime` replaces runtime files only). The same hash rule is available for documents on request: `--update-docs` regenerates the documents a previous run wrote and nobody edited, so a changed selection reaches `ROUTING.md` and the delegation matrix without `--force`; edited documents are kept and named. Docs and protocols go to `--dir` (default `./ai-orchestrator`); subagent definitions go to the project root your agent runs from (`--project`, default the current directory), because that is the only place Claude Code and Antigravity read them. It ends with an activation summary: what to copy where, which sign-ins, and one smoke command. Uninstall: delete the folder, the subagent folder it named, and `~/.ai-orchestrator/cli-run.log.jsonl` if you used `cli-run`.
22
+
23
+ ## The three levels
24
+
25
+ | Level | You have | You get |
26
+ |---|---|---|
27
+ | **1 · Beginner** | one LLM or one agent | tiers, task classification, the two build checkpoints, the protocols (build, propagate, gap analysis, deep research, numbers and logic, memory and record), a task-bundle template, and your agent set up to follow them |
28
+ | **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts), a delegation matrix generated from your selection, research triage across the lanes you have |
29
+ | **3 · Advanced** | a virtual machine | everything above, plus a gateway config rendered from the API keys you hold (asked separately from your CLIs), pinned images, box rules, privacy gates, and a weekly gap-analysis job with "what watches it" written down |
30
+
31
+ Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md).
32
+
33
+ ## The AIs it knows about
34
+
35
+ | Id | What | Level |
36
+ |---|---|---|
37
+ | `claude-code` | Claude Code CLI, the default orchestrator | 1+ |
38
+ | `codex` | Codex CLI on a ChatGPT plan: second coder, adversarial auditor | 1+ |
39
+ | `agy` | Antigravity CLI on a Google AI plan: research sweeps, concurrent fan-out | 1+ |
40
+ | `grok` | Grok CLI on X Premium: live X and web reads at $0 | 1+ |
41
+ | `hermes` | Hermes Agent: the free tier | 2+ |
42
+ | `qwen` | Qwen Code CLI with a cheap metered model: structured bulk | 2+ |
43
+ | `ollama` | local models: the privacy lane | 2+ |
44
+ | `claude-app`, `chatgpt-app`, `gemini-app` | chat apps with no CLI: level 1 via a paste block | 1 |
45
+
46
+ `npx github:aunysillyme/model-orchestrator#v0.1.4 --list` prints the catalog with install and sign-in notes. Details: [docs/catalog.md](docs/catalog.md).
47
+
48
+ ## Companion tools (both optional)
49
+
50
+ An orchestrator routes work. It does not make a model stop guessing numbers, and it does not give it a memory. Two tools from the same maintainer close those gaps. The installer asks about each one separately; selecting one writes a doc and config snippets, it installs nothing. `--tools codecalc,obsidian-tc` or `--no-tools` for scripted runs; `--yes` alone selects only the recommended one.
51
+
52
+ | Tool | Closes | Default | You need first |
53
+ |---|---|---|---|
54
+ | [codecalc](https://github.com/The-40-Thieves/codecalc) | guessed numbers, comparisons, complexity and equivalence claims: exact arithmetic, code execution in 31 languages, SMT logic checks, `verify_translation` / `verify_optimization`; offline, no key | yes | Python 3.10+ and `uv`. `uvx 'codecalc[full]' setup --write` registers it with Claude Code, Claude Desktop, Cursor, VS Code, Zed; snippets for Codex, Antigravity, Qwen Code are written for you |
55
+ | [obsidian-tc](https://github.com/The-40-Thieves/obsidian-tc) | no durable memory: hybrid search, backlinks, compare-and-swap writes with a confirmation gate, folder ACLs, a poison scan on inferred writes; 163 tools, local by default; AGPL-3.0 | no | an Obsidian vault folder; Node 24+ or Bun 1.1+ (stricter than this installer); Ollama with `nomic-embed-text` or a cloud embeddings key; the Obsidian app and its Local REST API plugin only for live bridge tools. Skip it if you do not keep notes in Obsidian |
56
+
57
+ Whether or not you select them, every level carries the two rules they serve: `protocols/numbers-and-logic.md` (when calling a calculator is mandatory, how to report a computed figure, why a thought log is not evidence) and `protocols/memory-and-record.md` (search before writing, the folder index is part of the change, one writer, inferred content marked as inferred).
58
+
59
+ ## Non-interactive
60
+
61
+ ```bash
62
+ npx github:aunysillyme/model-orchestrator#v0.1.4 --yes --level 2 --ais claude-code,codex,grok --primary claude-code --dir ./ai-orchestrator
63
+ npx github:aunysillyme/model-orchestrator#v0.1.4 --yes --level 3 --ais claude-code,codex,agy,grok,hermes,qwen,ollama --apis anthropic,openrouter --dry # print the plan, write nothing
64
+ npx github:aunysillyme/model-orchestrator#v0.1.4 --yes --level 2 --ais claude-code,codex --project ~/my-app --dir ~/my-app/ai-orchestrator # subagents into ~/my-app/.claude/agents
65
+ npx github:aunysillyme/model-orchestrator#v0.1.4 --yes --level 2 --ais claude-code,codex,grok --primary claude-code --update-docs # added a lane: regenerate the docs you never edited
66
+ ```
67
+
68
+ ## What gets written (level 3, everything)
69
+
70
+ ```
71
+ ai-orchestrator/
72
+ README.md start here, written for your level and your AIs
73
+ ORCHESTRATOR.md single-agent routing rules (level 1)
74
+ TASK_BUNDLE.md the brief every delegation carries
75
+ protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record
76
+ CODECALC.md OBSIDIAN-TC.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
77
+ <project>/.claude/agents/ five subagents, one per tier, at the PROJECT root (if Claude Code is primary)
78
+ CLAUDE.snippet.md the block to paste into your CLAUDE.md
79
+ ROUTING.md multi-lane decision tree (level 2+)
80
+ TIERS.md DELEGATION_MATRIX.md RESEARCH_TRIAGE.md CLI-RUN.md
81
+ bin/cli-run.mjs bin/lanes.json (node bin/cli-run.mjs --doctor is the smoke test)
82
+ vm/ gateway config, compose, box rules, privacy gates, jobs/ (level 3)
83
+ ```
84
+
85
+ ## Repo layout
86
+
87
+ | Folder | What |
88
+ |---|---|
89
+ | [`bin/`](bin/README.md) | `cli.js` (the installer) and `cli-run.mjs` (the lane runner) |
90
+ | [`src/`](src/README.md) | the catalog, the pure planner, detection, rendering |
91
+ | [`templates/`](templates/README.md) | everything the installer can write, by level, plus `tools/` for companions |
92
+ | [`docs/`](docs/README.md) | the three parts and the catalog |
93
+ | [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu and macOS, Node 18/20/22 |
94
+
95
+ ## What is enforced, what is delegated, what is an instruction
96
+
97
+ Most of what this package ships is text an agent is asked to follow. Be clear about which is which before relying on it unattended.
98
+
99
+ | Property | How it holds |
100
+ |---|---|
101
+ | Installer writes only inside `--dir` and `--project`, never a secret, never over a document without `--force` (or `--update-docs`, which touches only documents provably untouched since a previous run); machine-owned config always, runtime files only when provably untouched or with `--upgrade-runtime` | **enforced by code** (preflight, exclusive create, rollback, manifest hashes; tested) |
102
+ | `cli-run` exit codes, process-group kill on timeout and on SIGINT/SIGTERM, UTF-8-safe streaming, fixed-code durable log, `--expect-*` contracts with a pre-run snapshot | **enforced by code** (tested with stub lanes) |
103
+ | Codex audit lane runs read-only | **delegated to the vendor flag** (`--audit` → `--sandbox read-only`); commands and network still follow your codex config |
104
+ | Other lanes' permissions, sign-in state, model versions | **delegated to each vendor's own config**; `--doctor` checks presence, not versions |
105
+ | Gateway binds to loopback, keys by name only | **enforced in the generated files**; whether the gateway authenticates is your environment |
106
+ | Lane selection, tiers, privacy classes, one-writer, escalation, the protocols | **agent instructions**. Nothing here stops an agent that ignores its rules; the task bundle and the protocols make ignoring them visible, not impossible |
107
+ | Weekly audit bounded, previous report preserved | **enforced in the generated script and unit** (watchdog, temp-and-rename, `TimeoutStartSec`) |
108
+
109
+ If you need a property in the third row to be enforced, that is a router, a policy engine or a sandbox, and this package does not claim to be one.
110
+
111
+ ## Principles the whole thing rests on
112
+
113
+ 1. **Route by capability tier, not model name.** Default down, escalate on evidence.
114
+ 2. **A gate you cannot fail is not a gate.** Every checkpoint is a question that can come back wrong.
115
+ 3. **Exit 0 is not a deliverable.** Check for the artifact, not the status line. `cli-run` checks the response is structurally there; `--expect-file` checks the artifact.
116
+ 4. **Numbers are computed, never guessed.** A tool that calculates beats a model that feels finished.
117
+ 5. **A write nobody can find again did not happen.** Search first, keep the index true, one writer.
118
+ 6. **The orchestrator owns the main build.** Delegates hold none of your rules; they get bounded sub-parts and a brief.
119
+ 7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
120
+
121
+ ## Requirements
122
+
123
+ Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu.
124
+
125
+ **Privacy.** The installer makes no network call of its own and sends no telemetry; the only network activity is the `npm install -g` you approve per package. `cli-run` talks to nothing but the vendor CLI you name.
126
+
127
+ ## Contributing
128
+
129
+ Add an AI to `src/catalog.js` and every prompt, table, config and doc picks it up. Run `npm test`. Keep templates free of logic and free of anything that looks like a credential. The rest is in [CONTRIBUTING.md](CONTRIBUTING.md); releases in [RELEASING.md](RELEASING.md); security reports in [SECURITY.md](SECURITY.md).
130
+
131
+ ## License
132
+
133
+ [MIT](LICENSE)
package/SECURITY.md ADDED
@@ -0,0 +1,17 @@
1
+ # Security policy
2
+
3
+ This tool writes files into folders you name (`--dir` and `--project`) and, only after you say yes per package, runs `npm install -g <package>` for packages pinned in `src/catalog.js`. It never runs a vendor shell script, never writes a credential, and never overwrites a document without `--force`; the two stated exceptions (machine-owned config, hash-verified runtime files) are in the README. `bin/cli-run.mjs` spawns the agent CLI you name with the prompt you give it, in its own process group, and kills that group on timeout, overrun or signal. Nothing here makes a network call of its own; the vendor CLIs and `npm install` do.
4
+
5
+ Threat model and the audit rounds that shipped with each release: `docs/audit-brief.md` and `CHANGELOG.md`.
6
+
7
+ ## Reporting a vulnerability
8
+
9
+ Use GitHub's private vulnerability reporting on this repository (Security tab, "Report a vulnerability"). That opens a private advisory only the maintainer can read. Please do not open a public issue for anything that could let the installer write outside the named folders, run code it should not, expose a key, or let `cli-run` exit 0 for a lane that produced nothing.
10
+
11
+ You will get an acknowledgement within 7 days and a fix or a reasoned "won't fix" within 30. Only the latest release receives fixes. Credit is given in the changelog unless you ask otherwise.
12
+
13
+ Findings are reproduced before they are acted on; `CLEAN` is an acceptable outcome for a report that does not reproduce. Non-sensitive bugs go in a regular issue with the bug form.
14
+
15
+ ## Scope
16
+
17
+ In scope: `bin/`, `src/`, `scripts/`, the templates, and the files the installer writes from them (the weekly audit job, compose file, gateway config and setup script included). Out of scope: the agent CLIs, models and companion tools this package installs or links to; report those to their own projects.
package/bin/README.md ADDED
@@ -0,0 +1,10 @@
1
+ # bin/
2
+
3
+ Two executables. Both are plain Node, no dependencies.
4
+
5
+ | File | What it is | Who runs it |
6
+ |---|---|---|
7
+ | `cli.js` | The installer. `npx model-orchestrator` lands here. Asks level + access, writes files, offers npm installs one at a time. | You, once per machine or project |
8
+ | `cli-run.mjs` | The lane runner. Copied into your install at level 2 and above as `bin/cli-run.mjs`. Runs one agent CLI and exits non-zero unless it produced a deliverable. | Your orchestrator, every time it delegates to another CLI |
9
+
10
+ `cli-run.mjs` exports its judges (`judgeGrok`, `judgeCodex`, `judgeAgy`, `judgeHermes`, `judgeQwen`) so `test/` can prove each one goes red on the failure shapes it exists to catch, without spawning any CLI.