model-orchestrator 0.1.15 → 0.1.17
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +27 -1
- package/README.md +29 -14
- package/bin/cli-run.mjs +22 -7
- package/bin/cli.js +3 -3
- package/docs/README.md +1 -1
- package/docs/audit-brief.md +11 -0
- package/docs/catalog.md +1 -1
- package/docs/part-1-beginner.md +4 -4
- package/docs/part-2-intermediate.md +5 -5
- package/llms.txt +1 -1
- package/package.json +1 -1
- package/src/catalog.js +1 -1
- package/src/detect.js +28 -9
- package/src/install.js +31 -10
- package/templates/agents/agy/README.md +1 -1
- package/templates/agents/agy/builder.md +1 -1
- package/templates/agents/agy/finding-verifier.md +7 -1
- package/templates/agents/claude-code/README.md +2 -2
- package/templates/agents/claude-code/code-reviewer.md +9 -2
- package/templates/agents/claude-code/finding-verifier.md +9 -2
- package/templates/agents/snippets/chat.md +2 -2
- package/templates/agents/snippets/claude-code.md +2 -2
- package/templates/agents/snippets/generic.md +1 -1
- package/templates/agents/snippets/route-gate.mjs +1 -1
- package/templates/agents/snippets/route-metrics.mjs +356 -0
- package/templates/agents/snippets/settings.hooks.snippet.json +44 -0
- package/templates/beginner/ORCHESTRATOR.md +2 -2
- package/templates/common/README.md +1 -1
- package/templates/common/protocols/build-protocol.md +10 -10
- package/templates/common/protocols/deep-research.md +2 -2
- package/templates/common/protocols/gap-analysis.md +1 -1
- package/templates/common/protocols/propagate.md +2 -2
- package/templates/intermediate/CLI-RUN.md +3 -3
- package/templates/intermediate/DELEGATION_MATRIX.md +1 -1
- package/templates/intermediate/ROUTING.md +4 -4
- package/templates/intermediate/TIERS.md +12 -8
- package/templates/tools/codecalc/CODECALC.md +1 -1
- package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +3 -3
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,30 @@ All notable changes to this project are documented here. The format follows [Kee
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [0.1.17] - 2026-09-11
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
- **User-facing text now uses plain language instead of security-audit jargon.** Words like "risk", "attack lane", "adversarial", "blast radius", "fail closed" and "threat model" read as alarming to someone deciding whether to try the tool, so they scared off exactly the readers this project needs. No rule any of them described changed, only the wording: "risk" is now "stakes" everywhere it names a routing input (with a one-line definition added to `README.md` and `TIERS.md`), "attack lane" / "Stage 5 Attack" / "attack pass" are now "challenge lane" / "Stage 5 Challenge" / "challenge pass", "adversarial" (auditor, read, critique, turn, pass) is now "second-opinion", "blast radius" is now "everything it touches", "fail(s) closed" is now "refuses by default", and "threat model" is now "security notes" in the files that link to it. `test/prose.test.js` gained a permanent check (`no alarming security wording in user-facing text`) over the purely-prose, user-facing surface (`docs/`, `templates/`, `README.md`, `llms.txt`, `CONTRIBUTING.md`, the PR template) so the old wording cannot silently creep back in. `docs/audit-brief.md`, `SECURITY.md`, `CODE_OF_CONDUCT.md` and code identifiers/comments (for example the `ATTACK_LANE` render var) are unchanged, since these words are expected or load-bearing there.
|
|
12
|
+
|
|
13
|
+
## [0.1.16] - 2026-09-11
|
|
14
|
+
|
|
15
|
+
### Added
|
|
16
|
+
|
|
17
|
+
- **A third claude-code-only hook, `route-metrics.mjs`, answers "is my agent actually routing and delegating?"** A routing rule nobody measures is a rule nobody knows is followed. Wired to five events (`UserPromptSubmit`, `PreToolUse` on `Agent`/`Task`, `SubagentStart`, `SubagentStop`, `Stop`), it appends one JSON line per event to `~/.ai-orchestrator/route-metrics.jsonl` (the same directory and home resolution `bin/cli-run.mjs` already logs to): a turn, a subagent dispatch (`subagent_type`, background flag), a subagent start and stop (so a duration can be computed from a small state file keyed by `sha256(agent_id)`), and the lane parsed from a new hidden marker, `<!-- route: <lane> | <why> -->`, that the route-gate block now asks every reply to end with. Only named, charset-bounded fields ever reach the log; prompt text, tool descriptions, the raw assistant message, and the marker's "why" half never do. `node .claude/hooks/route-metrics.mjs --summary [--since <ISO date>]` reports turns, route-marker coverage, lanes by count, dispatches by `subagent_type`, dispatches with no matching start, and mean/max duration per agent type. Plain Node, zero deps, prints nothing to stdout on any event, fail-open (a miss is a missing log line, never a blocked turn). Installed and wired only when claude-code is the primary, same no-overwrite rules as the other two hooks. See `docs/audit-brief.md` for the full threat-model writeup.
|
|
18
|
+
- **CI now runs on `windows-latest` too, node 18/20/22, alongside Ubuntu and macOS.** `defaults.run.shell: bash` makes every workflow step Git Bash on the Windows runner instead of the default `pwsh`, so the same script runs on all three OSes with no parallel Windows rewrite.
|
|
19
|
+
|
|
20
|
+
### Fixed
|
|
21
|
+
|
|
22
|
+
- **`finding-verifier` and `code-reviewer` (claude-code) called themselves unqualified "Read-only" in their descriptions while carrying an unrestricted `Bash` grant**, the same overclaim `done-verifier` shipped with in 0.1.15 and was fixed there; nothing in that grant stops either from running a mutating command. Both descriptions and bodies now say plainly that they carry no file-editing tools and that Bash is bound by the prompt, not the tool grant. `templates/agents/claude-code/README.md` and `templates/agents/snippets/claude-code.md` are corrected the same way. The done-verifier-only test is replaced with one that walks every claude-code agent file: any agent whose `tools:` line includes `Bash` must qualify any "Read-only" claim, checked against its own file and against every generated doc surface.
|
|
23
|
+
- **`CLAUDE.md` and `CONTRIBUTING.md` each pinned a fixed test count that drifted the moment a test was added or removed**, the same class of drift `AGENTS.md` already avoided by saying the suite prints the current number instead. Both now say the same thing `AGENTS.md` does. A new test in `test/prose.test.js` fails if any top-level `.md` states a fixed count of test cases.
|
|
24
|
+
- **`which()` (`src/detect.js`, and its standalone copy in `bin/cli-run.mjs`) never found a Windows lane binary, because a PATH entry there never holds a bare `grok`: npm and vendor installers drop `grok.cmd` (or `.exe`/`.bat`), the same way any Windows shell resolves a bare command through `%PATHEXT%`.** `which()` now tries the bare name first (a no-op on POSIX, and still matching an already-extensioned name on Windows), then each `%PATHEXT%` suffix. `platform` is a parameter (default `process.platform`), the same pattern `killTree(pid, platform, deps)` already used, so the win32 branch has a real test (`test/detect.test.js`, new) on every OS this suite runs on. `home` is a parameter too, for the same testability reason: a dev machine with a real vendor CLI already on `~/.local/bin` made the new win32 tests false-negative until it was injectable.
|
|
25
|
+
- **Committed text files were checked out as CRLF on `windows-latest`, breaking every test that compares a file's exact bytes against a string built in memory with `\n`** (`docs/catalog.md` vs. `catalogMarkdown()`, the README's generated vendor-table section, `llms.txt`'s first line, and a backslash-continued shell example's line-rejoin regex in a test). New `.gitattributes` (`* text=auto eol=lf`) forces LF on checkout regardless of a contributor's or a runner's `core.autocrlf`; `test/fixtures/` (real captured vendor output, byte-exact on purpose) is marked `-text` so line-ending normalization never touches it.
|
|
26
|
+
- **A large block of `test/install.test.js` and `test/cli.test.js` compared a planned file's `f.rel` (built with `path.join`, so backslash-separated on win32) against a hardcoded forward-slash literal** (`'protocols/build-protocol.md'`, `'vm/README.md'`, `'.claude/agents/deep-planner.md'`, and about a dozen more), which is never equal on Windows; several other assertions hardcoded a POSIX `:` `PATH` delimiter and `/usr/bin`, `/bin` absolute paths, or replaced a spawned child's `env` outright and dropped `PATH`/`USERPROFILE`/the rest of the parent environment Windows itself needs. Path literals now go through `join(...)`; PATH construction goes through `node:path`'s `delimiter`; every replaced `env` object spreads `process.env` first and sets `USERPROFILE` alongside `HOME` (`os.homedir()` does not consult `HOME` on win32).
|
|
27
|
+
- **On a real Windows install, a reconfigure's "applied:" line always said "nothing" and the existing-runtime upgrade/conflict checks never fired**, because `bin/cli.js` checked a written file's `f.rel` (backslash-separated on win32, built by `path.join`) directly against `MACHINE_OWNED`/`RUNTIME` (`src/install.js`, hand-written with forward slashes), which never match on that OS. `fileClass()` already normalized before checking; it now exports that normalizer (`toPosixRel`) for `bin/cli.js`'s two direct checks to use too. Both take an optional `separator` parameter (default the real `path.sep`) so the win32 case has a test (`test/install.test.js`) provable from any host. Found via `test/cli.test.js`'s "#6: rerunning with an added lane..." on windows-latest CI.
|
|
28
|
+
- **A replaced child `env` object could carry both `PATH` and the host's own differently-cased `Path` key at once** (`{ ...process.env, PATH: x }` adds a new key next to whichever case the real environment block used; Windows env vars are case-insensitive, plain JS object keys are not), so which one a spawned child actually saw was implementation-defined, not last-key-wins. `test/cli.test.js`'s `mergeEnv()` replaces any existing case-variant of an overridden key instead of adding a second one.
|
|
29
|
+
- **`test/cli.test.js`'s fake lane binaries are `#!/bin/sh` scripts, which Windows cannot execute as `argv[0]`** (no shebang interpretation in `CreateProcess`). `writeShellStub()`/`writeNodeStub()` write a `.cmd` launcher beside the POSIX file on win32 that hands off to Git Bash's `sh.exe` or `node`; `which()`'s detection of the resulting binary works, but running one all the way through `cli-run.mjs`'s own spawn path is not yet reliable on Windows CI for a reason this pass did not fully root-cause. Those specific tests (and the `#12` upgrade-path pair, which diverges in its own, separately unclear way) are skipped on win32 with a stated reason rather than shipped flaky or silently broken; `killTree`'s win32 branch and `which()`'s `%PATHEXT%` resolution each keep their own direct, passing test. A handful of other tests assert an exact POSIX `--dir`/`--project` path string verbatim in output, which `path.resolve()` reinterprets as drive-relative on win32 (a real design question - should a level-3 `--dir` describing a remote Linux box's path ever go through the local host's path semantics at all? - out of scope to decide here) and are skipped the same way. `statSync().mode`'s executable bit is a POSIX-only assertion, dropped on win32 rather than asserted against a filesystem that has no equivalent concept.
|
|
30
|
+
|
|
7
31
|
## [0.1.15] - 2026-09-10
|
|
8
32
|
|
|
9
33
|
The portable parts of a live routing revision, delegate by default, gated on one verified fact rather than a guess: [code.claude.com/docs/en/sub-agents](https://code.claude.com/docs/en/sub-agents) states that a non-fork Claude Code subagent's initial context includes "every level of the CLAUDE.md hierarchy the main conversation loads", and that the built-in Explore and Plan agents skip it. No other lane in this catalog has that documented, so everything below is gated on `subagentsLoadRules(primary)`, currently true for claude-code alone; every other primary keeps its original wording unchanged.
|
|
@@ -238,7 +262,9 @@ First release.
|
|
|
238
262
|
- Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
|
|
239
263
|
- Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
|
|
240
264
|
|
|
241
|
-
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.
|
|
265
|
+
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.17...HEAD
|
|
266
|
+
[0.1.17]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.16...v0.1.17
|
|
267
|
+
[0.1.16]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.15...v0.1.16
|
|
242
268
|
[0.1.15]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...v0.1.15
|
|
243
269
|
[0.1.14]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...v0.1.14
|
|
244
270
|
[0.1.13]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.12...v0.1.13
|
package/README.md
CHANGED
|
@@ -44,7 +44,7 @@ Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/pa
|
|
|
44
44
|
| Id | What | Level |
|
|
45
45
|
|---|---|---|
|
|
46
46
|
| `claude-code` | Claude Code CLI, the default orchestrator | 1+ |
|
|
47
|
-
| `codex` | Codex CLI on a ChatGPT plan: second coder
|
|
47
|
+
| `codex` | Codex CLI on a ChatGPT plan: second coder and second-opinion reviewer (a different model family reading your diff) | 1+ |
|
|
48
48
|
| `agy` | Antigravity CLI on a Google AI plan: research sweeps, concurrent fan-out | 1+ |
|
|
49
49
|
| `grok` | Grok CLI on X Premium: live X and web reads at $0 | 1+ |
|
|
50
50
|
| `hermes` | Hermes Agent: the free tier | 2+ |
|
|
@@ -72,7 +72,7 @@ An install has two targets, and a scripted run should set both.
|
|
|
72
72
|
| Flag | Default | What lands there |
|
|
73
73
|
|---|---|---|
|
|
74
74
|
| `--dir` | `./ai-orchestrator` | the docs, protocols and (level 2+) `bin/cli-run.mjs`. Named after what it contains, not after this package, so a project can hold one without looking like a checkout of it. Pass `--dir ./model-orchestrator` if you prefer the package name. |
|
|
75
|
-
| `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets
|
|
75
|
+
| `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets three hook scripts in `.claude/hooks/`, wired by a settings snippet you merge yourself. |
|
|
76
76
|
|
|
77
77
|
`--project` defaulting to the current directory is the one that surprises people: run the command from your home folder with Claude Code as the primary and five agent files land in your home folder. The installer prints the resolved project path in the plan and says when you left it at the default. Set it.
|
|
78
78
|
|
|
@@ -100,7 +100,7 @@ ai-orchestrator/
|
|
|
100
100
|
protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record
|
|
101
101
|
CODECALC.md OBSIDIAN-TC.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
|
|
102
102
|
<project>/.claude/agents/ one per tier plus finding-verifier, done-verifier, reader, at the PROJECT root (if Claude Code is primary)
|
|
103
|
-
<project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart), Claude Code only
|
|
103
|
+
<project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart) + route-metrics.mjs (all five: see "Measuring routing" below), Claude Code only
|
|
104
104
|
CLAUDE.snippet.md the block to paste into your CLAUDE.md
|
|
105
105
|
settings.hooks.snippet.json the hooks block to merge into .claude/settings.json (Claude Code only)
|
|
106
106
|
ROUTING.md multi-lane decision tree (level 2+)
|
|
@@ -117,7 +117,7 @@ ai-orchestrator/
|
|
|
117
117
|
| [`src/`](src/README.md) | the catalog, the pure planner, detection, rendering |
|
|
118
118
|
| [`templates/`](templates/README.md) | everything the installer can write, by level, plus `tools/` for companions |
|
|
119
119
|
| [`docs/`](docs/README.md) | the three parts and the catalog |
|
|
120
|
-
| [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu and
|
|
120
|
+
| [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu, macOS and Windows, Node 18/20/22 |
|
|
121
121
|
|
|
122
122
|
## What is enforced, what is delegated, what is an instruction
|
|
123
123
|
|
|
@@ -161,7 +161,7 @@ One number per lane, and it is the same number the installer pins: where a lane
|
|
|
161
161
|
|
|
162
162
|
**The live canary runs on your machine, with your credentials.** That is what `node bin/cli-run.mjs --doctor --run` is: it sends every enabled lane one tiny prompt through your own sign-ins and reports `canary ok` or `canary FAILED rc=` per lane. Run it after install, and again after any vendor upgrade.
|
|
163
163
|
|
|
164
|
-
It deliberately does not run in this repository's CI. A canary is only meaningful against real credentials, and there are no credentials a maintainer could supply that would tell **you** anything about **your** lanes: your sign-ins, your quota, your vendor versions. A maintainer-credential canary in CI would prove one machine works and bill someone per run to do it. So CI runs the full suite against stub lanes on Ubuntu and
|
|
164
|
+
It deliberately does not run in this repository's CI. A canary is only meaningful against real credentials, and there are no credentials a maintainer could supply that would tell **you** anything about **your** lanes: your sign-ins, your quota, your vendor versions. A maintainer-credential canary in CI would prove one machine works and bill someone per run to do it. So CI runs the full suite against stub lanes on Ubuntu, macOS and Windows, Node 18/20/22, plus a packaged install into a clean consumer, and the live check ships to you instead.
|
|
165
165
|
|
|
166
166
|
## Principles the whole thing rests on
|
|
167
167
|
|
|
@@ -173,7 +173,7 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
|
|
|
173
173
|
6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, only the brief carries what this task needs. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
|
|
174
174
|
7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
|
|
175
175
|
|
|
176
|
-
## Routing by role, complexity and
|
|
176
|
+
## Routing by role, complexity and stakes
|
|
177
177
|
|
|
178
178
|
Role picks the agent. Two more inputs move the choice, and they move it in
|
|
179
179
|
different directions, so `TIERS.md` states them separately rather than folding
|
|
@@ -182,10 +182,14 @@ them into the role:
|
|
|
182
182
|
- **Complexity moves the effort.** A worker executing a finished plan needs less
|
|
183
183
|
reasoning than the reviewer judging its output. When the plan is airtight the
|
|
184
184
|
spec is carrying the thinking.
|
|
185
|
-
- **
|
|
186
|
-
irreversible changes buy the
|
|
187
|
-
human yes. A one-line change to an auth check is simple and high-
|
|
188
|
-
same time, and it is the
|
|
185
|
+
- **Stakes move the tier and the reader.** Security, privacy, data loss and
|
|
186
|
+
irreversible changes buy the challenge lane, a named check, a rollback path or
|
|
187
|
+
a human yes. A one-line change to an auth check is simple and high-stakes at
|
|
188
|
+
the same time, and it is the stakes that decide.
|
|
189
|
+
|
|
190
|
+
Stakes means what a mistake would cost: a security hole, leaked personal data,
|
|
191
|
+
lost data, or something you can't undo. Most tasks are low-stakes and route
|
|
192
|
+
normally.
|
|
189
193
|
|
|
190
194
|
The top of the ladder is bought with evidence: a reproduced failure, an
|
|
191
195
|
unresolved checkpoint, an irreversible change. A task that merely feels hard is
|
|
@@ -205,12 +209,23 @@ itself.
|
|
|
205
209
|
|
|
206
210
|
`done-verifier` probes the artifact a tracker item's done-signal names (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything itself. It carries no file-editing tools, but on claude-code it does carry `Bash` for those probes (`git log`, `grep`, `wc -l`, `test -f`); staying to read-only commands there is a rule in its prompt, not a restriction on the tool grant, and its own description says so. On agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `reader` is the one that is read-only by tool grant on both: no `Write`, `Edit`, or `Bash`. It reads and digests many files or notes and hands back exactly what the brief asked for, cited by `path:line`; it never classifies, tags or writes, which is what separates it from `bulk-worker`. Both ship in the claude-code and agy agent sets, at the fast tier.
|
|
207
211
|
|
|
212
|
+
## Measuring routing
|
|
213
|
+
|
|
214
|
+
A routing rule nobody measures is a rule nobody knows is followed. On a claude-code install, `route-metrics.mjs` (the third hook, wired to `UserPromptSubmit`, `PreToolUse` on `Agent`/`Task`, `SubagentStart`, `SubagentStop` and `Stop`) turns each of those into one JSON line under `~/.ai-orchestrator/route-metrics.jsonl`: a turn started, a subagent was dispatched (and with what, and in the background or not), a subagent started and stopped (so a duration can be computed), and the lane your agent named in its own hidden `<!-- route: <lane> | <why> -->` marker, which the route-gate block now asks for on every reply. It never logs prompt text, tool descriptions, or the "why" half of the marker: only the named fields above, charset-bounded, same principle as `cli-run.mjs`'s log.
|
|
215
|
+
|
|
216
|
+
```bash
|
|
217
|
+
node .claude/hooks/route-metrics.mjs --summary # since the log began
|
|
218
|
+
node .claude/hooks/route-metrics.mjs --summary --since 2026-09-01 # since a date
|
|
219
|
+
```
|
|
220
|
+
|
|
221
|
+
The report prints turns, **route-marker coverage** (the percentage of turns whose `Stop` event carried a real lane, not `missing`, which is the number that answers "is the agent actually tagging its routing decisions?"), lanes by count, dispatches by `subagent_type`, dispatches with no matching start (a hook or guard blocked the subagent before it launched), and mean/max duration per agent type. Fail-open by design, like the other two hooks: a miss here is a missing log line, never a blocked turn, and it prints nothing to stdout on any event since stdout on `UserPromptSubmit`/`SubagentStart` becomes model context.
|
|
222
|
+
|
|
208
223
|
## Pin the route, or know that you did not
|
|
209
224
|
|
|
210
225
|
A lane with no `--model`, no `--effort` and no `defaults` entry in
|
|
211
226
|
`bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
|
|
212
227
|
CLI configured months ago at a low reasoning effort keeps auditing at that
|
|
213
|
-
effort while your routing docs describe
|
|
228
|
+
effort while your routing docs describe a second-opinion pass.
|
|
214
229
|
|
|
215
230
|
```bash
|
|
216
231
|
node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
|
|
@@ -233,7 +248,7 @@ Install for the tools you have, then let the generated `ROUTING.md` decide the t
|
|
|
233
248
|
|
|
234
249
|
### How do I route tasks to cheaper models?
|
|
235
250
|
|
|
236
|
-
The rules route by role, complexity and
|
|
251
|
+
The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly never gets a frontier token.
|
|
237
252
|
|
|
238
253
|
### Is this an LLM router or an AI gateway?
|
|
239
254
|
|
|
@@ -245,7 +260,7 @@ Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run`
|
|
|
245
260
|
|
|
246
261
|
## Requirements
|
|
247
262
|
|
|
248
|
-
Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows
|
|
263
|
+
Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22). Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested there; a named set of tests is skipped on Windows, each with its reason in the test file, mainly running a lane end to end through `cli-run`, so treat lane execution on Windows as unproven until someone reports otherwise.
|
|
249
264
|
|
|
250
265
|
**Privacy.** The installer sends no telemetry and makes no network call of its own once it is running. Two things around that are worth being exact about:
|
|
251
266
|
|
|
@@ -261,7 +276,7 @@ Add an AI to `src/catalog.js` and every prompt, table, config and doc picks it u
|
|
|
261
276
|
## Credits
|
|
262
277
|
|
|
263
278
|
- [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for
|
|
264
|
-
separating role, complexity and
|
|
279
|
+
separating role, complexity and stakes instead of compressing them into one
|
|
265
280
|
scale, for recording the model and effort a lane was actually asked for, and
|
|
266
281
|
for verifying findings before they trigger repairs. All three shipped in
|
|
267
282
|
0.1.14.
|
package/bin/cli-run.mjs
CHANGED
|
@@ -75,18 +75,33 @@ export const REASONS = new Set([
|
|
|
75
75
|
|
|
76
76
|
const LOG = join(homedir(), '.ai-orchestrator', 'cli-run.log.jsonl');
|
|
77
77
|
|
|
78
|
+
// On win32, a PATH entry never holds a bare "grok": npm and vendor installers
|
|
79
|
+
// drop "grok.cmd" (or .exe/.bat/.ps1), the same way any Windows shell resolves
|
|
80
|
+
// a bare command through %PATHEXT%. Trying the bare name first keeps this a
|
|
81
|
+
// no-op on POSIX and matches an already-extensioned name (a .exe someone put
|
|
82
|
+
// on PATH directly) on Windows too. Kept in sync with src/detect.js's which(),
|
|
83
|
+
// which this file cannot import: it ships standalone into a user's install.
|
|
84
|
+
function candidateExtensions() {
|
|
85
|
+
if (process.platform !== 'win32') return [''];
|
|
86
|
+
const pathext = process.env.PATHEXT || '.COM;.EXE;.BAT;.CMD';
|
|
87
|
+
return ['', ...pathext.split(';').filter(Boolean)];
|
|
88
|
+
}
|
|
89
|
+
|
|
78
90
|
function which(bin) {
|
|
79
91
|
const dirs = (process.env['PATH'] || '').split(delimiter).filter(Boolean);
|
|
80
92
|
const home = homedir();
|
|
81
93
|
dirs.push(join(home, '.local', 'bin'), join(home, '.grok', 'bin'), join(home, '.npm-global', 'bin'));
|
|
94
|
+
const exts = candidateExtensions();
|
|
82
95
|
for (const d of dirs) {
|
|
83
|
-
const
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
96
|
+
for (const ext of exts) {
|
|
97
|
+
const p = join(d, bin + ext);
|
|
98
|
+
try {
|
|
99
|
+
if (!statSync(p).isFile()) continue; // a directory named like the binary is not the binary
|
|
100
|
+
accessSync(p, constants.X_OK);
|
|
101
|
+
return p;
|
|
102
|
+
} catch {
|
|
103
|
+
/* next */
|
|
104
|
+
}
|
|
90
105
|
}
|
|
91
106
|
}
|
|
92
107
|
return null;
|
package/bin/cli.js
CHANGED
|
@@ -10,7 +10,7 @@ import { spawnSync } from 'node:child_process';
|
|
|
10
10
|
import { resolve, join } from 'node:path';
|
|
11
11
|
import { which } from '../src/detect.js';
|
|
12
12
|
import { AIS, LEVELS, TOOLS, PROVIDERS, aisForLevel, agentCandidates, byId, npmSpec } from '../src/catalog.js';
|
|
13
|
-
import { planFiles, writeFiles, resolveSelection, resolveTools, resolveApis, dirProblems, readManifest, activationSteps, MACHINE_OWNED, RUNTIME, GENERATOR_VERSION } from '../src/install.js';
|
|
13
|
+
import { planFiles, writeFiles, resolveSelection, resolveTools, resolveApis, dirProblems, readManifest, activationSteps, MACHINE_OWNED, RUNTIME, toPosixRel, GENERATOR_VERSION } from '../src/install.js';
|
|
14
14
|
|
|
15
15
|
// One strict parse. Unknown flags, missing values and duplicates are usage
|
|
16
16
|
// errors (exit 2) before anything is planned, so a typo like --dryy can never
|
|
@@ -310,10 +310,10 @@ async function main() {
|
|
|
310
310
|
if (e && e.code === 'PREFLIGHT') bad(e.message);
|
|
311
311
|
throw e;
|
|
312
312
|
}
|
|
313
|
-
const ownedWritten = written.filter((w) => MACHINE_OWNED.has(w));
|
|
313
|
+
const ownedWritten = written.filter((w) => MACHINE_OWNED.has(toPosixRel(w)));
|
|
314
314
|
console.log(`\nWrote ${written.length} file(s)` + (skipped.length ? `, kept ${skipped.length} existing:` : '.'));
|
|
315
315
|
for (const s of skipped) console.log(' kept ' + s);
|
|
316
|
-
const existingRuntime = files.filter((f) => f.root !== 'project' && RUNTIME.has(f.rel)).length;
|
|
316
|
+
const existingRuntime = files.filter((f) => f.root !== 'project' && RUNTIME.has(toPosixRel(f.rel))).length;
|
|
317
317
|
if (prev || existingRuntime && (upgraded.length || conflicts.length || unverifiable.length) || docsUnverifiable.length) {
|
|
318
318
|
console.log(`\nExisting installation found${prev ? ` (MANIFEST.json from generator ${prev.generatorVersion || 'pre-0.1.1'}, ${prev.generatedAt || 'undated'}; this run is ${GENERATOR_VERSION})` : ' (no MANIFEST.json: it predates 0.1.1)'}.`);
|
|
319
319
|
if (prev) {
|
package/docs/README.md
CHANGED
|
@@ -7,7 +7,7 @@ The three parts, as reading. The installer writes the working files; these expla
|
|
|
7
7
|
| 1 Beginner | you use one LLM or one agent and want it to route well | [part-1-beginner.md](part-1-beginner.md) |
|
|
8
8
|
| 2 Intermediate | you have several AIs and want to call them through their CLIs from one orchestrator | [part-2-intermediate.md](part-2-intermediate.md) |
|
|
9
9
|
| 3 Advanced | you want the whole thing running unattended on a virtual machine | [part-3-advanced.md](part-3-advanced.md) |
|
|
10
|
-
| Audit brief | the
|
|
10
|
+
| Audit brief | the security notes and the two second-opinion audit rounds this shipped with | [audit-brief.md](audit-brief.md) |
|
|
11
11
|
| Catalog | what each AI and companion tool in the installer is for, how it installs, how it signs in | [catalog.md](catalog.md) |
|
|
12
12
|
|
|
13
13
|
Each part ends with "what the installer gives you at this level" so the doc and the files agree.
|
package/docs/audit-brief.md
CHANGED
|
@@ -102,3 +102,14 @@ Three findings reproduced against the 0.1.15 branch before it shipped, none of t
|
|
|
102
102
|
| 1 | HIGH. `readFileSync(0)` in both hooks blocked until stdin reached EOF (`sleep 3 \| ... node route-gate.mjs` still running past 1.5s); `route-gate.mjs` also read the whole rules file into memory before bounding it, so a FIFO planted at the rules path blocked forever on open. | Stdin is drained asynchronously against a 250ms hard cap in both hooks. `route-gate.mjs` refuses anything that is not `isFile()` via `statSync` before ever calling open, then reads through one fixed 64 KB buffer via `openSync`/`readSync`. Tests: an open, never-closed stdin pipe exits within 1s for both hooks; a FIFO at the rules path returns the fallback instead of hanging; a 200 MB sparse rules file completes in well under a second with output still capped. |
|
|
103
103
|
| 2 | MEDIUM. `done-verifier`'s description, both agent-folder READMEs, and the root README called it "read-only" without qualification, while its claude-code file carries an unrestricted `Bash` grant; nothing in that grant stops it from running a mutating command. | Every one of those surfaces now says plainly that `done-verifier` carries no file-editing tools and that its Bash use is bound by its own prompt, not by the tool grant; `reader` is named as the one that is read-only by tool grant (no Bash) on both formats. |
|
|
104
104
|
| 3 | MEDIUM. Three generated surfaces still stated the pre-0.1.15 premise on a claude-code install: `builder.md`'s description ("... or the main build itself"), `build-protocol.md`'s roles table and its "why the builder does not hand off" note, and `ROUTING.md`'s "Plan big, execute small" line ("the orchestrator executes"). | All three now render through `subagentsLoadRules(primary)`, the same gate the decision tree and "Who builds" already used; every other primary is unchanged. A semantic-regression test asserts a claude-code install contains none of the old phrasing and a codex install still does. |
|
|
105
|
+
|
|
106
|
+
## New in 0.1.16: a third hook, route-metrics.mjs
|
|
107
|
+
|
|
108
|
+
`route-metrics.mjs` ships to `.claude/hooks/` alongside `route-gate.mjs` and `subagent-context.mjs`, only when claude-code is the primary. Same shape as the other two: plain Node, zero deps, mode `0o755`. Unlike them, it is wired to five events at once (`UserPromptSubmit`, `PreToolUse` matched to `Agent|Task`, `SubagentStart`, `SubagentStop`, `Stop`), and it does write, deliberately: one JSON line per event, appended to `~/.ai-orchestrator/route-metrics.jsonl`.
|
|
109
|
+
|
|
110
|
+
- **Reads.** Only its own stdin (the JSON Claude Code sends per event) and, for `--summary`, its own log file. It never reads `transcript_path` even though that field is present on every event: the documented source for the route marker is `last_assistant_message`, and the docs say the transcript can lag, so a hook that read it instead could log a stale or absent marker as if it were current. It never reads the rules file, the task bundle, or any other project file.
|
|
111
|
+
- **Writes.** `~/.ai-orchestrator/route-metrics.jsonl` (append-only, rotated to `.jsonl.1` above 5 MB) and `~/.ai-orchestrator/route-metrics.state/<sha256(agent_id)>.json`, a small file recording a subagent's start time and type so `SubagentStop` can compute a duration; it is deleted on stop, and anything older than 24h is pruned on the next `SubagentStart`. Nothing outside `~/.ai-orchestrator/`. It never writes to stdout: on `UserPromptSubmit` and `SubagentStart`, stdout becomes model context, and this hook has nothing to say there, so it stays silent on every event, not just those two.
|
|
112
|
+
- **What it never logs.** Prompt text, tool descriptions, the full `tool_input`, `last_assistant_message` itself, or the "why" half of a route marker. Only six named fields ever reach a record: `session_id`, `subagent_type`, `agent_type`, and the parsed `lane`, each stripped to `[A-Za-z0-9_.+-]` and capped at 64 characters (128 for `session_id`) before being written, plus the event name and a `duration_s` number it computed itself. This mirrors `bin/cli-run.mjs`'s own log, which stores a fixed reason code and never a provider-supplied string.
|
|
113
|
+
- **Fail-open, on purpose.** Every code path that can fail (a malformed state file, a full disk, a rotation race, invalid JSON on stdin, an unrecognized event) is caught and produces no record rather than a thrown error or a non-zero exit; the process always exits 0. A miss here is a missing line in a telemetry log, never a blocked turn, so there is nothing to gate.
|
|
114
|
+
- **Bounded.** Stdin is drained asynchronously against a combined 1s time cap and 8 MB size cap; a payload that exceeds either is treated as truncated and parsed as nothing, never partially. `--summary` reads the log directly (never spawns anything, never executes a line in it).
|
|
115
|
+
- **Not yet attacked.** Untested here: two processes racing the same rotation at once (a rename plus an append landing on the same file); a state directory with thousands of leaked files from a long-lived session with a crashed hook (pruning runs, but only on `SubagentStart`, so an install that never starts a subagent again would never prune); behavior if `agent_id` collides across two concurrent subagents (sha256 makes this astronomically unlikely, not impossible).
|
package/docs/catalog.md
CHANGED
|
@@ -24,7 +24,7 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
|
|
|
24
24
|
### `codex` · Codex CLI (OpenAI, ChatGPT plan)
|
|
25
25
|
|
|
26
26
|
- **Kind:** agent-cli · **Access:** subscription · **Lane:** A · **Level:** 1+
|
|
27
|
-
- **Wins at:** second coder and
|
|
27
|
+
- **Wins at:** second coder and second-opinion reviewer (a different model family reading your diff)
|
|
28
28
|
- **Install:** `npm install -g @openai/codex@0.153.4`
|
|
29
29
|
- **Sign in:** `codex login` (add `--device-auth` on a machine with no browser)
|
|
30
30
|
- **Reads rules from:** `AGENTS.md`
|
package/docs/part-1-beginner.md
CHANGED
|
@@ -14,7 +14,7 @@ If your agent exposes model choice (Claude Code, Codex, Antigravity), map the ti
|
|
|
14
14
|
|
|
15
15
|
Three cost levers, always together: tier (price per token), token discipline (how many tokens: read only what you will touch, never re-read, deliverables not narration), effort (how hard each call thinks).
|
|
16
16
|
|
|
17
|
-
And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **
|
|
17
|
+
And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **Stakes** move the tier and who reads the result: security, privacy, data loss and irreversible changes are the four worth naming, because none of their failures can be fixed by editing the code afterwards. A one-line change to an auth check is simple and high-stakes at once, and it is the stakes that decide.
|
|
18
18
|
|
|
19
19
|
Robustness first, cost second. You split tiers because the split produces better work.
|
|
20
20
|
|
|
@@ -32,9 +32,9 @@ Modifiers: plan big, execute small · never silently retry a failed attempt at t
|
|
|
32
32
|
|
|
33
33
|
> A gate you cannot fail is not a gate.
|
|
34
34
|
|
|
35
|
-
"Does this look good?" passes every time. "Name
|
|
35
|
+
"Does this look good?" passes every time. "Name what is most likely to go wrong, and what the request as filed missed" can come back empty, which is how you know it worked.
|
|
36
36
|
|
|
37
|
-
Every build gets two checkpoints. **Before writing:** you map what it touches and what could break, then ask the deep tier on the finished map for
|
|
37
|
+
Every build gets two checkpoints. **Before writing:** you map what it touches and what could break, then ask the deep tier on the finished map for one named weak spot and one gap in the request. **After it is green:** a fresh context challenges it (bad input, failing dependency, drift from the plan), allowed to answer CLEAN, every finding reproduced before it reaches you. Then the ship, with a rollback named and an explicit yes. Then the loud negative: re-check the old name everywhere and expect zero.
|
|
38
38
|
|
|
39
39
|
## 4. Every hand-off carries a brief
|
|
40
40
|
|
|
@@ -46,7 +46,7 @@ After anything comprehensive, a fresh turn that hunts for what is **missing**, n
|
|
|
46
46
|
|
|
47
47
|
## 6. Deep research, single agent
|
|
48
48
|
|
|
49
|
-
Plan the sub-questions as their own turn and inspect them before spending anything. Sweep. Then a fresh
|
|
49
|
+
Plan the sub-questions as their own turn and inspect them before spending anything. Sweep. Then a fresh second-opinion turn told to question the premise. Plant one deliberately wrong figure and see whether it corrects it. Mark every claim CONFIRMED / DISAGREEMENT / REPORTED / UNVERIFIED. Agreement is weak evidence; disagreement is the signal.
|
|
50
50
|
|
|
51
51
|
## 7. Numbers and logic are computed, never guessed
|
|
52
52
|
|
|
@@ -13,7 +13,7 @@ Rule: never spend a frontier token on a task a cheap tier finishes correctly. Es
|
|
|
13
13
|
| Lane | Wins at |
|
|
14
14
|
|---|---|
|
|
15
15
|
| the orchestrator (Claude Code, or whichever you chose) | routes, maps, builds, verifies, records; drives the others as CLIs |
|
|
16
|
-
| Codex | second coder and
|
|
16
|
+
| Codex | second coder and second-opinion reviewer: a different model family reading your diff |
|
|
17
17
|
| Antigravity `agy` | deep research sweeps; concurrent fan-out (its subagent call takes an array) |
|
|
18
18
|
| Grok CLI | X and live web reads at $0 (the same search on the API bills per call) |
|
|
19
19
|
| Hermes | the free tier: rough drafts, first-pass summaries, divergent reads, cron jobs |
|
|
@@ -28,15 +28,15 @@ Every agent CLI can report success and deliver nothing. `bin/cli-run.mjs` builds
|
|
|
28
28
|
|
|
29
29
|
Every call goes through it. "This lane is flaky" becomes a query over its log instead of an argument. `node bin/cli-run.mjs --doctor` is the first thing to run after install: enabled lanes, binaries on PATH, the route each lane is pinned to, and with `--run` a one-word canary per lane.
|
|
30
30
|
|
|
31
|
-
There is a second thing a lane can be quietly wrong about. Left unpinned, it runs on **its own config file**, which the runner cannot see: a CLI set up months ago at a low reasoning effort keeps auditing at that effort while your routing docs describe
|
|
31
|
+
There is a second thing a lane can be quietly wrong about. Left unpinned, it runs on **its own config file**, which the runner cannot see: a CLI set up months ago at a low reasoning effort keeps auditing at that effort while your routing docs describe a second-opinion pass, and no error is ever raised. `--model` and `--effort` pin it per call, `defaults` in `bin/lanes.json` pins it per lane, and every run records the value requested and where it came from (`flag`, `lanes.json`, `lane_default`). The log claims no actual: grok reports a model id in its output, the other four lanes report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string the durable log never holds.
|
|
32
32
|
|
|
33
33
|
## 4. Every delegation carries a task bundle, on both surfaces
|
|
34
34
|
|
|
35
|
-
Subagents and CLI lanes are close to the same problem: something that may hold none of your rules, and broad tool access. A Claude Code subagent is the one documented exception, loading the project's CLAUDE.md hierarchy at start, so it keeps the standing rules but not this task's scope; a CLI lane and a fresh chat window get no such credit. The brief (purpose, task class, scope, capabilities, denied actions, conventions, report contract, exit parameters) goes in the prompt or in the file passed to `--brief` either way. If you can, gate it mechanically: a pre-dispatch hook that refuses a brief missing purpose, denied actions or a report contract. On claude-code, a `SubagentStart` hook can inject the essentials (where the rules and the brief format live) automatically; `.claude/hooks/subagent-context.mjs` is the generated example.
|
|
35
|
+
Subagents and CLI lanes are close to the same problem: something that may hold none of your rules, and broad tool access. A Claude Code subagent is the one documented exception, loading the project's CLAUDE.md hierarchy at start, so it keeps the standing rules but not this task's scope; a CLI lane and a fresh chat window get no such credit. The brief (purpose, task class, scope, capabilities, denied actions, conventions, report contract, exit parameters) goes in the prompt or in the file passed to `--brief` either way. If you can, gate it mechanically: a pre-dispatch hook that refuses a brief missing purpose, denied actions or a report contract. On claude-code, a `SubagentStart` hook can inject the essentials (where the rules and the brief format live) automatically; `.claude/hooks/subagent-context.mjs` is the generated example. A third hook, `.claude/hooks/route-metrics.mjs`, turns that same delegation into a measurement instead of an assumption: it logs every turn, dispatch, subagent start/stop and the lane named in the reply's hidden route marker, and `--summary` reports route-marker coverage, dispatches with no matching start, and duration per agent type.
|
|
36
36
|
|
|
37
37
|
## 5. Research: three engines, one triager
|
|
38
38
|
|
|
39
|
-
Fan the same plan to three model families (web sweep,
|
|
39
|
+
Fan the same plan to three model families (web sweep, second-opinion read, live data), each as one `cli-run` call. The orchestrator opens the primary sources itself, marks every claim, and writes the only durable record. Expect one engine to return confident unsourced numerics; downgrade it. Weight the engines that report their own gaps. Count dispositions, not briefs.
|
|
40
40
|
|
|
41
41
|
## 5a. A finding is a claim, not a fact
|
|
42
42
|
|
|
@@ -48,7 +48,7 @@ The second pass is now a different model reading the same artifact, in read-only
|
|
|
48
48
|
|
|
49
49
|
## 7. The build protocol, bound to lanes
|
|
50
50
|
|
|
51
|
-
Stage 1 Map: the orchestrator sweeps; CLI lanes critique the map at $0. Stage 2: deep tier,
|
|
51
|
+
Stage 1 Map: the orchestrator sweeps; CLI lanes critique the map at $0. Stage 2: deep tier, one named weak spot and one gap in the request. Stage 4: scanners on the added lines, refuses by default. Stage 5: security-shaped diff → the second coder in read-only audit mode; architecture-shaped → deep tier reviewing build against plan; never both. Two deep checkpoints per build; CLI lanes are uncapped.
|
|
52
52
|
|
|
53
53
|
## 8. Privacy gate
|
|
54
54
|
|
package/llms.txt
CHANGED
|
@@ -24,4 +24,4 @@ Levels: 1 beginner (one agent or chat app), 2 intermediate (several agent CLIs,
|
|
|
24
24
|
## Optional
|
|
25
25
|
|
|
26
26
|
- [Security policy](https://github.com/aunysillyme/model-orchestrator/blob/main/SECURITY.md)
|
|
27
|
-
- [Audit brief](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/audit-brief.md): the
|
|
27
|
+
- [Audit brief](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/audit-brief.md): the security notes and what has already been security-reviewed
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "model-orchestrator",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.17",
|
|
4
4
|
"description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
package/src/catalog.js
CHANGED
|
@@ -87,7 +87,7 @@ export const AIS = [
|
|
|
87
87
|
bin: 'codex',
|
|
88
88
|
access: 'subscription',
|
|
89
89
|
lane: 'A',
|
|
90
|
-
role: 'second coder and
|
|
90
|
+
role: 'second coder and second-opinion reviewer (a different model family reading your diff)',
|
|
91
91
|
minLevel: 1,
|
|
92
92
|
install: { npm: '@openai/codex', pin: '0.153.4' },
|
|
93
93
|
builtAgainst: '0.153.4',
|
package/src/detect.js
CHANGED
|
@@ -2,24 +2,43 @@ import { accessSync, statSync, constants } from 'node:fs';
|
|
|
2
2
|
import { delimiter, join } from 'node:path';
|
|
3
3
|
import { homedir } from 'node:os';
|
|
4
4
|
|
|
5
|
+
// On win32, a PATH entry never holds a bare "grok": npm and vendor installers
|
|
6
|
+
// drop "grok.cmd" (or .exe/.bat/.ps1), the same way any Windows shell resolves
|
|
7
|
+
// a bare command through %PATHEXT%. Trying the bare name first keeps this a
|
|
8
|
+
// no-op on POSIX and matches an already-extensioned name (a .exe someone put
|
|
9
|
+
// on PATH directly) on Windows too.
|
|
10
|
+
// platform is a parameter (default process.platform), not a hardcoded read,
|
|
11
|
+
// so the win32 branch has a test on every OS this suite runs on: the same
|
|
12
|
+
// pattern bin/cli-run.mjs's killTree(pid, platform, deps) already uses.
|
|
13
|
+
export function candidateExtensions(platform = process.platform, pathext = process.env.PATHEXT) {
|
|
14
|
+
if (platform !== 'win32') return [''];
|
|
15
|
+
return ['', ...(pathext || '.COM;.EXE;.BAT;.CMD').split(';').filter(Boolean)];
|
|
16
|
+
}
|
|
17
|
+
|
|
5
18
|
// PATH lookup plus the handful of places vendor installers drop binaries
|
|
6
19
|
// without touching PATH. Never a shell function, never a shell out.
|
|
7
20
|
// A directory with the binary's name is not a binary (X_OK passes on
|
|
8
21
|
// searchable directories), so the candidate must be a regular file.
|
|
9
|
-
|
|
22
|
+
// home is a parameter too (default homedir()), for the same reason platform
|
|
23
|
+
// is: a dev machine with a real ~/.local/bin/grok on it made the win32 tests
|
|
24
|
+
// here false-negative until this was injectable, since PATH alone was never
|
|
25
|
+
// the whole search.
|
|
26
|
+
export function which(bin, platform = process.platform, home = homedir()) {
|
|
10
27
|
if (!bin) return null;
|
|
11
|
-
const home = homedir();
|
|
12
28
|
const searchPath = process.env['PATH'] || '';
|
|
13
29
|
const dirs = searchPath.split(delimiter).filter(Boolean);
|
|
14
30
|
dirs.push(join(home, '.local', 'bin'), join(home, '.grok', 'bin'), join(home, '.npm-global', 'bin'));
|
|
31
|
+
const exts = candidateExtensions(platform);
|
|
15
32
|
for (const d of dirs) {
|
|
16
|
-
const
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
33
|
+
for (const ext of exts) {
|
|
34
|
+
const p = join(d, bin + ext);
|
|
35
|
+
try {
|
|
36
|
+
if (!statSync(p).isFile()) continue;
|
|
37
|
+
accessSync(p, constants.X_OK);
|
|
38
|
+
return p;
|
|
39
|
+
} catch {
|
|
40
|
+
/* keep looking */
|
|
41
|
+
}
|
|
23
42
|
}
|
|
24
43
|
}
|
|
25
44
|
return null;
|
package/src/install.js
CHANGED
|
@@ -155,21 +155,21 @@ export function laneVars(selected) {
|
|
|
155
155
|
if (has('hermes')) step0.push(`${cr('hermes')} (the free tier) for rough drafts and divergent reads`);
|
|
156
156
|
if (has('qwen')) step0.push(`${cr('qwen')} (the cheapest metered lane) for structured bulk, never for anything citing a line, number or source`);
|
|
157
157
|
if (has('grok')) step0.push(`${cr('grok')} for X and live web reads at $0`);
|
|
158
|
-
if (has('codex')) step0.push(`${cr('codex --audit')} for
|
|
158
|
+
if (has('codex')) step0.push(`${cr('codex --audit')} for a second-opinion read by a second model family`);
|
|
159
159
|
if (has('agy')) step0.push(`${cr('agy')} for research sweeps and concurrent fan-out`);
|
|
160
160
|
const stage1 = [];
|
|
161
|
-
if (has('codex')) stage1.push(`${cr('codex')} for
|
|
161
|
+
if (has('codex')) stage1.push(`${cr('codex')} for a second-opinion critique of the map`);
|
|
162
162
|
if (has('grok')) stage1.push(`${cr('grok')} to verify current API behaviour instead of trusting recall`);
|
|
163
163
|
if (has('hermes')) stage1.push(`${cr('hermes')} for a divergent read`);
|
|
164
164
|
if (has('agy')) stage1.push(`${cr('agy')} for a wide sweep of prior art`);
|
|
165
165
|
const examples = [];
|
|
166
166
|
examples.push(has('grok') ? `| "What is trending on X today" | ${cr('grok')} |` : '| "What is trending on X today" | live-researcher (standard tier with web tools) |');
|
|
167
|
-
examples.push(has('codex') ? `| "Audit this auth diff" | ${cr('codex --audit')} |` : '| "Audit this auth diff" | code-reviewer at deep tier, in a fresh context told to
|
|
167
|
+
examples.push(has('codex') ? `| "Audit this auth diff" | ${cr('codex --audit')} |` : '| "Audit this auth diff" | code-reviewer at deep tier, in a fresh context told to challenge |');
|
|
168
168
|
examples.push(has('qwen') ? `| "Classify these 200 items" | bulk-worker, or ${cr('qwen')} if the items may leave the machine |` : '| "Classify these 200 items" | bulk-worker |');
|
|
169
|
-
examples.push(enabled.length >= 2 ? '| "Research this topic properly" | several engines in parallel, see `RESEARCH_TRIAGE.md` |' : '| "Research this topic properly" | deep tier plans, standard tier sweeps, a fresh context
|
|
169
|
+
examples.push(enabled.length >= 2 ? '| "Research this topic properly" | several engines in parallel, see `RESEARCH_TRIAGE.md` |' : '| "Research this topic properly" | deep tier plans, standard tier sweeps, a fresh context challenges; see `RESEARCH_TRIAGE.md` |');
|
|
170
170
|
const roles = [];
|
|
171
171
|
if (has('agy')) roles.push('| Web sweep | `cli-run agy` | widest landscape pass |');
|
|
172
|
-
if (has('codex')) roles.push('|
|
|
172
|
+
if (has('codex')) roles.push('| Second-opinion read | `cli-run codex --audit` | question the premise, hunt for what the others would get wrong |');
|
|
173
173
|
if (has('grok')) roles.push('| Live data | `cli-run grok` | dated primary sources, real-time reads |');
|
|
174
174
|
if (has('hermes')) roles.push('| Cheap divergent read | `cli-run hermes` | another opinion at $0 |');
|
|
175
175
|
if (has('qwen')) roles.push('| Structured extraction | `cli-run qwen` | pull the facts into a table; never trust its citations without a check |');
|
|
@@ -183,12 +183,12 @@ export function laneVars(selected) {
|
|
|
183
183
|
return {
|
|
184
184
|
LANE_STEP0: step0.length ? step0.map((l) => ' - ' + l).join('\n') : ' - none selected yet: every task stays on your primary agent\'s tiers until you add a lane (re-run the installer with more AIs)',
|
|
185
185
|
STAGE1_LANES: stage1.length ? '; ' + stage1.join(', ') : '',
|
|
186
|
-
ATTACK_LANE: has('codex') ? '`cli-run codex --audit` (a second model family in a read-only sandbox)' : 'code-reviewer at deep tier, in a fresh context told to
|
|
186
|
+
ATTACK_LANE: has('codex') ? '`cli-run codex --audit` (a second model family in a read-only sandbox)' : 'code-reviewer at deep tier, in a fresh context told to challenge and allowed to answer CLEAN',
|
|
187
187
|
LIVE_LANE: has('grok') ? '`cli-run grok` first ($0), then' : '',
|
|
188
188
|
BULK_LANE: has('qwen') ? ', or `cli-run qwen` if the data may leave your machine' : has('hermes') ? ', or `cli-run hermes` for a free rough pass' : '',
|
|
189
189
|
LANE_EXAMPLES: examples.join('\n'),
|
|
190
190
|
RESEARCH_ROLES: roles.join('\n'),
|
|
191
|
-
RESEARCH_RUN: run.length ? run.join('\n') : '# no cli-run lane selected: run the sweep on your primary agent, then a fresh
|
|
191
|
+
RESEARCH_RUN: run.length ? run.join('\n') : '# no cli-run lane selected: run the sweep on your primary agent, then a fresh second-opinion turn (protocols/deep-research.md, level 1 shape)',
|
|
192
192
|
RESEARCH_ENGINES: String(run.length)
|
|
193
193
|
};
|
|
194
194
|
}
|
|
@@ -251,6 +251,8 @@ export function routeGateSection(selected) {
|
|
|
251
251
|
"Stay inline only when: (a) the brief would cost as much as the work itself, (b) the task needs this conversation's own context, (c) it is the human's decision or the final verification of delegated work (a delegate never verifies itself).",
|
|
252
252
|
'',
|
|
253
253
|
'Never: the built-in Explore or Plan agents for rule-bound work (they skip CLAUDE.md). general-purpose taking work a named agent already owns.',
|
|
254
|
+
'',
|
|
255
|
+
'End every reply with a hidden marker: `<!-- route: <lane> | <why, a few words> -->`. The route-metrics hook reads only the lane out of it, so routing coverage can be measured instead of assumed.',
|
|
254
256
|
'<!-- route-gate:end -->'
|
|
255
257
|
].join('\n');
|
|
256
258
|
}
|
|
@@ -360,7 +362,7 @@ export function activationSteps(opts) {
|
|
|
360
362
|
if (primary && primary.agentsDir) steps.push(`subagents are in ${join(projectAbs, primary.agentsDir)}; run ${primary.bin} from ${projectAbs} to pick them up`);
|
|
361
363
|
// Only claude-code ships hooks (route-gate, subagent-context): the wiring
|
|
362
364
|
// lives in a snippet, never written into a settings.json the user already has.
|
|
363
|
-
if (subagentsLoadRules(primary)) steps.push(`merge the hooks in ${join(dirAbs, 'settings.hooks.snippet.json')} into ${join(projectAbs, '.claude', 'settings.json')} (create it if missing) to wire the route-gate
|
|
365
|
+
if (subagentsLoadRules(primary)) steps.push(`merge the hooks in ${join(dirAbs, 'settings.hooks.snippet.json')} into ${join(projectAbs, '.claude', 'settings.json')} (create it if missing) to wire the route-gate, subagent-context and route-metrics hooks`);
|
|
364
366
|
for (const a of selected.filter((a) => a.bin && a.kind === 'agent-cli')) steps.push(`sign in to ${a.name}: ${a.auth}`);
|
|
365
367
|
// A local runtime has a bin but no sign-in, so the agent-cli loop above skips it
|
|
366
368
|
// and before this it appeared in no ordered list at any level (#26).
|
|
@@ -545,6 +547,11 @@ export function planFiles(opts) {
|
|
|
545
547
|
// written into a settings.json they already have.
|
|
546
548
|
add(join('.claude', 'hooks', 'route-gate.mjs'), render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'route-gate.mjs'), 'utf8'), v), 0o755, 'project');
|
|
547
549
|
add(join('.claude', 'hooks', 'subagent-context.mjs'), render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'subagent-context.mjs'), 'utf8'), v), 0o755, 'project');
|
|
550
|
+
// route-metrics.mjs (0.1.16), claude-code only: five events (UserPromptSubmit,
|
|
551
|
+
// PreToolUse on Agent|Task, SubagentStart, SubagentStop, Stop) turned into one
|
|
552
|
+
// JSON line each under ~/.ai-orchestrator/, so a routing rule nobody measures
|
|
553
|
+
// is not a rule nobody knows is followed.
|
|
554
|
+
add(join('.claude', 'hooks', 'route-metrics.mjs'), render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'route-metrics.mjs'), 'utf8'), v), 0o755, 'project');
|
|
548
555
|
add('settings.hooks.snippet.json', render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'settings.hooks.snippet.json'), 'utf8'), v));
|
|
549
556
|
} else if (primary && primary.id === 'agy') {
|
|
550
557
|
for (const f of walk(join(TEMPLATES, 'agents', 'agy'))) {
|
|
@@ -703,8 +710,22 @@ export const RUNTIME = new Set([
|
|
|
703
710
|
'vm/jobs/weekly-audit.service',
|
|
704
711
|
'vm/jobs/weekly-audit.timer'
|
|
705
712
|
]);
|
|
706
|
-
|
|
707
|
-
|
|
713
|
+
// MACHINE_OWNED and RUNTIME are keyed with forward slashes (they read as
|
|
714
|
+
// prose in the comment above them, and every caller needs the same one
|
|
715
|
+
// spelling regardless of host OS); an f.rel or a writeFiles() "written" path
|
|
716
|
+
// is built with path.join, so it is backslash-separated on win32. Both sets
|
|
717
|
+
// must be checked against the SAME normalized form, or a win32 install
|
|
718
|
+
// silently drops bin/lanes.json and every RUNTIME file from set membership
|
|
719
|
+
// (found: bin/cli.js's own "applied:"/existing-runtime checks did exactly
|
|
720
|
+
// that before this was exported for them to use too).
|
|
721
|
+
// separator is a parameter (default the real path.sep) so a test can prove
|
|
722
|
+
// the win32 case from any host, the same pattern which()'s platform
|
|
723
|
+
// parameter already uses.
|
|
724
|
+
export function toPosixRel(rel, separator = sep) {
|
|
725
|
+
return rel.split(separator).join('/');
|
|
726
|
+
}
|
|
727
|
+
export function fileClass(rel, separator = sep) {
|
|
728
|
+
const r = toPosixRel(rel, separator);
|
|
708
729
|
if (MACHINE_OWNED.has(r)) return 'owned';
|
|
709
730
|
if (RUNTIME.has(r)) return 'runtime';
|
|
710
731
|
return 'document';
|
|
@@ -2,4 +2,4 @@
|
|
|
2
2
|
|
|
3
3
|
Antigravity CLI custom agents, one per tier plus three checks (`finding-verifier`, `done-verifier`, `reader`), in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
|
|
4
4
|
|
|
5
|
-
`commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests;
|
|
5
|
+
`commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; deletes and other destructive commands still ask before running) and `off` for the read-only agents: `code-reviewer`, `finding-verifier`, `live-researcher`, `done-verifier`, `reader`. `model` is a tier: `pro` for deep-planner, `flash` for the rest. `done-verifier` and `reader` never write and, with `commandExecutionPolicy: off`, cannot execute any command at all here, mutating or not: unlike its claude-code counterpart, which does carry an unrestricted `Bash` and stays read-only by its prompt rather than by the tool grant, agy's `done-verifier` is mechanically blocked from shelling out and probes artifacts through whatever read or fetch capability it has instead. Neither is `bulk-worker`, which classifies, tags and transforms items and does write.
|
|
@@ -4,7 +4,7 @@ description: Well-specified execution of a bounded sub-part of a build.
|
|
|
4
4
|
model: flash
|
|
5
5
|
subagent: true
|
|
6
6
|
mainAgent: true
|
|
7
|
-
commandExecutionPolicy: auto # standard build/test commands run unattended;
|
|
7
|
+
commandExecutionPolicy: auto # standard build/test commands run unattended; destructive commands, like deletes, still ask before running
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# builder
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: finding-verifier
|
|
3
|
-
description:
|
|
3
|
+
description: Second-opinion verification of review findings; tries to disprove each one and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE. Read-only, never repairs.
|
|
4
4
|
model: flash
|
|
5
5
|
subagent: true
|
|
6
6
|
mainAgent: true
|
|
@@ -12,6 +12,12 @@ commandExecutionPolicy: off
|
|
|
12
12
|
A finding is a claim, not a fact. You try to disprove each one before it is
|
|
13
13
|
allowed to cause a repair.
|
|
14
14
|
|
|
15
|
+
No file-editing tools, and no command execution: this agent's
|
|
16
|
+
`commandExecutionPolicy` is `off`, so unlike its claude-code counterpart,
|
|
17
|
+
which carries an unrestricted `Bash` and stays read-only by its prompt rather
|
|
18
|
+
than by the tool grant, this agent is mechanically blocked from shelling out;
|
|
19
|
+
probe with whatever read or fetch capability you have instead.
|
|
20
|
+
|
|
15
21
|
For each finding you are given: read the cited file and line yourself, state the
|
|
16
22
|
input or sequence that would trigger it, then hunt for what makes it impossible
|
|
17
23
|
(a guard upstream, a caller that never passes that value, an existing test).
|