model-orchestrator 0.1.15 → 0.1.16

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,24 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.16] - 2026-09-11
8
+
9
+ ### Added
10
+
11
+ - **A third claude-code-only hook, `route-metrics.mjs`, answers "is my agent actually routing and delegating?"** A routing rule nobody measures is a rule nobody knows is followed. Wired to five events (`UserPromptSubmit`, `PreToolUse` on `Agent`/`Task`, `SubagentStart`, `SubagentStop`, `Stop`), it appends one JSON line per event to `~/.ai-orchestrator/route-metrics.jsonl` (the same directory and home resolution `bin/cli-run.mjs` already logs to): a turn, a subagent dispatch (`subagent_type`, background flag), a subagent start and stop (so a duration can be computed from a small state file keyed by `sha256(agent_id)`), and the lane parsed from a new hidden marker, `<!-- route: <lane> | <why> -->`, that the route-gate block now asks every reply to end with. Only named, charset-bounded fields ever reach the log; prompt text, tool descriptions, the raw assistant message, and the marker's "why" half never do. `node .claude/hooks/route-metrics.mjs --summary [--since <ISO date>]` reports turns, route-marker coverage, lanes by count, dispatches by `subagent_type`, dispatches with no matching start, and mean/max duration per agent type. Plain Node, zero deps, prints nothing to stdout on any event, fail-open (a miss is a missing log line, never a blocked turn). Installed and wired only when claude-code is the primary, same no-overwrite rules as the other two hooks. See `docs/audit-brief.md` for the full threat-model writeup.
12
+ - **CI now runs on `windows-latest` too, node 18/20/22, alongside Ubuntu and macOS.** `defaults.run.shell: bash` makes every workflow step Git Bash on the Windows runner instead of the default `pwsh`, so the same script runs on all three OSes with no parallel Windows rewrite.
13
+
14
+ ### Fixed
15
+
16
+ - **`finding-verifier` and `code-reviewer` (claude-code) called themselves unqualified "Read-only" in their descriptions while carrying an unrestricted `Bash` grant**, the same overclaim `done-verifier` shipped with in 0.1.15 and was fixed there; nothing in that grant stops either from running a mutating command. Both descriptions and bodies now say plainly that they carry no file-editing tools and that Bash is bound by the prompt, not the tool grant. `templates/agents/claude-code/README.md` and `templates/agents/snippets/claude-code.md` are corrected the same way. The done-verifier-only test is replaced with one that walks every claude-code agent file: any agent whose `tools:` line includes `Bash` must qualify any "Read-only" claim, checked against its own file and against every generated doc surface.
17
+ - **`CLAUDE.md` and `CONTRIBUTING.md` each pinned a fixed test count that drifted the moment a test was added or removed**, the same class of drift `AGENTS.md` already avoided by saying the suite prints the current number instead. Both now say the same thing `AGENTS.md` does. A new test in `test/prose.test.js` fails if any top-level `.md` states a fixed count of test cases.
18
+ - **`which()` (`src/detect.js`, and its standalone copy in `bin/cli-run.mjs`) never found a Windows lane binary, because a PATH entry there never holds a bare `grok`: npm and vendor installers drop `grok.cmd` (or `.exe`/`.bat`), the same way any Windows shell resolves a bare command through `%PATHEXT%`.** `which()` now tries the bare name first (a no-op on POSIX, and still matching an already-extensioned name on Windows), then each `%PATHEXT%` suffix. `platform` is a parameter (default `process.platform`), the same pattern `killTree(pid, platform, deps)` already used, so the win32 branch has a real test (`test/detect.test.js`, new) on every OS this suite runs on. `home` is a parameter too, for the same testability reason: a dev machine with a real vendor CLI already on `~/.local/bin` made the new win32 tests false-negative until it was injectable.
19
+ - **Committed text files were checked out as CRLF on `windows-latest`, breaking every test that compares a file's exact bytes against a string built in memory with `\n`** (`docs/catalog.md` vs. `catalogMarkdown()`, the README's generated vendor-table section, `llms.txt`'s first line, and a backslash-continued shell example's line-rejoin regex in a test). New `.gitattributes` (`* text=auto eol=lf`) forces LF on checkout regardless of a contributor's or a runner's `core.autocrlf`; `test/fixtures/` (real captured vendor output, byte-exact on purpose) is marked `-text` so line-ending normalization never touches it.
20
+ - **A large block of `test/install.test.js` and `test/cli.test.js` compared a planned file's `f.rel` (built with `path.join`, so backslash-separated on win32) against a hardcoded forward-slash literal** (`'protocols/build-protocol.md'`, `'vm/README.md'`, `'.claude/agents/deep-planner.md'`, and about a dozen more), which is never equal on Windows; several other assertions hardcoded a POSIX `:` `PATH` delimiter and `/usr/bin`, `/bin` absolute paths, or replaced a spawned child's `env` outright and dropped `PATH`/`USERPROFILE`/the rest of the parent environment Windows itself needs. Path literals now go through `join(...)`; PATH construction goes through `node:path`'s `delimiter`; every replaced `env` object spreads `process.env` first and sets `USERPROFILE` alongside `HOME` (`os.homedir()` does not consult `HOME` on win32).
21
+ - **On a real Windows install, a reconfigure's "applied:" line always said "nothing" and the existing-runtime upgrade/conflict checks never fired**, because `bin/cli.js` checked a written file's `f.rel` (backslash-separated on win32, built by `path.join`) directly against `MACHINE_OWNED`/`RUNTIME` (`src/install.js`, hand-written with forward slashes), which never match on that OS. `fileClass()` already normalized before checking; it now exports that normalizer (`toPosixRel`) for `bin/cli.js`'s two direct checks to use too. Both take an optional `separator` parameter (default the real `path.sep`) so the win32 case has a test (`test/install.test.js`) provable from any host. Found via `test/cli.test.js`'s "#6: rerunning with an added lane..." on windows-latest CI.
22
+ - **A replaced child `env` object could carry both `PATH` and the host's own differently-cased `Path` key at once** (`{ ...process.env, PATH: x }` adds a new key next to whichever case the real environment block used; Windows env vars are case-insensitive, plain JS object keys are not), so which one a spawned child actually saw was implementation-defined, not last-key-wins. `test/cli.test.js`'s `mergeEnv()` replaces any existing case-variant of an overridden key instead of adding a second one.
23
+ - **`test/cli.test.js`'s fake lane binaries are `#!/bin/sh` scripts, which Windows cannot execute as `argv[0]`** (no shebang interpretation in `CreateProcess`). `writeShellStub()`/`writeNodeStub()` write a `.cmd` launcher beside the POSIX file on win32 that hands off to Git Bash's `sh.exe` or `node`; `which()`'s detection of the resulting binary works, but running one all the way through `cli-run.mjs`'s own spawn path is not yet reliable on Windows CI for a reason this pass did not fully root-cause. Those specific tests (and the `#12` upgrade-path pair, which diverges in its own, separately unclear way) are skipped on win32 with a stated reason rather than shipped flaky or silently broken; `killTree`'s win32 branch and `which()`'s `%PATHEXT%` resolution each keep their own direct, passing test. A handful of other tests assert an exact POSIX `--dir`/`--project` path string verbatim in output, which `path.resolve()` reinterprets as drive-relative on win32 (a real design question - should a level-3 `--dir` describing a remote Linux box's path ever go through the local host's path semantics at all? - out of scope to decide here) and are skipped the same way. `statSync().mode`'s executable bit is a POSIX-only assertion, dropped on win32 rather than asserted against a filesystem that has no equivalent concept.
24
+
7
25
  ## [0.1.15] - 2026-09-10
8
26
 
9
27
  The portable parts of a live routing revision, delegate by default, gated on one verified fact rather than a guess: [code.claude.com/docs/en/sub-agents](https://code.claude.com/docs/en/sub-agents) states that a non-fork Claude Code subagent's initial context includes "every level of the CLAUDE.md hierarchy the main conversation loads", and that the built-in Explore and Plan agents skip it. No other lane in this catalog has that documented, so everything below is gated on `subagentsLoadRules(primary)`, currently true for claude-code alone; every other primary keeps its original wording unchanged.
@@ -238,7 +256,8 @@ First release.
238
256
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
239
257
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
240
258
 
241
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.15...HEAD
259
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.16...HEAD
260
+ [0.1.16]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.15...v0.1.16
242
261
  [0.1.15]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...v0.1.15
243
262
  [0.1.14]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...v0.1.14
244
263
  [0.1.13]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.12...v0.1.13
package/README.md CHANGED
@@ -72,7 +72,7 @@ An install has two targets, and a scripted run should set both.
72
72
  | Flag | Default | What lands there |
73
73
  |---|---|---|
74
74
  | `--dir` | `./ai-orchestrator` | the docs, protocols and (level 2+) `bin/cli-run.mjs`. Named after what it contains, not after this package, so a project can hold one without looking like a checkout of it. Pass `--dir ./model-orchestrator` if you prefer the package name. |
75
- | `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets two hook scripts in `.claude/hooks/`, wired by a settings snippet you merge yourself. |
75
+ | `--project` | the current directory | the subagent definitions, and the rules file your agent reads. Only Claude Code (`.claude/agents/`) and Antigravity (`.agents/agents/`) get files here, because that is the only place those CLIs look. Claude Code also gets three hook scripts in `.claude/hooks/`, wired by a settings snippet you merge yourself. |
76
76
 
77
77
  `--project` defaulting to the current directory is the one that surprises people: run the command from your home folder with Claude Code as the primary and five agent files land in your home folder. The installer prints the resolved project path in the plan and says when you left it at the default. Set it.
78
78
 
@@ -100,7 +100,7 @@ ai-orchestrator/
100
100
  protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record
101
101
  CODECALC.md OBSIDIAN-TC.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
102
102
  <project>/.claude/agents/ one per tier plus finding-verifier, done-verifier, reader, at the PROJECT root (if Claude Code is primary)
103
- <project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart), Claude Code only
103
+ <project>/.claude/hooks/ route-gate.mjs (UserPromptSubmit) + subagent-context.mjs (SubagentStart) + route-metrics.mjs (all five: see "Measuring routing" below), Claude Code only
104
104
  CLAUDE.snippet.md the block to paste into your CLAUDE.md
105
105
  settings.hooks.snippet.json the hooks block to merge into .claude/settings.json (Claude Code only)
106
106
  ROUTING.md multi-lane decision tree (level 2+)
@@ -117,7 +117,7 @@ ai-orchestrator/
117
117
  | [`src/`](src/README.md) | the catalog, the pure planner, detection, rendering |
118
118
  | [`templates/`](templates/README.md) | everything the installer can write, by level, plus `tools/` for companions |
119
119
  | [`docs/`](docs/README.md) | the three parts and the catalog |
120
- | [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu and macOS, Node 18/20/22 |
120
+ | [`test/`](test/README.md) | `npm test`: judges proven to go red, catalog integrity, planner, end-to-end install in a temp dir; `.github/workflows/test.yml` runs it on Ubuntu, macOS and Windows, Node 18/20/22 |
121
121
 
122
122
  ## What is enforced, what is delegated, what is an instruction
123
123
 
@@ -161,7 +161,7 @@ One number per lane, and it is the same number the installer pins: where a lane
161
161
 
162
162
  **The live canary runs on your machine, with your credentials.** That is what `node bin/cli-run.mjs --doctor --run` is: it sends every enabled lane one tiny prompt through your own sign-ins and reports `canary ok` or `canary FAILED rc=` per lane. Run it after install, and again after any vendor upgrade.
163
163
 
164
- It deliberately does not run in this repository's CI. A canary is only meaningful against real credentials, and there are no credentials a maintainer could supply that would tell **you** anything about **your** lanes: your sign-ins, your quota, your vendor versions. A maintainer-credential canary in CI would prove one machine works and bill someone per run to do it. So CI runs the full suite against stub lanes on Ubuntu and macOS, Node 18/20/22, plus a packaged install into a clean consumer, and the live check ships to you instead.
164
+ It deliberately does not run in this repository's CI. A canary is only meaningful against real credentials, and there are no credentials a maintainer could supply that would tell **you** anything about **your** lanes: your sign-ins, your quota, your vendor versions. A maintainer-credential canary in CI would prove one machine works and bill someone per run to do it. So CI runs the full suite against stub lanes on Ubuntu, macOS and Windows, Node 18/20/22, plus a packaged install into a clean consumer, and the live check ships to you instead.
165
165
 
166
166
  ## Principles the whole thing rests on
167
167
 
@@ -205,6 +205,17 @@ itself.
205
205
 
206
206
  `done-verifier` probes the artifact a tracker item's done-signal names (a file, a commit, a URL, a log line, a count) and returns MET, NOT_MET or UNVERIFIABLE; it never closes or edits anything itself. It carries no file-editing tools, but on claude-code it does carry `Bash` for those probes (`git log`, `grep`, `wc -l`, `test -f`); staying to read-only commands there is a rule in its prompt, not a restriction on the tool grant, and its own description says so. On agy, `commandExecutionPolicy: off` blocks command execution mechanically instead. `reader` is the one that is read-only by tool grant on both: no `Write`, `Edit`, or `Bash`. It reads and digests many files or notes and hands back exactly what the brief asked for, cited by `path:line`; it never classifies, tags or writes, which is what separates it from `bulk-worker`. Both ship in the claude-code and agy agent sets, at the fast tier.
207
207
 
208
+ ## Measuring routing
209
+
210
+ A routing rule nobody measures is a rule nobody knows is followed. On a claude-code install, `route-metrics.mjs` (the third hook, wired to `UserPromptSubmit`, `PreToolUse` on `Agent`/`Task`, `SubagentStart`, `SubagentStop` and `Stop`) turns each of those into one JSON line under `~/.ai-orchestrator/route-metrics.jsonl`: a turn started, a subagent was dispatched (and with what, and in the background or not), a subagent started and stopped (so a duration can be computed), and the lane your agent named in its own hidden `<!-- route: <lane> | <why> -->` marker, which the route-gate block now asks for on every reply. It never logs prompt text, tool descriptions, or the "why" half of the marker: only the named fields above, charset-bounded, same principle as `cli-run.mjs`'s log.
211
+
212
+ ```bash
213
+ node .claude/hooks/route-metrics.mjs --summary # since the log began
214
+ node .claude/hooks/route-metrics.mjs --summary --since 2026-09-01 # since a date
215
+ ```
216
+
217
+ The report prints turns, **route-marker coverage** (the percentage of turns whose `Stop` event carried a real lane, not `missing`, which is the number that answers "is the agent actually tagging its routing decisions?"), lanes by count, dispatches by `subagent_type`, dispatches with no matching start (a hook or guard blocked the subagent before it launched), and mean/max duration per agent type. Fail-open by design, like the other two hooks: a miss here is a missing log line, never a blocked turn, and it prints nothing to stdout on any event since stdout on `UserPromptSubmit`/`SubagentStart` becomes model context.
218
+
208
219
  ## Pin the route, or know that you did not
209
220
 
210
221
  A lane with no `--model`, no `--effort` and no `defaults` entry in
@@ -245,7 +256,7 @@ Yes. `--yes` with `--level`, `--ais` and `--project` runs headless, `--dry-run`
245
256
 
246
257
  ## Requirements
247
258
 
248
- Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows is untested: `cli-run` ends a lane's process tree there with `taskkill`, but nothing in CI runs on Windows, so treat it as unsupported until someone reports otherwise.
259
+ Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows: CI runs the suite on `windows-latest` (Node 18, 20, 22). Install, detection, the hooks and `cli-run`'s `taskkill` tree kill are tested there; a named set of tests is skipped on Windows, each with its reason in the test file, mainly running a lane end to end through `cli-run`, so treat lane execution on Windows as unproven until someone reports otherwise.
249
260
 
250
261
  **Privacy.** The installer sends no telemetry and makes no network call of its own once it is running. Two things around that are worth being exact about:
251
262
 
package/bin/cli-run.mjs CHANGED
@@ -75,18 +75,33 @@ export const REASONS = new Set([
75
75
 
76
76
  const LOG = join(homedir(), '.ai-orchestrator', 'cli-run.log.jsonl');
77
77
 
78
+ // On win32, a PATH entry never holds a bare "grok": npm and vendor installers
79
+ // drop "grok.cmd" (or .exe/.bat/.ps1), the same way any Windows shell resolves
80
+ // a bare command through %PATHEXT%. Trying the bare name first keeps this a
81
+ // no-op on POSIX and matches an already-extensioned name (a .exe someone put
82
+ // on PATH directly) on Windows too. Kept in sync with src/detect.js's which(),
83
+ // which this file cannot import: it ships standalone into a user's install.
84
+ function candidateExtensions() {
85
+ if (process.platform !== 'win32') return [''];
86
+ const pathext = process.env.PATHEXT || '.COM;.EXE;.BAT;.CMD';
87
+ return ['', ...pathext.split(';').filter(Boolean)];
88
+ }
89
+
78
90
  function which(bin) {
79
91
  const dirs = (process.env['PATH'] || '').split(delimiter).filter(Boolean);
80
92
  const home = homedir();
81
93
  dirs.push(join(home, '.local', 'bin'), join(home, '.grok', 'bin'), join(home, '.npm-global', 'bin'));
94
+ const exts = candidateExtensions();
82
95
  for (const d of dirs) {
83
- const p = join(d, bin);
84
- try {
85
- if (!statSync(p).isFile()) continue; // a directory named like the binary is not the binary
86
- accessSync(p, constants.X_OK);
87
- return p;
88
- } catch {
89
- /* next */
96
+ for (const ext of exts) {
97
+ const p = join(d, bin + ext);
98
+ try {
99
+ if (!statSync(p).isFile()) continue; // a directory named like the binary is not the binary
100
+ accessSync(p, constants.X_OK);
101
+ return p;
102
+ } catch {
103
+ /* next */
104
+ }
90
105
  }
91
106
  }
92
107
  return null;
package/bin/cli.js CHANGED
@@ -10,7 +10,7 @@ import { spawnSync } from 'node:child_process';
10
10
  import { resolve, join } from 'node:path';
11
11
  import { which } from '../src/detect.js';
12
12
  import { AIS, LEVELS, TOOLS, PROVIDERS, aisForLevel, agentCandidates, byId, npmSpec } from '../src/catalog.js';
13
- import { planFiles, writeFiles, resolveSelection, resolveTools, resolveApis, dirProblems, readManifest, activationSteps, MACHINE_OWNED, RUNTIME, GENERATOR_VERSION } from '../src/install.js';
13
+ import { planFiles, writeFiles, resolveSelection, resolveTools, resolveApis, dirProblems, readManifest, activationSteps, MACHINE_OWNED, RUNTIME, toPosixRel, GENERATOR_VERSION } from '../src/install.js';
14
14
 
15
15
  // One strict parse. Unknown flags, missing values and duplicates are usage
16
16
  // errors (exit 2) before anything is planned, so a typo like --dryy can never
@@ -310,10 +310,10 @@ async function main() {
310
310
  if (e && e.code === 'PREFLIGHT') bad(e.message);
311
311
  throw e;
312
312
  }
313
- const ownedWritten = written.filter((w) => MACHINE_OWNED.has(w));
313
+ const ownedWritten = written.filter((w) => MACHINE_OWNED.has(toPosixRel(w)));
314
314
  console.log(`\nWrote ${written.length} file(s)` + (skipped.length ? `, kept ${skipped.length} existing:` : '.'));
315
315
  for (const s of skipped) console.log(' kept ' + s);
316
- const existingRuntime = files.filter((f) => f.root !== 'project' && RUNTIME.has(f.rel)).length;
316
+ const existingRuntime = files.filter((f) => f.root !== 'project' && RUNTIME.has(toPosixRel(f.rel))).length;
317
317
  if (prev || existingRuntime && (upgraded.length || conflicts.length || unverifiable.length) || docsUnverifiable.length) {
318
318
  console.log(`\nExisting installation found${prev ? ` (MANIFEST.json from generator ${prev.generatorVersion || 'pre-0.1.1'}, ${prev.generatedAt || 'undated'}; this run is ${GENERATOR_VERSION})` : ' (no MANIFEST.json: it predates 0.1.1)'}.`);
319
319
  if (prev) {
@@ -102,3 +102,14 @@ Three findings reproduced against the 0.1.15 branch before it shipped, none of t
102
102
  | 1 | HIGH. `readFileSync(0)` in both hooks blocked until stdin reached EOF (`sleep 3 \| ... node route-gate.mjs` still running past 1.5s); `route-gate.mjs` also read the whole rules file into memory before bounding it, so a FIFO planted at the rules path blocked forever on open. | Stdin is drained asynchronously against a 250ms hard cap in both hooks. `route-gate.mjs` refuses anything that is not `isFile()` via `statSync` before ever calling open, then reads through one fixed 64 KB buffer via `openSync`/`readSync`. Tests: an open, never-closed stdin pipe exits within 1s for both hooks; a FIFO at the rules path returns the fallback instead of hanging; a 200 MB sparse rules file completes in well under a second with output still capped. |
103
103
  | 2 | MEDIUM. `done-verifier`'s description, both agent-folder READMEs, and the root README called it "read-only" without qualification, while its claude-code file carries an unrestricted `Bash` grant; nothing in that grant stops it from running a mutating command. | Every one of those surfaces now says plainly that `done-verifier` carries no file-editing tools and that its Bash use is bound by its own prompt, not by the tool grant; `reader` is named as the one that is read-only by tool grant (no Bash) on both formats. |
104
104
  | 3 | MEDIUM. Three generated surfaces still stated the pre-0.1.15 premise on a claude-code install: `builder.md`'s description ("... or the main build itself"), `build-protocol.md`'s roles table and its "why the builder does not hand off" note, and `ROUTING.md`'s "Plan big, execute small" line ("the orchestrator executes"). | All three now render through `subagentsLoadRules(primary)`, the same gate the decision tree and "Who builds" already used; every other primary is unchanged. A semantic-regression test asserts a claude-code install contains none of the old phrasing and a codex install still does. |
105
+
106
+ ## New in 0.1.16: a third hook, route-metrics.mjs
107
+
108
+ `route-metrics.mjs` ships to `.claude/hooks/` alongside `route-gate.mjs` and `subagent-context.mjs`, only when claude-code is the primary. Same shape as the other two: plain Node, zero deps, mode `0o755`. Unlike them, it is wired to five events at once (`UserPromptSubmit`, `PreToolUse` matched to `Agent|Task`, `SubagentStart`, `SubagentStop`, `Stop`), and it does write, deliberately: one JSON line per event, appended to `~/.ai-orchestrator/route-metrics.jsonl`.
109
+
110
+ - **Reads.** Only its own stdin (the JSON Claude Code sends per event) and, for `--summary`, its own log file. It never reads `transcript_path` even though that field is present on every event: the documented source for the route marker is `last_assistant_message`, and the docs say the transcript can lag, so a hook that read it instead could log a stale or absent marker as if it were current. It never reads the rules file, the task bundle, or any other project file.
111
+ - **Writes.** `~/.ai-orchestrator/route-metrics.jsonl` (append-only, rotated to `.jsonl.1` above 5 MB) and `~/.ai-orchestrator/route-metrics.state/<sha256(agent_id)>.json`, a small file recording a subagent's start time and type so `SubagentStop` can compute a duration; it is deleted on stop, and anything older than 24h is pruned on the next `SubagentStart`. Nothing outside `~/.ai-orchestrator/`. It never writes to stdout: on `UserPromptSubmit` and `SubagentStart`, stdout becomes model context, and this hook has nothing to say there, so it stays silent on every event, not just those two.
112
+ - **What it never logs.** Prompt text, tool descriptions, the full `tool_input`, `last_assistant_message` itself, or the "why" half of a route marker. Only six named fields ever reach a record: `session_id`, `subagent_type`, `agent_type`, and the parsed `lane`, each stripped to `[A-Za-z0-9_.+-]` and capped at 64 characters (128 for `session_id`) before being written, plus the event name and a `duration_s` number it computed itself. This mirrors `bin/cli-run.mjs`'s own log, which stores a fixed reason code and never a provider-supplied string.
113
+ - **Fail-open, on purpose.** Every code path that can fail (a malformed state file, a full disk, a rotation race, invalid JSON on stdin, an unrecognized event) is caught and produces no record rather than a thrown error or a non-zero exit; the process always exits 0. A miss here is a missing line in a telemetry log, never a blocked turn, so there is nothing to gate.
114
+ - **Bounded.** Stdin is drained asynchronously against a combined 1s time cap and 8 MB size cap; a payload that exceeds either is treated as truncated and parsed as nothing, never partially. `--summary` reads the log directly (never spawns anything, never executes a line in it).
115
+ - **Not yet attacked.** Untested here: two processes racing the same rotation at once (a rename plus an append landing on the same file); a state directory with thousands of leaked files from a long-lived session with a crashed hook (pruning runs, but only on `SubagentStart`, so an install that never starts a subagent again would never prune); behavior if `agent_id` collides across two concurrent subagents (sha256 makes this astronomically unlikely, not impossible).
@@ -32,7 +32,7 @@ There is a second thing a lane can be quietly wrong about. Left unpinned, it run
32
32
 
33
33
  ## 4. Every delegation carries a task bundle, on both surfaces
34
34
 
35
- Subagents and CLI lanes are close to the same problem: something that may hold none of your rules, and broad tool access. A Claude Code subagent is the one documented exception, loading the project's CLAUDE.md hierarchy at start, so it keeps the standing rules but not this task's scope; a CLI lane and a fresh chat window get no such credit. The brief (purpose, task class, scope, capabilities, denied actions, conventions, report contract, exit parameters) goes in the prompt or in the file passed to `--brief` either way. If you can, gate it mechanically: a pre-dispatch hook that refuses a brief missing purpose, denied actions or a report contract. On claude-code, a `SubagentStart` hook can inject the essentials (where the rules and the brief format live) automatically; `.claude/hooks/subagent-context.mjs` is the generated example.
35
+ Subagents and CLI lanes are close to the same problem: something that may hold none of your rules, and broad tool access. A Claude Code subagent is the one documented exception, loading the project's CLAUDE.md hierarchy at start, so it keeps the standing rules but not this task's scope; a CLI lane and a fresh chat window get no such credit. The brief (purpose, task class, scope, capabilities, denied actions, conventions, report contract, exit parameters) goes in the prompt or in the file passed to `--brief` either way. If you can, gate it mechanically: a pre-dispatch hook that refuses a brief missing purpose, denied actions or a report contract. On claude-code, a `SubagentStart` hook can inject the essentials (where the rules and the brief format live) automatically; `.claude/hooks/subagent-context.mjs` is the generated example. A third hook, `.claude/hooks/route-metrics.mjs`, turns that same delegation into a measurement instead of an assumption: it logs every turn, dispatch, subagent start/stop and the lane named in the reply's hidden route marker, and `--summary` reports route-marker coverage, dispatches with no matching start, and duration per agent type.
36
36
 
37
37
  ## 5. Research: three engines, one triager
38
38
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.15",
3
+ "version": "0.1.16",
4
4
  "description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
5
5
  "type": "module",
6
6
  "bin": {
package/src/detect.js CHANGED
@@ -2,24 +2,43 @@ import { accessSync, statSync, constants } from 'node:fs';
2
2
  import { delimiter, join } from 'node:path';
3
3
  import { homedir } from 'node:os';
4
4
 
5
+ // On win32, a PATH entry never holds a bare "grok": npm and vendor installers
6
+ // drop "grok.cmd" (or .exe/.bat/.ps1), the same way any Windows shell resolves
7
+ // a bare command through %PATHEXT%. Trying the bare name first keeps this a
8
+ // no-op on POSIX and matches an already-extensioned name (a .exe someone put
9
+ // on PATH directly) on Windows too.
10
+ // platform is a parameter (default process.platform), not a hardcoded read,
11
+ // so the win32 branch has a test on every OS this suite runs on: the same
12
+ // pattern bin/cli-run.mjs's killTree(pid, platform, deps) already uses.
13
+ export function candidateExtensions(platform = process.platform, pathext = process.env.PATHEXT) {
14
+ if (platform !== 'win32') return [''];
15
+ return ['', ...(pathext || '.COM;.EXE;.BAT;.CMD').split(';').filter(Boolean)];
16
+ }
17
+
5
18
  // PATH lookup plus the handful of places vendor installers drop binaries
6
19
  // without touching PATH. Never a shell function, never a shell out.
7
20
  // A directory with the binary's name is not a binary (X_OK passes on
8
21
  // searchable directories), so the candidate must be a regular file.
9
- export function which(bin) {
22
+ // home is a parameter too (default homedir()), for the same reason platform
23
+ // is: a dev machine with a real ~/.local/bin/grok on it made the win32 tests
24
+ // here false-negative until this was injectable, since PATH alone was never
25
+ // the whole search.
26
+ export function which(bin, platform = process.platform, home = homedir()) {
10
27
  if (!bin) return null;
11
- const home = homedir();
12
28
  const searchPath = process.env['PATH'] || '';
13
29
  const dirs = searchPath.split(delimiter).filter(Boolean);
14
30
  dirs.push(join(home, '.local', 'bin'), join(home, '.grok', 'bin'), join(home, '.npm-global', 'bin'));
31
+ const exts = candidateExtensions(platform);
15
32
  for (const d of dirs) {
16
- const p = join(d, bin);
17
- try {
18
- if (!statSync(p).isFile()) continue;
19
- accessSync(p, constants.X_OK);
20
- return p;
21
- } catch {
22
- /* keep looking */
33
+ for (const ext of exts) {
34
+ const p = join(d, bin + ext);
35
+ try {
36
+ if (!statSync(p).isFile()) continue;
37
+ accessSync(p, constants.X_OK);
38
+ return p;
39
+ } catch {
40
+ /* keep looking */
41
+ }
23
42
  }
24
43
  }
25
44
  return null;
package/src/install.js CHANGED
@@ -251,6 +251,8 @@ export function routeGateSection(selected) {
251
251
  "Stay inline only when: (a) the brief would cost as much as the work itself, (b) the task needs this conversation's own context, (c) it is the human's decision or the final verification of delegated work (a delegate never verifies itself).",
252
252
  '',
253
253
  'Never: the built-in Explore or Plan agents for rule-bound work (they skip CLAUDE.md). general-purpose taking work a named agent already owns.',
254
+ '',
255
+ 'End every reply with a hidden marker: `<!-- route: <lane> | <why, a few words> -->`. The route-metrics hook reads only the lane out of it, so routing coverage can be measured instead of assumed.',
254
256
  '<!-- route-gate:end -->'
255
257
  ].join('\n');
256
258
  }
@@ -360,7 +362,7 @@ export function activationSteps(opts) {
360
362
  if (primary && primary.agentsDir) steps.push(`subagents are in ${join(projectAbs, primary.agentsDir)}; run ${primary.bin} from ${projectAbs} to pick them up`);
361
363
  // Only claude-code ships hooks (route-gate, subagent-context): the wiring
362
364
  // lives in a snippet, never written into a settings.json the user already has.
363
- if (subagentsLoadRules(primary)) steps.push(`merge the hooks in ${join(dirAbs, 'settings.hooks.snippet.json')} into ${join(projectAbs, '.claude', 'settings.json')} (create it if missing) to wire the route-gate and subagent-context hooks`);
365
+ if (subagentsLoadRules(primary)) steps.push(`merge the hooks in ${join(dirAbs, 'settings.hooks.snippet.json')} into ${join(projectAbs, '.claude', 'settings.json')} (create it if missing) to wire the route-gate, subagent-context and route-metrics hooks`);
364
366
  for (const a of selected.filter((a) => a.bin && a.kind === 'agent-cli')) steps.push(`sign in to ${a.name}: ${a.auth}`);
365
367
  // A local runtime has a bin but no sign-in, so the agent-cli loop above skips it
366
368
  // and before this it appeared in no ordered list at any level (#26).
@@ -545,6 +547,11 @@ export function planFiles(opts) {
545
547
  // written into a settings.json they already have.
546
548
  add(join('.claude', 'hooks', 'route-gate.mjs'), render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'route-gate.mjs'), 'utf8'), v), 0o755, 'project');
547
549
  add(join('.claude', 'hooks', 'subagent-context.mjs'), render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'subagent-context.mjs'), 'utf8'), v), 0o755, 'project');
550
+ // route-metrics.mjs (0.1.16), claude-code only: five events (UserPromptSubmit,
551
+ // PreToolUse on Agent|Task, SubagentStart, SubagentStop, Stop) turned into one
552
+ // JSON line each under ~/.ai-orchestrator/, so a routing rule nobody measures
553
+ // is not a rule nobody knows is followed.
554
+ add(join('.claude', 'hooks', 'route-metrics.mjs'), render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'route-metrics.mjs'), 'utf8'), v), 0o755, 'project');
548
555
  add('settings.hooks.snippet.json', render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'settings.hooks.snippet.json'), 'utf8'), v));
549
556
  } else if (primary && primary.id === 'agy') {
550
557
  for (const f of walk(join(TEMPLATES, 'agents', 'agy'))) {
@@ -703,8 +710,22 @@ export const RUNTIME = new Set([
703
710
  'vm/jobs/weekly-audit.service',
704
711
  'vm/jobs/weekly-audit.timer'
705
712
  ]);
706
- export function fileClass(rel) {
707
- const r = rel.split(sep).join('/');
713
+ // MACHINE_OWNED and RUNTIME are keyed with forward slashes (they read as
714
+ // prose in the comment above them, and every caller needs the same one
715
+ // spelling regardless of host OS); an f.rel or a writeFiles() "written" path
716
+ // is built with path.join, so it is backslash-separated on win32. Both sets
717
+ // must be checked against the SAME normalized form, or a win32 install
718
+ // silently drops bin/lanes.json and every RUNTIME file from set membership
719
+ // (found: bin/cli.js's own "applied:"/existing-runtime checks did exactly
720
+ // that before this was exported for them to use too).
721
+ // separator is a parameter (default the real path.sep) so a test can prove
722
+ // the win32 case from any host, the same pattern which()'s platform
723
+ // parameter already uses.
724
+ export function toPosixRel(rel, separator = sep) {
725
+ return rel.split(separator).join('/');
726
+ }
727
+ export function fileClass(rel, separator = sep) {
728
+ const r = toPosixRel(rel, separator);
708
729
  if (MACHINE_OWNED.has(r)) return 'owned';
709
730
  if (RUNTIME.has(r)) return 'runtime';
710
731
  return 'document';
@@ -12,6 +12,12 @@ commandExecutionPolicy: off
12
12
  A finding is a claim, not a fact. You try to disprove each one before it is
13
13
  allowed to cause a repair.
14
14
 
15
+ No file-editing tools, and no command execution: this agent's
16
+ `commandExecutionPolicy` is `off`, so unlike its claude-code counterpart,
17
+ which carries an unrestricted `Bash` and stays read-only by its prompt rather
18
+ than by the tool grant, this agent is mechanically blocked from shelling out;
19
+ probe with whatever read or fetch capability you have instead.
20
+
15
21
  For each finding you are given: read the cited file and line yourself, state the
16
22
  input or sequence that would trigger it, then hunt for what makes it impossible
17
23
  (a guard upstream, a caller that never passes that value, an existing test).
@@ -6,11 +6,11 @@ One per tier, plus two checks and two agents with no file-editing tools: `findin
6
6
  |---|---|---|---|---|
7
7
  | deep-planner | deep | opus | xhigh | judges every build twice; never retrieves |
8
8
  | builder | standard | sonnet | high | executes; the default for everything that changes files |
9
- | code-reviewer | standard | sonnet | high | read-only findings |
9
+ | code-reviewer | standard | sonnet | high | findings only; no file-editing tools, Bash for checks only |
10
10
  | finding-verifier | standard | sonnet | high | tries to disprove a finding before it causes a repair |
11
11
  | live-researcher | standard | sonnet | medium | fresh data through tools |
12
12
  | bulk-worker | fast | haiku | low | mechanical volume, writes output |
13
13
  | done-verifier | fast | haiku | low | probes a tracker item's stated done-signal; no file-editing tools, Bash for probes only |
14
14
  | reader | fast | haiku | low | reads and digests many files or notes; read-only |
15
15
 
16
- Aliases resolve to the newest model in each family, so a version bump needs no edit here. Each agent carries its own token-discipline rule; the `effort` field is the third cost lever. Neither `done-verifier` nor `reader` carries `Write` or `Edit` in its `tools:` line. `reader` is read-only by tool grant as well: it carries no `Bash`. `done-verifier` does carry `Bash`, for its probes (`git log`, `grep`, `wc -l`, `test -f`); nothing in that grant stops it from running a command that changes state, so staying read-only there is a rule in its prompt, not a restriction on the tool, and its own file says so.
16
+ Aliases resolve to the newest model in each family, so a version bump needs no edit here. Each agent carries its own token-discipline rule; the `effort` field is the third cost lever. None of `done-verifier`, `finding-verifier`, `code-reviewer` or `reader` carries `Write` or `Edit` in its `tools:` line. `reader` is read-only by tool grant as well: it carries no `Bash`. `done-verifier`, `finding-verifier` and `code-reviewer` do carry `Bash`, for their probes and checks (`git log`, `grep`, `wc -l`, `test -f`); nothing in that grant stops any of them from running a command that changes state, so staying read-only there is a rule in each one's prompt, not a restriction on the tool, and each file says so.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: code-reviewer
3
- description: Code review. Use when asked to review code, a diff, or a repo for bugs, security issues, or quality. Read-only, returns findings. Do not use for writing or fixing code.
3
+ description: Code review. Use when asked to review code, a diff, or a repo for bugs, security issues, or quality. No file-editing tools; Bash is for read-only checks, bound by the prompt below, not by the tool grant. Returns findings. Do not use for writing or fixing code.
4
4
  tools: Read, Glob, Grep, Bash
5
5
  model: sonnet
6
6
  effort: high
@@ -10,10 +10,17 @@ You are the review tier of the model router.
10
10
 
11
11
  You review code for real bugs, security problems, and correctness issues.
12
12
 
13
+ You carry no Write or Edit tool, so you cannot touch a file. You do carry
14
+ Bash, and nothing in that grant stops you from running a command that changes
15
+ state; staying to read-only checks is a rule you follow below, not a
16
+ restriction you were given. Treat that boundary as load-bearing.
17
+
13
18
  Rules:
14
19
  - Report only findings you can defend with a concrete failure scenario. No style nitpicks unless asked.
15
20
  - Rank by severity. For each: file, line, what breaks, and the fix in one or two sentences.
16
21
  - Security findings (auth, secrets, injection, exposed endpoints) always rank first. Treat every endpoint as internet-facing.
17
- - You are read-only. Suggest fixes; do not apply them.
22
+ - Bash is for read-only checks only (`git log`, `grep`, `wc -l`, `test -f`, a
23
+ HEAD or GET request): never a command that changes state. Suggest fixes; do
24
+ not apply them.
18
25
  - If the code is clean, say so plainly. Do not invent findings.
19
26
  - Token discipline: read only the files under review, targeted sections where possible; report findings without restating the code; quote at most the few lines a finding needs.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: finding-verifier
3
- description: Adversarial verification of review findings. Use after a review or audit returns findings and before any of them trigger a repair. Read-only. Tries to DISPROVE each finding and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Do not use to find new problems, and do not use to fix anything.
3
+ description: Adversarial verification of review findings. Use after a review or audit returns findings and before any of them trigger a repair. No file-editing tools; Bash is for read-only checks, bound by the prompt below, not by the tool grant. Tries to DISPROVE each finding and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Do not use to find new problems, and do not use to fix anything.
4
4
  tools: Read, Glob, Grep, Bash
5
5
  model: sonnet
6
6
  effort: high
@@ -12,6 +12,11 @@ A finding is a claim, not a fact. Your job is to try to disprove each one before
12
12
  it is allowed to cause a change. A false finding is expensive twice: it buys a
13
13
  repair nobody needed, and it teaches everyone to skim the next report.
14
14
 
15
+ You carry no Write or Edit tool, so you cannot touch a file. You do carry
16
+ Bash, and nothing in that grant stops you from running a command that changes
17
+ state; staying to read-only checks is a rule you follow below, not a
18
+ restriction you were given. Treat that boundary as load-bearing.
19
+
15
20
  You are given findings from a review or an audit. For each one, independently:
16
21
 
17
22
  1. Read the cited file and line yourself. A citation that does not point at what
@@ -36,7 +41,9 @@ Return one verdict per finding, in the order you were given them:
36
41
  Rules:
37
42
  - Verify only the findings you were given. New problems you happen to notice go
38
43
  in a separate list at the end, clearly marked as unverified observations.
39
- - You are read-only. You never repair, and you never soften a finding's wording.
44
+ - Bash is for read-only checks only (`git log`, `grep`, `wc -l`, `test -f`, a
45
+ HEAD or GET request): never a command that changes state. You never repair,
46
+ and you never soften a finding's wording.
40
47
  - Verifying nothing is a real answer. If every finding is NOT_REPRODUCED, say
41
48
  that plainly; a verifier that always confirms something is a rubber stamp
42
49
  facing the other way.
@@ -9,7 +9,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
9
9
 
10
10
  1. Bulk, mechanical, many similar items -> bulk-worker (fast tier).
11
11
  2. Needs live data -> live-researcher (standard tier + tools).
12
- 3. Review without changing -> code-reviewer (standard, read-only).
12
+ 3. Review without changing -> code-reviewer (standard; no file-editing tools, Bash for checks only).
13
13
  3a. Holding findings from a review or scanner -> finding-verifier before any repair. Only CONFIRMED findings earn a change.
14
14
  3b. Checking a tracker item against its stated done-signal -> done-verifier. It never closes anything itself.
15
15
  4. Ambiguous, architectural, or expensive to get wrong -> deep-planner (deep tier), then hand the plan down.
@@ -0,0 +1,356 @@
1
+ #!/usr/bin/env node
2
+ // route-metrics.mjs: routing telemetry hook for {{PRIMARY_NAME}}.
3
+ //
4
+ // Answers "is my agent actually routing and delegating?" by turning five
5
+ // hook events into one JSON line each, appended to
6
+ // ~/.ai-orchestrator/route-metrics.jsonl (the same directory bin/cli-run.mjs
7
+ // already logs to, and the same os.homedir() resolution it uses):
8
+ //
9
+ // UserPromptSubmit -> {event:"turn"}
10
+ // PreToolUse (matcher Agent|Task) -> {event:"dispatch", subagent_type, background}
11
+ // SubagentStart -> {event:"start", agent_type}
12
+ // SubagentStop -> {event:"end", agent_type?, duration_s?}
13
+ // Stop -> {event:"route", lane}, parsed from the LAST
14
+ // <!-- route: <lane> | <why> --> marker in
15
+ // last_assistant_message (documented source;
16
+ // the transcript can lag, so that is never read)
17
+ //
18
+ // Pure telemetry, fail-open by design: this script prints NOTHING to stdout
19
+ // (stdout on UserPromptSubmit/SubagentStart becomes model context) and always
20
+ // exits 0, whether or not a line was written. A miss here is a missing log
21
+ // line, never a blocked turn.
22
+ //
23
+ // The durable log holds no provider-supplied string: prompt text, tool
24
+ // descriptions, and the "why" half of the route marker are never read into a
25
+ // field, only the named, charset-bounded values below. See docs/audit-brief.md.
26
+ //
27
+ // Second entry point: `node route-metrics.mjs --summary [--since <ISO date>]`
28
+ // prints a plain-text report from the log and exits 0 without touching stdin.
29
+ import { createHash } from 'node:crypto';
30
+ import {
31
+ existsSync, mkdirSync, appendFileSync, readFileSync, writeFileSync, unlinkSync, renameSync,
32
+ statSync, readdirSync
33
+ } from 'node:fs';
34
+ import { join } from 'node:path';
35
+ import { homedir } from 'node:os';
36
+
37
+ const HOME_DIR = join(homedir(), '.ai-orchestrator');
38
+ const LOG_FILE = join(HOME_DIR, 'route-metrics.jsonl');
39
+ const STATE_DIR = join(HOME_DIR, 'route-metrics.state');
40
+
41
+ const STDIN_MAX_BYTES = 8 * 1024 * 1024; // size cap: a giant or runaway payload is truncated, not parsed
42
+ const STDIN_DRAIN_MS = 1000; // hard cap: never let an open, never-closed stdin pipe hold this hook open
43
+ const LOG_ROTATE_BYTES = 5 * 1024 * 1024; // rotate to .1 above this size
44
+ const STATE_MAX_AGE_MS = 24 * 60 * 60 * 1000; // prune state files older than 24h
45
+ const TOKEN_CHARSET = /[^A-Za-z0-9_.+-]/g; // session_id, subagent_type, agent_type, lane tokens
46
+ const TOKEN_MAX_LEN = 64;
47
+ const SESSION_ID_MAX_LEN = 128;
48
+
49
+ // Strip anything outside the allowed charset and cap length, so no field in
50
+ // the durable log can carry an arbitrary provider- or model-supplied string
51
+ // (a newline, a control character, shell metacharacters, or just length).
52
+ function sanitize(raw, maxLen) {
53
+ if (typeof raw !== 'string' || raw.length === 0) return '';
54
+ return raw.replace(TOKEN_CHARSET, '').slice(0, maxLen);
55
+ }
56
+
57
+ function stateKeyFor(agentId) {
58
+ if (typeof agentId !== 'string' || agentId.length === 0) return null;
59
+ return createHash('sha256').update(agentId).digest('hex');
60
+ }
61
+
62
+ // Best-effort housekeeping: a leaked state file (a SubagentStop that never
63
+ // arrived) should not accumulate forever. Run on SubagentStart only, since
64
+ // that is the one event guaranteed to fire at least as often as starts happen.
65
+ function pruneOldState() {
66
+ let names;
67
+ try {
68
+ names = readdirSync(STATE_DIR);
69
+ } catch {
70
+ return; // no state dir yet: nothing to prune
71
+ }
72
+ const cutoff = Date.now() - STATE_MAX_AGE_MS;
73
+ for (const name of names) {
74
+ const p = join(STATE_DIR, name);
75
+ try {
76
+ if (statSync(p).mtimeMs < cutoff) unlinkSync(p);
77
+ } catch {
78
+ /* a race with another process touching the same file is not an error here */
79
+ }
80
+ }
81
+ }
82
+
83
+ function recordStart(agentId, agentType) {
84
+ const key = stateKeyFor(agentId);
85
+ if (!key) return;
86
+ try {
87
+ mkdirSync(STATE_DIR, { recursive: true });
88
+ writeFileSync(join(STATE_DIR, key + '.json'), JSON.stringify({ ts: Date.now(), agent_type: agentType }));
89
+ } catch {
90
+ /* telemetry never blocks the run */
91
+ }
92
+ }
93
+
94
+ // Reads and deletes the state file for this agent_id. Returns {agentType,
95
+ // durationS}, either possibly null: no agent_id and no state file both mean
96
+ // "none", which the caller reflects by omitting the field entirely.
97
+ function consumeStart(agentId) {
98
+ const key = stateKeyFor(agentId);
99
+ if (!key) return { agentType: null, durationS: null };
100
+ const p = join(STATE_DIR, key + '.json');
101
+ let agentType = null;
102
+ let durationS = null;
103
+ try {
104
+ const parsed = JSON.parse(readFileSync(p, 'utf8'));
105
+ if (parsed && typeof parsed.ts === 'number') durationS = Math.max(0, (Date.now() - parsed.ts) / 1000);
106
+ if (parsed && typeof parsed.agent_type === 'string' && parsed.agent_type) agentType = parsed.agent_type;
107
+ } catch {
108
+ /* no state file, or it was unreadable: none, not an error */
109
+ }
110
+ try {
111
+ unlinkSync(p);
112
+ } catch {
113
+ /* already gone */
114
+ }
115
+ return { agentType, durationS };
116
+ }
117
+
118
+ function appendLog(record) {
119
+ try {
120
+ mkdirSync(HOME_DIR, { recursive: true });
121
+ let size = 0;
122
+ try {
123
+ size = statSync(LOG_FILE).size;
124
+ } catch {
125
+ /* file does not exist yet: size stays 0 */
126
+ }
127
+ if (size > LOG_ROTATE_BYTES) {
128
+ try {
129
+ renameSync(LOG_FILE, LOG_FILE + '.1');
130
+ } catch {
131
+ /* a concurrent rotation losing this race is not worth failing over */
132
+ }
133
+ }
134
+ appendFileSync(LOG_FILE, JSON.stringify(record) + '\n');
135
+ } catch {
136
+ /* telemetry never blocks the run */
137
+ }
138
+ }
139
+
140
+ // Parses the LAST <!-- route: <lane> | <why> --> marker out of text. The
141
+ // "why" half is captured only to be discarded: it is never read into a
142
+ // variable that reaches the log. Returns an array of lane tokens (split on
143
+ // "+", the documented way to log more than one lane from a single marker),
144
+ // or ["missing"] when there is no marker at all.
145
+ export function extractLane(text) {
146
+ if (typeof text !== 'string' || text.length === 0) return ['missing'];
147
+ const re = /<!--\s*route:\s*([^|>]*)\|[^>]*-->/g;
148
+ let match;
149
+ let last = null;
150
+ while ((match = re.exec(text)) !== null) last = match;
151
+ if (!last) return ['missing'];
152
+ // A token carrying any character outside the charset is logged as
153
+ // "invalid", never stripped into a plausible-looking lane: stripping
154
+ // `main","evil":"1` would log a lane named "mainevil1" that nobody chose.
155
+ const parts = last[1]
156
+ .split('+')
157
+ .map((s) => s.trim())
158
+ .filter(Boolean)
159
+ .map((t) => (t.length > TOKEN_MAX_LEN || /[^A-Za-z0-9_.-]/.test(t) ? 'invalid' : t));
160
+ return parts.length ? parts : ['missing'];
161
+ }
162
+
163
+ // Turns one parsed hook-input object into a log record, or null when the
164
+ // event is not one this hook measures (or PreToolUse fired for a tool other
165
+ // than Agent/Task, which the settings matcher should already have excluded;
166
+ // this is a defensive second check, not the primary gate).
167
+ export function buildRecord(input, now = () => new Date().toISOString()) {
168
+ if (!input || typeof input !== 'object') return null;
169
+ const sessionId = sanitize(input.session_id, SESSION_ID_MAX_LEN) || 'unknown';
170
+ const ts = now();
171
+ const base = { ts, v: 1 };
172
+
173
+ switch (input.hook_event_name) {
174
+ case 'UserPromptSubmit':
175
+ return { ...base, event: 'turn', session_id: sessionId };
176
+
177
+ case 'PreToolUse': {
178
+ if (input.tool_name !== 'Agent' && input.tool_name !== 'Task') return null;
179
+ const toolInput = (input.tool_input && typeof input.tool_input === 'object') ? input.tool_input : {};
180
+ const subagentType = sanitize(toolInput.subagent_type, TOKEN_MAX_LEN) || 'general-purpose';
181
+ const background = toolInput.run_in_background === true;
182
+ return { ...base, event: 'dispatch', session_id: sessionId, subagent_type: subagentType, background };
183
+ }
184
+
185
+ case 'SubagentStart': {
186
+ pruneOldState();
187
+ const agentType = sanitize(input.agent_type, TOKEN_MAX_LEN) || 'unknown';
188
+ recordStart(input.agent_id, agentType);
189
+ return { ...base, event: 'start', session_id: sessionId, agent_type: agentType };
190
+ }
191
+
192
+ case 'SubagentStop': {
193
+ const { agentType, durationS } = consumeStart(input.agent_id);
194
+ const record = { ...base, event: 'end', session_id: sessionId };
195
+ if (agentType) record.agent_type = agentType;
196
+ if (durationS !== null) record.duration_s = durationS;
197
+ return record;
198
+ }
199
+
200
+ case 'Stop':
201
+ return { ...base, event: 'route', session_id: sessionId, lane: extractLane(input.last_assistant_message) };
202
+
203
+ default:
204
+ return null;
205
+ }
206
+ }
207
+
208
+ // Drain stdin without ever blocking on it, bounded by BOTH time and size. A
209
+ // bare `readFileSync(0)` waits for EOF, so a caller that pipes in and never
210
+ // closes its end left the process running indefinitely (the same class of
211
+ // bug route-gate.mjs and subagent-context.mjs already fix). The size cap is
212
+ // this hook's own addition: hook input is normally small, so a payload past
213
+ // the cap is treated as truncated and parsed as nothing, never partially.
214
+ function drainStdinBounded(timeoutMs, maxBytes) {
215
+ return new Promise((resolve) => {
216
+ let settled = false;
217
+ let bytes = 0;
218
+ let truncated = false;
219
+ const chunks = [];
220
+ const finish = () => {
221
+ if (settled) return;
222
+ settled = true;
223
+ clearTimeout(timer);
224
+ try {
225
+ process.stdin.removeAllListeners('data');
226
+ process.stdin.removeAllListeners('end');
227
+ process.stdin.removeAllListeners('error');
228
+ process.stdin.pause();
229
+ } catch {
230
+ /* stdin may already be gone */
231
+ }
232
+ resolve({ data: truncated ? null : Buffer.concat(chunks).toString('utf8'), truncated });
233
+ };
234
+ const timer = setTimeout(finish, timeoutMs);
235
+ if (timer.unref) timer.unref();
236
+ try {
237
+ process.stdin.on('data', (chunk) => {
238
+ if (truncated) return;
239
+ bytes += chunk.length;
240
+ if (bytes > maxBytes) {
241
+ truncated = true;
242
+ return finish();
243
+ }
244
+ chunks.push(chunk);
245
+ });
246
+ process.stdin.on('end', finish);
247
+ process.stdin.on('error', finish);
248
+ process.stdin.resume();
249
+ } catch {
250
+ finish();
251
+ }
252
+ });
253
+ }
254
+
255
+ async function runHook() {
256
+ const { data } = await drainStdinBounded(STDIN_DRAIN_MS, STDIN_MAX_BYTES);
257
+ if (data) {
258
+ let input;
259
+ try {
260
+ input = JSON.parse(data);
261
+ } catch {
262
+ input = null; // invalid JSON: log nothing
263
+ }
264
+ if (input) {
265
+ try {
266
+ const record = buildRecord(input);
267
+ if (record) appendLog(record);
268
+ } catch {
269
+ /* telemetry never blocks or fails the run */
270
+ }
271
+ }
272
+ }
273
+ process.exit(0); // fail-open, always: a miss here is a missing log line, never a blocked turn
274
+ }
275
+
276
+ // ---- --summary: a plain-text report, no stdin involved ----
277
+
278
+ function parseLines(text) {
279
+ const records = [];
280
+ for (const line of text.split('\n')) {
281
+ const trimmed = line.trim();
282
+ if (!trimmed) continue;
283
+ try {
284
+ records.push(JSON.parse(trimmed));
285
+ } catch {
286
+ /* one bad line (a torn write, a rotation race) does not sink the report */
287
+ }
288
+ }
289
+ return records;
290
+ }
291
+
292
+ function formatNumber(n) {
293
+ return Number.isInteger(n) ? String(n) : n.toFixed(2);
294
+ }
295
+
296
+ function runSummary(args) {
297
+ if (!existsSync(LOG_FILE)) {
298
+ console.log('route-metrics: no data yet (' + LOG_FILE + ' does not exist).');
299
+ return process.exit(0);
300
+ }
301
+ const sinceIdx = args.indexOf('--since');
302
+ const since = sinceIdx !== -1 ? Date.parse(args[sinceIdx + 1]) : NaN;
303
+ let records = parseLines(readFileSync(LOG_FILE, 'utf8'));
304
+ if (!Number.isNaN(since)) records = records.filter((r) => Date.parse(r.ts) >= since);
305
+
306
+ const turns = records.filter((r) => r.event === 'turn').length;
307
+ const routes = records.filter((r) => r.event === 'route');
308
+ const covered = routes.filter((r) => !(Array.isArray(r.lane) && r.lane.length === 1 && r.lane[0] === 'missing')).length;
309
+ const coveragePct = turns > 0 ? (covered / turns) * 100 : null;
310
+
311
+ const laneCounts = new Map();
312
+ for (const r of routes) {
313
+ for (const lane of Array.isArray(r.lane) ? r.lane : []) laneCounts.set(lane, (laneCounts.get(lane) || 0) + 1);
314
+ }
315
+
316
+ const dispatches = records.filter((r) => r.event === 'dispatch');
317
+ const dispatchCounts = new Map();
318
+ for (const r of dispatches) dispatchCounts.set(r.subagent_type, (dispatchCounts.get(r.subagent_type) || 0) + 1);
319
+
320
+ const starts = records.filter((r) => r.event === 'start').length;
321
+ const noMatchingStart = Math.max(0, dispatches.length - starts);
322
+
323
+ const ends = records.filter((r) => r.event === 'end' && r.agent_type && typeof r.duration_s === 'number');
324
+ const durationsByType = new Map();
325
+ for (const r of ends) {
326
+ if (!durationsByType.has(r.agent_type)) durationsByType.set(r.agent_type, []);
327
+ durationsByType.get(r.agent_type).push(r.duration_s);
328
+ }
329
+
330
+ const lines = [];
331
+ lines.push('route-metrics summary' + (Number.isNaN(since) ? '' : ' since ' + args[sinceIdx + 1]));
332
+ lines.push('turns: ' + turns);
333
+ lines.push('route-marker coverage: ' + (coveragePct === null ? 'no turns yet' : formatNumber(coveragePct) + '%') + ' (' + covered + '/' + turns + ')');
334
+ lines.push('lanes by count:');
335
+ if (laneCounts.size === 0) lines.push(' (none)');
336
+ for (const [lane, count] of [...laneCounts.entries()].sort((a, b) => b[1] - a[1])) lines.push(' ' + lane + ': ' + count);
337
+ lines.push('dispatches by subagent_type:');
338
+ if (dispatchCounts.size === 0) lines.push(' (none)');
339
+ for (const [type, count] of [...dispatchCounts.entries()].sort((a, b) => b[1] - a[1])) lines.push(' ' + type + ': ' + count);
340
+ lines.push('dispatches with no matching start: ' + noMatchingStart + ' (a hook or guard blocked them before launch)');
341
+ lines.push('duration by agent_type (mean / max, seconds):');
342
+ if (durationsByType.size === 0) lines.push(' (none)');
343
+ for (const [type, durs] of durationsByType) {
344
+ const mean = durs.reduce((a, b) => a + b, 0) / durs.length;
345
+ lines.push(' ' + type + ': ' + formatNumber(mean) + ' / ' + formatNumber(Math.max(...durs)));
346
+ }
347
+ console.log(lines.join('\n'));
348
+ process.exit(0);
349
+ }
350
+
351
+ const args = process.argv.slice(2);
352
+ if (args.includes('--summary')) {
353
+ runSummary(args);
354
+ } else {
355
+ runHook();
356
+ }
@@ -7,6 +7,23 @@
7
7
  "type": "command",
8
8
  "command": "node",
9
9
  "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-gate.mjs"]
10
+ },
11
+ {
12
+ "type": "command",
13
+ "command": "node",
14
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-metrics.mjs"]
15
+ }
16
+ ]
17
+ }
18
+ ],
19
+ "PreToolUse": [
20
+ {
21
+ "matcher": "Agent|Task",
22
+ "hooks": [
23
+ {
24
+ "type": "command",
25
+ "command": "node",
26
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-metrics.mjs"]
10
27
  }
11
28
  ]
12
29
  }
@@ -18,6 +35,33 @@
18
35
  "type": "command",
19
36
  "command": "node",
20
37
  "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/subagent-context.mjs"]
38
+ },
39
+ {
40
+ "type": "command",
41
+ "command": "node",
42
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-metrics.mjs"]
43
+ }
44
+ ]
45
+ }
46
+ ],
47
+ "SubagentStop": [
48
+ {
49
+ "hooks": [
50
+ {
51
+ "type": "command",
52
+ "command": "node",
53
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-metrics.mjs"]
54
+ }
55
+ ]
56
+ }
57
+ ],
58
+ "Stop": [
59
+ {
60
+ "hooks": [
61
+ {
62
+ "type": "command",
63
+ "command": "node",
64
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-metrics.mjs"]
21
65
  }
22
66
  ]
23
67
  }