model-orchestrator 0.1.15 → 0.1.17

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. package/CHANGELOG.md +27 -1
  2. package/README.md +29 -14
  3. package/bin/cli-run.mjs +22 -7
  4. package/bin/cli.js +3 -3
  5. package/docs/README.md +1 -1
  6. package/docs/audit-brief.md +11 -0
  7. package/docs/catalog.md +1 -1
  8. package/docs/part-1-beginner.md +4 -4
  9. package/docs/part-2-intermediate.md +5 -5
  10. package/llms.txt +1 -1
  11. package/package.json +1 -1
  12. package/src/catalog.js +1 -1
  13. package/src/detect.js +28 -9
  14. package/src/install.js +31 -10
  15. package/templates/agents/agy/README.md +1 -1
  16. package/templates/agents/agy/builder.md +1 -1
  17. package/templates/agents/agy/finding-verifier.md +7 -1
  18. package/templates/agents/claude-code/README.md +2 -2
  19. package/templates/agents/claude-code/code-reviewer.md +9 -2
  20. package/templates/agents/claude-code/finding-verifier.md +9 -2
  21. package/templates/agents/snippets/chat.md +2 -2
  22. package/templates/agents/snippets/claude-code.md +2 -2
  23. package/templates/agents/snippets/generic.md +1 -1
  24. package/templates/agents/snippets/route-gate.mjs +1 -1
  25. package/templates/agents/snippets/route-metrics.mjs +356 -0
  26. package/templates/agents/snippets/settings.hooks.snippet.json +44 -0
  27. package/templates/beginner/ORCHESTRATOR.md +2 -2
  28. package/templates/common/README.md +1 -1
  29. package/templates/common/protocols/build-protocol.md +10 -10
  30. package/templates/common/protocols/deep-research.md +2 -2
  31. package/templates/common/protocols/gap-analysis.md +1 -1
  32. package/templates/common/protocols/propagate.md +2 -2
  33. package/templates/intermediate/CLI-RUN.md +3 -3
  34. package/templates/intermediate/DELEGATION_MATRIX.md +1 -1
  35. package/templates/intermediate/ROUTING.md +4 -4
  36. package/templates/intermediate/TIERS.md +12 -8
  37. package/templates/tools/codecalc/CODECALC.md +1 -1
  38. package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +3 -3
@@ -6,11 +6,11 @@ One per tier, plus two checks and two agents with no file-editing tools: `findin
6
6
  |---|---|---|---|---|
7
7
  | deep-planner | deep | opus | xhigh | judges every build twice; never retrieves |
8
8
  | builder | standard | sonnet | high | executes; the default for everything that changes files |
9
- | code-reviewer | standard | sonnet | high | read-only findings |
9
+ | code-reviewer | standard | sonnet | high | findings only; no file-editing tools, Bash for checks only |
10
10
  | finding-verifier | standard | sonnet | high | tries to disprove a finding before it causes a repair |
11
11
  | live-researcher | standard | sonnet | medium | fresh data through tools |
12
12
  | bulk-worker | fast | haiku | low | mechanical volume, writes output |
13
13
  | done-verifier | fast | haiku | low | probes a tracker item's stated done-signal; no file-editing tools, Bash for probes only |
14
14
  | reader | fast | haiku | low | reads and digests many files or notes; read-only |
15
15
 
16
- Aliases resolve to the newest model in each family, so a version bump needs no edit here. Each agent carries its own token-discipline rule; the `effort` field is the third cost lever. Neither `done-verifier` nor `reader` carries `Write` or `Edit` in its `tools:` line. `reader` is read-only by tool grant as well: it carries no `Bash`. `done-verifier` does carry `Bash`, for its probes (`git log`, `grep`, `wc -l`, `test -f`); nothing in that grant stops it from running a command that changes state, so staying read-only there is a rule in its prompt, not a restriction on the tool, and its own file says so.
16
+ Aliases resolve to the newest model in each family, so a version bump needs no edit here. Each agent carries its own token-discipline rule; the `effort` field is the third cost lever. None of `done-verifier`, `finding-verifier`, `code-reviewer` or `reader` carries `Write` or `Edit` in its `tools:` line. `reader` is read-only by tool grant as well: it carries no `Bash`. `done-verifier`, `finding-verifier` and `code-reviewer` do carry `Bash`, for their probes and checks (`git log`, `grep`, `wc -l`, `test -f`); nothing in that grant stops any of them from running a command that changes state, so staying read-only there is a rule in each one's prompt, not a restriction on the tool, and each file says so.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: code-reviewer
3
- description: Code review. Use when asked to review code, a diff, or a repo for bugs, security issues, or quality. Read-only, returns findings. Do not use for writing or fixing code.
3
+ description: Code review. Use when asked to review code, a diff, or a repo for bugs, security issues, or quality. No file-editing tools; Bash is for read-only checks, bound by the prompt below, not by the tool grant. Returns findings. Do not use for writing or fixing code.
4
4
  tools: Read, Glob, Grep, Bash
5
5
  model: sonnet
6
6
  effort: high
@@ -10,10 +10,17 @@ You are the review tier of the model router.
10
10
 
11
11
  You review code for real bugs, security problems, and correctness issues.
12
12
 
13
+ You carry no Write or Edit tool, so you cannot touch a file. You do carry
14
+ Bash, and nothing in that grant stops you from running a command that changes
15
+ state; staying to read-only checks is a rule you follow below, not a
16
+ restriction you were given. Treat that boundary as load-bearing.
17
+
13
18
  Rules:
14
19
  - Report only findings you can defend with a concrete failure scenario. No style nitpicks unless asked.
15
20
  - Rank by severity. For each: file, line, what breaks, and the fix in one or two sentences.
16
21
  - Security findings (auth, secrets, injection, exposed endpoints) always rank first. Treat every endpoint as internet-facing.
17
- - You are read-only. Suggest fixes; do not apply them.
22
+ - Bash is for read-only checks only (`git log`, `grep`, `wc -l`, `test -f`, a
23
+ HEAD or GET request): never a command that changes state. Suggest fixes; do
24
+ not apply them.
18
25
  - If the code is clean, say so plainly. Do not invent findings.
19
26
  - Token discipline: read only the files under review, targeted sections where possible; report findings without restating the code; quote at most the few lines a finding needs.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: finding-verifier
3
- description: Adversarial verification of review findings. Use after a review or audit returns findings and before any of them trigger a repair. Read-only. Tries to DISPROVE each finding and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Do not use to find new problems, and do not use to fix anything.
3
+ description: Second-opinion verification of review findings. Use after a review or audit returns findings and before any of them trigger a repair. No file-editing tools; Bash is for read-only checks, bound by the prompt below, not by the tool grant. Tries to DISPROVE each finding and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Do not use to find new problems, and do not use to fix anything.
4
4
  tools: Read, Glob, Grep, Bash
5
5
  model: sonnet
6
6
  effort: high
@@ -12,6 +12,11 @@ A finding is a claim, not a fact. Your job is to try to disprove each one before
12
12
  it is allowed to cause a change. A false finding is expensive twice: it buys a
13
13
  repair nobody needed, and it teaches everyone to skim the next report.
14
14
 
15
+ You carry no Write or Edit tool, so you cannot touch a file. You do carry
16
+ Bash, and nothing in that grant stops you from running a command that changes
17
+ state; staying to read-only checks is a rule you follow below, not a
18
+ restriction you were given. Treat that boundary as load-bearing.
19
+
15
20
  You are given findings from a review or an audit. For each one, independently:
16
21
 
17
22
  1. Read the cited file and line yourself. A citation that does not point at what
@@ -36,7 +41,9 @@ Return one verdict per finding, in the order you were given them:
36
41
  Rules:
37
42
  - Verify only the findings you were given. New problems you happen to notice go
38
43
  in a separate list at the end, clearly marked as unverified observations.
39
- - You are read-only. You never repair, and you never soften a finding's wording.
44
+ - Bash is for read-only checks only (`git log`, `grep`, `wc -l`, `test -f`, a
45
+ HEAD or GET request): never a command that changes state. You never repair,
46
+ and you never soften a finding's wording.
40
47
  - Verifying nothing is a real answer. If every finding is NOT_REPRODUCED, say
41
48
  that plainly; a verifier that always confirms something is a rubber stamp
42
49
  facing the other way.
@@ -4,8 +4,8 @@
4
4
 
5
5
  ```
6
6
  You follow a model-orchestrator workflow inside this chat. Tiers describe effort, not automatic model switching or cost savings.
7
- Route first: bulk/formatting -> fast; live data -> standard with tools; review -> standard, read-only; ambiguous/high-risk -> deep; otherwise standard. State the tier. Escalate on failure instead of silently retrying.
8
- For builds: map affected parts; identify the biggest risk and any flaw in the request (none is valid with reasons); build and verify; use a fresh turn to challenge the result. Before irreversible actions, explain rollback and ask for approval.
7
+ Route first: bulk/formatting -> fast; live data -> standard with tools; review -> standard, read-only; ambiguous/high-stakes -> deep; otherwise standard. State the tier. Escalate on failure instead of silently retrying.
8
+ For builds: map affected parts; identify what is most likely to go wrong and any gap in the request (none is valid with reasons); build and verify; use a fresh turn to challenge the result. Before irreversible actions, explain rollback and ask for approval.
9
9
  For hand-offs: include purpose, scope, allowed and denied actions, required output, and stopping conditions. A fresh context has none of these instructions.
10
10
  After comprehensive work, check for omissions. Compute consequential numbers and comparisons with a tool; report what was checked and what remains unverified.
11
11
  Before durable writes, search existing records, update their index, use one writer, and label inferences.
@@ -9,7 +9,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
9
9
 
10
10
  1. Bulk, mechanical, many similar items -> bulk-worker (fast tier).
11
11
  2. Needs live data -> live-researcher (standard tier + tools).
12
- 3. Review without changing -> code-reviewer (standard, read-only).
12
+ 3. Review without changing -> code-reviewer (standard; no file-editing tools, Bash for checks only).
13
13
  3a. Holding findings from a review or scanner -> finding-verifier before any repair. Only CONFIRMED findings earn a change.
14
14
  3b. Checking a tracker item against its stated done-signal -> done-verifier. It never closes anything itself.
15
15
  4. Ambiguous, architectural, or expensive to get wrong -> deep-planner (deep tier), then hand the plan down.
@@ -17,7 +17,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
17
17
 
18
18
  A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Measure yours once: spawn a subagent with a one-line task and read its token count. Work smaller than that stays inline.
19
19
 
20
- Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one adversarial pass, an explicit human yes before anything irreversible, then the loud negative.
20
+ Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one challenge pass, an explicit human yes before anything irreversible, then the loud negative.
21
21
 
22
22
  Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A Claude Code subagent loads this CLAUDE.md hierarchy, so it holds the standing rules already, just not this task's scope; a second CLI or a fresh chat window may hold none of them. Absence is denial either way.
23
23
 
@@ -9,7 +9,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
9
9
 
10
10
  Route by capability tier, first match wins: bulk and mechanical -> fast tier · needs live data -> standard tier with tools · review without changing -> standard, read-only · ambiguous or expensive to get wrong -> deep tier, then hand the plan down · everything else -> build it directly at standard tier.
11
11
 
12
- Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: map the blast radius yourself, ask the deep tier for a named risk and a named flaw, build green, scan the added lines, one adversarial pass with every finding reproduced, an explicit human yes before anything irreversible, then re-grep the old identifier and expect zero.
12
+ Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: map everything it touches yourself, ask the deep tier for one named weak spot and one gap in the request, build green, scan the added lines, one challenge pass with every finding reproduced, an explicit human yes before anything irreversible, then re-grep the old identifier and expect zero.
13
13
 
14
14
  Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A fresh context holds none of these rules; absence is denial.
15
15
 
@@ -10,7 +10,7 @@
10
10
  // This script always exits 0, never blocks on stdin past a short bound,
11
11
  // reads at most 64 KB of the rules file through a fixed-size buffer (never
12
12
  // a full read of an arbitrarily large or non-regular file), and never
13
- // executes anything it reads. See docs/audit-brief.md for the threat model.
13
+ // executes anything it reads. See docs/audit-brief.md for the security notes.
14
14
  import { statSync, openSync, readSync, closeSync, realpathSync } from 'node:fs';
15
15
  import { join, isAbsolute } from 'node:path';
16
16
 
@@ -0,0 +1,356 @@
1
+ #!/usr/bin/env node
2
+ // route-metrics.mjs: routing telemetry hook for {{PRIMARY_NAME}}.
3
+ //
4
+ // Answers "is my agent actually routing and delegating?" by turning five
5
+ // hook events into one JSON line each, appended to
6
+ // ~/.ai-orchestrator/route-metrics.jsonl (the same directory bin/cli-run.mjs
7
+ // already logs to, and the same os.homedir() resolution it uses):
8
+ //
9
+ // UserPromptSubmit -> {event:"turn"}
10
+ // PreToolUse (matcher Agent|Task) -> {event:"dispatch", subagent_type, background}
11
+ // SubagentStart -> {event:"start", agent_type}
12
+ // SubagentStop -> {event:"end", agent_type?, duration_s?}
13
+ // Stop -> {event:"route", lane}, parsed from the LAST
14
+ // <!-- route: <lane> | <why> --> marker in
15
+ // last_assistant_message (documented source;
16
+ // the transcript can lag, so that is never read)
17
+ //
18
+ // Pure telemetry, fail-open by design: this script prints NOTHING to stdout
19
+ // (stdout on UserPromptSubmit/SubagentStart becomes model context) and always
20
+ // exits 0, whether or not a line was written. A miss here is a missing log
21
+ // line, never a blocked turn.
22
+ //
23
+ // The durable log holds no provider-supplied string: prompt text, tool
24
+ // descriptions, and the "why" half of the route marker are never read into a
25
+ // field, only the named, charset-bounded values below. See docs/audit-brief.md.
26
+ //
27
+ // Second entry point: `node route-metrics.mjs --summary [--since <ISO date>]`
28
+ // prints a plain-text report from the log and exits 0 without touching stdin.
29
+ import { createHash } from 'node:crypto';
30
+ import {
31
+ existsSync, mkdirSync, appendFileSync, readFileSync, writeFileSync, unlinkSync, renameSync,
32
+ statSync, readdirSync
33
+ } from 'node:fs';
34
+ import { join } from 'node:path';
35
+ import { homedir } from 'node:os';
36
+
37
+ const HOME_DIR = join(homedir(), '.ai-orchestrator');
38
+ const LOG_FILE = join(HOME_DIR, 'route-metrics.jsonl');
39
+ const STATE_DIR = join(HOME_DIR, 'route-metrics.state');
40
+
41
+ const STDIN_MAX_BYTES = 8 * 1024 * 1024; // size cap: a giant or runaway payload is truncated, not parsed
42
+ const STDIN_DRAIN_MS = 1000; // hard cap: never let an open, never-closed stdin pipe hold this hook open
43
+ const LOG_ROTATE_BYTES = 5 * 1024 * 1024; // rotate to .1 above this size
44
+ const STATE_MAX_AGE_MS = 24 * 60 * 60 * 1000; // prune state files older than 24h
45
+ const TOKEN_CHARSET = /[^A-Za-z0-9_.+-]/g; // session_id, subagent_type, agent_type, lane tokens
46
+ const TOKEN_MAX_LEN = 64;
47
+ const SESSION_ID_MAX_LEN = 128;
48
+
49
+ // Strip anything outside the allowed charset and cap length, so no field in
50
+ // the durable log can carry an arbitrary provider- or model-supplied string
51
+ // (a newline, a control character, shell metacharacters, or just length).
52
+ function sanitize(raw, maxLen) {
53
+ if (typeof raw !== 'string' || raw.length === 0) return '';
54
+ return raw.replace(TOKEN_CHARSET, '').slice(0, maxLen);
55
+ }
56
+
57
+ function stateKeyFor(agentId) {
58
+ if (typeof agentId !== 'string' || agentId.length === 0) return null;
59
+ return createHash('sha256').update(agentId).digest('hex');
60
+ }
61
+
62
+ // Best-effort housekeeping: a leaked state file (a SubagentStop that never
63
+ // arrived) should not accumulate forever. Run on SubagentStart only, since
64
+ // that is the one event guaranteed to fire at least as often as starts happen.
65
+ function pruneOldState() {
66
+ let names;
67
+ try {
68
+ names = readdirSync(STATE_DIR);
69
+ } catch {
70
+ return; // no state dir yet: nothing to prune
71
+ }
72
+ const cutoff = Date.now() - STATE_MAX_AGE_MS;
73
+ for (const name of names) {
74
+ const p = join(STATE_DIR, name);
75
+ try {
76
+ if (statSync(p).mtimeMs < cutoff) unlinkSync(p);
77
+ } catch {
78
+ /* a race with another process touching the same file is not an error here */
79
+ }
80
+ }
81
+ }
82
+
83
+ function recordStart(agentId, agentType) {
84
+ const key = stateKeyFor(agentId);
85
+ if (!key) return;
86
+ try {
87
+ mkdirSync(STATE_DIR, { recursive: true });
88
+ writeFileSync(join(STATE_DIR, key + '.json'), JSON.stringify({ ts: Date.now(), agent_type: agentType }));
89
+ } catch {
90
+ /* telemetry never blocks the run */
91
+ }
92
+ }
93
+
94
+ // Reads and deletes the state file for this agent_id. Returns {agentType,
95
+ // durationS}, either possibly null: no agent_id and no state file both mean
96
+ // "none", which the caller reflects by omitting the field entirely.
97
+ function consumeStart(agentId) {
98
+ const key = stateKeyFor(agentId);
99
+ if (!key) return { agentType: null, durationS: null };
100
+ const p = join(STATE_DIR, key + '.json');
101
+ let agentType = null;
102
+ let durationS = null;
103
+ try {
104
+ const parsed = JSON.parse(readFileSync(p, 'utf8'));
105
+ if (parsed && typeof parsed.ts === 'number') durationS = Math.max(0, (Date.now() - parsed.ts) / 1000);
106
+ if (parsed && typeof parsed.agent_type === 'string' && parsed.agent_type) agentType = parsed.agent_type;
107
+ } catch {
108
+ /* no state file, or it was unreadable: none, not an error */
109
+ }
110
+ try {
111
+ unlinkSync(p);
112
+ } catch {
113
+ /* already gone */
114
+ }
115
+ return { agentType, durationS };
116
+ }
117
+
118
+ function appendLog(record) {
119
+ try {
120
+ mkdirSync(HOME_DIR, { recursive: true });
121
+ let size = 0;
122
+ try {
123
+ size = statSync(LOG_FILE).size;
124
+ } catch {
125
+ /* file does not exist yet: size stays 0 */
126
+ }
127
+ if (size > LOG_ROTATE_BYTES) {
128
+ try {
129
+ renameSync(LOG_FILE, LOG_FILE + '.1');
130
+ } catch {
131
+ /* a concurrent rotation losing this race is not worth failing over */
132
+ }
133
+ }
134
+ appendFileSync(LOG_FILE, JSON.stringify(record) + '\n');
135
+ } catch {
136
+ /* telemetry never blocks the run */
137
+ }
138
+ }
139
+
140
+ // Parses the LAST <!-- route: <lane> | <why> --> marker out of text. The
141
+ // "why" half is captured only to be discarded: it is never read into a
142
+ // variable that reaches the log. Returns an array of lane tokens (split on
143
+ // "+", the documented way to log more than one lane from a single marker),
144
+ // or ["missing"] when there is no marker at all.
145
+ export function extractLane(text) {
146
+ if (typeof text !== 'string' || text.length === 0) return ['missing'];
147
+ const re = /<!--\s*route:\s*([^|>]*)\|[^>]*-->/g;
148
+ let match;
149
+ let last = null;
150
+ while ((match = re.exec(text)) !== null) last = match;
151
+ if (!last) return ['missing'];
152
+ // A token carrying any character outside the charset is logged as
153
+ // "invalid", never stripped into a plausible-looking lane: stripping
154
+ // `main","evil":"1` would log a lane named "mainevil1" that nobody chose.
155
+ const parts = last[1]
156
+ .split('+')
157
+ .map((s) => s.trim())
158
+ .filter(Boolean)
159
+ .map((t) => (t.length > TOKEN_MAX_LEN || /[^A-Za-z0-9_.-]/.test(t) ? 'invalid' : t));
160
+ return parts.length ? parts : ['missing'];
161
+ }
162
+
163
+ // Turns one parsed hook-input object into a log record, or null when the
164
+ // event is not one this hook measures (or PreToolUse fired for a tool other
165
+ // than Agent/Task, which the settings matcher should already have excluded;
166
+ // this is a defensive second check, not the primary gate).
167
+ export function buildRecord(input, now = () => new Date().toISOString()) {
168
+ if (!input || typeof input !== 'object') return null;
169
+ const sessionId = sanitize(input.session_id, SESSION_ID_MAX_LEN) || 'unknown';
170
+ const ts = now();
171
+ const base = { ts, v: 1 };
172
+
173
+ switch (input.hook_event_name) {
174
+ case 'UserPromptSubmit':
175
+ return { ...base, event: 'turn', session_id: sessionId };
176
+
177
+ case 'PreToolUse': {
178
+ if (input.tool_name !== 'Agent' && input.tool_name !== 'Task') return null;
179
+ const toolInput = (input.tool_input && typeof input.tool_input === 'object') ? input.tool_input : {};
180
+ const subagentType = sanitize(toolInput.subagent_type, TOKEN_MAX_LEN) || 'general-purpose';
181
+ const background = toolInput.run_in_background === true;
182
+ return { ...base, event: 'dispatch', session_id: sessionId, subagent_type: subagentType, background };
183
+ }
184
+
185
+ case 'SubagentStart': {
186
+ pruneOldState();
187
+ const agentType = sanitize(input.agent_type, TOKEN_MAX_LEN) || 'unknown';
188
+ recordStart(input.agent_id, agentType);
189
+ return { ...base, event: 'start', session_id: sessionId, agent_type: agentType };
190
+ }
191
+
192
+ case 'SubagentStop': {
193
+ const { agentType, durationS } = consumeStart(input.agent_id);
194
+ const record = { ...base, event: 'end', session_id: sessionId };
195
+ if (agentType) record.agent_type = agentType;
196
+ if (durationS !== null) record.duration_s = durationS;
197
+ return record;
198
+ }
199
+
200
+ case 'Stop':
201
+ return { ...base, event: 'route', session_id: sessionId, lane: extractLane(input.last_assistant_message) };
202
+
203
+ default:
204
+ return null;
205
+ }
206
+ }
207
+
208
+ // Drain stdin without ever blocking on it, bounded by BOTH time and size. A
209
+ // bare `readFileSync(0)` waits for EOF, so a caller that pipes in and never
210
+ // closes its end left the process running indefinitely (the same class of
211
+ // bug route-gate.mjs and subagent-context.mjs already fix). The size cap is
212
+ // this hook's own addition: hook input is normally small, so a payload past
213
+ // the cap is treated as truncated and parsed as nothing, never partially.
214
+ function drainStdinBounded(timeoutMs, maxBytes) {
215
+ return new Promise((resolve) => {
216
+ let settled = false;
217
+ let bytes = 0;
218
+ let truncated = false;
219
+ const chunks = [];
220
+ const finish = () => {
221
+ if (settled) return;
222
+ settled = true;
223
+ clearTimeout(timer);
224
+ try {
225
+ process.stdin.removeAllListeners('data');
226
+ process.stdin.removeAllListeners('end');
227
+ process.stdin.removeAllListeners('error');
228
+ process.stdin.pause();
229
+ } catch {
230
+ /* stdin may already be gone */
231
+ }
232
+ resolve({ data: truncated ? null : Buffer.concat(chunks).toString('utf8'), truncated });
233
+ };
234
+ const timer = setTimeout(finish, timeoutMs);
235
+ if (timer.unref) timer.unref();
236
+ try {
237
+ process.stdin.on('data', (chunk) => {
238
+ if (truncated) return;
239
+ bytes += chunk.length;
240
+ if (bytes > maxBytes) {
241
+ truncated = true;
242
+ return finish();
243
+ }
244
+ chunks.push(chunk);
245
+ });
246
+ process.stdin.on('end', finish);
247
+ process.stdin.on('error', finish);
248
+ process.stdin.resume();
249
+ } catch {
250
+ finish();
251
+ }
252
+ });
253
+ }
254
+
255
+ async function runHook() {
256
+ const { data } = await drainStdinBounded(STDIN_DRAIN_MS, STDIN_MAX_BYTES);
257
+ if (data) {
258
+ let input;
259
+ try {
260
+ input = JSON.parse(data);
261
+ } catch {
262
+ input = null; // invalid JSON: log nothing
263
+ }
264
+ if (input) {
265
+ try {
266
+ const record = buildRecord(input);
267
+ if (record) appendLog(record);
268
+ } catch {
269
+ /* telemetry never blocks or fails the run */
270
+ }
271
+ }
272
+ }
273
+ process.exit(0); // fail-open, always: a miss here is a missing log line, never a blocked turn
274
+ }
275
+
276
+ // ---- --summary: a plain-text report, no stdin involved ----
277
+
278
+ function parseLines(text) {
279
+ const records = [];
280
+ for (const line of text.split('\n')) {
281
+ const trimmed = line.trim();
282
+ if (!trimmed) continue;
283
+ try {
284
+ records.push(JSON.parse(trimmed));
285
+ } catch {
286
+ /* one bad line (a torn write, a rotation race) does not sink the report */
287
+ }
288
+ }
289
+ return records;
290
+ }
291
+
292
+ function formatNumber(n) {
293
+ return Number.isInteger(n) ? String(n) : n.toFixed(2);
294
+ }
295
+
296
+ function runSummary(args) {
297
+ if (!existsSync(LOG_FILE)) {
298
+ console.log('route-metrics: no data yet (' + LOG_FILE + ' does not exist).');
299
+ return process.exit(0);
300
+ }
301
+ const sinceIdx = args.indexOf('--since');
302
+ const since = sinceIdx !== -1 ? Date.parse(args[sinceIdx + 1]) : NaN;
303
+ let records = parseLines(readFileSync(LOG_FILE, 'utf8'));
304
+ if (!Number.isNaN(since)) records = records.filter((r) => Date.parse(r.ts) >= since);
305
+
306
+ const turns = records.filter((r) => r.event === 'turn').length;
307
+ const routes = records.filter((r) => r.event === 'route');
308
+ const covered = routes.filter((r) => !(Array.isArray(r.lane) && r.lane.length === 1 && r.lane[0] === 'missing')).length;
309
+ const coveragePct = turns > 0 ? (covered / turns) * 100 : null;
310
+
311
+ const laneCounts = new Map();
312
+ for (const r of routes) {
313
+ for (const lane of Array.isArray(r.lane) ? r.lane : []) laneCounts.set(lane, (laneCounts.get(lane) || 0) + 1);
314
+ }
315
+
316
+ const dispatches = records.filter((r) => r.event === 'dispatch');
317
+ const dispatchCounts = new Map();
318
+ for (const r of dispatches) dispatchCounts.set(r.subagent_type, (dispatchCounts.get(r.subagent_type) || 0) + 1);
319
+
320
+ const starts = records.filter((r) => r.event === 'start').length;
321
+ const noMatchingStart = Math.max(0, dispatches.length - starts);
322
+
323
+ const ends = records.filter((r) => r.event === 'end' && r.agent_type && typeof r.duration_s === 'number');
324
+ const durationsByType = new Map();
325
+ for (const r of ends) {
326
+ if (!durationsByType.has(r.agent_type)) durationsByType.set(r.agent_type, []);
327
+ durationsByType.get(r.agent_type).push(r.duration_s);
328
+ }
329
+
330
+ const lines = [];
331
+ lines.push('route-metrics summary' + (Number.isNaN(since) ? '' : ' since ' + args[sinceIdx + 1]));
332
+ lines.push('turns: ' + turns);
333
+ lines.push('route-marker coverage: ' + (coveragePct === null ? 'no turns yet' : formatNumber(coveragePct) + '%') + ' (' + covered + '/' + turns + ')');
334
+ lines.push('lanes by count:');
335
+ if (laneCounts.size === 0) lines.push(' (none)');
336
+ for (const [lane, count] of [...laneCounts.entries()].sort((a, b) => b[1] - a[1])) lines.push(' ' + lane + ': ' + count);
337
+ lines.push('dispatches by subagent_type:');
338
+ if (dispatchCounts.size === 0) lines.push(' (none)');
339
+ for (const [type, count] of [...dispatchCounts.entries()].sort((a, b) => b[1] - a[1])) lines.push(' ' + type + ': ' + count);
340
+ lines.push('dispatches with no matching start: ' + noMatchingStart + ' (a hook or guard blocked them before launch)');
341
+ lines.push('duration by agent_type (mean / max, seconds):');
342
+ if (durationsByType.size === 0) lines.push(' (none)');
343
+ for (const [type, durs] of durationsByType) {
344
+ const mean = durs.reduce((a, b) => a + b, 0) / durs.length;
345
+ lines.push(' ' + type + ': ' + formatNumber(mean) + ' / ' + formatNumber(Math.max(...durs)));
346
+ }
347
+ console.log(lines.join('\n'));
348
+ process.exit(0);
349
+ }
350
+
351
+ const args = process.argv.slice(2);
352
+ if (args.includes('--summary')) {
353
+ runSummary(args);
354
+ } else {
355
+ runHook();
356
+ }
@@ -7,6 +7,23 @@
7
7
  "type": "command",
8
8
  "command": "node",
9
9
  "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-gate.mjs"]
10
+ },
11
+ {
12
+ "type": "command",
13
+ "command": "node",
14
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-metrics.mjs"]
15
+ }
16
+ ]
17
+ }
18
+ ],
19
+ "PreToolUse": [
20
+ {
21
+ "matcher": "Agent|Task",
22
+ "hooks": [
23
+ {
24
+ "type": "command",
25
+ "command": "node",
26
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-metrics.mjs"]
10
27
  }
11
28
  ]
12
29
  }
@@ -18,6 +35,33 @@
18
35
  "type": "command",
19
36
  "command": "node",
20
37
  "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/subagent-context.mjs"]
38
+ },
39
+ {
40
+ "type": "command",
41
+ "command": "node",
42
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-metrics.mjs"]
43
+ }
44
+ ]
45
+ }
46
+ ],
47
+ "SubagentStop": [
48
+ {
49
+ "hooks": [
50
+ {
51
+ "type": "command",
52
+ "command": "node",
53
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-metrics.mjs"]
54
+ }
55
+ ]
56
+ }
57
+ ],
58
+ "Stop": [
59
+ {
60
+ "hooks": [
61
+ {
62
+ "type": "command",
63
+ "command": "node",
64
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-metrics.mjs"]
21
65
  }
22
66
  ]
23
67
  }
@@ -29,8 +29,8 @@ Modifiers:
29
29
 
30
30
  ## The two checkpoints (every build)
31
31
 
32
- - **Checkpoint 1, before writing anything.** You map the blast radius yourself (files, systems, docs, tickets). Then ask the deep tier, on the finished map: *is this the simplest way, what is the single biggest risk, where is the request as filed wrong?* It must return a named risk and a named flaw. Approval alone is not an answer.
33
- - **Checkpoint 2, after the build is green.** Security-shaped diffs get an adversarial read (in a fresh context, told to attack, allowed to answer CLEAN). Architecture-shaped diffs get the deep tier reviewing build against plan. Never both on one diff. Every finding reproduced before it reaches a human.
32
+ - **Checkpoint 1, before writing anything.** You map everything it touches yourself (files, systems, docs, tickets). Then ask the deep tier, on the finished map: *is this the simplest way, what is most likely to go wrong, what did the request miss?* It must return one named weak spot and one gap in the request. Approval alone is not an answer.
33
+ - **Checkpoint 2, after the build is green.** Security-shaped diffs get a second-opinion read (in a fresh context, told to challenge, allowed to answer CLEAN). Architecture-shaped diffs get the deep tier reviewing build against plan. Never both on one diff. Every finding reproduced before it reaches a human.
34
34
 
35
35
  Cap: two deep-tier consults per build. The full procedure is `protocols/build-protocol.md`.
36
36
 
@@ -47,7 +47,7 @@ Level 2 adds `ROUTING.md`, `TIERS.md`, `DELEGATION_MATRIX.md`, `RESEARCH_TRIAGE.
47
47
  ## The three rules that carry everything
48
48
 
49
49
  1. **Route by capability tier, not by model name.** deep = ambiguous or expensive to get wrong · standard = well-specified execution and review · fast = bulk and mechanical. Default down, escalate on evidence.
50
- 2. **A gate you cannot fail is not a gate.** "Does it look good?" passes every time. "Name the single biggest risk and the flaw in the request" can come back empty, which is how you know it worked.
50
+ 2. **A gate you cannot fail is not a gate.** "Does it look good?" passes every time. "Name what is most likely to go wrong, and what the request missed" can come back empty, which is how you know it worked.
51
51
  3. **Exit 0 is not a deliverable.** Any tool, CLI or subagent can report success and hand back nothing. Check for the artifact, not the status line.
52
52
 
53
53
  ## Where things went
@@ -2,7 +2,7 @@
2
2
 
3
3
  **Three phases, eight stages, and every gate is a question that can be answered wrong.**
4
4
 
5
- Fires on any task that builds, codes, implements, migrates or deploys. Rough test: if it would earn an adversarial audit or a tracker issue, it runs this.
5
+ Fires on any task that builds, codes, implements, migrates or deploys. Rough test: if it would earn a second-opinion audit or a tracker issue, it runs this.
6
6
 
7
7
  > **The one rule underneath:** a gate you cannot fail is not a gate. If a stage's exit reads like "confirm it looks good", it is written wrong and it will pass every time, including the times it should not.
8
8
 
@@ -14,7 +14,7 @@ Three corollaries:
14
14
  | Phase | Master question | Stages |
15
15
  |---|---|---|
16
16
  | 1 Pre-build | What exactly are we building, what do we need first, and what does this touch or break? | 0 Route · 1 Map · 2 Judge |
17
- | 2 Build | Is it secure, built on current code, and correct without hidden flaws? | 3 Build · 4 Scan · 5 Attack · 5b Ship gate |
17
+ | 2 Build | Is it secure, built on current code, and correct without hidden flaws? | 3 Build · 4 Scan · 5 Challenge · 5b Ship gate |
18
18
  | 3 Post-build | Did it land everywhere, is it proven against the real thing, and is it recorded? | 6 Verify · 7 Record |
19
19
 
20
20
  The two seams are the point. Pre-build to Build: nothing is written yet, changing your mind costs a conversation. Build to Post-build: the ship, the only irreversible step, the only one that needs an explicit human yes.
@@ -38,9 +38,9 @@ Four bounded questions, not four exhaustive scans. **The builder maps; the judgm
38
38
  ### Stage 2 · Judge (Checkpoint 1)
39
39
  Ask the judgment tier, on the finished map:
40
40
  1. Is this the simplest way to build it, or are we overcomplicating?
41
- 2. What is the single biggest risk, and where is the request as filed wrong?
41
+ 2. What is most likely to go wrong, and what did the request miss?
42
42
 
43
- **Gate:** a **named risk** and a **named flaw in the request**. Approval alone is not an exit; an advisor asked only to approve will approve. If a consult comes back mostly restating the map, the brief asked it to retrieve when it should have asked it to decide.
43
+ **Gate:** **one named weak spot** and **one gap in the request**. Approval alone is not an exit; an advisor asked only to approve will approve. If a consult comes back mostly restating the map, the brief asked it to retrieve when it should have asked it to decide.
44
44
 
45
45
  ## Phase 2 · Build
46
46
 
@@ -56,15 +56,15 @@ Ask the judgment tier, on the finished map:
56
56
  1. Any secret, key or token in the new code?
57
57
  2. Any vulnerability or vulnerable dependency in the lines we added?
58
58
 
59
- Secret detection, static analysis and dependency scanning, filtered to lines this diff added. Fail closed: a missing or erroring scanner exits non-zero, never a silent green.
59
+ Secret detection, static analysis and dependency scanning, filtered to lines this diff added. Refuses by default: a missing or erroring scanner exits non-zero, never a silent green.
60
60
 
61
61
  **Gate:** zero flags on added lines. Pre-existing flags are reported, never inherited as blockers, and never waved through unread. A scanner finding is a claim; read the code before calling it anything.
62
62
 
63
- ### Stage 5 · Attack (Checkpoint 2, one pass, never two)
63
+ ### Stage 5 · Challenge (Checkpoint 2, one pass, never two)
64
64
  1. Can bad input or a bad actor break it, and what happens when a dependency fails?
65
65
  2. Did the build stick to the approved plan, or did unintended changes sneak in?
66
66
 
67
- Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to an adversarial auditor, ideally a **different model family**. Architecture-shaped diffs go to the judgment tier reviewing build against plan. Never both on one diff.
67
+ Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to a second-opinion reviewer, ideally a **different model family**. Architecture-shaped diffs go to the judgment tier reviewing build against plan. Never both on one diff.
68
68
 
69
69
  Findings do not go straight to a repair. Hand them to the **finding-verifier**, whose job is to DISPROVE each one: read the cited line, state what would trigger it, then hunt for the guard, caller or test that makes it impossible. It returns one of three verdicts per finding, and rounding between them is the failure mode to watch for.
70
70
 
@@ -106,7 +106,7 @@ Use a different model family from the one that produced the finding where you ha
106
106
  |---|---|---|
107
107
  {{ROLES_BUILDER_ROW}}
108
108
  | Judgment tier | Stage 2 and the architectural arm of Stage 5. Argues with a finished map | Perform the retrieval |
109
- | Adversarial auditor | The security arm of Stage 5. Attacks the diff | Fix anything |
109
+ | Second-opinion reviewer | The security arm of Stage 5. Reviews the diff | Fix anything |
110
110
  | Mechanical gates | Stage 4 and any always-on guard | Be overridden without reading |
111
111
  | Cheap workers | Bounded sub-parts: bulk passes, wide searches, long loops | Own a stage |
112
112
  | Human | Stage 5b, and any irreversible or architectural call | Be the first line of review |
@@ -119,9 +119,9 @@ Use a different model family from the one that produced the finding where you ha
119
119
  PRE-BUILD
120
120
  [ ] 0 Inputs and access verified by live probe, not assumed
121
121
  [ ] 0 Confirmed this is a build and not a quick fix
122
- [ ] 1 Blast radius written: files, systems, issues
122
+ [ ] 1 Everything it touches written: files, systems, issues
123
123
  [ ] 1 Asked what could break, and whether this already exists
124
- [ ] 2 Judgment tier named a risk AND a flaw in the request
124
+ [ ] 2 Judgment tier named a weak spot AND a gap in the request
125
125
 
126
126
  BUILD
127
127
  [ ] 3 Repo clean, on a branch, base ref recorded