model-orchestrator 0.1.13 → 0.1.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (34) hide show
  1. package/AGENTS.md +26 -0
  2. package/CHANGELOG.md +41 -1
  3. package/README.md +95 -6
  4. package/bin/cli-run.mjs +144 -19
  5. package/bin/cli.js +1 -1
  6. package/docs/audit-brief.md +21 -0
  7. package/docs/part-1-beginner.md +3 -1
  8. package/docs/part-2-intermediate.md +8 -2
  9. package/llms.txt +27 -0
  10. package/package.json +32 -4
  11. package/src/README.md +1 -1
  12. package/src/catalog.js +9 -1
  13. package/src/install.js +188 -5
  14. package/templates/README.md +1 -1
  15. package/templates/agents/README.md +3 -1
  16. package/templates/agents/agy/README.md +2 -2
  17. package/templates/agents/agy/done-verifier.md +35 -0
  18. package/templates/agents/agy/finding-verifier.md +29 -0
  19. package/templates/agents/agy/reader.md +22 -0
  20. package/templates/agents/claude-code/README.md +7 -4
  21. package/templates/agents/claude-code/builder.md +6 -1
  22. package/templates/agents/claude-code/done-verifier.md +44 -0
  23. package/templates/agents/claude-code/finding-verifier.md +43 -0
  24. package/templates/agents/claude-code/reader.md +26 -0
  25. package/templates/agents/snippets/claude-code.md +9 -3
  26. package/templates/agents/snippets/route-gate.mjs +151 -0
  27. package/templates/agents/snippets/settings.hooks.snippet.json +26 -0
  28. package/templates/agents/snippets/subagent-context.mjs +76 -0
  29. package/templates/beginner/ORCHESTRATOR.md +4 -3
  30. package/templates/common/TASK_BUNDLE.md +2 -2
  31. package/templates/common/protocols/build-protocol.md +13 -3
  32. package/templates/intermediate/CLI-RUN.md +40 -3
  33. package/templates/intermediate/ROUTING.md +15 -11
  34. package/templates/intermediate/TIERS.md +44 -1
@@ -26,16 +26,22 @@ One driver, no second AI in the mix: the orchestrator invokes the CLIs; it never
26
26
 
27
27
  Every agent CLI can report success and deliver nothing. `bin/cli-run.mjs` builds the right invocation per lane, reads that lane's **native** terminal event, and exits `10` when a run produced no deliverable, `12` on timeout, `13` when the lane is missing. Byte count is not a check either; a run can emit hundreds of kilobytes and no conclusion. One lane's own success flags lie outright (an upstream 400 reported as success), so its judge reads the two honest signals instead.
28
28
 
29
- Every call goes through it. "This lane is flaky" becomes a query over its log instead of an argument. `node bin/cli-run.mjs --doctor` is the first thing to run after install: enabled lanes, binaries on PATH, and with `--run` a one-word canary per lane.
29
+ Every call goes through it. "This lane is flaky" becomes a query over its log instead of an argument. `node bin/cli-run.mjs --doctor` is the first thing to run after install: enabled lanes, binaries on PATH, the route each lane is pinned to, and with `--run` a one-word canary per lane.
30
+
31
+ There is a second thing a lane can be quietly wrong about. Left unpinned, it runs on **its own config file**, which the runner cannot see: a CLI set up months ago at a low reasoning effort keeps auditing at that effort while your routing docs describe an adversarial pass, and no error is ever raised. `--model` and `--effort` pin it per call, `defaults` in `bin/lanes.json` pins it per lane, and every run records the value requested and where it came from (`flag`, `lanes.json`, `lane_default`). The log claims no actual: grok reports a model id in its output, the other four lanes report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string the durable log never holds.
30
32
 
31
33
  ## 4. Every delegation carries a task bundle, on both surfaces
32
34
 
33
- Subagents and CLI lanes are the same problem: something with none of your rules and broad tool access. The brief (purpose, task class, scope, capabilities, denied actions, conventions, report contract, exit parameters) goes in the prompt or in the file passed to `--brief`. If you can, gate it mechanically: a pre-dispatch hook that refuses a brief missing purpose, denied actions or a report contract.
35
+ Subagents and CLI lanes are close to the same problem: something that may hold none of your rules, and broad tool access. A Claude Code subagent is the one documented exception, loading the project's CLAUDE.md hierarchy at start, so it keeps the standing rules but not this task's scope; a CLI lane and a fresh chat window get no such credit. The brief (purpose, task class, scope, capabilities, denied actions, conventions, report contract, exit parameters) goes in the prompt or in the file passed to `--brief` either way. If you can, gate it mechanically: a pre-dispatch hook that refuses a brief missing purpose, denied actions or a report contract. On claude-code, a `SubagentStart` hook can inject the essentials (where the rules and the brief format live) automatically; `.claude/hooks/subagent-context.mjs` is the generated example.
34
36
 
35
37
  ## 5. Research: three engines, one triager
36
38
 
37
39
  Fan the same plan to three model families (web sweep, adversarial read, live data), each as one `cli-run` call. The orchestrator opens the primary sources itself, marks every claim, and writes the only durable record. Expect one engine to return confident unsourced numerics; downgrade it. Weight the engines that report their own gaps. Count dispositions, not briefs.
38
40
 
41
+ ## 5a. A finding is a claim, not a fact
42
+
43
+ An audit that returns six findings has returned six claims. Hand them to `finding-verifier` before any of them causes a repair: it reads the cited line, states what would trigger the problem, then hunts for the guard, caller or test that makes it impossible, and answers CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE. Only CONFIRMED earns a change. Use a different family from the one that produced the finding, and let INCONCLUSIVE stand: rounding it up to be safe buys unnecessary repairs, rounding it down to be tidy hides real ones.
44
+
39
45
  ## 6. Gap analysis gets a second family
40
46
 
41
47
  The second pass is now a different model reading the same artifact, in read-only audit mode. Disagreement between families is the cheapest signal that something is soft.
package/llms.txt ADDED
@@ -0,0 +1,27 @@
1
+ # model-orchestrator
2
+
3
+ > Model orchestrator for AI coding agents and LLMs (Claude Code, Codex, Gemini, Grok, Qwen, Ollama). One installer writes routing rules, subagent definitions and a CLI lane runner for the AI tools you already have. The rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. It is not a proxy or gateway: it does not automatically compare prices or select models; your agent follows the rules and chooses.
4
+
5
+ Install and run: `npx model-orchestrator` (interactive), or headless: `npx model-orchestrator --yes --level 2 --ais claude-code,codex --project . --dir ./ai-orchestrator`. Preview without writing: add `--dry-run`. List every supported AI: `npx model-orchestrator --list`. Node 18 or newer, zero runtime dependencies, MIT licence.
6
+
7
+ Levels: 1 beginner (one agent or chat app), 2 intermediate (several agent CLIs, each called through `cli-run`), 3 advanced (adds a virtual machine with a gateway and a scheduled audit job). On Claude Code the install also delegates execution to subagents by default and ships two hooks (UserPromptSubmit, SubagentStart) that inject the routing table every turn.
8
+
9
+ ## Docs
10
+
11
+ - [README](https://github.com/aunysillyme/model-orchestrator/blob/main/README.md): what it writes, flags, principles, what is enforced versus instructed
12
+ - [Part 1: beginner](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-1-beginner.md): one agent, tiers, the task bundle every delegation carries
13
+ - [Part 2: intermediate](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-2-intermediate.md): several agent CLIs, the lane runner, pinning model and effort per lane
14
+ - [Part 3: advanced](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/part-3-advanced.md): the VM, the gateway, scheduled jobs, privacy gates
15
+ - [AI catalog](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/catalog.md): every supported AI, how to sign in, how each is detected
16
+
17
+ ## Reference
18
+
19
+ - [CLI runner](https://github.com/aunysillyme/model-orchestrator/blob/main/bin/README.md): `cli-run` lanes, exit codes, the route logged per run
20
+ - [Templates](https://github.com/aunysillyme/model-orchestrator/blob/main/templates/README.md): the routing, tiers, task bundle and protocol files the installer renders
21
+ - [Changelog](https://github.com/aunysillyme/model-orchestrator/blob/main/CHANGELOG.md): every release and the issue behind each fix
22
+ - [Agent instructions](https://github.com/aunysillyme/model-orchestrator/blob/main/AGENTS.md): running the installer from an agent, and contributing
23
+
24
+ ## Optional
25
+
26
+ - [Security policy](https://github.com/aunysillyme/model-orchestrator/blob/main/SECURITY.md)
27
+ - [Audit brief](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/audit-brief.md): the threat model and what has already been attacked
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.13",
4
- "description": "Routing instructions and a CLI runner for your AI tools. One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine.",
3
+ "version": "0.1.15",
4
+ "description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
5
5
  "type": "module",
6
6
  "bin": {
7
7
  "model-orchestrator": "bin/cli.js"
@@ -15,7 +15,9 @@
15
15
  "LICENSE",
16
16
  "CHANGELOG.md",
17
17
  "SECURITY.md",
18
- "scripts"
18
+ "scripts",
19
+ "llms.txt",
20
+ "AGENTS.md"
19
21
  ],
20
22
  "scripts": {
21
23
  "start": "node bin/cli.js",
@@ -51,7 +53,33 @@
51
53
  "agents",
52
54
  "subagents",
53
55
  "qwen",
54
- "mcp"
56
+ "mcp",
57
+ "reasoning-effort",
58
+ "model-routing",
59
+ "code-review",
60
+ "ai-agents",
61
+ "hooks",
62
+ "claude-code-hooks",
63
+ "delegation",
64
+ "llm-router",
65
+ "ai-router",
66
+ "model-switching",
67
+ "multi-model",
68
+ "token-optimization",
69
+ "token-efficiency",
70
+ "cost-optimization",
71
+ "llm-orchestration",
72
+ "agent-orchestration",
73
+ "ai-orchestration",
74
+ "claude",
75
+ "anthropic",
76
+ "openai",
77
+ "gemini",
78
+ "claude-code-router",
79
+ "agent-routing",
80
+ "prompt-routing",
81
+ "coding-agent",
82
+ "ai-coding"
55
83
  ],
56
84
  "author": "aunysillyme (https://github.com/aunysillyme)",
57
85
  "license": "MIT"
package/src/README.md CHANGED
@@ -4,6 +4,6 @@
4
4
  |---|---|
5
5
  | `catalog.js` | the single list of levels and AIs. Add an AI here and the prompts, docs tables, delegation matrix, gateway config and installer all pick it up. Nothing else lists AIs. |
6
6
  | `detect.js` | PATH lookup for a binary, plus the few places vendor installers drop binaries without touching PATH. No shell-outs. |
7
- | `install.js` | pure planner: turns (level, selection, primary) into a list of files to write, rendering templates and computing every generated table. `writeFiles` is the only thing that touches disk. `activationSteps()` and `snippetFor()` live here so the terminal summary and the generated README render the same list. |
7
+ | `install.js` | pure planner: turns (level, selection, primary) into a list of files to write, rendering templates and computing every generated table. `writeFiles` is the only thing that touches disk. `activationSteps()` and `snippetFor()` live here so the terminal summary and the generated README render the same list. `subagentsLoadRules(primary)` gates every delegate-by-default render var (builder-by-default wording, the route-gate table, the inline-threshold note) on the one verified premise: a claude-code subagent loads CLAUDE.md. |
8
8
  | `prompt.js` | line-buffered questions for the interactive path; piped answers are queued, EOF mid-prompt aborts instead of confirming a write. |
9
9
  | `render.js` | `{{KEY}}` substitution. Throws on an unknown key, so a template typo fails the test suite instead of shipping a literal placeholder. |
package/src/catalog.js CHANGED
@@ -69,7 +69,15 @@ export const AIS = [
69
69
  rulesFile: 'CLAUDE.md',
70
70
  agentsDir: '.claude/agents',
71
71
  cliRun: false,
72
- models: { deep: 'opus', standard: 'sonnet', fast: 'haiku' }
72
+ models: { deep: 'opus', standard: 'sonnet', fast: 'haiku' },
73
+ // Verified at code.claude.com/docs/en/sub-agents (fetched 2026-09-10): "A
74
+ // non-fork subagent's initial context contains: CLAUDE.md files: every
75
+ // level of the CLAUDE.md hierarchy the main conversation loads ... The
76
+ // built-in Explore and Plan agents skip this." No other lane in this
77
+ // catalog has that documented, so the builder-by-default routing, the
78
+ // route-gate hook and the inline-threshold note are gated on this field
79
+ // and stay claude-code only.
80
+ subagentsLoadRules: true
73
81
  },
74
82
  {
75
83
  id: 'codex',
package/src/install.js CHANGED
@@ -193,6 +193,146 @@ export function laneVars(selected) {
193
193
  };
194
194
  }
195
195
 
196
+ // Only claude-code has a verified sub-agents doc quote saying its subagents
197
+ // load the project's CLAUDE.md hierarchy (code.claude.com/docs/en/sub-agents,
198
+ // see the comment on the catalog entry). Every delegate-by-default surface below
199
+ // (builder-by-default wording, the route-gate hook, the inline-threshold
200
+ // note) is gated on this so a primary with no verified premise keeps the
201
+ // original, more conservative wording.
202
+ export function subagentsLoadRules(primary) {
203
+ return !!(primary && primary.subagentsLoadRules);
204
+ }
205
+
206
+ // Canonical agent order, tier-first. Used to render a stable, non-hardcoded
207
+ // "available as" list for the claude-code snippet from the files actually
208
+ // shipped, so a future agent addition or removal cannot leave the sentence
209
+ // stale the way the finding-verifier omission did.
210
+ const AGENT_ORDER = ['deep-planner', 'builder', 'code-reviewer', 'finding-verifier', 'live-researcher', 'bulk-worker', 'done-verifier', 'reader'];
211
+ export function claudeAgentIds() {
212
+ const dir = join(TEMPLATES, 'agents', 'claude-code');
213
+ if (!existsSync(dir)) return [];
214
+ const files = readdirSync(dir).filter((f) => f.endsWith('.md') && f !== 'README.md').map((f) => f.replace(/\.md$/, ''));
215
+ const set = new Set(files);
216
+ const ordered = AGENT_ORDER.filter((id) => set.has(id));
217
+ const extra = files.filter((id) => !AGENT_ORDER.includes(id)).sort();
218
+ return [...ordered, ...extra];
219
+ }
220
+
221
+ // The compact "pick the lane before acting" table, rendered from the AIs the
222
+ // user actually selected and the agents actually installed, never a second
223
+ // hand-typed copy of ROUTING.md's decision tree.
224
+ export function routeGateTable(selected) {
225
+ const rows = [
226
+ ['Bulk or mechanical, many similar items', 'bulk-worker'],
227
+ ['Needs live data', 'live-researcher'],
228
+ ['Review without changing', 'code-reviewer'],
229
+ ['Findings from a review or a scanner', 'finding-verifier, before any repair'],
230
+ ['Reading or digesting many files or notes', 'reader'],
231
+ ['Checking a tracker item against its stated done-signal', 'done-verifier'],
232
+ ['Ambiguous, architectural, expensive to get wrong', 'deep-planner'],
233
+ ['Everything else that changes files', 'builder, by default']
234
+ ];
235
+ for (const a of selected.filter((x) => x.cliRun)) rows.push([a.role, '`cli-run ' + a.id + '`']);
236
+ return table(rows, ['Task', 'Lane']);
237
+ }
238
+
239
+ // The marked block route-gate.mjs extracts at runtime. Installed only for
240
+ // claude-code so the hook always finds a block to read; other primaries get
241
+ // no hook and so get no block.
242
+ export function routeGateSection(selected) {
243
+ return [
244
+ '<!-- route-gate:start -->',
245
+ '## Route gate: pick the lane before acting',
246
+ '',
247
+ 'Injected on every turn by the `route-gate` hook, so this table is read at runtime rather than recalled from memory.',
248
+ '',
249
+ routeGateTable(selected),
250
+ '',
251
+ "Stay inline only when: (a) the brief would cost as much as the work itself, (b) the task needs this conversation's own context, (c) it is the human's decision or the final verification of delegated work (a delegate never verifies itself).",
252
+ '',
253
+ 'Never: the built-in Explore or Plan agents for rule-bound work (they skip CLAUDE.md). general-purpose taking work a named agent already owns.',
254
+ '<!-- route-gate:end -->'
255
+ ].join('\n');
256
+ }
257
+
258
+ // ROUTING.md / ORCHESTRATOR.md decision-tree rule 5 and the "Who builds"
259
+ // section read differently for claude-code, because only claude-code has the
260
+ // verified premise that its subagents load CLAUDE.md. Every other primary
261
+ // keeps the original wording: the orchestrator builds the main line directly
262
+ // and a subagent or second CLI is assumed to hold none of these rules.
263
+ export function decisionRule5(primary) {
264
+ return subagentsLoadRules(primary)
265
+ ? `5. **Everything else that changes files** → builder executes by default. The orchestrator plans, briefs, verifies and talks to the human; it stays inline only when (a) the brief would cost as much as the work, (b) the task needs this conversation's own context, or (c) it is the human's decision, or the final verification of delegated work (a delegate never verifies itself). Never route rule-bound work to the built-in Explore or Plan agents: both skip CLAUDE.md. general-purpose should not take work a named agent already owns.`
266
+ : `5. **Everything else that changes files** → the orchestrator builds it directly. Bounded sub-parts go to cheaper tiers; the main build is never handed off whole.`;
267
+ }
268
+ export function decisionRule5Beginner(primary) {
269
+ return subagentsLoadRules(primary)
270
+ ? `5. **Everything else that changes files or executes a known plan** → builder executes by default, at standard tier. The orchestrator plans, briefs, verifies and talks to you; it stays inline only when (a) the brief would cost as much as the work, (b) the task needs this conversation's own context, or (c) it is your decision, or the final verification of delegated work. Never route rule-bound work to the built-in Explore or Plan agents: both skip CLAUDE.md.`
271
+ : `5. **Everything else that changes files or executes a known plan** → you build it directly, at standard tier. The main build is never handed off whole; bounded sub-parts (a bulk pass, a wide search, a long audit loop) can go to cheaper tiers.`;
272
+ }
273
+ export function whoBuildsSection(primary) {
274
+ if (subagentsLoadRules(primary)) {
275
+ return [
276
+ '## Who builds',
277
+ '',
278
+ `**Builder executes by default.** A Claude Code subagent loads this project's CLAUDE.md hierarchy at start (verified: code.claude.com/docs/en/sub-agents), so it already carries the standing rules; the orchestrator's job is to plan, brief, verify and talk to the human, not to hold work a delegate can do. Stay inline only when: (a) the brief would cost as much as the work itself, (b) the task needs this conversation's own context, or (c) it is the human's decision to make, or the final verification of delegated work (a delegate never verifies its own output as final). Never route rule-bound work to the built-in Explore or Plan agents: both skip CLAUDE.md and the git status the router depends on. general-purpose should not take work a named agent already owns.`,
279
+ '',
280
+ `Delegate: the main build, background and long-running tasks, small tasks, scoping, verification, research, bounded sub-parts. Never delegate: the human's own decision, or the final sign-off on a delegate's work.`,
281
+ '',
282
+ `Every delegation carries \`TASK_BUNDLE.md\`. Its brief must restate this task's scope: a Claude Code subagent already has the standing rules, just not that.`
283
+ ].join('\n');
284
+ }
285
+ return [
286
+ '## Who builds',
287
+ '',
288
+ '**The orchestrator owns the main build.** It is the only surface that holds these rules: a subagent or a second CLI starts with none of them and cannot route. Handing the main build to one hands it to something the router cannot reach.',
289
+ '',
290
+ 'Delegate: background and long-running tasks, small tasks, scoping, verification, research, bounded sub-parts. Never delegate: the main build, or any step that must carry a house rule (secrets handling, the loud-negative verification, the durable record).',
291
+ '',
292
+ 'Every delegation carries `TASK_BUNDLE.md`. Its brief must restate every convention the delegate needs.'
293
+ ].join('\n');
294
+ }
295
+ export function addEndpointRow(primary) {
296
+ return subagentsLoadRules(primary)
297
+ ? '| "Add an endpoint" | builder, briefed and verified by the orchestrator |'
298
+ : '| "Add an endpoint" | the orchestrator builds it |';
299
+ }
300
+ export function inlineThresholdNote(primary) {
301
+ return subagentsLoadRules(primary)
302
+ ? '\n- **Measure your inline threshold once.** A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Spawn one with a one-line task and read its token count. Work smaller than that stays inline.'
303
+ : '';
304
+ }
305
+ export function delegateRulesNote(primary) {
306
+ return subagentsLoadRules(primary)
307
+ ? `Subagents, a fresh chat, a second window: a Claude Code subagent loads this project's CLAUDE.md hierarchy, so it holds the standing rules already, just not this task's scope; a second CLI or a fresh chat window may hold none of them.`
308
+ : 'Subagents, a fresh chat, a second window: each one holds none of these rules.';
309
+ }
310
+
311
+ // Pre-release audit finding 3: the delegate-by-default gate reached the
312
+ // decision tree and "Who builds" but missed three other generated surfaces
313
+ // stating the same old premise (the orchestrator writes the main build
314
+ // itself; a delegate inherits none of the session's rules). These three
315
+ // close that gap the same way: gated on subagentsLoadRules(primary), every
316
+ // other primary keeps the original wording unchanged.
317
+ export function planBigExecuteSmallLine(primary) {
318
+ return subagentsLoadRules(primary)
319
+ ? `- **Plan big, execute small**, within a build: deep tier plans at Checkpoint 1, builder executes from the orchestrator's brief, bulk and wide searches go down.`
320
+ : '- **Plan big, execute small**, within a build: deep tier plans at Checkpoint 1, the orchestrator executes, bulk and wide searches go down.';
321
+ }
322
+ export function rolesBuilderRow(primary) {
323
+ return subagentsLoadRules(primary)
324
+ ? [
325
+ '| Orchestrator | Routes, maps, briefs, verifies, records. Stages 0, 1, 2, 4, 5b, 6, 7 | Write the build |',
326
+ "| Builder | Executes Stage 3 from the orchestrator's brief | Route further, or verify its own work as final |"
327
+ ].join('\n')
328
+ : '| Builder / orchestrator | Routes, maps, writes, verifies, records. Stages 0, 1, 3, 6, 7 | Hand off the main build |';
329
+ }
330
+ export function builderHandoffNote(primary) {
331
+ return subagentsLoadRules(primary)
332
+ ? `**Why Stage 3 goes to builder by default:** a Claude Code subagent loads this project's CLAUDE.md hierarchy at start, so it already carries the standing rules; the orchestrator's brief only has to restate this task's scope (see \`TASK_BUNDLE.md\`). The orchestrator keeps Stage 3 for itself only when the brief would cost as much as the work, the task needs this conversation's own context, or it is the human's decision or the final verification of delegated work.`
333
+ : `**Why the builder does not hand off the main build:** a delegated agent does not inherit the session's standing rules and usually cannot delegate further. Any brief must restate every convention it needs (see \`TASK_BUNDLE.md\`), and that cost is itself a reason to build directly when the work fits.`;
334
+ }
335
+
196
336
  // Which activation file this primary gets. ONE decision, read by three
197
337
  // surfaces: planFiles writes the file, vars() names it in the generated README,
198
338
  // and bin/cli.js prints it in the terminal. Before 0.1.12 the README hardcoded
@@ -218,6 +358,9 @@ export function activationSteps(opts) {
218
358
  // replaces (#22).
219
359
  else if (snippet) steps.push(`open ${primary.chatName || primary.name} and paste the block in ${join(dirAbs, snippet)} into its ${primary.chatSurface || 'custom instructions'}`);
220
360
  if (primary && primary.agentsDir) steps.push(`subagents are in ${join(projectAbs, primary.agentsDir)}; run ${primary.bin} from ${projectAbs} to pick them up`);
361
+ // Only claude-code ships hooks (route-gate, subagent-context): the wiring
362
+ // lives in a snippet, never written into a settings.json the user already has.
363
+ if (subagentsLoadRules(primary)) steps.push(`merge the hooks in ${join(dirAbs, 'settings.hooks.snippet.json')} into ${join(projectAbs, '.claude', 'settings.json')} (create it if missing) to wire the route-gate and subagent-context hooks`);
221
364
  for (const a of selected.filter((a) => a.bin && a.kind === 'agent-cli')) steps.push(`sign in to ${a.name}: ${a.auth}`);
222
365
  // A local runtime has a bin but no sign-in, so the agent-cli loop above skips it
223
366
  // and before this it appeared in no ordered list at any level (#26).
@@ -232,15 +375,22 @@ export function activationSteps(opts) {
232
375
  // activationSteps is: level 1 writes no bin/, so a step naming cli-run.mjs or
233
376
  // lanes.json there described an install that did not happen (#27).
234
377
  export function proofSteps(opts) {
235
- const { level } = opts;
378
+ const { level, primary } = opts;
236
379
  const steps = [
237
380
  'Start a fresh agent session and ask: "Read the orchestrator instructions. Quote the routing rule you will use, then sort pear, apple, banana alphabetically. Name the tier and whether you delegated."',
238
381
  'Expect the fast tier and `apple, banana, pear`. If the agent cannot quote the routing rule, check the snippet location or chat instructions before continuing. This is a manual activation check, not proof that every future task follows the rules.'
239
382
  ];
240
383
  if (level >= 2) {
241
- steps.push('Run `node bin/cli-run.mjs --doctor` from this folder. It checks binary presence, not authentication or loaded instructions. `--doctor --run` additionally uses a little quota to test live responses. No enabled lanes means delegation is inactive.');
384
+ steps.push('Run `node bin/cli-run.mjs --doctor` from this folder. It checks binary presence, not authentication or loaded instructions, and prints the model and effort each lane is pinned to. `--doctor --run` additionally uses a little quota to test live responses. No enabled lanes means delegation is inactive.');
385
+ steps.push('Decide whether the route matters to you. Every lane starts unpinned, which means it runs on whatever its own config file says: a CLI configured months ago at a low reasoning effort will keep auditing at that effort while your docs describe something stronger. Pin it in `bin/lanes.json` under `defaults`, or per call with `--model` and `--effort`. Either way the run is recorded in the log with the value requested and where it came from.');
242
386
  steps.push('To test a real output contract, choose an enabled lane from `bin/lanes.json` and run `node bin/cli-run.mjs <lane> \'Return only {"sorted":["apple","banana","pear"]}\' --expect-json`. This uses quota. Expect JSON and exit 0; inspect the array yourself. A non-JSON response exits 10, a missing binary exits 13, and an authentication failure reports the vendor error. The explicit lane tests execution; your primary agent still makes delegation decisions.');
243
387
  }
388
+ // Only claude-code ships the route-gate hook, so only claude-code gets a
389
+ // proof step that checks it fired: the table must come from the hook's
390
+ // injected context, not from the agent reciting ROUTING.md from memory.
391
+ if (subagentsLoadRules(primary)) {
392
+ steps.push('Ask the agent: "Quote the route-gate table you were given this turn." It should quote the injected table verbatim, not recite it from memory. If it cannot, the hooks snippet was not merged into `.claude/settings.json`, or the hook found no rules file: check both before trusting the routing docs are actually reaching the agent.');
393
+ }
244
394
  return steps;
245
395
  }
246
396
 
@@ -259,7 +409,14 @@ function vars(opts) {
259
409
  const pinOf = (id) => (toolById[id] && toolById[id].pin) || 'latest';
260
410
  const snippet = snippetFor(primary);
261
411
  const steps = activationSteps({ level, selected, primary, tools, dir: opts.dir, project: opts.project });
262
- const proofs = proofSteps({ level });
412
+ const proofs = proofSteps({ level, primary });
413
+ const routingFile = level >= 2 ? 'ROUTING.md' : 'ORCHESTRATOR.md';
414
+ // The path route-gate.mjs and subagent-context.mjs resolve at runtime,
415
+ // relative to CLAUDE_PROJECT_DIR. Mirrors the RULES_PATH fallback below:
416
+ // outside the project, the honest path is absolute, never a hardcoded one.
417
+ const relJoin = (name) => (rulesPath === dirAbs ? join(dirAbs, name) : rulesPath === '.' ? name : rulesPath + '/' + name);
418
+ const rulesFileRel = relJoin(routingFile);
419
+ const taskBundleRel = relJoin('TASK_BUNDLE.md');
263
420
  // Only claude-code and agy put files under the project root. A chat primary
264
421
  // puts nothing there, so naming a project root would name a folder this run
265
422
  // never created (#21).
@@ -335,7 +492,24 @@ function vars(opts) {
335
492
  NPM_PACKAGES: selected.map(npmSpec).filter(Boolean).join(' ') || '""',
336
493
  SCRIPT_INSTALLERS: scriptInstallers(selected),
337
494
  COMPOSE_ENV: composeEnv(selected, apis),
338
- COMPOSE_OLLAMA: composeOllama(selected)
495
+ COMPOSE_OLLAMA: composeOllama(selected),
496
+ // Delegate by default (0.1.15): gated on subagentsLoadRules(primary), currently
497
+ // claude-code only. Every other primary keeps the original, more
498
+ // conservative wording these replace.
499
+ DECISION_RULE5: decisionRule5(primary),
500
+ DECISION_RULE5_L1: decisionRule5Beginner(primary),
501
+ WHO_BUILDS: whoBuildsSection(primary),
502
+ ADD_ENDPOINT_ROW: addEndpointRow(primary),
503
+ INLINE_THRESHOLD_NOTE: inlineThresholdNote(primary),
504
+ DELEGATE_RULES_NOTE: delegateRulesNote(primary),
505
+ PLAN_BIG_LINE: planBigExecuteSmallLine(primary),
506
+ ROLES_BUILDER_ROW: rolesBuilderRow(primary),
507
+ BUILDER_HANDOFF_NOTE: builderHandoffNote(primary),
508
+ ROUTE_GATE_SECTION: subagentsLoadRules(primary) ? '\n' + routeGateSection(selected) + '\n' : '',
509
+ AGENTS_LIST_LINE: claudeAgentIds().map((id) => '`' + id + '`').join(', '),
510
+ RULES_FILE_REL: rulesFileRel,
511
+ RULES_FILE_REL_JSON: JSON.stringify(rulesFileRel),
512
+ TASK_BUNDLE_REL_JSON: JSON.stringify(taskBundleRel)
339
513
  };
340
514
  }
341
515
 
@@ -365,6 +539,13 @@ export function planFiles(opts) {
365
539
  add(join('.claude', 'agents', f.rel), render(readFileSync(f.abs, 'utf8'), v), 0o644, 'project');
366
540
  }
367
541
  add('CLAUDE.snippet.md', render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'claude-code.md'), 'utf8'), v));
542
+ // Delegate-by-default hooks (0.1.15), claude-code only: route-gate.mjs (UserPromptSubmit)
543
+ // and subagent-context.mjs (SubagentStart) live where Claude Code looks for
544
+ // project hooks; the wiring snippet is a document the user merges in, never
545
+ // written into a settings.json they already have.
546
+ add(join('.claude', 'hooks', 'route-gate.mjs'), render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'route-gate.mjs'), 'utf8'), v), 0o755, 'project');
547
+ add(join('.claude', 'hooks', 'subagent-context.mjs'), render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'subagent-context.mjs'), 'utf8'), v), 0o755, 'project');
548
+ add('settings.hooks.snippet.json', render(readFileSync(join(TEMPLATES, 'agents', 'snippets', 'settings.hooks.snippet.json'), 'utf8'), v));
368
549
  } else if (primary && primary.id === 'agy') {
369
550
  for (const f of walk(join(TEMPLATES, 'agents', 'agy'))) {
370
551
  if (!installable('agents', f.rel)) continue;
@@ -389,7 +570,9 @@ export function planFiles(opts) {
389
570
  JSON.stringify(
390
571
  {
391
572
  enabled: selected.filter((a) => a.cliRun).map((a) => a.id),
392
- note: 'Lanes cli-run may call. Edit to enable or disable a lane. A lane not listed here exits 13 (unavailable).'
573
+ defaults: {},
574
+ note: 'Lanes cli-run may call. Edit to enable or disable a lane. A lane not listed here exits 13 (unavailable).',
575
+ defaultsNote: 'Pin what a lane runs with, so the route in your docs is the route that runs: "defaults": {"codex": {"model": "gpt-6-astra", "effort": "high"}}. Left empty, a lane inherits its own config file, which cli-run cannot see and does not guess. `--model` and `--effort` override this per call, and `--doctor` prints what each lane is pinned to. Every lane takes a model; every lane except qwen takes an effort.'
393
576
  },
394
577
  null,
395
578
  2
@@ -6,7 +6,7 @@ Everything the installer can write, organized by the level that adds it. Files a
6
6
  |---|---|---|
7
7
  | `common/` | every level | the start-here README, `TASK_BUNDLE.md`, `protocols/` (build, propagate, gap analysis, deep research, numbers and logic, memory and record) |
8
8
  | `beginner/` | every level | `ORCHESTRATOR.md`, the single-agent routing rules |
9
- | `agents/` | every level, one variant | the primary agent's loading surface: Claude Code subagents, Antigravity custom agents, or a paste snippet |
9
+ | `agents/` | every level, one variant | the primary agent's loading surface: Claude Code subagents (plus `.claude/hooks/route-gate.mjs` and `subagent-context.mjs`, and `settings.hooks.snippet.json` to wire them in), Antigravity custom agents, or a paste snippet |
10
10
  | `intermediate/` | level 2+ | `ROUTING.md`, `TIERS.md`, `DELEGATION_MATRIX.md`, `RESEARCH_TRIAGE.md`, `CLI-RUN.md` |
11
11
  | `advanced/` | level 3 | `vm/`: gateway config, compose file, box rules, privacy gates, scheduled jobs |
12
12
  | `tools/` | when selected | companion tools the AIs call: `codecalc/` and `obsidian-tc/` (install doc + MCP snippets each). See `tools/README.md` |
@@ -10,4 +10,6 @@ Loading surfaces for the primary agent. The installer writes exactly one of thes
10
10
  | `grok`, `hermes` | nothing agent-specific | rules travel with the prompt or the task bundle |
11
11
  | a chat app | `PASTE-INTO-YOUR-AGENT.md` | no files to load; paste into custom instructions |
12
12
 
13
- `snippets/` are rendered with the chosen agent's name and rules file. Nothing here is appended to a file the user already has.
13
+ `snippets/` are rendered with the chosen agent's name and rules file. Nothing here is appended to a file the user already has. `snippets/route-gate.mjs`, `snippets/subagent-context.mjs`, and `snippets/settings.hooks.snippet.json` are claude-code only: two hooks and the settings block that wires them, installed to `.claude/hooks/` and next to `CLAUDE.snippet.md`.
14
+
15
+ `claude-code/` and `agy/` both ship the same agent set: one per tier, plus `finding-verifier`, `done-verifier` and `reader`. Add an agent to one folder and its README, and the other.
@@ -1,5 +1,5 @@
1
1
  # .agents/agents/
2
2
 
3
- Antigravity CLI custom agents, one per tier, in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
3
+ Antigravity CLI custom agents, one per tier plus three checks (`finding-verifier`, `done-verifier`, `reader`), in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
4
4
 
5
- `commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; `auto` keeps deletes and other high-risk commands gated) and `off` for the read-only agents. `model` is a tier: `pro` for deep-planner, `flash` for the rest.
5
+ `commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; `auto` keeps deletes and other high-risk commands gated) and `off` for the read-only agents: `code-reviewer`, `finding-verifier`, `live-researcher`, `done-verifier`, `reader`. `model` is a tier: `pro` for deep-planner, `flash` for the rest. `done-verifier` and `reader` never write and, with `commandExecutionPolicy: off`, cannot execute any command at all here, mutating or not: unlike its claude-code counterpart, which does carry an unrestricted `Bash` and stays read-only by its prompt rather than by the tool grant, agy's `done-verifier` is mechanically blocked from shelling out and probes artifacts through whatever read or fetch capability it has instead. Neither is `bulk-worker`, which classifies, tags and transforms items and does write.
@@ -0,0 +1,35 @@
1
+ ---
2
+ name: done-verifier
3
+ description: Checks tracker items or tasks against their stated done-signal by probing the named artifact (a file, a commit, a URL, a log line, a count); no file-editing tools, no command execution (commandExecutionPolicy off); returns MET, NOT_MET or UNVERIFIABLE per item; never closes or edits anything.
4
+ model: flash
5
+ subagent: true
6
+ mainAgent: true
7
+ commandExecutionPolicy: off
8
+ ---
9
+
10
+ # done-verifier
11
+
12
+ Checks tracker items or tasks against their stated done-signal by probing the
13
+ named artifact. No file-editing tools, and no command execution: this agent's
14
+ `commandExecutionPolicy` is `off`, so unlike its claude-code counterpart it
15
+ cannot shell out at all, not even to a read-only command; probe with whatever
16
+ read or fetch capability you have instead.
17
+
18
+ For each item: read the stated done-signal, probe the exact artifact it
19
+ names, compare what you found against the claim.
20
+
21
+ Return one verdict per item:
22
+ - MET: the artifact matches the claim. Name what you checked.
23
+ - NOT_MET: the artifact is missing or contradicts the claim. Name what you
24
+ found instead.
25
+ - UNVERIFIABLE: you cannot probe it from here, no done-signal was stated, or
26
+ the check would need a command you are not able to run. Say what is
27
+ missing.
28
+
29
+ Rules:
30
+ - Stay inside the task bundle you were given. Anything not granted is denied.
31
+ - Never close, edit or comment on a tracker item; return verdicts only.
32
+ - If the only way to check something would mutate it, or would need command
33
+ execution you do not have, the item is UNVERIFIABLE, not MET.
34
+ - Token discipline: read only the cited artifact, hand back verdicts not
35
+ narration.
@@ -0,0 +1,29 @@
1
+ ---
2
+ name: finding-verifier
3
+ description: Adversarial verification of review findings; tries to disprove each one and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE. Read-only, never repairs.
4
+ model: flash
5
+ subagent: true
6
+ mainAgent: true
7
+ commandExecutionPolicy: off
8
+ ---
9
+
10
+ # finding-verifier
11
+
12
+ A finding is a claim, not a fact. You try to disprove each one before it is
13
+ allowed to cause a repair.
14
+
15
+ For each finding you are given: read the cited file and line yourself, state the
16
+ input or sequence that would trigger it, then hunt for what makes it impossible
17
+ (a guard upstream, a caller that never passes that value, an existing test).
18
+
19
+ Return one verdict per finding, in the order given:
20
+ - CONFIRMED: reproduced, or a concrete unblocked path. Give the path.
21
+ - NOT_REPRODUCED: you found what stops it. Name it and where it is.
22
+ - INCONCLUSIVE: not settleable read-only. Say what you would need.
23
+
24
+ Rules:
25
+ - Stay inside the task bundle you were given. Anything not granted is denied.
26
+ - Verify only the findings handed to you; anything else you notice goes at the end, marked unverified.
27
+ - Never round INCONCLUSIVE up to CONFIRMED to be safe, or down to NOT_REPRODUCED to be tidy.
28
+ - Read-only: you never repair and never reword a finding.
29
+ - Token discipline: read the cited code and its callers, not the repository.
@@ -0,0 +1,22 @@
1
+ ---
2
+ name: reader
3
+ description: Reads and digests many files or notes and returns facts, quotes with source, an index or a digest. Read-only. Different from bulk-worker, which classifies, tags and transforms items: reader only reads and reports.
4
+ model: flash
5
+ subagent: true
6
+ mainAgent: true
7
+ commandExecutionPolicy: off
8
+ ---
9
+
10
+ # reader
11
+
12
+ Reads and digests many files or notes and hands back exactly what the brief
13
+ asks for: facts, quotes, an index, a digest. Does not classify, tag,
14
+ transform or rewrite; that is bulk-worker's job, and reader never writes a
15
+ file.
16
+
17
+ Rules:
18
+ - Stay inside the task bundle you were given. Anything not granted is denied.
19
+ - Cite every fact or quote with its source (path or URL).
20
+ - Report what you did, what you did not do, and what you could not verify.
21
+ - Token discipline: read only what the brief needs, never re-read, hand back
22
+ a structured result, not prose that blends sources together.
@@ -1,13 +1,16 @@
1
1
  # .claude/agents/
2
2
 
3
- Five subagents, one per tier. Claude Code loads project-level agents from this folder automatically.
3
+ One per tier, plus two checks and two agents with no file-editing tools: `finding-verifier` sits between a review and a repair, `done-verifier` sits between a claim of "done" and a tracker close, and `reader` digests many files or notes without writing anything. Claude Code loads project-level agents from this folder automatically; the count is whatever this folder holds; `test/install.test.js` ties the claude-code snippet's agent list to the files actually shipped here, so this table cannot drift silently.
4
4
 
5
5
  | Agent | Tier | Model alias | Effort | Job |
6
6
  |---|---|---|---|---|
7
7
  | deep-planner | deep | opus | xhigh | judges every build twice; never retrieves |
8
- | builder | standard | sonnet | high | bounded sub-parts of a build |
8
+ | builder | standard | sonnet | high | executes; the default for everything that changes files |
9
9
  | code-reviewer | standard | sonnet | high | read-only findings |
10
+ | finding-verifier | standard | sonnet | high | tries to disprove a finding before it causes a repair |
10
11
  | live-researcher | standard | sonnet | medium | fresh data through tools |
11
- | bulk-worker | fast | haiku | low | mechanical volume |
12
+ | bulk-worker | fast | haiku | low | mechanical volume, writes output |
13
+ | done-verifier | fast | haiku | low | probes a tracker item's stated done-signal; no file-editing tools, Bash for probes only |
14
+ | reader | fast | haiku | low | reads and digests many files or notes; read-only |
12
15
 
13
- Aliases resolve to the newest model in each family, so a version bump needs no edit here. Each agent carries its own token-discipline rule; the `effort` field is the third cost lever.
16
+ Aliases resolve to the newest model in each family, so a version bump needs no edit here. Each agent carries its own token-discipline rule; the `effort` field is the third cost lever. Neither `done-verifier` nor `reader` carries `Write` or `Edit` in its `tools:` line. `reader` is read-only by tool grant as well: it carries no `Bash`. `done-verifier` does carry `Bash`, for its probes (`git log`, `grep`, `wc -l`, `test -f`); nothing in that grant stops it from running a command that changes state, so staying read-only there is a rule in its prompt, not a restriction on the tool, and its own file says so.
@@ -1,12 +1,17 @@
1
1
  ---
2
2
  name: builder
3
- description: Well-specified execution of a bounded sub-part. Use for writing code, editing files, wiring configs, and implementing a plan that already exists. Do not use for open-ended architecture questions, bulk classification, or the main build itself.
3
+ description: Executes builds by default on this router, including the main build, from a brief the orchestrator wrote. Use for writing code, editing files, wiring configs, running commands, and implementing a plan the orchestrator briefed. Do not use for open-ended architecture questions or bulk classification; those still go to deep-planner or bulk-worker.
4
4
  model: sonnet
5
5
  effort: high
6
6
  ---
7
7
 
8
8
  You are the execution tier of the model router.
9
9
 
10
+ The orchestrator stays inline only when the brief would cost as much as the
11
+ work, the task needs this conversation's own context, or it is the human's
12
+ decision or the final verification of delegated work. Everything else that
13
+ changes files, the main build included, comes to you.
14
+
10
15
  You implement specs and plans: write code, edit files, run commands.
11
16
 
12
17
  Rules:
@@ -0,0 +1,44 @@
1
+ ---
2
+ name: done-verifier
3
+ description: Checks tracker items or tasks against their stated done-signal. Use after work is claimed finished, to probe the named artifact (a file, a commit, a URL, a log line, a count) before a tracker item is closed. No file-editing tools; Bash is for read-only probes, bound by the prompt below, not by the tool grant. Returns MET, NOT_MET or UNVERIFIABLE per item, and never closes or edits anything itself.
4
+ tools: Read, Glob, Grep, Bash
5
+ model: haiku
6
+ effort: low
7
+ ---
8
+
9
+ You are the done-signal verification tier of the model router.
10
+
11
+ A tracker item is not done because someone said it is done; it is done because
12
+ its stated done-signal is true. Your job is to probe the artifact the
13
+ done-signal names, not to judge the work more broadly.
14
+
15
+ You carry no Write or Edit tool, so you cannot touch a file. You do carry
16
+ Bash, and nothing in that grant stops you from running a command that changes
17
+ state; staying to read-only checks is a rule you follow below, not a
18
+ restriction you were given. Treat that boundary as load-bearing.
19
+
20
+ For each item you are given:
21
+ 1. Read the stated done-signal. If there is none, or it only restates the
22
+ title, say so; that is a finding, not a thing to guess past.
23
+ 2. Probe the exact artifact it names: read the file, check the commit exists,
24
+ describe the URL, grep the log line, count what it says to count.
25
+ 3. Compare what you found against what the signal claims.
26
+
27
+ Return one verdict per item, in the order given:
28
+ - **MET**: the artifact exists and matches the claim. Name what you checked.
29
+ - **NOT_MET**: the artifact is missing, contradicts the claim, or the check
30
+ failed. Name what you found instead.
31
+ - **UNVERIFIABLE**: you cannot probe the artifact from here (behind a login,
32
+ on a machine you cannot reach, no done-signal stated). Say exactly what is
33
+ missing.
34
+
35
+ Rules:
36
+ - You never close, edit, or comment on a tracker item. You return verdicts;
37
+ something else acts on them.
38
+ - Verify only the items you were given. Anything else you notice goes in a
39
+ separate list at the end, marked unverified.
40
+ - Bash is for read-only checks only (`git log`, `grep`, `wc -l`, `test -f`, a
41
+ HEAD or GET request): never a command that changes state. If the only way
42
+ to check something would mutate it, that item is UNVERIFIABLE, not MET.
43
+ - Token discipline: read the cited artifact and nothing else; do not
44
+ summarize the whole tracker.