model-orchestrator 0.1.14 → 0.1.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -10,4 +10,6 @@ Loading surfaces for the primary agent. The installer writes exactly one of thes
10
10
  | `grok`, `hermes` | nothing agent-specific | rules travel with the prompt or the task bundle |
11
11
  | a chat app | `PASTE-INTO-YOUR-AGENT.md` | no files to load; paste into custom instructions |
12
12
 
13
- `snippets/` are rendered with the chosen agent's name and rules file. Nothing here is appended to a file the user already has.
13
+ `snippets/` are rendered with the chosen agent's name and rules file. Nothing here is appended to a file the user already has. `snippets/route-gate.mjs`, `snippets/subagent-context.mjs`, and `snippets/settings.hooks.snippet.json` are claude-code only: two hooks and the settings block that wires them, installed to `.claude/hooks/` and next to `CLAUDE.snippet.md`.
14
+
15
+ `claude-code/` and `agy/` both ship the same agent set: one per tier, plus `finding-verifier`, `done-verifier` and `reader`. Add an agent to one folder and its README, and the other.
@@ -1,5 +1,5 @@
1
1
  # .agents/agents/
2
2
 
3
- Antigravity CLI custom agents, one per tier plus a finding-verifier, in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
3
+ Antigravity CLI custom agents, one per tier plus three checks (`finding-verifier`, `done-verifier`, `reader`), in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
4
4
 
5
- `commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; `auto` keeps deletes and other high-risk commands gated) and `off` for the read-only agents. `model` is a tier: `pro` for deep-planner, `flash` for the rest.
5
+ `commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; `auto` keeps deletes and other high-risk commands gated) and `off` for the read-only agents: `code-reviewer`, `finding-verifier`, `live-researcher`, `done-verifier`, `reader`. `model` is a tier: `pro` for deep-planner, `flash` for the rest. `done-verifier` and `reader` never write and, with `commandExecutionPolicy: off`, cannot execute any command at all here, mutating or not: unlike its claude-code counterpart, which does carry an unrestricted `Bash` and stays read-only by its prompt rather than by the tool grant, agy's `done-verifier` is mechanically blocked from shelling out and probes artifacts through whatever read or fetch capability it has instead. Neither is `bulk-worker`, which classifies, tags and transforms items and does write.
@@ -0,0 +1,35 @@
1
+ ---
2
+ name: done-verifier
3
+ description: Checks tracker items or tasks against their stated done-signal by probing the named artifact (a file, a commit, a URL, a log line, a count); no file-editing tools, no command execution (commandExecutionPolicy off); returns MET, NOT_MET or UNVERIFIABLE per item; never closes or edits anything.
4
+ model: flash
5
+ subagent: true
6
+ mainAgent: true
7
+ commandExecutionPolicy: off
8
+ ---
9
+
10
+ # done-verifier
11
+
12
+ Checks tracker items or tasks against their stated done-signal by probing the
13
+ named artifact. No file-editing tools, and no command execution: this agent's
14
+ `commandExecutionPolicy` is `off`, so unlike its claude-code counterpart it
15
+ cannot shell out at all, not even to a read-only command; probe with whatever
16
+ read or fetch capability you have instead.
17
+
18
+ For each item: read the stated done-signal, probe the exact artifact it
19
+ names, compare what you found against the claim.
20
+
21
+ Return one verdict per item:
22
+ - MET: the artifact matches the claim. Name what you checked.
23
+ - NOT_MET: the artifact is missing or contradicts the claim. Name what you
24
+ found instead.
25
+ - UNVERIFIABLE: you cannot probe it from here, no done-signal was stated, or
26
+ the check would need a command you are not able to run. Say what is
27
+ missing.
28
+
29
+ Rules:
30
+ - Stay inside the task bundle you were given. Anything not granted is denied.
31
+ - Never close, edit or comment on a tracker item; return verdicts only.
32
+ - If the only way to check something would mutate it, or would need command
33
+ execution you do not have, the item is UNVERIFIABLE, not MET.
34
+ - Token discipline: read only the cited artifact, hand back verdicts not
35
+ narration.
@@ -0,0 +1,22 @@
1
+ ---
2
+ name: reader
3
+ description: Reads and digests many files or notes and returns facts, quotes with source, an index or a digest. Read-only. Different from bulk-worker, which classifies, tags and transforms items: reader only reads and reports.
4
+ model: flash
5
+ subagent: true
6
+ mainAgent: true
7
+ commandExecutionPolicy: off
8
+ ---
9
+
10
+ # reader
11
+
12
+ Reads and digests many files or notes and hands back exactly what the brief
13
+ asks for: facts, quotes, an index, a digest. Does not classify, tag,
14
+ transform or rewrite; that is bulk-worker's job, and reader never writes a
15
+ file.
16
+
17
+ Rules:
18
+ - Stay inside the task bundle you were given. Anything not granted is denied.
19
+ - Cite every fact or quote with its source (path or URL).
20
+ - Report what you did, what you did not do, and what you could not verify.
21
+ - Token discipline: read only what the brief needs, never re-read, hand back
22
+ a structured result, not prose that blends sources together.
@@ -1,14 +1,16 @@
1
1
  # .claude/agents/
2
2
 
3
- Six subagents. Five are one per tier; `finding-verifier` is the check that sits between a review and a repair. Claude Code loads project-level agents from this folder automatically.
3
+ One per tier, plus two checks and two agents with no file-editing tools: `finding-verifier` sits between a review and a repair, `done-verifier` sits between a claim of "done" and a tracker close, and `reader` digests many files or notes without writing anything. Claude Code loads project-level agents from this folder automatically; the count is whatever this folder holds; `test/install.test.js` ties the claude-code snippet's agent list to the files actually shipped here, so this table cannot drift silently.
4
4
 
5
5
  | Agent | Tier | Model alias | Effort | Job |
6
6
  |---|---|---|---|---|
7
7
  | deep-planner | deep | opus | xhigh | judges every build twice; never retrieves |
8
- | builder | standard | sonnet | high | bounded sub-parts of a build |
8
+ | builder | standard | sonnet | high | executes; the default for everything that changes files |
9
9
  | code-reviewer | standard | sonnet | high | read-only findings |
10
10
  | finding-verifier | standard | sonnet | high | tries to disprove a finding before it causes a repair |
11
11
  | live-researcher | standard | sonnet | medium | fresh data through tools |
12
- | bulk-worker | fast | haiku | low | mechanical volume |
12
+ | bulk-worker | fast | haiku | low | mechanical volume, writes output |
13
+ | done-verifier | fast | haiku | low | probes a tracker item's stated done-signal; no file-editing tools, Bash for probes only |
14
+ | reader | fast | haiku | low | reads and digests many files or notes; read-only |
13
15
 
14
- Aliases resolve to the newest model in each family, so a version bump needs no edit here. Each agent carries its own token-discipline rule; the `effort` field is the third cost lever.
16
+ Aliases resolve to the newest model in each family, so a version bump needs no edit here. Each agent carries its own token-discipline rule; the `effort` field is the third cost lever. Neither `done-verifier` nor `reader` carries `Write` or `Edit` in its `tools:` line. `reader` is read-only by tool grant as well: it carries no `Bash`. `done-verifier` does carry `Bash`, for its probes (`git log`, `grep`, `wc -l`, `test -f`); nothing in that grant stops it from running a command that changes state, so staying read-only there is a rule in its prompt, not a restriction on the tool, and its own file says so.
@@ -1,12 +1,17 @@
1
1
  ---
2
2
  name: builder
3
- description: Well-specified execution of a bounded sub-part. Use for writing code, editing files, wiring configs, and implementing a plan that already exists. Do not use for open-ended architecture questions, bulk classification, or the main build itself.
3
+ description: Executes builds by default on this router, including the main build, from a brief the orchestrator wrote. Use for writing code, editing files, wiring configs, running commands, and implementing a plan the orchestrator briefed. Do not use for open-ended architecture questions or bulk classification; those still go to deep-planner or bulk-worker.
4
4
  model: sonnet
5
5
  effort: high
6
6
  ---
7
7
 
8
8
  You are the execution tier of the model router.
9
9
 
10
+ The orchestrator stays inline only when the brief would cost as much as the
11
+ work, the task needs this conversation's own context, or it is the human's
12
+ decision or the final verification of delegated work. Everything else that
13
+ changes files, the main build included, comes to you.
14
+
10
15
  You implement specs and plans: write code, edit files, run commands.
11
16
 
12
17
  Rules:
@@ -0,0 +1,44 @@
1
+ ---
2
+ name: done-verifier
3
+ description: Checks tracker items or tasks against their stated done-signal. Use after work is claimed finished, to probe the named artifact (a file, a commit, a URL, a log line, a count) before a tracker item is closed. No file-editing tools; Bash is for read-only probes, bound by the prompt below, not by the tool grant. Returns MET, NOT_MET or UNVERIFIABLE per item, and never closes or edits anything itself.
4
+ tools: Read, Glob, Grep, Bash
5
+ model: haiku
6
+ effort: low
7
+ ---
8
+
9
+ You are the done-signal verification tier of the model router.
10
+
11
+ A tracker item is not done because someone said it is done; it is done because
12
+ its stated done-signal is true. Your job is to probe the artifact the
13
+ done-signal names, not to judge the work more broadly.
14
+
15
+ You carry no Write or Edit tool, so you cannot touch a file. You do carry
16
+ Bash, and nothing in that grant stops you from running a command that changes
17
+ state; staying to read-only checks is a rule you follow below, not a
18
+ restriction you were given. Treat that boundary as load-bearing.
19
+
20
+ For each item you are given:
21
+ 1. Read the stated done-signal. If there is none, or it only restates the
22
+ title, say so; that is a finding, not a thing to guess past.
23
+ 2. Probe the exact artifact it names: read the file, check the commit exists,
24
+ describe the URL, grep the log line, count what it says to count.
25
+ 3. Compare what you found against what the signal claims.
26
+
27
+ Return one verdict per item, in the order given:
28
+ - **MET**: the artifact exists and matches the claim. Name what you checked.
29
+ - **NOT_MET**: the artifact is missing, contradicts the claim, or the check
30
+ failed. Name what you found instead.
31
+ - **UNVERIFIABLE**: you cannot probe the artifact from here (behind a login,
32
+ on a machine you cannot reach, no done-signal stated). Say exactly what is
33
+ missing.
34
+
35
+ Rules:
36
+ - You never close, edit, or comment on a tracker item. You return verdicts;
37
+ something else acts on them.
38
+ - Verify only the items you were given. Anything else you notice goes in a
39
+ separate list at the end, marked unverified.
40
+ - Bash is for read-only checks only (`git log`, `grep`, `wc -l`, `test -f`, a
41
+ HEAD or GET request): never a command that changes state. If the only way
42
+ to check something would mutate it, that item is UNVERIFIABLE, not MET.
43
+ - Token discipline: read the cited artifact and nothing else; do not
44
+ summarize the whole tracker.
@@ -0,0 +1,26 @@
1
+ ---
2
+ name: reader
3
+ description: Reads and digests many files or notes and returns exactly what the brief asks for (facts, quotes with path:line, an index, a digest). Read-only. Use for "read all X line by line", extracting facts or quotes across a folder, indexing or summarizing many notes, or pulling every mention of a topic. Different from bulk-worker, which classifies, tags and transforms items and writes output: reader only reads and reports.
4
+ tools: Read, Glob, Grep
5
+ model: haiku
6
+ effort: low
7
+ ---
8
+
9
+ You are the reading tier of the model router.
10
+
11
+ You read and digest many files or notes and hand back exactly what the brief
12
+ asked for: facts, quotes, an index, a digest. You do not classify, tag,
13
+ transform or rewrite; that is bulk-worker's job, not yours, and you never
14
+ write a file.
15
+
16
+ Rules:
17
+ - Read the brief first and answer only what it asks. "Every mention of X"
18
+ means grep for X and read the hits, not the whole corpus.
19
+ - Cite every fact or quote with its source: `path:line` for code and notes, a
20
+ URL and a retrieval note for anything fetched.
21
+ - An index or digest is a structured list, one row or bullet per source, not
22
+ prose that blends sources together.
23
+ - If a source is missing, unreadable, or empty, say so by name; do not
24
+ silently skip it.
25
+ - Token discipline: read only what the brief needs, never re-read a file,
26
+ summarize as you go rather than holding full text for later.
@@ -11,12 +11,15 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
11
11
  2. Needs live data -> live-researcher (standard tier + tools).
12
12
  3. Review without changing -> code-reviewer (standard, read-only).
13
13
  3a. Holding findings from a review or scanner -> finding-verifier before any repair. Only CONFIRMED findings earn a change.
14
+ 3b. Checking a tracker item against its stated done-signal -> done-verifier. It never closes anything itself.
14
15
  4. Ambiguous, architectural, or expensive to get wrong -> deep-planner (deep tier), then hand the plan down.
15
- 5. Everything else that changes files -> build it directly. The main build is never handed off whole; bounded sub-parts go to builder.
16
+ 5. Everything else that changes files -> builder executes by default. The orchestrator plans, briefs, verifies and talks to you; it stays inline only when (a) the brief would cost as much as the work, (b) the task needs this conversation's own context, or (c) it is your decision, or the final verification of delegated work. Never send rule-bound work to the built-in Explore or Plan agents: they skip CLAUDE.md. general-purpose should not take work a named agent already owns.
17
+
18
+ A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Measure yours once: spawn a subagent with a one-line task and read its token count. Work smaller than that stays inline.
16
19
 
17
20
  Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one adversarial pass, an explicit human yes before anything irreversible, then the loud negative.
18
21
 
19
- Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A subagent holds none of these rules; absence is denial.
22
+ Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A Claude Code subagent loads this CLAUDE.md hierarchy, so it holds the standing rules already, just not this task's scope; a second CLI or a fresh chat window may hold none of them. Absence is denial either way.
20
23
 
21
24
  Never silently retry a failed attempt at the same tier. Escalate once and say so.
22
25
 
@@ -25,4 +28,6 @@ Numbers, comparisons, complexity and equivalence claims go through codecalc (or
25
28
  Anything durable is searched for before it is written and its folder index is corrected in the same pass; one writer per run: `{{RULES_PATH}}/protocols/memory-and-record.md`.
26
29
  ```
27
30
 
28
- Subagents were written to `{{AGENTS_DIR}}` (the project root, which is where Claude Code reads project-level agents; `--project` changes it). Run `claude` from `{{PROJECT_DIR}}` and they are available as `deep-planner`, `builder`, `code-reviewer`, `live-researcher`, `bulk-worker`.
31
+ Subagents were written to `{{AGENTS_DIR}}` (the project root, which is where Claude Code reads project-level agents; `--project` changes it). Run `claude` from `{{PROJECT_DIR}}` and they are available as {{AGENTS_LIST_LINE}}.
32
+
33
+ Two hooks were written to `{{AGENTS_DIR}}/../hooks/` (`.claude/hooks/`): `route-gate.mjs` injects the routing table on every prompt, and `subagent-context.mjs` reminds a spawned subagent where the rules and the task-bundle format live. Merge `settings.hooks.snippet.json`, written next to this file, into `.claude/settings.json` to wire them in.
@@ -0,0 +1,151 @@
1
+ #!/usr/bin/env node
2
+ // route-gate.mjs: UserPromptSubmit hook for {{PRIMARY_NAME}}.
3
+ //
4
+ // Reads the route-gate table out of {{RULES_FILE_REL}} and injects it as
5
+ // additionalContext on every turn, so the routing table is read at runtime
6
+ // from the one place it is generated (the rules file), never a second
7
+ // hand-typed copy that can drift from it.
8
+ //
9
+ // Fail-open by design: a miss here is a stray context string, not a gate.
10
+ // This script always exits 0, never blocks on stdin past a short bound,
11
+ // reads at most 64 KB of the rules file through a fixed-size buffer (never
12
+ // a full read of an arbitrarily large or non-regular file), and never
13
+ // executes anything it reads. See docs/audit-brief.md for the threat model.
14
+ import { statSync, openSync, readSync, closeSync, realpathSync } from 'node:fs';
15
+ import { join, isAbsolute } from 'node:path';
16
+
17
+ // Rendered at install time from the level and directory the user chose.
18
+ // Never hardcoded: a level 1 install points this at ORCHESTRATOR.md, level
19
+ // 2+ at ROUTING.md, and a --dir outside the project resolves to an absolute
20
+ // path instead of a relative one.
21
+ const RULES_FILE_REL = {{RULES_FILE_REL_JSON}};
22
+
23
+ const MAX_READ = 64 * 1024; // bounded read: this is a rules file, not a log
24
+ const MAX_CONTEXT = 4000; // bounded injection: a table, not the whole file
25
+ const STDIN_DRAIN_MS = 250; // hard cap: never let an open, never-closed stdin pipe hold this hook open
26
+ const START = '<!-- route-gate:start -->';
27
+ const END = '<!-- route-gate:end -->';
28
+
29
+ function fallback(reason) {
30
+ return 'route-gate: ' + reason + '. Pick the lane before acting: read ' + RULES_FILE_REL + ' yourself.';
31
+ }
32
+
33
+ function resolveRulesPath() {
34
+ if (isAbsolute(RULES_FILE_REL)) return RULES_FILE_REL;
35
+ const projectDir = process.env.CLAUDE_PROJECT_DIR;
36
+ if (!projectDir) return null;
37
+ // Resolve through whatever part of the project dir already exists, so a
38
+ // symlinked project folder still resolves to the real path the rules file
39
+ // was written under.
40
+ let root = projectDir;
41
+ try {
42
+ root = realpathSync(projectDir);
43
+ } catch {
44
+ /* keep the unresolved value; the read below reports the real failure */
45
+ }
46
+ return join(root, RULES_FILE_REL);
47
+ }
48
+
49
+ // Bounded, regular-file-only read. statSync (not lstatSync) follows a
50
+ // symlink to its target and reports what the target actually is, so a
51
+ // symlinked rules file still reads; a FIFO, socket, device or directory at
52
+ // the resolved path is refused before any open/read call touches it. That
53
+ // check matters: opening a FIFO for reading blocks until a writer opens the
54
+ // other end, and a plain readFileSync on any of these can hang or, for a
55
+ // huge or sparse regular file, allocate far more than this hook needs. The
56
+ // fixed-size buffer plus a single bounded readSync call means the on-disk
57
+ // size of the file never determines how much this hook reads or how long it
58
+ // takes.
59
+ function readBounded(path) {
60
+ let st;
61
+ try {
62
+ st = statSync(path);
63
+ } catch (e) {
64
+ throw Object.assign(new Error('could not stat ' + path + ' (' + ((e && e.code) || e) + ')'), { code: e && e.code });
65
+ }
66
+ if (!st.isFile()) throw new Error(path + ' is not a regular file');
67
+ const buf = Buffer.alloc(MAX_READ);
68
+ let fd;
69
+ try {
70
+ fd = openSync(path, 'r');
71
+ const bytesRead = readSync(fd, buf, 0, MAX_READ, 0);
72
+ return buf.toString('utf8', 0, bytesRead);
73
+ } finally {
74
+ if (fd !== undefined) closeSync(fd);
75
+ }
76
+ }
77
+
78
+ function computeContext() {
79
+ const path = resolveRulesPath();
80
+ if (!path) return fallback('CLAUDE_PROJECT_DIR is not set, so ' + RULES_FILE_REL + ' could not be located');
81
+
82
+ let text;
83
+ try {
84
+ text = readBounded(path);
85
+ } catch (e) {
86
+ return fallback((e && e.message) || String(e));
87
+ }
88
+
89
+ const s = text.indexOf(START);
90
+ const e = s === -1 ? -1 : text.indexOf(END, s);
91
+ if (s === -1 || e === -1) return fallback(path + ' has no route-gate block');
92
+
93
+ return text.slice(s, e + END.length).slice(0, MAX_CONTEXT);
94
+ }
95
+
96
+ // Drain stdin without ever blocking on it. A bare `readFileSync(0)` waits
97
+ // for stdin to reach EOF, so a caller that pipes into this hook and never
98
+ // closes its end of the pipe (or a bare TTY with no redirection at all)
99
+ // left the process running indefinitely. This races the real 'end' event
100
+ // against a hard timeout instead: whichever settles first wins, and the
101
+ // timer is unref'd so it can never itself be the reason the process stays
102
+ // alive past a normal exit.
103
+ function drainStdin(timeoutMs) {
104
+ return new Promise((resolve) => {
105
+ let settled = false;
106
+ const finish = () => {
107
+ if (settled) return;
108
+ settled = true;
109
+ clearTimeout(timer);
110
+ try {
111
+ process.stdin.removeAllListeners('data');
112
+ process.stdin.removeAllListeners('end');
113
+ process.stdin.removeAllListeners('error');
114
+ process.stdin.pause();
115
+ } catch {
116
+ /* stdin may already be gone; nothing left to clean up */
117
+ }
118
+ resolve();
119
+ };
120
+ const timer = setTimeout(finish, timeoutMs);
121
+ if (timer.unref) timer.unref();
122
+ try {
123
+ process.stdin.on('data', () => {});
124
+ process.stdin.on('end', finish);
125
+ process.stdin.on('error', finish);
126
+ process.stdin.resume();
127
+ } catch {
128
+ finish();
129
+ }
130
+ });
131
+ }
132
+
133
+ let additionalContext;
134
+ try {
135
+ additionalContext = computeContext();
136
+ } catch (err) {
137
+ additionalContext = fallback('route-gate.mjs failed unexpectedly (' + ((err && err.message) || err) + ')');
138
+ }
139
+
140
+ drainStdin(STDIN_DRAIN_MS).then(() => {
141
+ const payload = JSON.stringify({
142
+ hookSpecificOutput: {
143
+ hookEventName: 'UserPromptSubmit',
144
+ additionalContext
145
+ }
146
+ });
147
+ // Exit only after the write's callback fires, so a buffered write to a
148
+ // pipe (the common case on Windows, and possible anywhere output exceeds
149
+ // one write's worth) is not truncated by an exit that races ahead of it.
150
+ process.stdout.write(payload, () => process.exit(0));
151
+ });
@@ -0,0 +1,26 @@
1
+ {
2
+ "hooks": {
3
+ "UserPromptSubmit": [
4
+ {
5
+ "hooks": [
6
+ {
7
+ "type": "command",
8
+ "command": "node",
9
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-gate.mjs"]
10
+ }
11
+ ]
12
+ }
13
+ ],
14
+ "SubagentStart": [
15
+ {
16
+ "hooks": [
17
+ {
18
+ "type": "command",
19
+ "command": "node",
20
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/subagent-context.mjs"]
21
+ }
22
+ ]
23
+ }
24
+ ]
25
+ }
26
+ }
@@ -0,0 +1,76 @@
1
+ #!/usr/bin/env node
2
+ // subagent-context.mjs: SubagentStart hook for {{PRIMARY_NAME}}.
3
+ //
4
+ // A Claude Code subagent loads this project's CLAUDE.md hierarchy at start
5
+ // (code.claude.com/docs/en/sub-agents), so it already has the standing
6
+ // rules. What it does not have is this task's scope, and it can be tempted
7
+ // to route further work itself or to mark its own output verified. This
8
+ // hook injects a short, static reminder of where the rest lives and what
9
+ // "delegate" means.
10
+ //
11
+ // Fail-open by design: a miss here is a stray context string, not a gate.
12
+ // This script always exits 0, never executes anything it reads, and never
13
+ // blocks on stdin past a short bound (see drainStdin below).
14
+ import { isAbsolute } from 'node:path';
15
+
16
+ // Rendered at install time so a --dir outside the project still names an
17
+ // honest path rather than a hardcoded one.
18
+ const RULES_FILE_REL = {{RULES_FILE_REL_JSON}};
19
+ const TASK_BUNDLE_REL = {{TASK_BUNDLE_REL_JSON}};
20
+ const STDIN_DRAIN_MS = 250; // hard cap: never let an open, never-closed stdin pipe hold this hook open
21
+
22
+ const additionalContext = [
23
+ 'SUBAGENT CONTEXT (model-orchestrator).',
24
+ 'Routing rules: ' + RULES_FILE_REL + (isAbsolute(RULES_FILE_REL) ? '.' : ' (relative to the project root).'),
25
+ 'Task bundle format: ' + TASK_BUNDLE_REL + '.',
26
+ 'Report contract: say what you did, what you did NOT do, and what you could not verify. "Unverified" is acceptable; a confident guess is not. Stop at the bound your brief set, and never claim work you cannot show.',
27
+ 'You are a delegate: do not route further work to another subagent yourself, and do not mark your own output as the final verification of it.'
28
+ ].join(' ');
29
+
30
+ // Drain stdin without ever blocking on it. A bare `readFileSync(0)` waits
31
+ // for stdin to reach EOF, so a caller that pipes into this hook and never
32
+ // closes its end of the pipe left the process running indefinitely. This
33
+ // races the real 'end' event against a hard timeout instead: whichever
34
+ // settles first wins, and the timer is unref'd so it can never itself be
35
+ // the reason the process stays alive past a normal exit.
36
+ function drainStdin(timeoutMs) {
37
+ return new Promise((resolve) => {
38
+ let settled = false;
39
+ const finish = () => {
40
+ if (settled) return;
41
+ settled = true;
42
+ clearTimeout(timer);
43
+ try {
44
+ process.stdin.removeAllListeners('data');
45
+ process.stdin.removeAllListeners('end');
46
+ process.stdin.removeAllListeners('error');
47
+ process.stdin.pause();
48
+ } catch {
49
+ /* stdin may already be gone; nothing left to clean up */
50
+ }
51
+ resolve();
52
+ };
53
+ const timer = setTimeout(finish, timeoutMs);
54
+ if (timer.unref) timer.unref();
55
+ try {
56
+ process.stdin.on('data', () => {});
57
+ process.stdin.on('end', finish);
58
+ process.stdin.on('error', finish);
59
+ process.stdin.resume();
60
+ } catch {
61
+ finish();
62
+ }
63
+ });
64
+ }
65
+
66
+ drainStdin(STDIN_DRAIN_MS).then(() => {
67
+ const payload = JSON.stringify({
68
+ hookSpecificOutput: {
69
+ hookEventName: 'SubagentStart',
70
+ additionalContext
71
+ }
72
+ });
73
+ // Exit only after the write's callback fires, so a buffered write to a
74
+ // pipe is not truncated by an exit that races ahead of it.
75
+ process.stdout.write(payload, () => process.exit(0));
76
+ });
@@ -20,10 +20,10 @@ Robustness first, cost second. Split tiers because the split produces better wor
20
20
  2. **Needs live data?** trends, current docs, pricing, recent events → standard tier with tools; freshness comes from tools, not from a bigger model.
21
21
  3. **Reviewing without changing?** → standard tier, read-only, findings ranked by severity. Escalate to deep only for security-critical review.
22
22
  4. **Ambiguous, strategic, or expensive to get wrong?** "design my…", "figure out…", unknown cause → deep tier. Then hand the plan down.
23
- 5. **Everything else that changes files or executes a known plan** → you build it directly, at standard tier. The main build is never handed off whole; bounded sub-parts (a bulk pass, a wide search, a long audit loop) can go to cheaper tiers.
23
+ {{DECISION_RULE5_L1}}
24
24
 
25
25
  Modifiers:
26
- - **Plan big, execute small.** The expensive tier steers, the cheaper tier does the volume. Never make the fast tier design anything; never make the deep tier grind out bulk output.
26
+ - **Plan big, execute small.** The expensive tier steers, the cheaper tier does the volume. Never make the fast tier design anything; never make the deep tier grind out bulk output.{{INLINE_THRESHOLD_NOTE}}
27
27
  - **Never silently retry at the same tier after a failure.** Escalate one tier, or consult the deep tier once, and say which you did. If two consults do not unstick it, stop and tell the human.
28
28
  - **De-escalate.** If a request sounds deep but is a lookup or a small edit, route down. Default down, escalate on evidence.
29
29
 
@@ -36,7 +36,7 @@ Cap: two deep-tier consults per build. The full procedure is `protocols/build-pr
36
36
 
37
37
  ## Delegating inside one agent
38
38
 
39
- Subagents, a fresh chat, a second window: each one holds none of these rules. Every hand-off carries a `TASK_BUNDLE.md` brief: purpose, task class, granted scope, capabilities, denied actions, conventions it does not have, report contract, exit parameters. Absence is denial.
39
+ {{DELEGATE_RULES_NOTE}} Every hand-off carries a `TASK_BUNDLE.md` brief: purpose, task class, granted scope, capabilities, denied actions, conventions it does not have, report contract, exit parameters. Absence is denial.
40
40
 
41
41
  ## Numbers and logic go through a tool, never your head
42
42
 
@@ -53,3 +53,4 @@ Anything durable is searched for before it is written, its folder index is corre
53
53
  ## When you outgrow this
54
54
 
55
55
  You will know: you keep wanting a second model family to read your diff, a $0 lane for bulk, or a live-data lane your primary does not have. That is level 2. Re-run the installer with `--level 2`.
56
+ {{ROUTE_GATE_SECTION}}
@@ -1,6 +1,6 @@
1
1
  # Task Bundle: the brief every delegation carries
2
2
 
3
- A subagent, a second CLI, or a fresh chat window holds none of the rules your main session is holding. It cannot see your conventions, it cannot route, and it will read an unspecified edge as an open one.
3
+ A subagent, a second CLI, or a fresh chat window may hold none of the rules your main session is holding, and that is the default to assume. One exception: a Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already carries the standing rules, just not this task's scope. Either way, it cannot see this task's conventions and will read an unspecified edge as an open one.
4
4
 
5
5
  > A delegate gets an approved, bounded brief. Absence is not permission.
6
6
 
@@ -25,7 +25,7 @@ Copy this into the delegate's prompt. Delete nothing; write `none` where a field
25
25
  **Denied actions.** <explicit list: do not commit, push, deploy, delete, send, publish, close a ticket...>
26
26
  - Anything absent from Capabilities is denied. Absence is not permission.
27
27
 
28
- **Conventions you do not have.** <restate every house rule this task needs; the delegate holds none>
28
+ **Conventions you do not have.** <restate every house rule this task needs; even a delegate that loaded the standing rules still needs this task's scope, and a second CLI or a fresh chat window may hold none of it>
29
29
 
30
30
  **Report contract.** Return: <exactly what to hand back>. State plainly what you did NOT do
31
31
  and anything you could not verify. "Unverified" is an acceptable answer; a confident guess is not.
@@ -104,14 +104,14 @@ Use a different model family from the one that produced the finding where you ha
104
104
 
105
105
  | Role | Does | Does not |
106
106
  |---|---|---|
107
- | Builder / orchestrator | Routes, maps, writes, verifies, records. Stages 0, 1, 3, 6, 7 | Hand off the main build |
107
+ {{ROLES_BUILDER_ROW}}
108
108
  | Judgment tier | Stage 2 and the architectural arm of Stage 5. Argues with a finished map | Perform the retrieval |
109
109
  | Adversarial auditor | The security arm of Stage 5. Attacks the diff | Fix anything |
110
110
  | Mechanical gates | Stage 4 and any always-on guard | Be overridden without reading |
111
111
  | Cheap workers | Bounded sub-parts: bulk passes, wide searches, long loops | Own a stage |
112
112
  | Human | Stage 5b, and any irreversible or architectural call | Be the first line of review |
113
113
 
114
- **Why the builder does not hand off the main build:** a delegated agent does not inherit the session's standing rules and usually cannot delegate further. Any brief must restate every convention it needs (see `TASK_BUNDLE.md`), and that cost is itself a reason to build directly when the work fits.
114
+ {{BUILDER_HANDOFF_NOTE}}
115
115
 
116
116
  ## Checklist
117
117
 
@@ -15,19 +15,15 @@ Rule of thumb: never spend a frontier token on a task a cheap tier finishes corr
15
15
  0. **Is there a cheaper or better external lane for this?** Check `DELEGATION_MATRIX.md`. Your enabled lanes, every one called through `bin/cli-run.mjs`:
16
16
  {{LANE_STEP0}}
17
17
  1. **Bulk and mechanical?** → fast tier{{BULK_LANE}}. Many independent items each needing its own agent turn → a concurrent fan-out lane if you have one.
18
+ 1a. **Reading or digesting many files or notes, not writing?** → reader. Different from a bulk pass: reader reports, it does not classify, tag or transform.
18
19
  2. **Needs live data?** → {{LIVE_LANE}} standard tier with web tools.
19
20
  3. **Reviewing without changing?** → standard tier read-only. Security-critical → {{ATTACK_LANE}}.
20
21
  3a. **Holding findings from a review or a scanner?** → finding-verifier before any of them cause a repair. A finding is a claim, not a fact.
22
+ 3b. **Checking a tracker item or task against its stated done-signal?** → done-verifier. It probes the named artifact and returns MET, NOT_MET or UNVERIFIABLE; it never closes anything itself.
21
23
  4. **Ambiguous, strategic, expensive to get wrong?** → deep tier (deep-planner). Then hand the plan down.
22
- 5. **Everything else that changes files** → the orchestrator builds it directly. Bounded sub-parts go to cheaper tiers; the main build is never handed off whole.
24
+ {{DECISION_RULE5}}
23
25
 
24
- ## Who builds
25
-
26
- **The orchestrator owns the main build.** It is the only surface that holds these rules: a subagent or a second CLI starts with none of them and cannot route. Handing the main build to one hands it to something the router cannot reach.
27
-
28
- Delegate: background and long-running tasks, small tasks, scoping, verification, research, bounded sub-parts. Never delegate: the main build, or any step that must carry a house rule (secrets handling, the loud-negative verification, the durable record).
29
-
30
- Every delegation carries `TASK_BUNDLE.md`. Its brief must restate every convention the delegate needs.
26
+ {{WHO_BUILDS}}
31
27
 
32
28
  ## The Build Protocol, with lanes bound
33
29
 
@@ -56,7 +52,7 @@ One writer per run; every other lane proposes. Search before writing, index in t
56
52
 
57
53
  ## Modifier rules
58
54
 
59
- - **Plan big, execute small**, within a build: deep tier plans at Checkpoint 1, the orchestrator executes, bulk and wide searches go down.
55
+ {{PLAN_BIG_LINE}}{{INLINE_THRESHOLD_NOTE}}
60
56
  - **Escalation:** never silently retry at the same tier. Escalate one tier or consult deep once, and say which. Two consults that do not unstick it → stop and tell the human.
61
57
  - **De-escalation:** a request that sounds deep but is a lookup routes down.
62
58
  - **Long context:** mechanical digestion → fast tier in chunks; judgment over a long input → standard tier.
@@ -71,8 +67,11 @@ One writer per run; every other lane proposes. Search before writing, index in t
71
67
  |---|---|
72
68
  | "Design the architecture for X" | deep-planner |
73
69
  | "Review this service for bugs" | code-reviewer |
74
- | "Add an endpoint" | the orchestrator builds it |
70
+ {{ADD_ENDPOINT_ROW}}
75
71
  | "Why does this silently drop rows sometimes" | deep-planner (unknown cause), then build the fix directly |
76
72
  | "Summarize these 30 notes into one index" | bulk-worker |
73
+ | "Read every note in this folder and pull out every mention of X" | reader |
77
74
  | "The audit returned 6 findings" | finding-verifier first; repair only what comes back CONFIRMED |
75
+ | "Is issue #123 actually done" | done-verifier |
78
76
  {{LANE_EXAMPLES}}
77
+ {{ROUTE_GATE_SECTION}}
@@ -23,6 +23,8 @@ Tier sets the price per token. Token discipline sets how many tokens. **Effort s
23
23
  | builder | standard | high | a botched deploy is the costly failure |
24
24
  | live-researcher | standard | medium | tools do the retrieval |
25
25
  | bulk-worker | fast | low | the biggest cost win |
26
+ | done-verifier | fast | low | a done-signal check is a lookup, not a judgment call |
27
+ | reader | fast | low | digestion, not judgment |
26
28
 
27
29
  ## Three inputs, not one
28
30