model-orchestrator 0.1.13 → 0.1.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (34) hide show
  1. package/AGENTS.md +26 -0
  2. package/CHANGELOG.md +41 -1
  3. package/README.md +95 -6
  4. package/bin/cli-run.mjs +144 -19
  5. package/bin/cli.js +1 -1
  6. package/docs/audit-brief.md +21 -0
  7. package/docs/part-1-beginner.md +3 -1
  8. package/docs/part-2-intermediate.md +8 -2
  9. package/llms.txt +27 -0
  10. package/package.json +32 -4
  11. package/src/README.md +1 -1
  12. package/src/catalog.js +9 -1
  13. package/src/install.js +188 -5
  14. package/templates/README.md +1 -1
  15. package/templates/agents/README.md +3 -1
  16. package/templates/agents/agy/README.md +2 -2
  17. package/templates/agents/agy/done-verifier.md +35 -0
  18. package/templates/agents/agy/finding-verifier.md +29 -0
  19. package/templates/agents/agy/reader.md +22 -0
  20. package/templates/agents/claude-code/README.md +7 -4
  21. package/templates/agents/claude-code/builder.md +6 -1
  22. package/templates/agents/claude-code/done-verifier.md +44 -0
  23. package/templates/agents/claude-code/finding-verifier.md +43 -0
  24. package/templates/agents/claude-code/reader.md +26 -0
  25. package/templates/agents/snippets/claude-code.md +9 -3
  26. package/templates/agents/snippets/route-gate.mjs +151 -0
  27. package/templates/agents/snippets/settings.hooks.snippet.json +26 -0
  28. package/templates/agents/snippets/subagent-context.mjs +76 -0
  29. package/templates/beginner/ORCHESTRATOR.md +4 -3
  30. package/templates/common/TASK_BUNDLE.md +2 -2
  31. package/templates/common/protocols/build-protocol.md +13 -3
  32. package/templates/intermediate/CLI-RUN.md +40 -3
  33. package/templates/intermediate/ROUTING.md +15 -11
  34. package/templates/intermediate/TIERS.md +44 -1
@@ -0,0 +1,43 @@
1
+ ---
2
+ name: finding-verifier
3
+ description: Adversarial verification of review findings. Use after a review or audit returns findings and before any of them trigger a repair. Read-only. Tries to DISPROVE each finding and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Do not use to find new problems, and do not use to fix anything.
4
+ tools: Read, Glob, Grep, Bash
5
+ model: sonnet
6
+ effort: high
7
+ ---
8
+
9
+ You are the verification tier of the model router.
10
+
11
+ A finding is a claim, not a fact. Your job is to try to disprove each one before
12
+ it is allowed to cause a change. A false finding is expensive twice: it buys a
13
+ repair nobody needed, and it teaches everyone to skim the next report.
14
+
15
+ You are given findings from a review or an audit. For each one, independently:
16
+
17
+ 1. Read the cited file and line yourself. A citation that does not point at what
18
+ the finding describes is already a failure of the finding, not of the code.
19
+ 2. State the exact input, state or sequence that would make it happen.
20
+ 3. Look for what makes it impossible: a guard upstream, a type that cannot hold
21
+ that value, a caller that never passes it, a test that already covers it, a
22
+ framework guarantee.
23
+ 4. Where you can run something cheap and read-only that settles it, run it.
24
+
25
+ Return one verdict per finding, in the order you were given them:
26
+
27
+ - **CONFIRMED** you reproduced it, or traced a concrete path to it that nothing
28
+ prevents. Give the path in one or two sentences.
29
+ - **NOT_REPRODUCED** you found what stops it. Name that thing and where it is.
30
+ This is a success, not a failure to try.
31
+ - **INCONCLUSIVE** you could not settle it read-only. Say exactly what you would
32
+ need: a test run, a credential, a live environment, a decision from a human.
33
+ Never round this up to CONFIRMED to be safe, and never down to
34
+ NOT_REPRODUCED to be tidy.
35
+
36
+ Rules:
37
+ - Verify only the findings you were given. New problems you happen to notice go
38
+ in a separate list at the end, clearly marked as unverified observations.
39
+ - You are read-only. You never repair, and you never soften a finding's wording.
40
+ - Verifying nothing is a real answer. If every finding is NOT_REPRODUCED, say
41
+ that plainly; a verifier that always confirms something is a rubber stamp
42
+ facing the other way.
43
+ - Token discipline: read the cited code and its callers, not the repository.
@@ -0,0 +1,26 @@
1
+ ---
2
+ name: reader
3
+ description: Reads and digests many files or notes and returns exactly what the brief asks for (facts, quotes with path:line, an index, a digest). Read-only. Use for "read all X line by line", extracting facts or quotes across a folder, indexing or summarizing many notes, or pulling every mention of a topic. Different from bulk-worker, which classifies, tags and transforms items and writes output: reader only reads and reports.
4
+ tools: Read, Glob, Grep
5
+ model: haiku
6
+ effort: low
7
+ ---
8
+
9
+ You are the reading tier of the model router.
10
+
11
+ You read and digest many files or notes and hand back exactly what the brief
12
+ asked for: facts, quotes, an index, a digest. You do not classify, tag,
13
+ transform or rewrite; that is bulk-worker's job, not yours, and you never
14
+ write a file.
15
+
16
+ Rules:
17
+ - Read the brief first and answer only what it asks. "Every mention of X"
18
+ means grep for X and read the hits, not the whole corpus.
19
+ - Cite every fact or quote with its source: `path:line` for code and notes, a
20
+ URL and a retrieval note for anything fetched.
21
+ - An index or digest is a structured list, one row or bullet per source, not
22
+ prose that blends sources together.
23
+ - If a source is missing, unreadable, or empty, say so by name; do not
24
+ silently skip it.
25
+ - Token discipline: read only what the brief needs, never re-read a file,
26
+ summarize as you go rather than holding full text for later.
@@ -10,12 +10,16 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
10
10
  1. Bulk, mechanical, many similar items -> bulk-worker (fast tier).
11
11
  2. Needs live data -> live-researcher (standard tier + tools).
12
12
  3. Review without changing -> code-reviewer (standard, read-only).
13
+ 3a. Holding findings from a review or scanner -> finding-verifier before any repair. Only CONFIRMED findings earn a change.
14
+ 3b. Checking a tracker item against its stated done-signal -> done-verifier. It never closes anything itself.
13
15
  4. Ambiguous, architectural, or expensive to get wrong -> deep-planner (deep tier), then hand the plan down.
14
- 5. Everything else that changes files -> build it directly. The main build is never handed off whole; bounded sub-parts go to builder.
16
+ 5. Everything else that changes files -> builder executes by default. The orchestrator plans, briefs, verifies and talks to you; it stays inline only when (a) the brief would cost as much as the work, (b) the task needs this conversation's own context, or (c) it is your decision, or the final verification of delegated work. Never send rule-bound work to the built-in Explore or Plan agents: they skip CLAUDE.md. general-purpose should not take work a named agent already owns.
17
+
18
+ A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Measure yours once: spawn a subagent with a one-line task and read its token count. Work smaller than that stays inline.
15
19
 
16
20
  Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one adversarial pass, an explicit human yes before anything irreversible, then the loud negative.
17
21
 
18
- Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A subagent holds none of these rules; absence is denial.
22
+ Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A Claude Code subagent loads this CLAUDE.md hierarchy, so it holds the standing rules already, just not this task's scope; a second CLI or a fresh chat window may hold none of them. Absence is denial either way.
19
23
 
20
24
  Never silently retry a failed attempt at the same tier. Escalate once and say so.
21
25
 
@@ -24,4 +28,6 @@ Numbers, comparisons, complexity and equivalence claims go through codecalc (or
24
28
  Anything durable is searched for before it is written and its folder index is corrected in the same pass; one writer per run: `{{RULES_PATH}}/protocols/memory-and-record.md`.
25
29
  ```
26
30
 
27
- Subagents were written to `{{AGENTS_DIR}}` (the project root, which is where Claude Code reads project-level agents; `--project` changes it). Run `claude` from `{{PROJECT_DIR}}` and they are available as `deep-planner`, `builder`, `code-reviewer`, `live-researcher`, `bulk-worker`.
31
+ Subagents were written to `{{AGENTS_DIR}}` (the project root, which is where Claude Code reads project-level agents; `--project` changes it). Run `claude` from `{{PROJECT_DIR}}` and they are available as {{AGENTS_LIST_LINE}}.
32
+
33
+ Two hooks were written to `{{AGENTS_DIR}}/../hooks/` (`.claude/hooks/`): `route-gate.mjs` injects the routing table on every prompt, and `subagent-context.mjs` reminds a spawned subagent where the rules and the task-bundle format live. Merge `settings.hooks.snippet.json`, written next to this file, into `.claude/settings.json` to wire them in.
@@ -0,0 +1,151 @@
1
+ #!/usr/bin/env node
2
+ // route-gate.mjs: UserPromptSubmit hook for {{PRIMARY_NAME}}.
3
+ //
4
+ // Reads the route-gate table out of {{RULES_FILE_REL}} and injects it as
5
+ // additionalContext on every turn, so the routing table is read at runtime
6
+ // from the one place it is generated (the rules file), never a second
7
+ // hand-typed copy that can drift from it.
8
+ //
9
+ // Fail-open by design: a miss here is a stray context string, not a gate.
10
+ // This script always exits 0, never blocks on stdin past a short bound,
11
+ // reads at most 64 KB of the rules file through a fixed-size buffer (never
12
+ // a full read of an arbitrarily large or non-regular file), and never
13
+ // executes anything it reads. See docs/audit-brief.md for the threat model.
14
+ import { statSync, openSync, readSync, closeSync, realpathSync } from 'node:fs';
15
+ import { join, isAbsolute } from 'node:path';
16
+
17
+ // Rendered at install time from the level and directory the user chose.
18
+ // Never hardcoded: a level 1 install points this at ORCHESTRATOR.md, level
19
+ // 2+ at ROUTING.md, and a --dir outside the project resolves to an absolute
20
+ // path instead of a relative one.
21
+ const RULES_FILE_REL = {{RULES_FILE_REL_JSON}};
22
+
23
+ const MAX_READ = 64 * 1024; // bounded read: this is a rules file, not a log
24
+ const MAX_CONTEXT = 4000; // bounded injection: a table, not the whole file
25
+ const STDIN_DRAIN_MS = 250; // hard cap: never let an open, never-closed stdin pipe hold this hook open
26
+ const START = '<!-- route-gate:start -->';
27
+ const END = '<!-- route-gate:end -->';
28
+
29
+ function fallback(reason) {
30
+ return 'route-gate: ' + reason + '. Pick the lane before acting: read ' + RULES_FILE_REL + ' yourself.';
31
+ }
32
+
33
+ function resolveRulesPath() {
34
+ if (isAbsolute(RULES_FILE_REL)) return RULES_FILE_REL;
35
+ const projectDir = process.env.CLAUDE_PROJECT_DIR;
36
+ if (!projectDir) return null;
37
+ // Resolve through whatever part of the project dir already exists, so a
38
+ // symlinked project folder still resolves to the real path the rules file
39
+ // was written under.
40
+ let root = projectDir;
41
+ try {
42
+ root = realpathSync(projectDir);
43
+ } catch {
44
+ /* keep the unresolved value; the read below reports the real failure */
45
+ }
46
+ return join(root, RULES_FILE_REL);
47
+ }
48
+
49
+ // Bounded, regular-file-only read. statSync (not lstatSync) follows a
50
+ // symlink to its target and reports what the target actually is, so a
51
+ // symlinked rules file still reads; a FIFO, socket, device or directory at
52
+ // the resolved path is refused before any open/read call touches it. That
53
+ // check matters: opening a FIFO for reading blocks until a writer opens the
54
+ // other end, and a plain readFileSync on any of these can hang or, for a
55
+ // huge or sparse regular file, allocate far more than this hook needs. The
56
+ // fixed-size buffer plus a single bounded readSync call means the on-disk
57
+ // size of the file never determines how much this hook reads or how long it
58
+ // takes.
59
+ function readBounded(path) {
60
+ let st;
61
+ try {
62
+ st = statSync(path);
63
+ } catch (e) {
64
+ throw Object.assign(new Error('could not stat ' + path + ' (' + ((e && e.code) || e) + ')'), { code: e && e.code });
65
+ }
66
+ if (!st.isFile()) throw new Error(path + ' is not a regular file');
67
+ const buf = Buffer.alloc(MAX_READ);
68
+ let fd;
69
+ try {
70
+ fd = openSync(path, 'r');
71
+ const bytesRead = readSync(fd, buf, 0, MAX_READ, 0);
72
+ return buf.toString('utf8', 0, bytesRead);
73
+ } finally {
74
+ if (fd !== undefined) closeSync(fd);
75
+ }
76
+ }
77
+
78
+ function computeContext() {
79
+ const path = resolveRulesPath();
80
+ if (!path) return fallback('CLAUDE_PROJECT_DIR is not set, so ' + RULES_FILE_REL + ' could not be located');
81
+
82
+ let text;
83
+ try {
84
+ text = readBounded(path);
85
+ } catch (e) {
86
+ return fallback((e && e.message) || String(e));
87
+ }
88
+
89
+ const s = text.indexOf(START);
90
+ const e = s === -1 ? -1 : text.indexOf(END, s);
91
+ if (s === -1 || e === -1) return fallback(path + ' has no route-gate block');
92
+
93
+ return text.slice(s, e + END.length).slice(0, MAX_CONTEXT);
94
+ }
95
+
96
+ // Drain stdin without ever blocking on it. A bare `readFileSync(0)` waits
97
+ // for stdin to reach EOF, so a caller that pipes into this hook and never
98
+ // closes its end of the pipe (or a bare TTY with no redirection at all)
99
+ // left the process running indefinitely. This races the real 'end' event
100
+ // against a hard timeout instead: whichever settles first wins, and the
101
+ // timer is unref'd so it can never itself be the reason the process stays
102
+ // alive past a normal exit.
103
+ function drainStdin(timeoutMs) {
104
+ return new Promise((resolve) => {
105
+ let settled = false;
106
+ const finish = () => {
107
+ if (settled) return;
108
+ settled = true;
109
+ clearTimeout(timer);
110
+ try {
111
+ process.stdin.removeAllListeners('data');
112
+ process.stdin.removeAllListeners('end');
113
+ process.stdin.removeAllListeners('error');
114
+ process.stdin.pause();
115
+ } catch {
116
+ /* stdin may already be gone; nothing left to clean up */
117
+ }
118
+ resolve();
119
+ };
120
+ const timer = setTimeout(finish, timeoutMs);
121
+ if (timer.unref) timer.unref();
122
+ try {
123
+ process.stdin.on('data', () => {});
124
+ process.stdin.on('end', finish);
125
+ process.stdin.on('error', finish);
126
+ process.stdin.resume();
127
+ } catch {
128
+ finish();
129
+ }
130
+ });
131
+ }
132
+
133
+ let additionalContext;
134
+ try {
135
+ additionalContext = computeContext();
136
+ } catch (err) {
137
+ additionalContext = fallback('route-gate.mjs failed unexpectedly (' + ((err && err.message) || err) + ')');
138
+ }
139
+
140
+ drainStdin(STDIN_DRAIN_MS).then(() => {
141
+ const payload = JSON.stringify({
142
+ hookSpecificOutput: {
143
+ hookEventName: 'UserPromptSubmit',
144
+ additionalContext
145
+ }
146
+ });
147
+ // Exit only after the write's callback fires, so a buffered write to a
148
+ // pipe (the common case on Windows, and possible anywhere output exceeds
149
+ // one write's worth) is not truncated by an exit that races ahead of it.
150
+ process.stdout.write(payload, () => process.exit(0));
151
+ });
@@ -0,0 +1,26 @@
1
+ {
2
+ "hooks": {
3
+ "UserPromptSubmit": [
4
+ {
5
+ "hooks": [
6
+ {
7
+ "type": "command",
8
+ "command": "node",
9
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-gate.mjs"]
10
+ }
11
+ ]
12
+ }
13
+ ],
14
+ "SubagentStart": [
15
+ {
16
+ "hooks": [
17
+ {
18
+ "type": "command",
19
+ "command": "node",
20
+ "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/subagent-context.mjs"]
21
+ }
22
+ ]
23
+ }
24
+ ]
25
+ }
26
+ }
@@ -0,0 +1,76 @@
1
+ #!/usr/bin/env node
2
+ // subagent-context.mjs: SubagentStart hook for {{PRIMARY_NAME}}.
3
+ //
4
+ // A Claude Code subagent loads this project's CLAUDE.md hierarchy at start
5
+ // (code.claude.com/docs/en/sub-agents), so it already has the standing
6
+ // rules. What it does not have is this task's scope, and it can be tempted
7
+ // to route further work itself or to mark its own output verified. This
8
+ // hook injects a short, static reminder of where the rest lives and what
9
+ // "delegate" means.
10
+ //
11
+ // Fail-open by design: a miss here is a stray context string, not a gate.
12
+ // This script always exits 0, never executes anything it reads, and never
13
+ // blocks on stdin past a short bound (see drainStdin below).
14
+ import { isAbsolute } from 'node:path';
15
+
16
+ // Rendered at install time so a --dir outside the project still names an
17
+ // honest path rather than a hardcoded one.
18
+ const RULES_FILE_REL = {{RULES_FILE_REL_JSON}};
19
+ const TASK_BUNDLE_REL = {{TASK_BUNDLE_REL_JSON}};
20
+ const STDIN_DRAIN_MS = 250; // hard cap: never let an open, never-closed stdin pipe hold this hook open
21
+
22
+ const additionalContext = [
23
+ 'SUBAGENT CONTEXT (model-orchestrator).',
24
+ 'Routing rules: ' + RULES_FILE_REL + (isAbsolute(RULES_FILE_REL) ? '.' : ' (relative to the project root).'),
25
+ 'Task bundle format: ' + TASK_BUNDLE_REL + '.',
26
+ 'Report contract: say what you did, what you did NOT do, and what you could not verify. "Unverified" is acceptable; a confident guess is not. Stop at the bound your brief set, and never claim work you cannot show.',
27
+ 'You are a delegate: do not route further work to another subagent yourself, and do not mark your own output as the final verification of it.'
28
+ ].join(' ');
29
+
30
+ // Drain stdin without ever blocking on it. A bare `readFileSync(0)` waits
31
+ // for stdin to reach EOF, so a caller that pipes into this hook and never
32
+ // closes its end of the pipe left the process running indefinitely. This
33
+ // races the real 'end' event against a hard timeout instead: whichever
34
+ // settles first wins, and the timer is unref'd so it can never itself be
35
+ // the reason the process stays alive past a normal exit.
36
+ function drainStdin(timeoutMs) {
37
+ return new Promise((resolve) => {
38
+ let settled = false;
39
+ const finish = () => {
40
+ if (settled) return;
41
+ settled = true;
42
+ clearTimeout(timer);
43
+ try {
44
+ process.stdin.removeAllListeners('data');
45
+ process.stdin.removeAllListeners('end');
46
+ process.stdin.removeAllListeners('error');
47
+ process.stdin.pause();
48
+ } catch {
49
+ /* stdin may already be gone; nothing left to clean up */
50
+ }
51
+ resolve();
52
+ };
53
+ const timer = setTimeout(finish, timeoutMs);
54
+ if (timer.unref) timer.unref();
55
+ try {
56
+ process.stdin.on('data', () => {});
57
+ process.stdin.on('end', finish);
58
+ process.stdin.on('error', finish);
59
+ process.stdin.resume();
60
+ } catch {
61
+ finish();
62
+ }
63
+ });
64
+ }
65
+
66
+ drainStdin(STDIN_DRAIN_MS).then(() => {
67
+ const payload = JSON.stringify({
68
+ hookSpecificOutput: {
69
+ hookEventName: 'SubagentStart',
70
+ additionalContext
71
+ }
72
+ });
73
+ // Exit only after the write's callback fires, so a buffered write to a
74
+ // pipe is not truncated by an exit that races ahead of it.
75
+ process.stdout.write(payload, () => process.exit(0));
76
+ });
@@ -20,10 +20,10 @@ Robustness first, cost second. Split tiers because the split produces better wor
20
20
  2. **Needs live data?** trends, current docs, pricing, recent events → standard tier with tools; freshness comes from tools, not from a bigger model.
21
21
  3. **Reviewing without changing?** → standard tier, read-only, findings ranked by severity. Escalate to deep only for security-critical review.
22
22
  4. **Ambiguous, strategic, or expensive to get wrong?** "design my…", "figure out…", unknown cause → deep tier. Then hand the plan down.
23
- 5. **Everything else that changes files or executes a known plan** → you build it directly, at standard tier. The main build is never handed off whole; bounded sub-parts (a bulk pass, a wide search, a long audit loop) can go to cheaper tiers.
23
+ {{DECISION_RULE5_L1}}
24
24
 
25
25
  Modifiers:
26
- - **Plan big, execute small.** The expensive tier steers, the cheaper tier does the volume. Never make the fast tier design anything; never make the deep tier grind out bulk output.
26
+ - **Plan big, execute small.** The expensive tier steers, the cheaper tier does the volume. Never make the fast tier design anything; never make the deep tier grind out bulk output.{{INLINE_THRESHOLD_NOTE}}
27
27
  - **Never silently retry at the same tier after a failure.** Escalate one tier, or consult the deep tier once, and say which you did. If two consults do not unstick it, stop and tell the human.
28
28
  - **De-escalate.** If a request sounds deep but is a lookup or a small edit, route down. Default down, escalate on evidence.
29
29
 
@@ -36,7 +36,7 @@ Cap: two deep-tier consults per build. The full procedure is `protocols/build-pr
36
36
 
37
37
  ## Delegating inside one agent
38
38
 
39
- Subagents, a fresh chat, a second window: each one holds none of these rules. Every hand-off carries a `TASK_BUNDLE.md` brief: purpose, task class, granted scope, capabilities, denied actions, conventions it does not have, report contract, exit parameters. Absence is denial.
39
+ {{DELEGATE_RULES_NOTE}} Every hand-off carries a `TASK_BUNDLE.md` brief: purpose, task class, granted scope, capabilities, denied actions, conventions it does not have, report contract, exit parameters. Absence is denial.
40
40
 
41
41
  ## Numbers and logic go through a tool, never your head
42
42
 
@@ -53,3 +53,4 @@ Anything durable is searched for before it is written, its folder index is corre
53
53
  ## When you outgrow this
54
54
 
55
55
  You will know: you keep wanting a second model family to read your diff, a $0 lane for bulk, or a live-data lane your primary does not have. That is level 2. Re-run the installer with `--level 2`.
56
+ {{ROUTE_GATE_SECTION}}
@@ -1,6 +1,6 @@
1
1
  # Task Bundle: the brief every delegation carries
2
2
 
3
- A subagent, a second CLI, or a fresh chat window holds none of the rules your main session is holding. It cannot see your conventions, it cannot route, and it will read an unspecified edge as an open one.
3
+ A subagent, a second CLI, or a fresh chat window may hold none of the rules your main session is holding, and that is the default to assume. One exception: a Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already carries the standing rules, just not this task's scope. Either way, it cannot see this task's conventions and will read an unspecified edge as an open one.
4
4
 
5
5
  > A delegate gets an approved, bounded brief. Absence is not permission.
6
6
 
@@ -25,7 +25,7 @@ Copy this into the delegate's prompt. Delete nothing; write `none` where a field
25
25
  **Denied actions.** <explicit list: do not commit, push, deploy, delete, send, publish, close a ticket...>
26
26
  - Anything absent from Capabilities is denied. Absence is not permission.
27
27
 
28
- **Conventions you do not have.** <restate every house rule this task needs; the delegate holds none>
28
+ **Conventions you do not have.** <restate every house rule this task needs; even a delegate that loaded the standing rules still needs this task's scope, and a second CLI or a fresh chat window may hold none of it>
29
29
 
30
30
  **Report contract.** Return: <exactly what to hand back>. State plainly what you did NOT do
31
31
  and anything you could not verify. "Unverified" is an acceptable answer; a confident guess is not.
@@ -66,7 +66,17 @@ Secret detection, static analysis and dependency scanning, filtered to lines thi
66
66
 
67
67
  Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to an adversarial auditor, ideally a **different model family**. Architecture-shaped diffs go to the judgment tier reviewing build against plan. Never both on one diff.
68
68
 
69
- **Gate:** every finding **reproduced** before it reaches a human. Unreproduced items are dropped, not narrated. Hard cap one re-audit. `CLEAN` is a valid success state; an auditor that is not allowed to say so manufactures something.
69
+ Findings do not go straight to a repair. Hand them to the **finding-verifier**, whose job is to DISPROVE each one: read the cited line, state what would trigger it, then hunt for the guard, caller or test that makes it impossible. It returns one of three verdicts per finding, and rounding between them is the failure mode to watch for.
70
+
71
+ | Verdict | Meaning | What happens next |
72
+ |---|---|---|
73
+ | CONFIRMED | reproduced, or a concrete path nothing blocks | it earns a repair |
74
+ | NOT_REPRODUCED | something prevents it, named and located | dropped, and not narrated |
75
+ | INCONCLUSIVE | not settleable read-only | say what it would take; never round it to either side |
76
+
77
+ Use a different model family from the one that produced the finding where you have one: a family asked to check its own claim tends to agree with itself.
78
+
79
+ **Gate:** every finding **verified** before it reaches a human, and only CONFIRMED findings trigger a change. Hard cap one re-audit. `CLEAN` is a valid success state; an auditor that is not allowed to say so manufactures something, and so does a verifier that is expected to confirm.
70
80
 
71
81
  ### Stage 5b · Ship gate
72
82
  1. What is the rollback target? Record it before shipping.
@@ -94,14 +104,14 @@ Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk muta
94
104
 
95
105
  | Role | Does | Does not |
96
106
  |---|---|---|
97
- | Builder / orchestrator | Routes, maps, writes, verifies, records. Stages 0, 1, 3, 6, 7 | Hand off the main build |
107
+ {{ROLES_BUILDER_ROW}}
98
108
  | Judgment tier | Stage 2 and the architectural arm of Stage 5. Argues with a finished map | Perform the retrieval |
99
109
  | Adversarial auditor | The security arm of Stage 5. Attacks the diff | Fix anything |
100
110
  | Mechanical gates | Stage 4 and any always-on guard | Be overridden without reading |
101
111
  | Cheap workers | Bounded sub-parts: bulk passes, wide searches, long loops | Own a stage |
102
112
  | Human | Stage 5b, and any irreversible or architectural call | Be the first line of review |
103
113
 
104
- **Why the builder does not hand off the main build:** a delegated agent does not inherit the session's standing rules and usually cannot delegate further. Any brief must restate every convention it needs (see `TASK_BUNDLE.md`), and that cost is itself a reason to build directly when the work fits.
114
+ {{BUILDER_HANDOFF_NOTE}}
105
115
 
106
116
  ## Checklist
107
117
 
@@ -18,7 +18,8 @@ Enabled lanes (edit `bin/lanes.json`): {{CLI_RUN_LANES}}
18
18
  ```bash
19
19
  node bin/cli-run.mjs <grok|codex|agy|hermes|qwen> "<prompt>" [--brief FILE] [--timeout SECS] [--quiet]
20
20
  node bin/cli-run.mjs codex --audit "<prompt>" # read-only sandbox, the audit shape
21
- node bin/cli-run.mjs qwen [--model ID] [--safe-mode] "<prompt>"
21
+ node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
22
+ node bin/cli-run.mjs qwen [--safe-mode] "<prompt>" # qwen-only flag
22
23
  ```
23
24
 
24
25
  Put it on your PATH if you like: `ln -s "$PWD/bin/cli-run.mjs" ~/.local/bin/cli-run`.
@@ -65,13 +66,49 @@ qwen is the lane whose own success flags lie: an upstream 400 comes back as exit
65
66
  | 2 | usage error in cli-run itself |
66
67
  | N | the lane exited N != 0: passed through unchanged, verdict `exit_nonzero`, even when parseable text came back. The bounded head of the lane's stderr is shown on your terminal so an auth failure reads as one |
67
68
 
69
+ ## The route: which model, and how hard it thinks
70
+
71
+ A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe an adversarial pass, and nothing anywhere says so.
72
+
73
+ Pin it per call, or per lane:
74
+
75
+ ```bash
76
+ node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high # this call only
77
+ node bin/cli-run.mjs --doctor # prints what each lane is pinned to
78
+ ```
79
+
80
+ ```json
81
+ {
82
+ "enabled": ["codex", "grok"],
83
+ "defaults": { "codex": { "model": "gpt-6-astra", "effort": "high" } }
84
+ }
85
+ ```
86
+
87
+ A flag beats `defaults`; `defaults` beats nothing. Each vendor spells these differently and `cli-run` translates:
88
+
89
+ | Lane | Model | Reasoning effort |
90
+ |---|---|---|
91
+ | grok | `-m` | `--reasoning-effort` |
92
+ | codex | `-m` | `-c model_reasoning_effort="LEVEL"` |
93
+ | agy | `--model` | `--effort` (low, medium, high) |
94
+ | hermes | `-m` | `--reasoning` (none, minimal, ...) |
95
+ | qwen | `-m` | none: this lane has no reasoning flag |
96
+
97
+ Three rules that keep this honest:
98
+
99
+ - **A level `cli-run` does not recognise is not rejected here.** Levels are the vendor's, they change, and guessing the valid set would date this tool. An unknown level is refused by the lane and surfaces as that lane's own exit code and stderr.
100
+ - **`--effort` on qwen is a usage error, not a silent drop.** A flag that vanishes leaves you believing a route that never ran.
101
+ - **Values are charset-bounded** (letters, digits, and `. _ : @ / + -`, no leading dash, 64 characters). A model id becomes an argv element and, on codex, part of a TOML value; bounding it is what stops either from being escaped.
102
+
68
103
  ## Permissions are a separate layer
69
104
 
70
105
  `cli-run` never injects permission flags. Each CLI carries its own config, so every caller gets the same behaviour. Use each vendor's deny-list as the base layer; allow-lists only hold if every binary is enumerable in advance.
71
106
 
72
107
  ## Log
73
108
 
74
- `~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, rc, the lane's own exit code, signal, seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument.
109
+ `~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, rc, the lane's own exit code, signal, seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, the route (`model_requested`, `effort_requested`, and `model_source` / `effort_source`, each one of `flag`, `lanes.json` or `lane_default`), and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument, and so does "we route audits at high effort".
110
+
111
+ The log records what was **requested**, on every record including a run refused before the lane started. It does not record an actual. Reporting is inconsistent: grok returns a `modelUsage` block naming a model, the other four lanes return nothing of the kind, so an `actual` field would be populated for one lane and empty for four. It would also be a provider-supplied string, and this log holds fixed codes and bounded caller-supplied values only. `model_source: "lane_default"` is the honest way to say this run inherited something invisible from here.
75
112
 
76
113
  ## The prompt travels in argv
77
114
 
@@ -79,7 +116,7 @@ That is each vendor's documented headless shape (`-p`, `exec`). Two consequences
79
116
 
80
117
  ## lanes.json fails closed
81
118
 
82
- Absent: every lane enabled. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled.
119
+ Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all fail the whole file closed rather than being skipped quietly.
83
120
 
84
121
  ## A killed lane is not a deliverable
85
122
 
@@ -15,18 +15,15 @@ Rule of thumb: never spend a frontier token on a task a cheap tier finishes corr
15
15
  0. **Is there a cheaper or better external lane for this?** Check `DELEGATION_MATRIX.md`. Your enabled lanes, every one called through `bin/cli-run.mjs`:
16
16
  {{LANE_STEP0}}
17
17
  1. **Bulk and mechanical?** → fast tier{{BULK_LANE}}. Many independent items each needing its own agent turn → a concurrent fan-out lane if you have one.
18
+ 1a. **Reading or digesting many files or notes, not writing?** → reader. Different from a bulk pass: reader reports, it does not classify, tag or transform.
18
19
  2. **Needs live data?** → {{LIVE_LANE}} standard tier with web tools.
19
20
  3. **Reviewing without changing?** → standard tier read-only. Security-critical → {{ATTACK_LANE}}.
21
+ 3a. **Holding findings from a review or a scanner?** → finding-verifier before any of them cause a repair. A finding is a claim, not a fact.
22
+ 3b. **Checking a tracker item or task against its stated done-signal?** → done-verifier. It probes the named artifact and returns MET, NOT_MET or UNVERIFIABLE; it never closes anything itself.
20
23
  4. **Ambiguous, strategic, expensive to get wrong?** → deep tier (deep-planner). Then hand the plan down.
21
- 5. **Everything else that changes files** → the orchestrator builds it directly. Bounded sub-parts go to cheaper tiers; the main build is never handed off whole.
24
+ {{DECISION_RULE5}}
22
25
 
23
- ## Who builds
24
-
25
- **The orchestrator owns the main build.** It is the only surface that holds these rules: a subagent or a second CLI starts with none of them and cannot route. Handing the main build to one hands it to something the router cannot reach.
26
-
27
- Delegate: background and long-running tasks, small tasks, scoping, verification, research, bounded sub-parts. Never delegate: the main build, or any step that must carry a house rule (secrets handling, the loud-negative verification, the durable record).
28
-
29
- Every delegation carries `TASK_BUNDLE.md`. Its brief must restate every convention the delegate needs.
26
+ {{WHO_BUILDS}}
30
27
 
31
28
  ## The Build Protocol, with lanes bound
32
29
 
@@ -38,6 +35,7 @@ Every delegation carries `TASK_BUNDLE.md`. Its brief must restate every conventi
38
35
  | 3 Build | the orchestrator, against the installed dependency's source |
39
36
  | 4 Scan | secret + static + dependency scanners, diff-scoped, fail closed |
40
37
  | 5 Attack | security-shaped diff → {{ATTACK_LANE}}. Architecture-shaped → deep tier, build against plan. Never both |
38
+ | 5a Verify findings | finding-verifier, a different model family where you have one: CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a repair |
41
39
  | 5b Ship | rollback id recorded, explicit human yes |
42
40
  | 6 Verify | real test, negative test seen red, old identifier re-grepped to zero |
43
41
  | 7 Record | one end-to-end doc, tracker Done with evidence, plan doc deleted |
@@ -54,12 +52,14 @@ One writer per run; every other lane proposes. Search before writing, index in t
54
52
 
55
53
  ## Modifier rules
56
54
 
57
- - **Plan big, execute small**, within a build: deep tier plans at Checkpoint 1, the orchestrator executes, bulk and wide searches go down.
55
+ {{PLAN_BIG_LINE}}{{INLINE_THRESHOLD_NOTE}}
58
56
  - **Escalation:** never silently retry at the same tier. Escalate one tier or consult deep once, and say which. Two consults that do not unstick it → stop and tell the human.
59
57
  - **De-escalation:** a request that sounds deep but is a lookup routes down.
60
58
  - **Long context:** mechanical digestion → fast tier in chunks; judgment over a long input → standard tier.
61
59
  - **Token discipline on every delegation:** pass only the context the delegate needs, never the conversation.
62
- - **Effort per agent:** deep xhigh, review and build high, live research medium, bulk low.
60
+ - **Effort per agent:** deep xhigh, review, verification and build high, live research medium, bulk low.
61
+ - **Three inputs, not one:** role picks the agent, complexity moves the effort, risk moves the tier and who reads it. A one-line auth change is simple and high-risk at once, and the risk decides. See `TIERS.md`.
62
+ - **Pin the route when it matters:** a lane with no `--model`/`--effort` and no `defaults` entry in `bin/lanes.json` runs on its own config, which may be nothing like what this file describes. `cli-run --doctor` prints what each lane is pinned to, and every run logs the value requested and where it came from.
63
63
 
64
64
  ## Example routings
65
65
 
@@ -67,7 +67,11 @@ One writer per run; every other lane proposes. Search before writing, index in t
67
67
  |---|---|
68
68
  | "Design the architecture for X" | deep-planner |
69
69
  | "Review this service for bugs" | code-reviewer |
70
- | "Add an endpoint" | the orchestrator builds it |
70
+ {{ADD_ENDPOINT_ROW}}
71
71
  | "Why does this silently drop rows sometimes" | deep-planner (unknown cause), then build the fix directly |
72
72
  | "Summarize these 30 notes into one index" | bulk-worker |
73
+ | "Read every note in this folder and pull out every mention of X" | reader |
74
+ | "The audit returned 6 findings" | finding-verifier first; repair only what comes back CONFIRMED |
75
+ | "Is issue #123 actually done" | done-verifier |
73
76
  {{LANE_EXAMPLES}}
77
+ {{ROUTE_GATE_SECTION}}
@@ -19,11 +19,54 @@ Tier sets the price per token. Token discipline sets how many tokens. **Effort s
19
19
  |---|---|---|---|
20
20
  | deep-planner | deep | xhigh | judges every build twice; expensive to get wrong |
21
21
  | code-reviewer | standard | high | every endpoint is internet-facing |
22
+ | finding-verifier | standard | high | judging a claim is harder than producing it |
22
23
  | builder | standard | high | a botched deploy is the costly failure |
23
24
  | live-researcher | standard | medium | tools do the retrieval |
24
25
  | bulk-worker | fast | low | the biggest cost win |
26
+ | done-verifier | fast | low | a done-signal check is a lookup, not a judgment call |
27
+ | reader | fast | low | digestion, not judgment |
25
28
 
26
- Dials: drop builder to medium when the plan is airtight; raise code-reviewer to xhigh for a security-critical audit.
29
+ ## Three inputs, not one
30
+
31
+ Role alone does not decide a route. Two more inputs move it, and they move it in
32
+ opposite directions, so state them separately instead of folding them into the
33
+ role.
34
+
35
+ **Complexity moves the effort.** The same role does not need the same reasoning
36
+ on every task.
37
+
38
+ | Complexity | What it looks like | What moves |
39
+ |---|---|---|
40
+ | simple | one file, one obvious edit, no unknowns | drop one effort level |
41
+ | standard | the default | the table above |
42
+ | complex | several surfaces, or an unknown cause | keep effort, add the deep-tier checkpoint |
43
+ | critical | irreversible, or it rewrites a standing rule | the escalation rule below applies |
44
+
45
+ The dial that pays for itself: **a worker executing a finished plan needs less
46
+ reasoning than the reviewer judging its output.** When the plan is airtight the
47
+ spec is carrying the thinking, so builder drops to medium. When the plan is
48
+ vague, fix the plan; do not buy reasoning to paper over it.
49
+
50
+ **Risk moves the tier and the reader, never just the effort.** These four are
51
+ the ones worth naming, because their failures are not recoverable by editing the
52
+ code afterwards.
53
+
54
+ | Risk | Present when the change touches | What it buys |
55
+ |---|---|---|
56
+ | security | auth, tokens, sessions, routes, untrusted input | the attack pass, ideally a different model family |
57
+ | privacy | personal data, anything leaving the machine | the local lane, and a named check on what is sent |
58
+ | data loss | deletion, bulk mutation, migrations, overwrites | a reviewed rollback path before the change is written |
59
+ | irreversible | publishing, sending, rotating, anything with an audience | a human yes at Stage 5b, never an agent's |
60
+
61
+ A risk raises code-reviewer to xhigh, and a security-shaped diff goes to the
62
+ attack lane rather than to a second read by the same family. Risk is not a
63
+ synonym for difficulty: a one-line change to an auth check is simple and
64
+ high-risk at the same time, and it is the risk that decides the route.
65
+
66
+ **Reserve the top of the ladder for evidence.** xhigh and the escalation tier are
67
+ bought with a named reason: a reproduced failure, a checkpoint that came back
68
+ unresolved, an irreversible change. A task that merely feels hard is a deep-tier
69
+ task, not an escalation.
27
70
 
28
71
  ## Why split tiers: robustness first, cost second
29
72