model-orchestrator 0.1.14 → 0.1.15
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +26 -0
- package/CHANGELOG.md +24 -2
- package/README.md +38 -5
- package/bin/cli.js +1 -1
- package/docs/audit-brief.md +21 -0
- package/docs/part-1-beginner.md +1 -1
- package/docs/part-2-intermediate.md +1 -1
- package/llms.txt +27 -0
- package/package.json +28 -4
- package/src/README.md +1 -1
- package/src/catalog.js +9 -1
- package/src/install.js +183 -3
- package/templates/README.md +1 -1
- package/templates/agents/README.md +3 -1
- package/templates/agents/agy/README.md +2 -2
- package/templates/agents/agy/done-verifier.md +35 -0
- package/templates/agents/agy/reader.md +22 -0
- package/templates/agents/claude-code/README.md +6 -4
- package/templates/agents/claude-code/builder.md +6 -1
- package/templates/agents/claude-code/done-verifier.md +44 -0
- package/templates/agents/claude-code/reader.md +26 -0
- package/templates/agents/snippets/claude-code.md +8 -3
- package/templates/agents/snippets/route-gate.mjs +151 -0
- package/templates/agents/snippets/settings.hooks.snippet.json +26 -0
- package/templates/agents/snippets/subagent-context.mjs +76 -0
- package/templates/beginner/ORCHESTRATOR.md +4 -3
- package/templates/common/TASK_BUNDLE.md +2 -2
- package/templates/common/protocols/build-protocol.md +2 -2
- package/templates/intermediate/ROUTING.md +9 -10
- package/templates/intermediate/TIERS.md +2 -0
|
@@ -10,4 +10,6 @@ Loading surfaces for the primary agent. The installer writes exactly one of thes
|
|
|
10
10
|
| `grok`, `hermes` | nothing agent-specific | rules travel with the prompt or the task bundle |
|
|
11
11
|
| a chat app | `PASTE-INTO-YOUR-AGENT.md` | no files to load; paste into custom instructions |
|
|
12
12
|
|
|
13
|
-
`snippets/` are rendered with the chosen agent's name and rules file. Nothing here is appended to a file the user already has.
|
|
13
|
+
`snippets/` are rendered with the chosen agent's name and rules file. Nothing here is appended to a file the user already has. `snippets/route-gate.mjs`, `snippets/subagent-context.mjs`, and `snippets/settings.hooks.snippet.json` are claude-code only: two hooks and the settings block that wires them, installed to `.claude/hooks/` and next to `CLAUDE.snippet.md`.
|
|
14
|
+
|
|
15
|
+
`claude-code/` and `agy/` both ship the same agent set: one per tier, plus `finding-verifier`, `done-verifier` and `reader`. Add an agent to one folder and its README, and the other.
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
# .agents/agents/
|
|
2
2
|
|
|
3
|
-
Antigravity CLI custom agents, one per tier plus
|
|
3
|
+
Antigravity CLI custom agents, one per tier plus three checks (`finding-verifier`, `done-verifier`, `reader`), in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
|
|
4
4
|
|
|
5
|
-
`commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; `auto` keeps deletes and other high-risk commands gated) and `off` for the read-only agents
|
|
5
|
+
`commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; `auto` keeps deletes and other high-risk commands gated) and `off` for the read-only agents: `code-reviewer`, `finding-verifier`, `live-researcher`, `done-verifier`, `reader`. `model` is a tier: `pro` for deep-planner, `flash` for the rest. `done-verifier` and `reader` never write and, with `commandExecutionPolicy: off`, cannot execute any command at all here, mutating or not: unlike its claude-code counterpart, which does carry an unrestricted `Bash` and stays read-only by its prompt rather than by the tool grant, agy's `done-verifier` is mechanically blocked from shelling out and probes artifacts through whatever read or fetch capability it has instead. Neither is `bulk-worker`, which classifies, tags and transforms items and does write.
|
|
@@ -0,0 +1,35 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: done-verifier
|
|
3
|
+
description: Checks tracker items or tasks against their stated done-signal by probing the named artifact (a file, a commit, a URL, a log line, a count); no file-editing tools, no command execution (commandExecutionPolicy off); returns MET, NOT_MET or UNVERIFIABLE per item; never closes or edits anything.
|
|
4
|
+
model: flash
|
|
5
|
+
subagent: true
|
|
6
|
+
mainAgent: true
|
|
7
|
+
commandExecutionPolicy: off
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# done-verifier
|
|
11
|
+
|
|
12
|
+
Checks tracker items or tasks against their stated done-signal by probing the
|
|
13
|
+
named artifact. No file-editing tools, and no command execution: this agent's
|
|
14
|
+
`commandExecutionPolicy` is `off`, so unlike its claude-code counterpart it
|
|
15
|
+
cannot shell out at all, not even to a read-only command; probe with whatever
|
|
16
|
+
read or fetch capability you have instead.
|
|
17
|
+
|
|
18
|
+
For each item: read the stated done-signal, probe the exact artifact it
|
|
19
|
+
names, compare what you found against the claim.
|
|
20
|
+
|
|
21
|
+
Return one verdict per item:
|
|
22
|
+
- MET: the artifact matches the claim. Name what you checked.
|
|
23
|
+
- NOT_MET: the artifact is missing or contradicts the claim. Name what you
|
|
24
|
+
found instead.
|
|
25
|
+
- UNVERIFIABLE: you cannot probe it from here, no done-signal was stated, or
|
|
26
|
+
the check would need a command you are not able to run. Say what is
|
|
27
|
+
missing.
|
|
28
|
+
|
|
29
|
+
Rules:
|
|
30
|
+
- Stay inside the task bundle you were given. Anything not granted is denied.
|
|
31
|
+
- Never close, edit or comment on a tracker item; return verdicts only.
|
|
32
|
+
- If the only way to check something would mutate it, or would need command
|
|
33
|
+
execution you do not have, the item is UNVERIFIABLE, not MET.
|
|
34
|
+
- Token discipline: read only the cited artifact, hand back verdicts not
|
|
35
|
+
narration.
|
|
@@ -0,0 +1,22 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: reader
|
|
3
|
+
description: Reads and digests many files or notes and returns facts, quotes with source, an index or a digest. Read-only. Different from bulk-worker, which classifies, tags and transforms items: reader only reads and reports.
|
|
4
|
+
model: flash
|
|
5
|
+
subagent: true
|
|
6
|
+
mainAgent: true
|
|
7
|
+
commandExecutionPolicy: off
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# reader
|
|
11
|
+
|
|
12
|
+
Reads and digests many files or notes and hands back exactly what the brief
|
|
13
|
+
asks for: facts, quotes, an index, a digest. Does not classify, tag,
|
|
14
|
+
transform or rewrite; that is bulk-worker's job, and reader never writes a
|
|
15
|
+
file.
|
|
16
|
+
|
|
17
|
+
Rules:
|
|
18
|
+
- Stay inside the task bundle you were given. Anything not granted is denied.
|
|
19
|
+
- Cite every fact or quote with its source (path or URL).
|
|
20
|
+
- Report what you did, what you did not do, and what you could not verify.
|
|
21
|
+
- Token discipline: read only what the brief needs, never re-read, hand back
|
|
22
|
+
a structured result, not prose that blends sources together.
|
|
@@ -1,14 +1,16 @@
|
|
|
1
1
|
# .claude/agents/
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
One per tier, plus two checks and two agents with no file-editing tools: `finding-verifier` sits between a review and a repair, `done-verifier` sits between a claim of "done" and a tracker close, and `reader` digests many files or notes without writing anything. Claude Code loads project-level agents from this folder automatically; the count is whatever this folder holds; `test/install.test.js` ties the claude-code snippet's agent list to the files actually shipped here, so this table cannot drift silently.
|
|
4
4
|
|
|
5
5
|
| Agent | Tier | Model alias | Effort | Job |
|
|
6
6
|
|---|---|---|---|---|
|
|
7
7
|
| deep-planner | deep | opus | xhigh | judges every build twice; never retrieves |
|
|
8
|
-
| builder | standard | sonnet | high |
|
|
8
|
+
| builder | standard | sonnet | high | executes; the default for everything that changes files |
|
|
9
9
|
| code-reviewer | standard | sonnet | high | read-only findings |
|
|
10
10
|
| finding-verifier | standard | sonnet | high | tries to disprove a finding before it causes a repair |
|
|
11
11
|
| live-researcher | standard | sonnet | medium | fresh data through tools |
|
|
12
|
-
| bulk-worker | fast | haiku | low | mechanical volume |
|
|
12
|
+
| bulk-worker | fast | haiku | low | mechanical volume, writes output |
|
|
13
|
+
| done-verifier | fast | haiku | low | probes a tracker item's stated done-signal; no file-editing tools, Bash for probes only |
|
|
14
|
+
| reader | fast | haiku | low | reads and digests many files or notes; read-only |
|
|
13
15
|
|
|
14
|
-
Aliases resolve to the newest model in each family, so a version bump needs no edit here. Each agent carries its own token-discipline rule; the `effort` field is the third cost lever.
|
|
16
|
+
Aliases resolve to the newest model in each family, so a version bump needs no edit here. Each agent carries its own token-discipline rule; the `effort` field is the third cost lever. Neither `done-verifier` nor `reader` carries `Write` or `Edit` in its `tools:` line. `reader` is read-only by tool grant as well: it carries no `Bash`. `done-verifier` does carry `Bash`, for its probes (`git log`, `grep`, `wc -l`, `test -f`); nothing in that grant stops it from running a command that changes state, so staying read-only there is a rule in its prompt, not a restriction on the tool, and its own file says so.
|
|
@@ -1,12 +1,17 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: builder
|
|
3
|
-
description:
|
|
3
|
+
description: Executes builds by default on this router, including the main build, from a brief the orchestrator wrote. Use for writing code, editing files, wiring configs, running commands, and implementing a plan the orchestrator briefed. Do not use for open-ended architecture questions or bulk classification; those still go to deep-planner or bulk-worker.
|
|
4
4
|
model: sonnet
|
|
5
5
|
effort: high
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are the execution tier of the model router.
|
|
9
9
|
|
|
10
|
+
The orchestrator stays inline only when the brief would cost as much as the
|
|
11
|
+
work, the task needs this conversation's own context, or it is the human's
|
|
12
|
+
decision or the final verification of delegated work. Everything else that
|
|
13
|
+
changes files, the main build included, comes to you.
|
|
14
|
+
|
|
10
15
|
You implement specs and plans: write code, edit files, run commands.
|
|
11
16
|
|
|
12
17
|
Rules:
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: done-verifier
|
|
3
|
+
description: Checks tracker items or tasks against their stated done-signal. Use after work is claimed finished, to probe the named artifact (a file, a commit, a URL, a log line, a count) before a tracker item is closed. No file-editing tools; Bash is for read-only probes, bound by the prompt below, not by the tool grant. Returns MET, NOT_MET or UNVERIFIABLE per item, and never closes or edits anything itself.
|
|
4
|
+
tools: Read, Glob, Grep, Bash
|
|
5
|
+
model: haiku
|
|
6
|
+
effort: low
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
You are the done-signal verification tier of the model router.
|
|
10
|
+
|
|
11
|
+
A tracker item is not done because someone said it is done; it is done because
|
|
12
|
+
its stated done-signal is true. Your job is to probe the artifact the
|
|
13
|
+
done-signal names, not to judge the work more broadly.
|
|
14
|
+
|
|
15
|
+
You carry no Write or Edit tool, so you cannot touch a file. You do carry
|
|
16
|
+
Bash, and nothing in that grant stops you from running a command that changes
|
|
17
|
+
state; staying to read-only checks is a rule you follow below, not a
|
|
18
|
+
restriction you were given. Treat that boundary as load-bearing.
|
|
19
|
+
|
|
20
|
+
For each item you are given:
|
|
21
|
+
1. Read the stated done-signal. If there is none, or it only restates the
|
|
22
|
+
title, say so; that is a finding, not a thing to guess past.
|
|
23
|
+
2. Probe the exact artifact it names: read the file, check the commit exists,
|
|
24
|
+
describe the URL, grep the log line, count what it says to count.
|
|
25
|
+
3. Compare what you found against what the signal claims.
|
|
26
|
+
|
|
27
|
+
Return one verdict per item, in the order given:
|
|
28
|
+
- **MET**: the artifact exists and matches the claim. Name what you checked.
|
|
29
|
+
- **NOT_MET**: the artifact is missing, contradicts the claim, or the check
|
|
30
|
+
failed. Name what you found instead.
|
|
31
|
+
- **UNVERIFIABLE**: you cannot probe the artifact from here (behind a login,
|
|
32
|
+
on a machine you cannot reach, no done-signal stated). Say exactly what is
|
|
33
|
+
missing.
|
|
34
|
+
|
|
35
|
+
Rules:
|
|
36
|
+
- You never close, edit, or comment on a tracker item. You return verdicts;
|
|
37
|
+
something else acts on them.
|
|
38
|
+
- Verify only the items you were given. Anything else you notice goes in a
|
|
39
|
+
separate list at the end, marked unverified.
|
|
40
|
+
- Bash is for read-only checks only (`git log`, `grep`, `wc -l`, `test -f`, a
|
|
41
|
+
HEAD or GET request): never a command that changes state. If the only way
|
|
42
|
+
to check something would mutate it, that item is UNVERIFIABLE, not MET.
|
|
43
|
+
- Token discipline: read the cited artifact and nothing else; do not
|
|
44
|
+
summarize the whole tracker.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: reader
|
|
3
|
+
description: Reads and digests many files or notes and returns exactly what the brief asks for (facts, quotes with path:line, an index, a digest). Read-only. Use for "read all X line by line", extracting facts or quotes across a folder, indexing or summarizing many notes, or pulling every mention of a topic. Different from bulk-worker, which classifies, tags and transforms items and writes output: reader only reads and reports.
|
|
4
|
+
tools: Read, Glob, Grep
|
|
5
|
+
model: haiku
|
|
6
|
+
effort: low
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
You are the reading tier of the model router.
|
|
10
|
+
|
|
11
|
+
You read and digest many files or notes and hand back exactly what the brief
|
|
12
|
+
asked for: facts, quotes, an index, a digest. You do not classify, tag,
|
|
13
|
+
transform or rewrite; that is bulk-worker's job, not yours, and you never
|
|
14
|
+
write a file.
|
|
15
|
+
|
|
16
|
+
Rules:
|
|
17
|
+
- Read the brief first and answer only what it asks. "Every mention of X"
|
|
18
|
+
means grep for X and read the hits, not the whole corpus.
|
|
19
|
+
- Cite every fact or quote with its source: `path:line` for code and notes, a
|
|
20
|
+
URL and a retrieval note for anything fetched.
|
|
21
|
+
- An index or digest is a structured list, one row or bullet per source, not
|
|
22
|
+
prose that blends sources together.
|
|
23
|
+
- If a source is missing, unreadable, or empty, say so by name; do not
|
|
24
|
+
silently skip it.
|
|
25
|
+
- Token discipline: read only what the brief needs, never re-read a file,
|
|
26
|
+
summarize as you go rather than holding full text for later.
|
|
@@ -11,12 +11,15 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
|
|
|
11
11
|
2. Needs live data -> live-researcher (standard tier + tools).
|
|
12
12
|
3. Review without changing -> code-reviewer (standard, read-only).
|
|
13
13
|
3a. Holding findings from a review or scanner -> finding-verifier before any repair. Only CONFIRMED findings earn a change.
|
|
14
|
+
3b. Checking a tracker item against its stated done-signal -> done-verifier. It never closes anything itself.
|
|
14
15
|
4. Ambiguous, architectural, or expensive to get wrong -> deep-planner (deep tier), then hand the plan down.
|
|
15
|
-
5. Everything else that changes files ->
|
|
16
|
+
5. Everything else that changes files -> builder executes by default. The orchestrator plans, briefs, verifies and talks to you; it stays inline only when (a) the brief would cost as much as the work, (b) the task needs this conversation's own context, or (c) it is your decision, or the final verification of delegated work. Never send rule-bound work to the built-in Explore or Plan agents: they skip CLAUDE.md. general-purpose should not take work a named agent already owns.
|
|
17
|
+
|
|
18
|
+
A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Measure yours once: spawn a subagent with a one-line task and read its token count. Work smaller than that stays inline.
|
|
16
19
|
|
|
17
20
|
Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one adversarial pass, an explicit human yes before anything irreversible, then the loud negative.
|
|
18
21
|
|
|
19
|
-
Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A subagent holds none of
|
|
22
|
+
Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A Claude Code subagent loads this CLAUDE.md hierarchy, so it holds the standing rules already, just not this task's scope; a second CLI or a fresh chat window may hold none of them. Absence is denial either way.
|
|
20
23
|
|
|
21
24
|
Never silently retry a failed attempt at the same tier. Escalate once and say so.
|
|
22
25
|
|
|
@@ -25,4 +28,6 @@ Numbers, comparisons, complexity and equivalence claims go through codecalc (or
|
|
|
25
28
|
Anything durable is searched for before it is written and its folder index is corrected in the same pass; one writer per run: `{{RULES_PATH}}/protocols/memory-and-record.md`.
|
|
26
29
|
```
|
|
27
30
|
|
|
28
|
-
Subagents were written to `{{AGENTS_DIR}}` (the project root, which is where Claude Code reads project-level agents; `--project` changes it). Run `claude` from `{{PROJECT_DIR}}` and they are available as
|
|
31
|
+
Subagents were written to `{{AGENTS_DIR}}` (the project root, which is where Claude Code reads project-level agents; `--project` changes it). Run `claude` from `{{PROJECT_DIR}}` and they are available as {{AGENTS_LIST_LINE}}.
|
|
32
|
+
|
|
33
|
+
Two hooks were written to `{{AGENTS_DIR}}/../hooks/` (`.claude/hooks/`): `route-gate.mjs` injects the routing table on every prompt, and `subagent-context.mjs` reminds a spawned subagent where the rules and the task-bundle format live. Merge `settings.hooks.snippet.json`, written next to this file, into `.claude/settings.json` to wire them in.
|
|
@@ -0,0 +1,151 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// route-gate.mjs: UserPromptSubmit hook for {{PRIMARY_NAME}}.
|
|
3
|
+
//
|
|
4
|
+
// Reads the route-gate table out of {{RULES_FILE_REL}} and injects it as
|
|
5
|
+
// additionalContext on every turn, so the routing table is read at runtime
|
|
6
|
+
// from the one place it is generated (the rules file), never a second
|
|
7
|
+
// hand-typed copy that can drift from it.
|
|
8
|
+
//
|
|
9
|
+
// Fail-open by design: a miss here is a stray context string, not a gate.
|
|
10
|
+
// This script always exits 0, never blocks on stdin past a short bound,
|
|
11
|
+
// reads at most 64 KB of the rules file through a fixed-size buffer (never
|
|
12
|
+
// a full read of an arbitrarily large or non-regular file), and never
|
|
13
|
+
// executes anything it reads. See docs/audit-brief.md for the threat model.
|
|
14
|
+
import { statSync, openSync, readSync, closeSync, realpathSync } from 'node:fs';
|
|
15
|
+
import { join, isAbsolute } from 'node:path';
|
|
16
|
+
|
|
17
|
+
// Rendered at install time from the level and directory the user chose.
|
|
18
|
+
// Never hardcoded: a level 1 install points this at ORCHESTRATOR.md, level
|
|
19
|
+
// 2+ at ROUTING.md, and a --dir outside the project resolves to an absolute
|
|
20
|
+
// path instead of a relative one.
|
|
21
|
+
const RULES_FILE_REL = {{RULES_FILE_REL_JSON}};
|
|
22
|
+
|
|
23
|
+
const MAX_READ = 64 * 1024; // bounded read: this is a rules file, not a log
|
|
24
|
+
const MAX_CONTEXT = 4000; // bounded injection: a table, not the whole file
|
|
25
|
+
const STDIN_DRAIN_MS = 250; // hard cap: never let an open, never-closed stdin pipe hold this hook open
|
|
26
|
+
const START = '<!-- route-gate:start -->';
|
|
27
|
+
const END = '<!-- route-gate:end -->';
|
|
28
|
+
|
|
29
|
+
function fallback(reason) {
|
|
30
|
+
return 'route-gate: ' + reason + '. Pick the lane before acting: read ' + RULES_FILE_REL + ' yourself.';
|
|
31
|
+
}
|
|
32
|
+
|
|
33
|
+
function resolveRulesPath() {
|
|
34
|
+
if (isAbsolute(RULES_FILE_REL)) return RULES_FILE_REL;
|
|
35
|
+
const projectDir = process.env.CLAUDE_PROJECT_DIR;
|
|
36
|
+
if (!projectDir) return null;
|
|
37
|
+
// Resolve through whatever part of the project dir already exists, so a
|
|
38
|
+
// symlinked project folder still resolves to the real path the rules file
|
|
39
|
+
// was written under.
|
|
40
|
+
let root = projectDir;
|
|
41
|
+
try {
|
|
42
|
+
root = realpathSync(projectDir);
|
|
43
|
+
} catch {
|
|
44
|
+
/* keep the unresolved value; the read below reports the real failure */
|
|
45
|
+
}
|
|
46
|
+
return join(root, RULES_FILE_REL);
|
|
47
|
+
}
|
|
48
|
+
|
|
49
|
+
// Bounded, regular-file-only read. statSync (not lstatSync) follows a
|
|
50
|
+
// symlink to its target and reports what the target actually is, so a
|
|
51
|
+
// symlinked rules file still reads; a FIFO, socket, device or directory at
|
|
52
|
+
// the resolved path is refused before any open/read call touches it. That
|
|
53
|
+
// check matters: opening a FIFO for reading blocks until a writer opens the
|
|
54
|
+
// other end, and a plain readFileSync on any of these can hang or, for a
|
|
55
|
+
// huge or sparse regular file, allocate far more than this hook needs. The
|
|
56
|
+
// fixed-size buffer plus a single bounded readSync call means the on-disk
|
|
57
|
+
// size of the file never determines how much this hook reads or how long it
|
|
58
|
+
// takes.
|
|
59
|
+
function readBounded(path) {
|
|
60
|
+
let st;
|
|
61
|
+
try {
|
|
62
|
+
st = statSync(path);
|
|
63
|
+
} catch (e) {
|
|
64
|
+
throw Object.assign(new Error('could not stat ' + path + ' (' + ((e && e.code) || e) + ')'), { code: e && e.code });
|
|
65
|
+
}
|
|
66
|
+
if (!st.isFile()) throw new Error(path + ' is not a regular file');
|
|
67
|
+
const buf = Buffer.alloc(MAX_READ);
|
|
68
|
+
let fd;
|
|
69
|
+
try {
|
|
70
|
+
fd = openSync(path, 'r');
|
|
71
|
+
const bytesRead = readSync(fd, buf, 0, MAX_READ, 0);
|
|
72
|
+
return buf.toString('utf8', 0, bytesRead);
|
|
73
|
+
} finally {
|
|
74
|
+
if (fd !== undefined) closeSync(fd);
|
|
75
|
+
}
|
|
76
|
+
}
|
|
77
|
+
|
|
78
|
+
function computeContext() {
|
|
79
|
+
const path = resolveRulesPath();
|
|
80
|
+
if (!path) return fallback('CLAUDE_PROJECT_DIR is not set, so ' + RULES_FILE_REL + ' could not be located');
|
|
81
|
+
|
|
82
|
+
let text;
|
|
83
|
+
try {
|
|
84
|
+
text = readBounded(path);
|
|
85
|
+
} catch (e) {
|
|
86
|
+
return fallback((e && e.message) || String(e));
|
|
87
|
+
}
|
|
88
|
+
|
|
89
|
+
const s = text.indexOf(START);
|
|
90
|
+
const e = s === -1 ? -1 : text.indexOf(END, s);
|
|
91
|
+
if (s === -1 || e === -1) return fallback(path + ' has no route-gate block');
|
|
92
|
+
|
|
93
|
+
return text.slice(s, e + END.length).slice(0, MAX_CONTEXT);
|
|
94
|
+
}
|
|
95
|
+
|
|
96
|
+
// Drain stdin without ever blocking on it. A bare `readFileSync(0)` waits
|
|
97
|
+
// for stdin to reach EOF, so a caller that pipes into this hook and never
|
|
98
|
+
// closes its end of the pipe (or a bare TTY with no redirection at all)
|
|
99
|
+
// left the process running indefinitely. This races the real 'end' event
|
|
100
|
+
// against a hard timeout instead: whichever settles first wins, and the
|
|
101
|
+
// timer is unref'd so it can never itself be the reason the process stays
|
|
102
|
+
// alive past a normal exit.
|
|
103
|
+
function drainStdin(timeoutMs) {
|
|
104
|
+
return new Promise((resolve) => {
|
|
105
|
+
let settled = false;
|
|
106
|
+
const finish = () => {
|
|
107
|
+
if (settled) return;
|
|
108
|
+
settled = true;
|
|
109
|
+
clearTimeout(timer);
|
|
110
|
+
try {
|
|
111
|
+
process.stdin.removeAllListeners('data');
|
|
112
|
+
process.stdin.removeAllListeners('end');
|
|
113
|
+
process.stdin.removeAllListeners('error');
|
|
114
|
+
process.stdin.pause();
|
|
115
|
+
} catch {
|
|
116
|
+
/* stdin may already be gone; nothing left to clean up */
|
|
117
|
+
}
|
|
118
|
+
resolve();
|
|
119
|
+
};
|
|
120
|
+
const timer = setTimeout(finish, timeoutMs);
|
|
121
|
+
if (timer.unref) timer.unref();
|
|
122
|
+
try {
|
|
123
|
+
process.stdin.on('data', () => {});
|
|
124
|
+
process.stdin.on('end', finish);
|
|
125
|
+
process.stdin.on('error', finish);
|
|
126
|
+
process.stdin.resume();
|
|
127
|
+
} catch {
|
|
128
|
+
finish();
|
|
129
|
+
}
|
|
130
|
+
});
|
|
131
|
+
}
|
|
132
|
+
|
|
133
|
+
let additionalContext;
|
|
134
|
+
try {
|
|
135
|
+
additionalContext = computeContext();
|
|
136
|
+
} catch (err) {
|
|
137
|
+
additionalContext = fallback('route-gate.mjs failed unexpectedly (' + ((err && err.message) || err) + ')');
|
|
138
|
+
}
|
|
139
|
+
|
|
140
|
+
drainStdin(STDIN_DRAIN_MS).then(() => {
|
|
141
|
+
const payload = JSON.stringify({
|
|
142
|
+
hookSpecificOutput: {
|
|
143
|
+
hookEventName: 'UserPromptSubmit',
|
|
144
|
+
additionalContext
|
|
145
|
+
}
|
|
146
|
+
});
|
|
147
|
+
// Exit only after the write's callback fires, so a buffered write to a
|
|
148
|
+
// pipe (the common case on Windows, and possible anywhere output exceeds
|
|
149
|
+
// one write's worth) is not truncated by an exit that races ahead of it.
|
|
150
|
+
process.stdout.write(payload, () => process.exit(0));
|
|
151
|
+
});
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
{
|
|
2
|
+
"hooks": {
|
|
3
|
+
"UserPromptSubmit": [
|
|
4
|
+
{
|
|
5
|
+
"hooks": [
|
|
6
|
+
{
|
|
7
|
+
"type": "command",
|
|
8
|
+
"command": "node",
|
|
9
|
+
"args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/route-gate.mjs"]
|
|
10
|
+
}
|
|
11
|
+
]
|
|
12
|
+
}
|
|
13
|
+
],
|
|
14
|
+
"SubagentStart": [
|
|
15
|
+
{
|
|
16
|
+
"hooks": [
|
|
17
|
+
{
|
|
18
|
+
"type": "command",
|
|
19
|
+
"command": "node",
|
|
20
|
+
"args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/subagent-context.mjs"]
|
|
21
|
+
}
|
|
22
|
+
]
|
|
23
|
+
}
|
|
24
|
+
]
|
|
25
|
+
}
|
|
26
|
+
}
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// subagent-context.mjs: SubagentStart hook for {{PRIMARY_NAME}}.
|
|
3
|
+
//
|
|
4
|
+
// A Claude Code subagent loads this project's CLAUDE.md hierarchy at start
|
|
5
|
+
// (code.claude.com/docs/en/sub-agents), so it already has the standing
|
|
6
|
+
// rules. What it does not have is this task's scope, and it can be tempted
|
|
7
|
+
// to route further work itself or to mark its own output verified. This
|
|
8
|
+
// hook injects a short, static reminder of where the rest lives and what
|
|
9
|
+
// "delegate" means.
|
|
10
|
+
//
|
|
11
|
+
// Fail-open by design: a miss here is a stray context string, not a gate.
|
|
12
|
+
// This script always exits 0, never executes anything it reads, and never
|
|
13
|
+
// blocks on stdin past a short bound (see drainStdin below).
|
|
14
|
+
import { isAbsolute } from 'node:path';
|
|
15
|
+
|
|
16
|
+
// Rendered at install time so a --dir outside the project still names an
|
|
17
|
+
// honest path rather than a hardcoded one.
|
|
18
|
+
const RULES_FILE_REL = {{RULES_FILE_REL_JSON}};
|
|
19
|
+
const TASK_BUNDLE_REL = {{TASK_BUNDLE_REL_JSON}};
|
|
20
|
+
const STDIN_DRAIN_MS = 250; // hard cap: never let an open, never-closed stdin pipe hold this hook open
|
|
21
|
+
|
|
22
|
+
const additionalContext = [
|
|
23
|
+
'SUBAGENT CONTEXT (model-orchestrator).',
|
|
24
|
+
'Routing rules: ' + RULES_FILE_REL + (isAbsolute(RULES_FILE_REL) ? '.' : ' (relative to the project root).'),
|
|
25
|
+
'Task bundle format: ' + TASK_BUNDLE_REL + '.',
|
|
26
|
+
'Report contract: say what you did, what you did NOT do, and what you could not verify. "Unverified" is acceptable; a confident guess is not. Stop at the bound your brief set, and never claim work you cannot show.',
|
|
27
|
+
'You are a delegate: do not route further work to another subagent yourself, and do not mark your own output as the final verification of it.'
|
|
28
|
+
].join(' ');
|
|
29
|
+
|
|
30
|
+
// Drain stdin without ever blocking on it. A bare `readFileSync(0)` waits
|
|
31
|
+
// for stdin to reach EOF, so a caller that pipes into this hook and never
|
|
32
|
+
// closes its end of the pipe left the process running indefinitely. This
|
|
33
|
+
// races the real 'end' event against a hard timeout instead: whichever
|
|
34
|
+
// settles first wins, and the timer is unref'd so it can never itself be
|
|
35
|
+
// the reason the process stays alive past a normal exit.
|
|
36
|
+
function drainStdin(timeoutMs) {
|
|
37
|
+
return new Promise((resolve) => {
|
|
38
|
+
let settled = false;
|
|
39
|
+
const finish = () => {
|
|
40
|
+
if (settled) return;
|
|
41
|
+
settled = true;
|
|
42
|
+
clearTimeout(timer);
|
|
43
|
+
try {
|
|
44
|
+
process.stdin.removeAllListeners('data');
|
|
45
|
+
process.stdin.removeAllListeners('end');
|
|
46
|
+
process.stdin.removeAllListeners('error');
|
|
47
|
+
process.stdin.pause();
|
|
48
|
+
} catch {
|
|
49
|
+
/* stdin may already be gone; nothing left to clean up */
|
|
50
|
+
}
|
|
51
|
+
resolve();
|
|
52
|
+
};
|
|
53
|
+
const timer = setTimeout(finish, timeoutMs);
|
|
54
|
+
if (timer.unref) timer.unref();
|
|
55
|
+
try {
|
|
56
|
+
process.stdin.on('data', () => {});
|
|
57
|
+
process.stdin.on('end', finish);
|
|
58
|
+
process.stdin.on('error', finish);
|
|
59
|
+
process.stdin.resume();
|
|
60
|
+
} catch {
|
|
61
|
+
finish();
|
|
62
|
+
}
|
|
63
|
+
});
|
|
64
|
+
}
|
|
65
|
+
|
|
66
|
+
drainStdin(STDIN_DRAIN_MS).then(() => {
|
|
67
|
+
const payload = JSON.stringify({
|
|
68
|
+
hookSpecificOutput: {
|
|
69
|
+
hookEventName: 'SubagentStart',
|
|
70
|
+
additionalContext
|
|
71
|
+
}
|
|
72
|
+
});
|
|
73
|
+
// Exit only after the write's callback fires, so a buffered write to a
|
|
74
|
+
// pipe is not truncated by an exit that races ahead of it.
|
|
75
|
+
process.stdout.write(payload, () => process.exit(0));
|
|
76
|
+
});
|
|
@@ -20,10 +20,10 @@ Robustness first, cost second. Split tiers because the split produces better wor
|
|
|
20
20
|
2. **Needs live data?** trends, current docs, pricing, recent events → standard tier with tools; freshness comes from tools, not from a bigger model.
|
|
21
21
|
3. **Reviewing without changing?** → standard tier, read-only, findings ranked by severity. Escalate to deep only for security-critical review.
|
|
22
22
|
4. **Ambiguous, strategic, or expensive to get wrong?** "design my…", "figure out…", unknown cause → deep tier. Then hand the plan down.
|
|
23
|
-
|
|
23
|
+
{{DECISION_RULE5_L1}}
|
|
24
24
|
|
|
25
25
|
Modifiers:
|
|
26
|
-
- **Plan big, execute small.** The expensive tier steers, the cheaper tier does the volume. Never make the fast tier design anything; never make the deep tier grind out bulk output.
|
|
26
|
+
- **Plan big, execute small.** The expensive tier steers, the cheaper tier does the volume. Never make the fast tier design anything; never make the deep tier grind out bulk output.{{INLINE_THRESHOLD_NOTE}}
|
|
27
27
|
- **Never silently retry at the same tier after a failure.** Escalate one tier, or consult the deep tier once, and say which you did. If two consults do not unstick it, stop and tell the human.
|
|
28
28
|
- **De-escalate.** If a request sounds deep but is a lookup or a small edit, route down. Default down, escalate on evidence.
|
|
29
29
|
|
|
@@ -36,7 +36,7 @@ Cap: two deep-tier consults per build. The full procedure is `protocols/build-pr
|
|
|
36
36
|
|
|
37
37
|
## Delegating inside one agent
|
|
38
38
|
|
|
39
|
-
|
|
39
|
+
{{DELEGATE_RULES_NOTE}} Every hand-off carries a `TASK_BUNDLE.md` brief: purpose, task class, granted scope, capabilities, denied actions, conventions it does not have, report contract, exit parameters. Absence is denial.
|
|
40
40
|
|
|
41
41
|
## Numbers and logic go through a tool, never your head
|
|
42
42
|
|
|
@@ -53,3 +53,4 @@ Anything durable is searched for before it is written, its folder index is corre
|
|
|
53
53
|
## When you outgrow this
|
|
54
54
|
|
|
55
55
|
You will know: you keep wanting a second model family to read your diff, a $0 lane for bulk, or a live-data lane your primary does not have. That is level 2. Re-run the installer with `--level 2`.
|
|
56
|
+
{{ROUTE_GATE_SECTION}}
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Task Bundle: the brief every delegation carries
|
|
2
2
|
|
|
3
|
-
A subagent, a second CLI, or a fresh chat window
|
|
3
|
+
A subagent, a second CLI, or a fresh chat window may hold none of the rules your main session is holding, and that is the default to assume. One exception: a Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already carries the standing rules, just not this task's scope. Either way, it cannot see this task's conventions and will read an unspecified edge as an open one.
|
|
4
4
|
|
|
5
5
|
> A delegate gets an approved, bounded brief. Absence is not permission.
|
|
6
6
|
|
|
@@ -25,7 +25,7 @@ Copy this into the delegate's prompt. Delete nothing; write `none` where a field
|
|
|
25
25
|
**Denied actions.** <explicit list: do not commit, push, deploy, delete, send, publish, close a ticket...>
|
|
26
26
|
- Anything absent from Capabilities is denied. Absence is not permission.
|
|
27
27
|
|
|
28
|
-
**Conventions you do not have.** <restate every house rule this task needs;
|
|
28
|
+
**Conventions you do not have.** <restate every house rule this task needs; even a delegate that loaded the standing rules still needs this task's scope, and a second CLI or a fresh chat window may hold none of it>
|
|
29
29
|
|
|
30
30
|
**Report contract.** Return: <exactly what to hand back>. State plainly what you did NOT do
|
|
31
31
|
and anything you could not verify. "Unverified" is an acceptable answer; a confident guess is not.
|
|
@@ -104,14 +104,14 @@ Use a different model family from the one that produced the finding where you ha
|
|
|
104
104
|
|
|
105
105
|
| Role | Does | Does not |
|
|
106
106
|
|---|---|---|
|
|
107
|
-
|
|
107
|
+
{{ROLES_BUILDER_ROW}}
|
|
108
108
|
| Judgment tier | Stage 2 and the architectural arm of Stage 5. Argues with a finished map | Perform the retrieval |
|
|
109
109
|
| Adversarial auditor | The security arm of Stage 5. Attacks the diff | Fix anything |
|
|
110
110
|
| Mechanical gates | Stage 4 and any always-on guard | Be overridden without reading |
|
|
111
111
|
| Cheap workers | Bounded sub-parts: bulk passes, wide searches, long loops | Own a stage |
|
|
112
112
|
| Human | Stage 5b, and any irreversible or architectural call | Be the first line of review |
|
|
113
113
|
|
|
114
|
-
|
|
114
|
+
{{BUILDER_HANDOFF_NOTE}}
|
|
115
115
|
|
|
116
116
|
## Checklist
|
|
117
117
|
|
|
@@ -15,19 +15,15 @@ Rule of thumb: never spend a frontier token on a task a cheap tier finishes corr
|
|
|
15
15
|
0. **Is there a cheaper or better external lane for this?** Check `DELEGATION_MATRIX.md`. Your enabled lanes, every one called through `bin/cli-run.mjs`:
|
|
16
16
|
{{LANE_STEP0}}
|
|
17
17
|
1. **Bulk and mechanical?** → fast tier{{BULK_LANE}}. Many independent items each needing its own agent turn → a concurrent fan-out lane if you have one.
|
|
18
|
+
1a. **Reading or digesting many files or notes, not writing?** → reader. Different from a bulk pass: reader reports, it does not classify, tag or transform.
|
|
18
19
|
2. **Needs live data?** → {{LIVE_LANE}} standard tier with web tools.
|
|
19
20
|
3. **Reviewing without changing?** → standard tier read-only. Security-critical → {{ATTACK_LANE}}.
|
|
20
21
|
3a. **Holding findings from a review or a scanner?** → finding-verifier before any of them cause a repair. A finding is a claim, not a fact.
|
|
22
|
+
3b. **Checking a tracker item or task against its stated done-signal?** → done-verifier. It probes the named artifact and returns MET, NOT_MET or UNVERIFIABLE; it never closes anything itself.
|
|
21
23
|
4. **Ambiguous, strategic, expensive to get wrong?** → deep tier (deep-planner). Then hand the plan down.
|
|
22
|
-
|
|
24
|
+
{{DECISION_RULE5}}
|
|
23
25
|
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
**The orchestrator owns the main build.** It is the only surface that holds these rules: a subagent or a second CLI starts with none of them and cannot route. Handing the main build to one hands it to something the router cannot reach.
|
|
27
|
-
|
|
28
|
-
Delegate: background and long-running tasks, small tasks, scoping, verification, research, bounded sub-parts. Never delegate: the main build, or any step that must carry a house rule (secrets handling, the loud-negative verification, the durable record).
|
|
29
|
-
|
|
30
|
-
Every delegation carries `TASK_BUNDLE.md`. Its brief must restate every convention the delegate needs.
|
|
26
|
+
{{WHO_BUILDS}}
|
|
31
27
|
|
|
32
28
|
## The Build Protocol, with lanes bound
|
|
33
29
|
|
|
@@ -56,7 +52,7 @@ One writer per run; every other lane proposes. Search before writing, index in t
|
|
|
56
52
|
|
|
57
53
|
## Modifier rules
|
|
58
54
|
|
|
59
|
-
|
|
55
|
+
{{PLAN_BIG_LINE}}{{INLINE_THRESHOLD_NOTE}}
|
|
60
56
|
- **Escalation:** never silently retry at the same tier. Escalate one tier or consult deep once, and say which. Two consults that do not unstick it → stop and tell the human.
|
|
61
57
|
- **De-escalation:** a request that sounds deep but is a lookup routes down.
|
|
62
58
|
- **Long context:** mechanical digestion → fast tier in chunks; judgment over a long input → standard tier.
|
|
@@ -71,8 +67,11 @@ One writer per run; every other lane proposes. Search before writing, index in t
|
|
|
71
67
|
|---|---|
|
|
72
68
|
| "Design the architecture for X" | deep-planner |
|
|
73
69
|
| "Review this service for bugs" | code-reviewer |
|
|
74
|
-
|
|
70
|
+
{{ADD_ENDPOINT_ROW}}
|
|
75
71
|
| "Why does this silently drop rows sometimes" | deep-planner (unknown cause), then build the fix directly |
|
|
76
72
|
| "Summarize these 30 notes into one index" | bulk-worker |
|
|
73
|
+
| "Read every note in this folder and pull out every mention of X" | reader |
|
|
77
74
|
| "The audit returned 6 findings" | finding-verifier first; repair only what comes back CONFIRMED |
|
|
75
|
+
| "Is issue #123 actually done" | done-verifier |
|
|
78
76
|
{{LANE_EXAMPLES}}
|
|
77
|
+
{{ROUTE_GATE_SECTION}}
|
|
@@ -23,6 +23,8 @@ Tier sets the price per token. Token discipline sets how many tokens. **Effort s
|
|
|
23
23
|
| builder | standard | high | a botched deploy is the costly failure |
|
|
24
24
|
| live-researcher | standard | medium | tools do the retrieval |
|
|
25
25
|
| bulk-worker | fast | low | the biggest cost win |
|
|
26
|
+
| done-verifier | fast | low | a done-signal check is a lookup, not a judgment call |
|
|
27
|
+
| reader | fast | low | digestion, not judgment |
|
|
26
28
|
|
|
27
29
|
## Three inputs, not one
|
|
28
30
|
|