model-orchestrator 0.1.16 → 0.1.17

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,12 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.17] - 2026-09-11
8
+
9
+ ### Changed
10
+
11
+ - **User-facing text now uses plain language instead of security-audit jargon.** Words like "risk", "attack lane", "adversarial", "blast radius", "fail closed" and "threat model" read as alarming to someone deciding whether to try the tool, so they scared off exactly the readers this project needs. No rule any of them described changed, only the wording: "risk" is now "stakes" everywhere it names a routing input (with a one-line definition added to `README.md` and `TIERS.md`), "attack lane" / "Stage 5 Attack" / "attack pass" are now "challenge lane" / "Stage 5 Challenge" / "challenge pass", "adversarial" (auditor, read, critique, turn, pass) is now "second-opinion", "blast radius" is now "everything it touches", "fail(s) closed" is now "refuses by default", and "threat model" is now "security notes" in the files that link to it. `test/prose.test.js` gained a permanent check (`no alarming security wording in user-facing text`) over the purely-prose, user-facing surface (`docs/`, `templates/`, `README.md`, `llms.txt`, `CONTRIBUTING.md`, the PR template) so the old wording cannot silently creep back in. `docs/audit-brief.md`, `SECURITY.md`, `CODE_OF_CONDUCT.md` and code identifiers/comments (for example the `ATTACK_LANE` render var) are unchanged, since these words are expected or load-bearing there.
12
+
7
13
  ## [0.1.16] - 2026-09-11
8
14
 
9
15
  ### Added
@@ -256,7 +262,8 @@ First release.
256
262
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
257
263
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
258
264
 
259
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.16...HEAD
265
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.17...HEAD
266
+ [0.1.17]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.16...v0.1.17
260
267
  [0.1.16]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.15...v0.1.16
261
268
  [0.1.15]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...v0.1.15
262
269
  [0.1.14]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...v0.1.14
package/README.md CHANGED
@@ -44,7 +44,7 @@ Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/pa
44
44
  | Id | What | Level |
45
45
  |---|---|---|
46
46
  | `claude-code` | Claude Code CLI, the default orchestrator | 1+ |
47
- | `codex` | Codex CLI on a ChatGPT plan: second coder, adversarial auditor | 1+ |
47
+ | `codex` | Codex CLI on a ChatGPT plan: second coder and second-opinion reviewer (a different model family reading your diff) | 1+ |
48
48
  | `agy` | Antigravity CLI on a Google AI plan: research sweeps, concurrent fan-out | 1+ |
49
49
  | `grok` | Grok CLI on X Premium: live X and web reads at $0 | 1+ |
50
50
  | `hermes` | Hermes Agent: the free tier | 2+ |
@@ -173,7 +173,7 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
173
173
  6. **A delegate's brief carries this task's scope, whatever it already holds.** A Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already has the standing rules; a second CLI or a fresh chat window may hold none of them. Either way, only the brief carries what this task needs. On claude-code, that changes who executes: see "Who builds" in `ROUTING.md`.
174
174
  7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
175
175
 
176
- ## Routing by role, complexity and risk
176
+ ## Routing by role, complexity and stakes
177
177
 
178
178
  Role picks the agent. Two more inputs move the choice, and they move it in
179
179
  different directions, so `TIERS.md` states them separately rather than folding
@@ -182,10 +182,14 @@ them into the role:
182
182
  - **Complexity moves the effort.** A worker executing a finished plan needs less
183
183
  reasoning than the reviewer judging its output. When the plan is airtight the
184
184
  spec is carrying the thinking.
185
- - **Risk moves the tier and the reader.** Security, privacy, data loss and
186
- irreversible changes buy the attack lane, a named check, a rollback path or a
187
- human yes. A one-line change to an auth check is simple and high-risk at the
188
- same time, and it is the risk that decides.
185
+ - **Stakes move the tier and the reader.** Security, privacy, data loss and
186
+ irreversible changes buy the challenge lane, a named check, a rollback path or
187
+ a human yes. A one-line change to an auth check is simple and high-stakes at
188
+ the same time, and it is the stakes that decide.
189
+
190
+ Stakes means what a mistake would cost: a security hole, leaked personal data,
191
+ lost data, or something you can't undo. Most tasks are low-stakes and route
192
+ normally.
189
193
 
190
194
  The top of the ladder is bought with evidence: a reproduced failure, an
191
195
  unresolved checkpoint, an irreversible change. A task that merely feels hard is
@@ -221,7 +225,7 @@ The report prints turns, **route-marker coverage** (the percentage of turns whos
221
225
  A lane with no `--model`, no `--effort` and no `defaults` entry in
222
226
  `bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
223
227
  CLI configured months ago at a low reasoning effort keeps auditing at that
224
- effort while your routing docs describe an adversarial pass.
228
+ effort while your routing docs describe a second-opinion pass.
225
229
 
226
230
  ```bash
227
231
  node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
@@ -244,7 +248,7 @@ Install for the tools you have, then let the generated `ROUTING.md` decide the t
244
248
 
245
249
  ### How do I route tasks to cheaper models?
246
250
 
247
- The rules route by role, complexity and risk (see [Routing by role, complexity and risk](#routing-by-role-complexity-and-risk)). Role picks the agent, complexity moves the effort, risk moves the tier. A task a cheap tier finishes correctly never gets a frontier token.
251
+ The rules route by role, complexity and stakes (see [Routing by role, complexity and stakes](#routing-by-role-complexity-and-stakes)). Role picks the agent, complexity moves the effort, stakes move the tier. A task a cheap tier finishes correctly never gets a frontier token.
248
252
 
249
253
  ### Is this an LLM router or an AI gateway?
250
254
 
@@ -272,7 +276,7 @@ Add an AI to `src/catalog.js` and every prompt, table, config and doc picks it u
272
276
  ## Credits
273
277
 
274
278
  - [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for
275
- separating role, complexity and risk instead of compressing them into one
279
+ separating role, complexity and stakes instead of compressing them into one
276
280
  scale, for recording the model and effort a lane was actually asked for, and
277
281
  for verifying findings before they trigger repairs. All three shipped in
278
282
  0.1.14.
package/docs/README.md CHANGED
@@ -7,7 +7,7 @@ The three parts, as reading. The installer writes the working files; these expla
7
7
  | 1 Beginner | you use one LLM or one agent and want it to route well | [part-1-beginner.md](part-1-beginner.md) |
8
8
  | 2 Intermediate | you have several AIs and want to call them through their CLIs from one orchestrator | [part-2-intermediate.md](part-2-intermediate.md) |
9
9
  | 3 Advanced | you want the whole thing running unattended on a virtual machine | [part-3-advanced.md](part-3-advanced.md) |
10
- | Audit brief | the threat model and the two adversarial audit rounds this shipped with | [audit-brief.md](audit-brief.md) |
10
+ | Audit brief | the security notes and the two second-opinion audit rounds this shipped with | [audit-brief.md](audit-brief.md) |
11
11
  | Catalog | what each AI and companion tool in the installer is for, how it installs, how it signs in | [catalog.md](catalog.md) |
12
12
 
13
13
  Each part ends with "what the installer gives you at this level" so the doc and the files agree.
package/docs/catalog.md CHANGED
@@ -24,7 +24,7 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
24
24
  ### `codex` · Codex CLI (OpenAI, ChatGPT plan)
25
25
 
26
26
  - **Kind:** agent-cli · **Access:** subscription · **Lane:** A · **Level:** 1+
27
- - **Wins at:** second coder and adversarial auditor (a different model family reading your diff)
27
+ - **Wins at:** second coder and second-opinion reviewer (a different model family reading your diff)
28
28
  - **Install:** `npm install -g @openai/codex@0.153.4`
29
29
  - **Sign in:** `codex login` (add `--device-auth` on a machine with no browser)
30
30
  - **Reads rules from:** `AGENTS.md`
@@ -14,7 +14,7 @@ If your agent exposes model choice (Claude Code, Codex, Antigravity), map the ti
14
14
 
15
15
  Three cost levers, always together: tier (price per token), token discipline (how many tokens: read only what you will touch, never re-read, deliverables not narration), effort (how hard each call thinks).
16
16
 
17
- And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **Risk** moves the tier and who reads the result: security, privacy, data loss and irreversible changes are the four worth naming, because none of their failures can be fixed by editing the code afterwards. A one-line change to an auth check is simple and high-risk at once, and it is the risk that decides.
17
+ And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **Stakes** move the tier and who reads the result: security, privacy, data loss and irreversible changes are the four worth naming, because none of their failures can be fixed by editing the code afterwards. A one-line change to an auth check is simple and high-stakes at once, and it is the stakes that decide.
18
18
 
19
19
  Robustness first, cost second. You split tiers because the split produces better work.
20
20
 
@@ -32,9 +32,9 @@ Modifiers: plan big, execute small · never silently retry a failed attempt at t
32
32
 
33
33
  > A gate you cannot fail is not a gate.
34
34
 
35
- "Does this look good?" passes every time. "Name the single biggest risk and the flaw in the request as filed" can come back empty, which is how you know it worked.
35
+ "Does this look good?" passes every time. "Name what is most likely to go wrong, and what the request as filed missed" can come back empty, which is how you know it worked.
36
36
 
37
- Every build gets two checkpoints. **Before writing:** you map what it touches and what could break, then ask the deep tier on the finished map for a named risk and a named flaw. **After it is green:** a fresh context attacks it (bad input, failing dependency, drift from the plan), allowed to answer CLEAN, every finding reproduced before it reaches you. Then the ship, with a rollback named and an explicit yes. Then the loud negative: re-check the old name everywhere and expect zero.
37
+ Every build gets two checkpoints. **Before writing:** you map what it touches and what could break, then ask the deep tier on the finished map for one named weak spot and one gap in the request. **After it is green:** a fresh context challenges it (bad input, failing dependency, drift from the plan), allowed to answer CLEAN, every finding reproduced before it reaches you. Then the ship, with a rollback named and an explicit yes. Then the loud negative: re-check the old name everywhere and expect zero.
38
38
 
39
39
  ## 4. Every hand-off carries a brief
40
40
 
@@ -46,7 +46,7 @@ After anything comprehensive, a fresh turn that hunts for what is **missing**, n
46
46
 
47
47
  ## 6. Deep research, single agent
48
48
 
49
- Plan the sub-questions as their own turn and inspect them before spending anything. Sweep. Then a fresh adversarial turn told to attack the premise. Plant one deliberately wrong figure and see whether it corrects it. Mark every claim CONFIRMED / DISAGREEMENT / REPORTED / UNVERIFIED. Agreement is weak evidence; disagreement is the signal.
49
+ Plan the sub-questions as their own turn and inspect them before spending anything. Sweep. Then a fresh second-opinion turn told to question the premise. Plant one deliberately wrong figure and see whether it corrects it. Mark every claim CONFIRMED / DISAGREEMENT / REPORTED / UNVERIFIED. Agreement is weak evidence; disagreement is the signal.
50
50
 
51
51
  ## 7. Numbers and logic are computed, never guessed
52
52
 
@@ -13,7 +13,7 @@ Rule: never spend a frontier token on a task a cheap tier finishes correctly. Es
13
13
  | Lane | Wins at |
14
14
  |---|---|
15
15
  | the orchestrator (Claude Code, or whichever you chose) | routes, maps, builds, verifies, records; drives the others as CLIs |
16
- | Codex | second coder and adversarial auditor: a different model family reading your diff |
16
+ | Codex | second coder and second-opinion reviewer: a different model family reading your diff |
17
17
  | Antigravity `agy` | deep research sweeps; concurrent fan-out (its subagent call takes an array) |
18
18
  | Grok CLI | X and live web reads at $0 (the same search on the API bills per call) |
19
19
  | Hermes | the free tier: rough drafts, first-pass summaries, divergent reads, cron jobs |
@@ -28,7 +28,7 @@ Every agent CLI can report success and deliver nothing. `bin/cli-run.mjs` builds
28
28
 
29
29
  Every call goes through it. "This lane is flaky" becomes a query over its log instead of an argument. `node bin/cli-run.mjs --doctor` is the first thing to run after install: enabled lanes, binaries on PATH, the route each lane is pinned to, and with `--run` a one-word canary per lane.
30
30
 
31
- There is a second thing a lane can be quietly wrong about. Left unpinned, it runs on **its own config file**, which the runner cannot see: a CLI set up months ago at a low reasoning effort keeps auditing at that effort while your routing docs describe an adversarial pass, and no error is ever raised. `--model` and `--effort` pin it per call, `defaults` in `bin/lanes.json` pins it per lane, and every run records the value requested and where it came from (`flag`, `lanes.json`, `lane_default`). The log claims no actual: grok reports a model id in its output, the other four lanes report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string the durable log never holds.
31
+ There is a second thing a lane can be quietly wrong about. Left unpinned, it runs on **its own config file**, which the runner cannot see: a CLI set up months ago at a low reasoning effort keeps auditing at that effort while your routing docs describe a second-opinion pass, and no error is ever raised. `--model` and `--effort` pin it per call, `defaults` in `bin/lanes.json` pins it per lane, and every run records the value requested and where it came from (`flag`, `lanes.json`, `lane_default`). The log claims no actual: grok reports a model id in its output, the other four lanes report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string the durable log never holds.
32
32
 
33
33
  ## 4. Every delegation carries a task bundle, on both surfaces
34
34
 
@@ -36,7 +36,7 @@ Subagents and CLI lanes are close to the same problem: something that may hold n
36
36
 
37
37
  ## 5. Research: three engines, one triager
38
38
 
39
- Fan the same plan to three model families (web sweep, adversarial read, live data), each as one `cli-run` call. The orchestrator opens the primary sources itself, marks every claim, and writes the only durable record. Expect one engine to return confident unsourced numerics; downgrade it. Weight the engines that report their own gaps. Count dispositions, not briefs.
39
+ Fan the same plan to three model families (web sweep, second-opinion read, live data), each as one `cli-run` call. The orchestrator opens the primary sources itself, marks every claim, and writes the only durable record. Expect one engine to return confident unsourced numerics; downgrade it. Weight the engines that report their own gaps. Count dispositions, not briefs.
40
40
 
41
41
  ## 5a. A finding is a claim, not a fact
42
42
 
@@ -48,7 +48,7 @@ The second pass is now a different model reading the same artifact, in read-only
48
48
 
49
49
  ## 7. The build protocol, bound to lanes
50
50
 
51
- Stage 1 Map: the orchestrator sweeps; CLI lanes critique the map at $0. Stage 2: deep tier, a named risk and a named flaw. Stage 4: scanners on the added lines, fail closed. Stage 5: security-shaped diff → the second coder in read-only audit mode; architecture-shaped → deep tier reviewing build against plan; never both. Two deep checkpoints per build; CLI lanes are uncapped.
51
+ Stage 1 Map: the orchestrator sweeps; CLI lanes critique the map at $0. Stage 2: deep tier, one named weak spot and one gap in the request. Stage 4: scanners on the added lines, refuses by default. Stage 5: security-shaped diff → the second coder in read-only audit mode; architecture-shaped → deep tier reviewing build against plan; never both. Two deep checkpoints per build; CLI lanes are uncapped.
52
52
 
53
53
  ## 8. Privacy gate
54
54
 
package/llms.txt CHANGED
@@ -24,4 +24,4 @@ Levels: 1 beginner (one agent or chat app), 2 intermediate (several agent CLIs,
24
24
  ## Optional
25
25
 
26
26
  - [Security policy](https://github.com/aunysillyme/model-orchestrator/blob/main/SECURITY.md)
27
- - [Audit brief](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/audit-brief.md): the threat model and what has already been attacked
27
+ - [Audit brief](https://github.com/aunysillyme/model-orchestrator/blob/main/docs/audit-brief.md): the security notes and what has already been security-reviewed
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.16",
3
+ "version": "0.1.17",
4
4
  "description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
5
5
  "type": "module",
6
6
  "bin": {
package/src/catalog.js CHANGED
@@ -87,7 +87,7 @@ export const AIS = [
87
87
  bin: 'codex',
88
88
  access: 'subscription',
89
89
  lane: 'A',
90
- role: 'second coder and adversarial auditor (a different model family reading your diff)',
90
+ role: 'second coder and second-opinion reviewer (a different model family reading your diff)',
91
91
  minLevel: 1,
92
92
  install: { npm: '@openai/codex', pin: '0.153.4' },
93
93
  builtAgainst: '0.153.4',
package/src/install.js CHANGED
@@ -155,21 +155,21 @@ export function laneVars(selected) {
155
155
  if (has('hermes')) step0.push(`${cr('hermes')} (the free tier) for rough drafts and divergent reads`);
156
156
  if (has('qwen')) step0.push(`${cr('qwen')} (the cheapest metered lane) for structured bulk, never for anything citing a line, number or source`);
157
157
  if (has('grok')) step0.push(`${cr('grok')} for X and live web reads at $0`);
158
- if (has('codex')) step0.push(`${cr('codex --audit')} for an adversarial read by a second model family`);
158
+ if (has('codex')) step0.push(`${cr('codex --audit')} for a second-opinion read by a second model family`);
159
159
  if (has('agy')) step0.push(`${cr('agy')} for research sweeps and concurrent fan-out`);
160
160
  const stage1 = [];
161
- if (has('codex')) stage1.push(`${cr('codex')} for adversarial critique of the map`);
161
+ if (has('codex')) stage1.push(`${cr('codex')} for a second-opinion critique of the map`);
162
162
  if (has('grok')) stage1.push(`${cr('grok')} to verify current API behaviour instead of trusting recall`);
163
163
  if (has('hermes')) stage1.push(`${cr('hermes')} for a divergent read`);
164
164
  if (has('agy')) stage1.push(`${cr('agy')} for a wide sweep of prior art`);
165
165
  const examples = [];
166
166
  examples.push(has('grok') ? `| "What is trending on X today" | ${cr('grok')} |` : '| "What is trending on X today" | live-researcher (standard tier with web tools) |');
167
- examples.push(has('codex') ? `| "Audit this auth diff" | ${cr('codex --audit')} |` : '| "Audit this auth diff" | code-reviewer at deep tier, in a fresh context told to attack |');
167
+ examples.push(has('codex') ? `| "Audit this auth diff" | ${cr('codex --audit')} |` : '| "Audit this auth diff" | code-reviewer at deep tier, in a fresh context told to challenge |');
168
168
  examples.push(has('qwen') ? `| "Classify these 200 items" | bulk-worker, or ${cr('qwen')} if the items may leave the machine |` : '| "Classify these 200 items" | bulk-worker |');
169
- examples.push(enabled.length >= 2 ? '| "Research this topic properly" | several engines in parallel, see `RESEARCH_TRIAGE.md` |' : '| "Research this topic properly" | deep tier plans, standard tier sweeps, a fresh context attacks; see `RESEARCH_TRIAGE.md` |');
169
+ examples.push(enabled.length >= 2 ? '| "Research this topic properly" | several engines in parallel, see `RESEARCH_TRIAGE.md` |' : '| "Research this topic properly" | deep tier plans, standard tier sweeps, a fresh context challenges; see `RESEARCH_TRIAGE.md` |');
170
170
  const roles = [];
171
171
  if (has('agy')) roles.push('| Web sweep | `cli-run agy` | widest landscape pass |');
172
- if (has('codex')) roles.push('| Adversarial read | `cli-run codex --audit` | attack the premise, hunt for what the others would get wrong |');
172
+ if (has('codex')) roles.push('| Second-opinion read | `cli-run codex --audit` | question the premise, hunt for what the others would get wrong |');
173
173
  if (has('grok')) roles.push('| Live data | `cli-run grok` | dated primary sources, real-time reads |');
174
174
  if (has('hermes')) roles.push('| Cheap divergent read | `cli-run hermes` | another opinion at $0 |');
175
175
  if (has('qwen')) roles.push('| Structured extraction | `cli-run qwen` | pull the facts into a table; never trust its citations without a check |');
@@ -183,12 +183,12 @@ export function laneVars(selected) {
183
183
  return {
184
184
  LANE_STEP0: step0.length ? step0.map((l) => ' - ' + l).join('\n') : ' - none selected yet: every task stays on your primary agent\'s tiers until you add a lane (re-run the installer with more AIs)',
185
185
  STAGE1_LANES: stage1.length ? '; ' + stage1.join(', ') : '',
186
- ATTACK_LANE: has('codex') ? '`cli-run codex --audit` (a second model family in a read-only sandbox)' : 'code-reviewer at deep tier, in a fresh context told to attack and allowed to answer CLEAN',
186
+ ATTACK_LANE: has('codex') ? '`cli-run codex --audit` (a second model family in a read-only sandbox)' : 'code-reviewer at deep tier, in a fresh context told to challenge and allowed to answer CLEAN',
187
187
  LIVE_LANE: has('grok') ? '`cli-run grok` first ($0), then' : '',
188
188
  BULK_LANE: has('qwen') ? ', or `cli-run qwen` if the data may leave your machine' : has('hermes') ? ', or `cli-run hermes` for a free rough pass' : '',
189
189
  LANE_EXAMPLES: examples.join('\n'),
190
190
  RESEARCH_ROLES: roles.join('\n'),
191
- RESEARCH_RUN: run.length ? run.join('\n') : '# no cli-run lane selected: run the sweep on your primary agent, then a fresh adversarial turn (protocols/deep-research.md, level 1 shape)',
191
+ RESEARCH_RUN: run.length ? run.join('\n') : '# no cli-run lane selected: run the sweep on your primary agent, then a fresh second-opinion turn (protocols/deep-research.md, level 1 shape)',
192
192
  RESEARCH_ENGINES: String(run.length)
193
193
  };
194
194
  }
@@ -2,4 +2,4 @@
2
2
 
3
3
  Antigravity CLI custom agents, one per tier plus three checks (`finding-verifier`, `done-verifier`, `reader`), in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
4
4
 
5
- `commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; `auto` keeps deletes and other high-risk commands gated) and `off` for the read-only agents: `code-reviewer`, `finding-verifier`, `live-researcher`, `done-verifier`, `reader`. `model` is a tier: `pro` for deep-planner, `flash` for the rest. `done-verifier` and `reader` never write and, with `commandExecutionPolicy: off`, cannot execute any command at all here, mutating or not: unlike its claude-code counterpart, which does carry an unrestricted `Bash` and stays read-only by its prompt rather than by the tool grant, agy's `done-verifier` is mechanically blocked from shelling out and probes artifacts through whatever read or fetch capability it has instead. Neither is `bulk-worker`, which classifies, tags and transforms items and does write.
5
+ `commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; deletes and other destructive commands still ask before running) and `off` for the read-only agents: `code-reviewer`, `finding-verifier`, `live-researcher`, `done-verifier`, `reader`. `model` is a tier: `pro` for deep-planner, `flash` for the rest. `done-verifier` and `reader` never write and, with `commandExecutionPolicy: off`, cannot execute any command at all here, mutating or not: unlike its claude-code counterpart, which does carry an unrestricted `Bash` and stays read-only by its prompt rather than by the tool grant, agy's `done-verifier` is mechanically blocked from shelling out and probes artifacts through whatever read or fetch capability it has instead. Neither is `bulk-worker`, which classifies, tags and transforms items and does write.
@@ -4,7 +4,7 @@ description: Well-specified execution of a bounded sub-part of a build.
4
4
  model: flash
5
5
  subagent: true
6
6
  mainAgent: true
7
- commandExecutionPolicy: auto # standard build/test commands run unattended; high-risk commands stay gated
7
+ commandExecutionPolicy: auto # standard build/test commands run unattended; destructive commands, like deletes, still ask before running
8
8
  ---
9
9
 
10
10
  # builder
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: finding-verifier
3
- description: Adversarial verification of review findings; tries to disprove each one and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE. Read-only, never repairs.
3
+ description: Second-opinion verification of review findings; tries to disprove each one and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE. Read-only, never repairs.
4
4
  model: flash
5
5
  subagent: true
6
6
  mainAgent: true
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: finding-verifier
3
- description: Adversarial verification of review findings. Use after a review or audit returns findings and before any of them trigger a repair. No file-editing tools; Bash is for read-only checks, bound by the prompt below, not by the tool grant. Tries to DISPROVE each finding and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Do not use to find new problems, and do not use to fix anything.
3
+ description: Second-opinion verification of review findings. Use after a review or audit returns findings and before any of them trigger a repair. No file-editing tools; Bash is for read-only checks, bound by the prompt below, not by the tool grant. Tries to DISPROVE each finding and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Do not use to find new problems, and do not use to fix anything.
4
4
  tools: Read, Glob, Grep, Bash
5
5
  model: sonnet
6
6
  effort: high
@@ -4,8 +4,8 @@
4
4
 
5
5
  ```
6
6
  You follow a model-orchestrator workflow inside this chat. Tiers describe effort, not automatic model switching or cost savings.
7
- Route first: bulk/formatting -> fast; live data -> standard with tools; review -> standard, read-only; ambiguous/high-risk -> deep; otherwise standard. State the tier. Escalate on failure instead of silently retrying.
8
- For builds: map affected parts; identify the biggest risk and any flaw in the request (none is valid with reasons); build and verify; use a fresh turn to challenge the result. Before irreversible actions, explain rollback and ask for approval.
7
+ Route first: bulk/formatting -> fast; live data -> standard with tools; review -> standard, read-only; ambiguous/high-stakes -> deep; otherwise standard. State the tier. Escalate on failure instead of silently retrying.
8
+ For builds: map affected parts; identify what is most likely to go wrong and any gap in the request (none is valid with reasons); build and verify; use a fresh turn to challenge the result. Before irreversible actions, explain rollback and ask for approval.
9
9
  For hand-offs: include purpose, scope, allowed and denied actions, required output, and stopping conditions. A fresh context has none of these instructions.
10
10
  After comprehensive work, check for omissions. Compute consequential numbers and comparisons with a tool; report what was checked and what remains unverified.
11
11
  Before durable writes, search existing records, update their index, use one writer, and label inferences.
@@ -17,7 +17,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
17
17
 
18
18
  A subagent starts with your CLAUDE.md and tool definitions already loaded, so it has a fixed start-up cost before it does anything. Measure yours once: spawn a subagent with a one-line task and read its token count. Work smaller than that stays inline.
19
19
 
20
- Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one adversarial pass, an explicit human yes before anything irreversible, then the loud negative.
20
+ Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: two deep-tier checkpoints, a mechanical scan, one challenge pass, an explicit human yes before anything irreversible, then the loud negative.
21
21
 
22
22
  Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A Claude Code subagent loads this CLAUDE.md hierarchy, so it holds the standing rules already, just not this task's scope; a second CLI or a fresh chat window may hold none of them. Absence is denial either way.
23
23
 
@@ -9,7 +9,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
9
9
 
10
10
  Route by capability tier, first match wins: bulk and mechanical -> fast tier · needs live data -> standard tier with tools · review without changing -> standard, read-only · ambiguous or expensive to get wrong -> deep tier, then hand the plan down · everything else -> build it directly at standard tier.
11
11
 
12
- Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: map the blast radius yourself, ask the deep tier for a named risk and a named flaw, build green, scan the added lines, one adversarial pass with every finding reproduced, an explicit human yes before anything irreversible, then re-grep the old identifier and expect zero.
12
+ Every build runs `{{RULES_PATH}}/protocols/build-protocol.md`: map everything it touches yourself, ask the deep tier for one named weak spot and one gap in the request, build green, scan the added lines, one challenge pass with every finding reproduced, an explicit human yes before anything irreversible, then re-grep the old identifier and expect zero.
13
13
 
14
14
  Every delegation carries an `{{RULES_PATH}}/TASK_BUNDLE.md` brief. A fresh context holds none of these rules; absence is denial.
15
15
 
@@ -10,7 +10,7 @@
10
10
  // This script always exits 0, never blocks on stdin past a short bound,
11
11
  // reads at most 64 KB of the rules file through a fixed-size buffer (never
12
12
  // a full read of an arbitrarily large or non-regular file), and never
13
- // executes anything it reads. See docs/audit-brief.md for the threat model.
13
+ // executes anything it reads. See docs/audit-brief.md for the security notes.
14
14
  import { statSync, openSync, readSync, closeSync, realpathSync } from 'node:fs';
15
15
  import { join, isAbsolute } from 'node:path';
16
16
 
@@ -29,8 +29,8 @@ Modifiers:
29
29
 
30
30
  ## The two checkpoints (every build)
31
31
 
32
- - **Checkpoint 1, before writing anything.** You map the blast radius yourself (files, systems, docs, tickets). Then ask the deep tier, on the finished map: *is this the simplest way, what is the single biggest risk, where is the request as filed wrong?* It must return a named risk and a named flaw. Approval alone is not an answer.
33
- - **Checkpoint 2, after the build is green.** Security-shaped diffs get an adversarial read (in a fresh context, told to attack, allowed to answer CLEAN). Architecture-shaped diffs get the deep tier reviewing build against plan. Never both on one diff. Every finding reproduced before it reaches a human.
32
+ - **Checkpoint 1, before writing anything.** You map everything it touches yourself (files, systems, docs, tickets). Then ask the deep tier, on the finished map: *is this the simplest way, what is most likely to go wrong, what did the request miss?* It must return one named weak spot and one gap in the request. Approval alone is not an answer.
33
+ - **Checkpoint 2, after the build is green.** Security-shaped diffs get a second-opinion read (in a fresh context, told to challenge, allowed to answer CLEAN). Architecture-shaped diffs get the deep tier reviewing build against plan. Never both on one diff. Every finding reproduced before it reaches a human.
34
34
 
35
35
  Cap: two deep-tier consults per build. The full procedure is `protocols/build-protocol.md`.
36
36
 
@@ -47,7 +47,7 @@ Level 2 adds `ROUTING.md`, `TIERS.md`, `DELEGATION_MATRIX.md`, `RESEARCH_TRIAGE.
47
47
  ## The three rules that carry everything
48
48
 
49
49
  1. **Route by capability tier, not by model name.** deep = ambiguous or expensive to get wrong · standard = well-specified execution and review · fast = bulk and mechanical. Default down, escalate on evidence.
50
- 2. **A gate you cannot fail is not a gate.** "Does it look good?" passes every time. "Name the single biggest risk and the flaw in the request" can come back empty, which is how you know it worked.
50
+ 2. **A gate you cannot fail is not a gate.** "Does it look good?" passes every time. "Name what is most likely to go wrong, and what the request missed" can come back empty, which is how you know it worked.
51
51
  3. **Exit 0 is not a deliverable.** Any tool, CLI or subagent can report success and hand back nothing. Check for the artifact, not the status line.
52
52
 
53
53
  ## Where things went
@@ -2,7 +2,7 @@
2
2
 
3
3
  **Three phases, eight stages, and every gate is a question that can be answered wrong.**
4
4
 
5
- Fires on any task that builds, codes, implements, migrates or deploys. Rough test: if it would earn an adversarial audit or a tracker issue, it runs this.
5
+ Fires on any task that builds, codes, implements, migrates or deploys. Rough test: if it would earn a second-opinion audit or a tracker issue, it runs this.
6
6
 
7
7
  > **The one rule underneath:** a gate you cannot fail is not a gate. If a stage's exit reads like "confirm it looks good", it is written wrong and it will pass every time, including the times it should not.
8
8
 
@@ -14,7 +14,7 @@ Three corollaries:
14
14
  | Phase | Master question | Stages |
15
15
  |---|---|---|
16
16
  | 1 Pre-build | What exactly are we building, what do we need first, and what does this touch or break? | 0 Route · 1 Map · 2 Judge |
17
- | 2 Build | Is it secure, built on current code, and correct without hidden flaws? | 3 Build · 4 Scan · 5 Attack · 5b Ship gate |
17
+ | 2 Build | Is it secure, built on current code, and correct without hidden flaws? | 3 Build · 4 Scan · 5 Challenge · 5b Ship gate |
18
18
  | 3 Post-build | Did it land everywhere, is it proven against the real thing, and is it recorded? | 6 Verify · 7 Record |
19
19
 
20
20
  The two seams are the point. Pre-build to Build: nothing is written yet, changing your mind costs a conversation. Build to Post-build: the ship, the only irreversible step, the only one that needs an explicit human yes.
@@ -38,9 +38,9 @@ Four bounded questions, not four exhaustive scans. **The builder maps; the judgm
38
38
  ### Stage 2 · Judge (Checkpoint 1)
39
39
  Ask the judgment tier, on the finished map:
40
40
  1. Is this the simplest way to build it, or are we overcomplicating?
41
- 2. What is the single biggest risk, and where is the request as filed wrong?
41
+ 2. What is most likely to go wrong, and what did the request miss?
42
42
 
43
- **Gate:** a **named risk** and a **named flaw in the request**. Approval alone is not an exit; an advisor asked only to approve will approve. If a consult comes back mostly restating the map, the brief asked it to retrieve when it should have asked it to decide.
43
+ **Gate:** **one named weak spot** and **one gap in the request**. Approval alone is not an exit; an advisor asked only to approve will approve. If a consult comes back mostly restating the map, the brief asked it to retrieve when it should have asked it to decide.
44
44
 
45
45
  ## Phase 2 · Build
46
46
 
@@ -56,15 +56,15 @@ Ask the judgment tier, on the finished map:
56
56
  1. Any secret, key or token in the new code?
57
57
  2. Any vulnerability or vulnerable dependency in the lines we added?
58
58
 
59
- Secret detection, static analysis and dependency scanning, filtered to lines this diff added. Fail closed: a missing or erroring scanner exits non-zero, never a silent green.
59
+ Secret detection, static analysis and dependency scanning, filtered to lines this diff added. Refuses by default: a missing or erroring scanner exits non-zero, never a silent green.
60
60
 
61
61
  **Gate:** zero flags on added lines. Pre-existing flags are reported, never inherited as blockers, and never waved through unread. A scanner finding is a claim; read the code before calling it anything.
62
62
 
63
- ### Stage 5 · Attack (Checkpoint 2, one pass, never two)
63
+ ### Stage 5 · Challenge (Checkpoint 2, one pass, never two)
64
64
  1. Can bad input or a bad actor break it, and what happens when a dependency fails?
65
65
  2. Did the build stick to the approved plan, or did unintended changes sneak in?
66
66
 
67
- Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to an adversarial auditor, ideally a **different model family**. Architecture-shaped diffs go to the judgment tier reviewing build against plan. Never both on one diff.
67
+ Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to a second-opinion reviewer, ideally a **different model family**. Architecture-shaped diffs go to the judgment tier reviewing build against plan. Never both on one diff.
68
68
 
69
69
  Findings do not go straight to a repair. Hand them to the **finding-verifier**, whose job is to DISPROVE each one: read the cited line, state what would trigger it, then hunt for the guard, caller or test that makes it impossible. It returns one of three verdicts per finding, and rounding between them is the failure mode to watch for.
70
70
 
@@ -106,7 +106,7 @@ Use a different model family from the one that produced the finding where you ha
106
106
  |---|---|---|
107
107
  {{ROLES_BUILDER_ROW}}
108
108
  | Judgment tier | Stage 2 and the architectural arm of Stage 5. Argues with a finished map | Perform the retrieval |
109
- | Adversarial auditor | The security arm of Stage 5. Attacks the diff | Fix anything |
109
+ | Second-opinion reviewer | The security arm of Stage 5. Reviews the diff | Fix anything |
110
110
  | Mechanical gates | Stage 4 and any always-on guard | Be overridden without reading |
111
111
  | Cheap workers | Bounded sub-parts: bulk passes, wide searches, long loops | Own a stage |
112
112
  | Human | Stage 5b, and any irreversible or architectural call | Be the first line of review |
@@ -119,9 +119,9 @@ Use a different model family from the one that produced the finding where you ha
119
119
  PRE-BUILD
120
120
  [ ] 0 Inputs and access verified by live probe, not assumed
121
121
  [ ] 0 Confirmed this is a build and not a quick fix
122
- [ ] 1 Blast radius written: files, systems, issues
122
+ [ ] 1 Everything it touches written: files, systems, issues
123
123
  [ ] 1 Asked what could break, and whether this already exists
124
- [ ] 2 Judgment tier named a risk AND a flaw in the request
124
+ [ ] 2 Judgment tier named a weak spot AND a gap in the request
125
125
 
126
126
  BUILD
127
127
  [ ] 3 Repo clean, on a branch, base ref recorded
@@ -31,11 +31,11 @@ Agreement is weak evidence. Disagreement is the signal.
31
31
 
32
32
  ## Level 1: one agent
33
33
 
34
- You still get the shape. Run PLAN as its own turn and inspect it before spending anything. Run the sweep. Then run a **fresh-context adversarial turn** with a brief that says "attack the premise; list what this report would get wrong if its sources were stale". Plant one deliberately wrong figure in the brief and see whether it corrects it: if it does not, its confirmations are worth less than they look. Mark every claim.
34
+ You still get the shape. Run PLAN as its own turn and inspect it before spending anything. Run the sweep. Then run a **fresh-context second-opinion turn** with a brief that says "question the premise; list what this report would get wrong if its sources were stale". Plant one deliberately wrong figure in the brief and see whether it corrects it: if it does not, its confirmations are worth less than they look. Mark every claim.
35
35
 
36
36
  ## Level 2 and up: three engines, one triager
37
37
 
38
- Fan out the same PLAN to three different model families through their CLIs (a web-sweep lane, an adversarial-read lane, a live-data lane). Run them through `cli-run` so a run that produced nothing is caught as `rc=10` rather than read as an empty finding. The orchestrator triages: it opens the primary sources itself, marks each claim, and writes the brief. Only the orchestrator writes the durable record; every other engine proposes.
38
+ Fan out the same PLAN to three different model families through their CLIs (a web-sweep lane, a second-opinion-read lane, a live-data lane). Run them through `cli-run` so a run that produced nothing is caught as `rc=10` rather than read as an empty finding. The orchestrator triages: it opens the primary sources itself, marks each claim, and writes the brief. Only the orchestrator writes the durable record; every other engine proposes.
39
39
 
40
40
  Known failure shape: one engine will return confident unsourced numerics and claim full coverage. Downgrade those to hypothesis. The engines that report their own gaps honestly are the ones to weight.
41
41
 
@@ -14,7 +14,7 @@ Verification asks "is what I did correct?". Gap analysis asks "what did I not do
14
14
  ## Who runs it
15
15
 
16
16
  - **Level 1 (one agent):** the same agent, in a fresh turn, with a brief that says "you are looking for what is missing; do not re-verify what is present". Fresh context matters more than a different model.
17
- - **Level 2 and up:** a **different model family** reading the same artifact. Disagreement between two families is the cheapest available signal that something is soft. The adversarial coder lane (a second-opinion CLI in read-only mode) is the natural fit.
17
+ - **Level 2 and up:** a **different model family** reading the same artifact. Disagreement between two families is the cheapest available signal that something is soft. The second-opinion coder lane (read-only mode) is the natural fit.
18
18
  - **Level 3:** make it recurring. A weekly audit job enumerates live state (lanes, jobs, services, model lists), diffs it against the plan, and files a report. It catches the dead lane and the silently renamed model nobody noticed.
19
19
 
20
20
  ## The second half: analyze, compare, suggest
@@ -1,10 +1,10 @@
1
1
  # Propagate: change completeness
2
2
 
3
- **A rename is a refactor, not a single-file edit.** Any change to a name, term, path, slug, schema field, routing rule or shared convention has a blast radius, and the goal is zero silent strays.
3
+ **A rename is a refactor, not a single-file edit.** Any change to a name, term, path, slug, schema field, routing rule or shared convention reaches everything that uses it, and the goal is zero silent strays.
4
4
 
5
5
  This is retrieval work. It stays with the orchestrator (or a cheap worker for the grep sweep). It never goes to the deep tier: a judgment model re-deriving a file list is the most expensive routing mistake there is.
6
6
 
7
- ## 1. Map the blast radius (before editing anything)
7
+ ## 1. Map everything it touches (before editing anything)
8
8
 
9
9
  - **Docs and notes:** backlinks to the thing being renamed; literal search for the old term and its link forms. With obsidian-tc: `get_backlinks`, `search_text`, then `find_unresolved_links` after the change (`protocols/memory-and-record.md`).
10
10
  - **Memory / instructions:** grep every instructions file your agents read (`CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, `QWEN.md`, custom instructions) and any memory store.
@@ -68,7 +68,7 @@ qwen is the lane whose own success flags lie: an upstream 400 comes back as exit
68
68
 
69
69
  ## The route: which model, and how hard it thinks
70
70
 
71
- A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe an adversarial pass, and nothing anywhere says so.
71
+ A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe a second-opinion pass, and nothing anywhere says so.
72
72
 
73
73
  Pin it per call, or per lane:
74
74
 
@@ -114,9 +114,9 @@ The log records what was **requested**, on every record including a run refused
114
114
 
115
115
  That is each vendor's documented headless shape (`-p`, `exec`). Two consequences: argv is visible to other processes on the machine, so a prompt is never the place for a key; and argv is bounded by the OS (`ARG_MAX`), so a very large brief should be referenced by path inside the prompt rather than pasted whole.
116
116
 
117
- ## lanes.json fails closed
117
+ ## lanes.json refuses by default
118
118
 
119
- Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all fail the whole file closed rather than being skipped quietly.
119
+ Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all make the whole file refuse by default rather than being skipped quietly.
120
120
 
121
121
  ## A killed lane is not a deliverable
122
122
 
@@ -15,7 +15,7 @@ Generated {{DATE}} from the AIs you said you have: `{{AI_IDS}}`.
15
15
  | Many independent items each needing its own agent turn | a concurrent fan-out lane | one call, N children, on a subscription |
16
16
  | Live web or social reads | the live-data CLI | subscription-covered; the same search on the API bills per call |
17
17
  | Code review, no changes | standard tier, or the second-coder CLI | a different model family catches what one misses |
18
- | Adversarial audit of a security-shaped diff | the second-coder CLI in read-only audit mode | Claude writes, a second family attacks, the orchestrator reproduces |
18
+ | Second-opinion audit of a security-shaped diff | the second-coder CLI in read-only audit mode | Claude writes, a second family challenges, the orchestrator reproduces |
19
19
  | Deep architecture / planning | deep tier | expensive to get wrong |
20
20
  | Well-specified execution | the orchestrator | execution does not need the top tier |
21
21
  | Long-document analysis | the largest-context lane, or caching on the primary | window size vs re-query cost |
@@ -31,10 +31,10 @@ Rule of thumb: never spend a frontier token on a task a cheap tier finishes corr
31
31
  |---|---|
32
32
  | 0 Route | live probe for access; `cli-run` lanes are $0 and uncapped |
33
33
  | 1 Map | the orchestrator sweeps{{STAGE1_LANES}} |
34
- | 2 Judge | deep tier, on the finished map: a named risk and a named flaw |
34
+ | 2 Judge | deep tier, on the finished map: one named weak spot and one gap in the request |
35
35
  | 3 Build | the orchestrator, against the installed dependency's source |
36
- | 4 Scan | secret + static + dependency scanners, diff-scoped, fail closed |
37
- | 5 Attack | security-shaped diff → {{ATTACK_LANE}}. Architecture-shaped → deep tier, build against plan. Never both |
36
+ | 4 Scan | secret + static + dependency scanners, diff-scoped, refuses by default |
37
+ | 5 Challenge | security-shaped diff → {{ATTACK_LANE}}. Architecture-shaped → deep tier, build against plan. Never both |
38
38
  | 5a Verify findings | finding-verifier, a different model family where you have one: CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a repair |
39
39
  | 5b Ship | rollback id recorded, explicit human yes |
40
40
  | 6 Verify | real test, negative test seen red, old identifier re-grepped to zero |
@@ -58,7 +58,7 @@ One writer per run; every other lane proposes. Search before writing, index in t
58
58
  - **Long context:** mechanical digestion → fast tier in chunks; judgment over a long input → standard tier.
59
59
  - **Token discipline on every delegation:** pass only the context the delegate needs, never the conversation.
60
60
  - **Effort per agent:** deep xhigh, review, verification and build high, live research medium, bulk low.
61
- - **Three inputs, not one:** role picks the agent, complexity moves the effort, risk moves the tier and who reads it. A one-line auth change is simple and high-risk at once, and the risk decides. See `TIERS.md`.
61
+ - **Three inputs, not one:** role picks the agent, complexity moves the effort, stakes move the tier and who reads it. A one-line auth change is simple and high-stakes at once, and the stakes decide. See `TIERS.md`.
62
62
  - **Pin the route when it matters:** a lane with no `--model`/`--effort` and no `defaults` entry in `bin/lanes.json` runs on its own config, which may be nothing like what this file describes. `cli-run --doctor` prints what each lane is pinned to, and every run logs the value requested and where it came from.
63
63
 
64
64
  ## Example routings
@@ -47,21 +47,25 @@ reasoning than the reviewer judging its output.** When the plan is airtight the
47
47
  spec is carrying the thinking, so builder drops to medium. When the plan is
48
48
  vague, fix the plan; do not buy reasoning to paper over it.
49
49
 
50
- **Risk moves the tier and the reader, never just the effort.** These four are
50
+ **Stakes move the tier and the reader, never just the effort.** These four are
51
51
  the ones worth naming, because their failures are not recoverable by editing the
52
52
  code afterwards.
53
53
 
54
- | Risk | Present when the change touches | What it buys |
54
+ Stakes means what a mistake would cost: a security hole, leaked personal data,
55
+ lost data, or something you can't undo. Most tasks are low-stakes and route
56
+ normally.
57
+
58
+ | Stakes | Present when the change touches | What it buys |
55
59
  |---|---|---|
56
- | security | auth, tokens, sessions, routes, untrusted input | the attack pass, ideally a different model family |
60
+ | security | auth, tokens, sessions, routes, untrusted input | the challenge pass, ideally a different model family |
57
61
  | privacy | personal data, anything leaving the machine | the local lane, and a named check on what is sent |
58
62
  | data loss | deletion, bulk mutation, migrations, overwrites | a reviewed rollback path before the change is written |
59
63
  | irreversible | publishing, sending, rotating, anything with an audience | a human yes at Stage 5b, never an agent's |
60
64
 
61
- A risk raises code-reviewer to xhigh, and a security-shaped diff goes to the
62
- attack lane rather than to a second read by the same family. Risk is not a
63
- synonym for difficulty: a one-line change to an auth check is simple and
64
- high-risk at the same time, and it is the risk that decides the route.
65
+ High stakes raise code-reviewer to xhigh, and a security-shaped diff goes to
66
+ the challenge lane rather than to a second read by the same family. Stakes are
67
+ not a synonym for difficulty: a one-line change to an auth check is simple and
68
+ high-stakes at the same time, and it is the stakes that decide the route.
65
69
 
66
70
  **Reserve the top of the ladder for evidence.** xhigh and the escalation tier are
67
71
  bought with a named reason: a reproduced failure, a checkpoint that came back
@@ -70,7 +74,7 @@ task, not an escalation.
70
74
 
71
75
  ## Why split tiers: robustness first, cost second
72
76
 
73
- The split produces better work. The deep tier steers every build twice, and what it steers is **judgment, never retrieval**: the orchestrator sweeps the blast radius itself and hands the deep tier a finished map. Paying deep-tier rates for a file list is the most expensive routing mistake available.
77
+ The split produces better work. The deep tier steers every build twice, and what it steers is **judgment, never retrieval**: the orchestrator sweeps everything it touches itself and hands the deep tier a finished map. Paying deep-tier rates for a file list is the most expensive routing mistake available.
74
78
 
75
79
  Against a baseline of "standard tier with no consults", default checkpoints are a spend increase. That is the accepted trade, not a saving to claim.
76
80
 
@@ -36,7 +36,7 @@ The tools cannot help a model that never reaches for them. codecalc ships `SKILL
36
36
 
37
37
  ## What it is not
38
38
 
39
- Not a cloud sandbox for multi-tenant loads, not a replacement for a vendor's built-in interpreter when zero setup matters more than measurement. Its threat model is single-operator, local, stdio. It earns its keep when the correctness of a claim, not "it ran", is the point.
39
+ Not a cloud sandbox for multi-tenant loads, not a replacement for a vendor's built-in interpreter when zero setup matters more than measurement. It assumes a single-operator, local, stdio setup. It earns its keep when the correctness of a claim, not "it ran", is the point.
40
40
 
41
41
  ## On a box (level 3)
42
42
 
@@ -11,7 +11,7 @@ A durable, searchable, governed store that the protocols can call by name:
11
11
  | Need in the protocols | obsidian-tc tool |
12
12
  |---|---|
13
13
  | find what exists before writing (deep research dedupe, gap analysis) | `semantic_search`, `search_text`, `search_regex` |
14
- | map a rename's blast radius (propagate) | `get_backlinks`, `find_unresolved_links`, `rewrite_link` |
14
+ | map everything a rename touches (propagate) | `get_backlinks`, `find_unresolved_links`, `rewrite_link` |
15
15
  | record the end-to-end doc (build Stage 7) | `write_note` (compare-and-swap, confirmation on overwrite), `patch_note`, `append_note` |
16
16
  | keep inferred content honest | `write_note` with `provenance: "agent_synthesis"` runs a poison scan before the write lands |
17
17
  | keep a shared vault safe for several agents | JWT scopes, per-vault folder ACLs, a read-only kill switch, human-in-the-loop tokens |
@@ -56,9 +56,9 @@ Merge the block; do not replace the file.
56
56
 
57
57
  ## Security posture, read before a second agent touches it
58
58
 
59
- Zero-config mode boots with **auth off and no folder ACL**: anything that can reach the server has the same authority as raw filesystem access to the vault. That is acceptable only because the surface is local-only (the config fail-closes if you enable HTTP on a non-loopback host with auth off, and a DNS-rebinding guard protects loopback). Before exposing it to partially-trusted, remote or multi-agent callers, turn on `auth.mode: "jwt"` and set `acl.readPaths` / `writePaths` / `deletePaths` in the config file. Upstream `SECURITY.md` has the threat model and a private disclosure path.
59
+ Zero-config mode boots with **auth off and no folder ACL**: anything that can reach the server has the same authority as raw filesystem access to the vault. That is acceptable only because the surface is local-only (the config refuses by default if you enable HTTP on a non-loopback host with auth off, and a DNS-rebinding guard protects loopback). Before exposing it to partially-trusted, remote or multi-agent callers, turn on `auth.mode: "jwt"` and set `acl.readPaths` / `writePaths` / `deletePaths` in the config file. Upstream `SECURITY.md` has the security notes and a private disclosure path.
60
60
 
61
- Track record worth knowing: an independent code audit of v1.8.1 (July 2026) found three security-relevant gaps (an ACL fail-closed bypass in enumeration tools, a compare-and-swap bypass through `upsert`, a poison-eligibility gap in preference extraction). All three were fixed upstream before they were filed; verified against the v1.25.0 source on 2026-09-03.
61
+ Track record worth knowing: an independent code audit of v1.8.1 (July 2026) found three security-relevant gaps (an ACL bypass that let enumeration tools skip its refuse-by-default rule, a compare-and-swap bypass through `upsert`, a poison-eligibility gap in preference extraction). All three were fixed upstream before they were filed; verified against the v1.25.0 source on 2026-09-03.
62
62
 
63
63
  ## Level 3
64
64