model-orchestrator 0.1.12 → 0.1.14

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,38 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.14] - 2026-09-09
8
+
9
+ Three refinements to the routing model, from a review by [@shawnwows](https://x.com/shawnwows). The theme is the same in all three: a routing decision that was implied, inherited or asserted is now stated, pinned or checked.
10
+
11
+ ### Added
12
+
13
+ - **`--model` and `--effort` on every lane, and a route recorded per run.** A lane with no flag and no `defaults` entry in `bin/lanes.json` runs on its own config file, which `cli-run` cannot see: a CLI configured months ago at a low reasoning effort keeps auditing at that effort while the routing docs describe an adversarial pass, and nothing raises an error. Each vendor spells the flags differently and `cli-run` translates (`grok -m/--reasoning-effort`, `codex -m/-c model_reasoning_effort="X"`, `agy --model/--effort`, `hermes -m/--reasoning`, `qwen -m` and no reasoning flag), each one read from that CLI's own `--help`. Flags beat `defaults`, `defaults` beats nothing, `--doctor` prints what each lane is pinned to, and the log carries `model_requested`, `effort_requested`, `model_source` and `effort_source` on every record, including runs refused before the lane started. It records no "actual": one lane of five (grok) reports a model id in its own output and the other four report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string, which the durable log never holds. `--effort` on qwen is a usage error rather than a silent drop, and route values are charset-bounded because a model id becomes an argv element and, on codex, part of a TOML value.
14
+ - **`finding-verifier`, a sixth subagent, in both agent formats.** Review and scanner findings no longer go straight to a repair. It reads the cited line, states what would trigger the problem, hunts for the guard, caller or test that makes it impossible, and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a change; INCONCLUSIVE is never rounded up to be safe or down to be tidy. Bound into the build protocol as Stage 5a, into `ROUTING.md`, and into the Claude Code activation snippet. The reproduction rule already existed in Stage 5; it had no owner, no separate model family and no way to say "I could not settle this".
15
+ - **Complexity and risk as inputs, alongside role** (`TIERS.md`). Complexity moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output. Risk (security, privacy, data loss, irreversible) moves the tier and who reads the result, because none of those failures is fixable by editing the code afterwards. A one-line change to an auth check is simple and high-risk at once, and the risk decides. Deliberately two rules and two small tables rather than a role by complexity by risk matrix: an 80-cell table is not maintained, and an unmaintained routing table is worse than none because it is believed.
16
+
17
+ ### Changed
18
+
19
+ - `--model` is no longer qwen-only. `--safe-mode` still is.
20
+ - The route is resolved before the "lane disabled" and "binary missing" refusals, so those records carry it too. Found by the pre-release audit: a run refused for a missing binary is still a run that requested a route, and a failure record without one is the gap this release exists to close.
21
+ - `bin/lanes.json` gains an optional `defaults` block. It fails closed with the rest of the file: an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane with no reasoning flag refuses every lane until it is fixed, rather than being skipped quietly.
22
+ - The generated activation list gains a step about pinning the route, and `--doctor` output gains a route column with a plain sentence about what "not pinned" means.
23
+
24
+ ## [0.1.13] - 2026-09-08
25
+
26
+ Three issues from a fresh first-run walkthrough of 0.1.12 (#26, #27, #28). Same class as 0.1.12's five: a surface describing an install that did not happen. A fourth, #25, was filed and closed as a mistake on the reporter's side, not a defect: the warning it said was missing has been printed since 0.1.12 and the repro had been read through a truncated pipe.
27
+
28
+ ### Fixed
29
+
30
+ - **The "Then prove it took" list no longer sends a level 1 reader to a file level 1 never wrote** (#27). Step 4 told every reader, at every level, to pick a lane out of `bin/lanes.json` and run `node bin/cli-run.mjs`. Level 1 writes no `bin/` at all, and step 3 immediately above it hedged correctly with "At level 2+" while step 4 did not. The list is now `proofSteps()` in `src/install.js`, gated on level the same way `activationSteps()` is, and the template renders it. Two tests: the README's section must equal the array exactly for every level and primary, and no `bin/` path may appear in it that the plan did not write.
31
+ - **The box setup no longer tells you to sign in to CLIs you did not pick** (#26). `templates/advanced/vm/README.md` step 3 was a fixed sentence naming `codex login --device-auth`, `grok login --device-auth` and `agy`. A level 3 install of claude-code, codex, qwen and ollama was told to sign in to two CLIs it does not have and never told about the one it does. The step now renders each selected CLI's own `auth` string from the catalog. Everything else in that file was already computed from the selection, which is what made the one hardcoded line easy to miss.
32
+ - **A selected local runtime is finally told to install itself** (#26). `activationSteps()` filtered on `kind === 'agent-cli'`, so Ollama, which has a binary and a download page, appeared in no ordered list at any level. Its only mention was one row of a URL table in `DELEGATION_MATRIX.md`. It now gets a step naming the download page and the `ollama pull <model>` that has to follow it.
33
+ - **The tool block stopped saying the same word twice** (#28). Every run that selected a tool printed `optional: Optional. Needs Python 3.10+ and uv.`, because the label repeated the note's own first word. The label is `note:` now. The note keeps the word, because `--list` and the interactive picker print it bare with no label.
34
+
35
+ ### Changed
36
+
37
+ - `--primary` is documented as what it is. `--help` called it "required when several qualify", and then a `--yes` run with several candidates silently picked one in catalog order. The run now names the choice in the plan (`primary claude-code (chosen for you from claude-code, codex; pass --primary to decide it yourself)`) and the help says the same thing. Behaviour is unchanged: the default was sensible, only the promise was wrong.
38
+
7
39
  ## [0.1.12] - 2026-09-08
8
40
 
9
41
  Five issues from one first-run walkthrough of 0.1.11 (#20 to #24). Every one of them is the same failure: a page describing an install that did not happen. Each fix removes the second copy of a fact rather than correcting it.
@@ -185,7 +217,9 @@ First release.
185
217
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
186
218
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
187
219
 
188
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.12...HEAD
220
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...HEAD
221
+ [0.1.14]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...v0.1.14
222
+ [0.1.13]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.12...v0.1.13
189
223
  [0.1.12]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.11...v0.1.12
190
224
  [0.1.11]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.10...v0.1.11
191
225
  [0.1.10]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.9...v0.1.10
package/README.md CHANGED
@@ -25,7 +25,7 @@ It never writes a secret, never runs a vendor shell script for you, and never ov
25
25
  | Level | You have | You get |
26
26
  |---|---|---|
27
27
  | **1 · Beginner** | one LLM or one agent | tiers, task classification, the two build checkpoints, the protocols (build, propagate, gap analysis, deep research, numbers and logic, memory and record), a task-bundle template, and your agent set up to follow them |
28
- | **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts), a delegation matrix generated from your selection, research triage across the lanes you have |
28
+ | **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts; `--model` / `--effort` to pin the route and log it), a delegation matrix generated from your selection, research triage across the lanes you have |
29
29
  | **3 · Advanced** | a virtual machine | everything above, plus a gateway config rendered from the API keys you hold (asked separately from your CLIs), pinned images, box rules, privacy gates, and a weekly gap-analysis job with "what watches it" written down |
30
30
 
31
31
  Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md).
@@ -90,7 +90,7 @@ ai-orchestrator/
90
90
  TASK_BUNDLE.md the brief every delegation carries
91
91
  protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record
92
92
  CODECALC.md OBSIDIAN-TC.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
93
- <project>/.claude/agents/ five subagents, one per tier, at the PROJECT root (if Claude Code is primary)
93
+ <project>/.claude/agents/ six subagents, one per tier plus finding-verifier, at the PROJECT root (if Claude Code is primary)
94
94
  CLAUDE.snippet.md the block to paste into your CLAUDE.md
95
95
  ROUTING.md multi-lane decision tree (level 2+)
96
96
  TIERS.md DELEGATION_MATRIX.md RESEARCH_TRIAGE.md CLI-RUN.md
@@ -162,6 +162,54 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
162
162
  6. **The orchestrator owns the main build.** Delegates hold none of your rules; they get bounded sub-parts and a brief.
163
163
  7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
164
164
 
165
+ ## Routing by role, complexity and risk
166
+
167
+ Role picks the agent. Two more inputs move the choice, and they move it in
168
+ different directions, so `TIERS.md` states them separately rather than folding
169
+ them into the role:
170
+
171
+ - **Complexity moves the effort.** A worker executing a finished plan needs less
172
+ reasoning than the reviewer judging its output. When the plan is airtight the
173
+ spec is carrying the thinking.
174
+ - **Risk moves the tier and the reader.** Security, privacy, data loss and
175
+ irreversible changes buy the attack lane, a named check, a rollback path or a
176
+ human yes. A one-line change to an auth check is simple and high-risk at the
177
+ same time, and it is the risk that decides.
178
+
179
+ The top of the ladder is bought with evidence: a reproduced failure, an
180
+ unresolved checkpoint, an irreversible change. A task that merely feels hard is
181
+ a deep-tier task, not an escalation.
182
+
183
+ ## A finding is a claim, not a fact
184
+
185
+ Review findings do not go straight to a repair. `finding-verifier` reads the
186
+ cited line, states what would trigger the problem, then hunts for the guard,
187
+ caller or test that makes it impossible, and returns **CONFIRMED**,
188
+ **NOT_REPRODUCED** or **INCONCLUSIVE** per finding. Only CONFIRMED earns a
189
+ change. Use a different model family from the one that produced the finding
190
+ where you have one: a family asked to check its own claim tends to agree with
191
+ itself.
192
+
193
+ ## Pin the route, or know that you did not
194
+
195
+ A lane with no `--model`, no `--effort` and no `defaults` entry in
196
+ `bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
197
+ CLI configured months ago at a low reasoning effort keeps auditing at that
198
+ effort while your routing docs describe an adversarial pass.
199
+
200
+ ```bash
201
+ node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
202
+ node bin/cli-run.mjs --doctor # prints what each lane is pinned to, and what is not pinned
203
+ ```
204
+
205
+ Every run logs the model and effort **requested** and where the request came
206
+ from: `flag`, `lanes.json`, or `lane_default`, on every record including the
207
+ runs that never reached a lane. It does not log an actual. One lane of five
208
+ (grok) reports a model id in its own output and the other four report none, so
209
+ an actual field would be present for one lane and missing for four, and it
210
+ would be a provider-supplied string, which the durable log deliberately never
211
+ holds.
212
+
165
213
  ## Requirements
166
214
 
167
215
  Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows is untested: `cli-run` ends a lane's process tree there with `taskkill`, but nothing in CI runs on Windows, so treat it as unsupported until someone reports otherwise.
@@ -177,6 +225,14 @@ Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box tem
177
225
 
178
226
  Add an AI to `src/catalog.js` and every prompt, table, config and doc picks it up. Run `npm test`. Keep templates free of logic and free of anything that looks like a credential. The rest is in [CONTRIBUTING.md](CONTRIBUTING.md); releases in [RELEASING.md](RELEASING.md); security reports in [SECURITY.md](SECURITY.md).
179
227
 
228
+ ## Credits
229
+
230
+ - [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for
231
+ separating role, complexity and risk instead of compressing them into one
232
+ scale, for recording the model and effort a lane was actually asked for, and
233
+ for verifying findings before they trigger repairs. All three shipped in
234
+ 0.1.14.
235
+
180
236
  ## License
181
237
 
182
238
  [MIT](LICENSE)
package/bin/cli-run.mjs CHANGED
@@ -39,6 +39,17 @@
39
39
  //
40
40
  // The durable log stores a FIXED reason code per run (see REASONS), never a
41
41
  // provider-supplied string. Bounded vendor stderr goes to your terminal only.
42
+ //
43
+ // ROUTE: which model and reasoning effort a lane ran with.
44
+ // A lane with no --model and no lanes.json default inherits whatever its own
45
+ // config file says, which is invisible from here and is how a documented route
46
+ // silently stops being the route that runs. --model / --effort pin it per call,
47
+ // `defaults` in lanes.json pins it per lane, and every run logs the value that
48
+ // was REQUESTED plus where the request came from (flag, lanes.json, or nothing
49
+ // at all). It does not log an "actual". One lane of five (grok) does report a
50
+ // model id in its own output; the other four report none, and a field present
51
+ // for one lane and absent for four is worse than no field. It would also be a
52
+ // provider-supplied string, which this log deliberately never holds.
42
53
 
43
54
  import { spawn } from 'node:child_process';
44
55
  import { StringDecoder } from 'node:string_decoder';
@@ -190,28 +201,71 @@ export function judgeQwen(rc, out) {
190
201
  return pass(text, `subtype=success, totalErrors=0 across ${Object.keys(models).length} model(s)`);
191
202
  }
192
203
 
204
+ // --- route: model and effort per lane --------------------------------------
205
+ // Each vendor spells these differently, and the spelling was read from each
206
+ // CLI's own --help, not remembered. A lane with `effort: null` has no reasoning
207
+ // flag at all; asking for one there is a usage error, never a silent drop.
208
+ // grok -m MODEL --reasoning-effort EFFORT
209
+ // codex -m MODEL -c model_reasoning_effort="EFFORT" (a TOML override, hence the quotes)
210
+ // agy --model M --effort EFFORT (low|medium|high)
211
+ // hermes -m MODEL --reasoning LEVEL (none|minimal|...)
212
+ // qwen -m MODEL no reasoning flag
213
+ export const LANE_FLAGS = {
214
+ grok: { model: (v) => ['-m', v], effort: (v) => ['--reasoning-effort', v] },
215
+ codex: { model: (v) => ['-m', v], effort: (v) => ['-c', `model_reasoning_effort="${v}"`] },
216
+ agy: { model: (v) => ['--model', v], effort: (v) => ['--effort', v] },
217
+ hermes: { model: (v) => ['-m', v], effort: (v) => ['--reasoning', v] },
218
+ qwen: { model: (v) => ['-m', v], effort: null }
219
+ };
220
+
221
+ // A model id or effort level becomes an argv element and, for codex, part of a
222
+ // TOML value. Bounding the charset is what makes both safe: no leading dash (a
223
+ // value cannot become a flag), no quote, space or control character (a value
224
+ // cannot break out of the TOML string), and a length cap so a config file
225
+ // cannot push an unbounded string into the durable log.
226
+ export const ROUTE_VALUE = /^[A-Za-z0-9][A-Za-z0-9._:@/+-]{0,63}$/;
227
+ export function badRouteValue(kind, v) {
228
+ if (typeof v !== 'string' || !ROUTE_VALUE.test(v)) {
229
+ return `--${kind} must be 1 to 64 characters of letters, digits, dot, underscore, colon, at, slash, plus or dash, and may not start with a dash: ${JSON.stringify(v)}`;
230
+ }
231
+ return null;
232
+ }
233
+
193
234
  // --- adapters: build argv for a lane -------------------------------------
235
+ // Route flags go in front of the prompt for every lane, because two lanes
236
+ // (hermes, codex) take the prompt as a positional argument and a flag after it
237
+ // is either ignored or read as part of it.
238
+ function routeFlags(lane, opts) {
239
+ const spec = LANE_FLAGS[lane];
240
+ const out = [];
241
+ if (!spec) return out;
242
+ if (opts.model) out.push(...spec.model(opts.model));
243
+ if (opts.effort && spec.effort) out.push(...spec.effort(opts.effort));
244
+ return out;
245
+ }
246
+
194
247
  export function buildArgv(lane, binary, prompt, opts, tmp) {
195
248
  const timeout = opts.timeout;
249
+ const route = routeFlags(lane, opts);
196
250
  switch (lane) {
197
251
  case 'grok':
198
- return { argv: [binary, '--output-format', 'json', '-p', prompt] };
252
+ return { argv: [binary, '--output-format', 'json', ...route, '-p', prompt] };
199
253
  case 'codex': {
200
254
  const last = join(tmp, 'last.txt');
201
255
  const argv = [binary, 'exec', '--json', '--color', 'never', '--skip-git-repo-check', '-o', last];
202
256
  if (opts.audit) argv.push('--sandbox', 'read-only'); // an audit lane that can write is a bug
257
+ argv.push(...route);
203
258
  argv.push(prompt);
204
259
  return { argv, outFile: last };
205
260
  }
206
261
  case 'agy': {
207
262
  const mins = Math.max(1, Math.round(timeout / 60));
208
- return { argv: [binary, '--print-timeout', `${mins}m`, '--output-format', 'stream-json', '-p', prompt] };
263
+ return { argv: [binary, '--print-timeout', `${mins}m`, '--output-format', 'stream-json', ...route, '-p', prompt] };
209
264
  }
210
265
  case 'hermes':
211
- return { argv: [binary, '-z', prompt, '--usage-file', join(tmp, 'usage.json')] };
266
+ return { argv: [binary, '-z', ...route, prompt, '--usage-file', join(tmp, 'usage.json')] };
212
267
  case 'qwen': {
213
- const argv = [binary, '-o', 'json'];
214
- if (opts.model) argv.push('-m', opts.model);
268
+ const argv = [binary, '-o', 'json', ...route];
215
269
  if (opts.safeMode) argv.push('--safe-mode');
216
270
  argv.push('-p', prompt);
217
271
  return { argv };
@@ -370,26 +424,67 @@ function log(rec) {
370
424
  // documented default). PRESENT BUT UNREADABLE OR MALFORMED = no lane enabled:
371
425
  // a half-written config must fail closed, never re-enable what the installer
372
426
  // disabled. Returns null when the file is bad so the caller can say so.
373
- export function enabledLanes(here = dirname(fileURLToPath(import.meta.url))) {
427
+ export function laneConfig(here = dirname(fileURLToPath(import.meta.url))) {
374
428
  const p = join(here, 'lanes.json');
375
- if (!existsSync(p)) return LANES;
429
+ if (!existsSync(p)) return { enabled: LANES, defaults: {} };
376
430
  try {
377
431
  const j = JSON.parse(readFileSync(p, 'utf8'));
378
432
  if (!j || typeof j !== 'object' || !Array.isArray(j.enabled)) return null;
379
433
  if (!j.enabled.every((l) => typeof l === 'string' && LANES.includes(l))) return null;
380
- return j.enabled;
434
+ // `defaults` pins a model and effort per lane. It is optional; present and
435
+ // malformed fails closed with the rest of the file, because a half-written
436
+ // route is exactly the silent-inheritance problem this field exists to fix.
437
+ const defaults = {};
438
+ if (j.defaults !== undefined) {
439
+ if (!j.defaults || typeof j.defaults !== 'object' || Array.isArray(j.defaults)) return null;
440
+ for (const [lane, d] of Object.entries(j.defaults)) {
441
+ if (!LANES.includes(lane)) return null;
442
+ if (!d || typeof d !== 'object' || Array.isArray(d)) return null;
443
+ const { model, effort, ...rest } = d;
444
+ if (Object.keys(rest).length) return null;
445
+ if (model !== undefined && badRouteValue('model', model)) return null;
446
+ if (effort !== undefined) {
447
+ if (badRouteValue('effort', effort)) return null;
448
+ if (!LANE_FLAGS[lane] || !LANE_FLAGS[lane].effort) return null; // a lane with no reasoning flag cannot have one pinned
449
+ }
450
+ defaults[lane] = { model: model ?? null, effort: effort ?? null };
451
+ }
452
+ }
453
+ return { enabled: j.enabled, defaults };
381
454
  } catch {
382
455
  return null;
383
456
  }
384
457
  }
385
458
 
459
+ // Kept as the narrow question most callers ask. null still means malformed.
460
+ export function enabledLanes(here = dirname(fileURLToPath(import.meta.url))) {
461
+ const c = laneConfig(here);
462
+ return c === null ? null : c.enabled;
463
+ }
464
+
465
+ // Flag beats lanes.json beats nothing. `source` is what makes the log audit-worthy:
466
+ // 'lane_default' means this run inherited the vendor CLI's own config, unseen from here.
467
+ export function resolveRoute(lane, opts, defaults) {
468
+ const d = (defaults && defaults[lane]) || {};
469
+ const model = opts.model ?? d.model ?? null;
470
+ const effort = opts.effort ?? d.effort ?? null;
471
+ const src = (flag, def) => (flag != null ? 'flag' : def != null ? 'lanes.json' : 'lane_default');
472
+ return { model, effort, model_source: src(opts.model, d.model), effort_source: src(opts.effort, d.effort) };
473
+ }
474
+
386
475
  function usage(msg) {
387
476
  if (msg) console.error('cli-run: ' + msg);
388
477
  console.error(`usage: cli-run <${LANES.join('|')}> "<prompt>" [--brief FILE] [--timeout SECS] [--quiet]
389
- [--expect-file PATH] [--expect-json]
478
+ [--model ID] [--effort LEVEL] [--expect-file PATH] [--expect-json]
390
479
  cli-run codex --audit "<prompt>" read-only sandbox (audit shape)
391
- cli-run qwen [--model ID] [--safe-mode] "<prompt>"
392
- cli-run --doctor [--run] enabled lanes, binaries on PATH; --run sends each a tiny prompt`);
480
+ cli-run qwen [--safe-mode] "<prompt>" qwen-only flag
481
+ cli-run --doctor [--run] enabled lanes, binaries, and the route each one is pinned to
482
+
483
+ --model / --effort pin what a lane runs with, instead of letting it inherit its
484
+ own config. Every lane takes --model; every lane except qwen takes --effort.
485
+ Levels are the vendor's own (agy low|medium|high, hermes none|minimal|...): an
486
+ unknown level is rejected by the lane, and reported as that lane's exit code.
487
+ Pin them per lane instead of per call with "defaults" in bin/lanes.json.`);
393
488
  return USAGE;
394
489
  }
395
490
 
@@ -416,11 +511,12 @@ function installedPrimary(here = dirname(fileURLToPath(import.meta.url))) {
416
511
 
417
512
  // --doctor: the first thing to run after install.
418
513
  export async function doctor(run) {
419
- const enabled = enabledLanes();
420
- if (enabled === null) {
514
+ const cfg = laneConfig();
515
+ if (cfg === null) {
421
516
  console.error('doctor: lanes.json exists but is malformed; fix it first');
422
517
  return USAGE;
423
518
  }
519
+ const { enabled, defaults } = cfg;
424
520
  let bad = 0;
425
521
  console.log(`doctor: ${enabled.length} enabled lane(s): ${enabled.join(', ') || 'none'}`);
426
522
  const primary = installedPrimary();
@@ -432,7 +528,11 @@ export async function doctor(run) {
432
528
  for (const lane of LANES) {
433
529
  const on = enabled.includes(lane);
434
530
  const bin = which(lane);
435
- let line = ` ${lane.padEnd(7)} ${on ? 'enabled ' : 'disabled'} ${bin ? 'binary ok' : 'binary MISSING'}`;
531
+ const d = defaults[lane] || {};
532
+ // A disabled lane has no route worth reporting; saying "not pinned" there
533
+ // reads as a finding about a lane that is not going to run.
534
+ const route = !on ? '' : d.model || d.effort ? `route ${d.model || 'lane default'}/${d.effort || 'lane default'}` : 'route not pinned (inherits the lane\'s own config)';
535
+ let line = ` ${lane.padEnd(7)} ${on ? 'enabled ' : 'disabled'} ${bin ? 'binary ok' : 'binary MISSING'}${route ? ' ' + route : ''}`;
436
536
  if (on && !bin) bad++;
437
537
  if (on && bin && run) {
438
538
  const rc = await main([lane, 'Reply with exactly the word OK and nothing else.', '--timeout', '120', '--quiet']);
@@ -443,6 +543,7 @@ export async function doctor(run) {
443
543
  }
444
544
  console.log(bad ? `doctor: ${bad} problem(s)` : 'doctor: all enabled lanes ' + (run ? 'answered' : 'present'));
445
545
  console.log('doctor checks presence and, with --run, a one-word canary. It does not check vendor versions.');
546
+ console.log('"route not pinned" means that lane runs on whatever its own config file says, which this tool cannot see. Pin it in lanes.json "defaults" if the route matters.');
446
547
  return bad ? NO_DELIVERABLE : OK;
447
548
  }
448
549
 
@@ -485,10 +586,10 @@ export function checkContracts(opts, text, before) {
485
586
  }
486
587
 
487
588
  export async function main(argv) {
488
- const VALUE = new Set(['--brief', '--timeout', '--model', '--expect-file']);
589
+ const VALUE = new Set(['--brief', '--timeout', '--model', '--effort', '--expect-file']);
489
590
  const BOOL = new Set(['--quiet', '--audit', '--safe-mode', '--doctor', '--run', '--expect-json']);
490
591
  const args = [...argv];
491
- const opts = { timeout: 900, quiet: false, audit: false, model: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
592
+ const opts = { timeout: 900, quiet: false, audit: false, model: null, effort: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
492
593
  const positional = [];
493
594
  while (args.length) {
494
595
  const a = args.shift();
@@ -498,6 +599,7 @@ export async function main(argv) {
498
599
  if (a === '--brief') opts.brief = v;
499
600
  else if (a === '--timeout') opts.timeout = Number(v);
500
601
  else if (a === '--expect-file') opts.expectFile = v;
602
+ else if (a === '--effort') opts.effort = v;
501
603
  else opts.model = v;
502
604
  } else if (BOOL.has(a)) {
503
605
  if (a === '--quiet') opts.quiet = true;
@@ -530,11 +632,33 @@ export async function main(argv) {
530
632
  if (!prompt) return usage('give a prompt or --brief FILE');
531
633
  if (!Number.isFinite(opts.timeout) || opts.timeout <= 0) return usage('--timeout must be a positive number of seconds');
532
634
  if (opts.audit && lane !== 'codex') return usage('--audit is codex-only');
533
- if ((opts.model || opts.safeMode) && lane !== 'qwen') return usage('--model and --safe-mode are qwen-only');
635
+ if (opts.safeMode && lane !== 'qwen') return usage('--safe-mode is qwen-only');
636
+ for (const [kind, v] of [['model', opts.model], ['effort', opts.effort]]) {
637
+ if (v == null) continue;
638
+ const bad = badRouteValue(kind, v);
639
+ if (bad) return usage(bad);
640
+ }
641
+ // qwen has no reasoning flag. Dropping --effort silently would leave the caller
642
+ // believing a route that never happened, which is the defect this feature fixes.
643
+ if (opts.effort && !(LANE_FLAGS[lane] && LANE_FLAGS[lane].effort)) return usage(`${lane} has no reasoning-effort flag; --effort is not available on this lane`);
534
644
 
535
645
  const digest = createHash('sha256').update(prompt).digest('hex').slice(0, 12);
536
646
  const base = { lane, prompt_sha256_12: digest, prompt_chars: prompt.length };
537
- const enabled = enabledLanes();
647
+ const cfg = laneConfig();
648
+ // Resolve the route BEFORE the refusals below. A run that never reached a lane
649
+ // was still a request for one, and a failure record with no route is the exact
650
+ // gap this feature exists to close. A malformed lanes.json has no usable
651
+ // defaults, so the flags stand alone and say so.
652
+ const route = resolveRoute(lane, opts, cfg === null ? {} : cfg.defaults);
653
+ opts.model = route.model;
654
+ opts.effort = route.effort;
655
+ Object.assign(base, {
656
+ model_requested: route.model,
657
+ effort_requested: route.effort,
658
+ model_source: route.model_source,
659
+ effort_source: route.effort_source
660
+ });
661
+ const enabled = cfg === null ? null : cfg.enabled;
538
662
  if (enabled === null) {
539
663
  console.error('cli-run: lanes.json exists but is not a valid {"enabled": [...]} file; refusing every lane until it is fixed');
540
664
  log({ ...base, verdict: 'unavailable', rc: UNAVAILABLE, reason: 'lanes_json_malformed' });
@@ -601,7 +725,8 @@ export async function main(argv) {
601
725
  }
602
726
  }
603
727
  if (text && code === OK) process.stdout.write(text + '\n');
604
- if (!opts.quiet) console.error(`cli-run[${lane}] ${verdict} rc=${code} ${r.seconds.toFixed(1)}s raw=${r.outBytes || 0}B :: ${detail}`);
728
+ const routeNote = route.model || route.effort ? `${route.model || 'lane default'}/${route.effort || 'lane default'}` : 'lane default';
729
+ if (!opts.quiet) console.error(`cli-run[${lane}] ${verdict} rc=${code} ${r.seconds.toFixed(1)}s raw=${r.outBytes || 0}B route=${routeNote} :: ${detail}`);
605
730
  // Durable log: fixed reason code and structural numbers only.
606
731
  log({ ...base, verdict, rc: code, cli_rc: r.status, signal: r.signal || null, seconds: Math.round(r.seconds * 100) / 100, raw_bytes: r.outBytes || 0, deliverable_bytes: Buffer.byteLength(text), reason: REASONS.has(reason) ? reason : 'unknown' });
607
732
  return code;
package/bin/cli.js CHANGED
@@ -93,7 +93,8 @@ Usage
93
93
  Flags
94
94
  --level 1|2|3 1 beginner (one agent), 2 intermediate (many CLIs), 3 advanced (plus a VM)
95
95
  --ais a,b,c catalog ids you have access to (see --list)
96
- --primary id the agent that runs the system and receives the subagents (any level; required when several qualify)
96
+ --primary id the agent that runs the system and receives the subagents (any level). When several qualify
97
+ and --yes is set, the run picks one and says so in the plan; pass this to decide it yourself.
97
98
  --tools a,b companion tools to set up, all optional (default with --yes: codecalc only); --no-tools for none
98
99
  --apis a,b level 3 only: metered API keys you HOLD (anthropic,openai,google,xai,openrouter); --no-apis for none.
99
100
  Asked separately from the CLIs because a subscription is not an API key.
@@ -189,6 +190,7 @@ async function main() {
189
190
  // 3. Primary agent (the one that runs the system)
190
191
  const candidates = agentCandidates(selected);
191
192
  let primary = null;
193
+ let primaryAutoPicked = false;
192
194
  if (opt('primary')) {
193
195
  primary = byId[opt('primary')];
194
196
  if (!primary || !candidates.includes(primary)) bad('--primary must be one of: ' + candidates.map((a) => a.id).join(', '));
@@ -200,7 +202,10 @@ async function main() {
200
202
  // --yes picks for the user: claude-code if present, else the first agent that can load subagent
201
203
  // definitions (it gets five files written for it), else the first candidate. #19: codex listed
202
204
  // before agy used to win and nothing was written to the project root.
203
- if (yes) primary = candidates.find((a) => a.id === 'claude-code') || candidates.find((a) => a.agentsDir) || candidates[0];
205
+ if (yes) {
206
+ primary = candidates.find((a) => a.id === 'claude-code') || candidates.find((a) => a.agentsDir) || candidates[0];
207
+ primaryAutoPicked = true;
208
+ }
204
209
  else {
205
210
  console.log('\nWhich one is your primary agent (the one that runs the system)?');
206
211
  candidates.forEach((a, i) => console.log(` ${i + 1} ${a.name}`));
@@ -271,7 +276,7 @@ async function main() {
271
276
  const files = planFiles({ level, selected, primary, dir, project, tools, apis });
272
277
  const lvl = LEVELS.find((l) => l.id === level);
273
278
  const agentFiles = files.filter((f) => f.root === 'project');
274
- console.log(`\nPlan\n level ${lvl.id} ${lvl.name}\n access ${selected.map((a) => a.id).join(', ')}\n primary ${primary ? primary.id : 'none'}\n tools ${tools.map((t) => t.id).join(', ') || 'none'}` + (level >= 3 ? `\n api keys ${apis.map((p) => p.id).join(', ') || 'none'}` : '') + `\n folder ${dir}\n project ${project}${agentFiles.length ? ' (' + agentFiles.length + ' subagent files go here)' : ''}\n files ${files.length}`);
279
+ console.log(`\nPlan\n level ${lvl.id} ${lvl.name}\n access ${selected.map((a) => a.id).join(', ')}\n primary ${primary ? primary.id : 'none'}${primaryAutoPicked ? ` (chosen for you from ${candidates.map((a) => a.id).join(', ')}; pass --primary to decide it yourself)` : ''}\n tools ${tools.map((t) => t.id).join(', ') || 'none'}` + (level >= 3 ? `\n api keys ${apis.map((p) => p.id).join(', ') || 'none'}` : '') + `\n folder ${dir}\n project ${project}${agentFiles.length ? ' (' + agentFiles.length + ' subagent files go here)' : ''}\n files ${files.length}`);
275
280
  if (agentFiles.length && !opt('project')) {
276
281
  console.log(`\nNote: --project was not given, so the ${agentFiles.length} subagent file(s) go to the current directory (${project}). Pass --project to put them somewhere else.`);
277
282
  }
@@ -365,7 +370,7 @@ async function main() {
365
370
 
366
371
  for (const t of tools) {
367
372
  const doc = t.id.toUpperCase() + '.md';
368
- console.log(`\n${t.name}\n optional: ${t.optionalNote}\n needs: ${t.requires}\n run: ${t.install}\n one-click or self-registering for: ${t.autoClients.join(', ')}. Other agents and the details: ${dir}/${doc}`);
373
+ console.log(`\n${t.name}\n note: ${t.optionalNote}\n needs: ${t.requires}\n run: ${t.install}\n one-click or self-registering for: ${t.autoClients.join(', ')}. Other agents and the details: ${dir}/${doc}`);
369
374
  }
370
375
 
371
376
  // 7. Activation summary: writing the folder is half the job. Say exactly what
@@ -14,6 +14,8 @@ If your agent exposes model choice (Claude Code, Codex, Antigravity), map the ti
14
14
 
15
15
  Three cost levers, always together: tier (price per token), token discipline (how many tokens: read only what you will touch, never re-read, deliverables not narration), effort (how hard each call thinks).
16
16
 
17
+ And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **Risk** moves the tier and who reads the result: security, privacy, data loss and irreversible changes are the four worth naming, because none of their failures can be fixed by editing the code afterwards. A one-line change to an auth check is simple and high-risk at once, and it is the risk that decides.
18
+
17
19
  Robustness first, cost second. You split tiers because the split produces better work.
18
20
 
19
21
  ## 2. Classify every task, first match wins
@@ -26,7 +26,9 @@ One driver, no second AI in the mix: the orchestrator invokes the CLIs; it never
26
26
 
27
27
  Every agent CLI can report success and deliver nothing. `bin/cli-run.mjs` builds the right invocation per lane, reads that lane's **native** terminal event, and exits `10` when a run produced no deliverable, `12` on timeout, `13` when the lane is missing. Byte count is not a check either; a run can emit hundreds of kilobytes and no conclusion. One lane's own success flags lie outright (an upstream 400 reported as success), so its judge reads the two honest signals instead.
28
28
 
29
- Every call goes through it. "This lane is flaky" becomes a query over its log instead of an argument. `node bin/cli-run.mjs --doctor` is the first thing to run after install: enabled lanes, binaries on PATH, and with `--run` a one-word canary per lane.
29
+ Every call goes through it. "This lane is flaky" becomes a query over its log instead of an argument. `node bin/cli-run.mjs --doctor` is the first thing to run after install: enabled lanes, binaries on PATH, the route each lane is pinned to, and with `--run` a one-word canary per lane.
30
+
31
+ There is a second thing a lane can be quietly wrong about. Left unpinned, it runs on **its own config file**, which the runner cannot see: a CLI set up months ago at a low reasoning effort keeps auditing at that effort while your routing docs describe an adversarial pass, and no error is ever raised. `--model` and `--effort` pin it per call, `defaults` in `bin/lanes.json` pins it per lane, and every run records the value requested and where it came from (`flag`, `lanes.json`, `lane_default`). The log claims no actual: grok reports a model id in its output, the other four lanes report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string the durable log never holds.
30
32
 
31
33
  ## 4. Every delegation carries a task bundle, on both surfaces
32
34
 
@@ -36,6 +38,10 @@ Subagents and CLI lanes are the same problem: something with none of your rules
36
38
 
37
39
  Fan the same plan to three model families (web sweep, adversarial read, live data), each as one `cli-run` call. The orchestrator opens the primary sources itself, marks every claim, and writes the only durable record. Expect one engine to return confident unsourced numerics; downgrade it. Weight the engines that report their own gaps. Count dispositions, not briefs.
38
40
 
41
+ ## 5a. A finding is a claim, not a fact
42
+
43
+ An audit that returns six findings has returned six claims. Hand them to `finding-verifier` before any of them causes a repair: it reads the cited line, states what would trigger the problem, then hunts for the guard, caller or test that makes it impossible, and answers CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE. Only CONFIRMED earns a change. Use a different family from the one that produced the finding, and let INCONCLUSIVE stand: rounding it up to be safe buys unnecessary repairs, rounding it down to be tidy hides real ones.
44
+
39
45
  ## 6. Gap analysis gets a second family
40
46
 
41
47
  The second pass is now a different model reading the same artifact, in read-only audit mode. Disagreement between families is the cheapest signal that something is soft.
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.12",
4
- "description": "Routing instructions and a CLI runner for your AI tools. One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine.",
3
+ "version": "0.1.14",
4
+ "description": "Routing instructions and a CLI runner for your AI tools. One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine. Routes by role, complexity and risk; pins and logs the model and reasoning effort each lane runs with.",
5
5
  "type": "module",
6
6
  "bin": {
7
7
  "model-orchestrator": "bin/cli.js"
@@ -51,7 +51,11 @@
51
51
  "agents",
52
52
  "subagents",
53
53
  "qwen",
54
- "mcp"
54
+ "mcp",
55
+ "reasoning-effort",
56
+ "model-routing",
57
+ "code-review",
58
+ "ai-agents"
55
59
  ],
56
60
  "author": "aunysillyme (https://github.com/aunysillyme)",
57
61
  "license": "MIT"
package/src/install.js CHANGED
@@ -219,12 +219,32 @@ export function activationSteps(opts) {
219
219
  else if (snippet) steps.push(`open ${primary.chatName || primary.name} and paste the block in ${join(dirAbs, snippet)} into its ${primary.chatSurface || 'custom instructions'}`);
220
220
  if (primary && primary.agentsDir) steps.push(`subagents are in ${join(projectAbs, primary.agentsDir)}; run ${primary.bin} from ${projectAbs} to pick them up`);
221
221
  for (const a of selected.filter((a) => a.bin && a.kind === 'agent-cli')) steps.push(`sign in to ${a.name}: ${a.auth}`);
222
+ // A local runtime has a bin but no sign-in, so the agent-cli loop above skips it
223
+ // and before this it appeared in no ordered list at any level (#26).
224
+ for (const a of selected.filter((a) => a.bin && a.kind === 'local')) steps.push(`install ${a.name}: ${a.install.url}, then \`${a.bin} pull <model>\` before the local lane can answer`);
222
225
  for (const t of tools) steps.push(`${t.id}: ${t.install}`);
223
226
  if (level >= 2) steps.push(`smoke test: node ${join(dirAbs, 'bin', 'cli-run.mjs')} --doctor (add --run to send each lane one tiny prompt)`);
224
227
  if (level >= 3) steps.push(`box: read ${join(dirAbs, 'vm', 'README.md')}; keys named in vm/ENVIRONMENT.md go in your secrets manager, never a file`);
225
228
  return steps;
226
229
  }
227
230
 
231
+ // The verification list, in order. Gated on level for the same reason
232
+ // activationSteps is: level 1 writes no bin/, so a step naming cli-run.mjs or
233
+ // lanes.json there described an install that did not happen (#27).
234
+ export function proofSteps(opts) {
235
+ const { level } = opts;
236
+ const steps = [
237
+ 'Start a fresh agent session and ask: "Read the orchestrator instructions. Quote the routing rule you will use, then sort pear, apple, banana alphabetically. Name the tier and whether you delegated."',
238
+ 'Expect the fast tier and `apple, banana, pear`. If the agent cannot quote the routing rule, check the snippet location or chat instructions before continuing. This is a manual activation check, not proof that every future task follows the rules.'
239
+ ];
240
+ if (level >= 2) {
241
+ steps.push('Run `node bin/cli-run.mjs --doctor` from this folder. It checks binary presence, not authentication or loaded instructions, and prints the model and effort each lane is pinned to. `--doctor --run` additionally uses a little quota to test live responses. No enabled lanes means delegation is inactive.');
242
+ steps.push('Decide whether the route matters to you. Every lane starts unpinned, which means it runs on whatever its own config file says: a CLI configured months ago at a low reasoning effort will keep auditing at that effort while your docs describe something stronger. Pin it in `bin/lanes.json` under `defaults`, or per call with `--model` and `--effort`. Either way the run is recorded in the log with the value requested and where it came from.');
243
+ steps.push('To test a real output contract, choose an enabled lane from `bin/lanes.json` and run `node bin/cli-run.mjs <lane> \'Return only {"sorted":["apple","banana","pear"]}\' --expect-json`. This uses quota. Expect JSON and exit 0; inspect the array yourself. A non-JSON response exits 10, a missing binary exits 13, and an authentication failure reports the vendor error. The explicit lane tests execution; your primary agent still makes delegation decisions.');
244
+ }
245
+ return steps;
246
+ }
247
+
228
248
  function vars(opts) {
229
249
  const { level, selected, primary } = opts;
230
250
  const tools = opts.tools || [];
@@ -240,6 +260,7 @@ function vars(opts) {
240
260
  const pinOf = (id) => (toolById[id] && toolById[id].pin) || 'latest';
241
261
  const snippet = snippetFor(primary);
242
262
  const steps = activationSteps({ level, selected, primary, tools, dir: opts.dir, project: opts.project });
263
+ const proofs = proofSteps({ level });
243
264
  // Only claude-code and agy put files under the project root. A chat primary
244
265
  // puts nothing there, so naming a project root would name a folder this run
245
266
  // never created (#21).
@@ -253,6 +274,7 @@ function vars(opts) {
253
274
  return {
254
275
  ...laneVars(selected),
255
276
  ACTIVATION_STEPS: steps.map((st, i) => `${i + 1}. ${st}`).join('\n'),
277
+ PROOF_STEPS: proofs.map((st, i) => `${i + 1}. ${st}`).join('\n'),
256
278
  LOAD_IT: readsProjectRules
257
279
  ? `${primary.name} reads its rules from \`${primary.rulesFile}\` in the project root. The installer wrote \`${snippet}\` next to this README; copy its contents into \`${join(projectAbs, primary.rulesFile)}\`, creating that file if it does not exist. Nothing was appended to a file you already had.`
258
280
  : snippet
@@ -272,6 +294,12 @@ function vars(opts) {
272
294
  INSTALL_DIR: dirAbs,
273
295
  INSTALL_DIR_SH: shellQuote(dirAbs),
274
296
  INSTALL_DIR_SYSTEMD: systemdEscape(dirAbs),
297
+ // vm/README.md step 3 named `grok login` and `agy` whatever you picked (#26).
298
+ VM_SIGNIN: (() => {
299
+ const lines = selected.filter((a) => a.bin && a.kind === 'agent-cli').map((a) => ` - ${a.name}: ${a.auth}`);
300
+ for (const a of selected.filter((a) => a.bin && a.kind === 'local')) lines.push(` - ${a.name}: no sign-in. Install it from ${a.install.url}, then \`${a.bin} pull <model>\`.`);
301
+ return lines.length ? lines.join('\n') : ' - none: no CLI you selected needs a sign-in on the box.';
302
+ })(),
275
303
  AUDIT_LANE: lane || 'none',
276
304
  // Enforced boundary per lane: codex has a read-only sandbox flag; the others
277
305
  // run with whatever their own config allows, and the script says so.
@@ -362,7 +390,9 @@ export function planFiles(opts) {
362
390
  JSON.stringify(
363
391
  {
364
392
  enabled: selected.filter((a) => a.cliRun).map((a) => a.id),
365
- note: 'Lanes cli-run may call. Edit to enable or disable a lane. A lane not listed here exits 13 (unavailable).'
393
+ defaults: {},
394
+ note: 'Lanes cli-run may call. Edit to enable or disable a lane. A lane not listed here exits 13 (unavailable).',
395
+ defaultsNote: 'Pin what a lane runs with, so the route in your docs is the route that runs: "defaults": {"codex": {"model": "gpt-6-astra", "effort": "high"}}. Left empty, a lane inherits its own config file, which cli-run cannot see and does not guess. `--model` and `--effort` override this per call, and `--doctor` prints what each lane is pinned to. Every lane takes a model; every lane except qwen takes an effort.'
366
396
  },
367
397
  null,
368
398
  2
@@ -23,7 +23,8 @@ Generated {{DATE}} for: `{{AI_IDS}}`. Installed at `{{INSTALL_DIR}}`; the system
23
23
 
24
24
  1. Provision a box. Ubuntu, 2+ vCPU, 8 GB is comfortable. Put it on a private mesh network if you can; do not open ports to the internet.
25
25
  2. `bash setup-vm.sh`. It installs system deps and the npm-installable CLIs, then **prints** the vendor shell installers for the rest. Read those scripts before running them.
26
- 3. Sign each CLI in with its device-code flow (`codex login --device-auth`, `grok login --device-auth`, `agy` on first run). Run these inside `tmux` so a dropped SSH session does not kill the prompt. Headless Linux has no keyring by default; `setup-vm.sh` installs one so the CLIs stop re-prompting.
26
+ 3. Sign each CLI in, using the flow its vendor gives you. Run these inside `tmux` so a dropped SSH session does not kill the prompt. Headless Linux has no keyring by default; `setup-vm.sh` installs one so the CLIs stop re-prompting.
27
+ {{VM_SIGNIN}}
27
28
  4. Put provider keys in your secrets manager and export the names listed in `ENVIRONMENT.md` into the gateway's environment at start time. Never write a value into a file in this folder. The gateway config was rendered from the API keys you said you hold, not from your CLI subscriptions: those are different entitlements.
28
29
  5. `docker compose up -d`, then list the lanes without putting the key in argv (the key must be a single token, `^[A-Za-z0-9._-]+$`, because it is interpolated into curl's config grammar):
29
30
  ```bash
@@ -1,5 +1,5 @@
1
1
  # .agents/agents/
2
2
 
3
- Antigravity CLI custom agents, one per tier, in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
3
+ Antigravity CLI custom agents, one per tier plus a finding-verifier, in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
4
4
 
5
5
  `commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; `auto` keeps deletes and other high-risk commands gated) and `off` for the read-only agents. `model` is a tier: `pro` for deep-planner, `flash` for the rest.
@@ -0,0 +1,29 @@
1
+ ---
2
+ name: finding-verifier
3
+ description: Adversarial verification of review findings; tries to disprove each one and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE. Read-only, never repairs.
4
+ model: flash
5
+ subagent: true
6
+ mainAgent: true
7
+ commandExecutionPolicy: off
8
+ ---
9
+
10
+ # finding-verifier
11
+
12
+ A finding is a claim, not a fact. You try to disprove each one before it is
13
+ allowed to cause a repair.
14
+
15
+ For each finding you are given: read the cited file and line yourself, state the
16
+ input or sequence that would trigger it, then hunt for what makes it impossible
17
+ (a guard upstream, a caller that never passes that value, an existing test).
18
+
19
+ Return one verdict per finding, in the order given:
20
+ - CONFIRMED: reproduced, or a concrete unblocked path. Give the path.
21
+ - NOT_REPRODUCED: you found what stops it. Name it and where it is.
22
+ - INCONCLUSIVE: not settleable read-only. Say what you would need.
23
+
24
+ Rules:
25
+ - Stay inside the task bundle you were given. Anything not granted is denied.
26
+ - Verify only the findings handed to you; anything else you notice goes at the end, marked unverified.
27
+ - Never round INCONCLUSIVE up to CONFIRMED to be safe, or down to NOT_REPRODUCED to be tidy.
28
+ - Read-only: you never repair and never reword a finding.
29
+ - Token discipline: read the cited code and its callers, not the repository.
@@ -1,12 +1,13 @@
1
1
  # .claude/agents/
2
2
 
3
- Five subagents, one per tier. Claude Code loads project-level agents from this folder automatically.
3
+ Six subagents. Five are one per tier; `finding-verifier` is the check that sits between a review and a repair. Claude Code loads project-level agents from this folder automatically.
4
4
 
5
5
  | Agent | Tier | Model alias | Effort | Job |
6
6
  |---|---|---|---|---|
7
7
  | deep-planner | deep | opus | xhigh | judges every build twice; never retrieves |
8
8
  | builder | standard | sonnet | high | bounded sub-parts of a build |
9
9
  | code-reviewer | standard | sonnet | high | read-only findings |
10
+ | finding-verifier | standard | sonnet | high | tries to disprove a finding before it causes a repair |
10
11
  | live-researcher | standard | sonnet | medium | fresh data through tools |
11
12
  | bulk-worker | fast | haiku | low | mechanical volume |
12
13
 
@@ -0,0 +1,43 @@
1
+ ---
2
+ name: finding-verifier
3
+ description: Adversarial verification of review findings. Use after a review or audit returns findings and before any of them trigger a repair. Read-only. Tries to DISPROVE each finding and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Do not use to find new problems, and do not use to fix anything.
4
+ tools: Read, Glob, Grep, Bash
5
+ model: sonnet
6
+ effort: high
7
+ ---
8
+
9
+ You are the verification tier of the model router.
10
+
11
+ A finding is a claim, not a fact. Your job is to try to disprove each one before
12
+ it is allowed to cause a change. A false finding is expensive twice: it buys a
13
+ repair nobody needed, and it teaches everyone to skim the next report.
14
+
15
+ You are given findings from a review or an audit. For each one, independently:
16
+
17
+ 1. Read the cited file and line yourself. A citation that does not point at what
18
+ the finding describes is already a failure of the finding, not of the code.
19
+ 2. State the exact input, state or sequence that would make it happen.
20
+ 3. Look for what makes it impossible: a guard upstream, a type that cannot hold
21
+ that value, a caller that never passes it, a test that already covers it, a
22
+ framework guarantee.
23
+ 4. Where you can run something cheap and read-only that settles it, run it.
24
+
25
+ Return one verdict per finding, in the order you were given them:
26
+
27
+ - **CONFIRMED** you reproduced it, or traced a concrete path to it that nothing
28
+ prevents. Give the path in one or two sentences.
29
+ - **NOT_REPRODUCED** you found what stops it. Name that thing and where it is.
30
+ This is a success, not a failure to try.
31
+ - **INCONCLUSIVE** you could not settle it read-only. Say exactly what you would
32
+ need: a test run, a credential, a live environment, a decision from a human.
33
+ Never round this up to CONFIRMED to be safe, and never down to
34
+ NOT_REPRODUCED to be tidy.
35
+
36
+ Rules:
37
+ - Verify only the findings you were given. New problems you happen to notice go
38
+ in a separate list at the end, clearly marked as unverified observations.
39
+ - You are read-only. You never repair, and you never soften a finding's wording.
40
+ - Verifying nothing is a real answer. If every finding is NOT_REPRODUCED, say
41
+ that plainly; a verifier that always confirms something is a rubber stamp
42
+ facing the other way.
43
+ - Token discipline: read the cited code and its callers, not the repository.
@@ -10,6 +10,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
10
10
  1. Bulk, mechanical, many similar items -> bulk-worker (fast tier).
11
11
  2. Needs live data -> live-researcher (standard tier + tools).
12
12
  3. Review without changing -> code-reviewer (standard, read-only).
13
+ 3a. Holding findings from a review or scanner -> finding-verifier before any repair. Only CONFIRMED findings earn a change.
13
14
  4. Ambiguous, architectural, or expensive to get wrong -> deep-planner (deep tier), then hand the plan down.
14
15
  5. Everything else that changes files -> build it directly. The main build is never handed off whole; bounded sub-parts go to builder.
15
16
 
@@ -21,10 +21,7 @@ These are the same steps, in the same order, that the installer printed in your
21
21
 
22
22
  ## Then prove it took
23
23
 
24
- 1. Start a fresh agent session and ask: "Read the orchestrator instructions. Quote the routing rule you will use, then sort pear, apple, banana alphabetically. Name the tier and whether you delegated."
25
- 2. Expect the fast tier and `apple, banana, pear`. If the agent cannot quote the routing rule, check the snippet location or chat instructions before continuing. This is a manual activation check, not proof that every future task follows the rules.
26
- 3. At level 2+, run `node bin/cli-run.mjs --doctor` from this folder. It checks binary presence, not authentication or loaded instructions. `--doctor --run` additionally uses a little quota to test live responses. No enabled lanes means delegation is inactive.
27
- 4. To test a real output contract, choose an enabled lane from `bin/lanes.json` and run `node bin/cli-run.mjs <lane> 'Return only {"sorted":["apple","banana","pear"]}' --expect-json`. This uses quota. Expect JSON and exit 0; inspect the array yourself. A non-JSON response exits 10, a missing binary exits 13, and an authentication failure reports the vendor error. The explicit lane tests execution; your primary agent still makes delegation decisions.
24
+ {{PROOF_STEPS}}
28
25
 
29
26
  ## What is in this folder
30
27
 
@@ -66,7 +66,17 @@ Secret detection, static analysis and dependency scanning, filtered to lines thi
66
66
 
67
67
  Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to an adversarial auditor, ideally a **different model family**. Architecture-shaped diffs go to the judgment tier reviewing build against plan. Never both on one diff.
68
68
 
69
- **Gate:** every finding **reproduced** before it reaches a human. Unreproduced items are dropped, not narrated. Hard cap one re-audit. `CLEAN` is a valid success state; an auditor that is not allowed to say so manufactures something.
69
+ Findings do not go straight to a repair. Hand them to the **finding-verifier**, whose job is to DISPROVE each one: read the cited line, state what would trigger it, then hunt for the guard, caller or test that makes it impossible. It returns one of three verdicts per finding, and rounding between them is the failure mode to watch for.
70
+
71
+ | Verdict | Meaning | What happens next |
72
+ |---|---|---|
73
+ | CONFIRMED | reproduced, or a concrete path nothing blocks | it earns a repair |
74
+ | NOT_REPRODUCED | something prevents it, named and located | dropped, and not narrated |
75
+ | INCONCLUSIVE | not settleable read-only | say what it would take; never round it to either side |
76
+
77
+ Use a different model family from the one that produced the finding where you have one: a family asked to check its own claim tends to agree with itself.
78
+
79
+ **Gate:** every finding **verified** before it reaches a human, and only CONFIRMED findings trigger a change. Hard cap one re-audit. `CLEAN` is a valid success state; an auditor that is not allowed to say so manufactures something, and so does a verifier that is expected to confirm.
70
80
 
71
81
  ### Stage 5b · Ship gate
72
82
  1. What is the rollback target? Record it before shipping.
@@ -18,7 +18,8 @@ Enabled lanes (edit `bin/lanes.json`): {{CLI_RUN_LANES}}
18
18
  ```bash
19
19
  node bin/cli-run.mjs <grok|codex|agy|hermes|qwen> "<prompt>" [--brief FILE] [--timeout SECS] [--quiet]
20
20
  node bin/cli-run.mjs codex --audit "<prompt>" # read-only sandbox, the audit shape
21
- node bin/cli-run.mjs qwen [--model ID] [--safe-mode] "<prompt>"
21
+ node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
22
+ node bin/cli-run.mjs qwen [--safe-mode] "<prompt>" # qwen-only flag
22
23
  ```
23
24
 
24
25
  Put it on your PATH if you like: `ln -s "$PWD/bin/cli-run.mjs" ~/.local/bin/cli-run`.
@@ -65,13 +66,49 @@ qwen is the lane whose own success flags lie: an upstream 400 comes back as exit
65
66
  | 2 | usage error in cli-run itself |
66
67
  | N | the lane exited N != 0: passed through unchanged, verdict `exit_nonzero`, even when parseable text came back. The bounded head of the lane's stderr is shown on your terminal so an auth failure reads as one |
67
68
 
69
+ ## The route: which model, and how hard it thinks
70
+
71
+ A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe an adversarial pass, and nothing anywhere says so.
72
+
73
+ Pin it per call, or per lane:
74
+
75
+ ```bash
76
+ node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high # this call only
77
+ node bin/cli-run.mjs --doctor # prints what each lane is pinned to
78
+ ```
79
+
80
+ ```json
81
+ {
82
+ "enabled": ["codex", "grok"],
83
+ "defaults": { "codex": { "model": "gpt-6-astra", "effort": "high" } }
84
+ }
85
+ ```
86
+
87
+ A flag beats `defaults`; `defaults` beats nothing. Each vendor spells these differently and `cli-run` translates:
88
+
89
+ | Lane | Model | Reasoning effort |
90
+ |---|---|---|
91
+ | grok | `-m` | `--reasoning-effort` |
92
+ | codex | `-m` | `-c model_reasoning_effort="LEVEL"` |
93
+ | agy | `--model` | `--effort` (low, medium, high) |
94
+ | hermes | `-m` | `--reasoning` (none, minimal, ...) |
95
+ | qwen | `-m` | none: this lane has no reasoning flag |
96
+
97
+ Three rules that keep this honest:
98
+
99
+ - **A level `cli-run` does not recognise is not rejected here.** Levels are the vendor's, they change, and guessing the valid set would date this tool. An unknown level is refused by the lane and surfaces as that lane's own exit code and stderr.
100
+ - **`--effort` on qwen is a usage error, not a silent drop.** A flag that vanishes leaves you believing a route that never ran.
101
+ - **Values are charset-bounded** (letters, digits, and `. _ : @ / + -`, no leading dash, 64 characters). A model id becomes an argv element and, on codex, part of a TOML value; bounding it is what stops either from being escaped.
102
+
68
103
  ## Permissions are a separate layer
69
104
 
70
105
  `cli-run` never injects permission flags. Each CLI carries its own config, so every caller gets the same behaviour. Use each vendor's deny-list as the base layer; allow-lists only hold if every binary is enumerable in advance.
71
106
 
72
107
  ## Log
73
108
 
74
- `~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, rc, the lane's own exit code, signal, seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument.
109
+ `~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, rc, the lane's own exit code, signal, seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, the route (`model_requested`, `effort_requested`, and `model_source` / `effort_source`, each one of `flag`, `lanes.json` or `lane_default`), and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument, and so does "we route audits at high effort".
110
+
111
+ The log records what was **requested**, on every record including a run refused before the lane started. It does not record an actual. Reporting is inconsistent: grok returns a `modelUsage` block naming a model, the other four lanes return nothing of the kind, so an `actual` field would be populated for one lane and empty for four. It would also be a provider-supplied string, and this log holds fixed codes and bounded caller-supplied values only. `model_source: "lane_default"` is the honest way to say this run inherited something invisible from here.
75
112
 
76
113
  ## The prompt travels in argv
77
114
 
@@ -79,7 +116,7 @@ That is each vendor's documented headless shape (`-p`, `exec`). Two consequences
79
116
 
80
117
  ## lanes.json fails closed
81
118
 
82
- Absent: every lane enabled. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled.
119
+ Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all fail the whole file closed rather than being skipped quietly.
83
120
 
84
121
  ## A killed lane is not a deliverable
85
122
 
@@ -17,6 +17,7 @@ Rule of thumb: never spend a frontier token on a task a cheap tier finishes corr
17
17
  1. **Bulk and mechanical?** → fast tier{{BULK_LANE}}. Many independent items each needing its own agent turn → a concurrent fan-out lane if you have one.
18
18
  2. **Needs live data?** → {{LIVE_LANE}} standard tier with web tools.
19
19
  3. **Reviewing without changing?** → standard tier read-only. Security-critical → {{ATTACK_LANE}}.
20
+ 3a. **Holding findings from a review or a scanner?** → finding-verifier before any of them cause a repair. A finding is a claim, not a fact.
20
21
  4. **Ambiguous, strategic, expensive to get wrong?** → deep tier (deep-planner). Then hand the plan down.
21
22
  5. **Everything else that changes files** → the orchestrator builds it directly. Bounded sub-parts go to cheaper tiers; the main build is never handed off whole.
22
23
 
@@ -38,6 +39,7 @@ Every delegation carries `TASK_BUNDLE.md`. Its brief must restate every conventi
38
39
  | 3 Build | the orchestrator, against the installed dependency's source |
39
40
  | 4 Scan | secret + static + dependency scanners, diff-scoped, fail closed |
40
41
  | 5 Attack | security-shaped diff → {{ATTACK_LANE}}. Architecture-shaped → deep tier, build against plan. Never both |
42
+ | 5a Verify findings | finding-verifier, a different model family where you have one: CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a repair |
41
43
  | 5b Ship | rollback id recorded, explicit human yes |
42
44
  | 6 Verify | real test, negative test seen red, old identifier re-grepped to zero |
43
45
  | 7 Record | one end-to-end doc, tracker Done with evidence, plan doc deleted |
@@ -59,7 +61,9 @@ One writer per run; every other lane proposes. Search before writing, index in t
59
61
  - **De-escalation:** a request that sounds deep but is a lookup routes down.
60
62
  - **Long context:** mechanical digestion → fast tier in chunks; judgment over a long input → standard tier.
61
63
  - **Token discipline on every delegation:** pass only the context the delegate needs, never the conversation.
62
- - **Effort per agent:** deep xhigh, review and build high, live research medium, bulk low.
64
+ - **Effort per agent:** deep xhigh, review, verification and build high, live research medium, bulk low.
65
+ - **Three inputs, not one:** role picks the agent, complexity moves the effort, risk moves the tier and who reads it. A one-line auth change is simple and high-risk at once, and the risk decides. See `TIERS.md`.
66
+ - **Pin the route when it matters:** a lane with no `--model`/`--effort` and no `defaults` entry in `bin/lanes.json` runs on its own config, which may be nothing like what this file describes. `cli-run --doctor` prints what each lane is pinned to, and every run logs the value requested and where it came from.
63
67
 
64
68
  ## Example routings
65
69
 
@@ -70,4 +74,5 @@ One writer per run; every other lane proposes. Search before writing, index in t
70
74
  | "Add an endpoint" | the orchestrator builds it |
71
75
  | "Why does this silently drop rows sometimes" | deep-planner (unknown cause), then build the fix directly |
72
76
  | "Summarize these 30 notes into one index" | bulk-worker |
77
+ | "The audit returned 6 findings" | finding-verifier first; repair only what comes back CONFIRMED |
73
78
  {{LANE_EXAMPLES}}
@@ -19,11 +19,52 @@ Tier sets the price per token. Token discipline sets how many tokens. **Effort s
19
19
  |---|---|---|---|
20
20
  | deep-planner | deep | xhigh | judges every build twice; expensive to get wrong |
21
21
  | code-reviewer | standard | high | every endpoint is internet-facing |
22
+ | finding-verifier | standard | high | judging a claim is harder than producing it |
22
23
  | builder | standard | high | a botched deploy is the costly failure |
23
24
  | live-researcher | standard | medium | tools do the retrieval |
24
25
  | bulk-worker | fast | low | the biggest cost win |
25
26
 
26
- Dials: drop builder to medium when the plan is airtight; raise code-reviewer to xhigh for a security-critical audit.
27
+ ## Three inputs, not one
28
+
29
+ Role alone does not decide a route. Two more inputs move it, and they move it in
30
+ opposite directions, so state them separately instead of folding them into the
31
+ role.
32
+
33
+ **Complexity moves the effort.** The same role does not need the same reasoning
34
+ on every task.
35
+
36
+ | Complexity | What it looks like | What moves |
37
+ |---|---|---|
38
+ | simple | one file, one obvious edit, no unknowns | drop one effort level |
39
+ | standard | the default | the table above |
40
+ | complex | several surfaces, or an unknown cause | keep effort, add the deep-tier checkpoint |
41
+ | critical | irreversible, or it rewrites a standing rule | the escalation rule below applies |
42
+
43
+ The dial that pays for itself: **a worker executing a finished plan needs less
44
+ reasoning than the reviewer judging its output.** When the plan is airtight the
45
+ spec is carrying the thinking, so builder drops to medium. When the plan is
46
+ vague, fix the plan; do not buy reasoning to paper over it.
47
+
48
+ **Risk moves the tier and the reader, never just the effort.** These four are
49
+ the ones worth naming, because their failures are not recoverable by editing the
50
+ code afterwards.
51
+
52
+ | Risk | Present when the change touches | What it buys |
53
+ |---|---|---|
54
+ | security | auth, tokens, sessions, routes, untrusted input | the attack pass, ideally a different model family |
55
+ | privacy | personal data, anything leaving the machine | the local lane, and a named check on what is sent |
56
+ | data loss | deletion, bulk mutation, migrations, overwrites | a reviewed rollback path before the change is written |
57
+ | irreversible | publishing, sending, rotating, anything with an audience | a human yes at Stage 5b, never an agent's |
58
+
59
+ A risk raises code-reviewer to xhigh, and a security-shaped diff goes to the
60
+ attack lane rather than to a second read by the same family. Risk is not a
61
+ synonym for difficulty: a one-line change to an auth check is simple and
62
+ high-risk at the same time, and it is the risk that decides the route.
63
+
64
+ **Reserve the top of the ladder for evidence.** xhigh and the escalation tier are
65
+ bought with a named reason: a reproduced failure, a checkpoint that came back
66
+ unresolved, an irreversible change. A task that merely feels hard is a deep-tier
67
+ task, not an escalation.
27
68
 
28
69
  ## Why split tiers: robustness first, cost second
29
70