model-orchestrator 1.0.5 → 1.0.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,36 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [1.0.7] - 2026-09-30
8
+
9
+ ### Fixed
10
+
11
+ - Prefer an already-connected Claude worker MCP service in host routing, with the native CLI retained as a standalone fallback. Route suggestions and generated instructions describe the session, result, permission and environment boundaries; no MCP installation or automatic transport retry is added.
12
+ - Enable the native `claude` CLI worker with print-mode JSON completion checks, existing permission rules, model and effort overrides, failure classification and shared timeout/cancellation handling. Codex-primary stacks can assign independent review to Claude Code; Claude-primary stacks retain Codex review. Generated lanes use executable names while manifest AI identities retain installer IDs. Claude fixtures record native success and authentication failures.
13
+ - Keep supported lanes usable when an edited or unverifiable older runner is kept during an upgrade. Withhold unsupported lanes and their defaults from `lanes.json`, and report how to enable them with `--upgrade-runtime`.
14
+ - Classify Claude's expired OAuth session and `Failed to authenticate` errors as auth (exit 14).
15
+ - Match Claude HTTP codes only in API status fields or `API Error:` prefixes, so a 400 error mentioning 429 tokens stays rejected (exit 16).
16
+
17
+ ### Changed
18
+
19
+ - Check and pin the Claude lane at 2.1.285, and replace its hand-written success fixture with a captured restricted live success at that version.
20
+
21
+ ## [1.0.6] - 2026-09-29
22
+
23
+ ### Added
24
+
25
+ - The hermes lane takes `--provider` and `"provider"` in `bin/lanes.json` defaults, so a pinned model goes to a provider that serves it. A provider needs a model with it (a flag, a default or `HERMES_INFERENCE_MODEL`); otherwise cli-run refuses before the lane starts. Runs log `provider_requested` and `provider_source`.
26
+ - `--doctor` notes a hermes model pinned without a provider, and a provider pinned without a model. The terminal route line shows the provider.
27
+
28
+ ### Fixed
29
+
30
+ - A hermes model/provider mismatch (an upstream "model is not supported" or `model_not_found`, for example `grok-4.6` sent to `openai-codex`) is now class `rejected`, exit 16, with a problem line naming the mismatch and a fix pointing at `hermes model`. It was reported as a vague nonzero exit.
31
+ - A pinned hermes route (`--model` or `--effort`, by flag or `lanes.json` default) no longer fails every run. cli-run placed those flags between `-z` and the prompt, and `-z` takes the prompt as its value, so Hermes refused the arguments ("argument -z/--oneshot: expected one argument"). Route flags now come before `-z`, and a Hermes argument error is reported as such rather than as a toolsets error.
32
+
33
+ ### Changed
34
+
35
+ - The product website generates public Markdown mirrors, llms discovery files and agent navigation guidance from its existing documentation allowlist on each Git build.
36
+
7
37
  ## [1.0.5] - 2026-09-28
8
38
 
9
39
  ### Changed
@@ -557,7 +587,9 @@ First release.
557
587
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
558
588
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
559
589
 
560
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.5...HEAD
590
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.7...HEAD
591
+ [1.0.7]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.6...v1.0.7
592
+ [1.0.6]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.5...v1.0.6
561
593
  [1.0.5]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.4...v1.0.5
562
594
  [1.0.4]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.3...v1.0.4
563
595
  [1.0.3]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.2...v1.0.3
package/README.md CHANGED
@@ -175,6 +175,8 @@ It helps an agent match a task to a model, effort level and toolset. Run `npx mo
175
175
 
176
176
  Run `npx model-orchestrator --yes --level 2 --ais claude-code,codex --primary claude-code --project . --dir ./ai-orchestrator`, then follow the activation summary. Claude Code can dispatch scoped work through `aunx cli-run codex --brief TASK_BRIEF.md` and use a different model family for review.
177
177
 
178
+ For Codex as the main agent, set `--primary codex`. Independent review goes to Claude Code through a connected Claude worker MCP service when the host has one, or through `aunx cli-run claude --brief TASK_BRIEF.md`. The `claude` lane runs in `dontAsk` mode, so calls needing approval are denied; `--audit` stays Codex-only.
179
+
178
180
  ### How do I reduce Claude Code token usage?
179
181
 
180
182
  Install routing rules with `npx model-orchestrator`, so your agent has guidance for sending routine work to cheaper models and keeping reads scoped. Use `aunx route-metrics --summary` to measure where your work goes; savings depend on your tasks and model choices.
@@ -233,7 +235,7 @@ The lane wiring and the output judges were written against these versions, which
233
235
 
234
236
  | Lane | Vendor | Version this release was built against | Where that number is proved |
235
237
  |---|---|---|---|
236
- | `claude` | Anthropic | 2.1.226 | the npm pin the installer writes, `@anthropic-ai/claude-code@2.1.226` |
238
+ | `claude` | Anthropic | 2.1.285 | `test/fixtures/claude-2.1.285-success.json`, a recorded run |
237
239
  | `codex` | OpenAI | 0.153.4 | `test/fixtures/codex-0.153.4.jsonl`, a recorded run |
238
240
  | `agy` | Google | 1.1.27 | `test/fixtures/agy-1.1.27.jsonl`, a recorded run |
239
241
  | `grok` | xAI | 1.0.5 | `test/fixtures/grok-1.0.5.json`, a recorded run |
@@ -241,12 +243,14 @@ The lane wiring and the output judges were written against these versions, which
241
243
  | `qwen` | Alibaba | 0.22.3 | `test/fixtures/qwen-0.22.3-nokey.json`, a recorded run |
242
244
  | `ollama` | Ollama | 0.34.4 | the pinned image the level 3 box runs, `ollama/ollama:0.34.4` |
243
245
 
244
- Generated from `src/catalog.js` by `npm run gen:catalog`; `npm test` fails if this table and the catalog disagree. Fixtures were captured 2026-09-06.
246
+ Generated from `src/catalog.js` by `npm run gen:catalog`; `npm test` fails if this table and the catalog disagree. Original vendor fixtures were captured 2026-09-06; [fixture provenance](test/fixtures/README.md) records later captures and remaining gaps.
245
247
 
246
248
  <!-- vendor-table:end -->
247
249
 
248
250
  For npm-installed lanes, `builtAgainst` in the catalog supplies both the compatibility table and the install pin. Newer vendor versions may work or may change a flag the generated wiring uses. When a lane starts failing after a vendor upgrade, compare against this table first.
249
251
 
252
+ The `claude` lane is checked against recorded Claude Code 2.1.285 runs: a restricted success, a sign-in failure and an expired-session failure.
253
+
250
254
  **The live canary runs on your machine, with your credentials.** That is what `aunx cli-run --doctor --run` (direct form: `node bin/cli-run.mjs --doctor --run`) is: it sends every enabled lane one tiny prompt through your own sign-ins and reports `canary ok` or `canary FAILED rc=` per lane. Choose this optional live check after setup or a vendor upgrade when you want to verify actual responses.
251
255
 
252
256
  CI runs the full suite against stub lanes on Ubuntu, macOS and Windows, Node 18/20/22, plus a packaged install into a clean consumer. Run the live check locally to verify your own sign-ins, quota and vendor versions.
@@ -267,6 +271,7 @@ Run `npm test` with your change. Keep templates free of logic and credential val
267
271
  ## Credits
268
272
 
269
273
  - [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for separating role, complexity and stakes instead of compressing them into one scale, for recording the model and effort a lane was actually asked for, and for verifying findings before they trigger repairs. All three shipped in 0.1.14.
274
+ - [@shawnbissell](https://github.com/shawnbissell) added the native Claude Code worker lane (print-mode JSON judge, failure classes, dontAsk permissions) and the connected-MCP preference for Claude workers.
270
275
 
271
276
  ## License
272
277
 
package/bin/cli-run.mjs CHANGED
@@ -15,6 +15,7 @@
15
15
  // NATIVE terminal event and refuses to call an empty run a success.
16
16
  //
17
17
  // grok --output-format json -> stopReason == "end_turn" and text non-empty
18
+ // claude -p --output-format json -> result subtype success, is_error false, result non-empty
18
19
  // codex exec --json --color never -o F -> terminal {"type":"turn.completed"} and F non-empty
19
20
  // agy --output-format stream-json -> terminal {"event":"result"} status SUCCESS, response non-empty
20
21
  // hermes -z -> its exit code is already honest (0 ok / 1 none / 2 bad args)
@@ -53,15 +54,14 @@
53
54
  // problem/fix lines go to your terminal only, redacted.
54
55
  //
55
56
  // ROUTE: which model and reasoning effort a lane ran with.
57
+ // Hermes also accepts --provider to select the provider serving that model.
56
58
  // A lane with no --model and no lanes.json default inherits whatever its own
57
59
  // config file says, which is invisible from here and is how a documented route
58
60
  // silently stops being the route that runs. --model / --effort pin it per call,
59
61
  // `defaults` in lanes.json pins it per lane, and every run logs the value that
60
62
  // was REQUESTED plus where the request came from (flag, lanes.json, or nothing
61
- // at all). It does not log an "actual". One lane of five (grok) does report a
62
- // model id in its own output; the other four report none, and a field present
63
- // for one lane, and absent for four, is worse than no field. It would also be a
64
- // provider-supplied string, which this log deliberately never holds.
63
+ // at all). It does not log an "actual": model reporting varies by CLI and is
64
+ // provider-supplied text, which this log deliberately never holds.
65
65
 
66
66
  import { spawn, spawnSync } from 'node:child_process';
67
67
  import { StringDecoder } from 'node:string_decoder';
@@ -71,7 +71,7 @@ import { join, dirname, delimiter, resolve, relative, isAbsolute, sep } from 'no
71
71
  import { tmpdir, homedir } from 'node:os';
72
72
  import { fileURLToPath, pathToFileURL } from 'node:url';
73
73
 
74
- export const LANES = ['grok', 'codex', 'agy', 'hermes', 'qwen'];
74
+ export const LANES = ['grok', 'codex', 'agy', 'hermes', 'qwen', 'claude'];
75
75
  export const OK = 0, NO_DELIVERABLE = 10, NO_OUTPUT = 11, TIMEOUT = 12, UNAVAILABLE = 13, USAGE = 2;
76
76
  export const AUTH = 14, QUOTA = 15, REJECTED = 16, REFUSED = 17, CUT_SHORT = 18;
77
77
 
@@ -160,6 +160,31 @@ function lastJsonLine(out, needle, want) {
160
160
  const fail = (reason, detail) => ({ text: null, reason, detail });
161
161
  const pass = (text, detail) => ({ text, reason: 'ok', detail });
162
162
 
163
+ // Print-mode JSON is one result object, not the stream-json event sequence.
164
+ // Source: https://code.claude.com/docs/en/headless and /en/agent-sdk/agent-loop.
165
+ function claudeResultObject(out) {
166
+ try {
167
+ const o = JSON.parse(out);
168
+ return isObj(o) && o.type === 'result' ? o : null;
169
+ } catch { return null; }
170
+ }
171
+
172
+ export function judgeClaude(rc, out) {
173
+ let o;
174
+ try { o = JSON.parse(out); } catch { return fail('not_json', 'stdout was not JSON'); }
175
+ if (!isObj(o) || o.type !== 'result') return fail('not_result', 'JSON was not a result object');
176
+ if (o.subtype !== 'success') return fail('bad_subtype', 'result did not report successful completion');
177
+ if (o.is_error !== false) return fail('is_error', 'result did not report is_error=false');
178
+ if (o.errors !== undefined) {
179
+ if (!Array.isArray(o.errors) || !o.errors.every(e => typeof e === 'string')) return fail('error_message_not_string', 'errors was not an array of strings');
180
+ if (o.errors.length) return fail('is_error', 'result carried execution errors');
181
+ }
182
+ if (o.stop_reason != null && !['end_turn', 'stop_sequence', 'refusal'].includes(o.stop_reason)) return fail('bad_stop_reason', 'result stopped before completing a response');
183
+ if (typeof o.result !== 'string') return fail('result_not_string', 'result was not a string');
184
+ const text = o.result.trim();
185
+ return text ? pass(text, 'subtype=success, is_error=false') : fail('empty_result', 'success but empty result');
186
+ }
187
+
163
188
  export function judgeGrok(rc, out) {
164
189
  let o;
165
190
  try {
@@ -199,7 +224,11 @@ export function judgeHermes(rc, out, err) {
199
224
  // hermes collapses every upstream failure into one exit code; its stderr
200
225
  // is the only place the cause is named.
201
226
  const blob = String(err || '').toLowerCase(); // stderr only: stdout is the agent's own prose
202
- if (hermesQuotaText(blob)) why += ': upstream free tier degraded or limited, a retry is reasonable';
227
+ // A mismatch is checked first: a retry cannot fix it, and an unrelated
228
+ // "limit" warning on stderr must not turn it into a quota report.
229
+ if (hermesMismatchLine(err) || hermesMismatchLine(out, { vendorShape: true })) why += ': model/provider mismatch, the provider does not serve this model';
230
+ else if (hermesQuotaText(blob)) why += ': upstream free tier degraded or limited, a retry is reasonable';
231
+ else if (hermesArgError(blob)) why += ': hermes rejected its arguments, a caller bug and not a lane fault';
203
232
  else if (blob.includes('toolset')) why += ': invalid --toolsets value, a caller bug and not a lane fault';
204
233
  return fail('exit_nonzero', `hermes exit ${rc}: ${why}`);
205
234
  }
@@ -251,21 +280,28 @@ export function judgeQwen(rc, out) {
251
280
  // CLI's own --help, not remembered. A lane with `effort: null` has no reasoning
252
281
  // flag at all; asking for one there is a usage error, never a silent drop.
253
282
  // grok -m MODEL --reasoning-effort EFFORT
283
+ // claude --model M --effort EFFORT
254
284
  // codex -m MODEL -c model_reasoning_effort="EFFORT" (a TOML override, hence the quotes)
255
285
  // agy --model M --effort EFFORT (low|medium|high)
256
- // hermes -m MODEL --reasoning LEVEL (none|minimal|...)
286
+ // hermes -m MODEL --reasoning LEVEL --provider ID (none|minimal|...)
257
287
  // qwen -m MODEL no reasoning flag
288
+ // Only hermes takes a provider: one hermes install signs in to many providers,
289
+ // and a model id sent to a provider that does not serve it is an HTTP 400
290
+ // (grok-4.6 on openai-codex). `hermes -z --provider` without a model exits 2,
291
+ // so cli-run refuses that pairing before the lane starts.
258
292
  export const LANE_FLAGS = {
259
- grok: { model: (v) => ['-m', v], effort: (v) => ['--reasoning-effort', v] },
260
- codex: { model: (v) => ['-m', v], effort: (v) => ['-c', `model_reasoning_effort="${v}"`] },
261
- agy: { model: (v) => ['--model', v], effort: (v) => ['--effort', v] },
262
- hermes: { model: (v) => ['-m', v], effort: (v) => ['--reasoning', v] },
263
- qwen: { model: (v) => ['-m', v], effort: null }
293
+ claude: { model: (v) => ['--model', v], effort: (v) => ['--effort', v], provider: null },
294
+ grok: { model: (v) => ['-m', v], effort: (v) => ['--reasoning-effort', v], provider: null },
295
+ codex: { model: (v) => ['-m', v], effort: (v) => ['-c', `model_reasoning_effort="${v}"`], provider: null },
296
+ agy: { model: (v) => ['--model', v], effort: (v) => ['--effort', v], provider: null },
297
+ hermes: { model: (v) => ['-m', v], effort: (v) => ['--reasoning', v], provider: (v) => ['--provider', v] },
298
+ qwen: { model: (v) => ['-m', v], effort: null, provider: null }
264
299
  };
265
300
 
266
301
  // Auto is deliberately a small, static ladder. It is not a vendor capability
267
302
  // probe and it never chooses the top of a vendor's effort range.
268
303
  export const AUTO_EFFORT = {
304
+ claude: { small: 'medium', large: 'high' },
269
305
  codex: { small: 'medium', large: 'high' },
270
306
  grok: { small: 'medium', large: 'high' },
271
307
  agy: { small: 'medium', large: 'high' },
@@ -353,13 +389,16 @@ export function badRouteValue(kind, v) {
353
389
  }
354
390
 
355
391
  // --- adapters: build argv for a lane -------------------------------------
356
- // Route flags go in front of the prompt for every lane, because two lanes
357
- // (hermes, codex) take the prompt as a positional argument and a flag after it
358
- // is either ignored or read as part of it.
392
+ // Route flags go in front of the prompt for every lane, because codex takes the
393
+ // prompt as a positional argument and a flag after it is either ignored or read
394
+ // as part of it. hermes needs them in front of -z itself: -z takes the prompt as
395
+ // its own value, so `-z -m X prompt` is an argparse error ("argument -z/--oneshot:
396
+ // expected one argument") and every pinned hermes route failed that way.
359
397
  function routeFlags(lane, opts) {
360
398
  const spec = LANE_FLAGS[lane];
361
399
  const out = [];
362
400
  if (!spec) return out;
401
+ if (opts.provider && spec.provider) out.push(...spec.provider(opts.provider));
363
402
  if (opts.model) out.push(...spec.model(opts.model));
364
403
  if (opts.effort && spec.effort) out.push(...spec.effort(opts.effort));
365
404
  return out;
@@ -369,6 +408,11 @@ export function buildArgv(lane, binary, prompt, opts, tmp) {
369
408
  const timeout = opts.timeout;
370
409
  const route = routeFlags(lane, opts);
371
410
  switch (lane) {
411
+ case 'claude':
412
+ // dontAsk denies calls needing approval and retains configured allow/deny
413
+ // rules. It grants no new permissions and is not a filesystem sandbox.
414
+ // -- ends option parsing so even a brief starting with a dash is data.
415
+ return { argv: [binary, '-p', '--output-format', 'json', '--permission-mode', 'dontAsk', ...route, '--', prompt] };
372
416
  case 'grok':
373
417
  return { argv: [binary, '--output-format', 'json', ...route, '-p', prompt] };
374
418
  case 'codex': {
@@ -384,7 +428,7 @@ export function buildArgv(lane, binary, prompt, opts, tmp) {
384
428
  return { argv: [binary, '--print-timeout', `${mins}m`, '--output-format', 'stream-json', ...route, '-p', prompt] };
385
429
  }
386
430
  case 'hermes':
387
- return { argv: [binary, '-z', ...route, prompt, '--usage-file', join(tmp, 'usage.json')] };
431
+ return { argv: [binary, ...route, '-z', prompt, '--usage-file', join(tmp, 'usage.json')] };
388
432
  case 'qwen': {
389
433
  const argv = [binary, '-o', 'json', ...route];
390
434
  if (opts.safeMode) argv.push('--safe-mode');
@@ -398,6 +442,8 @@ export function buildArgv(lane, binary, prompt, opts, tmp) {
398
442
 
399
443
  export function judge(lane, rc, out, err, outFile) {
400
444
  switch (lane) {
445
+ case 'claude':
446
+ return judgeClaude(rc, out);
401
447
  case 'grok':
402
448
  return judgeGrok(rc, out);
403
449
  case 'codex':
@@ -527,6 +573,30 @@ function hermesQuotaText(lower) {
527
573
  return lower.includes('no usable content') || lower.includes('limit') || lower.includes('degraded');
528
574
  }
529
575
 
576
+ // argparse's own refusal. Its usage dump always lists `-t TOOLSETS`, so this is
577
+ // checked before the toolset match or every bad argument reads as a toolset error.
578
+ function hermesArgError(lower) {
579
+ return lower.includes('hermes: error:');
580
+ }
581
+
582
+ // A model id the chosen provider does not serve. `hermes -z` prints a failed
583
+ // turn's provider error on stdout and exits 2, so this reads both streams, but
584
+ // returns only the matching line: the rest of stdout is the agent's own prose.
585
+ // The wording is the upstream's own ("The 'grok-4.6' model is not supported
586
+ // when using Codex with a ChatGPT account"; OpenAI's model_not_found code).
587
+ export function hermesMismatchLine(text, { vendorShape = false } = {}) {
588
+ for (const line of String(text || '').split('\n')) {
589
+ const l = line.toLowerCase();
590
+ if (!(l.includes('model is not supported') || l.includes('model_not_found'))) continue;
591
+ // stdout is agent prose, so there only the vendor's own shape counts: an HTTP
592
+ // status, an error prefix, the Codex wording, or the OpenAI error code.
593
+ if (vendorShape && !(/^\s*(http [45]\d\d|error)\b/.test(l) || l.includes('model is not supported when using') || l.includes('model_not_found'))) continue;
594
+ // Terminal-safe: stdout can carry web-derived text, so no control characters.
595
+ return line.replace(/[\x00-\x08\x0b-\x1f\x7f]/g, '').trim().slice(0, 300);
596
+ }
597
+ return null;
598
+ }
599
+
530
600
  // qwen: the terminal event's full error text, and a result that is an API error.
531
601
  // Never the display detail, which is clipped and can carry a model name.
532
602
  export function qwenErrorText(out) {
@@ -545,16 +615,29 @@ export function qwenErrorText(out) {
545
615
  return parts.join('\n');
546
616
  }
547
617
 
618
+ // Only native failure fields are authoritative. A successful model answer
619
+ // mentioning login, quota or permissions must never become an error signal.
620
+ export function claudeErrorText(out) {
621
+ const o = claudeResultObject(out);
622
+ if (!o) return '';
623
+ const parts = Array.isArray(o.errors) ? o.errors.filter(e => typeof e === 'string') : [];
624
+ if (o.is_error === true && typeof o.result === 'string') parts.push(o.result);
625
+ if (Number.isInteger(o.api_error_status)) parts.push(`api_error_status=${o.api_error_status}`);
626
+ return parts.join('\n');
627
+ }
628
+
548
629
  function authoritativeBlob(lane, out, err, detail, rc) {
630
+ if (lane === 'claude') return `${claudeErrorText(out)}\n${err || ''}`;
549
631
  if (lane === 'codex') return `${codexErrorEventsText(out)}\n${err || ''}`;
550
632
  if (lane === 'agy') return `${agyResultFieldsText(out)}\n${err || ''}`;
551
- if (lane === 'hermes') return rc !== 0 ? String(err || '') : '';
633
+ if (lane === 'hermes') return rc !== 0 ? `${err || ''}\n${hermesMismatchLine(out, { vendorShape: true }) || ''}` : '';
552
634
  if (lane === 'qwen') return `${qwenErrorText(out)}\n${err || ''}`;
553
635
  return `${detail || ''}\n${err || ''}`; // grok: no auth, quota or rejected signal is defined
554
636
  }
555
637
 
556
638
  export function sigAuth(lane, blob) {
557
639
  const b = blob.toLowerCase();
640
+ if (lane === 'claude') return /authentication_error|permission_error|invalid api key|not logged in|please run \/login|failed to authenticate|oauth (?:token|session).*expired|api_error_status=(?:401|403)|api error: (?:401|403)\b/.test(b);
558
641
  if (lane === 'qwen') return b.includes('missing api key');
559
642
  if (lane === 'agy') return b.includes('you are not logged into antigravity') || b.includes('not authenticated');
560
643
  return false; // codex, grok, hermes: no documented native auth signal
@@ -562,6 +645,7 @@ export function sigAuth(lane, blob) {
562
645
 
563
646
  export function sigQuota(lane, blob) {
564
647
  const b = blob.toLowerCase();
648
+ if (lane === 'claude') return /rate_limit_error|rate limit|insufficient credits|credit balance|you['’]ve hit your limit|api_error_status=429|api error: 429\b/.test(b);
565
649
  if (lane === 'qwen') return b.includes('[api error: 402') || b.includes('requires more credits') || blob.includes(' 429') || b.includes('rate limit');
566
650
  if (lane === 'codex') return blob.includes('usage_limit_exceeded') || b.includes("you've hit your usage limit");
567
651
  if (lane === 'hermes') return hermesQuotaText(b);
@@ -570,15 +654,17 @@ export function sigQuota(lane, blob) {
570
654
 
571
655
  export function sigRejected(lane, blob) {
572
656
  const b = blob.toLowerCase();
657
+ if (lane === 'claude') return /invalid_request_error|unknown option|invalid (?:model|effort)|api_error_status=400|api error: 400\b/.test(b);
573
658
  if (lane === 'qwen') return b.includes('[api error: 400') || b.includes('no endpoints found') || b.includes('failed to parse grammar');
574
659
  if (lane === 'codex') return blob.includes('invalid_request_error');
575
- if (lane === 'hermes') return b.includes('toolset');
660
+ if (lane === 'hermes') return hermesArgError(b) || b.includes('toolset') || hermesMismatchLine(blob) !== null;
576
661
  return false;
577
662
  }
578
663
 
579
664
  // Which judge reasons mean the lane never reached a trustworthy finish, as
580
665
  // opposed to finishing cleanly with nothing in it (class empty).
581
666
  const CUT_SHORT_REASONS = {
667
+ claude: new Set(['not_json', 'not_result', 'bad_subtype', 'is_error', 'bad_stop_reason', 'result_not_string', 'error_message_not_string']),
582
668
  grok: new Set(['not_json', 'bad_stop_reason']),
583
669
  codex: new Set(['no_terminal_event']),
584
670
  agy: new Set(['no_terminal_event', 'bad_last_event']),
@@ -704,6 +790,10 @@ export function refusedGrok(out, { root = process.env.CLI_RUN_GROK_SESSIONS_ROOT
704
790
 
705
791
  export function countRefused(lane, out, err, opts) {
706
792
  try {
793
+ if (lane === 'claude') {
794
+ const pd = claudeResultObject(out)?.permission_denials;
795
+ return Array.isArray(pd) ? pd.length : null;
796
+ }
707
797
  if (lane === 'qwen') return refusedQwen(out);
708
798
  if (lane === 'agy') return refusedAgy(out);
709
799
  if (lane === 'codex') return refusedCodex(out, err);
@@ -752,6 +842,7 @@ const FIX = {
752
842
  auth: 'set the credential the message above names (its environment variable, or the lane\'s own login command), then rerun',
753
843
  quota: 'switch to another lane, or wait for the reset time if the message gave one',
754
844
  rejected: 'correct the model id, flag or request the upstream message names',
845
+ mismatch: 'pair the model with a provider that serves it: run `hermes model`, or pin both with --provider and --model (or "defaults": {"hermes": {"provider": ..., "model": ...}} in bin/lanes.json)',
755
846
  refused: 'adjust the hook or deny rule named above, or give this lane the tool it needs',
756
847
  cut_short: 'rerun once; if it recurs, run without --quiet and read the lane\'s stderr on the terminal',
757
848
  empty: 'rerun once, or use another lane',
@@ -769,7 +860,11 @@ export function problemAndFix(lane, cls, { out = '', err = '', detail = '', refu
769
860
  switch (cls) {
770
861
  case 'auth': return { problem: `${tag} auth: ${cause || d || 'missing or invalid credentials'}`, fix: FIX.auth };
771
862
  case 'quota': return { problem: `${tag} quota: ${cause || d || 'rate limit or credits exhausted'}`, fix: FIX.quota };
772
- case 'rejected': return { problem: `${tag} rejected: ${cause || d || 'the upstream rejected the request'}`, fix: FIX.rejected };
863
+ case 'rejected': {
864
+ const mismatch = lane === 'hermes' ? hermesMismatchLine(err) || hermesMismatchLine(out, { vendorShape: true }) : null;
865
+ if (mismatch) return { problem: `${tag} rejected: model/provider mismatch: ${redact(mismatch)}`, fix: FIX.mismatch };
866
+ return { problem: `${tag} rejected: ${cause || d || 'the upstream rejected the request'}`, fix: FIX.rejected };
867
+ }
773
868
  case 'refused': return { problem: `${tag} refused: ${denial() || 'a hook or deny rule blocked the call'}`, fix: FIX.refused };
774
869
  case 'cut_short': return { problem: `${tag} cut short: ${d || 'no terminal success event, cause not identifiable'}`, fix: FIX.cut_short };
775
870
  case 'empty': return { problem: `${tag} empty: ${d || 'completed but delivered nothing'}`, fix: FIX.empty };
@@ -801,6 +896,7 @@ export function classifyRun(lane, { rc = 0, out = '', err = '', reason = '', det
801
896
  else {
802
897
  const blob = authoritativeBlob(lane, out, err, detail, rc);
803
898
  if (sigAuth(lane, blob)) cls = 'auth';
899
+ else if (lane === 'hermes' && hermesMismatchLine(blob) !== null) cls = 'rejected'; // before quota: see judgeHermes
804
900
  else if (sigQuota(lane, blob)) cls = 'quota';
805
901
  else if (sigRejected(lane, blob)) cls = 'rejected';
806
902
  else if (Number.isInteger(refused) && refused > 0) cls = 'refused';
@@ -1080,15 +1176,19 @@ export function laneConfig(here = dirname(fileURLToPath(import.meta.url))) {
1080
1176
  for (const [lane, d] of Object.entries(j.defaults)) {
1081
1177
  if (!LANES.includes(lane)) return null;
1082
1178
  if (!d || typeof d !== 'object' || Array.isArray(d)) return null;
1083
- const { model, effort, ...rest } = d;
1179
+ const { model, effort, provider, ...rest } = d;
1084
1180
  if (Object.keys(rest).length) return null;
1085
1181
  if (model !== undefined && badRouteValue('model', model)) return null;
1182
+ if (provider !== undefined) {
1183
+ if (badRouteValue('provider', provider)) return null;
1184
+ if (!LANE_FLAGS[lane] || !LANE_FLAGS[lane].provider) return null; // only hermes routes by provider
1185
+ }
1086
1186
  if (effort !== undefined) {
1087
1187
  if (badRouteValue('effort', effort)) return null;
1088
1188
  if (effort.startsWith('auto') && effort !== 'auto') return null;
1089
1189
  if (!LANE_FLAGS[lane] || !LANE_FLAGS[lane].effort) return null; // a lane with no reasoning flag cannot have one pinned
1090
1190
  }
1091
- defaults[lane] = { model: model ?? null, effort: effort ?? null };
1191
+ defaults[lane] = { model: model ?? null, effort: effort ?? null, provider: provider ?? null };
1092
1192
  }
1093
1193
  }
1094
1194
  return { enabled: j.enabled, defaults };
@@ -1109,20 +1209,22 @@ export function resolveRoute(lane, opts, defaults) {
1109
1209
  const d = (defaults && defaults[lane]) || {};
1110
1210
  const model = opts.model ?? d.model ?? null;
1111
1211
  const effort = opts.effort ?? d.effort ?? null;
1212
+ const provider = opts.provider ?? d.provider ?? null;
1112
1213
  const src = (flag, def) => (flag != null ? 'flag' : def != null ? 'lanes.json' : 'lane_default');
1113
- return { model, effort, model_source: src(opts.model, d.model), effort_source: src(opts.effort, d.effort) };
1214
+ return { model, effort, provider, model_source: src(opts.model, d.model), effort_source: src(opts.effort, d.effort), provider_source: src(opts.provider, d.provider) };
1114
1215
  }
1115
1216
 
1116
1217
  function usage(msg) {
1117
1218
  if (msg) console.error('cli-run: ' + msg);
1118
1219
  console.error(`usage: cli-run <${LANES.join('|')}> "<prompt>" [--brief FILE] [--timeout SECS] [--quiet]
1119
- [--model ID] [--effort LEVEL] [--expect-file PATH] [--expect-json]
1220
+ [--model ID] [--effort LEVEL] [--provider ID] [--expect-file PATH] [--expect-json]
1120
1221
  cli-run codex --audit "<prompt>" read-only sandbox (audit shape)
1121
1222
  cli-run qwen [--safe-mode] "<prompt>" qwen-only flag
1122
1223
  cli-run --doctor [--run] enabled lanes, binaries, and the route each one is pinned to
1123
1224
 
1124
1225
  --model / --effort pin what a lane runs with, instead of letting it inherit its
1125
1226
  own config. Every lane takes --model; every lane except qwen takes --effort.
1227
+ Hermes also takes --provider; it needs --model, a defaults model, or HERMES_INFERENCE_MODEL.
1126
1228
  Levels are the vendor's own (agy low|medium|high, hermes none|minimal|...): an
1127
1229
  unknown level is rejected by the lane, and reported by class (codex: rejected, 16).
1128
1230
  Exit codes: 0 ok, 10 empty, 11 no output, 12 timeout, 13 unavailable, 14 auth,
@@ -1163,7 +1265,8 @@ export async function doctor(run, { here = dirname(fileURLToPath(import.meta.url
1163
1265
  const { enabled, defaults } = cfg;
1164
1266
  let bad = 0;
1165
1267
  if (!compact) console.log(`doctor: ${enabled.length} enabled lane(s): ${enabled.join(', ') || 'none'}`);
1166
- if (!compact && primary && !enabled.includes(primary)) console.log(` note: ${primary} is the main agent and is not an executable lane`);
1268
+ const primaryLane = primary === 'claude-code' ? 'claude' : primary;
1269
+ if (!compact && primary && !enabled.includes(primaryLane)) console.log(` note: ${primary} is the main agent and is not an executable lane`);
1167
1270
  if (!enabled.length) {
1168
1271
  if (compact) {
1169
1272
  console.log('doctor: no executable lanes enabled.');
@@ -1179,7 +1282,8 @@ export async function doctor(run, { here = dirname(fileURLToPath(import.meta.url
1179
1282
  const d = defaults[lane] || {};
1180
1283
  // A disabled lane has no route worth reporting; saying "not pinned" there
1181
1284
  // reads as a finding about a lane that is not going to run.
1182
- const route = !on ? '' : d.effort === 'auto' ? `route ${d.model || 'lane default'}/auto (sized per call)` : d.model || d.effort ? `route ${d.model || 'lane default'}/${d.effort || 'lane default'}` : 'route not pinned (inherits the lane\'s own config)';
1285
+ const model = `${d.provider ? d.provider + ':' : ''}${d.model || 'lane default'}`;
1286
+ const route = !on ? '' : d.effort === 'auto' ? `route ${model}/auto (sized per call)` : d.model || d.effort || d.provider ? `route ${model}/${d.effort || 'lane default'}` : 'route not pinned (inherits the lane\'s own config)';
1183
1287
  let line = ` ${lane.padEnd(7)} ${on ? 'enabled ' : 'disabled'} ${bin ? 'binary ok' : 'binary MISSING'}${route ? ' ' + route : ''}`;
1184
1288
  if (on && !bin) bad++;
1185
1289
  if (on && bin && run) {
@@ -1188,6 +1292,8 @@ export async function doctor(run, { here = dirname(fileURLToPath(import.meta.url
1188
1292
  if (rc !== OK) bad++;
1189
1293
  }
1190
1294
  console.log(compact ? ` ${lane}: ${bin ? 'present' : 'MISSING'}` : line);
1295
+ if (on && lane === 'hermes' && d.model && !d.provider) console.log(' note: model pinned with no provider: Hermes sends it to its default provider. A model that provider does not serve fails with HTTP 400; pin "provider" beside "model".');
1296
+ if (on && lane === 'hermes' && d.provider && !d.model) console.log(' note: provider pinned with no model: every run without --model will be refused unless HERMES_INFERENCE_MODEL supplies a model; pin "model" beside "provider".');
1191
1297
  }
1192
1298
  console.log(bad ? `doctor: ${bad} problem(s)` : 'doctor: all enabled lanes ' + (run ? 'answered' : 'present'));
1193
1299
  if (!compact) {
@@ -1236,10 +1342,10 @@ export function checkContracts(opts, text, before) {
1236
1342
  }
1237
1343
 
1238
1344
  export async function main(argv) {
1239
- const VALUE = new Set(['--brief', '--timeout', '--model', '--effort', '--expect-file']);
1345
+ const VALUE = new Set(['--brief', '--timeout', '--model', '--effort', '--provider', '--expect-file']);
1240
1346
  const BOOL = new Set(['--quiet', '--audit', '--safe-mode', '--doctor', '--run', '--expect-json']);
1241
1347
  const args = [...argv];
1242
- const opts = { timeout: 900, quiet: false, audit: false, model: null, effort: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
1348
+ const opts = { timeout: 900, quiet: false, audit: false, model: null, effort: null, provider: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
1243
1349
  const positional = [];
1244
1350
  while (args.length) {
1245
1351
  const a = args.shift();
@@ -1250,6 +1356,7 @@ export async function main(argv) {
1250
1356
  else if (a === '--timeout') opts.timeout = Number(v);
1251
1357
  else if (a === '--expect-file') opts.expectFile = v;
1252
1358
  else if (a === '--effort') opts.effort = v;
1359
+ else if (a === '--provider') opts.provider = v;
1253
1360
  else opts.model = v;
1254
1361
  } else if (BOOL.has(a)) {
1255
1362
  if (a === '--quiet') opts.quiet = true;
@@ -1283,7 +1390,7 @@ export async function main(argv) {
1283
1390
  if (!Number.isFinite(opts.timeout) || opts.timeout <= 0) return usage('--timeout must be a positive number of seconds');
1284
1391
  if (opts.audit && lane !== 'codex') return usage('--audit is codex-only');
1285
1392
  if (opts.safeMode && lane !== 'qwen') return usage('--safe-mode is qwen-only');
1286
- for (const [kind, v] of [['model', opts.model], ['effort', opts.effort]]) {
1393
+ for (const [kind, v] of [['model', opts.model], ['effort', opts.effort], ['provider', opts.provider]]) {
1287
1394
  if (v == null) continue;
1288
1395
  const bad = badRouteValue(kind, v);
1289
1396
  if (bad) return usage(bad);
@@ -1292,6 +1399,7 @@ export async function main(argv) {
1292
1399
  // qwen has no reasoning flag. Dropping --effort silently would leave the caller
1293
1400
  // believing a route that never happened, which is the defect this feature fixes.
1294
1401
  if (opts.effort && !(LANE_FLAGS[lane] && LANE_FLAGS[lane].effort)) return usage(`${lane} has no reasoning-effort flag; --effort is not available on this lane`);
1402
+ if (opts.provider && !LANE_FLAGS[lane].provider) return usage(`${lane} has no provider flag; --provider is hermes-only`);
1295
1403
 
1296
1404
  const digest = createHash('sha256').update(prompt).digest('hex').slice(0, 12);
1297
1405
  const base = { lane, prompt_sha256_12: digest, prompt_chars: prompt.length };
@@ -1301,11 +1409,15 @@ export async function main(argv) {
1301
1409
  // gap this feature exists to close. A malformed lanes.json has no usable
1302
1410
  // defaults, so the flags stand alone and say so.
1303
1411
  const route = resolveRoute(lane, opts, cfg === null ? {} : cfg.defaults);
1412
+ if (route.provider && !route.model && !(process.env.HERMES_INFERENCE_MODEL || '').trim()) return usage('--provider <p> needs a model: pass --model, or set "model" beside "provider" in lanes.json "defaults"');
1304
1413
  const sizing = resolveAutoEffort(lane, route.effort, prompt, opts.audit);
1305
1414
  opts.model = route.model;
1306
1415
  opts.effort = sizing.resolved;
1416
+ opts.provider = route.provider;
1307
1417
  Object.assign(base, {
1308
1418
  model_requested: route.model,
1419
+ provider_requested: route.provider,
1420
+ provider_source: route.provider_source,
1309
1421
  effort_requested: route.effort,
1310
1422
  model_source: route.model_source,
1311
1423
  effort_source: route.effort_source,
@@ -1384,10 +1496,11 @@ export async function main(argv) {
1384
1496
  }
1385
1497
  const code = cls === 'interrupted' ? 128 + (r.interrupted === 'SIGINT' ? 2 : 15) : CLASS_CODES[cls];
1386
1498
  if (text && code === OK) process.stdout.write(text + '\n');
1387
- const routeNote = route.model || route.effort ? `${route.model || 'lane default'}/${route.effort || 'lane default'}` : 'lane default';
1499
+ const routeNote = (route.provider ? `${route.provider}:` : '') + (route.model || route.effort ? `${route.model || 'lane default'}/${route.effort || 'lane default'}` : 'lane default');
1388
1500
  if (!opts.quiet) {
1389
1501
  console.error(`cli-run[${lane}] ${verdict} rc=${code} class=${cls} refused=${refused === null ? 'null' : refused} ${r.seconds.toFixed(1)}s raw=${r.outBytes || 0}B route=${routeNote} :: ${redact(detail)}`);
1390
1502
  let authoritative = null;
1503
+ if (lane === 'claude' && ['auth', 'quota', 'rejected'].includes(cls)) authoritative = stderrHead(claudeErrorText(out) || err);
1391
1504
  if (lane === 'codex' && (cls === 'quota' || cls === 'rejected')) authoritative = codexPrimaryError(out);
1392
1505
  const pf = problemAndFix(lane, cls, { out, err, detail, refused, authoritative });
1393
1506
  if (pf.problem) console.error('cli-run problem: ' + redact(pf.problem).slice(0, 600));
package/bin/cli.js CHANGED
@@ -498,6 +498,7 @@ async function main() {
498
498
  for (const path of preview.written) console.log(' would write ' + path);
499
499
  for (const path of [...preview.skipped, ...preview.conflicts, ...preview.unverifiable, ...preview.docsConflict, ...preview.docsUnverifiable]) console.log(' would keep ' + path);
500
500
  for (const path of preview.backups) console.log(' would back up ' + path);
501
+ if (preview.lanesWithheld.length) console.log(` lanes withheld: ${preview.lanesWithheld.join(', ')} (the kept bin/cli-run.mjs predates them; re-run with --upgrade-runtime to enable)`);
501
502
  }
502
503
  console.log('\n--dry: nothing written.');
503
504
  rl && rl.close();
@@ -548,9 +549,9 @@ async function main() {
548
549
  .concat(JSON.stringify(prev.roles || {}) !== JSON.stringify(plannedManifest.roles || {}) ? ['roles'] : [])
549
550
  : [];
550
551
 
551
- let written, skipped, upgraded, conflicts, unverifiable, docsUpdated, docsConflict, docsUnverifiable, docsRenamed;
552
+ let written, skipped, upgraded, conflicts, unverifiable, lanesWithheld, docsUpdated, docsConflict, docsUnverifiable, docsRenamed;
552
553
  try {
553
- ({ written, skipped, upgraded, conflicts, unverifiable, docsUpdated, docsConflict, docsUnverifiable, docsRenamed } = writeFiles(files, { dir, project, force: flag('force'), upgradeRuntime: flag('upgrade-runtime'), updateDocs: flag('update-docs'), prevManifest: prev, backupExisting: applySnippets, onBackup: (path) => console.log(' backup ' + path) }));
554
+ ({ written, skipped, upgraded, conflicts, unverifiable, lanesWithheld, docsUpdated, docsConflict, docsUnverifiable, docsRenamed } = writeFiles(files, { dir, project, force: flag('force'), upgradeRuntime: flag('upgrade-runtime'), updateDocs: flag('update-docs'), prevManifest: prev, backupExisting: applySnippets, onBackup: (path) => console.log(' backup ' + path) }));
554
555
  } catch (e) {
555
556
  if (e && e.code === 'PREFLIGHT') bad(e.message);
556
557
  throw e;
@@ -564,10 +565,11 @@ async function main() {
564
565
  if (prev) {
565
566
  if (changed.length) {
566
567
  console.log(` selection changed: ${changed.join(', ')}`);
567
- console.log(` applied: ${ownedWritten.join(', ') || 'nothing'} (machine-owned files are always rewritten, so the new lanes are live)`);
568
+ console.log(` applied: ${ownedWritten.join(', ') || 'nothing'} (machine-owned files are always rewritten)`);
568
569
  if (changed.includes('roles')) console.log(' the role assignment changed; MANIFEST.json and aunx route are current. Any kept documents may still carry the previous assignment.');
569
570
  } else console.log(' selection identical.');
570
571
  }
572
+ if (lanesWithheld.length) console.log(` lanes withheld: ${lanesWithheld.join(', ')} (the kept bin/cli-run.mjs predates them; re-run with --upgrade-runtime to enable)`);
571
573
  if (upgraded.length) console.log(` runtime upgraded: ${upgraded.join(', ')} ${flag('upgrade-runtime') ? '(--upgrade-runtime: replaced whether or not you had edited them)' : '(each installed copy matched the hash of a previous run, so nobody had edited it)'}`);
572
574
  if (conflicts.length) {
573
575
  console.log(` runtime CONFLICT, kept: ${conflicts.join(', ')}`);
package/docs/catalog.md CHANGED
@@ -16,14 +16,15 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
16
16
 
17
17
  - **Kind:** agent-cli · **Billing:** subscription · **Level:** 1+
18
18
  - **What it is:** Anthropic's terminal coding agent; its subagents load the project rules file
19
- - **Install:** `npm install -g @anthropic-ai/claude-code@2.1.226`
19
+ - **Install:** `npm install -g @anthropic-ai/claude-code@2.1.285`
20
20
  - **Sign in:** run `claude` once and sign in with your Anthropic account
21
21
  - **Reads rules from:** `CLAUDE.md` · subagents in `.claude/agents/`
22
+ - **cli-run lane:** yes
22
23
  - **Plans:**
23
24
  - Claude Pro (base headroom, checked 2026-09-12; tier models: unverified): https://claude.com/pricing
24
25
  - Claude Max 5x (high headroom, checked 2026-09-12; tier models: unverified): https://claude.com/pricing
25
26
  - Claude Max 20x (max headroom, checked 2026-09-12; tier models: unverified): https://claude.com/pricing
26
- - **Built against:** 2.1.226 (the same number the npm pin uses)
27
+ - **Built against:** 2.1.285 (the same number the npm pin uses)
27
28
 
28
29
  **Capability facts.** An unverified value needs a current capability check before use. Role assignments come from the selected stack, using these facts.
29
30
 
@@ -34,7 +35,7 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
34
35
  | `billing` | subscription |
35
36
  | `pricing` | unverified |
36
37
  | `headless` | yes |
37
- | `cliRun` | no |
38
+ | `cliRun` | yes |
38
39
  | `writesFiles` | yes |
39
40
  | `readOnlyMode` | no |
40
41
  | `liveWeb` | yes |
@@ -60,6 +60,8 @@ aunx cli-run --doctor
60
60
  # Direct form: node bin/cli-run.mjs --doctor
61
61
  ```
62
62
 
63
+ Hermes also takes `--provider '<provider-id>'` (or `"provider"` in its `lanes.json` defaults), always together with a model, because one Hermes install can reach several providers and a model sent to the wrong one fails with HTTP 400. `--doctor` notes a Hermes model pinned without a provider.
64
+
63
65
  Explicit flags override defaults in `bin/lanes.json`. Without either, the vendor CLI uses its own configuration. The runner records what was requested and the source of each request: `flag`, `lanes.json` or `lane_default`. These fields describe requested settings; the vendor's own reporting is the place to verify the actual model used.
64
66
 
65
67
  ## Share facts once, then scope each worker
@@ -6,6 +6,8 @@ This page summarizes completed reviews recorded in the [changelog](../CHANGELOG.
6
6
 
7
7
  | Release | Review recorded | What the review found | What changed and where it is checked |
8
8
  |---|---|---|---|
9
+ | 1.0.7 | Review of external PR 48 (native Claude Code lane) by a second model family, then one audit of the fix diff | An upgrade that kept an edited older runner could disable every lane; an expired Claude sign-in read as cut short; a bare HTTP number in an error could pick the wrong class | Kept runners withhold lanes they do not list; expired-session and status-anchored Claude signals; captured 2.1.285 fixtures; `test/install.test.js`, `test/classify.test.js`, `test/cli.test.js` |
10
+ | 1.0.6 | Pre-release review of the hermes provider route by a second model family, one round | Hermes stdout could print raw terminal control characters in a failure line; agent prose on stdout and an unrelated quota word could be misclassified; one provider test could not fail; route flags placed after `-z` broke every pinned hermes run | Control characters stripped, stdout matched only in the vendor's error shape, the mismatch checked before quota, route flags moved before `-z`; `test/classify.test.js`, `test/judges.test.js`, `test/cli.test.js` |
9
11
  | 1.0.2 | Independent review of 1.0.1 by a second model family, then an audit of its fix | A detached check descendant outlived the timeout while the notes said descendants stop; a symlinked-manifest refusal named no file or fix; refusal paths could carry terminal escapes; a directory got symlink advice | The protocol states the process-group limit; refusals name the path, escape control characters and give advice by file type; `test/security-messages.test.js` |
10
12
  | 1.0.1 | Codex Security scan and regression-backed remediation | A metrics summary executed project code; manifest reads and check descendants needed bounds; weekly jobs inherited gateway credentials | Packaged metrics dispatch, bounded file readers, check process-tree cleanup, stdin-only gateway probe and separate audit environment; `test/security-cli.test.js`, `test/security-vm.test.js`, `test/security-pins.test.js` |
11
13
  | 0.1.0 | Initial review and follow-up round, with a second model-family review | Paths could escape the write roots, partial writes could remain, malformed arguments could proceed, and empty results could appear successful | Containment preflight, exclusive writes with rollback, strict flag parsing and vendor-specific result checks; `test/install.test.js`, `test/cli.test.js`, `test/judges.test.js` |
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "1.0.5",
3
+ "version": "1.0.7",
4
4
  "description": "Model router for AI coding agents: installs routing rules, 8 subagents, hooks and a CLI runner so your AI picks model and effort per task and saves tokens",
5
5
  "type": "module",
6
6
  "bin": {
package/proof/README.md CHANGED
@@ -2,14 +2,14 @@
2
2
 
3
3
  Generated from [results.json](results.json). Each figure has a method, sample size, measurement date and expiry. Run the scripts on your own machine to compare.
4
4
 
5
- Environment: Node v22.22.3, darwin arm64. Timing varies with startup caches and other work on the machine. Synthetic cases show what those fixtures exercise.
5
+ Environment: Node v22.23.2, linux x64. Timing varies with startup caches and other work on the machine. Synthetic cases show what those fixtures exercise.
6
6
 
7
7
  | Measurement | Result | Sample size | Measured | Expires | Reproduce |
8
8
  |---|---|---|---|---|---|
9
- | Dry install wall time | 42.22 ms median | 7 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/install-time.js) |
10
- | Lane runner overhead | 63.74 ms median difference | 7 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/runner-overhead.js) |
11
- | Empty results flagged | 10 fixtures rejected | 10 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/missing-results.js) |
12
- | Acceptance failures blocked | 4 fixtures rejected | 4 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/check-gate.js) |
9
+ | Dry install wall time | 61.59 ms median | 7 | 2026-09-28 | 2026-10-12 | [script](../proof/scripts/install-time.js) |
10
+ | Lane runner overhead | 60.27 ms median difference | 7 | 2026-09-28 | 2026-10-12 | [script](../proof/scripts/runner-overhead.js) |
11
+ | Empty results flagged | 10 fixtures rejected | 10 | 2026-09-28 | 2026-10-12 | [script](../proof/scripts/missing-results.js) |
12
+ | Acceptance failures blocked | 4 fixtures rejected | 4 | 2026-09-28 | 2026-10-12 | [script](../proof/scripts/check-gate.js) |
13
13
 
14
14
  ## Run the proof scripts
15
15
 
@@ -40,7 +40,7 @@ The [recording script](scripts/record-gate.js) captures real command output into
40
40
 
41
41
  Spawn a fresh Node installer process per sample; level 2, Claude Code + Codex, no companions, --dry. Includes Node startup and planning; writes no install files. Isolated home and PATH, no real vendors.
42
42
 
43
- Kind: reproducible local measurement. Sample size: 7. Measured: 2026-09-27. Expires: 2026-10-11.
43
+ Kind: reproducible local measurement. Sample size: 7. Measured: 2026-09-28. Expires: 2026-10-12.
44
44
 
45
45
  Source: [proof/scripts/install-time.js](../proof/scripts/install-time.js).
46
46
 
@@ -48,7 +48,7 @@ Source: [proof/scripts/install-time.js](../proof/scripts/install-time.js).
48
48
 
49
49
  Paired fresh processes: direct Node stub versus cli-run hermes with the same stub. Alternates pair order. Includes wrapper startup, validation and local log writes; excludes vendor/network/model time.
50
50
 
51
- Kind: reproducible local measurement. Sample size: 7. Measured: 2026-09-27. Expires: 2026-10-11.
51
+ Kind: reproducible local measurement. Sample size: 7. Measured: 2026-09-28. Expires: 2026-10-12.
52
52
 
53
53
  Source: [proof/scripts/runner-overhead.js](../proof/scripts/runner-overhead.js).
54
54
 
@@ -56,7 +56,7 @@ Source: [proof/scripts/runner-overhead.js](../proof/scripts/runner-overhead.js).
56
56
 
57
57
  Run cli-run against an exit-0 stub for every supported lane, once with empty stdout and once with an empty native final result. Count exit 10/11 only. A successful Hermes response is the positive control. Synthetic fixtures measure these shapes only.
58
58
 
59
- Kind: reproducible local measurement. Sample size: 10. Measured: 2026-09-27. Expires: 2026-10-11.
59
+ Kind: reproducible local measurement. Sample size: 10. Measured: 2026-09-28. Expires: 2026-10-12.
60
60
 
61
61
  Source: [proof/scripts/missing-results.js](../proof/scripts/missing-results.js).
62
62
 
@@ -64,7 +64,7 @@ Source: [proof/scripts/missing-results.js](../proof/scripts/missing-results.js).
64
64
 
65
65
  Run aunx checks run against nonzero, manual, missing-program and timeout fixtures. Each must exit 1; a passing command must exit 0. This is a local command gate, activated by the user in their release sequence.
66
66
 
67
- Kind: reproducible local measurement. Sample size: 4. Measured: 2026-09-27. Expires: 2026-10-11.
67
+ Kind: reproducible local measurement. Sample size: 4. Measured: 2026-09-28. Expires: 2026-10-12.
68
68
 
69
69
  Source: [proof/scripts/check-gate.js](../proof/scripts/check-gate.js).
70
70
 
@@ -1,78 +1,78 @@
1
1
  {
2
2
  "schemaVersion": 1,
3
3
  "environment": {
4
- "node": "v22.22.3",
5
- "platform": "darwin",
6
- "arch": "arm64"
4
+ "node": "v22.23.2",
5
+ "platform": "linux",
6
+ "arch": "x64"
7
7
  },
8
8
  "entries": [
9
9
  {
10
10
  "id": "install-dry-ms",
11
11
  "label": "Dry install wall time",
12
12
  "kind": "reproducible",
13
- "value": 42.220875,
13
+ "value": 61.594306,
14
14
  "unit": "ms median",
15
- "measuredAt": "2026-09-27",
15
+ "measuredAt": "2026-09-28",
16
16
  "method": "Spawn a fresh Node installer process per sample; level 2, Claude Code + Codex, no companions, --dry. Includes Node startup and planning; writes no install files. Isolated home and PATH, no real vendors.",
17
17
  "sampleSize": 7,
18
18
  "script": "proof/scripts/install-time.js",
19
- "expiresAt": "2026-10-11",
19
+ "expiresAt": "2026-10-12",
20
20
  "samples": [
21
- 51.577042,
22
- 48.655542,
23
- 46.568834,
24
- 42.220875,
25
- 41.573667,
26
- 39.881333,
27
- 40.983417
21
+ 72.998866,
22
+ 61.594306,
23
+ 63.307248,
24
+ 60.59507,
25
+ 61.218156,
26
+ 66.993599,
27
+ 61.070192
28
28
  ]
29
29
  },
30
30
  {
31
31
  "id": "runner-overhead-ms",
32
32
  "label": "Lane runner overhead",
33
33
  "kind": "reproducible",
34
- "value": 63.742250999999996,
34
+ "value": 60.27329300000001,
35
35
  "unit": "ms median difference",
36
- "measuredAt": "2026-09-27",
36
+ "measuredAt": "2026-09-28",
37
37
  "method": "Paired fresh processes: direct Node stub versus cli-run hermes with the same stub. Alternates pair order. Includes wrapper startup, validation and local log writes; excludes vendor/network/model time.",
38
38
  "sampleSize": 7,
39
39
  "script": "proof/scripts/runner-overhead.js",
40
- "expiresAt": "2026-10-11",
40
+ "expiresAt": "2026-10-12",
41
41
  "pairs": [
42
42
  {
43
- "directMs": 26.925375,
44
- "wrappedMs": 409.562,
45
- "overheadMs": 382.63662500000004
43
+ "directMs": 27.602021,
44
+ "wrappedMs": 89.562129,
45
+ "overheadMs": 61.960108
46
46
  },
47
47
  {
48
- "directMs": 28.554875,
49
- "wrappedMs": 95.06725,
50
- "overheadMs": 66.512375
48
+ "directMs": 26.453384,
49
+ "wrappedMs": 85.145791,
50
+ "overheadMs": 58.692407
51
51
  },
52
52
  {
53
- "directMs": 28.981666,
54
- "wrappedMs": 94.610209,
55
- "overheadMs": 65.628543
53
+ "directMs": 26.104555,
54
+ "wrappedMs": 91.019659,
55
+ "overheadMs": 64.915104
56
56
  },
57
57
  {
58
- "directMs": 26.295583,
59
- "wrappedMs": 86.161958,
60
- "overheadMs": 59.866375
58
+ "directMs": 24.351936,
59
+ "wrappedMs": 84.625229,
60
+ "overheadMs": 60.27329300000001
61
61
  },
62
62
  {
63
- "directMs": 28.71575,
64
- "wrappedMs": 82.585583,
65
- "overheadMs": 53.869833
63
+ "directMs": 27.307321,
64
+ "wrappedMs": 86.025276,
65
+ "overheadMs": 58.717955
66
66
  },
67
67
  {
68
- "directMs": 26.668791,
69
- "wrappedMs": 90.411042,
70
- "overheadMs": 63.742250999999996
68
+ "directMs": 26.203506,
69
+ "wrappedMs": 84.365333,
70
+ "overheadMs": 58.161827
71
71
  },
72
72
  {
73
- "directMs": 34.081333,
74
- "wrappedMs": 82.568041,
75
- "overheadMs": 48.48670799999999
73
+ "directMs": 26.328261,
74
+ "wrappedMs": 86.629325,
75
+ "overheadMs": 60.301064
76
76
  }
77
77
  ]
78
78
  },
@@ -82,11 +82,11 @@
82
82
  "kind": "reproducible",
83
83
  "value": 10,
84
84
  "unit": "fixtures rejected",
85
- "measuredAt": "2026-09-27",
85
+ "measuredAt": "2026-09-28",
86
86
  "method": "Run cli-run against an exit-0 stub for every supported lane, once with empty stdout and once with an empty native final result. Count exit 10/11 only. A successful Hermes response is the positive control. Synthetic fixtures measure these shapes only.",
87
87
  "sampleSize": 10,
88
88
  "script": "proof/scripts/missing-results.js",
89
- "expiresAt": "2026-10-11",
89
+ "expiresAt": "2026-10-12",
90
90
  "outcomes": [
91
91
  {
92
92
  "lane": "grok",
@@ -146,11 +146,11 @@
146
146
  "kind": "reproducible",
147
147
  "value": 4,
148
148
  "unit": "fixtures rejected",
149
- "measuredAt": "2026-09-27",
149
+ "measuredAt": "2026-09-28",
150
150
  "method": "Run aunx checks run against nonzero, manual, missing-program and timeout fixtures. Each must exit 1; a passing command must exit 0. This is a local command gate, activated by the user in their release sequence.",
151
151
  "sampleSize": 4,
152
152
  "script": "proof/scripts/check-gate.js",
153
- "expiresAt": "2026-10-11",
153
+ "expiresAt": "2026-10-12",
154
154
  "outcomes": [
155
155
  {
156
156
  "fixture": "nonzero",
package/src/aunx.js CHANGED
@@ -205,6 +205,9 @@ function stackSuggestion(assignment) {
205
205
  if (assignment.command) via = '`' + clean(assignment.command.startsWith('cli-run ') ? `aunx ${assignment.command}` : assignment.command) + '`';
206
206
  else if (assignment.agent) via = '`' + clean(assignment.agent) + '` on your main agent';
207
207
  else via = { 'main-agent': 'your main agent', local: 'your local runtime', manual: 'manual handoff', subagent: 'a subagent', 'cli-run': 'cli-run' }[assignment.via];
208
+ if (assignment.ai === 'claude-code' && assignment.via === 'cli-run' && assignment.command === 'cli-run claude' && assignment.preferredTransport === 'mcp') {
209
+ via = `connected Claude worker MCP when available; fallback ${via}`;
210
+ }
208
211
  return `Your stack: ${name}, via ${via || 'your selected tool'}${reason ? ` (${clean(reason)})` : ''}.`;
209
212
  }
210
213
 
package/src/catalog.js CHANGED
@@ -71,7 +71,7 @@ export const AIS = [
71
71
  billing: 'subscription',
72
72
  pricing: null, // UNVERIFIED: check your provider's current rate.
73
73
  headless: true,
74
- cliRun: false,
74
+ cliRun: true,
75
75
  writesFiles: true,
76
76
  readOnlyMode: false,
77
77
  liveWeb: true, // Source: templates/agents/claude-code/live-researcher.md grants WebSearch and WebFetch.
@@ -93,8 +93,8 @@ export const AIS = [
93
93
  bin: 'claude',
94
94
  summary: 'Anthropic\'s terminal coding agent; its subagents load the project rules file',
95
95
  minLevel: 1,
96
- install: { npm: '@anthropic-ai/claude-code', url: 'https://code.claude.com/docs/en/setup', pin: '2.1.226' },
97
- builtAgainst: '2.1.226',
96
+ install: { npm: '@anthropic-ai/claude-code', url: 'https://code.claude.com/docs/en/setup', pin: '2.1.285' },
97
+ builtAgainst: '2.1.285',
98
98
  auth: 'run `claude` once and sign in with your Anthropic account',
99
99
  // Positive-only (Q1): an author-machine probe of a working, authenticated
100
100
  // session (2026-09-27) still returned {"loggedIn":false} with exit 1, so a
package/src/install.js CHANGED
@@ -70,7 +70,7 @@ export function lanesTable(selected, plans = {}, primary = inferPrimary(selected
70
70
  a.facts.billing,
71
71
  summaryWithEvidence(a),
72
72
  Object.entries(roles).filter(([, role]) => role.ai === a.id).map(([id]) => id).join(', ') || 'none',
73
- a.facts.cliRun ? '`cli-run ' + a.id + '`' : a.bin ? '`' + a.bin + '`' : 'the app',
73
+ a.facts.cliRun ? '`cli-run ' + a.bin + '`' : a.bin ? '`' + a.bin + '`' : 'the app',
74
74
  plans[a.id] ? `${plans[a.id].name} (${plans[a.id].headroom} headroom)` : 'not stated'
75
75
  ]);
76
76
  return table(rows, ['AI', 'Lane', 'What it is', 'Assigned roles', 'Call it with', 'Plan']);
@@ -180,7 +180,7 @@ export function dirProblems(dir) {
180
180
  export function auditLane(selected, primary = selected[0]) {
181
181
  const { roles } = assignRoles({ selected, primary });
182
182
  const id = roles.review.ai ?? roles.bulk.ai ?? null;
183
- return selected.some(a => a.id === id && a.facts.cliRun) ? id : null;
183
+ return selected.find(a => a.id === id && a.facts.cliRun)?.bin || null;
184
184
  }
185
185
 
186
186
  function stackContext(selected, primary, detected = new Set()) {
@@ -558,7 +558,7 @@ function vars(opts) {
558
558
  const assignment = assignRoles({ selected, primary, detected: opts.detected, plans });
559
559
  const stack = stackContext(selected, primary, opts.detected);
560
560
  const enabled = selected.filter(a => a.facts.cliRun);
561
- const exampleLane = enabled[0]?.id || '<lane>';
561
+ const exampleLane = enabled[0]?.bin || '<lane>';
562
562
  const exampleDefaults = { model: '<model-id>', ...(LANE_FLAGS[exampleLane]?.effort ? { effort: 'high' } : {}) };
563
563
  const auditExample = enabled.find(a => a.facts.readOnlyMode);
564
564
  const fallbackNote = 'When no separate lane qualifies, your main agent carries the job at its stated tier. Independent review and local-only work require an eligible lane.';
@@ -634,8 +634,8 @@ function vars(opts) {
634
634
  EXAMPLE_LANE: exampleLane,
635
635
  EXAMPLE_EFFORT_FLAGS: LANE_FLAGS[exampleLane]?.effort ? ' --effort high' : '',
636
636
  EXAMPLE_AUDIT_LANE: auditExample?.id || '',
637
- EXAMPLE_AUDIT_BLOCK: auditExample ? `When reviewing with an available read-only mode, select its audit shape:\n\n\x60\x60\x60bash\naunx cli-run ${auditExample.id} --audit --brief REVIEW.md\nnode bin/cli-run.mjs ${auditExample.id} --audit --brief REVIEW.md\n\x60\x60\x60` : 'When reviewing, verify the chosen lane permissions and request review-only work.',
638
- EXAMPLE_LANES_JSON: JSON.stringify({ enabled: enabled.map(a => a.id), defaults: { [exampleLane]: exampleDefaults } }, null, 2),
637
+ EXAMPLE_AUDIT_BLOCK: auditExample ? `When reviewing with an available read-only mode, select its audit shape:\n\n\x60\x60\x60bash\naunx cli-run ${auditExample.bin} --audit --brief REVIEW.md\nnode bin/cli-run.mjs ${auditExample.bin} --audit --brief REVIEW.md\n\x60\x60\x60` : 'When reviewing, verify the chosen lane permissions and request review-only work.',
638
+ EXAMPLE_LANES_JSON: JSON.stringify({ enabled: enabled.map(a => a.bin), defaults: { [exampleLane]: exampleDefaults } }, null, 2),
639
639
  // Renders only when Qwen is actually selected: the sentence names a flag
640
640
  // that is a usage error on every other lane (C1).
641
641
  QWEN_SAFE_MODE_NOTE: selected.some(a => a.id === 'qwen') ? "When Qwen's safe mode is required, pass `--safe-mode` to that lane. " : '',
@@ -686,8 +686,8 @@ function vars(opts) {
686
686
  AUDIT_LANE: lane || 'none',
687
687
  // Enforced boundary per lane: codex has a read-only sandbox flag; the others
688
688
  // run with whatever their own config allows, and the script says so.
689
- AUDIT_LANE_FLAGS: selected.find(a => a.id === lane)?.facts.readOnlyMode ? '--audit' : '',
690
- AUDIT_LANE_BOUNDARY_NOTE: selected.find(a => a.id === lane)?.facts.readOnlyMode
689
+ AUDIT_LANE_FLAGS: selected.find(a => a.bin === lane)?.facts.readOnlyMode ? '--audit' : '',
690
+ AUDIT_LANE_BOUNDARY_NOTE: selected.find(a => a.bin === lane)?.facts.readOnlyMode
691
691
  ? `${lane} --audit, a read-only filesystem sandbox; commands and network follow the ${lane} config`
692
692
  : lane
693
693
  ? `${lane} offers no sandbox flag cli-run can pass, so the denied-actions list is instruction-level only and enforcement is whatever ${lane}'s own permission config allows`
@@ -714,7 +714,10 @@ function vars(opts) {
714
714
  LANES_TABLE: lanesTable(selected, plans, primary),
715
715
  PLAN_GUIDANCE: planGuidance(selected, plans),
716
716
  INSTALL_TABLE: installTable(selected),
717
- CLI_RUN_LANES: selected.filter((a) => a.facts.cliRun).map((a) => a.id).join(', ') || 'none selected',
717
+ CLI_RUN_LANES: selected.filter((a) => a.facts.cliRun).map((a) => a.bin).join(', ') || 'none selected',
718
+ CLAUDE_WORKER_TRANSPORT: Object.values(assignment.roles).some(role => role.ai === 'claude-code' && role.via === 'cli-run')
719
+ ? '- When assigning a Claude Code worker, prefer a connected Claude worker MCP service exposed in the host tool catalog. Call its tools directly from the host, follow its session and permission workflow, and preserve the task scope. See `CLI-RUN.md` for dispatch and fallback boundaries.\n'
720
+ : '',
718
721
  GATEWAY_MODELS: gatewayModels(selected, apis),
719
722
  ENV_NAMES: envNames(selected, apis).map((n) => '- `' + n + '`').join('\n'),
720
723
  ENV_EXPORTS: envNames(selected, apis).map((n) => n + '=').join('\n'),
@@ -820,10 +823,10 @@ export function planFiles(opts) {
820
823
  join('bin', 'lanes.json'),
821
824
  JSON.stringify(
822
825
  {
823
- enabled: selected.filter((a) => a.facts.cliRun).map((a) => a.id),
824
- defaults: Object.fromEntries((opts.effortAuto || []).map((lane) => [lane, { effort: 'auto' }])),
826
+ enabled: selected.filter((a) => a.facts.cliRun).map((a) => a.bin),
827
+ defaults: Object.fromEntries((opts.effortAuto || []).map((id) => [selected.find(a => a.id === id)?.bin || id, { effort: 'auto' }])),
825
828
  note: 'Lanes cli-run may call. Edit to enable or disable a lane. A lane not listed here exits 13 (unavailable).',
826
- defaultsNote: 'Pin what a lane runs with, so the route in your docs is the route that runs: "defaults": {"' + (selected.find(a => a.facts.cliRun)?.id || '<lane>') + '": ' + JSON.stringify({ model: '<model-id>', ...(LANE_FLAGS[selected.find(a => a.facts.cliRun)?.id]?.effort ? { effort: 'high' } : {}) }) + '}. Left empty, a lane inherits its own config file, which cli-run cannot see and does not guess. `--model` and `--effort` override this per call, and `--doctor` prints what each lane is pinned to. Every enabled lane takes a model; the runner reports which lanes support an effort flag.'
829
+ defaultsNote: 'Pin what a lane runs with, so the route in your docs is the route that runs: "defaults": {"' + (selected.find(a => a.facts.cliRun)?.bin || '<lane>') + '": ' + JSON.stringify({ model: '<model-id>', ...(LANE_FLAGS[selected.find(a => a.facts.cliRun)?.bin]?.effort ? { effort: 'high' } : {}) }) + '}. Left empty, a lane inherits its own config file, which cli-run cannot see and does not guess. `--model` and `--effort` override this per call, and `--doctor` prints what each lane is pinned to. Every enabled lane takes a model; the runner reports which lanes support an effort flag. Hermes also takes "provider" beside "model", so the model goes to a provider that serves it.'
827
830
  },
828
831
  null,
829
832
  2
@@ -1117,6 +1120,8 @@ export function writeFiles(files, opts) {
1117
1120
  const upgraded = []; // runtime files replaced because the installed copy was an untouched generated one
1118
1121
  const conflicts = []; // runtime files kept because the installed copy differs from what we generated
1119
1122
  const unverifiable = []; // runtime files kept because there is no manifest to compare against
1123
+ const lanesWithheld = []; // selected lanes the kept runner does not support
1124
+ let lanesHash;
1120
1125
  const docsUpdated = []; // --update-docs: documents regenerated because the installed copy was an untouched generated one
1121
1126
  const docsConflict = []; // --update-docs: documents kept because you edited them
1122
1127
  const docsUnverifiable = []; // --update-docs: documents kept because there is no manifest to compare against
@@ -1206,6 +1211,24 @@ export function writeFiles(files, opts) {
1206
1211
  }
1207
1212
  }
1208
1213
  let content = f.content;
1214
+ if (k === 'dir' && key === 'bin/lanes.json' && keptKeys.has('bin/cli-run.mjs')) {
1215
+ // planFiles puts the runner first, so its keep decision is known here.
1216
+ let supported;
1217
+ try {
1218
+ const match = readFileSync(join(root, 'bin', 'cli-run.mjs'), 'utf8').match(/^\s*export const LANES = \[([\s\S]*?)\];/m);
1219
+ if (match && /^\s*(?:(?:'[^'\\]*'|"[^"\\]*")\s*(?:,\s*(?:'[^'\\]*'|"[^"\\]*")\s*)*,?\s*)?$/.test(match[1])) {
1220
+ supported = [...match[1].matchAll(/'([^'\\]*)'|"([^"\\]*)"/g)].map(m => m[1] ?? m[2]);
1221
+ }
1222
+ } catch { /* unknown runner contents are not evidence of unsupported lanes */ }
1223
+ if (supported) {
1224
+ const lanes = JSON.parse(content);
1225
+ lanesWithheld.push(...new Set([...lanes.enabled, ...Object.keys(lanes.defaults || {})].filter(lane => !supported.includes(lane))));
1226
+ lanes.enabled = lanes.enabled.filter(lane => supported.includes(lane));
1227
+ if (lanes.defaults) lanes.defaults = Object.fromEntries(Object.entries(lanes.defaults).filter(([lane]) => supported.includes(lane)));
1228
+ content = JSON.stringify(lanes, null, 2) + '\n';
1229
+ lanesHash = sha256(content);
1230
+ }
1231
+ }
1209
1232
  if (f.rel === 'MANIFEST.json') {
1210
1233
  if (hasLegacy) {
1211
1234
  const previousHash = prevHashes?.[LEGACY_BRIEF];
@@ -1239,6 +1262,7 @@ export function writeFiles(files, opts) {
1239
1262
  // formerly selected files, so uninstall still checks their original
1240
1263
  // installed hashes. New defaults do not erase a previous selection.
1241
1264
  m.files = { ...prevHashes, ...m.files };
1265
+ if (lanesHash) m.files['bin/lanes.json'] = lanesHash;
1242
1266
  if (Object.keys(activation).length) m.activation = activation;
1243
1267
  for (const removed of removedKeys) delete m.files[removed];
1244
1268
  for (const kk of Object.keys(m.files || {})) {
@@ -1304,7 +1328,7 @@ export function writeFiles(files, opts) {
1304
1328
  }
1305
1329
  throw e;
1306
1330
  }
1307
- return { written, skipped, upgraded, conflicts, unverifiable, docsUpdated, docsConflict, docsUnverifiable, docsRenamed, backups };
1331
+ return { written, skipped, upgraded, conflicts, unverifiable, lanesWithheld, docsUpdated, docsConflict, docsUnverifiable, docsRenamed, backups };
1308
1332
  }
1309
1333
 
1310
1334
  export function resolveSelection(ids) {
@@ -50,7 +50,7 @@ export async function installHealthCheck({ level, selected, primary, dir }) {
50
50
  // runner in the target directory. Only its data configuration is read.
51
51
  const rc = await doctor(false, {
52
52
  here: join(dir, 'bin'), primary: primary?.id, compact: true,
53
- ...(level < 2 ? { config: { enabled: selected.filter(ai => ai.facts.cliRun).map(ai => ai.id), defaults: {} } } : {})
53
+ ...(level < 2 ? { config: { enabled: selected.filter(ai => ai.facts.cliRun).map(ai => ai.bin), defaults: {} } } : {})
54
54
  });
55
55
  if (rc && rc !== 10 && rc !== 13) console.log(` doctor could not read the installed lane configuration (exit ${rc}).`);
56
56
  return rc;
package/src/roles.js CHANGED
@@ -136,7 +136,10 @@ export function roleRoute(roleId, assignment, { selected = [], primary = null, a
136
136
  if (role.ai === null) {
137
137
  out.reason = role.why;
138
138
  } else if (role.via === 'cli-run' && ai?.facts.cliRun === true) {
139
- out.command = `cli-run ${ai.id}${['review', 'verify'].includes(roleId) && ai.facts.readOnlyMode === true ? ' --audit' : ''}`;
139
+ out.command = `cli-run ${ai.bin}${['review', 'verify'].includes(roleId) && ai.facts.readOnlyMode === true ? ' --audit' : ''}`;
140
+ // The host knows its connected tools. The standalone CLI command remains
141
+ // usable without an MCP installation; it cannot call host-owned tools.
142
+ if (ai.id === 'claude-code') out.preferredTransport = 'mcp';
140
143
  } else if (role.ai === primary?.id && primary.facts.agentDefinitions && agents[roleId]) {
141
144
  out.agent = agents[roleId];
142
145
  }
@@ -151,6 +154,7 @@ export function roleHow(role, { selected = [], primary = null } = {}) {
151
154
  if (role.ai === null) return role.reason?.startsWith('Nothing in your stack') ? 'keep it off every lane here' : 'fresh-context self-check on your main agent';
152
155
  const ai = selected.find((item) => item.id === role.ai);
153
156
  const tier = `${role.tier} tier`;
157
+ if (role.preferredTransport === 'mcp') return `connected Claude worker MCP when available; fallback \`aunx ${role.command}\`, ${tier}`;
154
158
  if (role.command) return `\`aunx ${role.command}\`, ${tier}`;
155
159
  if (role.via === 'local') return `${ai?.bin ? `\`${ai.bin}\`` : 'local runtime'} on your machine, ${tier}`;
156
160
  if (ai?.facts.kind === 'chat') return `paste the work into your ${role.ai === primary?.id ? 'main agent' : 'chat app'}, ${tier}`;
@@ -18,7 +18,7 @@ The audit script builds an explicit allowed environment with shell builtins befo
18
18
 
19
19
  The allowed runtime names are `HOME`, `PATH`, `USER`, `LOGNAME`, `SHELL`; `LANG`, `LANGUAGE`, `TZ`; `LC_ALL`, `LC_CTYPE`, `LC_COLLATE`, `LC_MESSAGES`, `LC_MONETARY`, `LC_NUMERIC`, `LC_TIME`, `LC_ADDRESS`, `LC_IDENTIFICATION`, `LC_MEASUREMENT`, `LC_NAME`, `LC_PAPER`, `LC_TELEPHONE`; `TMPDIR`, `TMP`, `TEMP`; `XDG_CONFIG_HOME`, `XDG_DATA_HOME`, `XDG_STATE_HOME`, `XDG_CACHE_HOME`, `XDG_RUNTIME_DIR`, `DBUS_SESSION_BUS_ADDRESS`; and the Windows shell runtime names `SystemRoot`, `SYSTEMROOT`, `WINDIR`. User-service access and stored sign-ins use those same locations.
20
20
 
21
- Only the selected report worker additionally receives its location override: `CODEX_HOME` for `codex`, `HERMES_HOME` for `hermes`. These names are removed from collection/version probes. `agy`, `grok` and `qwen` use stored sign-ins in the common user/configuration locations, with no additional exported authentication name. Configure the selected vendor's sign-in in the systemd user's account, such as `hermes auth add <provider>` for Hermes. API-only setups that depend on exported provider keys need vendor-supported stored authentication before this job can run. Claude Code is not a cli-run worker in this catalog, so its session-token export is removed as well.
21
+ Only the selected report worker additionally receives its location override: `CODEX_HOME` for `codex`, `HERMES_HOME` for `hermes`. These names are removed from collection/version probes. `claude`, `agy`, `grok` and `qwen` must use stored sign-ins in the common user/configuration locations, with no additional exported authentication name. Configure the selected vendor's sign-in in the systemd user's account, such as `hermes auth add <provider>` for Hermes. Setups that depend on exported provider keys or session tokens need vendor-supported stored authentication before this job can run. `CLAUDE_CODE_OAUTH_TOKEN` remains excluded by the same boundary as other token variables.
22
22
 
23
23
  `PROBE_SECS` and `RUNNER_SECS` remain shell-only deadline overrides. Set documented runtime and location overrides in the service's `Environment=` settings or audit-only environment file. The key and any selected worker settings stay out of command arguments and logs. There is no arbitrary variable pass-through option.
24
24
 
@@ -34,7 +34,7 @@ Sign the selected vendor CLI in under the same user before enabling the timer. B
34
34
  | User service and keyring session | `XDG_RUNTIME_DIR`, `DBUS_SESSION_BUS_ADDRESS` |
35
35
  | Windows shell runtime compatibility | `SystemRoot`, `SYSTEMROOT`, `WINDIR` |
36
36
 
37
- The report worker additionally receives `CODEX_HOME` when the selected worker is `codex`, or `HERMES_HOME` when it is `hermes`, if that name was set. These location overrides are absent from collection and version probes. The `agy`, `grok` and `qwen` workers use their stored sign-ins under the common user/configuration directories. No token or provider-key variable is allowed for any current worker; `CLAUDE_CODE_OAUTH_TOKEN` is also removed because Claude Code is not a supported cli-run worker in this catalog. Configure authentication with the selected vendor's sign-in flow, such as `codex login`, `grok login`, `hermes auth add <provider>`, `agy`, or Qwen's `/auth`.
37
+ The report worker additionally receives `CODEX_HOME` when the selected worker is `codex`, or `HERMES_HOME` when it is `hermes`, if that name was set. These location overrides are absent from collection and version probes. The `claude`, `agy`, `grok` and `qwen` workers must use stored sign-ins under the common user/configuration directories. No token or provider-key variable is allowed for any current worker, including `CLAUDE_CODE_OAUTH_TOKEN`. Configure authentication with the selected vendor's sign-in flow, such as `claude`, `codex login`, `grok login`, `hermes auth add <provider>`, `agy`, or Qwen's `/auth`.
38
38
 
39
39
  `GATEWAY_MASTER_KEY` stays in a private shell variable until the stdin probe finishes and is then removed. `PROBE_SECS` and `RUNNER_SECS` override the shell's deadlines without becoming child exports. Configure the allowed names and worker location overrides through the service's `Environment=` settings or its audit-only environment file. Keep that file limited to the gateway key and the documented runtime/location settings; keep provider credentials in the gateway launch environment. The script provides no arbitrary extra-variable override.
40
40
 
@@ -6,6 +6,14 @@ Enabled lanes (edit `bin/lanes.json`): {{CLI_RUN_LANES}}.
6
6
 
7
7
  ## Call a lane with a task brief
8
8
 
9
+ - When Claude Code is assigned and a Claude worker MCP service is connected, prefer it. Inspect its schema and instructions, then call its tools from the host. Runner subprocesses cannot access the host's connected tools. The manifest's `preferredTransport: "mcp"` records the preference and `command` gives the CLI fallback; `aunx route` reports it without discovering or invoking servers.
10
+ - When starting an MCP session, send the scoped brief and project directory, keep the session ID and cursor, and poll as the service directs until the final result. Verify acceptance checks, handle errors and cancel on the task deadline or user cancellation. A start response or idle session alone is not proof of success.
11
+ - For either transport, preserve the caller's permissions and project scope. Do not auto-approve tools, broaden paths, add credentials or change authentication. Answer permission requests within existing authorization and deny requests outside it; keep review requests review-only. Check the MCP server's environment against the task's access limits, even when it differs from the command sandbox.
12
+ - When selecting MCP options, leave the model unset unless the task or lane defaults select it; translate effort and other defaults only when the service supports them.
13
+ - When a compatible service is absent, use the native CLI in an authorized environment with its existing sign-in. After either transport fails, report the cause before switching; do not retry automatically through the other transport. Use CLI doctor for CLI presence and optional canaries only; check MCP availability and authentication through the host.
14
+
15
+ Anthropic's built-in [`claude mcp serve`](https://code.claude.com/docs/en/mcp#use-claude-code-as-an-mcp-server) exposes tools; it does not by itself establish a Claude agent-review service.
16
+
9
17
  ```bash
10
18
  aunx cli-run {{EXAMPLE_LANE}} --brief TASK_BRIEF.md --timeout 900
11
19
  node bin/cli-run.mjs {{EXAMPLE_LANE}} --brief TASK_BRIEF.md --timeout 900
@@ -56,6 +64,7 @@ The runner supports these lanes whether or not you selected them.
56
64
 
57
65
  | Lane | Invocation built | Accepted response |
58
66
  |---|---|---|
67
+ | claude | `-p --output-format json --permission-mode dontAsk` | result object, `subtype == "success"`, `is_error == false`, non-empty `result`, no execution errors or truncated stop reason |
59
68
  | grok | `--output-format json -p` | `stopReason == "end_turn"` and non-empty `text` |
60
69
  | codex | `exec --json --color never --skip-git-repo-check -o FILE` | terminal `turn.completed` event and non-empty output file |
61
70
  | agy | `--print-timeout Nm --output-format stream-json -p` | terminal `result` event, `status == "SUCCESS"`, non-empty `response` |
@@ -64,6 +73,10 @@ The runner supports these lanes whether or not you selected them.
64
73
 
65
74
  When Qwen's error telemetry is absent, the runner refuses the response. These checks distinguish response structure from successful task execution.
66
75
 
76
+ Select Claude Code with installer ID `claude-code`; use `aunx cli-run claude` for its executable fallback. The native adapter needs the Claude CLI and its existing authentication, independently of optional MCP companions. It uses [Anthropic's print-mode JSON result](https://code.claude.com/docs/en/headless) and [result completion states](https://code.claude.com/docs/en/agent-sdk/agent-loop).
77
+
78
+ Claude runs with [the `dontAsk` permission mode](https://code.claude.com/docs/en/permissions): calls needing approval are denied, while existing allow rules and actions needing no approval still apply. The runner grants no additional tools or permissions. This is not a read-only filesystem sandbox; `--audit` remains Codex-only. Claude's `permission_denials` counts blocked tools; a non-empty completed answer with denials is accepted with a warning, while no answer with denials exits 17. Authentication or account-access errors exit 14. Error subtypes, malformed JSON and truncated responses exit 18 unless a more specific native error identifies the cause.
79
+
67
80
  ## Respond to the exit class
68
81
 
69
82
  | Code | Class | Action |
@@ -97,19 +110,22 @@ Replace the placeholder with a current vendor model ID before using this example
97
110
 
98
111
  The runner supports these lanes whether or not you selected them.
99
112
 
100
- | Lane | Model flag | Effort flag |
101
- |---|---|---|
102
- | grok | `-m` | `--reasoning-effort` |
103
- | codex | `-m` | `-c model_reasoning_effort="LEVEL"` |
104
- | agy | `--model` | `--effort` |
105
- | hermes | `-m` | `--reasoning` |
106
- | qwen | `-m` | Unsupported; an effort request is a usage error |
113
+ | Lane | Model flag | Effort flag | Provider flag |
114
+ |---|---|---|---|
115
+ | claude | `--model` | `--effort` | Unsupported |
116
+ | grok | `-m` | `--reasoning-effort` | Unsupported |
117
+ | codex | `-m` | `-c model_reasoning_effort="LEVEL"` | Unsupported |
118
+ | agy | `--model` | `--effort` | Unsupported |
119
+ | hermes | `-m` | `--reasoning` | `--provider` |
120
+ | qwen | `-m` | Unsupported; an effort request is a usage error | Unsupported |
121
+
122
+ Hermes signs in to many providers, and a model sent to a provider that does not serve it fails with HTTP 400. Pin both with `--provider` and `--model`, or with `"provider"` beside `"model"` under `"hermes"` in `bin/lanes.json` defaults. A provider with no model is a usage error. That failure is reported as `rejected` (exit 16) with a model/provider mismatch line; `hermes model` repairs the pairing in Hermes itself.
107
123
 
108
124
  When using `--effort auto`, treat its medium/high selection as a bounded heuristic; an audit has a high floor. Use explicit high for builds and xhigh where supported for security-critical or irreversible work. The vendor validates its own effort names and reports unsupported values through the failure class.
109
125
 
110
126
  ## Preserve permissions and secrets
111
127
 
112
- The runner never adds permission flags. Keep each vendor's permissions in its own configuration and route work within the task's granted scope. A denied write becomes a handoff to an authorized writer.
128
+ The runner grants no additional permissions. Claude uses `dontAsk` to deny calls needing approval; Codex `--audit` requests a read-only filesystem sandbox. Keep each vendor's permissions in its own configuration and route work within the task's granted scope. A denied write becomes a handoff to an authorized writer.
113
129
 
114
130
  Prompts travel in argv, which other processes may inspect. Never put secrets in a prompt. For large task inputs, give the worker a brief with authorized source paths instead of exceeding the operating system's argument-size limit.
115
131
 
@@ -58,6 +58,7 @@ When work runs in the background, check liveness and output growth every five mi
58
58
 
59
59
  ## Tools and fallbacks
60
60
 
61
+ {{CLAUDE_WORKER_TRANSPORT}}
61
62
  - When reporting consequential arithmetic or code equivalence, use a computing tool (`protocols/numbers-and-logic.md`). codecalc: {{CODECALC_STATUS}}. When absent, use the local runtime, test suite or spreadsheet.{{METERED_CITATION_NOTE}}
62
63
  - When writing durable records, search first, update the index and keep one writer (`protocols/memory-and-record.md`). obsidian-tc: {{OBSIDIAN_TC_STATUS}}. When absent, use project files, search and version control.
63
64
  - When using a changing library or API, read current documentation and verify behavior (`protocols/docs-then-prove.md`). Context7: {{CONTEXT7_STATUS}}. When absent, read official docs or installed source; use the local runtime when codecalc is absent.