model-orchestrator 1.0.6 → 1.0.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md CHANGED
@@ -25,14 +25,19 @@ When using the Claude Code plugin, follow [plugin/README.md](plugin/README.md).
25
25
 
26
26
  ## Change this repository
27
27
 
28
- - **Read first:** `CONTRIBUTING.md`, `src/README.md` and `docs/security-review-history.md`.
28
+ - **Read first:** `CONTRIBUTING.md`, `src/README.md`, `SECURITY.md` (security scope) and `docs/security-review-history.md`.
29
29
  - **Catalog:** when adding an AI or companion, edit `src/catalog.js`; keep templates free of logic.
30
30
  - **Templates:** when changing installed instructions, edit `templates/`. Use public terms: task brief, context file, acceptance checks, definition of done and result.
31
- - **Generated files:** run `npm run gen:catalog` after catalog changes and `npm run gen:plugin` after plugin-template changes. `test/plugin.test.js` checks the committed bundle.
31
+ - **Generated files:** run `npm run gen:catalog` after catalog changes and `npm run gen:plugin` after plugin-template changes. `plugin/` is generated from `templates/`; only `plugin/README.md` and `plugin/hooks/hooks.json` are hand-owned. `test/plugin.test.js` fails on drift. Before a release, `claude plugin validate --strict plugin` must pass.
32
32
  - **Plugin safety:** hooks may only read and emit context. No network, file writes, subprocesses or credential access.
33
- - **Installer safety:** writes remain inside `--dir` and `--project`; preserve user edits according to manifest hashes and explicit flags. Run no third-party installer.
34
- - **Runner safety:** preserve exit codes and the log schema. Before changing an output judge, add a failing case in `test/judges.test.js`.
33
+ - **Installer safety:** `bin/cli.js` writes only inside `--dir` and `--project`, refuses home-level agent configuration and runs no third-party installer or vendor script.
34
+ - **Installer ownership:** existing documents stay unless `--force` replaces them or `--update-docs` verifies their recorded hash. Machine-owned configuration refreshes, unchanged managed runtime files upgrade, and `--upgrade-runtime` explicitly replaces runtime files. Project activation merges supported entries with backups. Preserve these rules; `test/install.test.js` and `test/cli.test.js` check them.
35
+ - **Runner safety:** preserve exit codes and the log schema. `bin/cli-run.mjs` exits non-zero when a lane produced nothing. Every judge has a red case in `test/judges.test.js`; add one before changing a judge.
36
+ - **Systemd:** when touching `templates/advanced/vm/jobs/`, run `test/systemd/run-on-ubuntu.sh` by hand on a Linux host with systemd to check `TimeoutStartSec` and `KillMode`. This check sits outside `npm test`; follow `test/systemd/README.md`.
37
+ - **Fixtures:** `test/fixtures/` is real vendor output. When a vendor upgrade changes a shape, capture again at the new version and update `manifest.json` and the README compatibility table together.
35
38
  - **Secrets:** use environment-variable names only. Never add a credential value to code, examples or tests.
36
39
  - **Proof:** measure through `proof/scripts/`, store results in `proof/results.json` and regenerate the proof page. An expired entry warns and is re-measured; it never blocks tests or releases.
37
- - **Verify:** run `npm test` and report tests, pass, fail, skipped and exit code. The suite prints current counts. Use `npm pack --dry-run` to inspect publication contents.
40
+ - **Verify:** before proposing a change, run `npm test` and report tests, pass, fail, skipped and exit code. The suite prints current counts. Use `npm pack --dry-run` to inspect publication contents.
38
41
  - **Style:** use short condition-to-action instructions and no em dashes. `test/prose.test.js` checks public vocabulary and examples.
42
+
43
+ Each rule is the fix for a failure that reached an audit or CI. `CHANGELOG.md` names the issue behind each.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,28 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [1.0.8] - 2026-09-30
8
+
9
+ ### Changed
10
+
11
+ - AGENTS.md is the single source of agent instructions. CLAUDE.md imports it so Claude Code loads the same rules as every other agent.
12
+ - Merge the contributor rules that had diverged between AGENTS.md and CLAUDE.md.
13
+ - The feature request form asks for an observable success check.
14
+
15
+ ## [1.0.7] - 2026-09-30
16
+
17
+ ### Fixed
18
+
19
+ - Prefer an already-connected Claude worker MCP service in host routing, with the native CLI retained as a standalone fallback. Route suggestions and generated instructions describe the session, result, permission and environment boundaries; no MCP installation or automatic transport retry is added.
20
+ - Enable the native `claude` CLI worker with print-mode JSON completion checks, existing permission rules, model and effort overrides, failure classification and shared timeout/cancellation handling. Codex-primary stacks can assign independent review to Claude Code; Claude-primary stacks retain Codex review. Generated lanes use executable names while manifest AI identities retain installer IDs. Claude fixtures record native success and authentication failures.
21
+ - Keep supported lanes usable when an edited or unverifiable older runner is kept during an upgrade. Withhold unsupported lanes and their defaults from `lanes.json`, and report how to enable them with `--upgrade-runtime`.
22
+ - Classify Claude's expired OAuth session and `Failed to authenticate` errors as auth (exit 14).
23
+ - Match Claude HTTP codes only in API status fields or `API Error:` prefixes, so a 400 error mentioning 429 tokens stays rejected (exit 16).
24
+
25
+ ### Changed
26
+
27
+ - Check and pin the Claude lane at 2.1.285, and replace its hand-written success fixture with a captured restricted live success at that version.
28
+
7
29
  ## [1.0.6] - 2026-09-29
8
30
 
9
31
  ### Added
@@ -16,6 +38,10 @@ All notable changes to this project are documented here. The format follows [Kee
16
38
  - A hermes model/provider mismatch (an upstream "model is not supported" or `model_not_found`, for example `grok-4.6` sent to `openai-codex`) is now class `rejected`, exit 16, with a problem line naming the mismatch and a fix pointing at `hermes model`. It was reported as a vague nonzero exit.
17
39
  - A pinned hermes route (`--model` or `--effort`, by flag or `lanes.json` default) no longer fails every run. cli-run placed those flags between `-z` and the prompt, and `-z` takes the prompt as its value, so Hermes refused the arguments ("argument -z/--oneshot: expected one argument"). Route flags now come before `-z`, and a Hermes argument error is reported as such rather than as a toolsets error.
18
40
 
41
+ ### Changed
42
+
43
+ - The product website generates public Markdown mirrors, llms discovery files and agent navigation guidance from its existing documentation allowlist on each Git build.
44
+
19
45
  ## [1.0.5] - 2026-09-28
20
46
 
21
47
  ### Changed
@@ -569,7 +595,10 @@ First release.
569
595
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
570
596
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
571
597
 
572
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.5...HEAD
598
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.8...HEAD
599
+ [1.0.8]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.7...v1.0.8
600
+ [1.0.7]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.6...v1.0.7
601
+ [1.0.6]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.5...v1.0.6
573
602
  [1.0.5]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.4...v1.0.5
574
603
  [1.0.4]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.3...v1.0.4
575
604
  [1.0.3]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.2...v1.0.3
package/README.md CHANGED
@@ -175,6 +175,8 @@ It helps an agent match a task to a model, effort level and toolset. Run `npx mo
175
175
 
176
176
  Run `npx model-orchestrator --yes --level 2 --ais claude-code,codex --primary claude-code --project . --dir ./ai-orchestrator`, then follow the activation summary. Claude Code can dispatch scoped work through `aunx cli-run codex --brief TASK_BRIEF.md` and use a different model family for review.
177
177
 
178
+ For Codex as the main agent, set `--primary codex`. Independent review goes to Claude Code through a connected Claude worker MCP service when the host has one, or through `aunx cli-run claude --brief TASK_BRIEF.md`. The `claude` lane runs in `dontAsk` mode, so calls needing approval are denied; `--audit` stays Codex-only.
179
+
178
180
  ### How do I reduce Claude Code token usage?
179
181
 
180
182
  Install routing rules with `npx model-orchestrator`, so your agent has guidance for sending routine work to cheaper models and keeping reads scoped. Use `aunx route-metrics --summary` to measure where your work goes; savings depend on your tasks and model choices.
@@ -233,7 +235,7 @@ The lane wiring and the output judges were written against these versions, which
233
235
 
234
236
  | Lane | Vendor | Version this release was built against | Where that number is proved |
235
237
  |---|---|---|---|
236
- | `claude` | Anthropic | 2.1.226 | the npm pin the installer writes, `@anthropic-ai/claude-code@2.1.226` |
238
+ | `claude` | Anthropic | 2.1.285 | `test/fixtures/claude-2.1.285-success.json`, a recorded run |
237
239
  | `codex` | OpenAI | 0.153.4 | `test/fixtures/codex-0.153.4.jsonl`, a recorded run |
238
240
  | `agy` | Google | 1.1.27 | `test/fixtures/agy-1.1.27.jsonl`, a recorded run |
239
241
  | `grok` | xAI | 1.0.5 | `test/fixtures/grok-1.0.5.json`, a recorded run |
@@ -241,12 +243,14 @@ The lane wiring and the output judges were written against these versions, which
241
243
  | `qwen` | Alibaba | 0.22.3 | `test/fixtures/qwen-0.22.3-nokey.json`, a recorded run |
242
244
  | `ollama` | Ollama | 0.34.4 | the pinned image the level 3 box runs, `ollama/ollama:0.34.4` |
243
245
 
244
- Generated from `src/catalog.js` by `npm run gen:catalog`; `npm test` fails if this table and the catalog disagree. Fixtures were captured 2026-09-06.
246
+ Generated from `src/catalog.js` by `npm run gen:catalog`; `npm test` fails if this table and the catalog disagree. Original vendor fixtures were captured 2026-09-06; [fixture provenance](test/fixtures/README.md) records later captures and remaining gaps.
245
247
 
246
248
  <!-- vendor-table:end -->
247
249
 
248
250
  For npm-installed lanes, `builtAgainst` in the catalog supplies both the compatibility table and the install pin. Newer vendor versions may work or may change a flag the generated wiring uses. When a lane starts failing after a vendor upgrade, compare against this table first.
249
251
 
252
+ The `claude` lane is checked against recorded Claude Code 2.1.285 runs: a restricted success, a sign-in failure and an expired-session failure.
253
+
250
254
  **The live canary runs on your machine, with your credentials.** That is what `aunx cli-run --doctor --run` (direct form: `node bin/cli-run.mjs --doctor --run`) is: it sends every enabled lane one tiny prompt through your own sign-ins and reports `canary ok` or `canary FAILED rc=` per lane. Choose this optional live check after setup or a vendor upgrade when you want to verify actual responses.
251
255
 
252
256
  CI runs the full suite against stub lanes on Ubuntu, macOS and Windows, Node 18/20/22, plus a packaged install into a clean consumer. Run the live check locally to verify your own sign-ins, quota and vendor versions.
@@ -267,6 +271,7 @@ Run `npm test` with your change. Keep templates free of logic and credential val
267
271
  ## Credits
268
272
 
269
273
  - [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for separating role, complexity and stakes instead of compressing them into one scale, for recording the model and effort a lane was actually asked for, and for verifying findings before they trigger repairs. All three shipped in 0.1.14.
274
+ - [@shawnbissell](https://github.com/shawnbissell) added the native Claude Code worker lane (print-mode JSON judge, failure classes, dontAsk permissions) and the connected-MCP preference for Claude workers.
270
275
 
271
276
  ## License
272
277
 
package/bin/cli-run.mjs CHANGED
@@ -15,6 +15,7 @@
15
15
  // NATIVE terminal event and refuses to call an empty run a success.
16
16
  //
17
17
  // grok --output-format json -> stopReason == "end_turn" and text non-empty
18
+ // claude -p --output-format json -> result subtype success, is_error false, result non-empty
18
19
  // codex exec --json --color never -o F -> terminal {"type":"turn.completed"} and F non-empty
19
20
  // agy --output-format stream-json -> terminal {"event":"result"} status SUCCESS, response non-empty
20
21
  // hermes -z -> its exit code is already honest (0 ok / 1 none / 2 bad args)
@@ -59,10 +60,8 @@
59
60
  // silently stops being the route that runs. --model / --effort pin it per call,
60
61
  // `defaults` in lanes.json pins it per lane, and every run logs the value that
61
62
  // was REQUESTED plus where the request came from (flag, lanes.json, or nothing
62
- // at all). It does not log an "actual". One lane of five (grok) does report a
63
- // model id in its own output; the other four report none, and a field present
64
- // for one lane, and absent for four, is worse than no field. It would also be a
65
- // provider-supplied string, which this log deliberately never holds.
63
+ // at all). It does not log an "actual": model reporting varies by CLI and is
64
+ // provider-supplied text, which this log deliberately never holds.
66
65
 
67
66
  import { spawn, spawnSync } from 'node:child_process';
68
67
  import { StringDecoder } from 'node:string_decoder';
@@ -72,7 +71,7 @@ import { join, dirname, delimiter, resolve, relative, isAbsolute, sep } from 'no
72
71
  import { tmpdir, homedir } from 'node:os';
73
72
  import { fileURLToPath, pathToFileURL } from 'node:url';
74
73
 
75
- export const LANES = ['grok', 'codex', 'agy', 'hermes', 'qwen'];
74
+ export const LANES = ['grok', 'codex', 'agy', 'hermes', 'qwen', 'claude'];
76
75
  export const OK = 0, NO_DELIVERABLE = 10, NO_OUTPUT = 11, TIMEOUT = 12, UNAVAILABLE = 13, USAGE = 2;
77
76
  export const AUTH = 14, QUOTA = 15, REJECTED = 16, REFUSED = 17, CUT_SHORT = 18;
78
77
 
@@ -161,6 +160,31 @@ function lastJsonLine(out, needle, want) {
161
160
  const fail = (reason, detail) => ({ text: null, reason, detail });
162
161
  const pass = (text, detail) => ({ text, reason: 'ok', detail });
163
162
 
163
+ // Print-mode JSON is one result object, not the stream-json event sequence.
164
+ // Source: https://code.claude.com/docs/en/headless and /en/agent-sdk/agent-loop.
165
+ function claudeResultObject(out) {
166
+ try {
167
+ const o = JSON.parse(out);
168
+ return isObj(o) && o.type === 'result' ? o : null;
169
+ } catch { return null; }
170
+ }
171
+
172
+ export function judgeClaude(rc, out) {
173
+ let o;
174
+ try { o = JSON.parse(out); } catch { return fail('not_json', 'stdout was not JSON'); }
175
+ if (!isObj(o) || o.type !== 'result') return fail('not_result', 'JSON was not a result object');
176
+ if (o.subtype !== 'success') return fail('bad_subtype', 'result did not report successful completion');
177
+ if (o.is_error !== false) return fail('is_error', 'result did not report is_error=false');
178
+ if (o.errors !== undefined) {
179
+ if (!Array.isArray(o.errors) || !o.errors.every(e => typeof e === 'string')) return fail('error_message_not_string', 'errors was not an array of strings');
180
+ if (o.errors.length) return fail('is_error', 'result carried execution errors');
181
+ }
182
+ if (o.stop_reason != null && !['end_turn', 'stop_sequence', 'refusal'].includes(o.stop_reason)) return fail('bad_stop_reason', 'result stopped before completing a response');
183
+ if (typeof o.result !== 'string') return fail('result_not_string', 'result was not a string');
184
+ const text = o.result.trim();
185
+ return text ? pass(text, 'subtype=success, is_error=false') : fail('empty_result', 'success but empty result');
186
+ }
187
+
164
188
  export function judgeGrok(rc, out) {
165
189
  let o;
166
190
  try {
@@ -256,6 +280,7 @@ export function judgeQwen(rc, out) {
256
280
  // CLI's own --help, not remembered. A lane with `effort: null` has no reasoning
257
281
  // flag at all; asking for one there is a usage error, never a silent drop.
258
282
  // grok -m MODEL --reasoning-effort EFFORT
283
+ // claude --model M --effort EFFORT
259
284
  // codex -m MODEL -c model_reasoning_effort="EFFORT" (a TOML override, hence the quotes)
260
285
  // agy --model M --effort EFFORT (low|medium|high)
261
286
  // hermes -m MODEL --reasoning LEVEL --provider ID (none|minimal|...)
@@ -265,6 +290,7 @@ export function judgeQwen(rc, out) {
265
290
  // (grok-4.6 on openai-codex). `hermes -z --provider` without a model exits 2,
266
291
  // so cli-run refuses that pairing before the lane starts.
267
292
  export const LANE_FLAGS = {
293
+ claude: { model: (v) => ['--model', v], effort: (v) => ['--effort', v], provider: null },
268
294
  grok: { model: (v) => ['-m', v], effort: (v) => ['--reasoning-effort', v], provider: null },
269
295
  codex: { model: (v) => ['-m', v], effort: (v) => ['-c', `model_reasoning_effort="${v}"`], provider: null },
270
296
  agy: { model: (v) => ['--model', v], effort: (v) => ['--effort', v], provider: null },
@@ -275,6 +301,7 @@ export const LANE_FLAGS = {
275
301
  // Auto is deliberately a small, static ladder. It is not a vendor capability
276
302
  // probe and it never chooses the top of a vendor's effort range.
277
303
  export const AUTO_EFFORT = {
304
+ claude: { small: 'medium', large: 'high' },
278
305
  codex: { small: 'medium', large: 'high' },
279
306
  grok: { small: 'medium', large: 'high' },
280
307
  agy: { small: 'medium', large: 'high' },
@@ -381,6 +408,11 @@ export function buildArgv(lane, binary, prompt, opts, tmp) {
381
408
  const timeout = opts.timeout;
382
409
  const route = routeFlags(lane, opts);
383
410
  switch (lane) {
411
+ case 'claude':
412
+ // dontAsk denies calls needing approval and retains configured allow/deny
413
+ // rules. It grants no new permissions and is not a filesystem sandbox.
414
+ // -- ends option parsing so even a brief starting with a dash is data.
415
+ return { argv: [binary, '-p', '--output-format', 'json', '--permission-mode', 'dontAsk', ...route, '--', prompt] };
384
416
  case 'grok':
385
417
  return { argv: [binary, '--output-format', 'json', ...route, '-p', prompt] };
386
418
  case 'codex': {
@@ -410,6 +442,8 @@ export function buildArgv(lane, binary, prompt, opts, tmp) {
410
442
 
411
443
  export function judge(lane, rc, out, err, outFile) {
412
444
  switch (lane) {
445
+ case 'claude':
446
+ return judgeClaude(rc, out);
413
447
  case 'grok':
414
448
  return judgeGrok(rc, out);
415
449
  case 'codex':
@@ -581,7 +615,19 @@ export function qwenErrorText(out) {
581
615
  return parts.join('\n');
582
616
  }
583
617
 
618
+ // Only native failure fields are authoritative. A successful model answer
619
+ // mentioning login, quota or permissions must never become an error signal.
620
+ export function claudeErrorText(out) {
621
+ const o = claudeResultObject(out);
622
+ if (!o) return '';
623
+ const parts = Array.isArray(o.errors) ? o.errors.filter(e => typeof e === 'string') : [];
624
+ if (o.is_error === true && typeof o.result === 'string') parts.push(o.result);
625
+ if (Number.isInteger(o.api_error_status)) parts.push(`api_error_status=${o.api_error_status}`);
626
+ return parts.join('\n');
627
+ }
628
+
584
629
  function authoritativeBlob(lane, out, err, detail, rc) {
630
+ if (lane === 'claude') return `${claudeErrorText(out)}\n${err || ''}`;
585
631
  if (lane === 'codex') return `${codexErrorEventsText(out)}\n${err || ''}`;
586
632
  if (lane === 'agy') return `${agyResultFieldsText(out)}\n${err || ''}`;
587
633
  if (lane === 'hermes') return rc !== 0 ? `${err || ''}\n${hermesMismatchLine(out, { vendorShape: true }) || ''}` : '';
@@ -591,6 +637,7 @@ function authoritativeBlob(lane, out, err, detail, rc) {
591
637
 
592
638
  export function sigAuth(lane, blob) {
593
639
  const b = blob.toLowerCase();
640
+ if (lane === 'claude') return /authentication_error|permission_error|invalid api key|not logged in|please run \/login|failed to authenticate|oauth (?:token|session).*expired|api_error_status=(?:401|403)|api error: (?:401|403)\b/.test(b);
594
641
  if (lane === 'qwen') return b.includes('missing api key');
595
642
  if (lane === 'agy') return b.includes('you are not logged into antigravity') || b.includes('not authenticated');
596
643
  return false; // codex, grok, hermes: no documented native auth signal
@@ -598,6 +645,7 @@ export function sigAuth(lane, blob) {
598
645
 
599
646
  export function sigQuota(lane, blob) {
600
647
  const b = blob.toLowerCase();
648
+ if (lane === 'claude') return /rate_limit_error|rate limit|insufficient credits|credit balance|you['’]ve hit your limit|api_error_status=429|api error: 429\b/.test(b);
601
649
  if (lane === 'qwen') return b.includes('[api error: 402') || b.includes('requires more credits') || blob.includes(' 429') || b.includes('rate limit');
602
650
  if (lane === 'codex') return blob.includes('usage_limit_exceeded') || b.includes("you've hit your usage limit");
603
651
  if (lane === 'hermes') return hermesQuotaText(b);
@@ -606,6 +654,7 @@ export function sigQuota(lane, blob) {
606
654
 
607
655
  export function sigRejected(lane, blob) {
608
656
  const b = blob.toLowerCase();
657
+ if (lane === 'claude') return /invalid_request_error|unknown option|invalid (?:model|effort)|api_error_status=400|api error: 400\b/.test(b);
609
658
  if (lane === 'qwen') return b.includes('[api error: 400') || b.includes('no endpoints found') || b.includes('failed to parse grammar');
610
659
  if (lane === 'codex') return blob.includes('invalid_request_error');
611
660
  if (lane === 'hermes') return hermesArgError(b) || b.includes('toolset') || hermesMismatchLine(blob) !== null;
@@ -615,6 +664,7 @@ export function sigRejected(lane, blob) {
615
664
  // Which judge reasons mean the lane never reached a trustworthy finish, as
616
665
  // opposed to finishing cleanly with nothing in it (class empty).
617
666
  const CUT_SHORT_REASONS = {
667
+ claude: new Set(['not_json', 'not_result', 'bad_subtype', 'is_error', 'bad_stop_reason', 'result_not_string', 'error_message_not_string']),
618
668
  grok: new Set(['not_json', 'bad_stop_reason']),
619
669
  codex: new Set(['no_terminal_event']),
620
670
  agy: new Set(['no_terminal_event', 'bad_last_event']),
@@ -740,6 +790,10 @@ export function refusedGrok(out, { root = process.env.CLI_RUN_GROK_SESSIONS_ROOT
740
790
 
741
791
  export function countRefused(lane, out, err, opts) {
742
792
  try {
793
+ if (lane === 'claude') {
794
+ const pd = claudeResultObject(out)?.permission_denials;
795
+ return Array.isArray(pd) ? pd.length : null;
796
+ }
743
797
  if (lane === 'qwen') return refusedQwen(out);
744
798
  if (lane === 'agy') return refusedAgy(out);
745
799
  if (lane === 'codex') return refusedCodex(out, err);
@@ -1211,7 +1265,8 @@ export async function doctor(run, { here = dirname(fileURLToPath(import.meta.url
1211
1265
  const { enabled, defaults } = cfg;
1212
1266
  let bad = 0;
1213
1267
  if (!compact) console.log(`doctor: ${enabled.length} enabled lane(s): ${enabled.join(', ') || 'none'}`);
1214
- if (!compact && primary && !enabled.includes(primary)) console.log(` note: ${primary} is the main agent and is not an executable lane`);
1268
+ const primaryLane = primary === 'claude-code' ? 'claude' : primary;
1269
+ if (!compact && primary && !enabled.includes(primaryLane)) console.log(` note: ${primary} is the main agent and is not an executable lane`);
1215
1270
  if (!enabled.length) {
1216
1271
  if (compact) {
1217
1272
  console.log('doctor: no executable lanes enabled.');
@@ -1445,6 +1500,7 @@ export async function main(argv) {
1445
1500
  if (!opts.quiet) {
1446
1501
  console.error(`cli-run[${lane}] ${verdict} rc=${code} class=${cls} refused=${refused === null ? 'null' : refused} ${r.seconds.toFixed(1)}s raw=${r.outBytes || 0}B route=${routeNote} :: ${redact(detail)}`);
1447
1502
  let authoritative = null;
1503
+ if (lane === 'claude' && ['auth', 'quota', 'rejected'].includes(cls)) authoritative = stderrHead(claudeErrorText(out) || err);
1448
1504
  if (lane === 'codex' && (cls === 'quota' || cls === 'rejected')) authoritative = codexPrimaryError(out);
1449
1505
  const pf = problemAndFix(lane, cls, { out, err, detail, refused, authoritative });
1450
1506
  if (pf.problem) console.error('cli-run problem: ' + redact(pf.problem).slice(0, 600));
package/bin/cli.js CHANGED
@@ -498,6 +498,7 @@ async function main() {
498
498
  for (const path of preview.written) console.log(' would write ' + path);
499
499
  for (const path of [...preview.skipped, ...preview.conflicts, ...preview.unverifiable, ...preview.docsConflict, ...preview.docsUnverifiable]) console.log(' would keep ' + path);
500
500
  for (const path of preview.backups) console.log(' would back up ' + path);
501
+ if (preview.lanesWithheld.length) console.log(` lanes withheld: ${preview.lanesWithheld.join(', ')} (the kept bin/cli-run.mjs predates them; re-run with --upgrade-runtime to enable)`);
501
502
  }
502
503
  console.log('\n--dry: nothing written.');
503
504
  rl && rl.close();
@@ -548,9 +549,9 @@ async function main() {
548
549
  .concat(JSON.stringify(prev.roles || {}) !== JSON.stringify(plannedManifest.roles || {}) ? ['roles'] : [])
549
550
  : [];
550
551
 
551
- let written, skipped, upgraded, conflicts, unverifiable, docsUpdated, docsConflict, docsUnverifiable, docsRenamed;
552
+ let written, skipped, upgraded, conflicts, unverifiable, lanesWithheld, docsUpdated, docsConflict, docsUnverifiable, docsRenamed;
552
553
  try {
553
- ({ written, skipped, upgraded, conflicts, unverifiable, docsUpdated, docsConflict, docsUnverifiable, docsRenamed } = writeFiles(files, { dir, project, force: flag('force'), upgradeRuntime: flag('upgrade-runtime'), updateDocs: flag('update-docs'), prevManifest: prev, backupExisting: applySnippets, onBackup: (path) => console.log(' backup ' + path) }));
554
+ ({ written, skipped, upgraded, conflicts, unverifiable, lanesWithheld, docsUpdated, docsConflict, docsUnverifiable, docsRenamed } = writeFiles(files, { dir, project, force: flag('force'), upgradeRuntime: flag('upgrade-runtime'), updateDocs: flag('update-docs'), prevManifest: prev, backupExisting: applySnippets, onBackup: (path) => console.log(' backup ' + path) }));
554
555
  } catch (e) {
555
556
  if (e && e.code === 'PREFLIGHT') bad(e.message);
556
557
  throw e;
@@ -564,10 +565,11 @@ async function main() {
564
565
  if (prev) {
565
566
  if (changed.length) {
566
567
  console.log(` selection changed: ${changed.join(', ')}`);
567
- console.log(` applied: ${ownedWritten.join(', ') || 'nothing'} (machine-owned files are always rewritten, so the new lanes are live)`);
568
+ console.log(` applied: ${ownedWritten.join(', ') || 'nothing'} (machine-owned files are always rewritten)`);
568
569
  if (changed.includes('roles')) console.log(' the role assignment changed; MANIFEST.json and aunx route are current. Any kept documents may still carry the previous assignment.');
569
570
  } else console.log(' selection identical.');
570
571
  }
572
+ if (lanesWithheld.length) console.log(` lanes withheld: ${lanesWithheld.join(', ')} (the kept bin/cli-run.mjs predates them; re-run with --upgrade-runtime to enable)`);
571
573
  if (upgraded.length) console.log(` runtime upgraded: ${upgraded.join(', ')} ${flag('upgrade-runtime') ? '(--upgrade-runtime: replaced whether or not you had edited them)' : '(each installed copy matched the hash of a previous run, so nobody had edited it)'}`);
572
574
  if (conflicts.length) {
573
575
  console.log(` runtime CONFLICT, kept: ${conflicts.join(', ')}`);
package/docs/catalog.md CHANGED
@@ -16,14 +16,15 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
16
16
 
17
17
  - **Kind:** agent-cli · **Billing:** subscription · **Level:** 1+
18
18
  - **What it is:** Anthropic's terminal coding agent; its subagents load the project rules file
19
- - **Install:** `npm install -g @anthropic-ai/claude-code@2.1.226`
19
+ - **Install:** `npm install -g @anthropic-ai/claude-code@2.1.285`
20
20
  - **Sign in:** run `claude` once and sign in with your Anthropic account
21
21
  - **Reads rules from:** `CLAUDE.md` · subagents in `.claude/agents/`
22
+ - **cli-run lane:** yes
22
23
  - **Plans:**
23
24
  - Claude Pro (base headroom, checked 2026-09-12; tier models: unverified): https://claude.com/pricing
24
25
  - Claude Max 5x (high headroom, checked 2026-09-12; tier models: unverified): https://claude.com/pricing
25
26
  - Claude Max 20x (max headroom, checked 2026-09-12; tier models: unverified): https://claude.com/pricing
26
- - **Built against:** 2.1.226 (the same number the npm pin uses)
27
+ - **Built against:** 2.1.285 (the same number the npm pin uses)
27
28
 
28
29
  **Capability facts.** An unverified value needs a current capability check before use. Role assignments come from the selected stack, using these facts.
29
30
 
@@ -34,7 +35,7 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
34
35
  | `billing` | subscription |
35
36
  | `pricing` | unverified |
36
37
  | `headless` | yes |
37
- | `cliRun` | no |
38
+ | `cliRun` | yes |
38
39
  | `writesFiles` | yes |
39
40
  | `readOnlyMode` | no |
40
41
  | `liveWeb` | yes |
@@ -6,6 +6,8 @@ This page summarizes completed reviews recorded in the [changelog](../CHANGELOG.
6
6
 
7
7
  | Release | Review recorded | What the review found | What changed and where it is checked |
8
8
  |---|---|---|---|
9
+ | 1.0.7 | Review of external PR 48 (native Claude Code lane) by a second model family, then one audit of the fix diff | An upgrade that kept an edited older runner could disable every lane; an expired Claude sign-in read as cut short; a bare HTTP number in an error could pick the wrong class | Kept runners withhold lanes they do not list; expired-session and status-anchored Claude signals; captured 2.1.285 fixtures; `test/install.test.js`, `test/classify.test.js`, `test/cli.test.js` |
10
+ | 1.0.6 | Pre-release review of the hermes provider route by a second model family, one round | Hermes stdout could print raw terminal control characters in a failure line; agent prose on stdout and an unrelated quota word could be misclassified; one provider test could not fail; route flags placed after `-z` broke every pinned hermes run | Control characters stripped, stdout matched only in the vendor's error shape, the mismatch checked before quota, route flags moved before `-z`; `test/classify.test.js`, `test/judges.test.js`, `test/cli.test.js` |
9
11
  | 1.0.2 | Independent review of 1.0.1 by a second model family, then an audit of its fix | A detached check descendant outlived the timeout while the notes said descendants stop; a symlinked-manifest refusal named no file or fix; refusal paths could carry terminal escapes; a directory got symlink advice | The protocol states the process-group limit; refusals name the path, escape control characters and give advice by file type; `test/security-messages.test.js` |
10
12
  | 1.0.1 | Codex Security scan and regression-backed remediation | A metrics summary executed project code; manifest reads and check descendants needed bounds; weekly jobs inherited gateway credentials | Packaged metrics dispatch, bounded file readers, check process-tree cleanup, stdin-only gateway probe and separate audit environment; `test/security-cli.test.js`, `test/security-vm.test.js`, `test/security-pins.test.js` |
11
13
  | 0.1.0 | Initial review and follow-up round, with a second model-family review | Paths could escape the write roots, partial writes could remain, malformed arguments could proceed, and empty results could appear successful | Containment preflight, exclusive writes with rollback, strict flag parsing and vendor-specific result checks; `test/install.test.js`, `test/cli.test.js`, `test/judges.test.js` |
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "1.0.6",
3
+ "version": "1.0.8",
4
4
  "description": "Model router for AI coding agents: installs routing rules, 8 subagents, hooks and a CLI runner so your AI picks model and effort per task and saves tokens",
5
5
  "type": "module",
6
6
  "bin": {
package/proof/README.md CHANGED
@@ -2,14 +2,14 @@
2
2
 
3
3
  Generated from [results.json](results.json). Each figure has a method, sample size, measurement date and expiry. Run the scripts on your own machine to compare.
4
4
 
5
- Environment: Node v22.22.3, darwin arm64. Timing varies with startup caches and other work on the machine. Synthetic cases show what those fixtures exercise.
5
+ Environment: Node v22.23.2, linux x64. Timing varies with startup caches and other work on the machine. Synthetic cases show what those fixtures exercise.
6
6
 
7
7
  | Measurement | Result | Sample size | Measured | Expires | Reproduce |
8
8
  |---|---|---|---|---|---|
9
- | Dry install wall time | 42.22 ms median | 7 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/install-time.js) |
10
- | Lane runner overhead | 63.74 ms median difference | 7 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/runner-overhead.js) |
11
- | Empty results flagged | 10 fixtures rejected | 10 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/missing-results.js) |
12
- | Acceptance failures blocked | 4 fixtures rejected | 4 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/check-gate.js) |
9
+ | Dry install wall time | 61.59 ms median | 7 | 2026-09-28 | 2026-10-12 | [script](../proof/scripts/install-time.js) |
10
+ | Lane runner overhead | 60.27 ms median difference | 7 | 2026-09-28 | 2026-10-12 | [script](../proof/scripts/runner-overhead.js) |
11
+ | Empty results flagged | 10 fixtures rejected | 10 | 2026-09-28 | 2026-10-12 | [script](../proof/scripts/missing-results.js) |
12
+ | Acceptance failures blocked | 4 fixtures rejected | 4 | 2026-09-28 | 2026-10-12 | [script](../proof/scripts/check-gate.js) |
13
13
 
14
14
  ## Run the proof scripts
15
15
 
@@ -40,7 +40,7 @@ The [recording script](scripts/record-gate.js) captures real command output into
40
40
 
41
41
  Spawn a fresh Node installer process per sample; level 2, Claude Code + Codex, no companions, --dry. Includes Node startup and planning; writes no install files. Isolated home and PATH, no real vendors.
42
42
 
43
- Kind: reproducible local measurement. Sample size: 7. Measured: 2026-09-27. Expires: 2026-10-11.
43
+ Kind: reproducible local measurement. Sample size: 7. Measured: 2026-09-28. Expires: 2026-10-12.
44
44
 
45
45
  Source: [proof/scripts/install-time.js](../proof/scripts/install-time.js).
46
46
 
@@ -48,7 +48,7 @@ Source: [proof/scripts/install-time.js](../proof/scripts/install-time.js).
48
48
 
49
49
  Paired fresh processes: direct Node stub versus cli-run hermes with the same stub. Alternates pair order. Includes wrapper startup, validation and local log writes; excludes vendor/network/model time.
50
50
 
51
- Kind: reproducible local measurement. Sample size: 7. Measured: 2026-09-27. Expires: 2026-10-11.
51
+ Kind: reproducible local measurement. Sample size: 7. Measured: 2026-09-28. Expires: 2026-10-12.
52
52
 
53
53
  Source: [proof/scripts/runner-overhead.js](../proof/scripts/runner-overhead.js).
54
54
 
@@ -56,7 +56,7 @@ Source: [proof/scripts/runner-overhead.js](../proof/scripts/runner-overhead.js).
56
56
 
57
57
  Run cli-run against an exit-0 stub for every supported lane, once with empty stdout and once with an empty native final result. Count exit 10/11 only. A successful Hermes response is the positive control. Synthetic fixtures measure these shapes only.
58
58
 
59
- Kind: reproducible local measurement. Sample size: 10. Measured: 2026-09-27. Expires: 2026-10-11.
59
+ Kind: reproducible local measurement. Sample size: 10. Measured: 2026-09-28. Expires: 2026-10-12.
60
60
 
61
61
  Source: [proof/scripts/missing-results.js](../proof/scripts/missing-results.js).
62
62
 
@@ -64,7 +64,7 @@ Source: [proof/scripts/missing-results.js](../proof/scripts/missing-results.js).
64
64
 
65
65
  Run aunx checks run against nonzero, manual, missing-program and timeout fixtures. Each must exit 1; a passing command must exit 0. This is a local command gate, activated by the user in their release sequence.
66
66
 
67
- Kind: reproducible local measurement. Sample size: 4. Measured: 2026-09-27. Expires: 2026-10-11.
67
+ Kind: reproducible local measurement. Sample size: 4. Measured: 2026-09-28. Expires: 2026-10-12.
68
68
 
69
69
  Source: [proof/scripts/check-gate.js](../proof/scripts/check-gate.js).
70
70
 
@@ -1,78 +1,78 @@
1
1
  {
2
2
  "schemaVersion": 1,
3
3
  "environment": {
4
- "node": "v22.22.3",
5
- "platform": "darwin",
6
- "arch": "arm64"
4
+ "node": "v22.23.2",
5
+ "platform": "linux",
6
+ "arch": "x64"
7
7
  },
8
8
  "entries": [
9
9
  {
10
10
  "id": "install-dry-ms",
11
11
  "label": "Dry install wall time",
12
12
  "kind": "reproducible",
13
- "value": 42.220875,
13
+ "value": 61.594306,
14
14
  "unit": "ms median",
15
- "measuredAt": "2026-09-27",
15
+ "measuredAt": "2026-09-28",
16
16
  "method": "Spawn a fresh Node installer process per sample; level 2, Claude Code + Codex, no companions, --dry. Includes Node startup and planning; writes no install files. Isolated home and PATH, no real vendors.",
17
17
  "sampleSize": 7,
18
18
  "script": "proof/scripts/install-time.js",
19
- "expiresAt": "2026-10-11",
19
+ "expiresAt": "2026-10-12",
20
20
  "samples": [
21
- 51.577042,
22
- 48.655542,
23
- 46.568834,
24
- 42.220875,
25
- 41.573667,
26
- 39.881333,
27
- 40.983417
21
+ 72.998866,
22
+ 61.594306,
23
+ 63.307248,
24
+ 60.59507,
25
+ 61.218156,
26
+ 66.993599,
27
+ 61.070192
28
28
  ]
29
29
  },
30
30
  {
31
31
  "id": "runner-overhead-ms",
32
32
  "label": "Lane runner overhead",
33
33
  "kind": "reproducible",
34
- "value": 63.742250999999996,
34
+ "value": 60.27329300000001,
35
35
  "unit": "ms median difference",
36
- "measuredAt": "2026-09-27",
36
+ "measuredAt": "2026-09-28",
37
37
  "method": "Paired fresh processes: direct Node stub versus cli-run hermes with the same stub. Alternates pair order. Includes wrapper startup, validation and local log writes; excludes vendor/network/model time.",
38
38
  "sampleSize": 7,
39
39
  "script": "proof/scripts/runner-overhead.js",
40
- "expiresAt": "2026-10-11",
40
+ "expiresAt": "2026-10-12",
41
41
  "pairs": [
42
42
  {
43
- "directMs": 26.925375,
44
- "wrappedMs": 409.562,
45
- "overheadMs": 382.63662500000004
43
+ "directMs": 27.602021,
44
+ "wrappedMs": 89.562129,
45
+ "overheadMs": 61.960108
46
46
  },
47
47
  {
48
- "directMs": 28.554875,
49
- "wrappedMs": 95.06725,
50
- "overheadMs": 66.512375
48
+ "directMs": 26.453384,
49
+ "wrappedMs": 85.145791,
50
+ "overheadMs": 58.692407
51
51
  },
52
52
  {
53
- "directMs": 28.981666,
54
- "wrappedMs": 94.610209,
55
- "overheadMs": 65.628543
53
+ "directMs": 26.104555,
54
+ "wrappedMs": 91.019659,
55
+ "overheadMs": 64.915104
56
56
  },
57
57
  {
58
- "directMs": 26.295583,
59
- "wrappedMs": 86.161958,
60
- "overheadMs": 59.866375
58
+ "directMs": 24.351936,
59
+ "wrappedMs": 84.625229,
60
+ "overheadMs": 60.27329300000001
61
61
  },
62
62
  {
63
- "directMs": 28.71575,
64
- "wrappedMs": 82.585583,
65
- "overheadMs": 53.869833
63
+ "directMs": 27.307321,
64
+ "wrappedMs": 86.025276,
65
+ "overheadMs": 58.717955
66
66
  },
67
67
  {
68
- "directMs": 26.668791,
69
- "wrappedMs": 90.411042,
70
- "overheadMs": 63.742250999999996
68
+ "directMs": 26.203506,
69
+ "wrappedMs": 84.365333,
70
+ "overheadMs": 58.161827
71
71
  },
72
72
  {
73
- "directMs": 34.081333,
74
- "wrappedMs": 82.568041,
75
- "overheadMs": 48.48670799999999
73
+ "directMs": 26.328261,
74
+ "wrappedMs": 86.629325,
75
+ "overheadMs": 60.301064
76
76
  }
77
77
  ]
78
78
  },
@@ -82,11 +82,11 @@
82
82
  "kind": "reproducible",
83
83
  "value": 10,
84
84
  "unit": "fixtures rejected",
85
- "measuredAt": "2026-09-27",
85
+ "measuredAt": "2026-09-28",
86
86
  "method": "Run cli-run against an exit-0 stub for every supported lane, once with empty stdout and once with an empty native final result. Count exit 10/11 only. A successful Hermes response is the positive control. Synthetic fixtures measure these shapes only.",
87
87
  "sampleSize": 10,
88
88
  "script": "proof/scripts/missing-results.js",
89
- "expiresAt": "2026-10-11",
89
+ "expiresAt": "2026-10-12",
90
90
  "outcomes": [
91
91
  {
92
92
  "lane": "grok",
@@ -146,11 +146,11 @@
146
146
  "kind": "reproducible",
147
147
  "value": 4,
148
148
  "unit": "fixtures rejected",
149
- "measuredAt": "2026-09-27",
149
+ "measuredAt": "2026-09-28",
150
150
  "method": "Run aunx checks run against nonzero, manual, missing-program and timeout fixtures. Each must exit 1; a passing command must exit 0. This is a local command gate, activated by the user in their release sequence.",
151
151
  "sampleSize": 4,
152
152
  "script": "proof/scripts/check-gate.js",
153
- "expiresAt": "2026-10-11",
153
+ "expiresAt": "2026-10-12",
154
154
  "outcomes": [
155
155
  {
156
156
  "fixture": "nonzero",
package/src/aunx.js CHANGED
@@ -205,6 +205,9 @@ function stackSuggestion(assignment) {
205
205
  if (assignment.command) via = '`' + clean(assignment.command.startsWith('cli-run ') ? `aunx ${assignment.command}` : assignment.command) + '`';
206
206
  else if (assignment.agent) via = '`' + clean(assignment.agent) + '` on your main agent';
207
207
  else via = { 'main-agent': 'your main agent', local: 'your local runtime', manual: 'manual handoff', subagent: 'a subagent', 'cli-run': 'cli-run' }[assignment.via];
208
+ if (assignment.ai === 'claude-code' && assignment.via === 'cli-run' && assignment.command === 'cli-run claude' && assignment.preferredTransport === 'mcp') {
209
+ via = `connected Claude worker MCP when available; fallback ${via}`;
210
+ }
208
211
  return `Your stack: ${name}, via ${via || 'your selected tool'}${reason ? ` (${clean(reason)})` : ''}.`;
209
212
  }
210
213
 
package/src/catalog.js CHANGED
@@ -71,7 +71,7 @@ export const AIS = [
71
71
  billing: 'subscription',
72
72
  pricing: null, // UNVERIFIED: check your provider's current rate.
73
73
  headless: true,
74
- cliRun: false,
74
+ cliRun: true,
75
75
  writesFiles: true,
76
76
  readOnlyMode: false,
77
77
  liveWeb: true, // Source: templates/agents/claude-code/live-researcher.md grants WebSearch and WebFetch.
@@ -93,8 +93,8 @@ export const AIS = [
93
93
  bin: 'claude',
94
94
  summary: 'Anthropic\'s terminal coding agent; its subagents load the project rules file',
95
95
  minLevel: 1,
96
- install: { npm: '@anthropic-ai/claude-code', url: 'https://code.claude.com/docs/en/setup', pin: '2.1.226' },
97
- builtAgainst: '2.1.226',
96
+ install: { npm: '@anthropic-ai/claude-code', url: 'https://code.claude.com/docs/en/setup', pin: '2.1.285' },
97
+ builtAgainst: '2.1.285',
98
98
  auth: 'run `claude` once and sign in with your Anthropic account',
99
99
  // Positive-only (Q1): an author-machine probe of a working, authenticated
100
100
  // session (2026-09-27) still returned {"loggedIn":false} with exit 1, so a
package/src/install.js CHANGED
@@ -70,7 +70,7 @@ export function lanesTable(selected, plans = {}, primary = inferPrimary(selected
70
70
  a.facts.billing,
71
71
  summaryWithEvidence(a),
72
72
  Object.entries(roles).filter(([, role]) => role.ai === a.id).map(([id]) => id).join(', ') || 'none',
73
- a.facts.cliRun ? '`cli-run ' + a.id + '`' : a.bin ? '`' + a.bin + '`' : 'the app',
73
+ a.facts.cliRun ? '`cli-run ' + a.bin + '`' : a.bin ? '`' + a.bin + '`' : 'the app',
74
74
  plans[a.id] ? `${plans[a.id].name} (${plans[a.id].headroom} headroom)` : 'not stated'
75
75
  ]);
76
76
  return table(rows, ['AI', 'Lane', 'What it is', 'Assigned roles', 'Call it with', 'Plan']);
@@ -180,7 +180,7 @@ export function dirProblems(dir) {
180
180
  export function auditLane(selected, primary = selected[0]) {
181
181
  const { roles } = assignRoles({ selected, primary });
182
182
  const id = roles.review.ai ?? roles.bulk.ai ?? null;
183
- return selected.some(a => a.id === id && a.facts.cliRun) ? id : null;
183
+ return selected.find(a => a.id === id && a.facts.cliRun)?.bin || null;
184
184
  }
185
185
 
186
186
  function stackContext(selected, primary, detected = new Set()) {
@@ -558,7 +558,7 @@ function vars(opts) {
558
558
  const assignment = assignRoles({ selected, primary, detected: opts.detected, plans });
559
559
  const stack = stackContext(selected, primary, opts.detected);
560
560
  const enabled = selected.filter(a => a.facts.cliRun);
561
- const exampleLane = enabled[0]?.id || '<lane>';
561
+ const exampleLane = enabled[0]?.bin || '<lane>';
562
562
  const exampleDefaults = { model: '<model-id>', ...(LANE_FLAGS[exampleLane]?.effort ? { effort: 'high' } : {}) };
563
563
  const auditExample = enabled.find(a => a.facts.readOnlyMode);
564
564
  const fallbackNote = 'When no separate lane qualifies, your main agent carries the job at its stated tier. Independent review and local-only work require an eligible lane.';
@@ -634,8 +634,8 @@ function vars(opts) {
634
634
  EXAMPLE_LANE: exampleLane,
635
635
  EXAMPLE_EFFORT_FLAGS: LANE_FLAGS[exampleLane]?.effort ? ' --effort high' : '',
636
636
  EXAMPLE_AUDIT_LANE: auditExample?.id || '',
637
- EXAMPLE_AUDIT_BLOCK: auditExample ? `When reviewing with an available read-only mode, select its audit shape:\n\n\x60\x60\x60bash\naunx cli-run ${auditExample.id} --audit --brief REVIEW.md\nnode bin/cli-run.mjs ${auditExample.id} --audit --brief REVIEW.md\n\x60\x60\x60` : 'When reviewing, verify the chosen lane permissions and request review-only work.',
638
- EXAMPLE_LANES_JSON: JSON.stringify({ enabled: enabled.map(a => a.id), defaults: { [exampleLane]: exampleDefaults } }, null, 2),
637
+ EXAMPLE_AUDIT_BLOCK: auditExample ? `When reviewing with an available read-only mode, select its audit shape:\n\n\x60\x60\x60bash\naunx cli-run ${auditExample.bin} --audit --brief REVIEW.md\nnode bin/cli-run.mjs ${auditExample.bin} --audit --brief REVIEW.md\n\x60\x60\x60` : 'When reviewing, verify the chosen lane permissions and request review-only work.',
638
+ EXAMPLE_LANES_JSON: JSON.stringify({ enabled: enabled.map(a => a.bin), defaults: { [exampleLane]: exampleDefaults } }, null, 2),
639
639
  // Renders only when Qwen is actually selected: the sentence names a flag
640
640
  // that is a usage error on every other lane (C1).
641
641
  QWEN_SAFE_MODE_NOTE: selected.some(a => a.id === 'qwen') ? "When Qwen's safe mode is required, pass `--safe-mode` to that lane. " : '',
@@ -686,8 +686,8 @@ function vars(opts) {
686
686
  AUDIT_LANE: lane || 'none',
687
687
  // Enforced boundary per lane: codex has a read-only sandbox flag; the others
688
688
  // run with whatever their own config allows, and the script says so.
689
- AUDIT_LANE_FLAGS: selected.find(a => a.id === lane)?.facts.readOnlyMode ? '--audit' : '',
690
- AUDIT_LANE_BOUNDARY_NOTE: selected.find(a => a.id === lane)?.facts.readOnlyMode
689
+ AUDIT_LANE_FLAGS: selected.find(a => a.bin === lane)?.facts.readOnlyMode ? '--audit' : '',
690
+ AUDIT_LANE_BOUNDARY_NOTE: selected.find(a => a.bin === lane)?.facts.readOnlyMode
691
691
  ? `${lane} --audit, a read-only filesystem sandbox; commands and network follow the ${lane} config`
692
692
  : lane
693
693
  ? `${lane} offers no sandbox flag cli-run can pass, so the denied-actions list is instruction-level only and enforcement is whatever ${lane}'s own permission config allows`
@@ -714,7 +714,10 @@ function vars(opts) {
714
714
  LANES_TABLE: lanesTable(selected, plans, primary),
715
715
  PLAN_GUIDANCE: planGuidance(selected, plans),
716
716
  INSTALL_TABLE: installTable(selected),
717
- CLI_RUN_LANES: selected.filter((a) => a.facts.cliRun).map((a) => a.id).join(', ') || 'none selected',
717
+ CLI_RUN_LANES: selected.filter((a) => a.facts.cliRun).map((a) => a.bin).join(', ') || 'none selected',
718
+ CLAUDE_WORKER_TRANSPORT: Object.values(assignment.roles).some(role => role.ai === 'claude-code' && role.via === 'cli-run')
719
+ ? '- When assigning a Claude Code worker, prefer a connected Claude worker MCP service exposed in the host tool catalog. Call its tools directly from the host, follow its session and permission workflow, and preserve the task scope. See `CLI-RUN.md` for dispatch and fallback boundaries.\n'
720
+ : '',
718
721
  GATEWAY_MODELS: gatewayModels(selected, apis),
719
722
  ENV_NAMES: envNames(selected, apis).map((n) => '- `' + n + '`').join('\n'),
720
723
  ENV_EXPORTS: envNames(selected, apis).map((n) => n + '=').join('\n'),
@@ -820,10 +823,10 @@ export function planFiles(opts) {
820
823
  join('bin', 'lanes.json'),
821
824
  JSON.stringify(
822
825
  {
823
- enabled: selected.filter((a) => a.facts.cliRun).map((a) => a.id),
824
- defaults: Object.fromEntries((opts.effortAuto || []).map((lane) => [lane, { effort: 'auto' }])),
826
+ enabled: selected.filter((a) => a.facts.cliRun).map((a) => a.bin),
827
+ defaults: Object.fromEntries((opts.effortAuto || []).map((id) => [selected.find(a => a.id === id)?.bin || id, { effort: 'auto' }])),
825
828
  note: 'Lanes cli-run may call. Edit to enable or disable a lane. A lane not listed here exits 13 (unavailable).',
826
- defaultsNote: 'Pin what a lane runs with, so the route in your docs is the route that runs: "defaults": {"' + (selected.find(a => a.facts.cliRun)?.id || '<lane>') + '": ' + JSON.stringify({ model: '<model-id>', ...(LANE_FLAGS[selected.find(a => a.facts.cliRun)?.id]?.effort ? { effort: 'high' } : {}) }) + '}. Left empty, a lane inherits its own config file, which cli-run cannot see and does not guess. `--model` and `--effort` override this per call, and `--doctor` prints what each lane is pinned to. Every enabled lane takes a model; the runner reports which lanes support an effort flag. Hermes also takes "provider" beside "model", so the model goes to a provider that serves it.'
829
+ defaultsNote: 'Pin what a lane runs with, so the route in your docs is the route that runs: "defaults": {"' + (selected.find(a => a.facts.cliRun)?.bin || '<lane>') + '": ' + JSON.stringify({ model: '<model-id>', ...(LANE_FLAGS[selected.find(a => a.facts.cliRun)?.bin]?.effort ? { effort: 'high' } : {}) }) + '}. Left empty, a lane inherits its own config file, which cli-run cannot see and does not guess. `--model` and `--effort` override this per call, and `--doctor` prints what each lane is pinned to. Every enabled lane takes a model; the runner reports which lanes support an effort flag. Hermes also takes "provider" beside "model", so the model goes to a provider that serves it.'
827
830
  },
828
831
  null,
829
832
  2
@@ -1117,6 +1120,8 @@ export function writeFiles(files, opts) {
1117
1120
  const upgraded = []; // runtime files replaced because the installed copy was an untouched generated one
1118
1121
  const conflicts = []; // runtime files kept because the installed copy differs from what we generated
1119
1122
  const unverifiable = []; // runtime files kept because there is no manifest to compare against
1123
+ const lanesWithheld = []; // selected lanes the kept runner does not support
1124
+ let lanesHash;
1120
1125
  const docsUpdated = []; // --update-docs: documents regenerated because the installed copy was an untouched generated one
1121
1126
  const docsConflict = []; // --update-docs: documents kept because you edited them
1122
1127
  const docsUnverifiable = []; // --update-docs: documents kept because there is no manifest to compare against
@@ -1206,6 +1211,24 @@ export function writeFiles(files, opts) {
1206
1211
  }
1207
1212
  }
1208
1213
  let content = f.content;
1214
+ if (k === 'dir' && key === 'bin/lanes.json' && keptKeys.has('bin/cli-run.mjs')) {
1215
+ // planFiles puts the runner first, so its keep decision is known here.
1216
+ let supported;
1217
+ try {
1218
+ const match = readFileSync(join(root, 'bin', 'cli-run.mjs'), 'utf8').match(/^\s*export const LANES = \[([\s\S]*?)\];/m);
1219
+ if (match && /^\s*(?:(?:'[^'\\]*'|"[^"\\]*")\s*(?:,\s*(?:'[^'\\]*'|"[^"\\]*")\s*)*,?\s*)?$/.test(match[1])) {
1220
+ supported = [...match[1].matchAll(/'([^'\\]*)'|"([^"\\]*)"/g)].map(m => m[1] ?? m[2]);
1221
+ }
1222
+ } catch { /* unknown runner contents are not evidence of unsupported lanes */ }
1223
+ if (supported) {
1224
+ const lanes = JSON.parse(content);
1225
+ lanesWithheld.push(...new Set([...lanes.enabled, ...Object.keys(lanes.defaults || {})].filter(lane => !supported.includes(lane))));
1226
+ lanes.enabled = lanes.enabled.filter(lane => supported.includes(lane));
1227
+ if (lanes.defaults) lanes.defaults = Object.fromEntries(Object.entries(lanes.defaults).filter(([lane]) => supported.includes(lane)));
1228
+ content = JSON.stringify(lanes, null, 2) + '\n';
1229
+ lanesHash = sha256(content);
1230
+ }
1231
+ }
1209
1232
  if (f.rel === 'MANIFEST.json') {
1210
1233
  if (hasLegacy) {
1211
1234
  const previousHash = prevHashes?.[LEGACY_BRIEF];
@@ -1239,6 +1262,7 @@ export function writeFiles(files, opts) {
1239
1262
  // formerly selected files, so uninstall still checks their original
1240
1263
  // installed hashes. New defaults do not erase a previous selection.
1241
1264
  m.files = { ...prevHashes, ...m.files };
1265
+ if (lanesHash) m.files['bin/lanes.json'] = lanesHash;
1242
1266
  if (Object.keys(activation).length) m.activation = activation;
1243
1267
  for (const removed of removedKeys) delete m.files[removed];
1244
1268
  for (const kk of Object.keys(m.files || {})) {
@@ -1304,7 +1328,7 @@ export function writeFiles(files, opts) {
1304
1328
  }
1305
1329
  throw e;
1306
1330
  }
1307
- return { written, skipped, upgraded, conflicts, unverifiable, docsUpdated, docsConflict, docsUnverifiable, docsRenamed, backups };
1331
+ return { written, skipped, upgraded, conflicts, unverifiable, lanesWithheld, docsUpdated, docsConflict, docsUnverifiable, docsRenamed, backups };
1308
1332
  }
1309
1333
 
1310
1334
  export function resolveSelection(ids) {
@@ -50,7 +50,7 @@ export async function installHealthCheck({ level, selected, primary, dir }) {
50
50
  // runner in the target directory. Only its data configuration is read.
51
51
  const rc = await doctor(false, {
52
52
  here: join(dir, 'bin'), primary: primary?.id, compact: true,
53
- ...(level < 2 ? { config: { enabled: selected.filter(ai => ai.facts.cliRun).map(ai => ai.id), defaults: {} } } : {})
53
+ ...(level < 2 ? { config: { enabled: selected.filter(ai => ai.facts.cliRun).map(ai => ai.bin), defaults: {} } } : {})
54
54
  });
55
55
  if (rc && rc !== 10 && rc !== 13) console.log(` doctor could not read the installed lane configuration (exit ${rc}).`);
56
56
  return rc;
package/src/roles.js CHANGED
@@ -136,7 +136,10 @@ export function roleRoute(roleId, assignment, { selected = [], primary = null, a
136
136
  if (role.ai === null) {
137
137
  out.reason = role.why;
138
138
  } else if (role.via === 'cli-run' && ai?.facts.cliRun === true) {
139
- out.command = `cli-run ${ai.id}${['review', 'verify'].includes(roleId) && ai.facts.readOnlyMode === true ? ' --audit' : ''}`;
139
+ out.command = `cli-run ${ai.bin}${['review', 'verify'].includes(roleId) && ai.facts.readOnlyMode === true ? ' --audit' : ''}`;
140
+ // The host knows its connected tools. The standalone CLI command remains
141
+ // usable without an MCP installation; it cannot call host-owned tools.
142
+ if (ai.id === 'claude-code') out.preferredTransport = 'mcp';
140
143
  } else if (role.ai === primary?.id && primary.facts.agentDefinitions && agents[roleId]) {
141
144
  out.agent = agents[roleId];
142
145
  }
@@ -151,6 +154,7 @@ export function roleHow(role, { selected = [], primary = null } = {}) {
151
154
  if (role.ai === null) return role.reason?.startsWith('Nothing in your stack') ? 'keep it off every lane here' : 'fresh-context self-check on your main agent';
152
155
  const ai = selected.find((item) => item.id === role.ai);
153
156
  const tier = `${role.tier} tier`;
157
+ if (role.preferredTransport === 'mcp') return `connected Claude worker MCP when available; fallback \`aunx ${role.command}\`, ${tier}`;
154
158
  if (role.command) return `\`aunx ${role.command}\`, ${tier}`;
155
159
  if (role.via === 'local') return `${ai?.bin ? `\`${ai.bin}\`` : 'local runtime'} on your machine, ${tier}`;
156
160
  if (ai?.facts.kind === 'chat') return `paste the work into your ${role.ai === primary?.id ? 'main agent' : 'chat app'}, ${tier}`;
@@ -18,7 +18,7 @@ The audit script builds an explicit allowed environment with shell builtins befo
18
18
 
19
19
  The allowed runtime names are `HOME`, `PATH`, `USER`, `LOGNAME`, `SHELL`; `LANG`, `LANGUAGE`, `TZ`; `LC_ALL`, `LC_CTYPE`, `LC_COLLATE`, `LC_MESSAGES`, `LC_MONETARY`, `LC_NUMERIC`, `LC_TIME`, `LC_ADDRESS`, `LC_IDENTIFICATION`, `LC_MEASUREMENT`, `LC_NAME`, `LC_PAPER`, `LC_TELEPHONE`; `TMPDIR`, `TMP`, `TEMP`; `XDG_CONFIG_HOME`, `XDG_DATA_HOME`, `XDG_STATE_HOME`, `XDG_CACHE_HOME`, `XDG_RUNTIME_DIR`, `DBUS_SESSION_BUS_ADDRESS`; and the Windows shell runtime names `SystemRoot`, `SYSTEMROOT`, `WINDIR`. User-service access and stored sign-ins use those same locations.
20
20
 
21
- Only the selected report worker additionally receives its location override: `CODEX_HOME` for `codex`, `HERMES_HOME` for `hermes`. These names are removed from collection/version probes. `agy`, `grok` and `qwen` use stored sign-ins in the common user/configuration locations, with no additional exported authentication name. Configure the selected vendor's sign-in in the systemd user's account, such as `hermes auth add <provider>` for Hermes. API-only setups that depend on exported provider keys need vendor-supported stored authentication before this job can run. Claude Code is not a cli-run worker in this catalog, so its session-token export is removed as well.
21
+ Only the selected report worker additionally receives its location override: `CODEX_HOME` for `codex`, `HERMES_HOME` for `hermes`. These names are removed from collection/version probes. `claude`, `agy`, `grok` and `qwen` must use stored sign-ins in the common user/configuration locations, with no additional exported authentication name. Configure the selected vendor's sign-in in the systemd user's account, such as `hermes auth add <provider>` for Hermes. Setups that depend on exported provider keys or session tokens need vendor-supported stored authentication before this job can run. `CLAUDE_CODE_OAUTH_TOKEN` remains excluded by the same boundary as other token variables.
22
22
 
23
23
  `PROBE_SECS` and `RUNNER_SECS` remain shell-only deadline overrides. Set documented runtime and location overrides in the service's `Environment=` settings or audit-only environment file. The key and any selected worker settings stay out of command arguments and logs. There is no arbitrary variable pass-through option.
24
24
 
@@ -34,7 +34,7 @@ Sign the selected vendor CLI in under the same user before enabling the timer. B
34
34
  | User service and keyring session | `XDG_RUNTIME_DIR`, `DBUS_SESSION_BUS_ADDRESS` |
35
35
  | Windows shell runtime compatibility | `SystemRoot`, `SYSTEMROOT`, `WINDIR` |
36
36
 
37
- The report worker additionally receives `CODEX_HOME` when the selected worker is `codex`, or `HERMES_HOME` when it is `hermes`, if that name was set. These location overrides are absent from collection and version probes. The `agy`, `grok` and `qwen` workers use their stored sign-ins under the common user/configuration directories. No token or provider-key variable is allowed for any current worker; `CLAUDE_CODE_OAUTH_TOKEN` is also removed because Claude Code is not a supported cli-run worker in this catalog. Configure authentication with the selected vendor's sign-in flow, such as `codex login`, `grok login`, `hermes auth add <provider>`, `agy`, or Qwen's `/auth`.
37
+ The report worker additionally receives `CODEX_HOME` when the selected worker is `codex`, or `HERMES_HOME` when it is `hermes`, if that name was set. These location overrides are absent from collection and version probes. The `claude`, `agy`, `grok` and `qwen` workers must use stored sign-ins under the common user/configuration directories. No token or provider-key variable is allowed for any current worker, including `CLAUDE_CODE_OAUTH_TOKEN`. Configure authentication with the selected vendor's sign-in flow, such as `claude`, `codex login`, `grok login`, `hermes auth add <provider>`, `agy`, or Qwen's `/auth`.
38
38
 
39
39
  `GATEWAY_MASTER_KEY` stays in a private shell variable until the stdin probe finishes and is then removed. `PROBE_SECS` and `RUNNER_SECS` override the shell's deadlines without becoming child exports. Configure the allowed names and worker location overrides through the service's `Environment=` settings or its audit-only environment file. Keep that file limited to the gateway key and the documented runtime/location settings; keep provider credentials in the gateway launch environment. The script provides no arbitrary extra-variable override.
40
40
 
@@ -6,6 +6,14 @@ Enabled lanes (edit `bin/lanes.json`): {{CLI_RUN_LANES}}.
6
6
 
7
7
  ## Call a lane with a task brief
8
8
 
9
+ - When Claude Code is assigned and a Claude worker MCP service is connected, prefer it. Inspect its schema and instructions, then call its tools from the host. Runner subprocesses cannot access the host's connected tools. The manifest's `preferredTransport: "mcp"` records the preference and `command` gives the CLI fallback; `aunx route` reports it without discovering or invoking servers.
10
+ - When starting an MCP session, send the scoped brief and project directory, keep the session ID and cursor, and poll as the service directs until the final result. Verify acceptance checks, handle errors and cancel on the task deadline or user cancellation. A start response or idle session alone is not proof of success.
11
+ - For either transport, preserve the caller's permissions and project scope. Do not auto-approve tools, broaden paths, add credentials or change authentication. Answer permission requests within existing authorization and deny requests outside it; keep review requests review-only. Check the MCP server's environment against the task's access limits, even when it differs from the command sandbox.
12
+ - When selecting MCP options, leave the model unset unless the task or lane defaults select it; translate effort and other defaults only when the service supports them.
13
+ - When a compatible service is absent, use the native CLI in an authorized environment with its existing sign-in. After either transport fails, report the cause before switching; do not retry automatically through the other transport. Use CLI doctor for CLI presence and optional canaries only; check MCP availability and authentication through the host.
14
+
15
+ Anthropic's built-in [`claude mcp serve`](https://code.claude.com/docs/en/mcp#use-claude-code-as-an-mcp-server) exposes tools; it does not by itself establish a Claude agent-review service.
16
+
9
17
  ```bash
10
18
  aunx cli-run {{EXAMPLE_LANE}} --brief TASK_BRIEF.md --timeout 900
11
19
  node bin/cli-run.mjs {{EXAMPLE_LANE}} --brief TASK_BRIEF.md --timeout 900
@@ -56,6 +64,7 @@ The runner supports these lanes whether or not you selected them.
56
64
 
57
65
  | Lane | Invocation built | Accepted response |
58
66
  |---|---|---|
67
+ | claude | `-p --output-format json --permission-mode dontAsk` | result object, `subtype == "success"`, `is_error == false`, non-empty `result`, no execution errors or truncated stop reason |
59
68
  | grok | `--output-format json -p` | `stopReason == "end_turn"` and non-empty `text` |
60
69
  | codex | `exec --json --color never --skip-git-repo-check -o FILE` | terminal `turn.completed` event and non-empty output file |
61
70
  | agy | `--print-timeout Nm --output-format stream-json -p` | terminal `result` event, `status == "SUCCESS"`, non-empty `response` |
@@ -64,6 +73,10 @@ The runner supports these lanes whether or not you selected them.
64
73
 
65
74
  When Qwen's error telemetry is absent, the runner refuses the response. These checks distinguish response structure from successful task execution.
66
75
 
76
+ Select Claude Code with installer ID `claude-code`; use `aunx cli-run claude` for its executable fallback. The native adapter needs the Claude CLI and its existing authentication, independently of optional MCP companions. It uses [Anthropic's print-mode JSON result](https://code.claude.com/docs/en/headless) and [result completion states](https://code.claude.com/docs/en/agent-sdk/agent-loop).
77
+
78
+ Claude runs with [the `dontAsk` permission mode](https://code.claude.com/docs/en/permissions): calls needing approval are denied, while existing allow rules and actions needing no approval still apply. The runner grants no additional tools or permissions. This is not a read-only filesystem sandbox; `--audit` remains Codex-only. Claude's `permission_denials` counts blocked tools; a non-empty completed answer with denials is accepted with a warning, while no answer with denials exits 17. Authentication or account-access errors exit 14. Error subtypes, malformed JSON and truncated responses exit 18 unless a more specific native error identifies the cause.
79
+
67
80
  ## Respond to the exit class
68
81
 
69
82
  | Code | Class | Action |
@@ -99,6 +112,7 @@ The runner supports these lanes whether or not you selected them.
99
112
 
100
113
  | Lane | Model flag | Effort flag | Provider flag |
101
114
  |---|---|---|---|
115
+ | claude | `--model` | `--effort` | Unsupported |
102
116
  | grok | `-m` | `--reasoning-effort` | Unsupported |
103
117
  | codex | `-m` | `-c model_reasoning_effort="LEVEL"` | Unsupported |
104
118
  | agy | `--model` | `--effort` | Unsupported |
@@ -111,7 +125,7 @@ When using `--effort auto`, treat its medium/high selection as a bounded heurist
111
125
 
112
126
  ## Preserve permissions and secrets
113
127
 
114
- The runner never adds permission flags. Keep each vendor's permissions in its own configuration and route work within the task's granted scope. A denied write becomes a handoff to an authorized writer.
128
+ The runner grants no additional permissions. Claude uses `dontAsk` to deny calls needing approval; Codex `--audit` requests a read-only filesystem sandbox. Keep each vendor's permissions in its own configuration and route work within the task's granted scope. A denied write becomes a handoff to an authorized writer.
115
129
 
116
130
  Prompts travel in argv, which other processes may inspect. Never put secrets in a prompt. For large task inputs, give the worker a brief with authorized source paths instead of exceeding the operating system's argument-size limit.
117
131
 
@@ -58,6 +58,7 @@ When work runs in the background, check liveness and output growth every five mi
58
58
 
59
59
  ## Tools and fallbacks
60
60
 
61
+ {{CLAUDE_WORKER_TRANSPORT}}
61
62
  - When reporting consequential arithmetic or code equivalence, use a computing tool (`protocols/numbers-and-logic.md`). codecalc: {{CODECALC_STATUS}}. When absent, use the local runtime, test suite or spreadsheet.{{METERED_CITATION_NOTE}}
62
63
  - When writing durable records, search first, update the index and keep one writer (`protocols/memory-and-record.md`). obsidian-tc: {{OBSIDIAN_TC_STATUS}}. When absent, use project files, search and version control.
63
64
  - When using a changing library or API, read current documentation and verify behavior (`protocols/docs-then-prove.md`). Context7: {{CONTEXT7_STATUS}}. When absent, read official docs or installed source; use the local runtime when codecalc is absent.