model-orchestrator 0.1.21 → 0.1.23

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,28 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.23] - 2026-09-15
8
+
9
+ ### Added
10
+
11
+ - **`cli-run` says WHY a lane failed.** Every run lands in one class with its own exit code: `auth` 14, `quota` 15, `rejected` 16, `refused` 17, `cut_short` 18, next to the existing `empty` 10, `no_output` 11, `timeout` 12 and `unavailable` 13. A missing API key, a spent quota and an unknown model id used to share exit 10 or the vendor's own code, and each needs a different response: a missing key is not a model fault, and retrying a spent quota cannot help. Signals are read from each lane's authoritative error fields only, never the model's prose, with precedence auth, quota, rejected, refused, cut_short, empty.
12
+ - **`refused=N` on every run.** Tool calls a hook or deny rule blocked, counted from qwen's `permission_denials`, agy's deny-rule steps, codex's router `Rejected(` lines, and grok's session transcript (its `sessionId` is charset-checked, the path is contained to the sessions root, lines must name the same session, and the read is capped at 5 MiB). `null` when a lane gives no signal. A deliverable with refused calls is still exit 0.
13
+ - **A problem line and a fix line** on the terminal for every failure, and for an `ok` run with refused calls, so a calling agent can relay "this is what went wrong, this is the fix" and ask.
14
+ - **Terminal output is redacted** before it prints: JSON credential keys, `Authorization:` values, bearer values, URL query credentials, and common key prefixes, redacted before any clipping.
15
+ - **The durable log gains `class` and `refused`.** Both are fixed values; the log still never holds provider text, the problem line or stderr.
16
+
17
+ ### Changed
18
+
19
+ - **A nonzero vendor exit is no longer passed through as cli-run's exit code.** The class owns the code, and the vendor's own code stays in the log as `cli_rc`. A nonzero exit nothing else explains is `cut_short` (18), and it is still never `ok`. If a script compared `cli-run`'s exit code with a specific vendor code, compare `cli_rc` in the log instead; `!= 0` checks are unaffected.
20
+ - **A lane killed by a signal, or output past the 16 MiB buffer, is `cut_short` (18)**, not 10. Exit 10 now means only an empty run or an unmet `--expect-*` contract.
21
+ - **agy's judge refuses a non-object terminal `result` as `bad_last_event`** (was `bad_status`), and **qwen's judge refuses a non-string `error.message` as `error_message_not_string`**, so both classify as `cut_short`. hermes' stderr cause (degraded free tier, bad `--toolsets`) is now named in its detail line.
22
+
23
+ ## [0.1.22] - 2026-09-12
24
+
25
+ ### Added
26
+
27
+ - **SuperGrok Plus as a `grok` plan (high headroom).** xAI's pricing page lists it with "Significantly higher usage across Chat, Imagine, Voice & Build", so `--plans grok=supergrok-plus` now works and counts as a high-headroom lane for plan guidance and `--effort-auto`. SuperGrok Heavy stays out: the page states no Build usage for it. A test holds both.
28
+
7
29
  ## [0.1.21] - 2026-09-12
8
30
 
9
31
  ### Added
@@ -329,7 +351,9 @@ First release.
329
351
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
330
352
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
331
353
 
332
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.21...HEAD
354
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.23...HEAD
355
+ [0.1.23]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.22...v0.1.23
356
+ [0.1.22]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.21...v0.1.22
333
357
  [0.1.21]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.20...v0.1.21
334
358
  [0.1.20]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.19...v0.1.20
335
359
  [0.1.19]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.18...v0.1.19
package/bin/cli-run.mjs CHANGED
@@ -27,18 +27,30 @@
27
27
  // bounded by the OS ARG_MAX, so: no secrets in a prompt, and very large briefs
28
28
  // should be referenced by path in the prompt rather than pasted into it.
29
29
  //
30
- // Exit codes
31
- // 0 structurally accepted non-empty response (and every --expect-* contract met)
32
- // 10 ran, produced no deliverable, or a contract was not met, or killed by signal
33
- // 11 produced no output at all
34
- // 12 timed out (the lane AND its descendants are killed as a process group)
35
- // 13 lane unavailable (missing binary, disabled in lanes.json, or lanes.json malformed)
30
+ // Exit codes: one per failure CLASS, so the code says what to do next.
31
+ // 0 ok: structurally accepted non-empty response (and every --expect-* contract met)
32
+ // 10 empty: ran and delivered nothing, or a contract was not met
33
+ // 11 no_output: produced no output at all
34
+ // 12 timeout: the lane AND its descendants are killed as a process group
35
+ // 13 unavailable: missing binary, disabled in lanes.json, or lanes.json malformed
36
+ // 14 auth: the lane's own error says a credential is missing or not logged in
37
+ // 15 quota: the lane's own error says usage limit, credits or rate limit
38
+ // 16 rejected: the upstream rejected the request (bad model id, bad request)
39
+ // 17 refused: no deliverable, and the lane reports tool calls a hook or deny rule blocked
40
+ // 18 cut_short: no trustworthy finish: a missing or non-success terminal event, a lane
41
+ // killed by a signal, output past the 16 MiB buffer, or a nonzero vendor exit
36
42
  // 130 / 143 cli-run itself received SIGINT / SIGTERM: the lane's process group was killed first
37
43
  // 2 usage error in cli-run itself
38
- // N the lane exited N != 0: passed through, verdict exit_nonzero, even if text came back
44
+ // A run that delivered AND had tool calls refused is still 0, with refused=N and a
45
+ // problem/fix pair on the terminal. The vendor's own exit code is logged as cli_rc.
46
+ // Precedence when several signals are present: auth, quota, rejected, refused,
47
+ // cut_short, empty. Only the lane's authoritative error fields are searched,
48
+ // never the model's prose, so an answer that merely mentions "rate limit" is
49
+ // not a quota failure.
39
50
  //
40
- // The durable log stores a FIXED reason code per run (see REASONS), never a
41
- // provider-supplied string. Bounded vendor stderr goes to your terminal only.
51
+ // The durable log stores a FIXED reason code and class per run (see REASONS,
52
+ // CLASS_CODES), never a provider-supplied string. Bounded vendor stderr and the
53
+ // problem/fix lines go to your terminal only, redacted.
42
54
  //
43
55
  // ROUTE: which model and reasoning effort a lane ran with.
44
56
  // A lane with no --model and no lanes.json default inherits whatever its own
@@ -55,12 +67,21 @@ import { spawn, spawnSync } from 'node:child_process';
55
67
  import { StringDecoder } from 'node:string_decoder';
56
68
  import { createHash } from 'node:crypto';
57
69
  import { readFileSync, existsSync, mkdirSync, appendFileSync, mkdtempSync, rmSync, accessSync, constants, realpathSync, statSync, lstatSync, openSync, readSync, closeSync } from 'node:fs';
58
- import { join, dirname, delimiter, resolve } from 'node:path';
70
+ import { join, dirname, delimiter, resolve, relative, isAbsolute, sep } from 'node:path';
59
71
  import { tmpdir, homedir } from 'node:os';
60
72
  import { fileURLToPath, pathToFileURL } from 'node:url';
61
73
 
62
74
  export const LANES = ['grok', 'codex', 'agy', 'hermes', 'qwen'];
63
75
  export const OK = 0, NO_DELIVERABLE = 10, NO_OUTPUT = 11, TIMEOUT = 12, UNAVAILABLE = 13, USAGE = 2;
76
+ export const AUTH = 14, QUOTA = 15, REJECTED = 16, REFUSED = 17, CUT_SHORT = 18;
77
+
78
+ // The closed set of failure classes. Every judged run lands in exactly one, and
79
+ // the class owns the exit code. `interrupted` (cli-run itself was signalled)
80
+ // is logged as a class but exits 130 or 143.
81
+ export const CLASS_CODES = {
82
+ ok: OK, empty: NO_DELIVERABLE, no_output: NO_OUTPUT, timeout: TIMEOUT, unavailable: UNAVAILABLE,
83
+ auth: AUTH, quota: QUOTA, rejected: REJECTED, refused: REFUSED, cut_short: CUT_SHORT
84
+ };
64
85
 
65
86
  // Every reason that may reach the durable log. A judge or the wrapper picks
66
87
  // one of these; anything else is written as 'unknown'. Provider text never
@@ -69,8 +90,8 @@ export const REASONS = new Set([
69
90
  'ok', 'not_json', 'bad_stop_reason', 'empty_text', 'no_terminal_event', 'empty_output_file',
70
91
  'bad_status', 'empty_response', 'exit_nonzero', 'empty_stdout', 'bad_event_array', 'bad_last_event',
71
92
  'not_result', 'bad_subtype', 'is_error', 'result_not_string', 'empty_result', 'api_error_in_result',
72
- 'telemetry_absent', 'total_errors_unreadable', 'total_errors', 'contract_unmet',
73
- 'timeout', 'unavailable', 'killed', 'disabled', 'lanes_json_malformed', 'no_output', 'unknown'
93
+ 'telemetry_absent', 'total_errors_unreadable', 'total_errors', 'contract_unmet', 'error_message_not_string',
94
+ 'judge_raised', 'timeout', 'unavailable', 'killed', 'disabled', 'lanes_json_malformed', 'no_output', 'unknown'
74
95
  ]);
75
96
 
76
97
  const LOG = join(homedir(), '.ai-orchestrator', 'cli-run.log.jsonl');
@@ -163,7 +184,8 @@ export function judgeCodex(rc, out, err, fileText) {
163
184
  export function judgeAgy(rc, out) {
164
185
  const ev = lastJsonLine(out, '"result"', (o) => o && o.event === 'result');
165
186
  if (ev === null) return fail('no_terminal_event', 'no terminal result event');
166
- const term = ev.result && typeof ev.result === 'object' && !Array.isArray(ev.result) ? ev.result : {};
187
+ if (!ev.result || typeof ev.result !== 'object' || Array.isArray(ev.result)) return fail('bad_last_event', 'terminal result was not an object');
188
+ const term = ev.result;
167
189
  const status = term.status;
168
190
  const text = typeof term.response === 'string' ? term.response.trim() : '';
169
191
  if (status !== 'SUCCESS') return fail('bad_status', `status=${JSON.stringify(status)}`);
@@ -173,7 +195,12 @@ export function judgeAgy(rc, out) {
173
195
  export function judgeHermes(rc, out, err) {
174
196
  const text = String(out || '').trim();
175
197
  if (rc !== 0) {
176
- const why = { 1: 'no final response (agent produced nothing)', 2: 'bad args, or completed with an empty response' }[rc] || 'unknown failure';
198
+ let why = { 1: 'no final response (agent produced nothing)', 2: 'bad args, or completed with an empty response' }[rc] || 'unknown failure';
199
+ // hermes collapses every upstream failure into one exit code; its stderr
200
+ // is the only place the cause is named.
201
+ const blob = String(err || '').toLowerCase(); // stderr only: stdout is the agent's own prose
202
+ if (hermesQuotaText(blob)) why += ': upstream free tier degraded or limited, a retry is reasonable';
203
+ else if (blob.includes('toolset')) why += ': invalid --toolsets value, a caller bug and not a lane fault';
177
204
  return fail('exit_nonzero', `hermes exit ${rc}: ${why}`);
178
205
  }
179
206
  return text ? pass(text, 'exit 0') : fail('empty_stdout', 'exit 0 but empty stdout');
@@ -192,8 +219,11 @@ export function judgeQwen(rc, out) {
192
219
  if (term.type !== 'result') return fail('not_result', `last event was ${JSON.stringify(term.type)}, not result`);
193
220
  if (term.subtype !== 'success') {
194
221
  const e = term.error;
222
+ if (e && typeof e === 'object' && e.message !== undefined && typeof e.message !== 'string') {
223
+ return fail('error_message_not_string', `subtype=${JSON.stringify(term.subtype)}, error.message was not a string`);
224
+ }
195
225
  const msg = e && typeof e === 'object' && typeof e.message === 'string' ? e.message : e ? String(e) : '';
196
- return fail('bad_subtype', `subtype=${JSON.stringify(term.subtype)}` + (msg ? `: ${msg.slice(0, 120)}` : ''));
226
+ return fail('bad_subtype', `subtype=${JSON.stringify(term.subtype)}` + (msg ? `: ${redact(msg).slice(0, 120)}` : ''));
197
227
  }
198
228
  if (term.is_error) return fail('is_error', 'is_error true');
199
229
  if (term.result != null && typeof term.result !== 'string') return fail('result_not_string', `result was ${typeof term.result}, not a string`);
@@ -201,7 +231,7 @@ export function judgeQwen(rc, out) {
201
231
  if (!text) return fail('empty_result', 'success but empty result');
202
232
  // qwen reports success even when the upstream API rejected the call; the
203
233
  // error text lands in `result`. These two checks are the honest ones.
204
- if (text.startsWith('[API Error:')) return fail('api_error_in_result', `success flag lied, result is an API error: ${text.slice(0, 140)}`);
234
+ if (text.startsWith('[API Error:')) return fail('api_error_in_result', `success flag lied, result is an API error: ${redact(text).slice(0, 140)}`);
205
235
  const stats = term.stats;
206
236
  const models = stats && typeof stats === 'object' ? stats.models : null;
207
237
  if (!models || typeof models !== 'object' || Array.isArray(models) || Object.keys(models).length === 0) {
@@ -383,6 +413,406 @@ export function judge(lane, rc, out, err, outFile) {
383
413
  }
384
414
  }
385
415
 
416
+ // A judge is type-guarded, but this is the backstop: an exception from any
417
+ // judge would escape main() as cli-run's own crash and be misread as a
418
+ // wrapper bug. It becomes a classifiable verdict instead (class cut_short).
419
+ export function safeJudge(lane, rc, out, err, outFile) {
420
+ try {
421
+ return judge(lane, rc, out, err, outFile);
422
+ } catch (e) {
423
+ return fail('judge_raised', `judge raised ${(e && e.name) || 'Error'}: ${(e && e.message) || e}`);
424
+ }
425
+ }
426
+
427
+ // --- failure classes: WHY a run failed --------------------------------------
428
+ // Every function here returns a safe default instead of throwing. Signals are
429
+ // read from a lane's authoritative error fields only, never from assistant
430
+ // prose, and each one was taken from a captured vendor shape, not guessed.
431
+
432
+ // Terminal output only; nothing redacted here reaches the durable log anyway.
433
+ // Redact BEFORE clipping: a clip first can cut a long token so its tail no
434
+ // longer matches any pattern.
435
+ export const SECRET_PATTERNS = [
436
+ // A JSON string value, escapes included; an unterminated value (a clipped line) runs to the end.
437
+ // The optional backslashes also catch a JSON body escaped inside another string.
438
+ [/(\\?"(?:api[_-]?key|token|secret|password|access[_-]?token)\\?"\s*:\s*)\\?"(?:[^"\\]|\\.)*(?:"|$)/gi, '$1"[REDACTED]"'],
439
+ [/(authorization\s*:\s*)(\S+)\s+\S+/gi, '$1$2 [REDACTED]'],
440
+ [/([?&](?:token|key|api_key|access_token|sig)=)[^&\s"'<>]+/gi, '$1[REDACTED]'],
441
+ [/(bearer\s+)[A-Za-z0-9._+/=-]{8,}/gi, '$1[REDACTED]'],
442
+ [/sk-[A-Za-z0-9_-]{8,}/g, '[REDACTED]'],
443
+ [/xai-[A-Za-z0-9_-]{8,}/g, '[REDACTED]'],
444
+ [/ghp_[A-Za-z0-9_-]{8,}/g, '[REDACTED]'],
445
+ [/AIza[A-Za-z0-9_-]{8,}/g, '[REDACTED]']
446
+ ];
447
+ export function redact(text) {
448
+ if (typeof text !== 'string') return text;
449
+ let s = text;
450
+ for (const [pat, repl] of SECRET_PATTERNS) s = s.replace(pat, repl);
451
+ return s;
452
+ }
453
+
454
+ function* jsonLines(out) {
455
+ for (const raw of String(out || '').split('\n')) {
456
+ const line = raw.trim();
457
+ if (!line.startsWith('{') || line.length > 1_000_000) continue;
458
+ let o;
459
+ try {
460
+ o = JSON.parse(line);
461
+ } catch {
462
+ continue;
463
+ }
464
+ if (o && typeof o === 'object' && !Array.isArray(o)) yield o;
465
+ }
466
+ }
467
+
468
+ const isObj = (v) => !!v && typeof v === 'object' && !Array.isArray(v);
469
+
470
+ // codex: only its own `error`, `turn.failed` and error-item events. Never an
471
+ // agent_message, which is the model talking.
472
+ export function codexErrorEventsText(out) {
473
+ const parts = [];
474
+ for (const o of jsonLines(out)) {
475
+ if (o.type === 'error' && typeof o.message === 'string') parts.push(o.message);
476
+ else if (o.type === 'turn.failed') {
477
+ if (isObj(o.error) && typeof o.error.message === 'string') parts.push(o.error.message);
478
+ else if (typeof o.error === 'string') parts.push(o.error);
479
+ } else if (o.type === 'item.completed' && isObj(o.item) && o.item.type === 'error' && typeof o.item.message === 'string') {
480
+ parts.push(o.item.message);
481
+ }
482
+ }
483
+ return parts.join('\n');
484
+ }
485
+
486
+ // agy: only the terminal result event's own status and error fields.
487
+ export function agyResultFieldsText(out) {
488
+ const parts = [];
489
+ for (const o of jsonLines(out)) {
490
+ if (o.event !== 'result' || !isObj(o.result)) continue;
491
+ if (typeof o.result.status === 'string') parts.push(o.result.status);
492
+ const e = o.result.error;
493
+ if (typeof e === 'string') parts.push(e);
494
+ else if (isObj(e) && typeof e.message === 'string') parts.push(e.message);
495
+ }
496
+ return parts.join('\n');
497
+ }
498
+
499
+ // codex often carries the upstream's JSON error body as a string inside
500
+ // `message`; one level is unwrapped to the human-readable text underneath.
501
+ function unwrapJsonMessage(msg) {
502
+ try {
503
+ const p = JSON.parse(msg);
504
+ if (isObj(p) && isObj(p.error) && typeof p.error.message === 'string') return p.error.message;
505
+ } catch {
506
+ /* not JSON */
507
+ }
508
+ return msg;
509
+ }
510
+
511
+ // The single most specific native codex error, preferring a top-level `error`
512
+ // event over `turn.failed`, so a problem line names the real cause instead of
513
+ // "no terminal turn.completed event".
514
+ export function codexPrimaryError(out) {
515
+ const errors = [], failed = [];
516
+ for (const o of jsonLines(out)) {
517
+ if (o.type === 'error' && typeof o.message === 'string' && o.message) errors.push(unwrapJsonMessage(o.message));
518
+ else if (o.type === 'turn.failed') {
519
+ const m = isObj(o.error) ? o.error.message : o.error;
520
+ if (typeof m === 'string' && m) failed.push(unwrapJsonMessage(m));
521
+ }
522
+ }
523
+ return errors[0] || failed[0] || null;
524
+ }
525
+
526
+ function hermesQuotaText(lower) {
527
+ return lower.includes('no usable content') || lower.includes('limit') || lower.includes('degraded');
528
+ }
529
+
530
+ // qwen: the terminal event's full error text, and a result that is an API error.
531
+ // Never the display detail, which is clipped and can carry a model name.
532
+ export function qwenErrorText(out) {
533
+ let events;
534
+ try {
535
+ events = JSON.parse(out);
536
+ } catch {
537
+ return '';
538
+ }
539
+ const term = Array.isArray(events) && events.length ? events[events.length - 1] : null;
540
+ if (!isObj(term)) return '';
541
+ const parts = [];
542
+ if (typeof term.error === 'string') parts.push(term.error);
543
+ else if (isObj(term.error) && typeof term.error.message === 'string') parts.push(term.error.message);
544
+ if (typeof term.result === 'string' && term.result.trim().startsWith('[API Error:')) parts.push(term.result);
545
+ return parts.join('\n');
546
+ }
547
+
548
+ function authoritativeBlob(lane, out, err, detail, rc) {
549
+ if (lane === 'codex') return `${codexErrorEventsText(out)}\n${err || ''}`;
550
+ if (lane === 'agy') return `${agyResultFieldsText(out)}\n${err || ''}`;
551
+ if (lane === 'hermes') return rc !== 0 ? String(err || '') : '';
552
+ if (lane === 'qwen') return `${qwenErrorText(out)}\n${err || ''}`;
553
+ return `${detail || ''}\n${err || ''}`; // grok: no auth, quota or rejected signal is defined
554
+ }
555
+
556
+ export function sigAuth(lane, blob) {
557
+ const b = blob.toLowerCase();
558
+ if (lane === 'qwen') return b.includes('missing api key');
559
+ if (lane === 'agy') return b.includes('you are not logged into antigravity') || b.includes('not authenticated');
560
+ return false; // codex, grok, hermes: no documented native auth signal
561
+ }
562
+
563
+ export function sigQuota(lane, blob) {
564
+ const b = blob.toLowerCase();
565
+ if (lane === 'qwen') return b.includes('[api error: 402') || b.includes('requires more credits') || blob.includes(' 429') || b.includes('rate limit');
566
+ if (lane === 'codex') return blob.includes('usage_limit_exceeded') || b.includes("you've hit your usage limit");
567
+ if (lane === 'hermes') return hermesQuotaText(b);
568
+ return false;
569
+ }
570
+
571
+ export function sigRejected(lane, blob) {
572
+ const b = blob.toLowerCase();
573
+ if (lane === 'qwen') return b.includes('[api error: 400') || b.includes('no endpoints found') || b.includes('failed to parse grammar');
574
+ if (lane === 'codex') return blob.includes('invalid_request_error');
575
+ if (lane === 'hermes') return b.includes('toolset');
576
+ return false;
577
+ }
578
+
579
+ // Which judge reasons mean the lane never reached a trustworthy finish, as
580
+ // opposed to finishing cleanly with nothing in it (class empty).
581
+ const CUT_SHORT_REASONS = {
582
+ grok: new Set(['not_json', 'bad_stop_reason']),
583
+ codex: new Set(['no_terminal_event']),
584
+ agy: new Set(['no_terminal_event', 'bad_last_event']),
585
+ qwen: new Set(['not_json', 'bad_event_array', 'bad_last_event', 'not_result', 'error_message_not_string']),
586
+ hermes: new Set()
587
+ };
588
+ function isCutShort(lane, reason, rc) {
589
+ if (reason === 'judge_raised') return true;
590
+ if (reason === 'ok') return true; // text came back but the vendor exited nonzero
591
+ // A nonzero vendor exit no signal explains never finished on its own terms.
592
+ // hermes' exit 2 is the one honest vendor code for "bad args, or an empty response".
593
+ if (rc !== 0) return !(lane === 'hermes' && rc === 2);
594
+ return !!(CUT_SHORT_REASONS[lane] && CUT_SHORT_REASONS[lane].has(reason));
595
+ }
596
+
597
+ // --- refused: how many tool calls a hook or deny rule blocked ----------------
598
+ // null means the lane gave no readable signal, which is an unknown, never 0.
599
+ export function refusedQwen(out) {
600
+ let events;
601
+ try {
602
+ events = JSON.parse(out);
603
+ } catch {
604
+ return null;
605
+ }
606
+ if (!Array.isArray(events) || !events.length || !isObj(events[events.length - 1])) return null;
607
+ const pd = events[events.length - 1].permission_denials;
608
+ return Array.isArray(pd) ? pd.length : null;
609
+ }
610
+
611
+ export function refusedAgy(out) {
612
+ let found = false, count = 0;
613
+ for (const o of jsonLines(out)) {
614
+ found = true;
615
+ const e = isObj(o.step_update) && isObj(o.step_update.tool_info) ? o.step_update.tool_info.error : null;
616
+ if (isObj(e) && e.type === 'TOOL_ERROR' && typeof e.message === 'string') {
617
+ const m = e.message.toLowerCase();
618
+ if (m.includes('permission check failed') || m.includes('matches user-configured deny rule')) count++;
619
+ }
620
+ }
621
+ return found ? count : null;
622
+ }
623
+
624
+ // codex's router writes one `Rejected(` line per blocked command on stderr.
625
+ // Each line counts. Prose such as "operation not permitted" in an answer does
626
+ // not: it matches a model merely explaining the phrase.
627
+ const CODEX_ROUTER_REJECTED = /codex_core::tools::router:[^\n]*Rejected\(/gi;
628
+ export function refusedCodex(out, err) {
629
+ let count = (String(err || '').match(CODEX_ROUTER_REJECTED) || []).length;
630
+ for (const o of jsonLines(out)) {
631
+ // Forward-compatible only: a structural denial item, never prose.
632
+ if (o.type === 'item.completed' && isObj(o.item) && (o.item.type === 'command_denied' || o.item.type === 'denied')) count++;
633
+ }
634
+ return count || null;
635
+ }
636
+
637
+ // grok never reports a refusal in stdout. It lives in the session transcript,
638
+ // ~/.grok/sessions/<cwd, percent-encoded>/<sessionId>/updates.jsonl, one
639
+ // {"params":{"sessionId","update"}} envelope per line. Two denial shapes are
640
+ // counted: a PreToolUse hook run with status.blocked, and grok's own deny-rule
641
+ // engine ("Denied by permission policy"). The sessionId comes from lane output,
642
+ // so it is held to a narrow charset before any path is built, the resolved path
643
+ // must stay inside the sessions root, every counted line must name the same
644
+ // session, and the read is capped.
645
+ export const GROK_SESSION_ID = /^[A-Za-z0-9-]{8,64}$/;
646
+ export const GROK_TRANSCRIPT_CAP = 5 * 1024 * 1024;
647
+ // grok encodes the cwd the way Python's urllib.parse.quote(cwd, safe="") does.
648
+ function percentEncodePath(p) {
649
+ return encodeURIComponent(p).replace(/[!'()*]/g, (c) => '%' + c.charCodeAt(0).toString(16).toUpperCase());
650
+ }
651
+ // Path-aware containment. A string-prefix test accepts a sibling such as
652
+ // "sessions-2", or on POSIX a directory literally named "sessions\outside".
653
+ export function isInsideRoot(root, p, pathApi = { relative, isAbsolute, sep }) {
654
+ const rel = pathApi.relative(root, p);
655
+ return !!rel && !pathApi.isAbsolute(rel) && rel !== '..' && !rel.startsWith('..' + pathApi.sep);
656
+ }
657
+ export function refusedGrok(out, { root = process.env.CLI_RUN_GROK_SESSIONS_ROOT || join(homedir(), '.grok', 'sessions'), cwd = process.cwd(), cap = GROK_TRANSCRIPT_CAP } = {}) {
658
+ let o;
659
+ try {
660
+ o = JSON.parse(out);
661
+ } catch {
662
+ return null;
663
+ }
664
+ const sid = isObj(o) ? o.sessionId : null;
665
+ if (typeof sid !== 'string' || !GROK_SESSION_ID.test(sid)) return null;
666
+ const path = join(root, percentEncodePath(cwd), sid, 'updates.jsonl');
667
+ let data;
668
+ let fd;
669
+ try {
670
+ const rootReal = realpathSync(root);
671
+ const pathReal = realpathSync(path);
672
+ if (!isInsideRoot(rootReal, pathReal)) return null;
673
+ if (!statSync(pathReal).isFile()) return null;
674
+ fd = openSync(pathReal, 'r');
675
+ const buf = Buffer.alloc(Math.max(0, Math.min(cap, statSync(pathReal).size)));
676
+ let read = 0;
677
+ while (read < buf.length) {
678
+ const n = readSync(fd, buf, read, buf.length - read, read);
679
+ if (!n) break;
680
+ read += n;
681
+ }
682
+ data = buf.subarray(0, read).toString('utf8');
683
+ } catch {
684
+ return null;
685
+ } finally {
686
+ if (fd !== undefined) try { closeSync(fd); } catch {}
687
+ }
688
+ let count = 0, matched = false;
689
+ for (const rec of jsonLines(data)) {
690
+ if (!isObj(rec.params) || rec.params.sessionId !== sid) continue;
691
+ matched = true;
692
+ const u = rec.params.update;
693
+ if (!isObj(u)) continue;
694
+ if (u.sessionUpdate === 'hook_execution' && Array.isArray(u.runs)) {
695
+ for (const r of u.runs) if (isObj(r) && isObj(r.status) && r.status.blocked === true) count++;
696
+ } else if (u.sessionUpdate === 'tool_call_update' && u.status === 'failed') {
697
+ let blob = '';
698
+ try { blob = JSON.stringify(u.content ?? '').toLowerCase(); } catch {}
699
+ if (blob.includes('denied by permission policy') || blob.includes('hook denied')) count++;
700
+ }
701
+ }
702
+ return matched ? count : null;
703
+ }
704
+
705
+ export function countRefused(lane, out, err, opts) {
706
+ try {
707
+ if (lane === 'qwen') return refusedQwen(out);
708
+ if (lane === 'agy') return refusedAgy(out);
709
+ if (lane === 'codex') return refusedCodex(out, err);
710
+ if (lane === 'grok') return refusedGrok(out, opts);
711
+ return null; // hermes: no native refusal signal
712
+ } catch {
713
+ return null;
714
+ }
715
+ }
716
+
717
+ // Runs after the lane has exited, so adversarial output must not make it slow:
718
+ // the input is bounded and every repeat is bounded. Redacted before any match
719
+ // can clip a secret.
720
+ const DENIAL_SCAN_BYTES = 256 * 1024;
721
+ function firstDenialText(out, err) {
722
+ const blob = redact(`${String(out || '').slice(0, DENIAL_SCAN_BYTES)}\n${String(err || '').slice(0, DENIAL_SCAN_BYTES)}`);
723
+ for (const pat of [
724
+ /Permission denied for command\([^)\n]{0,200}\)\. Matches user-configured deny rule\./,
725
+ /Denied by permission policy:[^\n"]{0,160}/,
726
+ /Hook denied:[^\n"]{0,160}/,
727
+ /denied:[^\n"]{0,160}/,
728
+ /[Rr]ejected:[^\n"]{0,160}/,
729
+ /error=exec_command failed:[^\n]{0,160}/
730
+ ]) {
731
+ const m = pat.exec(blob);
732
+ if (m) return m[0].trim();
733
+ }
734
+ return null;
735
+ }
736
+
737
+ // A lane often paraphrases a refusal in its own words; a short head of its
738
+ // answer names it when no exact denial phrase matches.
739
+ function deliverableSnippet(out, limit = 160) {
740
+ let o;
741
+ try {
742
+ o = JSON.parse(out);
743
+ } catch {
744
+ return redact(String(out || '').slice(0, 4096)).trim().slice(0, limit) || null;
745
+ }
746
+ if (!isObj(o)) return null;
747
+ for (const k of ['text', 'response', 'result']) if (typeof o[k] === 'string' && o[k].trim()) return redact(o[k].slice(0, 4096)).trim().slice(0, limit);
748
+ return null;
749
+ }
750
+
751
+ const FIX = {
752
+ auth: 'set the credential the message above names (its environment variable, or the lane\'s own login command), then rerun',
753
+ quota: 'switch to another lane, or wait for the reset time if the message gave one',
754
+ rejected: 'correct the model id, flag or request the upstream message names',
755
+ refused: 'adjust the hook or deny rule named above, or give this lane the tool it needs',
756
+ cut_short: 'rerun once; if it recurs, run without --quiet and read the lane\'s stderr on the terminal',
757
+ empty: 'rerun once, or use another lane',
758
+ timeout: 'raise --timeout, or split the brief into smaller pieces',
759
+ no_output: 'rerun once; if it recurs, check that the lane runs on its own outside cli-run'
760
+ };
761
+
762
+ // A one-line problem naming the concrete cause, and a one-line fix. The caller
763
+ // relays both. Returns { problem: null, fix: null } when there is nothing to say.
764
+ export function problemAndFix(lane, cls, { out = '', err = '', detail = '', refused = null, authoritative = null } = {}) {
765
+ const cause = typeof authoritative === 'string' && authoritative.trim() ? redact(authoritative).trim() : null;
766
+ const d = typeof detail === 'string' && detail ? redact(detail) : null;
767
+ const tag = `cli-run[${lane}]`;
768
+ const denial = () => firstDenialText(out, err) || deliverableSnippet(out) || d;
769
+ switch (cls) {
770
+ case 'auth': return { problem: `${tag} auth: ${cause || d || 'missing or invalid credentials'}`, fix: FIX.auth };
771
+ case 'quota': return { problem: `${tag} quota: ${cause || d || 'rate limit or credits exhausted'}`, fix: FIX.quota };
772
+ case 'rejected': return { problem: `${tag} rejected: ${cause || d || 'the upstream rejected the request'}`, fix: FIX.rejected };
773
+ case 'refused': return { problem: `${tag} refused: ${denial() || 'a hook or deny rule blocked the call'}`, fix: FIX.refused };
774
+ case 'cut_short': return { problem: `${tag} cut short: ${d || 'no terminal success event, cause not identifiable'}`, fix: FIX.cut_short };
775
+ case 'empty': return { problem: `${tag} empty: ${d || 'completed but delivered nothing'}`, fix: FIX.empty };
776
+ case 'timeout': return { problem: `${tag} timeout: ${d || 'exceeded the wall clock'}`, fix: FIX.timeout };
777
+ case 'unavailable': return { problem: `${tag} unavailable: ${d || 'binary not found on PATH'}`, fix: `install the ${lane} CLI and put it on PATH, or enable it in lanes.json` };
778
+ case 'no_output': return { problem: `${tag} no output: the process wrote nothing to stdout or stderr`, fix: FIX.no_output };
779
+ case 'ok':
780
+ if (Number.isInteger(refused) && refused > 0) {
781
+ return { problem: `${tag} ok, but ${refused} call(s) were refused: ${denial() || 'a hook or deny rule blocked part of the call'}`, fix: FIX.refused };
782
+ }
783
+ return { problem: null, fix: null };
784
+ default:
785
+ return { problem: null, fix: null };
786
+ }
787
+ }
788
+
789
+ // Classify a judged run. `reason`/`detail` are the judge's own (not the
790
+ // wrapper's decorated detail). A nonzero vendor exit is never ok, whatever
791
+ // came back. Never throws.
792
+ export function classifyRun(lane, { rc = 0, out = '', err = '', reason = '', detail = '', text = '', refusedOpts } = {}) {
793
+ out = typeof out === 'string' ? out : '';
794
+ err = typeof err === 'string' ? err : '';
795
+ detail = typeof detail === 'string' ? detail : '';
796
+ const refused = countRefused(lane, out, err, refusedOpts);
797
+ let cls;
798
+ try {
799
+ if (text && rc === 0) cls = 'ok';
800
+ else if (!out.trim() && !err.trim()) cls = 'no_output';
801
+ else {
802
+ const blob = authoritativeBlob(lane, out, err, detail, rc);
803
+ if (sigAuth(lane, blob)) cls = 'auth';
804
+ else if (sigQuota(lane, blob)) cls = 'quota';
805
+ else if (sigRejected(lane, blob)) cls = 'rejected';
806
+ else if (Number.isInteger(refused) && refused > 0) cls = 'refused';
807
+ else if (isCutShort(lane, reason, rc)) cls = 'cut_short';
808
+ else cls = 'empty';
809
+ }
810
+ } catch {
811
+ cls = 'cut_short';
812
+ }
813
+ return { cls, refused };
814
+ }
815
+
386
816
  // --- the process boundary --------------------------------------------------
387
817
  // The lane runs DETACHED, so it leads its own process group. On timeout (or an
388
818
  // output-buffer overrun) the whole group is killed, not just the direct child:
@@ -694,7 +1124,9 @@ function usage(msg) {
694
1124
  --model / --effort pin what a lane runs with, instead of letting it inherit its
695
1125
  own config. Every lane takes --model; every lane except qwen takes --effort.
696
1126
  Levels are the vendor's own (agy low|medium|high, hermes none|minimal|...): an
697
- unknown level is rejected by the lane, and reported as that lane's exit code.
1127
+ unknown level is rejected by the lane, and reported by class (codex: rejected, 16).
1128
+ Exit codes: 0 ok, 10 empty, 11 no output, 12 timeout, 13 unavailable, 14 auth,
1129
+ 15 quota, 16 rejected, 17 refused, 18 cut short. Failures print a problem and a fix line.
698
1130
  auto sizes per call: below 4,000 prompt characters is medium, otherwise high; a codex audit is always high. Auto never resolves above high.
699
1131
  Pin them per lane instead of per call with "defaults" in bin/lanes.json.`);
700
1132
  return USAGE;
@@ -703,7 +1135,7 @@ function usage(msg) {
703
1135
  // Bounded, control-character-free head of vendor stderr for the terminal.
704
1136
  // Never logged: provider text can echo whatever the prompt contained.
705
1137
  function stderrHead(err, n = 300) {
706
- const s = String(err || '').replace(/[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]/g, '').trim();
1138
+ const s = redact(String(err || '').replace(/[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]/g, '')).trim();
707
1139
  if (!s) return '';
708
1140
  return s.length > n ? s.slice(0, n) + '…' : s;
709
1141
  }
@@ -877,21 +1309,21 @@ export async function main(argv) {
877
1309
  effort_truncated: sizing.truncated
878
1310
  });
879
1311
  const enabled = cfg === null ? null : cfg.enabled;
880
- if (enabled === null) {
881
- console.error('cli-run: lanes.json exists but is not a valid {"enabled": [...]} file; refusing every lane until it is fixed');
882
- log({ ...base, verdict: 'unavailable', rc: UNAVAILABLE, reason: 'lanes_json_malformed' });
1312
+ const unavailable = ({ msg, reason, fix }) => {
1313
+ console.error('cli-run: ' + msg);
1314
+ if (!opts.quiet) console.error(`cli-run fix: ${fix}`);
1315
+ log({ ...base, verdict: 'unavailable', class: 'unavailable', rc: UNAVAILABLE, refused: null, reason });
883
1316
  return UNAVAILABLE;
1317
+ };
1318
+ if (enabled === null) {
1319
+ return unavailable({ msg: 'lanes.json exists but is not a valid {"enabled": [...]} file; refusing every lane until it is fixed', reason: 'lanes_json_malformed', fix: 'fix bin/lanes.json, or delete it to enable every lane' });
884
1320
  }
885
1321
  if (!enabled.includes(lane)) {
886
- console.error(`cli-run: ${lane} is not enabled in lanes.json`);
887
- log({ ...base, verdict: 'unavailable', rc: UNAVAILABLE, reason: 'disabled' });
888
- return UNAVAILABLE;
1322
+ return unavailable({ msg: `${lane} is not enabled in lanes.json`, reason: 'disabled', fix: `add "${lane}" to "enabled" in bin/lanes.json, or use an enabled lane` });
889
1323
  }
890
1324
  const binary = which(lane);
891
1325
  if (!binary) {
892
- console.error(`cli-run: ${lane} not found on PATH`);
893
- log({ ...base, verdict: 'unavailable', rc: UNAVAILABLE, reason: 'unavailable' });
894
- return UNAVAILABLE;
1326
+ return unavailable({ msg: `${lane} not found on PATH`, reason: 'unavailable', fix: `install the ${lane} CLI and put it on PATH` });
895
1327
  }
896
1328
 
897
1329
  if (route.effort === 'auto' && !opts.quiet) console.error(`cli-run: effort auto -> ${sizing.resolved} (${sizing.basis})`);
@@ -903,52 +1335,60 @@ export async function main(argv) {
903
1335
  const r = await runBounded(cmd, opts.timeout);
904
1336
  const out = r.stdout || '';
905
1337
  const err = r.stderr || '';
906
- let verdict, reason, detail, code, text = '';
1338
+ let verdict, reason, detail, cls, text = '', refused = null;
907
1339
  if (r.interrupted) {
908
- verdict = 'interrupted'; reason = 'killed'; detail = `cli-run received ${r.interrupted}; the lane's process group was killed`; code = 128 + (r.interrupted === 'SIGINT' ? 2 : 15);
1340
+ verdict = 'interrupted'; reason = 'killed'; detail = `cli-run received ${r.interrupted}; the lane's process group was killed`; cls = 'interrupted';
909
1341
  } else if (r.timedOut) {
910
- verdict = 'timeout'; reason = 'timeout'; detail = `exceeded ${opts.timeout}s; process group killed`; code = TIMEOUT;
1342
+ verdict = 'timeout'; reason = 'timeout'; detail = `exceeded ${opts.timeout}s; process group killed`; cls = 'timeout';
911
1343
  } else if (r.overrun) {
912
- verdict = 'no_deliverable'; reason = 'no_output'; detail = 'output exceeded the 16 MiB buffer; process group killed'; code = NO_DELIVERABLE;
1344
+ verdict = 'no_deliverable'; reason = 'no_output'; detail = 'output exceeded the 16 MiB buffer; process group killed'; cls = 'cut_short';
913
1345
  } else if (r.error) {
914
- verdict = 'unavailable'; reason = 'unavailable'; detail = r.error.message; code = UNAVAILABLE;
1346
+ verdict = 'unavailable'; reason = 'unavailable'; detail = r.error.message; cls = 'unavailable';
915
1347
  } else if (r.signal || r.status === null) {
916
1348
  // A lane killed by a signal has no honest exit status. Whatever it printed
917
1349
  // before dying is not a deliverable; a null status must never become exit 0.
918
- verdict = 'killed'; reason = 'killed'; detail = `lane killed by ${r.signal || 'unknown signal'}`; code = NO_DELIVERABLE;
1350
+ verdict = 'killed'; reason = 'killed'; detail = `lane killed by ${r.signal || 'unknown signal'}`; cls = 'cut_short';
919
1351
  } else {
920
- const j = judge(lane, r.status, out, err, outFile);
921
- text = j.text || '';
1352
+ const j = safeJudge(lane, r.status, out, err, outFile);
922
1353
  reason = j.reason;
923
1354
  detail = j.detail;
1355
+ ({ cls, refused } = classifyRun(lane, { rc: r.status, out, err, reason: j.reason, detail: j.detail, text: j.text || '' }));
924
1356
  if (r.status !== 0) {
925
1357
  // A nonzero vendor exit is a failure on the vendor's own terms, whether or
926
- // not something parseable came back. Pass the code through, keep the
927
- // verdict honest, and show what the vendor said on stderr.
928
- verdict = 'exit_nonzero'; code = r.status;
1358
+ // not something parseable came back. The class names why; the vendor's
1359
+ // own code is kept as cli_rc, and its stderr head shows on the terminal.
1360
+ verdict = 'exit_nonzero';
929
1361
  if (reason === 'ok') reason = 'exit_nonzero';
930
1362
  const head = stderrHead(err);
931
1363
  detail = `lane exited ${r.status}` + (head ? `; stderr: ${head}` : '') + (j.reason !== 'ok' ? `; ${j.detail}` : '');
932
- } else if (text) {
933
- const unmet = checkContracts(opts, text, before);
1364
+ } else if (cls === 'ok') {
1365
+ const unmet = checkContracts(opts, j.text, before);
934
1366
  if (unmet) {
935
- verdict = 'no_deliverable'; reason = 'contract_unmet'; detail = unmet; code = NO_DELIVERABLE;
1367
+ verdict = 'no_deliverable'; reason = 'contract_unmet'; detail = unmet; cls = 'empty';
936
1368
  } else {
937
- verdict = 'ok'; code = OK;
1369
+ verdict = 'ok'; text = j.text;
938
1370
  }
939
- } else if (!out.trim() && !err.trim()) {
940
- verdict = 'no_output'; reason = 'no_output'; code = NO_OUTPUT;
1371
+ } else if (cls === 'no_output') {
1372
+ verdict = 'no_output'; reason = 'no_output';
941
1373
  } else {
942
- verdict = 'no_deliverable'; code = NO_DELIVERABLE;
1374
+ verdict = 'no_deliverable';
943
1375
  const head = stderrHead(err);
944
1376
  if (head) detail += `; stderr: ${head}`;
945
1377
  }
946
1378
  }
1379
+ const code = cls === 'interrupted' ? 128 + (r.interrupted === 'SIGINT' ? 2 : 15) : CLASS_CODES[cls];
947
1380
  if (text && code === OK) process.stdout.write(text + '\n');
948
1381
  const routeNote = route.model || route.effort ? `${route.model || 'lane default'}/${route.effort || 'lane default'}` : 'lane default';
949
- if (!opts.quiet) console.error(`cli-run[${lane}] ${verdict} rc=${code} ${r.seconds.toFixed(1)}s raw=${r.outBytes || 0}B route=${routeNote} :: ${detail}`);
950
- // Durable log: fixed reason code and structural numbers only.
951
- log({ ...base, verdict, rc: code, cli_rc: r.status, signal: r.signal || null, seconds: Math.round(r.seconds * 100) / 100, raw_bytes: r.outBytes || 0, deliverable_bytes: Buffer.byteLength(text), reason: REASONS.has(reason) ? reason : 'unknown' });
1382
+ if (!opts.quiet) {
1383
+ console.error(`cli-run[${lane}] ${verdict} rc=${code} class=${cls} refused=${refused === null ? 'null' : refused} ${r.seconds.toFixed(1)}s raw=${r.outBytes || 0}B route=${routeNote} :: ${redact(detail)}`);
1384
+ let authoritative = null;
1385
+ if (lane === 'codex' && (cls === 'quota' || cls === 'rejected')) authoritative = codexPrimaryError(out);
1386
+ const pf = problemAndFix(lane, cls, { out, err, detail, refused, authoritative });
1387
+ if (pf.problem) console.error('cli-run problem: ' + redact(pf.problem).slice(0, 600));
1388
+ if (pf.fix) console.error('cli-run fix: ' + pf.fix);
1389
+ }
1390
+ // Durable log: fixed reason code, fixed class and structural numbers only.
1391
+ log({ ...base, verdict, class: cls in CLASS_CODES || cls === 'interrupted' ? cls : 'unknown', rc: code, cli_rc: r.status, signal: r.signal || null, refused: Number.isInteger(refused) ? refused : null, seconds: Math.round(r.seconds * 100) / 100, raw_bytes: r.outBytes || 0, deliverable_bytes: Buffer.byteLength(text), reason: REASONS.has(reason) ? reason : 'unknown' });
952
1392
  return code;
953
1393
  } finally {
954
1394
  rmSync(tmp, { recursive: true, force: true });
@@ -126,3 +126,23 @@ Three findings reproduced against the 0.1.15 branch before it shipped, none of t
126
126
  - **Not yet attacked for real.** `windowsSpawnPlan`, `resolveCmdShim` and the escaping functions are unit-tested (pure string logic, runs on every CI host) and the resolved-shim path is exercised end to end on `windows-latest` through the fake-lane fixtures in `test/cli.test.js` (installed as a real npm-style `.cmd` shim). A lane can no longer reach `cmd.exe` at all (a unit test pins the refusal). The opt-in `cmd.exe` branch, used only by the installer's own `npm install -g`, is not exercised end to end through a live Windows process in this suite; its escaping is proven by exact-string unit tests only, and its arguments never include user text.
127
127
  - **A platform limit found the same way, unrelated to the spawn path itself: Windows has no OS-level signals at all.** `cli-run.mjs`'s graceful shutdown (`process.on('SIGTERM', ...)`, kill the lane's process group, then exit 143/130) is a POSIX guarantee only: `ChildProcess.kill(sig)` on Windows calls `TerminateProcess()` unconditionally for SIGTERM AND SIGINT alike, giving the target process no chance to run any handler at all, proven on `windows-latest` CI (the wrapper died as `{code: null, signal: sig}` for both; a hypothesis that SIGINT gets a real, catchable console-control event on Windows was tried first and measured false in this exact scenario, not assumed). `test/cli.test.js`'s `#13` now expects an unhandled termination for either signal on win32, and the original graceful-exit assertion elsewhere.
128
128
  - **Two narrow, individually-verified Windows skips remain, neither in the spawn path itself.** (1) A lane dying mid-run from a real POSIX signal cannot be reproduced on win32: a real Windows lane is a plain `node <script>` process, so it cannot die "by signal" any more than the product being tested can, and the only way a test fixture can even simulate one (a nested `sh -c "...; kill -TERM $$"`) puts an extra node process between cli-run.mjs and the dying shell, so cli-run.mjs observes only that node's translated exit code (measured: MSYS bash's self-kill status leaks through as a plain nonzero exit code, 3840, which this tool already handles honestly via `exit_nonzero`). (2) `weekly-audit.sh`'s watchdog (`bounded()`/`killtree()`, `pgrep -P` plus killing a backgrounded subshell's tree) relies on real bash job control this script only ever runs under on the Ubuntu box it targets; actually executing it against a genuinely hanging stub under Git Bash's job-control emulation hung past a 20s outer timeout on `windows-latest` CI, a known class of MSYS/Cygwin limitation (a `kill -KILL` not reliably reaching the underlying Windows process tree of a backgrounded subshell), not a defect in the generated script, which still renders and syntax-checks correctly.
129
+
130
+ ## New in 0.1.23: failure classes in cli-run
131
+
132
+ `bin/cli-run.mjs` now puts every run in a closed failure class that owns the exit code (`auth` 14, `quota` 15, `rejected` 16, `refused` 17, `cut_short` 18, beside `ok` 0, `empty` 10, `no_output` 11, `timeout` 12, `unavailable` 13), counts refused tool calls per lane (including a read of grok's session transcript, whose `sessionId` comes from lane stdout), and prints redacted problem and fix lines. The durable log gains only `class` and `refused`. Threat model: a false exit 0, provider text or a secret reaching the log or the terminal, a transcript read escaping the sessions root, misclassification that sends a user the wrong way, and anything that throws or stalls.
133
+
134
+ ### Round 1 (pre-release, GPT-6 Astra at xhigh), fixed before shipping
135
+
136
+ Every finding was reproduced as a failing case in `test/classify.test.js` against the unfixed code (all seven red), then fixed; the suite is green after. One round, by rule; the regression cases verify the fixes.
137
+
138
+ | # | Sev | Finding | Fix | Test |
139
+ |---|---|---|---|---|
140
+ | 1 | HIGH | qwen's display detail clipped an error message at 120 characters before redaction, so a long JSON password lost its closing quote and printed its prefix | redact before every clip: judge detail, API-error text, denial text, deliverable snippet, problem cause | R1 |
141
+ | 2 | HIGH | the JSON credential pattern ended at an escaped quote and missed unterminated values | match JSON string escapes, run an unterminated value to the end, and accept a JSON body escaped inside a string | R2 |
142
+ | 3 | HIGH | `firstDenialText` scanned 16 MiB with an unbounded `[^)]*`: 1 MiB of unclosed `Permission denied for command(` took 15.6 s after the lane had exited | scan at most 256 KiB of stdout and of stderr, with bounded repeats | R3 |
143
+ | 4 | MEDIUM | transcript containment was a string prefix with both separators, so on POSIX a directory literally named `sessions\outside` passed | `isInsideRoot()` via `path.relative`, rejecting `..`, absolute and cross-drive results; tested with `path.posix` and `path.win32` on every OS | R4 |
144
+ | 5 | MEDIUM | qwen signals were searched in the clipped display detail, so a quota error after 120 characters read as `empty`, and a model name in the detail could fake a signal | classify on `qwenErrorText()`: the terminal event's full error and an API-error result only | R5 |
145
+ | 6 | MEDIUM | hermes fell back to stdout when stderr was empty, so an answer saying "no rate limit" read as `quota` | stderr only | R6 |
146
+ | 7 | MEDIUM | a nonzero vendor exit with a recognised empty shape (agy exit 7, `SUCCESS`, empty response) returned `empty`, contradicting the documented `cut_short` fallback | an unexplained nonzero exit is `cut_short`, except hermes' own exit 2 | R7 |
147
+
148
+ Reported clean by the same audit: no false exit 0 across 675 malformed probes and five lanes; no provider text or secret fragment in probed log records; the real codex 0.153.4 fixture's non-fatal error item still succeeds; grok traversal, symlink, FIFO, session-match and read-cap guards; timeout, interrupt, overrun, killed lane, JSON contract, `--quiet` and the doctor canary. Not verified by it: native Windows execution, live vendors, filesystem races.
package/docs/catalog.md CHANGED
@@ -63,6 +63,7 @@ Generated from `src/catalog.js`. Do not hand-edit; `npm run gen:catalog` rewrite
63
63
  - **cli-run lane:** yes
64
64
  - **Plans:**
65
65
  - SuperGrok (base headroom, checked 2026-09-12): https://x.ai/news/grok-build-cli
66
+ - SuperGrok Plus (high headroom, checked 2026-09-12): https://x.ai/pricing
66
67
  - X Premium Plus (base headroom, checked 2026-09-12): https://x.ai/news/grok-build-cli
67
68
  - **Built against:** 1.0.5
68
69
 
@@ -28,7 +28,7 @@ One driver, no second AI in the mix: the orchestrator invokes the CLIs; it never
28
28
 
29
29
  ## 3. Exit 0 is a lie on every lane
30
30
 
31
- Every agent CLI can report success and deliver nothing. `bin/cli-run.mjs` builds the right invocation per lane, reads that lane's **native** terminal event, and exits `10` when a run produced no deliverable, `12` on timeout, `13` when the lane is missing. Byte count is not a check either; a run can emit hundreds of kilobytes and no conclusion. One lane's own success flags lie outright (an upstream 400 reported as success), so its judge reads the two honest signals instead.
31
+ Every agent CLI can report success and deliver nothing. `bin/cli-run.mjs` builds the right invocation per lane, reads that lane's **native** terminal event, and exits `10` when a run produced no deliverable, `12` on timeout, `13` when the lane is missing, and `14` to `18` when the lane's own error says why (auth, quota, rejected, refused, cut short), so a missing API key is never blamed on the model. Byte count is not a check either; a run can emit hundreds of kilobytes and no conclusion. One lane's own success flags lie outright (an upstream 400 reported as success), so its judge reads the two honest signals instead.
32
32
 
33
33
  Every call goes through it. "This lane is flaky" becomes a query over its log instead of an argument. `node bin/cli-run.mjs --doctor` is the first thing to run after install: enabled lanes, binaries on PATH, the route each lane is pinned to, and with `--run` a one-word canary per lane.
34
34
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "0.1.21",
3
+ "version": "0.1.23",
4
4
  "description": "Model orchestrator for AI coding agents and LLMs: Claude Code, Codex, Gemini, Grok, Qwen, Ollama. Routing rules tell your agent which model, subagent or CLI to use for each task, so small work goes to cheap tiers and fewer tokens go to frontier models. One installer, plus a CLI runner that logs every route.",
5
5
  "type": "module",
6
6
  "bin": {
package/src/catalog.js CHANGED
@@ -150,6 +150,7 @@ export const AIS = [
150
150
  cliRun: true
151
151
  , plans: [
152
152
  { id: 'supergrok', name: 'SuperGrok', headroom: 'base', source: 'https://x.ai/news/grok-build-cli', checked: '2026-09-12' },
153
+ { id: 'supergrok-plus', name: 'SuperGrok Plus', headroom: 'high', source: 'https://x.ai/pricing', checked: '2026-09-12' },
153
154
  { id: 'x-premium-plus', name: 'X Premium Plus', headroom: 'base', source: 'https://x.ai/news/grok-build-cli', checked: '2026-09-12' }
154
155
  ]
155
156
  },
@@ -7,7 +7,7 @@
7
7
  # - the previous successful report is NEVER truncated: output goes to a temp file and is
8
8
  # renamed into place only on a clean exit; failed output is kept beside it for diagnosis
9
9
  # - the lane runs with the strongest boundary it offers ({{AUDIT_LANE_BOUNDARY_NOTE}})
10
- # - rc 10/12/13 from cli-run means no report was produced; the timer's journal shows it
10
+ # - any nonzero rc from cli-run (10 to 18) means no report was produced; the timer's journal shows it
11
11
  set -uo pipefail
12
12
  INSTALL_DIR={{INSTALL_DIR_SH}}
13
13
  AUDIT_LANE="{{AUDIT_LANE}}"
@@ -56,16 +56,40 @@ qwen is the lane whose own success flags lie: an upstream 400 comes back as exit
56
56
 
57
57
  ## Exit codes
58
58
 
59
- | Code | Meaning |
60
- |---|---|
61
- | 0 | structurally accepted non-empty response, every `--expect-*` contract met |
62
- | 10 | ran and produced no deliverable, or a contract was unmet, or the lane was killed by a signal, or output overran the 16 MiB buffer |
63
- | 11 | no output at all |
64
- | 12 | timed out; the lane and every descendant in its process group were killed |
65
- | 13 | lane unavailable: binary missing, disabled in `lanes.json`, or `lanes.json` malformed |
66
- | 130 / 143 | cli-run itself received SIGINT / SIGTERM; the lane's process group was killed first, then the temp dir removed |
67
- | 2 | usage error in cli-run itself |
68
- | N | the lane exited N != 0: passed through unchanged, verdict `exit_nonzero`, even when parseable text came back. The bounded head of the lane's stderr is shown on your terminal so an auth failure reads as one |
59
+ | Code | Class | Meaning | What to do |
60
+ |---|---|---|---|
61
+ | 0 | `ok` | structurally accepted non-empty response, every `--expect-*` contract met | use it; if `refused=N` is above 0, read the problem line |
62
+ | 10 | `empty` | ran and delivered nothing, or a contract was unmet | rerun once, or use another lane |
63
+ | 11 | `no_output` | no output at all | rerun once; check the lane runs on its own |
64
+ | 12 | `timeout` | timed out; the lane and every descendant in its process group were killed | raise `--timeout` or split the brief |
65
+ | 13 | `unavailable` | binary missing, disabled in `lanes.json`, or `lanes.json` malformed | install or enable the lane |
66
+ | 14 | `auth` | the lane's own error says a credential is missing or it is not logged in | set the credential; retrying cannot help |
67
+ | 15 | `quota` | the lane's own error says usage limit, credits or rate limit | switch lanes or wait for the reset |
68
+ | 16 | `rejected` | the upstream rejected the request: an unknown model id, a malformed request | fix the id, flag or request it names |
69
+ | 17 | `refused` | no deliverable, and the lane reports tool calls a hook or deny rule blocked | adjust the rule, or give the lane the tool |
70
+ | 18 | `cut_short` | no trustworthy finish: a missing or non-success terminal event, a lane killed by a signal, output past the 16 MiB buffer, or a nonzero vendor exit nothing above explains (hermes' own exit 2, bad args or an empty response, stays `empty`) | rerun once, then read the lane's stderr |
71
+ | 130 / 143 | `interrupted` | cli-run itself received SIGINT / SIGTERM; the lane's process group was killed first, then the temp dir removed | |
72
+ | 2 | | usage error in cli-run itself | |
73
+
74
+ A nonzero vendor exit is never `ok`, even when parseable text came back; the vendor's own code is kept in the log as `cli_rc`, and the bounded head of its stderr is shown on your terminal.
75
+
76
+ ## Why a run failed: the class
77
+
78
+ Exit 10 used to cover causes that need opposite responses. A missing API key is not a model fault, and retrying a spent quota cannot help. So every run lands in exactly one class, and on the terminal every failure (and every `ok` with refused calls) gets two more lines:
79
+
80
+ ```
81
+ cli-run[qwen] exit_nonzero rc=14 class=auth refused=0 0.8s raw=212B route=lane default :: lane exited 1; subtype="error_during_execution": Missing API key ...
82
+ cli-run problem: cli-run[qwen] auth: Missing API key ...
83
+ cli-run fix: set the credential the message above names (its environment variable, or the lane's own login command), then rerun
84
+ ```
85
+
86
+ A calling agent relays both lines and asks before fixing anything.
87
+
88
+ - **Signals come from each lane's authoritative error fields only.** codex: its own `error`, `turn.failed` and error-item events plus stderr. agy: the terminal result's status and error plus stderr. qwen: the terminal event's full error text (never the clipped display line) plus stderr. hermes: stderr on a nonzero exit, never its stdout. Never the model's prose, so an answer that explains what "rate limit" means is not a quota failure. grok and hermes have no native auth signal, and none is invented.
89
+ - **Precedence** when several are present: `auth`, `quota`, `rejected`, `refused`, `cut_short`, `empty`. A missing credential explains everything downstream of it.
90
+ - **`refused`** counts tool calls a hook or deny rule blocked: qwen's `permission_denials`, agy's deny-rule `TOOL_ERROR` steps, one per codex router `Rejected(` line on stderr, and grok's session transcript (`~/.grok/sessions/<cwd>/<sessionId>/updates.jsonl`: hook runs with `blocked`, and "Denied by permission policy"). grok's `sessionId` comes from lane output, so it must match `^[A-Za-z0-9-]{8,64}$`, the resolved path must stay inside the sessions root, every counted line must name that session, and the read stops at 5 MiB. `null` means the lane gave no readable signal (hermes always), which is an unknown, never 0.
91
+ - **Refused with a deliverable is still exit 0.** The deliverable exists; `refused=N` and the problem and fix lines say what was blocked.
92
+ - **Redacted.** The status, problem and stderr lines pass through one redaction pass before they are printed: JSON credential keys (escaped quotes, unterminated values and JSON escaped inside a string included), `Authorization:` values of any scheme, bearer values, URL query credentials, and the `sk-`, `xai-`, `ghp_` and `AIza` key prefixes. Redaction runs before any clipping, so a long token is never cut into an unrecognisable fragment. Denial text is searched in at most the first 256 KiB of stdout and of stderr, with bounded patterns, so hostile output cannot stall the wrapper after the lane has exited. None of these lines is ever logged.
69
93
 
70
94
  ## The route: which model, and how hard it thinks
71
95
 
@@ -97,7 +121,7 @@ A flag beats `defaults`; `defaults` beats nothing. Each vendor spells these diff
97
121
 
98
122
  Three rules that keep this honest:
99
123
 
100
- - **A level `cli-run` does not recognise is not rejected here.** Levels are the vendor's, they change, and guessing the valid set would date this tool. An unknown level is refused by the lane and surfaces as that lane's own exit code and stderr.
124
+ - **A level `cli-run` does not recognise is not rejected here.** Levels are the vendor's, they change, and guessing the valid set would date this tool. An unknown level is refused by the lane and surfaces as a class (codex reports it as `rejected`, exit 16; a lane with no rejection signal as `cut_short`, exit 18), with the lane's own exit code in the log as `cli_rc` and its stderr on your terminal.
101
125
  - **`--effort` on qwen is a usage error, not a silent drop.** A flag that vanishes leaves you believing a route that never ran.
102
126
  - **`--effort auto` is bounded.** It uses prompt size, or an audit's changed-file evidence, and resolves only medium or high. An audit is always high. Auto is a heuristic, not a measurement: name `xhigh` explicitly for security-critical or irreversible work.
103
127
  - **Values are charset-bounded** (letters, digits, and `. _ : @ / + -`, no leading dash, 64 characters). A model id becomes an argv element and, on codex, part of a TOML value; bounding it is what stops either from being escaped.
@@ -108,7 +132,7 @@ Three rules that keep this honest:
108
132
 
109
133
  ## Log
110
134
 
111
- `~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, rc, the lane's own exit code, signal, seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, the route (`model_requested`, `effort_requested`, and `model_source` / `effort_source`, each one of `flag`, `lanes.json` or `lane_default`), auto-sizing evidence (`effort_resolved`, `effort_basis`, `effort_scope`), and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). `effort_basis` is exactly one of `explicit`, `prompt_chars`, `audit_floor`, or `none`. Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument, and so does "we route audits at high effort".
135
+ `~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, `class` (one of the classes above), rc, the lane's own exit code (`cli_rc`), signal, `refused` (an integer, or `null` when the lane gives no signal), seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, the route (`model_requested`, `effort_requested`, and `model_source` / `effort_source`, each one of `flag`, `lanes.json` or `lane_default`), auto-sizing evidence (`effort_resolved`, `effort_basis`, `effort_scope`), and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). `effort_basis` is exactly one of `explicit`, `prompt_chars`, `audit_floor`, or `none`. Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument, and so does "we route audits at high effort".
112
136
 
113
137
  The log records what was **requested**, on every record including a run refused before the lane started. It does not record an actual. Reporting is inconsistent: grok returns a `modelUsage` block naming a model, the other four lanes return nothing of the kind, so an `actual` field would be populated for one lane and empty for four. It would also be a provider-supplied string, and this log holds fixed codes and bounded caller-supplied values only. `model_source: "lane_default"` is the honest way to say this run inherited something invisible from here.
114
138
 
@@ -122,7 +146,7 @@ Absent: every lane enabled, nothing pinned. Present but malformed or unreadable:
122
146
 
123
147
  ## A killed lane is not a deliverable
124
148
 
125
- A lane that dies by signal has no honest exit status. Whatever it printed first is discarded; the run reports `killed` with exit 10.
149
+ A lane that dies by signal has no honest exit status. Whatever it printed first is discarded; the run reports `killed`, class `cut_short`, exit 18.
126
150
 
127
151
  ## Interrupting cli-run kills the lane too
128
152