model-orchestrator 1.0.5 → 1.0.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,18 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [1.0.6] - 2026-09-29
8
+
9
+ ### Added
10
+
11
+ - The hermes lane takes `--provider` and `"provider"` in `bin/lanes.json` defaults, so a pinned model goes to a provider that serves it. A provider needs a model with it (a flag, a default or `HERMES_INFERENCE_MODEL`); otherwise cli-run refuses before the lane starts. Runs log `provider_requested` and `provider_source`.
12
+ - `--doctor` notes a hermes model pinned without a provider, and a provider pinned without a model. The terminal route line shows the provider.
13
+
14
+ ### Fixed
15
+
16
+ - A hermes model/provider mismatch (an upstream "model is not supported" or `model_not_found`, for example `grok-4.6` sent to `openai-codex`) is now class `rejected`, exit 16, with a problem line naming the mismatch and a fix pointing at `hermes model`. It was reported as a vague nonzero exit.
17
+ - A pinned hermes route (`--model` or `--effort`, by flag or `lanes.json` default) no longer fails every run. cli-run placed those flags between `-z` and the prompt, and `-z` takes the prompt as its value, so Hermes refused the arguments ("argument -z/--oneshot: expected one argument"). Route flags now come before `-z`, and a Hermes argument error is reported as such rather than as a toolsets error.
18
+
7
19
  ## [1.0.5] - 2026-09-28
8
20
 
9
21
  ### Changed
package/bin/cli-run.mjs CHANGED
@@ -53,6 +53,7 @@
53
53
  // problem/fix lines go to your terminal only, redacted.
54
54
  //
55
55
  // ROUTE: which model and reasoning effort a lane ran with.
56
+ // Hermes also accepts --provider to select the provider serving that model.
56
57
  // A lane with no --model and no lanes.json default inherits whatever its own
57
58
  // config file says, which is invisible from here and is how a documented route
58
59
  // silently stops being the route that runs. --model / --effort pin it per call,
@@ -199,7 +200,11 @@ export function judgeHermes(rc, out, err) {
199
200
  // hermes collapses every upstream failure into one exit code; its stderr
200
201
  // is the only place the cause is named.
201
202
  const blob = String(err || '').toLowerCase(); // stderr only: stdout is the agent's own prose
202
- if (hermesQuotaText(blob)) why += ': upstream free tier degraded or limited, a retry is reasonable';
203
+ // A mismatch is checked first: a retry cannot fix it, and an unrelated
204
+ // "limit" warning on stderr must not turn it into a quota report.
205
+ if (hermesMismatchLine(err) || hermesMismatchLine(out, { vendorShape: true })) why += ': model/provider mismatch, the provider does not serve this model';
206
+ else if (hermesQuotaText(blob)) why += ': upstream free tier degraded or limited, a retry is reasonable';
207
+ else if (hermesArgError(blob)) why += ': hermes rejected its arguments, a caller bug and not a lane fault';
203
208
  else if (blob.includes('toolset')) why += ': invalid --toolsets value, a caller bug and not a lane fault';
204
209
  return fail('exit_nonzero', `hermes exit ${rc}: ${why}`);
205
210
  }
@@ -253,14 +258,18 @@ export function judgeQwen(rc, out) {
253
258
  // grok -m MODEL --reasoning-effort EFFORT
254
259
  // codex -m MODEL -c model_reasoning_effort="EFFORT" (a TOML override, hence the quotes)
255
260
  // agy --model M --effort EFFORT (low|medium|high)
256
- // hermes -m MODEL --reasoning LEVEL (none|minimal|...)
261
+ // hermes -m MODEL --reasoning LEVEL --provider ID (none|minimal|...)
257
262
  // qwen -m MODEL no reasoning flag
263
+ // Only hermes takes a provider: one hermes install signs in to many providers,
264
+ // and a model id sent to a provider that does not serve it is an HTTP 400
265
+ // (grok-4.6 on openai-codex). `hermes -z --provider` without a model exits 2,
266
+ // so cli-run refuses that pairing before the lane starts.
258
267
  export const LANE_FLAGS = {
259
- grok: { model: (v) => ['-m', v], effort: (v) => ['--reasoning-effort', v] },
260
- codex: { model: (v) => ['-m', v], effort: (v) => ['-c', `model_reasoning_effort="${v}"`] },
261
- agy: { model: (v) => ['--model', v], effort: (v) => ['--effort', v] },
262
- hermes: { model: (v) => ['-m', v], effort: (v) => ['--reasoning', v] },
263
- qwen: { model: (v) => ['-m', v], effort: null }
268
+ grok: { model: (v) => ['-m', v], effort: (v) => ['--reasoning-effort', v], provider: null },
269
+ codex: { model: (v) => ['-m', v], effort: (v) => ['-c', `model_reasoning_effort="${v}"`], provider: null },
270
+ agy: { model: (v) => ['--model', v], effort: (v) => ['--effort', v], provider: null },
271
+ hermes: { model: (v) => ['-m', v], effort: (v) => ['--reasoning', v], provider: (v) => ['--provider', v] },
272
+ qwen: { model: (v) => ['-m', v], effort: null, provider: null }
264
273
  };
265
274
 
266
275
  // Auto is deliberately a small, static ladder. It is not a vendor capability
@@ -353,13 +362,16 @@ export function badRouteValue(kind, v) {
353
362
  }
354
363
 
355
364
  // --- adapters: build argv for a lane -------------------------------------
356
- // Route flags go in front of the prompt for every lane, because two lanes
357
- // (hermes, codex) take the prompt as a positional argument and a flag after it
358
- // is either ignored or read as part of it.
365
+ // Route flags go in front of the prompt for every lane, because codex takes the
366
+ // prompt as a positional argument and a flag after it is either ignored or read
367
+ // as part of it. hermes needs them in front of -z itself: -z takes the prompt as
368
+ // its own value, so `-z -m X prompt` is an argparse error ("argument -z/--oneshot:
369
+ // expected one argument") and every pinned hermes route failed that way.
359
370
  function routeFlags(lane, opts) {
360
371
  const spec = LANE_FLAGS[lane];
361
372
  const out = [];
362
373
  if (!spec) return out;
374
+ if (opts.provider && spec.provider) out.push(...spec.provider(opts.provider));
363
375
  if (opts.model) out.push(...spec.model(opts.model));
364
376
  if (opts.effort && spec.effort) out.push(...spec.effort(opts.effort));
365
377
  return out;
@@ -384,7 +396,7 @@ export function buildArgv(lane, binary, prompt, opts, tmp) {
384
396
  return { argv: [binary, '--print-timeout', `${mins}m`, '--output-format', 'stream-json', ...route, '-p', prompt] };
385
397
  }
386
398
  case 'hermes':
387
- return { argv: [binary, '-z', ...route, prompt, '--usage-file', join(tmp, 'usage.json')] };
399
+ return { argv: [binary, ...route, '-z', prompt, '--usage-file', join(tmp, 'usage.json')] };
388
400
  case 'qwen': {
389
401
  const argv = [binary, '-o', 'json', ...route];
390
402
  if (opts.safeMode) argv.push('--safe-mode');
@@ -527,6 +539,30 @@ function hermesQuotaText(lower) {
527
539
  return lower.includes('no usable content') || lower.includes('limit') || lower.includes('degraded');
528
540
  }
529
541
 
542
+ // argparse's own refusal. Its usage dump always lists `-t TOOLSETS`, so this is
543
+ // checked before the toolset match or every bad argument reads as a toolset error.
544
+ function hermesArgError(lower) {
545
+ return lower.includes('hermes: error:');
546
+ }
547
+
548
+ // A model id the chosen provider does not serve. `hermes -z` prints a failed
549
+ // turn's provider error on stdout and exits 2, so this reads both streams, but
550
+ // returns only the matching line: the rest of stdout is the agent's own prose.
551
+ // The wording is the upstream's own ("The 'grok-4.6' model is not supported
552
+ // when using Codex with a ChatGPT account"; OpenAI's model_not_found code).
553
+ export function hermesMismatchLine(text, { vendorShape = false } = {}) {
554
+ for (const line of String(text || '').split('\n')) {
555
+ const l = line.toLowerCase();
556
+ if (!(l.includes('model is not supported') || l.includes('model_not_found'))) continue;
557
+ // stdout is agent prose, so there only the vendor's own shape counts: an HTTP
558
+ // status, an error prefix, the Codex wording, or the OpenAI error code.
559
+ if (vendorShape && !(/^\s*(http [45]\d\d|error)\b/.test(l) || l.includes('model is not supported when using') || l.includes('model_not_found'))) continue;
560
+ // Terminal-safe: stdout can carry web-derived text, so no control characters.
561
+ return line.replace(/[\x00-\x08\x0b-\x1f\x7f]/g, '').trim().slice(0, 300);
562
+ }
563
+ return null;
564
+ }
565
+
530
566
  // qwen: the terminal event's full error text, and a result that is an API error.
531
567
  // Never the display detail, which is clipped and can carry a model name.
532
568
  export function qwenErrorText(out) {
@@ -548,7 +584,7 @@ export function qwenErrorText(out) {
548
584
  function authoritativeBlob(lane, out, err, detail, rc) {
549
585
  if (lane === 'codex') return `${codexErrorEventsText(out)}\n${err || ''}`;
550
586
  if (lane === 'agy') return `${agyResultFieldsText(out)}\n${err || ''}`;
551
- if (lane === 'hermes') return rc !== 0 ? String(err || '') : '';
587
+ if (lane === 'hermes') return rc !== 0 ? `${err || ''}\n${hermesMismatchLine(out, { vendorShape: true }) || ''}` : '';
552
588
  if (lane === 'qwen') return `${qwenErrorText(out)}\n${err || ''}`;
553
589
  return `${detail || ''}\n${err || ''}`; // grok: no auth, quota or rejected signal is defined
554
590
  }
@@ -572,7 +608,7 @@ export function sigRejected(lane, blob) {
572
608
  const b = blob.toLowerCase();
573
609
  if (lane === 'qwen') return b.includes('[api error: 400') || b.includes('no endpoints found') || b.includes('failed to parse grammar');
574
610
  if (lane === 'codex') return blob.includes('invalid_request_error');
575
- if (lane === 'hermes') return b.includes('toolset');
611
+ if (lane === 'hermes') return hermesArgError(b) || b.includes('toolset') || hermesMismatchLine(blob) !== null;
576
612
  return false;
577
613
  }
578
614
 
@@ -752,6 +788,7 @@ const FIX = {
752
788
  auth: 'set the credential the message above names (its environment variable, or the lane\'s own login command), then rerun',
753
789
  quota: 'switch to another lane, or wait for the reset time if the message gave one',
754
790
  rejected: 'correct the model id, flag or request the upstream message names',
791
+ mismatch: 'pair the model with a provider that serves it: run `hermes model`, or pin both with --provider and --model (or "defaults": {"hermes": {"provider": ..., "model": ...}} in bin/lanes.json)',
755
792
  refused: 'adjust the hook or deny rule named above, or give this lane the tool it needs',
756
793
  cut_short: 'rerun once; if it recurs, run without --quiet and read the lane\'s stderr on the terminal',
757
794
  empty: 'rerun once, or use another lane',
@@ -769,7 +806,11 @@ export function problemAndFix(lane, cls, { out = '', err = '', detail = '', refu
769
806
  switch (cls) {
770
807
  case 'auth': return { problem: `${tag} auth: ${cause || d || 'missing or invalid credentials'}`, fix: FIX.auth };
771
808
  case 'quota': return { problem: `${tag} quota: ${cause || d || 'rate limit or credits exhausted'}`, fix: FIX.quota };
772
- case 'rejected': return { problem: `${tag} rejected: ${cause || d || 'the upstream rejected the request'}`, fix: FIX.rejected };
809
+ case 'rejected': {
810
+ const mismatch = lane === 'hermes' ? hermesMismatchLine(err) || hermesMismatchLine(out, { vendorShape: true }) : null;
811
+ if (mismatch) return { problem: `${tag} rejected: model/provider mismatch: ${redact(mismatch)}`, fix: FIX.mismatch };
812
+ return { problem: `${tag} rejected: ${cause || d || 'the upstream rejected the request'}`, fix: FIX.rejected };
813
+ }
773
814
  case 'refused': return { problem: `${tag} refused: ${denial() || 'a hook or deny rule blocked the call'}`, fix: FIX.refused };
774
815
  case 'cut_short': return { problem: `${tag} cut short: ${d || 'no terminal success event, cause not identifiable'}`, fix: FIX.cut_short };
775
816
  case 'empty': return { problem: `${tag} empty: ${d || 'completed but delivered nothing'}`, fix: FIX.empty };
@@ -801,6 +842,7 @@ export function classifyRun(lane, { rc = 0, out = '', err = '', reason = '', det
801
842
  else {
802
843
  const blob = authoritativeBlob(lane, out, err, detail, rc);
803
844
  if (sigAuth(lane, blob)) cls = 'auth';
845
+ else if (lane === 'hermes' && hermesMismatchLine(blob) !== null) cls = 'rejected'; // before quota: see judgeHermes
804
846
  else if (sigQuota(lane, blob)) cls = 'quota';
805
847
  else if (sigRejected(lane, blob)) cls = 'rejected';
806
848
  else if (Number.isInteger(refused) && refused > 0) cls = 'refused';
@@ -1080,15 +1122,19 @@ export function laneConfig(here = dirname(fileURLToPath(import.meta.url))) {
1080
1122
  for (const [lane, d] of Object.entries(j.defaults)) {
1081
1123
  if (!LANES.includes(lane)) return null;
1082
1124
  if (!d || typeof d !== 'object' || Array.isArray(d)) return null;
1083
- const { model, effort, ...rest } = d;
1125
+ const { model, effort, provider, ...rest } = d;
1084
1126
  if (Object.keys(rest).length) return null;
1085
1127
  if (model !== undefined && badRouteValue('model', model)) return null;
1128
+ if (provider !== undefined) {
1129
+ if (badRouteValue('provider', provider)) return null;
1130
+ if (!LANE_FLAGS[lane] || !LANE_FLAGS[lane].provider) return null; // only hermes routes by provider
1131
+ }
1086
1132
  if (effort !== undefined) {
1087
1133
  if (badRouteValue('effort', effort)) return null;
1088
1134
  if (effort.startsWith('auto') && effort !== 'auto') return null;
1089
1135
  if (!LANE_FLAGS[lane] || !LANE_FLAGS[lane].effort) return null; // a lane with no reasoning flag cannot have one pinned
1090
1136
  }
1091
- defaults[lane] = { model: model ?? null, effort: effort ?? null };
1137
+ defaults[lane] = { model: model ?? null, effort: effort ?? null, provider: provider ?? null };
1092
1138
  }
1093
1139
  }
1094
1140
  return { enabled: j.enabled, defaults };
@@ -1109,20 +1155,22 @@ export function resolveRoute(lane, opts, defaults) {
1109
1155
  const d = (defaults && defaults[lane]) || {};
1110
1156
  const model = opts.model ?? d.model ?? null;
1111
1157
  const effort = opts.effort ?? d.effort ?? null;
1158
+ const provider = opts.provider ?? d.provider ?? null;
1112
1159
  const src = (flag, def) => (flag != null ? 'flag' : def != null ? 'lanes.json' : 'lane_default');
1113
- return { model, effort, model_source: src(opts.model, d.model), effort_source: src(opts.effort, d.effort) };
1160
+ return { model, effort, provider, model_source: src(opts.model, d.model), effort_source: src(opts.effort, d.effort), provider_source: src(opts.provider, d.provider) };
1114
1161
  }
1115
1162
 
1116
1163
  function usage(msg) {
1117
1164
  if (msg) console.error('cli-run: ' + msg);
1118
1165
  console.error(`usage: cli-run <${LANES.join('|')}> "<prompt>" [--brief FILE] [--timeout SECS] [--quiet]
1119
- [--model ID] [--effort LEVEL] [--expect-file PATH] [--expect-json]
1166
+ [--model ID] [--effort LEVEL] [--provider ID] [--expect-file PATH] [--expect-json]
1120
1167
  cli-run codex --audit "<prompt>" read-only sandbox (audit shape)
1121
1168
  cli-run qwen [--safe-mode] "<prompt>" qwen-only flag
1122
1169
  cli-run --doctor [--run] enabled lanes, binaries, and the route each one is pinned to
1123
1170
 
1124
1171
  --model / --effort pin what a lane runs with, instead of letting it inherit its
1125
1172
  own config. Every lane takes --model; every lane except qwen takes --effort.
1173
+ Hermes also takes --provider; it needs --model, a defaults model, or HERMES_INFERENCE_MODEL.
1126
1174
  Levels are the vendor's own (agy low|medium|high, hermes none|minimal|...): an
1127
1175
  unknown level is rejected by the lane, and reported by class (codex: rejected, 16).
1128
1176
  Exit codes: 0 ok, 10 empty, 11 no output, 12 timeout, 13 unavailable, 14 auth,
@@ -1179,7 +1227,8 @@ export async function doctor(run, { here = dirname(fileURLToPath(import.meta.url
1179
1227
  const d = defaults[lane] || {};
1180
1228
  // A disabled lane has no route worth reporting; saying "not pinned" there
1181
1229
  // reads as a finding about a lane that is not going to run.
1182
- const route = !on ? '' : d.effort === 'auto' ? `route ${d.model || 'lane default'}/auto (sized per call)` : d.model || d.effort ? `route ${d.model || 'lane default'}/${d.effort || 'lane default'}` : 'route not pinned (inherits the lane\'s own config)';
1230
+ const model = `${d.provider ? d.provider + ':' : ''}${d.model || 'lane default'}`;
1231
+ const route = !on ? '' : d.effort === 'auto' ? `route ${model}/auto (sized per call)` : d.model || d.effort || d.provider ? `route ${model}/${d.effort || 'lane default'}` : 'route not pinned (inherits the lane\'s own config)';
1183
1232
  let line = ` ${lane.padEnd(7)} ${on ? 'enabled ' : 'disabled'} ${bin ? 'binary ok' : 'binary MISSING'}${route ? ' ' + route : ''}`;
1184
1233
  if (on && !bin) bad++;
1185
1234
  if (on && bin && run) {
@@ -1188,6 +1237,8 @@ export async function doctor(run, { here = dirname(fileURLToPath(import.meta.url
1188
1237
  if (rc !== OK) bad++;
1189
1238
  }
1190
1239
  console.log(compact ? ` ${lane}: ${bin ? 'present' : 'MISSING'}` : line);
1240
+ if (on && lane === 'hermes' && d.model && !d.provider) console.log(' note: model pinned with no provider: Hermes sends it to its default provider. A model that provider does not serve fails with HTTP 400; pin "provider" beside "model".');
1241
+ if (on && lane === 'hermes' && d.provider && !d.model) console.log(' note: provider pinned with no model: every run without --model will be refused unless HERMES_INFERENCE_MODEL supplies a model; pin "model" beside "provider".');
1191
1242
  }
1192
1243
  console.log(bad ? `doctor: ${bad} problem(s)` : 'doctor: all enabled lanes ' + (run ? 'answered' : 'present'));
1193
1244
  if (!compact) {
@@ -1236,10 +1287,10 @@ export function checkContracts(opts, text, before) {
1236
1287
  }
1237
1288
 
1238
1289
  export async function main(argv) {
1239
- const VALUE = new Set(['--brief', '--timeout', '--model', '--effort', '--expect-file']);
1290
+ const VALUE = new Set(['--brief', '--timeout', '--model', '--effort', '--provider', '--expect-file']);
1240
1291
  const BOOL = new Set(['--quiet', '--audit', '--safe-mode', '--doctor', '--run', '--expect-json']);
1241
1292
  const args = [...argv];
1242
- const opts = { timeout: 900, quiet: false, audit: false, model: null, effort: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
1293
+ const opts = { timeout: 900, quiet: false, audit: false, model: null, effort: null, provider: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
1243
1294
  const positional = [];
1244
1295
  while (args.length) {
1245
1296
  const a = args.shift();
@@ -1250,6 +1301,7 @@ export async function main(argv) {
1250
1301
  else if (a === '--timeout') opts.timeout = Number(v);
1251
1302
  else if (a === '--expect-file') opts.expectFile = v;
1252
1303
  else if (a === '--effort') opts.effort = v;
1304
+ else if (a === '--provider') opts.provider = v;
1253
1305
  else opts.model = v;
1254
1306
  } else if (BOOL.has(a)) {
1255
1307
  if (a === '--quiet') opts.quiet = true;
@@ -1283,7 +1335,7 @@ export async function main(argv) {
1283
1335
  if (!Number.isFinite(opts.timeout) || opts.timeout <= 0) return usage('--timeout must be a positive number of seconds');
1284
1336
  if (opts.audit && lane !== 'codex') return usage('--audit is codex-only');
1285
1337
  if (opts.safeMode && lane !== 'qwen') return usage('--safe-mode is qwen-only');
1286
- for (const [kind, v] of [['model', opts.model], ['effort', opts.effort]]) {
1338
+ for (const [kind, v] of [['model', opts.model], ['effort', opts.effort], ['provider', opts.provider]]) {
1287
1339
  if (v == null) continue;
1288
1340
  const bad = badRouteValue(kind, v);
1289
1341
  if (bad) return usage(bad);
@@ -1292,6 +1344,7 @@ export async function main(argv) {
1292
1344
  // qwen has no reasoning flag. Dropping --effort silently would leave the caller
1293
1345
  // believing a route that never happened, which is the defect this feature fixes.
1294
1346
  if (opts.effort && !(LANE_FLAGS[lane] && LANE_FLAGS[lane].effort)) return usage(`${lane} has no reasoning-effort flag; --effort is not available on this lane`);
1347
+ if (opts.provider && !LANE_FLAGS[lane].provider) return usage(`${lane} has no provider flag; --provider is hermes-only`);
1295
1348
 
1296
1349
  const digest = createHash('sha256').update(prompt).digest('hex').slice(0, 12);
1297
1350
  const base = { lane, prompt_sha256_12: digest, prompt_chars: prompt.length };
@@ -1301,11 +1354,15 @@ export async function main(argv) {
1301
1354
  // gap this feature exists to close. A malformed lanes.json has no usable
1302
1355
  // defaults, so the flags stand alone and say so.
1303
1356
  const route = resolveRoute(lane, opts, cfg === null ? {} : cfg.defaults);
1357
+ if (route.provider && !route.model && !(process.env.HERMES_INFERENCE_MODEL || '').trim()) return usage('--provider <p> needs a model: pass --model, or set "model" beside "provider" in lanes.json "defaults"');
1304
1358
  const sizing = resolveAutoEffort(lane, route.effort, prompt, opts.audit);
1305
1359
  opts.model = route.model;
1306
1360
  opts.effort = sizing.resolved;
1361
+ opts.provider = route.provider;
1307
1362
  Object.assign(base, {
1308
1363
  model_requested: route.model,
1364
+ provider_requested: route.provider,
1365
+ provider_source: route.provider_source,
1309
1366
  effort_requested: route.effort,
1310
1367
  model_source: route.model_source,
1311
1368
  effort_source: route.effort_source,
@@ -1384,7 +1441,7 @@ export async function main(argv) {
1384
1441
  }
1385
1442
  const code = cls === 'interrupted' ? 128 + (r.interrupted === 'SIGINT' ? 2 : 15) : CLASS_CODES[cls];
1386
1443
  if (text && code === OK) process.stdout.write(text + '\n');
1387
- const routeNote = route.model || route.effort ? `${route.model || 'lane default'}/${route.effort || 'lane default'}` : 'lane default';
1444
+ const routeNote = (route.provider ? `${route.provider}:` : '') + (route.model || route.effort ? `${route.model || 'lane default'}/${route.effort || 'lane default'}` : 'lane default');
1388
1445
  if (!opts.quiet) {
1389
1446
  console.error(`cli-run[${lane}] ${verdict} rc=${code} class=${cls} refused=${refused === null ? 'null' : refused} ${r.seconds.toFixed(1)}s raw=${r.outBytes || 0}B route=${routeNote} :: ${redact(detail)}`);
1390
1447
  let authoritative = null;
@@ -60,6 +60,8 @@ aunx cli-run --doctor
60
60
  # Direct form: node bin/cli-run.mjs --doctor
61
61
  ```
62
62
 
63
+ Hermes also takes `--provider '<provider-id>'` (or `"provider"` in its `lanes.json` defaults), always together with a model, because one Hermes install can reach several providers and a model sent to the wrong one fails with HTTP 400. `--doctor` notes a Hermes model pinned without a provider.
64
+
63
65
  Explicit flags override defaults in `bin/lanes.json`. Without either, the vendor CLI uses its own configuration. The runner records what was requested and the source of each request: `flag`, `lanes.json` or `lane_default`. These fields describe requested settings; the vendor's own reporting is the place to verify the actual model used.
64
66
 
65
67
  ## Share facts once, then scope each worker
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "1.0.5",
3
+ "version": "1.0.6",
4
4
  "description": "Model router for AI coding agents: installs routing rules, 8 subagents, hooks and a CLI runner so your AI picks model and effort per task and saves tokens",
5
5
  "type": "module",
6
6
  "bin": {
package/src/install.js CHANGED
@@ -823,7 +823,7 @@ export function planFiles(opts) {
823
823
  enabled: selected.filter((a) => a.facts.cliRun).map((a) => a.id),
824
824
  defaults: Object.fromEntries((opts.effortAuto || []).map((lane) => [lane, { effort: 'auto' }])),
825
825
  note: 'Lanes cli-run may call. Edit to enable or disable a lane. A lane not listed here exits 13 (unavailable).',
826
- defaultsNote: 'Pin what a lane runs with, so the route in your docs is the route that runs: "defaults": {"' + (selected.find(a => a.facts.cliRun)?.id || '<lane>') + '": ' + JSON.stringify({ model: '<model-id>', ...(LANE_FLAGS[selected.find(a => a.facts.cliRun)?.id]?.effort ? { effort: 'high' } : {}) }) + '}. Left empty, a lane inherits its own config file, which cli-run cannot see and does not guess. `--model` and `--effort` override this per call, and `--doctor` prints what each lane is pinned to. Every enabled lane takes a model; the runner reports which lanes support an effort flag.'
826
+ defaultsNote: 'Pin what a lane runs with, so the route in your docs is the route that runs: "defaults": {"' + (selected.find(a => a.facts.cliRun)?.id || '<lane>') + '": ' + JSON.stringify({ model: '<model-id>', ...(LANE_FLAGS[selected.find(a => a.facts.cliRun)?.id]?.effort ? { effort: 'high' } : {}) }) + '}. Left empty, a lane inherits its own config file, which cli-run cannot see and does not guess. `--model` and `--effort` override this per call, and `--doctor` prints what each lane is pinned to. Every enabled lane takes a model; the runner reports which lanes support an effort flag. Hermes also takes "provider" beside "model", so the model goes to a provider that serves it.'
827
827
  },
828
828
  null,
829
829
  2
@@ -97,13 +97,15 @@ Replace the placeholder with a current vendor model ID before using this example
97
97
 
98
98
  The runner supports these lanes whether or not you selected them.
99
99
 
100
- | Lane | Model flag | Effort flag |
101
- |---|---|---|
102
- | grok | `-m` | `--reasoning-effort` |
103
- | codex | `-m` | `-c model_reasoning_effort="LEVEL"` |
104
- | agy | `--model` | `--effort` |
105
- | hermes | `-m` | `--reasoning` |
106
- | qwen | `-m` | Unsupported; an effort request is a usage error |
100
+ | Lane | Model flag | Effort flag | Provider flag |
101
+ |---|---|---|---|
102
+ | grok | `-m` | `--reasoning-effort` | Unsupported |
103
+ | codex | `-m` | `-c model_reasoning_effort="LEVEL"` | Unsupported |
104
+ | agy | `--model` | `--effort` | Unsupported |
105
+ | hermes | `-m` | `--reasoning` | `--provider` |
106
+ | qwen | `-m` | Unsupported; an effort request is a usage error | Unsupported |
107
+
108
+ Hermes signs in to many providers, and a model sent to a provider that does not serve it fails with HTTP 400. Pin both with `--provider` and `--model`, or with `"provider"` beside `"model"` under `"hermes"` in `bin/lanes.json` defaults. A provider with no model is a usage error. That failure is reported as `rejected` (exit 16) with a model/provider mismatch line; `hermes model` repairs the pairing in Hermes itself.
107
109
 
108
110
  When using `--effort auto`, treat its medium/high selection as a bounded heuristic; an audit has a high floor. Use explicit high for builds and xhigh where supported for security-critical or irreversible work. The vendor validates its own effort names and reports unsupported values through the failure class.
109
111