model-orchestrator 1.0.4 → 1.0.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +1 -1
- package/CHANGELOG.md +21 -1
- package/bin/cli-run.mjs +80 -23
- package/docs/how-it-routes.md +2 -0
- package/package.json +1 -1
- package/proof/README.md +4 -22
- package/proof/results.json +0 -24
- package/proof/scripts/lib.js +3 -1
- package/proof/scripts/render.js +1 -1
- package/src/install.js +1 -1
- package/templates/intermediate/CLI-RUN.md +9 -7
package/AGENTS.md
CHANGED
|
@@ -33,6 +33,6 @@ When using the Claude Code plugin, follow [plugin/README.md](plugin/README.md).
|
|
|
33
33
|
- **Installer safety:** writes remain inside `--dir` and `--project`; preserve user edits according to manifest hashes and explicit flags. Run no third-party installer.
|
|
34
34
|
- **Runner safety:** preserve exit codes and the log schema. Before changing an output judge, add a failing case in `test/judges.test.js`.
|
|
35
35
|
- **Secrets:** use environment-variable names only. Never add a credential value to code, examples or tests.
|
|
36
|
-
- **Proof:** measure through `proof/scripts/`, store results in `proof/results.json` and regenerate the proof page.
|
|
36
|
+
- **Proof:** measure through `proof/scripts/`, store results in `proof/results.json` and regenerate the proof page. An expired entry warns and is re-measured; it never blocks tests or releases.
|
|
37
37
|
- **Verify:** run `npm test` and report tests, pass, fail, skipped and exit code. The suite prints current counts. Use `npm pack --dry-run` to inspect publication contents.
|
|
38
38
|
- **Style:** use short condition-to-action instructions and no em dashes. `test/prose.test.js` checks public vocabulary and examples.
|
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,25 @@ All notable changes to this project are documented here. The format follows [Kee
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [1.0.6] - 2026-09-29
|
|
8
|
+
|
|
9
|
+
### Added
|
|
10
|
+
|
|
11
|
+
- The hermes lane takes `--provider` and `"provider"` in `bin/lanes.json` defaults, so a pinned model goes to a provider that serves it. A provider needs a model with it (a flag, a default or `HERMES_INFERENCE_MODEL`); otherwise cli-run refuses before the lane starts. Runs log `provider_requested` and `provider_source`.
|
|
12
|
+
- `--doctor` notes a hermes model pinned without a provider, and a provider pinned without a model. The terminal route line shows the provider.
|
|
13
|
+
|
|
14
|
+
### Fixed
|
|
15
|
+
|
|
16
|
+
- A hermes model/provider mismatch (an upstream "model is not supported" or `model_not_found`, for example `grok-4.6` sent to `openai-codex`) is now class `rejected`, exit 16, with a problem line naming the mismatch and a fix pointing at `hermes model`. It was reported as a vague nonzero exit.
|
|
17
|
+
- A pinned hermes route (`--model` or `--effort`, by flag or `lanes.json` default) no longer fails every run. cli-run placed those flags between `-z` and the prompt, and `-z` takes the prompt as its value, so Hermes refused the arguments ("argument -z/--oneshot: expected one argument"). Route flags now come before `-z`, and a Hermes argument error is reported as such rather than as a toolsets error.
|
|
18
|
+
|
|
19
|
+
## [1.0.5] - 2026-09-28
|
|
20
|
+
|
|
21
|
+
### Changed
|
|
22
|
+
|
|
23
|
+
- The proof page keeps only figures that measure this package. The two browser token figures measured the author's own browser subagent, which model-orchestrator does not ship, and are removed from the proof data, the site and the trailer.
|
|
24
|
+
- An expired proof figure is a signal to re-measure: it warns and never fails tests, CI or a release. The product site leaves expired figures off.
|
|
25
|
+
|
|
7
26
|
## [1.0.4] - 2026-09-28
|
|
8
27
|
|
|
9
28
|
### Security
|
|
@@ -550,7 +569,8 @@ First release.
|
|
|
550
569
|
- Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
|
|
551
570
|
- Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
|
|
552
571
|
|
|
553
|
-
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.
|
|
572
|
+
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.5...HEAD
|
|
573
|
+
[1.0.5]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.4...v1.0.5
|
|
554
574
|
[1.0.4]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.3...v1.0.4
|
|
555
575
|
[1.0.3]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.2...v1.0.3
|
|
556
576
|
[1.0.2]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.1...v1.0.2
|
package/bin/cli-run.mjs
CHANGED
|
@@ -53,6 +53,7 @@
|
|
|
53
53
|
// problem/fix lines go to your terminal only, redacted.
|
|
54
54
|
//
|
|
55
55
|
// ROUTE: which model and reasoning effort a lane ran with.
|
|
56
|
+
// Hermes also accepts --provider to select the provider serving that model.
|
|
56
57
|
// A lane with no --model and no lanes.json default inherits whatever its own
|
|
57
58
|
// config file says, which is invisible from here and is how a documented route
|
|
58
59
|
// silently stops being the route that runs. --model / --effort pin it per call,
|
|
@@ -199,7 +200,11 @@ export function judgeHermes(rc, out, err) {
|
|
|
199
200
|
// hermes collapses every upstream failure into one exit code; its stderr
|
|
200
201
|
// is the only place the cause is named.
|
|
201
202
|
const blob = String(err || '').toLowerCase(); // stderr only: stdout is the agent's own prose
|
|
202
|
-
|
|
203
|
+
// A mismatch is checked first: a retry cannot fix it, and an unrelated
|
|
204
|
+
// "limit" warning on stderr must not turn it into a quota report.
|
|
205
|
+
if (hermesMismatchLine(err) || hermesMismatchLine(out, { vendorShape: true })) why += ': model/provider mismatch, the provider does not serve this model';
|
|
206
|
+
else if (hermesQuotaText(blob)) why += ': upstream free tier degraded or limited, a retry is reasonable';
|
|
207
|
+
else if (hermesArgError(blob)) why += ': hermes rejected its arguments, a caller bug and not a lane fault';
|
|
203
208
|
else if (blob.includes('toolset')) why += ': invalid --toolsets value, a caller bug and not a lane fault';
|
|
204
209
|
return fail('exit_nonzero', `hermes exit ${rc}: ${why}`);
|
|
205
210
|
}
|
|
@@ -253,14 +258,18 @@ export function judgeQwen(rc, out) {
|
|
|
253
258
|
// grok -m MODEL --reasoning-effort EFFORT
|
|
254
259
|
// codex -m MODEL -c model_reasoning_effort="EFFORT" (a TOML override, hence the quotes)
|
|
255
260
|
// agy --model M --effort EFFORT (low|medium|high)
|
|
256
|
-
// hermes -m MODEL --reasoning LEVEL
|
|
261
|
+
// hermes -m MODEL --reasoning LEVEL --provider ID (none|minimal|...)
|
|
257
262
|
// qwen -m MODEL no reasoning flag
|
|
263
|
+
// Only hermes takes a provider: one hermes install signs in to many providers,
|
|
264
|
+
// and a model id sent to a provider that does not serve it is an HTTP 400
|
|
265
|
+
// (grok-4.6 on openai-codex). `hermes -z --provider` without a model exits 2,
|
|
266
|
+
// so cli-run refuses that pairing before the lane starts.
|
|
258
267
|
export const LANE_FLAGS = {
|
|
259
|
-
grok: { model: (v) => ['-m', v], effort: (v) => ['--reasoning-effort', v] },
|
|
260
|
-
codex: { model: (v) => ['-m', v], effort: (v) => ['-c', `model_reasoning_effort="${v}"`] },
|
|
261
|
-
agy: { model: (v) => ['--model', v], effort: (v) => ['--effort', v] },
|
|
262
|
-
hermes: { model: (v) => ['-m', v], effort: (v) => ['--reasoning', v] },
|
|
263
|
-
qwen: { model: (v) => ['-m', v], effort: null }
|
|
268
|
+
grok: { model: (v) => ['-m', v], effort: (v) => ['--reasoning-effort', v], provider: null },
|
|
269
|
+
codex: { model: (v) => ['-m', v], effort: (v) => ['-c', `model_reasoning_effort="${v}"`], provider: null },
|
|
270
|
+
agy: { model: (v) => ['--model', v], effort: (v) => ['--effort', v], provider: null },
|
|
271
|
+
hermes: { model: (v) => ['-m', v], effort: (v) => ['--reasoning', v], provider: (v) => ['--provider', v] },
|
|
272
|
+
qwen: { model: (v) => ['-m', v], effort: null, provider: null }
|
|
264
273
|
};
|
|
265
274
|
|
|
266
275
|
// Auto is deliberately a small, static ladder. It is not a vendor capability
|
|
@@ -353,13 +362,16 @@ export function badRouteValue(kind, v) {
|
|
|
353
362
|
}
|
|
354
363
|
|
|
355
364
|
// --- adapters: build argv for a lane -------------------------------------
|
|
356
|
-
// Route flags go in front of the prompt for every lane, because
|
|
357
|
-
//
|
|
358
|
-
//
|
|
365
|
+
// Route flags go in front of the prompt for every lane, because codex takes the
|
|
366
|
+
// prompt as a positional argument and a flag after it is either ignored or read
|
|
367
|
+
// as part of it. hermes needs them in front of -z itself: -z takes the prompt as
|
|
368
|
+
// its own value, so `-z -m X prompt` is an argparse error ("argument -z/--oneshot:
|
|
369
|
+
// expected one argument") and every pinned hermes route failed that way.
|
|
359
370
|
function routeFlags(lane, opts) {
|
|
360
371
|
const spec = LANE_FLAGS[lane];
|
|
361
372
|
const out = [];
|
|
362
373
|
if (!spec) return out;
|
|
374
|
+
if (opts.provider && spec.provider) out.push(...spec.provider(opts.provider));
|
|
363
375
|
if (opts.model) out.push(...spec.model(opts.model));
|
|
364
376
|
if (opts.effort && spec.effort) out.push(...spec.effort(opts.effort));
|
|
365
377
|
return out;
|
|
@@ -384,7 +396,7 @@ export function buildArgv(lane, binary, prompt, opts, tmp) {
|
|
|
384
396
|
return { argv: [binary, '--print-timeout', `${mins}m`, '--output-format', 'stream-json', ...route, '-p', prompt] };
|
|
385
397
|
}
|
|
386
398
|
case 'hermes':
|
|
387
|
-
return { argv: [binary, '-z',
|
|
399
|
+
return { argv: [binary, ...route, '-z', prompt, '--usage-file', join(tmp, 'usage.json')] };
|
|
388
400
|
case 'qwen': {
|
|
389
401
|
const argv = [binary, '-o', 'json', ...route];
|
|
390
402
|
if (opts.safeMode) argv.push('--safe-mode');
|
|
@@ -527,6 +539,30 @@ function hermesQuotaText(lower) {
|
|
|
527
539
|
return lower.includes('no usable content') || lower.includes('limit') || lower.includes('degraded');
|
|
528
540
|
}
|
|
529
541
|
|
|
542
|
+
// argparse's own refusal. Its usage dump always lists `-t TOOLSETS`, so this is
|
|
543
|
+
// checked before the toolset match or every bad argument reads as a toolset error.
|
|
544
|
+
function hermesArgError(lower) {
|
|
545
|
+
return lower.includes('hermes: error:');
|
|
546
|
+
}
|
|
547
|
+
|
|
548
|
+
// A model id the chosen provider does not serve. `hermes -z` prints a failed
|
|
549
|
+
// turn's provider error on stdout and exits 2, so this reads both streams, but
|
|
550
|
+
// returns only the matching line: the rest of stdout is the agent's own prose.
|
|
551
|
+
// The wording is the upstream's own ("The 'grok-4.6' model is not supported
|
|
552
|
+
// when using Codex with a ChatGPT account"; OpenAI's model_not_found code).
|
|
553
|
+
export function hermesMismatchLine(text, { vendorShape = false } = {}) {
|
|
554
|
+
for (const line of String(text || '').split('\n')) {
|
|
555
|
+
const l = line.toLowerCase();
|
|
556
|
+
if (!(l.includes('model is not supported') || l.includes('model_not_found'))) continue;
|
|
557
|
+
// stdout is agent prose, so there only the vendor's own shape counts: an HTTP
|
|
558
|
+
// status, an error prefix, the Codex wording, or the OpenAI error code.
|
|
559
|
+
if (vendorShape && !(/^\s*(http [45]\d\d|error)\b/.test(l) || l.includes('model is not supported when using') || l.includes('model_not_found'))) continue;
|
|
560
|
+
// Terminal-safe: stdout can carry web-derived text, so no control characters.
|
|
561
|
+
return line.replace(/[\x00-\x08\x0b-\x1f\x7f]/g, '').trim().slice(0, 300);
|
|
562
|
+
}
|
|
563
|
+
return null;
|
|
564
|
+
}
|
|
565
|
+
|
|
530
566
|
// qwen: the terminal event's full error text, and a result that is an API error.
|
|
531
567
|
// Never the display detail, which is clipped and can carry a model name.
|
|
532
568
|
export function qwenErrorText(out) {
|
|
@@ -548,7 +584,7 @@ export function qwenErrorText(out) {
|
|
|
548
584
|
function authoritativeBlob(lane, out, err, detail, rc) {
|
|
549
585
|
if (lane === 'codex') return `${codexErrorEventsText(out)}\n${err || ''}`;
|
|
550
586
|
if (lane === 'agy') return `${agyResultFieldsText(out)}\n${err || ''}`;
|
|
551
|
-
if (lane === 'hermes') return rc !== 0 ?
|
|
587
|
+
if (lane === 'hermes') return rc !== 0 ? `${err || ''}\n${hermesMismatchLine(out, { vendorShape: true }) || ''}` : '';
|
|
552
588
|
if (lane === 'qwen') return `${qwenErrorText(out)}\n${err || ''}`;
|
|
553
589
|
return `${detail || ''}\n${err || ''}`; // grok: no auth, quota or rejected signal is defined
|
|
554
590
|
}
|
|
@@ -572,7 +608,7 @@ export function sigRejected(lane, blob) {
|
|
|
572
608
|
const b = blob.toLowerCase();
|
|
573
609
|
if (lane === 'qwen') return b.includes('[api error: 400') || b.includes('no endpoints found') || b.includes('failed to parse grammar');
|
|
574
610
|
if (lane === 'codex') return blob.includes('invalid_request_error');
|
|
575
|
-
if (lane === 'hermes') return b.includes('toolset');
|
|
611
|
+
if (lane === 'hermes') return hermesArgError(b) || b.includes('toolset') || hermesMismatchLine(blob) !== null;
|
|
576
612
|
return false;
|
|
577
613
|
}
|
|
578
614
|
|
|
@@ -752,6 +788,7 @@ const FIX = {
|
|
|
752
788
|
auth: 'set the credential the message above names (its environment variable, or the lane\'s own login command), then rerun',
|
|
753
789
|
quota: 'switch to another lane, or wait for the reset time if the message gave one',
|
|
754
790
|
rejected: 'correct the model id, flag or request the upstream message names',
|
|
791
|
+
mismatch: 'pair the model with a provider that serves it: run `hermes model`, or pin both with --provider and --model (or "defaults": {"hermes": {"provider": ..., "model": ...}} in bin/lanes.json)',
|
|
755
792
|
refused: 'adjust the hook or deny rule named above, or give this lane the tool it needs',
|
|
756
793
|
cut_short: 'rerun once; if it recurs, run without --quiet and read the lane\'s stderr on the terminal',
|
|
757
794
|
empty: 'rerun once, or use another lane',
|
|
@@ -769,7 +806,11 @@ export function problemAndFix(lane, cls, { out = '', err = '', detail = '', refu
|
|
|
769
806
|
switch (cls) {
|
|
770
807
|
case 'auth': return { problem: `${tag} auth: ${cause || d || 'missing or invalid credentials'}`, fix: FIX.auth };
|
|
771
808
|
case 'quota': return { problem: `${tag} quota: ${cause || d || 'rate limit or credits exhausted'}`, fix: FIX.quota };
|
|
772
|
-
case 'rejected':
|
|
809
|
+
case 'rejected': {
|
|
810
|
+
const mismatch = lane === 'hermes' ? hermesMismatchLine(err) || hermesMismatchLine(out, { vendorShape: true }) : null;
|
|
811
|
+
if (mismatch) return { problem: `${tag} rejected: model/provider mismatch: ${redact(mismatch)}`, fix: FIX.mismatch };
|
|
812
|
+
return { problem: `${tag} rejected: ${cause || d || 'the upstream rejected the request'}`, fix: FIX.rejected };
|
|
813
|
+
}
|
|
773
814
|
case 'refused': return { problem: `${tag} refused: ${denial() || 'a hook or deny rule blocked the call'}`, fix: FIX.refused };
|
|
774
815
|
case 'cut_short': return { problem: `${tag} cut short: ${d || 'no terminal success event, cause not identifiable'}`, fix: FIX.cut_short };
|
|
775
816
|
case 'empty': return { problem: `${tag} empty: ${d || 'completed but delivered nothing'}`, fix: FIX.empty };
|
|
@@ -801,6 +842,7 @@ export function classifyRun(lane, { rc = 0, out = '', err = '', reason = '', det
|
|
|
801
842
|
else {
|
|
802
843
|
const blob = authoritativeBlob(lane, out, err, detail, rc);
|
|
803
844
|
if (sigAuth(lane, blob)) cls = 'auth';
|
|
845
|
+
else if (lane === 'hermes' && hermesMismatchLine(blob) !== null) cls = 'rejected'; // before quota: see judgeHermes
|
|
804
846
|
else if (sigQuota(lane, blob)) cls = 'quota';
|
|
805
847
|
else if (sigRejected(lane, blob)) cls = 'rejected';
|
|
806
848
|
else if (Number.isInteger(refused) && refused > 0) cls = 'refused';
|
|
@@ -1080,15 +1122,19 @@ export function laneConfig(here = dirname(fileURLToPath(import.meta.url))) {
|
|
|
1080
1122
|
for (const [lane, d] of Object.entries(j.defaults)) {
|
|
1081
1123
|
if (!LANES.includes(lane)) return null;
|
|
1082
1124
|
if (!d || typeof d !== 'object' || Array.isArray(d)) return null;
|
|
1083
|
-
const { model, effort, ...rest } = d;
|
|
1125
|
+
const { model, effort, provider, ...rest } = d;
|
|
1084
1126
|
if (Object.keys(rest).length) return null;
|
|
1085
1127
|
if (model !== undefined && badRouteValue('model', model)) return null;
|
|
1128
|
+
if (provider !== undefined) {
|
|
1129
|
+
if (badRouteValue('provider', provider)) return null;
|
|
1130
|
+
if (!LANE_FLAGS[lane] || !LANE_FLAGS[lane].provider) return null; // only hermes routes by provider
|
|
1131
|
+
}
|
|
1086
1132
|
if (effort !== undefined) {
|
|
1087
1133
|
if (badRouteValue('effort', effort)) return null;
|
|
1088
1134
|
if (effort.startsWith('auto') && effort !== 'auto') return null;
|
|
1089
1135
|
if (!LANE_FLAGS[lane] || !LANE_FLAGS[lane].effort) return null; // a lane with no reasoning flag cannot have one pinned
|
|
1090
1136
|
}
|
|
1091
|
-
defaults[lane] = { model: model ?? null, effort: effort ?? null };
|
|
1137
|
+
defaults[lane] = { model: model ?? null, effort: effort ?? null, provider: provider ?? null };
|
|
1092
1138
|
}
|
|
1093
1139
|
}
|
|
1094
1140
|
return { enabled: j.enabled, defaults };
|
|
@@ -1109,20 +1155,22 @@ export function resolveRoute(lane, opts, defaults) {
|
|
|
1109
1155
|
const d = (defaults && defaults[lane]) || {};
|
|
1110
1156
|
const model = opts.model ?? d.model ?? null;
|
|
1111
1157
|
const effort = opts.effort ?? d.effort ?? null;
|
|
1158
|
+
const provider = opts.provider ?? d.provider ?? null;
|
|
1112
1159
|
const src = (flag, def) => (flag != null ? 'flag' : def != null ? 'lanes.json' : 'lane_default');
|
|
1113
|
-
return { model, effort, model_source: src(opts.model, d.model), effort_source: src(opts.effort, d.effort) };
|
|
1160
|
+
return { model, effort, provider, model_source: src(opts.model, d.model), effort_source: src(opts.effort, d.effort), provider_source: src(opts.provider, d.provider) };
|
|
1114
1161
|
}
|
|
1115
1162
|
|
|
1116
1163
|
function usage(msg) {
|
|
1117
1164
|
if (msg) console.error('cli-run: ' + msg);
|
|
1118
1165
|
console.error(`usage: cli-run <${LANES.join('|')}> "<prompt>" [--brief FILE] [--timeout SECS] [--quiet]
|
|
1119
|
-
[--model ID] [--effort LEVEL] [--expect-file PATH] [--expect-json]
|
|
1166
|
+
[--model ID] [--effort LEVEL] [--provider ID] [--expect-file PATH] [--expect-json]
|
|
1120
1167
|
cli-run codex --audit "<prompt>" read-only sandbox (audit shape)
|
|
1121
1168
|
cli-run qwen [--safe-mode] "<prompt>" qwen-only flag
|
|
1122
1169
|
cli-run --doctor [--run] enabled lanes, binaries, and the route each one is pinned to
|
|
1123
1170
|
|
|
1124
1171
|
--model / --effort pin what a lane runs with, instead of letting it inherit its
|
|
1125
1172
|
own config. Every lane takes --model; every lane except qwen takes --effort.
|
|
1173
|
+
Hermes also takes --provider; it needs --model, a defaults model, or HERMES_INFERENCE_MODEL.
|
|
1126
1174
|
Levels are the vendor's own (agy low|medium|high, hermes none|minimal|...): an
|
|
1127
1175
|
unknown level is rejected by the lane, and reported by class (codex: rejected, 16).
|
|
1128
1176
|
Exit codes: 0 ok, 10 empty, 11 no output, 12 timeout, 13 unavailable, 14 auth,
|
|
@@ -1179,7 +1227,8 @@ export async function doctor(run, { here = dirname(fileURLToPath(import.meta.url
|
|
|
1179
1227
|
const d = defaults[lane] || {};
|
|
1180
1228
|
// A disabled lane has no route worth reporting; saying "not pinned" there
|
|
1181
1229
|
// reads as a finding about a lane that is not going to run.
|
|
1182
|
-
const
|
|
1230
|
+
const model = `${d.provider ? d.provider + ':' : ''}${d.model || 'lane default'}`;
|
|
1231
|
+
const route = !on ? '' : d.effort === 'auto' ? `route ${model}/auto (sized per call)` : d.model || d.effort || d.provider ? `route ${model}/${d.effort || 'lane default'}` : 'route not pinned (inherits the lane\'s own config)';
|
|
1183
1232
|
let line = ` ${lane.padEnd(7)} ${on ? 'enabled ' : 'disabled'} ${bin ? 'binary ok' : 'binary MISSING'}${route ? ' ' + route : ''}`;
|
|
1184
1233
|
if (on && !bin) bad++;
|
|
1185
1234
|
if (on && bin && run) {
|
|
@@ -1188,6 +1237,8 @@ export async function doctor(run, { here = dirname(fileURLToPath(import.meta.url
|
|
|
1188
1237
|
if (rc !== OK) bad++;
|
|
1189
1238
|
}
|
|
1190
1239
|
console.log(compact ? ` ${lane}: ${bin ? 'present' : 'MISSING'}` : line);
|
|
1240
|
+
if (on && lane === 'hermes' && d.model && !d.provider) console.log(' note: model pinned with no provider: Hermes sends it to its default provider. A model that provider does not serve fails with HTTP 400; pin "provider" beside "model".');
|
|
1241
|
+
if (on && lane === 'hermes' && d.provider && !d.model) console.log(' note: provider pinned with no model: every run without --model will be refused unless HERMES_INFERENCE_MODEL supplies a model; pin "model" beside "provider".');
|
|
1191
1242
|
}
|
|
1192
1243
|
console.log(bad ? `doctor: ${bad} problem(s)` : 'doctor: all enabled lanes ' + (run ? 'answered' : 'present'));
|
|
1193
1244
|
if (!compact) {
|
|
@@ -1236,10 +1287,10 @@ export function checkContracts(opts, text, before) {
|
|
|
1236
1287
|
}
|
|
1237
1288
|
|
|
1238
1289
|
export async function main(argv) {
|
|
1239
|
-
const VALUE = new Set(['--brief', '--timeout', '--model', '--effort', '--expect-file']);
|
|
1290
|
+
const VALUE = new Set(['--brief', '--timeout', '--model', '--effort', '--provider', '--expect-file']);
|
|
1240
1291
|
const BOOL = new Set(['--quiet', '--audit', '--safe-mode', '--doctor', '--run', '--expect-json']);
|
|
1241
1292
|
const args = [...argv];
|
|
1242
|
-
const opts = { timeout: 900, quiet: false, audit: false, model: null, effort: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
|
|
1293
|
+
const opts = { timeout: 900, quiet: false, audit: false, model: null, effort: null, provider: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
|
|
1243
1294
|
const positional = [];
|
|
1244
1295
|
while (args.length) {
|
|
1245
1296
|
const a = args.shift();
|
|
@@ -1250,6 +1301,7 @@ export async function main(argv) {
|
|
|
1250
1301
|
else if (a === '--timeout') opts.timeout = Number(v);
|
|
1251
1302
|
else if (a === '--expect-file') opts.expectFile = v;
|
|
1252
1303
|
else if (a === '--effort') opts.effort = v;
|
|
1304
|
+
else if (a === '--provider') opts.provider = v;
|
|
1253
1305
|
else opts.model = v;
|
|
1254
1306
|
} else if (BOOL.has(a)) {
|
|
1255
1307
|
if (a === '--quiet') opts.quiet = true;
|
|
@@ -1283,7 +1335,7 @@ export async function main(argv) {
|
|
|
1283
1335
|
if (!Number.isFinite(opts.timeout) || opts.timeout <= 0) return usage('--timeout must be a positive number of seconds');
|
|
1284
1336
|
if (opts.audit && lane !== 'codex') return usage('--audit is codex-only');
|
|
1285
1337
|
if (opts.safeMode && lane !== 'qwen') return usage('--safe-mode is qwen-only');
|
|
1286
|
-
for (const [kind, v] of [['model', opts.model], ['effort', opts.effort]]) {
|
|
1338
|
+
for (const [kind, v] of [['model', opts.model], ['effort', opts.effort], ['provider', opts.provider]]) {
|
|
1287
1339
|
if (v == null) continue;
|
|
1288
1340
|
const bad = badRouteValue(kind, v);
|
|
1289
1341
|
if (bad) return usage(bad);
|
|
@@ -1292,6 +1344,7 @@ export async function main(argv) {
|
|
|
1292
1344
|
// qwen has no reasoning flag. Dropping --effort silently would leave the caller
|
|
1293
1345
|
// believing a route that never happened, which is the defect this feature fixes.
|
|
1294
1346
|
if (opts.effort && !(LANE_FLAGS[lane] && LANE_FLAGS[lane].effort)) return usage(`${lane} has no reasoning-effort flag; --effort is not available on this lane`);
|
|
1347
|
+
if (opts.provider && !LANE_FLAGS[lane].provider) return usage(`${lane} has no provider flag; --provider is hermes-only`);
|
|
1295
1348
|
|
|
1296
1349
|
const digest = createHash('sha256').update(prompt).digest('hex').slice(0, 12);
|
|
1297
1350
|
const base = { lane, prompt_sha256_12: digest, prompt_chars: prompt.length };
|
|
@@ -1301,11 +1354,15 @@ export async function main(argv) {
|
|
|
1301
1354
|
// gap this feature exists to close. A malformed lanes.json has no usable
|
|
1302
1355
|
// defaults, so the flags stand alone and say so.
|
|
1303
1356
|
const route = resolveRoute(lane, opts, cfg === null ? {} : cfg.defaults);
|
|
1357
|
+
if (route.provider && !route.model && !(process.env.HERMES_INFERENCE_MODEL || '').trim()) return usage('--provider <p> needs a model: pass --model, or set "model" beside "provider" in lanes.json "defaults"');
|
|
1304
1358
|
const sizing = resolveAutoEffort(lane, route.effort, prompt, opts.audit);
|
|
1305
1359
|
opts.model = route.model;
|
|
1306
1360
|
opts.effort = sizing.resolved;
|
|
1361
|
+
opts.provider = route.provider;
|
|
1307
1362
|
Object.assign(base, {
|
|
1308
1363
|
model_requested: route.model,
|
|
1364
|
+
provider_requested: route.provider,
|
|
1365
|
+
provider_source: route.provider_source,
|
|
1309
1366
|
effort_requested: route.effort,
|
|
1310
1367
|
model_source: route.model_source,
|
|
1311
1368
|
effort_source: route.effort_source,
|
|
@@ -1384,7 +1441,7 @@ export async function main(argv) {
|
|
|
1384
1441
|
}
|
|
1385
1442
|
const code = cls === 'interrupted' ? 128 + (r.interrupted === 'SIGINT' ? 2 : 15) : CLASS_CODES[cls];
|
|
1386
1443
|
if (text && code === OK) process.stdout.write(text + '\n');
|
|
1387
|
-
const routeNote = route.model || route.effort ? `${route.model || 'lane default'}/${route.effort || 'lane default'}` : 'lane default';
|
|
1444
|
+
const routeNote = (route.provider ? `${route.provider}:` : '') + (route.model || route.effort ? `${route.model || 'lane default'}/${route.effort || 'lane default'}` : 'lane default');
|
|
1388
1445
|
if (!opts.quiet) {
|
|
1389
1446
|
console.error(`cli-run[${lane}] ${verdict} rc=${code} class=${cls} refused=${refused === null ? 'null' : refused} ${r.seconds.toFixed(1)}s raw=${r.outBytes || 0}B route=${routeNote} :: ${redact(detail)}`);
|
|
1390
1447
|
let authoritative = null;
|
package/docs/how-it-routes.md
CHANGED
|
@@ -60,6 +60,8 @@ aunx cli-run --doctor
|
|
|
60
60
|
# Direct form: node bin/cli-run.mjs --doctor
|
|
61
61
|
```
|
|
62
62
|
|
|
63
|
+
Hermes also takes `--provider '<provider-id>'` (or `"provider"` in its `lanes.json` defaults), always together with a model, because one Hermes install can reach several providers and a model sent to the wrong one fails with HTTP 400. `--doctor` notes a Hermes model pinned without a provider.
|
|
64
|
+
|
|
63
65
|
Explicit flags override defaults in `bin/lanes.json`. Without either, the vendor CLI uses its own configuration. The runner records what was requested and the source of each request: `flag`, `lanes.json` or `lane_default`. These fields describe requested settings; the vendor's own reporting is the place to verify the actual model used.
|
|
64
66
|
|
|
65
67
|
## Share facts once, then scope each worker
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "model-orchestrator",
|
|
3
|
-
"version": "1.0.
|
|
3
|
+
"version": "1.0.6",
|
|
4
4
|
"description": "Model router for AI coding agents: installs routing rules, 8 subagents, hooks and a CLI runner so your AI picks model and effort per task and saves tokens",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
package/proof/README.md
CHANGED
|
@@ -10,8 +10,6 @@ Environment: Node v22.22.3, darwin arm64. Timing varies with startup caches and
|
|
|
10
10
|
| Lane runner overhead | 63.74 ms median difference | 7 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/runner-overhead.js) |
|
|
11
11
|
| Empty results flagged | 10 fixtures rejected | 10 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/missing-results.js) |
|
|
12
12
|
| Acceptance failures blocked | 4 fixtures rejected | 4 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/check-gate.js) |
|
|
13
|
-
| Main conversation browser tokens | 408147 tokens per browser step (median) | 9775 | 2026-09-26 | 2026-10-27 | author setup, re-measured locally |
|
|
14
|
-
| Small browser subagent tokens | 17197 tokens per step (highest of 5 runs) | 5 | 2026-09-27 | 2026-10-27 | author setup, re-measured locally |
|
|
15
13
|
|
|
16
14
|
## Run the proof scripts
|
|
17
15
|
|
|
@@ -21,7 +19,7 @@ node proof/scripts/render.js
|
|
|
21
19
|
npm test
|
|
22
20
|
```
|
|
23
21
|
|
|
24
|
-
The measurement command refreshes reproducible entries
|
|
22
|
+
The measurement command refreshes the reproducible entries. An expired entry is a signal to re-measure; the test suite rejects future-dated or incomplete entries and checks this page against the data. The weekly [refresh workflow](../.github/workflows/proof.yml) reruns the scripts and commits their data and generated page.
|
|
25
23
|
|
|
26
24
|
## Try the acceptance gate
|
|
27
25
|
|
|
@@ -70,22 +68,6 @@ Kind: reproducible local measurement. Sample size: 4. Measured: 2026-09-27. Expi
|
|
|
70
68
|
|
|
71
69
|
Source: [proof/scripts/check-gate.js](../proof/scripts/check-gate.js).
|
|
72
70
|
|
|
73
|
-
### Main conversation browser tokens
|
|
74
|
-
|
|
75
|
-
Measured on the author's Claude Code sessions: every browser tool call in the transcripts, counting tokens re-read by the main conversation per step. Median: 408,147 tokens per browser step across 145 sessions and 9,775 browser steps. Sample size counts browser steps.
|
|
76
|
-
|
|
77
|
-
Kind: measured on the author's setup. Sample size: 9775. Measured: 2026-09-26. Expires: 2026-10-27.
|
|
78
|
-
|
|
79
|
-
Source: author setup, re-measured locally.
|
|
80
|
-
|
|
81
|
-
### Small browser subagent tokens
|
|
82
|
-
|
|
83
|
-
At least 23x fewer tokens per step in this sample: the main conversation median of 408,147 divided by the highest subagent run of 17,197 is 23.73x. Measured on the author's Claude Code sessions, with the same kind of work handed to a small browser subagent. Five runs, tokens divided by steps or tool calls per run: 17,197 (206,369 tokens / 12 steps, 2026-09-26); 4,077 (93,773 / 23), 4,201 (105,034 / 25), 4,585 (91,709 / 20), and 4,489 (94,263 / 21), all four on 2026-09-27. Range: 4,077 to 17,197; median of 5 runs: 4,489. Typical context, using the median of 5 runs: about 91x fewer tokens per step. These are measurements from the author's own sessions, not a controlled comparison or a guarantee for other setups. The four 2026-09-27 runs shared one browser pane, so some steps were spent recovering a drifting tab, which raises the step count and lowers per-step tokens. The first run counted steps; the later runs counted tool calls. Sample size counts runs.
|
|
84
|
-
|
|
85
|
-
Kind: measured on the author's setup. Sample size: 5. Measured: 2026-09-27. Expires: 2026-10-27.
|
|
86
|
-
|
|
87
|
-
Source: author setup, re-measured locally.
|
|
88
|
-
|
|
89
71
|
## Operation and verification
|
|
90
72
|
|
|
91
73
|
- **What and why:** executable measurements keep public figures traceable to current output.
|
|
@@ -95,6 +77,6 @@ Source: author setup, re-measured locally.
|
|
|
95
77
|
- **Reads:** package scripts, the installer, runner and acceptance-check runner. Fixture tests use an isolated home and PATH.
|
|
96
78
|
- **Writes:** results.json, this generated page, temporary fixture directories and local fixture logs. The recorder writes gate-demo.cast and gate-demo.gif.
|
|
97
79
|
- **Closed loop:** the workflow fails when measurement or tests fail. GitHub Actions records the failure; repository notification settings decide who receives it. No separate alert service is configured.
|
|
98
|
-
- **Failure modes:** runner behavior changes,
|
|
99
|
-
- **Run and verify:** run the commands above, inspect sample arrays and fixture exit codes in results.json, and require npm test to pass.
|
|
100
|
-
- **Source of truth:** results.json and the scripts it names.
|
|
80
|
+
- **Failure modes:** runner behavior changes, missing runtime, unavailable write permission, or timing noise. Review the failed job, rerun locally, and send a reproducible issue to the repository maintainers.
|
|
81
|
+
- **Run and verify:** run the commands above, inspect sample arrays and fixture exit codes in results.json, and require npm test to pass.
|
|
82
|
+
- **Source of truth:** results.json and the scripts it names.
|
package/proof/results.json
CHANGED
|
@@ -169,30 +169,6 @@
|
|
|
169
169
|
"exitCode": 1
|
|
170
170
|
}
|
|
171
171
|
]
|
|
172
|
-
},
|
|
173
|
-
{
|
|
174
|
-
"id": "author-main-browser-tokens",
|
|
175
|
-
"label": "Main conversation browser tokens",
|
|
176
|
-
"kind": "author-setup",
|
|
177
|
-
"value": 408147,
|
|
178
|
-
"unit": "tokens per browser step (median)",
|
|
179
|
-
"measuredAt": "2026-09-26",
|
|
180
|
-
"method": "Measured on the author's Claude Code sessions: every browser tool call in the transcripts, counting tokens re-read by the main conversation per step. Median: 408,147 tokens per browser step across 145 sessions and 9,775 browser steps. Sample size counts browser steps.",
|
|
181
|
-
"sampleSize": 9775,
|
|
182
|
-
"script": "author setup, re-measured locally",
|
|
183
|
-
"expiresAt": "2026-10-27"
|
|
184
|
-
},
|
|
185
|
-
{
|
|
186
|
-
"id": "author-subagent-browser-tokens",
|
|
187
|
-
"label": "Small browser subagent tokens",
|
|
188
|
-
"kind": "author-setup",
|
|
189
|
-
"value": 17197,
|
|
190
|
-
"unit": "tokens per step (highest of 5 runs)",
|
|
191
|
-
"measuredAt": "2026-09-27",
|
|
192
|
-
"method": "At least 23x fewer tokens per step in this sample: the main conversation median of 408,147 divided by the highest subagent run of 17,197 is 23.73x. Measured on the author's Claude Code sessions, with the same kind of work handed to a small browser subagent. Five runs, tokens divided by steps or tool calls per run: 17,197 (206,369 tokens / 12 steps, 2026-09-26); 4,077 (93,773 / 23), 4,201 (105,034 / 25), 4,585 (91,709 / 20), and 4,489 (94,263 / 21), all four on 2026-09-27. Range: 4,077 to 17,197; median of 5 runs: 4,489. Typical context, using the median of 5 runs: about 91x fewer tokens per step. These are measurements from the author's own sessions, not a controlled comparison or a guarantee for other setups. The four 2026-09-27 runs shared one browser pane, so some steps were spent recovering a drifting tab, which raises the step count and lowers per-step tokens. The first run counted steps; the later runs counted tool calls. Sample size counts runs.",
|
|
193
|
-
"sampleSize": 5,
|
|
194
|
-
"script": "author setup, re-measured locally",
|
|
195
|
-
"expiresAt": "2026-10-27"
|
|
196
172
|
}
|
|
197
173
|
]
|
|
198
174
|
}
|
package/proof/scripts/lib.js
CHANGED
|
@@ -64,7 +64,9 @@ export function validateResults(data, now = new Date()) {
|
|
|
64
64
|
else {
|
|
65
65
|
if (item.measuredAt > today) errors.push(`${item.id}: measurement is in the future`);
|
|
66
66
|
if (item.expiresAt < item.measuredAt) errors.push(`${item.id}: expiry precedes measurement`);
|
|
67
|
-
|
|
67
|
+
// Expiry is a refresh signal, never a gate: an expired figure is reported and
|
|
68
|
+
// left off the pages, and tests, CI and releases keep running.
|
|
69
|
+
if (item.expiresAt < today) console.warn(`proof: ${item.id} expired ${item.expiresAt}; rerun node proof/scripts/measure.js`);
|
|
68
70
|
}
|
|
69
71
|
const localAuthorSource = item.kind === 'author-setup' && item.script === 'author setup, re-measured locally';
|
|
70
72
|
if (!localAuthorSource && (typeof item.script !== 'string' || !/^proof\/scripts\/[A-Za-z0-9_-]+\.js$/.test(item.script))) errors.push(`${item.id}: script must name a proof script`);
|
package/proof/scripts/render.js
CHANGED
|
@@ -7,7 +7,7 @@ export function proofMarkdown(data) {
|
|
|
7
7
|
const source = (e, label) => e.kind === 'author-setup' && e.script === 'author setup, re-measured locally' ? e.script : `[${label}](../${e.script})`;
|
|
8
8
|
const rows = data.entries.map(e => `| ${e.label} | ${e.value.toFixed(Number.isInteger(e.value) ? 0 : 2)} ${e.unit} | ${e.sampleSize} | ${e.measuredAt} | ${e.expiresAt} | ${source(e, 'script')} |`);
|
|
9
9
|
const methods = data.entries.map(e => `### ${e.label}\n\n${e.method}\n\nKind: ${e.kind === 'author-setup' ? "measured on the author's setup" : 'reproducible local measurement'}. Sample size: ${e.sampleSize}. Measured: ${e.measuredAt}. Expires: ${e.expiresAt}.\n\nSource: ${source(e, e.script)}.`).join('\n\n');
|
|
10
|
-
return `# Reproduce the measurements\n\nGenerated from [results.json](results.json). Each figure has a method, sample size, measurement date and expiry. Run the scripts on your own machine to compare.\n\nEnvironment: Node ${data.environment.node}, ${data.environment.platform} ${data.environment.arch}. Timing varies with startup caches and other work on the machine. Synthetic cases show what those fixtures exercise.\n\n| Measurement | Result | Sample size | Measured | Expires | Reproduce |\n|---|---|---|---|---|---|\n${rows.join('\n')}\n\n## Run the proof scripts\n\n\`\`\`sh\nnode proof/scripts/measure.js\nnode proof/scripts/render.js\nnpm test\n\`\`\`\n\nThe measurement command refreshes reproducible entries
|
|
10
|
+
return `# Reproduce the measurements\n\nGenerated from [results.json](results.json). Each figure has a method, sample size, measurement date and expiry. Run the scripts on your own machine to compare.\n\nEnvironment: Node ${data.environment.node}, ${data.environment.platform} ${data.environment.arch}. Timing varies with startup caches and other work on the machine. Synthetic cases show what those fixtures exercise.\n\n| Measurement | Result | Sample size | Measured | Expires | Reproduce |\n|---|---|---|---|---|---|\n${rows.join('\n')}\n\n## Run the proof scripts\n\n\`\`\`sh\nnode proof/scripts/measure.js\nnode proof/scripts/render.js\nnpm test\n\`\`\`\n\nThe measurement command refreshes the reproducible entries. An expired entry is a signal to re-measure; the test suite rejects future-dated or incomplete entries and checks this page against the data. The weekly [refresh workflow](../.github/workflows/proof.yml) reruns the scripts and commits their data and generated page.\n\n## Try the acceptance gate\n\n\`\`\`sh\naunx checks ACCEPTANCE_CHECKS.json\naunx checks run ACCEPTANCE_CHECKS.json\n\`\`\`\n\nThe scaffold starts red. Replace the sample with commands that prove your requirements, then put \`aunx checks run ACCEPTANCE_CHECKS.json && <your-release-command>\` in your own release sequence. Commands are local code you review before running. Manual evidence stays UNVERIFIED and blocks the gate.\n\n\n\nThe [recording script](scripts/record-gate.js) captures real command output into an asciicast, then renders it with an already installed agg. Companion tools are installed by their users.\n\n## Measurement methods\n\n${methods}\n\n## Operation and verification\n\n- **What and why:** executable measurements keep public figures traceable to current output.\n- **Trigger:** weekly schedule, workflow dispatch, or \`node proof/scripts/measure.js\`.\n- **Invocation chain:** workflow -> measurement functions -> isolated Node fixtures -> results.json -> this page -> npm test.\n- **Dependencies:** Node and the repository. The optional GIF recorder uses agg from the asciinema project.\n- **Reads:** package scripts, the installer, runner and acceptance-check runner. Fixture tests use an isolated home and PATH.\n- **Writes:** results.json, this generated page, temporary fixture directories and local fixture logs. The recorder writes gate-demo.cast and gate-demo.gif.\n- **Closed loop:** the workflow fails when measurement or tests fail. GitHub Actions records the failure; repository notification settings decide who receives it. No separate alert service is configured.\n- **Failure modes:** runner behavior changes, missing runtime, unavailable write permission, or timing noise. Review the failed job, rerun locally, and send a reproducible issue to the repository maintainers.\n- **Run and verify:** run the commands above, inspect sample arrays and fixture exit codes in results.json, and require npm test to pass.\n- **Source of truth:** results.json and the scripts it names.\n`;
|
|
11
11
|
}
|
|
12
12
|
export function renderPage() {
|
|
13
13
|
const data = readResults();
|
package/src/install.js
CHANGED
|
@@ -823,7 +823,7 @@ export function planFiles(opts) {
|
|
|
823
823
|
enabled: selected.filter((a) => a.facts.cliRun).map((a) => a.id),
|
|
824
824
|
defaults: Object.fromEntries((opts.effortAuto || []).map((lane) => [lane, { effort: 'auto' }])),
|
|
825
825
|
note: 'Lanes cli-run may call. Edit to enable or disable a lane. A lane not listed here exits 13 (unavailable).',
|
|
826
|
-
defaultsNote: 'Pin what a lane runs with, so the route in your docs is the route that runs: "defaults": {"' + (selected.find(a => a.facts.cliRun)?.id || '<lane>') + '": ' + JSON.stringify({ model: '<model-id>', ...(LANE_FLAGS[selected.find(a => a.facts.cliRun)?.id]?.effort ? { effort: 'high' } : {}) }) + '}. Left empty, a lane inherits its own config file, which cli-run cannot see and does not guess. `--model` and `--effort` override this per call, and `--doctor` prints what each lane is pinned to. Every enabled lane takes a model; the runner reports which lanes support an effort flag.'
|
|
826
|
+
defaultsNote: 'Pin what a lane runs with, so the route in your docs is the route that runs: "defaults": {"' + (selected.find(a => a.facts.cliRun)?.id || '<lane>') + '": ' + JSON.stringify({ model: '<model-id>', ...(LANE_FLAGS[selected.find(a => a.facts.cliRun)?.id]?.effort ? { effort: 'high' } : {}) }) + '}. Left empty, a lane inherits its own config file, which cli-run cannot see and does not guess. `--model` and `--effort` override this per call, and `--doctor` prints what each lane is pinned to. Every enabled lane takes a model; the runner reports which lanes support an effort flag. Hermes also takes "provider" beside "model", so the model goes to a provider that serves it.'
|
|
827
827
|
},
|
|
828
828
|
null,
|
|
829
829
|
2
|
|
@@ -97,13 +97,15 @@ Replace the placeholder with a current vendor model ID before using this example
|
|
|
97
97
|
|
|
98
98
|
The runner supports these lanes whether or not you selected them.
|
|
99
99
|
|
|
100
|
-
| Lane | Model flag | Effort flag |
|
|
101
|
-
|
|
102
|
-
| grok | `-m` | `--reasoning-effort` |
|
|
103
|
-
| codex | `-m` | `-c model_reasoning_effort="LEVEL"` |
|
|
104
|
-
| agy | `--model` | `--effort` |
|
|
105
|
-
| hermes | `-m` | `--reasoning` |
|
|
106
|
-
| qwen | `-m` | Unsupported; an effort request is a usage error |
|
|
100
|
+
| Lane | Model flag | Effort flag | Provider flag |
|
|
101
|
+
|---|---|---|---|
|
|
102
|
+
| grok | `-m` | `--reasoning-effort` | Unsupported |
|
|
103
|
+
| codex | `-m` | `-c model_reasoning_effort="LEVEL"` | Unsupported |
|
|
104
|
+
| agy | `--model` | `--effort` | Unsupported |
|
|
105
|
+
| hermes | `-m` | `--reasoning` | `--provider` |
|
|
106
|
+
| qwen | `-m` | Unsupported; an effort request is a usage error | Unsupported |
|
|
107
|
+
|
|
108
|
+
Hermes signs in to many providers, and a model sent to a provider that does not serve it fails with HTTP 400. Pin both with `--provider` and `--model`, or with `"provider"` beside `"model"` under `"hermes"` in `bin/lanes.json` defaults. A provider with no model is a usage error. That failure is reported as `rejected` (exit 16) with a model/provider mismatch line; `hermes model` repairs the pairing in Hermes itself.
|
|
107
109
|
|
|
108
110
|
When using `--effort auto`, treat its medium/high selection as a bounded heuristic; an audit has a high floor. Use explicit high for builds and xhigh where supported for security-critical or irreversible work. The vendor validates its own effort names and reports unsupported values through the failure class.
|
|
109
111
|
|