model-orchestrator 1.0.4 → 1.0.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md CHANGED
@@ -33,6 +33,6 @@ When using the Claude Code plugin, follow [plugin/README.md](plugin/README.md).
33
33
  - **Installer safety:** writes remain inside `--dir` and `--project`; preserve user edits according to manifest hashes and explicit flags. Run no third-party installer.
34
34
  - **Runner safety:** preserve exit codes and the log schema. Before changing an output judge, add a failing case in `test/judges.test.js`.
35
35
  - **Secrets:** use environment-variable names only. Never add a credential value to code, examples or tests.
36
- - **Proof:** measure through `proof/scripts/`, store results in `proof/results.json` and regenerate the proof page. The suite rejects expired entries.
36
+ - **Proof:** measure through `proof/scripts/`, store results in `proof/results.json` and regenerate the proof page. An expired entry warns and is re-measured; it never blocks tests or releases.
37
37
  - **Verify:** run `npm test` and report tests, pass, fail, skipped and exit code. The suite prints current counts. Use `npm pack --dry-run` to inspect publication contents.
38
38
  - **Style:** use short condition-to-action instructions and no em dashes. `test/prose.test.js` checks public vocabulary and examples.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,13 @@ All notable changes to this project are documented here. The format follows [Kee
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [1.0.5] - 2026-09-28
8
+
9
+ ### Changed
10
+
11
+ - The proof page keeps only figures that measure this package. The two browser token figures measured the author's own browser subagent, which model-orchestrator does not ship, and are removed from the proof data, the site and the trailer.
12
+ - An expired proof figure is a signal to re-measure: it warns and never fails tests, CI or a release. The product site leaves expired figures off.
13
+
7
14
  ## [1.0.4] - 2026-09-28
8
15
 
9
16
  ### Security
@@ -550,7 +557,8 @@ First release.
550
557
  - Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
551
558
  - Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
552
559
 
553
- [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.4...HEAD
560
+ [Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.5...HEAD
561
+ [1.0.5]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.4...v1.0.5
554
562
  [1.0.4]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.3...v1.0.4
555
563
  [1.0.3]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.2...v1.0.3
556
564
  [1.0.2]: https://github.com/aunysillyme/model-orchestrator/compare/v1.0.1...v1.0.2
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "model-orchestrator",
3
- "version": "1.0.4",
3
+ "version": "1.0.5",
4
4
  "description": "Model router for AI coding agents: installs routing rules, 8 subagents, hooks and a CLI runner so your AI picks model and effort per task and saves tokens",
5
5
  "type": "module",
6
6
  "bin": {
package/proof/README.md CHANGED
@@ -10,8 +10,6 @@ Environment: Node v22.22.3, darwin arm64. Timing varies with startup caches and
10
10
  | Lane runner overhead | 63.74 ms median difference | 7 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/runner-overhead.js) |
11
11
  | Empty results flagged | 10 fixtures rejected | 10 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/missing-results.js) |
12
12
  | Acceptance failures blocked | 4 fixtures rejected | 4 | 2026-09-27 | 2026-10-11 | [script](../proof/scripts/check-gate.js) |
13
- | Main conversation browser tokens | 408147 tokens per browser step (median) | 9775 | 2026-09-26 | 2026-10-27 | author setup, re-measured locally |
14
- | Small browser subagent tokens | 17197 tokens per step (highest of 5 runs) | 5 | 2026-09-27 | 2026-10-27 | author setup, re-measured locally |
15
13
 
16
14
  ## Run the proof scripts
17
15
 
@@ -21,7 +19,7 @@ node proof/scripts/render.js
21
19
  npm test
22
20
  ```
23
21
 
24
- The measurement command refreshes reproducible entries and preserves separately sourced author-setup entries. The test suite rejects expired, future-dated or incomplete entries and checks this page against the data. The weekly [refresh workflow](../.github/workflows/proof.yml) reruns the scripts and commits their data and generated page.
22
+ The measurement command refreshes the reproducible entries. An expired entry is a signal to re-measure; the test suite rejects future-dated or incomplete entries and checks this page against the data. The weekly [refresh workflow](../.github/workflows/proof.yml) reruns the scripts and commits their data and generated page.
25
23
 
26
24
  ## Try the acceptance gate
27
25
 
@@ -70,22 +68,6 @@ Kind: reproducible local measurement. Sample size: 4. Measured: 2026-09-27. Expi
70
68
 
71
69
  Source: [proof/scripts/check-gate.js](../proof/scripts/check-gate.js).
72
70
 
73
- ### Main conversation browser tokens
74
-
75
- Measured on the author's Claude Code sessions: every browser tool call in the transcripts, counting tokens re-read by the main conversation per step. Median: 408,147 tokens per browser step across 145 sessions and 9,775 browser steps. Sample size counts browser steps.
76
-
77
- Kind: measured on the author's setup. Sample size: 9775. Measured: 2026-09-26. Expires: 2026-10-27.
78
-
79
- Source: author setup, re-measured locally.
80
-
81
- ### Small browser subagent tokens
82
-
83
- At least 23x fewer tokens per step in this sample: the main conversation median of 408,147 divided by the highest subagent run of 17,197 is 23.73x. Measured on the author's Claude Code sessions, with the same kind of work handed to a small browser subagent. Five runs, tokens divided by steps or tool calls per run: 17,197 (206,369 tokens / 12 steps, 2026-09-26); 4,077 (93,773 / 23), 4,201 (105,034 / 25), 4,585 (91,709 / 20), and 4,489 (94,263 / 21), all four on 2026-09-27. Range: 4,077 to 17,197; median of 5 runs: 4,489. Typical context, using the median of 5 runs: about 91x fewer tokens per step. These are measurements from the author's own sessions, not a controlled comparison or a guarantee for other setups. The four 2026-09-27 runs shared one browser pane, so some steps were spent recovering a drifting tab, which raises the step count and lowers per-step tokens. The first run counted steps; the later runs counted tool calls. Sample size counts runs.
84
-
85
- Kind: measured on the author's setup. Sample size: 5. Measured: 2026-09-27. Expires: 2026-10-27.
86
-
87
- Source: author setup, re-measured locally.
88
-
89
71
  ## Operation and verification
90
72
 
91
73
  - **What and why:** executable measurements keep public figures traceable to current output.
@@ -95,6 +77,6 @@ Source: author setup, re-measured locally.
95
77
  - **Reads:** package scripts, the installer, runner and acceptance-check runner. Fixture tests use an isolated home and PATH.
96
78
  - **Writes:** results.json, this generated page, temporary fixture directories and local fixture logs. The recorder writes gate-demo.cast and gate-demo.gif.
97
79
  - **Closed loop:** the workflow fails when measurement or tests fail. GitHub Actions records the failure; repository notification settings decide who receives it. No separate alert service is configured.
98
- - **Failure modes:** runner behavior changes, an expired catalog snapshot, missing runtime, unavailable write permission, or timing noise. Review the failed job, rerun locally, and send a reproducible issue to the repository maintainers.
99
- - **Run and verify:** run the commands above, inspect sample arrays and fixture exit codes in results.json, and require npm test to pass. A future-time unit test proves expiry can fail.
100
- - **Source of truth:** results.json and the scripts it names. Author-setup evidence is added separately by its owner.
80
+ - **Failure modes:** runner behavior changes, missing runtime, unavailable write permission, or timing noise. Review the failed job, rerun locally, and send a reproducible issue to the repository maintainers.
81
+ - **Run and verify:** run the commands above, inspect sample arrays and fixture exit codes in results.json, and require npm test to pass.
82
+ - **Source of truth:** results.json and the scripts it names.
@@ -169,30 +169,6 @@
169
169
  "exitCode": 1
170
170
  }
171
171
  ]
172
- },
173
- {
174
- "id": "author-main-browser-tokens",
175
- "label": "Main conversation browser tokens",
176
- "kind": "author-setup",
177
- "value": 408147,
178
- "unit": "tokens per browser step (median)",
179
- "measuredAt": "2026-09-26",
180
- "method": "Measured on the author's Claude Code sessions: every browser tool call in the transcripts, counting tokens re-read by the main conversation per step. Median: 408,147 tokens per browser step across 145 sessions and 9,775 browser steps. Sample size counts browser steps.",
181
- "sampleSize": 9775,
182
- "script": "author setup, re-measured locally",
183
- "expiresAt": "2026-10-27"
184
- },
185
- {
186
- "id": "author-subagent-browser-tokens",
187
- "label": "Small browser subagent tokens",
188
- "kind": "author-setup",
189
- "value": 17197,
190
- "unit": "tokens per step (highest of 5 runs)",
191
- "measuredAt": "2026-09-27",
192
- "method": "At least 23x fewer tokens per step in this sample: the main conversation median of 408,147 divided by the highest subagent run of 17,197 is 23.73x. Measured on the author's Claude Code sessions, with the same kind of work handed to a small browser subagent. Five runs, tokens divided by steps or tool calls per run: 17,197 (206,369 tokens / 12 steps, 2026-09-26); 4,077 (93,773 / 23), 4,201 (105,034 / 25), 4,585 (91,709 / 20), and 4,489 (94,263 / 21), all four on 2026-09-27. Range: 4,077 to 17,197; median of 5 runs: 4,489. Typical context, using the median of 5 runs: about 91x fewer tokens per step. These are measurements from the author's own sessions, not a controlled comparison or a guarantee for other setups. The four 2026-09-27 runs shared one browser pane, so some steps were spent recovering a drifting tab, which raises the step count and lowers per-step tokens. The first run counted steps; the later runs counted tool calls. Sample size counts runs.",
193
- "sampleSize": 5,
194
- "script": "author setup, re-measured locally",
195
- "expiresAt": "2026-10-27"
196
172
  }
197
173
  ]
198
174
  }
@@ -64,7 +64,9 @@ export function validateResults(data, now = new Date()) {
64
64
  else {
65
65
  if (item.measuredAt > today) errors.push(`${item.id}: measurement is in the future`);
66
66
  if (item.expiresAt < item.measuredAt) errors.push(`${item.id}: expiry precedes measurement`);
67
- if (item.expiresAt < today) errors.push(`${item.id}: expired ${item.expiresAt}`);
67
+ // Expiry is a refresh signal, never a gate: an expired figure is reported and
68
+ // left off the pages, and tests, CI and releases keep running.
69
+ if (item.expiresAt < today) console.warn(`proof: ${item.id} expired ${item.expiresAt}; rerun node proof/scripts/measure.js`);
68
70
  }
69
71
  const localAuthorSource = item.kind === 'author-setup' && item.script === 'author setup, re-measured locally';
70
72
  if (!localAuthorSource && (typeof item.script !== 'string' || !/^proof\/scripts\/[A-Za-z0-9_-]+\.js$/.test(item.script))) errors.push(`${item.id}: script must name a proof script`);
@@ -7,7 +7,7 @@ export function proofMarkdown(data) {
7
7
  const source = (e, label) => e.kind === 'author-setup' && e.script === 'author setup, re-measured locally' ? e.script : `[${label}](../${e.script})`;
8
8
  const rows = data.entries.map(e => `| ${e.label} | ${e.value.toFixed(Number.isInteger(e.value) ? 0 : 2)} ${e.unit} | ${e.sampleSize} | ${e.measuredAt} | ${e.expiresAt} | ${source(e, 'script')} |`);
9
9
  const methods = data.entries.map(e => `### ${e.label}\n\n${e.method}\n\nKind: ${e.kind === 'author-setup' ? "measured on the author's setup" : 'reproducible local measurement'}. Sample size: ${e.sampleSize}. Measured: ${e.measuredAt}. Expires: ${e.expiresAt}.\n\nSource: ${source(e, e.script)}.`).join('\n\n');
10
- return `# Reproduce the measurements\n\nGenerated from [results.json](results.json). Each figure has a method, sample size, measurement date and expiry. Run the scripts on your own machine to compare.\n\nEnvironment: Node ${data.environment.node}, ${data.environment.platform} ${data.environment.arch}. Timing varies with startup caches and other work on the machine. Synthetic cases show what those fixtures exercise.\n\n| Measurement | Result | Sample size | Measured | Expires | Reproduce |\n|---|---|---|---|---|---|\n${rows.join('\n')}\n\n## Run the proof scripts\n\n\`\`\`sh\nnode proof/scripts/measure.js\nnode proof/scripts/render.js\nnpm test\n\`\`\`\n\nThe measurement command refreshes reproducible entries and preserves separately sourced author-setup entries. The test suite rejects expired, future-dated or incomplete entries and checks this page against the data. The weekly [refresh workflow](../.github/workflows/proof.yml) reruns the scripts and commits their data and generated page.\n\n## Try the acceptance gate\n\n\`\`\`sh\naunx checks ACCEPTANCE_CHECKS.json\naunx checks run ACCEPTANCE_CHECKS.json\n\`\`\`\n\nThe scaffold starts red. Replace the sample with commands that prove your requirements, then put \`aunx checks run ACCEPTANCE_CHECKS.json && <your-release-command>\` in your own release sequence. Commands are local code you review before running. Manual evidence stays UNVERIFIED and blocks the gate.\n\n![Acceptance gate rejects a missing output, then passes after the output exists](gate-demo.gif)\n\nThe [recording script](scripts/record-gate.js) captures real command output into an asciicast, then renders it with an already installed agg. Companion tools are installed by their users.\n\n## Measurement methods\n\n${methods}\n\n## Operation and verification\n\n- **What and why:** executable measurements keep public figures traceable to current output.\n- **Trigger:** weekly schedule, workflow dispatch, or \`node proof/scripts/measure.js\`.\n- **Invocation chain:** workflow -> measurement functions -> isolated Node fixtures -> results.json -> this page -> npm test.\n- **Dependencies:** Node and the repository. The optional GIF recorder uses agg from the asciinema project.\n- **Reads:** package scripts, the installer, runner and acceptance-check runner. Fixture tests use an isolated home and PATH.\n- **Writes:** results.json, this generated page, temporary fixture directories and local fixture logs. The recorder writes gate-demo.cast and gate-demo.gif.\n- **Closed loop:** the workflow fails when measurement or tests fail. GitHub Actions records the failure; repository notification settings decide who receives it. No separate alert service is configured.\n- **Failure modes:** runner behavior changes, an expired catalog snapshot, missing runtime, unavailable write permission, or timing noise. Review the failed job, rerun locally, and send a reproducible issue to the repository maintainers.\n- **Run and verify:** run the commands above, inspect sample arrays and fixture exit codes in results.json, and require npm test to pass. A future-time unit test proves expiry can fail.\n- **Source of truth:** results.json and the scripts it names. Author-setup evidence is added separately by its owner.\n`;
10
+ return `# Reproduce the measurements\n\nGenerated from [results.json](results.json). Each figure has a method, sample size, measurement date and expiry. Run the scripts on your own machine to compare.\n\nEnvironment: Node ${data.environment.node}, ${data.environment.platform} ${data.environment.arch}. Timing varies with startup caches and other work on the machine. Synthetic cases show what those fixtures exercise.\n\n| Measurement | Result | Sample size | Measured | Expires | Reproduce |\n|---|---|---|---|---|---|\n${rows.join('\n')}\n\n## Run the proof scripts\n\n\`\`\`sh\nnode proof/scripts/measure.js\nnode proof/scripts/render.js\nnpm test\n\`\`\`\n\nThe measurement command refreshes the reproducible entries. An expired entry is a signal to re-measure; the test suite rejects future-dated or incomplete entries and checks this page against the data. The weekly [refresh workflow](../.github/workflows/proof.yml) reruns the scripts and commits their data and generated page.\n\n## Try the acceptance gate\n\n\`\`\`sh\naunx checks ACCEPTANCE_CHECKS.json\naunx checks run ACCEPTANCE_CHECKS.json\n\`\`\`\n\nThe scaffold starts red. Replace the sample with commands that prove your requirements, then put \`aunx checks run ACCEPTANCE_CHECKS.json && <your-release-command>\` in your own release sequence. Commands are local code you review before running. Manual evidence stays UNVERIFIED and blocks the gate.\n\n![Acceptance gate rejects a missing output, then passes after the output exists](gate-demo.gif)\n\nThe [recording script](scripts/record-gate.js) captures real command output into an asciicast, then renders it with an already installed agg. Companion tools are installed by their users.\n\n## Measurement methods\n\n${methods}\n\n## Operation and verification\n\n- **What and why:** executable measurements keep public figures traceable to current output.\n- **Trigger:** weekly schedule, workflow dispatch, or \`node proof/scripts/measure.js\`.\n- **Invocation chain:** workflow -> measurement functions -> isolated Node fixtures -> results.json -> this page -> npm test.\n- **Dependencies:** Node and the repository. The optional GIF recorder uses agg from the asciinema project.\n- **Reads:** package scripts, the installer, runner and acceptance-check runner. Fixture tests use an isolated home and PATH.\n- **Writes:** results.json, this generated page, temporary fixture directories and local fixture logs. The recorder writes gate-demo.cast and gate-demo.gif.\n- **Closed loop:** the workflow fails when measurement or tests fail. GitHub Actions records the failure; repository notification settings decide who receives it. No separate alert service is configured.\n- **Failure modes:** runner behavior changes, missing runtime, unavailable write permission, or timing noise. Review the failed job, rerun locally, and send a reproducible issue to the repository maintainers.\n- **Run and verify:** run the commands above, inspect sample arrays and fixture exit codes in results.json, and require npm test to pass.\n- **Source of truth:** results.json and the scripts it names.\n`;
11
11
  }
12
12
  export function renderPage() {
13
13
  const data = readResults();