claude-autorouter 0.3.7 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.env.example +6 -3
- package/CODE_OF_CONDUCT.md +9 -0
- package/CONTRIBUTING.md +57 -0
- package/README.md +47 -70
- package/SECURITY.md +23 -0
- package/SUPPORT.md +18 -0
- package/bin/autorouter.mjs +40 -57
- package/docs/development.md +48 -2
- package/docs/hardware-benchmark.md +29 -0
- package/docs/hardware-comparison.md +55 -0
- package/docs/hardware-results-16gb.json +4002 -0
- package/docs/hardware-results-16gb.md +26 -0
- package/docs/hardware-results-64gb.json +4020 -0
- package/docs/reference.md +92 -37
- package/docs/releasing.md +79 -37
- package/docs/router-performance.json +1697 -0
- package/docs/router-performance.md +50 -0
- package/docs/status-performance.json +363 -0
- package/docs/status-performance.md +44 -0
- package/docs/subscription-integration.md +27 -0
- package/package.json +66 -10
- package/src/auto-routing.mjs +184 -24
- package/src/bounded-json.mjs +57 -0
- package/src/cli-help.mjs +90 -0
- package/src/config-command.mjs +158 -0
- package/src/config.mjs +53 -28
- package/src/contracts.mjs +123 -0
- package/src/evaluation-report.mjs +114 -0
- package/src/keychain.mjs +58 -0
- package/src/local-diagnostic.mjs +191 -0
- package/src/model-catalog.mjs +96 -0
- package/src/model-request.mjs +6 -7
- package/src/ollama-evaluator.mjs +9 -27
- package/src/onboarding.mjs +130 -26
- package/src/prompt-state.mjs +22 -7
- package/src/redaction.mjs +97 -0
- package/src/request-validation.mjs +54 -0
- package/src/response-observer.mjs +126 -18
- package/src/router.mjs +151 -61
- package/src/savings.mjs +74 -16
- package/src/server.mjs +79 -12
- package/src/session-history.mjs +262 -0
- package/src/session-log.mjs +9 -58
- package/src/status-state.mjs +110 -62
- package/src/statusline.mjs +57 -27
- package/src/telemetry-event.mjs +200 -0
- package/src/token-counter.mjs +3 -1
- package/src/turn-state.mjs +132 -0
- package/src/user-config.mjs +81 -10
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
# Local-model hardware validation
|
|
2
|
+
|
|
3
|
+
Actual Tev1 and Nimble evaluators have now been measured on a 16 GiB M4 and a 64 GiB M2 Ultra; see [the comparison and limitations](hardware-comparison.md). These instructions reproduce the opt-in benchmark on another Mac. It uses only checked-in synthetic tasks; no Jev or Anthropic calls, user configuration, user prompts or model downloads are involved.
|
|
4
|
+
|
|
5
|
+
Transfer `autorouter-hardware-benchmark.tar.gz` and its checksum to the other Mac. The bundle contains an explicit source allowlist and hashes; it excludes credentials, configurations, histories, git metadata and dependencies. Extract it:
|
|
6
|
+
|
|
7
|
+
```sh
|
|
8
|
+
shasum -a 256 -c autorouter-hardware-benchmark.tar.gz.sha256
|
|
9
|
+
tar -xzf autorouter-hardware-benchmark.tar.gz
|
|
10
|
+
cd autorouter-benchmark
|
|
11
|
+
node --version
|
|
12
|
+
node scripts/evaluate-ollama.mjs --help
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
Node 22+ and a running Ollama 0.35+ are required. No npm install is needed. Use already-installed model tags. Start when Ollama has no resident models; the benchmark refuses to displace unrelated resident models. It loads each selected candidate for a separate cold observation and repeated warm evaluations, then unloads that candidate from memory. It never deletes downloaded models. Keep other applications running at a representative load; CPU load averages and free memory are recorded, without process names or arguments.
|
|
16
|
+
|
|
17
|
+
Use the same two model tags and 30-second warm deadline as the 16 GiB baseline:
|
|
18
|
+
|
|
19
|
+
```sh
|
|
20
|
+
node scripts/evaluate-ollama.mjs --models tev1:4b-q4_K_M,nimble:9b-q4_K_M --split heldout --rounds 3 --stress-rounds 8 --timeout-ms 30000 --cold-timeout-ms 60000 --output hardware-report.json
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
For an installed Tev1 model, substitute its exact tag, such as `tev1:4b-q4_K_M`. Multiple installed candidates can be comma-separated. Use `--timeout-ms 0` to measure without the runtime timer; initial loading still has the separate cold deadline. Measure finite-deadline reliability in a separate report as well. Do not change label-agreement thresholds after seeing results: the default requires complete expected-label agreement, no under-routing and all three classified tiers. A nonzero exit can indicate a quality-gate failure while the JSON still contains valid measurements.
|
|
24
|
+
|
|
25
|
+
For a controlled comparison, match model artifact digests as well as tags, quantization, context allocation and runtime versions. The existing cross-host reports have different model digests and other conditions, so they are observational measurements rather than a RAM-only comparison.
|
|
26
|
+
|
|
27
|
+
The report contains hardware/Node/Ollama versions, initial/final background load/free memory, fixture and rubric hashes, model identity/quantization/resident memory, cold and repeated warm latency, full-budget stress cases, deadline failures and classification gates. Classification quality is agreement with the frozen synthetic rubric; selected and confirmed Claude models remain unmeasured. It does not prove end-user task quality or net savings.
|
|
28
|
+
|
|
29
|
+
Return `hardware-report.json` for comparison with the same workload on the 16 GiB Mac. Keep the bundle's `source-manifest.json` with it. Record whether the machine was in normal use or deliberately idle. A missing model or quality failure is reported explicitly rather than hidden by choosing a passing model or relabeling cases.
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# Local evaluator comparison: 16 GiB M4 and 64 GiB M2 Ultra
|
|
2
|
+
|
|
3
|
+
The October 5, 2026 measurements now cover both required memory classes. The returned 64 GiB report contains valid cold, warm and full-budget observations from a Mac in normal use with the usual applications open. Both candidates completed every request there, but both still failed the unchanged classification gate. Jev remains the default; these local measurements do not establish production routing accuracy.
|
|
4
|
+
|
|
5
|
+
The frozen workload is one cold request, three deterministically shuffled rounds of 30 held-out synthetic tasks, and eight separate 3,000-byte stress tasks per model. Warm and cold deadlines were 30 and 60 seconds respectively. The benchmark uses native `/v1/systemone` choice scoring, with no Claude generations, Jev requests or model downloads. The reported latency is end-to-end wall time, not isolated model computation.
|
|
6
|
+
|
|
7
|
+
| Measurement | Tev1 4B, 16 GiB | Tev1 4B, 64 GiB | Nimble 9B, 16 GiB | Nimble 9B, 64 GiB |
|
|
8
|
+
| --- | ---: | ---: | ---: | ---: |
|
|
9
|
+
| Cold wall time | 7.13 s | 3.57 s | 15.93 s | 3.77 s |
|
|
10
|
+
| Warm requests / valid results | 90 / 90 | 90 / 90 | 90 / 87 | 90 / 90 |
|
|
11
|
+
| Warm wall p50 / p95, including errors | 244 / 2,930 ms | 55 / 519 ms | 1,250 / 13,645 ms | 56 / 863 ms |
|
|
12
|
+
| Successful warm wall p50 / p95 | 244 / 2,930 ms | 55 / 519 ms | 1,148 / 11,249 ms | 56 / 863 ms |
|
|
13
|
+
| Exact frozen-label agreement | 93.3% | 93.3% | 93.3% | 96.7% |
|
|
14
|
+
| Under-routes / over-routes | 3 / 3 | 3 / 3 | 3 / 0 | 3 / 0 |
|
|
15
|
+
| Warm deadline errors | 0 | 0 | 3 | 0 |
|
|
16
|
+
| Full-budget stress p50 / p95 | 7,003 / 7,854 ms | 1,122 / 1,314 ms | 12,369 / 13,945 ms | 1,940 / 2,003 ms |
|
|
17
|
+
| Strict warm classification gate | Failed | Failed | Failed | Failed |
|
|
18
|
+
|
|
19
|
+
Tev1 under-routed `held-o-compiler` from Opus to Sonnet in all three rounds and over-routed `held-h-literal-opus` from Haiku to Opus. Nimble under-routed `held-o-injection` from Opus to Haiku in all three rounds. The 16 GiB Nimble run also had three timeout errors, which count against its agreement. Both models reached all three tiers and passed all eight stress labels on both hosts. Those stress tasks all expect Haiku; passing them does not establish broader accuracy. The preset warm gate remains 100% exact agreement, zero under-routing and all three classified tiers.
|
|
20
|
+
|
|
21
|
+
On the larger Mac, Tev1's stress p95 was 1.31 seconds and Nimble's was 2.00 seconds. The zero-error results were obtained with a 30-second warm deadline and do not certify reliability under a 1.5-second deadline. These observations support configurable model-specific deadlines, including the explicit disabled-deadline option. They do not determine an optimal deadline or justify changing the existing defaults.
|
|
22
|
+
|
|
23
|
+
## Conditions and comparability
|
|
24
|
+
|
|
25
|
+
| Condition | 16 GiB baseline | 64 GiB returned run |
|
|
26
|
+
| --- | --- | --- |
|
|
27
|
+
| Processor / logical CPUs | Apple M4 / 10 | Apple M2 Ultra / 24 |
|
|
28
|
+
| Unified memory | 16 GiB | 64 GiB |
|
|
29
|
+
| OS release / architecture | 25.6.0 / arm64 | 27.0.0 / arm64 |
|
|
30
|
+
| Node / Ollama | v26.10.0 / 0.35.0 | v22.14.0 / 0.35.1 |
|
|
31
|
+
| Initial 1/5/15-minute load | 3.55 / 3.35 / 3.48 | 3.87 / 5.12 / 4.93 |
|
|
32
|
+
| Final 1/5/15-minute load | 2.23 / 3.60 / 3.67 | 4.49 / 5.09 / 4.94 |
|
|
33
|
+
| Initial / final OS free memory | 842 / 122 MiB | 727 / 136 MiB |
|
|
34
|
+
| Background use | Applications running | Normal use, usual applications open |
|
|
35
|
+
|
|
36
|
+
Both hosts reported Q4_K_M quantization, 4.2B/9.0B parameters, context allocations of 2,050/8,194 tokens and model residency of 2.71/5.74 GiB for Tev1/Nimble. Residency is reported by Ollama; free memory is an OS snapshot, not total available memory or measured swap pressure. The JSON also records aggregate RSS across matching Ollama daemon/runner processes, which is not a pure model footprint. On the larger host, that aggregate rose from 2.94 to 7.99 GiB for Tev1 and 5.80 to 9.42 GiB for Nimble.
|
|
37
|
+
|
|
38
|
+
Model tags match but artifact digests differ:
|
|
39
|
+
|
|
40
|
+
| Model tag | 16 GiB digest | 64 GiB digest |
|
|
41
|
+
| --- | --- | --- |
|
|
42
|
+
| `tev1:4b-q4_K_M` | `3509ac7180e86e5fa8efc7b5745e32d55dd9d4e0a86bc9a88aba5323a5d29bc6` | `07be32e6e5e3dcc6b2deae7c68b89321e99daeedb08521e92771cc155473cbd9` |
|
|
43
|
+
| `nimble:9b-q4_K_M` | `3776806da5587387a996e28e75d5d07fbebe7410d47879e1f71782f36c896ce3` | `572f1f4c801d37171344b0520d14853a88887c249d11db3acfa6c81a34cdcfa3` |
|
|
44
|
+
|
|
45
|
+
Each returned download size is 28 bytes larger. This does not establish whether weights, templates or metadata changed. The comparison is observational: processor, runtime versions, background conditions and model artifacts differ, so faster timings cannot be attributed to RAM alone. A controlled hardware comparison would require matching model digests and runtime conditions. The 16 GiB timeout rows also recorded wall times exceeding their configured 30-second deadline, up to 969,208.59 ms; their cause was not established, so the timer should not be described as a strict wall-time upper bound in that run.
|
|
46
|
+
|
|
47
|
+
## Evidence and validation
|
|
48
|
+
|
|
49
|
+
[The 16 GiB report](hardware-results-16gb.json) and [the 64 GiB report](hardware-results-64gb.json) retain all per-case outcomes and unchanged gates. The larger report is complete while its overall `passed` field remains `false`, correctly reflecting failed classification gates. An independent offline review recomputed case/round coverage, labels, UTF-8 state sizes, percentile summaries, confusion matrices and acceptance gates.
|
|
50
|
+
|
|
51
|
+
All 36 hashes in the returned source manifest matched the frozen transfer sources when received. The fixture SHA-256 is `1ee6ab593e47d87a67e922c0afb0c9db5787d8a27d1a0ff3e24baa0543291fc5`; both model rubric hashes are `be151cedb4de4b7ef3f7162d751f70ce7d9dd14efc66fae1835f73ffd04027be`. Subsequent documentation and package-allowlist changes do not change those evaluated sources. The original transfer bundle is retained unchanged.
|
|
52
|
+
|
|
53
|
+
The returned report SHA-256 is `b76d25b3b062821fa2483e13847d79267734f1a83b35ffe6ea85c57f82f7f715`; its manifest SHA-256 is `dab377757912028fca5a929dd4ef2123a4c4001559f0d4ee26fd6bec092b04f1`. The public JSON includes this provenance and the user's background-use description; it contains synthetic case identifiers and metadata, without user prompts or process names. Reproduction instructions remain in [hardware-benchmark.md](hardware-benchmark.md).
|
|
54
|
+
|
|
55
|
+
The larger-memory measurement requirement is fulfilled. Classifier errors remain visible rather than being hidden by altered labels or thresholds. Neither this comparison nor the separate [router benchmark](router-performance.md) establishes parity with Jev, completed Claude task quality or subscription savings.
|