claude-autorouter 0.3.7 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. package/.env.example +4 -2
  2. package/CONTRIBUTING.md +37 -0
  3. package/README.md +43 -70
  4. package/bin/autorouter.mjs +40 -57
  5. package/docs/development.md +48 -2
  6. package/docs/hardware-benchmark.md +29 -0
  7. package/docs/hardware-comparison.md +55 -0
  8. package/docs/hardware-results-16gb.json +4002 -0
  9. package/docs/hardware-results-16gb.md +26 -0
  10. package/docs/hardware-results-64gb.json +4020 -0
  11. package/docs/reference.md +71 -32
  12. package/docs/releasing.md +74 -34
  13. package/docs/router-performance.json +1697 -0
  14. package/docs/router-performance.md +50 -0
  15. package/docs/status-performance.json +363 -0
  16. package/docs/status-performance.md +44 -0
  17. package/package.json +57 -9
  18. package/src/auto-routing.mjs +184 -24
  19. package/src/bounded-json.mjs +57 -0
  20. package/src/cli-help.mjs +87 -0
  21. package/src/config-command.mjs +141 -0
  22. package/src/config.mjs +52 -27
  23. package/src/contracts.mjs +123 -0
  24. package/src/evaluation-report.mjs +114 -0
  25. package/src/local-diagnostic.mjs +191 -0
  26. package/src/model-catalog.mjs +96 -0
  27. package/src/model-request.mjs +6 -7
  28. package/src/ollama-evaluator.mjs +9 -27
  29. package/src/onboarding.mjs +82 -23
  30. package/src/request-validation.mjs +54 -0
  31. package/src/response-observer.mjs +126 -18
  32. package/src/router.mjs +151 -61
  33. package/src/savings.mjs +74 -16
  34. package/src/server.mjs +79 -12
  35. package/src/session-history.mjs +261 -0
  36. package/src/session-log.mjs +9 -58
  37. package/src/status-state.mjs +110 -62
  38. package/src/statusline.mjs +57 -27
  39. package/src/telemetry-event.mjs +196 -0
  40. package/src/token-counter.mjs +3 -1
  41. package/src/turn-state.mjs +132 -0
  42. package/src/user-config.mjs +18 -8
@@ -0,0 +1,26 @@
1
+ # Actual local evaluator measurements: 16 GiB M4
2
+
3
+ Measured October 5, 2026 using Ollama 0.35.0, Node v26.10.0 and an Apple M4 with 10 logical CPUs and 16 GiB unified memory. Applications remained running: initial 1/5/15-minute load averages were 3.55/3.35/3.48, and reported free memory was 842 MiB. Free memory is an OS snapshot, not total available memory or a measurement of swap pressure.
4
+
5
+ Each already-installed Q4_K_M candidate received one cold request, then three deterministic shuffled rounds of 30 held-out synthetic tasks, followed by eight separate full-excerpt stress tasks. The warm deadline was 30 seconds; the cold deadline was 60 seconds. The benchmark began with no resident Ollama models and unloaded only its own candidates afterward. No downloads, Jev requests or Claude inference occurred.
6
+
7
+ | Measurement | Tev1 4B Q4_K_M | Nimble 9B Q4_K_M |
8
+ | --- | ---: | ---: |
9
+ | Cold wall time | 7.13 s | 15.93 s |
10
+ | Warm requests / valid results | 90 / 90 | 90 / 87 |
11
+ | Warm wall p50 / p95, including errors | 244 / 2,930 ms | 1,250 / 13,645 ms |
12
+ | Successful warm wall p50 / p95 | 244 / 2,930 ms | 1,148 / 11,249 ms |
13
+ | Exact frozen-label agreement | 93.3% | 93.3% |
14
+ | Under-routes / over-routes | 3 / 3 | 3 / 0 |
15
+ | Deadline errors | 0 | 3 |
16
+ | Reported model residency | 2.71 GiB | 5.74 GiB |
17
+ | Reported context allocation | 2,050 tokens | 8,194 tokens |
18
+ | Full-budget stress p50 / p95 | 7,003 / 7,854 ms | 12,369 / 13,945 ms |
19
+
20
+ Both candidates classified tasks into all three tiers. Neither passed the preset strict gate of 100% agreement and zero under-routing. Tev1 classified one complex fixture as Sonnet in all three rounds and one Haiku fixture as Opus. Nimble classified one complex fixture as Haiku in all three rounds and timed out on three other requests. Both passed all eight stress labels, but the stress set contains only Haiku tasks and cannot establish broader classification accuracy.
21
+
22
+ These results support retaining model-specific deadlines and the explicit disabled-deadline option. They do not establish an optimal deadline, parity with Jev, downstream Claude task quality or subscription savings. Timings include local service work, model computation and scheduling under this background load; they are separate from the [synthetic router benchmark](router-performance.md). Differences in context allocation and model residency must be considered when comparing candidates.
23
+
24
+ [Machine-readable results](hardware-results-16gb.json) include per-case outcomes, model digests, token metrics, residency snapshots and background-load metadata. No user prompts or process names are included. The frozen fixture SHA-256 is `1ee6ab593e47d87a67e922c0afb0c9db5787d8a27d1a0ff3e24baa0543291fc5`; the rubric SHA-256 is `be151cedb4de4b7ef3f7162d751f70ce7d9dd14efc66fae1835f73ffd04027be`.
25
+
26
+ The same workload has now been measured on a 64 GiB M2 Ultra. See [the comparison](hardware-comparison.md) for the results and differences in model artifacts and runtime conditions. The Nimble timeout rows on this baseline recorded wall times exceeding their configured 30-second deadline, up to 969,208.59 ms; the cause was not established. The original observations and acceptance thresholds remain unchanged.