claude-autorouter 0.2.0 → 0.3.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.env.example +42 -14
- package/README.md +39 -18
- package/bin/autorouter.mjs +10 -7
- package/docs/development.md +33 -4
- package/docs/ollama-evaluation.md +83 -62
- package/docs/reference.md +60 -24
- package/docs/releasing.md +19 -13
- package/package.json +4 -3
- package/src/config.mjs +11 -3
- package/src/ollama-evaluator.mjs +64 -35
- package/src/ollama-models.mjs +15 -9
- package/src/ollama-setup.mjs +18 -4
- package/src/onboarding.mjs +20 -9
- package/src/statusline.mjs +30 -6
package/.env.example
CHANGED
|
@@ -22,23 +22,51 @@ AUTOROUTER_JEV_TIMEOUT_MS=1500
|
|
|
22
22
|
AUTOROUTER_TOKEN_COUNT_TIMEOUT_MS=1500
|
|
23
23
|
AUTOROUTER_MIN_CONFIDENCE=0.75
|
|
24
24
|
|
|
25
|
-
# Experimental local evaluator
|
|
26
|
-
#
|
|
27
|
-
#
|
|
28
|
-
#
|
|
29
|
-
#
|
|
30
|
-
#
|
|
31
|
-
#
|
|
25
|
+
# Experimental local evaluator in AutoRouter 0.3.2+: native decision API only.
|
|
26
|
+
# Install/start Ollama 0.35+, then run:
|
|
27
|
+
# claude-autorouter setup --evaluator ollama --pull --force
|
|
28
|
+
# Defaults to nimble:9b-q4_K_M (~5.63 GB download); --force replaces user config.
|
|
29
|
+
# Existing downloads are kept. Replace Qwen config/environment values from 0.2.0.
|
|
30
|
+
# All local models use /v1/systemone; custom tags/aliases must support that API.
|
|
31
|
+
# Add --ollama-model LOCAL_TAG_OR_ALIAS to setup to choose another suitable model.
|
|
32
|
+
# Tev1 alternatives use the same native API; choose one explicitly:
|
|
33
|
+
# claude-autorouter setup --evaluator ollama --ollama-model tev1:0.8b --pull --force
|
|
34
|
+
# claude-autorouter setup --evaluator ollama --ollama-model tev1:4b-q4_K_M --pull --force
|
|
35
|
+
# Downloads: Tev1 0.8B Q8 ~812 MB; Tev1 4B Q4_K_M ~2.7 GB.
|
|
36
|
+
# tev1:latest / tev1:4b select ~4.5 GB Q8; model terms: https://ollama.com/library/tev1
|
|
37
|
+
# Or set AUTOROUTER_EVALUATOR=ollama above and configure an installed model:
|
|
32
38
|
# AUTOROUTER_OLLAMA_URL=http://127.0.0.1:11434
|
|
33
|
-
# AUTOROUTER_OLLAMA_MODEL=
|
|
34
|
-
#
|
|
39
|
+
# AUTOROUTER_OLLAMA_MODEL=nimble:9b-q4_K_M
|
|
40
|
+
# Version 0.3.2 defaults: Tev1 0.8B/custom 1500ms; Tev1 4B 15000ms; Nimble 30000ms.
|
|
41
|
+
# An explicit timeout (including an old saved 1500) always overrides the default.
|
|
42
|
+
# Setup, doctor and startup show the effective model/deadline.
|
|
43
|
+
# Set only when intentionally overriding; Jev's separate deadline is unchanged:
|
|
44
|
+
# AUTOROUTER_OLLAMA_TIMEOUT_MS=30000
|
|
45
|
+
# Set 0 to disable only the runtime evaluator timer:
|
|
46
|
+
# AUTOROUTER_OLLAMA_TIMEOUT_MS=0
|
|
47
|
+
# One launch without changing saved config:
|
|
48
|
+
# AUTOROUTER_OLLAMA_TIMEOUT_MS=0 claude-autorouter claude
|
|
49
|
+
# Persist it for an installed model (the setup flag overrides the environment):
|
|
50
|
+
# claude-autorouter setup --evaluator ollama --ollama-model tev1:4b --ollama-timeout-ms 0 --force
|
|
51
|
+
# Positive values 1..30000 retain a deadline; 2500 means a 2.5-second cutoff.
|
|
52
|
+
# Zero and --ollama-timeout-ms require AutoRouter 0.3.2 or newer.
|
|
35
53
|
# AUTOROUTER_OLLAMA_KEEP_ALIVE=5m
|
|
36
|
-
# The launcher primes the classifier before opening Claude, allowing up to 60s.
|
|
37
|
-
#
|
|
38
|
-
#
|
|
39
|
-
#
|
|
54
|
+
# The launcher primes the classifier before opening Claude, allowing up to 60s even with timeout 0.
|
|
55
|
+
# User/disconnect cancellation remains active; normal evaluator errors still use fallback.
|
|
56
|
+
# Warmup and metadata-only doctor checks do not certify speed or accuracy.
|
|
57
|
+
# Native model context is retained; evaluator state stays capped at 3,000 bytes.
|
|
58
|
+
# Local classification omits Claude executor system instructions; task/history remain.
|
|
59
|
+
# After the idle period, cold reloading may exceed an enabled deadline and use fallback.
|
|
40
60
|
# A longer keep-alive holds the model in memory longer but avoids some reloads.
|
|
41
|
-
#
|
|
61
|
+
# Native entropy confidence is not calibrated accuracy and does not use Jev's threshold.
|
|
62
|
+
# Historical measurements before 0.3.2:
|
|
63
|
+
# On the tested M4, all 12 Nimble tuning requests exceeded the 1,500 ms deadline.
|
|
64
|
+
# Tev1 0.8B: 18/24 labels, 450 ms median; 4B: 22/24, 3.15s at a 10s deadline.
|
|
65
|
+
# Both Tev1 tags have a 2,050-token native context window; see measured limits.
|
|
66
|
+
# Source-only regression: node scripts/test-ollama-routing.mjs --model tev1:0.8b
|
|
67
|
+
# Uses local inference only; no Anthropic/Jev calls, downloads or config writes.
|
|
68
|
+
# A fallback timeout is not a valid Sonnet decision; mismatches fail the regression.
|
|
69
|
+
# See docs/ollama-evaluation.md before allowing slower local classifications.
|
|
42
70
|
|
|
43
71
|
AUTOROUTER_PORT=8787
|
|
44
72
|
# Required only for standalone `serve`; the `claude` launcher generates one.
|
package/README.md
CHANGED
|
@@ -6,13 +6,7 @@ Requires Node.js 22+, macOS or Linux (including WSL), an installed `claude` comm
|
|
|
6
6
|
|
|
7
7
|
## Install and start
|
|
8
8
|
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
```sh
|
|
12
|
-
npm install -g ./claude-autorouter-0.2.0.tgz
|
|
13
|
-
```
|
|
14
|
-
|
|
15
|
-
Once the package is published, install it from the registry with:
|
|
9
|
+
Install from [npm](https://www.npmjs.com/package/claude-autorouter):
|
|
16
10
|
|
|
17
11
|
```sh
|
|
18
12
|
npm install -g claude-autorouter
|
|
@@ -56,31 +50,58 @@ Savings are an **API-equivalent estimate for the same token counts**, using Opus
|
|
|
56
50
|
|
|
57
51
|
## Experimental local evaluator
|
|
58
52
|
|
|
59
|
-
|
|
53
|
+
The local setup below requires AutoRouter 0.3.2 or newer. It uses Ollama's native `/v1/systemone` decision API with `nimble:9b-q4_K_M` by default. Jev remains the default evaluator. If upgrading from 0.2.0, replace the old Qwen model configuration using the [migration steps](docs/reference.md#migrating-an-older-ollama-config).
|
|
54
|
+
|
|
55
|
+
Version 0.3.2 excludes Claude's executor system instructions from the local classifier excerpt, retaining task and conversation excerpts. Runtime deadlines default to 1,500 ms for Tev1 0.8B/custom models, 15,000 ms for official Tev1 4B tags, and 30,000 ms for official Nimble tags. Explicit timeout settings, including a `1500` saved with 0.3.1, still override these defaults. Jev is unchanged.
|
|
56
|
+
|
|
57
|
+
Set `0` to disable AutoRouter's runtime evaluator deadline for one launch using your existing configuration:
|
|
58
|
+
|
|
59
|
+
```sh
|
|
60
|
+
AUTOROUTER_OLLAMA_TIMEOUT_MS=0 claude-autorouter claude
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
To save that setting for an installed Tev1 4B model:
|
|
60
64
|
|
|
61
65
|
```sh
|
|
62
|
-
claude-autorouter setup --evaluator ollama --ollama-
|
|
66
|
+
claude-autorouter setup --evaluator ollama --ollama-model tev1:4b --ollama-timeout-ms 0 --force
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
The setup flag overrides the timeout environment value and saves it. Cancellation and disconnected clients still stop evaluation, normal errors still use fallback, and startup priming keeps its separate 60-second deadline.
|
|
70
|
+
|
|
71
|
+
Install and start Ollama 0.35 or newer; [version 0.35.0](https://github.com/ollama/ollama/releases/tag/v0.35.0) is a prerelease as of September 29, 2026. Then run:
|
|
72
|
+
|
|
73
|
+
```sh
|
|
74
|
+
claude-autorouter setup --evaluator ollama --pull --force
|
|
63
75
|
claude-autorouter doctor
|
|
64
76
|
claude-autorouter claude
|
|
65
77
|
```
|
|
66
78
|
|
|
67
|
-
|
|
79
|
+
`--force` replaces existing AutoRouter configuration. `--pull` downloads the selected model only if missing. Setup does not install or start Ollama, or delete existing models. Select a native decision model explicitly with `--ollama-model`:
|
|
80
|
+
|
|
81
|
+
| Model | Approximate download | Selection |
|
|
82
|
+
| --- | ---: | --- |
|
|
83
|
+
| [Nimble 9B Q4_K_M](https://ollama.com/library/nimble) | 5.63 GB | Default: `nimble:9b-q4_K_M` |
|
|
84
|
+
| [Tev1 0.8B Q8](https://ollama.com/library/tev1) | 812 MB | `tev1:0.8b` |
|
|
85
|
+
| [Tev1 4B Q4_K_M](https://ollama.com/library/tev1) | 2.7 GB | `tev1:4b-q4_K_M` |
|
|
86
|
+
|
|
87
|
+
For example, select Tev1 0.8B with:
|
|
88
|
+
|
|
89
|
+
```sh
|
|
90
|
+
claude-autorouter setup --evaluator ollama --ollama-model tev1:0.8b --pull --force
|
|
91
|
+
```
|
|
68
92
|
|
|
69
|
-
|
|
93
|
+
Use `--ollama-model tev1:4b-q4_K_M` for the listed 4B variant; `tev1:latest` and `tev1:4b` select the larger Q8 download. Model terms are linked in the listings above; download size does not measure resident memory or routing quality. Custom native model tags and aliases also work.
|
|
70
94
|
|
|
71
|
-
|
|
95
|
+
No Jev key is needed for local classification. The launcher primes the evaluator before opening Claude's UI, and evaluation failures fall back to Sonnet or retain Opus without contacting Jev. Claude still answers through Anthropic, with the same routing guards and subscription limits.
|
|
72
96
|
|
|
73
|
-
|
|
74
|
-
| --- | ---: | ---: | ---: |
|
|
75
|
-
| `compact` (`qwen3:1.7b`) | 58.3% | 602 / 834 ms | 1.70 GB |
|
|
76
|
-
| `quality` (`qwen3:4b`) | 91.7% | 889 / 1,242 ms | 3.18 GB |
|
|
97
|
+
Setup, doctor, and startup show the effective model and deadline. Warmup and doctor do not certify classification speed or accuracy. `Ollama fallback: timeout` means no valid decision arrived in time; it is different from the evaluator choosing Sonnet. Source users can run the [local routing regression](docs/development.md#local-routing-regression) to check all three tiers without Claude or Jev calls.
|
|
77
98
|
|
|
78
|
-
|
|
99
|
+
Historical measurements before 0.3.2, on a 16 GiB M4: Tev1 0.8B matched 18/24 held-out labels with 450 ms median latency and no timeouts at 1,500 ms, including full-excerpt checks. Tev1 4B matched 22/24 with a 10-second diagnostic deadline and 3.15-second median latency. Nimble matched 23/24 with a 30-second deadline and 11.4-second median latency. Both larger models exceeded the then-default 1,500 ms. A separate six-case regression with the 0.3.2 fixes passed for both tested Tev1 4B variants and Nimble; Tev1 0.8B matched only three cases. These small tests do not establish general accuracy or Jev parity. See the [measurements and limits](docs/ollama-evaluation.md) and [Ollama reference](docs/reference.md#ollama-evaluator).
|
|
79
100
|
|
|
80
101
|
## Behavior and data
|
|
81
102
|
|
|
82
103
|
- The default client profile permits all three routing tiers. Tool continuations, thinking history, model-specific features, and context size can keep or upgrade a model even when the evaluator chooses a cheaper tier. [Routing policy](docs/reference.md#routing-policy).
|
|
83
|
-
- The selected evaluator receives bounded excerpts that can contain source code
|
|
104
|
+
- The selected evaluator receives bounded excerpts that can contain source code and tool results: TypeSafe with Jev, or the local service with Ollama. Jev also receives system-text excerpts; the local path excludes Claude's executor system instructions. Anthropic receives the complete request. Images, document payloads, and private thinking are omitted from classifier input. [Data flow and authentication](docs/reference.md#data-flow-and-authentication).
|
|
84
105
|
- Subscription access and usage limits still apply. Model switches can reduce cache reuse; cheaper token prices do not guarantee cheaper completed tasks. Run ordinary `claude` to bypass routing.
|
|
85
106
|
- The launcher is quiet by default. Use `AUTOROUTER_DEBUG=1` for metadata diagnostics or `AUTOROUTER_STATUSLINE=0` to retain your existing status line. [Troubleshooting](docs/reference.md#troubleshooting).
|
|
86
107
|
|
package/bin/autorouter.mjs
CHANGED
|
@@ -9,7 +9,7 @@ import { dirname } from 'node:path';
|
|
|
9
9
|
import { createStatusState } from '../src/status-state.mjs';
|
|
10
10
|
import { addStatusLineSettings } from '../src/status-settings.mjs';
|
|
11
11
|
import { loadUserConfig } from '../src/user-config.mjs';
|
|
12
|
-
import { setup, doctor } from '../src/onboarding.mjs';
|
|
12
|
+
import { setup, doctor, ollamaDeadlineText } from '../src/onboarding.mjs';
|
|
13
13
|
import { setupOllama } from '../src/ollama-setup.mjs';
|
|
14
14
|
|
|
15
15
|
const [command = 'help', ...args] = process.argv.slice(2);
|
|
@@ -21,8 +21,8 @@ if (['--version', '-v', 'version'].includes(command)) {
|
|
|
21
21
|
|
|
22
22
|
Usage:
|
|
23
23
|
claude-autorouter setup [--auth-mode subscription|api-key] [--force]
|
|
24
|
-
[--evaluator jev|ollama]
|
|
25
|
-
[--ollama-model MODEL] [--pull]
|
|
24
|
+
[--evaluator jev|ollama]
|
|
25
|
+
[--ollama-model MODEL] [--ollama-timeout-ms N] [--pull]
|
|
26
26
|
claude-autorouter doctor
|
|
27
27
|
claude-autorouter claude [Claude Code arguments]
|
|
28
28
|
claude-autorouter serve
|
|
@@ -35,11 +35,14 @@ AUTOROUTER_CONFIG selects a different file; environment variables take precedenc
|
|
|
35
35
|
Project .env files are never loaded automatically.
|
|
36
36
|
|
|
37
37
|
Jev is the default evaluator and requires TYPESAFE_API_KEY.
|
|
38
|
-
Ollama evaluates locally and requires
|
|
38
|
+
Ollama evaluates locally and requires Ollama 0.35+ with /v1/systemone.
|
|
39
39
|
Use setup --evaluator ollama --pull to detect Ollama and download a missing model.
|
|
40
|
-
|
|
40
|
+
The local default is nimble:9b-q4_K_M; --ollama-model selects another compatible model.
|
|
41
|
+
Smaller Tev1 options: --ollama-model tev1:0.8b or --ollama-model tev1:4b-q4_K_M.
|
|
42
|
+
Local routing deadlines: Tev1 0.8B/custom 1500 ms, Tev1 4B 15000 ms, Nimble 30000 ms.
|
|
43
|
+
Setup --ollama-timeout-ms N saves a routing deadline; use 0 to disable it.
|
|
44
|
+
AUTOROUTER_OLLAMA_TIMEOUT_MS also overrides the deadline; 0 disables it.
|
|
41
45
|
Local routing is experimental; see docs/ollama-evaluation.md for measured limits.
|
|
42
|
-
The auto preset selects using total RAM; compact is the default.
|
|
43
46
|
AUTOROUTER_AUTH_MODE=subscription uses your saved Claude Code login.
|
|
44
47
|
Without setup, AUTOROUTER_AUTH_MODE defaults to api-key and also requires ANTHROPIC_API_KEY.
|
|
45
48
|
AUTOROUTER_CLIENT_PROFILE=compatible (default) enables all three routing tiers.
|
|
@@ -84,7 +87,7 @@ Complete inference requests still go to Anthropic. See README.md.`);
|
|
|
84
87
|
config.localToken = randomBytes(32).toString('hex');
|
|
85
88
|
}
|
|
86
89
|
if (config.evaluator === 'ollama') {
|
|
87
|
-
console.error(`Preparing local Ollama evaluator (${config.ollamaModel})…`);
|
|
90
|
+
console.error(`Preparing local Ollama evaluator (${config.ollamaModel}); ${ollamaDeadlineText(config.ollamaTimeoutMs)}…`);
|
|
88
91
|
try { await setupOllama(config, { pull: false, warm: true, write: () => {} }); }
|
|
89
92
|
catch {
|
|
90
93
|
console.error('Ollama could not be prepared. Requests will use the conservative fallback while it is unavailable; run claude-autorouter doctor.');
|
package/docs/development.md
CHANGED
|
@@ -26,9 +26,37 @@ node --env-file=.env bin/autorouter.mjs claude
|
|
|
26
26
|
|
|
27
27
|
The explicit Node flag loads `.env`; the CLI itself does not auto-load project files. Environment values override the user config. Keep keys out of source control and command arguments.
|
|
28
28
|
|
|
29
|
+
With AutoRouter 0.3.2 or newer, launch an already configured Ollama evaluator without its runtime deadline using:
|
|
30
|
+
|
|
31
|
+
```sh
|
|
32
|
+
AUTOROUTER_OLLAMA_TIMEOUT_MS=0 claude-autorouter claude
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
Persist the setting with `claude-autorouter setup --evaluator ollama --ollama-model tev1:4b --ollama-timeout-ms 0 --force` for that installed model. The setup flag overrides the timeout environment value; later launch-time environment values still override saved configuration. Startup priming keeps its separate 60-second limit, cancellation remains active, and normal errors still use fallback.
|
|
36
|
+
|
|
37
|
+
## Local routing regression
|
|
38
|
+
|
|
39
|
+
The source-only harness below is opt-in and is not included in the npm package. It sends synthetic Claude-shaped requests through the real router and an installed local evaluator, checking task extraction, classifier choices, selected Claude tiers, and new human turns. It makes no Anthropic or Jev calls, downloads no models, and writes no user configuration.
|
|
40
|
+
|
|
41
|
+
```sh
|
|
42
|
+
node scripts/test-ollama-routing.mjs --model tev1:0.8b
|
|
43
|
+
node scripts/test-ollama-routing.mjs --model tev1:4b
|
|
44
|
+
node scripts/test-ollama-routing.mjs --model nimble:9b-q4_K_M
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
Start Ollama 0.35+ and install the selected model first. The harness refuses to run while another model is resident. It never explicitly unloads the selected model; its keep-alive setting controls residency. It warms once, then uses the production deadline for each uncached case: 1,500 ms for Tev1 0.8B/custom tags, 15,000 ms for official Tev1 4B tags, and 30,000 ms for official Nimble tags. Environment settings or `--timeout-ms N` can override the deadline; `--timeout-ms 0` disables the runtime timer while retaining cancellation and the separate warmup limit. This harness does not load saved user configuration. Use `--output artifacts/local-routing.json` to save a metadata report.
|
|
48
|
+
|
|
49
|
+
```sh
|
|
50
|
+
node scripts/test-ollama-routing.mjs --model tev1:4b --timeout-ms 0
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
The command fails on a wrong classification, fallback, unexpected guard override, or missing tier coverage. A Sonnet result with `source: ollama` and `classified_tier: sonnet` is a valid prediction; `source: fallback` and `classifier_error: timeout` means classification did not complete. Passing establishes these synthetic cases only. Warmup and metadata-only `doctor` checks do not establish speed or accuracy on real tasks.
|
|
54
|
+
|
|
55
|
+
Version 0.3.2 excludes Claude's executor system instructions from local classifier input before excerpt budgeting and retains task/history excerpts. Jev is unchanged. Version 0.3.1 included the executor background locally and defaulted every local model to 1,500 ms; see the [upgrade notes](reference.md#migrating-an-older-ollama-config). Historical benchmark results must remain labeled with their original excerpt policy and explicit deadlines.
|
|
56
|
+
|
|
29
57
|
## Live integration tests
|
|
30
58
|
|
|
31
|
-
|
|
59
|
+
The Claude integration tests below make real Claude calls and invoke the configured evaluator, consuming Claude usage and, with Jev, TypeSafe usage. They use temporary synthetic fixtures and disable unrelated customizations and MCP servers. With Ollama, start the local service and install the chosen model first; the harness does not install or download it.
|
|
32
60
|
|
|
33
61
|
```sh
|
|
34
62
|
npm run test:live
|
|
@@ -48,18 +76,19 @@ npm run eval
|
|
|
48
76
|
|
|
49
77
|
The bundled evaluation makes 12 classifier calls and no Claude generations. Jev is the default and incurs TypeSafe usage; set `AUTOROUTER_EVALUATOR=ollama` to evaluate an installed local model. It reports agreement with the starting rubric, fallback count, and p50/p95 routing latency. Edit `test/fixtures/routing.json` to represent the tasks you want to measure. Rubric agreement alone does not establish answer quality or net savings; compare completed tasks against fixed-model baselines.
|
|
50
78
|
|
|
51
|
-
For local evaluator measurements,
|
|
79
|
+
For local evaluator measurements, use Ollama 0.35+ and a model compatible with `/v1/systemone`. Distinguish cold model loading from warmed classification, and record the model tag, hardware, Ollama version, context size, prompt length, and resident memory. The launcher primes the classifier with a synthetic task before opening the UI, with a separate deadline of up to 60 seconds. Runtime and benchmark share the 3,000-character/3,000-UTF-8-byte state limit, so include non-ASCII cases and excerpts that fill the budget. Also measure the first request after keep-alive expiration: its reload can hit the normal deadline even when warm requests pass. Repeat on realistic prompt distributions instead of selecting a model from a single easy request. Disk download size is not resident RAM. Keep model downloads opt-in and respect each model's license.
|
|
52
80
|
|
|
53
81
|
The dedicated Ollama benchmark uses synthetic tuning/held-out fixtures and reports cold latency separately from repeated warm requests:
|
|
54
82
|
|
|
55
83
|
```sh
|
|
56
|
-
npm run eval:ollama -- --models
|
|
84
|
+
npm run eval:ollama -- --models nimble:9b-q4_K_M --split heldout --rounds 3 --stress-rounds 8
|
|
85
|
+
npm run eval:ollama -- --models tev1:0.8b,tev1:4b-q4_K_M --split heldout --rounds 1 --stress-rounds 8
|
|
57
86
|
node scripts/evaluate-ollama.mjs --help
|
|
58
87
|
```
|
|
59
88
|
|
|
60
89
|
Install each selected model first and use an idle Ollama instance with no resident models. The benchmark loads one candidate at a time and unloads it afterward. It does not download models or contact Claude or Jev. It reports classification errors and under/over-routing as well as latency; fixture labels are subjective rubric judgments, not measurements of completed task quality.
|
|
61
90
|
|
|
62
|
-
|
|
91
|
+
See the [local evaluator measurements](ollama-evaluation.md) for the hardware, results, and limits. The older Qwen chat adapter and its compact/quality/auto presets have been removed. Every local candidate now uses native choice scoring; context allocation comes from the model/server configuration. Native entropy confidence is not calibrated accuracy and is not used as Jev's confidence threshold. A successful launcher/Enterprise integration request establishes connectivity and model routing, not classifier accuracy or parity with Jev.
|
|
63
92
|
|
|
64
93
|
## Startup context diagnostics
|
|
65
94
|
|
|
@@ -1,100 +1,121 @@
|
|
|
1
|
-
# Local evaluator measurements
|
|
1
|
+
# Local decision evaluator measurements
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
AutoRouter supports `/v1/systemone` classification only: remote Jev with an API key, or local Ollama 0.35+ with a compatible decision model. Jev remains the default evaluator. The local default is `nimble:9b-q4_K_M`; the old Qwen chat adapter and compact/quality/auto presets have been removed. Existing downloaded models are not deleted.
|
|
4
4
|
|
|
5
|
-
##
|
|
5
|
+
## Version 0.3.2 routing regression (September 30, 2026)
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
The npm `0.3.1` runtime used a 1,500 ms deadline for every local model, even though startup priming allows 60 seconds. On the user's already-loaded `tev1:4b` (Q8), four synthetic tasks all timed out and selected Sonnet through `source=fallback`, `classifier_error=timeout`. With a separate 30-second diagnostic allowance, the same tasks selected Haiku, Haiku, Sonnet, and Opus. Three of those decisions took 5.6–6.1 seconds. This was a deadline failure, not a parser forcing every answer to Sonnet.
|
|
8
8
|
|
|
9
|
-
|
|
10
|
-
| --- | ---: | --- |
|
|
11
|
-
| `qwen3.5:0.8b` | Approximately 1.0 GB | Smallest candidate |
|
|
12
|
-
| `qwen3:1.7b` | Approximately 1.4 GB | Compact speed baseline |
|
|
13
|
-
| `qwen3.5:2b` | Approximately 2.7 GB | Additional compact candidate |
|
|
14
|
-
| `qwen3.5:4b` | Approximately 3.4 GB | Measured larger candidate; rejected for preset latency |
|
|
15
|
-
| `qwen3:4b` | Approximately 2.5 GB | Larger candidate selected for the `quality` option |
|
|
9
|
+
We captured three synthetic request shapes from the installed Claude client using an isolated local response stub, without paid provider calls. The human task survived intact and no compatibility guard forced Sonnet. The local evaluator nevertheless received 627–796 characters of Claude's general executor instructions. Tev1 0.8B classified all three captures as Sonnet. Removing only this system background changed the distributed-fencing case to Opus; the `[].length` example still selected Sonnet. Removing additional state fields did not improve that result, so the routing rubric and remaining state layout were retained.
|
|
16
10
|
|
|
17
|
-
|
|
11
|
+
Version 0.3.2 excludes executor system instructions from local classifier input, uses 15-second defaults for official Tev1 4B variants and 30 seconds for Nimble, and retains the 1.5-second default for Tev1 0.8B/custom models. Explicit timeout settings still win; `0` disables only the runtime evaluator timer. Setup accepts `--ollama-timeout-ms`; setup, doctor, and startup display the effective deadline. Compact status lines retain the fallback cause. Jev's input extraction and deadline are unchanged.
|
|
18
12
|
|
|
19
|
-
The `
|
|
13
|
+
The source-only, opt-in `npm run test:ollama -- --model TAG` checks actual `Router.route` results against six fixed synthetic Claude-shaped requests: literal output, `[].length`, a bounded feature, distributed fencing, a new mechanical task after a difficult task, and a new difficult task after a simple task. It checks the evaluator choice, selected Claude model, source, and reason, and fails on a mismatch, fallback, override, or missing tier coverage. Live cases use fresh routers to avoid decision-cache passes; persistent turn-cache transitions and compatibility guards are separately covered by offline regression tests. It is not included in the npm package, downloads nothing, and does not contact Claude or Jev.
|
|
20
14
|
|
|
21
|
-
|
|
15
|
+
One pass with the fixes on the same M4/16 GiB Mac produced:
|
|
22
16
|
|
|
23
|
-
|
|
17
|
+
| Installed model | Runtime deadline | Matching cases | Fallbacks | Decision latency range |
|
|
18
|
+
| --- | ---: | ---: | ---: | ---: |
|
|
19
|
+
| `tev1:0.8b` | 1,500 ms | 3 / 6 | 0 | 48–479 ms |
|
|
20
|
+
| `tev1:4b-q4_K_M` | 15,000 ms | 6 / 6 | 0 | 2.40–2.75 s |
|
|
21
|
+
| `tev1:4b` (Q8) | 15,000 ms | 6 / 6 | 0 | 2.38–2.70 s |
|
|
22
|
+
| `nimble:9b-q4_K_M` | 30,000 ms | 6 / 6 | 0 | 4.74–7.29 s |
|
|
24
23
|
|
|
25
|
-
The
|
|
24
|
+
Every model reached all three tiers. The 0.8B failures were genuine classifications: `[].length` → Sonnet, new mechanical task → Opus, new difficult task → Sonnet. Its live suite therefore **fails**, rather than treating valid native responses as proof of accuracy. These six cases are regression checks, not a new held-out quality benchmark; the samples, machine load, residency, and prompt-cache effects do not establish a quantization speed comparison or a worst-case latency bound. The larger-model budgets allow slower decisions; they do not make those models fast or prevent every timeout. No user configuration or model files were changed; the originally resident `tev1:4b` was restored after sequential model testing.
|
|
26
25
|
|
|
27
|
-
A
|
|
26
|
+
A separate check with `tev1:4b` and the runtime deadline disabled (`timeout_ms: 0`) passed the same six cases without fallback, with decision latencies of 4.37–5.42 seconds after 13.11 seconds of priming. It made no Claude or Jev calls and downloaded nothing. This validates disabled-timer routing on those cases; it adds no new quality cases or latency guarantee.
|
|
28
27
|
|
|
29
|
-
|
|
28
|
+
Fixture SHA-256: `88e12f962fb42b36a00edd6b8c2560e54620d90c978b1ff13a1ddf31ee46b585`. The questions remain unchanged at `be151cedb4de4b7ef3f7162d751f70ce7d9dd14efc66fae1835f73ffd04027be`.
|
|
30
29
|
|
|
31
|
-
|
|
30
|
+
## Historical benchmark method (before 0.3.2)
|
|
32
31
|
|
|
33
|
-
|
|
34
|
-
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
|
|
35
|
-
| `qwen3:1.7b` | 20 / 24 (83.3%) | 0 / 24 | 245 / 280 ms | 2.50 s | 1.36 GB | 1.70 GB |
|
|
36
|
-
| `qwen3:4b` | 24 / 24 (100%) | 0 / 24 | 743 / 889 ms | 9.62 s | 2.50 GB | 3.18 GB |
|
|
37
|
-
| `qwen3.5:0.8b` | 10 / 24 (41.7%) | 0 / 24 | 409 / 437 ms | 2.78 s | 1.04 GB | 1.09 GB |
|
|
38
|
-
| `qwen3.5:2b` | 18 / 24 (75.0%) | 0 / 24 | 918 / 1,009 ms | 6.49 s | 2.74 GB | 2.36 GB |
|
|
39
|
-
| `qwen3.5:4b` | No valid warm results | 24 / 24 | Deadline reached | 7.26 s | 3.39 GB | 3.14 GB |
|
|
32
|
+
Measurements were recorded on September 29, 2026, on an Apple M4 Mac with 16 GiB of unified memory alongside other applications. Nimble ran on an isolated Ollama 0.35.0 process; the installed Ollama 0.33.3 daemon was left unchanged during that test. Tev1 was tested later on the user's upgraded Ollama 0.35.0 service after its downloads finished. Version 0.35.0 was a prerelease at the time. These are observations under different application loads, not a controlled hardware comparison. See the [official release](https://github.com/ollama/ollama/releases/tag/v0.35.0), [Nimble catalog](https://ollama.com/library/nimble), and [Tev1 catalog](https://ollama.com/library/tev1).
|
|
40
33
|
|
|
41
|
-
The
|
|
34
|
+
The selected Nimble tag contains a 9B Q4_K_M model, approximately 5.63 GB to download, with an 8,194-token native context setting. `/v1/systemone` scores choices directly; it accepts no chat generation options or per-request context override. The model/server configuration determines allocation. The production excerpt remains bounded to 3,000 serialized characters and UTF-8 bytes, including for non-ASCII text. Native confidence measures the concentration of the choice distribution, not calibrated accuracy, and is not used as Jev's confidence threshold.
|
|
42
35
|
|
|
43
|
-
|
|
36
|
+
The fixture contains 36 balanced synthetic workloads: 12 tuning cases and 24 held-out cases, with equal numbers of Haiku, Sonnet, and Opus labels. Cases include mechanical edits, ordinary implementation, difficult correctness and security work, topic changes, tool results, short follow-ups, and misleading routing instructions. These labels are judgments under the routing policy, not proof of which Claude model would complete each task successfully. The native questions preserve the existing local routing policy and were frozen before Nimble testing; no changes were made from held-out results.
|
|
44
37
|
|
|
45
|
-
|
|
38
|
+
The historical harness used the then-current production state builder and evaluator, including local metadata checks in wall-clock latency. It bypasses AutoRouter's decision cache; Ollama's own caching remains enabled. A cold measurement starts with the model unloaded, but operating-system file caches and kernels may already be warm. Cold calls have a separate 60-second deadline. The normal evaluator deadline in that benchmark was 1,500 ms for every model. Reported model allocation comes from `/api/ps`; it is not a measurement of total process or system memory, and its GPU allocation is not additional independent RAM on this unified-memory Mac. Aggregate runtime RSS includes all Ollama and llama-server processes, including the original idle daemon.
|
|
46
39
|
|
|
47
|
-
|
|
40
|
+
## Historical Nimble results
|
|
48
41
|
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
42
|
+
The first production-deadline run returned Haiku correctly for its cold mechanical task in 17.64 seconds. **All 12 warm tuning requests timed out at 1,500 ms**, with cancellation p50/p95 of 1,503/1,513 ms. No warm classification accuracy can be inferred from that run. The model's reported allocation was 5.48 GB, with an 8,194-token context. This did not meet the desired fast-routing target on the tested Mac.
|
|
43
|
+
|
|
44
|
+
A separate first-load smoke call correctly classified `[].length` as Haiku in 25.23 seconds. It checks integration, not warm performance or classifier accuracy.
|
|
45
|
+
|
|
46
|
+
The separate held-out diagnostic used a 30,000 ms deadline and one pass over 24 distinct workloads. It did not change the then-default 1,500 ms budget or establish performance within it.
|
|
47
|
+
|
|
48
|
+
| Measurement | Result |
|
|
49
|
+
| --- | ---: |
|
|
50
|
+
| Held-out rubric agreement | 23 / 24 (95.8%) |
|
|
51
|
+
| Valid decisions / timeouts | 24 / 0 |
|
|
52
|
+
| Warm p50 / p95 wall time | 11,432 / 16,334 ms |
|
|
53
|
+
| Minimum / maximum wall time | 697 / 20,201 ms |
|
|
54
|
+
| Cold wall time | 15.47 s |
|
|
55
|
+
| Reported model allocation | 5.48 GB |
|
|
56
|
+
| Under-routes / over-routes | 1 / 0 |
|
|
53
57
|
|
|
54
|
-
The
|
|
58
|
+
All eight Haiku and eight Sonnet labels matched. Seven of eight Opus labels matched. The `held-o-injection` case, a difficult deadlock investigation containing an instruction to select Haiku, was incorrectly routed to Haiku. This is a concrete limitation against misleading routing instructions. These 24 decisions are a single pass, not repeated measurements or a claim of comparable performance to Jev.
|
|
55
59
|
|
|
56
|
-
|
|
57
|
-
| --- | ---: | ---: | ---: |
|
|
58
|
-
| Haiku | 12 | 6 | 6 |
|
|
59
|
-
| Sonnet | 0 | 24 | 0 |
|
|
60
|
-
| Opus | 0 | 18 | 6 |
|
|
60
|
+
All eight full-excerpt diagnostic requests completed at the 30,000 ms deadline and returned the expected Haiku tier. Their p50/p95 latency was 25,711/27,271 ms. These synthetic requests fill the 3,000-byte state budget with clearly mechanical tasks and unrelated tool output. They are separate performance checks, not eight additional held-out workloads. Longer excerpts can consume almost the entire diagnostic deadline on this machine.
|
|
61
61
|
|
|
62
|
-
The
|
|
62
|
+
An isolated live Claude Code test also passed using the saved Enterprise subscription login, a temporary configuration with a 30,000 ms evaluator deadline, and no Jev key. Nimble selected Haiku in 12.69 seconds; Anthropic returned HTTP 200 with the expected literal response and confirmed `claude-haiku-4-5-20251001`. Only a synthetic prompt was used, with no repository files or tools. This verifies the authentication and routing integration, not general classifier accuracy. The installed Ollama service and user configuration were left unchanged.
|
|
63
63
|
|
|
64
|
-
|
|
65
|
-
| --- | ---: | ---: | ---: |
|
|
66
|
-
| Haiku | 24 | 0 | 0 |
|
|
67
|
-
| Sonnet | 0 | 18 | 6 |
|
|
68
|
-
| Opus | 0 | 0 | 24 |
|
|
64
|
+
## Historical Tev1 results
|
|
69
65
|
|
|
70
|
-
|
|
66
|
+
Both [Tev1 variants](https://ollama.com/library/tev1) use the same production adapter and frozen questions, selected through `--ollama-model`. The 0.8B tag uses Q8_0 quantization and downloads approximately 812 MB; `tev1:4b-q4_K_M` downloads approximately 2.71 GB. The unqualified `tev1` tag selects the larger 4B Q8 model, which was not tested in this historical benchmark. It was tested separately in the 0.3.2 regression above. Jev remains the evaluator default and Nimble remains the local-model default.
|
|
71
67
|
|
|
72
|
-
|
|
68
|
+
Each measured tag ships `num_ctx:2050`. The window includes the template, routing criteria, and excerpt. The 3,000-byte state cap is not a guarantee that every possible input fits this smaller token window. The reported stress cases used 1,664–1,752 input tokens on 0.8B. Requests exceeding model limits use the usual fallback; AutoRouter does not switch protocols or silently truncate additional content for Tev1.
|
|
69
|
+
|
|
70
|
+
These measurements use one pass over the same 24 held-out cases and eight separate full-excerpt cases, after both downloads finished. An exploratory 0.8B run during the 4B download gave the same labels; its timings are excluded here. Neither questions nor expected labels were changed in response to Tev1 outputs.
|
|
71
|
+
|
|
72
|
+
| Model | Deadline | Held-out agreement | Timeouts | Warm p50 / p95 | Reported allocation |
|
|
73
|
+
| --- | ---: | ---: | ---: | ---: | ---: |
|
|
74
|
+
| `tev1:0.8b` | 1,500 ms | 18 / 24 (75.0%) | 0 / 24 | 450 / 488 ms | 0.89 GB |
|
|
75
|
+
| `tev1:4b-q4_K_M` | 1,500 ms | No valid warm decisions | 24 / 24 | Deadline reached | 2.91 GB |
|
|
76
|
+
| `tev1:4b-q4_K_M` diagnostic | 10,000 ms | 22 / 24 (91.7%) | 0 / 24 | 3,149 / 4,169 ms | 2.91 GB |
|
|
73
77
|
|
|
74
|
-
|
|
78
|
+
The 0.8B model matched four of eight Haiku labels, all eight Sonnet labels, and six of eight Opus labels. It over-routed four mechanical tasks and under-routed two difficult tasks to Sonnet, including a case with a misleading tier instruction. Cold wall time was 2.08 seconds for 0.8B and 6.04 seconds for 4B. Model size and fast responses do not establish sufficient accuracy for an engineering workload.
|
|
75
79
|
|
|
76
|
-
|
|
77
|
-
| --- | ---: | ---: |
|
|
78
|
-
| `qwen3:1.7b` | 8 / 8 | 1,502 / 1,503 ms |
|
|
79
|
-
| `qwen3:4b` | 8 / 8 | 1,502 / 1,505 ms |
|
|
80
|
+
With a separate 10-second deadline, 4B matched seven of eight Haiku labels, all eight Sonnet labels, and seven of eight Opus labels. It over-routed one mechanical case to Opus and under-routed one difficult case to Sonnet. Its cold diagnostic request took 3.69 seconds. The improved agreement comes with several seconds of classification latency; it is not performance at the then-default 1,500 ms deadline.
|
|
80
81
|
|
|
81
|
-
|
|
82
|
+
All eight 0.8B full-excerpt requests completed within 1,500 ms and returned Haiku, with p50/p95 of 1,079/1,150 ms. The 4B model timed out on all eight at 1,500 ms and again on all eight at 10,000 ms. Its 10-second cancellation p50/p95 was 10,006/10,081 ms; completed full-excerpt latency was not measured. The longer deadline therefore allows the reported short held-out decisions but does not guarantee completion for full excerpts.
|
|
82
83
|
|
|
83
|
-
|
|
84
|
+
The native API is the supported Ollama integration. Together's [publisher interface](https://huggingface.co/togethercomputer/Tev1-4B-experimental#intended-interface) describes a different training prompt layout from the schema rendered by [Ollama 0.35's compiler](https://github.com/ollama/ollama/blob/v0.35.0/decision/systemone.go). These results measure the actual Ollama native path, not a reproduction of Together's training-format evaluation or its published accuracy figures.
|
|
85
|
+
|
|
86
|
+
Both tags passed isolated live Claude Enterprise integration checks on the user's Ollama 0.35 service, with no Jev key and no repository tools or files. Claude returned the correct `0` for a synthetic `[].length` query and Anthropic returned HTTP 200. The 0.8B evaluator took 584 ms under the default deadline but selected Sonnet, over-routing this mechanical task. The 4B evaluator selected Haiku in 3.78 seconds using a temporary 10-second deadline. These checks verify setup, authentication, native evaluation, and generation; they do not imply that every routing decision is correct. Temporary configurations were removed, tested models were unloaded from memory, and the user's Ollama service and model files were retained.
|
|
87
|
+
|
|
88
|
+
Model digests:
|
|
89
|
+
|
|
90
|
+
- `tev1:0.8b`: `c0099a86fcbd81bc5876a0d1f94998d2b038f7f1b5f7329a3dba43a36903c652`
|
|
91
|
+
- `tev1:4b-q4_K_M`: `3509ac7180e86e5fa8efc7b5745e32d55dd9d4e0a86bc9a88aba5323a5d29bc6`
|
|
84
92
|
|
|
85
93
|
## Reproducing the evaluation
|
|
86
94
|
|
|
87
|
-
Use a source checkout; benchmark scripts and fixtures are
|
|
95
|
+
Use a source checkout; benchmark scripts and fixtures are not included in the npm package. The commands below evaluate the checked-out version. To reproduce the historical excerpt policy, use the `v0.3.1` checkout; 0.3.2 changes local input extraction. The explicit deadlines retain the historical budgets. Install Ollama 0.35+, start it, and explicitly download the model:
|
|
88
96
|
|
|
89
97
|
```sh
|
|
90
|
-
|
|
91
|
-
node scripts/evaluate-ollama.mjs --models
|
|
98
|
+
ollama pull nimble:9b-q4_K_M
|
|
99
|
+
node scripts/evaluate-ollama.mjs --models nimble:9b-q4_K_M --split tuning --rounds 1 --timeout-ms 1500 --output artifacts/nimble-tuning.json
|
|
100
|
+
node scripts/evaluate-ollama.mjs --models nimble:9b-q4_K_M --split heldout --rounds 1 --stress-rounds 8 --timeout-ms 30000 --output artifacts/nimble-diagnostic.json
|
|
101
|
+
ollama pull tev1:0.8b
|
|
102
|
+
ollama pull tev1:4b-q4_K_M
|
|
103
|
+
node scripts/evaluate-ollama.mjs --models tev1:0.8b,tev1:4b-q4_K_M --split heldout --rounds 1 --stress-rounds 8 --timeout-ms 1500 --output artifacts/tev1-default.json
|
|
104
|
+
node scripts/evaluate-ollama.mjs --models tev1:4b-q4_K_M --split heldout --rounds 1 --stress-rounds 8 --timeout-ms 10000 --output artifacts/tev1-diagnostic.json
|
|
92
105
|
```
|
|
93
106
|
|
|
94
|
-
|
|
107
|
+
Use `--endpoint http://127.0.0.1:PORT` for another local instance. The benchmark requires no resident models at startup, never downloads or deletes models, and unloads each tested model afterward. It sends only checked-in synthetic cases and does not contact Claude or Jev. Reports include model identity, protocol, fixture/question hashes, token counts, per-case results, timeouts, confusion matrices, latency, and reported allocation. The eight separate stress requests fill the excerpt budget and vary an early nonce to prevent reuse of the previous full state; they are performance checks, not held-out accuracy cases.
|
|
108
|
+
|
|
109
|
+
Reproducibility identifiers:
|
|
110
|
+
|
|
111
|
+
- Model digest: `3776806da5587387a996e28e75d5d07fbebe7410d47879e1f71782f36c896ce3`
|
|
112
|
+
- Fixture SHA-256: `1ef5111a6a36f0f4bc8d111d54c6985cac4c4a8b16357023a0958d285f7438ad`
|
|
113
|
+
- Native questions SHA-256: `be151cedb4de4b7ef3f7162d751f70ce7d9dd14efc66fae1835f73ffd04027be`
|
|
114
|
+
|
|
115
|
+
## Limits and previous measurements
|
|
95
116
|
|
|
96
|
-
|
|
117
|
+
This is a small synthetic rubric-agreement benchmark, not a downstream task-quality, savings, Jev-parity, or security evaluation. Classification can miss context outside the excerpt. The cases do not establish robust resistance to prompt injection. Performance depends on hardware, memory pressure, prompt length, and residency; a successful setup or simple request does not guarantee the runtime deadline.
|
|
97
118
|
|
|
98
|
-
|
|
119
|
+
The launcher primes the model before opening Claude, allowing up to 60 seconds for that synthetic classification. Idle unloading can still make later requests cold. Runtime timeouts and invalid responses use the existing conservative fallback, without contacting Jev. A longer `AUTOROUTER_OLLAMA_TIMEOUT_MS` trades added prompt latency for more completed local classifications; it does not make the evaluator faster. In 0.3.2, `0` disables the runtime timer while preserving user/disconnect cancellation, normal error fallback, and the separate startup limit.
|
|
99
120
|
|
|
100
|
-
|
|
121
|
+
Earlier Qwen results used a different `/api/chat` implementation and are not measurements of this native backend. They remain available in the [historical evaluation document](https://github.com/frapposelli/claude-autorouter/blob/548a175/docs/ollama-evaluation.md). Reproduce those results from that revision, not the current native-only harness.
|