claude-autorouter 0.2.0 → 0.3.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/.env.example CHANGED
@@ -22,23 +22,51 @@ AUTOROUTER_JEV_TIMEOUT_MS=1500
22
22
  AUTOROUTER_TOKEN_COUNT_TIMEOUT_MS=1500
23
23
  AUTOROUTER_MIN_CONFIDENCE=0.75
24
24
 
25
- # Experimental local evaluator: install/start Ollama, then configure with:
26
- # claude-autorouter setup --evaluator ollama --ollama-preset compact --pull
27
- # compact uses less memory; quality (qwen3:4b) had better synthetic rubric agreement.
28
- # Both were tested on a 16 GiB M4. Choose quality explicitly when memory permits.
29
- # auto selects quality above 24 GiB total RAM; this is only a headroom heuristic.
30
- # Add --force to replace an existing user config. Environment values win over it.
31
- # Or set AUTOROUTER_EVALUATOR=ollama above and configure an installed local model:
25
+ # Experimental local evaluator in AutoRouter 0.3.2+: native decision API only.
26
+ # Install/start Ollama 0.35+, then run:
27
+ # claude-autorouter setup --evaluator ollama --pull --force
28
+ # Defaults to nimble:9b-q4_K_M (~5.63 GB download); --force replaces user config.
29
+ # Existing downloads are kept. Replace Qwen config/environment values from 0.2.0.
30
+ # All local models use /v1/systemone; custom tags/aliases must support that API.
31
+ # Add --ollama-model LOCAL_TAG_OR_ALIAS to setup to choose another suitable model.
32
+ # Tev1 alternatives use the same native API; choose one explicitly:
33
+ # claude-autorouter setup --evaluator ollama --ollama-model tev1:0.8b --pull --force
34
+ # claude-autorouter setup --evaluator ollama --ollama-model tev1:4b-q4_K_M --pull --force
35
+ # Downloads: Tev1 0.8B Q8 ~812 MB; Tev1 4B Q4_K_M ~2.7 GB.
36
+ # tev1:latest / tev1:4b select ~4.5 GB Q8; model terms: https://ollama.com/library/tev1
37
+ # Or set AUTOROUTER_EVALUATOR=ollama above and configure an installed model:
32
38
  # AUTOROUTER_OLLAMA_URL=http://127.0.0.1:11434
33
- # AUTOROUTER_OLLAMA_MODEL=qwen3:1.7b
34
- # AUTOROUTER_OLLAMA_TIMEOUT_MS=1500
39
+ # AUTOROUTER_OLLAMA_MODEL=nimble:9b-q4_K_M
40
+ # Version 0.3.2 defaults: Tev1 0.8B/custom 1500ms; Tev1 4B 15000ms; Nimble 30000ms.
41
+ # An explicit timeout (including an old saved 1500) always overrides the default.
42
+ # Setup, doctor and startup show the effective model/deadline.
43
+ # Set only when intentionally overriding; Jev's separate deadline is unchanged:
44
+ # AUTOROUTER_OLLAMA_TIMEOUT_MS=30000
45
+ # Set 0 to disable only the runtime evaluator timer:
46
+ # AUTOROUTER_OLLAMA_TIMEOUT_MS=0
47
+ # One launch without changing saved config:
48
+ # AUTOROUTER_OLLAMA_TIMEOUT_MS=0 claude-autorouter claude
49
+ # Persist it for an installed model (the setup flag overrides the environment):
50
+ # claude-autorouter setup --evaluator ollama --ollama-model tev1:4b --ollama-timeout-ms 0 --force
51
+ # Positive values 1..30000 retain a deadline; 2500 means a 2.5-second cutoff.
52
+ # Zero and --ollama-timeout-ms require AutoRouter 0.3.2 or newer.
35
53
  # AUTOROUTER_OLLAMA_KEEP_ALIVE=5m
36
- # The launcher primes the classifier before opening Claude, allowing up to 60s.
37
- # Both presets timed out on full-excerpt stress tests at the runtime deadline.
38
- # See docs/ollama-evaluation.md for measured latency, memory, and accuracy limits.
39
- # After the idle period, cold reloading may exceed the deadline and use fallback.
54
+ # The launcher primes the classifier before opening Claude, allowing up to 60s even with timeout 0.
55
+ # User/disconnect cancellation remains active; normal evaluator errors still use fallback.
56
+ # Warmup and metadata-only doctor checks do not certify speed or accuracy.
57
+ # Native model context is retained; evaluator state stays capped at 3,000 bytes.
58
+ # Local classification omits Claude executor system instructions; task/history remain.
59
+ # After the idle period, cold reloading may exceed an enabled deadline and use fallback.
40
60
  # A longer keep-alive holds the model in memory longer but avoids some reloads.
41
- # The confidence threshold above applies only to Jev. Ollama returns a tier.
61
+ # Native entropy confidence is not calibrated accuracy and does not use Jev's threshold.
62
+ # Historical measurements before 0.3.2:
63
+ # On the tested M4, all 12 Nimble tuning requests exceeded the 1,500 ms deadline.
64
+ # Tev1 0.8B: 18/24 labels, 450 ms median; 4B: 22/24, 3.15s at a 10s deadline.
65
+ # Both Tev1 tags have a 2,050-token native context window; see measured limits.
66
+ # Source-only regression: node scripts/test-ollama-routing.mjs --model tev1:0.8b
67
+ # Uses local inference only; no Anthropic/Jev calls, downloads or config writes.
68
+ # A fallback timeout is not a valid Sonnet decision; mismatches fail the regression.
69
+ # See docs/ollama-evaluation.md before allowing slower local classifications.
42
70
 
43
71
  AUTOROUTER_PORT=8787
44
72
  # Required only for standalone `serve`; the `claude` launcher generates one.
package/README.md CHANGED
@@ -6,13 +6,7 @@ Requires Node.js 22+, macOS or Linux (including WSL), an installed `claude` comm
6
6
 
7
7
  ## Install and start
8
8
 
9
- **npm publication is pending.** Install the release tarball now:
10
-
11
- ```sh
12
- npm install -g ./claude-autorouter-0.2.0.tgz
13
- ```
14
-
15
- Once the package is published, install it from the registry with:
9
+ Install from [npm](https://www.npmjs.com/package/claude-autorouter):
16
10
 
17
11
  ```sh
18
12
  npm install -g claude-autorouter
@@ -56,31 +50,58 @@ Savings are an **API-equivalent estimate for the same token counts**, using Opus
56
50
 
57
51
  ## Experimental local evaluator
58
52
 
59
- [Install and start Ollama](https://docs.ollama.com/quickstart), then select a preset. The default `compact` uses less memory; `quality` showed better agreement with the routing rubric:
53
+ The local setup below requires AutoRouter 0.3.2 or newer. It uses Ollama's native `/v1/systemone` decision API with `nimble:9b-q4_K_M` by default. Jev remains the default evaluator. If upgrading from 0.2.0, replace the old Qwen model configuration using the [migration steps](docs/reference.md#migrating-an-older-ollama-config).
54
+
55
+ Version 0.3.2 excludes Claude's executor system instructions from the local classifier excerpt, retaining task and conversation excerpts. Runtime deadlines default to 1,500 ms for Tev1 0.8B/custom models, 15,000 ms for official Tev1 4B tags, and 30,000 ms for official Nimble tags. Explicit timeout settings, including a `1500` saved with 0.3.1, still override these defaults. Jev is unchanged.
56
+
57
+ Set `0` to disable AutoRouter's runtime evaluator deadline for one launch using your existing configuration:
58
+
59
+ ```sh
60
+ AUTOROUTER_OLLAMA_TIMEOUT_MS=0 claude-autorouter claude
61
+ ```
62
+
63
+ To save that setting for an installed Tev1 4B model:
60
64
 
61
65
  ```sh
62
- claude-autorouter setup --evaluator ollama --ollama-preset compact --pull
66
+ claude-autorouter setup --evaluator ollama --ollama-model tev1:4b --ollama-timeout-ms 0 --force
67
+ ```
68
+
69
+ The setup flag overrides the timeout environment value and saves it. Cancellation and disconnected clients still stop evaluation, normal errors still use fallback, and startup priming keeps its separate 60-second deadline.
70
+
71
+ Install and start Ollama 0.35 or newer; [version 0.35.0](https://github.com/ollama/ollama/releases/tag/v0.35.0) is a prerelease as of September 29, 2026. Then run:
72
+
73
+ ```sh
74
+ claude-autorouter setup --evaluator ollama --pull --force
63
75
  claude-autorouter doctor
64
76
  claude-autorouter claude
65
77
  ```
66
78
 
67
- Add `--force` when replacing an existing config. Setup downloads a missing selected model only with `--pull`; it does not install or start Ollama. No Jev key is needed. Claude still answers through Anthropic, with the same routing guards and subscription limits.
79
+ `--force` replaces existing AutoRouter configuration. `--pull` downloads the selected model only if missing. Setup does not install or start Ollama, or delete existing models. Select a native decision model explicitly with `--ollama-model`:
80
+
81
+ | Model | Approximate download | Selection |
82
+ | --- | ---: | --- |
83
+ | [Nimble 9B Q4_K_M](https://ollama.com/library/nimble) | 5.63 GB | Default: `nimble:9b-q4_K_M` |
84
+ | [Tev1 0.8B Q8](https://ollama.com/library/tev1) | 812 MB | `tev1:0.8b` |
85
+ | [Tev1 4B Q4_K_M](https://ollama.com/library/tev1) | 2.7 GB | `tev1:4b-q4_K_M` |
86
+
87
+ For example, select Tev1 0.8B with:
88
+
89
+ ```sh
90
+ claude-autorouter setup --evaluator ollama --ollama-model tev1:0.8b --pull --force
91
+ ```
68
92
 
69
- The launcher primes the local classifier before opening Claude's UI. Failed evaluations fall back to Sonnet or retain Opus without contacting Jev. See the [Ollama reference](docs/reference.md#ollama-evaluator) for setup options.
93
+ Use `--ollama-model tev1:4b-q4_K_M` for the listed 4B variant; `tev1:latest` and `tev1:4b` select the larger Q8 download. Model terms are linked in the listings above; download size does not measure resident memory or routing quality. Custom native model tags and aliases also work.
70
94
 
71
- Both presets were tested on a 16 GiB M4 Mac using 24 distinct held-out synthetic workloads repeated three times:
95
+ No Jev key is needed for local classification. The launcher primes the evaluator before opening Claude's UI, and evaluation failures fall back to Sonnet or retain Opus without contacting Jev. Claude still answers through Anthropic, with the same routing guards and subscription limits.
72
96
 
73
- | Preset | Rubric agreement | Warm p50 / p95 | Model allocation |
74
- | --- | ---: | ---: | ---: |
75
- | `compact` (`qwen3:1.7b`) | 58.3% | 602 / 834 ms | 1.70 GB |
76
- | `quality` (`qwen3:4b`) | 91.7% | 889 / 1,242 ms | 3.18 GB |
97
+ Setup, doctor, and startup show the effective model and deadline. Warmup and doctor do not certify classification speed or accuracy. `Ollama fallback: timeout` means no valid decision arrived in time; it is different from the evaluator choosing Sonnet. Source users can run the [local routing regression](docs/development.md#local-routing-regression) to check all three tiers without Claude or Jev calls.
77
98
 
78
- Use `--ollama-preset quality` to choose the larger model when memory permits, including on a 16 GiB machine. Each model timed out on all eight full-excerpt stress requests at the 1,500 ms deadline. These results do not establish parity with Jev or completed-task quality. Read the [measurements and limits](docs/ollama-evaluation.md) before choosing local classification.
99
+ Historical measurements before 0.3.2, on a 16 GiB M4: Tev1 0.8B matched 18/24 held-out labels with 450 ms median latency and no timeouts at 1,500 ms, including full-excerpt checks. Tev1 4B matched 22/24 with a 10-second diagnostic deadline and 3.15-second median latency. Nimble matched 23/24 with a 30-second deadline and 11.4-second median latency. Both larger models exceeded the then-default 1,500 ms. A separate six-case regression with the 0.3.2 fixes passed for both tested Tev1 4B variants and Nimble; Tev1 0.8B matched only three cases. These small tests do not establish general accuracy or Jev parity. See the [measurements and limits](docs/ollama-evaluation.md) and [Ollama reference](docs/reference.md#ollama-evaluator).
79
100
 
80
101
  ## Behavior and data
81
102
 
82
103
  - The default client profile permits all three routing tiers. Tool continuations, thinking history, model-specific features, and context size can keep or upgrade a model even when the evaluator chooses a cheaper tier. [Routing policy](docs/reference.md#routing-policy).
83
- - The selected evaluator receives bounded excerpts that can contain source code, system instructions, and tool results: TypeSafe with Jev, or the local service with Ollama. Anthropic receives the complete request. Images, document payloads, and private thinking are omitted from classifier input. [Data flow and authentication](docs/reference.md#data-flow-and-authentication).
104
+ - The selected evaluator receives bounded excerpts that can contain source code and tool results: TypeSafe with Jev, or the local service with Ollama. Jev also receives system-text excerpts; the local path excludes Claude's executor system instructions. Anthropic receives the complete request. Images, document payloads, and private thinking are omitted from classifier input. [Data flow and authentication](docs/reference.md#data-flow-and-authentication).
84
105
  - Subscription access and usage limits still apply. Model switches can reduce cache reuse; cheaper token prices do not guarantee cheaper completed tasks. Run ordinary `claude` to bypass routing.
85
106
  - The launcher is quiet by default. Use `AUTOROUTER_DEBUG=1` for metadata diagnostics or `AUTOROUTER_STATUSLINE=0` to retain your existing status line. [Troubleshooting](docs/reference.md#troubleshooting).
86
107
 
@@ -9,7 +9,7 @@ import { dirname } from 'node:path';
9
9
  import { createStatusState } from '../src/status-state.mjs';
10
10
  import { addStatusLineSettings } from '../src/status-settings.mjs';
11
11
  import { loadUserConfig } from '../src/user-config.mjs';
12
- import { setup, doctor } from '../src/onboarding.mjs';
12
+ import { setup, doctor, ollamaDeadlineText } from '../src/onboarding.mjs';
13
13
  import { setupOllama } from '../src/ollama-setup.mjs';
14
14
 
15
15
  const [command = 'help', ...args] = process.argv.slice(2);
@@ -21,8 +21,8 @@ if (['--version', '-v', 'version'].includes(command)) {
21
21
 
22
22
  Usage:
23
23
  claude-autorouter setup [--auth-mode subscription|api-key] [--force]
24
- [--evaluator jev|ollama] [--ollama-preset compact|quality|auto]
25
- [--ollama-model MODEL] [--pull]
24
+ [--evaluator jev|ollama]
25
+ [--ollama-model MODEL] [--ollama-timeout-ms N] [--pull]
26
26
  claude-autorouter doctor
27
27
  claude-autorouter claude [Claude Code arguments]
28
28
  claude-autorouter serve
@@ -35,11 +35,14 @@ AUTOROUTER_CONFIG selects a different file; environment variables take precedenc
35
35
  Project .env files are never loaded automatically.
36
36
 
37
37
  Jev is the default evaluator and requires TYPESAFE_API_KEY.
38
- Ollama evaluates locally and requires a running local Ollama service.
38
+ Ollama evaluates locally and requires Ollama 0.35+ with /v1/systemone.
39
39
  Use setup --evaluator ollama --pull to detect Ollama and download a missing model.
40
- Compact uses Qwen3 1.7B; quality offers Qwen3 4B for machines over 24 GB.
40
+ The local default is nimble:9b-q4_K_M; --ollama-model selects another compatible model.
41
+ Smaller Tev1 options: --ollama-model tev1:0.8b or --ollama-model tev1:4b-q4_K_M.
42
+ Local routing deadlines: Tev1 0.8B/custom 1500 ms, Tev1 4B 15000 ms, Nimble 30000 ms.
43
+ Setup --ollama-timeout-ms N saves a routing deadline; use 0 to disable it.
44
+ AUTOROUTER_OLLAMA_TIMEOUT_MS also overrides the deadline; 0 disables it.
41
45
  Local routing is experimental; see docs/ollama-evaluation.md for measured limits.
42
- The auto preset selects using total RAM; compact is the default.
43
46
  AUTOROUTER_AUTH_MODE=subscription uses your saved Claude Code login.
44
47
  Without setup, AUTOROUTER_AUTH_MODE defaults to api-key and also requires ANTHROPIC_API_KEY.
45
48
  AUTOROUTER_CLIENT_PROFILE=compatible (default) enables all three routing tiers.
@@ -84,7 +87,7 @@ Complete inference requests still go to Anthropic. See README.md.`);
84
87
  config.localToken = randomBytes(32).toString('hex');
85
88
  }
86
89
  if (config.evaluator === 'ollama') {
87
- console.error(`Preparing local Ollama evaluator (${config.ollamaModel})…`);
90
+ console.error(`Preparing local Ollama evaluator (${config.ollamaModel}); ${ollamaDeadlineText(config.ollamaTimeoutMs)}…`);
88
91
  try { await setupOllama(config, { pull: false, warm: true, write: () => {} }); }
89
92
  catch {
90
93
  console.error('Ollama could not be prepared. Requests will use the conservative fallback while it is unavailable; run claude-autorouter doctor.');
@@ -26,9 +26,37 @@ node --env-file=.env bin/autorouter.mjs claude
26
26
 
27
27
  The explicit Node flag loads `.env`; the CLI itself does not auto-load project files. Environment values override the user config. Keep keys out of source control and command arguments.
28
28
 
29
+ With AutoRouter 0.3.2 or newer, launch an already configured Ollama evaluator without its runtime deadline using:
30
+
31
+ ```sh
32
+ AUTOROUTER_OLLAMA_TIMEOUT_MS=0 claude-autorouter claude
33
+ ```
34
+
35
+ Persist the setting with `claude-autorouter setup --evaluator ollama --ollama-model tev1:4b --ollama-timeout-ms 0 --force` for that installed model. The setup flag overrides the timeout environment value; later launch-time environment values still override saved configuration. Startup priming keeps its separate 60-second limit, cancellation remains active, and normal errors still use fallback.
36
+
37
+ ## Local routing regression
38
+
39
+ The source-only harness below is opt-in and is not included in the npm package. It sends synthetic Claude-shaped requests through the real router and an installed local evaluator, checking task extraction, classifier choices, selected Claude tiers, and new human turns. It makes no Anthropic or Jev calls, downloads no models, and writes no user configuration.
40
+
41
+ ```sh
42
+ node scripts/test-ollama-routing.mjs --model tev1:0.8b
43
+ node scripts/test-ollama-routing.mjs --model tev1:4b
44
+ node scripts/test-ollama-routing.mjs --model nimble:9b-q4_K_M
45
+ ```
46
+
47
+ Start Ollama 0.35+ and install the selected model first. The harness refuses to run while another model is resident. It never explicitly unloads the selected model; its keep-alive setting controls residency. It warms once, then uses the production deadline for each uncached case: 1,500 ms for Tev1 0.8B/custom tags, 15,000 ms for official Tev1 4B tags, and 30,000 ms for official Nimble tags. Environment settings or `--timeout-ms N` can override the deadline; `--timeout-ms 0` disables the runtime timer while retaining cancellation and the separate warmup limit. This harness does not load saved user configuration. Use `--output artifacts/local-routing.json` to save a metadata report.
48
+
49
+ ```sh
50
+ node scripts/test-ollama-routing.mjs --model tev1:4b --timeout-ms 0
51
+ ```
52
+
53
+ The command fails on a wrong classification, fallback, unexpected guard override, or missing tier coverage. A Sonnet result with `source: ollama` and `classified_tier: sonnet` is a valid prediction; `source: fallback` and `classifier_error: timeout` means classification did not complete. Passing establishes these synthetic cases only. Warmup and metadata-only `doctor` checks do not establish speed or accuracy on real tasks.
54
+
55
+ Version 0.3.2 excludes Claude's executor system instructions from local classifier input before excerpt budgeting and retains task/history excerpts. Jev is unchanged. Version 0.3.1 included the executor background locally and defaulted every local model to 1,500 ms; see the [upgrade notes](reference.md#migrating-an-older-ollama-config). Historical benchmark results must remain labeled with their original excerpt policy and explicit deadlines.
56
+
29
57
  ## Live integration tests
30
58
 
31
- These tests make real Claude calls and invoke the configured evaluator, consuming Claude usage and, with Jev, TypeSafe usage. They use temporary synthetic fixtures and disable unrelated customizations and MCP servers. With Ollama, start the local service and install the chosen model first; the harness does not install or download it.
59
+ The Claude integration tests below make real Claude calls and invoke the configured evaluator, consuming Claude usage and, with Jev, TypeSafe usage. They use temporary synthetic fixtures and disable unrelated customizations and MCP servers. With Ollama, start the local service and install the chosen model first; the harness does not install or download it.
32
60
 
33
61
  ```sh
34
62
  npm run test:live
@@ -48,18 +76,19 @@ npm run eval
48
76
 
49
77
  The bundled evaluation makes 12 classifier calls and no Claude generations. Jev is the default and incurs TypeSafe usage; set `AUTOROUTER_EVALUATOR=ollama` to evaluate an installed local model. It reports agreement with the starting rubric, fallback count, and p50/p95 routing latency. Edit `test/fixtures/routing.json` to represent the tasks you want to measure. Rubric agreement alone does not establish answer quality or net savings; compare completed tasks against fixed-model baselines.
50
78
 
51
- For local evaluator measurements, distinguish cold model loading from warmed classification, and record the model tag, hardware, Ollama version, context size, prompt length, and resident memory. The launcher primes the classifier rubric with a synthetic task before opening the UI, with a separate deadline of up to 60 seconds. Runtime and benchmark share the 3,000-character/3,000-UTF-8-byte state limit, so include non-ASCII cases and excerpts that fill the budget. Also measure the first request after keep-alive expiration: its reload can hit the normal deadline even when warm requests pass. Repeat on realistic prompt distributions instead of selecting a model from a single easy request. Disk download size is not resident RAM, and the larger preset is not a speed guarantee. Keep model downloads opt-in and respect each model's license.
79
+ For local evaluator measurements, use Ollama 0.35+ and a model compatible with `/v1/systemone`. Distinguish cold model loading from warmed classification, and record the model tag, hardware, Ollama version, context size, prompt length, and resident memory. The launcher primes the classifier with a synthetic task before opening the UI, with a separate deadline of up to 60 seconds. Runtime and benchmark share the 3,000-character/3,000-UTF-8-byte state limit, so include non-ASCII cases and excerpts that fill the budget. Also measure the first request after keep-alive expiration: its reload can hit the normal deadline even when warm requests pass. Repeat on realistic prompt distributions instead of selecting a model from a single easy request. Disk download size is not resident RAM. Keep model downloads opt-in and respect each model's license.
52
80
 
53
81
  The dedicated Ollama benchmark uses synthetic tuning/held-out fixtures and reports cold latency separately from repeated warm requests:
54
82
 
55
83
  ```sh
56
- npm run eval:ollama -- --models qwen3:1.7b,qwen3:4b --split heldout --rounds 3 --stress-rounds 8
84
+ npm run eval:ollama -- --models nimble:9b-q4_K_M --split heldout --rounds 3 --stress-rounds 8
85
+ npm run eval:ollama -- --models tev1:0.8b,tev1:4b-q4_K_M --split heldout --rounds 1 --stress-rounds 8
57
86
  node scripts/evaluate-ollama.mjs --help
58
87
  ```
59
88
 
60
89
  Install each selected model first and use an idle Ollama instance with no resident models. The benchmark loads one candidate at a time and unloads it afterward. It does not download models or contact Claude or Jev. It reports classification errors and under/over-routing as well as latency; fixture labels are subjective rubric judgments, not measurements of completed task quality.
61
90
 
62
- On the 16 GiB M4 test Mac, compact matched 58.3% of held-out labels at warm p50/p95 latency of 602/834 ms and 1.70 GB model allocation. Quality (`qwen3:4b`) matched 91.7% at 889/1,242 ms and 3.18 GB. Each was tested on the same 24 distinct synthetic workloads repeated three times. Quality had no under-routing on this fixture, but routed two distinct Sonnet workloads to Opus on all repetitions. Both models exceeded the 1,500 ms deadline on all eight full-excerpt stress requests. No Jev comparison was run; local classification remains experimental. See the [local evaluator measurements](ollama-evaluation.md) for the hardware, candidate comparisons, and limits. A successful launcher/Enterprise integration request establishes connectivity and model routing, not classifier accuracy or parity with Jev.
91
+ See the [local evaluator measurements](ollama-evaluation.md) for the hardware, results, and limits. The older Qwen chat adapter and its compact/quality/auto presets have been removed. Every local candidate now uses native choice scoring; context allocation comes from the model/server configuration. Native entropy confidence is not calibrated accuracy and is not used as Jev's confidence threshold. A successful launcher/Enterprise integration request establishes connectivity and model routing, not classifier accuracy or parity with Jev.
63
92
 
64
93
  ## Startup context diagnostics
65
94
 
@@ -1,100 +1,121 @@
1
- # Local evaluator measurements
1
+ # Local decision evaluator measurements
2
2
 
3
- Jev remains the default evaluator. Ollama is an optional local classifier: Claude still generates the answer, and the router's capability and continuation guards still apply. This evaluation uses synthetic prompts only and does not measure the quality of Claude's completed work.
3
+ AutoRouter supports `/v1/systemone` classification only: remote Jev with an API key, or local Ollama 0.35+ with a compatible decision model. Jev remains the default evaluator. The local default is `nimble:9b-q4_K_M`; the old Qwen chat adapter and compact/quality/auto presets have been removed. Existing downloaded models are not deleted.
4
4
 
5
- ## Candidates and method
5
+ ## Version 0.3.2 routing regression (September 30, 2026)
6
6
 
7
- Measurements were recorded on September 29, 2026, on an Apple M4 Mac with 16 GiB of unified memory, running Ollama 0.33.3 alongside other applications. Candidates were loaded one at a time with a 4,096-token context. Download size is not the same as resident memory.
7
+ The npm `0.3.1` runtime used a 1,500 ms deadline for every local model, even though startup priming allows 60 seconds. On the user's already-loaded `tev1:4b` (Q8), four synthetic tasks all timed out and selected Sonnet through `source=fallback`, `classifier_error=timeout`. With a separate 30-second diagnostic allowance, the same tasks selected Haiku, Haiku, Sonnet, and Opus. Three of those decisions took 5.6–6.1 seconds. This was a deadline failure, not a parser forcing every answer to Sonnet.
8
8
 
9
- | Candidate | Published download size | Purpose |
10
- | --- | ---: | --- |
11
- | `qwen3.5:0.8b` | Approximately 1.0 GB | Smallest candidate |
12
- | `qwen3:1.7b` | Approximately 1.4 GB | Compact speed baseline |
13
- | `qwen3.5:2b` | Approximately 2.7 GB | Additional compact candidate |
14
- | `qwen3.5:4b` | Approximately 3.4 GB | Measured larger candidate; rejected for preset latency |
15
- | `qwen3:4b` | Approximately 2.5 GB | Larger candidate selected for the `quality` option |
9
+ We captured three synthetic request shapes from the installed Claude client using an isolated local response stub, without paid provider calls. The human task survived intact and no compatibility guard forced Sonnet. The local evaluator nevertheless received 627–796 characters of Claude's general executor instructions. Tev1 0.8B classified all three captures as Sonnet. Removing only this system background changed the distributed-fencing case to Opus; the `[].length` example still selected Sonnet. Removing additional state fields did not improve that result, so the routing rubric and remaining state layout were retained.
16
10
 
17
- Sizes and quantization vary by tag. Use explicit tags: untagged `qwen3.5` currently selects the much larger 9B model. See the official [Qwen3.5 model catalog](https://ollama.com/library/qwen3.5) and [Qwen3 1.7B listing](https://ollama.com/library/qwen3:1.7b).
11
+ Version 0.3.2 excludes executor system instructions from local classifier input, uses 15-second defaults for official Tev1 4B variants and 30 seconds for Nimble, and retains the 1.5-second default for Tev1 0.8B/custom models. Explicit timeout settings still win; `0` disables only the runtime evaluator timer. Setup accepts `--ollama-timeout-ms`; setup, doctor, and startup display the effective deadline. Compact status lines retain the fallback cause. Jev's input extraction and deadline are unchanged.
18
12
 
19
- The `compact` option selects `qwen3:1.7b` for its smaller allocation and faster measured classification. The `quality` option selects `qwen3:4b`, which agreed more often with the held-out labels while using more memory and time. The latter's official listing specifies a roughly 2.5 GB download and Q4_K_M quantization. Additional memory headroom is useful when other applications are open, but more RAM alone does not guarantee better accuracy or meeting the evaluator deadline. See the [official Qwen3 4B listing](https://ollama.com/library/qwen3:4b).
13
+ The source-only, opt-in `npm run test:ollama -- --model TAG` checks actual `Router.route` results against six fixed synthetic Claude-shaped requests: literal output, `[].length`, a bounded feature, distributed fencing, a new mechanical task after a difficult task, and a new difficult task after a simple task. It checks the evaluator choice, selected Claude model, source, and reason, and fails on a mismatch, fallback, override, or missing tier coverage. Live cases use fresh routers to avoid decision-cache passes; persistent turn-cache transitions and compatibility guards are separately covered by offline regression tests. It is not included in the npm package, downloads nothing, and does not contact Claude or Jev.
20
14
 
21
- The checked-in fixture contains 36 independently authored, balanced routing cases: 12 tuning cases and 24 held-out cases, with equal numbers of Haiku, Sonnet, and Opus labels. Cases cover mechanical edits, ordinary implementation, difficult correctness and security work, topic changes, tool results, short follow-ups, and misleading routing instructions. Expected labels are judgments under the routing rubric, not independently verified claims about which Claude model would succeed.
15
+ One pass with the fixes on the same M4/16 GiB Mac produced:
22
16
 
23
- The rubric was refined using the tuning set. The final held-out run uses the frozen rubric and is reported separately. Repeating each held-out case three times gives 72 decisions per model, but still only 24 distinct workloads.
17
+ | Installed model | Runtime deadline | Matching cases | Fallbacks | Decision latency range |
18
+ | --- | ---: | ---: | ---: | ---: |
19
+ | `tev1:0.8b` | 1,500 ms | 3 / 6 | 0 | 48–479 ms |
20
+ | `tev1:4b-q4_K_M` | 15,000 ms | 6 / 6 | 0 | 2.40–2.75 s |
21
+ | `tev1:4b` (Q8) | 15,000 ms | 6 / 6 | 0 | 2.38–2.70 s |
22
+ | `nimble:9b-q4_K_M` | 30,000 ms | 6 / 6 | 0 | 4.74–7.29 s |
24
23
 
25
- The harness calls the production `buildOllamaState` and `evaluateOllama` functions. It includes local model metadata checks in wall-clock latency and bypasses AutoRouter's decision cache. Ollama's normal shared-prefix caching remains enabled. Input excerpts have a 3,000-character/UTF-8-byte ceiling. Requests disable thinking, use temperature 0, seed 0, a 32-token output cap, and a JSON schema containing only the tier. A returned tier is not a calibrated confidence probability. See Ollama's [thinking controls](https://docs.ollama.com/capabilities/thinking) and [structured-output guidance](https://docs.ollama.com/capabilities/structured-outputs).
24
+ Every model reached all three tiers. The 0.8B failures were genuine classifications: `[].length` → Sonnet, new mechanical task → Opus, new difficult task → Sonnet. Its live suite therefore **fails**, rather than treating valid native responses as proof of accuracy. These six cases are regression checks, not a new held-out quality benchmark; the samples, machine load, residency, and prompt-cache effects do not establish a quantization speed comparison or a worst-case latency bound. The larger-model budgets allow slower decisions; they do not make those models fast or prevent every timeout. No user configuration or model files were changed; the originally resident `tev1:4b` was restored after sequential model testing.
26
25
 
27
- A cold measurement starts with the model unloaded from Ollama; operating-system file caches and compiled kernels may already be warm. Cold calls have a separate 60-second measurement deadline. Warm measurements use the production 1,500 ms deadline. Model allocation is read from `/api/ps`, not inferred from the download size. On unified-memory hardware, its GPU allocation is not additional independent RAM. See [Ollama's running-model API](https://docs.ollama.com/api/ps) and [context-memory guidance](https://docs.ollama.com/context-length).
26
+ A separate check with `tev1:4b` and the runtime deadline disabled (`timeout_ms: 0`) passed the same six cases without fallback, with decision latencies of 4.37–5.42 seconds after 13.11 seconds of priming. It made no Claude or Jev calls and downloaded nothing. This validates disabled-timer routing on those cases; it adds no new quality cases or latency guarantee.
28
27
 
29
- ## Results
28
+ Fixture SHA-256: `88e12f962fb42b36a00edd6b8c2560e54620d90c978b1ff13a1ddf31ee46b585`. The questions remain unchanged at `be151cedb4de4b7ef3f7162d751f70ce7d9dd14efc66fae1835f73ffd04027be`.
30
29
 
31
- The frozen-rubric tuning results below document model selection. They are not held-out performance. GB values use decimal bytes; allocation is what the local running-model API reported, rather than total process or system memory.
30
+ ## Historical benchmark method (before 0.3.2)
32
31
 
33
- | Model | Tuning agreement | Timeouts | Warm p50 / p95 | Cold wall time | Disk size | Model allocation |
34
- | --- | ---: | ---: | ---: | ---: | ---: | ---: |
35
- | `qwen3:1.7b` | 20 / 24 (83.3%) | 0 / 24 | 245 / 280 ms | 2.50 s | 1.36 GB | 1.70 GB |
36
- | `qwen3:4b` | 24 / 24 (100%) | 0 / 24 | 743 / 889 ms | 9.62 s | 2.50 GB | 3.18 GB |
37
- | `qwen3.5:0.8b` | 10 / 24 (41.7%) | 0 / 24 | 409 / 437 ms | 2.78 s | 1.04 GB | 1.09 GB |
38
- | `qwen3.5:2b` | 18 / 24 (75.0%) | 0 / 24 | 918 / 1,009 ms | 6.49 s | 2.74 GB | 2.36 GB |
39
- | `qwen3.5:4b` | No valid warm results | 24 / 24 | Deadline reached | 7.26 s | 3.39 GB | 3.14 GB |
32
+ Measurements were recorded on September 29, 2026, on an Apple M4 Mac with 16 GiB of unified memory alongside other applications. Nimble ran on an isolated Ollama 0.35.0 process; the installed Ollama 0.33.3 daemon was left unchanged during that test. Tev1 was tested later on the user's upgraded Ollama 0.35.0 service after its downloads finished. Version 0.35.0 was a prerelease at the time. These are observations under different application loads, not a controlled hardware comparison. See the [official release](https://github.com/ollama/ollama/releases/tag/v0.35.0), [Nimble catalog](https://ollama.com/library/nimble), and [Tev1 catalog](https://ollama.com/library/tev1).
40
33
 
41
- The smallest model was highly sensitive to rubric wording. It should not be selected solely because its download is small. The 2B candidate was slower and agreed less often than 1.7B on this tuning set. A separate diagnostic gave `qwen3.5:4b` a 5,000 ms deadline: all 12 tuning decisions agreed, but warm p50/p95 was 3,615/4,872 ms. **Qwen3.5 4B is rejected as a preset because it missed the production deadline on every warm tuning request.** Its diagnostic result demonstrates a latency tradeoff, not held-out quality. Extra RAM alone does not establish that this model will meet a 1,500 ms deadline on another machine.
34
+ The selected Nimble tag contains a 9B Q4_K_M model, approximately 5.63 GB to download, with an 8,194-token native context setting. `/v1/systemone` scores choices directly; it accepts no chat generation options or per-request context override. The model/server configuration determines allocation. The production excerpt remains bounded to 3,000 serialized characters and UTF-8 bytes, including for non-ASCII text. Native confidence measures the concentration of the choice distribution, not calibrated accuracy, and is not used as Jev's confidence threshold.
42
35
 
43
- An initial 4B run also exposed a metadata-response limit: its local `/api/show` response exceeded 64 KiB. That issue was fixed before the results above by bounding metadata separately at 1 MiB while retaining a 64 KiB classifier-output limit.
36
+ The fixture contains 36 balanced synthetic workloads: 12 tuning cases and 24 held-out cases, with equal numbers of Haiku, Sonnet, and Opus labels. Cases include mechanical edits, ordinary implementation, difficult correctness and security work, topic changes, tool results, short follow-ups, and misleading routing instructions. These labels are judgments under the routing policy, not proof of which Claude model would complete each task successfully. The native questions preserve the existing local routing policy and were frozen before Nimble testing; no changes were made from held-out results.
44
37
 
45
- ### Held-out results
38
+ The historical harness used the then-current production state builder and evaluator, including local metadata checks in wall-clock latency. It bypasses AutoRouter's decision cache; Ollama's own caching remains enabled. A cold measurement starts with the model unloaded, but operating-system file caches and kernels may already be warm. Cold calls have a separate 60-second deadline. The normal evaluator deadline in that benchmark was 1,500 ms for every model. Reported model allocation comes from `/api/ps`; it is not a measurement of total process or system memory, and its GPU allocation is not additional independent RAM on this unified-memory Mac. Aggregate runtime RSS includes all Ollama and llama-server processes, including the original idle daemon.
46
39
 
47
- Both selected candidates completed all 72 held-out decisions without a timeout. Qwen3 4B agreed more often than 1.7B, including every Opus-labeled workload. These measurements were taken later than tuning while other applications remained open; they do not represent an isolated hardware benchmark.
40
+ ## Historical Nimble results
48
41
 
49
- | Model | Agreement | Timeouts | Warm p50 / p95 | Cold wall time | Under-routes / over-routes |
50
- | --- | ---: | ---: | ---: | ---: | ---: |
51
- | `qwen3:1.7b` | 42 / 72 (58.3%) | 0 / 72 | 602 / 834 ms | 4.80 s | 18 / 12 |
52
- | `qwen3:4b` | 66 / 72 (91.7%) | 0 / 72 | 889 / 1,242 ms | 5.92 s | 0 / 6 |
42
+ The first production-deadline run returned Haiku correctly for its cold mechanical task in 17.64 seconds. **All 12 warm tuning requests timed out at 1,500 ms**, with cancellation p50/p95 of 1,503/1,513 ms. No warm classification accuracy can be inferred from that run. The model's reported allocation was 5.48 GB, with an 8,194-token context. This did not meet the desired fast-routing target on the tested Mac.
43
+
44
+ A separate first-load smoke call correctly classified `[].length` as Haiku in 25.23 seconds. It checks integration, not warm performance or classifier accuracy.
45
+
46
+ The separate held-out diagnostic used a 30,000 ms deadline and one pass over 24 distinct workloads. It did not change the then-default 1,500 ms budget or establish performance within it.
47
+
48
+ | Measurement | Result |
49
+ | --- | ---: |
50
+ | Held-out rubric agreement | 23 / 24 (95.8%) |
51
+ | Valid decisions / timeouts | 24 / 0 |
52
+ | Warm p50 / p95 wall time | 11,432 / 16,334 ms |
53
+ | Minimum / maximum wall time | 697 / 20,201 ms |
54
+ | Cold wall time | 15.47 s |
55
+ | Reported model allocation | 5.48 GB |
56
+ | Under-routes / over-routes | 1 / 0 |
53
57
 
54
- The 1.7B confusion matrix:
58
+ All eight Haiku and eight Sonnet labels matched. Seven of eight Opus labels matched. The `held-o-injection` case, a difficult deadlock investigation containing an instruction to select Haiku, was incorrectly routed to Haiku. This is a concrete limitation against misleading routing instructions. These 24 decisions are a single pass, not repeated measurements or a claim of comparable performance to Jev.
55
59
 
56
- | Expected tier | Returned Haiku | Returned Sonnet | Returned Opus |
57
- | --- | ---: | ---: | ---: |
58
- | Haiku | 12 | 6 | 6 |
59
- | Sonnet | 0 | 24 | 0 |
60
- | Opus | 0 | 18 | 6 |
60
+ All eight full-excerpt diagnostic requests completed at the 30,000 ms deadline and returned the expected Haiku tier. Their p50/p95 latency was 25,711/27,271 ms. These synthetic requests fill the 3,000-byte state budget with clearly mechanical tasks and unrelated tool output. They are separate performance checks, not eight additional held-out workloads. Longer excerpts can consume almost the entire diagnostic deadline on this machine.
61
61
 
62
- The 4B confusion matrix:
62
+ An isolated live Claude Code test also passed using the saved Enterprise subscription login, a temporary configuration with a 30,000 ms evaluator deadline, and no Jev key. Nimble selected Haiku in 12.69 seconds; Anthropic returned HTTP 200 with the expected literal response and confirmed `claude-haiku-4-5-20251001`. Only a synthetic prompt was used, with no repository files or tools. This verifies the authentication and routing integration, not general classifier accuracy. The installed Ollama service and user configuration were left unchanged.
63
63
 
64
- | Expected tier | Returned Haiku | Returned Sonnet | Returned Opus |
65
- | --- | ---: | ---: | ---: |
66
- | Haiku | 24 | 0 | 0 |
67
- | Sonnet | 0 | 18 | 6 |
68
- | Opus | 0 | 0 | 24 |
64
+ ## Historical Tev1 results
69
65
 
70
- Each row represents eight unique workloads repeated three times. The wrong labels were consistent across repetitions. For 1.7B, six of eight Opus workloads were sent to Sonnet, and four of eight Haiku workloads were sent to a higher tier. For 4B, two of eight Sonnet workloads were sent to Opus. The compact model's gap between tuning and held-out agreement limits what can be claimed about it. Both local options remain experimental; these results do not establish parity with Jev, which was not evaluated on this fixture.
66
+ Both [Tev1 variants](https://ollama.com/library/tev1) use the same production adapter and frozen questions, selected through `--ollama-model`. The 0.8B tag uses Q8_0 quantization and downloads approximately 812 MB; `tev1:4b-q4_K_M` downloads approximately 2.71 GB. The unqualified `tev1` tag selects the larger 4B Q8 model, which was not tested in this historical benchmark. It was tested separately in the 0.3.2 regression above. Jev remains the evaluator default and Nimble remains the local-model default.
71
67
 
72
- ### Full-excerpt performance
68
+ Each measured tag ships `num_ctx:2050`. The window includes the template, routing criteria, and excerpt. The 3,000-byte state cap is not a guarantee that every possible input fits this smaller token window. The reported stress cases used 1,664–1,752 input tokens on 0.8B. Requests exceeding model limits use the usual fallback; AutoRouter does not switch protocols or silently truncate additional content for Tev1.
69
+
70
+ These measurements use one pass over the same 24 held-out cases and eight separate full-excerpt cases, after both downloads finished. An exploratory 0.8B run during the 4B download gave the same labels; its timings are excluded here. Neither questions nor expected labels were changed in response to Tev1 outputs.
71
+
72
+ | Model | Deadline | Held-out agreement | Timeouts | Warm p50 / p95 | Reported allocation |
73
+ | --- | ---: | ---: | ---: | ---: | ---: |
74
+ | `tev1:0.8b` | 1,500 ms | 18 / 24 (75.0%) | 0 / 24 | 450 / 488 ms | 0.89 GB |
75
+ | `tev1:4b-q4_K_M` | 1,500 ms | No valid warm decisions | 24 / 24 | Deadline reached | 2.91 GB |
76
+ | `tev1:4b-q4_K_M` diagnostic | 10,000 ms | 22 / 24 (91.7%) | 0 / 24 | 3,149 / 4,169 ms | 2.91 GB |
73
77
 
74
- An additional eight synthetic requests per model filled the entire 3,000-byte state budget. They kept a clearly mechanical current task and included long synthetic tool output. A different nonce at the beginning of each serialized state prevented reuse of the previous full user-state prefix while preserving the common rubric prefix. Both models exceeded the 1,500 ms deadline on all eight requests. These are separate performance checks, not additional held-out accuracy cases. Short-prompt latency must not be treated as a bound for full excerpts.
78
+ The 0.8B model matched four of eight Haiku labels, all eight Sonnet labels, and six of eight Opus labels. It over-routed four mechanical tasks and under-routed two difficult tasks to Sonnet, including a case with a misleading tier instruction. Cold wall time was 2.08 seconds for 0.8B and 6.04 seconds for 4B. Model size and fast responses do not establish sufficient accuracy for an engineering workload.
75
79
 
76
- | Model | Timeouts at 1,500 ms | Wall p50 / p95 before cancellation |
77
- | --- | ---: | ---: |
78
- | `qwen3:1.7b` | 8 / 8 | 1,502 / 1,503 ms |
79
- | `qwen3:4b` | 8 / 8 | 1,502 / 1,505 ms |
80
+ With a separate 10-second deadline, 4B matched seven of eight Haiku labels, all eight Sonnet labels, and seven of eight Opus labels. It over-routed one mechanical case to Opus and under-routed one difficult case to Sonnet. Its cold diagnostic request took 3.69 seconds. The improved agreement comes with several seconds of classification latency; it is not performance at the then-default 1,500 ms deadline.
80
81
 
81
- A diagnostic rerun of 1.7B with a 5,000 ms deadline completed all eight and returned Haiku, with p50/p95 of 3,753/3,991 ms. Qwen3 4B still timed out on all eight at that longer deadline, with p50/p95 cancellation times of 5,002/5,007 ms. Its actual completion times for these full excerpts were not measured. A larger model and startup priming do not remove the need for a bounded timeout and fallback during longer requests.
82
+ All eight 0.8B full-excerpt requests completed within 1,500 ms and returned Haiku, with p50/p95 of 1,079/1,150 ms. The 4B model timed out on all eight at 1,500 ms and again on all eight at 10,000 ms. Its 10-second cancellation p50/p95 was 10,006/10,081 ms; completed full-excerpt latency was not measured. The longer deadline therefore allows the reported short held-out decisions but does not guarantee completion for full excerpts.
82
83
 
83
- All final comparisons use fixture SHA-256 `1ef5111a6a36f0f4bc8d111d54c6985cac4c4a8b16357023a0958d285f7438ad` and rubric SHA-256 `42c3e18ddcf7c9d8756100b740f686b3c2f1bbdfbcbaf3cbc3021a9c9b4c6ee6`.
84
+ The native API is the supported Ollama integration. Together's [publisher interface](https://huggingface.co/togethercomputer/Tev1-4B-experimental#intended-interface) describes a different training prompt layout from the schema rendered by [Ollama 0.35's compiler](https://github.com/ollama/ollama/blob/v0.35.0/decision/systemone.go). These results measure the actual Ollama native path, not a reproduction of Together's training-format evaluation or its published accuracy figures.
85
+
86
+ Both tags passed isolated live Claude Enterprise integration checks on the user's Ollama 0.35 service, with no Jev key and no repository tools or files. Claude returned the correct `0` for a synthetic `[].length` query and Anthropic returned HTTP 200. The 0.8B evaluator took 584 ms under the default deadline but selected Sonnet, over-routing this mechanical task. The 4B evaluator selected Haiku in 3.78 seconds using a temporary 10-second deadline. These checks verify setup, authentication, native evaluation, and generation; they do not imply that every routing decision is correct. Temporary configurations were removed, tested models were unloaded from memory, and the user's Ollama service and model files were retained.
87
+
88
+ Model digests:
89
+
90
+ - `tev1:0.8b`: `c0099a86fcbd81bc5876a0d1f94998d2b038f7f1b5f7329a3dba43a36903c652`
91
+ - `tev1:4b-q4_K_M`: `3509ac7180e86e5fa8efc7b5745e32d55dd9d4e0a86bc9a88aba5323a5d29bc6`
84
92
 
85
93
  ## Reproducing the evaluation
86
94
 
87
- Use a source checkout; benchmark scripts and fixtures are development files and are not bundled in the npm package. Start local Ollama and explicitly download the models you intend to test. The harness never downloads or deletes models, and refuses to begin while another model is resident. It unloads each tested model after its measurements.
95
+ Use a source checkout; benchmark scripts and fixtures are not included in the npm package. The commands below evaluate the checked-out version. To reproduce the historical excerpt policy, use the `v0.3.1` checkout; 0.3.2 changes local input extraction. The explicit deadlines retain the historical budgets. Install Ollama 0.35+, start it, and explicitly download the model:
88
96
 
89
97
  ```sh
90
- node scripts/evaluate-ollama.mjs --models qwen3:1.7b,qwen3:4b --split tuning --rounds 2 --output artifacts/ollama-tuning.json
91
- node scripts/evaluate-ollama.mjs --models qwen3:1.7b,qwen3:4b --split heldout --rounds 3 --stress-rounds 8 --output artifacts/ollama-heldout.json
98
+ ollama pull nimble:9b-q4_K_M
99
+ node scripts/evaluate-ollama.mjs --models nimble:9b-q4_K_M --split tuning --rounds 1 --timeout-ms 1500 --output artifacts/nimble-tuning.json
100
+ node scripts/evaluate-ollama.mjs --models nimble:9b-q4_K_M --split heldout --rounds 1 --stress-rounds 8 --timeout-ms 30000 --output artifacts/nimble-diagnostic.json
101
+ ollama pull tev1:0.8b
102
+ ollama pull tev1:4b-q4_K_M
103
+ node scripts/evaluate-ollama.mjs --models tev1:0.8b,tev1:4b-q4_K_M --split heldout --rounds 1 --stress-rounds 8 --timeout-ms 1500 --output artifacts/tev1-default.json
104
+ node scripts/evaluate-ollama.mjs --models tev1:4b-q4_K_M --split heldout --rounds 1 --stress-rounds 8 --timeout-ms 10000 --output artifacts/tev1-diagnostic.json
92
105
  ```
93
106
 
94
- JSON reports contain fixture/rubric hashes, model digests, quantization, per-case labels, timing counters, failures, confusion matrices, and allocation measurements. They contain no private repository prompts or API credentials. Raw reports are written only when `--output` is supplied; `artifacts/` is ignored by Git.
107
+ Use `--endpoint http://127.0.0.1:PORT` for another local instance. The benchmark requires no resident models at startup, never downloads or deletes models, and unloads each tested model afterward. It sends only checked-in synthetic cases and does not contact Claude or Jev. Reports include model identity, protocol, fixture/question hashes, token counts, per-case results, timeouts, confusion matrices, latency, and reported allocation. The eight separate stress requests fill the excerpt budget and vary an early nonce to prevent reuse of the previous full state; they are performance checks, not held-out accuracy cases.
108
+
109
+ Reproducibility identifiers:
110
+
111
+ - Model digest: `3776806da5587387a996e28e75d5d07fbebe7410d47879e1f71782f36c896ce3`
112
+ - Fixture SHA-256: `1ef5111a6a36f0f4bc8d111d54c6985cac4c4a8b16357023a0958d285f7438ad`
113
+ - Native questions SHA-256: `be151cedb4de4b7ef3f7162d751f70ce7d9dd14efc66fae1835f73ffd04027be`
114
+
115
+ ## Limits and previous measurements
95
116
 
96
- ## Limits
117
+ This is a small synthetic rubric-agreement benchmark, not a downstream task-quality, savings, Jev-parity, or security evaluation. Classification can miss context outside the excerpt. The cases do not establish robust resistance to prompt injection. Performance depends on hardware, memory pressure, prompt length, and residency; a successful setup or simple request does not guarantee the runtime deadline.
97
118
 
98
- This is a small synthetic rubric-agreement benchmark, not a downstream task-quality, cost-savings, or security evaluation. Its prompts cannot represent every repository or long conversation. A model can agree with the labels and still miss important context outside the excerpt. The tests do not establish robust resistance to prompt injection.
119
+ The launcher primes the model before opening Claude, allowing up to 60 seconds for that synthetic classification. Idle unloading can still make later requests cold. Runtime timeouts and invalid responses use the existing conservative fallback, without contacting Jev. A longer `AUTOROUTER_OLLAMA_TIMEOUT_MS` trades added prompt latency for more completed local classifications; it does not make the evaluator faster. In 0.3.2, `0` disables the runtime timer while preserving user/disconnect cancellation, normal error fallback, and the separate startup limit.
99
120
 
100
- Latency depends on hardware, current application load, model residency, and prompt length. Cold loading exceeds the normal evaluator deadline, which is why the launcher primes an installed local model with a synthetic classification using the production request settings before starting Claude. The warm benchmark measurements above already followed a full classification, so this startup improvement does not change those measurements. Timeouts and invalid responses use the router's existing conservative fallback policy. Re-run the evaluation before adopting different tags or changing the rubric; no result here guarantees accuracy on your workload.
121
+ Earlier Qwen results used a different `/api/chat` implementation and are not measurements of this native backend. They remain available in the [historical evaluation document](https://github.com/frapposelli/claude-autorouter/blob/548a175/docs/ollama-evaluation.md). Reproduce those results from that revision, not the current native-only harness.