claude-autorouter 0.3.7 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.env.example +4 -2
- package/CONTRIBUTING.md +37 -0
- package/README.md +43 -70
- package/bin/autorouter.mjs +40 -57
- package/docs/development.md +48 -2
- package/docs/hardware-benchmark.md +29 -0
- package/docs/hardware-comparison.md +55 -0
- package/docs/hardware-results-16gb.json +4002 -0
- package/docs/hardware-results-16gb.md +26 -0
- package/docs/hardware-results-64gb.json +4020 -0
- package/docs/reference.md +71 -32
- package/docs/releasing.md +74 -34
- package/docs/router-performance.json +1697 -0
- package/docs/router-performance.md +50 -0
- package/docs/status-performance.json +363 -0
- package/docs/status-performance.md +44 -0
- package/package.json +57 -9
- package/src/auto-routing.mjs +184 -24
- package/src/bounded-json.mjs +57 -0
- package/src/cli-help.mjs +87 -0
- package/src/config-command.mjs +141 -0
- package/src/config.mjs +52 -27
- package/src/contracts.mjs +123 -0
- package/src/evaluation-report.mjs +114 -0
- package/src/local-diagnostic.mjs +191 -0
- package/src/model-catalog.mjs +96 -0
- package/src/model-request.mjs +6 -7
- package/src/ollama-evaluator.mjs +9 -27
- package/src/onboarding.mjs +82 -23
- package/src/request-validation.mjs +54 -0
- package/src/response-observer.mjs +126 -18
- package/src/router.mjs +151 -61
- package/src/savings.mjs +74 -16
- package/src/server.mjs +79 -12
- package/src/session-history.mjs +261 -0
- package/src/session-log.mjs +9 -58
- package/src/status-state.mjs +110 -62
- package/src/statusline.mjs +57 -27
- package/src/telemetry-event.mjs +196 -0
- package/src/token-counter.mjs +3 -1
- package/src/turn-state.mjs +132 -0
- package/src/user-config.mjs +18 -8
package/.env.example
CHANGED
|
@@ -12,10 +12,12 @@ AUTOROUTER_STATUSLINE=1
|
|
|
12
12
|
# AUTOROUTER_CLIENT_PROFILE=auto
|
|
13
13
|
# Optional metadata logs on stderr. Redirect stderr to a file when using the UI.
|
|
14
14
|
# AUTOROUTER_DEBUG=1
|
|
15
|
-
# Optional persistent JSONL
|
|
15
|
+
# Optional persistent JSONL decisions/outcomes, one file per session per launch.
|
|
16
16
|
# Includes up to 500 characters of user prompt text; keep the directory local.
|
|
17
17
|
# Unset or empty disables logging. This does not print prompts in the terminal.
|
|
18
18
|
# AUTOROUTER_SESSION_LOG_DIR=/absolute/path/to/autorouter-sessions
|
|
19
|
+
# Omit all prompt excerpt fields (does not enable logging by itself).
|
|
20
|
+
# AUTOROUTER_SESSION_LOG_MODE=metadata
|
|
19
21
|
# Optional: allow two tool-free Stop-hook continuations, then end the turn on
|
|
20
22
|
# the third block. Applies to /goal and all Stop/SubagentStop hooks.
|
|
21
23
|
# Unset keeps Claude's default (currently 8); 0 DISABLES the cap.
|
|
@@ -40,7 +42,7 @@ AUTOROUTER_MIN_CONFIDENCE=0.75
|
|
|
40
42
|
# Experimental local evaluator in AutoRouter 0.3.2+: native decision API only.
|
|
41
43
|
# Install/start Ollama 0.35+, then run:
|
|
42
44
|
# claude-autorouter setup --evaluator ollama --pull --force
|
|
43
|
-
# Defaults to nimble:9b-q4_K_M (~5.63 GB download); --force
|
|
45
|
+
# Defaults to nimble:9b-q4_K_M (~5.63 GB download); --force preserves other settings.
|
|
44
46
|
# Existing downloads are kept. Replace Qwen config/environment values from 0.2.0.
|
|
45
47
|
# All local models use /v1/systemone; custom tags/aliases must support that API.
|
|
46
48
|
# Add --ollama-model LOCAL_TAG_OR_ALIAS to setup to choose another suitable model.
|
package/CONTRIBUTING.md
ADDED
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
# Contributing to AutoRouter
|
|
2
|
+
|
|
3
|
+
Use Node.js 22+ and macOS or Linux. Install pinned development tools, then run the local checks:
|
|
4
|
+
|
|
5
|
+
```sh
|
|
6
|
+
npm ci --ignore-scripts --no-audit --no-fund
|
|
7
|
+
npm run check
|
|
8
|
+
npm test
|
|
9
|
+
npm run test:package
|
|
10
|
+
```
|
|
11
|
+
|
|
12
|
+
Tests use synthetic local services and credentials. They require loopback binding, but make no paid provider calls, downloads or user-config changes. The package check installs and exercises the exact distributable archive. Opt-in provider/model canaries are described in [development and validation](docs/development.md).
|
|
13
|
+
|
|
14
|
+
## Changing behavior
|
|
15
|
+
|
|
16
|
+
Follow the [request lifecycle](docs/development.md#request-lifecycle-and-model-continuity). Keep authentication and permission decisions owned by Claude. Preserve provider bytes, signed history and unfamiliar extensions. Routing must check compatibility in every profile; new human tasks remain eligible to switch models. Active task state is separate from disposable classification caches, and only clean, successfully forwarded completion evidence can establish confirmed continuation state.
|
|
17
|
+
|
|
18
|
+
Update `src/contracts.mjs` alongside configuration, decision or telemetry changes. Keep the normalizer allowlist and privacy tests aligned. Logged selections are not successful outcomes, logging stays opt-in, and metadata mode must omit prompts. [Static/style checks](docs/development.md#static-contracts-and-style) run in CI; do not add runtime dependencies for developer tooling.
|
|
19
|
+
|
|
20
|
+
## Model and pricing updates
|
|
21
|
+
|
|
22
|
+
1. Identify the exact provider model ID; a family keyword or custom alias is not capability evidence.
|
|
23
|
+
2. Update `src/model-catalog.mjs` with the authoritative source and review date. Verify context/output limits, thinking, tool choice, native tool features and Auto eligibility.
|
|
24
|
+
3. Add valid-source compatibility fixtures for every affected profile and a regression for the new restriction or permitted switch. Keep unknown models/extensions conservative.
|
|
25
|
+
4. Update thinking adaptation only where documented; token counting and inference must apply the same compatibility policy.
|
|
26
|
+
5. If rates change, review `src/savings.mjs`, bump its pricing version/date and source, and test cache TTL/modifier/unknown-model coverage. Never silently price old logs using an unrecorded new table.
|
|
27
|
+
6. Run all three local checks and exact-package fixtures. Report separately any explicitly invoked real-provider observations and their limits.
|
|
28
|
+
|
|
29
|
+
## Evaluation and performance
|
|
30
|
+
|
|
31
|
+
Declare label agreement, acceptable tiers and under-routing thresholds before testing a candidate. Preserve fixture checksums and held-out cases; transport, evaluator availability, policy and task quality are separate gates. Profile coverage requires all three tiers for compatible routing and Sonnet/Opus for Auto.
|
|
32
|
+
|
|
33
|
+
Capture a baseline before changing overhead. Repeat the same workload and hardware with the [router/storage harnesses](docs/development.md#performance-regression-measurements), retaining call counts, latency distributions, memory and background load. Numerical timing gates are local and opt-in; CI checks deterministic cancellation/resource behavior. Real local-model benchmarks need representative memory sizes and observed cold/warm conditions.
|
|
34
|
+
|
|
35
|
+
## Releasing
|
|
36
|
+
|
|
37
|
+
Use the [release procedure](docs/releasing.md) and retain the tested immutable archive. Acceptance by npm and public availability are separate states. Verify registry integrity and an isolated install before calling a release verified. A pending submission is investigated or verified again, rather than blindly republished. New CLI interfaces and event-schema changes should be reviewed together as a minor release; correctness-only fixes can be independent patches.
|
package/README.md
CHANGED
|
@@ -1,126 +1,99 @@
|
|
|
1
1
|
# Claude AutoRouter
|
|
2
2
|
|
|
3
|
-
Use Haiku, Sonnet
|
|
3
|
+
Use Haiku, Sonnet and Opus in one Claude Code session. AutoRouter evaluates each coding request, checks model compatibility and context capacity, and forwards it through a local gateway. [TypeSafe Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) is the default evaluator; native Ollama `/v1/systemone` models provide an experimental local option. Claude owns authentication, tool permissions and safety review.
|
|
4
4
|
|
|
5
|
-
Requires Node.js 22+, macOS or Linux (including WSL), an installed `claude` command, and a Claude subscription login or Anthropic API key. The default evaluator also
|
|
5
|
+
Requires Node.js 22+, macOS or Linux (including WSL), an installed `claude` command, and a Claude subscription login or Anthropic API key. The default evaluator also needs a [TypeSafe API key](https://console.typesafe.ai). The installed CLI has no runtime dependencies.
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
Version 0.4.0 adds `config`, `sessions` and `doctor --evaluate-local`, durable task continuity, and clearer model outcomes. Upgrade from 0.3.x to use these commands. The [contributor guide](CONTRIBUTING.md) explains local verification, and the [release guide](docs/releasing.md) covers the changes and verified publication.
|
|
8
8
|
|
|
9
|
-
Install
|
|
9
|
+
## Install and start
|
|
10
10
|
|
|
11
11
|
```sh
|
|
12
12
|
npm install -g claude-autorouter
|
|
13
|
-
```
|
|
14
|
-
|
|
15
|
-
Set up once, then launch from any project directory:
|
|
16
|
-
|
|
17
|
-
```sh
|
|
18
13
|
claude-autorouter setup
|
|
19
14
|
claude-autorouter doctor
|
|
20
15
|
cd /path/to/project
|
|
21
16
|
claude-autorouter claude
|
|
22
17
|
```
|
|
23
18
|
|
|
24
|
-
Setup defaults to your Claude subscription and prompts for
|
|
19
|
+
Setup defaults to your Claude subscription and prompts privately for the Jev key. Run `claude auth login` if needed. Jev has separate credentials and billing; subscription mode needs no Anthropic API key. For API billing, use `setup --auth-mode api-key`.
|
|
25
20
|
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
Claude Code arguments pass through:
|
|
21
|
+
Configuration is saved privately at `~/.config/claude-autorouter/config.json`. Environment variables override it; project `.env` files are not loaded automatically. `setup --force` updates an existing configuration while preserving other settings. Use focused commands for later edits:
|
|
29
22
|
|
|
30
23
|
```sh
|
|
31
|
-
claude-autorouter
|
|
32
|
-
claude-autorouter
|
|
33
|
-
claude-autorouter
|
|
24
|
+
claude-autorouter config show
|
|
25
|
+
claude-autorouter config set AUTOROUTER_JEV_TIMEOUT_MS 2000
|
|
26
|
+
claude-autorouter config unset AUTOROUTER_JEV_TIMEOUT_MS
|
|
27
|
+
claude-autorouter help config
|
|
34
28
|
```
|
|
35
29
|
|
|
36
|
-
|
|
30
|
+
Secret updates use a hidden prompt or `--stdin`, never a command-line value. Claude arguments pass through, including `claude-autorouter claude --help`. [Configuration reference](docs/reference.md#configuration).
|
|
37
31
|
|
|
38
|
-
|
|
32
|
+
## Auto permission mode
|
|
39
33
|
|
|
40
34
|
```sh
|
|
41
35
|
claude-autorouter claude --permission-mode auto
|
|
42
36
|
```
|
|
43
37
|
|
|
44
|
-
This
|
|
38
|
+
This profile automatically switches between Sonnet 5.5 and Opus 5.5 for new human tasks. A Haiku verdict uses Sonnet. Tool and `/goal` continuations retain the task's execution model; a new task can switch up or down. Claude's native safety review and organization policies still apply. For Auto selected through Claude's UI, save `AUTOROUTER_CLIENT_PROFILE=auto` with `config set`. [Auto support and limitations](docs/reference.md#auto-permission-mode).
|
|
45
39
|
|
|
46
|
-
##
|
|
40
|
+
## Inspect decisions
|
|
47
41
|
|
|
48
|
-
The launcher adds a temporary status line
|
|
42
|
+
The launcher adds a temporary status line, preserving saved Claude settings:
|
|
49
43
|
|
|
50
44
|
```text
|
|
51
|
-
● AutoRouter ·
|
|
52
|
-
● AutoRouter · Sonnet 5 selected ·
|
|
45
|
+
● AutoRouter · Opus 5.5 · ready · Jev 210ms
|
|
46
|
+
● AutoRouter · Sonnet 5.5 selected · Auto floor from Haiku · Jev 220ms
|
|
53
47
|
```
|
|
54
48
|
|
|
55
|
-
|
|
49
|
+
`selected` means Anthropic has not reported the serving model yet. Guard reasons and errors stay visible before optional savings. Claude's own model label may show its starting model. [Status details](docs/reference.md#status-line-and-savings).
|
|
56
50
|
|
|
57
|
-
|
|
51
|
+
Persistent history is optional and disabled by default. Enable metadata-only records without prompt excerpts:
|
|
58
52
|
|
|
59
53
|
```sh
|
|
60
|
-
|
|
61
|
-
|
|
54
|
+
claude-autorouter config set AUTOROUTER_SESSION_LOG_MODE metadata
|
|
55
|
+
claude-autorouter config set AUTOROUTER_SESSION_LOG_DIR "$HOME/.local/state/claude-autorouter/sessions"
|
|
56
|
+
claude-autorouter claude
|
|
57
|
+
claude-autorouter sessions list
|
|
58
|
+
claude-autorouter sessions show ID --json
|
|
62
59
|
```
|
|
63
60
|
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
Savings are an **API-equivalent estimate for the same token counts**, using Opus as the baseline. They do not measure subscription bill reductions or quota credits and exclude Jev and local compute costs. [Status line and savings details](docs/reference.md#status-line-and-savings).
|
|
67
|
-
|
|
68
|
-
## Experimental local evaluator
|
|
69
|
-
|
|
70
|
-
The local setup below requires AutoRouter 0.3.2 or newer. It uses Ollama's native `/v1/systemone` decision API with `nimble:9b-q4_K_M` by default. Jev remains the default evaluator. If upgrading from 0.2.0, replace the old Qwen model configuration using the [migration steps](docs/reference.md#migrating-an-older-ollama-config).
|
|
61
|
+
Copy an `id` from `list`. History separates model decisions from outcomes and reports latency, fallbacks, failures and savings coverage. Choose `prompts` mode for bounded human-task excerpts. Files persist locally; no automatic deletion occurs. [History and privacy](docs/reference.md#session-decision-logs).
|
|
71
62
|
|
|
72
|
-
|
|
63
|
+
Savings are **API-equivalent estimates using the recorded Opus baseline and token counts**. They do not measure subscription bill reductions or quota credits, and exclude evaluator and local compute costs. Missing or unsupported usage stays unpriced.
|
|
73
64
|
|
|
74
|
-
|
|
65
|
+
## Local Ollama evaluator
|
|
75
66
|
|
|
76
|
-
|
|
77
|
-
AUTOROUTER_OLLAMA_TIMEOUT_MS=0 claude-autorouter claude
|
|
78
|
-
```
|
|
79
|
-
|
|
80
|
-
To save that setting for an installed Tev1 4B model:
|
|
67
|
+
Start Ollama 0.35+ with a model supporting its native decision endpoint, then configure it explicitly:
|
|
81
68
|
|
|
82
69
|
```sh
|
|
83
|
-
claude-autorouter setup --evaluator ollama --ollama-model tev1:4b --
|
|
70
|
+
claude-autorouter setup --evaluator ollama --ollama-model tev1:4b-q4_K_M --pull --force
|
|
71
|
+
claude-autorouter doctor --evaluate-local
|
|
72
|
+
claude-autorouter claude
|
|
84
73
|
```
|
|
85
74
|
|
|
86
|
-
|
|
75
|
+
`--pull` authorizes downloading the chosen model if missing. Setup keeps existing models and settings; ordinary launches download nothing. The default local model is `nimble:9b-q4_K_M`; `tev1:0.8b` is smaller and requires checking its accuracy on your tasks. Local classification needs no Jev key. Claude still answers through Anthropic. [Model choices, deadlines and historical measurements](docs/reference.md#ollama-evaluator).
|
|
87
76
|
|
|
88
|
-
|
|
77
|
+
To allow a slower local model to finish without AutoRouter's runtime deadline:
|
|
89
78
|
|
|
90
79
|
```sh
|
|
91
|
-
claude-autorouter
|
|
92
|
-
claude-autorouter doctor
|
|
93
|
-
claude-autorouter claude
|
|
80
|
+
claude-autorouter config set AUTOROUTER_OLLAMA_TIMEOUT_MS 0
|
|
94
81
|
```
|
|
95
82
|
|
|
96
|
-
|
|
83
|
+
Cancellation and response-size limits still apply. The local diagnostic uses synthetic prompts and reports observed latency, classification and fallback reasons; it makes no Anthropic/Jev calls or downloads and preserves unrelated resident models.
|
|
97
84
|
|
|
98
|
-
|
|
99
|
-
| --- | ---: | --- |
|
|
100
|
-
| [Nimble 9B Q4_K_M](https://ollama.com/library/nimble) | 5.63 GB | Default: `nimble:9b-q4_K_M` |
|
|
101
|
-
| [Tev1 0.8B Q8](https://ollama.com/library/tev1) | 812 MB | `tev1:0.8b` |
|
|
102
|
-
| [Tev1 4B Q4_K_M](https://ollama.com/library/tev1) | 2.7 GB | `tev1:4b-q4_K_M` |
|
|
85
|
+
## Troubleshoot and upgrade
|
|
103
86
|
|
|
104
|
-
|
|
87
|
+
Inspect `config show` for environment overrides, and `doctor` for setup health. A valid evaluator verdict can be overridden by continuity, context or model compatibility. `Ollama fallback: timeout` means evaluation failed to finish, rather than predicting Sonnet. Status errors and saved history explain these paths.
|
|
105
88
|
|
|
106
89
|
```sh
|
|
107
|
-
|
|
90
|
+
npm install -g claude-autorouter@latest
|
|
91
|
+
claude-autorouter --version
|
|
92
|
+
claude-autorouter doctor
|
|
108
93
|
```
|
|
109
94
|
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
No Jev key is needed for local classification. The launcher primes the evaluator before opening Claude's UI, and evaluation failures fall back to Sonnet or retain Opus without contacting Jev. Claude still answers through Anthropic, with the same routing guards and subscription limits.
|
|
113
|
-
|
|
114
|
-
Setup, doctor, and startup show the effective model and deadline. Warmup and doctor do not certify classification speed or accuracy. `Ollama fallback: timeout` means no valid decision arrived in time; it is different from the evaluator choosing Sonnet. Source users can run the [local routing regression](docs/development.md#local-routing-regression) to check all three tiers without Claude or Jev calls.
|
|
115
|
-
|
|
116
|
-
Historical measurements before 0.3.2, on a 16 GiB M4: Tev1 0.8B matched 18/24 held-out labels with 450 ms median latency and no timeouts at 1,500 ms, including full-excerpt checks. Tev1 4B matched 22/24 with a 10-second diagnostic deadline and 3.15-second median latency. Nimble matched 23/24 with a 30-second deadline and 11.4-second median latency. Both larger models exceeded the then-default 1,500 ms. A separate six-case regression with the 0.3.2 fixes passed for both tested Tev1 4B variants and Nimble; Tev1 0.8B matched only three cases. These small tests do not establish general accuracy or Jev parity. See the [measurements and limits](docs/ollama-evaluation.md) and [Ollama reference](docs/reference.md#ollama-evaluator).
|
|
117
|
-
|
|
118
|
-
## Behavior and data
|
|
95
|
+
Historical integration observations cover Claude Code 2.1.284–2.1.285. The versioned synthetic protocol fixtures test reviewed request/response contracts; they do not certify the current checkout against a live Claude version. Real-provider checks remain explicitly invoked. [Troubleshooting](docs/reference.md#troubleshooting) covers context use, blocked goals and logging. Run ordinary `claude` to bypass routing.
|
|
119
96
|
|
|
120
|
-
|
|
121
|
-
- The selected evaluator receives bounded excerpts that can contain source code and tool results: TypeSafe with Jev, or the local service with Ollama. Jev also receives system-text excerpts; the local path excludes Claude's executor system instructions. Anthropic receives the complete request. Images, document payloads, and private thinking are omitted from classifier input. [Data flow and authentication](docs/reference.md#data-flow-and-authentication).
|
|
122
|
-
- Subscription access and usage limits still apply. Model switches can reduce cache reuse; cheaper token prices do not guarantee cheaper completed tasks. Run ordinary `claude` to bypass routing.
|
|
123
|
-
- The launcher is quiet by default. Use `AUTOROUTER_DEBUG=1` for metadata diagnostics or `AUTOROUTER_STATUSLINE=0` to retain your existing status line. [Troubleshooting](docs/reference.md#troubleshooting).
|
|
124
|
-
- For blocked `/goal` loops, optionally launch with `env CLAUDE_CODE_STOP_HOOK_BLOCK_CAP=2 claude-autorouter claude`. Claude then ends the turn on the third consecutive blocking verdict without tool use, leaving the goal unmet. This also affects other Stop/SubagentStop hooks; defaults are unchanged. [Scope and saved configuration](docs/reference.md#shorter-stop-hook-loops-opt-in).
|
|
97
|
+
The evaluator receives bounded task/history excerpts that may contain code and tool results: TypeSafe for Jev, or your loopback Ollama service. Anthropic receives the complete request. Model switching can reduce cache reuse. [Data flow and authentication](docs/reference.md#data-flow-and-authentication).
|
|
125
98
|
|
|
126
|
-
[Reference](docs/reference.md) · [
|
|
99
|
+
[Reference](docs/reference.md) · [Contributing](CONTRIBUTING.md) · [Development](docs/development.md) · [Releases](docs/releasing.md) · [Apache-2.0](LICENSE)
|
package/bin/autorouter.mjs
CHANGED
|
@@ -12,66 +12,48 @@ import { createSessionLog } from '../src/session-log.mjs';
|
|
|
12
12
|
import { loadUserConfig } from '../src/user-config.mjs';
|
|
13
13
|
import { setup, doctor, ollamaDeadlineText } from '../src/onboarding.mjs';
|
|
14
14
|
import { setupOllama } from '../src/ollama-setup.mjs';
|
|
15
|
+
import { helpText } from '../src/cli-help.mjs';
|
|
16
|
+
import { configCommand } from '../src/config-command.mjs';
|
|
17
|
+
import { sessionsCommand } from '../src/session-history.mjs';
|
|
15
18
|
|
|
16
19
|
const [command = 'help', ...args] = process.argv.slice(2);
|
|
17
20
|
if (['--version', '-v', 'version'].includes(command)) {
|
|
18
21
|
console.log(JSON.parse(readFileSync(new URL('../package.json', import.meta.url), 'utf8')).version);
|
|
19
22
|
} else if (['help', '--help', '-h'].includes(command)
|
|
20
|
-
|| (['setup', 'doctor', 'serve'].includes(command) && args.some(arg => ['--help', '-h'].includes(arg)))) {
|
|
21
|
-
console.log(
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
claude-autorouter --version
|
|
34
|
-
|
|
35
|
-
Setup defaults to subscription authentication and prompts for keys without echoing.
|
|
36
|
-
For noninteractive setup, supply keys through environment variables.
|
|
37
|
-
User config: ~/.config/claude-autorouter/config.json (or XDG_CONFIG_HOME).
|
|
38
|
-
AUTOROUTER_CONFIG selects a different file; environment variables take precedence.
|
|
39
|
-
Project .env files are never loaded automatically.
|
|
40
|
-
|
|
41
|
-
Jev is the default evaluator and requires TYPESAFE_API_KEY.
|
|
42
|
-
Ollama evaluates locally and requires Ollama 0.35+ with /v1/systemone.
|
|
43
|
-
Use setup --evaluator ollama --pull to detect Ollama and download a missing model.
|
|
44
|
-
The local default is nimble:9b-q4_K_M; --ollama-model selects another compatible model.
|
|
45
|
-
Smaller Tev1 options: --ollama-model tev1:0.8b or --ollama-model tev1:4b-q4_K_M.
|
|
46
|
-
Local routing deadlines: Tev1 0.8B/custom 1500 ms, Tev1 4B 15000 ms, Nimble 30000 ms.
|
|
47
|
-
Setup --ollama-timeout-ms N saves a routing deadline; use 0 to disable it.
|
|
48
|
-
AUTOROUTER_OLLAMA_TIMEOUT_MS also overrides the deadline; 0 disables it.
|
|
49
|
-
Local routing is experimental; see docs/ollama-evaluation.md for measured limits.
|
|
50
|
-
AUTOROUTER_AUTH_MODE=subscription uses your saved Claude Code login.
|
|
51
|
-
Without setup, AUTOROUTER_AUTH_MODE defaults to api-key and also requires ANTHROPIC_API_KEY.
|
|
52
|
-
AUTOROUTER_CLIENT_PROFILE=compatible (default) enables all three routing tiers.
|
|
53
|
-
Use AUTOROUTER_CLIENT_PROFILE=native to retain Claude Code's own model/thinking settings.
|
|
54
|
-
Use AUTOROUTER_CLIENT_PROFILE=auto for Auto permission mode: Sonnet/Opus routing, native thinking.
|
|
55
|
-
Auto defaults to Sonnet 5.5 and Opus 5.5, switching on new human tasks and retaining tool turns.
|
|
56
|
-
An explicit claude --permission-mode auto selects the auto profile for that launch.
|
|
57
|
-
Claude's permission checks and organization policies still apply; Haiku does not support Auto mode.
|
|
58
|
-
Optional CLAUDE_CODE_STOP_HOOK_BLOCK_CAP=N limits consecutive tool-free Stop-hook continuations.
|
|
59
|
-
Use 2 to stop on the third block; applies to /goal and all Stop/SubagentStop hooks.
|
|
60
|
-
Unset preserves Claude's default; 0 disables the cap. Setup --stop-hook-block-cap N saves it.
|
|
61
|
-
Standalone serve also requires AUTOROUTER_TOKEN (at least 16 characters).
|
|
62
|
-
The claude launcher creates a temporary credential and an ephemeral port.
|
|
63
|
-
It enables an AutoRouter status line for this session (AUTOROUTER_STATUSLINE=0 to opt out).
|
|
64
|
-
Launcher logs are quiet by default; AUTOROUTER_DEBUG=1 enables diagnostic logs on stderr.
|
|
65
|
-
AUTOROUTER_SESSION_LOG_DIR writes private per-session JSONL decision logs with prompt excerpts.
|
|
66
|
-
Unset or empty disables session logs. Setup --session-log-dir DIR saves the directory.
|
|
67
|
-
Jev sends prompt excerpts to TypeSafe; Ollama keeps classification on this machine.
|
|
68
|
-
Complete inference requests still go to Anthropic. See README.md.`);
|
|
69
|
-
} else if (command === 'setup' || command === 'doctor') {
|
|
23
|
+
|| (['setup', 'doctor', 'serve', 'config', 'sessions'].includes(command) && args.some(arg => ['--help', '-h'].includes(arg)))) {
|
|
24
|
+
console.log(helpText(command === 'help' ? args[0] : command));
|
|
25
|
+
} else if (command === 'claude' && args.length === 1 && ['--help', '-h', '--version', '-v'].includes(args[0])) {
|
|
26
|
+
// Help/version are Claude-owned commands. No config file, evaluator keys,
|
|
27
|
+
// gateway, Ollama warmup or temporary status files are needed.
|
|
28
|
+
const env = { ...process.env };
|
|
29
|
+
delete env.TYPESAFE_API_KEY;
|
|
30
|
+
delete env.AUTOROUTER_TOKEN;
|
|
31
|
+
const child = spawn('claude', args, { stdio: 'inherit', env });
|
|
32
|
+
child.once('error', () => { console.error('Could not launch Claude Code. Ensure `claude` is installed and on PATH.'); process.exitCode = 1; });
|
|
33
|
+
child.once('exit', (code, signal) => { process.exitCode = code ?? (signal === 'SIGINT' ? 130 : signal === 'SIGTERM' ? 143 : 1); });
|
|
34
|
+
for (const signal of ['SIGINT', 'SIGTERM']) process.on(signal, () => child.kill(signal));
|
|
35
|
+
} else if (['setup', 'doctor', 'config', 'sessions'].includes(command)) {
|
|
70
36
|
try {
|
|
71
37
|
if (command === 'setup') await setup(args);
|
|
38
|
+
else if (command === 'config') {
|
|
39
|
+
if (await configCommand(args) === false) process.exitCode = 1;
|
|
40
|
+
}
|
|
41
|
+
else if (command === 'sessions') {
|
|
42
|
+
if (await sessionsCommand(args) === false) process.exitCode = 1;
|
|
43
|
+
}
|
|
72
44
|
else {
|
|
73
|
-
|
|
74
|
-
|
|
45
|
+
const evaluateLocal = args.includes('--evaluate-local');
|
|
46
|
+
const json = args.includes('--json');
|
|
47
|
+
if (args.some(arg => !['--evaluate-local', '--json'].includes(arg)) || new Set(args).size !== args.length
|
|
48
|
+
|| (json && !evaluateLocal)) throw new Error('Usage: claude-autorouter doctor [--evaluate-local [--json]]');
|
|
49
|
+
const controller = new AbortController();
|
|
50
|
+
const cancel = () => controller.abort();
|
|
51
|
+
for (const signal of ['SIGINT', 'SIGTERM']) process.once(signal, cancel);
|
|
52
|
+
try {
|
|
53
|
+
if (!await doctor({ evaluateLocal, json, signal: controller.signal })) process.exitCode = 1;
|
|
54
|
+
} finally {
|
|
55
|
+
for (const signal of ['SIGINT', 'SIGTERM']) process.removeListener(signal, cancel);
|
|
56
|
+
}
|
|
75
57
|
}
|
|
76
58
|
} catch (error) { console.error(error.message); process.exitCode = 1; }
|
|
77
59
|
} else if (!['claude', 'serve'].includes(command)) {
|
|
@@ -85,10 +67,9 @@ Complete inference requests still go to Anthropic. See README.md.`);
|
|
|
85
67
|
const stop = () => {
|
|
86
68
|
if (stopping) return stopping;
|
|
87
69
|
if (server) { server.close(); server.closeAllConnections(); }
|
|
88
|
-
status?.close();
|
|
89
70
|
// Drain accepted decision records before normal process exit. Pending
|
|
90
71
|
// filesystem writes keep Node alive; no timer or fire-and-forget buffer.
|
|
91
|
-
stopping = Promise.
|
|
72
|
+
stopping = Promise.allSettled([status?.close(), sessionLog?.close()]);
|
|
92
73
|
return stopping;
|
|
93
74
|
};
|
|
94
75
|
try {
|
|
@@ -120,15 +101,17 @@ Complete inference requests still go to Anthropic. See README.md.`);
|
|
|
120
101
|
let claudeArgs = args;
|
|
121
102
|
if (statusEnabled) {
|
|
122
103
|
status = createStatusState({ baselineModel: config.models.opus });
|
|
104
|
+
await status.ready;
|
|
123
105
|
if (status.path) {
|
|
124
106
|
try { claudeArgs = addStatusLineSettings(args, dirname(status.path)); }
|
|
125
107
|
catch {
|
|
126
|
-
status.close(); status = undefined;
|
|
108
|
+
await status.close(); status = undefined;
|
|
127
109
|
console.error('AutoRouter status line unavailable: could not safely prepare session settings. Passing your original settings to Claude.');
|
|
128
110
|
}
|
|
129
111
|
} else console.error('AutoRouter status line unavailable: could not create local status storage.');
|
|
130
112
|
}
|
|
131
113
|
if (config.sessionLogDir) sessionLog = await createSessionLog(config.sessionLogDir, {
|
|
114
|
+
includePrompts: config.sessionLogMode === 'prompts',
|
|
132
115
|
warn: message => console.error(message),
|
|
133
116
|
});
|
|
134
117
|
// Claude owns the terminal while its UI is running. Status updates use the
|
|
@@ -136,7 +119,7 @@ Complete inference requests still go to Anthropic. See README.md.`);
|
|
|
136
119
|
server = createRouterServer(config, {
|
|
137
120
|
log: diagnosticLogs ? undefined : () => {},
|
|
138
121
|
onStatus: event => status?.update(event),
|
|
139
|
-
|
|
122
|
+
onRecord: sessionLog ? entry => sessionLog.record(entry) : undefined,
|
|
140
123
|
});
|
|
141
124
|
const address = await listen(server, command === 'claude' ? 0 : config.port);
|
|
142
125
|
const baseUrl = `http://127.0.0.1:${address.port}`;
|
|
@@ -155,7 +138,7 @@ Complete inference requests still go to Anthropic. See README.md.`);
|
|
|
155
138
|
if (status?.path) env.AUTOROUTER_STATUS_FILE = status.path;
|
|
156
139
|
const child = spawn('claude', claudeArgs, { stdio: 'inherit', env });
|
|
157
140
|
child.once('error', () => { console.error('Could not launch Claude Code. Ensure `claude` is installed and on PATH.'); process.exitCode = 1; stop(); });
|
|
158
|
-
child.once('exit', (code, signal) => { process.exitCode = code ?? (signal === 'SIGINT' ? 130 : 1); stop(); });
|
|
141
|
+
child.once('exit', (code, signal) => { process.exitCode = code ?? (signal === 'SIGINT' ? 130 : signal === 'SIGTERM' ? 143 : 1); stop(); });
|
|
159
142
|
for (const signal of ['SIGINT', 'SIGTERM']) process.on(signal, () => child.kill(signal));
|
|
160
143
|
}
|
|
161
144
|
} catch (error) {
|
package/docs/development.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Development and validation
|
|
2
2
|
|
|
3
|
-
Use Node.js 22+ from a source checkout on macOS or Linux (including WSL). The
|
|
3
|
+
Use Node.js 22+ from a source checkout on macOS or Linux (including WSL). The installed CLI has no runtime package dependencies. Source checks use pinned TypeScript and Node type definitions; install these contributor tools with `npm ci --ignore-scripts --no-audit --no-fund`. Development scripts and tests are separate from the installed CLI; user setup is covered in the [README](../README.md).
|
|
4
4
|
|
|
5
5
|
## Local checks
|
|
6
6
|
|
|
@@ -16,6 +16,24 @@ Auto-mode regressions exercise Sonnet → Opus → Opus tool continuation → So
|
|
|
16
16
|
|
|
17
17
|
Package validation checks the distributable and installed command rather than relying on the source checkout's paths. Review the [release procedure](releasing.md) before distributing a tarball.
|
|
18
18
|
|
|
19
|
+
## Request lifecycle and model continuity
|
|
20
|
+
|
|
21
|
+
The gateway validates the request containers it consumes, evaluates the current task, applies continuity and capacity rules, then checks the proposed target against the shared model catalog. Unknown provider extensions remain intact; when their compatibility with another model is unknown, the source model is retained with a routing reason. Token-count requests apply the same compatibility rules before sending a request.
|
|
22
|
+
|
|
23
|
+
`src/model-catalog.mjs` contains exact model IDs, capability facts, source links, and a review date. A family name inside a custom alias does not establish capabilities. `src/model-request.mjs` contains explicit thinking adaptations; neither layer strips signed history or permission-review settings. Model updates should change the catalog, include a dated authoritative source, and add a request fixture demonstrating the restriction or new supported switch.
|
|
24
|
+
|
|
25
|
+
The evaluation cache and active execution state have separate lifetimes. `src/turn-state.mjs` keeps active human tasks and pending tools beyond the evaluator cache TTL. Retired tasks expire, aliases and records are bounded, and exhausting active-state capacity is reported rather than silently evicting another active task. State is process-local; after a gateway restart, missing continuity is reported as unknown until a successful response establishes it again.
|
|
26
|
+
|
|
27
|
+
The gateway stages a selected model under its request ID. `src/response-observer.mjs` observes serving models, provider fallback boundaries, closed tool calls, usage, and terminal response metadata without altering bytes. A serving-model observation alone is not a successful execution. Only clean completion evidence followed by successful HTTP forwarding commits the continuation model. Cancelled, failed, ambiguous, and superseded attempts cannot overwrite known state. New human tasks remain eligible for upward or downward switching, including Sonnet/Opus in Auto mode.
|
|
28
|
+
|
|
29
|
+
Direct `Router.route()` embedders that omit `requestId` retain selected, unconfirmed continuity for compatibility. Embedders that execute inference should supply a unique request ID and call `router.complete(id, evidence)` after successful delivery, or `router.complete(id)` on failure. The HTTP gateway owns that lifecycle automatically.
|
|
30
|
+
|
|
31
|
+
## Evaluation acceptance
|
|
32
|
+
|
|
33
|
+
Evaluation reports distinguish evaluator availability, rubric agreement, routing policy, tier coverage, transport, and independently checked task completion. An unmeasured gate is explicitly marked unmeasured. Normal runs cannot pass solely on classifier fallback or cached predictions; simulated-outage runs explicitly require fallback. Compatible routing requires all three selected tiers, while Auto requires Sonnet and Opus. Constrained fixtures declare expected guard overrides.
|
|
34
|
+
|
|
35
|
+
The general and Ollama evaluation scripts accept `--min-agreement` and `--max-under-route-rate`. Their defaults require complete expected-label agreement and no under-routing. Set any alternative thresholds **before** evaluating a candidate, retain the fixture checksum with the report, and keep tuning cases separate from held-out cases. Rubric labels are judgments about synthetic tasks; these reports do not prove end-user task quality or subscription savings. The expanded corpus includes multilingual tasks, short difficult follow-ups, ordinary work in long background context, and task text containing tier-selection instructions. No prompt-policy adjustment should be justified by rerunning and relabeling the held-out set.
|
|
36
|
+
|
|
19
37
|
## Run from source
|
|
20
38
|
|
|
21
39
|
```sh
|
|
@@ -76,7 +94,7 @@ For a classifier-only rubric evaluation:
|
|
|
76
94
|
npm run eval
|
|
77
95
|
```
|
|
78
96
|
|
|
79
|
-
The bundled evaluation makes
|
|
97
|
+
The bundled evaluation makes 19 classifier calls and no Claude generations. Jev is the default and incurs TypeSafe usage; set `AUTOROUTER_EVALUATOR=ollama` to evaluate an installed local model. It reports agreement with the starting rubric, fallback count, and p50/p95 routing latency. Edit `test/fixtures/routing.json` to represent the tasks you want to measure. Rubric agreement alone does not establish answer quality or net savings; compare completed tasks against fixed-model baselines.
|
|
80
98
|
|
|
81
99
|
For local evaluator measurements, use Ollama 0.35+ and a model compatible with `/v1/systemone`. Distinguish cold model loading from warmed classification, and record the model tag, hardware, Ollama version, context size, prompt length, and resident memory. The launcher primes the classifier with a synthetic task before opening the UI, with a separate deadline of up to 60 seconds. Runtime and benchmark share the 3,000-character/3,000-UTF-8-byte state limit, so include non-ASCII cases and excerpts that fill the budget. Also measure the first request after keep-alive expiration: its reload can hit the normal deadline even when warm requests pass. Repeat on realistic prompt distributions instead of selecting a model from a single easy request. Disk download size is not resident RAM. Keep model downloads opt-in and respect each model's license.
|
|
82
100
|
|
|
@@ -113,3 +131,31 @@ node --env-file=.env scripts/context-probe.mjs --cwd /path/to/synthetic-fixture
|
|
|
113
131
|
```
|
|
114
132
|
|
|
115
133
|
Private repository or connected-tool context may be present even when the typed prompt is harmless. Keep private-payload investigations local unless external processing is authorized. For shareable live regressions, prefer the isolated synthetic fixtures above. The probe report itself persists only metadata.
|
|
134
|
+
|
|
135
|
+
## Static contracts and style
|
|
136
|
+
|
|
137
|
+
`npm run check` checks JavaScript syntax, TypeScript/JSDoc contracts for configuration, classifier results, final routing decisions and normalized telemetry, and literal event producers in transport code. `src/contracts.mjs` is the shared development-time type contract; normalizers remain the runtime privacy boundary. Negative fixtures in `test/static-contracts.mts` and `test/static-checks.test.mjs` prove misspelled fields, invalid enums, payload fields and timing strings are rejected before execution. Provider request extensions remain opaque and are validated only where the router consumes them.
|
|
138
|
+
|
|
139
|
+
Use two-space indentation, LF endings, one final newline, semicolons and single quotes for ordinary strings. Compact pure helpers are allowed when readable; do not reformat unrelated code. The style check rejects trailing whitespace, tab indentation, `var`, and coercing comparisons except deliberate null/undefined checks. TypeScript is a contributor dependency only; public packages keep zero runtime dependencies. CI installs the pinned lockfile before checks and never runs provider inference automatically.
|
|
140
|
+
|
|
141
|
+
## Versioned protocol evidence
|
|
142
|
+
|
|
143
|
+
`test/fixtures/claude-protocol-v1.json` is a versioned, newly authored synthetic corpus reviewed against the Messages API, streaming, deferred-tool and fallback contracts. `test/protocol-fixtures.test.mjs` sends it through the real gateway, router and response observer with fake evaluator/upstream services. It covers Auto floors and switches, thinking adaptation and preservation, custom deferred tool references, conservative built-in server-tool history, compaction, scoped goal feedback, parallel agents, model fallback/tool ownership, usage and truncated responses.
|
|
144
|
+
|
|
145
|
+
Historical Claude Code 2.1.284/2.1.285 report versions, dates and hashes are separate metadata. The corpus does not copy captured prompts, invent provider signatures or certify a live client. Current-source real-provider canaries remain opt-in and unmeasured unless a separate report records them.
|
|
146
|
+
|
|
147
|
+
## Performance regression measurements
|
|
148
|
+
|
|
149
|
+
[Router measurements](router-performance.md) and [status-storage measurements](status-performance.md) record the pre-change baseline, repeated candidate runs, hardware/background load and baseline-derived gates. Source-only harnesses use synthetic inputs and providers; no credentials, prompts from user sessions or downloads are involved.
|
|
150
|
+
|
|
151
|
+
```sh
|
|
152
|
+
node --expose-gc scripts/benchmark-router.mjs baseline /tmp/router-comparison.json
|
|
153
|
+
node --expose-gc scripts/benchmark-router.mjs candidate /tmp/router-comparison.json --check
|
|
154
|
+
node scripts/benchmark-status.mjs --label local-check --check
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
Identical concurrent classifier inputs share one bounded evaluation. Each request applies its own continuity, capacity and compatibility checks. Cancelling one waiter preserves other waiters; cancelling all releases the shared evaluation. Cache identity retains the complete request, requested model floor, evaluator configuration and rubric hash. New tasks and sequential pinned continuations still evaluate; prior-pin classification reuse was deliberately not enabled. A 64 KiB response limit applies to both evaluators, and local model metadata is limited to 1 MiB. Disabling the Ollama timer does not disable cancellation or these byte limits.
|
|
158
|
+
|
|
159
|
+
Status writes use one asynchronous writer and a coalesced latest snapshot. Embedders await `state.ready` before using its path, `flush()` when they need persisted evidence, and `close()` before cleanup. Initial storage failures or a one-second readiness timeout disable the optional display. Accepted in-flight writes finish before directory removal, so shutdown cannot recreate files.
|
|
160
|
+
|
|
161
|
+
The router/storage measurements cover one 16 GiB M4 with synthetic providers. Actual Tev1 4B and Nimble 9B measurements on that Mac and a 64 GiB M2 Ultra are recorded separately in [the hardware comparison](hardware-comparison.md). Both candidates missed the unchanged strict quality gate on both hosts. Model digests and runtime conditions differ, so the cross-host results do not isolate RAM's effect. Use the [transfer bundle instructions](hardware-benchmark.md) to reproduce the workload, recording background load and observed residency. Do not infer model performance from router timings.
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
# Local-model hardware validation
|
|
2
|
+
|
|
3
|
+
Actual Tev1 and Nimble evaluators have now been measured on a 16 GiB M4 and a 64 GiB M2 Ultra; see [the comparison and limitations](hardware-comparison.md). These instructions reproduce the opt-in benchmark on another Mac. It uses only checked-in synthetic tasks; no Jev or Anthropic calls, user configuration, user prompts or model downloads are involved.
|
|
4
|
+
|
|
5
|
+
Transfer `autorouter-hardware-benchmark.tar.gz` and its checksum to the other Mac. The bundle contains an explicit source allowlist and hashes; it excludes credentials, configurations, histories, git metadata and dependencies. Extract it:
|
|
6
|
+
|
|
7
|
+
```sh
|
|
8
|
+
shasum -a 256 -c autorouter-hardware-benchmark.tar.gz.sha256
|
|
9
|
+
tar -xzf autorouter-hardware-benchmark.tar.gz
|
|
10
|
+
cd autorouter-benchmark
|
|
11
|
+
node --version
|
|
12
|
+
node scripts/evaluate-ollama.mjs --help
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
Node 22+ and a running Ollama 0.35+ are required. No npm install is needed. Use already-installed model tags. Start when Ollama has no resident models; the benchmark refuses to displace unrelated resident models. It loads each selected candidate for a separate cold observation and repeated warm evaluations, then unloads that candidate from memory. It never deletes downloaded models. Keep other applications running at a representative load; CPU load averages and free memory are recorded, without process names or arguments.
|
|
16
|
+
|
|
17
|
+
Use the same two model tags and 30-second warm deadline as the 16 GiB baseline:
|
|
18
|
+
|
|
19
|
+
```sh
|
|
20
|
+
node scripts/evaluate-ollama.mjs --models tev1:4b-q4_K_M,nimble:9b-q4_K_M --split heldout --rounds 3 --stress-rounds 8 --timeout-ms 30000 --cold-timeout-ms 60000 --output hardware-report.json
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
For an installed Tev1 model, substitute its exact tag, such as `tev1:4b-q4_K_M`. Multiple installed candidates can be comma-separated. Use `--timeout-ms 0` to measure without the runtime timer; initial loading still has the separate cold deadline. Measure finite-deadline reliability in a separate report as well. Do not change label-agreement thresholds after seeing results: the default requires complete expected-label agreement, no under-routing and all three classified tiers. A nonzero exit can indicate a quality-gate failure while the JSON still contains valid measurements.
|
|
24
|
+
|
|
25
|
+
For a controlled comparison, match model artifact digests as well as tags, quantization, context allocation and runtime versions. The existing cross-host reports have different model digests and other conditions, so they are observational measurements rather than a RAM-only comparison.
|
|
26
|
+
|
|
27
|
+
The report contains hardware/Node/Ollama versions, initial/final background load/free memory, fixture and rubric hashes, model identity/quantization/resident memory, cold and repeated warm latency, full-budget stress cases, deadline failures and classification gates. Classification quality is agreement with the frozen synthetic rubric; selected and confirmed Claude models remain unmeasured. It does not prove end-user task quality or net savings.
|
|
28
|
+
|
|
29
|
+
Return `hardware-report.json` for comparison with the same workload on the 16 GiB Mac. Keep the bundle's `source-manifest.json` with it. Record whether the machine was in normal use or deliberately idle. A missing model or quality failure is reported explicitly rather than hidden by choosing a passing model or relabeling cases.
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# Local evaluator comparison: 16 GiB M4 and 64 GiB M2 Ultra
|
|
2
|
+
|
|
3
|
+
The October 5, 2026 measurements now cover both required memory classes. The returned 64 GiB report contains valid cold, warm and full-budget observations from a Mac in normal use with the usual applications open. Both candidates completed every request there, but both still failed the unchanged classification gate. Jev remains the default; these local measurements do not establish production routing accuracy.
|
|
4
|
+
|
|
5
|
+
The frozen workload is one cold request, three deterministically shuffled rounds of 30 held-out synthetic tasks, and eight separate 3,000-byte stress tasks per model. Warm and cold deadlines were 30 and 60 seconds respectively. The benchmark uses native `/v1/systemone` choice scoring, with no Claude generations, Jev requests or model downloads. The reported latency is end-to-end wall time, not isolated model computation.
|
|
6
|
+
|
|
7
|
+
| Measurement | Tev1 4B, 16 GiB | Tev1 4B, 64 GiB | Nimble 9B, 16 GiB | Nimble 9B, 64 GiB |
|
|
8
|
+
| --- | ---: | ---: | ---: | ---: |
|
|
9
|
+
| Cold wall time | 7.13 s | 3.57 s | 15.93 s | 3.77 s |
|
|
10
|
+
| Warm requests / valid results | 90 / 90 | 90 / 90 | 90 / 87 | 90 / 90 |
|
|
11
|
+
| Warm wall p50 / p95, including errors | 244 / 2,930 ms | 55 / 519 ms | 1,250 / 13,645 ms | 56 / 863 ms |
|
|
12
|
+
| Successful warm wall p50 / p95 | 244 / 2,930 ms | 55 / 519 ms | 1,148 / 11,249 ms | 56 / 863 ms |
|
|
13
|
+
| Exact frozen-label agreement | 93.3% | 93.3% | 93.3% | 96.7% |
|
|
14
|
+
| Under-routes / over-routes | 3 / 3 | 3 / 3 | 3 / 0 | 3 / 0 |
|
|
15
|
+
| Warm deadline errors | 0 | 0 | 3 | 0 |
|
|
16
|
+
| Full-budget stress p50 / p95 | 7,003 / 7,854 ms | 1,122 / 1,314 ms | 12,369 / 13,945 ms | 1,940 / 2,003 ms |
|
|
17
|
+
| Strict warm classification gate | Failed | Failed | Failed | Failed |
|
|
18
|
+
|
|
19
|
+
Tev1 under-routed `held-o-compiler` from Opus to Sonnet in all three rounds and over-routed `held-h-literal-opus` from Haiku to Opus. Nimble under-routed `held-o-injection` from Opus to Haiku in all three rounds. The 16 GiB Nimble run also had three timeout errors, which count against its agreement. Both models reached all three tiers and passed all eight stress labels on both hosts. Those stress tasks all expect Haiku; passing them does not establish broader accuracy. The preset warm gate remains 100% exact agreement, zero under-routing and all three classified tiers.
|
|
20
|
+
|
|
21
|
+
On the larger Mac, Tev1's stress p95 was 1.31 seconds and Nimble's was 2.00 seconds. The zero-error results were obtained with a 30-second warm deadline and do not certify reliability under a 1.5-second deadline. These observations support configurable model-specific deadlines, including the explicit disabled-deadline option. They do not determine an optimal deadline or justify changing the existing defaults.
|
|
22
|
+
|
|
23
|
+
## Conditions and comparability
|
|
24
|
+
|
|
25
|
+
| Condition | 16 GiB baseline | 64 GiB returned run |
|
|
26
|
+
| --- | --- | --- |
|
|
27
|
+
| Processor / logical CPUs | Apple M4 / 10 | Apple M2 Ultra / 24 |
|
|
28
|
+
| Unified memory | 16 GiB | 64 GiB |
|
|
29
|
+
| OS release / architecture | 25.6.0 / arm64 | 27.0.0 / arm64 |
|
|
30
|
+
| Node / Ollama | v26.10.0 / 0.35.0 | v22.14.0 / 0.35.1 |
|
|
31
|
+
| Initial 1/5/15-minute load | 3.55 / 3.35 / 3.48 | 3.87 / 5.12 / 4.93 |
|
|
32
|
+
| Final 1/5/15-minute load | 2.23 / 3.60 / 3.67 | 4.49 / 5.09 / 4.94 |
|
|
33
|
+
| Initial / final OS free memory | 842 / 122 MiB | 727 / 136 MiB |
|
|
34
|
+
| Background use | Applications running | Normal use, usual applications open |
|
|
35
|
+
|
|
36
|
+
Both hosts reported Q4_K_M quantization, 4.2B/9.0B parameters, context allocations of 2,050/8,194 tokens and model residency of 2.71/5.74 GiB for Tev1/Nimble. Residency is reported by Ollama; free memory is an OS snapshot, not total available memory or measured swap pressure. The JSON also records aggregate RSS across matching Ollama daemon/runner processes, which is not a pure model footprint. On the larger host, that aggregate rose from 2.94 to 7.99 GiB for Tev1 and 5.80 to 9.42 GiB for Nimble.
|
|
37
|
+
|
|
38
|
+
Model tags match but artifact digests differ:
|
|
39
|
+
|
|
40
|
+
| Model tag | 16 GiB digest | 64 GiB digest |
|
|
41
|
+
| --- | --- | --- |
|
|
42
|
+
| `tev1:4b-q4_K_M` | `3509ac7180e86e5fa8efc7b5745e32d55dd9d4e0a86bc9a88aba5323a5d29bc6` | `07be32e6e5e3dcc6b2deae7c68b89321e99daeedb08521e92771cc155473cbd9` |
|
|
43
|
+
| `nimble:9b-q4_K_M` | `3776806da5587387a996e28e75d5d07fbebe7410d47879e1f71782f36c896ce3` | `572f1f4c801d37171344b0520d14853a88887c249d11db3acfa6c81a34cdcfa3` |
|
|
44
|
+
|
|
45
|
+
Each returned download size is 28 bytes larger. This does not establish whether weights, templates or metadata changed. The comparison is observational: processor, runtime versions, background conditions and model artifacts differ, so faster timings cannot be attributed to RAM alone. A controlled hardware comparison would require matching model digests and runtime conditions. The 16 GiB timeout rows also recorded wall times exceeding their configured 30-second deadline, up to 969,208.59 ms; their cause was not established, so the timer should not be described as a strict wall-time upper bound in that run.
|
|
46
|
+
|
|
47
|
+
## Evidence and validation
|
|
48
|
+
|
|
49
|
+
[The 16 GiB report](hardware-results-16gb.json) and [the 64 GiB report](hardware-results-64gb.json) retain all per-case outcomes and unchanged gates. The larger report is complete while its overall `passed` field remains `false`, correctly reflecting failed classification gates. An independent offline review recomputed case/round coverage, labels, UTF-8 state sizes, percentile summaries, confusion matrices and acceptance gates.
|
|
50
|
+
|
|
51
|
+
All 36 hashes in the returned source manifest matched the frozen transfer sources when received. The fixture SHA-256 is `1ee6ab593e47d87a67e922c0afb0c9db5787d8a27d1a0ff3e24baa0543291fc5`; both model rubric hashes are `be151cedb4de4b7ef3f7162d751f70ce7d9dd14efc66fae1835f73ffd04027be`. Subsequent documentation and package-allowlist changes do not change those evaluated sources. The original transfer bundle is retained unchanged.
|
|
52
|
+
|
|
53
|
+
The returned report SHA-256 is `b76d25b3b062821fa2483e13847d79267734f1a83b35ffe6ea85c57f82f7f715`; its manifest SHA-256 is `dab377757912028fca5a929dd4ef2123a4c4001559f0d4ee26fd6bec092b04f1`. The public JSON includes this provenance and the user's background-use description; it contains synthetic case identifiers and metadata, without user prompts or process names. Reproduction instructions remain in [hardware-benchmark.md](hardware-benchmark.md).
|
|
54
|
+
|
|
55
|
+
The larger-memory measurement requirement is fulfilled. Classifier errors remain visible rather than being hidden by altered labels or thresholds. Neither this comparison nor the separate [router benchmark](router-performance.md) establishes parity with Jev, completed Claude task quality or subscription savings.
|