claude-autorouter 0.3.7 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.env.example +6 -3
- package/CODE_OF_CONDUCT.md +9 -0
- package/CONTRIBUTING.md +57 -0
- package/README.md +47 -70
- package/SECURITY.md +23 -0
- package/SUPPORT.md +18 -0
- package/bin/autorouter.mjs +40 -57
- package/docs/development.md +48 -2
- package/docs/hardware-benchmark.md +29 -0
- package/docs/hardware-comparison.md +55 -0
- package/docs/hardware-results-16gb.json +4002 -0
- package/docs/hardware-results-16gb.md +26 -0
- package/docs/hardware-results-64gb.json +4020 -0
- package/docs/reference.md +92 -37
- package/docs/releasing.md +79 -37
- package/docs/router-performance.json +1697 -0
- package/docs/router-performance.md +50 -0
- package/docs/status-performance.json +363 -0
- package/docs/status-performance.md +44 -0
- package/docs/subscription-integration.md +27 -0
- package/package.json +66 -10
- package/src/auto-routing.mjs +184 -24
- package/src/bounded-json.mjs +57 -0
- package/src/cli-help.mjs +90 -0
- package/src/config-command.mjs +158 -0
- package/src/config.mjs +53 -28
- package/src/contracts.mjs +123 -0
- package/src/evaluation-report.mjs +114 -0
- package/src/keychain.mjs +58 -0
- package/src/local-diagnostic.mjs +191 -0
- package/src/model-catalog.mjs +96 -0
- package/src/model-request.mjs +6 -7
- package/src/ollama-evaluator.mjs +9 -27
- package/src/onboarding.mjs +130 -26
- package/src/prompt-state.mjs +22 -7
- package/src/redaction.mjs +97 -0
- package/src/request-validation.mjs +54 -0
- package/src/response-observer.mjs +126 -18
- package/src/router.mjs +151 -61
- package/src/savings.mjs +74 -16
- package/src/server.mjs +79 -12
- package/src/session-history.mjs +262 -0
- package/src/session-log.mjs +9 -58
- package/src/status-state.mjs +110 -62
- package/src/statusline.mjs +57 -27
- package/src/telemetry-event.mjs +200 -0
- package/src/token-counter.mjs +3 -1
- package/src/turn-state.mjs +132 -0
- package/src/user-config.mjs +81 -10
package/.env.example
CHANGED
|
@@ -1,7 +1,8 @@
|
|
|
1
1
|
# Copy to .env and load with: node --env-file=.env bin/autorouter.mjs claude
|
|
2
2
|
AUTOROUTER_AUTH_MODE=subscription
|
|
3
3
|
AUTOROUTER_CLIENT_PROFILE=compatible
|
|
4
|
-
#
|
|
4
|
+
# Local Ollama is the default evaluator (experimental; see settings below).
|
|
5
|
+
# Set jev to use TypeSafe's hosted evaluator, which needs TYPESAFE_API_KEY.
|
|
5
6
|
AUTOROUTER_EVALUATOR=jev
|
|
6
7
|
# The launcher enables the router status line for this session. Set 0 to keep your own.
|
|
7
8
|
AUTOROUTER_STATUSLINE=1
|
|
@@ -12,10 +13,12 @@ AUTOROUTER_STATUSLINE=1
|
|
|
12
13
|
# AUTOROUTER_CLIENT_PROFILE=auto
|
|
13
14
|
# Optional metadata logs on stderr. Redirect stderr to a file when using the UI.
|
|
14
15
|
# AUTOROUTER_DEBUG=1
|
|
15
|
-
# Optional persistent JSONL
|
|
16
|
+
# Optional persistent JSONL decisions/outcomes, one file per session per launch.
|
|
16
17
|
# Includes up to 500 characters of user prompt text; keep the directory local.
|
|
17
18
|
# Unset or empty disables logging. This does not print prompts in the terminal.
|
|
18
19
|
# AUTOROUTER_SESSION_LOG_DIR=/absolute/path/to/autorouter-sessions
|
|
20
|
+
# Omit all prompt excerpt fields (does not enable logging by itself).
|
|
21
|
+
# AUTOROUTER_SESSION_LOG_MODE=metadata
|
|
19
22
|
# Optional: allow two tool-free Stop-hook continuations, then end the turn on
|
|
20
23
|
# the third block. Applies to /goal and all Stop/SubagentStop hooks.
|
|
21
24
|
# Unset keeps Claude's default (currently 8); 0 DISABLES the cap.
|
|
@@ -40,7 +43,7 @@ AUTOROUTER_MIN_CONFIDENCE=0.75
|
|
|
40
43
|
# Experimental local evaluator in AutoRouter 0.3.2+: native decision API only.
|
|
41
44
|
# Install/start Ollama 0.35+, then run:
|
|
42
45
|
# claude-autorouter setup --evaluator ollama --pull --force
|
|
43
|
-
# Defaults to nimble:9b-q4_K_M (~5.63 GB download); --force
|
|
46
|
+
# Defaults to nimble:9b-q4_K_M (~5.63 GB download); --force preserves other settings.
|
|
44
47
|
# Existing downloads are kept. Replace Qwen config/environment values from 0.2.0.
|
|
45
48
|
# All local models use /v1/systemone; custom tags/aliases must support that API.
|
|
46
49
|
# Add --ollama-model LOCAL_TAG_OR_ALIAS to setup to choose another suitable model.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# Code of conduct
|
|
2
|
+
|
|
3
|
+
Be respectful and constructive in issues, pull requests, reviews and other project spaces. Welcome people with different backgrounds and experience levels, focus criticism on ideas and code, and respect requests to stop unwanted interaction.
|
|
4
|
+
|
|
5
|
+
Harassment, discriminatory or demeaning remarks, sexualized conduct, threats, personal attacks, and publishing someone's private information without consent are not acceptable. The same expectations apply when representing the project outside its repository.
|
|
6
|
+
|
|
7
|
+
Report concerns privately to maintainer Fabio Rapposelli at [fabio@rapposelli.org](mailto:fabio@rapposelli.org), with the subject `AutoRouter conduct report`. Include relevant links and enough context to investigate; avoid publishing the report or unrelated private information. Reports will be handled with discretion, sharing information only as needed to investigate and respond.
|
|
8
|
+
|
|
9
|
+
The maintainer may request changes, remove content, issue warnings, or temporarily or permanently restrict participation according to the severity and pattern of behavior. Requests to review a decision can be sent to the same address. This policy is maintained on a best-effort basis and does not promise a response deadline.
|
package/CONTRIBUTING.md
ADDED
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
# Contributing to AutoRouter
|
|
2
|
+
|
|
3
|
+
Bug reports, documentation improvements, reproducible routing cases and focused fixes are welcome. Read the [support guide](SUPPORT.md) before opening an issue, and use the [private security reporting process](SECURITY.md) for vulnerabilities. Participation follows the [code of conduct](CODE_OF_CONDUCT.md).
|
|
4
|
+
|
|
5
|
+
For a substantial behavior change, open an issue describing the problem and proposed scope before implementing it. Fabio Rapposelli ([@frapposelli](https://github.com/frapposelli)) maintains the project and reviews design and release decisions. Review is best effort; there is no guaranteed response time.
|
|
6
|
+
|
|
7
|
+
Coding agents working in a source checkout should follow [AGENTS.md](https://github.com/frapposelli/claude-autorouter/blob/main/AGENTS.md). [CLAUDE.md](https://github.com/frapposelli/claude-autorouter/blob/main/CLAUDE.md) imports the same guidance for Claude Code.
|
|
8
|
+
|
|
9
|
+
## Submit a change
|
|
10
|
+
|
|
11
|
+
Fork the repository, clone your fork and create a branch. While the repository is private, this requires access and permission to fork; existing collaborators can use a branch in their authorized checkout.
|
|
12
|
+
|
|
13
|
+
```sh
|
|
14
|
+
git clone https://github.com/YOUR-USERNAME/claude-autorouter.git
|
|
15
|
+
cd claude-autorouter
|
|
16
|
+
git switch -c describe-your-change
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
Use Node.js 22+ and macOS or Linux (including WSL). Install pinned development tools, then run the local checks:
|
|
20
|
+
|
|
21
|
+
```sh
|
|
22
|
+
npm ci --ignore-scripts --no-audit --no-fund
|
|
23
|
+
npm run check
|
|
24
|
+
npm test
|
|
25
|
+
npm run test:package
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
Tests use synthetic local services and credentials. They require loopback binding, but make no paid provider calls, downloads or user-config changes. The package check installs and exercises the exact distributable archive. Opt-in provider/model canaries are described in [development and validation](docs/development.md).
|
|
29
|
+
|
|
30
|
+
Open a pull request against `main`. Describe the problem, resulting behavior and relevant validation, linking an issue when one exists. Keep the change focused, include meaningful regressions for behavior changes, and update affected help or documentation. Report any checks you could not run. Real provider calls and model downloads are not required for ordinary contributions; label their results separately if deliberately run.
|
|
31
|
+
|
|
32
|
+
Use synthetic fixtures. Do not commit credentials, personal configuration, private prompts, transcripts or session logs. Metadata-only logs can still contain identifying information; inspect any material before sharing it. Contributions are accepted under the project's [Apache-2.0 license](LICENSE); submit only work you have the right to contribute. No CLA or sign-off workflow is required.
|
|
33
|
+
|
|
34
|
+
## Changing behavior
|
|
35
|
+
|
|
36
|
+
Follow the [request lifecycle](docs/development.md#request-lifecycle-and-model-continuity). Keep authentication and permission decisions owned by Claude. Preserve provider bytes, signed history and unfamiliar extensions. Routing must check compatibility in every profile; new human tasks remain eligible to switch models. Active task state is separate from disposable classification caches, and only clean, successfully forwarded completion evidence can establish confirmed continuation state.
|
|
37
|
+
|
|
38
|
+
Update `src/contracts.mjs` alongside configuration, decision or telemetry changes. Keep the normalizer allowlist and privacy tests aligned. Logged selections are not successful outcomes, logging stays opt-in, and metadata mode must omit prompts. [Static/style checks](docs/development.md#static-contracts-and-style) run in CI; do not add runtime dependencies for developer tooling.
|
|
39
|
+
|
|
40
|
+
## Model and pricing updates
|
|
41
|
+
|
|
42
|
+
1. Identify the exact provider model ID; a family keyword or custom alias is not capability evidence.
|
|
43
|
+
2. Update `src/model-catalog.mjs` with the authoritative source and review date. Verify context/output limits, thinking, tool choice, native tool features and Auto eligibility.
|
|
44
|
+
3. Add valid-source compatibility fixtures for every affected profile and a regression for the new restriction or permitted switch. Keep unknown models/extensions conservative.
|
|
45
|
+
4. Update thinking adaptation only where documented; token counting and inference must apply the same compatibility policy.
|
|
46
|
+
5. If rates change, review `src/savings.mjs`, bump its pricing version/date and source, and test cache TTL/modifier/unknown-model coverage. Never silently price old logs using an unrecorded new table.
|
|
47
|
+
6. Run all three local checks and exact-package fixtures. Report separately any explicitly invoked real-provider observations and their limits.
|
|
48
|
+
|
|
49
|
+
## Evaluation and performance
|
|
50
|
+
|
|
51
|
+
Declare label agreement, acceptable tiers and under-routing thresholds before testing a candidate. Preserve fixture checksums and held-out cases; transport, evaluator availability, policy and task quality are separate gates. Profile coverage requires all three tiers for compatible routing and Sonnet/Opus for Auto.
|
|
52
|
+
|
|
53
|
+
Capture a baseline before changing overhead. Repeat the same workload and hardware with the [router/storage harnesses](docs/development.md#performance-regression-measurements), retaining call counts, latency distributions, memory and background load. Numerical timing gates are local and opt-in; CI checks deterministic cancellation/resource behavior. Real local-model benchmarks need representative memory sizes and observed cold/warm conditions.
|
|
54
|
+
|
|
55
|
+
## Releasing
|
|
56
|
+
|
|
57
|
+
Use the [release procedure](docs/releasing.md) and retain the tested immutable archive. Acceptance by npm and public availability are separate states. Verify registry integrity and an isolated install before calling a release verified. A pending submission is investigated or verified again, rather than blindly republished. New CLI interfaces and event-schema changes should be reviewed together as a minor release; correctness-only fixes can be independent patches.
|
package/README.md
CHANGED
|
@@ -1,126 +1,103 @@
|
|
|
1
|
-
#
|
|
1
|
+
# AutoRouter
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
An independent local model-routing gateway for Claude Code. AutoRouter is not affiliated with, endorsed by, or sponsored by Anthropic. The existing npm package and command remain `claude-autorouter`.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
Use Haiku, Sonnet and Opus in one Claude Code session. AutoRouter evaluates each coding request, checks model compatibility and context capacity, and forwards it through a local gateway. Native Ollama `/v1/systemone` models are the default, experimental local evaluator, so task excerpts stay on your machine; [TypeSafe Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) is an optional hosted evaluator (`setup --evaluator jev`). Claude owns authentication, tool permissions and safety review.
|
|
6
6
|
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
Install from [npm](https://www.npmjs.com/package/claude-autorouter):
|
|
7
|
+
Requires Node.js 22+, macOS or Linux (including WSL), an installed `claude` command, and a Claude subscription login or Anthropic API key. The default evaluator also needs a [TypeSafe API key](https://console.typesafe.ai). The installed CLI has no runtime dependencies.
|
|
10
8
|
|
|
11
|
-
|
|
12
|
-
npm install -g claude-autorouter
|
|
13
|
-
```
|
|
9
|
+
Version 0.4.0 adds `config`, `sessions` and `doctor --evaluate-local`, durable task continuity, and clearer model outcomes. Upgrade from 0.3.x to use these commands. The [contributor guide](CONTRIBUTING.md) explains local verification, and the [release guide](docs/releasing.md) covers the changes and verified publication.
|
|
14
10
|
|
|
15
|
-
|
|
11
|
+
## Install and start
|
|
16
12
|
|
|
17
13
|
```sh
|
|
14
|
+
npm install -g claude-autorouter
|
|
18
15
|
claude-autorouter setup
|
|
19
16
|
claude-autorouter doctor
|
|
20
17
|
cd /path/to/project
|
|
21
18
|
claude-autorouter claude
|
|
22
19
|
```
|
|
23
20
|
|
|
24
|
-
Setup defaults to your Claude subscription and
|
|
21
|
+
Setup defaults to your Claude subscription and the local Ollama evaluator (Ollama 0.35+ with the default model; add `--pull` to download it), so it asks for no evaluator key. To use hosted Jev instead, run `claude-autorouter setup --evaluator jev`, which prompts privately for its key. Run `claude auth login` if needed. Jev has separate credentials and billing; subscription mode needs no Anthropic API key. For API billing, use `setup --auth-mode api-key`.
|
|
25
22
|
|
|
26
|
-
|
|
23
|
+
AutoRouter launches your installed, unmodified official Claude Code executable. Each user supplies their own login or API credentials. Subscription forwarding is a technical integration, not a claim of provider approval; review the [integration boundaries and current provider-policy notes](docs/subscription-integration.md) for your deployment.
|
|
27
24
|
|
|
28
|
-
|
|
25
|
+
Configuration is saved privately at `~/.config/claude-autorouter/config.json`. On macOS, new setups keep keys in the login Keychain; for an existing plaintext configuration, run `claude-autorouter config set AUTOROUTER_SECRET_STORE keychain` to move them. Environment variables override it; project `.env` files are not loaded automatically. `setup --force` updates an existing configuration while preserving other settings. Use focused commands for later edits:
|
|
29
26
|
|
|
30
27
|
```sh
|
|
31
|
-
claude-autorouter
|
|
32
|
-
claude-autorouter
|
|
33
|
-
claude-autorouter
|
|
28
|
+
claude-autorouter config show
|
|
29
|
+
claude-autorouter config set AUTOROUTER_JEV_TIMEOUT_MS 2000
|
|
30
|
+
claude-autorouter config unset AUTOROUTER_JEV_TIMEOUT_MS
|
|
31
|
+
claude-autorouter help config
|
|
34
32
|
```
|
|
35
33
|
|
|
36
|
-
|
|
34
|
+
Secret updates use a hidden prompt or `--stdin`, never a command-line value. Claude arguments pass through, including `claude-autorouter claude --help`. [Configuration reference](docs/reference.md#configuration).
|
|
37
35
|
|
|
38
|
-
|
|
36
|
+
## Auto permission mode
|
|
39
37
|
|
|
40
38
|
```sh
|
|
41
39
|
claude-autorouter claude --permission-mode auto
|
|
42
40
|
```
|
|
43
41
|
|
|
44
|
-
This
|
|
42
|
+
This profile automatically switches between Sonnet 5.5 and Opus 5.5 for new human tasks. A Haiku verdict uses Sonnet. Tool and `/goal` continuations retain the task's execution model; a new task can switch up or down. Claude's native safety review and organization policies still apply. For Auto selected through Claude's UI, save `AUTOROUTER_CLIENT_PROFILE=auto` with `config set`. [Auto support and limitations](docs/reference.md#auto-permission-mode).
|
|
45
43
|
|
|
46
|
-
##
|
|
44
|
+
## Inspect decisions
|
|
47
45
|
|
|
48
|
-
The launcher adds a temporary status line
|
|
46
|
+
The launcher adds a temporary status line, preserving saved Claude settings:
|
|
49
47
|
|
|
50
48
|
```text
|
|
51
|
-
● AutoRouter ·
|
|
52
|
-
● AutoRouter · Sonnet 5 selected ·
|
|
49
|
+
● AutoRouter · Opus 5.5 · ready · Jev 210ms
|
|
50
|
+
● AutoRouter · Sonnet 5.5 selected · Auto floor from Haiku · Jev 220ms
|
|
53
51
|
```
|
|
54
52
|
|
|
55
|
-
|
|
53
|
+
`selected` means Anthropic has not reported the serving model yet. Guard reasons and errors stay visible before optional savings. Claude's own model label may show its starting model. [Status details](docs/reference.md#status-line-and-savings).
|
|
56
54
|
|
|
57
|
-
|
|
55
|
+
Persistent history is optional and disabled by default. Enable metadata-only records without prompt excerpts:
|
|
58
56
|
|
|
59
57
|
```sh
|
|
60
|
-
|
|
61
|
-
|
|
58
|
+
claude-autorouter config set AUTOROUTER_SESSION_LOG_MODE metadata
|
|
59
|
+
claude-autorouter config set AUTOROUTER_SESSION_LOG_DIR "$HOME/.local/state/claude-autorouter/sessions"
|
|
60
|
+
claude-autorouter claude
|
|
61
|
+
claude-autorouter sessions list
|
|
62
|
+
claude-autorouter sessions show ID --json
|
|
62
63
|
```
|
|
63
64
|
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
Savings are an **API-equivalent estimate for the same token counts**, using Opus as the baseline. They do not measure subscription bill reductions or quota credits and exclude Jev and local compute costs. [Status line and savings details](docs/reference.md#status-line-and-savings).
|
|
67
|
-
|
|
68
|
-
## Experimental local evaluator
|
|
69
|
-
|
|
70
|
-
The local setup below requires AutoRouter 0.3.2 or newer. It uses Ollama's native `/v1/systemone` decision API with `nimble:9b-q4_K_M` by default. Jev remains the default evaluator. If upgrading from 0.2.0, replace the old Qwen model configuration using the [migration steps](docs/reference.md#migrating-an-older-ollama-config).
|
|
65
|
+
Copy an `id` from `list`. History separates model decisions from outcomes and reports latency, fallbacks, failures and savings coverage. Choose `prompts` mode for bounded human-task excerpts. Files persist locally; no automatic deletion occurs. [History and privacy](docs/reference.md#session-decision-logs).
|
|
71
66
|
|
|
72
|
-
|
|
67
|
+
Savings are **API-equivalent estimates using the recorded Opus baseline and token counts**. They do not measure subscription bill reductions or quota credits, and exclude evaluator and local compute costs. Missing or unsupported usage stays unpriced.
|
|
73
68
|
|
|
74
|
-
|
|
69
|
+
## Local Ollama evaluator
|
|
75
70
|
|
|
76
|
-
|
|
77
|
-
AUTOROUTER_OLLAMA_TIMEOUT_MS=0 claude-autorouter claude
|
|
78
|
-
```
|
|
79
|
-
|
|
80
|
-
To save that setting for an installed Tev1 4B model:
|
|
71
|
+
Ollama is the default evaluator. Start Ollama 0.35+ with a model supporting its native decision endpoint, then choose a model:
|
|
81
72
|
|
|
82
73
|
```sh
|
|
83
|
-
claude-autorouter setup --evaluator ollama --ollama-model tev1:4b --
|
|
74
|
+
claude-autorouter setup --evaluator ollama --ollama-model tev1:4b-q4_K_M --pull --force
|
|
75
|
+
claude-autorouter doctor --evaluate-local
|
|
76
|
+
claude-autorouter claude
|
|
84
77
|
```
|
|
85
78
|
|
|
86
|
-
|
|
79
|
+
`--pull` authorizes downloading the chosen model if missing. Setup keeps existing models and settings; ordinary launches download nothing. The default local model is `nimble:9b-q4_K_M`; `tev1:0.8b` is smaller and requires checking its accuracy on your tasks. Local classification needs no Jev key; Jev remains available with `setup --evaluator jev`. Claude still answers through Anthropic. [Model choices, deadlines and historical measurements](docs/reference.md#ollama-evaluator).
|
|
87
80
|
|
|
88
|
-
|
|
81
|
+
To allow a slower local model to finish without AutoRouter's runtime deadline:
|
|
89
82
|
|
|
90
83
|
```sh
|
|
91
|
-
claude-autorouter
|
|
92
|
-
claude-autorouter doctor
|
|
93
|
-
claude-autorouter claude
|
|
84
|
+
claude-autorouter config set AUTOROUTER_OLLAMA_TIMEOUT_MS 0
|
|
94
85
|
```
|
|
95
86
|
|
|
96
|
-
|
|
87
|
+
Cancellation and response-size limits still apply. The local diagnostic uses synthetic prompts and reports observed latency, classification and fallback reasons; it makes no Anthropic/Jev calls or downloads and preserves unrelated resident models.
|
|
97
88
|
|
|
98
|
-
|
|
99
|
-
| --- | ---: | --- |
|
|
100
|
-
| [Nimble 9B Q4_K_M](https://ollama.com/library/nimble) | 5.63 GB | Default: `nimble:9b-q4_K_M` |
|
|
101
|
-
| [Tev1 0.8B Q8](https://ollama.com/library/tev1) | 812 MB | `tev1:0.8b` |
|
|
102
|
-
| [Tev1 4B Q4_K_M](https://ollama.com/library/tev1) | 2.7 GB | `tev1:4b-q4_K_M` |
|
|
89
|
+
## Troubleshoot and upgrade
|
|
103
90
|
|
|
104
|
-
|
|
91
|
+
Inspect `config show` for environment overrides, and `doctor` for setup health. A valid evaluator verdict can be overridden by continuity, context or model compatibility. `Ollama fallback: timeout` means evaluation failed to finish, rather than predicting Sonnet. Status errors and saved history explain these paths.
|
|
105
92
|
|
|
106
93
|
```sh
|
|
107
|
-
|
|
94
|
+
npm install -g claude-autorouter@latest
|
|
95
|
+
claude-autorouter --version
|
|
96
|
+
claude-autorouter doctor
|
|
108
97
|
```
|
|
109
98
|
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
No Jev key is needed for local classification. The launcher primes the evaluator before opening Claude's UI, and evaluation failures fall back to Sonnet or retain Opus without contacting Jev. Claude still answers through Anthropic, with the same routing guards and subscription limits.
|
|
113
|
-
|
|
114
|
-
Setup, doctor, and startup show the effective model and deadline. Warmup and doctor do not certify classification speed or accuracy. `Ollama fallback: timeout` means no valid decision arrived in time; it is different from the evaluator choosing Sonnet. Source users can run the [local routing regression](docs/development.md#local-routing-regression) to check all three tiers without Claude or Jev calls.
|
|
115
|
-
|
|
116
|
-
Historical measurements before 0.3.2, on a 16 GiB M4: Tev1 0.8B matched 18/24 held-out labels with 450 ms median latency and no timeouts at 1,500 ms, including full-excerpt checks. Tev1 4B matched 22/24 with a 10-second diagnostic deadline and 3.15-second median latency. Nimble matched 23/24 with a 30-second deadline and 11.4-second median latency. Both larger models exceeded the then-default 1,500 ms. A separate six-case regression with the 0.3.2 fixes passed for both tested Tev1 4B variants and Nimble; Tev1 0.8B matched only three cases. These small tests do not establish general accuracy or Jev parity. See the [measurements and limits](docs/ollama-evaluation.md) and [Ollama reference](docs/reference.md#ollama-evaluator).
|
|
117
|
-
|
|
118
|
-
## Behavior and data
|
|
99
|
+
Historical integration observations cover Claude Code 2.1.284–2.1.285. The versioned synthetic protocol fixtures test reviewed request/response contracts; they do not certify the current checkout against a live Claude version. Real-provider checks remain explicitly invoked. [Troubleshooting](docs/reference.md#troubleshooting) covers context use, blocked goals and logging. Run ordinary `claude` to bypass routing.
|
|
119
100
|
|
|
120
|
-
|
|
121
|
-
- The selected evaluator receives bounded excerpts that can contain source code and tool results: TypeSafe with Jev, or the local service with Ollama. Jev also receives system-text excerpts; the local path excludes Claude's executor system instructions. Anthropic receives the complete request. Images, document payloads, and private thinking are omitted from classifier input. [Data flow and authentication](docs/reference.md#data-flow-and-authentication).
|
|
122
|
-
- Subscription access and usage limits still apply. Model switches can reduce cache reuse; cheaper token prices do not guarantee cheaper completed tasks. Run ordinary `claude` to bypass routing.
|
|
123
|
-
- The launcher is quiet by default. Use `AUTOROUTER_DEBUG=1` for metadata diagnostics or `AUTOROUTER_STATUSLINE=0` to retain your existing status line. [Troubleshooting](docs/reference.md#troubleshooting).
|
|
124
|
-
- For blocked `/goal` loops, optionally launch with `env CLAUDE_CODE_STOP_HOOK_BLOCK_CAP=2 claude-autorouter claude`. Claude then ends the turn on the third consecutive blocking verdict without tool use, leaving the goal unmet. This also affects other Stop/SubagentStop hooks; defaults are unchanged. [Scope and saved configuration](docs/reference.md#shorter-stop-hook-loops-opt-in).
|
|
101
|
+
The evaluator receives bounded task/history excerpts that may contain code and tool results: TypeSafe for Jev, or your loopback Ollama service. Recognizable credentials and personal identifiers are redacted from those excerpts first. Anthropic receives the complete request. Model switching can reduce cache reuse. [Data flow and authentication](docs/reference.md#data-flow-and-authentication).
|
|
125
102
|
|
|
126
|
-
[Reference](docs/reference.md) · [
|
|
103
|
+
[Reference](docs/reference.md) · [Integration and provider policy](docs/subscription-integration.md) · [Contributing](CONTRIBUTING.md) · [Development](docs/development.md) · [Releases](docs/releasing.md) · [Apache-2.0](LICENSE)
|
package/SECURITY.md
ADDED
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
# Security policy
|
|
2
|
+
|
|
3
|
+
## Report privately
|
|
4
|
+
|
|
5
|
+
Use GitHub's private [Report a vulnerability](https://github.com/frapposelli/claude-autorouter/security/advisories/new) form. Private vulnerability reporting is enabled for this repository. If the form is unavailable, email maintainer Fabio Rapposelli at [fabio@rapposelli.org](mailto:fabio@rapposelli.org) with the subject `AutoRouter security report`.
|
|
6
|
+
|
|
7
|
+
Do not open a public issue or pull request with exploit details, credentials or sensitive request data. Include the affected AutoRouter and Claude Code versions, operating system, evaluator/client profile, expected security boundary, observed impact and a minimal synthetic reproduction where possible. Do not include real API keys, OAuth tokens, private source code, prompts or transcripts. The maintainer can coordinate any additional evidence privately.
|
|
8
|
+
|
|
9
|
+
Reports are handled on a best-effort basis; there is no guaranteed response time or bounty program. We will coordinate investigation, remediation and disclosure with the reporter before publishing details.
|
|
10
|
+
|
|
11
|
+
## Supported versions and scope
|
|
12
|
+
|
|
13
|
+
Security fixes target the latest stable release. Older releases are not maintained as separate security branches; users may need to upgrade. Reports against `main` are also welcome.
|
|
14
|
+
|
|
15
|
+
Relevant reports include credential exposure, unauthorized access to the local gateway or saved configuration, unintended disclosure through logs, and routing or request changes that weaken Claude's authentication or permission boundaries. Provider accounts, billing and vulnerabilities in Claude Code, TypeSafe or Ollama should also be reported to the responsible provider when applicable.
|
|
16
|
+
|
|
17
|
+
## Handling diagnostic data
|
|
18
|
+
|
|
19
|
+
AutoRouter's default evaluator is local Ollama, which keeps classification on loopback. The optional hosted Jev evaluator (`--evaluator jev`) receives bounded task/history excerpts, which may contain private code or tool results. Recognizable credentials and personal identifiers are redacted from those excerpts first; this pattern-based filter reduces, but does not eliminate, disclosure. Anthropic still receives the full inference request. See [data flow and authentication](docs/reference.md#data-flow-and-authentication).
|
|
20
|
+
|
|
21
|
+
New macOS setups keep saved keys in the login Keychain. Existing and non-macOS configurations keep plaintext keys in the private configuration file until you run `claude-autorouter config set AUTOROUTER_SECRET_STORE keychain` (macOS only). See [credential storage](docs/reference.md#credential-storage).
|
|
22
|
+
|
|
23
|
+
Session logging is optional. Enabling only a log directory uses the default `prompts` mode, which includes bounded human-task excerpts with recognizable credentials and personal identifiers redacted (pattern-based, so not exhaustive). Select `AUTOROUTER_SESSION_LOG_MODE=metadata` before enabling a directory to omit those excerpts. Inspect even metadata-only output before sharing it; identifiers, paths or environment details can still be sensitive. Saved logs have no automatic deletion policy. See [history and privacy](docs/reference.md#session-decision-logs).
|
package/SUPPORT.md
ADDED
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
# Support
|
|
2
|
+
|
|
3
|
+
AutoRouter is an independent, community-maintained project. It does not provide official support for Anthropic, TypeSafe or Ollama, and has no response-time guarantee.
|
|
4
|
+
|
|
5
|
+
For setup and usage, start with the [README](README.md) and [troubleshooting reference](docs/reference.md#troubleshooting). Run `claude-autorouter --version`, `claude --version` and `claude-autorouter doctor` to identify the installation and configuration involved. Ordinary `doctor` performs configuration and local service checks; it does not verify paid-provider access.
|
|
6
|
+
|
|
7
|
+
Use [GitHub issues](https://github.com/frapposelli/claude-autorouter/issues) for reproducible bugs, feature proposals and questions not answered by the documentation. Repository access is required while the repository is private; if you cannot access it, contact [fabio@rapposelli.org](mailto:fabio@rapposelli.org). For vulnerabilities, follow [SECURITY.md](SECURITY.md) instead of opening an issue. Community conduct reports follow the [code of conduct](CODE_OF_CONDUCT.md).
|
|
8
|
+
|
|
9
|
+
## Make a report useful
|
|
10
|
+
|
|
11
|
+
- Include AutoRouter, Claude Code, Node.js and operating-system versions; include the Ollama version and model tag for local evaluation.
|
|
12
|
+
- Identify the authentication mode, evaluator and client profile without sharing credentials or full environment/configuration dumps.
|
|
13
|
+
- Describe expected and actual behavior, and provide minimal steps using a synthetic prompt or fixture when possible. Distinguish the selected model from the provider-confirmed serving model.
|
|
14
|
+
- Share only the relevant, inspected diagnostic excerpt. Remove secrets, private prompts, responses, source code, personal paths and organization identifiers. Screenshots can disclose this information too.
|
|
15
|
+
|
|
16
|
+
Logging is disabled by default. When enabling optional session history for a reproduction, explicitly choose `AUTOROUTER_SESSION_LOG_MODE=metadata`; the default `prompts` mode includes task excerpts. Metadata-only output still needs review before sharing. Debug stderr can also contain Claude's own diagnostics. See [session logs](docs/reference.md#session-decision-logs) and [troubleshooting](docs/reference.md#troubleshooting).
|
|
17
|
+
|
|
18
|
+
Account access, subscription/model eligibility, billing and provider outages belong with the responsible provider. AutoRouter reports can investigate how the gateway handles those failures, but cannot change a provider's account policies. Local-model accuracy and performance depend on the workload and hardware; include those conditions when reporting unexpected classifications.
|
package/bin/autorouter.mjs
CHANGED
|
@@ -12,66 +12,48 @@ import { createSessionLog } from '../src/session-log.mjs';
|
|
|
12
12
|
import { loadUserConfig } from '../src/user-config.mjs';
|
|
13
13
|
import { setup, doctor, ollamaDeadlineText } from '../src/onboarding.mjs';
|
|
14
14
|
import { setupOllama } from '../src/ollama-setup.mjs';
|
|
15
|
+
import { helpText } from '../src/cli-help.mjs';
|
|
16
|
+
import { configCommand } from '../src/config-command.mjs';
|
|
17
|
+
import { sessionsCommand } from '../src/session-history.mjs';
|
|
15
18
|
|
|
16
19
|
const [command = 'help', ...args] = process.argv.slice(2);
|
|
17
20
|
if (['--version', '-v', 'version'].includes(command)) {
|
|
18
21
|
console.log(JSON.parse(readFileSync(new URL('../package.json', import.meta.url), 'utf8')).version);
|
|
19
22
|
} else if (['help', '--help', '-h'].includes(command)
|
|
20
|
-
|| (['setup', 'doctor', 'serve'].includes(command) && args.some(arg => ['--help', '-h'].includes(arg)))) {
|
|
21
|
-
console.log(
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
claude-autorouter --version
|
|
34
|
-
|
|
35
|
-
Setup defaults to subscription authentication and prompts for keys without echoing.
|
|
36
|
-
For noninteractive setup, supply keys through environment variables.
|
|
37
|
-
User config: ~/.config/claude-autorouter/config.json (or XDG_CONFIG_HOME).
|
|
38
|
-
AUTOROUTER_CONFIG selects a different file; environment variables take precedence.
|
|
39
|
-
Project .env files are never loaded automatically.
|
|
40
|
-
|
|
41
|
-
Jev is the default evaluator and requires TYPESAFE_API_KEY.
|
|
42
|
-
Ollama evaluates locally and requires Ollama 0.35+ with /v1/systemone.
|
|
43
|
-
Use setup --evaluator ollama --pull to detect Ollama and download a missing model.
|
|
44
|
-
The local default is nimble:9b-q4_K_M; --ollama-model selects another compatible model.
|
|
45
|
-
Smaller Tev1 options: --ollama-model tev1:0.8b or --ollama-model tev1:4b-q4_K_M.
|
|
46
|
-
Local routing deadlines: Tev1 0.8B/custom 1500 ms, Tev1 4B 15000 ms, Nimble 30000 ms.
|
|
47
|
-
Setup --ollama-timeout-ms N saves a routing deadline; use 0 to disable it.
|
|
48
|
-
AUTOROUTER_OLLAMA_TIMEOUT_MS also overrides the deadline; 0 disables it.
|
|
49
|
-
Local routing is experimental; see docs/ollama-evaluation.md for measured limits.
|
|
50
|
-
AUTOROUTER_AUTH_MODE=subscription uses your saved Claude Code login.
|
|
51
|
-
Without setup, AUTOROUTER_AUTH_MODE defaults to api-key and also requires ANTHROPIC_API_KEY.
|
|
52
|
-
AUTOROUTER_CLIENT_PROFILE=compatible (default) enables all three routing tiers.
|
|
53
|
-
Use AUTOROUTER_CLIENT_PROFILE=native to retain Claude Code's own model/thinking settings.
|
|
54
|
-
Use AUTOROUTER_CLIENT_PROFILE=auto for Auto permission mode: Sonnet/Opus routing, native thinking.
|
|
55
|
-
Auto defaults to Sonnet 5.5 and Opus 5.5, switching on new human tasks and retaining tool turns.
|
|
56
|
-
An explicit claude --permission-mode auto selects the auto profile for that launch.
|
|
57
|
-
Claude's permission checks and organization policies still apply; Haiku does not support Auto mode.
|
|
58
|
-
Optional CLAUDE_CODE_STOP_HOOK_BLOCK_CAP=N limits consecutive tool-free Stop-hook continuations.
|
|
59
|
-
Use 2 to stop on the third block; applies to /goal and all Stop/SubagentStop hooks.
|
|
60
|
-
Unset preserves Claude's default; 0 disables the cap. Setup --stop-hook-block-cap N saves it.
|
|
61
|
-
Standalone serve also requires AUTOROUTER_TOKEN (at least 16 characters).
|
|
62
|
-
The claude launcher creates a temporary credential and an ephemeral port.
|
|
63
|
-
It enables an AutoRouter status line for this session (AUTOROUTER_STATUSLINE=0 to opt out).
|
|
64
|
-
Launcher logs are quiet by default; AUTOROUTER_DEBUG=1 enables diagnostic logs on stderr.
|
|
65
|
-
AUTOROUTER_SESSION_LOG_DIR writes private per-session JSONL decision logs with prompt excerpts.
|
|
66
|
-
Unset or empty disables session logs. Setup --session-log-dir DIR saves the directory.
|
|
67
|
-
Jev sends prompt excerpts to TypeSafe; Ollama keeps classification on this machine.
|
|
68
|
-
Complete inference requests still go to Anthropic. See README.md.`);
|
|
69
|
-
} else if (command === 'setup' || command === 'doctor') {
|
|
23
|
+
|| (['setup', 'doctor', 'serve', 'config', 'sessions'].includes(command) && args.some(arg => ['--help', '-h'].includes(arg)))) {
|
|
24
|
+
console.log(helpText(command === 'help' ? args[0] : command));
|
|
25
|
+
} else if (command === 'claude' && args.length === 1 && ['--help', '-h', '--version', '-v'].includes(args[0])) {
|
|
26
|
+
// Help/version are Claude-owned commands. No config file, evaluator keys,
|
|
27
|
+
// gateway, Ollama warmup or temporary status files are needed.
|
|
28
|
+
const env = { ...process.env };
|
|
29
|
+
delete env.TYPESAFE_API_KEY;
|
|
30
|
+
delete env.AUTOROUTER_TOKEN;
|
|
31
|
+
const child = spawn('claude', args, { stdio: 'inherit', env });
|
|
32
|
+
child.once('error', () => { console.error('Could not launch Claude Code. Ensure `claude` is installed and on PATH.'); process.exitCode = 1; });
|
|
33
|
+
child.once('exit', (code, signal) => { process.exitCode = code ?? (signal === 'SIGINT' ? 130 : signal === 'SIGTERM' ? 143 : 1); });
|
|
34
|
+
for (const signal of ['SIGINT', 'SIGTERM']) process.on(signal, () => child.kill(signal));
|
|
35
|
+
} else if (['setup', 'doctor', 'config', 'sessions'].includes(command)) {
|
|
70
36
|
try {
|
|
71
37
|
if (command === 'setup') await setup(args);
|
|
38
|
+
else if (command === 'config') {
|
|
39
|
+
if (await configCommand(args) === false) process.exitCode = 1;
|
|
40
|
+
}
|
|
41
|
+
else if (command === 'sessions') {
|
|
42
|
+
if (await sessionsCommand(args) === false) process.exitCode = 1;
|
|
43
|
+
}
|
|
72
44
|
else {
|
|
73
|
-
|
|
74
|
-
|
|
45
|
+
const evaluateLocal = args.includes('--evaluate-local');
|
|
46
|
+
const json = args.includes('--json');
|
|
47
|
+
if (args.some(arg => !['--evaluate-local', '--json'].includes(arg)) || new Set(args).size !== args.length
|
|
48
|
+
|| (json && !evaluateLocal)) throw new Error('Usage: claude-autorouter doctor [--evaluate-local [--json]]');
|
|
49
|
+
const controller = new AbortController();
|
|
50
|
+
const cancel = () => controller.abort();
|
|
51
|
+
for (const signal of ['SIGINT', 'SIGTERM']) process.once(signal, cancel);
|
|
52
|
+
try {
|
|
53
|
+
if (!await doctor({ evaluateLocal, json, signal: controller.signal })) process.exitCode = 1;
|
|
54
|
+
} finally {
|
|
55
|
+
for (const signal of ['SIGINT', 'SIGTERM']) process.removeListener(signal, cancel);
|
|
56
|
+
}
|
|
75
57
|
}
|
|
76
58
|
} catch (error) { console.error(error.message); process.exitCode = 1; }
|
|
77
59
|
} else if (!['claude', 'serve'].includes(command)) {
|
|
@@ -85,10 +67,9 @@ Complete inference requests still go to Anthropic. See README.md.`);
|
|
|
85
67
|
const stop = () => {
|
|
86
68
|
if (stopping) return stopping;
|
|
87
69
|
if (server) { server.close(); server.closeAllConnections(); }
|
|
88
|
-
status?.close();
|
|
89
70
|
// Drain accepted decision records before normal process exit. Pending
|
|
90
71
|
// filesystem writes keep Node alive; no timer or fire-and-forget buffer.
|
|
91
|
-
stopping = Promise.
|
|
72
|
+
stopping = Promise.allSettled([status?.close(), sessionLog?.close()]);
|
|
92
73
|
return stopping;
|
|
93
74
|
};
|
|
94
75
|
try {
|
|
@@ -120,15 +101,17 @@ Complete inference requests still go to Anthropic. See README.md.`);
|
|
|
120
101
|
let claudeArgs = args;
|
|
121
102
|
if (statusEnabled) {
|
|
122
103
|
status = createStatusState({ baselineModel: config.models.opus });
|
|
104
|
+
await status.ready;
|
|
123
105
|
if (status.path) {
|
|
124
106
|
try { claudeArgs = addStatusLineSettings(args, dirname(status.path)); }
|
|
125
107
|
catch {
|
|
126
|
-
status.close(); status = undefined;
|
|
108
|
+
await status.close(); status = undefined;
|
|
127
109
|
console.error('AutoRouter status line unavailable: could not safely prepare session settings. Passing your original settings to Claude.');
|
|
128
110
|
}
|
|
129
111
|
} else console.error('AutoRouter status line unavailable: could not create local status storage.');
|
|
130
112
|
}
|
|
131
113
|
if (config.sessionLogDir) sessionLog = await createSessionLog(config.sessionLogDir, {
|
|
114
|
+
includePrompts: config.sessionLogMode === 'prompts',
|
|
132
115
|
warn: message => console.error(message),
|
|
133
116
|
});
|
|
134
117
|
// Claude owns the terminal while its UI is running. Status updates use the
|
|
@@ -136,7 +119,7 @@ Complete inference requests still go to Anthropic. See README.md.`);
|
|
|
136
119
|
server = createRouterServer(config, {
|
|
137
120
|
log: diagnosticLogs ? undefined : () => {},
|
|
138
121
|
onStatus: event => status?.update(event),
|
|
139
|
-
|
|
122
|
+
onRecord: sessionLog ? entry => sessionLog.record(entry) : undefined,
|
|
140
123
|
});
|
|
141
124
|
const address = await listen(server, command === 'claude' ? 0 : config.port);
|
|
142
125
|
const baseUrl = `http://127.0.0.1:${address.port}`;
|
|
@@ -155,7 +138,7 @@ Complete inference requests still go to Anthropic. See README.md.`);
|
|
|
155
138
|
if (status?.path) env.AUTOROUTER_STATUS_FILE = status.path;
|
|
156
139
|
const child = spawn('claude', claudeArgs, { stdio: 'inherit', env });
|
|
157
140
|
child.once('error', () => { console.error('Could not launch Claude Code. Ensure `claude` is installed and on PATH.'); process.exitCode = 1; stop(); });
|
|
158
|
-
child.once('exit', (code, signal) => { process.exitCode = code ?? (signal === 'SIGINT' ? 130 : 1); stop(); });
|
|
141
|
+
child.once('exit', (code, signal) => { process.exitCode = code ?? (signal === 'SIGINT' ? 130 : signal === 'SIGTERM' ? 143 : 1); stop(); });
|
|
159
142
|
for (const signal of ['SIGINT', 'SIGTERM']) process.on(signal, () => child.kill(signal));
|
|
160
143
|
}
|
|
161
144
|
} catch (error) {
|
package/docs/development.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Development and validation
|
|
2
2
|
|
|
3
|
-
Use Node.js 22+ from a source checkout on macOS or Linux (including WSL). The
|
|
3
|
+
Use Node.js 22+ from a source checkout on macOS or Linux (including WSL). The installed CLI has no runtime package dependencies. Source checks use pinned TypeScript and Node type definitions; install these contributor tools with `npm ci --ignore-scripts --no-audit --no-fund`. Development scripts and tests are separate from the installed CLI; user setup is covered in the [README](../README.md).
|
|
4
4
|
|
|
5
5
|
## Local checks
|
|
6
6
|
|
|
@@ -16,6 +16,24 @@ Auto-mode regressions exercise Sonnet → Opus → Opus tool continuation → So
|
|
|
16
16
|
|
|
17
17
|
Package validation checks the distributable and installed command rather than relying on the source checkout's paths. Review the [release procedure](releasing.md) before distributing a tarball.
|
|
18
18
|
|
|
19
|
+
## Request lifecycle and model continuity
|
|
20
|
+
|
|
21
|
+
The gateway validates the request containers it consumes, evaluates the current task, applies continuity and capacity rules, then checks the proposed target against the shared model catalog. Unknown provider extensions remain intact; when their compatibility with another model is unknown, the source model is retained with a routing reason. Token-count requests apply the same compatibility rules before sending a request.
|
|
22
|
+
|
|
23
|
+
`src/model-catalog.mjs` contains exact model IDs, capability facts, source links, and a review date. A family name inside a custom alias does not establish capabilities. `src/model-request.mjs` contains explicit thinking adaptations; neither layer strips signed history or permission-review settings. Model updates should change the catalog, include a dated authoritative source, and add a request fixture demonstrating the restriction or new supported switch.
|
|
24
|
+
|
|
25
|
+
The evaluation cache and active execution state have separate lifetimes. `src/turn-state.mjs` keeps active human tasks and pending tools beyond the evaluator cache TTL. Retired tasks expire, aliases and records are bounded, and exhausting active-state capacity is reported rather than silently evicting another active task. State is process-local; after a gateway restart, missing continuity is reported as unknown until a successful response establishes it again.
|
|
26
|
+
|
|
27
|
+
The gateway stages a selected model under its request ID. `src/response-observer.mjs` observes serving models, provider fallback boundaries, closed tool calls, usage, and terminal response metadata without altering bytes. A serving-model observation alone is not a successful execution. Only clean completion evidence followed by successful HTTP forwarding commits the continuation model. Cancelled, failed, ambiguous, and superseded attempts cannot overwrite known state. New human tasks remain eligible for upward or downward switching, including Sonnet/Opus in Auto mode.
|
|
28
|
+
|
|
29
|
+
Direct `Router.route()` embedders that omit `requestId` retain selected, unconfirmed continuity for compatibility. Embedders that execute inference should supply a unique request ID and call `router.complete(id, evidence)` after successful delivery, or `router.complete(id)` on failure. The HTTP gateway owns that lifecycle automatically.
|
|
30
|
+
|
|
31
|
+
## Evaluation acceptance
|
|
32
|
+
|
|
33
|
+
Evaluation reports distinguish evaluator availability, rubric agreement, routing policy, tier coverage, transport, and independently checked task completion. An unmeasured gate is explicitly marked unmeasured. Normal runs cannot pass solely on classifier fallback or cached predictions; simulated-outage runs explicitly require fallback. Compatible routing requires all three selected tiers, while Auto requires Sonnet and Opus. Constrained fixtures declare expected guard overrides.
|
|
34
|
+
|
|
35
|
+
The general and Ollama evaluation scripts accept `--min-agreement` and `--max-under-route-rate`. Their defaults require complete expected-label agreement and no under-routing. Set any alternative thresholds **before** evaluating a candidate, retain the fixture checksum with the report, and keep tuning cases separate from held-out cases. Rubric labels are judgments about synthetic tasks; these reports do not prove end-user task quality or subscription savings. The expanded corpus includes multilingual tasks, short difficult follow-ups, ordinary work in long background context, and task text containing tier-selection instructions. No prompt-policy adjustment should be justified by rerunning and relabeling the held-out set.
|
|
36
|
+
|
|
19
37
|
## Run from source
|
|
20
38
|
|
|
21
39
|
```sh
|
|
@@ -76,7 +94,7 @@ For a classifier-only rubric evaluation:
|
|
|
76
94
|
npm run eval
|
|
77
95
|
```
|
|
78
96
|
|
|
79
|
-
The bundled evaluation makes
|
|
97
|
+
The bundled evaluation makes 19 classifier calls and no Claude generations. Jev is the default and incurs TypeSafe usage; set `AUTOROUTER_EVALUATOR=ollama` to evaluate an installed local model. It reports agreement with the starting rubric, fallback count, and p50/p95 routing latency. Edit `test/fixtures/routing.json` to represent the tasks you want to measure. Rubric agreement alone does not establish answer quality or net savings; compare completed tasks against fixed-model baselines.
|
|
80
98
|
|
|
81
99
|
For local evaluator measurements, use Ollama 0.35+ and a model compatible with `/v1/systemone`. Distinguish cold model loading from warmed classification, and record the model tag, hardware, Ollama version, context size, prompt length, and resident memory. The launcher primes the classifier with a synthetic task before opening the UI, with a separate deadline of up to 60 seconds. Runtime and benchmark share the 3,000-character/3,000-UTF-8-byte state limit, so include non-ASCII cases and excerpts that fill the budget. Also measure the first request after keep-alive expiration: its reload can hit the normal deadline even when warm requests pass. Repeat on realistic prompt distributions instead of selecting a model from a single easy request. Disk download size is not resident RAM. Keep model downloads opt-in and respect each model's license.
|
|
82
100
|
|
|
@@ -113,3 +131,31 @@ node --env-file=.env scripts/context-probe.mjs --cwd /path/to/synthetic-fixture
|
|
|
113
131
|
```
|
|
114
132
|
|
|
115
133
|
Private repository or connected-tool context may be present even when the typed prompt is harmless. Keep private-payload investigations local unless external processing is authorized. For shareable live regressions, prefer the isolated synthetic fixtures above. The probe report itself persists only metadata.
|
|
134
|
+
|
|
135
|
+
## Static contracts and style
|
|
136
|
+
|
|
137
|
+
`npm run check` checks JavaScript syntax, TypeScript/JSDoc contracts for configuration, classifier results, final routing decisions and normalized telemetry, and literal event producers in transport code. `src/contracts.mjs` is the shared development-time type contract; normalizers remain the runtime privacy boundary. Negative fixtures in `test/static-contracts.mts` and `test/static-checks.test.mjs` prove misspelled fields, invalid enums, payload fields and timing strings are rejected before execution. Provider request extensions remain opaque and are validated only where the router consumes them.
|
|
138
|
+
|
|
139
|
+
Use two-space indentation, LF endings, one final newline, semicolons and single quotes for ordinary strings. Compact pure helpers are allowed when readable; do not reformat unrelated code. The style check rejects trailing whitespace, tab indentation, `var`, and coercing comparisons except deliberate null/undefined checks. TypeScript is a contributor dependency only; public packages keep zero runtime dependencies. CI installs the pinned lockfile before checks and never runs provider inference automatically.
|
|
140
|
+
|
|
141
|
+
## Versioned protocol evidence
|
|
142
|
+
|
|
143
|
+
`test/fixtures/claude-protocol-v1.json` is a versioned, newly authored synthetic corpus reviewed against the Messages API, streaming, deferred-tool and fallback contracts. `test/protocol-fixtures.test.mjs` sends it through the real gateway, router and response observer with fake evaluator/upstream services. It covers Auto floors and switches, thinking adaptation and preservation, custom deferred tool references, conservative built-in server-tool history, compaction, scoped goal feedback, parallel agents, model fallback/tool ownership, usage and truncated responses.
|
|
144
|
+
|
|
145
|
+
Historical Claude Code 2.1.284/2.1.285 report versions, dates and hashes are separate metadata. The corpus does not copy captured prompts, invent provider signatures or certify a live client. Current-source real-provider canaries remain opt-in and unmeasured unless a separate report records them.
|
|
146
|
+
|
|
147
|
+
## Performance regression measurements
|
|
148
|
+
|
|
149
|
+
[Router measurements](router-performance.md) and [status-storage measurements](status-performance.md) record the pre-change baseline, repeated candidate runs, hardware/background load and baseline-derived gates. Source-only harnesses use synthetic inputs and providers; no credentials, prompts from user sessions or downloads are involved.
|
|
150
|
+
|
|
151
|
+
```sh
|
|
152
|
+
node --expose-gc scripts/benchmark-router.mjs baseline /tmp/router-comparison.json
|
|
153
|
+
node --expose-gc scripts/benchmark-router.mjs candidate /tmp/router-comparison.json --check
|
|
154
|
+
node scripts/benchmark-status.mjs --label local-check --check
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
Identical concurrent classifier inputs share one bounded evaluation. Each request applies its own continuity, capacity and compatibility checks. Cancelling one waiter preserves other waiters; cancelling all releases the shared evaluation. Cache identity retains the complete request, requested model floor, evaluator configuration and rubric hash. New tasks and sequential pinned continuations still evaluate; prior-pin classification reuse was deliberately not enabled. A 64 KiB response limit applies to both evaluators, and local model metadata is limited to 1 MiB. Disabling the Ollama timer does not disable cancellation or these byte limits.
|
|
158
|
+
|
|
159
|
+
Status writes use one asynchronous writer and a coalesced latest snapshot. Embedders await `state.ready` before using its path, `flush()` when they need persisted evidence, and `close()` before cleanup. Initial storage failures or a one-second readiness timeout disable the optional display. Accepted in-flight writes finish before directory removal, so shutdown cannot recreate files.
|
|
160
|
+
|
|
161
|
+
The router/storage measurements cover one 16 GiB M4 with synthetic providers. Actual Tev1 4B and Nimble 9B measurements on that Mac and a 64 GiB M2 Ultra are recorded separately in [the hardware comparison](hardware-comparison.md). Both candidates missed the unchanged strict quality gate on both hosts. Model digests and runtime conditions differ, so the cross-host results do not isolate RAM's effect. Use the [transfer bundle instructions](hardware-benchmark.md) to reproduce the workload, recording background load and observed residency. Do not infer model performance from router timings.
|