claude-autorouter 0.3.7 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (49) hide show
  1. package/.env.example +6 -3
  2. package/CODE_OF_CONDUCT.md +9 -0
  3. package/CONTRIBUTING.md +57 -0
  4. package/README.md +47 -70
  5. package/SECURITY.md +23 -0
  6. package/SUPPORT.md +18 -0
  7. package/bin/autorouter.mjs +40 -57
  8. package/docs/development.md +48 -2
  9. package/docs/hardware-benchmark.md +29 -0
  10. package/docs/hardware-comparison.md +55 -0
  11. package/docs/hardware-results-16gb.json +4002 -0
  12. package/docs/hardware-results-16gb.md +26 -0
  13. package/docs/hardware-results-64gb.json +4020 -0
  14. package/docs/reference.md +92 -37
  15. package/docs/releasing.md +79 -37
  16. package/docs/router-performance.json +1697 -0
  17. package/docs/router-performance.md +50 -0
  18. package/docs/status-performance.json +363 -0
  19. package/docs/status-performance.md +44 -0
  20. package/docs/subscription-integration.md +27 -0
  21. package/package.json +66 -10
  22. package/src/auto-routing.mjs +184 -24
  23. package/src/bounded-json.mjs +57 -0
  24. package/src/cli-help.mjs +90 -0
  25. package/src/config-command.mjs +158 -0
  26. package/src/config.mjs +53 -28
  27. package/src/contracts.mjs +123 -0
  28. package/src/evaluation-report.mjs +114 -0
  29. package/src/keychain.mjs +58 -0
  30. package/src/local-diagnostic.mjs +191 -0
  31. package/src/model-catalog.mjs +96 -0
  32. package/src/model-request.mjs +6 -7
  33. package/src/ollama-evaluator.mjs +9 -27
  34. package/src/onboarding.mjs +130 -26
  35. package/src/prompt-state.mjs +22 -7
  36. package/src/redaction.mjs +97 -0
  37. package/src/request-validation.mjs +54 -0
  38. package/src/response-observer.mjs +126 -18
  39. package/src/router.mjs +151 -61
  40. package/src/savings.mjs +74 -16
  41. package/src/server.mjs +79 -12
  42. package/src/session-history.mjs +262 -0
  43. package/src/session-log.mjs +9 -58
  44. package/src/status-state.mjs +110 -62
  45. package/src/statusline.mjs +57 -27
  46. package/src/telemetry-event.mjs +200 -0
  47. package/src/token-counter.mjs +3 -1
  48. package/src/turn-state.mjs +132 -0
  49. package/src/user-config.mjs +81 -10
package/.env.example CHANGED
@@ -1,7 +1,8 @@
1
1
  # Copy to .env and load with: node --env-file=.env bin/autorouter.mjs claude
2
2
  AUTOROUTER_AUTH_MODE=subscription
3
3
  AUTOROUTER_CLIENT_PROFILE=compatible
4
- # Jev remains the default evaluator. Ollama is experimental; see settings below.
4
+ # Local Ollama is the default evaluator (experimental; see settings below).
5
+ # Set jev to use TypeSafe's hosted evaluator, which needs TYPESAFE_API_KEY.
5
6
  AUTOROUTER_EVALUATOR=jev
6
7
  # The launcher enables the router status line for this session. Set 0 to keep your own.
7
8
  AUTOROUTER_STATUSLINE=1
@@ -12,10 +13,12 @@ AUTOROUTER_STATUSLINE=1
12
13
  # AUTOROUTER_CLIENT_PROFILE=auto
13
14
  # Optional metadata logs on stderr. Redirect stderr to a file when using the UI.
14
15
  # AUTOROUTER_DEBUG=1
15
- # Optional persistent JSONL decision logs, one file per session per launch.
16
+ # Optional persistent JSONL decisions/outcomes, one file per session per launch.
16
17
  # Includes up to 500 characters of user prompt text; keep the directory local.
17
18
  # Unset or empty disables logging. This does not print prompts in the terminal.
18
19
  # AUTOROUTER_SESSION_LOG_DIR=/absolute/path/to/autorouter-sessions
20
+ # Omit all prompt excerpt fields (does not enable logging by itself).
21
+ # AUTOROUTER_SESSION_LOG_MODE=metadata
19
22
  # Optional: allow two tool-free Stop-hook continuations, then end the turn on
20
23
  # the third block. Applies to /goal and all Stop/SubagentStop hooks.
21
24
  # Unset keeps Claude's default (currently 8); 0 DISABLES the cap.
@@ -40,7 +43,7 @@ AUTOROUTER_MIN_CONFIDENCE=0.75
40
43
  # Experimental local evaluator in AutoRouter 0.3.2+: native decision API only.
41
44
  # Install/start Ollama 0.35+, then run:
42
45
  # claude-autorouter setup --evaluator ollama --pull --force
43
- # Defaults to nimble:9b-q4_K_M (~5.63 GB download); --force replaces user config.
46
+ # Defaults to nimble:9b-q4_K_M (~5.63 GB download); --force preserves other settings.
44
47
  # Existing downloads are kept. Replace Qwen config/environment values from 0.2.0.
45
48
  # All local models use /v1/systemone; custom tags/aliases must support that API.
46
49
  # Add --ollama-model LOCAL_TAG_OR_ALIAS to setup to choose another suitable model.
@@ -0,0 +1,9 @@
1
+ # Code of conduct
2
+
3
+ Be respectful and constructive in issues, pull requests, reviews and other project spaces. Welcome people with different backgrounds and experience levels, focus criticism on ideas and code, and respect requests to stop unwanted interaction.
4
+
5
+ Harassment, discriminatory or demeaning remarks, sexualized conduct, threats, personal attacks, and publishing someone's private information without consent are not acceptable. The same expectations apply when representing the project outside its repository.
6
+
7
+ Report concerns privately to maintainer Fabio Rapposelli at [fabio@rapposelli.org](mailto:fabio@rapposelli.org), with the subject `AutoRouter conduct report`. Include relevant links and enough context to investigate; avoid publishing the report or unrelated private information. Reports will be handled with discretion, sharing information only as needed to investigate and respond.
8
+
9
+ The maintainer may request changes, remove content, issue warnings, or temporarily or permanently restrict participation according to the severity and pattern of behavior. Requests to review a decision can be sent to the same address. This policy is maintained on a best-effort basis and does not promise a response deadline.
@@ -0,0 +1,57 @@
1
+ # Contributing to AutoRouter
2
+
3
+ Bug reports, documentation improvements, reproducible routing cases and focused fixes are welcome. Read the [support guide](SUPPORT.md) before opening an issue, and use the [private security reporting process](SECURITY.md) for vulnerabilities. Participation follows the [code of conduct](CODE_OF_CONDUCT.md).
4
+
5
+ For a substantial behavior change, open an issue describing the problem and proposed scope before implementing it. Fabio Rapposelli ([@frapposelli](https://github.com/frapposelli)) maintains the project and reviews design and release decisions. Review is best effort; there is no guaranteed response time.
6
+
7
+ Coding agents working in a source checkout should follow [AGENTS.md](https://github.com/frapposelli/claude-autorouter/blob/main/AGENTS.md). [CLAUDE.md](https://github.com/frapposelli/claude-autorouter/blob/main/CLAUDE.md) imports the same guidance for Claude Code.
8
+
9
+ ## Submit a change
10
+
11
+ Fork the repository, clone your fork and create a branch. While the repository is private, this requires access and permission to fork; existing collaborators can use a branch in their authorized checkout.
12
+
13
+ ```sh
14
+ git clone https://github.com/YOUR-USERNAME/claude-autorouter.git
15
+ cd claude-autorouter
16
+ git switch -c describe-your-change
17
+ ```
18
+
19
+ Use Node.js 22+ and macOS or Linux (including WSL). Install pinned development tools, then run the local checks:
20
+
21
+ ```sh
22
+ npm ci --ignore-scripts --no-audit --no-fund
23
+ npm run check
24
+ npm test
25
+ npm run test:package
26
+ ```
27
+
28
+ Tests use synthetic local services and credentials. They require loopback binding, but make no paid provider calls, downloads or user-config changes. The package check installs and exercises the exact distributable archive. Opt-in provider/model canaries are described in [development and validation](docs/development.md).
29
+
30
+ Open a pull request against `main`. Describe the problem, resulting behavior and relevant validation, linking an issue when one exists. Keep the change focused, include meaningful regressions for behavior changes, and update affected help or documentation. Report any checks you could not run. Real provider calls and model downloads are not required for ordinary contributions; label their results separately if deliberately run.
31
+
32
+ Use synthetic fixtures. Do not commit credentials, personal configuration, private prompts, transcripts or session logs. Metadata-only logs can still contain identifying information; inspect any material before sharing it. Contributions are accepted under the project's [Apache-2.0 license](LICENSE); submit only work you have the right to contribute. No CLA or sign-off workflow is required.
33
+
34
+ ## Changing behavior
35
+
36
+ Follow the [request lifecycle](docs/development.md#request-lifecycle-and-model-continuity). Keep authentication and permission decisions owned by Claude. Preserve provider bytes, signed history and unfamiliar extensions. Routing must check compatibility in every profile; new human tasks remain eligible to switch models. Active task state is separate from disposable classification caches, and only clean, successfully forwarded completion evidence can establish confirmed continuation state.
37
+
38
+ Update `src/contracts.mjs` alongside configuration, decision or telemetry changes. Keep the normalizer allowlist and privacy tests aligned. Logged selections are not successful outcomes, logging stays opt-in, and metadata mode must omit prompts. [Static/style checks](docs/development.md#static-contracts-and-style) run in CI; do not add runtime dependencies for developer tooling.
39
+
40
+ ## Model and pricing updates
41
+
42
+ 1. Identify the exact provider model ID; a family keyword or custom alias is not capability evidence.
43
+ 2. Update `src/model-catalog.mjs` with the authoritative source and review date. Verify context/output limits, thinking, tool choice, native tool features and Auto eligibility.
44
+ 3. Add valid-source compatibility fixtures for every affected profile and a regression for the new restriction or permitted switch. Keep unknown models/extensions conservative.
45
+ 4. Update thinking adaptation only where documented; token counting and inference must apply the same compatibility policy.
46
+ 5. If rates change, review `src/savings.mjs`, bump its pricing version/date and source, and test cache TTL/modifier/unknown-model coverage. Never silently price old logs using an unrecorded new table.
47
+ 6. Run all three local checks and exact-package fixtures. Report separately any explicitly invoked real-provider observations and their limits.
48
+
49
+ ## Evaluation and performance
50
+
51
+ Declare label agreement, acceptable tiers and under-routing thresholds before testing a candidate. Preserve fixture checksums and held-out cases; transport, evaluator availability, policy and task quality are separate gates. Profile coverage requires all three tiers for compatible routing and Sonnet/Opus for Auto.
52
+
53
+ Capture a baseline before changing overhead. Repeat the same workload and hardware with the [router/storage harnesses](docs/development.md#performance-regression-measurements), retaining call counts, latency distributions, memory and background load. Numerical timing gates are local and opt-in; CI checks deterministic cancellation/resource behavior. Real local-model benchmarks need representative memory sizes and observed cold/warm conditions.
54
+
55
+ ## Releasing
56
+
57
+ Use the [release procedure](docs/releasing.md) and retain the tested immutable archive. Acceptance by npm and public availability are separate states. Verify registry integrity and an isolated install before calling a release verified. A pending submission is investigated or verified again, rather than blindly republished. New CLI interfaces and event-schema changes should be reviewed together as a minor release; correctness-only fixes can be independent patches.
package/README.md CHANGED
@@ -1,126 +1,103 @@
1
- # Claude AutoRouter
1
+ # AutoRouter
2
2
 
3
- Use Haiku, Sonnet, and Opus in one Claude Code session. A local gateway classifies coding requests with the selected evaluator, applies compatibility and context checks, and streams the selected model's response back to Claude Code. Claude's internal permission classifiers retain their selected model; execution requests keep their server safety-review settings and verdicts when routed. [TypeSafe Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) is the default; an experimental Ollama backend evaluates requests locally.
3
+ An independent local model-routing gateway for Claude Code. AutoRouter is not affiliated with, endorsed by, or sponsored by Anthropic. The existing npm package and command remain `claude-autorouter`.
4
4
 
5
- Requires Node.js 22+, macOS or Linux (including WSL), an installed `claude` command, and a Claude subscription login or Anthropic API key. The default evaluator also requires a [TypeSafe API key](https://console.typesafe.ai). There are no runtime package dependencies. Native Windows is not supported in this release.
5
+ Use Haiku, Sonnet and Opus in one Claude Code session. AutoRouter evaluates each coding request, checks model compatibility and context capacity, and forwards it through a local gateway. Native Ollama `/v1/systemone` models are the default, experimental local evaluator, so task excerpts stay on your machine; [TypeSafe Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) is an optional hosted evaluator (`setup --evaluator jev`). Claude owns authentication, tool permissions and safety review.
6
6
 
7
- ## Install and start
8
-
9
- Install from [npm](https://www.npmjs.com/package/claude-autorouter):
7
+ Requires Node.js 22+, macOS or Linux (including WSL), an installed `claude` command, and a Claude subscription login or Anthropic API key. The default evaluator also needs a [TypeSafe API key](https://console.typesafe.ai). The installed CLI has no runtime dependencies.
10
8
 
11
- ```sh
12
- npm install -g claude-autorouter
13
- ```
9
+ Version 0.4.0 adds `config`, `sessions` and `doctor --evaluate-local`, durable task continuity, and clearer model outcomes. Upgrade from 0.3.x to use these commands. The [contributor guide](CONTRIBUTING.md) explains local verification, and the [release guide](docs/releasing.md) covers the changes and verified publication.
14
10
 
15
- Set up once, then launch from any project directory:
11
+ ## Install and start
16
12
 
17
13
  ```sh
14
+ npm install -g claude-autorouter
18
15
  claude-autorouter setup
19
16
  claude-autorouter doctor
20
17
  cd /path/to/project
21
18
  claude-autorouter claude
22
19
  ```
23
20
 
24
- Setup defaults to your Claude subscription and prompts for your Jev key without echoing it. If Claude is not already signed in, run `claude auth login`. No Anthropic API key or exported subscription token is needed for subscription mode. Jev has separate credentials and billing.
21
+ Setup defaults to your Claude subscription and the local Ollama evaluator (Ollama 0.35+ with the default model; add `--pull` to download it), so it asks for no evaluator key. To use hosted Jev instead, run `claude-autorouter setup --evaluator jev`, which prompts privately for its key. Run `claude auth login` if needed. Jev has separate credentials and billing; subscription mode needs no Anthropic API key. For API billing, use `setup --auth-mode api-key`.
25
22
 
26
- Setup saves a private JSON config at `~/.config/claude-autorouter/config.json`; `XDG_CONFIG_HOME` and `AUTOROUTER_CONFIG` can change its location. Environment variables override saved configuration. Project `.env` files are not loaded automatically.
23
+ AutoRouter launches your installed, unmodified official Claude Code executable. Each user supplies their own login or API credentials. Subscription forwarding is a technical integration, not a claim of provider approval; review the [integration boundaries and current provider-policy notes](docs/subscription-integration.md) for your deployment.
27
24
 
28
- Claude Code arguments pass through:
25
+ Configuration is saved privately at `~/.config/claude-autorouter/config.json`. On macOS, new setups keep keys in the login Keychain; for an existing plaintext configuration, run `claude-autorouter config set AUTOROUTER_SECRET_STORE keychain` to move them. Environment variables override it; project `.env` files are not loaded automatically. `setup --force` updates an existing configuration while preserving other settings. Use focused commands for later edits:
29
26
 
30
27
  ```sh
31
- claude-autorouter claude -p "Fix the typo in README.md"
32
- claude-autorouter --help
33
- claude-autorouter --version
28
+ claude-autorouter config show
29
+ claude-autorouter config set AUTOROUTER_JEV_TIMEOUT_MS 2000
30
+ claude-autorouter config unset AUTOROUTER_JEV_TIMEOUT_MS
31
+ claude-autorouter help config
34
32
  ```
35
33
 
36
- For API billing, use `claude-autorouter setup --auth-mode api-key`. Use `--force` to replace an existing config. Automation can supply `TYPESAFE_API_KEY` and, in API-key mode, `ANTHROPIC_API_KEY` through the environment; keys are never command-line arguments. `doctor` checks local configuration and Claude installation/login state without paid requests. See the [configuration reference](docs/reference.md#configuration).
34
+ Secret updates use a hidden prompt or `--stdin`, never a command-line value. Claude arguments pass through, including `claude-autorouter claude --help`. [Configuration reference](docs/reference.md#configuration).
37
35
 
38
- To use automatic Sonnet/Opus routing with Claude's Auto permission mode (AutoRouter 0.3.7+):
36
+ ## Auto permission mode
39
37
 
40
38
  ```sh
41
39
  claude-autorouter claude --permission-mode auto
42
40
  ```
43
41
 
44
- This selects the Auto profile, defaulting to Sonnet 5.5 and Opus 5.5. The evaluator can choose again for each new human task, while a task's tool calls and goal continuations retain its selected model. A Haiku verdict uses Sonnet. Native safety review stays enabled; organization policies still apply. Use `AUTOROUTER_CLIENT_PROFILE=auto` for sessions where you select Auto in Claude's UI or saved settings. Version 0.3.6 enabled Auto permissions but passed server-reviewed execution through without routing; version 0.3.7 removes that restriction for compatible requests. [Auto-mode support and limitations](docs/reference.md#auto-permission-mode).
42
+ This profile automatically switches between Sonnet 5.5 and Opus 5.5 for new human tasks. A Haiku verdict uses Sonnet. Tool and `/goal` continuations retain the task's execution model; a new task can switch up or down. Claude's native safety review and organization policies still apply. For Auto selected through Claude's UI, save `AUTOROUTER_CLIENT_PROFILE=auto` with `config set`. [Auto support and limitations](docs/reference.md#auto-permission-mode).
45
43
 
46
- ## What you see
44
+ ## Inspect decisions
47
45
 
48
- The launcher adds a temporary status line and leaves saved Claude Code settings unchanged:
46
+ The launcher adds a temporary status line, preserving saved Claude settings:
49
47
 
50
48
  ```text
51
- ● AutoRouter · last Haiku 4.5 · ready · Jev 210ms · est saved $0.04 (75%) vs Opus
52
- ● AutoRouter · Sonnet 5 selected · connecting · Jev→Haiku 290ms · large context
49
+ ● AutoRouter · Opus 5.5 · ready · Jev 210ms
50
+ ● AutoRouter · Sonnet 5.5 selected · Auto floor from Haiku · Jev 220ms
53
51
  ```
54
52
 
55
- The confirmed model comes from Anthropic's response. Claude's own model label can still show its Haiku starting model. `API ctx` measures input against the actual model's known window; a different client limit remains visible as `CLI ctx`.
53
+ `selected` means Anthropic has not reported the serving model yet. Guard reasons and errors stay visible before optional savings. Claude's own model label may show its starting model. [Status details](docs/reference.md#status-line-and-savings).
56
54
 
57
- For a separate decision log per session (optional, disabled by default; AutoRouter 0.3.6+):
55
+ Persistent history is optional and disabled by default. Enable metadata-only records without prompt excerpts:
58
56
 
59
57
  ```sh
60
- env AUTOROUTER_SESSION_LOG_DIR="$HOME/.local/state/claude-autorouter/sessions" \
61
- claude-autorouter claude
58
+ claude-autorouter config set AUTOROUTER_SESSION_LOG_MODE metadata
59
+ claude-autorouter config set AUTOROUTER_SESSION_LOG_DIR "$HOME/.local/state/claude-autorouter/sessions"
60
+ claude-autorouter claude
61
+ claude-autorouter sessions list
62
+ claude-autorouter sessions show ID --json
62
63
  ```
63
64
 
64
- Each JSONL record includes a bounded prompt excerpt, selected model, decision latency, and routing reason. Files persist after Claude exits; terminal output stays quiet. [Session logs](docs/reference.md#session-decision-logs).
65
-
66
- Savings are an **API-equivalent estimate for the same token counts**, using Opus as the baseline. They do not measure subscription bill reductions or quota credits and exclude Jev and local compute costs. [Status line and savings details](docs/reference.md#status-line-and-savings).
67
-
68
- ## Experimental local evaluator
69
-
70
- The local setup below requires AutoRouter 0.3.2 or newer. It uses Ollama's native `/v1/systemone` decision API with `nimble:9b-q4_K_M` by default. Jev remains the default evaluator. If upgrading from 0.2.0, replace the old Qwen model configuration using the [migration steps](docs/reference.md#migrating-an-older-ollama-config).
65
+ Copy an `id` from `list`. History separates model decisions from outcomes and reports latency, fallbacks, failures and savings coverage. Choose `prompts` mode for bounded human-task excerpts. Files persist locally; no automatic deletion occurs. [History and privacy](docs/reference.md#session-decision-logs).
71
66
 
72
- Version 0.3.2 excludes Claude's executor system instructions from the local classifier excerpt, retaining task and conversation excerpts. Runtime deadlines default to 1,500 ms for Tev1 0.8B/custom models, 15,000 ms for official Tev1 4B tags, and 30,000 ms for official Nimble tags. Explicit timeout settings, including a `1500` saved with 0.3.1, still override these defaults. Jev is unchanged.
67
+ Savings are **API-equivalent estimates using the recorded Opus baseline and token counts**. They do not measure subscription bill reductions or quota credits, and exclude evaluator and local compute costs. Missing or unsupported usage stays unpriced.
73
68
 
74
- Set `0` to disable AutoRouter's runtime evaluator deadline for one launch using your existing configuration:
69
+ ## Local Ollama evaluator
75
70
 
76
- ```sh
77
- AUTOROUTER_OLLAMA_TIMEOUT_MS=0 claude-autorouter claude
78
- ```
79
-
80
- To save that setting for an installed Tev1 4B model:
71
+ Ollama is the default evaluator. Start Ollama 0.35+ with a model supporting its native decision endpoint, then choose a model:
81
72
 
82
73
  ```sh
83
- claude-autorouter setup --evaluator ollama --ollama-model tev1:4b --ollama-timeout-ms 0 --force
74
+ claude-autorouter setup --evaluator ollama --ollama-model tev1:4b-q4_K_M --pull --force
75
+ claude-autorouter doctor --evaluate-local
76
+ claude-autorouter claude
84
77
  ```
85
78
 
86
- The setup flag overrides the timeout environment value and saves it. Cancellation and disconnected clients still stop evaluation, normal errors still use fallback, and startup priming keeps its separate 60-second deadline.
79
+ `--pull` authorizes downloading the chosen model if missing. Setup keeps existing models and settings; ordinary launches download nothing. The default local model is `nimble:9b-q4_K_M`; `tev1:0.8b` is smaller and requires checking its accuracy on your tasks. Local classification needs no Jev key; Jev remains available with `setup --evaluator jev`. Claude still answers through Anthropic. [Model choices, deadlines and historical measurements](docs/reference.md#ollama-evaluator).
87
80
 
88
- Install and start Ollama 0.35 or newer; [version 0.35.0](https://github.com/ollama/ollama/releases/tag/v0.35.0) is a prerelease as of September 29, 2026. Then run:
81
+ To allow a slower local model to finish without AutoRouter's runtime deadline:
89
82
 
90
83
  ```sh
91
- claude-autorouter setup --evaluator ollama --pull --force
92
- claude-autorouter doctor
93
- claude-autorouter claude
84
+ claude-autorouter config set AUTOROUTER_OLLAMA_TIMEOUT_MS 0
94
85
  ```
95
86
 
96
- `--force` replaces existing AutoRouter configuration. `--pull` downloads the selected model only if missing. Setup does not install or start Ollama, or delete existing models. Select a native decision model explicitly with `--ollama-model`:
87
+ Cancellation and response-size limits still apply. The local diagnostic uses synthetic prompts and reports observed latency, classification and fallback reasons; it makes no Anthropic/Jev calls or downloads and preserves unrelated resident models.
97
88
 
98
- | Model | Approximate download | Selection |
99
- | --- | ---: | --- |
100
- | [Nimble 9B Q4_K_M](https://ollama.com/library/nimble) | 5.63 GB | Default: `nimble:9b-q4_K_M` |
101
- | [Tev1 0.8B Q8](https://ollama.com/library/tev1) | 812 MB | `tev1:0.8b` |
102
- | [Tev1 4B Q4_K_M](https://ollama.com/library/tev1) | 2.7 GB | `tev1:4b-q4_K_M` |
89
+ ## Troubleshoot and upgrade
103
90
 
104
- For example, select Tev1 0.8B with:
91
+ Inspect `config show` for environment overrides, and `doctor` for setup health. A valid evaluator verdict can be overridden by continuity, context or model compatibility. `Ollama fallback: timeout` means evaluation failed to finish, rather than predicting Sonnet. Status errors and saved history explain these paths.
105
92
 
106
93
  ```sh
107
- claude-autorouter setup --evaluator ollama --ollama-model tev1:0.8b --pull --force
94
+ npm install -g claude-autorouter@latest
95
+ claude-autorouter --version
96
+ claude-autorouter doctor
108
97
  ```
109
98
 
110
- Use `--ollama-model tev1:4b-q4_K_M` for the listed 4B variant; `tev1:latest` and `tev1:4b` select the larger Q8 download. Model terms are linked in the listings above; download size does not measure resident memory or routing quality. Custom native model tags and aliases also work.
111
-
112
- No Jev key is needed for local classification. The launcher primes the evaluator before opening Claude's UI, and evaluation failures fall back to Sonnet or retain Opus without contacting Jev. Claude still answers through Anthropic, with the same routing guards and subscription limits.
113
-
114
- Setup, doctor, and startup show the effective model and deadline. Warmup and doctor do not certify classification speed or accuracy. `Ollama fallback: timeout` means no valid decision arrived in time; it is different from the evaluator choosing Sonnet. Source users can run the [local routing regression](docs/development.md#local-routing-regression) to check all three tiers without Claude or Jev calls.
115
-
116
- Historical measurements before 0.3.2, on a 16 GiB M4: Tev1 0.8B matched 18/24 held-out labels with 450 ms median latency and no timeouts at 1,500 ms, including full-excerpt checks. Tev1 4B matched 22/24 with a 10-second diagnostic deadline and 3.15-second median latency. Nimble matched 23/24 with a 30-second deadline and 11.4-second median latency. Both larger models exceeded the then-default 1,500 ms. A separate six-case regression with the 0.3.2 fixes passed for both tested Tev1 4B variants and Nimble; Tev1 0.8B matched only three cases. These small tests do not establish general accuracy or Jev parity. See the [measurements and limits](docs/ollama-evaluation.md) and [Ollama reference](docs/reference.md#ollama-evaluator).
117
-
118
- ## Behavior and data
99
+ Historical integration observations cover Claude Code 2.1.284–2.1.285. The versioned synthetic protocol fixtures test reviewed request/response contracts; they do not certify the current checkout against a live Claude version. Real-provider checks remain explicitly invoked. [Troubleshooting](docs/reference.md#troubleshooting) covers context use, blocked goals and logging. Run ordinary `claude` to bypass routing.
119
100
 
120
- - The default client profile permits all three routing tiers. Tool continuations, thinking history, model-specific features, and context size can keep or upgrade a model even when the evaluator chooses a cheaper tier. [Routing policy](docs/reference.md#routing-policy).
121
- - The selected evaluator receives bounded excerpts that can contain source code and tool results: TypeSafe with Jev, or the local service with Ollama. Jev also receives system-text excerpts; the local path excludes Claude's executor system instructions. Anthropic receives the complete request. Images, document payloads, and private thinking are omitted from classifier input. [Data flow and authentication](docs/reference.md#data-flow-and-authentication).
122
- - Subscription access and usage limits still apply. Model switches can reduce cache reuse; cheaper token prices do not guarantee cheaper completed tasks. Run ordinary `claude` to bypass routing.
123
- - The launcher is quiet by default. Use `AUTOROUTER_DEBUG=1` for metadata diagnostics or `AUTOROUTER_STATUSLINE=0` to retain your existing status line. [Troubleshooting](docs/reference.md#troubleshooting).
124
- - For blocked `/goal` loops, optionally launch with `env CLAUDE_CODE_STOP_HOOK_BLOCK_CAP=2 claude-autorouter claude`. Claude then ends the turn on the third consecutive blocking verdict without tool use, leaving the goal unmet. This also affects other Stop/SubagentStop hooks; defaults are unchanged. [Scope and saved configuration](docs/reference.md#shorter-stop-hook-loops-opt-in).
101
+ The evaluator receives bounded task/history excerpts that may contain code and tool results: TypeSafe for Jev, or your loopback Ollama service. Recognizable credentials and personal identifiers are redacted from those excerpts first. Anthropic receives the complete request. Model switching can reduce cache reuse. [Data flow and authentication](docs/reference.md#data-flow-and-authentication).
125
102
 
126
- [Reference](docs/reference.md) · [Development and validation](docs/development.md) · [CI and npm release setup](docs/releasing.md) · [Apache-2.0 license](LICENSE)
103
+ [Reference](docs/reference.md) · [Integration and provider policy](docs/subscription-integration.md) · [Contributing](CONTRIBUTING.md) · [Development](docs/development.md) · [Releases](docs/releasing.md) · [Apache-2.0](LICENSE)
package/SECURITY.md ADDED
@@ -0,0 +1,23 @@
1
+ # Security policy
2
+
3
+ ## Report privately
4
+
5
+ Use GitHub's private [Report a vulnerability](https://github.com/frapposelli/claude-autorouter/security/advisories/new) form. Private vulnerability reporting is enabled for this repository. If the form is unavailable, email maintainer Fabio Rapposelli at [fabio@rapposelli.org](mailto:fabio@rapposelli.org) with the subject `AutoRouter security report`.
6
+
7
+ Do not open a public issue or pull request with exploit details, credentials or sensitive request data. Include the affected AutoRouter and Claude Code versions, operating system, evaluator/client profile, expected security boundary, observed impact and a minimal synthetic reproduction where possible. Do not include real API keys, OAuth tokens, private source code, prompts or transcripts. The maintainer can coordinate any additional evidence privately.
8
+
9
+ Reports are handled on a best-effort basis; there is no guaranteed response time or bounty program. We will coordinate investigation, remediation and disclosure with the reporter before publishing details.
10
+
11
+ ## Supported versions and scope
12
+
13
+ Security fixes target the latest stable release. Older releases are not maintained as separate security branches; users may need to upgrade. Reports against `main` are also welcome.
14
+
15
+ Relevant reports include credential exposure, unauthorized access to the local gateway or saved configuration, unintended disclosure through logs, and routing or request changes that weaken Claude's authentication or permission boundaries. Provider accounts, billing and vulnerabilities in Claude Code, TypeSafe or Ollama should also be reported to the responsible provider when applicable.
16
+
17
+ ## Handling diagnostic data
18
+
19
+ AutoRouter's default evaluator is local Ollama, which keeps classification on loopback. The optional hosted Jev evaluator (`--evaluator jev`) receives bounded task/history excerpts, which may contain private code or tool results. Recognizable credentials and personal identifiers are redacted from those excerpts first; this pattern-based filter reduces, but does not eliminate, disclosure. Anthropic still receives the full inference request. See [data flow and authentication](docs/reference.md#data-flow-and-authentication).
20
+
21
+ New macOS setups keep saved keys in the login Keychain. Existing and non-macOS configurations keep plaintext keys in the private configuration file until you run `claude-autorouter config set AUTOROUTER_SECRET_STORE keychain` (macOS only). See [credential storage](docs/reference.md#credential-storage).
22
+
23
+ Session logging is optional. Enabling only a log directory uses the default `prompts` mode, which includes bounded human-task excerpts with recognizable credentials and personal identifiers redacted (pattern-based, so not exhaustive). Select `AUTOROUTER_SESSION_LOG_MODE=metadata` before enabling a directory to omit those excerpts. Inspect even metadata-only output before sharing it; identifiers, paths or environment details can still be sensitive. Saved logs have no automatic deletion policy. See [history and privacy](docs/reference.md#session-decision-logs).
package/SUPPORT.md ADDED
@@ -0,0 +1,18 @@
1
+ # Support
2
+
3
+ AutoRouter is an independent, community-maintained project. It does not provide official support for Anthropic, TypeSafe or Ollama, and has no response-time guarantee.
4
+
5
+ For setup and usage, start with the [README](README.md) and [troubleshooting reference](docs/reference.md#troubleshooting). Run `claude-autorouter --version`, `claude --version` and `claude-autorouter doctor` to identify the installation and configuration involved. Ordinary `doctor` performs configuration and local service checks; it does not verify paid-provider access.
6
+
7
+ Use [GitHub issues](https://github.com/frapposelli/claude-autorouter/issues) for reproducible bugs, feature proposals and questions not answered by the documentation. Repository access is required while the repository is private; if you cannot access it, contact [fabio@rapposelli.org](mailto:fabio@rapposelli.org). For vulnerabilities, follow [SECURITY.md](SECURITY.md) instead of opening an issue. Community conduct reports follow the [code of conduct](CODE_OF_CONDUCT.md).
8
+
9
+ ## Make a report useful
10
+
11
+ - Include AutoRouter, Claude Code, Node.js and operating-system versions; include the Ollama version and model tag for local evaluation.
12
+ - Identify the authentication mode, evaluator and client profile without sharing credentials or full environment/configuration dumps.
13
+ - Describe expected and actual behavior, and provide minimal steps using a synthetic prompt or fixture when possible. Distinguish the selected model from the provider-confirmed serving model.
14
+ - Share only the relevant, inspected diagnostic excerpt. Remove secrets, private prompts, responses, source code, personal paths and organization identifiers. Screenshots can disclose this information too.
15
+
16
+ Logging is disabled by default. When enabling optional session history for a reproduction, explicitly choose `AUTOROUTER_SESSION_LOG_MODE=metadata`; the default `prompts` mode includes task excerpts. Metadata-only output still needs review before sharing. Debug stderr can also contain Claude's own diagnostics. See [session logs](docs/reference.md#session-decision-logs) and [troubleshooting](docs/reference.md#troubleshooting).
17
+
18
+ Account access, subscription/model eligibility, billing and provider outages belong with the responsible provider. AutoRouter reports can investigate how the gateway handles those failures, but cannot change a provider's account policies. Local-model accuracy and performance depend on the workload and hardware; include those conditions when reporting unexpected classifications.
@@ -12,66 +12,48 @@ import { createSessionLog } from '../src/session-log.mjs';
12
12
  import { loadUserConfig } from '../src/user-config.mjs';
13
13
  import { setup, doctor, ollamaDeadlineText } from '../src/onboarding.mjs';
14
14
  import { setupOllama } from '../src/ollama-setup.mjs';
15
+ import { helpText } from '../src/cli-help.mjs';
16
+ import { configCommand } from '../src/config-command.mjs';
17
+ import { sessionsCommand } from '../src/session-history.mjs';
15
18
 
16
19
  const [command = 'help', ...args] = process.argv.slice(2);
17
20
  if (['--version', '-v', 'version'].includes(command)) {
18
21
  console.log(JSON.parse(readFileSync(new URL('../package.json', import.meta.url), 'utf8')).version);
19
22
  } else if (['help', '--help', '-h'].includes(command)
20
- || (['setup', 'doctor', 'serve'].includes(command) && args.some(arg => ['--help', '-h'].includes(arg)))) {
21
- console.log(`Claude AutoRouter — routing for Haiku, Sonnet, and Opus
22
-
23
- Usage:
24
- claude-autorouter setup [--auth-mode subscription|api-key] [--force]
25
- [--client-profile compatible|native|auto]
26
- [--evaluator jev|ollama]
27
- [--ollama-model MODEL] [--ollama-timeout-ms N] [--pull]
28
- [--stop-hook-block-cap N]
29
- [--session-log-dir DIR]
30
- claude-autorouter doctor
31
- claude-autorouter claude [Claude Code arguments]
32
- claude-autorouter serve
33
- claude-autorouter --version
34
-
35
- Setup defaults to subscription authentication and prompts for keys without echoing.
36
- For noninteractive setup, supply keys through environment variables.
37
- User config: ~/.config/claude-autorouter/config.json (or XDG_CONFIG_HOME).
38
- AUTOROUTER_CONFIG selects a different file; environment variables take precedence.
39
- Project .env files are never loaded automatically.
40
-
41
- Jev is the default evaluator and requires TYPESAFE_API_KEY.
42
- Ollama evaluates locally and requires Ollama 0.35+ with /v1/systemone.
43
- Use setup --evaluator ollama --pull to detect Ollama and download a missing model.
44
- The local default is nimble:9b-q4_K_M; --ollama-model selects another compatible model.
45
- Smaller Tev1 options: --ollama-model tev1:0.8b or --ollama-model tev1:4b-q4_K_M.
46
- Local routing deadlines: Tev1 0.8B/custom 1500 ms, Tev1 4B 15000 ms, Nimble 30000 ms.
47
- Setup --ollama-timeout-ms N saves a routing deadline; use 0 to disable it.
48
- AUTOROUTER_OLLAMA_TIMEOUT_MS also overrides the deadline; 0 disables it.
49
- Local routing is experimental; see docs/ollama-evaluation.md for measured limits.
50
- AUTOROUTER_AUTH_MODE=subscription uses your saved Claude Code login.
51
- Without setup, AUTOROUTER_AUTH_MODE defaults to api-key and also requires ANTHROPIC_API_KEY.
52
- AUTOROUTER_CLIENT_PROFILE=compatible (default) enables all three routing tiers.
53
- Use AUTOROUTER_CLIENT_PROFILE=native to retain Claude Code's own model/thinking settings.
54
- Use AUTOROUTER_CLIENT_PROFILE=auto for Auto permission mode: Sonnet/Opus routing, native thinking.
55
- Auto defaults to Sonnet 5.5 and Opus 5.5, switching on new human tasks and retaining tool turns.
56
- An explicit claude --permission-mode auto selects the auto profile for that launch.
57
- Claude's permission checks and organization policies still apply; Haiku does not support Auto mode.
58
- Optional CLAUDE_CODE_STOP_HOOK_BLOCK_CAP=N limits consecutive tool-free Stop-hook continuations.
59
- Use 2 to stop on the third block; applies to /goal and all Stop/SubagentStop hooks.
60
- Unset preserves Claude's default; 0 disables the cap. Setup --stop-hook-block-cap N saves it.
61
- Standalone serve also requires AUTOROUTER_TOKEN (at least 16 characters).
62
- The claude launcher creates a temporary credential and an ephemeral port.
63
- It enables an AutoRouter status line for this session (AUTOROUTER_STATUSLINE=0 to opt out).
64
- Launcher logs are quiet by default; AUTOROUTER_DEBUG=1 enables diagnostic logs on stderr.
65
- AUTOROUTER_SESSION_LOG_DIR writes private per-session JSONL decision logs with prompt excerpts.
66
- Unset or empty disables session logs. Setup --session-log-dir DIR saves the directory.
67
- Jev sends prompt excerpts to TypeSafe; Ollama keeps classification on this machine.
68
- Complete inference requests still go to Anthropic. See README.md.`);
69
- } else if (command === 'setup' || command === 'doctor') {
23
+ || (['setup', 'doctor', 'serve', 'config', 'sessions'].includes(command) && args.some(arg => ['--help', '-h'].includes(arg)))) {
24
+ console.log(helpText(command === 'help' ? args[0] : command));
25
+ } else if (command === 'claude' && args.length === 1 && ['--help', '-h', '--version', '-v'].includes(args[0])) {
26
+ // Help/version are Claude-owned commands. No config file, evaluator keys,
27
+ // gateway, Ollama warmup or temporary status files are needed.
28
+ const env = { ...process.env };
29
+ delete env.TYPESAFE_API_KEY;
30
+ delete env.AUTOROUTER_TOKEN;
31
+ const child = spawn('claude', args, { stdio: 'inherit', env });
32
+ child.once('error', () => { console.error('Could not launch Claude Code. Ensure `claude` is installed and on PATH.'); process.exitCode = 1; });
33
+ child.once('exit', (code, signal) => { process.exitCode = code ?? (signal === 'SIGINT' ? 130 : signal === 'SIGTERM' ? 143 : 1); });
34
+ for (const signal of ['SIGINT', 'SIGTERM']) process.on(signal, () => child.kill(signal));
35
+ } else if (['setup', 'doctor', 'config', 'sessions'].includes(command)) {
70
36
  try {
71
37
  if (command === 'setup') await setup(args);
38
+ else if (command === 'config') {
39
+ if (await configCommand(args) === false) process.exitCode = 1;
40
+ }
41
+ else if (command === 'sessions') {
42
+ if (await sessionsCommand(args) === false) process.exitCode = 1;
43
+ }
72
44
  else {
73
- if (args.length) throw new Error('Usage: claude-autorouter doctor');
74
- if (!await doctor()) process.exitCode = 1;
45
+ const evaluateLocal = args.includes('--evaluate-local');
46
+ const json = args.includes('--json');
47
+ if (args.some(arg => !['--evaluate-local', '--json'].includes(arg)) || new Set(args).size !== args.length
48
+ || (json && !evaluateLocal)) throw new Error('Usage: claude-autorouter doctor [--evaluate-local [--json]]');
49
+ const controller = new AbortController();
50
+ const cancel = () => controller.abort();
51
+ for (const signal of ['SIGINT', 'SIGTERM']) process.once(signal, cancel);
52
+ try {
53
+ if (!await doctor({ evaluateLocal, json, signal: controller.signal })) process.exitCode = 1;
54
+ } finally {
55
+ for (const signal of ['SIGINT', 'SIGTERM']) process.removeListener(signal, cancel);
56
+ }
75
57
  }
76
58
  } catch (error) { console.error(error.message); process.exitCode = 1; }
77
59
  } else if (!['claude', 'serve'].includes(command)) {
@@ -85,10 +67,9 @@ Complete inference requests still go to Anthropic. See README.md.`);
85
67
  const stop = () => {
86
68
  if (stopping) return stopping;
87
69
  if (server) { server.close(); server.closeAllConnections(); }
88
- status?.close();
89
70
  // Drain accepted decision records before normal process exit. Pending
90
71
  // filesystem writes keep Node alive; no timer or fire-and-forget buffer.
91
- stopping = Promise.resolve().then(() => sessionLog?.close()).catch(() => {});
72
+ stopping = Promise.allSettled([status?.close(), sessionLog?.close()]);
92
73
  return stopping;
93
74
  };
94
75
  try {
@@ -120,15 +101,17 @@ Complete inference requests still go to Anthropic. See README.md.`);
120
101
  let claudeArgs = args;
121
102
  if (statusEnabled) {
122
103
  status = createStatusState({ baselineModel: config.models.opus });
104
+ await status.ready;
123
105
  if (status.path) {
124
106
  try { claudeArgs = addStatusLineSettings(args, dirname(status.path)); }
125
107
  catch {
126
- status.close(); status = undefined;
108
+ await status.close(); status = undefined;
127
109
  console.error('AutoRouter status line unavailable: could not safely prepare session settings. Passing your original settings to Claude.');
128
110
  }
129
111
  } else console.error('AutoRouter status line unavailable: could not create local status storage.');
130
112
  }
131
113
  if (config.sessionLogDir) sessionLog = await createSessionLog(config.sessionLogDir, {
114
+ includePrompts: config.sessionLogMode === 'prompts',
132
115
  warn: message => console.error(message),
133
116
  });
134
117
  // Claude owns the terminal while its UI is running. Status updates use the
@@ -136,7 +119,7 @@ Complete inference requests still go to Anthropic. See README.md.`);
136
119
  server = createRouterServer(config, {
137
120
  log: diagnosticLogs ? undefined : () => {},
138
121
  onStatus: event => status?.update(event),
139
- onDecision: sessionLog ? entry => sessionLog.record(entry) : undefined,
122
+ onRecord: sessionLog ? entry => sessionLog.record(entry) : undefined,
140
123
  });
141
124
  const address = await listen(server, command === 'claude' ? 0 : config.port);
142
125
  const baseUrl = `http://127.0.0.1:${address.port}`;
@@ -155,7 +138,7 @@ Complete inference requests still go to Anthropic. See README.md.`);
155
138
  if (status?.path) env.AUTOROUTER_STATUS_FILE = status.path;
156
139
  const child = spawn('claude', claudeArgs, { stdio: 'inherit', env });
157
140
  child.once('error', () => { console.error('Could not launch Claude Code. Ensure `claude` is installed and on PATH.'); process.exitCode = 1; stop(); });
158
- child.once('exit', (code, signal) => { process.exitCode = code ?? (signal === 'SIGINT' ? 130 : 1); stop(); });
141
+ child.once('exit', (code, signal) => { process.exitCode = code ?? (signal === 'SIGINT' ? 130 : signal === 'SIGTERM' ? 143 : 1); stop(); });
159
142
  for (const signal of ['SIGINT', 'SIGTERM']) process.on(signal, () => child.kill(signal));
160
143
  }
161
144
  } catch (error) {
@@ -1,6 +1,6 @@
1
1
  # Development and validation
2
2
 
3
- Use Node.js 22+ from a source checkout on macOS or Linux (including WSL). The project has no runtime package dependencies. Development scripts and tests are separate from the installed CLI; user setup is covered in the [README](../README.md).
3
+ Use Node.js 22+ from a source checkout on macOS or Linux (including WSL). The installed CLI has no runtime package dependencies. Source checks use pinned TypeScript and Node type definitions; install these contributor tools with `npm ci --ignore-scripts --no-audit --no-fund`. Development scripts and tests are separate from the installed CLI; user setup is covered in the [README](../README.md).
4
4
 
5
5
  ## Local checks
6
6
 
@@ -16,6 +16,24 @@ Auto-mode regressions exercise Sonnet → Opus → Opus tool continuation → So
16
16
 
17
17
  Package validation checks the distributable and installed command rather than relying on the source checkout's paths. Review the [release procedure](releasing.md) before distributing a tarball.
18
18
 
19
+ ## Request lifecycle and model continuity
20
+
21
+ The gateway validates the request containers it consumes, evaluates the current task, applies continuity and capacity rules, then checks the proposed target against the shared model catalog. Unknown provider extensions remain intact; when their compatibility with another model is unknown, the source model is retained with a routing reason. Token-count requests apply the same compatibility rules before sending a request.
22
+
23
+ `src/model-catalog.mjs` contains exact model IDs, capability facts, source links, and a review date. A family name inside a custom alias does not establish capabilities. `src/model-request.mjs` contains explicit thinking adaptations; neither layer strips signed history or permission-review settings. Model updates should change the catalog, include a dated authoritative source, and add a request fixture demonstrating the restriction or new supported switch.
24
+
25
+ The evaluation cache and active execution state have separate lifetimes. `src/turn-state.mjs` keeps active human tasks and pending tools beyond the evaluator cache TTL. Retired tasks expire, aliases and records are bounded, and exhausting active-state capacity is reported rather than silently evicting another active task. State is process-local; after a gateway restart, missing continuity is reported as unknown until a successful response establishes it again.
26
+
27
+ The gateway stages a selected model under its request ID. `src/response-observer.mjs` observes serving models, provider fallback boundaries, closed tool calls, usage, and terminal response metadata without altering bytes. A serving-model observation alone is not a successful execution. Only clean completion evidence followed by successful HTTP forwarding commits the continuation model. Cancelled, failed, ambiguous, and superseded attempts cannot overwrite known state. New human tasks remain eligible for upward or downward switching, including Sonnet/Opus in Auto mode.
28
+
29
+ Direct `Router.route()` embedders that omit `requestId` retain selected, unconfirmed continuity for compatibility. Embedders that execute inference should supply a unique request ID and call `router.complete(id, evidence)` after successful delivery, or `router.complete(id)` on failure. The HTTP gateway owns that lifecycle automatically.
30
+
31
+ ## Evaluation acceptance
32
+
33
+ Evaluation reports distinguish evaluator availability, rubric agreement, routing policy, tier coverage, transport, and independently checked task completion. An unmeasured gate is explicitly marked unmeasured. Normal runs cannot pass solely on classifier fallback or cached predictions; simulated-outage runs explicitly require fallback. Compatible routing requires all three selected tiers, while Auto requires Sonnet and Opus. Constrained fixtures declare expected guard overrides.
34
+
35
+ The general and Ollama evaluation scripts accept `--min-agreement` and `--max-under-route-rate`. Their defaults require complete expected-label agreement and no under-routing. Set any alternative thresholds **before** evaluating a candidate, retain the fixture checksum with the report, and keep tuning cases separate from held-out cases. Rubric labels are judgments about synthetic tasks; these reports do not prove end-user task quality or subscription savings. The expanded corpus includes multilingual tasks, short difficult follow-ups, ordinary work in long background context, and task text containing tier-selection instructions. No prompt-policy adjustment should be justified by rerunning and relabeling the held-out set.
36
+
19
37
  ## Run from source
20
38
 
21
39
  ```sh
@@ -76,7 +94,7 @@ For a classifier-only rubric evaluation:
76
94
  npm run eval
77
95
  ```
78
96
 
79
- The bundled evaluation makes 12 classifier calls and no Claude generations. Jev is the default and incurs TypeSafe usage; set `AUTOROUTER_EVALUATOR=ollama` to evaluate an installed local model. It reports agreement with the starting rubric, fallback count, and p50/p95 routing latency. Edit `test/fixtures/routing.json` to represent the tasks you want to measure. Rubric agreement alone does not establish answer quality or net savings; compare completed tasks against fixed-model baselines.
97
+ The bundled evaluation makes 19 classifier calls and no Claude generations. Jev is the default and incurs TypeSafe usage; set `AUTOROUTER_EVALUATOR=ollama` to evaluate an installed local model. It reports agreement with the starting rubric, fallback count, and p50/p95 routing latency. Edit `test/fixtures/routing.json` to represent the tasks you want to measure. Rubric agreement alone does not establish answer quality or net savings; compare completed tasks against fixed-model baselines.
80
98
 
81
99
  For local evaluator measurements, use Ollama 0.35+ and a model compatible with `/v1/systemone`. Distinguish cold model loading from warmed classification, and record the model tag, hardware, Ollama version, context size, prompt length, and resident memory. The launcher primes the classifier with a synthetic task before opening the UI, with a separate deadline of up to 60 seconds. Runtime and benchmark share the 3,000-character/3,000-UTF-8-byte state limit, so include non-ASCII cases and excerpts that fill the budget. Also measure the first request after keep-alive expiration: its reload can hit the normal deadline even when warm requests pass. Repeat on realistic prompt distributions instead of selecting a model from a single easy request. Disk download size is not resident RAM. Keep model downloads opt-in and respect each model's license.
82
100
 
@@ -113,3 +131,31 @@ node --env-file=.env scripts/context-probe.mjs --cwd /path/to/synthetic-fixture
113
131
  ```
114
132
 
115
133
  Private repository or connected-tool context may be present even when the typed prompt is harmless. Keep private-payload investigations local unless external processing is authorized. For shareable live regressions, prefer the isolated synthetic fixtures above. The probe report itself persists only metadata.
134
+
135
+ ## Static contracts and style
136
+
137
+ `npm run check` checks JavaScript syntax, TypeScript/JSDoc contracts for configuration, classifier results, final routing decisions and normalized telemetry, and literal event producers in transport code. `src/contracts.mjs` is the shared development-time type contract; normalizers remain the runtime privacy boundary. Negative fixtures in `test/static-contracts.mts` and `test/static-checks.test.mjs` prove misspelled fields, invalid enums, payload fields and timing strings are rejected before execution. Provider request extensions remain opaque and are validated only where the router consumes them.
138
+
139
+ Use two-space indentation, LF endings, one final newline, semicolons and single quotes for ordinary strings. Compact pure helpers are allowed when readable; do not reformat unrelated code. The style check rejects trailing whitespace, tab indentation, `var`, and coercing comparisons except deliberate null/undefined checks. TypeScript is a contributor dependency only; public packages keep zero runtime dependencies. CI installs the pinned lockfile before checks and never runs provider inference automatically.
140
+
141
+ ## Versioned protocol evidence
142
+
143
+ `test/fixtures/claude-protocol-v1.json` is a versioned, newly authored synthetic corpus reviewed against the Messages API, streaming, deferred-tool and fallback contracts. `test/protocol-fixtures.test.mjs` sends it through the real gateway, router and response observer with fake evaluator/upstream services. It covers Auto floors and switches, thinking adaptation and preservation, custom deferred tool references, conservative built-in server-tool history, compaction, scoped goal feedback, parallel agents, model fallback/tool ownership, usage and truncated responses.
144
+
145
+ Historical Claude Code 2.1.284/2.1.285 report versions, dates and hashes are separate metadata. The corpus does not copy captured prompts, invent provider signatures or certify a live client. Current-source real-provider canaries remain opt-in and unmeasured unless a separate report records them.
146
+
147
+ ## Performance regression measurements
148
+
149
+ [Router measurements](router-performance.md) and [status-storage measurements](status-performance.md) record the pre-change baseline, repeated candidate runs, hardware/background load and baseline-derived gates. Source-only harnesses use synthetic inputs and providers; no credentials, prompts from user sessions or downloads are involved.
150
+
151
+ ```sh
152
+ node --expose-gc scripts/benchmark-router.mjs baseline /tmp/router-comparison.json
153
+ node --expose-gc scripts/benchmark-router.mjs candidate /tmp/router-comparison.json --check
154
+ node scripts/benchmark-status.mjs --label local-check --check
155
+ ```
156
+
157
+ Identical concurrent classifier inputs share one bounded evaluation. Each request applies its own continuity, capacity and compatibility checks. Cancelling one waiter preserves other waiters; cancelling all releases the shared evaluation. Cache identity retains the complete request, requested model floor, evaluator configuration and rubric hash. New tasks and sequential pinned continuations still evaluate; prior-pin classification reuse was deliberately not enabled. A 64 KiB response limit applies to both evaluators, and local model metadata is limited to 1 MiB. Disabling the Ollama timer does not disable cancellation or these byte limits.
158
+
159
+ Status writes use one asynchronous writer and a coalesced latest snapshot. Embedders await `state.ready` before using its path, `flush()` when they need persisted evidence, and `close()` before cleanup. Initial storage failures or a one-second readiness timeout disable the optional display. Accepted in-flight writes finish before directory removal, so shutdown cannot recreate files.
160
+
161
+ The router/storage measurements cover one 16 GiB M4 with synthetic providers. Actual Tev1 4B and Nimble 9B measurements on that Mac and a 64 GiB M2 Ultra are recorded separately in [the hardware comparison](hardware-comparison.md). Both candidates missed the unchanged strict quality gate on both hosts. Model digests and runtime conditions differ, so the cross-host results do not isolate RAM's effect. Use the [transfer bundle instructions](hardware-benchmark.md) to reproduce the workload, recording background load and observed residency. Do not infer model performance from router timings.