claude-autorouter 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.env.example +45 -0
- package/LICENSE +202 -0
- package/README.md +87 -0
- package/bin/autorouter.mjs +136 -0
- package/bin/statusline.mjs +32 -0
- package/docs/development.md +84 -0
- package/docs/ollama-evaluation.md +100 -0
- package/docs/reference.md +227 -0
- package/docs/releasing.md +152 -0
- package/package.json +25 -0
- package/src/auth.mjs +51 -0
- package/src/config.mjs +80 -0
- package/src/model-request.mjs +13 -0
- package/src/ollama-evaluator.mjs +114 -0
- package/src/ollama-models.mjs +29 -0
- package/src/ollama-setup.mjs +184 -0
- package/src/onboarding.mjs +149 -0
- package/src/prompt-state.mjs +121 -0
- package/src/response-observer.mjs +176 -0
- package/src/router.mjs +325 -0
- package/src/savings.mjs +208 -0
- package/src/server.mjs +215 -0
- package/src/status-settings.mjs +88 -0
- package/src/status-state.mjs +174 -0
- package/src/statusline.mjs +185 -0
- package/src/token-counter.mjs +127 -0
- package/src/user-config.mjs +145 -0
|
@@ -0,0 +1,100 @@
|
|
|
1
|
+
# Local evaluator measurements
|
|
2
|
+
|
|
3
|
+
Jev remains the default evaluator. Ollama is an optional local classifier: Claude still generates the answer, and the router's capability and continuation guards still apply. This evaluation uses synthetic prompts only and does not measure the quality of Claude's completed work.
|
|
4
|
+
|
|
5
|
+
## Candidates and method
|
|
6
|
+
|
|
7
|
+
Measurements were recorded on September 29, 2026, on an Apple M4 Mac with 16 GiB of unified memory, running Ollama 0.33.3 alongside other applications. Candidates were loaded one at a time with a 4,096-token context. Download size is not the same as resident memory.
|
|
8
|
+
|
|
9
|
+
| Candidate | Published download size | Purpose |
|
|
10
|
+
| --- | ---: | --- |
|
|
11
|
+
| `qwen3.5:0.8b` | Approximately 1.0 GB | Smallest candidate |
|
|
12
|
+
| `qwen3:1.7b` | Approximately 1.4 GB | Compact speed baseline |
|
|
13
|
+
| `qwen3.5:2b` | Approximately 2.7 GB | Additional compact candidate |
|
|
14
|
+
| `qwen3.5:4b` | Approximately 3.4 GB | Measured larger candidate; rejected for preset latency |
|
|
15
|
+
| `qwen3:4b` | Approximately 2.5 GB | Larger candidate selected for the `quality` option |
|
|
16
|
+
|
|
17
|
+
Sizes and quantization vary by tag. Use explicit tags: untagged `qwen3.5` currently selects the much larger 9B model. See the official [Qwen3.5 model catalog](https://ollama.com/library/qwen3.5) and [Qwen3 1.7B listing](https://ollama.com/library/qwen3:1.7b).
|
|
18
|
+
|
|
19
|
+
The `compact` option selects `qwen3:1.7b` for its smaller allocation and faster measured classification. The `quality` option selects `qwen3:4b`, which agreed more often with the held-out labels while using more memory and time. The latter's official listing specifies a roughly 2.5 GB download and Q4_K_M quantization. Additional memory headroom is useful when other applications are open, but more RAM alone does not guarantee better accuracy or meeting the evaluator deadline. See the [official Qwen3 4B listing](https://ollama.com/library/qwen3:4b).
|
|
20
|
+
|
|
21
|
+
The checked-in fixture contains 36 independently authored, balanced routing cases: 12 tuning cases and 24 held-out cases, with equal numbers of Haiku, Sonnet, and Opus labels. Cases cover mechanical edits, ordinary implementation, difficult correctness and security work, topic changes, tool results, short follow-ups, and misleading routing instructions. Expected labels are judgments under the routing rubric, not independently verified claims about which Claude model would succeed.
|
|
22
|
+
|
|
23
|
+
The rubric was refined using the tuning set. The final held-out run uses the frozen rubric and is reported separately. Repeating each held-out case three times gives 72 decisions per model, but still only 24 distinct workloads.
|
|
24
|
+
|
|
25
|
+
The harness calls the production `buildOllamaState` and `evaluateOllama` functions. It includes local model metadata checks in wall-clock latency and bypasses AutoRouter's decision cache. Ollama's normal shared-prefix caching remains enabled. Input excerpts have a 3,000-character/UTF-8-byte ceiling. Requests disable thinking, use temperature 0, seed 0, a 32-token output cap, and a JSON schema containing only the tier. A returned tier is not a calibrated confidence probability. See Ollama's [thinking controls](https://docs.ollama.com/capabilities/thinking) and [structured-output guidance](https://docs.ollama.com/capabilities/structured-outputs).
|
|
26
|
+
|
|
27
|
+
A cold measurement starts with the model unloaded from Ollama; operating-system file caches and compiled kernels may already be warm. Cold calls have a separate 60-second measurement deadline. Warm measurements use the production 1,500 ms deadline. Model allocation is read from `/api/ps`, not inferred from the download size. On unified-memory hardware, its GPU allocation is not additional independent RAM. See [Ollama's running-model API](https://docs.ollama.com/api/ps) and [context-memory guidance](https://docs.ollama.com/context-length).
|
|
28
|
+
|
|
29
|
+
## Results
|
|
30
|
+
|
|
31
|
+
The frozen-rubric tuning results below document model selection. They are not held-out performance. GB values use decimal bytes; allocation is what the local running-model API reported, rather than total process or system memory.
|
|
32
|
+
|
|
33
|
+
| Model | Tuning agreement | Timeouts | Warm p50 / p95 | Cold wall time | Disk size | Model allocation |
|
|
34
|
+
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
|
|
35
|
+
| `qwen3:1.7b` | 20 / 24 (83.3%) | 0 / 24 | 245 / 280 ms | 2.50 s | 1.36 GB | 1.70 GB |
|
|
36
|
+
| `qwen3:4b` | 24 / 24 (100%) | 0 / 24 | 743 / 889 ms | 9.62 s | 2.50 GB | 3.18 GB |
|
|
37
|
+
| `qwen3.5:0.8b` | 10 / 24 (41.7%) | 0 / 24 | 409 / 437 ms | 2.78 s | 1.04 GB | 1.09 GB |
|
|
38
|
+
| `qwen3.5:2b` | 18 / 24 (75.0%) | 0 / 24 | 918 / 1,009 ms | 6.49 s | 2.74 GB | 2.36 GB |
|
|
39
|
+
| `qwen3.5:4b` | No valid warm results | 24 / 24 | Deadline reached | 7.26 s | 3.39 GB | 3.14 GB |
|
|
40
|
+
|
|
41
|
+
The smallest model was highly sensitive to rubric wording. It should not be selected solely because its download is small. The 2B candidate was slower and agreed less often than 1.7B on this tuning set. A separate diagnostic gave `qwen3.5:4b` a 5,000 ms deadline: all 12 tuning decisions agreed, but warm p50/p95 was 3,615/4,872 ms. **Qwen3.5 4B is rejected as a preset because it missed the production deadline on every warm tuning request.** Its diagnostic result demonstrates a latency tradeoff, not held-out quality. Extra RAM alone does not establish that this model will meet a 1,500 ms deadline on another machine.
|
|
42
|
+
|
|
43
|
+
An initial 4B run also exposed a metadata-response limit: its local `/api/show` response exceeded 64 KiB. That issue was fixed before the results above by bounding metadata separately at 1 MiB while retaining a 64 KiB classifier-output limit.
|
|
44
|
+
|
|
45
|
+
### Held-out results
|
|
46
|
+
|
|
47
|
+
Both selected candidates completed all 72 held-out decisions without a timeout. Qwen3 4B agreed more often than 1.7B, including every Opus-labeled workload. These measurements were taken later than tuning while other applications remained open; they do not represent an isolated hardware benchmark.
|
|
48
|
+
|
|
49
|
+
| Model | Agreement | Timeouts | Warm p50 / p95 | Cold wall time | Under-routes / over-routes |
|
|
50
|
+
| --- | ---: | ---: | ---: | ---: | ---: |
|
|
51
|
+
| `qwen3:1.7b` | 42 / 72 (58.3%) | 0 / 72 | 602 / 834 ms | 4.80 s | 18 / 12 |
|
|
52
|
+
| `qwen3:4b` | 66 / 72 (91.7%) | 0 / 72 | 889 / 1,242 ms | 5.92 s | 0 / 6 |
|
|
53
|
+
|
|
54
|
+
The 1.7B confusion matrix:
|
|
55
|
+
|
|
56
|
+
| Expected tier | Returned Haiku | Returned Sonnet | Returned Opus |
|
|
57
|
+
| --- | ---: | ---: | ---: |
|
|
58
|
+
| Haiku | 12 | 6 | 6 |
|
|
59
|
+
| Sonnet | 0 | 24 | 0 |
|
|
60
|
+
| Opus | 0 | 18 | 6 |
|
|
61
|
+
|
|
62
|
+
The 4B confusion matrix:
|
|
63
|
+
|
|
64
|
+
| Expected tier | Returned Haiku | Returned Sonnet | Returned Opus |
|
|
65
|
+
| --- | ---: | ---: | ---: |
|
|
66
|
+
| Haiku | 24 | 0 | 0 |
|
|
67
|
+
| Sonnet | 0 | 18 | 6 |
|
|
68
|
+
| Opus | 0 | 0 | 24 |
|
|
69
|
+
|
|
70
|
+
Each row represents eight unique workloads repeated three times. The wrong labels were consistent across repetitions. For 1.7B, six of eight Opus workloads were sent to Sonnet, and four of eight Haiku workloads were sent to a higher tier. For 4B, two of eight Sonnet workloads were sent to Opus. The compact model's gap between tuning and held-out agreement limits what can be claimed about it. Both local options remain experimental; these results do not establish parity with Jev, which was not evaluated on this fixture.
|
|
71
|
+
|
|
72
|
+
### Full-excerpt performance
|
|
73
|
+
|
|
74
|
+
An additional eight synthetic requests per model filled the entire 3,000-byte state budget. They kept a clearly mechanical current task and included long synthetic tool output. A different nonce at the beginning of each serialized state prevented reuse of the previous full user-state prefix while preserving the common rubric prefix. Both models exceeded the 1,500 ms deadline on all eight requests. These are separate performance checks, not additional held-out accuracy cases. Short-prompt latency must not be treated as a bound for full excerpts.
|
|
75
|
+
|
|
76
|
+
| Model | Timeouts at 1,500 ms | Wall p50 / p95 before cancellation |
|
|
77
|
+
| --- | ---: | ---: |
|
|
78
|
+
| `qwen3:1.7b` | 8 / 8 | 1,502 / 1,503 ms |
|
|
79
|
+
| `qwen3:4b` | 8 / 8 | 1,502 / 1,505 ms |
|
|
80
|
+
|
|
81
|
+
A diagnostic rerun of 1.7B with a 5,000 ms deadline completed all eight and returned Haiku, with p50/p95 of 3,753/3,991 ms. Qwen3 4B still timed out on all eight at that longer deadline, with p50/p95 cancellation times of 5,002/5,007 ms. Its actual completion times for these full excerpts were not measured. A larger model and startup priming do not remove the need for a bounded timeout and fallback during longer requests.
|
|
82
|
+
|
|
83
|
+
All final comparisons use fixture SHA-256 `1ef5111a6a36f0f4bc8d111d54c6985cac4c4a8b16357023a0958d285f7438ad` and rubric SHA-256 `42c3e18ddcf7c9d8756100b740f686b3c2f1bbdfbcbaf3cbc3021a9c9b4c6ee6`.
|
|
84
|
+
|
|
85
|
+
## Reproducing the evaluation
|
|
86
|
+
|
|
87
|
+
Use a source checkout; benchmark scripts and fixtures are development files and are not bundled in the npm package. Start local Ollama and explicitly download the models you intend to test. The harness never downloads or deletes models, and refuses to begin while another model is resident. It unloads each tested model after its measurements.
|
|
88
|
+
|
|
89
|
+
```sh
|
|
90
|
+
node scripts/evaluate-ollama.mjs --models qwen3:1.7b,qwen3:4b --split tuning --rounds 2 --output artifacts/ollama-tuning.json
|
|
91
|
+
node scripts/evaluate-ollama.mjs --models qwen3:1.7b,qwen3:4b --split heldout --rounds 3 --stress-rounds 8 --output artifacts/ollama-heldout.json
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
JSON reports contain fixture/rubric hashes, model digests, quantization, per-case labels, timing counters, failures, confusion matrices, and allocation measurements. They contain no private repository prompts or API credentials. Raw reports are written only when `--output` is supplied; `artifacts/` is ignored by Git.
|
|
95
|
+
|
|
96
|
+
## Limits
|
|
97
|
+
|
|
98
|
+
This is a small synthetic rubric-agreement benchmark, not a downstream task-quality, cost-savings, or security evaluation. Its prompts cannot represent every repository or long conversation. A model can agree with the labels and still miss important context outside the excerpt. The tests do not establish robust resistance to prompt injection.
|
|
99
|
+
|
|
100
|
+
Latency depends on hardware, current application load, model residency, and prompt length. Cold loading exceeds the normal evaluator deadline, which is why the launcher primes an installed local model with a synthetic classification using the production request settings before starting Claude. The warm benchmark measurements above already followed a full classification, so this startup improvement does not change those measurements. Timeouts and invalid responses use the router's existing conservative fallback policy. Re-run the evaluation before adopting different tags or changing the rubric; no result here guarantees accuracy on your workload.
|
|
@@ -0,0 +1,227 @@
|
|
|
1
|
+
# Reference
|
|
2
|
+
|
|
3
|
+
## Commands
|
|
4
|
+
|
|
5
|
+
| Command | Purpose |
|
|
6
|
+
| --- | --- |
|
|
7
|
+
| `claude-autorouter setup` | Save subscription-mode configuration and a Jev key |
|
|
8
|
+
| `claude-autorouter setup --auth-mode api-key` | Configure Jev and Anthropic API-key billing |
|
|
9
|
+
| `claude-autorouter setup --evaluator ollama --ollama-preset compact --pull` | Configure a local evaluator and download its selected model if missing |
|
|
10
|
+
| `claude-autorouter setup --force` | Replace an existing user config |
|
|
11
|
+
| `claude-autorouter doctor` | Check config, Claude executable/login, and the selected local Ollama model without paid calls |
|
|
12
|
+
| `claude-autorouter claude [arguments]` | Start a local router and pass arguments through to Claude Code |
|
|
13
|
+
| `claude-autorouter serve` | Run the router for separately configured clients |
|
|
14
|
+
| `claude-autorouter --help` | Show command help |
|
|
15
|
+
| `claude-autorouter --version` | Print the package version |
|
|
16
|
+
|
|
17
|
+
The launcher binds an ephemeral port on `127.0.0.1`, creates a temporary local credential, starts Claude with the gateway address, and shuts down when Claude exits. It works from any project directory on macOS or Linux, including WSL. Native Windows is unsupported in this release; the status-line command uses a POSIX shell. Standalone `serve` uses the configured port and requires a local token.
|
|
18
|
+
|
|
19
|
+
## Configuration
|
|
20
|
+
|
|
21
|
+
Setup defaults to subscription mode unless `--auth-mode` or `AUTOROUTER_AUTH_MODE` selects another mode. Jev remains the default evaluator; `--evaluator ollama` selects local classification. Setup prompts for required secrets without echoing them and writes a private JSON file. Stored keys are plaintext; keep the file private and out of source control. Supply keys through the environment when interactive input is unavailable. Subscription mode with Ollama requires no API keys. API-key authentication always requires `ANTHROPIC_API_KEY`, regardless of evaluator.
|
|
22
|
+
|
|
23
|
+
The config path is selected in this order:
|
|
24
|
+
|
|
25
|
+
1. `AUTOROUTER_CONFIG`, when set.
|
|
26
|
+
2. `$XDG_CONFIG_HOME/claude-autorouter/config.json`, when `XDG_CONFIG_HOME` is a nonempty absolute path.
|
|
27
|
+
3. `~/.config/claude-autorouter/config.json`.
|
|
28
|
+
|
|
29
|
+
The JSON file uses flat environment-style string keys, such as `AUTOROUTER_AUTH_MODE` and `TYPESAFE_API_KEY`. Environment values take precedence over the saved config. Use `setup --force` to replace existing configuration. The launcher does not discover or load a project's `.env` file. From a source checkout, explicitly loading one still works:
|
|
30
|
+
|
|
31
|
+
```sh
|
|
32
|
+
node --env-file=.env bin/autorouter.mjs claude
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
For an environment-only subscription launch, set `AUTOROUTER_AUTH_MODE=subscription` and either supply `TYPESAFE_API_KEY` or select `AUTOROUTER_EVALUATOR=ollama` with a running local model. For API-key mode, also supply `ANTHROPIC_API_KEY`. The shell variables are read by the router; Jev's key is removed from the Claude child environment.
|
|
36
|
+
|
|
37
|
+
| Variable | Default | Purpose |
|
|
38
|
+
| --- | --- | --- |
|
|
39
|
+
| `TYPESAFE_API_KEY` | required for Jev | Jev credential; unused by Ollama |
|
|
40
|
+
| `ANTHROPIC_API_KEY` | required in API-key mode | Upstream Anthropic credential |
|
|
41
|
+
| `AUTOROUTER_CONFIG` | see path order above | Explicit user config path |
|
|
42
|
+
| `AUTOROUTER_AUTH_MODE` | `api-key` without saved config; setup selects `subscription` | Authentication mode |
|
|
43
|
+
| `AUTOROUTER_EVALUATOR` | `jev` | `jev` or local `ollama` classification |
|
|
44
|
+
| `AUTOROUTER_CLIENT_PROFILE` | `compatible` | `native` retains Claude's own model and thinking settings |
|
|
45
|
+
| `AUTOROUTER_STATUSLINE` | enabled | `0` retains your existing status line |
|
|
46
|
+
| `AUTOROUTER_DEBUG` | off | `1` enables launcher metadata logs on stderr |
|
|
47
|
+
| `ENABLE_TOOL_SEARCH` | `true` in launcher when unset | Load MCP tool definitions on demand; explicit values are preserved |
|
|
48
|
+
| `AUTOROUTER_HAIKU_MODEL` | `claude-haiku-4-5-20251001` | Routine tier |
|
|
49
|
+
| `AUTOROUTER_SONNET_MODEL` | `claude-sonnet-5` | Standard tier |
|
|
50
|
+
| `AUTOROUTER_OPUS_MODEL` | `claude-opus-5-5` | Demanding tier and savings baseline |
|
|
51
|
+
| `AUTOROUTER_JEV_MODEL` | `jev-latest` | Classifier version |
|
|
52
|
+
| `AUTOROUTER_JEV_TIMEOUT_MS` | `1500` | Classifier deadline in milliseconds |
|
|
53
|
+
| `AUTOROUTER_OLLAMA_URL` | `http://127.0.0.1:11434` | Loopback Ollama base URL |
|
|
54
|
+
| `AUTOROUTER_OLLAMA_MODEL` | `qwen3:1.7b` | Installed local model tag; setup can choose a preset |
|
|
55
|
+
| `AUTOROUTER_OLLAMA_TIMEOUT_MS` | `1500` | Whole local classification deadline in milliseconds |
|
|
56
|
+
| `AUTOROUTER_OLLAMA_KEEP_ALIVE` | `5m` | How long Ollama retains the evaluator in memory |
|
|
57
|
+
| `AUTOROUTER_TOKEN_COUNT_TIMEOUT_MS` | `1500` | Context-check deadline; runs alongside classification |
|
|
58
|
+
| `AUTOROUTER_MIN_CONFIDENCE` | `0.75` | Jev confidence threshold; does not apply to Ollama |
|
|
59
|
+
| `AUTOROUTER_PORT` | `8787` | Standalone server port |
|
|
60
|
+
| `AUTOROUTER_TOKEN` | none | Standalone local credential, at least 16 characters |
|
|
61
|
+
| `AUTOROUTER_UPSTREAM_URL` | `https://api.anthropic.com` | Anthropic-compatible origin; fixed in subscription mode |
|
|
62
|
+
| `AUTOROUTER_JEV_URL` | `https://api.typesafe.ai/v1/systemone` | Jev endpoint |
|
|
63
|
+
|
|
64
|
+
Model access depends on your account. The policy recognizes specific Claude model versions; arbitrary gateway aliases do not automatically inherit their capabilities or context windows. Compare overrides with the [Anthropic model catalog](https://platform.claude.com/docs/en/models/overview).
|
|
65
|
+
|
|
66
|
+
## Ollama evaluator
|
|
67
|
+
|
|
68
|
+
Ollama is an experimental local classifier. It chooses a Claude tier; Haiku, Sonnet, or Opus still completes the task through Anthropic. Jev remains the default, and selecting Ollama never silently switches back to Jev.
|
|
69
|
+
|
|
70
|
+
[Install Ollama](https://docs.ollama.com/quickstart) and start its local service first. Open the Ollama app on macOS, or use `ollama serve` if a server is not already running. Then:
|
|
71
|
+
|
|
72
|
+
```sh
|
|
73
|
+
claude-autorouter setup --evaluator ollama --ollama-preset compact --pull
|
|
74
|
+
claude-autorouter doctor
|
|
75
|
+
claude-autorouter claude
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
Use `--force` to replace existing configuration. Setup detects the running local API. `--pull` authorizes downloading the chosen model when it is missing; without it, install the model yourself before setup. AutoRouter does not install Ollama, start its daemon, or download models during ordinary launches or `doctor` checks.
|
|
79
|
+
|
|
80
|
+
| Preset | Model | Selection |
|
|
81
|
+
| --- | --- | --- |
|
|
82
|
+
| `compact` | `qwen3:1.7b` | Default prioritizes lower memory use; lower held-out rubric agreement |
|
|
83
|
+
| `quality` | `qwen3:4b` | Better measured rubric agreement, with higher memory use and latency |
|
|
84
|
+
| `auto` | One of the above | `quality` only when reported total system memory exceeds 24 GiB; otherwise `compact` |
|
|
85
|
+
|
|
86
|
+
The memory heuristic estimates headroom using total system RAM, not currently free memory or a CPU/GPU benchmark. Both presets were tested on a 16 GiB Mac; `quality` can be selected explicitly on that capacity when memory permits. The automatic threshold is a conservative headroom choice, not a speed or accuracy guarantee. The `quality` model has an approximately 2.5 GB download; download size differs from resident memory, which includes runtime and context allocations. Concurrent applications also need memory. The classifier uses a 4,096-token context to bound that allocation. See [Ollama's context-memory guidance](https://docs.ollama.com/context-length). Downloaded models have their own licenses and are not included in this package.
|
|
87
|
+
|
|
88
|
+
The 16 GiB M4 comparison used 24 distinct held-out synthetic workloads, eight per tier, repeated three times for each model:
|
|
89
|
+
|
|
90
|
+
| Model | Rubric agreement | Warm p50 / p95 | Cold call | Model allocation |
|
|
91
|
+
| --- | ---: | ---: | ---: | ---: |
|
|
92
|
+
| `qwen3:1.7b` | 42 / 72 (58.3%) | 602 / 834 ms | 4.80 s | 1.70 GB |
|
|
93
|
+
| `qwen3:4b` | 66 / 72 (91.7%) | 889 / 1,242 ms | 5.92 s | 3.18 GB |
|
|
94
|
+
|
|
95
|
+
Both completed every short held-out request without a timeout. Compact routed six of eight distinct Opus-labeled workloads to Sonnet. Quality had no under-routing in this fixture; its six errors were two distinct Sonnet workloads routed to Opus on each repetition. Each model exceeded the 1,500 ms deadline on all eight requests in a separate full-excerpt stress test. The [local evaluation report](ollama-evaluation.md) records the test conditions and rejected candidates. There was no Jev comparison, so these results do not establish parity with Jev. Rubric agreement on synthetic cases does not establish the quality or cost of completed Claude tasks.
|
|
96
|
+
|
|
97
|
+
Use `--ollama-model LOCAL_TAG` to override the preset, for example with an already installed local model. The endpoint must be loopback (`127.0.0.1`, `localhost`, or `::1`), without a path, credentials, query, or fragment. Cloud model tags and metadata identifying a remote model are rejected before sending task text. Claude and Jev credentials are never attached to Ollama requests.
|
|
98
|
+
|
|
99
|
+
Local classification caps serialized evaluator state at both 3,000 characters and 3,000 UTF-8 bytes, including for non-ASCII prompts. It sends that state to `/api/chat`, requests a strict JSON tier, disables thinking, and caps output at 32 tokens. It does not invent a confidence probability; `AUTOROUTER_MIN_CONFIDENCE` applies only to Jev. All capability, tool-continuation, thinking, and context guards still apply.
|
|
100
|
+
|
|
101
|
+
Before opening Claude's UI, the launcher loads an installed model and primes the actual classifier rubric with a synthetic task, using a separate deadline of up to 60 seconds. Each normal evaluation has a 1,500 ms deadline covering local model metadata checks and classification. Priming reduces first-request overhead but does not guarantee that longer excerpts finish in time. `AUTOROUTER_OLLAMA_KEEP_ALIVE` defaults to `5m`. After five idle minutes, the next request may need to reload the model, exceed that deadline, and use the fallback. A longer positive keep-alive can reduce reloads while retaining memory longer; `0` unloads immediately and can make every evaluation cold. Supported values are `0` or a positive duration such as `30s`, `5m`, or `1h`, following the [Ollama API's keep-alive setting](https://docs.ollama.com/api/chat).
|
|
102
|
+
|
|
103
|
+
If startup priming fails, the launcher warns and continues. A missing model, unavailable service, malformed answer, or evaluation timeout falls back to Sonnet or retains an incoming Opus, subject to the usual compatibility policy. No Jev request is made. The status line identifies `Ollama fallback` and its error category. Run `doctor` to inspect local API availability and installed model metadata; it does not download or generate.
|
|
104
|
+
|
|
105
|
+
## Data flow and authentication
|
|
106
|
+
|
|
107
|
+
```text
|
|
108
|
+
Claude Code → authenticated local gateway → Jev or local Ollama classification
|
|
109
|
+
→ routing policy and optional token check
|
|
110
|
+
→ selected Claude model → streamed response
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
AutoRouter uses Claude Code's [gateway integration](https://code.claude.com/docs/en/llm-gateway-protocol), so it sees inference requests and tool continuations. It does not rely on a user-prompt hook.
|
|
114
|
+
|
|
115
|
+
The selected evaluator receives a bounded state containing the latest human request and excerpts of the original task, system text, and recent messages: up to 12,000 serialized characters sent to TypeSafe for Jev, or 3,000 UTF-8 bytes sent to the local Ollama service. These excerpts can include private source code and tool results. Images, document payloads, and signed thinking are omitted. Full tool schemas and full conversation history are not sent to either classifier. Anthropic receives the complete request, including its tools and attachments. Large or multimodal requests may also go to Anthropic's token-count endpoint before inference, including when classification is local.
|
|
116
|
+
|
|
117
|
+
In subscription mode, Claude Code owns login and OAuth refresh. AutoRouter forwards the current request's authorization and beta headers to Anthropic. It does not read keychain or saved login files, persist subscription tokens, or send them to Jev. A separate temporary `X-Autorouter-Token` authenticates the local connection and is stripped upstream. Subscription forwarding is restricted to `https://api.anthropic.com`. See [subscriptions and gateways](https://code.claude.com/docs/en/llm-gateway#subscriptions-and-gateways).
|
|
118
|
+
|
|
119
|
+
In API-key mode, the upstream key stays in the proxy and Claude receives a temporary local credential. Requests are billed to the supplied API key. Subscription requests remain subject to the subscription's model access and usage limits. AutoRouter never falls back from subscription authentication to API billing.
|
|
120
|
+
|
|
121
|
+
Routine logs contain route, model, timing, usage, and error-category metadata, not prompts, raw responses, or credentials. Status snapshots contain routing metadata and token counts in a private temporary directory and are deleted on normal launcher exit. Classification, turn, and token-count caches are held in memory. Claude Code and the external providers have their own storage and logging behavior.
|
|
122
|
+
|
|
123
|
+
## Routing policy
|
|
124
|
+
|
|
125
|
+
Each `/v1/messages` request is evaluated. Exact repeated bodies reuse a classification for five minutes. Both evaluators use a starting rubric choosing Haiku for routine work, Sonnet for ordinary engineering, and Opus for demanding reasoning. These choices require evaluation on your tasks; they are not quality guarantees.
|
|
126
|
+
|
|
127
|
+
The evaluator prioritizes the actual human request before startup metadata. Complete Claude reminder and tool-list blocks are excluded from that task excerpt, and long text retains its beginning and end. The outbound Anthropic request remains complete. Complexity outside the bounded excerpt can still be missed.
|
|
128
|
+
|
|
129
|
+
The following policy applies after classification:
|
|
130
|
+
|
|
131
|
+
- Jev's 1,500 ms deadline covers the response body and has no retry. Successful calls return immediately. Timeouts, HTTP errors, and invalid responses fall back to Sonnet or retain an existing stronger model.
|
|
132
|
+
- Jev confidence below 0.75 prevents a downgrade below Sonnet or the requested tier. Ollama returns a tier without calibrated confidence; its failure handling and compatibility guards still apply.
|
|
133
|
+
- Tool continuations retain the model chosen at the start of the human turn. Session, agent, and prompt headers identify turns; normalized conversation content provides a fallback. Moving prompt-cache markers does not create a new turn.
|
|
134
|
+
- Thinking history, fixed-budget thinking, server tools, context management, and other recognized model-specific features preserve the current model. Adaptive thinking, effort, and output above 64K prevent a Haiku choice. Fields are never stripped to force a downgrade.
|
|
135
|
+
- Mid-conversation `system` messages preserve the requested model and pass through unchanged. They do not count as a tool continuation by themselves.
|
|
136
|
+
- Recognized compaction and auxiliary requests otherwise preserve their requested model. Token counting and model discovery pass through without classification.
|
|
137
|
+
|
|
138
|
+
The default `compatible` profile starts Claude with Haiku-compatible requests and client-requested thinking disabled. When an upgrade to Opus 5/5.5 requires adaptive thinking, AutoRouter enables it. `AUTOROUTER_CLIENT_PROFILE=native` preserves normal client settings, which can constrain routing. An explicit Claude `--model` argument overrides the starting model, but `/model` and `--model` are requested models, not locks on the routed result.
|
|
139
|
+
|
|
140
|
+
The launcher enables `ENABLE_TOOL_SEARCH=true` when unset. Claude can otherwise disable on-demand MCP discovery when using a custom API address, loading connected-tool schemas into even a fresh conversation. Explicit values, including `false` or `auto:5`, are preserved. Managed settings and always-loaded tools can still affect deferral. See [Claude Code tool search](https://code.claude.com/docs/en/mcp#configure-tool-search).
|
|
141
|
+
|
|
142
|
+
### Context capacity
|
|
143
|
+
|
|
144
|
+
Context-relevant JSON above 150KB, or image/document blocks including those inside tool results, trigger a token check for otherwise compatible small models. This byte threshold is a trigger, not a token estimate. The check includes system instructions and active tool schemas. Unused deferred schemas are excluded from the trigger; discovered references and historical tool calls add them back. Unknown shapes are counted conservatively.
|
|
145
|
+
|
|
146
|
+
The check uses [Anthropic's token-count endpoint](https://platform.claude.com/docs/en/build-with-claude/token-counting), the current request's authentication, and the target model's tokenizer. It runs alongside the selected evaluator with a separate 1,500 ms deadline. Counts are cached for five minutes with at most 100 hash-and-number entries. Input at or below 190K can remain on a 200K model, allowing a margin for estimation differences. Unsupported input, timeouts, and failures fall back to the conservative capacity guard.
|
|
147
|
+
|
|
148
|
+
If the small model cannot fit, the default mapping upgrades it to Sonnet 5's native 1M window. A known capable Opus can be used when a configured Sonnet lacks that capacity. Compatible Haiku tool turns can upgrade as they grow; subsequent calls retain the upgrade. Thinking history and feature constraints remain pinned. Unknown capacities are not guessed. See [Anthropic context windows](https://platform.claude.com/docs/en/build-with-claude/context-windows).
|
|
149
|
+
|
|
150
|
+
Requests exceeding the selected model's capacity still receive the provider's error. AutoRouter does not truncate context or retry a generation on another model after an error or partial stream.
|
|
151
|
+
|
|
152
|
+
## Status line and savings
|
|
153
|
+
|
|
154
|
+
The `claude` launcher automatically adds a temporary [status-line command](https://code.claude.com/docs/en/statusline). Its fields show:
|
|
155
|
+
|
|
156
|
+
| Field | Meaning |
|
|
157
|
+
| --- | --- |
|
|
158
|
+
| `Sonnet 5 selected` | Routing chose this model; Anthropic has not confirmed it yet |
|
|
159
|
+
| `Opus 5.5` / `last Opus 5.5` | Provider-confirmed streaming or most recent model |
|
|
160
|
+
| `Jev` / `Ollama`, with `cache` or `fallback` when applicable | Classification source; timing includes concurrent context checks |
|
|
161
|
+
| `Jev→Haiku` or `Ollama→Haiku` beside Sonnet | A policy guard overrode the evaluator's Haiku choice |
|
|
162
|
+
| `large context` | Token count exceeded the small-model input budget |
|
|
163
|
+
| `size unverified` | Token checking failed or was unavailable; conservative guard applied |
|
|
164
|
+
| `API ctx` | Latest main request's input divided by the actual model's known window |
|
|
165
|
+
| `CLI ctx` | Claude's client accounting, shown when its window differs or API capacity is unknown |
|
|
166
|
+
| `est saved … vs Opus` | Cumulative API-equivalent token-cost estimate |
|
|
167
|
+
|
|
168
|
+
Background agents and auxiliary requests cannot replace the foreground model. Errors, fallback, cancellation, and stale/offline state remain visible. The command reads a local snapshot and makes no network requests. It respects terminal width and `NO_COLOR`; lower-priority fields disappear on narrow terminals.
|
|
169
|
+
|
|
170
|
+
Context includes uncached input, cache reads, and cache writes, excluding output to match [Claude's percentage formula](https://code.claude.com/docs/en/statusline#context-window-fields). It is current context rather than cumulative usage. Historical usage is marked `last`; compaction resets stale readings. The compatible client's 200K window can reach 100% while a routed Sonnet request uses only part of its 1M window. Displaying both does not change Claude's compaction threshold.
|
|
171
|
+
|
|
172
|
+
Savings compare the actual models' API token prices with the configured Opus model's prices for the **same reported counts and cache profile**. The percentage is `(Opus cost − routed cost) / Opus cost`. Input, output, cache reads, and 5-minute/1-hour cache writes are priced separately using the bundled table based on [Anthropic's published USD pricing](https://platform.claude.com/docs/en/about-claude/pricing).
|
|
173
|
+
|
|
174
|
+
Totals include completed main, agent, and auxiliary calls for the current session observed by this router process. Streaming usage is counted once, and totals reset when the router or session restarts. Unrecognized prices or unsupported usage produce `partial` or `savings unavailable`. Higher routed costs show `est extra`. The arithmetic runs locally.
|
|
175
|
+
|
|
176
|
+
This estimate does not measure subscription bill savings or quota credits. It excludes Jev charges, local compute costs, tool fees, negotiated discounts, and unpriced requests. A real Opus run can produce different tokens and cache hits. Incomplete streams, unknown cache-write TTLs, unsupported pricing modifiers, and unrecognized model versions are excluded rather than guessed. The rate table requires updates when prices change.
|
|
177
|
+
|
|
178
|
+
Set `AUTOROUTER_STATUSLINE=0` to retain an existing status line. Other `--settings` values are retained in the temporary overlay; source-relative Read/Edit rules keep their anchors. Ambiguous relative sandbox paths cause the launcher to skip the overlay and pass original settings through with a notice. Safe mode disables custom status lines; print mode has no status-line UI. Standalone `serve` does not install one.
|
|
179
|
+
|
|
180
|
+
## Troubleshooting
|
|
181
|
+
|
|
182
|
+
Run `claude-autorouter doctor` first. It performs local checks without Claude generations or Jev calls, including local HTTP checks for the selected Ollama model. It cannot establish Anthropic/TypeSafe availability, current quota, or whether a key will be accepted remotely.
|
|
183
|
+
|
|
184
|
+
For metadata logs without terminal noise:
|
|
185
|
+
|
|
186
|
+
```sh
|
|
187
|
+
AUTOROUTER_DEBUG=1 claude-autorouter claude 2>autorouter-debug.log
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
The launcher is quiet by default. Standalone `serve` logs to stderr by default. Claude still displays API errors normally.
|
|
191
|
+
|
|
192
|
+
**Authentication conflicts:** subscription launches remove inherited `ANTHROPIC_API_KEY`, `ANTHROPIC_AUTH_TOKEN`, and `CLAUDE_CODE_OAUTH_TOKEN` from the child environment, plus authentication-related custom headers. Claude settings can still supply conflicting `apiKeyHelper`, `env`, base-URL, or custom-header overrides. Remove those conflicts from the relevant settings, run `claude auth login` if needed, and relaunch. Parent-shell and saved Claude settings remain unchanged. `--bare` is incompatible with subscription OAuth.
|
|
193
|
+
|
|
194
|
+
**A simple prompt selects Sonnet:** inspect the reason and evaluator choice in the status line. Background context can be large even in a new session. Model-specific settings, tool continuity, Jev confidence, classifier failures, and context limits can override the classified tier. A fresh conversation picks up tool-loading changes; resumed history can retain already-loaded schemas.
|
|
195
|
+
|
|
196
|
+
**Claude says Haiku while AutoRouter says Opus:** the built-in model label is Claude's starting/requested model. The AutoRouter confirmed-model label comes from Anthropic. Claude's own token-cost estimate can likewise be attributed to the requested model.
|
|
197
|
+
|
|
198
|
+
**After restarting mid-conversation:** turn state is in memory and expires after 30 minutes. Unknown continuations preserve the incoming model. Start a fresh conversation when restarting around signed thinking; AutoRouter cannot reconstruct the prior actual model from lost turn state.
|
|
199
|
+
|
|
200
|
+
Switching models can lose prompt-cache reuse. A cheaper price per token does not guarantee a cheaper or faster task. Only requests using Claude's configured base URL are visible to this proxy. Alternate provider modes such as Bedrock, Vertex, Foundry, Mantle, and `ANTHROPIC_AWS` are unsupported; unset their enable flags before launching. Use ordinary `claude` to bypass routing.
|
|
201
|
+
|
|
202
|
+
## Standalone server
|
|
203
|
+
|
|
204
|
+
The launcher handles local credentials and environment settings automatically. For separate clients, configure `AUTOROUTER_TOKEN` with a random value of at least 16 characters and start:
|
|
205
|
+
|
|
206
|
+
```sh
|
|
207
|
+
claude-autorouter serve
|
|
208
|
+
```
|
|
209
|
+
|
|
210
|
+
For subscription mode, configure the client's environment in another terminal:
|
|
211
|
+
|
|
212
|
+
```sh
|
|
213
|
+
unset ANTHROPIC_API_KEY ANTHROPIC_AUTH_TOKEN CLAUDE_CODE_OAUTH_TOKEN
|
|
214
|
+
export ANTHROPIC_BASE_URL=http://127.0.0.1:8787
|
|
215
|
+
export ANTHROPIC_CUSTOM_HEADERS='X-Autorouter-Token: YOUR_LOCAL_ROUTER_TOKEN'
|
|
216
|
+
export CLAUDE_CODE_GATEWAY_HINT_HEADERS=1
|
|
217
|
+
export ANTHROPIC_MODEL=claude-haiku-4-5-20251001
|
|
218
|
+
export MAX_THINKING_TOKENS=0
|
|
219
|
+
export ENABLE_TOOL_SEARCH=true
|
|
220
|
+
claude
|
|
221
|
+
```
|
|
222
|
+
|
|
223
|
+
Replace the placeholder with the server's local token, never a subscription credential. Include other custom headers on separate lines if needed. Omit model/thinking variables to retain native client settings.
|
|
224
|
+
|
|
225
|
+
For API-key mode, set the base URL and gateway-hint variable above, and set both client `ANTHROPIC_API_KEY` and `ANTHROPIC_AUTH_TOKEN` to the local router token. Keep the real upstream API key in the server environment/config.
|
|
226
|
+
|
|
227
|
+
The server binds only to `127.0.0.1` and requires local authentication. `/health` checks the local process, not upstream provider availability.
|
|
@@ -0,0 +1,152 @@
|
|
|
1
|
+
# CI and npm releases
|
|
2
|
+
|
|
3
|
+
The package is `claude-autorouter`, licensed under [Apache-2.0](../LICENSE). **npm publication is pending.** The first release needs an interactive npm login; subsequent releases use GitHub Actions with npm trusted publishing. Preparing a tarball or merging a pull request does not publish it.
|
|
4
|
+
|
|
5
|
+
The GitHub repository is private. Publishing to npm makes the tarball's runtime source, README, configuration example, license, and shipped documentation public. Model weights, user configuration, credentials, transcripts, local artifacts, and test fixtures are excluded. Review the archive before the first publication and whenever the package allowlist changes.
|
|
6
|
+
|
|
7
|
+
## What runs automatically
|
|
8
|
+
|
|
9
|
+
| Workflow | Trigger | Behavior |
|
|
10
|
+
| --- | --- | --- |
|
|
11
|
+
| [ci.yml](https://github.com/frapposelli/claude-autorouter/blob/main/.github/workflows/ci.yml) | Pull requests, pushes to `main`, manual runs, and calls from the release workflow | Syntax checks, tests, and package smoke tests on Ubuntu/macOS with Node 22/24 |
|
|
12
|
+
| [publish.yml](https://github.com/frapposelli/claude-autorouter/blob/main/.github/workflows/publish.yml) | Push of a tag matching `v*` | Validate release, run CI, pack and test the candidate, then publish the verified archive |
|
|
13
|
+
|
|
14
|
+
A release tag must exactly equal `v` plus the version in `package.json`, and its commit must be reachable from `origin/main`. Package name and repository metadata must match `claude-autorouter` and `frapposelli/claude-autorouter`. Stable versions use npm's `latest` tag; prereleases such as `0.3.0-beta.1` use `next`.
|
|
15
|
+
|
|
16
|
+
The release workflow packs its candidate once and smoke-tests that exact `.tgz`. It uploads the archive and SHA-256 checksum as an Actions artifact. A separate publishing job downloads that artifact by its immutable ID, checks the checksum and every packaged file against the release checkout, then runs `npm publish` with scripts disabled. The publish job uses a GitHub-hosted Ubuntu runner, Node 24, and npm 11.19.1. Only that job has `id-token: write`; there is no `NPM_TOKEN` secret or required GitHub environment. Failed checks prevent publication.
|
|
17
|
+
|
|
18
|
+
The project has no package dependencies or lockfile, so CI runs its scripts directly without `npm ci`. Live Claude/Jev calls, Ollama downloads, and private repository probes are not CI checks.
|
|
19
|
+
|
|
20
|
+
## 1. Publish the first version interactively
|
|
21
|
+
|
|
22
|
+
Merge the release workflows and package metadata to `main`, push to GitHub, and ensure GitHub Actions is enabled for the repository. Its Actions policy must permit the pinned official GitHub actions and the reusable CI workflow in this repository. Use a clean checkout of that commit, with Node 24 and npm 11.19.1 to match the publisher. No release tag is needed for this bootstrap. First check the registry and account:
|
|
23
|
+
|
|
24
|
+
```sh
|
|
25
|
+
npm ping --registry https://registry.npmjs.org/
|
|
26
|
+
npm view claude-autorouter name version --registry https://registry.npmjs.org/
|
|
27
|
+
npm whoami --registry https://registry.npmjs.org/
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
The initial registry lookup returned HTTP 404; recheck immediately before release. If the name now exists, verify that your npm account owns it before continuing. A network or authentication failure is not evidence that a name is available. If another owner has claimed the name, choose an available name and update package metadata, release validation, documentation, and trust settings together.
|
|
31
|
+
|
|
32
|
+
If `whoami` reports `ENEEDAUTH`, sign in interactively and complete npm's browser/2FA prompts:
|
|
33
|
+
|
|
34
|
+
```sh
|
|
35
|
+
npm login --registry https://registry.npmjs.org/
|
|
36
|
+
npm whoami --registry https://registry.npmjs.org/
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
Use your own npm account with publishing permission, and keep credentials out of the repository and workflow secrets. See [npm login](https://docs.npmjs.com/cli/v11/commands/npm-login/) and [publishing a public package](https://docs.npmjs.com/creating-and-publishing-unscoped-public-packages/) for account requirements.
|
|
40
|
+
|
|
41
|
+
For the initial `0.2.0` release, run:
|
|
42
|
+
|
|
43
|
+
```sh
|
|
44
|
+
npm run check
|
|
45
|
+
npm test
|
|
46
|
+
npm run release:pack
|
|
47
|
+
npm run test:package -- --archive ./dist/claude-autorouter-0.2.0.tgz
|
|
48
|
+
tar -tzf ./dist/claude-autorouter-0.2.0.tgz
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
`release:pack` checks the allowed files and writes `dist/claude-autorouter-0.2.0.tgz` and its `.sha256` file. The `--archive` smoke test installs and exercises those bytes without repacking. Review the listed contents and retain both files. If anything changes, rebuild and repeat the exact-archive smoke test.
|
|
52
|
+
|
|
53
|
+
Publish that reviewed candidate, completing npm's interactive authentication challenge when requested:
|
|
54
|
+
|
|
55
|
+
```sh
|
|
56
|
+
npm publish ./dist/claude-autorouter-0.2.0.tgz --ignore-scripts --access public --tag latest --registry https://registry.npmjs.org/ --provenance=false
|
|
57
|
+
npm view claude-autorouter@0.2.0 version dist.integrity --registry https://registry.npmjs.org/
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
On a separate machine or disposable environment, verify the registry installation:
|
|
61
|
+
|
|
62
|
+
```sh
|
|
63
|
+
npm install -g claude-autorouter@0.2.0
|
|
64
|
+
claude-autorouter --version
|
|
65
|
+
claude-autorouter --help
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
Then run `setup`, `doctor`, and a launch from outside the source checkout as appropriate for that machine. `doctor` is local-only; a live prompt separately verifies provider access. Remove the README's pending-publication notice only after registry publication succeeds.
|
|
69
|
+
|
|
70
|
+
Do not push `v0.2.0` to test automation after this bootstrap: it would attempt to publish an existing version. npm name/version pairs cannot be reused, including after unpublishing. See the [npm publish reference](https://docs.npmjs.com/cli/v11/commands/npm-publish/).
|
|
71
|
+
|
|
72
|
+
## 2. Authorize this workflow on npm
|
|
73
|
+
|
|
74
|
+
Trusted publishing requires the package to exist first, which is why the initial version is published manually. The account configuring trust needs package write access and 2FA. See [npm trust prerequisites](https://docs.npmjs.com/cli/v11/commands/npm-trust/#prerequisites).
|
|
75
|
+
|
|
76
|
+
On npmjs.com, open the `claude-autorouter` package's **Settings → Trusted Publisher**, select **GitHub Actions**, and enter:
|
|
77
|
+
|
|
78
|
+
| Setting | Value |
|
|
79
|
+
| --- | --- |
|
|
80
|
+
| Organization or user | `frapposelli` |
|
|
81
|
+
| Repository | `claude-autorouter` |
|
|
82
|
+
| Workflow filename | `publish.yml` |
|
|
83
|
+
| Environment name | Leave blank |
|
|
84
|
+
| Allowed actions | Enable direct `npm publish` |
|
|
85
|
+
|
|
86
|
+
Use the filename only, not `.github/workflows/publish.yml`. No GitHub environment or npm token secret needs to be created. The owner, repository, and workflow must match exactly. New trust configurations default to permitting staged publication; **enable direct `npm publish`** for this workflow. See [npm trusted publishers](https://docs.npmjs.com/trusted-publishers/) and [staged publishing](https://docs.npmjs.com/staged-publishing/).
|
|
87
|
+
|
|
88
|
+
The workflow explicitly disables provenance because npm provenance is unsupported for private source repositories, even when the npm package is public. OIDC authentication still works. If the repository becomes public, review the workflow and metadata before enabling provenance. See [npm provenance requirements](https://docs.npmjs.com/generating-provenance-statements/).
|
|
89
|
+
|
|
90
|
+
After a successful trusted release, npm recommends the optional **Publishing access → Require two-factor authentication and disallow tokens** setting. It does not disable OIDC publishing. See [restricting token access](https://docs.npmjs.com/trusted-publishers/#recommended-restrict-token-access-when-using-trusted-publishers).
|
|
91
|
+
|
|
92
|
+
## 3. Release subsequent versions by tag
|
|
93
|
+
|
|
94
|
+
For the next patch after the bootstrap, prepare `0.2.1` on `main` or through a pull request:
|
|
95
|
+
|
|
96
|
+
```sh
|
|
97
|
+
git switch main
|
|
98
|
+
git pull --ff-only origin main
|
|
99
|
+
npm version 0.2.1 --no-git-tag-version
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
Review the version change and update any version-specific install examples or release notes. Check the candidate using the new filename:
|
|
103
|
+
|
|
104
|
+
```sh
|
|
105
|
+
npm run check
|
|
106
|
+
npm test
|
|
107
|
+
npm run release:pack
|
|
108
|
+
npm run test:package -- --archive ./dist/claude-autorouter-0.2.1.tgz
|
|
109
|
+
git diff --check
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
Commit the intended release changes and get that commit onto `main`, either through a pull request or a direct push allowed by the repository's branch rules. For a direct push with only the version changed:
|
|
113
|
+
|
|
114
|
+
```sh
|
|
115
|
+
git add package.json
|
|
116
|
+
git commit -m "Release 0.2.1"
|
|
117
|
+
git push origin main
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
Include any intentional documentation or release-note edits in that commit too. There is no publication from a branch push or PR merge. Wait for CI to pass, then tag that exact release commit:
|
|
121
|
+
|
|
122
|
+
```sh
|
|
123
|
+
git switch main
|
|
124
|
+
git pull --ff-only origin main
|
|
125
|
+
git tag -a v0.2.1 -m "Release 0.2.1"
|
|
126
|
+
git push origin v0.2.1
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
Before pushing, confirm `package.json` contains `0.2.1` and the tag points to the intended commit. For a prerelease, use a matching version/tag such as `0.3.0-beta.1` / `v0.3.0-beta.1`; it will publish under `next`, leaving `latest` unchanged.
|
|
130
|
+
|
|
131
|
+
Release stable versions in increasing version order, one tag at a time, and wait for each run to finish before pushing the next stable tag. The workflow queues releases without canceling an active run, but queue order does not sort semantic versions. Publishing an older stable version afterward could move `latest` backward; there is no registry version-order gate.
|
|
132
|
+
|
|
133
|
+
Open the tag's run under [GitHub Actions](https://github.com/frapposelli/claude-autorouter/actions). Under **Artifacts**, download `npm-package-<run-id>-<run-attempt>`, which contains the `.tgz` and checksum used for publication. Artifacts expire after 30 days, so retain them with the release record. After the publish job succeeds, verify the registry version and tags:
|
|
134
|
+
|
|
135
|
+
```sh
|
|
136
|
+
npm view claude-autorouter@0.2.1 version dist.integrity --registry https://registry.npmjs.org/
|
|
137
|
+
npm view claude-autorouter dist-tags --json --registry https://registry.npmjs.org/
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
Repeat the independent installation check for the released version. A GitHub Release page is optional; pushing the version tag is the publication trigger.
|
|
141
|
+
|
|
142
|
+
For local release diagnostics after the tag exists, `node scripts/release-check.mjs source v0.2.1` checks the tag, clean checkout, metadata, and ancestry. `node scripts/release-check.mjs archive v0.2.1` checks the candidate checksum and contents. These helpers are run automatically in the release workflow; the first untagged bootstrap uses the checks in step 1 instead.
|
|
143
|
+
|
|
144
|
+
## Recovering a failed release
|
|
145
|
+
|
|
146
|
+
- **Checks or archive validation failed:** nothing is published. Fix the cause and repeat validation before making a new release tag. Do not move a tag that already identifies a published version.
|
|
147
|
+
- **npm rejected OIDC authentication:** verify the npm trust fields, direct-publish permission, GitHub-hosted runner, and the publish job's OIDC permission. After correcting npm configuration, rerun the failed job if that version is still unpublished.
|
|
148
|
+
- **A publish timed out or the run was interrupted:** check `npm view` for the exact version before retrying. The registry may have accepted it before the connection failed.
|
|
149
|
+
- **The version already exists:** inspect the registry release; do not overwrite or unpublish to reuse it. Code or documentation corrections need a new version.
|
|
150
|
+
- **Local bootstrap authentication failed:** complete `npm login` and the account's 2FA flow in your terminal. CI trust cannot create the first package or substitute for that account step.
|
|
151
|
+
|
|
152
|
+
If release code or workflow changes are needed, commit the fix to `main` and prepare a new version/tag. Authentication-only corrections on npm can be retried against the unchanged, unpublished candidate.
|
package/package.json
ADDED
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
{
|
|
2
|
+
"name": "claude-autorouter",
|
|
3
|
+
"version": "0.2.0",
|
|
4
|
+
"license": "Apache-2.0",
|
|
5
|
+
"type": "module",
|
|
6
|
+
"description": "A local Claude Code model router with Jev and Ollama evaluators",
|
|
7
|
+
"repository": { "type": "git", "url": "git+https://github.com/frapposelli/claude-autorouter.git" },
|
|
8
|
+
"bin": { "claude-autorouter": "bin/autorouter.mjs" },
|
|
9
|
+
"engines": { "node": ">=22" },
|
|
10
|
+
"os": ["darwin", "linux"],
|
|
11
|
+
"keywords": ["claude", "claude-code", "model-routing", "jev", "ollama", "cli"],
|
|
12
|
+
"files": ["bin/*.mjs", "src/*.mjs", "docs/reference.md", "docs/development.md", "docs/releasing.md", "docs/ollama-evaluation.md", ".env.example", "LICENSE"],
|
|
13
|
+
"publishConfig": { "access": "public", "registry": "https://registry.npmjs.org/" },
|
|
14
|
+
"scripts": {
|
|
15
|
+
"start": "node bin/autorouter.mjs serve",
|
|
16
|
+
"claude": "node bin/autorouter.mjs claude",
|
|
17
|
+
"test": "node --test test/*.test.mjs",
|
|
18
|
+
"test:package": "node scripts/package-smoke.mjs",
|
|
19
|
+
"release:pack": "node scripts/release-pack.mjs",
|
|
20
|
+
"eval": "node --env-file=.env scripts/evaluate.mjs",
|
|
21
|
+
"eval:ollama": "node scripts/evaluate-ollama.mjs",
|
|
22
|
+
"test:live": "node --env-file=.env scripts/live-validation.mjs",
|
|
23
|
+
"check": "node scripts/check.mjs"
|
|
24
|
+
}
|
|
25
|
+
}
|
package/src/auth.mjs
ADDED
|
@@ -0,0 +1,51 @@
|
|
|
1
|
+
export const LOCAL_AUTH_HEADER = 'x-autorouter-token';
|
|
2
|
+
|
|
3
|
+
export function conflictingProviders(env = process.env) {
|
|
4
|
+
return ['CLAUDE_CODE_USE_BEDROCK', 'CLAUDE_CODE_USE_VERTEX', 'CLAUDE_CODE_USE_FOUNDRY',
|
|
5
|
+
'CLAUDE_CODE_USE_MANTLE', 'CLAUDE_CODE_USE_ANTHROPIC_AWS']
|
|
6
|
+
.filter(key => ['1', 'true'].includes(String(env[key]).toLowerCase()));
|
|
7
|
+
}
|
|
8
|
+
|
|
9
|
+
// This recognizes the subscription wire format, not the validity of the
|
|
10
|
+
// credential. Anthropic validates the bearer token on every forwarded request.
|
|
11
|
+
export function isSubscriptionRequest(headers) {
|
|
12
|
+
return !headers['x-api-key']
|
|
13
|
+
&& /^Bearer \S+$/i.test(headers.authorization ?? '')
|
|
14
|
+
&& String(headers['anthropic-beta'] ?? '').split(',').some(value => /^oauth-/.test(value.trim()));
|
|
15
|
+
}
|
|
16
|
+
|
|
17
|
+
export function buildClaudeEnv(config, baseUrl, parent = process.env) {
|
|
18
|
+
const env = { ...parent, ANTHROPIC_BASE_URL: baseUrl, CLAUDE_CODE_GATEWAY_HINT_HEADERS: '1' };
|
|
19
|
+
// Claude Code otherwise disables MCP tool search for a non-first-party
|
|
20
|
+
// base URL and loads every schema into context. This proxy preserves both
|
|
21
|
+
// tool_reference blocks and their beta headers. Respect explicit choices.
|
|
22
|
+
if (env.ENABLE_TOOL_SEARCH === undefined) env.ENABLE_TOOL_SEARCH = 'true';
|
|
23
|
+
// Start with a request format all three tiers accept. Jev still chooses the
|
|
24
|
+
// upstream model. Native mode preserves the client's full feature selection.
|
|
25
|
+
if (config.clientProfile === 'compatible') {
|
|
26
|
+
env.ANTHROPIC_MODEL = config.models.haiku;
|
|
27
|
+
env.MAX_THINKING_TOKENS = '0';
|
|
28
|
+
}
|
|
29
|
+
const subscription = config.authMode === 'subscription';
|
|
30
|
+
// Keep unrelated custom headers. Remove stale router credentials and, in
|
|
31
|
+
// subscription mode, inherited custom auth that could override the login.
|
|
32
|
+
const headers = String(env.ANTHROPIC_CUSTOM_HEADERS ?? '').split(/\r?\n/).filter(line => {
|
|
33
|
+
const name = line.split(':', 1)[0].trim().toLowerCase();
|
|
34
|
+
return line.trim() && name !== LOCAL_AUTH_HEADER
|
|
35
|
+
&& !(subscription && ['authorization', 'x-api-key'].includes(name));
|
|
36
|
+
});
|
|
37
|
+
if (subscription) {
|
|
38
|
+
delete env.ANTHROPIC_API_KEY;
|
|
39
|
+
delete env.ANTHROPIC_AUTH_TOKEN;
|
|
40
|
+
delete env.CLAUDE_CODE_OAUTH_TOKEN;
|
|
41
|
+
headers.push(`X-Autorouter-Token: ${config.localToken}`);
|
|
42
|
+
} else {
|
|
43
|
+
env.ANTHROPIC_API_KEY = config.localToken;
|
|
44
|
+
env.ANTHROPIC_AUTH_TOKEN = config.localToken;
|
|
45
|
+
}
|
|
46
|
+
if (headers.length) env.ANTHROPIC_CUSTOM_HEADERS = headers.join('\n');
|
|
47
|
+
else delete env.ANTHROPIC_CUSTOM_HEADERS;
|
|
48
|
+
delete env.TYPESAFE_API_KEY;
|
|
49
|
+
delete env.AUTOROUTER_TOKEN;
|
|
50
|
+
return env;
|
|
51
|
+
}
|