claude-autorouter 0.3.6 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.env.example +12 -7
- package/CONTRIBUTING.md +37 -0
- package/README.md +43 -70
- package/bin/autorouter.mjs +40 -56
- package/docs/development.md +50 -2
- package/docs/hardware-benchmark.md +29 -0
- package/docs/hardware-comparison.md +55 -0
- package/docs/hardware-results-16gb.json +4002 -0
- package/docs/hardware-results-16gb.md +26 -0
- package/docs/hardware-results-64gb.json +4020 -0
- package/docs/reference.md +83 -40
- package/docs/releasing.md +76 -34
- package/docs/router-performance.json +1697 -0
- package/docs/router-performance.md +50 -0
- package/docs/status-performance.json +363 -0
- package/docs/status-performance.md +44 -0
- package/package.json +57 -9
- package/src/auto-routing.mjs +214 -0
- package/src/bounded-json.mjs +57 -0
- package/src/cli-help.mjs +87 -0
- package/src/config-command.mjs +141 -0
- package/src/config.mjs +52 -27
- package/src/contracts.mjs +123 -0
- package/src/evaluation-report.mjs +114 -0
- package/src/local-diagnostic.mjs +191 -0
- package/src/model-catalog.mjs +96 -0
- package/src/model-request.mjs +10 -6
- package/src/ollama-evaluator.mjs +9 -27
- package/src/onboarding.mjs +82 -23
- package/src/request-validation.mjs +54 -0
- package/src/response-observer.mjs +126 -18
- package/src/router.mjs +174 -70
- package/src/savings.mjs +74 -16
- package/src/server.mjs +79 -12
- package/src/session-history.mjs +261 -0
- package/src/session-log.mjs +9 -58
- package/src/status-state.mjs +110 -62
- package/src/statusline.mjs +57 -27
- package/src/telemetry-event.mjs +196 -0
- package/src/token-counter.mjs +3 -1
- package/src/turn-state.mjs +132 -0
- package/src/user-config.mjs +18 -8
package/docs/reference.md
CHANGED
|
@@ -8,10 +8,17 @@
|
|
|
8
8
|
| `claude-autorouter setup --auth-mode api-key` | Configure Jev and Anthropic API-key billing |
|
|
9
9
|
| `claude-autorouter setup --client-profile auto` | Save a Sonnet/Opus profile compatible with Claude's Auto permission mode |
|
|
10
10
|
| `claude-autorouter setup --stop-hook-block-cap 2` | Opt into a shorter native Stop-hook continuation cap during setup |
|
|
11
|
-
| `claude-autorouter setup --session-log-dir DIR` | Save an opt-in directory for per-session JSONL
|
|
11
|
+
| `claude-autorouter setup --session-log-dir DIR` | Save an opt-in directory for per-session JSONL decisions and outcomes |
|
|
12
12
|
| `claude-autorouter setup --evaluator ollama --pull` | Configure the native local evaluator and download its selected model if missing |
|
|
13
13
|
| `claude-autorouter setup --evaluator ollama --ollama-timeout-ms 0 --force` | Save a disabled runtime evaluator deadline |
|
|
14
|
-
| `claude-autorouter setup --force` |
|
|
14
|
+
| `claude-autorouter setup --force` | Update an existing config while preserving unrelated saved settings |
|
|
15
|
+
| `claude-autorouter setup --replace` | Explicitly replace the saved configuration |
|
|
16
|
+
| `claude-autorouter config show --json` | Inspect effective settings and default/file/environment provenance; secrets are hidden |
|
|
17
|
+
| `claude-autorouter config set KEY VALUE` | Change one nonsecret saved setting |
|
|
18
|
+
| `claude-autorouter config unset KEY` | Remove one saved override |
|
|
19
|
+
| `claude-autorouter sessions list [--json]` | Inspect optional saved local history |
|
|
20
|
+
| `claude-autorouter sessions show ID [--json]` | Show correlated decisions, outcomes and pricing coverage |
|
|
21
|
+
| `claude-autorouter doctor --evaluate-local` | Run synthetic classifier checks on the installed local Ollama model |
|
|
15
22
|
| `claude-autorouter doctor` | Check config, Claude executable/login, and the selected local Ollama model without paid calls |
|
|
16
23
|
| `claude-autorouter claude [arguments]` | Start a local router and pass arguments through to Claude Code |
|
|
17
24
|
| `claude-autorouter serve` | Run the router for separately configured clients |
|
|
@@ -22,7 +29,7 @@ The launcher binds an ephemeral port on `127.0.0.1`, creates a temporary local c
|
|
|
22
29
|
|
|
23
30
|
## Configuration
|
|
24
31
|
|
|
25
|
-
|
|
32
|
+
First setup defaults to subscription mode unless `--auth-mode` or `AUTOROUTER_AUTH_MODE` selects another mode. Jev remains the default evaluator; `--evaluator ollama` selects local classification. An existing configuration updated with `--force` keeps its saved choices unless a command-line flag changes them; unrelated environment overrides remain temporary. Setup prompts for required secrets without echoing them and writes a private JSON file. Stored keys are plaintext; keep the file private and out of source control. Supply keys through the environment when interactive input is unavailable. Subscription mode with Ollama requires no API keys. API-key authentication always requires `ANTHROPIC_API_KEY`, regardless of evaluator.
|
|
26
33
|
|
|
27
34
|
The config path is selected in this order:
|
|
28
35
|
|
|
@@ -30,7 +37,7 @@ The config path is selected in this order:
|
|
|
30
37
|
2. `$XDG_CONFIG_HOME/claude-autorouter/config.json`, when `XDG_CONFIG_HOME` is a nonempty absolute path.
|
|
31
38
|
3. `~/.config/claude-autorouter/config.json`.
|
|
32
39
|
|
|
33
|
-
The JSON file uses flat environment-style string keys, such as `AUTOROUTER_AUTH_MODE` and `TYPESAFE_API_KEY`. Environment values take precedence over the saved config. Use `setup --force` to
|
|
40
|
+
The JSON file uses flat environment-style string keys, such as `AUTOROUTER_AUTH_MODE` and `TYPESAFE_API_KEY`. Environment values take precedence over the saved config. Use `config set` or `config unset` to edit a single saved setting, or `setup --force` to update an existing configuration while retaining unrelated settings. `setup --replace` explicitly rebuilds a readable supported saved configuration; malformed or unsupported files are left intact for manual repair. The launcher does not discover or load a project's `.env` file. From a source checkout, explicitly loading one still works:
|
|
34
41
|
|
|
35
42
|
```sh
|
|
36
43
|
node --env-file=.env bin/autorouter.mjs claude
|
|
@@ -48,11 +55,12 @@ For an environment-only subscription launch, set `AUTOROUTER_AUTH_MODE=subscript
|
|
|
48
55
|
| `AUTOROUTER_CLIENT_PROFILE` | `compatible` | `native` retains client model/thinking settings; `auto` starts with Sonnet when no explicit model is set and excludes Haiku from task routing |
|
|
49
56
|
| `AUTOROUTER_STATUSLINE` | enabled | `0` retains your existing status line |
|
|
50
57
|
| `AUTOROUTER_DEBUG` | off | `1` enables launcher metadata logs on stderr |
|
|
51
|
-
| `AUTOROUTER_SESSION_LOG_DIR` | off | Write per-session JSONL
|
|
58
|
+
| `AUTOROUTER_SESSION_LOG_DIR` | off | Write per-session JSONL decisions and outcomes into this directory; unset or empty disables it |
|
|
59
|
+
| `AUTOROUTER_SESSION_LOG_MODE` | `prompts` | `metadata` omits prompt excerpts; setting a mode alone does not enable logging |
|
|
52
60
|
| `CLAUDE_CODE_STOP_HOOK_BLOCK_CAP` | unset; Claude currently uses `8` | Optional cap on consecutive Stop/SubagentStop continuations without tool use; `0` disables the cap |
|
|
53
61
|
| `ENABLE_TOOL_SEARCH` | `true` in launcher when unset | Load MCP tool definitions on demand; explicit values are preserved |
|
|
54
62
|
| `AUTOROUTER_HAIKU_MODEL` | `claude-haiku-4-5-20251001` | Routine tier |
|
|
55
|
-
| `AUTOROUTER_SONNET_MODEL` | `claude-sonnet-5` | Standard tier |
|
|
63
|
+
| `AUTOROUTER_SONNET_MODEL` | `claude-sonnet-5`; `claude-sonnet-5-5` in the `auto` profile | Standard tier |
|
|
56
64
|
| `AUTOROUTER_OPUS_MODEL` | `claude-opus-5-5` | Demanding tier and savings baseline |
|
|
57
65
|
| `AUTOROUTER_JEV_MODEL` | `jev-latest` | Classifier version |
|
|
58
66
|
| `AUTOROUTER_JEV_TIMEOUT_MS` | `1500` | Classifier deadline in milliseconds |
|
|
@@ -69,6 +77,20 @@ For an environment-only subscription launch, set `AUTOROUTER_AUTH_MODE=subscript
|
|
|
69
77
|
|
|
70
78
|
Model access depends on your account. The policy recognizes specific Claude model versions; arbitrary gateway aliases do not automatically inherit their capabilities or context windows. Compare overrides with the [Anthropic model catalog](https://platform.claude.com/docs/en/models/overview).
|
|
71
79
|
|
|
80
|
+
### Inspect and change settings
|
|
81
|
+
|
|
82
|
+
```sh
|
|
83
|
+
claude-autorouter config show
|
|
84
|
+
claude-autorouter config show --json --check-all
|
|
85
|
+
claude-autorouter config set AUTOROUTER_OLLAMA_TIMEOUT_MS 0
|
|
86
|
+
claude-autorouter config unset AUTOROUTER_OLLAMA_TIMEOUT_MS
|
|
87
|
+
claude-autorouter config set TYPESAFE_API_KEY
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
`show` reports whether each setting comes from a default, the saved file, or an environment override. Secret values are never displayed. The last command uses a hidden prompt; scripts can pipe a secret to `config set TYPESAFE_API_KEY --stdin`. Secret values are not accepted as command arguments. Edits validate and atomically update only the named saved setting. Environment overrides still apply after a saved change. An unset saved deadline returns to the model default unless an environment value overrides it.
|
|
91
|
+
|
|
92
|
+
Normal startup validates the selected evaluator; stale settings for the inactive evaluator do not prevent it from starting. `show --check-all` explicitly checks both. Blank numeric settings fail with their setting name; zero retains its documented meaning. `claude-autorouter help COMMAND` gives focused command help. `claude-autorouter claude --help` and `--version` call Claude directly without router setup or credentials.
|
|
93
|
+
|
|
72
94
|
## Ollama evaluator
|
|
73
95
|
|
|
74
96
|
The local configuration documented here requires AutoRouter 0.3.2 or newer and remains experimental. It uses Ollama's native `/v1/systemone` decision endpoint for every model, replacing the chat backend from 0.2.0. Jev remains the default remote evaluator, using TypeSafe's `/v1/systemone` endpoint and a TypeSafe API key. Selecting Ollama never silently switches back to Jev. Haiku, Sonnet, or Opus still completes the task through Anthropic.
|
|
@@ -83,7 +105,7 @@ claude-autorouter doctor
|
|
|
83
105
|
claude-autorouter claude
|
|
84
106
|
```
|
|
85
107
|
|
|
86
|
-
`--force`
|
|
108
|
+
`--force` updates an existing user config and preserves its other settings and selected model unless explicitly changed. Setup detects the running local API. `--pull` authorizes downloading the chosen model when it is missing; without it, install the model yourself before setup. AutoRouter does not install Ollama, start its daemon, delete models, or download models during ordinary launches or `doctor` checks.
|
|
87
109
|
|
|
88
110
|
### Local model selection
|
|
89
111
|
|
|
@@ -115,7 +137,7 @@ The endpoint must be loopback (`127.0.0.1`, `localhost`, or `::1`), without a pa
|
|
|
115
137
|
|
|
116
138
|
Local classification caps serialized evaluator state at both 3,000 characters and 3,000 UTF-8 bytes, including for non-ASCII prompts. Claude's top-level executor system instructions are excluded before budgeting; the current task, original task, and recent conversation excerpts remain. `/v1/systemone` receives the bounded state and routing criteria and returns a tier directly. The router retains each model's native context setting: 8,194 tokens for the default Nimble tag and 2,050 for the listed Tev1 tags. Tev1's smaller window includes the routing criteria and template as well as the excerpt; the byte limit does not guarantee every possible input fits. Context errors use the normal fallback. Returned confidence scores summarize choice-distribution entropy; they are not calibrated accuracy probabilities. `AUTOROUTER_MIN_CONFIDENCE` applies only to Jev. All capability, tool-continuation, thinking, and context guards still apply.
|
|
117
139
|
|
|
118
|
-
The deadline covering local checks and classification defaults to 1,500 ms for Tev1 0.8B and custom/unrecognized tags, 15,000 ms for official Tev1 4B variants (including bare `tev1` and `latest`), and 30,000 ms for official Nimble variants. Official `library/` and `registry.ollama.ai/` aliases are recognized; a custom namespace such as `team/nimble` keeps the short default. An explicit timeout overrides the model default, including an old saved `1500`. Environment values override saved values on launch. Defaults are not written into the user config
|
|
140
|
+
The deadline covering local checks and classification defaults to 1,500 ms for Tev1 0.8B and custom/unrecognized tags, 15,000 ms for official Tev1 4B variants (including bare `tev1` and `latest`), and 30,000 ms for official Nimble variants. Official `library/` and `registry.ollama.ai/` aliases are recognized; a custom namespace such as `team/nimble` keeps the short default. An explicit timeout overrides the model default, including an old saved `1500`. Environment values override saved values on launch. Defaults are not written into the user config. To update a saved deadline, use `config set AUTOROUTER_OLLAMA_TIMEOUT_MS N` or `setup --ollama-timeout-ms N --force`; unrelated environment overrides remain temporary. Environment-provided Ollama settings are saved during first setup, replacement, or explicit `--evaluator ollama` selection. The command-line timeout flag takes precedence over the environment.
|
|
119
141
|
|
|
120
142
|
Set `AUTOROUTER_OLLAMA_TIMEOUT_MS=0` to remove AutoRouter's runtime evaluator timer while keeping the existing configuration:
|
|
121
143
|
|
|
@@ -137,11 +159,24 @@ If startup priming fails, the launcher warns and continues. An incompatible mode
|
|
|
137
159
|
|
|
138
160
|
Historical measurements before 0.3.2: Tev1 4B timed out on all eight full-excerpt checks even with a 10-second diagnostic allowance; its short-task results did not establish a full-excerpt latency bound. Tev1 0.8B completed all eight within 1,500 ms. On the tested 16 GiB M4, Nimble timed out on all 12 tuning requests at 1,500 ms. A separate 30-second diagnostic completed 24 held-out classifications with 23 matching labels, but median routing took 11.4 seconds. The one error followed a misleading tier instruction. These historical results precede the 0.3.2 excerpt changes and do not establish guarantees for the longer defaults. See the [measurements and limitations](ollama-evaluation.md).
|
|
139
161
|
|
|
162
|
+
### Test the installed local evaluator
|
|
163
|
+
|
|
164
|
+
```sh
|
|
165
|
+
claude-autorouter doctor --evaluate-local
|
|
166
|
+
claude-autorouter doctor --evaluate-local --json
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
This explicit diagnostic requires an installed local evaluator selected with `AUTOROUTER_EVALUATOR=ollama`. It uses synthetic tasks only, needs no Jev or Anthropic key, and makes no Claude inference calls. It checks local availability, measures a separate initial preparation call, then runs six uncached cases through the production classifier using the configured runtime deadline. A disabled runtime deadline remains disabled; Ctrl-C cancels the diagnostic.
|
|
170
|
+
|
|
171
|
+
The report separates availability, expected-label agreement, all-three-tier coverage, and latency. Residency is observed before calls, so it does not claim a controlled cold/warm benchmark. A pass establishes these six examples only. Incorrect predictions, fallback, or missing Haiku/Opus coverage fail even if Ollama answered successfully. Auto mode still checks all three raw evaluator labels; actual Auto routing applies its Sonnet floor separately.
|
|
172
|
+
|
|
173
|
+
The diagnostic does not download, unload, restart, or edit configuration. It refuses to run while unrelated models are resident and requires a positive keep-alive; normal routing still supports keep-alive `0`. Ordinary `doctor` remains a metadata check. A missing-model repair command uses the exact configured tag and preserves other settings.
|
|
174
|
+
|
|
140
175
|
### Migrating an older Ollama config
|
|
141
176
|
|
|
142
177
|
Version 0.3.1 used a 1,500 ms deadline for every local model. After upgrading to 0.3.2, an explicitly saved or exported `AUTOROUTER_OLLAMA_TIMEOUT_MS=1500` still wins over the new model-specific defaults. Remove that override to use the defaults, or rerun setup with the desired model and `--ollama-timeout-ms N --force`. The `0` value and setup timeout flag require 0.3.2 or newer.
|
|
143
178
|
|
|
144
|
-
Version 0.3.1 removed the Qwen chat backend and presets from 0.2.0. Existing downloaded models remain on disk, but an old Qwen model selection needs to be replaced with a native decision model. Run
|
|
179
|
+
Version 0.3.1 removed the Qwen chat backend and presets from 0.2.0. Existing downloaded models remain on disk, but an old Qwen model selection needs to be replaced with a native decision model. Run `claude-autorouter setup --evaluator ollama --ollama-model nimble:9b-q4_K_M --force`, or explicitly choose a Tev1 tag; merging with `--force` alone preserves the saved model. Remove or update any old `AUTOROUTER_OLLAMA_MODEL` environment value too, because environment variables override saved configuration. Update scripts to use `--ollama-model` when selecting a custom model.
|
|
145
180
|
|
|
146
181
|
## Data flow and authentication
|
|
147
182
|
|
|
@@ -163,7 +198,7 @@ Routine logs contain route, model, timing, usage, and error-category metadata, n
|
|
|
163
198
|
|
|
164
199
|
## Routing policy
|
|
165
200
|
|
|
166
|
-
Each `/v1/messages` request is evaluated. Exact repeated bodies reuse a classification for five minutes. Both evaluators use a starting rubric choosing Haiku for routine work, Sonnet for ordinary engineering, and Opus for demanding reasoning. These choices require evaluation on your tasks; they are not quality guarantees.
|
|
201
|
+
Each eligible `/v1/messages` request is evaluated. Exact repeated bodies reuse a classification for five minutes; concurrent identical evaluations share one request. Cache identity includes evaluator configuration, rubric and requested model floor. Internal permission classifiers and other documented pass-through paths skip evaluation. Both evaluators use a starting rubric choosing Haiku for routine work, Sonnet for ordinary engineering, and Opus for demanding reasoning. These choices require evaluation on your tasks; they are not quality guarantees.
|
|
167
202
|
|
|
168
203
|
The evaluator prioritizes the actual human request before startup metadata. Complete Claude reminder and tool-list blocks are excluded from that task excerpt, and long text retains its beginning and end. The outbound Anthropic request remains complete. Complexity outside the bounded excerpt can still be missed.
|
|
169
204
|
|
|
@@ -171,15 +206,15 @@ The following policy applies after classification:
|
|
|
171
206
|
|
|
172
207
|
- Jev's 1,500 ms deadline covers the response body and has no retry. Successful calls return immediately. Timeouts, HTTP errors, and invalid responses fall back to Sonnet or retain an existing stronger model.
|
|
173
208
|
- Jev confidence below 0.75 prevents a downgrade below Sonnet or the requested tier. Ollama returns a tier without calibrated confidence; its failure handling and compatibility guards still apply.
|
|
174
|
-
- Tool continuations retain the model
|
|
209
|
+
- Tool continuations retain the execution model confirmed by a successfully forwarded response. Active tasks and pending tools survive classification-cache expiry; retired task records expire separately. After restart, missing continuity is explicitly unknown. A selected model alone remains unconfirmed. Session, agent, and prompt headers identify turns; normalized conversation content provides a fallback. Text feedback from a Stop hook also retains the model when it serves the same gateway prompt ID and the client has not changed its requested model, subject to capability and context checks. Moving prompt-cache markers does not create a new turn.
|
|
175
210
|
- Claude's local `/goal` command can omit the prompt-ID header. For that path, an exact feedback label matching a preceding expanded `/goal` command keeps the original task and conversation anchor. This narrow text fallback also recognizes Claude's repeated-goal truncation format; arbitrary hook text is not treated as a goal. Feedback remains in the evaluator's recent conversation and the full API request. A new human message becomes the current task normally. The status line shows `prompt pinned` or `goal pinned` when either text-continuation rule applies.
|
|
176
|
-
- Thinking history, fixed-budget thinking, server tools, context management, and other recognized model-specific features preserve the current model. Adaptive thinking, effort, and output above 64K prevent a Haiku choice. Fields are never stripped to force a downgrade.
|
|
177
|
-
- Mid-conversation `system` messages preserve the requested model
|
|
178
|
-
- Auxiliary requests, including Claude's Auto permission classifier, pass through on their requested model without Jev/Ollama evaluation, token checks, or turn-state changes. Compaction retains its existing model and context-capacity policy.
|
|
211
|
+
- Thinking history, fixed-budget thinking, server tools, context management, and other recognized model-specific features preserve the current model except for the verified shared capabilities of the modern Auto-mode Sonnet/Opus pair described below. Adaptive thinking, effort, and output above 64K prevent a Haiku choice. Fields are never stripped to force a downgrade.
|
|
212
|
+
- Mid-conversation `system` messages preserve the requested model unless both Auto-mode models support them; they always pass through unchanged. They do not count as a tool continuation by themselves.
|
|
213
|
+
- Auxiliary requests, including Claude's Auto permission classifier, pass through on their requested model without Jev/Ollama evaluation, token checks, or turn-state changes. Compaction retains its existing model and context-capacity policy. Recognized server-reviewed execution requests can route between compatible Sonnet/Opus models while retaining `safeguards` and all verdicts unchanged. Unknown safeguards contracts pass through. Token counting and model discovery pass through without classification.
|
|
179
214
|
|
|
180
215
|
The default `compatible` profile starts Claude with Haiku-compatible requests and client-requested thinking disabled. AutoRouter uses adaptive thinking when upgrading these requests to Opus 5/5.5. Starting with 0.3.3, routing to exact `claude-sonnet-5-5` translates disabled thinking to `between_tools`, which skips up-front thinking but permits progress updates between tool calls. At `xhigh`/`max` effort, or when per-message effort differs from the top-level setting (default `high`), it uses adaptive thinking while preserving the effort settings. Token counting uses the same adaptation. Sonnet 5 still accepts disabled thinking and is unchanged. See [Sonnet 5.5 thinking requirements](https://platform.claude.com/docs/en/models/sonnet-5-5/migration-guide).
|
|
181
216
|
|
|
182
|
-
Explicit native `between_tools` and unknown thinking modes retain the incoming model on new human turns. Signed thinking blocks pass through unchanged and existing tool turns retain their model pin. `AUTOROUTER_CLIENT_PROFILE=native` preserves normal client settings, which can constrain routing. An explicit Claude `--model` argument overrides the starting model, but `/model` and `--model` are requested models, not locks on the routed result. Native same-model requests and unknown model aliases are not rewritten; clients must use settings supported by that model.
|
|
217
|
+
Explicit native `between_tools` and unknown thinking modes retain the incoming model on new human turns outside Auto routing. In Auto routing, a known Sonnet 5.5 `between_tools` request can upgrade to Opus with adaptive thinking. Signed thinking blocks pass through unchanged and existing tool turns retain their model pin. `AUTOROUTER_CLIENT_PROFILE=native` preserves normal client settings, which can constrain routing. An explicit Claude `--model` argument overrides the starting model, but `/model` and `--model` are requested models, not locks on the routed result. Native same-model requests and unknown model aliases are not rewritten; clients must use settings supported by that model.
|
|
183
218
|
|
|
184
219
|
The launcher enables `ENABLE_TOOL_SEARCH=true` when unset. Claude can otherwise disable on-demand MCP discovery when using a custom API address, loading connected-tool schemas into even a fresh conversation. Explicit values, including `false` or `auto:5`, are preserved. Managed settings and always-loaded tools can still affect deferral. See [Claude Code tool search](https://code.claude.com/docs/en/mcp#configure-tool-search).
|
|
185
220
|
|
|
@@ -187,7 +222,7 @@ The launcher enables `ENABLE_TOOL_SEARCH=true` when unset. Claude can otherwise
|
|
|
187
222
|
|
|
188
223
|
The default `compatible` profile starts Claude as Haiku to permit three-tier routing. Claude's Auto permission mode does not support Haiku, even if AutoRouter routes an API request to Sonnet. Eligibility is based on Claude's selected client model. Gateways themselves are supported. See [Claude's Auto-mode requirements](https://code.claude.com/docs/en/permission-modes#eliminate-permission-prompts-with-auto-mode).
|
|
189
224
|
|
|
190
|
-
AutoRouter 0.3.6
|
|
225
|
+
AutoRouter 0.3.6 introduced an `auto` client profile but bypassed evaluation for server-reviewed execution. Version 0.3.7 adds automatic switching on those requests. Launch with:
|
|
191
226
|
|
|
192
227
|
```sh
|
|
193
228
|
claude-autorouter claude --permission-mode auto
|
|
@@ -199,11 +234,15 @@ An explicit `--permission-mode auto` (or `--permission-mode=auto`) selects the p
|
|
|
199
234
|
env AUTOROUTER_CLIENT_PROFILE=auto claude-autorouter claude
|
|
200
235
|
```
|
|
201
236
|
|
|
202
|
-
The profile defaults
|
|
237
|
+
The profile defaults to Sonnet 5.5 and Opus 5.5, preserving explicit configured model IDs. The evaluator chooses Sonnet or Opus for each new human task; a routine Haiku verdict uses Sonnet and shows `Auto mode floor`. Tool and `/goal` continuations stay on the selected execution model. Claude's initial client model remains separate from the routed model. An explicit client `--model` or `ANTHROPIC_MODEL` can still make Auto unavailable if it selects Haiku or another unsupported model; choose a supported Sonnet or Opus instead.
|
|
238
|
+
|
|
239
|
+
Claude remains responsible for enabling the permission mode and enforcing organization settings, account availability, and tool rules. The profile does not enable Auto by itself or override `disableAutoMode`. AutoRouter does not reproduce Claude's settings precedence to infer a mode from settings files. `setup --client-profile auto --force` updates an existing configuration while retaining its other settings. `config set AUTOROUTER_CLIENT_PROFILE auto` changes just that setting.
|
|
240
|
+
|
|
241
|
+
**Safety review:** Claude's permission-classifier requests retain their exact requested model and skip AutoRouter's evaluator. Ordinary execution requests with the known `dangerous_tool_use` version-1 review contract are evaluated and routed, retaining the complete `safeguards` object, beta headers, and streamed safety verdicts. This also detects server review when Auto was selected in Claude's UI rather than through the launch flag. Unknown or malformed review contracts and safeguarded compaction pass through with `Auto safety`; a target that cannot accept the request shows `Auto model guard`. AutoRouter never turns off server review or converts denied actions to approvals. See [server-side classifier review](https://code.claude.com/docs/en/permission-modes#server-side-classifier-review).
|
|
203
242
|
|
|
204
|
-
|
|
243
|
+
**Shared execution capabilities:** automatic Auto routing supports exact Sonnet 5/5.5 and Opus 5/5.5 IDs. The default 5.5 pair shares a native 1M context window, adaptive thinking, native context-editing strategies, and mid-conversation system updates. These fields and existing signed thinking no longer pin every future human task. Sonnet 5 cannot accept mid-conversation system messages, per-message effort, or task budgets; use Sonnet 5.5 for those sessions. Unknown context-editing strategies, specialized server tools, fixed thinking budgets, fast mode, older/custom targets, and other incompatible requests still retain a compatible model. A Sonnet 5.5 `between_tools` request uses adaptive thinking when upgraded to Opus; effort and conversation history remain unchanged.
|
|
205
244
|
|
|
206
|
-
|
|
245
|
+
Thinking blocks stay verbatim in the conversation. Anthropic may drop blocks the selected model cannot read, so switching models does not preserve access to every model's private reasoning on every turn. User text, tool results, and prior answers remain available. Returning to a model can make its preserved thinking readable again. Model changes can also miss the prior model's prompt cache. See [preserved thinking and model switching](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking).
|
|
207
246
|
|
|
208
247
|
### Context capacity
|
|
209
248
|
|
|
@@ -223,7 +262,7 @@ The `claude` launcher automatically adds a temporary [status-line command](https
|
|
|
223
262
|
| --- | --- |
|
|
224
263
|
| `Sonnet 5 selected` | Routing chose this model; Anthropic has not confirmed it yet |
|
|
225
264
|
| `Opus 5.5` / `last Opus 5.5` | Provider-confirmed streaming or most recent model |
|
|
226
|
-
| `Jev` / `Ollama`, with `cache` or `fallback` when applicable | Classification source;
|
|
265
|
+
| `Jev` / `Ollama`, with `cache` or `fallback` when applicable | Classification source; displayed routing time includes evaluator waiting and context checks |
|
|
227
266
|
| `Jev→Haiku` or `Ollama→Haiku` beside Sonnet | A policy guard overrode the evaluator's Haiku choice |
|
|
228
267
|
| `large context` | Token count exceeded the small-model input budget |
|
|
229
268
|
| `size unverified` | Token checking failed or was unavailable; conservative guard applied |
|
|
@@ -231,51 +270,55 @@ The `claude` launcher automatically adds a temporary [status-line command](https
|
|
|
231
270
|
| `CLI ctx` | Claude's client accounting, shown when its window differs or API capacity is unknown |
|
|
232
271
|
| `est saved … vs Opus` | Cumulative API-equivalent token-cost estimate |
|
|
233
272
|
|
|
234
|
-
Background agents and auxiliary requests cannot replace the foreground model. Errors, fallback, cancellation, and stale/offline state remain visible. The command reads a local snapshot and makes no network requests. It respects terminal width and `NO_COLOR`; lower-priority fields disappear on narrow terminals.
|
|
273
|
+
Background agents and auxiliary requests cannot replace the foreground model. Errors, fallback, cancellation, incomplete-response evidence and stale/offline state remain visible. The command reads a local snapshot and makes no network requests. It respects terminal width and `NO_COLOR`; lower-priority fields disappear on narrow terminals.
|
|
235
274
|
|
|
236
275
|
Context includes uncached input, cache reads, and cache writes, excluding output to match [Claude's percentage formula](https://code.claude.com/docs/en/statusline#context-window-fields). It is current context rather than cumulative usage. Historical usage is marked `last`; compaction resets stale readings. The compatible client's 200K window can reach 100% while a routed Sonnet request uses only part of its 1M window. Displaying both does not change Claude's compaction threshold.
|
|
237
276
|
|
|
238
277
|
Savings compare the actual models' API token prices with the configured Opus model's prices for the **same reported counts and cache profile**. The percentage is `(Opus cost − routed cost) / Opus cost`. Input, output, cache reads, and 5-minute/1-hour cache writes are priced separately using the bundled table based on [Anthropic's published USD pricing](https://platform.claude.com/docs/en/about-claude/pricing).
|
|
239
278
|
|
|
240
|
-
Totals include completed main, agent, and auxiliary calls for the current session observed by this router process. Streaming usage is counted once, and totals reset when the router or session restarts. Unrecognized prices or unsupported usage produce `partial` or `savings unavailable`. Higher routed costs show `est extra`. The arithmetic runs locally.
|
|
279
|
+
Totals include completed main, agent, and auxiliary calls for the current session observed by this router process. Streaming usage is counted once, and totals reset when the router or session restarts. Unrecognized prices or unsupported usage produce `partial`, `unpriced N` or `savings unavailable`. Saved history identifies pricing table `2026-09-29.1` (reviewed September 29, 2026) and unpriced reason counts; unknown historical table versions are not repriced. Higher routed costs show `est extra`. The arithmetic runs locally.
|
|
241
280
|
|
|
242
281
|
This estimate does not measure subscription bill savings or quota credits. It excludes Jev charges, local compute costs, tool fees, negotiated discounts, and unpriced requests. A real Opus run can produce different tokens and cache hits. Incomplete streams, unknown cache-write TTLs, unsupported pricing modifiers, and unrecognized model versions are excluded rather than guessed. The rate table requires updates when prices change.
|
|
243
282
|
|
|
283
|
+
Historical integration observations cover Claude Code 2.1.284 and 2.1.285. The source-only versioned corpus in `test/fixtures/claude-protocol-v1.json` separates newly authored synthetic contracts from those dated observations and their artifact hashes. Gateway tests exercise Auto safeguards, thinking, deferred tools, compaction, goal scoping, fallback ownership and usage while preserving response bytes. They do not establish current-source live compatibility, evaluator accuracy or downstream task quality. Doctor reports the installed executable version separately; discovering a binary is not evidence that its protocol or Auto eligibility has been tested. Opt-in live validation supplements the synthetic fixtures.
|
|
284
|
+
|
|
244
285
|
Set `AUTOROUTER_STATUSLINE=0` to retain an existing status line. Other `--settings` values are retained in the temporary overlay; source-relative Read/Edit rules keep their anchors. Ambiguous relative sandbox paths cause the launcher to skip the overlay and pass original settings through with a notice. Safe mode disables custom status lines; print mode has no status-line UI. Standalone `serve` does not install one.
|
|
245
286
|
|
|
246
287
|
## Session decision logs
|
|
247
288
|
|
|
248
|
-
|
|
289
|
+
Logging is optional and disabled by default. Enable it for one launch:
|
|
249
290
|
|
|
250
291
|
```sh
|
|
251
292
|
env AUTOROUTER_SESSION_LOG_DIR="$HOME/.local/state/claude-autorouter/sessions" \
|
|
252
293
|
claude-autorouter claude
|
|
253
294
|
```
|
|
254
295
|
|
|
255
|
-
|
|
296
|
+
Or save metadata-only history, without prompt excerpts:
|
|
256
297
|
|
|
257
|
-
|
|
298
|
+
```sh
|
|
299
|
+
claude-autorouter config set AUTOROUTER_SESSION_LOG_MODE metadata
|
|
300
|
+
claude-autorouter config set AUTOROUTER_SESSION_LOG_DIR "$HOME/.local/state/claude-autorouter/sessions"
|
|
301
|
+
claude-autorouter sessions list
|
|
302
|
+
claude-autorouter sessions show autorouter-session-EXAMPLE
|
|
303
|
+
claude-autorouter sessions show autorouter-session-EXAMPLE --json
|
|
304
|
+
```
|
|
258
305
|
|
|
259
|
-
|
|
306
|
+
Use the exact `id` printed by `sessions list`. Commands need no evaluator credentials and do not contact providers. Human summaries distinguish selected models, observed serving models, confirmed completions, failures, cancellations, and pending/unconfirmed requests. They report routing latency, fallback and override counts, and API-equivalent savings coverage. A selected model or an HTTP 200 alone does not prove successful inference. Old schema-1 decision logs remain readable and explicitly lack outcome evidence.
|
|
260
307
|
|
|
261
|
-
|
|
262
|
-
{"schema_version":1,"event":"decision","timestamp":"2026-09-30T12:00:00.000Z","session_id":"example-session","prompt_excerpt":"Fix the typo in README.md","prompt_truncated":false,"requested_model":"claude-haiku-4-5-20251001","selected_model":"claude-sonnet-5","decision_latency_ms":214.37,"source":"jev","reason":"classified"}
|
|
263
|
-
```
|
|
308
|
+
`prompts` mode preserves the existing excerpt behavior when a log directory is enabled. Main requests retain at most 500 Unicode characters of the human task; recognized tool and goal continuations retain the originating task. Auxiliary, subagent, compaction, workflow, and attachment-only requests have empty excerpts. Metadata mode omits the excerpt fields entirely. Neither mode logs authentication headers, provider replies, full transcripts, or tool payloads. Text entered directly in a prompt can appear in an enabled prompt excerpt.
|
|
264
309
|
|
|
265
|
-
|
|
266
|
-
- `selected_model`: AutoRouter's final selected model after compatibility checks, before the upstream response. It does not confirm which model successfully answered.
|
|
267
|
-
- `decision_latency_ms`: time spent making the routing decision, including evaluator waiting, cache lookup, and any context checks. It excludes Claude generation time and log writing. A `passthrough` entry can be near zero because no evaluator was called.
|
|
268
|
-
- `source` and `reason`: distinguish evaluator choices, cache hits, fallbacks, turn/model constraints, and native safety pass-through. `classified_tier` is included when an evaluator returned a tier, which may differ from the final selected model.
|
|
310
|
+
The settings also work with `serve` and every client profile. `setup --session-log-dir DIR --session-log-mode metadata --force` updates an existing configuration. Setup resolves relative directories at setup time; environment-only paths resolve from the launch directory. `AUTOROUTER_SESSION_LOG_DIR=''` disables a saved directory for one launch. Setting only the mode never enables logging. `doctor` reports preferences without creating files.
|
|
269
311
|
|
|
270
|
-
|
|
312
|
+
Files are named `autorouter-session-*.jsonl`, one per observed Claude session in each router launch. Resuming a session in a new launch creates a new file. Requests without a session header share an anonymous file; subagents retain their agent IDs. Every line is a bounded schema-2 JSON record:
|
|
271
313
|
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
|
|
314
|
+
- `decision`: requested and selected models, evaluator verdict/source, policy reason, compatibility and continuity detail. `evaluation_latency_ms` measures evaluation/cache waiting; `routing_latency_ms` (also `decision_latency_ms`) includes compatibility and context checks.
|
|
315
|
+
- `outcome`: the same `request_id`, observed `confirmed_model` and bounded model transitions, safe error category, `completed`, `error` or `cancelled` status, and separate `completion_confirmed` evidence. It includes usage when available, `first_response_ms` from upstream forwarding to response headers, and total request latency. Early errors can have an outcome without a decision.
|
|
316
|
+
|
|
317
|
+
Outcomes record the configured Opus baseline and pricing-table version. History prices only successfully completed, supported usage with a recognized recorded table version and baseline. Unknown prices, missing or partial streams, ambiguous/mixed-model usage, and legacy records remain unpriced with a reason. Estimates never represent subscription charges. Status savings identify partial coverage; history exposes the version and counts.
|
|
275
318
|
|
|
276
|
-
|
|
319
|
+
New directories use `0700`, files use `0600`; existing directory permissions are unchanged. Asynchronous writes use a bounded 1 MiB queue and at most 128 session files per process. Normal shutdown drains accepted records. Storage failure or queue limits disable further logging with one generic warning while routing continues. Abrupt termination can lose unwritten records.
|
|
277
320
|
|
|
278
|
-
|
|
321
|
+
History reads at most 100 files, 4 MiB per file, 16 MiB total and 5,000 records per file. Limits, malformed records and partial tails are reported as partial coverage. Files are never automatically rotated or deleted; manage retention in your chosen directory. Log filenames are ignored by this repository and excluded from the npm package.
|
|
279
322
|
|
|
280
323
|
## Troubleshooting
|
|
281
324
|
|
|
@@ -315,7 +358,7 @@ The environment-only command also works on AutoRouter 0.3.4. Saved configuration
|
|
|
315
358
|
"CLAUDE_CODE_STOP_HOOK_BLOCK_CAP": "2"
|
|
316
359
|
```
|
|
317
360
|
|
|
318
|
-
For a new configuration, use `claude-autorouter setup --stop-hook-block-cap 2`; the flag works with either evaluator and overrides the environment during setup. Runtime environment values override saved configuration.
|
|
361
|
+
For a new configuration, use `claude-autorouter setup --stop-hook-block-cap 2`; the flag works with either evaluator and overrides the environment during setup. Runtime environment values override saved configuration. To change only a saved cap, use `config set CLAUDE_CODE_STOP_HOOK_BLOCK_CAP 2` or `setup --stop-hook-block-cap 2 --force`; both preserve unrelated saved settings. `doctor` reports the cap when configured. AutoRouter accepts nonnegative safe integers and leaves the setting absent unless you opt in.
|
|
319
362
|
|
|
320
363
|
### Other session issues
|
|
321
364
|
|
package/docs/releasing.md
CHANGED
|
@@ -14,6 +14,10 @@ Version `0.3.5` adds opt-in saved configuration for Claude's native `CLAUDE_CODE
|
|
|
14
14
|
|
|
15
15
|
Version `0.3.6` fixes Auto permission-mode launches with an Auto-compatible Sonnet/Opus profile while preserving Claude's permission classifiers and server safety-review requests. It also adds optional per-session JSONL decision logs containing a bounded human prompt excerpt, selected model, and routing latency. Logging is disabled by default; set `AUTOROUTER_SESSION_LOG_DIR` or use `setup --session-log-dir DIR` to enable it. See [Auto permission mode](reference.md#auto-permission-mode) and [session decision logs](reference.md#session-decision-logs).
|
|
16
16
|
|
|
17
|
+
Version `0.3.7` enables automatic Sonnet/Opus switching for compatible Auto-mode execution requests, including requests carrying the known server safety-review contract. The Auto profile defaults to Sonnet 5.5 and Opus 5.5, floors Haiku decisions to Sonnet, and retains the selected model through tool and goal continuations. Shared native context edits, mid-conversation system messages, and signed thinking history no longer pin new human tasks. Permission-classifier requests and safety verdicts remain unchanged; unknown contracts and incompatible model features still preserve a compatible model. Explicit model overrides remain in effect. See [Auto permission mode](reference.md#auto-permission-mode).
|
|
18
|
+
|
|
19
|
+
Version `0.4.0` completes the routing, configuration, history and performance improvement plan. Shared compatibility checks and durable task state preserve valid request features and confirmed tool/goal continuity across evaluator cache expiry, provider fallback and concurrent requests. New `config show/set/unset`, `sessions list/show` and `doctor --evaluate-local` commands support focused configuration edits, private metadata-only history and explicit local diagnostics. `setup --force` now merges saved settings; use `--replace` for deliberate replacement. Optional logging remains disabled by default; new schema-2 decision/outcome records separate selected and observed models, while the reader still accepts schema-1 files. Consumers parsing JSONL directly should account for both event kinds and the new schema. Identical concurrent evaluations are coalesced with independent cancellation, responses are bounded, and status persistence is asynchronous. Releases retain the tested archive and verify public npm availability and installation after submission. Actual 16 GiB/64 GiB Ollama results retain failed quality gates and comparison limits; Jev remains the default. See the [configuration/history reference](reference.md) and [hardware comparison](hardware-comparison.md).
|
|
20
|
+
|
|
17
21
|
The GitHub repository is private. Publishing to npm makes the tarball's runtime source, README, configuration example, license, and shipped documentation public. Model weights, user configuration, credentials, transcripts, session logs, local artifacts, and test fixtures are excluded. Review the archive before the first publication and whenever the package allowlist changes.
|
|
18
22
|
|
|
19
23
|
## What runs automatically
|
|
@@ -21,11 +25,15 @@ The GitHub repository is private. Publishing to npm makes the tarball's runtime
|
|
|
21
25
|
| Workflow | Trigger | Behavior |
|
|
22
26
|
| --- | --- | --- |
|
|
23
27
|
| [ci.yml](https://github.com/frapposelli/claude-autorouter/blob/main/.github/workflows/ci.yml) | Pull requests, pushes to `main`, manual runs, and calls from the release workflow | Syntax checks, tests, and package smoke tests on Ubuntu/macOS with Node 22/24 |
|
|
24
|
-
| [publish.yml](https://github.com/frapposelli/claude-autorouter/blob/main/.github/workflows/publish.yml) |
|
|
28
|
+
| [publish.yml](https://github.com/frapposelli/claude-autorouter/blob/main/.github/workflows/publish.yml) | Tag push; manual verification-only dispatch | Test and submit one canonical archive; independently verify registry availability and installation. Manual dispatch never publishes. |
|
|
25
29
|
|
|
26
30
|
A release tag must exactly equal `v` plus the version in `package.json`, and its commit must be reachable from `origin/main`. Package name and repository metadata must match `claude-autorouter` and `frapposelli/claude-autorouter`. Stable versions use npm's `latest` tag; prereleases such as `0.3.1-beta.1` use `next`.
|
|
27
31
|
|
|
28
|
-
The release workflow packs its candidate once and smoke-tests that exact `.tgz`. It
|
|
32
|
+
The release workflow packs its candidate once and smoke-tests that exact `.tgz`. It retains the archive, SHA-256 checksum, and commit-based release notes as Actions artifacts. The publishing job downloads the canonical artifact by immutable ID, checks every packaged file against the release checkout, and checks fresh npm metadata before submission. A new stable version must be greater than the current stable `latest`. An already-visible identical version skips publication; a different archive under that version is an error. Registry/network errors never count as proof that a version is unused.
|
|
33
|
+
|
|
34
|
+
Only the publishing job has `id-token: write`; there is no `NPM_TOKEN` or required GitHub environment. It runs on GitHub-hosted Ubuntu with Node 24 and npm 11.19.1. The separate verifier has read-only permissions and never changes npm distribution tags. The workflow serializes its publishers, but independent/manual publishers must coordinate: npm does not provide an atomic compare-and-swap for the `latest` tag. A newer `latest` observed during verification is reported as superseding this release and is never moved backward.
|
|
35
|
+
|
|
36
|
+
Verification polls uncached version metadata and the package document, checks the downloaded tarball against the tested archive, and installs the exact version from the public registry into a temporary prefix with an empty npm cache/config. It compares the installed file set and bytes with the canonical archive before invoking that executable’s `--version` and `--help` from an unrelated directory, with lifecycle scripts disabled and no evaluator credentials. Only a successful public install and matching artifact produce `verified`.
|
|
29
37
|
|
|
30
38
|
The project has no package dependencies or lockfile, so CI runs its scripts directly without `npm ci`. Live Claude/Jev calls, Ollama downloads, and private repository probes are not CI checks.
|
|
31
39
|
|
|
@@ -81,7 +89,7 @@ claude-autorouter --help
|
|
|
81
89
|
|
|
82
90
|
Then run `setup`, `doctor`, and a launch from outside the source checkout as appropriate for that machine. `doctor` is local-only; a live prompt separately verifies provider access. Keep the README's installation instructions aligned with the verified registry release.
|
|
83
91
|
|
|
84
|
-
Do not push `v0.2.0` to test automation after this bootstrap
|
|
92
|
+
Do not push `v0.2.0` to test automation after this bootstrap. Keep the original bootstrap archive; a current verifier only accepts releases matching its archive and repository validation rules. npm name/version pairs cannot be reused, including after unpublishing. See the [npm publish reference](https://docs.npmjs.com/cli/v11/commands/npm-publish/).
|
|
85
93
|
|
|
86
94
|
## 2. Authorize this workflow on npm
|
|
87
95
|
|
|
@@ -105,62 +113,96 @@ After a successful trusted release, npm recommends the optional **Publishing acc
|
|
|
105
113
|
|
|
106
114
|
## 3. Release subsequent versions by tag
|
|
107
115
|
|
|
108
|
-
The commands
|
|
116
|
+
Use the next unused version. The following commands use `0.4.0` as an example, not as a claim that it is currently available:
|
|
109
117
|
|
|
110
118
|
```sh
|
|
111
119
|
git switch main
|
|
112
120
|
git pull --ff-only origin main
|
|
113
|
-
npm version 0.
|
|
114
|
-
```
|
|
115
|
-
|
|
116
|
-
Review the version change and update any version-specific install examples or release notes. Check the candidate using the new filename:
|
|
117
|
-
|
|
118
|
-
```sh
|
|
121
|
+
npm version 0.4.0 --no-git-tag-version
|
|
119
122
|
npm run check
|
|
120
123
|
npm test
|
|
121
124
|
npm run release:pack
|
|
122
|
-
npm run test:package -- --archive ./dist/claude-autorouter-0.
|
|
125
|
+
npm run test:package -- --archive ./dist/claude-autorouter-0.4.0.tgz
|
|
123
126
|
git diff --check
|
|
124
127
|
```
|
|
125
128
|
|
|
126
|
-
|
|
129
|
+
Review the archive, version-specific documentation, and release notes. Commit all intended changes and get that commit onto `main` through the repository’s normal review process. Wait for CI to pass, then tag the exact release commit:
|
|
127
130
|
|
|
128
131
|
```sh
|
|
129
|
-
git
|
|
130
|
-
git
|
|
131
|
-
git
|
|
132
|
+
git switch main
|
|
133
|
+
git pull --ff-only origin main
|
|
134
|
+
git tag -a v0.4.0 -m "Release 0.4.0"
|
|
135
|
+
git push origin v0.4.0
|
|
132
136
|
```
|
|
133
137
|
|
|
134
|
-
|
|
138
|
+
The tag must match `package.json`. For a prerelease, use matching values such as `0.4.0-beta.1` / `v0.4.0-beta.1`; publication uses `next`, leaving `latest` unchanged. Release stable versions in increasing order, and wait for verification or investigate a pending submission before starting another stable release. The preflight blocks a new stable candidate that is not newer than the registry’s `latest`; queue order alone does not establish version order.
|
|
139
|
+
|
|
140
|
+
For local checks after a tag exists, `node scripts/release-check.mjs source v0.4.0` validates the clean checkout, tag, metadata, and main ancestry. `node scripts/release-check.mjs archive v0.4.0` validates the candidate checksum and contents. `dist/` must contain only that candidate’s `.tgz` and `.sha256`, so retain older artifacts elsewhere.
|
|
141
|
+
|
|
142
|
+
## 4. Inspect the release state and retain evidence
|
|
143
|
+
|
|
144
|
+
Open the tag’s run under [GitHub Actions](https://github.com/frapposelli/claude-autorouter/actions). Submission success is not proof that users can install the package. The verification job’s summary and retained JSON report distinguish:
|
|
145
|
+
|
|
146
|
+
| State | Meaning | Next step |
|
|
147
|
+
| --- | --- | --- |
|
|
148
|
+
| `preflight_ready` | The new candidate passed metadata and version-order checks; submission has not occurred | The initial workflow attempt may submit it |
|
|
149
|
+
| `submitted` | npm accepted the command, or an identical immutable version was already visible | Wait for independent verification |
|
|
150
|
+
| `validating_unavailable` | Metadata, tarball, distribution tag, or installation is still unavailable, or the registry is failing | Keep the original archive and rerun verification |
|
|
151
|
+
| `verified` | Exact archive integrity, distribution-tag state, and isolated public installation passed | Use the recorded upgrade command |
|
|
152
|
+
| `failed` | An input, provenance, integrity, version-order, or executable check failed | Investigate the report before changing anything |
|
|
153
|
+
|
|
154
|
+
The verifier polls for up to 15 minutes with backoff, then reports pending verification with a successful command exit. **A green workflow can therefore mean pending, not verified; read the recorded state.** An npm processing delay does not establish publication failure or justify a duplicate release. HTTP/network failures are reported separately from a missing version or tarball.
|
|
155
|
+
|
|
156
|
+
The workflow retains these artifacts for 90 days, subject to repository retention policy:
|
|
157
|
+
|
|
158
|
+
- `npm-package-<run-id>-<attempt>`: the canonical tested archive and SHA-256 file.
|
|
159
|
+
- `release-notes-<run-id>-<attempt>`: notes derived from the tagged source’s commits and archive identity.
|
|
160
|
+
- `release-submission-<run-id>-<attempt>`: preflight and, when accepted, submission reports.
|
|
161
|
+
- `release-verification-<run-id>-<attempt>`: the canonical archive, checksum, available notes, and verification report.
|
|
162
|
+
|
|
163
|
+
Download and retain the archive/checksum, notes, and report before Actions artifacts expire; attaching them to a GitHub Release is suitable for long-term retention. Artifact expiration is not a reason to repack a supposedly identical candidate for verification.
|
|
164
|
+
|
|
165
|
+
After the report says `verified`, use its exact version:
|
|
135
166
|
|
|
136
167
|
```sh
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
git push origin v0.3.2
|
|
168
|
+
npm install -g claude-autorouter@0.4.0
|
|
169
|
+
claude-autorouter --version
|
|
170
|
+
claude-autorouter --help
|
|
141
171
|
```
|
|
142
172
|
|
|
143
|
-
|
|
173
|
+
The verifier runs the equivalent exact-version registry install in isolation. If `latest` has since advanced, the report explicitly marks this release as superseded; it does not restore an older tag. An unqualified `npm install -g claude-autorouter` follows the registry’s current `latest` instead.
|
|
144
174
|
|
|
145
|
-
|
|
175
|
+
## 5. Rerun verification without publishing
|
|
146
176
|
|
|
147
|
-
|
|
177
|
+
Use the Actions **Run workflow** control for `publish.yml`, or the command below. Provide the release tag and the numeric run/artifact IDs from the original tag-triggered run:
|
|
148
178
|
|
|
149
179
|
```sh
|
|
150
|
-
|
|
151
|
-
|
|
180
|
+
gh workflow run publish.yml --ref main \
|
|
181
|
+
-f tag=v0.4.0 \
|
|
182
|
+
-f run_id=ORIGINAL_RUN_ID \
|
|
183
|
+
-f artifact_id=CANONICAL_NPM_PACKAGE_ARTIFACT_ID
|
|
152
184
|
```
|
|
153
185
|
|
|
154
|
-
|
|
186
|
+
This path validates that the artifact belongs to this repository’s original tag-triggered publishing workflow and matches the tag commit on `main`. It downloads those immutable bytes, checks their checksum and package metadata, and verifies npm availability. It cannot publish or alter npm tags and does not need npm OIDC permissions. Old runs may lack retained release notes; that does not prevent archive verification.
|
|
187
|
+
|
|
188
|
+
You can also download the original archive and its checksum and run the verifier independently from a current source checkout:
|
|
189
|
+
|
|
190
|
+
```sh
|
|
191
|
+
node scripts/release-verify.mjs verify v0.4.0 \
|
|
192
|
+
--archive /path/to/claude-autorouter-0.4.0.tgz \
|
|
193
|
+
--report artifacts/release-verification-0.4.0.json \
|
|
194
|
+
--timeout-ms 900000
|
|
195
|
+
```
|
|
155
196
|
|
|
156
|
-
|
|
197
|
+
Both files must retain their original names, and the archive must satisfy the public-package validation rules. This command performs only registry reads and a temporary isolated install; it neither publishes nor changes your installed CLI, npm login, or AutoRouter configuration. Exit `0` includes pending verification, so automation must inspect the report’s `state`; exit `1` means verification failed. A mismatch is never accepted as an already-published identical release.
|
|
157
198
|
|
|
158
|
-
## Recovering a
|
|
199
|
+
## Recovering a release
|
|
159
200
|
|
|
160
|
-
- **Checks
|
|
161
|
-
- **
|
|
162
|
-
- **
|
|
163
|
-
- **
|
|
164
|
-
- **
|
|
201
|
+
- **Checks/archive validation failed:** nothing was submitted by that failed path. Fix the cause, repeat validation, and follow the normal release process. Never move a published tag.
|
|
202
|
+
- **Registry preflight failed:** an HTTP/authentication/network failure is not evidence that the version is unused. Restore visibility before submitting.
|
|
203
|
+
- **npm rejected OIDC:** check the trusted-publisher fields, direct-publish permission, hosted runner, and OIDC permission. A job rerun deliberately does not resubmit a still-invisible version because an interrupted command may already have been accepted. Establish what happened before planning a new submission.
|
|
204
|
+
- **Submission timed out, the run was interrupted, or npm is processing it:** retain the original archive and use verification-only dispatch. A retry with a matching visible version skips `npm publish`; a retry with an invisible version only verifies it.
|
|
205
|
+
- **Existing version has different bytes:** stop. Never overwrite, unpublish, or replace the canonical artifact to reuse that version. Corrections require a new version.
|
|
206
|
+
- **Verification remains pending:** do not call it a failed publication. Rerun the verifier later with the original artifact, inspect npm’s public status if the registry is failing, and retain the evidence for support if validation remains stuck.
|
|
165
207
|
|
|
166
|
-
|
|
208
|
+
The workflow never repairs distribution tags automatically. npm’s [publish semantics](https://docs.npmjs.com/cli/commands/npm-publish/) make published name/version pairs immutable; [distribution tags](https://docs.npmjs.com/adding-dist-tags-to-packages/) are mutable references, so verification observes them without overwriting a newer release.
|