claude-autorouter 0.3.6 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. package/.env.example +12 -7
  2. package/CONTRIBUTING.md +37 -0
  3. package/README.md +43 -70
  4. package/bin/autorouter.mjs +40 -56
  5. package/docs/development.md +50 -2
  6. package/docs/hardware-benchmark.md +29 -0
  7. package/docs/hardware-comparison.md +55 -0
  8. package/docs/hardware-results-16gb.json +4002 -0
  9. package/docs/hardware-results-16gb.md +26 -0
  10. package/docs/hardware-results-64gb.json +4020 -0
  11. package/docs/reference.md +83 -40
  12. package/docs/releasing.md +76 -34
  13. package/docs/router-performance.json +1697 -0
  14. package/docs/router-performance.md +50 -0
  15. package/docs/status-performance.json +363 -0
  16. package/docs/status-performance.md +44 -0
  17. package/package.json +57 -9
  18. package/src/auto-routing.mjs +214 -0
  19. package/src/bounded-json.mjs +57 -0
  20. package/src/cli-help.mjs +87 -0
  21. package/src/config-command.mjs +141 -0
  22. package/src/config.mjs +52 -27
  23. package/src/contracts.mjs +123 -0
  24. package/src/evaluation-report.mjs +114 -0
  25. package/src/local-diagnostic.mjs +191 -0
  26. package/src/model-catalog.mjs +96 -0
  27. package/src/model-request.mjs +10 -6
  28. package/src/ollama-evaluator.mjs +9 -27
  29. package/src/onboarding.mjs +82 -23
  30. package/src/request-validation.mjs +54 -0
  31. package/src/response-observer.mjs +126 -18
  32. package/src/router.mjs +174 -70
  33. package/src/savings.mjs +74 -16
  34. package/src/server.mjs +79 -12
  35. package/src/session-history.mjs +261 -0
  36. package/src/session-log.mjs +9 -58
  37. package/src/status-state.mjs +110 -62
  38. package/src/statusline.mjs +57 -27
  39. package/src/telemetry-event.mjs +196 -0
  40. package/src/token-counter.mjs +3 -1
  41. package/src/turn-state.mjs +132 -0
  42. package/src/user-config.mjs +18 -8
package/docs/reference.md CHANGED
@@ -8,10 +8,17 @@
8
8
  | `claude-autorouter setup --auth-mode api-key` | Configure Jev and Anthropic API-key billing |
9
9
  | `claude-autorouter setup --client-profile auto` | Save a Sonnet/Opus profile compatible with Claude's Auto permission mode |
10
10
  | `claude-autorouter setup --stop-hook-block-cap 2` | Opt into a shorter native Stop-hook continuation cap during setup |
11
- | `claude-autorouter setup --session-log-dir DIR` | Save an opt-in directory for per-session JSONL decision logs |
11
+ | `claude-autorouter setup --session-log-dir DIR` | Save an opt-in directory for per-session JSONL decisions and outcomes |
12
12
  | `claude-autorouter setup --evaluator ollama --pull` | Configure the native local evaluator and download its selected model if missing |
13
13
  | `claude-autorouter setup --evaluator ollama --ollama-timeout-ms 0 --force` | Save a disabled runtime evaluator deadline |
14
- | `claude-autorouter setup --force` | Replace an existing user config |
14
+ | `claude-autorouter setup --force` | Update an existing config while preserving unrelated saved settings |
15
+ | `claude-autorouter setup --replace` | Explicitly replace the saved configuration |
16
+ | `claude-autorouter config show --json` | Inspect effective settings and default/file/environment provenance; secrets are hidden |
17
+ | `claude-autorouter config set KEY VALUE` | Change one nonsecret saved setting |
18
+ | `claude-autorouter config unset KEY` | Remove one saved override |
19
+ | `claude-autorouter sessions list [--json]` | Inspect optional saved local history |
20
+ | `claude-autorouter sessions show ID [--json]` | Show correlated decisions, outcomes and pricing coverage |
21
+ | `claude-autorouter doctor --evaluate-local` | Run synthetic classifier checks on the installed local Ollama model |
15
22
  | `claude-autorouter doctor` | Check config, Claude executable/login, and the selected local Ollama model without paid calls |
16
23
  | `claude-autorouter claude [arguments]` | Start a local router and pass arguments through to Claude Code |
17
24
  | `claude-autorouter serve` | Run the router for separately configured clients |
@@ -22,7 +29,7 @@ The launcher binds an ephemeral port on `127.0.0.1`, creates a temporary local c
22
29
 
23
30
  ## Configuration
24
31
 
25
- Setup defaults to subscription mode unless `--auth-mode` or `AUTOROUTER_AUTH_MODE` selects another mode. Jev remains the default evaluator; `--evaluator ollama` selects local classification. Setup prompts for required secrets without echoing them and writes a private JSON file. Stored keys are plaintext; keep the file private and out of source control. Supply keys through the environment when interactive input is unavailable. Subscription mode with Ollama requires no API keys. API-key authentication always requires `ANTHROPIC_API_KEY`, regardless of evaluator.
32
+ First setup defaults to subscription mode unless `--auth-mode` or `AUTOROUTER_AUTH_MODE` selects another mode. Jev remains the default evaluator; `--evaluator ollama` selects local classification. An existing configuration updated with `--force` keeps its saved choices unless a command-line flag changes them; unrelated environment overrides remain temporary. Setup prompts for required secrets without echoing them and writes a private JSON file. Stored keys are plaintext; keep the file private and out of source control. Supply keys through the environment when interactive input is unavailable. Subscription mode with Ollama requires no API keys. API-key authentication always requires `ANTHROPIC_API_KEY`, regardless of evaluator.
26
33
 
27
34
  The config path is selected in this order:
28
35
 
@@ -30,7 +37,7 @@ The config path is selected in this order:
30
37
  2. `$XDG_CONFIG_HOME/claude-autorouter/config.json`, when `XDG_CONFIG_HOME` is a nonempty absolute path.
31
38
  3. `~/.config/claude-autorouter/config.json`.
32
39
 
33
- The JSON file uses flat environment-style string keys, such as `AUTOROUTER_AUTH_MODE` and `TYPESAFE_API_KEY`. Environment values take precedence over the saved config. Use `setup --force` to replace existing configuration. The launcher does not discover or load a project's `.env` file. From a source checkout, explicitly loading one still works:
40
+ The JSON file uses flat environment-style string keys, such as `AUTOROUTER_AUTH_MODE` and `TYPESAFE_API_KEY`. Environment values take precedence over the saved config. Use `config set` or `config unset` to edit a single saved setting, or `setup --force` to update an existing configuration while retaining unrelated settings. `setup --replace` explicitly rebuilds a readable supported saved configuration; malformed or unsupported files are left intact for manual repair. The launcher does not discover or load a project's `.env` file. From a source checkout, explicitly loading one still works:
34
41
 
35
42
  ```sh
36
43
  node --env-file=.env bin/autorouter.mjs claude
@@ -48,11 +55,12 @@ For an environment-only subscription launch, set `AUTOROUTER_AUTH_MODE=subscript
48
55
  | `AUTOROUTER_CLIENT_PROFILE` | `compatible` | `native` retains client model/thinking settings; `auto` starts with Sonnet when no explicit model is set and excludes Haiku from task routing |
49
56
  | `AUTOROUTER_STATUSLINE` | enabled | `0` retains your existing status line |
50
57
  | `AUTOROUTER_DEBUG` | off | `1` enables launcher metadata logs on stderr |
51
- | `AUTOROUTER_SESSION_LOG_DIR` | off | Write per-session JSONL decision logs with prompt excerpts into this directory; unset or empty disables it |
58
+ | `AUTOROUTER_SESSION_LOG_DIR` | off | Write per-session JSONL decisions and outcomes into this directory; unset or empty disables it |
59
+ | `AUTOROUTER_SESSION_LOG_MODE` | `prompts` | `metadata` omits prompt excerpts; setting a mode alone does not enable logging |
52
60
  | `CLAUDE_CODE_STOP_HOOK_BLOCK_CAP` | unset; Claude currently uses `8` | Optional cap on consecutive Stop/SubagentStop continuations without tool use; `0` disables the cap |
53
61
  | `ENABLE_TOOL_SEARCH` | `true` in launcher when unset | Load MCP tool definitions on demand; explicit values are preserved |
54
62
  | `AUTOROUTER_HAIKU_MODEL` | `claude-haiku-4-5-20251001` | Routine tier |
55
- | `AUTOROUTER_SONNET_MODEL` | `claude-sonnet-5` | Standard tier |
63
+ | `AUTOROUTER_SONNET_MODEL` | `claude-sonnet-5`; `claude-sonnet-5-5` in the `auto` profile | Standard tier |
56
64
  | `AUTOROUTER_OPUS_MODEL` | `claude-opus-5-5` | Demanding tier and savings baseline |
57
65
  | `AUTOROUTER_JEV_MODEL` | `jev-latest` | Classifier version |
58
66
  | `AUTOROUTER_JEV_TIMEOUT_MS` | `1500` | Classifier deadline in milliseconds |
@@ -69,6 +77,20 @@ For an environment-only subscription launch, set `AUTOROUTER_AUTH_MODE=subscript
69
77
 
70
78
  Model access depends on your account. The policy recognizes specific Claude model versions; arbitrary gateway aliases do not automatically inherit their capabilities or context windows. Compare overrides with the [Anthropic model catalog](https://platform.claude.com/docs/en/models/overview).
71
79
 
80
+ ### Inspect and change settings
81
+
82
+ ```sh
83
+ claude-autorouter config show
84
+ claude-autorouter config show --json --check-all
85
+ claude-autorouter config set AUTOROUTER_OLLAMA_TIMEOUT_MS 0
86
+ claude-autorouter config unset AUTOROUTER_OLLAMA_TIMEOUT_MS
87
+ claude-autorouter config set TYPESAFE_API_KEY
88
+ ```
89
+
90
+ `show` reports whether each setting comes from a default, the saved file, or an environment override. Secret values are never displayed. The last command uses a hidden prompt; scripts can pipe a secret to `config set TYPESAFE_API_KEY --stdin`. Secret values are not accepted as command arguments. Edits validate and atomically update only the named saved setting. Environment overrides still apply after a saved change. An unset saved deadline returns to the model default unless an environment value overrides it.
91
+
92
+ Normal startup validates the selected evaluator; stale settings for the inactive evaluator do not prevent it from starting. `show --check-all` explicitly checks both. Blank numeric settings fail with their setting name; zero retains its documented meaning. `claude-autorouter help COMMAND` gives focused command help. `claude-autorouter claude --help` and `--version` call Claude directly without router setup or credentials.
93
+
72
94
  ## Ollama evaluator
73
95
 
74
96
  The local configuration documented here requires AutoRouter 0.3.2 or newer and remains experimental. It uses Ollama's native `/v1/systemone` decision endpoint for every model, replacing the chat backend from 0.2.0. Jev remains the default remote evaluator, using TypeSafe's `/v1/systemone` endpoint and a TypeSafe API key. Selecting Ollama never silently switches back to Jev. Haiku, Sonnet, or Opus still completes the task through Anthropic.
@@ -83,7 +105,7 @@ claude-autorouter doctor
83
105
  claude-autorouter claude
84
106
  ```
85
107
 
86
- `--force` replaces an existing user config. Setup detects the running local API. `--pull` authorizes downloading the chosen model when it is missing; without it, install the model yourself before setup. AutoRouter does not install Ollama, start its daemon, delete models, or download models during ordinary launches or `doctor` checks.
108
+ `--force` updates an existing user config and preserves its other settings and selected model unless explicitly changed. Setup detects the running local API. `--pull` authorizes downloading the chosen model when it is missing; without it, install the model yourself before setup. AutoRouter does not install Ollama, start its daemon, delete models, or download models during ordinary launches or `doctor` checks.
87
109
 
88
110
  ### Local model selection
89
111
 
@@ -115,7 +137,7 @@ The endpoint must be loopback (`127.0.0.1`, `localhost`, or `::1`), without a pa
115
137
 
116
138
  Local classification caps serialized evaluator state at both 3,000 characters and 3,000 UTF-8 bytes, including for non-ASCII prompts. Claude's top-level executor system instructions are excluded before budgeting; the current task, original task, and recent conversation excerpts remain. `/v1/systemone` receives the bounded state and routing criteria and returns a tier directly. The router retains each model's native context setting: 8,194 tokens for the default Nimble tag and 2,050 for the listed Tev1 tags. Tev1's smaller window includes the routing criteria and template as well as the excerpt; the byte limit does not guarantee every possible input fits. Context errors use the normal fallback. Returned confidence scores summarize choice-distribution entropy; they are not calibrated accuracy probabilities. `AUTOROUTER_MIN_CONFIDENCE` applies only to Jev. All capability, tool-continuation, thinking, and context guards still apply.
117
139
 
118
- The deadline covering local checks and classification defaults to 1,500 ms for Tev1 0.8B and custom/unrecognized tags, 15,000 ms for official Tev1 4B variants (including bare `tev1` and `latest`), and 30,000 ms for official Nimble variants. Official `library/` and `registry.ollama.ai/` aliases are recognized; a custom namespace such as `team/nimble` keeps the short default. An explicit timeout overrides the model default, including an old saved `1500`. Environment values override saved values on launch. Defaults are not written into the user config; `setup --force` replaces the config and saves an explicit timeout when supplied through `--ollama-timeout-ms` or the environment. During setup, the command-line flag takes precedence over the timeout environment value.
140
+ The deadline covering local checks and classification defaults to 1,500 ms for Tev1 0.8B and custom/unrecognized tags, 15,000 ms for official Tev1 4B variants (including bare `tev1` and `latest`), and 30,000 ms for official Nimble variants. Official `library/` and `registry.ollama.ai/` aliases are recognized; a custom namespace such as `team/nimble` keeps the short default. An explicit timeout overrides the model default, including an old saved `1500`. Environment values override saved values on launch. Defaults are not written into the user config. To update a saved deadline, use `config set AUTOROUTER_OLLAMA_TIMEOUT_MS N` or `setup --ollama-timeout-ms N --force`; unrelated environment overrides remain temporary. Environment-provided Ollama settings are saved during first setup, replacement, or explicit `--evaluator ollama` selection. The command-line timeout flag takes precedence over the environment.
119
141
 
120
142
  Set `AUTOROUTER_OLLAMA_TIMEOUT_MS=0` to remove AutoRouter's runtime evaluator timer while keeping the existing configuration:
121
143
 
@@ -137,11 +159,24 @@ If startup priming fails, the launcher warns and continues. An incompatible mode
137
159
 
138
160
  Historical measurements before 0.3.2: Tev1 4B timed out on all eight full-excerpt checks even with a 10-second diagnostic allowance; its short-task results did not establish a full-excerpt latency bound. Tev1 0.8B completed all eight within 1,500 ms. On the tested 16 GiB M4, Nimble timed out on all 12 tuning requests at 1,500 ms. A separate 30-second diagnostic completed 24 held-out classifications with 23 matching labels, but median routing took 11.4 seconds. The one error followed a misleading tier instruction. These historical results precede the 0.3.2 excerpt changes and do not establish guarantees for the longer defaults. See the [measurements and limitations](ollama-evaluation.md).
139
161
 
162
+ ### Test the installed local evaluator
163
+
164
+ ```sh
165
+ claude-autorouter doctor --evaluate-local
166
+ claude-autorouter doctor --evaluate-local --json
167
+ ```
168
+
169
+ This explicit diagnostic requires an installed local evaluator selected with `AUTOROUTER_EVALUATOR=ollama`. It uses synthetic tasks only, needs no Jev or Anthropic key, and makes no Claude inference calls. It checks local availability, measures a separate initial preparation call, then runs six uncached cases through the production classifier using the configured runtime deadline. A disabled runtime deadline remains disabled; Ctrl-C cancels the diagnostic.
170
+
171
+ The report separates availability, expected-label agreement, all-three-tier coverage, and latency. Residency is observed before calls, so it does not claim a controlled cold/warm benchmark. A pass establishes these six examples only. Incorrect predictions, fallback, or missing Haiku/Opus coverage fail even if Ollama answered successfully. Auto mode still checks all three raw evaluator labels; actual Auto routing applies its Sonnet floor separately.
172
+
173
+ The diagnostic does not download, unload, restart, or edit configuration. It refuses to run while unrelated models are resident and requires a positive keep-alive; normal routing still supports keep-alive `0`. Ordinary `doctor` remains a metadata check. A missing-model repair command uses the exact configured tag and preserves other settings.
174
+
140
175
  ### Migrating an older Ollama config
141
176
 
142
177
  Version 0.3.1 used a 1,500 ms deadline for every local model. After upgrading to 0.3.2, an explicitly saved or exported `AUTOROUTER_OLLAMA_TIMEOUT_MS=1500` still wins over the new model-specific defaults. Remove that override to use the defaults, or rerun setup with the desired model and `--ollama-timeout-ms N --force`. The `0` value and setup timeout flag require 0.3.2 or newer.
143
178
 
144
- Version 0.3.1 removed the Qwen chat backend and presets from 0.2.0. Existing downloaded models remain on disk, but an old Qwen model selection needs to be replaced with a native decision model. Run the setup command above with `--force`; it selects Nimble unless you pass `--ollama-model` or override the model through the environment. Remove or update any old `AUTOROUTER_OLLAMA_MODEL` environment value too, because environment variables override saved configuration. Update scripts to use `--ollama-model` when selecting a custom model.
179
+ Version 0.3.1 removed the Qwen chat backend and presets from 0.2.0. Existing downloaded models remain on disk, but an old Qwen model selection needs to be replaced with a native decision model. Run `claude-autorouter setup --evaluator ollama --ollama-model nimble:9b-q4_K_M --force`, or explicitly choose a Tev1 tag; merging with `--force` alone preserves the saved model. Remove or update any old `AUTOROUTER_OLLAMA_MODEL` environment value too, because environment variables override saved configuration. Update scripts to use `--ollama-model` when selecting a custom model.
145
180
 
146
181
  ## Data flow and authentication
147
182
 
@@ -163,7 +198,7 @@ Routine logs contain route, model, timing, usage, and error-category metadata, n
163
198
 
164
199
  ## Routing policy
165
200
 
166
- Each `/v1/messages` request is evaluated. Exact repeated bodies reuse a classification for five minutes. Both evaluators use a starting rubric choosing Haiku for routine work, Sonnet for ordinary engineering, and Opus for demanding reasoning. These choices require evaluation on your tasks; they are not quality guarantees.
201
+ Each eligible `/v1/messages` request is evaluated. Exact repeated bodies reuse a classification for five minutes; concurrent identical evaluations share one request. Cache identity includes evaluator configuration, rubric and requested model floor. Internal permission classifiers and other documented pass-through paths skip evaluation. Both evaluators use a starting rubric choosing Haiku for routine work, Sonnet for ordinary engineering, and Opus for demanding reasoning. These choices require evaluation on your tasks; they are not quality guarantees.
167
202
 
168
203
  The evaluator prioritizes the actual human request before startup metadata. Complete Claude reminder and tool-list blocks are excluded from that task excerpt, and long text retains its beginning and end. The outbound Anthropic request remains complete. Complexity outside the bounded excerpt can still be missed.
169
204
 
@@ -171,15 +206,15 @@ The following policy applies after classification:
171
206
 
172
207
  - Jev's 1,500 ms deadline covers the response body and has no retry. Successful calls return immediately. Timeouts, HTTP errors, and invalid responses fall back to Sonnet or retain an existing stronger model.
173
208
  - Jev confidence below 0.75 prevents a downgrade below Sonnet or the requested tier. Ollama returns a tier without calibrated confidence; its failure handling and compatibility guards still apply.
174
- - Tool continuations retain the model chosen at the start of the human turn. Session, agent, and prompt headers identify turns; normalized conversation content provides a fallback. Text feedback from a Stop hook also retains the model when it serves the same gateway prompt ID and the client has not changed its requested model, subject to capability and context checks. Moving prompt-cache markers does not create a new turn.
209
+ - Tool continuations retain the execution model confirmed by a successfully forwarded response. Active tasks and pending tools survive classification-cache expiry; retired task records expire separately. After restart, missing continuity is explicitly unknown. A selected model alone remains unconfirmed. Session, agent, and prompt headers identify turns; normalized conversation content provides a fallback. Text feedback from a Stop hook also retains the model when it serves the same gateway prompt ID and the client has not changed its requested model, subject to capability and context checks. Moving prompt-cache markers does not create a new turn.
175
210
  - Claude's local `/goal` command can omit the prompt-ID header. For that path, an exact feedback label matching a preceding expanded `/goal` command keeps the original task and conversation anchor. This narrow text fallback also recognizes Claude's repeated-goal truncation format; arbitrary hook text is not treated as a goal. Feedback remains in the evaluator's recent conversation and the full API request. A new human message becomes the current task normally. The status line shows `prompt pinned` or `goal pinned` when either text-continuation rule applies.
176
- - Thinking history, fixed-budget thinking, server tools, context management, and other recognized model-specific features preserve the current model. Adaptive thinking, effort, and output above 64K prevent a Haiku choice. Fields are never stripped to force a downgrade.
177
- - Mid-conversation `system` messages preserve the requested model and pass through unchanged. They do not count as a tool continuation by themselves.
178
- - Auxiliary requests, including Claude's Auto permission classifier, pass through on their requested model without Jev/Ollama evaluation, token checks, or turn-state changes. Compaction retains its existing model and context-capacity policy. Requests containing `safeguards` also pass through unchanged so the server's safety-review contract is preserved. Token counting and model discovery pass through without classification.
211
+ - Thinking history, fixed-budget thinking, server tools, context management, and other recognized model-specific features preserve the current model except for the verified shared capabilities of the modern Auto-mode Sonnet/Opus pair described below. Adaptive thinking, effort, and output above 64K prevent a Haiku choice. Fields are never stripped to force a downgrade.
212
+ - Mid-conversation `system` messages preserve the requested model unless both Auto-mode models support them; they always pass through unchanged. They do not count as a tool continuation by themselves.
213
+ - Auxiliary requests, including Claude's Auto permission classifier, pass through on their requested model without Jev/Ollama evaluation, token checks, or turn-state changes. Compaction retains its existing model and context-capacity policy. Recognized server-reviewed execution requests can route between compatible Sonnet/Opus models while retaining `safeguards` and all verdicts unchanged. Unknown safeguards contracts pass through. Token counting and model discovery pass through without classification.
179
214
 
180
215
  The default `compatible` profile starts Claude with Haiku-compatible requests and client-requested thinking disabled. AutoRouter uses adaptive thinking when upgrading these requests to Opus 5/5.5. Starting with 0.3.3, routing to exact `claude-sonnet-5-5` translates disabled thinking to `between_tools`, which skips up-front thinking but permits progress updates between tool calls. At `xhigh`/`max` effort, or when per-message effort differs from the top-level setting (default `high`), it uses adaptive thinking while preserving the effort settings. Token counting uses the same adaptation. Sonnet 5 still accepts disabled thinking and is unchanged. See [Sonnet 5.5 thinking requirements](https://platform.claude.com/docs/en/models/sonnet-5-5/migration-guide).
181
216
 
182
- Explicit native `between_tools` and unknown thinking modes retain the incoming model on new human turns. Signed thinking blocks pass through unchanged and existing tool turns retain their model pin. `AUTOROUTER_CLIENT_PROFILE=native` preserves normal client settings, which can constrain routing. An explicit Claude `--model` argument overrides the starting model, but `/model` and `--model` are requested models, not locks on the routed result. Native same-model requests and unknown model aliases are not rewritten; clients must use settings supported by that model.
217
+ Explicit native `between_tools` and unknown thinking modes retain the incoming model on new human turns outside Auto routing. In Auto routing, a known Sonnet 5.5 `between_tools` request can upgrade to Opus with adaptive thinking. Signed thinking blocks pass through unchanged and existing tool turns retain their model pin. `AUTOROUTER_CLIENT_PROFILE=native` preserves normal client settings, which can constrain routing. An explicit Claude `--model` argument overrides the starting model, but `/model` and `--model` are requested models, not locks on the routed result. Native same-model requests and unknown model aliases are not rewritten; clients must use settings supported by that model.
183
218
 
184
219
  The launcher enables `ENABLE_TOOL_SEARCH=true` when unset. Claude can otherwise disable on-demand MCP discovery when using a custom API address, loading connected-tool schemas into even a fresh conversation. Explicit values, including `false` or `auto:5`, are preserved. Managed settings and always-loaded tools can still affect deferral. See [Claude Code tool search](https://code.claude.com/docs/en/mcp#configure-tool-search).
185
220
 
@@ -187,7 +222,7 @@ The launcher enables `ENABLE_TOOL_SEARCH=true` when unset. Claude can otherwise
187
222
 
188
223
  The default `compatible` profile starts Claude as Haiku to permit three-tier routing. Claude's Auto permission mode does not support Haiku, even if AutoRouter routes an API request to Sonnet. Eligibility is based on Claude's selected client model. Gateways themselves are supported. See [Claude's Auto-mode requirements](https://code.claude.com/docs/en/permission-modes#eliminate-permission-prompts-with-auto-mode).
189
224
 
190
- AutoRouter 0.3.6 adds an `auto` client profile. Launch with:
225
+ AutoRouter 0.3.6 introduced an `auto` client profile but bypassed evaluation for server-reviewed execution. Version 0.3.7 adds automatic switching on those requests. Launch with:
191
226
 
192
227
  ```sh
193
228
  claude-autorouter claude --permission-mode auto
@@ -199,11 +234,15 @@ An explicit `--permission-mode auto` (or `--permission-mode=auto`) selects the p
199
234
  env AUTOROUTER_CLIENT_PROFILE=auto claude-autorouter claude
200
235
  ```
201
236
 
202
- The profile defaults the client to the configured Sonnet model and preserves native thinking and explicit model choices. It promotes a routine Haiku routing decision to Sonnet, while other model-feature and conversation-continuity guards still apply. Configured Sonnet/Opus targets must be known Auto-capable models. An explicit client `--model` or `ANTHROPIC_MODEL` can still make Auto unavailable if it selects Haiku or another unsupported model; choose a supported Sonnet or Opus instead.
237
+ The profile defaults to Sonnet 5.5 and Opus 5.5, preserving explicit configured model IDs. The evaluator chooses Sonnet or Opus for each new human task; a routine Haiku verdict uses Sonnet and shows `Auto mode floor`. Tool and `/goal` continuations stay on the selected execution model. Claude's initial client model remains separate from the routed model. An explicit client `--model` or `ANTHROPIC_MODEL` can still make Auto unavailable if it selects Haiku or another unsupported model; choose a supported Sonnet or Opus instead.
238
+
239
+ Claude remains responsible for enabling the permission mode and enforcing organization settings, account availability, and tool rules. The profile does not enable Auto by itself or override `disableAutoMode`. AutoRouter does not reproduce Claude's settings precedence to infer a mode from settings files. `setup --client-profile auto --force` updates an existing configuration while retaining its other settings. `config set AUTOROUTER_CLIENT_PROFILE auto` changes just that setting.
240
+
241
+ **Safety review:** Claude's permission-classifier requests retain their exact requested model and skip AutoRouter's evaluator. Ordinary execution requests with the known `dangerous_tool_use` version-1 review contract are evaluated and routed, retaining the complete `safeguards` object, beta headers, and streamed safety verdicts. This also detects server review when Auto was selected in Claude's UI rather than through the launch flag. Unknown or malformed review contracts and safeguarded compaction pass through with `Auto safety`; a target that cannot accept the request shows `Auto model guard`. AutoRouter never turns off server review or converts denied actions to approvals. See [server-side classifier review](https://code.claude.com/docs/en/permission-modes#server-side-classifier-review).
203
242
 
204
- Claude remains responsible for enabling the permission mode and enforcing organization settings, account availability, and tool rules. The profile does not enable Auto by itself or override `disableAutoMode`. AutoRouter does not reproduce Claude's settings precedence to infer a mode from settings files. For new saved configurations, `setup --client-profile auto` persists the profile; for an existing config, change only `AUTOROUTER_CLIENT_PROFILE` to `"auto"` to retain your other settings. `setup --force` replaces the config.
243
+ **Shared execution capabilities:** automatic Auto routing supports exact Sonnet 5/5.5 and Opus 5/5.5 IDs. The default 5.5 pair shares a native 1M context window, adaptive thinking, native context-editing strategies, and mid-conversation system updates. These fields and existing signed thinking no longer pin every future human task. Sonnet 5 cannot accept mid-conversation system messages, per-message effort, or task budgets; use Sonnet 5.5 for those sessions. Unknown context-editing strategies, specialized server tools, fixed thinking budgets, fast mode, older/custom targets, and other incompatible requests still retain a compatible model. A Sonnet 5.5 `between_tools` request uses adaptive thinking when upgraded to Opus; effort and conversation history remain unchanged.
205
244
 
206
- **Routing limits:** Claude's permission-classifier requests retain their exact requested model and skip AutoRouter's evaluator. Requests with server-side `safeguards` also retain their model and full body, and safety verdicts stream back unchanged. The status line shows `pass-through · Auto safety` for these requests. Recent Claude versions normally request server-side review through gateways, so Auto-mode sessions can stay on their client-selected model rather than switching between Sonnet and Opus. Ordinary routable requests use the Sonnet/Opus floor, shown as `Auto mode floor` when it changes a Haiku choice. AutoRouter does not disable server review to enable routing. See [server-side classifier review](https://code.claude.com/docs/en/permission-modes#server-side-classifier-review).
245
+ Thinking blocks stay verbatim in the conversation. Anthropic may drop blocks the selected model cannot read, so switching models does not preserve access to every model's private reasoning on every turn. User text, tool results, and prior answers remain available. Returning to a model can make its preserved thinking readable again. Model changes can also miss the prior model's prompt cache. See [preserved thinking and model switching](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking).
207
246
 
208
247
  ### Context capacity
209
248
 
@@ -223,7 +262,7 @@ The `claude` launcher automatically adds a temporary [status-line command](https
223
262
  | --- | --- |
224
263
  | `Sonnet 5 selected` | Routing chose this model; Anthropic has not confirmed it yet |
225
264
  | `Opus 5.5` / `last Opus 5.5` | Provider-confirmed streaming or most recent model |
226
- | `Jev` / `Ollama`, with `cache` or `fallback` when applicable | Classification source; timing includes concurrent context checks |
265
+ | `Jev` / `Ollama`, with `cache` or `fallback` when applicable | Classification source; displayed routing time includes evaluator waiting and context checks |
227
266
  | `Jev→Haiku` or `Ollama→Haiku` beside Sonnet | A policy guard overrode the evaluator's Haiku choice |
228
267
  | `large context` | Token count exceeded the small-model input budget |
229
268
  | `size unverified` | Token checking failed or was unavailable; conservative guard applied |
@@ -231,51 +270,55 @@ The `claude` launcher automatically adds a temporary [status-line command](https
231
270
  | `CLI ctx` | Claude's client accounting, shown when its window differs or API capacity is unknown |
232
271
  | `est saved … vs Opus` | Cumulative API-equivalent token-cost estimate |
233
272
 
234
- Background agents and auxiliary requests cannot replace the foreground model. Errors, fallback, cancellation, and stale/offline state remain visible. The command reads a local snapshot and makes no network requests. It respects terminal width and `NO_COLOR`; lower-priority fields disappear on narrow terminals.
273
+ Background agents and auxiliary requests cannot replace the foreground model. Errors, fallback, cancellation, incomplete-response evidence and stale/offline state remain visible. The command reads a local snapshot and makes no network requests. It respects terminal width and `NO_COLOR`; lower-priority fields disappear on narrow terminals.
235
274
 
236
275
  Context includes uncached input, cache reads, and cache writes, excluding output to match [Claude's percentage formula](https://code.claude.com/docs/en/statusline#context-window-fields). It is current context rather than cumulative usage. Historical usage is marked `last`; compaction resets stale readings. The compatible client's 200K window can reach 100% while a routed Sonnet request uses only part of its 1M window. Displaying both does not change Claude's compaction threshold.
237
276
 
238
277
  Savings compare the actual models' API token prices with the configured Opus model's prices for the **same reported counts and cache profile**. The percentage is `(Opus cost − routed cost) / Opus cost`. Input, output, cache reads, and 5-minute/1-hour cache writes are priced separately using the bundled table based on [Anthropic's published USD pricing](https://platform.claude.com/docs/en/about-claude/pricing).
239
278
 
240
- Totals include completed main, agent, and auxiliary calls for the current session observed by this router process. Streaming usage is counted once, and totals reset when the router or session restarts. Unrecognized prices or unsupported usage produce `partial` or `savings unavailable`. Higher routed costs show `est extra`. The arithmetic runs locally.
279
+ Totals include completed main, agent, and auxiliary calls for the current session observed by this router process. Streaming usage is counted once, and totals reset when the router or session restarts. Unrecognized prices or unsupported usage produce `partial`, `unpriced N` or `savings unavailable`. Saved history identifies pricing table `2026-09-29.1` (reviewed September 29, 2026) and unpriced reason counts; unknown historical table versions are not repriced. Higher routed costs show `est extra`. The arithmetic runs locally.
241
280
 
242
281
  This estimate does not measure subscription bill savings or quota credits. It excludes Jev charges, local compute costs, tool fees, negotiated discounts, and unpriced requests. A real Opus run can produce different tokens and cache hits. Incomplete streams, unknown cache-write TTLs, unsupported pricing modifiers, and unrecognized model versions are excluded rather than guessed. The rate table requires updates when prices change.
243
282
 
283
+ Historical integration observations cover Claude Code 2.1.284 and 2.1.285. The source-only versioned corpus in `test/fixtures/claude-protocol-v1.json` separates newly authored synthetic contracts from those dated observations and their artifact hashes. Gateway tests exercise Auto safeguards, thinking, deferred tools, compaction, goal scoping, fallback ownership and usage while preserving response bytes. They do not establish current-source live compatibility, evaluator accuracy or downstream task quality. Doctor reports the installed executable version separately; discovering a binary is not evidence that its protocol or Auto eligibility has been tested. Opt-in live validation supplements the synthetic fixtures.
284
+
244
285
  Set `AUTOROUTER_STATUSLINE=0` to retain an existing status line. Other `--settings` values are retained in the temporary overlay; source-relative Read/Edit rules keep their anchors. Ambiguous relative sandbox paths cause the launcher to skip the overlay and pass original settings through with a notice. Safe mode disables custom status lines; print mode has no status-line UI. Standalone `serve` does not install one.
245
286
 
246
287
  ## Session decision logs
247
288
 
248
- AutoRouter 0.3.6 adds optional persistent logs, separate from stderr and the temporary status-line snapshot. Logging is disabled by default. Enable it for one launch:
289
+ Logging is optional and disabled by default. Enable it for one launch:
249
290
 
250
291
  ```sh
251
292
  env AUTOROUTER_SESSION_LOG_DIR="$HOME/.local/state/claude-autorouter/sessions" \
252
293
  claude-autorouter claude
253
294
  ```
254
295
 
255
- The setting also works with `serve` and the Auto-compatible profile. For a new saved configuration, add `--session-log-dir DIR` to `setup`. For an existing config, add `AUTOROUTER_SESSION_LOG_DIR` with an absolute directory path to preserve your other settings. Setup resolves relative paths at setup time; an environment-only relative path resolves from the launch directory. Environment values override saved values; `AUTOROUTER_SESSION_LOG_DIR=''` disables a saved preference for one launch. `doctor` reports the setting without creating log files.
296
+ Or save metadata-only history, without prompt excerpts:
256
297
 
257
- Files are named `autorouter-session-*.jsonl`: one file per observed Claude session within a router launch, with a timestamp, random launch identifier, and hashed session identifier in the name. A resumed session in a new launch creates a new file. Requests without a session header share an anonymous file for that launch. Subagents with the same session ID share its file and retain their agent ID. Files are created only when a decision is recorded, and remain after the session ends.
298
+ ```sh
299
+ claude-autorouter config set AUTOROUTER_SESSION_LOG_MODE metadata
300
+ claude-autorouter config set AUTOROUTER_SESSION_LOG_DIR "$HOME/.local/state/claude-autorouter/sessions"
301
+ claude-autorouter sessions list
302
+ claude-autorouter sessions show autorouter-session-EXAMPLE
303
+ claude-autorouter sessions show autorouter-session-EXAMPLE --json
304
+ ```
258
305
 
259
- Every line is a standalone JSON object. The key fields look like this (additional IDs and routing metadata are included):
306
+ Use the exact `id` printed by `sessions list`. Commands need no evaluator credentials and do not contact providers. Human summaries distinguish selected models, observed serving models, confirmed completions, failures, cancellations, and pending/unconfirmed requests. They report routing latency, fallback and override counts, and API-equivalent savings coverage. A selected model or an HTTP 200 alone does not prove successful inference. Old schema-1 decision logs remain readable and explicitly lack outcome evidence.
260
307
 
261
- ```json
262
- {"schema_version":1,"event":"decision","timestamp":"2026-09-30T12:00:00.000Z","session_id":"example-session","prompt_excerpt":"Fix the typo in README.md","prompt_truncated":false,"requested_model":"claude-haiku-4-5-20251001","selected_model":"claude-sonnet-5","decision_latency_ms":214.37,"source":"jev","reason":"classified"}
263
- ```
308
+ `prompts` mode preserves the existing excerpt behavior when a log directory is enabled. Main requests retain at most 500 Unicode characters of the human task; recognized tool and goal continuations retain the originating task. Auxiliary, subagent, compaction, workflow, and attachment-only requests have empty excerpts. Metadata mode omits the excerpt fields entirely. Neither mode logs authentication headers, provider replies, full transcripts, or tool payloads. Text entered directly in a prompt can appear in an enabled prompt excerpt.
264
309
 
265
- - `prompt_excerpt`: up to 500 Unicode characters of the current human task for main requests or requests without a class header. Tool continuations and recognized `/goal` feedback keep the originating human task. System instructions, standalone reminder blocks, tool results, images, documents, and thinking are omitted. Auxiliary classifiers, compaction, subagents, and workflows have empty excerpts; a new attachment-only task also has an empty excerpt. `prompt_truncated` indicates that text exceeded the limit.
266
- - `selected_model`: AutoRouter's final selected model after compatibility checks, before the upstream response. It does not confirm which model successfully answered.
267
- - `decision_latency_ms`: time spent making the routing decision, including evaluator waiting, cache lookup, and any context checks. It excludes Claude generation time and log writing. A `passthrough` entry can be near zero because no evaluator was called.
268
- - `source` and `reason`: distinguish evaluator choices, cache hits, fallbacks, turn/model constraints, and native safety pass-through. `classified_tier` is included when an evaluator returned a tier, which may differ from the final selected model.
310
+ The settings also work with `serve` and every client profile. `setup --session-log-dir DIR --session-log-mode metadata --force` updates an existing configuration. Setup resolves relative directories at setup time; environment-only paths resolve from the launch directory. `AUTOROUTER_SESSION_LOG_DIR=''` disables a saved directory for one launch. Setting only the mode never enables logging. `doctor` reports preferences without creating files.
269
311
 
270
- Inspect a file with:
312
+ Files are named `autorouter-session-*.jsonl`, one per observed Claude session in each router launch. Resuming a session in a new launch creates a new file. Requests without a session header share an anonymous file; subagents retain their agent IDs. Every line is a bounded schema-2 JSON record:
271
313
 
272
- ```sh
273
- jq -c '{prompt_excerpt, selected_model, decision_latency_ms, source, reason}' /path/to/autorouter-session-EXAMPLE.jsonl
274
- ```
314
+ - `decision`: requested and selected models, evaluator verdict/source, policy reason, compatibility and continuity detail. `evaluation_latency_ms` measures evaluation/cache waiting; `routing_latency_ms` (also `decision_latency_ms`) includes compatibility and context checks.
315
+ - `outcome`: the same `request_id`, observed `confirmed_model` and bounded model transitions, safe error category, `completed`, `error` or `cancelled` status, and separate `completion_confirmed` evidence. It includes usage when available, `first_response_ms` from upstream forwarding to response headers, and total request latency. Early errors can have an outcome without a decision.
316
+
317
+ Outcomes record the configured Opus baseline and pricing-table version. History prices only successfully completed, supported usage with a recognized recorded table version and baseline. Unknown prices, missing or partial streams, ambiguous/mixed-model usage, and legacy records remain unpriced with a reason. Estimates never represent subscription charges. Status savings identify partial coverage; history exposes the version and counts.
275
318
 
276
- There is one record per completed routing decision, including requests whose upstream call later fails. Requests rejected before routing or cancelled before a decision are not recorded. Logs contain user text and are local plaintext: the feature is disabled by default, new directories use `0700`, and files use `0600`. Existing directory permissions are left unchanged. Authentication headers, provider replies, full transcripts, and tool payloads are not logged; text you put directly in a prompt can appear in its excerpt.
319
+ New directories use `0700`, files use `0600`; existing directory permissions are unchanged. Asynchronous writes use a bounded 1 MiB queue and at most 128 session files per process. Normal shutdown drains accepted records. Storage failure or queue limits disable further logging with one generic warning while routing continues. Abrupt termination can lose unwritten records.
277
320
 
278
- Writes run asynchronously through a bounded 1 MiB queue and support up to 128 session files per router process. Normal shutdown drains accepted records. Filesystem failure or a queue/session limit disables further logging with one generic warning while routing continues. An abrupt process kill or storage failure can lose unwritten records. Logs are retained without automatic rotation or deletion; manage them in your chosen directory. Log filenames are ignored by this repository and excluded from the npm package.
321
+ History reads at most 100 files, 4 MiB per file, 16 MiB total and 5,000 records per file. Limits, malformed records and partial tails are reported as partial coverage. Files are never automatically rotated or deleted; manage retention in your chosen directory. Log filenames are ignored by this repository and excluded from the npm package.
279
322
 
280
323
  ## Troubleshooting
281
324
 
@@ -315,7 +358,7 @@ The environment-only command also works on AutoRouter 0.3.4. Saved configuration
315
358
  "CLAUDE_CODE_STOP_HOOK_BLOCK_CAP": "2"
316
359
  ```
317
360
 
318
- For a new configuration, use `claude-autorouter setup --stop-hook-block-cap 2`; the flag works with either evaluator and overrides the environment during setup. Runtime environment values override saved configuration. `setup --force` replaces the entire config, so keep your existing evaluator/authentication options if using it. `doctor` reports the cap when configured. AutoRouter accepts nonnegative safe integers and leaves the setting absent unless you opt in.
361
+ For a new configuration, use `claude-autorouter setup --stop-hook-block-cap 2`; the flag works with either evaluator and overrides the environment during setup. Runtime environment values override saved configuration. To change only a saved cap, use `config set CLAUDE_CODE_STOP_HOOK_BLOCK_CAP 2` or `setup --stop-hook-block-cap 2 --force`; both preserve unrelated saved settings. `doctor` reports the cap when configured. AutoRouter accepts nonnegative safe integers and leaves the setting absent unless you opt in.
319
362
 
320
363
  ### Other session issues
321
364
 
package/docs/releasing.md CHANGED
@@ -14,6 +14,10 @@ Version `0.3.5` adds opt-in saved configuration for Claude's native `CLAUDE_CODE
14
14
 
15
15
  Version `0.3.6` fixes Auto permission-mode launches with an Auto-compatible Sonnet/Opus profile while preserving Claude's permission classifiers and server safety-review requests. It also adds optional per-session JSONL decision logs containing a bounded human prompt excerpt, selected model, and routing latency. Logging is disabled by default; set `AUTOROUTER_SESSION_LOG_DIR` or use `setup --session-log-dir DIR` to enable it. See [Auto permission mode](reference.md#auto-permission-mode) and [session decision logs](reference.md#session-decision-logs).
16
16
 
17
+ Version `0.3.7` enables automatic Sonnet/Opus switching for compatible Auto-mode execution requests, including requests carrying the known server safety-review contract. The Auto profile defaults to Sonnet 5.5 and Opus 5.5, floors Haiku decisions to Sonnet, and retains the selected model through tool and goal continuations. Shared native context edits, mid-conversation system messages, and signed thinking history no longer pin new human tasks. Permission-classifier requests and safety verdicts remain unchanged; unknown contracts and incompatible model features still preserve a compatible model. Explicit model overrides remain in effect. See [Auto permission mode](reference.md#auto-permission-mode).
18
+
19
+ Version `0.4.0` completes the routing, configuration, history and performance improvement plan. Shared compatibility checks and durable task state preserve valid request features and confirmed tool/goal continuity across evaluator cache expiry, provider fallback and concurrent requests. New `config show/set/unset`, `sessions list/show` and `doctor --evaluate-local` commands support focused configuration edits, private metadata-only history and explicit local diagnostics. `setup --force` now merges saved settings; use `--replace` for deliberate replacement. Optional logging remains disabled by default; new schema-2 decision/outcome records separate selected and observed models, while the reader still accepts schema-1 files. Consumers parsing JSONL directly should account for both event kinds and the new schema. Identical concurrent evaluations are coalesced with independent cancellation, responses are bounded, and status persistence is asynchronous. Releases retain the tested archive and verify public npm availability and installation after submission. Actual 16 GiB/64 GiB Ollama results retain failed quality gates and comparison limits; Jev remains the default. See the [configuration/history reference](reference.md) and [hardware comparison](hardware-comparison.md).
20
+
17
21
  The GitHub repository is private. Publishing to npm makes the tarball's runtime source, README, configuration example, license, and shipped documentation public. Model weights, user configuration, credentials, transcripts, session logs, local artifacts, and test fixtures are excluded. Review the archive before the first publication and whenever the package allowlist changes.
18
22
 
19
23
  ## What runs automatically
@@ -21,11 +25,15 @@ The GitHub repository is private. Publishing to npm makes the tarball's runtime
21
25
  | Workflow | Trigger | Behavior |
22
26
  | --- | --- | --- |
23
27
  | [ci.yml](https://github.com/frapposelli/claude-autorouter/blob/main/.github/workflows/ci.yml) | Pull requests, pushes to `main`, manual runs, and calls from the release workflow | Syntax checks, tests, and package smoke tests on Ubuntu/macOS with Node 22/24 |
24
- | [publish.yml](https://github.com/frapposelli/claude-autorouter/blob/main/.github/workflows/publish.yml) | Push of a tag matching `v*` | Validate release, run CI, pack and test the candidate, then publish the verified archive |
28
+ | [publish.yml](https://github.com/frapposelli/claude-autorouter/blob/main/.github/workflows/publish.yml) | Tag push; manual verification-only dispatch | Test and submit one canonical archive; independently verify registry availability and installation. Manual dispatch never publishes. |
25
29
 
26
30
  A release tag must exactly equal `v` plus the version in `package.json`, and its commit must be reachable from `origin/main`. Package name and repository metadata must match `claude-autorouter` and `frapposelli/claude-autorouter`. Stable versions use npm's `latest` tag; prereleases such as `0.3.1-beta.1` use `next`.
27
31
 
28
- The release workflow packs its candidate once and smoke-tests that exact `.tgz`. It uploads the archive and SHA-256 checksum as an Actions artifact. A separate publishing job downloads that artifact by its immutable ID, checks the checksum and every packaged file against the release checkout, then runs `npm publish` with scripts disabled. The publish job uses a GitHub-hosted Ubuntu runner, Node 24, and npm 11.19.1. Only that job has `id-token: write`; there is no `NPM_TOKEN` secret or required GitHub environment. Failed checks prevent publication.
32
+ The release workflow packs its candidate once and smoke-tests that exact `.tgz`. It retains the archive, SHA-256 checksum, and commit-based release notes as Actions artifacts. The publishing job downloads the canonical artifact by immutable ID, checks every packaged file against the release checkout, and checks fresh npm metadata before submission. A new stable version must be greater than the current stable `latest`. An already-visible identical version skips publication; a different archive under that version is an error. Registry/network errors never count as proof that a version is unused.
33
+
34
+ Only the publishing job has `id-token: write`; there is no `NPM_TOKEN` or required GitHub environment. It runs on GitHub-hosted Ubuntu with Node 24 and npm 11.19.1. The separate verifier has read-only permissions and never changes npm distribution tags. The workflow serializes its publishers, but independent/manual publishers must coordinate: npm does not provide an atomic compare-and-swap for the `latest` tag. A newer `latest` observed during verification is reported as superseding this release and is never moved backward.
35
+
36
+ Verification polls uncached version metadata and the package document, checks the downloaded tarball against the tested archive, and installs the exact version from the public registry into a temporary prefix with an empty npm cache/config. It compares the installed file set and bytes with the canonical archive before invoking that executable’s `--version` and `--help` from an unrelated directory, with lifecycle scripts disabled and no evaluator credentials. Only a successful public install and matching artifact produce `verified`.
29
37
 
30
38
  The project has no package dependencies or lockfile, so CI runs its scripts directly without `npm ci`. Live Claude/Jev calls, Ollama downloads, and private repository probes are not CI checks.
31
39
 
@@ -81,7 +89,7 @@ claude-autorouter --help
81
89
 
82
90
  Then run `setup`, `doctor`, and a launch from outside the source checkout as appropriate for that machine. `doctor` is local-only; a live prompt separately verifies provider access. Keep the README's installation instructions aligned with the verified registry release.
83
91
 
84
- Do not push `v0.2.0` to test automation after this bootstrap: it would attempt to publish an existing version. npm name/version pairs cannot be reused, including after unpublishing. See the [npm publish reference](https://docs.npmjs.com/cli/v11/commands/npm-publish/).
92
+ Do not push `v0.2.0` to test automation after this bootstrap. Keep the original bootstrap archive; a current verifier only accepts releases matching its archive and repository validation rules. npm name/version pairs cannot be reused, including after unpublishing. See the [npm publish reference](https://docs.npmjs.com/cli/v11/commands/npm-publish/).
85
93
 
86
94
  ## 2. Authorize this workflow on npm
87
95
 
@@ -105,62 +113,96 @@ After a successful trusted release, npm recommends the optional **Publishing acc
105
113
 
106
114
  ## 3. Release subsequent versions by tag
107
115
 
108
- The commands below illustrate the `0.3.2` release. For a new release, substitute the next unused version throughout; never reuse a published version:
116
+ Use the next unused version. The following commands use `0.4.0` as an example, not as a claim that it is currently available:
109
117
 
110
118
  ```sh
111
119
  git switch main
112
120
  git pull --ff-only origin main
113
- npm version 0.3.2 --no-git-tag-version
114
- ```
115
-
116
- Review the version change and update any version-specific install examples or release notes. Check the candidate using the new filename:
117
-
118
- ```sh
121
+ npm version 0.4.0 --no-git-tag-version
119
122
  npm run check
120
123
  npm test
121
124
  npm run release:pack
122
- npm run test:package -- --archive ./dist/claude-autorouter-0.3.2.tgz
125
+ npm run test:package -- --archive ./dist/claude-autorouter-0.4.0.tgz
123
126
  git diff --check
124
127
  ```
125
128
 
126
- Commit the intended release changes and get that commit onto `main`, either through a pull request or a direct push allowed by the repository's branch rules. For a direct push with only the version changed:
129
+ Review the archive, version-specific documentation, and release notes. Commit all intended changes and get that commit onto `main` through the repository’s normal review process. Wait for CI to pass, then tag the exact release commit:
127
130
 
128
131
  ```sh
129
- git add package.json
130
- git commit -m "Release 0.3.2"
131
- git push origin main
132
+ git switch main
133
+ git pull --ff-only origin main
134
+ git tag -a v0.4.0 -m "Release 0.4.0"
135
+ git push origin v0.4.0
132
136
  ```
133
137
 
134
- Include any intentional documentation or release-note edits in that commit too. There is no publication from a branch push or PR merge. Wait for CI to pass, then tag that exact release commit:
138
+ The tag must match `package.json`. For a prerelease, use matching values such as `0.4.0-beta.1` / `v0.4.0-beta.1`; publication uses `next`, leaving `latest` unchanged. Release stable versions in increasing order, and wait for verification or investigate a pending submission before starting another stable release. The preflight blocks a new stable candidate that is not newer than the registry’s `latest`; queue order alone does not establish version order.
139
+
140
+ For local checks after a tag exists, `node scripts/release-check.mjs source v0.4.0` validates the clean checkout, tag, metadata, and main ancestry. `node scripts/release-check.mjs archive v0.4.0` validates the candidate checksum and contents. `dist/` must contain only that candidate’s `.tgz` and `.sha256`, so retain older artifacts elsewhere.
141
+
142
+ ## 4. Inspect the release state and retain evidence
143
+
144
+ Open the tag’s run under [GitHub Actions](https://github.com/frapposelli/claude-autorouter/actions). Submission success is not proof that users can install the package. The verification job’s summary and retained JSON report distinguish:
145
+
146
+ | State | Meaning | Next step |
147
+ | --- | --- | --- |
148
+ | `preflight_ready` | The new candidate passed metadata and version-order checks; submission has not occurred | The initial workflow attempt may submit it |
149
+ | `submitted` | npm accepted the command, or an identical immutable version was already visible | Wait for independent verification |
150
+ | `validating_unavailable` | Metadata, tarball, distribution tag, or installation is still unavailable, or the registry is failing | Keep the original archive and rerun verification |
151
+ | `verified` | Exact archive integrity, distribution-tag state, and isolated public installation passed | Use the recorded upgrade command |
152
+ | `failed` | An input, provenance, integrity, version-order, or executable check failed | Investigate the report before changing anything |
153
+
154
+ The verifier polls for up to 15 minutes with backoff, then reports pending verification with a successful command exit. **A green workflow can therefore mean pending, not verified; read the recorded state.** An npm processing delay does not establish publication failure or justify a duplicate release. HTTP/network failures are reported separately from a missing version or tarball.
155
+
156
+ The workflow retains these artifacts for 90 days, subject to repository retention policy:
157
+
158
+ - `npm-package-<run-id>-<attempt>`: the canonical tested archive and SHA-256 file.
159
+ - `release-notes-<run-id>-<attempt>`: notes derived from the tagged source’s commits and archive identity.
160
+ - `release-submission-<run-id>-<attempt>`: preflight and, when accepted, submission reports.
161
+ - `release-verification-<run-id>-<attempt>`: the canonical archive, checksum, available notes, and verification report.
162
+
163
+ Download and retain the archive/checksum, notes, and report before Actions artifacts expire; attaching them to a GitHub Release is suitable for long-term retention. Artifact expiration is not a reason to repack a supposedly identical candidate for verification.
164
+
165
+ After the report says `verified`, use its exact version:
135
166
 
136
167
  ```sh
137
- git switch main
138
- git pull --ff-only origin main
139
- git tag -a v0.3.2 -m "Release 0.3.2"
140
- git push origin v0.3.2
168
+ npm install -g claude-autorouter@0.4.0
169
+ claude-autorouter --version
170
+ claude-autorouter --help
141
171
  ```
142
172
 
143
- Before pushing, confirm `package.json` contains `0.3.2` and the tag points to the intended commit. For a prerelease, use a matching version/tag such as `0.4.0-beta.1` / `v0.4.0-beta.1`; it will publish under `next`, leaving `latest` unchanged.
173
+ The verifier runs the equivalent exact-version registry install in isolation. If `latest` has since advanced, the report explicitly marks this release as superseded; it does not restore an older tag. An unqualified `npm install -g claude-autorouter` follows the registry’s current `latest` instead.
144
174
 
145
- Release stable versions in increasing version order, one tag at a time, and wait for each run to finish before pushing the next stable tag. The workflow queues releases without canceling an active run, but queue order does not sort semantic versions. Publishing an older stable version afterward could move `latest` backward; there is no registry version-order gate.
175
+ ## 5. Rerun verification without publishing
146
176
 
147
- Open the tag's run under [GitHub Actions](https://github.com/frapposelli/claude-autorouter/actions). Under **Artifacts**, download `npm-package-<run-id>-<run-attempt>`, which contains the `.tgz` and checksum used for publication. Artifacts expire after 30 days, so retain them with the release record. After the publish job succeeds, verify the registry version and tags:
177
+ Use the Actions **Run workflow** control for `publish.yml`, or the command below. Provide the release tag and the numeric run/artifact IDs from the original tag-triggered run:
148
178
 
149
179
  ```sh
150
- npm view claude-autorouter@0.3.2 version dist.integrity --registry https://registry.npmjs.org/
151
- npm view claude-autorouter dist-tags --json --registry https://registry.npmjs.org/
180
+ gh workflow run publish.yml --ref main \
181
+ -f tag=v0.4.0 \
182
+ -f run_id=ORIGINAL_RUN_ID \
183
+ -f artifact_id=CANONICAL_NPM_PACKAGE_ARTIFACT_ID
152
184
  ```
153
185
 
154
- Repeat the independent installation check for the released version. A GitHub Release page is optional; pushing the version tag is the publication trigger.
186
+ This path validates that the artifact belongs to this repository’s original tag-triggered publishing workflow and matches the tag commit on `main`. It downloads those immutable bytes, checks their checksum and package metadata, and verifies npm availability. It cannot publish or alter npm tags and does not need npm OIDC permissions. Old runs may lack retained release notes; that does not prevent archive verification.
187
+
188
+ You can also download the original archive and its checksum and run the verifier independently from a current source checkout:
189
+
190
+ ```sh
191
+ node scripts/release-verify.mjs verify v0.4.0 \
192
+ --archive /path/to/claude-autorouter-0.4.0.tgz \
193
+ --report artifacts/release-verification-0.4.0.json \
194
+ --timeout-ms 900000
195
+ ```
155
196
 
156
- For local release diagnostics after the tag exists, `node scripts/release-check.mjs source v0.3.2` checks the tag, clean checkout, metadata, and ancestry. `node scripts/release-check.mjs archive v0.3.2` checks the candidate checksum and contents; `dist/` must contain only that version's archive and checksum, so retain older artifacts elsewhere first. These helpers are run automatically in the release workflow; the first untagged bootstrap uses the checks in step 1 instead.
197
+ Both files must retain their original names, and the archive must satisfy the public-package validation rules. This command performs only registry reads and a temporary isolated install; it neither publishes nor changes your installed CLI, npm login, or AutoRouter configuration. Exit `0` includes pending verification, so automation must inspect the report’s `state`; exit `1` means verification failed. A mismatch is never accepted as an already-published identical release.
157
198
 
158
- ## Recovering a failed release
199
+ ## Recovering a release
159
200
 
160
- - **Checks or archive validation failed:** nothing is published. Fix the cause and repeat validation before making a new release tag. Do not move a tag that already identifies a published version.
161
- - **npm rejected OIDC authentication:** verify the npm trust fields, direct-publish permission, GitHub-hosted runner, and the publish job's OIDC permission. After correcting npm configuration, rerun the failed job if that version is still unpublished.
162
- - **A publish timed out or the run was interrupted:** check `npm view` for the exact version before retrying. The registry may have accepted it before the connection failed.
163
- - **The version already exists:** inspect the registry release; do not overwrite or unpublish to reuse it. Code or documentation corrections need a new version.
164
- - **Local bootstrap authentication failed:** complete `npm login` and the account's 2FA flow in your terminal. CI trust cannot create the first package or substitute for that account step.
201
+ - **Checks/archive validation failed:** nothing was submitted by that failed path. Fix the cause, repeat validation, and follow the normal release process. Never move a published tag.
202
+ - **Registry preflight failed:** an HTTP/authentication/network failure is not evidence that the version is unused. Restore visibility before submitting.
203
+ - **npm rejected OIDC:** check the trusted-publisher fields, direct-publish permission, hosted runner, and OIDC permission. A job rerun deliberately does not resubmit a still-invisible version because an interrupted command may already have been accepted. Establish what happened before planning a new submission.
204
+ - **Submission timed out, the run was interrupted, or npm is processing it:** retain the original archive and use verification-only dispatch. A retry with a matching visible version skips `npm publish`; a retry with an invisible version only verifies it.
205
+ - **Existing version has different bytes:** stop. Never overwrite, unpublish, or replace the canonical artifact to reuse that version. Corrections require a new version.
206
+ - **Verification remains pending:** do not call it a failed publication. Rerun the verifier later with the original artifact, inspect npm’s public status if the registry is failing, and retain the evidence for support if validation remains stuck.
165
207
 
166
- If release code or workflow changes are needed, commit the fix to `main` and prepare a new version/tag. Authentication-only corrections on npm can be retried against the unchanged, unpublished candidate.
208
+ The workflow never repairs distribution tags automatically. npm’s [publish semantics](https://docs.npmjs.com/cli/commands/npm-publish/) make published name/version pairs immutable; [distribution tags](https://docs.npmjs.com/adding-dist-tags-to-packages/) are mutable references, so verification observes them without overwriting a newer release.