claude-autorouter 0.3.7 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (49) hide show
  1. package/.env.example +6 -3
  2. package/CODE_OF_CONDUCT.md +9 -0
  3. package/CONTRIBUTING.md +57 -0
  4. package/README.md +47 -70
  5. package/SECURITY.md +23 -0
  6. package/SUPPORT.md +18 -0
  7. package/bin/autorouter.mjs +40 -57
  8. package/docs/development.md +48 -2
  9. package/docs/hardware-benchmark.md +29 -0
  10. package/docs/hardware-comparison.md +55 -0
  11. package/docs/hardware-results-16gb.json +4002 -0
  12. package/docs/hardware-results-16gb.md +26 -0
  13. package/docs/hardware-results-64gb.json +4020 -0
  14. package/docs/reference.md +92 -37
  15. package/docs/releasing.md +79 -37
  16. package/docs/router-performance.json +1697 -0
  17. package/docs/router-performance.md +50 -0
  18. package/docs/status-performance.json +363 -0
  19. package/docs/status-performance.md +44 -0
  20. package/docs/subscription-integration.md +27 -0
  21. package/package.json +66 -10
  22. package/src/auto-routing.mjs +184 -24
  23. package/src/bounded-json.mjs +57 -0
  24. package/src/cli-help.mjs +90 -0
  25. package/src/config-command.mjs +158 -0
  26. package/src/config.mjs +53 -28
  27. package/src/contracts.mjs +123 -0
  28. package/src/evaluation-report.mjs +114 -0
  29. package/src/keychain.mjs +58 -0
  30. package/src/local-diagnostic.mjs +191 -0
  31. package/src/model-catalog.mjs +96 -0
  32. package/src/model-request.mjs +6 -7
  33. package/src/ollama-evaluator.mjs +9 -27
  34. package/src/onboarding.mjs +130 -26
  35. package/src/prompt-state.mjs +22 -7
  36. package/src/redaction.mjs +97 -0
  37. package/src/request-validation.mjs +54 -0
  38. package/src/response-observer.mjs +126 -18
  39. package/src/router.mjs +151 -61
  40. package/src/savings.mjs +74 -16
  41. package/src/server.mjs +79 -12
  42. package/src/session-history.mjs +262 -0
  43. package/src/session-log.mjs +9 -58
  44. package/src/status-state.mjs +110 -62
  45. package/src/statusline.mjs +57 -27
  46. package/src/telemetry-event.mjs +200 -0
  47. package/src/token-counter.mjs +3 -1
  48. package/src/turn-state.mjs +132 -0
  49. package/src/user-config.mjs +81 -10
package/docs/reference.md CHANGED
@@ -1,5 +1,7 @@
1
1
  # Reference
2
2
 
3
+ AutoRouter is an independent gateway for Claude Code. Its existing `claude-autorouter` package, command, configuration paths, and repository identity remain unchanged. See [integration boundaries and provider policy](subscription-integration.md) for authentication ownership and the distinction between technical operation and provider authorization.
4
+
3
5
  ## Commands
4
6
 
5
7
  | Command | Purpose |
@@ -8,10 +10,17 @@
8
10
  | `claude-autorouter setup --auth-mode api-key` | Configure Jev and Anthropic API-key billing |
9
11
  | `claude-autorouter setup --client-profile auto` | Save a Sonnet/Opus profile compatible with Claude's Auto permission mode |
10
12
  | `claude-autorouter setup --stop-hook-block-cap 2` | Opt into a shorter native Stop-hook continuation cap during setup |
11
- | `claude-autorouter setup --session-log-dir DIR` | Save an opt-in directory for per-session JSONL decision logs |
13
+ | `claude-autorouter setup --session-log-dir DIR` | Save an opt-in directory for per-session JSONL decisions and outcomes |
12
14
  | `claude-autorouter setup --evaluator ollama --pull` | Configure the native local evaluator and download its selected model if missing |
13
15
  | `claude-autorouter setup --evaluator ollama --ollama-timeout-ms 0 --force` | Save a disabled runtime evaluator deadline |
14
- | `claude-autorouter setup --force` | Replace an existing user config |
16
+ | `claude-autorouter setup --force` | Update an existing config while preserving unrelated saved settings |
17
+ | `claude-autorouter setup --replace` | Explicitly replace the saved configuration |
18
+ | `claude-autorouter config show --json` | Inspect effective settings and default/file/environment provenance; secrets are hidden |
19
+ | `claude-autorouter config set KEY VALUE` | Change one nonsecret saved setting |
20
+ | `claude-autorouter config unset KEY` | Remove one saved override |
21
+ | `claude-autorouter sessions list [--json]` | Inspect optional saved local history |
22
+ | `claude-autorouter sessions show ID [--json]` | Show correlated decisions, outcomes and pricing coverage |
23
+ | `claude-autorouter doctor --evaluate-local` | Run synthetic classifier checks on the installed local Ollama model |
15
24
  | `claude-autorouter doctor` | Check config, Claude executable/login, and the selected local Ollama model without paid calls |
16
25
  | `claude-autorouter claude [arguments]` | Start a local router and pass arguments through to Claude Code |
17
26
  | `claude-autorouter serve` | Run the router for separately configured clients |
@@ -22,7 +31,7 @@ The launcher binds an ephemeral port on `127.0.0.1`, creates a temporary local c
22
31
 
23
32
  ## Configuration
24
33
 
25
- Setup defaults to subscription mode unless `--auth-mode` or `AUTOROUTER_AUTH_MODE` selects another mode. Jev remains the default evaluator; `--evaluator ollama` selects local classification. Setup prompts for required secrets without echoing them and writes a private JSON file. Stored keys are plaintext; keep the file private and out of source control. Supply keys through the environment when interactive input is unavailable. Subscription mode with Ollama requires no API keys. API-key authentication always requires `ANTHROPIC_API_KEY`, regardless of evaluator.
34
+ First setup defaults to subscription mode unless `--auth-mode` or `AUTOROUTER_AUTH_MODE` selects another mode. Local Ollama is the default evaluator (changed from Jev in 0.5.0; configurations created by `setup` always record their evaluator, so existing ones are unchanged, but an environment-only launch with no `AUTOROUTER_EVALUATOR` now selects Ollama); `--evaluator jev` selects TypeSafe's hosted evaluator. An existing configuration updated with `--force` keeps its saved choices unless a command-line flag changes them; unrelated environment overrides remain temporary. Setup prompts for required secrets without echoing them and writes a private JSON file. On macOS, new and `--replace` setups keep keys in the login Keychain by default (see [credential storage](#credential-storage)); elsewhere, and in existing configurations, stored keys are plaintext in that file, so keep it private and out of source control. Supply keys through the environment when interactive input is unavailable. Subscription mode with Ollama requires no API keys. API-key authentication always requires `ANTHROPIC_API_KEY`, regardless of evaluator.
26
35
 
27
36
  The config path is selected in this order:
28
37
 
@@ -30,12 +39,23 @@ The config path is selected in this order:
30
39
  2. `$XDG_CONFIG_HOME/claude-autorouter/config.json`, when `XDG_CONFIG_HOME` is a nonempty absolute path.
31
40
  3. `~/.config/claude-autorouter/config.json`.
32
41
 
33
- The JSON file uses flat environment-style string keys, such as `AUTOROUTER_AUTH_MODE` and `TYPESAFE_API_KEY`. Environment values take precedence over the saved config. Use `setup --force` to replace existing configuration. The launcher does not discover or load a project's `.env` file. From a source checkout, explicitly loading one still works:
42
+ The JSON file uses flat environment-style string keys, such as `AUTOROUTER_AUTH_MODE` and `TYPESAFE_API_KEY`. Environment values take precedence over the saved config. Use `config set` or `config unset` to edit a single saved setting, or `setup --force` to update an existing configuration while retaining unrelated settings. `setup --replace` explicitly rebuilds a readable supported saved configuration; malformed or unsupported files are left intact for manual repair. The launcher does not discover or load a project's `.env` file. From a source checkout, explicitly loading one still works:
34
43
 
35
44
  ```sh
36
45
  node --env-file=.env bin/autorouter.mjs claude
37
46
  ```
38
47
 
48
+ ### Credential storage
49
+
50
+ `AUTOROUTER_SECRET_STORE` selects where saved `ANTHROPIC_API_KEY`, `TYPESAFE_API_KEY` and `AUTOROUTER_TOKEN` values live. `file` stores them in the private JSON file; it is used on Linux and by configurations created before this option existed. `keychain`, available on macOS and chosen by default there for new setups, stores each one as a generic password in the login Keychain (service `claude-autorouter`, scoped to the configuration file path); the JSON file then contains only settings. Changing the setting moves saved keys:
51
+
52
+ ```sh
53
+ claude-autorouter config set AUTOROUTER_SECRET_STORE keychain # file → Keychain
54
+ claude-autorouter config set AUTOROUTER_SECRET_STORE file # Keychain → file
55
+ ```
56
+
57
+ If the Keychain cannot be used during a default setup (locked or headless), setup says so and saves keys to the file instead; an explicit `--secret-store keychain` fails rather than falling back. An existing plaintext configuration is never moved implicitly: `setup --force` and `doctor` flag plaintext keys on macOS and print the command above. Keys are written to the Keychain before the file is rewritten, so an interrupted move leaves a copy in both places rather than neither. Items are removed only after the file is saved. Values pass to the system `security` tool on standard input, never as process arguments, and each write is read back to confirm it. Keychain values must be printable single-line ASCII. Environment variables still take precedence: when the Keychain is locked (for example over SSH), a key supplied in the environment is used instead, but moving keys between stores waits until every saved key can be read. `sessions` never reads the Keychain.
58
+
39
59
  For an environment-only subscription launch, set `AUTOROUTER_AUTH_MODE=subscription` and either supply `TYPESAFE_API_KEY` or select `AUTOROUTER_EVALUATOR=ollama` with a running local model. For API-key mode, also supply `ANTHROPIC_API_KEY`. The shell variables are read by the router; Jev's key is removed from the Claude child environment.
40
60
 
41
61
  | Variable | Default | Purpose |
@@ -44,11 +64,13 @@ For an environment-only subscription launch, set `AUTOROUTER_AUTH_MODE=subscript
44
64
  | `ANTHROPIC_API_KEY` | required in API-key mode | Upstream Anthropic credential |
45
65
  | `AUTOROUTER_CONFIG` | see path order above | Explicit user config path |
46
66
  | `AUTOROUTER_AUTH_MODE` | `api-key` without saved config; setup selects `subscription` | Authentication mode |
47
- | `AUTOROUTER_EVALUATOR` | `jev` | `jev` or local `ollama` classification |
67
+ | `AUTOROUTER_EVALUATOR` | `ollama` | local `ollama` (default) or hosted `jev` classification |
68
+ | `AUTOROUTER_SECRET_STORE` | `keychain` for new macOS setups; otherwise `file` | Saved-config setting: `keychain` keeps saved keys in the macOS login Keychain; the environment cannot redirect it |
48
69
  | `AUTOROUTER_CLIENT_PROFILE` | `compatible` | `native` retains client model/thinking settings; `auto` starts with Sonnet when no explicit model is set and excludes Haiku from task routing |
49
70
  | `AUTOROUTER_STATUSLINE` | enabled | `0` retains your existing status line |
50
71
  | `AUTOROUTER_DEBUG` | off | `1` enables launcher metadata logs on stderr |
51
- | `AUTOROUTER_SESSION_LOG_DIR` | off | Write per-session JSONL decision logs with prompt excerpts into this directory; unset or empty disables it |
72
+ | `AUTOROUTER_SESSION_LOG_DIR` | off | Write per-session JSONL decisions and outcomes into this directory; unset or empty disables it |
73
+ | `AUTOROUTER_SESSION_LOG_MODE` | `prompts` | `metadata` omits prompt excerpts; setting a mode alone does not enable logging |
52
74
  | `CLAUDE_CODE_STOP_HOOK_BLOCK_CAP` | unset; Claude currently uses `8` | Optional cap on consecutive Stop/SubagentStop continuations without tool use; `0` disables the cap |
53
75
  | `ENABLE_TOOL_SEARCH` | `true` in launcher when unset | Load MCP tool definitions on demand; explicit values are preserved |
54
76
  | `AUTOROUTER_HAIKU_MODEL` | `claude-haiku-4-5-20251001` | Routine tier |
@@ -69,9 +91,23 @@ For an environment-only subscription launch, set `AUTOROUTER_AUTH_MODE=subscript
69
91
 
70
92
  Model access depends on your account. The policy recognizes specific Claude model versions; arbitrary gateway aliases do not automatically inherit their capabilities or context windows. Compare overrides with the [Anthropic model catalog](https://platform.claude.com/docs/en/models/overview).
71
93
 
94
+ ### Inspect and change settings
95
+
96
+ ```sh
97
+ claude-autorouter config show
98
+ claude-autorouter config show --json --check-all
99
+ claude-autorouter config set AUTOROUTER_OLLAMA_TIMEOUT_MS 0
100
+ claude-autorouter config unset AUTOROUTER_OLLAMA_TIMEOUT_MS
101
+ claude-autorouter config set TYPESAFE_API_KEY
102
+ ```
103
+
104
+ `show` reports whether each setting comes from a default, the saved file, or an environment override. Secret values are never displayed. The last command uses a hidden prompt; scripts can pipe a secret to `config set TYPESAFE_API_KEY --stdin`. Secret values are not accepted as command arguments. Edits validate and atomically update only the named saved setting. Environment overrides still apply after a saved change. An unset saved deadline returns to the model default unless an environment value overrides it.
105
+
106
+ Normal startup validates the selected evaluator; stale settings for the inactive evaluator do not prevent it from starting. `show --check-all` explicitly checks both. Blank numeric settings fail with their setting name; zero retains its documented meaning. `claude-autorouter help COMMAND` gives focused command help. `claude-autorouter claude --help` and `--version` call Claude directly without router setup or credentials.
107
+
72
108
  ## Ollama evaluator
73
109
 
74
- The local configuration documented here requires AutoRouter 0.3.2 or newer and remains experimental. It uses Ollama's native `/v1/systemone` decision endpoint for every model, replacing the chat backend from 0.2.0. Jev remains the default remote evaluator, using TypeSafe's `/v1/systemone` endpoint and a TypeSafe API key. Selecting Ollama never silently switches back to Jev. Haiku, Sonnet, or Opus still completes the task through Anthropic.
110
+ The local configuration documented here requires AutoRouter 0.3.2 or newer and remains experimental. It uses Ollama's native `/v1/systemone` decision endpoint for every model, replacing the chat backend from 0.2.0. Jev is the optional hosted evaluator, using TypeSafe's `/v1/systemone` endpoint and a TypeSafe API key. Selecting Ollama never silently switches back to Jev. Haiku, Sonnet, or Opus still completes the task through Anthropic.
75
111
 
76
112
  Version 0.3.2 excludes Claude's executor system instructions from local excerpts, uses model-specific runtime deadlines, and accepts `0` to disable that deadline. Setup, doctor, and startup show the effective model and deadline; setup accepts `--ollama-timeout-ms`. Jev is unchanged.
77
113
 
@@ -83,7 +119,7 @@ claude-autorouter doctor
83
119
  claude-autorouter claude
84
120
  ```
85
121
 
86
- `--force` replaces an existing user config. Setup detects the running local API. `--pull` authorizes downloading the chosen model when it is missing; without it, install the model yourself before setup. AutoRouter does not install Ollama, start its daemon, delete models, or download models during ordinary launches or `doctor` checks.
122
+ `--force` updates an existing user config and preserves its other settings and selected model unless explicitly changed. Setup detects the running local API. `--pull` authorizes downloading the chosen model when it is missing; without it, install the model yourself before setup. AutoRouter does not install Ollama, start its daemon, delete models, or download models during ordinary launches or `doctor` checks.
87
123
 
88
124
  ### Local model selection
89
125
 
@@ -115,7 +151,7 @@ The endpoint must be loopback (`127.0.0.1`, `localhost`, or `::1`), without a pa
115
151
 
116
152
  Local classification caps serialized evaluator state at both 3,000 characters and 3,000 UTF-8 bytes, including for non-ASCII prompts. Claude's top-level executor system instructions are excluded before budgeting; the current task, original task, and recent conversation excerpts remain. `/v1/systemone` receives the bounded state and routing criteria and returns a tier directly. The router retains each model's native context setting: 8,194 tokens for the default Nimble tag and 2,050 for the listed Tev1 tags. Tev1's smaller window includes the routing criteria and template as well as the excerpt; the byte limit does not guarantee every possible input fits. Context errors use the normal fallback. Returned confidence scores summarize choice-distribution entropy; they are not calibrated accuracy probabilities. `AUTOROUTER_MIN_CONFIDENCE` applies only to Jev. All capability, tool-continuation, thinking, and context guards still apply.
117
153
 
118
- The deadline covering local checks and classification defaults to 1,500 ms for Tev1 0.8B and custom/unrecognized tags, 15,000 ms for official Tev1 4B variants (including bare `tev1` and `latest`), and 30,000 ms for official Nimble variants. Official `library/` and `registry.ollama.ai/` aliases are recognized; a custom namespace such as `team/nimble` keeps the short default. An explicit timeout overrides the model default, including an old saved `1500`. Environment values override saved values on launch. Defaults are not written into the user config; `setup --force` replaces the config and saves an explicit timeout when supplied through `--ollama-timeout-ms` or the environment. During setup, the command-line flag takes precedence over the timeout environment value.
154
+ The deadline covering local checks and classification defaults to 1,500 ms for Tev1 0.8B and custom/unrecognized tags, 15,000 ms for official Tev1 4B variants (including bare `tev1` and `latest`), and 30,000 ms for official Nimble variants. Official `library/` and `registry.ollama.ai/` aliases are recognized; a custom namespace such as `team/nimble` keeps the short default. An explicit timeout overrides the model default, including an old saved `1500`. Environment values override saved values on launch. Defaults are not written into the user config. To update a saved deadline, use `config set AUTOROUTER_OLLAMA_TIMEOUT_MS N` or `setup --ollama-timeout-ms N --force`; unrelated environment overrides remain temporary. Environment-provided Ollama settings are saved during first setup, replacement, or explicit `--evaluator ollama` selection. The command-line timeout flag takes precedence over the environment.
119
155
 
120
156
  Set `AUTOROUTER_OLLAMA_TIMEOUT_MS=0` to remove AutoRouter's runtime evaluator timer while keeping the existing configuration:
121
157
 
@@ -137,11 +173,24 @@ If startup priming fails, the launcher warns and continues. An incompatible mode
137
173
 
138
174
  Historical measurements before 0.3.2: Tev1 4B timed out on all eight full-excerpt checks even with a 10-second diagnostic allowance; its short-task results did not establish a full-excerpt latency bound. Tev1 0.8B completed all eight within 1,500 ms. On the tested 16 GiB M4, Nimble timed out on all 12 tuning requests at 1,500 ms. A separate 30-second diagnostic completed 24 held-out classifications with 23 matching labels, but median routing took 11.4 seconds. The one error followed a misleading tier instruction. These historical results precede the 0.3.2 excerpt changes and do not establish guarantees for the longer defaults. See the [measurements and limitations](ollama-evaluation.md).
139
175
 
176
+ ### Test the installed local evaluator
177
+
178
+ ```sh
179
+ claude-autorouter doctor --evaluate-local
180
+ claude-autorouter doctor --evaluate-local --json
181
+ ```
182
+
183
+ This explicit diagnostic requires an installed local evaluator selected with `AUTOROUTER_EVALUATOR=ollama`. It uses synthetic tasks only, needs no Jev or Anthropic key, and makes no Claude inference calls. It checks local availability, measures a separate initial preparation call, then runs six uncached cases through the production classifier using the configured runtime deadline. A disabled runtime deadline remains disabled; Ctrl-C cancels the diagnostic.
184
+
185
+ The report separates availability, expected-label agreement, all-three-tier coverage, and latency. Residency is observed before calls, so it does not claim a controlled cold/warm benchmark. A pass establishes these six examples only. Incorrect predictions, fallback, or missing Haiku/Opus coverage fail even if Ollama answered successfully. Auto mode still checks all three raw evaluator labels; actual Auto routing applies its Sonnet floor separately.
186
+
187
+ The diagnostic does not download, unload, restart, or edit configuration. It refuses to run while unrelated models are resident and requires a positive keep-alive; normal routing still supports keep-alive `0`. Ordinary `doctor` remains a metadata check. A missing-model repair command uses the exact configured tag and preserves other settings.
188
+
140
189
  ### Migrating an older Ollama config
141
190
 
142
191
  Version 0.3.1 used a 1,500 ms deadline for every local model. After upgrading to 0.3.2, an explicitly saved or exported `AUTOROUTER_OLLAMA_TIMEOUT_MS=1500` still wins over the new model-specific defaults. Remove that override to use the defaults, or rerun setup with the desired model and `--ollama-timeout-ms N --force`. The `0` value and setup timeout flag require 0.3.2 or newer.
143
192
 
144
- Version 0.3.1 removed the Qwen chat backend and presets from 0.2.0. Existing downloaded models remain on disk, but an old Qwen model selection needs to be replaced with a native decision model. Run the setup command above with `--force`; it selects Nimble unless you pass `--ollama-model` or override the model through the environment. Remove or update any old `AUTOROUTER_OLLAMA_MODEL` environment value too, because environment variables override saved configuration. Update scripts to use `--ollama-model` when selecting a custom model.
193
+ Version 0.3.1 removed the Qwen chat backend and presets from 0.2.0. Existing downloaded models remain on disk, but an old Qwen model selection needs to be replaced with a native decision model. Run `claude-autorouter setup --evaluator ollama --ollama-model nimble:9b-q4_K_M --force`, or explicitly choose a Tev1 tag; merging with `--force` alone preserves the saved model. Remove or update any old `AUTOROUTER_OLLAMA_MODEL` environment value too, because environment variables override saved configuration. Update scripts to use `--ollama-model` when selecting a custom model.
145
194
 
146
195
  ## Data flow and authentication
147
196
 
@@ -151,19 +200,21 @@ Claude Code → authenticated local gateway → Jev or local Ollama classificati
151
200
  → selected Claude model → streamed response
152
201
  ```
153
202
 
154
- AutoRouter uses Claude Code's [gateway integration](https://code.claude.com/docs/en/llm-gateway-protocol), so it sees inference requests and tool continuations. It does not rely on a user-prompt hook.
203
+ AutoRouter launches the user's installed official Claude Code binary without patching it and uses Claude Code's [gateway integration](https://code.claude.com/docs/en/llm-gateway-protocol), so it sees inference requests and tool continuations. It does not rely on a user-prompt hook. Each user uses their own provider credentials; AutoRouter does not provide a Claude sign-in service or a shared provider account.
155
204
 
156
- The selected evaluator receives a bounded state containing the latest human request and excerpts of the original task and recent messages: up to 12,000 serialized characters sent to TypeSafe for Jev, or 3,000 UTF-8 bytes sent to the local Ollama service. Jev also receives system-text excerpts. The local path excludes Claude's top-level executor system instructions. These excerpts can include private source code and tool results. Images, document payloads, and signed thinking are omitted. Full tool schemas and full conversation history are not sent to either classifier. Anthropic receives the complete request, including its tools and attachments. Large or multimodal requests may also go to Anthropic's token-count endpoint before inference, including when classification is local.
205
+ The selected evaluator receives a bounded state containing the latest human request and excerpts of the original task and recent messages: up to 12,000 serialized characters sent to TypeSafe for Jev, or 3,000 UTF-8 bytes sent to the local Ollama service. Jev also receives system-text excerpts. The local path excludes Claude's top-level executor system instructions. These excerpts can include private source code and tool results. Before excerpting, recognizable sensitive values are replaced (the same filter applies to opt-in session-log prompt excerpts, including when old logs are read back) with markers such as `[REDACTED:secret]`: private keys, common provider token formats (Anthropic, OpenAI-style `sk-`, AWS, GitHub, GitLab, Slack, Google, Stripe, npm), JWTs, authorization headers, URL credentials, values assigned to password/secret/token/key-like names, email addresses, and checksum-valid IBANs and payment card numbers. Setting names stay visible. Redaction is pattern-based: unrecognized formats can remain, code resembling an assignment can be over-redacted, and it does not make arbitrary private source code safe to share. Images, document payloads, and signed thinking are omitted. Full tool schemas and full conversation history are not sent to either classifier. Anthropic receives the complete request, including its tools and attachments. Large or multimodal requests may also go to Anthropic's token-count endpoint before inference, including when classification is local.
157
206
 
158
207
  In subscription mode, Claude Code owns login and OAuth refresh. AutoRouter forwards the current request's authorization and beta headers to Anthropic. It does not read keychain or saved login files, persist subscription tokens, or send them to Jev. A separate temporary `X-Autorouter-Token` authenticates the local connection and is stripped upstream. Subscription forwarding is restricted to `https://api.anthropic.com`. See [subscriptions and gateways](https://code.claude.com/docs/en/llm-gateway#subscriptions-and-gateways).
159
208
 
160
209
  In API-key mode, the upstream key stays in the proxy and Claude receives a temporary local credential. Requests are billed to the supplied API key. Subscription requests remain subject to the subscription's model access and usage limits. AutoRouter never falls back from subscription authentication to API billing.
161
210
 
162
- Routine logs contain route, model, timing, usage, and error-category metadata, not prompts, raw responses, or credentials. Status snapshots contain routing metadata and token counts in a private temporary directory and are deleted on normal launcher exit. Classification, turn, and token-count caches are held in memory. Claude Code and the external providers have their own storage and logging behavior.
211
+ The proxy processes authenticated requests in memory, including their authorization headers. Preserving Claude's login flow does not by itself establish that every deployment is permitted. The [provider-policy note](subscription-integration.md#provider-guidance-and-unresolved-scope) records the current documentation and the unresolved scope of model-rewriting subscription forwarding. Jev requires its own TypeSafe credentials and billing, separate from Anthropic authentication.
212
+
213
+ Routine diagnostic logs contain route, model, timing, usage, and error-category metadata, not prompts, raw responses, or credentials. Opt-in session history is separate and includes task excerpts in its default `prompts` mode. Status snapshots contain routing metadata and token counts in a private temporary directory and are deleted on normal launcher exit. Classification, turn, and token-count caches are held in memory. Claude Code and the external providers have their own storage and logging behavior.
163
214
 
164
215
  ## Routing policy
165
216
 
166
- Each `/v1/messages` request is evaluated. Exact repeated bodies reuse a classification for five minutes. Both evaluators use a starting rubric choosing Haiku for routine work, Sonnet for ordinary engineering, and Opus for demanding reasoning. These choices require evaluation on your tasks; they are not quality guarantees.
217
+ Each eligible `/v1/messages` request is evaluated. Exact repeated bodies reuse a classification for five minutes; concurrent identical evaluations share one request. Cache identity includes evaluator configuration, rubric and requested model floor. Internal permission classifiers and other documented pass-through paths skip evaluation. Both evaluators use a starting rubric choosing Haiku for routine work, Sonnet for ordinary engineering, and Opus for demanding reasoning. These choices require evaluation on your tasks; they are not quality guarantees.
167
218
 
168
219
  The evaluator prioritizes the actual human request before startup metadata. Complete Claude reminder and tool-list blocks are excluded from that task excerpt, and long text retains its beginning and end. The outbound Anthropic request remains complete. Complexity outside the bounded excerpt can still be missed.
169
220
 
@@ -171,7 +222,7 @@ The following policy applies after classification:
171
222
 
172
223
  - Jev's 1,500 ms deadline covers the response body and has no retry. Successful calls return immediately. Timeouts, HTTP errors, and invalid responses fall back to Sonnet or retain an existing stronger model.
173
224
  - Jev confidence below 0.75 prevents a downgrade below Sonnet or the requested tier. Ollama returns a tier without calibrated confidence; its failure handling and compatibility guards still apply.
174
- - Tool continuations retain the model chosen at the start of the human turn. Session, agent, and prompt headers identify turns; normalized conversation content provides a fallback. Text feedback from a Stop hook also retains the model when it serves the same gateway prompt ID and the client has not changed its requested model, subject to capability and context checks. Moving prompt-cache markers does not create a new turn.
225
+ - Tool continuations retain the execution model confirmed by a successfully forwarded response. Active tasks and pending tools survive classification-cache expiry; retired task records expire separately. After restart, missing continuity is explicitly unknown. A selected model alone remains unconfirmed. Session, agent, and prompt headers identify turns; normalized conversation content provides a fallback. Text feedback from a Stop hook also retains the model when it serves the same gateway prompt ID and the client has not changed its requested model, subject to capability and context checks. Moving prompt-cache markers does not create a new turn.
175
226
  - Claude's local `/goal` command can omit the prompt-ID header. For that path, an exact feedback label matching a preceding expanded `/goal` command keeps the original task and conversation anchor. This narrow text fallback also recognizes Claude's repeated-goal truncation format; arbitrary hook text is not treated as a goal. Feedback remains in the evaluator's recent conversation and the full API request. A new human message becomes the current task normally. The status line shows `prompt pinned` or `goal pinned` when either text-continuation rule applies.
176
227
  - Thinking history, fixed-budget thinking, server tools, context management, and other recognized model-specific features preserve the current model except for the verified shared capabilities of the modern Auto-mode Sonnet/Opus pair described below. Adaptive thinking, effort, and output above 64K prevent a Haiku choice. Fields are never stripped to force a downgrade.
177
228
  - Mid-conversation `system` messages preserve the requested model unless both Auto-mode models support them; they always pass through unchanged. They do not count as a tool continuation by themselves.
@@ -201,7 +252,7 @@ env AUTOROUTER_CLIENT_PROFILE=auto claude-autorouter claude
201
252
 
202
253
  The profile defaults to Sonnet 5.5 and Opus 5.5, preserving explicit configured model IDs. The evaluator chooses Sonnet or Opus for each new human task; a routine Haiku verdict uses Sonnet and shows `Auto mode floor`. Tool and `/goal` continuations stay on the selected execution model. Claude's initial client model remains separate from the routed model. An explicit client `--model` or `ANTHROPIC_MODEL` can still make Auto unavailable if it selects Haiku or another unsupported model; choose a supported Sonnet or Opus instead.
203
254
 
204
- Claude remains responsible for enabling the permission mode and enforcing organization settings, account availability, and tool rules. The profile does not enable Auto by itself or override `disableAutoMode`. AutoRouter does not reproduce Claude's settings precedence to infer a mode from settings files. For new saved configurations, `setup --client-profile auto` persists the profile; for an existing config, change only `AUTOROUTER_CLIENT_PROFILE` to `"auto"` to retain your other settings. `setup --force` replaces the config.
255
+ Claude remains responsible for enabling the permission mode and enforcing organization settings, account availability, and tool rules. The profile does not enable Auto by itself or override `disableAutoMode`. AutoRouter does not reproduce Claude's settings precedence to infer a mode from settings files. `setup --client-profile auto --force` updates an existing configuration while retaining its other settings. `config set AUTOROUTER_CLIENT_PROFILE auto` changes just that setting.
205
256
 
206
257
  **Safety review:** Claude's permission-classifier requests retain their exact requested model and skip AutoRouter's evaluator. Ordinary execution requests with the known `dangerous_tool_use` version-1 review contract are evaluated and routed, retaining the complete `safeguards` object, beta headers, and streamed safety verdicts. This also detects server review when Auto was selected in Claude's UI rather than through the launch flag. Unknown or malformed review contracts and safeguarded compaction pass through with `Auto safety`; a target that cannot accept the request shows `Auto model guard`. AutoRouter never turns off server review or converts denied actions to approvals. See [server-side classifier review](https://code.claude.com/docs/en/permission-modes#server-side-classifier-review).
207
258
 
@@ -227,7 +278,7 @@ The `claude` launcher automatically adds a temporary [status-line command](https
227
278
  | --- | --- |
228
279
  | `Sonnet 5 selected` | Routing chose this model; Anthropic has not confirmed it yet |
229
280
  | `Opus 5.5` / `last Opus 5.5` | Provider-confirmed streaming or most recent model |
230
- | `Jev` / `Ollama`, with `cache` or `fallback` when applicable | Classification source; timing includes concurrent context checks |
281
+ | `Jev` / `Ollama`, with `cache` or `fallback` when applicable | Classification source; displayed routing time includes evaluator waiting and context checks |
231
282
  | `Jev→Haiku` or `Ollama→Haiku` beside Sonnet | A policy guard overrode the evaluator's Haiku choice |
232
283
  | `large context` | Token count exceeded the small-model input budget |
233
284
  | `size unverified` | Token checking failed or was unavailable; conservative guard applied |
@@ -235,51 +286,55 @@ The `claude` launcher automatically adds a temporary [status-line command](https
235
286
  | `CLI ctx` | Claude's client accounting, shown when its window differs or API capacity is unknown |
236
287
  | `est saved … vs Opus` | Cumulative API-equivalent token-cost estimate |
237
288
 
238
- Background agents and auxiliary requests cannot replace the foreground model. Errors, fallback, cancellation, and stale/offline state remain visible. The command reads a local snapshot and makes no network requests. It respects terminal width and `NO_COLOR`; lower-priority fields disappear on narrow terminals.
289
+ Background agents and auxiliary requests cannot replace the foreground model. Errors, fallback, cancellation, incomplete-response evidence and stale/offline state remain visible. The command reads a local snapshot and makes no network requests. It respects terminal width and `NO_COLOR`; lower-priority fields disappear on narrow terminals.
239
290
 
240
291
  Context includes uncached input, cache reads, and cache writes, excluding output to match [Claude's percentage formula](https://code.claude.com/docs/en/statusline#context-window-fields). It is current context rather than cumulative usage. Historical usage is marked `last`; compaction resets stale readings. The compatible client's 200K window can reach 100% while a routed Sonnet request uses only part of its 1M window. Displaying both does not change Claude's compaction threshold.
241
292
 
242
293
  Savings compare the actual models' API token prices with the configured Opus model's prices for the **same reported counts and cache profile**. The percentage is `(Opus cost − routed cost) / Opus cost`. Input, output, cache reads, and 5-minute/1-hour cache writes are priced separately using the bundled table based on [Anthropic's published USD pricing](https://platform.claude.com/docs/en/about-claude/pricing).
243
294
 
244
- Totals include completed main, agent, and auxiliary calls for the current session observed by this router process. Streaming usage is counted once, and totals reset when the router or session restarts. Unrecognized prices or unsupported usage produce `partial` or `savings unavailable`. Higher routed costs show `est extra`. The arithmetic runs locally.
295
+ Totals include completed main, agent, and auxiliary calls for the current session observed by this router process. Streaming usage is counted once, and totals reset when the router or session restarts. Unrecognized prices or unsupported usage produce `partial`, `unpriced N` or `savings unavailable`. Saved history identifies pricing table `2026-09-29.1` (reviewed September 29, 2026) and unpriced reason counts; unknown historical table versions are not repriced. Higher routed costs show `est extra`. The arithmetic runs locally.
245
296
 
246
297
  This estimate does not measure subscription bill savings or quota credits. It excludes Jev charges, local compute costs, tool fees, negotiated discounts, and unpriced requests. A real Opus run can produce different tokens and cache hits. Incomplete streams, unknown cache-write TTLs, unsupported pricing modifiers, and unrecognized model versions are excluded rather than guessed. The rate table requires updates when prices change.
247
298
 
299
+ Historical integration observations cover Claude Code 2.1.284 and 2.1.285. The source-only versioned corpus in `test/fixtures/claude-protocol-v1.json` separates newly authored synthetic contracts from those dated observations and their artifact hashes. Gateway tests exercise Auto safeguards, thinking, deferred tools, compaction, goal scoping, fallback ownership and usage while preserving response bytes. They do not establish current-source live compatibility, evaluator accuracy or downstream task quality. Doctor reports the installed executable version separately; discovering a binary is not evidence that its protocol or Auto eligibility has been tested. Opt-in live validation supplements the synthetic fixtures.
300
+
248
301
  Set `AUTOROUTER_STATUSLINE=0` to retain an existing status line. Other `--settings` values are retained in the temporary overlay; source-relative Read/Edit rules keep their anchors. Ambiguous relative sandbox paths cause the launcher to skip the overlay and pass original settings through with a notice. Safe mode disables custom status lines; print mode has no status-line UI. Standalone `serve` does not install one.
249
302
 
250
303
  ## Session decision logs
251
304
 
252
- AutoRouter 0.3.6 adds optional persistent logs, separate from stderr and the temporary status-line snapshot. Logging is disabled by default. Enable it for one launch:
305
+ Logging is optional and disabled by default. Enable it for one launch:
253
306
 
254
307
  ```sh
255
308
  env AUTOROUTER_SESSION_LOG_DIR="$HOME/.local/state/claude-autorouter/sessions" \
256
309
  claude-autorouter claude
257
310
  ```
258
311
 
259
- The setting also works with `serve` and the Auto-compatible profile. For a new saved configuration, add `--session-log-dir DIR` to `setup`. For an existing config, add `AUTOROUTER_SESSION_LOG_DIR` with an absolute directory path to preserve your other settings. Setup resolves relative paths at setup time; an environment-only relative path resolves from the launch directory. Environment values override saved values; `AUTOROUTER_SESSION_LOG_DIR=''` disables a saved preference for one launch. `doctor` reports the setting without creating log files.
312
+ Or save metadata-only history, without prompt excerpts:
260
313
 
261
- Files are named `autorouter-session-*.jsonl`: one file per observed Claude session within a router launch, with a timestamp, random launch identifier, and hashed session identifier in the name. A resumed session in a new launch creates a new file. Requests without a session header share an anonymous file for that launch. Subagents with the same session ID share its file and retain their agent ID. Files are created only when a decision is recorded, and remain after the session ends.
314
+ ```sh
315
+ claude-autorouter config set AUTOROUTER_SESSION_LOG_MODE metadata
316
+ claude-autorouter config set AUTOROUTER_SESSION_LOG_DIR "$HOME/.local/state/claude-autorouter/sessions"
317
+ claude-autorouter sessions list
318
+ claude-autorouter sessions show autorouter-session-EXAMPLE
319
+ claude-autorouter sessions show autorouter-session-EXAMPLE --json
320
+ ```
262
321
 
263
- Every line is a standalone JSON object. The key fields look like this (additional IDs and routing metadata are included):
322
+ Use the exact `id` printed by `sessions list`. Commands need no evaluator credentials and do not contact providers. Human summaries distinguish selected models, observed serving models, confirmed completions, failures, cancellations, and pending/unconfirmed requests. They report routing latency, fallback and override counts, and API-equivalent savings coverage. A selected model or an HTTP 200 alone does not prove successful inference. Old schema-1 decision logs remain readable and explicitly lack outcome evidence.
264
323
 
265
- ```json
266
- {"schema_version":1,"event":"decision","timestamp":"2026-09-30T12:00:00.000Z","session_id":"example-session","prompt_excerpt":"Fix the typo in README.md","prompt_truncated":false,"requested_model":"claude-haiku-4-5-20251001","selected_model":"claude-sonnet-5","decision_latency_ms":214.37,"source":"jev","reason":"classified"}
267
- ```
324
+ `prompts` mode preserves the existing excerpt behavior when a log directory is enabled. Main requests retain at most 500 Unicode characters of the human task; recognized tool and goal continuations retain the originating task. Auxiliary, subagent, compaction, workflow, and attachment-only requests have empty excerpts. Metadata mode omits the excerpt fields entirely. Neither mode logs authentication headers, provider replies, full transcripts, or tool payloads. Text entered directly in a prompt can appear in an enabled prompt excerpt.
268
325
 
269
- - `prompt_excerpt`: up to 500 Unicode characters of the current human task for main requests or requests without a class header. Tool continuations and recognized `/goal` feedback keep the originating human task. System instructions, standalone reminder blocks, tool results, images, documents, and thinking are omitted. Auxiliary classifiers, compaction, subagents, and workflows have empty excerpts; a new attachment-only task also has an empty excerpt. `prompt_truncated` indicates that text exceeded the limit.
270
- - `selected_model`: AutoRouter's final selected model after compatibility checks, before the upstream response. It does not confirm which model successfully answered.
271
- - `decision_latency_ms`: time spent making the routing decision, including evaluator waiting, cache lookup, and any context checks. It excludes Claude generation time and log writing. A `passthrough` entry can be near zero because no evaluator was called.
272
- - `source` and `reason`: distinguish evaluator choices, cache hits, fallbacks, turn/model constraints, and native safety pass-through. `classified_tier` is included when an evaluator returned a tier, which may differ from the final selected model.
326
+ The settings also work with `serve` and every client profile. `setup --session-log-dir DIR --session-log-mode metadata --force` updates an existing configuration. Setup resolves relative directories at setup time; environment-only paths resolve from the launch directory. `AUTOROUTER_SESSION_LOG_DIR=''` disables a saved directory for one launch. Setting only the mode never enables logging. `doctor` reports preferences without creating files.
273
327
 
274
- Inspect a file with:
328
+ Files are named `autorouter-session-*.jsonl`, one per observed Claude session in each router launch. Resuming a session in a new launch creates a new file. Requests without a session header share an anonymous file; subagents retain their agent IDs. Every line is a bounded schema-2 JSON record:
275
329
 
276
- ```sh
277
- jq -c '{prompt_excerpt, selected_model, decision_latency_ms, source, reason}' /path/to/autorouter-session-EXAMPLE.jsonl
278
- ```
330
+ - `decision`: requested and selected models, evaluator verdict/source, policy reason, compatibility and continuity detail. `evaluation_latency_ms` measures evaluation/cache waiting; `routing_latency_ms` (also `decision_latency_ms`) includes compatibility and context checks.
331
+ - `outcome`: the same `request_id`, observed `confirmed_model` and bounded model transitions, safe error category, `completed`, `error` or `cancelled` status, and separate `completion_confirmed` evidence. It includes usage when available, `first_response_ms` from upstream forwarding to response headers, and total request latency. Early errors can have an outcome without a decision.
332
+
333
+ Outcomes record the configured Opus baseline and pricing-table version. History prices only successfully completed, supported usage with a recognized recorded table version and baseline. Unknown prices, missing or partial streams, ambiguous/mixed-model usage, and legacy records remain unpriced with a reason. Estimates never represent subscription charges. Status savings identify partial coverage; history exposes the version and counts.
279
334
 
280
- There is one record per completed routing decision, including requests whose upstream call later fails. Requests rejected before routing or cancelled before a decision are not recorded. Logs contain user text and are local plaintext: the feature is disabled by default, new directories use `0700`, and files use `0600`. Existing directory permissions are left unchanged. Authentication headers, provider replies, full transcripts, and tool payloads are not logged; text you put directly in a prompt can appear in its excerpt.
335
+ New directories use `0700`, files use `0600`; existing directory permissions are unchanged. Asynchronous writes use a bounded 1 MiB queue and at most 128 session files per process. Normal shutdown drains accepted records. Storage failure or queue limits disable further logging with one generic warning while routing continues. Abrupt termination can lose unwritten records.
281
336
 
282
- Writes run asynchronously through a bounded 1 MiB queue and support up to 128 session files per router process. Normal shutdown drains accepted records. Filesystem failure or a queue/session limit disables further logging with one generic warning while routing continues. An abrupt process kill or storage failure can lose unwritten records. Logs are retained without automatic rotation or deletion; manage them in your chosen directory. Log filenames are ignored by this repository and excluded from the npm package.
337
+ History reads at most 100 files, 4 MiB per file, 16 MiB total and 5,000 records per file. Limits, malformed records and partial tails are reported as partial coverage. Files are never automatically rotated or deleted; manage retention in your chosen directory. Log filenames are ignored by this repository and excluded from the npm package.
283
338
 
284
339
  ## Troubleshooting
285
340
 
@@ -319,7 +374,7 @@ The environment-only command also works on AutoRouter 0.3.4. Saved configuration
319
374
  "CLAUDE_CODE_STOP_HOOK_BLOCK_CAP": "2"
320
375
  ```
321
376
 
322
- For a new configuration, use `claude-autorouter setup --stop-hook-block-cap 2`; the flag works with either evaluator and overrides the environment during setup. Runtime environment values override saved configuration. `setup --force` replaces the entire config, so keep your existing evaluator/authentication options if using it. `doctor` reports the cap when configured. AutoRouter accepts nonnegative safe integers and leaves the setting absent unless you opt in.
377
+ For a new configuration, use `claude-autorouter setup --stop-hook-block-cap 2`; the flag works with either evaluator and overrides the environment during setup. Runtime environment values override saved configuration. To change only a saved cap, use `config set CLAUDE_CODE_STOP_HOOK_BLOCK_CAP 2` or `setup --stop-hook-block-cap 2 --force`; both preserve unrelated saved settings. `doctor` reports the cap when configured. AutoRouter accepts nonnegative safe integers and leaves the setting absent unless you opt in.
323
378
 
324
379
  ### Other session issues
325
380
 
package/docs/releasing.md CHANGED
@@ -16,20 +16,28 @@ Version `0.3.6` fixes Auto permission-mode launches with an Auto-compatible Sonn
16
16
 
17
17
  Version `0.3.7` enables automatic Sonnet/Opus switching for compatible Auto-mode execution requests, including requests carrying the known server safety-review contract. The Auto profile defaults to Sonnet 5.5 and Opus 5.5, floors Haiku decisions to Sonnet, and retains the selected model through tool and goal continuations. Shared native context edits, mid-conversation system messages, and signed thinking history no longer pin new human tasks. Permission-classifier requests and safety verdicts remain unchanged; unknown contracts and incompatible model features still preserve a compatible model. Explicit model overrides remain in effect. See [Auto permission mode](reference.md#auto-permission-mode).
18
18
 
19
- The GitHub repository is private. Publishing to npm makes the tarball's runtime source, README, configuration example, license, and shipped documentation public. Model weights, user configuration, credentials, transcripts, session logs, local artifacts, and test fixtures are excluded. Review the archive before the first publication and whenever the package allowlist changes.
19
+ Version `0.4.0` completes the routing, configuration, history and performance improvement plan. Shared compatibility checks and durable task state preserve valid request features and confirmed tool/goal continuity across evaluator cache expiry, provider fallback and concurrent requests. New `config show/set/unset`, `sessions list/show` and `doctor --evaluate-local` commands support focused configuration edits, private metadata-only history and explicit local diagnostics. `setup --force` now merges saved settings; use `--replace` for deliberate replacement. Optional logging remains disabled by default; new schema-2 decision/outcome records separate selected and observed models, while the reader still accepts schema-1 files. Consumers parsing JSONL directly should account for both event kinds and the new schema. Identical concurrent evaluations are coalesced with independent cancellation, responses are bounded, and status persistence is asynchronous. Releases retain the tested archive and verify public npm availability and installation after submission. Actual 16 GiB/64 GiB Ollama results retain failed quality gates and comparison limits; Jev remains the default. See the [configuration/history reference](reference.md) and [hardware comparison](hardware-comparison.md).
20
+
21
+ Version `0.5.0` makes local Ollama the default evaluator and TypeSafe Jev an explicit option (`setup --evaluator jev`), so evaluator excerpts stay on the machine unless the user opts in. **Breaking for environment-only launches** that relied on the implicit Jev default; configurations created by `setup` record their evaluator and are unchanged. Evaluator excerpts and opt-in session-log prompt excerpts are now redacted for recognizable credentials and personal identifiers (pattern-based, not exhaustive). On macOS, new `setup` runs keep saved keys in the login Keychain by default; existing plaintext configurations are not moved implicitly, and `doctor` prints the `config set AUTOROUTER_SECRET_STORE keychain` command to move them. See [credential storage](reference.md#credential-storage) and the [data flow](reference.md#data-flow-and-authentication).
22
+
23
+ The GitHub repository became public on October 6, 2026, after preparation PR #1 merged. That launch created no release tag and published no new npm version. npm publication remains a separate release operation. Its tarball includes runtime source, README, configuration example, license, and shipped documentation; model weights, user configuration, credentials, transcripts, session logs, local artifacts, and test fixtures are excluded. Review each release archive, especially when the package allowlist changes. The public Git repository also exposes history, development scripts and tests; the npm archive allowlist does not govern that material.
20
24
 
21
25
  ## What runs automatically
22
26
 
23
27
  | Workflow | Trigger | Behavior |
24
28
  | --- | --- | --- |
25
29
  | [ci.yml](https://github.com/frapposelli/claude-autorouter/blob/main/.github/workflows/ci.yml) | Pull requests, pushes to `main`, manual runs, and calls from the release workflow | Syntax checks, tests, and package smoke tests on Ubuntu/macOS with Node 22/24 |
26
- | [publish.yml](https://github.com/frapposelli/claude-autorouter/blob/main/.github/workflows/publish.yml) | Push of a tag matching `v*` | Validate release, run CI, pack and test the candidate, then publish the verified archive |
30
+ | [publish.yml](https://github.com/frapposelli/claude-autorouter/blob/main/.github/workflows/publish.yml) | Tag push; manual verification-only dispatch | Test and submit one canonical archive; independently verify registry availability and installation. Manual dispatch never publishes. |
27
31
 
28
32
  A release tag must exactly equal `v` plus the version in `package.json`, and its commit must be reachable from `origin/main`. Package name and repository metadata must match `claude-autorouter` and `frapposelli/claude-autorouter`. Stable versions use npm's `latest` tag; prereleases such as `0.3.1-beta.1` use `next`.
29
33
 
30
- The release workflow packs its candidate once and smoke-tests that exact `.tgz`. It uploads the archive and SHA-256 checksum as an Actions artifact. A separate publishing job downloads that artifact by its immutable ID, checks the checksum and every packaged file against the release checkout, then runs `npm publish` with scripts disabled. The publish job uses a GitHub-hosted Ubuntu runner, Node 24, and npm 11.19.1. Only that job has `id-token: write`; there is no `NPM_TOKEN` secret or required GitHub environment. Failed checks prevent publication.
34
+ The release workflow packs its candidate once and smoke-tests that exact `.tgz`. It retains the archive, SHA-256 checksum, and commit-based release notes as Actions artifacts. The publishing job downloads the canonical artifact by immutable ID, checks every packaged file against the release checkout, and checks fresh npm metadata before submission. A new stable version must be greater than the current stable `latest`. An already-visible identical version skips publication; a different archive under that version is an error. Registry/network errors never count as proof that a version is unused.
35
+
36
+ Only the publishing job has `id-token: write`; there is no `NPM_TOKEN` or required GitHub environment. It runs on GitHub-hosted Ubuntu with Node 24 and npm 11.19.1. The separate verifier has read-only permissions and never changes npm distribution tags. The workflow serializes its publishers, but independent/manual publishers must coordinate: npm does not provide an atomic compare-and-swap for the `latest` tag. A newer `latest` observed during verification is reported as superseding this release and is never moved backward.
37
+
38
+ Verification polls uncached version metadata and the package document, checks the downloaded tarball against the tested archive, and installs the exact version from the public registry into a temporary prefix with an empty npm cache/config. It compares the installed file set and bytes with the canonical archive before invoking that executable’s `--version` and `--help` from an unrelated directory, with lifecycle scripts disabled and no evaluator credentials. Only a successful public install and matching artifact produce `verified`.
31
39
 
32
- The project has no package dependencies or lockfile, so CI runs its scripts directly without `npm ci`. Live Claude/Jev calls, Ollama downloads, and private repository probes are not CI checks.
40
+ The installed CLI has no runtime dependencies. Contributor tooling uses pinned development dependencies and a committed `package-lock.json`; CI installs them with `npm ci --ignore-scripts --no-audit --no-fund` before checks. Live Claude/Jev calls, Ollama downloads, and private repository probes are not CI checks.
33
41
 
34
42
  ## 1. Publish the first version interactively
35
43
 
@@ -83,7 +91,7 @@ claude-autorouter --help
83
91
 
84
92
  Then run `setup`, `doctor`, and a launch from outside the source checkout as appropriate for that machine. `doctor` is local-only; a live prompt separately verifies provider access. Keep the README's installation instructions aligned with the verified registry release.
85
93
 
86
- Do not push `v0.2.0` to test automation after this bootstrap: it would attempt to publish an existing version. npm name/version pairs cannot be reused, including after unpublishing. See the [npm publish reference](https://docs.npmjs.com/cli/v11/commands/npm-publish/).
94
+ Do not push `v0.2.0` to test automation after this bootstrap. Keep the original bootstrap archive; a current verifier only accepts releases matching its archive and repository validation rules. npm name/version pairs cannot be reused, including after unpublishing. See the [npm publish reference](https://docs.npmjs.com/cli/v11/commands/npm-publish/).
87
95
 
88
96
  ## 2. Authorize this workflow on npm
89
97
 
@@ -101,68 +109,102 @@ On npmjs.com, open the `claude-autorouter` package's **Settings → Trusted Publ
101
109
 
102
110
  Use the filename only, not `.github/workflows/publish.yml`. No GitHub environment or npm token secret needs to be created. The owner, repository, and workflow must match exactly. New trust configurations default to permitting staged publication; **enable direct `npm publish`** for this workflow. See [npm trusted publishers](https://docs.npmjs.com/trusted-publishers/) and [staged publishing](https://docs.npmjs.com/staged-publishing/).
103
111
 
104
- The workflow explicitly disables provenance because npm provenance is unsupported for private source repositories, even when the npm package is public. OIDC authentication still works. If the repository becomes public, review the workflow and metadata before enabling provenance. See [npm provenance requirements](https://docs.npmjs.com/generating-provenance-statements/).
112
+ The publishing workflow selects provenance from repository visibility. With the source now public, the workflow requests provenance on future publication; it would disable provenance for a private source repository. OIDC authentication works independently of provenance. The package preparation job's dry run always disables provenance because it has no OIDC publishing permission. Before the first public-source release, review repository metadata and npm trust settings, then verify the resulting provenance statement. The public launch itself created no attestation, and changing visibility does not add attestations to historical releases. See [npm provenance requirements](https://docs.npmjs.com/generating-provenance-statements/).
105
113
 
106
114
  After a successful trusted release, npm recommends the optional **Publishing access → Require two-factor authentication and disallow tokens** setting. It does not disable OIDC publishing. See [restricting token access](https://docs.npmjs.com/trusted-publishers/#recommended-restrict-token-access-when-using-trusted-publishers).
107
115
 
108
116
  ## 3. Release subsequent versions by tag
109
117
 
110
- The commands below illustrate the `0.3.2` release. For a new release, substitute the next unused version throughout; never reuse a published version:
118
+ Use the next unused version. The following commands use `0.4.0` as an example, not as a claim that it is currently available:
111
119
 
112
120
  ```sh
113
121
  git switch main
114
122
  git pull --ff-only origin main
115
- npm version 0.3.2 --no-git-tag-version
116
- ```
117
-
118
- Review the version change and update any version-specific install examples or release notes. Check the candidate using the new filename:
119
-
120
- ```sh
123
+ npm version 0.4.0 --no-git-tag-version
121
124
  npm run check
122
125
  npm test
123
126
  npm run release:pack
124
- npm run test:package -- --archive ./dist/claude-autorouter-0.3.2.tgz
127
+ npm run test:package -- --archive ./dist/claude-autorouter-0.4.0.tgz
125
128
  git diff --check
126
129
  ```
127
130
 
128
- Commit the intended release changes and get that commit onto `main`, either through a pull request or a direct push allowed by the repository's branch rules. For a direct push with only the version changed:
131
+ Review the archive, version-specific documentation, and release notes. Commit all intended changes and get that commit onto `main` through the repository’s normal review process. Wait for CI to pass, then tag the exact release commit:
129
132
 
130
133
  ```sh
131
- git add package.json
132
- git commit -m "Release 0.3.2"
133
- git push origin main
134
+ git switch main
135
+ git pull --ff-only origin main
136
+ git tag -a v0.4.0 -m "Release 0.4.0"
137
+ git push origin v0.4.0
134
138
  ```
135
139
 
136
- Include any intentional documentation or release-note edits in that commit too. There is no publication from a branch push or PR merge. Wait for CI to pass, then tag that exact release commit:
140
+ The tag must match `package.json`. For a prerelease, use matching values such as `0.4.0-beta.1` / `v0.4.0-beta.1`; publication uses `next`, leaving `latest` unchanged. Release stable versions in increasing order, and wait for verification or investigate a pending submission before starting another stable release. The preflight blocks a new stable candidate that is not newer than the registry’s `latest`; queue order alone does not establish version order.
141
+
142
+ For local checks after a tag exists, `node scripts/release-check.mjs source v0.4.0` validates the clean checkout, tag, metadata, and main ancestry. `node scripts/release-check.mjs archive v0.4.0` validates the candidate checksum and contents. `dist/` must contain only that candidate’s `.tgz` and `.sha256`, so retain older artifacts elsewhere.
143
+
144
+ ## 4. Inspect the release state and retain evidence
145
+
146
+ Open the tag’s run under [GitHub Actions](https://github.com/frapposelli/claude-autorouter/actions). Submission success is not proof that users can install the package. The verification job’s summary and retained JSON report distinguish:
147
+
148
+ | State | Meaning | Next step |
149
+ | --- | --- | --- |
150
+ | `preflight_ready` | The new candidate passed metadata and version-order checks; submission has not occurred | The initial workflow attempt may submit it |
151
+ | `submitted` | npm accepted the command, or an identical immutable version was already visible | Wait for independent verification |
152
+ | `validating_unavailable` | Metadata, tarball, distribution tag, or installation is still unavailable, or the registry is failing | Keep the original archive and rerun verification |
153
+ | `verified` | Exact archive integrity, distribution-tag state, and isolated public installation passed | Use the recorded upgrade command |
154
+ | `failed` | An input, provenance, integrity, version-order, or executable check failed | Investigate the report before changing anything |
155
+
156
+ The verifier polls for up to 15 minutes with backoff, then reports pending verification with a successful command exit. **A green workflow can therefore mean pending, not verified; read the recorded state.** An npm processing delay does not establish publication failure or justify a duplicate release. HTTP/network failures are reported separately from a missing version or tarball.
157
+
158
+ The workflow retains these artifacts for 90 days, subject to repository retention policy:
159
+
160
+ - `npm-package-<run-id>-<attempt>`: the canonical tested archive and SHA-256 file.
161
+ - `release-notes-<run-id>-<attempt>`: notes derived from the tagged source’s commits and archive identity.
162
+ - `release-submission-<run-id>-<attempt>`: preflight and, when accepted, submission reports.
163
+ - `release-verification-<run-id>-<attempt>`: the canonical archive, checksum, available notes, and verification report.
164
+
165
+ Download and retain the archive/checksum, notes, and report before Actions artifacts expire; attaching them to a GitHub Release is suitable for long-term retention. Artifact expiration is not a reason to repack a supposedly identical candidate for verification.
166
+
167
+ After the report says `verified`, use its exact version:
137
168
 
138
169
  ```sh
139
- git switch main
140
- git pull --ff-only origin main
141
- git tag -a v0.3.2 -m "Release 0.3.2"
142
- git push origin v0.3.2
170
+ npm install -g claude-autorouter@0.4.0
171
+ claude-autorouter --version
172
+ claude-autorouter --help
143
173
  ```
144
174
 
145
- Before pushing, confirm `package.json` contains `0.3.2` and the tag points to the intended commit. For a prerelease, use a matching version/tag such as `0.4.0-beta.1` / `v0.4.0-beta.1`; it will publish under `next`, leaving `latest` unchanged.
175
+ The verifier runs the equivalent exact-version registry install in isolation. If `latest` has since advanced, the report explicitly marks this release as superseded; it does not restore an older tag. An unqualified `npm install -g claude-autorouter` follows the registry’s current `latest` instead.
146
176
 
147
- Release stable versions in increasing version order, one tag at a time, and wait for each run to finish before pushing the next stable tag. The workflow queues releases without canceling an active run, but queue order does not sort semantic versions. Publishing an older stable version afterward could move `latest` backward; there is no registry version-order gate.
177
+ ## 5. Rerun verification without publishing
148
178
 
149
- Open the tag's run under [GitHub Actions](https://github.com/frapposelli/claude-autorouter/actions). Under **Artifacts**, download `npm-package-<run-id>-<run-attempt>`, which contains the `.tgz` and checksum used for publication. Artifacts expire after 30 days, so retain them with the release record. After the publish job succeeds, verify the registry version and tags:
179
+ Use the Actions **Run workflow** control for `publish.yml`, or the command below. Provide the release tag and the numeric run/artifact IDs from the original tag-triggered run:
150
180
 
151
181
  ```sh
152
- npm view claude-autorouter@0.3.2 version dist.integrity --registry https://registry.npmjs.org/
153
- npm view claude-autorouter dist-tags --json --registry https://registry.npmjs.org/
182
+ gh workflow run publish.yml --ref main \
183
+ -f tag=v0.4.0 \
184
+ -f run_id=ORIGINAL_RUN_ID \
185
+ -f artifact_id=CANONICAL_NPM_PACKAGE_ARTIFACT_ID
154
186
  ```
155
187
 
156
- Repeat the independent installation check for the released version. A GitHub Release page is optional; pushing the version tag is the publication trigger.
188
+ This path validates that the artifact belongs to this repository’s original tag-triggered publishing workflow and matches the tag commit on `main`. It downloads those immutable bytes, checks their checksum and package metadata, and verifies npm availability. It cannot publish or alter npm tags and does not need npm OIDC permissions. Old runs may lack retained release notes; that does not prevent archive verification.
189
+
190
+ You can also download the original archive and its checksum and run the verifier independently from a current source checkout:
191
+
192
+ ```sh
193
+ node scripts/release-verify.mjs verify v0.4.0 \
194
+ --archive /path/to/claude-autorouter-0.4.0.tgz \
195
+ --report artifacts/release-verification-0.4.0.json \
196
+ --timeout-ms 900000
197
+ ```
157
198
 
158
- For local release diagnostics after the tag exists, `node scripts/release-check.mjs source v0.3.2` checks the tag, clean checkout, metadata, and ancestry. `node scripts/release-check.mjs archive v0.3.2` checks the candidate checksum and contents; `dist/` must contain only that version's archive and checksum, so retain older artifacts elsewhere first. These helpers are run automatically in the release workflow; the first untagged bootstrap uses the checks in step 1 instead.
199
+ Both files must retain their original names, and the archive must satisfy the public-package validation rules. This command performs only registry reads and a temporary isolated install; it neither publishes nor changes your installed CLI, npm login, or AutoRouter configuration. Exit `0` includes pending verification, so automation must inspect the report’s `state`; exit `1` means verification failed. A mismatch is never accepted as an already-published identical release.
159
200
 
160
- ## Recovering a failed release
201
+ ## Recovering a release
161
202
 
162
- - **Checks or archive validation failed:** nothing is published. Fix the cause and repeat validation before making a new release tag. Do not move a tag that already identifies a published version.
163
- - **npm rejected OIDC authentication:** verify the npm trust fields, direct-publish permission, GitHub-hosted runner, and the publish job's OIDC permission. After correcting npm configuration, rerun the failed job if that version is still unpublished.
164
- - **A publish timed out or the run was interrupted:** check `npm view` for the exact version before retrying. The registry may have accepted it before the connection failed.
165
- - **The version already exists:** inspect the registry release; do not overwrite or unpublish to reuse it. Code or documentation corrections need a new version.
166
- - **Local bootstrap authentication failed:** complete `npm login` and the account's 2FA flow in your terminal. CI trust cannot create the first package or substitute for that account step.
203
+ - **Checks/archive validation failed:** nothing was submitted by that failed path. Fix the cause, repeat validation, and follow the normal release process. Never move a published tag.
204
+ - **Registry preflight failed:** an HTTP/authentication/network failure is not evidence that the version is unused. Restore visibility before submitting.
205
+ - **npm rejected OIDC:** check the trusted-publisher fields, direct-publish permission, hosted runner, and OIDC permission. A job rerun deliberately does not resubmit a still-invisible version because an interrupted command may already have been accepted. Establish what happened before planning a new submission.
206
+ - **Submission timed out, the run was interrupted, or npm is processing it:** retain the original archive and use verification-only dispatch. A retry with a matching visible version skips `npm publish`; a retry with an invisible version only verifies it.
207
+ - **Existing version has different bytes:** stop. Never overwrite, unpublish, or replace the canonical artifact to reuse that version. Corrections require a new version.
208
+ - **Verification remains pending:** do not call it a failed publication. Rerun the verifier later with the original artifact, inspect npm’s public status if the registry is failing, and retain the evidence for support if validation remains stuck.
167
209
 
168
- If release code or workflow changes are needed, commit the fix to `main` and prepare a new version/tag. Authentication-only corrections on npm can be retried against the unchanged, unpublished candidate.
210
+ The workflow never repairs distribution tags automatically. npm’s [publish semantics](https://docs.npmjs.com/cli/commands/npm-publish/) make published name/version pairs immutable; [distribution tags](https://docs.npmjs.com/adding-dist-tags-to-packages/) are mutable references, so verification observes them without overwriting a newer release.