claude-autorouter 0.3.4 → 0.3.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.env.example +12 -0
- package/README.md +19 -1
- package/bin/autorouter.mjs +28 -2
- package/docs/reference.md +83 -2
- package/docs/releasing.md +5 -1
- package/package.json +1 -1
- package/src/auth.mjs +21 -0
- package/src/config.mjs +43 -7
- package/src/onboarding.mjs +30 -3
- package/src/prompt-state.mjs +48 -0
- package/src/router.mjs +30 -3
- package/src/server.mjs +17 -1
- package/src/session-log.mjs +171 -0
- package/src/status-state.mjs +1 -1
- package/src/statusline.mjs +3 -2
- package/src/user-config.mjs +2 -1
package/.env.example
CHANGED
|
@@ -5,8 +5,20 @@ AUTOROUTER_CLIENT_PROFILE=compatible
|
|
|
5
5
|
AUTOROUTER_EVALUATOR=jev
|
|
6
6
|
# The launcher enables the router status line for this session. Set 0 to keep your own.
|
|
7
7
|
AUTOROUTER_STATUSLINE=1
|
|
8
|
+
# Claude's Auto permission mode needs a supported Sonnet or Opus client.
|
|
9
|
+
# This profile defaults to Sonnet and excludes Haiku from task routing.
|
|
10
|
+
# AutoRouter also selects it for: claude-autorouter claude --permission-mode auto
|
|
11
|
+
# AUTOROUTER_CLIENT_PROFILE=auto
|
|
8
12
|
# Optional metadata logs on stderr. Redirect stderr to a file when using the UI.
|
|
9
13
|
# AUTOROUTER_DEBUG=1
|
|
14
|
+
# Optional persistent JSONL decision logs, one file per session per launch.
|
|
15
|
+
# Includes up to 500 characters of user prompt text; keep the directory local.
|
|
16
|
+
# Unset or empty disables logging. This does not print prompts in the terminal.
|
|
17
|
+
# AUTOROUTER_SESSION_LOG_DIR=/absolute/path/to/autorouter-sessions
|
|
18
|
+
# Optional: allow two tool-free Stop-hook continuations, then end the turn on
|
|
19
|
+
# the third block. Applies to /goal and all Stop/SubagentStop hooks.
|
|
20
|
+
# Unset keeps Claude's default (currently 8); 0 DISABLES the cap.
|
|
21
|
+
# CLAUDE_CODE_STOP_HOOK_BLOCK_CAP=2
|
|
10
22
|
# Required for Jev only; subscription + Ollama needs no API keys.
|
|
11
23
|
TYPESAFE_API_KEY=
|
|
12
24
|
# For API billing instead, set AUTOROUTER_AUTH_MODE=api-key and fill this in.
|
package/README.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Claude AutoRouter
|
|
2
2
|
|
|
3
|
-
Use Haiku, Sonnet, and Opus in one Claude Code session. A local gateway classifies
|
|
3
|
+
Use Haiku, Sonnet, and Opus in one Claude Code session. A local gateway classifies coding requests with the selected evaluator, applies compatibility and context checks, and streams the selected model's response back to Claude Code. Claude's internal classifiers and server safety-review requests pass through unchanged. [TypeSafe Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) is the default; an experimental Ollama backend evaluates requests locally.
|
|
4
4
|
|
|
5
5
|
Requires Node.js 22+, macOS or Linux (including WSL), an installed `claude` command, and a Claude subscription login or Anthropic API key. The default evaluator also requires a [TypeSafe API key](https://console.typesafe.ai). There are no runtime package dependencies. Native Windows is not supported in this release.
|
|
6
6
|
|
|
@@ -35,6 +35,14 @@ claude-autorouter --version
|
|
|
35
35
|
|
|
36
36
|
For API billing, use `claude-autorouter setup --auth-mode api-key`. Use `--force` to replace an existing config. Automation can supply `TYPESAFE_API_KEY` and, in API-key mode, `ANTHROPIC_API_KEY` through the environment; keys are never command-line arguments. `doctor` checks local configuration and Claude installation/login state without paid requests. See the [configuration reference](docs/reference.md#configuration).
|
|
37
37
|
|
|
38
|
+
To use Claude's Auto permission mode (AutoRouter 0.3.6+):
|
|
39
|
+
|
|
40
|
+
```sh
|
|
41
|
+
claude-autorouter claude --permission-mode auto
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
This selects the Auto-compatible profile: a Sonnet starting model, native thinking, and Sonnet/Opus task routing. Haiku does not support Auto mode. Claude's own permission classifiers and requests carrying server safety-review settings pass through with their selected model unchanged. Organization policies still apply. Use `AUTOROUTER_CLIENT_PROFILE=auto` for sessions where you select Auto in Claude's UI or saved settings. [Auto-mode support and limitations](docs/reference.md#auto-permission-mode).
|
|
45
|
+
|
|
38
46
|
## What you see
|
|
39
47
|
|
|
40
48
|
The launcher adds a temporary status line and leaves saved Claude Code settings unchanged:
|
|
@@ -46,6 +54,15 @@ The launcher adds a temporary status line and leaves saved Claude Code settings
|
|
|
46
54
|
|
|
47
55
|
The confirmed model comes from Anthropic's response. Claude's own model label can still show its Haiku starting model. `API ctx` measures input against the actual model's known window; a different client limit remains visible as `CLI ctx`.
|
|
48
56
|
|
|
57
|
+
For a separate decision log per session (optional, disabled by default; AutoRouter 0.3.6+):
|
|
58
|
+
|
|
59
|
+
```sh
|
|
60
|
+
env AUTOROUTER_SESSION_LOG_DIR="$HOME/.local/state/claude-autorouter/sessions" \
|
|
61
|
+
claude-autorouter claude
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
Each JSONL record includes a bounded prompt excerpt, selected model, decision latency, and routing reason. Files persist after Claude exits; terminal output stays quiet. [Session logs](docs/reference.md#session-decision-logs).
|
|
65
|
+
|
|
49
66
|
Savings are an **API-equivalent estimate for the same token counts**, using Opus as the baseline. They do not measure subscription bill reductions or quota credits and exclude Jev and local compute costs. [Status line and savings details](docs/reference.md#status-line-and-savings).
|
|
50
67
|
|
|
51
68
|
## Experimental local evaluator
|
|
@@ -104,5 +121,6 @@ Historical measurements before 0.3.2, on a 16 GiB M4: Tev1 0.8B matched 18/24 he
|
|
|
104
121
|
- The selected evaluator receives bounded excerpts that can contain source code and tool results: TypeSafe with Jev, or the local service with Ollama. Jev also receives system-text excerpts; the local path excludes Claude's executor system instructions. Anthropic receives the complete request. Images, document payloads, and private thinking are omitted from classifier input. [Data flow and authentication](docs/reference.md#data-flow-and-authentication).
|
|
105
122
|
- Subscription access and usage limits still apply. Model switches can reduce cache reuse; cheaper token prices do not guarantee cheaper completed tasks. Run ordinary `claude` to bypass routing.
|
|
106
123
|
- The launcher is quiet by default. Use `AUTOROUTER_DEBUG=1` for metadata diagnostics or `AUTOROUTER_STATUSLINE=0` to retain your existing status line. [Troubleshooting](docs/reference.md#troubleshooting).
|
|
124
|
+
- For blocked `/goal` loops, optionally launch with `env CLAUDE_CODE_STOP_HOOK_BLOCK_CAP=2 claude-autorouter claude`. Claude then ends the turn on the third consecutive blocking verdict without tool use, leaving the goal unmet. This also affects other Stop/SubagentStop hooks; defaults are unchanged. [Scope and saved configuration](docs/reference.md#shorter-stop-hook-loops-opt-in).
|
|
107
125
|
|
|
108
126
|
[Reference](docs/reference.md) · [Development and validation](docs/development.md) · [CI and npm release setup](docs/releasing.md) · [Apache-2.0 license](LICENSE)
|
package/bin/autorouter.mjs
CHANGED
|
@@ -4,10 +4,11 @@ import { spawn } from 'node:child_process';
|
|
|
4
4
|
import { readFileSync } from 'node:fs';
|
|
5
5
|
import { readConfig, requireKeys } from '../src/config.mjs';
|
|
6
6
|
import { createRouterServer, listen } from '../src/server.mjs';
|
|
7
|
-
import { buildClaudeEnv, conflictingProviders } from '../src/auth.mjs';
|
|
7
|
+
import { buildClaudeEnv, clientProfileForLaunch, conflictingProviders } from '../src/auth.mjs';
|
|
8
8
|
import { dirname } from 'node:path';
|
|
9
9
|
import { createStatusState } from '../src/status-state.mjs';
|
|
10
10
|
import { addStatusLineSettings } from '../src/status-settings.mjs';
|
|
11
|
+
import { createSessionLog } from '../src/session-log.mjs';
|
|
11
12
|
import { loadUserConfig } from '../src/user-config.mjs';
|
|
12
13
|
import { setup, doctor, ollamaDeadlineText } from '../src/onboarding.mjs';
|
|
13
14
|
import { setupOllama } from '../src/ollama-setup.mjs';
|
|
@@ -21,8 +22,11 @@ if (['--version', '-v', 'version'].includes(command)) {
|
|
|
21
22
|
|
|
22
23
|
Usage:
|
|
23
24
|
claude-autorouter setup [--auth-mode subscription|api-key] [--force]
|
|
25
|
+
[--client-profile compatible|native|auto]
|
|
24
26
|
[--evaluator jev|ollama]
|
|
25
27
|
[--ollama-model MODEL] [--ollama-timeout-ms N] [--pull]
|
|
28
|
+
[--stop-hook-block-cap N]
|
|
29
|
+
[--session-log-dir DIR]
|
|
26
30
|
claude-autorouter doctor
|
|
27
31
|
claude-autorouter claude [Claude Code arguments]
|
|
28
32
|
claude-autorouter serve
|
|
@@ -47,10 +51,18 @@ AUTOROUTER_AUTH_MODE=subscription uses your saved Claude Code login.
|
|
|
47
51
|
Without setup, AUTOROUTER_AUTH_MODE defaults to api-key and also requires ANTHROPIC_API_KEY.
|
|
48
52
|
AUTOROUTER_CLIENT_PROFILE=compatible (default) enables all three routing tiers.
|
|
49
53
|
Use AUTOROUTER_CLIENT_PROFILE=native to retain Claude Code's own model/thinking settings.
|
|
54
|
+
Use AUTOROUTER_CLIENT_PROFILE=auto for Auto permission mode: Sonnet/Opus routing, native thinking.
|
|
55
|
+
An explicit claude --permission-mode auto selects the auto profile for that launch.
|
|
56
|
+
Claude's permission checks and organization policies still apply; Haiku does not support Auto mode.
|
|
57
|
+
Optional CLAUDE_CODE_STOP_HOOK_BLOCK_CAP=N limits consecutive tool-free Stop-hook continuations.
|
|
58
|
+
Use 2 to stop on the third block; applies to /goal and all Stop/SubagentStop hooks.
|
|
59
|
+
Unset preserves Claude's default; 0 disables the cap. Setup --stop-hook-block-cap N saves it.
|
|
50
60
|
Standalone serve also requires AUTOROUTER_TOKEN (at least 16 characters).
|
|
51
61
|
The claude launcher creates a temporary credential and an ephemeral port.
|
|
52
62
|
It enables an AutoRouter status line for this session (AUTOROUTER_STATUSLINE=0 to opt out).
|
|
53
63
|
Launcher logs are quiet by default; AUTOROUTER_DEBUG=1 enables diagnostic logs on stderr.
|
|
64
|
+
AUTOROUTER_SESSION_LOG_DIR writes private per-session JSONL decision logs with prompt excerpts.
|
|
65
|
+
Unset or empty disables session logs. Setup --session-log-dir DIR saves the directory.
|
|
54
66
|
Jev sends prompt excerpts to TypeSafe; Ollama keeps classification on this machine.
|
|
55
67
|
Complete inference requests still go to Anthropic. See README.md.`);
|
|
56
68
|
} else if (command === 'setup' || command === 'doctor') {
|
|
@@ -67,14 +79,24 @@ Complete inference requests still go to Anthropic. See README.md.`);
|
|
|
67
79
|
} else {
|
|
68
80
|
let server;
|
|
69
81
|
let status;
|
|
82
|
+
let sessionLog;
|
|
83
|
+
let stopping;
|
|
70
84
|
const stop = () => {
|
|
85
|
+
if (stopping) return stopping;
|
|
71
86
|
if (server) { server.close(); server.closeAllConnections(); }
|
|
72
87
|
status?.close();
|
|
88
|
+
// Drain accepted decision records before normal process exit. Pending
|
|
89
|
+
// filesystem writes keep Node alive; no timer or fire-and-forget buffer.
|
|
90
|
+
stopping = Promise.resolve().then(() => sessionLog?.close()).catch(() => {});
|
|
91
|
+
return stopping;
|
|
73
92
|
};
|
|
74
93
|
try {
|
|
75
94
|
if (command === 'serve' && args.length) throw new Error('Usage: claude-autorouter serve');
|
|
76
95
|
const runtimeEnv = loadUserConfig().env;
|
|
77
|
-
const config = readConfig(
|
|
96
|
+
const config = readConfig(command === 'claude' ? {
|
|
97
|
+
...runtimeEnv,
|
|
98
|
+
AUTOROUTER_CLIENT_PROFILE: clientProfileForLaunch(runtimeEnv.AUTOROUTER_CLIENT_PROFILE ?? 'compatible', args),
|
|
99
|
+
} : runtimeEnv);
|
|
78
100
|
requireKeys(config);
|
|
79
101
|
const diagnosticLogs = command === 'serve' || runtimeEnv.AUTOROUTER_DEBUG === '1';
|
|
80
102
|
if (command === 'claude') {
|
|
@@ -105,11 +127,15 @@ Complete inference requests still go to Anthropic. See README.md.`);
|
|
|
105
127
|
}
|
|
106
128
|
} else console.error('AutoRouter status line unavailable: could not create local status storage.');
|
|
107
129
|
}
|
|
130
|
+
if (config.sessionLogDir) sessionLog = await createSessionLog(config.sessionLogDir, {
|
|
131
|
+
warn: message => console.error(message),
|
|
132
|
+
});
|
|
108
133
|
// Claude owns the terminal while its UI is running. Status updates use the
|
|
109
134
|
// local snapshot independently; proxy JSON must not write over the UI.
|
|
110
135
|
server = createRouterServer(config, {
|
|
111
136
|
log: diagnosticLogs ? undefined : () => {},
|
|
112
137
|
onStatus: event => status?.update(event),
|
|
138
|
+
onDecision: sessionLog ? entry => sessionLog.record(entry) : undefined,
|
|
113
139
|
});
|
|
114
140
|
const address = await listen(server, command === 'claude' ? 0 : config.port);
|
|
115
141
|
const baseUrl = `http://127.0.0.1:${address.port}`;
|
package/docs/reference.md
CHANGED
|
@@ -6,6 +6,9 @@
|
|
|
6
6
|
| --- | --- |
|
|
7
7
|
| `claude-autorouter setup` | Save subscription-mode configuration and a Jev key |
|
|
8
8
|
| `claude-autorouter setup --auth-mode api-key` | Configure Jev and Anthropic API-key billing |
|
|
9
|
+
| `claude-autorouter setup --client-profile auto` | Save a Sonnet/Opus profile compatible with Claude's Auto permission mode |
|
|
10
|
+
| `claude-autorouter setup --stop-hook-block-cap 2` | Opt into a shorter native Stop-hook continuation cap during setup |
|
|
11
|
+
| `claude-autorouter setup --session-log-dir DIR` | Save an opt-in directory for per-session JSONL decision logs |
|
|
9
12
|
| `claude-autorouter setup --evaluator ollama --pull` | Configure the native local evaluator and download its selected model if missing |
|
|
10
13
|
| `claude-autorouter setup --evaluator ollama --ollama-timeout-ms 0 --force` | Save a disabled runtime evaluator deadline |
|
|
11
14
|
| `claude-autorouter setup --force` | Replace an existing user config |
|
|
@@ -42,9 +45,11 @@ For an environment-only subscription launch, set `AUTOROUTER_AUTH_MODE=subscript
|
|
|
42
45
|
| `AUTOROUTER_CONFIG` | see path order above | Explicit user config path |
|
|
43
46
|
| `AUTOROUTER_AUTH_MODE` | `api-key` without saved config; setup selects `subscription` | Authentication mode |
|
|
44
47
|
| `AUTOROUTER_EVALUATOR` | `jev` | `jev` or local `ollama` classification |
|
|
45
|
-
| `AUTOROUTER_CLIENT_PROFILE` | `compatible` | `native` retains
|
|
48
|
+
| `AUTOROUTER_CLIENT_PROFILE` | `compatible` | `native` retains client model/thinking settings; `auto` starts with Sonnet when no explicit model is set and excludes Haiku from task routing |
|
|
46
49
|
| `AUTOROUTER_STATUSLINE` | enabled | `0` retains your existing status line |
|
|
47
50
|
| `AUTOROUTER_DEBUG` | off | `1` enables launcher metadata logs on stderr |
|
|
51
|
+
| `AUTOROUTER_SESSION_LOG_DIR` | off | Write per-session JSONL decision logs with prompt excerpts into this directory; unset or empty disables it |
|
|
52
|
+
| `CLAUDE_CODE_STOP_HOOK_BLOCK_CAP` | unset; Claude currently uses `8` | Optional cap on consecutive Stop/SubagentStop continuations without tool use; `0` disables the cap |
|
|
48
53
|
| `ENABLE_TOOL_SEARCH` | `true` in launcher when unset | Load MCP tool definitions on demand; explicit values are preserved |
|
|
49
54
|
| `AUTOROUTER_HAIKU_MODEL` | `claude-haiku-4-5-20251001` | Routine tier |
|
|
50
55
|
| `AUTOROUTER_SONNET_MODEL` | `claude-sonnet-5` | Standard tier |
|
|
@@ -170,7 +175,7 @@ The following policy applies after classification:
|
|
|
170
175
|
- Claude's local `/goal` command can omit the prompt-ID header. For that path, an exact feedback label matching a preceding expanded `/goal` command keeps the original task and conversation anchor. This narrow text fallback also recognizes Claude's repeated-goal truncation format; arbitrary hook text is not treated as a goal. Feedback remains in the evaluator's recent conversation and the full API request. A new human message becomes the current task normally. The status line shows `prompt pinned` or `goal pinned` when either text-continuation rule applies.
|
|
171
176
|
- Thinking history, fixed-budget thinking, server tools, context management, and other recognized model-specific features preserve the current model. Adaptive thinking, effort, and output above 64K prevent a Haiku choice. Fields are never stripped to force a downgrade.
|
|
172
177
|
- Mid-conversation `system` messages preserve the requested model and pass through unchanged. They do not count as a tool continuation by themselves.
|
|
173
|
-
-
|
|
178
|
+
- Auxiliary requests, including Claude's Auto permission classifier, pass through on their requested model without Jev/Ollama evaluation, token checks, or turn-state changes. Compaction retains its existing model and context-capacity policy. Requests containing `safeguards` also pass through unchanged so the server's safety-review contract is preserved. Token counting and model discovery pass through without classification.
|
|
174
179
|
|
|
175
180
|
The default `compatible` profile starts Claude with Haiku-compatible requests and client-requested thinking disabled. AutoRouter uses adaptive thinking when upgrading these requests to Opus 5/5.5. Starting with 0.3.3, routing to exact `claude-sonnet-5-5` translates disabled thinking to `between_tools`, which skips up-front thinking but permits progress updates between tool calls. At `xhigh`/`max` effort, or when per-message effort differs from the top-level setting (default `high`), it uses adaptive thinking while preserving the effort settings. Token counting uses the same adaptation. Sonnet 5 still accepts disabled thinking and is unchanged. See [Sonnet 5.5 thinking requirements](https://platform.claude.com/docs/en/models/sonnet-5-5/migration-guide).
|
|
176
181
|
|
|
@@ -178,6 +183,28 @@ Explicit native `between_tools` and unknown thinking modes retain the incoming m
|
|
|
178
183
|
|
|
179
184
|
The launcher enables `ENABLE_TOOL_SEARCH=true` when unset. Claude can otherwise disable on-demand MCP discovery when using a custom API address, loading connected-tool schemas into even a fresh conversation. Explicit values, including `false` or `auto:5`, are preserved. Managed settings and always-loaded tools can still affect deferral. See [Claude Code tool search](https://code.claude.com/docs/en/mcp#configure-tool-search).
|
|
180
185
|
|
|
186
|
+
### Auto permission mode
|
|
187
|
+
|
|
188
|
+
The default `compatible` profile starts Claude as Haiku to permit three-tier routing. Claude's Auto permission mode does not support Haiku, even if AutoRouter routes an API request to Sonnet. Eligibility is based on Claude's selected client model. Gateways themselves are supported. See [Claude's Auto-mode requirements](https://code.claude.com/docs/en/permission-modes#eliminate-permission-prompts-with-auto-mode).
|
|
189
|
+
|
|
190
|
+
AutoRouter 0.3.6 adds an `auto` client profile. Launch with:
|
|
191
|
+
|
|
192
|
+
```sh
|
|
193
|
+
claude-autorouter claude --permission-mode auto
|
|
194
|
+
```
|
|
195
|
+
|
|
196
|
+
An explicit `--permission-mode auto` (or `--permission-mode=auto`) selects the profile for that launch, including when your saved profile is `native`. All arguments still go to Claude unchanged. For Auto chosen from Claude's UI or existing settings instead, use:
|
|
197
|
+
|
|
198
|
+
```sh
|
|
199
|
+
env AUTOROUTER_CLIENT_PROFILE=auto claude-autorouter claude
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
The profile defaults the client to the configured Sonnet model and preserves native thinking and explicit model choices. It promotes a routine Haiku routing decision to Sonnet, while other model-feature and conversation-continuity guards still apply. Configured Sonnet/Opus targets must be known Auto-capable models. An explicit client `--model` or `ANTHROPIC_MODEL` can still make Auto unavailable if it selects Haiku or another unsupported model; choose a supported Sonnet or Opus instead.
|
|
203
|
+
|
|
204
|
+
Claude remains responsible for enabling the permission mode and enforcing organization settings, account availability, and tool rules. The profile does not enable Auto by itself or override `disableAutoMode`. AutoRouter does not reproduce Claude's settings precedence to infer a mode from settings files. For new saved configurations, `setup --client-profile auto` persists the profile; for an existing config, change only `AUTOROUTER_CLIENT_PROFILE` to `"auto"` to retain your other settings. `setup --force` replaces the config.
|
|
205
|
+
|
|
206
|
+
**Routing limits:** Claude's permission-classifier requests retain their exact requested model and skip AutoRouter's evaluator. Requests with server-side `safeguards` also retain their model and full body, and safety verdicts stream back unchanged. The status line shows `pass-through · Auto safety` for these requests. Recent Claude versions normally request server-side review through gateways, so Auto-mode sessions can stay on their client-selected model rather than switching between Sonnet and Opus. Ordinary routable requests use the Sonnet/Opus floor, shown as `Auto mode floor` when it changes a Haiku choice. AutoRouter does not disable server review to enable routing. See [server-side classifier review](https://code.claude.com/docs/en/permission-modes#server-side-classifier-review).
|
|
207
|
+
|
|
181
208
|
### Context capacity
|
|
182
209
|
|
|
183
210
|
Context-relevant JSON above 150KB, or image/document blocks including those inside tool results, trigger a token check for otherwise compatible small models. This byte threshold is a trigger, not a token estimate. The check includes system instructions and active tool schemas. Unused deferred schemas are excluded from the trigger; discovered references and historical tool calls add them back. Unknown shapes are counted conservatively.
|
|
@@ -216,6 +243,40 @@ This estimate does not measure subscription bill savings or quota credits. It ex
|
|
|
216
243
|
|
|
217
244
|
Set `AUTOROUTER_STATUSLINE=0` to retain an existing status line. Other `--settings` values are retained in the temporary overlay; source-relative Read/Edit rules keep their anchors. Ambiguous relative sandbox paths cause the launcher to skip the overlay and pass original settings through with a notice. Safe mode disables custom status lines; print mode has no status-line UI. Standalone `serve` does not install one.
|
|
218
245
|
|
|
246
|
+
## Session decision logs
|
|
247
|
+
|
|
248
|
+
AutoRouter 0.3.6 adds optional persistent logs, separate from stderr and the temporary status-line snapshot. Logging is disabled by default. Enable it for one launch:
|
|
249
|
+
|
|
250
|
+
```sh
|
|
251
|
+
env AUTOROUTER_SESSION_LOG_DIR="$HOME/.local/state/claude-autorouter/sessions" \
|
|
252
|
+
claude-autorouter claude
|
|
253
|
+
```
|
|
254
|
+
|
|
255
|
+
The setting also works with `serve` and the Auto-compatible profile. For a new saved configuration, add `--session-log-dir DIR` to `setup`. For an existing config, add `AUTOROUTER_SESSION_LOG_DIR` with an absolute directory path to preserve your other settings. Setup resolves relative paths at setup time; an environment-only relative path resolves from the launch directory. Environment values override saved values; `AUTOROUTER_SESSION_LOG_DIR=''` disables a saved preference for one launch. `doctor` reports the setting without creating log files.
|
|
256
|
+
|
|
257
|
+
Files are named `autorouter-session-*.jsonl`: one file per observed Claude session within a router launch, with a timestamp, random launch identifier, and hashed session identifier in the name. A resumed session in a new launch creates a new file. Requests without a session header share an anonymous file for that launch. Subagents with the same session ID share its file and retain their agent ID. Files are created only when a decision is recorded, and remain after the session ends.
|
|
258
|
+
|
|
259
|
+
Every line is a standalone JSON object. The key fields look like this (additional IDs and routing metadata are included):
|
|
260
|
+
|
|
261
|
+
```json
|
|
262
|
+
{"schema_version":1,"event":"decision","timestamp":"2026-09-30T12:00:00.000Z","session_id":"example-session","prompt_excerpt":"Fix the typo in README.md","prompt_truncated":false,"requested_model":"claude-haiku-4-5-20251001","selected_model":"claude-sonnet-5","decision_latency_ms":214.37,"source":"jev","reason":"classified"}
|
|
263
|
+
```
|
|
264
|
+
|
|
265
|
+
- `prompt_excerpt`: up to 500 Unicode characters of the current human task for main requests or requests without a class header. Tool continuations and recognized `/goal` feedback keep the originating human task. System instructions, standalone reminder blocks, tool results, images, documents, and thinking are omitted. Auxiliary classifiers, compaction, subagents, and workflows have empty excerpts; a new attachment-only task also has an empty excerpt. `prompt_truncated` indicates that text exceeded the limit.
|
|
266
|
+
- `selected_model`: AutoRouter's final selected model after compatibility checks, before the upstream response. It does not confirm which model successfully answered.
|
|
267
|
+
- `decision_latency_ms`: time spent making the routing decision, including evaluator waiting, cache lookup, and any context checks. It excludes Claude generation time and log writing. A `passthrough` entry can be near zero because no evaluator was called.
|
|
268
|
+
- `source` and `reason`: distinguish evaluator choices, cache hits, fallbacks, turn/model constraints, and native safety pass-through. `classified_tier` is included when an evaluator returned a tier, which may differ from the final selected model.
|
|
269
|
+
|
|
270
|
+
Inspect a file with:
|
|
271
|
+
|
|
272
|
+
```sh
|
|
273
|
+
jq -c '{prompt_excerpt, selected_model, decision_latency_ms, source, reason}' /path/to/autorouter-session-EXAMPLE.jsonl
|
|
274
|
+
```
|
|
275
|
+
|
|
276
|
+
There is one record per completed routing decision, including requests whose upstream call later fails. Requests rejected before routing or cancelled before a decision are not recorded. Logs contain user text and are local plaintext: the feature is disabled by default, new directories use `0700`, and files use `0600`. Existing directory permissions are left unchanged. Authentication headers, provider replies, full transcripts, and tool payloads are not logged; text you put directly in a prompt can appear in its excerpt.
|
|
277
|
+
|
|
278
|
+
Writes run asynchronously through a bounded 1 MiB queue and support up to 128 session files per router process. Normal shutdown drains accepted records. Filesystem failure or a queue/session limit disables further logging with one generic warning while routing continues. An abrupt process kill or storage failure can lose unwritten records. Logs are retained without automatic rotation or deletion; manage them in your chosen directory. Log filenames are ignored by this repository and excluded from the npm package.
|
|
279
|
+
|
|
219
280
|
## Troubleshooting
|
|
220
281
|
|
|
221
282
|
Run `claude-autorouter doctor` first. It performs local checks without Claude generations or Jev calls, including local HTTP checks for the selected Ollama model. It cannot establish Anthropic/TypeSafe availability, current quota, or whether a key will be accepted remotely.
|
|
@@ -238,6 +299,26 @@ The launcher is quiet by default. Standalone `serve` logs to stderr by default.
|
|
|
238
299
|
|
|
239
300
|
If the worker reports a blocker but the checker keeps returning “not yet met,” Claude can repeat its answer until its no-progress guard pauses the goal. Repeated tool calls can keep the loop running longer. Use `/goal clear` to end the loop, resolve the external blocker, and set the goal again. For tasks that may require human action, explicitly allow reporting a blocker as an alternative end condition, for example: `/goal Verify the discrepancy against upstream main and run the relevant tests, or report an external authorization blocker and stop.` This changes what counts as completion; AutoRouter does not declare blocked work successful or rewrite goal instructions. See [Claude Code goal evaluation](https://code.claude.com/docs/en/goal#how-evaluation-works).
|
|
240
301
|
|
|
302
|
+
### Shorter Stop-hook loops (opt-in)
|
|
303
|
+
|
|
304
|
+
To return control sooner when a goal keeps reporting the same unmet condition, set Claude's native continuation cap for one launch:
|
|
305
|
+
|
|
306
|
+
```sh
|
|
307
|
+
env CLAUDE_CODE_STOP_HOOK_BLOCK_CAP=2 claude-autorouter claude
|
|
308
|
+
```
|
|
309
|
+
|
|
310
|
+
This permits two consecutive continuations without tool use; the third blocking verdict ends the turn. The goal remains set and unmet, and a new message can resume it. Tool activity resets the counter, so this is not a total turn or request limit and cannot bound repeated failed tool calls. It applies to **all Stop and SubagentStop hooks**, including `/goal`. A smaller cap can pause useful work sooner. Unset preserves Claude's default (currently `8`); **`0` disables the guard**. AutoRouter does not install a Stop hook or change completion verdicts. See [Claude's environment-variable reference](https://code.claude.com/docs/en/env-vars) and [Stop-hook loop behavior](https://code.claude.com/docs/en/hooks#stop).
|
|
311
|
+
|
|
312
|
+
The environment-only command also works on AutoRouter 0.3.4. Saved configuration and the setup flag require AutoRouter 0.3.5 or newer. To save the preference, add the following property to your existing AutoRouter config JSON, preserving its other values:
|
|
313
|
+
|
|
314
|
+
```json
|
|
315
|
+
"CLAUDE_CODE_STOP_HOOK_BLOCK_CAP": "2"
|
|
316
|
+
```
|
|
317
|
+
|
|
318
|
+
For a new configuration, use `claude-autorouter setup --stop-hook-block-cap 2`; the flag works with either evaluator and overrides the environment during setup. Runtime environment values override saved configuration. `setup --force` replaces the entire config, so keep your existing evaluator/authentication options if using it. `doctor` reports the cap when configured. AutoRouter accepts nonnegative safe integers and leaves the setting absent unless you opt in.
|
|
319
|
+
|
|
320
|
+
### Other session issues
|
|
321
|
+
|
|
241
322
|
**After restarting mid-conversation:** turn state is in memory and expires after 30 minutes. Unknown continuations preserve the incoming model. Start a fresh conversation when restarting around signed thinking; AutoRouter cannot reconstruct the prior actual model from lost turn state.
|
|
242
323
|
|
|
243
324
|
Switching models can lose prompt-cache reuse. A cheaper price per token does not guarantee a cheaper or faster task. Only requests using Claude's configured base URL are visible to this proxy. Alternate provider modes such as Bedrock, Vertex, Foundry, Mantle, and `ANTHROPIC_AWS` are unsupported; unset their enable flags before launching. Use ordinary `claude` to bypass routing.
|
package/docs/releasing.md
CHANGED
|
@@ -10,7 +10,11 @@ Version `0.3.3` fixes HTTP 400 errors when a compatible request with disabled th
|
|
|
10
10
|
|
|
11
11
|
Version `0.3.4` preserves the selected model across Stop-hook feedback for the same prompt, including `/goal` commands that omit the gateway prompt-ID header. Recognized goal feedback remains conversation context rather than replacing the human task in evaluator excerpts. Goal-checker verdicts remain unchanged; external authorization blockers can still cause Claude's own goal loop to repeat. See [goal troubleshooting](reference.md#troubleshooting).
|
|
12
12
|
|
|
13
|
-
|
|
13
|
+
Version `0.3.5` adds opt-in saved configuration for Claude's native `CLAUDE_CODE_STOP_HOOK_BLOCK_CAP`, with `setup --stop-hook-block-cap N`, validation, and `doctor` reporting. A value of `2` permits two consecutive Stop-hook continuations without tool use and ends the turn on the third blocking verdict, leaving an unmet goal set. This affects all Stop/SubagentStop hooks, and tool activity resets the counter. Defaults and completion verdicts are unchanged; `0` disables the guard. See [shorter Stop-hook loops](reference.md#shorter-stop-hook-loops-opt-in).
|
|
14
|
+
|
|
15
|
+
Version `0.3.6` fixes Auto permission-mode launches with an Auto-compatible Sonnet/Opus profile while preserving Claude's permission classifiers and server safety-review requests. It also adds optional per-session JSONL decision logs containing a bounded human prompt excerpt, selected model, and routing latency. Logging is disabled by default; set `AUTOROUTER_SESSION_LOG_DIR` or use `setup --session-log-dir DIR` to enable it. See [Auto permission mode](reference.md#auto-permission-mode) and [session decision logs](reference.md#session-decision-logs).
|
|
16
|
+
|
|
17
|
+
The GitHub repository is private. Publishing to npm makes the tarball's runtime source, README, configuration example, license, and shipped documentation public. Model weights, user configuration, credentials, transcripts, session logs, local artifacts, and test fixtures are excluded. Review the archive before the first publication and whenever the package allowlist changes.
|
|
14
18
|
|
|
15
19
|
## What runs automatically
|
|
16
20
|
|
package/package.json
CHANGED
package/src/auth.mjs
CHANGED
|
@@ -1,5 +1,17 @@
|
|
|
1
1
|
export const LOCAL_AUTH_HEADER = 'x-autorouter-token';
|
|
2
2
|
|
|
3
|
+
// Leave Claude's permission selection and policy enforcement to Claude. This
|
|
4
|
+
// only chooses a compatible routing profile for an explicit Auto-mode launch.
|
|
5
|
+
export function clientProfileForLaunch(profile, args) {
|
|
6
|
+
let permissionMode;
|
|
7
|
+
for (let i = 0; i < args.length; i++) {
|
|
8
|
+
if (args[i] === '--') break;
|
|
9
|
+
if (args[i] === '--permission-mode') permissionMode = args[++i];
|
|
10
|
+
else if (args[i].startsWith('--permission-mode=')) permissionMode = args[i].slice('--permission-mode='.length);
|
|
11
|
+
}
|
|
12
|
+
return permissionMode === 'auto' ? 'auto' : profile;
|
|
13
|
+
}
|
|
14
|
+
|
|
3
15
|
export function conflictingProviders(env = process.env) {
|
|
4
16
|
return ['CLAUDE_CODE_USE_BEDROCK', 'CLAUDE_CODE_USE_VERTEX', 'CLAUDE_CODE_USE_FOUNDRY',
|
|
5
17
|
'CLAUDE_CODE_USE_MANTLE', 'CLAUDE_CODE_USE_ANTHROPIC_AWS']
|
|
@@ -16,6 +28,11 @@ export function isSubscriptionRequest(headers) {
|
|
|
16
28
|
|
|
17
29
|
export function buildClaudeEnv(config, baseUrl, parent = process.env) {
|
|
18
30
|
const env = { ...parent, ANTHROPIC_BASE_URL: baseUrl, CLAUDE_CODE_GATEWAY_HINT_HEADERS: '1' };
|
|
31
|
+
// Opt into Claude's own loop guard without installing hooks or altering
|
|
32
|
+
// their verdicts. An unset cap leaves Claude's default in control.
|
|
33
|
+
if (env.CLAUDE_CODE_STOP_HOOK_BLOCK_CAP === undefined && config.stopHookBlockCap !== undefined) {
|
|
34
|
+
env.CLAUDE_CODE_STOP_HOOK_BLOCK_CAP = String(config.stopHookBlockCap);
|
|
35
|
+
}
|
|
19
36
|
// Claude Code otherwise disables MCP tool search for a non-first-party
|
|
20
37
|
// base URL and loads every schema into context. This proxy preserves both
|
|
21
38
|
// tool_reference blocks and their beta headers. Respect explicit choices.
|
|
@@ -25,6 +42,10 @@ export function buildClaudeEnv(config, baseUrl, parent = process.env) {
|
|
|
25
42
|
if (config.clientProfile === 'compatible') {
|
|
26
43
|
env.ANTHROPIC_MODEL = config.models.haiku;
|
|
27
44
|
env.MAX_THINKING_TOKENS = '0';
|
|
45
|
+
} else if (config.clientProfile === 'auto') {
|
|
46
|
+
// Haiku cannot run Auto permission mode. Leave explicit client choices
|
|
47
|
+
// and thinking settings to Claude; it still enforces account/admin gates.
|
|
48
|
+
env.ANTHROPIC_MODEL ??= config.models.sonnet;
|
|
28
49
|
}
|
|
29
50
|
const subscription = config.authMode === 'subscription';
|
|
30
51
|
// Keep unrelated custom headers. Remove stale router credentials and, in
|
package/src/config.mjs
CHANGED
|
@@ -1,6 +1,31 @@
|
|
|
1
1
|
import { DEFAULT_OLLAMA_MODEL, defaultOllamaTimeoutMs, validateOllamaEndpoint, validateOllamaModel } from './ollama-models.mjs';
|
|
2
|
+
import { resolve } from 'node:path';
|
|
2
3
|
|
|
3
4
|
export const TIERS = ['haiku', 'sonnet', 'opus'];
|
|
5
|
+
export const CLIENT_PROFILES = ['compatible', 'native', 'auto'];
|
|
6
|
+
// Exact Anthropic models whose Auto permission-mode support is documented.
|
|
7
|
+
// Custom aliases are not proof of the capabilities of their upstream model.
|
|
8
|
+
const AUTO_MODE_MODELS = new Set([
|
|
9
|
+
'claude-sonnet-4-6', 'claude-sonnet-5', 'claude-sonnet-5-5',
|
|
10
|
+
'claude-opus-4-6', 'claude-opus-4-7', 'claude-opus-4-8', 'claude-opus-5', 'claude-opus-5-5',
|
|
11
|
+
]);
|
|
12
|
+
|
|
13
|
+
export function parseStopHookBlockCap(value, name = 'CLAUDE_CODE_STOP_HOOK_BLOCK_CAP') {
|
|
14
|
+
if (!['string', 'number'].includes(typeof value)
|
|
15
|
+
|| (typeof value === 'string' && !/^[0-9]+$/.test(value.trim()))
|
|
16
|
+
|| !Number.isSafeInteger(Number(value)) || Number(value) < 0) {
|
|
17
|
+
throw new Error(`${name} requires a nonnegative safe integer (0 disables the Stop-hook continuation cap)`);
|
|
18
|
+
}
|
|
19
|
+
return Number(value);
|
|
20
|
+
}
|
|
21
|
+
|
|
22
|
+
export function parseSessionLogDir(value, name = 'AUTOROUTER_SESSION_LOG_DIR') {
|
|
23
|
+
if (value === undefined || value === '') return undefined;
|
|
24
|
+
if (typeof value !== 'string' || !value.trim() || /[\u0000-\u001f\u007f]/.test(value)) {
|
|
25
|
+
throw new Error(`${name} must be a directory path, or an empty string to disable session logging`);
|
|
26
|
+
}
|
|
27
|
+
return resolve(value);
|
|
28
|
+
}
|
|
4
29
|
|
|
5
30
|
function number(env, key, fallback, min, max, integer = true) {
|
|
6
31
|
const value = Number(env[key] ?? fallback);
|
|
@@ -39,8 +64,20 @@ export function readConfig(env = process.env) {
|
|
|
39
64
|
throw new Error('AUTOROUTER_AUTH_MODE must be api-key or subscription');
|
|
40
65
|
}
|
|
41
66
|
const clientProfile = env.AUTOROUTER_CLIENT_PROFILE ?? 'compatible';
|
|
42
|
-
if (!
|
|
43
|
-
throw new Error('AUTOROUTER_CLIENT_PROFILE must be compatible or
|
|
67
|
+
if (!CLIENT_PROFILES.includes(clientProfile)) {
|
|
68
|
+
throw new Error('AUTOROUTER_CLIENT_PROFILE must be compatible, native or auto');
|
|
69
|
+
}
|
|
70
|
+
const models = {
|
|
71
|
+
haiku: env.AUTOROUTER_HAIKU_MODEL ?? 'claude-haiku-4-5-20251001',
|
|
72
|
+
sonnet: env.AUTOROUTER_SONNET_MODEL ?? 'claude-sonnet-5',
|
|
73
|
+
opus: env.AUTOROUTER_OPUS_MODEL ?? 'claude-opus-5-5',
|
|
74
|
+
};
|
|
75
|
+
if (clientProfile === 'auto') {
|
|
76
|
+
for (const tier of ['sonnet', 'opus']) {
|
|
77
|
+
if (!AUTO_MODE_MODELS.has(models[tier])) {
|
|
78
|
+
throw new Error(`AUTOROUTER_${tier.toUpperCase()}_MODEL must be a known Auto-mode-capable Sonnet or Opus model for the auto profile`);
|
|
79
|
+
}
|
|
80
|
+
}
|
|
44
81
|
}
|
|
45
82
|
const upstream = endpoint(env.AUTOROUTER_UPSTREAM_URL ?? 'https://api.anthropic.com', 'AUTOROUTER_UPSTREAM_URL');
|
|
46
83
|
if (authMode === 'subscription' && upstream !== 'https://api.anthropic.com') {
|
|
@@ -50,6 +87,9 @@ export function readConfig(env = process.env) {
|
|
|
50
87
|
evaluator,
|
|
51
88
|
authMode,
|
|
52
89
|
clientProfile,
|
|
90
|
+
sessionLogDir: parseSessionLogDir(env.AUTOROUTER_SESSION_LOG_DIR),
|
|
91
|
+
stopHookBlockCap: env.CLAUDE_CODE_STOP_HOOK_BLOCK_CAP === undefined
|
|
92
|
+
? undefined : parseStopHookBlockCap(env.CLAUDE_CODE_STOP_HOOK_BLOCK_CAP),
|
|
53
93
|
anthropicKey: authMode === 'api-key' ? env.ANTHROPIC_API_KEY : undefined,
|
|
54
94
|
jevKey: env.TYPESAFE_API_KEY,
|
|
55
95
|
localToken: env.AUTOROUTER_TOKEN,
|
|
@@ -61,11 +101,7 @@ export function readConfig(env = process.env) {
|
|
|
61
101
|
ollamaTimeoutMs: number(env, 'AUTOROUTER_OLLAMA_TIMEOUT_MS', defaultOllamaTimeoutMs(ollamaModel), 0, 30000),
|
|
62
102
|
ollamaStateChars: 3000,
|
|
63
103
|
ollamaKeepAlive,
|
|
64
|
-
models
|
|
65
|
-
haiku: env.AUTOROUTER_HAIKU_MODEL ?? 'claude-haiku-4-5-20251001',
|
|
66
|
-
sonnet: env.AUTOROUTER_SONNET_MODEL ?? 'claude-sonnet-5',
|
|
67
|
-
opus: env.AUTOROUTER_OPUS_MODEL ?? 'claude-opus-5-5',
|
|
68
|
-
},
|
|
104
|
+
models,
|
|
69
105
|
port: number(env, 'AUTOROUTER_PORT', 8787, 0, 65535),
|
|
70
106
|
jevTimeoutMs: number(env, 'AUTOROUTER_JEV_TIMEOUT_MS', 1500, 1, 10000),
|
|
71
107
|
tokenCountTimeoutMs: number(env, 'AUTOROUTER_TOKEN_COUNT_TIMEOUT_MS', 1500, 1, 10000),
|
package/src/onboarding.mjs
CHANGED
|
@@ -3,7 +3,7 @@ import { existsSync } from 'node:fs';
|
|
|
3
3
|
import { createInterface } from 'node:readline';
|
|
4
4
|
import { Writable } from 'node:stream';
|
|
5
5
|
import { promisify } from 'node:util';
|
|
6
|
-
import { readConfig, requireKeys } from './config.mjs';
|
|
6
|
+
import { CLIENT_PROFILES, readConfig, requireKeys, parseStopHookBlockCap, parseSessionLogDir } from './config.mjs';
|
|
7
7
|
import { buildClaudeEnv, conflictingProviders, LOCAL_AUTH_HEADER } from './auth.mjs';
|
|
8
8
|
import { getConfigPath, loadUserConfig, saveUserConfig } from './user-config.mjs';
|
|
9
9
|
import { DEFAULT_OLLAMA_MODEL, validateOllamaModel } from './ollama-models.mjs';
|
|
@@ -14,6 +14,10 @@ const execute = promisify(execFile);
|
|
|
14
14
|
export const ollamaDeadlineText = timeoutMs => timeoutMs === 0
|
|
15
15
|
? 'routing deadline disabled' : `routing deadline ${timeoutMs} ms per request`;
|
|
16
16
|
|
|
17
|
+
const stopHookCapText = cap => cap === 0
|
|
18
|
+
? 'Claude Stop/SubagentStop continuation cap disabled (0).'
|
|
19
|
+
: `Claude Stop/SubagentStop cap: ${cap} continuations without tool use.`;
|
|
20
|
+
|
|
17
21
|
// Readline manages editing and restores terminal state; its output is discarded
|
|
18
22
|
// so neither typing nor pasted credentials are echoed to the terminal.
|
|
19
23
|
export async function askSecret(label, { input = process.stdin, output = process.stderr } = {}) {
|
|
@@ -37,13 +41,17 @@ export async function setup(args, {
|
|
|
37
41
|
env = process.env, write = console.log, prompt = askSecret, fetchImpl = fetch, signal,
|
|
38
42
|
} = {}) {
|
|
39
43
|
let authMode = env.AUTOROUTER_AUTH_MODE ?? 'subscription';
|
|
44
|
+
let clientProfile = env.AUTOROUTER_CLIENT_PROFILE ?? 'compatible';
|
|
40
45
|
let evaluator = env.AUTOROUTER_EVALUATOR ?? 'jev';
|
|
41
46
|
let model;
|
|
42
47
|
let ollamaTimeoutMs;
|
|
48
|
+
let stopHookBlockCap = env.CLAUDE_CODE_STOP_HOOK_BLOCK_CAP;
|
|
49
|
+
let sessionLogDir = env.AUTOROUTER_SESSION_LOG_DIR;
|
|
43
50
|
let pull = false;
|
|
44
51
|
let overwrite = false;
|
|
45
52
|
for (let i = 0; i < args.length; i++) {
|
|
46
53
|
if (args[i] === '--auth-mode') authMode = args[++i];
|
|
54
|
+
else if (args[i] === '--client-profile') clientProfile = args[++i];
|
|
47
55
|
else if (args[i] === '--evaluator') evaluator = args[++i];
|
|
48
56
|
else if (args[i] === '--ollama-model') { model = args[++i]; if (model === undefined) throw new Error('--ollama-model requires a model tag'); }
|
|
49
57
|
else if (args[i] === '--ollama-timeout-ms') {
|
|
@@ -53,19 +61,32 @@ export async function setup(args, {
|
|
|
53
61
|
}
|
|
54
62
|
ollamaTimeoutMs = String(Number(value));
|
|
55
63
|
}
|
|
64
|
+
else if (args[i] === '--stop-hook-block-cap') {
|
|
65
|
+
stopHookBlockCap = parseStopHookBlockCap(args[++i], '--stop-hook-block-cap');
|
|
66
|
+
}
|
|
67
|
+
else if (args[i] === '--session-log-dir') {
|
|
68
|
+
sessionLogDir = args[++i];
|
|
69
|
+
if (sessionLogDir === undefined || sessionLogDir.startsWith('--')) throw new Error('--session-log-dir requires a directory path');
|
|
70
|
+
parseSessionLogDir(sessionLogDir, '--session-log-dir');
|
|
71
|
+
}
|
|
56
72
|
else if (args[i] === '--pull') pull = true;
|
|
57
73
|
else if (args[i] === '--force') overwrite = true;
|
|
58
|
-
else throw new Error('Usage: claude-autorouter setup [--auth-mode subscription|api-key] [--evaluator jev|ollama] [--ollama-model TAG] [--ollama-timeout-ms N] [--pull] [--force]');
|
|
74
|
+
else throw new Error('Usage: claude-autorouter setup [--auth-mode subscription|api-key] [--client-profile compatible|native|auto] [--evaluator jev|ollama] [--ollama-model TAG] [--ollama-timeout-ms N] [--stop-hook-block-cap N] [--session-log-dir DIR] [--pull] [--force]');
|
|
59
75
|
}
|
|
60
76
|
if (!['subscription', 'api-key'].includes(authMode)) throw new Error('--auth-mode must be subscription or api-key');
|
|
77
|
+
if (!CLIENT_PROFILES.includes(clientProfile)) throw new Error('--client-profile must be compatible, native or auto');
|
|
61
78
|
if (!['jev', 'ollama'].includes(evaluator)) throw new Error('--evaluator must be jev or ollama');
|
|
62
79
|
if (evaluator !== 'ollama' && (model !== undefined || ollamaTimeoutMs !== undefined || pull)) throw new Error('Ollama model, deadline and download options require --evaluator ollama');
|
|
80
|
+
if (stopHookBlockCap !== undefined) stopHookBlockCap = parseStopHookBlockCap(stopHookBlockCap);
|
|
81
|
+
if (sessionLogDir !== undefined) sessionLogDir = parseSessionLogDir(sessionLogDir) ?? '';
|
|
63
82
|
const path = getConfigPath(env);
|
|
64
83
|
if (!overwrite && existsSync(path)) throw new Error('AutoRouter configuration already exists. Use setup --force to replace it.');
|
|
65
84
|
write(evaluator === 'ollama'
|
|
66
85
|
? 'AutoRouter evaluates bounded prompt excerpts locally with Ollama. Complete requests still go to Anthropic.'
|
|
67
86
|
: 'AutoRouter sends bounded prompt excerpts to TypeSafe Jev and complete requests to Anthropic.');
|
|
68
|
-
const values = { AUTOROUTER_AUTH_MODE: authMode, AUTOROUTER_CLIENT_PROFILE:
|
|
87
|
+
const values = { AUTOROUTER_AUTH_MODE: authMode, AUTOROUTER_CLIENT_PROFILE: clientProfile, AUTOROUTER_EVALUATOR: evaluator };
|
|
88
|
+
if (stopHookBlockCap !== undefined) values.CLAUDE_CODE_STOP_HOOK_BLOCK_CAP = String(stopHookBlockCap);
|
|
89
|
+
if (sessionLogDir !== undefined) values.AUTOROUTER_SESSION_LOG_DIR = sessionLogDir;
|
|
69
90
|
if (evaluator === 'ollama') {
|
|
70
91
|
values.AUTOROUTER_OLLAMA_MODEL = validateOllamaModel(model ?? env.AUTOROUTER_OLLAMA_MODEL ?? DEFAULT_OLLAMA_MODEL);
|
|
71
92
|
for (const key of ['AUTOROUTER_OLLAMA_URL', 'AUTOROUTER_OLLAMA_TIMEOUT_MS', 'AUTOROUTER_OLLAMA_KEEP_ALIVE']) {
|
|
@@ -81,6 +102,9 @@ export async function setup(args, {
|
|
|
81
102
|
}
|
|
82
103
|
const config = readConfig(values);
|
|
83
104
|
requireKeys(config);
|
|
105
|
+
if (config.clientProfile === 'auto') write('Auto-compatible profile: Sonnet/Opus task routing. Claude controls permission-mode availability and safety checks.');
|
|
106
|
+
if (config.stopHookBlockCap !== undefined) write(stopHookCapText(config.stopHookBlockCap));
|
|
107
|
+
if (config.sessionLogDir) write('Session decision logs enabled; files include up to 500 characters of user prompt text per decision.');
|
|
84
108
|
if (evaluator === 'ollama') {
|
|
85
109
|
write(`Local evaluator: ${config.ollamaModel}; ${ollamaDeadlineText(config.ollamaTimeoutMs)}.`);
|
|
86
110
|
const controller = new AbortController();
|
|
@@ -115,6 +139,9 @@ export async function doctor({ env = process.env, write = console.log, run = exe
|
|
|
115
139
|
for (const key of conflictingProviders(effectiveEnv)) {
|
|
116
140
|
report(false, `Unset ${key}; AutoRouter uses the Anthropic Messages API`);
|
|
117
141
|
}
|
|
142
|
+
if (config?.stopHookBlockCap !== undefined) write(stopHookCapText(config.stopHookBlockCap));
|
|
143
|
+
if (config?.clientProfile === 'auto') write('Auto-compatible profile: Sonnet/Opus task routing. Claude controls permission-mode availability and safety checks.');
|
|
144
|
+
if (config?.sessionLogDir) write('Session decision logs enabled; files include up to 500 characters of user prompt text per decision.');
|
|
118
145
|
if (config?.evaluator === 'ollama') {
|
|
119
146
|
write(`Local evaluator: ${config.ollamaModel}; ${ollamaDeadlineText(config.ollamaTimeoutMs)}.`);
|
|
120
147
|
write('Model availability is checked below; classification speed and accuracy are not tested.');
|
package/src/prompt-state.mjs
CHANGED
|
@@ -122,6 +122,54 @@ export function goalFeedbackIndexes(messages) {
|
|
|
122
122
|
return indexes;
|
|
123
123
|
}
|
|
124
124
|
|
|
125
|
+
// Optional diagnostic logs need only the current human text, not evaluator
|
|
126
|
+
// history or non-text placeholders. Bound collection before joining strings,
|
|
127
|
+
// and never visit tool input/output, attachments, or reasoning payloads.
|
|
128
|
+
export function promptExcerpt(body, maxChars = 500) {
|
|
129
|
+
if (!Number.isSafeInteger(maxChars) || maxChars < 0) throw new TypeError('maxChars must be a nonnegative safe integer');
|
|
130
|
+
const messages = body?.messages;
|
|
131
|
+
if (!maxChars || !Array.isArray(messages)) return '';
|
|
132
|
+
let feedbackIndexes;
|
|
133
|
+
for (let index = messages.length - 1; index >= 0; index--) {
|
|
134
|
+
const message = messages[index];
|
|
135
|
+
if (message?.role !== 'user') continue;
|
|
136
|
+
const content = message.content;
|
|
137
|
+
if (Array.isArray(content) && content.some(block => block?.type === 'tool_result')) continue;
|
|
138
|
+
const standalone = typeof content === 'string' ? content
|
|
139
|
+
: Array.isArray(content) && content.length === 1 && content[0]?.type === 'text'
|
|
140
|
+
&& typeof content[0].text === 'string' ? content[0].text : undefined;
|
|
141
|
+
if (standalone?.startsWith('Stop hook feedback:\n[')) {
|
|
142
|
+
feedbackIndexes ??= goalFeedbackIndexes(messages);
|
|
143
|
+
if (feedbackIndexes.has(index)) continue;
|
|
144
|
+
}
|
|
145
|
+
const blocks = typeof content === 'string' ? [{ type: 'text', text: content }]
|
|
146
|
+
: Array.isArray(content) ? content : [];
|
|
147
|
+
const characters = [];
|
|
148
|
+
let nonText = false;
|
|
149
|
+
for (const block of blocks) {
|
|
150
|
+
if (block?.type !== 'text' || typeof block.text !== 'string') {
|
|
151
|
+
nonText = true;
|
|
152
|
+
continue;
|
|
153
|
+
}
|
|
154
|
+
const value = block.text;
|
|
155
|
+
// Also omit complete wrappers in string messages. Unlike classifier
|
|
156
|
+
// input, a diagnostic excerpt should never log a reminder-only turn.
|
|
157
|
+
if (!/\S/.test(value) || isReminderBlock(value)) continue;
|
|
158
|
+
if (characters.length) characters.push('\n');
|
|
159
|
+
for (const character of value) {
|
|
160
|
+
if (characters.length >= maxChars) break;
|
|
161
|
+
characters.push(character);
|
|
162
|
+
}
|
|
163
|
+
if (characters.length >= maxChars) break;
|
|
164
|
+
}
|
|
165
|
+
if (characters.length) return characters.join('').toWellFormed();
|
|
166
|
+
// An image/document-only human turn is a new task with no safe excerpt;
|
|
167
|
+
// do not incorrectly label it with the preceding human task's text.
|
|
168
|
+
if (nonText) return '';
|
|
169
|
+
}
|
|
170
|
+
return '';
|
|
171
|
+
}
|
|
172
|
+
|
|
125
173
|
export function buildState(body, limit = 12000) {
|
|
126
174
|
const messages = body.messages ?? [];
|
|
127
175
|
const feedbackIndexes = goalFeedbackIndexes(messages);
|
package/src/router.mjs
CHANGED
|
@@ -226,6 +226,26 @@ export class Router {
|
|
|
226
226
|
async route(body, { scope = '', signal, requestClass = '', promptId = '', countTokens } = {}) {
|
|
227
227
|
const start = performance.now();
|
|
228
228
|
const c = this.config;
|
|
229
|
+
// Claude owns auxiliary permission checks and the server-side safeguards
|
|
230
|
+
// contract. Never evaluate, adapt, or count these requests: changing
|
|
231
|
+
// their model can change the safety decision or invalidate its context.
|
|
232
|
+
if (requestClass === 'auxiliary' || body.safeguards !== undefined) {
|
|
233
|
+
// A safeguarded main request still produces the next tool turn. Replace
|
|
234
|
+
// any older routing pin with the actual preserved model so a later
|
|
235
|
+
// request that omits safeguards cannot restore that stale model. Side
|
|
236
|
+
// classifiers and compaction never take ownership of the main turn.
|
|
237
|
+
if (!['auxiliary', 'compaction'].includes(requestClass)) {
|
|
238
|
+
const turn = turnInfo(body, scope, promptId);
|
|
239
|
+
if (turn.index >= 0 || promptId) {
|
|
240
|
+
const pin = { model: body.model, requestedModel: body.model };
|
|
241
|
+
this.turns.set(turn.key, pin);
|
|
242
|
+
if (turn.contentKey !== turn.key) this.turns.set(turn.contentKey, pin);
|
|
243
|
+
}
|
|
244
|
+
}
|
|
245
|
+
return { model: body.model, source: 'passthrough',
|
|
246
|
+
reason: requestClass === 'auxiliary' ? 'internal_request' : 'auto_mode_safeguards',
|
|
247
|
+
latency_ms: Math.round((performance.now() - start) * 100) / 100 };
|
|
248
|
+
}
|
|
229
249
|
const hasSystemMessage = body.messages.some(m => m.role === 'system');
|
|
230
250
|
const unknownModel = rank(body.model) < 0 && !Object.values(c.models).includes(body.model);
|
|
231
251
|
const modelSpecificThinking = body.thinking && !['disabled', 'adaptive'].includes(body.thinking.type);
|
|
@@ -243,12 +263,19 @@ export class Router {
|
|
|
243
263
|
// Check suspicious input in parallel with Jev. Byte size only triggers a
|
|
244
264
|
// check: common tool catalogs can be 200KB yet occupy far less than 200K
|
|
245
265
|
// tokens. Tiny requests keep the one-call fast path.
|
|
246
|
-
const earlyCount = !capacityLocked && countTokens && CAPACITY_UPGRADE_MODELS.has(c.models.haiku)
|
|
266
|
+
const earlyCount = c.clientProfile !== 'auto' && !capacityLocked && countTokens && CAPACITY_UPGRADE_MODELS.has(c.models.haiku)
|
|
247
267
|
&& (contextSizeBytes(body, c.models.haiku) > 150000 || hasAttachments)
|
|
248
268
|
? safelyCount(c.models.haiku) : undefined;
|
|
249
269
|
const decision = await this.classify(body, signal);
|
|
250
270
|
let model = c.models[decision.tier];
|
|
251
271
|
let reason = decision.reason;
|
|
272
|
+
// Auto permission mode requires a supported execution model. Retain the
|
|
273
|
+
// evaluator's verdict for observability; stronger compatibility and turn
|
|
274
|
+
// constraints below still decide whether this ordinary choice can apply.
|
|
275
|
+
if (c.clientProfile === 'auto' && decision.tier === 'haiku') {
|
|
276
|
+
model = c.models.sonnet;
|
|
277
|
+
reason = 'auto_mode_floor';
|
|
278
|
+
}
|
|
252
279
|
const turn = turnInfo(body, scope, promptId);
|
|
253
280
|
const promptPin = promptId ? this.turns.get(turn.key) : undefined;
|
|
254
281
|
const turnPin = promptPin ?? this.turns.get(turn.contentKey);
|
|
@@ -269,7 +296,7 @@ export class Router {
|
|
|
269
296
|
|
|
270
297
|
// A tool result belongs to the model that requested it. Do not bounce the
|
|
271
298
|
// agent between models partway through one human turn.
|
|
272
|
-
if (requestClass === 'compaction'
|
|
299
|
+
if (requestClass === 'compaction') preserve(body.model, 'internal_request');
|
|
273
300
|
// Mid-conversation system messages are only supported by certain models.
|
|
274
301
|
// Keep the client's capable model and all message fields (including
|
|
275
302
|
// clear_at, tool changes, and output_config) instead of down-routing.
|
|
@@ -336,7 +363,7 @@ export class Router {
|
|
|
336
363
|
}
|
|
337
364
|
}
|
|
338
365
|
const identifiableUpgrade = capacityUpgraded && (turn.index >= 0 || promptId);
|
|
339
|
-
if ((!turn.continuation || previous || identifiableUpgrade || reason === 'mid_conversation_system') &&
|
|
366
|
+
if ((!turn.continuation || previous || identifiableUpgrade || reason === 'mid_conversation_system') && requestClass !== 'compaction') {
|
|
340
367
|
const pin = { model, requestedModel: body.model };
|
|
341
368
|
this.turns.set(turn.key, pin);
|
|
342
369
|
// Keep the content key too: later human turns carry signed thinking but
|
package/src/server.mjs
CHANGED
|
@@ -7,6 +7,7 @@ import { LOCAL_AUTH_HEADER, isSubscriptionRequest } from './auth.mjs';
|
|
|
7
7
|
import { prepareRequest } from './model-request.mjs';
|
|
8
8
|
import { createResponseObserver } from './response-observer.mjs';
|
|
9
9
|
import { createTokenCounter } from './token-counter.mjs';
|
|
10
|
+
import { promptExcerpt } from './prompt-state.mjs';
|
|
10
11
|
|
|
11
12
|
function cleanHeaders(headers) {
|
|
12
13
|
const blocked = new Set(['host', 'connection', 'keep-alive', 'proxy-authenticate', 'proxy-authorization', 'te', 'trailer', 'transfer-encoding', 'upgrade', 'content-length']);
|
|
@@ -101,7 +102,7 @@ async function forward(url, req, res, body, config, signal, log, status) {
|
|
|
101
102
|
}
|
|
102
103
|
}
|
|
103
104
|
|
|
104
|
-
export function createRouterServer(config, { router = new Router(config), tokenCounter = createTokenCounter(config), log = entry => process.stderr.write(`${JSON.stringify(entry)}\n`), onStatus = () => {} } = {}) {
|
|
105
|
+
export function createRouterServer(config, { router = new Router(config), tokenCounter = createTokenCounter(config), log = entry => process.stderr.write(`${JSON.stringify(entry)}\n`), onStatus = () => {}, onDecision } = {}) {
|
|
105
106
|
if (!config.localToken || config.localToken.length < 16) throw new Error('AUTOROUTER_TOKEN must contain at least 16 characters');
|
|
106
107
|
const server = http.createServer(async (req, res) => {
|
|
107
108
|
const controller = new AbortController();
|
|
@@ -175,6 +176,21 @@ export function createRouterServer(config, { router = new Router(config), tokenC
|
|
|
175
176
|
if (controller.signal.aborted) { status('request_cancelled'); return; }
|
|
176
177
|
const prepared = prepareRequest(parsed, decision.model);
|
|
177
178
|
body = Buffer.from(JSON.stringify(prepared.request));
|
|
179
|
+
// Prompt excerpts go only to this explicit opt-in sink, never to
|
|
180
|
+
// ordinary diagnostics or the status snapshot. Optional logging
|
|
181
|
+
// cannot delay or fail forwarding, including an async sink failure.
|
|
182
|
+
if (onDecision) {
|
|
183
|
+
try {
|
|
184
|
+
const chars = [...(!context.request_class || context.request_class === 'main' ? promptExcerpt(parsed, 501) : '')];
|
|
185
|
+
Promise.resolve(onDecision({
|
|
186
|
+
schema_version: 1, event: 'decision', timestamp: new Date().toISOString(), ...context,
|
|
187
|
+
prompt_excerpt: chars.slice(0, 500).join(''), prompt_truncated: chars.length > 500,
|
|
188
|
+
requested_model: parsed.model, selected_model: decision.model, decision_latency_ms: decision.latency_ms,
|
|
189
|
+
source: decision.source, reason: decision.reason, evaluator: decision.evaluator,
|
|
190
|
+
classified_tier: decision.classified_tier, classifier_error: decision.classifier_error,
|
|
191
|
+
})).catch(() => {});
|
|
192
|
+
} catch {}
|
|
193
|
+
}
|
|
178
194
|
log({ event: 'route', requested_model: parsed.model, ...decision, request_adjustments: prepared.adjustments });
|
|
179
195
|
const { model, source, evaluator, reason, latency_ms, classifier_error, classifier_status, classified_tier, context_check, counted_input_tokens } = decision;
|
|
180
196
|
const pricingValue = (field, allowed, fallback) => prepared.request[field] === undefined ? fallback
|
|
@@ -0,0 +1,171 @@
|
|
|
1
|
+
import { constants } from 'node:fs';
|
|
2
|
+
import fs from 'node:fs/promises';
|
|
3
|
+
import { createHash, randomBytes } from 'node:crypto';
|
|
4
|
+
import { join, resolve } from 'node:path';
|
|
5
|
+
|
|
6
|
+
const MAX_PENDING_BYTES = 1024 * 1024;
|
|
7
|
+
const MAX_SESSIONS = 128;
|
|
8
|
+
const WARNING = 'AutoRouter session logging disabled.';
|
|
9
|
+
const SOURCES = new Set(['jev', 'ollama', 'cache', 'fallback', 'passthrough']);
|
|
10
|
+
const ERRORS = new Set(['timeout', 'http_error', 'invalid_response', 'network_error']);
|
|
11
|
+
const identifier = value => typeof value === 'string' && /^[A-Za-z0-9_.:-]{1,200}$/.test(value) ? value : undefined;
|
|
12
|
+
const model = value => typeof value === 'string' && /^[A-Za-z0-9_.:/-]{1,120}$/.test(value) ? value : undefined;
|
|
13
|
+
const code = value => typeof value === 'string' && /^[a-z][a-z0-9_]{0,79}$/.test(value) ? value : undefined;
|
|
14
|
+
|
|
15
|
+
function excerpt(value) {
|
|
16
|
+
if (typeof value !== 'string') return { text: '', truncated: false };
|
|
17
|
+
let text = '', count = 0;
|
|
18
|
+
for (const character of value) {
|
|
19
|
+
if (count++ === 500) return { text: text.toWellFormed(), truncated: true };
|
|
20
|
+
text += character;
|
|
21
|
+
}
|
|
22
|
+
return { text: text.toWellFormed(), truncated: false };
|
|
23
|
+
}
|
|
24
|
+
|
|
25
|
+
function normalize(entry) {
|
|
26
|
+
if (!entry || typeof entry !== 'object' || Array.isArray(entry) || entry.event !== 'decision') return;
|
|
27
|
+
const requestId = identifier(entry.request_id);
|
|
28
|
+
const selectedModel = model(entry.selected_model);
|
|
29
|
+
const requestedModel = model(entry.requested_model);
|
|
30
|
+
if (!requestId || !selectedModel || !requestedModel) return;
|
|
31
|
+
const anonymous = entry.session_id === undefined || entry.session_id === null || entry.session_id === '';
|
|
32
|
+
const sessionId = anonymous ? undefined : identifier(entry.session_id);
|
|
33
|
+
// An invalid explicit identity must not mix records into an anonymous file.
|
|
34
|
+
if (!anonymous && !sessionId) return;
|
|
35
|
+
const requestClass = code(entry.request_class);
|
|
36
|
+
const foreground = entry.request_class === undefined || entry.request_class === null || entry.request_class === '' || entry.request_class === 'main';
|
|
37
|
+
const prompt = foreground ? excerpt(entry.prompt_excerpt) : { text: '', truncated: false };
|
|
38
|
+
const timestamp = typeof entry.timestamp === 'string' && /^\d{4}-\d\d-\d\dT\d\d:\d\d:\d\d\.\d{3}Z$/.test(entry.timestamp)
|
|
39
|
+
&& Number.isFinite(Date.parse(entry.timestamp)) ? entry.timestamp : new Date().toISOString();
|
|
40
|
+
const row = { schema_version: 1, event: 'decision', timestamp, request_id: requestId,
|
|
41
|
+
...(sessionId ? { session_id: sessionId } : {}),
|
|
42
|
+
prompt_excerpt: prompt.text, prompt_truncated: prompt.truncated || (foreground && entry.prompt_truncated === true),
|
|
43
|
+
requested_model: requestedModel, selected_model: selectedModel };
|
|
44
|
+
for (const field of ['agent_id', 'prompt_id']) {
|
|
45
|
+
const value = identifier(entry[field]);
|
|
46
|
+
if (value) row[field] = value;
|
|
47
|
+
}
|
|
48
|
+
if (requestClass) row.request_class = requestClass;
|
|
49
|
+
if (typeof entry.decision_latency_ms === 'number' && Number.isFinite(entry.decision_latency_ms)
|
|
50
|
+
&& entry.decision_latency_ms >= 0) row.decision_latency_ms = entry.decision_latency_ms;
|
|
51
|
+
if (SOURCES.has(entry.source)) row.source = entry.source;
|
|
52
|
+
const reason = code(entry.reason);
|
|
53
|
+
if (reason) row.reason = reason;
|
|
54
|
+
if (['jev', 'ollama'].includes(entry.evaluator)) row.evaluator = entry.evaluator;
|
|
55
|
+
if (['haiku', 'sonnet', 'opus'].includes(entry.classified_tier)) row.classified_tier = entry.classified_tier;
|
|
56
|
+
if (ERRORS.has(entry.classifier_error)) row.classifier_error = entry.classifier_error;
|
|
57
|
+
return { sessionKey: sessionId ? `session:${sessionId}` : 'anonymous', row };
|
|
58
|
+
}
|
|
59
|
+
|
|
60
|
+
// Inference never waits on this writer. Each launch creates new files; only
|
|
61
|
+
// accepted, bounded JSON lines are retained until the serialized writer drains.
|
|
62
|
+
export async function createSessionLog(directory, { warn = () => {} } = {}) {
|
|
63
|
+
let accepting = true, failed = false, warned = false, pendingBytes = 0;
|
|
64
|
+
let root, directoryIdentity, pump, closePromise;
|
|
65
|
+
const sessions = new Map(), queue = [];
|
|
66
|
+
const launch = `${new Date().toISOString().replace(/[-:.]/g, '')}-${randomBytes(12).toString('hex')}`;
|
|
67
|
+
const disable = () => {
|
|
68
|
+
accepting = false;
|
|
69
|
+
if (warned) return;
|
|
70
|
+
warned = true;
|
|
71
|
+
try { Promise.resolve(warn(WARNING)).catch(() => {}); } catch {}
|
|
72
|
+
};
|
|
73
|
+
async function checkDirectory() {
|
|
74
|
+
const current = await fs.lstat(root);
|
|
75
|
+
if (!current.isDirectory() || current.isSymbolicLink()
|
|
76
|
+
|| (directoryIdentity && (current.dev !== directoryIdentity.dev || current.ino !== directoryIdentity.ino))) {
|
|
77
|
+
throw new Error('Invalid session log directory');
|
|
78
|
+
}
|
|
79
|
+
return current;
|
|
80
|
+
}
|
|
81
|
+
try {
|
|
82
|
+
if (typeof directory !== 'string' || !directory.trim() || typeof constants.O_NOFOLLOW !== 'number') throw new Error('Invalid session log directory');
|
|
83
|
+
root = resolve(directory);
|
|
84
|
+
try { await checkDirectory(); }
|
|
85
|
+
catch (error) { if (error.code !== 'ENOENT') throw error; }
|
|
86
|
+
await fs.mkdir(root, { recursive: true, mode: 0o700 });
|
|
87
|
+
directoryIdentity = await checkDirectory();
|
|
88
|
+
} catch { failed = true; disable(); }
|
|
89
|
+
|
|
90
|
+
async function handleFor(session) {
|
|
91
|
+
if (session.handle) return session.handle;
|
|
92
|
+
await checkDirectory();
|
|
93
|
+
const handle = await fs.open(session.path, constants.O_WRONLY | constants.O_APPEND | constants.O_CREAT | constants.O_EXCL | constants.O_NOFOLLOW, 0o600);
|
|
94
|
+
// Track immediately, including when a subsequent check fails, so shutdown
|
|
95
|
+
// always closes the descriptor. Never reopen or follow an existing path.
|
|
96
|
+
session.handle = handle;
|
|
97
|
+
const stat = await handle.stat();
|
|
98
|
+
if (!stat.isFile() || stat.nlink !== 1) throw new Error('Invalid session log file');
|
|
99
|
+
await handle.chmod(0o600);
|
|
100
|
+
await checkDirectory();
|
|
101
|
+
return handle;
|
|
102
|
+
}
|
|
103
|
+
async function drain() {
|
|
104
|
+
let index = 0;
|
|
105
|
+
try {
|
|
106
|
+
while (index < queue.length && !failed) {
|
|
107
|
+
const item = queue[index];
|
|
108
|
+
queue[index++] = undefined;
|
|
109
|
+
const handle = await handleFor(item.session);
|
|
110
|
+
await handle.writeFile(item.line);
|
|
111
|
+
pendingBytes -= item.bytes;
|
|
112
|
+
// A steady producer can keep the queue nonempty indefinitely. Bound
|
|
113
|
+
// processed slots as well as the pending strings they once held.
|
|
114
|
+
if (index >= 256) { queue.splice(0, index); index = 0; }
|
|
115
|
+
}
|
|
116
|
+
} catch {
|
|
117
|
+
failed = true;
|
|
118
|
+
disable();
|
|
119
|
+
} finally {
|
|
120
|
+
queue.length = 0;
|
|
121
|
+
pendingBytes = 0;
|
|
122
|
+
}
|
|
123
|
+
}
|
|
124
|
+
function schedule() {
|
|
125
|
+
if (pump || failed || !queue.length) return;
|
|
126
|
+
pump = Promise.resolve().then(drain).finally(() => {
|
|
127
|
+
pump = undefined;
|
|
128
|
+
// A record can arrive between drain resolving and this continuation.
|
|
129
|
+
// Keep it scheduled even if shutdown has already stopped new records.
|
|
130
|
+
schedule();
|
|
131
|
+
});
|
|
132
|
+
}
|
|
133
|
+
function record(entry) {
|
|
134
|
+
if (!accepting) return false;
|
|
135
|
+
try {
|
|
136
|
+
const normalized = normalize(entry);
|
|
137
|
+
if (!normalized) return false;
|
|
138
|
+
const line = `${JSON.stringify(normalized.row)}\n`;
|
|
139
|
+
const bytes = Buffer.byteLength(line);
|
|
140
|
+
let session = sessions.get(normalized.sessionKey);
|
|
141
|
+
if (pendingBytes + bytes > MAX_PENDING_BYTES || (!session && sessions.size >= MAX_SESSIONS)) {
|
|
142
|
+
disable();
|
|
143
|
+
return false;
|
|
144
|
+
}
|
|
145
|
+
if (!session) {
|
|
146
|
+
const sessionHash = createHash('sha256').update(normalized.sessionKey).digest('hex');
|
|
147
|
+
session = { path: join(root, `autorouter-session-${launch}-${sessionHash}.jsonl`) };
|
|
148
|
+
sessions.set(normalized.sessionKey, session);
|
|
149
|
+
}
|
|
150
|
+
queue.push({ session, line, bytes });
|
|
151
|
+
pendingBytes += bytes;
|
|
152
|
+
schedule();
|
|
153
|
+
return true;
|
|
154
|
+
} catch { return false; }
|
|
155
|
+
}
|
|
156
|
+
function close() {
|
|
157
|
+
if (!closePromise) {
|
|
158
|
+
accepting = false;
|
|
159
|
+
closePromise = (async () => {
|
|
160
|
+
while (pump) await pump;
|
|
161
|
+
for (const session of sessions.values()) {
|
|
162
|
+
if (!session.handle) continue;
|
|
163
|
+
try { await session.handle.close(); } catch { disable(); }
|
|
164
|
+
session.handle = undefined;
|
|
165
|
+
}
|
|
166
|
+
})().catch(() => { disable(); });
|
|
167
|
+
}
|
|
168
|
+
return closePromise;
|
|
169
|
+
}
|
|
170
|
+
return { record, close };
|
|
171
|
+
}
|
package/src/status-state.mjs
CHANGED
|
@@ -102,7 +102,7 @@ export function createStatusState(options = {}) {
|
|
|
102
102
|
switch (event.event) {
|
|
103
103
|
case 'route': {
|
|
104
104
|
const fields = { requested_model: modelName(event.requested_model), selected_model: modelName(event.model),
|
|
105
|
-
source: ['jev', 'ollama', 'cache', 'fallback'].includes(event.source) ? event.source : undefined,
|
|
105
|
+
source: ['jev', 'ollama', 'cache', 'fallback', 'passthrough'].includes(event.source) ? event.source : undefined,
|
|
106
106
|
evaluator: ['jev', 'ollama'].includes(event.evaluator) ? event.evaluator : undefined,
|
|
107
107
|
reason: code(event.reason), latency_ms: latency(event.latency_ms),
|
|
108
108
|
classified_tier: ['haiku', 'sonnet', 'opus'].includes(event.classified_tier) ? event.classified_tier : undefined,
|
package/src/statusline.mjs
CHANGED
|
@@ -6,6 +6,7 @@ const REASONS = {
|
|
|
6
6
|
model_specific_features: 'model features', large_or_multimodal_request: 'large request',
|
|
7
7
|
context_capacity: 'large context',
|
|
8
8
|
internal_request: 'internal request', unknown_model: 'custom model', low_confidence: 'low confidence',
|
|
9
|
+
auto_mode_floor: 'Auto mode floor', auto_mode_safeguards: 'Auto safety',
|
|
9
10
|
};
|
|
10
11
|
const CLASSIFIER_ERRORS = {
|
|
11
12
|
timeout: 'timeout', http_error: 'HTTP error', invalid_response: 'invalid response', network_error: 'network error',
|
|
@@ -128,14 +129,14 @@ export function renderStatusLine(input, snapshot, { now = Date.now(), color = tr
|
|
|
128
129
|
? `error ${state.status}` : errorType ? `error ${errorType}` : 'error';
|
|
129
130
|
|
|
130
131
|
const details = [];
|
|
131
|
-
const source = ['jev', 'ollama', 'cache', 'fallback'].includes(state?.source) ? state.source : undefined;
|
|
132
|
+
const source = ['jev', 'ollama', 'cache', 'fallback', 'passthrough'].includes(state?.source) ? state.source : undefined;
|
|
132
133
|
const evaluator = ['jev', 'ollama'].includes(state?.evaluator) ? state.evaluator : undefined;
|
|
133
134
|
const evaluatorLabel = evaluator === 'ollama' ? 'Ollama' : evaluator === 'jev' ? 'Jev' : '';
|
|
134
135
|
let fallbackCause = '';
|
|
135
136
|
let fallbackPhase = '';
|
|
136
137
|
let compactFallback = false;
|
|
137
138
|
if (source) {
|
|
138
|
-
const sourceLabel = source === 'jev' ? 'Jev' : source === 'ollama' ? 'Ollama'
|
|
139
|
+
const sourceLabel = source === 'passthrough' ? 'pass-through' : source === 'jev' ? 'Jev' : source === 'ollama' ? 'Ollama'
|
|
139
140
|
: evaluatorLabel ? `${evaluatorLabel} ${source}` : source;
|
|
140
141
|
const timing = Number.isFinite(state.latency_ms) && state.latency_ms >= 0 ? ` ${Math.round(Math.min(state.latency_ms, 999999))}ms` : '';
|
|
141
142
|
const classified = ['haiku', 'sonnet', 'opus'].includes(state.classified_tier) ? state.classified_tier : undefined;
|
package/src/user-config.mjs
CHANGED
|
@@ -13,7 +13,8 @@ const CONFIG_KEYS = new Set([
|
|
|
13
13
|
'AUTOROUTER_HAIKU_MODEL', 'AUTOROUTER_SONNET_MODEL', 'AUTOROUTER_OPUS_MODEL',
|
|
14
14
|
'AUTOROUTER_PORT', 'AUTOROUTER_JEV_TIMEOUT_MS', 'AUTOROUTER_TOKEN_COUNT_TIMEOUT_MS',
|
|
15
15
|
'AUTOROUTER_MIN_CONFIDENCE', 'AUTOROUTER_STATUSLINE', 'AUTOROUTER_DEBUG',
|
|
16
|
-
'
|
|
16
|
+
'AUTOROUTER_SESSION_LOG_DIR',
|
|
17
|
+
'ENABLE_TOOL_SEARCH', 'CLAUDE_CODE_STOP_HOOK_BLOCK_CAP',
|
|
17
18
|
'AUTOROUTER_EVALUATOR', 'AUTOROUTER_OLLAMA_URL', 'AUTOROUTER_OLLAMA_MODEL',
|
|
18
19
|
'AUTOROUTER_OLLAMA_TIMEOUT_MS', 'AUTOROUTER_OLLAMA_KEEP_ALIVE',
|
|
19
20
|
]);
|