claude-autorouter 0.3.2 → 0.3.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/docs/reference.md CHANGED
@@ -166,12 +166,15 @@ The following policy applies after classification:
166
166
 
167
167
  - Jev's 1,500 ms deadline covers the response body and has no retry. Successful calls return immediately. Timeouts, HTTP errors, and invalid responses fall back to Sonnet or retain an existing stronger model.
168
168
  - Jev confidence below 0.75 prevents a downgrade below Sonnet or the requested tier. Ollama returns a tier without calibrated confidence; its failure handling and compatibility guards still apply.
169
- - Tool continuations retain the model chosen at the start of the human turn. Session, agent, and prompt headers identify turns; normalized conversation content provides a fallback. Moving prompt-cache markers does not create a new turn.
169
+ - Tool continuations retain the model chosen at the start of the human turn. Session, agent, and prompt headers identify turns; normalized conversation content provides a fallback. Text feedback from a Stop hook also retains the model when it serves the same gateway prompt ID and the client has not changed its requested model, subject to capability and context checks. Moving prompt-cache markers does not create a new turn.
170
+ - Claude's local `/goal` command can omit the prompt-ID header. For that path, an exact feedback label matching a preceding expanded `/goal` command keeps the original task and conversation anchor. This narrow text fallback also recognizes Claude's repeated-goal truncation format; arbitrary hook text is not treated as a goal. Feedback remains in the evaluator's recent conversation and the full API request. A new human message becomes the current task normally. The status line shows `prompt pinned` or `goal pinned` when either text-continuation rule applies.
170
171
  - Thinking history, fixed-budget thinking, server tools, context management, and other recognized model-specific features preserve the current model. Adaptive thinking, effort, and output above 64K prevent a Haiku choice. Fields are never stripped to force a downgrade.
171
172
  - Mid-conversation `system` messages preserve the requested model and pass through unchanged. They do not count as a tool continuation by themselves.
172
173
  - Recognized compaction and auxiliary requests otherwise preserve their requested model. Token counting and model discovery pass through without classification.
173
174
 
174
- The default `compatible` profile starts Claude with Haiku-compatible requests and client-requested thinking disabled. When an upgrade to Opus 5/5.5 requires adaptive thinking, AutoRouter enables it. `AUTOROUTER_CLIENT_PROFILE=native` preserves normal client settings, which can constrain routing. An explicit Claude `--model` argument overrides the starting model, but `/model` and `--model` are requested models, not locks on the routed result.
175
+ The default `compatible` profile starts Claude with Haiku-compatible requests and client-requested thinking disabled. AutoRouter uses adaptive thinking when upgrading these requests to Opus 5/5.5. Starting with 0.3.3, routing to exact `claude-sonnet-5-5` translates disabled thinking to `between_tools`, which skips up-front thinking but permits progress updates between tool calls. At `xhigh`/`max` effort, or when per-message effort differs from the top-level setting (default `high`), it uses adaptive thinking while preserving the effort settings. Token counting uses the same adaptation. Sonnet 5 still accepts disabled thinking and is unchanged. See [Sonnet 5.5 thinking requirements](https://platform.claude.com/docs/en/models/sonnet-5-5/migration-guide).
176
+
177
+ Explicit native `between_tools` and unknown thinking modes retain the incoming model on new human turns. Signed thinking blocks pass through unchanged and existing tool turns retain their model pin. `AUTOROUTER_CLIENT_PROFILE=native` preserves normal client settings, which can constrain routing. An explicit Claude `--model` argument overrides the starting model, but `/model` and `--model` are requested models, not locks on the routed result. Native same-model requests and unknown model aliases are not rewritten; clients must use settings supported by that model.
175
178
 
176
179
  The launcher enables `ENABLE_TOOL_SEARCH=true` when unset. Claude can otherwise disable on-demand MCP discovery when using a custom API address, loading connected-tool schemas into even a fresh conversation. Explicit values, including `false` or `auto:5`, are preserved. Managed settings and always-loaded tools can still affect deferral. See [Claude Code tool search](https://code.claude.com/docs/en/mcp#configure-tool-search).
177
180
 
@@ -231,6 +234,10 @@ The launcher is quiet by default. Standalone `serve` logs to stderr by default.
231
234
 
232
235
  **Claude says Haiku while AutoRouter says Opus:** the built-in model label is Claude's starting/requested model. The AutoRouter confirmed-model label comes from Anthropic. Claude's own token-cost estimate can likewise be attributed to the requested model.
233
236
 
237
+ **`/goal` repeats a question or says it is blocked:** Claude Code runs its own completion checker after each turn, separately from AutoRouter's Jev/Ollama classifier. AutoRouter preserves the requested model for that auxiliary check and forwards its verdict unchanged. The check cannot authorize GitHub SAML, approve a tool, or resolve an external dependency. A tool's authentication error is also different from a Claude API authentication failure.
238
+
239
+ If the worker reports a blocker but the checker keeps returning “not yet met,” Claude can repeat its answer until its no-progress guard pauses the goal. Repeated tool calls can keep the loop running longer. Use `/goal clear` to end the loop, resolve the external blocker, and set the goal again. For tasks that may require human action, explicitly allow reporting a blocker as an alternative end condition, for example: `/goal Verify the discrepancy against upstream main and run the relevant tests, or report an external authorization blocker and stop.` This changes what counts as completion; AutoRouter does not declare blocked work successful or rewrite goal instructions. See [Claude Code goal evaluation](https://code.claude.com/docs/en/goal#how-evaluation-works).
240
+
234
241
  **After restarting mid-conversation:** turn state is in memory and expires after 30 minutes. Unknown continuations preserve the incoming model. Start a fresh conversation when restarting around signed thinking; AutoRouter cannot reconstruct the prior actual model from lost turn state.
235
242
 
236
243
  Switching models can lose prompt-cache reuse. A cheaper price per token does not guarantee a cheaper or faster task. Only requests using Claude's configured base URL are visible to this proxy. Alternate provider modes such as Bedrock, Vertex, Foundry, Mantle, and `ANTHROPIC_AWS` are unsupported; unset their enable flags before launching. Use ordinary `claude` to bypass routing.
package/docs/releasing.md CHANGED
@@ -6,6 +6,10 @@ Version `0.3.1` replaces the old Ollama chat evaluator and Qwen presets with the
6
6
 
7
7
  Version `0.3.2` fixes local timeout fallbacks with model-specific deadlines, removes Claude executor instructions from local evaluator excerpts, and keeps fallback causes visible in compact status lines. It also adds `AUTOROUTER_OLLAMA_TIMEOUT_MS=0` and `setup --ollama-timeout-ms 0` to disable the runtime evaluation deadline while preserving caller cancellation and the separate startup warmup limit. Existing explicit timeout settings still override the defaults; Jev is unchanged.
8
8
 
9
+ Version `0.3.3` fixes HTTP 400 errors when a compatible request with disabled thinking is routed to Sonnet 5.5. Inference and token counting translate that setting to `between_tools`, or adaptive thinking when effort settings require it. Model defaults are unchanged; select Sonnet 5.5 with `AUTOROUTER_SONNET_MODEL=claude-sonnet-5-5`.
10
+
11
+ Version `0.3.4` preserves the selected model across Stop-hook feedback for the same prompt, including `/goal` commands that omit the gateway prompt-ID header. Recognized goal feedback remains conversation context rather than replacing the human task in evaluator excerpts. Goal-checker verdicts remain unchanged; external authorization blockers can still cause Claude's own goal loop to repeat. See [goal troubleshooting](reference.md#troubleshooting).
12
+
9
13
  The GitHub repository is private. Publishing to npm makes the tarball's runtime source, README, configuration example, license, and shipped documentation public. Model weights, user configuration, credentials, transcripts, local artifacts, and test fixtures are excluded. Review the archive before the first publication and whenever the package allowlist changes.
10
14
 
11
15
  ## What runs automatically
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "claude-autorouter",
3
- "version": "0.3.2",
3
+ "version": "0.3.4",
4
4
  "license": "Apache-2.0",
5
5
  "type": "module",
6
6
  "description": "A Claude Code model router with Jev and local Ollama System One evaluators",
@@ -1,13 +1,28 @@
1
- // These model IDs require adaptive thinking even when the client's source
2
- // model supports disabling it. Keep the adaptation explicit and narrow.
3
- const ALWAYS_ADAPTIVE = new Set(['claude-opus-5', 'claude-opus-5-5']);
1
+ // Keep adaptations explicit: models in the same family can have different
2
+ // thinking contracts. Preserve the existing adaptive Opus upgrade behavior.
3
+ const ADAPTIVE_TARGETS = new Set(['claude-opus-5', 'claude-opus-5-5']);
4
+
5
+ function sonnetNeedsAdaptive(body) {
6
+ const effort = body.output_config?.effort ?? 'high';
7
+ // Sonnet 5.5's between_tools mode accepts high effort or below, and cannot
8
+ // change effort through per-message overrides. Preserve those overrides by
9
+ // selecting adaptive thinking instead of dropping or lowering the effort.
10
+ return ['xhigh', 'max'].includes(effort) || body.messages?.some(message =>
11
+ message.output_config?.effort !== undefined && message.output_config.effort !== effort);
12
+ }
4
13
 
5
14
  export function prepareRequest(body, model) {
6
15
  const request = { ...body, model };
7
16
  const adjustments = [];
8
- if (model !== body.model && ALWAYS_ADAPTIVE.has(model) && body.thinking?.type === 'disabled') {
9
- request.thinking = { type: 'adaptive' };
10
- adjustments.push('adaptive_thinking_required');
17
+ if (model !== body.model && body.thinking?.type === 'disabled') {
18
+ if (model === 'claude-sonnet-5-5') {
19
+ const type = sonnetNeedsAdaptive(body) ? 'adaptive' : 'between_tools';
20
+ request.thinking = { type };
21
+ adjustments.push(type === 'adaptive' ? 'adaptive_thinking_required' : 'between_tools_thinking_required');
22
+ } else if (ADAPTIVE_TARGETS.has(model)) {
23
+ request.thinking = { type: 'adaptive' };
24
+ adjustments.push('adaptive_thinking_required');
25
+ }
11
26
  }
12
27
  return { request, adjustments };
13
28
  }
@@ -76,12 +76,60 @@ function humanTask(message) {
76
76
  return contentText(message.content, true);
77
77
  }
78
78
 
79
+ // Claude Code emits /goal Stop-hook feedback as user text. Recognize only its
80
+ // observed wrapper and a condition established by an earlier expanded command;
81
+ // ordinary messages mentioning hooks remain human tasks. This does not alter
82
+ // message content, and the feedback remains available as classifier history.
83
+ export function goalFeedbackIndexes(messages) {
84
+ const indexes = new Set();
85
+ if (!Array.isArray(messages)) return indexes;
86
+ let condition;
87
+ let shortCondition;
88
+ let sawFullFeedback = false;
89
+ for (let index = 0; index < messages.length; index++) {
90
+ const message = messages[index];
91
+ if (message?.role !== 'user') continue;
92
+ const command = /^\s*<command-name>\/goal<\/command-name>\s*<command-message>goal<\/command-message>\s*<command-args>([\s\S]*?)<\/command-args>(?:\s|$)/.exec(humanTask(message));
93
+ if (command) {
94
+ const value = command[1].trim();
95
+ // /goal with no arguments is a status query, not a replacement goal.
96
+ if (!value) continue;
97
+ condition = value.length <= 4000 && !/^(?:clear|stop|off|reset|none|cancel)$/i.test(value) ? value : undefined;
98
+ shortCondition = undefined;
99
+ sawFullFeedback = false;
100
+ if (condition?.length > 500) {
101
+ let prefix = condition.slice(0, 500);
102
+ const last = prefix.charCodeAt(prefix.length - 1);
103
+ if (last >= 0xd800 && last <= 0xdbff) prefix = prefix.slice(0, -1);
104
+ shortCondition = `${prefix}… [+${condition.length - prefix.length} chars]`;
105
+ }
106
+ continue;
107
+ }
108
+ if (!condition || messages[index - 1]?.role !== 'assistant') continue;
109
+ const text = typeof message.content === 'string' ? message.content
110
+ : Array.isArray(message.content) && message.content.length === 1 && message.content[0]?.type === 'text'
111
+ && typeof message.content[0].text === 'string' ? message.content[0].text : undefined;
112
+ if (text === undefined) continue;
113
+ const matches = label => {
114
+ const prefix = `Stop hook feedback:\n[${label}]: `;
115
+ return text.startsWith(prefix) && Boolean(text.slice(prefix.length).trim());
116
+ };
117
+ if (matches(condition)) {
118
+ indexes.add(index);
119
+ sawFullFeedback = true;
120
+ } else if (sawFullFeedback && shortCondition && matches(shortCondition)) indexes.add(index);
121
+ }
122
+ return indexes;
123
+ }
124
+
79
125
  export function buildState(body, limit = 12000) {
80
126
  const messages = body.messages ?? [];
127
+ const feedbackIndexes = goalFeedbackIndexes(messages);
81
128
  let firstTask = '';
82
129
  let currentTask = '';
83
130
  let currentIndex = -1;
84
131
  for (let index = 0; index < messages.length; index++) {
132
+ if (feedbackIndexes.has(index)) continue;
85
133
  const task = humanTask(messages[index]);
86
134
  if (!task.trim()) continue;
87
135
  if (!firstTask) firstTask = task;
package/src/router.mjs CHANGED
@@ -1,6 +1,6 @@
1
1
  import { createHash } from 'node:crypto';
2
2
  import { TIERS } from './config.mjs';
3
- import { buildState } from './prompt-state.mjs';
3
+ import { buildState, goalFeedbackIndexes } from './prompt-state.mjs';
4
4
  import { buildOllamaState, evaluateOllama } from './ollama-evaluator.mjs';
5
5
  export { buildState } from './prompt-state.mjs';
6
6
 
@@ -134,10 +134,11 @@ function turnContent(content) {
134
134
 
135
135
  function turnInfo(body, scope, promptId = '') {
136
136
  const messages = body.messages ?? [];
137
+ const feedback = goalFeedbackIndexes(messages);
137
138
  let index = -1;
138
139
  for (let i = messages.length - 1; i >= 0; i--) {
139
140
  const message = messages[i];
140
- if (message.role === 'user' && !(Array.isArray(message.content) && message.content.some(b => b.type === 'tool_result'))) {
141
+ if (message.role === 'user' && !feedback.has(i) && !(Array.isArray(message.content) && message.content.some(b => b.type === 'tool_result'))) {
141
142
  index = i; break;
142
143
  }
143
144
  }
@@ -147,6 +148,7 @@ function turnInfo(body, scope, promptId = '') {
147
148
  index,
148
149
  key: promptId ? hash(['prompt', scope, promptId]) : contentKey,
149
150
  contentKey,
151
+ goalFeedback: feedback.has(messages.length - 1),
150
152
  // Claude Code can append turn-scoped system instructions after the human
151
153
  // prompt. These do not start an assistant/tool continuation.
152
154
  continuation: index < 0 || messages.slice(index + 1).some(m => m.role !== 'system'),
@@ -226,7 +228,8 @@ export class Router {
226
228
  const c = this.config;
227
229
  const hasSystemMessage = body.messages.some(m => m.role === 'system');
228
230
  const unknownModel = rank(body.model) < 0 && !Object.values(c.models).includes(body.model);
229
- const modelSpecificFeatures = body.thinking?.type === 'enabled' || body.context_management || body.speed || body.container || body.mcp_servers || body.tools?.some(t => t.type && t.type !== 'custom');
231
+ const modelSpecificThinking = body.thinking && !['disabled', 'adaptive'].includes(body.thinking.type);
232
+ const modelSpecificFeatures = modelSpecificThinking || body.context_management || body.speed || body.container || body.mcp_servers || body.tools?.some(t => t.type && t.type !== 'custom');
230
233
  const thinkingHistory = hasContentBlock(body, ['thinking', 'redacted_thinking']);
231
234
  const knownSourceModel = LARGE_CONTEXT_MODELS.has(body.model) || CAPACITY_UPGRADE_MODELS.has(body.model);
232
235
  const capacityLocked = hasSystemMessage || unknownModel || !knownSourceModel || modelSpecificFeatures || thinkingHistory;
@@ -247,11 +250,15 @@ export class Router {
247
250
  let model = c.models[decision.tier];
248
251
  let reason = decision.reason;
249
252
  const turn = turnInfo(body, scope, promptId);
250
- let previous = this.turns.get(turn.key) ?? this.turns.get(turn.contentKey);
253
+ const promptPin = promptId ? this.turns.get(turn.key) : undefined;
254
+ const turnPin = promptPin ?? this.turns.get(turn.contentKey);
255
+ let previous = turnPin?.model;
256
+ const textTurn = !turn.continuation || turn.goalFeedback;
257
+ const textPin = promptPin ?? (!promptId && turn.goalFeedback ? turnPin : undefined);
251
258
  // A new human prompt can still carry signed thinking from the preceding
252
259
  // turn. Recover that turn's actual routed model when it is known.
253
260
  if (!turn.continuation && body.messages.length > 1 && !previous) {
254
- previous = this.turns.get(turnInfo({ ...body, messages: body.messages.slice(0, turn.index) }, scope).key);
261
+ previous = this.turns.get(turnInfo({ ...body, messages: body.messages.slice(0, turn.index) }, scope).key)?.model;
255
262
  }
256
263
  let preserved = false;
257
264
  const preserve = (chosen, why) => { model = chosen; reason = why; preserved = true; };
@@ -268,7 +275,24 @@ export class Router {
268
275
  // clear_at, tool changes, and output_config) instead of down-routing.
269
276
  else if (hasSystemMessage) preserve(body.model, 'mid_conversation_system');
270
277
  else if (unknownModel) preserve(body.model, 'unknown_model');
271
- else if (turn.continuation) keep(previous ? 'tool_turn_pinned' : 'unknown_continuation');
278
+ // A new native request can explicitly select a model-specific thinking
279
+ // mode, including between_tools. An earlier turn's model is not evidence
280
+ // that it accepts that mode. Existing tool turns retain their pin below.
281
+ else if (modelSpecificThinking && body.thinking.type !== 'enabled' && textTurn) preserve(body.model, 'model_specific_features');
282
+ // Stop hooks (including /goal) return feedback as user-role text, even
283
+ // though it still serves the same human prompt. Trust the scoped gateway
284
+ // identity instead of treating that text as a new task. A client model
285
+ // change can be an explicit fallback after a failure; do not undo it.
286
+ // Local /goal commands can omit the gateway prompt ID. Exact feedback for
287
+ // a known goal then uses the original conversation anchor as a fallback.
288
+ else if (textTurn && textPin?.requestedModel === body.model && !modelSpecificFeatures) {
289
+ const needsSonnet = body.thinking?.type === 'adaptive' || body.output_config?.effort || body.max_tokens > 64000;
290
+ if (needsSonnet && (textPin.model === c.models.haiku || rank(textPin.model) === 0)) {
291
+ preserve(c.models.sonnet, 'requires_sonnet_capabilities');
292
+ } else preserve(textPin.model, promptPin ? 'prompt_turn_pinned' : 'goal_turn_pinned');
293
+ }
294
+ else if (turn.goalFeedback && !turnPin) preserve(body.model, 'unknown_continuation');
295
+ else if (turn.continuation && !turn.goalFeedback) keep(previous ? 'tool_turn_pinned' : 'unknown_continuation');
272
296
  // Unknown or model-specific features are preserved, never silently removed.
273
297
  else if (decision.source === 'fallback' && rank(body.model) >= 1) keep('classifier_unavailable');
274
298
  else if (modelSpecificFeatures) keep('model_specific_features');
@@ -313,12 +337,13 @@ export class Router {
313
337
  }
314
338
  const identifiableUpgrade = capacityUpgraded && (turn.index >= 0 || promptId);
315
339
  if ((!turn.continuation || previous || identifiableUpgrade || reason === 'mid_conversation_system') && !['compaction', 'auxiliary'].includes(requestClass)) {
316
- this.turns.set(turn.key, model);
340
+ const pin = { model, requestedModel: body.model };
341
+ this.turns.set(turn.key, pin);
317
342
  // Keep the content key too: later human turns carry signed thinking but
318
343
  // have a new prompt ID, so they must recover the preceding routed model.
319
344
  // Refresh this alias during known continuations too, since tool discovery
320
345
  // can change the content key without changing the gateway prompt ID.
321
- if (turn.contentKey !== turn.key) this.turns.set(turn.contentKey, model);
346
+ if (turn.contentKey !== turn.key) this.turns.set(turn.contentKey, pin);
322
347
  }
323
348
  return { ...decision, ...contextCheck, model, reason, latency_ms: Math.round((performance.now() - start) * 100) / 100 };
324
349
  }
@@ -1,6 +1,7 @@
1
1
  const PHASES = new Set(['routing', 'connecting', 'streaming', 'ready', 'error', 'cancelled']);
2
2
  const REASONS = {
3
- tool_turn_pinned: 'turn pinned', thinking_history: 'thinking pinned', unknown_continuation: 'continuation pinned',
3
+ tool_turn_pinned: 'turn pinned', prompt_turn_pinned: 'prompt pinned', goal_turn_pinned: 'goal pinned',
4
+ thinking_history: 'thinking pinned', unknown_continuation: 'continuation pinned',
4
5
  mid_conversation_system: 'system features', requires_sonnet_capabilities: 'capability guard',
5
6
  model_specific_features: 'model features', large_or_multimodal_request: 'large request',
6
7
  context_capacity: 'large context',
@@ -1,4 +1,5 @@
1
1
  import { createHash } from 'node:crypto';
2
+ import { prepareRequest } from './model-request.mjs';
2
3
 
3
4
  // Count-token API fields, including the beta fields used by Claude Code.
4
5
  // Keep unknown extensions out of the count path: dropping new input context
@@ -80,7 +81,10 @@ export function createTokenCounter(config, { fetchImpl = fetch } = {}) {
80
81
  let onAbort;
81
82
  const controller = new AbortController();
82
83
  try {
83
- const payload = { ...Object.fromEntries(Object.entries(body).filter(([key]) => COUNT_FIELDS.has(key))), model };
84
+ // Apply the same target-model compatibility changes as inference before
85
+ // projecting count fields. Keep the source model available to the adapter.
86
+ const prepared = prepareRequest(body, model).request;
87
+ const payload = Object.fromEntries(Object.entries(prepared).filter(([key]) => COUNT_FIELDS.has(key)));
84
88
  const serialized = JSON.stringify(payload);
85
89
  const requestHeaders = new Headers(headers);
86
90
  requestHeaders.delete('content-length');