claude-autorouter 0.3.3 → 0.3.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/docs/reference.md CHANGED
@@ -166,7 +166,8 @@ The following policy applies after classification:
166
166
 
167
167
  - Jev's 1,500 ms deadline covers the response body and has no retry. Successful calls return immediately. Timeouts, HTTP errors, and invalid responses fall back to Sonnet or retain an existing stronger model.
168
168
  - Jev confidence below 0.75 prevents a downgrade below Sonnet or the requested tier. Ollama returns a tier without calibrated confidence; its failure handling and compatibility guards still apply.
169
- - Tool continuations retain the model chosen at the start of the human turn. Session, agent, and prompt headers identify turns; normalized conversation content provides a fallback. Moving prompt-cache markers does not create a new turn.
169
+ - Tool continuations retain the model chosen at the start of the human turn. Session, agent, and prompt headers identify turns; normalized conversation content provides a fallback. Text feedback from a Stop hook also retains the model when it serves the same gateway prompt ID and the client has not changed its requested model, subject to capability and context checks. Moving prompt-cache markers does not create a new turn.
170
+ - Claude's local `/goal` command can omit the prompt-ID header. For that path, an exact feedback label matching a preceding expanded `/goal` command keeps the original task and conversation anchor. This narrow text fallback also recognizes Claude's repeated-goal truncation format; arbitrary hook text is not treated as a goal. Feedback remains in the evaluator's recent conversation and the full API request. A new human message becomes the current task normally. The status line shows `prompt pinned` or `goal pinned` when either text-continuation rule applies.
170
171
  - Thinking history, fixed-budget thinking, server tools, context management, and other recognized model-specific features preserve the current model. Adaptive thinking, effort, and output above 64K prevent a Haiku choice. Fields are never stripped to force a downgrade.
171
172
  - Mid-conversation `system` messages preserve the requested model and pass through unchanged. They do not count as a tool continuation by themselves.
172
173
  - Recognized compaction and auxiliary requests otherwise preserve their requested model. Token counting and model discovery pass through without classification.
@@ -233,6 +234,10 @@ The launcher is quiet by default. Standalone `serve` logs to stderr by default.
233
234
 
234
235
  **Claude says Haiku while AutoRouter says Opus:** the built-in model label is Claude's starting/requested model. The AutoRouter confirmed-model label comes from Anthropic. Claude's own token-cost estimate can likewise be attributed to the requested model.
235
236
 
237
+ **`/goal` repeats a question or says it is blocked:** Claude Code runs its own completion checker after each turn, separately from AutoRouter's Jev/Ollama classifier. AutoRouter preserves the requested model for that auxiliary check and forwards its verdict unchanged. The check cannot authorize GitHub SAML, approve a tool, or resolve an external dependency. A tool's authentication error is also different from a Claude API authentication failure.
238
+
239
+ If the worker reports a blocker but the checker keeps returning “not yet met,” Claude can repeat its answer until its no-progress guard pauses the goal. Repeated tool calls can keep the loop running longer. Use `/goal clear` to end the loop, resolve the external blocker, and set the goal again. For tasks that may require human action, explicitly allow reporting a blocker as an alternative end condition, for example: `/goal Verify the discrepancy against upstream main and run the relevant tests, or report an external authorization blocker and stop.` This changes what counts as completion; AutoRouter does not declare blocked work successful or rewrite goal instructions. See [Claude Code goal evaluation](https://code.claude.com/docs/en/goal#how-evaluation-works).
240
+
236
241
  **After restarting mid-conversation:** turn state is in memory and expires after 30 minutes. Unknown continuations preserve the incoming model. Start a fresh conversation when restarting around signed thinking; AutoRouter cannot reconstruct the prior actual model from lost turn state.
237
242
 
238
243
  Switching models can lose prompt-cache reuse. A cheaper price per token does not guarantee a cheaper or faster task. Only requests using Claude's configured base URL are visible to this proxy. Alternate provider modes such as Bedrock, Vertex, Foundry, Mantle, and `ANTHROPIC_AWS` are unsupported; unset their enable flags before launching. Use ordinary `claude` to bypass routing.
package/docs/releasing.md CHANGED
@@ -8,6 +8,8 @@ Version `0.3.2` fixes local timeout fallbacks with model-specific deadlines, rem
8
8
 
9
9
  Version `0.3.3` fixes HTTP 400 errors when a compatible request with disabled thinking is routed to Sonnet 5.5. Inference and token counting translate that setting to `between_tools`, or adaptive thinking when effort settings require it. Model defaults are unchanged; select Sonnet 5.5 with `AUTOROUTER_SONNET_MODEL=claude-sonnet-5-5`.
10
10
 
11
+ Version `0.3.4` preserves the selected model across Stop-hook feedback for the same prompt, including `/goal` commands that omit the gateway prompt-ID header. Recognized goal feedback remains conversation context rather than replacing the human task in evaluator excerpts. Goal-checker verdicts remain unchanged; external authorization blockers can still cause Claude's own goal loop to repeat. See [goal troubleshooting](reference.md#troubleshooting).
12
+
11
13
  The GitHub repository is private. Publishing to npm makes the tarball's runtime source, README, configuration example, license, and shipped documentation public. Model weights, user configuration, credentials, transcripts, local artifacts, and test fixtures are excluded. Review the archive before the first publication and whenever the package allowlist changes.
12
14
 
13
15
  ## What runs automatically
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "claude-autorouter",
3
- "version": "0.3.3",
3
+ "version": "0.3.4",
4
4
  "license": "Apache-2.0",
5
5
  "type": "module",
6
6
  "description": "A Claude Code model router with Jev and local Ollama System One evaluators",
@@ -76,12 +76,60 @@ function humanTask(message) {
76
76
  return contentText(message.content, true);
77
77
  }
78
78
 
79
+ // Claude Code emits /goal Stop-hook feedback as user text. Recognize only its
80
+ // observed wrapper and a condition established by an earlier expanded command;
81
+ // ordinary messages mentioning hooks remain human tasks. This does not alter
82
+ // message content, and the feedback remains available as classifier history.
83
+ export function goalFeedbackIndexes(messages) {
84
+ const indexes = new Set();
85
+ if (!Array.isArray(messages)) return indexes;
86
+ let condition;
87
+ let shortCondition;
88
+ let sawFullFeedback = false;
89
+ for (let index = 0; index < messages.length; index++) {
90
+ const message = messages[index];
91
+ if (message?.role !== 'user') continue;
92
+ const command = /^\s*<command-name>\/goal<\/command-name>\s*<command-message>goal<\/command-message>\s*<command-args>([\s\S]*?)<\/command-args>(?:\s|$)/.exec(humanTask(message));
93
+ if (command) {
94
+ const value = command[1].trim();
95
+ // /goal with no arguments is a status query, not a replacement goal.
96
+ if (!value) continue;
97
+ condition = value.length <= 4000 && !/^(?:clear|stop|off|reset|none|cancel)$/i.test(value) ? value : undefined;
98
+ shortCondition = undefined;
99
+ sawFullFeedback = false;
100
+ if (condition?.length > 500) {
101
+ let prefix = condition.slice(0, 500);
102
+ const last = prefix.charCodeAt(prefix.length - 1);
103
+ if (last >= 0xd800 && last <= 0xdbff) prefix = prefix.slice(0, -1);
104
+ shortCondition = `${prefix}… [+${condition.length - prefix.length} chars]`;
105
+ }
106
+ continue;
107
+ }
108
+ if (!condition || messages[index - 1]?.role !== 'assistant') continue;
109
+ const text = typeof message.content === 'string' ? message.content
110
+ : Array.isArray(message.content) && message.content.length === 1 && message.content[0]?.type === 'text'
111
+ && typeof message.content[0].text === 'string' ? message.content[0].text : undefined;
112
+ if (text === undefined) continue;
113
+ const matches = label => {
114
+ const prefix = `Stop hook feedback:\n[${label}]: `;
115
+ return text.startsWith(prefix) && Boolean(text.slice(prefix.length).trim());
116
+ };
117
+ if (matches(condition)) {
118
+ indexes.add(index);
119
+ sawFullFeedback = true;
120
+ } else if (sawFullFeedback && shortCondition && matches(shortCondition)) indexes.add(index);
121
+ }
122
+ return indexes;
123
+ }
124
+
79
125
  export function buildState(body, limit = 12000) {
80
126
  const messages = body.messages ?? [];
127
+ const feedbackIndexes = goalFeedbackIndexes(messages);
81
128
  let firstTask = '';
82
129
  let currentTask = '';
83
130
  let currentIndex = -1;
84
131
  for (let index = 0; index < messages.length; index++) {
132
+ if (feedbackIndexes.has(index)) continue;
85
133
  const task = humanTask(messages[index]);
86
134
  if (!task.trim()) continue;
87
135
  if (!firstTask) firstTask = task;
package/src/router.mjs CHANGED
@@ -1,6 +1,6 @@
1
1
  import { createHash } from 'node:crypto';
2
2
  import { TIERS } from './config.mjs';
3
- import { buildState } from './prompt-state.mjs';
3
+ import { buildState, goalFeedbackIndexes } from './prompt-state.mjs';
4
4
  import { buildOllamaState, evaluateOllama } from './ollama-evaluator.mjs';
5
5
  export { buildState } from './prompt-state.mjs';
6
6
 
@@ -134,10 +134,11 @@ function turnContent(content) {
134
134
 
135
135
  function turnInfo(body, scope, promptId = '') {
136
136
  const messages = body.messages ?? [];
137
+ const feedback = goalFeedbackIndexes(messages);
137
138
  let index = -1;
138
139
  for (let i = messages.length - 1; i >= 0; i--) {
139
140
  const message = messages[i];
140
- if (message.role === 'user' && !(Array.isArray(message.content) && message.content.some(b => b.type === 'tool_result'))) {
141
+ if (message.role === 'user' && !feedback.has(i) && !(Array.isArray(message.content) && message.content.some(b => b.type === 'tool_result'))) {
141
142
  index = i; break;
142
143
  }
143
144
  }
@@ -147,6 +148,7 @@ function turnInfo(body, scope, promptId = '') {
147
148
  index,
148
149
  key: promptId ? hash(['prompt', scope, promptId]) : contentKey,
149
150
  contentKey,
151
+ goalFeedback: feedback.has(messages.length - 1),
150
152
  // Claude Code can append turn-scoped system instructions after the human
151
153
  // prompt. These do not start an assistant/tool continuation.
152
154
  continuation: index < 0 || messages.slice(index + 1).some(m => m.role !== 'system'),
@@ -248,11 +250,15 @@ export class Router {
248
250
  let model = c.models[decision.tier];
249
251
  let reason = decision.reason;
250
252
  const turn = turnInfo(body, scope, promptId);
251
- let previous = this.turns.get(turn.key) ?? this.turns.get(turn.contentKey);
253
+ const promptPin = promptId ? this.turns.get(turn.key) : undefined;
254
+ const turnPin = promptPin ?? this.turns.get(turn.contentKey);
255
+ let previous = turnPin?.model;
256
+ const textTurn = !turn.continuation || turn.goalFeedback;
257
+ const textPin = promptPin ?? (!promptId && turn.goalFeedback ? turnPin : undefined);
252
258
  // A new human prompt can still carry signed thinking from the preceding
253
259
  // turn. Recover that turn's actual routed model when it is known.
254
260
  if (!turn.continuation && body.messages.length > 1 && !previous) {
255
- previous = this.turns.get(turnInfo({ ...body, messages: body.messages.slice(0, turn.index) }, scope).key);
261
+ previous = this.turns.get(turnInfo({ ...body, messages: body.messages.slice(0, turn.index) }, scope).key)?.model;
256
262
  }
257
263
  let preserved = false;
258
264
  const preserve = (chosen, why) => { model = chosen; reason = why; preserved = true; };
@@ -272,8 +278,21 @@ export class Router {
272
278
  // A new native request can explicitly select a model-specific thinking
273
279
  // mode, including between_tools. An earlier turn's model is not evidence
274
280
  // that it accepts that mode. Existing tool turns retain their pin below.
275
- else if (modelSpecificThinking && body.thinking.type !== 'enabled' && !turn.continuation) preserve(body.model, 'model_specific_features');
276
- else if (turn.continuation) keep(previous ? 'tool_turn_pinned' : 'unknown_continuation');
281
+ else if (modelSpecificThinking && body.thinking.type !== 'enabled' && textTurn) preserve(body.model, 'model_specific_features');
282
+ // Stop hooks (including /goal) return feedback as user-role text, even
283
+ // though it still serves the same human prompt. Trust the scoped gateway
284
+ // identity instead of treating that text as a new task. A client model
285
+ // change can be an explicit fallback after a failure; do not undo it.
286
+ // Local /goal commands can omit the gateway prompt ID. Exact feedback for
287
+ // a known goal then uses the original conversation anchor as a fallback.
288
+ else if (textTurn && textPin?.requestedModel === body.model && !modelSpecificFeatures) {
289
+ const needsSonnet = body.thinking?.type === 'adaptive' || body.output_config?.effort || body.max_tokens > 64000;
290
+ if (needsSonnet && (textPin.model === c.models.haiku || rank(textPin.model) === 0)) {
291
+ preserve(c.models.sonnet, 'requires_sonnet_capabilities');
292
+ } else preserve(textPin.model, promptPin ? 'prompt_turn_pinned' : 'goal_turn_pinned');
293
+ }
294
+ else if (turn.goalFeedback && !turnPin) preserve(body.model, 'unknown_continuation');
295
+ else if (turn.continuation && !turn.goalFeedback) keep(previous ? 'tool_turn_pinned' : 'unknown_continuation');
277
296
  // Unknown or model-specific features are preserved, never silently removed.
278
297
  else if (decision.source === 'fallback' && rank(body.model) >= 1) keep('classifier_unavailable');
279
298
  else if (modelSpecificFeatures) keep('model_specific_features');
@@ -318,12 +337,13 @@ export class Router {
318
337
  }
319
338
  const identifiableUpgrade = capacityUpgraded && (turn.index >= 0 || promptId);
320
339
  if ((!turn.continuation || previous || identifiableUpgrade || reason === 'mid_conversation_system') && !['compaction', 'auxiliary'].includes(requestClass)) {
321
- this.turns.set(turn.key, model);
340
+ const pin = { model, requestedModel: body.model };
341
+ this.turns.set(turn.key, pin);
322
342
  // Keep the content key too: later human turns carry signed thinking but
323
343
  // have a new prompt ID, so they must recover the preceding routed model.
324
344
  // Refresh this alias during known continuations too, since tool discovery
325
345
  // can change the content key without changing the gateway prompt ID.
326
- if (turn.contentKey !== turn.key) this.turns.set(turn.contentKey, model);
346
+ if (turn.contentKey !== turn.key) this.turns.set(turn.contentKey, pin);
327
347
  }
328
348
  return { ...decision, ...contextCheck, model, reason, latency_ms: Math.round((performance.now() - start) * 100) / 100 };
329
349
  }
@@ -1,6 +1,7 @@
1
1
  const PHASES = new Set(['routing', 'connecting', 'streaming', 'ready', 'error', 'cancelled']);
2
2
  const REASONS = {
3
- tool_turn_pinned: 'turn pinned', thinking_history: 'thinking pinned', unknown_continuation: 'continuation pinned',
3
+ tool_turn_pinned: 'turn pinned', prompt_turn_pinned: 'prompt pinned', goal_turn_pinned: 'goal pinned',
4
+ thinking_history: 'thinking pinned', unknown_continuation: 'continuation pinned',
4
5
  mid_conversation_system: 'system features', requires_sonnet_capabilities: 'capability guard',
5
6
  model_specific_features: 'model features', large_or_multimodal_request: 'large request',
6
7
  context_capacity: 'large context',