@gajae-code/ai 0.12.12 → 0.12.13

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,23 @@
2
2
 
3
3
  ## [Unreleased]
4
4
 
5
+ ## [0.12.13] - 2026-08-06
6
+
7
+ ### Changed
8
+
9
+ - Anthropic prompt caching now defaults to top-level automatic caching (`cache_control: { type: "ephemeral" }`) for every Claude-family model, including through non-canonical Anthropic-compatible gateways (Cloudflare AI Gateway, GitHub Copilot, GitLab Duo, Vercel AI Gateway, zenmux, etc.), instead of only `api.anthropic.com`. Non-Claude models on unknown compatible endpoints keep the previous no-cache default; `compat.promptCacheMode: "none"`, `compat.promptCacheMode: "explicit"`, and configured or per-request `cacheRetention: "none"` still opt out. Non-canonical Claude models get the default ~5m cache lifetime unless the endpoint sets `compat.supportsLongCacheRetention: true`.
10
+
11
+ ### Fixed
12
+
13
+ - `todo_write` raw argument rejections now carry bounded, authority-controlled correction codes for each rejected shape: unknown root keys, unknown operation-entry keys, done/drop entries missing a task or phase target, and unknown init list-entry keys. Each code maps to a fixed correction message naming the accepted shape (never echoing the offending input), so invalid calls surface specific guidance while valid payloads keep the existing passthrough/coercion path (#3916).
14
+ - Anthropic Sonnet 5 now exposes Anthropic's real `xhigh` and `max` thinking efforts on the Messages API (`minimal`/`low`/`medium`/`high`/`xhigh`/`max`), matching official support. The previous generic `kind === opus` gate excluded it from the full preset range; the capability predicate is now an explicit version-scoped list (Opus 4.7+, Sonnet 5+), so older Sonnet generations and Bedrock Converse routes stay fail-closed at their previously advertised levels (issue #3913).
15
+ - Alibaba Token Plan now exposes Qwen 3.8 Max under the provider-supported `qwen3.8-max` wire id instead of the rejected `qwen-3.8-max` spelling; catalog regeneration canonicalizes a legacy discovered alias rather than retaining a broken duplicate (#3909).
16
+ - Canonicalized first-class MiniMax M3 catalog ids (issue #3896). The bundled catalog previously shipped stale lowercase `minimax-m3` duplicates (512K) next to the canonical `MiniMax-M3` (1M) on all four first-class MiniMax providers, plus a non-official `minimax-v3` entry under `minimax-code`. The lowercase `minimax-m3` entries and `minimax-v3` are removed; `MiniMax-M3` is the single canonical first-class id (the regen-safe 1M pin in `applyGeneratedModelPolicy` now keys on `MiniMax-M3` / `MiniMax-M3[1m]` instead of the removed lowercase id), `DEFAULT_MODEL_PER_PROVIDER` points at `MiniMax-M3`, and the official Anthropic Token Plan id `MiniMax-M3[1m]` is first-class on the `minimax` / `minimax-cn` Anthropic routes with 1M context semantics. Unrelated catalog providers keep their own `minimax-m3` contracts.
17
+ - Anthropic thinking-replay repair now also triggers when the mutation/signature `invalid_request_error` arrives as a statusless in-stream SSE `error` event (issue #3900). Proxies such as CLIProxyAPI forward the upstream 400 body over an HTTP 200 SSE stream, so the thrown error carries no HTTP status; the classifiers previously required `status === 400` and let the session loop on an unrecoverable replay rejection. Statusless errors still require the full `invalid_request_error` thinking wording, so unrelated transport failures never claim the one-shot repair.
18
+ - Anthropic thinking-replay repair now also recovers when a proxy masks the rejection entirely (issue #3900). Live CLIProxyAPI captures replace the upstream 400 body with a generic `{"type":"api_error","message":"An error occurred while processing the request."}` SSE event on an HTTP 200 response, which names no cause and matches no transient phrase, so the turn died on the first attempt. Such a masked rejection now takes the same one-shot latest-then-full-history repair, but only before the first token and only while the request actually replays signed `thinking`/`redacted_thinking` blocks; masked failures on requests without replayed thinking still surface immediately. The classifier is exported as `isAnthropicMaskedProxyRejection`.
19
+
20
+ - Anthropic cache-control resolution now falls back to `model.cacheRetention` at the provider boundary, preserving configured retention and request-over-model precedence through special dispatch wrappers such as GitLab Duo. A configured `cacheRetention: "none"` can no longer be dropped and replaced by the new automatic Claude-family cache marker.
21
+ - Anthropic explicit prompt caching now advances its conversation breakpoint during tool-use loops by marking the latest completed assistant tool-use turn while leaving the newest tool result uncached. Previously it kept refreshing only the original human message until another human turn arrived, pinning proxy cache reads to the static tools/system prefix throughout long agentic runs.
5
22
  ## [0.12.12] - 2026-08-05
6
23
 
7
24
  ### Fixed
@@ -9,6 +26,7 @@
9
26
  - OpenAI Responses and Azure OpenAI Responses now map the first-event timeout into the SDK request/setup timeout the same way Completions does, so a never-resolving pre-headers fetch on a provider-owned lazy stream cannot wait the SDK's 10-minute default before any transport watchdog exists. Alibaba Responses honors an explicit shorter first-event override before headers; Azure/env-pinned setup timeouts normalize to the typed `stream_first_event_timeout` failure.
10
27
  - OpenAI Codex cost estimates now treat an explicit response `service_tier` as authoritative, so a request for priority processing that the provider serves at the default tier is no longer charged the priority multiplier; the requested tier remains the fallback when the terminal response omits the field.
11
28
  - Added shared `isReasoningContentReplayError` classifier and `stripUnusableReasoningItems` repair for the DeepSeek-family reasoning-content replay rejection ("The `reasoning_content` in the thinking mode must be passed back to the API"). The classifier detects the error across message carrier shapes; the repair removes only `reasoning` items whose `encrypted_content` a proxy stripped to empty, preserving all non-reasoning history (text, tool calls, tool outputs). The agent loop consumes both for a bounded repair-and-resend circuit breaker.
29
+ - Codex statusless HTTP 200 SSE `invalid_request_error` events retry once without a forced named function choice only when the exact rejected name is still present in the request's serialized tools, before any output is emitted (#3669).
12
30
 
13
31
  ## [0.12.11] - 2026-08-03
14
32
 
@@ -50,6 +50,19 @@ export declare function isAnthropicThinkingBlockMutationError(error: unknown): b
50
50
  * than only the latest one.
51
51
  */
52
52
  export declare function isAnthropicThinkingSignatureInvalidError(error: unknown): boolean;
53
+ /**
54
+ * CLIProxyAPI replaces Anthropic's rejection body wholesale instead of forwarding
55
+ * it: the client only ever sees
56
+ * `{"type":"error","error":{"type":"api_error","message":"An error occurred while
57
+ * processing the request."}}`, delivered as an in-stream SSE `error` event on an
58
+ * HTTP 200 response, so neither the status nor the message survives. Captured CPA
59
+ * traces for that masked shape carry the thinking-integrity 400 upstream (issue
60
+ * #3900), and the generic body matches no transient phrase either, so the turn
61
+ * dies unrecoverably. Nothing in the payload names the cause; callers must pair
62
+ * this with a request that actually replays signed thinking blocks before
63
+ * treating it as a thinking-replay rejection.
64
+ */
65
+ export declare function isAnthropicMaskedProxyRejection(error: unknown): boolean;
53
66
  export declare const claudeCodeVersion = "2.1.219";
54
67
  export declare const claudeCodeEntrypoint = "sdk-cli";
55
68
  export declare const claudeToolPrefix: string;
@@ -537,7 +537,7 @@ export type TSchema = ZodType | TJsonSchema;
537
537
  export type Static<S> = S extends ZodType ? z.infer<S> : S extends {
538
538
  static: infer T;
539
539
  } ? T : unknown;
540
- export type RawArgumentRejectionCode = "ask-intent-review-requires-positive-round" | "ask-intent-contract-requires-non-empty-authority" | "ask-deep-interview-metadata-requires-deep-interview-gate";
540
+ export type RawArgumentRejectionCode = "ask-intent-review-requires-positive-round" | "ask-intent-contract-requires-non-empty-authority" | "ask-deep-interview-metadata-requires-deep-interview-gate" | "todo-write-unknown-root-key" | "todo-write-unknown-op-entry-key" | "todo-write-done-drop-requires-target" | "todo-write-unknown-init-entry-key";
541
541
  export type RawArgumentValidationResult = {
542
542
  outcome: "passthrough";
543
543
  } | {
@@ -790,8 +790,10 @@ export interface AnthropicCompat extends ToolChoiceCompat {
790
790
  supportsLongCacheRetention?: boolean;
791
791
  /**
792
792
  * Prompt-cache transport accepted by this Anthropic-compatible endpoint.
793
- * Canonical Anthropic defaults to `"automatic"`; noncanonical endpoints default
794
- * to `"none"` and must explicitly opt into generated `"explicit"` markers.
793
+ * Canonical Anthropic and Claude-family models default to `"automatic"`;
794
+ * noncanonical non-Claude endpoints default to `"none"`. Set `"automatic"` to
795
+ * opt an otherwise unknown compatible endpoint into top-level caching, `"none"`
796
+ * to opt out, or `"explicit"` for endpoints that require block-level markers.
795
797
  */
796
798
  promptCacheMode?: "none" | "explicit" | "automatic";
797
799
  }
@@ -26,6 +26,11 @@ export declare function markToolChoiceIncapability(model: Model<Api>, maxSupport
26
26
  export declare function resolveToolChoice(model: Model<Api>, requested: ToolChoice | undefined, compat?: ToolChoiceCompat): ResolveToolChoiceResult;
27
27
  /** Detects provider errors indicating forced tool_choice is unsupported. */
28
28
  export declare function isForcedToolChoiceUnsupportedError(error: unknown, sentForcedToolChoice: boolean): boolean;
29
+ /**
30
+ * Detects Codex's statusless SSE rejection for a named function tool choice.
31
+ * This is intentionally separate from the shared HTTP-400 classifier.
32
+ */
33
+ export declare function isCodexStatuslessNamedToolChoiceNotFoundError(error: unknown, forcedToolName: string | undefined, sentToolNames: readonly string[]): boolean;
29
34
  export type { ToolChoiceCompat, ToolChoiceSupport, ToolChoiceSupportSource } from "../types";
30
35
  export interface ResolveToolChoiceResult {
31
36
  requestedChoice: ToolChoice | undefined;
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "type": "module",
3
3
  "name": "@gajae-code/ai",
4
- "version": "0.12.12",
4
+ "version": "0.12.13",
5
5
  "description": "Unified LLM API with automatic model discovery and provider configuration",
6
6
  "homepage": "https://gajae-code.com",
7
7
  "author": "Yeachan-Heo and Gajae Code Contributors",
@@ -40,7 +40,7 @@
40
40
  "dependencies": {
41
41
  "@anthropic-ai/sdk": "^0.94.0",
42
42
  "@bufbuild/protobuf": "^2.12.0",
43
- "@gajae-code/utils": "0.12.12",
43
+ "@gajae-code/utils": "0.12.13",
44
44
  "openai": "^6.36.0",
45
45
  "partial-json": "^0.1.7",
46
46
  "zod": "4.4.3"
@@ -379,8 +379,18 @@ export function supportsAnthropicAdaptiveThinkingDisplay(modelId: string): boole
379
379
  function anthropicModelHasRealXHighEffort<TApi extends Api>(model: ApiModel<TApi>): boolean {
380
380
  if (model.api !== "anthropic-messages") return false;
381
381
  const parsedModel = parseKnownModel(model.id);
382
- if (parsedModel.family !== "anthropic" || parsedModel.kind !== "opus") return false;
383
- return semverGte(parsedModel.version, "4.7");
382
+ if (parsedModel.family !== "anthropic") return false;
383
+ // Explicit capability predicate instead of a generic `kind === opus` gate:
384
+ // Sonnet 5 officially exposes Anthropic's real xhigh and max presets on
385
+ // the Messages API just like Opus 4.7+. Older Sonnet generations do not,
386
+ // so the predicate stays fail-closed for them.
387
+ if (parsedModel.kind === "opus") {
388
+ return semverGte(parsedModel.version, "4.7");
389
+ }
390
+ if (parsedModel.kind === "sonnet") {
391
+ return semverGte(parsedModel.version, "5.0");
392
+ }
393
+ return false;
384
394
  }
385
395
 
386
396
  function applyGeneratedModelPolicy(model: ApiModel<Api>): void {
@@ -457,10 +467,11 @@ function applyGeneratedModelPolicy(model: ApiModel<Api>): void {
457
467
  };
458
468
  }
459
469
  // MiniMax-M3's official Token Plan routes expose a 1M context window.
460
- // Scope the correction to the four first-class regional MiniMax routes;
461
- // unrelated catalog aliases and providers keep their own contracts.
470
+ // Scope the correction to the four first-class regional MiniMax routes
471
+ // (canonical id plus the Anthropic Token Plan `[1m]` id); unrelated
472
+ // catalog aliases and providers keep their own contracts.
462
473
  if (
463
- model.id === "minimax-m3" &&
474
+ (model.id === "MiniMax-M3" || model.id === "MiniMax-M3[1m]") &&
464
475
  (model.provider === "minimax" ||
465
476
  model.provider === "minimax-cn" ||
466
477
  model.provider === "minimax-code" ||
@@ -690,10 +701,16 @@ function inferAnthropicSupportedEfforts<TApi extends Api>(
690
701
  // Converse lacks it (same split as Opus 4.7+ below).
691
702
  return model.api === "anthropic-messages" ? DEFAULT_REASONING_EFFORTS_WITH_XHIGH : DEFAULT_REASONING_EFFORTS;
692
703
  }
693
- if (parsedModel.kind !== "opus") return DEFAULT_REASONING_EFFORTS;
694
- return anthropicModelHasRealXHighEffort(model)
695
- ? DEFAULT_REASONING_EFFORTS_WITH_XHIGH_AND_MAX
696
- : DEFAULT_REASONING_EFFORTS_WITH_MAX;
704
+ if (anthropicModelHasRealXHighEffort(model)) {
705
+ // Opus 4.7+ and Sonnet 5 expose both Anthropic's real xhigh and
706
+ // max presets on the Messages API.
707
+ return DEFAULT_REASONING_EFFORTS_WITH_XHIGH_AND_MAX;
708
+ }
709
+ if (parsedModel.kind === "opus") {
710
+ // Opus 4.6 exposes max but not the newer xhigh literal.
711
+ return DEFAULT_REASONING_EFFORTS_WITH_MAX;
712
+ }
713
+ return DEFAULT_REASONING_EFFORTS;
697
714
  }
698
715
  return inferFallbackEfforts(model);
699
716
  }
package/src/models.json CHANGED
@@ -89,8 +89,8 @@
89
89
  "maxLevel": "xhigh"
90
90
  }
91
91
  },
92
- "qwen-3.8-max": {
93
- "id": "qwen-3.8-max",
92
+ "qwen3.8-max": {
93
+ "id": "qwen3.8-max",
94
94
  "name": "Qwen3.8 Max",
95
95
  "api": "openai-responses",
96
96
  "provider": "alibaba-token-plan",
@@ -3952,7 +3952,7 @@
3952
3952
  "thinking": {
3953
3953
  "mode": "anthropic-adaptive",
3954
3954
  "minLevel": "minimal",
3955
- "maxLevel": "high"
3955
+ "maxLevel": "max"
3956
3956
  }
3957
3957
  },
3958
3958
  "claude-opus-5": {
@@ -4740,7 +4740,7 @@
4740
4740
  "thinking": {
4741
4741
  "mode": "anthropic-adaptive",
4742
4742
  "minLevel": "minimal",
4743
- "maxLevel": "high"
4743
+ "maxLevel": "max"
4744
4744
  }
4745
4745
  },
4746
4746
  "claude-sonnet-4-5": {
@@ -9866,7 +9866,7 @@
9866
9866
  "thinking": {
9867
9867
  "mode": "anthropic-adaptive",
9868
9868
  "minLevel": "minimal",
9869
- "maxLevel": "high"
9869
+ "maxLevel": "max"
9870
9870
  }
9871
9871
  },
9872
9872
  "gemini-2.5-pro": {
@@ -40201,8 +40201,8 @@
40201
40201
  "maxLevel": "xhigh"
40202
40202
  }
40203
40203
  },
40204
- "minimax-m3": {
40205
- "id": "minimax-m3",
40204
+ "MiniMax-M3": {
40205
+ "id": "MiniMax-M3",
40206
40206
  "name": "MiniMax-M3",
40207
40207
  "api": "anthropic-messages",
40208
40208
  "provider": "minimax",
@@ -40213,12 +40213,12 @@
40213
40213
  "image"
40214
40214
  ],
40215
40215
  "cost": {
40216
- "input": 0.6,
40217
- "output": 2.4,
40218
- "cacheRead": 0.12,
40216
+ "input": 0.3,
40217
+ "output": 1.2,
40218
+ "cacheRead": 0.06,
40219
40219
  "cacheWrite": 0
40220
40220
  },
40221
- "contextWindow": 512000,
40221
+ "contextWindow": 1000000,
40222
40222
  "maxTokens": 128000,
40223
40223
  "thinking": {
40224
40224
  "mode": "budget",
@@ -40226,9 +40226,9 @@
40226
40226
  "maxLevel": "xhigh"
40227
40227
  }
40228
40228
  },
40229
- "MiniMax-M3": {
40230
- "id": "MiniMax-M3",
40231
- "name": "MiniMax-M3",
40229
+ "MiniMax-M3[1m]": {
40230
+ "id": "MiniMax-M3[1m]",
40231
+ "name": "MiniMax-M3[1m]",
40232
40232
  "api": "anthropic-messages",
40233
40233
  "provider": "minimax",
40234
40234
  "baseUrl": "https://api.minimax.io/anthropic",
@@ -40421,8 +40421,8 @@
40421
40421
  "maxLevel": "xhigh"
40422
40422
  }
40423
40423
  },
40424
- "minimax-m3": {
40425
- "id": "minimax-m3",
40424
+ "MiniMax-M3": {
40425
+ "id": "MiniMax-M3",
40426
40426
  "name": "MiniMax-M3",
40427
40427
  "api": "anthropic-messages",
40428
40428
  "provider": "minimax-cn",
@@ -40433,12 +40433,12 @@
40433
40433
  "image"
40434
40434
  ],
40435
40435
  "cost": {
40436
- "input": 0.6,
40437
- "output": 2.4,
40438
- "cacheRead": 0.12,
40436
+ "input": 0.3,
40437
+ "output": 1.2,
40438
+ "cacheRead": 0.06,
40439
40439
  "cacheWrite": 0
40440
40440
  },
40441
- "contextWindow": 512000,
40441
+ "contextWindow": 1000000,
40442
40442
  "maxTokens": 128000,
40443
40443
  "thinking": {
40444
40444
  "mode": "budget",
@@ -40446,9 +40446,9 @@
40446
40446
  "maxLevel": "xhigh"
40447
40447
  }
40448
40448
  },
40449
- "MiniMax-M3": {
40450
- "id": "MiniMax-M3",
40451
- "name": "MiniMax-M3",
40449
+ "MiniMax-M3[1m]": {
40450
+ "id": "MiniMax-M3[1m]",
40451
+ "name": "MiniMax-M3[1m]",
40452
40452
  "api": "anthropic-messages",
40453
40453
  "provider": "minimax-cn",
40454
40454
  "baseUrl": "https://api.minimaxi.com/anthropic",
@@ -40713,37 +40713,6 @@
40713
40713
  "maxLevel": "high"
40714
40714
  }
40715
40715
  },
40716
- "minimax-m3": {
40717
- "id": "minimax-m3",
40718
- "name": "MiniMax-M3",
40719
- "api": "openai-completions",
40720
- "provider": "minimax-code",
40721
- "baseUrl": "https://api.minimax.io/v1",
40722
- "reasoning": true,
40723
- "input": [
40724
- "text",
40725
- "image"
40726
- ],
40727
- "cost": {
40728
- "input": 0,
40729
- "output": 0,
40730
- "cacheRead": 0,
40731
- "cacheWrite": 0
40732
- },
40733
- "contextWindow": 512000,
40734
- "maxTokens": 128000,
40735
- "compat": {
40736
- "supportsStore": false,
40737
- "supportsDeveloperRole": false,
40738
- "supportsReasoningEffort": false,
40739
- "reasoningContentField": "reasoning_content"
40740
- },
40741
- "thinking": {
40742
- "mode": "effort",
40743
- "minLevel": "minimal",
40744
- "maxLevel": "high"
40745
- }
40746
- },
40747
40716
  "MiniMax-M3": {
40748
40717
  "id": "MiniMax-M3",
40749
40718
  "name": "MiniMax-M3",
@@ -40774,37 +40743,6 @@
40774
40743
  "minLevel": "minimal",
40775
40744
  "maxLevel": "high"
40776
40745
  }
40777
- },
40778
- "minimax-v3": {
40779
- "id": "minimax-v3",
40780
- "name": "MiniMax-V3",
40781
- "api": "openai-completions",
40782
- "provider": "minimax-code",
40783
- "baseUrl": "https://api.minimax.io/v1",
40784
- "reasoning": true,
40785
- "input": [
40786
- "text",
40787
- "image"
40788
- ],
40789
- "cost": {
40790
- "input": 0,
40791
- "output": 0,
40792
- "cacheRead": 0,
40793
- "cacheWrite": 0
40794
- },
40795
- "contextWindow": 512000,
40796
- "maxTokens": 128000,
40797
- "compat": {
40798
- "supportsStore": false,
40799
- "supportsDeveloperRole": false,
40800
- "supportsReasoningEffort": false,
40801
- "reasoningContentField": "reasoning_content"
40802
- },
40803
- "thinking": {
40804
- "mode": "effort",
40805
- "minLevel": "minimal",
40806
- "maxLevel": "high"
40807
- }
40808
40746
  }
40809
40747
  },
40810
40748
  "minimax-code-cn": {
@@ -41048,37 +40986,6 @@
41048
40986
  "maxLevel": "high"
41049
40987
  }
41050
40988
  },
41051
- "minimax-m3": {
41052
- "id": "minimax-m3",
41053
- "name": "MiniMax-M3",
41054
- "api": "openai-completions",
41055
- "provider": "minimax-code-cn",
41056
- "baseUrl": "https://api.minimaxi.com/v1",
41057
- "reasoning": true,
41058
- "input": [
41059
- "text",
41060
- "image"
41061
- ],
41062
- "cost": {
41063
- "input": 0,
41064
- "output": 0,
41065
- "cacheRead": 0,
41066
- "cacheWrite": 0
41067
- },
41068
- "contextWindow": 512000,
41069
- "maxTokens": 128000,
41070
- "compat": {
41071
- "supportsStore": false,
41072
- "supportsDeveloperRole": false,
41073
- "supportsReasoningEffort": false,
41074
- "reasoningContentField": "reasoning_content"
41075
- },
41076
- "thinking": {
41077
- "mode": "effort",
41078
- "minLevel": "minimal",
41079
- "maxLevel": "high"
41080
- }
41081
- },
41082
40989
  "MiniMax-M3": {
41083
40990
  "id": "MiniMax-M3",
41084
40991
  "name": "MiniMax-M3",
@@ -60445,7 +60352,7 @@
60445
60352
  "thinking": {
60446
60353
  "mode": "anthropic-adaptive",
60447
60354
  "minLevel": "minimal",
60448
- "maxLevel": "high"
60355
+ "maxLevel": "max"
60449
60356
  }
60450
60357
  },
60451
60358
  "deepseek-v4-flash": {
@@ -75442,7 +75349,7 @@
75442
75349
  "thinking": {
75443
75350
  "mode": "anthropic-adaptive",
75444
75351
  "minLevel": "minimal",
75445
- "maxLevel": "high"
75352
+ "maxLevel": "max"
75446
75353
  }
75447
75354
  },
75448
75355
  "arcee-ai/trinity-large-preview": {
@@ -81719,7 +81626,7 @@
81719
81626
  "thinking": {
81720
81627
  "mode": "anthropic-adaptive",
81721
81628
  "minLevel": "minimal",
81722
- "maxLevel": "high"
81629
+ "maxLevel": "max"
81723
81630
  }
81724
81631
  },
81725
81632
  "anthropic/claude-sonnet-5-free": {
@@ -81744,7 +81651,7 @@
81744
81651
  "thinking": {
81745
81652
  "mode": "anthropic-adaptive",
81746
81653
  "minLevel": "minimal",
81747
- "maxLevel": "high"
81654
+ "maxLevel": "max"
81748
81655
  }
81749
81656
  },
81750
81657
  "baidu/ernie-5.0-thinking-preview": {
@@ -365,9 +365,9 @@ export const DEFAULT_MODEL_PER_PROVIDER: Record<KnownProvider, string> = {
365
365
  "google-antigravity": "gemini-3-pro-high",
366
366
  "google-gemini-cli": "gemini-2.5-pro",
367
367
  "google-vertex": "gemini-3-pro-preview",
368
- minimax: "minimax-m3",
369
- "minimax-code": "minimax-m3",
370
- "minimax-code-cn": "minimax-m3",
368
+ minimax: "MiniMax-M3",
369
+ "minimax-code": "MiniMax-M3",
370
+ "minimax-code-cn": "MiniMax-M3",
371
371
  "openai-codex": "gpt-5.5",
372
372
  "gitlab-duo": "duo-chat-sonnet-4-5",
373
373
  } as Record<KnownProvider, string>;
@@ -395,8 +395,20 @@ export function isAnthropicFastModeUnsupportedError(error: unknown): boolean {
395
395
  return false;
396
396
  }
397
397
 
398
+ /**
399
+ * Proxies (e.g. CLIProxyAPI) can deliver Anthropic's 400 body as an in-stream
400
+ * SSE `error` event on an HTTP 200 response; the thrown error then carries no
401
+ * HTTP status at all (issue #3900). Accept both the direct 400 and the
402
+ * statusless SSE shape — the strict `invalid_request_error` message checks in
403
+ * each matcher keep the statusless branch from claiming unrelated failures.
404
+ */
405
+ function isAnthropicInvalidRequestStatus(error: unknown): boolean {
406
+ const status = extractHttpStatusFromError(error);
407
+ return status === 400 || status === undefined;
408
+ }
409
+
398
410
  export function isAnthropicThinkingBlockMutationError(error: unknown): boolean {
399
- if (extractHttpStatusFromError(error) !== 400) return false;
411
+ if (!isAnthropicInvalidRequestStatus(error)) return false;
400
412
  const message = error instanceof Error ? error.message : String(error);
401
413
  return (
402
414
  /invalid_request_error/i.test(message) &&
@@ -414,7 +426,7 @@ export function isAnthropicThinkingBlockMutationError(error: unknown): boolean {
414
426
  * than only the latest one.
415
427
  */
416
428
  export function isAnthropicThinkingSignatureInvalidError(error: unknown): boolean {
417
- if (extractHttpStatusFromError(error) !== 400) return false;
429
+ if (!isAnthropicInvalidRequestStatus(error)) return false;
418
430
  const message = error instanceof Error ? error.message : String(error);
419
431
  return (
420
432
  /invalid_request_error/i.test(message) &&
@@ -423,6 +435,27 @@ export function isAnthropicThinkingSignatureInvalidError(error: unknown): boolea
423
435
  );
424
436
  }
425
437
 
438
+ /**
439
+ * CLIProxyAPI replaces Anthropic's rejection body wholesale instead of forwarding
440
+ * it: the client only ever sees
441
+ * `{"type":"error","error":{"type":"api_error","message":"An error occurred while
442
+ * processing the request."}}`, delivered as an in-stream SSE `error` event on an
443
+ * HTTP 200 response, so neither the status nor the message survives. Captured CPA
444
+ * traces for that masked shape carry the thinking-integrity 400 upstream (issue
445
+ * #3900), and the generic body matches no transient phrase either, so the turn
446
+ * dies unrecoverably. Nothing in the payload names the cause; callers must pair
447
+ * this with a request that actually replays signed thinking blocks before
448
+ * treating it as a thinking-replay rejection.
449
+ */
450
+ export function isAnthropicMaskedProxyRejection(error: unknown): boolean {
451
+ const status = extractHttpStatusFromError(error);
452
+ if (status !== undefined && status !== 400) return false;
453
+ const message = error instanceof Error ? error.message : String(error);
454
+ // A body that still names its error type is classified by the strict matchers.
455
+ if (/invalid_request_error/i.test(message)) return false;
456
+ return /"type"\s*:\s*"api_error"/.test(message) && /an error occurred while processing/i.test(message);
457
+ }
458
+
426
459
  function hasStrictAnthropicTools(params: MessageCreateParamsStreaming): boolean {
427
460
  const tools = params.tools as Array<{ strict?: unknown }> | undefined;
428
461
  return tools?.some(tool => tool.strict === true) ?? false;
@@ -447,12 +480,21 @@ function dropAnthropicStrictTools(params: MessageCreateParamsStreaming): void {
447
480
  }
448
481
  }
449
482
 
483
+ function isClaudeFamilyModel(model: Model<"anthropic-messages">): boolean {
484
+ // Classify the same identifier the request body serializes (`params.model =
485
+ // model.id` in buildParams); a differing `wireModelId` is not dispatched by
486
+ // this transport, so it must not drive the cache decision either.
487
+ const id = model.id;
488
+ const shortId = id.includes("/") ? id.slice(id.lastIndexOf("/") + 1) : id;
489
+ return shortId.toLowerCase().startsWith("claude-");
490
+ }
491
+
450
492
  function getCacheControl(
451
493
  model: Model<"anthropic-messages">,
452
494
  baseUrl: string,
453
495
  cacheRetention?: CacheRetention,
454
496
  ): { mode: AnthropicCacheMode; cacheControl?: AnthropicCacheControl } {
455
- const retention = resolveCacheRetention(cacheRetention, "long");
497
+ const retention = resolveCacheRetention(cacheRetention ?? model.cacheRetention, "long");
456
498
  if (retention === "none") return { mode: "none" };
457
499
 
458
500
  const isCanonicalApi = isAnthropicApiBaseUrl(baseUrl);
@@ -462,7 +504,7 @@ function getCacheControl(
462
504
  ? "none"
463
505
  : promptCacheMode === "explicit"
464
506
  ? "explicit"
465
- : isCanonicalApi
507
+ : promptCacheMode === "automatic" || isCanonicalApi || isClaudeFamilyModel(model)
466
508
  ? "automatic"
467
509
  : "none";
468
510
  if (mode === "none") return { mode };
@@ -1847,7 +1889,12 @@ export const streamAnthropic: StreamFunction<"anthropic-messages"> = (
1847
1889
  !options?.fallbackManaged &&
1848
1890
  !repairAllAssistantThinking &&
1849
1891
  firstTokenTime === undefined &&
1850
- (thinkingSignatureInvalid || isAnthropicThinkingBlockMutationError(streamFailure))
1892
+ (thinkingSignatureInvalid ||
1893
+ isAnthropicThinkingBlockMutationError(streamFailure) ||
1894
+ // Masked proxy rejection: unclassifiable on its own, so the replayed
1895
+ // request shape is the evidence. Without signed thinking blocks in
1896
+ // flight there is nothing to repair and the error must surface.
1897
+ (isAnthropicMaskedProxyRejection(streamFailure) && hasNativeThinkingBlocks(params.messages)))
1851
1898
  ) {
1852
1899
  // The mutation 400 blames the "latest assistant message", but its cited
1853
1900
  // `messages.N.content.M` path can point at an EARLIER replayed turn, so the
@@ -2285,10 +2332,11 @@ function applyExplicitPromptCaching(params: AnthropicCacheParams, cacheControl:
2285
2332
  const currentUser = params.messages[currentUserIndex];
2286
2333
  if (!currentUser) return;
2287
2334
 
2288
- // A tool result is encoded as role "user" on the wire, but belongs to the
2289
- // preceding assistant turn. Anchor that assistant turn, not the tool result,
2290
- // so changing tool output does not invalidate the reusable conversation prefix.
2291
- for (let index = currentUserIndex - 1; index >= 0; index--) {
2335
+ // Tool results are encoded as role "user" on the wire but belong to the
2336
+ // assistant tool-use turn immediately before them. Anchor the latest completed
2337
+ // assistant turn so the reusable prefix advances during an agent tool loop,
2338
+ // while keeping the newest tool output outside the cache boundary.
2339
+ for (let index = params.messages.length - 1; index >= 0; index--) {
2292
2340
  const message = params.messages[index];
2293
2341
  if (message?.role !== "assistant" || !Array.isArray(message.content)) continue;
2294
2342
  if (
@@ -66,6 +66,7 @@ import {
66
66
  toolWireSchema,
67
67
  } from "../utils/schema";
68
68
  import {
69
+ isCodexStatuslessNamedToolChoiceNotFoundError,
69
70
  isForcedToolChoiceUnsupportedError,
70
71
  markToolChoiceIncapability,
71
72
  resolveToolChoice,
@@ -247,9 +248,7 @@ async function retryCodexInitialTransportWithoutToolChoice(
247
248
  requestBodyForState: RequestBody;
248
249
  transport: CodexTransport;
249
250
  }> {
250
- if (
251
- !isForcedToolChoiceUnsupportedError(error, isForcedCodexToolChoice(requestContext.transformedBody.tool_choice))
252
- ) {
251
+ if (!isCodexForcedToolChoiceUnsupportedError(error, requestContext.transformedBody)) {
253
252
  throw error;
254
253
  }
255
254
  const reason = await finalizeErrorMessage(error, requestContext.rawRequestDump);
@@ -1527,7 +1526,7 @@ async function tryRetryWithoutForcedToolChoice(
1527
1526
  context.output.content.length > 0 ||
1528
1527
  context.firstTokenTime !== undefined ||
1529
1528
  context.options?.signal?.aborted ||
1530
- !isForcedToolChoiceUnsupportedError(error, isForcedCodexToolChoice(runtime.requestBodyForState.tool_choice))
1529
+ !isCodexForcedToolChoiceUnsupportedError(error, runtime.requestBodyForState)
1531
1530
  ) {
1532
1531
  return false;
1533
1532
  }
@@ -1579,6 +1578,30 @@ async function tryRetryWithoutForcedToolChoice(
1579
1578
  function isForcedCodexToolChoice(choice: RequestBody["tool_choice"]): boolean {
1580
1579
  return !!choice && choice !== "none" && choice !== "auto";
1581
1580
  }
1581
+ function isCodexForcedToolChoiceUnsupportedError(error: unknown, body: RequestBody): boolean {
1582
+ if (isForcedToolChoiceUnsupportedError(error, isForcedCodexToolChoice(body.tool_choice))) {
1583
+ return true;
1584
+ }
1585
+ return isCodexStatuslessNamedToolChoiceNotFoundError(
1586
+ error,
1587
+ codexNamedFunctionToolChoiceName(body.tool_choice),
1588
+ codexSerializedToolNames(body.tools),
1589
+ );
1590
+ }
1591
+
1592
+ function codexNamedFunctionToolChoiceName(choice: RequestBody["tool_choice"]): string | undefined {
1593
+ if (!choice || typeof choice !== "object") return undefined;
1594
+ const namedChoice = choice as { type?: unknown; name?: unknown };
1595
+ return namedChoice.type === "function" && typeof namedChoice.name === "string" ? namedChoice.name : undefined;
1596
+ }
1597
+
1598
+ function codexSerializedToolNames(tools: RequestBody["tools"]): string[] {
1599
+ if (!Array.isArray(tools)) return [];
1600
+ return tools.flatMap(tool => {
1601
+ const name = (tool as { name?: unknown }).name;
1602
+ return typeof name === "string" ? [name] : [];
1603
+ });
1604
+ }
1582
1605
 
1583
1606
  /**
1584
1607
  * Handles `websocket_connection_limit_reached` errors by closing the stale connection
package/src/types.ts CHANGED
@@ -743,7 +743,11 @@ export type Static<S> = S extends ZodType ? z.infer<S> : S extends { static: inf
743
743
  export type RawArgumentRejectionCode =
744
744
  | "ask-intent-review-requires-positive-round"
745
745
  | "ask-intent-contract-requires-non-empty-authority"
746
- | "ask-deep-interview-metadata-requires-deep-interview-gate";
746
+ | "ask-deep-interview-metadata-requires-deep-interview-gate"
747
+ | "todo-write-unknown-root-key"
748
+ | "todo-write-unknown-op-entry-key"
749
+ | "todo-write-done-drop-requires-target"
750
+ | "todo-write-unknown-init-entry-key";
747
751
 
748
752
  export type RawArgumentValidationResult =
749
753
  | { outcome: "passthrough" }
@@ -947,8 +951,10 @@ export interface AnthropicCompat extends ToolChoiceCompat {
947
951
  supportsLongCacheRetention?: boolean;
948
952
  /**
949
953
  * Prompt-cache transport accepted by this Anthropic-compatible endpoint.
950
- * Canonical Anthropic defaults to `"automatic"`; noncanonical endpoints default
951
- * to `"none"` and must explicitly opt into generated `"explicit"` markers.
954
+ * Canonical Anthropic and Claude-family models default to `"automatic"`;
955
+ * noncanonical non-Claude endpoints default to `"none"`. Set `"automatic"` to
956
+ * opt an otherwise unknown compatible endpoint into top-level caching, `"none"`
957
+ * to opt out, or `"explicit"` for endpoints that require block-level markers.
952
958
  */
953
959
  promptCacheMode?: "none" | "explicit" | "automatic";
954
960
  }
@@ -165,7 +165,10 @@ export function resolveToolChoice(
165
165
 
166
166
  /** Detects provider errors indicating forced tool_choice is unsupported. */
167
167
  export function isForcedToolChoiceUnsupportedError(error: unknown, sentForcedToolChoice: boolean): boolean {
168
- if (!sentForcedToolChoice || extractHttpStatusFromError(error) !== 400) return false;
168
+ const status = extractHttpStatusFromError(error);
169
+ if (!sentForcedToolChoice || status !== 400) {
170
+ return false;
171
+ }
169
172
  const message = errorMessage(error);
170
173
  return (
171
174
  // `by <something>` continuations ("not supported by billing") describe a
@@ -179,6 +182,35 @@ export function isForcedToolChoiceUnsupportedError(error: unknown, sentForcedToo
179
182
  /tool[_\s-]?choices?\s+['"`][^'"`\r\n]+['"`]\s+not\s+found\s+in\s+['"`]tools['"`]\s+parameter\b/is.test(message)
180
183
  );
181
184
  }
185
+ /**
186
+ * Detects Codex's statusless SSE rejection for a named function tool choice.
187
+ * This is intentionally separate from the shared HTTP-400 classifier.
188
+ */
189
+ export function isCodexStatuslessNamedToolChoiceNotFoundError(
190
+ error: unknown,
191
+ forcedToolName: string | undefined,
192
+ sentToolNames: readonly string[],
193
+ ): boolean {
194
+ if (
195
+ extractHttpStatusFromError(error) !== undefined ||
196
+ extractProviderErrorCode(error) !== "invalid_request_error" ||
197
+ !forcedToolName
198
+ ) {
199
+ return false;
200
+ }
201
+ const match =
202
+ /^Tool choice '([^']+)' not found in 'tools' parameter\.$/.exec(errorMessage(error)) ??
203
+ /^Codex error event: Tool choice '([^']+)' not found in 'tools' parameter\. \(code=invalid_request_error\)$/.exec(
204
+ errorMessage(error),
205
+ );
206
+ return match?.[1] === forcedToolName && sentToolNames.includes(forcedToolName);
207
+ }
208
+
209
+ function extractProviderErrorCode(error: unknown): string | undefined {
210
+ if (!error || typeof error !== "object") return undefined;
211
+ const code = (error as { code?: unknown }).code;
212
+ return typeof code === "string" ? code : undefined;
213
+ }
182
214
 
183
215
  export type { ToolChoiceCompat, ToolChoiceSupport, ToolChoiceSupportSource } from "../types";
184
216
 
@@ -965,6 +965,11 @@ const RAW_ARGUMENT_REJECTION_MESSAGES: Record<RawArgumentRejectionCode, string>
965
965
  "deepInterview.intent_contract requires non-empty items and confirmation_options",
966
966
  "ask-deep-interview-metadata-requires-deep-interview-gate":
967
967
  "deepInterview metadata cannot be combined with a non-deep-interview workflowGate",
968
+ "todo-write-unknown-root-key": "todo_write root accepts only an ops array of operation entries",
969
+ "todo-write-unknown-op-entry-key":
970
+ "todo_write operation entries accept only op, list, task, phase, items, and text keys",
971
+ "todo-write-done-drop-requires-target": "todo_write done and drop entries require a task or phase target",
972
+ "todo-write-unknown-init-entry-key": "todo_write init list entries accept only phase and items keys",
968
973
  };
969
974
 
970
975
  /**