opencode-cache-engine 0.3.6 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/docs/cache-policy-inventory.md +313 -0
- package/package.json +1 -1
- package/src/cache-engine-core.mjs +16 -50
- package/src/cache-policy-core.mjs +440 -0
- package/test/cache-engine.test.mjs +235 -0
|
@@ -0,0 +1,313 @@
|
|
|
1
|
+
# CacheEngine Cache-Policy Compatibility Inventory
|
|
2
|
+
|
|
3
|
+
Release: **v0.4.0 — research and documentation only.**
|
|
4
|
+
|
|
5
|
+
This document inventories documented cache behavior for the model boundaries that
|
|
6
|
+
CacheEngine classifies, and states whether CacheEngine's current treatment is
|
|
7
|
+
supported by first-party documentation. It makes **no runtime claims** and does
|
|
8
|
+
**not** propose or apply code changes.
|
|
9
|
+
|
|
10
|
+
- Repository revision for "CacheEngine current treatment": `19b87f2` (`master`).
|
|
11
|
+
- All sources were consulted on **2026-09-26**.
|
|
12
|
+
- Runtime behavior, model detection, provider configuration, telemetry, and GPT
|
|
13
|
+
limits are unchanged by this release.
|
|
14
|
+
|
|
15
|
+
## Evidence classification legend
|
|
16
|
+
|
|
17
|
+
Every substantive statement below is tagged:
|
|
18
|
+
|
|
19
|
+
| Tag | Meaning |
|
|
20
|
+
| --- | --- |
|
|
21
|
+
| **[D]** | Documented fact — stated in a first-party source. |
|
|
22
|
+
| **[O]** | Observed behavior — observed in this repository (tests or hook logic). |
|
|
23
|
+
| **[I]** | Inference — reasoned from documented facts, not stated directly. |
|
|
24
|
+
| **[U]** | Unknown — first-party documentation is insufficient or silent. |
|
|
25
|
+
|
|
26
|
+
Two non-negotiable rules applied throughout:
|
|
27
|
+
|
|
28
|
+
1. A model being newer is **not** evidence that it inherits an older cache policy.
|
|
29
|
+
2. A shared model-name prefix is **not** evidence that two models share cache controls.
|
|
30
|
+
|
|
31
|
+
Caching behavior is never inferred from pricing alone, from one SDK's type
|
|
32
|
+
declarations, or from a third-party blog when first-party documentation exists.
|
|
33
|
+
"unknown — first-party docs insufficient" is used instead of a guess.
|
|
34
|
+
|
|
35
|
+
## CacheEngine current treatment (baseline for the matrix)
|
|
36
|
+
|
|
37
|
+
Source: `src/cache-engine-core.mjs` (detection, transforms) and
|
|
38
|
+
`src/cache-engine.ts` (hooks), revision `19b87f2`.
|
|
39
|
+
|
|
40
|
+
| Family | Detection (verbatim) | Current treatment | Affinity header |
|
|
41
|
+
| --- | --- | --- | --- |
|
|
42
|
+
| DeepSeek | `/deepseek/i` on `${apiID} ${modelID}` or `providerID` | Passive; no mutation | none |
|
|
43
|
+
| GPT-5.6 | `/gpt-5\.6(?![\d.])/i` on slug **and** `isOpenAIish` (provider `openai`/`azure`, slug `openai/`/`azure/`, or npm `@ai-sdk/openai`/`@ai-sdk/azure`) | Inject missing `promptCacheKey` + `promptCacheOptions` (`implicit`, `30m`) | none |
|
|
44
|
+
| GLM-5.3 | `/glm-5\.3(?![\d.])/i` on slug | Relocate identifiable `<env>` block to system tail | `x-session-id` only when `providerID === "openrouter"` |
|
|
45
|
+
| MiMo-V2.6 | `/mimo-v2\.6-(flash\|pro)(?![\w-])/i` on slug | Relocate identifiable `<env>` block to system tail; provider-change telemetry | `x-session-id` only when `providerID === "openrouter"` |
|
|
46
|
+
| Neutral | everything else | Byte-untouched | none |
|
|
47
|
+
|
|
48
|
+
Detection consequences worth stating explicitly:
|
|
49
|
+
|
|
50
|
+
- `gpt-6*` / `gpt-6-*` is **not** matched → neutral. **[O]**
|
|
51
|
+
- `deepseek-v5` (or any future `*deepseek*` id) matches the passive DeepSeek
|
|
52
|
+
branch because the regex is a bare substring test. **[O]**
|
|
53
|
+
- `mimo-v2.6-pro-ultraspeed` is **not** matched: the `(?![\w-])` lookahead fails
|
|
54
|
+
on the following `-`. Test `MiMo V2.5 and Pro-UltraSpeed do NOT match MiMo policy`
|
|
55
|
+
confirms this. **[O]**
|
|
56
|
+
- `mimo-v2.5*` and `glm-5.2`/`glm-4.x` are neutral. **[O]**
|
|
57
|
+
|
|
58
|
+
---
|
|
59
|
+
|
|
60
|
+
## 1. OpenAI (GPT-5.6 and later)
|
|
61
|
+
|
|
62
|
+
OpenAI documents a **generation-boundary** cache policy, literally named
|
|
63
|
+
"GPT-5.6 and later", plus a distinct class for "GPT-5.5 and GPT-5.5 Pro" and a
|
|
64
|
+
catch-all "earlier models". [D]
|
|
65
|
+
|
|
66
|
+
| # | Item | Finding | Tag |
|
|
67
|
+
| --- | --- | --- | --- |
|
|
68
|
+
| 1 | Model names/aliases | GPT-5.6 family: `gpt-5.6-sol` (flagship), `gpt-5.6-terra`, `gpt-5.6-luna`, `gpt-5.6-cyber`; `gpt-5.6` is an alias for `gpt-5.6-sol`. GPT-6 family: `gpt-6-astra` (flagship), `gpt-6-sol`, `gpt-6-luna`. Aliases `gpt-daybreak-blue-latest` → `gpt-5.6-sol`, `gpt-daybreak-red-latest` → `gpt-5.6-cyber`. | [D] |
|
|
69
|
+
| 2 | Family identification | Pattern `gpt-<generation>[-<codename>][-<variant>]`; unsuffixed id maps to one tier. The generalized rule is inference; the concrete IDs/aliases are documented. | [D]/[I] |
|
|
70
|
+
| 3 | Scope | Generation-boundary-specific ("GPT-5.6 and later" covers GPT-5.6 and GPT-6). Not creator-wide, not exact-model. Cache availability remains a per-model supported feature. | [D] |
|
|
71
|
+
| 4 | Automatic vs controls | Enabled by default (implicit). GPT-5.6+ additionally supports explicit opt-in via `prompt_cache_options.mode="explicit"` + per-block `prompt_cache_breakpoint`, and `prewarm`. | [D] |
|
|
72
|
+
| 5 | Cacheable-prefix rules | Reuse requires the **entire rendered prefix** to match, including hidden instructions, developer messages, tool definitions (names/descriptions/schemas/order), and history. `model`, `tools`, `parallel_tool_calls`, `text.format`, `reasoning.effort`, `text.verbosity`, `context_management`, `service_tier`, `prompt_cache_key` affect matching. | [D] |
|
|
73
|
+
| 6 | Minimum cacheable length | **1,024 visible input tokens** for GPT-5.6 and later (fixed). Pre-5.6 length varies by request settings; exact numbers not enumerated. | [D]/[U] |
|
|
74
|
+
| 7 | Implicit vs explicit breakpoints | GPT-5.6+ supports both; explicit mode with no developer breakpoints caches nothing; up to 4 cache writes/request. Breakpoints valid on `input_text`/`input_image`/`input_file` and `function_call_output`, **not** on top-level `instructions` or `additional_tools`. Pre-5.6 is implicit only (e.g. GPT-5.5/Pro at 2,048-token intervals). | [D] |
|
|
75
|
+
| 8 | Cache key behavior | `prompt_cache_key` is optional on GPT-5.6+ (routing is automatic; key used for per-customer accounting/anti-probing). On pre-5.6 it is the routing-optimization control. A key influences routing; it does not pin a machine or guarantee a hit. | [D] |
|
|
76
|
+
| 9 | Retention/TTL | GPT-5.6+: `prompt_cache_options.ttl`, only supported value `30m` (also default); reuse refreshes TTL with no re-write charge. Pre-5.6: `prompt_cache_retention` = `in_memory` or `24h`. GPT-5.5/Pro: `24h` only. | [D] |
|
|
77
|
+
| 10 | Usage fields | Responses: `usage.input_tokens_details.cached_tokens`, `usage.input_tokens_details.cache_write_tokens`, plus `prompt_cache_diagnostics` (`cache_hit`/`cache_miss`, `reason`, `comparison_reusable_tokens`, `cache_missed_tokens`). Chat Completions field naming not confirmed first-party. | [D]/[U] |
|
|
78
|
+
| 11 | Generation differences | gpt-4o: 0.5× cached, no write charge. gpt-5/5.1/5.5: 0.1×, no write charge. GPT-5.6+: write 1.25×, read 0.1×, min 1,024, TTL 30m, explicit+both modes, exact (non-rounded) boundary reporting. GPT-6 inherits the GPT-5.6 policy. | [D] |
|
|
79
|
+
| 12 | Prefix-stability guidance | Put stable developer instructions first; move timestamps/user-specific content later. Append messages; do not rewrite/compact/truncate earlier turns. Keep tool definitions/order/schemas stable; disable tools via `tool_choice`/`allowed_tools` rather than removing definitions. | [D] |
|
|
80
|
+
| 13 | Affinity/routing | Cache entries are machine-local; routing is automatic and depends on machine load and a hash of initial tokens (incl. tool definitions), plus `prompt_cache_key` pre-5.6. Traffic >15 rpm can overflow to other machines; keys do not pin machines. | [D] |
|
|
81
|
+
| 14 | Context/limits (cache-relevant) | `gpt-5.6-sol` and GPT-6: 1,050,000 context / 922,000 max input / 128,000 max output. Input >272K tokens priced at 2× input (and 2× cache rates) and 1.5× output for the whole request. | [D] |
|
|
82
|
+
|
|
83
|
+
**OpenAI sources** (title — URL, consulted 2026-09-26; scope in parentheses):
|
|
84
|
+
|
|
85
|
+
- OpenAI, *Prompt caching* — https://developers.openai.com/api/docs/guides/prompt-caching (all OpenAI models; the authority for the items above).
|
|
86
|
+
- OpenAI, *Prompt cache diagnostics* — https://developers.openai.com/api/docs/guides/prompt-caching/diagnostics.
|
|
87
|
+
- OpenAI, *Models* — https://developers.openai.com/api/docs/models; *Compare models* — https://developers.openai.com/api/docs/models/compare.
|
|
88
|
+
- OpenAI, *GPT-5.6 Sol* — https://developers.openai.com/api/docs/models/gpt-5.6-sol (exact model page: `gpt-5.6-sol`).
|
|
89
|
+
- OpenAI, *GPT-6 Astra* — https://developers.openai.com/api/docs/models/gpt-6-astra (exact model page: `gpt-6-astra`).
|
|
90
|
+
- OpenAI, *Using GPT-6* — https://developers.openai.com/api/docs/guides/latest-model (GPT-6 family).
|
|
91
|
+
- OpenAI, Responses API reference — https://developers.openai.com/api/reference/resources/responses/methods/create.
|
|
92
|
+
- OpenAI, *Pricing* — https://developers.openai.com/api/docs/pricing.
|
|
93
|
+
|
|
94
|
+
**CacheEngine compatibility:** GPT-5.6 handling is consistent with the documented
|
|
95
|
+
baseline: it injects only `promptCacheKey` + `promptCacheOptions{mode:implicit,
|
|
96
|
+
ttl:"30m"}`, preserves runtime-supplied values, and never sets context/output
|
|
97
|
+
limits. It does not use explicit breakpoints or `prewarm`, which is a subset of
|
|
98
|
+
the documented capability. GPT-6 is **not** classified and therefore gets no GPT
|
|
99
|
+
cache metadata — a documented-scope gap, not a documented incompatibility.
|
|
100
|
+
|
|
101
|
+
---
|
|
102
|
+
|
|
103
|
+
## 2. DeepSeek (V4 and later)
|
|
104
|
+
|
|
105
|
+
DeepSeek documents context caching as **provider-wide, automatic, implicit disk
|
|
106
|
+
caching** with no request control. [D]
|
|
107
|
+
|
|
108
|
+
| # | Item | Finding | Tag |
|
|
109
|
+
| --- | --- | --- | --- |
|
|
110
|
+
| 1 | Model names/aliases | Current API ids: `deepseek-flash`, `deepseek-v4-pro`. Legacy (`deepseek-v4-flash`, `deepseek-v4-flash-vision-exp`) retired but still routed to DeepSeek-V4.1-Flash. `deepseek-chat`/`deepseek-reasoner` retired 2026-07-24. `deepseek-v4.1` and a bare `deepseek-v4` are **not** request identifiers ("DeepSeek-V4.1-Flash" is only a MODEL VERSION string). | [D] |
|
|
111
|
+
| 2 | Family identification | Human-readable MODEL VERSION row and change-log categories ("V4 model family"). No machine-readable family field is documented; `system_fingerprint` is described only as "backend configuration". | [D]/[U] |
|
|
112
|
+
| 3 | Scope | Caching guide is provider-wide ("default for all users"); cache-hit pricing listed for both current models. No doc states V4+/exact-model scoping. | [D]/[U] |
|
|
113
|
+
| 4 | Automatic vs controls | "Enabled by default… without needing to modify their code"; no flag/field documented. OpenRouter independently states DeepSeek caching is automated and needs no configuration (transport page). | [D] |
|
|
114
|
+
| 5 | Cacheable-prefix rules | Hits require a full match of a **cache prefix unit** (independent, complete units). Units persist at user-input end, model-output end, on detected common prefixes, and at fixed token intervals. Historical rule: only identical prefixes "starting from the 0th token" hit. Post-SWA matching differs from before. | [D] |
|
|
115
|
+
| 6 | Minimum cacheable length | "**64 tokens** as a storage unit; content less than 64 tokens will not be cached" — stated on the 2024 announcement page only, not restated on the current guide/pricing page. Possible staleness for V4+. | [D]/[I] |
|
|
116
|
+
| 7 | Implicit vs explicit breakpoints | Implicit only. No `cache_control`/breakpoint/cache-creation endpoint appears in the request schema. | [D] |
|
|
117
|
+
| 8 | Cache key behavior | No cache-key field. `user_id` is documented for KVCache isolation and scheduling isolation; not a cache key per se. | [D]/[U] |
|
|
118
|
+
| 9 | Retention/TTL | "Cleared… usually within a few hours to a few days"; "cache construction takes seconds". No numeric TTL. | [D] |
|
|
119
|
+
| 10 | Usage fields | `usage.prompt_cache_hit_tokens`, `usage.prompt_cache_miss_tokens`; `usage.prompt_tokens = hit + miss`; `usage.prompt_tokens_details.cached_tokens` = same as hit tokens. | [D] |
|
|
120
|
+
| 11 | Generation differences | V2: MLA caching feasibility. Sliding Window Attention changed prefix storage/matching (transition version not named in the cache guide; change log shows V3.2-Exp as the SWA/DSA point). V4: token-wise compression + DSA, 1M standard. V4.1-Flash: "smallest model in our new architecture family". No doc states cache semantics differ between V4.1-Flash and V4-Pro. | [D]/[I] |
|
|
121
|
+
| 12 | Prefix-stability guidance | Mechanism only; no explicit "keep system/tools/order stable" instruction. Documented examples: `A+B`→`A+B+C` hits; `A+B`→`A+C` does not (until common prefix `A` is recognized). | [D]/[U] |
|
|
122
|
+
| 13 | Affinity/routing | `user_id` isolates KVCache; no cache-affinity routing guarantee documented. Each user's cache is isolated. | [D] |
|
|
123
|
+
| 14 | Direct vs OpenRouter | Direct API: automatic disk caching on both current models. OpenRouter: lists DeepSeek under automated caching, cache reads 0.1× input, `cached_tokens`/`cache_write_tokens` in `prompt_tokens_details`; OpenRouter does not enumerate the current V4 names. | [D] |
|
|
124
|
+
|
|
125
|
+
**DeepSeek sources** (consulted 2026-09-26; scope in parentheses):
|
|
126
|
+
|
|
127
|
+
- DeepSeek, *Context Caching* — https://api-docs.deepseek.com/guides/kv_cache (provider-wide; default for all users).
|
|
128
|
+
- DeepSeek, *Models & Pricing* — https://api-docs.deepseek.com/quick_start/pricing (`deepseek-flash`, `deepseek-v4-pro`).
|
|
129
|
+
- DeepSeek, *Chat Completions API* — https://api-docs.deepseek.com/api/create-chat-completion (request/usage schema).
|
|
130
|
+
- DeepSeek, *Rate Limit & Isolation* — https://api-docs.deepseek.com/quick_start/rate_limit (`user_id` isolation).
|
|
131
|
+
- DeepSeek, *Change Log* — https://api-docs.deepseek.com/updates; *news260424* — https://api-docs.deepseek.com/news/news260424; *news260910* — https://api-docs.deepseek.com/news/news260910; *news0802* — https://api-docs.deepseek.com/news/news0802 (2024 64-token note; older matching rule).
|
|
132
|
+
- OpenRouter, *Prompt Caching* — https://openrouter.ai/docs/features/prompt-caching (transport layer only).
|
|
133
|
+
|
|
134
|
+
**CacheEngine compatibility:** Passive DeepSeek treatment matches the documented
|
|
135
|
+
baseline exactly (automatic, no controls, no key). The `/deepseek/i` substring
|
|
136
|
+
detector means any future `deepseek-v5`-style id also becomes passive, which is
|
|
137
|
+
the safe direction. No CacheEngine optimization is required for DeepSeek.
|
|
138
|
+
|
|
139
|
+
---
|
|
140
|
+
|
|
141
|
+
## 3. Z.AI GLM (5.3 and later)
|
|
142
|
+
|
|
143
|
+
Z.AI documents **implicit, automatic context caching**; no explicit cache
|
|
144
|
+
control, cache key, or numeric TTL is documented. [D]
|
|
145
|
+
|
|
146
|
+
| # | Item | Finding | Tag |
|
|
147
|
+
| --- | --- | --- | --- |
|
|
148
|
+
| 1 | Model names/aliases | "GLM-5.3" is documented as current flagship. Enum includes `glm-5.3`, `glm-5.2`, `glm-5.1`, `glm-5`, `glm-4.7`, `glm-4.7-flash`, `glm-4.7-flashx`, `glm-4.6`, `glm-4.5`, `glm-4.5-air/x/airx/flash`, `glm-4-32b-0414-128k`; vision adds `glm-5.3-flashx`, `glm-5.3-flash`, `glm-4.6v*`, `glm-4.5v`. Release order 4.5→4.6→4.7→5→5.1→5.2→5.3. No alias mechanism documented. | [D] |
|
|
149
|
+
| 2 | Family identification | Literal `model` id string only; no family field. Param docs group "GLM-5.3/5.2/5.1/5/4.7/4.6 series" vs "GLM-4.5 series". GLM-5.3 shares a base model with GLM-5.2, differing by post-training. | [D] |
|
|
150
|
+
| 3 | Scope | Implicit mechanism is service-wide, but cacheability and cached-input price are **exact-model-specific** (e.g. `GLM-4-32B-0414-128K` has no cached-input price). | [D]/[I] |
|
|
151
|
+
| 4 | Automatic vs controls | "Automatic Cache Recognition: Implicit caching… without manual configuration." No `cache_control`, `prompt_cache_key`, breakpoint, or TTL parameter in the API. | [D] |
|
|
152
|
+
| 5 | Cacheable-prefix rules | Identifies content "identical or highly similar to previous requests"; identical content hits best; "minor formatting differences may affect cache effectiveness". Recommended layout: stable instructions first in the system prompt, variable content last. | [D] |
|
|
153
|
+
| 6 | Minimum cacheable length | No hard minimum documented. Soft guideline only: repeated prefix "recommend 500+ tokens"; two or three short system sentences usually will not hit. | [D]/[U] |
|
|
154
|
+
| 7 | Implicit vs explicit breakpoints | Implicit only. No explicit breakpoint documented. | [D] |
|
|
155
|
+
| 8 | Cache key behavior | No user-supplied or derived cache key documented by Z.AI. | [U] |
|
|
156
|
+
| 9 | Retention/TTL | "Cache has reasonable time limits, will recalculate after expiration"; asynchronous effect. No numeric expiry. Cached-input storage listed as limited-time free. | [D]/[U] |
|
|
157
|
+
| 10 | Usage fields | `usage.prompt_tokens_details.cached_tokens` ("tokens served from cache"). No cache-write/creation field documented (writes are not billed/reported). | [D] |
|
|
158
|
+
| 11 | Generation differences | No cache-semantics difference stated per generation. Differences are pricing (cached-input $0.26 for 5.3/5.2/5.1, $0.2 for 5, $0.11 for 4.7/4.6/4.5, etc.), support (no cache price for `GLM-4-32B-0414-128K`), and unrelated context/output changes. | [D] |
|
|
159
|
+
| 12 | Prefix-stability guidance | Use stable system-prompt templates; put long documents as system messages; manage history; avoid frequent content changes; monitor hit rate. Stable rules/knowledge first, variable content last. No tool-order guidance. | [D] |
|
|
160
|
+
| 13 | Affinity/routing | Z.AI documents **no** affinity/session/routing parameter; direct-endpoint cache-reuse controls are unknown. (OpenRouter's sticky routing is transport, not Z.AI semantics.) | [D]/[U] |
|
|
161
|
+
| 14 | Direct vs OpenRouter | Direct endpoints: cache docs example `api.z.ai/api/paas/v4/chat/completions` and `open.bigmodel.cn/api/paas/v4/chat/completions`, returning `cached_tokens`. OpenRouter route: automated, no configuration, reads reported in `cached_tokens`, plus sticky routing. | [D] |
|
|
162
|
+
|
|
163
|
+
**Z.AI GLM sources** (consulted 2026-09-26; scope in parentheses):
|
|
164
|
+
|
|
165
|
+
- Z.AI, *Context Caching* — https://docs.z.ai/guides/capabilities/cache.md (service-wide; GLM families).
|
|
166
|
+
- BigModel (Z.AI studio), *上下文缓存* — https://docs.bigmodel.cn/cn/guide/capabilities/cache.md (same feature; 500+ token recommendation; layout advice).
|
|
167
|
+
- Z.AI, *Chat Completion* API — https://docs.z.ai/api-reference/llm/chat-completion.md (model enum; `prompt_tokens_details`).
|
|
168
|
+
- Z.AI, *Pricing* — https://docs.z.ai/guides/overview/pricing.md (per-model cached-input rates).
|
|
169
|
+
- Z.AI, *GLM-5.3* — https://docs.z.ai/guides/llm/glm-5.3.md; *Migrate to GLM-5.3* — https://docs.z.ai/guides/overview/migrate-to-glm-new.md; *New Released* — https://docs.z.ai/release-notes/new-released.md.
|
|
170
|
+
- OpenRouter, *Prompt Caching* — https://openrouter.ai/docs/features/prompt-caching (transport only).
|
|
171
|
+
|
|
172
|
+
**CacheEngine compatibility:** GLM env-block relocation is a CacheEngine
|
|
173
|
+
optimization consistent with Z.AI's documented "stable prefix first, variable
|
|
174
|
+
content last" guidance, but it is **not** a Z.AI-documented control — it is an
|
|
175
|
+
exact-model overlay. `x-session-id` is an OpenRouter transport affordance, not a
|
|
176
|
+
Z.AI-documented request parameter.
|
|
177
|
+
|
|
178
|
+
---
|
|
179
|
+
|
|
180
|
+
## 4. Xiaomi MiMo (V2.6 and later)
|
|
181
|
+
|
|
182
|
+
Xiaomi documents per-model "Context Caching" and prefix-hit billing but describes
|
|
183
|
+
**no** request control, cache key, minimum length, TTL, or prefix-stability
|
|
184
|
+
guidance. [D]/[U]
|
|
185
|
+
|
|
186
|
+
| # | Item | Finding | Tag |
|
|
187
|
+
| --- | --- | --- | --- |
|
|
188
|
+
| 1 | Model names/aliases | API ids: `mimo-v2.6-pro`, `mimo-v2.6-flash`, `mimo-v2.6-pro-ultraspeed`, `mimo-v2.5-pro`, `mimo-v2.5`. OpenRouter slugs: `xiaomi/mimo-v2.6-flash`, `xiaomi/mimo-v2.6-pro`, `xiaomi/mimo-v2.6-pro-ultraspeed`. HF checkpoints: `MiMo-V2.6-Pro-RL`, `MiMo-V2.6-Flash-RL`, `MiMo-V2.6-Distill-Qwen-9B`. UltraSpeed is documented as a **mode** of Pro, not a separate checkpoint. | [D] |
|
|
189
|
+
| 2 | Family identification | "MiMo-V2.6 series" = Pro + Flash (both native multimodal); Pro also offers UltraSpeed mode. HF collection "MiMo-V2.6". | [D] |
|
|
190
|
+
| 3 | Scope | Per-model capability list shows "Context Caching" for v2.6-flash (and indexed model pages for pro/ultraspeed/v2.5). Whether the implementation is shared family-wide vs per exact model is not stated. | [D]/[U] |
|
|
191
|
+
| 4 | Automatic vs controls | No cache-control field in any of the three protocols (OpenAI, Anthropic, Responses). Pricing describes cache hits purely from prefix content hitting the cache. Operation is provider-managed/implicit. | [D]/[I] |
|
|
192
|
+
| 5 | Cacheable-prefix rules | Only "the requested prefix content" is stated; no prefix-boundary/ordering/role definition. | [D]/[U] |
|
|
193
|
+
| 6 | Minimum cacheable length | Not documented. | [U] |
|
|
194
|
+
| 7 | Implicit vs explicit breakpoints | No explicit breakpoint mechanism documented in any request schema. | [D]/[U] |
|
|
195
|
+
| 8 | Cache key behavior | No user-supplied cache key or derivation rule documented. | [U] |
|
|
196
|
+
| 9 | Retention/TTL | No prompt-cache TTL/expiry documented. (The docs' "5-minute cache period" refers to the Web Search plugin toggle — **not** the prompt cache.) | [U] |
|
|
197
|
+
| 10 | Usage fields | OpenAI Chat: `usage.prompt_tokens_details.cached_tokens`. Responses: `usage.input_tokens_details.cached_tokens`. Anthropic: `usage.cache_read_input_tokens`. No Xiaomi cache-**write** field (Cache Write is limited-time free). | [D] |
|
|
198
|
+
| 11 | Generation differences | v2.5-pro/v2.5 deprecated 2026-10-21; v2.6 current. UltraSpeed exists as a V2.5-Pro limited-time mode and as `mimo-v2.6-pro-ultraspeed` ("up to 20x" speed). Batch API supports only `mimo-v2.6-pro`/`mimo-v2.6-flash`. Cache-hit pricing differs per model. No documented caching-**mechanism** difference across flash/pro/ultraspeed or v2.5/v2.6. | [D]/[U] |
|
|
199
|
+
| 12 | Prefix-stability guidance | None found in Xiaomi docs. | [U] |
|
|
200
|
+
| 13 | Affinity/routing | Xiaomi documents no session-affinity parameter; only `api-key`/Bearer auth. OpenRouter sticky routing is transport, not Xiaomi semantics. | [D]/[U] |
|
|
201
|
+
| 14 | Direct vs OpenRouter | OpenRouter's own Prompt Caching page has per-provider sections for OpenAI/Anthropic/Groq/Grok/Moonshot/Alibaba/DeepSeek/Z.AI/Google — **MiMo is absent**, so OpenRouter documents no MiMo-specific cache semantics. | [D] |
|
|
202
|
+
|
|
203
|
+
**Xiaomi MiMo sources** (consulted 2026-09-26; scope in parentheses):
|
|
204
|
+
|
|
205
|
+
- MiMo, *OpenAI Chat Completions API Compatibility* — https://mimo.mi.com/docs/en-US/api/chat/openai-api (all MiMo API ids; usage fields).
|
|
206
|
+
- MiMo, *Anthropic Messages API Compatibility* — https://mimo.mi.com/docs/en-US/api/chat/anthropic-api.
|
|
207
|
+
- MiMo, *OpenAI Responses API Compatibility* — https://mimo.mi.com/docs/en-US/api/chat/responses.
|
|
208
|
+
- MiMo, *API Pricing* — https://mimo.mi.com/docs/en-US/price/pay-as-you-go (per-model cache-hit/write pricing).
|
|
209
|
+
- MiMo, *Models* — https://mimo.mi.com/docs/en-US/quick-start/summary/model (v2.5 deprecation; capabilities).
|
|
210
|
+
- MiMo, *MiMo-V2.6-Flash* — https://mimo.mi.com/models/en-US/mimo-v2.6-flash ("Context Caching" capability).
|
|
211
|
+
- MiMo, *MiMo-V2.6: Scaling Up RL for Self-Improvement* — https://mimo.mi.com/docs/en-US/news/latest/v2-6 (series naming, UltraSpeed mode).
|
|
212
|
+
- Hugging Face, `XiaomiMiMo/MiMo-V2.6-Flash-RL` and `MiMo-V2.6-Pro-RL` model cards (checkpoint names).
|
|
213
|
+
- OpenRouter, *Prompt Caching* — https://openrouter.ai/docs/features/prompt-caching; *Provider Routing* — https://openrouter.ai/docs/features/provider-routing (transport only).
|
|
214
|
+
|
|
215
|
+
**CacheEngine compatibility:** MiMo env-block relocation is not supported by any
|
|
216
|
+
Xiaomi prefix-stability instruction ([U]); it is an exact-model overlay carried
|
|
217
|
+
over from the GLM family. MiMo `cachedTokens/promptTokens` telemetry matches the
|
|
218
|
+
documented `cached_tokens`/`prompt_tokens` fields. `x-session-id` is documented
|
|
219
|
+
by OpenRouter (transport) but not by Xiaomi.
|
|
220
|
+
|
|
221
|
+
---
|
|
222
|
+
|
|
223
|
+
## 5. OpenRouter transport facts (routing only, not model semantics)
|
|
224
|
+
|
|
225
|
+
These are OpenRouter routing/transport facts. They are **not** evidence of any
|
|
226
|
+
creator's cache semantics.
|
|
227
|
+
|
|
228
|
+
- OpenRouter uses **provider sticky routing** to keep follow-up requests on the
|
|
229
|
+
provider that served a cached request; it activates when cache-read pricing is
|
|
230
|
+
below normal prompt pricing, and falls back if the sticky provider is
|
|
231
|
+
unavailable. [D]
|
|
232
|
+
- Sticky granularity is account-level, per model, per conversation. The default
|
|
233
|
+
conversation key is a hash of the first system/developer message plus the first
|
|
234
|
+
non-system message. [D]
|
|
235
|
+
- An explicit top-level body `session_id` **or** `x-session-id` header replaces
|
|
236
|
+
the derived key (body wins if both; max 256 chars). Without it, stickiness
|
|
237
|
+
activates only after a cache hit. Sticky sessions expire after **10 minutes of
|
|
238
|
+
inactivity** and are disabled when manual `provider.order` is set. [D]
|
|
239
|
+
- Source: OpenRouter, *Prompt Caching* — https://openrouter.ai/docs/features/prompt-caching
|
|
240
|
+
(consulted 2026-09-26).
|
|
241
|
+
|
|
242
|
+
CacheEngine's `x-session-id` injection is therefore a **transport-specific**
|
|
243
|
+
behavior for `providerID === "openrouter"` only, which matches OpenRouter's
|
|
244
|
+
documented header name. It is not a creator-documented cache control.
|
|
245
|
+
|
|
246
|
+
---
|
|
247
|
+
|
|
248
|
+
## 6. CacheEngine Compatibility Matrix
|
|
249
|
+
|
|
250
|
+
Legend for "recommended family inheritance": **keep** = current treatment
|
|
251
|
+
matches docs; **extend** = docs support broadening scope (a future change, not
|
|
252
|
+
made here); **hold** = do not inherit without first-party evidence.
|
|
253
|
+
|
|
254
|
+
| Creator | Model / example pattern | Cache policy (documented) | CacheEngine current treatment | Recommended family inheritance | Recommended exact-model exception | Confidence | Source | Verified |
|
|
255
|
+
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
|
|
256
|
+
| OpenAI | `gpt-5.6`, `gpt-5.6-*` (sol/terra/luna/cyber) | "GPT-5.6 and later": implicit default, optional explicit breakpoints, min 1,024, TTL 30m, write 1.25×/read 0.1× | GPT policy: inject `promptCacheKey` + `promptCacheOptions{implicit,30m}` | keep | none documented; CacheEngine's subset is valid | High (docs) / Medium (treatment) | OpenAI *Prompt caching*; *GPT-5.6 Sol* | 2026-09-26 |
|
|
257
|
+
| OpenAI | GPT-6 / current later GPT family: `gpt-6`, `gpt-6-*` (astra/sol/luna) | Inherits the GPT-5.6-and-later policy | **Not matched** → neutral (no GPT metadata) | **extend**: docs classify GPT-6 as "later", so it inherits the same policy | none documented | High (docs) / High (gap) | OpenAI *Using GPT-6*; *GPT-6 Astra*; *Prompt caching* | 2026-09-26 |
|
|
258
|
+
| OpenAI | pre-5.6 negative controls: `gpt-5.5`, `gpt-5.4`, `gpt-5.2`, `gpt-5.1`, `gpt-5`, `gpt-4.1`, `gpt-4o` | Implicit only; different min-length class; `in_memory`/`24h` retention; `prompt_cache_key` for routing | neutral | **hold** — do not inherit 5.6 policy | n/a | High | OpenAI *Prompt caching*; *Pricing* | 2026-09-26 |
|
|
259
|
+
| DeepSeek | V4: `deepseek-v4-pro`, legacy `deepseek-v4-flash` | Provider-wide automatic disk cache; implicit; prefix-unit matching; hit/miss token fields | Passive (no mutation) | keep | none | High | DeepSeek *Context Caching*; *Models & Pricing* | 2026-09-26 |
|
|
260
|
+
| DeepSeek | V4.1 / current V4-family: `deepseek-flash` (MODEL VERSION "DeepSeek-V4.1-Flash") | Same provider-wide automatic policy; cache-hit pricing listed for both current models | Passive | keep (creator/family baseline) | none documented | High | DeepSeek *Models & Pricing*; *news260910* | 2026-09-26 |
|
|
261
|
+
| DeepSeek | future-looking V4+ identifiers: `deepseek-v4.1`, `deepseek-v4`, `deepseek-v5` | Not documented as request ids (`deepseek-v4.1`/`deepseek-v4` invalid or version-string only) | Passive via `/deepseek/i` substring (e.g. `deepseek-v5` already passive) | keep passive; treat as unknown-friendly | none | Medium (detection) / Low (future ids) | DeepSeek *Models & Pricing*; *Chat Completions API* | 2026-09-26 |
|
|
262
|
+
| DeepSeek | pre-V4 negative controls: `deepseek-chat`, `deepseek-reasoner` | Retired names (retired 2026-07-24); no separate V4+ cache policy claimed | Passive | hold | n/a | High | DeepSeek *Change Log*; *news260424* | 2026-09-26 |
|
|
263
|
+
| Z.AI | GLM 5.3: `glm-5.3`, `glm-5.3-flash`, `glm-5.3-flashx` | Implicit automatic caching; `cached_tokens`; stable-prompt-first guidance; no documented min/TTL/key | GLM policy: `<env>` relocation; OpenRouter `x-session-id` | keep (env relocation is an exact overlay, not a Z.AI control) | env relocation is the overlay; keep scoped to GLM-5.3 | Medium | Z.AI *Context Caching*; *Chat Completion*; *Pricing* | 2026-09-26 |
|
|
264
|
+
| Z.AI | current later GLM generations (documented): none newer than 5.3; newest below is `glm-5.2`/`glm-5.1`/`glm-5`/`glm-4.7` | Same implicit mechanism documented service-wide; cached-input price per model | neutral (only `glm-5.3` matched) | **hold** — no doc says 5.3 overlay extends upward; none newer documented | n/a | High (no later gens documented) | Z.AI *New Released*; *Pricing* | 2026-09-26 |
|
|
265
|
+
| Z.AI | 5.2 and earlier negative controls: `glm-5.2`, `glm-5.1`, `glm-5`, `glm-4.7`, `glm-4.6`, `glm-4.5`, `glm-4-32b-*` | Cacheable (except `glm-4-32b-0414-128k`), different cached-input pricing; no cache-semantics difference documented | neutral | hold | n/a | High | Z.AI *Pricing*; *Chat Completion* (enum) | 2026-09-26 |
|
|
266
|
+
| Xiaomi | MiMo V2.6 Flash: `mimo-v2.6-flash` (OR `xiaomi/mimo-v2.6-flash`) | Provider-managed implicit caching; `cached_tokens`; no documented min/TTL/key/prefix rules | MiMo policy: `<env>` relocation; provider-change telemetry; OpenRouter `x-session-id` | keep scoped as exact-model overlay | env relocation unsupported by docs → treat as overlay | Medium | MiMo *Models*; *Pricing*; *openai-api* | 2026-09-26 |
|
|
267
|
+
| Xiaomi | MiMo V2.6 Pro: `mimo-v2.6-pro` (OR `xiaomi/mimo-v2.6-pro`) | Same documented per-model implicit caching; per-model pricing | MiMo policy (same as Flash) | keep | none documented | Medium | MiMo *Models*; *Pricing* | 2026-09-26 |
|
|
268
|
+
| Xiaomi | MiMo V2.6 Pro UltraSpeed: `mimo-v2.6-pro-ultraspeed` (OR `xiaomi/mimo-v2.6-pro-ultraspeed`) | Documented as a Pro **mode**, same V2.6 series; "Context Caching" listed; no documented mechanism difference from Pro | **Not matched** → neutral (`(?![\w-])` blocks it) | **extend OR document as exception** — docs treat it as the same V2.6 series, but no doc proves identical cache controls | explicit decision needed; current behavior is a de-facto exact-model exception | Medium | MiMo *news/latest/v2-6*; *Models*; *Pricing* | 2026-09-26 |
|
|
269
|
+
| Xiaomi | later MiMo generations: none documented beyond V2.6; V2.5 deprecates 2026-10-21 | Not documented | neutral | hold (no docs) | n/a | High (no later gens documented) | MiMo *Models* | 2026-09-26 |
|
|
270
|
+
| Xiaomi | V2.5 negative control: `mimo-v2.5`, `mimo-v2.5-pro` | Documented as deprecated 2026-10-21; cache capability listed; no V2.6 policy inheritance claimed | neutral | hold | n/a | High | MiMo *Models*; *news/latest/v2-6* | 2026-09-26 |
|
|
271
|
+
|
|
272
|
+
---
|
|
273
|
+
|
|
274
|
+
## 7. Cases where documentation is insufficient to establish inheritance
|
|
275
|
+
|
|
276
|
+
These are explicit gaps. For each, first-party docs do **not** establish whether
|
|
277
|
+
a newer model inherits an older policy. They must not be resolved by guessing.
|
|
278
|
+
|
|
279
|
+
1. **OpenAI GPT-6 → GPT-5.6 policy inheritance.** The cache guide literally
|
|
280
|
+
groups "GPT-5.6 and later", and the *Using GPT-6* guide treats GPT-6 as the
|
|
281
|
+
current flagship, so inheritance is supported by the generation-boundary
|
|
282
|
+
wording. However, no GPT-6 model page enumerates a cache-behavior delta beyond
|
|
283
|
+
inherited features; the inventory treats inheritance as **[D]** but notes no
|
|
284
|
+
GPT-6-specific cache documentation exists.
|
|
285
|
+
2. **DeepSeek V4.1 vs V4-Pro cache-semantics difference.** Not documented. Only
|
|
286
|
+
pricing differs; caching is described provider-wide.
|
|
287
|
+
3. **DeepSeek minimum cacheable length for V4+.** The only number (64 tokens) is
|
|
288
|
+
from a 2024 page and is not restated for V4+; treat as **stale/unverified**.
|
|
289
|
+
4. **DeepSeek TTL.** "Hours to days" only; no numeric TTL.
|
|
290
|
+
5. **Z.AI minimum cacheable length and TTL.** No hard minimum (only a 500+ token
|
|
291
|
+
recommendation); no numeric TTL.
|
|
292
|
+
6. **Z.AI cache key.** No key documented.
|
|
293
|
+
7. **Z.AI generations newer than 5.3.** None documented; cannot establish
|
|
294
|
+
upward inheritance.
|
|
295
|
+
8. **Xiaomi prompt-cache prefix rules, minimum length, key, and TTL.** All
|
|
296
|
+
undocumented.
|
|
297
|
+
9. **Xiaomi V2.6 Pro UltraSpeed cache mechanism.** Documented as the same V2.6
|
|
298
|
+
series and a Pro mode, but no first-party statement proves identical cache
|
|
299
|
+
controls to Pro/Flash. This is the one place where CacheEngine's current
|
|
300
|
+
non-match (`mimo-v2.6-pro-ultraspeed` → neutral) is a real scope decision
|
|
301
|
+
rather than a documented fact.
|
|
302
|
+
10. **MiMo on OpenRouter.** OpenRouter's Prompt Caching page documents no
|
|
303
|
+
MiMo-specific cache section; only generic sticky-routing behavior applies.
|
|
304
|
+
11. **OpenAI Chat Completions usage-field naming** (`prompt_tokens_details` vs
|
|
305
|
+
Responses `input_tokens_details`) was not confirmed first-party.
|
|
306
|
+
|
|
307
|
+
## 8. What did not change
|
|
308
|
+
|
|
309
|
+
- No runtime file, test file, configuration, provider config, detection regex,
|
|
310
|
+
policy logic, telemetry field, or GPT context/output limit was modified.
|
|
311
|
+
- The existing test suite was run only to confirm the repository remains green
|
|
312
|
+
(116/116).
|
|
313
|
+
- This release contains research and documentation only.
|
package/package.json
CHANGED
|
@@ -20,6 +20,7 @@ import { appendFileSync, existsSync, mkdirSync, readFileSync } from "node:fs"
|
|
|
20
20
|
import { createHash } from "node:crypto"
|
|
21
21
|
import { homedir } from "node:os"
|
|
22
22
|
import { dirname, join } from "node:path"
|
|
23
|
+
import { resolveLegacyFamily } from "./cache-policy-core.mjs"
|
|
23
24
|
|
|
24
25
|
export const CONFIG_FILENAME = "cache-engine.json"
|
|
25
26
|
export const DEFAULT_CONFIG_PATH = join(homedir(), ".config/opencode", CONFIG_FILENAME)
|
|
@@ -205,61 +206,26 @@ export function canonicalStringify(value) {
|
|
|
205
206
|
|
|
206
207
|
// ---------------------------------------------------------------------------
|
|
207
208
|
// Model detection -> cache-policy family
|
|
209
|
+
//
|
|
210
|
+
// Classification now lives in the structured registry/resolver in
|
|
211
|
+
// cache-policy-core.mjs (traceable to docs/cache-policy-inventory.md).
|
|
212
|
+
// detectPolicy() remains a compatibility wrapper over resolveLegacyFamily() so
|
|
213
|
+
// runtime behavior is unchanged until the resolver is wired in a later release.
|
|
208
214
|
// ---------------------------------------------------------------------------
|
|
209
215
|
|
|
210
|
-
|
|
211
|
-
|
|
212
|
-
|
|
213
|
-
|
|
214
|
-
|
|
215
|
-
|
|
216
|
-
const providerID = String(m.providerID ?? m.provider ?? "")
|
|
217
|
-
const apiID = String(m.modelID ?? api.id ?? m.id ?? m.apiID ?? "")
|
|
218
|
-
const modelID = String(m.id ?? "")
|
|
219
|
-
const npm = String(api.npm ?? m.npm ?? "")
|
|
220
|
-
const name = String(m.name ?? "")
|
|
221
|
-
const slug = `${apiID} ${modelID}`.trim()
|
|
222
|
-
return { providerID, apiID, modelID, npm, name, slug: slug.toLowerCase() }
|
|
223
|
-
}
|
|
224
|
-
|
|
225
|
-
// OpenAI-ish context is required before we apply GPT-5.6 options, so we never
|
|
226
|
-
// send GPT-5.6-only fields to a non-OpenAI endpoint merely because a model
|
|
227
|
-
// string contains "gpt-5.6". A slug that explicitly starts with openai/ or
|
|
228
|
-
// azure/ (typical for openrouter/azure/openai-compatible routes) also counts
|
|
229
|
-
// because the upstream IS OpenAI. A bare openai-compatible provider with no
|
|
230
|
-
// such slug does NOT count: we must not guess.
|
|
231
|
-
function isOpenAIish(s) {
|
|
232
|
-
const { providerID, slug, npm } = s
|
|
233
|
-
const p = providerID.toLowerCase()
|
|
234
|
-
if (p === "openai" || p === "azure") return true
|
|
235
|
-
if (slug.startsWith("openai/") || slug.startsWith("azure/")) return true
|
|
236
|
-
if (/@ai-sdk\/openai|@ai-sdk\/azure/.test(npm)) return true
|
|
237
|
-
return false
|
|
238
|
-
}
|
|
239
|
-
|
|
240
|
-
// The gpt-5.6 family, tolerating OpenCode's variants: gpt-5.6, gpt-5.6-luna,
|
|
241
|
-
// gpt-5.6-luna:flex, gpt-5.6-luna-pro:flex, gpt-5.6-<anything>.
|
|
242
|
-
// A trailing digit guard avoids matching hypothetical "gpt-5.60" etc.
|
|
243
|
-
const GPT56_RE = /gpt-5\.6(?![\d.])/i
|
|
244
|
-
// GLM 5.3 family only (not glm-4.x / glm-4.6 etc).
|
|
245
|
-
const GLM53_RE = /glm-5\.3(?![\d.])/i
|
|
246
|
-
// Xiaomi MiMo V2.6 explicitly targets Flash + Pro only. The trailing
|
|
247
|
-
// (?![\w-]) guard prevents matching a hypothetical "mimo-v2.6-pro-ultraspeed"
|
|
248
|
-
// or "mimo-v2.6-flashx", and the v2\.6 literal excludes V2.5 / V2.
|
|
249
|
-
const MIMO26_RE = /mimo-v2\.6-(flash|pro)(?![\w-])/i
|
|
250
|
-
const DEEPSEEK_RE = /deepseek/i
|
|
216
|
+
const LEGACY_FAMILY_TO_POLICY = {
|
|
217
|
+
"gpt-5.6": POLICY_GPT56,
|
|
218
|
+
"glm-5.3": POLICY_GLM53,
|
|
219
|
+
"mimo-v2.6": POLICY_MIMO26,
|
|
220
|
+
deepseek: POLICY_DEEPSEEK,
|
|
221
|
+
}
|
|
251
222
|
|
|
252
223
|
// Pure classifier. Returns one of the POLICY_* keys. `model` may be a full
|
|
253
|
-
// OpenCode Model, or {providerID, modelID|apiID|id}.
|
|
224
|
+
// OpenCode Model, or {providerID, modelID|apiID|id}. Newer generations (e.g.
|
|
225
|
+
// gpt-6) and intra-family exact overlays intentionally resolve to neutral here,
|
|
226
|
+
// exactly as before v0.4.0; use resolvePolicy() for the richer structured result.
|
|
254
227
|
export function detectPolicy(model) {
|
|
255
|
-
|
|
256
|
-
const s = modelSignals(model)
|
|
257
|
-
if (!s.slug) return POLICY_NEUTRAL
|
|
258
|
-
if (GPT56_RE.test(s.slug) && isOpenAIish(s)) return POLICY_GPT56
|
|
259
|
-
if (GLM53_RE.test(s.slug)) return POLICY_GLM53
|
|
260
|
-
if (MIMO26_RE.test(s.slug)) return POLICY_MIMO26
|
|
261
|
-
if (DEEPSEEK_RE.test(s.slug) || DEEPSEEK_RE.test(s.providerID)) return POLICY_DEEPSEEK
|
|
262
|
-
return POLICY_NEUTRAL
|
|
228
|
+
return LEGACY_FAMILY_TO_POLICY[resolveLegacyFamily(model)] ?? POLICY_NEUTRAL
|
|
263
229
|
}
|
|
264
230
|
|
|
265
231
|
export function policyEnabled(cfg, family) {
|
|
@@ -0,0 +1,440 @@
|
|
|
1
|
+
// cache-policy-core.mjs
|
|
2
|
+
//
|
|
3
|
+
// Pure policy registry + resolver for CacheEngine (v0.4.0).
|
|
4
|
+
//
|
|
5
|
+
// This module is the structured policy-resolution layer described by
|
|
6
|
+
// docs/cache-policy-inventory.md. It separates four concerns that used to be
|
|
7
|
+
// entangled in a single model-name-to-behavior branch:
|
|
8
|
+
//
|
|
9
|
+
// 1. creator / family classification
|
|
10
|
+
// 2. baseline cache policy (documented facts)
|
|
11
|
+
// 3. model-specific overlays (CacheEngine code behaviors, NOT implied by family)
|
|
12
|
+
// 4. transport capabilities (e.g. OpenRouter affinity), kept separate from
|
|
13
|
+
// creator cache semantics
|
|
14
|
+
//
|
|
15
|
+
// It performs NO network calls and NO runtime documentation lookups. Every
|
|
16
|
+
// registry entry is traceable to docs/cache-policy-inventory.md via
|
|
17
|
+
// `inventoryRef`. Inheritance is always explicit (`inheritsFrom`); "newer means
|
|
18
|
+
// same behavior" is never an unconditional rule.
|
|
19
|
+
//
|
|
20
|
+
// This release does not wire the resolver into runtime hooks. detectPolicy()
|
|
21
|
+
// remains the compatibility classifier until wiring is approved.
|
|
22
|
+
|
|
23
|
+
// ---------------------------------------------------------------------------
|
|
24
|
+
// Model normalization (shared with the legacy classifier)
|
|
25
|
+
// ---------------------------------------------------------------------------
|
|
26
|
+
|
|
27
|
+
// Normalize a model-like object into a searchable haystack. Accepts both the
|
|
28
|
+
// full OpenCode Model ({providerID, id, api:{id,npm}, name}) and slim test
|
|
29
|
+
// objects ({providerID, modelID/apiID}).
|
|
30
|
+
export function modelSignals(model) {
|
|
31
|
+
const m = model && typeof model === "object" ? model : {}
|
|
32
|
+
const api = m.api && typeof m.api === "object" ? m.api : {}
|
|
33
|
+
const providerID = String(m.providerID ?? m.provider ?? "")
|
|
34
|
+
const apiID = String(m.modelID ?? api.id ?? m.id ?? m.apiID ?? "")
|
|
35
|
+
const modelID = String(m.id ?? "")
|
|
36
|
+
const npm = String(api.npm ?? m.npm ?? "")
|
|
37
|
+
const name = String(m.name ?? "")
|
|
38
|
+
const slug = `${apiID} ${modelID}`.trim()
|
|
39
|
+
return { providerID, apiID, modelID, npm, name, slug: slug.toLowerCase() }
|
|
40
|
+
}
|
|
41
|
+
|
|
42
|
+
// OpenAI-ish context is required before we apply GPT-5.6 options, so we never
|
|
43
|
+
// send GPT-5.6-only fields to a non-OpenAI endpoint merely because a model
|
|
44
|
+
// string contains "gpt-5.6". A slug that explicitly starts with openai/ or
|
|
45
|
+
// azure/ (typical for openrouter/azure/openai-compatible routes) also counts
|
|
46
|
+
// because the upstream IS OpenAI. A bare openai-compatible provider with no
|
|
47
|
+
// such slug does NOT count: we must not guess.
|
|
48
|
+
export function isOpenAIish(s) {
|
|
49
|
+
const { providerID, slug, npm } = s
|
|
50
|
+
const p = providerID.toLowerCase()
|
|
51
|
+
if (p === "openai" || p === "azure") return true
|
|
52
|
+
if (slug.startsWith("openai/") || slug.startsWith("azure/")) return true
|
|
53
|
+
if (/@ai-sdk\/openai|@ai-sdk\/azure/.test(npm)) return true
|
|
54
|
+
return false
|
|
55
|
+
}
|
|
56
|
+
|
|
57
|
+
// Candidate ids for exact/alias lookup. Includes the raw apiID/modelID, the
|
|
58
|
+
// lower-cased forms, and a single stripped transport/vendor prefix
|
|
59
|
+
// (e.g. "openai/gpt-5.6-luna" -> "gpt-5.6-luna", "xiaomi/mimo-v2.6-flash" ->
|
|
60
|
+
// "mimo-v2.6-flash"). Prefix stripping is a lookup convenience only and never
|
|
61
|
+
// implies cache semantics.
|
|
62
|
+
function candidateIds(s) {
|
|
63
|
+
const raw = [s.apiID, s.modelID].filter(Boolean).map((v) => String(v).toLowerCase())
|
|
64
|
+
const out = new Set()
|
|
65
|
+
for (const v of raw) {
|
|
66
|
+
out.add(v)
|
|
67
|
+
const stripped = v.replace(/^[a-z0-9._-]+\//, "")
|
|
68
|
+
if (stripped) out.add(stripped)
|
|
69
|
+
}
|
|
70
|
+
return [...out]
|
|
71
|
+
}
|
|
72
|
+
|
|
73
|
+
// ---------------------------------------------------------------------------
|
|
74
|
+
// Baseline policies (documented cache-policy facts, not code behavior)
|
|
75
|
+
// ---------------------------------------------------------------------------
|
|
76
|
+
|
|
77
|
+
export const BASELINES = {
|
|
78
|
+
"openai.gpt56.cache": {
|
|
79
|
+
id: "openai.gpt56.cache",
|
|
80
|
+
creator: "openai",
|
|
81
|
+
appliesTo: "GPT-5.6 and later (OpenAI-documented generation boundary)",
|
|
82
|
+
automatic: true,
|
|
83
|
+
defaultMode: "implicit",
|
|
84
|
+
supportsExplicitBreakpoints: true,
|
|
85
|
+
minCacheTokens: 1024,
|
|
86
|
+
ttl: "30m",
|
|
87
|
+
cacheKeyOptional: true,
|
|
88
|
+
cacheWriteBilled: true,
|
|
89
|
+
usageFields: [
|
|
90
|
+
"input_tokens_details.cached_tokens",
|
|
91
|
+
"input_tokens_details.cache_write_tokens",
|
|
92
|
+
],
|
|
93
|
+
inventoryRef: "§1 OpenAI",
|
|
94
|
+
},
|
|
95
|
+
"deepseek.kv-cache": {
|
|
96
|
+
id: "deepseek.kv-cache",
|
|
97
|
+
creator: "deepseek",
|
|
98
|
+
appliesTo: "DeepSeek provider-wide (documented default for all users)",
|
|
99
|
+
automatic: true,
|
|
100
|
+
defaultMode: "implicit",
|
|
101
|
+
supportsExplicitBreakpoints: false,
|
|
102
|
+
minCacheTokens: null,
|
|
103
|
+
ttl: null,
|
|
104
|
+
cacheKeyOptional: false,
|
|
105
|
+
cacheWriteBilled: false,
|
|
106
|
+
usageFields: ["prompt_cache_hit_tokens", "prompt_cache_miss_tokens"],
|
|
107
|
+
inventoryRef: "§2 DeepSeek",
|
|
108
|
+
},
|
|
109
|
+
"zai.implicit-cache": {
|
|
110
|
+
id: "zai.implicit-cache",
|
|
111
|
+
creator: "z.ai",
|
|
112
|
+
appliesTo: "Z.AI service-wide implicit context caching",
|
|
113
|
+
automatic: true,
|
|
114
|
+
defaultMode: "implicit",
|
|
115
|
+
supportsExplicitBreakpoints: false,
|
|
116
|
+
minCacheTokens: null,
|
|
117
|
+
ttl: null,
|
|
118
|
+
cacheKeyOptional: false,
|
|
119
|
+
cacheWriteBilled: false,
|
|
120
|
+
usageFields: ["prompt_tokens_details.cached_tokens"],
|
|
121
|
+
inventoryRef: "§3 Z.AI GLM",
|
|
122
|
+
},
|
|
123
|
+
"xiaomi.implicit-cache": {
|
|
124
|
+
id: "xiaomi.implicit-cache",
|
|
125
|
+
creator: "xiaomi",
|
|
126
|
+
appliesTo: "Xiaomi MiMo provider-managed implicit caching",
|
|
127
|
+
automatic: true,
|
|
128
|
+
defaultMode: "implicit",
|
|
129
|
+
supportsExplicitBreakpoints: false,
|
|
130
|
+
minCacheTokens: null,
|
|
131
|
+
ttl: null,
|
|
132
|
+
cacheKeyOptional: false,
|
|
133
|
+
cacheWriteBilled: false,
|
|
134
|
+
usageFields: ["prompt_tokens_details.cached_tokens", "cache_read_input_tokens"],
|
|
135
|
+
inventoryRef: "§4 Xiaomi MiMo",
|
|
136
|
+
},
|
|
137
|
+
"neutral.none": {
|
|
138
|
+
id: "neutral.none",
|
|
139
|
+
creator: "unknown",
|
|
140
|
+
appliesTo: "No registered cache policy; no model-specific optimization implied",
|
|
141
|
+
automatic: false,
|
|
142
|
+
defaultMode: null,
|
|
143
|
+
supportsExplicitBreakpoints: false,
|
|
144
|
+
minCacheTokens: null,
|
|
145
|
+
ttl: null,
|
|
146
|
+
cacheKeyOptional: false,
|
|
147
|
+
cacheWriteBilled: false,
|
|
148
|
+
usageFields: [],
|
|
149
|
+
inventoryRef: "§6 Compatibility Matrix",
|
|
150
|
+
},
|
|
151
|
+
}
|
|
152
|
+
|
|
153
|
+
// ---------------------------------------------------------------------------
|
|
154
|
+
// Model-specific overlays (CacheEngine code behaviors)
|
|
155
|
+
//
|
|
156
|
+
// An overlay is a concrete CacheEngine mutation. It is attached to a family
|
|
157
|
+
// ONLY by explicit registration; being classified into a creator/family never
|
|
158
|
+
// implies an overlay. This is what keeps "is GLM" from automatically meaning
|
|
159
|
+
// "<env> relocation".
|
|
160
|
+
// ---------------------------------------------------------------------------
|
|
161
|
+
|
|
162
|
+
export const OVERLAYS = {
|
|
163
|
+
"gpt56.prompt-cache-options": {
|
|
164
|
+
id: "gpt56.prompt-cache-options",
|
|
165
|
+
family: "gpt-5.6",
|
|
166
|
+
hook: "chat.params",
|
|
167
|
+
behavior: "inject missing promptCacheKey + promptCacheOptions(implicit, 30m)",
|
|
168
|
+
inventoryRef: "§1 OpenAI",
|
|
169
|
+
},
|
|
170
|
+
"glm53.env-relocation": {
|
|
171
|
+
id: "glm53.env-relocation",
|
|
172
|
+
family: "glm-5.3",
|
|
173
|
+
hook: "experimental.chat.system.transform",
|
|
174
|
+
behavior: "relocate the identifiable <env> block to the system tail",
|
|
175
|
+
inventoryRef: "§3 Z.AI GLM",
|
|
176
|
+
},
|
|
177
|
+
"mimo26.env-relocation": {
|
|
178
|
+
id: "mimo26.env-relocation",
|
|
179
|
+
family: "mimo-v2.6",
|
|
180
|
+
hook: "experimental.chat.system.transform",
|
|
181
|
+
behavior: "relocate the identifiable <env> block to the system tail",
|
|
182
|
+
inventoryRef: "§4 Xiaomi MiMo",
|
|
183
|
+
},
|
|
184
|
+
}
|
|
185
|
+
|
|
186
|
+
// ---------------------------------------------------------------------------
|
|
187
|
+
// Transport capabilities (routing), deliberately independent of cache policy
|
|
188
|
+
// ---------------------------------------------------------------------------
|
|
189
|
+
|
|
190
|
+
export const TRANSPORTS = {
|
|
191
|
+
openrouter: {
|
|
192
|
+
id: "openrouter",
|
|
193
|
+
kind: "openrouter",
|
|
194
|
+
sessionAffinityHeader: "x-session-id",
|
|
195
|
+
stickyRouting: true,
|
|
196
|
+
inventoryRef: "§5 OpenRouter transport",
|
|
197
|
+
},
|
|
198
|
+
}
|
|
199
|
+
|
|
200
|
+
function resolveTransport(s) {
|
|
201
|
+
const p = String(s.providerID ?? "").toLowerCase()
|
|
202
|
+
if (p === "openrouter") return { ...TRANSPORTS.openrouter }
|
|
203
|
+
if (!p) {
|
|
204
|
+
return { id: "unknown", kind: "unknown", sessionAffinityHeader: null, stickyRouting: false, inventoryRef: "§5 OpenRouter transport" }
|
|
205
|
+
}
|
|
206
|
+
return { id: p, kind: "direct", sessionAffinityHeader: null, stickyRouting: false, inventoryRef: "§5 OpenRouter transport" }
|
|
207
|
+
}
|
|
208
|
+
|
|
209
|
+
// ---------------------------------------------------------------------------
|
|
210
|
+
// Registry
|
|
211
|
+
//
|
|
212
|
+
// Entries are evaluated in array order, which encodes detection priority and
|
|
213
|
+
// preserves the legacy classifier's precedence (gpt-5.6 > glm-5.3 > mimo-v2.6 >
|
|
214
|
+
// deepseek). `legacy: true` entries reproduce the pre-v0.4.0 detectPolicy()
|
|
215
|
+
// behavior exactly. `legacy: false` entries (gpt-6, Pro UltraSpeed) are
|
|
216
|
+
// available to the new resolver only, so runtime behavior is unchanged until
|
|
217
|
+
// wiring is approved.
|
|
218
|
+
// ---------------------------------------------------------------------------
|
|
219
|
+
|
|
220
|
+
export const POLICY_REGISTRY = [
|
|
221
|
+
{
|
|
222
|
+
id: "openai.gpt-5.6",
|
|
223
|
+
creator: "openai",
|
|
224
|
+
family: "gpt-5.6",
|
|
225
|
+
kind: "family",
|
|
226
|
+
pattern: /gpt-5\.6(?![\d.])/i,
|
|
227
|
+
requiresOpenAIish: true,
|
|
228
|
+
exactIds: ["gpt-5.6-sol", "gpt-5.6-terra", "gpt-5.6-luna", "gpt-5.6-cyber"],
|
|
229
|
+
baseline: "openai.gpt56.cache",
|
|
230
|
+
overlays: ["gpt56.prompt-cache-options"],
|
|
231
|
+
legacy: true,
|
|
232
|
+
inventoryRef: "§1 OpenAI",
|
|
233
|
+
},
|
|
234
|
+
{
|
|
235
|
+
id: "openai.gpt-6",
|
|
236
|
+
creator: "openai",
|
|
237
|
+
family: "gpt-6",
|
|
238
|
+
kind: "family",
|
|
239
|
+
pattern: /gpt-6(?![\d.])/i,
|
|
240
|
+
requiresOpenAIish: true,
|
|
241
|
+
exactIds: ["gpt-6-astra", "gpt-6-sol", "gpt-6-luna"],
|
|
242
|
+
baseline: "openai.gpt56.cache",
|
|
243
|
+
inheritsFrom: "gpt-5.6",
|
|
244
|
+
overlays: [],
|
|
245
|
+
legacy: false,
|
|
246
|
+
note: "Documented inheritance of the GPT-5.6-and-later baseline. No CacheEngine overlay is registered for gpt-6 yet.",
|
|
247
|
+
inventoryRef: "§1 OpenAI",
|
|
248
|
+
},
|
|
249
|
+
{
|
|
250
|
+
id: "zai.glm-5.3",
|
|
251
|
+
creator: "z.ai",
|
|
252
|
+
family: "glm-5.3",
|
|
253
|
+
kind: "family",
|
|
254
|
+
pattern: /glm-5\.3(?![\d.])/i,
|
|
255
|
+
exactIds: ["glm-5.3", "glm-5.3-flash", "glm-5.3-flashx"],
|
|
256
|
+
baseline: "zai.implicit-cache",
|
|
257
|
+
overlays: ["glm53.env-relocation"],
|
|
258
|
+
legacy: true,
|
|
259
|
+
inventoryRef: "§3 Z.AI GLM",
|
|
260
|
+
},
|
|
261
|
+
{
|
|
262
|
+
id: "xiaomi.mimo-v2.6",
|
|
263
|
+
creator: "xiaomi",
|
|
264
|
+
family: "mimo-v2.6",
|
|
265
|
+
kind: "family",
|
|
266
|
+
pattern: /mimo-v2\.6-(flash|pro)(?![\w-])/i,
|
|
267
|
+
exactIds: ["mimo-v2.6-flash", "mimo-v2.6-pro"],
|
|
268
|
+
baseline: "xiaomi.implicit-cache",
|
|
269
|
+
overlays: ["mimo26.env-relocation"],
|
|
270
|
+
legacy: true,
|
|
271
|
+
inventoryRef: "§4 Xiaomi MiMo",
|
|
272
|
+
},
|
|
273
|
+
{
|
|
274
|
+
id: "xiaomi.mimo-v2.6-pro-ultraspeed",
|
|
275
|
+
creator: "xiaomi",
|
|
276
|
+
family: "mimo-v2.6",
|
|
277
|
+
kind: "exact",
|
|
278
|
+
exactIds: ["mimo-v2.6-pro-ultraspeed"],
|
|
279
|
+
baseline: "xiaomi.implicit-cache",
|
|
280
|
+
overlays: [],
|
|
281
|
+
legacy: false,
|
|
282
|
+
policyStatus: "documented-series-member-without-registered-overlay",
|
|
283
|
+
note: "Documented as a Pro mode in the same V2.6 series, but the inventory does not establish identical cache controls and CacheEngine registers no overlay for it.",
|
|
284
|
+
inventoryRef: "§4 Xiaomi MiMo",
|
|
285
|
+
},
|
|
286
|
+
{
|
|
287
|
+
id: "deepseek.baseline",
|
|
288
|
+
creator: "deepseek",
|
|
289
|
+
family: "deepseek",
|
|
290
|
+
kind: "creator",
|
|
291
|
+
pattern: /deepseek/i,
|
|
292
|
+
providerPattern: /deepseek/i,
|
|
293
|
+
baseline: "deepseek.kv-cache",
|
|
294
|
+
overlays: [],
|
|
295
|
+
legacy: true,
|
|
296
|
+
inventoryRef: "§2 DeepSeek",
|
|
297
|
+
},
|
|
298
|
+
]
|
|
299
|
+
|
|
300
|
+
// ---------------------------------------------------------------------------
|
|
301
|
+
// Explicit aliases identified by the inventory
|
|
302
|
+
// ---------------------------------------------------------------------------
|
|
303
|
+
|
|
304
|
+
export const MODEL_ALIASES = {
|
|
305
|
+
"gpt-5.6": { canonicalId: "gpt-5.6-sol", family: "gpt-5.6", creator: "openai", inventoryRef: "§1 OpenAI" },
|
|
306
|
+
"gpt-daybreak-blue-latest": { canonicalId: "gpt-5.6-sol", family: "gpt-5.6", creator: "openai", inventoryRef: "§1 OpenAI" },
|
|
307
|
+
"gpt-daybreak-red-latest": { canonicalId: "gpt-5.6-cyber", family: "gpt-5.6", creator: "openai", inventoryRef: "§1 OpenAI" },
|
|
308
|
+
"deepseek-v4-flash": { canonicalId: "deepseek-flash", family: "deepseek", creator: "deepseek", status: "retired-legacy-id", inventoryRef: "§2 DeepSeek" },
|
|
309
|
+
"deepseek-chat": { canonicalId: null, family: "deepseek", creator: "deepseek", status: "retired", inventoryRef: "§2 DeepSeek" },
|
|
310
|
+
"deepseek-reasoner": { canonicalId: null, family: "deepseek", creator: "deepseek", status: "retired", inventoryRef: "§2 DeepSeek" },
|
|
311
|
+
}
|
|
312
|
+
|
|
313
|
+
// ---------------------------------------------------------------------------
|
|
314
|
+
// Resolution
|
|
315
|
+
// ---------------------------------------------------------------------------
|
|
316
|
+
|
|
317
|
+
function neutralResult(reason, transport) {
|
|
318
|
+
return {
|
|
319
|
+
creator: "unknown",
|
|
320
|
+
family: "neutral",
|
|
321
|
+
baseline: BASELINES["neutral.none"],
|
|
322
|
+
overlays: [],
|
|
323
|
+
transport,
|
|
324
|
+
matchType: "neutral",
|
|
325
|
+
matchReason: reason,
|
|
326
|
+
matchedId: null,
|
|
327
|
+
inventoryRef: null,
|
|
328
|
+
note: null,
|
|
329
|
+
}
|
|
330
|
+
}
|
|
331
|
+
|
|
332
|
+
function overlaysFor(ids) {
|
|
333
|
+
return (ids ?? []).map((id) => OVERLAYS[id]).filter(Boolean)
|
|
334
|
+
}
|
|
335
|
+
|
|
336
|
+
function resultFromEntry(entry, matchType, matchReason, matchedId, transport) {
|
|
337
|
+
return {
|
|
338
|
+
creator: entry.creator,
|
|
339
|
+
family: entry.family,
|
|
340
|
+
baseline: BASELINES[entry.baseline] ?? null,
|
|
341
|
+
overlays: overlaysFor(entry.overlays),
|
|
342
|
+
transport,
|
|
343
|
+
matchType,
|
|
344
|
+
matchReason,
|
|
345
|
+
matchedId: matchedId ?? null,
|
|
346
|
+
inventoryRef: entry.inventoryRef ?? null,
|
|
347
|
+
note: entry.note ?? null,
|
|
348
|
+
}
|
|
349
|
+
}
|
|
350
|
+
|
|
351
|
+
function resultFromFamily(family, creator, matchType, matchReason, matchedId, inventoryRef, note, transport) {
|
|
352
|
+
const entry = POLICY_REGISTRY.find((e) => e.family === family && e.kind !== "exact")
|
|
353
|
+
return {
|
|
354
|
+
creator,
|
|
355
|
+
family,
|
|
356
|
+
baseline: entry ? BASELINES[entry.baseline] ?? null : null,
|
|
357
|
+
overlays: entry ? overlaysFor(entry.overlays) : [],
|
|
358
|
+
transport,
|
|
359
|
+
matchType,
|
|
360
|
+
matchReason,
|
|
361
|
+
matchedId: matchedId ?? null,
|
|
362
|
+
inventoryRef: inventoryRef ?? null,
|
|
363
|
+
note: note ?? null,
|
|
364
|
+
}
|
|
365
|
+
}
|
|
366
|
+
|
|
367
|
+
// Structured resolver. Returns creator, family, baseline, overlays, transport,
|
|
368
|
+
// matchType, and matchReason. Unknown models resolve to a neutral result with
|
|
369
|
+
// an empty overlay list and no model-specific optimization.
|
|
370
|
+
export function resolvePolicy(model) {
|
|
371
|
+
if (!model || typeof model !== "object") {
|
|
372
|
+
return neutralResult("neutral:invalid-model", resolveTransport({}))
|
|
373
|
+
}
|
|
374
|
+
const s = modelSignals(model)
|
|
375
|
+
const transport = resolveTransport(s)
|
|
376
|
+
if (!s.slug) return neutralResult("neutral:no-model-identity", transport)
|
|
377
|
+
|
|
378
|
+
const ids = candidateIds(s)
|
|
379
|
+
|
|
380
|
+
// 1. Explicit aliases (inventory-identified). An alias inherits its target
|
|
381
|
+
// family's context gate, so e.g. an OpenAI alias on a non-OpenAI gateway does
|
|
382
|
+
// not gain a cache policy it would not otherwise have.
|
|
383
|
+
for (const id of ids) {
|
|
384
|
+
const alias = MODEL_ALIASES[id]
|
|
385
|
+
if (!alias) continue
|
|
386
|
+
const familyEntry = POLICY_REGISTRY.find((e) => e.family === alias.family && e.kind !== "exact")
|
|
387
|
+
if (familyEntry?.requiresOpenAIish && !isOpenAIish(s)) continue
|
|
388
|
+
const reason = `alias:${id}->${alias.canonicalId ?? alias.family}`
|
|
389
|
+
return resultFromFamily(alias.family, alias.creator, "exact", reason, id, alias.inventoryRef, alias.status ?? null, transport)
|
|
390
|
+
}
|
|
391
|
+
|
|
392
|
+
// 2. Exact model ids (documented models).
|
|
393
|
+
for (const entry of POLICY_REGISTRY) {
|
|
394
|
+
if (!entry.exactIds || entry.exactIds.length === 0) continue
|
|
395
|
+
const hit = ids.find((id) => entry.exactIds.includes(id))
|
|
396
|
+
if (hit) return resultFromEntry(entry, "exact", `exact-id:${hit}`, hit, transport)
|
|
397
|
+
}
|
|
398
|
+
|
|
399
|
+
// 3. Model family / range patterns.
|
|
400
|
+
for (const entry of POLICY_REGISTRY) {
|
|
401
|
+
if (entry.kind !== "family" || !entry.pattern) continue
|
|
402
|
+
if (!entry.pattern.test(s.slug)) continue
|
|
403
|
+
if (entry.requiresOpenAIish && !isOpenAIish(s)) continue
|
|
404
|
+
return resultFromEntry(entry, "family", `family-pattern:${entry.id}`, null, transport)
|
|
405
|
+
}
|
|
406
|
+
|
|
407
|
+
// 4. Creator baseline.
|
|
408
|
+
for (const entry of POLICY_REGISTRY) {
|
|
409
|
+
if (entry.kind !== "creator") continue
|
|
410
|
+
const slugHit = entry.pattern ? entry.pattern.test(s.slug) : false
|
|
411
|
+
const providerHit = entry.providerPattern ? entry.providerPattern.test(s.providerID) : false
|
|
412
|
+
if (slugHit || providerHit) return resultFromEntry(entry, "creator", `creator-baseline:${entry.id}`, null, transport)
|
|
413
|
+
}
|
|
414
|
+
|
|
415
|
+
return neutralResult("neutral:no-match", transport)
|
|
416
|
+
}
|
|
417
|
+
|
|
418
|
+
// Compatibility classification used by detectPolicy(). Reproduces the
|
|
419
|
+
// pre-v0.4.0 behavior exactly: it considers only `legacy` registry entries,
|
|
420
|
+
// excludes newer-generation/alias/exact-overlay additions, and returns a family
|
|
421
|
+
// string (or "neutral"). Callers map the family to a POLICY_* constant.
|
|
422
|
+
export function resolveLegacyFamily(model) {
|
|
423
|
+
if (!model || typeof model !== "object") return "neutral"
|
|
424
|
+
const s = modelSignals(model)
|
|
425
|
+
if (!s.slug) return "neutral"
|
|
426
|
+
for (const entry of POLICY_REGISTRY) {
|
|
427
|
+
if (!entry.legacy) continue
|
|
428
|
+
if (entry.kind === "family") {
|
|
429
|
+
if (!entry.pattern.test(s.slug)) continue
|
|
430
|
+
if (entry.requiresOpenAIish && !isOpenAIish(s)) continue
|
|
431
|
+
return entry.family
|
|
432
|
+
}
|
|
433
|
+
if (entry.kind === "creator") {
|
|
434
|
+
const slugHit = entry.pattern ? entry.pattern.test(s.slug) : false
|
|
435
|
+
const providerHit = entry.providerPattern ? entry.providerPattern.test(s.providerID) : false
|
|
436
|
+
if (slugHit || providerHit) return entry.family
|
|
437
|
+
}
|
|
438
|
+
}
|
|
439
|
+
return "neutral"
|
|
440
|
+
}
|
|
@@ -48,6 +48,14 @@ import {
|
|
|
48
48
|
toolFingerprint,
|
|
49
49
|
toolWireFingerprint,
|
|
50
50
|
} from "../src/cache-engine-core.mjs"
|
|
51
|
+
import {
|
|
52
|
+
BASELINES,
|
|
53
|
+
MODEL_ALIASES,
|
|
54
|
+
OVERLAYS,
|
|
55
|
+
POLICY_REGISTRY,
|
|
56
|
+
resolveLegacyFamily,
|
|
57
|
+
resolvePolicy,
|
|
58
|
+
} from "../src/cache-policy-core.mjs"
|
|
51
59
|
|
|
52
60
|
const asst = (id, read, write) => ({
|
|
53
61
|
info: { id, role: "assistant", tokens: { cache: { read, write } } },
|
|
@@ -1316,3 +1324,230 @@ test("affinity telemetry preserves non-OpenRouter families and reports provider
|
|
|
1316
1324
|
assert.equal(glmChange.to.providerID, "zai")
|
|
1317
1325
|
assert.equal(glmChange.policy, POLICY_GLM53)
|
|
1318
1326
|
})
|
|
1327
|
+
|
|
1328
|
+
// ===========================================================================
|
|
1329
|
+
// v0.4.0 policy registry + resolvePolicy() (research-backed, pure, unwired)
|
|
1330
|
+
//
|
|
1331
|
+
// The inventory (docs/cache-policy-inventory.md) is the authority. These tests
|
|
1332
|
+
// assert the structured resolution layer and that detectPolicy() remains
|
|
1333
|
+
// byte-compatible with its pre-v0.4.0 behavior.
|
|
1334
|
+
// ===========================================================================
|
|
1335
|
+
|
|
1336
|
+
const overlayIds = (r) => r.overlays.map((o) => o.id)
|
|
1337
|
+
const baseId = (r) => (r.baseline ? r.baseline.id : null)
|
|
1338
|
+
|
|
1339
|
+
test("resolvePolicy: GPT-5.6 exact + inventory aliases resolve to the gpt-5.6 family", () => {
|
|
1340
|
+
const exact = resolvePolicy(M("openai", "gpt-5.6-sol"))
|
|
1341
|
+
assert.equal(exact.creator, "openai")
|
|
1342
|
+
assert.equal(exact.family, "gpt-5.6")
|
|
1343
|
+
assert.equal(exact.matchType, "exact")
|
|
1344
|
+
assert.equal(baseId(exact), "openai.gpt56.cache")
|
|
1345
|
+
assert.deepEqual(overlayIds(exact), ["gpt56.prompt-cache-options"])
|
|
1346
|
+
assert.equal(exact.transport.kind, "direct")
|
|
1347
|
+
|
|
1348
|
+
// Inventory alias: gpt-5.6 -> gpt-5.6-sol
|
|
1349
|
+
const alias = resolvePolicy(M("openai", "gpt-5.6"))
|
|
1350
|
+
assert.equal(alias.family, "gpt-5.6")
|
|
1351
|
+
assert.equal(alias.matchType, "exact")
|
|
1352
|
+
assert.ok(alias.matchReason.startsWith("alias:gpt-5.6"))
|
|
1353
|
+
assert.deepEqual(overlayIds(alias), ["gpt56.prompt-cache-options"])
|
|
1354
|
+
|
|
1355
|
+
// OpenRouter-prefixed documented variant resolves by exact id (prefix stripped)
|
|
1356
|
+
const orVariant = resolvePolicy(M("openrouter", "openai/gpt-5.6-luna"))
|
|
1357
|
+
assert.equal(orVariant.family, "gpt-5.6")
|
|
1358
|
+
assert.equal(orVariant.matchType, "exact")
|
|
1359
|
+
})
|
|
1360
|
+
|
|
1361
|
+
test("resolvePolicy: GPT-6 resolves via explicit documented inheritance, with no overlay", () => {
|
|
1362
|
+
const r = resolvePolicy(M("openrouter", "openai/gpt-6-luna"))
|
|
1363
|
+
assert.equal(r.creator, "openai")
|
|
1364
|
+
assert.equal(r.family, "gpt-6")
|
|
1365
|
+
assert.equal(baseId(r), "openai.gpt56.cache")
|
|
1366
|
+
assert.deepEqual(overlayIds(r), [])
|
|
1367
|
+
assert.equal(resolvePolicy(M("openai", "gpt-6-astra")).family, "gpt-6")
|
|
1368
|
+
|
|
1369
|
+
// Inheritance is explicit, not "newer means same": the registry records it.
|
|
1370
|
+
const gpt6 = POLICY_REGISTRY.find((e) => e.family === "gpt-6")
|
|
1371
|
+
assert.equal(gpt6.inheritsFrom, "gpt-5.6")
|
|
1372
|
+
assert.ok(gpt6.inventoryRef)
|
|
1373
|
+
// And no overlay is implied by that inheritance.
|
|
1374
|
+
assert.deepEqual(gpt6.overlays, [])
|
|
1375
|
+
})
|
|
1376
|
+
|
|
1377
|
+
test("resolvePolicy: GPT-6 is not classified by the legacy detectPolicy wrapper", () => {
|
|
1378
|
+
// Intentional: the new layer knows gpt-6, the compatibility wrapper does not.
|
|
1379
|
+
assert.equal(detectPolicy(M("openai", "gpt-6-astra")), POLICY_NEUTRAL)
|
|
1380
|
+
assert.equal(resolvePolicy(M("openai", "gpt-6-astra")).family, "gpt-6")
|
|
1381
|
+
})
|
|
1382
|
+
|
|
1383
|
+
test("resolvePolicy: pre-5.6 GPT negative controls are neutral with no overlays", () => {
|
|
1384
|
+
for (const id of ["gpt-5.5", "gpt-5.4", "gpt-5.2", "gpt-5.1", "gpt-5", "gpt-4.1", "gpt-4o"]) {
|
|
1385
|
+
const r = resolvePolicy(M("openai", id))
|
|
1386
|
+
assert.equal(r.family, "neutral", `${id} must be neutral`)
|
|
1387
|
+
assert.deepEqual(overlayIds(r), [])
|
|
1388
|
+
assert.equal(r.matchType, "neutral")
|
|
1389
|
+
}
|
|
1390
|
+
})
|
|
1391
|
+
|
|
1392
|
+
test("resolvePolicy: DeepSeek V4 / V4.1 resolve to the creator baseline (passive, no overlay)", () => {
|
|
1393
|
+
const v4 = resolvePolicy(M("deepseek", "deepseek-v4-pro"))
|
|
1394
|
+
assert.equal(v4.creator, "deepseek")
|
|
1395
|
+
assert.equal(v4.family, "deepseek")
|
|
1396
|
+
assert.equal(v4.matchType, "creator")
|
|
1397
|
+
assert.equal(baseId(v4), "deepseek.kv-cache")
|
|
1398
|
+
assert.deepEqual(overlayIds(v4), [])
|
|
1399
|
+
|
|
1400
|
+
assert.equal(resolvePolicy(M("deepseek", "deepseek-flash")).family, "deepseek")
|
|
1401
|
+
|
|
1402
|
+
// Inventory alias: retired deepseek-v4-flash -> deepseek-flash
|
|
1403
|
+
const legacy = resolvePolicy(M("deepseek", "deepseek-v4-flash"))
|
|
1404
|
+
assert.equal(legacy.family, "deepseek")
|
|
1405
|
+
assert.ok(legacy.matchReason.startsWith("alias:deepseek-v4-flash"))
|
|
1406
|
+
|
|
1407
|
+
// pre-V4 negative control still resolves to the creator baseline
|
|
1408
|
+
assert.equal(resolvePolicy(M("deepseek", "deepseek-chat")).family, "deepseek")
|
|
1409
|
+
})
|
|
1410
|
+
|
|
1411
|
+
test("resolvePolicy: GLM-5.3 + documented 5.3 variants carry the overlay; 5.2 and earlier do not", () => {
|
|
1412
|
+
const r = resolvePolicy(M("zai", "glm-5.3"))
|
|
1413
|
+
assert.equal(r.creator, "z.ai")
|
|
1414
|
+
assert.equal(r.family, "glm-5.3")
|
|
1415
|
+
assert.equal(baseId(r), "zai.implicit-cache")
|
|
1416
|
+
assert.deepEqual(overlayIds(r), ["glm53.env-relocation"])
|
|
1417
|
+
|
|
1418
|
+
assert.equal(resolvePolicy(M("zai", "glm-5.3-flash")).family, "glm-5.3")
|
|
1419
|
+
assert.equal(resolvePolicy(M("z-ai", "glm-5.3-flashx")).family, "glm-5.3")
|
|
1420
|
+
|
|
1421
|
+
for (const id of ["glm-5.2", "glm-5.1", "glm-5", "glm-4.7", "glm-4.6", "glm-4.5"]) {
|
|
1422
|
+
const g = resolvePolicy(M("zai", id))
|
|
1423
|
+
assert.equal(g.family, "neutral", `${id} must be neutral`)
|
|
1424
|
+
assert.deepEqual(overlayIds(g), [])
|
|
1425
|
+
}
|
|
1426
|
+
})
|
|
1427
|
+
|
|
1428
|
+
test("resolvePolicy: MiMo V2.6 Flash/Pro carry the overlay; Pro UltraSpeed is an explicit no-overlay series member", () => {
|
|
1429
|
+
const flash = resolvePolicy(M("xiaomi", "mimo-v2.6-flash"))
|
|
1430
|
+
assert.equal(flash.creator, "xiaomi")
|
|
1431
|
+
assert.equal(flash.family, "mimo-v2.6")
|
|
1432
|
+
assert.equal(baseId(flash), "xiaomi.implicit-cache")
|
|
1433
|
+
assert.deepEqual(overlayIds(flash), ["mimo26.env-relocation"])
|
|
1434
|
+
|
|
1435
|
+
const pro = resolvePolicy(M("xiaomi", "mimo-v2.6-pro"))
|
|
1436
|
+
assert.equal(pro.family, "mimo-v2.6")
|
|
1437
|
+
assert.deepEqual(overlayIds(pro), ["mimo26.env-relocation"])
|
|
1438
|
+
|
|
1439
|
+
// Documented as the same V2.6 series but with NO registered overlay.
|
|
1440
|
+
const ultraspeed = resolvePolicy(M("xiaomi", "mimo-v2.6-pro-ultraspeed"))
|
|
1441
|
+
assert.equal(ultraspeed.family, "mimo-v2.6")
|
|
1442
|
+
assert.equal(ultraspeed.matchType, "exact")
|
|
1443
|
+
assert.deepEqual(overlayIds(ultraspeed), [])
|
|
1444
|
+
const uEntry = POLICY_REGISTRY.find((e) => e.id === "xiaomi.mimo-v2.6-pro-ultraspeed")
|
|
1445
|
+
assert.equal(uEntry.policyStatus, "documented-series-member-without-registered-overlay")
|
|
1446
|
+
|
|
1447
|
+
const v25 = resolvePolicy(M("xiaomi", "mimo-v2.5"))
|
|
1448
|
+
assert.equal(v25.family, "neutral")
|
|
1449
|
+
assert.deepEqual(overlayIds(v25), [])
|
|
1450
|
+
})
|
|
1451
|
+
|
|
1452
|
+
test("resolvePolicy: overlays are never implied by creator/family classification", () => {
|
|
1453
|
+
// Being GLM/MiMo/DeepSeek does not by itself grant a transformation overlay.
|
|
1454
|
+
assert.deepEqual(overlayIds(resolvePolicy(M("zai", "glm-4.6"))), [])
|
|
1455
|
+
assert.deepEqual(overlayIds(resolvePolicy(M("xiaomi", "mimo-v2.5"))), [])
|
|
1456
|
+
assert.deepEqual(overlayIds(resolvePolicy(M("deepseek", "deepseek-v4-pro"))), [])
|
|
1457
|
+
// A hypothetical future GLM does not silently inherit the 5.3 overlay.
|
|
1458
|
+
assert.deepEqual(overlayIds(resolvePolicy(M("zai", "glm-5.4"))), [])
|
|
1459
|
+
})
|
|
1460
|
+
|
|
1461
|
+
test("resolvePolicy: malformed and unknown model ids resolve safely", () => {
|
|
1462
|
+
for (const bad of [undefined, null, {}, 42, "gpt-5.6", [], true]) {
|
|
1463
|
+
const r = resolvePolicy(bad)
|
|
1464
|
+
assert.equal(r.family, "neutral")
|
|
1465
|
+
assert.equal(r.matchType, "neutral")
|
|
1466
|
+
assert.equal(baseId(r), "neutral.none")
|
|
1467
|
+
assert.deepEqual(r.overlays, [])
|
|
1468
|
+
}
|
|
1469
|
+
// A provider with no model identity gets no family.
|
|
1470
|
+
assert.equal(resolvePolicy({ providerID: "openai" }).family, "neutral")
|
|
1471
|
+
})
|
|
1472
|
+
|
|
1473
|
+
test("resolvePolicy: unknown providers/creators do not gain policy from transport identity", () => {
|
|
1474
|
+
const orUnknown = resolvePolicy(M("openrouter", "acme/mystery-model-9"))
|
|
1475
|
+
assert.equal(orUnknown.family, "neutral")
|
|
1476
|
+
assert.deepEqual(orUnknown.overlays, [])
|
|
1477
|
+
assert.equal(orUnknown.transport.kind, "openrouter")
|
|
1478
|
+
assert.equal(orUnknown.transport.sessionAffinityHeader, "x-session-id")
|
|
1479
|
+
|
|
1480
|
+
// gpt-5.6 on a non-OpenAI gateway stays neutral even though the id looks OpenAI
|
|
1481
|
+
assert.equal(resolvePolicy(M("some-gateway", "gpt-5.6")).family, "neutral")
|
|
1482
|
+
assert.equal(resolvePolicy(M("xiaomi", "not-a-mimo")).family, "neutral")
|
|
1483
|
+
})
|
|
1484
|
+
|
|
1485
|
+
test("resolvePolicy: transport is computed independently from cache policy", () => {
|
|
1486
|
+
const ds = resolvePolicy(M("deepseek", "deepseek-v4-pro"))
|
|
1487
|
+
assert.equal(ds.transport.kind, "direct")
|
|
1488
|
+
assert.equal(ds.transport.sessionAffinityHeader, null)
|
|
1489
|
+
|
|
1490
|
+
const or = resolvePolicy(M("openrouter", "openai/gpt-5.6-luna"))
|
|
1491
|
+
assert.equal(or.family, "gpt-5.6")
|
|
1492
|
+
assert.equal(or.transport.kind, "openrouter")
|
|
1493
|
+
assert.equal(or.transport.sessionAffinityHeader, "x-session-id")
|
|
1494
|
+
|
|
1495
|
+
const zai = resolvePolicy(M("zai", "glm-5.3"))
|
|
1496
|
+
assert.equal(zai.transport.kind, "direct")
|
|
1497
|
+
assert.equal(zai.transport.sessionAffinityHeader, null)
|
|
1498
|
+
|
|
1499
|
+
const noProvider = resolvePolicy({ modelID: "gpt-6-astra" })
|
|
1500
|
+
assert.equal(noProvider.transport.kind, "unknown")
|
|
1501
|
+
})
|
|
1502
|
+
|
|
1503
|
+
test("resolvePolicy is pure and does not mutate its input", () => {
|
|
1504
|
+
const model = Object.freeze({ providerID: "openai", modelID: "gpt-5.6-sol" })
|
|
1505
|
+
const r = resolvePolicy(model)
|
|
1506
|
+
assert.equal(r.family, "gpt-5.6")
|
|
1507
|
+
assert.equal(model.modelID, "gpt-5.6-sol")
|
|
1508
|
+
})
|
|
1509
|
+
|
|
1510
|
+
test("detectPolicy remains compatible with the legacy family resolution", () => {
|
|
1511
|
+
const legacyMap = {
|
|
1512
|
+
"gpt-5.6": POLICY_GPT56,
|
|
1513
|
+
"glm-5.3": POLICY_GLM53,
|
|
1514
|
+
"mimo-v2.6": POLICY_MIMO26,
|
|
1515
|
+
deepseek: POLICY_DEEPSEEK,
|
|
1516
|
+
}
|
|
1517
|
+
const samples = [
|
|
1518
|
+
M("openrouter", "openai/gpt-5.6-luna"),
|
|
1519
|
+
M("openai", "gpt-5.6"),
|
|
1520
|
+
M("openai-compatible", "gpt-5.6"),
|
|
1521
|
+
M("openai", "gpt-5.5"),
|
|
1522
|
+
M("zai", "glm-5.3-flash"),
|
|
1523
|
+
M("zai", "glm-4.6"),
|
|
1524
|
+
M("xiaomi", "mimo-v2.6-flash"),
|
|
1525
|
+
M("xiaomi", "mimo-v2.6-pro-ultraspeed"),
|
|
1526
|
+
M("xiaomi", "mimo-v2.5"),
|
|
1527
|
+
M("deepseek", "deepseek-chat"),
|
|
1528
|
+
M("openrouter", "x-ai/grok-4"),
|
|
1529
|
+
M("anthropic", "claude-sonnet-4-5"),
|
|
1530
|
+
{},
|
|
1531
|
+
null,
|
|
1532
|
+
undefined,
|
|
1533
|
+
42,
|
|
1534
|
+
]
|
|
1535
|
+
for (const m of samples) {
|
|
1536
|
+
const expected = legacyMap[resolveLegacyFamily(m)] ?? POLICY_NEUTRAL
|
|
1537
|
+
assert.equal(detectPolicy(m), expected)
|
|
1538
|
+
}
|
|
1539
|
+
})
|
|
1540
|
+
|
|
1541
|
+
test("registry is traceable and internally consistent", () => {
|
|
1542
|
+
for (const entry of POLICY_REGISTRY) {
|
|
1543
|
+
assert.ok(entry.inventoryRef, `${entry.id} must cite the inventory`)
|
|
1544
|
+
assert.ok(BASELINES[entry.baseline], `${entry.id} baseline must exist`)
|
|
1545
|
+
for (const overlayId of entry.overlays) {
|
|
1546
|
+
assert.ok(OVERLAYS[overlayId], `${entry.id} overlay ${overlayId} must exist`)
|
|
1547
|
+
}
|
|
1548
|
+
}
|
|
1549
|
+
for (const [id, alias] of Object.entries(MODEL_ALIASES)) {
|
|
1550
|
+
assert.ok(alias.inventoryRef, `alias ${id} must cite the inventory`)
|
|
1551
|
+
assert.ok(alias.family)
|
|
1552
|
+
}
|
|
1553
|
+
})
|