opencode-cache-engine 0.4.0 → 0.4.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -18,15 +18,30 @@ plugin under `~/.config/opencode/plugins/`. For released installs where
18
18
  reproducibility matters, pin an exact package version rather than relying on
19
19
  `@latest` resolution or a moving cache entry; see [Installation](#installation).
20
20
 
21
+ Quick Installation (TUI):
22
+ ``` text
23
+ opencode plugin opencode-cache-engine
24
+ ```
25
+
21
26
  `CacheEngine` is an OpenCode plugin designed for long-running agent sessions where prompt-cache efficiency affects both latency and cost. It keeps the harness conservative for providers whose cache behavior is already automatic, while applying provider-specific optimizations where the provider exposes useful cache controls or where prompt structure can be safely improved.
22
27
 
23
28
  The plugin currently has four cache-policy families:
24
29
 
25
30
  * **DeepSeek** — passive cache observability; request structure is preserved.
26
- * **GPT-5.6** — documented cache-key/options metadata, with prompt text unchanged.
31
+ * **GPT-5.6 and later** — documented cache-key/options metadata, with prompt text
32
+ unchanged. GPT-6 and future 5.6+/6+/7+ versions resolve through the same
33
+ documented boundary.
27
34
  * **GLM-5.3** — narrow, content-preserving `<env>` relocation and diagnostics.
28
35
  * **MiMo-V2.6** — narrow, content-preserving `<env>` relocation and diagnostics.
29
36
 
37
+ Family classification is not hard-coded in the runtime. A pure policy registry
38
+ and resolver in `src/cache-policy-core.mjs` returns a structured result
39
+ (`creator`, `family`, `baseline`, `overlays`, `transport`, `matchType`,
40
+ `matchReason`), and the hooks gate their behavior on that result. The registry is
41
+ the single runtime source of policy classification. The first-party research
42
+ behind each registry entry is recorded in
43
+ [docs/cache-policy-inventory.md](docs/cache-policy-inventory.md).
44
+
30
45
  For both MiMo-V2.6 and GLM-5.3, CacheEngine adds its deterministic
31
46
  `x-session-id` request header only when OpenCode identifies the actual provider
32
47
  as `openrouter`. It does not add that OpenRouter-specific header for
@@ -44,12 +59,12 @@ The plugin operates at the OpenCode harness level rather than implementing a pro
44
59
 
45
60
  It:
46
61
 
47
- 1. Detects the model/provider family in use.
48
- 2. Applies only the policy appropriate for that family.
62
+ 1. Resolves the model/provider policy through the registry resolver.
63
+ 2. Applies only the mutations registered for that policy.
49
64
  3. Observes system-prompt and tool-definition stability.
50
65
  4. Records provider-reported cache token usage.
51
66
  5. Adds a deterministic compaction continuation block.
52
- 6. Applies GPT-5.6 cache-control metadata.
67
+ 6. Applies GPT-5.6-and-later cache-control metadata.
53
68
  7. Applies the GLM-5.3 and MiMo-V2.6 volatile-environment relocation.
54
69
  8. Records diagnostics that help determine whether prompt-shape changes correlate with cache behavior.
55
70
  9. Records MiMo/GLM affinity outcomes and provider-identity changes.
@@ -115,7 +130,7 @@ DeepSeek:
115
130
 
116
131
  ### Policy: active cache control
117
132
 
118
- GPT-5.6 is the only current policy that actively injects cache-control request metadata.
133
+ GPT-5.6 and later is the only policy family that actively injects cache-control request metadata.
119
134
 
120
135
  The plugin adds:
121
136
 
@@ -405,7 +420,7 @@ and are reported diagnostically; the message content is left untouched.
405
420
  | Policy family | Detection | Prompt text changed? | Cache metadata changed? | OpenRouter affinity header | Primary cache signal |
406
421
  | ------------- | --------- | ------------------- | ----------------------- | -------------------------- | -------------------- |
407
422
  | DeepSeek | `deepseek` | No | No | None | provider `cache.read` / `cache.write` |
408
- | GPT-5.6 | `gpt-5.6*` on OpenAI-ish endpoints | No | Yes: `prompt_cache_key` + options | None | provider cache tokens |
423
+ | GPT-5.6 and later | version boundary `gpt-<major>[.<minor>] ≥ 5.6` on OpenAI-ish endpoints (includes GPT-6) | No | Yes: `prompt_cache_key` + options | None | provider cache tokens |
409
424
  | GLM-5.3 | `glm-5.3*` | Yes, narrowly (`<env>` tail) | No provider cache key | `x-session-id` on OpenRouter only | provider cache tokens (GLM ratio) |
410
425
  | MiMo-V2.6 | Flash / Pro only | Yes, narrowly (`<env>` tail) | No: implicit caching only | `x-session-id` on OpenRouter only | `cached_tokens / prompt_tokens` |
411
426
 
@@ -923,7 +938,7 @@ The model detector recognizes:
923
938
  * GLM-5.3 variants
924
939
  * MiMo-V2.6 Flash and Pro (`xiaomi/mimo-v2.6-flash`, `mimo-v2.6-pro`, ...)
925
940
 
926
- GPT-5.6 has an additional OpenAI/Azure-context check so a string containing `gpt-5.6` does not automatically cause GPT-specific fields to be sent to an unrelated endpoint.
941
+ The GPT-5.6-and-later family has an additional OpenAI/Azure-context check, so a string containing a qualifying GPT version (for example `gpt-5.6` or `gpt-6`) does not automatically cause GPT-specific fields to be sent to an unrelated endpoint.
927
942
 
928
943
  MiMo detection targets exactly Flash and Pro: it excludes `mimo-v2.5`,
929
944
  `mimo-v2.5-pro`, and `mimo-v2.6-pro-ultraspeed`.
@@ -953,7 +968,8 @@ For that reason, a stable provider route is preferable when your goal is to meas
953
968
 
954
969
  # Architecture
955
970
 
956
- The implementation is split into two layers.
971
+ The implementation is split across a hook entry point, a pure logic core, and a
972
+ pure policy registry.
957
973
 
958
974
  ## `cache-engine.ts`
959
975
 
@@ -989,7 +1005,7 @@ This contains dependency-light pure logic.
989
1005
 
990
1006
  It owns:
991
1007
 
992
- * provider classification
1008
+ * the legacy `detectPolicy()` compatibility wrapper (delegating to the registry)
993
1009
  * configuration parsing
994
1010
  * hashing
995
1011
  * canonicalization
@@ -1006,6 +1022,42 @@ Keeping these functions in plain JavaScript allows the logic to be tested indepe
1006
1022
 
1007
1023
  ---
1008
1024
 
1025
+ ## `cache-policy-core.mjs`
1026
+
1027
+ This is the pure policy registry and resolver. It separates cache policy from
1028
+ request mutation:
1029
+
1030
+ * creator / family classification
1031
+ * baseline cache-policy descriptors (documented facts)
1032
+ * model-specific overlays (for example GLM/MiMo `<env>` relocation)
1033
+ * transport capabilities (for example OpenRouter `x-session-id` affinity)
1034
+ * explicit, inventory-traceable inheritance (`inheritsFrom`)
1035
+ * safe neutral fallback for unknown or future models
1036
+
1037
+ `resolvePolicy(model)` returns `creator`, `family`, `baseline`, `overlays`,
1038
+ `transport`, `matchType`, and `matchReason`. `resolveRuntimePolicy(model)`
1039
+ returns the runtime-facing descriptor the hooks consume: the legacy policy
1040
+ string plus explicit capability flags.
1041
+
1042
+ Only registry entries marked `legacy` enable runtime behavior; documented but
1043
+ non-legacy entries (for example `mimo-v2.6-pro-ultraspeed`) and all unknown
1044
+ models resolve to a neutral runtime. A newer or unknown model therefore never
1045
+ inherits a current model's mutation unless the registry explicitly registers it.
1046
+ The GPT family is a documented exception in the sense that its boundary is
1047
+ version-based (`GPT-5.6 and later`), so GPT-6 and future 5.6+/6+/7+ versions are
1048
+ covered by the registered boundary rather than by an exact-model list.
1049
+
1050
+ Transport is kept separate from cache policy: OpenRouter affinity is a transport
1051
+ capability, not part of a creator's cache semantics. Overlays are also explicit,
1052
+ so being classified into a family does not by itself enable a prompt
1053
+ transformation.
1054
+
1055
+ The module is pure: no network calls and no runtime documentation lookups. The
1056
+ legacy `detectPolicy()` in `cache-engine-core.mjs` remains a thin compatibility
1057
+ wrapper over the resolver's legacy path.
1058
+
1059
+ ---
1060
+
1009
1061
  ## Tests
1010
1062
 
1011
1063
  The repository's test suite validates the provider-independent and provider-specific logic.
@@ -1013,6 +1065,7 @@ The repository's test suite validates the provider-independent and provider-spec
1013
1065
  Coverage includes:
1014
1066
 
1015
1067
  * model detection
1068
+ * policy registry resolution and runtime-policy equivalence
1016
1069
  * GPT cache-key stability
1017
1070
  * GPT cache-option defaults
1018
1071
  * protection against overwriting existing cache options
@@ -1176,11 +1229,14 @@ opencode-cache-engine/
1176
1229
  ├── src/
1177
1230
  │ ├── cache-engine.ts
1178
1231
  │ ├── cache-engine-core.mjs
1232
+ │ ├── cache-policy-core.mjs
1179
1233
  │ └── tui.mjs
1180
1234
  ├── test/
1181
1235
  │ └── cache-engine.test.mjs
1182
1236
  ├── examples/
1183
1237
  │ └── cache-engine.json
1238
+ ├── docs/
1239
+ │ └── cache-policy-inventory.md
1184
1240
  ├── package.json
1185
1241
  ├── README.md
1186
1242
  └── LICENSE
@@ -1219,7 +1275,7 @@ release, use:
1219
1275
  ```json
1220
1276
  {
1221
1277
  "plugin": [
1222
- "opencode-cache-engine@0.3.6"
1278
+ "opencode-cache-engine@0.4.1"
1223
1279
  ]
1224
1280
  }
1225
1281
  ```
@@ -1286,7 +1342,7 @@ A prefix change is a diagnostic signal, not automatic proof of a cache miss.
1286
1342
 
1287
1343
  ## GPT-5.6 cache options are missing
1288
1344
 
1289
- Verify that the model is actually classified as GPT-5.6 and that the endpoint is recognized as OpenAI/Azure-compatible.
1345
+ Verify that the model is within the documented GPT-5.6-and-later boundary (for example `gpt-5.6-*` or `gpt-6-*`) and that the endpoint is recognized as OpenAI/Azure-compatible.
1290
1346
 
1291
1347
  The detector intentionally rejects ambiguous OpenAI-compatible providers rather than guessing.
1292
1348
 
@@ -32,22 +32,47 @@ Caching behavior is never inferred from pricing alone, from one SDK's type
32
32
  declarations, or from a third-party blog when first-party documentation exists.
33
33
  "unknown — first-party docs insufficient" is used instead of a guess.
34
34
 
35
+ ## Runtime integration
36
+
37
+ This inventory is the research input for the policy registry in
38
+ `src/cache-policy-core.mjs`. As of **v0.4.1** the runtime hooks consume
39
+ `resolveRuntimePolicy(model)` and gate behavior on the registry's explicit
40
+ capabilities, so the "CacheEngine current treatment" column below describes
41
+ resolver-driven behavior.
42
+
43
+ Two guarantees follow from that migration:
44
+
45
+ - `resolveRuntimePolicy(model).policy` equals the legacy `detectPolicy(model)`
46
+ string, so telemetry and gating are unchanged for every model supported in
47
+ v0.3.6.
48
+ - Registry entries marked non-`legacy` (for example
49
+ `mimo-v2.6-pro-ultraspeed`) and all unknown models resolve to a neutral
50
+ runtime, so no documented-but-unwired model gains a current model's mutation.
51
+
52
+ The runtime reads the registry at classification time only; there are no network
53
+ calls and no runtime documentation lookups.
54
+
35
55
  ## CacheEngine current treatment (baseline for the matrix)
36
56
 
37
- Source: `src/cache-engine-core.mjs` (detection, transforms) and
38
- `src/cache-engine.ts` (hooks), revision `19b87f2`.
57
+ Source: the pure registry/resolver (`src/cache-policy-core.mjs`) and the hook
58
+ entry (`src/cache-engine.ts`), as wired in v0.4.1. The behavior described here is
59
+ the same as at the pre-resolver revision `19b87f2`; only the classification
60
+ source changed.
39
61
 
40
62
  | Family | Detection (verbatim) | Current treatment | Affinity header |
41
63
  | --- | --- | --- | --- |
42
64
  | DeepSeek | `/deepseek/i` on `${apiID} ${modelID}` or `providerID` | Passive; no mutation | none |
43
- | GPT-5.6 | `/gpt-5\.6(?![\d.])/i` on slug **and** `isOpenAIish` (provider `openai`/`azure`, slug `openai/`/`azure/`, or npm `@ai-sdk/openai`/`@ai-sdk/azure`) | Inject missing `promptCacheKey` + `promptCacheOptions` (`implicit`, `30m`) | none |
65
+ | GPT-5.6 | version boundary `gpt-<major>[.<minor>] ≥ 5.6` on slug **and** `isOpenAIish` (provider `openai`/`azure`, slug `openai/`/`azure/`, or npm `@ai-sdk/openai`/`@ai-sdk/azure`); since v0.4.2 covers GPT-6 and later | Inject missing `promptCacheKey` + `promptCacheOptions` (`implicit`, `30m`) | none |
44
66
  | GLM-5.3 | `/glm-5\.3(?![\d.])/i` on slug | Relocate identifiable `<env>` block to system tail | `x-session-id` only when `providerID === "openrouter"` |
45
67
  | MiMo-V2.6 | `/mimo-v2\.6-(flash\|pro)(?![\w-])/i` on slug | Relocate identifiable `<env>` block to system tail; provider-change telemetry | `x-session-id` only when `providerID === "openrouter"` |
46
68
  | Neutral | everything else | Byte-untouched | none |
47
69
 
48
70
  Detection consequences worth stating explicitly:
49
71
 
50
- - `gpt-6*` / `gpt-6-*` is **not** matched → neutral. **[O]**
72
+ - `gpt-6` / `gpt-6-*` (and any future 5.6+/6+/7+ version) is matched by the
73
+ documented GPT-5.6-and-later boundary → GPT policy. **[O]** (v0.4.2)
74
+ - `gpt-5.5`, `gpt-5.2`, `gpt-4o`, and the malformed `gpt-5.60` are **not**
75
+ matched → neutral. **[O]**
51
76
  - `deepseek-v5` (or any future `*deepseek*` id) matches the passive DeepSeek
52
77
  branch because the regex is a bare substring test. **[O]**
53
78
  - `mimo-v2.6-pro-ultraspeed` is **not** matched: the `(?![\w-])` lookahead fails
@@ -95,8 +120,10 @@ catch-all "earlier models". [D]
95
120
  baseline: it injects only `promptCacheKey` + `promptCacheOptions{mode:implicit,
96
121
  ttl:"30m"}`, preserves runtime-supplied values, and never sets context/output
97
122
  limits. It does not use explicit breakpoints or `prewarm`, which is a subset of
98
- the documented capability. GPT-6 is **not** classified and therefore gets no GPT
99
- cache metadata — a documented-scope gap, not a documented incompatibility.
123
+ the documented capability. Since v0.4.2 the policy resolves by the documented
124
+ "GPT-5.6 and later" boundary, so GPT-6 (astra/sol/luna) receives the same
125
+ baseline; no GPT-6 cache-control exception is documented (OpenAI *Prompt
126
+ caching* guide, re-verified 2026-09-27), and none is coded.
100
127
 
101
128
  ---
102
129
 
@@ -254,7 +281,7 @@ made here); **hold** = do not inherit without first-party evidence.
254
281
  | Creator | Model / example pattern | Cache policy (documented) | CacheEngine current treatment | Recommended family inheritance | Recommended exact-model exception | Confidence | Source | Verified |
255
282
  | --- | --- | --- | --- | --- | --- | --- | --- | --- |
256
283
  | OpenAI | `gpt-5.6`, `gpt-5.6-*` (sol/terra/luna/cyber) | "GPT-5.6 and later": implicit default, optional explicit breakpoints, min 1,024, TTL 30m, write 1.25×/read 0.1× | GPT policy: inject `promptCacheKey` + `promptCacheOptions{implicit,30m}` | keep | none documented; CacheEngine's subset is valid | High (docs) / Medium (treatment) | OpenAI *Prompt caching*; *GPT-5.6 Sol* | 2026-09-26 |
257
- | OpenAI | GPT-6 / current later GPT family: `gpt-6`, `gpt-6-*` (astra/sol/luna) | Inherits the GPT-5.6-and-later policy | **Not matched** → neutral (no GPT metadata) | **extend**: docs classify GPT-6 as "later", so it inherits the same policy | none documented | High (docs) / High (gap) | OpenAI *Using GPT-6*; *GPT-6 Astra*; *Prompt caching* | 2026-09-26 |
284
+ | OpenAI | GPT-6 / current later GPT family: `gpt-6`, `gpt-6-*` (astra/sol/luna) | Inherits the GPT-5.6-and-later policy | **GPT policy** via the GPT-5.6-and-later boundary (v0.4.2) | **keep** — covered by the boundary predicate; no exact-model entry needed | none documented | High (docs) / High (treatment) | OpenAI *Using GPT-6*; *GPT-6 Astra*; *Prompt caching* | 2026-09-27 |
258
285
  | OpenAI | pre-5.6 negative controls: `gpt-5.5`, `gpt-5.4`, `gpt-5.2`, `gpt-5.1`, `gpt-5`, `gpt-4.1`, `gpt-4o` | Implicit only; different min-length class; `in_memory`/`24h` retention; `prompt_cache_key` for routing | neutral | **hold** — do not inherit 5.6 policy | n/a | High | OpenAI *Prompt caching*; *Pricing* | 2026-09-26 |
259
286
  | DeepSeek | V4: `deepseek-v4-pro`, legacy `deepseek-v4-flash` | Provider-wide automatic disk cache; implicit; prefix-unit matching; hit/miss token fields | Passive (no mutation) | keep | none | High | DeepSeek *Context Caching*; *Models & Pricing* | 2026-09-26 |
260
287
  | DeepSeek | V4.1 / current V4-family: `deepseek-flash` (MODEL VERSION "DeepSeek-V4.1-Flash") | Same provider-wide automatic policy; cache-hit pricing listed for both current models | Passive | keep (creator/family baseline) | none documented | High | DeepSeek *Models & Pricing*; *news260910* | 2026-09-26 |
@@ -311,3 +338,27 @@ a newer model inherits an older policy. They must not be resolved by guessing.
311
338
  - The existing test suite was run only to confirm the repository remains green
312
339
  (116/116).
313
340
  - This release contains research and documentation only.
341
+
342
+ ### Follow-up: v0.4.1 runtime integration
343
+
344
+ - v0.4.0 added the registry (`src/cache-policy-core.mjs`) without wiring it.
345
+ - v0.4.1 wired the runtime to `resolveRuntimePolicy()` and preserved behavior:
346
+ every model supported in v0.3.6 keeps its prior treatment, and non-legacy or
347
+ unknown models remain neutral.
348
+ - The v0.4.0 research statements above are unchanged; only the runtime now reads
349
+ this registry as its single source of policy classification.
350
+
351
+ ### Follow-up: v0.4.2 GPT-5.6-and-later boundary
352
+
353
+ - OpenAI's documented "GPT-5.6 and later" boundary was re-verified against the
354
+ first-party *Prompt caching* guide and *Using GPT-6* guide on **2026-09-27**.
355
+ GPT-6 (astra/sol/luna) is documented in the same cache regime with no
356
+ cache-control exception.
357
+ - v0.4.2 replaces the exact `gpt-5.6` string match with a version-boundary
358
+ predicate (`gpt-<major>[.<minor>] ≥ 5.6`), still gated on the OpenAI-ish
359
+ provider check. GPT-6 and future 5.6+/6+/7+ models therefore need no
360
+ exact-model registry entry, while `gpt-5.5` and earlier and malformed ids such
361
+ as `gpt-5.60` stay neutral.
362
+ - The only OpenAI cache controls injected remain `promptCacheKey` +
363
+ `promptCacheOptions{implicit,30m}`; no breakpoint or prewarm behavior was
364
+ added. DeepSeek, GLM, MiMo, and OpenRouter affinity behavior are unchanged.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "opencode-cache-engine",
3
- "version": "0.4.0",
3
+ "version": "0.4.2",
4
4
  "private": false,
5
5
  "description": "Provider-aware prompt-cache optimization and telemetry for OpenCode",
6
6
  "keywords": [
@@ -4,12 +4,7 @@ import {
4
4
  DEFAULT_CONFIG_PATH,
5
5
  DIGEST_TEMPLATE,
6
6
  affinityTelemetryFields,
7
- POLICY_GLM53,
8
- POLICY_GPT56,
9
- POLICY_MIMO26,
10
- POLICY_NEUTRAL,
11
7
  createRecorder,
12
- detectPolicy,
13
8
  detectReasoningIssues,
14
9
  digestDecision,
15
10
  ensureMetricsDir,
@@ -37,6 +32,7 @@ import {
37
32
  toolFingerprint,
38
33
  toolWireFingerprint,
39
34
  } from "./cache-engine-core.mjs"
35
+ import { resolveRuntimePolicy } from "./cache-policy-core.mjs"
40
36
 
41
37
  // ---------------------------------------------------------------------------
42
38
  // cache-engine
@@ -101,7 +97,21 @@ type Shape = {
101
97
  toolCount: number | null
102
98
  }
103
99
 
104
- type ModelInfo = { family: string; providerID: string; modelID: string }
100
+ // Resolved runtime policy (the registry's runtime descriptor). `policy` is the
101
+ // legacy telemetry/state string emitted by the pre-v0.4.0 classifier.
102
+ type PolicyRuntime = {
103
+ policy: string
104
+ isNeutral: boolean
105
+ gptCacheMetadata: boolean
106
+ envRelocation: "glm" | "mimo" | null
107
+ thinkingIntegrity: boolean
108
+ cacheRatio: "glm" | "mimo" | null
109
+ providerChange: "glm" | "mimo" | null
110
+ prefixDiagnostics: boolean
111
+ openRouterAffinity: boolean
112
+ }
113
+
114
+ type ModelInfo = { family: string; providerID: string; modelID: string; caps: PolicyRuntime }
105
115
 
106
116
  type ToolCache = { semanticToolsHash: string | null; wireToolsHash: string | null }
107
117
 
@@ -258,19 +268,21 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
258
268
  // gpt56/glm53/deepseek classification once established.
259
269
  const rememberModel = (sid: string, model: ChatParamsModel | undefined): ModelInfo | null => {
260
270
  if (!model) return null
261
- const family = detectPolicy(model)
271
+ // Single runtime source of policy classification: the registry resolver.
272
+ const caps = resolveRuntimePolicy(model) as PolicyRuntime
262
273
  const s = get(sid)
263
274
  const info: ModelInfo = {
264
- family,
275
+ family: caps.policy,
265
276
  providerID: String(model.providerID ?? ""),
266
277
  modelID: String(model.api?.id ?? model.id ?? ""),
278
+ caps,
267
279
  }
268
- if (s.modelInfo == null || (s.modelInfo.family === POLICY_NEUTRAL && family !== POLICY_NEUTRAL)) {
280
+ if (s.modelInfo == null || (s.modelInfo.caps.isNeutral && !caps.isNeutral)) {
269
281
  s.modelInfo = info
270
282
  // A title/summary request (neutral small model) may have established the
271
283
  // system baseline first. Its "powered by the model named ..." env line
272
284
  // differs from the real model's, so re-baseline on upgrade.
273
- if (s.baselineSystem !== null && s.modelInfo.family !== POLICY_NEUTRAL) {
285
+ if (s.baselineSystem !== null && !caps.isNeutral) {
274
286
  s.baselineSystem = null
275
287
  s.shape = null
276
288
  }
@@ -350,11 +362,11 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
350
362
 
351
363
  s.lastProcessedMessageID = nextProcessedCursor(firstPage, startCursor)
352
364
 
353
- const family = s.modelInfo?.family
365
+ const caps = s.modelInfo?.caps
354
366
  const glmIntegrity =
355
- family === POLICY_GLM53 &&
356
- policyEnabled(cfg, POLICY_GLM53) &&
357
- cfg.policies?.[POLICY_GLM53]?.preserveThinkingIntegrity === true
367
+ caps?.thinkingIntegrity === true &&
368
+ policyEnabled(cfg, caps.policy) &&
369
+ cfg.policies?.[caps.policy]?.preserveThinkingIntegrity === true
358
370
 
359
371
  if (glmIntegrity && reasoningNewestFirst.length > 0) {
360
372
  // messages arrive newest-first; process oldest->newest so `seen` grows
@@ -417,11 +429,11 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
417
429
  recFields.model = s.modelInfo.modelID
418
430
  recFields.policy = s.modelInfo.family
419
431
  }
420
- if (family === POLICY_GLM53) {
432
+ if (caps?.cacheRatio === "glm") {
421
433
  recFields.promptTokens = read + write + input
422
434
  recFields.glmHitRate = glmHitRatio(read, write, input)
423
435
  }
424
- if (family === POLICY_MIMO26) {
436
+ if (caps?.cacheRatio === "mimo") {
425
437
  // Provider-reported cached tokens / total prompt tokens. The runtime's
426
438
  // `input` is the non-cached prompt input and `cache.read` is the
427
439
  // cached prompt input, so total prompt tokens are derived as read +
@@ -431,7 +443,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
431
443
  recFields.promptTokens = promptTokens
432
444
  recFields.cachedTokens = read
433
445
  recFields.cacheHitRate = mimoHitRate(read, promptTokens)
434
- if (cfg.policies?.[POLICY_MIMO26]?.stickySession === true) {
446
+ if (cfg.policies?.[caps.policy]?.stickySession === true) {
435
447
  recFields.stickySessionId = mimoSessionIdFor(sid)
436
448
  }
437
449
  // Prefer the latest live provider identity over the latched one.
@@ -440,10 +452,10 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
440
452
  recFields.model = s.mimoProvider.modelID
441
453
  }
442
454
  }
443
- if (family === POLICY_GPT56 && s.gptInjected) {
455
+ if (caps?.gptCacheMetadata === true && s.gptInjected) {
444
456
  recFields.keyStrategy = "session"
445
- recFields.mode = cfg.policies?.[POLICY_GPT56]?.mode
446
- recFields.ttl = cfg.policies?.[POLICY_GPT56]?.ttl
457
+ recFields.mode = cfg.policies?.[caps.policy]?.mode
458
+ recFields.ttl = cfg.policies?.[caps.policy]?.ttl
447
459
  }
448
460
  rec.record(recFields)
449
461
  } else {
@@ -495,8 +507,9 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
495
507
  try {
496
508
  const model = input.model as unknown as ChatParamsModel
497
509
  const providerID = String(model?.providerID ?? "")
498
- const family = detectPolicy(model)
499
- if (family !== POLICY_MIMO26 && family !== POLICY_GLM53) return
510
+ const caps = resolveRuntimePolicy(model) as PolicyRuntime
511
+ if (!caps.openRouterAffinity) return
512
+ const family = caps.policy
500
513
 
501
514
  const hasSessionIDHeader = (headers?: Record<string, string>) =>
502
515
  Object.keys(headers ?? {}).some((name) => name.toLowerCase() === "x-session-id")
@@ -536,6 +549,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
536
549
  try {
537
550
  const info = rememberModel(input.sessionID, input.model as unknown as ChatParamsModel)
538
551
  const family = info?.family
552
+ const caps = info?.caps
539
553
 
540
554
  // ---- MiMo-V2.6: provider-switch diagnostics (telemetry only) ---------
541
555
  // MiMo cache lives at the provider side, so a provider change within one
@@ -545,7 +559,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
545
559
  // NOTE: OpenRouter's *upstream* provider selection (e.g. xiaomi/fp8) is
546
560
  // not exposed to plugins; only the OpenCode providerID/modelID are
547
561
  // observable here.
548
- if (family === POLICY_MIMO26 && policyEnabled(cfg, POLICY_MIMO26) && info) {
562
+ if (caps?.providerChange === "mimo" && policyEnabled(cfg, caps.policy) && info) {
549
563
  // Use the LIVE model identity (not the latched one) so a provider
550
564
  // switch within the session is actually observable.
551
565
  const live = input.model as unknown as ChatParamsModel
@@ -558,7 +572,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
558
572
  const ev = providerChangeEvent(s.mimoProvider, cur)
559
573
  if (ev.changed) {
560
574
  const sticky =
561
- cfg.policies?.[POLICY_MIMO26]?.stickySession === true
575
+ cfg.policies?.[caps.policy]?.stickySession === true
562
576
  ? { stickySessionId: mimoSessionIdFor(input.sessionID) }
563
577
  : {}
564
578
  rec.record({
@@ -566,7 +580,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
566
580
  sid: input.sessionID,
567
581
  ts: Date.now(),
568
582
  reason: "mimo_provider_changed",
569
- policy: POLICY_MIMO26,
583
+ policy: caps.policy,
570
584
  from: ev.from,
571
585
  to: ev.to,
572
586
  ...sticky,
@@ -579,7 +593,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
579
593
  }
580
594
 
581
595
  // ---- GLM-5.3 provider identity observation (telemetry only) ----------
582
- if (family === POLICY_GLM53 && policyEnabled(cfg, POLICY_GLM53) && info) {
596
+ if (caps?.providerChange === "glm" && policyEnabled(cfg, caps.policy) && info) {
583
597
  const live = input.model as unknown as ChatParamsModel
584
598
  const cur = {
585
599
  providerID: String(live?.providerID ?? ""),
@@ -594,7 +608,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
594
608
  sid: input.sessionID,
595
609
  ts: Date.now(),
596
610
  reason: "glm_provider_changed",
597
- policy: POLICY_GLM53,
611
+ policy: caps.policy,
598
612
  from: ev.from,
599
613
  to: ev.to,
600
614
  note: "OpenCode providerID changed; provider-specific upstream routing is not plugin-visible",
@@ -605,12 +619,12 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
605
619
  return
606
620
  }
607
621
 
608
- if (!(family === POLICY_GPT56 && policyEnabled(cfg, POLICY_GPT56))) {
622
+ if (!(caps?.gptCacheMetadata === true && policyEnabled(cfg, caps.policy))) {
609
623
  // DeepSeek / GLM / neutral: nothing to inject. GLM has no cache-key API;
610
624
  // DeepSeek caching is fully passive; we never mutate requests for them.
611
625
  return
612
626
  }
613
- const gpol = cfg.policies?.[POLICY_GPT56]
627
+ const gpol = cfg.policies?.[caps.policy]
614
628
  const applyRoot = gpol?.cacheRootKey !== false
615
629
  const compaction = input.agent === "compaction"
616
630
  const isolated = compaction && gpol?.compactionCacheIsolation === true
@@ -678,7 +692,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
678
692
  kind: "cache-options",
679
693
  sid: input.sessionID,
680
694
  ts: Date.now(),
681
- policy: POLICY_GPT56,
695
+ policy: caps.policy,
682
696
  provider: info.providerID,
683
697
  model: info.modelID,
684
698
  keyStrategy: applyRoot ? "cache-root" : "session",
@@ -784,43 +798,40 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
784
798
  const model = input.model as unknown as ChatParamsModel
785
799
  rememberModel(sid, model)
786
800
  const s = get(sid)
787
- const family = s.modelInfo?.family
801
+ const caps = s.modelInfo?.caps
788
802
 
789
803
  // ---- GLM-5.3 / MiMo-V2.6 input-shape stabilization ------------------
790
804
  // Relocate the identifiable volatile env block (per-day date) to the
791
805
  // tail of the single system string, content-preserving, ONLY when the
792
- // family opts into system stabilization and the block markers are
793
- // present exactly. Never touches other content/order; never applied to
794
- // other families. The runtime passes a single-element system array
795
- // (verified against the installed runtime), so no generic reordering is
796
- // involved.
806
+ // resolved runtime policy opts into system stabilization and the block
807
+ // markers are present exactly. Never touches other content/order; never
808
+ // applied to other families. The runtime passes a single-element system
809
+ // array (verified against the installed runtime), so no generic
810
+ // reordering is involved.
797
811
  //
798
812
  // In-place mutation note: request.ts keeps using its own local `system`
799
813
  // array after the hook (the trigger's returned output is ignored), so
800
814
  // reassigning `output.system = [...]` would be lost. We rewrite the
801
815
  // single element in place instead.
802
816
  let systemText = output.system.join("\n")
803
- const glmStabilize =
804
- family === POLICY_GLM53 &&
805
- policyEnabled(cfg, POLICY_GLM53) &&
806
- cfg.policies?.[POLICY_GLM53]?.stabilizeSystem === true
807
- const mimoStabilize =
808
- family === POLICY_MIMO26 &&
809
- policyEnabled(cfg, POLICY_MIMO26) &&
810
- cfg.policies?.[POLICY_MIMO26]?.stabilizeSystem === true
811
- if ((glmStabilize || mimoStabilize) && output.system.length === 1) {
817
+ const envVariant = caps?.envRelocation ?? null
818
+ const stabilize =
819
+ envVariant !== null &&
820
+ policyEnabled(cfg, caps!.policy) &&
821
+ cfg.policies?.[caps!.policy]?.stabilizeSystem === true
822
+ if (stabilize && output.system.length === 1) {
812
823
  const rel = relocateVolatileEnvBlock(output.system[0])
813
824
  if (rel.changed) {
814
825
  output.system[0] = rel.text
815
826
  systemText = rel.text
816
- if (family === POLICY_MIMO26) {
827
+ if (envVariant === "mimo") {
817
828
  log("debug", "mimo system env block relocated to suffix", { sid })
818
829
  rec.record({
819
830
  kind: "boundary",
820
831
  sid,
821
832
  ts: Date.now(),
822
833
  reason: "mimo_system_env_relocated",
823
- policy: POLICY_MIMO26,
834
+ policy: caps!.policy,
824
835
  provider: s.modelInfo?.providerID,
825
836
  model: s.modelInfo?.modelID,
826
837
  })
@@ -906,13 +917,13 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
906
917
  // MiMo-specific explicit diagnostic: the STABLE prefix changed (not
907
918
  // just the relocated volatile env suffix). Reported only; the new
908
919
  // content is never overwritten with a stale snapshot.
909
- if (family === POLICY_MIMO26 && reasons.includes("system_stable_prefix_changed")) {
920
+ if (caps?.prefixDiagnostics === true && reasons.includes("system_stable_prefix_changed")) {
910
921
  rec.record({
911
922
  kind: "boundary",
912
923
  sid,
913
924
  ts: Date.now(),
914
925
  reason: "mimo_system_prefix_changed",
915
- policy: POLICY_MIMO26,
926
+ policy: caps!.policy,
916
927
  provider: s.modelInfo?.providerID,
917
928
  model: s.modelInfo?.modelID,
918
929
  changedFields: granular,
@@ -17,8 +17,10 @@
17
17
  // `inventoryRef`. Inheritance is always explicit (`inheritsFrom`); "newer means
18
18
  // same behavior" is never an unconditional rule.
19
19
  //
20
- // This release does not wire the resolver into runtime hooks. detectPolicy()
21
- // remains the compatibility classifier until wiring is approved.
20
+ // As of v0.4.1 the runtime hook layer (cache-engine.ts) consumes
21
+ // resolveRuntimePolicy() as its single source of policy classification.
22
+ // detectPolicy() is retained as the compatibility classifier for the legacy
23
+ // POLICY_* strings.
22
24
 
23
25
  // ---------------------------------------------------------------------------
24
26
  // Model normalization (shared with the legacy classifier)
@@ -54,6 +56,31 @@ export function isOpenAIish(s) {
54
56
  return false
55
57
  }
56
58
 
59
+ // The documented OpenAI cache-policy boundary is the generation phrase
60
+ // "GPT-5.6 and later" (docs/cache-policy-inventory.md §1; OpenAI *Prompt
61
+ // caching* guide, re-verified 2026-09-27). This matcher expresses that boundary
62
+ // by version rather than by an exact-model string, so future 5.6+/6+/7+ models
63
+ // need no registry entry:
64
+ // - major > 5 -> in family
65
+ // - major === 5 && minor >= 6 -> in family
66
+ // - everything else -> out
67
+ // The token must be followed by a non-digit/non-dot boundary, so malformed ids
68
+ // such as "gpt-5.60" and "gpt-5.6.1" do not match (same guard as pre-v0.4.2).
69
+ // OpenAI minor versions are single-digit, so a multi-digit minor is treated as
70
+ // malformed rather than as a higher version.
71
+ export function isGpt56OrLater(slug) {
72
+ const text = String(slug ?? "").toLowerCase()
73
+ const re = /gpt-(\d{1,3})(?:\.(\d))?(?![\d.])/g
74
+ let m
75
+ while ((m = re.exec(text)) !== null) {
76
+ const major = Number(m[1])
77
+ const minor = m[2] === undefined ? 0 : Number(m[2])
78
+ if (major > 5) return true
79
+ if (major === 5 && minor >= 6) return true
80
+ }
81
+ return false
82
+ }
83
+
57
84
  // Candidate ids for exact/alias lookup. Includes the raw apiID/modelID, the
58
85
  // lower-cased forms, and a single stripped transport/vendor prefix
59
86
  // (e.g. "openai/gpt-5.6-luna" -> "gpt-5.6-luna", "xiaomi/mimo-v2.6-flash" ->
@@ -206,6 +233,36 @@ function resolveTransport(s) {
206
233
  return { id: p, kind: "direct", sessionAffinityHeader: null, stickyRouting: false, inventoryRef: "§5 OpenRouter transport" }
207
234
  }
208
235
 
236
+ // ---------------------------------------------------------------------------
237
+ // Runtime capability descriptors
238
+ //
239
+ // The runtime consumes `resolvePolicy(...).runtime` for gating. `policy` is the
240
+ // legacy telemetry/state string, so telemetry stays byte-identical. Every
241
+ // capability is explicit per registry entry: classification into a creator or
242
+ // family never implies a mutation. `legacy: false` entries always resolve to
243
+ // NEUTRAL_RUNTIME, so a future-looking model gains nothing until the registry
244
+ // explicitly says so.
245
+ // ---------------------------------------------------------------------------
246
+
247
+ const NEUTRAL_RUNTIME = Object.freeze({
248
+ policy: "neutral",
249
+ isNeutral: true,
250
+ gptCacheMetadata: false,
251
+ envRelocation: null,
252
+ thinkingIntegrity: false,
253
+ cacheRatio: null,
254
+ providerChange: null,
255
+ prefixDiagnostics: false,
256
+ openRouterAffinity: false,
257
+ })
258
+
259
+ const rt = (policy, overrides = {}) => ({
260
+ ...NEUTRAL_RUNTIME,
261
+ ...overrides,
262
+ policy,
263
+ isNeutral: policy === "neutral",
264
+ })
265
+
209
266
  // ---------------------------------------------------------------------------
210
267
  // Registry
211
268
  //
@@ -219,31 +276,32 @@ function resolveTransport(s) {
219
276
 
220
277
  export const POLICY_REGISTRY = [
221
278
  {
222
- id: "openai.gpt-5.6",
279
+ // v0.4.2: one documented GPT-5.6-and-later family, matched by the version
280
+ // boundary rather than an exact model string. GPT-6 (astra/sol/luna) is
281
+ // documented in the same regime with no cache-control exception, so it
282
+ // inherits this baseline and overlay. Future 5.6+/6+/7+ models resolve here
283
+ // without a new registry entry.
284
+ id: "openai.gpt-5.6-plus",
223
285
  creator: "openai",
224
286
  family: "gpt-5.6",
225
287
  kind: "family",
226
- pattern: /gpt-5\.6(?![\d.])/i,
288
+ predicate: isGpt56OrLater,
227
289
  requiresOpenAIish: true,
228
- exactIds: ["gpt-5.6-sol", "gpt-5.6-terra", "gpt-5.6-luna", "gpt-5.6-cyber"],
290
+ exactIds: [
291
+ "gpt-5.6-sol",
292
+ "gpt-5.6-terra",
293
+ "gpt-5.6-luna",
294
+ "gpt-5.6-cyber",
295
+ "gpt-6-astra",
296
+ "gpt-6-sol",
297
+ "gpt-6-luna",
298
+ ],
229
299
  baseline: "openai.gpt56.cache",
230
300
  overlays: ["gpt56.prompt-cache-options"],
231
301
  legacy: true,
232
- inventoryRef: "§1 OpenAI",
233
- },
234
- {
235
- id: "openai.gpt-6",
236
- creator: "openai",
237
- family: "gpt-6",
238
- kind: "family",
239
- pattern: /gpt-6(?![\d.])/i,
240
- requiresOpenAIish: true,
241
- exactIds: ["gpt-6-astra", "gpt-6-sol", "gpt-6-luna"],
242
- baseline: "openai.gpt56.cache",
243
- inheritsFrom: "gpt-5.6",
244
- overlays: [],
245
- legacy: false,
246
- note: "Documented inheritance of the GPT-5.6-and-later baseline. No CacheEngine overlay is registered for gpt-6 yet.",
302
+ runtime: rt("gpt56", { gptCacheMetadata: true }),
303
+ boundary: "GPT-5.6 and later",
304
+ note: "Documented boundary 'GPT-5.6 and later' (OpenAI Prompt caching guide, re-verified 2026-09-27) includes GPT-6 with no documented cache-control exception. No explicit breakpoint or prewarm behavior is registered.",
247
305
  inventoryRef: "§1 OpenAI",
248
306
  },
249
307
  {
@@ -256,6 +314,13 @@ export const POLICY_REGISTRY = [
256
314
  baseline: "zai.implicit-cache",
257
315
  overlays: ["glm53.env-relocation"],
258
316
  legacy: true,
317
+ runtime: rt("glm53", {
318
+ envRelocation: "glm",
319
+ thinkingIntegrity: true,
320
+ cacheRatio: "glm",
321
+ providerChange: "glm",
322
+ openRouterAffinity: true,
323
+ }),
259
324
  inventoryRef: "§3 Z.AI GLM",
260
325
  },
261
326
  {
@@ -268,6 +333,13 @@ export const POLICY_REGISTRY = [
268
333
  baseline: "xiaomi.implicit-cache",
269
334
  overlays: ["mimo26.env-relocation"],
270
335
  legacy: true,
336
+ runtime: rt("mimo26", {
337
+ envRelocation: "mimo",
338
+ cacheRatio: "mimo",
339
+ providerChange: "mimo",
340
+ prefixDiagnostics: true,
341
+ openRouterAffinity: true,
342
+ }),
271
343
  inventoryRef: "§4 Xiaomi MiMo",
272
344
  },
273
345
  {
@@ -279,6 +351,7 @@ export const POLICY_REGISTRY = [
279
351
  baseline: "xiaomi.implicit-cache",
280
352
  overlays: [],
281
353
  legacy: false,
354
+ runtime: rt("neutral"),
282
355
  policyStatus: "documented-series-member-without-registered-overlay",
283
356
  note: "Documented as a Pro mode in the same V2.6 series, but the inventory does not establish identical cache controls and CacheEngine registers no overlay for it.",
284
357
  inventoryRef: "§4 Xiaomi MiMo",
@@ -293,6 +366,7 @@ export const POLICY_REGISTRY = [
293
366
  baseline: "deepseek.kv-cache",
294
367
  overlays: [],
295
368
  legacy: true,
369
+ runtime: rt("deepseek"),
296
370
  inventoryRef: "§2 DeepSeek",
297
371
  },
298
372
  ]
@@ -301,13 +375,16 @@ export const POLICY_REGISTRY = [
301
375
  // Explicit aliases identified by the inventory
302
376
  // ---------------------------------------------------------------------------
303
377
 
378
+ // `legacy` records whether the pre-v0.4.0 classifier already matched this alias.
379
+ // Only legacy aliases carry runtime capabilities; newer documented aliases are
380
+ // resolved for information but stay runtime-neutral (no new optimization).
304
381
  export const MODEL_ALIASES = {
305
- "gpt-5.6": { canonicalId: "gpt-5.6-sol", family: "gpt-5.6", creator: "openai", inventoryRef: "§1 OpenAI" },
306
- "gpt-daybreak-blue-latest": { canonicalId: "gpt-5.6-sol", family: "gpt-5.6", creator: "openai", inventoryRef: "§1 OpenAI" },
307
- "gpt-daybreak-red-latest": { canonicalId: "gpt-5.6-cyber", family: "gpt-5.6", creator: "openai", inventoryRef: "§1 OpenAI" },
308
- "deepseek-v4-flash": { canonicalId: "deepseek-flash", family: "deepseek", creator: "deepseek", status: "retired-legacy-id", inventoryRef: "§2 DeepSeek" },
309
- "deepseek-chat": { canonicalId: null, family: "deepseek", creator: "deepseek", status: "retired", inventoryRef: "§2 DeepSeek" },
310
- "deepseek-reasoner": { canonicalId: null, family: "deepseek", creator: "deepseek", status: "retired", inventoryRef: "§2 DeepSeek" },
382
+ "gpt-5.6": { canonicalId: "gpt-5.6-sol", family: "gpt-5.6", creator: "openai", legacy: true, inventoryRef: "§1 OpenAI" },
383
+ "gpt-daybreak-blue-latest": { canonicalId: "gpt-5.6-sol", family: "gpt-5.6", creator: "openai", legacy: false, inventoryRef: "§1 OpenAI" },
384
+ "gpt-daybreak-red-latest": { canonicalId: "gpt-5.6-cyber", family: "gpt-5.6", creator: "openai", legacy: false, inventoryRef: "§1 OpenAI" },
385
+ "deepseek-v4-flash": { canonicalId: "deepseek-flash", family: "deepseek", creator: "deepseek", legacy: true, status: "retired-legacy-id", inventoryRef: "§2 DeepSeek" },
386
+ "deepseek-chat": { canonicalId: null, family: "deepseek", creator: "deepseek", legacy: true, status: "retired", inventoryRef: "§2 DeepSeek" },
387
+ "deepseek-reasoner": { canonicalId: null, family: "deepseek", creator: "deepseek", legacy: true, status: "retired", inventoryRef: "§2 DeepSeek" },
311
388
  }
312
389
 
313
390
  // ---------------------------------------------------------------------------
@@ -320,6 +397,7 @@ function neutralResult(reason, transport) {
320
397
  family: "neutral",
321
398
  baseline: BASELINES["neutral.none"],
322
399
  overlays: [],
400
+ runtime: NEUTRAL_RUNTIME,
323
401
  transport,
324
402
  matchType: "neutral",
325
403
  matchReason: reason,
@@ -333,12 +411,25 @@ function overlaysFor(ids) {
333
411
  return (ids ?? []).map((id) => OVERLAYS[id]).filter(Boolean)
334
412
  }
335
413
 
414
+ // A family entry matches by regex `pattern` or by a pure `predicate(slug)`.
415
+ function familyMatches(entry, slug) {
416
+ if (typeof entry.predicate === "function") return entry.predicate(slug)
417
+ return entry.pattern ? entry.pattern.test(slug) : false
418
+ }
419
+
420
+ // Only legacy entries carry runtime capabilities. A non-legacy entry (gpt-6,
421
+ // Pro UltraSpeed) resolves for information but stays neutral at runtime.
422
+ function runtimeForEntry(entry) {
423
+ return entry.legacy === false ? NEUTRAL_RUNTIME : entry.runtime ?? NEUTRAL_RUNTIME
424
+ }
425
+
336
426
  function resultFromEntry(entry, matchType, matchReason, matchedId, transport) {
337
427
  return {
338
428
  creator: entry.creator,
339
429
  family: entry.family,
340
430
  baseline: BASELINES[entry.baseline] ?? null,
341
431
  overlays: overlaysFor(entry.overlays),
432
+ runtime: runtimeForEntry(entry),
342
433
  transport,
343
434
  matchType,
344
435
  matchReason,
@@ -348,13 +439,17 @@ function resultFromEntry(entry, matchType, matchReason, matchedId, transport) {
348
439
  }
349
440
  }
350
441
 
351
- function resultFromFamily(family, creator, matchType, matchReason, matchedId, inventoryRef, note, transport) {
442
+ function resultFromFamily(family, creator, matchType, matchReason, matchedId, inventoryRef, note, transport, aliasLegacy) {
352
443
  const entry = POLICY_REGISTRY.find((e) => e.family === family && e.kind !== "exact")
444
+ // An alias is runtime-active only when both the alias and its target family
445
+ // were recognized before v0.4.0.
446
+ const active = aliasLegacy !== false && (!entry || entry.legacy !== false)
353
447
  return {
354
448
  creator,
355
449
  family,
356
450
  baseline: entry ? BASELINES[entry.baseline] ?? null : null,
357
451
  overlays: entry ? overlaysFor(entry.overlays) : [],
452
+ runtime: active && entry ? entry.runtime ?? NEUTRAL_RUNTIME : NEUTRAL_RUNTIME,
358
453
  transport,
359
454
  matchType,
360
455
  matchReason,
@@ -386,20 +481,23 @@ export function resolvePolicy(model) {
386
481
  const familyEntry = POLICY_REGISTRY.find((e) => e.family === alias.family && e.kind !== "exact")
387
482
  if (familyEntry?.requiresOpenAIish && !isOpenAIish(s)) continue
388
483
  const reason = `alias:${id}->${alias.canonicalId ?? alias.family}`
389
- return resultFromFamily(alias.family, alias.creator, "exact", reason, id, alias.inventoryRef, alias.status ?? null, transport)
484
+ return resultFromFamily(alias.family, alias.creator, "exact", reason, id, alias.inventoryRef, alias.status ?? null, transport, alias.legacy)
390
485
  }
391
486
 
392
- // 2. Exact model ids (documented models).
487
+ // 2. Exact model ids (documented models). The entry's context gate still
488
+ // applies, so an exact OpenAI id on a non-OpenAI endpoint is never guessed.
393
489
  for (const entry of POLICY_REGISTRY) {
394
490
  if (!entry.exactIds || entry.exactIds.length === 0) continue
395
491
  const hit = ids.find((id) => entry.exactIds.includes(id))
396
- if (hit) return resultFromEntry(entry, "exact", `exact-id:${hit}`, hit, transport)
492
+ if (!hit) continue
493
+ if (entry.requiresOpenAIish && !isOpenAIish(s)) continue
494
+ return resultFromEntry(entry, "exact", `exact-id:${hit}`, hit, transport)
397
495
  }
398
496
 
399
- // 3. Model family / range patterns.
497
+ // 3. Model family / range matchers (regex pattern or version predicate).
400
498
  for (const entry of POLICY_REGISTRY) {
401
- if (entry.kind !== "family" || !entry.pattern) continue
402
- if (!entry.pattern.test(s.slug)) continue
499
+ if (entry.kind !== "family") continue
500
+ if (!familyMatches(entry, s.slug)) continue
403
501
  if (entry.requiresOpenAIish && !isOpenAIish(s)) continue
404
502
  return resultFromEntry(entry, "family", `family-pattern:${entry.id}`, null, transport)
405
503
  }
@@ -415,6 +513,13 @@ export function resolvePolicy(model) {
415
513
  return neutralResult("neutral:no-match", transport)
416
514
  }
417
515
 
516
+ // Convenience accessor for the runtime: the legacy policy string + explicit
517
+ // capability flags. This is the single source the runtime gates on; it is
518
+ // guaranteed equal to the pre-v0.4.0 detectPolicy() classification.
519
+ export function resolveRuntimePolicy(model) {
520
+ return resolvePolicy(model).runtime
521
+ }
522
+
418
523
  // Compatibility classification used by detectPolicy(). Reproduces the
419
524
  // pre-v0.4.0 behavior exactly: it considers only `legacy` registry entries,
420
525
  // excludes newer-generation/alias/exact-overlay additions, and returns a family
@@ -426,7 +531,7 @@ export function resolveLegacyFamily(model) {
426
531
  for (const entry of POLICY_REGISTRY) {
427
532
  if (!entry.legacy) continue
428
533
  if (entry.kind === "family") {
429
- if (!entry.pattern.test(s.slug)) continue
534
+ if (!familyMatches(entry, s.slug)) continue
430
535
  if (entry.requiresOpenAIish && !isOpenAIish(s)) continue
431
536
  return entry.family
432
537
  }
@@ -53,8 +53,10 @@ import {
53
53
  MODEL_ALIASES,
54
54
  OVERLAYS,
55
55
  POLICY_REGISTRY,
56
+ isGpt56OrLater,
56
57
  resolveLegacyFamily,
57
58
  resolvePolicy,
59
+ resolveRuntimePolicy,
58
60
  } from "../src/cache-policy-core.mjs"
59
61
 
60
62
  const asst = (id, read, write) => ({
@@ -1358,26 +1360,27 @@ test("resolvePolicy: GPT-5.6 exact + inventory aliases resolve to the gpt-5.6 fa
1358
1360
  assert.equal(orVariant.matchType, "exact")
1359
1361
  })
1360
1362
 
1361
- test("resolvePolicy: GPT-6 resolves via explicit documented inheritance, with no overlay", () => {
1363
+ test("v0.4.2: GPT-6 resolves through the documented GPT-5.6-and-later boundary", () => {
1362
1364
  const r = resolvePolicy(M("openrouter", "openai/gpt-6-luna"))
1363
1365
  assert.equal(r.creator, "openai")
1364
- assert.equal(r.family, "gpt-6")
1366
+ assert.equal(r.family, "gpt-5.6")
1365
1367
  assert.equal(baseId(r), "openai.gpt56.cache")
1366
- assert.deepEqual(overlayIds(r), [])
1367
- assert.equal(resolvePolicy(M("openai", "gpt-6-astra")).family, "gpt-6")
1368
+ // GPT-6 is documented in the same cache regime, so it gets the same overlay.
1369
+ assert.deepEqual(overlayIds(r), ["gpt56.prompt-cache-options"])
1370
+ assert.equal(resolvePolicy(M("openai", "gpt-6-astra")).family, "gpt-5.6")
1368
1371
 
1369
- // Inheritance is explicit, not "newer means same": the registry records it.
1370
- const gpt6 = POLICY_REGISTRY.find((e) => e.family === "gpt-6")
1371
- assert.equal(gpt6.inheritsFrom, "gpt-5.6")
1372
- assert.ok(gpt6.inventoryRef)
1373
- // And no overlay is implied by that inheritance.
1374
- assert.deepEqual(gpt6.overlays, [])
1372
+ // The boundary is version-based, not an exact-model list.
1373
+ const entry = POLICY_REGISTRY.find((e) => e.id === "openai.gpt-5.6-plus")
1374
+ assert.equal(entry.boundary, "GPT-5.6 and later")
1375
+ assert.equal(typeof entry.predicate, "function")
1376
+ assert.ok(entry.inventoryRef)
1375
1377
  })
1376
1378
 
1377
- test("resolvePolicy: GPT-6 is not classified by the legacy detectPolicy wrapper", () => {
1378
- // Intentional: the new layer knows gpt-6, the compatibility wrapper does not.
1379
- assert.equal(detectPolicy(M("openai", "gpt-6-astra")), POLICY_NEUTRAL)
1380
- assert.equal(resolvePolicy(M("openai", "gpt-6-astra")).family, "gpt-6")
1379
+ test("v0.4.2: the legacy detectPolicy wrapper follows the same boundary", () => {
1380
+ assert.equal(detectPolicy(M("openai", "gpt-6-astra")), POLICY_GPT56)
1381
+ assert.equal(detectPolicy(M("openai", "gpt-5.6")), POLICY_GPT56)
1382
+ assert.equal(detectPolicy(M("openai", "gpt-5.5")), POLICY_NEUTRAL)
1383
+ assert.equal(detectPolicy(M("openai", "gpt-5.60")), POLICY_NEUTRAL)
1381
1384
  })
1382
1385
 
1383
1386
  test("resolvePolicy: pre-5.6 GPT negative controls are neutral with no overlays", () => {
@@ -1551,3 +1554,268 @@ test("registry is traceable and internally consistent", () => {
1551
1554
  assert.ok(alias.family)
1552
1555
  }
1553
1556
  })
1557
+
1558
+ // ===========================================================================
1559
+ // v0.4.1 runtime policy migration (behavior preservation)
1560
+ //
1561
+ // The runtime now classifies via resolveRuntimePolicy(). These tests prove the
1562
+ // resolved policy equals the legacy detectPolicy() string for every supported
1563
+ // model, and that the hook-observable behavior (GPT cache options, <env>
1564
+ // relocation, OpenRouter affinity) is unchanged. Newer/unknown models must
1565
+ // gain no mutation.
1566
+ // ===========================================================================
1567
+
1568
+ const policyCoreURL = new URL("../src/cache-policy-core.mjs", import.meta.url).href
1569
+
1570
+ test("v0.4.1: resolveRuntimePolicy.policy matches detectPolicy across a broad matrix", () => {
1571
+ const samples = [
1572
+ M("openai", "gpt-5.6"),
1573
+ M("openai", "gpt-5.6-sol"),
1574
+ M("openrouter", "openai/gpt-5.6-luna"),
1575
+ M("openai-compatible", "gpt-5.6"),
1576
+ M("openai", "gpt-6-astra"),
1577
+ M("openai", "gpt-5.5"),
1578
+ M("openai", "gpt-daybreak-blue-latest"),
1579
+ M("openai", "gpt-daybreak-red-latest"),
1580
+ M("zai", "glm-5.3"),
1581
+ M("zai", "glm-5.3-flash"),
1582
+ M("zai", "glm-5.2"),
1583
+ M("xiaomi", "mimo-v2.6-flash"),
1584
+ M("xiaomi", "mimo-v2.6-pro"),
1585
+ M("xiaomi", "mimo-v2.6-pro-ultraspeed"),
1586
+ M("xiaomi", "mimo-v2.5"),
1587
+ M("deepseek", "deepseek-v4-pro"),
1588
+ M("deepseek", "deepseek-flash"),
1589
+ M("deepseek", "deepseek-v4-flash"),
1590
+ M("deepseek", "deepseek-chat"),
1591
+ M("openrouter", "x-ai/grok-4"),
1592
+ {},
1593
+ null,
1594
+ undefined,
1595
+ 42,
1596
+ ]
1597
+ for (const m of samples) {
1598
+ assert.equal(resolveRuntimePolicy(m).policy, detectPolicy(m))
1599
+ }
1600
+ })
1601
+
1602
+ test("v0.4.1: GPT-5.6 alias keeps the overlay but a newer alias stays runtime-neutral", () => {
1603
+ // Legacy alias: gpt-5.6 -> gpt-5.6-sol (was matched by the old regex).
1604
+ const legacy = resolveRuntimePolicy(M("openai", "gpt-5.6"))
1605
+ assert.equal(legacy.policy, "gpt56")
1606
+ assert.equal(legacy.gptCacheMetadata, true)
1607
+ // Newer documented alias that the old classifier did NOT match: no mutation.
1608
+ const newer = resolveRuntimePolicy(M("openai", "gpt-daybreak-blue-latest"))
1609
+ assert.equal(newer.policy, "neutral")
1610
+ assert.equal(newer.gptCacheMetadata, false)
1611
+ assert.equal(resolvePolicy(M("openai", "gpt-daybreak-blue-latest")).family, "gpt-5.6")
1612
+ })
1613
+
1614
+ async function runPolicyMigrationProbe() {
1615
+ const home = mkdtempSync(join(tmpdir(), "ce-policy-migration-"))
1616
+ const pluginURL = new URL("../src/cache-engine.ts", import.meta.url).href
1617
+ const coreURL = new URL("../src/cache-engine-core.mjs", import.meta.url).href
1618
+ const script = `
1619
+ import assert from "node:assert/strict"
1620
+ process.env.CACHE_ENGINE_METRICS_FILE = process.env.HOME + "/policy-migration.jsonl"
1621
+ const { CacheEngine } = await import(${JSON.stringify(pluginURL)})
1622
+ const { detectPolicy } = await import(${JSON.stringify(coreURL)})
1623
+ const { resolvePolicy, resolveRuntimePolicy } = await import(${JSON.stringify(policyCoreURL)})
1624
+ const client = {
1625
+ app: { log: async () => ({}) },
1626
+ session: { get: async () => ({ data: { parentID: null } }) },
1627
+ tool: { list: async () => ({ data: [] }) },
1628
+ }
1629
+ const hooks = await CacheEngine({ client, directory: process.env.HOME })
1630
+ assert.equal(typeof hooks["chat.params"], "function")
1631
+ assert.equal(typeof hooks["chat.headers"], "function")
1632
+ assert.equal(typeof hooks["experimental.chat.system.transform"], "function")
1633
+ const SYS = ["A: keep1", "B: You are powered by the model named x. The exact model ID is acme/x", "C: <env>", "D: Today's date: 2026-08-17", "E: </env>", "F: keep2"].join("\\n")
1634
+ const CASES = [
1635
+ { name: "gpt-5.6", model: { providerID: "openai", id: "gpt-5.6", api: { id: "gpt-5.6", npm: "@ai-sdk/openai" } }, expect: { policy: "gpt56", env: false, gpt: true, header: false } },
1636
+ { name: "gpt-5.6-openrouter", model: { providerID: "openrouter", id: "openai/gpt-5.6-sol", api: { id: "openai/gpt-5.6-sol" } }, expect: { policy: "gpt56", env: false, gpt: true, header: false } },
1637
+ { name: "gpt-6-astra", model: { providerID: "openai", id: "gpt-6-astra", api: { id: "gpt-6-astra", npm: "@ai-sdk/openai" } }, expect: { policy: "gpt56", env: false, gpt: true, header: false } },
1638
+ { name: "gpt-6-openrouter", model: { providerID: "openrouter", id: "openai/gpt-6-luna", api: { id: "openai/gpt-6-luna" } }, expect: { policy: "gpt56", env: false, gpt: true, header: false } },
1639
+ { name: "gpt-6-openai-compatible", model: { providerID: "openai-compatible", id: "gpt-6-astra", api: { id: "gpt-6-astra" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
1640
+ { name: "gpt-5.60-malformed", model: { providerID: "openai", id: "gpt-5.60", api: { id: "gpt-5.60", npm: "@ai-sdk/openai" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
1641
+ { name: "gpt-5.5", model: { providerID: "openai", id: "gpt-5.5", api: { id: "gpt-5.5", npm: "@ai-sdk/openai" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
1642
+ { name: "gpt-daybreak-alias", model: { providerID: "openai", id: "gpt-daybreak-blue-latest", api: { id: "gpt-daybreak-blue-latest", npm: "@ai-sdk/openai" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
1643
+ { name: "deepseek-v4-pro", model: { providerID: "deepseek", id: "deepseek-v4-pro", api: { id: "deepseek-v4-pro" } }, expect: { policy: "deepseek", env: false, gpt: false, header: false } },
1644
+ { name: "deepseek-flash", model: { providerID: "deepseek", id: "deepseek-flash", api: { id: "deepseek-flash" } }, expect: { policy: "deepseek", env: false, gpt: false, header: false } },
1645
+ { name: "glm-5.3-direct", model: { providerID: "zai", id: "glm-5.3", api: { id: "glm-5.3" } }, expect: { policy: "glm53", env: true, gpt: false, header: false } },
1646
+ { name: "glm-5.3-openrouter", model: { providerID: "openrouter", id: "z-ai/glm-5.3-flash", api: { id: "z-ai/glm-5.3-flash" } }, expect: { policy: "glm53", env: true, gpt: false, header: true } },
1647
+ { name: "glm-5.2", model: { providerID: "zai", id: "glm-5.2", api: { id: "glm-5.2" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
1648
+ { name: "mimo-v2.6-flash-direct", model: { providerID: "xiaomi", id: "mimo-v2.6-flash", api: { id: "mimo-v2.6-flash" } }, expect: { policy: "mimo26", env: true, gpt: false, header: false } },
1649
+ { name: "mimo-v2.6-pro-openrouter", model: { providerID: "openrouter", id: "xiaomi/mimo-v2.6-pro", api: { id: "xiaomi/mimo-v2.6-pro" } }, expect: { policy: "mimo26", env: true, gpt: false, header: true } },
1650
+ { name: "mimo-v2.6-pro-ultraspeed", model: { providerID: "xiaomi", id: "mimo-v2.6-pro-ultraspeed", api: { id: "mimo-v2.6-pro-ultraspeed" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
1651
+ { name: "mimo-v2.5", model: { providerID: "xiaomi", id: "mimo-v2.5", api: { id: "mimo-v2.5" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
1652
+ { name: "unknown-provider", model: { providerID: "mystery-provider", id: "xiaomi/mimo-v2.6-flash", api: { id: "xiaomi/mimo-v2.6-flash" } }, expect: { policy: "mimo26", env: true, gpt: false, header: false } },
1653
+ { name: "unknown-openrouter", model: { providerID: "openrouter", id: "acme/mystery-9", api: { id: "acme/mystery-9" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
1654
+ ]
1655
+ const results = []
1656
+ for (const c of CASES) {
1657
+ const model = c.model
1658
+ const sid = "ses_" + c.name
1659
+ const provider = { source: "config", info: { id: String(model.providerID ?? "") }, options: {} }
1660
+ const existing = { "User-Agent": "preserve", "x-custom": "preserve" }
1661
+ const paramsOut = { options: {} }
1662
+ const headersOut = { headers: { ...existing } }
1663
+ const sysOut = { system: [SYS] }
1664
+ await hooks["chat.params"]({ sessionID: sid, agent: "build", model, provider, message: { id: "msg-" + c.name, sessionID: sid, role: "user", content: "probe" } }, paramsOut)
1665
+ await hooks["experimental.chat.system.transform"]({ sessionID: sid, model, provider }, sysOut)
1666
+ await hooks["chat.headers"]({ sessionID: sid, agent: "build", model, provider, message: { id: "msg-" + c.name, sessionID: sid, role: "user", content: "probe" } }, headersOut)
1667
+ const rt = resolveRuntimePolicy(model)
1668
+ const existingHeadersPreserved = Object.entries(existing).every(([k, v]) => headersOut.headers[k] === v)
1669
+ results.push({
1670
+ name: c.name,
1671
+ runtimePolicy: rt.policy,
1672
+ detectPolicy: detectPolicy(model),
1673
+ richFamily: resolvePolicy(model).family,
1674
+ overlays: resolvePolicy(model).overlays.map((o) => o.id),
1675
+ systemRelocated: sysOut.system[0] !== SYS,
1676
+ gptOptionInjected: paramsOut.options.promptCacheKey !== undefined,
1677
+ gptOptions: paramsOut.options.promptCacheOptions ?? null,
1678
+ affinityHeaderAttached: headersOut.headers["x-session-id"] !== undefined,
1679
+ existingHeadersPreserved,
1680
+ expect: c.expect,
1681
+ })
1682
+ }
1683
+ process.stdout.write(JSON.stringify({ results }))
1684
+ `
1685
+ const stdout = execFileSync(process.execPath, ["--experimental-strip-types", "--input-type=module", "-e", script], {
1686
+ cwd: process.cwd(),
1687
+ env: { ...process.env, HOME: home },
1688
+ encoding: "utf8",
1689
+ })
1690
+ return JSON.parse(stdout.trim())
1691
+ }
1692
+
1693
+ let policyMigrationProbe
1694
+ const policyMigrationResults = async () => (policyMigrationProbe ??= runPolicyMigrationProbe())
1695
+
1696
+ test("v0.4.1: runtime policy is the resolver's and stays equal to detectPolicy (hook path)", async () => {
1697
+ const { results } = await policyMigrationResults()
1698
+ assert.ok(results.length >= 15)
1699
+ for (const r of results) {
1700
+ assert.equal(r.runtimePolicy, r.detectPolicy, `${r.name}: resolver policy must equal legacy detectPolicy`)
1701
+ assert.equal(r.runtimePolicy, r.expect.policy, `${r.name}: unexpected policy`)
1702
+ }
1703
+ })
1704
+
1705
+ test("v0.4.1: every supported model keeps its pre-migration hook behavior", async () => {
1706
+ const { results } = await policyMigrationResults()
1707
+ for (const r of results) {
1708
+ assert.equal(r.systemRelocated, r.expect.env, `${r.name}: <env> relocation`)
1709
+ assert.equal(r.gptOptionInjected, r.expect.gpt, `${r.name}: GPT cache-options injection`)
1710
+ assert.equal(r.affinityHeaderAttached, r.expect.header, `${r.name}: OpenRouter affinity header`)
1711
+ // No provider in the matrix mutates or drops pre-existing headers.
1712
+ assert.equal(r.existingHeadersPreserved, true, `${r.name}: existing headers preserved`)
1713
+ }
1714
+ })
1715
+
1716
+ test("v0.4.1: GPT-5.6 keeps promptCacheOptions implicit/30m through the resolver", async () => {
1717
+ const { results } = await policyMigrationResults()
1718
+ const gpt = results.find((r) => r.name === "gpt-5.6")
1719
+ assert.deepEqual(gpt.gptOptions, { mode: "implicit", ttl: "30m" })
1720
+ })
1721
+
1722
+ test("v0.4.1: future-looking and unknown models gain no mutation", async () => {
1723
+ const { results } = await policyMigrationResults()
1724
+ const noMutation = [
1725
+ "gpt-daybreak-alias",
1726
+ "gpt-5.5",
1727
+ "gpt-6-openai-compatible",
1728
+ "gpt-5.60-malformed",
1729
+ "glm-5.2",
1730
+ "mimo-v2.6-pro-ultraspeed",
1731
+ "mimo-v2.5",
1732
+ "unknown-openrouter",
1733
+ ]
1734
+ for (const name of noMutation) {
1735
+ const r = results.find((x) => x.name === name)
1736
+ assert.ok(r, `${name} present`)
1737
+ assert.equal(r.gptOptionInjected, false, `${name}: no GPT options`)
1738
+ assert.equal(r.systemRelocated, false, `${name}: no <env> relocation`)
1739
+ assert.equal(r.affinityHeaderAttached, false, `${name}: no affinity header`)
1740
+ }
1741
+ })
1742
+
1743
+ test("v0.4.1: non-OpenRouter models never receive the OpenRouter header", async () => {
1744
+ const { results } = await policyMigrationResults()
1745
+ for (const name of ["glm-5.3-direct", "mimo-v2.6-flash-direct", "unknown-provider", "gpt-5.6", "deepseek-v4-pro"]) {
1746
+ const r = results.find((x) => x.name === name)
1747
+ assert.equal(r.affinityHeaderAttached, false, `${name}: no affinity header off OpenRouter`)
1748
+ }
1749
+ })
1750
+
1751
+ // ===========================================================================
1752
+ // v0.4.2 GPT-5.6-and-later boundary
1753
+ //
1754
+ // Source: OpenAI *Prompt caching* guide, "GPT-5.6 and later" generation
1755
+ // boundary, re-verified 2026-09-27 (docs/cache-policy-inventory.md §1). GPT-6
1756
+ // is documented in the same regime with no cache-control exception.
1757
+ // ===========================================================================
1758
+
1759
+ test("v0.4.2: isGpt56OrLater matches the documented boundary by version", () => {
1760
+ const inFamily = [
1761
+ "gpt-5.6",
1762
+ "gpt-5.6-luna",
1763
+ "openai/gpt-5.6-sol",
1764
+ "gpt-6",
1765
+ "gpt-6-astra",
1766
+ "gpt-6-sol",
1767
+ "gpt-6-luna",
1768
+ "gpt-5.7",
1769
+ "gpt-7",
1770
+ "gpt-6.1",
1771
+ ]
1772
+ for (const id of inFamily) assert.equal(isGpt56OrLater(id), true, `${id} is in family`)
1773
+ const outOfFamily = [
1774
+ "gpt-5.5",
1775
+ "gpt-5.4",
1776
+ "gpt-5.2",
1777
+ "gpt-5.1",
1778
+ "gpt-5",
1779
+ "gpt-4.1",
1780
+ "gpt-4o",
1781
+ "gpt-5.60",
1782
+ "gpt-5.6.1",
1783
+ "gpt-4",
1784
+ "claude-sonnet-4-5",
1785
+ ]
1786
+ for (const id of outOfFamily) assert.equal(isGpt56OrLater(id), false, `${id} is out of family`)
1787
+ })
1788
+
1789
+ test("v0.4.2: GPT boundary respects OpenAI/provider gating", () => {
1790
+ assert.equal(resolveRuntimePolicy(M("openai", "gpt-5.6")).gptCacheMetadata, true)
1791
+ assert.equal(resolveRuntimePolicy(M("openai", "gpt-6-astra")).gptCacheMetadata, true)
1792
+ assert.equal(resolveRuntimePolicy(M("azure", "gpt-6-sol")).gptCacheMetadata, true)
1793
+ assert.equal(resolveRuntimePolicy(M("openrouter", "openai/gpt-6-luna")).gptCacheMetadata, true)
1794
+ assert.equal(resolveRuntimePolicy(M("openai-compatible", "gpt-6-astra")).gptCacheMetadata, false)
1795
+ assert.equal(resolveRuntimePolicy(M("llama.cpp", "gpt-5.6")).gptCacheMetadata, false)
1796
+ assert.equal(resolveRuntimePolicy(M("openai", "gpt-5.5")).gptCacheMetadata, false)
1797
+ assert.equal(resolveRuntimePolicy(M("openai", "gpt-5.60")).gptCacheMetadata, false)
1798
+ })
1799
+
1800
+ test("v0.4.2: covered later GPT requests receive the same documented baseline at runtime", async () => {
1801
+ const { results } = await policyMigrationResults()
1802
+ const gpt6 = results.find((r) => r.name === "gpt-6-astra")
1803
+ assert.equal(gpt6.runtimePolicy, "gpt56")
1804
+ assert.equal(gpt6.gptOptionInjected, true)
1805
+ assert.deepEqual(gpt6.gptOptions, { mode: "implicit", ttl: "30m" })
1806
+ // Same baseline as GPT-5.6, no new mechanism.
1807
+ const gpt56 = results.find((r) => r.name === "gpt-5.6")
1808
+ assert.deepEqual(gpt6.gptOptions, gpt56.gptOptions)
1809
+ // No prompt transformation or affinity was introduced for gpt-6.
1810
+ assert.equal(gpt6.systemRelocated, false)
1811
+ assert.equal(gpt6.affinityHeaderAttached, false)
1812
+ })
1813
+
1814
+ test("v0.4.2: pre-5.6 and out-of-family GPT ids get no GPT options", async () => {
1815
+ const { results } = await policyMigrationResults()
1816
+ for (const name of ["gpt-5.5", "gpt-5.60-malformed", "gpt-6-openai-compatible"]) {
1817
+ const r = results.find((x) => x.name === name)
1818
+ assert.equal(r.gptOptionInjected, false, `${name}: no GPT options`)
1819
+ assert.equal(r.runtimePolicy, "neutral", `${name}: neutral runtime`)
1820
+ }
1821
+ })