opencode-cache-engine 0.4.0 → 0.4.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +67 -11
- package/docs/cache-policy-inventory.md +58 -7
- package/package.json +1 -1
- package/src/cache-engine.ts +60 -49
- package/src/cache-policy-core.mjs +139 -34
- package/test/cache-engine.test.mjs +282 -14
package/README.md
CHANGED
|
@@ -18,15 +18,30 @@ plugin under `~/.config/opencode/plugins/`. For released installs where
|
|
|
18
18
|
reproducibility matters, pin an exact package version rather than relying on
|
|
19
19
|
`@latest` resolution or a moving cache entry; see [Installation](#installation).
|
|
20
20
|
|
|
21
|
+
Quick Installation (TUI):
|
|
22
|
+
``` text
|
|
23
|
+
opencode plugin opencode-cache-engine
|
|
24
|
+
```
|
|
25
|
+
|
|
21
26
|
`CacheEngine` is an OpenCode plugin designed for long-running agent sessions where prompt-cache efficiency affects both latency and cost. It keeps the harness conservative for providers whose cache behavior is already automatic, while applying provider-specific optimizations where the provider exposes useful cache controls or where prompt structure can be safely improved.
|
|
22
27
|
|
|
23
28
|
The plugin currently has four cache-policy families:
|
|
24
29
|
|
|
25
30
|
* **DeepSeek** — passive cache observability; request structure is preserved.
|
|
26
|
-
* **GPT-5.6** — documented cache-key/options metadata, with prompt text
|
|
31
|
+
* **GPT-5.6 and later** — documented cache-key/options metadata, with prompt text
|
|
32
|
+
unchanged. GPT-6 and future 5.6+/6+/7+ versions resolve through the same
|
|
33
|
+
documented boundary.
|
|
27
34
|
* **GLM-5.3** — narrow, content-preserving `<env>` relocation and diagnostics.
|
|
28
35
|
* **MiMo-V2.6** — narrow, content-preserving `<env>` relocation and diagnostics.
|
|
29
36
|
|
|
37
|
+
Family classification is not hard-coded in the runtime. A pure policy registry
|
|
38
|
+
and resolver in `src/cache-policy-core.mjs` returns a structured result
|
|
39
|
+
(`creator`, `family`, `baseline`, `overlays`, `transport`, `matchType`,
|
|
40
|
+
`matchReason`), and the hooks gate their behavior on that result. The registry is
|
|
41
|
+
the single runtime source of policy classification. The first-party research
|
|
42
|
+
behind each registry entry is recorded in
|
|
43
|
+
[docs/cache-policy-inventory.md](docs/cache-policy-inventory.md).
|
|
44
|
+
|
|
30
45
|
For both MiMo-V2.6 and GLM-5.3, CacheEngine adds its deterministic
|
|
31
46
|
`x-session-id` request header only when OpenCode identifies the actual provider
|
|
32
47
|
as `openrouter`. It does not add that OpenRouter-specific header for
|
|
@@ -44,12 +59,12 @@ The plugin operates at the OpenCode harness level rather than implementing a pro
|
|
|
44
59
|
|
|
45
60
|
It:
|
|
46
61
|
|
|
47
|
-
1.
|
|
48
|
-
2. Applies only the
|
|
62
|
+
1. Resolves the model/provider policy through the registry resolver.
|
|
63
|
+
2. Applies only the mutations registered for that policy.
|
|
49
64
|
3. Observes system-prompt and tool-definition stability.
|
|
50
65
|
4. Records provider-reported cache token usage.
|
|
51
66
|
5. Adds a deterministic compaction continuation block.
|
|
52
|
-
6. Applies GPT-5.6 cache-control metadata.
|
|
67
|
+
6. Applies GPT-5.6-and-later cache-control metadata.
|
|
53
68
|
7. Applies the GLM-5.3 and MiMo-V2.6 volatile-environment relocation.
|
|
54
69
|
8. Records diagnostics that help determine whether prompt-shape changes correlate with cache behavior.
|
|
55
70
|
9. Records MiMo/GLM affinity outcomes and provider-identity changes.
|
|
@@ -115,7 +130,7 @@ DeepSeek:
|
|
|
115
130
|
|
|
116
131
|
### Policy: active cache control
|
|
117
132
|
|
|
118
|
-
GPT-5.6 is the only
|
|
133
|
+
GPT-5.6 and later is the only policy family that actively injects cache-control request metadata.
|
|
119
134
|
|
|
120
135
|
The plugin adds:
|
|
121
136
|
|
|
@@ -405,7 +420,7 @@ and are reported diagnostically; the message content is left untouched.
|
|
|
405
420
|
| Policy family | Detection | Prompt text changed? | Cache metadata changed? | OpenRouter affinity header | Primary cache signal |
|
|
406
421
|
| ------------- | --------- | ------------------- | ----------------------- | -------------------------- | -------------------- |
|
|
407
422
|
| DeepSeek | `deepseek` | No | No | None | provider `cache.read` / `cache.write` |
|
|
408
|
-
| GPT-5.6 | `gpt
|
|
423
|
+
| GPT-5.6 and later | version boundary `gpt-<major>[.<minor>] ≥ 5.6` on OpenAI-ish endpoints (includes GPT-6) | No | Yes: `prompt_cache_key` + options | None | provider cache tokens |
|
|
409
424
|
| GLM-5.3 | `glm-5.3*` | Yes, narrowly (`<env>` tail) | No provider cache key | `x-session-id` on OpenRouter only | provider cache tokens (GLM ratio) |
|
|
410
425
|
| MiMo-V2.6 | Flash / Pro only | Yes, narrowly (`<env>` tail) | No: implicit caching only | `x-session-id` on OpenRouter only | `cached_tokens / prompt_tokens` |
|
|
411
426
|
|
|
@@ -923,7 +938,7 @@ The model detector recognizes:
|
|
|
923
938
|
* GLM-5.3 variants
|
|
924
939
|
* MiMo-V2.6 Flash and Pro (`xiaomi/mimo-v2.6-flash`, `mimo-v2.6-pro`, ...)
|
|
925
940
|
|
|
926
|
-
GPT-5.6 has an additional OpenAI/Azure-context check so a string containing `gpt-5.6` does not automatically cause GPT-specific fields to be sent to an unrelated endpoint.
|
|
941
|
+
The GPT-5.6-and-later family has an additional OpenAI/Azure-context check, so a string containing a qualifying GPT version (for example `gpt-5.6` or `gpt-6`) does not automatically cause GPT-specific fields to be sent to an unrelated endpoint.
|
|
927
942
|
|
|
928
943
|
MiMo detection targets exactly Flash and Pro: it excludes `mimo-v2.5`,
|
|
929
944
|
`mimo-v2.5-pro`, and `mimo-v2.6-pro-ultraspeed`.
|
|
@@ -953,7 +968,8 @@ For that reason, a stable provider route is preferable when your goal is to meas
|
|
|
953
968
|
|
|
954
969
|
# Architecture
|
|
955
970
|
|
|
956
|
-
The implementation is split
|
|
971
|
+
The implementation is split across a hook entry point, a pure logic core, and a
|
|
972
|
+
pure policy registry.
|
|
957
973
|
|
|
958
974
|
## `cache-engine.ts`
|
|
959
975
|
|
|
@@ -989,7 +1005,7 @@ This contains dependency-light pure logic.
|
|
|
989
1005
|
|
|
990
1006
|
It owns:
|
|
991
1007
|
|
|
992
|
-
*
|
|
1008
|
+
* the legacy `detectPolicy()` compatibility wrapper (delegating to the registry)
|
|
993
1009
|
* configuration parsing
|
|
994
1010
|
* hashing
|
|
995
1011
|
* canonicalization
|
|
@@ -1006,6 +1022,42 @@ Keeping these functions in plain JavaScript allows the logic to be tested indepe
|
|
|
1006
1022
|
|
|
1007
1023
|
---
|
|
1008
1024
|
|
|
1025
|
+
## `cache-policy-core.mjs`
|
|
1026
|
+
|
|
1027
|
+
This is the pure policy registry and resolver. It separates cache policy from
|
|
1028
|
+
request mutation:
|
|
1029
|
+
|
|
1030
|
+
* creator / family classification
|
|
1031
|
+
* baseline cache-policy descriptors (documented facts)
|
|
1032
|
+
* model-specific overlays (for example GLM/MiMo `<env>` relocation)
|
|
1033
|
+
* transport capabilities (for example OpenRouter `x-session-id` affinity)
|
|
1034
|
+
* explicit, inventory-traceable inheritance (`inheritsFrom`)
|
|
1035
|
+
* safe neutral fallback for unknown or future models
|
|
1036
|
+
|
|
1037
|
+
`resolvePolicy(model)` returns `creator`, `family`, `baseline`, `overlays`,
|
|
1038
|
+
`transport`, `matchType`, and `matchReason`. `resolveRuntimePolicy(model)`
|
|
1039
|
+
returns the runtime-facing descriptor the hooks consume: the legacy policy
|
|
1040
|
+
string plus explicit capability flags.
|
|
1041
|
+
|
|
1042
|
+
Only registry entries marked `legacy` enable runtime behavior; documented but
|
|
1043
|
+
non-legacy entries (for example `mimo-v2.6-pro-ultraspeed`) and all unknown
|
|
1044
|
+
models resolve to a neutral runtime. A newer or unknown model therefore never
|
|
1045
|
+
inherits a current model's mutation unless the registry explicitly registers it.
|
|
1046
|
+
The GPT family is a documented exception in the sense that its boundary is
|
|
1047
|
+
version-based (`GPT-5.6 and later`), so GPT-6 and future 5.6+/6+/7+ versions are
|
|
1048
|
+
covered by the registered boundary rather than by an exact-model list.
|
|
1049
|
+
|
|
1050
|
+
Transport is kept separate from cache policy: OpenRouter affinity is a transport
|
|
1051
|
+
capability, not part of a creator's cache semantics. Overlays are also explicit,
|
|
1052
|
+
so being classified into a family does not by itself enable a prompt
|
|
1053
|
+
transformation.
|
|
1054
|
+
|
|
1055
|
+
The module is pure: no network calls and no runtime documentation lookups. The
|
|
1056
|
+
legacy `detectPolicy()` in `cache-engine-core.mjs` remains a thin compatibility
|
|
1057
|
+
wrapper over the resolver's legacy path.
|
|
1058
|
+
|
|
1059
|
+
---
|
|
1060
|
+
|
|
1009
1061
|
## Tests
|
|
1010
1062
|
|
|
1011
1063
|
The repository's test suite validates the provider-independent and provider-specific logic.
|
|
@@ -1013,6 +1065,7 @@ The repository's test suite validates the provider-independent and provider-spec
|
|
|
1013
1065
|
Coverage includes:
|
|
1014
1066
|
|
|
1015
1067
|
* model detection
|
|
1068
|
+
* policy registry resolution and runtime-policy equivalence
|
|
1016
1069
|
* GPT cache-key stability
|
|
1017
1070
|
* GPT cache-option defaults
|
|
1018
1071
|
* protection against overwriting existing cache options
|
|
@@ -1176,11 +1229,14 @@ opencode-cache-engine/
|
|
|
1176
1229
|
├── src/
|
|
1177
1230
|
│ ├── cache-engine.ts
|
|
1178
1231
|
│ ├── cache-engine-core.mjs
|
|
1232
|
+
│ ├── cache-policy-core.mjs
|
|
1179
1233
|
│ └── tui.mjs
|
|
1180
1234
|
├── test/
|
|
1181
1235
|
│ └── cache-engine.test.mjs
|
|
1182
1236
|
├── examples/
|
|
1183
1237
|
│ └── cache-engine.json
|
|
1238
|
+
├── docs/
|
|
1239
|
+
│ └── cache-policy-inventory.md
|
|
1184
1240
|
├── package.json
|
|
1185
1241
|
├── README.md
|
|
1186
1242
|
└── LICENSE
|
|
@@ -1219,7 +1275,7 @@ release, use:
|
|
|
1219
1275
|
```json
|
|
1220
1276
|
{
|
|
1221
1277
|
"plugin": [
|
|
1222
|
-
"opencode-cache-engine@0.
|
|
1278
|
+
"opencode-cache-engine@0.4.1"
|
|
1223
1279
|
]
|
|
1224
1280
|
}
|
|
1225
1281
|
```
|
|
@@ -1286,7 +1342,7 @@ A prefix change is a diagnostic signal, not automatic proof of a cache miss.
|
|
|
1286
1342
|
|
|
1287
1343
|
## GPT-5.6 cache options are missing
|
|
1288
1344
|
|
|
1289
|
-
Verify that the model is
|
|
1345
|
+
Verify that the model is within the documented GPT-5.6-and-later boundary (for example `gpt-5.6-*` or `gpt-6-*`) and that the endpoint is recognized as OpenAI/Azure-compatible.
|
|
1290
1346
|
|
|
1291
1347
|
The detector intentionally rejects ambiguous OpenAI-compatible providers rather than guessing.
|
|
1292
1348
|
|
|
@@ -32,22 +32,47 @@ Caching behavior is never inferred from pricing alone, from one SDK's type
|
|
|
32
32
|
declarations, or from a third-party blog when first-party documentation exists.
|
|
33
33
|
"unknown — first-party docs insufficient" is used instead of a guess.
|
|
34
34
|
|
|
35
|
+
## Runtime integration
|
|
36
|
+
|
|
37
|
+
This inventory is the research input for the policy registry in
|
|
38
|
+
`src/cache-policy-core.mjs`. As of **v0.4.1** the runtime hooks consume
|
|
39
|
+
`resolveRuntimePolicy(model)` and gate behavior on the registry's explicit
|
|
40
|
+
capabilities, so the "CacheEngine current treatment" column below describes
|
|
41
|
+
resolver-driven behavior.
|
|
42
|
+
|
|
43
|
+
Two guarantees follow from that migration:
|
|
44
|
+
|
|
45
|
+
- `resolveRuntimePolicy(model).policy` equals the legacy `detectPolicy(model)`
|
|
46
|
+
string, so telemetry and gating are unchanged for every model supported in
|
|
47
|
+
v0.3.6.
|
|
48
|
+
- Registry entries marked non-`legacy` (for example
|
|
49
|
+
`mimo-v2.6-pro-ultraspeed`) and all unknown models resolve to a neutral
|
|
50
|
+
runtime, so no documented-but-unwired model gains a current model's mutation.
|
|
51
|
+
|
|
52
|
+
The runtime reads the registry at classification time only; there are no network
|
|
53
|
+
calls and no runtime documentation lookups.
|
|
54
|
+
|
|
35
55
|
## CacheEngine current treatment (baseline for the matrix)
|
|
36
56
|
|
|
37
|
-
Source: `src/cache-
|
|
38
|
-
`src/cache-engine.ts`
|
|
57
|
+
Source: the pure registry/resolver (`src/cache-policy-core.mjs`) and the hook
|
|
58
|
+
entry (`src/cache-engine.ts`), as wired in v0.4.1. The behavior described here is
|
|
59
|
+
the same as at the pre-resolver revision `19b87f2`; only the classification
|
|
60
|
+
source changed.
|
|
39
61
|
|
|
40
62
|
| Family | Detection (verbatim) | Current treatment | Affinity header |
|
|
41
63
|
| --- | --- | --- | --- |
|
|
42
64
|
| DeepSeek | `/deepseek/i` on `${apiID} ${modelID}` or `providerID` | Passive; no mutation | none |
|
|
43
|
-
| GPT-5.6 |
|
|
65
|
+
| GPT-5.6 | version boundary `gpt-<major>[.<minor>] ≥ 5.6` on slug **and** `isOpenAIish` (provider `openai`/`azure`, slug `openai/`/`azure/`, or npm `@ai-sdk/openai`/`@ai-sdk/azure`); since v0.4.2 covers GPT-6 and later | Inject missing `promptCacheKey` + `promptCacheOptions` (`implicit`, `30m`) | none |
|
|
44
66
|
| GLM-5.3 | `/glm-5\.3(?![\d.])/i` on slug | Relocate identifiable `<env>` block to system tail | `x-session-id` only when `providerID === "openrouter"` |
|
|
45
67
|
| MiMo-V2.6 | `/mimo-v2\.6-(flash\|pro)(?![\w-])/i` on slug | Relocate identifiable `<env>` block to system tail; provider-change telemetry | `x-session-id` only when `providerID === "openrouter"` |
|
|
46
68
|
| Neutral | everything else | Byte-untouched | none |
|
|
47
69
|
|
|
48
70
|
Detection consequences worth stating explicitly:
|
|
49
71
|
|
|
50
|
-
- `gpt-6
|
|
72
|
+
- `gpt-6` / `gpt-6-*` (and any future 5.6+/6+/7+ version) is matched by the
|
|
73
|
+
documented GPT-5.6-and-later boundary → GPT policy. **[O]** (v0.4.2)
|
|
74
|
+
- `gpt-5.5`, `gpt-5.2`, `gpt-4o`, and the malformed `gpt-5.60` are **not**
|
|
75
|
+
matched → neutral. **[O]**
|
|
51
76
|
- `deepseek-v5` (or any future `*deepseek*` id) matches the passive DeepSeek
|
|
52
77
|
branch because the regex is a bare substring test. **[O]**
|
|
53
78
|
- `mimo-v2.6-pro-ultraspeed` is **not** matched: the `(?![\w-])` lookahead fails
|
|
@@ -95,8 +120,10 @@ catch-all "earlier models". [D]
|
|
|
95
120
|
baseline: it injects only `promptCacheKey` + `promptCacheOptions{mode:implicit,
|
|
96
121
|
ttl:"30m"}`, preserves runtime-supplied values, and never sets context/output
|
|
97
122
|
limits. It does not use explicit breakpoints or `prewarm`, which is a subset of
|
|
98
|
-
the documented capability.
|
|
99
|
-
|
|
123
|
+
the documented capability. Since v0.4.2 the policy resolves by the documented
|
|
124
|
+
"GPT-5.6 and later" boundary, so GPT-6 (astra/sol/luna) receives the same
|
|
125
|
+
baseline; no GPT-6 cache-control exception is documented (OpenAI *Prompt
|
|
126
|
+
caching* guide, re-verified 2026-09-27), and none is coded.
|
|
100
127
|
|
|
101
128
|
---
|
|
102
129
|
|
|
@@ -254,7 +281,7 @@ made here); **hold** = do not inherit without first-party evidence.
|
|
|
254
281
|
| Creator | Model / example pattern | Cache policy (documented) | CacheEngine current treatment | Recommended family inheritance | Recommended exact-model exception | Confidence | Source | Verified |
|
|
255
282
|
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
|
|
256
283
|
| OpenAI | `gpt-5.6`, `gpt-5.6-*` (sol/terra/luna/cyber) | "GPT-5.6 and later": implicit default, optional explicit breakpoints, min 1,024, TTL 30m, write 1.25×/read 0.1× | GPT policy: inject `promptCacheKey` + `promptCacheOptions{implicit,30m}` | keep | none documented; CacheEngine's subset is valid | High (docs) / Medium (treatment) | OpenAI *Prompt caching*; *GPT-5.6 Sol* | 2026-09-26 |
|
|
257
|
-
| OpenAI | GPT-6 / current later GPT family: `gpt-6`, `gpt-6-*` (astra/sol/luna) | Inherits the GPT-5.6-and-later policy | **
|
|
284
|
+
| OpenAI | GPT-6 / current later GPT family: `gpt-6`, `gpt-6-*` (astra/sol/luna) | Inherits the GPT-5.6-and-later policy | **GPT policy** via the GPT-5.6-and-later boundary (v0.4.2) | **keep** — covered by the boundary predicate; no exact-model entry needed | none documented | High (docs) / High (treatment) | OpenAI *Using GPT-6*; *GPT-6 Astra*; *Prompt caching* | 2026-09-27 |
|
|
258
285
|
| OpenAI | pre-5.6 negative controls: `gpt-5.5`, `gpt-5.4`, `gpt-5.2`, `gpt-5.1`, `gpt-5`, `gpt-4.1`, `gpt-4o` | Implicit only; different min-length class; `in_memory`/`24h` retention; `prompt_cache_key` for routing | neutral | **hold** — do not inherit 5.6 policy | n/a | High | OpenAI *Prompt caching*; *Pricing* | 2026-09-26 |
|
|
259
286
|
| DeepSeek | V4: `deepseek-v4-pro`, legacy `deepseek-v4-flash` | Provider-wide automatic disk cache; implicit; prefix-unit matching; hit/miss token fields | Passive (no mutation) | keep | none | High | DeepSeek *Context Caching*; *Models & Pricing* | 2026-09-26 |
|
|
260
287
|
| DeepSeek | V4.1 / current V4-family: `deepseek-flash` (MODEL VERSION "DeepSeek-V4.1-Flash") | Same provider-wide automatic policy; cache-hit pricing listed for both current models | Passive | keep (creator/family baseline) | none documented | High | DeepSeek *Models & Pricing*; *news260910* | 2026-09-26 |
|
|
@@ -311,3 +338,27 @@ a newer model inherits an older policy. They must not be resolved by guessing.
|
|
|
311
338
|
- The existing test suite was run only to confirm the repository remains green
|
|
312
339
|
(116/116).
|
|
313
340
|
- This release contains research and documentation only.
|
|
341
|
+
|
|
342
|
+
### Follow-up: v0.4.1 runtime integration
|
|
343
|
+
|
|
344
|
+
- v0.4.0 added the registry (`src/cache-policy-core.mjs`) without wiring it.
|
|
345
|
+
- v0.4.1 wired the runtime to `resolveRuntimePolicy()` and preserved behavior:
|
|
346
|
+
every model supported in v0.3.6 keeps its prior treatment, and non-legacy or
|
|
347
|
+
unknown models remain neutral.
|
|
348
|
+
- The v0.4.0 research statements above are unchanged; only the runtime now reads
|
|
349
|
+
this registry as its single source of policy classification.
|
|
350
|
+
|
|
351
|
+
### Follow-up: v0.4.2 GPT-5.6-and-later boundary
|
|
352
|
+
|
|
353
|
+
- OpenAI's documented "GPT-5.6 and later" boundary was re-verified against the
|
|
354
|
+
first-party *Prompt caching* guide and *Using GPT-6* guide on **2026-09-27**.
|
|
355
|
+
GPT-6 (astra/sol/luna) is documented in the same cache regime with no
|
|
356
|
+
cache-control exception.
|
|
357
|
+
- v0.4.2 replaces the exact `gpt-5.6` string match with a version-boundary
|
|
358
|
+
predicate (`gpt-<major>[.<minor>] ≥ 5.6`), still gated on the OpenAI-ish
|
|
359
|
+
provider check. GPT-6 and future 5.6+/6+/7+ models therefore need no
|
|
360
|
+
exact-model registry entry, while `gpt-5.5` and earlier and malformed ids such
|
|
361
|
+
as `gpt-5.60` stay neutral.
|
|
362
|
+
- The only OpenAI cache controls injected remain `promptCacheKey` +
|
|
363
|
+
`promptCacheOptions{implicit,30m}`; no breakpoint or prewarm behavior was
|
|
364
|
+
added. DeepSeek, GLM, MiMo, and OpenRouter affinity behavior are unchanged.
|
package/package.json
CHANGED
package/src/cache-engine.ts
CHANGED
|
@@ -4,12 +4,7 @@ import {
|
|
|
4
4
|
DEFAULT_CONFIG_PATH,
|
|
5
5
|
DIGEST_TEMPLATE,
|
|
6
6
|
affinityTelemetryFields,
|
|
7
|
-
POLICY_GLM53,
|
|
8
|
-
POLICY_GPT56,
|
|
9
|
-
POLICY_MIMO26,
|
|
10
|
-
POLICY_NEUTRAL,
|
|
11
7
|
createRecorder,
|
|
12
|
-
detectPolicy,
|
|
13
8
|
detectReasoningIssues,
|
|
14
9
|
digestDecision,
|
|
15
10
|
ensureMetricsDir,
|
|
@@ -37,6 +32,7 @@ import {
|
|
|
37
32
|
toolFingerprint,
|
|
38
33
|
toolWireFingerprint,
|
|
39
34
|
} from "./cache-engine-core.mjs"
|
|
35
|
+
import { resolveRuntimePolicy } from "./cache-policy-core.mjs"
|
|
40
36
|
|
|
41
37
|
// ---------------------------------------------------------------------------
|
|
42
38
|
// cache-engine
|
|
@@ -101,7 +97,21 @@ type Shape = {
|
|
|
101
97
|
toolCount: number | null
|
|
102
98
|
}
|
|
103
99
|
|
|
104
|
-
|
|
100
|
+
// Resolved runtime policy (the registry's runtime descriptor). `policy` is the
|
|
101
|
+
// legacy telemetry/state string emitted by the pre-v0.4.0 classifier.
|
|
102
|
+
type PolicyRuntime = {
|
|
103
|
+
policy: string
|
|
104
|
+
isNeutral: boolean
|
|
105
|
+
gptCacheMetadata: boolean
|
|
106
|
+
envRelocation: "glm" | "mimo" | null
|
|
107
|
+
thinkingIntegrity: boolean
|
|
108
|
+
cacheRatio: "glm" | "mimo" | null
|
|
109
|
+
providerChange: "glm" | "mimo" | null
|
|
110
|
+
prefixDiagnostics: boolean
|
|
111
|
+
openRouterAffinity: boolean
|
|
112
|
+
}
|
|
113
|
+
|
|
114
|
+
type ModelInfo = { family: string; providerID: string; modelID: string; caps: PolicyRuntime }
|
|
105
115
|
|
|
106
116
|
type ToolCache = { semanticToolsHash: string | null; wireToolsHash: string | null }
|
|
107
117
|
|
|
@@ -258,19 +268,21 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
258
268
|
// gpt56/glm53/deepseek classification once established.
|
|
259
269
|
const rememberModel = (sid: string, model: ChatParamsModel | undefined): ModelInfo | null => {
|
|
260
270
|
if (!model) return null
|
|
261
|
-
|
|
271
|
+
// Single runtime source of policy classification: the registry resolver.
|
|
272
|
+
const caps = resolveRuntimePolicy(model) as PolicyRuntime
|
|
262
273
|
const s = get(sid)
|
|
263
274
|
const info: ModelInfo = {
|
|
264
|
-
family,
|
|
275
|
+
family: caps.policy,
|
|
265
276
|
providerID: String(model.providerID ?? ""),
|
|
266
277
|
modelID: String(model.api?.id ?? model.id ?? ""),
|
|
278
|
+
caps,
|
|
267
279
|
}
|
|
268
|
-
if (s.modelInfo == null || (s.modelInfo.
|
|
280
|
+
if (s.modelInfo == null || (s.modelInfo.caps.isNeutral && !caps.isNeutral)) {
|
|
269
281
|
s.modelInfo = info
|
|
270
282
|
// A title/summary request (neutral small model) may have established the
|
|
271
283
|
// system baseline first. Its "powered by the model named ..." env line
|
|
272
284
|
// differs from the real model's, so re-baseline on upgrade.
|
|
273
|
-
if (s.baselineSystem !== null &&
|
|
285
|
+
if (s.baselineSystem !== null && !caps.isNeutral) {
|
|
274
286
|
s.baselineSystem = null
|
|
275
287
|
s.shape = null
|
|
276
288
|
}
|
|
@@ -350,11 +362,11 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
350
362
|
|
|
351
363
|
s.lastProcessedMessageID = nextProcessedCursor(firstPage, startCursor)
|
|
352
364
|
|
|
353
|
-
const
|
|
365
|
+
const caps = s.modelInfo?.caps
|
|
354
366
|
const glmIntegrity =
|
|
355
|
-
|
|
356
|
-
policyEnabled(cfg,
|
|
357
|
-
cfg.policies?.[
|
|
367
|
+
caps?.thinkingIntegrity === true &&
|
|
368
|
+
policyEnabled(cfg, caps.policy) &&
|
|
369
|
+
cfg.policies?.[caps.policy]?.preserveThinkingIntegrity === true
|
|
358
370
|
|
|
359
371
|
if (glmIntegrity && reasoningNewestFirst.length > 0) {
|
|
360
372
|
// messages arrive newest-first; process oldest->newest so `seen` grows
|
|
@@ -417,11 +429,11 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
417
429
|
recFields.model = s.modelInfo.modelID
|
|
418
430
|
recFields.policy = s.modelInfo.family
|
|
419
431
|
}
|
|
420
|
-
if (
|
|
432
|
+
if (caps?.cacheRatio === "glm") {
|
|
421
433
|
recFields.promptTokens = read + write + input
|
|
422
434
|
recFields.glmHitRate = glmHitRatio(read, write, input)
|
|
423
435
|
}
|
|
424
|
-
if (
|
|
436
|
+
if (caps?.cacheRatio === "mimo") {
|
|
425
437
|
// Provider-reported cached tokens / total prompt tokens. The runtime's
|
|
426
438
|
// `input` is the non-cached prompt input and `cache.read` is the
|
|
427
439
|
// cached prompt input, so total prompt tokens are derived as read +
|
|
@@ -431,7 +443,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
431
443
|
recFields.promptTokens = promptTokens
|
|
432
444
|
recFields.cachedTokens = read
|
|
433
445
|
recFields.cacheHitRate = mimoHitRate(read, promptTokens)
|
|
434
|
-
if (cfg.policies?.[
|
|
446
|
+
if (cfg.policies?.[caps.policy]?.stickySession === true) {
|
|
435
447
|
recFields.stickySessionId = mimoSessionIdFor(sid)
|
|
436
448
|
}
|
|
437
449
|
// Prefer the latest live provider identity over the latched one.
|
|
@@ -440,10 +452,10 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
440
452
|
recFields.model = s.mimoProvider.modelID
|
|
441
453
|
}
|
|
442
454
|
}
|
|
443
|
-
if (
|
|
455
|
+
if (caps?.gptCacheMetadata === true && s.gptInjected) {
|
|
444
456
|
recFields.keyStrategy = "session"
|
|
445
|
-
recFields.mode = cfg.policies?.[
|
|
446
|
-
recFields.ttl = cfg.policies?.[
|
|
457
|
+
recFields.mode = cfg.policies?.[caps.policy]?.mode
|
|
458
|
+
recFields.ttl = cfg.policies?.[caps.policy]?.ttl
|
|
447
459
|
}
|
|
448
460
|
rec.record(recFields)
|
|
449
461
|
} else {
|
|
@@ -495,8 +507,9 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
495
507
|
try {
|
|
496
508
|
const model = input.model as unknown as ChatParamsModel
|
|
497
509
|
const providerID = String(model?.providerID ?? "")
|
|
498
|
-
const
|
|
499
|
-
if (
|
|
510
|
+
const caps = resolveRuntimePolicy(model) as PolicyRuntime
|
|
511
|
+
if (!caps.openRouterAffinity) return
|
|
512
|
+
const family = caps.policy
|
|
500
513
|
|
|
501
514
|
const hasSessionIDHeader = (headers?: Record<string, string>) =>
|
|
502
515
|
Object.keys(headers ?? {}).some((name) => name.toLowerCase() === "x-session-id")
|
|
@@ -536,6 +549,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
536
549
|
try {
|
|
537
550
|
const info = rememberModel(input.sessionID, input.model as unknown as ChatParamsModel)
|
|
538
551
|
const family = info?.family
|
|
552
|
+
const caps = info?.caps
|
|
539
553
|
|
|
540
554
|
// ---- MiMo-V2.6: provider-switch diagnostics (telemetry only) ---------
|
|
541
555
|
// MiMo cache lives at the provider side, so a provider change within one
|
|
@@ -545,7 +559,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
545
559
|
// NOTE: OpenRouter's *upstream* provider selection (e.g. xiaomi/fp8) is
|
|
546
560
|
// not exposed to plugins; only the OpenCode providerID/modelID are
|
|
547
561
|
// observable here.
|
|
548
|
-
if (
|
|
562
|
+
if (caps?.providerChange === "mimo" && policyEnabled(cfg, caps.policy) && info) {
|
|
549
563
|
// Use the LIVE model identity (not the latched one) so a provider
|
|
550
564
|
// switch within the session is actually observable.
|
|
551
565
|
const live = input.model as unknown as ChatParamsModel
|
|
@@ -558,7 +572,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
558
572
|
const ev = providerChangeEvent(s.mimoProvider, cur)
|
|
559
573
|
if (ev.changed) {
|
|
560
574
|
const sticky =
|
|
561
|
-
cfg.policies?.[
|
|
575
|
+
cfg.policies?.[caps.policy]?.stickySession === true
|
|
562
576
|
? { stickySessionId: mimoSessionIdFor(input.sessionID) }
|
|
563
577
|
: {}
|
|
564
578
|
rec.record({
|
|
@@ -566,7 +580,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
566
580
|
sid: input.sessionID,
|
|
567
581
|
ts: Date.now(),
|
|
568
582
|
reason: "mimo_provider_changed",
|
|
569
|
-
policy:
|
|
583
|
+
policy: caps.policy,
|
|
570
584
|
from: ev.from,
|
|
571
585
|
to: ev.to,
|
|
572
586
|
...sticky,
|
|
@@ -579,7 +593,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
579
593
|
}
|
|
580
594
|
|
|
581
595
|
// ---- GLM-5.3 provider identity observation (telemetry only) ----------
|
|
582
|
-
if (
|
|
596
|
+
if (caps?.providerChange === "glm" && policyEnabled(cfg, caps.policy) && info) {
|
|
583
597
|
const live = input.model as unknown as ChatParamsModel
|
|
584
598
|
const cur = {
|
|
585
599
|
providerID: String(live?.providerID ?? ""),
|
|
@@ -594,7 +608,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
594
608
|
sid: input.sessionID,
|
|
595
609
|
ts: Date.now(),
|
|
596
610
|
reason: "glm_provider_changed",
|
|
597
|
-
policy:
|
|
611
|
+
policy: caps.policy,
|
|
598
612
|
from: ev.from,
|
|
599
613
|
to: ev.to,
|
|
600
614
|
note: "OpenCode providerID changed; provider-specific upstream routing is not plugin-visible",
|
|
@@ -605,12 +619,12 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
605
619
|
return
|
|
606
620
|
}
|
|
607
621
|
|
|
608
|
-
if (!(
|
|
622
|
+
if (!(caps?.gptCacheMetadata === true && policyEnabled(cfg, caps.policy))) {
|
|
609
623
|
// DeepSeek / GLM / neutral: nothing to inject. GLM has no cache-key API;
|
|
610
624
|
// DeepSeek caching is fully passive; we never mutate requests for them.
|
|
611
625
|
return
|
|
612
626
|
}
|
|
613
|
-
const gpol = cfg.policies?.[
|
|
627
|
+
const gpol = cfg.policies?.[caps.policy]
|
|
614
628
|
const applyRoot = gpol?.cacheRootKey !== false
|
|
615
629
|
const compaction = input.agent === "compaction"
|
|
616
630
|
const isolated = compaction && gpol?.compactionCacheIsolation === true
|
|
@@ -678,7 +692,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
678
692
|
kind: "cache-options",
|
|
679
693
|
sid: input.sessionID,
|
|
680
694
|
ts: Date.now(),
|
|
681
|
-
policy:
|
|
695
|
+
policy: caps.policy,
|
|
682
696
|
provider: info.providerID,
|
|
683
697
|
model: info.modelID,
|
|
684
698
|
keyStrategy: applyRoot ? "cache-root" : "session",
|
|
@@ -784,43 +798,40 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
784
798
|
const model = input.model as unknown as ChatParamsModel
|
|
785
799
|
rememberModel(sid, model)
|
|
786
800
|
const s = get(sid)
|
|
787
|
-
const
|
|
801
|
+
const caps = s.modelInfo?.caps
|
|
788
802
|
|
|
789
803
|
// ---- GLM-5.3 / MiMo-V2.6 input-shape stabilization ------------------
|
|
790
804
|
// Relocate the identifiable volatile env block (per-day date) to the
|
|
791
805
|
// tail of the single system string, content-preserving, ONLY when the
|
|
792
|
-
//
|
|
793
|
-
// present exactly. Never touches other content/order; never
|
|
794
|
-
// other families. The runtime passes a single-element system
|
|
795
|
-
// (verified against the installed runtime), so no generic
|
|
796
|
-
// involved.
|
|
806
|
+
// resolved runtime policy opts into system stabilization and the block
|
|
807
|
+
// markers are present exactly. Never touches other content/order; never
|
|
808
|
+
// applied to other families. The runtime passes a single-element system
|
|
809
|
+
// array (verified against the installed runtime), so no generic
|
|
810
|
+
// reordering is involved.
|
|
797
811
|
//
|
|
798
812
|
// In-place mutation note: request.ts keeps using its own local `system`
|
|
799
813
|
// array after the hook (the trigger's returned output is ignored), so
|
|
800
814
|
// reassigning `output.system = [...]` would be lost. We rewrite the
|
|
801
815
|
// single element in place instead.
|
|
802
816
|
let systemText = output.system.join("\n")
|
|
803
|
-
const
|
|
804
|
-
|
|
805
|
-
|
|
806
|
-
cfg
|
|
807
|
-
|
|
808
|
-
|
|
809
|
-
policyEnabled(cfg, POLICY_MIMO26) &&
|
|
810
|
-
cfg.policies?.[POLICY_MIMO26]?.stabilizeSystem === true
|
|
811
|
-
if ((glmStabilize || mimoStabilize) && output.system.length === 1) {
|
|
817
|
+
const envVariant = caps?.envRelocation ?? null
|
|
818
|
+
const stabilize =
|
|
819
|
+
envVariant !== null &&
|
|
820
|
+
policyEnabled(cfg, caps!.policy) &&
|
|
821
|
+
cfg.policies?.[caps!.policy]?.stabilizeSystem === true
|
|
822
|
+
if (stabilize && output.system.length === 1) {
|
|
812
823
|
const rel = relocateVolatileEnvBlock(output.system[0])
|
|
813
824
|
if (rel.changed) {
|
|
814
825
|
output.system[0] = rel.text
|
|
815
826
|
systemText = rel.text
|
|
816
|
-
if (
|
|
827
|
+
if (envVariant === "mimo") {
|
|
817
828
|
log("debug", "mimo system env block relocated to suffix", { sid })
|
|
818
829
|
rec.record({
|
|
819
830
|
kind: "boundary",
|
|
820
831
|
sid,
|
|
821
832
|
ts: Date.now(),
|
|
822
833
|
reason: "mimo_system_env_relocated",
|
|
823
|
-
policy:
|
|
834
|
+
policy: caps!.policy,
|
|
824
835
|
provider: s.modelInfo?.providerID,
|
|
825
836
|
model: s.modelInfo?.modelID,
|
|
826
837
|
})
|
|
@@ -906,13 +917,13 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
906
917
|
// MiMo-specific explicit diagnostic: the STABLE prefix changed (not
|
|
907
918
|
// just the relocated volatile env suffix). Reported only; the new
|
|
908
919
|
// content is never overwritten with a stale snapshot.
|
|
909
|
-
if (
|
|
920
|
+
if (caps?.prefixDiagnostics === true && reasons.includes("system_stable_prefix_changed")) {
|
|
910
921
|
rec.record({
|
|
911
922
|
kind: "boundary",
|
|
912
923
|
sid,
|
|
913
924
|
ts: Date.now(),
|
|
914
925
|
reason: "mimo_system_prefix_changed",
|
|
915
|
-
policy:
|
|
926
|
+
policy: caps!.policy,
|
|
916
927
|
provider: s.modelInfo?.providerID,
|
|
917
928
|
model: s.modelInfo?.modelID,
|
|
918
929
|
changedFields: granular,
|
|
@@ -17,8 +17,10 @@
|
|
|
17
17
|
// `inventoryRef`. Inheritance is always explicit (`inheritsFrom`); "newer means
|
|
18
18
|
// same behavior" is never an unconditional rule.
|
|
19
19
|
//
|
|
20
|
-
//
|
|
21
|
-
//
|
|
20
|
+
// As of v0.4.1 the runtime hook layer (cache-engine.ts) consumes
|
|
21
|
+
// resolveRuntimePolicy() as its single source of policy classification.
|
|
22
|
+
// detectPolicy() is retained as the compatibility classifier for the legacy
|
|
23
|
+
// POLICY_* strings.
|
|
22
24
|
|
|
23
25
|
// ---------------------------------------------------------------------------
|
|
24
26
|
// Model normalization (shared with the legacy classifier)
|
|
@@ -54,6 +56,31 @@ export function isOpenAIish(s) {
|
|
|
54
56
|
return false
|
|
55
57
|
}
|
|
56
58
|
|
|
59
|
+
// The documented OpenAI cache-policy boundary is the generation phrase
|
|
60
|
+
// "GPT-5.6 and later" (docs/cache-policy-inventory.md §1; OpenAI *Prompt
|
|
61
|
+
// caching* guide, re-verified 2026-09-27). This matcher expresses that boundary
|
|
62
|
+
// by version rather than by an exact-model string, so future 5.6+/6+/7+ models
|
|
63
|
+
// need no registry entry:
|
|
64
|
+
// - major > 5 -> in family
|
|
65
|
+
// - major === 5 && minor >= 6 -> in family
|
|
66
|
+
// - everything else -> out
|
|
67
|
+
// The token must be followed by a non-digit/non-dot boundary, so malformed ids
|
|
68
|
+
// such as "gpt-5.60" and "gpt-5.6.1" do not match (same guard as pre-v0.4.2).
|
|
69
|
+
// OpenAI minor versions are single-digit, so a multi-digit minor is treated as
|
|
70
|
+
// malformed rather than as a higher version.
|
|
71
|
+
export function isGpt56OrLater(slug) {
|
|
72
|
+
const text = String(slug ?? "").toLowerCase()
|
|
73
|
+
const re = /gpt-(\d{1,3})(?:\.(\d))?(?![\d.])/g
|
|
74
|
+
let m
|
|
75
|
+
while ((m = re.exec(text)) !== null) {
|
|
76
|
+
const major = Number(m[1])
|
|
77
|
+
const minor = m[2] === undefined ? 0 : Number(m[2])
|
|
78
|
+
if (major > 5) return true
|
|
79
|
+
if (major === 5 && minor >= 6) return true
|
|
80
|
+
}
|
|
81
|
+
return false
|
|
82
|
+
}
|
|
83
|
+
|
|
57
84
|
// Candidate ids for exact/alias lookup. Includes the raw apiID/modelID, the
|
|
58
85
|
// lower-cased forms, and a single stripped transport/vendor prefix
|
|
59
86
|
// (e.g. "openai/gpt-5.6-luna" -> "gpt-5.6-luna", "xiaomi/mimo-v2.6-flash" ->
|
|
@@ -206,6 +233,36 @@ function resolveTransport(s) {
|
|
|
206
233
|
return { id: p, kind: "direct", sessionAffinityHeader: null, stickyRouting: false, inventoryRef: "§5 OpenRouter transport" }
|
|
207
234
|
}
|
|
208
235
|
|
|
236
|
+
// ---------------------------------------------------------------------------
|
|
237
|
+
// Runtime capability descriptors
|
|
238
|
+
//
|
|
239
|
+
// The runtime consumes `resolvePolicy(...).runtime` for gating. `policy` is the
|
|
240
|
+
// legacy telemetry/state string, so telemetry stays byte-identical. Every
|
|
241
|
+
// capability is explicit per registry entry: classification into a creator or
|
|
242
|
+
// family never implies a mutation. `legacy: false` entries always resolve to
|
|
243
|
+
// NEUTRAL_RUNTIME, so a future-looking model gains nothing until the registry
|
|
244
|
+
// explicitly says so.
|
|
245
|
+
// ---------------------------------------------------------------------------
|
|
246
|
+
|
|
247
|
+
const NEUTRAL_RUNTIME = Object.freeze({
|
|
248
|
+
policy: "neutral",
|
|
249
|
+
isNeutral: true,
|
|
250
|
+
gptCacheMetadata: false,
|
|
251
|
+
envRelocation: null,
|
|
252
|
+
thinkingIntegrity: false,
|
|
253
|
+
cacheRatio: null,
|
|
254
|
+
providerChange: null,
|
|
255
|
+
prefixDiagnostics: false,
|
|
256
|
+
openRouterAffinity: false,
|
|
257
|
+
})
|
|
258
|
+
|
|
259
|
+
const rt = (policy, overrides = {}) => ({
|
|
260
|
+
...NEUTRAL_RUNTIME,
|
|
261
|
+
...overrides,
|
|
262
|
+
policy,
|
|
263
|
+
isNeutral: policy === "neutral",
|
|
264
|
+
})
|
|
265
|
+
|
|
209
266
|
// ---------------------------------------------------------------------------
|
|
210
267
|
// Registry
|
|
211
268
|
//
|
|
@@ -219,31 +276,32 @@ function resolveTransport(s) {
|
|
|
219
276
|
|
|
220
277
|
export const POLICY_REGISTRY = [
|
|
221
278
|
{
|
|
222
|
-
|
|
279
|
+
// v0.4.2: one documented GPT-5.6-and-later family, matched by the version
|
|
280
|
+
// boundary rather than an exact model string. GPT-6 (astra/sol/luna) is
|
|
281
|
+
// documented in the same regime with no cache-control exception, so it
|
|
282
|
+
// inherits this baseline and overlay. Future 5.6+/6+/7+ models resolve here
|
|
283
|
+
// without a new registry entry.
|
|
284
|
+
id: "openai.gpt-5.6-plus",
|
|
223
285
|
creator: "openai",
|
|
224
286
|
family: "gpt-5.6",
|
|
225
287
|
kind: "family",
|
|
226
|
-
|
|
288
|
+
predicate: isGpt56OrLater,
|
|
227
289
|
requiresOpenAIish: true,
|
|
228
|
-
exactIds: [
|
|
290
|
+
exactIds: [
|
|
291
|
+
"gpt-5.6-sol",
|
|
292
|
+
"gpt-5.6-terra",
|
|
293
|
+
"gpt-5.6-luna",
|
|
294
|
+
"gpt-5.6-cyber",
|
|
295
|
+
"gpt-6-astra",
|
|
296
|
+
"gpt-6-sol",
|
|
297
|
+
"gpt-6-luna",
|
|
298
|
+
],
|
|
229
299
|
baseline: "openai.gpt56.cache",
|
|
230
300
|
overlays: ["gpt56.prompt-cache-options"],
|
|
231
301
|
legacy: true,
|
|
232
|
-
|
|
233
|
-
|
|
234
|
-
|
|
235
|
-
id: "openai.gpt-6",
|
|
236
|
-
creator: "openai",
|
|
237
|
-
family: "gpt-6",
|
|
238
|
-
kind: "family",
|
|
239
|
-
pattern: /gpt-6(?![\d.])/i,
|
|
240
|
-
requiresOpenAIish: true,
|
|
241
|
-
exactIds: ["gpt-6-astra", "gpt-6-sol", "gpt-6-luna"],
|
|
242
|
-
baseline: "openai.gpt56.cache",
|
|
243
|
-
inheritsFrom: "gpt-5.6",
|
|
244
|
-
overlays: [],
|
|
245
|
-
legacy: false,
|
|
246
|
-
note: "Documented inheritance of the GPT-5.6-and-later baseline. No CacheEngine overlay is registered for gpt-6 yet.",
|
|
302
|
+
runtime: rt("gpt56", { gptCacheMetadata: true }),
|
|
303
|
+
boundary: "GPT-5.6 and later",
|
|
304
|
+
note: "Documented boundary 'GPT-5.6 and later' (OpenAI Prompt caching guide, re-verified 2026-09-27) includes GPT-6 with no documented cache-control exception. No explicit breakpoint or prewarm behavior is registered.",
|
|
247
305
|
inventoryRef: "§1 OpenAI",
|
|
248
306
|
},
|
|
249
307
|
{
|
|
@@ -256,6 +314,13 @@ export const POLICY_REGISTRY = [
|
|
|
256
314
|
baseline: "zai.implicit-cache",
|
|
257
315
|
overlays: ["glm53.env-relocation"],
|
|
258
316
|
legacy: true,
|
|
317
|
+
runtime: rt("glm53", {
|
|
318
|
+
envRelocation: "glm",
|
|
319
|
+
thinkingIntegrity: true,
|
|
320
|
+
cacheRatio: "glm",
|
|
321
|
+
providerChange: "glm",
|
|
322
|
+
openRouterAffinity: true,
|
|
323
|
+
}),
|
|
259
324
|
inventoryRef: "§3 Z.AI GLM",
|
|
260
325
|
},
|
|
261
326
|
{
|
|
@@ -268,6 +333,13 @@ export const POLICY_REGISTRY = [
|
|
|
268
333
|
baseline: "xiaomi.implicit-cache",
|
|
269
334
|
overlays: ["mimo26.env-relocation"],
|
|
270
335
|
legacy: true,
|
|
336
|
+
runtime: rt("mimo26", {
|
|
337
|
+
envRelocation: "mimo",
|
|
338
|
+
cacheRatio: "mimo",
|
|
339
|
+
providerChange: "mimo",
|
|
340
|
+
prefixDiagnostics: true,
|
|
341
|
+
openRouterAffinity: true,
|
|
342
|
+
}),
|
|
271
343
|
inventoryRef: "§4 Xiaomi MiMo",
|
|
272
344
|
},
|
|
273
345
|
{
|
|
@@ -279,6 +351,7 @@ export const POLICY_REGISTRY = [
|
|
|
279
351
|
baseline: "xiaomi.implicit-cache",
|
|
280
352
|
overlays: [],
|
|
281
353
|
legacy: false,
|
|
354
|
+
runtime: rt("neutral"),
|
|
282
355
|
policyStatus: "documented-series-member-without-registered-overlay",
|
|
283
356
|
note: "Documented as a Pro mode in the same V2.6 series, but the inventory does not establish identical cache controls and CacheEngine registers no overlay for it.",
|
|
284
357
|
inventoryRef: "§4 Xiaomi MiMo",
|
|
@@ -293,6 +366,7 @@ export const POLICY_REGISTRY = [
|
|
|
293
366
|
baseline: "deepseek.kv-cache",
|
|
294
367
|
overlays: [],
|
|
295
368
|
legacy: true,
|
|
369
|
+
runtime: rt("deepseek"),
|
|
296
370
|
inventoryRef: "§2 DeepSeek",
|
|
297
371
|
},
|
|
298
372
|
]
|
|
@@ -301,13 +375,16 @@ export const POLICY_REGISTRY = [
|
|
|
301
375
|
// Explicit aliases identified by the inventory
|
|
302
376
|
// ---------------------------------------------------------------------------
|
|
303
377
|
|
|
378
|
+
// `legacy` records whether the pre-v0.4.0 classifier already matched this alias.
|
|
379
|
+
// Only legacy aliases carry runtime capabilities; newer documented aliases are
|
|
380
|
+
// resolved for information but stay runtime-neutral (no new optimization).
|
|
304
381
|
export const MODEL_ALIASES = {
|
|
305
|
-
"gpt-5.6": { canonicalId: "gpt-5.6-sol", family: "gpt-5.6", creator: "openai", inventoryRef: "§1 OpenAI" },
|
|
306
|
-
"gpt-daybreak-blue-latest": { canonicalId: "gpt-5.6-sol", family: "gpt-5.6", creator: "openai", inventoryRef: "§1 OpenAI" },
|
|
307
|
-
"gpt-daybreak-red-latest": { canonicalId: "gpt-5.6-cyber", family: "gpt-5.6", creator: "openai", inventoryRef: "§1 OpenAI" },
|
|
308
|
-
"deepseek-v4-flash": { canonicalId: "deepseek-flash", family: "deepseek", creator: "deepseek", status: "retired-legacy-id", inventoryRef: "§2 DeepSeek" },
|
|
309
|
-
"deepseek-chat": { canonicalId: null, family: "deepseek", creator: "deepseek", status: "retired", inventoryRef: "§2 DeepSeek" },
|
|
310
|
-
"deepseek-reasoner": { canonicalId: null, family: "deepseek", creator: "deepseek", status: "retired", inventoryRef: "§2 DeepSeek" },
|
|
382
|
+
"gpt-5.6": { canonicalId: "gpt-5.6-sol", family: "gpt-5.6", creator: "openai", legacy: true, inventoryRef: "§1 OpenAI" },
|
|
383
|
+
"gpt-daybreak-blue-latest": { canonicalId: "gpt-5.6-sol", family: "gpt-5.6", creator: "openai", legacy: false, inventoryRef: "§1 OpenAI" },
|
|
384
|
+
"gpt-daybreak-red-latest": { canonicalId: "gpt-5.6-cyber", family: "gpt-5.6", creator: "openai", legacy: false, inventoryRef: "§1 OpenAI" },
|
|
385
|
+
"deepseek-v4-flash": { canonicalId: "deepseek-flash", family: "deepseek", creator: "deepseek", legacy: true, status: "retired-legacy-id", inventoryRef: "§2 DeepSeek" },
|
|
386
|
+
"deepseek-chat": { canonicalId: null, family: "deepseek", creator: "deepseek", legacy: true, status: "retired", inventoryRef: "§2 DeepSeek" },
|
|
387
|
+
"deepseek-reasoner": { canonicalId: null, family: "deepseek", creator: "deepseek", legacy: true, status: "retired", inventoryRef: "§2 DeepSeek" },
|
|
311
388
|
}
|
|
312
389
|
|
|
313
390
|
// ---------------------------------------------------------------------------
|
|
@@ -320,6 +397,7 @@ function neutralResult(reason, transport) {
|
|
|
320
397
|
family: "neutral",
|
|
321
398
|
baseline: BASELINES["neutral.none"],
|
|
322
399
|
overlays: [],
|
|
400
|
+
runtime: NEUTRAL_RUNTIME,
|
|
323
401
|
transport,
|
|
324
402
|
matchType: "neutral",
|
|
325
403
|
matchReason: reason,
|
|
@@ -333,12 +411,25 @@ function overlaysFor(ids) {
|
|
|
333
411
|
return (ids ?? []).map((id) => OVERLAYS[id]).filter(Boolean)
|
|
334
412
|
}
|
|
335
413
|
|
|
414
|
+
// A family entry matches by regex `pattern` or by a pure `predicate(slug)`.
|
|
415
|
+
function familyMatches(entry, slug) {
|
|
416
|
+
if (typeof entry.predicate === "function") return entry.predicate(slug)
|
|
417
|
+
return entry.pattern ? entry.pattern.test(slug) : false
|
|
418
|
+
}
|
|
419
|
+
|
|
420
|
+
// Only legacy entries carry runtime capabilities. A non-legacy entry (gpt-6,
|
|
421
|
+
// Pro UltraSpeed) resolves for information but stays neutral at runtime.
|
|
422
|
+
function runtimeForEntry(entry) {
|
|
423
|
+
return entry.legacy === false ? NEUTRAL_RUNTIME : entry.runtime ?? NEUTRAL_RUNTIME
|
|
424
|
+
}
|
|
425
|
+
|
|
336
426
|
function resultFromEntry(entry, matchType, matchReason, matchedId, transport) {
|
|
337
427
|
return {
|
|
338
428
|
creator: entry.creator,
|
|
339
429
|
family: entry.family,
|
|
340
430
|
baseline: BASELINES[entry.baseline] ?? null,
|
|
341
431
|
overlays: overlaysFor(entry.overlays),
|
|
432
|
+
runtime: runtimeForEntry(entry),
|
|
342
433
|
transport,
|
|
343
434
|
matchType,
|
|
344
435
|
matchReason,
|
|
@@ -348,13 +439,17 @@ function resultFromEntry(entry, matchType, matchReason, matchedId, transport) {
|
|
|
348
439
|
}
|
|
349
440
|
}
|
|
350
441
|
|
|
351
|
-
function resultFromFamily(family, creator, matchType, matchReason, matchedId, inventoryRef, note, transport) {
|
|
442
|
+
function resultFromFamily(family, creator, matchType, matchReason, matchedId, inventoryRef, note, transport, aliasLegacy) {
|
|
352
443
|
const entry = POLICY_REGISTRY.find((e) => e.family === family && e.kind !== "exact")
|
|
444
|
+
// An alias is runtime-active only when both the alias and its target family
|
|
445
|
+
// were recognized before v0.4.0.
|
|
446
|
+
const active = aliasLegacy !== false && (!entry || entry.legacy !== false)
|
|
353
447
|
return {
|
|
354
448
|
creator,
|
|
355
449
|
family,
|
|
356
450
|
baseline: entry ? BASELINES[entry.baseline] ?? null : null,
|
|
357
451
|
overlays: entry ? overlaysFor(entry.overlays) : [],
|
|
452
|
+
runtime: active && entry ? entry.runtime ?? NEUTRAL_RUNTIME : NEUTRAL_RUNTIME,
|
|
358
453
|
transport,
|
|
359
454
|
matchType,
|
|
360
455
|
matchReason,
|
|
@@ -386,20 +481,23 @@ export function resolvePolicy(model) {
|
|
|
386
481
|
const familyEntry = POLICY_REGISTRY.find((e) => e.family === alias.family && e.kind !== "exact")
|
|
387
482
|
if (familyEntry?.requiresOpenAIish && !isOpenAIish(s)) continue
|
|
388
483
|
const reason = `alias:${id}->${alias.canonicalId ?? alias.family}`
|
|
389
|
-
return resultFromFamily(alias.family, alias.creator, "exact", reason, id, alias.inventoryRef, alias.status ?? null, transport)
|
|
484
|
+
return resultFromFamily(alias.family, alias.creator, "exact", reason, id, alias.inventoryRef, alias.status ?? null, transport, alias.legacy)
|
|
390
485
|
}
|
|
391
486
|
|
|
392
|
-
// 2. Exact model ids (documented models).
|
|
487
|
+
// 2. Exact model ids (documented models). The entry's context gate still
|
|
488
|
+
// applies, so an exact OpenAI id on a non-OpenAI endpoint is never guessed.
|
|
393
489
|
for (const entry of POLICY_REGISTRY) {
|
|
394
490
|
if (!entry.exactIds || entry.exactIds.length === 0) continue
|
|
395
491
|
const hit = ids.find((id) => entry.exactIds.includes(id))
|
|
396
|
-
if (hit)
|
|
492
|
+
if (!hit) continue
|
|
493
|
+
if (entry.requiresOpenAIish && !isOpenAIish(s)) continue
|
|
494
|
+
return resultFromEntry(entry, "exact", `exact-id:${hit}`, hit, transport)
|
|
397
495
|
}
|
|
398
496
|
|
|
399
|
-
// 3. Model family / range
|
|
497
|
+
// 3. Model family / range matchers (regex pattern or version predicate).
|
|
400
498
|
for (const entry of POLICY_REGISTRY) {
|
|
401
|
-
if (entry.kind !== "family"
|
|
402
|
-
if (!entry
|
|
499
|
+
if (entry.kind !== "family") continue
|
|
500
|
+
if (!familyMatches(entry, s.slug)) continue
|
|
403
501
|
if (entry.requiresOpenAIish && !isOpenAIish(s)) continue
|
|
404
502
|
return resultFromEntry(entry, "family", `family-pattern:${entry.id}`, null, transport)
|
|
405
503
|
}
|
|
@@ -415,6 +513,13 @@ export function resolvePolicy(model) {
|
|
|
415
513
|
return neutralResult("neutral:no-match", transport)
|
|
416
514
|
}
|
|
417
515
|
|
|
516
|
+
// Convenience accessor for the runtime: the legacy policy string + explicit
|
|
517
|
+
// capability flags. This is the single source the runtime gates on; it is
|
|
518
|
+
// guaranteed equal to the pre-v0.4.0 detectPolicy() classification.
|
|
519
|
+
export function resolveRuntimePolicy(model) {
|
|
520
|
+
return resolvePolicy(model).runtime
|
|
521
|
+
}
|
|
522
|
+
|
|
418
523
|
// Compatibility classification used by detectPolicy(). Reproduces the
|
|
419
524
|
// pre-v0.4.0 behavior exactly: it considers only `legacy` registry entries,
|
|
420
525
|
// excludes newer-generation/alias/exact-overlay additions, and returns a family
|
|
@@ -426,7 +531,7 @@ export function resolveLegacyFamily(model) {
|
|
|
426
531
|
for (const entry of POLICY_REGISTRY) {
|
|
427
532
|
if (!entry.legacy) continue
|
|
428
533
|
if (entry.kind === "family") {
|
|
429
|
-
if (!entry
|
|
534
|
+
if (!familyMatches(entry, s.slug)) continue
|
|
430
535
|
if (entry.requiresOpenAIish && !isOpenAIish(s)) continue
|
|
431
536
|
return entry.family
|
|
432
537
|
}
|
|
@@ -53,8 +53,10 @@ import {
|
|
|
53
53
|
MODEL_ALIASES,
|
|
54
54
|
OVERLAYS,
|
|
55
55
|
POLICY_REGISTRY,
|
|
56
|
+
isGpt56OrLater,
|
|
56
57
|
resolveLegacyFamily,
|
|
57
58
|
resolvePolicy,
|
|
59
|
+
resolveRuntimePolicy,
|
|
58
60
|
} from "../src/cache-policy-core.mjs"
|
|
59
61
|
|
|
60
62
|
const asst = (id, read, write) => ({
|
|
@@ -1358,26 +1360,27 @@ test("resolvePolicy: GPT-5.6 exact + inventory aliases resolve to the gpt-5.6 fa
|
|
|
1358
1360
|
assert.equal(orVariant.matchType, "exact")
|
|
1359
1361
|
})
|
|
1360
1362
|
|
|
1361
|
-
test("
|
|
1363
|
+
test("v0.4.2: GPT-6 resolves through the documented GPT-5.6-and-later boundary", () => {
|
|
1362
1364
|
const r = resolvePolicy(M("openrouter", "openai/gpt-6-luna"))
|
|
1363
1365
|
assert.equal(r.creator, "openai")
|
|
1364
|
-
assert.equal(r.family, "gpt-6")
|
|
1366
|
+
assert.equal(r.family, "gpt-5.6")
|
|
1365
1367
|
assert.equal(baseId(r), "openai.gpt56.cache")
|
|
1366
|
-
|
|
1367
|
-
assert.
|
|
1368
|
+
// GPT-6 is documented in the same cache regime, so it gets the same overlay.
|
|
1369
|
+
assert.deepEqual(overlayIds(r), ["gpt56.prompt-cache-options"])
|
|
1370
|
+
assert.equal(resolvePolicy(M("openai", "gpt-6-astra")).family, "gpt-5.6")
|
|
1368
1371
|
|
|
1369
|
-
//
|
|
1370
|
-
const
|
|
1371
|
-
assert.equal(
|
|
1372
|
-
assert.
|
|
1373
|
-
|
|
1374
|
-
assert.deepEqual(gpt6.overlays, [])
|
|
1372
|
+
// The boundary is version-based, not an exact-model list.
|
|
1373
|
+
const entry = POLICY_REGISTRY.find((e) => e.id === "openai.gpt-5.6-plus")
|
|
1374
|
+
assert.equal(entry.boundary, "GPT-5.6 and later")
|
|
1375
|
+
assert.equal(typeof entry.predicate, "function")
|
|
1376
|
+
assert.ok(entry.inventoryRef)
|
|
1375
1377
|
})
|
|
1376
1378
|
|
|
1377
|
-
test("
|
|
1378
|
-
|
|
1379
|
-
assert.equal(detectPolicy(M("openai", "gpt-6
|
|
1380
|
-
assert.equal(
|
|
1379
|
+
test("v0.4.2: the legacy detectPolicy wrapper follows the same boundary", () => {
|
|
1380
|
+
assert.equal(detectPolicy(M("openai", "gpt-6-astra")), POLICY_GPT56)
|
|
1381
|
+
assert.equal(detectPolicy(M("openai", "gpt-5.6")), POLICY_GPT56)
|
|
1382
|
+
assert.equal(detectPolicy(M("openai", "gpt-5.5")), POLICY_NEUTRAL)
|
|
1383
|
+
assert.equal(detectPolicy(M("openai", "gpt-5.60")), POLICY_NEUTRAL)
|
|
1381
1384
|
})
|
|
1382
1385
|
|
|
1383
1386
|
test("resolvePolicy: pre-5.6 GPT negative controls are neutral with no overlays", () => {
|
|
@@ -1551,3 +1554,268 @@ test("registry is traceable and internally consistent", () => {
|
|
|
1551
1554
|
assert.ok(alias.family)
|
|
1552
1555
|
}
|
|
1553
1556
|
})
|
|
1557
|
+
|
|
1558
|
+
// ===========================================================================
|
|
1559
|
+
// v0.4.1 runtime policy migration (behavior preservation)
|
|
1560
|
+
//
|
|
1561
|
+
// The runtime now classifies via resolveRuntimePolicy(). These tests prove the
|
|
1562
|
+
// resolved policy equals the legacy detectPolicy() string for every supported
|
|
1563
|
+
// model, and that the hook-observable behavior (GPT cache options, <env>
|
|
1564
|
+
// relocation, OpenRouter affinity) is unchanged. Newer/unknown models must
|
|
1565
|
+
// gain no mutation.
|
|
1566
|
+
// ===========================================================================
|
|
1567
|
+
|
|
1568
|
+
const policyCoreURL = new URL("../src/cache-policy-core.mjs", import.meta.url).href
|
|
1569
|
+
|
|
1570
|
+
test("v0.4.1: resolveRuntimePolicy.policy matches detectPolicy across a broad matrix", () => {
|
|
1571
|
+
const samples = [
|
|
1572
|
+
M("openai", "gpt-5.6"),
|
|
1573
|
+
M("openai", "gpt-5.6-sol"),
|
|
1574
|
+
M("openrouter", "openai/gpt-5.6-luna"),
|
|
1575
|
+
M("openai-compatible", "gpt-5.6"),
|
|
1576
|
+
M("openai", "gpt-6-astra"),
|
|
1577
|
+
M("openai", "gpt-5.5"),
|
|
1578
|
+
M("openai", "gpt-daybreak-blue-latest"),
|
|
1579
|
+
M("openai", "gpt-daybreak-red-latest"),
|
|
1580
|
+
M("zai", "glm-5.3"),
|
|
1581
|
+
M("zai", "glm-5.3-flash"),
|
|
1582
|
+
M("zai", "glm-5.2"),
|
|
1583
|
+
M("xiaomi", "mimo-v2.6-flash"),
|
|
1584
|
+
M("xiaomi", "mimo-v2.6-pro"),
|
|
1585
|
+
M("xiaomi", "mimo-v2.6-pro-ultraspeed"),
|
|
1586
|
+
M("xiaomi", "mimo-v2.5"),
|
|
1587
|
+
M("deepseek", "deepseek-v4-pro"),
|
|
1588
|
+
M("deepseek", "deepseek-flash"),
|
|
1589
|
+
M("deepseek", "deepseek-v4-flash"),
|
|
1590
|
+
M("deepseek", "deepseek-chat"),
|
|
1591
|
+
M("openrouter", "x-ai/grok-4"),
|
|
1592
|
+
{},
|
|
1593
|
+
null,
|
|
1594
|
+
undefined,
|
|
1595
|
+
42,
|
|
1596
|
+
]
|
|
1597
|
+
for (const m of samples) {
|
|
1598
|
+
assert.equal(resolveRuntimePolicy(m).policy, detectPolicy(m))
|
|
1599
|
+
}
|
|
1600
|
+
})
|
|
1601
|
+
|
|
1602
|
+
test("v0.4.1: GPT-5.6 alias keeps the overlay but a newer alias stays runtime-neutral", () => {
|
|
1603
|
+
// Legacy alias: gpt-5.6 -> gpt-5.6-sol (was matched by the old regex).
|
|
1604
|
+
const legacy = resolveRuntimePolicy(M("openai", "gpt-5.6"))
|
|
1605
|
+
assert.equal(legacy.policy, "gpt56")
|
|
1606
|
+
assert.equal(legacy.gptCacheMetadata, true)
|
|
1607
|
+
// Newer documented alias that the old classifier did NOT match: no mutation.
|
|
1608
|
+
const newer = resolveRuntimePolicy(M("openai", "gpt-daybreak-blue-latest"))
|
|
1609
|
+
assert.equal(newer.policy, "neutral")
|
|
1610
|
+
assert.equal(newer.gptCacheMetadata, false)
|
|
1611
|
+
assert.equal(resolvePolicy(M("openai", "gpt-daybreak-blue-latest")).family, "gpt-5.6")
|
|
1612
|
+
})
|
|
1613
|
+
|
|
1614
|
+
async function runPolicyMigrationProbe() {
|
|
1615
|
+
const home = mkdtempSync(join(tmpdir(), "ce-policy-migration-"))
|
|
1616
|
+
const pluginURL = new URL("../src/cache-engine.ts", import.meta.url).href
|
|
1617
|
+
const coreURL = new URL("../src/cache-engine-core.mjs", import.meta.url).href
|
|
1618
|
+
const script = `
|
|
1619
|
+
import assert from "node:assert/strict"
|
|
1620
|
+
process.env.CACHE_ENGINE_METRICS_FILE = process.env.HOME + "/policy-migration.jsonl"
|
|
1621
|
+
const { CacheEngine } = await import(${JSON.stringify(pluginURL)})
|
|
1622
|
+
const { detectPolicy } = await import(${JSON.stringify(coreURL)})
|
|
1623
|
+
const { resolvePolicy, resolveRuntimePolicy } = await import(${JSON.stringify(policyCoreURL)})
|
|
1624
|
+
const client = {
|
|
1625
|
+
app: { log: async () => ({}) },
|
|
1626
|
+
session: { get: async () => ({ data: { parentID: null } }) },
|
|
1627
|
+
tool: { list: async () => ({ data: [] }) },
|
|
1628
|
+
}
|
|
1629
|
+
const hooks = await CacheEngine({ client, directory: process.env.HOME })
|
|
1630
|
+
assert.equal(typeof hooks["chat.params"], "function")
|
|
1631
|
+
assert.equal(typeof hooks["chat.headers"], "function")
|
|
1632
|
+
assert.equal(typeof hooks["experimental.chat.system.transform"], "function")
|
|
1633
|
+
const SYS = ["A: keep1", "B: You are powered by the model named x. The exact model ID is acme/x", "C: <env>", "D: Today's date: 2026-08-17", "E: </env>", "F: keep2"].join("\\n")
|
|
1634
|
+
const CASES = [
|
|
1635
|
+
{ name: "gpt-5.6", model: { providerID: "openai", id: "gpt-5.6", api: { id: "gpt-5.6", npm: "@ai-sdk/openai" } }, expect: { policy: "gpt56", env: false, gpt: true, header: false } },
|
|
1636
|
+
{ name: "gpt-5.6-openrouter", model: { providerID: "openrouter", id: "openai/gpt-5.6-sol", api: { id: "openai/gpt-5.6-sol" } }, expect: { policy: "gpt56", env: false, gpt: true, header: false } },
|
|
1637
|
+
{ name: "gpt-6-astra", model: { providerID: "openai", id: "gpt-6-astra", api: { id: "gpt-6-astra", npm: "@ai-sdk/openai" } }, expect: { policy: "gpt56", env: false, gpt: true, header: false } },
|
|
1638
|
+
{ name: "gpt-6-openrouter", model: { providerID: "openrouter", id: "openai/gpt-6-luna", api: { id: "openai/gpt-6-luna" } }, expect: { policy: "gpt56", env: false, gpt: true, header: false } },
|
|
1639
|
+
{ name: "gpt-6-openai-compatible", model: { providerID: "openai-compatible", id: "gpt-6-astra", api: { id: "gpt-6-astra" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
|
|
1640
|
+
{ name: "gpt-5.60-malformed", model: { providerID: "openai", id: "gpt-5.60", api: { id: "gpt-5.60", npm: "@ai-sdk/openai" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
|
|
1641
|
+
{ name: "gpt-5.5", model: { providerID: "openai", id: "gpt-5.5", api: { id: "gpt-5.5", npm: "@ai-sdk/openai" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
|
|
1642
|
+
{ name: "gpt-daybreak-alias", model: { providerID: "openai", id: "gpt-daybreak-blue-latest", api: { id: "gpt-daybreak-blue-latest", npm: "@ai-sdk/openai" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
|
|
1643
|
+
{ name: "deepseek-v4-pro", model: { providerID: "deepseek", id: "deepseek-v4-pro", api: { id: "deepseek-v4-pro" } }, expect: { policy: "deepseek", env: false, gpt: false, header: false } },
|
|
1644
|
+
{ name: "deepseek-flash", model: { providerID: "deepseek", id: "deepseek-flash", api: { id: "deepseek-flash" } }, expect: { policy: "deepseek", env: false, gpt: false, header: false } },
|
|
1645
|
+
{ name: "glm-5.3-direct", model: { providerID: "zai", id: "glm-5.3", api: { id: "glm-5.3" } }, expect: { policy: "glm53", env: true, gpt: false, header: false } },
|
|
1646
|
+
{ name: "glm-5.3-openrouter", model: { providerID: "openrouter", id: "z-ai/glm-5.3-flash", api: { id: "z-ai/glm-5.3-flash" } }, expect: { policy: "glm53", env: true, gpt: false, header: true } },
|
|
1647
|
+
{ name: "glm-5.2", model: { providerID: "zai", id: "glm-5.2", api: { id: "glm-5.2" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
|
|
1648
|
+
{ name: "mimo-v2.6-flash-direct", model: { providerID: "xiaomi", id: "mimo-v2.6-flash", api: { id: "mimo-v2.6-flash" } }, expect: { policy: "mimo26", env: true, gpt: false, header: false } },
|
|
1649
|
+
{ name: "mimo-v2.6-pro-openrouter", model: { providerID: "openrouter", id: "xiaomi/mimo-v2.6-pro", api: { id: "xiaomi/mimo-v2.6-pro" } }, expect: { policy: "mimo26", env: true, gpt: false, header: true } },
|
|
1650
|
+
{ name: "mimo-v2.6-pro-ultraspeed", model: { providerID: "xiaomi", id: "mimo-v2.6-pro-ultraspeed", api: { id: "mimo-v2.6-pro-ultraspeed" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
|
|
1651
|
+
{ name: "mimo-v2.5", model: { providerID: "xiaomi", id: "mimo-v2.5", api: { id: "mimo-v2.5" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
|
|
1652
|
+
{ name: "unknown-provider", model: { providerID: "mystery-provider", id: "xiaomi/mimo-v2.6-flash", api: { id: "xiaomi/mimo-v2.6-flash" } }, expect: { policy: "mimo26", env: true, gpt: false, header: false } },
|
|
1653
|
+
{ name: "unknown-openrouter", model: { providerID: "openrouter", id: "acme/mystery-9", api: { id: "acme/mystery-9" } }, expect: { policy: "neutral", env: false, gpt: false, header: false } },
|
|
1654
|
+
]
|
|
1655
|
+
const results = []
|
|
1656
|
+
for (const c of CASES) {
|
|
1657
|
+
const model = c.model
|
|
1658
|
+
const sid = "ses_" + c.name
|
|
1659
|
+
const provider = { source: "config", info: { id: String(model.providerID ?? "") }, options: {} }
|
|
1660
|
+
const existing = { "User-Agent": "preserve", "x-custom": "preserve" }
|
|
1661
|
+
const paramsOut = { options: {} }
|
|
1662
|
+
const headersOut = { headers: { ...existing } }
|
|
1663
|
+
const sysOut = { system: [SYS] }
|
|
1664
|
+
await hooks["chat.params"]({ sessionID: sid, agent: "build", model, provider, message: { id: "msg-" + c.name, sessionID: sid, role: "user", content: "probe" } }, paramsOut)
|
|
1665
|
+
await hooks["experimental.chat.system.transform"]({ sessionID: sid, model, provider }, sysOut)
|
|
1666
|
+
await hooks["chat.headers"]({ sessionID: sid, agent: "build", model, provider, message: { id: "msg-" + c.name, sessionID: sid, role: "user", content: "probe" } }, headersOut)
|
|
1667
|
+
const rt = resolveRuntimePolicy(model)
|
|
1668
|
+
const existingHeadersPreserved = Object.entries(existing).every(([k, v]) => headersOut.headers[k] === v)
|
|
1669
|
+
results.push({
|
|
1670
|
+
name: c.name,
|
|
1671
|
+
runtimePolicy: rt.policy,
|
|
1672
|
+
detectPolicy: detectPolicy(model),
|
|
1673
|
+
richFamily: resolvePolicy(model).family,
|
|
1674
|
+
overlays: resolvePolicy(model).overlays.map((o) => o.id),
|
|
1675
|
+
systemRelocated: sysOut.system[0] !== SYS,
|
|
1676
|
+
gptOptionInjected: paramsOut.options.promptCacheKey !== undefined,
|
|
1677
|
+
gptOptions: paramsOut.options.promptCacheOptions ?? null,
|
|
1678
|
+
affinityHeaderAttached: headersOut.headers["x-session-id"] !== undefined,
|
|
1679
|
+
existingHeadersPreserved,
|
|
1680
|
+
expect: c.expect,
|
|
1681
|
+
})
|
|
1682
|
+
}
|
|
1683
|
+
process.stdout.write(JSON.stringify({ results }))
|
|
1684
|
+
`
|
|
1685
|
+
const stdout = execFileSync(process.execPath, ["--experimental-strip-types", "--input-type=module", "-e", script], {
|
|
1686
|
+
cwd: process.cwd(),
|
|
1687
|
+
env: { ...process.env, HOME: home },
|
|
1688
|
+
encoding: "utf8",
|
|
1689
|
+
})
|
|
1690
|
+
return JSON.parse(stdout.trim())
|
|
1691
|
+
}
|
|
1692
|
+
|
|
1693
|
+
let policyMigrationProbe
|
|
1694
|
+
const policyMigrationResults = async () => (policyMigrationProbe ??= runPolicyMigrationProbe())
|
|
1695
|
+
|
|
1696
|
+
test("v0.4.1: runtime policy is the resolver's and stays equal to detectPolicy (hook path)", async () => {
|
|
1697
|
+
const { results } = await policyMigrationResults()
|
|
1698
|
+
assert.ok(results.length >= 15)
|
|
1699
|
+
for (const r of results) {
|
|
1700
|
+
assert.equal(r.runtimePolicy, r.detectPolicy, `${r.name}: resolver policy must equal legacy detectPolicy`)
|
|
1701
|
+
assert.equal(r.runtimePolicy, r.expect.policy, `${r.name}: unexpected policy`)
|
|
1702
|
+
}
|
|
1703
|
+
})
|
|
1704
|
+
|
|
1705
|
+
test("v0.4.1: every supported model keeps its pre-migration hook behavior", async () => {
|
|
1706
|
+
const { results } = await policyMigrationResults()
|
|
1707
|
+
for (const r of results) {
|
|
1708
|
+
assert.equal(r.systemRelocated, r.expect.env, `${r.name}: <env> relocation`)
|
|
1709
|
+
assert.equal(r.gptOptionInjected, r.expect.gpt, `${r.name}: GPT cache-options injection`)
|
|
1710
|
+
assert.equal(r.affinityHeaderAttached, r.expect.header, `${r.name}: OpenRouter affinity header`)
|
|
1711
|
+
// No provider in the matrix mutates or drops pre-existing headers.
|
|
1712
|
+
assert.equal(r.existingHeadersPreserved, true, `${r.name}: existing headers preserved`)
|
|
1713
|
+
}
|
|
1714
|
+
})
|
|
1715
|
+
|
|
1716
|
+
test("v0.4.1: GPT-5.6 keeps promptCacheOptions implicit/30m through the resolver", async () => {
|
|
1717
|
+
const { results } = await policyMigrationResults()
|
|
1718
|
+
const gpt = results.find((r) => r.name === "gpt-5.6")
|
|
1719
|
+
assert.deepEqual(gpt.gptOptions, { mode: "implicit", ttl: "30m" })
|
|
1720
|
+
})
|
|
1721
|
+
|
|
1722
|
+
test("v0.4.1: future-looking and unknown models gain no mutation", async () => {
|
|
1723
|
+
const { results } = await policyMigrationResults()
|
|
1724
|
+
const noMutation = [
|
|
1725
|
+
"gpt-daybreak-alias",
|
|
1726
|
+
"gpt-5.5",
|
|
1727
|
+
"gpt-6-openai-compatible",
|
|
1728
|
+
"gpt-5.60-malformed",
|
|
1729
|
+
"glm-5.2",
|
|
1730
|
+
"mimo-v2.6-pro-ultraspeed",
|
|
1731
|
+
"mimo-v2.5",
|
|
1732
|
+
"unknown-openrouter",
|
|
1733
|
+
]
|
|
1734
|
+
for (const name of noMutation) {
|
|
1735
|
+
const r = results.find((x) => x.name === name)
|
|
1736
|
+
assert.ok(r, `${name} present`)
|
|
1737
|
+
assert.equal(r.gptOptionInjected, false, `${name}: no GPT options`)
|
|
1738
|
+
assert.equal(r.systemRelocated, false, `${name}: no <env> relocation`)
|
|
1739
|
+
assert.equal(r.affinityHeaderAttached, false, `${name}: no affinity header`)
|
|
1740
|
+
}
|
|
1741
|
+
})
|
|
1742
|
+
|
|
1743
|
+
test("v0.4.1: non-OpenRouter models never receive the OpenRouter header", async () => {
|
|
1744
|
+
const { results } = await policyMigrationResults()
|
|
1745
|
+
for (const name of ["glm-5.3-direct", "mimo-v2.6-flash-direct", "unknown-provider", "gpt-5.6", "deepseek-v4-pro"]) {
|
|
1746
|
+
const r = results.find((x) => x.name === name)
|
|
1747
|
+
assert.equal(r.affinityHeaderAttached, false, `${name}: no affinity header off OpenRouter`)
|
|
1748
|
+
}
|
|
1749
|
+
})
|
|
1750
|
+
|
|
1751
|
+
// ===========================================================================
|
|
1752
|
+
// v0.4.2 GPT-5.6-and-later boundary
|
|
1753
|
+
//
|
|
1754
|
+
// Source: OpenAI *Prompt caching* guide, "GPT-5.6 and later" generation
|
|
1755
|
+
// boundary, re-verified 2026-09-27 (docs/cache-policy-inventory.md §1). GPT-6
|
|
1756
|
+
// is documented in the same regime with no cache-control exception.
|
|
1757
|
+
// ===========================================================================
|
|
1758
|
+
|
|
1759
|
+
test("v0.4.2: isGpt56OrLater matches the documented boundary by version", () => {
|
|
1760
|
+
const inFamily = [
|
|
1761
|
+
"gpt-5.6",
|
|
1762
|
+
"gpt-5.6-luna",
|
|
1763
|
+
"openai/gpt-5.6-sol",
|
|
1764
|
+
"gpt-6",
|
|
1765
|
+
"gpt-6-astra",
|
|
1766
|
+
"gpt-6-sol",
|
|
1767
|
+
"gpt-6-luna",
|
|
1768
|
+
"gpt-5.7",
|
|
1769
|
+
"gpt-7",
|
|
1770
|
+
"gpt-6.1",
|
|
1771
|
+
]
|
|
1772
|
+
for (const id of inFamily) assert.equal(isGpt56OrLater(id), true, `${id} is in family`)
|
|
1773
|
+
const outOfFamily = [
|
|
1774
|
+
"gpt-5.5",
|
|
1775
|
+
"gpt-5.4",
|
|
1776
|
+
"gpt-5.2",
|
|
1777
|
+
"gpt-5.1",
|
|
1778
|
+
"gpt-5",
|
|
1779
|
+
"gpt-4.1",
|
|
1780
|
+
"gpt-4o",
|
|
1781
|
+
"gpt-5.60",
|
|
1782
|
+
"gpt-5.6.1",
|
|
1783
|
+
"gpt-4",
|
|
1784
|
+
"claude-sonnet-4-5",
|
|
1785
|
+
]
|
|
1786
|
+
for (const id of outOfFamily) assert.equal(isGpt56OrLater(id), false, `${id} is out of family`)
|
|
1787
|
+
})
|
|
1788
|
+
|
|
1789
|
+
test("v0.4.2: GPT boundary respects OpenAI/provider gating", () => {
|
|
1790
|
+
assert.equal(resolveRuntimePolicy(M("openai", "gpt-5.6")).gptCacheMetadata, true)
|
|
1791
|
+
assert.equal(resolveRuntimePolicy(M("openai", "gpt-6-astra")).gptCacheMetadata, true)
|
|
1792
|
+
assert.equal(resolveRuntimePolicy(M("azure", "gpt-6-sol")).gptCacheMetadata, true)
|
|
1793
|
+
assert.equal(resolveRuntimePolicy(M("openrouter", "openai/gpt-6-luna")).gptCacheMetadata, true)
|
|
1794
|
+
assert.equal(resolveRuntimePolicy(M("openai-compatible", "gpt-6-astra")).gptCacheMetadata, false)
|
|
1795
|
+
assert.equal(resolveRuntimePolicy(M("llama.cpp", "gpt-5.6")).gptCacheMetadata, false)
|
|
1796
|
+
assert.equal(resolveRuntimePolicy(M("openai", "gpt-5.5")).gptCacheMetadata, false)
|
|
1797
|
+
assert.equal(resolveRuntimePolicy(M("openai", "gpt-5.60")).gptCacheMetadata, false)
|
|
1798
|
+
})
|
|
1799
|
+
|
|
1800
|
+
test("v0.4.2: covered later GPT requests receive the same documented baseline at runtime", async () => {
|
|
1801
|
+
const { results } = await policyMigrationResults()
|
|
1802
|
+
const gpt6 = results.find((r) => r.name === "gpt-6-astra")
|
|
1803
|
+
assert.equal(gpt6.runtimePolicy, "gpt56")
|
|
1804
|
+
assert.equal(gpt6.gptOptionInjected, true)
|
|
1805
|
+
assert.deepEqual(gpt6.gptOptions, { mode: "implicit", ttl: "30m" })
|
|
1806
|
+
// Same baseline as GPT-5.6, no new mechanism.
|
|
1807
|
+
const gpt56 = results.find((r) => r.name === "gpt-5.6")
|
|
1808
|
+
assert.deepEqual(gpt6.gptOptions, gpt56.gptOptions)
|
|
1809
|
+
// No prompt transformation or affinity was introduced for gpt-6.
|
|
1810
|
+
assert.equal(gpt6.systemRelocated, false)
|
|
1811
|
+
assert.equal(gpt6.affinityHeaderAttached, false)
|
|
1812
|
+
})
|
|
1813
|
+
|
|
1814
|
+
test("v0.4.2: pre-5.6 and out-of-family GPT ids get no GPT options", async () => {
|
|
1815
|
+
const { results } = await policyMigrationResults()
|
|
1816
|
+
for (const name of ["gpt-5.5", "gpt-5.60-malformed", "gpt-6-openai-compatible"]) {
|
|
1817
|
+
const r = results.find((x) => x.name === name)
|
|
1818
|
+
assert.equal(r.gptOptionInjected, false, `${name}: no GPT options`)
|
|
1819
|
+
assert.equal(r.runtimePolicy, "neutral", `${name}: neutral runtime`)
|
|
1820
|
+
}
|
|
1821
|
+
})
|