opencode-cache-engine 0.2.0 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +295 -11
- package/examples/cache-engine.json +6 -0
- package/package.json +3 -2
- package/src/cache-engine-core.mjs +76 -9
- package/src/cache-engine.ts +124 -10
- package/test/cache-engine.test.mjs +206 -1
- package/README_rewritten.md +0 -965
package/README.md
CHANGED
|
@@ -12,11 +12,12 @@ The server target handles cache optimization, prompt-shape diagnostics, compacti
|
|
|
12
12
|
|
|
13
13
|
`CacheEngine` is an OpenCode plugin designed for long-running agent sessions where prompt-cache efficiency affects both latency and cost. It keeps the harness conservative for providers whose cache behavior is already automatic, while applying provider-specific optimizations where the provider exposes useful cache controls or where prompt structure can be safely improved.
|
|
14
14
|
|
|
15
|
-
The plugin currently has
|
|
15
|
+
The plugin currently has four cache-policy families:
|
|
16
16
|
|
|
17
17
|
* **DeepSeek V4.1 Flash** — passive cache-stability and observability
|
|
18
18
|
* **GPT-5.6 Luna** — active cache-control configuration
|
|
19
19
|
* **GLM-5.3 Flash** — conservative system-prompt stabilization
|
|
20
|
+
* **MiMo-V2.6 (Flash / Pro)** — prefix stability and OpenRouter session-affinity diagnostics
|
|
20
21
|
|
|
21
22
|
The central design principle is:
|
|
22
23
|
|
|
@@ -35,8 +36,9 @@ It:
|
|
|
35
36
|
4. Records provider-reported cache token usage.
|
|
36
37
|
5. Adds a deterministic compaction continuation block.
|
|
37
38
|
6. Applies GPT-5.6 cache-control metadata.
|
|
38
|
-
7. Applies the GLM-5.3 volatile-environment relocation.
|
|
39
|
+
7. Applies the GLM-5.3 and MiMo-V2.6 volatile-environment relocation.
|
|
39
40
|
8. Records diagnostics that help determine whether prompt-shape changes correlate with cache behavior.
|
|
41
|
+
9. Records MiMo-V2.6 provider identity and provider-switch diagnostics.
|
|
40
42
|
|
|
41
43
|
The plugin deliberately avoids pretending that a local hash is proof of a provider cache hit. Provider-reported token usage remains the authoritative signal.
|
|
42
44
|
|
|
@@ -224,15 +226,167 @@ It only occurs when:
|
|
|
224
226
|
The plugin does not arbitrarily rearrange unrelated prompt content.
|
|
225
227
|
|
|
226
228
|
|
|
227
|
-
|
|
229
|
+
## MiMo-V2.6 (Flash / Pro)
|
|
230
|
+
|
|
231
|
+
### Policy: prefix stability + OpenRouter session affinity
|
|
232
|
+
|
|
233
|
+
MiMo-V2.6 is Xiaomi's current model family. The plugin targets exactly two
|
|
234
|
+
identifiers:
|
|
235
|
+
|
|
236
|
+
* `xiaomi/mimo-v2.6-flash` / `mimo-v2.6-flash`
|
|
237
|
+
* `xiaomi/mimo-v2.6-pro` / `mimo-v2.6-pro`
|
|
238
|
+
|
|
239
|
+
Detection also tolerates `provider/model` shapes where `api.id` contains those
|
|
240
|
+
slugs. It deliberately does **not** match `mimo-v2.5`, `mimo-v2.5-pro`,
|
|
241
|
+
`mimo-v2.6-pro-ultraspeed`, or unrelated MiMo models.
|
|
242
|
+
|
|
243
|
+
### Implicit context caching
|
|
244
|
+
|
|
245
|
+
Xiaomi documents context caching for both V2.6 Flash and Pro, and exposes
|
|
246
|
+
`usage.prompt_tokens_details.cached_tokens` as the number of prompt tokens
|
|
247
|
+
served from cache. The V2.6 API documents implicit context caching, not a
|
|
248
|
+
user-supplied cache key or explicit breakpoint.
|
|
249
|
+
|
|
250
|
+
Accordingly the plugin **injects no cache-control parameter** for MiMo. It does
|
|
251
|
+
not send `promptCacheKey`, `cacheControl`, `cacheBreakpoint`, or `ttl`.
|
|
252
|
+
Implicit caching is the default assumption.
|
|
253
|
+
|
|
254
|
+
### Environment-block stabilization
|
|
255
|
+
|
|
256
|
+
MiMo uses the same narrow, content-preserving transformation as GLM-5.3: the
|
|
257
|
+
identifiable volatile `<env>` block is relocated to the **tail** of the single
|
|
258
|
+
system string. Contents are preserved byte-for-byte; only position changes. This
|
|
259
|
+
keeps the large reusable prefix stable when only the environment/date changes.
|
|
260
|
+
|
|
261
|
+
The transformation is applied only when:
|
|
262
|
+
|
|
263
|
+
* the selected model is MiMo-V2.6 Flash/Pro
|
|
264
|
+
* `mimo26.stabilizeSystem` is `true`
|
|
265
|
+
* there is exactly one system string
|
|
266
|
+
* the expected `<env>` markers exist and the block is identified unambiguously
|
|
267
|
+
* the block is not already at the tail
|
|
268
|
+
|
|
269
|
+
### No generic system-prompt freezing
|
|
270
|
+
|
|
271
|
+
MiMo-Code's own harness freezes its per-session system prefix. This plugin does
|
|
272
|
+
**not** copy that mechanism. System instructions can legitimately change because
|
|
273
|
+
of permissions, tools, agent mode, skills, MCP state, or project configuration;
|
|
274
|
+
a plugin-level snapshot must never override a legitimate change.
|
|
275
|
+
|
|
276
|
+
Instead the plugin:
|
|
277
|
+
|
|
278
|
+
* records a first-seen system baseline per session;
|
|
279
|
+
* computes the full system hash, stable prefix hash, and volatile suffix hash;
|
|
280
|
+
* records changes for MiMo sessions;
|
|
281
|
+
* allows the `<env>` relocation when that is the only identified volatility;
|
|
282
|
+
* reports other system changes diagnostically and never overwrites the new
|
|
283
|
+
content.
|
|
284
|
+
|
|
285
|
+
Explicit telemetry events:
|
|
286
|
+
|
|
287
|
+
* `mimo_system_env_relocated`
|
|
288
|
+
* `mimo_system_prefix_changed`
|
|
289
|
+
|
|
290
|
+
### OpenRouter sticky session — derived but not injected
|
|
291
|
+
|
|
292
|
+
OpenRouter documents a top-level `session_id` request field for sticky provider
|
|
293
|
+
routing, which keeps a session's requests on the same upstream provider so
|
|
294
|
+
provider-side prompt caches stay warm.
|
|
295
|
+
|
|
296
|
+
The plugin includes a pure, session-scoped derivation (`mimoSessionIdFor`):
|
|
297
|
+
deterministic, distinct per session, printable/no-whitespace, and well under the
|
|
298
|
+
256-character cap. However, **the derived id is not injected into requests**.
|
|
299
|
+
|
|
300
|
+
Rationale (verified against the installed runtime): OpenCode's OpenRouter
|
|
301
|
+
request adapter forwards only `usage`, `reasoning`, and `prompt_cache_key` from
|
|
302
|
+
provider options, and exposes no top-level `session_id` path. Adding an
|
|
303
|
+
unsupported field would be guessing, so the id is recorded as telemetry only,
|
|
304
|
+
and `mimo26.stickySession` currently gates that recording. If a future runtime
|
|
305
|
+
gains a verified `session_id` path, the helper is already in place.
|
|
306
|
+
|
|
307
|
+
Note that OpenCode itself sets `x-session-affinity` / `X-Session-Id` HTTP
|
|
308
|
+
headers for non-opencode providers, and can set a flat `promptCacheKey` for
|
|
309
|
+
OpenRouter when `setCacheKey: true` is configured. Those are HTTP routing
|
|
310
|
+
headers and an OpenAI-style cache key respectively — they are not OpenRouter's
|
|
311
|
+
documented body `session_id`.
|
|
312
|
+
|
|
313
|
+
### MiMo cache metrics
|
|
314
|
+
|
|
315
|
+
MiMo caches are provider-managed, so the authoritative metric is provider
|
|
316
|
+
reported. For MiMo the plugin emits the preferred ratio:
|
|
317
|
+
|
|
318
|
+
```text
|
|
319
|
+
cacheHitRate = cachedTokens / promptTokens
|
|
320
|
+
```
|
|
321
|
+
|
|
322
|
+
This is intentionally **not** the `read / (read + write)` form used by other
|
|
323
|
+
families. It is not GLM's `read / (read + write + input)` either.
|
|
324
|
+
|
|
325
|
+
Derivation: the runtime exposes assistant tokens as `{ input, output,
|
|
326
|
+
cache:{ read, write } }`, where `input` is the non-cached prompt input and
|
|
327
|
+
`cache.read` is the cached prompt input. Total prompt tokens are therefore
|
|
328
|
+
derived as `read + input`, and `cachedTokens = read`. `cache.write` is a
|
|
329
|
+
separate accounting bucket and is not folded in; no cache-write value is
|
|
330
|
+
fabricated, and the ratio is `null` when `promptTokens` is zero.
|
|
331
|
+
|
|
332
|
+
A MiMo usage record looks conceptually like:
|
|
333
|
+
|
|
334
|
+
```json
|
|
335
|
+
{
|
|
336
|
+
"kind": "usage",
|
|
337
|
+
"policy": "mimo26",
|
|
338
|
+
"provider": "openrouter",
|
|
339
|
+
"model": "xiaomi/mimo-v2.6-flash",
|
|
340
|
+
"promptTokens": 50000,
|
|
341
|
+
"cachedTokens": 47000,
|
|
342
|
+
"cacheHitRate": 94
|
|
343
|
+
}
|
|
344
|
+
```
|
|
345
|
+
|
|
346
|
+
### Provider-switch diagnostics
|
|
347
|
+
|
|
348
|
+
Because MiMo caches live at the provider side, a provider change within one
|
|
349
|
+
session can silently invalidate them. The plugin records provider identity on
|
|
350
|
+
every MiMo request and emits a `mimo_provider_changed` boundary event when the
|
|
351
|
+
OpenCode `providerID` changes within a session. It never forces or overrides the
|
|
352
|
+
user's provider selection.
|
|
353
|
+
|
|
354
|
+
Limitation: OpenRouter's *upstream* provider selection (for example
|
|
355
|
+
`xiaomi/fp8` vs `atlas-cloud/fp8`) is not exposed to plugins, so only the
|
|
356
|
+
OpenCode `providerID`/`modelID` are observable.
|
|
357
|
+
|
|
358
|
+
### Reasoning / thinking
|
|
228
359
|
|
|
229
|
-
|
|
360
|
+
MiMo-V2.6 supports deep thinking and reports reasoning tokens. The plugin does
|
|
361
|
+
not treat reasoning replay as a cache requirement: reasoning diagnostics are
|
|
362
|
+
instrumentation only, and the plugin never rewrites, duplicates, reorders, or
|
|
363
|
+
re-injects reasoning content, nor changes reasoning effort for caching.
|
|
230
364
|
|
|
231
|
-
|
|
232
|
-
|
|
233
|
-
|
|
234
|
-
|
|
235
|
-
|
|
365
|
+
### Skill-catalog / history limitation
|
|
366
|
+
|
|
367
|
+
MiMo-Code moved skill catalogs out of repeatedly rewritten user messages and
|
|
368
|
+
toward the system tail. In this OpenCode runtime the skill guidance
|
|
369
|
+
(`<available_skills>`) and MCP instructions already live in the **system
|
|
370
|
+
prefix**, not in user-message history. The plugin therefore performs no
|
|
371
|
+
message-history rewrite. Skill/MCP changes simply appear as system-prefix changes
|
|
372
|
+
and are reported diagnostically; the message content is left untouched.
|
|
373
|
+
|
|
374
|
+
|
|
375
|
+
# Provider comparison
|
|
376
|
+
|
|
377
|
+
| Provider | Detection | Prompt text changed? | Cache metadata changed? | Primary cache signal |
|
|
378
|
+
| ------------------- | ---------------------------------- | --------------------------- | ---------------------------------- | --------------------------------------- |
|
|
379
|
+
| DeepSeek V4.1 Flash | `deepseek` | No | No | provider `cache.read`/`cache.write` |
|
|
380
|
+
| GPT-5.6 Luna | `gpt-5.6*` on OpenAI-ish endpoints | No | Yes: `prompt_cache_key` + options | provider cache tokens |
|
|
381
|
+
| GLM-5.3 Flash | `glm-5.3*` | Yes, narrowly (`<env>` tail) | No provider cache key | provider cache tokens (GLM ratio) |
|
|
382
|
+
| MiMo-V2.6 Flash/Pro | `mimo-v2.6-flash` / `mimo-v2.6-pro` | Yes, narrowly (`<env>` tail) | No: implicit caching only | `cached_tokens / prompt_tokens` |
|
|
383
|
+
|
|
384
|
+
|
|
385
|
+
# Prompt-cache strategy
|
|
386
|
+
|
|
387
|
+
The plugin uses different strategies because cache mechanisms differ by provider.
|
|
388
|
+
The table above summarises them; the essential point is the distinction between
|
|
389
|
+
*changing prompt text* and *changing cache metadata*.
|
|
236
390
|
|
|
237
391
|
This distinction is fundamental.
|
|
238
392
|
|
|
@@ -349,6 +503,15 @@ read / (read + write + input)
|
|
|
349
503
|
|
|
350
504
|
as implemented by `glmHitRatio()`.
|
|
351
505
|
|
|
506
|
+
For MiMo, the implementation uses the provider-documented prompt-cache ratio:
|
|
507
|
+
|
|
508
|
+
```text
|
|
509
|
+
cacheHitRate = cachedTokens / promptTokens
|
|
510
|
+
```
|
|
511
|
+
|
|
512
|
+
implemented by `mimoHitRate()`. `hitRatePct()` itself is left untouched so other
|
|
513
|
+
providers are unaffected.
|
|
514
|
+
|
|
352
515
|
### Important metric distinction
|
|
353
516
|
|
|
354
517
|
These ratios answer different questions.
|
|
@@ -361,7 +524,12 @@ These ratios answer different questions.
|
|
|
361
524
|
|
|
362
525
|
> How much of the total prompt-token accounting was represented by cached reads?
|
|
363
526
|
|
|
364
|
-
|
|
527
|
+
`cachedTokens / promptTokens` (MiMo) answers:
|
|
528
|
+
|
|
529
|
+
> Of the prompt tokens the provider processed, what fraction was served from
|
|
530
|
+
> cache?
|
|
531
|
+
|
|
532
|
+
Do not treat these percentages as interchangeable.
|
|
365
533
|
|
|
366
534
|
---
|
|
367
535
|
|
|
@@ -444,6 +612,29 @@ Telemetry is intended to answer questions such as:
|
|
|
444
612
|
* Did a compaction occur?
|
|
445
613
|
* Which provider/model/policy was active?
|
|
446
614
|
* Did the GLM system stabilization actually change the observed prompt shape?
|
|
615
|
+
* Did MiMo's environment relocation fire (`mimo_system_env_relocated`)?
|
|
616
|
+
* Did MiMo's stable system prefix change (`mimo_system_prefix_changed`)?
|
|
617
|
+
* Did the MiMo provider change within a session (`mimo_provider_changed`)?
|
|
618
|
+
* What was MiMo's provider-reported cache hit rate (`cacheHitRate`)?
|
|
619
|
+
|
|
620
|
+
A MiMo usage record adds the provider-reported cache fields:
|
|
621
|
+
|
|
622
|
+
```json
|
|
623
|
+
{
|
|
624
|
+
"kind": "usage",
|
|
625
|
+
"sid": "session-id",
|
|
626
|
+
"ts": 1750000000000,
|
|
627
|
+
"policy": "mimo26",
|
|
628
|
+
"provider": "openrouter",
|
|
629
|
+
"model": "xiaomi/mimo-v2.6-flash",
|
|
630
|
+
"read": 47000,
|
|
631
|
+
"input": 3000,
|
|
632
|
+
"promptTokens": 50000,
|
|
633
|
+
"cachedTokens": 47000,
|
|
634
|
+
"cacheHitRate": 94,
|
|
635
|
+
"stickySessionId": "mimo-ses-0123456789abcdef"
|
|
636
|
+
}
|
|
637
|
+
```
|
|
447
638
|
|
|
448
639
|
|
|
449
640
|
# Configuration
|
|
@@ -473,6 +664,12 @@ The default configuration is:
|
|
|
473
664
|
"enabled": true,
|
|
474
665
|
"stabilizeSystem": true,
|
|
475
666
|
"preserveThinkingIntegrity": true
|
|
667
|
+
},
|
|
668
|
+
"mimo26": {
|
|
669
|
+
"enabled": true,
|
|
670
|
+
"stabilizeSystem": true,
|
|
671
|
+
"stickySession": true,
|
|
672
|
+
"preserveThinkingIntegrity": true
|
|
476
673
|
}
|
|
477
674
|
}
|
|
478
675
|
}
|
|
@@ -623,6 +820,45 @@ The reasoning instrumentation is intended to identify anomalies such as:
|
|
|
623
820
|
It is diagnostic rather than a reason to rewrite or fabricate reasoning content. The implementation maps these conditions to explicit diagnostic reasons.
|
|
624
821
|
|
|
625
822
|
|
|
823
|
+
# MiMo-V2.6 configuration
|
|
824
|
+
|
|
825
|
+
```json
|
|
826
|
+
{
|
|
827
|
+
"mimo26": {
|
|
828
|
+
"enabled": true,
|
|
829
|
+
"stabilizeSystem": true,
|
|
830
|
+
"stickySession": true,
|
|
831
|
+
"preserveThinkingIntegrity": true
|
|
832
|
+
}
|
|
833
|
+
}
|
|
834
|
+
```
|
|
835
|
+
|
|
836
|
+
### `enabled`
|
|
837
|
+
|
|
838
|
+
Enables the MiMo-V2.6 policy.
|
|
839
|
+
|
|
840
|
+
### `stabilizeSystem`
|
|
841
|
+
|
|
842
|
+
Enables relocation of the volatile `<env>` section to the system-prompt tail
|
|
843
|
+
(same narrow, content-preserving transformation as GLM-5.3).
|
|
844
|
+
|
|
845
|
+
### `stickySession`
|
|
846
|
+
|
|
847
|
+
Gates derivation/recording of the OpenRouter sticky-session id
|
|
848
|
+
(`mimoSessionIdFor`). The id is recorded as telemetry; it is **not** injected
|
|
849
|
+
into the request because this runtime exposes no verified OpenRouter top-level
|
|
850
|
+
`session_id` path. See "OpenRouter sticky session — derived but not injected".
|
|
851
|
+
|
|
852
|
+
### `preserveThinkingIntegrity`
|
|
853
|
+
|
|
854
|
+
Enables reasoning diagnostics as instrumentation. It never rewrites, duplicates,
|
|
855
|
+
reorders, or re-injects reasoning content, and it is not a cache requirement.
|
|
856
|
+
|
|
857
|
+
No `cacheBlockSize`, `cacheTTL`, `cacheBreakpoint`, or `minimumCacheTokens`
|
|
858
|
+
knobs are exposed: those values are not established by authoritative V2.6
|
|
859
|
+
documentation.
|
|
860
|
+
|
|
861
|
+
|
|
626
862
|
# Model detection
|
|
627
863
|
|
|
628
864
|
The plugin classifies requests into:
|
|
@@ -631,6 +867,7 @@ The plugin classifies requests into:
|
|
|
631
867
|
deepseek
|
|
632
868
|
gpt56
|
|
633
869
|
glm53
|
|
870
|
+
mimo26
|
|
634
871
|
neutral
|
|
635
872
|
```
|
|
636
873
|
|
|
@@ -639,9 +876,13 @@ The model detector recognizes:
|
|
|
639
876
|
* DeepSeek model/provider identifiers
|
|
640
877
|
* GPT-5.6 variants
|
|
641
878
|
* GLM-5.3 variants
|
|
879
|
+
* MiMo-V2.6 Flash and Pro (`xiaomi/mimo-v2.6-flash`, `mimo-v2.6-pro`, ...)
|
|
642
880
|
|
|
643
881
|
GPT-5.6 has an additional OpenAI/Azure-context check so a string containing `gpt-5.6` does not automatically cause GPT-specific fields to be sent to an unrelated endpoint.
|
|
644
882
|
|
|
883
|
+
MiMo detection targets exactly Flash and Pro: it excludes `mimo-v2.5`,
|
|
884
|
+
`mimo-v2.5-pro`, and `mimo-v2.6-pro-ultraspeed`.
|
|
885
|
+
|
|
645
886
|
Unknown models use the neutral policy.
|
|
646
887
|
|
|
647
888
|
Neutral means:
|
|
@@ -657,7 +898,7 @@ This plugin is compatible with OpenRouter because the cache policy is based on t
|
|
|
657
898
|
|
|
658
899
|
For cache-sensitive workloads, provider stability remains important.
|
|
659
900
|
|
|
660
|
-
The plugin does not attempt to compensate for provider switching by rewriting prompts.
|
|
901
|
+
The plugin does not attempt to compensate for provider switching by rewriting prompts. It records MiMo provider identity and provider-switch diagnostics so routing instability is at least observable.
|
|
661
902
|
|
|
662
903
|
For that reason, a stable provider route is preferable when your goal is to measure and maximize prefix reuse.
|
|
663
904
|
|
|
@@ -964,6 +1205,12 @@ GPT-5.6:
|
|
|
964
1205
|
|
|
965
1206
|
GLM-5.3:
|
|
966
1207
|
volatile env block relocated when eligible
|
|
1208
|
+
|
|
1209
|
+
MiMo-V2.6:
|
|
1210
|
+
volatile env block relocated when eligible
|
|
1211
|
+
no GPT/GLM-only cache fields present
|
|
1212
|
+
no OpenRouter top-level session_id injected (unsupported by this runtime)
|
|
1213
|
+
telemetry carries provider/model/promptTokens/cachedTokens/cacheHitRate
|
|
967
1214
|
```
|
|
968
1215
|
|
|
969
1216
|
---
|
|
@@ -1006,6 +1253,42 @@ The relevant block must contain the expected beginning and closing marker, and t
|
|
|
1006
1253
|
|
|
1007
1254
|
---
|
|
1008
1255
|
|
|
1256
|
+
## MiMo-V2.6 prompt is not being changed
|
|
1257
|
+
|
|
1258
|
+
MiMo uses the same eligibility rules as GLM-5.3: exactly one system string, both
|
|
1259
|
+
`<env>` markers present, block identified unambiguously, and
|
|
1260
|
+
`mimo26.stabilizeSystem` enabled. If the block is already at the tail, the
|
|
1261
|
+
operation is a no-op.
|
|
1262
|
+
|
|
1263
|
+
---
|
|
1264
|
+
|
|
1265
|
+
## MiMo provider is not classified as `mimo26`
|
|
1266
|
+
|
|
1267
|
+
Verify the model identifier is exactly Flash or Pro:
|
|
1268
|
+
|
|
1269
|
+
```text
|
|
1270
|
+
mimo-v2.6-flash
|
|
1271
|
+
mimo-v2.6-pro
|
|
1272
|
+
xiaomi/mimo-v2.6-flash
|
|
1273
|
+
xiaomi/mimo-v2.6-pro
|
|
1274
|
+
```
|
|
1275
|
+
|
|
1276
|
+
`mimo-v2.5`, `mimo-v2.5-pro`, and `mimo-v2.6-pro-ultraspeed` are intentionally
|
|
1277
|
+
not matched.
|
|
1278
|
+
|
|
1279
|
+
---
|
|
1280
|
+
|
|
1281
|
+
## No OpenRouter `session_id` is sent for MiMo
|
|
1282
|
+
|
|
1283
|
+
This is expected. The installed OpenCode runtime's OpenRouter request adapter
|
|
1284
|
+
forwards only `usage`, `reasoning`, and `prompt_cache_key` from provider options
|
|
1285
|
+
and exposes no top-level `session_id` path. The plugin derives a stable
|
|
1286
|
+
`stickySessionId` and records it as telemetry, but does not inject it rather than
|
|
1287
|
+
send an unsupported field. This may change if a future runtime exposes a verified
|
|
1288
|
+
path.
|
|
1289
|
+
|
|
1290
|
+
---
|
|
1291
|
+
|
|
1009
1292
|
## Metrics file is missing
|
|
1010
1293
|
|
|
1011
1294
|
Telemetry is best-effort.
|
|
@@ -1030,6 +1313,7 @@ The current implementation is intentionally conservative:
|
|
|
1030
1313
|
DeepSeek -> preserve and measure
|
|
1031
1314
|
GPT-5.6 -> configure cache controls
|
|
1032
1315
|
GLM-5.3 -> isolate volatile prompt content
|
|
1316
|
+
MiMo-V2.6 -> stabilize prefix + observe provider/cache reality
|
|
1033
1317
|
```
|
|
1034
1318
|
|
|
1035
1319
|
That separation is the core design of the project.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "opencode-cache-engine",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.3.0",
|
|
4
4
|
"private": false,
|
|
5
5
|
"description": "Provider-aware prompt-cache optimization and telemetry for OpenCode",
|
|
6
6
|
"keywords": [
|
|
@@ -9,7 +9,8 @@
|
|
|
9
9
|
"prompt-cache",
|
|
10
10
|
"deepseek",
|
|
11
11
|
"glm",
|
|
12
|
-
"gpt"
|
|
12
|
+
"gpt",
|
|
13
|
+
"mimo"
|
|
13
14
|
],
|
|
14
15
|
"homepage": "https://github.com/AlexJoaquimPereira/opencode-cache-engine#readme",
|
|
15
16
|
"bugs": {
|
|
@@ -5,9 +5,10 @@
|
|
|
5
5
|
// TypeScript compiler. The plugin entry (cache-engine.ts) imports this module.
|
|
6
6
|
//
|
|
7
7
|
// This module is PROVIDER-AWARE: it classifies a model into a cache-policy
|
|
8
|
-
// family (deepseek | gpt56 | glm53 | neutral) and exposes small pure
|
|
9
|
-
// each family's strategy. The plugin entry (cache-engine.ts) remains
|
|
10
|
-
// place that touches OpenCode hooks; every decision here is testable in
|
|
8
|
+
// family (deepseek | gpt56 | glm53 | mimo26 | neutral) and exposes small pure
|
|
9
|
+
// helpers for each family's strategy. The plugin entry (cache-engine.ts) remains
|
|
10
|
+
// the only place that touches OpenCode hooks; every decision here is testable in
|
|
11
|
+
// Node.
|
|
11
12
|
//
|
|
12
13
|
// Terminology note: these functions deal with the *observed* system/tool
|
|
13
14
|
// prefix shape. An observed change means the request's prefix bytes changed; it
|
|
@@ -35,6 +36,7 @@ export const DIGEST_TEMPLATE = `## Session digest (cache-stable continuation blo
|
|
|
35
36
|
export const POLICY_DEEPSEEK = "deepseek"
|
|
36
37
|
export const POLICY_GPT56 = "gpt56"
|
|
37
38
|
export const POLICY_GLM53 = "glm53"
|
|
39
|
+
export const POLICY_MIMO26 = "mimo26"
|
|
38
40
|
export const POLICY_NEUTRAL = "neutral"
|
|
39
41
|
|
|
40
42
|
export const GPT56_DEFAULT_TTL = "30m"
|
|
@@ -65,6 +67,15 @@ function defaultPolicies() {
|
|
|
65
67
|
stabilizeSystem: true,
|
|
66
68
|
preserveThinkingIntegrity: true,
|
|
67
69
|
},
|
|
70
|
+
mimo26: {
|
|
71
|
+
enabled: true,
|
|
72
|
+
stabilizeSystem: true,
|
|
73
|
+
stickySession: true,
|
|
74
|
+
// MiMo reasoning diagnostics are instrumentation only. Unlike GLM
|
|
75
|
+
// preserved thinking, there is no evidence that MiMo prompt-cache reuse
|
|
76
|
+
// depends on reasoning replay, so this never rewrites reasoning content.
|
|
77
|
+
preserveThinkingIntegrity: true,
|
|
78
|
+
},
|
|
68
79
|
}
|
|
69
80
|
}
|
|
70
81
|
|
|
@@ -118,7 +129,7 @@ export function parseConfig(raw, env) {
|
|
|
118
129
|
if (typeof raw.logPrefixChanges === "boolean") cfg.logPrefixChanges = raw.logPrefixChanges
|
|
119
130
|
if (raw.policies && typeof raw.policies === "object") {
|
|
120
131
|
const d = defaultPolicies()
|
|
121
|
-
for (const fam of ["deepseek", "gpt56", "glm53"]) {
|
|
132
|
+
for (const fam of ["deepseek", "gpt56", "glm53", "mimo26"]) {
|
|
122
133
|
if (raw.policies[fam]) cfg.policies[fam] = parsePolicy(raw.policies[fam], d[fam])
|
|
123
134
|
}
|
|
124
135
|
}
|
|
@@ -232,6 +243,10 @@ function isOpenAIish(s) {
|
|
|
232
243
|
const GPT56_RE = /gpt-5\.6(?![\d.])/i
|
|
233
244
|
// GLM 5.3 family only (not glm-4.x / glm-4.6 etc).
|
|
234
245
|
const GLM53_RE = /glm-5\.3(?![\d.])/i
|
|
246
|
+
// Xiaomi MiMo V2.6 explicitly targets Flash + Pro only. The trailing
|
|
247
|
+
// (?![\w-]) guard prevents matching a hypothetical "mimo-v2.6-pro-ultraspeed"
|
|
248
|
+
// or "mimo-v2.6-flashx", and the v2\.6 literal excludes V2.5 / V2.
|
|
249
|
+
const MIMO26_RE = /mimo-v2\.6-(flash|pro)(?![\w-])/i
|
|
235
250
|
const DEEPSEEK_RE = /deepseek/i
|
|
236
251
|
|
|
237
252
|
// Pure classifier. Returns one of the POLICY_* keys. `model` may be a full
|
|
@@ -242,6 +257,7 @@ export function detectPolicy(model) {
|
|
|
242
257
|
if (!s.slug) return POLICY_NEUTRAL
|
|
243
258
|
if (GPT56_RE.test(s.slug) && isOpenAIish(s)) return POLICY_GPT56
|
|
244
259
|
if (GLM53_RE.test(s.slug)) return POLICY_GLM53
|
|
260
|
+
if (MIMO26_RE.test(s.slug)) return POLICY_MIMO26
|
|
245
261
|
if (DEEPSEEK_RE.test(s.slug) || DEEPSEEK_RE.test(s.providerID)) return POLICY_DEEPSEEK
|
|
246
262
|
return POLICY_NEUTRAL
|
|
247
263
|
}
|
|
@@ -284,16 +300,17 @@ export function gptCacheOptionsDelta(existingOptions, { key, mode = GPT56_DEFAUL
|
|
|
284
300
|
|
|
285
301
|
// The OpenCode system string begins with the agent prompt, then an env block
|
|
286
302
|
// ("You are powered by the model named ... Today's date: ... </env>") whose
|
|
287
|
-
// only per-day volatile byte is the date line. For GLM-5.3
|
|
288
|
-
// whole identifiable env block to the END of the system string so
|
|
289
|
-
// change only invalidates the tail of the prompt, leaving the long
|
|
290
|
-
// prefix intact. Content is preserved byte-for-byte (only position
|
|
303
|
+
// only per-day volatile byte is the date line. For GLM-5.3 and MiMo-V2.6 we
|
|
304
|
+
// relocate that whole identifiable env block to the END of the system string so
|
|
305
|
+
// a daily date change only invalidates the tail of the prompt, leaving the long
|
|
306
|
+
// stable prefix intact. Content is preserved byte-for-byte (only position
|
|
307
|
+
// changes).
|
|
291
308
|
//
|
|
292
309
|
// Returns { text, changed }. When the block cannot be identified unambiguously,
|
|
293
310
|
// returns the input unchanged (changed:false). This is a content-preserving
|
|
294
311
|
// reorder of clearly volatile metadata only -- it never reorders arbitrary
|
|
295
312
|
// instructions. This function is ONLY applied when the caller has already
|
|
296
|
-
// classified the model
|
|
313
|
+
// classified the model into a family that opts into system stabilization.
|
|
297
314
|
export function relocateVolatileEnvBlock(text) {
|
|
298
315
|
if (typeof text !== "string") return { text, changed: false }
|
|
299
316
|
const START = "You are powered by the model named "
|
|
@@ -438,6 +455,56 @@ export function glmHitRatio(read, write, input) {
|
|
|
438
455
|
return Math.round((100 * read) / denom)
|
|
439
456
|
}
|
|
440
457
|
|
|
458
|
+
// ---------------------------------------------------------------------------
|
|
459
|
+
// MiMo-V2.6 cache metrics + sticky-session identity
|
|
460
|
+
//
|
|
461
|
+
// MiMo caching is provider-managed (implicit context caching). Xiaomi documents
|
|
462
|
+
// usage.prompt_tokens_details.cached_tokens as the number of PROMPT tokens
|
|
463
|
+
// served from cache and prompt_tokens as the total prompt-token count, so the
|
|
464
|
+
// authoritative cache metric is cachedTokens / promptTokens -- NOT the
|
|
465
|
+
// read/(read+write) form used by other families. `hitRatePct` is intentionally
|
|
466
|
+
// left untouched so existing providers are unaffected.
|
|
467
|
+
//
|
|
468
|
+
// The runtime exposes `Message.info.tokens` as { input, output, cache:{read,
|
|
469
|
+
// write} } where `input` is the NON-cached prompt input and `cache.read` is the
|
|
470
|
+
// cached prompt input. Total prompt tokens are therefore derived as
|
|
471
|
+
// read + input (cache.write is a separate accounting bucket and is NOT folded
|
|
472
|
+
// in). We never fabricate cache-write values.
|
|
473
|
+
// ---------------------------------------------------------------------------
|
|
474
|
+
|
|
475
|
+
// cachedTokens / promptTokens, rounded to a percentage. Returns null when the
|
|
476
|
+
// denominator is unknown/zero or the inputs are not finite numbers, so no
|
|
477
|
+
// fabricated hit rate is ever emitted.
|
|
478
|
+
export function mimoHitRate(cachedTokens, promptTokens) {
|
|
479
|
+
if (!Number.isFinite(cachedTokens) || !Number.isFinite(promptTokens)) return null
|
|
480
|
+
if (promptTokens <= 0 || cachedTokens < 0) return null
|
|
481
|
+
return Math.round((100 * cachedTokens) / promptTokens)
|
|
482
|
+
}
|
|
483
|
+
|
|
484
|
+
// Derive a stable, session-scoped identifier suitable for OpenRouter's
|
|
485
|
+
// documented `session_id` sticky-routing key. Pure function of the OpenCode
|
|
486
|
+
// session id only: identical sessions map to identical ids, distinct sessions
|
|
487
|
+
// map to distinct ids, and transient request contents cannot influence it.
|
|
488
|
+
// The value is printable, contains no whitespace, and is far below the 256-char
|
|
489
|
+
// cap (25 chars). NOTE: this runtime's OpenRouter request adapter does not emit
|
|
490
|
+
// a top-level `session_id` (it forwards only usage/reasoning/prompt_cache_key),
|
|
491
|
+
// so this id is currently recorded as telemetry only and is never injected.
|
|
492
|
+
export function mimoSessionIdFor(sessionID) {
|
|
493
|
+
if (typeof sessionID !== "string" || sessionID.length === 0) return null
|
|
494
|
+
return `mimo-ses-${shorthash(sessionID)}`
|
|
495
|
+
}
|
|
496
|
+
|
|
497
|
+
// Detect a provider switch within the same session. `previous` and `current`
|
|
498
|
+
// are {providerID, modelID} observations. Returns {changed:false} until two
|
|
499
|
+
// real observations exist; a change is only reported when both are known and
|
|
500
|
+
// the providerID differs. Never forces or overrides provider selection.
|
|
501
|
+
export function providerChangeEvent(previous, current) {
|
|
502
|
+
const prev = previous && typeof previous.providerID === "string" ? previous.providerID : null
|
|
503
|
+
const cur = current && typeof current.providerID === "string" ? current.providerID : null
|
|
504
|
+
if (prev == null || cur == null || prev === cur) return { changed: false, from: null, to: null }
|
|
505
|
+
return { changed: true, from: previous, to: current }
|
|
506
|
+
}
|
|
507
|
+
|
|
441
508
|
// Decide whether a `usage` record should be emitted for an aggregation sample.
|
|
442
509
|
// We must not fabricate a zero-valued cache event merely because the session
|
|
443
510
|
// became idle: a sample only counts when at least one assistant message with
|