opencode-cache-engine 0.2.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -12,11 +12,12 @@ The server target handles cache optimization, prompt-shape diagnostics, compacti
12
12
 
13
13
  `CacheEngine` is an OpenCode plugin designed for long-running agent sessions where prompt-cache efficiency affects both latency and cost. It keeps the harness conservative for providers whose cache behavior is already automatic, while applying provider-specific optimizations where the provider exposes useful cache controls or where prompt structure can be safely improved.
14
14
 
15
- The plugin currently has three cache-policy families:
15
+ The plugin currently has four cache-policy families:
16
16
 
17
17
  * **DeepSeek V4.1 Flash** — passive cache-stability and observability
18
18
  * **GPT-5.6 Luna** — active cache-control configuration
19
19
  * **GLM-5.3 Flash** — conservative system-prompt stabilization
20
+ * **MiMo-V2.6 (Flash / Pro)** — prefix stability and OpenRouter session-affinity diagnostics
20
21
 
21
22
  The central design principle is:
22
23
 
@@ -35,8 +36,9 @@ It:
35
36
  4. Records provider-reported cache token usage.
36
37
  5. Adds a deterministic compaction continuation block.
37
38
  6. Applies GPT-5.6 cache-control metadata.
38
- 7. Applies the GLM-5.3 volatile-environment relocation.
39
+ 7. Applies the GLM-5.3 and MiMo-V2.6 volatile-environment relocation.
39
40
  8. Records diagnostics that help determine whether prompt-shape changes correlate with cache behavior.
41
+ 9. Records MiMo-V2.6 provider identity and provider-switch diagnostics.
40
42
 
41
43
  The plugin deliberately avoids pretending that a local hash is proof of a provider cache hit. Provider-reported token usage remains the authoritative signal.
42
44
 
@@ -224,15 +226,167 @@ It only occurs when:
224
226
  The plugin does not arbitrarily rearrange unrelated prompt content.
225
227
 
226
228
 
227
- # Prompt-cache strategy
229
+ ## MiMo-V2.6 (Flash / Pro)
230
+
231
+ ### Policy: prefix stability + OpenRouter session affinity
232
+
233
+ MiMo-V2.6 is Xiaomi's current model family. The plugin targets exactly two
234
+ identifiers:
235
+
236
+ * `xiaomi/mimo-v2.6-flash` / `mimo-v2.6-flash`
237
+ * `xiaomi/mimo-v2.6-pro` / `mimo-v2.6-pro`
238
+
239
+ Detection also tolerates `provider/model` shapes where `api.id` contains those
240
+ slugs. It deliberately does **not** match `mimo-v2.5`, `mimo-v2.5-pro`,
241
+ `mimo-v2.6-pro-ultraspeed`, or unrelated MiMo models.
242
+
243
+ ### Implicit context caching
244
+
245
+ Xiaomi documents context caching for both V2.6 Flash and Pro, and exposes
246
+ `usage.prompt_tokens_details.cached_tokens` as the number of prompt tokens
247
+ served from cache. The V2.6 API documents implicit context caching, not a
248
+ user-supplied cache key or explicit breakpoint.
249
+
250
+ Accordingly the plugin **injects no cache-control parameter** for MiMo. It does
251
+ not send `promptCacheKey`, `cacheControl`, `cacheBreakpoint`, or `ttl`.
252
+ Implicit caching is the default assumption.
253
+
254
+ ### Environment-block stabilization
255
+
256
+ MiMo uses the same narrow, content-preserving transformation as GLM-5.3: the
257
+ identifiable volatile `<env>` block is relocated to the **tail** of the single
258
+ system string. Contents are preserved byte-for-byte; only position changes. This
259
+ keeps the large reusable prefix stable when only the environment/date changes.
260
+
261
+ The transformation is applied only when:
262
+
263
+ * the selected model is MiMo-V2.6 Flash/Pro
264
+ * `mimo26.stabilizeSystem` is `true`
265
+ * there is exactly one system string
266
+ * the expected `<env>` markers exist and the block is identified unambiguously
267
+ * the block is not already at the tail
268
+
269
+ ### No generic system-prompt freezing
270
+
271
+ MiMo-Code's own harness freezes its per-session system prefix. This plugin does
272
+ **not** copy that mechanism. System instructions can legitimately change because
273
+ of permissions, tools, agent mode, skills, MCP state, or project configuration;
274
+ a plugin-level snapshot must never override a legitimate change.
275
+
276
+ Instead the plugin:
277
+
278
+ * records a first-seen system baseline per session;
279
+ * computes the full system hash, stable prefix hash, and volatile suffix hash;
280
+ * records changes for MiMo sessions;
281
+ * allows the `<env>` relocation when that is the only identified volatility;
282
+ * reports other system changes diagnostically and never overwrites the new
283
+ content.
284
+
285
+ Explicit telemetry events:
286
+
287
+ * `mimo_system_env_relocated`
288
+ * `mimo_system_prefix_changed`
289
+
290
+ ### OpenRouter sticky session — derived but not injected
291
+
292
+ OpenRouter documents a top-level `session_id` request field for sticky provider
293
+ routing, which keeps a session's requests on the same upstream provider so
294
+ provider-side prompt caches stay warm.
295
+
296
+ The plugin includes a pure, session-scoped derivation (`mimoSessionIdFor`):
297
+ deterministic, distinct per session, printable/no-whitespace, and well under the
298
+ 256-character cap. However, **the derived id is not injected into requests**.
299
+
300
+ Rationale (verified against the installed runtime): OpenCode's OpenRouter
301
+ request adapter forwards only `usage`, `reasoning`, and `prompt_cache_key` from
302
+ provider options, and exposes no top-level `session_id` path. Adding an
303
+ unsupported field would be guessing, so the id is recorded as telemetry only,
304
+ and `mimo26.stickySession` currently gates that recording. If a future runtime
305
+ gains a verified `session_id` path, the helper is already in place.
306
+
307
+ Note that OpenCode itself sets `x-session-affinity` / `X-Session-Id` HTTP
308
+ headers for non-opencode providers, and can set a flat `promptCacheKey` for
309
+ OpenRouter when `setCacheKey: true` is configured. Those are HTTP routing
310
+ headers and an OpenAI-style cache key respectively — they are not OpenRouter's
311
+ documented body `session_id`.
312
+
313
+ ### MiMo cache metrics
314
+
315
+ MiMo caches are provider-managed, so the authoritative metric is provider
316
+ reported. For MiMo the plugin emits the preferred ratio:
317
+
318
+ ```text
319
+ cacheHitRate = cachedTokens / promptTokens
320
+ ```
321
+
322
+ This is intentionally **not** the `read / (read + write)` form used by other
323
+ families. It is not GLM's `read / (read + write + input)` either.
324
+
325
+ Derivation: the runtime exposes assistant tokens as `{ input, output,
326
+ cache:{ read, write } }`, where `input` is the non-cached prompt input and
327
+ `cache.read` is the cached prompt input. Total prompt tokens are therefore
328
+ derived as `read + input`, and `cachedTokens = read`. `cache.write` is a
329
+ separate accounting bucket and is not folded in; no cache-write value is
330
+ fabricated, and the ratio is `null` when `promptTokens` is zero.
331
+
332
+ A MiMo usage record looks conceptually like:
333
+
334
+ ```json
335
+ {
336
+ "kind": "usage",
337
+ "policy": "mimo26",
338
+ "provider": "openrouter",
339
+ "model": "xiaomi/mimo-v2.6-flash",
340
+ "promptTokens": 50000,
341
+ "cachedTokens": 47000,
342
+ "cacheHitRate": 94
343
+ }
344
+ ```
345
+
346
+ ### Provider-switch diagnostics
347
+
348
+ Because MiMo caches live at the provider side, a provider change within one
349
+ session can silently invalidate them. The plugin records provider identity on
350
+ every MiMo request and emits a `mimo_provider_changed` boundary event when the
351
+ OpenCode `providerID` changes within a session. It never forces or overrides the
352
+ user's provider selection.
353
+
354
+ Limitation: OpenRouter's *upstream* provider selection (for example
355
+ `xiaomi/fp8` vs `atlas-cloud/fp8`) is not exposed to plugins, so only the
356
+ OpenCode `providerID`/`modelID` are observable.
357
+
358
+ ### Reasoning / thinking
228
359
 
229
- The plugin uses three different strategies because cache mechanisms differ by provider.
360
+ MiMo-V2.6 supports deep thinking and reports reasoning tokens. The plugin does
361
+ not treat reasoning replay as a cache requirement: reasoning diagnostics are
362
+ instrumentation only, and the plugin never rewrites, duplicates, reorders, or
363
+ re-injects reasoning content, nor changes reasoning effort for caching.
230
364
 
231
- | Provider | Prompt text changed? | Cache metadata changed? | Main strategy |
232
- | ----------------- | -------------------: | ----------------------: | --------------------------------- |
233
- | DeepSeek V4.1 Flash | No | No | Preserve stable harness + observe |
234
- | GPT-5.6 Luna | No | Yes | Stable cache key + cache options |
235
- | GLM-5.3 Flash | Yes, narrowly | No provider cache key | Isolate volatile system content |
365
+ ### Skill-catalog / history limitation
366
+
367
+ MiMo-Code moved skill catalogs out of repeatedly rewritten user messages and
368
+ toward the system tail. In this OpenCode runtime the skill guidance
369
+ (`<available_skills>`) and MCP instructions already live in the **system
370
+ prefix**, not in user-message history. The plugin therefore performs no
371
+ message-history rewrite. Skill/MCP changes simply appear as system-prefix changes
372
+ and are reported diagnostically; the message content is left untouched.
373
+
374
+
375
+ # Provider comparison
376
+
377
+ | Provider | Detection | Prompt text changed? | Cache metadata changed? | Primary cache signal |
378
+ | ------------------- | ---------------------------------- | --------------------------- | ---------------------------------- | --------------------------------------- |
379
+ | DeepSeek V4.1 Flash | `deepseek` | No | No | provider `cache.read`/`cache.write` |
380
+ | GPT-5.6 Luna | `gpt-5.6*` on OpenAI-ish endpoints | No | Yes: `prompt_cache_key` + options | provider cache tokens |
381
+ | GLM-5.3 Flash | `glm-5.3*` | Yes, narrowly (`<env>` tail) | No provider cache key | provider cache tokens (GLM ratio) |
382
+ | MiMo-V2.6 Flash/Pro | `mimo-v2.6-flash` / `mimo-v2.6-pro` | Yes, narrowly (`<env>` tail) | No: implicit caching only | `cached_tokens / prompt_tokens` |
383
+
384
+
385
+ # Prompt-cache strategy
386
+
387
+ The plugin uses different strategies because cache mechanisms differ by provider.
388
+ The table above summarises them; the essential point is the distinction between
389
+ *changing prompt text* and *changing cache metadata*.
236
390
 
237
391
  This distinction is fundamental.
238
392
 
@@ -349,6 +503,15 @@ read / (read + write + input)
349
503
 
350
504
  as implemented by `glmHitRatio()`.
351
505
 
506
+ For MiMo, the implementation uses the provider-documented prompt-cache ratio:
507
+
508
+ ```text
509
+ cacheHitRate = cachedTokens / promptTokens
510
+ ```
511
+
512
+ implemented by `mimoHitRate()`. `hitRatePct()` itself is left untouched so other
513
+ providers are unaffected.
514
+
352
515
  ### Important metric distinction
353
516
 
354
517
  These ratios answer different questions.
@@ -361,7 +524,12 @@ These ratios answer different questions.
361
524
 
362
525
  > How much of the total prompt-token accounting was represented by cached reads?
363
526
 
364
- Do not treat the two percentages as interchangeable.
527
+ `cachedTokens / promptTokens` (MiMo) answers:
528
+
529
+ > Of the prompt tokens the provider processed, what fraction was served from
530
+ > cache?
531
+
532
+ Do not treat these percentages as interchangeable.
365
533
 
366
534
  ---
367
535
 
@@ -444,6 +612,29 @@ Telemetry is intended to answer questions such as:
444
612
  * Did a compaction occur?
445
613
  * Which provider/model/policy was active?
446
614
  * Did the GLM system stabilization actually change the observed prompt shape?
615
+ * Did MiMo's environment relocation fire (`mimo_system_env_relocated`)?
616
+ * Did MiMo's stable system prefix change (`mimo_system_prefix_changed`)?
617
+ * Did the MiMo provider change within a session (`mimo_provider_changed`)?
618
+ * What was MiMo's provider-reported cache hit rate (`cacheHitRate`)?
619
+
620
+ A MiMo usage record adds the provider-reported cache fields:
621
+
622
+ ```json
623
+ {
624
+ "kind": "usage",
625
+ "sid": "session-id",
626
+ "ts": 1750000000000,
627
+ "policy": "mimo26",
628
+ "provider": "openrouter",
629
+ "model": "xiaomi/mimo-v2.6-flash",
630
+ "read": 47000,
631
+ "input": 3000,
632
+ "promptTokens": 50000,
633
+ "cachedTokens": 47000,
634
+ "cacheHitRate": 94,
635
+ "stickySessionId": "mimo-ses-0123456789abcdef"
636
+ }
637
+ ```
447
638
 
448
639
 
449
640
  # Configuration
@@ -473,6 +664,12 @@ The default configuration is:
473
664
  "enabled": true,
474
665
  "stabilizeSystem": true,
475
666
  "preserveThinkingIntegrity": true
667
+ },
668
+ "mimo26": {
669
+ "enabled": true,
670
+ "stabilizeSystem": true,
671
+ "stickySession": true,
672
+ "preserveThinkingIntegrity": true
476
673
  }
477
674
  }
478
675
  }
@@ -623,6 +820,45 @@ The reasoning instrumentation is intended to identify anomalies such as:
623
820
  It is diagnostic rather than a reason to rewrite or fabricate reasoning content. The implementation maps these conditions to explicit diagnostic reasons.
624
821
 
625
822
 
823
+ # MiMo-V2.6 configuration
824
+
825
+ ```json
826
+ {
827
+ "mimo26": {
828
+ "enabled": true,
829
+ "stabilizeSystem": true,
830
+ "stickySession": true,
831
+ "preserveThinkingIntegrity": true
832
+ }
833
+ }
834
+ ```
835
+
836
+ ### `enabled`
837
+
838
+ Enables the MiMo-V2.6 policy.
839
+
840
+ ### `stabilizeSystem`
841
+
842
+ Enables relocation of the volatile `<env>` section to the system-prompt tail
843
+ (same narrow, content-preserving transformation as GLM-5.3).
844
+
845
+ ### `stickySession`
846
+
847
+ Gates derivation/recording of the OpenRouter sticky-session id
848
+ (`mimoSessionIdFor`). The id is recorded as telemetry; it is **not** injected
849
+ into the request because this runtime exposes no verified OpenRouter top-level
850
+ `session_id` path. See "OpenRouter sticky session — derived but not injected".
851
+
852
+ ### `preserveThinkingIntegrity`
853
+
854
+ Enables reasoning diagnostics as instrumentation. It never rewrites, duplicates,
855
+ reorders, or re-injects reasoning content, and it is not a cache requirement.
856
+
857
+ No `cacheBlockSize`, `cacheTTL`, `cacheBreakpoint`, or `minimumCacheTokens`
858
+ knobs are exposed: those values are not established by authoritative V2.6
859
+ documentation.
860
+
861
+
626
862
  # Model detection
627
863
 
628
864
  The plugin classifies requests into:
@@ -631,6 +867,7 @@ The plugin classifies requests into:
631
867
  deepseek
632
868
  gpt56
633
869
  glm53
870
+ mimo26
634
871
  neutral
635
872
  ```
636
873
 
@@ -639,9 +876,13 @@ The model detector recognizes:
639
876
  * DeepSeek model/provider identifiers
640
877
  * GPT-5.6 variants
641
878
  * GLM-5.3 variants
879
+ * MiMo-V2.6 Flash and Pro (`xiaomi/mimo-v2.6-flash`, `mimo-v2.6-pro`, ...)
642
880
 
643
881
  GPT-5.6 has an additional OpenAI/Azure-context check so a string containing `gpt-5.6` does not automatically cause GPT-specific fields to be sent to an unrelated endpoint.
644
882
 
883
+ MiMo detection targets exactly Flash and Pro: it excludes `mimo-v2.5`,
884
+ `mimo-v2.5-pro`, and `mimo-v2.6-pro-ultraspeed`.
885
+
645
886
  Unknown models use the neutral policy.
646
887
 
647
888
  Neutral means:
@@ -657,7 +898,7 @@ This plugin is compatible with OpenRouter because the cache policy is based on t
657
898
 
658
899
  For cache-sensitive workloads, provider stability remains important.
659
900
 
660
- The plugin does not attempt to compensate for provider switching by rewriting prompts.
901
+ The plugin does not attempt to compensate for provider switching by rewriting prompts. It records MiMo provider identity and provider-switch diagnostics so routing instability is at least observable.
661
902
 
662
903
  For that reason, a stable provider route is preferable when your goal is to measure and maximize prefix reuse.
663
904
 
@@ -964,6 +1205,12 @@ GPT-5.6:
964
1205
 
965
1206
  GLM-5.3:
966
1207
  volatile env block relocated when eligible
1208
+
1209
+ MiMo-V2.6:
1210
+ volatile env block relocated when eligible
1211
+ no GPT/GLM-only cache fields present
1212
+ no OpenRouter top-level session_id injected (unsupported by this runtime)
1213
+ telemetry carries provider/model/promptTokens/cachedTokens/cacheHitRate
967
1214
  ```
968
1215
 
969
1216
  ---
@@ -1006,6 +1253,42 @@ The relevant block must contain the expected beginning and closing marker, and t
1006
1253
 
1007
1254
  ---
1008
1255
 
1256
+ ## MiMo-V2.6 prompt is not being changed
1257
+
1258
+ MiMo uses the same eligibility rules as GLM-5.3: exactly one system string, both
1259
+ `<env>` markers present, block identified unambiguously, and
1260
+ `mimo26.stabilizeSystem` enabled. If the block is already at the tail, the
1261
+ operation is a no-op.
1262
+
1263
+ ---
1264
+
1265
+ ## MiMo provider is not classified as `mimo26`
1266
+
1267
+ Verify the model identifier is exactly Flash or Pro:
1268
+
1269
+ ```text
1270
+ mimo-v2.6-flash
1271
+ mimo-v2.6-pro
1272
+ xiaomi/mimo-v2.6-flash
1273
+ xiaomi/mimo-v2.6-pro
1274
+ ```
1275
+
1276
+ `mimo-v2.5`, `mimo-v2.5-pro`, and `mimo-v2.6-pro-ultraspeed` are intentionally
1277
+ not matched.
1278
+
1279
+ ---
1280
+
1281
+ ## No OpenRouter `session_id` is sent for MiMo
1282
+
1283
+ This is expected. The installed OpenCode runtime's OpenRouter request adapter
1284
+ forwards only `usage`, `reasoning`, and `prompt_cache_key` from provider options
1285
+ and exposes no top-level `session_id` path. The plugin derives a stable
1286
+ `stickySessionId` and records it as telemetry, but does not inject it rather than
1287
+ send an unsupported field. This may change if a future runtime exposes a verified
1288
+ path.
1289
+
1290
+ ---
1291
+
1009
1292
  ## Metrics file is missing
1010
1293
 
1011
1294
  Telemetry is best-effort.
@@ -1030,6 +1313,7 @@ The current implementation is intentionally conservative:
1030
1313
  DeepSeek -> preserve and measure
1031
1314
  GPT-5.6 -> configure cache controls
1032
1315
  GLM-5.3 -> isolate volatile prompt content
1316
+ MiMo-V2.6 -> stabilize prefix + observe provider/cache reality
1033
1317
  ```
1034
1318
 
1035
1319
  That separation is the core design of the project.
@@ -20,6 +20,12 @@
20
20
  "enabled": true,
21
21
  "stabilizeSystem": true,
22
22
  "preserveThinkingIntegrity": true
23
+ },
24
+ "mimo26": {
25
+ "enabled": true,
26
+ "stabilizeSystem": true,
27
+ "stickySession": true,
28
+ "preserveThinkingIntegrity": true
23
29
  }
24
30
  }
25
31
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "opencode-cache-engine",
3
- "version": "0.2.0",
3
+ "version": "0.3.0",
4
4
  "private": false,
5
5
  "description": "Provider-aware prompt-cache optimization and telemetry for OpenCode",
6
6
  "keywords": [
@@ -9,7 +9,8 @@
9
9
  "prompt-cache",
10
10
  "deepseek",
11
11
  "glm",
12
- "gpt"
12
+ "gpt",
13
+ "mimo"
13
14
  ],
14
15
  "homepage": "https://github.com/AlexJoaquimPereira/opencode-cache-engine#readme",
15
16
  "bugs": {
@@ -5,9 +5,10 @@
5
5
  // TypeScript compiler. The plugin entry (cache-engine.ts) imports this module.
6
6
  //
7
7
  // This module is PROVIDER-AWARE: it classifies a model into a cache-policy
8
- // family (deepseek | gpt56 | glm53 | neutral) and exposes small pure helpers for
9
- // each family's strategy. The plugin entry (cache-engine.ts) remains the only
10
- // place that touches OpenCode hooks; every decision here is testable in Node.
8
+ // family (deepseek | gpt56 | glm53 | mimo26 | neutral) and exposes small pure
9
+ // helpers for each family's strategy. The plugin entry (cache-engine.ts) remains
10
+ // the only place that touches OpenCode hooks; every decision here is testable in
11
+ // Node.
11
12
  //
12
13
  // Terminology note: these functions deal with the *observed* system/tool
13
14
  // prefix shape. An observed change means the request's prefix bytes changed; it
@@ -35,6 +36,7 @@ export const DIGEST_TEMPLATE = `## Session digest (cache-stable continuation blo
35
36
  export const POLICY_DEEPSEEK = "deepseek"
36
37
  export const POLICY_GPT56 = "gpt56"
37
38
  export const POLICY_GLM53 = "glm53"
39
+ export const POLICY_MIMO26 = "mimo26"
38
40
  export const POLICY_NEUTRAL = "neutral"
39
41
 
40
42
  export const GPT56_DEFAULT_TTL = "30m"
@@ -65,6 +67,15 @@ function defaultPolicies() {
65
67
  stabilizeSystem: true,
66
68
  preserveThinkingIntegrity: true,
67
69
  },
70
+ mimo26: {
71
+ enabled: true,
72
+ stabilizeSystem: true,
73
+ stickySession: true,
74
+ // MiMo reasoning diagnostics are instrumentation only. Unlike GLM
75
+ // preserved thinking, there is no evidence that MiMo prompt-cache reuse
76
+ // depends on reasoning replay, so this never rewrites reasoning content.
77
+ preserveThinkingIntegrity: true,
78
+ },
68
79
  }
69
80
  }
70
81
 
@@ -118,7 +129,7 @@ export function parseConfig(raw, env) {
118
129
  if (typeof raw.logPrefixChanges === "boolean") cfg.logPrefixChanges = raw.logPrefixChanges
119
130
  if (raw.policies && typeof raw.policies === "object") {
120
131
  const d = defaultPolicies()
121
- for (const fam of ["deepseek", "gpt56", "glm53"]) {
132
+ for (const fam of ["deepseek", "gpt56", "glm53", "mimo26"]) {
122
133
  if (raw.policies[fam]) cfg.policies[fam] = parsePolicy(raw.policies[fam], d[fam])
123
134
  }
124
135
  }
@@ -232,6 +243,10 @@ function isOpenAIish(s) {
232
243
  const GPT56_RE = /gpt-5\.6(?![\d.])/i
233
244
  // GLM 5.3 family only (not glm-4.x / glm-4.6 etc).
234
245
  const GLM53_RE = /glm-5\.3(?![\d.])/i
246
+ // Xiaomi MiMo V2.6 explicitly targets Flash + Pro only. The trailing
247
+ // (?![\w-]) guard prevents matching a hypothetical "mimo-v2.6-pro-ultraspeed"
248
+ // or "mimo-v2.6-flashx", and the v2\.6 literal excludes V2.5 / V2.
249
+ const MIMO26_RE = /mimo-v2\.6-(flash|pro)(?![\w-])/i
235
250
  const DEEPSEEK_RE = /deepseek/i
236
251
 
237
252
  // Pure classifier. Returns one of the POLICY_* keys. `model` may be a full
@@ -242,6 +257,7 @@ export function detectPolicy(model) {
242
257
  if (!s.slug) return POLICY_NEUTRAL
243
258
  if (GPT56_RE.test(s.slug) && isOpenAIish(s)) return POLICY_GPT56
244
259
  if (GLM53_RE.test(s.slug)) return POLICY_GLM53
260
+ if (MIMO26_RE.test(s.slug)) return POLICY_MIMO26
245
261
  if (DEEPSEEK_RE.test(s.slug) || DEEPSEEK_RE.test(s.providerID)) return POLICY_DEEPSEEK
246
262
  return POLICY_NEUTRAL
247
263
  }
@@ -284,16 +300,17 @@ export function gptCacheOptionsDelta(existingOptions, { key, mode = GPT56_DEFAUL
284
300
 
285
301
  // The OpenCode system string begins with the agent prompt, then an env block
286
302
  // ("You are powered by the model named ... Today's date: ... </env>") whose
287
- // only per-day volatile byte is the date line. For GLM-5.3 we relocate that
288
- // whole identifiable env block to the END of the system string so a daily date
289
- // change only invalidates the tail of the prompt, leaving the long stable
290
- // prefix intact. Content is preserved byte-for-byte (only position changes).
303
+ // only per-day volatile byte is the date line. For GLM-5.3 and MiMo-V2.6 we
304
+ // relocate that whole identifiable env block to the END of the system string so
305
+ // a daily date change only invalidates the tail of the prompt, leaving the long
306
+ // stable prefix intact. Content is preserved byte-for-byte (only position
307
+ // changes).
291
308
  //
292
309
  // Returns { text, changed }. When the block cannot be identified unambiguously,
293
310
  // returns the input unchanged (changed:false). This is a content-preserving
294
311
  // reorder of clearly volatile metadata only -- it never reorders arbitrary
295
312
  // instructions. This function is ONLY applied when the caller has already
296
- // classified the model as GLM-5.3.
313
+ // classified the model into a family that opts into system stabilization.
297
314
  export function relocateVolatileEnvBlock(text) {
298
315
  if (typeof text !== "string") return { text, changed: false }
299
316
  const START = "You are powered by the model named "
@@ -438,6 +455,56 @@ export function glmHitRatio(read, write, input) {
438
455
  return Math.round((100 * read) / denom)
439
456
  }
440
457
 
458
+ // ---------------------------------------------------------------------------
459
+ // MiMo-V2.6 cache metrics + sticky-session identity
460
+ //
461
+ // MiMo caching is provider-managed (implicit context caching). Xiaomi documents
462
+ // usage.prompt_tokens_details.cached_tokens as the number of PROMPT tokens
463
+ // served from cache and prompt_tokens as the total prompt-token count, so the
464
+ // authoritative cache metric is cachedTokens / promptTokens -- NOT the
465
+ // read/(read+write) form used by other families. `hitRatePct` is intentionally
466
+ // left untouched so existing providers are unaffected.
467
+ //
468
+ // The runtime exposes `Message.info.tokens` as { input, output, cache:{read,
469
+ // write} } where `input` is the NON-cached prompt input and `cache.read` is the
470
+ // cached prompt input. Total prompt tokens are therefore derived as
471
+ // read + input (cache.write is a separate accounting bucket and is NOT folded
472
+ // in). We never fabricate cache-write values.
473
+ // ---------------------------------------------------------------------------
474
+
475
+ // cachedTokens / promptTokens, rounded to a percentage. Returns null when the
476
+ // denominator is unknown/zero or the inputs are not finite numbers, so no
477
+ // fabricated hit rate is ever emitted.
478
+ export function mimoHitRate(cachedTokens, promptTokens) {
479
+ if (!Number.isFinite(cachedTokens) || !Number.isFinite(promptTokens)) return null
480
+ if (promptTokens <= 0 || cachedTokens < 0) return null
481
+ return Math.round((100 * cachedTokens) / promptTokens)
482
+ }
483
+
484
+ // Derive a stable, session-scoped identifier suitable for OpenRouter's
485
+ // documented `session_id` sticky-routing key. Pure function of the OpenCode
486
+ // session id only: identical sessions map to identical ids, distinct sessions
487
+ // map to distinct ids, and transient request contents cannot influence it.
488
+ // The value is printable, contains no whitespace, and is far below the 256-char
489
+ // cap (25 chars). NOTE: this runtime's OpenRouter request adapter does not emit
490
+ // a top-level `session_id` (it forwards only usage/reasoning/prompt_cache_key),
491
+ // so this id is currently recorded as telemetry only and is never injected.
492
+ export function mimoSessionIdFor(sessionID) {
493
+ if (typeof sessionID !== "string" || sessionID.length === 0) return null
494
+ return `mimo-ses-${shorthash(sessionID)}`
495
+ }
496
+
497
+ // Detect a provider switch within the same session. `previous` and `current`
498
+ // are {providerID, modelID} observations. Returns {changed:false} until two
499
+ // real observations exist; a change is only reported when both are known and
500
+ // the providerID differs. Never forces or overrides provider selection.
501
+ export function providerChangeEvent(previous, current) {
502
+ const prev = previous && typeof previous.providerID === "string" ? previous.providerID : null
503
+ const cur = current && typeof current.providerID === "string" ? current.providerID : null
504
+ if (prev == null || cur == null || prev === cur) return { changed: false, from: null, to: null }
505
+ return { changed: true, from: previous, to: current }
506
+ }
507
+
441
508
  // Decide whether a `usage` record should be emitted for an aggregation sample.
442
509
  // We must not fabricate a zero-valued cache event merely because the session
443
510
  // became idle: a sample only counts when at least one assistant message with