opencode-cache-engine 0.1.1 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,20 +1,28 @@
1
1
  # OpenCode Cache Engine
2
2
 
3
- Provider-aware prompt-cache optimization and observability for [OpenCode](https://opencode.ai).
3
+ Provider-aware prompt-cache optimization and observability for
4
+ [OpenCode](https://opencode.ai).
5
+
6
+ `opencode-cache-engine` is an OpenCode npm plugin with two targets:
7
+
8
+ - **Server target** — the actual cache-engine runtime and provider policies.
9
+ - **TUI target** — registration with OpenCode's TUI plugin manager.
10
+
11
+ The server target handles cache optimization, prompt-shape diagnostics, compaction handling, and cache telemetry. The TUI target provides the plugin-manager integration and enable/disable state for the TUI-facing plugin entry.
4
12
 
5
13
  `CacheEngine` is an OpenCode plugin designed for long-running agent sessions where prompt-cache efficiency affects both latency and cost. It keeps the harness conservative for providers whose cache behavior is already automatic, while applying provider-specific optimizations where the provider exposes useful cache controls or where prompt structure can be safely improved.
6
14
 
7
- The plugin currently has three cache-policy families:
15
+ The plugin currently has four cache-policy families:
8
16
 
9
17
  * **DeepSeek V4.1 Flash** — passive cache-stability and observability
10
18
  * **GPT-5.6 Luna** — active cache-control configuration
11
19
  * **GLM-5.3 Flash** — conservative system-prompt stabilization
20
+ * **MiMo-V2.6 (Flash / Pro)** — prefix stability and OpenRouter session-affinity diagnostics
12
21
 
13
22
  The central design principle is:
14
23
 
15
24
  > Optimize the request structure only when there is a clear provider-specific reason to do so. Otherwise, preserve OpenCode's native request behavior and measure what the provider actually reports.
16
25
 
17
- ---
18
26
 
19
27
  ## What this plugin does
20
28
 
@@ -28,12 +36,12 @@ It:
28
36
  4. Records provider-reported cache token usage.
29
37
  5. Adds a deterministic compaction continuation block.
30
38
  6. Applies GPT-5.6 cache-control metadata.
31
- 7. Applies the GLM-5.3 volatile-environment relocation.
39
+ 7. Applies the GLM-5.3 and MiMo-V2.6 volatile-environment relocation.
32
40
  8. Records diagnostics that help determine whether prompt-shape changes correlate with cache behavior.
41
+ 9. Records MiMo-V2.6 provider identity and provider-switch diagnostics.
33
42
 
34
43
  The plugin deliberately avoids pretending that a local hash is proof of a provider cache hit. Provider-reported token usage remains the authoritative signal.
35
44
 
36
- ---
37
45
 
38
46
  # Provider behavior
39
47
 
@@ -88,7 +96,6 @@ DeepSeek:
88
96
  force cache behavior
89
97
  ```
90
98
 
91
- ---
92
99
 
93
100
  ## GPT-5.6 Luna
94
101
 
@@ -153,7 +160,6 @@ compaction:
153
160
 
154
161
  This prevents a compaction-specific prompt from sharing the same GPT cache namespace as the normal live-session prompt. The behavior is deterministic and tested explicitly.
155
162
 
156
- ---
157
163
 
158
164
  ## GLM-5.3 Flash
159
165
 
@@ -219,17 +225,168 @@ It only occurs when:
219
225
 
220
226
  The plugin does not arbitrarily rearrange unrelated prompt content.
221
227
 
222
- ---
223
228
 
224
- # Prompt-cache strategy
229
+ ## MiMo-V2.6 (Flash / Pro)
230
+
231
+ ### Policy: prefix stability + OpenRouter session affinity
232
+
233
+ MiMo-V2.6 is Xiaomi's current model family. The plugin targets exactly two
234
+ identifiers:
235
+
236
+ * `xiaomi/mimo-v2.6-flash` / `mimo-v2.6-flash`
237
+ * `xiaomi/mimo-v2.6-pro` / `mimo-v2.6-pro`
238
+
239
+ Detection also tolerates `provider/model` shapes where `api.id` contains those
240
+ slugs. It deliberately does **not** match `mimo-v2.5`, `mimo-v2.5-pro`,
241
+ `mimo-v2.6-pro-ultraspeed`, or unrelated MiMo models.
242
+
243
+ ### Implicit context caching
244
+
245
+ Xiaomi documents context caching for both V2.6 Flash and Pro, and exposes
246
+ `usage.prompt_tokens_details.cached_tokens` as the number of prompt tokens
247
+ served from cache. The V2.6 API documents implicit context caching, not a
248
+ user-supplied cache key or explicit breakpoint.
249
+
250
+ Accordingly the plugin **injects no cache-control parameter** for MiMo. It does
251
+ not send `promptCacheKey`, `cacheControl`, `cacheBreakpoint`, or `ttl`.
252
+ Implicit caching is the default assumption.
253
+
254
+ ### Environment-block stabilization
255
+
256
+ MiMo uses the same narrow, content-preserving transformation as GLM-5.3: the
257
+ identifiable volatile `<env>` block is relocated to the **tail** of the single
258
+ system string. Contents are preserved byte-for-byte; only position changes. This
259
+ keeps the large reusable prefix stable when only the environment/date changes.
225
260
 
226
- The plugin uses three different strategies because cache mechanisms differ by provider.
261
+ The transformation is applied only when:
227
262
 
228
- | Provider | Prompt text changed? | Cache metadata changed? | Main strategy |
229
- | ----------------- | -------------------: | ----------------------: | --------------------------------- |
230
- | DeepSeek V4.1 Flash | No | No | Preserve stable harness + observe |
231
- | GPT-5.6 Luna | No | Yes | Stable cache key + cache options |
232
- | GLM-5.3 Flash | Yes, narrowly | No provider cache key | Isolate volatile system content |
263
+ * the selected model is MiMo-V2.6 Flash/Pro
264
+ * `mimo26.stabilizeSystem` is `true`
265
+ * there is exactly one system string
266
+ * the expected `<env>` markers exist and the block is identified unambiguously
267
+ * the block is not already at the tail
268
+
269
+ ### No generic system-prompt freezing
270
+
271
+ MiMo-Code's own harness freezes its per-session system prefix. This plugin does
272
+ **not** copy that mechanism. System instructions can legitimately change because
273
+ of permissions, tools, agent mode, skills, MCP state, or project configuration;
274
+ a plugin-level snapshot must never override a legitimate change.
275
+
276
+ Instead the plugin:
277
+
278
+ * records a first-seen system baseline per session;
279
+ * computes the full system hash, stable prefix hash, and volatile suffix hash;
280
+ * records changes for MiMo sessions;
281
+ * allows the `<env>` relocation when that is the only identified volatility;
282
+ * reports other system changes diagnostically and never overwrites the new
283
+ content.
284
+
285
+ Explicit telemetry events:
286
+
287
+ * `mimo_system_env_relocated`
288
+ * `mimo_system_prefix_changed`
289
+
290
+ ### OpenRouter sticky session — derived but not injected
291
+
292
+ OpenRouter documents a top-level `session_id` request field for sticky provider
293
+ routing, which keeps a session's requests on the same upstream provider so
294
+ provider-side prompt caches stay warm.
295
+
296
+ The plugin includes a pure, session-scoped derivation (`mimoSessionIdFor`):
297
+ deterministic, distinct per session, printable/no-whitespace, and well under the
298
+ 256-character cap. However, **the derived id is not injected into requests**.
299
+
300
+ Rationale (verified against the installed runtime): OpenCode's OpenRouter
301
+ request adapter forwards only `usage`, `reasoning`, and `prompt_cache_key` from
302
+ provider options, and exposes no top-level `session_id` path. Adding an
303
+ unsupported field would be guessing, so the id is recorded as telemetry only,
304
+ and `mimo26.stickySession` currently gates that recording. If a future runtime
305
+ gains a verified `session_id` path, the helper is already in place.
306
+
307
+ Note that OpenCode itself sets `x-session-affinity` / `X-Session-Id` HTTP
308
+ headers for non-opencode providers, and can set a flat `promptCacheKey` for
309
+ OpenRouter when `setCacheKey: true` is configured. Those are HTTP routing
310
+ headers and an OpenAI-style cache key respectively — they are not OpenRouter's
311
+ documented body `session_id`.
312
+
313
+ ### MiMo cache metrics
314
+
315
+ MiMo caches are provider-managed, so the authoritative metric is provider
316
+ reported. For MiMo the plugin emits the preferred ratio:
317
+
318
+ ```text
319
+ cacheHitRate = cachedTokens / promptTokens
320
+ ```
321
+
322
+ This is intentionally **not** the `read / (read + write)` form used by other
323
+ families. It is not GLM's `read / (read + write + input)` either.
324
+
325
+ Derivation: the runtime exposes assistant tokens as `{ input, output,
326
+ cache:{ read, write } }`, where `input` is the non-cached prompt input and
327
+ `cache.read` is the cached prompt input. Total prompt tokens are therefore
328
+ derived as `read + input`, and `cachedTokens = read`. `cache.write` is a
329
+ separate accounting bucket and is not folded in; no cache-write value is
330
+ fabricated, and the ratio is `null` when `promptTokens` is zero.
331
+
332
+ A MiMo usage record looks conceptually like:
333
+
334
+ ```json
335
+ {
336
+ "kind": "usage",
337
+ "policy": "mimo26",
338
+ "provider": "openrouter",
339
+ "model": "xiaomi/mimo-v2.6-flash",
340
+ "promptTokens": 50000,
341
+ "cachedTokens": 47000,
342
+ "cacheHitRate": 94
343
+ }
344
+ ```
345
+
346
+ ### Provider-switch diagnostics
347
+
348
+ Because MiMo caches live at the provider side, a provider change within one
349
+ session can silently invalidate them. The plugin records provider identity on
350
+ every MiMo request and emits a `mimo_provider_changed` boundary event when the
351
+ OpenCode `providerID` changes within a session. It never forces or overrides the
352
+ user's provider selection.
353
+
354
+ Limitation: OpenRouter's *upstream* provider selection (for example
355
+ `xiaomi/fp8` vs `atlas-cloud/fp8`) is not exposed to plugins, so only the
356
+ OpenCode `providerID`/`modelID` are observable.
357
+
358
+ ### Reasoning / thinking
359
+
360
+ MiMo-V2.6 supports deep thinking and reports reasoning tokens. The plugin does
361
+ not treat reasoning replay as a cache requirement: reasoning diagnostics are
362
+ instrumentation only, and the plugin never rewrites, duplicates, reorders, or
363
+ re-injects reasoning content, nor changes reasoning effort for caching.
364
+
365
+ ### Skill-catalog / history limitation
366
+
367
+ MiMo-Code moved skill catalogs out of repeatedly rewritten user messages and
368
+ toward the system tail. In this OpenCode runtime the skill guidance
369
+ (`<available_skills>`) and MCP instructions already live in the **system
370
+ prefix**, not in user-message history. The plugin therefore performs no
371
+ message-history rewrite. Skill/MCP changes simply appear as system-prefix changes
372
+ and are reported diagnostically; the message content is left untouched.
373
+
374
+
375
+ # Provider comparison
376
+
377
+ | Provider | Detection | Prompt text changed? | Cache metadata changed? | Primary cache signal |
378
+ | ------------------- | ---------------------------------- | --------------------------- | ---------------------------------- | --------------------------------------- |
379
+ | DeepSeek V4.1 Flash | `deepseek` | No | No | provider `cache.read`/`cache.write` |
380
+ | GPT-5.6 Luna | `gpt-5.6*` on OpenAI-ish endpoints | No | Yes: `prompt_cache_key` + options | provider cache tokens |
381
+ | GLM-5.3 Flash | `glm-5.3*` | Yes, narrowly (`<env>` tail) | No provider cache key | provider cache tokens (GLM ratio) |
382
+ | MiMo-V2.6 Flash/Pro | `mimo-v2.6-flash` / `mimo-v2.6-pro` | Yes, narrowly (`<env>` tail) | No: implicit caching only | `cached_tokens / prompt_tokens` |
383
+
384
+
385
+ # Prompt-cache strategy
386
+
387
+ The plugin uses different strategies because cache mechanisms differ by provider.
388
+ The table above summarises them; the essential point is the distinction between
389
+ *changing prompt text* and *changing cache metadata*.
233
390
 
234
391
  This distinction is fundamental.
235
392
 
@@ -237,7 +394,6 @@ The plugin is **not** a generic "rewrite every prompt for caching" engine.
237
394
 
238
395
  It is a provider-aware cache policy engine.
239
396
 
240
- ---
241
397
 
242
398
  # System-prompt diagnostics
243
399
 
@@ -261,7 +417,6 @@ It does **not** mean:
261
417
 
262
418
  This distinction is intentional. Provider-reported cache token counts are the authoritative cache signal.
263
419
 
264
- ---
265
420
 
266
421
  # Tool-definition diagnostics
267
422
 
@@ -294,7 +449,6 @@ This distinction matters because semantic equality and byte-level request equali
294
449
 
295
450
  The plugin uses these fingerprints for **diagnostics only**. It does not reorder the tools to force a particular fingerprint.
296
451
 
297
- ---
298
452
 
299
453
  # Compaction handling
300
454
 
@@ -313,7 +467,6 @@ The plugin adds a deterministic continuation template:
313
467
  The digest is inserted once per compaction operation using a guard that prevents duplicate insertion if the compaction hook fires multiple times.
314
468
  The objective is to provide a deterministic continuation structure rather than generating a different arbitrary cache-affecting block on every compaction.
315
469
 
316
- ---
317
470
 
318
471
  # Cache metrics
319
472
 
@@ -350,6 +503,15 @@ read / (read + write + input)
350
503
 
351
504
  as implemented by `glmHitRatio()`.
352
505
 
506
+ For MiMo, the implementation uses the provider-documented prompt-cache ratio:
507
+
508
+ ```text
509
+ cacheHitRate = cachedTokens / promptTokens
510
+ ```
511
+
512
+ implemented by `mimoHitRate()`. `hitRatePct()` itself is left untouched so other
513
+ providers are unaffected.
514
+
353
515
  ### Important metric distinction
354
516
 
355
517
  These ratios answer different questions.
@@ -362,7 +524,12 @@ These ratios answer different questions.
362
524
 
363
525
  > How much of the total prompt-token accounting was represented by cached reads?
364
526
 
365
- Do not treat the two percentages as interchangeable.
527
+ `cachedTokens / promptTokens` (MiMo) answers:
528
+
529
+ > Of the prompt tokens the provider processed, what fraction was served from
530
+ > cache?
531
+
532
+ Do not treat these percentages as interchangeable.
366
533
 
367
534
  ---
368
535
 
@@ -388,7 +555,6 @@ Telemetry is best-effort.
388
555
 
389
556
  A failed metrics write must never break an OpenCode request. The recorder catches write failures rather than allowing telemetry failures to affect execution.
390
557
 
391
- ---
392
558
 
393
559
  # Metrics examples
394
560
 
@@ -446,8 +612,30 @@ Telemetry is intended to answer questions such as:
446
612
  * Did a compaction occur?
447
613
  * Which provider/model/policy was active?
448
614
  * Did the GLM system stabilization actually change the observed prompt shape?
615
+ * Did MiMo's environment relocation fire (`mimo_system_env_relocated`)?
616
+ * Did MiMo's stable system prefix change (`mimo_system_prefix_changed`)?
617
+ * Did the MiMo provider change within a session (`mimo_provider_changed`)?
618
+ * What was MiMo's provider-reported cache hit rate (`cacheHitRate`)?
619
+
620
+ A MiMo usage record adds the provider-reported cache fields:
621
+
622
+ ```json
623
+ {
624
+ "kind": "usage",
625
+ "sid": "session-id",
626
+ "ts": 1750000000000,
627
+ "policy": "mimo26",
628
+ "provider": "openrouter",
629
+ "model": "xiaomi/mimo-v2.6-flash",
630
+ "read": 47000,
631
+ "input": 3000,
632
+ "promptTokens": 50000,
633
+ "cachedTokens": 47000,
634
+ "cacheHitRate": 94,
635
+ "stickySessionId": "mimo-ses-0123456789abcdef"
636
+ }
637
+ ```
449
638
 
450
- ---
451
639
 
452
640
  # Configuration
453
641
 
@@ -476,6 +664,12 @@ The default configuration is:
476
664
  "enabled": true,
477
665
  "stabilizeSystem": true,
478
666
  "preserveThinkingIntegrity": true
667
+ },
668
+ "mimo26": {
669
+ "enabled": true,
670
+ "stabilizeSystem": true,
671
+ "stickySession": true,
672
+ "preserveThinkingIntegrity": true
479
673
  }
480
674
  }
481
675
  }
@@ -483,7 +677,6 @@ The default configuration is:
483
677
 
484
678
  The configuration parser starts from these defaults and applies valid file/environment overrides without mutating the caller's configuration object.
485
679
 
486
- ---
487
680
 
488
681
  # Configuration options
489
682
 
@@ -535,7 +728,6 @@ Controls whether the deterministic compaction continuation block is inserted.
535
728
 
536
729
  Controls warning logs for observed prefix-shape changes.
537
730
 
538
- ---
539
731
 
540
732
  # DeepSeek configuration
541
733
 
@@ -549,7 +741,6 @@ There are intentionally very few settings here.
549
741
 
550
742
  DeepSeek is treated as the conservative/passive policy.
551
743
 
552
- ---
553
744
 
554
745
  # GPT-5.6 configuration
555
746
 
@@ -601,7 +792,6 @@ Defaults to:
601
792
 
602
793
  Existing request options are not overwritten by the plugin.
603
794
 
604
- ---
605
795
 
606
796
  # GLM-5.3 configuration
607
797
 
@@ -629,7 +819,45 @@ The reasoning instrumentation is intended to identify anomalies such as:
629
819
 
630
820
  It is diagnostic rather than a reason to rewrite or fabricate reasoning content. The implementation maps these conditions to explicit diagnostic reasons.
631
821
 
632
- ---
822
+
823
+ # MiMo-V2.6 configuration
824
+
825
+ ```json
826
+ {
827
+ "mimo26": {
828
+ "enabled": true,
829
+ "stabilizeSystem": true,
830
+ "stickySession": true,
831
+ "preserveThinkingIntegrity": true
832
+ }
833
+ }
834
+ ```
835
+
836
+ ### `enabled`
837
+
838
+ Enables the MiMo-V2.6 policy.
839
+
840
+ ### `stabilizeSystem`
841
+
842
+ Enables relocation of the volatile `<env>` section to the system-prompt tail
843
+ (same narrow, content-preserving transformation as GLM-5.3).
844
+
845
+ ### `stickySession`
846
+
847
+ Gates derivation/recording of the OpenRouter sticky-session id
848
+ (`mimoSessionIdFor`). The id is recorded as telemetry; it is **not** injected
849
+ into the request because this runtime exposes no verified OpenRouter top-level
850
+ `session_id` path. See "OpenRouter sticky session — derived but not injected".
851
+
852
+ ### `preserveThinkingIntegrity`
853
+
854
+ Enables reasoning diagnostics as instrumentation. It never rewrites, duplicates,
855
+ reorders, or re-injects reasoning content, and it is not a cache requirement.
856
+
857
+ No `cacheBlockSize`, `cacheTTL`, `cacheBreakpoint`, or `minimumCacheTokens`
858
+ knobs are exposed: those values are not established by authoritative V2.6
859
+ documentation.
860
+
633
861
 
634
862
  # Model detection
635
863
 
@@ -639,6 +867,7 @@ The plugin classifies requests into:
639
867
  deepseek
640
868
  gpt56
641
869
  glm53
870
+ mimo26
642
871
  neutral
643
872
  ```
644
873
 
@@ -647,9 +876,13 @@ The model detector recognizes:
647
876
  * DeepSeek model/provider identifiers
648
877
  * GPT-5.6 variants
649
878
  * GLM-5.3 variants
879
+ * MiMo-V2.6 Flash and Pro (`xiaomi/mimo-v2.6-flash`, `mimo-v2.6-pro`, ...)
650
880
 
651
881
  GPT-5.6 has an additional OpenAI/Azure-context check so a string containing `gpt-5.6` does not automatically cause GPT-specific fields to be sent to an unrelated endpoint.
652
882
 
883
+ MiMo detection targets exactly Flash and Pro: it excludes `mimo-v2.5`,
884
+ `mimo-v2.5-pro`, and `mimo-v2.6-pro-ultraspeed`.
885
+
653
886
  Unknown models use the neutral policy.
654
887
 
655
888
  Neutral means:
@@ -658,7 +891,6 @@ Neutral means:
658
891
  no provider-specific request mutation
659
892
  ```
660
893
 
661
- ---
662
894
 
663
895
  # OpenRouter usage
664
896
 
@@ -666,11 +898,10 @@ This plugin is compatible with OpenRouter because the cache policy is based on t
666
898
 
667
899
  For cache-sensitive workloads, provider stability remains important.
668
900
 
669
- The plugin does not attempt to compensate for provider switching by rewriting prompts.
901
+ The plugin does not attempt to compensate for provider switching by rewriting prompts. It records MiMo provider identity and provider-switch diagnostics so routing instability is at least observable.
670
902
 
671
903
  For that reason, a stable provider route is preferable when your goal is to measure and maximize prefix reuse.
672
904
 
673
- ---
674
905
 
675
906
  # Architecture
676
907
 
@@ -893,14 +1124,15 @@ A typical standalone repository can use:
893
1124
  opencode-cache-engine/
894
1125
  ├── src/
895
1126
  │ ├── cache-engine.ts
896
- │ └── cache-engine-core.mjs
1127
+ │ ├── cache-engine-core.mjs
1128
+ │ └── tui.mjs
897
1129
  ├── test/
898
1130
  │ └── cache-engine.test.mjs
899
1131
  ├── examples/
900
1132
  │ └── cache-engine.json
1133
+ ├── package.json
901
1134
  ├── README.md
902
- ├── LICENSE
903
- └── package.json
1135
+ └── LICENSE
904
1136
  ```
905
1137
 
906
1138
  The OpenCode plugin export remains:
@@ -925,6 +1157,20 @@ without changing the `CacheEngine` export identifier.
925
1157
 
926
1158
  Install the plugin into the OpenCode plugins directory according to your OpenCode plugin-loading setup.
927
1159
 
1160
+ `opencode-cache-engine` is distributed as an npm package.
1161
+
1162
+ ## Server/runtime plugin
1163
+
1164
+ Add the package to the OpenCode runtime plugin configuration:
1165
+
1166
+ ```json
1167
+ {
1168
+ "plugin": [
1169
+ "opencode-cache-engine"
1170
+ ]
1171
+ }
1172
+ ```
1173
+
928
1174
  The runtime entry should expose:
929
1175
 
930
1176
  ```ts
@@ -959,6 +1205,12 @@ GPT-5.6:
959
1205
 
960
1206
  GLM-5.3:
961
1207
  volatile env block relocated when eligible
1208
+
1209
+ MiMo-V2.6:
1210
+ volatile env block relocated when eligible
1211
+ no GPT/GLM-only cache fields present
1212
+ no OpenRouter top-level session_id injected (unsupported by this runtime)
1213
+ telemetry carries provider/model/promptTokens/cachedTokens/cacheHitRate
962
1214
  ```
963
1215
 
964
1216
  ---
@@ -1001,6 +1253,42 @@ The relevant block must contain the expected beginning and closing marker, and t
1001
1253
 
1002
1254
  ---
1003
1255
 
1256
+ ## MiMo-V2.6 prompt is not being changed
1257
+
1258
+ MiMo uses the same eligibility rules as GLM-5.3: exactly one system string, both
1259
+ `<env>` markers present, block identified unambiguously, and
1260
+ `mimo26.stabilizeSystem` enabled. If the block is already at the tail, the
1261
+ operation is a no-op.
1262
+
1263
+ ---
1264
+
1265
+ ## MiMo provider is not classified as `mimo26`
1266
+
1267
+ Verify the model identifier is exactly Flash or Pro:
1268
+
1269
+ ```text
1270
+ mimo-v2.6-flash
1271
+ mimo-v2.6-pro
1272
+ xiaomi/mimo-v2.6-flash
1273
+ xiaomi/mimo-v2.6-pro
1274
+ ```
1275
+
1276
+ `mimo-v2.5`, `mimo-v2.5-pro`, and `mimo-v2.6-pro-ultraspeed` are intentionally
1277
+ not matched.
1278
+
1279
+ ---
1280
+
1281
+ ## No OpenRouter `session_id` is sent for MiMo
1282
+
1283
+ This is expected. The installed OpenCode runtime's OpenRouter request adapter
1284
+ forwards only `usage`, `reasoning`, and `prompt_cache_key` from provider options
1285
+ and exposes no top-level `session_id` path. The plugin derives a stable
1286
+ `stickySessionId` and records it as telemetry, but does not inject it rather than
1287
+ send an unsupported field. This may change if a future runtime exposes a verified
1288
+ path.
1289
+
1290
+ ---
1291
+
1004
1292
  ## Metrics file is missing
1005
1293
 
1006
1294
  Telemetry is best-effort.
@@ -1025,6 +1313,7 @@ The current implementation is intentionally conservative:
1025
1313
  DeepSeek -> preserve and measure
1026
1314
  GPT-5.6 -> configure cache controls
1027
1315
  GLM-5.3 -> isolate volatile prompt content
1316
+ MiMo-V2.6 -> stabilize prefix + observe provider/cache reality
1028
1317
  ```
1029
1318
 
1030
1319
  That separation is the core design of the project.
@@ -20,6 +20,12 @@
20
20
  "enabled": true,
21
21
  "stabilizeSystem": true,
22
22
  "preserveThinkingIntegrity": true
23
+ },
24
+ "mimo26": {
25
+ "enabled": true,
26
+ "stabilizeSystem": true,
27
+ "stickySession": true,
28
+ "preserveThinkingIntegrity": true
23
29
  }
24
30
  }
25
31
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "opencode-cache-engine",
3
- "version": "0.1.1",
3
+ "version": "0.3.0",
4
4
  "private": false,
5
5
  "description": "Provider-aware prompt-cache optimization and telemetry for OpenCode",
6
6
  "keywords": [
@@ -9,7 +9,8 @@
9
9
  "prompt-cache",
10
10
  "deepseek",
11
11
  "glm",
12
- "gpt"
12
+ "gpt",
13
+ "mimo"
13
14
  ],
14
15
  "homepage": "https://github.com/AlexJoaquimPereira/opencode-cache-engine#readme",
15
16
  "bugs": {
@@ -27,6 +28,10 @@
27
28
  "example": "examples",
28
29
  "test": "test"
29
30
  },
31
+ "exports": {
32
+ "./server": "./src/cache-engine.ts",
33
+ "./tui": "./src/tui.mjs"
34
+ },
30
35
  "scripts": {
31
36
  "test": "node --test test/cache-engine.test.mjs",
32
37
  "postpublish": "VERSION=$(node -p \"require('./package').version\") && TAG=v$VERSION && echo \"Creating git tag $TAG for npm $VERSION\" && git tag $TAG && git push origin $TAG || echo \"Tag $TAG may already exist or push failed; continuing.\"",
@@ -5,9 +5,10 @@
5
5
  // TypeScript compiler. The plugin entry (cache-engine.ts) imports this module.
6
6
  //
7
7
  // This module is PROVIDER-AWARE: it classifies a model into a cache-policy
8
- // family (deepseek | gpt56 | glm53 | neutral) and exposes small pure helpers for
9
- // each family's strategy. The plugin entry (cache-engine.ts) remains the only
10
- // place that touches OpenCode hooks; every decision here is testable in Node.
8
+ // family (deepseek | gpt56 | glm53 | mimo26 | neutral) and exposes small pure
9
+ // helpers for each family's strategy. The plugin entry (cache-engine.ts) remains
10
+ // the only place that touches OpenCode hooks; every decision here is testable in
11
+ // Node.
11
12
  //
12
13
  // Terminology note: these functions deal with the *observed* system/tool
13
14
  // prefix shape. An observed change means the request's prefix bytes changed; it
@@ -35,6 +36,7 @@ export const DIGEST_TEMPLATE = `## Session digest (cache-stable continuation blo
35
36
  export const POLICY_DEEPSEEK = "deepseek"
36
37
  export const POLICY_GPT56 = "gpt56"
37
38
  export const POLICY_GLM53 = "glm53"
39
+ export const POLICY_MIMO26 = "mimo26"
38
40
  export const POLICY_NEUTRAL = "neutral"
39
41
 
40
42
  export const GPT56_DEFAULT_TTL = "30m"
@@ -65,6 +67,15 @@ function defaultPolicies() {
65
67
  stabilizeSystem: true,
66
68
  preserveThinkingIntegrity: true,
67
69
  },
70
+ mimo26: {
71
+ enabled: true,
72
+ stabilizeSystem: true,
73
+ stickySession: true,
74
+ // MiMo reasoning diagnostics are instrumentation only. Unlike GLM
75
+ // preserved thinking, there is no evidence that MiMo prompt-cache reuse
76
+ // depends on reasoning replay, so this never rewrites reasoning content.
77
+ preserveThinkingIntegrity: true,
78
+ },
68
79
  }
69
80
  }
70
81
 
@@ -118,7 +129,7 @@ export function parseConfig(raw, env) {
118
129
  if (typeof raw.logPrefixChanges === "boolean") cfg.logPrefixChanges = raw.logPrefixChanges
119
130
  if (raw.policies && typeof raw.policies === "object") {
120
131
  const d = defaultPolicies()
121
- for (const fam of ["deepseek", "gpt56", "glm53"]) {
132
+ for (const fam of ["deepseek", "gpt56", "glm53", "mimo26"]) {
122
133
  if (raw.policies[fam]) cfg.policies[fam] = parsePolicy(raw.policies[fam], d[fam])
123
134
  }
124
135
  }
@@ -232,6 +243,10 @@ function isOpenAIish(s) {
232
243
  const GPT56_RE = /gpt-5\.6(?![\d.])/i
233
244
  // GLM 5.3 family only (not glm-4.x / glm-4.6 etc).
234
245
  const GLM53_RE = /glm-5\.3(?![\d.])/i
246
+ // Xiaomi MiMo V2.6 explicitly targets Flash + Pro only. The trailing
247
+ // (?![\w-]) guard prevents matching a hypothetical "mimo-v2.6-pro-ultraspeed"
248
+ // or "mimo-v2.6-flashx", and the v2\.6 literal excludes V2.5 / V2.
249
+ const MIMO26_RE = /mimo-v2\.6-(flash|pro)(?![\w-])/i
235
250
  const DEEPSEEK_RE = /deepseek/i
236
251
 
237
252
  // Pure classifier. Returns one of the POLICY_* keys. `model` may be a full
@@ -242,6 +257,7 @@ export function detectPolicy(model) {
242
257
  if (!s.slug) return POLICY_NEUTRAL
243
258
  if (GPT56_RE.test(s.slug) && isOpenAIish(s)) return POLICY_GPT56
244
259
  if (GLM53_RE.test(s.slug)) return POLICY_GLM53
260
+ if (MIMO26_RE.test(s.slug)) return POLICY_MIMO26
245
261
  if (DEEPSEEK_RE.test(s.slug) || DEEPSEEK_RE.test(s.providerID)) return POLICY_DEEPSEEK
246
262
  return POLICY_NEUTRAL
247
263
  }
@@ -284,16 +300,17 @@ export function gptCacheOptionsDelta(existingOptions, { key, mode = GPT56_DEFAUL
284
300
 
285
301
  // The OpenCode system string begins with the agent prompt, then an env block
286
302
  // ("You are powered by the model named ... Today's date: ... </env>") whose
287
- // only per-day volatile byte is the date line. For GLM-5.3 we relocate that
288
- // whole identifiable env block to the END of the system string so a daily date
289
- // change only invalidates the tail of the prompt, leaving the long stable
290
- // prefix intact. Content is preserved byte-for-byte (only position changes).
303
+ // only per-day volatile byte is the date line. For GLM-5.3 and MiMo-V2.6 we
304
+ // relocate that whole identifiable env block to the END of the system string so
305
+ // a daily date change only invalidates the tail of the prompt, leaving the long
306
+ // stable prefix intact. Content is preserved byte-for-byte (only position
307
+ // changes).
291
308
  //
292
309
  // Returns { text, changed }. When the block cannot be identified unambiguously,
293
310
  // returns the input unchanged (changed:false). This is a content-preserving
294
311
  // reorder of clearly volatile metadata only -- it never reorders arbitrary
295
312
  // instructions. This function is ONLY applied when the caller has already
296
- // classified the model as GLM-5.3.
313
+ // classified the model into a family that opts into system stabilization.
297
314
  export function relocateVolatileEnvBlock(text) {
298
315
  if (typeof text !== "string") return { text, changed: false }
299
316
  const START = "You are powered by the model named "
@@ -438,6 +455,56 @@ export function glmHitRatio(read, write, input) {
438
455
  return Math.round((100 * read) / denom)
439
456
  }
440
457
 
458
+ // ---------------------------------------------------------------------------
459
+ // MiMo-V2.6 cache metrics + sticky-session identity
460
+ //
461
+ // MiMo caching is provider-managed (implicit context caching). Xiaomi documents
462
+ // usage.prompt_tokens_details.cached_tokens as the number of PROMPT tokens
463
+ // served from cache and prompt_tokens as the total prompt-token count, so the
464
+ // authoritative cache metric is cachedTokens / promptTokens -- NOT the
465
+ // read/(read+write) form used by other families. `hitRatePct` is intentionally
466
+ // left untouched so existing providers are unaffected.
467
+ //
468
+ // The runtime exposes `Message.info.tokens` as { input, output, cache:{read,
469
+ // write} } where `input` is the NON-cached prompt input and `cache.read` is the
470
+ // cached prompt input. Total prompt tokens are therefore derived as
471
+ // read + input (cache.write is a separate accounting bucket and is NOT folded
472
+ // in). We never fabricate cache-write values.
473
+ // ---------------------------------------------------------------------------
474
+
475
+ // cachedTokens / promptTokens, rounded to a percentage. Returns null when the
476
+ // denominator is unknown/zero or the inputs are not finite numbers, so no
477
+ // fabricated hit rate is ever emitted.
478
+ export function mimoHitRate(cachedTokens, promptTokens) {
479
+ if (!Number.isFinite(cachedTokens) || !Number.isFinite(promptTokens)) return null
480
+ if (promptTokens <= 0 || cachedTokens < 0) return null
481
+ return Math.round((100 * cachedTokens) / promptTokens)
482
+ }
483
+
484
+ // Derive a stable, session-scoped identifier suitable for OpenRouter's
485
+ // documented `session_id` sticky-routing key. Pure function of the OpenCode
486
+ // session id only: identical sessions map to identical ids, distinct sessions
487
+ // map to distinct ids, and transient request contents cannot influence it.
488
+ // The value is printable, contains no whitespace, and is far below the 256-char
489
+ // cap (25 chars). NOTE: this runtime's OpenRouter request adapter does not emit
490
+ // a top-level `session_id` (it forwards only usage/reasoning/prompt_cache_key),
491
+ // so this id is currently recorded as telemetry only and is never injected.
492
+ export function mimoSessionIdFor(sessionID) {
493
+ if (typeof sessionID !== "string" || sessionID.length === 0) return null
494
+ return `mimo-ses-${shorthash(sessionID)}`
495
+ }
496
+
497
+ // Detect a provider switch within the same session. `previous` and `current`
498
+ // are {providerID, modelID} observations. Returns {changed:false} until two
499
+ // real observations exist; a change is only reported when both are known and
500
+ // the providerID differs. Never forces or overrides provider selection.
501
+ export function providerChangeEvent(previous, current) {
502
+ const prev = previous && typeof previous.providerID === "string" ? previous.providerID : null
503
+ const cur = current && typeof current.providerID === "string" ? current.providerID : null
504
+ if (prev == null || cur == null || prev === cur) return { changed: false, from: null, to: null }
505
+ return { changed: true, from: previous, to: current }
506
+ }
507
+
441
508
  // Decide whether a `usage` record should be emitted for an aggregation sample.
442
509
  // We must not fabricate a zero-valued cache event merely because the session
443
510
  // became idle: a sample only counts when at least one assistant message with
@@ -5,6 +5,7 @@ import {
5
5
  DIGEST_TEMPLATE,
6
6
  POLICY_GLM53,
7
7
  POLICY_GPT56,
8
+ POLICY_MIMO26,
8
9
  POLICY_NEUTRAL,
9
10
  createRecorder,
10
11
  detectPolicy,
@@ -16,10 +17,13 @@ import {
16
17
  gptCacheOptionsDelta,
17
18
  hitRatePct,
18
19
  loadConfig,
20
+ mimoHitRate,
21
+ mimoSessionIdFor,
19
22
  nextProcessedCursor,
20
23
  observeReasoningEffort,
21
24
  policyEnabled,
22
25
  prefixChangeReasons,
26
+ providerChangeEvent,
23
27
  reasoningEffortFromOptions,
24
28
  reasoningIssueReasons,
25
29
  relocateVolatileEnvBlock,
@@ -37,7 +41,7 @@ import {
37
41
  // cache-engine
38
42
  //
39
43
  // Provider-aware prompt-cache observability + conservative cache-shape
40
- // preservation for ONE OpenCode TUI across three model families:
44
+ // preservation for ONE OpenCode TUI across four model families:
41
45
  //
42
46
  // DeepSeek V4.1 Flash -> pure passive. >99.66% hit rate is preserved by never
43
47
  // mutating system/options/requests. Observability only.
@@ -52,16 +56,28 @@ import {
52
56
  // history stable, and INSTRUMENT preserved-thinking
53
57
  // integrity (duplicate/reorder/modified reasoning). No
54
58
  // invented cache key (Z.ai exposes none).
59
+ // MiMo-V2.6 -> prefix stability + OpenRouter session affinity. The
60
+ // safe env-block relocation is applied (stabilizeSystem)
61
+ // and provider switches within a session are diagnosed.
62
+ // MiMo caching is provider-managed implicit context
63
+ // caching; no cache key/breakpoint/TTL is invented. The
64
+ // OpenRouter sticky-session id is derived but currently
65
+ // NOT injected: this runtime's OpenRouter request
66
+ // adapter forwards only usage/reasoning/prompt_cache_key
67
+ // and exposes no top-level `session_id` path (verified
68
+ // against the installed runtime; see README).
55
69
  //
56
70
  // The engine remains conservative: it observes, hashes, compares, records,
57
71
  // appends a compaction continuation template, and (for GPT-5.6 only) injects
58
72
  // documented cache options. It never rewrites message history, reorders tools,
59
73
  // or alters user content. DeepSeek and neutral models are byte-untouched.
74
+ // MiMo/GLM only relocate the identifiable volatile env block, content-preserving.
60
75
  //
61
76
  // IMPORTANT (terminology): local hashes describe the *observed* prefix shape.
62
77
  // A changed hash means request bytes changed; it is NOT proof the provider's
63
78
  // cache key changed or that a cache miss occurred. Provider-reported cache
64
- // token counts are authoritative; hashes are diagnostics only.
79
+ // token counts are authoritative; hashes are diagnostics only. For MiMo, the
80
+ // authoritative cache signal is provider-reported cached_tokens.
65
81
  // ---------------------------------------------------------------------------
66
82
 
67
83
  const TOOL_FETCH_TTL_MS = 1500
@@ -113,6 +129,7 @@ type SessionState = {
113
129
  cacheRootAt: number | null
114
130
  reasoningSeen: Map<string, number>
115
131
  reasoningLastSeq: string[] | null
132
+ mimoProvider: { providerID: string; modelID: string } | null
116
133
  }
117
134
 
118
135
  const emptyShape = (): Shape => ({
@@ -154,6 +171,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
154
171
  cacheRootAt: null,
155
172
  reasoningSeen: new Map(),
156
173
  reasoningLastSeq: null,
174
+ mimoProvider: null,
157
175
  }
158
176
  sessions.set(sid, s)
159
177
  }
@@ -397,6 +415,25 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
397
415
  recFields.promptTokens = read + write + input
398
416
  recFields.glmHitRate = glmHitRatio(read, write, input)
399
417
  }
418
+ if (family === POLICY_MIMO26) {
419
+ // Provider-reported cached tokens / total prompt tokens. The runtime's
420
+ // `input` is the non-cached prompt input and `cache.read` is the
421
+ // cached prompt input, so total prompt tokens are derived as read +
422
+ // input (cache.write is a separate accounting bucket). No cache-write
423
+ // value is fabricated; the ratio is null when prompt tokens are 0.
424
+ const promptTokens = read + input
425
+ recFields.promptTokens = promptTokens
426
+ recFields.cachedTokens = read
427
+ recFields.cacheHitRate = mimoHitRate(read, promptTokens)
428
+ if (cfg.policies?.[POLICY_MIMO26]?.stickySession === true) {
429
+ recFields.stickySessionId = mimoSessionIdFor(sid)
430
+ }
431
+ // Prefer the latest live provider identity over the latched one.
432
+ if (s.mimoProvider) {
433
+ recFields.provider = s.mimoProvider.providerID
434
+ recFields.model = s.mimoProvider.modelID
435
+ }
436
+ }
400
437
  if (family === POLICY_GPT56 && s.gptInjected) {
401
438
  recFields.keyStrategy = "session"
402
439
  recFields.mode = cfg.policies?.[POLICY_GPT56]?.mode
@@ -445,6 +482,48 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
445
482
  try {
446
483
  const info = rememberModel(input.sessionID, input.model as unknown as ChatParamsModel)
447
484
  const family = info?.family
485
+
486
+ // ---- MiMo-V2.6: provider-switch diagnostics (telemetry only) ---------
487
+ // MiMo cache lives at the provider side, so a provider change within one
488
+ // OpenCode session can silently invalidate it. We record identity on every
489
+ // MiMo request and emit a diagnostic when the OpenCode providerID changes.
490
+ // This never forces or overrides provider routing.
491
+ // NOTE: OpenRouter's *upstream* provider selection (e.g. xiaomi/fp8) is
492
+ // not exposed to plugins; only the OpenCode providerID/modelID are
493
+ // observable here.
494
+ if (family === POLICY_MIMO26 && policyEnabled(cfg, POLICY_MIMO26) && info) {
495
+ // Use the LIVE model identity (not the latched one) so a provider
496
+ // switch within the session is actually observable.
497
+ const live = input.model as unknown as ChatParamsModel
498
+ const cur = {
499
+ providerID: String(live?.providerID ?? ""),
500
+ modelID: String(live?.api?.id ?? live?.id ?? ""),
501
+ }
502
+ if (cur.providerID) {
503
+ const s = get(input.sessionID)
504
+ const ev = providerChangeEvent(s.mimoProvider, cur)
505
+ if (ev.changed) {
506
+ const sticky =
507
+ cfg.policies?.[POLICY_MIMO26]?.stickySession === true
508
+ ? { stickySessionId: mimoSessionIdFor(input.sessionID) }
509
+ : {}
510
+ rec.record({
511
+ kind: "boundary",
512
+ sid: input.sessionID,
513
+ ts: Date.now(),
514
+ reason: "mimo_provider_changed",
515
+ policy: POLICY_MIMO26,
516
+ from: ev.from,
517
+ to: ev.to,
518
+ ...sticky,
519
+ note: "OpenCode providerID changed; upstream routing is not plugin-visible",
520
+ })
521
+ }
522
+ s.mimoProvider = cur
523
+ }
524
+ return
525
+ }
526
+
448
527
  if (!(family === POLICY_GPT56 && policyEnabled(cfg, POLICY_GPT56))) {
449
528
  // DeepSeek / GLM / neutral: nothing to inject. GLM has no cache-key API;
450
529
  // DeepSeek caching is fully passive; we never mutate requests for them.
@@ -626,28 +705,47 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
626
705
  const s = get(sid)
627
706
  const family = s.modelInfo?.family
628
707
 
629
- // ---- GLM-5.3 input-shape stabilization -----------------------------
708
+ // ---- GLM-5.3 / MiMo-V2.6 input-shape stabilization ------------------
630
709
  // Relocate the identifiable volatile env block (per-day date) to the
631
710
  // tail of the single system string, content-preserving, ONLY when the
632
- // model is GLM-5.3 and the block markers are present exactly. Never
633
- // touches other content/order; never applied to other families.
711
+ // family opts into system stabilization and the block markers are
712
+ // present exactly. Never touches other content/order; never applied to
713
+ // other families. The runtime passes a single-element system array
714
+ // (verified against the installed runtime), so no generic reordering is
715
+ // involved.
634
716
  //
635
717
  // In-place mutation note: request.ts keeps using its own local `system`
636
718
  // array after the hook (the trigger's returned output is ignored), so
637
719
  // reassigning `output.system = [...]` would be lost. We rewrite the
638
720
  // single element in place instead.
639
721
  let systemText = output.system.join("\n")
640
- if (
722
+ const glmStabilize =
641
723
  family === POLICY_GLM53 &&
642
724
  policyEnabled(cfg, POLICY_GLM53) &&
643
- cfg.policies?.[POLICY_GLM53]?.stabilizeSystem === true &&
644
- output.system.length === 1
645
- ) {
725
+ cfg.policies?.[POLICY_GLM53]?.stabilizeSystem === true
726
+ const mimoStabilize =
727
+ family === POLICY_MIMO26 &&
728
+ policyEnabled(cfg, POLICY_MIMO26) &&
729
+ cfg.policies?.[POLICY_MIMO26]?.stabilizeSystem === true
730
+ if ((glmStabilize || mimoStabilize) && output.system.length === 1) {
646
731
  const rel = relocateVolatileEnvBlock(output.system[0])
647
732
  if (rel.changed) {
648
733
  output.system[0] = rel.text
649
734
  systemText = rel.text
650
- log("debug", "glm system env block relocated to suffix", { sid })
735
+ if (family === POLICY_MIMO26) {
736
+ log("debug", "mimo system env block relocated to suffix", { sid })
737
+ rec.record({
738
+ kind: "boundary",
739
+ sid,
740
+ ts: Date.now(),
741
+ reason: "mimo_system_env_relocated",
742
+ policy: POLICY_MIMO26,
743
+ provider: s.modelInfo?.providerID,
744
+ model: s.modelInfo?.modelID,
745
+ })
746
+ } else {
747
+ log("debug", "glm system env block relocated to suffix", { sid })
748
+ }
651
749
  }
652
750
  }
653
751
 
@@ -724,6 +822,22 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
724
822
  ? { toolCount, prevToolCount: prevCount, semanticToolsChanged: semanticChanged, wireToolsChanged: wireChanged }
725
823
  : {}),
726
824
  })
825
+ // MiMo-specific explicit diagnostic: the STABLE prefix changed (not
826
+ // just the relocated volatile env suffix). Reported only; the new
827
+ // content is never overwritten with a stale snapshot.
828
+ if (family === POLICY_MIMO26 && reasons.includes("system_stable_prefix_changed")) {
829
+ rec.record({
830
+ kind: "boundary",
831
+ sid,
832
+ ts: Date.now(),
833
+ reason: "mimo_system_prefix_changed",
834
+ policy: POLICY_MIMO26,
835
+ provider: s.modelInfo?.providerID,
836
+ model: s.modelInfo?.modelID,
837
+ changedFields: granular,
838
+ reasons,
839
+ })
840
+ }
727
841
  if (cfg.logPrefixChanges) {
728
842
  log("warn", "observed prefix shape change", {
729
843
  sid,
package/src/tui.mjs ADDED
@@ -0,0 +1,11 @@
1
+ const plugin = {
2
+ id: "opencode-cache-engine",
3
+
4
+ async tui() {
5
+ // CacheEngine has no TUI UI of its own.
6
+ // This target exists so the npm package can be registered,
7
+ // displayed, and enabled/disabled by the OpenCode plugin manager.
8
+ },
9
+ }
10
+
11
+ export default plugin
@@ -16,6 +16,7 @@ import {
16
16
  POLICY_DEEPSEEK,
17
17
  POLICY_GLM53,
18
18
  POLICY_GPT56,
19
+ POLICY_MIMO26,
19
20
  POLICY_NEUTRAL,
20
21
  canonicalStringify,
21
22
  commonPrefixLength,
@@ -28,8 +29,11 @@ import {
28
29
  gptCacheOptionsDelta,
29
30
  hitRatePct,
30
31
  loadConfig,
32
+ mimoHitRate,
33
+ mimoSessionIdFor,
31
34
  nextProcessedCursor,
32
35
  parseConfig,
36
+ providerChangeEvent,
33
37
  relocateVolatileEnvBlock,
34
38
  scanPage,
35
39
  shapeDiff,
@@ -221,6 +225,30 @@ test("recorder writes valid JSONL to a real file", () => {
221
225
  assert.deepEqual(JSON.parse(lines[0]), { kind: "prefix-change", sid: "s", dimensions: ["system"] })
222
226
  })
223
227
 
228
+ // --- 14. TUI target test -----------------------------------------------------
229
+
230
+ test("TUI target exports the expected plugin module", async () => {
231
+ const mod = await import("../src/tui.mjs")
232
+ assert.equal(mod.default.id, "opencode-cache-engine")
233
+ assert.equal(typeof mod.default.tui, "function")
234
+ assert.equal("server" in mod.default, false)
235
+ })
236
+
237
+ // --- 15. Package manifest test with TUI --------------------------------------
238
+
239
+ test("package exposes separate server and TUI targets", async () => {
240
+ const { readFileSync } = await import("node:fs")
241
+ const { join } = await import("node:path")
242
+
243
+ const pkg = JSON.parse(
244
+ readFileSync(join(process.cwd(), "package.json"), "utf8")
245
+ )
246
+
247
+ assert.equal(pkg.exports["./server"], "./src/cache-engine.ts")
248
+ assert.equal(pkg.exports["./tui"], "./src/tui.mjs")
249
+ })
250
+
251
+
224
252
  // ===========================================================================
225
253
  // Provider-aware model detection
226
254
  // ===========================================================================
@@ -269,6 +297,32 @@ test("unrelated GLM models do NOT match GLM-5.3 policy", () => {
269
297
  assert.equal(detectPolicy(M("zai", "glm-4.5")), POLICY_NEUTRAL)
270
298
  })
271
299
 
300
+ test("MiMo V2.6 Flash/Pro match MiMo policy (openrouter + direct)", () => {
301
+ assert.equal(detectPolicy(M("openrouter", "xiaomi/mimo-v2.6-flash")), POLICY_MIMO26)
302
+ assert.equal(detectPolicy(M("openrouter", "xiaomi/mimo-v2.6-pro")), POLICY_MIMO26)
303
+ assert.equal(detectPolicy(M("xiaomi", "mimo-v2.6-flash")), POLICY_MIMO26)
304
+ assert.equal(detectPolicy(M("xiaomi", "mimo-v2.6-pro")), POLICY_MIMO26)
305
+ // full Model shape via api.id
306
+ assert.equal(
307
+ detectPolicy({ providerID: "openrouter", api: { id: "xiaomi/mimo-v2.6-flash", npm: "@openrouter/ai-sdk-provider" } }),
308
+ POLICY_MIMO26,
309
+ )
310
+ })
311
+
312
+ test("MiMo V2.5 and Pro-UltraSpeed do NOT match MiMo policy", () => {
313
+ assert.equal(detectPolicy(M("openrouter", "xiaomi/mimo-v2.5")), POLICY_NEUTRAL)
314
+ assert.equal(detectPolicy(M("openrouter", "xiaomi/mimo-v2.5-pro")), POLICY_NEUTRAL)
315
+ assert.equal(detectPolicy(M("xiaomi", "mimo-v2.5")), POLICY_NEUTRAL)
316
+ assert.equal(detectPolicy(M("xiaomi", "mimo-v2.6-pro-ultraspeed")), POLICY_NEUTRAL)
317
+ assert.equal(detectPolicy(M("xiaomi", "mimo-v2.6-flashx")), POLICY_NEUTRAL)
318
+ })
319
+
320
+ test("unrelated MiMo/other models do NOT match MiMo policy", () => {
321
+ assert.equal(detectPolicy(M("openrouter", "xiaomi/mimo-v2")), POLICY_NEUTRAL)
322
+ assert.equal(detectPolicy(M("openrouter", "xiaomi/mimo-v2-flash")), POLICY_NEUTRAL)
323
+ assert.equal(detectPolicy(M("anthropic", "claude-sonnet-4-5")), POLICY_NEUTRAL)
324
+ })
325
+
272
326
  test("unrelated models match neutral policy", () => {
273
327
  assert.equal(detectPolicy(M("anthropic", "claude-sonnet-4-5")), POLICY_NEUTRAL)
274
328
  assert.equal(detectPolicy(M("openrouter", "x-ai/grok-4")), POLICY_NEUTRAL)
@@ -535,10 +589,16 @@ test("empty current sequence -> no anomalies", () => {
535
589
  // Config: provider policies
536
590
  // ===========================================================================
537
591
 
538
- test("config defaults enable all three policies", () => {
592
+ test("config defaults enable all four policies", () => {
539
593
  const cfg = parseConfig({}, {})
540
594
  assert.deepEqual(cfg.policies.deepseek, { enabled: true })
541
595
  assert.deepEqual(cfg.policies.glm53, { enabled: true, stabilizeSystem: true, preserveThinkingIntegrity: true })
596
+ assert.deepEqual(cfg.policies.mimo26, {
597
+ enabled: true,
598
+ stabilizeSystem: true,
599
+ stickySession: true,
600
+ preserveThinkingIntegrity: true,
601
+ })
542
602
  assert.deepEqual(cfg.policies.gpt56, {
543
603
  enabled: true,
544
604
  promptCacheKey: true,
@@ -565,6 +625,7 @@ test("config policy overrides are honored", () => {
565
625
  },
566
626
  glm53: { stabilizeSystem: false },
567
627
  deepseek: { enabled: true },
628
+ mimo26: { enabled: false, stabilizeSystem: false, stickySession: false },
568
629
  },
569
630
  },
570
631
  {},
@@ -578,6 +639,10 @@ test("config policy overrides are honored", () => {
578
639
  assert.equal(cfg.policies.gpt56.ttl, "1h")
579
640
  assert.equal(cfg.policies.glm53.stabilizeSystem, false)
580
641
  assert.equal(cfg.policies.glm53.preserveThinkingIntegrity, true)
642
+ assert.equal(cfg.policies.mimo26.enabled, false)
643
+ assert.equal(cfg.policies.mimo26.stabilizeSystem, false)
644
+ assert.equal(cfg.policies.mimo26.stickySession, false)
645
+ assert.equal(cfg.policies.mimo26.preserveThinkingIntegrity, true)
581
646
  })
582
647
 
583
648
  test("config invalid policy values fall back to defaults", () => {
@@ -763,3 +828,143 @@ test("boundary: reasoning-integrity reason tokens", () => {
763
828
  ])
764
829
  assert.deepEqual(reasoningIssueReasons(null), [])
765
830
  })
831
+
832
+ // ===========================================================================
833
+ // MiMo-V2.6: system env relocation (shared content-preserving helper)
834
+ // ===========================================================================
835
+
836
+ const MIMO_SYSTEM = [
837
+ "You are a senior software engineer.",
838
+ "You are powered by the model named mimo-v2.6-flash. The exact model ID is xiaomi/mimo-v2.6-flash",
839
+ "Here is some useful information about the environment you are running in:",
840
+ "<env>",
841
+ "Working directory: /home/dev/project",
842
+ "Today's date: 2026-08-17",
843
+ "</env>",
844
+ "You MUST follow AGENTS.md instructions and keep your responses concise.",
845
+ ].join("\n")
846
+
847
+ test("MiMo env block is relocated to the tail, contents preserved exactly", () => {
848
+ const { text, changed } = relocateVolatileEnvBlock(MIMO_SYSTEM)
849
+ assert.equal(changed, true)
850
+ assert.ok(text.endsWith("</env>"))
851
+ const norm = (t) => t.split("\n").filter((l) => l).sort().join("\n")
852
+ assert.equal(norm(text), norm(MIMO_SYSTEM))
853
+ assert.ok(text.indexOf("You MUST follow") < text.indexOf("You are powered"))
854
+ })
855
+
856
+ test("MiMo relocation is a no-op when the block is already at the tail", () => {
857
+ const already = relocateVolatileEnvBlock(MIMO_SYSTEM).text
858
+ const again = relocateVolatileEnvBlock(already)
859
+ assert.equal(again.changed, false)
860
+ assert.equal(again.text, already)
861
+ })
862
+
863
+ test("MiMo relocation is a no-op when start/end markers are missing", () => {
864
+ const missingStart = "Just a system prompt.\n<env>\nToday's date: x\n</env>"
865
+ assert.equal(relocateVolatileEnvBlock(missingStart).changed, false)
866
+ const missingEnd = "You are powered by the model named mimo-v2.6-flash\nbut never closed"
867
+ assert.equal(relocateVolatileEnvBlock(missingEnd).changed, false)
868
+ })
869
+
870
+ // ===========================================================================
871
+ // MiMo-V2.6: sticky-session identity (pure, derived but not injected)
872
+ // ===========================================================================
873
+
874
+ test("mimoSessionIdFor: deterministic + stable for the same session", () => {
875
+ assert.equal(mimoSessionIdFor("ses_abc123"), mimoSessionIdFor("ses_abc123"))
876
+ })
877
+
878
+ test("mimoSessionIdFor: distinct sessions produce distinct ids", () => {
879
+ assert.notEqual(mimoSessionIdFor("ses_abc"), mimoSessionIdFor("ses_xyz"))
880
+ })
881
+
882
+ test("mimoSessionIdFor: printable, no whitespace, within 256 chars", () => {
883
+ const id = mimoSessionIdFor("ses_" + "a".repeat(500))
884
+ assert.ok(id.length <= 256)
885
+ assert.ok(!/\s/.test(id))
886
+ assert.ok(/^[\x20-\x7E]+$/.test(id))
887
+ })
888
+
889
+ test("mimoSessionIdFor: invalid input -> null (no fabrication)", () => {
890
+ assert.equal(mimoSessionIdFor(undefined), null)
891
+ assert.equal(mimoSessionIdFor(null), null)
892
+ assert.equal(mimoSessionIdFor(""), null)
893
+ assert.equal(mimoSessionIdFor(123), null)
894
+ })
895
+
896
+ test("mimoSessionIdFor: transient request fields cannot alter the id", () => {
897
+ // the helper is a pure function of the session id; extra args are ignored
898
+ assert.equal(mimoSessionIdFor("ses_stable"), mimoSessionIdFor("ses_stable", { turn: 7, temperature: 0.9 }))
899
+ })
900
+
901
+ // ===========================================================================
902
+ // MiMo-V2.6: cached/prompt token metrics
903
+ // ===========================================================================
904
+
905
+ test("mimoHitRate = cachedTokens / promptTokens", () => {
906
+ assert.equal(mimoHitRate(47000, 50000), 94)
907
+ assert.equal(mimoHitRate(0, 100), 0)
908
+ assert.equal(mimoHitRate(100, 100), 100)
909
+ })
910
+
911
+ test("mimoHitRate returns null for zero/unknown denominators (no fabrication)", () => {
912
+ assert.equal(mimoHitRate(0, 0), null)
913
+ assert.equal(mimoHitRate(10, 0), null)
914
+ assert.equal(mimoHitRate(NaN, 100), null)
915
+ assert.equal(mimoHitRate(10, NaN), null)
916
+ assert.equal(mimoHitRate(undefined, 100), null)
917
+ assert.equal(mimoHitRate(10, undefined), null)
918
+ })
919
+
920
+ test("mimoHitRate is NOT the read/(read+write) form", () => {
921
+ // read=80, input=20 -> promptTokens derived 100 -> 80%. The read/(read+write)
922
+ // form would give 100% for (80, 0); they must not be conflated.
923
+ assert.equal(mimoHitRate(80, 80 + 20), 80)
924
+ assert.equal(hitRatePct(80, 0), 100)
925
+ })
926
+
927
+ // ===========================================================================
928
+ // MiMo-V2.6: provider-switch diagnostics
929
+ // ===========================================================================
930
+
931
+ test("providerChangeEvent: no event until two real observations exist", () => {
932
+ assert.deepEqual(providerChangeEvent(null, { providerID: "openrouter", modelID: "xiaomi/mimo-v2.6-flash" }), {
933
+ changed: false,
934
+ from: null,
935
+ to: null,
936
+ })
937
+ assert.equal(providerChangeEvent({ providerID: "openrouter" }, null).changed, false)
938
+ })
939
+
940
+ test("providerChangeEvent: same provider -> no change", () => {
941
+ const a = { providerID: "openrouter", modelID: "xiaomi/mimo-v2.6-flash" }
942
+ const b = { providerID: "openrouter", modelID: "xiaomi/mimo-v2.6-pro" }
943
+ assert.equal(providerChangeEvent(a, b).changed, false)
944
+ })
945
+
946
+ test("providerChangeEvent: different provider -> change with from/to", () => {
947
+ const a = { providerID: "openrouter", modelID: "xiaomi/mimo-v2.6-flash" }
948
+ const b = { providerID: "xiaomi", modelID: "mimo-v2.6-flash" }
949
+ const ev = providerChangeEvent(a, b)
950
+ assert.equal(ev.changed, true)
951
+ assert.deepEqual(ev.from, a)
952
+ assert.deepEqual(ev.to, b)
953
+ })
954
+
955
+ // ===========================================================================
956
+ // MiMo-V2.6: no tool-definition mutation
957
+ // ===========================================================================
958
+
959
+ test("MiMo policy performs no tool mutation (fingerprints are pure inputs)", () => {
960
+ // The plugin never adds a MiMo tool-ordering pass: this runtime already sorts
961
+ // tools alphabetically before the wire. Classifying a model as MiMo must not
962
+ // affect tool fingerprints, which are a pure function of the tool definitions.
963
+ const model = { providerID: "openrouter", modelID: "xiaomi/mimo-v2.6-flash" }
964
+ assert.equal(detectPolicy(model), POLICY_MIMO26)
965
+ const before = { sem: toolFingerprint(TOOLS), wire: toolWireFingerprint(TOOLS) }
966
+ const after = { sem: toolFingerprint(TOOLS), wire: toolWireFingerprint(TOOLS) }
967
+ assert.deepEqual(after, before)
968
+ // semantic fingerprint stays order-insensitive regardless of the policy
969
+ assert.equal(toolFingerprint([...TOOLS].reverse()), before.sem)
970
+ })