opencode-cache-engine 0.1.1 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +322 -33
- package/examples/cache-engine.json +6 -0
- package/package.json +7 -2
- package/src/cache-engine-core.mjs +76 -9
- package/src/cache-engine.ts +124 -10
- package/src/tui.mjs +11 -0
- package/test/cache-engine.test.mjs +206 -1
package/README.md
CHANGED
|
@@ -1,20 +1,28 @@
|
|
|
1
1
|
# OpenCode Cache Engine
|
|
2
2
|
|
|
3
|
-
Provider-aware prompt-cache optimization and observability for
|
|
3
|
+
Provider-aware prompt-cache optimization and observability for
|
|
4
|
+
[OpenCode](https://opencode.ai).
|
|
5
|
+
|
|
6
|
+
`opencode-cache-engine` is an OpenCode npm plugin with two targets:
|
|
7
|
+
|
|
8
|
+
- **Server target** — the actual cache-engine runtime and provider policies.
|
|
9
|
+
- **TUI target** — registration with OpenCode's TUI plugin manager.
|
|
10
|
+
|
|
11
|
+
The server target handles cache optimization, prompt-shape diagnostics, compaction handling, and cache telemetry. The TUI target provides the plugin-manager integration and enable/disable state for the TUI-facing plugin entry.
|
|
4
12
|
|
|
5
13
|
`CacheEngine` is an OpenCode plugin designed for long-running agent sessions where prompt-cache efficiency affects both latency and cost. It keeps the harness conservative for providers whose cache behavior is already automatic, while applying provider-specific optimizations where the provider exposes useful cache controls or where prompt structure can be safely improved.
|
|
6
14
|
|
|
7
|
-
The plugin currently has
|
|
15
|
+
The plugin currently has four cache-policy families:
|
|
8
16
|
|
|
9
17
|
* **DeepSeek V4.1 Flash** — passive cache-stability and observability
|
|
10
18
|
* **GPT-5.6 Luna** — active cache-control configuration
|
|
11
19
|
* **GLM-5.3 Flash** — conservative system-prompt stabilization
|
|
20
|
+
* **MiMo-V2.6 (Flash / Pro)** — prefix stability and OpenRouter session-affinity diagnostics
|
|
12
21
|
|
|
13
22
|
The central design principle is:
|
|
14
23
|
|
|
15
24
|
> Optimize the request structure only when there is a clear provider-specific reason to do so. Otherwise, preserve OpenCode's native request behavior and measure what the provider actually reports.
|
|
16
25
|
|
|
17
|
-
---
|
|
18
26
|
|
|
19
27
|
## What this plugin does
|
|
20
28
|
|
|
@@ -28,12 +36,12 @@ It:
|
|
|
28
36
|
4. Records provider-reported cache token usage.
|
|
29
37
|
5. Adds a deterministic compaction continuation block.
|
|
30
38
|
6. Applies GPT-5.6 cache-control metadata.
|
|
31
|
-
7. Applies the GLM-5.3 volatile-environment relocation.
|
|
39
|
+
7. Applies the GLM-5.3 and MiMo-V2.6 volatile-environment relocation.
|
|
32
40
|
8. Records diagnostics that help determine whether prompt-shape changes correlate with cache behavior.
|
|
41
|
+
9. Records MiMo-V2.6 provider identity and provider-switch diagnostics.
|
|
33
42
|
|
|
34
43
|
The plugin deliberately avoids pretending that a local hash is proof of a provider cache hit. Provider-reported token usage remains the authoritative signal.
|
|
35
44
|
|
|
36
|
-
---
|
|
37
45
|
|
|
38
46
|
# Provider behavior
|
|
39
47
|
|
|
@@ -88,7 +96,6 @@ DeepSeek:
|
|
|
88
96
|
force cache behavior
|
|
89
97
|
```
|
|
90
98
|
|
|
91
|
-
---
|
|
92
99
|
|
|
93
100
|
## GPT-5.6 Luna
|
|
94
101
|
|
|
@@ -153,7 +160,6 @@ compaction:
|
|
|
153
160
|
|
|
154
161
|
This prevents a compaction-specific prompt from sharing the same GPT cache namespace as the normal live-session prompt. The behavior is deterministic and tested explicitly.
|
|
155
162
|
|
|
156
|
-
---
|
|
157
163
|
|
|
158
164
|
## GLM-5.3 Flash
|
|
159
165
|
|
|
@@ -219,17 +225,168 @@ It only occurs when:
|
|
|
219
225
|
|
|
220
226
|
The plugin does not arbitrarily rearrange unrelated prompt content.
|
|
221
227
|
|
|
222
|
-
---
|
|
223
228
|
|
|
224
|
-
|
|
229
|
+
## MiMo-V2.6 (Flash / Pro)
|
|
230
|
+
|
|
231
|
+
### Policy: prefix stability + OpenRouter session affinity
|
|
232
|
+
|
|
233
|
+
MiMo-V2.6 is Xiaomi's current model family. The plugin targets exactly two
|
|
234
|
+
identifiers:
|
|
235
|
+
|
|
236
|
+
* `xiaomi/mimo-v2.6-flash` / `mimo-v2.6-flash`
|
|
237
|
+
* `xiaomi/mimo-v2.6-pro` / `mimo-v2.6-pro`
|
|
238
|
+
|
|
239
|
+
Detection also tolerates `provider/model` shapes where `api.id` contains those
|
|
240
|
+
slugs. It deliberately does **not** match `mimo-v2.5`, `mimo-v2.5-pro`,
|
|
241
|
+
`mimo-v2.6-pro-ultraspeed`, or unrelated MiMo models.
|
|
242
|
+
|
|
243
|
+
### Implicit context caching
|
|
244
|
+
|
|
245
|
+
Xiaomi documents context caching for both V2.6 Flash and Pro, and exposes
|
|
246
|
+
`usage.prompt_tokens_details.cached_tokens` as the number of prompt tokens
|
|
247
|
+
served from cache. The V2.6 API documents implicit context caching, not a
|
|
248
|
+
user-supplied cache key or explicit breakpoint.
|
|
249
|
+
|
|
250
|
+
Accordingly the plugin **injects no cache-control parameter** for MiMo. It does
|
|
251
|
+
not send `promptCacheKey`, `cacheControl`, `cacheBreakpoint`, or `ttl`.
|
|
252
|
+
Implicit caching is the default assumption.
|
|
253
|
+
|
|
254
|
+
### Environment-block stabilization
|
|
255
|
+
|
|
256
|
+
MiMo uses the same narrow, content-preserving transformation as GLM-5.3: the
|
|
257
|
+
identifiable volatile `<env>` block is relocated to the **tail** of the single
|
|
258
|
+
system string. Contents are preserved byte-for-byte; only position changes. This
|
|
259
|
+
keeps the large reusable prefix stable when only the environment/date changes.
|
|
225
260
|
|
|
226
|
-
The
|
|
261
|
+
The transformation is applied only when:
|
|
227
262
|
|
|
228
|
-
|
|
229
|
-
|
|
230
|
-
|
|
231
|
-
|
|
232
|
-
|
|
263
|
+
* the selected model is MiMo-V2.6 Flash/Pro
|
|
264
|
+
* `mimo26.stabilizeSystem` is `true`
|
|
265
|
+
* there is exactly one system string
|
|
266
|
+
* the expected `<env>` markers exist and the block is identified unambiguously
|
|
267
|
+
* the block is not already at the tail
|
|
268
|
+
|
|
269
|
+
### No generic system-prompt freezing
|
|
270
|
+
|
|
271
|
+
MiMo-Code's own harness freezes its per-session system prefix. This plugin does
|
|
272
|
+
**not** copy that mechanism. System instructions can legitimately change because
|
|
273
|
+
of permissions, tools, agent mode, skills, MCP state, or project configuration;
|
|
274
|
+
a plugin-level snapshot must never override a legitimate change.
|
|
275
|
+
|
|
276
|
+
Instead the plugin:
|
|
277
|
+
|
|
278
|
+
* records a first-seen system baseline per session;
|
|
279
|
+
* computes the full system hash, stable prefix hash, and volatile suffix hash;
|
|
280
|
+
* records changes for MiMo sessions;
|
|
281
|
+
* allows the `<env>` relocation when that is the only identified volatility;
|
|
282
|
+
* reports other system changes diagnostically and never overwrites the new
|
|
283
|
+
content.
|
|
284
|
+
|
|
285
|
+
Explicit telemetry events:
|
|
286
|
+
|
|
287
|
+
* `mimo_system_env_relocated`
|
|
288
|
+
* `mimo_system_prefix_changed`
|
|
289
|
+
|
|
290
|
+
### OpenRouter sticky session — derived but not injected
|
|
291
|
+
|
|
292
|
+
OpenRouter documents a top-level `session_id` request field for sticky provider
|
|
293
|
+
routing, which keeps a session's requests on the same upstream provider so
|
|
294
|
+
provider-side prompt caches stay warm.
|
|
295
|
+
|
|
296
|
+
The plugin includes a pure, session-scoped derivation (`mimoSessionIdFor`):
|
|
297
|
+
deterministic, distinct per session, printable/no-whitespace, and well under the
|
|
298
|
+
256-character cap. However, **the derived id is not injected into requests**.
|
|
299
|
+
|
|
300
|
+
Rationale (verified against the installed runtime): OpenCode's OpenRouter
|
|
301
|
+
request adapter forwards only `usage`, `reasoning`, and `prompt_cache_key` from
|
|
302
|
+
provider options, and exposes no top-level `session_id` path. Adding an
|
|
303
|
+
unsupported field would be guessing, so the id is recorded as telemetry only,
|
|
304
|
+
and `mimo26.stickySession` currently gates that recording. If a future runtime
|
|
305
|
+
gains a verified `session_id` path, the helper is already in place.
|
|
306
|
+
|
|
307
|
+
Note that OpenCode itself sets `x-session-affinity` / `X-Session-Id` HTTP
|
|
308
|
+
headers for non-opencode providers, and can set a flat `promptCacheKey` for
|
|
309
|
+
OpenRouter when `setCacheKey: true` is configured. Those are HTTP routing
|
|
310
|
+
headers and an OpenAI-style cache key respectively — they are not OpenRouter's
|
|
311
|
+
documented body `session_id`.
|
|
312
|
+
|
|
313
|
+
### MiMo cache metrics
|
|
314
|
+
|
|
315
|
+
MiMo caches are provider-managed, so the authoritative metric is provider
|
|
316
|
+
reported. For MiMo the plugin emits the preferred ratio:
|
|
317
|
+
|
|
318
|
+
```text
|
|
319
|
+
cacheHitRate = cachedTokens / promptTokens
|
|
320
|
+
```
|
|
321
|
+
|
|
322
|
+
This is intentionally **not** the `read / (read + write)` form used by other
|
|
323
|
+
families. It is not GLM's `read / (read + write + input)` either.
|
|
324
|
+
|
|
325
|
+
Derivation: the runtime exposes assistant tokens as `{ input, output,
|
|
326
|
+
cache:{ read, write } }`, where `input` is the non-cached prompt input and
|
|
327
|
+
`cache.read` is the cached prompt input. Total prompt tokens are therefore
|
|
328
|
+
derived as `read + input`, and `cachedTokens = read`. `cache.write` is a
|
|
329
|
+
separate accounting bucket and is not folded in; no cache-write value is
|
|
330
|
+
fabricated, and the ratio is `null` when `promptTokens` is zero.
|
|
331
|
+
|
|
332
|
+
A MiMo usage record looks conceptually like:
|
|
333
|
+
|
|
334
|
+
```json
|
|
335
|
+
{
|
|
336
|
+
"kind": "usage",
|
|
337
|
+
"policy": "mimo26",
|
|
338
|
+
"provider": "openrouter",
|
|
339
|
+
"model": "xiaomi/mimo-v2.6-flash",
|
|
340
|
+
"promptTokens": 50000,
|
|
341
|
+
"cachedTokens": 47000,
|
|
342
|
+
"cacheHitRate": 94
|
|
343
|
+
}
|
|
344
|
+
```
|
|
345
|
+
|
|
346
|
+
### Provider-switch diagnostics
|
|
347
|
+
|
|
348
|
+
Because MiMo caches live at the provider side, a provider change within one
|
|
349
|
+
session can silently invalidate them. The plugin records provider identity on
|
|
350
|
+
every MiMo request and emits a `mimo_provider_changed` boundary event when the
|
|
351
|
+
OpenCode `providerID` changes within a session. It never forces or overrides the
|
|
352
|
+
user's provider selection.
|
|
353
|
+
|
|
354
|
+
Limitation: OpenRouter's *upstream* provider selection (for example
|
|
355
|
+
`xiaomi/fp8` vs `atlas-cloud/fp8`) is not exposed to plugins, so only the
|
|
356
|
+
OpenCode `providerID`/`modelID` are observable.
|
|
357
|
+
|
|
358
|
+
### Reasoning / thinking
|
|
359
|
+
|
|
360
|
+
MiMo-V2.6 supports deep thinking and reports reasoning tokens. The plugin does
|
|
361
|
+
not treat reasoning replay as a cache requirement: reasoning diagnostics are
|
|
362
|
+
instrumentation only, and the plugin never rewrites, duplicates, reorders, or
|
|
363
|
+
re-injects reasoning content, nor changes reasoning effort for caching.
|
|
364
|
+
|
|
365
|
+
### Skill-catalog / history limitation
|
|
366
|
+
|
|
367
|
+
MiMo-Code moved skill catalogs out of repeatedly rewritten user messages and
|
|
368
|
+
toward the system tail. In this OpenCode runtime the skill guidance
|
|
369
|
+
(`<available_skills>`) and MCP instructions already live in the **system
|
|
370
|
+
prefix**, not in user-message history. The plugin therefore performs no
|
|
371
|
+
message-history rewrite. Skill/MCP changes simply appear as system-prefix changes
|
|
372
|
+
and are reported diagnostically; the message content is left untouched.
|
|
373
|
+
|
|
374
|
+
|
|
375
|
+
# Provider comparison
|
|
376
|
+
|
|
377
|
+
| Provider | Detection | Prompt text changed? | Cache metadata changed? | Primary cache signal |
|
|
378
|
+
| ------------------- | ---------------------------------- | --------------------------- | ---------------------------------- | --------------------------------------- |
|
|
379
|
+
| DeepSeek V4.1 Flash | `deepseek` | No | No | provider `cache.read`/`cache.write` |
|
|
380
|
+
| GPT-5.6 Luna | `gpt-5.6*` on OpenAI-ish endpoints | No | Yes: `prompt_cache_key` + options | provider cache tokens |
|
|
381
|
+
| GLM-5.3 Flash | `glm-5.3*` | Yes, narrowly (`<env>` tail) | No provider cache key | provider cache tokens (GLM ratio) |
|
|
382
|
+
| MiMo-V2.6 Flash/Pro | `mimo-v2.6-flash` / `mimo-v2.6-pro` | Yes, narrowly (`<env>` tail) | No: implicit caching only | `cached_tokens / prompt_tokens` |
|
|
383
|
+
|
|
384
|
+
|
|
385
|
+
# Prompt-cache strategy
|
|
386
|
+
|
|
387
|
+
The plugin uses different strategies because cache mechanisms differ by provider.
|
|
388
|
+
The table above summarises them; the essential point is the distinction between
|
|
389
|
+
*changing prompt text* and *changing cache metadata*.
|
|
233
390
|
|
|
234
391
|
This distinction is fundamental.
|
|
235
392
|
|
|
@@ -237,7 +394,6 @@ The plugin is **not** a generic "rewrite every prompt for caching" engine.
|
|
|
237
394
|
|
|
238
395
|
It is a provider-aware cache policy engine.
|
|
239
396
|
|
|
240
|
-
---
|
|
241
397
|
|
|
242
398
|
# System-prompt diagnostics
|
|
243
399
|
|
|
@@ -261,7 +417,6 @@ It does **not** mean:
|
|
|
261
417
|
|
|
262
418
|
This distinction is intentional. Provider-reported cache token counts are the authoritative cache signal.
|
|
263
419
|
|
|
264
|
-
---
|
|
265
420
|
|
|
266
421
|
# Tool-definition diagnostics
|
|
267
422
|
|
|
@@ -294,7 +449,6 @@ This distinction matters because semantic equality and byte-level request equali
|
|
|
294
449
|
|
|
295
450
|
The plugin uses these fingerprints for **diagnostics only**. It does not reorder the tools to force a particular fingerprint.
|
|
296
451
|
|
|
297
|
-
---
|
|
298
452
|
|
|
299
453
|
# Compaction handling
|
|
300
454
|
|
|
@@ -313,7 +467,6 @@ The plugin adds a deterministic continuation template:
|
|
|
313
467
|
The digest is inserted once per compaction operation using a guard that prevents duplicate insertion if the compaction hook fires multiple times.
|
|
314
468
|
The objective is to provide a deterministic continuation structure rather than generating a different arbitrary cache-affecting block on every compaction.
|
|
315
469
|
|
|
316
|
-
---
|
|
317
470
|
|
|
318
471
|
# Cache metrics
|
|
319
472
|
|
|
@@ -350,6 +503,15 @@ read / (read + write + input)
|
|
|
350
503
|
|
|
351
504
|
as implemented by `glmHitRatio()`.
|
|
352
505
|
|
|
506
|
+
For MiMo, the implementation uses the provider-documented prompt-cache ratio:
|
|
507
|
+
|
|
508
|
+
```text
|
|
509
|
+
cacheHitRate = cachedTokens / promptTokens
|
|
510
|
+
```
|
|
511
|
+
|
|
512
|
+
implemented by `mimoHitRate()`. `hitRatePct()` itself is left untouched so other
|
|
513
|
+
providers are unaffected.
|
|
514
|
+
|
|
353
515
|
### Important metric distinction
|
|
354
516
|
|
|
355
517
|
These ratios answer different questions.
|
|
@@ -362,7 +524,12 @@ These ratios answer different questions.
|
|
|
362
524
|
|
|
363
525
|
> How much of the total prompt-token accounting was represented by cached reads?
|
|
364
526
|
|
|
365
|
-
|
|
527
|
+
`cachedTokens / promptTokens` (MiMo) answers:
|
|
528
|
+
|
|
529
|
+
> Of the prompt tokens the provider processed, what fraction was served from
|
|
530
|
+
> cache?
|
|
531
|
+
|
|
532
|
+
Do not treat these percentages as interchangeable.
|
|
366
533
|
|
|
367
534
|
---
|
|
368
535
|
|
|
@@ -388,7 +555,6 @@ Telemetry is best-effort.
|
|
|
388
555
|
|
|
389
556
|
A failed metrics write must never break an OpenCode request. The recorder catches write failures rather than allowing telemetry failures to affect execution.
|
|
390
557
|
|
|
391
|
-
---
|
|
392
558
|
|
|
393
559
|
# Metrics examples
|
|
394
560
|
|
|
@@ -446,8 +612,30 @@ Telemetry is intended to answer questions such as:
|
|
|
446
612
|
* Did a compaction occur?
|
|
447
613
|
* Which provider/model/policy was active?
|
|
448
614
|
* Did the GLM system stabilization actually change the observed prompt shape?
|
|
615
|
+
* Did MiMo's environment relocation fire (`mimo_system_env_relocated`)?
|
|
616
|
+
* Did MiMo's stable system prefix change (`mimo_system_prefix_changed`)?
|
|
617
|
+
* Did the MiMo provider change within a session (`mimo_provider_changed`)?
|
|
618
|
+
* What was MiMo's provider-reported cache hit rate (`cacheHitRate`)?
|
|
619
|
+
|
|
620
|
+
A MiMo usage record adds the provider-reported cache fields:
|
|
621
|
+
|
|
622
|
+
```json
|
|
623
|
+
{
|
|
624
|
+
"kind": "usage",
|
|
625
|
+
"sid": "session-id",
|
|
626
|
+
"ts": 1750000000000,
|
|
627
|
+
"policy": "mimo26",
|
|
628
|
+
"provider": "openrouter",
|
|
629
|
+
"model": "xiaomi/mimo-v2.6-flash",
|
|
630
|
+
"read": 47000,
|
|
631
|
+
"input": 3000,
|
|
632
|
+
"promptTokens": 50000,
|
|
633
|
+
"cachedTokens": 47000,
|
|
634
|
+
"cacheHitRate": 94,
|
|
635
|
+
"stickySessionId": "mimo-ses-0123456789abcdef"
|
|
636
|
+
}
|
|
637
|
+
```
|
|
449
638
|
|
|
450
|
-
---
|
|
451
639
|
|
|
452
640
|
# Configuration
|
|
453
641
|
|
|
@@ -476,6 +664,12 @@ The default configuration is:
|
|
|
476
664
|
"enabled": true,
|
|
477
665
|
"stabilizeSystem": true,
|
|
478
666
|
"preserveThinkingIntegrity": true
|
|
667
|
+
},
|
|
668
|
+
"mimo26": {
|
|
669
|
+
"enabled": true,
|
|
670
|
+
"stabilizeSystem": true,
|
|
671
|
+
"stickySession": true,
|
|
672
|
+
"preserveThinkingIntegrity": true
|
|
479
673
|
}
|
|
480
674
|
}
|
|
481
675
|
}
|
|
@@ -483,7 +677,6 @@ The default configuration is:
|
|
|
483
677
|
|
|
484
678
|
The configuration parser starts from these defaults and applies valid file/environment overrides without mutating the caller's configuration object.
|
|
485
679
|
|
|
486
|
-
---
|
|
487
680
|
|
|
488
681
|
# Configuration options
|
|
489
682
|
|
|
@@ -535,7 +728,6 @@ Controls whether the deterministic compaction continuation block is inserted.
|
|
|
535
728
|
|
|
536
729
|
Controls warning logs for observed prefix-shape changes.
|
|
537
730
|
|
|
538
|
-
---
|
|
539
731
|
|
|
540
732
|
# DeepSeek configuration
|
|
541
733
|
|
|
@@ -549,7 +741,6 @@ There are intentionally very few settings here.
|
|
|
549
741
|
|
|
550
742
|
DeepSeek is treated as the conservative/passive policy.
|
|
551
743
|
|
|
552
|
-
---
|
|
553
744
|
|
|
554
745
|
# GPT-5.6 configuration
|
|
555
746
|
|
|
@@ -601,7 +792,6 @@ Defaults to:
|
|
|
601
792
|
|
|
602
793
|
Existing request options are not overwritten by the plugin.
|
|
603
794
|
|
|
604
|
-
---
|
|
605
795
|
|
|
606
796
|
# GLM-5.3 configuration
|
|
607
797
|
|
|
@@ -629,7 +819,45 @@ The reasoning instrumentation is intended to identify anomalies such as:
|
|
|
629
819
|
|
|
630
820
|
It is diagnostic rather than a reason to rewrite or fabricate reasoning content. The implementation maps these conditions to explicit diagnostic reasons.
|
|
631
821
|
|
|
632
|
-
|
|
822
|
+
|
|
823
|
+
# MiMo-V2.6 configuration
|
|
824
|
+
|
|
825
|
+
```json
|
|
826
|
+
{
|
|
827
|
+
"mimo26": {
|
|
828
|
+
"enabled": true,
|
|
829
|
+
"stabilizeSystem": true,
|
|
830
|
+
"stickySession": true,
|
|
831
|
+
"preserveThinkingIntegrity": true
|
|
832
|
+
}
|
|
833
|
+
}
|
|
834
|
+
```
|
|
835
|
+
|
|
836
|
+
### `enabled`
|
|
837
|
+
|
|
838
|
+
Enables the MiMo-V2.6 policy.
|
|
839
|
+
|
|
840
|
+
### `stabilizeSystem`
|
|
841
|
+
|
|
842
|
+
Enables relocation of the volatile `<env>` section to the system-prompt tail
|
|
843
|
+
(same narrow, content-preserving transformation as GLM-5.3).
|
|
844
|
+
|
|
845
|
+
### `stickySession`
|
|
846
|
+
|
|
847
|
+
Gates derivation/recording of the OpenRouter sticky-session id
|
|
848
|
+
(`mimoSessionIdFor`). The id is recorded as telemetry; it is **not** injected
|
|
849
|
+
into the request because this runtime exposes no verified OpenRouter top-level
|
|
850
|
+
`session_id` path. See "OpenRouter sticky session — derived but not injected".
|
|
851
|
+
|
|
852
|
+
### `preserveThinkingIntegrity`
|
|
853
|
+
|
|
854
|
+
Enables reasoning diagnostics as instrumentation. It never rewrites, duplicates,
|
|
855
|
+
reorders, or re-injects reasoning content, and it is not a cache requirement.
|
|
856
|
+
|
|
857
|
+
No `cacheBlockSize`, `cacheTTL`, `cacheBreakpoint`, or `minimumCacheTokens`
|
|
858
|
+
knobs are exposed: those values are not established by authoritative V2.6
|
|
859
|
+
documentation.
|
|
860
|
+
|
|
633
861
|
|
|
634
862
|
# Model detection
|
|
635
863
|
|
|
@@ -639,6 +867,7 @@ The plugin classifies requests into:
|
|
|
639
867
|
deepseek
|
|
640
868
|
gpt56
|
|
641
869
|
glm53
|
|
870
|
+
mimo26
|
|
642
871
|
neutral
|
|
643
872
|
```
|
|
644
873
|
|
|
@@ -647,9 +876,13 @@ The model detector recognizes:
|
|
|
647
876
|
* DeepSeek model/provider identifiers
|
|
648
877
|
* GPT-5.6 variants
|
|
649
878
|
* GLM-5.3 variants
|
|
879
|
+
* MiMo-V2.6 Flash and Pro (`xiaomi/mimo-v2.6-flash`, `mimo-v2.6-pro`, ...)
|
|
650
880
|
|
|
651
881
|
GPT-5.6 has an additional OpenAI/Azure-context check so a string containing `gpt-5.6` does not automatically cause GPT-specific fields to be sent to an unrelated endpoint.
|
|
652
882
|
|
|
883
|
+
MiMo detection targets exactly Flash and Pro: it excludes `mimo-v2.5`,
|
|
884
|
+
`mimo-v2.5-pro`, and `mimo-v2.6-pro-ultraspeed`.
|
|
885
|
+
|
|
653
886
|
Unknown models use the neutral policy.
|
|
654
887
|
|
|
655
888
|
Neutral means:
|
|
@@ -658,7 +891,6 @@ Neutral means:
|
|
|
658
891
|
no provider-specific request mutation
|
|
659
892
|
```
|
|
660
893
|
|
|
661
|
-
---
|
|
662
894
|
|
|
663
895
|
# OpenRouter usage
|
|
664
896
|
|
|
@@ -666,11 +898,10 @@ This plugin is compatible with OpenRouter because the cache policy is based on t
|
|
|
666
898
|
|
|
667
899
|
For cache-sensitive workloads, provider stability remains important.
|
|
668
900
|
|
|
669
|
-
The plugin does not attempt to compensate for provider switching by rewriting prompts.
|
|
901
|
+
The plugin does not attempt to compensate for provider switching by rewriting prompts. It records MiMo provider identity and provider-switch diagnostics so routing instability is at least observable.
|
|
670
902
|
|
|
671
903
|
For that reason, a stable provider route is preferable when your goal is to measure and maximize prefix reuse.
|
|
672
904
|
|
|
673
|
-
---
|
|
674
905
|
|
|
675
906
|
# Architecture
|
|
676
907
|
|
|
@@ -893,14 +1124,15 @@ A typical standalone repository can use:
|
|
|
893
1124
|
opencode-cache-engine/
|
|
894
1125
|
├── src/
|
|
895
1126
|
│ ├── cache-engine.ts
|
|
896
|
-
│
|
|
1127
|
+
│ ├── cache-engine-core.mjs
|
|
1128
|
+
│ └── tui.mjs
|
|
897
1129
|
├── test/
|
|
898
1130
|
│ └── cache-engine.test.mjs
|
|
899
1131
|
├── examples/
|
|
900
1132
|
│ └── cache-engine.json
|
|
1133
|
+
├── package.json
|
|
901
1134
|
├── README.md
|
|
902
|
-
|
|
903
|
-
└── package.json
|
|
1135
|
+
└── LICENSE
|
|
904
1136
|
```
|
|
905
1137
|
|
|
906
1138
|
The OpenCode plugin export remains:
|
|
@@ -925,6 +1157,20 @@ without changing the `CacheEngine` export identifier.
|
|
|
925
1157
|
|
|
926
1158
|
Install the plugin into the OpenCode plugins directory according to your OpenCode plugin-loading setup.
|
|
927
1159
|
|
|
1160
|
+
`opencode-cache-engine` is distributed as an npm package.
|
|
1161
|
+
|
|
1162
|
+
## Server/runtime plugin
|
|
1163
|
+
|
|
1164
|
+
Add the package to the OpenCode runtime plugin configuration:
|
|
1165
|
+
|
|
1166
|
+
```json
|
|
1167
|
+
{
|
|
1168
|
+
"plugin": [
|
|
1169
|
+
"opencode-cache-engine"
|
|
1170
|
+
]
|
|
1171
|
+
}
|
|
1172
|
+
```
|
|
1173
|
+
|
|
928
1174
|
The runtime entry should expose:
|
|
929
1175
|
|
|
930
1176
|
```ts
|
|
@@ -959,6 +1205,12 @@ GPT-5.6:
|
|
|
959
1205
|
|
|
960
1206
|
GLM-5.3:
|
|
961
1207
|
volatile env block relocated when eligible
|
|
1208
|
+
|
|
1209
|
+
MiMo-V2.6:
|
|
1210
|
+
volatile env block relocated when eligible
|
|
1211
|
+
no GPT/GLM-only cache fields present
|
|
1212
|
+
no OpenRouter top-level session_id injected (unsupported by this runtime)
|
|
1213
|
+
telemetry carries provider/model/promptTokens/cachedTokens/cacheHitRate
|
|
962
1214
|
```
|
|
963
1215
|
|
|
964
1216
|
---
|
|
@@ -1001,6 +1253,42 @@ The relevant block must contain the expected beginning and closing marker, and t
|
|
|
1001
1253
|
|
|
1002
1254
|
---
|
|
1003
1255
|
|
|
1256
|
+
## MiMo-V2.6 prompt is not being changed
|
|
1257
|
+
|
|
1258
|
+
MiMo uses the same eligibility rules as GLM-5.3: exactly one system string, both
|
|
1259
|
+
`<env>` markers present, block identified unambiguously, and
|
|
1260
|
+
`mimo26.stabilizeSystem` enabled. If the block is already at the tail, the
|
|
1261
|
+
operation is a no-op.
|
|
1262
|
+
|
|
1263
|
+
---
|
|
1264
|
+
|
|
1265
|
+
## MiMo provider is not classified as `mimo26`
|
|
1266
|
+
|
|
1267
|
+
Verify the model identifier is exactly Flash or Pro:
|
|
1268
|
+
|
|
1269
|
+
```text
|
|
1270
|
+
mimo-v2.6-flash
|
|
1271
|
+
mimo-v2.6-pro
|
|
1272
|
+
xiaomi/mimo-v2.6-flash
|
|
1273
|
+
xiaomi/mimo-v2.6-pro
|
|
1274
|
+
```
|
|
1275
|
+
|
|
1276
|
+
`mimo-v2.5`, `mimo-v2.5-pro`, and `mimo-v2.6-pro-ultraspeed` are intentionally
|
|
1277
|
+
not matched.
|
|
1278
|
+
|
|
1279
|
+
---
|
|
1280
|
+
|
|
1281
|
+
## No OpenRouter `session_id` is sent for MiMo
|
|
1282
|
+
|
|
1283
|
+
This is expected. The installed OpenCode runtime's OpenRouter request adapter
|
|
1284
|
+
forwards only `usage`, `reasoning`, and `prompt_cache_key` from provider options
|
|
1285
|
+
and exposes no top-level `session_id` path. The plugin derives a stable
|
|
1286
|
+
`stickySessionId` and records it as telemetry, but does not inject it rather than
|
|
1287
|
+
send an unsupported field. This may change if a future runtime exposes a verified
|
|
1288
|
+
path.
|
|
1289
|
+
|
|
1290
|
+
---
|
|
1291
|
+
|
|
1004
1292
|
## Metrics file is missing
|
|
1005
1293
|
|
|
1006
1294
|
Telemetry is best-effort.
|
|
@@ -1025,6 +1313,7 @@ The current implementation is intentionally conservative:
|
|
|
1025
1313
|
DeepSeek -> preserve and measure
|
|
1026
1314
|
GPT-5.6 -> configure cache controls
|
|
1027
1315
|
GLM-5.3 -> isolate volatile prompt content
|
|
1316
|
+
MiMo-V2.6 -> stabilize prefix + observe provider/cache reality
|
|
1028
1317
|
```
|
|
1029
1318
|
|
|
1030
1319
|
That separation is the core design of the project.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "opencode-cache-engine",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.3.0",
|
|
4
4
|
"private": false,
|
|
5
5
|
"description": "Provider-aware prompt-cache optimization and telemetry for OpenCode",
|
|
6
6
|
"keywords": [
|
|
@@ -9,7 +9,8 @@
|
|
|
9
9
|
"prompt-cache",
|
|
10
10
|
"deepseek",
|
|
11
11
|
"glm",
|
|
12
|
-
"gpt"
|
|
12
|
+
"gpt",
|
|
13
|
+
"mimo"
|
|
13
14
|
],
|
|
14
15
|
"homepage": "https://github.com/AlexJoaquimPereira/opencode-cache-engine#readme",
|
|
15
16
|
"bugs": {
|
|
@@ -27,6 +28,10 @@
|
|
|
27
28
|
"example": "examples",
|
|
28
29
|
"test": "test"
|
|
29
30
|
},
|
|
31
|
+
"exports": {
|
|
32
|
+
"./server": "./src/cache-engine.ts",
|
|
33
|
+
"./tui": "./src/tui.mjs"
|
|
34
|
+
},
|
|
30
35
|
"scripts": {
|
|
31
36
|
"test": "node --test test/cache-engine.test.mjs",
|
|
32
37
|
"postpublish": "VERSION=$(node -p \"require('./package').version\") && TAG=v$VERSION && echo \"Creating git tag $TAG for npm $VERSION\" && git tag $TAG && git push origin $TAG || echo \"Tag $TAG may already exist or push failed; continuing.\"",
|
|
@@ -5,9 +5,10 @@
|
|
|
5
5
|
// TypeScript compiler. The plugin entry (cache-engine.ts) imports this module.
|
|
6
6
|
//
|
|
7
7
|
// This module is PROVIDER-AWARE: it classifies a model into a cache-policy
|
|
8
|
-
// family (deepseek | gpt56 | glm53 | neutral) and exposes small pure
|
|
9
|
-
// each family's strategy. The plugin entry (cache-engine.ts) remains
|
|
10
|
-
// place that touches OpenCode hooks; every decision here is testable in
|
|
8
|
+
// family (deepseek | gpt56 | glm53 | mimo26 | neutral) and exposes small pure
|
|
9
|
+
// helpers for each family's strategy. The plugin entry (cache-engine.ts) remains
|
|
10
|
+
// the only place that touches OpenCode hooks; every decision here is testable in
|
|
11
|
+
// Node.
|
|
11
12
|
//
|
|
12
13
|
// Terminology note: these functions deal with the *observed* system/tool
|
|
13
14
|
// prefix shape. An observed change means the request's prefix bytes changed; it
|
|
@@ -35,6 +36,7 @@ export const DIGEST_TEMPLATE = `## Session digest (cache-stable continuation blo
|
|
|
35
36
|
export const POLICY_DEEPSEEK = "deepseek"
|
|
36
37
|
export const POLICY_GPT56 = "gpt56"
|
|
37
38
|
export const POLICY_GLM53 = "glm53"
|
|
39
|
+
export const POLICY_MIMO26 = "mimo26"
|
|
38
40
|
export const POLICY_NEUTRAL = "neutral"
|
|
39
41
|
|
|
40
42
|
export const GPT56_DEFAULT_TTL = "30m"
|
|
@@ -65,6 +67,15 @@ function defaultPolicies() {
|
|
|
65
67
|
stabilizeSystem: true,
|
|
66
68
|
preserveThinkingIntegrity: true,
|
|
67
69
|
},
|
|
70
|
+
mimo26: {
|
|
71
|
+
enabled: true,
|
|
72
|
+
stabilizeSystem: true,
|
|
73
|
+
stickySession: true,
|
|
74
|
+
// MiMo reasoning diagnostics are instrumentation only. Unlike GLM
|
|
75
|
+
// preserved thinking, there is no evidence that MiMo prompt-cache reuse
|
|
76
|
+
// depends on reasoning replay, so this never rewrites reasoning content.
|
|
77
|
+
preserveThinkingIntegrity: true,
|
|
78
|
+
},
|
|
68
79
|
}
|
|
69
80
|
}
|
|
70
81
|
|
|
@@ -118,7 +129,7 @@ export function parseConfig(raw, env) {
|
|
|
118
129
|
if (typeof raw.logPrefixChanges === "boolean") cfg.logPrefixChanges = raw.logPrefixChanges
|
|
119
130
|
if (raw.policies && typeof raw.policies === "object") {
|
|
120
131
|
const d = defaultPolicies()
|
|
121
|
-
for (const fam of ["deepseek", "gpt56", "glm53"]) {
|
|
132
|
+
for (const fam of ["deepseek", "gpt56", "glm53", "mimo26"]) {
|
|
122
133
|
if (raw.policies[fam]) cfg.policies[fam] = parsePolicy(raw.policies[fam], d[fam])
|
|
123
134
|
}
|
|
124
135
|
}
|
|
@@ -232,6 +243,10 @@ function isOpenAIish(s) {
|
|
|
232
243
|
const GPT56_RE = /gpt-5\.6(?![\d.])/i
|
|
233
244
|
// GLM 5.3 family only (not glm-4.x / glm-4.6 etc).
|
|
234
245
|
const GLM53_RE = /glm-5\.3(?![\d.])/i
|
|
246
|
+
// Xiaomi MiMo V2.6 explicitly targets Flash + Pro only. The trailing
|
|
247
|
+
// (?![\w-]) guard prevents matching a hypothetical "mimo-v2.6-pro-ultraspeed"
|
|
248
|
+
// or "mimo-v2.6-flashx", and the v2\.6 literal excludes V2.5 / V2.
|
|
249
|
+
const MIMO26_RE = /mimo-v2\.6-(flash|pro)(?![\w-])/i
|
|
235
250
|
const DEEPSEEK_RE = /deepseek/i
|
|
236
251
|
|
|
237
252
|
// Pure classifier. Returns one of the POLICY_* keys. `model` may be a full
|
|
@@ -242,6 +257,7 @@ export function detectPolicy(model) {
|
|
|
242
257
|
if (!s.slug) return POLICY_NEUTRAL
|
|
243
258
|
if (GPT56_RE.test(s.slug) && isOpenAIish(s)) return POLICY_GPT56
|
|
244
259
|
if (GLM53_RE.test(s.slug)) return POLICY_GLM53
|
|
260
|
+
if (MIMO26_RE.test(s.slug)) return POLICY_MIMO26
|
|
245
261
|
if (DEEPSEEK_RE.test(s.slug) || DEEPSEEK_RE.test(s.providerID)) return POLICY_DEEPSEEK
|
|
246
262
|
return POLICY_NEUTRAL
|
|
247
263
|
}
|
|
@@ -284,16 +300,17 @@ export function gptCacheOptionsDelta(existingOptions, { key, mode = GPT56_DEFAUL
|
|
|
284
300
|
|
|
285
301
|
// The OpenCode system string begins with the agent prompt, then an env block
|
|
286
302
|
// ("You are powered by the model named ... Today's date: ... </env>") whose
|
|
287
|
-
// only per-day volatile byte is the date line. For GLM-5.3
|
|
288
|
-
// whole identifiable env block to the END of the system string so
|
|
289
|
-
// change only invalidates the tail of the prompt, leaving the long
|
|
290
|
-
// prefix intact. Content is preserved byte-for-byte (only position
|
|
303
|
+
// only per-day volatile byte is the date line. For GLM-5.3 and MiMo-V2.6 we
|
|
304
|
+
// relocate that whole identifiable env block to the END of the system string so
|
|
305
|
+
// a daily date change only invalidates the tail of the prompt, leaving the long
|
|
306
|
+
// stable prefix intact. Content is preserved byte-for-byte (only position
|
|
307
|
+
// changes).
|
|
291
308
|
//
|
|
292
309
|
// Returns { text, changed }. When the block cannot be identified unambiguously,
|
|
293
310
|
// returns the input unchanged (changed:false). This is a content-preserving
|
|
294
311
|
// reorder of clearly volatile metadata only -- it never reorders arbitrary
|
|
295
312
|
// instructions. This function is ONLY applied when the caller has already
|
|
296
|
-
// classified the model
|
|
313
|
+
// classified the model into a family that opts into system stabilization.
|
|
297
314
|
export function relocateVolatileEnvBlock(text) {
|
|
298
315
|
if (typeof text !== "string") return { text, changed: false }
|
|
299
316
|
const START = "You are powered by the model named "
|
|
@@ -438,6 +455,56 @@ export function glmHitRatio(read, write, input) {
|
|
|
438
455
|
return Math.round((100 * read) / denom)
|
|
439
456
|
}
|
|
440
457
|
|
|
458
|
+
// ---------------------------------------------------------------------------
|
|
459
|
+
// MiMo-V2.6 cache metrics + sticky-session identity
|
|
460
|
+
//
|
|
461
|
+
// MiMo caching is provider-managed (implicit context caching). Xiaomi documents
|
|
462
|
+
// usage.prompt_tokens_details.cached_tokens as the number of PROMPT tokens
|
|
463
|
+
// served from cache and prompt_tokens as the total prompt-token count, so the
|
|
464
|
+
// authoritative cache metric is cachedTokens / promptTokens -- NOT the
|
|
465
|
+
// read/(read+write) form used by other families. `hitRatePct` is intentionally
|
|
466
|
+
// left untouched so existing providers are unaffected.
|
|
467
|
+
//
|
|
468
|
+
// The runtime exposes `Message.info.tokens` as { input, output, cache:{read,
|
|
469
|
+
// write} } where `input` is the NON-cached prompt input and `cache.read` is the
|
|
470
|
+
// cached prompt input. Total prompt tokens are therefore derived as
|
|
471
|
+
// read + input (cache.write is a separate accounting bucket and is NOT folded
|
|
472
|
+
// in). We never fabricate cache-write values.
|
|
473
|
+
// ---------------------------------------------------------------------------
|
|
474
|
+
|
|
475
|
+
// cachedTokens / promptTokens, rounded to a percentage. Returns null when the
|
|
476
|
+
// denominator is unknown/zero or the inputs are not finite numbers, so no
|
|
477
|
+
// fabricated hit rate is ever emitted.
|
|
478
|
+
export function mimoHitRate(cachedTokens, promptTokens) {
|
|
479
|
+
if (!Number.isFinite(cachedTokens) || !Number.isFinite(promptTokens)) return null
|
|
480
|
+
if (promptTokens <= 0 || cachedTokens < 0) return null
|
|
481
|
+
return Math.round((100 * cachedTokens) / promptTokens)
|
|
482
|
+
}
|
|
483
|
+
|
|
484
|
+
// Derive a stable, session-scoped identifier suitable for OpenRouter's
|
|
485
|
+
// documented `session_id` sticky-routing key. Pure function of the OpenCode
|
|
486
|
+
// session id only: identical sessions map to identical ids, distinct sessions
|
|
487
|
+
// map to distinct ids, and transient request contents cannot influence it.
|
|
488
|
+
// The value is printable, contains no whitespace, and is far below the 256-char
|
|
489
|
+
// cap (25 chars). NOTE: this runtime's OpenRouter request adapter does not emit
|
|
490
|
+
// a top-level `session_id` (it forwards only usage/reasoning/prompt_cache_key),
|
|
491
|
+
// so this id is currently recorded as telemetry only and is never injected.
|
|
492
|
+
export function mimoSessionIdFor(sessionID) {
|
|
493
|
+
if (typeof sessionID !== "string" || sessionID.length === 0) return null
|
|
494
|
+
return `mimo-ses-${shorthash(sessionID)}`
|
|
495
|
+
}
|
|
496
|
+
|
|
497
|
+
// Detect a provider switch within the same session. `previous` and `current`
|
|
498
|
+
// are {providerID, modelID} observations. Returns {changed:false} until two
|
|
499
|
+
// real observations exist; a change is only reported when both are known and
|
|
500
|
+
// the providerID differs. Never forces or overrides provider selection.
|
|
501
|
+
export function providerChangeEvent(previous, current) {
|
|
502
|
+
const prev = previous && typeof previous.providerID === "string" ? previous.providerID : null
|
|
503
|
+
const cur = current && typeof current.providerID === "string" ? current.providerID : null
|
|
504
|
+
if (prev == null || cur == null || prev === cur) return { changed: false, from: null, to: null }
|
|
505
|
+
return { changed: true, from: previous, to: current }
|
|
506
|
+
}
|
|
507
|
+
|
|
441
508
|
// Decide whether a `usage` record should be emitted for an aggregation sample.
|
|
442
509
|
// We must not fabricate a zero-valued cache event merely because the session
|
|
443
510
|
// became idle: a sample only counts when at least one assistant message with
|
package/src/cache-engine.ts
CHANGED
|
@@ -5,6 +5,7 @@ import {
|
|
|
5
5
|
DIGEST_TEMPLATE,
|
|
6
6
|
POLICY_GLM53,
|
|
7
7
|
POLICY_GPT56,
|
|
8
|
+
POLICY_MIMO26,
|
|
8
9
|
POLICY_NEUTRAL,
|
|
9
10
|
createRecorder,
|
|
10
11
|
detectPolicy,
|
|
@@ -16,10 +17,13 @@ import {
|
|
|
16
17
|
gptCacheOptionsDelta,
|
|
17
18
|
hitRatePct,
|
|
18
19
|
loadConfig,
|
|
20
|
+
mimoHitRate,
|
|
21
|
+
mimoSessionIdFor,
|
|
19
22
|
nextProcessedCursor,
|
|
20
23
|
observeReasoningEffort,
|
|
21
24
|
policyEnabled,
|
|
22
25
|
prefixChangeReasons,
|
|
26
|
+
providerChangeEvent,
|
|
23
27
|
reasoningEffortFromOptions,
|
|
24
28
|
reasoningIssueReasons,
|
|
25
29
|
relocateVolatileEnvBlock,
|
|
@@ -37,7 +41,7 @@ import {
|
|
|
37
41
|
// cache-engine
|
|
38
42
|
//
|
|
39
43
|
// Provider-aware prompt-cache observability + conservative cache-shape
|
|
40
|
-
// preservation for ONE OpenCode TUI across
|
|
44
|
+
// preservation for ONE OpenCode TUI across four model families:
|
|
41
45
|
//
|
|
42
46
|
// DeepSeek V4.1 Flash -> pure passive. >99.66% hit rate is preserved by never
|
|
43
47
|
// mutating system/options/requests. Observability only.
|
|
@@ -52,16 +56,28 @@ import {
|
|
|
52
56
|
// history stable, and INSTRUMENT preserved-thinking
|
|
53
57
|
// integrity (duplicate/reorder/modified reasoning). No
|
|
54
58
|
// invented cache key (Z.ai exposes none).
|
|
59
|
+
// MiMo-V2.6 -> prefix stability + OpenRouter session affinity. The
|
|
60
|
+
// safe env-block relocation is applied (stabilizeSystem)
|
|
61
|
+
// and provider switches within a session are diagnosed.
|
|
62
|
+
// MiMo caching is provider-managed implicit context
|
|
63
|
+
// caching; no cache key/breakpoint/TTL is invented. The
|
|
64
|
+
// OpenRouter sticky-session id is derived but currently
|
|
65
|
+
// NOT injected: this runtime's OpenRouter request
|
|
66
|
+
// adapter forwards only usage/reasoning/prompt_cache_key
|
|
67
|
+
// and exposes no top-level `session_id` path (verified
|
|
68
|
+
// against the installed runtime; see README).
|
|
55
69
|
//
|
|
56
70
|
// The engine remains conservative: it observes, hashes, compares, records,
|
|
57
71
|
// appends a compaction continuation template, and (for GPT-5.6 only) injects
|
|
58
72
|
// documented cache options. It never rewrites message history, reorders tools,
|
|
59
73
|
// or alters user content. DeepSeek and neutral models are byte-untouched.
|
|
74
|
+
// MiMo/GLM only relocate the identifiable volatile env block, content-preserving.
|
|
60
75
|
//
|
|
61
76
|
// IMPORTANT (terminology): local hashes describe the *observed* prefix shape.
|
|
62
77
|
// A changed hash means request bytes changed; it is NOT proof the provider's
|
|
63
78
|
// cache key changed or that a cache miss occurred. Provider-reported cache
|
|
64
|
-
// token counts are authoritative; hashes are diagnostics only.
|
|
79
|
+
// token counts are authoritative; hashes are diagnostics only. For MiMo, the
|
|
80
|
+
// authoritative cache signal is provider-reported cached_tokens.
|
|
65
81
|
// ---------------------------------------------------------------------------
|
|
66
82
|
|
|
67
83
|
const TOOL_FETCH_TTL_MS = 1500
|
|
@@ -113,6 +129,7 @@ type SessionState = {
|
|
|
113
129
|
cacheRootAt: number | null
|
|
114
130
|
reasoningSeen: Map<string, number>
|
|
115
131
|
reasoningLastSeq: string[] | null
|
|
132
|
+
mimoProvider: { providerID: string; modelID: string } | null
|
|
116
133
|
}
|
|
117
134
|
|
|
118
135
|
const emptyShape = (): Shape => ({
|
|
@@ -154,6 +171,7 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
154
171
|
cacheRootAt: null,
|
|
155
172
|
reasoningSeen: new Map(),
|
|
156
173
|
reasoningLastSeq: null,
|
|
174
|
+
mimoProvider: null,
|
|
157
175
|
}
|
|
158
176
|
sessions.set(sid, s)
|
|
159
177
|
}
|
|
@@ -397,6 +415,25 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
397
415
|
recFields.promptTokens = read + write + input
|
|
398
416
|
recFields.glmHitRate = glmHitRatio(read, write, input)
|
|
399
417
|
}
|
|
418
|
+
if (family === POLICY_MIMO26) {
|
|
419
|
+
// Provider-reported cached tokens / total prompt tokens. The runtime's
|
|
420
|
+
// `input` is the non-cached prompt input and `cache.read` is the
|
|
421
|
+
// cached prompt input, so total prompt tokens are derived as read +
|
|
422
|
+
// input (cache.write is a separate accounting bucket). No cache-write
|
|
423
|
+
// value is fabricated; the ratio is null when prompt tokens are 0.
|
|
424
|
+
const promptTokens = read + input
|
|
425
|
+
recFields.promptTokens = promptTokens
|
|
426
|
+
recFields.cachedTokens = read
|
|
427
|
+
recFields.cacheHitRate = mimoHitRate(read, promptTokens)
|
|
428
|
+
if (cfg.policies?.[POLICY_MIMO26]?.stickySession === true) {
|
|
429
|
+
recFields.stickySessionId = mimoSessionIdFor(sid)
|
|
430
|
+
}
|
|
431
|
+
// Prefer the latest live provider identity over the latched one.
|
|
432
|
+
if (s.mimoProvider) {
|
|
433
|
+
recFields.provider = s.mimoProvider.providerID
|
|
434
|
+
recFields.model = s.mimoProvider.modelID
|
|
435
|
+
}
|
|
436
|
+
}
|
|
400
437
|
if (family === POLICY_GPT56 && s.gptInjected) {
|
|
401
438
|
recFields.keyStrategy = "session"
|
|
402
439
|
recFields.mode = cfg.policies?.[POLICY_GPT56]?.mode
|
|
@@ -445,6 +482,48 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
445
482
|
try {
|
|
446
483
|
const info = rememberModel(input.sessionID, input.model as unknown as ChatParamsModel)
|
|
447
484
|
const family = info?.family
|
|
485
|
+
|
|
486
|
+
// ---- MiMo-V2.6: provider-switch diagnostics (telemetry only) ---------
|
|
487
|
+
// MiMo cache lives at the provider side, so a provider change within one
|
|
488
|
+
// OpenCode session can silently invalidate it. We record identity on every
|
|
489
|
+
// MiMo request and emit a diagnostic when the OpenCode providerID changes.
|
|
490
|
+
// This never forces or overrides provider routing.
|
|
491
|
+
// NOTE: OpenRouter's *upstream* provider selection (e.g. xiaomi/fp8) is
|
|
492
|
+
// not exposed to plugins; only the OpenCode providerID/modelID are
|
|
493
|
+
// observable here.
|
|
494
|
+
if (family === POLICY_MIMO26 && policyEnabled(cfg, POLICY_MIMO26) && info) {
|
|
495
|
+
// Use the LIVE model identity (not the latched one) so a provider
|
|
496
|
+
// switch within the session is actually observable.
|
|
497
|
+
const live = input.model as unknown as ChatParamsModel
|
|
498
|
+
const cur = {
|
|
499
|
+
providerID: String(live?.providerID ?? ""),
|
|
500
|
+
modelID: String(live?.api?.id ?? live?.id ?? ""),
|
|
501
|
+
}
|
|
502
|
+
if (cur.providerID) {
|
|
503
|
+
const s = get(input.sessionID)
|
|
504
|
+
const ev = providerChangeEvent(s.mimoProvider, cur)
|
|
505
|
+
if (ev.changed) {
|
|
506
|
+
const sticky =
|
|
507
|
+
cfg.policies?.[POLICY_MIMO26]?.stickySession === true
|
|
508
|
+
? { stickySessionId: mimoSessionIdFor(input.sessionID) }
|
|
509
|
+
: {}
|
|
510
|
+
rec.record({
|
|
511
|
+
kind: "boundary",
|
|
512
|
+
sid: input.sessionID,
|
|
513
|
+
ts: Date.now(),
|
|
514
|
+
reason: "mimo_provider_changed",
|
|
515
|
+
policy: POLICY_MIMO26,
|
|
516
|
+
from: ev.from,
|
|
517
|
+
to: ev.to,
|
|
518
|
+
...sticky,
|
|
519
|
+
note: "OpenCode providerID changed; upstream routing is not plugin-visible",
|
|
520
|
+
})
|
|
521
|
+
}
|
|
522
|
+
s.mimoProvider = cur
|
|
523
|
+
}
|
|
524
|
+
return
|
|
525
|
+
}
|
|
526
|
+
|
|
448
527
|
if (!(family === POLICY_GPT56 && policyEnabled(cfg, POLICY_GPT56))) {
|
|
449
528
|
// DeepSeek / GLM / neutral: nothing to inject. GLM has no cache-key API;
|
|
450
529
|
// DeepSeek caching is fully passive; we never mutate requests for them.
|
|
@@ -626,28 +705,47 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
626
705
|
const s = get(sid)
|
|
627
706
|
const family = s.modelInfo?.family
|
|
628
707
|
|
|
629
|
-
// ---- GLM-5.3 input-shape stabilization
|
|
708
|
+
// ---- GLM-5.3 / MiMo-V2.6 input-shape stabilization ------------------
|
|
630
709
|
// Relocate the identifiable volatile env block (per-day date) to the
|
|
631
710
|
// tail of the single system string, content-preserving, ONLY when the
|
|
632
|
-
//
|
|
633
|
-
// touches other content/order; never applied to
|
|
711
|
+
// family opts into system stabilization and the block markers are
|
|
712
|
+
// present exactly. Never touches other content/order; never applied to
|
|
713
|
+
// other families. The runtime passes a single-element system array
|
|
714
|
+
// (verified against the installed runtime), so no generic reordering is
|
|
715
|
+
// involved.
|
|
634
716
|
//
|
|
635
717
|
// In-place mutation note: request.ts keeps using its own local `system`
|
|
636
718
|
// array after the hook (the trigger's returned output is ignored), so
|
|
637
719
|
// reassigning `output.system = [...]` would be lost. We rewrite the
|
|
638
720
|
// single element in place instead.
|
|
639
721
|
let systemText = output.system.join("\n")
|
|
640
|
-
|
|
722
|
+
const glmStabilize =
|
|
641
723
|
family === POLICY_GLM53 &&
|
|
642
724
|
policyEnabled(cfg, POLICY_GLM53) &&
|
|
643
|
-
cfg.policies?.[POLICY_GLM53]?.stabilizeSystem === true
|
|
644
|
-
|
|
645
|
-
|
|
725
|
+
cfg.policies?.[POLICY_GLM53]?.stabilizeSystem === true
|
|
726
|
+
const mimoStabilize =
|
|
727
|
+
family === POLICY_MIMO26 &&
|
|
728
|
+
policyEnabled(cfg, POLICY_MIMO26) &&
|
|
729
|
+
cfg.policies?.[POLICY_MIMO26]?.stabilizeSystem === true
|
|
730
|
+
if ((glmStabilize || mimoStabilize) && output.system.length === 1) {
|
|
646
731
|
const rel = relocateVolatileEnvBlock(output.system[0])
|
|
647
732
|
if (rel.changed) {
|
|
648
733
|
output.system[0] = rel.text
|
|
649
734
|
systemText = rel.text
|
|
650
|
-
|
|
735
|
+
if (family === POLICY_MIMO26) {
|
|
736
|
+
log("debug", "mimo system env block relocated to suffix", { sid })
|
|
737
|
+
rec.record({
|
|
738
|
+
kind: "boundary",
|
|
739
|
+
sid,
|
|
740
|
+
ts: Date.now(),
|
|
741
|
+
reason: "mimo_system_env_relocated",
|
|
742
|
+
policy: POLICY_MIMO26,
|
|
743
|
+
provider: s.modelInfo?.providerID,
|
|
744
|
+
model: s.modelInfo?.modelID,
|
|
745
|
+
})
|
|
746
|
+
} else {
|
|
747
|
+
log("debug", "glm system env block relocated to suffix", { sid })
|
|
748
|
+
}
|
|
651
749
|
}
|
|
652
750
|
}
|
|
653
751
|
|
|
@@ -724,6 +822,22 @@ export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
|
724
822
|
? { toolCount, prevToolCount: prevCount, semanticToolsChanged: semanticChanged, wireToolsChanged: wireChanged }
|
|
725
823
|
: {}),
|
|
726
824
|
})
|
|
825
|
+
// MiMo-specific explicit diagnostic: the STABLE prefix changed (not
|
|
826
|
+
// just the relocated volatile env suffix). Reported only; the new
|
|
827
|
+
// content is never overwritten with a stale snapshot.
|
|
828
|
+
if (family === POLICY_MIMO26 && reasons.includes("system_stable_prefix_changed")) {
|
|
829
|
+
rec.record({
|
|
830
|
+
kind: "boundary",
|
|
831
|
+
sid,
|
|
832
|
+
ts: Date.now(),
|
|
833
|
+
reason: "mimo_system_prefix_changed",
|
|
834
|
+
policy: POLICY_MIMO26,
|
|
835
|
+
provider: s.modelInfo?.providerID,
|
|
836
|
+
model: s.modelInfo?.modelID,
|
|
837
|
+
changedFields: granular,
|
|
838
|
+
reasons,
|
|
839
|
+
})
|
|
840
|
+
}
|
|
727
841
|
if (cfg.logPrefixChanges) {
|
|
728
842
|
log("warn", "observed prefix shape change", {
|
|
729
843
|
sid,
|
package/src/tui.mjs
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
const plugin = {
|
|
2
|
+
id: "opencode-cache-engine",
|
|
3
|
+
|
|
4
|
+
async tui() {
|
|
5
|
+
// CacheEngine has no TUI UI of its own.
|
|
6
|
+
// This target exists so the npm package can be registered,
|
|
7
|
+
// displayed, and enabled/disabled by the OpenCode plugin manager.
|
|
8
|
+
},
|
|
9
|
+
}
|
|
10
|
+
|
|
11
|
+
export default plugin
|
|
@@ -16,6 +16,7 @@ import {
|
|
|
16
16
|
POLICY_DEEPSEEK,
|
|
17
17
|
POLICY_GLM53,
|
|
18
18
|
POLICY_GPT56,
|
|
19
|
+
POLICY_MIMO26,
|
|
19
20
|
POLICY_NEUTRAL,
|
|
20
21
|
canonicalStringify,
|
|
21
22
|
commonPrefixLength,
|
|
@@ -28,8 +29,11 @@ import {
|
|
|
28
29
|
gptCacheOptionsDelta,
|
|
29
30
|
hitRatePct,
|
|
30
31
|
loadConfig,
|
|
32
|
+
mimoHitRate,
|
|
33
|
+
mimoSessionIdFor,
|
|
31
34
|
nextProcessedCursor,
|
|
32
35
|
parseConfig,
|
|
36
|
+
providerChangeEvent,
|
|
33
37
|
relocateVolatileEnvBlock,
|
|
34
38
|
scanPage,
|
|
35
39
|
shapeDiff,
|
|
@@ -221,6 +225,30 @@ test("recorder writes valid JSONL to a real file", () => {
|
|
|
221
225
|
assert.deepEqual(JSON.parse(lines[0]), { kind: "prefix-change", sid: "s", dimensions: ["system"] })
|
|
222
226
|
})
|
|
223
227
|
|
|
228
|
+
// --- 14. TUI target test -----------------------------------------------------
|
|
229
|
+
|
|
230
|
+
test("TUI target exports the expected plugin module", async () => {
|
|
231
|
+
const mod = await import("../src/tui.mjs")
|
|
232
|
+
assert.equal(mod.default.id, "opencode-cache-engine")
|
|
233
|
+
assert.equal(typeof mod.default.tui, "function")
|
|
234
|
+
assert.equal("server" in mod.default, false)
|
|
235
|
+
})
|
|
236
|
+
|
|
237
|
+
// --- 15. Package manifest test with TUI --------------------------------------
|
|
238
|
+
|
|
239
|
+
test("package exposes separate server and TUI targets", async () => {
|
|
240
|
+
const { readFileSync } = await import("node:fs")
|
|
241
|
+
const { join } = await import("node:path")
|
|
242
|
+
|
|
243
|
+
const pkg = JSON.parse(
|
|
244
|
+
readFileSync(join(process.cwd(), "package.json"), "utf8")
|
|
245
|
+
)
|
|
246
|
+
|
|
247
|
+
assert.equal(pkg.exports["./server"], "./src/cache-engine.ts")
|
|
248
|
+
assert.equal(pkg.exports["./tui"], "./src/tui.mjs")
|
|
249
|
+
})
|
|
250
|
+
|
|
251
|
+
|
|
224
252
|
// ===========================================================================
|
|
225
253
|
// Provider-aware model detection
|
|
226
254
|
// ===========================================================================
|
|
@@ -269,6 +297,32 @@ test("unrelated GLM models do NOT match GLM-5.3 policy", () => {
|
|
|
269
297
|
assert.equal(detectPolicy(M("zai", "glm-4.5")), POLICY_NEUTRAL)
|
|
270
298
|
})
|
|
271
299
|
|
|
300
|
+
test("MiMo V2.6 Flash/Pro match MiMo policy (openrouter + direct)", () => {
|
|
301
|
+
assert.equal(detectPolicy(M("openrouter", "xiaomi/mimo-v2.6-flash")), POLICY_MIMO26)
|
|
302
|
+
assert.equal(detectPolicy(M("openrouter", "xiaomi/mimo-v2.6-pro")), POLICY_MIMO26)
|
|
303
|
+
assert.equal(detectPolicy(M("xiaomi", "mimo-v2.6-flash")), POLICY_MIMO26)
|
|
304
|
+
assert.equal(detectPolicy(M("xiaomi", "mimo-v2.6-pro")), POLICY_MIMO26)
|
|
305
|
+
// full Model shape via api.id
|
|
306
|
+
assert.equal(
|
|
307
|
+
detectPolicy({ providerID: "openrouter", api: { id: "xiaomi/mimo-v2.6-flash", npm: "@openrouter/ai-sdk-provider" } }),
|
|
308
|
+
POLICY_MIMO26,
|
|
309
|
+
)
|
|
310
|
+
})
|
|
311
|
+
|
|
312
|
+
test("MiMo V2.5 and Pro-UltraSpeed do NOT match MiMo policy", () => {
|
|
313
|
+
assert.equal(detectPolicy(M("openrouter", "xiaomi/mimo-v2.5")), POLICY_NEUTRAL)
|
|
314
|
+
assert.equal(detectPolicy(M("openrouter", "xiaomi/mimo-v2.5-pro")), POLICY_NEUTRAL)
|
|
315
|
+
assert.equal(detectPolicy(M("xiaomi", "mimo-v2.5")), POLICY_NEUTRAL)
|
|
316
|
+
assert.equal(detectPolicy(M("xiaomi", "mimo-v2.6-pro-ultraspeed")), POLICY_NEUTRAL)
|
|
317
|
+
assert.equal(detectPolicy(M("xiaomi", "mimo-v2.6-flashx")), POLICY_NEUTRAL)
|
|
318
|
+
})
|
|
319
|
+
|
|
320
|
+
test("unrelated MiMo/other models do NOT match MiMo policy", () => {
|
|
321
|
+
assert.equal(detectPolicy(M("openrouter", "xiaomi/mimo-v2")), POLICY_NEUTRAL)
|
|
322
|
+
assert.equal(detectPolicy(M("openrouter", "xiaomi/mimo-v2-flash")), POLICY_NEUTRAL)
|
|
323
|
+
assert.equal(detectPolicy(M("anthropic", "claude-sonnet-4-5")), POLICY_NEUTRAL)
|
|
324
|
+
})
|
|
325
|
+
|
|
272
326
|
test("unrelated models match neutral policy", () => {
|
|
273
327
|
assert.equal(detectPolicy(M("anthropic", "claude-sonnet-4-5")), POLICY_NEUTRAL)
|
|
274
328
|
assert.equal(detectPolicy(M("openrouter", "x-ai/grok-4")), POLICY_NEUTRAL)
|
|
@@ -535,10 +589,16 @@ test("empty current sequence -> no anomalies", () => {
|
|
|
535
589
|
// Config: provider policies
|
|
536
590
|
// ===========================================================================
|
|
537
591
|
|
|
538
|
-
test("config defaults enable all
|
|
592
|
+
test("config defaults enable all four policies", () => {
|
|
539
593
|
const cfg = parseConfig({}, {})
|
|
540
594
|
assert.deepEqual(cfg.policies.deepseek, { enabled: true })
|
|
541
595
|
assert.deepEqual(cfg.policies.glm53, { enabled: true, stabilizeSystem: true, preserveThinkingIntegrity: true })
|
|
596
|
+
assert.deepEqual(cfg.policies.mimo26, {
|
|
597
|
+
enabled: true,
|
|
598
|
+
stabilizeSystem: true,
|
|
599
|
+
stickySession: true,
|
|
600
|
+
preserveThinkingIntegrity: true,
|
|
601
|
+
})
|
|
542
602
|
assert.deepEqual(cfg.policies.gpt56, {
|
|
543
603
|
enabled: true,
|
|
544
604
|
promptCacheKey: true,
|
|
@@ -565,6 +625,7 @@ test("config policy overrides are honored", () => {
|
|
|
565
625
|
},
|
|
566
626
|
glm53: { stabilizeSystem: false },
|
|
567
627
|
deepseek: { enabled: true },
|
|
628
|
+
mimo26: { enabled: false, stabilizeSystem: false, stickySession: false },
|
|
568
629
|
},
|
|
569
630
|
},
|
|
570
631
|
{},
|
|
@@ -578,6 +639,10 @@ test("config policy overrides are honored", () => {
|
|
|
578
639
|
assert.equal(cfg.policies.gpt56.ttl, "1h")
|
|
579
640
|
assert.equal(cfg.policies.glm53.stabilizeSystem, false)
|
|
580
641
|
assert.equal(cfg.policies.glm53.preserveThinkingIntegrity, true)
|
|
642
|
+
assert.equal(cfg.policies.mimo26.enabled, false)
|
|
643
|
+
assert.equal(cfg.policies.mimo26.stabilizeSystem, false)
|
|
644
|
+
assert.equal(cfg.policies.mimo26.stickySession, false)
|
|
645
|
+
assert.equal(cfg.policies.mimo26.preserveThinkingIntegrity, true)
|
|
581
646
|
})
|
|
582
647
|
|
|
583
648
|
test("config invalid policy values fall back to defaults", () => {
|
|
@@ -763,3 +828,143 @@ test("boundary: reasoning-integrity reason tokens", () => {
|
|
|
763
828
|
])
|
|
764
829
|
assert.deepEqual(reasoningIssueReasons(null), [])
|
|
765
830
|
})
|
|
831
|
+
|
|
832
|
+
// ===========================================================================
|
|
833
|
+
// MiMo-V2.6: system env relocation (shared content-preserving helper)
|
|
834
|
+
// ===========================================================================
|
|
835
|
+
|
|
836
|
+
const MIMO_SYSTEM = [
|
|
837
|
+
"You are a senior software engineer.",
|
|
838
|
+
"You are powered by the model named mimo-v2.6-flash. The exact model ID is xiaomi/mimo-v2.6-flash",
|
|
839
|
+
"Here is some useful information about the environment you are running in:",
|
|
840
|
+
"<env>",
|
|
841
|
+
"Working directory: /home/dev/project",
|
|
842
|
+
"Today's date: 2026-08-17",
|
|
843
|
+
"</env>",
|
|
844
|
+
"You MUST follow AGENTS.md instructions and keep your responses concise.",
|
|
845
|
+
].join("\n")
|
|
846
|
+
|
|
847
|
+
test("MiMo env block is relocated to the tail, contents preserved exactly", () => {
|
|
848
|
+
const { text, changed } = relocateVolatileEnvBlock(MIMO_SYSTEM)
|
|
849
|
+
assert.equal(changed, true)
|
|
850
|
+
assert.ok(text.endsWith("</env>"))
|
|
851
|
+
const norm = (t) => t.split("\n").filter((l) => l).sort().join("\n")
|
|
852
|
+
assert.equal(norm(text), norm(MIMO_SYSTEM))
|
|
853
|
+
assert.ok(text.indexOf("You MUST follow") < text.indexOf("You are powered"))
|
|
854
|
+
})
|
|
855
|
+
|
|
856
|
+
test("MiMo relocation is a no-op when the block is already at the tail", () => {
|
|
857
|
+
const already = relocateVolatileEnvBlock(MIMO_SYSTEM).text
|
|
858
|
+
const again = relocateVolatileEnvBlock(already)
|
|
859
|
+
assert.equal(again.changed, false)
|
|
860
|
+
assert.equal(again.text, already)
|
|
861
|
+
})
|
|
862
|
+
|
|
863
|
+
test("MiMo relocation is a no-op when start/end markers are missing", () => {
|
|
864
|
+
const missingStart = "Just a system prompt.\n<env>\nToday's date: x\n</env>"
|
|
865
|
+
assert.equal(relocateVolatileEnvBlock(missingStart).changed, false)
|
|
866
|
+
const missingEnd = "You are powered by the model named mimo-v2.6-flash\nbut never closed"
|
|
867
|
+
assert.equal(relocateVolatileEnvBlock(missingEnd).changed, false)
|
|
868
|
+
})
|
|
869
|
+
|
|
870
|
+
// ===========================================================================
|
|
871
|
+
// MiMo-V2.6: sticky-session identity (pure, derived but not injected)
|
|
872
|
+
// ===========================================================================
|
|
873
|
+
|
|
874
|
+
test("mimoSessionIdFor: deterministic + stable for the same session", () => {
|
|
875
|
+
assert.equal(mimoSessionIdFor("ses_abc123"), mimoSessionIdFor("ses_abc123"))
|
|
876
|
+
})
|
|
877
|
+
|
|
878
|
+
test("mimoSessionIdFor: distinct sessions produce distinct ids", () => {
|
|
879
|
+
assert.notEqual(mimoSessionIdFor("ses_abc"), mimoSessionIdFor("ses_xyz"))
|
|
880
|
+
})
|
|
881
|
+
|
|
882
|
+
test("mimoSessionIdFor: printable, no whitespace, within 256 chars", () => {
|
|
883
|
+
const id = mimoSessionIdFor("ses_" + "a".repeat(500))
|
|
884
|
+
assert.ok(id.length <= 256)
|
|
885
|
+
assert.ok(!/\s/.test(id))
|
|
886
|
+
assert.ok(/^[\x20-\x7E]+$/.test(id))
|
|
887
|
+
})
|
|
888
|
+
|
|
889
|
+
test("mimoSessionIdFor: invalid input -> null (no fabrication)", () => {
|
|
890
|
+
assert.equal(mimoSessionIdFor(undefined), null)
|
|
891
|
+
assert.equal(mimoSessionIdFor(null), null)
|
|
892
|
+
assert.equal(mimoSessionIdFor(""), null)
|
|
893
|
+
assert.equal(mimoSessionIdFor(123), null)
|
|
894
|
+
})
|
|
895
|
+
|
|
896
|
+
test("mimoSessionIdFor: transient request fields cannot alter the id", () => {
|
|
897
|
+
// the helper is a pure function of the session id; extra args are ignored
|
|
898
|
+
assert.equal(mimoSessionIdFor("ses_stable"), mimoSessionIdFor("ses_stable", { turn: 7, temperature: 0.9 }))
|
|
899
|
+
})
|
|
900
|
+
|
|
901
|
+
// ===========================================================================
|
|
902
|
+
// MiMo-V2.6: cached/prompt token metrics
|
|
903
|
+
// ===========================================================================
|
|
904
|
+
|
|
905
|
+
test("mimoHitRate = cachedTokens / promptTokens", () => {
|
|
906
|
+
assert.equal(mimoHitRate(47000, 50000), 94)
|
|
907
|
+
assert.equal(mimoHitRate(0, 100), 0)
|
|
908
|
+
assert.equal(mimoHitRate(100, 100), 100)
|
|
909
|
+
})
|
|
910
|
+
|
|
911
|
+
test("mimoHitRate returns null for zero/unknown denominators (no fabrication)", () => {
|
|
912
|
+
assert.equal(mimoHitRate(0, 0), null)
|
|
913
|
+
assert.equal(mimoHitRate(10, 0), null)
|
|
914
|
+
assert.equal(mimoHitRate(NaN, 100), null)
|
|
915
|
+
assert.equal(mimoHitRate(10, NaN), null)
|
|
916
|
+
assert.equal(mimoHitRate(undefined, 100), null)
|
|
917
|
+
assert.equal(mimoHitRate(10, undefined), null)
|
|
918
|
+
})
|
|
919
|
+
|
|
920
|
+
test("mimoHitRate is NOT the read/(read+write) form", () => {
|
|
921
|
+
// read=80, input=20 -> promptTokens derived 100 -> 80%. The read/(read+write)
|
|
922
|
+
// form would give 100% for (80, 0); they must not be conflated.
|
|
923
|
+
assert.equal(mimoHitRate(80, 80 + 20), 80)
|
|
924
|
+
assert.equal(hitRatePct(80, 0), 100)
|
|
925
|
+
})
|
|
926
|
+
|
|
927
|
+
// ===========================================================================
|
|
928
|
+
// MiMo-V2.6: provider-switch diagnostics
|
|
929
|
+
// ===========================================================================
|
|
930
|
+
|
|
931
|
+
test("providerChangeEvent: no event until two real observations exist", () => {
|
|
932
|
+
assert.deepEqual(providerChangeEvent(null, { providerID: "openrouter", modelID: "xiaomi/mimo-v2.6-flash" }), {
|
|
933
|
+
changed: false,
|
|
934
|
+
from: null,
|
|
935
|
+
to: null,
|
|
936
|
+
})
|
|
937
|
+
assert.equal(providerChangeEvent({ providerID: "openrouter" }, null).changed, false)
|
|
938
|
+
})
|
|
939
|
+
|
|
940
|
+
test("providerChangeEvent: same provider -> no change", () => {
|
|
941
|
+
const a = { providerID: "openrouter", modelID: "xiaomi/mimo-v2.6-flash" }
|
|
942
|
+
const b = { providerID: "openrouter", modelID: "xiaomi/mimo-v2.6-pro" }
|
|
943
|
+
assert.equal(providerChangeEvent(a, b).changed, false)
|
|
944
|
+
})
|
|
945
|
+
|
|
946
|
+
test("providerChangeEvent: different provider -> change with from/to", () => {
|
|
947
|
+
const a = { providerID: "openrouter", modelID: "xiaomi/mimo-v2.6-flash" }
|
|
948
|
+
const b = { providerID: "xiaomi", modelID: "mimo-v2.6-flash" }
|
|
949
|
+
const ev = providerChangeEvent(a, b)
|
|
950
|
+
assert.equal(ev.changed, true)
|
|
951
|
+
assert.deepEqual(ev.from, a)
|
|
952
|
+
assert.deepEqual(ev.to, b)
|
|
953
|
+
})
|
|
954
|
+
|
|
955
|
+
// ===========================================================================
|
|
956
|
+
// MiMo-V2.6: no tool-definition mutation
|
|
957
|
+
// ===========================================================================
|
|
958
|
+
|
|
959
|
+
test("MiMo policy performs no tool mutation (fingerprints are pure inputs)", () => {
|
|
960
|
+
// The plugin never adds a MiMo tool-ordering pass: this runtime already sorts
|
|
961
|
+
// tools alphabetically before the wire. Classifying a model as MiMo must not
|
|
962
|
+
// affect tool fingerprints, which are a pure function of the tool definitions.
|
|
963
|
+
const model = { providerID: "openrouter", modelID: "xiaomi/mimo-v2.6-flash" }
|
|
964
|
+
assert.equal(detectPolicy(model), POLICY_MIMO26)
|
|
965
|
+
const before = { sem: toolFingerprint(TOOLS), wire: toolWireFingerprint(TOOLS) }
|
|
966
|
+
const after = { sem: toolFingerprint(TOOLS), wire: toolWireFingerprint(TOOLS) }
|
|
967
|
+
assert.deepEqual(after, before)
|
|
968
|
+
// semantic fingerprint stays order-insensitive regardless of the policy
|
|
969
|
+
assert.equal(toolFingerprint([...TOOLS].reverse()), before.sem)
|
|
970
|
+
})
|