opencode-cache-engine 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md ADDED
@@ -0,0 +1,1032 @@
1
+ # OpenCode Cache Engine
2
+
3
+ Provider-aware prompt-cache optimization and observability for [OpenCode](https://opencode.ai).
4
+
5
+ `CacheEngine` is an OpenCode plugin designed for long-running agent sessions where prompt-cache efficiency affects both latency and cost. It keeps the harness conservative for providers whose cache behavior is already automatic, while applying provider-specific optimizations where the provider exposes useful cache controls or where prompt structure can be safely improved.
6
+
7
+ The plugin currently has three cache-policy families:
8
+
9
+ * **DeepSeek V4 Flash** — passive cache-stability and observability
10
+ * **GPT-5.6 Luna** — active cache-control configuration
11
+ * **GLM-5.3 Flash** — conservative system-prompt stabilization
12
+
13
+ The central design principle is:
14
+
15
+ > Optimize the request structure only when there is a clear provider-specific reason to do so. Otherwise, preserve OpenCode's native request behavior and measure what the provider actually reports.
16
+
17
+ ---
18
+
19
+ ## What this plugin does
20
+
21
+ The plugin operates at the OpenCode harness level rather than implementing a provider-specific client.
22
+
23
+ It:
24
+
25
+ 1. Detects the model/provider family in use.
26
+ 2. Applies only the policy appropriate for that family.
27
+ 3. Observes system-prompt and tool-definition stability.
28
+ 4. Records provider-reported cache token usage.
29
+ 5. Adds a deterministic compaction continuation block.
30
+ 6. Applies GPT-5.6 cache-control metadata.
31
+ 7. Applies the GLM-5.3 volatile-environment relocation.
32
+ 8. Records diagnostics that help determine whether prompt-shape changes correlate with cache behavior.
33
+
34
+ The plugin deliberately avoids pretending that a local hash is proof of a provider cache hit. Provider-reported token usage remains the authoritative signal.
35
+
36
+ ---
37
+
38
+ # Provider behavior
39
+
40
+ ## DeepSeek V4 Flash
41
+
42
+ ### Policy: passive
43
+
44
+ DeepSeek receives **no cache-specific request mutation**.
45
+
46
+ The plugin does not:
47
+
48
+ * rewrite the system prompt
49
+ * reorder tools
50
+ * modify messages
51
+ * inject cache-control fields
52
+ * inject a prompt-cache key
53
+ * alter provider request options
54
+
55
+ The DeepSeek branch exists primarily to preserve a stable harness while providing observability around the prefix structure and cache usage.
56
+
57
+ This is intentional. The implementation describes DeepSeek as a passive policy whose purpose is to preserve the existing high-cache-rate behavior rather than introduce new request mutations.
58
+
59
+ The plugin still observes:
60
+
61
+ * system-prompt shape
62
+ * semantic tool definitions
63
+ * wire-order tool definitions
64
+ * prefix changes
65
+ * cache read tokens
66
+ * cache write tokens
67
+ * compaction boundaries
68
+
69
+ ### Why passive?
70
+
71
+ DeepSeek's cache behavior is provider-managed. Introducing unnecessary prompt mutations would risk changing the prefix that the provider can reuse.
72
+
73
+ Therefore the plugin follows a simple rule:
74
+
75
+ ```text
76
+ DeepSeek:
77
+ preserve request
78
+ preserve prefix
79
+ measure cache
80
+ ```
81
+
82
+ rather than:
83
+
84
+ ```text
85
+ DeepSeek:
86
+ rewrite request
87
+ guess cache key
88
+ force cache behavior
89
+ ```
90
+
91
+ ---
92
+
93
+ ## GPT-5.6 Luna
94
+
95
+ ### Policy: active cache control
96
+
97
+ GPT-5.6 is the only current policy that actively injects cache-control request metadata.
98
+
99
+ The plugin adds:
100
+
101
+ ```json
102
+ {
103
+ "promptCacheKey": "<stable-session-key>",
104
+ "promptCacheOptions": {
105
+ "mode": "implicit",
106
+ "ttl": "30m"
107
+ }
108
+ }
109
+ ```
110
+
111
+ The key is derived from the OpenCode session identity and is independent of transient request data. The implementation also preserves existing provider-supplied cache settings rather than overwriting them.
112
+
113
+ ### Important: the prompt text is not rewritten
114
+
115
+ For GPT-5.6:
116
+
117
+ ```text
118
+ system prompt -> unchanged
119
+ conversation -> unchanged
120
+ tool definitions -> unchanged
121
+
122
+ request metadata -> cache key/options added
123
+ ```
124
+
125
+ This means the plugin is controlling the cache namespace and cache behavior without performing prompt surgery.
126
+
127
+ ### Default GPT configuration
128
+
129
+ ```json
130
+ {
131
+ "promptCacheKey": true,
132
+ "cacheRootKey": false,
133
+ "compactionCacheIsolation": true,
134
+ "reasoningEffortDiagnostics": true,
135
+ "mode": "implicit",
136
+ "ttl": "30m"
137
+ }
138
+ ```
139
+
140
+ The current implementation intentionally leaves `cacheRootKey` disabled because the OpenCode runtime does not currently expose sufficiently reliable fork lineage for safe parent-cache inheritance. The code path remains available for a future runtime that exposes reliable parent relationships.
141
+
142
+ ### Compaction isolation
143
+
144
+ Compaction uses a deterministic separate cache-key namespace:
145
+
146
+ ```text
147
+ live session:
148
+ ses_abc123
149
+
150
+ compaction:
151
+ ses_abc123:compact
152
+ ```
153
+
154
+ This prevents a compaction-specific prompt from sharing the same GPT cache namespace as the normal live-session prompt. The behavior is deterministic and tested explicitly.
155
+
156
+ ---
157
+
158
+ ## GLM-5.3 Flash
159
+
160
+ ### Policy: input-shape optimization
161
+
162
+ GLM-5.3 receives the only prompt-text transformation in the current plugin.
163
+
164
+ The plugin identifies OpenCode's volatile `<env>` section and moves it to the **tail of the system prompt**.
165
+
166
+ Conceptually:
167
+
168
+ ```text
169
+ BEFORE
170
+
171
+ [large stable instructions]
172
+ [volatile environment/date block]
173
+ [more stable instructions]
174
+ ```
175
+
176
+ becomes:
177
+
178
+ ```text
179
+ AFTER
180
+
181
+ [large stable instructions]
182
+ [more stable instructions]
183
+ [volatile environment/date block]
184
+ ```
185
+
186
+ The contents of the environment block are preserved exactly. The operation changes its location, not its contents.
187
+
188
+ ### Why?
189
+
190
+ The environment block can contain volatile information such as a changing date.
191
+
192
+ Keeping that material at the end allows the earlier portion of the system prompt to remain stable across requests.
193
+
194
+ The plugin therefore attempts to isolate volatility:
195
+
196
+ ```text
197
+ stable prefix
198
+ ---------------------------
199
+ unchanged across requests
200
+
201
+ volatile suffix
202
+ ---------------------------
203
+ allowed to change
204
+ ```
205
+
206
+ The system-shape diagnostics explicitly distinguish the stable prefix from the volatile suffix for this purpose.
207
+
208
+ ### GLM safety constraints
209
+
210
+ The transformation is deliberately narrow.
211
+
212
+ It only occurs when:
213
+
214
+ * the selected model is GLM-5.3
215
+ * GLM stabilization is enabled
216
+ * there is exactly one system string
217
+ * the expected environment markers exist
218
+ * the block can be identified unambiguously
219
+
220
+ The plugin does not arbitrarily rearrange unrelated prompt content.
221
+
222
+ ---
223
+
224
+ # Prompt-cache strategy
225
+
226
+ The plugin uses three different strategies because cache mechanisms differ by provider.
227
+
228
+ | Provider | Prompt text changed? | Cache metadata changed? | Main strategy |
229
+ | ----------------- | -------------------: | ----------------------: | --------------------------------- |
230
+ | DeepSeek V4 Flash | No | No | Preserve stable harness + observe |
231
+ | GPT-5.6 Luna | No | Yes | Stable cache key + cache options |
232
+ | GLM-5.3 Flash | Yes, narrowly | No provider cache key | Isolate volatile system content |
233
+
234
+ This distinction is fundamental.
235
+
236
+ The plugin is **not** a generic "rewrite every prompt for caching" engine.
237
+
238
+ It is a provider-aware cache policy engine.
239
+
240
+ ---
241
+
242
+ # System-prompt diagnostics
243
+
244
+ The plugin fingerprints the system prompt to detect structural changes between requests.
245
+
246
+ For newer provider-aware diagnostics it tracks:
247
+
248
+ * full system hash
249
+ * stable system-prefix hash
250
+ * volatile system-suffix hash
251
+
252
+ The stable/volatile decomposition is based on the longest common prefix against the session baseline.
253
+
254
+ A change in a hash means:
255
+
256
+ > The observed request bytes changed.
257
+
258
+ It does **not** mean:
259
+
260
+ > The provider definitely generated a cache miss.
261
+
262
+ This distinction is intentional. Provider-reported cache token counts are the authoritative cache signal.
263
+
264
+ ---
265
+
266
+ # Tool-definition diagnostics
267
+
268
+ Tool definitions are normalized before fingerprinting.
269
+
270
+ Runtime-only fields such as:
271
+
272
+ * object identity
273
+ * function references
274
+ * timestamps
275
+ * arbitrary runtime metadata
276
+
277
+ are excluded.
278
+
279
+ The semantic fingerprint is order-insensitive and represents the model-visible tool definitions.
280
+
281
+ The plugin also tracks wire-order fingerprints so that it can distinguish:
282
+
283
+ ```text
284
+ same tools, different ordering
285
+ ```
286
+
287
+ from:
288
+
289
+ ```text
290
+ different tool definitions
291
+ ```
292
+
293
+ This distinction matters because semantic equality and byte-level request equality are not necessarily the same thing.
294
+
295
+ The plugin uses these fingerprints for **diagnostics only**. It does not reorder the tools to force a particular fingerprint.
296
+
297
+ ---
298
+
299
+ # Compaction handling
300
+
301
+ OpenCode sessions eventually undergo compaction as their conversation history grows.
302
+
303
+ The plugin adds a deterministic continuation template:
304
+
305
+ ```text
306
+ ## Session digest (cache-stable continuation block)
307
+ - Goal:
308
+ - Decisions made:
309
+ - Pending:
310
+ - Active files:
311
+ ```
312
+
313
+ The digest is inserted once per compaction operation using a guard that prevents duplicate insertion if the compaction hook fires multiple times.
314
+ The objective is to provide a deterministic continuation structure rather than generating a different arbitrary cache-affecting block on every compaction.
315
+
316
+ ---
317
+
318
+ # Cache metrics
319
+
320
+ The plugin records cache usage from OpenCode assistant-message token data.
321
+
322
+ At minimum it tracks:
323
+
324
+ ```text
325
+ cache.read
326
+ cache.write
327
+ ```
328
+
329
+ and aggregates those values across the session.
330
+
331
+ The default cache ratio reported by the core helper is:
332
+
333
+ ```text
334
+ hit rate = read / (read + write)
335
+ ```
336
+
337
+ This is deliberately an accounting metric based on cache read/write tokens.
338
+
339
+ For GLM, the implementation additionally calculates a prompt-token ratio:
340
+
341
+ ```text
342
+ cached / (cached + cache-write + input)
343
+ ```
344
+
345
+ using:
346
+
347
+ ```text
348
+ read / (read + write + input)
349
+ ```
350
+
351
+ as implemented by `glmHitRatio()`.
352
+
353
+ ### Important metric distinction
354
+
355
+ These ratios answer different questions.
356
+
357
+ `read / (read + write)` answers approximately:
358
+
359
+ > Of the tokens represented as cache reads/writes, how much was reused?
360
+
361
+ `read / (read + write + input)` answers:
362
+
363
+ > How much of the total prompt-token accounting was represented by cached reads?
364
+
365
+ Do not treat the two percentages as interchangeable.
366
+
367
+ ---
368
+
369
+ # Telemetry
370
+
371
+ Metrics are written as JSONL.
372
+
373
+ The default location is:
374
+
375
+ ```text
376
+ ~/.cache/opencode/cache-metrics.jsonl
377
+ ```
378
+
379
+ The default configuration path is:
380
+
381
+ ```text
382
+ ~/.config/opencode/cache-engine.json
383
+ ```
384
+
385
+ These paths are defined by the plugin core.
386
+
387
+ Telemetry is best-effort.
388
+
389
+ A failed metrics write must never break an OpenCode request. The recorder catches write failures rather than allowing telemetry failures to affect execution.
390
+
391
+ ---
392
+
393
+ # Metrics examples
394
+
395
+ A usage record can contain fields such as:
396
+
397
+ ```json
398
+ {
399
+ "kind": "usage-event",
400
+ "sid": "session-id",
401
+ "ts": 1750000000000,
402
+ "read": 120000,
403
+ "write": 3000,
404
+ "cost": 0.0123,
405
+ "provider": "z-ai",
406
+ "model": "glm-5.3-flash",
407
+ "policy": "glm53"
408
+ }
409
+ ```
410
+
411
+ A prefix-change record can look like:
412
+
413
+ ```json
414
+ {
415
+ "kind": "prefix-change",
416
+ "sid": "session-id",
417
+ "ts": 1750000000000,
418
+ "dimensions": [
419
+ "system"
420
+ ]
421
+ }
422
+ ```
423
+
424
+ A compaction record can contain:
425
+
426
+ ```json
427
+ {
428
+ "kind": "compaction",
429
+ "sid": "session-id",
430
+ "ts": 1750000000000,
431
+ "reason": "compaction",
432
+ "usageSamples": 7,
433
+ "cumulative": {
434
+ "read": 900000,
435
+ "write": 12000
436
+ }
437
+ }
438
+ ```
439
+
440
+ Telemetry is intended to answer questions such as:
441
+
442
+ * Did the system prompt change?
443
+ * Did the tool definitions change?
444
+ * Did cache reads increase?
445
+ * Did cache writes increase?
446
+ * Did a compaction occur?
447
+ * Which provider/model/policy was active?
448
+ * Did the GLM system stabilization actually change the observed prompt shape?
449
+
450
+ ---
451
+
452
+ # Configuration
453
+
454
+ The default configuration is:
455
+
456
+ ```json
457
+ {
458
+ "enabled": true,
459
+ "metricsFile": "~/.cache/opencode/cache-metrics.jsonl",
460
+ "compactTemplate": true,
461
+ "logPrefixChanges": true,
462
+ "policies": {
463
+ "deepseek": {
464
+ "enabled": true
465
+ },
466
+ "gpt56": {
467
+ "enabled": true,
468
+ "promptCacheKey": true,
469
+ "cacheRootKey": false,
470
+ "compactionCacheIsolation": true,
471
+ "reasoningEffortDiagnostics": true,
472
+ "mode": "implicit",
473
+ "ttl": "30m"
474
+ },
475
+ "glm53": {
476
+ "enabled": true,
477
+ "stabilizeSystem": true,
478
+ "preserveThinkingIntegrity": true
479
+ }
480
+ }
481
+ }
482
+ ```
483
+
484
+ The configuration parser starts from these defaults and applies valid file/environment overrides without mutating the caller's configuration object.
485
+
486
+ ---
487
+
488
+ # Configuration options
489
+
490
+ ## Global
491
+
492
+ ### `enabled`
493
+
494
+ ```json
495
+ {
496
+ "enabled": true
497
+ }
498
+ ```
499
+
500
+ Enables or disables the entire plugin.
501
+
502
+ ---
503
+
504
+ ### `metricsFile`
505
+
506
+ ```json
507
+ {
508
+ "metricsFile": "~/.cache/opencode/cache-metrics.jsonl"
509
+ }
510
+ ```
511
+
512
+ Controls where JSONL telemetry is written.
513
+
514
+ ---
515
+
516
+ ### `compactTemplate`
517
+
518
+ ```json
519
+ {
520
+ "compactTemplate": true
521
+ }
522
+ ```
523
+
524
+ Controls whether the deterministic compaction continuation block is inserted.
525
+
526
+ ---
527
+
528
+ ### `logPrefixChanges`
529
+
530
+ ```json
531
+ {
532
+ "logPrefixChanges": true
533
+ }
534
+ ```
535
+
536
+ Controls warning logs for observed prefix-shape changes.
537
+
538
+ ---
539
+
540
+ # DeepSeek configuration
541
+
542
+ ```json
543
+ "deepseek": {
544
+ "enabled": true
545
+ }
546
+ ```
547
+
548
+ There are intentionally very few settings here.
549
+
550
+ DeepSeek is treated as the conservative/passive policy.
551
+
552
+ ---
553
+
554
+ # GPT-5.6 configuration
555
+
556
+ ```json
557
+ "gpt56": {
558
+ "enabled": true,
559
+ "promptCacheKey": true,
560
+ "cacheRootKey": false,
561
+ "compactionCacheIsolation": true,
562
+ "reasoningEffortDiagnostics": true,
563
+ "mode": "implicit",
564
+ "ttl": "30m"
565
+ }
566
+ ```
567
+
568
+ ### `promptCacheKey`
569
+
570
+ Controls whether the plugin provides a stable session-derived GPT cache key.
571
+
572
+ ### `cacheRootKey`
573
+
574
+ Controls whether a parent/fork cache root is used.
575
+
576
+ Disabled by default because reliable fork lineage is not currently guaranteed by the runtime.
577
+
578
+ ### `compactionCacheIsolation`
579
+
580
+ Uses a separate deterministic cache namespace for compaction requests.
581
+
582
+ ### `reasoningEffortDiagnostics`
583
+
584
+ Tracks GPT reasoning-effort changes for diagnostics.
585
+
586
+ ### `mode`
587
+
588
+ Defaults to:
589
+
590
+ ```text
591
+ implicit
592
+ ```
593
+
594
+ ### `ttl`
595
+
596
+ Defaults to:
597
+
598
+ ```text
599
+ 30m
600
+ ```
601
+
602
+ Existing request options are not overwritten by the plugin.
603
+
604
+ ---
605
+
606
+ # GLM-5.3 configuration
607
+
608
+ ```json
609
+ "glm53": {
610
+ "enabled": true,
611
+ "stabilizeSystem": true,
612
+ "preserveThinkingIntegrity": true
613
+ }
614
+ ```
615
+
616
+ ### `stabilizeSystem`
617
+
618
+ Enables relocation of the volatile `<env>` section to the system-prompt tail.
619
+
620
+ ### `preserveThinkingIntegrity`
621
+
622
+ Enables diagnostic checks around reasoning continuity.
623
+
624
+ The reasoning instrumentation is intended to identify anomalies such as:
625
+
626
+ * duplicate reasoning
627
+ * reordered reasoning
628
+ * modified reasoning
629
+
630
+ It is diagnostic rather than a reason to rewrite or fabricate reasoning content. The implementation maps these conditions to explicit diagnostic reasons.
631
+
632
+ ---
633
+
634
+ # Model detection
635
+
636
+ The plugin classifies requests into:
637
+
638
+ ```text
639
+ deepseek
640
+ gpt56
641
+ glm53
642
+ neutral
643
+ ```
644
+
645
+ The model detector recognizes:
646
+
647
+ * DeepSeek model/provider identifiers
648
+ * GPT-5.6 variants
649
+ * GLM-5.3 variants
650
+
651
+ GPT-5.6 has an additional OpenAI/Azure-context check so a string containing `gpt-5.6` does not automatically cause GPT-specific fields to be sent to an unrelated endpoint.
652
+
653
+ Unknown models use the neutral policy.
654
+
655
+ Neutral means:
656
+
657
+ ```text
658
+ no provider-specific request mutation
659
+ ```
660
+
661
+ ---
662
+
663
+ # OpenRouter usage
664
+
665
+ This plugin is compatible with OpenRouter because the cache policy is based on the model/provider signals available to OpenCode.
666
+
667
+ For cache-sensitive workloads, provider stability remains important.
668
+
669
+ The plugin does not attempt to compensate for provider switching by rewriting prompts.
670
+
671
+ For that reason, a stable provider route is preferable when your goal is to measure and maximize prefix reuse.
672
+
673
+ ---
674
+
675
+ # Architecture
676
+
677
+ The implementation is split into two layers.
678
+
679
+ ## `cache-engine.ts`
680
+
681
+ This is the OpenCode plugin entry point.
682
+
683
+ It owns:
684
+
685
+ * OpenCode hooks
686
+ * session state
687
+ * provider-policy selection
688
+ * telemetry integration
689
+ * request mutation
690
+ * system-prompt transformation
691
+ * compaction handling
692
+
693
+ The exported plugin is:
694
+
695
+ ```ts
696
+ export const CacheEngine: Plugin = async ({ client, directory }) => {
697
+ // ...
698
+ }
699
+ ```
700
+
701
+ The identifier `CacheEngine` is the OpenCode plugin export name. It does not determine the eventual npm package name.
702
+
703
+ ---
704
+
705
+ ## `cache-engine-core.mjs`
706
+
707
+ This contains dependency-light pure logic.
708
+
709
+ It owns:
710
+
711
+ * provider classification
712
+ * configuration parsing
713
+ * hashing
714
+ * canonicalization
715
+ * tool fingerprints
716
+ * system-shape decomposition
717
+ * GPT cache-key generation
718
+ * cache-option generation
719
+ * GLM environment relocation
720
+ * reasoning diagnostics
721
+ * usage aggregation
722
+ * compaction guards
723
+
724
+ Keeping these functions in plain JavaScript allows the logic to be tested independently with Node's built-in test runner.
725
+
726
+ ---
727
+
728
+ ## Tests
729
+
730
+ The repository's test suite validates the provider-independent and provider-specific logic.
731
+
732
+ Coverage includes:
733
+
734
+ * model detection
735
+ * GPT cache-key stability
736
+ * GPT cache-option defaults
737
+ * protection against overwriting existing cache options
738
+ * GLM environment relocation
739
+ * deterministic hashing
740
+ * system-prefix decomposition
741
+ * tool fingerprints
742
+ * reasoning diagnostics
743
+ * compaction isolation
744
+ * cache-hit calculations
745
+ * configuration behavior
746
+ * JSONL telemetry behavior
747
+
748
+ The tests are designed around the pure core logic, while OpenCode runtime behavior is validated separately through actual plugin loading.
749
+
750
+ ---
751
+
752
+ # Design principles
753
+
754
+ ## 1. Provider-specific behavior
755
+
756
+ Different providers expose different cache mechanisms.
757
+
758
+ The plugin therefore does not assume that one strategy is optimal everywhere.
759
+
760
+ ---
761
+
762
+ ## 2. Preserve working behavior
763
+
764
+ The plugin should not modify a provider's request merely because a mutation is technically possible.
765
+
766
+ This is especially important for DeepSeek, where the current policy is intentionally passive.
767
+
768
+ ---
769
+
770
+ ## 3. Measure provider reality
771
+
772
+ Local hashes are diagnostics.
773
+
774
+ Provider-reported cache token counts are the authoritative signal.
775
+
776
+ The implementation explicitly distinguishes:
777
+
778
+ ```text
779
+ observed prefix change
780
+ ```
781
+
782
+ from:
783
+
784
+ ```text
785
+ confirmed provider cache miss
786
+ ```
787
+
788
+ because the plugin cannot infer the latter reliably from local prompt hashes alone.
789
+
790
+ ---
791
+
792
+ ## 4. Never overwrite explicit provider configuration
793
+
794
+ Where GPT cache options already exist, the plugin leaves them alone.
795
+
796
+ This allows the runtime or user configuration to remain authoritative.
797
+
798
+ ---
799
+
800
+ ## 5. Keep mutations deterministic
801
+
802
+ When the plugin does transform the request, the transformation should be:
803
+
804
+ * narrow
805
+ * deterministic
806
+ * content-preserving where possible
807
+ * provider-specific
808
+ * easy to disable
809
+
810
+ The GLM environment relocation follows these rules.
811
+
812
+ ---
813
+
814
+ ## 6. Keep telemetry out of the critical path
815
+
816
+ A metrics failure must not break model execution.
817
+
818
+ Telemetry is therefore best-effort.
819
+
820
+ ---
821
+
822
+ # What the plugin does NOT do
823
+
824
+ The plugin does not:
825
+
826
+ * invent cache hits
827
+ * claim a local hash proves a provider cache hit
828
+ * rewrite DeepSeek prompts
829
+ * reorder tools
830
+ * fabricate reasoning
831
+ * modify conversation history arbitrarily
832
+ * force explicit GPT cache breakpoints by default
833
+ * silently overwrite existing GPT cache options
834
+ * assume every model named `gpt-5.6` is an OpenAI-compatible endpoint
835
+ * use fork inheritance unless reliable lineage is available
836
+
837
+ ---
838
+
839
+ # Cost optimization philosophy
840
+
841
+ Cache hit rate is useful, but it is not the only cost metric.
842
+
843
+ The economic objective is:
844
+
845
+ ```text
846
+ total task cost
847
+ =
848
+ prompt/cache cost
849
+ +
850
+ output/reasoning cost
851
+ +
852
+ additional requests
853
+ ```
854
+
855
+ A model with a slightly lower cache hit rate can still be cheaper if it completes the task with fewer tokens or fewer model calls.
856
+
857
+ For that reason, this plugin is primarily an **instrumentation + targeted optimization layer**, not a cache-rate maximizer at any cost.
858
+
859
+ The recommended evaluation unit is:
860
+
861
+ ```text
862
+ cost per completed task
863
+ ```
864
+
865
+ rather than:
866
+
867
+ ```text
868
+ cache percentage alone
869
+ ```
870
+
871
+ ---
872
+
873
+ # Operational recommendations
874
+
875
+ For reliable cache measurements:
876
+
877
+ 1. Keep the provider fixed whenever possible.
878
+ 2. Avoid changing unrelated system-prompt content during a benchmark.
879
+ 3. Keep tool definitions stable.
880
+ 4. Compare equivalent tasks across models.
881
+ 5. Record actual provider cache token counts.
882
+ 6. Compare total task cost, not only cache percentage.
883
+ 7. Treat compaction as a separate cache boundary when analyzing results.
884
+ 8. Avoid interpreting a local prefix hash change as definitive proof of a cache miss.
885
+
886
+ ---
887
+
888
+ # File layout
889
+
890
+ A typical standalone repository can use:
891
+
892
+ ```text
893
+ opencode-cache-engine/
894
+ ├── src/
895
+ │ ├── cache-engine.ts
896
+ │ └── cache-engine-core.mjs
897
+ ├── test/
898
+ │ └── cache-engine.test.mjs
899
+ ├── examples/
900
+ │ └── cache-engine.json
901
+ ├── README.md
902
+ ├── LICENSE
903
+ └── package.json
904
+ ```
905
+
906
+ The OpenCode plugin export remains:
907
+
908
+ ```ts
909
+ export const CacheEngine
910
+ ```
911
+
912
+ regardless of the eventual npm package name.
913
+
914
+ For example, the npm package could be named:
915
+
916
+ ```text
917
+ opencode-cache-engine
918
+ ```
919
+
920
+ without changing the `CacheEngine` export identifier.
921
+
922
+ ---
923
+
924
+ # Installation
925
+
926
+ Install the plugin into the OpenCode plugins directory according to your OpenCode plugin-loading setup.
927
+
928
+ The runtime entry should expose:
929
+
930
+ ```ts
931
+ export const CacheEngine: Plugin = async ({ client, directory }) => {
932
+ // ...
933
+ }
934
+ ```
935
+
936
+ After installation, verify that OpenCode loads the plugin successfully before benchmarking cache behavior.
937
+
938
+ ---
939
+
940
+ # Validation
941
+
942
+ The core test suite can be run with Node:
943
+
944
+ ```bash
945
+ node --test test/cache-engine.test.mjs
946
+ ```
947
+
948
+ The tests are intentionally dependency-light and exercise the pure logic independently of the OpenCode runtime.
949
+
950
+ Runtime validation should additionally confirm:
951
+
952
+ ```text
953
+ DeepSeek:
954
+ no request mutation
955
+
956
+ GPT-5.6:
957
+ promptCacheKey present
958
+ promptCacheOptions present
959
+
960
+ GLM-5.3:
961
+ volatile env block relocated when eligible
962
+ ```
963
+
964
+ ---
965
+
966
+ # Troubleshooting
967
+
968
+ ## DeepSeek cache rate dropped
969
+
970
+ First check provider stability and whether OpenCode's system/tool prefix changed.
971
+
972
+ The plugin itself does not intentionally mutate DeepSeek request options.
973
+
974
+ Inspect the telemetry for:
975
+
976
+ ```text
977
+ prefix-change
978
+ usage-event
979
+ compaction
980
+ ```
981
+
982
+ A prefix change is a diagnostic signal, not automatic proof of a cache miss.
983
+
984
+ ---
985
+
986
+ ## GPT-5.6 cache options are missing
987
+
988
+ Verify that the model is actually classified as GPT-5.6 and that the endpoint is recognized as OpenAI/Azure-compatible.
989
+
990
+ The detector intentionally rejects ambiguous OpenAI-compatible providers rather than guessing.
991
+
992
+ Also check whether the outgoing request already supplied its own cache options. Existing settings are intentionally preserved.
993
+
994
+ ---
995
+
996
+ ## GLM-5.3 prompt is not being changed
997
+
998
+ The environment relocation only occurs when the plugin can identify the expected block unambiguously.
999
+
1000
+ The relevant block must contain the expected beginning and closing marker, and the system structure must meet the plugin's eligibility rules.
1001
+
1002
+ ---
1003
+
1004
+ ## Metrics file is missing
1005
+
1006
+ Telemetry is best-effort.
1007
+
1008
+ Check:
1009
+
1010
+ ```text
1011
+ ~/.cache/opencode/cache-metrics.jsonl
1012
+ ```
1013
+
1014
+ and verify that the configured parent directory is writable.
1015
+
1016
+ A telemetry failure is intentionally swallowed so it does not break model execution.
1017
+
1018
+ ---
1019
+
1020
+ # Status
1021
+
1022
+ The current implementation is intentionally conservative:
1023
+
1024
+ ```text
1025
+ DeepSeek -> preserve and measure
1026
+ GPT-5.6 -> configure cache controls
1027
+ GLM-5.3 -> isolate volatile prompt content
1028
+ ```
1029
+
1030
+ That separation is the core design of the project.
1031
+
1032
+ The plugin should be evaluated using real provider-reported usage and real task cost rather than assuming that any particular local transformation guarantees a cache hit.