opencode-cache-engine 0.1.1 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,6 +1,14 @@
1
1
  # OpenCode Cache Engine
2
2
 
3
- Provider-aware prompt-cache optimization and observability for [OpenCode](https://opencode.ai).
3
+ Provider-aware prompt-cache optimization and observability for
4
+ [OpenCode](https://opencode.ai).
5
+
6
+ `opencode-cache-engine` is an OpenCode npm plugin with two targets:
7
+
8
+ - **Server target** — the actual cache-engine runtime and provider policies.
9
+ - **TUI target** — registration with OpenCode's TUI plugin manager.
10
+
11
+ The server target handles cache optimization, prompt-shape diagnostics, compaction handling, and cache telemetry. The TUI target provides the plugin-manager integration and enable/disable state for the TUI-facing plugin entry.
4
12
 
5
13
  `CacheEngine` is an OpenCode plugin designed for long-running agent sessions where prompt-cache efficiency affects both latency and cost. It keeps the harness conservative for providers whose cache behavior is already automatic, while applying provider-specific optimizations where the provider exposes useful cache controls or where prompt structure can be safely improved.
6
14
 
@@ -14,7 +22,6 @@ The central design principle is:
14
22
 
15
23
  > Optimize the request structure only when there is a clear provider-specific reason to do so. Otherwise, preserve OpenCode's native request behavior and measure what the provider actually reports.
16
24
 
17
- ---
18
25
 
19
26
  ## What this plugin does
20
27
 
@@ -33,7 +40,6 @@ It:
33
40
 
34
41
  The plugin deliberately avoids pretending that a local hash is proof of a provider cache hit. Provider-reported token usage remains the authoritative signal.
35
42
 
36
- ---
37
43
 
38
44
  # Provider behavior
39
45
 
@@ -88,7 +94,6 @@ DeepSeek:
88
94
  force cache behavior
89
95
  ```
90
96
 
91
- ---
92
97
 
93
98
  ## GPT-5.6 Luna
94
99
 
@@ -153,7 +158,6 @@ compaction:
153
158
 
154
159
  This prevents a compaction-specific prompt from sharing the same GPT cache namespace as the normal live-session prompt. The behavior is deterministic and tested explicitly.
155
160
 
156
- ---
157
161
 
158
162
  ## GLM-5.3 Flash
159
163
 
@@ -219,7 +223,6 @@ It only occurs when:
219
223
 
220
224
  The plugin does not arbitrarily rearrange unrelated prompt content.
221
225
 
222
- ---
223
226
 
224
227
  # Prompt-cache strategy
225
228
 
@@ -237,7 +240,6 @@ The plugin is **not** a generic "rewrite every prompt for caching" engine.
237
240
 
238
241
  It is a provider-aware cache policy engine.
239
242
 
240
- ---
241
243
 
242
244
  # System-prompt diagnostics
243
245
 
@@ -261,7 +263,6 @@ It does **not** mean:
261
263
 
262
264
  This distinction is intentional. Provider-reported cache token counts are the authoritative cache signal.
263
265
 
264
- ---
265
266
 
266
267
  # Tool-definition diagnostics
267
268
 
@@ -294,7 +295,6 @@ This distinction matters because semantic equality and byte-level request equali
294
295
 
295
296
  The plugin uses these fingerprints for **diagnostics only**. It does not reorder the tools to force a particular fingerprint.
296
297
 
297
- ---
298
298
 
299
299
  # Compaction handling
300
300
 
@@ -313,7 +313,6 @@ The plugin adds a deterministic continuation template:
313
313
  The digest is inserted once per compaction operation using a guard that prevents duplicate insertion if the compaction hook fires multiple times.
314
314
  The objective is to provide a deterministic continuation structure rather than generating a different arbitrary cache-affecting block on every compaction.
315
315
 
316
- ---
317
316
 
318
317
  # Cache metrics
319
318
 
@@ -388,7 +387,6 @@ Telemetry is best-effort.
388
387
 
389
388
  A failed metrics write must never break an OpenCode request. The recorder catches write failures rather than allowing telemetry failures to affect execution.
390
389
 
391
- ---
392
390
 
393
391
  # Metrics examples
394
392
 
@@ -447,7 +445,6 @@ Telemetry is intended to answer questions such as:
447
445
  * Which provider/model/policy was active?
448
446
  * Did the GLM system stabilization actually change the observed prompt shape?
449
447
 
450
- ---
451
448
 
452
449
  # Configuration
453
450
 
@@ -483,7 +480,6 @@ The default configuration is:
483
480
 
484
481
  The configuration parser starts from these defaults and applies valid file/environment overrides without mutating the caller's configuration object.
485
482
 
486
- ---
487
483
 
488
484
  # Configuration options
489
485
 
@@ -535,7 +531,6 @@ Controls whether the deterministic compaction continuation block is inserted.
535
531
 
536
532
  Controls warning logs for observed prefix-shape changes.
537
533
 
538
- ---
539
534
 
540
535
  # DeepSeek configuration
541
536
 
@@ -549,7 +544,6 @@ There are intentionally very few settings here.
549
544
 
550
545
  DeepSeek is treated as the conservative/passive policy.
551
546
 
552
- ---
553
547
 
554
548
  # GPT-5.6 configuration
555
549
 
@@ -601,7 +595,6 @@ Defaults to:
601
595
 
602
596
  Existing request options are not overwritten by the plugin.
603
597
 
604
- ---
605
598
 
606
599
  # GLM-5.3 configuration
607
600
 
@@ -629,7 +622,6 @@ The reasoning instrumentation is intended to identify anomalies such as:
629
622
 
630
623
  It is diagnostic rather than a reason to rewrite or fabricate reasoning content. The implementation maps these conditions to explicit diagnostic reasons.
631
624
 
632
- ---
633
625
 
634
626
  # Model detection
635
627
 
@@ -658,7 +650,6 @@ Neutral means:
658
650
  no provider-specific request mutation
659
651
  ```
660
652
 
661
- ---
662
653
 
663
654
  # OpenRouter usage
664
655
 
@@ -670,7 +661,6 @@ The plugin does not attempt to compensate for provider switching by rewriting pr
670
661
 
671
662
  For that reason, a stable provider route is preferable when your goal is to measure and maximize prefix reuse.
672
663
 
673
- ---
674
664
 
675
665
  # Architecture
676
666
 
@@ -893,14 +883,15 @@ A typical standalone repository can use:
893
883
  opencode-cache-engine/
894
884
  ├── src/
895
885
  │ ├── cache-engine.ts
896
- │ └── cache-engine-core.mjs
886
+ │ ├── cache-engine-core.mjs
887
+ │ └── tui.mjs
897
888
  ├── test/
898
889
  │ └── cache-engine.test.mjs
899
890
  ├── examples/
900
891
  │ └── cache-engine.json
892
+ ├── package.json
901
893
  ├── README.md
902
- ├── LICENSE
903
- └── package.json
894
+ └── LICENSE
904
895
  ```
905
896
 
906
897
  The OpenCode plugin export remains:
@@ -925,6 +916,20 @@ without changing the `CacheEngine` export identifier.
925
916
 
926
917
  Install the plugin into the OpenCode plugins directory according to your OpenCode plugin-loading setup.
927
918
 
919
+ `opencode-cache-engine` is distributed as an npm package.
920
+
921
+ ## Server/runtime plugin
922
+
923
+ Add the package to the OpenCode runtime plugin configuration:
924
+
925
+ ```json
926
+ {
927
+ "plugin": [
928
+ "opencode-cache-engine"
929
+ ]
930
+ }
931
+ ```
932
+
928
933
  The runtime entry should expose:
929
934
 
930
935
  ```ts
@@ -0,0 +1,965 @@
1
+ # OpenCode Cache Engine
2
+
3
+ Provider-aware prompt-cache optimization and observability for [OpenCode](https://opencode.ai).
4
+
5
+ `opencode-cache-engine` is an OpenCode npm plugin designed for long-running agent sessions where prompt-cache efficiency affects latency, token usage, and total task cost.
6
+
7
+ The package has two targets:
8
+
9
+ - **Server target** — the actual cache engine, provider policies, prompt transformations, compaction handling, and telemetry.
10
+ - **TUI target** — registration with OpenCode's TUI plugin manager.
11
+
12
+ The two targets are intentionally separate. The TUI target does not duplicate the cache-engine implementation.
13
+
14
+ ## Design principle
15
+
16
+ > Optimize the request structure only when there is a clear provider-specific reason to do so. Otherwise, preserve OpenCode's native request behavior and measure what the provider actually reports.
17
+
18
+ Local hashes and structural diagnostics describe observed request shape. They are not proof of a provider cache hit or miss. Provider-reported cache-token usage is the authoritative signal.
19
+
20
+ ---
21
+
22
+ # Provider policies
23
+
24
+ The current implementation has three cache-policy families.
25
+
26
+ | Provider | Policy | Prompt text changed? | Cache metadata changed? | Main strategy |
27
+ | --- | --- | --- | --- | --- |
28
+ | DeepSeek V4.1 Flash | Passive | No | No | Preserve stable harness + observe |
29
+ | GPT-5.6 Luna | Active cache control | No | Yes | Stable cache key + cache options |
30
+ | GLM-5.3 Flash | Input-shape optimization | Yes, narrowly | No provider cache key | Isolate volatile system content |
31
+
32
+ The plugin is intentionally **not** a generic prompt-rewriter. Each provider receives only the behavior justified by its cache model.
33
+
34
+ ---
35
+
36
+ # What the plugin does
37
+
38
+ The server target operates at the OpenCode harness level rather than implementing a provider-specific client.
39
+
40
+ It:
41
+
42
+ 1. Detects the model/provider family in use.
43
+ 2. Applies only the policy appropriate for that family.
44
+ 3. Observes system-prompt and tool-definition stability.
45
+ 4. Records provider-reported cache-token usage.
46
+ 5. Adds a deterministic compaction continuation block.
47
+ 6. Applies GPT-5.6 cache-control metadata.
48
+ 7. Applies the GLM-5.3 volatile-environment relocation.
49
+ 8. Records diagnostics that help correlate request-shape changes with provider cache behavior.
50
+
51
+ Telemetry is best-effort and must never become a dependency of model execution.
52
+
53
+ ---
54
+
55
+ # DeepSeek V4.1 Flash
56
+
57
+ ## Policy: passive
58
+
59
+ DeepSeek receives no cache-specific request mutation.
60
+
61
+ The plugin does not:
62
+
63
+ - rewrite the system prompt
64
+ - reorder tools
65
+ - modify messages
66
+ - inject cache-control fields
67
+ - inject a prompt-cache key
68
+ - alter provider request options
69
+
70
+ The DeepSeek branch exists primarily to preserve a stable harness while providing observability around prefix structure and actual cache usage.
71
+
72
+ The plugin observes:
73
+
74
+ - system-prompt shape
75
+ - semantic tool definitions
76
+ - wire-order tool definitions
77
+ - prefix changes
78
+ - cache read tokens
79
+ - cache write tokens
80
+ - compaction boundaries
81
+
82
+ The rule is deliberately simple:
83
+
84
+ ```text
85
+ DeepSeek:
86
+ preserve request
87
+ preserve prefix
88
+ measure cache
89
+ ```
90
+
91
+ rather than attempting to guess or force the provider's cache behavior.
92
+
93
+ ---
94
+
95
+ # GPT-5.6 Luna
96
+
97
+ ## Policy: active cache control
98
+
99
+ GPT-5.6 is the current policy that actively injects cache-control request metadata.
100
+
101
+ The plugin adds, when the corresponding fields are not already present:
102
+
103
+ ```json
104
+ {
105
+ "promptCacheKey": "<stable-session-key>",
106
+ "promptCacheOptions": {
107
+ "mode": "implicit",
108
+ "ttl": "30m"
109
+ }
110
+ }
111
+ ```
112
+
113
+ The key is derived from stable OpenCode session identity and is independent of transient request data.
114
+
115
+ Existing provider-supplied cache options are not overwritten.
116
+
117
+ ## The prompt text is not rewritten
118
+
119
+ For GPT-5.6:
120
+
121
+ ```text
122
+ system prompt -> unchanged
123
+ conversation -> unchanged
124
+ tool definitions -> unchanged
125
+
126
+ request metadata -> cache key/options added
127
+ ```
128
+
129
+ This controls cache behavior without performing prompt surgery.
130
+
131
+ ## Default GPT configuration
132
+
133
+ ```json
134
+ {
135
+ "promptCacheKey": true,
136
+ "cacheRootKey": false,
137
+ "compactionCacheIsolation": true,
138
+ "reasoningEffortDiagnostics": true,
139
+ "mode": "implicit",
140
+ "ttl": "30m"
141
+ }
142
+ ```
143
+
144
+ `cacheRootKey` is disabled by default because the current OpenCode runtime does not expose sufficiently reliable fork lineage for safe cross-fork cache-root inheritance.
145
+
146
+ ## Compaction isolation
147
+
148
+ Compaction uses a deterministic separate cache-key namespace:
149
+
150
+ ```text
151
+ live session:
152
+ ses_abc123
153
+
154
+ compaction:
155
+ ses_abc123:compact
156
+ ```
157
+
158
+ This prevents compaction-specific cache writes from sharing the normal live-session namespace.
159
+
160
+ ---
161
+
162
+ # GLM-5.3 Flash
163
+
164
+ ## Policy: input-shape optimization
165
+
166
+ GLM-5.3 receives the only current prompt-text transformation.
167
+
168
+ The plugin identifies OpenCode's volatile `<env>` section and moves it to the **tail of the system prompt**.
169
+
170
+ Conceptually:
171
+
172
+ ```text
173
+ BEFORE
174
+
175
+ [large stable instructions]
176
+ [volatile environment/date block]
177
+ [more stable instructions]
178
+ ```
179
+
180
+ becomes:
181
+
182
+ ```text
183
+ AFTER
184
+
185
+ [large stable instructions]
186
+ [more stable instructions]
187
+ [volatile environment/date block]
188
+ ```
189
+
190
+ The environment block's contents are preserved. Only its position changes.
191
+
192
+ The objective is to isolate volatile information so that a changing date or environment value does not unnecessarily disturb the earlier stable prefix.
193
+
194
+ ## GLM safety constraints
195
+
196
+ The transformation occurs only when:
197
+
198
+ - the selected model is GLM-5.3;
199
+ - GLM stabilization is enabled;
200
+ - there is exactly one system string;
201
+ - the expected environment markers exist; and
202
+ - the block can be identified unambiguously.
203
+
204
+ The plugin does not arbitrarily reorder unrelated system-prompt content.
205
+
206
+ ---
207
+
208
+ # System-prompt diagnostics
209
+
210
+ The plugin fingerprints the system prompt to detect structural changes between requests.
211
+
212
+ Provider-aware diagnostics track:
213
+
214
+ - full system hash
215
+ - stable system-prefix hash
216
+ - volatile system-suffix hash
217
+
218
+ The stable/volatile decomposition is based on the longest common prefix against the session baseline.
219
+
220
+ A changed hash means:
221
+
222
+ ```text
223
+ The observed request bytes changed.
224
+ ```
225
+
226
+ It does **not** mean:
227
+
228
+ ```text
229
+ The provider definitely generated a cache miss.
230
+ ```
231
+
232
+ Provider-reported cache-token usage remains authoritative.
233
+
234
+ ---
235
+
236
+ # Tool-definition diagnostics
237
+
238
+ Tool definitions are normalized before fingerprinting.
239
+
240
+ Runtime-only fields such as object identity, function references, timestamps, and arbitrary runtime metadata are excluded.
241
+
242
+ The plugin maintains two useful fingerprints:
243
+
244
+ - **semantic fingerprint** — order-insensitive representation of model-visible tool definitions;
245
+ - **wire-order fingerprint** — order-sensitive diagnostic representation of the closest deterministic pre-wire tool ordering available to the plugin.
246
+
247
+ These fingerprints are diagnostic only. The plugin does not reorder tools merely to force a particular fingerprint.
248
+
249
+ ---
250
+
251
+ # Compaction handling
252
+
253
+ OpenCode sessions eventually undergo compaction as conversation history grows.
254
+
255
+ The plugin adds a deterministic continuation template:
256
+
257
+ ```text
258
+ ## Session digest (cache-stable continuation block)
259
+ - Goal:
260
+ - Decisions made:
261
+ - Pending:
262
+ - Active files:
263
+ ```
264
+
265
+ The digest is inserted once per compaction invocation using a guard that prevents duplicate insertion if the hook fires more than once.
266
+
267
+ The objective is to provide a deterministic continuation structure rather than generating a different arbitrary cache-affecting block on every compaction.
268
+
269
+ ---
270
+
271
+ # Cache metrics
272
+
273
+ The plugin records cache usage from OpenCode assistant-message token data.
274
+
275
+ At minimum it tracks:
276
+
277
+ ```text
278
+ cache.read
279
+ cache.write
280
+ ```
281
+
282
+ and aggregates those values across the session.
283
+
284
+ ## Generic cache ratio
285
+
286
+ The core helper reports:
287
+
288
+ ```text
289
+ hit rate = read / (read + write)
290
+ ```
291
+
292
+ This is an accounting metric based on cache read/write tokens.
293
+
294
+ ## GLM prompt-token ratio
295
+
296
+ For GLM, the implementation additionally calculates:
297
+
298
+ ```text
299
+ read / (read + write + input)
300
+ ```
301
+
302
+ These ratios answer different questions and should not be treated as interchangeable.
303
+
304
+ ---
305
+
306
+ # Telemetry
307
+
308
+ Metrics are written as JSONL.
309
+
310
+ Default metrics path:
311
+
312
+ ```text
313
+ ~/.cache/opencode/cache-metrics.jsonl
314
+ ```
315
+
316
+ Default plugin configuration path:
317
+
318
+ ```text
319
+ ~/.config/opencode/cache-engine.json
320
+ ```
321
+
322
+ Telemetry is best-effort. A failed metrics write must never break an OpenCode request.
323
+
324
+ ## Example usage record
325
+
326
+ ```json
327
+ {
328
+ "kind": "usage-event",
329
+ "sid": "session-id",
330
+ "ts": 1750000000000,
331
+ "read": 120000,
332
+ "write": 3000,
333
+ "cost": 0.0123,
334
+ "provider": "z-ai",
335
+ "model": "glm-5.3-flash",
336
+ "policy": "glm53"
337
+ }
338
+ ```
339
+
340
+ ## Example prefix-change record
341
+
342
+ ```json
343
+ {
344
+ "kind": "prefix-change",
345
+ "sid": "session-id",
346
+ "ts": 1750000000000,
347
+ "dimensions": [
348
+ "system"
349
+ ]
350
+ }
351
+ ```
352
+
353
+ ## Example compaction record
354
+
355
+ ```json
356
+ {
357
+ "kind": "compaction",
358
+ "sid": "session-id",
359
+ "ts": 1750000000000,
360
+ "reason": "compaction",
361
+ "usageSamples": 7,
362
+ "cumulative": {
363
+ "read": 900000,
364
+ "write": 12000
365
+ }
366
+ }
367
+ ```
368
+
369
+ Telemetry is intended to answer questions such as:
370
+
371
+ - Did the system prompt change?
372
+ - Did the tool definitions change?
373
+ - Did cache reads increase?
374
+ - Did cache writes increase?
375
+ - Did a compaction occur?
376
+ - Which provider/model/policy was active?
377
+ - Did provider-specific system stabilization change the observed prompt shape?
378
+
379
+ ---
380
+
381
+ # Configuration
382
+
383
+ A representative default configuration is:
384
+
385
+ ```json
386
+ {
387
+ "enabled": true,
388
+ "metricsFile": "~/.cache/opencode/cache-metrics.jsonl",
389
+ "compactTemplate": true,
390
+ "logPrefixChanges": true,
391
+ "policies": {
392
+ "deepseek": {
393
+ "enabled": true
394
+ },
395
+ "gpt56": {
396
+ "enabled": true,
397
+ "promptCacheKey": true,
398
+ "cacheRootKey": false,
399
+ "compactionCacheIsolation": true,
400
+ "reasoningEffortDiagnostics": true,
401
+ "mode": "implicit",
402
+ "ttl": "30m"
403
+ },
404
+ "glm53": {
405
+ "enabled": true,
406
+ "stabilizeSystem": true,
407
+ "preserveThinkingIntegrity": true
408
+ }
409
+ }
410
+ }
411
+ ```
412
+
413
+ The configuration parser starts from defaults and applies valid file/environment overrides without mutating the caller's object.
414
+
415
+ ---
416
+
417
+ # Configuration options
418
+
419
+ ## Global
420
+
421
+ ### `enabled`
422
+
423
+ Enables or disables the entire server plugin.
424
+
425
+ ### `metricsFile`
426
+
427
+ Controls where JSONL telemetry is written.
428
+
429
+ ### `compactTemplate`
430
+
431
+ Controls whether the deterministic compaction continuation block is inserted.
432
+
433
+ ### `logPrefixChanges`
434
+
435
+ Controls warning logs for observed prefix-shape changes.
436
+
437
+ ## DeepSeek
438
+
439
+ ```json
440
+ "deepseek": {
441
+ "enabled": true
442
+ }
443
+ ```
444
+
445
+ There are intentionally few DeepSeek settings because the policy is passive.
446
+
447
+ ## GPT-5.6
448
+
449
+ ```json
450
+ "gpt56": {
451
+ "enabled": true,
452
+ "promptCacheKey": true,
453
+ "cacheRootKey": false,
454
+ "compactionCacheIsolation": true,
455
+ "reasoningEffortDiagnostics": true,
456
+ "mode": "implicit",
457
+ "ttl": "30m"
458
+ }
459
+ ```
460
+
461
+ ### `promptCacheKey`
462
+
463
+ Controls whether the plugin provides a stable session-derived GPT cache key.
464
+
465
+ ### `cacheRootKey`
466
+
467
+ Controls parent/fork cache-root inheritance. Disabled by default because reliable fork lineage is not currently guaranteed by the runtime.
468
+
469
+ ### `compactionCacheIsolation`
470
+
471
+ Uses a separate deterministic cache namespace for compaction requests.
472
+
473
+ ### `reasoningEffortDiagnostics`
474
+
475
+ Tracks GPT reasoning-effort changes for diagnostics.
476
+
477
+ ### `mode`
478
+
479
+ Defaults to `implicit`.
480
+
481
+ ### `ttl`
482
+
483
+ Defaults to `30m`.
484
+
485
+ Existing request options are not overwritten.
486
+
487
+ ## GLM-5.3
488
+
489
+ ```json
490
+ "glm53": {
491
+ "enabled": true,
492
+ "stabilizeSystem": true,
493
+ "preserveThinkingIntegrity": true
494
+ }
495
+ ```
496
+
497
+ ### `stabilizeSystem`
498
+
499
+ Enables relocation of the volatile `<env>` section to the system-prompt tail.
500
+
501
+ ### `preserveThinkingIntegrity`
502
+
503
+ Enables diagnostic checks around reasoning continuity.
504
+
505
+ The reasoning instrumentation is diagnostic only. It does not rewrite, fabricate, or reorder reasoning content.
506
+
507
+ ---
508
+
509
+ # Model detection
510
+
511
+ The server plugin classifies requests into:
512
+
513
+ ```text
514
+ deepseek
515
+ gpt56
516
+ glm53
517
+ neutral
518
+ ```
519
+
520
+ The detector recognizes provider/model identifiers for the supported policy families.
521
+
522
+ GPT-5.6 has an additional OpenAI/Azure-context check so a string containing `gpt-5.6` does not automatically cause GPT-specific fields to be sent to an unrelated endpoint.
523
+
524
+ Unknown models use the neutral policy:
525
+
526
+ ```text
527
+ no provider-specific request mutation
528
+ ```
529
+
530
+ ---
531
+
532
+ # OpenRouter usage
533
+
534
+ The plugin is compatible with OpenRouter because policy selection is based on the provider/model information exposed by OpenCode.
535
+
536
+ For cache-sensitive workloads, provider stability is important.
537
+
538
+ The plugin does not attempt to compensate for provider switching by rewriting prompts. Stable provider routing is preferable when the objective is to preserve reusable prefixes.
539
+
540
+ ---
541
+
542
+ # Package architecture
543
+
544
+ `opencode-cache-engine` is an npm package with separate server and TUI targets.
545
+
546
+ ```text
547
+ opencode-cache-engine/
548
+ ├── src/
549
+ │ ├── cache-engine.ts
550
+ │ ├── cache-engine-core.mjs
551
+ │ └── tui.mjs
552
+ ├── test/
553
+ │ └── cache-engine.test.mjs
554
+ ├── examples/
555
+ │ └── cache-engine.json
556
+ ├── package.json
557
+ ├── README.md
558
+ └── LICENSE
559
+ ```
560
+
561
+ The package exports two OpenCode targets:
562
+
563
+ ```json
564
+ {
565
+ "exports": {
566
+ "./server": "./src/cache-engine.ts",
567
+ "./tui": "./src/tui.mjs"
568
+ }
569
+ }
570
+ ```
571
+
572
+ ## Server target
573
+
574
+ `./server` points to the actual CacheEngine implementation.
575
+
576
+ It owns:
577
+
578
+ - OpenCode server hooks
579
+ - session state
580
+ - provider-policy selection
581
+ - request mutation
582
+ - system-prompt transformation
583
+ - telemetry
584
+ - compaction handling
585
+
586
+ The implementation exports:
587
+
588
+ ```ts
589
+ export const CacheEngine: Plugin = async ({ client, directory }) => {
590
+ // ...
591
+ }
592
+ ```
593
+
594
+ The identifier `CacheEngine` is the code export name. It does not determine the npm package name.
595
+
596
+ ## TUI target
597
+
598
+ `./tui` is a small TUI-only registration module.
599
+
600
+ Its purpose is to make the npm package discoverable by OpenCode's TUI plugin manager.
601
+
602
+ The TUI target does not duplicate or contain cache-engine logic.
603
+
604
+ TUI activation state and server-runtime activation are separate concepts unless an explicit state bridge is implemented.
605
+
606
+ ---
607
+
608
+ # Installation
609
+
610
+ ## npm package
611
+
612
+ The package name is:
613
+
614
+ ```text
615
+ opencode-cache-engine
616
+ ```
617
+
618
+ For the normal runtime/release configuration, add the package to the OpenCode server plugin configuration:
619
+
620
+ ```jsonc
621
+ {
622
+ "plugin": [
623
+ "opencode-cache-engine"
624
+ ]
625
+ }
626
+ ```
627
+
628
+ Add the same package to the TUI plugin registry when you want it displayed and managed by the OpenCode TUI plugin manager:
629
+
630
+ ```json
631
+ {
632
+ "plugin": [
633
+ "opencode-cache-engine"
634
+ ],
635
+ "$schema": "https://www.opencode.ai/tui.json"
636
+ }
637
+ ```
638
+
639
+ The package must expose both server and TUI targets for the TUI registry to accept it.
640
+
641
+ ## No direct local plugin copy
642
+
643
+ Do not install a second CacheEngine copy by placing `cache-engine.ts` or related source files under:
644
+
645
+ ```text
646
+ ~/.config/opencode/plugins/
647
+ ```
648
+
649
+ when using the configured npm package.
650
+
651
+ A duplicate local copy can cause multiple plugin instances to be loaded from different sources.
652
+
653
+ ---
654
+
655
+ # Development from Git
656
+
657
+ The Git repository is the development source of truth.
658
+
659
+ During active development, work directly in the checkout:
660
+
661
+ ```text
662
+ ~/Desktop/opencode-cache-engine/
663
+ ```
664
+
665
+ The development workflow is:
666
+
667
+ ```text
668
+ Git working tree
669
+ ↓
670
+ unit tests
671
+ ↓
672
+ OpenCode local development target
673
+ ↓
674
+ runtime validation
675
+ ↓
676
+ git commit
677
+ ↓
678
+ npm release
679
+ ```
680
+
681
+ For local development, the server target can be loaded directly from the Git checkout through OpenCode's local plugin configuration rather than publishing an npm version for every change.
682
+
683
+ The repository remains the only place where source code is edited.
684
+
685
+ ---
686
+
687
+ # Testing
688
+
689
+ ## Unit tests
690
+
691
+ Run the test suite from the repository root:
692
+
693
+ ```bash
694
+ node --test test/cache-engine.test.mjs
695
+ ```
696
+
697
+ The test file imports the pure core module from:
698
+
699
+ ```js
700
+ from "../src/cache-engine-core.mjs"
701
+ ```
702
+
703
+ The tests validate provider-independent and provider-specific logic, including:
704
+
705
+ - model detection
706
+ - GPT cache-key stability
707
+ - GPT cache-option defaults
708
+ - protection against overwriting existing cache options
709
+ - GLM environment relocation
710
+ - deterministic hashing
711
+ - system-prefix decomposition
712
+ - tool fingerprints
713
+ - reasoning diagnostics
714
+ - compaction isolation
715
+ - cache-hit calculations
716
+ - configuration behavior
717
+ - JSONL telemetry behavior
718
+
719
+ Runtime behavior is validated separately through actual OpenCode plugin loading.
720
+
721
+ ## Package validation
722
+
723
+ Before publishing, inspect the package contents:
724
+
725
+ ```bash
726
+ npm pack --dry-run
727
+ ```
728
+
729
+ The package should contain the runtime source and documentation needed by OpenCode, including:
730
+
731
+ ```text
732
+ src/cache-engine.ts
733
+ src/cache-engine-core.mjs
734
+ src/tui.mjs
735
+ package.json
736
+ README.md
737
+ LICENSE
738
+ examples/...
739
+ ```
740
+
741
+ Test the TUI target directly with Node when appropriate:
742
+
743
+ ```bash
744
+ node -e 'import("./src/tui.mjs").then(m => console.log(m.default))'
745
+ ```
746
+
747
+ ---
748
+
749
+ # Design principles
750
+
751
+ ## 1. Provider-specific behavior
752
+
753
+ Different providers expose different cache mechanisms.
754
+
755
+ The plugin therefore does not assume that one strategy is optimal everywhere.
756
+
757
+ ## 2. Preserve working behavior
758
+
759
+ The plugin should not modify a provider's request merely because a mutation is technically possible.
760
+
761
+ This is especially important for DeepSeek, where the current policy is intentionally passive.
762
+
763
+ ## 3. Measure provider reality
764
+
765
+ Local hashes are diagnostics.
766
+
767
+ Provider-reported cache-token counts are authoritative.
768
+
769
+ The implementation explicitly distinguishes:
770
+
771
+ ```text
772
+ observed prefix change
773
+ ```
774
+
775
+ from:
776
+
777
+ ```text
778
+ confirmed provider cache miss
779
+ ```
780
+
781
+ because the plugin cannot infer the latter reliably from local hashes alone.
782
+
783
+ ## 4. Never overwrite explicit provider configuration
784
+
785
+ Where GPT cache options already exist, the plugin leaves them alone.
786
+
787
+ ## 5. Keep mutations deterministic
788
+
789
+ When the plugin transforms a request, the transformation should be:
790
+
791
+ - narrow
792
+ - deterministic
793
+ - content-preserving where possible
794
+ - provider-specific
795
+ - easy to disable
796
+
797
+ ## 6. Keep telemetry out of the critical path
798
+
799
+ A metrics failure must not break model execution.
800
+
801
+ ---
802
+
803
+ # What the plugin does NOT do
804
+
805
+ The plugin does not:
806
+
807
+ - invent cache hits;
808
+ - claim a local hash proves a provider cache hit;
809
+ - rewrite DeepSeek prompts;
810
+ - reorder tools merely to force cache reuse;
811
+ - fabricate reasoning;
812
+ - modify conversation history arbitrarily;
813
+ - force explicit GPT cache breakpoints by default;
814
+ - silently overwrite existing GPT cache options;
815
+ - assume every model named `gpt-5.6` is an OpenAI-compatible endpoint; or
816
+ - use fork inheritance unless reliable lineage is available.
817
+
818
+ ---
819
+
820
+ # Cost optimization philosophy
821
+
822
+ Cache hit rate is useful, but it is not the only cost metric.
823
+
824
+ The economic objective is:
825
+
826
+ ```text
827
+ total task cost
828
+ =
829
+ prompt/cache cost
830
+ +
831
+ output/reasoning cost
832
+ +
833
+ additional requests
834
+ ```
835
+
836
+ A model with a slightly lower cache hit rate can still be cheaper if it completes the task with fewer tokens or fewer model calls.
837
+
838
+ For that reason, this plugin is primarily an **instrumentation + targeted optimization layer**, not a cache-rate maximizer at any cost.
839
+
840
+ The preferred evaluation unit is:
841
+
842
+ ```text
843
+ cost per completed task
844
+ ```
845
+
846
+ rather than cache percentage alone.
847
+
848
+ ---
849
+
850
+ # Operational recommendations
851
+
852
+ For reliable cache measurements:
853
+
854
+ 1. Keep the provider fixed whenever possible.
855
+ 2. Avoid changing unrelated system-prompt content during a benchmark.
856
+ 3. Keep tool definitions stable.
857
+ 4. Compare equivalent tasks across models.
858
+ 5. Record actual provider cache-token counts.
859
+ 6. Compare total task cost, not only cache percentage.
860
+ 7. Treat compaction as a separate cache boundary when analyzing results.
861
+ 8. Avoid interpreting a local prefix hash change as definitive proof of a cache miss.
862
+
863
+ ---
864
+
865
+ # Troubleshooting
866
+
867
+ ## `there is no TUI target`
868
+
869
+ This means the npm package does not expose a TUI target that your installed OpenCode version recognizes.
870
+
871
+ Verify that the package contains a TUI entrypoint and that `package.json` exposes the required `./tui` target.
872
+
873
+ After changing the package, update/reinstall the npm package and restart OpenCode.
874
+
875
+ ## The old `oc-plugin-caching` still appears
876
+
877
+ Search all OpenCode configuration and local-plugin locations for the old plugin name:
878
+
879
+ ```bash
880
+ grep -R "oc-plugin-caching" ~/.config/opencode ~/.opencode .opencode 2>/dev/null
881
+ ```
882
+
883
+ Remove the old configuration entry and remove the obsolete local plugin source copy.
884
+
885
+ ## DeepSeek cache rate dropped
886
+
887
+ First check provider stability and whether OpenCode's system/tool prefix changed.
888
+
889
+ The plugin does not intentionally mutate DeepSeek request options.
890
+
891
+ Inspect telemetry for:
892
+
893
+ ```text
894
+ prefix-change
895
+ usage-event
896
+ compaction
897
+ ```
898
+
899
+ A prefix change is a diagnostic signal, not automatic proof of a cache miss.
900
+
901
+ ## GPT-5.6 cache options are missing
902
+
903
+ Verify that the model is classified as GPT-5.6 and that the endpoint is recognized as OpenAI/Azure-compatible.
904
+
905
+ Also check whether the outgoing request already supplied its own cache options. Existing settings are intentionally preserved.
906
+
907
+ ## GLM-5.3 prompt is not being changed
908
+
909
+ The environment relocation only occurs when the plugin can identify the expected block unambiguously.
910
+
911
+ The relevant block must contain the expected beginning and closing marker, and the system structure must meet the plugin's eligibility rules.
912
+
913
+ ## Metrics file is missing
914
+
915
+ Telemetry is best-effort.
916
+
917
+ Check:
918
+
919
+ ```text
920
+ ~/.cache/opencode/cache-metrics.jsonl
921
+ ```
922
+
923
+ and verify that the configured parent directory is writable.
924
+
925
+ ---
926
+
927
+ # Release workflow
928
+
929
+ The Git repository is the source of truth for development.
930
+
931
+ A normal release workflow is:
932
+
933
+ ```text
934
+ edit source
935
+ ↓
936
+ run tests
937
+ ↓
938
+ run OpenCode runtime validation
939
+ ↓
940
+ git commit
941
+ ↓
942
+ bump package version
943
+ ↓
944
+ npm publish
945
+ ↓
946
+ use published opencode-cache-engine package
947
+ ```
948
+
949
+ Do not edit the installed npm cache as a substitute for repository development.
950
+
951
+ ---
952
+
953
+ # Status
954
+
955
+ The current implementation is intentionally conservative:
956
+
957
+ ```text
958
+ DeepSeek -> preserve and measure
959
+ GPT-5.6 -> configure cache controls
960
+ GLM-5.3 -> isolate volatile prompt content
961
+ ```
962
+
963
+ That separation is the core design of the project.
964
+
965
+ The plugin should be evaluated using real provider-reported usage and real task cost rather than assuming that any particular local transformation guarantees a cache hit.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "opencode-cache-engine",
3
- "version": "0.1.1",
3
+ "version": "0.2.0",
4
4
  "private": false,
5
5
  "description": "Provider-aware prompt-cache optimization and telemetry for OpenCode",
6
6
  "keywords": [
@@ -27,6 +27,10 @@
27
27
  "example": "examples",
28
28
  "test": "test"
29
29
  },
30
+ "exports": {
31
+ "./server": "./src/cache-engine.ts",
32
+ "./tui": "./src/tui.mjs"
33
+ },
30
34
  "scripts": {
31
35
  "test": "node --test test/cache-engine.test.mjs",
32
36
  "postpublish": "VERSION=$(node -p \"require('./package').version\") && TAG=v$VERSION && echo \"Creating git tag $TAG for npm $VERSION\" && git tag $TAG && git push origin $TAG || echo \"Tag $TAG may already exist or push failed; continuing.\"",
package/src/tui.mjs ADDED
@@ -0,0 +1,11 @@
1
+ const plugin = {
2
+ id: "opencode-cache-engine",
3
+
4
+ async tui() {
5
+ // CacheEngine has no TUI UI of its own.
6
+ // This target exists so the npm package can be registered,
7
+ // displayed, and enabled/disabled by the OpenCode plugin manager.
8
+ },
9
+ }
10
+
11
+ export default plugin