opencode-cache-engine 0.1.1 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +27 -22
- package/README_rewritten.md +965 -0
- package/package.json +5 -1
- package/src/tui.mjs +11 -0
package/README.md
CHANGED
|
@@ -1,6 +1,14 @@
|
|
|
1
1
|
# OpenCode Cache Engine
|
|
2
2
|
|
|
3
|
-
Provider-aware prompt-cache optimization and observability for
|
|
3
|
+
Provider-aware prompt-cache optimization and observability for
|
|
4
|
+
[OpenCode](https://opencode.ai).
|
|
5
|
+
|
|
6
|
+
`opencode-cache-engine` is an OpenCode npm plugin with two targets:
|
|
7
|
+
|
|
8
|
+
- **Server target** — the actual cache-engine runtime and provider policies.
|
|
9
|
+
- **TUI target** — registration with OpenCode's TUI plugin manager.
|
|
10
|
+
|
|
11
|
+
The server target handles cache optimization, prompt-shape diagnostics, compaction handling, and cache telemetry. The TUI target provides the plugin-manager integration and enable/disable state for the TUI-facing plugin entry.
|
|
4
12
|
|
|
5
13
|
`CacheEngine` is an OpenCode plugin designed for long-running agent sessions where prompt-cache efficiency affects both latency and cost. It keeps the harness conservative for providers whose cache behavior is already automatic, while applying provider-specific optimizations where the provider exposes useful cache controls or where prompt structure can be safely improved.
|
|
6
14
|
|
|
@@ -14,7 +22,6 @@ The central design principle is:
|
|
|
14
22
|
|
|
15
23
|
> Optimize the request structure only when there is a clear provider-specific reason to do so. Otherwise, preserve OpenCode's native request behavior and measure what the provider actually reports.
|
|
16
24
|
|
|
17
|
-
---
|
|
18
25
|
|
|
19
26
|
## What this plugin does
|
|
20
27
|
|
|
@@ -33,7 +40,6 @@ It:
|
|
|
33
40
|
|
|
34
41
|
The plugin deliberately avoids pretending that a local hash is proof of a provider cache hit. Provider-reported token usage remains the authoritative signal.
|
|
35
42
|
|
|
36
|
-
---
|
|
37
43
|
|
|
38
44
|
# Provider behavior
|
|
39
45
|
|
|
@@ -88,7 +94,6 @@ DeepSeek:
|
|
|
88
94
|
force cache behavior
|
|
89
95
|
```
|
|
90
96
|
|
|
91
|
-
---
|
|
92
97
|
|
|
93
98
|
## GPT-5.6 Luna
|
|
94
99
|
|
|
@@ -153,7 +158,6 @@ compaction:
|
|
|
153
158
|
|
|
154
159
|
This prevents a compaction-specific prompt from sharing the same GPT cache namespace as the normal live-session prompt. The behavior is deterministic and tested explicitly.
|
|
155
160
|
|
|
156
|
-
---
|
|
157
161
|
|
|
158
162
|
## GLM-5.3 Flash
|
|
159
163
|
|
|
@@ -219,7 +223,6 @@ It only occurs when:
|
|
|
219
223
|
|
|
220
224
|
The plugin does not arbitrarily rearrange unrelated prompt content.
|
|
221
225
|
|
|
222
|
-
---
|
|
223
226
|
|
|
224
227
|
# Prompt-cache strategy
|
|
225
228
|
|
|
@@ -237,7 +240,6 @@ The plugin is **not** a generic "rewrite every prompt for caching" engine.
|
|
|
237
240
|
|
|
238
241
|
It is a provider-aware cache policy engine.
|
|
239
242
|
|
|
240
|
-
---
|
|
241
243
|
|
|
242
244
|
# System-prompt diagnostics
|
|
243
245
|
|
|
@@ -261,7 +263,6 @@ It does **not** mean:
|
|
|
261
263
|
|
|
262
264
|
This distinction is intentional. Provider-reported cache token counts are the authoritative cache signal.
|
|
263
265
|
|
|
264
|
-
---
|
|
265
266
|
|
|
266
267
|
# Tool-definition diagnostics
|
|
267
268
|
|
|
@@ -294,7 +295,6 @@ This distinction matters because semantic equality and byte-level request equali
|
|
|
294
295
|
|
|
295
296
|
The plugin uses these fingerprints for **diagnostics only**. It does not reorder the tools to force a particular fingerprint.
|
|
296
297
|
|
|
297
|
-
---
|
|
298
298
|
|
|
299
299
|
# Compaction handling
|
|
300
300
|
|
|
@@ -313,7 +313,6 @@ The plugin adds a deterministic continuation template:
|
|
|
313
313
|
The digest is inserted once per compaction operation using a guard that prevents duplicate insertion if the compaction hook fires multiple times.
|
|
314
314
|
The objective is to provide a deterministic continuation structure rather than generating a different arbitrary cache-affecting block on every compaction.
|
|
315
315
|
|
|
316
|
-
---
|
|
317
316
|
|
|
318
317
|
# Cache metrics
|
|
319
318
|
|
|
@@ -388,7 +387,6 @@ Telemetry is best-effort.
|
|
|
388
387
|
|
|
389
388
|
A failed metrics write must never break an OpenCode request. The recorder catches write failures rather than allowing telemetry failures to affect execution.
|
|
390
389
|
|
|
391
|
-
---
|
|
392
390
|
|
|
393
391
|
# Metrics examples
|
|
394
392
|
|
|
@@ -447,7 +445,6 @@ Telemetry is intended to answer questions such as:
|
|
|
447
445
|
* Which provider/model/policy was active?
|
|
448
446
|
* Did the GLM system stabilization actually change the observed prompt shape?
|
|
449
447
|
|
|
450
|
-
---
|
|
451
448
|
|
|
452
449
|
# Configuration
|
|
453
450
|
|
|
@@ -483,7 +480,6 @@ The default configuration is:
|
|
|
483
480
|
|
|
484
481
|
The configuration parser starts from these defaults and applies valid file/environment overrides without mutating the caller's configuration object.
|
|
485
482
|
|
|
486
|
-
---
|
|
487
483
|
|
|
488
484
|
# Configuration options
|
|
489
485
|
|
|
@@ -535,7 +531,6 @@ Controls whether the deterministic compaction continuation block is inserted.
|
|
|
535
531
|
|
|
536
532
|
Controls warning logs for observed prefix-shape changes.
|
|
537
533
|
|
|
538
|
-
---
|
|
539
534
|
|
|
540
535
|
# DeepSeek configuration
|
|
541
536
|
|
|
@@ -549,7 +544,6 @@ There are intentionally very few settings here.
|
|
|
549
544
|
|
|
550
545
|
DeepSeek is treated as the conservative/passive policy.
|
|
551
546
|
|
|
552
|
-
---
|
|
553
547
|
|
|
554
548
|
# GPT-5.6 configuration
|
|
555
549
|
|
|
@@ -601,7 +595,6 @@ Defaults to:
|
|
|
601
595
|
|
|
602
596
|
Existing request options are not overwritten by the plugin.
|
|
603
597
|
|
|
604
|
-
---
|
|
605
598
|
|
|
606
599
|
# GLM-5.3 configuration
|
|
607
600
|
|
|
@@ -629,7 +622,6 @@ The reasoning instrumentation is intended to identify anomalies such as:
|
|
|
629
622
|
|
|
630
623
|
It is diagnostic rather than a reason to rewrite or fabricate reasoning content. The implementation maps these conditions to explicit diagnostic reasons.
|
|
631
624
|
|
|
632
|
-
---
|
|
633
625
|
|
|
634
626
|
# Model detection
|
|
635
627
|
|
|
@@ -658,7 +650,6 @@ Neutral means:
|
|
|
658
650
|
no provider-specific request mutation
|
|
659
651
|
```
|
|
660
652
|
|
|
661
|
-
---
|
|
662
653
|
|
|
663
654
|
# OpenRouter usage
|
|
664
655
|
|
|
@@ -670,7 +661,6 @@ The plugin does not attempt to compensate for provider switching by rewriting pr
|
|
|
670
661
|
|
|
671
662
|
For that reason, a stable provider route is preferable when your goal is to measure and maximize prefix reuse.
|
|
672
663
|
|
|
673
|
-
---
|
|
674
664
|
|
|
675
665
|
# Architecture
|
|
676
666
|
|
|
@@ -893,14 +883,15 @@ A typical standalone repository can use:
|
|
|
893
883
|
opencode-cache-engine/
|
|
894
884
|
├── src/
|
|
895
885
|
│ ├── cache-engine.ts
|
|
896
|
-
│
|
|
886
|
+
│ ├── cache-engine-core.mjs
|
|
887
|
+
│ └── tui.mjs
|
|
897
888
|
├── test/
|
|
898
889
|
│ └── cache-engine.test.mjs
|
|
899
890
|
├── examples/
|
|
900
891
|
│ └── cache-engine.json
|
|
892
|
+
├── package.json
|
|
901
893
|
├── README.md
|
|
902
|
-
|
|
903
|
-
└── package.json
|
|
894
|
+
└── LICENSE
|
|
904
895
|
```
|
|
905
896
|
|
|
906
897
|
The OpenCode plugin export remains:
|
|
@@ -925,6 +916,20 @@ without changing the `CacheEngine` export identifier.
|
|
|
925
916
|
|
|
926
917
|
Install the plugin into the OpenCode plugins directory according to your OpenCode plugin-loading setup.
|
|
927
918
|
|
|
919
|
+
`opencode-cache-engine` is distributed as an npm package.
|
|
920
|
+
|
|
921
|
+
## Server/runtime plugin
|
|
922
|
+
|
|
923
|
+
Add the package to the OpenCode runtime plugin configuration:
|
|
924
|
+
|
|
925
|
+
```json
|
|
926
|
+
{
|
|
927
|
+
"plugin": [
|
|
928
|
+
"opencode-cache-engine"
|
|
929
|
+
]
|
|
930
|
+
}
|
|
931
|
+
```
|
|
932
|
+
|
|
928
933
|
The runtime entry should expose:
|
|
929
934
|
|
|
930
935
|
```ts
|
|
@@ -0,0 +1,965 @@
|
|
|
1
|
+
# OpenCode Cache Engine
|
|
2
|
+
|
|
3
|
+
Provider-aware prompt-cache optimization and observability for [OpenCode](https://opencode.ai).
|
|
4
|
+
|
|
5
|
+
`opencode-cache-engine` is an OpenCode npm plugin designed for long-running agent sessions where prompt-cache efficiency affects latency, token usage, and total task cost.
|
|
6
|
+
|
|
7
|
+
The package has two targets:
|
|
8
|
+
|
|
9
|
+
- **Server target** — the actual cache engine, provider policies, prompt transformations, compaction handling, and telemetry.
|
|
10
|
+
- **TUI target** — registration with OpenCode's TUI plugin manager.
|
|
11
|
+
|
|
12
|
+
The two targets are intentionally separate. The TUI target does not duplicate the cache-engine implementation.
|
|
13
|
+
|
|
14
|
+
## Design principle
|
|
15
|
+
|
|
16
|
+
> Optimize the request structure only when there is a clear provider-specific reason to do so. Otherwise, preserve OpenCode's native request behavior and measure what the provider actually reports.
|
|
17
|
+
|
|
18
|
+
Local hashes and structural diagnostics describe observed request shape. They are not proof of a provider cache hit or miss. Provider-reported cache-token usage is the authoritative signal.
|
|
19
|
+
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
# Provider policies
|
|
23
|
+
|
|
24
|
+
The current implementation has three cache-policy families.
|
|
25
|
+
|
|
26
|
+
| Provider | Policy | Prompt text changed? | Cache metadata changed? | Main strategy |
|
|
27
|
+
| --- | --- | --- | --- | --- |
|
|
28
|
+
| DeepSeek V4.1 Flash | Passive | No | No | Preserve stable harness + observe |
|
|
29
|
+
| GPT-5.6 Luna | Active cache control | No | Yes | Stable cache key + cache options |
|
|
30
|
+
| GLM-5.3 Flash | Input-shape optimization | Yes, narrowly | No provider cache key | Isolate volatile system content |
|
|
31
|
+
|
|
32
|
+
The plugin is intentionally **not** a generic prompt-rewriter. Each provider receives only the behavior justified by its cache model.
|
|
33
|
+
|
|
34
|
+
---
|
|
35
|
+
|
|
36
|
+
# What the plugin does
|
|
37
|
+
|
|
38
|
+
The server target operates at the OpenCode harness level rather than implementing a provider-specific client.
|
|
39
|
+
|
|
40
|
+
It:
|
|
41
|
+
|
|
42
|
+
1. Detects the model/provider family in use.
|
|
43
|
+
2. Applies only the policy appropriate for that family.
|
|
44
|
+
3. Observes system-prompt and tool-definition stability.
|
|
45
|
+
4. Records provider-reported cache-token usage.
|
|
46
|
+
5. Adds a deterministic compaction continuation block.
|
|
47
|
+
6. Applies GPT-5.6 cache-control metadata.
|
|
48
|
+
7. Applies the GLM-5.3 volatile-environment relocation.
|
|
49
|
+
8. Records diagnostics that help correlate request-shape changes with provider cache behavior.
|
|
50
|
+
|
|
51
|
+
Telemetry is best-effort and must never become a dependency of model execution.
|
|
52
|
+
|
|
53
|
+
---
|
|
54
|
+
|
|
55
|
+
# DeepSeek V4.1 Flash
|
|
56
|
+
|
|
57
|
+
## Policy: passive
|
|
58
|
+
|
|
59
|
+
DeepSeek receives no cache-specific request mutation.
|
|
60
|
+
|
|
61
|
+
The plugin does not:
|
|
62
|
+
|
|
63
|
+
- rewrite the system prompt
|
|
64
|
+
- reorder tools
|
|
65
|
+
- modify messages
|
|
66
|
+
- inject cache-control fields
|
|
67
|
+
- inject a prompt-cache key
|
|
68
|
+
- alter provider request options
|
|
69
|
+
|
|
70
|
+
The DeepSeek branch exists primarily to preserve a stable harness while providing observability around prefix structure and actual cache usage.
|
|
71
|
+
|
|
72
|
+
The plugin observes:
|
|
73
|
+
|
|
74
|
+
- system-prompt shape
|
|
75
|
+
- semantic tool definitions
|
|
76
|
+
- wire-order tool definitions
|
|
77
|
+
- prefix changes
|
|
78
|
+
- cache read tokens
|
|
79
|
+
- cache write tokens
|
|
80
|
+
- compaction boundaries
|
|
81
|
+
|
|
82
|
+
The rule is deliberately simple:
|
|
83
|
+
|
|
84
|
+
```text
|
|
85
|
+
DeepSeek:
|
|
86
|
+
preserve request
|
|
87
|
+
preserve prefix
|
|
88
|
+
measure cache
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
rather than attempting to guess or force the provider's cache behavior.
|
|
92
|
+
|
|
93
|
+
---
|
|
94
|
+
|
|
95
|
+
# GPT-5.6 Luna
|
|
96
|
+
|
|
97
|
+
## Policy: active cache control
|
|
98
|
+
|
|
99
|
+
GPT-5.6 is the current policy that actively injects cache-control request metadata.
|
|
100
|
+
|
|
101
|
+
The plugin adds, when the corresponding fields are not already present:
|
|
102
|
+
|
|
103
|
+
```json
|
|
104
|
+
{
|
|
105
|
+
"promptCacheKey": "<stable-session-key>",
|
|
106
|
+
"promptCacheOptions": {
|
|
107
|
+
"mode": "implicit",
|
|
108
|
+
"ttl": "30m"
|
|
109
|
+
}
|
|
110
|
+
}
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
The key is derived from stable OpenCode session identity and is independent of transient request data.
|
|
114
|
+
|
|
115
|
+
Existing provider-supplied cache options are not overwritten.
|
|
116
|
+
|
|
117
|
+
## The prompt text is not rewritten
|
|
118
|
+
|
|
119
|
+
For GPT-5.6:
|
|
120
|
+
|
|
121
|
+
```text
|
|
122
|
+
system prompt -> unchanged
|
|
123
|
+
conversation -> unchanged
|
|
124
|
+
tool definitions -> unchanged
|
|
125
|
+
|
|
126
|
+
request metadata -> cache key/options added
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
This controls cache behavior without performing prompt surgery.
|
|
130
|
+
|
|
131
|
+
## Default GPT configuration
|
|
132
|
+
|
|
133
|
+
```json
|
|
134
|
+
{
|
|
135
|
+
"promptCacheKey": true,
|
|
136
|
+
"cacheRootKey": false,
|
|
137
|
+
"compactionCacheIsolation": true,
|
|
138
|
+
"reasoningEffortDiagnostics": true,
|
|
139
|
+
"mode": "implicit",
|
|
140
|
+
"ttl": "30m"
|
|
141
|
+
}
|
|
142
|
+
```
|
|
143
|
+
|
|
144
|
+
`cacheRootKey` is disabled by default because the current OpenCode runtime does not expose sufficiently reliable fork lineage for safe cross-fork cache-root inheritance.
|
|
145
|
+
|
|
146
|
+
## Compaction isolation
|
|
147
|
+
|
|
148
|
+
Compaction uses a deterministic separate cache-key namespace:
|
|
149
|
+
|
|
150
|
+
```text
|
|
151
|
+
live session:
|
|
152
|
+
ses_abc123
|
|
153
|
+
|
|
154
|
+
compaction:
|
|
155
|
+
ses_abc123:compact
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
This prevents compaction-specific cache writes from sharing the normal live-session namespace.
|
|
159
|
+
|
|
160
|
+
---
|
|
161
|
+
|
|
162
|
+
# GLM-5.3 Flash
|
|
163
|
+
|
|
164
|
+
## Policy: input-shape optimization
|
|
165
|
+
|
|
166
|
+
GLM-5.3 receives the only current prompt-text transformation.
|
|
167
|
+
|
|
168
|
+
The plugin identifies OpenCode's volatile `<env>` section and moves it to the **tail of the system prompt**.
|
|
169
|
+
|
|
170
|
+
Conceptually:
|
|
171
|
+
|
|
172
|
+
```text
|
|
173
|
+
BEFORE
|
|
174
|
+
|
|
175
|
+
[large stable instructions]
|
|
176
|
+
[volatile environment/date block]
|
|
177
|
+
[more stable instructions]
|
|
178
|
+
```
|
|
179
|
+
|
|
180
|
+
becomes:
|
|
181
|
+
|
|
182
|
+
```text
|
|
183
|
+
AFTER
|
|
184
|
+
|
|
185
|
+
[large stable instructions]
|
|
186
|
+
[more stable instructions]
|
|
187
|
+
[volatile environment/date block]
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
The environment block's contents are preserved. Only its position changes.
|
|
191
|
+
|
|
192
|
+
The objective is to isolate volatile information so that a changing date or environment value does not unnecessarily disturb the earlier stable prefix.
|
|
193
|
+
|
|
194
|
+
## GLM safety constraints
|
|
195
|
+
|
|
196
|
+
The transformation occurs only when:
|
|
197
|
+
|
|
198
|
+
- the selected model is GLM-5.3;
|
|
199
|
+
- GLM stabilization is enabled;
|
|
200
|
+
- there is exactly one system string;
|
|
201
|
+
- the expected environment markers exist; and
|
|
202
|
+
- the block can be identified unambiguously.
|
|
203
|
+
|
|
204
|
+
The plugin does not arbitrarily reorder unrelated system-prompt content.
|
|
205
|
+
|
|
206
|
+
---
|
|
207
|
+
|
|
208
|
+
# System-prompt diagnostics
|
|
209
|
+
|
|
210
|
+
The plugin fingerprints the system prompt to detect structural changes between requests.
|
|
211
|
+
|
|
212
|
+
Provider-aware diagnostics track:
|
|
213
|
+
|
|
214
|
+
- full system hash
|
|
215
|
+
- stable system-prefix hash
|
|
216
|
+
- volatile system-suffix hash
|
|
217
|
+
|
|
218
|
+
The stable/volatile decomposition is based on the longest common prefix against the session baseline.
|
|
219
|
+
|
|
220
|
+
A changed hash means:
|
|
221
|
+
|
|
222
|
+
```text
|
|
223
|
+
The observed request bytes changed.
|
|
224
|
+
```
|
|
225
|
+
|
|
226
|
+
It does **not** mean:
|
|
227
|
+
|
|
228
|
+
```text
|
|
229
|
+
The provider definitely generated a cache miss.
|
|
230
|
+
```
|
|
231
|
+
|
|
232
|
+
Provider-reported cache-token usage remains authoritative.
|
|
233
|
+
|
|
234
|
+
---
|
|
235
|
+
|
|
236
|
+
# Tool-definition diagnostics
|
|
237
|
+
|
|
238
|
+
Tool definitions are normalized before fingerprinting.
|
|
239
|
+
|
|
240
|
+
Runtime-only fields such as object identity, function references, timestamps, and arbitrary runtime metadata are excluded.
|
|
241
|
+
|
|
242
|
+
The plugin maintains two useful fingerprints:
|
|
243
|
+
|
|
244
|
+
- **semantic fingerprint** — order-insensitive representation of model-visible tool definitions;
|
|
245
|
+
- **wire-order fingerprint** — order-sensitive diagnostic representation of the closest deterministic pre-wire tool ordering available to the plugin.
|
|
246
|
+
|
|
247
|
+
These fingerprints are diagnostic only. The plugin does not reorder tools merely to force a particular fingerprint.
|
|
248
|
+
|
|
249
|
+
---
|
|
250
|
+
|
|
251
|
+
# Compaction handling
|
|
252
|
+
|
|
253
|
+
OpenCode sessions eventually undergo compaction as conversation history grows.
|
|
254
|
+
|
|
255
|
+
The plugin adds a deterministic continuation template:
|
|
256
|
+
|
|
257
|
+
```text
|
|
258
|
+
## Session digest (cache-stable continuation block)
|
|
259
|
+
- Goal:
|
|
260
|
+
- Decisions made:
|
|
261
|
+
- Pending:
|
|
262
|
+
- Active files:
|
|
263
|
+
```
|
|
264
|
+
|
|
265
|
+
The digest is inserted once per compaction invocation using a guard that prevents duplicate insertion if the hook fires more than once.
|
|
266
|
+
|
|
267
|
+
The objective is to provide a deterministic continuation structure rather than generating a different arbitrary cache-affecting block on every compaction.
|
|
268
|
+
|
|
269
|
+
---
|
|
270
|
+
|
|
271
|
+
# Cache metrics
|
|
272
|
+
|
|
273
|
+
The plugin records cache usage from OpenCode assistant-message token data.
|
|
274
|
+
|
|
275
|
+
At minimum it tracks:
|
|
276
|
+
|
|
277
|
+
```text
|
|
278
|
+
cache.read
|
|
279
|
+
cache.write
|
|
280
|
+
```
|
|
281
|
+
|
|
282
|
+
and aggregates those values across the session.
|
|
283
|
+
|
|
284
|
+
## Generic cache ratio
|
|
285
|
+
|
|
286
|
+
The core helper reports:
|
|
287
|
+
|
|
288
|
+
```text
|
|
289
|
+
hit rate = read / (read + write)
|
|
290
|
+
```
|
|
291
|
+
|
|
292
|
+
This is an accounting metric based on cache read/write tokens.
|
|
293
|
+
|
|
294
|
+
## GLM prompt-token ratio
|
|
295
|
+
|
|
296
|
+
For GLM, the implementation additionally calculates:
|
|
297
|
+
|
|
298
|
+
```text
|
|
299
|
+
read / (read + write + input)
|
|
300
|
+
```
|
|
301
|
+
|
|
302
|
+
These ratios answer different questions and should not be treated as interchangeable.
|
|
303
|
+
|
|
304
|
+
---
|
|
305
|
+
|
|
306
|
+
# Telemetry
|
|
307
|
+
|
|
308
|
+
Metrics are written as JSONL.
|
|
309
|
+
|
|
310
|
+
Default metrics path:
|
|
311
|
+
|
|
312
|
+
```text
|
|
313
|
+
~/.cache/opencode/cache-metrics.jsonl
|
|
314
|
+
```
|
|
315
|
+
|
|
316
|
+
Default plugin configuration path:
|
|
317
|
+
|
|
318
|
+
```text
|
|
319
|
+
~/.config/opencode/cache-engine.json
|
|
320
|
+
```
|
|
321
|
+
|
|
322
|
+
Telemetry is best-effort. A failed metrics write must never break an OpenCode request.
|
|
323
|
+
|
|
324
|
+
## Example usage record
|
|
325
|
+
|
|
326
|
+
```json
|
|
327
|
+
{
|
|
328
|
+
"kind": "usage-event",
|
|
329
|
+
"sid": "session-id",
|
|
330
|
+
"ts": 1750000000000,
|
|
331
|
+
"read": 120000,
|
|
332
|
+
"write": 3000,
|
|
333
|
+
"cost": 0.0123,
|
|
334
|
+
"provider": "z-ai",
|
|
335
|
+
"model": "glm-5.3-flash",
|
|
336
|
+
"policy": "glm53"
|
|
337
|
+
}
|
|
338
|
+
```
|
|
339
|
+
|
|
340
|
+
## Example prefix-change record
|
|
341
|
+
|
|
342
|
+
```json
|
|
343
|
+
{
|
|
344
|
+
"kind": "prefix-change",
|
|
345
|
+
"sid": "session-id",
|
|
346
|
+
"ts": 1750000000000,
|
|
347
|
+
"dimensions": [
|
|
348
|
+
"system"
|
|
349
|
+
]
|
|
350
|
+
}
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
## Example compaction record
|
|
354
|
+
|
|
355
|
+
```json
|
|
356
|
+
{
|
|
357
|
+
"kind": "compaction",
|
|
358
|
+
"sid": "session-id",
|
|
359
|
+
"ts": 1750000000000,
|
|
360
|
+
"reason": "compaction",
|
|
361
|
+
"usageSamples": 7,
|
|
362
|
+
"cumulative": {
|
|
363
|
+
"read": 900000,
|
|
364
|
+
"write": 12000
|
|
365
|
+
}
|
|
366
|
+
}
|
|
367
|
+
```
|
|
368
|
+
|
|
369
|
+
Telemetry is intended to answer questions such as:
|
|
370
|
+
|
|
371
|
+
- Did the system prompt change?
|
|
372
|
+
- Did the tool definitions change?
|
|
373
|
+
- Did cache reads increase?
|
|
374
|
+
- Did cache writes increase?
|
|
375
|
+
- Did a compaction occur?
|
|
376
|
+
- Which provider/model/policy was active?
|
|
377
|
+
- Did provider-specific system stabilization change the observed prompt shape?
|
|
378
|
+
|
|
379
|
+
---
|
|
380
|
+
|
|
381
|
+
# Configuration
|
|
382
|
+
|
|
383
|
+
A representative default configuration is:
|
|
384
|
+
|
|
385
|
+
```json
|
|
386
|
+
{
|
|
387
|
+
"enabled": true,
|
|
388
|
+
"metricsFile": "~/.cache/opencode/cache-metrics.jsonl",
|
|
389
|
+
"compactTemplate": true,
|
|
390
|
+
"logPrefixChanges": true,
|
|
391
|
+
"policies": {
|
|
392
|
+
"deepseek": {
|
|
393
|
+
"enabled": true
|
|
394
|
+
},
|
|
395
|
+
"gpt56": {
|
|
396
|
+
"enabled": true,
|
|
397
|
+
"promptCacheKey": true,
|
|
398
|
+
"cacheRootKey": false,
|
|
399
|
+
"compactionCacheIsolation": true,
|
|
400
|
+
"reasoningEffortDiagnostics": true,
|
|
401
|
+
"mode": "implicit",
|
|
402
|
+
"ttl": "30m"
|
|
403
|
+
},
|
|
404
|
+
"glm53": {
|
|
405
|
+
"enabled": true,
|
|
406
|
+
"stabilizeSystem": true,
|
|
407
|
+
"preserveThinkingIntegrity": true
|
|
408
|
+
}
|
|
409
|
+
}
|
|
410
|
+
}
|
|
411
|
+
```
|
|
412
|
+
|
|
413
|
+
The configuration parser starts from defaults and applies valid file/environment overrides without mutating the caller's object.
|
|
414
|
+
|
|
415
|
+
---
|
|
416
|
+
|
|
417
|
+
# Configuration options
|
|
418
|
+
|
|
419
|
+
## Global
|
|
420
|
+
|
|
421
|
+
### `enabled`
|
|
422
|
+
|
|
423
|
+
Enables or disables the entire server plugin.
|
|
424
|
+
|
|
425
|
+
### `metricsFile`
|
|
426
|
+
|
|
427
|
+
Controls where JSONL telemetry is written.
|
|
428
|
+
|
|
429
|
+
### `compactTemplate`
|
|
430
|
+
|
|
431
|
+
Controls whether the deterministic compaction continuation block is inserted.
|
|
432
|
+
|
|
433
|
+
### `logPrefixChanges`
|
|
434
|
+
|
|
435
|
+
Controls warning logs for observed prefix-shape changes.
|
|
436
|
+
|
|
437
|
+
## DeepSeek
|
|
438
|
+
|
|
439
|
+
```json
|
|
440
|
+
"deepseek": {
|
|
441
|
+
"enabled": true
|
|
442
|
+
}
|
|
443
|
+
```
|
|
444
|
+
|
|
445
|
+
There are intentionally few DeepSeek settings because the policy is passive.
|
|
446
|
+
|
|
447
|
+
## GPT-5.6
|
|
448
|
+
|
|
449
|
+
```json
|
|
450
|
+
"gpt56": {
|
|
451
|
+
"enabled": true,
|
|
452
|
+
"promptCacheKey": true,
|
|
453
|
+
"cacheRootKey": false,
|
|
454
|
+
"compactionCacheIsolation": true,
|
|
455
|
+
"reasoningEffortDiagnostics": true,
|
|
456
|
+
"mode": "implicit",
|
|
457
|
+
"ttl": "30m"
|
|
458
|
+
}
|
|
459
|
+
```
|
|
460
|
+
|
|
461
|
+
### `promptCacheKey`
|
|
462
|
+
|
|
463
|
+
Controls whether the plugin provides a stable session-derived GPT cache key.
|
|
464
|
+
|
|
465
|
+
### `cacheRootKey`
|
|
466
|
+
|
|
467
|
+
Controls parent/fork cache-root inheritance. Disabled by default because reliable fork lineage is not currently guaranteed by the runtime.
|
|
468
|
+
|
|
469
|
+
### `compactionCacheIsolation`
|
|
470
|
+
|
|
471
|
+
Uses a separate deterministic cache namespace for compaction requests.
|
|
472
|
+
|
|
473
|
+
### `reasoningEffortDiagnostics`
|
|
474
|
+
|
|
475
|
+
Tracks GPT reasoning-effort changes for diagnostics.
|
|
476
|
+
|
|
477
|
+
### `mode`
|
|
478
|
+
|
|
479
|
+
Defaults to `implicit`.
|
|
480
|
+
|
|
481
|
+
### `ttl`
|
|
482
|
+
|
|
483
|
+
Defaults to `30m`.
|
|
484
|
+
|
|
485
|
+
Existing request options are not overwritten.
|
|
486
|
+
|
|
487
|
+
## GLM-5.3
|
|
488
|
+
|
|
489
|
+
```json
|
|
490
|
+
"glm53": {
|
|
491
|
+
"enabled": true,
|
|
492
|
+
"stabilizeSystem": true,
|
|
493
|
+
"preserveThinkingIntegrity": true
|
|
494
|
+
}
|
|
495
|
+
```
|
|
496
|
+
|
|
497
|
+
### `stabilizeSystem`
|
|
498
|
+
|
|
499
|
+
Enables relocation of the volatile `<env>` section to the system-prompt tail.
|
|
500
|
+
|
|
501
|
+
### `preserveThinkingIntegrity`
|
|
502
|
+
|
|
503
|
+
Enables diagnostic checks around reasoning continuity.
|
|
504
|
+
|
|
505
|
+
The reasoning instrumentation is diagnostic only. It does not rewrite, fabricate, or reorder reasoning content.
|
|
506
|
+
|
|
507
|
+
---
|
|
508
|
+
|
|
509
|
+
# Model detection
|
|
510
|
+
|
|
511
|
+
The server plugin classifies requests into:
|
|
512
|
+
|
|
513
|
+
```text
|
|
514
|
+
deepseek
|
|
515
|
+
gpt56
|
|
516
|
+
glm53
|
|
517
|
+
neutral
|
|
518
|
+
```
|
|
519
|
+
|
|
520
|
+
The detector recognizes provider/model identifiers for the supported policy families.
|
|
521
|
+
|
|
522
|
+
GPT-5.6 has an additional OpenAI/Azure-context check so a string containing `gpt-5.6` does not automatically cause GPT-specific fields to be sent to an unrelated endpoint.
|
|
523
|
+
|
|
524
|
+
Unknown models use the neutral policy:
|
|
525
|
+
|
|
526
|
+
```text
|
|
527
|
+
no provider-specific request mutation
|
|
528
|
+
```
|
|
529
|
+
|
|
530
|
+
---
|
|
531
|
+
|
|
532
|
+
# OpenRouter usage
|
|
533
|
+
|
|
534
|
+
The plugin is compatible with OpenRouter because policy selection is based on the provider/model information exposed by OpenCode.
|
|
535
|
+
|
|
536
|
+
For cache-sensitive workloads, provider stability is important.
|
|
537
|
+
|
|
538
|
+
The plugin does not attempt to compensate for provider switching by rewriting prompts. Stable provider routing is preferable when the objective is to preserve reusable prefixes.
|
|
539
|
+
|
|
540
|
+
---
|
|
541
|
+
|
|
542
|
+
# Package architecture
|
|
543
|
+
|
|
544
|
+
`opencode-cache-engine` is an npm package with separate server and TUI targets.
|
|
545
|
+
|
|
546
|
+
```text
|
|
547
|
+
opencode-cache-engine/
|
|
548
|
+
├── src/
|
|
549
|
+
│ ├── cache-engine.ts
|
|
550
|
+
│ ├── cache-engine-core.mjs
|
|
551
|
+
│ └── tui.mjs
|
|
552
|
+
├── test/
|
|
553
|
+
│ └── cache-engine.test.mjs
|
|
554
|
+
├── examples/
|
|
555
|
+
│ └── cache-engine.json
|
|
556
|
+
├── package.json
|
|
557
|
+
├── README.md
|
|
558
|
+
└── LICENSE
|
|
559
|
+
```
|
|
560
|
+
|
|
561
|
+
The package exports two OpenCode targets:
|
|
562
|
+
|
|
563
|
+
```json
|
|
564
|
+
{
|
|
565
|
+
"exports": {
|
|
566
|
+
"./server": "./src/cache-engine.ts",
|
|
567
|
+
"./tui": "./src/tui.mjs"
|
|
568
|
+
}
|
|
569
|
+
}
|
|
570
|
+
```
|
|
571
|
+
|
|
572
|
+
## Server target
|
|
573
|
+
|
|
574
|
+
`./server` points to the actual CacheEngine implementation.
|
|
575
|
+
|
|
576
|
+
It owns:
|
|
577
|
+
|
|
578
|
+
- OpenCode server hooks
|
|
579
|
+
- session state
|
|
580
|
+
- provider-policy selection
|
|
581
|
+
- request mutation
|
|
582
|
+
- system-prompt transformation
|
|
583
|
+
- telemetry
|
|
584
|
+
- compaction handling
|
|
585
|
+
|
|
586
|
+
The implementation exports:
|
|
587
|
+
|
|
588
|
+
```ts
|
|
589
|
+
export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
590
|
+
// ...
|
|
591
|
+
}
|
|
592
|
+
```
|
|
593
|
+
|
|
594
|
+
The identifier `CacheEngine` is the code export name. It does not determine the npm package name.
|
|
595
|
+
|
|
596
|
+
## TUI target
|
|
597
|
+
|
|
598
|
+
`./tui` is a small TUI-only registration module.
|
|
599
|
+
|
|
600
|
+
Its purpose is to make the npm package discoverable by OpenCode's TUI plugin manager.
|
|
601
|
+
|
|
602
|
+
The TUI target does not duplicate or contain cache-engine logic.
|
|
603
|
+
|
|
604
|
+
TUI activation state and server-runtime activation are separate concepts unless an explicit state bridge is implemented.
|
|
605
|
+
|
|
606
|
+
---
|
|
607
|
+
|
|
608
|
+
# Installation
|
|
609
|
+
|
|
610
|
+
## npm package
|
|
611
|
+
|
|
612
|
+
The package name is:
|
|
613
|
+
|
|
614
|
+
```text
|
|
615
|
+
opencode-cache-engine
|
|
616
|
+
```
|
|
617
|
+
|
|
618
|
+
For the normal runtime/release configuration, add the package to the OpenCode server plugin configuration:
|
|
619
|
+
|
|
620
|
+
```jsonc
|
|
621
|
+
{
|
|
622
|
+
"plugin": [
|
|
623
|
+
"opencode-cache-engine"
|
|
624
|
+
]
|
|
625
|
+
}
|
|
626
|
+
```
|
|
627
|
+
|
|
628
|
+
Add the same package to the TUI plugin registry when you want it displayed and managed by the OpenCode TUI plugin manager:
|
|
629
|
+
|
|
630
|
+
```json
|
|
631
|
+
{
|
|
632
|
+
"plugin": [
|
|
633
|
+
"opencode-cache-engine"
|
|
634
|
+
],
|
|
635
|
+
"$schema": "https://www.opencode.ai/tui.json"
|
|
636
|
+
}
|
|
637
|
+
```
|
|
638
|
+
|
|
639
|
+
The package must expose both server and TUI targets for the TUI registry to accept it.
|
|
640
|
+
|
|
641
|
+
## No direct local plugin copy
|
|
642
|
+
|
|
643
|
+
Do not install a second CacheEngine copy by placing `cache-engine.ts` or related source files under:
|
|
644
|
+
|
|
645
|
+
```text
|
|
646
|
+
~/.config/opencode/plugins/
|
|
647
|
+
```
|
|
648
|
+
|
|
649
|
+
when using the configured npm package.
|
|
650
|
+
|
|
651
|
+
A duplicate local copy can cause multiple plugin instances to be loaded from different sources.
|
|
652
|
+
|
|
653
|
+
---
|
|
654
|
+
|
|
655
|
+
# Development from Git
|
|
656
|
+
|
|
657
|
+
The Git repository is the development source of truth.
|
|
658
|
+
|
|
659
|
+
During active development, work directly in the checkout:
|
|
660
|
+
|
|
661
|
+
```text
|
|
662
|
+
~/Desktop/opencode-cache-engine/
|
|
663
|
+
```
|
|
664
|
+
|
|
665
|
+
The development workflow is:
|
|
666
|
+
|
|
667
|
+
```text
|
|
668
|
+
Git working tree
|
|
669
|
+
↓
|
|
670
|
+
unit tests
|
|
671
|
+
↓
|
|
672
|
+
OpenCode local development target
|
|
673
|
+
↓
|
|
674
|
+
runtime validation
|
|
675
|
+
↓
|
|
676
|
+
git commit
|
|
677
|
+
↓
|
|
678
|
+
npm release
|
|
679
|
+
```
|
|
680
|
+
|
|
681
|
+
For local development, the server target can be loaded directly from the Git checkout through OpenCode's local plugin configuration rather than publishing an npm version for every change.
|
|
682
|
+
|
|
683
|
+
The repository remains the only place where source code is edited.
|
|
684
|
+
|
|
685
|
+
---
|
|
686
|
+
|
|
687
|
+
# Testing
|
|
688
|
+
|
|
689
|
+
## Unit tests
|
|
690
|
+
|
|
691
|
+
Run the test suite from the repository root:
|
|
692
|
+
|
|
693
|
+
```bash
|
|
694
|
+
node --test test/cache-engine.test.mjs
|
|
695
|
+
```
|
|
696
|
+
|
|
697
|
+
The test file imports the pure core module from:
|
|
698
|
+
|
|
699
|
+
```js
|
|
700
|
+
from "../src/cache-engine-core.mjs"
|
|
701
|
+
```
|
|
702
|
+
|
|
703
|
+
The tests validate provider-independent and provider-specific logic, including:
|
|
704
|
+
|
|
705
|
+
- model detection
|
|
706
|
+
- GPT cache-key stability
|
|
707
|
+
- GPT cache-option defaults
|
|
708
|
+
- protection against overwriting existing cache options
|
|
709
|
+
- GLM environment relocation
|
|
710
|
+
- deterministic hashing
|
|
711
|
+
- system-prefix decomposition
|
|
712
|
+
- tool fingerprints
|
|
713
|
+
- reasoning diagnostics
|
|
714
|
+
- compaction isolation
|
|
715
|
+
- cache-hit calculations
|
|
716
|
+
- configuration behavior
|
|
717
|
+
- JSONL telemetry behavior
|
|
718
|
+
|
|
719
|
+
Runtime behavior is validated separately through actual OpenCode plugin loading.
|
|
720
|
+
|
|
721
|
+
## Package validation
|
|
722
|
+
|
|
723
|
+
Before publishing, inspect the package contents:
|
|
724
|
+
|
|
725
|
+
```bash
|
|
726
|
+
npm pack --dry-run
|
|
727
|
+
```
|
|
728
|
+
|
|
729
|
+
The package should contain the runtime source and documentation needed by OpenCode, including:
|
|
730
|
+
|
|
731
|
+
```text
|
|
732
|
+
src/cache-engine.ts
|
|
733
|
+
src/cache-engine-core.mjs
|
|
734
|
+
src/tui.mjs
|
|
735
|
+
package.json
|
|
736
|
+
README.md
|
|
737
|
+
LICENSE
|
|
738
|
+
examples/...
|
|
739
|
+
```
|
|
740
|
+
|
|
741
|
+
Test the TUI target directly with Node when appropriate:
|
|
742
|
+
|
|
743
|
+
```bash
|
|
744
|
+
node -e 'import("./src/tui.mjs").then(m => console.log(m.default))'
|
|
745
|
+
```
|
|
746
|
+
|
|
747
|
+
---
|
|
748
|
+
|
|
749
|
+
# Design principles
|
|
750
|
+
|
|
751
|
+
## 1. Provider-specific behavior
|
|
752
|
+
|
|
753
|
+
Different providers expose different cache mechanisms.
|
|
754
|
+
|
|
755
|
+
The plugin therefore does not assume that one strategy is optimal everywhere.
|
|
756
|
+
|
|
757
|
+
## 2. Preserve working behavior
|
|
758
|
+
|
|
759
|
+
The plugin should not modify a provider's request merely because a mutation is technically possible.
|
|
760
|
+
|
|
761
|
+
This is especially important for DeepSeek, where the current policy is intentionally passive.
|
|
762
|
+
|
|
763
|
+
## 3. Measure provider reality
|
|
764
|
+
|
|
765
|
+
Local hashes are diagnostics.
|
|
766
|
+
|
|
767
|
+
Provider-reported cache-token counts are authoritative.
|
|
768
|
+
|
|
769
|
+
The implementation explicitly distinguishes:
|
|
770
|
+
|
|
771
|
+
```text
|
|
772
|
+
observed prefix change
|
|
773
|
+
```
|
|
774
|
+
|
|
775
|
+
from:
|
|
776
|
+
|
|
777
|
+
```text
|
|
778
|
+
confirmed provider cache miss
|
|
779
|
+
```
|
|
780
|
+
|
|
781
|
+
because the plugin cannot infer the latter reliably from local hashes alone.
|
|
782
|
+
|
|
783
|
+
## 4. Never overwrite explicit provider configuration
|
|
784
|
+
|
|
785
|
+
Where GPT cache options already exist, the plugin leaves them alone.
|
|
786
|
+
|
|
787
|
+
## 5. Keep mutations deterministic
|
|
788
|
+
|
|
789
|
+
When the plugin transforms a request, the transformation should be:
|
|
790
|
+
|
|
791
|
+
- narrow
|
|
792
|
+
- deterministic
|
|
793
|
+
- content-preserving where possible
|
|
794
|
+
- provider-specific
|
|
795
|
+
- easy to disable
|
|
796
|
+
|
|
797
|
+
## 6. Keep telemetry out of the critical path
|
|
798
|
+
|
|
799
|
+
A metrics failure must not break model execution.
|
|
800
|
+
|
|
801
|
+
---
|
|
802
|
+
|
|
803
|
+
# What the plugin does NOT do
|
|
804
|
+
|
|
805
|
+
The plugin does not:
|
|
806
|
+
|
|
807
|
+
- invent cache hits;
|
|
808
|
+
- claim a local hash proves a provider cache hit;
|
|
809
|
+
- rewrite DeepSeek prompts;
|
|
810
|
+
- reorder tools merely to force cache reuse;
|
|
811
|
+
- fabricate reasoning;
|
|
812
|
+
- modify conversation history arbitrarily;
|
|
813
|
+
- force explicit GPT cache breakpoints by default;
|
|
814
|
+
- silently overwrite existing GPT cache options;
|
|
815
|
+
- assume every model named `gpt-5.6` is an OpenAI-compatible endpoint; or
|
|
816
|
+
- use fork inheritance unless reliable lineage is available.
|
|
817
|
+
|
|
818
|
+
---
|
|
819
|
+
|
|
820
|
+
# Cost optimization philosophy
|
|
821
|
+
|
|
822
|
+
Cache hit rate is useful, but it is not the only cost metric.
|
|
823
|
+
|
|
824
|
+
The economic objective is:
|
|
825
|
+
|
|
826
|
+
```text
|
|
827
|
+
total task cost
|
|
828
|
+
=
|
|
829
|
+
prompt/cache cost
|
|
830
|
+
+
|
|
831
|
+
output/reasoning cost
|
|
832
|
+
+
|
|
833
|
+
additional requests
|
|
834
|
+
```
|
|
835
|
+
|
|
836
|
+
A model with a slightly lower cache hit rate can still be cheaper if it completes the task with fewer tokens or fewer model calls.
|
|
837
|
+
|
|
838
|
+
For that reason, this plugin is primarily an **instrumentation + targeted optimization layer**, not a cache-rate maximizer at any cost.
|
|
839
|
+
|
|
840
|
+
The preferred evaluation unit is:
|
|
841
|
+
|
|
842
|
+
```text
|
|
843
|
+
cost per completed task
|
|
844
|
+
```
|
|
845
|
+
|
|
846
|
+
rather than cache percentage alone.
|
|
847
|
+
|
|
848
|
+
---
|
|
849
|
+
|
|
850
|
+
# Operational recommendations
|
|
851
|
+
|
|
852
|
+
For reliable cache measurements:
|
|
853
|
+
|
|
854
|
+
1. Keep the provider fixed whenever possible.
|
|
855
|
+
2. Avoid changing unrelated system-prompt content during a benchmark.
|
|
856
|
+
3. Keep tool definitions stable.
|
|
857
|
+
4. Compare equivalent tasks across models.
|
|
858
|
+
5. Record actual provider cache-token counts.
|
|
859
|
+
6. Compare total task cost, not only cache percentage.
|
|
860
|
+
7. Treat compaction as a separate cache boundary when analyzing results.
|
|
861
|
+
8. Avoid interpreting a local prefix hash change as definitive proof of a cache miss.
|
|
862
|
+
|
|
863
|
+
---
|
|
864
|
+
|
|
865
|
+
# Troubleshooting
|
|
866
|
+
|
|
867
|
+
## `there is no TUI target`
|
|
868
|
+
|
|
869
|
+
This means the npm package does not expose a TUI target that your installed OpenCode version recognizes.
|
|
870
|
+
|
|
871
|
+
Verify that the package contains a TUI entrypoint and that `package.json` exposes the required `./tui` target.
|
|
872
|
+
|
|
873
|
+
After changing the package, update/reinstall the npm package and restart OpenCode.
|
|
874
|
+
|
|
875
|
+
## The old `oc-plugin-caching` still appears
|
|
876
|
+
|
|
877
|
+
Search all OpenCode configuration and local-plugin locations for the old plugin name:
|
|
878
|
+
|
|
879
|
+
```bash
|
|
880
|
+
grep -R "oc-plugin-caching" ~/.config/opencode ~/.opencode .opencode 2>/dev/null
|
|
881
|
+
```
|
|
882
|
+
|
|
883
|
+
Remove the old configuration entry and remove the obsolete local plugin source copy.
|
|
884
|
+
|
|
885
|
+
## DeepSeek cache rate dropped
|
|
886
|
+
|
|
887
|
+
First check provider stability and whether OpenCode's system/tool prefix changed.
|
|
888
|
+
|
|
889
|
+
The plugin does not intentionally mutate DeepSeek request options.
|
|
890
|
+
|
|
891
|
+
Inspect telemetry for:
|
|
892
|
+
|
|
893
|
+
```text
|
|
894
|
+
prefix-change
|
|
895
|
+
usage-event
|
|
896
|
+
compaction
|
|
897
|
+
```
|
|
898
|
+
|
|
899
|
+
A prefix change is a diagnostic signal, not automatic proof of a cache miss.
|
|
900
|
+
|
|
901
|
+
## GPT-5.6 cache options are missing
|
|
902
|
+
|
|
903
|
+
Verify that the model is classified as GPT-5.6 and that the endpoint is recognized as OpenAI/Azure-compatible.
|
|
904
|
+
|
|
905
|
+
Also check whether the outgoing request already supplied its own cache options. Existing settings are intentionally preserved.
|
|
906
|
+
|
|
907
|
+
## GLM-5.3 prompt is not being changed
|
|
908
|
+
|
|
909
|
+
The environment relocation only occurs when the plugin can identify the expected block unambiguously.
|
|
910
|
+
|
|
911
|
+
The relevant block must contain the expected beginning and closing marker, and the system structure must meet the plugin's eligibility rules.
|
|
912
|
+
|
|
913
|
+
## Metrics file is missing
|
|
914
|
+
|
|
915
|
+
Telemetry is best-effort.
|
|
916
|
+
|
|
917
|
+
Check:
|
|
918
|
+
|
|
919
|
+
```text
|
|
920
|
+
~/.cache/opencode/cache-metrics.jsonl
|
|
921
|
+
```
|
|
922
|
+
|
|
923
|
+
and verify that the configured parent directory is writable.
|
|
924
|
+
|
|
925
|
+
---
|
|
926
|
+
|
|
927
|
+
# Release workflow
|
|
928
|
+
|
|
929
|
+
The Git repository is the source of truth for development.
|
|
930
|
+
|
|
931
|
+
A normal release workflow is:
|
|
932
|
+
|
|
933
|
+
```text
|
|
934
|
+
edit source
|
|
935
|
+
↓
|
|
936
|
+
run tests
|
|
937
|
+
↓
|
|
938
|
+
run OpenCode runtime validation
|
|
939
|
+
↓
|
|
940
|
+
git commit
|
|
941
|
+
↓
|
|
942
|
+
bump package version
|
|
943
|
+
↓
|
|
944
|
+
npm publish
|
|
945
|
+
↓
|
|
946
|
+
use published opencode-cache-engine package
|
|
947
|
+
```
|
|
948
|
+
|
|
949
|
+
Do not edit the installed npm cache as a substitute for repository development.
|
|
950
|
+
|
|
951
|
+
---
|
|
952
|
+
|
|
953
|
+
# Status
|
|
954
|
+
|
|
955
|
+
The current implementation is intentionally conservative:
|
|
956
|
+
|
|
957
|
+
```text
|
|
958
|
+
DeepSeek -> preserve and measure
|
|
959
|
+
GPT-5.6 -> configure cache controls
|
|
960
|
+
GLM-5.3 -> isolate volatile prompt content
|
|
961
|
+
```
|
|
962
|
+
|
|
963
|
+
That separation is the core design of the project.
|
|
964
|
+
|
|
965
|
+
The plugin should be evaluated using real provider-reported usage and real task cost rather than assuming that any particular local transformation guarantees a cache hit.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "opencode-cache-engine",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.2.0",
|
|
4
4
|
"private": false,
|
|
5
5
|
"description": "Provider-aware prompt-cache optimization and telemetry for OpenCode",
|
|
6
6
|
"keywords": [
|
|
@@ -27,6 +27,10 @@
|
|
|
27
27
|
"example": "examples",
|
|
28
28
|
"test": "test"
|
|
29
29
|
},
|
|
30
|
+
"exports": {
|
|
31
|
+
"./server": "./src/cache-engine.ts",
|
|
32
|
+
"./tui": "./src/tui.mjs"
|
|
33
|
+
},
|
|
30
34
|
"scripts": {
|
|
31
35
|
"test": "node --test test/cache-engine.test.mjs",
|
|
32
36
|
"postpublish": "VERSION=$(node -p \"require('./package').version\") && TAG=v$VERSION && echo \"Creating git tag $TAG for npm $VERSION\" && git tag $TAG && git push origin $TAG || echo \"Tag $TAG may already exist or push failed; continuing.\"",
|
package/src/tui.mjs
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
const plugin = {
|
|
2
|
+
id: "opencode-cache-engine",
|
|
3
|
+
|
|
4
|
+
async tui() {
|
|
5
|
+
// CacheEngine has no TUI UI of its own.
|
|
6
|
+
// This target exists so the npm package can be registered,
|
|
7
|
+
// displayed, and enabled/disabled by the OpenCode plugin manager.
|
|
8
|
+
},
|
|
9
|
+
}
|
|
10
|
+
|
|
11
|
+
export default plugin
|