opencode-cache-engine 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +1032 -0
- package/examples/cache-engine.json +25 -0
- package/opencode-cache-engine-0.1.0.tgz +0 -0
- package/package.json +33 -0
- package/src/cache-engine-core.mjs +728 -0
- package/src/cache-engine.ts +753 -0
- package/test/cache-engine.test.mjs +765 -0
package/README.md
ADDED
|
@@ -0,0 +1,1032 @@
|
|
|
1
|
+
# OpenCode Cache Engine
|
|
2
|
+
|
|
3
|
+
Provider-aware prompt-cache optimization and observability for [OpenCode](https://opencode.ai).
|
|
4
|
+
|
|
5
|
+
`CacheEngine` is an OpenCode plugin designed for long-running agent sessions where prompt-cache efficiency affects both latency and cost. It keeps the harness conservative for providers whose cache behavior is already automatic, while applying provider-specific optimizations where the provider exposes useful cache controls or where prompt structure can be safely improved.
|
|
6
|
+
|
|
7
|
+
The plugin currently has three cache-policy families:
|
|
8
|
+
|
|
9
|
+
* **DeepSeek V4 Flash** — passive cache-stability and observability
|
|
10
|
+
* **GPT-5.6 Luna** — active cache-control configuration
|
|
11
|
+
* **GLM-5.3 Flash** — conservative system-prompt stabilization
|
|
12
|
+
|
|
13
|
+
The central design principle is:
|
|
14
|
+
|
|
15
|
+
> Optimize the request structure only when there is a clear provider-specific reason to do so. Otherwise, preserve OpenCode's native request behavior and measure what the provider actually reports.
|
|
16
|
+
|
|
17
|
+
---
|
|
18
|
+
|
|
19
|
+
## What this plugin does
|
|
20
|
+
|
|
21
|
+
The plugin operates at the OpenCode harness level rather than implementing a provider-specific client.
|
|
22
|
+
|
|
23
|
+
It:
|
|
24
|
+
|
|
25
|
+
1. Detects the model/provider family in use.
|
|
26
|
+
2. Applies only the policy appropriate for that family.
|
|
27
|
+
3. Observes system-prompt and tool-definition stability.
|
|
28
|
+
4. Records provider-reported cache token usage.
|
|
29
|
+
5. Adds a deterministic compaction continuation block.
|
|
30
|
+
6. Applies GPT-5.6 cache-control metadata.
|
|
31
|
+
7. Applies the GLM-5.3 volatile-environment relocation.
|
|
32
|
+
8. Records diagnostics that help determine whether prompt-shape changes correlate with cache behavior.
|
|
33
|
+
|
|
34
|
+
The plugin deliberately avoids pretending that a local hash is proof of a provider cache hit. Provider-reported token usage remains the authoritative signal.
|
|
35
|
+
|
|
36
|
+
---
|
|
37
|
+
|
|
38
|
+
# Provider behavior
|
|
39
|
+
|
|
40
|
+
## DeepSeek V4 Flash
|
|
41
|
+
|
|
42
|
+
### Policy: passive
|
|
43
|
+
|
|
44
|
+
DeepSeek receives **no cache-specific request mutation**.
|
|
45
|
+
|
|
46
|
+
The plugin does not:
|
|
47
|
+
|
|
48
|
+
* rewrite the system prompt
|
|
49
|
+
* reorder tools
|
|
50
|
+
* modify messages
|
|
51
|
+
* inject cache-control fields
|
|
52
|
+
* inject a prompt-cache key
|
|
53
|
+
* alter provider request options
|
|
54
|
+
|
|
55
|
+
The DeepSeek branch exists primarily to preserve a stable harness while providing observability around the prefix structure and cache usage.
|
|
56
|
+
|
|
57
|
+
This is intentional. The implementation describes DeepSeek as a passive policy whose purpose is to preserve the existing high-cache-rate behavior rather than introduce new request mutations.
|
|
58
|
+
|
|
59
|
+
The plugin still observes:
|
|
60
|
+
|
|
61
|
+
* system-prompt shape
|
|
62
|
+
* semantic tool definitions
|
|
63
|
+
* wire-order tool definitions
|
|
64
|
+
* prefix changes
|
|
65
|
+
* cache read tokens
|
|
66
|
+
* cache write tokens
|
|
67
|
+
* compaction boundaries
|
|
68
|
+
|
|
69
|
+
### Why passive?
|
|
70
|
+
|
|
71
|
+
DeepSeek's cache behavior is provider-managed. Introducing unnecessary prompt mutations would risk changing the prefix that the provider can reuse.
|
|
72
|
+
|
|
73
|
+
Therefore the plugin follows a simple rule:
|
|
74
|
+
|
|
75
|
+
```text
|
|
76
|
+
DeepSeek:
|
|
77
|
+
preserve request
|
|
78
|
+
preserve prefix
|
|
79
|
+
measure cache
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
rather than:
|
|
83
|
+
|
|
84
|
+
```text
|
|
85
|
+
DeepSeek:
|
|
86
|
+
rewrite request
|
|
87
|
+
guess cache key
|
|
88
|
+
force cache behavior
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
---
|
|
92
|
+
|
|
93
|
+
## GPT-5.6 Luna
|
|
94
|
+
|
|
95
|
+
### Policy: active cache control
|
|
96
|
+
|
|
97
|
+
GPT-5.6 is the only current policy that actively injects cache-control request metadata.
|
|
98
|
+
|
|
99
|
+
The plugin adds:
|
|
100
|
+
|
|
101
|
+
```json
|
|
102
|
+
{
|
|
103
|
+
"promptCacheKey": "<stable-session-key>",
|
|
104
|
+
"promptCacheOptions": {
|
|
105
|
+
"mode": "implicit",
|
|
106
|
+
"ttl": "30m"
|
|
107
|
+
}
|
|
108
|
+
}
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
The key is derived from the OpenCode session identity and is independent of transient request data. The implementation also preserves existing provider-supplied cache settings rather than overwriting them.
|
|
112
|
+
|
|
113
|
+
### Important: the prompt text is not rewritten
|
|
114
|
+
|
|
115
|
+
For GPT-5.6:
|
|
116
|
+
|
|
117
|
+
```text
|
|
118
|
+
system prompt -> unchanged
|
|
119
|
+
conversation -> unchanged
|
|
120
|
+
tool definitions -> unchanged
|
|
121
|
+
|
|
122
|
+
request metadata -> cache key/options added
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
This means the plugin is controlling the cache namespace and cache behavior without performing prompt surgery.
|
|
126
|
+
|
|
127
|
+
### Default GPT configuration
|
|
128
|
+
|
|
129
|
+
```json
|
|
130
|
+
{
|
|
131
|
+
"promptCacheKey": true,
|
|
132
|
+
"cacheRootKey": false,
|
|
133
|
+
"compactionCacheIsolation": true,
|
|
134
|
+
"reasoningEffortDiagnostics": true,
|
|
135
|
+
"mode": "implicit",
|
|
136
|
+
"ttl": "30m"
|
|
137
|
+
}
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
The current implementation intentionally leaves `cacheRootKey` disabled because the OpenCode runtime does not currently expose sufficiently reliable fork lineage for safe parent-cache inheritance. The code path remains available for a future runtime that exposes reliable parent relationships.
|
|
141
|
+
|
|
142
|
+
### Compaction isolation
|
|
143
|
+
|
|
144
|
+
Compaction uses a deterministic separate cache-key namespace:
|
|
145
|
+
|
|
146
|
+
```text
|
|
147
|
+
live session:
|
|
148
|
+
ses_abc123
|
|
149
|
+
|
|
150
|
+
compaction:
|
|
151
|
+
ses_abc123:compact
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
This prevents a compaction-specific prompt from sharing the same GPT cache namespace as the normal live-session prompt. The behavior is deterministic and tested explicitly.
|
|
155
|
+
|
|
156
|
+
---
|
|
157
|
+
|
|
158
|
+
## GLM-5.3 Flash
|
|
159
|
+
|
|
160
|
+
### Policy: input-shape optimization
|
|
161
|
+
|
|
162
|
+
GLM-5.3 receives the only prompt-text transformation in the current plugin.
|
|
163
|
+
|
|
164
|
+
The plugin identifies OpenCode's volatile `<env>` section and moves it to the **tail of the system prompt**.
|
|
165
|
+
|
|
166
|
+
Conceptually:
|
|
167
|
+
|
|
168
|
+
```text
|
|
169
|
+
BEFORE
|
|
170
|
+
|
|
171
|
+
[large stable instructions]
|
|
172
|
+
[volatile environment/date block]
|
|
173
|
+
[more stable instructions]
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
becomes:
|
|
177
|
+
|
|
178
|
+
```text
|
|
179
|
+
AFTER
|
|
180
|
+
|
|
181
|
+
[large stable instructions]
|
|
182
|
+
[more stable instructions]
|
|
183
|
+
[volatile environment/date block]
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
The contents of the environment block are preserved exactly. The operation changes its location, not its contents.
|
|
187
|
+
|
|
188
|
+
### Why?
|
|
189
|
+
|
|
190
|
+
The environment block can contain volatile information such as a changing date.
|
|
191
|
+
|
|
192
|
+
Keeping that material at the end allows the earlier portion of the system prompt to remain stable across requests.
|
|
193
|
+
|
|
194
|
+
The plugin therefore attempts to isolate volatility:
|
|
195
|
+
|
|
196
|
+
```text
|
|
197
|
+
stable prefix
|
|
198
|
+
---------------------------
|
|
199
|
+
unchanged across requests
|
|
200
|
+
|
|
201
|
+
volatile suffix
|
|
202
|
+
---------------------------
|
|
203
|
+
allowed to change
|
|
204
|
+
```
|
|
205
|
+
|
|
206
|
+
The system-shape diagnostics explicitly distinguish the stable prefix from the volatile suffix for this purpose.
|
|
207
|
+
|
|
208
|
+
### GLM safety constraints
|
|
209
|
+
|
|
210
|
+
The transformation is deliberately narrow.
|
|
211
|
+
|
|
212
|
+
It only occurs when:
|
|
213
|
+
|
|
214
|
+
* the selected model is GLM-5.3
|
|
215
|
+
* GLM stabilization is enabled
|
|
216
|
+
* there is exactly one system string
|
|
217
|
+
* the expected environment markers exist
|
|
218
|
+
* the block can be identified unambiguously
|
|
219
|
+
|
|
220
|
+
The plugin does not arbitrarily rearrange unrelated prompt content.
|
|
221
|
+
|
|
222
|
+
---
|
|
223
|
+
|
|
224
|
+
# Prompt-cache strategy
|
|
225
|
+
|
|
226
|
+
The plugin uses three different strategies because cache mechanisms differ by provider.
|
|
227
|
+
|
|
228
|
+
| Provider | Prompt text changed? | Cache metadata changed? | Main strategy |
|
|
229
|
+
| ----------------- | -------------------: | ----------------------: | --------------------------------- |
|
|
230
|
+
| DeepSeek V4 Flash | No | No | Preserve stable harness + observe |
|
|
231
|
+
| GPT-5.6 Luna | No | Yes | Stable cache key + cache options |
|
|
232
|
+
| GLM-5.3 Flash | Yes, narrowly | No provider cache key | Isolate volatile system content |
|
|
233
|
+
|
|
234
|
+
This distinction is fundamental.
|
|
235
|
+
|
|
236
|
+
The plugin is **not** a generic "rewrite every prompt for caching" engine.
|
|
237
|
+
|
|
238
|
+
It is a provider-aware cache policy engine.
|
|
239
|
+
|
|
240
|
+
---
|
|
241
|
+
|
|
242
|
+
# System-prompt diagnostics
|
|
243
|
+
|
|
244
|
+
The plugin fingerprints the system prompt to detect structural changes between requests.
|
|
245
|
+
|
|
246
|
+
For newer provider-aware diagnostics it tracks:
|
|
247
|
+
|
|
248
|
+
* full system hash
|
|
249
|
+
* stable system-prefix hash
|
|
250
|
+
* volatile system-suffix hash
|
|
251
|
+
|
|
252
|
+
The stable/volatile decomposition is based on the longest common prefix against the session baseline.
|
|
253
|
+
|
|
254
|
+
A change in a hash means:
|
|
255
|
+
|
|
256
|
+
> The observed request bytes changed.
|
|
257
|
+
|
|
258
|
+
It does **not** mean:
|
|
259
|
+
|
|
260
|
+
> The provider definitely generated a cache miss.
|
|
261
|
+
|
|
262
|
+
This distinction is intentional. Provider-reported cache token counts are the authoritative cache signal.
|
|
263
|
+
|
|
264
|
+
---
|
|
265
|
+
|
|
266
|
+
# Tool-definition diagnostics
|
|
267
|
+
|
|
268
|
+
Tool definitions are normalized before fingerprinting.
|
|
269
|
+
|
|
270
|
+
Runtime-only fields such as:
|
|
271
|
+
|
|
272
|
+
* object identity
|
|
273
|
+
* function references
|
|
274
|
+
* timestamps
|
|
275
|
+
* arbitrary runtime metadata
|
|
276
|
+
|
|
277
|
+
are excluded.
|
|
278
|
+
|
|
279
|
+
The semantic fingerprint is order-insensitive and represents the model-visible tool definitions.
|
|
280
|
+
|
|
281
|
+
The plugin also tracks wire-order fingerprints so that it can distinguish:
|
|
282
|
+
|
|
283
|
+
```text
|
|
284
|
+
same tools, different ordering
|
|
285
|
+
```
|
|
286
|
+
|
|
287
|
+
from:
|
|
288
|
+
|
|
289
|
+
```text
|
|
290
|
+
different tool definitions
|
|
291
|
+
```
|
|
292
|
+
|
|
293
|
+
This distinction matters because semantic equality and byte-level request equality are not necessarily the same thing.
|
|
294
|
+
|
|
295
|
+
The plugin uses these fingerprints for **diagnostics only**. It does not reorder the tools to force a particular fingerprint.
|
|
296
|
+
|
|
297
|
+
---
|
|
298
|
+
|
|
299
|
+
# Compaction handling
|
|
300
|
+
|
|
301
|
+
OpenCode sessions eventually undergo compaction as their conversation history grows.
|
|
302
|
+
|
|
303
|
+
The plugin adds a deterministic continuation template:
|
|
304
|
+
|
|
305
|
+
```text
|
|
306
|
+
## Session digest (cache-stable continuation block)
|
|
307
|
+
- Goal:
|
|
308
|
+
- Decisions made:
|
|
309
|
+
- Pending:
|
|
310
|
+
- Active files:
|
|
311
|
+
```
|
|
312
|
+
|
|
313
|
+
The digest is inserted once per compaction operation using a guard that prevents duplicate insertion if the compaction hook fires multiple times.
|
|
314
|
+
The objective is to provide a deterministic continuation structure rather than generating a different arbitrary cache-affecting block on every compaction.
|
|
315
|
+
|
|
316
|
+
---
|
|
317
|
+
|
|
318
|
+
# Cache metrics
|
|
319
|
+
|
|
320
|
+
The plugin records cache usage from OpenCode assistant-message token data.
|
|
321
|
+
|
|
322
|
+
At minimum it tracks:
|
|
323
|
+
|
|
324
|
+
```text
|
|
325
|
+
cache.read
|
|
326
|
+
cache.write
|
|
327
|
+
```
|
|
328
|
+
|
|
329
|
+
and aggregates those values across the session.
|
|
330
|
+
|
|
331
|
+
The default cache ratio reported by the core helper is:
|
|
332
|
+
|
|
333
|
+
```text
|
|
334
|
+
hit rate = read / (read + write)
|
|
335
|
+
```
|
|
336
|
+
|
|
337
|
+
This is deliberately an accounting metric based on cache read/write tokens.
|
|
338
|
+
|
|
339
|
+
For GLM, the implementation additionally calculates a prompt-token ratio:
|
|
340
|
+
|
|
341
|
+
```text
|
|
342
|
+
cached / (cached + cache-write + input)
|
|
343
|
+
```
|
|
344
|
+
|
|
345
|
+
using:
|
|
346
|
+
|
|
347
|
+
```text
|
|
348
|
+
read / (read + write + input)
|
|
349
|
+
```
|
|
350
|
+
|
|
351
|
+
as implemented by `glmHitRatio()`.
|
|
352
|
+
|
|
353
|
+
### Important metric distinction
|
|
354
|
+
|
|
355
|
+
These ratios answer different questions.
|
|
356
|
+
|
|
357
|
+
`read / (read + write)` answers approximately:
|
|
358
|
+
|
|
359
|
+
> Of the tokens represented as cache reads/writes, how much was reused?
|
|
360
|
+
|
|
361
|
+
`read / (read + write + input)` answers:
|
|
362
|
+
|
|
363
|
+
> How much of the total prompt-token accounting was represented by cached reads?
|
|
364
|
+
|
|
365
|
+
Do not treat the two percentages as interchangeable.
|
|
366
|
+
|
|
367
|
+
---
|
|
368
|
+
|
|
369
|
+
# Telemetry
|
|
370
|
+
|
|
371
|
+
Metrics are written as JSONL.
|
|
372
|
+
|
|
373
|
+
The default location is:
|
|
374
|
+
|
|
375
|
+
```text
|
|
376
|
+
~/.cache/opencode/cache-metrics.jsonl
|
|
377
|
+
```
|
|
378
|
+
|
|
379
|
+
The default configuration path is:
|
|
380
|
+
|
|
381
|
+
```text
|
|
382
|
+
~/.config/opencode/cache-engine.json
|
|
383
|
+
```
|
|
384
|
+
|
|
385
|
+
These paths are defined by the plugin core.
|
|
386
|
+
|
|
387
|
+
Telemetry is best-effort.
|
|
388
|
+
|
|
389
|
+
A failed metrics write must never break an OpenCode request. The recorder catches write failures rather than allowing telemetry failures to affect execution.
|
|
390
|
+
|
|
391
|
+
---
|
|
392
|
+
|
|
393
|
+
# Metrics examples
|
|
394
|
+
|
|
395
|
+
A usage record can contain fields such as:
|
|
396
|
+
|
|
397
|
+
```json
|
|
398
|
+
{
|
|
399
|
+
"kind": "usage-event",
|
|
400
|
+
"sid": "session-id",
|
|
401
|
+
"ts": 1750000000000,
|
|
402
|
+
"read": 120000,
|
|
403
|
+
"write": 3000,
|
|
404
|
+
"cost": 0.0123,
|
|
405
|
+
"provider": "z-ai",
|
|
406
|
+
"model": "glm-5.3-flash",
|
|
407
|
+
"policy": "glm53"
|
|
408
|
+
}
|
|
409
|
+
```
|
|
410
|
+
|
|
411
|
+
A prefix-change record can look like:
|
|
412
|
+
|
|
413
|
+
```json
|
|
414
|
+
{
|
|
415
|
+
"kind": "prefix-change",
|
|
416
|
+
"sid": "session-id",
|
|
417
|
+
"ts": 1750000000000,
|
|
418
|
+
"dimensions": [
|
|
419
|
+
"system"
|
|
420
|
+
]
|
|
421
|
+
}
|
|
422
|
+
```
|
|
423
|
+
|
|
424
|
+
A compaction record can contain:
|
|
425
|
+
|
|
426
|
+
```json
|
|
427
|
+
{
|
|
428
|
+
"kind": "compaction",
|
|
429
|
+
"sid": "session-id",
|
|
430
|
+
"ts": 1750000000000,
|
|
431
|
+
"reason": "compaction",
|
|
432
|
+
"usageSamples": 7,
|
|
433
|
+
"cumulative": {
|
|
434
|
+
"read": 900000,
|
|
435
|
+
"write": 12000
|
|
436
|
+
}
|
|
437
|
+
}
|
|
438
|
+
```
|
|
439
|
+
|
|
440
|
+
Telemetry is intended to answer questions such as:
|
|
441
|
+
|
|
442
|
+
* Did the system prompt change?
|
|
443
|
+
* Did the tool definitions change?
|
|
444
|
+
* Did cache reads increase?
|
|
445
|
+
* Did cache writes increase?
|
|
446
|
+
* Did a compaction occur?
|
|
447
|
+
* Which provider/model/policy was active?
|
|
448
|
+
* Did the GLM system stabilization actually change the observed prompt shape?
|
|
449
|
+
|
|
450
|
+
---
|
|
451
|
+
|
|
452
|
+
# Configuration
|
|
453
|
+
|
|
454
|
+
The default configuration is:
|
|
455
|
+
|
|
456
|
+
```json
|
|
457
|
+
{
|
|
458
|
+
"enabled": true,
|
|
459
|
+
"metricsFile": "~/.cache/opencode/cache-metrics.jsonl",
|
|
460
|
+
"compactTemplate": true,
|
|
461
|
+
"logPrefixChanges": true,
|
|
462
|
+
"policies": {
|
|
463
|
+
"deepseek": {
|
|
464
|
+
"enabled": true
|
|
465
|
+
},
|
|
466
|
+
"gpt56": {
|
|
467
|
+
"enabled": true,
|
|
468
|
+
"promptCacheKey": true,
|
|
469
|
+
"cacheRootKey": false,
|
|
470
|
+
"compactionCacheIsolation": true,
|
|
471
|
+
"reasoningEffortDiagnostics": true,
|
|
472
|
+
"mode": "implicit",
|
|
473
|
+
"ttl": "30m"
|
|
474
|
+
},
|
|
475
|
+
"glm53": {
|
|
476
|
+
"enabled": true,
|
|
477
|
+
"stabilizeSystem": true,
|
|
478
|
+
"preserveThinkingIntegrity": true
|
|
479
|
+
}
|
|
480
|
+
}
|
|
481
|
+
}
|
|
482
|
+
```
|
|
483
|
+
|
|
484
|
+
The configuration parser starts from these defaults and applies valid file/environment overrides without mutating the caller's configuration object.
|
|
485
|
+
|
|
486
|
+
---
|
|
487
|
+
|
|
488
|
+
# Configuration options
|
|
489
|
+
|
|
490
|
+
## Global
|
|
491
|
+
|
|
492
|
+
### `enabled`
|
|
493
|
+
|
|
494
|
+
```json
|
|
495
|
+
{
|
|
496
|
+
"enabled": true
|
|
497
|
+
}
|
|
498
|
+
```
|
|
499
|
+
|
|
500
|
+
Enables or disables the entire plugin.
|
|
501
|
+
|
|
502
|
+
---
|
|
503
|
+
|
|
504
|
+
### `metricsFile`
|
|
505
|
+
|
|
506
|
+
```json
|
|
507
|
+
{
|
|
508
|
+
"metricsFile": "~/.cache/opencode/cache-metrics.jsonl"
|
|
509
|
+
}
|
|
510
|
+
```
|
|
511
|
+
|
|
512
|
+
Controls where JSONL telemetry is written.
|
|
513
|
+
|
|
514
|
+
---
|
|
515
|
+
|
|
516
|
+
### `compactTemplate`
|
|
517
|
+
|
|
518
|
+
```json
|
|
519
|
+
{
|
|
520
|
+
"compactTemplate": true
|
|
521
|
+
}
|
|
522
|
+
```
|
|
523
|
+
|
|
524
|
+
Controls whether the deterministic compaction continuation block is inserted.
|
|
525
|
+
|
|
526
|
+
---
|
|
527
|
+
|
|
528
|
+
### `logPrefixChanges`
|
|
529
|
+
|
|
530
|
+
```json
|
|
531
|
+
{
|
|
532
|
+
"logPrefixChanges": true
|
|
533
|
+
}
|
|
534
|
+
```
|
|
535
|
+
|
|
536
|
+
Controls warning logs for observed prefix-shape changes.
|
|
537
|
+
|
|
538
|
+
---
|
|
539
|
+
|
|
540
|
+
# DeepSeek configuration
|
|
541
|
+
|
|
542
|
+
```json
|
|
543
|
+
"deepseek": {
|
|
544
|
+
"enabled": true
|
|
545
|
+
}
|
|
546
|
+
```
|
|
547
|
+
|
|
548
|
+
There are intentionally very few settings here.
|
|
549
|
+
|
|
550
|
+
DeepSeek is treated as the conservative/passive policy.
|
|
551
|
+
|
|
552
|
+
---
|
|
553
|
+
|
|
554
|
+
# GPT-5.6 configuration
|
|
555
|
+
|
|
556
|
+
```json
|
|
557
|
+
"gpt56": {
|
|
558
|
+
"enabled": true,
|
|
559
|
+
"promptCacheKey": true,
|
|
560
|
+
"cacheRootKey": false,
|
|
561
|
+
"compactionCacheIsolation": true,
|
|
562
|
+
"reasoningEffortDiagnostics": true,
|
|
563
|
+
"mode": "implicit",
|
|
564
|
+
"ttl": "30m"
|
|
565
|
+
}
|
|
566
|
+
```
|
|
567
|
+
|
|
568
|
+
### `promptCacheKey`
|
|
569
|
+
|
|
570
|
+
Controls whether the plugin provides a stable session-derived GPT cache key.
|
|
571
|
+
|
|
572
|
+
### `cacheRootKey`
|
|
573
|
+
|
|
574
|
+
Controls whether a parent/fork cache root is used.
|
|
575
|
+
|
|
576
|
+
Disabled by default because reliable fork lineage is not currently guaranteed by the runtime.
|
|
577
|
+
|
|
578
|
+
### `compactionCacheIsolation`
|
|
579
|
+
|
|
580
|
+
Uses a separate deterministic cache namespace for compaction requests.
|
|
581
|
+
|
|
582
|
+
### `reasoningEffortDiagnostics`
|
|
583
|
+
|
|
584
|
+
Tracks GPT reasoning-effort changes for diagnostics.
|
|
585
|
+
|
|
586
|
+
### `mode`
|
|
587
|
+
|
|
588
|
+
Defaults to:
|
|
589
|
+
|
|
590
|
+
```text
|
|
591
|
+
implicit
|
|
592
|
+
```
|
|
593
|
+
|
|
594
|
+
### `ttl`
|
|
595
|
+
|
|
596
|
+
Defaults to:
|
|
597
|
+
|
|
598
|
+
```text
|
|
599
|
+
30m
|
|
600
|
+
```
|
|
601
|
+
|
|
602
|
+
Existing request options are not overwritten by the plugin.
|
|
603
|
+
|
|
604
|
+
---
|
|
605
|
+
|
|
606
|
+
# GLM-5.3 configuration
|
|
607
|
+
|
|
608
|
+
```json
|
|
609
|
+
"glm53": {
|
|
610
|
+
"enabled": true,
|
|
611
|
+
"stabilizeSystem": true,
|
|
612
|
+
"preserveThinkingIntegrity": true
|
|
613
|
+
}
|
|
614
|
+
```
|
|
615
|
+
|
|
616
|
+
### `stabilizeSystem`
|
|
617
|
+
|
|
618
|
+
Enables relocation of the volatile `<env>` section to the system-prompt tail.
|
|
619
|
+
|
|
620
|
+
### `preserveThinkingIntegrity`
|
|
621
|
+
|
|
622
|
+
Enables diagnostic checks around reasoning continuity.
|
|
623
|
+
|
|
624
|
+
The reasoning instrumentation is intended to identify anomalies such as:
|
|
625
|
+
|
|
626
|
+
* duplicate reasoning
|
|
627
|
+
* reordered reasoning
|
|
628
|
+
* modified reasoning
|
|
629
|
+
|
|
630
|
+
It is diagnostic rather than a reason to rewrite or fabricate reasoning content. The implementation maps these conditions to explicit diagnostic reasons.
|
|
631
|
+
|
|
632
|
+
---
|
|
633
|
+
|
|
634
|
+
# Model detection
|
|
635
|
+
|
|
636
|
+
The plugin classifies requests into:
|
|
637
|
+
|
|
638
|
+
```text
|
|
639
|
+
deepseek
|
|
640
|
+
gpt56
|
|
641
|
+
glm53
|
|
642
|
+
neutral
|
|
643
|
+
```
|
|
644
|
+
|
|
645
|
+
The model detector recognizes:
|
|
646
|
+
|
|
647
|
+
* DeepSeek model/provider identifiers
|
|
648
|
+
* GPT-5.6 variants
|
|
649
|
+
* GLM-5.3 variants
|
|
650
|
+
|
|
651
|
+
GPT-5.6 has an additional OpenAI/Azure-context check so a string containing `gpt-5.6` does not automatically cause GPT-specific fields to be sent to an unrelated endpoint.
|
|
652
|
+
|
|
653
|
+
Unknown models use the neutral policy.
|
|
654
|
+
|
|
655
|
+
Neutral means:
|
|
656
|
+
|
|
657
|
+
```text
|
|
658
|
+
no provider-specific request mutation
|
|
659
|
+
```
|
|
660
|
+
|
|
661
|
+
---
|
|
662
|
+
|
|
663
|
+
# OpenRouter usage
|
|
664
|
+
|
|
665
|
+
This plugin is compatible with OpenRouter because the cache policy is based on the model/provider signals available to OpenCode.
|
|
666
|
+
|
|
667
|
+
For cache-sensitive workloads, provider stability remains important.
|
|
668
|
+
|
|
669
|
+
The plugin does not attempt to compensate for provider switching by rewriting prompts.
|
|
670
|
+
|
|
671
|
+
For that reason, a stable provider route is preferable when your goal is to measure and maximize prefix reuse.
|
|
672
|
+
|
|
673
|
+
---
|
|
674
|
+
|
|
675
|
+
# Architecture
|
|
676
|
+
|
|
677
|
+
The implementation is split into two layers.
|
|
678
|
+
|
|
679
|
+
## `cache-engine.ts`
|
|
680
|
+
|
|
681
|
+
This is the OpenCode plugin entry point.
|
|
682
|
+
|
|
683
|
+
It owns:
|
|
684
|
+
|
|
685
|
+
* OpenCode hooks
|
|
686
|
+
* session state
|
|
687
|
+
* provider-policy selection
|
|
688
|
+
* telemetry integration
|
|
689
|
+
* request mutation
|
|
690
|
+
* system-prompt transformation
|
|
691
|
+
* compaction handling
|
|
692
|
+
|
|
693
|
+
The exported plugin is:
|
|
694
|
+
|
|
695
|
+
```ts
|
|
696
|
+
export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
697
|
+
// ...
|
|
698
|
+
}
|
|
699
|
+
```
|
|
700
|
+
|
|
701
|
+
The identifier `CacheEngine` is the OpenCode plugin export name. It does not determine the eventual npm package name.
|
|
702
|
+
|
|
703
|
+
---
|
|
704
|
+
|
|
705
|
+
## `cache-engine-core.mjs`
|
|
706
|
+
|
|
707
|
+
This contains dependency-light pure logic.
|
|
708
|
+
|
|
709
|
+
It owns:
|
|
710
|
+
|
|
711
|
+
* provider classification
|
|
712
|
+
* configuration parsing
|
|
713
|
+
* hashing
|
|
714
|
+
* canonicalization
|
|
715
|
+
* tool fingerprints
|
|
716
|
+
* system-shape decomposition
|
|
717
|
+
* GPT cache-key generation
|
|
718
|
+
* cache-option generation
|
|
719
|
+
* GLM environment relocation
|
|
720
|
+
* reasoning diagnostics
|
|
721
|
+
* usage aggregation
|
|
722
|
+
* compaction guards
|
|
723
|
+
|
|
724
|
+
Keeping these functions in plain JavaScript allows the logic to be tested independently with Node's built-in test runner.
|
|
725
|
+
|
|
726
|
+
---
|
|
727
|
+
|
|
728
|
+
## Tests
|
|
729
|
+
|
|
730
|
+
The repository's test suite validates the provider-independent and provider-specific logic.
|
|
731
|
+
|
|
732
|
+
Coverage includes:
|
|
733
|
+
|
|
734
|
+
* model detection
|
|
735
|
+
* GPT cache-key stability
|
|
736
|
+
* GPT cache-option defaults
|
|
737
|
+
* protection against overwriting existing cache options
|
|
738
|
+
* GLM environment relocation
|
|
739
|
+
* deterministic hashing
|
|
740
|
+
* system-prefix decomposition
|
|
741
|
+
* tool fingerprints
|
|
742
|
+
* reasoning diagnostics
|
|
743
|
+
* compaction isolation
|
|
744
|
+
* cache-hit calculations
|
|
745
|
+
* configuration behavior
|
|
746
|
+
* JSONL telemetry behavior
|
|
747
|
+
|
|
748
|
+
The tests are designed around the pure core logic, while OpenCode runtime behavior is validated separately through actual plugin loading.
|
|
749
|
+
|
|
750
|
+
---
|
|
751
|
+
|
|
752
|
+
# Design principles
|
|
753
|
+
|
|
754
|
+
## 1. Provider-specific behavior
|
|
755
|
+
|
|
756
|
+
Different providers expose different cache mechanisms.
|
|
757
|
+
|
|
758
|
+
The plugin therefore does not assume that one strategy is optimal everywhere.
|
|
759
|
+
|
|
760
|
+
---
|
|
761
|
+
|
|
762
|
+
## 2. Preserve working behavior
|
|
763
|
+
|
|
764
|
+
The plugin should not modify a provider's request merely because a mutation is technically possible.
|
|
765
|
+
|
|
766
|
+
This is especially important for DeepSeek, where the current policy is intentionally passive.
|
|
767
|
+
|
|
768
|
+
---
|
|
769
|
+
|
|
770
|
+
## 3. Measure provider reality
|
|
771
|
+
|
|
772
|
+
Local hashes are diagnostics.
|
|
773
|
+
|
|
774
|
+
Provider-reported cache token counts are the authoritative signal.
|
|
775
|
+
|
|
776
|
+
The implementation explicitly distinguishes:
|
|
777
|
+
|
|
778
|
+
```text
|
|
779
|
+
observed prefix change
|
|
780
|
+
```
|
|
781
|
+
|
|
782
|
+
from:
|
|
783
|
+
|
|
784
|
+
```text
|
|
785
|
+
confirmed provider cache miss
|
|
786
|
+
```
|
|
787
|
+
|
|
788
|
+
because the plugin cannot infer the latter reliably from local prompt hashes alone.
|
|
789
|
+
|
|
790
|
+
---
|
|
791
|
+
|
|
792
|
+
## 4. Never overwrite explicit provider configuration
|
|
793
|
+
|
|
794
|
+
Where GPT cache options already exist, the plugin leaves them alone.
|
|
795
|
+
|
|
796
|
+
This allows the runtime or user configuration to remain authoritative.
|
|
797
|
+
|
|
798
|
+
---
|
|
799
|
+
|
|
800
|
+
## 5. Keep mutations deterministic
|
|
801
|
+
|
|
802
|
+
When the plugin does transform the request, the transformation should be:
|
|
803
|
+
|
|
804
|
+
* narrow
|
|
805
|
+
* deterministic
|
|
806
|
+
* content-preserving where possible
|
|
807
|
+
* provider-specific
|
|
808
|
+
* easy to disable
|
|
809
|
+
|
|
810
|
+
The GLM environment relocation follows these rules.
|
|
811
|
+
|
|
812
|
+
---
|
|
813
|
+
|
|
814
|
+
## 6. Keep telemetry out of the critical path
|
|
815
|
+
|
|
816
|
+
A metrics failure must not break model execution.
|
|
817
|
+
|
|
818
|
+
Telemetry is therefore best-effort.
|
|
819
|
+
|
|
820
|
+
---
|
|
821
|
+
|
|
822
|
+
# What the plugin does NOT do
|
|
823
|
+
|
|
824
|
+
The plugin does not:
|
|
825
|
+
|
|
826
|
+
* invent cache hits
|
|
827
|
+
* claim a local hash proves a provider cache hit
|
|
828
|
+
* rewrite DeepSeek prompts
|
|
829
|
+
* reorder tools
|
|
830
|
+
* fabricate reasoning
|
|
831
|
+
* modify conversation history arbitrarily
|
|
832
|
+
* force explicit GPT cache breakpoints by default
|
|
833
|
+
* silently overwrite existing GPT cache options
|
|
834
|
+
* assume every model named `gpt-5.6` is an OpenAI-compatible endpoint
|
|
835
|
+
* use fork inheritance unless reliable lineage is available
|
|
836
|
+
|
|
837
|
+
---
|
|
838
|
+
|
|
839
|
+
# Cost optimization philosophy
|
|
840
|
+
|
|
841
|
+
Cache hit rate is useful, but it is not the only cost metric.
|
|
842
|
+
|
|
843
|
+
The economic objective is:
|
|
844
|
+
|
|
845
|
+
```text
|
|
846
|
+
total task cost
|
|
847
|
+
=
|
|
848
|
+
prompt/cache cost
|
|
849
|
+
+
|
|
850
|
+
output/reasoning cost
|
|
851
|
+
+
|
|
852
|
+
additional requests
|
|
853
|
+
```
|
|
854
|
+
|
|
855
|
+
A model with a slightly lower cache hit rate can still be cheaper if it completes the task with fewer tokens or fewer model calls.
|
|
856
|
+
|
|
857
|
+
For that reason, this plugin is primarily an **instrumentation + targeted optimization layer**, not a cache-rate maximizer at any cost.
|
|
858
|
+
|
|
859
|
+
The recommended evaluation unit is:
|
|
860
|
+
|
|
861
|
+
```text
|
|
862
|
+
cost per completed task
|
|
863
|
+
```
|
|
864
|
+
|
|
865
|
+
rather than:
|
|
866
|
+
|
|
867
|
+
```text
|
|
868
|
+
cache percentage alone
|
|
869
|
+
```
|
|
870
|
+
|
|
871
|
+
---
|
|
872
|
+
|
|
873
|
+
# Operational recommendations
|
|
874
|
+
|
|
875
|
+
For reliable cache measurements:
|
|
876
|
+
|
|
877
|
+
1. Keep the provider fixed whenever possible.
|
|
878
|
+
2. Avoid changing unrelated system-prompt content during a benchmark.
|
|
879
|
+
3. Keep tool definitions stable.
|
|
880
|
+
4. Compare equivalent tasks across models.
|
|
881
|
+
5. Record actual provider cache token counts.
|
|
882
|
+
6. Compare total task cost, not only cache percentage.
|
|
883
|
+
7. Treat compaction as a separate cache boundary when analyzing results.
|
|
884
|
+
8. Avoid interpreting a local prefix hash change as definitive proof of a cache miss.
|
|
885
|
+
|
|
886
|
+
---
|
|
887
|
+
|
|
888
|
+
# File layout
|
|
889
|
+
|
|
890
|
+
A typical standalone repository can use:
|
|
891
|
+
|
|
892
|
+
```text
|
|
893
|
+
opencode-cache-engine/
|
|
894
|
+
├── src/
|
|
895
|
+
│ ├── cache-engine.ts
|
|
896
|
+
│ └── cache-engine-core.mjs
|
|
897
|
+
├── test/
|
|
898
|
+
│ └── cache-engine.test.mjs
|
|
899
|
+
├── examples/
|
|
900
|
+
│ └── cache-engine.json
|
|
901
|
+
├── README.md
|
|
902
|
+
├── LICENSE
|
|
903
|
+
└── package.json
|
|
904
|
+
```
|
|
905
|
+
|
|
906
|
+
The OpenCode plugin export remains:
|
|
907
|
+
|
|
908
|
+
```ts
|
|
909
|
+
export const CacheEngine
|
|
910
|
+
```
|
|
911
|
+
|
|
912
|
+
regardless of the eventual npm package name.
|
|
913
|
+
|
|
914
|
+
For example, the npm package could be named:
|
|
915
|
+
|
|
916
|
+
```text
|
|
917
|
+
opencode-cache-engine
|
|
918
|
+
```
|
|
919
|
+
|
|
920
|
+
without changing the `CacheEngine` export identifier.
|
|
921
|
+
|
|
922
|
+
---
|
|
923
|
+
|
|
924
|
+
# Installation
|
|
925
|
+
|
|
926
|
+
Install the plugin into the OpenCode plugins directory according to your OpenCode plugin-loading setup.
|
|
927
|
+
|
|
928
|
+
The runtime entry should expose:
|
|
929
|
+
|
|
930
|
+
```ts
|
|
931
|
+
export const CacheEngine: Plugin = async ({ client, directory }) => {
|
|
932
|
+
// ...
|
|
933
|
+
}
|
|
934
|
+
```
|
|
935
|
+
|
|
936
|
+
After installation, verify that OpenCode loads the plugin successfully before benchmarking cache behavior.
|
|
937
|
+
|
|
938
|
+
---
|
|
939
|
+
|
|
940
|
+
# Validation
|
|
941
|
+
|
|
942
|
+
The core test suite can be run with Node:
|
|
943
|
+
|
|
944
|
+
```bash
|
|
945
|
+
node --test test/cache-engine.test.mjs
|
|
946
|
+
```
|
|
947
|
+
|
|
948
|
+
The tests are intentionally dependency-light and exercise the pure logic independently of the OpenCode runtime.
|
|
949
|
+
|
|
950
|
+
Runtime validation should additionally confirm:
|
|
951
|
+
|
|
952
|
+
```text
|
|
953
|
+
DeepSeek:
|
|
954
|
+
no request mutation
|
|
955
|
+
|
|
956
|
+
GPT-5.6:
|
|
957
|
+
promptCacheKey present
|
|
958
|
+
promptCacheOptions present
|
|
959
|
+
|
|
960
|
+
GLM-5.3:
|
|
961
|
+
volatile env block relocated when eligible
|
|
962
|
+
```
|
|
963
|
+
|
|
964
|
+
---
|
|
965
|
+
|
|
966
|
+
# Troubleshooting
|
|
967
|
+
|
|
968
|
+
## DeepSeek cache rate dropped
|
|
969
|
+
|
|
970
|
+
First check provider stability and whether OpenCode's system/tool prefix changed.
|
|
971
|
+
|
|
972
|
+
The plugin itself does not intentionally mutate DeepSeek request options.
|
|
973
|
+
|
|
974
|
+
Inspect the telemetry for:
|
|
975
|
+
|
|
976
|
+
```text
|
|
977
|
+
prefix-change
|
|
978
|
+
usage-event
|
|
979
|
+
compaction
|
|
980
|
+
```
|
|
981
|
+
|
|
982
|
+
A prefix change is a diagnostic signal, not automatic proof of a cache miss.
|
|
983
|
+
|
|
984
|
+
---
|
|
985
|
+
|
|
986
|
+
## GPT-5.6 cache options are missing
|
|
987
|
+
|
|
988
|
+
Verify that the model is actually classified as GPT-5.6 and that the endpoint is recognized as OpenAI/Azure-compatible.
|
|
989
|
+
|
|
990
|
+
The detector intentionally rejects ambiguous OpenAI-compatible providers rather than guessing.
|
|
991
|
+
|
|
992
|
+
Also check whether the outgoing request already supplied its own cache options. Existing settings are intentionally preserved.
|
|
993
|
+
|
|
994
|
+
---
|
|
995
|
+
|
|
996
|
+
## GLM-5.3 prompt is not being changed
|
|
997
|
+
|
|
998
|
+
The environment relocation only occurs when the plugin can identify the expected block unambiguously.
|
|
999
|
+
|
|
1000
|
+
The relevant block must contain the expected beginning and closing marker, and the system structure must meet the plugin's eligibility rules.
|
|
1001
|
+
|
|
1002
|
+
---
|
|
1003
|
+
|
|
1004
|
+
## Metrics file is missing
|
|
1005
|
+
|
|
1006
|
+
Telemetry is best-effort.
|
|
1007
|
+
|
|
1008
|
+
Check:
|
|
1009
|
+
|
|
1010
|
+
```text
|
|
1011
|
+
~/.cache/opencode/cache-metrics.jsonl
|
|
1012
|
+
```
|
|
1013
|
+
|
|
1014
|
+
and verify that the configured parent directory is writable.
|
|
1015
|
+
|
|
1016
|
+
A telemetry failure is intentionally swallowed so it does not break model execution.
|
|
1017
|
+
|
|
1018
|
+
---
|
|
1019
|
+
|
|
1020
|
+
# Status
|
|
1021
|
+
|
|
1022
|
+
The current implementation is intentionally conservative:
|
|
1023
|
+
|
|
1024
|
+
```text
|
|
1025
|
+
DeepSeek -> preserve and measure
|
|
1026
|
+
GPT-5.6 -> configure cache controls
|
|
1027
|
+
GLM-5.3 -> isolate volatile prompt content
|
|
1028
|
+
```
|
|
1029
|
+
|
|
1030
|
+
That separation is the core design of the project.
|
|
1031
|
+
|
|
1032
|
+
The plugin should be evaluated using real provider-reported usage and real task cost rather than assuming that any particular local transformation guarantees a cache hit.
|