@yadsh/dsh-kv-persist 0.2.3 → 0.2.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. package/README.md +8 -5
  2. package/compatibility.json +1 -3
  3. package/lib/backends/llama-cpp/backend.d.ts.map +1 -1
  4. package/lib/backends/llama-cpp/backend.js +3 -1
  5. package/lib/backends/llama-cpp/backend.js.map +1 -1
  6. package/lib/backends/llama-cpp/client.d.ts.map +1 -1
  7. package/lib/backends/llama-cpp/client.js +21 -11
  8. package/lib/backends/llama-cpp/client.js.map +1 -1
  9. package/lib/backends/types.d.ts.map +1 -1
  10. package/lib/config.d.ts.map +1 -1
  11. package/lib/config.js +19 -6
  12. package/lib/config.js.map +1 -1
  13. package/lib/coordinator/checkpoint-policy.d.ts.map +1 -1
  14. package/lib/coordinator/checkpoint-policy.js +1 -1
  15. package/lib/coordinator/checkpoint-policy.js.map +1 -1
  16. package/lib/coordinator/coordinator.d.ts.map +1 -1
  17. package/lib/coordinator/coordinator.js +90 -15
  18. package/lib/coordinator/coordinator.js.map +1 -1
  19. package/lib/errors.d.ts.map +1 -1
  20. package/lib/errors.js.map +1 -1
  21. package/lib/index.d.ts +3 -3
  22. package/lib/index.d.ts.map +1 -1
  23. package/lib/index.js +1 -1
  24. package/lib/index.js.map +1 -1
  25. package/lib/observability/diagnostics.d.ts.map +1 -1
  26. package/lib/observability/diagnostics.js +4 -2
  27. package/lib/observability/diagnostics.js.map +1 -1
  28. package/lib/observability/metrics.d.ts.map +1 -1
  29. package/lib/observability/metrics.js.map +1 -1
  30. package/lib/service.d.ts.map +1 -1
  31. package/lib/service.js +13 -5
  32. package/lib/service.js.map +1 -1
  33. package/lib/snapshots/fingerprint.d.ts.map +1 -1
  34. package/lib/snapshots/fingerprint.js.map +1 -1
  35. package/lib/snapshots/manifest.d.ts.map +1 -1
  36. package/lib/snapshots/manifest.js +0 -0
  37. package/lib/snapshots/manifest.js.map +1 -1
  38. package/lib/snapshots/repository.d.ts.map +1 -1
  39. package/lib/snapshots/repository.js +28 -10
  40. package/lib/snapshots/repository.js.map +1 -1
  41. package/package.json +7 -6
  42. package/docs/dsh-kv-persist.md +0 -3312
@@ -1,3312 +0,0 @@
1
- # dsh-kv-persist
2
-
3
- > Persistent KV-cache/session-state manager for DeepSeek Harness.
4
-
5
- **Status:** Draft / Initial Specification
6
- **Target:** DeepSeek Harness + llama.cpp `llama-server`
7
- **Initial version:** `0.1.0`
8
- **Primary backend:** llama.cpp slot save/restore API
9
- **Future backends:** vLLM / SGLang / other providers exposing reusable prefix/session state
10
-
11
- ---
12
-
13
- ## 1. Summary
14
-
15
- `dsh-kv-persist` is a DeepSeek Harness infrastructure plugin that persists LLM runtime cache state between sessions and process restarts.
16
-
17
- The first implementation targets `llama-server` and its slot persistence API:
18
-
19
- - `GET /slots`
20
- - `POST /slots/{id}?action=save`
21
- - `POST /slots/{id}?action=restore`
22
- - `POST /slots/{id}?action=erase`
23
-
24
- llama.cpp can persist a slot's prompt/KV state into a file under `--slot-save-path`, then restore that state later.
25
-
26
- The plugin maps:
27
-
28
- ```text
29
- DSH session
30
- ↓
31
- provider + model + server instance
32
- ↓
33
- llama.cpp slot
34
- ↓
35
- persistent snapshot
36
- ```
37
-
38
- Its main purpose is to avoid re-prefilling tens or hundreds of thousands of tokens whenever:
39
-
40
- - the user switches between DSH sessions;
41
- - `llama-server` evicts a session from its active slot;
42
- - DSH is restarted;
43
- - `llama-server` is restarted;
44
- - another request temporarily pollutes the active slot;
45
- - multiple projects share the same local model server.
46
-
47
- For a large agent prompt this can turn:
48
-
49
- ```text
50
- restore session
51
- → process 40K–100K prompt tokens again
52
- → wait tens/hundreds of seconds
53
- ```
54
-
55
- into:
56
-
57
- ```text
58
- restore session
59
- → load KV snapshot
60
- → process only changed suffix
61
- ```
62
-
63
- ---
64
-
65
- # 2. Goals
66
-
67
- The plugin MUST:
68
-
69
- 1. Associate persistent cache snapshots with DSH `session.id`.
70
- 2. Automatically restore the appropriate snapshot before a session resumes.
71
- 3. Automatically save dirty cache state according to a configurable checkpoint policy.
72
- 4. Prevent one DSH session from accidentally reusing another session's slot as authoritative state.
73
- 5. Detect incompatible or stale snapshots and fail safely.
74
- 6. Never make the model-visible conversation depend on the snapshot.
75
- 7. Treat KV persistence purely as an optimization.
76
- 8. Fall back to normal cold prompt processing whenever persistence is unavailable.
77
- 9. Work with a normal OpenAI-compatible `llama-server`.
78
- 10. Provide useful observability:
79
- - restore hit/miss;
80
- - bytes saved/read;
81
- - save/restore latency;
82
- - cache tokens;
83
- - slot ownership;
84
- - snapshot age;
85
- - cold fallback count.
86
- 11. Survive plugin hot reload and normal Cordis disposal cleanly.
87
- 12. Be backend-agnostic internally even though llama.cpp is the first backend.
88
-
89
- DeepSeek Harness already treats sessions as append-only durable state and derives LLM messages from that state, so persisted KV MUST remain a disposable acceleration layer rather than a second source of truth.
90
-
91
- ---
92
-
93
- # 3. Non-goals
94
-
95
- Initial versions MUST NOT:
96
-
97
- - replace DSH session persistence;
98
- - store actual conversation history as the authoritative state;
99
- - attempt to reconstruct missing DSH events from KV;
100
- - modify model-visible prompts to improve cache hits;
101
- - expose KV management as an LLM tool;
102
- - assume snapshots are portable between different models;
103
- - assume snapshots are portable between arbitrary llama.cpp builds;
104
- - manage llama-server process startup itself;
105
- - transparently migrate a snapshot between machines;
106
- - promise persistence support for every recurrent/hybrid model;
107
- - require DSH UI modifications.
108
-
109
- The model should ideally never know this plugin exists.
110
-
111
- ---
112
-
113
- # 4. Core design principle
114
-
115
- The following relationship is fundamental:
116
-
117
- ```text
118
- DSH session log = truth
119
- KV snapshot = derived cache
120
- ```
121
-
122
- A snapshot can always be deleted.
123
-
124
- Deleting:
125
-
126
- ```text
127
- $DSH_HOME/cache/dsh-kv-persist/...
128
- ```
129
-
130
- must never destroy the conversation.
131
-
132
- The worst possible consequence of missing or invalid KV state must be:
133
-
134
- ```text
135
- cold prefill
136
- ```
137
-
138
- not:
139
-
140
- ```text
141
- corrupted conversation
142
- wrong session
143
- missing messages
144
- incorrect continuation
145
- ```
146
-
147
- ---
148
-
149
- # 5. Relevant DSH architecture
150
-
151
- DSH exposes an in-memory `ctx.sessions` service.
152
-
153
- A `Session` has a stable `session.id` and an append-only sequence of `SessionEvent`s. LLM history is derived from those events rather than stored as a separate mutable chat history. DSH persistence plugins can observe `session/event`, `session/flush`, `session/created`, and `session/disposed`.
154
-
155
- Model calls eventually pass through:
156
-
157
- ```text
158
- agent loop
159
- ↓
160
- agent/request
161
- ↓
162
- prepareCall()
163
- ↓
164
- llm/stream
165
- ↓
166
- LLM adapter
167
- ↓
168
- llama-server
169
- ```
170
-
171
- `llm/stream` is explicitly intended to support infrastructure such as caching/logging/routing. Loop-generated `GenerateOptions` are deep-frozen and must be observed rather than mutated.
172
-
173
- `GenerateOptions` also carries:
174
-
175
- ```ts
176
- sessionId?: SessionId
177
- purpose?: 'compaction' | 'session-title'
178
- ```
179
-
180
- which gives the plugin enough information to distinguish ordinary session inference from auxiliary LLM requests.
181
-
182
- ---
183
-
184
- # 6. llama.cpp backend
185
-
186
- ## 6.1 Server requirements
187
-
188
- The llama server should be started with at least:
189
-
190
- ```bash
191
- llama-server \
192
- ... \
193
- --slots \
194
- --slot-save-path /some/path
195
- ```
196
-
197
- For the initial MVP:
198
-
199
- ```bash
200
- --parallel 1
201
- ```
202
-
203
- is strongly recommended.
204
-
205
- Example:
206
-
207
- ```bash
208
- ./llama-server \
209
- -m Qwen3.8-27B-GSQ-RCO-IQ2_XS.gguf \
210
- --ctx-size 131072 \
211
- --parallel 1 \
212
- --flash-attn on \
213
- --cache-type-k q4_0 \
214
- --cache-type-v q4_0 \
215
- --n-gpu-layers all \
216
- --slot-save-path E:/LLM/kv-cache \
217
- --host 127.0.0.1 \
218
- --port 8080
219
- ```
220
-
221
- ---
222
-
223
- # 7. Why the MVP should use one slot
224
-
225
- With:
226
-
227
- ```text
228
- --parallel 1
229
- ```
230
-
231
- there is one inference slot:
232
-
233
- ```text
234
- slot 0
235
- ```
236
-
237
- Therefore the plugin doesn't need to modify OpenAI request bodies.
238
-
239
- It can treat slot `0` as an expensive hardware-backed working register:
240
-
241
- ```text
242
- ┌──────────────┐
243
- session A ─────► │ │
244
- session B ─────► │ slot 0 │
245
- session C ─────► │ │
246
- └──────────────┘
247
- ↕
248
- snapshots
249
- ```
250
-
251
- Switching A → B becomes:
252
-
253
- ```text
254
- save A if dirty
255
- restore B
256
- run B
257
- ```
258
-
259
- Switching B → A:
260
-
261
- ```text
262
- save B if dirty
263
- restore A
264
- run A
265
- ```
266
-
267
- This makes an excellent first implementation because:
268
-
269
- - there is no slot allocator;
270
- - there are no slot races;
271
- - no llama-specific fields have to enter DSH `GenerateOptions`;
272
- - the existing DSH LLM adapter can remain untouched.
273
-
274
- ---
275
-
276
- # 8. Multi-slot limitation
277
-
278
- llama.cpp supports specifying:
279
-
280
- ```json
281
- {
282
- "id_slot": 2
283
- }
284
- ```
285
-
286
- on inference requests, allowing a client to explicitly select a slot.
287
-
288
- However DSH's `GenerateOptions` deliberately contains provider-neutral fields and loop requests are immutable. `llm/stream` middleware therefore cannot safely do:
289
-
290
- ```ts
291
- options.id_slot = 2
292
- ```
293
-
294
- or otherwise inject arbitrary llama.cpp transport fields.
295
-
296
- Consequently:
297
-
298
- ```text
299
- v0.1:
300
- parallel=1
301
- generic middleware
302
- slot 0
303
-
304
- v0.2+:
305
- managed multi-slot mode
306
- custom llama.cpp transport adapter
307
- ```
308
-
309
- This separation should be intentional.
310
-
311
- ---
312
-
313
- # 9. High-level architecture
314
-
315
- ```text
316
- ┌──────────────────────────────────────────────────────┐
317
- │ DeepSeek Harness │
318
- │ │
319
- │ ctx.sessions │
320
- │ │ │
321
- │ ├──── session lifecycle ───────┐ │
322
- │ │ │ │
323
- │ ctx.llm │ │
324
- │ │ │ │
325
- │ └──── llm/stream ──────────────┤ │
326
- │ ▼ │
327
- │ KvPersistService │
328
- │ / │ \ │
329
- │ / │ \ │
330
- │ Coordinator Metadata Metrics │
331
- │ │ │
332
- │ ▼ │
333
- │ KvBackend interface │
334
- │ │ │
335
- │ ▼ │
336
- │ LlamaCppBackend │
337
- └───────────────────────┼──────────────────────────────┘
338
- │ HTTP
339
- ▼
340
- ┌───────────────┐
341
- │ llama-server │
342
- │ │
343
- │ slot 0 │
344
- │ /slots API │
345
- └───────┬───────┘
346
- │
347
- ▼
348
- --slot-save-path
349
- │
350
- ┌────────────────┼────────────────┐
351
- ▼ ▼ ▼
352
- session-A.bin session-B.bin session-C.bin
353
- ```
354
-
355
- ---
356
-
357
- # 10. Package layout
358
-
359
- Recommended repository structure:
360
-
361
- ```text
362
- dsh-kv-persist/
363
- ├─ src/
364
- │ ├─ index.ts
365
- │ ├─ config.ts
366
- │ ├─ service.ts
367
- │ │
368
- │ ├─ coordinator/
369
- │ │ ├─ coordinator.ts
370
- │ │ ├─ slot-lease.ts
371
- │ │ ├─ state-machine.ts
372
- │ │ └─ checkpoint-policy.ts
373
- │ │
374
- │ ├─ backends/
375
- │ │ ├─ types.ts
376
- │ │ └─ llama-cpp/
377
- │ │ ├─ backend.ts
378
- │ │ ├─ client.ts
379
- │ │ ├─ discovery.ts
380
- │ │ ├─ compatibility.ts
381
- │ │ └─ types.ts
382
- │ │
383
- │ ├─ snapshots/
384
- │ │ ├─ manifest.ts
385
- │ │ ├─ naming.ts
386
- │ │ ├─ fingerprint.ts
387
- │ │ └─ repository.ts
388
- │ │
389
- │ ├─ observability/
390
- │ │ ├─ metrics.ts
391
- │ │ └─ diagnostics.ts
392
- │ │
393
- │ └─ errors.ts
394
- │
395
- ├─ test/
396
- │ ├─ unit/
397
- │ ├─ integration/
398
- │ └─ fixtures/
399
- │
400
- ├─ cordis.patch.yml
401
- ├─ package.json
402
- ├─ tsconfig.json
403
- ├─ README.md
404
- └─ SPEC.md
405
- ```
406
-
407
- Do not put the entire implementation into `index.ts`.
408
-
409
- ---
410
-
411
- # 11. Public service
412
-
413
- The plugin SHOULD expose:
414
-
415
- ```ts
416
- ctx.kvPersist
417
- ```
418
-
419
- through a Cordis `Service`.
420
-
421
- Approximate API:
422
-
423
- ```ts
424
- interface KvPersistService {
425
- status(): Promise<KvPersistStatus>
426
-
427
- getSessionState(
428
- sessionId: string,
429
- ): Promise<SessionKvState | undefined>
430
-
431
- save(
432
- sessionId: string,
433
- options?: SaveOptions,
434
- ): Promise<SnapshotResult>
435
-
436
- restore(
437
- sessionId: string,
438
- options?: RestoreOptions,
439
- ): Promise<RestoreResult>
440
-
441
- invalidate(
442
- sessionId: string,
443
- reason?: string,
444
- ): Promise<void>
445
-
446
- purge(
447
- sessionId: string,
448
- ): Promise<void>
449
-
450
- flush(): Promise<void>
451
- }
452
- ```
453
-
454
- The service makes future UI/CLI plugins possible without depending directly on llama.cpp.
455
-
456
- ---
457
-
458
- # 12. Backend abstraction
459
-
460
- Core code must NOT contain direct `/slots` HTTP calls.
461
-
462
- Define:
463
-
464
- ```ts
465
- interface KvPersistenceBackend {
466
- readonly kind: string
467
-
468
- probe(signal?: AbortSignal): Promise<BackendCapabilities>
469
-
470
- inspectSlots(
471
- signal?: AbortSignal,
472
- ): Promise<BackendSlot[]>
473
-
474
- saveSlot(
475
- slotId: number,
476
- snapshotKey: string,
477
- signal?: AbortSignal,
478
- ): Promise<BackendSaveResult>
479
-
480
- restoreSlot(
481
- slotId: number,
482
- snapshotKey: string,
483
- signal?: AbortSignal,
484
- ): Promise<BackendRestoreResult>
485
-
486
- eraseSlot(
487
- slotId: number,
488
- signal?: AbortSignal,
489
- ): Promise<BackendEraseResult>
490
- }
491
- ```
492
-
493
- Initial implementation:
494
-
495
- ```text
496
- KvPersistenceBackend
497
- └── LlamaCppBackend
498
- ```
499
-
500
- Potential future implementations:
501
-
502
- ```text
503
- KvPersistenceBackend
504
- ├── LlamaCppBackend
505
- ├── SglangBackend
506
- ├── VllmBackend
507
- └── CustomGatewayBackend
508
- ```
509
-
510
- ---
511
-
512
- # 13. Snapshot identity
513
-
514
- A snapshot must never be identified by `sessionId` alone.
515
-
516
- Conceptual key:
517
-
518
- ```text
519
- SnapshotKey =
520
- server instance
521
- + provider
522
- + model
523
- + DSH session
524
- + compatibility generation
525
- ```
526
-
527
- Example:
528
-
529
- ```ts
530
- interface SnapshotIdentity {
531
- sessionId: string
532
-
533
- provider: string
534
- model: string
535
-
536
- backend: 'llama.cpp'
537
-
538
- serverInstanceKey: string
539
- modelFingerprint: string
540
-
541
- compatibilityVersion: number
542
- }
543
- ```
544
-
545
- ---
546
-
547
- # 14. Snapshot metadata
548
-
549
- The plugin keeps its own small metadata record independently from the potentially huge `.bin` file.
550
-
551
- Example:
552
-
553
- ```json
554
- {
555
- "schemaVersion": 1,
556
-
557
- "sessionId": "01991d...",
558
- "sessionSeq": 341,
559
-
560
- "provider": "local-qwen",
561
- "model": "Qwen3.8-27B-GSQ-RCO-IQ2_XS.gguf",
562
-
563
- "backend": "llama.cpp",
564
- "serverInstanceKey": "local-3060",
565
- "serverEndpointHash": "sha256:...",
566
-
567
- "modelFingerprint": "sha256:...",
568
- "runtimeFingerprint": "sha256:...",
569
-
570
- "slotId": 0,
571
-
572
- "createdAt": "2026-08-29T20:00:00Z",
573
- "updatedAt": "2026-08-29T20:42:00Z",
574
-
575
- "tokens": 48321,
576
- "bytes": 2384203912,
577
-
578
- "snapshotFilename": "b3-c5-....bin",
579
-
580
- "state": "ready"
581
- }
582
- ```
583
-
584
- ---
585
-
586
- # 15. Runtime fingerprint
587
-
588
- Restoring arbitrary binary model state is dangerous.
589
-
590
- Snapshots should therefore have a compatibility fingerprint.
591
-
592
- Suggested fingerprint inputs:
593
-
594
- ```text
595
- backend type
596
- llama.cpp server/build version, when discoverable
597
- model identifier
598
- model file fingerprint/configured model key
599
- context size
600
- KV K type
601
- KV V type
602
- parallel/slot geometry
603
- relevant recurrent/hybrid mode
604
- LoRA configuration
605
- speculative decoding configuration
606
- plugin snapshot schema generation
607
- ```
608
-
609
- For values the server cannot expose reliably, configuration should allow an explicit:
610
-
611
- ```yaml
612
- runtimeKey: qwen38-27b-iq2xs-ctx128k-q4kv-v1
613
- ```
614
-
615
- Changing `runtimeKey` makes old snapshots invisible.
616
-
617
- This provides a simple manual escape hatch.
618
-
619
- ---
620
-
621
- # 16. Filename generation
622
-
623
- Raw session IDs or titles SHOULD NOT become filenames.
624
-
625
- Use:
626
-
627
- ```text
628
- sha256(
629
- backendInstanceKey +
630
- provider +
631
- model +
632
- sessionId
633
- )
634
- ```
635
-
636
- Example:
637
-
638
- ```text
639
- 7c/7c856dc594.........bin
640
- ```
641
-
642
- Benefits:
643
-
644
- - fixed filename length;
645
- - no unsafe characters;
646
- - no path traversal;
647
- - no leaking chat titles;
648
- - no collisions in normal operation.
649
-
650
- ---
651
-
652
- # 17. Snapshot storage
653
-
654
- Important distinction:
655
-
656
- ```text
657
- metadata storage
658
- ≠
659
- KV binary storage
660
- ```
661
-
662
- The KV binary is created by `llama-server` itself inside:
663
-
664
- ```text
665
- --slot-save-path
666
- ```
667
-
668
- The DSH plugin may not even have filesystem access to that directory.
669
-
670
- Therefore the architecture must support:
671
-
672
- ### Metadata
673
-
674
- Stored locally by the plugin:
675
-
676
- ```text
677
- <DSH home>/cache/dsh-kv-persist/
678
- ```
679
-
680
- ### Binary
681
-
682
- Stored by llama-server:
683
-
684
- ```text
685
- <LLAMA_SLOT_SAVE_PATH>/
686
- ```
687
-
688
- The plugin communicates using only the generated filename.
689
-
690
- This allows:
691
-
692
- ```text
693
- DSH in container A
694
- llama-server in container B
695
- ```
696
-
697
- provided the server owns its snapshot directory.
698
-
699
- ---
700
-
701
- # 18. Session state machine
702
-
703
- Each session can be:
704
-
705
- ```text
706
- NONE
707
- no known snapshot
708
-
709
- COLD
710
- session exists but no state loaded
711
-
712
- RESTORING
713
- snapshot restore in progress
714
-
715
- ACTIVE_CLEAN
716
- current server slot corresponds to snapshot
717
-
718
- ACTIVE_DIRTY
719
- model has advanced beyond saved snapshot
720
-
721
- SAVING
722
- persistence operation in progress
723
-
724
- SAVED
725
- latest runtime state has durable snapshot
726
-
727
- INVALID
728
- snapshot exists but is incompatible/corrupted
729
- ```
730
-
731
- Typical transitions:
732
-
733
- ```text
734
- NONE
735
- ↓ first request
736
- COLD
737
- ↓ normal prefill
738
- ACTIVE_DIRTY
739
- ↓ checkpoint
740
- SAVING
741
- ↓
742
- SAVED
743
- ```
744
-
745
- Resume:
746
-
747
- ```text
748
- SAVED
749
- ↓ session request
750
- RESTORING
751
- ↓
752
- ACTIVE_CLEAN
753
- ↓ inference
754
- ACTIVE_DIRTY
755
- ```
756
-
757
- Session switch:
758
-
759
- ```text
760
- A ACTIVE_DIRTY
761
- ↓
762
- save A
763
- ↓
764
- A SAVED
765
- ↓
766
- restore B
767
- ↓
768
- B ACTIVE_CLEAN
769
- ```
770
-
771
- ---
772
-
773
- # 19. Slot state
774
-
775
- The coordinator separately tracks the physical slot.
776
-
777
- ```ts
778
- interface ManagedSlot {
779
- id: number
780
-
781
- ownerSessionId?: string
782
-
783
- snapshotRevision?: string
784
-
785
- state:
786
- | 'unknown'
787
- | 'idle'
788
- | 'restoring'
789
- | 'ready'
790
- | 'inference'
791
- | 'dirty'
792
- | 'saving'
793
- | 'broken'
794
-
795
- lastUsedAt?: number
796
- }
797
- ```
798
-
799
- For v0.1 there is simply:
800
-
801
- ```ts
802
- slot[0]
803
- ```
804
-
805
- ---
806
-
807
- # 20. Critical concurrency rule
808
-
809
- No two operations may concurrently mutate the same llama slot.
810
-
811
- The following must share one lock:
812
-
813
- ```text
814
- restore
815
- erase
816
- inference
817
- save
818
- ```
819
-
820
- For MVP:
821
-
822
- ```ts
823
- serverMutex.runExclusive(...)
824
- ```
825
-
826
- is sufficient.
827
-
828
- Conceptually:
829
-
830
- ```text
831
- acquire slot lease
832
- ↓
833
- prepare slot
834
- ↓
835
- run inference
836
- ↓
837
- update dirty state
838
- ↓
839
- optional checkpoint
840
- ↓
841
- release slot lease
842
- ```
843
-
844
- This lock is absolutely critical.
845
-
846
- Without it:
847
-
848
- ```text
849
- session A restoring
850
- +
851
- session B starting request
852
- =
853
- undefined cache ownership
854
- ```
855
-
856
- ---
857
-
858
- # 21. Request interception
859
-
860
- The plugin listens to:
861
-
862
- ```text
863
- llm/stream
864
- ```
865
-
866
- For each request:
867
-
868
- ```ts
869
- if (!isManagedProvider(options.provider)) {
870
- return next()
871
- }
872
-
873
- if (!options.sessionId) {
874
- return handleAuxiliaryRequest(...)
875
- }
876
-
877
- return coordinator.runSessionRequest({
878
- sessionId: options.sessionId,
879
- provider: options.provider,
880
- model: options.model,
881
- purpose: options.purpose,
882
- next,
883
- })
884
- ```
885
-
886
- `llm/stream` is an appropriate seam because it wraps every actual streaming model invocation, including adapter dispatch.
887
-
888
- ---
889
-
890
- # 22. Auxiliary requests
891
-
892
- DSH may make model calls for:
893
-
894
- ```text
895
- session-title
896
- compaction
897
- other future purposes
898
- ```
899
-
900
- They can pollute slot 0.
901
-
902
- Therefore they must explicitly participate in coordination.
903
-
904
- Default:
905
-
906
- ```yaml
907
- auxiliaryRequests: isolate
908
- ```
909
-
910
- In single-slot mode:
911
-
912
- ```text
913
- main session dirty
914
- ↓
915
- save current main session
916
- ↓
917
- auxiliary request
918
- ↓
919
- slot becomes unowned
920
- ↓
921
- next main request restores its snapshot
922
- ```
923
-
924
- Do NOT accidentally assign an auxiliary request to the currently active DSH session snapshot.
925
-
926
- `GenerateOptions.purpose` exists specifically to identify current auxiliary call categories.
927
-
928
- ---
929
-
930
- # 23. Restore algorithm
931
-
932
- Pseudo-code:
933
-
934
- ```ts
935
- async function prepareSession(sessionId, route) {
936
- const slot = slot0
937
-
938
- if (
939
- slot.ownerSessionId === sessionId &&
940
- slot.state !== 'broken'
941
- ) {
942
- return { kind: 'already-active' }
943
- }
944
-
945
- if (slot.ownerSessionId) {
946
- await checkpointIfNeeded(slot.ownerSessionId)
947
- }
948
-
949
- const snapshot = await repository.findCompatible(
950
- sessionId,
951
- route,
952
- )
953
-
954
- if (!snapshot) {
955
- await backend.eraseSlot(slot.id)
956
-
957
- slot.ownerSessionId = sessionId
958
- slot.state = 'idle'
959
-
960
- return { kind: 'cold' }
961
- }
962
-
963
- try {
964
- const result = await backend.restoreSlot(
965
- slot.id,
966
- snapshot.filename,
967
- )
968
-
969
- verifyRestore(result, snapshot)
970
-
971
- slot.ownerSessionId = sessionId
972
- slot.state = 'ready'
973
-
974
- return {
975
- kind: 'restored',
976
- tokens: result.tokens,
977
- bytes: result.bytes,
978
- durationMs: result.durationMs,
979
- }
980
- } catch (error) {
981
- markSnapshotInvalid(snapshot, error)
982
-
983
- await backend.eraseSlot(slot.id)
984
-
985
- slot.ownerSessionId = sessionId
986
- slot.state = 'idle'
987
-
988
- return {
989
- kind: 'cold-fallback',
990
- error,
991
- }
992
- }
993
- }
994
- ```
995
-
996
- The key policy:
997
-
998
- > Restore failure is never fatal to ordinary inference unless strict mode is explicitly enabled.
999
-
1000
- ---
1001
-
1002
- # 24. Post-restore validation
1003
-
1004
- An HTTP `200` from llama.cpp is not enough.
1005
-
1006
- The plugin should validate:
1007
-
1008
- ```text
1009
- n_restored > 0
1010
- expected snapshot existed
1011
- slot endpoint remains healthy
1012
- reported token count is plausible
1013
- ```
1014
-
1015
- Optionally inspect:
1016
-
1017
- ```text
1018
- GET /slots
1019
- ```
1020
-
1021
- after restore.
1022
-
1023
- For hybrid/recurrent models, introduce:
1024
-
1025
- ```yaml
1026
- restoreVerification: strict
1027
- ```
1028
-
1029
- which can require slot state to reflect restored token count before considering restore successful.
1030
-
1031
- This is useful because persistent state support can vary by llama.cpp model architecture/build.
1032
-
1033
- ---
1034
-
1035
- # 25. Inference algorithm
1036
-
1037
- After preparing the slot:
1038
-
1039
- ```text
1040
- slot owner = session
1041
- ↓
1042
- call downstream LLM adapter
1043
- ↓
1044
- consume stream normally
1045
- ↓
1046
- terminal successful finish
1047
- ↓
1048
- mark slot dirty
1049
- ```
1050
-
1051
- Important:
1052
-
1053
- The plugin MUST preserve streaming.
1054
-
1055
- It must not buffer the full model response merely to implement persistence.
1056
-
1057
- Conceptually:
1058
-
1059
- ```ts
1060
- const downstream = next()
1061
-
1062
- return async function* () {
1063
- let succeeded = false
1064
-
1065
- try {
1066
- for await (const chunk of downstream) {
1067
- if (
1068
- chunk.type === 'finish' &&
1069
- chunk.reason.kind !== 'error' &&
1070
- chunk.reason.kind !== 'aborted'
1071
- ) {
1072
- succeeded = true
1073
- }
1074
-
1075
- yield chunk
1076
- }
1077
- } finally {
1078
- if (succeeded) {
1079
- markDirty(sessionId)
1080
- }
1081
- }
1082
- }
1083
- ```
1084
-
1085
- Care must be taken with Cordis waterfall semantics: resolve `next()` at the correct point and don't accidentally call a consumed waterfall continuation lazily.
1086
-
1087
- ---
1088
-
1089
- # 26. Checkpoint strategy
1090
-
1091
- Saving a multi-gigabyte KV file after every generation can destroy the performance win.
1092
-
1093
- Therefore checkpointing must be configurable.
1094
-
1095
- Supported strategies:
1096
-
1097
- ```text
1098
- switch
1099
- turn
1100
- step
1101
- idle
1102
- shutdown
1103
- manual
1104
- ```
1105
-
1106
- ### `switch`
1107
-
1108
- Save only when the physical slot must be reassigned.
1109
-
1110
- This should be the primary default.
1111
-
1112
- Example:
1113
-
1114
- ```text
1115
- A → A → A → A
1116
- ```
1117
-
1118
- No disk writes.
1119
-
1120
- Then:
1121
-
1122
- ```text
1123
- A → B
1124
- ```
1125
-
1126
- causes one save of A.
1127
-
1128
- ### `turn`
1129
-
1130
- Persist after every completed user turn.
1131
-
1132
- Higher durability, much more disk I/O.
1133
-
1134
- ### `step`
1135
-
1136
- Persist after every successful agent model step.
1137
-
1138
- Useful for crash-sensitive long autonomous jobs.
1139
-
1140
- Potentially very expensive.
1141
-
1142
- ### `idle`
1143
-
1144
- Save dirty state after N seconds without another request.
1145
-
1146
- Recommended:
1147
-
1148
- ```yaml
1149
- idleCheckpointMs: 30000
1150
- ```
1151
-
1152
- ### `shutdown`
1153
-
1154
- Attempt a final checkpoint during normal Cordis disposal.
1155
-
1156
- ### Recommended default
1157
-
1158
- ```text
1159
- switch + idle + shutdown
1160
- ```
1161
-
1162
- This provides a good balance between speed and durability.
1163
-
1164
- ---
1165
-
1166
- # 27. Checkpoint coalescing
1167
-
1168
- Multiple save requests for the same revision should collapse.
1169
-
1170
- Example:
1171
-
1172
- ```text
1173
- turn/end
1174
- idle timer
1175
- session/flush
1176
- plugin shutdown
1177
- ```
1178
-
1179
- may all happen near one another.
1180
-
1181
- Use:
1182
-
1183
- ```ts
1184
- interface SaveGeneration {
1185
- dirtyRevision: number
1186
- persistedRevision: number
1187
- inFlight?: Promise<SnapshotResult>
1188
- }
1189
- ```
1190
-
1191
- If:
1192
-
1193
- ```text
1194
- dirtyRevision === persistedRevision
1195
- ```
1196
-
1197
- do nothing.
1198
-
1199
- If a save is already running for the current revision:
1200
-
1201
- ```text
1202
- await existing save
1203
- ```
1204
-
1205
- rather than starting another.
1206
-
1207
- ---
1208
-
1209
- # 28. Dirty generations
1210
-
1211
- Never use a single boolean if avoidable.
1212
-
1213
- Use monotonic generations:
1214
-
1215
- ```text
1216
- dirtyRevision = 31
1217
- savedRevision = 30
1218
- ```
1219
-
1220
- Then:
1221
-
1222
- ```text
1223
- dirtyRevision > savedRevision
1224
- ```
1225
-
1226
- means dirty.
1227
-
1228
- If inference completes while a save is in progress:
1229
-
1230
- ```text
1231
- save revision 30
1232
- new inference → revision 31
1233
- save finishes
1234
- savedRevision = 30
1235
- ```
1236
-
1237
- and state correctly remains dirty.
1238
-
1239
- ---
1240
-
1241
- # 29. Session sequence tracking
1242
-
1243
- Snapshot metadata SHOULD include the latest known DSH:
1244
-
1245
- ```text
1246
- session.seq
1247
- ```
1248
-
1249
- Example:
1250
-
1251
- ```json
1252
- {
1253
- "sessionSeq": 482
1254
- }
1255
- ```
1256
-
1257
- This does NOT mean KV state is equivalent to every event through seq 482.
1258
-
1259
- It is primarily:
1260
-
1261
- - diagnostic metadata;
1262
- - freshness information;
1263
- - useful for invalidation;
1264
- - useful for future exact request fingerprinting.
1265
-
1266
- ---
1267
-
1268
- # 30. Prompt/request fingerprints
1269
-
1270
- Future versions SHOULD record a fingerprint of the request that produced the snapshot.
1271
-
1272
- Conceptually:
1273
-
1274
- ```text
1275
- hash(
1276
- provider
1277
- model
1278
- system prompt
1279
- normalized messages
1280
- tool schemas
1281
- relevant template configuration
1282
- )
1283
- ```
1284
-
1285
- However this SHOULD NOT be mandatory in v0.1.
1286
-
1287
- Why?
1288
-
1289
- Because llama.cpp itself compares the restored prompt cache with the next incoming prompt and can reuse the common prefix while processing the changed suffix.
1290
-
1291
- The plugin mainly needs to prevent gross incompatibility.
1292
-
1293
- ---
1294
-
1295
- # 31. Snapshot invalidation
1296
-
1297
- Snapshots must become invalid when any important runtime identity changes.
1298
-
1299
- Examples:
1300
-
1301
- ```text
1302
- model changed
1303
- model GGUF replaced
1304
- KV layout changed
1305
- runtimeKey changed
1306
- llama backend changed
1307
- LoRA changed
1308
- server instance changed incompatibly
1309
- snapshot restore failed
1310
- manifest malformed
1311
- snapshot file missing
1312
- ```
1313
-
1314
- Invalidation must NOT delete data immediately.
1315
-
1316
- Prefer:
1317
-
1318
- ```text
1319
- state = invalid
1320
- reason = MODEL_FINGERPRINT_CHANGED
1321
- ```
1322
-
1323
- Cleanup can happen separately.
1324
-
1325
- ---
1326
-
1327
- # 32. Failure policy
1328
-
1329
- Persistence is an optimization.
1330
-
1331
- Default failure behavior:
1332
-
1333
- ```text
1334
- save failed
1335
- → log warning
1336
- → continue session
1337
-
1338
- restore failed
1339
- → mark snapshot invalid
1340
- → erase slot
1341
- → cold prefill
1342
- → continue session
1343
-
1344
- server /slots unavailable
1345
- → disable persistence temporarily
1346
- → continue ordinary inference
1347
- ```
1348
-
1349
- Only configuration:
1350
-
1351
- ```yaml
1352
- strict: true
1353
- ```
1354
-
1355
- should turn persistence failure into request failure.
1356
-
1357
- Default:
1358
-
1359
- ```yaml
1360
- strict: false
1361
- ```
1362
-
1363
- ---
1364
-
1365
- # 33. Circuit breaker
1366
-
1367
- If llama slot persistence is broken, we don't want every request to waste seconds retrying it.
1368
-
1369
- Backend state:
1370
-
1371
- ```text
1372
- HEALTHY
1373
- DEGRADED
1374
- OPEN
1375
- HALF_OPEN
1376
- ```
1377
-
1378
- Suggested behavior:
1379
-
1380
- ```text
1381
- 3 consecutive persistence failures
1382
- → disable save/restore for 60s
1383
-
1384
- after 60s
1385
- → probe
1386
-
1387
- probe successful
1388
- → resume
1389
-
1390
- probe fails
1391
- → backoff
1392
- ```
1393
-
1394
- Inference itself remains active.
1395
-
1396
- ---
1397
-
1398
- # 34. Backend health probe
1399
-
1400
- At plugin initialization:
1401
-
1402
- ```text
1403
- GET /slots
1404
- ```
1405
-
1406
- Validate:
1407
-
1408
- ```text
1409
- server reachable
1410
- slots endpoint available
1411
- expected slot count
1412
- slot 0 exists for single-slot mode
1413
- ```
1414
-
1415
- Optionally test persistence capability using a non-destructive mechanism where possible.
1416
-
1417
- The plugin SHOULD clearly distinguish:
1418
-
1419
- ```text
1420
- LLM endpoint alive
1421
- ```
1422
-
1423
- from:
1424
-
1425
- ```text
1426
- KV persistence supported
1427
- ```
1428
-
1429
- ---
1430
-
1431
- # 35. Configuration
1432
-
1433
- Initial configuration proposal:
1434
-
1435
- ```yaml
1436
- kv-persist:
1437
- enabled: true
1438
-
1439
- backend:
1440
- type: llama.cpp
1441
- baseURL: http://127.0.0.1:8080
1442
-
1443
- providers:
1444
- - local-qwen
1445
-
1446
- mode: single-slot
1447
- slotId: 0
1448
-
1449
- runtimeKey: qwen38-27b-iq2xs-ctx128k-q4kv
1450
-
1451
- checkpoint:
1452
- onSwitch: true
1453
- onShutdown: true
1454
- onSessionFlush: true
1455
-
1456
- idleMs: 30000
1457
-
1458
- onTurnEnd: false
1459
- onStepEnd: false
1460
-
1461
- restore:
1462
- enabled: true
1463
- verify: true
1464
-
1465
- failure:
1466
- strict: false
1467
- maxConsecutiveFailures: 3
1468
- cooldownMs: 60000
1469
-
1470
- metadata:
1471
- path: ${DSH_HOME}/cache/dsh-kv-persist
1472
-
1473
- logging:
1474
- level: info
1475
- ```
1476
-
1477
- ---
1478
-
1479
- # 36. Config schema
1480
-
1481
- Using Schemastery-style validation:
1482
-
1483
- ```ts
1484
- interface Config {
1485
- enabled?: boolean
1486
-
1487
- backend: {
1488
- type: 'llama.cpp'
1489
- baseURL: string
1490
- apiKey?: string
1491
- requestTimeoutMs?: number
1492
- }
1493
-
1494
- providers?: string[]
1495
-
1496
- mode?: 'single-slot' | 'managed-slots'
1497
-
1498
- slotId?: number
1499
-
1500
- runtimeKey?: string
1501
-
1502
- checkpoint?: {
1503
- onSwitch?: boolean
1504
- onShutdown?: boolean
1505
- onSessionFlush?: boolean
1506
-
1507
- idleMs?: number
1508
-
1509
- onTurnEnd?: boolean
1510
- onStepEnd?: boolean
1511
- }
1512
-
1513
- restore?: {
1514
- enabled?: boolean
1515
- verify?: boolean
1516
- }
1517
-
1518
- failure?: {
1519
- strict?: boolean
1520
- maxConsecutiveFailures?: number
1521
- cooldownMs?: number
1522
- }
1523
-
1524
- metadata?: {
1525
- path?: string
1526
- }
1527
- }
1528
- ```
1529
-
1530
- Recommended defaults:
1531
-
1532
- ```text
1533
- enabled = true
1534
-
1535
- mode = single-slot
1536
- slotId = 0
1537
-
1538
- onSwitch = true
1539
- onShutdown = true
1540
- onSessionFlush = true
1541
- idleMs = 30000
1542
-
1543
- onTurnEnd = false
1544
- onStepEnd = false
1545
-
1546
- restore.enabled = true
1547
- restore.verify = true
1548
-
1549
- strict = false
1550
- ```
1551
-
1552
- ---
1553
-
1554
- # 37. Automatic provider filtering
1555
-
1556
- The plugin MUST NOT touch cloud providers by default.
1557
-
1558
- Example:
1559
-
1560
- ```yaml
1561
- providers:
1562
- - local-qwen
1563
- - local-coder
1564
- ```
1565
-
1566
- A request to:
1567
-
1568
- ```text
1569
- deepseek
1570
- openai
1571
- anthropic
1572
- ```
1573
-
1574
- passes straight through.
1575
-
1576
- Future configuration may support:
1577
-
1578
- ```yaml
1579
- providers:
1580
- local-qwen:
1581
- backend: local-llama
1582
- ```
1583
-
1584
- for multiple servers.
1585
-
1586
- ---
1587
-
1588
- # 38. Multiple llama servers
1589
-
1590
- Architecture should support this eventually:
1591
-
1592
- ```yaml
1593
- servers:
1594
- rtx3060:
1595
- type: llama.cpp
1596
- baseURL: http://127.0.0.1:8080
1597
- runtimeKey: qwen38
1598
-
1599
- rtx4090:
1600
- type: llama.cpp
1601
- baseURL: http://192.168.1.42:8080
1602
- runtimeKey: qwen-coder
1603
-
1604
- routes:
1605
- local-qwen:
1606
- server: rtx3060
1607
-
1608
- local-coder:
1609
- server: rtx4090
1610
- ```
1611
-
1612
- Internally:
1613
-
1614
- ```text
1615
- Coordinator
1616
- ↓
1617
- ServerRuntime[]
1618
- ↓
1619
- independent slot pools + locks
1620
- ```
1621
-
1622
- Do not use one global mutex across different servers.
1623
-
1624
- ---
1625
-
1626
- # 39. Metadata repository
1627
-
1628
- Default path (shared DSH home convention; a non-blank `$DSH_HOME` wins,
1629
- otherwise DSH home is `~/.dsh`):
1630
-
1631
- ```text
1632
- $DSH_HOME/cache/dsh-kv-persist/
1633
- ```
1634
-
1635
- Layout:
1636
-
1637
- ```text
1638
- dsh-kv-persist/
1639
- ├─ instances/
1640
- │ └─ local-3060/
1641
- │ └─ sessions/
1642
- │ ├─ 7c856d....json
1643
- │ └─ d084fd....json
1644
- │
1645
- ├─ index.json
1646
- └─ plugin-state.json
1647
- ```
1648
-
1649
- Binary KV files SHOULD NOT be copied here automatically.
1650
-
1651
- ---
1652
-
1653
- # 40. Atomic metadata writes
1654
-
1655
- Manifest updates should use:
1656
-
1657
- ```text
1658
- write temp
1659
- fsync/close
1660
- rename
1661
- ```
1662
-
1663
- rather than overwriting JSON directly.
1664
-
1665
- Example:
1666
-
1667
- ```text
1668
- session.json.tmp
1669
- ↓
1670
- rename
1671
- ↓
1672
- session.json
1673
- ```
1674
-
1675
- A crash must not leave half-written metadata treated as valid.
1676
-
1677
- ---
1678
-
1679
- # 41. Snapshot naming generations
1680
-
1681
- Default strategy:
1682
-
1683
- ```text
1684
- one rolling snapshot per session
1685
- ```
1686
-
1687
- rather than:
1688
-
1689
- ```text
1690
- snapshot-000001.bin
1691
- snapshot-000002.bin
1692
- snapshot-000003.bin
1693
- ...
1694
- ```
1695
-
1696
- because KV snapshots can be enormous.
1697
-
1698
- Conceptually:
1699
-
1700
- ```text
1701
- <sessionHash>.bin
1702
- ```
1703
-
1704
- If llama.cpp cannot safely atomically replace an existing file in a particular build/backend, support two rotating names:
1705
-
1706
- ```text
1707
- <hash>.a.bin
1708
- <hash>.b.bin
1709
- ```
1710
-
1711
- Manifest points to the latest completed generation.
1712
-
1713
- This prevents an interrupted write from destroying the previous valid checkpoint.
1714
-
1715
- ---
1716
-
1717
- # 42. Local vs remote server cleanup
1718
-
1719
- The llama slots API manages save/restore/erase of slot state, but binary file lifecycle may not always be remotely manageable.
1720
-
1721
- Therefore define:
1722
-
1723
- ```text
1724
- snapshotBinaryManagement:
1725
- server-owned
1726
- shared-filesystem
1727
- ```
1728
-
1729
- ### `server-owned`
1730
-
1731
- Plugin only knows filenames.
1732
-
1733
- No direct delete.
1734
-
1735
- Use rolling filenames to limit growth.
1736
-
1737
- ### `shared-filesystem`
1738
-
1739
- Plugin is configured with the same physical save directory and may:
1740
-
1741
- - inspect file size;
1742
- - delete invalid snapshots;
1743
- - enforce disk quota;
1744
- - perform atomic rotation.
1745
-
1746
- This mode MUST be opt-in.
1747
-
1748
- ---
1749
-
1750
- # 43. Disk quota
1751
-
1752
- Future version:
1753
-
1754
- ```yaml
1755
- retention:
1756
- maxTotalBytes: 100GB
1757
- maxSessions: 50
1758
- maxAgeDays: 30
1759
- ```
1760
-
1761
- Eviction policy:
1762
-
1763
- ```text
1764
- invalid first
1765
- then oldest unused
1766
- then LRU
1767
- ```
1768
-
1769
- Never remove the active slot state as part of disk cleanup.
1770
-
1771
- Only stored snapshots.
1772
-
1773
- ---
1774
-
1775
- # 44. Security
1776
-
1777
- The plugin MUST assume the llama management endpoint is privileged.
1778
-
1779
- Recommended setup:
1780
-
1781
- ```text
1782
- 127.0.0.1
1783
- or
1784
- trusted private network
1785
- or
1786
- authenticated reverse proxy
1787
- ```
1788
-
1789
- Do not expose slot management APIs publicly.
1790
-
1791
- Snapshot filenames must be generated by the plugin and sanitized.
1792
-
1793
- Never accept:
1794
-
1795
- ```text
1796
- ../../foo
1797
- C:\whatever
1798
- /etc/passwd
1799
- ```
1800
-
1801
- as a raw snapshot filename.
1802
-
1803
- Only backend-generated opaque keys may reach:
1804
-
1805
- ```text
1806
- action=save
1807
- action=restore
1808
- ```
1809
-
1810
- ---
1811
-
1812
- # 45. Logging
1813
-
1814
- Recommended structured events:
1815
-
1816
- ```text
1817
- kv.backend.ready
1818
- kv.backend.unavailable
1819
-
1820
- kv.slot.acquire
1821
- kv.slot.release
1822
-
1823
- kv.session.cold
1824
- kv.session.restore.start
1825
- kv.session.restore.success
1826
- kv.session.restore.failed
1827
-
1828
- kv.session.save.start
1829
- kv.session.save.success
1830
- kv.session.save.failed
1831
-
1832
- kv.session.switch
1833
-
1834
- kv.snapshot.invalidated
1835
-
1836
- kv.persistence.circuit_open
1837
- kv.persistence.circuit_recovered
1838
- ```
1839
-
1840
- Example:
1841
-
1842
- ```text
1843
- [kv-persist] restored
1844
- session=7c856d
1845
- slot=0
1846
- tokens=48192
1847
- bytes=2.31GiB
1848
- duration=418ms
1849
- ```
1850
-
1851
- Avoid logging full session IDs at normal verbosity if unnecessary.
1852
-
1853
- Use abbreviated hashes.
1854
-
1855
- ---
1856
-
1857
- # 46. Metrics
1858
-
1859
- Expose internal counters through the service and later optional Prometheus integration:
1860
-
1861
- ```text
1862
- dsh_kv_restore_total
1863
- dsh_kv_restore_hit_total
1864
- dsh_kv_restore_miss_total
1865
- dsh_kv_restore_failure_total
1866
-
1867
- dsh_kv_save_total
1868
- dsh_kv_save_failure_total
1869
-
1870
- dsh_kv_restore_bytes_total
1871
- dsh_kv_save_bytes_total
1872
-
1873
- dsh_kv_restore_duration_ms
1874
- dsh_kv_save_duration_ms
1875
-
1876
- dsh_kv_cold_prefill_total
1877
-
1878
- dsh_kv_slot_switch_total
1879
-
1880
- dsh_kv_snapshot_count
1881
- dsh_kv_snapshot_bytes
1882
- ```
1883
-
1884
- Especially useful derived metric:
1885
-
1886
- ```text
1887
- persistent-cache restore hit rate
1888
- ```
1889
-
1890
- ---
1891
-
1892
- # 47. Diagnostics API
1893
-
1894
- `ctx.kvPersist.status()` should return something like:
1895
-
1896
- ```json
1897
- {
1898
- "enabled": true,
1899
- "backend": {
1900
- "kind": "llama.cpp",
1901
- "state": "healthy",
1902
- "endpoint": "http://127.0.0.1:8080"
1903
- },
1904
-
1905
- "mode": "single-slot",
1906
-
1907
- "slots": [
1908
- {
1909
- "id": 0,
1910
- "owner": "7c856d",
1911
- "state": "dirty"
1912
- }
1913
- ],
1914
-
1915
- "snapshots": {
1916
- "known": 14,
1917
- "valid": 13,
1918
- "invalid": 1
1919
- },
1920
-
1921
- "stats": {
1922
- "restores": 23,
1923
- "restoreHits": 21,
1924
- "coldStarts": 2,
1925
- "saves": 18
1926
- }
1927
- }
1928
- ```
1929
-
1930
- ---
1931
-
1932
- # 48. Optional CLI
1933
-
1934
- Eventually expose human-facing commands:
1935
-
1936
- ```text
1937
- dsh kv status
1938
- dsh kv list
1939
- dsh kv save <session>
1940
- dsh kv restore <session>
1941
- dsh kv invalidate <session>
1942
- dsh kv purge <session>
1943
- dsh kv gc
1944
- dsh kv doctor
1945
- ```
1946
-
1947
- `doctor` should be particularly useful.
1948
-
1949
- Example:
1950
-
1951
- ```text
1952
- $ dsh kv doctor
1953
-
1954
- Backend: llama.cpp
1955
- Endpoint: http://127.0.0.1:8080
1956
- Slots API: OK
1957
- Slots: 1
1958
- Configured mode: single-slot
1959
- Slot 0: idle
1960
- Save path capability: OK
1961
- Restore verification: OK
1962
- Hybrid model persistence: unverified
1963
- Metadata directory: writable
1964
- Result: READY
1965
- ```
1966
-
1967
- ---
1968
-
1969
- # 49. No model-facing tool by default
1970
-
1971
- Do NOT register:
1972
-
1973
- ```text
1974
- save_kv_cache
1975
- restore_kv_cache
1976
- ```
1977
-
1978
- with `ctx.tools`.
1979
-
1980
- There is almost no reason for the LLM itself to manage its infrastructure cache.
1981
-
1982
- It wastes tool schema tokens and introduces failure modes such as:
1983
-
1984
- ```text
1985
- model decides to purge its own cache
1986
- ```
1987
-
1988
- Management belongs to:
1989
-
1990
- ```text
1991
- plugin
1992
- user CLI
1993
- UI
1994
- ```
1995
-
1996
- not to the model.
1997
-
1998
- ---
1999
-
2000
- # 50. Session lifecycle integration
2001
-
2002
- Subscribe to session lifecycle for:
2003
-
2004
- ```text
2005
- session created
2006
- session flush
2007
- session disposed
2008
- turn end
2009
- ```
2010
-
2011
- Use these events as persistence hints.
2012
-
2013
- DSH explicitly exposes session persistence hooks and `session/flush` as the durability checkpoint seam.
2014
-
2015
- Recommended semantics:
2016
-
2017
- ### session created/resumed
2018
-
2019
- Do not restore immediately.
2020
-
2021
- Lazy restore on first actual LLM request.
2022
-
2023
- Reason:
2024
-
2025
- ```text
2026
- opening a chat in UI
2027
- ```
2028
-
2029
- should not evict another active KV session unless inference actually occurs.
2030
-
2031
- ### session flush
2032
-
2033
- If the session currently owns a dirty slot:
2034
-
2035
- ```text
2036
- checkpoint
2037
- ```
2038
-
2039
- ### session disposed
2040
-
2041
- If dirty:
2042
-
2043
- ```text
2044
- checkpoint if configured
2045
- ```
2046
-
2047
- then remove runtime ownership.
2048
-
2049
- ### turn/end
2050
-
2051
- Checkpoint only when:
2052
-
2053
- ```yaml
2054
- checkpoint.onTurnEnd: true
2055
- ```
2056
-
2057
- ---
2058
-
2059
- # 51. Lazy restore
2060
-
2061
- This is important enough to be explicit.
2062
-
2063
- Bad:
2064
-
2065
- ```text
2066
- user clicks session B
2067
- → immediately save A
2068
- → restore 3GB B
2069
- → user clicks session C
2070
- → immediately save B
2071
- → restore C
2072
- ```
2073
-
2074
- Good:
2075
-
2076
- ```text
2077
- user clicks session B
2078
- → nothing
2079
-
2080
- user actually sends message in B
2081
- → switch slot
2082
- ```
2083
-
2084
- Snapshot management follows inference, not UI navigation.
2085
-
2086
- ---
2087
-
2088
- # 52. Save-before-evict invariant
2089
-
2090
- Before assigning a dirty slot to another owner:
2091
-
2092
- ```text
2093
- MUST attempt save
2094
- ```
2095
-
2096
- unless configured:
2097
-
2098
- ```yaml
2099
- checkpoint.onSwitch: false
2100
- ```
2101
-
2102
- Default invariant:
2103
-
2104
- ```text
2105
- dirty A
2106
- +
2107
- need B
2108
- =
2109
- save A before erase/restore B
2110
- ```
2111
-
2112
- This is the core of session switching.
2113
-
2114
- ---
2115
-
2116
- # 53. Cold-session behavior
2117
-
2118
- If no snapshot exists:
2119
-
2120
- ```text
2121
- erase slot
2122
- assign owner
2123
- let llama.cpp process full request
2124
- mark dirty
2125
- ```
2126
-
2127
- Do not try to build a snapshot before inference.
2128
-
2129
- Snapshot will naturally be created by the next checkpoint.
2130
-
2131
- ---
2132
-
2133
- # 54. Snapshot restore and prompt divergence
2134
-
2135
- A restored KV snapshot is not assumed to perfectly equal the next DSH request.
2136
-
2137
- Example:
2138
-
2139
- Snapshot:
2140
-
2141
- ```text
2142
- system
2143
- A
2144
- assistant A
2145
- B
2146
- assistant B
2147
- ```
2148
-
2149
- Current request:
2150
-
2151
- ```text
2152
- system
2153
- A
2154
- assistant A
2155
- B
2156
- assistant B
2157
- C
2158
- ```
2159
-
2160
- Ideal outcome:
2161
-
2162
- ```text
2163
- reuse existing prefix
2164
- process only C
2165
- ```
2166
-
2167
- If the system prompt/tool schema changed:
2168
-
2169
- ```text
2170
- old prefix
2171
- ↓
2172
- divergence detected by llama.cpp
2173
- ↓
2174
- recompute changed suffix
2175
- ```
2176
-
2177
- Thus plugin-side fingerprints are mainly for runtime compatibility and diagnostics rather than replacing llama.cpp's prompt matching.
2178
-
2179
- ---
2180
-
2181
- # 55. Context compaction
2182
-
2183
- Compaction changes model-visible history substantially.
2184
-
2185
- The plugin does not need special correctness logic.
2186
-
2187
- After compaction:
2188
-
2189
- ```text
2190
- restored old snapshot
2191
- ↓
2192
- incoming compacted prompt differs
2193
- ↓
2194
- llama.cpp finds smaller common prefix
2195
- ↓
2196
- new prompt is processed
2197
- ↓
2198
- slot becomes dirty
2199
- ↓
2200
- next checkpoint replaces snapshot
2201
- ```
2202
-
2203
- However the plugin should emit diagnostics:
2204
-
2205
- ```text
2206
- large restored cache
2207
- low subsequent cache reuse
2208
- possible compaction/prompt mutation
2209
- ```
2210
-
2211
- Future versions can proactively invalidate on known compaction events.
2212
-
2213
- ---
2214
-
2215
- # 56. Provider/model changes inside a session
2216
-
2217
- DSH allows the request route to change between steps.
2218
-
2219
- Therefore one DSH session can theoretically contain:
2220
-
2221
- ```text
2222
- model A
2223
- → model B
2224
- → model A
2225
- ```
2226
-
2227
- KV identity must therefore include:
2228
-
2229
- ```text
2230
- provider + model
2231
- ```
2232
-
2233
- not only session.
2234
-
2235
- Conceptually:
2236
-
2237
- ```text
2238
- session X
2239
- ├─ qwen snapshot
2240
- └─ coder snapshot
2241
- ```
2242
-
2243
- MVP may simplify by allowing one current snapshot per:
2244
-
2245
- ```text
2246
- (session, provider, model)
2247
- ```
2248
-
2249
- ---
2250
-
2251
- # 57. LoRA and runtime mutations
2252
-
2253
- If llama-server changes:
2254
-
2255
- ```text
2256
- LoRA
2257
- model
2258
- chat template
2259
- KV representation
2260
- other state that affects serialized cache
2261
- ```
2262
-
2263
- the runtime fingerprint must change or snapshots must be invalidated.
2264
-
2265
- Never silently restore across obviously different model states.
2266
-
2267
- ---
2268
-
2269
- # 58. Plugin lifecycle
2270
-
2271
- Cordis plugins may be unloaded through configuration changes, HMR, explicit disposal, or loss of dependencies; resources external to Cordis should be tied to `ctx.effect()` and cleaned on unload.
2272
-
2273
- The plugin should therefore register:
2274
-
2275
- ```text
2276
- timers
2277
- HTTP resources
2278
- backend lifecycle
2279
- shutdown save
2280
- ```
2281
-
2282
- through proper Cordis effects.
2283
-
2284
- On dispose:
2285
-
2286
- ```text
2287
- stop accepting new persistence work
2288
- ↓
2289
- cancel idle timers
2290
- ↓
2291
- wait for/abort safe pending operations
2292
- ↓
2293
- checkpoint active dirty slot if configured
2294
- ↓
2295
- dispose service
2296
- ```
2297
-
2298
- ---
2299
-
2300
- # 59. Cancellation
2301
-
2302
- User cancellation of generation MUST NOT be blocked by a slow snapshot write.
2303
-
2304
- Inference `AbortSignal` belongs to inference.
2305
-
2306
- Persistence operations should use their own bounded timeout.
2307
-
2308
- For example:
2309
-
2310
- ```yaml
2311
- backend:
2312
- requestTimeoutMs: 15000
2313
- ```
2314
-
2315
- If save exceeds timeout:
2316
-
2317
- ```text
2318
- log
2319
- mark persistence degraded
2320
- release workflow
2321
- ```
2322
-
2323
- Don't leave the agent permanently stuck because an NVMe/cache filesystem is unhappy.
2324
-
2325
- ---
2326
-
2327
- # 60. Crash semantics
2328
-
2329
- There are three relevant crashes:
2330
-
2331
- ### DSH crashes
2332
-
2333
- llama-server remains alive.
2334
-
2335
- Active slot may still contain valid state.
2336
-
2337
- v0.1 may ignore this unsaved in-memory opportunity and restore the last durable snapshot.
2338
-
2339
- Future optimization:
2340
-
2341
- ```text
2342
- inspect current slot metadata
2343
- re-associate if ownership can be proven
2344
- ```
2345
-
2346
- ### llama-server crashes
2347
-
2348
- Only durable snapshots survive.
2349
-
2350
- After restart:
2351
-
2352
- ```text
2353
- probe
2354
- restore snapshot
2355
- ```
2356
-
2357
- ### Machine crashes during snapshot save
2358
-
2359
- Manifest must continue referencing the previous known-good generation.
2360
-
2361
- This is why two-file rotation may eventually be useful.
2362
-
2363
- ---
2364
-
2365
- # 61. Hybrid/recurrent model compatibility
2366
-
2367
- Qwen3.x hybrid/recurrent architectures make this especially important.
2368
-
2369
- Define compatibility states:
2370
-
2371
- ```text
2372
- supported
2373
- experimental
2374
- broken
2375
- unknown
2376
- ```
2377
-
2378
- Example metadata:
2379
-
2380
- ```json
2381
- {
2382
- "persistenceCompatibility": "experimental"
2383
- }
2384
- ```
2385
-
2386
- `dsh kv doctor` can perform an opt-in verification:
2387
-
2388
- ```text
2389
- 1. cold prompt
2390
- 2. save
2391
- 3. erase
2392
- 4. restore
2393
- 5. inspect
2394
- 6. identical prompt
2395
- 7. confirm cache reuse
2396
- ```
2397
-
2398
- An even stronger test:
2399
-
2400
- ```text
2401
- save
2402
- restart server manually
2403
- restore
2404
- same prompt
2405
- verify hit
2406
- ```
2407
-
2408
- The plugin should never assume that receiving `n_restored` means a particular model/build definitely restored usable recurrent state.
2409
-
2410
- ---
2411
-
2412
- # 62. Compatibility database
2413
-
2414
- Future versions may contain small rules:
2415
-
2416
- ```ts
2417
- interface CompatibilityRule {
2418
- backend: 'llama.cpp'
2419
- architecture?: string
2420
- minBuild?: number
2421
- maxBuild?: number
2422
- status: 'supported' | 'experimental' | 'broken'
2423
- note?: string
2424
- }
2425
- ```
2426
-
2427
- But avoid hardcoding large brittle version tables initially.
2428
-
2429
- Prefer runtime verification.
2430
-
2431
- ---
2432
-
2433
- # 63. Multi-slot architecture
2434
-
2435
- After MVP, support:
2436
-
2437
- ```text
2438
- --parallel N
2439
- ```
2440
-
2441
- with a real slot pool.
2442
-
2443
- Example N=4:
2444
-
2445
- ```text
2446
- slot 0 → session A
2447
- slot 1 → session B
2448
- slot 2 → session C
2449
- slot 3 → session D
2450
- ```
2451
-
2452
- Session E arrives:
2453
-
2454
- ```text
2455
- choose LRU clean/dirty slot
2456
- ↓
2457
- save old owner if dirty
2458
- ↓
2459
- restore E
2460
- ↓
2461
- bind slot to E
2462
- ```
2463
-
2464
- ---
2465
-
2466
- # 64. Multi-slot slot selection
2467
-
2468
- Selection order:
2469
-
2470
- ```text
2471
- 1. slot already owned by requested session
2472
- 2. empty slot
2473
- 3. clean least-recently-used slot
2474
- 4. dirty least-recently-used slot
2475
- ```
2476
-
2477
- Evicting a dirty slot requires save.
2478
-
2479
- Pseudo-code:
2480
-
2481
- ```ts
2482
- function selectSlot(sessionId) {
2483
- return (
2484
- ownedBy(sessionId) ??
2485
- emptySlot() ??
2486
- lruClean() ??
2487
- lruDirty()
2488
- )
2489
- }
2490
- ```
2491
-
2492
- ---
2493
-
2494
- # 65. Multi-slot transport
2495
-
2496
- For managed multi-slot support, introduce:
2497
-
2498
- ```text
2499
- dsh-llama.cpp adapter
2500
- ```
2501
-
2502
- or a transport backend capable of injecting:
2503
-
2504
- ```json
2505
- {
2506
- "id_slot": 2,
2507
- "cache_prompt": true
2508
- }
2509
- ```
2510
-
2511
- into llama-server requests.
2512
-
2513
- Potential package architecture:
2514
-
2515
- ```text
2516
- dsh-kv-persist
2517
- └─ coordination/service
2518
-
2519
- dsh-llm-llama-cpp
2520
- └─ llama-specific transport
2521
- ```
2522
-
2523
- The two can communicate through:
2524
-
2525
- ```text
2526
- ctx.kvPersist
2527
- ```
2528
-
2529
- This is preferable to making the persistence plugin own all OpenAI serialization logic.
2530
-
2531
- ---
2532
-
2533
- # 66. Alternative multi-slot sidecar
2534
-
2535
- Another possible backend:
2536
-
2537
- ```text
2538
- DSH
2539
- ↓
2540
- normal OpenAI adapter
2541
- ↓
2542
- local KV-aware reverse proxy
2543
- ↓
2544
- llama-server
2545
- ```
2546
-
2547
- Proxy receives a hidden session identifier and injects:
2548
-
2549
- ```text
2550
- id_slot
2551
- ```
2552
-
2553
- This is useful if DSH's adapter layer remains intentionally provider-neutral.
2554
-
2555
- However a native adapter is probably cleaner.
2556
-
2557
- ---
2558
-
2559
- # 67. Future upstream opportunity
2560
-
2561
- Potential DSH upstream proposal:
2562
-
2563
- ```ts
2564
- GenerateOptions.transportMetadata?
2565
- ```
2566
-
2567
- or an adapter-private request context carrying:
2568
-
2569
- ```text
2570
- sessionId
2571
- ```
2572
-
2573
- all the way into adapters.
2574
-
2575
- DSH already provides `sessionId` as model-hidden routing metadata, so a llama-specific adapter can naturally use that for slot assignment without exposing it to the model.
2576
-
2577
- Avoid adding llama-specific fields to core DSH vocabulary.
2578
-
2579
- ---
2580
-
2581
- # 68. MVP scope — v0.1
2582
-
2583
- The first usable release should contain only:
2584
-
2585
- ```text
2586
- llama.cpp backend
2587
- single server
2588
- single slot
2589
- explicit managed provider list
2590
- sessionId → snapshot mapping
2591
- save on switch
2592
- idle save
2593
- save on shutdown/flush
2594
- lazy restore
2595
- restore fallback
2596
- metadata manifests
2597
- global slot mutex
2598
- logging
2599
- status API
2600
- basic doctor/probe
2601
- ```
2602
-
2603
- Explicitly NOT in v0.1:
2604
-
2605
- ```text
2606
- multi-slot
2607
- UI
2608
- disk GC
2609
- multiple servers
2610
- snapshot migration
2611
- Prometheus
2612
- custom adapter
2613
- automatic server startup
2614
- ```
2615
-
2616
- Keep v0.1 small enough to actually ship.
2617
-
2618
- ---
2619
-
2620
- # 69. v0.1 request flow
2621
-
2622
- Example: first ever session A request.
2623
-
2624
- ```text
2625
- DSH llm/stream(A)
2626
- ↓
2627
- plugin sees managed provider
2628
- ↓
2629
- acquire slot 0
2630
- ↓
2631
- no current owner
2632
- ↓
2633
- no snapshot A
2634
- ↓
2635
- erase slot
2636
- ↓
2637
- owner = A
2638
- ↓
2639
- next()
2640
- ↓
2641
- llama processes prompt
2642
- ↓
2643
- stream response
2644
- ↓
2645
- mark A dirty
2646
- ↓
2647
- release
2648
- ```
2649
-
2650
- Second request A:
2651
-
2652
- ```text
2653
- llm/stream(A)
2654
- ↓
2655
- slot already belongs to A
2656
- ↓
2657
- no save/restore
2658
- ↓
2659
- next()
2660
- ↓
2661
- normal in-memory cache hit
2662
- ```
2663
-
2664
- This is very important:
2665
-
2666
- > The plugin must not save/restore when the requested session is already resident.
2667
-
2668
- Persistence must not make the happy path slower.
2669
-
2670
- ---
2671
-
2672
- # 70. v0.1 session switch
2673
-
2674
- A → B:
2675
-
2676
- ```text
2677
- request B
2678
- ↓
2679
- acquire slot
2680
- ↓
2681
- slot owner = A, A dirty
2682
- ↓
2683
- save A.bin
2684
- ↓
2685
- mark A saved
2686
- ↓
2687
- find B snapshot
2688
- ↓
2689
- restore B.bin
2690
- ↓
2691
- owner = B
2692
- ↓
2693
- request B
2694
- ```
2695
-
2696
- B → A:
2697
-
2698
- ```text
2699
- save B if dirty
2700
- restore A
2701
- run A
2702
- ```
2703
-
2704
- ---
2705
-
2706
- # 71. v0.1 idle save
2707
-
2708
- After request A finishes:
2709
-
2710
- ```text
2711
- A dirty
2712
- ↓
2713
- start/reset 30s timer
2714
- ```
2715
-
2716
- If another A request comes within 30 seconds:
2717
-
2718
- ```text
2719
- cancel/reset timer
2720
- ```
2721
-
2722
- If idle timer fires:
2723
-
2724
- ```text
2725
- acquire slot
2726
- ↓
2727
- confirm slot still owned by A
2728
- ↓
2729
- confirm same dirty generation
2730
- ↓
2731
- save
2732
- ↓
2733
- release
2734
- ```
2735
-
2736
- Never save based purely on an old timer callback without rechecking ownership.
2737
-
2738
- ---
2739
-
2740
- # 72. v0.1 auxiliary request flow
2741
-
2742
- Suppose A is active and DSH starts session-title generation.
2743
-
2744
- ```text
2745
- A dirty
2746
- ↓
2747
- aux request detected
2748
- ↓
2749
- save A
2750
- ↓
2751
- slot owner cleared
2752
- ↓
2753
- run title request
2754
- ↓
2755
- slot owner = auxiliary/unowned
2756
- ```
2757
-
2758
- Next A request:
2759
-
2760
- ```text
2761
- restore A
2762
- ```
2763
-
2764
- This is slower than having a separate aux slot but correct.
2765
-
2766
- v0.2 multi-slot can reserve:
2767
-
2768
- ```text
2769
- slot N-1 = auxiliary
2770
- ```
2771
-
2772
- ---
2773
-
2774
- # 73. Performance targets
2775
-
2776
- MVP should add almost zero overhead when a session remains resident.
2777
-
2778
- Resident request overhead target:
2779
-
2780
- ```text
2781
- < 1 ms plugin CPU overhead
2782
- 0 disk I/O
2783
- 0 extra llama management calls
2784
- ```
2785
-
2786
- Session restore cost is dominated by backend I/O.
2787
-
2788
- The plugin should record:
2789
-
2790
- ```text
2791
- save latency
2792
- restore latency
2793
- bytes
2794
- tokens
2795
- ```
2796
-
2797
- so the user can compare:
2798
-
2799
- ```text
2800
- cold prefill time
2801
- vs
2802
- restore time
2803
- ```
2804
-
2805
- ---
2806
-
2807
- # 74. Acceptance criteria for v0.1
2808
-
2809
- Release `0.1.0` is acceptable when all of the following work:
2810
-
2811
- 1. Start llama-server with slot save path.
2812
- 2. Start DSH with plugin.
2813
- 3. Open session A.
2814
- 4. Send large prompt.
2815
- 5. Slot becomes owned by A.
2816
- 6. Send another A turn.
2817
- 7. No disk save/restore occurs.
2818
- 8. Switch to session B.
2819
- 9. A is saved automatically.
2820
- 10. B runs.
2821
- 11. Switch back to A.
2822
- 12. A snapshot is restored.
2823
- 13. Next request demonstrates substantial prompt-cache reuse.
2824
- 14. Restart DSH.
2825
- 15. Open A and send another message.
2826
- 16. Plugin restores A snapshot.
2827
- 17. Conversation remains correct if snapshot file is manually deleted.
2828
- 18. Conversation remains correct if restore returns an error.
2829
- 19. Unmanaged providers are completely unaffected.
2830
- 20. Plugin hot unload cleans timers/resources.
2831
-
2832
- ---
2833
-
2834
- # 75. Integration tests
2835
-
2836
- Minimum integration test suite:
2837
-
2838
- ### Cold start
2839
-
2840
- ```text
2841
- snapshot absent
2842
- → request succeeds
2843
- → state dirty
2844
- ```
2845
-
2846
- ### Resident reuse
2847
-
2848
- ```text
2849
- A request
2850
- A request
2851
- → no save
2852
- → no restore
2853
- ```
2854
-
2855
- ### Switch
2856
-
2857
- ```text
2858
- A
2859
- B
2860
- → save A
2861
- ```
2862
-
2863
- ### Restore
2864
-
2865
- ```text
2866
- A
2867
- B
2868
- A
2869
- → restore A
2870
- ```
2871
-
2872
- ### Corrupt snapshot
2873
-
2874
- ```text
2875
- restore fails
2876
- → snapshot invalidated
2877
- → cold request succeeds
2878
- ```
2879
-
2880
- ### Backend unavailable
2881
-
2882
- ```text
2883
- /slots unreachable
2884
- → ordinary LLM request still works
2885
- ```
2886
-
2887
- ### Save failure
2888
-
2889
- ```text
2890
- save A fails
2891
- → B still eventually runs in non-strict mode
2892
- ```
2893
-
2894
- ### Auxiliary request
2895
-
2896
- ```text
2897
- A
2898
- session-title
2899
- A
2900
- → no incorrect slot ownership
2901
- ```
2902
-
2903
- ### Cancellation
2904
-
2905
- ```text
2906
- cancel model request
2907
- → lock released
2908
- → next session still works
2909
- ```
2910
-
2911
- ### HMR/disposal
2912
-
2913
- ```text
2914
- reload plugin
2915
- → no orphan timer
2916
- → no dead mutex
2917
- ```
2918
-
2919
- ---
2920
-
2921
- # 76. Unit tests
2922
-
2923
- Unit test:
2924
-
2925
- ```text
2926
- state-machine transitions
2927
- slot selection
2928
- snapshot compatibility
2929
- filename sanitization
2930
- fingerprint stability
2931
- dirty revision logic
2932
- save coalescing
2933
- idle timer invalidation
2934
- failure circuit breaker
2935
- manifest atomicity
2936
- provider filtering
2937
- ```
2938
-
2939
- No network should be necessary for these.
2940
-
2941
- ---
2942
-
2943
- # 77. Fake backend
2944
-
2945
- Create:
2946
-
2947
- ```ts
2948
- class FakeKvBackend
2949
- ```
2950
-
2951
- with deterministic state.
2952
-
2953
- Example capabilities:
2954
-
2955
- ```ts
2956
- backend.failNextSave()
2957
- backend.failNextRestore()
2958
- backend.delayRestore(100)
2959
- backend.removeSnapshot(key)
2960
- backend.corruptSnapshot(key)
2961
- ```
2962
-
2963
- Most coordinator tests should run against this rather than launching llama-server.
2964
-
2965
- ---
2966
-
2967
- # 78. Real llama integration test
2968
-
2969
- Optional test profile:
2970
-
2971
- ```text
2972
- DSH_KV_TEST_LLAMA_URL=http://127.0.0.1:8080
2973
- ```
2974
-
2975
- Tests only run when explicitly enabled.
2976
-
2977
- Never require a GPU in the ordinary CI pipeline.
2978
-
2979
- ---
2980
-
2981
- # 79. Error taxonomy
2982
-
2983
- Use stable codes.
2984
-
2985
- Suggested:
2986
-
2987
- ```text
2988
- KV_BACKEND_UNAVAILABLE
2989
- KV_BACKEND_UNSUPPORTED
2990
-
2991
- KV_SLOT_NOT_FOUND
2992
- KV_SLOT_BUSY
2993
- KV_SLOT_STATE_INVALID
2994
-
2995
- KV_SNAPSHOT_NOT_FOUND
2996
- KV_SNAPSHOT_INCOMPATIBLE
2997
- KV_SNAPSHOT_CORRUPT
2998
-
2999
- KV_SAVE_FAILED
3000
- KV_RESTORE_FAILED
3001
- KV_ERASE_FAILED
3002
-
3003
- KV_MANIFEST_INVALID
3004
- KV_METADATA_IO
3005
-
3006
- KV_OPERATION_TIMEOUT
3007
-
3008
- KV_INVARIANT
3009
- ```
3010
-
3011
- Infrastructure diagnostics become much easier than matching error strings.
3012
-
3013
- ---
3014
-
3015
- # 80. Example logs
3016
-
3017
- First request:
3018
-
3019
- ```text
3020
- [kv-persist] session cold
3021
- session=7c856d slot=0
3022
- ```
3023
-
3024
- Idle checkpoint:
3025
-
3026
- ```text
3027
- [kv-persist] snapshot saved
3028
- session=7c856d
3029
- tokens=48712
3030
- bytes=2.42GiB
3031
- save=531ms
3032
- ```
3033
-
3034
- Resume:
3035
-
3036
- ```text
3037
- [kv-persist] snapshot restored
3038
- session=7c856d
3039
- tokens=48712
3040
- bytes=2.42GiB
3041
- restore=188ms
3042
- ```
3043
-
3044
- Failure:
3045
-
3046
- ```text
3047
- [kv-persist] restore failed; falling back to cold prefill
3048
- session=7c856d
3049
- code=KV_RESTORE_FAILED
3050
- ```
3051
-
3052
- ---
3053
-
3054
- # 81. User-visible UX
3055
-
3056
- Most of the time:
3057
-
3058
- ```text
3059
- nothing
3060
- ```
3061
-
3062
- It should simply make old local-model sessions resume quickly.
3063
-
3064
- Potential status line later:
3065
-
3066
- ```text
3067
- KV: restored 48.7K · 188ms
3068
- ```
3069
-
3070
- or:
3071
-
3072
- ```text
3073
- KV: resident
3074
- ```
3075
-
3076
- or:
3077
-
3078
- ```text
3079
- KV: cold
3080
- ```
3081
-
3082
- But this belongs to a later UI integration and should not block the core plugin.
3083
-
3084
- ---
3085
-
3086
- # 82. Suggested README pitch
3087
-
3088
- > `dsh-kv-persist` keeps local LLM sessions warm across session switches and restarts.
3089
- >
3090
- > It maps DeepSeek Harness sessions to persistent inference-cache snapshots and restores them when a session becomes active again. The initial backend uses llama.cpp's slot save/restore API, allowing large agent contexts to resume without repeating a full prompt prefill.
3091
- >
3092
- > KV state is treated strictly as an optimization: DSH's session log remains the source of truth, and any missing, stale, or incompatible cache automatically falls back to normal inference.
3093
-
3094
- ---
3095
-
3096
- # 83. Roadmap
3097
-
3098
- ## Phase 0 — research/prototype
3099
-
3100
- - Validate llama.cpp save/restore with target Qwen3.8 build.
3101
- - Verify restore within same server process.
3102
- - Verify restore across llama-server restart.
3103
- - Measure snapshot sizes.
3104
- - Measure save/restore throughput.
3105
- - Confirm cache hit after restore.
3106
- - Document hybrid-model behavior.
3107
-
3108
- ## Phase 1 — MVP / `0.1`
3109
-
3110
- - Cordis service.
3111
- - llama.cpp client.
3112
- - backend probe.
3113
- - single-slot coordinator.
3114
- - session mapping.
3115
- - `llm/stream` wrapper.
3116
- - lazy restore.
3117
- - save-before-switch.
3118
- - idle checkpoint.
3119
- - shutdown/session-flush checkpoint.
3120
- - local metadata.
3121
- - logging.
3122
- - cold fallback.
3123
- - fake backend tests.
3124
-
3125
- ## Phase 2 — reliability / `0.2`
3126
-
3127
- - compatibility fingerprints.
3128
- - circuit breaker.
3129
- - snapshot verification.
3130
- - atomic snapshot rotation.
3131
- - diagnostics API.
3132
- - `doctor`.
3133
- - cleanup tooling.
3134
- - improved hybrid/recurrent testing.
3135
-
3136
- ## Phase 3 — multi-slot / `0.3`
3137
-
3138
- - slot pool.
3139
- - LRU assignment.
3140
- - explicit slot leases.
3141
- - llama-specific transport integration.
3142
- - request `id_slot`.
3143
- - auxiliary slot reservation.
3144
- - concurrent sessions.
3145
-
3146
- ## Phase 4 — observability / `0.4`
3147
-
3148
- - metrics.
3149
- - cache hit statistics.
3150
- - storage statistics.
3151
- - performance comparisons.
3152
- - optional DSH UI panel.
3153
-
3154
- ## Phase 5 — generalized persistence / `1.0`
3155
-
3156
- - stable backend interface.
3157
- - multiple servers.
3158
- - multiple backends.
3159
- - retention policies.
3160
- - documented API for external plugins.
3161
- - production-hardening.
3162
-
3163
- ---
3164
-
3165
- # 84. First implementation milestone
3166
-
3167
- The first prototype should intentionally do almost nothing clever.
3168
-
3169
- Hardcode/test:
3170
-
3171
- ```text
3172
- provider = local-qwen
3173
- slot = 0
3174
- server = localhost:8080
3175
- ```
3176
-
3177
- Implement only:
3178
-
3179
- ```text
3180
- request A
3181
- request A
3182
- request B
3183
- request A
3184
- ```
3185
-
3186
- Expected management calls:
3187
-
3188
- ```text
3189
- A #1:
3190
- erase
3191
-
3192
- A #2:
3193
- none
3194
-
3195
- B:
3196
- save A
3197
- erase/restore B
3198
-
3199
- A #3:
3200
- save B
3201
- restore A
3202
- ```
3203
-
3204
- Once this works reliably, abstract it.
3205
-
3206
- Do not start by implementing:
3207
-
3208
- ```text
3209
- multi-server
3210
- multi-slot
3211
- GC
3212
- UI
3213
- dynamic provider discovery
3214
- ```
3215
-
3216
- before proving the fundamental cache lifecycle.
3217
-
3218
- ---
3219
-
3220
- # 85. Key architectural invariants
3221
-
3222
- These should eventually exist as comments/tests.
3223
-
3224
- **Invariant 1**
3225
-
3226
- ```text
3227
- DSH session state never depends on KV persistence.
3228
- ```
3229
-
3230
- **Invariant 2**
3231
-
3232
- ```text
3233
- At most one owner controls a physical slot at a time.
3234
- ```
3235
-
3236
- **Invariant 3**
3237
-
3238
- ```text
3239
- A dirty slot is checkpointed before reassignment unless policy explicitly disables it.
3240
- ```
3241
-
3242
- **Invariant 4**
3243
-
3244
- ```text
3245
- A snapshot is restored only when its runtime identity is compatible.
3246
- ```
3247
-
3248
- **Invariant 5**
3249
-
3250
- ```text
3251
- Persistence failure defaults to cold inference.
3252
- ```
3253
-
3254
- **Invariant 6**
3255
-
3256
- ```text
3257
- Resident-session requests incur no disk I/O.
3258
- ```
3259
-
3260
- **Invariant 7**
3261
-
3262
- ```text
3263
- Auxiliary LLM requests never become authoritative state for a conversation session.
3264
- ```
3265
-
3266
- **Invariant 8**
3267
-
3268
- ```text
3269
- All backend mutation operations are serialized per physical slot.
3270
- ```
3271
-
3272
- **Invariant 9**
3273
-
3274
- ```text
3275
- Snapshot filenames are plugin-generated opaque identifiers.
3276
- ```
3277
-
3278
- **Invariant 10**
3279
-
3280
- ```text
3281
- A successful HTTP restore is not automatically equivalent to a verified usable restore.
3282
- ```
3283
-
3284
- ---
3285
-
3286
- # 86. Recommended initial technical direction
3287
-
3288
- For the first release, use:
3289
-
3290
- ```text
3291
- Cordis plugin
3292
- +
3293
- ctx.sessions lifecycle
3294
- +
3295
- llm/stream observation
3296
- +
3297
- single llama slot
3298
- +
3299
- server-side snapshot files
3300
- ```
3301
-
3302
- Do NOT fork or patch DeepSeek Harness.
3303
-
3304
- Do NOT replace the existing OpenAI-compatible provider.
3305
-
3306
- Do NOT modify prompts.
3307
-
3308
- Do NOT make cache state part of SessionEvent history.
3309
-
3310
- Once the single-slot implementation proves useful, introduce the llama-specific transport adapter needed for explicit `id_slot` and proper multi-session concurrency.
3311
-
3312
- This yields a plugin that starts as a small, useful local optimization but has a clean path toward becoming a general persistence/cache coordinator for local inference runtimes.