ds4-context-engine 0.1.2 → 0.2.0-beta.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -16,7 +16,7 @@ bounded active context with provenance
16
16
  Pi provider
17
17
  ```
18
18
 
19
- > **Project status:** M0–M13 are implemented. The Pi adapter and standalone `ds4-context-core` package are version `0.1.2`; the adapter targets Pi `0.84.3`.
19
+ > **Project status:** M0–M19 are implemented. Prerelease `0.2.0-beta.1` adds quality measurement, structural/hybrid retrieval, cross-session project memory, learned-ranking shadow evaluation, and the runtime adapter kit. Stable `0.1.2` remains available; both lines target Pi `0.84.3`. M20 local KV reuse is not included in this prerelease.
20
20
 
21
21
  ## Why DS4
22
22
 
@@ -27,13 +27,18 @@ It provides:
27
27
  - deterministic token budgeting with soft and hard input limits;
28
28
  - preservation of the current request, recent turns and atomic tool call/result groups;
29
29
  - exact and FTS5 historical retrieval with source provenance;
30
- - trust-gated project indexing, Git-aware invalidation and bounded source snippets;
30
+ - opt-in hybrid semantic retrieval with a deterministic local embedding and lexical fallback;
31
+ - trust-gated structural project indexing, Git-aware invalidation and bounded source snippets;
31
32
  - hierarchical, validated, non-destructive compaction summaries;
32
33
  - persistent pins and append-only durable memory stored canonically in Pi JSONL;
34
+ - opt-in checkpointed project-memory replay across exact trusted Pi project sessions;
33
35
  - content-addressed storage and bounded references for large tool results;
34
36
  - privacy classifications, secret redaction and provider-specific allow rules;
35
37
  - model-specific calibration and adaptive context allocation;
36
38
  - optional verified continuation for eligible OpenAI Responses profiles;
39
+ - opt-in metadata-only context-quality metrics and deterministic replay comparisons;
40
+ - checksummed metadata-only learned ranking with shadow mode, canonical classified feedback and static fallback;
41
+ - a versioned runtime adapter contract, reusable conformance kit and non-Pi callback/JSONL reference adapter;
37
42
  - an inspectable Context Manifest explaining included and excluded material;
38
43
  - fail-open recovery to Pi's native context path for operational failures.
39
44
 
@@ -74,12 +79,20 @@ pi -e git:github.com/Alucard24/ds4-context-engine
74
79
 
75
80
  ### From npm
76
81
 
77
- Install the public npm package with:
82
+ Install the latest stable public npm package with:
78
83
 
79
84
  ```bash
80
85
  pi install npm:ds4-context-engine
81
86
  ```
82
87
 
88
+ Install the opt-in 0.2 prerelease from the `beta` dist-tag with:
89
+
90
+ ```bash
91
+ pi install npm:ds4-context-engine@beta
92
+ ```
93
+
94
+ All prerelease packages (`ds4-context-engine`, `ds4-context-core`, and `ds4-context-reference-adapter`) use the same exact version.
95
+
83
96
  ### Local checkout
84
97
 
85
98
  ```bash
@@ -106,6 +119,7 @@ After loading the extension, inspect its state:
106
119
 
107
120
  ```text
108
121
  /context status
122
+ /context adapter
109
123
  /context tokens
110
124
  /context health
111
125
  ```
@@ -149,6 +163,7 @@ Project configuration and project source indexing are disabled when Pi reports t
149
163
  | Command | Purpose |
150
164
  | --- | --- |
151
165
  | `/context` or `/context status` | Runtime, session, planner and subsystem status |
166
+ | `/context adapter` | Runtime contract and per-capability negotiation diagnostics |
152
167
  | `/context tokens` | Token budget and active-context composition |
153
168
  | `/context manifest` | Latest Context Manifest |
154
169
  | `/context explain` | Human-readable planning explanation |
@@ -159,6 +174,8 @@ Project configuration and project source indexing are disabled when Pi reports t
159
174
  | `/context project` | Project index and retrieval status |
160
175
  | `/context privacy` | Classification and provider-policy status |
161
176
  | `/context model` | Active model profile and calibration |
177
+ | `/context quality` | Metadata-only context-quality scores and sample counts |
178
+ | `/context ranking` | Learned model, promotion gate, aggregate shadow comparison and feedback counts |
162
179
  | `/context continuation` | Native continuation decisions and counters |
163
180
  | `/context artifacts` | Artifact storage and integrity status |
164
181
  | `/context compaction` | Last compaction status |
@@ -179,10 +196,22 @@ Project configuration and project source indexing are disabled when Pi reports t
179
196
  /context memory supersede MEMORY_ID [--source ID,ID] <new claim>
180
197
  /context memory invalidate MEMORY_ID [reason]
181
198
  /context memory expire MEMORY_ID [reason]
199
+ /context memory sources
200
+ /context memory exclude SESSION_ID [reason]
201
+ /context memory include SESSION_ID
182
202
  ```
183
203
 
184
204
  Valid privacy classifications are `normal`, `internal`, `sensitive` and `local-only`.
185
205
 
206
+ Learned-ranking feedback and local training are explicit:
207
+
208
+ ```text
209
+ /context ranking feedback useful|irrelevant CANDIDATE_ID [--classification LEVEL]
210
+ /context ranking train
211
+ ```
212
+
213
+ See [`docs/LEARNED_RANKING.md`](docs/LEARNED_RANKING.md).
214
+
186
215
  ## Configuration reference
187
216
 
188
217
  The following example shows the main configuration groups. Omitted values use the defaults in [`packages/core/src/config/config.ts`](packages/core/src/config/config.ts).
@@ -208,7 +237,19 @@ The following example shows the main configuration groups. Omitted values use th
208
237
  "exact": true,
209
238
  "fts": true,
210
239
  "semantic": false,
211
- "maxResults": 12
240
+ "maxResults": 12,
241
+ "embedding": {
242
+ "mode": "local",
243
+ "provider": "ds4-local",
244
+ "model": "feature-hash-v1",
245
+ "dimensions": 256,
246
+ "remoteProfiles": [],
247
+ "maxSources": 50000,
248
+ "candidatePool": 80,
249
+ "batchSize": 64,
250
+ "queryCacheSize": 64,
251
+ "timeoutMs": 2000
252
+ }
212
253
  },
213
254
  "project": {
214
255
  "enabled": true,
@@ -221,6 +262,8 @@ The following example shows the main configuration groups. Omitted values use th
221
262
  },
222
263
  "memory": {
223
264
  "enabled": true,
265
+ "crossSession": false,
266
+ "maxProjectSessions": 250,
224
267
  "maxPinChars": 4000,
225
268
  "maxClaimChars": 2000,
226
269
  "maxResults": 12
@@ -271,13 +314,27 @@ The following example shows the main configuration groups. Omitted values use th
271
314
  "maxStateAgeMs": 1800000,
272
315
  "retryManagedReplay": true
273
316
  },
317
+ "quality": {
318
+ "enabled": false,
319
+ "maxSamples": 1000
320
+ },
321
+ "ranking": {
322
+ "mode": "off",
323
+ "modelPath": "ds4-context/ranking-model.json",
324
+ "minimumTrainingSamples": 20,
325
+ "maxTrainingSamples": 10000,
326
+ "maxLatencyMs": 10
327
+ },
274
328
  "diagnostics": {
275
329
  "storeContextManifest": true,
276
330
  "storeFullRenderedContext": false,
277
331
  "logLevel": "info"
278
332
  },
279
333
  "storage": {
280
- "databasePath": "ds4-context/context.db"
334
+ "databasePath": "ds4-context/context.db",
335
+ "busyTimeoutMs": 5000,
336
+ "writeRetryTimeoutMs": 30000,
337
+ "projectIndexLeaseMs": 120000
281
338
  }
282
339
  }
283
340
  ```
@@ -311,10 +368,11 @@ By default, derived state is stored below Pi's agent directory:
311
368
  ```text
312
369
  ~/.pi/agent/ds4-context/
313
370
  ├── context.db
371
+ ├── ranking-model.json
314
372
  └── artifacts/
315
373
  ```
316
374
 
317
- The database contains rebuildable indexes, summary metadata, manifests, project projections and calibration data. Canonical memory and pin mutations remain append-only entries in Pi JSONL. Project files remain canonical for project knowledge. Complete tool results remain in Pi JSONL while the artifact store keeps verified, content-addressed copies for bounded retrieval.
375
+ The database contains rebuildable indexes, summary metadata, manifests, project projections and calibration data. The optional checksummed learned-ranking model is also derived local state; its classified metadata-only labels remain canonical Pi custom entries. All Pi sessions share this WAL database: writes use bounded busy-aware transaction replay, and a renewable project lease prevents multiple Pi processes from indexing the same project concurrently. `busyTimeoutMs` controls each SQLite lock wait, while `writeRetryTimeoutMs` bounds the total replay window. Canonical memory and pin mutations remain append-only entries in Pi JSONL. Project files remain canonical for project knowledge. Complete tool results remain in Pi JSONL while the artifact store keeps verified, content-addressed copies for bounded retrieval.
318
376
 
319
377
  To validate or rebuild derived state:
320
378
 
@@ -323,32 +381,36 @@ To validate or rebuild derived state:
323
381
  /context rebuild-index
324
382
  ```
325
383
 
326
- Deleting DS4's database must not alter a Pi session or project, although derived indexes and calibration data will be regenerated.
384
+ Deleting DS4's database must not alter a Pi session or project, although derived indexes and calibration data will be regenerated. When `memory.crossSession` is enabled for a trusted project, DS4 discovers bounded sibling Pi JSONL files by exact canonical header identity, incrementally replays their explicit project mutations, and excludes missing or unverifiable sources.
327
385
 
328
386
  ## Development
329
387
 
330
388
  ```bash
331
389
  npm ci
332
390
  npm run build:core
391
+ npm run build:adapters
333
392
  npm run typecheck
334
393
  npm test
335
394
  npm run check
395
+ npm run quality:compare
336
396
  npm run pack:check
337
397
  npm pack --dry-run
338
398
  npm pack --dry-run --workspace ds4-context-core
399
+ npm pack --dry-run --workspace ds4-context-reference-adapter
339
400
  ```
340
401
 
341
- The test suite covers configuration, migrations, canonical JSONL projection, planning, atomic tool groups, retrieval, compaction, project knowledge, artifacts, memory, privacy, model awareness, continuation, the portable-core dependency boundary and Pi extension lifecycle behavior. The package check also builds both tarballs, installs them in a clean temporary consumer and starts the packaged extension with isolated Pi RPC state.
402
+ The test suite covers configuration, migrations, canonical JSONL projection, planning, atomic tool groups, retrieval, compaction, project knowledge, artifacts, memory, privacy, model awareness, continuation, runtime-adapter conformance, the portable-core dependency boundary and Pi extension lifecycle behavior. The package check builds all three tarballs, installs them in a clean temporary consumer, reruns compiled reference-adapter conformance and starts the packaged Pi extension with isolated RPC state.
342
403
 
343
404
  ### Portable core
344
405
 
345
- `ds4-context-core` is a compiled ESM package with no Pi dependency. It owns runtime-neutral policy, storage and projections; agent adapters translate native sessions and lifecycle hooks at the boundary. The root `ds4-context-engine` package is the Pi adapter and depends one-way on the core workspace.
406
+ `ds4-context-core` is a compiled ESM package with no runtime SDK dependency. It owns runtime-neutral policy, storage, adapter contracts and projections; agent adapters translate native sessions and lifecycle hooks at the boundary. The root `ds4-context-engine` package is the Pi adapter. `ds4-context-reference-adapter` is a separately compiled non-Pi callback/JSONL implementation. Both depend one-way and exactly on matching core.
346
407
 
347
408
  ### Repository layout
348
409
 
349
410
  ```text
350
- packages/core/src portable policy, planning, compaction, retrieval and storage
351
- src/pi-adapter Pi JSONL projection, summary completion and provider integration
411
+ packages/core/src portable policy, adapter kit, planning, retrieval and storage
412
+ packages/reference-adapter/src non-Pi callback/JSONL reference runtime boundary
413
+ src/pi-adapter Pi JSONL projection, summary completion and provider integration
352
414
  src/extension Pi hooks, commands and fail-open orchestration
353
415
  tests core contract, unit, integration, golden and benchmark coverage
354
416
  scripts package and release-readiness checks
@@ -359,10 +421,13 @@ scripts package and release-readiness checks
359
421
 
360
422
  - [Architecture](docs/ARCHITECTURE.md)
361
423
  - [Context planner](docs/CONTEXT_PLANNER.md)
424
+ - [Context quality](docs/CONTEXT_QUALITY.md)
425
+ - [Learned ranking](docs/LEARNED_RANKING.md)
362
426
  - [Context Manifest](docs/CONTEXT_MANIFEST.md)
363
427
  - [Compaction](docs/COMPACTION.md)
364
428
  - [Summary graph](docs/SUMMARY_GRAPH.md)
365
429
  - [Historical retrieval](docs/RETRIEVAL.md)
430
+ - [Hybrid semantic retrieval](docs/HYBRID_RETRIEVAL.md)
366
431
  - [Project knowledge](docs/PROJECT_KNOWLEDGE.md)
367
432
  - [Artifacts](docs/ARTIFACTS.md)
368
433
  - [Memory and pins](docs/MEMORY_AND_PINS.md)
@@ -370,6 +435,7 @@ scripts package and release-readiness checks
370
435
  - [Model awareness](docs/MODEL_AWARENESS.md)
371
436
  - [Native continuation](docs/NATIVE_CONTINUATION.md)
372
437
  - [Portable core](docs/PORTABLE_CORE.md)
438
+ - [Runtime adapter kit](docs/RUNTIME_ADAPTER_KIT.md)
373
439
  - [Storage](docs/STORAGE.md)
374
440
  - [Roadmap 0.2.0](docs/ROADMAP_0.2.0.md)
375
441
  - [Release process](docs/RELEASING.md)
@@ -378,9 +444,9 @@ scripts package and release-readiness checks
378
444
 
379
445
  ## Roadmap
380
446
 
381
- The original M0–M13 roadmap is complete. `ds4-context-core` now contains the compiled Pi-independent implementation, while runtime-specific behavior remains in the Pi adapter.
447
+ The original M0–M13 roadmap is complete. `ds4-context-core` contains the compiled runtime-neutral implementation. M14 context-quality metrics, M15 rich symbol indexing, M16 hybrid semantic retrieval, M17 cross-session project memory, M18 learned-ranking shadow evaluation and M19's runtime adapter/conformance kit plus non-Pi reference adapter are implemented on `main`. Learned active ranking remains promotion-gated and static ranking stays authoritative on every failure.
382
448
 
383
- The planned [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) covers context-quality metrics, richer symbol indexing, hybrid semantic retrieval, cross-session project memory, optional learned ranking, a runtime adapter kit with one reference adapter, and optional local KV reuse. Sensitive or transport-specific behavior remains opt-in, and the 0.1 lexical planner stays available as the deterministic fallback.
449
+ The remaining [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) covers optional local KV reuse. Sensitive or transport-specific behavior remains opt-in, and the 0.1 lexical planner stays available as the deterministic fallback.
384
450
 
385
451
  ## Contributing
386
452
 
@@ -9,7 +9,12 @@ Pi session_start
9
9
  -> validate Pi JSONL v3 header
10
10
  -> full index or checkpointed append sync
11
11
  -> replay versioned memory/pin custom-entry mutations into transactional projections
12
+ -> when opted in and trusted, checkpoint/replay explicit project mutations from exact-identity sibling Pi sessions
12
13
  -> if trusted, canonicalize project root and incrementally index bounded text files
14
+ -> parse TypeScript/JavaScript/Python/Go declaration boundaries through the runtime-neutral parser interface
15
+ -> fall back deterministically to bounded text windows when adapters, syntax, or language support are unavailable
16
+ -> persist only rebuildable symbol/signature/parent/import/reference projections
17
+ -> when semantic retrieval is opted in, refresh source-hash/model-keyed vectors through the runtime embedding port
13
18
  -> snapshot Git root/branch/HEAD/dirty paths
14
19
 
15
20
  Pi context hook
@@ -30,16 +35,22 @@ Pi context hook
30
35
  -> select a contiguous, model-adaptive recent tail
31
36
  -> derive current-task identifiers, files, errors, phrases and keywords
32
37
  -> query exact matches and FTS5 over canonical indexed entries
38
+ -> optionally add bounded cosine-ranked history candidates and deterministic lexical/vector rank fusion
33
39
  -> reject active-context duplicates and all alternate-branch candidates
34
40
  -> rank, deduplicate, quote and budget historical evidence groups
35
- -> query exact path/symbol/phrase and project FTS5 candidates
36
- -> live-validate candidate SHA-256 and reindex changed files
41
+ -> query exact literal path and qualified/simple declaration indexes ahead of phrase and project FTS5 candidates
42
+ -> optionally add bounded project vectors while retaining exact path/symbol priority
43
+ -> live-validate candidate SHA-256 and reindex only changed-file projections
37
44
  -> rank, overlap-deduplicate, quote and budget project source groups
38
45
  -> omit prohibited history/project/pin/memory supplements with metadata-only privacy reasons
39
46
  -> fit active Pi summaries in the remaining budget
40
47
  -> validate hard limit, current request, and tool call/results
41
48
  -> return selected messages or fail open to the already privacy-sanitized native context
42
49
  -> persist metadata-only Context Manifest, privacy counters, and prompt hash
50
+ -> when opted in, queue the completed metadata-only manifest for deferred quality measurement
51
+
52
+ Pi agent_settled / shutdown
53
+ -> materialize and persist bounded content-free quality counts without changing the plan
43
54
 
44
55
  before_provider_request
45
56
  -> recheck provider-specific serialized system/messages/tools/content
@@ -109,10 +120,10 @@ session_tree / shutdown
109
120
  -> immutable graph nodes, ordered edges, roots, levels and current-branch active path
110
121
 
111
122
  /context retrieved
112
- -> query terms, candidate counts, branch blocks, budget decisions and injected excerpts
123
+ -> query terms, lexical/vector/fused counts, embedding profile/freshness/fallback, branch blocks, budgets and excerpts
113
124
 
114
125
  /context project
115
- -> trust, Git revision, file/snippet/stale counts, retrieval decisions and local excerpts
126
+ -> trust, Git revision, file/snippet/stale/vector counts, retrieval decisions and local excerpts
116
127
 
117
128
  /context pins | pin | unpin
118
129
  -> inspect or append immutable session/branch/project pin mutations
@@ -141,16 +152,18 @@ session_tree / shutdown
141
152
  Dependency direction is one-way:
142
153
 
143
154
  ```text
144
- Pi native types and lifecycle
145
- ↓
146
- ds4-context-engine (src/pi-adapter + src/extension)
147
- ↓
148
- ds4-context-core (packages/core)
155
+ Pi native types and lifecycle callback/JSONL runtime
156
+ ↓ ↓
157
+ ds4-context-engine ds4-context-reference-adapter
158
+ └──────────────────────┬───────────────────────┘
159
+ ↓
160
+ ds4-context-core (packages/core)
149
161
  ```
150
162
 
151
163
  `ds4-context-core` is compiled ESM and has no dependency on Pi. Its workspace contains:
152
164
 
153
- - `packages/core/src/core`: portable model profiles, robust calibration, adaptive category limits, budgets and token-estimation policy;
165
+ - `packages/core/src/adapter`: versioned runtime contract, canonical tool-group validation, isolated capability negotiation and framework-neutral conformance runner;
166
+ - `packages/core/src/core`: portable canonical messages, model profiles, robust calibration, adaptive category limits, budgets and token-estimation policy;
154
167
  - `packages/core/src/continuation`: hashed-prefix continuation decisions without provider transport or response APIs;
155
168
  - `packages/core/src/config`: runtime-neutral configuration model and filesystem loader;
156
169
  - `packages/core/src/planner`: atomic grouping, deterministic ranking, fitting, validation and privacy-aware plans;
@@ -158,21 +171,25 @@ ds4-context-core (packages/core)
158
171
  - `packages/core/src/memory`: mutation projections, conservative contradiction/key detection, scope selection, prompt boundaries, ranking and diagnostics;
159
172
  - `packages/core/src/artifacts`: atomic content-addressed files, deterministic condensation, redaction, branch-safe literal search, reconciliation and garbage collection;
160
173
  - `packages/core/src/compaction`: structured summary contract, hierarchical graph model, validation, lifecycle metadata and source hashing;
161
- - `packages/core/src/retrieval`: task descriptors, safe FTS queries, deterministic ranking, evidence quoting, deduplication and token fitting;
174
+ - `packages/core/src/retrieval`: task descriptors, safe FTS queries, runtime-neutral embedding port, semantic index orchestration, deterministic rank fusion, evidence quoting, quality comparison, deduplication and token fitting;
162
175
  - `packages/core/src/project`: trust-gated file discovery, hashing, Git state, symbol/chunk extraction, invalidation, retrieval and source quoting;
163
- - `packages/core/src/persistence`: rebuildable session/project/memory/pin SQLite state, repositories, FTS5, event replay and transactional migrations;
176
+ - `packages/core/src/quality`: versioned replay fixtures/contracts, deterministic metrics, static/candidate comparison and metadata-only aggregation;
177
+ - `packages/core/src/ranking`: bounded metadata-only features, classified label contracts, deterministic local training, checksummed model artifacts, aggregate shadow comparison and promotion-gated inference;
178
+ - `packages/core/src/persistence`: rebuildable session/project/vector/memory/pin/quality SQLite state, repositories, FTS5, event replay and transactional migrations;
164
179
  - `packages/core/src/manifest` and `packages/core/src/shared`: runtime-neutral projections, provenance, hashing, stable serialization and logging.
165
180
 
181
+ The `packages/reference-adapter` workspace is the non-Pi reference adapter: it reads bounded append-only canonical JSONL, injects completion through a host callback, enforces privacy at that callback boundary, rebuilds disposable snapshots and explicitly disables unsupported native features.
182
+
166
183
  The root `ds4-context-engine` package is the Pi adapter:
167
184
 
168
- - `src/pi-adapter`: byte-safe Pi JSONL reading, provenance mapping, custom mutation projection, active label discovery, checkpoints, runtime snapshots, Pi model completion for summaries and the narrow Pi-AI OpenAI Responses transport wrapper;
185
+ - `src/pi-adapter`: byte-safe Pi JSONL reading, provenance mapping, memory/pin and learned-ranking label projection, active label discovery, checkpoints, runtime snapshots, Pi model completion for summaries and the narrow Pi-AI OpenAI Responses transport wrapper;
169
186
  - `src/extension`: Pi hooks, lifecycle, command presentation and fail-open/fail-closed orchestration.
170
187
 
171
- The adapter may import core exports. Core source must never import `@earendil-works/pi-ai`, `@earendil-works/pi-coding-agent`, `src/pi-adapter` or `src/extension`; an automated boundary test enforces this rule.
188
+ Adapters may import core exports. Core source must never import `@earendil-works/pi-ai`, `@earendil-works/pi-coding-agent`, `src/pi-adapter`, `src/extension` or a reference-adapter source path; an automated boundary test enforces this rule. Runtime SDK dependencies belong only to their adapter package.
172
189
 
173
190
  ## Canonical and derived state
174
191
 
175
- The Pi session JSONL remains canonical for conversation/tool state, inline classification markers, and append-only classified memory/pin custom mutations; live files remain canonical for project knowledge. Native continuation keeps only volatile request/response-item hashes plus the minimum response handle and creates no continuation table or custom entry. SQLite and content-addressed object files store only rebuildable indexes, summary nodes/edges, metadata-only manifests, project file/snippet projections, artifact copies/references, materialized memory/pins, and calibration data. Each aggregate's active text is the Pi compaction summary; non-active nodes created by the same operation are embedded in its details, while older ancestors remain in earlier entries. Deleting the database must never damage or alter a Pi session or project. Reopening a source session replays its memory/pin mutations. Ephemeral sessions keep manifests and graph nodes in memory, disable durable memory/pins/artifacts, and may share the project index because files—not session JSONL—are its durable source.
192
+ The Pi session JSONL remains canonical for conversation/tool state, inline classification markers, append-only classified memory/pin mutations and metadata-only learned-ranking feedback/replay labels; live files remain canonical for project knowledge. Native continuation keeps only volatile request/response-item hashes plus the minimum response handle and creates no continuation table or custom entry. SQLite and content-addressed object files store only rebuildable indexes, source-hash/model-keyed vectors, summary nodes/edges, metadata-only manifests, project file/snippet projections, artifact copies/references, materialized memory/pins, calibration data, and bounded metadata-only quality samples. The checksummed learned-ranking model is a separate disposable local artifact reconstructed from canonical labels; it contains bounded weights and aggregate gate metadata, never raw text. Each aggregate's active text is the Pi compaction summary; non-active nodes created by the same operation are embedded in its details, while older ancestors remain in earlier entries. Deleting the database must never damage or alter a Pi session or project. Reopening a source session replays its memory/pin mutations. Ephemeral sessions keep manifests and graph nodes in memory, disable durable memory/pins/artifacts, and may share the project index because files—not session JSONL—are its durable source.
176
193
 
177
194
  ## Lifecycle
178
195
 
@@ -195,6 +212,6 @@ Database settings:
195
212
 
196
213
  ## Failure policy
197
214
 
198
- Configuration, database, session/project indexing, memory/pin replay, artifact offload/search, retrieval, planning, observer, native continuation, and diagnostics failures are caught at the extension boundary. Session index failures retain the previous transactional snapshot. Historical and project FTS errors degrade to exact matches; project subsystem failure contributes no snippets without disabling session management. Expected planning hazards produce an explicit fallback manifest and discard synthetic evidence.
215
+ Configuration, database, session/project indexing, memory/pin replay, artifact offload/search, retrieval, planning, observer, native continuation, quality measurement, and diagnostics failures are caught at the extension boundary. Session index failures retain the previous transactional snapshot. Cross-session source failures exclude only the unverifiable source and retain explicit diagnostics; they do not disable current-session memory. Historical and project FTS errors degrade to exact matches; embedding consent/privacy/model/timeout/corruption/provider failures degrade to lexical results; project subsystem failure contributes no snippets without disabling session management. Expected planning hazards produce an explicit fallback manifest and discard synthetic evidence.
199
216
 
200
217
  Privacy is the exception to ordinary fail-open behavior. Once enabled, planner failures return the sanitized native array, preparation failures replace message content with structural placeholders, and provider-payload sanitizer failures return an empty object so the remote request fails rather than receiving unchecked content. Pi 0.84.3 runs provider-payload handlers in extension load order, so DS4 should be loaded last when other extensions can rewrite provider payloads.
@@ -20,13 +20,14 @@ A Context Manifest explains the context visible at DS4's Pi `context` hook witho
20
20
  - artifact IDs, SHA-256, bytes, MIME, classification, exact source entry/tool IDs, error state, and before/after token estimates;
21
21
  - provider destination and allow-set names, selected classification counts, blocked/excluded/redacted counts, final provider-check count, and enforcement stage;
22
22
  - planner mode/version, original and selected counts, group counts, internal budgets, duration, and fallback reason;
23
+ - learned-ranking mode/status, feature/model versions, candidate count, aggregate disagreement/rank shift, duration, and generic static-fallback reason;
23
24
  - planner and policy versions;
24
25
  - deterministic SHA-256 over system prompt, active tools, and messages;
25
26
  - Pi's reported context usage when available;
26
27
  - finalized uncached input, cache-read, cache-write, total provider input, and cache shares when available;
27
28
  - optional native-continuation eligibility, storage-consent state, request mode, full/sent/omitted input-item counts, state age, generic fallback/invalidation reason, and managed-replay retry outcome.
28
29
 
29
- The manifest does **not** contain system instructions, message text, classified spans, pin content, memory claims, project snippets, artifact content/excerpts, tool arguments/results, image data, provider payloads, provider response/conversation IDs, API keys, or headers.
30
+ The manifest does **not** contain system instructions, message text, classified spans, pin content, memory claims, project snippets, artifact content/excerpts, learned-ranking feature vectors/labels/candidate IDs/model weights, tool arguments/results, image data, provider payloads, provider response/conversation IDs, API keys, or headers.
30
31
 
31
32
  ## Provenance mapping
32
33
 
@@ -62,6 +63,7 @@ Use:
62
63
  /context pins
63
64
  /context memory
64
65
  /context privacy
66
+ /context ranking
65
67
  /context continuation
66
68
  /context artifacts
67
69
  ```
@@ -76,6 +76,14 @@ With privacy disabled, DS4 returns Pi's original `AgentMessage[]` when:
76
76
 
77
77
  Expected fallbacks are recorded in the Context Manifest. With privacy enabled, the fallback baseline is the sanitized native array—not raw Pi messages—and an unexpected privacy failure replaces content/payload fields instead of sending unchecked data. Observer mode disables planning but still enforces enabled privacy policy and records manifests/usage calibration.
78
78
 
79
+ ## Quality measurement
80
+
81
+ M14 can queue the finalized manifest after planning when `quality.enabled` is true; materialization runs after `agent_settled`, outside provider planning latency. It records only counts, ratios, normalized reason codes, budget utilization and separate timing; it does not inspect or persist message/evidence text and cannot alter the active selection. Live samples remain unlabeled for evidence recall. The versioned synthetic replay corpus supplies expected source IDs for deterministic baseline/candidate comparisons. Any quality failure is isolated and the 0.1 plan remains active. See [`CONTEXT_QUALITY.md`](CONTEXT_QUALITY.md).
82
+
83
+ ## Learned ranking
84
+
85
+ M18 can evaluate bounded metadata-only features after privacy exclusion and before supplemental candidates enter category fitting. `shadow` keeps every static score/order authoritative and records aggregate disagreement only. `active` is accepted only for a compatible checksummed model carrying an eligible held-out promotion report. Privacy exclusions, mandatory pins/current turns, atomic groups and hard budgets cannot be overridden. See [`LEARNED_RANKING.md`](LEARNED_RANKING.md).
86
+
79
87
  ## Current limits
80
88
 
81
- The planner does not call a model inside the `context` hook. Model calibration uses only finalized provider usage and deterministic local statistics. Historical/project retrieval and memory ranking are lexical; semantic reranking is intentionally disabled even if configured. Project symbol extraction is heuristic, artifact search is literal, and memory/pin creation is manual-first. Automatic memory extraction remains disabled; M10 supplies policy enforcement but not an automatic classifier or confirmation workflow. Provider-payload coverage targets Pi 0.84.3's supported serializers, and DS4 must load after any extension allowed to replace payloads when strict final ordering is required.
89
+ The planner does not call a model inside the `context` hook. Model calibration uses only finalized provider usage and deterministic local statistics. Historical/project retrieval can opt into derived semantic candidates; learned supplemental reranking remains off by default and active mode is promotion-gated. Project symbol extraction is heuristic, artifact search is literal, and memory/pin creation is manual-first. Automatic memory extraction remains disabled; M10 supplies policy enforcement but not an automatic classifier or confirmation workflow. Provider-payload coverage targets Pi 0.84.3's supported serializers, and DS4 must load after any extension allowed to replace payloads when strict final ordering is required.
@@ -0,0 +1,91 @@
1
+ # Context Quality Metrics
2
+
3
+ M14 adds an opt-in, metadata-only quality layer around the deterministic 0.1 planner. The context hook only queues the completed metadata-only manifest; metric materialization and storage run after `agent_settled` (or during shutdown flush), outside provider planning latency. Measurement never changes selection, ranking, privacy enforcement, provider payloads, or Pi JSONL.
4
+
5
+ ## Configuration
6
+
7
+ Quality sampling is disabled by default:
8
+
9
+ ```json
10
+ {
11
+ "quality": {
12
+ "enabled": true,
13
+ "maxSamples": 1000
14
+ }
15
+ }
16
+ ```
17
+
18
+ `maxSamples` must be between 1 and 100,000. Retention is bounded in SQLite. Disabling the feature returns before sample scheduling. Enabling it adds only a bounded metadata-reference queue operation to the context path; the more expensive metric pass is deferred.
19
+
20
+ ## Metrics
21
+
22
+ `context-quality-v1` records deterministic counts and ratios for:
23
+
24
+ - expected evidence recall;
25
+ - irrelevant selected-token ratio;
26
+ - duplicate evidence references;
27
+ - provenance coverage;
28
+ - current-request retention;
29
+ - atomic-group validity;
30
+ - overflow and planner-fallback rates;
31
+ - selected/dropped source-kind counts;
32
+ - category budget utilization;
33
+ - normalized selection/drop reason counts.
34
+
35
+ The primary score is versioned with the metric contract:
36
+
37
+ ```text
38
+ 35% evidence recall
39
+ 20% relevant-token share
40
+ 15% provenance coverage
41
+ 15% current-request retention
42
+ 10% atomic-group validity
43
+ 5% no-overflow/no-fallback reliability
44
+ ```
45
+
46
+ A ratio with no applicable denominator is reported as `null` and contributes a neutral value once an aggregate contains labeled evidence. An aggregate with no labeled evidence reports a zero primary score rather than claiming success. Live requests do not have human expected-evidence labels, so their evidence recall remains explicitly unlabeled rather than being assigned a tautological success. Versioned replay fixtures supply expected source IDs and produce labeled recall.
47
+
48
+ Planning duration is stored and reported separately. It is excluded from deterministic aggregates and golden output because wall-clock time is not byte-stable.
49
+
50
+ ## Privacy and storage
51
+
52
+ `context_quality_samples` is a disposable SQLite v11 projection. A stored sample contains only:
53
+
54
+ - schema, metric, corpus, planner and profile versions;
55
+ - aggregate source-kind, token, budget and decision counts;
56
+ - outcome labels;
57
+ - normalized timing values.
58
+
59
+ It does **not** contain prompts, messages, summaries, memory claims, artifact text, project paths, evidence text, provider payloads, provider response IDs, or raw evidence source IDs. Live source IDs are hashed only while constructing the volatile metric input; only resulting counts are persisted.
60
+
61
+ Malformed, incomplete, unknown-version, or structurally inconsistent rows are ignored during aggregation. Quality write/read failures are caught independently and cannot replace or block the 0.1 context plan.
62
+
63
+ Deleting SQLite discards samples without affecting canonical state. Replaying the same versioned local corpus reconstructs byte-identical non-timing aggregates.
64
+
65
+ ## Replay corpus and comparisons
66
+
67
+ [`quality/corpus-v1.json`](../quality/corpus-v1.json) contains synthetic, sanitized metadata fixtures with task descriptors, expected evidence source IDs, atomic groups, token costs, and planner budgets. It contains no captured user or project text.
68
+
69
+ Run the 0.1 static-ranking baseline against the task-weighted 0.2 candidate interface:
70
+
71
+ ```bash
72
+ npm run quality:compare
73
+ ```
74
+
75
+ The final stdout line is stable JSON. Add `-- --timing` to report wall-clock duration separately on stderr. The task-weighted candidate remains evaluation-only. M18 adds a separate sanitized learned-ranking promotion fixture contract that enforces quality, exact-recall, privacy, atomicity, overflow, latency and determinism gates before active ordering is eligible; see [`LEARNED_RANKING.md`](LEARNED_RANKING.md).
76
+
77
+ ## Diagnostics
78
+
79
+ ```text
80
+ /context quality
81
+ ```
82
+
83
+ The command reports sample counts, labeled coverage, aggregate scores, rates, category utilization, normalized reasons, and separate mean/p95 planning duration. It never renders source content.
84
+
85
+ ## Verification
86
+
87
+ - `tests/unit/context-quality.test.ts` covers every metric, weighted aggregation, deterministic replay, and live-sample redaction.
88
+ - `tests/integration/context-quality-repository.test.ts` covers delete/rebuild equivalence, bounded retention, and corrupt-row isolation.
89
+ - `tests/golden/context-quality-comparison.test.ts` locks byte-stable non-timing comparison output.
90
+ - `tests/integration/extension.test.ts` covers opt-in runtime recording and `/context quality`.
91
+ - `tests/benchmarks/context-quality.bench.ts` measures disabled/enabled context-path scheduling separately from deferred 1,000-item materialization. On the development host, a 1,000-message planner measured `2.5680 ms` p99 with metrics disabled and `2.5929 ms` with enabled scheduling (about `0.97%` overhead); the deferred 1,000-item pass measured `11.6510 ms` p99. These are observational, not portable guarantees.
@@ -0,0 +1,117 @@
1
+ # Hybrid Semantic Retrieval
2
+
3
+ M16 adds opt-in vector candidate generation to historical and trusted-project retrieval. Exact identifiers and FTS5 remain active and authoritative; semantic candidates are fused into the same bounded, deterministic ranking and never replace source provenance or live-hash checks.
4
+
5
+ ## Configuration
6
+
7
+ Semantic retrieval remains disabled for upgrades from 0.1:
8
+
9
+ ```json
10
+ {
11
+ "retrieval": {
12
+ "semantic": true
13
+ }
14
+ }
15
+ ```
16
+
17
+ The supported default is the runtime-owned local feature-hash embedding:
18
+
19
+ ```json
20
+ {
21
+ "retrieval": {
22
+ "semantic": true,
23
+ "embedding": {
24
+ "mode": "local",
25
+ "provider": "ds4-local",
26
+ "model": "feature-hash-v1",
27
+ "dimensions": 256,
28
+ "maxSources": 50000,
29
+ "candidatePool": 80,
30
+ "batchSize": 64,
31
+ "queryCacheSize": 64,
32
+ "timeoutMs": 2000
33
+ }
34
+ }
35
+ }
36
+ ```
37
+
38
+ It is deterministic, has no native dependency and performs no network access. Core defines `EmbeddingPort`; the Pi adapter supplies the implementation. Other runtimes can inject a compatible local, WASM or remote port without adding model invocation to core.
39
+
40
+ Remote embedding requires all of the following:
41
+
42
+ - `mode: "remote"`;
43
+ - an exact `provider/model` entry in `remoteProfiles` (wildcards are rejected);
44
+ - `privacy.enabled: true`;
45
+ - provider-specific privacy allow rules;
46
+ - a runtime-injected `EmbeddingPort` whose provider, model, dimensions and remote destination exactly match configuration.
47
+
48
+ ```json
49
+ {
50
+ "retrieval": {
51
+ "semantic": true,
52
+ "embedding": {
53
+ "mode": "remote",
54
+ "provider": "embedding.example",
55
+ "model": "semantic-v2",
56
+ "dimensions": 768,
57
+ "remoteProfiles": ["embedding.example/semantic-v2"]
58
+ }
59
+ },
60
+ "privacy": {
61
+ "enabled": true,
62
+ "remoteProviders": {
63
+ "embedding.example": ["normal", "internal"]
64
+ }
65
+ }
66
+ }
67
+ ```
68
+
69
+ The packaged Pi adapter intentionally includes no default remote embedding client. Missing or mismatched ports produce lexical-only results.
70
+
71
+ ## Privacy
72
+
73
+ Every remote source and query passes through `PrivacyPolicyEngine` before `EmbeddingPort.embed`. A value whose effective classification is `local-only`, or which contains a provider-blocked span, is excluded as a whole and never reaches the remote port. Allowed text is secret-redacted before invocation. Local mode does not cross a provider boundary.
74
+
75
+ Embedding diagnostics contain only counts, model identity, dimensions, destination, freshness, cache status, timings and normalized fallback reasons. They never contain query text, source text, vectors, provider response IDs or remote handles.
76
+
77
+ ## Derived storage
78
+
79
+ SQLite schema v13 adds `derived_embeddings`. Each row is keyed by:
80
+
81
+ ```text
82
+ source kind + scope + source key + source hash + chunking version
83
+ + embedding provider + embedding model + dimensions
84
+ ```
85
+
86
+ Rows contain the source group, numeric vector JSON and indexing time, but no copied query or evidence text. Session entries use `pi-session-entry-v1`; project chunks use their parser version or `text-window-v1`.
87
+
88
+ A source edit prunes only vectors whose key/hash/chunk version is no longer current. Unchanged source rows and unrelated model/dimension profiles remain intact. A model or dimension switch selects a separate profile instead of rewriting compatible rows. Deleting SQLite discards the whole vector projection; canonical Pi JSONL and live project files rebuild it.
89
+
90
+ ## Candidate generation and rank fusion
91
+
92
+ Candidate pools are bounded. Generation order is:
93
+
94
+ 1. exact project path, qualified symbol, simple symbol, quoted phrase or historical identifier;
95
+ 2. escaped FTS5 candidates;
96
+ 3. cosine-ranked vector candidates.
97
+
98
+ The lexical and vector ranks are combined with deterministic reciprocal-rank fusion, semantic similarity and stable source-ID tie-breaking. Exact matches retain a score tier that vectors cannot displace. Historical candidates still require active-branch membership and exclusion from the active native context. Project candidates still require project trust, sensitive-file exclusion and a live SHA-256 match before injection.
99
+
100
+ Source vectors are generated during session/project index sync. Query vectors use a bounded volatile hash-keyed cache. Repeating a request with current source vectors performs no embedding call. Query hashes and vectors are not persisted as conversation state.
101
+
102
+ ## Failure behavior
103
+
104
+ Missing models, consent failure, privacy exclusion, invalid dimensions, corrupt vectors, synchronous-port timeout, provider exceptions, vector storage errors and vector search errors all return the available exact/FTS result. They cannot fail context planning. Diagnostics expose the fallback reason without source content.
105
+
106
+ ## Quality gate
107
+
108
+ [`quality/semantic-corpus-v1.json`](../quality/semantic-corpus-v1.json) is a synthetic extension of the M14 replay methodology. It measures the same evidence-recall and irrelevant-token signals for semantic synonyms and exact-match preservation. The byte-stable golden report records:
109
+
110
+ ```text
111
+ lexical evidence recall 0.25
112
+ hybrid evidence recall 1.00
113
+ recall delta +0.75
114
+ irrelevant-token delta 0.00
115
+ ```
116
+
117
+ Verification covers local query reuse, source-vector caching, profile isolation, changed-source pruning, remote `local-only` exclusion, provider failure, exact priority, historical/project fusion and schema v13 rebuild behavior.