ds4-context-engine 0.3.9 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -16,7 +16,7 @@ bounded active context with provenance
16
16
  Pi provider
17
17
  ```
18
18
 
19
- > **Project status:** The coordinated `0.3.9` release adds hard estimated-input limits for every compaction request and the complete compaction operation, including retries and concurrent segments. The conservative context-fill input budget is now the default, while the prior hard-limit budget remains opt-in. Canonical history, SQLite schema 16 and runtime contracts remain unchanged from `0.3.8`; Pi remains pinned to `0.84.3`. See the [0.3.9 release record](docs/releases/0.3.9.md) for validation and publication status.
19
+ > **Project status:** The coordinated `0.4.0` release adds `storage.scope: "agent" | "project"`, now defaulting to per-project SQLite projections with shared token calibration in the agent database (opt out with `storage.scope: "agent"`). The opt-in BPE estimation and bounded auto-tuning from `0.3.10` keep `chars-v1` and disabled auto-tuning as their defaults. The bounded compaction controls from `0.3.9` remain in place; canonical history, SQLite schema 16 and runtime contracts are unchanged. Pi remains pinned to `0.84.3`. See the [0.4.0 release record](docs/releases/0.4.0.md), [ADR 064](docs/ADR/064-per-project-databases-with-shared-calibration.md) and [model-awareness validation](docs/MODEL_AWARENESS.md).
20
20
 
21
21
  **Current compaction defaults:** `compaction.directUpdate=true`, `compaction.inputBudget="context"`, `compaction.segmentTargetTokens=30000`, `compaction.maxRequestInputTokens=64000`, `compaction.maxOperationInputTokens=2000000`, `compaction.maxConcurrentSegments=2`. Every DS4 provider attempt is bounded by the effective request limit, and the operation limit includes retries; `inputBudget="summary"` remains an explicit throughput-oriented opt-in. Existing compaction/master switches still apply. See [latency controls and compatibility](docs/COMPACTION.md#latency-controls). No real-provider speedup is claimed from mock tests. The five optional editing/reading/artifact/job features introduced in `0.3.4` remain default-off.
22
22
 
@@ -385,6 +385,7 @@ The following example shows the main configuration groups. Omitted values use th
385
385
  "logLevel": "info"
386
386
  },
387
387
  "storage": {
388
+ "scope": "project",
388
389
  "databasePath": "ds4-context/context.db",
389
390
  "busyTimeoutMs": 5000,
390
391
  "writeRetryTimeoutMs": 30000,
@@ -514,6 +515,8 @@ scripts package and release-readiness checks
514
515
  - [Roadmap 0.2.0](docs/ROADMAP_0.2.0.md)
515
516
  - [Release process](docs/RELEASING.md)
516
517
  - [0.2.0 release readiness](docs/RELEASE_READINESS_0.2.0.md)
518
+ - [0.4.0 release notes](docs/releases/0.4.0.md)
519
+ - [0.3.10 release notes](docs/releases/0.3.10.md)
517
520
  - [0.3.9 release notes](docs/releases/0.3.9.md)
518
521
  - [0.3.8 release notes](docs/releases/0.3.8.md)
519
522
  - [0.3.7 release notes](docs/releases/0.3.7.md)
@@ -541,7 +544,7 @@ scripts package and release-readiness checks
541
544
 
542
545
  The original M0–M13 roadmap is complete. `ds4-context-core` contains the compiled runtime-neutral implementation. M14 context-quality metrics, M15 rich symbol indexing, M16 hybrid semantic retrieval, M17 cross-session project memory, M18 learned-ranking shadow evaluation, M19's runtime adapter/conformance kit, and M20 opt-in local KV eligibility/replay are implemented on `main`. Learned active ranking remains promotion-gated, Pi reports local KV as unsupported, and static ranking/native completion stay authoritative on every failure.
543
546
 
544
- The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.3.9 extends the bounded compaction updates, summary input headroom, concurrent segments and phase timings introduced in 0.3.5 with per-request and cumulative operation input limits; 0.3.8 adds indexed FTS key deletion without changing search results. The opt-in [anchored editing](docs/ANCHORED_EDITING.md) and [portable agent tools](docs/PORTABLE_AGENT_TOOLS.md) from 0.3.4 remain default-off, without backend rewind, forced sampling or operational KV integration. Confirmation, provenance, Pi fallback and canonical/configuration/SQLite/runtime contracts remain unchanged. The [0.2 readiness record](docs/RELEASE_READINESS_0.2.0.md) remains the compatibility baseline; the lexical planner stays available as the deterministic fallback.
547
+ The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.4.0 introduces `storage.scope` (`"agent" | "project"`, default `"project"`): one rebuildable SQLite projection per trusted canonical project root, with token calibration shared in the agent database; `storage.scope: "agent"` restores the previous single-file layout. Version 0.3.10 adds opt-in BPE estimation and bounded model-budget auto-tuning while retaining the `chars-v1` default. Version 0.3.9 extends the bounded compaction updates, summary input headroom, concurrent segments and phase timings introduced in 0.3.5 with per-request and cumulative operation input limits; 0.3.8 adds indexed FTS key deletion without changing search results. The opt-in [anchored editing](docs/ANCHORED_EDITING.md) and [portable agent tools](docs/PORTABLE_AGENT_TOOLS.md) from 0.3.4 remain default-off, without backend rewind, forced sampling or operational KV integration. Confirmation, provenance, Pi fallback and canonical/configuration/SQLite/runtime contracts remain unchanged. The [0.2 readiness record](docs/RELEASE_READINESS_0.2.0.md) remains the compatibility baseline; the lexical planner stays available as the deterministic fallback.
545
548
 
546
549
  ## Contributing
547
550
 
@@ -0,0 +1,88 @@
1
+ # 064 — Per-project databases with shared token calibration
2
+
3
+ **Date:** 2026-09-25
4
+ **Status:** Accepted
5
+ **Related:** [002](002-pi-jsonl-canonical-sqlite-rebuildable.md), [058](058-bounded-manifest-storage.md), [063](063-fts-rowid-key-mappings.md)
6
+
7
+ ## Context
8
+
9
+ The storage plan decided **D1**: keep one shared SQLite projection at
10
+ `~/.pi/agent/ds4-context/context.db` for every Pi session and explicitly avoid
11
+ per-session *or per-project* databases. That projection is deliberately
12
+ derived, rebuildable and disposable.
13
+
14
+ The session index (`entries`, `entries_fts`) is the table that actually grows
15
+ with total indexed history, and automatic eviction remains deferred until an
16
+ on-demand rehydration path is verified. A single file therefore also means a
17
+ single growth boundary, a single write lock and a single point of physical
18
+ reset shared by unrelated projects. Content is already partitioned logically
19
+ by `sessions.project_path`, `project_states`, `project_files` and
20
+ `project_memory_sessions`; the file boundary was the only thing missing.
21
+
22
+ The one piece of genuinely global learning is token calibration. Schema v10's
23
+ `token_calibration` is keyed by exact `provider + model + estimator_version`
24
+ and has **no project column**: a naive per-project split would restart
25
+ calibration for every project (minimum 3 samples, 8 accepted for the opt-in
26
+ `autoTune` expansion), weakening exactly the BPE/auto-tuning path it should
27
+ protect.
28
+
29
+ ## Decision
30
+
31
+ Add `storage.scope: "agent" | "project"` (default `"project"`) to the storage
32
+ configuration:
33
+
34
+ - `agent` keeps the previous shared-database behavior and remains selectable.
35
+ - `project` derives one database per trusted canonical project root:
36
+ `projects/<sha256(canonicalRoot)[0..32]>.db`, next to the configured agent
37
+ database. Untrusted projects, broad roots (home directory, filesystem root)
38
+ and any resolution failure fall back to the agent database.
39
+
40
+ Both files receive the **same schema and the same migrations**; there is no
41
+ schema fork and migrations 1–15 are untouched. The split changes only which
42
+ repository each handle is used for:
43
+
44
+ - the agent database keeps `token_calibration` (shared learning);
45
+ - the project database keeps the session index, project index, context
46
+ manifests, summary graph, memory/pin projections, embeddings, quality
47
+ samples and artifact metadata; artifact object bytes move to
48
+ `projects/artifacts/<project-digest>/` so the orphan garbage collector,
49
+ which only sees references in the current database, can never delete
50
+ another project's objects;
51
+ - `resource_leases` and the client lease stay per file, protecting each
52
+ database independently.
53
+
54
+ With `project`, a calibration sample is derived from the project manifest and
55
+ inserted into the agent database with `manifest_id = NULL` (the column and its
56
+ partial unique index already allow this). The two writes are intentionally
57
+ **not** one cross-database transaction: losing one calibration sample is
58
+ harmless, whereas losing manifest/usage consistency is not. In `agent` scope
59
+ the previous single transaction is unchanged. Manifest pruning only detaches
60
+ calibration rows in `agent` scope; the project database's calibration table
61
+ stays empty.
62
+
63
+ Every project database starts empty and is rebuilt from canonical Pi JSONL,
64
+ project files and memory/pin `CustomEntry` records. The previous shared
65
+ database is left untouched: with the default change, existing users cold-start
66
+ their per-project indexes while their existing calibration remains available
67
+ in the agent database.
68
+
69
+ ## Consequences
70
+
71
+ - Physical isolation per project: separate growth, separate write lock,
72
+ "reset project state" = remove one file, and no eviction needed to bound a
73
+ single project's index.
74
+ - Calibration stays global: a sample learned in project A is immediately
75
+ visible in project B for the same provider/model/estimator.
76
+ - Default behavior changes. A pre-existing shared database becomes the agent
77
+ database (calibration and old manifests) and is no longer the active
78
+ projection for new sessions. This requires a minor release and release
79
+ notes; `storage.scope: "agent"` restores the old layout.
80
+ - Maintenance and diagnostics become per file: `/context storage` reports the
81
+ active project database and, when split, the shared agent database;
82
+ `ds4-context-storage inspect|compact|recover --database <path>` must be
83
+ pointed at each file.
84
+ - `storage.databasePath` now names the agent database; project databases
85
+ derive from its directory. A manually configured per-project path keeps
86
+ working, but the derived `projects/` directory is the supported layout.
87
+ - D1 is superseded by this ADR; the development plan keeps the original text
88
+ with an explicit amendment pointer.
@@ -67,5 +67,6 @@ The initial decisions from the development plan are accepted:
67
67
  | [061](061-compaction-latency.md) | Bound compaction update calls, input budgets, concurrent segments and phase timings | Accepted |
68
68
  | [062](062-cache-aware-context-planning.md) | Opt-in cache-aware tail planning using model pricing and observed cache shares | Accepted |
69
69
  | [063](063-fts-rowid-key-mappings.md) | Resolve FTS key deletes through rowid mapping tables | Accepted |
70
+ | [064](064-per-project-databases-with-shared-calibration.md) | Split project state into per-project databases and keep token calibration shared | Accepted |
70
71
 
71
72
  Each decision will receive a dedicated record when implementation pressure introduces alternatives or consequences not already covered by the development plan.
@@ -45,7 +45,7 @@ actualInputTokens = input + cacheRead + cacheWrite
45
45
  rawCalibrationRatio = actualInputTokens / estimatedInputTokens
46
46
  ```
47
47
 
48
- The existing scalar SQLite columns are authoritative for persisted usage. `message_end` updates those columns and the calibration sample without reading or rewriting `manifest_json`; repository reads hydrate the usage fields from the scalars. The raw estimate is retained even after calibration so ratios cannot recursively calibrate already-corrected values. The next call reads only the exact provider/model window and applies bounded median/MAD outlier rejection. Error, aborted, missing, zero-usage, duplicate, or uncorrelated responses do not create calibration samples. Every manifest is correlated with at most one assistant response. If a concurrently pruned or oversize-skipped manifest cannot be correlated, DS4 keeps only bounded volatile calibration and does not retry against another row. See [`MODEL_AWARENESS.md`](MODEL_AWARENESS.md).
48
+ The existing scalar SQLite columns are authoritative for persisted usage. `message_end` updates those columns and the calibration sample without reading or rewriting `manifest_json`; repository reads hydrate the usage fields from the scalars. The raw estimate is retained even after calibration so ratios cannot recursively calibrate already-corrected values. The next call reads only the exact provider/model/**estimator-version** window and applies bounded median/MAD outlier rejection. When present, `modelAwareness.drift` and `modelAwareness.autoTune` contain only aggregate warning/tuning metadata (no message contents). Error, aborted, missing, zero-usage, duplicate, or uncorrelated responses do not create calibration samples. Every manifest is correlated with at most one assistant response. If a concurrently pruned or oversize-skipped manifest cannot be correlated, DS4 keeps only bounded volatile calibration and does not retry against another row. See [`MODEL_AWARENESS.md`](MODEL_AWARENESS.md).
49
49
 
50
50
  ## Retention
51
51
 
@@ -1,5 +1,7 @@
1
1
  # Advanced Model Awareness
2
2
 
3
+ **Validation status (OpenRouter, synthetic DS4/Pi sessions):** GPT-6-sol, GPT-6-luna and GPT-5.6-terra have each demonstrated multi-turn BPE calibration, a genuinely larger opt-in tail budget, provider usage above the native-window headroom threshold, and withdrawal of that expansion on the next turn in the same session. Historical retrieval and a live tool cycle were verified separately for all three. These bounded experiments do not establish universal tokenizer accuracy, provider-side compaction behavior, or production retrieval quality. `chars-v1` remains the default; BPE and `autoTune` remain opt-in. See the measured runs below.
4
+
3
5
  M11 resolves an independent deterministic planning profile for every exact `provider/model`. Pi's model descriptor remains the default source for context window, output ceiling, reasoning, and image support; explicit DS4 overrides can repair provider metadata or tune category limits without changing canonical session state.
4
6
 
5
7
  ## Profile resolution
@@ -68,10 +70,10 @@ Each successful assistant response is correlated with one pending Context Manife
68
70
 
69
71
  ```text
70
72
  actualInputTokens = input + cacheRead + cacheWrite
71
- ratio = actualInputTokens / raw chars-v1 estimate
73
+ ratio = actualInputTokens / raw selected-estimator estimate
72
74
  ```
73
75
 
74
- Calibration is isolated by exact provider and model ID. The latest configured window is validated and processed deterministically:
76
+ Calibration is isolated by exact provider, model ID **and estimator version**; switching from `chars-v1` to BPE never reuses the old ratios. The latest configured window is validated and processed deterministically:
75
77
 
76
78
  1. reject invalid or zero samples;
77
79
  2. reject ratios outside the configured hard bounds;
@@ -82,7 +84,40 @@ Calibration is isolated by exact provider and model ID. The latest configured wi
82
84
 
83
85
  Until enough samples exist, the multiplier remains `1.0`. The default window is 24 samples, minimum is 3, and hard ratio bounds are 0.5–2.0. Outliers remain counted in metadata but do not influence the applied ratio.
84
86
 
85
- Provider-token capacities are computed first from context window, output reserve, safety margin, and policy ratios. Planner limits are then converted into local-estimator units by dividing by the accepted multiplier. Raw manifest estimates remain uncalibrated so future samples do not feed a corrected estimate back into itself. Adaptive tail/history/project budgets use the same conversion.
87
+ Provider-token capacities are computed first from context window, output reserve, safety margin, and policy ratios. Global hard/soft/preferred input limits are converted into local-estimator units using **at least** a multiplier of 1: observed underestimation may reduce them, but apparent overestimation on short prompts never raises them above nominal model limits. Raw manifest estimates remain uncalibrated so future samples do not feed a corrected estimate back into itself. Adaptive tail/history/project budgets use the accepted multiplier but remain capped by the configured `context.*` maxima even after conversion.
88
+
89
+ Calibration samples are global learning, not project data: with the default `storage.scope: "project"` the manifests live in the per-project database while `token_calibration` stays in the shared agent database (see [ADR 064](ADR/064-per-project-databases-with-shared-calibration.md)). A sample learned in one project therefore applies to every other project for the same provider/model/estimator.
90
+
91
+ ### Optional BPE estimator and measured budget tuning
92
+
93
+ The default remains `chars-v1`. To opt a profile into local OpenAI `o200k_base` BPE text counting (without fetching vocabulary over the network):
94
+
95
+ ```json
96
+ {
97
+ "modelAwareness": {
98
+ "overrides": { "openai/gpt-4o": { "tokenEstimator": "o200k-base-v1" } },
99
+ "autoTune": true
100
+ }
101
+ }
102
+ ```
103
+
104
+ Use BPE only for a model known to share that encoding. The estimator counts text with BPE in the Pi observer, managed planner and fixed prompt/tool estimates; message wrappers, images, reasoning, provider-specific serialization, indexed retrieval hints, and persisted summaries remain estimates. It does **not** make token counts exact or enable provider-side continuation. If the provider's actual input tokens drift persistently from estimates (median ≥1.25 or ≤0.80 after the calibration minimum, or repeated hard-bound outliers), `/context model` and status surface a metadata-only warning; no prompt content is logged.
105
+
106
+ `autoTune` is opt-in and starts neutral. After at least eight accepted calibration samples, it expands *automatic* recent-tail, history and project ceilings by at most 12.5% of their nominal values, only if **every valid provider-usage sample** in the calibration window — including ratio outliers excluded from ratio calibration — remains at or below 60% of the current preferred/hard input target. It never raises a ceiling beyond the configured context cap, never overrides an explicit model-specific category limit, and never changes the provider's window, output reserve, safety margin or hard/soft input limits. Without enough evidence or headroom it stays neutral. Any calibration multiplier is applied after this small expansion; tune decisions and warnings are recorded as metadata in the manifest.
107
+
108
+ ## Validation goal for BPE and auto-tuning
109
+
110
+ **User-requested outcome:** verify long conversations, tools, retrieval, calibration across turns, and auto-tuning safety on the context actually built by DS4/Pi. Single-turn synthetic token probes are preliminary evidence only; they do **not** satisfy this goal or justify declaring the feature complete.
111
+
112
+ The validation must cover, and report results separately for:
113
+
114
+ 1. Long multi-turn sessions, including large context windows, compaction boundaries, and preservation of the current request.
115
+ 2. Tool calls and results (including large results): atomic selection, offload, and the estimated versus observed token count.
116
+ 3. Retrieved history and project context: relevance, budget selection, and whether relevant older material survives planning.
117
+ 4. Consecutive turns on the **same** provider/model and estimator profile: provider usage correlation, accepted/rejected calibration samples, changes to the applied ratio, and isolation when models or estimators change.
118
+ 5. Opt-in auto-tuning: evidence thresholds, configured and hard limits, outliers, insufficient headroom, and no regression in retrieval or context integrity when limits expand.
119
+
120
+ Use local, deterministic integration tests where possible. If real provider calls are needed, require a **new explicit call and input-size limit** before sending anything; the earlier single-turn budgets are exhausted. Report any scenario not verified as *not verified*, rather than treating the synthetic pilot as an end-to-end result. Do not change the default estimator or enable auto-tuning by default on the strength of that pilot.
86
121
 
87
122
  ## Provider cache metrics
88
123
 
@@ -126,6 +161,42 @@ Use:
126
161
  /context status
127
162
  ```
128
163
 
164
+ ## Local validation progress and outstanding provider evidence
165
+
166
+ `tests/integration/model-aware-real-context.test.ts` exercises the Pi `context` and `message_end` hooks over canonical multi-turn JSONL without any network transport. A long branch with a tool call/result verifies atomic selection and retrieval of an older decision under a BPE-managed budget. Successive turns verify that eight correlated usage records are required before opt-in expansion, that estimator/model switches isolate calibration, and that a new near-limit usage record withdraws the expansion even when its ratio is excluded from calibration. `tests/unit/model-aware-estimator.test.ts` also checks explicit overrides, configured caps, and the high-usage outlier regression.
167
+
168
+ **Local integration checks inject usage; they are not independent provider measurements.** A separately authorized, bounded live probe (`scripts/verify-model-aware-real-session.mjs`) used a temporary Pi JSONL session, DS4's actual hooks, synthetic content, `openrouter/openai/gpt-4o-mini`, and Pi-normalized provider usage. Eight successive short calls accumulated calibration samples; on the ninth call the manifest showed eight accepted samples, applied ratio 0.677157, and opt-in `autoTune: expanded`. That ninth call did **not** contain the intended appended long branch: Pi had already constructed its runner, and its manifest estimated only 1,940 input tokens. It is not evidence for long-context tuning.
169
+
170
+ A second, fresh Pi session with the long branch seeded *before* runner construction produced 83 original messages; DS4 selected 21 groups, excluded 21, retrieved the older `cobalt-713` decision, and included both the synthetic tool-call assistant entry and tool result. For that request, the BPE-based manifest estimate was **23,492** input tokens and provider usage was **23,196**. This is **one** live, untuned long-session/tool/retrieval measurement; it does not prove stable drift across lengths or models. The earlier [single-turn pilots](PROVIDER_TOKEN_DRIFT_BENCHMARK.md) remain separate. The two runs together attempted 10 calls and conservatively reserved 217,678 of the authorized 250,000 estimated-input-token limit (per-call ceiling 64,000; 16-call ceiling). All temporary session data was deleted.
171
+
172
+ A subsequent, separately authorized probe (`scripts/verify-model-aware-calibrated-session.mjs`) seeded a 28-turn synthetic history and tool call/result **before starting the same Pi session**. Eight real provider calls calibrated BPE on that long context; the ninth measured **23,926 estimated / 23,486 actual input tokens**, eight accepted samples, `autoTune: expanded`, and retrieval of the old decision with the historical tool call and result included. This verifies expansion on a real long context after calibration in one session, not just across separate sessions. Across this probe **11 calls** reserved **588,052/600,000** estimated input tokens (80,000 per-call limit); no further provider requests were made under that authorization.
173
+
174
+ The attempted higher-occupancy request used **37,552 estimated / 37,464 actual tokens**, below the 60%-of-target withdrawal threshold (**53,760** for this model/configuration). The following turn still reported `expanded`: **withdrawal is not verified live**. That larger turn also excluded the historical tool/retrieval groups, an observed quality limit under that selected context, not proof of an auto-tuning regression. The local Pi compaction-boundary test covers summary preservation under BPE. A separately bounded **one-request** live compaction probe (`scripts/verify-model-aware-compaction-session.mjs`) preseeded the canonical Pi compaction entry before starting Pi; the DS4 managed manifest included its summary and measured **250 estimated / 244 actual** provider input tokens. This used one further call under the third authorization, bringing that budget to **10 calls / 181,069 estimated input tokens reserved**. It checks the provider-bound context *after* a synthetic compaction entry, not live summary generation. At that point, live tool execution and high-usage withdrawal remained unverified; both were tested in the later authorized round below. Model variety and production retrieval quality still remain outside these bounded probes. A third bounded probe (`scripts/verify-model-aware-safety-session.mjs`) reached **60,509 actual input tokens** (>53,760), but only **seven** of eight preceding short-call samples had been accepted; the manifest at the high-usage request showed seven accepted samples and `insufficient-samples`, not an expanded policy. Thus it **does not** test withdrawal. The safety probe attempted **9 calls**, reserving **169,006** estimated input tokens, and stopped without a follow-up call. It now gates the expensive request on eight *accepted* samples and seeds a stable synthetic prefix, but has not been rerun. Its temporary session was deleted; the remaining third-round budget after the separate compaction request cannot fit a new calibration/high-usage/follow-up sequence under the preflight limits. No live tool execution was attempted in that round because the automatic multi-request cycle was not yet safely capped per request. Those results alone did not verify withdrawal or tool execution.
175
+
176
+ A final, separately authorized round verified the outstanding **synthetic** provider-path scenarios on the same OpenRouter/GPT-4o-mini profile. After a stable synthetic prefix, eight usage samples were accepted; the ninth short turn showed `expanded`. The next request reported **60,708 estimated / 60,640 actual** input tokens, above the **53,760** headroom threshold; the following turn reported **60,795 estimated / 60,722 actual** and `no-headroom` (recent-tail limit decreased from **27,699** to **24,564**). `scripts/verify-model-aware-safety-session.mjs` used **11/15** calls and **292,818/350,000** estimated-input tokens reserved (maximum 80,000 per call). This demonstrates withdrawal on measured provider usage, not merely a synthetic injected sample.
177
+
178
+ With the remaining budget, `scripts/verify-model-aware-live-tool.mjs` gated **each** Pi provider request at `ModelRuntime.streamSimple`, disabled cache warming/retry/compaction, and executed a single synthetic tool. The first request included its tool schema and used **105** provider input tokens; the second included the actual tool result and used **139**. Exactly **two** requests and one tool execution occurred. The final fourth-round total was **13/15 calls, 317,119/350,000** estimated-input tokens reserved. Only aggregate counters and booleans were reported; temporary synthetic sessions were removed.
179
+
180
+ These bounded live measurements, together with the canonical Pi JSONL integration tests, cover long context, historical and executed tools, retrieval, between-turn calibration, compaction-boundary context, and auto-tuning withdrawal **for this synthetic OpenRouter/GPT-4o-mini profile**. They do not establish universal accuracy for other providers/models, real private sessions, arbitrary compaction generation, or production retrieval quality.
181
+
182
+ ### Native-window checks for Sol, Luna and Terra — separately authorized
183
+
184
+ The Pi catalog exposed a **1,050,000-token window** for each OpenRouter model `openai/gpt-6-sol`, `openai/gpt-6-luna`, and `openai/gpt-5.6-terra`. One additional authorization limited their *combined* probes to **42 provider attempts, 600,000 estimated input tokens / 2,500,000 controlled characters per request, and 4,800,000 estimated tokens / 15,000,000 controlled characters total**. `scripts/verify-model-aware-triad.mjs` and `scripts/verify-model-aware-triad-followup.mjs` gated **every** Pi request at `ModelRuntime.streamSimple`, including tool continuations; no real sessions, credentials, prompt bodies, responses or raw upstream errors were logged. Temporary canonical Pi JSONL sessions were deleted.
185
+
186
+ The first probe consumed **30 calls / 1,433,445 tokens / 3,679,430 characters reserved**. For *each* model, nine calibration calls in the same seeded long session led to at least eight accepted usage samples. The tenth call measured **36,027 BPE-estimated / 35,503 Pi-normalized actual input tokens**, `autoTune: expanded`, with the historical tool call and result included. **This did not verify retrieval**: the decision was still in the 64k recent tail and therefore had not been fetched by retrieval; the probe correctly stopped before tool execution and the high-occupancy turn. At this native profile, the reported `expanded` status did **not** demonstrate a larger recent-tail limit: its configured 64k cap was already saturated.
187
+
188
+ The second probe seeded **70 turns before Pi session construction**, putting the old decision outside the default recent tail. It consumed nine more calls (one retrieval request and an actual **two-request, one-execution** synthetic tool cycle per model). The decision was both retrieved and included, the historical tool call/result stayed included, and the new tool schema appeared in the first provider request with its actual result in the second. Per-model provider usage for the retrieval request was **63,050 / 63,049 / 63,053** tokens (Sol/Luna/Terra), against BPE estimates **63,699 / 63,698 / 63,702**. All three tool cycles completed. These are measurements of the DS4/Pi managed context, not three single-turn tokenizer prompts.
189
+
190
+ The three remaining authorized requests measured large BPE-managed input, **one per model**. The selected current user turn was included, and Pi reported **381,501 / 381,502 / 381,502** input tokens against manifest estimates **381,593 / 381,594 / 381,594**. The historical call/result were present, but the old decision was **not** selected; the oversized prompt did not ask about that decision, so this is **not** a valid retrieval-quality test. The request generator targeted 462k from the *raw* canonical history, but DS4 excluded older history and the selected provider input remained ~381.5k. Thus all three requests were **below** the native-window auto-tuning withdrawal threshold of **441,000** (60% of the 735k preferred target), and far below the full 1.05M window. They were fresh sessions without eight accepted calibration samples; **none tests `expanded → high provider usage → no-headroom` on these models**. Do not report the triad's auto-tuning safety or category expansion as live-verified at its native window; the deterministic local matrix in `tests/unit/model-aware-estimator.test.ts` covers only injected samples and caps. The entire authorization was used: **42/42 calls**, **3,296,138/4,800,000** estimated input tokens and **9,357,266/15,000,000** controlled characters reserved. No more provider calls are authorized under it. The historical probes now reject `--live` under this exhausted authorization. **After** those measurements, their local high-input sizing was corrected to require at least 470k BPE tokens in the selected current user text (rather than in raw canonical history), with a second gate immediately before provider transport; unit tests cover the discarded-history regression and per-call preflight. This correction was **not** run against a provider and cannot retroactively validate withdrawal.
191
+
192
+ Across these specific synthetic shapes, BPE manifests slightly overestimated provider usage, but no universal correction follows. At this earlier point native-window compaction generation, production retrieval quality and live high-usage auto-tune withdrawal for Sol/Luna/Terra were unverified; the later withdrawal runs below close **only the last of those gaps**. The later probes check actual category growth (not merely `expanded` status), a selected current turn above 470k BPE tokens, usage above 441k, and withdrawal in the same Pi session. Their temporary opt-in category ceilings are 80k/40k/40k because the default native tail ceiling is already equal to its configured maximum of 64k and cannot grow; the probes do not establish expansion under unchanged category defaults. BPE and auto-tuning remain opt-in; no default or provider-storage policy was changed.
193
+
194
+ ### Native-window withdrawal round (later, opt-in categories)
195
+
196
+ With a separate explicit cap of 39 calls, 3,800,000 reserved input tokens and 11,000,000 controlled characters, the same-session Pi/DS4 probe used temporary category ceilings 80k/40k/40k to expose a **real** tail expansion. It ran 34 provider requests (3,375,536 estimated input tokens reserved; 9,845,274 controlled characters). For **GPT-6-sol and GPT-6-luna**, each session accepted at least eight calibration samples; the expanded tail exceeded the same-ratio no-tune baseline (about 72,922 versus 64,819 estimator tokens). A later request used **475,703 provider input tokens** (BPE estimate 475,790), above the 441,000 provider-token headroom threshold. On the next request, `autoTune` was `no-headroom` and the tail equalled its no-tune baseline (64,805); actual provider input was 471,846/471,844. The selected current turn, historical decision, call and result were present before the large request. The large input and follow-up did **not** retain the unrelated old decision/tool group; these measurements do not establish production retrieval quality.
197
+
198
+ For **GPT-5.6-terra**, calibration and real expansion passed in that first round, but its high-input request was blocked **before transport** by the cumulative budget/character caps: the probe had incorrectly budgeted the follow-up as short even though Pi resends the large preceding turn. The temporary session was then disposed. A **separately authorized Terra-only run** used a fresh synthetic Pi/DS4 session and counted *both* large requests. It made 12 provider calls (1,449,079 estimated tokens and 4,310,843 characters reserved, under separate 13-call / 1,900,000-token / 6,000,000-character maxima). After at least eight accepted calibration samples, Terra's tail grew to 72,922 estimator tokens versus a same-ratio untuned baseline of 64,819. The large request used **475,707 provider input tokens** (BPE estimate 475,794), above the 441,000 threshold. On the next turn, provider input was **471,852** (estimate 471,929), `autoTune` was `no-headroom`, and the tail returned to **64,805**, exactly its untuned baseline at that turn's ratio. The historical decision/call/result and current turn were included before the large request; unrelated older groups were excluded during the large request. **This verifies the defined live high-usage expansion/withdrawal scenario for all three models, not a general retrieval or compaction guarantee.** Both probes are now locked against replay; defaults remain unchanged.
199
+
129
200
  ## Performance and tests
130
201
 
131
- `tests/benchmarks/model-awareness.bench.ts` measures a bounded 200-sample calibration analysis and repeated 32k/128k/200k profile resolution. Unit and golden tests cover deterministic tiers, override precedence, robust outlier rejection, cache accounting, and calibrated budgets. Integration tests switch local/remote providers and 32k/128k/200k models while checking profile isolation, privacy re-enforcement, canonical JSONL preservation, SQLite cache metrics, and profile reuse.
202
+ `tests/benchmarks/model-awareness.bench.ts` measures a bounded 200-sample calibration analysis and repeated 32k/128k/200k profile resolution. Unit and golden tests cover deterministic tiers, override precedence, robust outlier rejection, cache accounting, and calibrated budgets. Integration tests switch local/remote providers and 32k/128k/200k models while checking profile isolation, privacy re-enforcement, canonical JSONL preservation, SQLite cache metrics, and profile reuse. For a separately authorized, bounded real-provider measurement of both estimators against Pi SDK usage, see [PROVIDER_TOKEN_DRIFT_BENCHMARK.md](PROVIDER_TOKEN_DRIFT_BENCHMARK.md); it does not substitute for DS4 manifest measurements in a live session.
@@ -0,0 +1,85 @@
1
+ # Live provider token-drift probe (opt-in)
2
+
3
+ This developer-only, **source-checkout-only** probe (the script is not included in the published npm package) compares DS4's raw `chars-v1` and `o200k-base-v1` estimates with **Pi SDK provider usage** for synthetic, single-turn text. It does not start a Pi session, read session JSONL or SQLite, record responses, or change DS4 configuration. It never prints API keys, requests, response bodies or raw upstream error messages. It requires local Pi provider authentication and network access. The probe runs only with `--live`; without that flag it prints the bounded plan.
4
+
5
+ ```bash
6
+ npm run build:core
7
+ node scripts/compare-provider-token-drift.mjs \
8
+ --model openrouter/openai/gpt-4o-mini \
9
+ --model deepseek/deepseek-v4-flash \
10
+ --model openai-codex/gpt-5.4-mini
11
+ # After inspecting the dry-run budget, add --live to make real calls.
12
+ ```
13
+
14
+ The script accepts repeated `--model provider/model-id` and optional `--sizes 512,4096,48000`. **Per invocation**, it refuses more than 12 model calls or 200,000 total input characters (system prompt included), disables SDK request retries, and imposes a 45-second deadline per call. Multiple invocations have separate caps: track their cumulative spend yourself. Output is one metadata-only JSON report on stdout. The SDK's catalog-derived USD cost is an estimate, not a bill; a provider can still charge for failures or retries outside this script's control. If you pipe output to a file, keep it outside the repo unless you intentionally want to publish the aggregate metadata.
15
+
16
+ For each successful call, `actualInputTokens = usage.input + usage.cacheRead + usage.cacheWrite`; the cached fractions are **not** extra tokens on top of that sum. `residualTokens = actual - rawEstimate` (positive means underestimation), and `underestimationPctOfActual = max(0, residual) / actual × 100`. Adjacent slopes use differences across prompt sizes for a single exact provider/model, removing most fixed framing. BPE counts only text; both estimators use the same DS4 message/system overhead. No raw provider payload is available from this probe, so the result does **not** prove the accuracy of DS4's full observer in a live Pi extension chain (tools, images, privacy, cache/continuation, and later extensions differ). Compare `/context model`, `/context tokens`, and manifest `actualInputTokens` versus `estimatedInputTokens` in an ordinary consented Pi session for that second step. Never copy session text or auth files into the report.
17
+
18
+ ## First observed run — 24 September 2026, 10:09–10:13 UTC
19
+
20
+ User-approved cumulative envelope: 12 call attempts and 200,000 input characters. Three invocations totalled **12 attempts and 195,176 planned input characters**: 9 calls/158,382 characters, a single Codex diagnostic call/574 characters, then 2 calls/36,220 characters. Success: 8; failure/missing usage: 4. Synthetic multilingual/code-like repeated text; no private session data. Sizes below are user-message characters; system prompt is included in estimates and the total envelope. These are observations, **not** a representative multi-session cost or quality benchmark.
21
+
22
+ | Provider / requested model | User chars | Real input¹ | `chars-v1` | BPE `o200k-base-v1` | BPE residual (real − estimate) |
23
+ |---|---:|---:|---:|---:|---:|
24
+ | OpenRouter / `openai/gpt-4o-mini` | 512 | 190 | 161 | 196 | −6 |
25
+ | OpenRouter / `openai/gpt-4o-mini` | 4,096 | 1,436 | 1,057 | 1,442 | −6 |
26
+ | OpenRouter / `openai/gpt-4o-mini` | 48,000 | 16,668 | 12,033 | 16,674 | −6 |
27
+ | DeepSeek / `deepseek-v4-flash`² | 512 | 179 | 161 | 196 | −17 |
28
+ | DeepSeek / `deepseek-v4-flash`² | 4,096 | 1,388 | 1,057 | 1,442 | −54 |
29
+ | DeepSeek / `deepseek-v4-flash`² | 48,000 | 16,172 | 12,033 | 16,674 | −502 |
30
+ | OpenRouter / `deepseek/deepseek-v4-flash` | 4,096 | 1,387 | 1,057 | 1,442 | −55 |
31
+ | OpenRouter / `deepseek/deepseek-v4-flash` | 32,000 | 10,785 | 8,033 | 11,124 | −339 |
32
+
33
+ ¹ Pi's normalized provider usage: input + cache read + cache write. Some long requests reported cached reads (up to 1,408 tokens); they were included exactly once. Successful responses produced 1–3 output tokens. ² DeepSeek reported response model `deepseek-flash`, an alias of the requested ID; no equivalence to OpenRouter routing is assumed.
34
+
35
+ - OpenRouter GPT-4o-mini: raw `chars-v1` median actual/estimate **1.358562**; raw BPE median **0.995839**. At 48k chars, chars/4 underestimated by 4,635 tokens (27.81% of actual), while BPE overestimated by 6 tokens. Adjacent BPE slopes were 1.000 and 1.000.
36
+ - Direct DeepSeek: raw chars median **1.313150**; BPE median **0.962552**. At 48k chars, chars/4 underestimated by 4,139 tokens (25.59% of actual), while BPE overestimated by 502 tokens. Adjacent BPE slopes were 0.970305 and 0.970588. Close agreement here does **not** establish that DeepSeek uses OpenAI's tokenizer or justify turning BPE on for that model by default.
37
+ - OpenRouter DeepSeek: two valid samples; raw chars median **1.327396**, BPE median **0.965692**. At 32k chars, chars/4 underestimated by 2,752 tokens; BPE overestimated by 339. Two samples are insufficient to validate a persistent drift warning or tuning.
38
+ - `openai-codex/gpt-5.4-mini`: three initial attempts produced no usable usage; one later 512-character diagnostic returned `request-failed` with sanitized category `other` and no HTTP status. The local Pi auth resolver did return a credential, but the reason for the probe failure remains **unverified**. Do not infer its tokenizer drift, account availability in the interactive TUI, or provider-side token consumption from these calls. Direct `openai/gpt-4o-mini` API auth was not configured in the Pi runtime used by this probe; OpenRouter is a distinct route.
39
+
40
+ **Interpretation:** Three sizes with one call each do not establish statistical reliability or cover large context windows, tools, images, prefixes reused across turns, or output-heavy workflows. `chars-v1` remains the default; existing per-model calibration may compensate after enough **accepted same-profile samples**. The first small sample can be excluded by MAD filtering, so three different-size wire calls need not become three accepted DS4 calibration samples. Keep BPE and `autoTune` opt-in; do not port Hub's fixed 3.5% margin from these measurements. For promotion or automatic budget changes, collect a repeated same-size and mixed-size sample with matching DS4 manifest estimates and actual provider usage under an explicitly authorized, separately bounded run.
41
+
42
+ ## Second observed run — 24 September 2026, 10:35 UTC
43
+
44
+ A separate, explicitly approved envelope covered **12 call attempts and up to 100,000 input characters**. The dry run planned exactly 12 attempts and 99,816 characters: `--sizes 512,16000` for three requested OpenAI models through each of OpenRouter and Codex. The live probe produced six OpenRouter usages and six Codex failures; it did not read session data.
45
+
46
+ | Route / requested model | User chars | Real input¹ | `chars-v1` | BPE `o200k-base-v1` | BPE residual |
47
+ |---|---:|---:|---:|---:|---:|
48
+ | OpenRouter / `openai/gpt-6-sol` | 512 | 189 | 161 | 196 | −7 |
49
+ | OpenRouter / `openai/gpt-6-sol` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
50
+ | OpenRouter / `openai/gpt-6-luna` | 512 | 189 | 161 | 196 | −7 |
51
+ | OpenRouter / `openai/gpt-6-luna` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
52
+ | OpenRouter / `openai/gpt-5.6-terra` | 512 | 189 | 161 | 196 | −7 |
53
+ | OpenRouter / `openai/gpt-5.6-terra` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
54
+
55
+ ¹ Pi-normalized input includes cached tokens once. Each 16,000-character call reported `cacheWriteTokens = 5,561` and zero cache-read tokens. Those were **writes**, not cache hits. Each successful response reported five output tokens. The six equal input counts describe these identical synthetic prompts on this route; they are not proof of identical tokenizers or of how the direct Codex route would count a DS4 session. For each model, the adjacent BPE size slope was exactly 1.000; at 16k characters, raw `chars-v1` underestimated by 1,531 tokens (27.52% of actual), versus BPE overestimating by seven.
56
+
57
+ For `openai-codex/gpt-6-sol`, `gpt-6-luna`, and `gpt-5.6-terra`, **both sizes failed** without provider usage. The sanitized error classifier returned `quota` for all six, with no HTTP status captured. This is evidence of an SDK-path quota/limit error classification, **not** a verified account balance, an HTTP rejection, or a tokenizer measurement. It does not retroactively identify the earlier `gpt-5.4-mini` failure (classified `other`). No further calls were made beyond this run's approved envelope. Two sizes per exact model remain insufficient for DS4 calibration or auto-tuning conclusions.
58
+
59
+ ## Isolated Pi + DS4 managed-context pilot — 24 September 2026
60
+
61
+ For a future, **separately authorized** run: build the core (`npm run build:core`), inspect the planned sizes and character count with `node scripts/compare-ds4-manifest-usage.mjs`, then use `node scripts/compare-ds4-manifest-usage.mjs --live` only after setting a new call/character budget. The default mode is pinned to OpenRouter `openai/gpt-6-sol` and a maximum of 12 calls / 120,000 controlled prompt characters per invocation. `--comparison` instead preflights **both** OpenRouter `openai/gpt-6-luna` and `openai/gpt-5.6-terra`, capped at 24 calls / 240,000 controlled characters combined; `--comparison --live` requires its own authorization. A new authorization is required even when repeating either exact plan.
62
+
63
+ The user separately authorized **up to 12 calls / 120,000 controlled input characters** to OpenRouter `openai/gpt-6-sol` only. `scripts/compare-ds4-manifest-usage.mjs` first checked a local sandbox without provider traffic, then made **12 single-turn calls**, with three repeats at each size (512, 4,096, 12,000 and 20,000 synthetic user characters). Planned synthetic user + fixed system text was **110,568 characters**. This limit counts controlled prompt text, **not** Pi-generated framing or serialized protocol bytes. No existing Pi session history was used; each Pi session and DS4 database was isolated in a temporary directory, with no tools, project content, native continuation, or auto-tuning. DS4's managed `context` hook and its `o200k-base-v1` profile override were active. The script stops on a missing/mismatched manifest or missing usage rather than fabricating a measurement.
64
+
65
+ | Synthetic user chars | First observed DS4 manifest estimate | First actual input¹ | First residual (actual − estimate) | Repeats |
66
+ |---:|---:|---:|---:|---:|
67
+ | 512 | 220 | 209 | −11 | 3 |
68
+ | 4,096 | 1,467 | 1,456 | −11 | 3 |
69
+ | 12,000 | 4,210 | 4,199 | −11 | 3 |
70
+ | 20,000 | 6,983 | 6,972 | −11 | 3 |
71
+
72
+ All **12** calls had matching `managed` manifests and Pi-normalized provider usage, with **exactly −11 tokens** of residual in each call. Token counts varied by about one or two between identically sized repeats, but the residual stayed fixed. The largest relative overestimate was **5.26% of actual input** on the first 512-character call (11/209); at 20,000 characters it was about **0.16%**. `cacheReadTokens` was zero, while longer calls reported mostly cache **writes**; no claim of cache hits or exact billed usage follows from those fields.
73
+
74
+ ¹ `totalInputTokens` in the DS4 manifest (the same Pi-normalized `input + cacheRead + cacheWrite` metric used by model-awareness calibration), not independently obtained provider invoice data. Because the selected estimator was BPE, this pilot did **not** produce an alternative `chars-v1` DS4 manifest for the same wire requests. These are fresh one-turn sessions with one synthetic prompt shape and a maximum of ~7k actual tokens; they do not validate multi-turn prefixes, tool schemas/results, retrieval, real project text, larger windows, other requested models, or production auto-tuning. The 12 samples are independent sessions, **not** 12 accepted samples in one persistent DS4 calibration profile. Keep BPE opt-in and `autoTune` off by default; do not apply an inferred fixed framing correction or Hub's fixed 3.5% drift allowance from this pilot.
75
+
76
+ ## Separately authorized Luna/Terra managed-context comparison — 24 September 2026
77
+
78
+ A subsequent authorization covered **at most 24 calls / 240,000 controlled synthetic input characters combined** on OpenRouter `openai/gpt-6-luna` and `openai/gpt-5.6-terra`; no Codex traffic was authorized. `node scripts/compare-ds4-manifest-usage.mjs --comparison` preflighted **24 calls / 221,136 controlled characters**. `--comparison --live` completed **12 calls per model**, with three fresh, isolated Pi+DS4 managed sessions per model at each synthetic user size (512, 4,096, 12,000, 20,000 characters). No existing session history, project content, tools, auto-tuning, or native continuation entered the requests. All 24 calls had matching manifests and Pi-normalized provider input usage; there were no reported failed rows.
79
+
80
+ | Model | First 512-character DS4 BPE estimate | First actual input¹ | Repeats per size | Residual on all 12 calls |
81
+ |---|---:|---:|---:|---:|
82
+ | OpenRouter / `openai/gpt-6-luna` | 220 | 209 | 3 | −11 tokens |
83
+ | OpenRouter / `openai/gpt-5.6-terra` | 219 | 208 | 3 | −11 tokens |
84
+
85
+ In every size group for both models, each of the three residuals was **−11 tokens**; each group reported zero `cacheReadTokens`. This replicates the small, constant **overestimate** observed for Sol on this specific one-turn synthetic shape. It does **not** establish a provider-independent 11-token correction, tokenizer equivalence across arbitrary text, long-window accuracy, model calibration from a continuous session, or safe budget auto-tuning. The combined 24-call authorization is exhausted. A later, separate authorization covered multi-turn, historical and executed tools, retrieval, and large selected inputs for all three OpenRouter models; see [native-window checks](MODEL_AWARENESS.md#native-window-checks-for-sol-luna-and-terra--separately-authorized) for results and the remaining unverified high-usage auto-tune withdrawal. Codex diagnosis still needs separate approval; keep BPE opt-in and `autoTune` off by default.
package/docs/STORAGE.md CHANGED
@@ -10,6 +10,19 @@ Pi's session JSONL is canonical for conversations and live project files are can
10
10
 
11
11
  The extension never edits or rewrites Pi JSONL or project source files. Manual memory/pin commands, confirmed `context_persistence` canonical writes, and learned-ranking feedback append versioned classified Pi `CustomEntry` records through Pi's official `appendEntry()` API. The tool does not write SQLite as a substitute for a canonical Pin or Memory append.
12
12
 
13
+ ## Storage scope
14
+
15
+ `storage.scope` selects where the disposable projection lives (see [ADR 064](ADR/064-per-project-databases-with-shared-calibration.md)):
16
+
17
+ - `agent` — the previous behavior: one shared `context.db` for every session and project.
18
+ - `project` (**default**) — one database per trusted canonical project root, derived as `projects/<sha256(root)[0..32]>.db` next to the agent database. Untrusted projects and broad roots (home directory, filesystem root) fall back to the agent database.
19
+
20
+ Both files receive the same schema and migrations. The agent database keeps only `token_calibration`, so a sample learned in one project is visible to every other project for the same provider/model/estimator; project databases keep the session index, project index, manifests, summary graph, memory/pin projections, embeddings, quality samples and artifact metadata. Project artifact object bytes move under `projects/artifacts/<project-digest>/` so garbage collection stays scoped to one project. A project database starts empty and is rebuilt from canonical JSONL and project files; the split never rewrites the previous shared database.
21
+
22
+ Calibration and manifest writes are intentionally not one cross-database transaction: a failed calibration insert loses at most one sample, while manifest/usage consistency stays inside the project database. Manifest pruning detaches calibration rows only in `agent` scope; in `project` scope the project database's calibration table stays empty. `storage.databasePath` names the agent database; the derived `projects/` directory is the supported layout.
23
+
24
+ Maintenance and diagnostics are per file: `/context storage` reports the active project database and, when split, the shared agent database; `ds4-context-storage inspect|compact|recover --database <path>` must be pointed at each file.
25
+
13
26
  The M19 non-Pi reference adapter owns a separate `ds4-runtime-session-v1` JSONL source selected by its host runtime. Its header binds runtime/session identity and the exact canonical project root; following records contain provenance-checked canonical messages. DS4 snapshots and capability diagnostics are disposable. `createReferenceHistory()` refuses overwrite, append uses a dedicated provenance-checked operation, files are mode `0600` where supported, and rebuild never edits this runtime-owned canonical file. Reference JSONL is not imported into Pi or `context.db`.
14
27
 
15
28
  M20 local KV state is entirely runtime-owned and volatile. Core returns only an in-memory eligibility fingerprint to the runtime port; it has no cache-handle field or serialization API. Prefixes, fingerprints, handles and provider outputs are absent from Pi/reference JSONL, Context Manifests, ranking artifacts and every SQLite table. Aggregate hit/miss/prefill counters live only on the adapter controller, and a restart safely resets them with the runtime cache. M20 adds no database migration.
@@ -83,7 +96,7 @@ Custom entries have empty lexical search text and never enter Pi context directl
83
96
 
84
97
  Schema v10 extends `context_manifests` and `token_calibration` with separate uncached-input, cache-read, and cache-write token columns. New calibration rows also carry the correlated manifest ID and explicit estimator version. Legacy pre-v10 samples migrate as `chars-v1` with their prior total stored as uncached input and zero cache fields; this preserves historical ratio behavior without inventing cache hits.
85
98
 
86
- Calibration rows are derived telemetry, isolated by exact provider/model and bounded to the latest configured window at read time. The runtime recomputes median/MAD outlier filtering deterministically; no learned model or mutable provider state is stored. Deleting the database loses calibration and cache history but never session content. Ephemeral sessions and configurations that disable manifest persistence keep only a bounded in-memory window.
99
+ Calibration rows are derived telemetry, isolated by exact provider/model and bounded to the latest configured window at read time. The runtime recomputes median/MAD outlier filtering deterministically; no learned model or mutable provider state is stored. Deleting the database loses calibration and cache history but never session content. Ephemeral sessions and configurations that disable manifest persistence keep only a bounded in-memory window. With `storage.scope: "project"`, calibration lives in the agent database while the manifest that produced the sample lives in the project database; the sample therefore carries no `manifest_id` and the two writes are separate transactions.
87
100
 
88
101
  ## Context quality samples
89
102
 
@@ -107,7 +120,7 @@ The volatile state is cleared on lifecycle/model/branch/compaction boundaries an
107
120
 
108
121
  Schema v8 splits content objects from source references. `artifact_objects` is keyed by SHA-256 and stores the private file path, MIME, byte size, verification timestamps, and integrity status. `artifacts` is keyed by a deterministic source-specific ID and references session/entry/tool identity plus original/condensed token estimates and an optional derived privacy classification in `metadata_json`. Equal bytes across calls or sessions deduplicate to one object while retaining independent provenance.
109
122
 
110
- Objects live under `ds4-context/artifacts/<sha-prefix>/<sha256>` with private permissions and atomic writes. Pi's full JSONL tool result remains canonical; the object file is a rebuildable local cache. No artifact content is stored in Context Manifests. Search recomputes SHA-256 and returns only bounded, redacted, JSON-quoted literal-match windows for a current-branch reference. The runtime reapplies the stored artifact classification before returning excerpts to the active provider; prohibited remote searches return no content.
123
+ Objects live under `ds4-context/artifacts/<sha-prefix>/<sha256>` with private permissions and atomic writes. With `storage.scope: "project"` the store root becomes `projects/artifacts/<project-digest>/...` so artifact bytes and their metadata share one boundary; the orphan garbage collector only sees references in the current database and must never delete another project's objects. Pi's full JSONL tool result remains canonical; the object file is a rebuildable local cache. No artifact content is stored in Context Manifests. Search recomputes SHA-256 and returns only bounded, redacted, JSON-quoted literal-match windows for a current-branch reference. The runtime reapplies the stored artifact classification before returning excerpts to the active provider; prohibited remote searches return no content.
111
124
 
112
125
  A full index rebuild replays all message entries, recreates missing qualifying objects, removes stale session references, and garbage-collects object rows/files with no references. Missing/corrupt states are reported by `/context health` without blocking Pi.
113
126
 
@@ -117,7 +130,7 @@ For persisted sessions, each `context` hook stores a metadata-only manifest cont
117
130
 
118
131
  `before_provider_request` updates the pending in-memory manifest with final-check/redaction counters but never the provider payload. The following finalized assistant response updates only the existing scalar usage columns (`actual_tokens`, `input_tokens`, `cache_read_tokens`, and `cache_write_tokens`) and adds at most one exact-model calibration sample. It does not read or rewrite `manifest_json`. Repository reads hydrate authoritative usage from those columns. Ephemeral, oversize-skipped, concurrently pruned, and otherwise uncorrelated manifests retain bounded calibration only in memory.
119
132
 
120
- Retention is bounded without a schema change: SQLite keeps the latest 128 manifests globally and at most 200 calibration samples for each provider/model/estimator profile. A manifest prune first detaches its small calibration row, then removes the large diagnostic JSON; calibration has its own per-profile retention. Save and prune are one transaction. Existing oversized stores are reduced incrementally by at most 32 rows and 8 MiB of serialized manifest payload per subsequent manifest write; one individually oversized oldest row may be removed to guarantee progress. There is no startup purge.
133
+ Retention is bounded without a schema change: SQLite keeps the latest 128 manifests globally and at most 200 calibration samples for each provider/model/estimator profile. A manifest prune first detaches its small calibration row, then removes the large diagnostic JSON; calibration has its own per-profile retention. Save and prune are one transaction. In `project` scope the manifest transaction runs in the project database while the calibration sample is inserted separately into the agent database. Existing oversized stores are reduced incrementally by at most 32 rows and 8 MiB of serialized manifest payload per subsequent manifest write; one individually oversized oldest row may be removed to guarantee progress. There is no startup purge.
121
134
 
122
135
  New manifest persistence is byte-bounded. Payloads up to 256 KiB remain complete. Larger payloads preserve all `included` provenance and replace only the `excluded` inventory with a deterministic first/last sample of at most 256 details plus explicit `ds4-context-manifest-inventory-v1` counts, token/classification/kind rollups, and digests. The wrapper returned by `getStored()` declares `complete` or `excluded-rollup`; the live runtime manifest remains complete. A projected payload over 1 MiB is skipped without affecting the model request. Deleted pages become reusable by SQLite but do not promise an immediate reduction in filesystem size. Manifests and calibration remain disposable; Pi JSONL and project files are untouched.
123
136
 
@@ -152,3 +165,25 @@ A full rebuild does not blindly delete unchanged entries. It upserts all observe
152
165
  Session reconciliation is transactional. Memory/pin mutation replacement, checkpoint update, source exclusion and full materialization each occur under the shared write coordinator. Each manifest upsert and dual-bound incremental retention prune share one transaction; each scalar usage/calibration update and its independent per-profile prune do the same. Each quality upsert and bounded-retention prune also share one transaction; quality failures do not affect manifests or planning. Each changed project file is replaced transactionally with its snippets and FTS rows; embedding upserts and canonical-source pruning are transactional; artifact object/reference metadata and project deletion batches are atomic. A filesystem artifact write precedes its metadata transaction, so an interrupted metadata write may leave only an unreferenced content-addressed cache file; canonical JSONL remains sufficient for recovery. If manifest serialization, projection, retention, or SQLite writing fails, the complete current manifest remains in memory and the provider request is unchanged. Other artifact/project failures contribute no replacement/snippets; planner failures discard all synthetic evidence; Pi continues with its native context.
153
166
 
154
167
  After bounded busy-aware replay is exhausted, DS4 emits `database.write_lock_timeout` with only the coordinator operation name, attempt count, elapsed/configured waits, and SQLite primary code. The thrown error repeats the operation and categorical lock status but never includes SQL, bound values, provider content, or the raw SQLite message. Retry and rollback diagnostics follow the same metadata-only rule.
168
+
169
+ ## Growth measurements
170
+
171
+ `tests/benchmarks/storage-scale.bench.ts` seeds session indexes of 5,000 / 50,000 / 200,000 entries and measures the paths a session actually pays. Run it on demand:
172
+
173
+ ```bash
174
+ npx vitest bench tests/benchmarks/storage-scale.bench.ts
175
+ ```
176
+
177
+ Measured means on this development machine (Node.js 26.5.1, `node:sqlite`, one session per database):
178
+
179
+ | Path | 5k entries | 50k entries | 200k entries |
180
+ | --- | ---: | ---: | ---: |
181
+ | Exact identifier scan (`instr` over one session) | 0.97 ms | 12.5 ms | 52.7 ms |
182
+ | Exact phrase scan | 0.98 ms | 12.8 ms | 53.2 ms |
183
+ | FTS retrieval, common token | 0.14 ms | 2.3 ms | 9.0 ms |
184
+ | FTS retrieval, rare token | 0.06 ms | 0.15 ms | 0.94 ms |
185
+ | Per-session stats (`COUNT`/`SUM` for one session) | 0.33 ms | 4.8 ms | 21.8 ms |
186
+ | Storage diagnostics | 0.10 ms | 0.10 ms | 0.10 ms |
187
+ | Append-only unchanged re-check (1,000 entries) | 1.4 ms | 2.9 ms | 7.2 ms |
188
+
189
+ Growth is **real but bounded**: exact identifier and phrase scans are literal `instr()` scans over the session's rows and grow roughly linearly, passing the 50 ms typical-operation target only around 200k indexed entries in a single session. FTS retrieval and per-session aggregate SQL grow sublinearly and stay in single-digit milliseconds; bounded manifest/calibration diagnostics are flat. The dominant cost tracks a single session's size, not the file size, so `storage.scope: "project"` bounds physical growth, lock scope and reset per project but does not by itself change this per-session scan profile. Long single sessions near or above the 200k-entry range are the case where exact-identifier retrieval latency becomes measurable.
@@ -1,6 +1,6 @@
1
1
  # Offline SQLite Storage Maintenance
2
2
 
3
- DS4 keeps `context.db` as disposable derived state, while Pi session JSONL and live project files remain canonical. Normal runtime retention stops unbounded manifest growth and makes deleted pages reusable. It does not promise that an existing high-water SQLite file shrinks physically.
3
+ DS4 keeps `context.db` as disposable derived state, while Pi session JSONL and live project files remain canonical. Normal runtime retention stops unbounded manifest growth and makes deleted pages reusable. It does not promise that an existing high-water SQLite file shrinks physically. With the default `storage.scope: "project"` each trusted project has its own database under `ds4-context/projects/` and the agent database keeps shared token calibration: run the maintenance commands per file, not only on the agent database (see [`STORAGE.md`](STORAGE.md)).
4
4
 
5
5
  Physical compaction is therefore an explicit offline operation. It is never model-callable, never runs at startup, and never edits Pi JSONL or project files.
6
6
 
@@ -0,0 +1,29 @@
1
+ # Release 0.3.10 — Opt-in BPE estimation and bounded model budget auto-tuning
2
+
3
+ **Coordinated packages:** `ds4-context-core`, `ds4-context-reference-adapter`, and `ds4-context-engine` 0.3.10.
4
+ **Implementation commit:** `0607238`.
5
+
6
+ ## Summary
7
+
8
+ Adds a selectable `o200k-base-v1` text estimator, provider/model/estimator-specific calibration, and opt-in evidence-gated expansion of automatic context-category ceilings. `chars-v1` remains the default estimator; `modelAwareness.autoTune` remains off unless explicitly enabled. The reference adapter and Pi extension continue to depend exactly on the matching core version.
9
+
10
+ ## Changes
11
+
12
+ - Portable core accepts a BPE estimator through its existing runtime-neutral interface; the Pi adapter lazy-loads `js-tiktoken`. Non-text content still uses bounded heuristics.
13
+ - Model profiles can select `tokenEstimator` by exact `provider/model`, provider wildcard, or global override. Calibration histories are isolated by estimator version as well as provider and model.
14
+ - Auto-tuning uses accepted same-profile samples, rejects outliers, respects configured category limits and hard input ceilings, and withdraws expansion after insufficient headroom. Calibrated estimator-unit limits cannot exceed nominal model limits.
15
+ - Manifests and model diagnostics identify estimator version and calibration/auto-tuning decisions. Existing golden compatibility remains anchored to the unchanged default behavior.
16
+
17
+ ## Measured scope and limitations
18
+
19
+ Bounded synthetic multi-turn Pi/DS4 sessions on OpenRouter `openai/gpt-6-sol`, `openai/gpt-6-luna`, and `openai/gpt-5.6-terra` demonstrated calibration, a genuinely larger opt-in recent tail, provider usage above the native-window headroom threshold, and withdrawal to `no-headroom` on the following turn in each model's session. Retrieval and live tool cycles were also exercised separately. See [model-awareness measurements](../MODEL_AWARENESS.md) and [provider-token drift](../PROVIDER_TOKEN_DRIFT_BENCHMARK.md) for methodology, numbers, and bounds.
20
+
21
+ These measurements do not establish production retrieval quality, provider-side compaction behavior, or accuracy for arbitrary models/routes. In particular, equivalent Codex-route measurements did not yield usage and are not claimed as validated. No provider calls are required for this release procedure.
22
+
23
+ ## Compatibility
24
+
25
+ No new defaults are enabled: `chars-v1`, disabled `autoTune`, and disabled DS4 native continuation remain unchanged. No SQLite migration or change to canonical Pi JSONL, privacy consent, native continuation, or the portable runtime-adapter contract is introduced. BPE remains adapter-injected rather than adding a third-party tokenizer dependency to portable core. Restart Pi after upgrading so the new compiled core and session configuration are loaded.
26
+
27
+ ## Validation and publication
28
+
29
+ On Node 26.5.1, the coordinated release passed a clean `npm ci`, `npm run check` (97 Vitest files, 607 tests, TypeScript builds and root typecheck), deterministic `npm run quality:compare`, the `npm run schema:context-persistence` size bound, and `npm run pack:check` in a clean consumer after synchronizing both exported runtime version constants. All typecheck/tests were also rerun after that correction. Dry-run tarball review contained 243 core files, 7 reference-adapter files, and 97 extension files; none included untracked local state. All three packages were then published to npm at 0.3.10 in dependency order. After registry propagation, `npm run registry:check -- 0.3.10` verified all three exact-version artifacts in a fresh consumer.
@@ -1,7 +1,7 @@
1
1
  # Release 0.3.9 — Bounded compaction request and operation input
2
2
 
3
3
  **Version analyzed:** DS4 Context Engine `0.3.9`
4
- **Commit:** (recorded after pack and registry verification)
4
+ **Commit:** `01407cd`
5
5
  **Coordinated packages:** `ds4-context-core` 0.3.9, `ds4-context-reference-adapter` 0.3.9, `ds4-context-engine` 0.3.9
6
6
 
7
7
  ## Summary
@@ -0,0 +1,86 @@
1
+ # Release 0.4.0 — Per-project databases with shared token calibration
2
+
3
+ **Coordinated packages:** `ds4-context-core`, `ds4-context-reference-adapter`, and `ds4-context-engine` 0.4.0.
4
+ **Implementation commit:** `2c2e248`.
5
+ **Decision record:** [ADR 064](../ADR/064-per-project-databases-with-shared-calibration.md).
6
+
7
+ ## Summary
8
+
9
+ Adds the opt-out `storage.scope: "agent" | "project"` setting and defaults it to
10
+ `project`. Each trusted canonical project root now gets its own SQLite
11
+ projection next to the configured agent database, while token calibration stays
12
+ in the agent database and remains shared across projects. `storage.scope:
13
+ "agent"` restores the previous single-file layout. The reference adapter and Pi
14
+ extension continue to depend exactly on the matching core version.
15
+
16
+ ## Changes
17
+
18
+ - `storage.scope: "project"` (new default) derives one database per trusted
19
+ canonical project root at `projects/<sha256(root)[0..32]>.db`, next to
20
+ `storage.databasePath`. Untrusted projects, broad roots (home directory,
21
+ filesystem root) and any resolution failure fall back to the agent database.
22
+ - Both files receive the same schema and the same migrations; migrations 1–15
23
+ are untouched and there is no schema fork. The project database holds the
24
+ session index, project index, context manifests, summary graph, memory/pin
25
+ projections, embeddings, quality samples and artifact metadata; the agent
26
+ database holds `token_calibration`.
27
+ - With the split, a calibration sample is written to the agent database with
28
+ `manifest_id = NULL`, exactly once per manifest. The two writes are
29
+ intentionally not one cross-database transaction: a lost calibration sample is
30
+ harmless, a lost manifest/usage consistency is not. In `agent` scope the
31
+ previous single transaction is unchanged.
32
+ - Artifact object bytes move to `projects/artifacts/<project-digest>/` under
33
+ project scope so the orphan garbage collector, which only sees references in
34
+ the current database, can never delete another project's objects.
35
+ - `/context storage` and `/context diagnostics` report the active project
36
+ database and, when split, the shared agent database. Storage maintenance
37
+ stays per file: `ds4-context-storage inspect|compact|recover --database <path>`
38
+ must be pointed at each database.
39
+ - The previous shared database is left untouched. With the new default, existing
40
+ users cold-start per-project indexes while existing calibration remains
41
+ available in the agent database; project indexes rebuild from canonical Pi
42
+ JSONL, project files and memory/pin `CustomEntry` records.
43
+
44
+ ## Breaking behavior
45
+
46
+ The default storage layout changes. A pre-existing
47
+ `~/.pi/agent/ds4-context/context.db` becomes the agent database and is no longer
48
+ the active projection for new sessions. `storage.scope: "agent"` restores the
49
+ old single-database behavior. Rows are not migrated between files; the project
50
+ databases are derived and rebuildable. Decision D1 of the storage plan ("no
51
+ per-session or per-project databases") is superseded by ADR 064 and the plan
52
+ keeps the original text with an explicit amendment pointer.
53
+
54
+ ## Measured scope and limitations
55
+
56
+ `tests/benchmarks/storage-scale.bench.ts` seeds one session with 5,000 /
57
+ 50,000 / 200,000 entries and measures the paths a session pays. Measured means
58
+ on the development machine (Node.js 26.5.1, `node:sqlite`): exact identifier
59
+ scan 0.97 / 12.5 / 52.7 ms, exact phrase scan 0.98 / 12.8 / 53.2 ms, FTS
60
+ common token 0.14 / 2.3 / 9.0 ms, per-session stats 0.33 / 4.8 / 21.8 ms,
61
+ storage diagnostics flat at 0.10 ms, unchanged re-check 1.4 / 2.9 / 7.2 ms.
62
+ Growth is real but bounded and tracks single-session size, not file size:
63
+ project scope bounds physical growth, lock scope and reset per project, but does
64
+ not by itself change the per-session exact-scan profile. See
65
+ [storage growth measurements](../STORAGE.md#growth-measurements).
66
+
67
+ These are local measurements on one host, not portable guarantees. No provider
68
+ calls are involved in this release procedure.
69
+
70
+ ## Compatibility
71
+
72
+ No provider-facing default changes: `chars-v1` remains the default estimator,
73
+ `modelAwareness.autoTune` and DS4 native continuation stay disabled. No SQLite
74
+ migration, no change to canonical Pi JSONL, privacy consent, native continuation
75
+ or the portable runtime-adapter contract. Restart Pi after upgrading so the new
76
+ compiled core and session configuration are loaded.
77
+
78
+ ## Validation and publication
79
+
80
+ On Node 26.5.1, the coordinated release passed `npm run check` (99 Vitest files,
81
+ 615 tests, TypeScript builds and root typecheck), deterministic
82
+ `npm run quality:compare`, the `npm run schema:context-persistence` size bound,
83
+ and `npm run pack:check` in a clean consumer after synchronizing both exported
84
+ runtime version constants. `git diff --check` was clean and dry-run tarball
85
+ review contained 247 core files, 7 reference-adapter files and 98 extension
86
+ files, none including untracked local state.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ds4-context-engine",
3
- "version": "0.3.9",
3
+ "version": "0.4.0",
4
4
  "description": "Non-destructive, provider-independent context management for Pi.",
5
5
  "type": "module",
6
6
  "license": "MIT",
@@ -52,6 +52,7 @@
52
52
  "quality:compare": "node scripts/compare-context-quality.mjs",
53
53
  "schema:context-persistence": "node scripts/measure-context-persistence-schema.mjs",
54
54
  "latency:check": "npm run build:core && node scripts/compare-disabled-planning-latency.mjs",
55
+ "storage:scale": "npm run build:core && vitest bench tests/benchmarks/storage-scale.bench.ts",
55
56
  "pack:check": "node scripts/verify-packages.mjs",
56
57
  "registry:check": "node scripts/verify-registry-packages.mjs",
57
58
  "prepare": "npm run build:core && npm run build:adapters"
@@ -62,7 +63,8 @@
62
63
  ]
63
64
  },
64
65
  "dependencies": {
65
- "ds4-context-core": "0.3.9"
66
+ "ds4-context-core": "0.4.0",
67
+ "js-tiktoken": "1.0.21"
66
68
  },
67
69
  "peerDependencies": {
68
70
  "@earendil-works/pi-ai": "0.84.3",
@@ -250,6 +250,9 @@ function formatStatus(diagnostics: RuntimeDiagnostics): string {
250
250
  `Privacy blocked/redacted: ${count(diagnostics.privacy.blockedBlocks)} / ${count(diagnostics.privacy.secretRedactions)}`,
251
251
  `Estimator calibration: ${diagnostics.modelAwareness?.calibration.calibrated ? `x${diagnostics.modelAwareness.calibration.appliedRatio.toFixed(3)}` : "collecting/neutral"}`,
252
252
  `Calibration samples: ${count(diagnostics.modelAwareness?.calibration.acceptedSamples)} accepted`,
253
+ ...(diagnostics.modelAwareness?.drift ? [
254
+ `Token estimate warning: ${diagnostics.modelAwareness.drift.code} (${diagnostics.modelAwareness.drift.severity}, ${count(diagnostics.modelAwareness.drift.sampleCount)} samples)`,
255
+ ] : []),
253
256
  `Native continuation: ${diagnostics.nativeContinuation.status} (${diagnostics.nativeContinuation.last?.mode ?? "no request"})`,
254
257
  `Continuation saved items: ${count(diagnostics.nativeContinuation.last?.omittedInputItems)}`,
255
258
  `Quality metrics: ${diagnostics.quality.enabled ? `${count(diagnostics.quality.storedSamples)} sample(s)` : "disabled"}`,
@@ -365,6 +368,9 @@ function formatManifest(diagnostics: RuntimeDiagnostics): string {
365
368
  `Artifacts: ${count(manifest.artifacts?.length ?? 0)}`,
366
369
  `Privacy: ${manifest.privacy?.enforcement ?? "disabled"}${manifest.privacy ? ` (${manifest.privacy.destination})` : ""}`,
367
370
  `Model calibration: ${manifest.modelAwareness?.calibration.calibrated ? `x${manifest.modelAwareness.calibration.appliedRatio.toFixed(3)}` : "neutral/collecting"}`,
371
+ ...(manifest.modelAwareness?.drift ? [
372
+ `Token estimate warning: ${manifest.modelAwareness.drift.code} (${manifest.modelAwareness.drift.severity})`,
373
+ ] : []),
368
374
  `Adaptive tail/hist/project: ${count(manifest.modelAwareness?.adaptive.recentTailTokens)} / ${count(manifest.modelAwareness?.adaptive.maxRetrievedHistoryTokens)} / ${count(manifest.modelAwareness?.adaptive.maxProjectTokens)}`,
369
375
  `Continuation: ${manifest.nativeContinuation?.mode ?? "disabled"}; sent/full ${count(manifest.nativeContinuation?.sentInputItems)} / ${count(manifest.nativeContinuation?.fullInputItems)}`,
370
376
  "",
@@ -664,6 +670,12 @@ function formatModelAwareness(diagnostics: RuntimeDiagnostics): string {
664
670
  `Samples observed/accepted: ${count(calibration.observedSamples)} / ${count(calibration.acceptedSamples)}`,
665
671
  `Samples rejected/outliers: ${count(calibration.rejectedSamples)} / ${count(calibration.outlierSamples)}`,
666
672
  `Calibration bounds/window: ${calibration.lowerRatioBound.toFixed(2)}-${calibration.upperRatioBound.toFixed(2)} / ${count(calibration.windowSize)}`,
673
+ ...(awareness.drift ? [
674
+ `Token estimate warning: ${awareness.drift.code} (${awareness.drift.severity}, ${count(awareness.drift.sampleCount)} samples${awareness.drift.medianRatio === undefined ? "" : `, x${awareness.drift.medianRatio.toFixed(3)}`})`,
675
+ ] : []),
676
+ ...(awareness.autoTune ? [
677
+ `Budget auto-tuning: ${awareness.autoTune.status} (${count(awareness.autoTune.acceptedSamples)} samples, x${awareness.autoTune.boostFactor.toFixed(3)})`,
678
+ ] : []),
667
679
  `Adaptive recent tail: ${count(awareness.adaptive.recentTailTokens)} (nominal ${count(awareness.adaptive.nominalRecentTailTokens)})`,
668
680
  `Adaptive history retrieval: ${count(awareness.adaptive.maxRetrievedHistoryTokens)} (nominal ${count(awareness.adaptive.nominalRetrievedHistoryTokens)})`,
669
681
  `Adaptive project retrieval: ${count(awareness.adaptive.maxProjectTokens)} (nominal ${count(awareness.adaptive.nominalProjectTokens)})`,
@@ -823,13 +835,20 @@ function formatAdapter(diagnostics: RuntimeDiagnostics): string {
823
835
  ].join("\n");
824
836
  }
825
837
 
826
- function formatStorage(storage: StorageDiagnostics, databasePath?: string): string {
838
+ function formatStorage(
839
+ storage: StorageDiagnostics,
840
+ databasePath?: string,
841
+ agentDatabasePath?: string,
842
+ ): string {
843
+ const agentLine = agentDatabasePath && agentDatabasePath !== databasePath
844
+ ? [`Agent database (shared): ${agentDatabasePath}`] : [];
827
845
  if (storage.status === "unavailable") {
828
846
  return [
829
847
  "DS4 Storage",
830
848
  "",
831
849
  "Status: unavailable",
832
850
  `Database: ${databasePath ?? "unavailable"}`,
851
+ ...agentLine,
833
852
  "Pi fallback remains active; no storage mutation was attempted.",
834
853
  ].join("\n");
835
854
  }
@@ -838,6 +857,7 @@ function formatStorage(storage: StorageDiagnostics, databasePath?: string): stri
838
857
  "",
839
858
  `Status: ${storage.status}`,
840
859
  `Database: ${databasePath ?? "unavailable"}`,
860
+ ...agentLine,
841
861
  `Schema / journal: ${storage.schemaVersion ?? "n/a"} / ${storage.journalMode ?? "n/a"}`,
842
862
  `Database / WAL / SHM: ${bytes(storage.databaseBytes)} / ${bytes(storage.walBytes)} / ${bytes(storage.shmBytes)}`,
843
863
  `Allocated / reusable: ${bytes(storage.allocatedBytes)} / ${bytes(storage.reusableBytes)}`,
@@ -1222,7 +1242,7 @@ export function registerContextCommand(pi: ExtensionAPI, runtime: Ds4ContextRunt
1222
1242
  const storage = runtime.storageDiagnostics();
1223
1243
  present(
1224
1244
  ctx,
1225
- formatStorage(storage, diagnostics.databasePath),
1245
+ formatStorage(storage, diagnostics.databasePath, diagnostics.agentDatabasePath),
1226
1246
  storage.status === "ok" ? "info" : "warning",
1227
1247
  );
1228
1248
  return;
@@ -1,7 +1,7 @@
1
1
  import { randomUUID } from "node:crypto";
2
2
  import { existsSync, mkdirSync, readFileSync, realpathSync, writeFileSync } from "node:fs";
3
3
  import { homedir } from "node:os";
4
- import { dirname, join, parse, resolve } from "node:path";
4
+ import { basename, dirname, join, parse, resolve } from "node:path";
5
5
  import type {
6
6
  Api,
7
7
  AssistantMessage,
@@ -46,11 +46,12 @@ export type {
46
46
  import {
47
47
  loadConfig,
48
48
  resolveDatabasePath,
49
+ resolveProjectDatabasePath as deriveProjectDatabasePath,
49
50
  resolveRankingModelPath,
50
51
  validateConfigFile,
51
52
  type LoadedConfig,
52
53
  } from "ds4-context-core/config/config-loader";
53
- import { CONFIG_SCHEMA_VERSION, createDefaultConfig, type Ds4ContextConfig } from "ds4-context-core/config/config";
54
+ import { CONFIG_SCHEMA_VERSION, createDefaultConfig, type Ds4ContextConfig, type StorageScope } from "ds4-context-core/config/config";
54
55
  import {
55
56
  applyConfigValue,
56
57
  findConfigField,
@@ -60,11 +61,12 @@ import { calculateContextBudget, type ContextBudget } from "ds4-context-core/cor
60
61
  import {
61
62
  modelProfileKey,
62
63
  resolveModelAwareness,
64
+ tokenEstimatorVersion,
63
65
  type ResolvedModelAwareness,
64
66
  type TokenCalibrationSample,
65
67
  } from "ds4-context-core/core/model-awareness";
66
68
  import type { ModelDescriptor } from "ds4-context-core/core/model-profile";
67
- import { estimateMessagesTokens, estimateTextTokens } from "ds4-context-core/core/token-estimator";
69
+ import { CHARS_ESTIMATOR, type TokenEstimator } from "ds4-context-core/core/token-estimator";
68
70
  import type {
69
71
  ContextManifest,
70
72
  ModelAwarenessManifest,
@@ -164,6 +166,7 @@ import {
164
166
  findPiSourceEntryIds,
165
167
  fingerprint,
166
168
  } from "../pi-adapter/context-observer.ts";
169
+ import { O200K_ESTIMATOR } from "../pi-adapter/bpe-token-estimator.ts";
167
170
  import { projectSessionFileMutations } from "../pi-adapter/memory-adapter.ts";
168
171
  import { ProjectMemorySynchronizer } from "../pi-adapter/project-memory-sync.ts";
169
172
  import { projectRankingLabels } from "../pi-adapter/ranking-adapter.ts";
@@ -385,6 +388,8 @@ export interface RuntimeDiagnostics {
385
388
  session?: PiSessionSnapshot;
386
389
  model?: { provider: string; id: string };
387
390
  databasePath?: string;
391
+ /** Present only when storage.scope splits calibration from project state. */
392
+ agentDatabasePath?: string;
388
393
  databaseSchemaVersion?: number;
389
394
  indexed?: SessionIndexStats;
390
395
  observation?: ContextObservation;
@@ -424,8 +429,10 @@ export class Ds4ContextRuntime {
424
429
  private loadedConfig?: LoadedConfig;
425
430
  private session?: PiSessionSnapshot;
426
431
  private database?: ContextDatabase;
432
+ private agentDatabase?: ContextDatabase;
427
433
  private indexer?: PiSessionIndexer;
428
434
  private databasePath?: string;
435
+ private agentDatabasePath?: string;
429
436
  private observation?: ContextObservation;
430
437
  private lastManifest?: ContextManifest;
431
438
  private lastPersistedInventory?: PersistedManifestInventory;
@@ -577,17 +584,33 @@ export class Ds4ContextRuntime {
577
584
  return;
578
585
  }
579
586
 
580
- this.databasePath = resolveDatabasePath(
587
+ const agentDatabasePath = resolveDatabasePath(
581
588
  this.config.storage.databasePath,
582
589
  this.dependencies.agentDir,
583
590
  this.dependencies.homeDir,
584
591
  );
592
+ this.agentDatabasePath = agentDatabasePath;
593
+ const projectDatabasePath = this.resolveProjectDatabasePath(
594
+ this.config.storage.scope,
595
+ ctx,
596
+ agentDatabasePath,
597
+ );
598
+ this.databasePath = projectDatabasePath ?? agentDatabasePath;
599
+ if (projectDatabasePath) {
600
+ this.agentDatabase = ContextDatabase.open(agentDatabasePath, {
601
+ logger: this.logger,
602
+ now: this.now(),
603
+ busyTimeoutMs: this.config.storage.busyTimeoutMs,
604
+ writeRetryTimeoutMs: this.config.storage.writeRetryTimeoutMs,
605
+ });
606
+ }
585
607
  this.database = ContextDatabase.open(this.databasePath, {
586
608
  logger: this.logger,
587
609
  now: this.now(),
588
610
  busyTimeoutMs: this.config.storage.busyTimeoutMs,
589
611
  writeRetryTimeoutMs: this.config.storage.writeRetryTimeoutMs,
590
612
  });
613
+ this.agentDatabase ??= this.database;
591
614
  this.initializeRanking();
592
615
 
593
616
  this.indexer = new PiSessionIndexer(this.database.sessionIndex, {
@@ -758,15 +781,37 @@ export class Ds4ContextRuntime {
758
781
  };
759
782
  }
760
783
 
784
+ private estimatorForModel(model: ModelDescriptor): TokenEstimator {
785
+ return tokenEstimatorVersion(model, this.config.modelAwareness) === "o200k-base-v1"
786
+ ? O200K_ESTIMATOR : CHARS_ESTIMATOR;
787
+ }
788
+
789
+ private resolveProjectDatabasePath(
790
+ scope: StorageScope,
791
+ ctx: ExtensionContext,
792
+ agentDatabasePath: string,
793
+ ): string | undefined {
794
+ if (scope !== "project") return undefined;
795
+ // Untrusted projects ignore project configuration; they also fall back to
796
+ // the shared agent database instead of creating a new physical boundary.
797
+ if (!ctx.isProjectTrusted()) return undefined;
798
+ const projectRoot = canonicalProjectPath(ctx.cwd);
799
+ if (isBroadProjectRoot(projectRoot, this.dependencies.homeDir ?? homedir())) return undefined;
800
+ const projectDatabasePath = deriveProjectDatabasePath(agentDatabasePath, projectRoot);
801
+ return projectDatabasePath === agentDatabasePath ? undefined : projectDatabasePath;
802
+ }
803
+
761
804
  private calibrationSamples(model: ModelDescriptor): TokenCalibrationSample[] {
805
+ const version = this.estimatorForModel(model).version;
762
806
  if (this.database && this.session?.sessionFile && this.config.diagnostics.storeContextManifest) {
763
- return this.database.manifests.listCalibrationSamples(
807
+ return (this.agentDatabase ?? this.database).calibrations.list(
764
808
  model.provider,
765
809
  model.id,
766
810
  this.config.modelAwareness.calibrationWindow,
811
+ version,
767
812
  );
768
813
  }
769
- return [...(this.volatileCalibration.get(modelProfileKey(model.provider, model.id)) ?? [])];
814
+ return [...(this.volatileCalibration.get(`${modelProfileKey(model.provider, model.id)}\0${version}`) ?? [])];
770
815
  }
771
816
 
772
817
  private switchForModel(model: ModelDescriptor): ModelSwitchManifest {
@@ -831,6 +876,7 @@ export class Ds4ContextRuntime {
831
876
  requestsPerTurn: number,
832
877
  turnsPerEpoch: number,
833
878
  modelKey: string,
879
+ estimator: TokenEstimator,
834
880
  ): number | undefined {
835
881
  const pricing = Ds4ContextRuntime.cachePricing(cost);
836
882
  if (!pricing || plan.mode !== "managed") return undefined;
@@ -839,7 +885,7 @@ export class Ds4ContextRuntime {
839
885
  let total = fixedTokens;
840
886
  for (const message of plan.messages) {
841
887
  hashes.push(fingerprint(message));
842
- const estimate = estimateMessagesTokens([message]);
888
+ const estimate = estimator.estimateMessagesTokens([message]);
843
889
  tokens.push(estimate);
844
890
  total += estimate;
845
891
  }
@@ -908,6 +954,8 @@ export class Ds4ContextRuntime {
908
954
  ...awareness.calibration,
909
955
  cache: { ...awareness.calibration.cache },
910
956
  },
957
+ ...(awareness.drift ? { drift: awareness.drift } : {}),
958
+ ...(awareness.autoTune ? { autoTune: awareness.autoTune } : {}),
911
959
  adaptive: { ...awareness.limits },
912
960
  switch: modelSwitch,
913
961
  };
@@ -922,7 +970,7 @@ export class Ds4ContextRuntime {
922
970
  usage: ProviderUsageManifest,
923
971
  createdAt: number,
924
972
  ): void {
925
- const key = modelProfileKey(manifest.provider, manifest.model);
973
+ const key = `${modelProfileKey(manifest.provider, manifest.model)}\0${manifest.modelAwareness?.calibration.estimator ?? "chars-v1"}`;
926
974
  const samples = this.volatileCalibration.get(key) ?? [];
927
975
  samples.unshift({
928
976
  estimatedTokens: manifest.estimatedInputTokens,
@@ -980,6 +1028,7 @@ export class Ds4ContextRuntime {
980
1028
  }
981
1029
  const model = snapshotModel(ctx);
982
1030
  const activeModel = model ? this.resolveActiveModel(model) : undefined;
1031
+ const estimator = model ? this.estimatorForModel(model) : CHARS_ESTIMATOR;
983
1032
  const budget = activeModel?.budget;
984
1033
  let effectiveEvent = preparedPrivacy.event;
985
1034
  let artifactReferences = [] as NonNullable<ContextManifest["artifacts"]>;
@@ -993,8 +1042,8 @@ export class Ds4ContextRuntime {
993
1042
  preparedPrivacy.messageClassifications,
994
1043
  this.config.artifacts.adaptiveBudget && budget ? {
995
1044
  inputTokens: budget.activeInputBudget,
996
- fixedTokens: estimateTextTokens(preparedPrivacy.systemPrompt) + 8
997
- + preparedPrivacy.tools.reduce((sum, tool) => sum + estimateObservedToolTokens(tool), 0),
1045
+ fixedTokens: estimator.estimateTextTokens(preparedPrivacy.systemPrompt) + 8
1046
+ + preparedPrivacy.tools.reduce((sum, tool) => sum + estimateObservedToolTokens(tool, estimator), 0),
998
1047
  } : undefined,
999
1048
  );
1000
1049
  effectiveEvent = { type: "context", messages: transformed.messages };
@@ -1043,6 +1092,7 @@ export class Ds4ContextRuntime {
1043
1092
  createdAt: observedAt,
1044
1093
  policyVersion: POLICY_VERSION,
1045
1094
  plannerVersion: OBSERVER_PLANNER_VERSION,
1095
+ tokenEstimator: estimator,
1046
1096
  ...(activeModel ? {
1047
1097
  profile: activeModel.awareness.profile,
1048
1098
  budget: activeModel.budget,
@@ -1107,6 +1157,7 @@ export class Ds4ContextRuntime {
1107
1157
  fixedTokens,
1108
1158
  budget,
1109
1159
  config: effectiveContextConfig,
1160
+ tokenEstimator: estimator,
1110
1161
  pinnedMessageIndices,
1111
1162
  supplementalMessages: dedupSupplementalMessages,
1112
1163
  });
@@ -1138,13 +1189,14 @@ export class Ds4ContextRuntime {
1138
1189
  fixedTokens,
1139
1190
  budget,
1140
1191
  config: effectiveContextConfig,
1192
+ tokenEstimator: estimator,
1141
1193
  pinnedMessageIndices,
1142
1194
  supplementalMessages: dedupSupplementalMessages,
1143
1195
  cacheAwareTailTokens: decision.recentTailTokens,
1144
1196
  });
1145
1197
  const modelKey = modelProfileKey(model.provider, model.id);
1146
- const nominalEpoch = this.cacheAwareEpochCost(nominalDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey);
1147
- const extendedEpoch = this.cacheAwareEpochCost(extendedDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey);
1198
+ const nominalEpoch = this.cacheAwareEpochCost(nominalDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey, estimator);
1199
+ const extendedEpoch = this.cacheAwareEpochCost(extendedDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey, estimator);
1148
1200
  this.logger.debug("context.cache_aware_candidate", {
1149
1201
  eligible: decision.eligible,
1150
1202
  tailExtended: decision.tailExtended,
@@ -1181,6 +1233,7 @@ export class Ds4ContextRuntime {
1181
1233
  fixedTokens,
1182
1234
  budget,
1183
1235
  config: effectiveContextConfig,
1236
+ tokenEstimator: estimator,
1184
1237
  pinnedMessageIndices,
1185
1238
  supplementalMessages: dedupSupplementalMessages,
1186
1239
  cacheAwareTailTokens,
@@ -1275,7 +1328,7 @@ export class Ds4ContextRuntime {
1275
1328
  if (sanitized.blockedBlocks > 0) {
1276
1329
  privacyExcludedSources.push({
1277
1330
  sourceId: supplement.sourceIds[0],
1278
- tokens: estimateMessagesTokens([supplement.message]),
1331
+ tokens: estimator.estimateMessagesTokens([supplement.message]),
1279
1332
  kind: supplement.kind,
1280
1333
  classification: sanitized.classification,
1281
1334
  score: supplement.score,
@@ -1328,6 +1381,7 @@ export class Ds4ContextRuntime {
1328
1381
  fixedTokens,
1329
1382
  budget,
1330
1383
  config: effectiveContextConfig,
1384
+ tokenEstimator: estimator,
1331
1385
  pinnedMessageIndices,
1332
1386
  supplementalMessages: rankedSupplementalMessages,
1333
1387
  ...(cacheAwareTailTokens !== undefined ? { cacheAwareTailTokens } : {}),
@@ -1408,7 +1462,7 @@ export class Ds4ContextRuntime {
1408
1462
  const currentTokens: number[] = [];
1409
1463
  for (const message of plan.messages) {
1410
1464
  currentHashes.push(fingerprint(message));
1411
- currentTokens.push(estimateMessagesTokens([message]));
1465
+ currentTokens.push(estimator.estimateMessagesTokens([message]));
1412
1466
  }
1413
1467
  const reusablePrefixTokens = estimateReusablePrefixTokens(
1414
1468
  previousHashes,
@@ -1495,6 +1549,7 @@ export class Ds4ContextRuntime {
1495
1549
  createdAt: observedAt,
1496
1550
  policyVersion: POLICY_VERSION,
1497
1551
  plannerVersion: PLANNER_VERSION,
1552
+ tokenEstimator: estimator,
1498
1553
  ...(activeModel ? {
1499
1554
  profile: activeModel.awareness.profile,
1500
1555
  budget: activeModel.budget,
@@ -1542,7 +1597,7 @@ export class Ds4ContextRuntime {
1542
1597
  estimatedMessageTokens: manifest.composition.messageTokens,
1543
1598
  originalMessageCount: manifest.planning?.originalMessageCount ?? historyEvent.messages.length,
1544
1599
  originalEstimatedMessageTokens: manifest.planning?.originalMessageTokens
1545
- ?? estimateMessagesTokens(historyEvent.messages),
1600
+ ?? estimator.estimateMessagesTokens(historyEvent.messages),
1546
1601
  ...(usage?.tokens !== null && usage?.tokens !== undefined ? { reportedTokens: usage.tokens } : {}),
1547
1602
  ...(manifest.planning?.durationMs !== undefined
1548
1603
  ? { planningDurationMs: manifest.planning.durationMs }
@@ -1865,13 +1920,39 @@ export class Ds4ContextRuntime {
1865
1920
  const createdAt = this.now();
1866
1921
  try {
1867
1922
  if (manifestPersisted && this.database) {
1868
- const updated = this.database.manifests.recordProviderUsage(
1923
+ const estimatorVersion = manifest?.modelAwareness?.calibration.estimator ?? "chars-v1";
1924
+ // With storage.scope=project the manifest lives in the project database
1925
+ // while calibration stays shared in the agent database. The two writes
1926
+ // are intentionally not one transaction: losing one sample is harmless,
1927
+ // losing manifest/usage consistency is not.
1928
+ const calibrationDatabase = this.agentDatabase !== this.database
1929
+ ? this.agentDatabase : undefined;
1930
+ const updated = this.database.manifests.recordProviderUsageOutcome(
1869
1931
  manifestId,
1870
1932
  providerUsage,
1871
1933
  createdAt,
1872
- manifest?.modelAwareness?.calibration.estimator ?? "chars-v1",
1934
+ estimatorVersion,
1935
+ { writeCalibration: calibrationDatabase === undefined },
1873
1936
  );
1874
- if (!updated && manifest?.estimatedInputTokens) {
1937
+ if (updated.outcome === "recorded" && calibrationDatabase) {
1938
+ const source = this.database.manifests.calibrationSource(manifestId);
1939
+ const recorded = source
1940
+ ? calibrationDatabase.calibrations.record({
1941
+ provider: source.provider,
1942
+ model: source.model,
1943
+ estimatedTokens: source.estimatedTokens,
1944
+ actualInputTokens: providerUsage.totalInputTokens,
1945
+ inputTokens: providerUsage.inputTokens,
1946
+ cacheReadTokens: providerUsage.cacheReadTokens,
1947
+ cacheWriteTokens: providerUsage.cacheWriteTokens,
1948
+ createdAt,
1949
+ estimatorVersion,
1950
+ })
1951
+ : false;
1952
+ if (!recorded && manifest?.estimatedInputTokens) {
1953
+ this.rememberVolatileCalibration(manifest, providerUsage, createdAt);
1954
+ }
1955
+ } else if (!updated.manifest && manifest?.estimatedInputTokens) {
1875
1956
  this.rememberVolatileCalibration(manifest, providerUsage, createdAt);
1876
1957
  }
1877
1958
  } else if (manifest?.estimatedInputTokens) {
@@ -2688,10 +2769,13 @@ export class Ds4ContextRuntime {
2688
2769
  return;
2689
2770
  }
2690
2771
  try {
2691
- const store = new FileArtifactStore(
2692
- join(this.dependencies.agentDir, "ds4-context", "artifacts"),
2693
- this.now,
2694
- );
2772
+ // Artifact bytes and their metadata must share one boundary: the orphan
2773
+ // garbage collector only sees references in the current database, so a
2774
+ // shared store would let one project delete another project's objects.
2775
+ const storeRoot = this.agentDatabase !== this.database && this.databasePath
2776
+ ? join(dirname(this.databasePath), "artifacts", basename(this.databasePath, ".db"))
2777
+ : join(this.dependencies.agentDir, "ds4-context", "artifacts");
2778
+ const store = new FileArtifactStore(storeRoot, this.now);
2695
2779
  this.artifactManager = new ArtifactManager(
2696
2780
  store,
2697
2781
  this.database.artifacts,
@@ -3276,6 +3360,8 @@ export class Ds4ContextRuntime {
3276
3360
  session: currentSession,
3277
3361
  ...(ctx.model ? { model: { provider: ctx.model.provider, id: ctx.model.id } } : {}),
3278
3362
  ...(this.databasePath ? { databasePath: this.databasePath } : {}),
3363
+ ...(this.agentDatabasePath && this.agentDatabasePath !== this.databasePath
3364
+ ? { agentDatabasePath: this.agentDatabasePath } : {}),
3279
3365
  ...(this.database ? { databaseSchemaVersion: this.database.schemaVersion } : {}),
3280
3366
  ...(indexed ? { indexed } : {}),
3281
3367
  ...(this.observation ? { observation: this.observation } : {}),
@@ -3313,7 +3399,9 @@ export class Ds4ContextRuntime {
3313
3399
  }
3314
3400
 
3315
3401
  storageDiagnostics(): StorageDiagnostics {
3316
- return this.database?.storageDiagnostics(this.session?.projectPath)
3402
+ const calibrationDatabase = this.agentDatabase && this.agentDatabase !== this.database
3403
+ ? this.agentDatabase : undefined;
3404
+ return this.database?.storageDiagnostics(this.session?.projectPath, calibrationDatabase)
3317
3405
  ?? unavailableStorageDiagnostics();
3318
3406
  }
3319
3407
 
@@ -3440,17 +3528,25 @@ export class Ds4ContextRuntime {
3440
3528
  try {
3441
3529
  this.database?.close();
3442
3530
  } finally {
3443
- this.compaction = undefined;
3444
- this.retrievalEngine = undefined;
3445
- this.projectKnowledge = undefined;
3446
- this.projectRefreshPending = false;
3447
- this.memoryManager = undefined;
3448
- this.projectMemorySynchronizer = undefined;
3449
- this.lastCrossSessionMemory = disabledCrossSessionMemoryDiagnostics();
3450
- this.lastMemoryMutationSignature = undefined;
3451
- this.artifactManager = undefined;
3452
- this.indexer = undefined;
3453
- this.database = undefined;
3531
+ try {
3532
+ if (this.agentDatabase && this.agentDatabase !== this.database) {
3533
+ this.agentDatabase.close();
3534
+ }
3535
+ } finally {
3536
+ this.agentDatabase = undefined;
3537
+ this.agentDatabasePath = undefined;
3538
+ this.compaction = undefined;
3539
+ this.retrievalEngine = undefined;
3540
+ this.projectKnowledge = undefined;
3541
+ this.projectRefreshPending = false;
3542
+ this.memoryManager = undefined;
3543
+ this.projectMemorySynchronizer = undefined;
3544
+ this.lastCrossSessionMemory = disabledCrossSessionMemoryDiagnostics();
3545
+ this.lastMemoryMutationSignature = undefined;
3546
+ this.artifactManager = undefined;
3547
+ this.indexer = undefined;
3548
+ this.database = undefined;
3549
+ }
3454
3550
  }
3455
3551
  }
3456
3552
 
@@ -0,0 +1,24 @@
1
+ import { createRequire } from "node:module";
2
+ import type { Tiktoken } from "js-tiktoken/lite";
3
+ import { createO200kEstimator } from "ds4-context-core/core/bpe-token-estimator";
4
+
5
+ const load = createRequire(import.meta.url);
6
+
7
+ let encoder: Tiktoken | undefined;
8
+
9
+ /**
10
+ * Opt-in OpenAI o200k text estimator. Model-specific serializers, image tokens,
11
+ * reasoning, and provider-specific wrappers are still estimated, not exact.
12
+ * No remote vocabulary download or provider request is performed.
13
+ */
14
+ export const O200K_ESTIMATOR = createO200kEstimator((text) => {
15
+ if (!encoder) {
16
+ // CJS exports are loaded on first opt-in use, not for every Pi session.
17
+ const { Tiktoken: Encoder } = load("js-tiktoken/lite") as typeof import("js-tiktoken/lite");
18
+ const ranks = load("js-tiktoken/ranks/o200k_base") as {
19
+ pat_str: string; special_tokens: Record<string, number>; bpe_ranks: string;
20
+ };
21
+ encoder = new Encoder(ranks);
22
+ }
23
+ return encoder.encode(text).length;
24
+ });
@@ -8,7 +8,7 @@ import {
8
8
  import type { ContextConfig } from "ds4-context-core/config/config";
9
9
  import { calculateContextBudget, type ContextBudget } from "ds4-context-core/core/budget-manager";
10
10
  import { createModelProfile, type ModelProfile } from "ds4-context-core/core/model-profile";
11
- import { estimateMessageTokens } from "ds4-context-core/core/token-estimator";
11
+ import { estimateMessageTokens, type TokenEstimator } from "ds4-context-core/core/token-estimator";
12
12
  import type {
13
13
  ArtifactManifestRef,
14
14
  ContextManifest,
@@ -54,6 +54,7 @@ export interface BuildPiObserverManifestOptions {
54
54
  profile?: ModelProfile;
55
55
  budget?: ContextBudget;
56
56
  modelAwareness?: ModelAwarenessManifest;
57
+ tokenEstimator?: TokenEstimator;
57
58
  plan?: ManagedContextPlan<PiAgentMessage>;
58
59
  projectRevision?: ProjectRevision;
59
60
  pins?: readonly PinManifestRef[];
@@ -386,6 +387,7 @@ export function buildPiObserverManifest(options: BuildPiObserverManifestOptions)
386
387
  profile,
387
388
  budget,
388
389
  systemPrompt: options.systemPrompt ?? options.ctx.getSystemPrompt(),
390
+ ...(options.tokenEstimator ? { tokenEstimator: options.tokenEstimator } : {}),
389
391
  ...(options.systemClassification ? { systemClassification: options.systemClassification } : {}),
390
392
  ...(options.systemPrivacyReason ? { systemPrivacyReason: options.systemPrivacyReason } : {}),
391
393
  tools: options.tools ?? activeTools(options.pi),
@@ -1,4 +1,4 @@
1
- export const EXTENSION_VERSION = "0.3.9";
1
+ export const EXTENSION_VERSION = "0.4.0";
2
2
  export const SUPPORTED_PI_VERSION = "0.84.3";
3
3
  export const OBSERVER_PLANNER_VERSION = "observer-model-aware-v1";
4
4
  export const PLANNER_VERSION = "managed-learned-ranking-v1";