ds4-context-engine 0.3.9 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +5 -2
- package/docs/ADR/064-per-project-databases-with-shared-calibration.md +88 -0
- package/docs/ADR/README.md +1 -0
- package/docs/CONTEXT_MANIFEST.md +1 -1
- package/docs/MODEL_AWARENESS.md +75 -4
- package/docs/PROVIDER_TOKEN_DRIFT_BENCHMARK.md +85 -0
- package/docs/STORAGE.md +38 -3
- package/docs/STORAGE_MAINTENANCE.md +1 -1
- package/docs/releases/0.3.10.md +29 -0
- package/docs/releases/0.3.9.md +1 -1
- package/docs/releases/0.4.0.md +86 -0
- package/package.json +4 -2
- package/src/extension/commands.ts +22 -2
- package/src/extension/runtime.ts +130 -34
- package/src/pi-adapter/bpe-token-estimator.ts +24 -0
- package/src/pi-adapter/context-observer.ts +3 -1
- package/src/pi-adapter/version.ts +1 -1
package/README.md
CHANGED
|
@@ -16,7 +16,7 @@ bounded active context with provenance
|
|
|
16
16
|
Pi provider
|
|
17
17
|
```
|
|
18
18
|
|
|
19
|
-
> **Project status:** The coordinated `0.
|
|
19
|
+
> **Project status:** The coordinated `0.4.0` release adds `storage.scope: "agent" | "project"`, now defaulting to per-project SQLite projections with shared token calibration in the agent database (opt out with `storage.scope: "agent"`). The opt-in BPE estimation and bounded auto-tuning from `0.3.10` keep `chars-v1` and disabled auto-tuning as their defaults. The bounded compaction controls from `0.3.9` remain in place; canonical history, SQLite schema 16 and runtime contracts are unchanged. Pi remains pinned to `0.84.3`. See the [0.4.0 release record](docs/releases/0.4.0.md), [ADR 064](docs/ADR/064-per-project-databases-with-shared-calibration.md) and [model-awareness validation](docs/MODEL_AWARENESS.md).
|
|
20
20
|
|
|
21
21
|
**Current compaction defaults:** `compaction.directUpdate=true`, `compaction.inputBudget="context"`, `compaction.segmentTargetTokens=30000`, `compaction.maxRequestInputTokens=64000`, `compaction.maxOperationInputTokens=2000000`, `compaction.maxConcurrentSegments=2`. Every DS4 provider attempt is bounded by the effective request limit, and the operation limit includes retries; `inputBudget="summary"` remains an explicit throughput-oriented opt-in. Existing compaction/master switches still apply. See [latency controls and compatibility](docs/COMPACTION.md#latency-controls). No real-provider speedup is claimed from mock tests. The five optional editing/reading/artifact/job features introduced in `0.3.4` remain default-off.
|
|
22
22
|
|
|
@@ -385,6 +385,7 @@ The following example shows the main configuration groups. Omitted values use th
|
|
|
385
385
|
"logLevel": "info"
|
|
386
386
|
},
|
|
387
387
|
"storage": {
|
|
388
|
+
"scope": "project",
|
|
388
389
|
"databasePath": "ds4-context/context.db",
|
|
389
390
|
"busyTimeoutMs": 5000,
|
|
390
391
|
"writeRetryTimeoutMs": 30000,
|
|
@@ -514,6 +515,8 @@ scripts package and release-readiness checks
|
|
|
514
515
|
- [Roadmap 0.2.0](docs/ROADMAP_0.2.0.md)
|
|
515
516
|
- [Release process](docs/RELEASING.md)
|
|
516
517
|
- [0.2.0 release readiness](docs/RELEASE_READINESS_0.2.0.md)
|
|
518
|
+
- [0.4.0 release notes](docs/releases/0.4.0.md)
|
|
519
|
+
- [0.3.10 release notes](docs/releases/0.3.10.md)
|
|
517
520
|
- [0.3.9 release notes](docs/releases/0.3.9.md)
|
|
518
521
|
- [0.3.8 release notes](docs/releases/0.3.8.md)
|
|
519
522
|
- [0.3.7 release notes](docs/releases/0.3.7.md)
|
|
@@ -541,7 +544,7 @@ scripts package and release-readiness checks
|
|
|
541
544
|
|
|
542
545
|
The original M0–M13 roadmap is complete. `ds4-context-core` contains the compiled runtime-neutral implementation. M14 context-quality metrics, M15 rich symbol indexing, M16 hybrid semantic retrieval, M17 cross-session project memory, M18 learned-ranking shadow evaluation, M19's runtime adapter/conformance kit, and M20 opt-in local KV eligibility/replay are implemented on `main`. Learned active ranking remains promotion-gated, Pi reports local KV as unsupported, and static ranking/native completion stay authoritative on every failure.
|
|
543
546
|
|
|
544
|
-
The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.3.9 extends the bounded compaction updates, summary input headroom, concurrent segments and phase timings introduced in 0.3.5 with per-request and cumulative operation input limits; 0.3.8 adds indexed FTS key deletion without changing search results. The opt-in [anchored editing](docs/ANCHORED_EDITING.md) and [portable agent tools](docs/PORTABLE_AGENT_TOOLS.md) from 0.3.4 remain default-off, without backend rewind, forced sampling or operational KV integration. Confirmation, provenance, Pi fallback and canonical/configuration/SQLite/runtime contracts remain unchanged. The [0.2 readiness record](docs/RELEASE_READINESS_0.2.0.md) remains the compatibility baseline; the lexical planner stays available as the deterministic fallback.
|
|
547
|
+
The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.4.0 introduces `storage.scope` (`"agent" | "project"`, default `"project"`): one rebuildable SQLite projection per trusted canonical project root, with token calibration shared in the agent database; `storage.scope: "agent"` restores the previous single-file layout. Version 0.3.10 adds opt-in BPE estimation and bounded model-budget auto-tuning while retaining the `chars-v1` default. Version 0.3.9 extends the bounded compaction updates, summary input headroom, concurrent segments and phase timings introduced in 0.3.5 with per-request and cumulative operation input limits; 0.3.8 adds indexed FTS key deletion without changing search results. The opt-in [anchored editing](docs/ANCHORED_EDITING.md) and [portable agent tools](docs/PORTABLE_AGENT_TOOLS.md) from 0.3.4 remain default-off, without backend rewind, forced sampling or operational KV integration. Confirmation, provenance, Pi fallback and canonical/configuration/SQLite/runtime contracts remain unchanged. The [0.2 readiness record](docs/RELEASE_READINESS_0.2.0.md) remains the compatibility baseline; the lexical planner stays available as the deterministic fallback.
|
|
545
548
|
|
|
546
549
|
## Contributing
|
|
547
550
|
|
|
@@ -0,0 +1,88 @@
|
|
|
1
|
+
# 064 — Per-project databases with shared token calibration
|
|
2
|
+
|
|
3
|
+
**Date:** 2026-09-25
|
|
4
|
+
**Status:** Accepted
|
|
5
|
+
**Related:** [002](002-pi-jsonl-canonical-sqlite-rebuildable.md), [058](058-bounded-manifest-storage.md), [063](063-fts-rowid-key-mappings.md)
|
|
6
|
+
|
|
7
|
+
## Context
|
|
8
|
+
|
|
9
|
+
The storage plan decided **D1**: keep one shared SQLite projection at
|
|
10
|
+
`~/.pi/agent/ds4-context/context.db` for every Pi session and explicitly avoid
|
|
11
|
+
per-session *or per-project* databases. That projection is deliberately
|
|
12
|
+
derived, rebuildable and disposable.
|
|
13
|
+
|
|
14
|
+
The session index (`entries`, `entries_fts`) is the table that actually grows
|
|
15
|
+
with total indexed history, and automatic eviction remains deferred until an
|
|
16
|
+
on-demand rehydration path is verified. A single file therefore also means a
|
|
17
|
+
single growth boundary, a single write lock and a single point of physical
|
|
18
|
+
reset shared by unrelated projects. Content is already partitioned logically
|
|
19
|
+
by `sessions.project_path`, `project_states`, `project_files` and
|
|
20
|
+
`project_memory_sessions`; the file boundary was the only thing missing.
|
|
21
|
+
|
|
22
|
+
The one piece of genuinely global learning is token calibration. Schema v10's
|
|
23
|
+
`token_calibration` is keyed by exact `provider + model + estimator_version`
|
|
24
|
+
and has **no project column**: a naive per-project split would restart
|
|
25
|
+
calibration for every project (minimum 3 samples, 8 accepted for the opt-in
|
|
26
|
+
`autoTune` expansion), weakening exactly the BPE/auto-tuning path it should
|
|
27
|
+
protect.
|
|
28
|
+
|
|
29
|
+
## Decision
|
|
30
|
+
|
|
31
|
+
Add `storage.scope: "agent" | "project"` (default `"project"`) to the storage
|
|
32
|
+
configuration:
|
|
33
|
+
|
|
34
|
+
- `agent` keeps the previous shared-database behavior and remains selectable.
|
|
35
|
+
- `project` derives one database per trusted canonical project root:
|
|
36
|
+
`projects/<sha256(canonicalRoot)[0..32]>.db`, next to the configured agent
|
|
37
|
+
database. Untrusted projects, broad roots (home directory, filesystem root)
|
|
38
|
+
and any resolution failure fall back to the agent database.
|
|
39
|
+
|
|
40
|
+
Both files receive the **same schema and the same migrations**; there is no
|
|
41
|
+
schema fork and migrations 1–15 are untouched. The split changes only which
|
|
42
|
+
repository each handle is used for:
|
|
43
|
+
|
|
44
|
+
- the agent database keeps `token_calibration` (shared learning);
|
|
45
|
+
- the project database keeps the session index, project index, context
|
|
46
|
+
manifests, summary graph, memory/pin projections, embeddings, quality
|
|
47
|
+
samples and artifact metadata; artifact object bytes move to
|
|
48
|
+
`projects/artifacts/<project-digest>/` so the orphan garbage collector,
|
|
49
|
+
which only sees references in the current database, can never delete
|
|
50
|
+
another project's objects;
|
|
51
|
+
- `resource_leases` and the client lease stay per file, protecting each
|
|
52
|
+
database independently.
|
|
53
|
+
|
|
54
|
+
With `project`, a calibration sample is derived from the project manifest and
|
|
55
|
+
inserted into the agent database with `manifest_id = NULL` (the column and its
|
|
56
|
+
partial unique index already allow this). The two writes are intentionally
|
|
57
|
+
**not** one cross-database transaction: losing one calibration sample is
|
|
58
|
+
harmless, whereas losing manifest/usage consistency is not. In `agent` scope
|
|
59
|
+
the previous single transaction is unchanged. Manifest pruning only detaches
|
|
60
|
+
calibration rows in `agent` scope; the project database's calibration table
|
|
61
|
+
stays empty.
|
|
62
|
+
|
|
63
|
+
Every project database starts empty and is rebuilt from canonical Pi JSONL,
|
|
64
|
+
project files and memory/pin `CustomEntry` records. The previous shared
|
|
65
|
+
database is left untouched: with the default change, existing users cold-start
|
|
66
|
+
their per-project indexes while their existing calibration remains available
|
|
67
|
+
in the agent database.
|
|
68
|
+
|
|
69
|
+
## Consequences
|
|
70
|
+
|
|
71
|
+
- Physical isolation per project: separate growth, separate write lock,
|
|
72
|
+
"reset project state" = remove one file, and no eviction needed to bound a
|
|
73
|
+
single project's index.
|
|
74
|
+
- Calibration stays global: a sample learned in project A is immediately
|
|
75
|
+
visible in project B for the same provider/model/estimator.
|
|
76
|
+
- Default behavior changes. A pre-existing shared database becomes the agent
|
|
77
|
+
database (calibration and old manifests) and is no longer the active
|
|
78
|
+
projection for new sessions. This requires a minor release and release
|
|
79
|
+
notes; `storage.scope: "agent"` restores the old layout.
|
|
80
|
+
- Maintenance and diagnostics become per file: `/context storage` reports the
|
|
81
|
+
active project database and, when split, the shared agent database;
|
|
82
|
+
`ds4-context-storage inspect|compact|recover --database <path>` must be
|
|
83
|
+
pointed at each file.
|
|
84
|
+
- `storage.databasePath` now names the agent database; project databases
|
|
85
|
+
derive from its directory. A manually configured per-project path keeps
|
|
86
|
+
working, but the derived `projects/` directory is the supported layout.
|
|
87
|
+
- D1 is superseded by this ADR; the development plan keeps the original text
|
|
88
|
+
with an explicit amendment pointer.
|
package/docs/ADR/README.md
CHANGED
|
@@ -67,5 +67,6 @@ The initial decisions from the development plan are accepted:
|
|
|
67
67
|
| [061](061-compaction-latency.md) | Bound compaction update calls, input budgets, concurrent segments and phase timings | Accepted |
|
|
68
68
|
| [062](062-cache-aware-context-planning.md) | Opt-in cache-aware tail planning using model pricing and observed cache shares | Accepted |
|
|
69
69
|
| [063](063-fts-rowid-key-mappings.md) | Resolve FTS key deletes through rowid mapping tables | Accepted |
|
|
70
|
+
| [064](064-per-project-databases-with-shared-calibration.md) | Split project state into per-project databases and keep token calibration shared | Accepted |
|
|
70
71
|
|
|
71
72
|
Each decision will receive a dedicated record when implementation pressure introduces alternatives or consequences not already covered by the development plan.
|
package/docs/CONTEXT_MANIFEST.md
CHANGED
|
@@ -45,7 +45,7 @@ actualInputTokens = input + cacheRead + cacheWrite
|
|
|
45
45
|
rawCalibrationRatio = actualInputTokens / estimatedInputTokens
|
|
46
46
|
```
|
|
47
47
|
|
|
48
|
-
The existing scalar SQLite columns are authoritative for persisted usage. `message_end` updates those columns and the calibration sample without reading or rewriting `manifest_json`; repository reads hydrate the usage fields from the scalars. The raw estimate is retained even after calibration so ratios cannot recursively calibrate already-corrected values. The next call reads only the exact provider/model window and applies bounded median/MAD outlier rejection. Error, aborted, missing, zero-usage, duplicate, or uncorrelated responses do not create calibration samples. Every manifest is correlated with at most one assistant response. If a concurrently pruned or oversize-skipped manifest cannot be correlated, DS4 keeps only bounded volatile calibration and does not retry against another row. See [`MODEL_AWARENESS.md`](MODEL_AWARENESS.md).
|
|
48
|
+
The existing scalar SQLite columns are authoritative for persisted usage. `message_end` updates those columns and the calibration sample without reading or rewriting `manifest_json`; repository reads hydrate the usage fields from the scalars. The raw estimate is retained even after calibration so ratios cannot recursively calibrate already-corrected values. The next call reads only the exact provider/model/**estimator-version** window and applies bounded median/MAD outlier rejection. When present, `modelAwareness.drift` and `modelAwareness.autoTune` contain only aggregate warning/tuning metadata (no message contents). Error, aborted, missing, zero-usage, duplicate, or uncorrelated responses do not create calibration samples. Every manifest is correlated with at most one assistant response. If a concurrently pruned or oversize-skipped manifest cannot be correlated, DS4 keeps only bounded volatile calibration and does not retry against another row. See [`MODEL_AWARENESS.md`](MODEL_AWARENESS.md).
|
|
49
49
|
|
|
50
50
|
## Retention
|
|
51
51
|
|
package/docs/MODEL_AWARENESS.md
CHANGED
|
@@ -1,5 +1,7 @@
|
|
|
1
1
|
# Advanced Model Awareness
|
|
2
2
|
|
|
3
|
+
**Validation status (OpenRouter, synthetic DS4/Pi sessions):** GPT-6-sol, GPT-6-luna and GPT-5.6-terra have each demonstrated multi-turn BPE calibration, a genuinely larger opt-in tail budget, provider usage above the native-window headroom threshold, and withdrawal of that expansion on the next turn in the same session. Historical retrieval and a live tool cycle were verified separately for all three. These bounded experiments do not establish universal tokenizer accuracy, provider-side compaction behavior, or production retrieval quality. `chars-v1` remains the default; BPE and `autoTune` remain opt-in. See the measured runs below.
|
|
4
|
+
|
|
3
5
|
M11 resolves an independent deterministic planning profile for every exact `provider/model`. Pi's model descriptor remains the default source for context window, output ceiling, reasoning, and image support; explicit DS4 overrides can repair provider metadata or tune category limits without changing canonical session state.
|
|
4
6
|
|
|
5
7
|
## Profile resolution
|
|
@@ -68,10 +70,10 @@ Each successful assistant response is correlated with one pending Context Manife
|
|
|
68
70
|
|
|
69
71
|
```text
|
|
70
72
|
actualInputTokens = input + cacheRead + cacheWrite
|
|
71
|
-
ratio = actualInputTokens / raw
|
|
73
|
+
ratio = actualInputTokens / raw selected-estimator estimate
|
|
72
74
|
```
|
|
73
75
|
|
|
74
|
-
Calibration is isolated by exact provider and
|
|
76
|
+
Calibration is isolated by exact provider, model ID **and estimator version**; switching from `chars-v1` to BPE never reuses the old ratios. The latest configured window is validated and processed deterministically:
|
|
75
77
|
|
|
76
78
|
1. reject invalid or zero samples;
|
|
77
79
|
2. reject ratios outside the configured hard bounds;
|
|
@@ -82,7 +84,40 @@ Calibration is isolated by exact provider and model ID. The latest configured wi
|
|
|
82
84
|
|
|
83
85
|
Until enough samples exist, the multiplier remains `1.0`. The default window is 24 samples, minimum is 3, and hard ratio bounds are 0.5–2.0. Outliers remain counted in metadata but do not influence the applied ratio.
|
|
84
86
|
|
|
85
|
-
Provider-token capacities are computed first from context window, output reserve, safety margin, and policy ratios.
|
|
87
|
+
Provider-token capacities are computed first from context window, output reserve, safety margin, and policy ratios. Global hard/soft/preferred input limits are converted into local-estimator units using **at least** a multiplier of 1: observed underestimation may reduce them, but apparent overestimation on short prompts never raises them above nominal model limits. Raw manifest estimates remain uncalibrated so future samples do not feed a corrected estimate back into itself. Adaptive tail/history/project budgets use the accepted multiplier but remain capped by the configured `context.*` maxima even after conversion.
|
|
88
|
+
|
|
89
|
+
Calibration samples are global learning, not project data: with the default `storage.scope: "project"` the manifests live in the per-project database while `token_calibration` stays in the shared agent database (see [ADR 064](ADR/064-per-project-databases-with-shared-calibration.md)). A sample learned in one project therefore applies to every other project for the same provider/model/estimator.
|
|
90
|
+
|
|
91
|
+
### Optional BPE estimator and measured budget tuning
|
|
92
|
+
|
|
93
|
+
The default remains `chars-v1`. To opt a profile into local OpenAI `o200k_base` BPE text counting (without fetching vocabulary over the network):
|
|
94
|
+
|
|
95
|
+
```json
|
|
96
|
+
{
|
|
97
|
+
"modelAwareness": {
|
|
98
|
+
"overrides": { "openai/gpt-4o": { "tokenEstimator": "o200k-base-v1" } },
|
|
99
|
+
"autoTune": true
|
|
100
|
+
}
|
|
101
|
+
}
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
Use BPE only for a model known to share that encoding. The estimator counts text with BPE in the Pi observer, managed planner and fixed prompt/tool estimates; message wrappers, images, reasoning, provider-specific serialization, indexed retrieval hints, and persisted summaries remain estimates. It does **not** make token counts exact or enable provider-side continuation. If the provider's actual input tokens drift persistently from estimates (median ≥1.25 or ≤0.80 after the calibration minimum, or repeated hard-bound outliers), `/context model` and status surface a metadata-only warning; no prompt content is logged.
|
|
105
|
+
|
|
106
|
+
`autoTune` is opt-in and starts neutral. After at least eight accepted calibration samples, it expands *automatic* recent-tail, history and project ceilings by at most 12.5% of their nominal values, only if **every valid provider-usage sample** in the calibration window — including ratio outliers excluded from ratio calibration — remains at or below 60% of the current preferred/hard input target. It never raises a ceiling beyond the configured context cap, never overrides an explicit model-specific category limit, and never changes the provider's window, output reserve, safety margin or hard/soft input limits. Without enough evidence or headroom it stays neutral. Any calibration multiplier is applied after this small expansion; tune decisions and warnings are recorded as metadata in the manifest.
|
|
107
|
+
|
|
108
|
+
## Validation goal for BPE and auto-tuning
|
|
109
|
+
|
|
110
|
+
**User-requested outcome:** verify long conversations, tools, retrieval, calibration across turns, and auto-tuning safety on the context actually built by DS4/Pi. Single-turn synthetic token probes are preliminary evidence only; they do **not** satisfy this goal or justify declaring the feature complete.
|
|
111
|
+
|
|
112
|
+
The validation must cover, and report results separately for:
|
|
113
|
+
|
|
114
|
+
1. Long multi-turn sessions, including large context windows, compaction boundaries, and preservation of the current request.
|
|
115
|
+
2. Tool calls and results (including large results): atomic selection, offload, and the estimated versus observed token count.
|
|
116
|
+
3. Retrieved history and project context: relevance, budget selection, and whether relevant older material survives planning.
|
|
117
|
+
4. Consecutive turns on the **same** provider/model and estimator profile: provider usage correlation, accepted/rejected calibration samples, changes to the applied ratio, and isolation when models or estimators change.
|
|
118
|
+
5. Opt-in auto-tuning: evidence thresholds, configured and hard limits, outliers, insufficient headroom, and no regression in retrieval or context integrity when limits expand.
|
|
119
|
+
|
|
120
|
+
Use local, deterministic integration tests where possible. If real provider calls are needed, require a **new explicit call and input-size limit** before sending anything; the earlier single-turn budgets are exhausted. Report any scenario not verified as *not verified*, rather than treating the synthetic pilot as an end-to-end result. Do not change the default estimator or enable auto-tuning by default on the strength of that pilot.
|
|
86
121
|
|
|
87
122
|
## Provider cache metrics
|
|
88
123
|
|
|
@@ -126,6 +161,42 @@ Use:
|
|
|
126
161
|
/context status
|
|
127
162
|
```
|
|
128
163
|
|
|
164
|
+
## Local validation progress and outstanding provider evidence
|
|
165
|
+
|
|
166
|
+
`tests/integration/model-aware-real-context.test.ts` exercises the Pi `context` and `message_end` hooks over canonical multi-turn JSONL without any network transport. A long branch with a tool call/result verifies atomic selection and retrieval of an older decision under a BPE-managed budget. Successive turns verify that eight correlated usage records are required before opt-in expansion, that estimator/model switches isolate calibration, and that a new near-limit usage record withdraws the expansion even when its ratio is excluded from calibration. `tests/unit/model-aware-estimator.test.ts` also checks explicit overrides, configured caps, and the high-usage outlier regression.
|
|
167
|
+
|
|
168
|
+
**Local integration checks inject usage; they are not independent provider measurements.** A separately authorized, bounded live probe (`scripts/verify-model-aware-real-session.mjs`) used a temporary Pi JSONL session, DS4's actual hooks, synthetic content, `openrouter/openai/gpt-4o-mini`, and Pi-normalized provider usage. Eight successive short calls accumulated calibration samples; on the ninth call the manifest showed eight accepted samples, applied ratio 0.677157, and opt-in `autoTune: expanded`. That ninth call did **not** contain the intended appended long branch: Pi had already constructed its runner, and its manifest estimated only 1,940 input tokens. It is not evidence for long-context tuning.
|
|
169
|
+
|
|
170
|
+
A second, fresh Pi session with the long branch seeded *before* runner construction produced 83 original messages; DS4 selected 21 groups, excluded 21, retrieved the older `cobalt-713` decision, and included both the synthetic tool-call assistant entry and tool result. For that request, the BPE-based manifest estimate was **23,492** input tokens and provider usage was **23,196**. This is **one** live, untuned long-session/tool/retrieval measurement; it does not prove stable drift across lengths or models. The earlier [single-turn pilots](PROVIDER_TOKEN_DRIFT_BENCHMARK.md) remain separate. The two runs together attempted 10 calls and conservatively reserved 217,678 of the authorized 250,000 estimated-input-token limit (per-call ceiling 64,000; 16-call ceiling). All temporary session data was deleted.
|
|
171
|
+
|
|
172
|
+
A subsequent, separately authorized probe (`scripts/verify-model-aware-calibrated-session.mjs`) seeded a 28-turn synthetic history and tool call/result **before starting the same Pi session**. Eight real provider calls calibrated BPE on that long context; the ninth measured **23,926 estimated / 23,486 actual input tokens**, eight accepted samples, `autoTune: expanded`, and retrieval of the old decision with the historical tool call and result included. This verifies expansion on a real long context after calibration in one session, not just across separate sessions. Across this probe **11 calls** reserved **588,052/600,000** estimated input tokens (80,000 per-call limit); no further provider requests were made under that authorization.
|
|
173
|
+
|
|
174
|
+
The attempted higher-occupancy request used **37,552 estimated / 37,464 actual tokens**, below the 60%-of-target withdrawal threshold (**53,760** for this model/configuration). The following turn still reported `expanded`: **withdrawal is not verified live**. That larger turn also excluded the historical tool/retrieval groups, an observed quality limit under that selected context, not proof of an auto-tuning regression. The local Pi compaction-boundary test covers summary preservation under BPE. A separately bounded **one-request** live compaction probe (`scripts/verify-model-aware-compaction-session.mjs`) preseeded the canonical Pi compaction entry before starting Pi; the DS4 managed manifest included its summary and measured **250 estimated / 244 actual** provider input tokens. This used one further call under the third authorization, bringing that budget to **10 calls / 181,069 estimated input tokens reserved**. It checks the provider-bound context *after* a synthetic compaction entry, not live summary generation. At that point, live tool execution and high-usage withdrawal remained unverified; both were tested in the later authorized round below. Model variety and production retrieval quality still remain outside these bounded probes. A third bounded probe (`scripts/verify-model-aware-safety-session.mjs`) reached **60,509 actual input tokens** (>53,760), but only **seven** of eight preceding short-call samples had been accepted; the manifest at the high-usage request showed seven accepted samples and `insufficient-samples`, not an expanded policy. Thus it **does not** test withdrawal. The safety probe attempted **9 calls**, reserving **169,006** estimated input tokens, and stopped without a follow-up call. It now gates the expensive request on eight *accepted* samples and seeds a stable synthetic prefix, but has not been rerun. Its temporary session was deleted; the remaining third-round budget after the separate compaction request cannot fit a new calibration/high-usage/follow-up sequence under the preflight limits. No live tool execution was attempted in that round because the automatic multi-request cycle was not yet safely capped per request. Those results alone did not verify withdrawal or tool execution.
|
|
175
|
+
|
|
176
|
+
A final, separately authorized round verified the outstanding **synthetic** provider-path scenarios on the same OpenRouter/GPT-4o-mini profile. After a stable synthetic prefix, eight usage samples were accepted; the ninth short turn showed `expanded`. The next request reported **60,708 estimated / 60,640 actual** input tokens, above the **53,760** headroom threshold; the following turn reported **60,795 estimated / 60,722 actual** and `no-headroom` (recent-tail limit decreased from **27,699** to **24,564**). `scripts/verify-model-aware-safety-session.mjs` used **11/15** calls and **292,818/350,000** estimated-input tokens reserved (maximum 80,000 per call). This demonstrates withdrawal on measured provider usage, not merely a synthetic injected sample.
|
|
177
|
+
|
|
178
|
+
With the remaining budget, `scripts/verify-model-aware-live-tool.mjs` gated **each** Pi provider request at `ModelRuntime.streamSimple`, disabled cache warming/retry/compaction, and executed a single synthetic tool. The first request included its tool schema and used **105** provider input tokens; the second included the actual tool result and used **139**. Exactly **two** requests and one tool execution occurred. The final fourth-round total was **13/15 calls, 317,119/350,000** estimated-input tokens reserved. Only aggregate counters and booleans were reported; temporary synthetic sessions were removed.
|
|
179
|
+
|
|
180
|
+
These bounded live measurements, together with the canonical Pi JSONL integration tests, cover long context, historical and executed tools, retrieval, between-turn calibration, compaction-boundary context, and auto-tuning withdrawal **for this synthetic OpenRouter/GPT-4o-mini profile**. They do not establish universal accuracy for other providers/models, real private sessions, arbitrary compaction generation, or production retrieval quality.
|
|
181
|
+
|
|
182
|
+
### Native-window checks for Sol, Luna and Terra — separately authorized
|
|
183
|
+
|
|
184
|
+
The Pi catalog exposed a **1,050,000-token window** for each OpenRouter model `openai/gpt-6-sol`, `openai/gpt-6-luna`, and `openai/gpt-5.6-terra`. One additional authorization limited their *combined* probes to **42 provider attempts, 600,000 estimated input tokens / 2,500,000 controlled characters per request, and 4,800,000 estimated tokens / 15,000,000 controlled characters total**. `scripts/verify-model-aware-triad.mjs` and `scripts/verify-model-aware-triad-followup.mjs` gated **every** Pi request at `ModelRuntime.streamSimple`, including tool continuations; no real sessions, credentials, prompt bodies, responses or raw upstream errors were logged. Temporary canonical Pi JSONL sessions were deleted.
|
|
185
|
+
|
|
186
|
+
The first probe consumed **30 calls / 1,433,445 tokens / 3,679,430 characters reserved**. For *each* model, nine calibration calls in the same seeded long session led to at least eight accepted usage samples. The tenth call measured **36,027 BPE-estimated / 35,503 Pi-normalized actual input tokens**, `autoTune: expanded`, with the historical tool call and result included. **This did not verify retrieval**: the decision was still in the 64k recent tail and therefore had not been fetched by retrieval; the probe correctly stopped before tool execution and the high-occupancy turn. At this native profile, the reported `expanded` status did **not** demonstrate a larger recent-tail limit: its configured 64k cap was already saturated.
|
|
187
|
+
|
|
188
|
+
The second probe seeded **70 turns before Pi session construction**, putting the old decision outside the default recent tail. It consumed nine more calls (one retrieval request and an actual **two-request, one-execution** synthetic tool cycle per model). The decision was both retrieved and included, the historical tool call/result stayed included, and the new tool schema appeared in the first provider request with its actual result in the second. Per-model provider usage for the retrieval request was **63,050 / 63,049 / 63,053** tokens (Sol/Luna/Terra), against BPE estimates **63,699 / 63,698 / 63,702**. All three tool cycles completed. These are measurements of the DS4/Pi managed context, not three single-turn tokenizer prompts.
|
|
189
|
+
|
|
190
|
+
The three remaining authorized requests measured large BPE-managed input, **one per model**. The selected current user turn was included, and Pi reported **381,501 / 381,502 / 381,502** input tokens against manifest estimates **381,593 / 381,594 / 381,594**. The historical call/result were present, but the old decision was **not** selected; the oversized prompt did not ask about that decision, so this is **not** a valid retrieval-quality test. The request generator targeted 462k from the *raw* canonical history, but DS4 excluded older history and the selected provider input remained ~381.5k. Thus all three requests were **below** the native-window auto-tuning withdrawal threshold of **441,000** (60% of the 735k preferred target), and far below the full 1.05M window. They were fresh sessions without eight accepted calibration samples; **none tests `expanded → high provider usage → no-headroom` on these models**. Do not report the triad's auto-tuning safety or category expansion as live-verified at its native window; the deterministic local matrix in `tests/unit/model-aware-estimator.test.ts` covers only injected samples and caps. The entire authorization was used: **42/42 calls**, **3,296,138/4,800,000** estimated input tokens and **9,357,266/15,000,000** controlled characters reserved. No more provider calls are authorized under it. The historical probes now reject `--live` under this exhausted authorization. **After** those measurements, their local high-input sizing was corrected to require at least 470k BPE tokens in the selected current user text (rather than in raw canonical history), with a second gate immediately before provider transport; unit tests cover the discarded-history regression and per-call preflight. This correction was **not** run against a provider and cannot retroactively validate withdrawal.
|
|
191
|
+
|
|
192
|
+
Across these specific synthetic shapes, BPE manifests slightly overestimated provider usage, but no universal correction follows. At this earlier point native-window compaction generation, production retrieval quality and live high-usage auto-tune withdrawal for Sol/Luna/Terra were unverified; the later withdrawal runs below close **only the last of those gaps**. The later probes check actual category growth (not merely `expanded` status), a selected current turn above 470k BPE tokens, usage above 441k, and withdrawal in the same Pi session. Their temporary opt-in category ceilings are 80k/40k/40k because the default native tail ceiling is already equal to its configured maximum of 64k and cannot grow; the probes do not establish expansion under unchanged category defaults. BPE and auto-tuning remain opt-in; no default or provider-storage policy was changed.
|
|
193
|
+
|
|
194
|
+
### Native-window withdrawal round (later, opt-in categories)
|
|
195
|
+
|
|
196
|
+
With a separate explicit cap of 39 calls, 3,800,000 reserved input tokens and 11,000,000 controlled characters, the same-session Pi/DS4 probe used temporary category ceilings 80k/40k/40k to expose a **real** tail expansion. It ran 34 provider requests (3,375,536 estimated input tokens reserved; 9,845,274 controlled characters). For **GPT-6-sol and GPT-6-luna**, each session accepted at least eight calibration samples; the expanded tail exceeded the same-ratio no-tune baseline (about 72,922 versus 64,819 estimator tokens). A later request used **475,703 provider input tokens** (BPE estimate 475,790), above the 441,000 provider-token headroom threshold. On the next request, `autoTune` was `no-headroom` and the tail equalled its no-tune baseline (64,805); actual provider input was 471,846/471,844. The selected current turn, historical decision, call and result were present before the large request. The large input and follow-up did **not** retain the unrelated old decision/tool group; these measurements do not establish production retrieval quality.
|
|
197
|
+
|
|
198
|
+
For **GPT-5.6-terra**, calibration and real expansion passed in that first round, but its high-input request was blocked **before transport** by the cumulative budget/character caps: the probe had incorrectly budgeted the follow-up as short even though Pi resends the large preceding turn. The temporary session was then disposed. A **separately authorized Terra-only run** used a fresh synthetic Pi/DS4 session and counted *both* large requests. It made 12 provider calls (1,449,079 estimated tokens and 4,310,843 characters reserved, under separate 13-call / 1,900,000-token / 6,000,000-character maxima). After at least eight accepted calibration samples, Terra's tail grew to 72,922 estimator tokens versus a same-ratio untuned baseline of 64,819. The large request used **475,707 provider input tokens** (BPE estimate 475,794), above the 441,000 threshold. On the next turn, provider input was **471,852** (estimate 471,929), `autoTune` was `no-headroom`, and the tail returned to **64,805**, exactly its untuned baseline at that turn's ratio. The historical decision/call/result and current turn were included before the large request; unrelated older groups were excluded during the large request. **This verifies the defined live high-usage expansion/withdrawal scenario for all three models, not a general retrieval or compaction guarantee.** Both probes are now locked against replay; defaults remain unchanged.
|
|
199
|
+
|
|
129
200
|
## Performance and tests
|
|
130
201
|
|
|
131
|
-
`tests/benchmarks/model-awareness.bench.ts` measures a bounded 200-sample calibration analysis and repeated 32k/128k/200k profile resolution. Unit and golden tests cover deterministic tiers, override precedence, robust outlier rejection, cache accounting, and calibrated budgets. Integration tests switch local/remote providers and 32k/128k/200k models while checking profile isolation, privacy re-enforcement, canonical JSONL preservation, SQLite cache metrics, and profile reuse.
|
|
202
|
+
`tests/benchmarks/model-awareness.bench.ts` measures a bounded 200-sample calibration analysis and repeated 32k/128k/200k profile resolution. Unit and golden tests cover deterministic tiers, override precedence, robust outlier rejection, cache accounting, and calibrated budgets. Integration tests switch local/remote providers and 32k/128k/200k models while checking profile isolation, privacy re-enforcement, canonical JSONL preservation, SQLite cache metrics, and profile reuse. For a separately authorized, bounded real-provider measurement of both estimators against Pi SDK usage, see [PROVIDER_TOKEN_DRIFT_BENCHMARK.md](PROVIDER_TOKEN_DRIFT_BENCHMARK.md); it does not substitute for DS4 manifest measurements in a live session.
|
|
@@ -0,0 +1,85 @@
|
|
|
1
|
+
# Live provider token-drift probe (opt-in)
|
|
2
|
+
|
|
3
|
+
This developer-only, **source-checkout-only** probe (the script is not included in the published npm package) compares DS4's raw `chars-v1` and `o200k-base-v1` estimates with **Pi SDK provider usage** for synthetic, single-turn text. It does not start a Pi session, read session JSONL or SQLite, record responses, or change DS4 configuration. It never prints API keys, requests, response bodies or raw upstream error messages. It requires local Pi provider authentication and network access. The probe runs only with `--live`; without that flag it prints the bounded plan.
|
|
4
|
+
|
|
5
|
+
```bash
|
|
6
|
+
npm run build:core
|
|
7
|
+
node scripts/compare-provider-token-drift.mjs \
|
|
8
|
+
--model openrouter/openai/gpt-4o-mini \
|
|
9
|
+
--model deepseek/deepseek-v4-flash \
|
|
10
|
+
--model openai-codex/gpt-5.4-mini
|
|
11
|
+
# After inspecting the dry-run budget, add --live to make real calls.
|
|
12
|
+
```
|
|
13
|
+
|
|
14
|
+
The script accepts repeated `--model provider/model-id` and optional `--sizes 512,4096,48000`. **Per invocation**, it refuses more than 12 model calls or 200,000 total input characters (system prompt included), disables SDK request retries, and imposes a 45-second deadline per call. Multiple invocations have separate caps: track their cumulative spend yourself. Output is one metadata-only JSON report on stdout. The SDK's catalog-derived USD cost is an estimate, not a bill; a provider can still charge for failures or retries outside this script's control. If you pipe output to a file, keep it outside the repo unless you intentionally want to publish the aggregate metadata.
|
|
15
|
+
|
|
16
|
+
For each successful call, `actualInputTokens = usage.input + usage.cacheRead + usage.cacheWrite`; the cached fractions are **not** extra tokens on top of that sum. `residualTokens = actual - rawEstimate` (positive means underestimation), and `underestimationPctOfActual = max(0, residual) / actual × 100`. Adjacent slopes use differences across prompt sizes for a single exact provider/model, removing most fixed framing. BPE counts only text; both estimators use the same DS4 message/system overhead. No raw provider payload is available from this probe, so the result does **not** prove the accuracy of DS4's full observer in a live Pi extension chain (tools, images, privacy, cache/continuation, and later extensions differ). Compare `/context model`, `/context tokens`, and manifest `actualInputTokens` versus `estimatedInputTokens` in an ordinary consented Pi session for that second step. Never copy session text or auth files into the report.
|
|
17
|
+
|
|
18
|
+
## First observed run — 24 September 2026, 10:09–10:13 UTC
|
|
19
|
+
|
|
20
|
+
User-approved cumulative envelope: 12 call attempts and 200,000 input characters. Three invocations totalled **12 attempts and 195,176 planned input characters**: 9 calls/158,382 characters, a single Codex diagnostic call/574 characters, then 2 calls/36,220 characters. Success: 8; failure/missing usage: 4. Synthetic multilingual/code-like repeated text; no private session data. Sizes below are user-message characters; system prompt is included in estimates and the total envelope. These are observations, **not** a representative multi-session cost or quality benchmark.
|
|
21
|
+
|
|
22
|
+
| Provider / requested model | User chars | Real input¹ | `chars-v1` | BPE `o200k-base-v1` | BPE residual (real − estimate) |
|
|
23
|
+
|---|---:|---:|---:|---:|---:|
|
|
24
|
+
| OpenRouter / `openai/gpt-4o-mini` | 512 | 190 | 161 | 196 | −6 |
|
|
25
|
+
| OpenRouter / `openai/gpt-4o-mini` | 4,096 | 1,436 | 1,057 | 1,442 | −6 |
|
|
26
|
+
| OpenRouter / `openai/gpt-4o-mini` | 48,000 | 16,668 | 12,033 | 16,674 | −6 |
|
|
27
|
+
| DeepSeek / `deepseek-v4-flash`² | 512 | 179 | 161 | 196 | −17 |
|
|
28
|
+
| DeepSeek / `deepseek-v4-flash`² | 4,096 | 1,388 | 1,057 | 1,442 | −54 |
|
|
29
|
+
| DeepSeek / `deepseek-v4-flash`² | 48,000 | 16,172 | 12,033 | 16,674 | −502 |
|
|
30
|
+
| OpenRouter / `deepseek/deepseek-v4-flash` | 4,096 | 1,387 | 1,057 | 1,442 | −55 |
|
|
31
|
+
| OpenRouter / `deepseek/deepseek-v4-flash` | 32,000 | 10,785 | 8,033 | 11,124 | −339 |
|
|
32
|
+
|
|
33
|
+
¹ Pi's normalized provider usage: input + cache read + cache write. Some long requests reported cached reads (up to 1,408 tokens); they were included exactly once. Successful responses produced 1–3 output tokens. ² DeepSeek reported response model `deepseek-flash`, an alias of the requested ID; no equivalence to OpenRouter routing is assumed.
|
|
34
|
+
|
|
35
|
+
- OpenRouter GPT-4o-mini: raw `chars-v1` median actual/estimate **1.358562**; raw BPE median **0.995839**. At 48k chars, chars/4 underestimated by 4,635 tokens (27.81% of actual), while BPE overestimated by 6 tokens. Adjacent BPE slopes were 1.000 and 1.000.
|
|
36
|
+
- Direct DeepSeek: raw chars median **1.313150**; BPE median **0.962552**. At 48k chars, chars/4 underestimated by 4,139 tokens (25.59% of actual), while BPE overestimated by 502 tokens. Adjacent BPE slopes were 0.970305 and 0.970588. Close agreement here does **not** establish that DeepSeek uses OpenAI's tokenizer or justify turning BPE on for that model by default.
|
|
37
|
+
- OpenRouter DeepSeek: two valid samples; raw chars median **1.327396**, BPE median **0.965692**. At 32k chars, chars/4 underestimated by 2,752 tokens; BPE overestimated by 339. Two samples are insufficient to validate a persistent drift warning or tuning.
|
|
38
|
+
- `openai-codex/gpt-5.4-mini`: three initial attempts produced no usable usage; one later 512-character diagnostic returned `request-failed` with sanitized category `other` and no HTTP status. The local Pi auth resolver did return a credential, but the reason for the probe failure remains **unverified**. Do not infer its tokenizer drift, account availability in the interactive TUI, or provider-side token consumption from these calls. Direct `openai/gpt-4o-mini` API auth was not configured in the Pi runtime used by this probe; OpenRouter is a distinct route.
|
|
39
|
+
|
|
40
|
+
**Interpretation:** Three sizes with one call each do not establish statistical reliability or cover large context windows, tools, images, prefixes reused across turns, or output-heavy workflows. `chars-v1` remains the default; existing per-model calibration may compensate after enough **accepted same-profile samples**. The first small sample can be excluded by MAD filtering, so three different-size wire calls need not become three accepted DS4 calibration samples. Keep BPE and `autoTune` opt-in; do not port Hub's fixed 3.5% margin from these measurements. For promotion or automatic budget changes, collect a repeated same-size and mixed-size sample with matching DS4 manifest estimates and actual provider usage under an explicitly authorized, separately bounded run.
|
|
41
|
+
|
|
42
|
+
## Second observed run — 24 September 2026, 10:35 UTC
|
|
43
|
+
|
|
44
|
+
A separate, explicitly approved envelope covered **12 call attempts and up to 100,000 input characters**. The dry run planned exactly 12 attempts and 99,816 characters: `--sizes 512,16000` for three requested OpenAI models through each of OpenRouter and Codex. The live probe produced six OpenRouter usages and six Codex failures; it did not read session data.
|
|
45
|
+
|
|
46
|
+
| Route / requested model | User chars | Real input¹ | `chars-v1` | BPE `o200k-base-v1` | BPE residual |
|
|
47
|
+
|---|---:|---:|---:|---:|---:|
|
|
48
|
+
| OpenRouter / `openai/gpt-6-sol` | 512 | 189 | 161 | 196 | −7 |
|
|
49
|
+
| OpenRouter / `openai/gpt-6-sol` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
|
|
50
|
+
| OpenRouter / `openai/gpt-6-luna` | 512 | 189 | 161 | 196 | −7 |
|
|
51
|
+
| OpenRouter / `openai/gpt-6-luna` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
|
|
52
|
+
| OpenRouter / `openai/gpt-5.6-terra` | 512 | 189 | 161 | 196 | −7 |
|
|
53
|
+
| OpenRouter / `openai/gpt-5.6-terra` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
|
|
54
|
+
|
|
55
|
+
¹ Pi-normalized input includes cached tokens once. Each 16,000-character call reported `cacheWriteTokens = 5,561` and zero cache-read tokens. Those were **writes**, not cache hits. Each successful response reported five output tokens. The six equal input counts describe these identical synthetic prompts on this route; they are not proof of identical tokenizers or of how the direct Codex route would count a DS4 session. For each model, the adjacent BPE size slope was exactly 1.000; at 16k characters, raw `chars-v1` underestimated by 1,531 tokens (27.52% of actual), versus BPE overestimating by seven.
|
|
56
|
+
|
|
57
|
+
For `openai-codex/gpt-6-sol`, `gpt-6-luna`, and `gpt-5.6-terra`, **both sizes failed** without provider usage. The sanitized error classifier returned `quota` for all six, with no HTTP status captured. This is evidence of an SDK-path quota/limit error classification, **not** a verified account balance, an HTTP rejection, or a tokenizer measurement. It does not retroactively identify the earlier `gpt-5.4-mini` failure (classified `other`). No further calls were made beyond this run's approved envelope. Two sizes per exact model remain insufficient for DS4 calibration or auto-tuning conclusions.
|
|
58
|
+
|
|
59
|
+
## Isolated Pi + DS4 managed-context pilot — 24 September 2026
|
|
60
|
+
|
|
61
|
+
For a future, **separately authorized** run: build the core (`npm run build:core`), inspect the planned sizes and character count with `node scripts/compare-ds4-manifest-usage.mjs`, then use `node scripts/compare-ds4-manifest-usage.mjs --live` only after setting a new call/character budget. The default mode is pinned to OpenRouter `openai/gpt-6-sol` and a maximum of 12 calls / 120,000 controlled prompt characters per invocation. `--comparison` instead preflights **both** OpenRouter `openai/gpt-6-luna` and `openai/gpt-5.6-terra`, capped at 24 calls / 240,000 controlled characters combined; `--comparison --live` requires its own authorization. A new authorization is required even when repeating either exact plan.
|
|
62
|
+
|
|
63
|
+
The user separately authorized **up to 12 calls / 120,000 controlled input characters** to OpenRouter `openai/gpt-6-sol` only. `scripts/compare-ds4-manifest-usage.mjs` first checked a local sandbox without provider traffic, then made **12 single-turn calls**, with three repeats at each size (512, 4,096, 12,000 and 20,000 synthetic user characters). Planned synthetic user + fixed system text was **110,568 characters**. This limit counts controlled prompt text, **not** Pi-generated framing or serialized protocol bytes. No existing Pi session history was used; each Pi session and DS4 database was isolated in a temporary directory, with no tools, project content, native continuation, or auto-tuning. DS4's managed `context` hook and its `o200k-base-v1` profile override were active. The script stops on a missing/mismatched manifest or missing usage rather than fabricating a measurement.
|
|
64
|
+
|
|
65
|
+
| Synthetic user chars | First observed DS4 manifest estimate | First actual input¹ | First residual (actual − estimate) | Repeats |
|
|
66
|
+
|---:|---:|---:|---:|---:|
|
|
67
|
+
| 512 | 220 | 209 | −11 | 3 |
|
|
68
|
+
| 4,096 | 1,467 | 1,456 | −11 | 3 |
|
|
69
|
+
| 12,000 | 4,210 | 4,199 | −11 | 3 |
|
|
70
|
+
| 20,000 | 6,983 | 6,972 | −11 | 3 |
|
|
71
|
+
|
|
72
|
+
All **12** calls had matching `managed` manifests and Pi-normalized provider usage, with **exactly −11 tokens** of residual in each call. Token counts varied by about one or two between identically sized repeats, but the residual stayed fixed. The largest relative overestimate was **5.26% of actual input** on the first 512-character call (11/209); at 20,000 characters it was about **0.16%**. `cacheReadTokens` was zero, while longer calls reported mostly cache **writes**; no claim of cache hits or exact billed usage follows from those fields.
|
|
73
|
+
|
|
74
|
+
¹ `totalInputTokens` in the DS4 manifest (the same Pi-normalized `input + cacheRead + cacheWrite` metric used by model-awareness calibration), not independently obtained provider invoice data. Because the selected estimator was BPE, this pilot did **not** produce an alternative `chars-v1` DS4 manifest for the same wire requests. These are fresh one-turn sessions with one synthetic prompt shape and a maximum of ~7k actual tokens; they do not validate multi-turn prefixes, tool schemas/results, retrieval, real project text, larger windows, other requested models, or production auto-tuning. The 12 samples are independent sessions, **not** 12 accepted samples in one persistent DS4 calibration profile. Keep BPE opt-in and `autoTune` off by default; do not apply an inferred fixed framing correction or Hub's fixed 3.5% drift allowance from this pilot.
|
|
75
|
+
|
|
76
|
+
## Separately authorized Luna/Terra managed-context comparison — 24 September 2026
|
|
77
|
+
|
|
78
|
+
A subsequent authorization covered **at most 24 calls / 240,000 controlled synthetic input characters combined** on OpenRouter `openai/gpt-6-luna` and `openai/gpt-5.6-terra`; no Codex traffic was authorized. `node scripts/compare-ds4-manifest-usage.mjs --comparison` preflighted **24 calls / 221,136 controlled characters**. `--comparison --live` completed **12 calls per model**, with three fresh, isolated Pi+DS4 managed sessions per model at each synthetic user size (512, 4,096, 12,000, 20,000 characters). No existing session history, project content, tools, auto-tuning, or native continuation entered the requests. All 24 calls had matching manifests and Pi-normalized provider input usage; there were no reported failed rows.
|
|
79
|
+
|
|
80
|
+
| Model | First 512-character DS4 BPE estimate | First actual input¹ | Repeats per size | Residual on all 12 calls |
|
|
81
|
+
|---|---:|---:|---:|---:|
|
|
82
|
+
| OpenRouter / `openai/gpt-6-luna` | 220 | 209 | 3 | −11 tokens |
|
|
83
|
+
| OpenRouter / `openai/gpt-5.6-terra` | 219 | 208 | 3 | −11 tokens |
|
|
84
|
+
|
|
85
|
+
In every size group for both models, each of the three residuals was **−11 tokens**; each group reported zero `cacheReadTokens`. This replicates the small, constant **overestimate** observed for Sol on this specific one-turn synthetic shape. It does **not** establish a provider-independent 11-token correction, tokenizer equivalence across arbitrary text, long-window accuracy, model calibration from a continuous session, or safe budget auto-tuning. The combined 24-call authorization is exhausted. A later, separate authorization covered multi-turn, historical and executed tools, retrieval, and large selected inputs for all three OpenRouter models; see [native-window checks](MODEL_AWARENESS.md#native-window-checks-for-sol-luna-and-terra--separately-authorized) for results and the remaining unverified high-usage auto-tune withdrawal. Codex diagnosis still needs separate approval; keep BPE opt-in and `autoTune` off by default.
|
package/docs/STORAGE.md
CHANGED
|
@@ -10,6 +10,19 @@ Pi's session JSONL is canonical for conversations and live project files are can
|
|
|
10
10
|
|
|
11
11
|
The extension never edits or rewrites Pi JSONL or project source files. Manual memory/pin commands, confirmed `context_persistence` canonical writes, and learned-ranking feedback append versioned classified Pi `CustomEntry` records through Pi's official `appendEntry()` API. The tool does not write SQLite as a substitute for a canonical Pin or Memory append.
|
|
12
12
|
|
|
13
|
+
## Storage scope
|
|
14
|
+
|
|
15
|
+
`storage.scope` selects where the disposable projection lives (see [ADR 064](ADR/064-per-project-databases-with-shared-calibration.md)):
|
|
16
|
+
|
|
17
|
+
- `agent` — the previous behavior: one shared `context.db` for every session and project.
|
|
18
|
+
- `project` (**default**) — one database per trusted canonical project root, derived as `projects/<sha256(root)[0..32]>.db` next to the agent database. Untrusted projects and broad roots (home directory, filesystem root) fall back to the agent database.
|
|
19
|
+
|
|
20
|
+
Both files receive the same schema and migrations. The agent database keeps only `token_calibration`, so a sample learned in one project is visible to every other project for the same provider/model/estimator; project databases keep the session index, project index, manifests, summary graph, memory/pin projections, embeddings, quality samples and artifact metadata. Project artifact object bytes move under `projects/artifacts/<project-digest>/` so garbage collection stays scoped to one project. A project database starts empty and is rebuilt from canonical JSONL and project files; the split never rewrites the previous shared database.
|
|
21
|
+
|
|
22
|
+
Calibration and manifest writes are intentionally not one cross-database transaction: a failed calibration insert loses at most one sample, while manifest/usage consistency stays inside the project database. Manifest pruning detaches calibration rows only in `agent` scope; in `project` scope the project database's calibration table stays empty. `storage.databasePath` names the agent database; the derived `projects/` directory is the supported layout.
|
|
23
|
+
|
|
24
|
+
Maintenance and diagnostics are per file: `/context storage` reports the active project database and, when split, the shared agent database; `ds4-context-storage inspect|compact|recover --database <path>` must be pointed at each file.
|
|
25
|
+
|
|
13
26
|
The M19 non-Pi reference adapter owns a separate `ds4-runtime-session-v1` JSONL source selected by its host runtime. Its header binds runtime/session identity and the exact canonical project root; following records contain provenance-checked canonical messages. DS4 snapshots and capability diagnostics are disposable. `createReferenceHistory()` refuses overwrite, append uses a dedicated provenance-checked operation, files are mode `0600` where supported, and rebuild never edits this runtime-owned canonical file. Reference JSONL is not imported into Pi or `context.db`.
|
|
14
27
|
|
|
15
28
|
M20 local KV state is entirely runtime-owned and volatile. Core returns only an in-memory eligibility fingerprint to the runtime port; it has no cache-handle field or serialization API. Prefixes, fingerprints, handles and provider outputs are absent from Pi/reference JSONL, Context Manifests, ranking artifacts and every SQLite table. Aggregate hit/miss/prefill counters live only on the adapter controller, and a restart safely resets them with the runtime cache. M20 adds no database migration.
|
|
@@ -83,7 +96,7 @@ Custom entries have empty lexical search text and never enter Pi context directl
|
|
|
83
96
|
|
|
84
97
|
Schema v10 extends `context_manifests` and `token_calibration` with separate uncached-input, cache-read, and cache-write token columns. New calibration rows also carry the correlated manifest ID and explicit estimator version. Legacy pre-v10 samples migrate as `chars-v1` with their prior total stored as uncached input and zero cache fields; this preserves historical ratio behavior without inventing cache hits.
|
|
85
98
|
|
|
86
|
-
Calibration rows are derived telemetry, isolated by exact provider/model and bounded to the latest configured window at read time. The runtime recomputes median/MAD outlier filtering deterministically; no learned model or mutable provider state is stored. Deleting the database loses calibration and cache history but never session content. Ephemeral sessions and configurations that disable manifest persistence keep only a bounded in-memory window.
|
|
99
|
+
Calibration rows are derived telemetry, isolated by exact provider/model and bounded to the latest configured window at read time. The runtime recomputes median/MAD outlier filtering deterministically; no learned model or mutable provider state is stored. Deleting the database loses calibration and cache history but never session content. Ephemeral sessions and configurations that disable manifest persistence keep only a bounded in-memory window. With `storage.scope: "project"`, calibration lives in the agent database while the manifest that produced the sample lives in the project database; the sample therefore carries no `manifest_id` and the two writes are separate transactions.
|
|
87
100
|
|
|
88
101
|
## Context quality samples
|
|
89
102
|
|
|
@@ -107,7 +120,7 @@ The volatile state is cleared on lifecycle/model/branch/compaction boundaries an
|
|
|
107
120
|
|
|
108
121
|
Schema v8 splits content objects from source references. `artifact_objects` is keyed by SHA-256 and stores the private file path, MIME, byte size, verification timestamps, and integrity status. `artifacts` is keyed by a deterministic source-specific ID and references session/entry/tool identity plus original/condensed token estimates and an optional derived privacy classification in `metadata_json`. Equal bytes across calls or sessions deduplicate to one object while retaining independent provenance.
|
|
109
122
|
|
|
110
|
-
Objects live under `ds4-context/artifacts/<sha-prefix>/<sha256>` with private permissions and atomic writes. Pi's full JSONL tool result remains canonical; the object file is a rebuildable local cache. No artifact content is stored in Context Manifests. Search recomputes SHA-256 and returns only bounded, redacted, JSON-quoted literal-match windows for a current-branch reference. The runtime reapplies the stored artifact classification before returning excerpts to the active provider; prohibited remote searches return no content.
|
|
123
|
+
Objects live under `ds4-context/artifacts/<sha-prefix>/<sha256>` with private permissions and atomic writes. With `storage.scope: "project"` the store root becomes `projects/artifacts/<project-digest>/...` so artifact bytes and their metadata share one boundary; the orphan garbage collector only sees references in the current database and must never delete another project's objects. Pi's full JSONL tool result remains canonical; the object file is a rebuildable local cache. No artifact content is stored in Context Manifests. Search recomputes SHA-256 and returns only bounded, redacted, JSON-quoted literal-match windows for a current-branch reference. The runtime reapplies the stored artifact classification before returning excerpts to the active provider; prohibited remote searches return no content.
|
|
111
124
|
|
|
112
125
|
A full index rebuild replays all message entries, recreates missing qualifying objects, removes stale session references, and garbage-collects object rows/files with no references. Missing/corrupt states are reported by `/context health` without blocking Pi.
|
|
113
126
|
|
|
@@ -117,7 +130,7 @@ For persisted sessions, each `context` hook stores a metadata-only manifest cont
|
|
|
117
130
|
|
|
118
131
|
`before_provider_request` updates the pending in-memory manifest with final-check/redaction counters but never the provider payload. The following finalized assistant response updates only the existing scalar usage columns (`actual_tokens`, `input_tokens`, `cache_read_tokens`, and `cache_write_tokens`) and adds at most one exact-model calibration sample. It does not read or rewrite `manifest_json`. Repository reads hydrate authoritative usage from those columns. Ephemeral, oversize-skipped, concurrently pruned, and otherwise uncorrelated manifests retain bounded calibration only in memory.
|
|
119
132
|
|
|
120
|
-
Retention is bounded without a schema change: SQLite keeps the latest 128 manifests globally and at most 200 calibration samples for each provider/model/estimator profile. A manifest prune first detaches its small calibration row, then removes the large diagnostic JSON; calibration has its own per-profile retention. Save and prune are one transaction. Existing oversized stores are reduced incrementally by at most 32 rows and 8 MiB of serialized manifest payload per subsequent manifest write; one individually oversized oldest row may be removed to guarantee progress. There is no startup purge.
|
|
133
|
+
Retention is bounded without a schema change: SQLite keeps the latest 128 manifests globally and at most 200 calibration samples for each provider/model/estimator profile. A manifest prune first detaches its small calibration row, then removes the large diagnostic JSON; calibration has its own per-profile retention. Save and prune are one transaction. In `project` scope the manifest transaction runs in the project database while the calibration sample is inserted separately into the agent database. Existing oversized stores are reduced incrementally by at most 32 rows and 8 MiB of serialized manifest payload per subsequent manifest write; one individually oversized oldest row may be removed to guarantee progress. There is no startup purge.
|
|
121
134
|
|
|
122
135
|
New manifest persistence is byte-bounded. Payloads up to 256 KiB remain complete. Larger payloads preserve all `included` provenance and replace only the `excluded` inventory with a deterministic first/last sample of at most 256 details plus explicit `ds4-context-manifest-inventory-v1` counts, token/classification/kind rollups, and digests. The wrapper returned by `getStored()` declares `complete` or `excluded-rollup`; the live runtime manifest remains complete. A projected payload over 1 MiB is skipped without affecting the model request. Deleted pages become reusable by SQLite but do not promise an immediate reduction in filesystem size. Manifests and calibration remain disposable; Pi JSONL and project files are untouched.
|
|
123
136
|
|
|
@@ -152,3 +165,25 @@ A full rebuild does not blindly delete unchanged entries. It upserts all observe
|
|
|
152
165
|
Session reconciliation is transactional. Memory/pin mutation replacement, checkpoint update, source exclusion and full materialization each occur under the shared write coordinator. Each manifest upsert and dual-bound incremental retention prune share one transaction; each scalar usage/calibration update and its independent per-profile prune do the same. Each quality upsert and bounded-retention prune also share one transaction; quality failures do not affect manifests or planning. Each changed project file is replaced transactionally with its snippets and FTS rows; embedding upserts and canonical-source pruning are transactional; artifact object/reference metadata and project deletion batches are atomic. A filesystem artifact write precedes its metadata transaction, so an interrupted metadata write may leave only an unreferenced content-addressed cache file; canonical JSONL remains sufficient for recovery. If manifest serialization, projection, retention, or SQLite writing fails, the complete current manifest remains in memory and the provider request is unchanged. Other artifact/project failures contribute no replacement/snippets; planner failures discard all synthetic evidence; Pi continues with its native context.
|
|
153
166
|
|
|
154
167
|
After bounded busy-aware replay is exhausted, DS4 emits `database.write_lock_timeout` with only the coordinator operation name, attempt count, elapsed/configured waits, and SQLite primary code. The thrown error repeats the operation and categorical lock status but never includes SQL, bound values, provider content, or the raw SQLite message. Retry and rollback diagnostics follow the same metadata-only rule.
|
|
168
|
+
|
|
169
|
+
## Growth measurements
|
|
170
|
+
|
|
171
|
+
`tests/benchmarks/storage-scale.bench.ts` seeds session indexes of 5,000 / 50,000 / 200,000 entries and measures the paths a session actually pays. Run it on demand:
|
|
172
|
+
|
|
173
|
+
```bash
|
|
174
|
+
npx vitest bench tests/benchmarks/storage-scale.bench.ts
|
|
175
|
+
```
|
|
176
|
+
|
|
177
|
+
Measured means on this development machine (Node.js 26.5.1, `node:sqlite`, one session per database):
|
|
178
|
+
|
|
179
|
+
| Path | 5k entries | 50k entries | 200k entries |
|
|
180
|
+
| --- | ---: | ---: | ---: |
|
|
181
|
+
| Exact identifier scan (`instr` over one session) | 0.97 ms | 12.5 ms | 52.7 ms |
|
|
182
|
+
| Exact phrase scan | 0.98 ms | 12.8 ms | 53.2 ms |
|
|
183
|
+
| FTS retrieval, common token | 0.14 ms | 2.3 ms | 9.0 ms |
|
|
184
|
+
| FTS retrieval, rare token | 0.06 ms | 0.15 ms | 0.94 ms |
|
|
185
|
+
| Per-session stats (`COUNT`/`SUM` for one session) | 0.33 ms | 4.8 ms | 21.8 ms |
|
|
186
|
+
| Storage diagnostics | 0.10 ms | 0.10 ms | 0.10 ms |
|
|
187
|
+
| Append-only unchanged re-check (1,000 entries) | 1.4 ms | 2.9 ms | 7.2 ms |
|
|
188
|
+
|
|
189
|
+
Growth is **real but bounded**: exact identifier and phrase scans are literal `instr()` scans over the session's rows and grow roughly linearly, passing the 50 ms typical-operation target only around 200k indexed entries in a single session. FTS retrieval and per-session aggregate SQL grow sublinearly and stay in single-digit milliseconds; bounded manifest/calibration diagnostics are flat. The dominant cost tracks a single session's size, not the file size, so `storage.scope: "project"` bounds physical growth, lock scope and reset per project but does not by itself change this per-session scan profile. Long single sessions near or above the 200k-entry range are the case where exact-identifier retrieval latency becomes measurable.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Offline SQLite Storage Maintenance
|
|
2
2
|
|
|
3
|
-
DS4 keeps `context.db` as disposable derived state, while Pi session JSONL and live project files remain canonical. Normal runtime retention stops unbounded manifest growth and makes deleted pages reusable. It does not promise that an existing high-water SQLite file shrinks physically.
|
|
3
|
+
DS4 keeps `context.db` as disposable derived state, while Pi session JSONL and live project files remain canonical. Normal runtime retention stops unbounded manifest growth and makes deleted pages reusable. It does not promise that an existing high-water SQLite file shrinks physically. With the default `storage.scope: "project"` each trusted project has its own database under `ds4-context/projects/` and the agent database keeps shared token calibration: run the maintenance commands per file, not only on the agent database (see [`STORAGE.md`](STORAGE.md)).
|
|
4
4
|
|
|
5
5
|
Physical compaction is therefore an explicit offline operation. It is never model-callable, never runs at startup, and never edits Pi JSONL or project files.
|
|
6
6
|
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
# Release 0.3.10 — Opt-in BPE estimation and bounded model budget auto-tuning
|
|
2
|
+
|
|
3
|
+
**Coordinated packages:** `ds4-context-core`, `ds4-context-reference-adapter`, and `ds4-context-engine` 0.3.10.
|
|
4
|
+
**Implementation commit:** `0607238`.
|
|
5
|
+
|
|
6
|
+
## Summary
|
|
7
|
+
|
|
8
|
+
Adds a selectable `o200k-base-v1` text estimator, provider/model/estimator-specific calibration, and opt-in evidence-gated expansion of automatic context-category ceilings. `chars-v1` remains the default estimator; `modelAwareness.autoTune` remains off unless explicitly enabled. The reference adapter and Pi extension continue to depend exactly on the matching core version.
|
|
9
|
+
|
|
10
|
+
## Changes
|
|
11
|
+
|
|
12
|
+
- Portable core accepts a BPE estimator through its existing runtime-neutral interface; the Pi adapter lazy-loads `js-tiktoken`. Non-text content still uses bounded heuristics.
|
|
13
|
+
- Model profiles can select `tokenEstimator` by exact `provider/model`, provider wildcard, or global override. Calibration histories are isolated by estimator version as well as provider and model.
|
|
14
|
+
- Auto-tuning uses accepted same-profile samples, rejects outliers, respects configured category limits and hard input ceilings, and withdraws expansion after insufficient headroom. Calibrated estimator-unit limits cannot exceed nominal model limits.
|
|
15
|
+
- Manifests and model diagnostics identify estimator version and calibration/auto-tuning decisions. Existing golden compatibility remains anchored to the unchanged default behavior.
|
|
16
|
+
|
|
17
|
+
## Measured scope and limitations
|
|
18
|
+
|
|
19
|
+
Bounded synthetic multi-turn Pi/DS4 sessions on OpenRouter `openai/gpt-6-sol`, `openai/gpt-6-luna`, and `openai/gpt-5.6-terra` demonstrated calibration, a genuinely larger opt-in recent tail, provider usage above the native-window headroom threshold, and withdrawal to `no-headroom` on the following turn in each model's session. Retrieval and live tool cycles were also exercised separately. See [model-awareness measurements](../MODEL_AWARENESS.md) and [provider-token drift](../PROVIDER_TOKEN_DRIFT_BENCHMARK.md) for methodology, numbers, and bounds.
|
|
20
|
+
|
|
21
|
+
These measurements do not establish production retrieval quality, provider-side compaction behavior, or accuracy for arbitrary models/routes. In particular, equivalent Codex-route measurements did not yield usage and are not claimed as validated. No provider calls are required for this release procedure.
|
|
22
|
+
|
|
23
|
+
## Compatibility
|
|
24
|
+
|
|
25
|
+
No new defaults are enabled: `chars-v1`, disabled `autoTune`, and disabled DS4 native continuation remain unchanged. No SQLite migration or change to canonical Pi JSONL, privacy consent, native continuation, or the portable runtime-adapter contract is introduced. BPE remains adapter-injected rather than adding a third-party tokenizer dependency to portable core. Restart Pi after upgrading so the new compiled core and session configuration are loaded.
|
|
26
|
+
|
|
27
|
+
## Validation and publication
|
|
28
|
+
|
|
29
|
+
On Node 26.5.1, the coordinated release passed a clean `npm ci`, `npm run check` (97 Vitest files, 607 tests, TypeScript builds and root typecheck), deterministic `npm run quality:compare`, the `npm run schema:context-persistence` size bound, and `npm run pack:check` in a clean consumer after synchronizing both exported runtime version constants. All typecheck/tests were also rerun after that correction. Dry-run tarball review contained 243 core files, 7 reference-adapter files, and 97 extension files; none included untracked local state. All three packages were then published to npm at 0.3.10 in dependency order. After registry propagation, `npm run registry:check -- 0.3.10` verified all three exact-version artifacts in a fresh consumer.
|
package/docs/releases/0.3.9.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Release 0.3.9 — Bounded compaction request and operation input
|
|
2
2
|
|
|
3
3
|
**Version analyzed:** DS4 Context Engine `0.3.9`
|
|
4
|
-
**Commit:**
|
|
4
|
+
**Commit:** `01407cd`
|
|
5
5
|
**Coordinated packages:** `ds4-context-core` 0.3.9, `ds4-context-reference-adapter` 0.3.9, `ds4-context-engine` 0.3.9
|
|
6
6
|
|
|
7
7
|
## Summary
|
|
@@ -0,0 +1,86 @@
|
|
|
1
|
+
# Release 0.4.0 — Per-project databases with shared token calibration
|
|
2
|
+
|
|
3
|
+
**Coordinated packages:** `ds4-context-core`, `ds4-context-reference-adapter`, and `ds4-context-engine` 0.4.0.
|
|
4
|
+
**Implementation commit:** `2c2e248`.
|
|
5
|
+
**Decision record:** [ADR 064](../ADR/064-per-project-databases-with-shared-calibration.md).
|
|
6
|
+
|
|
7
|
+
## Summary
|
|
8
|
+
|
|
9
|
+
Adds the opt-out `storage.scope: "agent" | "project"` setting and defaults it to
|
|
10
|
+
`project`. Each trusted canonical project root now gets its own SQLite
|
|
11
|
+
projection next to the configured agent database, while token calibration stays
|
|
12
|
+
in the agent database and remains shared across projects. `storage.scope:
|
|
13
|
+
"agent"` restores the previous single-file layout. The reference adapter and Pi
|
|
14
|
+
extension continue to depend exactly on the matching core version.
|
|
15
|
+
|
|
16
|
+
## Changes
|
|
17
|
+
|
|
18
|
+
- `storage.scope: "project"` (new default) derives one database per trusted
|
|
19
|
+
canonical project root at `projects/<sha256(root)[0..32]>.db`, next to
|
|
20
|
+
`storage.databasePath`. Untrusted projects, broad roots (home directory,
|
|
21
|
+
filesystem root) and any resolution failure fall back to the agent database.
|
|
22
|
+
- Both files receive the same schema and the same migrations; migrations 1–15
|
|
23
|
+
are untouched and there is no schema fork. The project database holds the
|
|
24
|
+
session index, project index, context manifests, summary graph, memory/pin
|
|
25
|
+
projections, embeddings, quality samples and artifact metadata; the agent
|
|
26
|
+
database holds `token_calibration`.
|
|
27
|
+
- With the split, a calibration sample is written to the agent database with
|
|
28
|
+
`manifest_id = NULL`, exactly once per manifest. The two writes are
|
|
29
|
+
intentionally not one cross-database transaction: a lost calibration sample is
|
|
30
|
+
harmless, a lost manifest/usage consistency is not. In `agent` scope the
|
|
31
|
+
previous single transaction is unchanged.
|
|
32
|
+
- Artifact object bytes move to `projects/artifacts/<project-digest>/` under
|
|
33
|
+
project scope so the orphan garbage collector, which only sees references in
|
|
34
|
+
the current database, can never delete another project's objects.
|
|
35
|
+
- `/context storage` and `/context diagnostics` report the active project
|
|
36
|
+
database and, when split, the shared agent database. Storage maintenance
|
|
37
|
+
stays per file: `ds4-context-storage inspect|compact|recover --database <path>`
|
|
38
|
+
must be pointed at each database.
|
|
39
|
+
- The previous shared database is left untouched. With the new default, existing
|
|
40
|
+
users cold-start per-project indexes while existing calibration remains
|
|
41
|
+
available in the agent database; project indexes rebuild from canonical Pi
|
|
42
|
+
JSONL, project files and memory/pin `CustomEntry` records.
|
|
43
|
+
|
|
44
|
+
## Breaking behavior
|
|
45
|
+
|
|
46
|
+
The default storage layout changes. A pre-existing
|
|
47
|
+
`~/.pi/agent/ds4-context/context.db` becomes the agent database and is no longer
|
|
48
|
+
the active projection for new sessions. `storage.scope: "agent"` restores the
|
|
49
|
+
old single-database behavior. Rows are not migrated between files; the project
|
|
50
|
+
databases are derived and rebuildable. Decision D1 of the storage plan ("no
|
|
51
|
+
per-session or per-project databases") is superseded by ADR 064 and the plan
|
|
52
|
+
keeps the original text with an explicit amendment pointer.
|
|
53
|
+
|
|
54
|
+
## Measured scope and limitations
|
|
55
|
+
|
|
56
|
+
`tests/benchmarks/storage-scale.bench.ts` seeds one session with 5,000 /
|
|
57
|
+
50,000 / 200,000 entries and measures the paths a session pays. Measured means
|
|
58
|
+
on the development machine (Node.js 26.5.1, `node:sqlite`): exact identifier
|
|
59
|
+
scan 0.97 / 12.5 / 52.7 ms, exact phrase scan 0.98 / 12.8 / 53.2 ms, FTS
|
|
60
|
+
common token 0.14 / 2.3 / 9.0 ms, per-session stats 0.33 / 4.8 / 21.8 ms,
|
|
61
|
+
storage diagnostics flat at 0.10 ms, unchanged re-check 1.4 / 2.9 / 7.2 ms.
|
|
62
|
+
Growth is real but bounded and tracks single-session size, not file size:
|
|
63
|
+
project scope bounds physical growth, lock scope and reset per project, but does
|
|
64
|
+
not by itself change the per-session exact-scan profile. See
|
|
65
|
+
[storage growth measurements](../STORAGE.md#growth-measurements).
|
|
66
|
+
|
|
67
|
+
These are local measurements on one host, not portable guarantees. No provider
|
|
68
|
+
calls are involved in this release procedure.
|
|
69
|
+
|
|
70
|
+
## Compatibility
|
|
71
|
+
|
|
72
|
+
No provider-facing default changes: `chars-v1` remains the default estimator,
|
|
73
|
+
`modelAwareness.autoTune` and DS4 native continuation stay disabled. No SQLite
|
|
74
|
+
migration, no change to canonical Pi JSONL, privacy consent, native continuation
|
|
75
|
+
or the portable runtime-adapter contract. Restart Pi after upgrading so the new
|
|
76
|
+
compiled core and session configuration are loaded.
|
|
77
|
+
|
|
78
|
+
## Validation and publication
|
|
79
|
+
|
|
80
|
+
On Node 26.5.1, the coordinated release passed `npm run check` (99 Vitest files,
|
|
81
|
+
615 tests, TypeScript builds and root typecheck), deterministic
|
|
82
|
+
`npm run quality:compare`, the `npm run schema:context-persistence` size bound,
|
|
83
|
+
and `npm run pack:check` in a clean consumer after synchronizing both exported
|
|
84
|
+
runtime version constants. `git diff --check` was clean and dry-run tarball
|
|
85
|
+
review contained 247 core files, 7 reference-adapter files and 98 extension
|
|
86
|
+
files, none including untracked local state.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "ds4-context-engine",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.4.0",
|
|
4
4
|
"description": "Non-destructive, provider-independent context management for Pi.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|
|
@@ -52,6 +52,7 @@
|
|
|
52
52
|
"quality:compare": "node scripts/compare-context-quality.mjs",
|
|
53
53
|
"schema:context-persistence": "node scripts/measure-context-persistence-schema.mjs",
|
|
54
54
|
"latency:check": "npm run build:core && node scripts/compare-disabled-planning-latency.mjs",
|
|
55
|
+
"storage:scale": "npm run build:core && vitest bench tests/benchmarks/storage-scale.bench.ts",
|
|
55
56
|
"pack:check": "node scripts/verify-packages.mjs",
|
|
56
57
|
"registry:check": "node scripts/verify-registry-packages.mjs",
|
|
57
58
|
"prepare": "npm run build:core && npm run build:adapters"
|
|
@@ -62,7 +63,8 @@
|
|
|
62
63
|
]
|
|
63
64
|
},
|
|
64
65
|
"dependencies": {
|
|
65
|
-
"ds4-context-core": "0.
|
|
66
|
+
"ds4-context-core": "0.4.0",
|
|
67
|
+
"js-tiktoken": "1.0.21"
|
|
66
68
|
},
|
|
67
69
|
"peerDependencies": {
|
|
68
70
|
"@earendil-works/pi-ai": "0.84.3",
|
|
@@ -250,6 +250,9 @@ function formatStatus(diagnostics: RuntimeDiagnostics): string {
|
|
|
250
250
|
`Privacy blocked/redacted: ${count(diagnostics.privacy.blockedBlocks)} / ${count(diagnostics.privacy.secretRedactions)}`,
|
|
251
251
|
`Estimator calibration: ${diagnostics.modelAwareness?.calibration.calibrated ? `x${diagnostics.modelAwareness.calibration.appliedRatio.toFixed(3)}` : "collecting/neutral"}`,
|
|
252
252
|
`Calibration samples: ${count(diagnostics.modelAwareness?.calibration.acceptedSamples)} accepted`,
|
|
253
|
+
...(diagnostics.modelAwareness?.drift ? [
|
|
254
|
+
`Token estimate warning: ${diagnostics.modelAwareness.drift.code} (${diagnostics.modelAwareness.drift.severity}, ${count(diagnostics.modelAwareness.drift.sampleCount)} samples)`,
|
|
255
|
+
] : []),
|
|
253
256
|
`Native continuation: ${diagnostics.nativeContinuation.status} (${diagnostics.nativeContinuation.last?.mode ?? "no request"})`,
|
|
254
257
|
`Continuation saved items: ${count(diagnostics.nativeContinuation.last?.omittedInputItems)}`,
|
|
255
258
|
`Quality metrics: ${diagnostics.quality.enabled ? `${count(diagnostics.quality.storedSamples)} sample(s)` : "disabled"}`,
|
|
@@ -365,6 +368,9 @@ function formatManifest(diagnostics: RuntimeDiagnostics): string {
|
|
|
365
368
|
`Artifacts: ${count(manifest.artifacts?.length ?? 0)}`,
|
|
366
369
|
`Privacy: ${manifest.privacy?.enforcement ?? "disabled"}${manifest.privacy ? ` (${manifest.privacy.destination})` : ""}`,
|
|
367
370
|
`Model calibration: ${manifest.modelAwareness?.calibration.calibrated ? `x${manifest.modelAwareness.calibration.appliedRatio.toFixed(3)}` : "neutral/collecting"}`,
|
|
371
|
+
...(manifest.modelAwareness?.drift ? [
|
|
372
|
+
`Token estimate warning: ${manifest.modelAwareness.drift.code} (${manifest.modelAwareness.drift.severity})`,
|
|
373
|
+
] : []),
|
|
368
374
|
`Adaptive tail/hist/project: ${count(manifest.modelAwareness?.adaptive.recentTailTokens)} / ${count(manifest.modelAwareness?.adaptive.maxRetrievedHistoryTokens)} / ${count(manifest.modelAwareness?.adaptive.maxProjectTokens)}`,
|
|
369
375
|
`Continuation: ${manifest.nativeContinuation?.mode ?? "disabled"}; sent/full ${count(manifest.nativeContinuation?.sentInputItems)} / ${count(manifest.nativeContinuation?.fullInputItems)}`,
|
|
370
376
|
"",
|
|
@@ -664,6 +670,12 @@ function formatModelAwareness(diagnostics: RuntimeDiagnostics): string {
|
|
|
664
670
|
`Samples observed/accepted: ${count(calibration.observedSamples)} / ${count(calibration.acceptedSamples)}`,
|
|
665
671
|
`Samples rejected/outliers: ${count(calibration.rejectedSamples)} / ${count(calibration.outlierSamples)}`,
|
|
666
672
|
`Calibration bounds/window: ${calibration.lowerRatioBound.toFixed(2)}-${calibration.upperRatioBound.toFixed(2)} / ${count(calibration.windowSize)}`,
|
|
673
|
+
...(awareness.drift ? [
|
|
674
|
+
`Token estimate warning: ${awareness.drift.code} (${awareness.drift.severity}, ${count(awareness.drift.sampleCount)} samples${awareness.drift.medianRatio === undefined ? "" : `, x${awareness.drift.medianRatio.toFixed(3)}`})`,
|
|
675
|
+
] : []),
|
|
676
|
+
...(awareness.autoTune ? [
|
|
677
|
+
`Budget auto-tuning: ${awareness.autoTune.status} (${count(awareness.autoTune.acceptedSamples)} samples, x${awareness.autoTune.boostFactor.toFixed(3)})`,
|
|
678
|
+
] : []),
|
|
667
679
|
`Adaptive recent tail: ${count(awareness.adaptive.recentTailTokens)} (nominal ${count(awareness.adaptive.nominalRecentTailTokens)})`,
|
|
668
680
|
`Adaptive history retrieval: ${count(awareness.adaptive.maxRetrievedHistoryTokens)} (nominal ${count(awareness.adaptive.nominalRetrievedHistoryTokens)})`,
|
|
669
681
|
`Adaptive project retrieval: ${count(awareness.adaptive.maxProjectTokens)} (nominal ${count(awareness.adaptive.nominalProjectTokens)})`,
|
|
@@ -823,13 +835,20 @@ function formatAdapter(diagnostics: RuntimeDiagnostics): string {
|
|
|
823
835
|
].join("\n");
|
|
824
836
|
}
|
|
825
837
|
|
|
826
|
-
function formatStorage(
|
|
838
|
+
function formatStorage(
|
|
839
|
+
storage: StorageDiagnostics,
|
|
840
|
+
databasePath?: string,
|
|
841
|
+
agentDatabasePath?: string,
|
|
842
|
+
): string {
|
|
843
|
+
const agentLine = agentDatabasePath && agentDatabasePath !== databasePath
|
|
844
|
+
? [`Agent database (shared): ${agentDatabasePath}`] : [];
|
|
827
845
|
if (storage.status === "unavailable") {
|
|
828
846
|
return [
|
|
829
847
|
"DS4 Storage",
|
|
830
848
|
"",
|
|
831
849
|
"Status: unavailable",
|
|
832
850
|
`Database: ${databasePath ?? "unavailable"}`,
|
|
851
|
+
...agentLine,
|
|
833
852
|
"Pi fallback remains active; no storage mutation was attempted.",
|
|
834
853
|
].join("\n");
|
|
835
854
|
}
|
|
@@ -838,6 +857,7 @@ function formatStorage(storage: StorageDiagnostics, databasePath?: string): stri
|
|
|
838
857
|
"",
|
|
839
858
|
`Status: ${storage.status}`,
|
|
840
859
|
`Database: ${databasePath ?? "unavailable"}`,
|
|
860
|
+
...agentLine,
|
|
841
861
|
`Schema / journal: ${storage.schemaVersion ?? "n/a"} / ${storage.journalMode ?? "n/a"}`,
|
|
842
862
|
`Database / WAL / SHM: ${bytes(storage.databaseBytes)} / ${bytes(storage.walBytes)} / ${bytes(storage.shmBytes)}`,
|
|
843
863
|
`Allocated / reusable: ${bytes(storage.allocatedBytes)} / ${bytes(storage.reusableBytes)}`,
|
|
@@ -1222,7 +1242,7 @@ export function registerContextCommand(pi: ExtensionAPI, runtime: Ds4ContextRunt
|
|
|
1222
1242
|
const storage = runtime.storageDiagnostics();
|
|
1223
1243
|
present(
|
|
1224
1244
|
ctx,
|
|
1225
|
-
formatStorage(storage, diagnostics.databasePath),
|
|
1245
|
+
formatStorage(storage, diagnostics.databasePath, diagnostics.agentDatabasePath),
|
|
1226
1246
|
storage.status === "ok" ? "info" : "warning",
|
|
1227
1247
|
);
|
|
1228
1248
|
return;
|
package/src/extension/runtime.ts
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
import { randomUUID } from "node:crypto";
|
|
2
2
|
import { existsSync, mkdirSync, readFileSync, realpathSync, writeFileSync } from "node:fs";
|
|
3
3
|
import { homedir } from "node:os";
|
|
4
|
-
import { dirname, join, parse, resolve } from "node:path";
|
|
4
|
+
import { basename, dirname, join, parse, resolve } from "node:path";
|
|
5
5
|
import type {
|
|
6
6
|
Api,
|
|
7
7
|
AssistantMessage,
|
|
@@ -46,11 +46,12 @@ export type {
|
|
|
46
46
|
import {
|
|
47
47
|
loadConfig,
|
|
48
48
|
resolveDatabasePath,
|
|
49
|
+
resolveProjectDatabasePath as deriveProjectDatabasePath,
|
|
49
50
|
resolveRankingModelPath,
|
|
50
51
|
validateConfigFile,
|
|
51
52
|
type LoadedConfig,
|
|
52
53
|
} from "ds4-context-core/config/config-loader";
|
|
53
|
-
import { CONFIG_SCHEMA_VERSION, createDefaultConfig, type Ds4ContextConfig } from "ds4-context-core/config/config";
|
|
54
|
+
import { CONFIG_SCHEMA_VERSION, createDefaultConfig, type Ds4ContextConfig, type StorageScope } from "ds4-context-core/config/config";
|
|
54
55
|
import {
|
|
55
56
|
applyConfigValue,
|
|
56
57
|
findConfigField,
|
|
@@ -60,11 +61,12 @@ import { calculateContextBudget, type ContextBudget } from "ds4-context-core/cor
|
|
|
60
61
|
import {
|
|
61
62
|
modelProfileKey,
|
|
62
63
|
resolveModelAwareness,
|
|
64
|
+
tokenEstimatorVersion,
|
|
63
65
|
type ResolvedModelAwareness,
|
|
64
66
|
type TokenCalibrationSample,
|
|
65
67
|
} from "ds4-context-core/core/model-awareness";
|
|
66
68
|
import type { ModelDescriptor } from "ds4-context-core/core/model-profile";
|
|
67
|
-
import {
|
|
69
|
+
import { CHARS_ESTIMATOR, type TokenEstimator } from "ds4-context-core/core/token-estimator";
|
|
68
70
|
import type {
|
|
69
71
|
ContextManifest,
|
|
70
72
|
ModelAwarenessManifest,
|
|
@@ -164,6 +166,7 @@ import {
|
|
|
164
166
|
findPiSourceEntryIds,
|
|
165
167
|
fingerprint,
|
|
166
168
|
} from "../pi-adapter/context-observer.ts";
|
|
169
|
+
import { O200K_ESTIMATOR } from "../pi-adapter/bpe-token-estimator.ts";
|
|
167
170
|
import { projectSessionFileMutations } from "../pi-adapter/memory-adapter.ts";
|
|
168
171
|
import { ProjectMemorySynchronizer } from "../pi-adapter/project-memory-sync.ts";
|
|
169
172
|
import { projectRankingLabels } from "../pi-adapter/ranking-adapter.ts";
|
|
@@ -385,6 +388,8 @@ export interface RuntimeDiagnostics {
|
|
|
385
388
|
session?: PiSessionSnapshot;
|
|
386
389
|
model?: { provider: string; id: string };
|
|
387
390
|
databasePath?: string;
|
|
391
|
+
/** Present only when storage.scope splits calibration from project state. */
|
|
392
|
+
agentDatabasePath?: string;
|
|
388
393
|
databaseSchemaVersion?: number;
|
|
389
394
|
indexed?: SessionIndexStats;
|
|
390
395
|
observation?: ContextObservation;
|
|
@@ -424,8 +429,10 @@ export class Ds4ContextRuntime {
|
|
|
424
429
|
private loadedConfig?: LoadedConfig;
|
|
425
430
|
private session?: PiSessionSnapshot;
|
|
426
431
|
private database?: ContextDatabase;
|
|
432
|
+
private agentDatabase?: ContextDatabase;
|
|
427
433
|
private indexer?: PiSessionIndexer;
|
|
428
434
|
private databasePath?: string;
|
|
435
|
+
private agentDatabasePath?: string;
|
|
429
436
|
private observation?: ContextObservation;
|
|
430
437
|
private lastManifest?: ContextManifest;
|
|
431
438
|
private lastPersistedInventory?: PersistedManifestInventory;
|
|
@@ -577,17 +584,33 @@ export class Ds4ContextRuntime {
|
|
|
577
584
|
return;
|
|
578
585
|
}
|
|
579
586
|
|
|
580
|
-
|
|
587
|
+
const agentDatabasePath = resolveDatabasePath(
|
|
581
588
|
this.config.storage.databasePath,
|
|
582
589
|
this.dependencies.agentDir,
|
|
583
590
|
this.dependencies.homeDir,
|
|
584
591
|
);
|
|
592
|
+
this.agentDatabasePath = agentDatabasePath;
|
|
593
|
+
const projectDatabasePath = this.resolveProjectDatabasePath(
|
|
594
|
+
this.config.storage.scope,
|
|
595
|
+
ctx,
|
|
596
|
+
agentDatabasePath,
|
|
597
|
+
);
|
|
598
|
+
this.databasePath = projectDatabasePath ?? agentDatabasePath;
|
|
599
|
+
if (projectDatabasePath) {
|
|
600
|
+
this.agentDatabase = ContextDatabase.open(agentDatabasePath, {
|
|
601
|
+
logger: this.logger,
|
|
602
|
+
now: this.now(),
|
|
603
|
+
busyTimeoutMs: this.config.storage.busyTimeoutMs,
|
|
604
|
+
writeRetryTimeoutMs: this.config.storage.writeRetryTimeoutMs,
|
|
605
|
+
});
|
|
606
|
+
}
|
|
585
607
|
this.database = ContextDatabase.open(this.databasePath, {
|
|
586
608
|
logger: this.logger,
|
|
587
609
|
now: this.now(),
|
|
588
610
|
busyTimeoutMs: this.config.storage.busyTimeoutMs,
|
|
589
611
|
writeRetryTimeoutMs: this.config.storage.writeRetryTimeoutMs,
|
|
590
612
|
});
|
|
613
|
+
this.agentDatabase ??= this.database;
|
|
591
614
|
this.initializeRanking();
|
|
592
615
|
|
|
593
616
|
this.indexer = new PiSessionIndexer(this.database.sessionIndex, {
|
|
@@ -758,15 +781,37 @@ export class Ds4ContextRuntime {
|
|
|
758
781
|
};
|
|
759
782
|
}
|
|
760
783
|
|
|
784
|
+
private estimatorForModel(model: ModelDescriptor): TokenEstimator {
|
|
785
|
+
return tokenEstimatorVersion(model, this.config.modelAwareness) === "o200k-base-v1"
|
|
786
|
+
? O200K_ESTIMATOR : CHARS_ESTIMATOR;
|
|
787
|
+
}
|
|
788
|
+
|
|
789
|
+
private resolveProjectDatabasePath(
|
|
790
|
+
scope: StorageScope,
|
|
791
|
+
ctx: ExtensionContext,
|
|
792
|
+
agentDatabasePath: string,
|
|
793
|
+
): string | undefined {
|
|
794
|
+
if (scope !== "project") return undefined;
|
|
795
|
+
// Untrusted projects ignore project configuration; they also fall back to
|
|
796
|
+
// the shared agent database instead of creating a new physical boundary.
|
|
797
|
+
if (!ctx.isProjectTrusted()) return undefined;
|
|
798
|
+
const projectRoot = canonicalProjectPath(ctx.cwd);
|
|
799
|
+
if (isBroadProjectRoot(projectRoot, this.dependencies.homeDir ?? homedir())) return undefined;
|
|
800
|
+
const projectDatabasePath = deriveProjectDatabasePath(agentDatabasePath, projectRoot);
|
|
801
|
+
return projectDatabasePath === agentDatabasePath ? undefined : projectDatabasePath;
|
|
802
|
+
}
|
|
803
|
+
|
|
761
804
|
private calibrationSamples(model: ModelDescriptor): TokenCalibrationSample[] {
|
|
805
|
+
const version = this.estimatorForModel(model).version;
|
|
762
806
|
if (this.database && this.session?.sessionFile && this.config.diagnostics.storeContextManifest) {
|
|
763
|
-
return this.database.
|
|
807
|
+
return (this.agentDatabase ?? this.database).calibrations.list(
|
|
764
808
|
model.provider,
|
|
765
809
|
model.id,
|
|
766
810
|
this.config.modelAwareness.calibrationWindow,
|
|
811
|
+
version,
|
|
767
812
|
);
|
|
768
813
|
}
|
|
769
|
-
return [...(this.volatileCalibration.get(modelProfileKey(model.provider, model.id)) ?? [])];
|
|
814
|
+
return [...(this.volatileCalibration.get(`${modelProfileKey(model.provider, model.id)}\0${version}`) ?? [])];
|
|
770
815
|
}
|
|
771
816
|
|
|
772
817
|
private switchForModel(model: ModelDescriptor): ModelSwitchManifest {
|
|
@@ -831,6 +876,7 @@ export class Ds4ContextRuntime {
|
|
|
831
876
|
requestsPerTurn: number,
|
|
832
877
|
turnsPerEpoch: number,
|
|
833
878
|
modelKey: string,
|
|
879
|
+
estimator: TokenEstimator,
|
|
834
880
|
): number | undefined {
|
|
835
881
|
const pricing = Ds4ContextRuntime.cachePricing(cost);
|
|
836
882
|
if (!pricing || plan.mode !== "managed") return undefined;
|
|
@@ -839,7 +885,7 @@ export class Ds4ContextRuntime {
|
|
|
839
885
|
let total = fixedTokens;
|
|
840
886
|
for (const message of plan.messages) {
|
|
841
887
|
hashes.push(fingerprint(message));
|
|
842
|
-
const estimate = estimateMessagesTokens([message]);
|
|
888
|
+
const estimate = estimator.estimateMessagesTokens([message]);
|
|
843
889
|
tokens.push(estimate);
|
|
844
890
|
total += estimate;
|
|
845
891
|
}
|
|
@@ -908,6 +954,8 @@ export class Ds4ContextRuntime {
|
|
|
908
954
|
...awareness.calibration,
|
|
909
955
|
cache: { ...awareness.calibration.cache },
|
|
910
956
|
},
|
|
957
|
+
...(awareness.drift ? { drift: awareness.drift } : {}),
|
|
958
|
+
...(awareness.autoTune ? { autoTune: awareness.autoTune } : {}),
|
|
911
959
|
adaptive: { ...awareness.limits },
|
|
912
960
|
switch: modelSwitch,
|
|
913
961
|
};
|
|
@@ -922,7 +970,7 @@ export class Ds4ContextRuntime {
|
|
|
922
970
|
usage: ProviderUsageManifest,
|
|
923
971
|
createdAt: number,
|
|
924
972
|
): void {
|
|
925
|
-
const key = modelProfileKey(manifest.provider, manifest.model)
|
|
973
|
+
const key = `${modelProfileKey(manifest.provider, manifest.model)}\0${manifest.modelAwareness?.calibration.estimator ?? "chars-v1"}`;
|
|
926
974
|
const samples = this.volatileCalibration.get(key) ?? [];
|
|
927
975
|
samples.unshift({
|
|
928
976
|
estimatedTokens: manifest.estimatedInputTokens,
|
|
@@ -980,6 +1028,7 @@ export class Ds4ContextRuntime {
|
|
|
980
1028
|
}
|
|
981
1029
|
const model = snapshotModel(ctx);
|
|
982
1030
|
const activeModel = model ? this.resolveActiveModel(model) : undefined;
|
|
1031
|
+
const estimator = model ? this.estimatorForModel(model) : CHARS_ESTIMATOR;
|
|
983
1032
|
const budget = activeModel?.budget;
|
|
984
1033
|
let effectiveEvent = preparedPrivacy.event;
|
|
985
1034
|
let artifactReferences = [] as NonNullable<ContextManifest["artifacts"]>;
|
|
@@ -993,8 +1042,8 @@ export class Ds4ContextRuntime {
|
|
|
993
1042
|
preparedPrivacy.messageClassifications,
|
|
994
1043
|
this.config.artifacts.adaptiveBudget && budget ? {
|
|
995
1044
|
inputTokens: budget.activeInputBudget,
|
|
996
|
-
fixedTokens: estimateTextTokens(preparedPrivacy.systemPrompt) + 8
|
|
997
|
-
+ preparedPrivacy.tools.reduce((sum, tool) => sum + estimateObservedToolTokens(tool), 0),
|
|
1045
|
+
fixedTokens: estimator.estimateTextTokens(preparedPrivacy.systemPrompt) + 8
|
|
1046
|
+
+ preparedPrivacy.tools.reduce((sum, tool) => sum + estimateObservedToolTokens(tool, estimator), 0),
|
|
998
1047
|
} : undefined,
|
|
999
1048
|
);
|
|
1000
1049
|
effectiveEvent = { type: "context", messages: transformed.messages };
|
|
@@ -1043,6 +1092,7 @@ export class Ds4ContextRuntime {
|
|
|
1043
1092
|
createdAt: observedAt,
|
|
1044
1093
|
policyVersion: POLICY_VERSION,
|
|
1045
1094
|
plannerVersion: OBSERVER_PLANNER_VERSION,
|
|
1095
|
+
tokenEstimator: estimator,
|
|
1046
1096
|
...(activeModel ? {
|
|
1047
1097
|
profile: activeModel.awareness.profile,
|
|
1048
1098
|
budget: activeModel.budget,
|
|
@@ -1107,6 +1157,7 @@ export class Ds4ContextRuntime {
|
|
|
1107
1157
|
fixedTokens,
|
|
1108
1158
|
budget,
|
|
1109
1159
|
config: effectiveContextConfig,
|
|
1160
|
+
tokenEstimator: estimator,
|
|
1110
1161
|
pinnedMessageIndices,
|
|
1111
1162
|
supplementalMessages: dedupSupplementalMessages,
|
|
1112
1163
|
});
|
|
@@ -1138,13 +1189,14 @@ export class Ds4ContextRuntime {
|
|
|
1138
1189
|
fixedTokens,
|
|
1139
1190
|
budget,
|
|
1140
1191
|
config: effectiveContextConfig,
|
|
1192
|
+
tokenEstimator: estimator,
|
|
1141
1193
|
pinnedMessageIndices,
|
|
1142
1194
|
supplementalMessages: dedupSupplementalMessages,
|
|
1143
1195
|
cacheAwareTailTokens: decision.recentTailTokens,
|
|
1144
1196
|
});
|
|
1145
1197
|
const modelKey = modelProfileKey(model.provider, model.id);
|
|
1146
|
-
const nominalEpoch = this.cacheAwareEpochCost(nominalDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey);
|
|
1147
|
-
const extendedEpoch = this.cacheAwareEpochCost(extendedDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey);
|
|
1198
|
+
const nominalEpoch = this.cacheAwareEpochCost(nominalDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey, estimator);
|
|
1199
|
+
const extendedEpoch = this.cacheAwareEpochCost(extendedDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey, estimator);
|
|
1148
1200
|
this.logger.debug("context.cache_aware_candidate", {
|
|
1149
1201
|
eligible: decision.eligible,
|
|
1150
1202
|
tailExtended: decision.tailExtended,
|
|
@@ -1181,6 +1233,7 @@ export class Ds4ContextRuntime {
|
|
|
1181
1233
|
fixedTokens,
|
|
1182
1234
|
budget,
|
|
1183
1235
|
config: effectiveContextConfig,
|
|
1236
|
+
tokenEstimator: estimator,
|
|
1184
1237
|
pinnedMessageIndices,
|
|
1185
1238
|
supplementalMessages: dedupSupplementalMessages,
|
|
1186
1239
|
cacheAwareTailTokens,
|
|
@@ -1275,7 +1328,7 @@ export class Ds4ContextRuntime {
|
|
|
1275
1328
|
if (sanitized.blockedBlocks > 0) {
|
|
1276
1329
|
privacyExcludedSources.push({
|
|
1277
1330
|
sourceId: supplement.sourceIds[0],
|
|
1278
|
-
tokens: estimateMessagesTokens([supplement.message]),
|
|
1331
|
+
tokens: estimator.estimateMessagesTokens([supplement.message]),
|
|
1279
1332
|
kind: supplement.kind,
|
|
1280
1333
|
classification: sanitized.classification,
|
|
1281
1334
|
score: supplement.score,
|
|
@@ -1328,6 +1381,7 @@ export class Ds4ContextRuntime {
|
|
|
1328
1381
|
fixedTokens,
|
|
1329
1382
|
budget,
|
|
1330
1383
|
config: effectiveContextConfig,
|
|
1384
|
+
tokenEstimator: estimator,
|
|
1331
1385
|
pinnedMessageIndices,
|
|
1332
1386
|
supplementalMessages: rankedSupplementalMessages,
|
|
1333
1387
|
...(cacheAwareTailTokens !== undefined ? { cacheAwareTailTokens } : {}),
|
|
@@ -1408,7 +1462,7 @@ export class Ds4ContextRuntime {
|
|
|
1408
1462
|
const currentTokens: number[] = [];
|
|
1409
1463
|
for (const message of plan.messages) {
|
|
1410
1464
|
currentHashes.push(fingerprint(message));
|
|
1411
|
-
currentTokens.push(estimateMessagesTokens([message]));
|
|
1465
|
+
currentTokens.push(estimator.estimateMessagesTokens([message]));
|
|
1412
1466
|
}
|
|
1413
1467
|
const reusablePrefixTokens = estimateReusablePrefixTokens(
|
|
1414
1468
|
previousHashes,
|
|
@@ -1495,6 +1549,7 @@ export class Ds4ContextRuntime {
|
|
|
1495
1549
|
createdAt: observedAt,
|
|
1496
1550
|
policyVersion: POLICY_VERSION,
|
|
1497
1551
|
plannerVersion: PLANNER_VERSION,
|
|
1552
|
+
tokenEstimator: estimator,
|
|
1498
1553
|
...(activeModel ? {
|
|
1499
1554
|
profile: activeModel.awareness.profile,
|
|
1500
1555
|
budget: activeModel.budget,
|
|
@@ -1542,7 +1597,7 @@ export class Ds4ContextRuntime {
|
|
|
1542
1597
|
estimatedMessageTokens: manifest.composition.messageTokens,
|
|
1543
1598
|
originalMessageCount: manifest.planning?.originalMessageCount ?? historyEvent.messages.length,
|
|
1544
1599
|
originalEstimatedMessageTokens: manifest.planning?.originalMessageTokens
|
|
1545
|
-
?? estimateMessagesTokens(historyEvent.messages),
|
|
1600
|
+
?? estimator.estimateMessagesTokens(historyEvent.messages),
|
|
1546
1601
|
...(usage?.tokens !== null && usage?.tokens !== undefined ? { reportedTokens: usage.tokens } : {}),
|
|
1547
1602
|
...(manifest.planning?.durationMs !== undefined
|
|
1548
1603
|
? { planningDurationMs: manifest.planning.durationMs }
|
|
@@ -1865,13 +1920,39 @@ export class Ds4ContextRuntime {
|
|
|
1865
1920
|
const createdAt = this.now();
|
|
1866
1921
|
try {
|
|
1867
1922
|
if (manifestPersisted && this.database) {
|
|
1868
|
-
const
|
|
1923
|
+
const estimatorVersion = manifest?.modelAwareness?.calibration.estimator ?? "chars-v1";
|
|
1924
|
+
// With storage.scope=project the manifest lives in the project database
|
|
1925
|
+
// while calibration stays shared in the agent database. The two writes
|
|
1926
|
+
// are intentionally not one transaction: losing one sample is harmless,
|
|
1927
|
+
// losing manifest/usage consistency is not.
|
|
1928
|
+
const calibrationDatabase = this.agentDatabase !== this.database
|
|
1929
|
+
? this.agentDatabase : undefined;
|
|
1930
|
+
const updated = this.database.manifests.recordProviderUsageOutcome(
|
|
1869
1931
|
manifestId,
|
|
1870
1932
|
providerUsage,
|
|
1871
1933
|
createdAt,
|
|
1872
|
-
|
|
1934
|
+
estimatorVersion,
|
|
1935
|
+
{ writeCalibration: calibrationDatabase === undefined },
|
|
1873
1936
|
);
|
|
1874
|
-
if (
|
|
1937
|
+
if (updated.outcome === "recorded" && calibrationDatabase) {
|
|
1938
|
+
const source = this.database.manifests.calibrationSource(manifestId);
|
|
1939
|
+
const recorded = source
|
|
1940
|
+
? calibrationDatabase.calibrations.record({
|
|
1941
|
+
provider: source.provider,
|
|
1942
|
+
model: source.model,
|
|
1943
|
+
estimatedTokens: source.estimatedTokens,
|
|
1944
|
+
actualInputTokens: providerUsage.totalInputTokens,
|
|
1945
|
+
inputTokens: providerUsage.inputTokens,
|
|
1946
|
+
cacheReadTokens: providerUsage.cacheReadTokens,
|
|
1947
|
+
cacheWriteTokens: providerUsage.cacheWriteTokens,
|
|
1948
|
+
createdAt,
|
|
1949
|
+
estimatorVersion,
|
|
1950
|
+
})
|
|
1951
|
+
: false;
|
|
1952
|
+
if (!recorded && manifest?.estimatedInputTokens) {
|
|
1953
|
+
this.rememberVolatileCalibration(manifest, providerUsage, createdAt);
|
|
1954
|
+
}
|
|
1955
|
+
} else if (!updated.manifest && manifest?.estimatedInputTokens) {
|
|
1875
1956
|
this.rememberVolatileCalibration(manifest, providerUsage, createdAt);
|
|
1876
1957
|
}
|
|
1877
1958
|
} else if (manifest?.estimatedInputTokens) {
|
|
@@ -2688,10 +2769,13 @@ export class Ds4ContextRuntime {
|
|
|
2688
2769
|
return;
|
|
2689
2770
|
}
|
|
2690
2771
|
try {
|
|
2691
|
-
|
|
2692
|
-
|
|
2693
|
-
|
|
2694
|
-
|
|
2772
|
+
// Artifact bytes and their metadata must share one boundary: the orphan
|
|
2773
|
+
// garbage collector only sees references in the current database, so a
|
|
2774
|
+
// shared store would let one project delete another project's objects.
|
|
2775
|
+
const storeRoot = this.agentDatabase !== this.database && this.databasePath
|
|
2776
|
+
? join(dirname(this.databasePath), "artifacts", basename(this.databasePath, ".db"))
|
|
2777
|
+
: join(this.dependencies.agentDir, "ds4-context", "artifacts");
|
|
2778
|
+
const store = new FileArtifactStore(storeRoot, this.now);
|
|
2695
2779
|
this.artifactManager = new ArtifactManager(
|
|
2696
2780
|
store,
|
|
2697
2781
|
this.database.artifacts,
|
|
@@ -3276,6 +3360,8 @@ export class Ds4ContextRuntime {
|
|
|
3276
3360
|
session: currentSession,
|
|
3277
3361
|
...(ctx.model ? { model: { provider: ctx.model.provider, id: ctx.model.id } } : {}),
|
|
3278
3362
|
...(this.databasePath ? { databasePath: this.databasePath } : {}),
|
|
3363
|
+
...(this.agentDatabasePath && this.agentDatabasePath !== this.databasePath
|
|
3364
|
+
? { agentDatabasePath: this.agentDatabasePath } : {}),
|
|
3279
3365
|
...(this.database ? { databaseSchemaVersion: this.database.schemaVersion } : {}),
|
|
3280
3366
|
...(indexed ? { indexed } : {}),
|
|
3281
3367
|
...(this.observation ? { observation: this.observation } : {}),
|
|
@@ -3313,7 +3399,9 @@ export class Ds4ContextRuntime {
|
|
|
3313
3399
|
}
|
|
3314
3400
|
|
|
3315
3401
|
storageDiagnostics(): StorageDiagnostics {
|
|
3316
|
-
|
|
3402
|
+
const calibrationDatabase = this.agentDatabase && this.agentDatabase !== this.database
|
|
3403
|
+
? this.agentDatabase : undefined;
|
|
3404
|
+
return this.database?.storageDiagnostics(this.session?.projectPath, calibrationDatabase)
|
|
3317
3405
|
?? unavailableStorageDiagnostics();
|
|
3318
3406
|
}
|
|
3319
3407
|
|
|
@@ -3440,17 +3528,25 @@ export class Ds4ContextRuntime {
|
|
|
3440
3528
|
try {
|
|
3441
3529
|
this.database?.close();
|
|
3442
3530
|
} finally {
|
|
3443
|
-
|
|
3444
|
-
|
|
3445
|
-
|
|
3446
|
-
|
|
3447
|
-
|
|
3448
|
-
|
|
3449
|
-
|
|
3450
|
-
|
|
3451
|
-
|
|
3452
|
-
|
|
3453
|
-
|
|
3531
|
+
try {
|
|
3532
|
+
if (this.agentDatabase && this.agentDatabase !== this.database) {
|
|
3533
|
+
this.agentDatabase.close();
|
|
3534
|
+
}
|
|
3535
|
+
} finally {
|
|
3536
|
+
this.agentDatabase = undefined;
|
|
3537
|
+
this.agentDatabasePath = undefined;
|
|
3538
|
+
this.compaction = undefined;
|
|
3539
|
+
this.retrievalEngine = undefined;
|
|
3540
|
+
this.projectKnowledge = undefined;
|
|
3541
|
+
this.projectRefreshPending = false;
|
|
3542
|
+
this.memoryManager = undefined;
|
|
3543
|
+
this.projectMemorySynchronizer = undefined;
|
|
3544
|
+
this.lastCrossSessionMemory = disabledCrossSessionMemoryDiagnostics();
|
|
3545
|
+
this.lastMemoryMutationSignature = undefined;
|
|
3546
|
+
this.artifactManager = undefined;
|
|
3547
|
+
this.indexer = undefined;
|
|
3548
|
+
this.database = undefined;
|
|
3549
|
+
}
|
|
3454
3550
|
}
|
|
3455
3551
|
}
|
|
3456
3552
|
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
import { createRequire } from "node:module";
|
|
2
|
+
import type { Tiktoken } from "js-tiktoken/lite";
|
|
3
|
+
import { createO200kEstimator } from "ds4-context-core/core/bpe-token-estimator";
|
|
4
|
+
|
|
5
|
+
const load = createRequire(import.meta.url);
|
|
6
|
+
|
|
7
|
+
let encoder: Tiktoken | undefined;
|
|
8
|
+
|
|
9
|
+
/**
|
|
10
|
+
* Opt-in OpenAI o200k text estimator. Model-specific serializers, image tokens,
|
|
11
|
+
* reasoning, and provider-specific wrappers are still estimated, not exact.
|
|
12
|
+
* No remote vocabulary download or provider request is performed.
|
|
13
|
+
*/
|
|
14
|
+
export const O200K_ESTIMATOR = createO200kEstimator((text) => {
|
|
15
|
+
if (!encoder) {
|
|
16
|
+
// CJS exports are loaded on first opt-in use, not for every Pi session.
|
|
17
|
+
const { Tiktoken: Encoder } = load("js-tiktoken/lite") as typeof import("js-tiktoken/lite");
|
|
18
|
+
const ranks = load("js-tiktoken/ranks/o200k_base") as {
|
|
19
|
+
pat_str: string; special_tokens: Record<string, number>; bpe_ranks: string;
|
|
20
|
+
};
|
|
21
|
+
encoder = new Encoder(ranks);
|
|
22
|
+
}
|
|
23
|
+
return encoder.encode(text).length;
|
|
24
|
+
});
|
|
@@ -8,7 +8,7 @@ import {
|
|
|
8
8
|
import type { ContextConfig } from "ds4-context-core/config/config";
|
|
9
9
|
import { calculateContextBudget, type ContextBudget } from "ds4-context-core/core/budget-manager";
|
|
10
10
|
import { createModelProfile, type ModelProfile } from "ds4-context-core/core/model-profile";
|
|
11
|
-
import { estimateMessageTokens } from "ds4-context-core/core/token-estimator";
|
|
11
|
+
import { estimateMessageTokens, type TokenEstimator } from "ds4-context-core/core/token-estimator";
|
|
12
12
|
import type {
|
|
13
13
|
ArtifactManifestRef,
|
|
14
14
|
ContextManifest,
|
|
@@ -54,6 +54,7 @@ export interface BuildPiObserverManifestOptions {
|
|
|
54
54
|
profile?: ModelProfile;
|
|
55
55
|
budget?: ContextBudget;
|
|
56
56
|
modelAwareness?: ModelAwarenessManifest;
|
|
57
|
+
tokenEstimator?: TokenEstimator;
|
|
57
58
|
plan?: ManagedContextPlan<PiAgentMessage>;
|
|
58
59
|
projectRevision?: ProjectRevision;
|
|
59
60
|
pins?: readonly PinManifestRef[];
|
|
@@ -386,6 +387,7 @@ export function buildPiObserverManifest(options: BuildPiObserverManifestOptions)
|
|
|
386
387
|
profile,
|
|
387
388
|
budget,
|
|
388
389
|
systemPrompt: options.systemPrompt ?? options.ctx.getSystemPrompt(),
|
|
390
|
+
...(options.tokenEstimator ? { tokenEstimator: options.tokenEstimator } : {}),
|
|
389
391
|
...(options.systemClassification ? { systemClassification: options.systemClassification } : {}),
|
|
390
392
|
...(options.systemPrivacyReason ? { systemPrivacyReason: options.systemPrivacyReason } : {}),
|
|
391
393
|
tools: options.tools ?? activeTools(options.pi),
|