ds4-context-engine 0.3.9 → 0.3.10
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +3 -2
- package/docs/CONTEXT_MANIFEST.md +1 -1
- package/docs/MODEL_AWARENESS.md +73 -4
- package/docs/PROVIDER_TOKEN_DRIFT_BENCHMARK.md +85 -0
- package/docs/releases/0.3.10.md +29 -0
- package/docs/releases/0.3.9.md +1 -1
- package/package.json +3 -2
- package/src/extension/commands.ts +12 -0
- package/src/extension/runtime.ts +30 -11
- package/src/pi-adapter/bpe-token-estimator.ts +24 -0
- package/src/pi-adapter/context-observer.ts +3 -1
- package/src/pi-adapter/version.ts +1 -1
package/README.md
CHANGED
|
@@ -16,7 +16,7 @@ bounded active context with provenance
|
|
|
16
16
|
Pi provider
|
|
17
17
|
```
|
|
18
18
|
|
|
19
|
-
> **Project status:** The coordinated `0.3.
|
|
19
|
+
> **Project status:** The coordinated `0.3.10` release adds opt-in BPE token estimation, estimator-specific provider calibration and bounded auto-tuning of context category ceilings. `chars-v1` and disabled auto-tuning remain the defaults. The bounded compaction controls from `0.3.9` remain in place; canonical history, SQLite schema 16 and runtime contracts are unchanged. Pi remains pinned to `0.84.3`. See the [0.3.10 release record](docs/releases/0.3.10.md) and [model-awareness validation](docs/MODEL_AWARENESS.md).
|
|
20
20
|
|
|
21
21
|
**Current compaction defaults:** `compaction.directUpdate=true`, `compaction.inputBudget="context"`, `compaction.segmentTargetTokens=30000`, `compaction.maxRequestInputTokens=64000`, `compaction.maxOperationInputTokens=2000000`, `compaction.maxConcurrentSegments=2`. Every DS4 provider attempt is bounded by the effective request limit, and the operation limit includes retries; `inputBudget="summary"` remains an explicit throughput-oriented opt-in. Existing compaction/master switches still apply. See [latency controls and compatibility](docs/COMPACTION.md#latency-controls). No real-provider speedup is claimed from mock tests. The five optional editing/reading/artifact/job features introduced in `0.3.4` remain default-off.
|
|
22
22
|
|
|
@@ -514,6 +514,7 @@ scripts package and release-readiness checks
|
|
|
514
514
|
- [Roadmap 0.2.0](docs/ROADMAP_0.2.0.md)
|
|
515
515
|
- [Release process](docs/RELEASING.md)
|
|
516
516
|
- [0.2.0 release readiness](docs/RELEASE_READINESS_0.2.0.md)
|
|
517
|
+
- [0.3.10 release notes](docs/releases/0.3.10.md)
|
|
517
518
|
- [0.3.9 release notes](docs/releases/0.3.9.md)
|
|
518
519
|
- [0.3.8 release notes](docs/releases/0.3.8.md)
|
|
519
520
|
- [0.3.7 release notes](docs/releases/0.3.7.md)
|
|
@@ -541,7 +542,7 @@ scripts package and release-readiness checks
|
|
|
541
542
|
|
|
542
543
|
The original M0–M13 roadmap is complete. `ds4-context-core` contains the compiled runtime-neutral implementation. M14 context-quality metrics, M15 rich symbol indexing, M16 hybrid semantic retrieval, M17 cross-session project memory, M18 learned-ranking shadow evaluation, M19's runtime adapter/conformance kit, and M20 opt-in local KV eligibility/replay are implemented on `main`. Learned active ranking remains promotion-gated, Pi reports local KV as unsupported, and static ranking/native completion stay authoritative on every failure.
|
|
543
544
|
|
|
544
|
-
The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.3.9 extends the bounded compaction updates, summary input headroom, concurrent segments and phase timings introduced in 0.3.5 with per-request and cumulative operation input limits; 0.3.8 adds indexed FTS key deletion without changing search results. The opt-in [anchored editing](docs/ANCHORED_EDITING.md) and [portable agent tools](docs/PORTABLE_AGENT_TOOLS.md) from 0.3.4 remain default-off, without backend rewind, forced sampling or operational KV integration. Confirmation, provenance, Pi fallback and canonical/configuration/SQLite/runtime contracts remain unchanged. The [0.2 readiness record](docs/RELEASE_READINESS_0.2.0.md) remains the compatibility baseline; the lexical planner stays available as the deterministic fallback.
|
|
545
|
+
The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.3.10 adds opt-in BPE estimation and bounded model-budget auto-tuning while retaining the `chars-v1` default. Version 0.3.9 extends the bounded compaction updates, summary input headroom, concurrent segments and phase timings introduced in 0.3.5 with per-request and cumulative operation input limits; 0.3.8 adds indexed FTS key deletion without changing search results. The opt-in [anchored editing](docs/ANCHORED_EDITING.md) and [portable agent tools](docs/PORTABLE_AGENT_TOOLS.md) from 0.3.4 remain default-off, without backend rewind, forced sampling or operational KV integration. Confirmation, provenance, Pi fallback and canonical/configuration/SQLite/runtime contracts remain unchanged. The [0.2 readiness record](docs/RELEASE_READINESS_0.2.0.md) remains the compatibility baseline; the lexical planner stays available as the deterministic fallback.
|
|
545
546
|
|
|
546
547
|
## Contributing
|
|
547
548
|
|
package/docs/CONTEXT_MANIFEST.md
CHANGED
|
@@ -45,7 +45,7 @@ actualInputTokens = input + cacheRead + cacheWrite
|
|
|
45
45
|
rawCalibrationRatio = actualInputTokens / estimatedInputTokens
|
|
46
46
|
```
|
|
47
47
|
|
|
48
|
-
The existing scalar SQLite columns are authoritative for persisted usage. `message_end` updates those columns and the calibration sample without reading or rewriting `manifest_json`; repository reads hydrate the usage fields from the scalars. The raw estimate is retained even after calibration so ratios cannot recursively calibrate already-corrected values. The next call reads only the exact provider/model window and applies bounded median/MAD outlier rejection. Error, aborted, missing, zero-usage, duplicate, or uncorrelated responses do not create calibration samples. Every manifest is correlated with at most one assistant response. If a concurrently pruned or oversize-skipped manifest cannot be correlated, DS4 keeps only bounded volatile calibration and does not retry against another row. See [`MODEL_AWARENESS.md`](MODEL_AWARENESS.md).
|
|
48
|
+
The existing scalar SQLite columns are authoritative for persisted usage. `message_end` updates those columns and the calibration sample without reading or rewriting `manifest_json`; repository reads hydrate the usage fields from the scalars. The raw estimate is retained even after calibration so ratios cannot recursively calibrate already-corrected values. The next call reads only the exact provider/model/**estimator-version** window and applies bounded median/MAD outlier rejection. When present, `modelAwareness.drift` and `modelAwareness.autoTune` contain only aggregate warning/tuning metadata (no message contents). Error, aborted, missing, zero-usage, duplicate, or uncorrelated responses do not create calibration samples. Every manifest is correlated with at most one assistant response. If a concurrently pruned or oversize-skipped manifest cannot be correlated, DS4 keeps only bounded volatile calibration and does not retry against another row. See [`MODEL_AWARENESS.md`](MODEL_AWARENESS.md).
|
|
49
49
|
|
|
50
50
|
## Retention
|
|
51
51
|
|
package/docs/MODEL_AWARENESS.md
CHANGED
|
@@ -1,5 +1,7 @@
|
|
|
1
1
|
# Advanced Model Awareness
|
|
2
2
|
|
|
3
|
+
**Validation status (OpenRouter, synthetic DS4/Pi sessions):** GPT-6-sol, GPT-6-luna and GPT-5.6-terra have each demonstrated multi-turn BPE calibration, a genuinely larger opt-in tail budget, provider usage above the native-window headroom threshold, and withdrawal of that expansion on the next turn in the same session. Historical retrieval and a live tool cycle were verified separately for all three. These bounded experiments do not establish universal tokenizer accuracy, provider-side compaction behavior, or production retrieval quality. `chars-v1` remains the default; BPE and `autoTune` remain opt-in. See the measured runs below.
|
|
4
|
+
|
|
3
5
|
M11 resolves an independent deterministic planning profile for every exact `provider/model`. Pi's model descriptor remains the default source for context window, output ceiling, reasoning, and image support; explicit DS4 overrides can repair provider metadata or tune category limits without changing canonical session state.
|
|
4
6
|
|
|
5
7
|
## Profile resolution
|
|
@@ -68,10 +70,10 @@ Each successful assistant response is correlated with one pending Context Manife
|
|
|
68
70
|
|
|
69
71
|
```text
|
|
70
72
|
actualInputTokens = input + cacheRead + cacheWrite
|
|
71
|
-
ratio = actualInputTokens / raw
|
|
73
|
+
ratio = actualInputTokens / raw selected-estimator estimate
|
|
72
74
|
```
|
|
73
75
|
|
|
74
|
-
Calibration is isolated by exact provider and
|
|
76
|
+
Calibration is isolated by exact provider, model ID **and estimator version**; switching from `chars-v1` to BPE never reuses the old ratios. The latest configured window is validated and processed deterministically:
|
|
75
77
|
|
|
76
78
|
1. reject invalid or zero samples;
|
|
77
79
|
2. reject ratios outside the configured hard bounds;
|
|
@@ -82,7 +84,38 @@ Calibration is isolated by exact provider and model ID. The latest configured wi
|
|
|
82
84
|
|
|
83
85
|
Until enough samples exist, the multiplier remains `1.0`. The default window is 24 samples, minimum is 3, and hard ratio bounds are 0.5–2.0. Outliers remain counted in metadata but do not influence the applied ratio.
|
|
84
86
|
|
|
85
|
-
Provider-token capacities are computed first from context window, output reserve, safety margin, and policy ratios.
|
|
87
|
+
Provider-token capacities are computed first from context window, output reserve, safety margin, and policy ratios. Global hard/soft/preferred input limits are converted into local-estimator units using **at least** a multiplier of 1: observed underestimation may reduce them, but apparent overestimation on short prompts never raises them above nominal model limits. Raw manifest estimates remain uncalibrated so future samples do not feed a corrected estimate back into itself. Adaptive tail/history/project budgets use the accepted multiplier but remain capped by the configured `context.*` maxima even after conversion.
|
|
88
|
+
|
|
89
|
+
### Optional BPE estimator and measured budget tuning
|
|
90
|
+
|
|
91
|
+
The default remains `chars-v1`. To opt a profile into local OpenAI `o200k_base` BPE text counting (without fetching vocabulary over the network):
|
|
92
|
+
|
|
93
|
+
```json
|
|
94
|
+
{
|
|
95
|
+
"modelAwareness": {
|
|
96
|
+
"overrides": { "openai/gpt-4o": { "tokenEstimator": "o200k-base-v1" } },
|
|
97
|
+
"autoTune": true
|
|
98
|
+
}
|
|
99
|
+
}
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
Use BPE only for a model known to share that encoding. The estimator counts text with BPE in the Pi observer, managed planner and fixed prompt/tool estimates; message wrappers, images, reasoning, provider-specific serialization, indexed retrieval hints, and persisted summaries remain estimates. It does **not** make token counts exact or enable provider-side continuation. If the provider's actual input tokens drift persistently from estimates (median ≥1.25 or ≤0.80 after the calibration minimum, or repeated hard-bound outliers), `/context model` and status surface a metadata-only warning; no prompt content is logged.
|
|
103
|
+
|
|
104
|
+
`autoTune` is opt-in and starts neutral. After at least eight accepted calibration samples, it expands *automatic* recent-tail, history and project ceilings by at most 12.5% of their nominal values, only if **every valid provider-usage sample** in the calibration window — including ratio outliers excluded from ratio calibration — remains at or below 60% of the current preferred/hard input target. It never raises a ceiling beyond the configured context cap, never overrides an explicit model-specific category limit, and never changes the provider's window, output reserve, safety margin or hard/soft input limits. Without enough evidence or headroom it stays neutral. Any calibration multiplier is applied after this small expansion; tune decisions and warnings are recorded as metadata in the manifest.
|
|
105
|
+
|
|
106
|
+
## Validation goal for BPE and auto-tuning
|
|
107
|
+
|
|
108
|
+
**User-requested outcome:** verify long conversations, tools, retrieval, calibration across turns, and auto-tuning safety on the context actually built by DS4/Pi. Single-turn synthetic token probes are preliminary evidence only; they do **not** satisfy this goal or justify declaring the feature complete.
|
|
109
|
+
|
|
110
|
+
The validation must cover, and report results separately for:
|
|
111
|
+
|
|
112
|
+
1. Long multi-turn sessions, including large context windows, compaction boundaries, and preservation of the current request.
|
|
113
|
+
2. Tool calls and results (including large results): atomic selection, offload, and the estimated versus observed token count.
|
|
114
|
+
3. Retrieved history and project context: relevance, budget selection, and whether relevant older material survives planning.
|
|
115
|
+
4. Consecutive turns on the **same** provider/model and estimator profile: provider usage correlation, accepted/rejected calibration samples, changes to the applied ratio, and isolation when models or estimators change.
|
|
116
|
+
5. Opt-in auto-tuning: evidence thresholds, configured and hard limits, outliers, insufficient headroom, and no regression in retrieval or context integrity when limits expand.
|
|
117
|
+
|
|
118
|
+
Use local, deterministic integration tests where possible. If real provider calls are needed, require a **new explicit call and input-size limit** before sending anything; the earlier single-turn budgets are exhausted. Report any scenario not verified as *not verified*, rather than treating the synthetic pilot as an end-to-end result. Do not change the default estimator or enable auto-tuning by default on the strength of that pilot.
|
|
86
119
|
|
|
87
120
|
## Provider cache metrics
|
|
88
121
|
|
|
@@ -126,6 +159,42 @@ Use:
|
|
|
126
159
|
/context status
|
|
127
160
|
```
|
|
128
161
|
|
|
162
|
+
## Local validation progress and outstanding provider evidence
|
|
163
|
+
|
|
164
|
+
`tests/integration/model-aware-real-context.test.ts` exercises the Pi `context` and `message_end` hooks over canonical multi-turn JSONL without any network transport. A long branch with a tool call/result verifies atomic selection and retrieval of an older decision under a BPE-managed budget. Successive turns verify that eight correlated usage records are required before opt-in expansion, that estimator/model switches isolate calibration, and that a new near-limit usage record withdraws the expansion even when its ratio is excluded from calibration. `tests/unit/model-aware-estimator.test.ts` also checks explicit overrides, configured caps, and the high-usage outlier regression.
|
|
165
|
+
|
|
166
|
+
**Local integration checks inject usage; they are not independent provider measurements.** A separately authorized, bounded live probe (`scripts/verify-model-aware-real-session.mjs`) used a temporary Pi JSONL session, DS4's actual hooks, synthetic content, `openrouter/openai/gpt-4o-mini`, and Pi-normalized provider usage. Eight successive short calls accumulated calibration samples; on the ninth call the manifest showed eight accepted samples, applied ratio 0.677157, and opt-in `autoTune: expanded`. That ninth call did **not** contain the intended appended long branch: Pi had already constructed its runner, and its manifest estimated only 1,940 input tokens. It is not evidence for long-context tuning.
|
|
167
|
+
|
|
168
|
+
A second, fresh Pi session with the long branch seeded *before* runner construction produced 83 original messages; DS4 selected 21 groups, excluded 21, retrieved the older `cobalt-713` decision, and included both the synthetic tool-call assistant entry and tool result. For that request, the BPE-based manifest estimate was **23,492** input tokens and provider usage was **23,196**. This is **one** live, untuned long-session/tool/retrieval measurement; it does not prove stable drift across lengths or models. The earlier [single-turn pilots](PROVIDER_TOKEN_DRIFT_BENCHMARK.md) remain separate. The two runs together attempted 10 calls and conservatively reserved 217,678 of the authorized 250,000 estimated-input-token limit (per-call ceiling 64,000; 16-call ceiling). All temporary session data was deleted.
|
|
169
|
+
|
|
170
|
+
A subsequent, separately authorized probe (`scripts/verify-model-aware-calibrated-session.mjs`) seeded a 28-turn synthetic history and tool call/result **before starting the same Pi session**. Eight real provider calls calibrated BPE on that long context; the ninth measured **23,926 estimated / 23,486 actual input tokens**, eight accepted samples, `autoTune: expanded`, and retrieval of the old decision with the historical tool call and result included. This verifies expansion on a real long context after calibration in one session, not just across separate sessions. Across this probe **11 calls** reserved **588,052/600,000** estimated input tokens (80,000 per-call limit); no further provider requests were made under that authorization.
|
|
171
|
+
|
|
172
|
+
The attempted higher-occupancy request used **37,552 estimated / 37,464 actual tokens**, below the 60%-of-target withdrawal threshold (**53,760** for this model/configuration). The following turn still reported `expanded`: **withdrawal is not verified live**. That larger turn also excluded the historical tool/retrieval groups, an observed quality limit under that selected context, not proof of an auto-tuning regression. The local Pi compaction-boundary test covers summary preservation under BPE. A separately bounded **one-request** live compaction probe (`scripts/verify-model-aware-compaction-session.mjs`) preseeded the canonical Pi compaction entry before starting Pi; the DS4 managed manifest included its summary and measured **250 estimated / 244 actual** provider input tokens. This used one further call under the third authorization, bringing that budget to **10 calls / 181,069 estimated input tokens reserved**. It checks the provider-bound context *after* a synthetic compaction entry, not live summary generation. At that point, live tool execution and high-usage withdrawal remained unverified; both were tested in the later authorized round below. Model variety and production retrieval quality still remain outside these bounded probes. A third bounded probe (`scripts/verify-model-aware-safety-session.mjs`) reached **60,509 actual input tokens** (>53,760), but only **seven** of eight preceding short-call samples had been accepted; the manifest at the high-usage request showed seven accepted samples and `insufficient-samples`, not an expanded policy. Thus it **does not** test withdrawal. The safety probe attempted **9 calls**, reserving **169,006** estimated input tokens, and stopped without a follow-up call. It now gates the expensive request on eight *accepted* samples and seeds a stable synthetic prefix, but has not been rerun. Its temporary session was deleted; the remaining third-round budget after the separate compaction request cannot fit a new calibration/high-usage/follow-up sequence under the preflight limits. No live tool execution was attempted in that round because the automatic multi-request cycle was not yet safely capped per request. Those results alone did not verify withdrawal or tool execution.
|
|
173
|
+
|
|
174
|
+
A final, separately authorized round verified the outstanding **synthetic** provider-path scenarios on the same OpenRouter/GPT-4o-mini profile. After a stable synthetic prefix, eight usage samples were accepted; the ninth short turn showed `expanded`. The next request reported **60,708 estimated / 60,640 actual** input tokens, above the **53,760** headroom threshold; the following turn reported **60,795 estimated / 60,722 actual** and `no-headroom` (recent-tail limit decreased from **27,699** to **24,564**). `scripts/verify-model-aware-safety-session.mjs` used **11/15** calls and **292,818/350,000** estimated-input tokens reserved (maximum 80,000 per call). This demonstrates withdrawal on measured provider usage, not merely a synthetic injected sample.
|
|
175
|
+
|
|
176
|
+
With the remaining budget, `scripts/verify-model-aware-live-tool.mjs` gated **each** Pi provider request at `ModelRuntime.streamSimple`, disabled cache warming/retry/compaction, and executed a single synthetic tool. The first request included its tool schema and used **105** provider input tokens; the second included the actual tool result and used **139**. Exactly **two** requests and one tool execution occurred. The final fourth-round total was **13/15 calls, 317,119/350,000** estimated-input tokens reserved. Only aggregate counters and booleans were reported; temporary synthetic sessions were removed.
|
|
177
|
+
|
|
178
|
+
These bounded live measurements, together with the canonical Pi JSONL integration tests, cover long context, historical and executed tools, retrieval, between-turn calibration, compaction-boundary context, and auto-tuning withdrawal **for this synthetic OpenRouter/GPT-4o-mini profile**. They do not establish universal accuracy for other providers/models, real private sessions, arbitrary compaction generation, or production retrieval quality.
|
|
179
|
+
|
|
180
|
+
### Native-window checks for Sol, Luna and Terra — separately authorized
|
|
181
|
+
|
|
182
|
+
The Pi catalog exposed a **1,050,000-token window** for each OpenRouter model `openai/gpt-6-sol`, `openai/gpt-6-luna`, and `openai/gpt-5.6-terra`. One additional authorization limited their *combined* probes to **42 provider attempts, 600,000 estimated input tokens / 2,500,000 controlled characters per request, and 4,800,000 estimated tokens / 15,000,000 controlled characters total**. `scripts/verify-model-aware-triad.mjs` and `scripts/verify-model-aware-triad-followup.mjs` gated **every** Pi request at `ModelRuntime.streamSimple`, including tool continuations; no real sessions, credentials, prompt bodies, responses or raw upstream errors were logged. Temporary canonical Pi JSONL sessions were deleted.
|
|
183
|
+
|
|
184
|
+
The first probe consumed **30 calls / 1,433,445 tokens / 3,679,430 characters reserved**. For *each* model, nine calibration calls in the same seeded long session led to at least eight accepted usage samples. The tenth call measured **36,027 BPE-estimated / 35,503 Pi-normalized actual input tokens**, `autoTune: expanded`, with the historical tool call and result included. **This did not verify retrieval**: the decision was still in the 64k recent tail and therefore had not been fetched by retrieval; the probe correctly stopped before tool execution and the high-occupancy turn. At this native profile, the reported `expanded` status did **not** demonstrate a larger recent-tail limit: its configured 64k cap was already saturated.
|
|
185
|
+
|
|
186
|
+
The second probe seeded **70 turns before Pi session construction**, putting the old decision outside the default recent tail. It consumed nine more calls (one retrieval request and an actual **two-request, one-execution** synthetic tool cycle per model). The decision was both retrieved and included, the historical tool call/result stayed included, and the new tool schema appeared in the first provider request with its actual result in the second. Per-model provider usage for the retrieval request was **63,050 / 63,049 / 63,053** tokens (Sol/Luna/Terra), against BPE estimates **63,699 / 63,698 / 63,702**. All three tool cycles completed. These are measurements of the DS4/Pi managed context, not three single-turn tokenizer prompts.
|
|
187
|
+
|
|
188
|
+
The three remaining authorized requests measured large BPE-managed input, **one per model**. The selected current user turn was included, and Pi reported **381,501 / 381,502 / 381,502** input tokens against manifest estimates **381,593 / 381,594 / 381,594**. The historical call/result were present, but the old decision was **not** selected; the oversized prompt did not ask about that decision, so this is **not** a valid retrieval-quality test. The request generator targeted 462k from the *raw* canonical history, but DS4 excluded older history and the selected provider input remained ~381.5k. Thus all three requests were **below** the native-window auto-tuning withdrawal threshold of **441,000** (60% of the 735k preferred target), and far below the full 1.05M window. They were fresh sessions without eight accepted calibration samples; **none tests `expanded → high provider usage → no-headroom` on these models**. Do not report the triad's auto-tuning safety or category expansion as live-verified at its native window; the deterministic local matrix in `tests/unit/model-aware-estimator.test.ts` covers only injected samples and caps. The entire authorization was used: **42/42 calls**, **3,296,138/4,800,000** estimated input tokens and **9,357,266/15,000,000** controlled characters reserved. No more provider calls are authorized under it. The historical probes now reject `--live` under this exhausted authorization. **After** those measurements, their local high-input sizing was corrected to require at least 470k BPE tokens in the selected current user text (rather than in raw canonical history), with a second gate immediately before provider transport; unit tests cover the discarded-history regression and per-call preflight. This correction was **not** run against a provider and cannot retroactively validate withdrawal.
|
|
189
|
+
|
|
190
|
+
Across these specific synthetic shapes, BPE manifests slightly overestimated provider usage, but no universal correction follows. At this earlier point native-window compaction generation, production retrieval quality and live high-usage auto-tune withdrawal for Sol/Luna/Terra were unverified; the later withdrawal runs below close **only the last of those gaps**. The later probes check actual category growth (not merely `expanded` status), a selected current turn above 470k BPE tokens, usage above 441k, and withdrawal in the same Pi session. Their temporary opt-in category ceilings are 80k/40k/40k because the default native tail ceiling is already equal to its configured maximum of 64k and cannot grow; the probes do not establish expansion under unchanged category defaults. BPE and auto-tuning remain opt-in; no default or provider-storage policy was changed.
|
|
191
|
+
|
|
192
|
+
### Native-window withdrawal round (later, opt-in categories)
|
|
193
|
+
|
|
194
|
+
With a separate explicit cap of 39 calls, 3,800,000 reserved input tokens and 11,000,000 controlled characters, the same-session Pi/DS4 probe used temporary category ceilings 80k/40k/40k to expose a **real** tail expansion. It ran 34 provider requests (3,375,536 estimated input tokens reserved; 9,845,274 controlled characters). For **GPT-6-sol and GPT-6-luna**, each session accepted at least eight calibration samples; the expanded tail exceeded the same-ratio no-tune baseline (about 72,922 versus 64,819 estimator tokens). A later request used **475,703 provider input tokens** (BPE estimate 475,790), above the 441,000 provider-token headroom threshold. On the next request, `autoTune` was `no-headroom` and the tail equalled its no-tune baseline (64,805); actual provider input was 471,846/471,844. The selected current turn, historical decision, call and result were present before the large request. The large input and follow-up did **not** retain the unrelated old decision/tool group; these measurements do not establish production retrieval quality.
|
|
195
|
+
|
|
196
|
+
For **GPT-5.6-terra**, calibration and real expansion passed in that first round, but its high-input request was blocked **before transport** by the cumulative budget/character caps: the probe had incorrectly budgeted the follow-up as short even though Pi resends the large preceding turn. The temporary session was then disposed. A **separately authorized Terra-only run** used a fresh synthetic Pi/DS4 session and counted *both* large requests. It made 12 provider calls (1,449,079 estimated tokens and 4,310,843 characters reserved, under separate 13-call / 1,900,000-token / 6,000,000-character maxima). After at least eight accepted calibration samples, Terra's tail grew to 72,922 estimator tokens versus a same-ratio untuned baseline of 64,819. The large request used **475,707 provider input tokens** (BPE estimate 475,794), above the 441,000 threshold. On the next turn, provider input was **471,852** (estimate 471,929), `autoTune` was `no-headroom`, and the tail returned to **64,805**, exactly its untuned baseline at that turn's ratio. The historical decision/call/result and current turn were included before the large request; unrelated older groups were excluded during the large request. **This verifies the defined live high-usage expansion/withdrawal scenario for all three models, not a general retrieval or compaction guarantee.** Both probes are now locked against replay; defaults remain unchanged.
|
|
197
|
+
|
|
129
198
|
## Performance and tests
|
|
130
199
|
|
|
131
|
-
`tests/benchmarks/model-awareness.bench.ts` measures a bounded 200-sample calibration analysis and repeated 32k/128k/200k profile resolution. Unit and golden tests cover deterministic tiers, override precedence, robust outlier rejection, cache accounting, and calibrated budgets. Integration tests switch local/remote providers and 32k/128k/200k models while checking profile isolation, privacy re-enforcement, canonical JSONL preservation, SQLite cache metrics, and profile reuse.
|
|
200
|
+
`tests/benchmarks/model-awareness.bench.ts` measures a bounded 200-sample calibration analysis and repeated 32k/128k/200k profile resolution. Unit and golden tests cover deterministic tiers, override precedence, robust outlier rejection, cache accounting, and calibrated budgets. Integration tests switch local/remote providers and 32k/128k/200k models while checking profile isolation, privacy re-enforcement, canonical JSONL preservation, SQLite cache metrics, and profile reuse. For a separately authorized, bounded real-provider measurement of both estimators against Pi SDK usage, see [PROVIDER_TOKEN_DRIFT_BENCHMARK.md](PROVIDER_TOKEN_DRIFT_BENCHMARK.md); it does not substitute for DS4 manifest measurements in a live session.
|
|
@@ -0,0 +1,85 @@
|
|
|
1
|
+
# Live provider token-drift probe (opt-in)
|
|
2
|
+
|
|
3
|
+
This developer-only, **source-checkout-only** probe (the script is not included in the published npm package) compares DS4's raw `chars-v1` and `o200k-base-v1` estimates with **Pi SDK provider usage** for synthetic, single-turn text. It does not start a Pi session, read session JSONL or SQLite, record responses, or change DS4 configuration. It never prints API keys, requests, response bodies or raw upstream error messages. It requires local Pi provider authentication and network access. The probe runs only with `--live`; without that flag it prints the bounded plan.
|
|
4
|
+
|
|
5
|
+
```bash
|
|
6
|
+
npm run build:core
|
|
7
|
+
node scripts/compare-provider-token-drift.mjs \
|
|
8
|
+
--model openrouter/openai/gpt-4o-mini \
|
|
9
|
+
--model deepseek/deepseek-v4-flash \
|
|
10
|
+
--model openai-codex/gpt-5.4-mini
|
|
11
|
+
# After inspecting the dry-run budget, add --live to make real calls.
|
|
12
|
+
```
|
|
13
|
+
|
|
14
|
+
The script accepts repeated `--model provider/model-id` and optional `--sizes 512,4096,48000`. **Per invocation**, it refuses more than 12 model calls or 200,000 total input characters (system prompt included), disables SDK request retries, and imposes a 45-second deadline per call. Multiple invocations have separate caps: track their cumulative spend yourself. Output is one metadata-only JSON report on stdout. The SDK's catalog-derived USD cost is an estimate, not a bill; a provider can still charge for failures or retries outside this script's control. If you pipe output to a file, keep it outside the repo unless you intentionally want to publish the aggregate metadata.
|
|
15
|
+
|
|
16
|
+
For each successful call, `actualInputTokens = usage.input + usage.cacheRead + usage.cacheWrite`; the cached fractions are **not** extra tokens on top of that sum. `residualTokens = actual - rawEstimate` (positive means underestimation), and `underestimationPctOfActual = max(0, residual) / actual × 100`. Adjacent slopes use differences across prompt sizes for a single exact provider/model, removing most fixed framing. BPE counts only text; both estimators use the same DS4 message/system overhead. No raw provider payload is available from this probe, so the result does **not** prove the accuracy of DS4's full observer in a live Pi extension chain (tools, images, privacy, cache/continuation, and later extensions differ). Compare `/context model`, `/context tokens`, and manifest `actualInputTokens` versus `estimatedInputTokens` in an ordinary consented Pi session for that second step. Never copy session text or auth files into the report.
|
|
17
|
+
|
|
18
|
+
## First observed run — 24 September 2026, 10:09–10:13 UTC
|
|
19
|
+
|
|
20
|
+
User-approved cumulative envelope: 12 call attempts and 200,000 input characters. Three invocations totalled **12 attempts and 195,176 planned input characters**: 9 calls/158,382 characters, a single Codex diagnostic call/574 characters, then 2 calls/36,220 characters. Success: 8; failure/missing usage: 4. Synthetic multilingual/code-like repeated text; no private session data. Sizes below are user-message characters; system prompt is included in estimates and the total envelope. These are observations, **not** a representative multi-session cost or quality benchmark.
|
|
21
|
+
|
|
22
|
+
| Provider / requested model | User chars | Real input¹ | `chars-v1` | BPE `o200k-base-v1` | BPE residual (real − estimate) |
|
|
23
|
+
|---|---:|---:|---:|---:|---:|
|
|
24
|
+
| OpenRouter / `openai/gpt-4o-mini` | 512 | 190 | 161 | 196 | −6 |
|
|
25
|
+
| OpenRouter / `openai/gpt-4o-mini` | 4,096 | 1,436 | 1,057 | 1,442 | −6 |
|
|
26
|
+
| OpenRouter / `openai/gpt-4o-mini` | 48,000 | 16,668 | 12,033 | 16,674 | −6 |
|
|
27
|
+
| DeepSeek / `deepseek-v4-flash`² | 512 | 179 | 161 | 196 | −17 |
|
|
28
|
+
| DeepSeek / `deepseek-v4-flash`² | 4,096 | 1,388 | 1,057 | 1,442 | −54 |
|
|
29
|
+
| DeepSeek / `deepseek-v4-flash`² | 48,000 | 16,172 | 12,033 | 16,674 | −502 |
|
|
30
|
+
| OpenRouter / `deepseek/deepseek-v4-flash` | 4,096 | 1,387 | 1,057 | 1,442 | −55 |
|
|
31
|
+
| OpenRouter / `deepseek/deepseek-v4-flash` | 32,000 | 10,785 | 8,033 | 11,124 | −339 |
|
|
32
|
+
|
|
33
|
+
¹ Pi's normalized provider usage: input + cache read + cache write. Some long requests reported cached reads (up to 1,408 tokens); they were included exactly once. Successful responses produced 1–3 output tokens. ² DeepSeek reported response model `deepseek-flash`, an alias of the requested ID; no equivalence to OpenRouter routing is assumed.
|
|
34
|
+
|
|
35
|
+
- OpenRouter GPT-4o-mini: raw `chars-v1` median actual/estimate **1.358562**; raw BPE median **0.995839**. At 48k chars, chars/4 underestimated by 4,635 tokens (27.81% of actual), while BPE overestimated by 6 tokens. Adjacent BPE slopes were 1.000 and 1.000.
|
|
36
|
+
- Direct DeepSeek: raw chars median **1.313150**; BPE median **0.962552**. At 48k chars, chars/4 underestimated by 4,139 tokens (25.59% of actual), while BPE overestimated by 502 tokens. Adjacent BPE slopes were 0.970305 and 0.970588. Close agreement here does **not** establish that DeepSeek uses OpenAI's tokenizer or justify turning BPE on for that model by default.
|
|
37
|
+
- OpenRouter DeepSeek: two valid samples; raw chars median **1.327396**, BPE median **0.965692**. At 32k chars, chars/4 underestimated by 2,752 tokens; BPE overestimated by 339. Two samples are insufficient to validate a persistent drift warning or tuning.
|
|
38
|
+
- `openai-codex/gpt-5.4-mini`: three initial attempts produced no usable usage; one later 512-character diagnostic returned `request-failed` with sanitized category `other` and no HTTP status. The local Pi auth resolver did return a credential, but the reason for the probe failure remains **unverified**. Do not infer its tokenizer drift, account availability in the interactive TUI, or provider-side token consumption from these calls. Direct `openai/gpt-4o-mini` API auth was not configured in the Pi runtime used by this probe; OpenRouter is a distinct route.
|
|
39
|
+
|
|
40
|
+
**Interpretation:** Three sizes with one call each do not establish statistical reliability or cover large context windows, tools, images, prefixes reused across turns, or output-heavy workflows. `chars-v1` remains the default; existing per-model calibration may compensate after enough **accepted same-profile samples**. The first small sample can be excluded by MAD filtering, so three different-size wire calls need not become three accepted DS4 calibration samples. Keep BPE and `autoTune` opt-in; do not port Hub's fixed 3.5% margin from these measurements. For promotion or automatic budget changes, collect a repeated same-size and mixed-size sample with matching DS4 manifest estimates and actual provider usage under an explicitly authorized, separately bounded run.
|
|
41
|
+
|
|
42
|
+
## Second observed run — 24 September 2026, 10:35 UTC
|
|
43
|
+
|
|
44
|
+
A separate, explicitly approved envelope covered **12 call attempts and up to 100,000 input characters**. The dry run planned exactly 12 attempts and 99,816 characters: `--sizes 512,16000` for three requested OpenAI models through each of OpenRouter and Codex. The live probe produced six OpenRouter usages and six Codex failures; it did not read session data.
|
|
45
|
+
|
|
46
|
+
| Route / requested model | User chars | Real input¹ | `chars-v1` | BPE `o200k-base-v1` | BPE residual |
|
|
47
|
+
|---|---:|---:|---:|---:|---:|
|
|
48
|
+
| OpenRouter / `openai/gpt-6-sol` | 512 | 189 | 161 | 196 | −7 |
|
|
49
|
+
| OpenRouter / `openai/gpt-6-sol` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
|
|
50
|
+
| OpenRouter / `openai/gpt-6-luna` | 512 | 189 | 161 | 196 | −7 |
|
|
51
|
+
| OpenRouter / `openai/gpt-6-luna` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
|
|
52
|
+
| OpenRouter / `openai/gpt-5.6-terra` | 512 | 189 | 161 | 196 | −7 |
|
|
53
|
+
| OpenRouter / `openai/gpt-5.6-terra` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
|
|
54
|
+
|
|
55
|
+
¹ Pi-normalized input includes cached tokens once. Each 16,000-character call reported `cacheWriteTokens = 5,561` and zero cache-read tokens. Those were **writes**, not cache hits. Each successful response reported five output tokens. The six equal input counts describe these identical synthetic prompts on this route; they are not proof of identical tokenizers or of how the direct Codex route would count a DS4 session. For each model, the adjacent BPE size slope was exactly 1.000; at 16k characters, raw `chars-v1` underestimated by 1,531 tokens (27.52% of actual), versus BPE overestimating by seven.
|
|
56
|
+
|
|
57
|
+
For `openai-codex/gpt-6-sol`, `gpt-6-luna`, and `gpt-5.6-terra`, **both sizes failed** without provider usage. The sanitized error classifier returned `quota` for all six, with no HTTP status captured. This is evidence of an SDK-path quota/limit error classification, **not** a verified account balance, an HTTP rejection, or a tokenizer measurement. It does not retroactively identify the earlier `gpt-5.4-mini` failure (classified `other`). No further calls were made beyond this run's approved envelope. Two sizes per exact model remain insufficient for DS4 calibration or auto-tuning conclusions.
|
|
58
|
+
|
|
59
|
+
## Isolated Pi + DS4 managed-context pilot — 24 September 2026
|
|
60
|
+
|
|
61
|
+
For a future, **separately authorized** run: build the core (`npm run build:core`), inspect the planned sizes and character count with `node scripts/compare-ds4-manifest-usage.mjs`, then use `node scripts/compare-ds4-manifest-usage.mjs --live` only after setting a new call/character budget. The default mode is pinned to OpenRouter `openai/gpt-6-sol` and a maximum of 12 calls / 120,000 controlled prompt characters per invocation. `--comparison` instead preflights **both** OpenRouter `openai/gpt-6-luna` and `openai/gpt-5.6-terra`, capped at 24 calls / 240,000 controlled characters combined; `--comparison --live` requires its own authorization. A new authorization is required even when repeating either exact plan.
|
|
62
|
+
|
|
63
|
+
The user separately authorized **up to 12 calls / 120,000 controlled input characters** to OpenRouter `openai/gpt-6-sol` only. `scripts/compare-ds4-manifest-usage.mjs` first checked a local sandbox without provider traffic, then made **12 single-turn calls**, with three repeats at each size (512, 4,096, 12,000 and 20,000 synthetic user characters). Planned synthetic user + fixed system text was **110,568 characters**. This limit counts controlled prompt text, **not** Pi-generated framing or serialized protocol bytes. No existing Pi session history was used; each Pi session and DS4 database was isolated in a temporary directory, with no tools, project content, native continuation, or auto-tuning. DS4's managed `context` hook and its `o200k-base-v1` profile override were active. The script stops on a missing/mismatched manifest or missing usage rather than fabricating a measurement.
|
|
64
|
+
|
|
65
|
+
| Synthetic user chars | First observed DS4 manifest estimate | First actual input¹ | First residual (actual − estimate) | Repeats |
|
|
66
|
+
|---:|---:|---:|---:|---:|
|
|
67
|
+
| 512 | 220 | 209 | −11 | 3 |
|
|
68
|
+
| 4,096 | 1,467 | 1,456 | −11 | 3 |
|
|
69
|
+
| 12,000 | 4,210 | 4,199 | −11 | 3 |
|
|
70
|
+
| 20,000 | 6,983 | 6,972 | −11 | 3 |
|
|
71
|
+
|
|
72
|
+
All **12** calls had matching `managed` manifests and Pi-normalized provider usage, with **exactly −11 tokens** of residual in each call. Token counts varied by about one or two between identically sized repeats, but the residual stayed fixed. The largest relative overestimate was **5.26% of actual input** on the first 512-character call (11/209); at 20,000 characters it was about **0.16%**. `cacheReadTokens` was zero, while longer calls reported mostly cache **writes**; no claim of cache hits or exact billed usage follows from those fields.
|
|
73
|
+
|
|
74
|
+
¹ `totalInputTokens` in the DS4 manifest (the same Pi-normalized `input + cacheRead + cacheWrite` metric used by model-awareness calibration), not independently obtained provider invoice data. Because the selected estimator was BPE, this pilot did **not** produce an alternative `chars-v1` DS4 manifest for the same wire requests. These are fresh one-turn sessions with one synthetic prompt shape and a maximum of ~7k actual tokens; they do not validate multi-turn prefixes, tool schemas/results, retrieval, real project text, larger windows, other requested models, or production auto-tuning. The 12 samples are independent sessions, **not** 12 accepted samples in one persistent DS4 calibration profile. Keep BPE opt-in and `autoTune` off by default; do not apply an inferred fixed framing correction or Hub's fixed 3.5% drift allowance from this pilot.
|
|
75
|
+
|
|
76
|
+
## Separately authorized Luna/Terra managed-context comparison — 24 September 2026
|
|
77
|
+
|
|
78
|
+
A subsequent authorization covered **at most 24 calls / 240,000 controlled synthetic input characters combined** on OpenRouter `openai/gpt-6-luna` and `openai/gpt-5.6-terra`; no Codex traffic was authorized. `node scripts/compare-ds4-manifest-usage.mjs --comparison` preflighted **24 calls / 221,136 controlled characters**. `--comparison --live` completed **12 calls per model**, with three fresh, isolated Pi+DS4 managed sessions per model at each synthetic user size (512, 4,096, 12,000, 20,000 characters). No existing session history, project content, tools, auto-tuning, or native continuation entered the requests. All 24 calls had matching manifests and Pi-normalized provider input usage; there were no reported failed rows.
|
|
79
|
+
|
|
80
|
+
| Model | First 512-character DS4 BPE estimate | First actual input¹ | Repeats per size | Residual on all 12 calls |
|
|
81
|
+
|---|---:|---:|---:|---:|
|
|
82
|
+
| OpenRouter / `openai/gpt-6-luna` | 220 | 209 | 3 | −11 tokens |
|
|
83
|
+
| OpenRouter / `openai/gpt-5.6-terra` | 219 | 208 | 3 | −11 tokens |
|
|
84
|
+
|
|
85
|
+
In every size group for both models, each of the three residuals was **−11 tokens**; each group reported zero `cacheReadTokens`. This replicates the small, constant **overestimate** observed for Sol on this specific one-turn synthetic shape. It does **not** establish a provider-independent 11-token correction, tokenizer equivalence across arbitrary text, long-window accuracy, model calibration from a continuous session, or safe budget auto-tuning. The combined 24-call authorization is exhausted. A later, separate authorization covered multi-turn, historical and executed tools, retrieval, and large selected inputs for all three OpenRouter models; see [native-window checks](MODEL_AWARENESS.md#native-window-checks-for-sol-luna-and-terra--separately-authorized) for results and the remaining unverified high-usage auto-tune withdrawal. Codex diagnosis still needs separate approval; keep BPE opt-in and `autoTune` off by default.
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
# Release 0.3.10 — Opt-in BPE estimation and bounded model budget auto-tuning
|
|
2
|
+
|
|
3
|
+
**Coordinated packages:** `ds4-context-core`, `ds4-context-reference-adapter`, and `ds4-context-engine` 0.3.10.
|
|
4
|
+
**Implementation commit:** `0607238`.
|
|
5
|
+
|
|
6
|
+
## Summary
|
|
7
|
+
|
|
8
|
+
Adds a selectable `o200k-base-v1` text estimator, provider/model/estimator-specific calibration, and opt-in evidence-gated expansion of automatic context-category ceilings. `chars-v1` remains the default estimator; `modelAwareness.autoTune` remains off unless explicitly enabled. The reference adapter and Pi extension continue to depend exactly on the matching core version.
|
|
9
|
+
|
|
10
|
+
## Changes
|
|
11
|
+
|
|
12
|
+
- Portable core accepts a BPE estimator through its existing runtime-neutral interface; the Pi adapter lazy-loads `js-tiktoken`. Non-text content still uses bounded heuristics.
|
|
13
|
+
- Model profiles can select `tokenEstimator` by exact `provider/model`, provider wildcard, or global override. Calibration histories are isolated by estimator version as well as provider and model.
|
|
14
|
+
- Auto-tuning uses accepted same-profile samples, rejects outliers, respects configured category limits and hard input ceilings, and withdraws expansion after insufficient headroom. Calibrated estimator-unit limits cannot exceed nominal model limits.
|
|
15
|
+
- Manifests and model diagnostics identify estimator version and calibration/auto-tuning decisions. Existing golden compatibility remains anchored to the unchanged default behavior.
|
|
16
|
+
|
|
17
|
+
## Measured scope and limitations
|
|
18
|
+
|
|
19
|
+
Bounded synthetic multi-turn Pi/DS4 sessions on OpenRouter `openai/gpt-6-sol`, `openai/gpt-6-luna`, and `openai/gpt-5.6-terra` demonstrated calibration, a genuinely larger opt-in recent tail, provider usage above the native-window headroom threshold, and withdrawal to `no-headroom` on the following turn in each model's session. Retrieval and live tool cycles were also exercised separately. See [model-awareness measurements](../MODEL_AWARENESS.md) and [provider-token drift](../PROVIDER_TOKEN_DRIFT_BENCHMARK.md) for methodology, numbers, and bounds.
|
|
20
|
+
|
|
21
|
+
These measurements do not establish production retrieval quality, provider-side compaction behavior, or accuracy for arbitrary models/routes. In particular, equivalent Codex-route measurements did not yield usage and are not claimed as validated. No provider calls are required for this release procedure.
|
|
22
|
+
|
|
23
|
+
## Compatibility
|
|
24
|
+
|
|
25
|
+
No new defaults are enabled: `chars-v1`, disabled `autoTune`, and disabled DS4 native continuation remain unchanged. No SQLite migration or change to canonical Pi JSONL, privacy consent, native continuation, or the portable runtime-adapter contract is introduced. BPE remains adapter-injected rather than adding a third-party tokenizer dependency to portable core. Restart Pi after upgrading so the new compiled core and session configuration are loaded.
|
|
26
|
+
|
|
27
|
+
## Validation and publication
|
|
28
|
+
|
|
29
|
+
On Node 26.5.1, the coordinated release passed a clean `npm ci`, `npm run check` (97 Vitest files, 607 tests, TypeScript builds and root typecheck), deterministic `npm run quality:compare`, the `npm run schema:context-persistence` size bound, and `npm run pack:check` in a clean consumer after synchronizing both exported runtime version constants. All typecheck/tests were also rerun after that correction. Dry-run tarball review contained 243 core files, 7 reference-adapter files, and 97 extension files; none included untracked local state. Registry publication and exact-artifact verification are still pending.
|
package/docs/releases/0.3.9.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Release 0.3.9 — Bounded compaction request and operation input
|
|
2
2
|
|
|
3
3
|
**Version analyzed:** DS4 Context Engine `0.3.9`
|
|
4
|
-
**Commit:**
|
|
4
|
+
**Commit:** `01407cd`
|
|
5
5
|
**Coordinated packages:** `ds4-context-core` 0.3.9, `ds4-context-reference-adapter` 0.3.9, `ds4-context-engine` 0.3.9
|
|
6
6
|
|
|
7
7
|
## Summary
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "ds4-context-engine",
|
|
3
|
-
"version": "0.3.
|
|
3
|
+
"version": "0.3.10",
|
|
4
4
|
"description": "Non-destructive, provider-independent context management for Pi.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|
|
@@ -62,7 +62,8 @@
|
|
|
62
62
|
]
|
|
63
63
|
},
|
|
64
64
|
"dependencies": {
|
|
65
|
-
"ds4-context-core": "0.3.
|
|
65
|
+
"ds4-context-core": "0.3.10",
|
|
66
|
+
"js-tiktoken": "1.0.21"
|
|
66
67
|
},
|
|
67
68
|
"peerDependencies": {
|
|
68
69
|
"@earendil-works/pi-ai": "0.84.3",
|
|
@@ -250,6 +250,9 @@ function formatStatus(diagnostics: RuntimeDiagnostics): string {
|
|
|
250
250
|
`Privacy blocked/redacted: ${count(diagnostics.privacy.blockedBlocks)} / ${count(diagnostics.privacy.secretRedactions)}`,
|
|
251
251
|
`Estimator calibration: ${diagnostics.modelAwareness?.calibration.calibrated ? `x${diagnostics.modelAwareness.calibration.appliedRatio.toFixed(3)}` : "collecting/neutral"}`,
|
|
252
252
|
`Calibration samples: ${count(diagnostics.modelAwareness?.calibration.acceptedSamples)} accepted`,
|
|
253
|
+
...(diagnostics.modelAwareness?.drift ? [
|
|
254
|
+
`Token estimate warning: ${diagnostics.modelAwareness.drift.code} (${diagnostics.modelAwareness.drift.severity}, ${count(diagnostics.modelAwareness.drift.sampleCount)} samples)`,
|
|
255
|
+
] : []),
|
|
253
256
|
`Native continuation: ${diagnostics.nativeContinuation.status} (${diagnostics.nativeContinuation.last?.mode ?? "no request"})`,
|
|
254
257
|
`Continuation saved items: ${count(diagnostics.nativeContinuation.last?.omittedInputItems)}`,
|
|
255
258
|
`Quality metrics: ${diagnostics.quality.enabled ? `${count(diagnostics.quality.storedSamples)} sample(s)` : "disabled"}`,
|
|
@@ -365,6 +368,9 @@ function formatManifest(diagnostics: RuntimeDiagnostics): string {
|
|
|
365
368
|
`Artifacts: ${count(manifest.artifacts?.length ?? 0)}`,
|
|
366
369
|
`Privacy: ${manifest.privacy?.enforcement ?? "disabled"}${manifest.privacy ? ` (${manifest.privacy.destination})` : ""}`,
|
|
367
370
|
`Model calibration: ${manifest.modelAwareness?.calibration.calibrated ? `x${manifest.modelAwareness.calibration.appliedRatio.toFixed(3)}` : "neutral/collecting"}`,
|
|
371
|
+
...(manifest.modelAwareness?.drift ? [
|
|
372
|
+
`Token estimate warning: ${manifest.modelAwareness.drift.code} (${manifest.modelAwareness.drift.severity})`,
|
|
373
|
+
] : []),
|
|
368
374
|
`Adaptive tail/hist/project: ${count(manifest.modelAwareness?.adaptive.recentTailTokens)} / ${count(manifest.modelAwareness?.adaptive.maxRetrievedHistoryTokens)} / ${count(manifest.modelAwareness?.adaptive.maxProjectTokens)}`,
|
|
369
375
|
`Continuation: ${manifest.nativeContinuation?.mode ?? "disabled"}; sent/full ${count(manifest.nativeContinuation?.sentInputItems)} / ${count(manifest.nativeContinuation?.fullInputItems)}`,
|
|
370
376
|
"",
|
|
@@ -664,6 +670,12 @@ function formatModelAwareness(diagnostics: RuntimeDiagnostics): string {
|
|
|
664
670
|
`Samples observed/accepted: ${count(calibration.observedSamples)} / ${count(calibration.acceptedSamples)}`,
|
|
665
671
|
`Samples rejected/outliers: ${count(calibration.rejectedSamples)} / ${count(calibration.outlierSamples)}`,
|
|
666
672
|
`Calibration bounds/window: ${calibration.lowerRatioBound.toFixed(2)}-${calibration.upperRatioBound.toFixed(2)} / ${count(calibration.windowSize)}`,
|
|
673
|
+
...(awareness.drift ? [
|
|
674
|
+
`Token estimate warning: ${awareness.drift.code} (${awareness.drift.severity}, ${count(awareness.drift.sampleCount)} samples${awareness.drift.medianRatio === undefined ? "" : `, x${awareness.drift.medianRatio.toFixed(3)}`})`,
|
|
675
|
+
] : []),
|
|
676
|
+
...(awareness.autoTune ? [
|
|
677
|
+
`Budget auto-tuning: ${awareness.autoTune.status} (${count(awareness.autoTune.acceptedSamples)} samples, x${awareness.autoTune.boostFactor.toFixed(3)})`,
|
|
678
|
+
] : []),
|
|
667
679
|
`Adaptive recent tail: ${count(awareness.adaptive.recentTailTokens)} (nominal ${count(awareness.adaptive.nominalRecentTailTokens)})`,
|
|
668
680
|
`Adaptive history retrieval: ${count(awareness.adaptive.maxRetrievedHistoryTokens)} (nominal ${count(awareness.adaptive.nominalRetrievedHistoryTokens)})`,
|
|
669
681
|
`Adaptive project retrieval: ${count(awareness.adaptive.maxProjectTokens)} (nominal ${count(awareness.adaptive.nominalProjectTokens)})`,
|
package/src/extension/runtime.ts
CHANGED
|
@@ -60,11 +60,12 @@ import { calculateContextBudget, type ContextBudget } from "ds4-context-core/cor
|
|
|
60
60
|
import {
|
|
61
61
|
modelProfileKey,
|
|
62
62
|
resolveModelAwareness,
|
|
63
|
+
tokenEstimatorVersion,
|
|
63
64
|
type ResolvedModelAwareness,
|
|
64
65
|
type TokenCalibrationSample,
|
|
65
66
|
} from "ds4-context-core/core/model-awareness";
|
|
66
67
|
import type { ModelDescriptor } from "ds4-context-core/core/model-profile";
|
|
67
|
-
import {
|
|
68
|
+
import { CHARS_ESTIMATOR, type TokenEstimator } from "ds4-context-core/core/token-estimator";
|
|
68
69
|
import type {
|
|
69
70
|
ContextManifest,
|
|
70
71
|
ModelAwarenessManifest,
|
|
@@ -164,6 +165,7 @@ import {
|
|
|
164
165
|
findPiSourceEntryIds,
|
|
165
166
|
fingerprint,
|
|
166
167
|
} from "../pi-adapter/context-observer.ts";
|
|
168
|
+
import { O200K_ESTIMATOR } from "../pi-adapter/bpe-token-estimator.ts";
|
|
167
169
|
import { projectSessionFileMutations } from "../pi-adapter/memory-adapter.ts";
|
|
168
170
|
import { ProjectMemorySynchronizer } from "../pi-adapter/project-memory-sync.ts";
|
|
169
171
|
import { projectRankingLabels } from "../pi-adapter/ranking-adapter.ts";
|
|
@@ -758,15 +760,22 @@ export class Ds4ContextRuntime {
|
|
|
758
760
|
};
|
|
759
761
|
}
|
|
760
762
|
|
|
763
|
+
private estimatorForModel(model: ModelDescriptor): TokenEstimator {
|
|
764
|
+
return tokenEstimatorVersion(model, this.config.modelAwareness) === "o200k-base-v1"
|
|
765
|
+
? O200K_ESTIMATOR : CHARS_ESTIMATOR;
|
|
766
|
+
}
|
|
767
|
+
|
|
761
768
|
private calibrationSamples(model: ModelDescriptor): TokenCalibrationSample[] {
|
|
769
|
+
const version = this.estimatorForModel(model).version;
|
|
762
770
|
if (this.database && this.session?.sessionFile && this.config.diagnostics.storeContextManifest) {
|
|
763
771
|
return this.database.manifests.listCalibrationSamples(
|
|
764
772
|
model.provider,
|
|
765
773
|
model.id,
|
|
766
774
|
this.config.modelAwareness.calibrationWindow,
|
|
775
|
+
version,
|
|
767
776
|
);
|
|
768
777
|
}
|
|
769
|
-
return [...(this.volatileCalibration.get(modelProfileKey(model.provider, model.id)) ?? [])];
|
|
778
|
+
return [...(this.volatileCalibration.get(`${modelProfileKey(model.provider, model.id)}\0${version}`) ?? [])];
|
|
770
779
|
}
|
|
771
780
|
|
|
772
781
|
private switchForModel(model: ModelDescriptor): ModelSwitchManifest {
|
|
@@ -831,6 +840,7 @@ export class Ds4ContextRuntime {
|
|
|
831
840
|
requestsPerTurn: number,
|
|
832
841
|
turnsPerEpoch: number,
|
|
833
842
|
modelKey: string,
|
|
843
|
+
estimator: TokenEstimator,
|
|
834
844
|
): number | undefined {
|
|
835
845
|
const pricing = Ds4ContextRuntime.cachePricing(cost);
|
|
836
846
|
if (!pricing || plan.mode !== "managed") return undefined;
|
|
@@ -839,7 +849,7 @@ export class Ds4ContextRuntime {
|
|
|
839
849
|
let total = fixedTokens;
|
|
840
850
|
for (const message of plan.messages) {
|
|
841
851
|
hashes.push(fingerprint(message));
|
|
842
|
-
const estimate = estimateMessagesTokens([message]);
|
|
852
|
+
const estimate = estimator.estimateMessagesTokens([message]);
|
|
843
853
|
tokens.push(estimate);
|
|
844
854
|
total += estimate;
|
|
845
855
|
}
|
|
@@ -908,6 +918,8 @@ export class Ds4ContextRuntime {
|
|
|
908
918
|
...awareness.calibration,
|
|
909
919
|
cache: { ...awareness.calibration.cache },
|
|
910
920
|
},
|
|
921
|
+
...(awareness.drift ? { drift: awareness.drift } : {}),
|
|
922
|
+
...(awareness.autoTune ? { autoTune: awareness.autoTune } : {}),
|
|
911
923
|
adaptive: { ...awareness.limits },
|
|
912
924
|
switch: modelSwitch,
|
|
913
925
|
};
|
|
@@ -922,7 +934,7 @@ export class Ds4ContextRuntime {
|
|
|
922
934
|
usage: ProviderUsageManifest,
|
|
923
935
|
createdAt: number,
|
|
924
936
|
): void {
|
|
925
|
-
const key = modelProfileKey(manifest.provider, manifest.model)
|
|
937
|
+
const key = `${modelProfileKey(manifest.provider, manifest.model)}\0${manifest.modelAwareness?.calibration.estimator ?? "chars-v1"}`;
|
|
926
938
|
const samples = this.volatileCalibration.get(key) ?? [];
|
|
927
939
|
samples.unshift({
|
|
928
940
|
estimatedTokens: manifest.estimatedInputTokens,
|
|
@@ -980,6 +992,7 @@ export class Ds4ContextRuntime {
|
|
|
980
992
|
}
|
|
981
993
|
const model = snapshotModel(ctx);
|
|
982
994
|
const activeModel = model ? this.resolveActiveModel(model) : undefined;
|
|
995
|
+
const estimator = model ? this.estimatorForModel(model) : CHARS_ESTIMATOR;
|
|
983
996
|
const budget = activeModel?.budget;
|
|
984
997
|
let effectiveEvent = preparedPrivacy.event;
|
|
985
998
|
let artifactReferences = [] as NonNullable<ContextManifest["artifacts"]>;
|
|
@@ -993,8 +1006,8 @@ export class Ds4ContextRuntime {
|
|
|
993
1006
|
preparedPrivacy.messageClassifications,
|
|
994
1007
|
this.config.artifacts.adaptiveBudget && budget ? {
|
|
995
1008
|
inputTokens: budget.activeInputBudget,
|
|
996
|
-
fixedTokens: estimateTextTokens(preparedPrivacy.systemPrompt) + 8
|
|
997
|
-
+ preparedPrivacy.tools.reduce((sum, tool) => sum + estimateObservedToolTokens(tool), 0),
|
|
1009
|
+
fixedTokens: estimator.estimateTextTokens(preparedPrivacy.systemPrompt) + 8
|
|
1010
|
+
+ preparedPrivacy.tools.reduce((sum, tool) => sum + estimateObservedToolTokens(tool, estimator), 0),
|
|
998
1011
|
} : undefined,
|
|
999
1012
|
);
|
|
1000
1013
|
effectiveEvent = { type: "context", messages: transformed.messages };
|
|
@@ -1043,6 +1056,7 @@ export class Ds4ContextRuntime {
|
|
|
1043
1056
|
createdAt: observedAt,
|
|
1044
1057
|
policyVersion: POLICY_VERSION,
|
|
1045
1058
|
plannerVersion: OBSERVER_PLANNER_VERSION,
|
|
1059
|
+
tokenEstimator: estimator,
|
|
1046
1060
|
...(activeModel ? {
|
|
1047
1061
|
profile: activeModel.awareness.profile,
|
|
1048
1062
|
budget: activeModel.budget,
|
|
@@ -1107,6 +1121,7 @@ export class Ds4ContextRuntime {
|
|
|
1107
1121
|
fixedTokens,
|
|
1108
1122
|
budget,
|
|
1109
1123
|
config: effectiveContextConfig,
|
|
1124
|
+
tokenEstimator: estimator,
|
|
1110
1125
|
pinnedMessageIndices,
|
|
1111
1126
|
supplementalMessages: dedupSupplementalMessages,
|
|
1112
1127
|
});
|
|
@@ -1138,13 +1153,14 @@ export class Ds4ContextRuntime {
|
|
|
1138
1153
|
fixedTokens,
|
|
1139
1154
|
budget,
|
|
1140
1155
|
config: effectiveContextConfig,
|
|
1156
|
+
tokenEstimator: estimator,
|
|
1141
1157
|
pinnedMessageIndices,
|
|
1142
1158
|
supplementalMessages: dedupSupplementalMessages,
|
|
1143
1159
|
cacheAwareTailTokens: decision.recentTailTokens,
|
|
1144
1160
|
});
|
|
1145
1161
|
const modelKey = modelProfileKey(model.provider, model.id);
|
|
1146
|
-
const nominalEpoch = this.cacheAwareEpochCost(nominalDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey);
|
|
1147
|
-
const extendedEpoch = this.cacheAwareEpochCost(extendedDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey);
|
|
1162
|
+
const nominalEpoch = this.cacheAwareEpochCost(nominalDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey, estimator);
|
|
1163
|
+
const extendedEpoch = this.cacheAwareEpochCost(extendedDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey, estimator);
|
|
1148
1164
|
this.logger.debug("context.cache_aware_candidate", {
|
|
1149
1165
|
eligible: decision.eligible,
|
|
1150
1166
|
tailExtended: decision.tailExtended,
|
|
@@ -1181,6 +1197,7 @@ export class Ds4ContextRuntime {
|
|
|
1181
1197
|
fixedTokens,
|
|
1182
1198
|
budget,
|
|
1183
1199
|
config: effectiveContextConfig,
|
|
1200
|
+
tokenEstimator: estimator,
|
|
1184
1201
|
pinnedMessageIndices,
|
|
1185
1202
|
supplementalMessages: dedupSupplementalMessages,
|
|
1186
1203
|
cacheAwareTailTokens,
|
|
@@ -1275,7 +1292,7 @@ export class Ds4ContextRuntime {
|
|
|
1275
1292
|
if (sanitized.blockedBlocks > 0) {
|
|
1276
1293
|
privacyExcludedSources.push({
|
|
1277
1294
|
sourceId: supplement.sourceIds[0],
|
|
1278
|
-
tokens: estimateMessagesTokens([supplement.message]),
|
|
1295
|
+
tokens: estimator.estimateMessagesTokens([supplement.message]),
|
|
1279
1296
|
kind: supplement.kind,
|
|
1280
1297
|
classification: sanitized.classification,
|
|
1281
1298
|
score: supplement.score,
|
|
@@ -1328,6 +1345,7 @@ export class Ds4ContextRuntime {
|
|
|
1328
1345
|
fixedTokens,
|
|
1329
1346
|
budget,
|
|
1330
1347
|
config: effectiveContextConfig,
|
|
1348
|
+
tokenEstimator: estimator,
|
|
1331
1349
|
pinnedMessageIndices,
|
|
1332
1350
|
supplementalMessages: rankedSupplementalMessages,
|
|
1333
1351
|
...(cacheAwareTailTokens !== undefined ? { cacheAwareTailTokens } : {}),
|
|
@@ -1408,7 +1426,7 @@ export class Ds4ContextRuntime {
|
|
|
1408
1426
|
const currentTokens: number[] = [];
|
|
1409
1427
|
for (const message of plan.messages) {
|
|
1410
1428
|
currentHashes.push(fingerprint(message));
|
|
1411
|
-
currentTokens.push(estimateMessagesTokens([message]));
|
|
1429
|
+
currentTokens.push(estimator.estimateMessagesTokens([message]));
|
|
1412
1430
|
}
|
|
1413
1431
|
const reusablePrefixTokens = estimateReusablePrefixTokens(
|
|
1414
1432
|
previousHashes,
|
|
@@ -1495,6 +1513,7 @@ export class Ds4ContextRuntime {
|
|
|
1495
1513
|
createdAt: observedAt,
|
|
1496
1514
|
policyVersion: POLICY_VERSION,
|
|
1497
1515
|
plannerVersion: PLANNER_VERSION,
|
|
1516
|
+
tokenEstimator: estimator,
|
|
1498
1517
|
...(activeModel ? {
|
|
1499
1518
|
profile: activeModel.awareness.profile,
|
|
1500
1519
|
budget: activeModel.budget,
|
|
@@ -1542,7 +1561,7 @@ export class Ds4ContextRuntime {
|
|
|
1542
1561
|
estimatedMessageTokens: manifest.composition.messageTokens,
|
|
1543
1562
|
originalMessageCount: manifest.planning?.originalMessageCount ?? historyEvent.messages.length,
|
|
1544
1563
|
originalEstimatedMessageTokens: manifest.planning?.originalMessageTokens
|
|
1545
|
-
?? estimateMessagesTokens(historyEvent.messages),
|
|
1564
|
+
?? estimator.estimateMessagesTokens(historyEvent.messages),
|
|
1546
1565
|
...(usage?.tokens !== null && usage?.tokens !== undefined ? { reportedTokens: usage.tokens } : {}),
|
|
1547
1566
|
...(manifest.planning?.durationMs !== undefined
|
|
1548
1567
|
? { planningDurationMs: manifest.planning.durationMs }
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
import { createRequire } from "node:module";
|
|
2
|
+
import type { Tiktoken } from "js-tiktoken/lite";
|
|
3
|
+
import { createO200kEstimator } from "ds4-context-core/core/bpe-token-estimator";
|
|
4
|
+
|
|
5
|
+
const load = createRequire(import.meta.url);
|
|
6
|
+
|
|
7
|
+
let encoder: Tiktoken | undefined;
|
|
8
|
+
|
|
9
|
+
/**
|
|
10
|
+
* Opt-in OpenAI o200k text estimator. Model-specific serializers, image tokens,
|
|
11
|
+
* reasoning, and provider-specific wrappers are still estimated, not exact.
|
|
12
|
+
* No remote vocabulary download or provider request is performed.
|
|
13
|
+
*/
|
|
14
|
+
export const O200K_ESTIMATOR = createO200kEstimator((text) => {
|
|
15
|
+
if (!encoder) {
|
|
16
|
+
// CJS exports are loaded on first opt-in use, not for every Pi session.
|
|
17
|
+
const { Tiktoken: Encoder } = load("js-tiktoken/lite") as typeof import("js-tiktoken/lite");
|
|
18
|
+
const ranks = load("js-tiktoken/ranks/o200k_base") as {
|
|
19
|
+
pat_str: string; special_tokens: Record<string, number>; bpe_ranks: string;
|
|
20
|
+
};
|
|
21
|
+
encoder = new Encoder(ranks);
|
|
22
|
+
}
|
|
23
|
+
return encoder.encode(text).length;
|
|
24
|
+
});
|
|
@@ -8,7 +8,7 @@ import {
|
|
|
8
8
|
import type { ContextConfig } from "ds4-context-core/config/config";
|
|
9
9
|
import { calculateContextBudget, type ContextBudget } from "ds4-context-core/core/budget-manager";
|
|
10
10
|
import { createModelProfile, type ModelProfile } from "ds4-context-core/core/model-profile";
|
|
11
|
-
import { estimateMessageTokens } from "ds4-context-core/core/token-estimator";
|
|
11
|
+
import { estimateMessageTokens, type TokenEstimator } from "ds4-context-core/core/token-estimator";
|
|
12
12
|
import type {
|
|
13
13
|
ArtifactManifestRef,
|
|
14
14
|
ContextManifest,
|
|
@@ -54,6 +54,7 @@ export interface BuildPiObserverManifestOptions {
|
|
|
54
54
|
profile?: ModelProfile;
|
|
55
55
|
budget?: ContextBudget;
|
|
56
56
|
modelAwareness?: ModelAwarenessManifest;
|
|
57
|
+
tokenEstimator?: TokenEstimator;
|
|
57
58
|
plan?: ManagedContextPlan<PiAgentMessage>;
|
|
58
59
|
projectRevision?: ProjectRevision;
|
|
59
60
|
pins?: readonly PinManifestRef[];
|
|
@@ -386,6 +387,7 @@ export function buildPiObserverManifest(options: BuildPiObserverManifestOptions)
|
|
|
386
387
|
profile,
|
|
387
388
|
budget,
|
|
388
389
|
systemPrompt: options.systemPrompt ?? options.ctx.getSystemPrompt(),
|
|
390
|
+
...(options.tokenEstimator ? { tokenEstimator: options.tokenEstimator } : {}),
|
|
389
391
|
...(options.systemClassification ? { systemClassification: options.systemClassification } : {}),
|
|
390
392
|
...(options.systemPrivacyReason ? { systemPrivacyReason: options.systemPrivacyReason } : {}),
|
|
391
393
|
tools: options.tools ?? activeTools(options.pi),
|