ds4-context-engine 0.3.8 → 0.3.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -16,9 +16,9 @@ bounded active context with provenance
16
16
  Pi provider
17
17
  ```
18
18
 
19
- > **Project status:** The coordinated `0.3.5` release adds bounded compaction optimizations: one validated previous-summary/source update when the complete prompt fits, a dedicated calibrated input budget, up to two concurrent segments, and metadata-only phase timings in `/context compaction`. Canonical history, SQLite schema 15 and runtime contracts remain unchanged; Pi remains pinned to `0.84.3`. See the [0.3.5 release record](docs/releases/0.3.5.md) for validation and publication status.
19
+ > **Project status:** The coordinated `0.3.10` release adds opt-in BPE token estimation, estimator-specific provider calibration and bounded auto-tuning of context category ceilings. `chars-v1` and disabled auto-tuning remain the defaults. The bounded compaction controls from `0.3.9` remain in place; canonical history, SQLite schema 16 and runtime contracts are unchanged. Pi remains pinned to `0.84.3`. See the [0.3.10 release record](docs/releases/0.3.10.md) and [model-awareness validation](docs/MODEL_AWARENESS.md).
20
20
 
21
- **New compaction defaults:** `compaction.directUpdate=true`, `compaction.inputBudget="summary"`, `compaction.maxConcurrentSegments=2`. Existing compaction/master switches still apply. See [latency controls and compatibility](docs/COMPACTION.md#latency-controls) for the legacy-path settings. No real-provider speedup is claimed from mock tests. The five optional editing/reading/artifact/job features introduced in `0.3.4` remain default-off.
21
+ **Current compaction defaults:** `compaction.directUpdate=true`, `compaction.inputBudget="context"`, `compaction.segmentTargetTokens=30000`, `compaction.maxRequestInputTokens=64000`, `compaction.maxOperationInputTokens=2000000`, `compaction.maxConcurrentSegments=2`. Every DS4 provider attempt is bounded by the effective request limit, and the operation limit includes retries; `inputBudget="summary"` remains an explicit throughput-oriented opt-in. Existing compaction/master switches still apply. See [latency controls and compatibility](docs/COMPACTION.md#latency-controls). No real-provider speedup is claimed from mock tests. The five optional editing/reading/artifact/job features introduced in `0.3.4` remain default-off.
22
22
 
23
23
  ## Why DS4
24
24
 
@@ -331,9 +331,11 @@ The following example shows the main configuration groups. Omitted values use th
331
331
  "mode": "hierarchical",
332
332
  "validate": true,
333
333
  "segmentTargetTokens": 30000,
334
+ "maxRequestInputTokens": 64000,
335
+ "maxOperationInputTokens": 2000000,
334
336
  "preserveRecentVerbatim": true,
335
337
  "directUpdate": true,
336
- "inputBudget": "summary",
338
+ "inputBudget": "context",
337
339
  "maxConcurrentSegments": 2
338
340
  },
339
341
  "privacy": {
@@ -512,6 +514,11 @@ scripts package and release-readiness checks
512
514
  - [Roadmap 0.2.0](docs/ROADMAP_0.2.0.md)
513
515
  - [Release process](docs/RELEASING.md)
514
516
  - [0.2.0 release readiness](docs/RELEASE_READINESS_0.2.0.md)
517
+ - [0.3.10 release notes](docs/releases/0.3.10.md)
518
+ - [0.3.9 release notes](docs/releases/0.3.9.md)
519
+ - [0.3.8 release notes](docs/releases/0.3.8.md)
520
+ - [0.3.7 release notes](docs/releases/0.3.7.md)
521
+ - [0.3.6 release notes](docs/releases/0.3.6.md)
515
522
  - [0.3.5 release notes](docs/releases/0.3.5.md)
516
523
  - [0.3.4 release notes](docs/releases/0.3.4.md)
517
524
  - [0.3.3 release notes](docs/releases/0.3.3.md)
@@ -535,7 +542,7 @@ scripts package and release-readiness checks
535
542
 
536
543
  The original M0–M13 roadmap is complete. `ds4-context-core` contains the compiled runtime-neutral implementation. M14 context-quality metrics, M15 rich symbol indexing, M16 hybrid semantic retrieval, M17 cross-session project memory, M18 learned-ranking shadow evaluation, M19's runtime adapter/conformance kit, and M20 opt-in local KV eligibility/replay are implemented on `main`. Learned active ranking remains promotion-gated, Pi reports local KV as unsupported, and static ranking/native completion stay authoritative on every failure.
537
544
 
538
- The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.3.5 adds bounded compaction updates, summary input headroom, concurrent segments and phase timings. The opt-in [anchored editing](docs/ANCHORED_EDITING.md) and [portable agent tools](docs/PORTABLE_AGENT_TOOLS.md) from 0.3.4 remain default-off, without backend rewind, forced sampling or operational KV integration. Confirmation, provenance, Pi fallback and canonical/configuration/SQLite/runtime contracts remain unchanged. The [0.2 readiness record](docs/RELEASE_READINESS_0.2.0.md) remains the compatibility baseline; the lexical planner stays available as the deterministic fallback.
545
+ The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.3.10 adds opt-in BPE estimation and bounded model-budget auto-tuning while retaining the `chars-v1` default. Version 0.3.9 extends the bounded compaction updates, summary input headroom, concurrent segments and phase timings introduced in 0.3.5 with per-request and cumulative operation input limits; 0.3.8 adds indexed FTS key deletion without changing search results. The opt-in [anchored editing](docs/ANCHORED_EDITING.md) and [portable agent tools](docs/PORTABLE_AGENT_TOOLS.md) from 0.3.4 remain default-off, without backend rewind, forced sampling or operational KV integration. Confirmation, provenance, Pi fallback and canonical/configuration/SQLite/runtime contracts remain unchanged. The [0.2 readiness record](docs/RELEASE_READINESS_0.2.0.md) remains the compatibility baseline; the lexical planner stays available as the deterministic fallback.
539
546
 
540
547
  ## Contributing
541
548
 
@@ -14,6 +14,14 @@ Pi normally updates previous summary plus discarded messages in one request. DS4
14
14
  - Expose process-local metadata-only timings for preparation, segment/direct generation, aggregation, graph preparation/persistence and total DS4 hook duration, plus chosen path, effective provider/model, direct prompt size and concurrency. Timings use a monotonic clock, include retries and local validation within generation, and are wall times (not sums of overlapping calls). They are not canonical evidence, are not persisted in JSONL, and do not measure a subsequent native Pi fallback.
15
15
  - Preserve schema-v2 compaction metadata, the existing `task-state` kind, source/classification/validation contracts, Pi cut points, atomic tool exchanges, fresh routing IDs, retry policy, canonical JSONL and all-or-nothing graph preparation. `segmentSummaryId` refers to the update node itself on a direct update; do not invent a segment that was never generated.
16
16
 
17
+ ## Subsequent default revision
18
+
19
+ After operational testing, the default for `compaction.inputBudget` was changed from `summary` to `context`. The original `summary` mode remains available as an explicit throughput-oriented opt-in. The conservative default limits direct updates and indivisible atomic groups to the ordinary active input target, reducing request-size peaks at the cost of potentially more segment and aggregate calls. This revision changes only the default: calibrated hard limits, output headroom, atomicity, validation, and fail-closed behavior remain intact.
20
+
21
+ The hierarchical partitioner applies `min(compaction.segmentTargetTokens, effective request input limit)` to ordinary segments. A single indivisible message or complete tool exchange may exceed the target, but it is isolated and must still fit the effective request input limit.
22
+
23
+ Post-release hardening adds two additive safeguards. `maxRequestInputTokens` caps the estimated prompt size of every direct-update, segment, aggregate, and retry attempt after intersecting it with the selected model input budget. `maxOperationInputTokens` caps the cumulative estimated prompt input reserved across the whole DS4 operation, including retries and concurrent segment attempts. Crossing either limit fails closed to Pi without committing a partial summary graph.
24
+
17
25
  ## Compatibility and validation
18
26
 
19
27
  Configuration is additive; absent fields use the new defaults. The old scheduling path can be compared using `directUpdate=false`, `inputBudget=context`, and `maxConcurrentSegments=1`. The compaction master switch still delegates to Pi when disabled. The original local implementation excluded versioning and publication; the user subsequently authorized the coordinated 0.3.5 release. No dependency upgrade, live database maintenance or Pi upgrade is part of this change.
@@ -8,12 +8,12 @@ DS4 intercepts Pi's `session_before_compact` event but preserves Pi's cut-point
8
8
  2. DS4 maps every source message by exact fingerprint to a canonical branch entry ID.
9
9
  3. Pi's serializer converts the newly discarded span to bounded conversation text; enabled privacy policy sanitizes conversation, previous summary, custom instructions, and file paths for the effective compaction provider (dedicated model when configured and eligible).
10
10
  4. DS4 estimates the **complete** sanitized request against a calibrated summary-specific input budget, including framing, instructions, file inventories and output contract. With `compaction.directUpdate` enabled, a previous summary plus new source that fits is updated and validated in **one call**, producing an immutable `task-state` node. No predecessor needs only the existing one-segment call. Oversized updates fall through to hierarchical planning; no additional source is truncated to force a fit.
11
- 5. The hierarchical path partitions oversized source into ordered contiguous segments. Individual messages are indivisible, and every tool call remains in the same atomic group as all matching results. Up to `compaction.maxConcurrentSegments` independent segment requests run concurrently (default 2). Each summary is validated against only its own sanitized evidence. Identities, source/child order and usage accumulation follow source order, not completion order. Cache retention stays disabled and every attempt has a fresh routing session ID.
11
+ 5. The hierarchical path partitions source into ordered contiguous segments capped at `min(compaction.segmentTargetTokens, requestInputLimitTokens)`, where the effective request limit is also bounded by the selected model input budget. The target is soft only for one indivisible atomic group: an individual message or complete tool call/result exchange may exceed it, but must still fit the effective request input limit and is isolated in its own segment. Up to `compaction.maxConcurrentSegments` independent segment requests run concurrently (default 2). Each summary is validated against only its own sanitized evidence. Identities, source/child order and usage accumulation follow source order, not completion order. Cache retention stays disabled and every attempt has a fresh routing session ID.
12
12
  6. Hierarchical requests recursively aggregate ordered children until one root remains: previous branch summary first, then new segments. Aggregation stays sequential and budget-checked. Direct updates instead link their single node to the predecessor and new canonical source IDs, without generating a synthetic segment or a separate aggregate. DS4-generated IDs, hashes, kinds, and graph levels never become model-visible evidence. A Pi-native predecessor is imported as an explicitly unverified branch node.
13
13
  7. The highest input classification wraps each generated node, then all nodes are persisted atomically as one `prepared` graph batch. Usage includes every returned direct/segment/aggregate request and transport replay. Pi receives only the final root text and still appends exactly one canonical `CompactionEntry` with `fromHook: true`.
14
14
  8. `session_compact` commits all nodes and associates the active root with the Pi entry; failure marks the complete prepared batch `failed`.
15
15
 
16
- Fan-out and fan-in are bounded to 32 segment requests, 64 aggregate requests, and 16 aggregate passes. The DS4 transport replay policy is: `compaction.transport.maxAttempts` (default 3) total attempts and `compaction.transport.baseDelayMs` (default 2000 ms, capped at 60 s) backoff, doubling per attempt, abort-aware. Replay never applies to input, usage, rate, authentication, validation, or output-limit failures. A base prompt, individual message, atomic tool exchange, pair of child summaries, or total operation that cannot fit within those limits fails closed. Any mapping, budget, model, output-limit, validation, abort, or storage error returns `undefined` from the hook, allowing Pi's default compaction to run. On segment failure or cancellation DS4 stops scheduling, aborts siblings and awaits all started workers before fallback; no partial graph is installed. Cancellation is cooperative: accepted provider work may still cost tokens, and a provider that ignores abort can delay settlement.
16
+ Fan-out and fan-in are bounded to 32 segment requests, 64 aggregate requests, and 16 aggregate passes. The DS4 transport replay policy is: `compaction.transport.maxAttempts` (default 4, matching Pi's initial call plus three retries) total attempts and `compaction.transport.baseDelayMs` (default 2000 ms, capped at 60 s) backoff, doubling per attempt, abort-aware. Replay never applies to input, usage, rate, authentication, validation, or output-limit failures. A base prompt, individual message, atomic tool exchange, pair of child summaries, or total operation that cannot fit within those limits fails closed. Any mapping, budget, model, output-limit, validation, abort, or storage error returns `undefined` from the hook, allowing Pi's default compaction to run. On segment failure or cancellation DS4 stops scheduling, aborts siblings and awaits all started workers before fallback; no partial graph is installed. Cancellation is cooperative: accepted provider work may still cost tokens, and a provider that ignores abort can delay settlement.
17
17
 
18
18
  ## Required summary contract
19
19
 
@@ -80,42 +80,49 @@ Semantics:
80
80
 
81
81
  ## Latency controls
82
82
 
83
- The coordinated `0.3.5` release introduces these additive defaults (absent from `0.3.4`):
83
+ The current defaults favor bounded request size while retaining direct updates and bounded parallelism:
84
84
 
85
85
  ```json
86
86
  {
87
87
  "compaction": {
88
88
  "directUpdate": true,
89
- "inputBudget": "summary",
89
+ "inputBudget": "context",
90
+ "segmentTargetTokens": 30000,
91
+ "maxRequestInputTokens": 64000,
92
+ "maxOperationInputTokens": 2000000,
90
93
  "maxConcurrentSegments": 2
91
94
  }
92
95
  }
93
96
  ```
94
97
 
95
- - `directUpdate`: one validated previous-summary plus new-source request when the entire prompt fits. Set false to always retain the segment-then-aggregate route. Validation or provider failures still fall back to Pi, not an unvalidated update.
96
- - `inputBudget`: `summary` uses calibrated `hardInputLimit` rather than the ordinary context fill target (`activeInputBudget`). `context` selects that legacy fill target. Both are additionally capped by `(context window - safety margin - actual summary output cap) / calibration ratio`, rounded down, and never exceed the configured/model hard limit. Existing output reservations remain conservative; ordinary session planning and proactive thresholds are unchanged. This reduces avoidable fragmentation, not a guarantee of provider fit or faster processing for a larger prompt.
98
+ `segmentTargetTokens` is the soft partition target for source segments. `maxRequestInputTokens` is the hard estimated-input cap applied to every direct-update, segment, aggregate, and retry attempt; the effective request limit is the minimum of this value and the safe model input budget. An indivisible message or complete tool exchange above that limit fails closed to Pi's native compaction rather than creating an oversized DS4 request.
99
+
100
+ `maxOperationInputTokens` bounds the sum of estimated prompt tokens reserved immediately before all provider attempts in one DS4 compaction, including transport retries and concurrently scheduled segments. Exceeding it aborts remaining DS4 work and falls back without committing a partial summary graph. Defaults are 64,000 per request and 2,000,000 per operation; both must be positive integers, and the operation limit must be at least the configured request limit.
101
+
102
+ - `directUpdate`: one validated previous-summary plus new-source request when the entire prompt fits the effective request input limit. Set false to always retain the segment-then-aggregate route. Validation or provider failures still fall back to Pi, not an unvalidated update.
103
+ - `inputBudget`: `context` (default) uses the ordinary context fill target (`activeInputBudget`). `summary` is an explicit throughput-oriented opt-in that permits the calibrated `hardInputLimit`. Both are additionally capped by `(context window - safety margin - actual summary output cap) / calibration ratio`, rounded down, and never exceed the configured/model hard limit. Existing output reservations remain conservative; ordinary session planning and proactive thresholds are unchanged. The conservative default reduces request-size peaks but can produce more segments and calls; it is not a guarantee of provider fit or lower total token use.
97
104
  - `maxConcurrentSegments`: integer **1–2**, default 2. Only independent segments overlap, including their retries. Aggregates do not run until their children have completed. Use 1 for sequential execution or providers with restrictive concurrent-request limits. Rate-limit failures are not transport-retried.
98
105
 
99
- For an old-path comparison set `directUpdate=false`, `inputBudget=context`, `maxConcurrentSegments=1`. All features remain behind the existing compaction/master switches. Settings are applied on session load; after upgrading the package or rebuilding a development checkout, fully restart Pi to avoid stale compiled-core modules. No schema migration is required.
106
+ For an old sequential-path comparison set `directUpdate=false`, `inputBudget=context`, `maxConcurrentSegments=1`. To reproduce the `0.3.5` throughput-oriented budget, set `inputBudget=summary`. All features remain behind the existing compaction/master switches. Settings are applied on session load; after upgrading the package or rebuilding a development checkout, fully restart Pi to avoid stale compiled-core modules. No schema migration is required.
100
107
 
101
108
  Mock-provider regression tests verify fewer calls, bounded overlap, exact budget boundaries, validation, privacy, immutable provenance and JSONL rebuild. They do **not** establish real-provider wall-time gains, semantic equivalence of generated summaries, or a guaranteed completion time. See [ADR-061](ADR/061-compaction-latency.md).
102
109
 
103
110
  ## Transport retry policy
104
111
 
105
- Summary requests are replayed only for transport-classified failures (thrown transport errors or `stopReason: "error"` responses whose message matches network/timeout patterns). The DS4 replay policy uses **three total attempts**, not three retries after the initial call, and can be tuned per deployment:
112
+ Summary requests are replayed only for transport-classified failures (thrown transport errors or `stopReason: "error"` responses whose message matches network/timeout patterns). The DS4 replay policy uses **four total attempts** (the initial call plus up to three retries), matching Pi's standard retry count, and can be tuned per deployment:
106
113
 
107
114
  ```json
108
115
  {
109
116
  "compaction": {
110
117
  "transport": {
111
- "maxAttempts": 3,
118
+ "maxAttempts": 4,
112
119
  "baseDelayMs": 2000
113
120
  }
114
121
  }
115
122
  }
116
123
  ```
117
124
 
118
- - `compaction.transport.maxAttempts`: total attempts per direct update, segment or aggregate call, integer 1–10, default 3. With 1, no transport failure is retried.
125
+ - `compaction.transport.maxAttempts`: total attempts per direct update, segment or aggregate call, integer 1–10, default 4. With 1, no transport failure is retried.
119
126
  - `compaction.transport.baseDelayMs`: base backoff before the first replay, integer 0–60000, default 2000. The delay doubles per attempt (2000, 4000, 8000, …) and is capped at 60 s.
120
127
  - Replays use a fresh routing session per attempt; diagnostics expose only stage, failed/next attempt, max attempts, and delay.
121
128
  - Aborts (including during backoff) never trigger replay; non-transport failures are never retried; usage is summed across replayed responses.
@@ -45,7 +45,7 @@ actualInputTokens = input + cacheRead + cacheWrite
45
45
  rawCalibrationRatio = actualInputTokens / estimatedInputTokens
46
46
  ```
47
47
 
48
- The existing scalar SQLite columns are authoritative for persisted usage. `message_end` updates those columns and the calibration sample without reading or rewriting `manifest_json`; repository reads hydrate the usage fields from the scalars. The raw estimate is retained even after calibration so ratios cannot recursively calibrate already-corrected values. The next call reads only the exact provider/model window and applies bounded median/MAD outlier rejection. Error, aborted, missing, zero-usage, duplicate, or uncorrelated responses do not create calibration samples. Every manifest is correlated with at most one assistant response. If a concurrently pruned or oversize-skipped manifest cannot be correlated, DS4 keeps only bounded volatile calibration and does not retry against another row. See [`MODEL_AWARENESS.md`](MODEL_AWARENESS.md).
48
+ The existing scalar SQLite columns are authoritative for persisted usage. `message_end` updates those columns and the calibration sample without reading or rewriting `manifest_json`; repository reads hydrate the usage fields from the scalars. The raw estimate is retained even after calibration so ratios cannot recursively calibrate already-corrected values. The next call reads only the exact provider/model/**estimator-version** window and applies bounded median/MAD outlier rejection. When present, `modelAwareness.drift` and `modelAwareness.autoTune` contain only aggregate warning/tuning metadata (no message contents). Error, aborted, missing, zero-usage, duplicate, or uncorrelated responses do not create calibration samples. Every manifest is correlated with at most one assistant response. If a concurrently pruned or oversize-skipped manifest cannot be correlated, DS4 keeps only bounded volatile calibration and does not retry against another row. See [`MODEL_AWARENESS.md`](MODEL_AWARENESS.md).
49
49
 
50
50
  ## Retention
51
51
 
@@ -76,7 +76,7 @@ With privacy disabled, DS4 returns Pi's original `AgentMessage[]` when:
76
76
  - final estimated input exceeds the hard limit;
77
77
  - an unexpected adapter or planner exception occurs.
78
78
 
79
- Expected fallbacks are recorded in the Context Manifest. With privacy enabled, the fallback baseline is the sanitized native array—not raw Pi messages—and an unexpected privacy failure replaces content/payload fields instead of sending unchecked data. Observer mode disables planning but still enforces enabled privacy policy and records manifests/usage calibration.
79
+ Expected fallbacks are recorded in the Context Manifest. Because fail-open preserves the complete native array, DS4's hard input limit is not guaranteed in fallback when the current request, fixed overhead, or another mandatory atomic group is itself oversized. With privacy enabled, the fallback baseline is the sanitized native array—not raw Pi messages—and an unexpected privacy failure replaces content/payload fields instead of sending unchecked data. Observer mode disables planning but still enforces enabled privacy policy and records manifests/usage calibration.
80
80
 
81
81
  ## Quality measurement
82
82
 
@@ -1,5 +1,7 @@
1
1
  # Advanced Model Awareness
2
2
 
3
+ **Validation status (OpenRouter, synthetic DS4/Pi sessions):** GPT-6-sol, GPT-6-luna and GPT-5.6-terra have each demonstrated multi-turn BPE calibration, a genuinely larger opt-in tail budget, provider usage above the native-window headroom threshold, and withdrawal of that expansion on the next turn in the same session. Historical retrieval and a live tool cycle were verified separately for all three. These bounded experiments do not establish universal tokenizer accuracy, provider-side compaction behavior, or production retrieval quality. `chars-v1` remains the default; BPE and `autoTune` remain opt-in. See the measured runs below.
4
+
3
5
  M11 resolves an independent deterministic planning profile for every exact `provider/model`. Pi's model descriptor remains the default source for context window, output ceiling, reasoning, and image support; explicit DS4 overrides can repair provider metadata or tune category limits without changing canonical session state.
4
6
 
5
7
  ## Profile resolution
@@ -68,10 +70,10 @@ Each successful assistant response is correlated with one pending Context Manife
68
70
 
69
71
  ```text
70
72
  actualInputTokens = input + cacheRead + cacheWrite
71
- ratio = actualInputTokens / raw chars-v1 estimate
73
+ ratio = actualInputTokens / raw selected-estimator estimate
72
74
  ```
73
75
 
74
- Calibration is isolated by exact provider and model ID. The latest configured window is validated and processed deterministically:
76
+ Calibration is isolated by exact provider, model ID **and estimator version**; switching from `chars-v1` to BPE never reuses the old ratios. The latest configured window is validated and processed deterministically:
75
77
 
76
78
  1. reject invalid or zero samples;
77
79
  2. reject ratios outside the configured hard bounds;
@@ -82,7 +84,38 @@ Calibration is isolated by exact provider and model ID. The latest configured wi
82
84
 
83
85
  Until enough samples exist, the multiplier remains `1.0`. The default window is 24 samples, minimum is 3, and hard ratio bounds are 0.5–2.0. Outliers remain counted in metadata but do not influence the applied ratio.
84
86
 
85
- Provider-token capacities are computed first from context window, output reserve, safety margin, and policy ratios. Planner limits are then converted into local-estimator units by dividing by the accepted multiplier. Raw manifest estimates remain uncalibrated so future samples do not feed a corrected estimate back into itself. Adaptive tail/history/project budgets use the same conversion.
87
+ Provider-token capacities are computed first from context window, output reserve, safety margin, and policy ratios. Global hard/soft/preferred input limits are converted into local-estimator units using **at least** a multiplier of 1: observed underestimation may reduce them, but apparent overestimation on short prompts never raises them above nominal model limits. Raw manifest estimates remain uncalibrated so future samples do not feed a corrected estimate back into itself. Adaptive tail/history/project budgets use the accepted multiplier but remain capped by the configured `context.*` maxima even after conversion.
88
+
89
+ ### Optional BPE estimator and measured budget tuning
90
+
91
+ The default remains `chars-v1`. To opt a profile into local OpenAI `o200k_base` BPE text counting (without fetching vocabulary over the network):
92
+
93
+ ```json
94
+ {
95
+ "modelAwareness": {
96
+ "overrides": { "openai/gpt-4o": { "tokenEstimator": "o200k-base-v1" } },
97
+ "autoTune": true
98
+ }
99
+ }
100
+ ```
101
+
102
+ Use BPE only for a model known to share that encoding. The estimator counts text with BPE in the Pi observer, managed planner and fixed prompt/tool estimates; message wrappers, images, reasoning, provider-specific serialization, indexed retrieval hints, and persisted summaries remain estimates. It does **not** make token counts exact or enable provider-side continuation. If the provider's actual input tokens drift persistently from estimates (median ≥1.25 or ≤0.80 after the calibration minimum, or repeated hard-bound outliers), `/context model` and status surface a metadata-only warning; no prompt content is logged.
103
+
104
+ `autoTune` is opt-in and starts neutral. After at least eight accepted calibration samples, it expands *automatic* recent-tail, history and project ceilings by at most 12.5% of their nominal values, only if **every valid provider-usage sample** in the calibration window — including ratio outliers excluded from ratio calibration — remains at or below 60% of the current preferred/hard input target. It never raises a ceiling beyond the configured context cap, never overrides an explicit model-specific category limit, and never changes the provider's window, output reserve, safety margin or hard/soft input limits. Without enough evidence or headroom it stays neutral. Any calibration multiplier is applied after this small expansion; tune decisions and warnings are recorded as metadata in the manifest.
105
+
106
+ ## Validation goal for BPE and auto-tuning
107
+
108
+ **User-requested outcome:** verify long conversations, tools, retrieval, calibration across turns, and auto-tuning safety on the context actually built by DS4/Pi. Single-turn synthetic token probes are preliminary evidence only; they do **not** satisfy this goal or justify declaring the feature complete.
109
+
110
+ The validation must cover, and report results separately for:
111
+
112
+ 1. Long multi-turn sessions, including large context windows, compaction boundaries, and preservation of the current request.
113
+ 2. Tool calls and results (including large results): atomic selection, offload, and the estimated versus observed token count.
114
+ 3. Retrieved history and project context: relevance, budget selection, and whether relevant older material survives planning.
115
+ 4. Consecutive turns on the **same** provider/model and estimator profile: provider usage correlation, accepted/rejected calibration samples, changes to the applied ratio, and isolation when models or estimators change.
116
+ 5. Opt-in auto-tuning: evidence thresholds, configured and hard limits, outliers, insufficient headroom, and no regression in retrieval or context integrity when limits expand.
117
+
118
+ Use local, deterministic integration tests where possible. If real provider calls are needed, require a **new explicit call and input-size limit** before sending anything; the earlier single-turn budgets are exhausted. Report any scenario not verified as *not verified*, rather than treating the synthetic pilot as an end-to-end result. Do not change the default estimator or enable auto-tuning by default on the strength of that pilot.
86
119
 
87
120
  ## Provider cache metrics
88
121
 
@@ -126,6 +159,42 @@ Use:
126
159
  /context status
127
160
  ```
128
161
 
162
+ ## Local validation progress and outstanding provider evidence
163
+
164
+ `tests/integration/model-aware-real-context.test.ts` exercises the Pi `context` and `message_end` hooks over canonical multi-turn JSONL without any network transport. A long branch with a tool call/result verifies atomic selection and retrieval of an older decision under a BPE-managed budget. Successive turns verify that eight correlated usage records are required before opt-in expansion, that estimator/model switches isolate calibration, and that a new near-limit usage record withdraws the expansion even when its ratio is excluded from calibration. `tests/unit/model-aware-estimator.test.ts` also checks explicit overrides, configured caps, and the high-usage outlier regression.
165
+
166
+ **Local integration checks inject usage; they are not independent provider measurements.** A separately authorized, bounded live probe (`scripts/verify-model-aware-real-session.mjs`) used a temporary Pi JSONL session, DS4's actual hooks, synthetic content, `openrouter/openai/gpt-4o-mini`, and Pi-normalized provider usage. Eight successive short calls accumulated calibration samples; on the ninth call the manifest showed eight accepted samples, applied ratio 0.677157, and opt-in `autoTune: expanded`. That ninth call did **not** contain the intended appended long branch: Pi had already constructed its runner, and its manifest estimated only 1,940 input tokens. It is not evidence for long-context tuning.
167
+
168
+ A second, fresh Pi session with the long branch seeded *before* runner construction produced 83 original messages; DS4 selected 21 groups, excluded 21, retrieved the older `cobalt-713` decision, and included both the synthetic tool-call assistant entry and tool result. For that request, the BPE-based manifest estimate was **23,492** input tokens and provider usage was **23,196**. This is **one** live, untuned long-session/tool/retrieval measurement; it does not prove stable drift across lengths or models. The earlier [single-turn pilots](PROVIDER_TOKEN_DRIFT_BENCHMARK.md) remain separate. The two runs together attempted 10 calls and conservatively reserved 217,678 of the authorized 250,000 estimated-input-token limit (per-call ceiling 64,000; 16-call ceiling). All temporary session data was deleted.
169
+
170
+ A subsequent, separately authorized probe (`scripts/verify-model-aware-calibrated-session.mjs`) seeded a 28-turn synthetic history and tool call/result **before starting the same Pi session**. Eight real provider calls calibrated BPE on that long context; the ninth measured **23,926 estimated / 23,486 actual input tokens**, eight accepted samples, `autoTune: expanded`, and retrieval of the old decision with the historical tool call and result included. This verifies expansion on a real long context after calibration in one session, not just across separate sessions. Across this probe **11 calls** reserved **588,052/600,000** estimated input tokens (80,000 per-call limit); no further provider requests were made under that authorization.
171
+
172
+ The attempted higher-occupancy request used **37,552 estimated / 37,464 actual tokens**, below the 60%-of-target withdrawal threshold (**53,760** for this model/configuration). The following turn still reported `expanded`: **withdrawal is not verified live**. That larger turn also excluded the historical tool/retrieval groups, an observed quality limit under that selected context, not proof of an auto-tuning regression. The local Pi compaction-boundary test covers summary preservation under BPE. A separately bounded **one-request** live compaction probe (`scripts/verify-model-aware-compaction-session.mjs`) preseeded the canonical Pi compaction entry before starting Pi; the DS4 managed manifest included its summary and measured **250 estimated / 244 actual** provider input tokens. This used one further call under the third authorization, bringing that budget to **10 calls / 181,069 estimated input tokens reserved**. It checks the provider-bound context *after* a synthetic compaction entry, not live summary generation. At that point, live tool execution and high-usage withdrawal remained unverified; both were tested in the later authorized round below. Model variety and production retrieval quality still remain outside these bounded probes. A third bounded probe (`scripts/verify-model-aware-safety-session.mjs`) reached **60,509 actual input tokens** (>53,760), but only **seven** of eight preceding short-call samples had been accepted; the manifest at the high-usage request showed seven accepted samples and `insufficient-samples`, not an expanded policy. Thus it **does not** test withdrawal. The safety probe attempted **9 calls**, reserving **169,006** estimated input tokens, and stopped without a follow-up call. It now gates the expensive request on eight *accepted* samples and seeds a stable synthetic prefix, but has not been rerun. Its temporary session was deleted; the remaining third-round budget after the separate compaction request cannot fit a new calibration/high-usage/follow-up sequence under the preflight limits. No live tool execution was attempted in that round because the automatic multi-request cycle was not yet safely capped per request. Those results alone did not verify withdrawal or tool execution.
173
+
174
+ A final, separately authorized round verified the outstanding **synthetic** provider-path scenarios on the same OpenRouter/GPT-4o-mini profile. After a stable synthetic prefix, eight usage samples were accepted; the ninth short turn showed `expanded`. The next request reported **60,708 estimated / 60,640 actual** input tokens, above the **53,760** headroom threshold; the following turn reported **60,795 estimated / 60,722 actual** and `no-headroom` (recent-tail limit decreased from **27,699** to **24,564**). `scripts/verify-model-aware-safety-session.mjs` used **11/15** calls and **292,818/350,000** estimated-input tokens reserved (maximum 80,000 per call). This demonstrates withdrawal on measured provider usage, not merely a synthetic injected sample.
175
+
176
+ With the remaining budget, `scripts/verify-model-aware-live-tool.mjs` gated **each** Pi provider request at `ModelRuntime.streamSimple`, disabled cache warming/retry/compaction, and executed a single synthetic tool. The first request included its tool schema and used **105** provider input tokens; the second included the actual tool result and used **139**. Exactly **two** requests and one tool execution occurred. The final fourth-round total was **13/15 calls, 317,119/350,000** estimated-input tokens reserved. Only aggregate counters and booleans were reported; temporary synthetic sessions were removed.
177
+
178
+ These bounded live measurements, together with the canonical Pi JSONL integration tests, cover long context, historical and executed tools, retrieval, between-turn calibration, compaction-boundary context, and auto-tuning withdrawal **for this synthetic OpenRouter/GPT-4o-mini profile**. They do not establish universal accuracy for other providers/models, real private sessions, arbitrary compaction generation, or production retrieval quality.
179
+
180
+ ### Native-window checks for Sol, Luna and Terra — separately authorized
181
+
182
+ The Pi catalog exposed a **1,050,000-token window** for each OpenRouter model `openai/gpt-6-sol`, `openai/gpt-6-luna`, and `openai/gpt-5.6-terra`. One additional authorization limited their *combined* probes to **42 provider attempts, 600,000 estimated input tokens / 2,500,000 controlled characters per request, and 4,800,000 estimated tokens / 15,000,000 controlled characters total**. `scripts/verify-model-aware-triad.mjs` and `scripts/verify-model-aware-triad-followup.mjs` gated **every** Pi request at `ModelRuntime.streamSimple`, including tool continuations; no real sessions, credentials, prompt bodies, responses or raw upstream errors were logged. Temporary canonical Pi JSONL sessions were deleted.
183
+
184
+ The first probe consumed **30 calls / 1,433,445 tokens / 3,679,430 characters reserved**. For *each* model, nine calibration calls in the same seeded long session led to at least eight accepted usage samples. The tenth call measured **36,027 BPE-estimated / 35,503 Pi-normalized actual input tokens**, `autoTune: expanded`, with the historical tool call and result included. **This did not verify retrieval**: the decision was still in the 64k recent tail and therefore had not been fetched by retrieval; the probe correctly stopped before tool execution and the high-occupancy turn. At this native profile, the reported `expanded` status did **not** demonstrate a larger recent-tail limit: its configured 64k cap was already saturated.
185
+
186
+ The second probe seeded **70 turns before Pi session construction**, putting the old decision outside the default recent tail. It consumed nine more calls (one retrieval request and an actual **two-request, one-execution** synthetic tool cycle per model). The decision was both retrieved and included, the historical tool call/result stayed included, and the new tool schema appeared in the first provider request with its actual result in the second. Per-model provider usage for the retrieval request was **63,050 / 63,049 / 63,053** tokens (Sol/Luna/Terra), against BPE estimates **63,699 / 63,698 / 63,702**. All three tool cycles completed. These are measurements of the DS4/Pi managed context, not three single-turn tokenizer prompts.
187
+
188
+ The three remaining authorized requests measured large BPE-managed input, **one per model**. The selected current user turn was included, and Pi reported **381,501 / 381,502 / 381,502** input tokens against manifest estimates **381,593 / 381,594 / 381,594**. The historical call/result were present, but the old decision was **not** selected; the oversized prompt did not ask about that decision, so this is **not** a valid retrieval-quality test. The request generator targeted 462k from the *raw* canonical history, but DS4 excluded older history and the selected provider input remained ~381.5k. Thus all three requests were **below** the native-window auto-tuning withdrawal threshold of **441,000** (60% of the 735k preferred target), and far below the full 1.05M window. They were fresh sessions without eight accepted calibration samples; **none tests `expanded → high provider usage → no-headroom` on these models**. Do not report the triad's auto-tuning safety or category expansion as live-verified at its native window; the deterministic local matrix in `tests/unit/model-aware-estimator.test.ts` covers only injected samples and caps. The entire authorization was used: **42/42 calls**, **3,296,138/4,800,000** estimated input tokens and **9,357,266/15,000,000** controlled characters reserved. No more provider calls are authorized under it. The historical probes now reject `--live` under this exhausted authorization. **After** those measurements, their local high-input sizing was corrected to require at least 470k BPE tokens in the selected current user text (rather than in raw canonical history), with a second gate immediately before provider transport; unit tests cover the discarded-history regression and per-call preflight. This correction was **not** run against a provider and cannot retroactively validate withdrawal.
189
+
190
+ Across these specific synthetic shapes, BPE manifests slightly overestimated provider usage, but no universal correction follows. At this earlier point native-window compaction generation, production retrieval quality and live high-usage auto-tune withdrawal for Sol/Luna/Terra were unverified; the later withdrawal runs below close **only the last of those gaps**. The later probes check actual category growth (not merely `expanded` status), a selected current turn above 470k BPE tokens, usage above 441k, and withdrawal in the same Pi session. Their temporary opt-in category ceilings are 80k/40k/40k because the default native tail ceiling is already equal to its configured maximum of 64k and cannot grow; the probes do not establish expansion under unchanged category defaults. BPE and auto-tuning remain opt-in; no default or provider-storage policy was changed.
191
+
192
+ ### Native-window withdrawal round (later, opt-in categories)
193
+
194
+ With a separate explicit cap of 39 calls, 3,800,000 reserved input tokens and 11,000,000 controlled characters, the same-session Pi/DS4 probe used temporary category ceilings 80k/40k/40k to expose a **real** tail expansion. It ran 34 provider requests (3,375,536 estimated input tokens reserved; 9,845,274 controlled characters). For **GPT-6-sol and GPT-6-luna**, each session accepted at least eight calibration samples; the expanded tail exceeded the same-ratio no-tune baseline (about 72,922 versus 64,819 estimator tokens). A later request used **475,703 provider input tokens** (BPE estimate 475,790), above the 441,000 provider-token headroom threshold. On the next request, `autoTune` was `no-headroom` and the tail equalled its no-tune baseline (64,805); actual provider input was 471,846/471,844. The selected current turn, historical decision, call and result were present before the large request. The large input and follow-up did **not** retain the unrelated old decision/tool group; these measurements do not establish production retrieval quality.
195
+
196
+ For **GPT-5.6-terra**, calibration and real expansion passed in that first round, but its high-input request was blocked **before transport** by the cumulative budget/character caps: the probe had incorrectly budgeted the follow-up as short even though Pi resends the large preceding turn. The temporary session was then disposed. A **separately authorized Terra-only run** used a fresh synthetic Pi/DS4 session and counted *both* large requests. It made 12 provider calls (1,449,079 estimated tokens and 4,310,843 characters reserved, under separate 13-call / 1,900,000-token / 6,000,000-character maxima). After at least eight accepted calibration samples, Terra's tail grew to 72,922 estimator tokens versus a same-ratio untuned baseline of 64,819. The large request used **475,707 provider input tokens** (BPE estimate 475,794), above the 441,000 threshold. On the next turn, provider input was **471,852** (estimate 471,929), `autoTune` was `no-headroom`, and the tail returned to **64,805**, exactly its untuned baseline at that turn's ratio. The historical decision/call/result and current turn were included before the large request; unrelated older groups were excluded during the large request. **This verifies the defined live high-usage expansion/withdrawal scenario for all three models, not a general retrieval or compaction guarantee.** Both probes are now locked against replay; defaults remain unchanged.
197
+
129
198
  ## Performance and tests
130
199
 
131
- `tests/benchmarks/model-awareness.bench.ts` measures a bounded 200-sample calibration analysis and repeated 32k/128k/200k profile resolution. Unit and golden tests cover deterministic tiers, override precedence, robust outlier rejection, cache accounting, and calibrated budgets. Integration tests switch local/remote providers and 32k/128k/200k models while checking profile isolation, privacy re-enforcement, canonical JSONL preservation, SQLite cache metrics, and profile reuse.
200
+ `tests/benchmarks/model-awareness.bench.ts` measures a bounded 200-sample calibration analysis and repeated 32k/128k/200k profile resolution. Unit and golden tests cover deterministic tiers, override precedence, robust outlier rejection, cache accounting, and calibrated budgets. Integration tests switch local/remote providers and 32k/128k/200k models while checking profile isolation, privacy re-enforcement, canonical JSONL preservation, SQLite cache metrics, and profile reuse. For a separately authorized, bounded real-provider measurement of both estimators against Pi SDK usage, see [PROVIDER_TOKEN_DRIFT_BENCHMARK.md](PROVIDER_TOKEN_DRIFT_BENCHMARK.md); it does not substitute for DS4 manifest measurements in a live session.
@@ -0,0 +1,85 @@
1
+ # Live provider token-drift probe (opt-in)
2
+
3
+ This developer-only, **source-checkout-only** probe (the script is not included in the published npm package) compares DS4's raw `chars-v1` and `o200k-base-v1` estimates with **Pi SDK provider usage** for synthetic, single-turn text. It does not start a Pi session, read session JSONL or SQLite, record responses, or change DS4 configuration. It never prints API keys, requests, response bodies or raw upstream error messages. It requires local Pi provider authentication and network access. The probe runs only with `--live`; without that flag it prints the bounded plan.
4
+
5
+ ```bash
6
+ npm run build:core
7
+ node scripts/compare-provider-token-drift.mjs \
8
+ --model openrouter/openai/gpt-4o-mini \
9
+ --model deepseek/deepseek-v4-flash \
10
+ --model openai-codex/gpt-5.4-mini
11
+ # After inspecting the dry-run budget, add --live to make real calls.
12
+ ```
13
+
14
+ The script accepts repeated `--model provider/model-id` and optional `--sizes 512,4096,48000`. **Per invocation**, it refuses more than 12 model calls or 200,000 total input characters (system prompt included), disables SDK request retries, and imposes a 45-second deadline per call. Multiple invocations have separate caps: track their cumulative spend yourself. Output is one metadata-only JSON report on stdout. The SDK's catalog-derived USD cost is an estimate, not a bill; a provider can still charge for failures or retries outside this script's control. If you pipe output to a file, keep it outside the repo unless you intentionally want to publish the aggregate metadata.
15
+
16
+ For each successful call, `actualInputTokens = usage.input + usage.cacheRead + usage.cacheWrite`; the cached fractions are **not** extra tokens on top of that sum. `residualTokens = actual - rawEstimate` (positive means underestimation), and `underestimationPctOfActual = max(0, residual) / actual × 100`. Adjacent slopes use differences across prompt sizes for a single exact provider/model, removing most fixed framing. BPE counts only text; both estimators use the same DS4 message/system overhead. No raw provider payload is available from this probe, so the result does **not** prove the accuracy of DS4's full observer in a live Pi extension chain (tools, images, privacy, cache/continuation, and later extensions differ). Compare `/context model`, `/context tokens`, and manifest `actualInputTokens` versus `estimatedInputTokens` in an ordinary consented Pi session for that second step. Never copy session text or auth files into the report.
17
+
18
+ ## First observed run — 24 September 2026, 10:09–10:13 UTC
19
+
20
+ User-approved cumulative envelope: 12 call attempts and 200,000 input characters. Three invocations totalled **12 attempts and 195,176 planned input characters**: 9 calls/158,382 characters, a single Codex diagnostic call/574 characters, then 2 calls/36,220 characters. Success: 8; failure/missing usage: 4. Synthetic multilingual/code-like repeated text; no private session data. Sizes below are user-message characters; system prompt is included in estimates and the total envelope. These are observations, **not** a representative multi-session cost or quality benchmark.
21
+
22
+ | Provider / requested model | User chars | Real input¹ | `chars-v1` | BPE `o200k-base-v1` | BPE residual (real − estimate) |
23
+ |---|---:|---:|---:|---:|---:|
24
+ | OpenRouter / `openai/gpt-4o-mini` | 512 | 190 | 161 | 196 | −6 |
25
+ | OpenRouter / `openai/gpt-4o-mini` | 4,096 | 1,436 | 1,057 | 1,442 | −6 |
26
+ | OpenRouter / `openai/gpt-4o-mini` | 48,000 | 16,668 | 12,033 | 16,674 | −6 |
27
+ | DeepSeek / `deepseek-v4-flash`² | 512 | 179 | 161 | 196 | −17 |
28
+ | DeepSeek / `deepseek-v4-flash`² | 4,096 | 1,388 | 1,057 | 1,442 | −54 |
29
+ | DeepSeek / `deepseek-v4-flash`² | 48,000 | 16,172 | 12,033 | 16,674 | −502 |
30
+ | OpenRouter / `deepseek/deepseek-v4-flash` | 4,096 | 1,387 | 1,057 | 1,442 | −55 |
31
+ | OpenRouter / `deepseek/deepseek-v4-flash` | 32,000 | 10,785 | 8,033 | 11,124 | −339 |
32
+
33
+ ¹ Pi's normalized provider usage: input + cache read + cache write. Some long requests reported cached reads (up to 1,408 tokens); they were included exactly once. Successful responses produced 1–3 output tokens. ² DeepSeek reported response model `deepseek-flash`, an alias of the requested ID; no equivalence to OpenRouter routing is assumed.
34
+
35
+ - OpenRouter GPT-4o-mini: raw `chars-v1` median actual/estimate **1.358562**; raw BPE median **0.995839**. At 48k chars, chars/4 underestimated by 4,635 tokens (27.81% of actual), while BPE overestimated by 6 tokens. Adjacent BPE slopes were 1.000 and 1.000.
36
+ - Direct DeepSeek: raw chars median **1.313150**; BPE median **0.962552**. At 48k chars, chars/4 underestimated by 4,139 tokens (25.59% of actual), while BPE overestimated by 502 tokens. Adjacent BPE slopes were 0.970305 and 0.970588. Close agreement here does **not** establish that DeepSeek uses OpenAI's tokenizer or justify turning BPE on for that model by default.
37
+ - OpenRouter DeepSeek: two valid samples; raw chars median **1.327396**, BPE median **0.965692**. At 32k chars, chars/4 underestimated by 2,752 tokens; BPE overestimated by 339. Two samples are insufficient to validate a persistent drift warning or tuning.
38
+ - `openai-codex/gpt-5.4-mini`: three initial attempts produced no usable usage; one later 512-character diagnostic returned `request-failed` with sanitized category `other` and no HTTP status. The local Pi auth resolver did return a credential, but the reason for the probe failure remains **unverified**. Do not infer its tokenizer drift, account availability in the interactive TUI, or provider-side token consumption from these calls. Direct `openai/gpt-4o-mini` API auth was not configured in the Pi runtime used by this probe; OpenRouter is a distinct route.
39
+
40
+ **Interpretation:** Three sizes with one call each do not establish statistical reliability or cover large context windows, tools, images, prefixes reused across turns, or output-heavy workflows. `chars-v1` remains the default; existing per-model calibration may compensate after enough **accepted same-profile samples**. The first small sample can be excluded by MAD filtering, so three different-size wire calls need not become three accepted DS4 calibration samples. Keep BPE and `autoTune` opt-in; do not port Hub's fixed 3.5% margin from these measurements. For promotion or automatic budget changes, collect a repeated same-size and mixed-size sample with matching DS4 manifest estimates and actual provider usage under an explicitly authorized, separately bounded run.
41
+
42
+ ## Second observed run — 24 September 2026, 10:35 UTC
43
+
44
+ A separate, explicitly approved envelope covered **12 call attempts and up to 100,000 input characters**. The dry run planned exactly 12 attempts and 99,816 characters: `--sizes 512,16000` for three requested OpenAI models through each of OpenRouter and Codex. The live probe produced six OpenRouter usages and six Codex failures; it did not read session data.
45
+
46
+ | Route / requested model | User chars | Real input¹ | `chars-v1` | BPE `o200k-base-v1` | BPE residual |
47
+ |---|---:|---:|---:|---:|---:|
48
+ | OpenRouter / `openai/gpt-6-sol` | 512 | 189 | 161 | 196 | −7 |
49
+ | OpenRouter / `openai/gpt-6-sol` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
50
+ | OpenRouter / `openai/gpt-6-luna` | 512 | 189 | 161 | 196 | −7 |
51
+ | OpenRouter / `openai/gpt-6-luna` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
52
+ | OpenRouter / `openai/gpt-5.6-terra` | 512 | 189 | 161 | 196 | −7 |
53
+ | OpenRouter / `openai/gpt-5.6-terra` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
54
+
55
+ ¹ Pi-normalized input includes cached tokens once. Each 16,000-character call reported `cacheWriteTokens = 5,561` and zero cache-read tokens. Those were **writes**, not cache hits. Each successful response reported five output tokens. The six equal input counts describe these identical synthetic prompts on this route; they are not proof of identical tokenizers or of how the direct Codex route would count a DS4 session. For each model, the adjacent BPE size slope was exactly 1.000; at 16k characters, raw `chars-v1` underestimated by 1,531 tokens (27.52% of actual), versus BPE overestimating by seven.
56
+
57
+ For `openai-codex/gpt-6-sol`, `gpt-6-luna`, and `gpt-5.6-terra`, **both sizes failed** without provider usage. The sanitized error classifier returned `quota` for all six, with no HTTP status captured. This is evidence of an SDK-path quota/limit error classification, **not** a verified account balance, an HTTP rejection, or a tokenizer measurement. It does not retroactively identify the earlier `gpt-5.4-mini` failure (classified `other`). No further calls were made beyond this run's approved envelope. Two sizes per exact model remain insufficient for DS4 calibration or auto-tuning conclusions.
58
+
59
+ ## Isolated Pi + DS4 managed-context pilot — 24 September 2026
60
+
61
+ For a future, **separately authorized** run: build the core (`npm run build:core`), inspect the planned sizes and character count with `node scripts/compare-ds4-manifest-usage.mjs`, then use `node scripts/compare-ds4-manifest-usage.mjs --live` only after setting a new call/character budget. The default mode is pinned to OpenRouter `openai/gpt-6-sol` and a maximum of 12 calls / 120,000 controlled prompt characters per invocation. `--comparison` instead preflights **both** OpenRouter `openai/gpt-6-luna` and `openai/gpt-5.6-terra`, capped at 24 calls / 240,000 controlled characters combined; `--comparison --live` requires its own authorization. A new authorization is required even when repeating either exact plan.
62
+
63
+ The user separately authorized **up to 12 calls / 120,000 controlled input characters** to OpenRouter `openai/gpt-6-sol` only. `scripts/compare-ds4-manifest-usage.mjs` first checked a local sandbox without provider traffic, then made **12 single-turn calls**, with three repeats at each size (512, 4,096, 12,000 and 20,000 synthetic user characters). Planned synthetic user + fixed system text was **110,568 characters**. This limit counts controlled prompt text, **not** Pi-generated framing or serialized protocol bytes. No existing Pi session history was used; each Pi session and DS4 database was isolated in a temporary directory, with no tools, project content, native continuation, or auto-tuning. DS4's managed `context` hook and its `o200k-base-v1` profile override were active. The script stops on a missing/mismatched manifest or missing usage rather than fabricating a measurement.
64
+
65
+ | Synthetic user chars | First observed DS4 manifest estimate | First actual input¹ | First residual (actual − estimate) | Repeats |
66
+ |---:|---:|---:|---:|---:|
67
+ | 512 | 220 | 209 | −11 | 3 |
68
+ | 4,096 | 1,467 | 1,456 | −11 | 3 |
69
+ | 12,000 | 4,210 | 4,199 | −11 | 3 |
70
+ | 20,000 | 6,983 | 6,972 | −11 | 3 |
71
+
72
+ All **12** calls had matching `managed` manifests and Pi-normalized provider usage, with **exactly −11 tokens** of residual in each call. Token counts varied by about one or two between identically sized repeats, but the residual stayed fixed. The largest relative overestimate was **5.26% of actual input** on the first 512-character call (11/209); at 20,000 characters it was about **0.16%**. `cacheReadTokens` was zero, while longer calls reported mostly cache **writes**; no claim of cache hits or exact billed usage follows from those fields.
73
+
74
+ ¹ `totalInputTokens` in the DS4 manifest (the same Pi-normalized `input + cacheRead + cacheWrite` metric used by model-awareness calibration), not independently obtained provider invoice data. Because the selected estimator was BPE, this pilot did **not** produce an alternative `chars-v1` DS4 manifest for the same wire requests. These are fresh one-turn sessions with one synthetic prompt shape and a maximum of ~7k actual tokens; they do not validate multi-turn prefixes, tool schemas/results, retrieval, real project text, larger windows, other requested models, or production auto-tuning. The 12 samples are independent sessions, **not** 12 accepted samples in one persistent DS4 calibration profile. Keep BPE opt-in and `autoTune` off by default; do not apply an inferred fixed framing correction or Hub's fixed 3.5% drift allowance from this pilot.
75
+
76
+ ## Separately authorized Luna/Terra managed-context comparison — 24 September 2026
77
+
78
+ A subsequent authorization covered **at most 24 calls / 240,000 controlled synthetic input characters combined** on OpenRouter `openai/gpt-6-luna` and `openai/gpt-5.6-terra`; no Codex traffic was authorized. `node scripts/compare-ds4-manifest-usage.mjs --comparison` preflighted **24 calls / 221,136 controlled characters**. `--comparison --live` completed **12 calls per model**, with three fresh, isolated Pi+DS4 managed sessions per model at each synthetic user size (512, 4,096, 12,000, 20,000 characters). No existing session history, project content, tools, auto-tuning, or native continuation entered the requests. All 24 calls had matching manifests and Pi-normalized provider input usage; there were no reported failed rows.
79
+
80
+ | Model | First 512-character DS4 BPE estimate | First actual input¹ | Repeats per size | Residual on all 12 calls |
81
+ |---|---:|---:|---:|---:|
82
+ | OpenRouter / `openai/gpt-6-luna` | 220 | 209 | 3 | −11 tokens |
83
+ | OpenRouter / `openai/gpt-5.6-terra` | 219 | 208 | 3 | −11 tokens |
84
+
85
+ In every size group for both models, each of the three residuals was **−11 tokens**; each group reported zero `cacheReadTokens`. This replicates the small, constant **overestimate** observed for Sol on this specific one-turn synthetic shape. It does **not** establish a provider-independent 11-token correction, tokenizer equivalence across arbitrary text, long-window accuracy, model calibration from a continuous session, or safe budget auto-tuning. The combined 24-call authorization is exhausted. A later, separate authorization covered multi-turn, historical and executed tools, retrieval, and large selected inputs for all three OpenRouter models; see [native-window checks](MODEL_AWARENESS.md#native-window-checks-for-sol-luna-and-terra--separately-authorized) for results and the remaining unverified high-usage auto-tune withdrawal. Codex diagnosis still needs separate approval; keep BPE opt-in and `autoTune` off by default.
@@ -0,0 +1,29 @@
1
+ # Release 0.3.10 — Opt-in BPE estimation and bounded model budget auto-tuning
2
+
3
+ **Coordinated packages:** `ds4-context-core`, `ds4-context-reference-adapter`, and `ds4-context-engine` 0.3.10.
4
+ **Implementation commit:** `0607238`.
5
+
6
+ ## Summary
7
+
8
+ Adds a selectable `o200k-base-v1` text estimator, provider/model/estimator-specific calibration, and opt-in evidence-gated expansion of automatic context-category ceilings. `chars-v1` remains the default estimator; `modelAwareness.autoTune` remains off unless explicitly enabled. The reference adapter and Pi extension continue to depend exactly on the matching core version.
9
+
10
+ ## Changes
11
+
12
+ - Portable core accepts a BPE estimator through its existing runtime-neutral interface; the Pi adapter lazy-loads `js-tiktoken`. Non-text content still uses bounded heuristics.
13
+ - Model profiles can select `tokenEstimator` by exact `provider/model`, provider wildcard, or global override. Calibration histories are isolated by estimator version as well as provider and model.
14
+ - Auto-tuning uses accepted same-profile samples, rejects outliers, respects configured category limits and hard input ceilings, and withdraws expansion after insufficient headroom. Calibrated estimator-unit limits cannot exceed nominal model limits.
15
+ - Manifests and model diagnostics identify estimator version and calibration/auto-tuning decisions. Existing golden compatibility remains anchored to the unchanged default behavior.
16
+
17
+ ## Measured scope and limitations
18
+
19
+ Bounded synthetic multi-turn Pi/DS4 sessions on OpenRouter `openai/gpt-6-sol`, `openai/gpt-6-luna`, and `openai/gpt-5.6-terra` demonstrated calibration, a genuinely larger opt-in recent tail, provider usage above the native-window headroom threshold, and withdrawal to `no-headroom` on the following turn in each model's session. Retrieval and live tool cycles were also exercised separately. See [model-awareness measurements](../MODEL_AWARENESS.md) and [provider-token drift](../PROVIDER_TOKEN_DRIFT_BENCHMARK.md) for methodology, numbers, and bounds.
20
+
21
+ These measurements do not establish production retrieval quality, provider-side compaction behavior, or accuracy for arbitrary models/routes. In particular, equivalent Codex-route measurements did not yield usage and are not claimed as validated. No provider calls are required for this release procedure.
22
+
23
+ ## Compatibility
24
+
25
+ No new defaults are enabled: `chars-v1`, disabled `autoTune`, and disabled DS4 native continuation remain unchanged. No SQLite migration or change to canonical Pi JSONL, privacy consent, native continuation, or the portable runtime-adapter contract is introduced. BPE remains adapter-injected rather than adding a third-party tokenizer dependency to portable core. Restart Pi after upgrading so the new compiled core and session configuration are loaded.
26
+
27
+ ## Validation and publication
28
+
29
+ On Node 26.5.1, the coordinated release passed a clean `npm ci`, `npm run check` (97 Vitest files, 607 tests, TypeScript builds and root typecheck), deterministic `npm run quality:compare`, the `npm run schema:context-persistence` size bound, and `npm run pack:check` in a clean consumer after synchronizing both exported runtime version constants. All typecheck/tests were also rerun after that correction. Dry-run tarball review contained 243 core files, 7 reference-adapter files, and 97 extension files; none included untracked local state. Registry publication and exact-artifact verification are still pending.
@@ -1,7 +1,7 @@
1
1
  # Release 0.3.8 — Indexed FTS key deletes via rowid mappings (schema 16)
2
2
 
3
3
  **Version analyzed:** DS4 Context Engine `0.3.8`
4
- **Commit:** (recorded after pack verification)
4
+ **Commit:** `18f36e1`
5
5
  **Coordinated packages:** `ds4-context-core` 0.3.8, `ds4-context-reference-adapter` 0.3.8, `ds4-context-engine` 0.3.8
6
6
 
7
7
  ## Summary
@@ -0,0 +1,80 @@
1
+ # Release 0.3.9 — Bounded compaction request and operation input
2
+
3
+ **Version analyzed:** DS4 Context Engine `0.3.9`
4
+ **Commit:** `01407cd`
5
+ **Coordinated packages:** `ds4-context-core` 0.3.9, `ds4-context-reference-adapter` 0.3.9, `ds4-context-engine` 0.3.9
6
+
7
+ ## Summary
8
+
9
+ Hardens custom compaction against request-size and total-operation token spikes.
10
+ Every direct-update, segment, aggregate and retry attempt is now bounded by an
11
+ effective estimated-input limit. A second cumulative budget covers the complete
12
+ compaction operation, including concurrent work and transport retries.
13
+
14
+ The default compaction input budget changes from the summary hard limit to the
15
+ ordinary context-fill budget. The former behavior remains available explicitly
16
+ for users who accept larger provider requests in exchange for fewer hierarchy
17
+ levels.
18
+
19
+ ## Changes
20
+
21
+ - Added `compaction.maxRequestInputTokens`, default 64,000, as a hard ceiling on
22
+ estimated input for each compaction provider request. The effective ceiling is
23
+ the minimum of this setting and the selected model-aware compaction budget.
24
+ - Added `compaction.maxOperationInputTokens`, default 2,000,000, for cumulative
25
+ estimated input across the complete operation. Every accepted attempt reserves
26
+ its input before provider dispatch, so retries and concurrent segments count.
27
+ - Configuration validation requires both limits to be positive integers and the
28
+ operation limit to be at least the configured request limit.
29
+ - Changed the default `compaction.inputBudget` from `summary` to `context`.
30
+ Explicit `summary` configurations retain the previous calibrated-hard-limit
31
+ behavior.
32
+ - Changed the default `compaction.transport.maxAttempts` from 3 to 4. Attempts
33
+ remain covered by the cumulative operation budget.
34
+ - Segment partitioning now uses the minimum of `segmentTargetTokens` and the
35
+ effective request limit. An indivisible message or complete tool exchange that
36
+ cannot fit fails closed to Pi's native compaction.
37
+ - `/context compaction` reports the effective request limit and cumulative
38
+ operation input consumed.
39
+ - Direct update, segment and aggregate paths continue to use strict validation,
40
+ privacy filtering, immutable provenance and all-or-nothing graph installation.
41
+
42
+ ## Validation
43
+
44
+ Before publication, the coordinated release passed:
45
+
46
+ - TypeScript builds and typecheck for the root, core and reference adapter;
47
+ - 85 Vitest files with 561 tests;
48
+ - deterministic context-quality comparison;
49
+ - context-persistence schema measurement;
50
+ - clean-consumer package verification for all three tarballs;
51
+ - whitespace/error-marker checks and coordinated manifest inspection.
52
+
53
+ The package verifier also checks the new defaults and both input-budget modes
54
+ from the packed `ds4-context-core` artifact.
55
+
56
+ ## Compatibility and migration
57
+
58
+ - No SQLite migration: schema 16 remains current.
59
+ - No Pi, runtime-adapter, history, persistence-tool or summary-contract version
60
+ changes.
61
+ - Existing configurations that explicitly set `inputBudget` or transport retry
62
+ attempts retain those values.
63
+ - New limit fields are additive. Configurations that omit them receive the
64
+ documented defaults.
65
+ - Pi JSONL remains canonical and custom-compaction failure still delegates to
66
+ Pi's native fallback without installing a partial Summary Graph.
67
+ - A full Pi restart is required after upgrading so compiled core modules and
68
+ session configuration are reloaded.
69
+
70
+ ## Known limits
71
+
72
+ - If the current request, fixed overhead or another mandatory atomic group is
73
+ itself above the planner hard limit, fail-open context planning preserves the
74
+ native context; DS4 cannot transparently split an in-flight Pi tool loop.
75
+ - Provider work accepted before a concurrent sibling fails may still be billed,
76
+ even though DS4 aborts and awaits all started workers before fallback.
77
+ - Limits use DS4's calibrated token estimator. Provider-side tokenization and
78
+ billing remain authoritative.
79
+ - Mock-provider tests establish bounded scheduling and request accounting, not
80
+ real-provider latency or semantic-equivalence guarantees.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ds4-context-engine",
3
- "version": "0.3.8",
3
+ "version": "0.3.10",
4
4
  "description": "Non-destructive, provider-independent context management for Pi.",
5
5
  "type": "module",
6
6
  "license": "MIT",
@@ -62,7 +62,8 @@
62
62
  ]
63
63
  },
64
64
  "dependencies": {
65
- "ds4-context-core": "0.3.8"
65
+ "ds4-context-core": "0.3.10",
66
+ "js-tiktoken": "1.0.21"
66
67
  },
67
68
  "peerDependencies": {
68
69
  "@earendil-works/pi-ai": "0.84.3",
@@ -250,6 +250,9 @@ function formatStatus(diagnostics: RuntimeDiagnostics): string {
250
250
  `Privacy blocked/redacted: ${count(diagnostics.privacy.blockedBlocks)} / ${count(diagnostics.privacy.secretRedactions)}`,
251
251
  `Estimator calibration: ${diagnostics.modelAwareness?.calibration.calibrated ? `x${diagnostics.modelAwareness.calibration.appliedRatio.toFixed(3)}` : "collecting/neutral"}`,
252
252
  `Calibration samples: ${count(diagnostics.modelAwareness?.calibration.acceptedSamples)} accepted`,
253
+ ...(diagnostics.modelAwareness?.drift ? [
254
+ `Token estimate warning: ${diagnostics.modelAwareness.drift.code} (${diagnostics.modelAwareness.drift.severity}, ${count(diagnostics.modelAwareness.drift.sampleCount)} samples)`,
255
+ ] : []),
253
256
  `Native continuation: ${diagnostics.nativeContinuation.status} (${diagnostics.nativeContinuation.last?.mode ?? "no request"})`,
254
257
  `Continuation saved items: ${count(diagnostics.nativeContinuation.last?.omittedInputItems)}`,
255
258
  `Quality metrics: ${diagnostics.quality.enabled ? `${count(diagnostics.quality.storedSamples)} sample(s)` : "disabled"}`,
@@ -365,6 +368,9 @@ function formatManifest(diagnostics: RuntimeDiagnostics): string {
365
368
  `Artifacts: ${count(manifest.artifacts?.length ?? 0)}`,
366
369
  `Privacy: ${manifest.privacy?.enforcement ?? "disabled"}${manifest.privacy ? ` (${manifest.privacy.destination})` : ""}`,
367
370
  `Model calibration: ${manifest.modelAwareness?.calibration.calibrated ? `x${manifest.modelAwareness.calibration.appliedRatio.toFixed(3)}` : "neutral/collecting"}`,
371
+ ...(manifest.modelAwareness?.drift ? [
372
+ `Token estimate warning: ${manifest.modelAwareness.drift.code} (${manifest.modelAwareness.drift.severity})`,
373
+ ] : []),
368
374
  `Adaptive tail/hist/project: ${count(manifest.modelAwareness?.adaptive.recentTailTokens)} / ${count(manifest.modelAwareness?.adaptive.maxRetrievedHistoryTokens)} / ${count(manifest.modelAwareness?.adaptive.maxProjectTokens)}`,
369
375
  `Continuation: ${manifest.nativeContinuation?.mode ?? "disabled"}; sent/full ${count(manifest.nativeContinuation?.sentInputItems)} / ${count(manifest.nativeContinuation?.fullInputItems)}`,
370
376
  "",
@@ -664,6 +670,12 @@ function formatModelAwareness(diagnostics: RuntimeDiagnostics): string {
664
670
  `Samples observed/accepted: ${count(calibration.observedSamples)} / ${count(calibration.acceptedSamples)}`,
665
671
  `Samples rejected/outliers: ${count(calibration.rejectedSamples)} / ${count(calibration.outlierSamples)}`,
666
672
  `Calibration bounds/window: ${calibration.lowerRatioBound.toFixed(2)}-${calibration.upperRatioBound.toFixed(2)} / ${count(calibration.windowSize)}`,
673
+ ...(awareness.drift ? [
674
+ `Token estimate warning: ${awareness.drift.code} (${awareness.drift.severity}, ${count(awareness.drift.sampleCount)} samples${awareness.drift.medianRatio === undefined ? "" : `, x${awareness.drift.medianRatio.toFixed(3)}`})`,
675
+ ] : []),
676
+ ...(awareness.autoTune ? [
677
+ `Budget auto-tuning: ${awareness.autoTune.status} (${count(awareness.autoTune.acceptedSamples)} samples, x${awareness.autoTune.boostFactor.toFixed(3)})`,
678
+ ] : []),
667
679
  `Adaptive recent tail: ${count(awareness.adaptive.recentTailTokens)} (nominal ${count(awareness.adaptive.nominalRecentTailTokens)})`,
668
680
  `Adaptive history retrieval: ${count(awareness.adaptive.maxRetrievedHistoryTokens)} (nominal ${count(awareness.adaptive.nominalRetrievedHistoryTokens)})`,
669
681
  `Adaptive project retrieval: ${count(awareness.adaptive.maxProjectTokens)} (nominal ${count(awareness.adaptive.nominalProjectTokens)})`,
@@ -881,6 +893,8 @@ function formatCompaction(diagnostics: RuntimeDiagnostics, preview: boolean): st
881
893
  `Chosen path: ${compaction.path ?? "n/a"}`,
882
894
  `Input budget mode: ${compaction.inputBudgetMode ?? "n/a"}`,
883
895
  `Input budget: ${count(compaction.inputBudgetTokens)}`,
896
+ `Request input limit: ${count(compaction.requestInputLimitTokens)} (configured max ${count(compaction.maxRequestInputTokens)})`,
897
+ `Operation input: ${count(compaction.operationInputTokens)} / ${count(compaction.maxOperationInputTokens)}`,
884
898
  `Whole-source prompt: ${count(compaction.sourcePromptTokens)}`,
885
899
  `Direct-update prompt: ${count(compaction.directPromptTokens)}`,
886
900
  `Segment concurrency cap: ${count(compaction.maxConcurrentSegments)}`,
@@ -60,11 +60,12 @@ import { calculateContextBudget, type ContextBudget } from "ds4-context-core/cor
60
60
  import {
61
61
  modelProfileKey,
62
62
  resolveModelAwareness,
63
+ tokenEstimatorVersion,
63
64
  type ResolvedModelAwareness,
64
65
  type TokenCalibrationSample,
65
66
  } from "ds4-context-core/core/model-awareness";
66
67
  import type { ModelDescriptor } from "ds4-context-core/core/model-profile";
67
- import { estimateMessagesTokens, estimateTextTokens } from "ds4-context-core/core/token-estimator";
68
+ import { CHARS_ESTIMATOR, type TokenEstimator } from "ds4-context-core/core/token-estimator";
68
69
  import type {
69
70
  ContextManifest,
70
71
  ModelAwarenessManifest,
@@ -164,6 +165,7 @@ import {
164
165
  findPiSourceEntryIds,
165
166
  fingerprint,
166
167
  } from "../pi-adapter/context-observer.ts";
168
+ import { O200K_ESTIMATOR } from "../pi-adapter/bpe-token-estimator.ts";
167
169
  import { projectSessionFileMutations } from "../pi-adapter/memory-adapter.ts";
168
170
  import { ProjectMemorySynchronizer } from "../pi-adapter/project-memory-sync.ts";
169
171
  import { projectRankingLabels } from "../pi-adapter/ranking-adapter.ts";
@@ -758,15 +760,22 @@ export class Ds4ContextRuntime {
758
760
  };
759
761
  }
760
762
 
763
+ private estimatorForModel(model: ModelDescriptor): TokenEstimator {
764
+ return tokenEstimatorVersion(model, this.config.modelAwareness) === "o200k-base-v1"
765
+ ? O200K_ESTIMATOR : CHARS_ESTIMATOR;
766
+ }
767
+
761
768
  private calibrationSamples(model: ModelDescriptor): TokenCalibrationSample[] {
769
+ const version = this.estimatorForModel(model).version;
762
770
  if (this.database && this.session?.sessionFile && this.config.diagnostics.storeContextManifest) {
763
771
  return this.database.manifests.listCalibrationSamples(
764
772
  model.provider,
765
773
  model.id,
766
774
  this.config.modelAwareness.calibrationWindow,
775
+ version,
767
776
  );
768
777
  }
769
- return [...(this.volatileCalibration.get(modelProfileKey(model.provider, model.id)) ?? [])];
778
+ return [...(this.volatileCalibration.get(`${modelProfileKey(model.provider, model.id)}\0${version}`) ?? [])];
770
779
  }
771
780
 
772
781
  private switchForModel(model: ModelDescriptor): ModelSwitchManifest {
@@ -831,6 +840,7 @@ export class Ds4ContextRuntime {
831
840
  requestsPerTurn: number,
832
841
  turnsPerEpoch: number,
833
842
  modelKey: string,
843
+ estimator: TokenEstimator,
834
844
  ): number | undefined {
835
845
  const pricing = Ds4ContextRuntime.cachePricing(cost);
836
846
  if (!pricing || plan.mode !== "managed") return undefined;
@@ -839,7 +849,7 @@ export class Ds4ContextRuntime {
839
849
  let total = fixedTokens;
840
850
  for (const message of plan.messages) {
841
851
  hashes.push(fingerprint(message));
842
- const estimate = estimateMessagesTokens([message]);
852
+ const estimate = estimator.estimateMessagesTokens([message]);
843
853
  tokens.push(estimate);
844
854
  total += estimate;
845
855
  }
@@ -908,6 +918,8 @@ export class Ds4ContextRuntime {
908
918
  ...awareness.calibration,
909
919
  cache: { ...awareness.calibration.cache },
910
920
  },
921
+ ...(awareness.drift ? { drift: awareness.drift } : {}),
922
+ ...(awareness.autoTune ? { autoTune: awareness.autoTune } : {}),
911
923
  adaptive: { ...awareness.limits },
912
924
  switch: modelSwitch,
913
925
  };
@@ -922,7 +934,7 @@ export class Ds4ContextRuntime {
922
934
  usage: ProviderUsageManifest,
923
935
  createdAt: number,
924
936
  ): void {
925
- const key = modelProfileKey(manifest.provider, manifest.model);
937
+ const key = `${modelProfileKey(manifest.provider, manifest.model)}\0${manifest.modelAwareness?.calibration.estimator ?? "chars-v1"}`;
926
938
  const samples = this.volatileCalibration.get(key) ?? [];
927
939
  samples.unshift({
928
940
  estimatedTokens: manifest.estimatedInputTokens,
@@ -980,6 +992,7 @@ export class Ds4ContextRuntime {
980
992
  }
981
993
  const model = snapshotModel(ctx);
982
994
  const activeModel = model ? this.resolveActiveModel(model) : undefined;
995
+ const estimator = model ? this.estimatorForModel(model) : CHARS_ESTIMATOR;
983
996
  const budget = activeModel?.budget;
984
997
  let effectiveEvent = preparedPrivacy.event;
985
998
  let artifactReferences = [] as NonNullable<ContextManifest["artifacts"]>;
@@ -993,8 +1006,8 @@ export class Ds4ContextRuntime {
993
1006
  preparedPrivacy.messageClassifications,
994
1007
  this.config.artifacts.adaptiveBudget && budget ? {
995
1008
  inputTokens: budget.activeInputBudget,
996
- fixedTokens: estimateTextTokens(preparedPrivacy.systemPrompt) + 8
997
- + preparedPrivacy.tools.reduce((sum, tool) => sum + estimateObservedToolTokens(tool), 0),
1009
+ fixedTokens: estimator.estimateTextTokens(preparedPrivacy.systemPrompt) + 8
1010
+ + preparedPrivacy.tools.reduce((sum, tool) => sum + estimateObservedToolTokens(tool, estimator), 0),
998
1011
  } : undefined,
999
1012
  );
1000
1013
  effectiveEvent = { type: "context", messages: transformed.messages };
@@ -1043,6 +1056,7 @@ export class Ds4ContextRuntime {
1043
1056
  createdAt: observedAt,
1044
1057
  policyVersion: POLICY_VERSION,
1045
1058
  plannerVersion: OBSERVER_PLANNER_VERSION,
1059
+ tokenEstimator: estimator,
1046
1060
  ...(activeModel ? {
1047
1061
  profile: activeModel.awareness.profile,
1048
1062
  budget: activeModel.budget,
@@ -1107,6 +1121,7 @@ export class Ds4ContextRuntime {
1107
1121
  fixedTokens,
1108
1122
  budget,
1109
1123
  config: effectiveContextConfig,
1124
+ tokenEstimator: estimator,
1110
1125
  pinnedMessageIndices,
1111
1126
  supplementalMessages: dedupSupplementalMessages,
1112
1127
  });
@@ -1138,13 +1153,14 @@ export class Ds4ContextRuntime {
1138
1153
  fixedTokens,
1139
1154
  budget,
1140
1155
  config: effectiveContextConfig,
1156
+ tokenEstimator: estimator,
1141
1157
  pinnedMessageIndices,
1142
1158
  supplementalMessages: dedupSupplementalMessages,
1143
1159
  cacheAwareTailTokens: decision.recentTailTokens,
1144
1160
  });
1145
1161
  const modelKey = modelProfileKey(model.provider, model.id);
1146
- const nominalEpoch = this.cacheAwareEpochCost(nominalDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey);
1147
- const extendedEpoch = this.cacheAwareEpochCost(extendedDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey);
1162
+ const nominalEpoch = this.cacheAwareEpochCost(nominalDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey, estimator);
1163
+ const extendedEpoch = this.cacheAwareEpochCost(extendedDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey, estimator);
1148
1164
  this.logger.debug("context.cache_aware_candidate", {
1149
1165
  eligible: decision.eligible,
1150
1166
  tailExtended: decision.tailExtended,
@@ -1181,6 +1197,7 @@ export class Ds4ContextRuntime {
1181
1197
  fixedTokens,
1182
1198
  budget,
1183
1199
  config: effectiveContextConfig,
1200
+ tokenEstimator: estimator,
1184
1201
  pinnedMessageIndices,
1185
1202
  supplementalMessages: dedupSupplementalMessages,
1186
1203
  cacheAwareTailTokens,
@@ -1275,7 +1292,7 @@ export class Ds4ContextRuntime {
1275
1292
  if (sanitized.blockedBlocks > 0) {
1276
1293
  privacyExcludedSources.push({
1277
1294
  sourceId: supplement.sourceIds[0],
1278
- tokens: estimateMessagesTokens([supplement.message]),
1295
+ tokens: estimator.estimateMessagesTokens([supplement.message]),
1279
1296
  kind: supplement.kind,
1280
1297
  classification: sanitized.classification,
1281
1298
  score: supplement.score,
@@ -1328,6 +1345,7 @@ export class Ds4ContextRuntime {
1328
1345
  fixedTokens,
1329
1346
  budget,
1330
1347
  config: effectiveContextConfig,
1348
+ tokenEstimator: estimator,
1331
1349
  pinnedMessageIndices,
1332
1350
  supplementalMessages: rankedSupplementalMessages,
1333
1351
  ...(cacheAwareTailTokens !== undefined ? { cacheAwareTailTokens } : {}),
@@ -1408,7 +1426,7 @@ export class Ds4ContextRuntime {
1408
1426
  const currentTokens: number[] = [];
1409
1427
  for (const message of plan.messages) {
1410
1428
  currentHashes.push(fingerprint(message));
1411
- currentTokens.push(estimateMessagesTokens([message]));
1429
+ currentTokens.push(estimator.estimateMessagesTokens([message]));
1412
1430
  }
1413
1431
  const reusablePrefixTokens = estimateReusablePrefixTokens(
1414
1432
  previousHashes,
@@ -1495,6 +1513,7 @@ export class Ds4ContextRuntime {
1495
1513
  createdAt: observedAt,
1496
1514
  policyVersion: POLICY_VERSION,
1497
1515
  plannerVersion: PLANNER_VERSION,
1516
+ tokenEstimator: estimator,
1498
1517
  ...(activeModel ? {
1499
1518
  profile: activeModel.awareness.profile,
1500
1519
  budget: activeModel.budget,
@@ -1542,7 +1561,7 @@ export class Ds4ContextRuntime {
1542
1561
  estimatedMessageTokens: manifest.composition.messageTokens,
1543
1562
  originalMessageCount: manifest.planning?.originalMessageCount ?? historyEvent.messages.length,
1544
1563
  originalEstimatedMessageTokens: manifest.planning?.originalMessageTokens
1545
- ?? estimateMessagesTokens(historyEvent.messages),
1564
+ ?? estimator.estimateMessagesTokens(historyEvent.messages),
1546
1565
  ...(usage?.tokens !== null && usage?.tokens !== undefined ? { reportedTokens: usage.tokens } : {}),
1547
1566
  ...(manifest.planning?.durationMs !== undefined
1548
1567
  ? { planningDurationMs: manifest.planning.durationMs }
@@ -0,0 +1,24 @@
1
+ import { createRequire } from "node:module";
2
+ import type { Tiktoken } from "js-tiktoken/lite";
3
+ import { createO200kEstimator } from "ds4-context-core/core/bpe-token-estimator";
4
+
5
+ const load = createRequire(import.meta.url);
6
+
7
+ let encoder: Tiktoken | undefined;
8
+
9
+ /**
10
+ * Opt-in OpenAI o200k text estimator. Model-specific serializers, image tokens,
11
+ * reasoning, and provider-specific wrappers are still estimated, not exact.
12
+ * No remote vocabulary download or provider request is performed.
13
+ */
14
+ export const O200K_ESTIMATOR = createO200kEstimator((text) => {
15
+ if (!encoder) {
16
+ // CJS exports are loaded on first opt-in use, not for every Pi session.
17
+ const { Tiktoken: Encoder } = load("js-tiktoken/lite") as typeof import("js-tiktoken/lite");
18
+ const ranks = load("js-tiktoken/ranks/o200k_base") as {
19
+ pat_str: string; special_tokens: Record<string, number>; bpe_ranks: string;
20
+ };
21
+ encoder = new Encoder(ranks);
22
+ }
23
+ return encoder.encode(text).length;
24
+ });
@@ -9,7 +9,12 @@ import type {
9
9
  SessionEntry,
10
10
  } from "@earendil-works/pi-coding-agent";
11
11
  import type { Api, Model } from "@earendil-works/pi-ai";
12
- import type { CompactionThinkingLevel, Ds4ContextConfig } from "ds4-context-core/config/config";
12
+ import {
13
+ DEFAULT_COMPACTION_MAX_OPERATION_INPUT_TOKENS,
14
+ DEFAULT_COMPACTION_MAX_REQUEST_INPUT_TOKENS,
15
+ type CompactionThinkingLevel,
16
+ type Ds4ContextConfig,
17
+ } from "ds4-context-core/config/config";
13
18
  import { calculateContextBudget, type ContextBudget } from "ds4-context-core/core/budget-manager";
14
19
  import { createModelProfile, type ModelDescriptor } from "ds4-context-core/core/model-profile";
15
20
  import type { ContextManifest } from "ds4-context-core/manifest/context-manifest";
@@ -80,6 +85,8 @@ export interface CompactionDiagnostics {
80
85
  validate: boolean;
81
86
  preserveRecentVerbatim: boolean;
82
87
  segmentTargetTokens: number;
88
+ maxRequestInputTokens: number;
89
+ maxOperationInputTokens: number;
83
90
  phase: CompactionPhase;
84
91
  trigger?: CompactionTrigger;
85
92
  summaryId?: string;
@@ -91,6 +98,8 @@ export interface CompactionDiagnostics {
91
98
  completedAt?: number;
92
99
  lastError?: string;
93
100
  inputBudgetTokens?: number;
101
+ requestInputLimitTokens?: number;
102
+ operationInputTokens?: number;
94
103
  sourcePromptTokens?: number;
95
104
  segmentCount?: number;
96
105
  aggregateCalls?: number;
@@ -189,9 +198,20 @@ interface CompactionCoordinatorDependencies {
189
198
 
190
199
  type MutableCompactionState = Omit<
191
200
  CompactionDiagnostics,
192
- "enabled" | "validate" | "preserveRecentVerbatim" | "segmentTargetTokens" | "proactiveEligible"
201
+ | "enabled"
202
+ | "validate"
203
+ | "preserveRecentVerbatim"
204
+ | "segmentTargetTokens"
205
+ | "maxRequestInputTokens"
206
+ | "maxOperationInputTokens"
207
+ | "proactiveEligible"
193
208
  >;
194
209
 
210
+ interface CompactionOperationInputBudget {
211
+ limitTokens: number;
212
+ usedTokens: number;
213
+ }
214
+
195
215
  function classifiedSummary(content: string, classification: PrivacyClassification): string {
196
216
  return classification === "normal"
197
217
  ? content
@@ -260,6 +280,14 @@ export class CompactionCoordinator {
260
280
  }
261
281
  };
262
282
  const maxConcurrentSegments = config.compaction.maxConcurrentSegments ?? 2;
283
+ const maxRequestInputTokens = config.compaction.maxRequestInputTokens
284
+ ?? DEFAULT_COMPACTION_MAX_REQUEST_INPUT_TOKENS;
285
+ const maxOperationInputTokens = config.compaction.maxOperationInputTokens
286
+ ?? DEFAULT_COMPACTION_MAX_OPERATION_INPUT_TOKENS;
287
+ const operationInputBudget: CompactionOperationInputBudget = {
288
+ limitTokens: maxOperationInputTokens,
289
+ usedTokens: 0,
290
+ };
263
291
  const trigger: CompactionTrigger = this.proactiveRequested ? "proactive" : event.reason;
264
292
  this.state = {
265
293
  phase: "generating",
@@ -272,18 +300,20 @@ export class CompactionCoordinator {
272
300
  summaryCalls: 0,
273
301
  provider: model.provider,
274
302
  model: model.id,
275
- inputBudgetMode: config.compaction.inputBudget ?? "summary",
303
+ inputBudgetMode: config.compaction.inputBudget ?? "context",
276
304
  maxConcurrentSegments,
305
+ operationInputTokens: 0,
277
306
  timings,
278
307
  };
279
308
 
280
309
  try {
281
- const { source, inputBudgetTokens, wholePlan, directPlan, segmentPlans } = await measure("preparationMs", () => {
310
+ const { source, inputBudgetTokens, requestInputLimitTokens, wholePlan, directPlan, segmentPlans } = await measure("preparationMs", () => {
282
311
  if (event.signal.aborted) throw new Error("Compaction summary generation aborted");
283
312
  this.dependencies.syncSessionIndex(ctx);
284
313
  const source = prepareCompactionSource(event);
285
314
  const inputBudgetTokens = this.inputBudgetTokens(model);
286
315
  if (inputBudgetTokens <= 0) throw new Error("Active model has no safe compaction input budget");
316
+ const requestInputLimitTokens = Math.min(inputBudgetTokens, maxRequestInputTokens);
287
317
  const wholePlan = this.buildSegmentPlan(source, event, model.provider);
288
318
  const update = (config.compaction.directUpdate ?? true) && source.previousSummary
289
319
  ? this.buildSegmentPlan({
@@ -292,18 +322,21 @@ export class CompactionCoordinator {
292
322
  segmentModifiedFiles: source.modifiedFiles,
293
323
  }, event, model.provider, source.previousSummary)
294
324
  : undefined;
295
- const directPlan = update && update.promptTokens <= inputBudgetTokens ? update : undefined;
325
+ const directPlan = update && update.promptTokens <= requestInputLimitTokens ? update : undefined;
296
326
  this.state = {
297
327
  ...this.state,
298
328
  path: directPlan ? "direct-update" : "hierarchical",
299
329
  sourceEntries: source.sourceEntryIds.length,
300
330
  inputBudgetTokens,
331
+ requestInputLimitTokens,
301
332
  sourcePromptTokens: wholePlan.promptTokens,
302
333
  ...(update ? { directPromptTokens: update.promptTokens } : {}),
303
334
  };
304
- const segmentPlans = directPlan ? [] : this.partitionSegmentPlans(source, wholePlan, event, model.provider, inputBudgetTokens);
335
+ const segmentPlans = directPlan
336
+ ? []
337
+ : this.partitionSegmentPlans(source, wholePlan, event, model.provider, requestInputLimitTokens);
305
338
  this.state.segmentCount = segmentPlans.length;
306
- return { source, inputBudgetTokens, wholePlan, directPlan, segmentPlans };
339
+ return { source, inputBudgetTokens, requestInputLimitTokens, wholePlan, directPlan, segmentPlans };
307
340
  });
308
341
 
309
342
  const usedIds = new Set(this.graphRecords.keys());
@@ -352,6 +385,9 @@ export class CompactionCoordinator {
352
385
  ctx,
353
386
  model,
354
387
  thinking: config.compaction.summary?.thinking,
388
+ promptTokens: plan.promptTokens,
389
+ requestInputLimitTokens,
390
+ operationInputBudget,
355
391
  }),
356
392
  ));
357
393
  const generatedNodes: EmbeddedSummaryNode[] = results.map((generated, index) => {
@@ -385,7 +421,8 @@ export class CompactionCoordinator {
385
421
  model: model.id,
386
422
  readFiles: source.readFiles,
387
423
  modifiedFiles: source.modifiedFiles,
388
- inputBudgetTokens,
424
+ requestInputLimitTokens,
425
+ operationInputBudget,
389
426
  nextId,
390
427
  createdNodes,
391
428
  usages,
@@ -698,6 +735,10 @@ export class CompactionCoordinator {
698
735
  validate: config.compaction.validate,
699
736
  preserveRecentVerbatim: config.compaction.preserveRecentVerbatim,
700
737
  segmentTargetTokens: config.compaction.segmentTargetTokens,
738
+ maxRequestInputTokens: config.compaction.maxRequestInputTokens
739
+ ?? DEFAULT_COMPACTION_MAX_REQUEST_INPUT_TOKENS,
740
+ maxOperationInputTokens: config.compaction.maxOperationInputTokens
741
+ ?? DEFAULT_COMPACTION_MAX_OPERATION_INPUT_TOKENS,
701
742
  ...this.state,
702
743
  ...(contextTokens !== undefined ? { contextTokens } : {}),
703
744
  ...(budget ? { softLimitTokens: this.providerSoftLimit(budget) } : {}),
@@ -800,7 +841,7 @@ export class CompactionCoordinator {
800
841
  ),
801
842
  };
802
843
  const maxOutputTokens = Math.max(1, Math.min(config.context.maxSummaryTokens, model.maxTokens ?? config.context.maxSummaryTokens));
803
- return compactionInputBudget(resolved.budget, maxOutputTokens, config.compaction.inputBudget ?? "summary");
844
+ return compactionInputBudget(resolved.budget, maxOutputTokens, config.compaction.inputBudget ?? "context");
804
845
  }
805
846
 
806
847
  private classify(text: string, provider: string): {
@@ -870,9 +911,13 @@ export class CompactionCoordinator {
870
911
  wholePlan: SegmentGenerationPlan,
871
912
  event: SessionBeforeCompactEvent,
872
913
  provider: string,
873
- inputBudgetTokens: number,
914
+ requestInputLimitTokens: number,
874
915
  ): SegmentGenerationPlan[] {
875
- if (wholePlan.promptTokens <= inputBudgetTokens) return [wholePlan];
916
+ const segmentTargetTokens = Math.min(
917
+ requestInputLimitTokens,
918
+ this.dependencies.config.compaction.segmentTargetTokens,
919
+ );
920
+ if (wholePlan.promptTokens <= segmentTargetTokens) return [wholePlan];
876
921
 
877
922
  const groups = buildCompactionAtomicGroups(source.messages);
878
923
  const plans: SegmentGenerationPlan[] = [];
@@ -886,7 +931,7 @@ export class CompactionCoordinator {
886
931
  event,
887
932
  provider,
888
933
  );
889
- if (candidatePlan.promptTokens <= inputBudgetTokens) {
934
+ if (candidatePlan.promptTokens <= segmentTargetTokens) {
890
935
  currentIndices = candidateIndices;
891
936
  currentPlan = candidatePlan;
892
937
  continue;
@@ -903,9 +948,9 @@ export class CompactionCoordinator {
903
948
  event,
904
949
  provider,
905
950
  );
906
- if (atomicPlan.promptTokens > inputBudgetTokens) {
951
+ if (atomicPlan.promptTokens > requestInputLimitTokens) {
907
952
  throw new Error(
908
- `Compaction source contains an indivisible atomic group above the model input budget (promptTokens=${atomicPlan.promptTokens}; inputBudgetTokens=${inputBudgetTokens})`,
953
+ `Compaction source contains an indivisible atomic group above the request input limit (promptTokens=${atomicPlan.promptTokens}; requestInputLimitTokens=${requestInputLimitTokens})`,
909
954
  );
910
955
  }
911
956
  currentIndices = [...group.messageIndices];
@@ -974,7 +1019,8 @@ export class CompactionCoordinator {
974
1019
  modelObject: Model<Api>;
975
1020
  readFiles: readonly string[];
976
1021
  modifiedFiles: readonly string[];
977
- inputBudgetTokens: number;
1022
+ requestInputLimitTokens: number;
1023
+ operationInputBudget: CompactionOperationInputBudget;
978
1024
  nextId: () => string;
979
1025
  createdNodes: EmbeddedSummaryNode[];
980
1026
  usages: GeneratedSummary["usage"][];
@@ -1003,13 +1049,13 @@ export class CompactionCoordinator {
1003
1049
  input.readFiles,
1004
1050
  input.modifiedFiles,
1005
1051
  );
1006
- if (candidatePlan.promptTokens <= input.inputBudgetTokens) {
1052
+ if (candidatePlan.promptTokens <= input.requestInputLimitTokens) {
1007
1053
  current = candidate;
1008
1054
  continue;
1009
1055
  }
1010
1056
  if (current.length === 1) {
1011
1057
  throw new Error(
1012
- `Compaction child summaries cannot be aggregated within the model input budget (inputBudgetTokens=${input.inputBudgetTokens})`,
1058
+ `Compaction child summaries cannot be aggregated within the request input limit (requestInputLimitTokens=${input.requestInputLimitTokens})`,
1013
1059
  );
1014
1060
  }
1015
1061
  batches.push(current);
@@ -1034,7 +1080,7 @@ export class CompactionCoordinator {
1034
1080
  input.readFiles,
1035
1081
  input.modifiedFiles,
1036
1082
  );
1037
- if (plan.promptTokens > input.inputBudgetTokens) {
1083
+ if (plan.promptTokens > input.requestInputLimitTokens) {
1038
1084
  throw new Error("Compaction aggregate prompt exceeded its preflight input budget");
1039
1085
  }
1040
1086
  this.state.aggregateCalls = aggregateCalls;
@@ -1048,6 +1094,9 @@ export class CompactionCoordinator {
1048
1094
  ctx: input.ctx,
1049
1095
  model: input.modelObject,
1050
1096
  thinking: this.dependencies.config.compaction.summary?.thinking,
1097
+ promptTokens: plan.promptTokens,
1098
+ requestInputLimitTokens: input.requestInputLimitTokens,
1099
+ operationInputBudget: input.operationInputBudget,
1051
1100
  });
1052
1101
  input.usages.push(generated.usage);
1053
1102
  const aggregateNode: EmbeddedSummaryNode = {
@@ -1088,7 +1137,15 @@ export class CompactionCoordinator {
1088
1137
  ctx: ExtensionContext;
1089
1138
  model: Model<Api>;
1090
1139
  thinking?: CompactionThinkingLevel;
1140
+ promptTokens: number;
1141
+ requestInputLimitTokens: number;
1142
+ operationInputBudget: CompactionOperationInputBudget;
1091
1143
  }): Promise<GeneratedSummary> {
1144
+ if (input.promptTokens > input.requestInputLimitTokens) {
1145
+ throw new Error(
1146
+ `Compaction ${input.stage} prompt exceeded the request input limit (promptTokens=${input.promptTokens}; requestInputLimitTokens=${input.requestInputLimitTokens})`,
1147
+ );
1148
+ }
1092
1149
  this.state.summaryCalls = (this.state.summaryCalls ?? 0) + 1;
1093
1150
  return generateValidatedSummary({
1094
1151
  ...input,
@@ -1096,6 +1153,16 @@ export class CompactionCoordinator {
1096
1153
  maxSummaryTokens: this.dependencies.config.context.maxSummaryTokens,
1097
1154
  transport: this.dependencies.config.compaction.transport,
1098
1155
  now: this.dependencies.now,
1156
+ onAttempt: (diagnostic) => {
1157
+ const nextInputTokens = input.operationInputBudget.usedTokens + input.promptTokens;
1158
+ if (nextInputTokens > input.operationInputBudget.limitTokens) {
1159
+ throw new Error(
1160
+ `Compaction operation input limit exceeded before ${diagnostic.stage} attempt ${diagnostic.attempt} (nextInputTokens=${nextInputTokens}; maxOperationInputTokens=${input.operationInputBudget.limitTokens})`,
1161
+ );
1162
+ }
1163
+ input.operationInputBudget.usedTokens = nextInputTokens;
1164
+ this.state.operationInputTokens = nextInputTokens;
1165
+ },
1099
1166
  onTransportRetry: (diagnostic) => {
1100
1167
  this.state.transportRetries = (this.state.transportRetries ?? 0) + 1;
1101
1168
  this.dependencies.logger.debug("compaction.transport_retry", { ...diagnostic });
@@ -1205,6 +1272,10 @@ export function defaultCompactionDiagnostics(config: Ds4ContextConfig): Compacti
1205
1272
  validate: config.compaction.validate,
1206
1273
  preserveRecentVerbatim: config.compaction.preserveRecentVerbatim,
1207
1274
  segmentTargetTokens: config.compaction.segmentTargetTokens,
1275
+ maxRequestInputTokens: config.compaction.maxRequestInputTokens
1276
+ ?? DEFAULT_COMPACTION_MAX_REQUEST_INPUT_TOKENS,
1277
+ maxOperationInputTokens: config.compaction.maxOperationInputTokens
1278
+ ?? DEFAULT_COMPACTION_MAX_OPERATION_INPUT_TOKENS,
1208
1279
  phase: "idle",
1209
1280
  proactiveEligible: false,
1210
1281
  };
@@ -8,7 +8,7 @@ import {
8
8
  import type { ContextConfig } from "ds4-context-core/config/config";
9
9
  import { calculateContextBudget, type ContextBudget } from "ds4-context-core/core/budget-manager";
10
10
  import { createModelProfile, type ModelProfile } from "ds4-context-core/core/model-profile";
11
- import { estimateMessageTokens } from "ds4-context-core/core/token-estimator";
11
+ import { estimateMessageTokens, type TokenEstimator } from "ds4-context-core/core/token-estimator";
12
12
  import type {
13
13
  ArtifactManifestRef,
14
14
  ContextManifest,
@@ -54,6 +54,7 @@ export interface BuildPiObserverManifestOptions {
54
54
  profile?: ModelProfile;
55
55
  budget?: ContextBudget;
56
56
  modelAwareness?: ModelAwarenessManifest;
57
+ tokenEstimator?: TokenEstimator;
57
58
  plan?: ManagedContextPlan<PiAgentMessage>;
58
59
  projectRevision?: ProjectRevision;
59
60
  pins?: readonly PinManifestRef[];
@@ -386,6 +387,7 @@ export function buildPiObserverManifest(options: BuildPiObserverManifestOptions)
386
387
  profile,
387
388
  budget,
388
389
  systemPrompt: options.systemPrompt ?? options.ctx.getSystemPrompt(),
390
+ ...(options.tokenEstimator ? { tokenEstimator: options.tokenEstimator } : {}),
389
391
  ...(options.systemClassification ? { systemClassification: options.systemClassification } : {}),
390
392
  ...(options.systemPrivacyReason ? { systemPrivacyReason: options.systemPrivacyReason } : {}),
391
393
  tools: options.tools ?? activeTools(options.pi),
@@ -13,16 +13,16 @@ import {
13
13
  type SummaryValidationResult,
14
14
  } from "ds4-context-core/compaction/summary-contract";
15
15
 
16
- export const DEFAULT_COMPACTION_TRANSPORT_MAX_ATTEMPTS = 3;
16
+ export const DEFAULT_COMPACTION_TRANSPORT_MAX_ATTEMPTS = 4;
17
17
  export const DEFAULT_COMPACTION_TRANSPORT_BASE_DELAY_MS = 2000;
18
18
  export const COMPACTION_TRANSPORT_MAX_DELAY_MS = 60_000;
19
19
 
20
20
  /**
21
- * Transport retry policy for compaction summary requests: three total attempts
22
- * (not three retries), 2000 ms base delay, exponential backoff, abort-aware.
21
+ * Transport retry policy for compaction summary requests: four total attempts
22
+ * (the initial call plus three retries), 2000 ms base delay, exponential backoff, abort-aware.
23
23
  */
24
24
  export interface CompactionTransportPolicy {
25
- /** Total attempts for transport-classified failures. Default: 3. */
25
+ /** Total attempts for transport-classified failures. Default: 4. */
26
26
  maxAttempts?: number;
27
27
  /** Base backoff delay in ms, doubled per attempt. Default: 2000. */
28
28
  baseDelayMs?: number;
@@ -54,6 +54,12 @@ export interface CompactionTransportRetryDiagnostic {
54
54
  delayMs: number;
55
55
  }
56
56
 
57
+ export interface CompactionAttemptDiagnostic {
58
+ stage: "segment" | "aggregate" | "update";
59
+ attempt: number;
60
+ maxAttempts: number;
61
+ }
62
+
57
63
  export interface GenerateValidatedSummaryInput {
58
64
  stage: "segment" | "aggregate" | "update";
59
65
  prompt: string;
@@ -68,9 +74,11 @@ export interface GenerateValidatedSummaryInput {
68
74
  model?: Model<Api>;
69
75
  /** Reasoning level for the summary request; `off` (default) keeps the pre-existing request shape. */
70
76
  thinking?: CompactionThinkingLevel;
71
- /** Transport-only retry policy; three total attempts by default. */
77
+ /** Transport-only retry policy; four total attempts by default. */
72
78
  transport?: CompactionTransportPolicy;
73
79
  now: () => number;
80
+ /** Called synchronously before every provider attempt, including retries. */
81
+ onAttempt?: (diagnostic: CompactionAttemptDiagnostic) => void;
74
82
  onTransportRetry?: (diagnostic: CompactionTransportRetryDiagnostic) => void;
75
83
  }
76
84
 
@@ -201,6 +209,7 @@ export async function generateValidatedSummary(
201
209
  for (;;) {
202
210
  if (input.event.signal.aborted) throw abortedError();
203
211
  attempt++;
212
+ input.onAttempt?.({ stage: input.stage, attempt, maxAttempts });
204
213
  try {
205
214
  response = await input.ctx.modelRegistry.complete(
206
215
  model,
@@ -1,4 +1,4 @@
1
- export const EXTENSION_VERSION = "0.3.8";
1
+ export const EXTENSION_VERSION = "0.3.10";
2
2
  export const SUPPORTED_PI_VERSION = "0.84.3";
3
3
  export const OBSERVER_PLANNER_VERSION = "observer-model-aware-v1";
4
4
  export const PLANNER_VERSION = "managed-learned-ranking-v1";