ds4-context-engine 0.3.8 → 0.3.10
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +11 -4
- package/docs/ADR/061-compaction-latency.md +8 -0
- package/docs/COMPACTION.md +17 -10
- package/docs/CONTEXT_MANIFEST.md +1 -1
- package/docs/CONTEXT_PLANNER.md +1 -1
- package/docs/MODEL_AWARENESS.md +73 -4
- package/docs/PROVIDER_TOKEN_DRIFT_BENCHMARK.md +85 -0
- package/docs/releases/0.3.10.md +29 -0
- package/docs/releases/0.3.8.md +1 -1
- package/docs/releases/0.3.9.md +80 -0
- package/package.json +3 -2
- package/src/extension/commands.ts +14 -0
- package/src/extension/runtime.ts +30 -11
- package/src/pi-adapter/bpe-token-estimator.ts +24 -0
- package/src/pi-adapter/compaction-coordinator.ts +89 -18
- package/src/pi-adapter/context-observer.ts +3 -1
- package/src/pi-adapter/summary-generator.ts +14 -5
- package/src/pi-adapter/version.ts +1 -1
package/README.md
CHANGED
|
@@ -16,9 +16,9 @@ bounded active context with provenance
|
|
|
16
16
|
Pi provider
|
|
17
17
|
```
|
|
18
18
|
|
|
19
|
-
> **Project status:** The coordinated `0.3.
|
|
19
|
+
> **Project status:** The coordinated `0.3.10` release adds opt-in BPE token estimation, estimator-specific provider calibration and bounded auto-tuning of context category ceilings. `chars-v1` and disabled auto-tuning remain the defaults. The bounded compaction controls from `0.3.9` remain in place; canonical history, SQLite schema 16 and runtime contracts are unchanged. Pi remains pinned to `0.84.3`. See the [0.3.10 release record](docs/releases/0.3.10.md) and [model-awareness validation](docs/MODEL_AWARENESS.md).
|
|
20
20
|
|
|
21
|
-
**
|
|
21
|
+
**Current compaction defaults:** `compaction.directUpdate=true`, `compaction.inputBudget="context"`, `compaction.segmentTargetTokens=30000`, `compaction.maxRequestInputTokens=64000`, `compaction.maxOperationInputTokens=2000000`, `compaction.maxConcurrentSegments=2`. Every DS4 provider attempt is bounded by the effective request limit, and the operation limit includes retries; `inputBudget="summary"` remains an explicit throughput-oriented opt-in. Existing compaction/master switches still apply. See [latency controls and compatibility](docs/COMPACTION.md#latency-controls). No real-provider speedup is claimed from mock tests. The five optional editing/reading/artifact/job features introduced in `0.3.4` remain default-off.
|
|
22
22
|
|
|
23
23
|
## Why DS4
|
|
24
24
|
|
|
@@ -331,9 +331,11 @@ The following example shows the main configuration groups. Omitted values use th
|
|
|
331
331
|
"mode": "hierarchical",
|
|
332
332
|
"validate": true,
|
|
333
333
|
"segmentTargetTokens": 30000,
|
|
334
|
+
"maxRequestInputTokens": 64000,
|
|
335
|
+
"maxOperationInputTokens": 2000000,
|
|
334
336
|
"preserveRecentVerbatim": true,
|
|
335
337
|
"directUpdate": true,
|
|
336
|
-
"inputBudget": "
|
|
338
|
+
"inputBudget": "context",
|
|
337
339
|
"maxConcurrentSegments": 2
|
|
338
340
|
},
|
|
339
341
|
"privacy": {
|
|
@@ -512,6 +514,11 @@ scripts package and release-readiness checks
|
|
|
512
514
|
- [Roadmap 0.2.0](docs/ROADMAP_0.2.0.md)
|
|
513
515
|
- [Release process](docs/RELEASING.md)
|
|
514
516
|
- [0.2.0 release readiness](docs/RELEASE_READINESS_0.2.0.md)
|
|
517
|
+
- [0.3.10 release notes](docs/releases/0.3.10.md)
|
|
518
|
+
- [0.3.9 release notes](docs/releases/0.3.9.md)
|
|
519
|
+
- [0.3.8 release notes](docs/releases/0.3.8.md)
|
|
520
|
+
- [0.3.7 release notes](docs/releases/0.3.7.md)
|
|
521
|
+
- [0.3.6 release notes](docs/releases/0.3.6.md)
|
|
515
522
|
- [0.3.5 release notes](docs/releases/0.3.5.md)
|
|
516
523
|
- [0.3.4 release notes](docs/releases/0.3.4.md)
|
|
517
524
|
- [0.3.3 release notes](docs/releases/0.3.3.md)
|
|
@@ -535,7 +542,7 @@ scripts package and release-readiness checks
|
|
|
535
542
|
|
|
536
543
|
The original M0–M13 roadmap is complete. `ds4-context-core` contains the compiled runtime-neutral implementation. M14 context-quality metrics, M15 rich symbol indexing, M16 hybrid semantic retrieval, M17 cross-session project memory, M18 learned-ranking shadow evaluation, M19's runtime adapter/conformance kit, and M20 opt-in local KV eligibility/replay are implemented on `main`. Learned active ranking remains promotion-gated, Pi reports local KV as unsupported, and static ranking/native completion stay authoritative on every failure.
|
|
537
544
|
|
|
538
|
-
The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.3.
|
|
545
|
+
The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.3.10 adds opt-in BPE estimation and bounded model-budget auto-tuning while retaining the `chars-v1` default. Version 0.3.9 extends the bounded compaction updates, summary input headroom, concurrent segments and phase timings introduced in 0.3.5 with per-request and cumulative operation input limits; 0.3.8 adds indexed FTS key deletion without changing search results. The opt-in [anchored editing](docs/ANCHORED_EDITING.md) and [portable agent tools](docs/PORTABLE_AGENT_TOOLS.md) from 0.3.4 remain default-off, without backend rewind, forced sampling or operational KV integration. Confirmation, provenance, Pi fallback and canonical/configuration/SQLite/runtime contracts remain unchanged. The [0.2 readiness record](docs/RELEASE_READINESS_0.2.0.md) remains the compatibility baseline; the lexical planner stays available as the deterministic fallback.
|
|
539
546
|
|
|
540
547
|
## Contributing
|
|
541
548
|
|
|
@@ -14,6 +14,14 @@ Pi normally updates previous summary plus discarded messages in one request. DS4
|
|
|
14
14
|
- Expose process-local metadata-only timings for preparation, segment/direct generation, aggregation, graph preparation/persistence and total DS4 hook duration, plus chosen path, effective provider/model, direct prompt size and concurrency. Timings use a monotonic clock, include retries and local validation within generation, and are wall times (not sums of overlapping calls). They are not canonical evidence, are not persisted in JSONL, and do not measure a subsequent native Pi fallback.
|
|
15
15
|
- Preserve schema-v2 compaction metadata, the existing `task-state` kind, source/classification/validation contracts, Pi cut points, atomic tool exchanges, fresh routing IDs, retry policy, canonical JSONL and all-or-nothing graph preparation. `segmentSummaryId` refers to the update node itself on a direct update; do not invent a segment that was never generated.
|
|
16
16
|
|
|
17
|
+
## Subsequent default revision
|
|
18
|
+
|
|
19
|
+
After operational testing, the default for `compaction.inputBudget` was changed from `summary` to `context`. The original `summary` mode remains available as an explicit throughput-oriented opt-in. The conservative default limits direct updates and indivisible atomic groups to the ordinary active input target, reducing request-size peaks at the cost of potentially more segment and aggregate calls. This revision changes only the default: calibrated hard limits, output headroom, atomicity, validation, and fail-closed behavior remain intact.
|
|
20
|
+
|
|
21
|
+
The hierarchical partitioner applies `min(compaction.segmentTargetTokens, effective request input limit)` to ordinary segments. A single indivisible message or complete tool exchange may exceed the target, but it is isolated and must still fit the effective request input limit.
|
|
22
|
+
|
|
23
|
+
Post-release hardening adds two additive safeguards. `maxRequestInputTokens` caps the estimated prompt size of every direct-update, segment, aggregate, and retry attempt after intersecting it with the selected model input budget. `maxOperationInputTokens` caps the cumulative estimated prompt input reserved across the whole DS4 operation, including retries and concurrent segment attempts. Crossing either limit fails closed to Pi without committing a partial summary graph.
|
|
24
|
+
|
|
17
25
|
## Compatibility and validation
|
|
18
26
|
|
|
19
27
|
Configuration is additive; absent fields use the new defaults. The old scheduling path can be compared using `directUpdate=false`, `inputBudget=context`, and `maxConcurrentSegments=1`. The compaction master switch still delegates to Pi when disabled. The original local implementation excluded versioning and publication; the user subsequently authorized the coordinated 0.3.5 release. No dependency upgrade, live database maintenance or Pi upgrade is part of this change.
|
package/docs/COMPACTION.md
CHANGED
|
@@ -8,12 +8,12 @@ DS4 intercepts Pi's `session_before_compact` event but preserves Pi's cut-point
|
|
|
8
8
|
2. DS4 maps every source message by exact fingerprint to a canonical branch entry ID.
|
|
9
9
|
3. Pi's serializer converts the newly discarded span to bounded conversation text; enabled privacy policy sanitizes conversation, previous summary, custom instructions, and file paths for the effective compaction provider (dedicated model when configured and eligible).
|
|
10
10
|
4. DS4 estimates the **complete** sanitized request against a calibrated summary-specific input budget, including framing, instructions, file inventories and output contract. With `compaction.directUpdate` enabled, a previous summary plus new source that fits is updated and validated in **one call**, producing an immutable `task-state` node. No predecessor needs only the existing one-segment call. Oversized updates fall through to hierarchical planning; no additional source is truncated to force a fit.
|
|
11
|
-
5. The hierarchical path partitions
|
|
11
|
+
5. The hierarchical path partitions source into ordered contiguous segments capped at `min(compaction.segmentTargetTokens, requestInputLimitTokens)`, where the effective request limit is also bounded by the selected model input budget. The target is soft only for one indivisible atomic group: an individual message or complete tool call/result exchange may exceed it, but must still fit the effective request input limit and is isolated in its own segment. Up to `compaction.maxConcurrentSegments` independent segment requests run concurrently (default 2). Each summary is validated against only its own sanitized evidence. Identities, source/child order and usage accumulation follow source order, not completion order. Cache retention stays disabled and every attempt has a fresh routing session ID.
|
|
12
12
|
6. Hierarchical requests recursively aggregate ordered children until one root remains: previous branch summary first, then new segments. Aggregation stays sequential and budget-checked. Direct updates instead link their single node to the predecessor and new canonical source IDs, without generating a synthetic segment or a separate aggregate. DS4-generated IDs, hashes, kinds, and graph levels never become model-visible evidence. A Pi-native predecessor is imported as an explicitly unverified branch node.
|
|
13
13
|
7. The highest input classification wraps each generated node, then all nodes are persisted atomically as one `prepared` graph batch. Usage includes every returned direct/segment/aggregate request and transport replay. Pi receives only the final root text and still appends exactly one canonical `CompactionEntry` with `fromHook: true`.
|
|
14
14
|
8. `session_compact` commits all nodes and associates the active root with the Pi entry; failure marks the complete prepared batch `failed`.
|
|
15
15
|
|
|
16
|
-
Fan-out and fan-in are bounded to 32 segment requests, 64 aggregate requests, and 16 aggregate passes. The DS4 transport replay policy is: `compaction.transport.maxAttempts` (default
|
|
16
|
+
Fan-out and fan-in are bounded to 32 segment requests, 64 aggregate requests, and 16 aggregate passes. The DS4 transport replay policy is: `compaction.transport.maxAttempts` (default 4, matching Pi's initial call plus three retries) total attempts and `compaction.transport.baseDelayMs` (default 2000 ms, capped at 60 s) backoff, doubling per attempt, abort-aware. Replay never applies to input, usage, rate, authentication, validation, or output-limit failures. A base prompt, individual message, atomic tool exchange, pair of child summaries, or total operation that cannot fit within those limits fails closed. Any mapping, budget, model, output-limit, validation, abort, or storage error returns `undefined` from the hook, allowing Pi's default compaction to run. On segment failure or cancellation DS4 stops scheduling, aborts siblings and awaits all started workers before fallback; no partial graph is installed. Cancellation is cooperative: accepted provider work may still cost tokens, and a provider that ignores abort can delay settlement.
|
|
17
17
|
|
|
18
18
|
## Required summary contract
|
|
19
19
|
|
|
@@ -80,42 +80,49 @@ Semantics:
|
|
|
80
80
|
|
|
81
81
|
## Latency controls
|
|
82
82
|
|
|
83
|
-
The
|
|
83
|
+
The current defaults favor bounded request size while retaining direct updates and bounded parallelism:
|
|
84
84
|
|
|
85
85
|
```json
|
|
86
86
|
{
|
|
87
87
|
"compaction": {
|
|
88
88
|
"directUpdate": true,
|
|
89
|
-
"inputBudget": "
|
|
89
|
+
"inputBudget": "context",
|
|
90
|
+
"segmentTargetTokens": 30000,
|
|
91
|
+
"maxRequestInputTokens": 64000,
|
|
92
|
+
"maxOperationInputTokens": 2000000,
|
|
90
93
|
"maxConcurrentSegments": 2
|
|
91
94
|
}
|
|
92
95
|
}
|
|
93
96
|
```
|
|
94
97
|
|
|
95
|
-
|
|
96
|
-
|
|
98
|
+
`segmentTargetTokens` is the soft partition target for source segments. `maxRequestInputTokens` is the hard estimated-input cap applied to every direct-update, segment, aggregate, and retry attempt; the effective request limit is the minimum of this value and the safe model input budget. An indivisible message or complete tool exchange above that limit fails closed to Pi's native compaction rather than creating an oversized DS4 request.
|
|
99
|
+
|
|
100
|
+
`maxOperationInputTokens` bounds the sum of estimated prompt tokens reserved immediately before all provider attempts in one DS4 compaction, including transport retries and concurrently scheduled segments. Exceeding it aborts remaining DS4 work and falls back without committing a partial summary graph. Defaults are 64,000 per request and 2,000,000 per operation; both must be positive integers, and the operation limit must be at least the configured request limit.
|
|
101
|
+
|
|
102
|
+
- `directUpdate`: one validated previous-summary plus new-source request when the entire prompt fits the effective request input limit. Set false to always retain the segment-then-aggregate route. Validation or provider failures still fall back to Pi, not an unvalidated update.
|
|
103
|
+
- `inputBudget`: `context` (default) uses the ordinary context fill target (`activeInputBudget`). `summary` is an explicit throughput-oriented opt-in that permits the calibrated `hardInputLimit`. Both are additionally capped by `(context window - safety margin - actual summary output cap) / calibration ratio`, rounded down, and never exceed the configured/model hard limit. Existing output reservations remain conservative; ordinary session planning and proactive thresholds are unchanged. The conservative default reduces request-size peaks but can produce more segments and calls; it is not a guarantee of provider fit or lower total token use.
|
|
97
104
|
- `maxConcurrentSegments`: integer **1–2**, default 2. Only independent segments overlap, including their retries. Aggregates do not run until their children have completed. Use 1 for sequential execution or providers with restrictive concurrent-request limits. Rate-limit failures are not transport-retried.
|
|
98
105
|
|
|
99
|
-
For an old-path comparison set `directUpdate=false`, `inputBudget=context`, `maxConcurrentSegments=1`. All features remain behind the existing compaction/master switches. Settings are applied on session load; after upgrading the package or rebuilding a development checkout, fully restart Pi to avoid stale compiled-core modules. No schema migration is required.
|
|
106
|
+
For an old sequential-path comparison set `directUpdate=false`, `inputBudget=context`, `maxConcurrentSegments=1`. To reproduce the `0.3.5` throughput-oriented budget, set `inputBudget=summary`. All features remain behind the existing compaction/master switches. Settings are applied on session load; after upgrading the package or rebuilding a development checkout, fully restart Pi to avoid stale compiled-core modules. No schema migration is required.
|
|
100
107
|
|
|
101
108
|
Mock-provider regression tests verify fewer calls, bounded overlap, exact budget boundaries, validation, privacy, immutable provenance and JSONL rebuild. They do **not** establish real-provider wall-time gains, semantic equivalence of generated summaries, or a guaranteed completion time. See [ADR-061](ADR/061-compaction-latency.md).
|
|
102
109
|
|
|
103
110
|
## Transport retry policy
|
|
104
111
|
|
|
105
|
-
Summary requests are replayed only for transport-classified failures (thrown transport errors or `stopReason: "error"` responses whose message matches network/timeout patterns). The DS4 replay policy uses **
|
|
112
|
+
Summary requests are replayed only for transport-classified failures (thrown transport errors or `stopReason: "error"` responses whose message matches network/timeout patterns). The DS4 replay policy uses **four total attempts** (the initial call plus up to three retries), matching Pi's standard retry count, and can be tuned per deployment:
|
|
106
113
|
|
|
107
114
|
```json
|
|
108
115
|
{
|
|
109
116
|
"compaction": {
|
|
110
117
|
"transport": {
|
|
111
|
-
"maxAttempts":
|
|
118
|
+
"maxAttempts": 4,
|
|
112
119
|
"baseDelayMs": 2000
|
|
113
120
|
}
|
|
114
121
|
}
|
|
115
122
|
}
|
|
116
123
|
```
|
|
117
124
|
|
|
118
|
-
- `compaction.transport.maxAttempts`: total attempts per direct update, segment or aggregate call, integer 1–10, default
|
|
125
|
+
- `compaction.transport.maxAttempts`: total attempts per direct update, segment or aggregate call, integer 1–10, default 4. With 1, no transport failure is retried.
|
|
119
126
|
- `compaction.transport.baseDelayMs`: base backoff before the first replay, integer 0–60000, default 2000. The delay doubles per attempt (2000, 4000, 8000, …) and is capped at 60 s.
|
|
120
127
|
- Replays use a fresh routing session per attempt; diagnostics expose only stage, failed/next attempt, max attempts, and delay.
|
|
121
128
|
- Aborts (including during backoff) never trigger replay; non-transport failures are never retried; usage is summed across replayed responses.
|
package/docs/CONTEXT_MANIFEST.md
CHANGED
|
@@ -45,7 +45,7 @@ actualInputTokens = input + cacheRead + cacheWrite
|
|
|
45
45
|
rawCalibrationRatio = actualInputTokens / estimatedInputTokens
|
|
46
46
|
```
|
|
47
47
|
|
|
48
|
-
The existing scalar SQLite columns are authoritative for persisted usage. `message_end` updates those columns and the calibration sample without reading or rewriting `manifest_json`; repository reads hydrate the usage fields from the scalars. The raw estimate is retained even after calibration so ratios cannot recursively calibrate already-corrected values. The next call reads only the exact provider/model window and applies bounded median/MAD outlier rejection. Error, aborted, missing, zero-usage, duplicate, or uncorrelated responses do not create calibration samples. Every manifest is correlated with at most one assistant response. If a concurrently pruned or oversize-skipped manifest cannot be correlated, DS4 keeps only bounded volatile calibration and does not retry against another row. See [`MODEL_AWARENESS.md`](MODEL_AWARENESS.md).
|
|
48
|
+
The existing scalar SQLite columns are authoritative for persisted usage. `message_end` updates those columns and the calibration sample without reading or rewriting `manifest_json`; repository reads hydrate the usage fields from the scalars. The raw estimate is retained even after calibration so ratios cannot recursively calibrate already-corrected values. The next call reads only the exact provider/model/**estimator-version** window and applies bounded median/MAD outlier rejection. When present, `modelAwareness.drift` and `modelAwareness.autoTune` contain only aggregate warning/tuning metadata (no message contents). Error, aborted, missing, zero-usage, duplicate, or uncorrelated responses do not create calibration samples. Every manifest is correlated with at most one assistant response. If a concurrently pruned or oversize-skipped manifest cannot be correlated, DS4 keeps only bounded volatile calibration and does not retry against another row. See [`MODEL_AWARENESS.md`](MODEL_AWARENESS.md).
|
|
49
49
|
|
|
50
50
|
## Retention
|
|
51
51
|
|
package/docs/CONTEXT_PLANNER.md
CHANGED
|
@@ -76,7 +76,7 @@ With privacy disabled, DS4 returns Pi's original `AgentMessage[]` when:
|
|
|
76
76
|
- final estimated input exceeds the hard limit;
|
|
77
77
|
- an unexpected adapter or planner exception occurs.
|
|
78
78
|
|
|
79
|
-
Expected fallbacks are recorded in the Context Manifest. With privacy enabled, the fallback baseline is the sanitized native array—not raw Pi messages—and an unexpected privacy failure replaces content/payload fields instead of sending unchecked data. Observer mode disables planning but still enforces enabled privacy policy and records manifests/usage calibration.
|
|
79
|
+
Expected fallbacks are recorded in the Context Manifest. Because fail-open preserves the complete native array, DS4's hard input limit is not guaranteed in fallback when the current request, fixed overhead, or another mandatory atomic group is itself oversized. With privacy enabled, the fallback baseline is the sanitized native array—not raw Pi messages—and an unexpected privacy failure replaces content/payload fields instead of sending unchecked data. Observer mode disables planning but still enforces enabled privacy policy and records manifests/usage calibration.
|
|
80
80
|
|
|
81
81
|
## Quality measurement
|
|
82
82
|
|
package/docs/MODEL_AWARENESS.md
CHANGED
|
@@ -1,5 +1,7 @@
|
|
|
1
1
|
# Advanced Model Awareness
|
|
2
2
|
|
|
3
|
+
**Validation status (OpenRouter, synthetic DS4/Pi sessions):** GPT-6-sol, GPT-6-luna and GPT-5.6-terra have each demonstrated multi-turn BPE calibration, a genuinely larger opt-in tail budget, provider usage above the native-window headroom threshold, and withdrawal of that expansion on the next turn in the same session. Historical retrieval and a live tool cycle were verified separately for all three. These bounded experiments do not establish universal tokenizer accuracy, provider-side compaction behavior, or production retrieval quality. `chars-v1` remains the default; BPE and `autoTune` remain opt-in. See the measured runs below.
|
|
4
|
+
|
|
3
5
|
M11 resolves an independent deterministic planning profile for every exact `provider/model`. Pi's model descriptor remains the default source for context window, output ceiling, reasoning, and image support; explicit DS4 overrides can repair provider metadata or tune category limits without changing canonical session state.
|
|
4
6
|
|
|
5
7
|
## Profile resolution
|
|
@@ -68,10 +70,10 @@ Each successful assistant response is correlated with one pending Context Manife
|
|
|
68
70
|
|
|
69
71
|
```text
|
|
70
72
|
actualInputTokens = input + cacheRead + cacheWrite
|
|
71
|
-
ratio = actualInputTokens / raw
|
|
73
|
+
ratio = actualInputTokens / raw selected-estimator estimate
|
|
72
74
|
```
|
|
73
75
|
|
|
74
|
-
Calibration is isolated by exact provider and
|
|
76
|
+
Calibration is isolated by exact provider, model ID **and estimator version**; switching from `chars-v1` to BPE never reuses the old ratios. The latest configured window is validated and processed deterministically:
|
|
75
77
|
|
|
76
78
|
1. reject invalid or zero samples;
|
|
77
79
|
2. reject ratios outside the configured hard bounds;
|
|
@@ -82,7 +84,38 @@ Calibration is isolated by exact provider and model ID. The latest configured wi
|
|
|
82
84
|
|
|
83
85
|
Until enough samples exist, the multiplier remains `1.0`. The default window is 24 samples, minimum is 3, and hard ratio bounds are 0.5–2.0. Outliers remain counted in metadata but do not influence the applied ratio.
|
|
84
86
|
|
|
85
|
-
Provider-token capacities are computed first from context window, output reserve, safety margin, and policy ratios.
|
|
87
|
+
Provider-token capacities are computed first from context window, output reserve, safety margin, and policy ratios. Global hard/soft/preferred input limits are converted into local-estimator units using **at least** a multiplier of 1: observed underestimation may reduce them, but apparent overestimation on short prompts never raises them above nominal model limits. Raw manifest estimates remain uncalibrated so future samples do not feed a corrected estimate back into itself. Adaptive tail/history/project budgets use the accepted multiplier but remain capped by the configured `context.*` maxima even after conversion.
|
|
88
|
+
|
|
89
|
+
### Optional BPE estimator and measured budget tuning
|
|
90
|
+
|
|
91
|
+
The default remains `chars-v1`. To opt a profile into local OpenAI `o200k_base` BPE text counting (without fetching vocabulary over the network):
|
|
92
|
+
|
|
93
|
+
```json
|
|
94
|
+
{
|
|
95
|
+
"modelAwareness": {
|
|
96
|
+
"overrides": { "openai/gpt-4o": { "tokenEstimator": "o200k-base-v1" } },
|
|
97
|
+
"autoTune": true
|
|
98
|
+
}
|
|
99
|
+
}
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
Use BPE only for a model known to share that encoding. The estimator counts text with BPE in the Pi observer, managed planner and fixed prompt/tool estimates; message wrappers, images, reasoning, provider-specific serialization, indexed retrieval hints, and persisted summaries remain estimates. It does **not** make token counts exact or enable provider-side continuation. If the provider's actual input tokens drift persistently from estimates (median ≥1.25 or ≤0.80 after the calibration minimum, or repeated hard-bound outliers), `/context model` and status surface a metadata-only warning; no prompt content is logged.
|
|
103
|
+
|
|
104
|
+
`autoTune` is opt-in and starts neutral. After at least eight accepted calibration samples, it expands *automatic* recent-tail, history and project ceilings by at most 12.5% of their nominal values, only if **every valid provider-usage sample** in the calibration window — including ratio outliers excluded from ratio calibration — remains at or below 60% of the current preferred/hard input target. It never raises a ceiling beyond the configured context cap, never overrides an explicit model-specific category limit, and never changes the provider's window, output reserve, safety margin or hard/soft input limits. Without enough evidence or headroom it stays neutral. Any calibration multiplier is applied after this small expansion; tune decisions and warnings are recorded as metadata in the manifest.
|
|
105
|
+
|
|
106
|
+
## Validation goal for BPE and auto-tuning
|
|
107
|
+
|
|
108
|
+
**User-requested outcome:** verify long conversations, tools, retrieval, calibration across turns, and auto-tuning safety on the context actually built by DS4/Pi. Single-turn synthetic token probes are preliminary evidence only; they do **not** satisfy this goal or justify declaring the feature complete.
|
|
109
|
+
|
|
110
|
+
The validation must cover, and report results separately for:
|
|
111
|
+
|
|
112
|
+
1. Long multi-turn sessions, including large context windows, compaction boundaries, and preservation of the current request.
|
|
113
|
+
2. Tool calls and results (including large results): atomic selection, offload, and the estimated versus observed token count.
|
|
114
|
+
3. Retrieved history and project context: relevance, budget selection, and whether relevant older material survives planning.
|
|
115
|
+
4. Consecutive turns on the **same** provider/model and estimator profile: provider usage correlation, accepted/rejected calibration samples, changes to the applied ratio, and isolation when models or estimators change.
|
|
116
|
+
5. Opt-in auto-tuning: evidence thresholds, configured and hard limits, outliers, insufficient headroom, and no regression in retrieval or context integrity when limits expand.
|
|
117
|
+
|
|
118
|
+
Use local, deterministic integration tests where possible. If real provider calls are needed, require a **new explicit call and input-size limit** before sending anything; the earlier single-turn budgets are exhausted. Report any scenario not verified as *not verified*, rather than treating the synthetic pilot as an end-to-end result. Do not change the default estimator or enable auto-tuning by default on the strength of that pilot.
|
|
86
119
|
|
|
87
120
|
## Provider cache metrics
|
|
88
121
|
|
|
@@ -126,6 +159,42 @@ Use:
|
|
|
126
159
|
/context status
|
|
127
160
|
```
|
|
128
161
|
|
|
162
|
+
## Local validation progress and outstanding provider evidence
|
|
163
|
+
|
|
164
|
+
`tests/integration/model-aware-real-context.test.ts` exercises the Pi `context` and `message_end` hooks over canonical multi-turn JSONL without any network transport. A long branch with a tool call/result verifies atomic selection and retrieval of an older decision under a BPE-managed budget. Successive turns verify that eight correlated usage records are required before opt-in expansion, that estimator/model switches isolate calibration, and that a new near-limit usage record withdraws the expansion even when its ratio is excluded from calibration. `tests/unit/model-aware-estimator.test.ts` also checks explicit overrides, configured caps, and the high-usage outlier regression.
|
|
165
|
+
|
|
166
|
+
**Local integration checks inject usage; they are not independent provider measurements.** A separately authorized, bounded live probe (`scripts/verify-model-aware-real-session.mjs`) used a temporary Pi JSONL session, DS4's actual hooks, synthetic content, `openrouter/openai/gpt-4o-mini`, and Pi-normalized provider usage. Eight successive short calls accumulated calibration samples; on the ninth call the manifest showed eight accepted samples, applied ratio 0.677157, and opt-in `autoTune: expanded`. That ninth call did **not** contain the intended appended long branch: Pi had already constructed its runner, and its manifest estimated only 1,940 input tokens. It is not evidence for long-context tuning.
|
|
167
|
+
|
|
168
|
+
A second, fresh Pi session with the long branch seeded *before* runner construction produced 83 original messages; DS4 selected 21 groups, excluded 21, retrieved the older `cobalt-713` decision, and included both the synthetic tool-call assistant entry and tool result. For that request, the BPE-based manifest estimate was **23,492** input tokens and provider usage was **23,196**. This is **one** live, untuned long-session/tool/retrieval measurement; it does not prove stable drift across lengths or models. The earlier [single-turn pilots](PROVIDER_TOKEN_DRIFT_BENCHMARK.md) remain separate. The two runs together attempted 10 calls and conservatively reserved 217,678 of the authorized 250,000 estimated-input-token limit (per-call ceiling 64,000; 16-call ceiling). All temporary session data was deleted.
|
|
169
|
+
|
|
170
|
+
A subsequent, separately authorized probe (`scripts/verify-model-aware-calibrated-session.mjs`) seeded a 28-turn synthetic history and tool call/result **before starting the same Pi session**. Eight real provider calls calibrated BPE on that long context; the ninth measured **23,926 estimated / 23,486 actual input tokens**, eight accepted samples, `autoTune: expanded`, and retrieval of the old decision with the historical tool call and result included. This verifies expansion on a real long context after calibration in one session, not just across separate sessions. Across this probe **11 calls** reserved **588,052/600,000** estimated input tokens (80,000 per-call limit); no further provider requests were made under that authorization.
|
|
171
|
+
|
|
172
|
+
The attempted higher-occupancy request used **37,552 estimated / 37,464 actual tokens**, below the 60%-of-target withdrawal threshold (**53,760** for this model/configuration). The following turn still reported `expanded`: **withdrawal is not verified live**. That larger turn also excluded the historical tool/retrieval groups, an observed quality limit under that selected context, not proof of an auto-tuning regression. The local Pi compaction-boundary test covers summary preservation under BPE. A separately bounded **one-request** live compaction probe (`scripts/verify-model-aware-compaction-session.mjs`) preseeded the canonical Pi compaction entry before starting Pi; the DS4 managed manifest included its summary and measured **250 estimated / 244 actual** provider input tokens. This used one further call under the third authorization, bringing that budget to **10 calls / 181,069 estimated input tokens reserved**. It checks the provider-bound context *after* a synthetic compaction entry, not live summary generation. At that point, live tool execution and high-usage withdrawal remained unverified; both were tested in the later authorized round below. Model variety and production retrieval quality still remain outside these bounded probes. A third bounded probe (`scripts/verify-model-aware-safety-session.mjs`) reached **60,509 actual input tokens** (>53,760), but only **seven** of eight preceding short-call samples had been accepted; the manifest at the high-usage request showed seven accepted samples and `insufficient-samples`, not an expanded policy. Thus it **does not** test withdrawal. The safety probe attempted **9 calls**, reserving **169,006** estimated input tokens, and stopped without a follow-up call. It now gates the expensive request on eight *accepted* samples and seeds a stable synthetic prefix, but has not been rerun. Its temporary session was deleted; the remaining third-round budget after the separate compaction request cannot fit a new calibration/high-usage/follow-up sequence under the preflight limits. No live tool execution was attempted in that round because the automatic multi-request cycle was not yet safely capped per request. Those results alone did not verify withdrawal or tool execution.
|
|
173
|
+
|
|
174
|
+
A final, separately authorized round verified the outstanding **synthetic** provider-path scenarios on the same OpenRouter/GPT-4o-mini profile. After a stable synthetic prefix, eight usage samples were accepted; the ninth short turn showed `expanded`. The next request reported **60,708 estimated / 60,640 actual** input tokens, above the **53,760** headroom threshold; the following turn reported **60,795 estimated / 60,722 actual** and `no-headroom` (recent-tail limit decreased from **27,699** to **24,564**). `scripts/verify-model-aware-safety-session.mjs` used **11/15** calls and **292,818/350,000** estimated-input tokens reserved (maximum 80,000 per call). This demonstrates withdrawal on measured provider usage, not merely a synthetic injected sample.
|
|
175
|
+
|
|
176
|
+
With the remaining budget, `scripts/verify-model-aware-live-tool.mjs` gated **each** Pi provider request at `ModelRuntime.streamSimple`, disabled cache warming/retry/compaction, and executed a single synthetic tool. The first request included its tool schema and used **105** provider input tokens; the second included the actual tool result and used **139**. Exactly **two** requests and one tool execution occurred. The final fourth-round total was **13/15 calls, 317,119/350,000** estimated-input tokens reserved. Only aggregate counters and booleans were reported; temporary synthetic sessions were removed.
|
|
177
|
+
|
|
178
|
+
These bounded live measurements, together with the canonical Pi JSONL integration tests, cover long context, historical and executed tools, retrieval, between-turn calibration, compaction-boundary context, and auto-tuning withdrawal **for this synthetic OpenRouter/GPT-4o-mini profile**. They do not establish universal accuracy for other providers/models, real private sessions, arbitrary compaction generation, or production retrieval quality.
|
|
179
|
+
|
|
180
|
+
### Native-window checks for Sol, Luna and Terra — separately authorized
|
|
181
|
+
|
|
182
|
+
The Pi catalog exposed a **1,050,000-token window** for each OpenRouter model `openai/gpt-6-sol`, `openai/gpt-6-luna`, and `openai/gpt-5.6-terra`. One additional authorization limited their *combined* probes to **42 provider attempts, 600,000 estimated input tokens / 2,500,000 controlled characters per request, and 4,800,000 estimated tokens / 15,000,000 controlled characters total**. `scripts/verify-model-aware-triad.mjs` and `scripts/verify-model-aware-triad-followup.mjs` gated **every** Pi request at `ModelRuntime.streamSimple`, including tool continuations; no real sessions, credentials, prompt bodies, responses or raw upstream errors were logged. Temporary canonical Pi JSONL sessions were deleted.
|
|
183
|
+
|
|
184
|
+
The first probe consumed **30 calls / 1,433,445 tokens / 3,679,430 characters reserved**. For *each* model, nine calibration calls in the same seeded long session led to at least eight accepted usage samples. The tenth call measured **36,027 BPE-estimated / 35,503 Pi-normalized actual input tokens**, `autoTune: expanded`, with the historical tool call and result included. **This did not verify retrieval**: the decision was still in the 64k recent tail and therefore had not been fetched by retrieval; the probe correctly stopped before tool execution and the high-occupancy turn. At this native profile, the reported `expanded` status did **not** demonstrate a larger recent-tail limit: its configured 64k cap was already saturated.
|
|
185
|
+
|
|
186
|
+
The second probe seeded **70 turns before Pi session construction**, putting the old decision outside the default recent tail. It consumed nine more calls (one retrieval request and an actual **two-request, one-execution** synthetic tool cycle per model). The decision was both retrieved and included, the historical tool call/result stayed included, and the new tool schema appeared in the first provider request with its actual result in the second. Per-model provider usage for the retrieval request was **63,050 / 63,049 / 63,053** tokens (Sol/Luna/Terra), against BPE estimates **63,699 / 63,698 / 63,702**. All three tool cycles completed. These are measurements of the DS4/Pi managed context, not three single-turn tokenizer prompts.
|
|
187
|
+
|
|
188
|
+
The three remaining authorized requests measured large BPE-managed input, **one per model**. The selected current user turn was included, and Pi reported **381,501 / 381,502 / 381,502** input tokens against manifest estimates **381,593 / 381,594 / 381,594**. The historical call/result were present, but the old decision was **not** selected; the oversized prompt did not ask about that decision, so this is **not** a valid retrieval-quality test. The request generator targeted 462k from the *raw* canonical history, but DS4 excluded older history and the selected provider input remained ~381.5k. Thus all three requests were **below** the native-window auto-tuning withdrawal threshold of **441,000** (60% of the 735k preferred target), and far below the full 1.05M window. They were fresh sessions without eight accepted calibration samples; **none tests `expanded → high provider usage → no-headroom` on these models**. Do not report the triad's auto-tuning safety or category expansion as live-verified at its native window; the deterministic local matrix in `tests/unit/model-aware-estimator.test.ts` covers only injected samples and caps. The entire authorization was used: **42/42 calls**, **3,296,138/4,800,000** estimated input tokens and **9,357,266/15,000,000** controlled characters reserved. No more provider calls are authorized under it. The historical probes now reject `--live` under this exhausted authorization. **After** those measurements, their local high-input sizing was corrected to require at least 470k BPE tokens in the selected current user text (rather than in raw canonical history), with a second gate immediately before provider transport; unit tests cover the discarded-history regression and per-call preflight. This correction was **not** run against a provider and cannot retroactively validate withdrawal.
|
|
189
|
+
|
|
190
|
+
Across these specific synthetic shapes, BPE manifests slightly overestimated provider usage, but no universal correction follows. At this earlier point native-window compaction generation, production retrieval quality and live high-usage auto-tune withdrawal for Sol/Luna/Terra were unverified; the later withdrawal runs below close **only the last of those gaps**. The later probes check actual category growth (not merely `expanded` status), a selected current turn above 470k BPE tokens, usage above 441k, and withdrawal in the same Pi session. Their temporary opt-in category ceilings are 80k/40k/40k because the default native tail ceiling is already equal to its configured maximum of 64k and cannot grow; the probes do not establish expansion under unchanged category defaults. BPE and auto-tuning remain opt-in; no default or provider-storage policy was changed.
|
|
191
|
+
|
|
192
|
+
### Native-window withdrawal round (later, opt-in categories)
|
|
193
|
+
|
|
194
|
+
With a separate explicit cap of 39 calls, 3,800,000 reserved input tokens and 11,000,000 controlled characters, the same-session Pi/DS4 probe used temporary category ceilings 80k/40k/40k to expose a **real** tail expansion. It ran 34 provider requests (3,375,536 estimated input tokens reserved; 9,845,274 controlled characters). For **GPT-6-sol and GPT-6-luna**, each session accepted at least eight calibration samples; the expanded tail exceeded the same-ratio no-tune baseline (about 72,922 versus 64,819 estimator tokens). A later request used **475,703 provider input tokens** (BPE estimate 475,790), above the 441,000 provider-token headroom threshold. On the next request, `autoTune` was `no-headroom` and the tail equalled its no-tune baseline (64,805); actual provider input was 471,846/471,844. The selected current turn, historical decision, call and result were present before the large request. The large input and follow-up did **not** retain the unrelated old decision/tool group; these measurements do not establish production retrieval quality.
|
|
195
|
+
|
|
196
|
+
For **GPT-5.6-terra**, calibration and real expansion passed in that first round, but its high-input request was blocked **before transport** by the cumulative budget/character caps: the probe had incorrectly budgeted the follow-up as short even though Pi resends the large preceding turn. The temporary session was then disposed. A **separately authorized Terra-only run** used a fresh synthetic Pi/DS4 session and counted *both* large requests. It made 12 provider calls (1,449,079 estimated tokens and 4,310,843 characters reserved, under separate 13-call / 1,900,000-token / 6,000,000-character maxima). After at least eight accepted calibration samples, Terra's tail grew to 72,922 estimator tokens versus a same-ratio untuned baseline of 64,819. The large request used **475,707 provider input tokens** (BPE estimate 475,794), above the 441,000 threshold. On the next turn, provider input was **471,852** (estimate 471,929), `autoTune` was `no-headroom`, and the tail returned to **64,805**, exactly its untuned baseline at that turn's ratio. The historical decision/call/result and current turn were included before the large request; unrelated older groups were excluded during the large request. **This verifies the defined live high-usage expansion/withdrawal scenario for all three models, not a general retrieval or compaction guarantee.** Both probes are now locked against replay; defaults remain unchanged.
|
|
197
|
+
|
|
129
198
|
## Performance and tests
|
|
130
199
|
|
|
131
|
-
`tests/benchmarks/model-awareness.bench.ts` measures a bounded 200-sample calibration analysis and repeated 32k/128k/200k profile resolution. Unit and golden tests cover deterministic tiers, override precedence, robust outlier rejection, cache accounting, and calibrated budgets. Integration tests switch local/remote providers and 32k/128k/200k models while checking profile isolation, privacy re-enforcement, canonical JSONL preservation, SQLite cache metrics, and profile reuse.
|
|
200
|
+
`tests/benchmarks/model-awareness.bench.ts` measures a bounded 200-sample calibration analysis and repeated 32k/128k/200k profile resolution. Unit and golden tests cover deterministic tiers, override precedence, robust outlier rejection, cache accounting, and calibrated budgets. Integration tests switch local/remote providers and 32k/128k/200k models while checking profile isolation, privacy re-enforcement, canonical JSONL preservation, SQLite cache metrics, and profile reuse. For a separately authorized, bounded real-provider measurement of both estimators against Pi SDK usage, see [PROVIDER_TOKEN_DRIFT_BENCHMARK.md](PROVIDER_TOKEN_DRIFT_BENCHMARK.md); it does not substitute for DS4 manifest measurements in a live session.
|
|
@@ -0,0 +1,85 @@
|
|
|
1
|
+
# Live provider token-drift probe (opt-in)
|
|
2
|
+
|
|
3
|
+
This developer-only, **source-checkout-only** probe (the script is not included in the published npm package) compares DS4's raw `chars-v1` and `o200k-base-v1` estimates with **Pi SDK provider usage** for synthetic, single-turn text. It does not start a Pi session, read session JSONL or SQLite, record responses, or change DS4 configuration. It never prints API keys, requests, response bodies or raw upstream error messages. It requires local Pi provider authentication and network access. The probe runs only with `--live`; without that flag it prints the bounded plan.
|
|
4
|
+
|
|
5
|
+
```bash
|
|
6
|
+
npm run build:core
|
|
7
|
+
node scripts/compare-provider-token-drift.mjs \
|
|
8
|
+
--model openrouter/openai/gpt-4o-mini \
|
|
9
|
+
--model deepseek/deepseek-v4-flash \
|
|
10
|
+
--model openai-codex/gpt-5.4-mini
|
|
11
|
+
# After inspecting the dry-run budget, add --live to make real calls.
|
|
12
|
+
```
|
|
13
|
+
|
|
14
|
+
The script accepts repeated `--model provider/model-id` and optional `--sizes 512,4096,48000`. **Per invocation**, it refuses more than 12 model calls or 200,000 total input characters (system prompt included), disables SDK request retries, and imposes a 45-second deadline per call. Multiple invocations have separate caps: track their cumulative spend yourself. Output is one metadata-only JSON report on stdout. The SDK's catalog-derived USD cost is an estimate, not a bill; a provider can still charge for failures or retries outside this script's control. If you pipe output to a file, keep it outside the repo unless you intentionally want to publish the aggregate metadata.
|
|
15
|
+
|
|
16
|
+
For each successful call, `actualInputTokens = usage.input + usage.cacheRead + usage.cacheWrite`; the cached fractions are **not** extra tokens on top of that sum. `residualTokens = actual - rawEstimate` (positive means underestimation), and `underestimationPctOfActual = max(0, residual) / actual × 100`. Adjacent slopes use differences across prompt sizes for a single exact provider/model, removing most fixed framing. BPE counts only text; both estimators use the same DS4 message/system overhead. No raw provider payload is available from this probe, so the result does **not** prove the accuracy of DS4's full observer in a live Pi extension chain (tools, images, privacy, cache/continuation, and later extensions differ). Compare `/context model`, `/context tokens`, and manifest `actualInputTokens` versus `estimatedInputTokens` in an ordinary consented Pi session for that second step. Never copy session text or auth files into the report.
|
|
17
|
+
|
|
18
|
+
## First observed run — 24 September 2026, 10:09–10:13 UTC
|
|
19
|
+
|
|
20
|
+
User-approved cumulative envelope: 12 call attempts and 200,000 input characters. Three invocations totalled **12 attempts and 195,176 planned input characters**: 9 calls/158,382 characters, a single Codex diagnostic call/574 characters, then 2 calls/36,220 characters. Success: 8; failure/missing usage: 4. Synthetic multilingual/code-like repeated text; no private session data. Sizes below are user-message characters; system prompt is included in estimates and the total envelope. These are observations, **not** a representative multi-session cost or quality benchmark.
|
|
21
|
+
|
|
22
|
+
| Provider / requested model | User chars | Real input¹ | `chars-v1` | BPE `o200k-base-v1` | BPE residual (real − estimate) |
|
|
23
|
+
|---|---:|---:|---:|---:|---:|
|
|
24
|
+
| OpenRouter / `openai/gpt-4o-mini` | 512 | 190 | 161 | 196 | −6 |
|
|
25
|
+
| OpenRouter / `openai/gpt-4o-mini` | 4,096 | 1,436 | 1,057 | 1,442 | −6 |
|
|
26
|
+
| OpenRouter / `openai/gpt-4o-mini` | 48,000 | 16,668 | 12,033 | 16,674 | −6 |
|
|
27
|
+
| DeepSeek / `deepseek-v4-flash`² | 512 | 179 | 161 | 196 | −17 |
|
|
28
|
+
| DeepSeek / `deepseek-v4-flash`² | 4,096 | 1,388 | 1,057 | 1,442 | −54 |
|
|
29
|
+
| DeepSeek / `deepseek-v4-flash`² | 48,000 | 16,172 | 12,033 | 16,674 | −502 |
|
|
30
|
+
| OpenRouter / `deepseek/deepseek-v4-flash` | 4,096 | 1,387 | 1,057 | 1,442 | −55 |
|
|
31
|
+
| OpenRouter / `deepseek/deepseek-v4-flash` | 32,000 | 10,785 | 8,033 | 11,124 | −339 |
|
|
32
|
+
|
|
33
|
+
¹ Pi's normalized provider usage: input + cache read + cache write. Some long requests reported cached reads (up to 1,408 tokens); they were included exactly once. Successful responses produced 1–3 output tokens. ² DeepSeek reported response model `deepseek-flash`, an alias of the requested ID; no equivalence to OpenRouter routing is assumed.
|
|
34
|
+
|
|
35
|
+
- OpenRouter GPT-4o-mini: raw `chars-v1` median actual/estimate **1.358562**; raw BPE median **0.995839**. At 48k chars, chars/4 underestimated by 4,635 tokens (27.81% of actual), while BPE overestimated by 6 tokens. Adjacent BPE slopes were 1.000 and 1.000.
|
|
36
|
+
- Direct DeepSeek: raw chars median **1.313150**; BPE median **0.962552**. At 48k chars, chars/4 underestimated by 4,139 tokens (25.59% of actual), while BPE overestimated by 502 tokens. Adjacent BPE slopes were 0.970305 and 0.970588. Close agreement here does **not** establish that DeepSeek uses OpenAI's tokenizer or justify turning BPE on for that model by default.
|
|
37
|
+
- OpenRouter DeepSeek: two valid samples; raw chars median **1.327396**, BPE median **0.965692**. At 32k chars, chars/4 underestimated by 2,752 tokens; BPE overestimated by 339. Two samples are insufficient to validate a persistent drift warning or tuning.
|
|
38
|
+
- `openai-codex/gpt-5.4-mini`: three initial attempts produced no usable usage; one later 512-character diagnostic returned `request-failed` with sanitized category `other` and no HTTP status. The local Pi auth resolver did return a credential, but the reason for the probe failure remains **unverified**. Do not infer its tokenizer drift, account availability in the interactive TUI, or provider-side token consumption from these calls. Direct `openai/gpt-4o-mini` API auth was not configured in the Pi runtime used by this probe; OpenRouter is a distinct route.
|
|
39
|
+
|
|
40
|
+
**Interpretation:** Three sizes with one call each do not establish statistical reliability or cover large context windows, tools, images, prefixes reused across turns, or output-heavy workflows. `chars-v1` remains the default; existing per-model calibration may compensate after enough **accepted same-profile samples**. The first small sample can be excluded by MAD filtering, so three different-size wire calls need not become three accepted DS4 calibration samples. Keep BPE and `autoTune` opt-in; do not port Hub's fixed 3.5% margin from these measurements. For promotion or automatic budget changes, collect a repeated same-size and mixed-size sample with matching DS4 manifest estimates and actual provider usage under an explicitly authorized, separately bounded run.
|
|
41
|
+
|
|
42
|
+
## Second observed run — 24 September 2026, 10:35 UTC
|
|
43
|
+
|
|
44
|
+
A separate, explicitly approved envelope covered **12 call attempts and up to 100,000 input characters**. The dry run planned exactly 12 attempts and 99,816 characters: `--sizes 512,16000` for three requested OpenAI models through each of OpenRouter and Codex. The live probe produced six OpenRouter usages and six Codex failures; it did not read session data.
|
|
45
|
+
|
|
46
|
+
| Route / requested model | User chars | Real input¹ | `chars-v1` | BPE `o200k-base-v1` | BPE residual |
|
|
47
|
+
|---|---:|---:|---:|---:|---:|
|
|
48
|
+
| OpenRouter / `openai/gpt-6-sol` | 512 | 189 | 161 | 196 | −7 |
|
|
49
|
+
| OpenRouter / `openai/gpt-6-sol` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
|
|
50
|
+
| OpenRouter / `openai/gpt-6-luna` | 512 | 189 | 161 | 196 | −7 |
|
|
51
|
+
| OpenRouter / `openai/gpt-6-luna` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
|
|
52
|
+
| OpenRouter / `openai/gpt-5.6-terra` | 512 | 189 | 161 | 196 | −7 |
|
|
53
|
+
| OpenRouter / `openai/gpt-5.6-terra` | 16,000 | 5,564 | 4,033 | 5,571 | −7 |
|
|
54
|
+
|
|
55
|
+
¹ Pi-normalized input includes cached tokens once. Each 16,000-character call reported `cacheWriteTokens = 5,561` and zero cache-read tokens. Those were **writes**, not cache hits. Each successful response reported five output tokens. The six equal input counts describe these identical synthetic prompts on this route; they are not proof of identical tokenizers or of how the direct Codex route would count a DS4 session. For each model, the adjacent BPE size slope was exactly 1.000; at 16k characters, raw `chars-v1` underestimated by 1,531 tokens (27.52% of actual), versus BPE overestimating by seven.
|
|
56
|
+
|
|
57
|
+
For `openai-codex/gpt-6-sol`, `gpt-6-luna`, and `gpt-5.6-terra`, **both sizes failed** without provider usage. The sanitized error classifier returned `quota` for all six, with no HTTP status captured. This is evidence of an SDK-path quota/limit error classification, **not** a verified account balance, an HTTP rejection, or a tokenizer measurement. It does not retroactively identify the earlier `gpt-5.4-mini` failure (classified `other`). No further calls were made beyond this run's approved envelope. Two sizes per exact model remain insufficient for DS4 calibration or auto-tuning conclusions.
|
|
58
|
+
|
|
59
|
+
## Isolated Pi + DS4 managed-context pilot — 24 September 2026
|
|
60
|
+
|
|
61
|
+
For a future, **separately authorized** run: build the core (`npm run build:core`), inspect the planned sizes and character count with `node scripts/compare-ds4-manifest-usage.mjs`, then use `node scripts/compare-ds4-manifest-usage.mjs --live` only after setting a new call/character budget. The default mode is pinned to OpenRouter `openai/gpt-6-sol` and a maximum of 12 calls / 120,000 controlled prompt characters per invocation. `--comparison` instead preflights **both** OpenRouter `openai/gpt-6-luna` and `openai/gpt-5.6-terra`, capped at 24 calls / 240,000 controlled characters combined; `--comparison --live` requires its own authorization. A new authorization is required even when repeating either exact plan.
|
|
62
|
+
|
|
63
|
+
The user separately authorized **up to 12 calls / 120,000 controlled input characters** to OpenRouter `openai/gpt-6-sol` only. `scripts/compare-ds4-manifest-usage.mjs` first checked a local sandbox without provider traffic, then made **12 single-turn calls**, with three repeats at each size (512, 4,096, 12,000 and 20,000 synthetic user characters). Planned synthetic user + fixed system text was **110,568 characters**. This limit counts controlled prompt text, **not** Pi-generated framing or serialized protocol bytes. No existing Pi session history was used; each Pi session and DS4 database was isolated in a temporary directory, with no tools, project content, native continuation, or auto-tuning. DS4's managed `context` hook and its `o200k-base-v1` profile override were active. The script stops on a missing/mismatched manifest or missing usage rather than fabricating a measurement.
|
|
64
|
+
|
|
65
|
+
| Synthetic user chars | First observed DS4 manifest estimate | First actual input¹ | First residual (actual − estimate) | Repeats |
|
|
66
|
+
|---:|---:|---:|---:|---:|
|
|
67
|
+
| 512 | 220 | 209 | −11 | 3 |
|
|
68
|
+
| 4,096 | 1,467 | 1,456 | −11 | 3 |
|
|
69
|
+
| 12,000 | 4,210 | 4,199 | −11 | 3 |
|
|
70
|
+
| 20,000 | 6,983 | 6,972 | −11 | 3 |
|
|
71
|
+
|
|
72
|
+
All **12** calls had matching `managed` manifests and Pi-normalized provider usage, with **exactly −11 tokens** of residual in each call. Token counts varied by about one or two between identically sized repeats, but the residual stayed fixed. The largest relative overestimate was **5.26% of actual input** on the first 512-character call (11/209); at 20,000 characters it was about **0.16%**. `cacheReadTokens` was zero, while longer calls reported mostly cache **writes**; no claim of cache hits or exact billed usage follows from those fields.
|
|
73
|
+
|
|
74
|
+
¹ `totalInputTokens` in the DS4 manifest (the same Pi-normalized `input + cacheRead + cacheWrite` metric used by model-awareness calibration), not independently obtained provider invoice data. Because the selected estimator was BPE, this pilot did **not** produce an alternative `chars-v1` DS4 manifest for the same wire requests. These are fresh one-turn sessions with one synthetic prompt shape and a maximum of ~7k actual tokens; they do not validate multi-turn prefixes, tool schemas/results, retrieval, real project text, larger windows, other requested models, or production auto-tuning. The 12 samples are independent sessions, **not** 12 accepted samples in one persistent DS4 calibration profile. Keep BPE opt-in and `autoTune` off by default; do not apply an inferred fixed framing correction or Hub's fixed 3.5% drift allowance from this pilot.
|
|
75
|
+
|
|
76
|
+
## Separately authorized Luna/Terra managed-context comparison — 24 September 2026
|
|
77
|
+
|
|
78
|
+
A subsequent authorization covered **at most 24 calls / 240,000 controlled synthetic input characters combined** on OpenRouter `openai/gpt-6-luna` and `openai/gpt-5.6-terra`; no Codex traffic was authorized. `node scripts/compare-ds4-manifest-usage.mjs --comparison` preflighted **24 calls / 221,136 controlled characters**. `--comparison --live` completed **12 calls per model**, with three fresh, isolated Pi+DS4 managed sessions per model at each synthetic user size (512, 4,096, 12,000, 20,000 characters). No existing session history, project content, tools, auto-tuning, or native continuation entered the requests. All 24 calls had matching manifests and Pi-normalized provider input usage; there were no reported failed rows.
|
|
79
|
+
|
|
80
|
+
| Model | First 512-character DS4 BPE estimate | First actual input¹ | Repeats per size | Residual on all 12 calls |
|
|
81
|
+
|---|---:|---:|---:|---:|
|
|
82
|
+
| OpenRouter / `openai/gpt-6-luna` | 220 | 209 | 3 | −11 tokens |
|
|
83
|
+
| OpenRouter / `openai/gpt-5.6-terra` | 219 | 208 | 3 | −11 tokens |
|
|
84
|
+
|
|
85
|
+
In every size group for both models, each of the three residuals was **−11 tokens**; each group reported zero `cacheReadTokens`. This replicates the small, constant **overestimate** observed for Sol on this specific one-turn synthetic shape. It does **not** establish a provider-independent 11-token correction, tokenizer equivalence across arbitrary text, long-window accuracy, model calibration from a continuous session, or safe budget auto-tuning. The combined 24-call authorization is exhausted. A later, separate authorization covered multi-turn, historical and executed tools, retrieval, and large selected inputs for all three OpenRouter models; see [native-window checks](MODEL_AWARENESS.md#native-window-checks-for-sol-luna-and-terra--separately-authorized) for results and the remaining unverified high-usage auto-tune withdrawal. Codex diagnosis still needs separate approval; keep BPE opt-in and `autoTune` off by default.
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
# Release 0.3.10 — Opt-in BPE estimation and bounded model budget auto-tuning
|
|
2
|
+
|
|
3
|
+
**Coordinated packages:** `ds4-context-core`, `ds4-context-reference-adapter`, and `ds4-context-engine` 0.3.10.
|
|
4
|
+
**Implementation commit:** `0607238`.
|
|
5
|
+
|
|
6
|
+
## Summary
|
|
7
|
+
|
|
8
|
+
Adds a selectable `o200k-base-v1` text estimator, provider/model/estimator-specific calibration, and opt-in evidence-gated expansion of automatic context-category ceilings. `chars-v1` remains the default estimator; `modelAwareness.autoTune` remains off unless explicitly enabled. The reference adapter and Pi extension continue to depend exactly on the matching core version.
|
|
9
|
+
|
|
10
|
+
## Changes
|
|
11
|
+
|
|
12
|
+
- Portable core accepts a BPE estimator through its existing runtime-neutral interface; the Pi adapter lazy-loads `js-tiktoken`. Non-text content still uses bounded heuristics.
|
|
13
|
+
- Model profiles can select `tokenEstimator` by exact `provider/model`, provider wildcard, or global override. Calibration histories are isolated by estimator version as well as provider and model.
|
|
14
|
+
- Auto-tuning uses accepted same-profile samples, rejects outliers, respects configured category limits and hard input ceilings, and withdraws expansion after insufficient headroom. Calibrated estimator-unit limits cannot exceed nominal model limits.
|
|
15
|
+
- Manifests and model diagnostics identify estimator version and calibration/auto-tuning decisions. Existing golden compatibility remains anchored to the unchanged default behavior.
|
|
16
|
+
|
|
17
|
+
## Measured scope and limitations
|
|
18
|
+
|
|
19
|
+
Bounded synthetic multi-turn Pi/DS4 sessions on OpenRouter `openai/gpt-6-sol`, `openai/gpt-6-luna`, and `openai/gpt-5.6-terra` demonstrated calibration, a genuinely larger opt-in recent tail, provider usage above the native-window headroom threshold, and withdrawal to `no-headroom` on the following turn in each model's session. Retrieval and live tool cycles were also exercised separately. See [model-awareness measurements](../MODEL_AWARENESS.md) and [provider-token drift](../PROVIDER_TOKEN_DRIFT_BENCHMARK.md) for methodology, numbers, and bounds.
|
|
20
|
+
|
|
21
|
+
These measurements do not establish production retrieval quality, provider-side compaction behavior, or accuracy for arbitrary models/routes. In particular, equivalent Codex-route measurements did not yield usage and are not claimed as validated. No provider calls are required for this release procedure.
|
|
22
|
+
|
|
23
|
+
## Compatibility
|
|
24
|
+
|
|
25
|
+
No new defaults are enabled: `chars-v1`, disabled `autoTune`, and disabled DS4 native continuation remain unchanged. No SQLite migration or change to canonical Pi JSONL, privacy consent, native continuation, or the portable runtime-adapter contract is introduced. BPE remains adapter-injected rather than adding a third-party tokenizer dependency to portable core. Restart Pi after upgrading so the new compiled core and session configuration are loaded.
|
|
26
|
+
|
|
27
|
+
## Validation and publication
|
|
28
|
+
|
|
29
|
+
On Node 26.5.1, the coordinated release passed a clean `npm ci`, `npm run check` (97 Vitest files, 607 tests, TypeScript builds and root typecheck), deterministic `npm run quality:compare`, the `npm run schema:context-persistence` size bound, and `npm run pack:check` in a clean consumer after synchronizing both exported runtime version constants. All typecheck/tests were also rerun after that correction. Dry-run tarball review contained 243 core files, 7 reference-adapter files, and 97 extension files; none included untracked local state. Registry publication and exact-artifact verification are still pending.
|
package/docs/releases/0.3.8.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Release 0.3.8 — Indexed FTS key deletes via rowid mappings (schema 16)
|
|
2
2
|
|
|
3
3
|
**Version analyzed:** DS4 Context Engine `0.3.8`
|
|
4
|
-
**Commit:**
|
|
4
|
+
**Commit:** `18f36e1`
|
|
5
5
|
**Coordinated packages:** `ds4-context-core` 0.3.8, `ds4-context-reference-adapter` 0.3.8, `ds4-context-engine` 0.3.8
|
|
6
6
|
|
|
7
7
|
## Summary
|
|
@@ -0,0 +1,80 @@
|
|
|
1
|
+
# Release 0.3.9 — Bounded compaction request and operation input
|
|
2
|
+
|
|
3
|
+
**Version analyzed:** DS4 Context Engine `0.3.9`
|
|
4
|
+
**Commit:** `01407cd`
|
|
5
|
+
**Coordinated packages:** `ds4-context-core` 0.3.9, `ds4-context-reference-adapter` 0.3.9, `ds4-context-engine` 0.3.9
|
|
6
|
+
|
|
7
|
+
## Summary
|
|
8
|
+
|
|
9
|
+
Hardens custom compaction against request-size and total-operation token spikes.
|
|
10
|
+
Every direct-update, segment, aggregate and retry attempt is now bounded by an
|
|
11
|
+
effective estimated-input limit. A second cumulative budget covers the complete
|
|
12
|
+
compaction operation, including concurrent work and transport retries.
|
|
13
|
+
|
|
14
|
+
The default compaction input budget changes from the summary hard limit to the
|
|
15
|
+
ordinary context-fill budget. The former behavior remains available explicitly
|
|
16
|
+
for users who accept larger provider requests in exchange for fewer hierarchy
|
|
17
|
+
levels.
|
|
18
|
+
|
|
19
|
+
## Changes
|
|
20
|
+
|
|
21
|
+
- Added `compaction.maxRequestInputTokens`, default 64,000, as a hard ceiling on
|
|
22
|
+
estimated input for each compaction provider request. The effective ceiling is
|
|
23
|
+
the minimum of this setting and the selected model-aware compaction budget.
|
|
24
|
+
- Added `compaction.maxOperationInputTokens`, default 2,000,000, for cumulative
|
|
25
|
+
estimated input across the complete operation. Every accepted attempt reserves
|
|
26
|
+
its input before provider dispatch, so retries and concurrent segments count.
|
|
27
|
+
- Configuration validation requires both limits to be positive integers and the
|
|
28
|
+
operation limit to be at least the configured request limit.
|
|
29
|
+
- Changed the default `compaction.inputBudget` from `summary` to `context`.
|
|
30
|
+
Explicit `summary` configurations retain the previous calibrated-hard-limit
|
|
31
|
+
behavior.
|
|
32
|
+
- Changed the default `compaction.transport.maxAttempts` from 3 to 4. Attempts
|
|
33
|
+
remain covered by the cumulative operation budget.
|
|
34
|
+
- Segment partitioning now uses the minimum of `segmentTargetTokens` and the
|
|
35
|
+
effective request limit. An indivisible message or complete tool exchange that
|
|
36
|
+
cannot fit fails closed to Pi's native compaction.
|
|
37
|
+
- `/context compaction` reports the effective request limit and cumulative
|
|
38
|
+
operation input consumed.
|
|
39
|
+
- Direct update, segment and aggregate paths continue to use strict validation,
|
|
40
|
+
privacy filtering, immutable provenance and all-or-nothing graph installation.
|
|
41
|
+
|
|
42
|
+
## Validation
|
|
43
|
+
|
|
44
|
+
Before publication, the coordinated release passed:
|
|
45
|
+
|
|
46
|
+
- TypeScript builds and typecheck for the root, core and reference adapter;
|
|
47
|
+
- 85 Vitest files with 561 tests;
|
|
48
|
+
- deterministic context-quality comparison;
|
|
49
|
+
- context-persistence schema measurement;
|
|
50
|
+
- clean-consumer package verification for all three tarballs;
|
|
51
|
+
- whitespace/error-marker checks and coordinated manifest inspection.
|
|
52
|
+
|
|
53
|
+
The package verifier also checks the new defaults and both input-budget modes
|
|
54
|
+
from the packed `ds4-context-core` artifact.
|
|
55
|
+
|
|
56
|
+
## Compatibility and migration
|
|
57
|
+
|
|
58
|
+
- No SQLite migration: schema 16 remains current.
|
|
59
|
+
- No Pi, runtime-adapter, history, persistence-tool or summary-contract version
|
|
60
|
+
changes.
|
|
61
|
+
- Existing configurations that explicitly set `inputBudget` or transport retry
|
|
62
|
+
attempts retain those values.
|
|
63
|
+
- New limit fields are additive. Configurations that omit them receive the
|
|
64
|
+
documented defaults.
|
|
65
|
+
- Pi JSONL remains canonical and custom-compaction failure still delegates to
|
|
66
|
+
Pi's native fallback without installing a partial Summary Graph.
|
|
67
|
+
- A full Pi restart is required after upgrading so compiled core modules and
|
|
68
|
+
session configuration are reloaded.
|
|
69
|
+
|
|
70
|
+
## Known limits
|
|
71
|
+
|
|
72
|
+
- If the current request, fixed overhead or another mandatory atomic group is
|
|
73
|
+
itself above the planner hard limit, fail-open context planning preserves the
|
|
74
|
+
native context; DS4 cannot transparently split an in-flight Pi tool loop.
|
|
75
|
+
- Provider work accepted before a concurrent sibling fails may still be billed,
|
|
76
|
+
even though DS4 aborts and awaits all started workers before fallback.
|
|
77
|
+
- Limits use DS4's calibrated token estimator. Provider-side tokenization and
|
|
78
|
+
billing remain authoritative.
|
|
79
|
+
- Mock-provider tests establish bounded scheduling and request accounting, not
|
|
80
|
+
real-provider latency or semantic-equivalence guarantees.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "ds4-context-engine",
|
|
3
|
-
"version": "0.3.
|
|
3
|
+
"version": "0.3.10",
|
|
4
4
|
"description": "Non-destructive, provider-independent context management for Pi.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|
|
@@ -62,7 +62,8 @@
|
|
|
62
62
|
]
|
|
63
63
|
},
|
|
64
64
|
"dependencies": {
|
|
65
|
-
"ds4-context-core": "0.3.
|
|
65
|
+
"ds4-context-core": "0.3.10",
|
|
66
|
+
"js-tiktoken": "1.0.21"
|
|
66
67
|
},
|
|
67
68
|
"peerDependencies": {
|
|
68
69
|
"@earendil-works/pi-ai": "0.84.3",
|
|
@@ -250,6 +250,9 @@ function formatStatus(diagnostics: RuntimeDiagnostics): string {
|
|
|
250
250
|
`Privacy blocked/redacted: ${count(diagnostics.privacy.blockedBlocks)} / ${count(diagnostics.privacy.secretRedactions)}`,
|
|
251
251
|
`Estimator calibration: ${diagnostics.modelAwareness?.calibration.calibrated ? `x${diagnostics.modelAwareness.calibration.appliedRatio.toFixed(3)}` : "collecting/neutral"}`,
|
|
252
252
|
`Calibration samples: ${count(diagnostics.modelAwareness?.calibration.acceptedSamples)} accepted`,
|
|
253
|
+
...(diagnostics.modelAwareness?.drift ? [
|
|
254
|
+
`Token estimate warning: ${diagnostics.modelAwareness.drift.code} (${diagnostics.modelAwareness.drift.severity}, ${count(diagnostics.modelAwareness.drift.sampleCount)} samples)`,
|
|
255
|
+
] : []),
|
|
253
256
|
`Native continuation: ${diagnostics.nativeContinuation.status} (${diagnostics.nativeContinuation.last?.mode ?? "no request"})`,
|
|
254
257
|
`Continuation saved items: ${count(diagnostics.nativeContinuation.last?.omittedInputItems)}`,
|
|
255
258
|
`Quality metrics: ${diagnostics.quality.enabled ? `${count(diagnostics.quality.storedSamples)} sample(s)` : "disabled"}`,
|
|
@@ -365,6 +368,9 @@ function formatManifest(diagnostics: RuntimeDiagnostics): string {
|
|
|
365
368
|
`Artifacts: ${count(manifest.artifacts?.length ?? 0)}`,
|
|
366
369
|
`Privacy: ${manifest.privacy?.enforcement ?? "disabled"}${manifest.privacy ? ` (${manifest.privacy.destination})` : ""}`,
|
|
367
370
|
`Model calibration: ${manifest.modelAwareness?.calibration.calibrated ? `x${manifest.modelAwareness.calibration.appliedRatio.toFixed(3)}` : "neutral/collecting"}`,
|
|
371
|
+
...(manifest.modelAwareness?.drift ? [
|
|
372
|
+
`Token estimate warning: ${manifest.modelAwareness.drift.code} (${manifest.modelAwareness.drift.severity})`,
|
|
373
|
+
] : []),
|
|
368
374
|
`Adaptive tail/hist/project: ${count(manifest.modelAwareness?.adaptive.recentTailTokens)} / ${count(manifest.modelAwareness?.adaptive.maxRetrievedHistoryTokens)} / ${count(manifest.modelAwareness?.adaptive.maxProjectTokens)}`,
|
|
369
375
|
`Continuation: ${manifest.nativeContinuation?.mode ?? "disabled"}; sent/full ${count(manifest.nativeContinuation?.sentInputItems)} / ${count(manifest.nativeContinuation?.fullInputItems)}`,
|
|
370
376
|
"",
|
|
@@ -664,6 +670,12 @@ function formatModelAwareness(diagnostics: RuntimeDiagnostics): string {
|
|
|
664
670
|
`Samples observed/accepted: ${count(calibration.observedSamples)} / ${count(calibration.acceptedSamples)}`,
|
|
665
671
|
`Samples rejected/outliers: ${count(calibration.rejectedSamples)} / ${count(calibration.outlierSamples)}`,
|
|
666
672
|
`Calibration bounds/window: ${calibration.lowerRatioBound.toFixed(2)}-${calibration.upperRatioBound.toFixed(2)} / ${count(calibration.windowSize)}`,
|
|
673
|
+
...(awareness.drift ? [
|
|
674
|
+
`Token estimate warning: ${awareness.drift.code} (${awareness.drift.severity}, ${count(awareness.drift.sampleCount)} samples${awareness.drift.medianRatio === undefined ? "" : `, x${awareness.drift.medianRatio.toFixed(3)}`})`,
|
|
675
|
+
] : []),
|
|
676
|
+
...(awareness.autoTune ? [
|
|
677
|
+
`Budget auto-tuning: ${awareness.autoTune.status} (${count(awareness.autoTune.acceptedSamples)} samples, x${awareness.autoTune.boostFactor.toFixed(3)})`,
|
|
678
|
+
] : []),
|
|
667
679
|
`Adaptive recent tail: ${count(awareness.adaptive.recentTailTokens)} (nominal ${count(awareness.adaptive.nominalRecentTailTokens)})`,
|
|
668
680
|
`Adaptive history retrieval: ${count(awareness.adaptive.maxRetrievedHistoryTokens)} (nominal ${count(awareness.adaptive.nominalRetrievedHistoryTokens)})`,
|
|
669
681
|
`Adaptive project retrieval: ${count(awareness.adaptive.maxProjectTokens)} (nominal ${count(awareness.adaptive.nominalProjectTokens)})`,
|
|
@@ -881,6 +893,8 @@ function formatCompaction(diagnostics: RuntimeDiagnostics, preview: boolean): st
|
|
|
881
893
|
`Chosen path: ${compaction.path ?? "n/a"}`,
|
|
882
894
|
`Input budget mode: ${compaction.inputBudgetMode ?? "n/a"}`,
|
|
883
895
|
`Input budget: ${count(compaction.inputBudgetTokens)}`,
|
|
896
|
+
`Request input limit: ${count(compaction.requestInputLimitTokens)} (configured max ${count(compaction.maxRequestInputTokens)})`,
|
|
897
|
+
`Operation input: ${count(compaction.operationInputTokens)} / ${count(compaction.maxOperationInputTokens)}`,
|
|
884
898
|
`Whole-source prompt: ${count(compaction.sourcePromptTokens)}`,
|
|
885
899
|
`Direct-update prompt: ${count(compaction.directPromptTokens)}`,
|
|
886
900
|
`Segment concurrency cap: ${count(compaction.maxConcurrentSegments)}`,
|
package/src/extension/runtime.ts
CHANGED
|
@@ -60,11 +60,12 @@ import { calculateContextBudget, type ContextBudget } from "ds4-context-core/cor
|
|
|
60
60
|
import {
|
|
61
61
|
modelProfileKey,
|
|
62
62
|
resolveModelAwareness,
|
|
63
|
+
tokenEstimatorVersion,
|
|
63
64
|
type ResolvedModelAwareness,
|
|
64
65
|
type TokenCalibrationSample,
|
|
65
66
|
} from "ds4-context-core/core/model-awareness";
|
|
66
67
|
import type { ModelDescriptor } from "ds4-context-core/core/model-profile";
|
|
67
|
-
import {
|
|
68
|
+
import { CHARS_ESTIMATOR, type TokenEstimator } from "ds4-context-core/core/token-estimator";
|
|
68
69
|
import type {
|
|
69
70
|
ContextManifest,
|
|
70
71
|
ModelAwarenessManifest,
|
|
@@ -164,6 +165,7 @@ import {
|
|
|
164
165
|
findPiSourceEntryIds,
|
|
165
166
|
fingerprint,
|
|
166
167
|
} from "../pi-adapter/context-observer.ts";
|
|
168
|
+
import { O200K_ESTIMATOR } from "../pi-adapter/bpe-token-estimator.ts";
|
|
167
169
|
import { projectSessionFileMutations } from "../pi-adapter/memory-adapter.ts";
|
|
168
170
|
import { ProjectMemorySynchronizer } from "../pi-adapter/project-memory-sync.ts";
|
|
169
171
|
import { projectRankingLabels } from "../pi-adapter/ranking-adapter.ts";
|
|
@@ -758,15 +760,22 @@ export class Ds4ContextRuntime {
|
|
|
758
760
|
};
|
|
759
761
|
}
|
|
760
762
|
|
|
763
|
+
private estimatorForModel(model: ModelDescriptor): TokenEstimator {
|
|
764
|
+
return tokenEstimatorVersion(model, this.config.modelAwareness) === "o200k-base-v1"
|
|
765
|
+
? O200K_ESTIMATOR : CHARS_ESTIMATOR;
|
|
766
|
+
}
|
|
767
|
+
|
|
761
768
|
private calibrationSamples(model: ModelDescriptor): TokenCalibrationSample[] {
|
|
769
|
+
const version = this.estimatorForModel(model).version;
|
|
762
770
|
if (this.database && this.session?.sessionFile && this.config.diagnostics.storeContextManifest) {
|
|
763
771
|
return this.database.manifests.listCalibrationSamples(
|
|
764
772
|
model.provider,
|
|
765
773
|
model.id,
|
|
766
774
|
this.config.modelAwareness.calibrationWindow,
|
|
775
|
+
version,
|
|
767
776
|
);
|
|
768
777
|
}
|
|
769
|
-
return [...(this.volatileCalibration.get(modelProfileKey(model.provider, model.id)) ?? [])];
|
|
778
|
+
return [...(this.volatileCalibration.get(`${modelProfileKey(model.provider, model.id)}\0${version}`) ?? [])];
|
|
770
779
|
}
|
|
771
780
|
|
|
772
781
|
private switchForModel(model: ModelDescriptor): ModelSwitchManifest {
|
|
@@ -831,6 +840,7 @@ export class Ds4ContextRuntime {
|
|
|
831
840
|
requestsPerTurn: number,
|
|
832
841
|
turnsPerEpoch: number,
|
|
833
842
|
modelKey: string,
|
|
843
|
+
estimator: TokenEstimator,
|
|
834
844
|
): number | undefined {
|
|
835
845
|
const pricing = Ds4ContextRuntime.cachePricing(cost);
|
|
836
846
|
if (!pricing || plan.mode !== "managed") return undefined;
|
|
@@ -839,7 +849,7 @@ export class Ds4ContextRuntime {
|
|
|
839
849
|
let total = fixedTokens;
|
|
840
850
|
for (const message of plan.messages) {
|
|
841
851
|
hashes.push(fingerprint(message));
|
|
842
|
-
const estimate = estimateMessagesTokens([message]);
|
|
852
|
+
const estimate = estimator.estimateMessagesTokens([message]);
|
|
843
853
|
tokens.push(estimate);
|
|
844
854
|
total += estimate;
|
|
845
855
|
}
|
|
@@ -908,6 +918,8 @@ export class Ds4ContextRuntime {
|
|
|
908
918
|
...awareness.calibration,
|
|
909
919
|
cache: { ...awareness.calibration.cache },
|
|
910
920
|
},
|
|
921
|
+
...(awareness.drift ? { drift: awareness.drift } : {}),
|
|
922
|
+
...(awareness.autoTune ? { autoTune: awareness.autoTune } : {}),
|
|
911
923
|
adaptive: { ...awareness.limits },
|
|
912
924
|
switch: modelSwitch,
|
|
913
925
|
};
|
|
@@ -922,7 +934,7 @@ export class Ds4ContextRuntime {
|
|
|
922
934
|
usage: ProviderUsageManifest,
|
|
923
935
|
createdAt: number,
|
|
924
936
|
): void {
|
|
925
|
-
const key = modelProfileKey(manifest.provider, manifest.model)
|
|
937
|
+
const key = `${modelProfileKey(manifest.provider, manifest.model)}\0${manifest.modelAwareness?.calibration.estimator ?? "chars-v1"}`;
|
|
926
938
|
const samples = this.volatileCalibration.get(key) ?? [];
|
|
927
939
|
samples.unshift({
|
|
928
940
|
estimatedTokens: manifest.estimatedInputTokens,
|
|
@@ -980,6 +992,7 @@ export class Ds4ContextRuntime {
|
|
|
980
992
|
}
|
|
981
993
|
const model = snapshotModel(ctx);
|
|
982
994
|
const activeModel = model ? this.resolveActiveModel(model) : undefined;
|
|
995
|
+
const estimator = model ? this.estimatorForModel(model) : CHARS_ESTIMATOR;
|
|
983
996
|
const budget = activeModel?.budget;
|
|
984
997
|
let effectiveEvent = preparedPrivacy.event;
|
|
985
998
|
let artifactReferences = [] as NonNullable<ContextManifest["artifacts"]>;
|
|
@@ -993,8 +1006,8 @@ export class Ds4ContextRuntime {
|
|
|
993
1006
|
preparedPrivacy.messageClassifications,
|
|
994
1007
|
this.config.artifacts.adaptiveBudget && budget ? {
|
|
995
1008
|
inputTokens: budget.activeInputBudget,
|
|
996
|
-
fixedTokens: estimateTextTokens(preparedPrivacy.systemPrompt) + 8
|
|
997
|
-
+ preparedPrivacy.tools.reduce((sum, tool) => sum + estimateObservedToolTokens(tool), 0),
|
|
1009
|
+
fixedTokens: estimator.estimateTextTokens(preparedPrivacy.systemPrompt) + 8
|
|
1010
|
+
+ preparedPrivacy.tools.reduce((sum, tool) => sum + estimateObservedToolTokens(tool, estimator), 0),
|
|
998
1011
|
} : undefined,
|
|
999
1012
|
);
|
|
1000
1013
|
effectiveEvent = { type: "context", messages: transformed.messages };
|
|
@@ -1043,6 +1056,7 @@ export class Ds4ContextRuntime {
|
|
|
1043
1056
|
createdAt: observedAt,
|
|
1044
1057
|
policyVersion: POLICY_VERSION,
|
|
1045
1058
|
plannerVersion: OBSERVER_PLANNER_VERSION,
|
|
1059
|
+
tokenEstimator: estimator,
|
|
1046
1060
|
...(activeModel ? {
|
|
1047
1061
|
profile: activeModel.awareness.profile,
|
|
1048
1062
|
budget: activeModel.budget,
|
|
@@ -1107,6 +1121,7 @@ export class Ds4ContextRuntime {
|
|
|
1107
1121
|
fixedTokens,
|
|
1108
1122
|
budget,
|
|
1109
1123
|
config: effectiveContextConfig,
|
|
1124
|
+
tokenEstimator: estimator,
|
|
1110
1125
|
pinnedMessageIndices,
|
|
1111
1126
|
supplementalMessages: dedupSupplementalMessages,
|
|
1112
1127
|
});
|
|
@@ -1138,13 +1153,14 @@ export class Ds4ContextRuntime {
|
|
|
1138
1153
|
fixedTokens,
|
|
1139
1154
|
budget,
|
|
1140
1155
|
config: effectiveContextConfig,
|
|
1156
|
+
tokenEstimator: estimator,
|
|
1141
1157
|
pinnedMessageIndices,
|
|
1142
1158
|
supplementalMessages: dedupSupplementalMessages,
|
|
1143
1159
|
cacheAwareTailTokens: decision.recentTailTokens,
|
|
1144
1160
|
});
|
|
1145
1161
|
const modelKey = modelProfileKey(model.provider, model.id);
|
|
1146
|
-
const nominalEpoch = this.cacheAwareEpochCost(nominalDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey);
|
|
1147
|
-
const extendedEpoch = this.cacheAwareEpochCost(extendedDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey);
|
|
1162
|
+
const nominalEpoch = this.cacheAwareEpochCost(nominalDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey, estimator);
|
|
1163
|
+
const extendedEpoch = this.cacheAwareEpochCost(extendedDedupPlan, fixedTokens, model.cost, effectiveContextConfig.cacheAware.expectedRequestsPerTurn, effectiveContextConfig.cacheAware.expectedTurnsPerEpoch, modelKey, estimator);
|
|
1148
1164
|
this.logger.debug("context.cache_aware_candidate", {
|
|
1149
1165
|
eligible: decision.eligible,
|
|
1150
1166
|
tailExtended: decision.tailExtended,
|
|
@@ -1181,6 +1197,7 @@ export class Ds4ContextRuntime {
|
|
|
1181
1197
|
fixedTokens,
|
|
1182
1198
|
budget,
|
|
1183
1199
|
config: effectiveContextConfig,
|
|
1200
|
+
tokenEstimator: estimator,
|
|
1184
1201
|
pinnedMessageIndices,
|
|
1185
1202
|
supplementalMessages: dedupSupplementalMessages,
|
|
1186
1203
|
cacheAwareTailTokens,
|
|
@@ -1275,7 +1292,7 @@ export class Ds4ContextRuntime {
|
|
|
1275
1292
|
if (sanitized.blockedBlocks > 0) {
|
|
1276
1293
|
privacyExcludedSources.push({
|
|
1277
1294
|
sourceId: supplement.sourceIds[0],
|
|
1278
|
-
tokens: estimateMessagesTokens([supplement.message]),
|
|
1295
|
+
tokens: estimator.estimateMessagesTokens([supplement.message]),
|
|
1279
1296
|
kind: supplement.kind,
|
|
1280
1297
|
classification: sanitized.classification,
|
|
1281
1298
|
score: supplement.score,
|
|
@@ -1328,6 +1345,7 @@ export class Ds4ContextRuntime {
|
|
|
1328
1345
|
fixedTokens,
|
|
1329
1346
|
budget,
|
|
1330
1347
|
config: effectiveContextConfig,
|
|
1348
|
+
tokenEstimator: estimator,
|
|
1331
1349
|
pinnedMessageIndices,
|
|
1332
1350
|
supplementalMessages: rankedSupplementalMessages,
|
|
1333
1351
|
...(cacheAwareTailTokens !== undefined ? { cacheAwareTailTokens } : {}),
|
|
@@ -1408,7 +1426,7 @@ export class Ds4ContextRuntime {
|
|
|
1408
1426
|
const currentTokens: number[] = [];
|
|
1409
1427
|
for (const message of plan.messages) {
|
|
1410
1428
|
currentHashes.push(fingerprint(message));
|
|
1411
|
-
currentTokens.push(estimateMessagesTokens([message]));
|
|
1429
|
+
currentTokens.push(estimator.estimateMessagesTokens([message]));
|
|
1412
1430
|
}
|
|
1413
1431
|
const reusablePrefixTokens = estimateReusablePrefixTokens(
|
|
1414
1432
|
previousHashes,
|
|
@@ -1495,6 +1513,7 @@ export class Ds4ContextRuntime {
|
|
|
1495
1513
|
createdAt: observedAt,
|
|
1496
1514
|
policyVersion: POLICY_VERSION,
|
|
1497
1515
|
plannerVersion: PLANNER_VERSION,
|
|
1516
|
+
tokenEstimator: estimator,
|
|
1498
1517
|
...(activeModel ? {
|
|
1499
1518
|
profile: activeModel.awareness.profile,
|
|
1500
1519
|
budget: activeModel.budget,
|
|
@@ -1542,7 +1561,7 @@ export class Ds4ContextRuntime {
|
|
|
1542
1561
|
estimatedMessageTokens: manifest.composition.messageTokens,
|
|
1543
1562
|
originalMessageCount: manifest.planning?.originalMessageCount ?? historyEvent.messages.length,
|
|
1544
1563
|
originalEstimatedMessageTokens: manifest.planning?.originalMessageTokens
|
|
1545
|
-
?? estimateMessagesTokens(historyEvent.messages),
|
|
1564
|
+
?? estimator.estimateMessagesTokens(historyEvent.messages),
|
|
1546
1565
|
...(usage?.tokens !== null && usage?.tokens !== undefined ? { reportedTokens: usage.tokens } : {}),
|
|
1547
1566
|
...(manifest.planning?.durationMs !== undefined
|
|
1548
1567
|
? { planningDurationMs: manifest.planning.durationMs }
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
import { createRequire } from "node:module";
|
|
2
|
+
import type { Tiktoken } from "js-tiktoken/lite";
|
|
3
|
+
import { createO200kEstimator } from "ds4-context-core/core/bpe-token-estimator";
|
|
4
|
+
|
|
5
|
+
const load = createRequire(import.meta.url);
|
|
6
|
+
|
|
7
|
+
let encoder: Tiktoken | undefined;
|
|
8
|
+
|
|
9
|
+
/**
|
|
10
|
+
* Opt-in OpenAI o200k text estimator. Model-specific serializers, image tokens,
|
|
11
|
+
* reasoning, and provider-specific wrappers are still estimated, not exact.
|
|
12
|
+
* No remote vocabulary download or provider request is performed.
|
|
13
|
+
*/
|
|
14
|
+
export const O200K_ESTIMATOR = createO200kEstimator((text) => {
|
|
15
|
+
if (!encoder) {
|
|
16
|
+
// CJS exports are loaded on first opt-in use, not for every Pi session.
|
|
17
|
+
const { Tiktoken: Encoder } = load("js-tiktoken/lite") as typeof import("js-tiktoken/lite");
|
|
18
|
+
const ranks = load("js-tiktoken/ranks/o200k_base") as {
|
|
19
|
+
pat_str: string; special_tokens: Record<string, number>; bpe_ranks: string;
|
|
20
|
+
};
|
|
21
|
+
encoder = new Encoder(ranks);
|
|
22
|
+
}
|
|
23
|
+
return encoder.encode(text).length;
|
|
24
|
+
});
|
|
@@ -9,7 +9,12 @@ import type {
|
|
|
9
9
|
SessionEntry,
|
|
10
10
|
} from "@earendil-works/pi-coding-agent";
|
|
11
11
|
import type { Api, Model } from "@earendil-works/pi-ai";
|
|
12
|
-
import
|
|
12
|
+
import {
|
|
13
|
+
DEFAULT_COMPACTION_MAX_OPERATION_INPUT_TOKENS,
|
|
14
|
+
DEFAULT_COMPACTION_MAX_REQUEST_INPUT_TOKENS,
|
|
15
|
+
type CompactionThinkingLevel,
|
|
16
|
+
type Ds4ContextConfig,
|
|
17
|
+
} from "ds4-context-core/config/config";
|
|
13
18
|
import { calculateContextBudget, type ContextBudget } from "ds4-context-core/core/budget-manager";
|
|
14
19
|
import { createModelProfile, type ModelDescriptor } from "ds4-context-core/core/model-profile";
|
|
15
20
|
import type { ContextManifest } from "ds4-context-core/manifest/context-manifest";
|
|
@@ -80,6 +85,8 @@ export interface CompactionDiagnostics {
|
|
|
80
85
|
validate: boolean;
|
|
81
86
|
preserveRecentVerbatim: boolean;
|
|
82
87
|
segmentTargetTokens: number;
|
|
88
|
+
maxRequestInputTokens: number;
|
|
89
|
+
maxOperationInputTokens: number;
|
|
83
90
|
phase: CompactionPhase;
|
|
84
91
|
trigger?: CompactionTrigger;
|
|
85
92
|
summaryId?: string;
|
|
@@ -91,6 +98,8 @@ export interface CompactionDiagnostics {
|
|
|
91
98
|
completedAt?: number;
|
|
92
99
|
lastError?: string;
|
|
93
100
|
inputBudgetTokens?: number;
|
|
101
|
+
requestInputLimitTokens?: number;
|
|
102
|
+
operationInputTokens?: number;
|
|
94
103
|
sourcePromptTokens?: number;
|
|
95
104
|
segmentCount?: number;
|
|
96
105
|
aggregateCalls?: number;
|
|
@@ -189,9 +198,20 @@ interface CompactionCoordinatorDependencies {
|
|
|
189
198
|
|
|
190
199
|
type MutableCompactionState = Omit<
|
|
191
200
|
CompactionDiagnostics,
|
|
192
|
-
|
|
201
|
+
| "enabled"
|
|
202
|
+
| "validate"
|
|
203
|
+
| "preserveRecentVerbatim"
|
|
204
|
+
| "segmentTargetTokens"
|
|
205
|
+
| "maxRequestInputTokens"
|
|
206
|
+
| "maxOperationInputTokens"
|
|
207
|
+
| "proactiveEligible"
|
|
193
208
|
>;
|
|
194
209
|
|
|
210
|
+
interface CompactionOperationInputBudget {
|
|
211
|
+
limitTokens: number;
|
|
212
|
+
usedTokens: number;
|
|
213
|
+
}
|
|
214
|
+
|
|
195
215
|
function classifiedSummary(content: string, classification: PrivacyClassification): string {
|
|
196
216
|
return classification === "normal"
|
|
197
217
|
? content
|
|
@@ -260,6 +280,14 @@ export class CompactionCoordinator {
|
|
|
260
280
|
}
|
|
261
281
|
};
|
|
262
282
|
const maxConcurrentSegments = config.compaction.maxConcurrentSegments ?? 2;
|
|
283
|
+
const maxRequestInputTokens = config.compaction.maxRequestInputTokens
|
|
284
|
+
?? DEFAULT_COMPACTION_MAX_REQUEST_INPUT_TOKENS;
|
|
285
|
+
const maxOperationInputTokens = config.compaction.maxOperationInputTokens
|
|
286
|
+
?? DEFAULT_COMPACTION_MAX_OPERATION_INPUT_TOKENS;
|
|
287
|
+
const operationInputBudget: CompactionOperationInputBudget = {
|
|
288
|
+
limitTokens: maxOperationInputTokens,
|
|
289
|
+
usedTokens: 0,
|
|
290
|
+
};
|
|
263
291
|
const trigger: CompactionTrigger = this.proactiveRequested ? "proactive" : event.reason;
|
|
264
292
|
this.state = {
|
|
265
293
|
phase: "generating",
|
|
@@ -272,18 +300,20 @@ export class CompactionCoordinator {
|
|
|
272
300
|
summaryCalls: 0,
|
|
273
301
|
provider: model.provider,
|
|
274
302
|
model: model.id,
|
|
275
|
-
inputBudgetMode: config.compaction.inputBudget ?? "
|
|
303
|
+
inputBudgetMode: config.compaction.inputBudget ?? "context",
|
|
276
304
|
maxConcurrentSegments,
|
|
305
|
+
operationInputTokens: 0,
|
|
277
306
|
timings,
|
|
278
307
|
};
|
|
279
308
|
|
|
280
309
|
try {
|
|
281
|
-
const { source, inputBudgetTokens, wholePlan, directPlan, segmentPlans } = await measure("preparationMs", () => {
|
|
310
|
+
const { source, inputBudgetTokens, requestInputLimitTokens, wholePlan, directPlan, segmentPlans } = await measure("preparationMs", () => {
|
|
282
311
|
if (event.signal.aborted) throw new Error("Compaction summary generation aborted");
|
|
283
312
|
this.dependencies.syncSessionIndex(ctx);
|
|
284
313
|
const source = prepareCompactionSource(event);
|
|
285
314
|
const inputBudgetTokens = this.inputBudgetTokens(model);
|
|
286
315
|
if (inputBudgetTokens <= 0) throw new Error("Active model has no safe compaction input budget");
|
|
316
|
+
const requestInputLimitTokens = Math.min(inputBudgetTokens, maxRequestInputTokens);
|
|
287
317
|
const wholePlan = this.buildSegmentPlan(source, event, model.provider);
|
|
288
318
|
const update = (config.compaction.directUpdate ?? true) && source.previousSummary
|
|
289
319
|
? this.buildSegmentPlan({
|
|
@@ -292,18 +322,21 @@ export class CompactionCoordinator {
|
|
|
292
322
|
segmentModifiedFiles: source.modifiedFiles,
|
|
293
323
|
}, event, model.provider, source.previousSummary)
|
|
294
324
|
: undefined;
|
|
295
|
-
const directPlan = update && update.promptTokens <=
|
|
325
|
+
const directPlan = update && update.promptTokens <= requestInputLimitTokens ? update : undefined;
|
|
296
326
|
this.state = {
|
|
297
327
|
...this.state,
|
|
298
328
|
path: directPlan ? "direct-update" : "hierarchical",
|
|
299
329
|
sourceEntries: source.sourceEntryIds.length,
|
|
300
330
|
inputBudgetTokens,
|
|
331
|
+
requestInputLimitTokens,
|
|
301
332
|
sourcePromptTokens: wholePlan.promptTokens,
|
|
302
333
|
...(update ? { directPromptTokens: update.promptTokens } : {}),
|
|
303
334
|
};
|
|
304
|
-
const segmentPlans = directPlan
|
|
335
|
+
const segmentPlans = directPlan
|
|
336
|
+
? []
|
|
337
|
+
: this.partitionSegmentPlans(source, wholePlan, event, model.provider, requestInputLimitTokens);
|
|
305
338
|
this.state.segmentCount = segmentPlans.length;
|
|
306
|
-
return { source, inputBudgetTokens, wholePlan, directPlan, segmentPlans };
|
|
339
|
+
return { source, inputBudgetTokens, requestInputLimitTokens, wholePlan, directPlan, segmentPlans };
|
|
307
340
|
});
|
|
308
341
|
|
|
309
342
|
const usedIds = new Set(this.graphRecords.keys());
|
|
@@ -352,6 +385,9 @@ export class CompactionCoordinator {
|
|
|
352
385
|
ctx,
|
|
353
386
|
model,
|
|
354
387
|
thinking: config.compaction.summary?.thinking,
|
|
388
|
+
promptTokens: plan.promptTokens,
|
|
389
|
+
requestInputLimitTokens,
|
|
390
|
+
operationInputBudget,
|
|
355
391
|
}),
|
|
356
392
|
));
|
|
357
393
|
const generatedNodes: EmbeddedSummaryNode[] = results.map((generated, index) => {
|
|
@@ -385,7 +421,8 @@ export class CompactionCoordinator {
|
|
|
385
421
|
model: model.id,
|
|
386
422
|
readFiles: source.readFiles,
|
|
387
423
|
modifiedFiles: source.modifiedFiles,
|
|
388
|
-
|
|
424
|
+
requestInputLimitTokens,
|
|
425
|
+
operationInputBudget,
|
|
389
426
|
nextId,
|
|
390
427
|
createdNodes,
|
|
391
428
|
usages,
|
|
@@ -698,6 +735,10 @@ export class CompactionCoordinator {
|
|
|
698
735
|
validate: config.compaction.validate,
|
|
699
736
|
preserveRecentVerbatim: config.compaction.preserveRecentVerbatim,
|
|
700
737
|
segmentTargetTokens: config.compaction.segmentTargetTokens,
|
|
738
|
+
maxRequestInputTokens: config.compaction.maxRequestInputTokens
|
|
739
|
+
?? DEFAULT_COMPACTION_MAX_REQUEST_INPUT_TOKENS,
|
|
740
|
+
maxOperationInputTokens: config.compaction.maxOperationInputTokens
|
|
741
|
+
?? DEFAULT_COMPACTION_MAX_OPERATION_INPUT_TOKENS,
|
|
701
742
|
...this.state,
|
|
702
743
|
...(contextTokens !== undefined ? { contextTokens } : {}),
|
|
703
744
|
...(budget ? { softLimitTokens: this.providerSoftLimit(budget) } : {}),
|
|
@@ -800,7 +841,7 @@ export class CompactionCoordinator {
|
|
|
800
841
|
),
|
|
801
842
|
};
|
|
802
843
|
const maxOutputTokens = Math.max(1, Math.min(config.context.maxSummaryTokens, model.maxTokens ?? config.context.maxSummaryTokens));
|
|
803
|
-
return compactionInputBudget(resolved.budget, maxOutputTokens, config.compaction.inputBudget ?? "
|
|
844
|
+
return compactionInputBudget(resolved.budget, maxOutputTokens, config.compaction.inputBudget ?? "context");
|
|
804
845
|
}
|
|
805
846
|
|
|
806
847
|
private classify(text: string, provider: string): {
|
|
@@ -870,9 +911,13 @@ export class CompactionCoordinator {
|
|
|
870
911
|
wholePlan: SegmentGenerationPlan,
|
|
871
912
|
event: SessionBeforeCompactEvent,
|
|
872
913
|
provider: string,
|
|
873
|
-
|
|
914
|
+
requestInputLimitTokens: number,
|
|
874
915
|
): SegmentGenerationPlan[] {
|
|
875
|
-
|
|
916
|
+
const segmentTargetTokens = Math.min(
|
|
917
|
+
requestInputLimitTokens,
|
|
918
|
+
this.dependencies.config.compaction.segmentTargetTokens,
|
|
919
|
+
);
|
|
920
|
+
if (wholePlan.promptTokens <= segmentTargetTokens) return [wholePlan];
|
|
876
921
|
|
|
877
922
|
const groups = buildCompactionAtomicGroups(source.messages);
|
|
878
923
|
const plans: SegmentGenerationPlan[] = [];
|
|
@@ -886,7 +931,7 @@ export class CompactionCoordinator {
|
|
|
886
931
|
event,
|
|
887
932
|
provider,
|
|
888
933
|
);
|
|
889
|
-
if (candidatePlan.promptTokens <=
|
|
934
|
+
if (candidatePlan.promptTokens <= segmentTargetTokens) {
|
|
890
935
|
currentIndices = candidateIndices;
|
|
891
936
|
currentPlan = candidatePlan;
|
|
892
937
|
continue;
|
|
@@ -903,9 +948,9 @@ export class CompactionCoordinator {
|
|
|
903
948
|
event,
|
|
904
949
|
provider,
|
|
905
950
|
);
|
|
906
|
-
if (atomicPlan.promptTokens >
|
|
951
|
+
if (atomicPlan.promptTokens > requestInputLimitTokens) {
|
|
907
952
|
throw new Error(
|
|
908
|
-
`Compaction source contains an indivisible atomic group above the
|
|
953
|
+
`Compaction source contains an indivisible atomic group above the request input limit (promptTokens=${atomicPlan.promptTokens}; requestInputLimitTokens=${requestInputLimitTokens})`,
|
|
909
954
|
);
|
|
910
955
|
}
|
|
911
956
|
currentIndices = [...group.messageIndices];
|
|
@@ -974,7 +1019,8 @@ export class CompactionCoordinator {
|
|
|
974
1019
|
modelObject: Model<Api>;
|
|
975
1020
|
readFiles: readonly string[];
|
|
976
1021
|
modifiedFiles: readonly string[];
|
|
977
|
-
|
|
1022
|
+
requestInputLimitTokens: number;
|
|
1023
|
+
operationInputBudget: CompactionOperationInputBudget;
|
|
978
1024
|
nextId: () => string;
|
|
979
1025
|
createdNodes: EmbeddedSummaryNode[];
|
|
980
1026
|
usages: GeneratedSummary["usage"][];
|
|
@@ -1003,13 +1049,13 @@ export class CompactionCoordinator {
|
|
|
1003
1049
|
input.readFiles,
|
|
1004
1050
|
input.modifiedFiles,
|
|
1005
1051
|
);
|
|
1006
|
-
if (candidatePlan.promptTokens <= input.
|
|
1052
|
+
if (candidatePlan.promptTokens <= input.requestInputLimitTokens) {
|
|
1007
1053
|
current = candidate;
|
|
1008
1054
|
continue;
|
|
1009
1055
|
}
|
|
1010
1056
|
if (current.length === 1) {
|
|
1011
1057
|
throw new Error(
|
|
1012
|
-
`Compaction child summaries cannot be aggregated within the
|
|
1058
|
+
`Compaction child summaries cannot be aggregated within the request input limit (requestInputLimitTokens=${input.requestInputLimitTokens})`,
|
|
1013
1059
|
);
|
|
1014
1060
|
}
|
|
1015
1061
|
batches.push(current);
|
|
@@ -1034,7 +1080,7 @@ export class CompactionCoordinator {
|
|
|
1034
1080
|
input.readFiles,
|
|
1035
1081
|
input.modifiedFiles,
|
|
1036
1082
|
);
|
|
1037
|
-
if (plan.promptTokens > input.
|
|
1083
|
+
if (plan.promptTokens > input.requestInputLimitTokens) {
|
|
1038
1084
|
throw new Error("Compaction aggregate prompt exceeded its preflight input budget");
|
|
1039
1085
|
}
|
|
1040
1086
|
this.state.aggregateCalls = aggregateCalls;
|
|
@@ -1048,6 +1094,9 @@ export class CompactionCoordinator {
|
|
|
1048
1094
|
ctx: input.ctx,
|
|
1049
1095
|
model: input.modelObject,
|
|
1050
1096
|
thinking: this.dependencies.config.compaction.summary?.thinking,
|
|
1097
|
+
promptTokens: plan.promptTokens,
|
|
1098
|
+
requestInputLimitTokens: input.requestInputLimitTokens,
|
|
1099
|
+
operationInputBudget: input.operationInputBudget,
|
|
1051
1100
|
});
|
|
1052
1101
|
input.usages.push(generated.usage);
|
|
1053
1102
|
const aggregateNode: EmbeddedSummaryNode = {
|
|
@@ -1088,7 +1137,15 @@ export class CompactionCoordinator {
|
|
|
1088
1137
|
ctx: ExtensionContext;
|
|
1089
1138
|
model: Model<Api>;
|
|
1090
1139
|
thinking?: CompactionThinkingLevel;
|
|
1140
|
+
promptTokens: number;
|
|
1141
|
+
requestInputLimitTokens: number;
|
|
1142
|
+
operationInputBudget: CompactionOperationInputBudget;
|
|
1091
1143
|
}): Promise<GeneratedSummary> {
|
|
1144
|
+
if (input.promptTokens > input.requestInputLimitTokens) {
|
|
1145
|
+
throw new Error(
|
|
1146
|
+
`Compaction ${input.stage} prompt exceeded the request input limit (promptTokens=${input.promptTokens}; requestInputLimitTokens=${input.requestInputLimitTokens})`,
|
|
1147
|
+
);
|
|
1148
|
+
}
|
|
1092
1149
|
this.state.summaryCalls = (this.state.summaryCalls ?? 0) + 1;
|
|
1093
1150
|
return generateValidatedSummary({
|
|
1094
1151
|
...input,
|
|
@@ -1096,6 +1153,16 @@ export class CompactionCoordinator {
|
|
|
1096
1153
|
maxSummaryTokens: this.dependencies.config.context.maxSummaryTokens,
|
|
1097
1154
|
transport: this.dependencies.config.compaction.transport,
|
|
1098
1155
|
now: this.dependencies.now,
|
|
1156
|
+
onAttempt: (diagnostic) => {
|
|
1157
|
+
const nextInputTokens = input.operationInputBudget.usedTokens + input.promptTokens;
|
|
1158
|
+
if (nextInputTokens > input.operationInputBudget.limitTokens) {
|
|
1159
|
+
throw new Error(
|
|
1160
|
+
`Compaction operation input limit exceeded before ${diagnostic.stage} attempt ${diagnostic.attempt} (nextInputTokens=${nextInputTokens}; maxOperationInputTokens=${input.operationInputBudget.limitTokens})`,
|
|
1161
|
+
);
|
|
1162
|
+
}
|
|
1163
|
+
input.operationInputBudget.usedTokens = nextInputTokens;
|
|
1164
|
+
this.state.operationInputTokens = nextInputTokens;
|
|
1165
|
+
},
|
|
1099
1166
|
onTransportRetry: (diagnostic) => {
|
|
1100
1167
|
this.state.transportRetries = (this.state.transportRetries ?? 0) + 1;
|
|
1101
1168
|
this.dependencies.logger.debug("compaction.transport_retry", { ...diagnostic });
|
|
@@ -1205,6 +1272,10 @@ export function defaultCompactionDiagnostics(config: Ds4ContextConfig): Compacti
|
|
|
1205
1272
|
validate: config.compaction.validate,
|
|
1206
1273
|
preserveRecentVerbatim: config.compaction.preserveRecentVerbatim,
|
|
1207
1274
|
segmentTargetTokens: config.compaction.segmentTargetTokens,
|
|
1275
|
+
maxRequestInputTokens: config.compaction.maxRequestInputTokens
|
|
1276
|
+
?? DEFAULT_COMPACTION_MAX_REQUEST_INPUT_TOKENS,
|
|
1277
|
+
maxOperationInputTokens: config.compaction.maxOperationInputTokens
|
|
1278
|
+
?? DEFAULT_COMPACTION_MAX_OPERATION_INPUT_TOKENS,
|
|
1208
1279
|
phase: "idle",
|
|
1209
1280
|
proactiveEligible: false,
|
|
1210
1281
|
};
|
|
@@ -8,7 +8,7 @@ import {
|
|
|
8
8
|
import type { ContextConfig } from "ds4-context-core/config/config";
|
|
9
9
|
import { calculateContextBudget, type ContextBudget } from "ds4-context-core/core/budget-manager";
|
|
10
10
|
import { createModelProfile, type ModelProfile } from "ds4-context-core/core/model-profile";
|
|
11
|
-
import { estimateMessageTokens } from "ds4-context-core/core/token-estimator";
|
|
11
|
+
import { estimateMessageTokens, type TokenEstimator } from "ds4-context-core/core/token-estimator";
|
|
12
12
|
import type {
|
|
13
13
|
ArtifactManifestRef,
|
|
14
14
|
ContextManifest,
|
|
@@ -54,6 +54,7 @@ export interface BuildPiObserverManifestOptions {
|
|
|
54
54
|
profile?: ModelProfile;
|
|
55
55
|
budget?: ContextBudget;
|
|
56
56
|
modelAwareness?: ModelAwarenessManifest;
|
|
57
|
+
tokenEstimator?: TokenEstimator;
|
|
57
58
|
plan?: ManagedContextPlan<PiAgentMessage>;
|
|
58
59
|
projectRevision?: ProjectRevision;
|
|
59
60
|
pins?: readonly PinManifestRef[];
|
|
@@ -386,6 +387,7 @@ export function buildPiObserverManifest(options: BuildPiObserverManifestOptions)
|
|
|
386
387
|
profile,
|
|
387
388
|
budget,
|
|
388
389
|
systemPrompt: options.systemPrompt ?? options.ctx.getSystemPrompt(),
|
|
390
|
+
...(options.tokenEstimator ? { tokenEstimator: options.tokenEstimator } : {}),
|
|
389
391
|
...(options.systemClassification ? { systemClassification: options.systemClassification } : {}),
|
|
390
392
|
...(options.systemPrivacyReason ? { systemPrivacyReason: options.systemPrivacyReason } : {}),
|
|
391
393
|
tools: options.tools ?? activeTools(options.pi),
|
|
@@ -13,16 +13,16 @@ import {
|
|
|
13
13
|
type SummaryValidationResult,
|
|
14
14
|
} from "ds4-context-core/compaction/summary-contract";
|
|
15
15
|
|
|
16
|
-
export const DEFAULT_COMPACTION_TRANSPORT_MAX_ATTEMPTS =
|
|
16
|
+
export const DEFAULT_COMPACTION_TRANSPORT_MAX_ATTEMPTS = 4;
|
|
17
17
|
export const DEFAULT_COMPACTION_TRANSPORT_BASE_DELAY_MS = 2000;
|
|
18
18
|
export const COMPACTION_TRANSPORT_MAX_DELAY_MS = 60_000;
|
|
19
19
|
|
|
20
20
|
/**
|
|
21
|
-
* Transport retry policy for compaction summary requests:
|
|
22
|
-
* (
|
|
21
|
+
* Transport retry policy for compaction summary requests: four total attempts
|
|
22
|
+
* (the initial call plus three retries), 2000 ms base delay, exponential backoff, abort-aware.
|
|
23
23
|
*/
|
|
24
24
|
export interface CompactionTransportPolicy {
|
|
25
|
-
/** Total attempts for transport-classified failures. Default:
|
|
25
|
+
/** Total attempts for transport-classified failures. Default: 4. */
|
|
26
26
|
maxAttempts?: number;
|
|
27
27
|
/** Base backoff delay in ms, doubled per attempt. Default: 2000. */
|
|
28
28
|
baseDelayMs?: number;
|
|
@@ -54,6 +54,12 @@ export interface CompactionTransportRetryDiagnostic {
|
|
|
54
54
|
delayMs: number;
|
|
55
55
|
}
|
|
56
56
|
|
|
57
|
+
export interface CompactionAttemptDiagnostic {
|
|
58
|
+
stage: "segment" | "aggregate" | "update";
|
|
59
|
+
attempt: number;
|
|
60
|
+
maxAttempts: number;
|
|
61
|
+
}
|
|
62
|
+
|
|
57
63
|
export interface GenerateValidatedSummaryInput {
|
|
58
64
|
stage: "segment" | "aggregate" | "update";
|
|
59
65
|
prompt: string;
|
|
@@ -68,9 +74,11 @@ export interface GenerateValidatedSummaryInput {
|
|
|
68
74
|
model?: Model<Api>;
|
|
69
75
|
/** Reasoning level for the summary request; `off` (default) keeps the pre-existing request shape. */
|
|
70
76
|
thinking?: CompactionThinkingLevel;
|
|
71
|
-
/** Transport-only retry policy;
|
|
77
|
+
/** Transport-only retry policy; four total attempts by default. */
|
|
72
78
|
transport?: CompactionTransportPolicy;
|
|
73
79
|
now: () => number;
|
|
80
|
+
/** Called synchronously before every provider attempt, including retries. */
|
|
81
|
+
onAttempt?: (diagnostic: CompactionAttemptDiagnostic) => void;
|
|
74
82
|
onTransportRetry?: (diagnostic: CompactionTransportRetryDiagnostic) => void;
|
|
75
83
|
}
|
|
76
84
|
|
|
@@ -201,6 +209,7 @@ export async function generateValidatedSummary(
|
|
|
201
209
|
for (;;) {
|
|
202
210
|
if (input.event.signal.aborted) throw abortedError();
|
|
203
211
|
attempt++;
|
|
212
|
+
input.onAttempt?.({ stage: input.stage, attempt, maxAttempts });
|
|
204
213
|
try {
|
|
205
214
|
response = await input.ctx.modelRegistry.complete(
|
|
206
215
|
model,
|