ds4-context-engine 0.3.7 → 0.3.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -16,9 +16,9 @@ bounded active context with provenance
16
16
  Pi provider
17
17
  ```
18
18
 
19
- > **Project status:** The coordinated `0.3.5` release adds bounded compaction optimizations: one validated previous-summary/source update when the complete prompt fits, a dedicated calibrated input budget, up to two concurrent segments, and metadata-only phase timings in `/context compaction`. Canonical history, SQLite schema 15 and runtime contracts remain unchanged; Pi remains pinned to `0.84.3`. See the [0.3.5 release record](docs/releases/0.3.5.md) for validation and publication status.
19
+ > **Project status:** The coordinated `0.3.9` release adds hard estimated-input limits for every compaction request and the complete compaction operation, including retries and concurrent segments. The conservative context-fill input budget is now the default, while the prior hard-limit budget remains opt-in. Canonical history, SQLite schema 16 and runtime contracts remain unchanged from `0.3.8`; Pi remains pinned to `0.84.3`. See the [0.3.9 release record](docs/releases/0.3.9.md) for validation and publication status.
20
20
 
21
- **New compaction defaults:** `compaction.directUpdate=true`, `compaction.inputBudget="summary"`, `compaction.maxConcurrentSegments=2`. Existing compaction/master switches still apply. See [latency controls and compatibility](docs/COMPACTION.md#latency-controls) for the legacy-path settings. No real-provider speedup is claimed from mock tests. The five optional editing/reading/artifact/job features introduced in `0.3.4` remain default-off.
21
+ **Current compaction defaults:** `compaction.directUpdate=true`, `compaction.inputBudget="context"`, `compaction.segmentTargetTokens=30000`, `compaction.maxRequestInputTokens=64000`, `compaction.maxOperationInputTokens=2000000`, `compaction.maxConcurrentSegments=2`. Every DS4 provider attempt is bounded by the effective request limit, and the operation limit includes retries; `inputBudget="summary"` remains an explicit throughput-oriented opt-in. Existing compaction/master switches still apply. See [latency controls and compatibility](docs/COMPACTION.md#latency-controls). No real-provider speedup is claimed from mock tests. The five optional editing/reading/artifact/job features introduced in `0.3.4` remain default-off.
22
22
 
23
23
  ## Why DS4
24
24
 
@@ -331,9 +331,11 @@ The following example shows the main configuration groups. Omitted values use th
331
331
  "mode": "hierarchical",
332
332
  "validate": true,
333
333
  "segmentTargetTokens": 30000,
334
+ "maxRequestInputTokens": 64000,
335
+ "maxOperationInputTokens": 2000000,
334
336
  "preserveRecentVerbatim": true,
335
337
  "directUpdate": true,
336
- "inputBudget": "summary",
338
+ "inputBudget": "context",
337
339
  "maxConcurrentSegments": 2
338
340
  },
339
341
  "privacy": {
@@ -512,6 +514,10 @@ scripts package and release-readiness checks
512
514
  - [Roadmap 0.2.0](docs/ROADMAP_0.2.0.md)
513
515
  - [Release process](docs/RELEASING.md)
514
516
  - [0.2.0 release readiness](docs/RELEASE_READINESS_0.2.0.md)
517
+ - [0.3.9 release notes](docs/releases/0.3.9.md)
518
+ - [0.3.8 release notes](docs/releases/0.3.8.md)
519
+ - [0.3.7 release notes](docs/releases/0.3.7.md)
520
+ - [0.3.6 release notes](docs/releases/0.3.6.md)
515
521
  - [0.3.5 release notes](docs/releases/0.3.5.md)
516
522
  - [0.3.4 release notes](docs/releases/0.3.4.md)
517
523
  - [0.3.3 release notes](docs/releases/0.3.3.md)
@@ -535,7 +541,7 @@ scripts package and release-readiness checks
535
541
 
536
542
  The original M0–M13 roadmap is complete. `ds4-context-core` contains the compiled runtime-neutral implementation. M14 context-quality metrics, M15 rich symbol indexing, M16 hybrid semantic retrieval, M17 cross-session project memory, M18 learned-ranking shadow evaluation, M19's runtime adapter/conformance kit, and M20 opt-in local KV eligibility/replay are implemented on `main`. Learned active ranking remains promotion-gated, Pi reports local KV as unsupported, and static ranking/native completion stay authoritative on every failure.
537
543
 
538
- The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.3.5 adds bounded compaction updates, summary input headroom, concurrent segments and phase timings. The opt-in [anchored editing](docs/ANCHORED_EDITING.md) and [portable agent tools](docs/PORTABLE_AGENT_TOOLS.md) from 0.3.4 remain default-off, without backend rewind, forced sampling or operational KV integration. Confirmation, provenance, Pi fallback and canonical/configuration/SQLite/runtime contracts remain unchanged. The [0.2 readiness record](docs/RELEASE_READINESS_0.2.0.md) remains the compatibility baseline; the lexical planner stays available as the deterministic fallback.
544
+ The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.3.9 extends the bounded compaction updates, summary input headroom, concurrent segments and phase timings introduced in 0.3.5 with per-request and cumulative operation input limits; 0.3.8 adds indexed FTS key deletion without changing search results. The opt-in [anchored editing](docs/ANCHORED_EDITING.md) and [portable agent tools](docs/PORTABLE_AGENT_TOOLS.md) from 0.3.4 remain default-off, without backend rewind, forced sampling or operational KV integration. Confirmation, provenance, Pi fallback and canonical/configuration/SQLite/runtime contracts remain unchanged. The [0.2 readiness record](docs/RELEASE_READINESS_0.2.0.md) remains the compatibility baseline; the lexical planner stays available as the deterministic fallback.
539
545
 
540
546
  ## Contributing
541
547
 
@@ -14,6 +14,14 @@ Pi normally updates previous summary plus discarded messages in one request. DS4
14
14
  - Expose process-local metadata-only timings for preparation, segment/direct generation, aggregation, graph preparation/persistence and total DS4 hook duration, plus chosen path, effective provider/model, direct prompt size and concurrency. Timings use a monotonic clock, include retries and local validation within generation, and are wall times (not sums of overlapping calls). They are not canonical evidence, are not persisted in JSONL, and do not measure a subsequent native Pi fallback.
15
15
  - Preserve schema-v2 compaction metadata, the existing `task-state` kind, source/classification/validation contracts, Pi cut points, atomic tool exchanges, fresh routing IDs, retry policy, canonical JSONL and all-or-nothing graph preparation. `segmentSummaryId` refers to the update node itself on a direct update; do not invent a segment that was never generated.
16
16
 
17
+ ## Subsequent default revision
18
+
19
+ After operational testing, the default for `compaction.inputBudget` was changed from `summary` to `context`. The original `summary` mode remains available as an explicit throughput-oriented opt-in. The conservative default limits direct updates and indivisible atomic groups to the ordinary active input target, reducing request-size peaks at the cost of potentially more segment and aggregate calls. This revision changes only the default: calibrated hard limits, output headroom, atomicity, validation, and fail-closed behavior remain intact.
20
+
21
+ The hierarchical partitioner applies `min(compaction.segmentTargetTokens, effective request input limit)` to ordinary segments. A single indivisible message or complete tool exchange may exceed the target, but it is isolated and must still fit the effective request input limit.
22
+
23
+ Post-release hardening adds two additive safeguards. `maxRequestInputTokens` caps the estimated prompt size of every direct-update, segment, aggregate, and retry attempt after intersecting it with the selected model input budget. `maxOperationInputTokens` caps the cumulative estimated prompt input reserved across the whole DS4 operation, including retries and concurrent segment attempts. Crossing either limit fails closed to Pi without committing a partial summary graph.
24
+
17
25
  ## Compatibility and validation
18
26
 
19
27
  Configuration is additive; absent fields use the new defaults. The old scheduling path can be compared using `directUpdate=false`, `inputBudget=context`, and `maxConcurrentSegments=1`. The compaction master switch still delegates to Pi when disabled. The original local implementation excluded versioning and publication; the user subsequently authorized the coordinated 0.3.5 release. No dependency upgrade, live database maintenance or Pi upgrade is part of this change.
@@ -0,0 +1,53 @@
1
+ # 063 — Resolve FTS key deletes through rowid mapping tables
2
+
3
+ **Date:** 2026-09-06
4
+ **Status:** Accepted
5
+ **Related:** [002](002-pi-jsonl-canonical-sqlite-rebuildable.md), [058](058-bounded-manifest-storage.md)
6
+
7
+ ## Context
8
+
9
+ The FTS5 shadow tables used by the session index and the project knowledge
10
+ index declare `entry_key` / `snippet_id` as `UNINDEXED` columns. FTS5 cannot
11
+ use the token vocabulary to locate a row by an `UNINDEXED` column, so every
12
+ `DELETE FROM …_fts WHERE entry_key = ?` is a full scan of the whole virtual
13
+ table, whose cost is proportional to the size of the entire FTS index.
14
+
15
+ For incremental appends the cost is small, but a session-index **rebuild**
16
+ (fork, path change, truncation, missing checkpoint) re-writes every entry:
17
+ a large session (11k entries, ~600MB database) paid thousands of full scans
18
+ and stalled for hours, surfacing as an indefinite hang of the agent session.
19
+
20
+ A first candidate — making `entry_key` an indexed FTS5 column and recreating
21
+ the tables in a migration — was benchmarked and rejected: FTS5 then treats the
22
+ key as searchable vocabulary, changing the `MATCH` surface (a term present
23
+ only in an entry key becomes a false-positive hit) and the `bm25` column
24
+ weight mapping of existing queries.
25
+
26
+ ## Decision
27
+
28
+ Keep the FTS table definitions byte-identical and add dedicated mapping
29
+ tables that derive the SQLite rowid of each FTS row:
30
+
31
+ - `entries_fts_keys(entry_key PRIMARY KEY, fts_rowid)`
32
+ - `project_snippets_fts_keys(snippet_id, project_path, fts_rowid)`
33
+
34
+ Deletes become `DELETE FROM …_fts WHERE rowid = <mapped>` (O(log n)) and
35
+ rebuild cleanups resolve stale rows through the same mapping joined on the
36
+ base tables. Insert paths upsert the mapping from the FTS `last_insert_rowid`
37
+ inside the same write transaction as the FTS insert, so mapping and index
38
+ cannot diverge (a crash rolls back both).
39
+
40
+ The mapping tables are purely derived, rebuildable state: migration 16
41
+ backfills them from the live FTS rows with one scan, in a single
42
+ transaction — a failure rolls back and the previous schema keeps working.
43
+
44
+ ## Consequences
45
+
46
+ - Per-row FTS deletes: ~0.23 ms instead of ~40 ms–2 s (cold) per row;
47
+ a 11k-row rebuild drops from hours to ~3 s.
48
+ - The FTS schema, the `MATCH` surface and the `bm25` weight mapping are
49
+ unchanged (verified by the migration and repository tests).
50
+ - Existing databases upgrade non-destructively on first open after the
51
+ extension update; no user data is stored in the mapping tables.
52
+ - Future migrations that drop/recreate FTS tables must rebuild the mapping
53
+ tables in the same migration.
@@ -66,5 +66,6 @@ The initial decisions from the development plan are accepted:
66
66
  | [060](060-optional-portable-agent-tools.md) | Opt in to edit reports, adaptive reads/results and session-owned local jobs | Accepted |
67
67
  | [061](061-compaction-latency.md) | Bound compaction update calls, input budgets, concurrent segments and phase timings | Accepted |
68
68
  | [062](062-cache-aware-context-planning.md) | Opt-in cache-aware tail planning using model pricing and observed cache shares | Accepted |
69
+ | [063](063-fts-rowid-key-mappings.md) | Resolve FTS key deletes through rowid mapping tables | Accepted |
69
70
 
70
71
  Each decision will receive a dedicated record when implementation pressure introduces alternatives or consequences not already covered by the development plan.
@@ -8,12 +8,12 @@ DS4 intercepts Pi's `session_before_compact` event but preserves Pi's cut-point
8
8
  2. DS4 maps every source message by exact fingerprint to a canonical branch entry ID.
9
9
  3. Pi's serializer converts the newly discarded span to bounded conversation text; enabled privacy policy sanitizes conversation, previous summary, custom instructions, and file paths for the effective compaction provider (dedicated model when configured and eligible).
10
10
  4. DS4 estimates the **complete** sanitized request against a calibrated summary-specific input budget, including framing, instructions, file inventories and output contract. With `compaction.directUpdate` enabled, a previous summary plus new source that fits is updated and validated in **one call**, producing an immutable `task-state` node. No predecessor needs only the existing one-segment call. Oversized updates fall through to hierarchical planning; no additional source is truncated to force a fit.
11
- 5. The hierarchical path partitions oversized source into ordered contiguous segments. Individual messages are indivisible, and every tool call remains in the same atomic group as all matching results. Up to `compaction.maxConcurrentSegments` independent segment requests run concurrently (default 2). Each summary is validated against only its own sanitized evidence. Identities, source/child order and usage accumulation follow source order, not completion order. Cache retention stays disabled and every attempt has a fresh routing session ID.
11
+ 5. The hierarchical path partitions source into ordered contiguous segments capped at `min(compaction.segmentTargetTokens, requestInputLimitTokens)`, where the effective request limit is also bounded by the selected model input budget. The target is soft only for one indivisible atomic group: an individual message or complete tool call/result exchange may exceed it, but must still fit the effective request input limit and is isolated in its own segment. Up to `compaction.maxConcurrentSegments` independent segment requests run concurrently (default 2). Each summary is validated against only its own sanitized evidence. Identities, source/child order and usage accumulation follow source order, not completion order. Cache retention stays disabled and every attempt has a fresh routing session ID.
12
12
  6. Hierarchical requests recursively aggregate ordered children until one root remains: previous branch summary first, then new segments. Aggregation stays sequential and budget-checked. Direct updates instead link their single node to the predecessor and new canonical source IDs, without generating a synthetic segment or a separate aggregate. DS4-generated IDs, hashes, kinds, and graph levels never become model-visible evidence. A Pi-native predecessor is imported as an explicitly unverified branch node.
13
13
  7. The highest input classification wraps each generated node, then all nodes are persisted atomically as one `prepared` graph batch. Usage includes every returned direct/segment/aggregate request and transport replay. Pi receives only the final root text and still appends exactly one canonical `CompactionEntry` with `fromHook: true`.
14
14
  8. `session_compact` commits all nodes and associates the active root with the Pi entry; failure marks the complete prepared batch `failed`.
15
15
 
16
- Fan-out and fan-in are bounded to 32 segment requests, 64 aggregate requests, and 16 aggregate passes. The DS4 transport replay policy is: `compaction.transport.maxAttempts` (default 3) total attempts and `compaction.transport.baseDelayMs` (default 2000 ms, capped at 60 s) backoff, doubling per attempt, abort-aware. Replay never applies to input, usage, rate, authentication, validation, or output-limit failures. A base prompt, individual message, atomic tool exchange, pair of child summaries, or total operation that cannot fit within those limits fails closed. Any mapping, budget, model, output-limit, validation, abort, or storage error returns `undefined` from the hook, allowing Pi's default compaction to run. On segment failure or cancellation DS4 stops scheduling, aborts siblings and awaits all started workers before fallback; no partial graph is installed. Cancellation is cooperative: accepted provider work may still cost tokens, and a provider that ignores abort can delay settlement.
16
+ Fan-out and fan-in are bounded to 32 segment requests, 64 aggregate requests, and 16 aggregate passes. The DS4 transport replay policy is: `compaction.transport.maxAttempts` (default 4, matching Pi's initial call plus three retries) total attempts and `compaction.transport.baseDelayMs` (default 2000 ms, capped at 60 s) backoff, doubling per attempt, abort-aware. Replay never applies to input, usage, rate, authentication, validation, or output-limit failures. A base prompt, individual message, atomic tool exchange, pair of child summaries, or total operation that cannot fit within those limits fails closed. Any mapping, budget, model, output-limit, validation, abort, or storage error returns `undefined` from the hook, allowing Pi's default compaction to run. On segment failure or cancellation DS4 stops scheduling, aborts siblings and awaits all started workers before fallback; no partial graph is installed. Cancellation is cooperative: accepted provider work may still cost tokens, and a provider that ignores abort can delay settlement.
17
17
 
18
18
  ## Required summary contract
19
19
 
@@ -80,42 +80,49 @@ Semantics:
80
80
 
81
81
  ## Latency controls
82
82
 
83
- The coordinated `0.3.5` release introduces these additive defaults (absent from `0.3.4`):
83
+ The current defaults favor bounded request size while retaining direct updates and bounded parallelism:
84
84
 
85
85
  ```json
86
86
  {
87
87
  "compaction": {
88
88
  "directUpdate": true,
89
- "inputBudget": "summary",
89
+ "inputBudget": "context",
90
+ "segmentTargetTokens": 30000,
91
+ "maxRequestInputTokens": 64000,
92
+ "maxOperationInputTokens": 2000000,
90
93
  "maxConcurrentSegments": 2
91
94
  }
92
95
  }
93
96
  ```
94
97
 
95
- - `directUpdate`: one validated previous-summary plus new-source request when the entire prompt fits. Set false to always retain the segment-then-aggregate route. Validation or provider failures still fall back to Pi, not an unvalidated update.
96
- - `inputBudget`: `summary` uses calibrated `hardInputLimit` rather than the ordinary context fill target (`activeInputBudget`). `context` selects that legacy fill target. Both are additionally capped by `(context window - safety margin - actual summary output cap) / calibration ratio`, rounded down, and never exceed the configured/model hard limit. Existing output reservations remain conservative; ordinary session planning and proactive thresholds are unchanged. This reduces avoidable fragmentation, not a guarantee of provider fit or faster processing for a larger prompt.
98
+ `segmentTargetTokens` is the soft partition target for source segments. `maxRequestInputTokens` is the hard estimated-input cap applied to every direct-update, segment, aggregate, and retry attempt; the effective request limit is the minimum of this value and the safe model input budget. An indivisible message or complete tool exchange above that limit fails closed to Pi's native compaction rather than creating an oversized DS4 request.
99
+
100
+ `maxOperationInputTokens` bounds the sum of estimated prompt tokens reserved immediately before all provider attempts in one DS4 compaction, including transport retries and concurrently scheduled segments. Exceeding it aborts remaining DS4 work and falls back without committing a partial summary graph. Defaults are 64,000 per request and 2,000,000 per operation; both must be positive integers, and the operation limit must be at least the configured request limit.
101
+
102
+ - `directUpdate`: one validated previous-summary plus new-source request when the entire prompt fits the effective request input limit. Set false to always retain the segment-then-aggregate route. Validation or provider failures still fall back to Pi, not an unvalidated update.
103
+ - `inputBudget`: `context` (default) uses the ordinary context fill target (`activeInputBudget`). `summary` is an explicit throughput-oriented opt-in that permits the calibrated `hardInputLimit`. Both are additionally capped by `(context window - safety margin - actual summary output cap) / calibration ratio`, rounded down, and never exceed the configured/model hard limit. Existing output reservations remain conservative; ordinary session planning and proactive thresholds are unchanged. The conservative default reduces request-size peaks but can produce more segments and calls; it is not a guarantee of provider fit or lower total token use.
97
104
  - `maxConcurrentSegments`: integer **1–2**, default 2. Only independent segments overlap, including their retries. Aggregates do not run until their children have completed. Use 1 for sequential execution or providers with restrictive concurrent-request limits. Rate-limit failures are not transport-retried.
98
105
 
99
- For an old-path comparison set `directUpdate=false`, `inputBudget=context`, `maxConcurrentSegments=1`. All features remain behind the existing compaction/master switches. Settings are applied on session load; after upgrading the package or rebuilding a development checkout, fully restart Pi to avoid stale compiled-core modules. No schema migration is required.
106
+ For an old sequential-path comparison set `directUpdate=false`, `inputBudget=context`, `maxConcurrentSegments=1`. To reproduce the `0.3.5` throughput-oriented budget, set `inputBudget=summary`. All features remain behind the existing compaction/master switches. Settings are applied on session load; after upgrading the package or rebuilding a development checkout, fully restart Pi to avoid stale compiled-core modules. No schema migration is required.
100
107
 
101
108
  Mock-provider regression tests verify fewer calls, bounded overlap, exact budget boundaries, validation, privacy, immutable provenance and JSONL rebuild. They do **not** establish real-provider wall-time gains, semantic equivalence of generated summaries, or a guaranteed completion time. See [ADR-061](ADR/061-compaction-latency.md).
102
109
 
103
110
  ## Transport retry policy
104
111
 
105
- Summary requests are replayed only for transport-classified failures (thrown transport errors or `stopReason: "error"` responses whose message matches network/timeout patterns). The DS4 replay policy uses **three total attempts**, not three retries after the initial call, and can be tuned per deployment:
112
+ Summary requests are replayed only for transport-classified failures (thrown transport errors or `stopReason: "error"` responses whose message matches network/timeout patterns). The DS4 replay policy uses **four total attempts** (the initial call plus up to three retries), matching Pi's standard retry count, and can be tuned per deployment:
106
113
 
107
114
  ```json
108
115
  {
109
116
  "compaction": {
110
117
  "transport": {
111
- "maxAttempts": 3,
118
+ "maxAttempts": 4,
112
119
  "baseDelayMs": 2000
113
120
  }
114
121
  }
115
122
  }
116
123
  ```
117
124
 
118
- - `compaction.transport.maxAttempts`: total attempts per direct update, segment or aggregate call, integer 1–10, default 3. With 1, no transport failure is retried.
125
+ - `compaction.transport.maxAttempts`: total attempts per direct update, segment or aggregate call, integer 1–10, default 4. With 1, no transport failure is retried.
119
126
  - `compaction.transport.baseDelayMs`: base backoff before the first replay, integer 0–60000, default 2000. The delay doubles per attempt (2000, 4000, 8000, …) and is capped at 60 s.
120
127
  - Replays use a fresh routing session per attempt; diagnostics expose only stage, failed/next attempt, max attempts, and delay.
121
128
  - Aborts (including during backoff) never trigger replay; non-transport failures are never retried; usage is summed across replayed responses.
@@ -76,7 +76,7 @@ With privacy disabled, DS4 returns Pi's original `AgentMessage[]` when:
76
76
  - final estimated input exceeds the hard limit;
77
77
  - an unexpected adapter or planner exception occurs.
78
78
 
79
- Expected fallbacks are recorded in the Context Manifest. With privacy enabled, the fallback baseline is the sanitized native array—not raw Pi messages—and an unexpected privacy failure replaces content/payload fields instead of sending unchecked data. Observer mode disables planning but still enforces enabled privacy policy and records manifests/usage calibration.
79
+ Expected fallbacks are recorded in the Context Manifest. Because fail-open preserves the complete native array, DS4's hard input limit is not guaranteed in fallback when the current request, fixed overhead, or another mandatory atomic group is itself oversized. With privacy enabled, the fallback baseline is the sanitized native array—not raw Pi messages—and an unexpected privacy failure replaces content/payload fields instead of sending unchecked data. Observer mode disables planning but still enforces enabled privacy policy and records manifests/usage calibration.
80
80
 
81
81
  ## Quality measurement
82
82
 
@@ -70,6 +70,15 @@ same tail caps as 0.3.6.
70
70
  - Prefix-cache simulator: 7 passed.
71
71
  - Cache-aware runtime integration: 4 passed.
72
72
  - Config catalog/loader validation: passed.
73
+ - `pack:check` from a clean HEAD snapshot: core 239 files,
74
+ reference-adapter 7 files, engine 91 files; consumer install clean.
75
+ - Published and registry-verified: `ds4-context-core` 0.3.7,
76
+ `ds4-context-reference-adapter` 0.3.7, `ds4-context-engine` 0.3.7
77
+ (`registry:check` PASS at exact version; one transient registry replica
78
+ miss on first attempt, confirmed present via `npm view`).
79
+ - Release fix `7404a9a`: root `package.json` must depend exactly on
80
+ `ds4-context-core@0.3.7` (caught by the `pack:check` gate on the
81
+ coordinated-version requirement).
73
82
 
74
83
  ## Known limits
75
84
 
@@ -0,0 +1,64 @@
1
+ # Release 0.3.8 — Indexed FTS key deletes via rowid mappings (schema 16)
2
+
3
+ **Version analyzed:** DS4 Context Engine `0.3.8`
4
+ **Commit:** `18f36e1`
5
+ **Coordinated packages:** `ds4-context-core` 0.3.8, `ds4-context-reference-adapter` 0.3.8, `ds4-context-engine` 0.3.8
6
+
7
+ ## Summary
8
+
9
+ Bugfix release: FTS deletes were full virtual-table scans (FTS5 `UNINDEXED`
10
+ key columns), making session-index rebuilds on large sessions stall for
11
+ hours. The FTS tables are now rebuilt over rowid mapping tables, turning
12
+ every keyed delete into an indexed lookup. No behavior, configuration or
13
+ search-surface change: the fix is transparent and addresses the
14
+ fork/resume stall on large sessions.
15
+
16
+ ## Root cause
17
+
18
+ `DELETE FROM entries_fts WHERE entry_key = ?` cannot use the FTS5
19
+ vocabulary because `entry_key` (and `project_snippets_fts.snippet_id`) are
20
+ declared `UNINDEXED`. Every row removal scanned the entire FTS index.
21
+ A rebuild of a session with ~11k entries performed thousands of those
22
+ scans: measured 65k `pread64`/s against a 615MB database and ~100 full
23
+ index passes per ~140s, with an estimated **hours** to completion —
24
+ surfacing as an indefinite hang of the agent session after a session
25
+ fork/resume. The behavior class pre-dates 0.3.7; the 0.3.7 work sessions
26
+ made it visible by pushing the database into the hundreds of MB.
27
+
28
+ Indexing the keys directly inside FTS5 was benchmarked and **rejected**:
29
+ it would make key text part of the `MATCH` vocabulary (false positives)
30
+ and shift the `bm25` column weights. See ADR 063.
31
+
32
+ ## Changes
33
+
34
+ - Migration **16** `fts-rowid-key-mappings`: two derived mapping tables
35
+ (`entries_fts_keys`, `project_snippets_fts_keys`) backfilled from the
36
+ live FTS rows in one scan, inside a single transaction.
37
+ - `SessionIndexRepository`: per-row deletes go through the mapping
38
+ (`DELETE … WHERE rowid = ?`); the rebuild stale-row cleanup joins the
39
+ mapping with `entries` and uses the session index; the mapping is
40
+ upserted from the FTS `last_insert_rowid` in the same transaction.
41
+ - `ProjectKnowledgeRepository`: same pattern for per-snippet deletes
42
+ (`replaceFile`) and per-project cleanup (`clearProject`).
43
+ - The FTS schemas, the `MATCH` queries and the `bm25` weights are
44
+ byte-identical to 0.3.7; search results do not change.
45
+ - Tests: `tests/integration/fts-rowid-keys.test.ts` covering the
46
+ v15 → v16 upgrade (data preservation, key ↔ rowid correctness, search
47
+ surface unchanged), rebuild/append/stale-cleanup consistency and the
48
+ snippet `replaceFile`/`clearProject` paths; version-pinned tests moved
49
+ to `CURRENT_SCHEMA_VERSION`.
50
+
51
+ ## Compatibility and migration
52
+
53
+ - **Non-destructive by design:** the FTS indexes are derived state; the
54
+ migration is transactional, so on any failure the previous schema
55
+ remains intact and the extension keeps working with the old behavior.
56
+ - **No data stored in the mapping tables:** the base tables remain the
57
+ single source of truth.
58
+ - Upgrade runs automatically on first open after the extension update
59
+ (measured 467 ms on a 615MB database with 31k entries / 62k snippets;
60
+ no user-visible step required).
61
+ - The mapping backfill is one scan; a very large database adds only
62
+ seconds to the first open after update.
63
+ - Existing guarantees (fail-open, atomicity, privacy, pins, compaction,
64
+ hard limits) are unchanged.
@@ -0,0 +1,80 @@
1
+ # Release 0.3.9 — Bounded compaction request and operation input
2
+
3
+ **Version analyzed:** DS4 Context Engine `0.3.9`
4
+ **Commit:** (recorded after pack and registry verification)
5
+ **Coordinated packages:** `ds4-context-core` 0.3.9, `ds4-context-reference-adapter` 0.3.9, `ds4-context-engine` 0.3.9
6
+
7
+ ## Summary
8
+
9
+ Hardens custom compaction against request-size and total-operation token spikes.
10
+ Every direct-update, segment, aggregate and retry attempt is now bounded by an
11
+ effective estimated-input limit. A second cumulative budget covers the complete
12
+ compaction operation, including concurrent work and transport retries.
13
+
14
+ The default compaction input budget changes from the summary hard limit to the
15
+ ordinary context-fill budget. The former behavior remains available explicitly
16
+ for users who accept larger provider requests in exchange for fewer hierarchy
17
+ levels.
18
+
19
+ ## Changes
20
+
21
+ - Added `compaction.maxRequestInputTokens`, default 64,000, as a hard ceiling on
22
+ estimated input for each compaction provider request. The effective ceiling is
23
+ the minimum of this setting and the selected model-aware compaction budget.
24
+ - Added `compaction.maxOperationInputTokens`, default 2,000,000, for cumulative
25
+ estimated input across the complete operation. Every accepted attempt reserves
26
+ its input before provider dispatch, so retries and concurrent segments count.
27
+ - Configuration validation requires both limits to be positive integers and the
28
+ operation limit to be at least the configured request limit.
29
+ - Changed the default `compaction.inputBudget` from `summary` to `context`.
30
+ Explicit `summary` configurations retain the previous calibrated-hard-limit
31
+ behavior.
32
+ - Changed the default `compaction.transport.maxAttempts` from 3 to 4. Attempts
33
+ remain covered by the cumulative operation budget.
34
+ - Segment partitioning now uses the minimum of `segmentTargetTokens` and the
35
+ effective request limit. An indivisible message or complete tool exchange that
36
+ cannot fit fails closed to Pi's native compaction.
37
+ - `/context compaction` reports the effective request limit and cumulative
38
+ operation input consumed.
39
+ - Direct update, segment and aggregate paths continue to use strict validation,
40
+ privacy filtering, immutable provenance and all-or-nothing graph installation.
41
+
42
+ ## Validation
43
+
44
+ Before publication, the coordinated release passed:
45
+
46
+ - TypeScript builds and typecheck for the root, core and reference adapter;
47
+ - 85 Vitest files with 561 tests;
48
+ - deterministic context-quality comparison;
49
+ - context-persistence schema measurement;
50
+ - clean-consumer package verification for all three tarballs;
51
+ - whitespace/error-marker checks and coordinated manifest inspection.
52
+
53
+ The package verifier also checks the new defaults and both input-budget modes
54
+ from the packed `ds4-context-core` artifact.
55
+
56
+ ## Compatibility and migration
57
+
58
+ - No SQLite migration: schema 16 remains current.
59
+ - No Pi, runtime-adapter, history, persistence-tool or summary-contract version
60
+ changes.
61
+ - Existing configurations that explicitly set `inputBudget` or transport retry
62
+ attempts retain those values.
63
+ - New limit fields are additive. Configurations that omit them receive the
64
+ documented defaults.
65
+ - Pi JSONL remains canonical and custom-compaction failure still delegates to
66
+ Pi's native fallback without installing a partial Summary Graph.
67
+ - A full Pi restart is required after upgrading so compiled core modules and
68
+ session configuration are reloaded.
69
+
70
+ ## Known limits
71
+
72
+ - If the current request, fixed overhead or another mandatory atomic group is
73
+ itself above the planner hard limit, fail-open context planning preserves the
74
+ native context; DS4 cannot transparently split an in-flight Pi tool loop.
75
+ - Provider work accepted before a concurrent sibling fails may still be billed,
76
+ even though DS4 aborts and awaits all started workers before fallback.
77
+ - Limits use DS4's calibrated token estimator. Provider-side tokenization and
78
+ billing remain authoritative.
79
+ - Mock-provider tests establish bounded scheduling and request accounting, not
80
+ real-provider latency or semantic-equivalence guarantees.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ds4-context-engine",
3
- "version": "0.3.7",
3
+ "version": "0.3.9",
4
4
  "description": "Non-destructive, provider-independent context management for Pi.",
5
5
  "type": "module",
6
6
  "license": "MIT",
@@ -62,7 +62,7 @@
62
62
  ]
63
63
  },
64
64
  "dependencies": {
65
- "ds4-context-core": "0.3.7"
65
+ "ds4-context-core": "0.3.9"
66
66
  },
67
67
  "peerDependencies": {
68
68
  "@earendil-works/pi-ai": "0.84.3",
@@ -77,5 +77,9 @@
77
77
  },
78
78
  "engines": {
79
79
  "node": ">=22.19.0"
80
+ },
81
+ "allowScripts": {
82
+ "@google/genai@1.52.0": true,
83
+ "protobufjs@7.6.5": true
80
84
  }
81
85
  }
@@ -881,6 +881,8 @@ function formatCompaction(diagnostics: RuntimeDiagnostics, preview: boolean): st
881
881
  `Chosen path: ${compaction.path ?? "n/a"}`,
882
882
  `Input budget mode: ${compaction.inputBudgetMode ?? "n/a"}`,
883
883
  `Input budget: ${count(compaction.inputBudgetTokens)}`,
884
+ `Request input limit: ${count(compaction.requestInputLimitTokens)} (configured max ${count(compaction.maxRequestInputTokens)})`,
885
+ `Operation input: ${count(compaction.operationInputTokens)} / ${count(compaction.maxOperationInputTokens)}`,
884
886
  `Whole-source prompt: ${count(compaction.sourcePromptTokens)}`,
885
887
  `Direct-update prompt: ${count(compaction.directPromptTokens)}`,
886
888
  `Segment concurrency cap: ${count(compaction.maxConcurrentSegments)}`,
@@ -9,7 +9,12 @@ import type {
9
9
  SessionEntry,
10
10
  } from "@earendil-works/pi-coding-agent";
11
11
  import type { Api, Model } from "@earendil-works/pi-ai";
12
- import type { CompactionThinkingLevel, Ds4ContextConfig } from "ds4-context-core/config/config";
12
+ import {
13
+ DEFAULT_COMPACTION_MAX_OPERATION_INPUT_TOKENS,
14
+ DEFAULT_COMPACTION_MAX_REQUEST_INPUT_TOKENS,
15
+ type CompactionThinkingLevel,
16
+ type Ds4ContextConfig,
17
+ } from "ds4-context-core/config/config";
13
18
  import { calculateContextBudget, type ContextBudget } from "ds4-context-core/core/budget-manager";
14
19
  import { createModelProfile, type ModelDescriptor } from "ds4-context-core/core/model-profile";
15
20
  import type { ContextManifest } from "ds4-context-core/manifest/context-manifest";
@@ -80,6 +85,8 @@ export interface CompactionDiagnostics {
80
85
  validate: boolean;
81
86
  preserveRecentVerbatim: boolean;
82
87
  segmentTargetTokens: number;
88
+ maxRequestInputTokens: number;
89
+ maxOperationInputTokens: number;
83
90
  phase: CompactionPhase;
84
91
  trigger?: CompactionTrigger;
85
92
  summaryId?: string;
@@ -91,6 +98,8 @@ export interface CompactionDiagnostics {
91
98
  completedAt?: number;
92
99
  lastError?: string;
93
100
  inputBudgetTokens?: number;
101
+ requestInputLimitTokens?: number;
102
+ operationInputTokens?: number;
94
103
  sourcePromptTokens?: number;
95
104
  segmentCount?: number;
96
105
  aggregateCalls?: number;
@@ -189,9 +198,20 @@ interface CompactionCoordinatorDependencies {
189
198
 
190
199
  type MutableCompactionState = Omit<
191
200
  CompactionDiagnostics,
192
- "enabled" | "validate" | "preserveRecentVerbatim" | "segmentTargetTokens" | "proactiveEligible"
201
+ | "enabled"
202
+ | "validate"
203
+ | "preserveRecentVerbatim"
204
+ | "segmentTargetTokens"
205
+ | "maxRequestInputTokens"
206
+ | "maxOperationInputTokens"
207
+ | "proactiveEligible"
193
208
  >;
194
209
 
210
+ interface CompactionOperationInputBudget {
211
+ limitTokens: number;
212
+ usedTokens: number;
213
+ }
214
+
195
215
  function classifiedSummary(content: string, classification: PrivacyClassification): string {
196
216
  return classification === "normal"
197
217
  ? content
@@ -260,6 +280,14 @@ export class CompactionCoordinator {
260
280
  }
261
281
  };
262
282
  const maxConcurrentSegments = config.compaction.maxConcurrentSegments ?? 2;
283
+ const maxRequestInputTokens = config.compaction.maxRequestInputTokens
284
+ ?? DEFAULT_COMPACTION_MAX_REQUEST_INPUT_TOKENS;
285
+ const maxOperationInputTokens = config.compaction.maxOperationInputTokens
286
+ ?? DEFAULT_COMPACTION_MAX_OPERATION_INPUT_TOKENS;
287
+ const operationInputBudget: CompactionOperationInputBudget = {
288
+ limitTokens: maxOperationInputTokens,
289
+ usedTokens: 0,
290
+ };
263
291
  const trigger: CompactionTrigger = this.proactiveRequested ? "proactive" : event.reason;
264
292
  this.state = {
265
293
  phase: "generating",
@@ -272,18 +300,20 @@ export class CompactionCoordinator {
272
300
  summaryCalls: 0,
273
301
  provider: model.provider,
274
302
  model: model.id,
275
- inputBudgetMode: config.compaction.inputBudget ?? "summary",
303
+ inputBudgetMode: config.compaction.inputBudget ?? "context",
276
304
  maxConcurrentSegments,
305
+ operationInputTokens: 0,
277
306
  timings,
278
307
  };
279
308
 
280
309
  try {
281
- const { source, inputBudgetTokens, wholePlan, directPlan, segmentPlans } = await measure("preparationMs", () => {
310
+ const { source, inputBudgetTokens, requestInputLimitTokens, wholePlan, directPlan, segmentPlans } = await measure("preparationMs", () => {
282
311
  if (event.signal.aborted) throw new Error("Compaction summary generation aborted");
283
312
  this.dependencies.syncSessionIndex(ctx);
284
313
  const source = prepareCompactionSource(event);
285
314
  const inputBudgetTokens = this.inputBudgetTokens(model);
286
315
  if (inputBudgetTokens <= 0) throw new Error("Active model has no safe compaction input budget");
316
+ const requestInputLimitTokens = Math.min(inputBudgetTokens, maxRequestInputTokens);
287
317
  const wholePlan = this.buildSegmentPlan(source, event, model.provider);
288
318
  const update = (config.compaction.directUpdate ?? true) && source.previousSummary
289
319
  ? this.buildSegmentPlan({
@@ -292,18 +322,21 @@ export class CompactionCoordinator {
292
322
  segmentModifiedFiles: source.modifiedFiles,
293
323
  }, event, model.provider, source.previousSummary)
294
324
  : undefined;
295
- const directPlan = update && update.promptTokens <= inputBudgetTokens ? update : undefined;
325
+ const directPlan = update && update.promptTokens <= requestInputLimitTokens ? update : undefined;
296
326
  this.state = {
297
327
  ...this.state,
298
328
  path: directPlan ? "direct-update" : "hierarchical",
299
329
  sourceEntries: source.sourceEntryIds.length,
300
330
  inputBudgetTokens,
331
+ requestInputLimitTokens,
301
332
  sourcePromptTokens: wholePlan.promptTokens,
302
333
  ...(update ? { directPromptTokens: update.promptTokens } : {}),
303
334
  };
304
- const segmentPlans = directPlan ? [] : this.partitionSegmentPlans(source, wholePlan, event, model.provider, inputBudgetTokens);
335
+ const segmentPlans = directPlan
336
+ ? []
337
+ : this.partitionSegmentPlans(source, wholePlan, event, model.provider, requestInputLimitTokens);
305
338
  this.state.segmentCount = segmentPlans.length;
306
- return { source, inputBudgetTokens, wholePlan, directPlan, segmentPlans };
339
+ return { source, inputBudgetTokens, requestInputLimitTokens, wholePlan, directPlan, segmentPlans };
307
340
  });
308
341
 
309
342
  const usedIds = new Set(this.graphRecords.keys());
@@ -352,6 +385,9 @@ export class CompactionCoordinator {
352
385
  ctx,
353
386
  model,
354
387
  thinking: config.compaction.summary?.thinking,
388
+ promptTokens: plan.promptTokens,
389
+ requestInputLimitTokens,
390
+ operationInputBudget,
355
391
  }),
356
392
  ));
357
393
  const generatedNodes: EmbeddedSummaryNode[] = results.map((generated, index) => {
@@ -385,7 +421,8 @@ export class CompactionCoordinator {
385
421
  model: model.id,
386
422
  readFiles: source.readFiles,
387
423
  modifiedFiles: source.modifiedFiles,
388
- inputBudgetTokens,
424
+ requestInputLimitTokens,
425
+ operationInputBudget,
389
426
  nextId,
390
427
  createdNodes,
391
428
  usages,
@@ -698,6 +735,10 @@ export class CompactionCoordinator {
698
735
  validate: config.compaction.validate,
699
736
  preserveRecentVerbatim: config.compaction.preserveRecentVerbatim,
700
737
  segmentTargetTokens: config.compaction.segmentTargetTokens,
738
+ maxRequestInputTokens: config.compaction.maxRequestInputTokens
739
+ ?? DEFAULT_COMPACTION_MAX_REQUEST_INPUT_TOKENS,
740
+ maxOperationInputTokens: config.compaction.maxOperationInputTokens
741
+ ?? DEFAULT_COMPACTION_MAX_OPERATION_INPUT_TOKENS,
701
742
  ...this.state,
702
743
  ...(contextTokens !== undefined ? { contextTokens } : {}),
703
744
  ...(budget ? { softLimitTokens: this.providerSoftLimit(budget) } : {}),
@@ -800,7 +841,7 @@ export class CompactionCoordinator {
800
841
  ),
801
842
  };
802
843
  const maxOutputTokens = Math.max(1, Math.min(config.context.maxSummaryTokens, model.maxTokens ?? config.context.maxSummaryTokens));
803
- return compactionInputBudget(resolved.budget, maxOutputTokens, config.compaction.inputBudget ?? "summary");
844
+ return compactionInputBudget(resolved.budget, maxOutputTokens, config.compaction.inputBudget ?? "context");
804
845
  }
805
846
 
806
847
  private classify(text: string, provider: string): {
@@ -870,9 +911,13 @@ export class CompactionCoordinator {
870
911
  wholePlan: SegmentGenerationPlan,
871
912
  event: SessionBeforeCompactEvent,
872
913
  provider: string,
873
- inputBudgetTokens: number,
914
+ requestInputLimitTokens: number,
874
915
  ): SegmentGenerationPlan[] {
875
- if (wholePlan.promptTokens <= inputBudgetTokens) return [wholePlan];
916
+ const segmentTargetTokens = Math.min(
917
+ requestInputLimitTokens,
918
+ this.dependencies.config.compaction.segmentTargetTokens,
919
+ );
920
+ if (wholePlan.promptTokens <= segmentTargetTokens) return [wholePlan];
876
921
 
877
922
  const groups = buildCompactionAtomicGroups(source.messages);
878
923
  const plans: SegmentGenerationPlan[] = [];
@@ -886,7 +931,7 @@ export class CompactionCoordinator {
886
931
  event,
887
932
  provider,
888
933
  );
889
- if (candidatePlan.promptTokens <= inputBudgetTokens) {
934
+ if (candidatePlan.promptTokens <= segmentTargetTokens) {
890
935
  currentIndices = candidateIndices;
891
936
  currentPlan = candidatePlan;
892
937
  continue;
@@ -903,9 +948,9 @@ export class CompactionCoordinator {
903
948
  event,
904
949
  provider,
905
950
  );
906
- if (atomicPlan.promptTokens > inputBudgetTokens) {
951
+ if (atomicPlan.promptTokens > requestInputLimitTokens) {
907
952
  throw new Error(
908
- `Compaction source contains an indivisible atomic group above the model input budget (promptTokens=${atomicPlan.promptTokens}; inputBudgetTokens=${inputBudgetTokens})`,
953
+ `Compaction source contains an indivisible atomic group above the request input limit (promptTokens=${atomicPlan.promptTokens}; requestInputLimitTokens=${requestInputLimitTokens})`,
909
954
  );
910
955
  }
911
956
  currentIndices = [...group.messageIndices];
@@ -974,7 +1019,8 @@ export class CompactionCoordinator {
974
1019
  modelObject: Model<Api>;
975
1020
  readFiles: readonly string[];
976
1021
  modifiedFiles: readonly string[];
977
- inputBudgetTokens: number;
1022
+ requestInputLimitTokens: number;
1023
+ operationInputBudget: CompactionOperationInputBudget;
978
1024
  nextId: () => string;
979
1025
  createdNodes: EmbeddedSummaryNode[];
980
1026
  usages: GeneratedSummary["usage"][];
@@ -1003,13 +1049,13 @@ export class CompactionCoordinator {
1003
1049
  input.readFiles,
1004
1050
  input.modifiedFiles,
1005
1051
  );
1006
- if (candidatePlan.promptTokens <= input.inputBudgetTokens) {
1052
+ if (candidatePlan.promptTokens <= input.requestInputLimitTokens) {
1007
1053
  current = candidate;
1008
1054
  continue;
1009
1055
  }
1010
1056
  if (current.length === 1) {
1011
1057
  throw new Error(
1012
- `Compaction child summaries cannot be aggregated within the model input budget (inputBudgetTokens=${input.inputBudgetTokens})`,
1058
+ `Compaction child summaries cannot be aggregated within the request input limit (requestInputLimitTokens=${input.requestInputLimitTokens})`,
1013
1059
  );
1014
1060
  }
1015
1061
  batches.push(current);
@@ -1034,7 +1080,7 @@ export class CompactionCoordinator {
1034
1080
  input.readFiles,
1035
1081
  input.modifiedFiles,
1036
1082
  );
1037
- if (plan.promptTokens > input.inputBudgetTokens) {
1083
+ if (plan.promptTokens > input.requestInputLimitTokens) {
1038
1084
  throw new Error("Compaction aggregate prompt exceeded its preflight input budget");
1039
1085
  }
1040
1086
  this.state.aggregateCalls = aggregateCalls;
@@ -1048,6 +1094,9 @@ export class CompactionCoordinator {
1048
1094
  ctx: input.ctx,
1049
1095
  model: input.modelObject,
1050
1096
  thinking: this.dependencies.config.compaction.summary?.thinking,
1097
+ promptTokens: plan.promptTokens,
1098
+ requestInputLimitTokens: input.requestInputLimitTokens,
1099
+ operationInputBudget: input.operationInputBudget,
1051
1100
  });
1052
1101
  input.usages.push(generated.usage);
1053
1102
  const aggregateNode: EmbeddedSummaryNode = {
@@ -1088,7 +1137,15 @@ export class CompactionCoordinator {
1088
1137
  ctx: ExtensionContext;
1089
1138
  model: Model<Api>;
1090
1139
  thinking?: CompactionThinkingLevel;
1140
+ promptTokens: number;
1141
+ requestInputLimitTokens: number;
1142
+ operationInputBudget: CompactionOperationInputBudget;
1091
1143
  }): Promise<GeneratedSummary> {
1144
+ if (input.promptTokens > input.requestInputLimitTokens) {
1145
+ throw new Error(
1146
+ `Compaction ${input.stage} prompt exceeded the request input limit (promptTokens=${input.promptTokens}; requestInputLimitTokens=${input.requestInputLimitTokens})`,
1147
+ );
1148
+ }
1092
1149
  this.state.summaryCalls = (this.state.summaryCalls ?? 0) + 1;
1093
1150
  return generateValidatedSummary({
1094
1151
  ...input,
@@ -1096,6 +1153,16 @@ export class CompactionCoordinator {
1096
1153
  maxSummaryTokens: this.dependencies.config.context.maxSummaryTokens,
1097
1154
  transport: this.dependencies.config.compaction.transport,
1098
1155
  now: this.dependencies.now,
1156
+ onAttempt: (diagnostic) => {
1157
+ const nextInputTokens = input.operationInputBudget.usedTokens + input.promptTokens;
1158
+ if (nextInputTokens > input.operationInputBudget.limitTokens) {
1159
+ throw new Error(
1160
+ `Compaction operation input limit exceeded before ${diagnostic.stage} attempt ${diagnostic.attempt} (nextInputTokens=${nextInputTokens}; maxOperationInputTokens=${input.operationInputBudget.limitTokens})`,
1161
+ );
1162
+ }
1163
+ input.operationInputBudget.usedTokens = nextInputTokens;
1164
+ this.state.operationInputTokens = nextInputTokens;
1165
+ },
1099
1166
  onTransportRetry: (diagnostic) => {
1100
1167
  this.state.transportRetries = (this.state.transportRetries ?? 0) + 1;
1101
1168
  this.dependencies.logger.debug("compaction.transport_retry", { ...diagnostic });
@@ -1205,6 +1272,10 @@ export function defaultCompactionDiagnostics(config: Ds4ContextConfig): Compacti
1205
1272
  validate: config.compaction.validate,
1206
1273
  preserveRecentVerbatim: config.compaction.preserveRecentVerbatim,
1207
1274
  segmentTargetTokens: config.compaction.segmentTargetTokens,
1275
+ maxRequestInputTokens: config.compaction.maxRequestInputTokens
1276
+ ?? DEFAULT_COMPACTION_MAX_REQUEST_INPUT_TOKENS,
1277
+ maxOperationInputTokens: config.compaction.maxOperationInputTokens
1278
+ ?? DEFAULT_COMPACTION_MAX_OPERATION_INPUT_TOKENS,
1208
1279
  phase: "idle",
1209
1280
  proactiveEligible: false,
1210
1281
  };
@@ -13,16 +13,16 @@ import {
13
13
  type SummaryValidationResult,
14
14
  } from "ds4-context-core/compaction/summary-contract";
15
15
 
16
- export const DEFAULT_COMPACTION_TRANSPORT_MAX_ATTEMPTS = 3;
16
+ export const DEFAULT_COMPACTION_TRANSPORT_MAX_ATTEMPTS = 4;
17
17
  export const DEFAULT_COMPACTION_TRANSPORT_BASE_DELAY_MS = 2000;
18
18
  export const COMPACTION_TRANSPORT_MAX_DELAY_MS = 60_000;
19
19
 
20
20
  /**
21
- * Transport retry policy for compaction summary requests: three total attempts
22
- * (not three retries), 2000 ms base delay, exponential backoff, abort-aware.
21
+ * Transport retry policy for compaction summary requests: four total attempts
22
+ * (the initial call plus three retries), 2000 ms base delay, exponential backoff, abort-aware.
23
23
  */
24
24
  export interface CompactionTransportPolicy {
25
- /** Total attempts for transport-classified failures. Default: 3. */
25
+ /** Total attempts for transport-classified failures. Default: 4. */
26
26
  maxAttempts?: number;
27
27
  /** Base backoff delay in ms, doubled per attempt. Default: 2000. */
28
28
  baseDelayMs?: number;
@@ -54,6 +54,12 @@ export interface CompactionTransportRetryDiagnostic {
54
54
  delayMs: number;
55
55
  }
56
56
 
57
+ export interface CompactionAttemptDiagnostic {
58
+ stage: "segment" | "aggregate" | "update";
59
+ attempt: number;
60
+ maxAttempts: number;
61
+ }
62
+
57
63
  export interface GenerateValidatedSummaryInput {
58
64
  stage: "segment" | "aggregate" | "update";
59
65
  prompt: string;
@@ -68,9 +74,11 @@ export interface GenerateValidatedSummaryInput {
68
74
  model?: Model<Api>;
69
75
  /** Reasoning level for the summary request; `off` (default) keeps the pre-existing request shape. */
70
76
  thinking?: CompactionThinkingLevel;
71
- /** Transport-only retry policy; three total attempts by default. */
77
+ /** Transport-only retry policy; four total attempts by default. */
72
78
  transport?: CompactionTransportPolicy;
73
79
  now: () => number;
80
+ /** Called synchronously before every provider attempt, including retries. */
81
+ onAttempt?: (diagnostic: CompactionAttemptDiagnostic) => void;
74
82
  onTransportRetry?: (diagnostic: CompactionTransportRetryDiagnostic) => void;
75
83
  }
76
84
 
@@ -201,6 +209,7 @@ export async function generateValidatedSummary(
201
209
  for (;;) {
202
210
  if (input.event.signal.aborted) throw abortedError();
203
211
  attempt++;
212
+ input.onAttempt?.({ stage: input.stage, attempt, maxAttempts });
204
213
  try {
205
214
  response = await input.ctx.modelRegistry.complete(
206
215
  model,
@@ -1,4 +1,4 @@
1
- export const EXTENSION_VERSION = "0.3.7";
1
+ export const EXTENSION_VERSION = "0.3.9";
2
2
  export const SUPPORTED_PI_VERSION = "0.84.3";
3
3
  export const OBSERVER_PLANNER_VERSION = "observer-model-aware-v1";
4
4
  export const PLANNER_VERSION = "managed-learned-ranking-v1";