ds4-context-engine 0.4.2 → 0.4.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -16,7 +16,7 @@ bounded active context with provenance
16
16
  Pi provider
17
17
  ```
18
18
 
19
- > **Project status:** The coordinated `0.4.0` release adds `storage.scope: "agent" | "project"`, now defaulting to per-project SQLite projections with shared token calibration in the agent database (opt out with `storage.scope: "agent"`). The opt-in BPE estimation and bounded auto-tuning from `0.3.10` keep `chars-v1` and disabled auto-tuning as their defaults. The bounded compaction controls from `0.3.9` remain in place; canonical history, SQLite schema 16 and runtime contracts are unchanged. Pi remains pinned to `0.84.3`. Patch `0.4.1` states the contiguous-span rule for compaction summaries in the summarizer prompt while leaving validation strictness, the eight-bullet repair bound and every provider-facing default unchanged. Patch `0.4.2` adds class-only diagnostics for rejected exact-value spans: the existing `custom_fallback` warning carries `unsupportedSpanClasses` as class names and counters, never span text, with validation, repair bounds and provider-facing defaults unchanged. See the [0.4.2 release record](docs/releases/0.4.2.md), the [0.4.1 release record](docs/releases/0.4.1.md) and the [0.4.0 release record](docs/releases/0.4.0.md), [ADR 064](docs/ADR/064-per-project-databases-with-shared-calibration.md) and [model-awareness validation](docs/MODEL_AWARENESS.md).
19
+ > **Project status:** The coordinated `0.4.0` release adds `storage.scope: "agent" | "project"`, now defaulting to per-project SQLite projections with shared token calibration in the agent database (opt out with `storage.scope: "agent"`). The opt-in BPE estimation and bounded auto-tuning from `0.3.10` keep `chars-v1` and disabled auto-tuning as their defaults. The bounded compaction controls from `0.3.9` remain in place; canonical history, SQLite schema 16 and runtime contracts are unchanged. Pi remains pinned to `0.84.3`. Patch `0.4.1` states the contiguous-span rule for compaction summaries in the summarizer prompt while leaving validation strictness, the eight-bullet repair bound and every provider-facing default unchanged. Patch `0.4.2` adds class-only diagnostics for rejected exact-value spans: the existing `custom_fallback` warning carries `unsupportedSpanClasses` as class names and counters, never span text, with validation, repair bounds and provider-facing defaults unchanged. Patch `0.4.3` makes those diagnostics interpretable: every distinct span receives the cheap transformed-form lookups before the bounded near-miss analysis starts, each near-miss candidate may spend only an equal share of what remains, and spans that were not analysed report `not-classified-length`, `not-classified-partial` or `not-classified-budget` instead of `no-near-miss`. See the [0.4.3 release record](docs/releases/0.4.3.md), the [0.4.2 release record](docs/releases/0.4.2.md), the [0.4.1 release record](docs/releases/0.4.1.md) and the [0.4.0 release record](docs/releases/0.4.0.md), [ADR 064](docs/ADR/064-per-project-databases-with-shared-calibration.md) and [model-awareness validation](docs/MODEL_AWARENESS.md).
20
20
 
21
21
  **Current compaction defaults:** `compaction.directUpdate=true`, `compaction.inputBudget="context"`, `compaction.segmentTargetTokens=30000`, `compaction.maxRequestInputTokens=64000`, `compaction.maxOperationInputTokens=2000000`, `compaction.maxConcurrentSegments=2`. Every DS4 provider attempt is bounded by the effective request limit, and the operation limit includes retries; `inputBudget="summary"` remains an explicit throughput-oriented opt-in. Existing compaction/master switches still apply. See [latency controls and compatibility](docs/COMPACTION.md#latency-controls). No real-provider speedup is claimed from mock tests. The five optional editing/reading/artifact/job features introduced in `0.3.4` remain default-off.
22
22
 
@@ -515,6 +515,7 @@ scripts package and release-readiness checks
515
515
  - [Roadmap 0.2.0](docs/ROADMAP_0.2.0.md)
516
516
  - [Release process](docs/RELEASING.md)
517
517
  - [0.2.0 release readiness](docs/RELEASE_READINESS_0.2.0.md)
518
+ - [0.4.3 release notes](docs/releases/0.4.3.md)
518
519
  - [0.4.2 release notes](docs/releases/0.4.2.md)
519
520
  - [0.4.1 release notes](docs/releases/0.4.1.md)
520
521
  - [0.4.0 release notes](docs/releases/0.4.0.md)
@@ -36,7 +36,9 @@ Every section must occur once, in order, and contain content or `- None`. DS4 re
36
36
 
37
37
  An unrepaired exact-value failure reports only the stage, issue code, categorical repair status, unsupported-span count, and affected-bullet count. Repair statuses distinguish an unsupported location, more than eight bullets, removal above 25%, and an unexpected invalid second validation. The disputed text is intentionally absent from logs, UI notifications, and diagnostics because it may contain sensitive source material.
38
38
 
39
- The same failure carries a class-only span report on the fallback warning (`unsupportedSpanClasses`): rejected-span count, affected bullets, length buckets, character shapes (spaces, backslashes, escape sequences, separators, quotes, typographic characters, JSON punctuation), and how each distinct span relates to the evidence — escaped or unescaped rendering, collapsed whitespace, case or typographic variant, one-character deletion, two separately present values joined into one span, no near-miss, or not classified because the bounded probe budget was exhausted. The report contains class names and counters only, never span text, and the classifier is observation-only: it cannot change validation, repair bounds, or the fail-closed decision. It exists to separate a summarizer that invents values from one whose rendering or composition rules differ from the validation domain, which require opposite fixes.
39
+ The same failure carries a class-only span report on the fallback warning (`unsupportedSpanClasses`): rejected-span count, affected bullets, length buckets, character shapes (spaces, backslashes, escape sequences, separators, quotes, typographic characters, JSON punctuation), and how each distinct span relates to the evidence — escaped or unescaped rendering, collapsed whitespace, case or typographic variant, one-character deletion, two separately present values joined into one span, or no near-miss (every applicable lookup ran and found nothing). The report contains class names and counters only, never span text, and the classifier is observation-only: it cannot change validation, repair bounds, or the fail-closed decision. It exists to separate a summarizer that invents values from one whose rendering or composition rules differ from the validation domain, which require opposite fixes.
40
+
41
+ Three relations mark spans that were *not* analysed, and they must never be read as invention. `not-classified-length` means the span exceeds the near-miss length limit (96 characters); `not-classified-partial` means its share of the budget ran out; `not-classified-budget` means the shared budget was already exhausted. To keep the histogram interpretable, transformed-form lookups cover every distinct span before any near-miss analysis starts, and each near-miss candidate may spend only an equal share of what remains (at least 32 lookups, never more than the budget left). The report also carries `spansClassifiedCheap` (span occurrences attributed by a transformed-form relation), `corpusSources`, `probeBudget` and `probesUsed`, so a run can be read without guessing how much of the budget was consumed. `classificationComplete` is false when the shared budget or a per-span share stopped an analysis; a `not-classified-length` span does not clear it, because that limit is deterministic rather than a resource shortfall.
40
42
 
41
43
  ## Provenance and recovery
42
44
 
@@ -110,4 +110,14 @@ Vitest files and 623 tests), deterministic `npm run quality:compare` (candidate
110
110
  `0.9875` against baseline `0.808156`), `npm run schema:context-persistence`
111
111
  (`passed: true`), and `npm run pack:check` in a clean consumer for all three
112
112
  packages at 0.4.2 (247 core files, 7 reference-adapter files, 100 extension
113
- files). No provider calls are involved in this release procedure.
113
+ files). The coordinated publication then proceeded in dependency order
114
+ (`ds4-context-core`, then `ds4-context-reference-adapter`, then
115
+ `ds4-context-engine`), each command reporting the exact 0.4.2 artifact,
116
+ followed by `npm run registry:check -- 0.4.2`. Registry propagation lagged the
117
+ publishes: the first `npm run registry:check -- 0.4.2` failed with
118
+ `ETARGET ... No matching version found for ds4-context-core@0.4.2`, and six
119
+ retries over roughly four minutes reported the same error while the registry
120
+ metadata already listed 0.4.2. Once propagation completed, the repeated
121
+ `npm run registry:check -- 0.4.2` reported "Verified all DS4 registry packages
122
+ at exact version 0.4.2", and the `latest` dist-tag resolved to 0.4.2 for all
123
+ three packages. No provider calls are involved in this release procedure.
@@ -0,0 +1,159 @@
1
+ # Release 0.4.3 — Fair span coverage and honest "unanalysed" classes
2
+
3
+ **Coordinated packages:** `ds4-context-core`, `ds4-context-reference-adapter`, and `ds4-context-engine` 0.4.3.
4
+ **Implementation commit:** `47100ad`.
5
+
6
+ ## Summary
7
+
8
+ Diagnostics patch. The class-only report that `0.4.2` attached to a failed
9
+ compaction could not answer the question it exists for: in its first production
10
+ run, 13 of 23 rejected spans came back as `not-classified-budget`, and the two
11
+ `no-near-miss` entries could not be distinguished from spans the classifier
12
+ never analysed. The classifier now covers every distinct span with the cheap
13
+ transformed-form lookups before any near-miss analysis starts, gives each
14
+ near-miss candidate an equal share of what remains, and reports spans it did not
15
+ analyse as unanalysed instead of guessing.
16
+
17
+ Validation strictness, the eight-bullet and 25% repair bounds, the fail-closed
18
+ fallback to Pi default compaction, the provider-facing defaults, the SQLite
19
+ schema and every configuration key are unchanged. The report still contains
20
+ class names and counters only, so it stays safe to log.
21
+
22
+ ## Changes
23
+
24
+ - `packages/core/src/compaction/summary-contract.ts`:
25
+ - Transformed-form lookups (`escaped-form-present`, `unescaped-form-present`,
26
+ `whitespace-collapsed-present`, `case-variant-present`,
27
+ `typographic-variant-present`) run for every distinct span before the first
28
+ near-miss lookup, so span order can no longer decide which spans are
29
+ classified.
30
+ - Each remaining near-miss candidate may spend an equal share of the lookups
31
+ left, at least 32 and never more than the budget remaining, so one long span
32
+ cannot consume the budget of the others.
33
+ - New relations distinguish "not analysed" from "absent": `not-classified-length`
34
+ above the 96-character near-miss limit, `not-classified-partial` when the
35
+ span's share ran out, and the existing `not-classified-budget` when the
36
+ shared budget was already exhausted. `no-near-miss` now means every
37
+ applicable lookup ran and found nothing.
38
+ - The report adds `spansClassifiedCheap` (span occurrences attributed by a
39
+ transformed-form relation), `corpusSources`, `probeBudget` and `probesUsed`.
40
+ `classificationComplete` is false when the shared budget or a per-span share
41
+ stopped an analysis; a `not-classified-length` span does not clear it,
42
+ because that limit is deterministic rather than a resource shortfall.
43
+ - `UnsupportedSpanClassOptions.probeBudget` can override the 4000-lookup
44
+ budget; values below one or non-finite values are ignored. No configuration
45
+ key was added.
46
+ - `tests/unit/summary-contract.test.ts`: the existing bound test now asserts
47
+ `not-classified-partial` and `probesUsed <= probeBudget`, plus new coverage for
48
+ the phase split, the length limit, budget exhaustion and unusable budget
49
+ values.
50
+ - `docs/COMPACTION.md`: the diagnostics section documents the three
51
+ not-classified classes, the phase split and the new counters.
52
+
53
+ No change was needed in `src/pi-adapter/summary-generator.ts` or
54
+ `src/pi-adapter/compaction-coordinator.ts`: the existing error and the existing
55
+ `custom_fallback` warning already carry the report, so the new fields appear in
56
+ `unsupportedSpanClasses` automatically.
57
+
58
+ ## Why
59
+
60
+ The first production run of the `0.4.2` report (a manual compaction, 23 rejected
61
+ spans across 16 bullets) came back as:
62
+
63
+ | Class | Count |
64
+ | --- | --- |
65
+ | `escaped-form-present` | 8 |
66
+ | `no-near-miss` | 2 |
67
+ | `not-classified-budget` | 13 |
68
+
69
+ Length buckets were dominated by long spans (10 in `len-33-64`, 7 in
70
+ `len-65-120`, 2 in `len-121-plus`) and shapes by prose (20 spans containing
71
+ spaces, 10 quotes, 10 colons), so the entries looked like quoted excerpts rather
72
+ than identifiers. Only the eight escaped-form spans were actionable: the value
73
+ was present, only its rendering differed. Two defects in the instrument blocked
74
+ the rest:
75
+
76
+ - **Coverage.** Every lookup costs one scan of each corpus source
77
+ (`sourceText`, `readFiles`, `modifiedFiles`) and the budget was spent in span
78
+ order, while `single-deletion-present` costs up to one lookup per character
79
+ and `composed-two-present-parts` up to two more per length. A handful of long
80
+ spans, or their complete analysis, could therefore exhaust the budget before
81
+ the remaining spans received a single cheap lookup.
82
+ - **Ambiguity.** `no-near-miss` was also returned for spans above
83
+ `MAX_NEAR_MISS_SPAN_LENGTH`, where the near-miss analysis is never attempted.
84
+
85
+ Neither defect weakened validation: the repair bound and the fail-closed
86
+ fallback were respected in every one of these runs. The problem was that counts
87
+ alone could not separate a summarizer that invents values from one whose
88
+ rendering or composition rules differ from the validation domain, and those two
89
+ findings call for opposite fixes. This release makes the next report
90
+ interpretable before any change to validation, repair, or the summarizer prompt
91
+ is considered.
92
+
93
+ ## Privacy and safety
94
+
95
+ - The classifier remains observation-only and cannot alter validation, repair
96
+ bounds, or the fail-closed decision.
97
+ - The report still contains class names and counters only. The new fields are
98
+ integers (occurrence, source and lookup counts); no span text, no evidence
99
+ text, and no hashes of either are added to logs, warnings, or errors.
100
+ - No provider call, network access, or configuration change is involved in this
101
+ release; the classifier runs only after a validation failure has already been
102
+ decided.
103
+
104
+ ## Compatibility and persistence
105
+
106
+ - No migrations, no schema change (still 16), no new configuration key, and no
107
+ change to any default.
108
+ - The report shape is additive: existing consumers keep working, but a reader of
109
+ `relations` must tolerate the new keys `not-classified-length` and
110
+ `not-classified-partial`, and can no longer assume that `no-near-miss` covers
111
+ every unanalysed span.
112
+ - `classificationComplete: false` is now also possible with the shared budget
113
+ intact, when a per-span share stops an analysis.
114
+
115
+ ## Measured scope and limitations
116
+
117
+ Measured locally on the development host with the unit fixtures, no provider
118
+ calls:
119
+
120
+ - A 96-character span with no transformed-form match costs 185 lookups when
121
+ analysed completely (96 single-deletion plus 89 prefix probes), confirming that
122
+ two or three such spans could exhaust the whole budget under the `0.4.2` order.
123
+ - With 80 distinct spans of 91-92 characters and the default budget, the cheap
124
+ phase spends 80 lookups and each span then receives a 49-lookup share:
125
+ `probesUsed` equals `probeBudget` exactly and every span is reported as
126
+ `not-classified-partial`.
127
+ - The knife edge is observable: in a combined fixture, a budget of 190 reports
128
+ `not-classified-partial` while 191 reports `no-near-miss`, because the share is
129
+ exhausted on the lookup that completes the analysis.
130
+
131
+ Limitations that remain:
132
+
133
+ - Distinct spans are deduplicated, so `spansClassifiedCheap` counts occurrences
134
+ while the budget is spent per distinct value.
135
+ - Spans above 96 characters are still never near-miss analysed; they now say so
136
+ instead of claiming absence.
137
+ - The default budget is still 4000 lookups per failure, unchanged from `0.4.2`.
138
+ - A `not-classified-*` class still cannot say whether the span was invented;
139
+ it says only that the classifier did not look.
140
+
141
+ ## Validation
142
+
143
+ - `npm run build:core`, `npm run build:adapters`, and `npx tsc --noEmit` clean.
144
+ - `npm run check`: 99 test files, 627 tests, all passing.
145
+ - `npm run quality:compare`, `npm run schema:context-persistence` (passed) and
146
+ `npm run pack:check`: core 0.4.3 (247 files), reference adapter 0.4.3
147
+ (7 files) and engine 0.4.3 (101 files) verified in a clean consumer.
148
+ - Classifier behavior was additionally checked against the `0.4.2` semantics
149
+ with a synthetic reconstruction of the production histogram shape (long spans
150
+ first, cheap matches after), without touching session data.
151
+ - Not verified: the improved histogram on a real failed compaction. That
152
+ requires a natural failure and is expected to appear on the next
153
+ `compaction.custom_fallback` warning.
154
+
155
+ ## Publication
156
+
157
+ Published manually in dependency order (core, reference adapter, engine) with
158
+ `npm publish`, after `npm run pack:check` on the same commit. Registry
159
+ verification is recorded in the follow-up documentation commit.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ds4-context-engine",
3
- "version": "0.4.2",
3
+ "version": "0.4.3",
4
4
  "description": "Non-destructive, provider-independent context management for Pi.",
5
5
  "type": "module",
6
6
  "license": "MIT",
@@ -63,7 +63,7 @@
63
63
  ]
64
64
  },
65
65
  "dependencies": {
66
- "ds4-context-core": "0.4.2",
66
+ "ds4-context-core": "0.4.3",
67
67
  "js-tiktoken": "1.0.21"
68
68
  },
69
69
  "peerDependencies": {
@@ -1,4 +1,4 @@
1
- export const EXTENSION_VERSION = "0.4.2";
1
+ export const EXTENSION_VERSION = "0.4.3";
2
2
  export const SUPPORTED_PI_VERSION = "0.84.3";
3
3
  export const OBSERVER_PLANNER_VERSION = "observer-model-aware-v1";
4
4
  export const PLANNER_VERSION = "managed-learned-ranking-v1";