ds4-context-engine 0.4.1 → 0.4.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +4 -2
- package/docs/COMPACTION.md +5 -1
- package/docs/releases/0.4.1.md +10 -4
- package/docs/releases/0.4.2.md +123 -0
- package/docs/releases/0.4.3.md +159 -0
- package/package.json +2 -2
- package/src/pi-adapter/compaction-coordinator.ts +9 -1
- package/src/pi-adapter/summary-generator.ts +33 -2
- package/src/pi-adapter/version.ts +1 -1
package/README.md
CHANGED
|
@@ -16,7 +16,7 @@ bounded active context with provenance
|
|
|
16
16
|
Pi provider
|
|
17
17
|
```
|
|
18
18
|
|
|
19
|
-
> **Project status:** The coordinated `0.4.0` release adds `storage.scope: "agent" | "project"`, now defaulting to per-project SQLite projections with shared token calibration in the agent database (opt out with `storage.scope: "agent"`). The opt-in BPE estimation and bounded auto-tuning from `0.3.10` keep `chars-v1` and disabled auto-tuning as their defaults. The bounded compaction controls from `0.3.9` remain in place; canonical history, SQLite schema 16 and runtime contracts are unchanged. Pi remains pinned to `0.84.3`. Patch `0.4.1` states the contiguous-span rule for compaction summaries in the summarizer prompt while leaving validation strictness, the eight-bullet repair bound and every provider-facing default unchanged. See the [0.4.1 release record](docs/releases/0.4.1.md) and the [0.4.0 release record](docs/releases/0.4.0.md), [ADR 064](docs/ADR/064-per-project-databases-with-shared-calibration.md) and [model-awareness validation](docs/MODEL_AWARENESS.md).
|
|
19
|
+
> **Project status:** The coordinated `0.4.0` release adds `storage.scope: "agent" | "project"`, now defaulting to per-project SQLite projections with shared token calibration in the agent database (opt out with `storage.scope: "agent"`). The opt-in BPE estimation and bounded auto-tuning from `0.3.10` keep `chars-v1` and disabled auto-tuning as their defaults. The bounded compaction controls from `0.3.9` remain in place; canonical history, SQLite schema 16 and runtime contracts are unchanged. Pi remains pinned to `0.84.3`. Patch `0.4.1` states the contiguous-span rule for compaction summaries in the summarizer prompt while leaving validation strictness, the eight-bullet repair bound and every provider-facing default unchanged. Patch `0.4.2` adds class-only diagnostics for rejected exact-value spans: the existing `custom_fallback` warning carries `unsupportedSpanClasses` as class names and counters, never span text, with validation, repair bounds and provider-facing defaults unchanged. Patch `0.4.3` makes those diagnostics interpretable: every distinct span receives the cheap transformed-form lookups before the bounded near-miss analysis starts, each near-miss candidate may spend only an equal share of what remains, and spans that were not analysed report `not-classified-length`, `not-classified-partial` or `not-classified-budget` instead of `no-near-miss`. See the [0.4.3 release record](docs/releases/0.4.3.md), the [0.4.2 release record](docs/releases/0.4.2.md), the [0.4.1 release record](docs/releases/0.4.1.md) and the [0.4.0 release record](docs/releases/0.4.0.md), [ADR 064](docs/ADR/064-per-project-databases-with-shared-calibration.md) and [model-awareness validation](docs/MODEL_AWARENESS.md).
|
|
20
20
|
|
|
21
21
|
**Current compaction defaults:** `compaction.directUpdate=true`, `compaction.inputBudget="context"`, `compaction.segmentTargetTokens=30000`, `compaction.maxRequestInputTokens=64000`, `compaction.maxOperationInputTokens=2000000`, `compaction.maxConcurrentSegments=2`. Every DS4 provider attempt is bounded by the effective request limit, and the operation limit includes retries; `inputBudget="summary"` remains an explicit throughput-oriented opt-in. Existing compaction/master switches still apply. See [latency controls and compatibility](docs/COMPACTION.md#latency-controls). No real-provider speedup is claimed from mock tests. The five optional editing/reading/artifact/job features introduced in `0.3.4` remain default-off.
|
|
22
22
|
|
|
@@ -515,6 +515,8 @@ scripts package and release-readiness checks
|
|
|
515
515
|
- [Roadmap 0.2.0](docs/ROADMAP_0.2.0.md)
|
|
516
516
|
- [Release process](docs/RELEASING.md)
|
|
517
517
|
- [0.2.0 release readiness](docs/RELEASE_READINESS_0.2.0.md)
|
|
518
|
+
- [0.4.3 release notes](docs/releases/0.4.3.md)
|
|
519
|
+
- [0.4.2 release notes](docs/releases/0.4.2.md)
|
|
518
520
|
- [0.4.1 release notes](docs/releases/0.4.1.md)
|
|
519
521
|
- [0.4.0 release notes](docs/releases/0.4.0.md)
|
|
520
522
|
- [0.3.10 release notes](docs/releases/0.3.10.md)
|
|
@@ -545,7 +547,7 @@ scripts package and release-readiness checks
|
|
|
545
547
|
|
|
546
548
|
The original M0–M13 roadmap is complete. `ds4-context-core` contains the compiled runtime-neutral implementation. M14 context-quality metrics, M15 rich symbol indexing, M16 hybrid semantic retrieval, M17 cross-session project memory, M18 learned-ranking shadow evaluation, M19's runtime adapter/conformance kit, and M20 opt-in local KV eligibility/replay are implemented on `main`. Learned active ranking remains promotion-gated, Pi reports local KV as unsupported, and static ranking/native completion stay authoritative on every failure.
|
|
547
549
|
|
|
548
|
-
The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.4.0 introduces `storage.scope` (`"agent" | "project"`, default `"project"`): one rebuildable SQLite projection per trusted canonical project root, with token calibration shared in the agent database; `storage.scope: "agent"` restores the previous single-file layout. Version 0.4.1 states the contiguous-span rule for compaction summaries in the summarizer prompt without changing deterministic validation or the repair bounds. Version 0.3.10 adds opt-in BPE estimation and bounded model-budget auto-tuning while retaining the `chars-v1` default. Version 0.3.9 extends the bounded compaction updates, summary input headroom, concurrent segments and phase timings introduced in 0.3.5 with per-request and cumulative operation input limits; 0.3.8 adds indexed FTS key deletion without changing search results. The opt-in [anchored editing](docs/ANCHORED_EDITING.md) and [portable agent tools](docs/PORTABLE_AGENT_TOOLS.md) from 0.3.4 remain default-off, without backend rewind, forced sampling or operational KV integration. Confirmation, provenance, Pi fallback and canonical/configuration/SQLite/runtime contracts remain unchanged. The [0.2 readiness record](docs/RELEASE_READINESS_0.2.0.md) remains the compatibility baseline; the lexical planner stays available as the deterministic fallback.
|
|
550
|
+
The [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) is complete. The stable 0.3 line carries forward the [context persistence tool](docs/CONTEXT_PERSISTENCE_TOOL.md), privacy-safe [compaction](docs/COMPACTION.md), bounded persisted manifests, cooperative client leases and recoverable offline maintenance. Version 0.4.0 introduces `storage.scope` (`"agent" | "project"`, default `"project"`): one rebuildable SQLite projection per trusted canonical project root, with token calibration shared in the agent database; `storage.scope: "agent"` restores the previous single-file layout. Version 0.4.1 states the contiguous-span rule for compaction summaries in the summarizer prompt without changing deterministic validation or the repair bounds. Version 0.4.2 adds class-only diagnostics to the same fail-closed path: the `custom_fallback` warning reports `unsupportedSpanClasses` as class names and counters only, never span text, without changing validation strictness, repair bounds or defaults. Version 0.3.10 adds opt-in BPE estimation and bounded model-budget auto-tuning while retaining the `chars-v1` default. Version 0.3.9 extends the bounded compaction updates, summary input headroom, concurrent segments and phase timings introduced in 0.3.5 with per-request and cumulative operation input limits; 0.3.8 adds indexed FTS key deletion without changing search results. The opt-in [anchored editing](docs/ANCHORED_EDITING.md) and [portable agent tools](docs/PORTABLE_AGENT_TOOLS.md) from 0.3.4 remain default-off, without backend rewind, forced sampling or operational KV integration. Confirmation, provenance, Pi fallback and canonical/configuration/SQLite/runtime contracts remain unchanged. The [0.2 readiness record](docs/RELEASE_READINESS_0.2.0.md) remains the compatibility baseline; the lexical planner stays available as the deterministic fallback.
|
|
549
551
|
|
|
550
552
|
## Contributing
|
|
551
553
|
|
package/docs/COMPACTION.md
CHANGED
|
@@ -36,6 +36,10 @@ Every section must occur once, in order, and contain content or `- None`. DS4 re
|
|
|
36
36
|
|
|
37
37
|
An unrepaired exact-value failure reports only the stage, issue code, categorical repair status, unsupported-span count, and affected-bullet count. Repair statuses distinguish an unsupported location, more than eight bullets, removal above 25%, and an unexpected invalid second validation. The disputed text is intentionally absent from logs, UI notifications, and diagnostics because it may contain sensitive source material.
|
|
38
38
|
|
|
39
|
+
The same failure carries a class-only span report on the fallback warning (`unsupportedSpanClasses`): rejected-span count, affected bullets, length buckets, character shapes (spaces, backslashes, escape sequences, separators, quotes, typographic characters, JSON punctuation), and how each distinct span relates to the evidence — escaped or unescaped rendering, collapsed whitespace, case or typographic variant, one-character deletion, two separately present values joined into one span, or no near-miss (every applicable lookup ran and found nothing). The report contains class names and counters only, never span text, and the classifier is observation-only: it cannot change validation, repair bounds, or the fail-closed decision. It exists to separate a summarizer that invents values from one whose rendering or composition rules differ from the validation domain, which require opposite fixes.
|
|
40
|
+
|
|
41
|
+
Three relations mark spans that were *not* analysed, and they must never be read as invention. `not-classified-length` means the span exceeds the near-miss length limit (96 characters); `not-classified-partial` means its share of the budget ran out; `not-classified-budget` means the shared budget was already exhausted. To keep the histogram interpretable, transformed-form lookups cover every distinct span before any near-miss analysis starts, and each near-miss candidate may spend only an equal share of what remains (at least 32 lookups, never more than the budget left). The report also carries `spansClassifiedCheap` (span occurrences attributed by a transformed-form relation), `corpusSources`, `probeBudget` and `probesUsed`, so a run can be read without guessing how much of the budget was consumed. `classificationComplete` is false when the shared budget or a per-span share stopped an analysis; a `not-classified-length` span does not clear it, because that limit is deterministic rather than a resource shortfall.
|
|
42
|
+
|
|
39
43
|
## Provenance and recovery
|
|
40
44
|
|
|
41
45
|
`CompactionEntry.details` contains cumulative `readFiles` and `modifiedFiles` plus:
|
|
@@ -146,4 +150,4 @@ It requests compaction at most once per session leaf. Pi's native threshold and
|
|
|
146
150
|
/context summaries
|
|
147
151
|
```
|
|
148
152
|
|
|
149
|
-
The preview reports thresholds and eligibility; Pi remains authoritative for the exact cut point. Runtime diagnostics additionally report the effective provider/model, chosen path (`direct-update` or `hierarchical`), budget mode, calibrated input budget, estimated new-source and full-update prompt sizes, logical summary calls (excluding retries), generated segment count, aggregate-call count, concurrency cap, and completed transport-retry count without source content. Monotonic wall timings cover preparation, generation (direct or parallel segments, including validation and retries), aggregation, graph preparation/persistence, and total DS4 hook time. Parallel generation time is elapsed wall time, not a sum of overlapping requests. Total time excludes Pi's subsequent canonical append or native fallback. Timings survive the in-process commit/fallback notification, but are not persisted in JSONL or reconstructed as fake durations after restart. Metadata-only `compaction.timings` is emitted at `debug` for successful and failed attempts. Each retry emits only stage, failed/next attempt, maximum attempts, and delay at `debug`. Routine `summary_graph_prepared` and `summary_graph_committed` lifecycle events are also emitted only at `debug`; fallback, failure, persistence, and reconciliation problems remain actionable warnings. Oversized-group and bounded-operation errors expose only numeric budgets/counts, never rejected source text. Provider failures are reduced to metadata-only categories such as `input-limit`, `usage-limit`, `rate-limit`, `authentication`, or `transport`; raw provider error details are not logged or shown. The proactive-threshold TUI notification remains a user-visible `info` notice because it explains why an automatic compaction started.
|
|
153
|
+
The preview reports thresholds and eligibility; Pi remains authoritative for the exact cut point. Runtime diagnostics additionally report the effective provider/model, chosen path (`direct-update` or `hierarchical`), budget mode, calibrated input budget, estimated new-source and full-update prompt sizes, logical summary calls (excluding retries), generated segment count, aggregate-call count, concurrency cap, and completed transport-retry count without source content. Monotonic wall timings cover preparation, generation (direct or parallel segments, including validation and retries), aggregation, graph preparation/persistence, and total DS4 hook time. Parallel generation time is elapsed wall time, not a sum of overlapping requests. Total time excludes Pi's subsequent canonical append or native fallback. Timings survive the in-process commit/fallback notification, but are not persisted in JSONL or reconstructed as fake durations after restart. Metadata-only `compaction.timings` is emitted at `debug` for successful and failed attempts. Failed validation additionally embeds the class-only `unsupportedSpanClasses` report in the `compaction.custom_fallback` warning, so a failing compaction can be diagnosed from logs alone without reproducing the session or reading its content. Each retry emits only stage, failed/next attempt, maximum attempts, and delay at `debug`. Routine `summary_graph_prepared` and `summary_graph_committed` lifecycle events are also emitted only at `debug`; fallback, failure, persistence, and reconciliation problems remain actionable warnings. Oversized-group and bounded-operation errors expose only numeric budgets/counts, never rejected source text. Provider failures are reduced to metadata-only categories such as `input-limit`, `usage-limit`, `rate-limit`, `authentication`, or `transport`; raw provider error details are not logged or shown. The proactive-threshold TUI notification remains a user-visible `info` notice because it explains why an automatic compaction started.
|
package/docs/releases/0.4.1.md
CHANGED
|
@@ -80,8 +80,14 @@ files, 615 tests, TypeScript builds and root typecheck), deterministic
|
|
|
80
80
|
in a clean consumer for all three packages at 0.4.1 (247 core files, 7
|
|
81
81
|
reference-adapter files, 99 extension files). Tarball review found no session,
|
|
82
82
|
`.pi`, database, JSONL or credential files, and `git diff --check` was clean.
|
|
83
|
-
The coordinated publication then
|
|
83
|
+
The coordinated publication then proceeded in dependency order
|
|
84
84
|
(`ds4-context-core`, then `ds4-context-reference-adapter`, then
|
|
85
|
-
`ds4-context-engine`)
|
|
86
|
-
|
|
87
|
-
|
|
85
|
+
`ds4-context-engine`), each command reporting the exact 0.4.1 artifact.
|
|
86
|
+
Registry propagation lagged the publishes: the first
|
|
87
|
+
`npm run registry:check -- 0.4.1` failed with
|
|
88
|
+
`ETARGET ... No matching version found for ds4-context-core@0.4.1`, and
|
|
89
|
+
`ds4-context-engine` still resolved to 0.4.0 after core, reference adapter and
|
|
90
|
+
core again were visible. Once propagation completed, the repeated
|
|
91
|
+
`npm run registry:check -- 0.4.1` reported "Verified all DS4 registry packages
|
|
92
|
+
at exact version 0.4.1", and the `latest` dist-tag resolved to 0.4.1 for all
|
|
93
|
+
three packages. No provider calls are involved in this release procedure.
|
|
@@ -0,0 +1,123 @@
|
|
|
1
|
+
# Release 0.4.2 — Class-only diagnostics for rejected exact-value spans
|
|
2
|
+
|
|
3
|
+
**Coordinated packages:** `ds4-context-core`, `ds4-context-reference-adapter`, and `ds4-context-engine` 0.4.2.
|
|
4
|
+
**Implementation commit:** `6ecd2af`.
|
|
5
|
+
|
|
6
|
+
## Summary
|
|
7
|
+
|
|
8
|
+
Diagnostics patch. Fail-closed exact-value validation now produces a class-only
|
|
9
|
+
report of the rejected backticked spans, and the existing
|
|
10
|
+
`compaction.custom_fallback` warning carries it as `unsupportedSpanClasses`.
|
|
11
|
+
Validation strictness, the eight-bullet and 25% repair bounds, the fail-closed
|
|
12
|
+
fallback to Pi default compaction, and every provider-facing default are
|
|
13
|
+
unchanged. The report contains class names and counters only: it never contains
|
|
14
|
+
span text, so it is safe to log and to quote in reports.
|
|
15
|
+
|
|
16
|
+
## Changes
|
|
17
|
+
|
|
18
|
+
- `packages/core/src/compaction/summary-contract.ts` adds
|
|
19
|
+
`classifyUnsupportedExactValueSpans()`, a pure observation-only classifier
|
|
20
|
+
returning an `UnsupportedSpanClassReport`:
|
|
21
|
+
- `relations` counts near-miss relations between the rejected span and the
|
|
22
|
+
segment source: `escaped-form-present`, `unescaped-form-present`,
|
|
23
|
+
`whitespace-collapsed-present`, `case-variant-present`,
|
|
24
|
+
`typographic-variant-present`, `single-deletion-present`,
|
|
25
|
+
`composed-two-present-parts`, `no-near-miss`, `not-classified-budget`.
|
|
26
|
+
- `shapes` counts structural shapes: `space`, `backslash`, `double-backslash`,
|
|
27
|
+
`colon`, `equals`, `double-dash`, `slash`, `quote`, `typographic-char`,
|
|
28
|
+
`digit-group-separator`, `json-punctuation`.
|
|
29
|
+
- `lengthBuckets` counts span lengths: `len-1-8`, `len-9-16`, `len-17-32`,
|
|
30
|
+
`len-33-64`, `len-65-120`, `len-121-plus`.
|
|
31
|
+
- `spans`, optional `bullets`, and `classificationComplete` report the totals
|
|
32
|
+
and whether the evidence-probe budget stopped the near-miss classification.
|
|
33
|
+
- The classifier is bounded: probe budget 4000 evidence lookups, near-miss
|
|
34
|
+
spans limited to 96 characters, composed parts required to be at least four
|
|
35
|
+
characters. It cannot alter validation or repair.
|
|
36
|
+
- `src/pi-adapter/summary-generator.ts`: `SummaryValidationError` carries the
|
|
37
|
+
report and machine-readable issue codes. Error message and behavior are
|
|
38
|
+
unchanged.
|
|
39
|
+
- `src/pi-adapter/compaction-coordinator.ts`: the existing
|
|
40
|
+
`custom_fallback` warning gains `unsupportedSpanClasses` when a report exists.
|
|
41
|
+
- Tests: core classifier coverage, generator wiring, coordinator logging, and a
|
|
42
|
+
guard asserting that logged diagnostics contain no span text.
|
|
43
|
+
- `docs/COMPACTION.md`: diagnostics and validation sections describe the report
|
|
44
|
+
and its limits.
|
|
45
|
+
|
|
46
|
+
## Why
|
|
47
|
+
|
|
48
|
+
Production sessions on separate hosts failed closed with
|
|
49
|
+
`compaction.custom_fallback`, `unsupported-exact-value`,
|
|
50
|
+
`repair=too-many-bullets`, and observed counts of 30/20, 23/19, 19/15, and
|
|
51
|
+
17/11 unsupported spans and affected bullets. The bounded repair removes at most
|
|
52
|
+
eight whole bullets and refuses without attempting any removal above that, so
|
|
53
|
+
every one of those runs fell back to Pi default compaction unmodified.
|
|
54
|
+
|
|
55
|
+
Counts alone cannot attribute the failure: a summarizer that invents values and
|
|
56
|
+
a summarizer whose rendering or composition rules differ from the validation
|
|
57
|
+
domain need opposite fixes, and the two are indistinguishable in a bare
|
|
58
|
+
`unsupportedSpans` number. 0.4.1 changed the prompt for the composition
|
|
59
|
+
hypothesis; the following observed runs kept the same failure class and the same
|
|
60
|
+
order of magnitude, so the next step is attribution rather than another prompt
|
|
61
|
+
edit. This release adds only the measurement.
|
|
62
|
+
|
|
63
|
+
## Privacy and safety
|
|
64
|
+
|
|
65
|
+
- No conversation text is added to any log, warning detail, or error message.
|
|
66
|
+
The report is a fixed taxonomy of class names plus integer counters.
|
|
67
|
+
- The classifier performs bounded lookups against the already-sanitized segment
|
|
68
|
+
source used by validation; it does not read session files, SQLite, or any
|
|
69
|
+
other data, and it does not run on successful summaries.
|
|
70
|
+
- Privacy enforcement, provider-facing payloads, and `compaction` phase timings
|
|
71
|
+
are untouched.
|
|
72
|
+
|
|
73
|
+
## Compatibility and persistence
|
|
74
|
+
|
|
75
|
+
- No SQLite migration or schema change (still 16), no canonical Pi JSONL change,
|
|
76
|
+
no privacy consent change, no portable runtime-adapter contract change.
|
|
77
|
+
- Validation strictness, the eight-bullet bound, the 25% removal bound, and the
|
|
78
|
+
fail-open behavior of the extension remain exactly as in 0.4.1.
|
|
79
|
+
- Provider-facing defaults are unchanged: `chars-v1` remains the default
|
|
80
|
+
estimator, `modelAwareness.autoTune` and DS4 native continuation stay
|
|
81
|
+
disabled, `storage.scope` keeps its `project` default.
|
|
82
|
+
|
|
83
|
+
## Measured scope and limitations
|
|
84
|
+
|
|
85
|
+
The classifier is diagnostic: it does not fix the observed fallback and no
|
|
86
|
+
reduction of the fallback rate is claimed. Its output is only as good as the
|
|
87
|
+
rejected spans it is given, and `classificationComplete: false` marks the runs
|
|
88
|
+
where the probe budget stopped the near-miss classification.
|
|
89
|
+
|
|
90
|
+
## Related operational finding (not a change in this release)
|
|
91
|
+
|
|
92
|
+
On the same host an independent failure mode was identified outside DS4: Pi's
|
|
93
|
+
own summarization request caps output at
|
|
94
|
+
`Math.min(Math.floor(0.8 * reserveTokens), model.maxTokens)` with
|
|
95
|
+
`reserveTokens` defaulting to 16384, i.e. 13107 output tokens for the history
|
|
96
|
+
summary and 8192 for the turn-prefix summary, and Pi passes the session thinking
|
|
97
|
+
level to that request. Both the dedicated-model attempt and the fallback attempt
|
|
98
|
+
reported the same stop condition
|
|
99
|
+
(`Summarization failed: generation hit the token cap and the summary is
|
|
100
|
+
incomplete`, raised for `stopReason === "length"`), so compaction ended in an
|
|
101
|
+
error even after DS4 and the dedicated compaction model had both declined.
|
|
102
|
+
Raising `compaction.reserveTokens` globally, or per model through
|
|
103
|
+
`compaction.modelOverrides`, is an operator setting in Pi and is not part of
|
|
104
|
+
this release.
|
|
105
|
+
|
|
106
|
+
## Validation
|
|
107
|
+
|
|
108
|
+
On Node 26.5.1 the release passed `npm run check` (builds, `tsc --noEmit`, 99
|
|
109
|
+
Vitest files and 623 tests), deterministic `npm run quality:compare` (candidate
|
|
110
|
+
`0.9875` against baseline `0.808156`), `npm run schema:context-persistence`
|
|
111
|
+
(`passed: true`), and `npm run pack:check` in a clean consumer for all three
|
|
112
|
+
packages at 0.4.2 (247 core files, 7 reference-adapter files, 100 extension
|
|
113
|
+
files). The coordinated publication then proceeded in dependency order
|
|
114
|
+
(`ds4-context-core`, then `ds4-context-reference-adapter`, then
|
|
115
|
+
`ds4-context-engine`), each command reporting the exact 0.4.2 artifact,
|
|
116
|
+
followed by `npm run registry:check -- 0.4.2`. Registry propagation lagged the
|
|
117
|
+
publishes: the first `npm run registry:check -- 0.4.2` failed with
|
|
118
|
+
`ETARGET ... No matching version found for ds4-context-core@0.4.2`, and six
|
|
119
|
+
retries over roughly four minutes reported the same error while the registry
|
|
120
|
+
metadata already listed 0.4.2. Once propagation completed, the repeated
|
|
121
|
+
`npm run registry:check -- 0.4.2` reported "Verified all DS4 registry packages
|
|
122
|
+
at exact version 0.4.2", and the `latest` dist-tag resolved to 0.4.2 for all
|
|
123
|
+
three packages. No provider calls are involved in this release procedure.
|
|
@@ -0,0 +1,159 @@
|
|
|
1
|
+
# Release 0.4.3 — Fair span coverage and honest "unanalysed" classes
|
|
2
|
+
|
|
3
|
+
**Coordinated packages:** `ds4-context-core`, `ds4-context-reference-adapter`, and `ds4-context-engine` 0.4.3.
|
|
4
|
+
**Implementation commit:** `47100ad`.
|
|
5
|
+
|
|
6
|
+
## Summary
|
|
7
|
+
|
|
8
|
+
Diagnostics patch. The class-only report that `0.4.2` attached to a failed
|
|
9
|
+
compaction could not answer the question it exists for: in its first production
|
|
10
|
+
run, 13 of 23 rejected spans came back as `not-classified-budget`, and the two
|
|
11
|
+
`no-near-miss` entries could not be distinguished from spans the classifier
|
|
12
|
+
never analysed. The classifier now covers every distinct span with the cheap
|
|
13
|
+
transformed-form lookups before any near-miss analysis starts, gives each
|
|
14
|
+
near-miss candidate an equal share of what remains, and reports spans it did not
|
|
15
|
+
analyse as unanalysed instead of guessing.
|
|
16
|
+
|
|
17
|
+
Validation strictness, the eight-bullet and 25% repair bounds, the fail-closed
|
|
18
|
+
fallback to Pi default compaction, the provider-facing defaults, the SQLite
|
|
19
|
+
schema and every configuration key are unchanged. The report still contains
|
|
20
|
+
class names and counters only, so it stays safe to log.
|
|
21
|
+
|
|
22
|
+
## Changes
|
|
23
|
+
|
|
24
|
+
- `packages/core/src/compaction/summary-contract.ts`:
|
|
25
|
+
- Transformed-form lookups (`escaped-form-present`, `unescaped-form-present`,
|
|
26
|
+
`whitespace-collapsed-present`, `case-variant-present`,
|
|
27
|
+
`typographic-variant-present`) run for every distinct span before the first
|
|
28
|
+
near-miss lookup, so span order can no longer decide which spans are
|
|
29
|
+
classified.
|
|
30
|
+
- Each remaining near-miss candidate may spend an equal share of the lookups
|
|
31
|
+
left, at least 32 and never more than the budget remaining, so one long span
|
|
32
|
+
cannot consume the budget of the others.
|
|
33
|
+
- New relations distinguish "not analysed" from "absent": `not-classified-length`
|
|
34
|
+
above the 96-character near-miss limit, `not-classified-partial` when the
|
|
35
|
+
span's share ran out, and the existing `not-classified-budget` when the
|
|
36
|
+
shared budget was already exhausted. `no-near-miss` now means every
|
|
37
|
+
applicable lookup ran and found nothing.
|
|
38
|
+
- The report adds `spansClassifiedCheap` (span occurrences attributed by a
|
|
39
|
+
transformed-form relation), `corpusSources`, `probeBudget` and `probesUsed`.
|
|
40
|
+
`classificationComplete` is false when the shared budget or a per-span share
|
|
41
|
+
stopped an analysis; a `not-classified-length` span does not clear it,
|
|
42
|
+
because that limit is deterministic rather than a resource shortfall.
|
|
43
|
+
- `UnsupportedSpanClassOptions.probeBudget` can override the 4000-lookup
|
|
44
|
+
budget; values below one or non-finite values are ignored. No configuration
|
|
45
|
+
key was added.
|
|
46
|
+
- `tests/unit/summary-contract.test.ts`: the existing bound test now asserts
|
|
47
|
+
`not-classified-partial` and `probesUsed <= probeBudget`, plus new coverage for
|
|
48
|
+
the phase split, the length limit, budget exhaustion and unusable budget
|
|
49
|
+
values.
|
|
50
|
+
- `docs/COMPACTION.md`: the diagnostics section documents the three
|
|
51
|
+
not-classified classes, the phase split and the new counters.
|
|
52
|
+
|
|
53
|
+
No change was needed in `src/pi-adapter/summary-generator.ts` or
|
|
54
|
+
`src/pi-adapter/compaction-coordinator.ts`: the existing error and the existing
|
|
55
|
+
`custom_fallback` warning already carry the report, so the new fields appear in
|
|
56
|
+
`unsupportedSpanClasses` automatically.
|
|
57
|
+
|
|
58
|
+
## Why
|
|
59
|
+
|
|
60
|
+
The first production run of the `0.4.2` report (a manual compaction, 23 rejected
|
|
61
|
+
spans across 16 bullets) came back as:
|
|
62
|
+
|
|
63
|
+
| Class | Count |
|
|
64
|
+
| --- | --- |
|
|
65
|
+
| `escaped-form-present` | 8 |
|
|
66
|
+
| `no-near-miss` | 2 |
|
|
67
|
+
| `not-classified-budget` | 13 |
|
|
68
|
+
|
|
69
|
+
Length buckets were dominated by long spans (10 in `len-33-64`, 7 in
|
|
70
|
+
`len-65-120`, 2 in `len-121-plus`) and shapes by prose (20 spans containing
|
|
71
|
+
spaces, 10 quotes, 10 colons), so the entries looked like quoted excerpts rather
|
|
72
|
+
than identifiers. Only the eight escaped-form spans were actionable: the value
|
|
73
|
+
was present, only its rendering differed. Two defects in the instrument blocked
|
|
74
|
+
the rest:
|
|
75
|
+
|
|
76
|
+
- **Coverage.** Every lookup costs one scan of each corpus source
|
|
77
|
+
(`sourceText`, `readFiles`, `modifiedFiles`) and the budget was spent in span
|
|
78
|
+
order, while `single-deletion-present` costs up to one lookup per character
|
|
79
|
+
and `composed-two-present-parts` up to two more per length. A handful of long
|
|
80
|
+
spans, or their complete analysis, could therefore exhaust the budget before
|
|
81
|
+
the remaining spans received a single cheap lookup.
|
|
82
|
+
- **Ambiguity.** `no-near-miss` was also returned for spans above
|
|
83
|
+
`MAX_NEAR_MISS_SPAN_LENGTH`, where the near-miss analysis is never attempted.
|
|
84
|
+
|
|
85
|
+
Neither defect weakened validation: the repair bound and the fail-closed
|
|
86
|
+
fallback were respected in every one of these runs. The problem was that counts
|
|
87
|
+
alone could not separate a summarizer that invents values from one whose
|
|
88
|
+
rendering or composition rules differ from the validation domain, and those two
|
|
89
|
+
findings call for opposite fixes. This release makes the next report
|
|
90
|
+
interpretable before any change to validation, repair, or the summarizer prompt
|
|
91
|
+
is considered.
|
|
92
|
+
|
|
93
|
+
## Privacy and safety
|
|
94
|
+
|
|
95
|
+
- The classifier remains observation-only and cannot alter validation, repair
|
|
96
|
+
bounds, or the fail-closed decision.
|
|
97
|
+
- The report still contains class names and counters only. The new fields are
|
|
98
|
+
integers (occurrence, source and lookup counts); no span text, no evidence
|
|
99
|
+
text, and no hashes of either are added to logs, warnings, or errors.
|
|
100
|
+
- No provider call, network access, or configuration change is involved in this
|
|
101
|
+
release; the classifier runs only after a validation failure has already been
|
|
102
|
+
decided.
|
|
103
|
+
|
|
104
|
+
## Compatibility and persistence
|
|
105
|
+
|
|
106
|
+
- No migrations, no schema change (still 16), no new configuration key, and no
|
|
107
|
+
change to any default.
|
|
108
|
+
- The report shape is additive: existing consumers keep working, but a reader of
|
|
109
|
+
`relations` must tolerate the new keys `not-classified-length` and
|
|
110
|
+
`not-classified-partial`, and can no longer assume that `no-near-miss` covers
|
|
111
|
+
every unanalysed span.
|
|
112
|
+
- `classificationComplete: false` is now also possible with the shared budget
|
|
113
|
+
intact, when a per-span share stops an analysis.
|
|
114
|
+
|
|
115
|
+
## Measured scope and limitations
|
|
116
|
+
|
|
117
|
+
Measured locally on the development host with the unit fixtures, no provider
|
|
118
|
+
calls:
|
|
119
|
+
|
|
120
|
+
- A 96-character span with no transformed-form match costs 185 lookups when
|
|
121
|
+
analysed completely (96 single-deletion plus 89 prefix probes), confirming that
|
|
122
|
+
two or three such spans could exhaust the whole budget under the `0.4.2` order.
|
|
123
|
+
- With 80 distinct spans of 91-92 characters and the default budget, the cheap
|
|
124
|
+
phase spends 80 lookups and each span then receives a 49-lookup share:
|
|
125
|
+
`probesUsed` equals `probeBudget` exactly and every span is reported as
|
|
126
|
+
`not-classified-partial`.
|
|
127
|
+
- The knife edge is observable: in a combined fixture, a budget of 190 reports
|
|
128
|
+
`not-classified-partial` while 191 reports `no-near-miss`, because the share is
|
|
129
|
+
exhausted on the lookup that completes the analysis.
|
|
130
|
+
|
|
131
|
+
Limitations that remain:
|
|
132
|
+
|
|
133
|
+
- Distinct spans are deduplicated, so `spansClassifiedCheap` counts occurrences
|
|
134
|
+
while the budget is spent per distinct value.
|
|
135
|
+
- Spans above 96 characters are still never near-miss analysed; they now say so
|
|
136
|
+
instead of claiming absence.
|
|
137
|
+
- The default budget is still 4000 lookups per failure, unchanged from `0.4.2`.
|
|
138
|
+
- A `not-classified-*` class still cannot say whether the span was invented;
|
|
139
|
+
it says only that the classifier did not look.
|
|
140
|
+
|
|
141
|
+
## Validation
|
|
142
|
+
|
|
143
|
+
- `npm run build:core`, `npm run build:adapters`, and `npx tsc --noEmit` clean.
|
|
144
|
+
- `npm run check`: 99 test files, 627 tests, all passing.
|
|
145
|
+
- `npm run quality:compare`, `npm run schema:context-persistence` (passed) and
|
|
146
|
+
`npm run pack:check`: core 0.4.3 (247 files), reference adapter 0.4.3
|
|
147
|
+
(7 files) and engine 0.4.3 (101 files) verified in a clean consumer.
|
|
148
|
+
- Classifier behavior was additionally checked against the `0.4.2` semantics
|
|
149
|
+
with a synthetic reconstruction of the production histogram shape (long spans
|
|
150
|
+
first, cheap matches after), without touching session data.
|
|
151
|
+
- Not verified: the improved histogram on a real failed compaction. That
|
|
152
|
+
requires a natural failure and is expected to appear on the next
|
|
153
|
+
`compaction.custom_fallback` warning.
|
|
154
|
+
|
|
155
|
+
## Publication
|
|
156
|
+
|
|
157
|
+
Published manually in dependency order (core, reference adapter, engine) with
|
|
158
|
+
`npm publish`, after `npm run pack:check` on the same commit. Registry
|
|
159
|
+
verification is recorded in the follow-up documentation commit.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "ds4-context-engine",
|
|
3
|
-
"version": "0.4.
|
|
3
|
+
"version": "0.4.3",
|
|
4
4
|
"description": "Non-destructive, provider-independent context management for Pi.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|
|
@@ -63,7 +63,7 @@
|
|
|
63
63
|
]
|
|
64
64
|
},
|
|
65
65
|
"dependencies": {
|
|
66
|
-
"ds4-context-core": "0.4.
|
|
66
|
+
"ds4-context-core": "0.4.3",
|
|
67
67
|
"js-tiktoken": "1.0.21"
|
|
68
68
|
},
|
|
69
69
|
"peerDependencies": {
|
|
@@ -43,6 +43,7 @@ import {
|
|
|
43
43
|
import {
|
|
44
44
|
generateValidatedSummary,
|
|
45
45
|
sumUsage,
|
|
46
|
+
SummaryValidationError,
|
|
46
47
|
type GeneratedSummary,
|
|
47
48
|
} from "./summary-generator.ts";
|
|
48
49
|
import {
|
|
@@ -501,6 +502,9 @@ export class CompactionCoordinator {
|
|
|
501
502
|
};
|
|
502
503
|
} catch (error) {
|
|
503
504
|
const message = error instanceof Error ? error.message : String(error);
|
|
505
|
+
const spanClasses = error instanceof SummaryValidationError
|
|
506
|
+
? error.spanClassReport
|
|
507
|
+
: undefined;
|
|
504
508
|
this.state = {
|
|
505
509
|
...this.state,
|
|
506
510
|
phase: "failed",
|
|
@@ -508,7 +512,11 @@ export class CompactionCoordinator {
|
|
|
508
512
|
completedAt: this.dependencies.now(),
|
|
509
513
|
lastError: message,
|
|
510
514
|
};
|
|
511
|
-
this.dependencies.logger.warn("compaction.custom_fallback", {
|
|
515
|
+
this.dependencies.logger.warn("compaction.custom_fallback", {
|
|
516
|
+
trigger,
|
|
517
|
+
error: message,
|
|
518
|
+
...(spanClasses ? { unsupportedSpanClasses: spanClasses } : {}),
|
|
519
|
+
});
|
|
512
520
|
if (!event.signal.aborted && ctx.hasUI) {
|
|
513
521
|
ctx.ui.notify(`DS4 compaction unavailable; using Pi default. ${message}`, "warning");
|
|
514
522
|
}
|
|
@@ -7,16 +7,38 @@ import type {
|
|
|
7
7
|
import type { CompactionThinkingLevel } from "ds4-context-core/config/config";
|
|
8
8
|
import {
|
|
9
9
|
analyzeUnsupportedExactValueBullets,
|
|
10
|
+
classifyUnsupportedExactValueSpans,
|
|
10
11
|
groundSummaryFileSections,
|
|
11
12
|
validateSummary,
|
|
12
13
|
type SummaryValidationInput,
|
|
13
14
|
type SummaryValidationResult,
|
|
15
|
+
type UnsupportedSpanClassReport,
|
|
14
16
|
} from "ds4-context-core/compaction/summary-contract";
|
|
15
17
|
|
|
16
18
|
export const DEFAULT_COMPACTION_TRANSPORT_MAX_ATTEMPTS = 4;
|
|
17
19
|
export const DEFAULT_COMPACTION_TRANSPORT_BASE_DELAY_MS = 2000;
|
|
18
20
|
export const COMPACTION_TRANSPORT_MAX_DELAY_MS = 60_000;
|
|
19
21
|
|
|
22
|
+
/**
|
|
23
|
+
* Fail-closed summary rejection. Carries the class-only span diagnostics so the
|
|
24
|
+
* coordinator can log why spans were rejected without logging span text.
|
|
25
|
+
*/
|
|
26
|
+
export class SummaryValidationError extends Error {
|
|
27
|
+
readonly codes: readonly string[];
|
|
28
|
+
readonly spanClassReport?: UnsupportedSpanClassReport;
|
|
29
|
+
|
|
30
|
+
constructor(
|
|
31
|
+
message: string,
|
|
32
|
+
codes: readonly string[],
|
|
33
|
+
spanClassReport?: UnsupportedSpanClassReport,
|
|
34
|
+
) {
|
|
35
|
+
super(message);
|
|
36
|
+
this.name = "SummaryValidationError";
|
|
37
|
+
this.codes = codes;
|
|
38
|
+
this.spanClassReport = spanClassReport;
|
|
39
|
+
}
|
|
40
|
+
}
|
|
41
|
+
|
|
20
42
|
/**
|
|
21
43
|
* Transport retry policy for compaction summary requests: four total attempts
|
|
22
44
|
* (the initial call plus three retries), 2000 ms base delay, exponential backoff, abort-aware.
|
|
@@ -330,11 +352,20 @@ export async function generateValidatedSummary(
|
|
|
330
352
|
}
|
|
331
353
|
}
|
|
332
354
|
if (validation.status === "invalid") {
|
|
333
|
-
const codes = unique(validation.issues.map((issue) => issue.code))
|
|
355
|
+
const codes = unique(validation.issues.map((issue) => issue.code));
|
|
334
356
|
const repairDiagnostics = exactRepair
|
|
335
357
|
? `; repair=${exactRepairFailure ?? exactRepair.status}; unsupportedSpans=${exactRepair.unsupportedSpans}; affectedBullets=${exactRepair.affectedBullets}`
|
|
336
358
|
: "";
|
|
337
|
-
|
|
359
|
+
const spanClassReport = validation.issues.some((issue) => issue.code === "unsupported-exact-value")
|
|
360
|
+
? classifyUnsupportedExactValueSpans(content, validationInput, {
|
|
361
|
+
...(exactRepair ? { affectedBullets: exactRepair.affectedBullets } : {}),
|
|
362
|
+
})
|
|
363
|
+
: undefined;
|
|
364
|
+
throw new SummaryValidationError(
|
|
365
|
+
`Compaction ${input.stage} summary validation failed: ${codes.join(", ")}${repairDiagnostics}`,
|
|
366
|
+
codes,
|
|
367
|
+
spanClassReport,
|
|
368
|
+
);
|
|
338
369
|
}
|
|
339
370
|
return { content, validation, usage: sumUsage([...retryUsages, response.usage]) };
|
|
340
371
|
}
|