dsh-jev-prune 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,48 @@
1
+ # Contributing
2
+
3
+ Thanks for improving `dsh-jev-prune`. The repository targets a pre-release DSH runtime, so small host-shape assumptions can have large effects. Keep changes narrow and make those assumptions testable.
4
+
5
+ ## Start from current `main`
6
+
7
+ ```bash
8
+ git fetch origin
9
+ git switch main
10
+ git pull --ff-only
11
+ git switch -c <type>/<short-name>
12
+ npm ci
13
+ ```
14
+
15
+ Do not build a PR on an older feature branch. DSH-facing fixes often touch the same hook and event-shape code, and stale branches caused earlier contributions to conflict after review.
16
+
17
+ ## Validate changes
18
+
19
+ Run the checks that match CI:
20
+
21
+ ```bash
22
+ npm run check
23
+ npm run smoke
24
+ npm run coverage
25
+ ```
26
+
27
+ - `check.js` covers pure helpers and safety invariants.
28
+ - `smoke_apply.mjs` runs the real `apply()` wiring against DSH-shaped fakes.
29
+ - Coverage is measured for `jev.js`, `prune.js`, `receipt.js`, and `state.js`; CI enforces minimum line, branch, and function coverage.
30
+ - CI runs the unit suite on Ubuntu and Windows with Node 22 and 24. The integration job uses the locked DSH 0.1.5-rc.2 fixture.
31
+
32
+ When changing a host assumption, also run `jev_probe_shapes` in a real DSH session when possible and include the observed shape in the PR description without secrets or full session content.
33
+
34
+ ## Pull requests
35
+
36
+ Keep one coherent concern per PR. Explain the concrete failure trigger, the new behaviour, and the validation performed. Update README/configuration tables when defaults or user-visible behaviour change.
37
+
38
+ Before opening a PR:
39
+
40
+ - Rebase or merge the latest `main` and resolve conflicts locally.
41
+ - Add a regression that fails for the old implementation for behavioural fixes.
42
+ - Preserve tool-call pairing, recent-node protection, evidence guards, and fallback behaviour.
43
+ - Confirm every unit-matrix, coverage, and integration check passes.
44
+ - Link the issue with `Fixes #N` only when the PR fully resolves it; use `Refs #N` for partial work.
45
+
46
+ ## Reporting issues
47
+
48
+ Use the issue templates and include the DSH version, Node version, relevant plugin configuration, and a redacted status/heartbeat excerpt. Never post API keys or complete private session logs.
@@ -0,0 +1,60 @@
1
+ # Porting contract
2
+
3
+ This document describes the host surface required to run `dsh-jev-prune` outside the tested DSH 0.1.5-rc.2 bundle. Start by loading the plugin and calling `jev_probe_shapes` in a disposable session. Compare the result with the assumptions below before enabling writes.
4
+
5
+ ## Services
6
+
7
+ The adapter in `index.js` consumes these services:
8
+
9
+ | Service | Required behaviour |
10
+ |---|---|
11
+ | `toolResultPruner` | Exposes synchronous `pruneSession(session)`, `measureContent(blocks)`, `pruneContent(blocks)`, and a token estimator through its context. The plugin replaces `pruneSession`. |
12
+ | `compaction` | Exposes async `compactRegion(startSeq, endSeq, agent, signal)` and `summarize(input, agent, signal)`. The plugin temporarily wraps `summarize` to inject a deterministic receipt. |
13
+ | `tools` | Registers `jev_prune_status`, `jev_prune_now`, `jev_compact_now`, `jev_restore`, and `jev_probe_shapes`. Missing tools must not disable pruning. |
14
+ | `commands` | Optionally registers `/jev`. |
15
+ | `llm` | `resolveModelInfo(provider, model)` returns `context.contextWindow` for ratio pressure gates. |
16
+ | `tokenMeter` | `measure(session)` returns `totalTokens` and preferably per-node token entries. Required for ratio pressure comparisons and receipt-size checks. |
17
+
18
+ `@deepseek-ai/dsh-llm` is optional at module-load time. If its `freezeMessage` export cannot be loaded, the plugin uses a shallow-copy fallback.
19
+
20
+ ## Hook semantics
21
+
22
+ The host must provide waterfall-style `agent/pre-step` hooks with `next()` chaining. Registration with the prepend flag must run the judge hook before the host's normal compaction hook; otherwise synchronous pruning sees stale verdicts.
23
+
24
+ The event payload must expose `agent`, optional `signal`, and `agent.session`. Cancellation signals are passed to judge and compaction calls.
25
+
26
+ ## Session and event shape
27
+
28
+ The active session must provide:
29
+
30
+ - `surface.nodes`: ordered active seq values.
31
+ - `eventAt(seq)`: returns an event whose `event.seq` equals the requested seq.
32
+ - `deriveEventMessage(event)`: returns the DSH message represented by an event.
33
+ - `append(type, data, options)`: appends an event and applies `surfaceOp`.
34
+
35
+ The plugin recognizes:
36
+
37
+ - `assistant/message` with `data.message.content[]` blocks of type `tool-call`; call identity may be in `id`, `callId`, or compatible fields handled by `state.js`.
38
+ - `tool/result` with `data.message.source.callId` and a `tool-result` content block.
39
+ - Replacement events with `sourceEventSeqs` and `{ surfaceOp: { op: 'replace', startSeq, endSeq } }`.
40
+ - `compaction/summary` and checkpoint events carrying `shadowedRange`, `shadowedSeqs`, and a compaction identifier.
41
+
42
+ Old events must remain addressable through `eventAt` after a surface replacement. Evidence scanning and verdict inheritance depend on that append-only log.
43
+
44
+ ## Compaction protocol
45
+
46
+ `compactRegion` must reject or avoid unbalanced tool pairs and must dynamically dispatch through its current `summarize` method. If it captures a private summarizer before the plugin wraps the method, deterministic receipt injection cannot work and layer 2 must remain disabled.
47
+
48
+ The return value should expose `shadowedSeqs` and `compactionId`. Missing optional fields reduce observability but should not corrupt the surface.
49
+
50
+ ## Porting checklist
51
+
52
+ 1. Install the plugin with `compactReceipts: false` and `dryRun: true`.
53
+ 2. Run `jev_probe_shapes`; verify event types, content block names, call IDs, and tool names.
54
+ 3. Run `node test/check.js` and `node test/smoke_apply.mjs` against the target dependency tree.
55
+ 4. Confirm the judge hook runs before the host calls `pruneSession`.
56
+ 5. Enable layer 1 and inspect `jev_prune_status` plus the heartbeat file.
57
+ 6. Enable layer 2 in dry-run mode; confirm selected ranges begin and end on balanced cuts.
58
+ 7. Perform one real receipt compaction and verify its provider is `jev-receipt` and the original events remain restorable.
59
+
60
+ Document every adapter needed for a new host version. Do not silently coerce unknown event shapes into the tested DSH format.
@@ -0,0 +1,27 @@
1
+ # Fixed-version release preparation
2
+
3
+ Release tags, GitHub releases and npm packages are created from a verified merged commit. `package.json` identifies version `0.1.0`; check the GitHub release and npm registry for distribution availability.
4
+
5
+ ## Before tagging
6
+
7
+ 1. Resolve or explicitly accept the [runtime review findings](review-2026-10-05.md). Use an experimental prerelease while they remain; do not imply production readiness.
8
+ 2. Review PR #45's changelog/install changes against this PR; reconcile overlapping README changes. Keep issue #35 open until published artifacts exist.
9
+ 3. From the merged commit, run the check, smoke, coverage and deterministic demo commands. CI must pass; validate the intended DSH profile in a real host separately.
10
+ 4. Choose a version, update `package.json` and the lockfile together, and change the changelog's prepared heading only when actually publishing.
11
+ 5. Run `npm pack --dry-run`; check documentation, replay and demo assets ship. Never include private logs or credentials.
12
+
13
+ ## Pin and release the verified commit
14
+
15
+ ```bash
16
+ # Replace both placeholders with the chosen version and VERIFIED merged SHA.
17
+ git tag -a <version-tag> <verified-commit-sha> -m "Experimental release"
18
+ git push origin <version-tag>
19
+ # Create a GitHub release for that tag, with experimental status and known limits.
20
+ dsh plugin --profile web add github:yangyu666/dsh-jev-prune#<version-tag>
21
+ ```
22
+
23
+ Verify that the pinned installation succeeds in a clean DSH profile. A tag must identify the validated merge result; it must not point to an older branch or to this PR's unmerged head by accident. The GIF/JSON/cast can be attached to the GitHub release.
24
+
25
+ ## Optional npm distribution
26
+
27
+ Use the owner's npm login and required authentication to publish the exact version. Check `npm view dsh-jev-prune@<version> version` and install that exact published version in a clean profile before adding an npm-first command to the README. GitHub tag publication and npm publication are separate operations.
@@ -0,0 +1,310 @@
1
+ # Implementation reference
2
+
3
+ Moved from the original README. For current limits, read [the review findings](review-2026-10-05.md); the ownership mechanism below does not establish external transaction identity.
4
+
5
+ ## The problem it solves
6
+
7
+ DSH's built-in context reclamation is **purely volumetric**. Once a tool result crosses a size threshold, its middle is chopped out and the head and tail are kept; region compaction, meanwhile, has the model **write a summary** to stand in for old history. The first approach cannot tell "this result is large but I still need it" from "this one is spent", and the second one invites summary hallucination.
8
+
9
+ This plugin replaces the decision in both places with Jev's structured output (`noul` / `choice`, returning calibrated probabilities), under one design rule:
10
+
11
+ > **What should not be generated by a model is not generated by a model.** Trimming only ever decides *keep* or *discard*; the original text is preserved verbatim. Region compaction injects a **deterministic receipt** produced by code, containing no model inference at all.
12
+
13
+ ## The two layers
14
+
15
+ ![The two layers: result trimming and receipt compaction](../assets/two-layers.png)
16
+
17
+ | Layer | Interception point | DSH default | This plugin |
18
+ |---|---|---|---|
19
+ | **1 · Result trimming** | `ctx.toolResultPruner.pruneSession` | Chops the middle once `thresholdChars` is exceeded | Jev decides, per tool result, whether it will still be needed. Needed ones are **never trimmed, however large**; stale ones are trimmed **however small** (unless shorter than `minCharsToPrune`); with no judgment available it falls back to DSH's original behaviour |
20
+ | **2 · Receipt compaction** | `ctx.compaction.summarize` + `compactRegion`, or a single-result surface replacement for mixed batches | The model reads the raw history and writes a summary | Moves fully eligible read-only steps out of the surface. In a parallel batch where only some results qualify, it keeps every call/result envelope and replaces only the eligible result bodies with **deterministic receipts**. Tool name, command, path, character count and seq are all computed by code |
21
+
22
+ A layer-2 receipt looks like this:
23
+
24
+ ```
25
+ [已压缩 · 确定性回执] 原历史 s25–s27 是 1 次工具调用(共约 16489 字符输出),
26
+ 为释放上下文已移出。以下为事实清单(工具名/入参/字符数/seq 由代码算出;模型原话逐字引用,不含任何推断):
27
+ · s25 模型原话(逐字引用):identifiers:把版本字符串拆成标识符。Next inc.js.
28
+ · s27 read:C:\Users\you\project\src\state.js → 16489 字符输出
29
+ 原始事件仍完整保存在会话日志中(seqs 25–27)。需要内容时重跑相同命令/读取相同文件即可;本回执不含对内容的解释。
30
+ ```
31
+
32
+ Parallel tool batches are evaluated per call/result pair. DSH 0.1.5 can replace one surface node or one balanced contiguous region, but cannot remove five pairs from a six-call assistant message in one atomic operation. For a mixed batch, the plugin therefore replaces each eligible `tool/result` body with a short receipt through DSH's native single-node `replace` protocol; rejected and recent results remain byte-for-byte unchanged, the assistant head stays intact, and tool pairing remains valid after every replacement. Results are matched to calls by `callId`, so completion order may differ from declaration order. A fully eligible batch still uses `compactRegion` and removes the whole balanced step.
33
+
34
+ > The receipt body is emitted by the plugin's own JavaScript, so its wording is Chinese today — as is the `/jev` status output. The `jev_*` tool descriptions are already English. Localizing the runtime strings is a separate change.
35
+
36
+ ## Measured results
37
+
38
+ Numbers below come from a three-arm measurement of a 37-step "read every file in a directory, strictly one at a time" long task over `semver@6e05b76`: **A** = vanilla DSH (baseline), **B** = the plugin with `compactReceipts: false` (layer 1 only), **C** = the plugin with defaults (both layers). Every number is taken from the `usage` object the API actually returned in the DSH session logs — no estimator is involved in scoring.
39
+
40
+ **Layer 1 acts on every long run; the baseline never does.** Layer 1 trimmed stale results in every plugin-arm run (4–7 nodes per long run; arm A trims 0 by definition), and answers on the short/medium/long tasks were correct in every arm.
41
+
42
+ **Long tasks run on a fraction of the full-price input.** Uncached ("full-price") input tokens per completed 37-step run: baseline 75,781–87,872 vs full plugin 25,516–45,071 — roughly a third of the baseline at equal task scope. Report token classes separately: the cache-hit share of input is high (78–94%) and hit/miss prices differ by ~50×, so total-token comparisons mislead by design.
43
+
44
+ **Fewer destructive compactions — some of them receipts instead of summaries.** Per completed long run, DSH's own model-written summary compactions dropped from 31–40 (baseline) to 15–23 with the full plugin, of which 6–8 were replaced by deterministic receipts rendered by code. Layer 2 stays deliberately silent on short/medium tasks — its gates require a contiguous run of eligible read-only steps — so on those workloads the effect comes from layer 1. Even when layer 2 fires, DSH-initiated compactions still exist and fall back to model summaries: the plugin reduces them, it does not eliminate them.
45
+
46
+ **A failure mode we found and fixed.** Receipts replace a whole "call + result" step, including the assistant message that carried it; in one long run the model's intermediate notes were erased step by step (12 of 13) and it stopped issuing tool calls mid-task. Receipts now carry each step's **assistant-visible text verbatim** (`text` blocks only, `reasoning` drafts excluded, zero model generation; bounded by `receiptTextChars`, default 400, `0` disables). After the fix, the same long task completed in every run with correct answers.
47
+
48
+ ## Gating (layer 2)
49
+
50
+ Moving a whole pair out of the surface is destructive, so the default is deliberately conservative. Every one of the following must hold:
51
+
52
+ - **Intersection of two axes**: `result` (is the content still needed) and `effect` (did the call change state outside the session) must *each* fall inside this session's trailing `compactQuantile`
53
+ - The tool is not in `neverCompactTools` (write-type calls are excluded by a hard rule, never by a probability)
54
+ - **Evidence guard**: results matching `error` / `assert` / `fail` / `todo` and friends are never moved out. If layer 1 has already trimmed a result, the guard follows `sourceEventSeqs` and scans the original event too
55
+ - Steps whose assistant **text** exceeds `maxStepTextChars`, or whose **`reasoning`** exceeds `maxStepReasoningChars`, are never moved out. The two are measured separately on purpose: long `text` means the step is delivering a conclusion worth keeping, while long `reasoning` is just scratch work — merging them into one budget let reasoning length alone silently shut layer 2 off
56
+ - Anything within the most recent `compactPreserveRecent` nodes is skipped by layer 2 (layer 1 uses `preserveRecent`)
57
+ - Both ends of the range must satisfy DSH's tool-pairing balance; the span must save at least `compactMinChars` characters; and the receipt must stay below `receiptMaxRatio` of the original content's tokens
58
+
59
+ Probabilities are consumed as **relative quantiles**, never as a fixed threshold: the output distribution of a small judge model is narrow, and only the relative ordering *within one session* carries stable information.
60
+
61
+ **Degradation on small populations.** Read-only tools are often a minority in write/execute-heavy sessions (measured: 1 in 6), which can leave a quantile population of only two or three items — too few for ordering to mean anything. Rather than giving up, the mode degrades to an **absolute floor**: both axes must fall below `floorThreshold` (default `0.2`, materially stricter than `compactThreshold`, compensating for the missing relative information). If the population is below `minCandidatesForFloor` (default `2`), nothing is moved out — a single sample is not a distribution. The default was lowered from `3` to `2` so that the two-candidate populations that batch-read sessions really produce are not skipped outright; one sample still never acts. Degradation is always reported in the report and the heartbeat; it never happens silently.
62
+
63
+ ## Receipt ownership (fence)
64
+
65
+ The plugin assigns a token to each pending receipt, permits only one consumption, and clears only its own pending entry. This guards against duplicate consumption, but does not establish ownership of an external concurrent summary: external host paths do not update the plugin's active token. The original smoke C3 test checks a competing call after consumption, not competing arrival before consumption.
66
+
67
+ Until transaction identity is propagated or all entry paths are serialized, do not claim concurrency safety. [Reproduction and suggested remedy](review-2026-10-05.md#p1--external-summary-can-consume-another-transactions-receipt)
68
+
69
+ ## Requirements
70
+
71
+ - Node `^22.19.0 || >=24.0.0`
72
+ - `dsh` (`@deepseek-ai/dsh`), with the base bundle loaded into the profile (`tool-result-pruner` and `compaction-basic` are included by default)
73
+ - A TypeSafe API key (`TYPESAFE_API_KEY` environment variable)
74
+ - Runtime peer dependencies: `@deepseek-ai/schemastery`, `@deepseek-ai/dsh-tools` (provided by the host)
75
+ - Optional dynamic dependency: `freezeMessage` from `@deepseek-ai/dsh-llm` (falls back to a shallow copy when absent; the plugin keeps working)
76
+ - Host services consumed: `toolResultPruner`, `compaction`, `tools`, `commands`, `llm`, and **`tokenMeter`**. `tokenMeter` is used by the pressure gates; if your host does not register it, ratio-based gating cannot compare anything — set `softLimit` / `compactSoftLimit` to an **absolute token count**, or use `judgeOn: 'always'` / `compactOn: 'always'` (see *Pressure-gate failure direction*). A missing meter never silently disables a layer: the gate is left un-armed for that pass and the reason is logged at `warn` level.
77
+
78
+ **Version alignment**: the peer range for `@deepseek-ai/dsh-tools` is `^0.1.5-rc.2` — the tested version `0.1.5-rc.2` sits on npm's `next` tag, not `latest`. A lockfile is committed (devDependencies pin the tested versions), so `npm ci` reproduces the exact test conditions.
79
+
80
+ **Compatibility**: tested against `@deepseek-ai/dsh@0.1.5-rc.2`. DSH 0.1.x is a pre-release line, and event shapes and service names can shift between rc versions. After upgrading DSH, re-run `npm run check` and the smoke test, and call `jev_probe_shapes` once in a real session to verify the field assumptions.
81
+
82
+ Project documentation: [contributing](CONTRIBUTING.md), [architecture](ARCHITECTURE.md), [porting contract](PORTING.md), and [configuration examples](../examples/README.md).
83
+
84
+ ## Install
85
+
86
+ ```bash
87
+ # Install from a local directory (the repo directory name matches the package name)
88
+ dsh plugin --profile web add link:/absolute/path/to/dsh-jev-prune
89
+
90
+ # Confirm it made it into the config tree
91
+ dsh --profile web --dump-config | grep jev-prune
92
+ ```
93
+
94
+ On a machine without pnpm, the equivalent manual wiring (idempotent) is:
95
+
96
+ ```bash
97
+ node scripts/wire_profile.mjs <DSH_HOME> <profile-name>
98
+ ```
99
+
100
+ ## Configuration
101
+
102
+ | Key | Default | Description |
103
+ |---|---|---|
104
+ | `enabled` | `true` | Master switch |
105
+ | `model` | `jev-latest` | Judge model |
106
+ | `keepMode` | `budget` | Layer 1 decision rule. `budget`: *how much* to trim is set by the pressure-gap ratio, *which* results by Jev's ranking (see below). `absolute`: the legacy fixed-threshold behaviour |
107
+ | `keepThreshold` | `0.5` | Layer 1: in `absolute` mode, `P(keep)` ≥ this means no trimming; in `budget` mode it is a **protection ceiling** only (results at or above it never enter the candidate pool) |
108
+ | `alwaysTrimRatio` | `0.5` | Layer 1: fixed trim ratio used **only** under `judgeOn: 'always'` (that mode has no pressure signal to derive one from). The budget is this fraction of the candidate pool's total character gain. `pressure` mode computes the ratio from the gap and ignores this key |
109
+ | `keepFloorThreshold` / `minCandidatesForBudget` | `0.2` / `4` | `budget` mode small-population fallback: with fewer than 4 judged candidates, only results with `P(keep) < 0.2` are eligible (same degraded-mode shape as layer 2) |
110
+ | `resultExcerptChars` | `240` | Layer 1: per-result excerpt budget copied into the judge's state (see below); `0` restores the blind `ok, N chars` line |
111
+ | `preserveRecent` | `4` | Layer 1 leaves the most recent N surface nodes alone |
112
+ | `headChars` / `tailChars` | `600` / `200` | Layer 1: how many head/tail characters a trim keeps |
113
+ | `minCharsToPrune` | `400` | Layer 1: anything shorter is never trimmed |
114
+ | `judgeOn` / `softLimit` | `pressure` / `55%` | Layer 1: when to judge, and the pressure line |
115
+ | `compactReceipts` / `compactOn` | `true` / `pressure` | Layer 2: switch and pressure line (`compactSoftLimit`, default 70%) |
116
+ | `compactMode` | `relative` | `relative` (recommended) or `absolute` (with `compactThreshold`) |
117
+ | `compactQuantile` | `0.34` | The trailing fraction taken on each of the two axes; the intersection is used |
118
+ | `compactPreserveRecent` | `1` | Layer 2 leaves the most recent N surface nodes alone; independent from layer 1's wider recent window |
119
+ | `minCandidatesForRelative` | `4` | Minimum population for relative quantiles; below it the mode **degrades** to an absolute floor (see below) rather than giving up |
120
+ | `floorThreshold` / `minCandidatesForFloor` | `0.2` / `2` | Absolute floor used in the degraded mode (materially stricter than `compactThreshold`) and its minimum sample size |
121
+ | `neverCompactTools` | write-type tools | Layer 2 never moves these out; comparison is normalized (`Edit` ≡ `edit`) |
122
+ | `neverPruneTools` | `Write` / `NotebookEdit` | Layer **1** never touches these. Narrower than the row above on purpose: layer 1 only truncates (reversible, the original stays in the session log), so the arguments of diff-style editors (`Edit`/`ApplyPatch`…) are fair game; layer 2 removes the pair outright, so it keeps guarding all of them |
123
+ | `compactTools` | read-only set | Allow-list, **non-empty by default** (`DSH_READONLY_TOOLS`: `read`/`glob`/`grep`/`list`/`fetch`… plus PowerShell read-only cmdlets such as `getchilditem`/`selectstring`). Setting it to `[]` relaxes the gate to the deny-list only — shell calls then become movable too, which is an explicit opt-in into an unsafe mode |
124
+ | `evidenceGuard` / `evidencePatterns` | `true` / built-in list | Evidence guard |
125
+ | `compactMinChars` / `receiptMaxRatio` | `2000` / `0.5` | Layer 2 economical floors |
126
+ | `receiptArgChars` / `receiptTextChars` | `120` / `400` | Maximum code-point lengths for each rendered call argument and assistant-visible-text excerpt in a receipt; set the latter to `0` to omit assistant excerpts |
127
+ | `maxCompactionsPerPass` | `3` | How many compaction transactions one pass may run. Raised to 3 so a large context converges in a **single** pass instead of being squeezed across many pre-steps; set to `1` for the old behaviour |
128
+ | `judgeMaxRetries` / `judgeRetryBaseMs` | `2` / `300` | Retry count and backoff base for judge requests (see below); `0` disables retries |
129
+ | `dryRun` | `false` | Both layers only judge and account; nothing is changed |
130
+ | `heartbeatFile` | `''` | Where to persist state (the host swallows plugin logs, so a file is the only external observation channel) |
131
+
132
+ > Deprecated keys kept only for compatibility (they no longer take effect): `volumeBudgetThresholdChars`, `budgetMinChars` — early `budget` mode anchored the budget on the volume rule; it now uses the pressure-gap ratio.
133
+
134
+ ### Retrying judge requests
135
+
136
+ A single network hiccup used to void the **entire round** of judging — no candidate got a probability and both layers silently did nothing. Failures are now classified:
137
+
138
+ | Failure | Handling |
139
+ |---|---|
140
+ | Network error / timeout | **Retry** with exponential backoff (`300ms` → `600ms`, up to 2 retries by default) |
141
+ | `429` / `5xx` | **Retry** (server temporarily unavailable) |
142
+ | Other `4xx` (`401` bad key, `400` malformed request) | **No retry** — immediately fatal; retrying only burns quota |
143
+ | Response missing `answers` | **No retry** (a retry would most likely return the same broken body) |
144
+ | External `signal` already aborted | **No retry**, and no new request is issued |
145
+
146
+ Batches are isolated too: one failed batch no longer discards the remaining ones, and the count shows up as `失败批次 N 个` in the status report. Only when **every** batch fails is the round treated as failed.
147
+
148
+ **Counter semantics**: `client.lastRetries` is reset at the start of every `ask` and is what the status report shows, while `client.retries` is the lifetime total for the client (useful for "has this client ever had to retry?"). `client.requests` counts **HTTP attempts actually issued**, including failed ones, so the identity `requests === successful asks + retries` holds.
149
+
150
+ ### Token-estimate accuracy
151
+
152
+ `estimateTokens` is a heuristic (the plugin ships no tokenizer), but its constants are no longer guesses: they were grid-searched against a real BPE tokenizer over 22 samples (English prose, camelCase identifiers, JSON, Windows and Unix paths, git diffs, Chinese, mixed Chinese/English, code blocks, logs, pure punctuation, hex/UUID, table rows, single glyphs, whitespace), scoring on a **weighted fit + holdout** objective to avoid overfitting.
153
+
154
+ Mean absolute error drops from **20.5% to 10.7%** (holdout 20.5% → 14.4%), and the **direction** was corrected: the old formula over-estimated pure English by **+37%** and Unix paths by **+44%**, and since both layers use this value in a ratio, it was tightening both gates. The new estimate is essentially unbiased (−0.3%).
155
+
156
+ `npm run check` asserts accuracy against a **holdout set** (5 samples that took no part in the fit, with reference lengths measured from the real tokenizer; current MAE 5.5%): a hard `MAE ≤ 15%` bound plus a directional assertion that English prose must not be over-estimated. The holdout is what gives this assertion teeth — an assertion built on the calibration data itself is a tautology that can only catch "someone hand-edited the constants", never "the constants overfit the fitting set"; with a real holdout, pushing `wordSlope` to 0.9 jumps the MAE to 38.7% and fails immediately.
157
+
158
+ ### Pressure-gate failure direction
159
+
160
+ Both layers fail **closed** and in the same direction: **when the threshold itself cannot be computed** (a `ratio` soft limit with an unresolvable context window), **neither layer acts**.
161
+
162
+ The old behaviour was asymmetric — layer 2 skipped when it could not resolve a threshold, while layer 1 simply **fell through and proceeded**; more subtly, a missing meter left `used` at `0`, so `0 < threshold` was always true and judging ran **every single round**, i.e. the gate did not exist. For a gate whose purpose is to avoid spending Jev calls, "if we cannot tell, do not spend" is the safe direction.
163
+
164
+ **The boundary**: failing closed justifies declining to spend, but it must not turn into silently switching the feature off. When the soft limit is an **absolute token count** (`softLimit: 3000`), the threshold comes straight from `limit.value` and **the meter is irrelevant** — so if the meter is missing or throws, the gate is simply left **un-armed for that pass** (reported as `压力门本次不设防` at `warn` level) and judging proceeds. `smoke_apply.mjs` pins both directions.
165
+
166
+ Note that `tokenMeter` is a **host-provided** service; if your host does not expose it, configure `softLimit` as an absolute token count (or set `judgeOn: 'always'` / `compactOn: 'always'`) rather than relying on ratio-based pressure gating.
167
+
168
+ ### Host compaction threshold vs. `softLimit`
169
+
170
+ Layer 1 does not schedule `pruneSession` itself. The host's `compaction-basic` bundle calls it when the host reaches its own pressure threshold (`thresholdRatio`, 0.8 by default) or on context overflow. This plugin's `softLimit` controls when Jev judging starts and how large the trimming budget is; it does not replace the host threshold.
171
+
172
+ In pressure mode, a layer-1 trim therefore needs both conditions:
173
+
174
+ ```text
175
+ host calls pruneSession
176
+ AND
177
+ used tokens exceed softLimit (so the pressure-gap budget is greater than zero)
178
+ ```
179
+
180
+ Keep `softLimit` at or below the host's `thresholdRatio` unless the delayed behaviour is intentional. For example, with host `thresholdRatio: 0.8` and plugin `softLimit: 90%`, host calls between 80% and 90% produce a zero plugin budget; trimming starts only after usage reaches 90%. With the default `softLimit: 55%`, judging is ready before the host's normal 80% compaction call.
181
+
182
+ For `@deepseek-ai/dsh-llm-deepseek@0.1.5-rc.2`, configure a smaller context window on the matching model entry:
183
+
184
+ ```yaml
185
+ - id: llm-deepseek
186
+ config:
187
+ models:
188
+ - id: deepseek-flash
189
+ contextWindow: 10000
190
+ ```
191
+
192
+ Setting only `defaultContextWindow` does not override catalog models that already carry their own `contextWindow`; the model entry wins. If the plugin cannot resolve the effective window, ratio-based gates stop and report the reason instead of guessing.
193
+
194
+ ### Layer 1: pressure-quantile trimming
195
+
196
+ The layer-1 decision used to be a bare fixed threshold: `keep = P(keep) ≥ 0.5`. Live-host measurement broke that assumption: **every judged candidate scored below 0.5** (42/42 in a 132k-token session, 5/5 in a short one; median ≈ 0.13–0.17). Jev's probabilities live in a narrow band — the exact trap layer 2 had already escaped by switching to relative quantiles, except nobody applied the lesson to layer 1. Under the fixed threshold, the first layer's real-world behaviour was *"trim everything that was judged"*, including results the session still needed.
197
+
198
+ `budget` mode (the default) decouples the two questions:
199
+
200
+ - **How much to trim** comes from the **pressure-gap ratio**: `ratio = (used − threshold) / window`, computed automatically each judge pass (the fraction of the context window that is over the soft limit); the budget is `ratio × the pool's total character savings`. Zero gap ⇒ trim nothing; the closer to the ceiling, the more is trimmed.
201
+ - **Which results** comes from Jev: candidates are sorted by `P(keep)` ascending and trimmed until the budget is met. Results with `P(keep) ≥ keepThreshold` (0.5) are a protection ceiling and never enter the pool; candidates whose savings are already counted stop there — the rest are recorded as `预算已用尽` (`keptByBudget`) rather than silently kept or trimmed.
202
+ - With a small population (< `minCandidatesForBudget`), the mode **degrades** to an absolute floor (`keepFloorThreshold`, 0.2) instead of inventing a ranking from 2–3 samples — the same degraded-mode shape layer 2 uses.
203
+
204
+ The old behaviour stays available as `keepMode: 'absolute'`.
205
+
206
+ ### Result excerpts (giving the judge eyes)
207
+
208
+ The judge's state used to describe every tool result as `ok, 16489 chars (内容省略)` — the judge knew *that something big existed* but not *what was in it*. Blind judging plus a fixed threshold degenerates into "trim whatever is large".
209
+
210
+ With `resultExcerptChars` (default `240`), each result line in the state carries a bounded excerpt. Lines are picked by **informativeness**, not position, because the two naive rules both failed a real-session A/B:
211
+
212
+ 1. error/evidence-pattern lines (the obvious candidate), plus
213
+ 2. **salient lines**: constant identifiers (`THRESHOLD_DISCOUNT_PCT`, `E2001_BASE_IMAGE`), assignments/keys (`timeout = 4800`), file paths — the critical config line in a 20 KB module is neither at the head nor an error, and rule 1 alone missed it (measured: judging outcomes identical to no excerpt at all), plus
214
+ 3. a middle-line fallback for pure-prose results (the middle is exactly what "cut the middle" loses).
215
+
216
+ The excerpt is hard-bounded per result and counted against the state budget, so it cannot blow up the request size. One trade-off to know: excerpts **share** the fixed `maxStateTokens` budget with history lines — at ~70 tokens per excerpted result, a 100-result session spends ~28% of the default 25k-token budget on excerpts, and the squeeze logic compensates by dropping more history lines. If you run very long sessions, raise `maxStateTokens` (Jev's ceiling is 32k) or lower `resultExcerptChars` rather than disabling excerpts entirely. Note the interaction with `budget` mode: excerpts shift *probabilities*; only the ranking-based decision converts better information into *different trimming*. Under a fixed threshold both A/B arms behaved identically — the excerpt's value presupposes the ranking rule.
217
+
218
+ ### Judgement observability (heartbeat)
219
+
220
+ The heartbeat now records **decision evidence**, not just counters — because both the 0.5-threshold failure above and the upstream `#25–#29` regressions were invisible in a stats-only heartbeat:
221
+
222
+ | Field | Content |
223
+ |---|---|
224
+ | `keep` | Layer-1 decision rule as configured (mode, ceiling, floor, budget parameters) |
225
+ | `gate` | Last pressure-gate evaluation: `used` / resolved window / threshold / `skip` + reason / candidate count |
226
+ | `probSummary` | Distribution of `P(keep)`: p10–p90, mean, counts above/below the ceiling |
227
+ | `probSamples` | Last 200 raw probabilities (for histograms) |
228
+ | `lastJudgePass.rows` | Per-candidate detail: seq, tool, chars, `prob`, `effectProb` |
229
+ | `stats.preStepEvents` / `stats.judgePassSkipped` + `lastJudgeSkipReason` | Distinguishes "the event never fired" / "no candidates" / "gate skipped" — three failures that used to look identical from outside |
230
+
231
+ One structural fix came out of this: the heartbeat is merge-written, but the pre-step hook itself never called `writeHeartbeat` — so gate-skipped passes left the file frozen at the boot snapshot (`bootedAt == now`) and every skipped path was unobservable. The hook now persists after every step.
232
+
233
+ ### Out-of-range configuration
234
+
235
+ Every numeric option has a valid range, and out-of-range values are **never** passed through to the runtime:
236
+
237
+ - **Through the `Config` schema** (the host's normal load path) → a `ValidationError` is thrown. A loud refusal.
238
+ - **Without schema normalization** (a config object injected directly by `cordis.patch.yml`, or the `PLUGIN_CFG` built by the smoke test) → `resolveConfig` falls the value back to its **default** (not to the nearest bound, because "how far off was it" isn't interpretable), and records a `configWarnings` entry in the status report and heartbeat.
239
+
240
+ These are the real consequences, all of which used to happen **silently**:
241
+
242
+ | Setting | Behavior before the fix |
243
+ |---|---|
244
+ | `preserveRecent = -5` | `lastAllowed` grew instead of shrinking → **recent-node protection completely defeated** (in-flight tool calls could be touched) |
245
+ | `maxStepTextChars = -1` | every step judged "text too long" → **layer 2 permanently and silently dead** |
246
+ | `compactMinChars = -100` | the gate ceased to exist |
247
+ | `receiptMaxRatio = 5` | a receipt 5× larger than the original was allowed through (safety gate defeated) |
248
+ | `keepThreshold = 2` | layer 1 pruned everything (`prob >= 2` is never true) |
249
+
250
+ Note that `0` is a **legal** value for most keys (`headChars = 0` keeps no head; `preserveRecent = 0` protects nothing) — it is not treated as "unset".
251
+
252
+ ## In-session usage
253
+
254
+ | Entry point | Purpose |
255
+ |---|---|
256
+ | `/jev`, `jev_prune_status` | Ledger for both layers: cached judgments, cumulative savings, takeover state, tool-name index |
257
+ | `jev_prune_now` | Force one layer-1 trimming pass |
258
+ | `jev_compact_now` (supports `dryRun`) | Force one layer-2 receipt compaction and list every gate's exclusion counts; in `dryRun` it also prints the full receipt text |
259
+ | `jev_restore` | Safety valve: fetch back the original text that a checkpoint moved out (read-only) |
260
+ | `jev_probe_shapes` | Print the real event shapes and the resolved tool names, for adapting to a different DSH version |
261
+
262
+ In normal operation both layers are driven automatically by context pressure; no manual step is needed.
263
+
264
+ ## Design notes
265
+
266
+ - **Reads probabilities, never generated text**: judgments always read `answers[id].noul` from the structured response
267
+ - **The state carries the task goal**: "is this still useful" really means "useful relative to the goal", so the state header carries the most recent user instruction
268
+ - **Structure by code, semantics by the model**: write-type calls and the recent window are guaranteed by hard rules, never entrusted to a probability
269
+ - **The shadow-price protocol is aligned verbatim** with DSH's `compaction/prune` + `surfaceOp: replace`, so pure consumers can reuse the same token accounting
270
+ - **Slices by Unicode code point**, never splitting a surrogate pair; token estimation uses a per-word correction algorithm that works for mixed CJK/Latin text
271
+ - **Hook ordering is load-bearing**: the judge hook is `prepend`ed (`ctx.on(..., true)`) so it runs **before the base bundle's `compaction-basic`**. That package is the *only* caller of `pruner.pruneSession` (both call sites live in it — `:888` for context-overflow, `:902` for pressure), so it is also the only place layer 1's verdicts get consumed. Registered without `prepend`, pruning would read the *previous* round's verdicts and every fresh result would fall back to the size rules — layer 1 silently inert, no error anywhere. `smoke_apply.mjs` block **M** pins this by observing the judge counter at the moment `pruneSession` is called.
272
+
273
+ ## Testing and layout
274
+
275
+ ```bash
276
+ npm ci
277
+ npm run check
278
+ npm run smoke
279
+ npm run coverage
280
+ npm run demo
281
+ node scripts/verify_layout.mjs
282
+ ```
283
+
284
+ The package includes `src/`, `test/`, `scripts/`, `docs/`, `demo/` and `assets/`. CI copies the entry and `src/` together, preserving relative imports. Public imports `dsh-jev-prune`, `dsh-jev-prune/index.js` and the four legacy helper subpaths still resolve. Direct helper commands now use `scripts/`; direct test commands now use `test/`.
285
+
286
+ Use the committed lockfile for local verification. Fresh npm dependency resolution against the upstream RC peers can select partially published rc.3 packages and fail with ERESOLVE; that upstream dependency-tree issue is independent of this layout migration. The package archive itself was verified with the locked dependency fixture, including helper tests, smoke checks and manual profile wiring.
287
+
288
+ ```text
289
+ index.js plugin entry
290
+ src/ runtime helper modules
291
+ test/ checks and simulated host
292
+ scripts/ profile and session utilities
293
+ docs/ architecture and reference
294
+ demo/ reproducible demonstration
295
+ assets/ images and recordings
296
+ examples/ profile snippets
297
+ cordis.patch.yml installation contract
298
+ ```
299
+
300
+ ## Privacy
301
+
302
+ The plugin sends session history text — including file paths, code snippets and command output — to the TypeSafe API for judgment. Assess this yourself before working on sensitive code. If data must not leave the machine, swap the judgment backend for a self-hosted model: the judgment and the compaction machinery are decoupled, and the replacement points are `jev.js` and `state.js`.
303
+
304
+ ## License
305
+
306
+ [MIT](../LICENSE)
307
+
308
+ ---
309
+
310
+ **English** · [简体中文](../README_zh.md)
@@ -0,0 +1,16 @@
1
+ ## Measured results
2
+
3
+ Numbers below come from a three-arm measurement of a 37-step "read every file in a directory, strictly one at a time" long task over `semver@6e05b76`: **A** = vanilla DSH (baseline), **B** = the plugin with `compactReceipts: false` (layer 1 only), **C** = the plugin with defaults (both layers). Every number is taken from the `usage` object the API actually returned in the DSH session logs — no estimator is involved in scoring.
4
+
5
+ **Layer 1 acts on every long run; the baseline never does.** Layer 1 trimmed stale results in every plugin-arm run (4–7 nodes per long run; arm A trims 0 by definition), and answers on the short/medium/long tasks were correct in every arm.
6
+
7
+ **Long tasks run on a fraction of the full-price input.** Uncached ("full-price") input tokens per completed 37-step run: baseline 75,781–87,872 vs full plugin 25,516–45,071 — roughly a third of the baseline at equal task scope. Report token classes separately: the cache-hit share of input is high (78–94%) and hit/miss prices differ by ~50×, so total-token comparisons mislead by design.
8
+
9
+ **Fewer destructive compactions — some of them receipts instead of summaries.** Per completed long run, DSH's own model-written summary compactions dropped from 31–40 (baseline) to 15–23 with the full plugin, of which 6–8 were replaced by deterministic receipts rendered by code. Layer 2 stays deliberately silent on short/medium tasks — its gates require a contiguous run of eligible read-only steps — so on those workloads the effect comes from layer 1. Even when layer 2 fires, DSH-initiated compactions still exist and fall back to model summaries: the plugin reduces them, it does not eliminate them.
10
+
11
+ **A failure mode we found and fixed.** Receipts replace a whole "call + result" step, including the assistant message that carried it; in one long run the model's intermediate notes were erased step by step (12 of 13) and it stopped issuing tool calls mid-task. Receipts now carry each step's **assistant-visible text verbatim** (`text` blocks only, `reasoning` drafts excluded, zero model generation; bounded by `receiptTextChars`, default 400, `0` disables). After the fix, the same long task completed in every run with correct answers.
12
+
13
+
14
+ ## Evidence boundary
15
+
16
+ These figures were reported in the upstream README before this documentation change. Raw session logs, repetition counts, full scoring criteria, model/judge versions and judge spend are not committed. Treat the ranges as workload-specific observations, not a general cost-saving guarantee. The deterministic demo is a behavior check and is not the measured experiment.
@@ -0,0 +1,44 @@
1
+ # Source review and release readiness
2
+
3
+ Reviewed starting at `d9e85fd0734e4a542b1c07ed6c1c0125dc9e52a8` on 2026-10-05. The unchanged baseline passes `npm run check` and all 113 smoke checks. Additional diagnostic probes expose gaps; this documentation/demo PR does not fix runtime behavior.
4
+
5
+ ## Open PR #45 and issue #35
6
+
7
+ [PR #45](https://github.com/yangyu666/dsh-jev-prune/pull/45) adds a changelog and npm-first installation. [Issue #35](https://github.com/yangyu666/dsh-jev-prune/issues/35) requests npm publication, a tag and a release.
8
+
9
+ - **P1: unpublished distribution presented as available.** The current npm identity check returns `ENEEDAUTH`; the repository has no version tags. Do not advertise bare-package installation until registry existence and a clean installation are verified. A prepared changelog must not claim a release has already happened.
10
+ - **P1: changelog concurrency guarantee is unsupported.** Claim-once consumption is implemented, but transaction ownership is not established. See reproduction below.
11
+ - **P2: empty quantile intersection is described as floor degradation.** `computeEligibleSeqs` uses the floor for a population below the minimum; a large population with an empty intersection does not automatically enter that branch. Release notes should describe the actual condition.
12
+ - Issue #35 should remain open until the intended published artifacts exist. A preparation PR alone does not complete it. This PR supplies release instructions and a prepared changelog, with no automatic issue-closing keyword.
13
+
14
+ ## Runtime findings
15
+
16
+ ### P1 — external summary can consume another transaction's receipt
17
+
18
+ `index.js` stores a pending receipt by session and compares its fence to one apply-instance-wide `activeFence`. External host summaries do not update that fence. In smoke block C3, the alleged competing call actually executes **after** the plugin summary has consumed the receipt.
19
+
20
+ Reproduction: before calling `innerCompact` inside C3's `compactRegion` wrapper, await a microtask and invoke `compaction.summarize({messages:['elsewhere']}, agent, signal)`. Then call `innerCompact`. The external call returns `provider=jev-receipt`; the intended compaction returns `provider=fake`. Both relevant assertions fail. A later fence check cannot undo incorrect injection.
21
+
22
+ Suggested fix: bind pending receipts to a transaction identity propagated through the actual call chain, or serialize every compaction entry path. Validate both arrival orders and multiple sessions, not only a second consumption attempt.
23
+
24
+ ### P1 — result judgments survive changes in task goal
25
+
26
+ `judgePass` identifies fresh candidates by missing probability axes, and reads the current goal only after the fresh-candidate early return. Reproduction: judge three results, append a new user goal requiring the earlier Read result, and run pre-step again. Judge requests remain `1 -> 1`.
27
+
28
+ Suggested fix: invalidate result-axis judgments when the relevant goal/context version changes. Effect judgments may be cached separately. Document what counts as a goal change.
29
+
30
+ ### P2 — recovery API omits replacement events
31
+
32
+ `findCompactionRecord` searches `compaction/summary` and checkpoint references only. Passing a layer-1 replacement seq to `jev_restore` returns no matching checkpoint. Mixed-batch replacements use the same `compaction/prune` protocol and lack a direct recovery path too. Even region retrieval caps each event at 4,000 UTF-16 units and does not follow chains back through earlier trims.
33
+
34
+ Suggested fix: resolve replacement source chains and prune records, provide bounded pagination, and distinguish current hidden-region content from the original pre-trim body. Preserve read-only semantics.
35
+
36
+ ### P2 — small-population fallback ignores zero budget
37
+
38
+ `planTrims` enters its floor branch before testing `budget > 0`. One candidate with `prob=0.1`, `gain=5000`, `chars=6000` and `pressureRatio=0` produces `mode=floor`, `selected=[1]`, `spent=5000`.
39
+
40
+ Suggested fix: define whether zero pressure disables both normal and floor modes, then place the guard accordingly. Test `alwaysTrimRatio=0`, unknown window and zero gap separately.
41
+
42
+ ## Release boundary
43
+
44
+ The deterministic demo exercises one serial path with fixed scores. It is not evidence that these issues are resolved. Runtime findings should be addressed before promoting a production-ready release. An explicitly experimental prerelease can document these limitations, but should not advertise concurrency or complete recovery guarantees.
@@ -0,0 +1,7 @@
1
+ # Configuration examples
2
+
3
+ `minimal.yml` is a Cordis patch fragment for the tested DSH profile. Install the plugin first, then apply or merge the fragment into the profile configuration.
4
+
5
+ Keep the TypeSafe key in the `TYPESAFE_API_KEY` environment variable. Do not commit it to the configuration file.
6
+
7
+ For a smaller DeepSeek context window, configure `models[].contextWindow` on the provider's matching model entry; `defaultContextWindow` does not override catalog entries that already define a window.
@@ -0,0 +1,10 @@
1
+ - insert:
2
+ - id: jev-prune
3
+ name: dsh-jev-prune
4
+ config:
5
+ judgeOn: pressure
6
+ softLimit: "55%"
7
+ compactReceipts: true
8
+ compactOn: pressure
9
+ compactSoftLimit: "70%"
10
+ dryRun: true