@deepseek-ai/dsh-output-retention 0.1.1-rc.2 → 0.1.2-alpha.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.i18n.yaml CHANGED
@@ -2,5 +2,5 @@
2
2
  # side as of the last confirmed-consistent state. Both languages carry equal authority;
3
3
  # after editing either side, bring the other along and re-record with:
4
4
  # pnpm run verify-translation-pairing --write packages/util/output-retention/README.md
5
- README.md: e1d2c828fd4c6526cc391e7f93d7fe20df26e268
6
- README.zh.md: 51055dd33e2cce3b82a1d15d6994f9d9f37a0579
5
+ README.md: 6614f82149728b12419a4750a59772ef9e5de440
6
+ README.zh.md: 4f8d6bb235b7ee4c0b79a094308ad6c4a5913e36
package/README.md CHANGED
@@ -1,97 +1,165 @@
1
- # dsh-output-retention
1
+ ---
2
+ description: "Bounded model-facing output for tools that must cap how much context they return: item and text retainers plus a standardized omission footer."
3
+ kind: "package-library"
4
+ ---
5
+
6
+ # @deepseek-ai/dsh-output-retention
2
7
 
3
8
  English | [中文](README.zh.md)
4
9
 
5
- A dependency-light **retention** library: bounded model-facing output for tools that must cap how much context they return. A caller feeds items or text chunks into a bounded object, then gets the retained content plus exact omission metadata.
10
+ ## Summary
6
11
 
7
- The library owns **only** the mechanical question *"what did we keep, and what did we omit?"*. Tool-specific code keeps its business semantics: file grouping, line numbering, exit codes, provider error states, per-line preview truncation, spill files, and the model-facing prose. This is the boundary the [Agent Note](../../../.agents/notes/implemented/architecture/2026-07-06-tool-result-retention-library.md) draws.
12
+ `dsh-output-retention` bounds how much context a tool returns to the model: a caller feeds items or text chunks into a retainer, then gets back the retained content plus exact omission metadata. `ItemRetainer` caps an ordered list of logical units (paths, matches, sources) at a head budget; `TextRetainer` caps a byte-oriented text stream with head, tail, or head-and-tail windows and keeps UTF-8 boundaries valid at every cut. A standardized omission clause and a notice formatter give tools a consistent "results capped" footer while the tool owns the recovery guidance. The library answers only the mechanical question of what was kept and what was omitted — grouping, line numbering, spill files, and provider error states stay in the tool. It is a dependency-light library that tool packages import directly; a `cordis.yml` cannot load it.
8
13
 
9
- It is a **library, not a service or plugin**: no `ctx`, registers nothing, emits no events. The only state is per-retainer (one accumulation), never cross-call. Tool packages import it directly.
14
+ ## Table of Contents
10
15
 
11
- ## API
16
+ - [Use this package](#use-this-package)
17
+ - [Understand the implementation](#understand-the-implementation)
18
+ - [Further Exploration](#further-exploration)
19
+ - [Model Experience](#model-experience)
20
+ - [Known Limitations and Deferred Work](#known-limitations-and-deferred-work)
21
+ - [Dev Note](#dev-note)
12
22
 
13
- ```ts
14
- import {
15
- ItemRetainer, TextRetainer,
16
- describeOmitted, formatRetentionNotice,
17
- } from '@deepseek-ai/dsh-output-retention'
18
- import type {
19
- Omitted, PushDecision, RetainedItems, RetainedText,
20
- ItemRetentionStrategy, TextRetentionStrategy, RetentionNotice,
21
- } from '@deepseek-ai/dsh-output-retention'
22
- ```
23
+ -----
23
24
 
24
- | Export | Role |
25
- |---|---|
26
- | `ItemRetainer<T>` | Bounds ordered logical units (paths, grep matches, sources). `head` only. `push()` → `PushDecision`; `finish()` → `RetainedItems<T>`. |
27
- | `TextRetainer` | Bounds a byte-oriented text stream. `head` / `tail` / `headTail`, UTF-8 boundaries preserved at `finish()`. `push()` → `PushDecision`; `finish()` → `RetainedText`. |
28
- | `describeOmitted(omitted, unit)` | Standardized omission clause (`exact` prints a count; `unknown` does not). |
29
- | `formatRetentionNotice(notice, recovery)` | Joins the standardized omission clause with the tool's own recovery guidance. |
30
- | `Omitted` | `none` / `exact` / `unknown` — how much was omitted. |
31
- | `PushDecision` | `{ kept, truncated }` — the per-push retention result. |
25
+ <a id="use-this-package"></a>
26
+ ## Use this package
32
27
 
33
- ## Resource Modes
28
+ Use a retainer wherever a tool must cap how much of its result reaches the model, and report honestly what was dropped. Choose `ItemRetainer` for ordered logical units and `TextRetainer` for byte-oriented streams.
34
29
 
35
- The two retainers are separate names, not one generic collector, because they differ in **resource model**.
30
+ ### Bounding a list of items
36
31
 
37
- - **`ItemRetainer` bounds ordered logical units.** A search tool can collect a full result set for spill-file recovery while retaining only the first `maxItems` for the model-facing preview. The omission count is exact because the caller keeps feeding every observed item.
38
- - **`TextRetainer` bounds byte-oriented text.** `head`, `tail`, and `headTail` preserve UTF-8 boundaries at `finish()`; `headTail` is the shape `dsh-spill-policy` uses to build a bounded preview around a spill-file notice.
39
-
40
- ## `truncated` is a budget fact, never "incomplete"
32
+ ```ts
33
+ import { ItemRetainer } from '@deepseek-ai/dsh-output-retention'
41
34
 
42
- `truncated` means *the retainer omitted otherwise-available content because of a budget*. It does **not** mean the upstream was incomplete. Permission failures, skipped binary files, provider partial failures, unreadable candidates, and invalid UTF-8 stay in tool-domain fields — never folded into `truncated`. Conflating the two is the bug this library's naming most invites; keep them separate.
35
+ declare const globMaxResults: number
36
+ declare const candidates: AsyncIterable<{ path: string }>
37
+ const retainer = new ItemRetainer<{ path: string }>({ kind: 'head', maxItems: globMaxResults })
38
+ for await (const entry of candidates) {
39
+ retainer.push(entry) // keep draining past the cap for an exact count
40
+ }
41
+ const { items, truncated, omitted } = retainer.finish()
42
+ ```
43
43
 
44
- ## Bytes, not characters
44
+ `push()` reports per item whether it was kept, and `finish()` returns the retained items plus `omitted` — an exact count when the caller kept feeding every observed unit. A search tool can collect the full result set for a spill file while retaining only the first page for the model.
45
45
 
46
- Text caps and `omittedBytes` count **bytes**, for process/body safety (a child's pipe and an HTTP body are byte streams). A chunk that straddles a codepoint is handled: `finish()` trims a partial codepoint at each cut so the returned text never introduces a replacement char at the boundary, and the two sides are decoded separately so a codepoint is never reconstructed across the omitted middle. Character- or line-level preview budgets are a separate, tool-owned concern.
46
+ ### Bounding a text stream
47
47
 
48
- ## Tool mappings
48
+ ```text
49
+ import { TextRetainer } from '@deepseek-ai/dsh-output-retention'
49
50
 
50
- Current retention consumers use these mappings:
51
+ const out = new TextRetainer({ kind: 'headTail', headBytes: headCap, tailBytes: tailCap })
52
+ child.stdout.on('data', (chunk: Buffer) => { out.push(chunk) })
53
+ const { text, omittedBytes } = out.finish()
54
+ ```
51
55
 
52
- | Tool | Retainer & strategy | Notes |
53
- |---|---|---|
54
- | `glob` | `ItemRetainer<FsGlobEntry>`, `head` | Collect the full sorted path list for a spill file while retaining the first page inline. Path mapping, skipped candidates, and `incomplete` stay outside. |
55
- | `grep` | `ItemRetainer<FlatGrepMatch>`, `head` | Collect matches for a spill file while retaining the first page inline. Per-match preview truncation, grouping, sorting, and `incomplete` stay outside. |
56
- | `bash` | `TextRetainer`, `tail` or `headTail` | Executor still owns spill files, exit status, signal, timeout, and background jobs. |
57
- | `web_fetch` | `TextRetainer`, `head` or `headTail` | Provider/resource caps stay provider facts; the retainer supplies only retained text and omission metadata. |
58
- | `web_search` | `ItemRetainer<WebSearchSource>`, `head` | Standardizes the "sources capped" notice when providers return more sources than the model-facing result should include. |
56
+ `head`, `tail`, and `headTail` count bytes, not characters or lines: a child's pipe and an HTTP body are byte streams. `finish()` trims a partial codepoint at each cut, so the returned text never carries a replacement character introduced by the cut, and a codepoint is never reconstructed across the omitted middle.
59
57
 
60
- `read` remains outside this generic library. Its `read-render` helper owns a file-specific pagination contract — `offset`/`limit`, line numbers, `totalLines`, offset-out-of-range errors, per-line preview truncation, and a byte cap over the selected window — which is a line-window renderer. A single `Omitted` count cannot represent both sides of that window.
58
+ ### Building the omission footer
61
59
 
62
- ## Usage shape
60
+ ```ts
61
+ import { formatRetentionNotice } from '@deepseek-ai/dsh-output-retention'
63
62
 
64
- ```ts ignore-check
65
- // glob: keep the first page inline while still collecting the full list for spill.
66
- const retainer = new ItemRetainer<FsGlobEntry>({ kind: 'head', maxItems: globMaxResults })
67
- const allEntries: FsGlobEntry[] = []
68
- for await (const entry of candidates) {
69
- allEntries.push(entry)
70
- retainer.push(entry)
71
- }
72
- const { items, truncated, omitted } = retainer.finish()
63
+ declare const grepMaxMatches: number
64
+ declare const items: { length: number }
65
+ import type { Omitted } from '@deepseek-ai/dsh-output-retention'
73
66
 
74
- // bash: keep a head + tail, read to process exit.
75
- const out = new TextRetainer({ kind: 'headTail', headBytes: headCap, tailBytes: tailCap })
76
- child.stdout.on('data', (chunk: Buffer) => { out.push(chunk) })
77
- const { text, omittedBytes } = out.finish()
67
+ declare const omitted: Omitted
78
68
 
79
- // A footer: the library standardizes the omission clause; the tool owns recovery words.
80
69
  const footer = formatRetentionNotice(
81
70
  { scope: 'grep', strategy: 'head', unit: 'items', limit: grepMaxMatches, kept: items.length, omitted },
82
71
  ({ kept }) => `Results capped at ${kept}. Narrow the pattern, path, or include to see more.`,
83
72
  )
84
73
  ```
85
74
 
75
+ The library standardizes the omission clause (`Omitted 3 items.`) and joins it with the tool's own recovery guidance; only the tool knows the recovery action, so the tool supplies those words.
76
+
77
+ ### What `truncated` means
78
+
79
+ `truncated` is a budget fact: the retainer omitted otherwise-available content because of a cap. It never means the upstream was incomplete — permission failures, skipped binary files, provider partial failures, and unreadable candidates stay in tool-domain fields, never folded into `truncated`.
80
+
81
+ ### How the current tools use it
82
+
83
+ | Tool | Retainer | What the tool still owns |
84
+ |---|---|---|
85
+ | `glob` | `ItemRetainer`, `head` | Spill-file collection, path mapping, skipped candidates, `incomplete` |
86
+ | `grep` | `ItemRetainer`, `head` | Spill-file collection, per-match preview truncation, grouping, sorting |
87
+ | `bash` | `TextRetainer`, `tail` or `headTail` | Spill files, exit status, signal, timeout, background jobs |
88
+ | `web_fetch` | `TextRetainer`, `head` or `headTail` | Provider and resource caps, error states |
89
+ | `web_search` | `ItemRetainer`, `head` | The "sources capped" notice wording and provider facts |
90
+
91
+ `read` stays outside this library: its line-window pagination (`offset`/`limit`, line numbers, `totalLines`) is a file-specific renderer that a single omission count cannot represent.
92
+
93
+ -----
94
+
95
+ <a id="understand-the-implementation"></a>
96
+ ## Understand the implementation
97
+
98
+ <details>
99
+ <summary>Implementation internals — click to expand</summary>
100
+
101
+ The library is built on one separation: it owns the mechanical question of what was kept and what was omitted; tool packages own every business meaning.
102
+
103
+ ### Source map
104
+
105
+ | File | Role |
106
+ |---|---|
107
+ | [`src/index.ts`](src/index.ts) | `ItemRetainer`, `TextRetainer`, `describeOmitted`, and `formatRetentionNotice` |
108
+ | [`src/invariant.ts`](src/invariant.ts) | Invariant companion (no runtime invariant; the retention algebra is exercised by unit tests) |
109
+
110
+ ### Two retainers, two resource models
111
+
112
+ `ItemRetainer` bounds ordered logical units and keeps only the first `maxItems`; the caller keeps pushing every observed unit so the omission count is exact. `TextRetainer` bounds bytes with one shared prefix/suffix accumulator: `head` is prefix-only, `tail` is suffix-only, `headTail` is both, and the accumulator holds at most `headBytes + tailBytes + one chunk` in memory, so a large stream does not accumulate unbounded.
113
+
114
+ ### How the budget facts stay honest
115
+
116
+ `push()` returns `kept` (this unit or chunk fully retained) and `truncated` (anything dropped yet). `finish()` reports omission against the bytes actually returned, so a UTF-8 boundary trim that drops partial-codepoint bytes is counted too — a notice built from the budget alone would overstate the retained text. `describeOmitted` prints a count only for `exact`; `unknown` prints no count because the caller provided none.
117
+
118
+ ### The read-render exclusion
119
+
120
+ `read`'s `offset`/`limit` pagination is a line-window renderer with its own byte cap over the selected window; a single `Omitted` value cannot represent both sides of that window, so it stays out of this library.
121
+
122
+ </details>
123
+
124
+ -----
125
+
126
+ <a id="further-exploration"></a>
127
+ ## Further Exploration
128
+
129
+ Read these pages when you need the consumers or the boundary decision behind the library.
130
+
131
+ - [Tool-result retention library Agent Note](../../../.agents/notes/implemented/architecture/2026-07-06-tool-result-retention-library.md) — the boundary the library draws around tool semantics.
132
+ - [Spill policy](../../spill/spill-policy/README.md) — composes `TextRetainer` for a bounded preview around a spill-file notice.
133
+ - [Spill subsystem](../../../docs/subsystems/spill.md) — the spill vocabulary this library's preview mechanics serve.
134
+ - [File search tool](../../fs/tool-fs-search/README.md) — an `ItemRetainer` consumer collecting full results for spill.
135
+
136
+ -----
137
+
138
+ <a id="model-experience"></a>
86
139
  ## Model Experience
87
140
 
88
- Indirectly, through tool consumers that render retained content and omission metadata.
141
+ Indirectly, through the retention consumers that render retained content and omission metadata.
89
142
 
90
143
  #### KV Cache effect
91
144
 
92
- No direct invalidation; the named consumer owns any request-prefix changes.
145
+ No direct invalidation; the retention consumers own any request-prefix changes.
93
146
 
94
147
  ## Known Limitations and Deferred Work
95
148
 
149
+ <a id="known-limitations-and-deferred-work"></a>
150
+
151
+
152
+ These limits define what the retainers deliberately do not cover. They are current package constraints, not a task backlog.
153
+
96
154
  - **Item retention supports `head` only** — tail, head/tail, pagination, grouping, and provider-completeness semantics remain tool-owned.
97
- - **Text retention is byte-oriented** — line and character windows such as `read` pagination require a separate renderer, and a cut may discard partial UTF-8 boundary bytes to keep returned text valid.
155
+ - **Text retention is byte-oriented** — line and character windows such as `read` pagination require a separate renderer, and a cut may discard partial UTF-8 boundary bytes to keep the returned text valid.
156
+
157
+ <a id="dev-note"></a>
158
+ ### Dev Note
159
+
160
+ <details>
161
+ <summary>Working context for maintainers — click to expand</summary>
162
+
163
+ None.
164
+
165
+ </details>
package/README.zh.md CHANGED
@@ -1,97 +1,165 @@
1
- # dsh-output-retention
1
+ ---
2
+ description: "为必须限制返回上下文量的工具提供有界的面向模型输出:项与文本 retainer,以及标准化的省略页脚。"
3
+ kind: "package-library"
4
+ ---
5
+
6
+ # @deepseek-ai/dsh-output-retention
2
7
 
3
8
  [English](README.md) | 中文
4
9
 
5
- 一个轻依赖的**保留**库:为必须限制返回上下文量的工具提供有界的面向模型输出。调用方将项或文本分片送入有界对象,然后取回保留的内容和精确的省略元数据。
10
+ ## 概述
6
11
 
7
- 该库**只**负责这个机制问题:*「我们保留了什么,又省略了什么?」*。工具专用代码保留其业务语义:文件分组、行号、退出码、提供方错误状态、每行预览截断、spill 文件以及面向模型的文案。这就是 [Agent Note](../../../.agents/notes/implemented/architecture/2026-07-06-tool-result-retention-library.zh.md) 划定的边界。
12
+ `dsh-output-retention` 限制工具返回给模型的上下文量:调用方把项或文本分片送入 retainer,然后取回保留的内容与精确的省略元数据。`ItemRetainer` 以头部预算限制有序逻辑单元列表(路径、匹配项、来源);`TextRetainer` 以 head、tail 或 head+tail 窗口限制面向字节的文本流,并在每个切割处保持 UTF-8 边界有效。标准化的省略子句与通知格式化器让工具获得一致的「结果已达上限」页脚,而恢复指引由工具自己提供。该库只回答「保留了什么、省略了什么」这个机制问题——分组、行号、spill 文件与提供方错误状态都留在工具侧。它是轻依赖库,由工具包直接导入;`cordis.yml` 无法加载它。
8
13
 
9
- 它是**库,而非服务或插件**:没有 `ctx`,不注册任何内容,不发出任何事件。状态只存在于每个 retainer(一次累积)中,绝不跨调用。工具包直接导入它。
14
+ ## 目录
10
15
 
11
- ## 对外接口
16
+ - [使用本包](#use-this-package)
17
+ - [理解实现](#understand-the-implementation)
18
+ - [进一步探索](#further-exploration)
19
+ - [模型体验](#model-experience)
20
+ - [已知限制与延期工作](#known-limitations-and-deferred-work)
21
+ - [开发备注](#dev-note)
12
22
 
13
- ```ts
14
- import {
15
- ItemRetainer, TextRetainer,
16
- describeOmitted, formatRetentionNotice,
17
- } from '@deepseek-ai/dsh-output-retention'
18
- import type {
19
- Omitted, PushDecision, RetainedItems, RetainedText,
20
- ItemRetentionStrategy, TextRetentionStrategy, RetentionNotice,
21
- } from '@deepseek-ai/dsh-output-retention'
22
- ```
23
+ -----
23
24
 
24
- | 导出项 | 职责 |
25
- |---|---|
26
- | `ItemRetainer<T>` | 限制有序逻辑单元(路径、grep 匹配项、来源)。只支持 `head`。`push()` → `PushDecision`;`finish()` → `RetainedItems<T>`。 |
27
- | `TextRetainer` | 限制面向字节的文本流。`head` / `tail` / `headTail`,并在 `finish()` 时保留 UTF-8 边界。`push()` → `PushDecision`;`finish()` → `RetainedText`。 |
28
- | `describeOmitted(omitted, unit)` | 标准化的省略子句(`exact` 输出数量;`unknown` 不输出)。 |
29
- | `formatRetentionNotice(notice, recovery)` | 将标准化的省略子句与工具自有的恢复指引连接起来。 |
30
- | `Omitted` | `none` / `exact` / `unknown`:省略了多少内容。 |
31
- | `PushDecision` | `{ kept, truncated }`:每次 push 的保留结果。 |
25
+ <a id="use-this-package"></a>
26
+ ## 使用本包
32
27
 
33
- ## 资源模式
28
+ 凡是工具必须限制其结果到达模型的数量、并如实报告丢弃内容的地方,都使用 retainer。有序逻辑单元选 `ItemRetainer`,面向字节的流选 `TextRetainer`。
34
29
 
35
- 两个 retainer 使用独立名称,而不是同一个通用收集器,因为它们的**资源模型**不同。
30
+ ### 限制项列表
36
31
 
37
- - **`ItemRetainer` 限制有序逻辑单元**。搜索工具可收集完整结果集用于 spill 文件恢复,同时只为面向模型的预览保留前 `maxItems` 项。因为调用方会继续送入每个已观察到的项,所以省略数量是精确的。
38
- - **`TextRetainer` 限制面向字节的文本**。`head`、`tail` 和 `headTail` 在 `finish()` 时保留 UTF-8 边界;`headTail` 是 `dsh-spill-policy` 用于围绕 spill 文件通知构建有界预览的形态。
39
-
40
- ## `truncated` 是预算事实,绝不表示「不完整」
32
+ ```ts
33
+ import { ItemRetainer } from '@deepseek-ai/dsh-output-retention'
41
34
 
42
- `truncated` 表示*因为预算限制,retainer 省略了本可获得的内容*。它**不**表示上游不完整。权限失败、跳过二进制文件、提供方部分失败、不可读候选项和无效 UTF-8 保留在工具领域字段中,绝不合并到 `truncated`。将两者混为一谈是该库命名最容易诱发的缺陷;务必保持分离。
35
+ declare const globMaxResults: number
36
+ declare const candidates: AsyncIterable<{ path: string }>
37
+ const retainer = new ItemRetainer<{ path: string }>({ kind: 'head', maxItems: globMaxResults })
38
+ for await (const entry of candidates) {
39
+ retainer.push(entry) // keep draining past the cap for an exact count
40
+ }
41
+ const { items, truncated, omitted } = retainer.finish()
42
+ ```
43
43
 
44
- ## 字节,而非字符
44
+ `push()` 逐项报告该项是否被保留,`finish()` 返回保留的项与 `omitted`——当调用方持续送入每个已观察单元时,这是一个精确计数。搜索工具可以收集完整结果集用于 spill 文件,同时只为模型保留第一页。
45
45
 
46
- 文本上限和 `omittedBytes` 按**字节**计数,以保证进程/正文安全(子进程管道和 HTTP 正文都是字节流)。跨越码点的分片会被正确处理:`finish()` 会修剪每个切割位置的不完整码点,使返回文本绝不在边界引入替换字符;首尾两侧会分开解码,因此绝不会跨越被省略的中间部分重建码点。按字符或行限制的预览预算属于独立的工具职责。
46
+ ### 限制文本流
47
47
 
48
- ## 工具映射
48
+ ```text
49
+ import { TextRetainer } from '@deepseek-ai/dsh-output-retention'
49
50
 
50
- 当前的保留机制消费方采用以下映射:
51
+ const out = new TextRetainer({ kind: 'headTail', headBytes: headCap, tailBytes: tailCap })
52
+ child.stdout.on('data', (chunk: Buffer) => { out.push(chunk) })
53
+ const { text, omittedBytes } = out.finish()
54
+ ```
51
55
 
52
- | 工具 | Retainer 与策略 | 说明 |
53
- |---|---|---|
54
- | `glob` | `ItemRetainer<FsGlobEntry>`,`head` | 收集完整的已排序路径列表用于 spill 文件,同时在内联位置保留第一页。路径映射、已跳过候选项和 `incomplete` 保留在外部。 |
55
- | `grep` | `ItemRetainer<FlatGrepMatch>`,`head` | 收集匹配项用于 spill 文件,同时在内联位置保留第一页。每个匹配项的预览截断、分组、排序和 `incomplete` 保留在外部。 |
56
- | `bash` | `TextRetainer`,`tail` 或 `headTail` | 执行器仍负责 spill 文件、退出状态、信号、超时和后台任务。 |
57
- | `web_fetch` | `TextRetainer`,`head` 或 `headTail` | 提供方/资源上限保留为提供方事实;retainer 只提供保留文本和省略元数据。 |
58
- | `web_search` | `ItemRetainer<WebSearchSource>`,`head` | 当提供方返回的来源超过面向模型的结果应包含的数量时,标准化「来源已达上限」通知。 |
56
+ `head`、`tail` 与 `headTail` 按字节而非字符或行计数:子进程管道与 HTTP 正文都是字节流。`finish()` 会在每个切割处修剪不完整的码点,因此返回的文本绝不会携带由切割引入的替换字符,码点也绝不会跨被省略的中间部分重建。
59
57
 
60
- `read` 仍不属于这个通用库。其 `read-render` 辅助工具负责文件专用的分页约定:`offset`/`limit`、行号、`totalLines`、偏移越界错误、每行预览截断,以及所选窗口的字节上限。该辅助工具是一个行窗口渲染器。单个 `Omitted` 数量无法表示该窗口两侧。
58
+ ### 构建省略页脚
61
59
 
62
- ## 使用形态
60
+ ```ts
61
+ import { formatRetentionNotice } from '@deepseek-ai/dsh-output-retention'
63
62
 
64
- ```ts ignore-check
65
- // glob: keep the first page inline while still collecting the full list for spill.
66
- const retainer = new ItemRetainer<FsGlobEntry>({ kind: 'head', maxItems: globMaxResults })
67
- const allEntries: FsGlobEntry[] = []
68
- for await (const entry of candidates) {
69
- allEntries.push(entry)
70
- retainer.push(entry)
71
- }
72
- const { items, truncated, omitted } = retainer.finish()
63
+ declare const grepMaxMatches: number
64
+ declare const items: { length: number }
65
+ import type { Omitted } from '@deepseek-ai/dsh-output-retention'
73
66
 
74
- // bash: keep a head + tail, read to process exit.
75
- const out = new TextRetainer({ kind: 'headTail', headBytes: headCap, tailBytes: tailCap })
76
- child.stdout.on('data', (chunk: Buffer) => { out.push(chunk) })
77
- const { text, omittedBytes } = out.finish()
67
+ declare const omitted: Omitted
78
68
 
79
- // A footer: the library standardizes the omission clause; the tool owns recovery words.
80
69
  const footer = formatRetentionNotice(
81
70
  { scope: 'grep', strategy: 'head', unit: 'items', limit: grepMaxMatches, kept: items.length, omitted },
82
71
  ({ kept }) => `Results capped at ${kept}. Narrow the pattern, path, or include to see more.`,
83
72
  )
84
73
  ```
85
74
 
75
+ 库负责标准化省略子句(`Omitted 3 items.`)并把它与工具自有的恢复指引拼接;只有工具知道恢复动作,因此这些措辞由工具提供。
76
+
77
+ ### `truncated` 意味着什么
78
+
79
+ `truncated` 是预算事实:retainer 因上限而省略了本可获得的内容。它绝不表示上游不完整——权限失败、跳过二进制文件、提供方部分失败与不可读候选项都留在工具领域字段中,绝不并入 `truncated`。
80
+
81
+ ### 当前工具如何使用它
82
+
83
+ | 工具 | Retainer | 工具仍负责什么 |
84
+ |---|---|---|
85
+ | `glob` | `ItemRetainer`,`head` | spill 文件收集、路径映射、已跳过候选项、`incomplete` |
86
+ | `grep` | `ItemRetainer`,`head` | spill 文件收集、逐匹配预览截断、分组、排序 |
87
+ | `bash` | `TextRetainer`,`tail` 或 `headTail` | spill 文件、退出状态、信号、超时、后台任务 |
88
+ | `web_fetch` | `TextRetainer`,`head` 或 `headTail` | 提供方与资源上限、错误状态 |
89
+ | `web_search` | `ItemRetainer`,`head` | 「来源已达上限」通知措辞与提供方事实 |
90
+
91
+ `read` 不属于本库:它的行窗口分页(`offset`/`limit`、行号、`totalLines`)是文件专属渲染器,单个省略计数无法表示该窗口的两侧。
92
+
93
+ -----
94
+
95
+ <a id="understand-the-implementation"></a>
96
+ ## 理解实现
97
+
98
+ <details>
99
+ <summary>实现细节——点击展开</summary>
100
+
101
+ 本库建立在一个分离之上:它负责「保留了什么、省略了什么」这个机制问题;业务含义全部归工具包所有。
102
+
103
+ ### 源码地图
104
+
105
+ | 文件 | 职责 |
106
+ |---|---|
107
+ | [`src/index.ts`](src/index.ts) | `ItemRetainer`、`TextRetainer`、`describeOmitted` 与 `formatRetentionNotice` |
108
+ | [`src/invariant.ts`](src/invariant.ts) | 不变式伴生插件(无运行时不变式;保留运算由单元测试覆盖) |
109
+
110
+ ### 两个 retainer,两种资源模型
111
+
112
+ `ItemRetainer` 限制有序逻辑单元,只保留前 `maxItems` 个;调用方持续送入每个已观察单元,因此省略计数是精确的。`TextRetainer` 用同一个前缀/后缀累加器限制字节:`head` 只留前缀,`tail` 只留后缀,`headTail` 两者都留;累加器在内存中至多持有 `headBytes + tailBytes + 一个分片`,因此大流不会无界累积。
113
+
114
+ ### 预算事实如何保持诚实
115
+
116
+ `push()` 返回 `kept`(该单元或分片是否完整保留)与 `truncated`(是否已丢弃任何内容)。`finish()` 按实际返回的字节报告省略,因此丢弃部分码点字节的 UTF-8 边界修剪也会被计入——仅按预算推导的通知会高估保留文本。`describeOmitted` 只为 `exact` 打印计数;`unknown` 不打印计数,因为调用方没有提供。
117
+
118
+ ### read 渲染的排除
119
+
120
+ `read` 的 `offset`/`limit` 分页是行窗口渲染器,对所选窗口有自己的字节上限;单个 `Omitted` 值无法表示该窗口两侧,因此它不属于本库。
121
+
122
+ </details>
123
+
124
+ -----
125
+
126
+ <a id="further-exploration"></a>
127
+ ## 进一步探索
128
+
129
+ 当你需要消费方或库背后的边界决策时,阅读以下页面。
130
+
131
+ - [工具结果保留库 Agent Note](../../../.agents/notes/implemented/architecture/2026-07-06-tool-result-retention-library.zh.md)——库围绕工具语义划定的边界。
132
+ - [spill 策略](../../spill/spill-policy/README.zh.md)——组合 `TextRetainer`,围绕 spill 文件通知构建有界预览。
133
+ - [spill 子系统](../../../docs/subsystems/spill.zh.md)——本库预览机制所服务的 spill 词汇。
134
+ - [文件搜索工具](../../fs/tool-fs-search/README.zh.md)——为 spill 收集完整结果的 `ItemRetainer` 消费方。
135
+
136
+ -----
137
+
138
+ <a id="model-experience"></a>
86
139
  ## 模型体验
87
140
 
88
- 通过渲染保留内容和省略元数据的工具消费方间接影响模型。
141
+ 通过渲染保留内容与省略元数据的保留消费方间接影响模型。
89
142
 
90
143
  #### KV Cache 影响
91
144
 
92
- 不会直接导致 KV Cache 失效;请求前缀变更由上述消费方负责。
145
+ 不会直接导致失效;请求前缀的任何变更由保留消费方负责。
146
+
147
+ ## 已知限制与延期工作
148
+
149
+ <a id="known-limitations-and-deferred-work"></a>
150
+
151
+
152
+ 这些限制说明 retainer 刻意不覆盖什么。它们是当前包约束,不是任务积压。
153
+
154
+ - **项保留只支持 `head`**——tail、head/tail、分页、分组与提供方完整性语义仍归工具所有。
155
+ - **文本保留面向字节**——`read` 分页等行窗口与字符窗口需要单独的渲染器;切割可能丢弃部分 UTF-8 边界字节,以保持返回文本有效。
156
+
157
+ <a id="dev-note"></a>
158
+ ### 开发备注
159
+
160
+ <details>
161
+ <summary>维护者的工作上下文——点击展开</summary>
93
162
 
94
- ## 已知限制与暂缓事项
163
+ 无。
95
164
 
96
- - **项保留只支持 `head`**:tail、head/tail、分页、分组和提供方完整性语义仍由工具负责。
97
- - **文本保留面向字节**:`read` 分页等行窗口和字符窗口需要单独的渲染器;切割可能会丢弃部分 UTF-8 边界字节,以保持返回文本有效。
165
+ </details>
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "@deepseek-ai/dsh-output-retention",
3
3
  "description": "Zero-dependency bounded-retention primitive: ItemRetainer/TextRetainer + neutral notice helpers (what did we keep, what did we omit)",
4
- "version": "0.1.1-rc.2",
4
+ "version": "0.1.2-alpha.2",
5
5
  "publishConfig": {
6
6
  "access": "public"
7
7
  },
@@ -32,11 +32,11 @@
32
32
  ],
33
33
  "license": "MIT",
34
34
  "peerDependencies": {
35
- "@deepseek-ai/dsh-invariants": "^0.1.1-rc.2",
36
- "@deepseek-ai/cordis": "^4.0.1"
35
+ "@deepseek-ai/dsh-invariants": "^0.1.2-alpha.2",
36
+ "@deepseek-ai/cordis": "^4.0.2"
37
37
  },
38
38
  "devDependencies": {
39
- "@deepseek-ai/cordis": "^4.0.1",
40
- "@deepseek-ai/dsh-invariants": "^0.1.1-rc.2"
39
+ "@deepseek-ai/dsh-invariants": "^0.1.2-alpha.2",
40
+ "@deepseek-ai/cordis": "^4.0.2"
41
41
  }
42
42
  }