@deepseek-ai/dsh-output-retention 0.1.1-rc.2 → 0.1.2-alpha.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.i18n.yaml +2 -2
- package/README.md +128 -60
- package/README.zh.md +130 -62
- package/package.json +5 -5
package/README.i18n.yaml
CHANGED
|
@@ -2,5 +2,5 @@
|
|
|
2
2
|
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
|
3
3
|
# after editing either side, bring the other along and re-record with:
|
|
4
4
|
# pnpm run verify-translation-pairing --write packages/util/output-retention/README.md
|
|
5
|
-
README.md:
|
|
6
|
-
README.zh.md:
|
|
5
|
+
README.md: 6614f82149728b12419a4750a59772ef9e5de440
|
|
6
|
+
README.zh.md: 4f8d6bb235b7ee4c0b79a094308ad6c4a5913e36
|
package/README.md
CHANGED
|
@@ -1,97 +1,165 @@
|
|
|
1
|
-
|
|
1
|
+
---
|
|
2
|
+
description: "Bounded model-facing output for tools that must cap how much context they return: item and text retainers plus a standardized omission footer."
|
|
3
|
+
kind: "package-library"
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# @deepseek-ai/dsh-output-retention
|
|
2
7
|
|
|
3
8
|
English | [中文](README.zh.md)
|
|
4
9
|
|
|
5
|
-
|
|
10
|
+
## Summary
|
|
6
11
|
|
|
7
|
-
|
|
12
|
+
`dsh-output-retention` bounds how much context a tool returns to the model: a caller feeds items or text chunks into a retainer, then gets back the retained content plus exact omission metadata. `ItemRetainer` caps an ordered list of logical units (paths, matches, sources) at a head budget; `TextRetainer` caps a byte-oriented text stream with head, tail, or head-and-tail windows and keeps UTF-8 boundaries valid at every cut. A standardized omission clause and a notice formatter give tools a consistent "results capped" footer while the tool owns the recovery guidance. The library answers only the mechanical question of what was kept and what was omitted — grouping, line numbering, spill files, and provider error states stay in the tool. It is a dependency-light library that tool packages import directly; a `cordis.yml` cannot load it.
|
|
8
13
|
|
|
9
|
-
|
|
14
|
+
## Table of Contents
|
|
10
15
|
|
|
11
|
-
|
|
16
|
+
- [Use this package](#use-this-package)
|
|
17
|
+
- [Understand the implementation](#understand-the-implementation)
|
|
18
|
+
- [Further Exploration](#further-exploration)
|
|
19
|
+
- [Model Experience](#model-experience)
|
|
20
|
+
- [Known Limitations and Deferred Work](#known-limitations-and-deferred-work)
|
|
21
|
+
- [Dev Note](#dev-note)
|
|
12
22
|
|
|
13
|
-
|
|
14
|
-
import {
|
|
15
|
-
ItemRetainer, TextRetainer,
|
|
16
|
-
describeOmitted, formatRetentionNotice,
|
|
17
|
-
} from '@deepseek-ai/dsh-output-retention'
|
|
18
|
-
import type {
|
|
19
|
-
Omitted, PushDecision, RetainedItems, RetainedText,
|
|
20
|
-
ItemRetentionStrategy, TextRetentionStrategy, RetentionNotice,
|
|
21
|
-
} from '@deepseek-ai/dsh-output-retention'
|
|
22
|
-
```
|
|
23
|
+
-----
|
|
23
24
|
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
| `ItemRetainer<T>` | Bounds ordered logical units (paths, grep matches, sources). `head` only. `push()` → `PushDecision`; `finish()` → `RetainedItems<T>`. |
|
|
27
|
-
| `TextRetainer` | Bounds a byte-oriented text stream. `head` / `tail` / `headTail`, UTF-8 boundaries preserved at `finish()`. `push()` → `PushDecision`; `finish()` → `RetainedText`. |
|
|
28
|
-
| `describeOmitted(omitted, unit)` | Standardized omission clause (`exact` prints a count; `unknown` does not). |
|
|
29
|
-
| `formatRetentionNotice(notice, recovery)` | Joins the standardized omission clause with the tool's own recovery guidance. |
|
|
30
|
-
| `Omitted` | `none` / `exact` / `unknown` — how much was omitted. |
|
|
31
|
-
| `PushDecision` | `{ kept, truncated }` — the per-push retention result. |
|
|
25
|
+
<a id="use-this-package"></a>
|
|
26
|
+
## Use this package
|
|
32
27
|
|
|
33
|
-
|
|
28
|
+
Use a retainer wherever a tool must cap how much of its result reaches the model, and report honestly what was dropped. Choose `ItemRetainer` for ordered logical units and `TextRetainer` for byte-oriented streams.
|
|
34
29
|
|
|
35
|
-
|
|
30
|
+
### Bounding a list of items
|
|
36
31
|
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
## `truncated` is a budget fact, never "incomplete"
|
|
32
|
+
```ts
|
|
33
|
+
import { ItemRetainer } from '@deepseek-ai/dsh-output-retention'
|
|
41
34
|
|
|
42
|
-
|
|
35
|
+
declare const globMaxResults: number
|
|
36
|
+
declare const candidates: AsyncIterable<{ path: string }>
|
|
37
|
+
const retainer = new ItemRetainer<{ path: string }>({ kind: 'head', maxItems: globMaxResults })
|
|
38
|
+
for await (const entry of candidates) {
|
|
39
|
+
retainer.push(entry) // keep draining past the cap for an exact count
|
|
40
|
+
}
|
|
41
|
+
const { items, truncated, omitted } = retainer.finish()
|
|
42
|
+
```
|
|
43
43
|
|
|
44
|
-
|
|
44
|
+
`push()` reports per item whether it was kept, and `finish()` returns the retained items plus `omitted` — an exact count when the caller kept feeding every observed unit. A search tool can collect the full result set for a spill file while retaining only the first page for the model.
|
|
45
45
|
|
|
46
|
-
|
|
46
|
+
### Bounding a text stream
|
|
47
47
|
|
|
48
|
-
|
|
48
|
+
```text
|
|
49
|
+
import { TextRetainer } from '@deepseek-ai/dsh-output-retention'
|
|
49
50
|
|
|
50
|
-
|
|
51
|
+
const out = new TextRetainer({ kind: 'headTail', headBytes: headCap, tailBytes: tailCap })
|
|
52
|
+
child.stdout.on('data', (chunk: Buffer) => { out.push(chunk) })
|
|
53
|
+
const { text, omittedBytes } = out.finish()
|
|
54
|
+
```
|
|
51
55
|
|
|
52
|
-
|
|
53
|
-
|---|---|---|
|
|
54
|
-
| `glob` | `ItemRetainer<FsGlobEntry>`, `head` | Collect the full sorted path list for a spill file while retaining the first page inline. Path mapping, skipped candidates, and `incomplete` stay outside. |
|
|
55
|
-
| `grep` | `ItemRetainer<FlatGrepMatch>`, `head` | Collect matches for a spill file while retaining the first page inline. Per-match preview truncation, grouping, sorting, and `incomplete` stay outside. |
|
|
56
|
-
| `bash` | `TextRetainer`, `tail` or `headTail` | Executor still owns spill files, exit status, signal, timeout, and background jobs. |
|
|
57
|
-
| `web_fetch` | `TextRetainer`, `head` or `headTail` | Provider/resource caps stay provider facts; the retainer supplies only retained text and omission metadata. |
|
|
58
|
-
| `web_search` | `ItemRetainer<WebSearchSource>`, `head` | Standardizes the "sources capped" notice when providers return more sources than the model-facing result should include. |
|
|
56
|
+
`head`, `tail`, and `headTail` count bytes, not characters or lines: a child's pipe and an HTTP body are byte streams. `finish()` trims a partial codepoint at each cut, so the returned text never carries a replacement character introduced by the cut, and a codepoint is never reconstructed across the omitted middle.
|
|
59
57
|
|
|
60
|
-
|
|
58
|
+
### Building the omission footer
|
|
61
59
|
|
|
62
|
-
|
|
60
|
+
```ts
|
|
61
|
+
import { formatRetentionNotice } from '@deepseek-ai/dsh-output-retention'
|
|
63
62
|
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
const allEntries: FsGlobEntry[] = []
|
|
68
|
-
for await (const entry of candidates) {
|
|
69
|
-
allEntries.push(entry)
|
|
70
|
-
retainer.push(entry)
|
|
71
|
-
}
|
|
72
|
-
const { items, truncated, omitted } = retainer.finish()
|
|
63
|
+
declare const grepMaxMatches: number
|
|
64
|
+
declare const items: { length: number }
|
|
65
|
+
import type { Omitted } from '@deepseek-ai/dsh-output-retention'
|
|
73
66
|
|
|
74
|
-
|
|
75
|
-
const out = new TextRetainer({ kind: 'headTail', headBytes: headCap, tailBytes: tailCap })
|
|
76
|
-
child.stdout.on('data', (chunk: Buffer) => { out.push(chunk) })
|
|
77
|
-
const { text, omittedBytes } = out.finish()
|
|
67
|
+
declare const omitted: Omitted
|
|
78
68
|
|
|
79
|
-
// A footer: the library standardizes the omission clause; the tool owns recovery words.
|
|
80
69
|
const footer = formatRetentionNotice(
|
|
81
70
|
{ scope: 'grep', strategy: 'head', unit: 'items', limit: grepMaxMatches, kept: items.length, omitted },
|
|
82
71
|
({ kept }) => `Results capped at ${kept}. Narrow the pattern, path, or include to see more.`,
|
|
83
72
|
)
|
|
84
73
|
```
|
|
85
74
|
|
|
75
|
+
The library standardizes the omission clause (`Omitted 3 items.`) and joins it with the tool's own recovery guidance; only the tool knows the recovery action, so the tool supplies those words.
|
|
76
|
+
|
|
77
|
+
### What `truncated` means
|
|
78
|
+
|
|
79
|
+
`truncated` is a budget fact: the retainer omitted otherwise-available content because of a cap. It never means the upstream was incomplete — permission failures, skipped binary files, provider partial failures, and unreadable candidates stay in tool-domain fields, never folded into `truncated`.
|
|
80
|
+
|
|
81
|
+
### How the current tools use it
|
|
82
|
+
|
|
83
|
+
| Tool | Retainer | What the tool still owns |
|
|
84
|
+
|---|---|---|
|
|
85
|
+
| `glob` | `ItemRetainer`, `head` | Spill-file collection, path mapping, skipped candidates, `incomplete` |
|
|
86
|
+
| `grep` | `ItemRetainer`, `head` | Spill-file collection, per-match preview truncation, grouping, sorting |
|
|
87
|
+
| `bash` | `TextRetainer`, `tail` or `headTail` | Spill files, exit status, signal, timeout, background jobs |
|
|
88
|
+
| `web_fetch` | `TextRetainer`, `head` or `headTail` | Provider and resource caps, error states |
|
|
89
|
+
| `web_search` | `ItemRetainer`, `head` | The "sources capped" notice wording and provider facts |
|
|
90
|
+
|
|
91
|
+
`read` stays outside this library: its line-window pagination (`offset`/`limit`, line numbers, `totalLines`) is a file-specific renderer that a single omission count cannot represent.
|
|
92
|
+
|
|
93
|
+
-----
|
|
94
|
+
|
|
95
|
+
<a id="understand-the-implementation"></a>
|
|
96
|
+
## Understand the implementation
|
|
97
|
+
|
|
98
|
+
<details>
|
|
99
|
+
<summary>Implementation internals — click to expand</summary>
|
|
100
|
+
|
|
101
|
+
The library is built on one separation: it owns the mechanical question of what was kept and what was omitted; tool packages own every business meaning.
|
|
102
|
+
|
|
103
|
+
### Source map
|
|
104
|
+
|
|
105
|
+
| File | Role |
|
|
106
|
+
|---|---|
|
|
107
|
+
| [`src/index.ts`](src/index.ts) | `ItemRetainer`, `TextRetainer`, `describeOmitted`, and `formatRetentionNotice` |
|
|
108
|
+
| [`src/invariant.ts`](src/invariant.ts) | Invariant companion (no runtime invariant; the retention algebra is exercised by unit tests) |
|
|
109
|
+
|
|
110
|
+
### Two retainers, two resource models
|
|
111
|
+
|
|
112
|
+
`ItemRetainer` bounds ordered logical units and keeps only the first `maxItems`; the caller keeps pushing every observed unit so the omission count is exact. `TextRetainer` bounds bytes with one shared prefix/suffix accumulator: `head` is prefix-only, `tail` is suffix-only, `headTail` is both, and the accumulator holds at most `headBytes + tailBytes + one chunk` in memory, so a large stream does not accumulate unbounded.
|
|
113
|
+
|
|
114
|
+
### How the budget facts stay honest
|
|
115
|
+
|
|
116
|
+
`push()` returns `kept` (this unit or chunk fully retained) and `truncated` (anything dropped yet). `finish()` reports omission against the bytes actually returned, so a UTF-8 boundary trim that drops partial-codepoint bytes is counted too — a notice built from the budget alone would overstate the retained text. `describeOmitted` prints a count only for `exact`; `unknown` prints no count because the caller provided none.
|
|
117
|
+
|
|
118
|
+
### The read-render exclusion
|
|
119
|
+
|
|
120
|
+
`read`'s `offset`/`limit` pagination is a line-window renderer with its own byte cap over the selected window; a single `Omitted` value cannot represent both sides of that window, so it stays out of this library.
|
|
121
|
+
|
|
122
|
+
</details>
|
|
123
|
+
|
|
124
|
+
-----
|
|
125
|
+
|
|
126
|
+
<a id="further-exploration"></a>
|
|
127
|
+
## Further Exploration
|
|
128
|
+
|
|
129
|
+
Read these pages when you need the consumers or the boundary decision behind the library.
|
|
130
|
+
|
|
131
|
+
- [Tool-result retention library Agent Note](../../../.agents/notes/implemented/architecture/2026-07-06-tool-result-retention-library.md) — the boundary the library draws around tool semantics.
|
|
132
|
+
- [Spill policy](../../spill/spill-policy/README.md) — composes `TextRetainer` for a bounded preview around a spill-file notice.
|
|
133
|
+
- [Spill subsystem](../../../docs/subsystems/spill.md) — the spill vocabulary this library's preview mechanics serve.
|
|
134
|
+
- [File search tool](../../fs/tool-fs-search/README.md) — an `ItemRetainer` consumer collecting full results for spill.
|
|
135
|
+
|
|
136
|
+
-----
|
|
137
|
+
|
|
138
|
+
<a id="model-experience"></a>
|
|
86
139
|
## Model Experience
|
|
87
140
|
|
|
88
|
-
Indirectly, through
|
|
141
|
+
Indirectly, through the retention consumers that render retained content and omission metadata.
|
|
89
142
|
|
|
90
143
|
#### KV Cache effect
|
|
91
144
|
|
|
92
|
-
No direct invalidation; the
|
|
145
|
+
No direct invalidation; the retention consumers own any request-prefix changes.
|
|
93
146
|
|
|
94
147
|
## Known Limitations and Deferred Work
|
|
95
148
|
|
|
149
|
+
<a id="known-limitations-and-deferred-work"></a>
|
|
150
|
+
|
|
151
|
+
|
|
152
|
+
These limits define what the retainers deliberately do not cover. They are current package constraints, not a task backlog.
|
|
153
|
+
|
|
96
154
|
- **Item retention supports `head` only** — tail, head/tail, pagination, grouping, and provider-completeness semantics remain tool-owned.
|
|
97
|
-
- **Text retention is byte-oriented** — line and character windows such as `read` pagination require a separate renderer, and a cut may discard partial UTF-8 boundary bytes to keep returned text valid.
|
|
155
|
+
- **Text retention is byte-oriented** — line and character windows such as `read` pagination require a separate renderer, and a cut may discard partial UTF-8 boundary bytes to keep the returned text valid.
|
|
156
|
+
|
|
157
|
+
<a id="dev-note"></a>
|
|
158
|
+
### Dev Note
|
|
159
|
+
|
|
160
|
+
<details>
|
|
161
|
+
<summary>Working context for maintainers — click to expand</summary>
|
|
162
|
+
|
|
163
|
+
None.
|
|
164
|
+
|
|
165
|
+
</details>
|
package/README.zh.md
CHANGED
|
@@ -1,97 +1,165 @@
|
|
|
1
|
-
|
|
1
|
+
---
|
|
2
|
+
description: "为必须限制返回上下文量的工具提供有界的面向模型输出:项与文本 retainer,以及标准化的省略页脚。"
|
|
3
|
+
kind: "package-library"
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# @deepseek-ai/dsh-output-retention
|
|
2
7
|
|
|
3
8
|
[English](README.md) | 中文
|
|
4
9
|
|
|
5
|
-
|
|
10
|
+
## 概述
|
|
6
11
|
|
|
7
|
-
|
|
12
|
+
`dsh-output-retention` 限制工具返回给模型的上下文量:调用方把项或文本分片送入 retainer,然后取回保留的内容与精确的省略元数据。`ItemRetainer` 以头部预算限制有序逻辑单元列表(路径、匹配项、来源);`TextRetainer` 以 head、tail 或 head+tail 窗口限制面向字节的文本流,并在每个切割处保持 UTF-8 边界有效。标准化的省略子句与通知格式化器让工具获得一致的「结果已达上限」页脚,而恢复指引由工具自己提供。该库只回答「保留了什么、省略了什么」这个机制问题——分组、行号、spill 文件与提供方错误状态都留在工具侧。它是轻依赖库,由工具包直接导入;`cordis.yml` 无法加载它。
|
|
8
13
|
|
|
9
|
-
|
|
14
|
+
## 目录
|
|
10
15
|
|
|
11
|
-
|
|
16
|
+
- [使用本包](#use-this-package)
|
|
17
|
+
- [理解实现](#understand-the-implementation)
|
|
18
|
+
- [进一步探索](#further-exploration)
|
|
19
|
+
- [模型体验](#model-experience)
|
|
20
|
+
- [已知限制与延期工作](#known-limitations-and-deferred-work)
|
|
21
|
+
- [开发备注](#dev-note)
|
|
12
22
|
|
|
13
|
-
|
|
14
|
-
import {
|
|
15
|
-
ItemRetainer, TextRetainer,
|
|
16
|
-
describeOmitted, formatRetentionNotice,
|
|
17
|
-
} from '@deepseek-ai/dsh-output-retention'
|
|
18
|
-
import type {
|
|
19
|
-
Omitted, PushDecision, RetainedItems, RetainedText,
|
|
20
|
-
ItemRetentionStrategy, TextRetentionStrategy, RetentionNotice,
|
|
21
|
-
} from '@deepseek-ai/dsh-output-retention'
|
|
22
|
-
```
|
|
23
|
+
-----
|
|
23
24
|
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
| `ItemRetainer<T>` | 限制有序逻辑单元(路径、grep 匹配项、来源)。只支持 `head`。`push()` → `PushDecision`;`finish()` → `RetainedItems<T>`。 |
|
|
27
|
-
| `TextRetainer` | 限制面向字节的文本流。`head` / `tail` / `headTail`,并在 `finish()` 时保留 UTF-8 边界。`push()` → `PushDecision`;`finish()` → `RetainedText`。 |
|
|
28
|
-
| `describeOmitted(omitted, unit)` | 标准化的省略子句(`exact` 输出数量;`unknown` 不输出)。 |
|
|
29
|
-
| `formatRetentionNotice(notice, recovery)` | 将标准化的省略子句与工具自有的恢复指引连接起来。 |
|
|
30
|
-
| `Omitted` | `none` / `exact` / `unknown`:省略了多少内容。 |
|
|
31
|
-
| `PushDecision` | `{ kept, truncated }`:每次 push 的保留结果。 |
|
|
25
|
+
<a id="use-this-package"></a>
|
|
26
|
+
## 使用本包
|
|
32
27
|
|
|
33
|
-
|
|
28
|
+
凡是工具必须限制其结果到达模型的数量、并如实报告丢弃内容的地方,都使用 retainer。有序逻辑单元选 `ItemRetainer`,面向字节的流选 `TextRetainer`。
|
|
34
29
|
|
|
35
|
-
|
|
30
|
+
### 限制项列表
|
|
36
31
|
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
## `truncated` 是预算事实,绝不表示「不完整」
|
|
32
|
+
```ts
|
|
33
|
+
import { ItemRetainer } from '@deepseek-ai/dsh-output-retention'
|
|
41
34
|
|
|
42
|
-
|
|
35
|
+
declare const globMaxResults: number
|
|
36
|
+
declare const candidates: AsyncIterable<{ path: string }>
|
|
37
|
+
const retainer = new ItemRetainer<{ path: string }>({ kind: 'head', maxItems: globMaxResults })
|
|
38
|
+
for await (const entry of candidates) {
|
|
39
|
+
retainer.push(entry) // keep draining past the cap for an exact count
|
|
40
|
+
}
|
|
41
|
+
const { items, truncated, omitted } = retainer.finish()
|
|
42
|
+
```
|
|
43
43
|
|
|
44
|
-
|
|
44
|
+
`push()` 逐项报告该项是否被保留,`finish()` 返回保留的项与 `omitted`——当调用方持续送入每个已观察单元时,这是一个精确计数。搜索工具可以收集完整结果集用于 spill 文件,同时只为模型保留第一页。
|
|
45
45
|
|
|
46
|
-
|
|
46
|
+
### 限制文本流
|
|
47
47
|
|
|
48
|
-
|
|
48
|
+
```text
|
|
49
|
+
import { TextRetainer } from '@deepseek-ai/dsh-output-retention'
|
|
49
50
|
|
|
50
|
-
|
|
51
|
+
const out = new TextRetainer({ kind: 'headTail', headBytes: headCap, tailBytes: tailCap })
|
|
52
|
+
child.stdout.on('data', (chunk: Buffer) => { out.push(chunk) })
|
|
53
|
+
const { text, omittedBytes } = out.finish()
|
|
54
|
+
```
|
|
51
55
|
|
|
52
|
-
|
|
53
|
-
|---|---|---|
|
|
54
|
-
| `glob` | `ItemRetainer<FsGlobEntry>`,`head` | 收集完整的已排序路径列表用于 spill 文件,同时在内联位置保留第一页。路径映射、已跳过候选项和 `incomplete` 保留在外部。 |
|
|
55
|
-
| `grep` | `ItemRetainer<FlatGrepMatch>`,`head` | 收集匹配项用于 spill 文件,同时在内联位置保留第一页。每个匹配项的预览截断、分组、排序和 `incomplete` 保留在外部。 |
|
|
56
|
-
| `bash` | `TextRetainer`,`tail` 或 `headTail` | 执行器仍负责 spill 文件、退出状态、信号、超时和后台任务。 |
|
|
57
|
-
| `web_fetch` | `TextRetainer`,`head` 或 `headTail` | 提供方/资源上限保留为提供方事实;retainer 只提供保留文本和省略元数据。 |
|
|
58
|
-
| `web_search` | `ItemRetainer<WebSearchSource>`,`head` | 当提供方返回的来源超过面向模型的结果应包含的数量时,标准化「来源已达上限」通知。 |
|
|
56
|
+
`head`、`tail` 与 `headTail` 按字节而非字符或行计数:子进程管道与 HTTP 正文都是字节流。`finish()` 会在每个切割处修剪不完整的码点,因此返回的文本绝不会携带由切割引入的替换字符,码点也绝不会跨被省略的中间部分重建。
|
|
59
57
|
|
|
60
|
-
|
|
58
|
+
### 构建省略页脚
|
|
61
59
|
|
|
62
|
-
|
|
60
|
+
```ts
|
|
61
|
+
import { formatRetentionNotice } from '@deepseek-ai/dsh-output-retention'
|
|
63
62
|
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
const allEntries: FsGlobEntry[] = []
|
|
68
|
-
for await (const entry of candidates) {
|
|
69
|
-
allEntries.push(entry)
|
|
70
|
-
retainer.push(entry)
|
|
71
|
-
}
|
|
72
|
-
const { items, truncated, omitted } = retainer.finish()
|
|
63
|
+
declare const grepMaxMatches: number
|
|
64
|
+
declare const items: { length: number }
|
|
65
|
+
import type { Omitted } from '@deepseek-ai/dsh-output-retention'
|
|
73
66
|
|
|
74
|
-
|
|
75
|
-
const out = new TextRetainer({ kind: 'headTail', headBytes: headCap, tailBytes: tailCap })
|
|
76
|
-
child.stdout.on('data', (chunk: Buffer) => { out.push(chunk) })
|
|
77
|
-
const { text, omittedBytes } = out.finish()
|
|
67
|
+
declare const omitted: Omitted
|
|
78
68
|
|
|
79
|
-
// A footer: the library standardizes the omission clause; the tool owns recovery words.
|
|
80
69
|
const footer = formatRetentionNotice(
|
|
81
70
|
{ scope: 'grep', strategy: 'head', unit: 'items', limit: grepMaxMatches, kept: items.length, omitted },
|
|
82
71
|
({ kept }) => `Results capped at ${kept}. Narrow the pattern, path, or include to see more.`,
|
|
83
72
|
)
|
|
84
73
|
```
|
|
85
74
|
|
|
75
|
+
库负责标准化省略子句(`Omitted 3 items.`)并把它与工具自有的恢复指引拼接;只有工具知道恢复动作,因此这些措辞由工具提供。
|
|
76
|
+
|
|
77
|
+
### `truncated` 意味着什么
|
|
78
|
+
|
|
79
|
+
`truncated` 是预算事实:retainer 因上限而省略了本可获得的内容。它绝不表示上游不完整——权限失败、跳过二进制文件、提供方部分失败与不可读候选项都留在工具领域字段中,绝不并入 `truncated`。
|
|
80
|
+
|
|
81
|
+
### 当前工具如何使用它
|
|
82
|
+
|
|
83
|
+
| 工具 | Retainer | 工具仍负责什么 |
|
|
84
|
+
|---|---|---|
|
|
85
|
+
| `glob` | `ItemRetainer`,`head` | spill 文件收集、路径映射、已跳过候选项、`incomplete` |
|
|
86
|
+
| `grep` | `ItemRetainer`,`head` | spill 文件收集、逐匹配预览截断、分组、排序 |
|
|
87
|
+
| `bash` | `TextRetainer`,`tail` 或 `headTail` | spill 文件、退出状态、信号、超时、后台任务 |
|
|
88
|
+
| `web_fetch` | `TextRetainer`,`head` 或 `headTail` | 提供方与资源上限、错误状态 |
|
|
89
|
+
| `web_search` | `ItemRetainer`,`head` | 「来源已达上限」通知措辞与提供方事实 |
|
|
90
|
+
|
|
91
|
+
`read` 不属于本库:它的行窗口分页(`offset`/`limit`、行号、`totalLines`)是文件专属渲染器,单个省略计数无法表示该窗口的两侧。
|
|
92
|
+
|
|
93
|
+
-----
|
|
94
|
+
|
|
95
|
+
<a id="understand-the-implementation"></a>
|
|
96
|
+
## 理解实现
|
|
97
|
+
|
|
98
|
+
<details>
|
|
99
|
+
<summary>实现细节——点击展开</summary>
|
|
100
|
+
|
|
101
|
+
本库建立在一个分离之上:它负责「保留了什么、省略了什么」这个机制问题;业务含义全部归工具包所有。
|
|
102
|
+
|
|
103
|
+
### 源码地图
|
|
104
|
+
|
|
105
|
+
| 文件 | 职责 |
|
|
106
|
+
|---|---|
|
|
107
|
+
| [`src/index.ts`](src/index.ts) | `ItemRetainer`、`TextRetainer`、`describeOmitted` 与 `formatRetentionNotice` |
|
|
108
|
+
| [`src/invariant.ts`](src/invariant.ts) | 不变式伴生插件(无运行时不变式;保留运算由单元测试覆盖) |
|
|
109
|
+
|
|
110
|
+
### 两个 retainer,两种资源模型
|
|
111
|
+
|
|
112
|
+
`ItemRetainer` 限制有序逻辑单元,只保留前 `maxItems` 个;调用方持续送入每个已观察单元,因此省略计数是精确的。`TextRetainer` 用同一个前缀/后缀累加器限制字节:`head` 只留前缀,`tail` 只留后缀,`headTail` 两者都留;累加器在内存中至多持有 `headBytes + tailBytes + 一个分片`,因此大流不会无界累积。
|
|
113
|
+
|
|
114
|
+
### 预算事实如何保持诚实
|
|
115
|
+
|
|
116
|
+
`push()` 返回 `kept`(该单元或分片是否完整保留)与 `truncated`(是否已丢弃任何内容)。`finish()` 按实际返回的字节报告省略,因此丢弃部分码点字节的 UTF-8 边界修剪也会被计入——仅按预算推导的通知会高估保留文本。`describeOmitted` 只为 `exact` 打印计数;`unknown` 不打印计数,因为调用方没有提供。
|
|
117
|
+
|
|
118
|
+
### read 渲染的排除
|
|
119
|
+
|
|
120
|
+
`read` 的 `offset`/`limit` 分页是行窗口渲染器,对所选窗口有自己的字节上限;单个 `Omitted` 值无法表示该窗口两侧,因此它不属于本库。
|
|
121
|
+
|
|
122
|
+
</details>
|
|
123
|
+
|
|
124
|
+
-----
|
|
125
|
+
|
|
126
|
+
<a id="further-exploration"></a>
|
|
127
|
+
## 进一步探索
|
|
128
|
+
|
|
129
|
+
当你需要消费方或库背后的边界决策时,阅读以下页面。
|
|
130
|
+
|
|
131
|
+
- [工具结果保留库 Agent Note](../../../.agents/notes/implemented/architecture/2026-07-06-tool-result-retention-library.zh.md)——库围绕工具语义划定的边界。
|
|
132
|
+
- [spill 策略](../../spill/spill-policy/README.zh.md)——组合 `TextRetainer`,围绕 spill 文件通知构建有界预览。
|
|
133
|
+
- [spill 子系统](../../../docs/subsystems/spill.zh.md)——本库预览机制所服务的 spill 词汇。
|
|
134
|
+
- [文件搜索工具](../../fs/tool-fs-search/README.zh.md)——为 spill 收集完整结果的 `ItemRetainer` 消费方。
|
|
135
|
+
|
|
136
|
+
-----
|
|
137
|
+
|
|
138
|
+
<a id="model-experience"></a>
|
|
86
139
|
## 模型体验
|
|
87
140
|
|
|
88
|
-
|
|
141
|
+
通过渲染保留内容与省略元数据的保留消费方间接影响模型。
|
|
89
142
|
|
|
90
143
|
#### KV Cache 影响
|
|
91
144
|
|
|
92
|
-
|
|
145
|
+
不会直接导致失效;请求前缀的任何变更由保留消费方负责。
|
|
146
|
+
|
|
147
|
+
## 已知限制与延期工作
|
|
148
|
+
|
|
149
|
+
<a id="known-limitations-and-deferred-work"></a>
|
|
150
|
+
|
|
151
|
+
|
|
152
|
+
这些限制说明 retainer 刻意不覆盖什么。它们是当前包约束,不是任务积压。
|
|
153
|
+
|
|
154
|
+
- **项保留只支持 `head`**——tail、head/tail、分页、分组与提供方完整性语义仍归工具所有。
|
|
155
|
+
- **文本保留面向字节**——`read` 分页等行窗口与字符窗口需要单独的渲染器;切割可能丢弃部分 UTF-8 边界字节,以保持返回文本有效。
|
|
156
|
+
|
|
157
|
+
<a id="dev-note"></a>
|
|
158
|
+
### 开发备注
|
|
159
|
+
|
|
160
|
+
<details>
|
|
161
|
+
<summary>维护者的工作上下文——点击展开</summary>
|
|
93
162
|
|
|
94
|
-
|
|
163
|
+
无。
|
|
95
164
|
|
|
96
|
-
|
|
97
|
-
- **文本保留面向字节**:`read` 分页等行窗口和字符窗口需要单独的渲染器;切割可能会丢弃部分 UTF-8 边界字节,以保持返回文本有效。
|
|
165
|
+
</details>
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@deepseek-ai/dsh-output-retention",
|
|
3
3
|
"description": "Zero-dependency bounded-retention primitive: ItemRetainer/TextRetainer + neutral notice helpers (what did we keep, what did we omit)",
|
|
4
|
-
"version": "0.1.
|
|
4
|
+
"version": "0.1.2-alpha.2",
|
|
5
5
|
"publishConfig": {
|
|
6
6
|
"access": "public"
|
|
7
7
|
},
|
|
@@ -32,11 +32,11 @@
|
|
|
32
32
|
],
|
|
33
33
|
"license": "MIT",
|
|
34
34
|
"peerDependencies": {
|
|
35
|
-
"@deepseek-ai/dsh-invariants": "^0.1.
|
|
36
|
-
"@deepseek-ai/cordis": "^4.0.
|
|
35
|
+
"@deepseek-ai/dsh-invariants": "^0.1.2-alpha.2",
|
|
36
|
+
"@deepseek-ai/cordis": "^4.0.2"
|
|
37
37
|
},
|
|
38
38
|
"devDependencies": {
|
|
39
|
-
"@deepseek-ai/
|
|
40
|
-
"@deepseek-ai/
|
|
39
|
+
"@deepseek-ai/dsh-invariants": "^0.1.2-alpha.2",
|
|
40
|
+
"@deepseek-ai/cordis": "^4.0.2"
|
|
41
41
|
}
|
|
42
42
|
}
|