@crazx/dsh-compaction-basic 0.1.2-rc.1.zw.2 → 0.1.5-alpha.1.zw.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.i18n.yaml +2 -2
- package/README.md +12 -9
- package/README.zh.md +17 -14
- package/lib/index.js +73 -60
- package/lib/types/region.d.ts +4 -2
- package/lib/types/summarizer.d.ts +16 -4
- package/package.json +22 -22
package/README.i18n.yaml
CHANGED
|
@@ -2,5 +2,5 @@
|
|
|
2
2
|
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
|
3
3
|
# after editing either side, bring the other along and re-record with:
|
|
4
4
|
# pnpm run verify-translation-pairing --write packages/compaction/compaction-basic/README.md
|
|
5
|
-
README.md:
|
|
6
|
-
README.zh.md:
|
|
5
|
+
README.md: d70bc758c5bba4a9791a1f5017beb80783b1017e
|
|
6
|
+
README.zh.md: 086dc4a49d423daa661a0521e50be51db8db7872
|
package/README.md
CHANGED
|
@@ -9,7 +9,7 @@ English | [中文](README.zh.md)
|
|
|
9
9
|
|
|
10
10
|
## Summary
|
|
11
11
|
|
|
12
|
-
|
|
12
|
+
This package keeps long agent conversations working near the model's context limit. As token pressure builds, it condenses the oldest history into a summary while preserving recent messages; after a context-overflow error, it condenses and retries. You can also request condensation with `/compact` and optionally trim oversized tool outputs first. Fitting input uses one extra model request; oversized input uses bounded hierarchical map-reduce calls, and only the final summary text is retained. It cannot reduce the system prompt, tools, or session prefix, or split one indivisible unit such as a single huge tool call.
|
|
13
13
|
|
|
14
14
|
## Table of Contents
|
|
15
15
|
|
|
@@ -109,18 +109,18 @@ The backend is built on four commitments:
|
|
|
109
109
|
|
|
110
110
|
- **One measurement service prices every decision.** The singleton `ctx.tokenMeter` measures the latest canonical logged envelope and current surface at one consumed-log revision. When the routed adapter declares request-image pricing, the meter applies it to image history. Pressure, recent-tail retention, range selection, and shrink validation use the same route-priced node figures; logged replacement shadow prices stay on the route-independent heuristic so pure projection folds remain consistent.
|
|
111
111
|
- **The log-recorded bracket is the transaction.** All entry points share one bracket-first region transaction: validate the range and live lock, append `compaction/start` synchronously, prepare and await the summary, revalidate, append `compaction/summary` plus the replacement, and make exactly one closing attempt. Automatic and explicit-region calls require a numeric open-turn owner and whole-surface stability; `compactNow()` reserves idle admission, uses `turn: null`, accepts append-only context outside its selected span, flushes every closed attempt, and releases admission in `finally`.
|
|
112
|
-
- **Summarization
|
|
112
|
+
- **Summarization preserves the cache-reusing fast path and bounds oversized work.** A fitting one-shot replays the system prompt held by the `system/message` at surface node 0, the last routed request's tools, and the shadowed-region messages byte-for-byte. Hierarchical calls replay that same fixed system head exactly once while bounding chronological map spans and recursive reductions.
|
|
113
113
|
- **`summarize()` is the sole subclass hook.** A template- or remote-summarizer subclass can override it while pressure, retention, cited source events, shrink validation, and shadowed-token accounting stay on the token meter.
|
|
114
114
|
|
|
115
115
|
### Automatic triggers and overflow recovery
|
|
116
116
|
|
|
117
|
-
With `auto: true`, a serial `agent/pre-step` listener checks pressure before request derivation: it prices the latest durable routed request envelope through `ctx.tokenMeter`, and when pressure crosses the routed model's threshold it prunes, then summarizes the oldest balanced span while keeping a priced recent tail. The `agent/request-error` listener reacts to a provider-confirmed `CONTEXT_WINDOW_EXCEEDED`: it bypasses the normal threshold and retention policy, attempts one maximal balanced head reduction, and authorizes a retry only after the surface replacement generation advances. Cancellation stays authoritative throughout.
|
|
117
|
+
With `auto: true`, a serial `agent/pre-step` listener checks pressure before request derivation: it prices the latest durable routed request envelope through `ctx.tokenMeter`, and when pressure crosses the routed model's threshold it prunes, then summarizes the oldest balanced span while keeping a priced recent tail. Every selected range starts at the first surface node that is not a `system/message`, so a system prompt at surface node 0 is never shadowed; a later `system/message` appended by an in-history prompt update is ordinary history that the range may shadow, and the agent loop's projection then replaces node 0 with the current prompt when their text differs ([decision rule](../../core/agent-loop/README.md#understand-the-implementation)). The `agent/request-error` listener reacts to a provider-confirmed `CONTEXT_WINDOW_EXCEEDED`: it bypasses the normal threshold and retention policy, attempts one maximal balanced head reduction, and authorizes a retry only after the surface replacement generation advances. Cancellation stays authoritative throughout.
|
|
118
118
|
|
|
119
119
|
Pressure policy resolves capacity from the adapter that owns the durable route. An adapter that returns no capacity for a valid dynamic route makes the manual pressure path throw a target-specific configuration error; the automatic listener warns once for that exact target and continues with full history.
|
|
120
120
|
|
|
121
121
|
### Summarization mechanics
|
|
122
122
|
|
|
123
|
-
A direct `ctx.llm.stream()` call uses the configured provider/model pair and cap, falling back to the latest logged request target and then the `AgentOptions` pair, without running the loop-only `agent/request` extension point.
|
|
123
|
+
A direct `ctx.llm.stream()` call uses the configured provider/model pair and cap, falling back to the latest logged request target and then the `AgentOptions` pair, without running the loop-only `agent/request` extension point. When the complete request fits, the call replays the derived `system/message` at surface node 0 as the leading entry of `messages`, followed by the shadowed-region messages (including a shadowed in-history `system/message` in its surface position), carries the header's tools verbatim, and appends the compaction instruction as the final user message. Oversized or Provider-rejected input is grouped into tool-balanced chronological map spans and recursively reduced; every map and reduce request replays the current non-empty system head exactly once as its first message, while later dynamic `system/message` entries remain ordinary chronological map input and are never promoted or deleted. The fixed-head, optional-tool, and instruction cost is included in every stage budget. Tool schemas accompany hierarchy calls only when `replayTools: true`. An empty-content system head contributes no message but remains outside the compacted range. Every call sets `GenerateOptions.purpose` to `compaction`; only returned text enters the checkpoint, excluding reasoning and tool calls. Image output fails with `UNSUPPORTED_CONTENT` rather than disappearing. The replacement user message frames the summary with `<compacted-summary>` tags; the raw final summary remains on the `compaction/summary` event.
|
|
124
124
|
|
|
125
125
|
### The region transaction
|
|
126
126
|
|
|
@@ -136,7 +136,10 @@ The transaction validates the surface span and the durable lock, appends `compac
|
|
|
136
136
|
|---|---|
|
|
137
137
|
| [`src/index.ts`](src/index.ts) | Plugin entry: `BasicCompactionEngine`, automatic listeners, entry-point dispatch |
|
|
138
138
|
| [`src/region.ts`](src/region.ts) | Retention selection and the shared bracket-first compaction transaction |
|
|
139
|
-
| [`src/summarizer.ts`](src/summarizer.ts) | Default `ctx.llm.stream()` summarization, checkpoint framing, safe-summary projection |
|
|
139
|
+
| [`src/summarizer.ts`](src/summarizer.ts) | Default `ctx.llm.stream()` one-shot summarization, checkpoint framing, safe-summary projection |
|
|
140
|
+
| [`src/hierarchical.ts`](src/hierarchical.ts) | Bounded map-reduce fallback, adaptive splitting, stage usage aggregation |
|
|
141
|
+
| [`src/hierarchical-planner.ts`](src/hierarchical-planner.ts) | Tool-balanced units and greedy token-budget planning |
|
|
142
|
+
| [`src/hierarchical-prompts.ts`](src/hierarchical-prompts.ts) | Structured map/reduce prompts and output validation |
|
|
140
143
|
| [`src/config.ts`](src/config.ts) | Load-time validation and routed-model policy resolution |
|
|
141
144
|
| [`src/types.ts`](src/types.ts) | `BasicCompactionConfig` and resolved policy vocabulary |
|
|
142
145
|
| — | No runtime invariant companion is published; this package exposes no independent event sequence or mutable data relation beyond contracts enforced at its owning seam. |
|
|
@@ -186,7 +189,7 @@ Replacing rather than append-only. Each checkpoint invalidates reuse from the fi
|
|
|
186
189
|
|
|
187
190
|
#### What the model sees
|
|
188
191
|
|
|
189
|
-
When the complete request fits, the summarization model receives the conversation replayed verbatim — the same system
|
|
192
|
+
When the complete request fits, the summarization model receives the conversation replayed verbatim — the same system head, tool schemas, and messages the last routed request sent for the shadowed region — followed by one final user message: the compaction instruction below. For hierarchy, each map request receives the same system head exactly once, an ordered tool-balanced source span, and a structured map instruction; reduce requests receive that same head exactly once, ordered `<partial-summary>` frames, and a structured reduce instruction. Later system updates remain in their original map chronology instead of becoming a second fixed head. Tool schemas accompany hierarchy calls only when `replayTools: true`. The conversation model never sees these private requests or their reasoning; only the final text is stored.
|
|
190
193
|
|
|
191
194
|
##### Compaction instruction (final user message)
|
|
192
195
|
|
|
@@ -233,7 +236,7 @@ A fitting input costs one separate model call: the replayed conversation prefix
|
|
|
233
236
|
|
|
234
237
|
#### KV Cache effect
|
|
235
238
|
|
|
236
|
-
The fitting one-shot
|
|
239
|
+
The fitting one-shot preserves the complete warm prefix through the shadowed region. Hierarchical calls preserve prefix reuse for the fixed system head and, when enabled, tool schemas; map payloads and reduction frames diverge after that shared prefix. Routing the summarizer to a different provider/model forgoes the conversation route's cache reuse.
|
|
237
240
|
|
|
238
241
|
## Known Limitations and Deferred Work
|
|
239
242
|
|
|
@@ -244,9 +247,9 @@ These limits define when automatic condensation is a poor fit or needs special c
|
|
|
244
247
|
|
|
245
248
|
- **Meter accuracy follows the fixed heuristic** — missing reusable provider usage falls back to character count plus structural overhead rather than exact tokenization; image occurrences carry provider-exact visual tokens only on routes whose adapter declares request-image pricing.
|
|
246
249
|
- **Overflow classification is adapter-maintained** — provider wording can change; both DeepSeek adapters normalize recognized context-limit failures to `CONTEXT_WINDOW_EXCEEDED`.
|
|
247
|
-
- **Bounded recovery requires summary-model capacity metadata** — an adapter that omits `contextWindow` keeps the legacy one-shot path. If that request succeeds, behavior is unchanged; if it overflows, hierarchy cannot derive safe chunk budgets and fails with an actionable capacity error.
|
|
248
|
-
- **Hierarchy output is a strict checkpoint protocol** — every map and reduce stage must return all required headings. Truncation, visual output, malformed structure, exhausted `maxDepth`, or an indivisible source/partial that still overflows fails the complete compaction transaction without installing a partial checkpoint.
|
|
249
250
|
- **Some indivisible-unit and envelope-only overflow remains outside surface compaction** — recovery cannot shrink system/tools/prefix, split an indivisible non-tool node, or repair a tool unit whose non-prunable remainder still exceeds the window. The optional pruner can shrink text-bearing tool-result bulk inside an otherwise indivisible pair.
|
|
251
|
+
- **Hierarchy requires declared summary-model capacity** — without a positive integer `contextWindow`, fitting input still uses one-shot summarization, but a confirmed one-shot overflow fails clearly because bounded chunk planning has no trustworthy window.
|
|
252
|
+
- **Hierarchy is intentionally bounded** — a Provider-rejected indivisible tool-balanced span, fixed system/tools/instruction overhead that exhausts the stage budget, a reduction round that does not reduce partial count, or reaching `maxDepth` fails instead of looping.
|
|
250
253
|
- **`compactRegion` requires an open turn** — a manual call on a fully-closed session throws ("no open turn") rather than compacting.
|
|
251
254
|
- **Summarization failure preserves the latest durable surface** — before any replacement, the auto path logs a warning and proceeds with full over-budget history. If pruning already landed, a later summarization failure proceeds from that durable pruned surface. Summarization truncation at `maxTokens`, which hidden reasoning tokens can consume, follows the same rule.
|
|
252
255
|
|
package/README.zh.md
CHANGED
|
@@ -9,7 +9,7 @@ kind: "package-reference"
|
|
|
9
9
|
|
|
10
10
|
## 概述
|
|
11
11
|
|
|
12
|
-
|
|
12
|
+
本包让长时 agent 会话在接近模型上下文上限时仍能正常工作。token 压力上升时,它会把最旧的历史压缩为摘要并保留近期消息;上下文溢出错误发生后,它会压缩并重试。你也可以通过 `/compact` 按需压缩,并选择先修剪超大工具输出。能容纳的输入使用一次额外模型请求;超大输入使用有界层次 map-reduce 调用,并且只保留最终摘要文本。它无法缩减系统提示词、工具或会话前缀,也无法拆分单个不可分单元(例如一次超大工具调用)。
|
|
13
13
|
|
|
14
14
|
## 目录
|
|
15
15
|
|
|
@@ -29,7 +29,7 @@ kind: "package-reference"
|
|
|
29
29
|
|
|
30
30
|
### 你会得到什么
|
|
31
31
|
|
|
32
|
-
默认设置下你会获得四种行为:会话向模型上下文上限增长时自动压缩;提供方确认上下文溢出错误后的恢复(先压缩再重试该请求);通过 `/compact`
|
|
32
|
+
默认设置下你会获得四种行为:会话向模型上下文上限增长时自动压缩;提供方确认上下文溢出错误后的恢复(先压缩再重试该请求);通过 `/compact` 命令按需压缩;以及——挂载修剪器时——压缩前对超大工具输出的修剪。能放入摘要模型窗口的输入保留复用 cache 的 one-shot 路径;超大输入或提供方拒绝的输入回退到有界时序 map-reduce。
|
|
33
33
|
|
|
34
34
|
### 最小可用组合
|
|
35
35
|
|
|
@@ -71,11 +71,11 @@ kind: "package-reference"
|
|
|
71
71
|
| `maxTokens` | `8192` | 摘要请求的输出上限;可包含推理 token。 |
|
|
72
72
|
| `compactionRetries` | `1` | 压力仍高于阈值时,在首次压缩后进行的额外尝试次数。 |
|
|
73
73
|
| `maxOverflowRetries` | `1` | 已确认上下文窗口溢出后的最大重试次数;`0` 只禁用恢复。 |
|
|
74
|
-
| `chunkInputRatio` | `0.6` |
|
|
74
|
+
| `chunkInputRatio` | `0.6` | 每个层次阶段输入可使用的摘要模型窗口比例;有效范围为 `[0.1, 0.9]`。 |
|
|
75
75
|
| `mapMaxTokens` | `4096` | 单次层次 map 调用的提供方生成上限。 |
|
|
76
76
|
| `reduceMaxTokens` | `8192` | 单次层次 reduce 调用的提供方生成上限。 |
|
|
77
|
-
| `maxDepth` | `4` |
|
|
78
|
-
| `replayTools` | `false` |
|
|
77
|
+
| `maxDepth` | `4` | 最大递归 reduce 轮数;有效范围为 `1..8`。 |
|
|
78
|
+
| `replayTools` | `false` | 在层次阶段重放工具 schema。严格提供方可能需要启用它,但它会占用分块输入并降低前缀复用。 |
|
|
79
79
|
| `modelPolicies` | `[]` | 针对个别模型路由的精确 `{ provider, model, ...partialPolicy }` 覆盖。 |
|
|
80
80
|
| `auto` | `true` | 启用自动压缩与溢出恢复;设为 `false` 则仅手动执行。 |
|
|
81
81
|
|
|
@@ -109,18 +109,18 @@ kind: "package-reference"
|
|
|
109
109
|
|
|
110
110
|
- **一个测量服务为每个决策定价。** 单例 `ctx.tokenMeter` 会在同一个已消费日志 revision 上测量最新规范已记录 envelope 与当前表层。路由适配器声明请求图片定价时,meter 会将其应用于图片历史。压力、近期尾部保留、范围选择与缩减验证使用同一套路由定价的节点数值;已记录的替换影子价仍使用与路由无关的启发式规则,使纯投影 fold 保持一致。
|
|
111
111
|
- **日志记录的标记对就是事务。** 所有入口点共享一个先记录标记的区域事务:验证范围与活动锁,同步追加 `compaction/start`,准备并等待摘要,重新验证,再追加 `compaction/summary` 与替换,最后恰好进行一次闭合尝试。自动调用与显式范围调用要求数字标识的开放轮次归属与整个表层稳定;`compactNow()` 会预留空闲接纳,使用 `turn: null`,允许所选 span 之外追加仅追加上下文,flush 每次已闭合尝试,并在 `finally` 中释放接纳预留。
|
|
112
|
-
-
|
|
112
|
+
- **摘要保留复用 cache 的快速路径,并约束超大工作。** 能容纳的 one-shot 会逐字回放 surface 节点 0 处 `system/message` 所承载的系统提示词、上次已路由请求的工具与已遮蔽区域消息。层次调用会恰好一次重放同一个固定 system head,同时约束时序 map span 与递归 reduce。
|
|
113
113
|
- **`summarize()` 是唯一的子类钩子。** 基于模板或远程摘要器的子类可以覆盖它,同时压力、保留、被引用的源事件、缩减验证与已遮蔽 token 计量仍由 token meter 负责。
|
|
114
114
|
|
|
115
115
|
### 自动触发与溢出恢复
|
|
116
116
|
|
|
117
|
-
当 `auto: true` 时,串行 `agent/pre-step` listener 会在请求派生前检查压力:它通过 `ctx.tokenMeter` 为最新持久路由请求 envelope
|
|
117
|
+
当 `auto: true` 时,串行 `agent/pre-step` listener 会在请求派生前检查压力:它通过 `ctx.tokenMeter` 为最新持久路由请求 envelope 定价,当压力越过路由模型的阈值时,先剪枝,再在保留已定价近期尾部的同时摘要最旧的平衡范围。每个选定范围都从第一个不是 `system/message` 的 surface 节点开始,因此位于 surface 节点 0 的系统提示词永不会被遮蔽;由历史内提示词更新追加的后续 `system/message` 是普通历史,范围可以遮蔽它,agent loop 的投影随后会在二者文本不同时用当前提示词替换节点 0([决策规则](../../core/agent-loop/README.zh.md#understand-the-implementation))。`agent/request-error` listener 响应提供方确认的 `CONTEXT_WINDOW_EXCEEDED`:它绕过常规阈值与保留策略,尝试一次最大平衡头部缩减,并且只在表层替换 generation 前进后才授权重试。取消全程保持最终决定权。
|
|
118
118
|
|
|
119
119
|
压力策略从拥有持久路由的适配器解析容量。适配器无法为有效动态路由返回容量时,手动压力路径会抛出目标特定配置错误;自动 listener 会对该精确目标警告一次,并携带完整历史继续。
|
|
120
120
|
|
|
121
121
|
### 摘要机制
|
|
122
122
|
|
|
123
|
-
直接 `ctx.llm.stream()` 调用使用已配置的提供方/模型对与上限,回退到最新已记录请求目标,然后再回退到 `AgentOptions` 对,而不运行仅用于 agent loop 的 `agent/request`
|
|
123
|
+
直接 `ctx.llm.stream()` 调用使用已配置的提供方/模型对与上限,回退到最新已记录请求目标,然后再回退到 `AgentOptions` 对,而不运行仅用于 agent loop 的 `agent/request` 扩展点。完整请求能容纳时,该调用将 surface 节点 0 处派生的 `system/message` 作为 `messages` 的首项回放,后接已遮蔽区域消息(包括位于其 surface 位置的被遮蔽历史内 `system/message`),逐字携带 header 的工具,并将压缩指令作为最后一条 user 消息追加。超大输入或提供方拒绝的输入会被组合成工具配对平衡的时序 map span,再递归 reduce;每个 map 与 reduce 请求都把当前非空 system head 恰好一次作为首条消息重放,后续动态 `system/message` 仍是普通时序 map 输入,绝不会被提升或删除。每阶段预算都包含固定 head、可选工具与指令成本。只有 `replayTools: true` 时,工具 schema 才随层次调用发送。空内容系统头节点不贡献消息,但仍处于压缩范围之外。每次调用都将 `GenerateOptions.purpose` 设为 `compaction`;只有返回文本进入检查点,推理与工具调用都会被排除。图片输出会以 `UNSUPPORTED_CONTENT` 失败,而不是消失。替换 user 消息用 `<compacted-summary>` 标签框定摘要;最终原始摘要保留在 `compaction/summary` 事件上。
|
|
124
124
|
|
|
125
125
|
### 区域事务
|
|
126
126
|
|
|
@@ -136,7 +136,10 @@ kind: "package-reference"
|
|
|
136
136
|
|---|---|
|
|
137
137
|
| [`src/index.ts`](src/index.ts) | 插件入口:`BasicCompactionEngine`、自动 listener、入口点分发 |
|
|
138
138
|
| [`src/region.ts`](src/region.ts) | 保留选择与共享的先记录标记压缩事务 |
|
|
139
|
-
| [`src/summarizer.ts`](src/summarizer.ts) | 默认 `ctx.llm.stream()` 摘要、检查点框定、安全摘要投影 |
|
|
139
|
+
| [`src/summarizer.ts`](src/summarizer.ts) | 默认 `ctx.llm.stream()` one-shot 摘要、检查点框定、安全摘要投影 |
|
|
140
|
+
| [`src/hierarchical.ts`](src/hierarchical.ts) | 有界 map-reduce 回退、自适应拆分与阶段用量聚合 |
|
|
141
|
+
| [`src/hierarchical-planner.ts`](src/hierarchical-planner.ts) | 工具配对平衡单元与贪心 token 预算规划 |
|
|
142
|
+
| [`src/hierarchical-prompts.ts`](src/hierarchical-prompts.ts) | 结构化 map/reduce 提示词与输出验证 |
|
|
140
143
|
| [`src/config.ts`](src/config.ts) | 加载时验证与路由模型策略解析 |
|
|
141
144
|
| [`src/types.ts`](src/types.ts) | `BasicCompactionConfig` 与已解析策略词汇 |
|
|
142
145
|
| — | 不发布运行时不变式伴生入口;持久标记对可在会话日志中观察。 |
|
|
@@ -186,7 +189,7 @@ This is an automatically generated checkpoint condensing an earlier span of the
|
|
|
186
189
|
|
|
187
190
|
#### 模型看到的内容
|
|
188
191
|
|
|
189
|
-
|
|
192
|
+
完整请求能容纳时,摘要模型会接收逐字回放的会话:与上次已路由请求为已遮蔽区域发送的相同 system head、工具 schema 与消息,后面跟随一条最终 user 消息,即下方压缩指令。层次模式下,每个 map 请求接收恰好一次相同 system head、一个有序且工具配对平衡的源 span,以及结构化 map 指令;reduce 请求接收恰好一次相同 head、有序 `<partial-summary>` frame 和结构化 reduce 指令。后续 system 更新保留在原始 map 时序中,不会成为第二个固定 head。只有 `replayTools: true` 时,工具 schema 才随层次调用发送。会话模型绝不会看到这些私有请求或其推理;只有最终返回文本会被存储。
|
|
190
193
|
|
|
191
194
|
##### 压缩指令(最终 user 消息)
|
|
192
195
|
|
|
@@ -229,11 +232,11 @@ Rules:
|
|
|
229
232
|
|
|
230
233
|
#### Token 影响
|
|
231
234
|
|
|
232
|
-
|
|
235
|
+
能容纳的输入会产生一次独立模型调用:输入是已回放会话前缀加固定指令,输出受 `maxTokens` 限制。层次模式每个 map span 调用一次模型,再进行一次或多次 reduce,分别受 `mapMaxTokens` 与 `reduceMaxTokens` 限制;提供方确认的 overflow 可能在本地二分前增加失败尝试。收敛重试可能多次支付任一种成本。
|
|
233
236
|
|
|
234
237
|
#### KV Cache 影响
|
|
235
238
|
|
|
236
|
-
|
|
239
|
+
能容纳的 one-shot 会保留穿过已遮蔽区域的完整热前缀。层次调用会保留固定 system head,以及启用时工具 schema 的前缀复用;map payload 与 reduce frame 会在该共享前缀之后分叉。将摘要器路由到不同提供方/模型会放弃会话路由的 cache 复用。
|
|
237
240
|
|
|
238
241
|
## 已知限制与延期工作
|
|
239
242
|
|
|
@@ -244,9 +247,9 @@ Rules:
|
|
|
244
247
|
|
|
245
248
|
- **计量准确度取决于固定启发式规则**——可复用提供方用量缺失时,会回退到字符数加结构开销,而非精确的 token 化;只有在适配器声明了请求图片定价的路由上,图片出现处才携带提供方精确的视觉 token。
|
|
246
249
|
- **溢出分类由适配器维护**——提供方措辞可能改变;两个 DeepSeek 适配器将可识别的上下文限制失败规范化为 `CONTEXT_WINDOW_EXCEEDED`。
|
|
247
|
-
- **有界恢复需要摘要模型的容量元数据**——省略 `contextWindow` 的适配器继续走旧的一次性路径。该请求成功则行为不变;若溢出,层次无法推导安全 chunk 预算,并以可操作的容量错误失败。
|
|
248
|
-
- **层次输出是严格的检查点协议**——每个 map 与 reduce 阶段必须返回全部必需标题。截断、视觉输出、结构畸形、耗尽 `maxDepth`,或仍溢出的不可分源/部分摘要,都会让整次压缩事务失败,不安装部分检查点。
|
|
249
250
|
- **部分不可分单元与仅 envelope 溢出仍不在表层压缩范围内**——恢复无法缩减系统/工具/前缀、拆分不可分的非工具节点,或修复不可剪枝剩余部分仍超出窗口的工具单元。可选 pruner 可以缩减原本不可分工具对内的文本型工具结果主体。
|
|
251
|
+
- **层次模式要求声明摘要模型容量**——没有正整数 `contextWindow` 时,能容纳的输入仍可 one-shot 摘要;但 one-shot 被确认 overflow 后会清晰失败,因为有界分块规划没有可信窗口。
|
|
252
|
+
- **层次模式刻意有界**——提供方拒绝不可分的工具配对平衡 span、固定 system/tools/instruction 开销耗尽阶段预算、reduce 轮次未减少 partial 数量,或达到 `maxDepth` 时都会失败,而不会循环。
|
|
250
253
|
- **`compactRegion` 要求存在未结束的轮次**——在完全关闭的会话上手动调用会抛出异常(「no open turn」),而不是执行压缩。
|
|
251
254
|
- **摘要失败会保留最新持久表层**——任何替换前,自动路径会记录警告,并携带完整超预算历史继续。如果剪枝已落地,后续摘要失败会从该持久剪枝表层继续。因达到 `maxTokens` 而发生的摘要截断(隐藏推理 token 可能会耗尽该额度)遵循同一规则。
|
|
252
255
|
|
package/lib/index.js
CHANGED
|
@@ -337,7 +337,6 @@ async function summarizeWithLlm(ctx, config, input, agent, signal) {
|
|
|
337
337
|
provider: target.provider,
|
|
338
338
|
model: target.model,
|
|
339
339
|
messages,
|
|
340
|
-
...input.system === void 0 ? {} : { system: input.system },
|
|
341
340
|
...input.tools === void 0 ? {} : { tools: [...input.tools] },
|
|
342
341
|
maxTokens: config.maxTokens,
|
|
343
342
|
sessionId: agent.session.id,
|
|
@@ -345,7 +344,7 @@ async function summarizeWithLlm(ctx, config, input, agent, signal) {
|
|
|
345
344
|
...signal === void 0 ? {} : { signal }
|
|
346
345
|
};
|
|
347
346
|
for await (const chunk of ctx.llm.stream(options)) assembler.push(chunk);
|
|
348
|
-
const error = finishError
|
|
347
|
+
const error = finishError(assembler.finish);
|
|
349
348
|
if (error !== void 0) throw error;
|
|
350
349
|
const rawOutput = assembler.blocks();
|
|
351
350
|
const summary = summaryText(rawOutput);
|
|
@@ -378,8 +377,12 @@ function frameSummary(summary) {
|
|
|
378
377
|
}
|
|
379
378
|
];
|
|
380
379
|
}
|
|
381
|
-
/**
|
|
382
|
-
|
|
380
|
+
/**
|
|
381
|
+
* Map a terminal summarization finish to its fail-closed error.
|
|
382
|
+
* @param finish - terminal stream finish emitted by the summary request.
|
|
383
|
+
* @returns the corresponding error, or `undefined` for a complete stop.
|
|
384
|
+
*/
|
|
385
|
+
function finishError(finish) {
|
|
383
386
|
switch (finish.kind) {
|
|
384
387
|
case "error":
|
|
385
388
|
case "aborted": {
|
|
@@ -395,7 +398,11 @@ function finishError$1(finish) {
|
|
|
395
398
|
default: return;
|
|
396
399
|
}
|
|
397
400
|
}
|
|
398
|
-
/**
|
|
401
|
+
/**
|
|
402
|
+
* Reject visual output and keep only text before synthesizing a user message.
|
|
403
|
+
* @param blocks - raw content blocks emitted by the summary request.
|
|
404
|
+
* @returns the text-only blocks safe to persist as a compaction checkpoint.
|
|
405
|
+
*/
|
|
399
406
|
function summaryText(blocks) {
|
|
400
407
|
if (contentHasImage(blocks)) throw new LlmError("compaction summary cannot contain image output", "UNSUPPORTED_CONTENT");
|
|
401
408
|
return blocks.filter((block) => block.type === "text");
|
|
@@ -415,8 +422,21 @@ function summaryText(blocks) {
|
|
|
415
422
|
*/
|
|
416
423
|
var SurfaceChangedError = class extends Error {};
|
|
417
424
|
/**
|
|
418
|
-
*
|
|
419
|
-
*
|
|
425
|
+
* The `system/message` holding surface node 0, or `undefined` when another
|
|
426
|
+
* message-producing event starts the surface.
|
|
427
|
+
* @param session - session supplying the log behind the current surface.
|
|
428
|
+
* @param headSeq - seq at surface node 0 of a non-empty surface.
|
|
429
|
+
* @returns the system head event, or `undefined` without one.
|
|
430
|
+
*/
|
|
431
|
+
function systemHead(session, headSeq) {
|
|
432
|
+
const head = session.eventAt(headSeq);
|
|
433
|
+
return head.type === "system/message" ? head : void 0;
|
|
434
|
+
}
|
|
435
|
+
/**
|
|
436
|
+
* Resolve the next range starting at the first non-system surface node while
|
|
437
|
+
* retaining a priced recent tail and never splitting an assistant
|
|
438
|
+
* tool-call/result pair. A `system/message` at surface node 0 is never inside
|
|
439
|
+
* the range; without one the range starts at node 0.
|
|
420
440
|
* @param session - session supplying authoritative current surface positions.
|
|
421
441
|
* @param measurement - unified pressure and surface measurement from the conversation meter.
|
|
422
442
|
* @param retainTokens - minimum recent tail budget retained verbatim.
|
|
@@ -427,6 +447,7 @@ function selectCompactableRange(session, measurement, retainTokens) {
|
|
|
427
447
|
if (pricedNodes.length === 0) return null;
|
|
428
448
|
const surfaceNodes = session.surface.nodes;
|
|
429
449
|
if (surfaceNodes.length !== pricedNodes.length || surfaceNodes.some((seq, index) => seq !== pricedNodes[index]?.seq)) throw new Error("compaction: token-meter surface does not match the current session surface");
|
|
450
|
+
const firstIdx = systemHead(session, surfaceNodes[0]) === void 0 ? 0 : 1;
|
|
430
451
|
let accumulated = 0;
|
|
431
452
|
let keepFromIdx = pricedNodes.length;
|
|
432
453
|
for (let index = pricedNodes.length - 1; index >= 0; index -= 1) {
|
|
@@ -434,14 +455,14 @@ function selectCompactableRange(session, measurement, retainTokens) {
|
|
|
434
455
|
keepFromIdx = index;
|
|
435
456
|
if (accumulated >= retainTokens) break;
|
|
436
457
|
}
|
|
437
|
-
if (keepFromIdx
|
|
438
|
-
while (keepFromIdx >
|
|
458
|
+
if (keepFromIdx <= firstIdx) return null;
|
|
459
|
+
while (keepFromIdx > firstIdx) {
|
|
439
460
|
if (toolPairingBalancedBefore(session, surfaceNodes[keepFromIdx])) break;
|
|
440
461
|
keepFromIdx -= 1;
|
|
441
462
|
}
|
|
442
|
-
if (keepFromIdx
|
|
463
|
+
if (keepFromIdx <= firstIdx) return null;
|
|
443
464
|
return {
|
|
444
|
-
start: surfaceNodes[
|
|
465
|
+
start: surfaceNodes[firstIdx],
|
|
445
466
|
end: surfaceNodes[keepFromIdx - 1]
|
|
446
467
|
};
|
|
447
468
|
}
|
|
@@ -652,8 +673,8 @@ function commitCompactionBody(session, startEvent, summarized) {
|
|
|
652
673
|
session.append("user/message", checkpointMessage, {
|
|
653
674
|
surfaceOp: {
|
|
654
675
|
op: "replace",
|
|
655
|
-
start,
|
|
656
|
-
end
|
|
676
|
+
startSeq: start,
|
|
677
|
+
endSeq: end
|
|
657
678
|
},
|
|
658
679
|
sourceEventSeqs: [
|
|
659
680
|
startEvent.seq,
|
|
@@ -684,21 +705,24 @@ function completeCompaction(pending, endEvent) {
|
|
|
684
705
|
}
|
|
685
706
|
/**
|
|
686
707
|
* Reconstruct the last routed request's cacheable prefix for the shadowed
|
|
687
|
-
* region:
|
|
688
|
-
*
|
|
689
|
-
*
|
|
690
|
-
* and reuses the provider's
|
|
691
|
-
*
|
|
708
|
+
* region: the system prompt held by the `system/message` at surface node 0,
|
|
709
|
+
* the header's tool schemas, then the region's own derived messages in surface
|
|
710
|
+
* order. The summarizer appends only the compaction instruction after this, so
|
|
711
|
+
* the call is a genuine prefix of the conversation and reuses the provider's
|
|
712
|
+
* KV cache. A surface without a system head, or whose head projects to no
|
|
713
|
+
* message, contributes no leading system message.
|
|
714
|
+
* @param session - session supplying the surface head, request header, and per-node projection.
|
|
692
715
|
* @param shadowedSeqs - the surface-node seqs, in order, being compacted.
|
|
693
716
|
* @returns the replayed conversation prefix to condense.
|
|
694
717
|
*/
|
|
695
718
|
function buildSummarizationInput(session, shadowedSeqs) {
|
|
696
719
|
const header = session.requestHeader();
|
|
720
|
+
const head = systemHead(session, session.surface.nodes[0]);
|
|
721
|
+
const system = head === void 0 ? null : session.deriveEventMessage(head);
|
|
697
722
|
const regionMessages = shadowedSeqs.map((seq) => session.deriveEventMessage(session.eventAt(seq))).filter((message) => message !== null);
|
|
698
723
|
return {
|
|
699
|
-
...header?.system === void 0 ? {} : { system: header.system },
|
|
700
724
|
...header?.tools === void 0 ? {} : { tools: header.tools },
|
|
701
|
-
messages: regionMessages
|
|
725
|
+
messages: system === null ? regionMessages : [system, ...regionMessages]
|
|
702
726
|
};
|
|
703
727
|
}
|
|
704
728
|
/** Inspect open-turn, unmatched-compaction, and latest seed-boundary state independently. */
|
|
@@ -978,7 +1002,8 @@ var HierarchicalSummarizer = class {
|
|
|
978
1002
|
/* v8 ignore next -- LlmRuntime validates defined capacity before returning model info. */
|
|
979
1003
|
if (!Number.isSafeInteger(contextWindow) || contextWindow < 1) throw new Error(`hierarchical compaction: no positive integer context capacity for summary target ${target.provider}/${target.model}`);
|
|
980
1004
|
const estimate = (message) => this.ctx.tokenMeter.estimateMessage(message);
|
|
981
|
-
const
|
|
1005
|
+
const replay = hierarchyReplay(input.messages);
|
|
1006
|
+
const oneShotTokens = this.estimateCallInput(input, replay, COMPACTION_INSTRUCTION, true, estimate);
|
|
982
1007
|
let hadFailedLlmAttempt = false;
|
|
983
1008
|
if (oneShotTokens + target.oneShotMaxTokens <= contextWindow) try {
|
|
984
1009
|
return await oneShot();
|
|
@@ -988,10 +1013,10 @@ var HierarchicalSummarizer = class {
|
|
|
988
1013
|
}
|
|
989
1014
|
const inputBudget = Math.floor(contextWindow * this.hierarchy.chunkInputRatio);
|
|
990
1015
|
this.assertStageOutputReserve(contextWindow, inputBudget, this.hierarchy.mapMaxTokens, "map");
|
|
991
|
-
const totalUnits = toolBalancedUnits(
|
|
992
|
-
const mapReserve = this.estimateFixedInput(input, mapInstruction(totalUnits, totalUnits, totalUnits), this.hierarchy.replayTools, estimate);
|
|
1016
|
+
const totalUnits = toolBalancedUnits(replay.sourceMessages).length;
|
|
1017
|
+
const mapReserve = this.estimateFixedInput(input, replay.systemHead, mapInstruction(totalUnits, totalUnits, totalUnits), this.hierarchy.replayTools, estimate);
|
|
993
1018
|
const mapMessageBudget = this.messageBudget(inputBudget, mapReserve, "map");
|
|
994
|
-
const chunks = planMessageChunks(
|
|
1019
|
+
const chunks = planMessageChunks(replay.sourceMessages, mapMessageBudget, estimate);
|
|
995
1020
|
/* v8 ignore next -- stock range selection never submits an empty shadowed region. */
|
|
996
1021
|
if (chunks.length === 0) throw new Error("hierarchical compaction: oversized input produced no map chunks");
|
|
997
1022
|
const calls = [];
|
|
@@ -1003,10 +1028,7 @@ var HierarchicalSummarizer = class {
|
|
|
1003
1028
|
/* v8 ignore next -- the loop condition proves shift has an entry. */
|
|
1004
1029
|
if (span === void 0) break;
|
|
1005
1030
|
try {
|
|
1006
|
-
const result = await this.runStage(
|
|
1007
|
-
...input,
|
|
1008
|
-
messages: span.messages
|
|
1009
|
-
}, mapInstruction(span.start, span.end, totalUnits), target, this.hierarchy.mapMaxTokens, agent, signal);
|
|
1031
|
+
const result = await this.runStage(input, span.messages, replay.systemHead, mapInstruction(span.start, span.end, totalUnits), target, this.hierarchy.mapMaxTokens, agent, signal);
|
|
1010
1032
|
calls.push(result);
|
|
1011
1033
|
partials.push(this.partial(result, span.start, span.end, `map source units ${span.start}-${span.end}`));
|
|
1012
1034
|
} catch (error) {
|
|
@@ -1023,7 +1045,7 @@ var HierarchicalSummarizer = class {
|
|
|
1023
1045
|
let usedReduce = false;
|
|
1024
1046
|
for (let round = 1; partials.length > 1; round += 1) {
|
|
1025
1047
|
if (round > this.hierarchy.maxDepth) throw new Error(`hierarchical compaction: reduction did not converge within ${this.hierarchy.maxDepth} round(s)`);
|
|
1026
|
-
const reduceReserve = this.estimateFixedInput(input, reduceInstruction(round, totalUnits, totalUnits, totalUnits), this.hierarchy.replayTools, estimate);
|
|
1048
|
+
const reduceReserve = this.estimateFixedInput(input, replay.systemHead, reduceInstruction(round, totalUnits, totalUnits, totalUnits), this.hierarchy.replayTools, estimate);
|
|
1027
1049
|
const reduceMessageBudget = this.messageBudget(inputBudget, reduceReserve, `reduce round ${round}`);
|
|
1028
1050
|
let groups;
|
|
1029
1051
|
try {
|
|
@@ -1044,10 +1066,7 @@ var HierarchicalSummarizer = class {
|
|
|
1044
1066
|
/* v8 ignore next -- the loop condition proves shift has an entry. */
|
|
1045
1067
|
if (span === void 0) break;
|
|
1046
1068
|
try {
|
|
1047
|
-
const result = await this.runStage(
|
|
1048
|
-
...input,
|
|
1049
|
-
messages: span.messages
|
|
1050
|
-
}, reduceInstruction(round, span.start, span.end, totalUnits), target, this.hierarchy.reduceMaxTokens, agent, signal);
|
|
1069
|
+
const result = await this.runStage(input, span.messages, replay.systemHead, reduceInstruction(round, span.start, span.end, totalUnits), target, this.hierarchy.reduceMaxTokens, agent, signal);
|
|
1051
1070
|
calls.push(result);
|
|
1052
1071
|
next.push(this.partial(result, span.start, span.end, `reduce round ${round} source units ${span.start}-${span.end}`));
|
|
1053
1072
|
} catch (error) {
|
|
@@ -1171,12 +1190,12 @@ var HierarchicalSummarizer = class {
|
|
|
1171
1190
|
if (inputBudget + outputTokens > contextWindow) throw new Error(`hierarchical compaction: ${stage} input budget ${inputBudget} plus output reserve ${outputTokens} exceeds summary context ${contextWindow}`);
|
|
1172
1191
|
}
|
|
1173
1192
|
/** Price a complete auxiliary call input. */
|
|
1174
|
-
estimateCallInput(input, instruction, includeTools, estimate) {
|
|
1175
|
-
return this.estimateFixedInput(input, instruction, includeTools, estimate) + estimateMessages(
|
|
1193
|
+
estimateCallInput(input, replay, instruction, includeTools, estimate) {
|
|
1194
|
+
return this.estimateFixedInput(input, replay.systemHead, instruction, includeTools, estimate) + estimateMessages(replay.sourceMessages, estimate);
|
|
1176
1195
|
}
|
|
1177
|
-
/** Price the repeated
|
|
1178
|
-
estimateFixedInput(input, instruction, includeTools, estimate) {
|
|
1179
|
-
return (
|
|
1196
|
+
/** Price the repeated system head, optional tools, and final stage instruction. */
|
|
1197
|
+
estimateFixedInput(input, systemHead, instruction, includeTools, estimate) {
|
|
1198
|
+
return (systemHead === void 0 ? 0 : estimate(systemHead)) + (!includeTools || input.tools === void 0 || input.tools.length === 0 ? 0 : Math.ceil(JSON.stringify(input.tools).length / CHARS_PER_TOKEN) + ENVELOPE_OVERHEAD) + estimate(this.instructionMessage(instruction));
|
|
1180
1199
|
}
|
|
1181
1200
|
/** Derive positive room for stage messages after fixed input. */
|
|
1182
1201
|
messageBudget(inputBudget, fixedTokens, stage) {
|
|
@@ -1185,14 +1204,17 @@ var HierarchicalSummarizer = class {
|
|
|
1185
1204
|
return budget;
|
|
1186
1205
|
}
|
|
1187
1206
|
/** Run one private map or reduce model call and require structured text. */
|
|
1188
|
-
async runStage(input, instruction, target, maxTokens, agent, signal) {
|
|
1207
|
+
async runStage(input, messages, systemHead, instruction, target, maxTokens, agent, signal) {
|
|
1189
1208
|
signal?.throwIfAborted();
|
|
1190
1209
|
const assembler = new BlockAssembler();
|
|
1191
1210
|
const options = {
|
|
1192
1211
|
provider: target.provider,
|
|
1193
1212
|
model: target.model,
|
|
1194
|
-
messages: [
|
|
1195
|
-
|
|
1213
|
+
messages: [
|
|
1214
|
+
...systemHead === void 0 ? [] : [systemHead],
|
|
1215
|
+
...messages,
|
|
1216
|
+
this.instructionMessage(instruction)
|
|
1217
|
+
],
|
|
1196
1218
|
...this.hierarchy.replayTools && input.tools !== void 0 ? { tools: [...input.tools] } : {},
|
|
1197
1219
|
maxTokens,
|
|
1198
1220
|
sessionId: agent.session.id,
|
|
@@ -1203,8 +1225,7 @@ var HierarchicalSummarizer = class {
|
|
|
1203
1225
|
const finishFailure = finishError(assembler.finish);
|
|
1204
1226
|
if (finishFailure !== void 0) throw finishFailure;
|
|
1205
1227
|
const rawOutput = assembler.blocks();
|
|
1206
|
-
|
|
1207
|
-
const summary = rawOutput.filter((block) => block.type === "text");
|
|
1228
|
+
const summary = summaryText(rawOutput);
|
|
1208
1229
|
validateStructuredSummary(summary, "hierarchical compaction stage");
|
|
1209
1230
|
return {
|
|
1210
1231
|
summary,
|
|
@@ -1244,6 +1265,15 @@ var HierarchicalSummarizer = class {
|
|
|
1244
1265
|
});
|
|
1245
1266
|
}
|
|
1246
1267
|
};
|
|
1268
|
+
/** Separate the fixed surface system head from chronological source history. */
|
|
1269
|
+
function hierarchyReplay(messages) {
|
|
1270
|
+
const [first, ...rest] = messages;
|
|
1271
|
+
if (first?.role === "system") return {
|
|
1272
|
+
systemHead: first,
|
|
1273
|
+
sourceMessages: rest
|
|
1274
|
+
};
|
|
1275
|
+
return { sourceMessages: messages };
|
|
1276
|
+
}
|
|
1247
1277
|
/** Build the terminal diagnostic for a provider-rejected atomic span. */
|
|
1248
1278
|
function indivisibleOverflow(stage, cause) {
|
|
1249
1279
|
const error = new OversizedCompactionUnitError(`hierarchical compaction: ${stage} still exceeds the provider context window and is indivisible`, { cause });
|
|
@@ -1254,23 +1284,6 @@ function indivisibleOverflow(stage, cause) {
|
|
|
1254
1284
|
function hasErrorCode(error, code) {
|
|
1255
1285
|
return typeof error === "object" && error !== null && "code" in error && error.code === code;
|
|
1256
1286
|
}
|
|
1257
|
-
/** Map a terminal stage finish to a fail-closed error. */
|
|
1258
|
-
function finishError(finish) {
|
|
1259
|
-
switch (finish.kind) {
|
|
1260
|
-
case "error":
|
|
1261
|
-
case "aborted": {
|
|
1262
|
-
const error = new Error(finish.failure.message);
|
|
1263
|
-
error.code = finish.failure.code;
|
|
1264
|
-
return error;
|
|
1265
|
-
}
|
|
1266
|
-
case "max-tokens": {
|
|
1267
|
-
const error = /* @__PURE__ */ new Error("hierarchical compaction stage truncated at the token cap");
|
|
1268
|
-
error.code = "MAX_TOKENS";
|
|
1269
|
-
return error;
|
|
1270
|
-
}
|
|
1271
|
-
default: return;
|
|
1272
|
-
}
|
|
1273
|
-
}
|
|
1274
1287
|
/**
|
|
1275
1288
|
* Sum disjoint provider usage across every successful map and reduce call.
|
|
1276
1289
|
* @param usages - stage usage values in call order.
|
package/lib/types/region.d.ts
CHANGED
|
@@ -25,8 +25,10 @@ interface CompactionTransactionOptions {
|
|
|
25
25
|
readonly sourceCommandId?: CommandId;
|
|
26
26
|
}
|
|
27
27
|
/**
|
|
28
|
-
* Resolve the next
|
|
29
|
-
* and never splitting an assistant
|
|
28
|
+
* Resolve the next range starting at the first non-system surface node while
|
|
29
|
+
* retaining a priced recent tail and never splitting an assistant
|
|
30
|
+
* tool-call/result pair. A `system/message` at surface node 0 is never inside
|
|
31
|
+
* the range; without one the range starts at node 0.
|
|
30
32
|
* @param session - session supplying authoritative current surface positions.
|
|
31
33
|
* @param measurement - unified pressure and surface measurement from the conversation meter.
|
|
32
34
|
* @param retainTokens - minimum recent tail budget retained verbatim.
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
* @module @deepseek-ai/dsh-compaction-basic/summarizer
|
|
5
5
|
*/
|
|
6
6
|
import type { Context } from '@deepseek-ai/cordis';
|
|
7
|
-
import type { ContentBlock, Message, TokenUsage, ToolSchema } from '@deepseek-ai/dsh-llm';
|
|
7
|
+
import type { ContentBlock, FinishReason, Message, TokenUsage, ToolSchema } from '@deepseek-ai/dsh-llm';
|
|
8
8
|
import type { Agent } from '@deepseek-ai/dsh-agent';
|
|
9
9
|
interface SummaryConfig {
|
|
10
10
|
readonly summarizationProvider: string;
|
|
@@ -26,11 +26,9 @@ export declare const COMPACTION_INSTRUCTION: string;
|
|
|
26
26
|
* compaction instruction is then the only novel input.
|
|
27
27
|
*/
|
|
28
28
|
export interface SummarizationInput {
|
|
29
|
-
/** The conversation's own system prompt, reused for prefix-cache alignment; absent for a system-less request. */
|
|
30
|
-
readonly system?: string;
|
|
31
29
|
/** The conversation's tool schemas, reused for prefix-cache alignment; absent when the request carried none. */
|
|
32
30
|
readonly tools?: readonly ToolSchema[];
|
|
33
|
-
/** The
|
|
31
|
+
/** The derived system head, when present, followed by the shadowed region in surface order. */
|
|
34
32
|
readonly messages: readonly Message[];
|
|
35
33
|
}
|
|
36
34
|
/** Safe summary content plus the exact auxiliary call envelope recorded with it. */
|
|
@@ -70,5 +68,19 @@ export declare function summarizeWithLlm(ctx: Context, config: SummaryConfig, in
|
|
|
70
68
|
* @returns content for the synthesized replacement user message.
|
|
71
69
|
*/
|
|
72
70
|
export declare function frameSummary(summary: readonly ContentBlock[]): ContentBlock[];
|
|
71
|
+
/**
|
|
72
|
+
* Map a terminal summarization finish to its fail-closed error.
|
|
73
|
+
* @param finish - terminal stream finish emitted by the summary request.
|
|
74
|
+
* @returns the corresponding error, or `undefined` for a complete stop.
|
|
75
|
+
*/
|
|
76
|
+
export declare function finishError(finish: FinishReason): Error | undefined;
|
|
77
|
+
/**
|
|
78
|
+
* Reject visual output and keep only text before synthesizing a user message.
|
|
79
|
+
* @param blocks - raw content blocks emitted by the summary request.
|
|
80
|
+
* @returns the text-only blocks safe to persist as a compaction checkpoint.
|
|
81
|
+
*/
|
|
82
|
+
export declare function summaryText(blocks: readonly ContentBlock[]): Array<Extract<ContentBlock, {
|
|
83
|
+
type: 'text';
|
|
84
|
+
}>>;
|
|
73
85
|
export {};
|
|
74
86
|
//# sourceMappingURL=summarizer.d.ts.map
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@crazx/dsh-compaction-basic",
|
|
3
3
|
"description": "Token-meter-driven compaction policy and LLM summarization backend for the DeepSeek Harness",
|
|
4
|
-
"version": "0.1.
|
|
4
|
+
"version": "0.1.5-alpha.1.zw.1",
|
|
5
5
|
"publishConfig": {
|
|
6
6
|
"access": "public"
|
|
7
7
|
},
|
|
@@ -28,13 +28,13 @@
|
|
|
28
28
|
"license": "MIT",
|
|
29
29
|
"peerDependencies": {
|
|
30
30
|
"@deepseek-ai/cordis": "^4.0.2",
|
|
31
|
-
"@deepseek-ai/dsh-agent": "^0.1.
|
|
32
|
-
"@deepseek-ai/dsh-commands": "^0.1.
|
|
33
|
-
"@deepseek-ai/dsh-compaction": "^0.1.
|
|
34
|
-
"@deepseek-ai/dsh-compaction-tool-result-pruner": "^0.1.
|
|
35
|
-
"@deepseek-ai/dsh-llm": "^0.1.
|
|
36
|
-
"@deepseek-ai/dsh-session": "^0.1.
|
|
37
|
-
"@deepseek-ai/dsh-token-meter": "^0.1.
|
|
31
|
+
"@deepseek-ai/dsh-agent": "^0.1.5-alpha.1",
|
|
32
|
+
"@deepseek-ai/dsh-commands": "^0.1.5-alpha.1",
|
|
33
|
+
"@deepseek-ai/dsh-compaction": "^0.1.5-alpha.1",
|
|
34
|
+
"@deepseek-ai/dsh-compaction-tool-result-pruner": "^0.1.5-alpha.1",
|
|
35
|
+
"@deepseek-ai/dsh-llm": "^0.1.5-alpha.1",
|
|
36
|
+
"@deepseek-ai/dsh-session": "^0.1.5-alpha.1",
|
|
37
|
+
"@deepseek-ai/dsh-token-meter": "^0.1.5-alpha.1"
|
|
38
38
|
},
|
|
39
39
|
"peerDependenciesMeta": {
|
|
40
40
|
"@deepseek-ai/dsh-compaction-tool-result-pruner": {
|
|
@@ -42,25 +42,25 @@
|
|
|
42
42
|
}
|
|
43
43
|
},
|
|
44
44
|
"dependencies": {
|
|
45
|
-
"@deepseek-ai/dsh-util-values": "^0.1.
|
|
45
|
+
"@deepseek-ai/dsh-util-values": "^0.1.5-alpha.1",
|
|
46
46
|
"@deepseek-ai/schemastery": "^3.18.2"
|
|
47
47
|
},
|
|
48
48
|
"devDependencies": {
|
|
49
49
|
"@deepseek-ai/cordis": "^4.0.2",
|
|
50
50
|
"@deepseek-ai/cordis-plugin-include": "^1.0.7",
|
|
51
51
|
"@deepseek-ai/cordis-plugin-loader": "^1.0.3",
|
|
52
|
-
"@deepseek-ai/dsh-agent": "^0.1.
|
|
53
|
-
"@deepseek-ai/dsh-agent-loop": "^0.1.
|
|
54
|
-
"@deepseek-ai/dsh-agent-loop-testkit": "^0.1.
|
|
55
|
-
"@deepseek-ai/dsh-commands": "^0.1.
|
|
56
|
-
"@deepseek-ai/dsh-compaction": "^0.1.
|
|
57
|
-
"@deepseek-ai/dsh-compaction-tool-result-pruner": "^0.1.
|
|
58
|
-
"@deepseek-ai/dsh-invariants": "^0.1.
|
|
59
|
-
"@deepseek-ai/dsh-llm": "^0.1.
|
|
60
|
-
"@deepseek-ai/dsh-llm-retry": "^0.1.
|
|
61
|
-
"@deepseek-ai/dsh-session": "^0.1.
|
|
62
|
-
"@deepseek-ai/dsh-session-projection": "^0.1.
|
|
63
|
-
"@deepseek-ai/dsh-token-meter": "^0.1.
|
|
64
|
-
"@deepseek-ai/dsh-tools": "^0.1.
|
|
52
|
+
"@deepseek-ai/dsh-agent": "^0.1.5-alpha.1",
|
|
53
|
+
"@deepseek-ai/dsh-agent-loop": "^0.1.5-alpha.1",
|
|
54
|
+
"@deepseek-ai/dsh-agent-loop-testkit": "^0.1.5-alpha.1",
|
|
55
|
+
"@deepseek-ai/dsh-commands": "^0.1.5-alpha.1",
|
|
56
|
+
"@deepseek-ai/dsh-compaction": "^0.1.5-alpha.1",
|
|
57
|
+
"@deepseek-ai/dsh-compaction-tool-result-pruner": "^0.1.5-alpha.1",
|
|
58
|
+
"@deepseek-ai/dsh-invariants": "^0.1.5-alpha.1",
|
|
59
|
+
"@deepseek-ai/dsh-llm": "^0.1.5-alpha.1",
|
|
60
|
+
"@deepseek-ai/dsh-llm-retry": "^0.1.5-alpha.1",
|
|
61
|
+
"@deepseek-ai/dsh-session": "^0.1.5-alpha.1",
|
|
62
|
+
"@deepseek-ai/dsh-session-projection": "^0.1.5-alpha.1",
|
|
63
|
+
"@deepseek-ai/dsh-token-meter": "^0.1.5-alpha.1",
|
|
64
|
+
"@deepseek-ai/dsh-tools": "^0.1.5-alpha.1"
|
|
65
65
|
}
|
|
66
66
|
}
|