@crazx/dsh-compaction-basic 0.1.2-rc.1.zw.1 → 0.1.5-alpha.1.zw.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.i18n.yaml CHANGED
@@ -2,5 +2,5 @@
2
2
  # side as of the last confirmed-consistent state. Both languages carry equal authority;
3
3
  # after editing either side, bring the other along and re-record with:
4
4
  # pnpm run verify-translation-pairing --write packages/compaction/compaction-basic/README.md
5
- README.md: f57fcc16ab27d461e4371506af32d09c7db2a698
6
- README.zh.md: a71a63f1973d87096d1ea0d94353c6e74cdb10e2
5
+ README.md: d70bc758c5bba4a9791a1f5017beb80783b1017e
6
+ README.zh.md: 086dc4a49d423daa661a0521e50be51db8db7872
package/README.md CHANGED
@@ -9,7 +9,7 @@ English | [中文](README.zh.md)
9
9
 
10
10
  ## Summary
11
11
 
12
- `dsh-compaction-basic` keeps long agent conversations working near the model's context limit. As token pressure builds, it automatically condenses the oldest part of the conversation into a summary and keeps the recent part intact; after a context-overflow error it condenses and retries. You can also condense on demand with `/compact` from `dsh-command-compact`, and mount `dsh-compaction-tool-result-pruner` to trim oversized tool outputs first. Condensation costs one extra model request that reads the selected history and writes the summary; only the summary text is kept. It condenses derived history only — it cannot shrink the system prompt, tools, or session prefix, and one indivisible unit such as a single huge tool call cannot be split.
12
+ This package keeps long agent conversations working near the model's context limit. As token pressure builds, it condenses the oldest history into a summary while preserving recent messages; after a context-overflow error, it condenses and retries. You can also request condensation with `/compact` and optionally trim oversized tool outputs first. Fitting input uses one extra model request; oversized input uses bounded hierarchical map-reduce calls, and only the final summary text is retained. It cannot reduce the system prompt, tools, or session prefix, or split one indivisible unit such as a single huge tool call.
13
13
 
14
14
  ## Table of Contents
15
15
 
@@ -109,18 +109,18 @@ The backend is built on four commitments:
109
109
 
110
110
  - **One measurement service prices every decision.** The singleton `ctx.tokenMeter` measures the latest canonical logged envelope and current surface at one consumed-log revision. When the routed adapter declares request-image pricing, the meter applies it to image history. Pressure, recent-tail retention, range selection, and shrink validation use the same route-priced node figures; logged replacement shadow prices stay on the route-independent heuristic so pure projection folds remain consistent.
111
111
  - **The log-recorded bracket is the transaction.** All entry points share one bracket-first region transaction: validate the range and live lock, append `compaction/start` synchronously, prepare and await the summary, revalidate, append `compaction/summary` plus the replacement, and make exactly one closing attempt. Automatic and explicit-region calls require a numeric open-turn owner and whole-surface stability; `compactNow()` reserves idle admission, uses `turn: null`, accepts append-only context outside its selected span, flushes every closed attempt, and releases admission in `finally`.
112
- - **Summarization reuses the provider's warm prefix.** Replaying the last routed request's system prompt, tools, and shadowed-region messages byte-for-byte makes the auxiliary call a genuine prefix of the conversation, so only the trailing instruction and the summary output are uncached.
112
+ - **Summarization preserves the cache-reusing fast path and bounds oversized work.** A fitting one-shot replays the system prompt held by the `system/message` at surface node 0, the last routed request's tools, and the shadowed-region messages byte-for-byte. Hierarchical calls replay that same fixed system head exactly once while bounding chronological map spans and recursive reductions.
113
113
  - **`summarize()` is the sole subclass hook.** A template- or remote-summarizer subclass can override it while pressure, retention, cited source events, shrink validation, and shadowed-token accounting stay on the token meter.
114
114
 
115
115
  ### Automatic triggers and overflow recovery
116
116
 
117
- With `auto: true`, a serial `agent/pre-step` listener checks pressure before request derivation: it prices the latest durable routed request envelope through `ctx.tokenMeter`, and when pressure crosses the routed model's threshold it prunes, then summarizes the oldest balanced span while keeping a priced recent tail. The `agent/request-error` listener reacts to a provider-confirmed `CONTEXT_WINDOW_EXCEEDED`: it bypasses the normal threshold and retention policy, attempts one maximal balanced head reduction, and authorizes a retry only after the surface replacement generation advances. Cancellation stays authoritative throughout.
117
+ With `auto: true`, a serial `agent/pre-step` listener checks pressure before request derivation: it prices the latest durable routed request envelope through `ctx.tokenMeter`, and when pressure crosses the routed model's threshold it prunes, then summarizes the oldest balanced span while keeping a priced recent tail. Every selected range starts at the first surface node that is not a `system/message`, so a system prompt at surface node 0 is never shadowed; a later `system/message` appended by an in-history prompt update is ordinary history that the range may shadow, and the agent loop's projection then replaces node 0 with the current prompt when their text differs ([decision rule](../../core/agent-loop/README.md#understand-the-implementation)). The `agent/request-error` listener reacts to a provider-confirmed `CONTEXT_WINDOW_EXCEEDED`: it bypasses the normal threshold and retention policy, attempts one maximal balanced head reduction, and authorizes a retry only after the surface replacement generation advances. Cancellation stays authoritative throughout.
118
118
 
119
119
  Pressure policy resolves capacity from the adapter that owns the durable route. An adapter that returns no capacity for a valid dynamic route makes the manual pressure path throw a target-specific configuration error; the automatic listener warns once for that exact target and continues with full history.
120
120
 
121
121
  ### Summarization mechanics
122
122
 
123
- A direct `ctx.llm.stream()` call uses the configured provider/model pair and cap, falling back to the latest logged request target and then the `AgentOptions` pair, without running the loop-only `agent/request` extension point. The call replays the conversation's own system prompt, tools, and shadowed-region messages verbatim — including image references, which the selected adapter must resolve or explicitly reject — and appends the compaction instruction as the final user message, so it reuses the provider's warm prefix cache instead of invalidating it. The call sets `GenerateOptions.purpose` to `compaction`; only returned text enters the checkpoint, excluding reasoning and tool calls. Image output fails with `UNSUPPORTED_CONTENT` rather than disappearing. The replacement user message frames the summary with `<compacted-summary>` tags; the raw summary remains on the `compaction/summary` event.
123
+ A direct `ctx.llm.stream()` call uses the configured provider/model pair and cap, falling back to the latest logged request target and then the `AgentOptions` pair, without running the loop-only `agent/request` extension point. When the complete request fits, the call replays the derived `system/message` at surface node 0 as the leading entry of `messages`, followed by the shadowed-region messages (including a shadowed in-history `system/message` in its surface position), carries the header's tools verbatim, and appends the compaction instruction as the final user message. Oversized or Provider-rejected input is grouped into tool-balanced chronological map spans and recursively reduced; every map and reduce request replays the current non-empty system head exactly once as its first message, while later dynamic `system/message` entries remain ordinary chronological map input and are never promoted or deleted. The fixed-head, optional-tool, and instruction cost is included in every stage budget. Tool schemas accompany hierarchy calls only when `replayTools: true`. An empty-content system head contributes no message but remains outside the compacted range. Every call sets `GenerateOptions.purpose` to `compaction`; only returned text enters the checkpoint, excluding reasoning and tool calls. Image output fails with `UNSUPPORTED_CONTENT` rather than disappearing. The replacement user message frames the summary with `<compacted-summary>` tags; the raw final summary remains on the `compaction/summary` event.
124
124
 
125
125
  ### The region transaction
126
126
 
@@ -136,7 +136,10 @@ The transaction validates the surface span and the durable lock, appends `compac
136
136
  |---|---|
137
137
  | [`src/index.ts`](src/index.ts) | Plugin entry: `BasicCompactionEngine`, automatic listeners, entry-point dispatch |
138
138
  | [`src/region.ts`](src/region.ts) | Retention selection and the shared bracket-first compaction transaction |
139
- | [`src/summarizer.ts`](src/summarizer.ts) | Default `ctx.llm.stream()` summarization, checkpoint framing, safe-summary projection |
139
+ | [`src/summarizer.ts`](src/summarizer.ts) | Default `ctx.llm.stream()` one-shot summarization, checkpoint framing, safe-summary projection |
140
+ | [`src/hierarchical.ts`](src/hierarchical.ts) | Bounded map-reduce fallback, adaptive splitting, stage usage aggregation |
141
+ | [`src/hierarchical-planner.ts`](src/hierarchical-planner.ts) | Tool-balanced units and greedy token-budget planning |
142
+ | [`src/hierarchical-prompts.ts`](src/hierarchical-prompts.ts) | Structured map/reduce prompts and output validation |
140
143
  | [`src/config.ts`](src/config.ts) | Load-time validation and routed-model policy resolution |
141
144
  | [`src/types.ts`](src/types.ts) | `BasicCompactionConfig` and resolved policy vocabulary |
142
145
  | — | No runtime invariant companion is published; this package exposes no independent event sequence or mutable data relation beyond contracts enforced at its owning seam. |
@@ -186,7 +189,7 @@ Replacing rather than append-only. Each checkpoint invalidates reuse from the fi
186
189
 
187
190
  #### What the model sees
188
191
 
189
- When the complete request fits, the summarization model receives the conversation replayed verbatim — the same system prompt, tool schemas, and messages the last routed request sent for the shadowed region — followed by one final user message: the compaction instruction below. For hierarchy, each map request receives the same system prompt, an ordered tool-balanced source span, and a structured map instruction; reduce requests receive ordered `<partial-summary>` frames and a structured reduce instruction. Tool schemas accompany hierarchy calls only when `replayTools: true`. The conversation model never sees these private requests or their reasoning; only the final text is stored.
192
+ When the complete request fits, the summarization model receives the conversation replayed verbatim — the same system head, tool schemas, and messages the last routed request sent for the shadowed region — followed by one final user message: the compaction instruction below. For hierarchy, each map request receives the same system head exactly once, an ordered tool-balanced source span, and a structured map instruction; reduce requests receive that same head exactly once, ordered `<partial-summary>` frames, and a structured reduce instruction. Later system updates remain in their original map chronology instead of becoming a second fixed head. Tool schemas accompany hierarchy calls only when `replayTools: true`. The conversation model never sees these private requests or their reasoning; only the final text is stored.
190
193
 
191
194
  ##### Compaction instruction (final user message)
192
195
 
@@ -233,7 +236,7 @@ A fitting input costs one separate model call: the replayed conversation prefix
233
236
 
234
237
  #### KV Cache effect
235
238
 
236
- The fitting one-shot request matches the conversation's replayed system prompt, tools, and shadowed-region messages byte-for-byte, so the provider's warm prefix cache is reused up to the trailing instruction. Routing to another model or compacting a non-head range forgoes that reuse. Hierarchy intentionally bounds each call and therefore cannot preserve one full warm prefix: map calls may reuse their leading system/message prefix where the Provider permits, while reduce calls operate on newly generated partials. `replayTools: false` also omits the tool-schema prefix to leave more room for source messages.
239
+ The fitting one-shot preserves the complete warm prefix through the shadowed region. Hierarchical calls preserve prefix reuse for the fixed system head and, when enabled, tool schemas; map payloads and reduction frames diverge after that shared prefix. Routing the summarizer to a different provider/model forgoes the conversation route's cache reuse.
237
240
 
238
241
  ## Known Limitations and Deferred Work
239
242
 
@@ -244,9 +247,9 @@ These limits define when automatic condensation is a poor fit or needs special c
244
247
 
245
248
  - **Meter accuracy follows the fixed heuristic** — missing reusable provider usage falls back to character count plus structural overhead rather than exact tokenization; image occurrences carry provider-exact visual tokens only on routes whose adapter declares request-image pricing.
246
249
  - **Overflow classification is adapter-maintained** — provider wording can change; both DeepSeek adapters normalize recognized context-limit failures to `CONTEXT_WINDOW_EXCEEDED`.
247
- - **Bounded recovery requires summary-model capacity metadata** — an adapter that omits `contextWindow` keeps the legacy one-shot path. If that request succeeds, behavior is unchanged; if it overflows, hierarchy cannot derive safe chunk budgets and fails with an actionable capacity error.
248
- - **Hierarchy output is a strict checkpoint protocol** — every map and reduce stage must return all required headings. Truncation, visual output, malformed structure, exhausted `maxDepth`, or an indivisible source/partial that still overflows fails the complete compaction transaction without installing a partial checkpoint.
249
250
  - **Some indivisible-unit and envelope-only overflow remains outside surface compaction** — recovery cannot shrink system/tools/prefix, split an indivisible non-tool node, or repair a tool unit whose non-prunable remainder still exceeds the window. The optional pruner can shrink text-bearing tool-result bulk inside an otherwise indivisible pair.
251
+ - **Hierarchy requires declared summary-model capacity** — without a positive integer `contextWindow`, fitting input still uses one-shot summarization, but a confirmed one-shot overflow fails clearly because bounded chunk planning has no trustworthy window.
252
+ - **Hierarchy is intentionally bounded** — a Provider-rejected indivisible tool-balanced span, fixed system/tools/instruction overhead that exhausts the stage budget, a reduction round that does not reduce partial count, or reaching `maxDepth` fails instead of looping.
250
253
  - **`compactRegion` requires an open turn** — a manual call on a fully-closed session throws ("no open turn") rather than compacting.
251
254
  - **Summarization failure preserves the latest durable surface** — before any replacement, the auto path logs a warning and proceeds with full over-budget history. If pruning already landed, a later summarization failure proceeds from that durable pruned surface. Summarization truncation at `maxTokens`, which hidden reasoning tokens can consume, follows the same rule.
252
255
 
package/README.zh.md CHANGED
@@ -9,7 +9,7 @@ kind: "package-reference"
9
9
 
10
10
  ## 概述
11
11
 
12
- `dsh-compaction-basic` 让长时 agent 会话在接近模型上下文上限时仍能正常工作。token 压力上升时,它会自动把对话最旧的部分压缩为摘要,并保持近期部分完整;上下文溢出错误发生后,它会压缩并重试。你也可以通过 `dsh-command-compact` 的 `/compact` 按需压缩,并挂载 `dsh-compaction-tool-result-pruner` 先修剪超大工具输出。压缩的代价是一次额外的模型请求,它读取所选历史并写出摘要;只有摘要文本会被保留。它只压缩派生历史——无法缩减系统提示词、工具或会话前缀,也无法拆分单个不可分单元(例如一次超大工具调用)。
12
+ 本包让长时 agent 会话在接近模型上下文上限时仍能正常工作。token 压力上升时,它会把最旧的历史压缩为摘要并保留近期消息;上下文溢出错误发生后,它会压缩并重试。你也可以通过 `/compact` 按需压缩,并选择先修剪超大工具输出。能容纳的输入使用一次额外模型请求;超大输入使用有界层次 map-reduce 调用,并且只保留最终摘要文本。它无法缩减系统提示词、工具或会话前缀,也无法拆分单个不可分单元(例如一次超大工具调用)。
13
13
 
14
14
  ## 目录
15
15
 
@@ -29,7 +29,7 @@ kind: "package-reference"
29
29
 
30
30
  ### 你会得到什么
31
31
 
32
- 默认设置下你会获得四种行为:会话向模型上下文上限增长时自动压缩;提供方确认上下文溢出错误后的恢复(先压缩再重试该请求);通过 `/compact` 命令按需压缩;以及——挂载修剪器时——压缩前对超大工具输出的修剪。
32
+ 默认设置下你会获得四种行为:会话向模型上下文上限增长时自动压缩;提供方确认上下文溢出错误后的恢复(先压缩再重试该请求);通过 `/compact` 命令按需压缩;以及——挂载修剪器时——压缩前对超大工具输出的修剪。能放入摘要模型窗口的输入保留复用 cache 的 one-shot 路径;超大输入或提供方拒绝的输入回退到有界时序 map-reduce。
33
33
 
34
34
  ### 最小可用组合
35
35
 
@@ -71,11 +71,11 @@ kind: "package-reference"
71
71
  | `maxTokens` | `8192` | 摘要请求的输出上限;可包含推理 token。 |
72
72
  | `compactionRetries` | `1` | 压力仍高于阈值时,在首次压缩后进行的额外尝试次数。 |
73
73
  | `maxOverflowRetries` | `1` | 已确认上下文窗口溢出后的最大重试次数;`0` 只禁用恢复。 |
74
- | `chunkInputRatio` | `0.6` | 每个层次阶段输入可用的摘要模型窗口比例;有效范围 `[0.1, 0.9]`。 |
74
+ | `chunkInputRatio` | `0.6` | 每个层次阶段输入可使用的摘要模型窗口比例;有效范围为 `[0.1, 0.9]`。 |
75
75
  | `mapMaxTokens` | `4096` | 单次层次 map 调用的提供方生成上限。 |
76
76
  | `reduceMaxTokens` | `8192` | 单次层次 reduce 调用的提供方生成上限。 |
77
- | `maxDepth` | `4` | 递归 reduce 轮次上限;有效范围 `1..8`。 |
78
- | `replayTools` | `false` | 在层次阶段回放工具 schema。严格提供方可能需要打开此项,但会占用 chunk 输入并降低前缀复用。 |
77
+ | `maxDepth` | `4` | 最大递归 reduce 轮数;有效范围为 `1..8`。 |
78
+ | `replayTools` | `false` | 在层次阶段重放工具 schema。严格提供方可能需要启用它,但它会占用分块输入并降低前缀复用。 |
79
79
  | `modelPolicies` | `[]` | 针对个别模型路由的精确 `{ provider, model, ...partialPolicy }` 覆盖。 |
80
80
  | `auto` | `true` | 启用自动压缩与溢出恢复;设为 `false` 则仅手动执行。 |
81
81
 
@@ -109,18 +109,18 @@ kind: "package-reference"
109
109
 
110
110
  - **一个测量服务为每个决策定价。** 单例 `ctx.tokenMeter` 会在同一个已消费日志 revision 上测量最新规范已记录 envelope 与当前表层。路由适配器声明请求图片定价时,meter 会将其应用于图片历史。压力、近期尾部保留、范围选择与缩减验证使用同一套路由定价的节点数值;已记录的替换影子价仍使用与路由无关的启发式规则,使纯投影 fold 保持一致。
111
111
  - **日志记录的标记对就是事务。** 所有入口点共享一个先记录标记的区域事务:验证范围与活动锁,同步追加 `compaction/start`,准备并等待摘要,重新验证,再追加 `compaction/summary` 与替换,最后恰好进行一次闭合尝试。自动调用与显式范围调用要求数字标识的开放轮次归属与整个表层稳定;`compactNow()` 会预留空闲接纳,使用 `turn: null`,允许所选 span 之外追加仅追加上下文,flush 每次已闭合尝试,并在 `finally` 中释放接纳预留。
112
- - **摘要复用提供方的热前缀。** 逐字回放上次已路由请求的系统提示词、工具与已遮蔽区域消息,使辅助调用成为会话的真正前缀,因此只有尾随指令与摘要输出未缓存。
112
+ - **摘要保留复用 cache 的快速路径,并约束超大工作。** 能容纳的 one-shot 会逐字回放 surface 节点 0 处 `system/message` 所承载的系统提示词、上次已路由请求的工具与已遮蔽区域消息。层次调用会恰好一次重放同一个固定 system head,同时约束时序 map span 与递归 reduce。
113
113
  - **`summarize()` 是唯一的子类钩子。** 基于模板或远程摘要器的子类可以覆盖它,同时压力、保留、被引用的源事件、缩减验证与已遮蔽 token 计量仍由 token meter 负责。
114
114
 
115
115
  ### 自动触发与溢出恢复
116
116
 
117
- 当 `auto: true` 时,串行 `agent/pre-step` listener 会在请求派生前检查压力:它通过 `ctx.tokenMeter` 为最新持久路由请求 envelope 定价,当压力越过路由模型的阈值时,先剪枝,再在保留已定价近期尾部的同时摘要最旧的平衡范围。`agent/request-error` listener 响应提供方确认的 `CONTEXT_WINDOW_EXCEEDED`:它绕过常规阈值与保留策略,尝试一次最大平衡头部缩减,并且只在表层替换 generation 前进后才授权重试。取消全程保持最终决定权。
117
+ 当 `auto: true` 时,串行 `agent/pre-step` listener 会在请求派生前检查压力:它通过 `ctx.tokenMeter` 为最新持久路由请求 envelope 定价,当压力越过路由模型的阈值时,先剪枝,再在保留已定价近期尾部的同时摘要最旧的平衡范围。每个选定范围都从第一个不是 `system/message` 的 surface 节点开始,因此位于 surface 节点 0 的系统提示词永不会被遮蔽;由历史内提示词更新追加的后续 `system/message` 是普通历史,范围可以遮蔽它,agent loop 的投影随后会在二者文本不同时用当前提示词替换节点 0([决策规则](../../core/agent-loop/README.zh.md#understand-the-implementation))。`agent/request-error` listener 响应提供方确认的 `CONTEXT_WINDOW_EXCEEDED`:它绕过常规阈值与保留策略,尝试一次最大平衡头部缩减,并且只在表层替换 generation 前进后才授权重试。取消全程保持最终决定权。
118
118
 
119
119
  压力策略从拥有持久路由的适配器解析容量。适配器无法为有效动态路由返回容量时,手动压力路径会抛出目标特定配置错误;自动 listener 会对该精确目标警告一次,并携带完整历史继续。
120
120
 
121
121
  ### 摘要机制
122
122
 
123
- 直接 `ctx.llm.stream()` 调用使用已配置的提供方/模型对与上限,回退到最新已记录请求目标,然后再回退到 `AgentOptions` 对,而不运行仅用于 agent loop 的 `agent/request` 扩展点。该调用逐字回放会话自身的系统提示词、工具与已遮蔽区域消息——包括所选适配器必须解析或明确拒绝的图片引用——并将压缩指令作为最后一条 user 消息追加,从而复用提供方的热前缀 cache,而非使它失效。调用将 `GenerateOptions.purpose` 设为 `compaction`;只有返回文本进入检查点,推理与工具调用都会被排除。图片输出会以 `UNSUPPORTED_CONTENT` 失败,而不是消失。替换 user 消息用 `<compacted-summary>` 标签框定摘要;原始摘要保留在 `compaction/summary` 事件上。
123
+ 直接 `ctx.llm.stream()` 调用使用已配置的提供方/模型对与上限,回退到最新已记录请求目标,然后再回退到 `AgentOptions` 对,而不运行仅用于 agent loop 的 `agent/request` 扩展点。完整请求能容纳时,该调用将 surface 节点 0 处派生的 `system/message` 作为 `messages` 的首项回放,后接已遮蔽区域消息(包括位于其 surface 位置的被遮蔽历史内 `system/message`),逐字携带 header 的工具,并将压缩指令作为最后一条 user 消息追加。超大输入或提供方拒绝的输入会被组合成工具配对平衡的时序 map span,再递归 reduce;每个 map 与 reduce 请求都把当前非空 system head 恰好一次作为首条消息重放,后续动态 `system/message` 仍是普通时序 map 输入,绝不会被提升或删除。每阶段预算都包含固定 head、可选工具与指令成本。只有 `replayTools: true` 时,工具 schema 才随层次调用发送。空内容系统头节点不贡献消息,但仍处于压缩范围之外。每次调用都将 `GenerateOptions.purpose` 设为 `compaction`;只有返回文本进入检查点,推理与工具调用都会被排除。图片输出会以 `UNSUPPORTED_CONTENT` 失败,而不是消失。替换 user 消息用 `<compacted-summary>` 标签框定摘要;最终原始摘要保留在 `compaction/summary` 事件上。
124
124
 
125
125
  ### 区域事务
126
126
 
@@ -136,7 +136,10 @@ kind: "package-reference"
136
136
  |---|---|
137
137
  | [`src/index.ts`](src/index.ts) | 插件入口:`BasicCompactionEngine`、自动 listener、入口点分发 |
138
138
  | [`src/region.ts`](src/region.ts) | 保留选择与共享的先记录标记压缩事务 |
139
- | [`src/summarizer.ts`](src/summarizer.ts) | 默认 `ctx.llm.stream()` 摘要、检查点框定、安全摘要投影 |
139
+ | [`src/summarizer.ts`](src/summarizer.ts) | 默认 `ctx.llm.stream()` one-shot 摘要、检查点框定、安全摘要投影 |
140
+ | [`src/hierarchical.ts`](src/hierarchical.ts) | 有界 map-reduce 回退、自适应拆分与阶段用量聚合 |
141
+ | [`src/hierarchical-planner.ts`](src/hierarchical-planner.ts) | 工具配对平衡单元与贪心 token 预算规划 |
142
+ | [`src/hierarchical-prompts.ts`](src/hierarchical-prompts.ts) | 结构化 map/reduce 提示词与输出验证 |
140
143
  | [`src/config.ts`](src/config.ts) | 加载时验证与路由模型策略解析 |
141
144
  | [`src/types.ts`](src/types.ts) | `BasicCompactionConfig` 与已解析策略词汇 |
142
145
  | — | 不发布运行时不变式伴生入口;持久标记对可在会话日志中观察。 |
@@ -186,7 +189,7 @@ This is an automatically generated checkpoint condensing an earlier span of the
186
189
 
187
190
  #### 模型看到的内容
188
191
 
189
- 摘要模型会接收逐字回放的会话:与上次已路由请求为已遮蔽区域发送的相同系统提示词、工具 schema 与消息,后面跟随一条最终 user 消息,即下方压缩指令。会话模型绝不会看到该私有请求或其推理;只有返回文本会被存储。
192
+ 完整请求能容纳时,摘要模型会接收逐字回放的会话:与上次已路由请求为已遮蔽区域发送的相同 system head、工具 schema 与消息,后面跟随一条最终 user 消息,即下方压缩指令。层次模式下,每个 map 请求接收恰好一次相同 system head、一个有序且工具配对平衡的源 span,以及结构化 map 指令;reduce 请求接收恰好一次相同 head、有序 `<partial-summary>` frame 和结构化 reduce 指令。后续 system 更新保留在原始 map 时序中,不会成为第二个固定 head。只有 `replayTools: true` 时,工具 schema 才随层次调用发送。会话模型绝不会看到这些私有请求或其推理;只有最终返回文本会被存储。
190
193
 
191
194
  ##### 压缩指令(最终 user 消息)
192
195
 
@@ -229,11 +232,11 @@ Rules:
229
232
 
230
233
  #### Token 影响
231
234
 
232
- 这是一次独立模型调用:输入是已回放会话前缀加固定指令,输出受 `maxTokens` 限制。收敛重试可能多次支付这项成本。
235
+ 能容纳的输入会产生一次独立模型调用:输入是已回放会话前缀加固定指令,输出受 `maxTokens` 限制。层次模式每个 map span 调用一次模型,再进行一次或多次 reduce,分别受 `mapMaxTokens` 与 `reduceMaxTokens` 限制;提供方确认的 overflow 可能在本地二分前增加失败尝试。收敛重试可能多次支付任一种成本。
233
236
 
234
237
  #### KV Cache 影响
235
238
 
236
- 已回放系统提示词、工具与已遮蔽区域消息与会话最后一个已路由请求逐字匹配,因此提供方的热前缀 cache 可复用至尾随指令之前;只有该指令与摘要输出未缓存。将摘要器路由到不同提供方/模型,或压缩非头部范围,都会放弃该复用。
239
+ 能容纳的 one-shot 会保留穿过已遮蔽区域的完整热前缀。层次调用会保留固定 system head,以及启用时工具 schema 的前缀复用;map payload 与 reduce frame 会在该共享前缀之后分叉。将摘要器路由到不同提供方/模型会放弃会话路由的 cache 复用。
237
240
 
238
241
  ## 已知限制与延期工作
239
242
 
@@ -244,9 +247,9 @@ Rules:
244
247
 
245
248
  - **计量准确度取决于固定启发式规则**——可复用提供方用量缺失时,会回退到字符数加结构开销,而非精确的 token 化;只有在适配器声明了请求图片定价的路由上,图片出现处才携带提供方精确的视觉 token。
246
249
  - **溢出分类由适配器维护**——提供方措辞可能改变;两个 DeepSeek 适配器将可识别的上下文限制失败规范化为 `CONTEXT_WINDOW_EXCEEDED`。
247
- - **有界恢复需要摘要模型的容量元数据**——省略 `contextWindow` 的适配器继续走旧的一次性路径。该请求成功则行为不变;若溢出,层次无法推导安全 chunk 预算,并以可操作的容量错误失败。
248
- - **层次输出是严格的检查点协议**——每个 map 与 reduce 阶段必须返回全部必需标题。截断、视觉输出、结构畸形、耗尽 `maxDepth`,或仍溢出的不可分源/部分摘要,都会让整次压缩事务失败,不安装部分检查点。
249
250
  - **部分不可分单元与仅 envelope 溢出仍不在表层压缩范围内**——恢复无法缩减系统/工具/前缀、拆分不可分的非工具节点,或修复不可剪枝剩余部分仍超出窗口的工具单元。可选 pruner 可以缩减原本不可分工具对内的文本型工具结果主体。
251
+ - **层次模式要求声明摘要模型容量**——没有正整数 `contextWindow` 时,能容纳的输入仍可 one-shot 摘要;但 one-shot 被确认 overflow 后会清晰失败,因为有界分块规划没有可信窗口。
252
+ - **层次模式刻意有界**——提供方拒绝不可分的工具配对平衡 span、固定 system/tools/instruction 开销耗尽阶段预算、reduce 轮次未减少 partial 数量,或达到 `maxDepth` 时都会失败,而不会循环。
250
253
  - **`compactRegion` 要求存在未结束的轮次**——在完全关闭的会话上手动调用会抛出异常(「no open turn」),而不是执行压缩。
251
254
  - **摘要失败会保留最新持久表层**——任何替换前,自动路径会记录警告,并携带完整超预算历史继续。如果剪枝已落地,后续摘要失败会从该持久剪枝表层继续。因达到 `maxTokens` 而发生的摘要截断(隐藏推理 token 可能会耗尽该额度)遵循同一规则。
252
255
 
package/lib/index.js CHANGED
@@ -337,7 +337,6 @@ async function summarizeWithLlm(ctx, config, input, agent, signal) {
337
337
  provider: target.provider,
338
338
  model: target.model,
339
339
  messages,
340
- ...input.system === void 0 ? {} : { system: input.system },
341
340
  ...input.tools === void 0 ? {} : { tools: [...input.tools] },
342
341
  maxTokens: config.maxTokens,
343
342
  sessionId: agent.session.id,
@@ -345,7 +344,7 @@ async function summarizeWithLlm(ctx, config, input, agent, signal) {
345
344
  ...signal === void 0 ? {} : { signal }
346
345
  };
347
346
  for await (const chunk of ctx.llm.stream(options)) assembler.push(chunk);
348
- const error = finishError$1(assembler.finish);
347
+ const error = finishError(assembler.finish);
349
348
  if (error !== void 0) throw error;
350
349
  const rawOutput = assembler.blocks();
351
350
  const summary = summaryText(rawOutput);
@@ -378,8 +377,12 @@ function frameSummary(summary) {
378
377
  }
379
378
  ];
380
379
  }
381
- /** Map a terminal summarization finish to its fail-closed error. */
382
- function finishError$1(finish) {
380
+ /**
381
+ * Map a terminal summarization finish to its fail-closed error.
382
+ * @param finish - terminal stream finish emitted by the summary request.
383
+ * @returns the corresponding error, or `undefined` for a complete stop.
384
+ */
385
+ function finishError(finish) {
383
386
  switch (finish.kind) {
384
387
  case "error":
385
388
  case "aborted": {
@@ -395,7 +398,11 @@ function finishError$1(finish) {
395
398
  default: return;
396
399
  }
397
400
  }
398
- /** Reject visual output and keep only text before synthesizing a user message. */
401
+ /**
402
+ * Reject visual output and keep only text before synthesizing a user message.
403
+ * @param blocks - raw content blocks emitted by the summary request.
404
+ * @returns the text-only blocks safe to persist as a compaction checkpoint.
405
+ */
399
406
  function summaryText(blocks) {
400
407
  if (contentHasImage(blocks)) throw new LlmError("compaction summary cannot contain image output", "UNSUPPORTED_CONTENT");
401
408
  return blocks.filter((block) => block.type === "text");
@@ -415,8 +422,21 @@ function summaryText(blocks) {
415
422
  */
416
423
  var SurfaceChangedError = class extends Error {};
417
424
  /**
418
- * Resolve the next head-anchored range while retaining a priced recent tail
419
- * and never splitting an assistant tool-call/result pair.
425
+ * The `system/message` holding surface node 0, or `undefined` when another
426
+ * message-producing event starts the surface.
427
+ * @param session - session supplying the log behind the current surface.
428
+ * @param headSeq - seq at surface node 0 of a non-empty surface.
429
+ * @returns the system head event, or `undefined` without one.
430
+ */
431
+ function systemHead(session, headSeq) {
432
+ const head = session.eventAt(headSeq);
433
+ return head.type === "system/message" ? head : void 0;
434
+ }
435
+ /**
436
+ * Resolve the next range starting at the first non-system surface node while
437
+ * retaining a priced recent tail and never splitting an assistant
438
+ * tool-call/result pair. A `system/message` at surface node 0 is never inside
439
+ * the range; without one the range starts at node 0.
420
440
  * @param session - session supplying authoritative current surface positions.
421
441
  * @param measurement - unified pressure and surface measurement from the conversation meter.
422
442
  * @param retainTokens - minimum recent tail budget retained verbatim.
@@ -427,6 +447,7 @@ function selectCompactableRange(session, measurement, retainTokens) {
427
447
  if (pricedNodes.length === 0) return null;
428
448
  const surfaceNodes = session.surface.nodes;
429
449
  if (surfaceNodes.length !== pricedNodes.length || surfaceNodes.some((seq, index) => seq !== pricedNodes[index]?.seq)) throw new Error("compaction: token-meter surface does not match the current session surface");
450
+ const firstIdx = systemHead(session, surfaceNodes[0]) === void 0 ? 0 : 1;
430
451
  let accumulated = 0;
431
452
  let keepFromIdx = pricedNodes.length;
432
453
  for (let index = pricedNodes.length - 1; index >= 0; index -= 1) {
@@ -434,14 +455,14 @@ function selectCompactableRange(session, measurement, retainTokens) {
434
455
  keepFromIdx = index;
435
456
  if (accumulated >= retainTokens) break;
436
457
  }
437
- if (keepFromIdx === 0) return null;
438
- while (keepFromIdx > 0) {
458
+ if (keepFromIdx <= firstIdx) return null;
459
+ while (keepFromIdx > firstIdx) {
439
460
  if (toolPairingBalancedBefore(session, surfaceNodes[keepFromIdx])) break;
440
461
  keepFromIdx -= 1;
441
462
  }
442
- if (keepFromIdx === 0) return null;
463
+ if (keepFromIdx <= firstIdx) return null;
443
464
  return {
444
- start: surfaceNodes[0],
465
+ start: surfaceNodes[firstIdx],
445
466
  end: surfaceNodes[keepFromIdx - 1]
446
467
  };
447
468
  }
@@ -652,8 +673,8 @@ function commitCompactionBody(session, startEvent, summarized) {
652
673
  session.append("user/message", checkpointMessage, {
653
674
  surfaceOp: {
654
675
  op: "replace",
655
- start,
656
- end
676
+ startSeq: start,
677
+ endSeq: end
657
678
  },
658
679
  sourceEventSeqs: [
659
680
  startEvent.seq,
@@ -684,21 +705,24 @@ function completeCompaction(pending, endEvent) {
684
705
  }
685
706
  /**
686
707
  * Reconstruct the last routed request's cacheable prefix for the shadowed
687
- * region: its system prompt and tool schemas, then the region's own derived
688
- * messages in surface order. The summarizer appends only the compaction
689
- * instruction after this, so the call is a genuine prefix of the conversation
690
- * and reuses the provider's KV cache.
691
- * @param session - session supplying the request header and per-node projection.
708
+ * region: the system prompt held by the `system/message` at surface node 0,
709
+ * the header's tool schemas, then the region's own derived messages in surface
710
+ * order. The summarizer appends only the compaction instruction after this, so
711
+ * the call is a genuine prefix of the conversation and reuses the provider's
712
+ * KV cache. A surface without a system head, or whose head projects to no
713
+ * message, contributes no leading system message.
714
+ * @param session - session supplying the surface head, request header, and per-node projection.
692
715
  * @param shadowedSeqs - the surface-node seqs, in order, being compacted.
693
716
  * @returns the replayed conversation prefix to condense.
694
717
  */
695
718
  function buildSummarizationInput(session, shadowedSeqs) {
696
719
  const header = session.requestHeader();
720
+ const head = systemHead(session, session.surface.nodes[0]);
721
+ const system = head === void 0 ? null : session.deriveEventMessage(head);
697
722
  const regionMessages = shadowedSeqs.map((seq) => session.deriveEventMessage(session.eventAt(seq))).filter((message) => message !== null);
698
723
  return {
699
- ...header?.system === void 0 ? {} : { system: header.system },
700
724
  ...header?.tools === void 0 ? {} : { tools: header.tools },
701
- messages: regionMessages
725
+ messages: system === null ? regionMessages : [system, ...regionMessages]
702
726
  };
703
727
  }
704
728
  /** Inspect open-turn, unmatched-compaction, and latest seed-boundary state independently. */
@@ -978,7 +1002,8 @@ var HierarchicalSummarizer = class {
978
1002
  /* v8 ignore next -- LlmRuntime validates defined capacity before returning model info. */
979
1003
  if (!Number.isSafeInteger(contextWindow) || contextWindow < 1) throw new Error(`hierarchical compaction: no positive integer context capacity for summary target ${target.provider}/${target.model}`);
980
1004
  const estimate = (message) => this.ctx.tokenMeter.estimateMessage(message);
981
- const oneShotTokens = this.estimateCallInput(input, COMPACTION_INSTRUCTION, true, estimate);
1005
+ const replay = hierarchyReplay(input.messages);
1006
+ const oneShotTokens = this.estimateCallInput(input, replay, COMPACTION_INSTRUCTION, true, estimate);
982
1007
  let hadFailedLlmAttempt = false;
983
1008
  if (oneShotTokens + target.oneShotMaxTokens <= contextWindow) try {
984
1009
  return await oneShot();
@@ -988,10 +1013,10 @@ var HierarchicalSummarizer = class {
988
1013
  }
989
1014
  const inputBudget = Math.floor(contextWindow * this.hierarchy.chunkInputRatio);
990
1015
  this.assertStageOutputReserve(contextWindow, inputBudget, this.hierarchy.mapMaxTokens, "map");
991
- const totalUnits = toolBalancedUnits(input.messages).length;
992
- const mapReserve = this.estimateFixedInput(input, mapInstruction(totalUnits, totalUnits, totalUnits), this.hierarchy.replayTools, estimate);
1016
+ const totalUnits = toolBalancedUnits(replay.sourceMessages).length;
1017
+ const mapReserve = this.estimateFixedInput(input, replay.systemHead, mapInstruction(totalUnits, totalUnits, totalUnits), this.hierarchy.replayTools, estimate);
993
1018
  const mapMessageBudget = this.messageBudget(inputBudget, mapReserve, "map");
994
- const chunks = planMessageChunks(input.messages, mapMessageBudget, estimate);
1019
+ const chunks = planMessageChunks(replay.sourceMessages, mapMessageBudget, estimate);
995
1020
  /* v8 ignore next -- stock range selection never submits an empty shadowed region. */
996
1021
  if (chunks.length === 0) throw new Error("hierarchical compaction: oversized input produced no map chunks");
997
1022
  const calls = [];
@@ -1003,10 +1028,7 @@ var HierarchicalSummarizer = class {
1003
1028
  /* v8 ignore next -- the loop condition proves shift has an entry. */
1004
1029
  if (span === void 0) break;
1005
1030
  try {
1006
- const result = await this.runStage({
1007
- ...input,
1008
- messages: span.messages
1009
- }, mapInstruction(span.start, span.end, totalUnits), target, this.hierarchy.mapMaxTokens, agent, signal);
1031
+ const result = await this.runStage(input, span.messages, replay.systemHead, mapInstruction(span.start, span.end, totalUnits), target, this.hierarchy.mapMaxTokens, agent, signal);
1010
1032
  calls.push(result);
1011
1033
  partials.push(this.partial(result, span.start, span.end, `map source units ${span.start}-${span.end}`));
1012
1034
  } catch (error) {
@@ -1023,7 +1045,7 @@ var HierarchicalSummarizer = class {
1023
1045
  let usedReduce = false;
1024
1046
  for (let round = 1; partials.length > 1; round += 1) {
1025
1047
  if (round > this.hierarchy.maxDepth) throw new Error(`hierarchical compaction: reduction did not converge within ${this.hierarchy.maxDepth} round(s)`);
1026
- const reduceReserve = this.estimateFixedInput(input, reduceInstruction(round, totalUnits, totalUnits, totalUnits), this.hierarchy.replayTools, estimate);
1048
+ const reduceReserve = this.estimateFixedInput(input, replay.systemHead, reduceInstruction(round, totalUnits, totalUnits, totalUnits), this.hierarchy.replayTools, estimate);
1027
1049
  const reduceMessageBudget = this.messageBudget(inputBudget, reduceReserve, `reduce round ${round}`);
1028
1050
  let groups;
1029
1051
  try {
@@ -1044,10 +1066,7 @@ var HierarchicalSummarizer = class {
1044
1066
  /* v8 ignore next -- the loop condition proves shift has an entry. */
1045
1067
  if (span === void 0) break;
1046
1068
  try {
1047
- const result = await this.runStage({
1048
- ...input,
1049
- messages: span.messages
1050
- }, reduceInstruction(round, span.start, span.end, totalUnits), target, this.hierarchy.reduceMaxTokens, agent, signal);
1069
+ const result = await this.runStage(input, span.messages, replay.systemHead, reduceInstruction(round, span.start, span.end, totalUnits), target, this.hierarchy.reduceMaxTokens, agent, signal);
1051
1070
  calls.push(result);
1052
1071
  next.push(this.partial(result, span.start, span.end, `reduce round ${round} source units ${span.start}-${span.end}`));
1053
1072
  } catch (error) {
@@ -1171,12 +1190,12 @@ var HierarchicalSummarizer = class {
1171
1190
  if (inputBudget + outputTokens > contextWindow) throw new Error(`hierarchical compaction: ${stage} input budget ${inputBudget} plus output reserve ${outputTokens} exceeds summary context ${contextWindow}`);
1172
1191
  }
1173
1192
  /** Price a complete auxiliary call input. */
1174
- estimateCallInput(input, instruction, includeTools, estimate) {
1175
- return this.estimateFixedInput(input, instruction, includeTools, estimate) + estimateMessages(input.messages, estimate);
1193
+ estimateCallInput(input, replay, instruction, includeTools, estimate) {
1194
+ return this.estimateFixedInput(input, replay.systemHead, instruction, includeTools, estimate) + estimateMessages(replay.sourceMessages, estimate);
1176
1195
  }
1177
- /** Price the repeated header and final instruction for one stage. */
1178
- estimateFixedInput(input, instruction, includeTools, estimate) {
1179
- return (input.system === void 0 ? 0 : Math.ceil(input.system.length / CHARS_PER_TOKEN) + ENVELOPE_OVERHEAD) + (!includeTools || input.tools === void 0 || input.tools.length === 0 ? 0 : Math.ceil(JSON.stringify(input.tools).length / CHARS_PER_TOKEN) + ENVELOPE_OVERHEAD) + estimate(this.instructionMessage(instruction));
1196
+ /** Price the repeated system head, optional tools, and final stage instruction. */
1197
+ estimateFixedInput(input, systemHead, instruction, includeTools, estimate) {
1198
+ return (systemHead === void 0 ? 0 : estimate(systemHead)) + (!includeTools || input.tools === void 0 || input.tools.length === 0 ? 0 : Math.ceil(JSON.stringify(input.tools).length / CHARS_PER_TOKEN) + ENVELOPE_OVERHEAD) + estimate(this.instructionMessage(instruction));
1180
1199
  }
1181
1200
  /** Derive positive room for stage messages after fixed input. */
1182
1201
  messageBudget(inputBudget, fixedTokens, stage) {
@@ -1185,14 +1204,17 @@ var HierarchicalSummarizer = class {
1185
1204
  return budget;
1186
1205
  }
1187
1206
  /** Run one private map or reduce model call and require structured text. */
1188
- async runStage(input, instruction, target, maxTokens, agent, signal) {
1207
+ async runStage(input, messages, systemHead, instruction, target, maxTokens, agent, signal) {
1189
1208
  signal?.throwIfAborted();
1190
1209
  const assembler = new BlockAssembler();
1191
1210
  const options = {
1192
1211
  provider: target.provider,
1193
1212
  model: target.model,
1194
- messages: [...input.messages, this.instructionMessage(instruction)],
1195
- ...input.system === void 0 ? {} : { system: input.system },
1213
+ messages: [
1214
+ ...systemHead === void 0 ? [] : [systemHead],
1215
+ ...messages,
1216
+ this.instructionMessage(instruction)
1217
+ ],
1196
1218
  ...this.hierarchy.replayTools && input.tools !== void 0 ? { tools: [...input.tools] } : {},
1197
1219
  maxTokens,
1198
1220
  sessionId: agent.session.id,
@@ -1203,8 +1225,7 @@ var HierarchicalSummarizer = class {
1203
1225
  const finishFailure = finishError(assembler.finish);
1204
1226
  if (finishFailure !== void 0) throw finishFailure;
1205
1227
  const rawOutput = assembler.blocks();
1206
- if (contentHasImage(rawOutput)) throw new LlmError("hierarchical compaction summary cannot contain image output", "UNSUPPORTED_CONTENT");
1207
- const summary = rawOutput.filter((block) => block.type === "text");
1228
+ const summary = summaryText(rawOutput);
1208
1229
  validateStructuredSummary(summary, "hierarchical compaction stage");
1209
1230
  return {
1210
1231
  summary,
@@ -1244,6 +1265,15 @@ var HierarchicalSummarizer = class {
1244
1265
  });
1245
1266
  }
1246
1267
  };
1268
+ /** Separate the fixed surface system head from chronological source history. */
1269
+ function hierarchyReplay(messages) {
1270
+ const [first, ...rest] = messages;
1271
+ if (first?.role === "system") return {
1272
+ systemHead: first,
1273
+ sourceMessages: rest
1274
+ };
1275
+ return { sourceMessages: messages };
1276
+ }
1247
1277
  /** Build the terminal diagnostic for a provider-rejected atomic span. */
1248
1278
  function indivisibleOverflow(stage, cause) {
1249
1279
  const error = new OversizedCompactionUnitError(`hierarchical compaction: ${stage} still exceeds the provider context window and is indivisible`, { cause });
@@ -1254,23 +1284,6 @@ function indivisibleOverflow(stage, cause) {
1254
1284
  function hasErrorCode(error, code) {
1255
1285
  return typeof error === "object" && error !== null && "code" in error && error.code === code;
1256
1286
  }
1257
- /** Map a terminal stage finish to a fail-closed error. */
1258
- function finishError(finish) {
1259
- switch (finish.kind) {
1260
- case "error":
1261
- case "aborted": {
1262
- const error = new Error(finish.failure.message);
1263
- error.code = finish.failure.code;
1264
- return error;
1265
- }
1266
- case "max-tokens": {
1267
- const error = /* @__PURE__ */ new Error("hierarchical compaction stage truncated at the token cap");
1268
- error.code = "MAX_TOKENS";
1269
- return error;
1270
- }
1271
- default: return;
1272
- }
1273
- }
1274
1287
  /**
1275
1288
  * Sum disjoint provider usage across every successful map and reduce call.
1276
1289
  * @param usages - stage usage values in call order.
@@ -25,8 +25,10 @@ interface CompactionTransactionOptions {
25
25
  readonly sourceCommandId?: CommandId;
26
26
  }
27
27
  /**
28
- * Resolve the next head-anchored range while retaining a priced recent tail
29
- * and never splitting an assistant tool-call/result pair.
28
+ * Resolve the next range starting at the first non-system surface node while
29
+ * retaining a priced recent tail and never splitting an assistant
30
+ * tool-call/result pair. A `system/message` at surface node 0 is never inside
31
+ * the range; without one the range starts at node 0.
30
32
  * @param session - session supplying authoritative current surface positions.
31
33
  * @param measurement - unified pressure and surface measurement from the conversation meter.
32
34
  * @param retainTokens - minimum recent tail budget retained verbatim.
@@ -4,7 +4,7 @@
4
4
  * @module @deepseek-ai/dsh-compaction-basic/summarizer
5
5
  */
6
6
  import type { Context } from '@deepseek-ai/cordis';
7
- import type { ContentBlock, Message, TokenUsage, ToolSchema } from '@deepseek-ai/dsh-llm';
7
+ import type { ContentBlock, FinishReason, Message, TokenUsage, ToolSchema } from '@deepseek-ai/dsh-llm';
8
8
  import type { Agent } from '@deepseek-ai/dsh-agent';
9
9
  interface SummaryConfig {
10
10
  readonly summarizationProvider: string;
@@ -26,11 +26,9 @@ export declare const COMPACTION_INSTRUCTION: string;
26
26
  * compaction instruction is then the only novel input.
27
27
  */
28
28
  export interface SummarizationInput {
29
- /** The conversation's own system prompt, reused for prefix-cache alignment; absent for a system-less request. */
30
- readonly system?: string;
31
29
  /** The conversation's tool schemas, reused for prefix-cache alignment; absent when the request carried none. */
32
30
  readonly tools?: readonly ToolSchema[];
33
- /** The shadowed region, in surface order, that precedes the compaction instruction. */
31
+ /** The derived system head, when present, followed by the shadowed region in surface order. */
34
32
  readonly messages: readonly Message[];
35
33
  }
36
34
  /** Safe summary content plus the exact auxiliary call envelope recorded with it. */
@@ -70,5 +68,19 @@ export declare function summarizeWithLlm(ctx: Context, config: SummaryConfig, in
70
68
  * @returns content for the synthesized replacement user message.
71
69
  */
72
70
  export declare function frameSummary(summary: readonly ContentBlock[]): ContentBlock[];
71
+ /**
72
+ * Map a terminal summarization finish to its fail-closed error.
73
+ * @param finish - terminal stream finish emitted by the summary request.
74
+ * @returns the corresponding error, or `undefined` for a complete stop.
75
+ */
76
+ export declare function finishError(finish: FinishReason): Error | undefined;
77
+ /**
78
+ * Reject visual output and keep only text before synthesizing a user message.
79
+ * @param blocks - raw content blocks emitted by the summary request.
80
+ * @returns the text-only blocks safe to persist as a compaction checkpoint.
81
+ */
82
+ export declare function summaryText(blocks: readonly ContentBlock[]): Array<Extract<ContentBlock, {
83
+ type: 'text';
84
+ }>>;
73
85
  export {};
74
86
  //# sourceMappingURL=summarizer.d.ts.map
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "@crazx/dsh-compaction-basic",
3
3
  "description": "Token-meter-driven compaction policy and LLM summarization backend for the DeepSeek Harness",
4
- "version": "0.1.2-rc.1.zw.1",
4
+ "version": "0.1.5-alpha.1.zw.1",
5
5
  "publishConfig": {
6
6
  "access": "public"
7
7
  },
@@ -28,13 +28,13 @@
28
28
  "license": "MIT",
29
29
  "peerDependencies": {
30
30
  "@deepseek-ai/cordis": "^4.0.2",
31
- "@deepseek-ai/dsh-agent": "^0.1.2-rc.1",
32
- "@deepseek-ai/dsh-commands": "^0.1.2-rc.1",
33
- "@deepseek-ai/dsh-compaction": "^0.1.2-rc.1",
34
- "@deepseek-ai/dsh-compaction-tool-result-pruner": "^0.1.2-rc.1",
35
- "@deepseek-ai/dsh-llm": "^0.1.2-rc.1",
36
- "@deepseek-ai/dsh-session": "^0.1.2-rc.1",
37
- "@deepseek-ai/dsh-token-meter": "^0.1.2-rc.1"
31
+ "@deepseek-ai/dsh-agent": "^0.1.5-alpha.1",
32
+ "@deepseek-ai/dsh-commands": "^0.1.5-alpha.1",
33
+ "@deepseek-ai/dsh-compaction": "^0.1.5-alpha.1",
34
+ "@deepseek-ai/dsh-compaction-tool-result-pruner": "^0.1.5-alpha.1",
35
+ "@deepseek-ai/dsh-llm": "^0.1.5-alpha.1",
36
+ "@deepseek-ai/dsh-session": "^0.1.5-alpha.1",
37
+ "@deepseek-ai/dsh-token-meter": "^0.1.5-alpha.1"
38
38
  },
39
39
  "peerDependenciesMeta": {
40
40
  "@deepseek-ai/dsh-compaction-tool-result-pruner": {
@@ -42,25 +42,25 @@
42
42
  }
43
43
  },
44
44
  "dependencies": {
45
- "@deepseek-ai/dsh-util-values": "^0.1.2-rc.1",
45
+ "@deepseek-ai/dsh-util-values": "^0.1.5-alpha.1",
46
46
  "@deepseek-ai/schemastery": "^3.18.2"
47
47
  },
48
48
  "devDependencies": {
49
49
  "@deepseek-ai/cordis": "^4.0.2",
50
50
  "@deepseek-ai/cordis-plugin-include": "^1.0.7",
51
51
  "@deepseek-ai/cordis-plugin-loader": "^1.0.3",
52
- "@deepseek-ai/dsh-agent": "^0.1.2-rc.1",
53
- "@deepseek-ai/dsh-agent-loop": "^0.1.2-rc.1",
54
- "@deepseek-ai/dsh-agent-loop-testkit": "^0.1.2-rc.1",
55
- "@deepseek-ai/dsh-commands": "^0.1.2-rc.1",
56
- "@deepseek-ai/dsh-compaction": "^0.1.2-rc.1",
57
- "@deepseek-ai/dsh-compaction-tool-result-pruner": "^0.1.2-rc.1",
58
- "@deepseek-ai/dsh-invariants": "^0.1.2-rc.1",
59
- "@deepseek-ai/dsh-llm": "^0.1.2-rc.1",
60
- "@deepseek-ai/dsh-llm-retry": "^0.1.2-rc.1",
61
- "@deepseek-ai/dsh-session": "^0.1.2-rc.1",
62
- "@deepseek-ai/dsh-session-projection": "^0.1.2-rc.1",
63
- "@deepseek-ai/dsh-token-meter": "^0.1.2-rc.1",
64
- "@deepseek-ai/dsh-tools": "^0.1.2-rc.1"
52
+ "@deepseek-ai/dsh-agent": "^0.1.5-alpha.1",
53
+ "@deepseek-ai/dsh-agent-loop": "^0.1.5-alpha.1",
54
+ "@deepseek-ai/dsh-agent-loop-testkit": "^0.1.5-alpha.1",
55
+ "@deepseek-ai/dsh-commands": "^0.1.5-alpha.1",
56
+ "@deepseek-ai/dsh-compaction": "^0.1.5-alpha.1",
57
+ "@deepseek-ai/dsh-compaction-tool-result-pruner": "^0.1.5-alpha.1",
58
+ "@deepseek-ai/dsh-invariants": "^0.1.5-alpha.1",
59
+ "@deepseek-ai/dsh-llm": "^0.1.5-alpha.1",
60
+ "@deepseek-ai/dsh-llm-retry": "^0.1.5-alpha.1",
61
+ "@deepseek-ai/dsh-session": "^0.1.5-alpha.1",
62
+ "@deepseek-ai/dsh-session-projection": "^0.1.5-alpha.1",
63
+ "@deepseek-ai/dsh-token-meter": "^0.1.5-alpha.1",
64
+ "@deepseek-ai/dsh-tools": "^0.1.5-alpha.1"
65
65
  }
66
66
  }