@deepseek-ai/dsh-compaction-basic 0.1.6-alpha.1 → 0.1.7-alpha.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.i18n.yaml CHANGED
@@ -2,5 +2,5 @@
2
2
  # side as of the last confirmed-consistent state. Both languages carry equal authority;
3
3
  # after editing either side, bring the other along and re-record with:
4
4
  # pnpm run verify-translation-pairing --write packages/compaction/compaction-basic/README.md
5
- README.md: 68d5a321e04a90bfb726392d1fb179612f6e0626
6
- README.zh.md: f941cac74053a680cad859b68c7836e181065bc9
5
+ README.md: e0a7d90fecf308fa9f83bf87c2bfb5c4ff13f2e3
6
+ README.zh.md: 32d4f528e46206c5939dda038b710e01ce912d99
package/README.md CHANGED
@@ -59,22 +59,23 @@ You can verify success by watching the conversation continue past the point wher
59
59
 
60
60
  ### Tuning when condensation starts
61
61
 
62
- All settings are optional. The defaults start condensing at 80% of the routed model's context window and keep the newest 16% verbatim; the table below is the complete policy surface, and the generated [configuration catalog](../../../docs/config-catalog.md#deepseek-aidsh-compaction-basic) is the exhaustive source.
62
+ All settings are optional. With context window `W`, effective request output cap `O`, and headroom `B`, the default trigger is `floor(min(W × 0.8, W − O − B))`, where `B = 65,536` tokens. Retention keeps the newest 16% of `W − O` verbatim. The table below lists every setting; the generated [configuration catalog](../../../docs/config-catalog.md#deepseek-aidsh-compaction-basic) also includes their types.
63
63
 
64
64
  | Field | Default | Meaning |
65
65
  |---|---|---|
66
- | `thresholdRatio` | `0.8` | Start condensing at `floor(routedContextWindow × ratio)`. |
67
- | `retainRatio` | `0.16` | Recent conversation kept verbatim as a fraction of the routed context window; mutually exclusive with `retainTokens`. |
66
+ | `thresholdRatio` | `0.8` | Window fraction used in `floor(min(W × thresholdRatio, W − O − headroomTokens))`. |
67
+ | `headroomTokens` | `65536` | Additional pressure headroom beyond the routed output reservation; a non-negative integer. |
68
+ | `retainRatio` | `0.16` | Recent conversation kept verbatim as a fraction of `W − O`; mutually exclusive with `retainTokens`. |
68
69
  | `retainTokens` | — | Absolute recent-conversation budget kept verbatim; mutually exclusive with `retainRatio` and must be below the resolved threshold. |
69
70
  | `summarizationProvider` | `''` | Set together with `summarizationModel`; an empty pair uses the latest routed request target, then the `AgentOptions` pair. |
70
71
  | `summarizationModel` | `''` | Set together with `summarizationProvider`; an empty pair uses the latest routed request target, then the `AgentOptions` pair. |
71
- | `maxTokens` | `8192` | Output cap for the summarization request; may include reasoning tokens. |
72
+ | `maxTokens` | `headroomTokens` (`65536`) | Positive summary output cap, including any provider-counted reasoning tokens. Explicit per-model caps override explicit global caps; otherwise the cap follows the resolved headroom. |
72
73
  | `compactionRetries` | `1` | Extra condensation attempts after the first when pressure remains above threshold. |
73
74
  | `maxOverflowRetries` | `1` | Maximum retries after a confirmed context-window overflow; `0` disables recovery only. |
74
75
  | `modelPolicies` | `[]` | Exact `{ provider, model, ...partialPolicy }` overrides for individual model routes. |
75
76
  | `auto` | `true` | Enable automatic condensation and overflow recovery; set `false` for manual-only operation. |
76
77
 
77
- Misconfiguration fails fast: an unknown setting, a duplicate per-model override, both retention forms together, or a ratio retention that is not below the threshold all reject the plugin at load. An absolute `retainTokens` budget — top-level or per-model — that is not below its threshold fails when that model is first used, because the comparison needs the model's context size.
78
+ Misconfiguration fails fast: unknown settings, duplicate per-model overrides, invalid token counts, both retention forms together, or a retention ratio at least as large as the threshold ratio reject the plugin at load. When the model is first used, `W − O − B` must be positive and the resolved retained budget must be below the trigger. Zero headroom requires an explicit positive `maxTokens`, globally or in that model policy. Small-window deployments must configure headroom that fits their capacity; lower `thresholdRatio` to compact earlier.
78
79
 
79
80
  ### What happens when condensation runs
80
81
 
@@ -111,11 +112,11 @@ The backend is built on four commitments:
111
112
 
112
113
  With `auto: true`, a serial `agent/pre-step` listener checks pressure before request derivation: it prices the latest durable routed request envelope through `ctx.tokenMeter`, and when pressure crosses the routed model's threshold it prunes, then summarizes the oldest balanced span while keeping a priced recent tail. Every selected range starts at the first surface node that is not a `system/message`, so a system prompt at surface node 0 is never shadowed; a later `system/message` appended by an in-history prompt update is ordinary history that the range may shadow, and the agent loop's projection then replaces node 0 with the current prompt when their text differs ([decision rule](../../core/agent-loop/README.md#understand-the-implementation)). The `agent/request-error` listener reacts to a provider-confirmed `CONTEXT_WINDOW_EXCEEDED`: it bypasses the normal threshold and retention policy, attempts one maximal balanced head reduction, and authorizes a retry only after the surface replacement generation advances. Cancellation stays authoritative throughout.
113
114
 
114
- Pressure policy resolves capacity from the adapter that owns the durable route. An adapter that returns no capacity for a valid dynamic route makes the manual pressure path throw a target-specific configuration error; the automatic listener warns once for that exact target and continues with full history.
115
+ Pressure policy resolves capacity from the adapter that owns the durable route. Missing capacity, output plus headroom exhausting the window, or a retained budget at least as large as the threshold makes the manual pressure path throw a target-specific configuration error. The automatic listener warns once for that exact target and skips proactive compaction until its configuration is corrected; provider-confirmed overflow recovery remains available.
115
116
 
116
117
  ### Summarization mechanics
117
118
 
118
- A direct `ctx.llm.stream()` call uses the configured provider/model pair and cap, falling back to the latest logged request target and then the `AgentOptions` pair, without running the loop-only `agent/request` extension point. The call replays the derived `system/message` at surface node 0 as the leading entry of `messages`, followed by the shadowed-region messages (including a shadowed in-history `system/message` in its surface position), and carries the header's tools verbatim — including image references, which the selected adapter must resolve or explicitly reject — and appends the compaction instruction as the final user message, so it reuses the provider's warm prefix cache instead of invalidating it. An empty-content system head contributes no message but remains outside the compacted range. The call sets `GenerateOptions.purpose` to `compaction`; only returned text enters the checkpoint, excluding reasoning and tool calls. Image output fails with `UNSUPPORTED_CONTENT` rather than disappearing. The replacement user message frames the summary with `<compacted-summary>` tags; the raw summary remains on the `compaction/summary` event.
119
+ A direct `ctx.llm.stream()` call uses the configured provider/model pair and cap, falling back to the latest logged request target and then the `AgentOptions` pair, without running the loop-only `agent/request` extension point. The call replays the derived `system/message` at surface node 0 as the leading entry of `messages`, followed by the shadowed-region messages (including a shadowed in-history `system/message` in its surface position), and carries the header's tools verbatim — including image references, which the selected adapter must resolve or explicitly reject — and appends the compaction instruction as the final user message, so it reuses the provider's warm prefix cache instead of invalidating it. An empty-content system head contributes no message but remains outside the compacted range. The final instruction is a frozen `RequestUserInput` without durable identity or source; the replayed history and persisted checkpoint remain durable messages. The call sets `GenerateOptions.purpose` to `compaction`; only returned text enters the checkpoint, excluding reasoning and tool calls. Image output fails with `UNSUPPORTED_CONTENT` rather than disappearing. The replacement user message frames the summary with `<compacted-summary>` tags; the raw summary remains on the `compaction/summary` event.
119
120
 
120
121
  ### The region transaction
121
122
 
@@ -125,7 +126,7 @@ The transaction validates the surface span and the durable lock, appends `compac
125
126
 
126
127
  ### Config resolution
127
128
 
128
- `resolveConfig` validates and detaches the defaults, `resolveTargetPolicy` merges an exact provider/model override over them, and `resolveCompactSpec` scales the merged policy into concrete token budgets using the adapter-owned context capacity. Model discovery (`listModels()`) is never consulted for policy; only the durable route's capacity matters.
129
+ `resolveConfig` validates and detaches defaults, `resolveTargetPolicy` merges exact provider/model overrides, and `resolveCompactSpec` requires explicit adapter capacity and routed output reservation to resolve the trigger and retained budget. The effective envelope’s `maxTokens` supplies that reservation, falling back to the adapter default and then zero. Model discovery (`listModels()`) is never consulted for policy; only the durable route's capacity matters.
129
130
 
130
131
  ### Source map
131
132
 
package/README.zh.md CHANGED
@@ -59,22 +59,23 @@ kind: "package-reference"
59
59
 
60
60
  ### 调整压缩开始的时机
61
61
 
62
- 所有设置都可选。默认在已路由模型上下文窗口的 80% 处开始压缩,并逐字保留最新的 16%;下表是完整的策略面,生成的[配置目录](../../../docs/config-catalog.zh.md#deepseek-aidsh-compaction-basic)是穷尽式真源。
62
+ 所有设置都可选。设上下文窗口为 `W`、生效请求输出上限为 `O`、余量为 `B`,默认触发阈值为 `floor(min(W × 0.8, W − O − B))`,其中 `B = 65,536` tokens。逐字保留的近期历史预算仍为 `W − O` 的 16%。下表列出全部设置;生成的[配置目录](../../../docs/config-catalog.zh.md#deepseek-aidsh-compaction-basic)还包含字段类型。
63
63
 
64
64
  | 字段 | 默认值 | 含义 |
65
65
  |---|---|---|
66
- | `thresholdRatio` | `0.8` | 在 `floor(routedContextWindow × ratio)` 处开始压缩。 |
67
- | `retainRatio` | `0.16` | 以已路由上下文窗口的一部分表示逐字保留的近期对话;与 `retainTokens` 互斥。 |
66
+ | `thresholdRatio` | `0.8` | 用于 `floor(min(W × thresholdRatio, W − O − headroomTokens))` 的窗口比例。 |
67
+ | `headroomTokens` | `65536` | 路由请求输出预留之外的额外压力余量;必须为非负整数。 |
68
+ | `retainRatio` | `0.16` | 以 `W − O` 的一部分表示逐字保留的近期对话;与 `retainTokens` 互斥。 |
68
69
  | `retainTokens` | — | 逐字保留的近期对话绝对预算;与 `retainRatio` 互斥,并且必须低于已解析阈值。 |
69
70
  | `summarizationProvider` | `''` | 与 `summarizationModel` 一起设置;空对使用最新已路由请求目标,再回退到 `AgentOptions` 对。 |
70
71
  | `summarizationModel` | `''` | 与 `summarizationProvider` 一起设置;空对使用最新已路由请求目标,再回退到 `AgentOptions` 对。 |
71
- | `maxTokens` | `8192` | 摘要请求的输出上限;可包含推理 token。 |
72
+ | `maxTokens` | `headroomTokens`(`65536`) | 正数摘要输出上限,包含提供方计入的推理 token。显式模型上限覆盖显式全局上限;否则跟随解析后的余量。 |
72
73
  | `compactionRetries` | `1` | 压力仍高于阈值时,在首次压缩后进行的额外尝试次数。 |
73
74
  | `maxOverflowRetries` | `1` | 已确认上下文窗口溢出后的最大重试次数;`0` 只禁用恢复。 |
74
75
  | `modelPolicies` | `[]` | 针对个别模型路由的精确 `{ provider, model, ...partialPolicy }` 覆盖。 |
75
76
  | `auto` | `true` | 启用自动压缩与溢出恢复;设为 `false` 则仅手动执行。 |
76
77
 
77
- 配置错误会快速失败:未知设置、重复的按模型覆盖、两种保留形式同时出现,或比例保留量不低于阈值,都会在加载时拒绝插件。任何绝对 `retainTokens` 预算——顶层或按模型——不低于其阈值时,都会在该模型首次使用时失败,因为该比较需要模型的上下文大小。
78
+ 配置错误会快速失败:未知设置、重复的按模型覆盖、无效 token 数、两种保留形式同时出现,或保留比例不小于阈值比例,都会在加载时拒绝插件。模型首次使用时,`W − O − B` 必须为正,且解析出的保留预算必须低于触发阈值。余量为零时,必须在全局或对应模型策略中显式设置正数 `maxTokens`。小窗口部署必须配置适合其容量的余量;降低 `thresholdRatio` 可以提早压缩。
78
79
 
79
80
  ### 压缩运行时会发生什么
80
81
 
@@ -111,11 +112,11 @@ kind: "package-reference"
111
112
 
112
113
  当 `auto: true` 时,串行 `agent/pre-step` listener 会在请求派生前检查压力:它通过 `ctx.tokenMeter` 为最新持久路由请求 envelope 定价,当压力越过路由模型的阈值时,先剪枝,再在保留已定价近期尾部的同时摘要最旧的平衡范围。每个选定范围都从第一个不是 `system/message` 的 surface 节点开始,因此位于 surface 节点 0 的系统提示词永不会被遮蔽;由历史内提示词更新追加的后续 `system/message` 是普通历史,范围可以遮蔽它,agent loop(智能体循环)的投影随后会在二者文本不同时用当前提示词替换节点 0([决策规则](../../core/agent-loop/README.zh.md#understand-the-implementation))。`agent/request-error` listener 响应提供方确认的 `CONTEXT_WINDOW_EXCEEDED`:它绕过常规阈值与保留策略,尝试一次最大平衡头部缩减,并且只在表层替换 generation 前进后才授权重试。取消全程保持最终决定权。
113
114
 
114
- 压力策略从拥有持久路由的适配器解析容量。适配器无法为有效动态路由返回容量时,手动压力路径会抛出目标特定配置错误;自动 listener 会对该精确目标警告一次,并携带完整历史继续。
115
+ 压力策略从拥有持久路由的适配器解析容量。容量缺失、输出预留与余量耗尽窗口,或保留预算不小于阈值时,手动压力路径会抛出目标特定配置错误。自动 listener 会对该精确目标警告一次,并在配置修正前跳过主动压缩;提供方确认溢出后的恢复仍然可用。
115
116
 
116
117
  ### 摘要机制
117
118
 
118
- 直接 `ctx.llm.stream()` 调用使用已配置的提供方/模型对与上限,回退到最新已记录请求目标,然后再回退到 `AgentOptions` 对,而不运行仅用于 agent loop 的 `agent/request` 扩展点。该调用将 surface 节点 0 处派生的 `system/message` 作为 `messages` 的首项回放,后接已遮蔽区域消息(包括位于其 surface 位置的被遮蔽历史内 `system/message`),并逐字携带 header 的工具——包括所选适配器必须解析或明确拒绝的图片引用——并将压缩指令作为最后一条 user 消息追加,从而复用提供方的热前缀 cache,而非使它失效。空内容系统头节点不贡献消息,但仍处于压缩范围之外。调用将 `GenerateOptions.purpose` 设为 `compaction`;只有返回文本进入检查点,推理与工具调用都会被排除。图片输出会以 `UNSUPPORTED_CONTENT` 失败,而不是消失。替换 user 消息用 `<compacted-summary>` 标签框定摘要;原始摘要保留在 `compaction/summary` 事件上。
119
+ 直接 `ctx.llm.stream()` 调用使用已配置的提供方/模型对与上限,回退到最新已记录请求目标,然后再回退到 `AgentOptions` 对,而不运行仅用于 agent loop 的 `agent/request` 扩展点。该调用将 surface 节点 0 处派生的 `system/message` 作为 `messages` 的首项回放,后接已遮蔽区域消息(包括位于其 surface 位置的被遮蔽历史内 `system/message`),并逐字携带 header 的工具——包括所选适配器必须解析或明确拒绝的图片引用——并将压缩指令作为最后一条 user 消息追加,从而复用提供方的热前缀 cache,而非使它失效。空内容系统头节点不贡献消息,但仍处于压缩范围之外。最终指令是冻结的 `RequestUserInput`,不含持久身份或来源;回放历史和持久化检查点仍然使用持久消息。调用将 `GenerateOptions.purpose` 设为 `compaction`;只有返回文本进入检查点,推理与工具调用都会被排除。图片输出会以 `UNSUPPORTED_CONTENT` 失败,而不是消失。替换 user 消息用 `<compacted-summary>` 标签框定摘要;原始摘要保留在 `compaction/summary` 事件上。
119
120
 
120
121
  ### 区域事务
121
122
 
@@ -125,7 +126,7 @@ kind: "package-reference"
125
126
 
126
127
  ### 配置解析
127
128
 
128
- `resolveConfig` 验证并分离默认值,`resolveTargetPolicy` 将精确的提供方/模型覆盖合并到默认值之上,`resolveCompactSpec` 使用适配器拥有的上下文容量将合并后的策略缩放为具体 token 预算。策略解析绝不咨询模型发现(`listModels()`);只有持久路由的容量才重要。
129
+ `resolveConfig` 验证并分离默认值,`resolveTargetPolicy` 合并精确的提供方/模型覆盖,`resolveCompactSpec` 要求显式传入适配器容量与路由请求的输出预留,以解析触发阈值和保留预算。预留取自生效信封的 `maxTokens`,否则回退到适配器默认值,再回退到零。策略解析绝不咨询模型发现(`listModels()`);只有持久路由的容量才重要。
129
130
 
130
131
  ### 源码地图
131
132
 
package/lib/index.js CHANGED
@@ -18,6 +18,7 @@ const DEFAULT_RETAIN_RATIO = .16;
18
18
  /** Fields shared by top-level defaults and exact-target overrides. */
19
19
  const POLICY_CONFIG_KEYS = [
20
20
  "thresholdRatio",
21
+ "headroomTokens",
21
22
  "retainRatio",
22
23
  "retainTokens",
23
24
  "summarizationProvider",
@@ -59,17 +60,25 @@ function resolveConfig(config = {}) {
59
60
  validateKeys(config, BASIC_COMPACT_CONFIG_KEYS, "BasicCompactionConfig");
60
61
  validatePolicy(config, "BasicCompactionConfig");
61
62
  if (config.auto !== void 0 && typeof config.auto !== "boolean") throw new Error("BasicCompactionConfig: auto must be a boolean");
63
+ const headroomTokens = config.headroomTokens ?? 65536;
64
+ const maxTokens = config.maxTokens ?? headroomTokens;
65
+ assertPositiveInteger("BasicCompactionConfig.maxTokens (explicit or from headroomTokens)", maxTokens);
62
66
  const thresholdRatio = config.thresholdRatio ?? DEFAULT_THRESHOLD_RATIO;
63
67
  const retention = resolveRetention(config, { retainRatio: DEFAULT_RETAIN_RATIO });
64
68
  validateRatioRetention(thresholdRatio, retention, "BasicCompactionConfig");
65
69
  const modelPolicies = resolveModelPolicies(config.modelPolicies);
66
- for (const [index, policy] of modelPolicies.entries()) validateRatioRetention(policy.thresholdRatio ?? thresholdRatio, resolveRetention(policy, retention), `BasicCompactionConfig: modelPolicies[${index}]`);
70
+ for (const [index, policy] of modelPolicies.entries()) {
71
+ if (policy.maxTokens === void 0 && config.maxTokens === void 0 && policy.headroomTokens !== void 0) policy.maxTokens = policy.headroomTokens;
72
+ assertPositiveInteger(`BasicCompactionConfig: modelPolicies[${index}].maxTokens (explicit or from headroomTokens)`, policy.maxTokens ?? maxTokens);
73
+ validateRatioRetention(policy.thresholdRatio ?? thresholdRatio, resolveRetention(policy, retention), `BasicCompactionConfig: modelPolicies[${index}]`);
74
+ }
67
75
  return deepFreeze({
68
76
  thresholdRatio,
77
+ headroomTokens,
69
78
  ...retention,
70
79
  summarizationProvider: config.summarizationProvider ?? "",
71
80
  summarizationModel: config.summarizationModel ?? "",
72
- maxTokens: config.maxTokens ?? 8192,
81
+ maxTokens,
73
82
  compactionRetries: config.compactionRetries ?? 1,
74
83
  maxOverflowRetries: config.maxOverflowRetries ?? 1,
75
84
  modelPolicies,
@@ -91,6 +100,7 @@ function resolveTargetPolicy(config, target) {
91
100
  model: target.model
92
101
  },
93
102
  thresholdRatio: override?.thresholdRatio ?? config.thresholdRatio,
103
+ headroomTokens: override?.headroomTokens ?? config.headroomTokens,
94
104
  ...resolveRetention(override ?? {}, inheritedRetention),
95
105
  summarizationProvider: override?.summarizationProvider ?? config.summarizationProvider,
96
106
  summarizationModel: override?.summarizationModel ?? config.summarizationModel,
@@ -101,15 +111,26 @@ function resolveTargetPolicy(config, target) {
101
111
  }
102
112
  /**
103
113
  * Scale one routed policy into concrete token budgets for its model capacity.
114
+ *
115
+ * Pressure is capped by both the window fraction and the capacity remaining
116
+ * after the routed output reservation plus compaction headroom. Retention scales
117
+ * the message budget before headroom is deducted.
118
+ *
104
119
  * @param policy - merged policy for the exact routed target.
105
120
  * @param contextWindow - positive adapter-owned capacity for that target.
121
+ * @param reservedCompletionTokens - output tokens one routed request reserves.
106
122
  * @returns detached immutable pressure and retention budgets.
107
123
  */
108
- function resolveCompactSpec(policy, contextWindow) {
124
+ function resolveCompactSpec(policy, contextWindow, reservedCompletionTokens) {
109
125
  const targetKey = `${policy.target.provider}/${policy.target.model}`;
110
126
  if (!Number.isInteger(contextWindow) || contextWindow <= 0) throw new TargetPressureConfigError(targetKey, `BasicCompactionConfig: contextWindow (${contextWindow}) must be a positive integer`);
111
- const thresholdTokens = Math.floor(contextWindow * policy.thresholdRatio);
112
- const retainTokens = policy.retainTokens === void 0 ? Math.floor(contextWindow * policy.retainRatio) : policy.retainTokens;
127
+ if (!Number.isInteger(reservedCompletionTokens) || reservedCompletionTokens < 0) throw new TargetPressureConfigError(targetKey, `BasicCompactionConfig: reservedCompletionTokens (${reservedCompletionTokens}) must be a non-negative integer`);
128
+ const messageBudgetTokens = contextWindow - reservedCompletionTokens;
129
+ if (messageBudgetTokens <= 0) throw new TargetPressureConfigError(targetKey, `compaction-basic: ${targetKey} reserves ${reservedCompletionTokens} completion tokens of its ${contextWindow}-token context window, leaving no message budget; configure the adapter model's contextWindow above the effective request maxTokens`);
130
+ const pressureBudgetTokens = messageBudgetTokens - policy.headroomTokens;
131
+ if (pressureBudgetTokens <= 0) throw new TargetPressureConfigError(targetKey, `compaction-basic: ${targetKey} reserves ${reservedCompletionTokens} completion tokens and ${policy.headroomTokens} headroom tokens of its ${contextWindow}-token context window, leaving no pressure budget; reduce the effective request maxTokens or compaction headroomTokens, or configure a larger adapter model contextWindow`);
132
+ const thresholdTokens = Math.floor(Math.min(contextWindow * policy.thresholdRatio, pressureBudgetTokens));
133
+ const retainTokens = policy.retainTokens === void 0 ? Math.floor(messageBudgetTokens * policy.retainRatio) : policy.retainTokens;
113
134
  if (retainTokens >= thresholdTokens) throw new TargetPressureConfigError(targetKey, `BasicCompactionConfig: ${policy.target.provider}/${policy.target.model} retainTokens (${retainTokens}) must be less than threshold tokens ${thresholdTokens}`);
114
135
  return deepFreeze({
115
136
  target: { ...policy.target },
@@ -158,12 +179,14 @@ function assertModelPolicy(source, name) {
158
179
  /** Validate the fields common to defaults and exact-target partial overrides. */
159
180
  function validatePolicy(config, name) {
160
181
  const thresholdRatio = config.thresholdRatio;
182
+ const headroomTokens = config.headroomTokens;
161
183
  const retainRatio = config.retainRatio;
162
184
  const retainTokens = config.retainTokens;
163
185
  const maxTokens = config.maxTokens;
164
186
  const compactionRetries = config.compactionRetries;
165
187
  const maxOverflowRetries = config.maxOverflowRetries;
166
188
  if (thresholdRatio !== void 0) assertRatio(`${name}.thresholdRatio`, thresholdRatio);
189
+ if (headroomTokens !== void 0) assertNonNegativeInteger(`${name}.headroomTokens`, headroomTokens);
167
190
  if (retainRatio !== void 0) assertRatio(`${name}.retainRatio`, retainRatio);
168
191
  if (retainTokens !== void 0) assertNonNegativeInteger(`${name}.retainTokens`, retainTokens);
169
192
  if (retainRatio !== void 0 && retainTokens !== void 0) throw new Error(`${name}: retainRatio and retainTokens are mutually exclusive`);
@@ -279,15 +302,12 @@ async function summarizeWithLlm(ctx, config, input, agent, signal) {
279
302
  const target = configured ?? latest ?? agentTarget;
280
303
  if (target === void 0) throw new Error("no provider/model available for summarization: set both BasicCompactionConfig summarization fields, route one request, or set both AgentOptions fields");
281
304
  const assembler = new BlockAssembler();
282
- const messages = [...input.messages, createUserMessage({
305
+ const messages = [...input.messages, deepFreeze({
306
+ role: "user",
283
307
  content: [{
284
308
  type: "text",
285
309
  text: COMPACTION_INSTRUCTION
286
- }],
287
- source: {
288
- kind: "plugin",
289
- plugin: "dsh-compaction-basic"
290
- }
310
+ }]
291
311
  })];
292
312
  const options = {
293
313
  provider: target.provider,
@@ -727,6 +747,15 @@ function routedTarget(session) {
727
747
  model: config.model
728
748
  };
729
749
  }
750
+ /**
751
+ * Output tokens the routed request reserves, which the provider charges to the
752
+ * same window as the prompt. The effective envelope's own cap wins; otherwise
753
+ * the adapter's per-request default, which the adapter materializes when that
754
+ * envelope omits one. No declared cap means no reservation.
755
+ */
756
+ function reservedCompletionTokens(agent, defaultMaxTokens) {
757
+ return agent.session.requestHeader()?.config.maxTokens ?? defaultMaxTokens ?? 0;
758
+ }
730
759
  /** Resolve the conversation target used to select an optional policy override. */
731
760
  function conversationTarget(agent) {
732
761
  const routed = routedTarget(agent.session);
@@ -738,6 +767,7 @@ function conversationTarget(agent) {
738
767
  };
739
768
  }
740
769
  const thresholdRatioSchema = z.number();
770
+ const headroomTokensSchema = z.number().step(1).min(0);
741
771
  const retainRatioSchema = z.number();
742
772
  const retainTokensSchema = z.number().step(1).min(0);
743
773
  const summarizationProviderSchema = z.string();
@@ -749,6 +779,7 @@ const modelPolicy = z.object({
749
779
  provider: z.string().required(),
750
780
  model: z.string().required(),
751
781
  thresholdRatio: thresholdRatioSchema,
782
+ headroomTokens: headroomTokensSchema,
752
783
  retainRatio: retainRatioSchema,
753
784
  retainTokens: retainTokensSchema,
754
785
  summarizationProvider: summarizationProviderSchema,
@@ -773,6 +804,7 @@ var BasicCompactionEngine = class extends CompactionEngine {
773
804
  ];
774
805
  static Config = z.object({
775
806
  thresholdRatio: thresholdRatioSchema,
807
+ headroomTokens: headroomTokensSchema,
776
808
  retainRatio: retainRatioSchema,
777
809
  retainTokens: retainTokensSchema,
778
810
  summarizationProvider: summarizationProviderSchema,
@@ -900,11 +932,11 @@ var BasicCompactionEngine = class extends CompactionEngine {
900
932
  if (range === null) return null;
901
933
  return this.compactRegion(range.start, range.end, agent, signal);
902
934
  }
903
- const context = (await this.ctx.llm.resolveModelInfo(target.provider, target.model, signal)).context;
935
+ const info = await this.ctx.llm.resolveModelInfo(target.provider, target.model, signal);
904
936
  assertNoActiveCompaction(agent.session, "automatic pressure compaction");
905
937
  const targetKey = `${target.provider}/${target.model}`;
906
- if (context === void 0) throw new TargetPressureConfigError(targetKey, `compaction-basic: no context capacity for ${targetKey}; configure contextWindow on that adapter model`);
907
- const spec = resolveCompactSpec(policy, context.contextWindow);
938
+ if (info.context === void 0) throw new TargetPressureConfigError(targetKey, `compaction-basic: no context capacity for ${targetKey}; configure contextWindow on that adapter model`);
939
+ const spec = resolveCompactSpec(policy, info.context.contextWindow, reservedCompletionTokens(agent, info.defaultMaxTokens));
908
940
  if (measurement.totalTokens < spec.thresholdTokens) return null;
909
941
  if (prune !== void 0) {
910
942
  prune.pruneSession(agent.session);
@@ -29,9 +29,15 @@ export declare function resolveConfig(config?: BasicCompactionConfig): ResolvedC
29
29
  export declare function resolveTargetPolicy(config: ResolvedConfig, target: Pick<LlmCallConfig, 'provider' | 'model'>): ResolvedTargetPolicy;
30
30
  /**
31
31
  * Scale one routed policy into concrete token budgets for its model capacity.
32
+ *
33
+ * Pressure is capped by both the window fraction and the capacity remaining
34
+ * after the routed output reservation plus compaction headroom. Retention scales
35
+ * the message budget before headroom is deducted.
36
+ *
32
37
  * @param policy - merged policy for the exact routed target.
33
38
  * @param contextWindow - positive adapter-owned capacity for that target.
39
+ * @param reservedCompletionTokens - output tokens one routed request reserves.
34
40
  * @returns detached immutable pressure and retention budgets.
35
41
  */
36
- export declare function resolveCompactSpec(policy: ResolvedTargetPolicy, contextWindow: number): ResolvedCompactSpec;
42
+ export declare function resolveCompactSpec(policy: ResolvedTargetPolicy, contextWindow: number, reservedCompletionTokens: number): ResolvedCompactSpec;
37
43
  //# sourceMappingURL=config.d.ts.map
@@ -6,9 +6,11 @@
6
6
  import type { LlmCallConfig } from '@deepseek-ai/dsh-llm';
7
7
  /** Policy fields shared by the default policy and exact model overrides. */
8
8
  export interface CompactionPolicyConfig {
9
- /** Compact at this fraction of the model's context window. Defaults to `0.8`. */
9
+ /** Window fraction for pressure; capped at context window minus reserved output and `headroomTokens`. Defaults to `0.8`. */
10
10
  thresholdRatio?: number;
11
- /** Recent context retained as a fraction of the model's window. Defaults to `0.16`. */
11
+ /** Additional pressure headroom beyond the routed output reservation. Non-negative integer; defaults to `65536`. */
12
+ headroomTokens?: number;
13
+ /** Recent context retained as a fraction of context window minus reserved output tokens. Defaults to `0.16`. */
12
14
  retainRatio?: number;
13
15
  /** Absolute recent-context budget; mutually exclusive with `retainRatio`. */
14
16
  retainTokens?: number;
@@ -16,7 +18,7 @@ export interface CompactionPolicyConfig {
16
18
  summarizationProvider?: string;
17
19
  /** Summary model; set together with `summarizationProvider`, or inherit the conversation target. */
18
20
  summarizationModel?: string;
19
- /** Provider generation cap for summarization. Defaults to `8192`. */
21
+ /** Provider generation cap for summarization. Defaults to the resolved `headroomTokens`; an explicit cap must be positive. */
20
22
  maxTokens?: number;
21
23
  /** Extra attempts after the first compaction when pressure remains above threshold. Defaults to `1`. */
22
24
  compactionRetries?: number;
@@ -48,6 +50,7 @@ export type ResolvedRetention = {
48
50
  /** Validated policy fields shared before and after exact-target matching. */
49
51
  interface ResolvedPolicyFields {
50
52
  readonly thresholdRatio: number;
53
+ readonly headroomTokens: number;
51
54
  readonly summarizationProvider: string;
52
55
  readonly summarizationModel: string;
53
56
  readonly maxTokens: number;
@@ -64,7 +67,8 @@ export type ResolvedTargetPolicy = ResolvedPolicyFields & ResolvedRetention & {
64
67
  readonly target: Pick<LlmCallConfig, 'provider' | 'model'>;
65
68
  };
66
69
  /** One routed model's concrete pressure and retention budget. */
67
- export type ResolvedCompactSpec = Omit<ResolvedTargetPolicy, 'retainRatio' | 'retainTokens'> & {
70
+ export type ResolvedCompactSpec = Omit<ResolvedTargetPolicy, 'retainRatio' | 'retainTokens' | 'headroomTokens'> & {
71
+ /** Adapter-declared full window; token budgets below exclude reserved output tokens. */
68
72
  readonly contextWindow: number;
69
73
  readonly thresholdTokens: number;
70
74
  readonly retainTokens: number;
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "@deepseek-ai/dsh-compaction-basic",
3
3
  "description": "Token-meter-driven compaction policy and LLM summarization backend for the DeepSeek Harness",
4
- "version": "0.1.6-alpha.1",
4
+ "version": "0.1.7-alpha.1",
5
5
  "publishConfig": {
6
6
  "access": "public"
7
7
  },
@@ -27,14 +27,14 @@
27
27
  ],
28
28
  "license": "MIT",
29
29
  "peerDependencies": {
30
- "@deepseek-ai/cordis": "^4.0.2",
31
- "@deepseek-ai/dsh-agent": "^0.1.6-alpha.1",
32
- "@deepseek-ai/dsh-compaction": "^0.1.6-alpha.1",
33
- "@deepseek-ai/dsh-llm": "^0.1.6-alpha.1",
34
- "@deepseek-ai/dsh-session": "^0.1.6-alpha.1",
35
- "@deepseek-ai/dsh-compaction-tool-result-pruner": "^0.1.6-alpha.1",
36
- "@deepseek-ai/dsh-token-meter": "^0.1.6-alpha.1",
37
- "@deepseek-ai/dsh-commands": "^0.1.6-alpha.1"
30
+ "@deepseek-ai/cordis": "^4.0.3",
31
+ "@deepseek-ai/dsh-agent": "^0.1.7-alpha.1",
32
+ "@deepseek-ai/dsh-compaction": "^0.1.7-alpha.1",
33
+ "@deepseek-ai/dsh-compaction-tool-result-pruner": "^0.1.7-alpha.1",
34
+ "@deepseek-ai/dsh-llm": "^0.1.7-alpha.1",
35
+ "@deepseek-ai/dsh-commands": "^0.1.7-alpha.1",
36
+ "@deepseek-ai/dsh-session": "^0.1.7-alpha.1",
37
+ "@deepseek-ai/dsh-token-meter": "^0.1.7-alpha.1"
38
38
  },
39
39
  "peerDependenciesMeta": {
40
40
  "@deepseek-ai/dsh-compaction-tool-result-pruner": {
@@ -42,26 +42,26 @@
42
42
  }
43
43
  },
44
44
  "dependencies": {
45
- "@deepseek-ai/dsh-util-values": "^0.1.6-alpha.1",
46
- "@deepseek-ai/schemastery": "^3.18.2"
45
+ "@deepseek-ai/dsh-util-values": "^0.1.7-alpha.1",
46
+ "@deepseek-ai/schemastery": "^3.18.3"
47
47
  },
48
48
  "devDependencies": {
49
- "@deepseek-ai/cordis": "^4.0.2",
50
- "@deepseek-ai/cordis-plugin-include": "^1.0.7",
51
- "@deepseek-ai/cordis-plugin-loader": "^1.0.3",
52
- "@deepseek-ai/dsh-agent": "^0.1.6-alpha.1",
53
- "@deepseek-ai/dsh-agent-loop": "^0.1.6-alpha.1",
54
- "@deepseek-ai/dsh-commands": "^0.1.6-alpha.1",
55
- "@deepseek-ai/dsh-agent-loop-testkit": "^0.1.6-alpha.1",
56
- "@deepseek-ai/dsh-compaction-tool-result-pruner": "^0.1.6-alpha.1",
57
- "@deepseek-ai/dsh-invariants": "^0.1.6-alpha.1",
58
- "@deepseek-ai/dsh-llm": "^0.1.6-alpha.1",
59
- "@deepseek-ai/dsh-compaction": "^0.1.6-alpha.1",
60
- "@deepseek-ai/dsh-session": "^0.1.6-alpha.1",
61
- "@deepseek-ai/dsh-llm-retry": "^0.1.6-alpha.1",
62
- "@deepseek-ai/dsh-token-meter": "^0.1.6-alpha.1",
63
- "@deepseek-ai/dsh-session-projection": "^0.1.6-alpha.1",
64
- "@deepseek-ai/dsh-tools": "^0.1.6-alpha.1",
65
- "@deepseek-ai/dsh-compaction-image-offload": "^0.1.6-alpha.1"
49
+ "@deepseek-ai/cordis": "^4.0.3",
50
+ "@deepseek-ai/cordis-plugin-include": "^1.0.8",
51
+ "@deepseek-ai/dsh-agent": "^0.1.7-alpha.1",
52
+ "@deepseek-ai/cordis-plugin-loader": "^1.0.4",
53
+ "@deepseek-ai/dsh-agent-loop": "^0.1.7-alpha.1",
54
+ "@deepseek-ai/dsh-agent-loop-testkit": "^0.1.7-alpha.1",
55
+ "@deepseek-ai/dsh-commands": "^0.1.7-alpha.1",
56
+ "@deepseek-ai/dsh-compaction-tool-result-pruner": "^0.1.7-alpha.1",
57
+ "@deepseek-ai/dsh-compaction": "^0.1.7-alpha.1",
58
+ "@deepseek-ai/dsh-invariants": "^0.1.7-alpha.1",
59
+ "@deepseek-ai/dsh-llm": "^0.1.7-alpha.1",
60
+ "@deepseek-ai/dsh-llm-retry": "^0.1.7-alpha.1",
61
+ "@deepseek-ai/dsh-session": "^0.1.7-alpha.1",
62
+ "@deepseek-ai/dsh-session-projection": "^0.1.7-alpha.1",
63
+ "@deepseek-ai/dsh-token-meter": "^0.1.7-alpha.1",
64
+ "@deepseek-ai/dsh-tools": "^0.1.7-alpha.1",
65
+ "@deepseek-ai/dsh-compaction-image-offload": "^0.1.7-alpha.1"
66
66
  }
67
67
  }