dsh-dupguard 1.6.2 → 1.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,44 @@
2
2
 
3
3
  本文件遵循 [Keep a Changelog](https://keepachangelog.com/zh-CN/1.1.0/),版本号遵循 [SemVer](https://semver.org/lang/zh-CN/)。
4
4
 
5
+ ## [1.7.0] - 2026-09-24
6
+
7
+ ### Added
8
+
9
+ - **多字符片段白名单 `ignoredSubstrings`**(默认 `[]`,显式开启):
10
+ - 整段**字面量**匹配(区分大小写、不支持正则),命中时先整段剔除,再做去空白与逐字符剔除;
11
+ 用于 `|---|`、`<br>` 这类由多个字符组成的固定片段(逐字符白名单只能忽略单个字符);
12
+ - **长片段优先**匹配,避免 `---` 抢先破坏 `-----`;
13
+ - **跨增量安全**:每块最多保留(实际最长片段 − 1)个码点,等下一增量拼回后再匹配;
14
+ 块结束/流结束时补投尾巴,尾部文本仍参与检测;
15
+ - 上限:每项 ≤ 64 码点、最多 64 项;空值/超长/超量条目运行时丢弃并各告警一次(同样内容只告警一次);
16
+ - 片段表为空时走零开销快路径,默认行为与未提供该功能时**逐字节一致**;
17
+ - 设置页新增「忽略片段」列表(整段加入、不按码点拆分,含空值/超长/重复/超量校验),
18
+ 与字符白名单共用写通道,可持久化并可被「恢复默认」清空。
19
+
20
+ ### Notes
21
+
22
+ - 实测开销(10 万增量近失配流):空表 2.6 µs/增量 → 配置 4 条不匹配片段 3.3 µs/增量(约 1.3×);
23
+ - 双入口(`lib/index.js` / `plugin/host.js`)实现保持一致,共享用例分别驱动两边;
24
+ - 宿主启动日志新增一行**生效参数清单**(阈值 / 窗口 / 最大单元 / 字符与片段白名单项数),
25
+ 设置页底部常驻「设置通道 + 构建标记」诊断行——用于区分「宿主未加载新代码」「设置没下发」「实现有问题」三类故障。
26
+
27
+ ## [1.6.3] - 2026-09-24
28
+
29
+ ### Fixed
30
+
31
+ - **写入成功却被判为「保存失败」**:`remote.settings.mutate()` 返回的命名空间视图是**写入前**的快照
32
+ (配置由 loader 异步重载后才更新),此前的实现拿它与本地期望值做回读比对,于是每次写入都
33
+ 必然不一致——面板报错、看起来「没保存」,而宿主其实已经落盘。现改为**只以宿主应答为准**
34
+ (`response.ok === false` 才算失败),与官方 `ConfigForm.mutate` 的语义一致;恢复默认同样按
35
+ 「全部 unset 是否未被拒绝」判定。
36
+
37
+ ### Added
38
+
39
+ - 设置页底部常驻诊断行 `设置通道:<状态>|<诊断>|构建 <标记>`,并在写入失败时附带宿主原始原因
40
+ 与构建标记;浏览器控制台同步打印 `[dupguard] mutate → …` / `mutate 被拒…`,用于快速区分
41
+ 「宿主拒绝」「通道未就绪」「回读不一致」三类问题。
42
+
5
43
  ## [1.6.2] - 2026-09-24
6
44
 
7
45
  ### Fixed
package/README.md CHANGED
@@ -49,6 +49,8 @@ When triggered, the already-generated text is committed as a normal assistant me
49
49
  - **真正的服务端停止**:提前关闭流迭代 → 适配器 `consumer.abort()` → 中断 HTTP 连接,模型在服务端停止生成。
50
50
  - **安全停止**:绝不 `abort()` agent 步骤信号;补发协议合规的 `block-end` + `finish(stop)`,消息正常提交。
51
51
  - **Markdown 表格友好**:默认忽略连字符与竖线(`ignoredChars` 白名单),表格分隔行与长分隔线不会被误判为复读。
52
+ - **片段白名单**:`ignoredSubstrings` 可按**整段**忽略多字符片段(如 `|---|`、`<br>`)——逐字符白名单只能
53
+ 忽略单个字符,组合片段的重复仍会被计入;命中时先整段剔除再走常规清洗,且跨增量切分也能正确剔除。
52
54
  - **图形化设置页**(npm 常驻版):在 DSH 设置面板注册与「通用设置 / 模型 / 插件 / Agent 预设」并列的
53
55
  「重复守卫」分节,可视化编辑白名单与全部检测参数(阈值、最小/最大单元长度、检测窗口、空白与
54
56
  reasoning、工具参数开关)并持久化(`dsh-dupguard` 设置命名空间),修改即时生效;窗口小于
@@ -198,6 +200,7 @@ and persisted to `settings.yaml`; the dynamic build uses the constants.
198
200
  | `codeBlockMultiplier` | `3` | 代码块内处理分三档:**≥2** 按「阈值 × 本倍数」判定(越大越不易误杀正常代码,但也越晚兜住块内失控复读);**1** 与块外同样严格;**0** 完全不检测代码块内。范围 0–100 / how code blocks are handled: >=2 = threshold x this multiplier, 1 = same as outside, 0 = do not detect inside code blocks at all. Range 0–100 |
199
201
  | `stripWhitespace` | `true` | 检测前移除空白/换行,识别带分隔符的复读 / strip whitespace so `"x x x"` and `"x\nx\nx"` are caught |
200
202
  | `ignoredChars` | `['-', '\|']` | 检测时忽略的字符白名单:Markdown 表格分隔行(连字符与竖线)不参与重复统计。条目必须**是单个字符**(按 Unicode 码点匹配,emoji 也算一个);设置页一次输入多个字符会逐个加入,多字符/空条目在运行时被丢弃并告警 / whitelist of characters ignored during detection, so Markdown table separators don't count. Entries must be a **single character** (matched per Unicode code point); the settings page splits multi-character input into individual entries, and invalid entries are dropped at runtime with a warning |
203
+ | `ignoredSubstrings` | `[]` | **片段白名单(多字符)**:整段字面量匹配(区分大小写、不支持正则),命中时**先整段剔除**,再做去空白与逐字符剔除;长片段优先。用于 `\|---\|`、`<br>` 这类由多个字符组成的固定片段。每项 ≤ 64 码点、最多 64 项,超限项运行时丢弃并告警。代价:为跨增量匹配,每块最多保留(最长片段 − 1)个字符不参与检测 / whitelist of **multi-character substrings**, matched literally (case-sensitive, no regex); matches are stripped first, then whitespace and per-character rules apply, longest first. Entries ≤ 64 code points, at most 64 entries (invalid entries dropped with a warning). Cost: up to (longest entry − 1) trailing characters per block are held back for cross-delta matching |
201
204
  | `monitorReasoning` | `true` | 是否检测思考文本(思考中的复读同样消耗 token,默认截停;只检测可见输出时置 `false`)/ also guard reasoning (thinking) text — on by default; set `false` to guard visible output only |
202
205
  | `monitorToolArguments` | `false` | 是否检测工具调用参数 / also guard tool-call JSON args — off by default (base64/JSON repeats are common) |
203
206
  | `fixStandingMountConflict` | `true` | DSH ≤ 0.1.6-alpha.2 兼容补丁:幂等化 `cordisInspect.register`,修复 preset standing-mount 多代并存冲突(仅代码常量)/ idempotent `cordisInspect.register` patch for the DSH ≤ 0.1.6-alpha.2 standing-mount conflict (code constant only) |
@@ -225,6 +228,8 @@ Listens to the `llm/stream` waterfall (wraps every streaming model call) and ret
225
228
  ### 2. 检测算法 / Detection
226
229
 
227
230
  - 按块索引(`chunk.index`)分别累积文本,多块交替输出互不干扰;
231
+ - 清洗顺序:**先按片段白名单整段剔除**(`ignoredSubstrings`,长片段优先,跨增量尾巴由块结束时的补投兜住)
232
+ → 再去空白(可关)→ 最后按单字符白名单剔除(`ignoredChars`)。片段逐次替换为空串,非正则匹配;
228
233
  - 去空白后做**尾部连续重复检测**:文本以某个单元(长度 `minUnitLength`..`maxUnitLength`,默认 1..80)
229
234
  连续重复 ≥ `threshold` 次结尾即触发。模型一旦复读,重复必然在尾部,因此尾部检测即可实时捕获所有
230
235
  循环,同时避免全窗口词频的误报(如正常中文里高频的"的")。
@@ -280,9 +285,9 @@ default), Markdown table separator rows and horizontal rules (whitelisted by def
280
285
  │ ├── index.js # npm/组合常驻形式(package.json main 入口,含设置集成)
281
286
  │ └── client.js # 浏览器端设置页(ModuleLoader 格式,dsh.client 入口)
282
287
  ├── tests/
283
- │ ├── detector.test.js # 端到端测试:双入口防漂移 + reasoning 开关 + settings/Config 集成(79 项)
284
- │ ├── client.test.js # 设置页组件测试:最小 React/DSH 桩(旧版 settingsScope + 新版 remote,23 项)
285
- │ ├── stress-host-adversarial.js # 压力:边界/协议交错/热更新 churn/畸形输入
288
+ │ ├── detector.test.js # 端到端测试:双入口防漂移 + reasoning 开关 + settings/Config 集成(82 项)
289
+ │ ├── client.test.js # 设置页组件测试:最小 React/DSH 桩(旧版 settingsScope + 新版 remote,29 项)
290
+ │ ├── stress-host-adversarial.js # 压力:边界/协议交错/围栏与片段白名单/热更新 churn/畸形输入
286
291
  │ ├── stress-host-throughput.js # 压力:吞吐/内存/200 路并发/参数极值(METRIC 指标)
287
292
  │ ├── stress-client-ui.js # 压力:设置页高频交互、乱序应答、挂载泄漏
288
293
  │ ├── stress-real-invariant.mjs # 压力:真实 DSH llm-invariant + BlockAssembler 端到端校验
@@ -300,8 +305,8 @@ default), Markdown table separator rows and horizontal rules (whitelisted by def
300
305
 
301
306
  ```bash
302
307
  npm test # 功能测试(两个文件)
303
- node tests/detector.test.js # 检测端到端(79 项)
304
- node tests/client.test.js # 设置页组件(25 项)
308
+ node tests/detector.test.js # 检测端到端(82 项)
309
+ node tests/client.test.js # 设置页组件(29 项)
305
310
  ```
306
311
 
307
312
  同一套用例分别驱动两个入口(`plugin/host.js` 经 `new Function` 求值、`lib/index.js` 经
@@ -353,6 +358,9 @@ when DSH is not installed. CI runs on Node 20/22/24 (matching DSH; Node 18 is no
353
358
  - **代码块分档只覆盖围栏代码块**:行内代码(`` `x` ``)与缩进代码块(4 空格)仍按普通阈值判定;
354
359
  模型忘记闭合围栏时,其后内容一律按代码块处理。倍数为 `0`(完全不检测)时,代码块内的失控复读
355
360
  不会被截停——这是「代码再长也不误杀」的代价;默认倍数 `3` 则会在「阈值 × 3」处兜底。
361
+ - **片段白名单的代价**:为跨增量匹配,启用后每块最多保留(最长片段 − 1)个字符不参与检测(块结束时补投),
362
+ 即检测最多延迟这么多字符;条目上限 64 项 × 64 字符,逐条字面量替换、不支持正则;片段表为空时走
363
+ 零开销快路径(默认不启用,行为与未提供该功能时一致)。
356
364
  - 检测窗口上限 1,048,576 字符:每个增量都要重写一次缓冲,成本随窗口线性增长——缓冲填满后
357
365
  1 MiB 窗口约 0.13 ms/增量,实测 4 字符增量的平均值为 33 µs/增量(含缓冲填充期)。默认 8192 无感
358
366
  (1.9 µs/增量,模型侧毫秒级的 token 间隔下可忽略)。
package/lib/client.js CHANGED
@@ -26,6 +26,8 @@ window.__ModuleLoader__.load({
26
26
  const React = require('react')
27
27
 
28
28
  const NS = 'dsh-dupguard'
29
+ /** 构建标记:显示在设置页底部,用于判断浏览器实际加载的是哪一版(排查缓存/旧包问题)。 */
30
+ const BUILD_MARK = '1.7.0'
29
31
  /** 旧版(≤ 0.1.6)register 出来的命名空间名。 */
30
32
  const SETTINGS_NS = 'dsh-dupguard'
31
33
  /** 新版(≥ 0.1.7)命名空间取自 loader entry id:本 bundle 插入的行 id 为 dupguard。 */
@@ -41,6 +43,15 @@ window.__ModuleLoader__.load({
41
43
  empty: '白名单为空:所有字符都参与重复统计。',
42
44
  addPlaceholder: '输入要忽略的字符(可多个)',
43
45
  add: '添加',
46
+ substrings: '忽略片段(多字符,按整段匹配)',
47
+ substringsEmpty: '片段白名单为空:没有整段被忽略的字符串。',
48
+ substringAddPlaceholder: '输入要整段忽略的字符串(如 |---|)',
49
+ substringAdd: '添加片段',
50
+ substringHint: '整段字面量匹配(区分大小写、不支持正则):命中时先整段剔除,再做去空白与逐字符剔除。长片段优先。每项 ≤ {maxLen} 字符、最多 {maxCount} 项。',
51
+ errSubstringEmpty: '请输入非空片段。',
52
+ errSubstringTooLong: '片段过长:每项最多 {maxLen} 个字符。',
53
+ errSubstringDuplicate: '该片段已在列表中。',
54
+ errSubstringLimit: '片段数量已达上限({maxCount} 项)。',
44
55
  params: '检测参数',
45
56
  threshold: '触发阈值(连续重复次数)',
46
57
  thresholdHint: '同一字符串连续重复达到该次数即截停(≥ 语义)。范围 {min}–{max}。',
@@ -83,6 +94,15 @@ window.__ModuleLoader__.load({
83
94
  empty: 'Whitelist is empty: every character counts.',
84
95
  addPlaceholder: 'Characters to ignore (one or more)',
85
96
  add: 'Add',
97
+ substrings: 'Ignored substrings (multi-character, matched whole)',
98
+ substringsEmpty: 'No ignored substrings: nothing is stripped as a whole.',
99
+ substringAddPlaceholder: 'String to ignore as a whole (e.g. |---|)',
100
+ substringAdd: 'Add substring',
101
+ substringHint: 'Literal match of the whole substring (case-sensitive, no regex): matches are stripped first, then whitespace and per-character rules apply. Longer substrings win. Each entry ≤ {maxLen} characters, at most {maxCount} entries.',
102
+ errSubstringEmpty: 'Enter a non-empty substring.',
103
+ errSubstringTooLong: 'Substring too long: at most {maxLen} characters each.',
104
+ errSubstringDuplicate: 'That substring is already in the list.',
105
+ errSubstringLimit: 'Substring limit reached ({maxCount} entries).',
86
106
  params: 'Detection parameters',
87
107
  threshold: 'Threshold (consecutive repeats)',
88
108
  thresholdHint: 'Stop once a string repeats this many times in a row (>= semantics). Range {min}–{max}.',
@@ -175,6 +195,7 @@ window.__ModuleLoader__.load({
175
195
  const BOOL_FIELDS = ['stripWhitespace', 'skipCodeBlocks', 'monitorReasoning', 'monitorToolArguments']
176
196
  const DEFAULTS = {
177
197
  ignoredChars: ['-', '|'],
198
+ ignoredSubstrings: [],
178
199
  threshold: 10,
179
200
  codeBlockMultiplier: 3,
180
201
  minUnitLength: 1,
@@ -185,7 +206,11 @@ window.__ModuleLoader__.load({
185
206
  monitorReasoning: true,
186
207
  monitorToolArguments: false,
187
208
  }
188
- const ALL_FIELDS = ['ignoredChars'].concat(NUMERIC_FIELDS.map((field) => field.key), BOOL_FIELDS)
209
+ const ALL_FIELDS = ['ignoredChars', 'ignoredSubstrings'].concat(NUMERIC_FIELDS.map((field) => field.key), BOOL_FIELDS)
210
+
211
+ /** 与宿主 lib/index.js 的常量保持一致:片段上限(每项码点数 / 条目数)。 */
212
+ const MAX_SUBSTRING_LENGTH = 64
213
+ const MAX_SUBSTRINGS = 64
189
214
 
190
215
  /** 模板占位符替换:{name} → vars.name。 */
191
216
  const fmt = (text, vars) =>
@@ -196,7 +221,10 @@ window.__ModuleLoader__.load({
196
221
  * 保证任何异常形状到达视图时仍是 ready 且可渲染。
197
222
  */
198
223
  function normalizeValue(raw) {
199
- const out = { ignoredChars: [...DEFAULTS.ignoredChars] }
224
+ const out = {
225
+ ignoredChars: [...DEFAULTS.ignoredChars],
226
+ ignoredSubstrings: [...DEFAULTS.ignoredSubstrings],
227
+ }
200
228
  for (const field of NUMERIC_FIELDS) out[field.key] = DEFAULTS[field.key]
201
229
  for (const key of BOOL_FIELDS) out[key] = DEFAULTS[key]
202
230
  try {
@@ -204,6 +232,11 @@ window.__ModuleLoader__.load({
204
232
  if (Array.isArray(raw.ignoredChars)) {
205
233
  out.ignoredChars = raw.ignoredChars.filter((ch) => typeof ch === 'string')
206
234
  }
235
+ if (Array.isArray(raw.ignoredSubstrings)) {
236
+ out.ignoredSubstrings = raw.ignoredSubstrings
237
+ .filter((item) => typeof item === 'string' && item.length > 0)
238
+ .slice(0, MAX_SUBSTRINGS)
239
+ }
207
240
  for (const field of NUMERIC_FIELDS) {
208
241
  const value = raw[field.key]
209
242
  if (Number.isSafeInteger(value) && value >= field.min && value <= field.max) out[field.key] = value
@@ -219,7 +252,10 @@ window.__ModuleLoader__.load({
219
252
  /** 归一化值为编辑表单:数值字段存字符串,便于输入中途的状态。 */
220
253
  function toForm(value) {
221
254
  const normalized = normalizeValue(value)
222
- const form = { ignoredChars: normalized.ignoredChars }
255
+ const form = {
256
+ ignoredChars: normalized.ignoredChars,
257
+ ignoredSubstrings: normalized.ignoredSubstrings,
258
+ }
223
259
  for (const field of NUMERIC_FIELDS) form[field.key] = String(normalized[field.key])
224
260
  for (const key of BOOL_FIELDS) form[key] = normalized[key]
225
261
  return form
@@ -248,6 +284,7 @@ window.__ModuleLoader__.load({
248
284
  const [form, setForm] = React.useState(null)
249
285
  const [writeState, setWriteState] = React.useState(null)
250
286
  const [errors, setErrors] = React.useState({})
287
+ const [draftSub, setDraftSub] = React.useState('')
251
288
  const dirty = React.useRef({})
252
289
  const lastRemote = React.useRef(null)
253
290
 
@@ -344,28 +381,35 @@ window.__ModuleLoader__.load({
344
381
  }
345
382
  : null
346
383
 
347
- // ---- 写路径:settingsScope 控制器(DSH 自己的设置写通道,自动携带 revision、
348
- // 串行化并发写、把宿主应答折叠回镜像)。DSH 0.1.2 起客户端不再暴露
349
- // connection.api,控制器接口自 0.1.1 起稳定,故不直接触碰 wire 面。 ----
384
+ // ---- 写路径:设置控制器(新版 remote.settings,旧版 settingsScope)。 ----
385
+ // 写成功与否**只以宿主应答为准**:宿主返回的视图是写入前的快照(配置由 loader
386
+ // 异步重载),拿它跟本地期望值比对必然不一致——这正是此前「明明保存成功却报
387
+ // 保存失败、重开后新值不见了」的原因。官方实现同样是信任 response.ok。
350
388
  const snapshotValue = () => {
351
389
  const current = controller.getSnapshot()
352
390
  return current !== undefined && current.value !== undefined ? current.value : {}
353
391
  }
354
- const sameList = (a, b) => Array.isArray(a) && Array.isArray(b) &&
355
- a.length === b.length && a.every((item, index) => item === b[index])
356
392
  const currentForm = () => (form === null ? remoteForm : form)
357
- const runWrite = (operation, verify) => {
393
+ const runWrite = (operation) => {
358
394
  setWriteState('saving')
395
+ // 失败时带上宿主返回的原始原因(writeReason)与构建标记:
396
+ // 标记同时用于判断浏览器加载的是哪一版(排查旧包/缓存)。
397
+ const failureText = () => {
398
+ const reason = typeof props.writeReason === 'function' ? props.writeReason() : null
399
+ const suffix = reason === null || reason === undefined ? '(宿主未给出原因)' : String(reason)
400
+ return t('saveFailed') + ' [diag ' + BUILD_MARK + '] ' + suffix
401
+ }
359
402
  Promise.resolve()
360
403
  .then(() => operation())
361
- .then(() => setWriteState(verify() ? 'saved' : 'error:' + t('saveFailed')), (error) => {
404
+ .then((accepted) => setWriteState(accepted === false ? 'error:' + failureText() : 'saved'), (error) => {
362
405
  setWriteState('error:' + String((error && error.message) || error))
363
406
  })
364
407
  }
365
- const commitList = (next) => {
366
- dirty.current.ignoredChars = true
367
- setForm({ ...currentForm(), ignoredChars: next })
368
- runWrite(() => controller.set('ignoredChars', next), () => sameList(snapshotValue().ignoredChars, next))
408
+ /** 通用列表字段写入(字符白名单 / 片段白名单共用)。 */
409
+ const commitFieldList = (field, next) => {
410
+ dirty.current[field] = true
411
+ setForm({ ...currentForm(), [field]: next })
412
+ runWrite(() => controller.set(field, next))
369
413
  }
370
414
  const commitNumber = (field) => {
371
415
  const message = fieldError(field)
@@ -376,7 +420,7 @@ window.__ModuleLoader__.load({
376
420
  const value = parsed[field.key]
377
421
  setErrors((prev) => ({ ...prev, [field.key]: null }))
378
422
  // 值未变化时不写入:避免仅仅聚焦/失焦就把默认值写进用户层,
379
- // 污染 settings.yaml 并让「恢复默认」失去意义。
423
+ // 污染配置文件并让「恢复默认」失去意义。
380
424
  if (snapshotValue()[field.key] === value) {
381
425
  dirty.current[field.key] = false
382
426
  setForm({ ...currentForm(), [field.key]: String(value) })
@@ -384,27 +428,20 @@ window.__ModuleLoader__.load({
384
428
  }
385
429
  dirty.current[field.key] = true
386
430
  setForm({ ...currentForm(), [field.key]: String(value) })
387
- runWrite(() => controller.set(field.key, value), () => snapshotValue()[field.key] === value)
431
+ runWrite(() => controller.set(field.key, value))
388
432
  }
389
433
  const commitBool = (key) => {
390
434
  const value = shown[key] !== true
391
435
  dirty.current[key] = true
392
436
  setForm({ ...currentForm(), [key]: value })
393
- runWrite(() => controller.set(key, value), () => snapshotValue()[key] === value)
437
+ runWrite(() => controller.set(key, value))
394
438
  }
395
439
  const resetAll = () => {
396
440
  dirty.current = {}
397
441
  setErrors({})
398
- runWrite(
399
- () => Promise.all(ALL_FIELDS.map((key) => controller.unset(key))),
400
- () => {
401
- setForm(toForm(snapshotValue()))
402
- const current = controller.getSnapshot()
403
- const user = current !== undefined ? current.user : undefined
404
- if (user === undefined || user === null) return true
405
- return ALL_FIELDS.every((key) => user[key] === undefined)
406
- },
407
- )
442
+ // 恢复默认 = 逐字段 unset;接受与否同样只看宿主应答。
443
+ runWrite(() => Promise.all(ALL_FIELDS.map((key) => controller.unset(key)))
444
+ .then((results) => results.every((item) => item !== false)))
408
445
  }
409
446
 
410
447
  const add = () => {
@@ -421,9 +458,40 @@ window.__ModuleLoader__.load({
421
458
  next.push(ch)
422
459
  changed = true
423
460
  }
424
- if (changed) commitList(next)
461
+ if (changed) commitFieldList('ignoredChars', next)
425
462
  }
426
- const remove = (ch) => commitList(shown.ignoredChars.filter((item) => item !== ch))
463
+ const remove = (ch) => commitFieldList('ignoredChars', shown.ignoredChars.filter((item) => item !== ch))
464
+
465
+ /**
466
+ * 片段白名单:与字符白名单不同,**不按码点拆分**——整段就是一个条目。
467
+ * 校验空值/超长/重复/超量,错误就地显示且不写入。
468
+ */
469
+ const addSubstring = () => {
470
+ const text = draftSub.trim()
471
+ const fail = (message) => setErrors((prev) => ({ ...prev, substrings: message }))
472
+ if (text.length === 0) {
473
+ fail(t('errSubstringEmpty'))
474
+ return
475
+ }
476
+ const length = [...text].length
477
+ if (length > MAX_SUBSTRING_LENGTH) {
478
+ fail(fmt(t('errSubstringTooLong'), { maxLen: MAX_SUBSTRING_LENGTH }))
479
+ return
480
+ }
481
+ if (shown.ignoredSubstrings.indexOf(text) !== -1) {
482
+ fail(t('errSubstringDuplicate'))
483
+ return
484
+ }
485
+ if (shown.ignoredSubstrings.length >= MAX_SUBSTRINGS) {
486
+ fail(fmt(t('errSubstringLimit'), { maxCount: MAX_SUBSTRINGS }))
487
+ return
488
+ }
489
+ setDraftSub('')
490
+ setErrors((prev) => ({ ...prev, substrings: null }))
491
+ commitFieldList('ignoredSubstrings', [...shown.ignoredSubstrings, text])
492
+ }
493
+ const removeSubstring = (item) =>
494
+ commitFieldList('ignoredSubstrings', shown.ignoredSubstrings.filter((entry) => entry !== item))
427
495
 
428
496
  const fieldRow = (field) => {
429
497
  const message = errors[field.key] !== undefined && errors[field.key] !== null
@@ -497,6 +565,36 @@ window.__ModuleLoader__.load({
497
565
  React.createElement('button', { className: 'dg-btn', type: 'button', onClick: add }, t('add')),
498
566
  ),
499
567
 
568
+ React.createElement('p', { className: 'dg-note' }, t('substrings')),
569
+ shown.ignoredSubstrings.length === 0
570
+ ? React.createElement('p', { className: 'dg-empty' }, t('substringsEmpty'))
571
+ : React.createElement('div', { className: 'dg-chips' },
572
+ shown.ignoredSubstrings.map((item) => React.createElement('span', { className: 'dg-chip', key: item },
573
+ item,
574
+ React.createElement('button', {
575
+ className: 'dg-chip-remove',
576
+ type: 'button',
577
+ onClick: () => removeSubstring(item),
578
+ 'aria-label': 'remove',
579
+ }, '\u00d7')))),
580
+ React.createElement('div', { className: 'dg-add' },
581
+ React.createElement('input', {
582
+ className: 'dg-input',
583
+ placeholder: t('substringAddPlaceholder'),
584
+ value: draftSub,
585
+ onChange: (event) => setDraftSub(event.target.value),
586
+ onKeyDown: (event) => {
587
+ if (event.key === 'Enter') addSubstring()
588
+ },
589
+ }),
590
+ React.createElement('button', { className: 'dg-btn', type: 'button', onClick: addSubstring }, t('substringAdd')),
591
+ ),
592
+ React.createElement('p', { className: 'dg-field-hint' },
593
+ fmt(t('substringHint'), { maxLen: MAX_SUBSTRING_LENGTH, maxCount: MAX_SUBSTRINGS })),
594
+ errors.substrings === undefined || errors.substrings === null
595
+ ? null
596
+ : React.createElement('p', { className: 'dg-field-error' }, errors.substrings),
597
+
500
598
  React.createElement('p', { className: 'dg-note' }, t('params')),
501
599
  React.createElement('div', { className: 'dg-fields' },
502
600
  NUMERIC_FIELDS.map((field) => fieldRow(field)),
@@ -513,6 +611,11 @@ window.__ModuleLoader__.load({
513
611
  : writeState === 'saved'
514
612
  ? React.createElement('p', { className: 'dg-note' }, t('saved'))
515
613
  : React.createElement('p', { className: 'dg-note dg-error' }, writeState),
614
+ React.createElement('p', { className: 'dg-note' },
615
+ fmt(t('loadingDiag'), {
616
+ state: mirrorSnap ? String(mirrorSnap.status) : 'unknown',
617
+ diag: (mirrorSnap && mirrorSnap.diag ? String(mirrorSnap.diag) : '') + '|构建 ' + BUILD_MARK,
618
+ })),
516
619
  ),
517
620
  )
518
621
  }
@@ -657,6 +760,32 @@ window.__ModuleLoader__.load({
657
760
  if (!closed && mine === generation) loaded = true
658
761
  }
659
762
  }
763
+ /** 写入失败时面板上展示的宿主原始原因(诊断用)。 */
764
+ let lastWriteError = null
765
+
766
+ /** 只尝试属于本插件的命名空间名:describe 报出的 ns 与其去掉组合前缀的形式。 */
767
+ const namespaceCandidates = () => {
768
+ const list = [resolvedNamespace]
769
+ const cut = resolvedNamespace.lastIndexOf(':')
770
+ if (cut !== -1 && cut + 1 < resolvedNamespace.length) {
771
+ const stripped = resolvedNamespace.slice(cut + 1)
772
+ if (list.indexOf(stripped) === -1) list.push(stripped)
773
+ }
774
+ return list
775
+ }
776
+
777
+ const attempt = async (namespace, owned, expected) => {
778
+ let response
779
+ try {
780
+ response = await settings.mutate(namespace, owned, expected)
781
+ } catch (error) {
782
+ return { ok: false, error: errorText(error) }
783
+ }
784
+ if (response === null || typeof response !== 'object') return { ok: false, error: '宿主应答无效' }
785
+ if (response.ok === true) return { ok: true, value: response.value }
786
+ return { ok: false, error: responseError(response, '宿主拒绝了该写入') }
787
+ }
788
+
660
789
  const mutate = (ops) => {
661
790
  const owned = ops.map((op) => (op.op === 'set'
662
791
  ? { op: 'set', path: [...op.path], value: op.value }
@@ -665,26 +794,40 @@ window.__ModuleLoader__.load({
665
794
  if (closed || channel === null) return false
666
795
  // 首次写入前先完成一次 describe:否则命名空间与 revision 都还是初始值。
667
796
  if (!loaded) await load()
668
- let response
669
- try {
670
- response = await settings.mutate(resolvedNamespace, owned, revision)
671
- } catch (_error) {
672
- await load()
673
- return false
797
+ const failures = []
798
+ for (const namespace of namespaceCandidates()) {
799
+ console.info('[dupguard] mutate → ' + namespace + ' revision=' + String(revision) +
800
+ ' paths=' + JSON.stringify(owned.map((op) => op.path)))
801
+ let result = await attempt(namespace, owned, revision)
802
+ if (!result.ok) {
803
+ // 常见于并发写入导致的 revision 过期:回读后按新 revision 重试一次。
804
+ console.warn('[dupguard] mutate 被拒(' + namespace + '):' + result.error + ',回读后重试')
805
+ await load()
806
+ result = await attempt(namespace, owned, revision)
807
+ }
808
+ if (result.ok) {
809
+ lastWriteError = null
810
+ if (!closed) acceptView(result.value)
811
+ publish(undefined, { status: mirror.status, error: null, diag: '已写入命名空间 ' + namespace })
812
+ return true
813
+ }
814
+ failures.push(namespace + ':' + result.error)
674
815
  }
675
- if (response === null || typeof response !== 'object' || response.ok !== true) {
676
- await load() // 写失败:回读宿主真实状态,避免界面停留在错误的本地值
677
- return false
678
- }
679
- if (!closed) acceptView(response.value)
680
- return true
816
+ lastWriteError = failures.join(';')
817
+ // 写失败:回读宿主真实状态,避免界面停留在错误的本地值;并把原因显示到面板。
818
+ await load()
819
+ publish(undefined, { status: mirror.status, error: '写入被拒绝 —— ' + lastWriteError, diag: 'mutate 失败' })
820
+ return false
681
821
  })
682
822
  tail = task.then(() => {}, () => {})
683
823
  return task
684
824
  }
825
+ /** 供组件展示的最近一次写入失败原因。 */
826
+ const writeError = () => lastWriteError
685
827
  return {
686
828
  load,
687
829
  mutate,
830
+ writeError,
688
831
  close: () => {
689
832
  closed = true
690
833
  },
@@ -763,12 +906,20 @@ window.__ModuleLoader__.load({
763
906
  controller: {
764
907
  subscribe,
765
908
  getSnapshot: () => snapshot,
766
- set: (field, value) => (channel === null
767
- ? Promise.resolve(false)
768
- : channel.mutate([{ op: 'set', path: [field], value }])),
769
- unset: (field) => (channel === null
770
- ? Promise.resolve(false)
771
- : channel.mutate([{ op: 'unset', path: [field] }])),
909
+ set: (field, value) => {
910
+ if (channel === null) {
911
+ lastWriteError = '设置通道未就绪(channel=null)'
912
+ return Promise.resolve(false)
913
+ }
914
+ return channel.mutate([{ op: 'set', path: [field], value }])
915
+ },
916
+ unset: (field) => {
917
+ if (channel === null) {
918
+ lastWriteError = '设置通道未就绪(channel=null)'
919
+ return Promise.resolve(false)
920
+ }
921
+ return channel.mutate([{ op: 'unset', path: [field] }])
922
+ },
772
923
  },
773
924
  mirror: {
774
925
  subscribe,
@@ -781,6 +932,8 @@ window.__ModuleLoader__.load({
781
932
  },
782
933
  /** 更新面板上的通道诊断文本(不改变数据状态)。 */
783
934
  note: (text) => publish(undefined, { status: mirror.status, error: mirror.error, diag: text }),
935
+ /** 最近一次写入失败的宿主原因(旧版通道不提供时返回 null)。 */
936
+ writeReason: () => (channel !== null && typeof channel.writeError === 'function' ? channel.writeError() : null),
784
937
  disposed: () => disposed,
785
938
  /** 是否已有可用通道(旧版接上后即可停止新版探测)。 */
786
939
  hasChannel: () => channel !== null,
@@ -921,6 +1074,7 @@ window.__ModuleLoader__.load({
921
1074
  // 传函数而非快照:远程/只读状态可能由 describe 的 writable 或
922
1075
  // remote.$host.isLoopback 稍后确定,渲染时必须取最新值。
923
1076
  isLoopback: () => bridge.isLoopback(),
1077
+ writeReason: () => bridge.writeReason(),
924
1078
  })
925
1079
  ctx.slots.inject('settings.section', () => ctx.slots.register({
926
1080
  name: 'settings.section',
package/lib/index.js CHANGED
@@ -35,6 +35,15 @@ const LIMITS = {
35
35
  codeBlockMultiplier: { min: 0, max: 100 },
36
36
  }
37
37
 
38
+ /**
39
+ * 片段白名单(ignoredSubstrings)的上限。
40
+ *
41
+ * 两者同时界定「跨增量保留的尾巴长度(≤ 最长片段 - 1 码点)」与单增量匹配成本,
42
+ * 因此是代码常量而非可调设置。
43
+ */
44
+ const IGNORED_SUBSTRINGS_MAX_COUNT = 64
45
+ const IGNORED_SUBSTRING_MAX_LENGTH = 64
46
+
38
47
  /**
39
48
  * 把字段标记为 volatile(DSH 的「可热改」约定)。
40
49
  *
@@ -71,6 +80,7 @@ function markVolatile(schema) {
71
80
  function createSettingsSchema(z) {
72
81
  return z.object({
73
82
  ignoredChars: markVolatile(z.array(z.string()).default([...CONFIG.ignoredChars])),
83
+ ignoredSubstrings: markVolatile(z.array(z.string()).default([...CONFIG.ignoredSubstrings])),
74
84
  threshold: markVolatile(z.number().min(LIMITS.threshold.min).max(LIMITS.threshold.max).step(1).default(CONFIG.threshold)),
75
85
  minUnitLength: markVolatile(z.number().min(LIMITS.minUnitLength.min).max(LIMITS.minUnitLength.max).step(1).default(CONFIG.minUnitLength)),
76
86
  maxUnitLength: markVolatile(z.number().min(LIMITS.maxUnitLength.min).max(LIMITS.maxUnitLength.max).step(1).default(CONFIG.maxUnitLength)),
@@ -114,6 +124,13 @@ const CONFIG = {
114
124
  // 模型忘记闭合围栏时,其后内容都按代码块处理。
115
125
  skipCodeBlocks: true,
116
126
  codeBlockMultiplier: 3,
127
+ // 片段白名单:整段匹配的多字符串(字面量、区分大小写、不支持正则)。
128
+ // 命中时先整段剔除,再做去空白与逐字符剔除 —— 用于 `|---|`、`------` 这类
129
+ // 由多个字符组成的固定片段:逐字符白名单只能忽略单个字符,组合片段的重复
130
+ // 仍会被计入。长片段优先匹配,避免 `---` 抢先破坏 `-----`。
131
+ // 上限:每项 ≤ IGNORED_SUBSTRING_MAX_LENGTH 码点、数组 ≤ IGNORED_SUBSTRINGS_MAX_COUNT 项。
132
+ // 代价:为跨增量匹配,每块最多保留(最长片段 - 1)个码点不参与检测(tail 延迟)。
133
+ ignoredSubstrings: [],
117
134
  // 是否同时检测思考(reasoning)文本。默认开启:思考中的复读同样消耗 token,
118
135
  // 应立即截停。注意:正常思考中若连续重复同一字符串 10 次以上(如"等等等等"),
119
136
  // 也会被截停,属预期行为;需要只检测可见输出时可置为 false。
@@ -213,13 +230,87 @@ function stripIgnoredChars(text, ignored) {
213
230
  return out
214
231
  }
215
232
 
216
- /** 供检测使用的增量清洗:去空白 + 移除白名单字符。 */
217
- function sanitizeDelta(delta, config) {
218
- let piece = config.stripWhitespace ? stripWhitespace(delta) : delta
233
+ /** 供检测使用的增量清洗:去空白 + 移除白名单字符(片段剔除在此之前完成)。 */
234
+ function sanitizePiece(text, config) {
235
+ let piece = config.stripWhitespace ? stripWhitespace(text) : text
219
236
  if (config.ignoredChars.length > 0) piece = stripIgnoredChars(piece, config.ignoredChars)
220
237
  return piece
221
238
  }
222
239
 
240
+ /**
241
+ * 增量片段剔除器(每个文本块一份)。
242
+ *
243
+ * 语义:把配置里的片段当**字面量子串**整段剔除,长片段优先(避免 `---` 抢先破坏 `-----`),
244
+ * 剔除后再交给 sanitizePiece(去空白 → 逐字符白名单)。
245
+ *
246
+ * 跨增量:为避免片段被增量边界切断,末尾最多保留(最长片段 - 1)个**码点**不输出,
247
+ * 等下一次增量拼回后再匹配;flush() 吐出尾巴(块结束/流结束时调用),保证尾部文本仍参与检测。
248
+ * 代价:检测最多延迟(最长片段 - 1)个码点。
249
+ *
250
+ * 片段表为空时走零开销快路径(不缓冲、原样返回),因此默认行为与未启用该功能时完全一致。
251
+ *
252
+ * @param getPatterns - 读取当前片段表(支持热更新)。
253
+ * @param maxLength - 片段长度上限(决定保留的尾巴长度)。
254
+ */
255
+ function createSubstringStripper(getPatterns, maxLength) {
256
+ let buffer = ''
257
+ let cachedSource = null
258
+ let cachedSorted = []
259
+ const keepLength = () => Math.max(1, maxLength) - 1
260
+
261
+ /** 长片段优先的片段表(按数组引用缓存,热更新后自动重建)。 */
262
+ const sortedPatterns = () => {
263
+ const list = getPatterns()
264
+ if (list !== cachedSource || list.length !== cachedSorted.length) {
265
+ cachedSource = list
266
+ cachedSorted = [...list].sort((left, right) => [...right].length - [...left].length)
267
+ }
268
+ return cachedSorted
269
+ }
270
+
271
+ /** 从 text 中移除所有片段出现(字面量,非正则)。 */
272
+ const removePatterns = (text, patterns) => {
273
+ let out = text
274
+ for (const pattern of patterns) {
275
+ if (pattern.length === 0) continue
276
+ if (out.indexOf(pattern) !== -1) out = out.split(pattern).join('')
277
+ }
278
+ return out
279
+ }
280
+
281
+ return {
282
+ /** 消费一段增量,返回可交给后续清洗的文本(可能为空串)。 */
283
+ push(delta) {
284
+ const patterns = sortedPatterns()
285
+ if (patterns.length === 0) {
286
+ // 快路径:未配置片段时不缓冲;若此前残留尾巴(刚被清空配置),先吐出。
287
+ if (buffer.length === 0) return delta
288
+ const carried = buffer
289
+ buffer = ''
290
+ return carried + delta
291
+ }
292
+ buffer += delta
293
+ const cleaned = removePatterns(buffer, patterns)
294
+ const codePoints = [...cleaned]
295
+ // 保留长度按**实际最长片段**计算(而非配置上限),否则检测会被无谓地拖后。
296
+ const longest = [...patterns[0]].length
297
+ const keep = Math.min(Math.max(0, longest - 1), keepLength(), codePoints.length)
298
+ if (keep === 0) {
299
+ buffer = ''
300
+ return cleaned
301
+ }
302
+ buffer = codePoints.slice(codePoints.length - keep).join('')
303
+ return codePoints.slice(0, codePoints.length - keep).join('')
304
+ },
305
+ /** 吐出保留的尾巴(块结束/流结束)。 */
306
+ flush() {
307
+ const out = buffer
308
+ buffer = ''
309
+ return out
310
+ },
311
+ }
312
+ }
313
+
223
314
  /**
224
315
  * 流式围栏代码块过滤器(每个文本块一份状态)。
225
316
  *
@@ -416,6 +507,50 @@ function normalizeIgnoredChars(raw, fallback) {
416
507
  /** 上一次「无效白名单条目」告警的签名,用于去重(参数每次调用都会重算)。 */
417
508
  let lastIgnoredCharsWarning = ''
418
509
 
510
+ /**
511
+ * 片段白名单归一化:丢弃空值/非字符串、超长(> 64 码点)与超量(> 64 项)条目,
512
+ * 并去重。丢弃项只告警一次(同样内容),避免每次模型调用刷屏。
513
+ */
514
+ function normalizeIgnoredSubstrings(raw, fallback) {
515
+ if (!Array.isArray(raw)) return [...fallback]
516
+ const kept = []
517
+ const dropped = []
518
+ for (const entry of raw) {
519
+ if (typeof entry !== 'string' || entry.length === 0) {
520
+ dropped.push(String(entry))
521
+ continue
522
+ }
523
+ if ([...entry].length > IGNORED_SUBSTRING_MAX_LENGTH) {
524
+ dropped.push(entry)
525
+ continue
526
+ }
527
+ if (kept.indexOf(entry) !== -1) continue
528
+ if (kept.length >= IGNORED_SUBSTRINGS_MAX_COUNT) {
529
+ dropped.push(entry)
530
+ continue
531
+ }
532
+ kept.push(entry)
533
+ }
534
+ if (dropped.length > 0) {
535
+ const signature = JSON.stringify(dropped)
536
+ if (signature !== lastSubstringWarning) {
537
+ lastSubstringWarning = signature
538
+ console.warn(
539
+ '[dupguard] 片段白名单条目无效,已忽略:' + JSON.stringify(dropped) +
540
+ '(要求非空字符串、每项 ≤ ' + String(IGNORED_SUBSTRING_MAX_LENGTH) +
541
+ ' 码点、最多 ' + String(IGNORED_SUBSTRINGS_MAX_COUNT) + ' 项)。'
542
+ )
543
+ }
544
+ }
545
+ return kept
546
+ }
547
+
548
+ /** 上一次「无效片段白名单条目」告警的签名,用于去重。 */
549
+ let lastSubstringWarning = ''
550
+
551
+ /** 上一次打印的生效参数摘要,用于去重(旧版设置通道每次变更都会回调)。 */
552
+ let lastRuntimeSummary = ''
553
+
419
554
  /**
420
555
  * 为一次 llm/stream 调用创建守卫。
421
556
  * 每次模型调用都会新建一份状态,互不干扰。
@@ -439,6 +574,7 @@ function createStreamGuard(options, runtime) {
439
574
  stripped: '', // 去空白后的滚动窗口:仅用于检测
440
575
  fence: createFenceFilter(), // 围栏代码块状态(skipCodeBlocks 开启时使用)
441
576
  lastRunCode: undefined, // 上一片段的代码块归属(用于跨围栏边界清空缓冲)
577
+ substrings: createSubstringStripper(() => runtime.ignoredSubstrings, IGNORED_SUBSTRING_MAX_LENGTH),
442
578
  toolCallId: undefined,
443
579
  toolCallName: undefined,
444
580
  toolCallArguments: '',
@@ -467,7 +603,7 @@ function createStreamGuard(options, runtime) {
467
603
  if (skipInsideCode && b.lastRunCode !== undefined && run.code !== b.lastRunCode) b.stripped = ''
468
604
  b.lastRunCode = run.code
469
605
  if (skipInsideCode && run.code) continue
470
- const piece = sanitizeDelta(run.text, runtime)
606
+ const piece = sanitizePiece(b.substrings.push(run.text), runtime)
471
607
  if (piece.length === 0) continue
472
608
  b.stripped = (b.stripped + piece).slice(-runtime.detectionWindow)
473
609
  const threshold = run.code ? codeThreshold : runtime.threshold
@@ -477,6 +613,21 @@ function createStreamGuard(options, runtime) {
477
613
  return null
478
614
  }
479
615
 
616
+ /**
617
+ * 块结束时吐出片段剔除器保留的尾巴(≤ 最长片段 - 1 码点),
618
+ * 让它仍参与检测;阈值沿用该块最后一次片段的代码块归属。
619
+ */
620
+ function flushSubstringTail(b) {
621
+ if (b === undefined || b.substrings === undefined) return null
622
+ const tail = sanitizePiece(b.substrings.flush(), runtime)
623
+ if (tail.length === 0) return null
624
+ b.stripped = (b.stripped + tail).slice(-runtime.detectionWindow)
625
+ const threshold = b.lastRunCode === true
626
+ ? runtime.threshold * (runtime.skipCodeBlocks === true ? runtime.codeBlockMultiplier : 1)
627
+ : runtime.threshold
628
+ return findRepeatedTail(b.stripped, threshold, runtime.minUnitLength, runtime.maxUnitLength)
629
+ }
630
+
480
631
  /** 依据 StreamChunk 协议累积状态;命中时置 stopped。 */
481
632
  function feed(chunk) {
482
633
  switch (chunk.type) {
@@ -507,7 +658,7 @@ function createStreamGuard(options, runtime) {
507
658
  if (chunk.name !== undefined) b.toolCallName = chunk.name
508
659
  b.toolCallArguments += chunk.argumentsDelta
509
660
  if (runtime.monitorToolArguments) {
510
- const piece = sanitizeDelta(chunk.argumentsDelta, runtime)
661
+ const piece = sanitizePiece(b.substrings.push(chunk.argumentsDelta), runtime)
511
662
  b.stripped = (b.stripped + piece).slice(-runtime.detectionWindow)
512
663
  const hit = findRepeatedTail(b.stripped, runtime.threshold, runtime.minUnitLength, runtime.maxUnitLength)
513
664
  if (hit !== null) stopped = hit
@@ -515,11 +666,21 @@ function createStreamGuard(options, runtime) {
515
666
  return
516
667
  }
517
668
  case 'block-end': {
669
+ const b = blocks.get(chunk.index)
670
+ const tailHit = flushSubstringTail(b)
671
+ if (tailHit !== null) stopped = tailHit
518
672
  blocks.delete(chunk.index)
519
673
  return
520
674
  }
521
675
  case 'usage':
522
- case 'finish':
676
+ case 'finish': {
677
+ // 流结束时可能还有未收到 block-end 的块:把尾巴补上,避免尾部文本漏检。
678
+ for (const b of blocks.values()) {
679
+ const tailHit = flushSubstringTail(b)
680
+ if (tailHit !== null) stopped = tailHit
681
+ }
682
+ return
683
+ }
523
684
  default:
524
685
  return
525
686
  }
@@ -621,6 +782,7 @@ function installStandingMountPatch(ctx) {
621
782
  function settingsBase() {
622
783
  return {
623
784
  ignoredChars: [...CONFIG.ignoredChars],
785
+ ignoredSubstrings: [...CONFIG.ignoredSubstrings],
624
786
  threshold: CONFIG.threshold,
625
787
  minUnitLength: CONFIG.minUnitLength,
626
788
  maxUnitLength: CONFIG.maxUnitLength,
@@ -714,6 +876,7 @@ function applySettingsToRuntime(runtime, value) {
714
876
  }
715
877
  const pickBool = (key) => (typeof source[key] === 'boolean' ? source[key] : fallback[key])
716
878
  runtime.ignoredChars = normalizeIgnoredChars(source.ignoredChars, fallback.ignoredChars)
879
+ runtime.ignoredSubstrings = normalizeIgnoredSubstrings(source.ignoredSubstrings, fallback.ignoredSubstrings)
717
880
  runtime.threshold = pickInt('threshold')
718
881
  runtime.minUnitLength = pickInt('minUnitLength')
719
882
  runtime.maxUnitLength = pickInt('maxUnitLength')
@@ -748,7 +911,10 @@ function readConfigValue(config, key) {
748
911
 
749
912
  /** 把插件 Config 归一化成 applySettingsToRuntime 需要的设置值对象。 */
750
913
  function settingsValueFromConfig(config) {
751
- const value = { ignoredChars: readConfigValue(config, 'ignoredChars') }
914
+ const value = {
915
+ ignoredChars: readConfigValue(config, 'ignoredChars'),
916
+ ignoredSubstrings: readConfigValue(config, 'ignoredSubstrings'),
917
+ }
752
918
  for (const key of Object.keys(LIMITS)) value[key] = readConfigValue(config, key)
753
919
  for (const key of ['stripWhitespace', 'skipCodeBlocks', 'monitorReasoning', 'monitorToolArguments']) {
754
920
  value[key] = readConfigValue(config, key)
@@ -756,6 +922,18 @@ function settingsValueFromConfig(config) {
756
922
  return value
757
923
  }
758
924
 
925
+ /**
926
+ * 生效参数摘要(宿主日志):确认宿主加载的版本与「设置是否真的进来了」。
927
+ * 片段白名单是否被读到,从这里一眼可见。
928
+ */
929
+ function summarizeRuntime(runtime) { return '阈值 ' + String(runtime.threshold) +
930
+ '|窗口 ' + String(runtime.detectionWindow) +
931
+ '|最大单元 ' + String(runtime.maxUnitLength) +
932
+ '|字符白名单 ' + String(runtime.ignoredChars.length) + ' 项' +
933
+ '|片段白名单 ' + String(runtime.ignoredSubstrings.length) + ' 项' +
934
+ (runtime.ignoredSubstrings.length > 0 ? '(' + runtime.ignoredSubstrings.join('、') + ')' : '')
935
+ }
936
+
759
937
  /**
760
938
  * 从 DSH 的 settings 服务注册检测参数(仅常驻版);设置页可动态调整全部字段。
761
939
  * 返回注销器;settings 服务缺失(无头环境)时不做任何事。
@@ -780,6 +958,12 @@ async function installSettings(ctx, runtime) {
780
958
  })
781
959
  const sync = () => {
782
960
  applySettingsToRuntime(runtime, scope.get())
961
+ // 生效参数清单(同内容只打印一次):便于在宿主日志确认设置是否真的进来了。
962
+ const summary = summarizeRuntime(runtime)
963
+ if (summary !== lastRuntimeSummary) {
964
+ lastRuntimeSummary = summary
965
+ console.log('[dupguard] 生效参数:' + summary)
966
+ }
783
967
  }
784
968
  sync()
785
969
  const stop = scope.watch(sync)
@@ -828,9 +1012,11 @@ function apply(ctx, config) {
828
1012
  applySettingsToRuntime(next, settingsValueFromConfig(config))
829
1013
  return next
830
1014
  }
831
- fromConfig() // 启动时先算一次:日志中暴露非法配置
1015
+ const startup = fromConfig() // 启动时先算一次:日志中暴露非法配置
832
1016
  runtimeSource = fromConfig
833
1017
  console.log('[dupguard] 设置参数取自插件 Config(DSH ≥ 0.1.7 模型;命名空间为 loader entry id)')
1018
+ // 生效参数清单:用于确认宿主加载的版本与「设置是否真的进来了」(片段白名单是否生效靠它判断)。
1019
+ console.log('[dupguard] 生效参数:' + summarizeRuntime(startup))
834
1020
  })
835
1021
  // llm/stream:包裹每次流式模型调用的瀑布事件。
836
1022
  // 监听器返回包装后的 AsyncIterable,即成为本次调用对消费方可见的流。
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "dsh-dupguard",
3
- "version": "1.6.2",
3
+ "version": "1.7.0",
4
4
  "description": "Real-time repetition guard for DeepSeek Harness (DSH): stops model generation when the same string repeats >=10 times in the streamed output. 实时检测 DSH 大模型流式输出中的重复内容,同一字符串重复十次以上立即停止生成。",
5
5
  "keywords": [
6
6
  "dsh-plugin",
package/plugin/host.js CHANGED
@@ -17,6 +17,15 @@
17
17
  // 均可在设置中动态调整并持久化;两版默认值保持一致。
18
18
  // ============================================================================
19
19
 
20
+ /**
21
+ * 片段白名单(ignoredSubstrings)的上限。
22
+ *
23
+ * 两者同时界定「跨增量保留的尾巴长度(≤ 最长片段 - 1 码点)」与单增量匹配成本,
24
+ * 因此是代码常量而非可调设置。
25
+ */
26
+ const IGNORED_SUBSTRINGS_MAX_COUNT = 64
27
+ const IGNORED_SUBSTRING_MAX_LENGTH = 64
28
+
20
29
  const CONFIG = {
21
30
  // 触发阈值:同一字符串连续重复次数达到该值时停止输出(用户需求:重复十次以上)。
22
31
  // 语义为「>= threshold」,即第 10 次重复出现时就触发。
@@ -35,6 +44,13 @@ const CONFIG = {
35
44
  // (如 "|---|---|"),正常表格输出会大量连续出现,不应视为复读。
36
45
  // 默认忽略连字符与竖线;需要更严格的检测时可改为空数组 []。
37
46
  ignoredChars: ['-', '|'],
47
+ // 片段白名单:整段匹配的多字符串(字面量、区分大小写、不支持正则)。
48
+ // 命中时先整段剔除,再做去空白与逐字符剔除 —— 用于 `|---|`、`------` 这类
49
+ // 由多个字符组成的固定片段:逐字符白名单只能忽略单个字符,组合片段的重复
50
+ // 仍会被计入。长片段优先匹配,避免 `---` 抢先破坏 `-----`。
51
+ // 上限:每项 ≤ IGNORED_SUBSTRING_MAX_LENGTH 码点、数组 ≤ IGNORED_SUBSTRINGS_MAX_COUNT 项。
52
+ // 代价:为跨增量匹配,每块最多保留(最长片段 - 1)个码点不参与检测(tail 延迟)。
53
+ ignoredSubstrings: [],
38
54
  // 围栏代码块(``` / ~~~)内的重复检测按倍数分三档:
39
55
  // ≥2 → 放宽:块内改用 threshold × codeBlockMultiplier 判定(默认 3),
40
56
  // 既能放过正常代码,又能兜住真正的失控复读;
@@ -108,13 +124,130 @@ function stripIgnoredChars(text, ignored) {
108
124
  return out
109
125
  }
110
126
 
111
- /** 供检测使用的增量清洗:去空白 + 移除白名单字符。 */
112
- function sanitizeDelta(delta, config) {
113
- let piece = config.stripWhitespace ? stripWhitespace(delta) : delta
127
+ /** 供检测使用的增量清洗:去空白 + 移除白名单字符(片段剔除在此之前完成)。 */
128
+ function sanitizePiece(text, config) {
129
+ let piece = config.stripWhitespace ? stripWhitespace(text) : text
114
130
  if (config.ignoredChars.length > 0) piece = stripIgnoredChars(piece, config.ignoredChars)
115
131
  return piece
116
132
  }
117
133
 
134
+ /**
135
+ * 增量片段剔除器(每个文本块一份)。
136
+ *
137
+ * 语义:把配置里的片段当**字面量子串**整段剔除,长片段优先(避免 `---` 抢先破坏 `-----`),
138
+ * 剔除后再交给 sanitizePiece(去空白 → 逐字符白名单)。
139
+ *
140
+ * 跨增量:为避免片段被增量边界切断,末尾最多保留(最长片段 - 1)个**码点**不输出,
141
+ * 等下一次增量拼回后再匹配;flush() 吐出尾巴(块结束/流结束时调用),保证尾部文本仍参与检测。
142
+ * 代价:检测最多延迟(最长片段 - 1)个码点。
143
+ *
144
+ * 片段表为空时走零开销快路径(不缓冲、原样返回),因此默认行为与未启用该功能时完全一致。
145
+ *
146
+ * 与 lib/index.js 中的同名实现保持一致(两个入口行为必须相同)。
147
+ *
148
+ * @param getPatterns - 读取当前片段表(支持热更新)。
149
+ * @param maxLength - 片段长度上限(决定保留的尾巴长度)。
150
+ */
151
+ function createSubstringStripper(getPatterns, maxLength) {
152
+ let buffer = ''
153
+ let cachedSource = null
154
+ let cachedSorted = []
155
+ const keepLength = () => Math.max(1, maxLength) - 1
156
+
157
+ /** 长片段优先的片段表(按数组引用缓存,热更新后自动重建)。 */
158
+ const sortedPatterns = () => {
159
+ const list = getPatterns()
160
+ if (list !== cachedSource || list.length !== cachedSorted.length) {
161
+ cachedSource = list
162
+ cachedSorted = [...list].sort((left, right) => [...right].length - [...left].length)
163
+ }
164
+ return cachedSorted
165
+ }
166
+
167
+ /** 从 text 中移除所有片段出现(字面量,非正则)。 */
168
+ const removePatterns = (text, patterns) => {
169
+ let out = text
170
+ for (const pattern of patterns) {
171
+ if (pattern.length === 0) continue
172
+ if (out.indexOf(pattern) !== -1) out = out.split(pattern).join('')
173
+ }
174
+ return out
175
+ }
176
+
177
+ return {
178
+ /** 消费一段增量,返回可交给后续清洗的文本(可能为空串)。 */
179
+ push(delta) {
180
+ const patterns = sortedPatterns()
181
+ if (patterns.length === 0) {
182
+ // 快路径:未配置片段时不缓冲;若此前残留尾巴(刚被清空配置),先吐出。
183
+ if (buffer.length === 0) return delta
184
+ const carried = buffer
185
+ buffer = ''
186
+ return carried + delta
187
+ }
188
+ buffer += delta
189
+ const cleaned = removePatterns(buffer, patterns)
190
+ const codePoints = [...cleaned]
191
+ // 保留长度按**实际最长片段**计算(而非配置上限),否则检测会被无谓地拖后。
192
+ const longest = [...patterns[0]].length
193
+ const keep = Math.min(Math.max(0, longest - 1), keepLength(), codePoints.length)
194
+ if (keep === 0) {
195
+ buffer = ''
196
+ return cleaned
197
+ }
198
+ buffer = codePoints.slice(codePoints.length - keep).join('')
199
+ return codePoints.slice(0, codePoints.length - keep).join('')
200
+ },
201
+ /** 吐出保留的尾巴(块结束/流结束)。 */
202
+ flush() {
203
+ const out = buffer
204
+ buffer = ''
205
+ return out
206
+ },
207
+ }
208
+ }
209
+
210
+ /**
211
+ * 片段白名单归一化:丢弃空值/非字符串、超长(> 64 码点)与超量(> 64 项)条目,
212
+ * 并去重。丢弃项只告警一次(同样内容),避免每次模型调用刷屏。
213
+ */
214
+ function normalizeIgnoredSubstrings(raw, fallback) {
215
+ if (!Array.isArray(raw)) return [...fallback]
216
+ const kept = []
217
+ const dropped = []
218
+ for (const entry of raw) {
219
+ if (typeof entry !== 'string' || entry.length === 0) {
220
+ dropped.push(String(entry))
221
+ continue
222
+ }
223
+ if ([...entry].length > IGNORED_SUBSTRING_MAX_LENGTH) {
224
+ dropped.push(entry)
225
+ continue
226
+ }
227
+ if (kept.indexOf(entry) !== -1) continue
228
+ if (kept.length >= IGNORED_SUBSTRINGS_MAX_COUNT) {
229
+ dropped.push(entry)
230
+ continue
231
+ }
232
+ kept.push(entry)
233
+ }
234
+ if (dropped.length > 0) {
235
+ const signature = JSON.stringify(dropped)
236
+ if (signature !== lastSubstringWarning) {
237
+ lastSubstringWarning = signature
238
+ console.warn(
239
+ '[dupguard] 片段白名单条目无效,已忽略:' + JSON.stringify(dropped) +
240
+ '(要求非空字符串、每项 ≤ ' + String(IGNORED_SUBSTRING_MAX_LENGTH) +
241
+ ' 码点、最多 ' + String(IGNORED_SUBSTRINGS_MAX_COUNT) + ' 项)。'
242
+ )
243
+ }
244
+ }
245
+ return kept
246
+ }
247
+
248
+ /** 上一次「无效片段白名单条目」告警的签名,用于去重。 */
249
+ let lastSubstringWarning = ''
250
+
118
251
  /**
119
252
  * 流式围栏代码块过滤器(每个文本块一份状态)。
120
253
  *
@@ -286,6 +419,7 @@ function createStreamGuard(options) {
286
419
  stripped: '', // 去空白后的滚动窗口:仅用于检测
287
420
  fence: createFenceFilter(), // 围栏代码块状态(skipCodeBlocks 开启时使用)
288
421
  lastRunCode: undefined, // 上一片段的代码块归属(用于跨围栏边界清空缓冲)
422
+ substrings: createSubstringStripper(() => CONFIG.ignoredSubstrings, IGNORED_SUBSTRING_MAX_LENGTH),
289
423
  toolCallId: undefined,
290
424
  toolCallName: undefined,
291
425
  toolCallArguments: '',
@@ -314,7 +448,7 @@ function createStreamGuard(options) {
314
448
  if (skipInsideCode && b.lastRunCode !== undefined && run.code !== b.lastRunCode) b.stripped = ''
315
449
  b.lastRunCode = run.code
316
450
  if (skipInsideCode && run.code) continue
317
- const piece = sanitizeDelta(run.text, CONFIG)
451
+ const piece = sanitizePiece(b.substrings.push(run.text), CONFIG)
318
452
  if (piece.length === 0) continue
319
453
  b.stripped = (b.stripped + piece).slice(-CONFIG.detectionWindow)
320
454
  const threshold = run.code ? codeThreshold : CONFIG.threshold
@@ -324,6 +458,21 @@ function createStreamGuard(options) {
324
458
  return null
325
459
  }
326
460
 
461
+ /**
462
+ * 块结束时吐出片段剔除器保留的尾巴(≤ 最长片段 - 1 码点),
463
+ * 让它仍参与检测;阈值沿用该块最后一次片段的代码块归属。
464
+ */
465
+ function flushSubstringTail(b) {
466
+ if (b === undefined || b.substrings === undefined) return null
467
+ const tail = sanitizePiece(b.substrings.flush(), CONFIG)
468
+ if (tail.length === 0) return null
469
+ b.stripped = (b.stripped + tail).slice(-CONFIG.detectionWindow)
470
+ const threshold = b.lastRunCode === true
471
+ ? CONFIG.threshold * (CONFIG.skipCodeBlocks === true ? CONFIG.codeBlockMultiplier : 1)
472
+ : CONFIG.threshold
473
+ return findRepeatedTail(b.stripped, threshold, CONFIG.minUnitLength, CONFIG.maxUnitLength)
474
+ }
475
+
327
476
  /** 依据 StreamChunk 协议累积状态;命中时置 stopped。 */
328
477
  function feed(chunk) {
329
478
  switch (chunk.type) {
@@ -354,7 +503,7 @@ function createStreamGuard(options) {
354
503
  if (chunk.name !== undefined) b.toolCallName = chunk.name
355
504
  b.toolCallArguments += chunk.argumentsDelta
356
505
  if (CONFIG.monitorToolArguments) {
357
- const piece = sanitizeDelta(chunk.argumentsDelta, CONFIG)
506
+ const piece = sanitizePiece(b.substrings.push(chunk.argumentsDelta), CONFIG)
358
507
  b.stripped = (b.stripped + piece).slice(-CONFIG.detectionWindow)
359
508
  const hit = findRepeatedTail(b.stripped, CONFIG.threshold, CONFIG.minUnitLength, CONFIG.maxUnitLength)
360
509
  if (hit !== null) stopped = hit
@@ -362,11 +511,21 @@ function createStreamGuard(options) {
362
511
  return
363
512
  }
364
513
  case 'block-end': {
514
+ const b = blocks.get(chunk.index)
515
+ const tailHit = flushSubstringTail(b)
516
+ if (tailHit !== null) stopped = tailHit
365
517
  blocks.delete(chunk.index)
366
518
  return
367
519
  }
368
520
  case 'usage':
369
- case 'finish':
521
+ case 'finish': {
522
+ // 流结束时可能还有未收到 block-end 的块:把尾巴补上,避免尾部文本漏检。
523
+ for (const b of blocks.values()) {
524
+ const tailHit = flushSubstringTail(b)
525
+ if (tailHit !== null) stopped = tailHit
526
+ }
527
+ return
528
+ }
370
529
  default:
371
530
  return
372
531
  }