dsh-dupguard 1.6.2 → 1.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +38 -0
- package/README.md +13 -5
- package/lib/client.js +201 -47
- package/lib/index.js +194 -8
- package/package.json +1 -1
- package/plugin/host.js +165 -6
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,44 @@
|
|
|
2
2
|
|
|
3
3
|
本文件遵循 [Keep a Changelog](https://keepachangelog.com/zh-CN/1.1.0/),版本号遵循 [SemVer](https://semver.org/lang/zh-CN/)。
|
|
4
4
|
|
|
5
|
+
## [1.7.0] - 2026-09-24
|
|
6
|
+
|
|
7
|
+
### Added
|
|
8
|
+
|
|
9
|
+
- **多字符片段白名单 `ignoredSubstrings`**(默认 `[]`,显式开启):
|
|
10
|
+
- 整段**字面量**匹配(区分大小写、不支持正则),命中时先整段剔除,再做去空白与逐字符剔除;
|
|
11
|
+
用于 `|---|`、`<br>` 这类由多个字符组成的固定片段(逐字符白名单只能忽略单个字符);
|
|
12
|
+
- **长片段优先**匹配,避免 `---` 抢先破坏 `-----`;
|
|
13
|
+
- **跨增量安全**:每块最多保留(实际最长片段 − 1)个码点,等下一增量拼回后再匹配;
|
|
14
|
+
块结束/流结束时补投尾巴,尾部文本仍参与检测;
|
|
15
|
+
- 上限:每项 ≤ 64 码点、最多 64 项;空值/超长/超量条目运行时丢弃并各告警一次(同样内容只告警一次);
|
|
16
|
+
- 片段表为空时走零开销快路径,默认行为与未提供该功能时**逐字节一致**;
|
|
17
|
+
- 设置页新增「忽略片段」列表(整段加入、不按码点拆分,含空值/超长/重复/超量校验),
|
|
18
|
+
与字符白名单共用写通道,可持久化并可被「恢复默认」清空。
|
|
19
|
+
|
|
20
|
+
### Notes
|
|
21
|
+
|
|
22
|
+
- 实测开销(10 万增量近失配流):空表 2.6 µs/增量 → 配置 4 条不匹配片段 3.3 µs/增量(约 1.3×);
|
|
23
|
+
- 双入口(`lib/index.js` / `plugin/host.js`)实现保持一致,共享用例分别驱动两边;
|
|
24
|
+
- 宿主启动日志新增一行**生效参数清单**(阈值 / 窗口 / 最大单元 / 字符与片段白名单项数),
|
|
25
|
+
设置页底部常驻「设置通道 + 构建标记」诊断行——用于区分「宿主未加载新代码」「设置没下发」「实现有问题」三类故障。
|
|
26
|
+
|
|
27
|
+
## [1.6.3] - 2026-09-24
|
|
28
|
+
|
|
29
|
+
### Fixed
|
|
30
|
+
|
|
31
|
+
- **写入成功却被判为「保存失败」**:`remote.settings.mutate()` 返回的命名空间视图是**写入前**的快照
|
|
32
|
+
(配置由 loader 异步重载后才更新),此前的实现拿它与本地期望值做回读比对,于是每次写入都
|
|
33
|
+
必然不一致——面板报错、看起来「没保存」,而宿主其实已经落盘。现改为**只以宿主应答为准**
|
|
34
|
+
(`response.ok === false` 才算失败),与官方 `ConfigForm.mutate` 的语义一致;恢复默认同样按
|
|
35
|
+
「全部 unset 是否未被拒绝」判定。
|
|
36
|
+
|
|
37
|
+
### Added
|
|
38
|
+
|
|
39
|
+
- 设置页底部常驻诊断行 `设置通道:<状态>|<诊断>|构建 <标记>`,并在写入失败时附带宿主原始原因
|
|
40
|
+
与构建标记;浏览器控制台同步打印 `[dupguard] mutate → …` / `mutate 被拒…`,用于快速区分
|
|
41
|
+
「宿主拒绝」「通道未就绪」「回读不一致」三类问题。
|
|
42
|
+
|
|
5
43
|
## [1.6.2] - 2026-09-24
|
|
6
44
|
|
|
7
45
|
### Fixed
|
package/README.md
CHANGED
|
@@ -49,6 +49,8 @@ When triggered, the already-generated text is committed as a normal assistant me
|
|
|
49
49
|
- **真正的服务端停止**:提前关闭流迭代 → 适配器 `consumer.abort()` → 中断 HTTP 连接,模型在服务端停止生成。
|
|
50
50
|
- **安全停止**:绝不 `abort()` agent 步骤信号;补发协议合规的 `block-end` + `finish(stop)`,消息正常提交。
|
|
51
51
|
- **Markdown 表格友好**:默认忽略连字符与竖线(`ignoredChars` 白名单),表格分隔行与长分隔线不会被误判为复读。
|
|
52
|
+
- **片段白名单**:`ignoredSubstrings` 可按**整段**忽略多字符片段(如 `|---|`、`<br>`)——逐字符白名单只能
|
|
53
|
+
忽略单个字符,组合片段的重复仍会被计入;命中时先整段剔除再走常规清洗,且跨增量切分也能正确剔除。
|
|
52
54
|
- **图形化设置页**(npm 常驻版):在 DSH 设置面板注册与「通用设置 / 模型 / 插件 / Agent 预设」并列的
|
|
53
55
|
「重复守卫」分节,可视化编辑白名单与全部检测参数(阈值、最小/最大单元长度、检测窗口、空白与
|
|
54
56
|
reasoning、工具参数开关)并持久化(`dsh-dupguard` 设置命名空间),修改即时生效;窗口小于
|
|
@@ -198,6 +200,7 @@ and persisted to `settings.yaml`; the dynamic build uses the constants.
|
|
|
198
200
|
| `codeBlockMultiplier` | `3` | 代码块内处理分三档:**≥2** 按「阈值 × 本倍数」判定(越大越不易误杀正常代码,但也越晚兜住块内失控复读);**1** 与块外同样严格;**0** 完全不检测代码块内。范围 0–100 / how code blocks are handled: >=2 = threshold x this multiplier, 1 = same as outside, 0 = do not detect inside code blocks at all. Range 0–100 |
|
|
199
201
|
| `stripWhitespace` | `true` | 检测前移除空白/换行,识别带分隔符的复读 / strip whitespace so `"x x x"` and `"x\nx\nx"` are caught |
|
|
200
202
|
| `ignoredChars` | `['-', '\|']` | 检测时忽略的字符白名单:Markdown 表格分隔行(连字符与竖线)不参与重复统计。条目必须**是单个字符**(按 Unicode 码点匹配,emoji 也算一个);设置页一次输入多个字符会逐个加入,多字符/空条目在运行时被丢弃并告警 / whitelist of characters ignored during detection, so Markdown table separators don't count. Entries must be a **single character** (matched per Unicode code point); the settings page splits multi-character input into individual entries, and invalid entries are dropped at runtime with a warning |
|
|
203
|
+
| `ignoredSubstrings` | `[]` | **片段白名单(多字符)**:整段字面量匹配(区分大小写、不支持正则),命中时**先整段剔除**,再做去空白与逐字符剔除;长片段优先。用于 `\|---\|`、`<br>` 这类由多个字符组成的固定片段。每项 ≤ 64 码点、最多 64 项,超限项运行时丢弃并告警。代价:为跨增量匹配,每块最多保留(最长片段 − 1)个字符不参与检测 / whitelist of **multi-character substrings**, matched literally (case-sensitive, no regex); matches are stripped first, then whitespace and per-character rules apply, longest first. Entries ≤ 64 code points, at most 64 entries (invalid entries dropped with a warning). Cost: up to (longest entry − 1) trailing characters per block are held back for cross-delta matching |
|
|
201
204
|
| `monitorReasoning` | `true` | 是否检测思考文本(思考中的复读同样消耗 token,默认截停;只检测可见输出时置 `false`)/ also guard reasoning (thinking) text — on by default; set `false` to guard visible output only |
|
|
202
205
|
| `monitorToolArguments` | `false` | 是否检测工具调用参数 / also guard tool-call JSON args — off by default (base64/JSON repeats are common) |
|
|
203
206
|
| `fixStandingMountConflict` | `true` | DSH ≤ 0.1.6-alpha.2 兼容补丁:幂等化 `cordisInspect.register`,修复 preset standing-mount 多代并存冲突(仅代码常量)/ idempotent `cordisInspect.register` patch for the DSH ≤ 0.1.6-alpha.2 standing-mount conflict (code constant only) |
|
|
@@ -225,6 +228,8 @@ Listens to the `llm/stream` waterfall (wraps every streaming model call) and ret
|
|
|
225
228
|
### 2. 检测算法 / Detection
|
|
226
229
|
|
|
227
230
|
- 按块索引(`chunk.index`)分别累积文本,多块交替输出互不干扰;
|
|
231
|
+
- 清洗顺序:**先按片段白名单整段剔除**(`ignoredSubstrings`,长片段优先,跨增量尾巴由块结束时的补投兜住)
|
|
232
|
+
→ 再去空白(可关)→ 最后按单字符白名单剔除(`ignoredChars`)。片段逐次替换为空串,非正则匹配;
|
|
228
233
|
- 去空白后做**尾部连续重复检测**:文本以某个单元(长度 `minUnitLength`..`maxUnitLength`,默认 1..80)
|
|
229
234
|
连续重复 ≥ `threshold` 次结尾即触发。模型一旦复读,重复必然在尾部,因此尾部检测即可实时捕获所有
|
|
230
235
|
循环,同时避免全窗口词频的误报(如正常中文里高频的"的")。
|
|
@@ -280,9 +285,9 @@ default), Markdown table separator rows and horizontal rules (whitelisted by def
|
|
|
280
285
|
│ ├── index.js # npm/组合常驻形式(package.json main 入口,含设置集成)
|
|
281
286
|
│ └── client.js # 浏览器端设置页(ModuleLoader 格式,dsh.client 入口)
|
|
282
287
|
├── tests/
|
|
283
|
-
│ ├── detector.test.js # 端到端测试:双入口防漂移 + reasoning 开关 + settings/Config 集成(
|
|
284
|
-
│ ├── client.test.js # 设置页组件测试:最小 React/DSH 桩(旧版 settingsScope + 新版 remote,
|
|
285
|
-
│ ├── stress-host-adversarial.js #
|
|
288
|
+
│ ├── detector.test.js # 端到端测试:双入口防漂移 + reasoning 开关 + settings/Config 集成(82 项)
|
|
289
|
+
│ ├── client.test.js # 设置页组件测试:最小 React/DSH 桩(旧版 settingsScope + 新版 remote,29 项)
|
|
290
|
+
│ ├── stress-host-adversarial.js # 压力:边界/协议交错/围栏与片段白名单/热更新 churn/畸形输入
|
|
286
291
|
│ ├── stress-host-throughput.js # 压力:吞吐/内存/200 路并发/参数极值(METRIC 指标)
|
|
287
292
|
│ ├── stress-client-ui.js # 压力:设置页高频交互、乱序应答、挂载泄漏
|
|
288
293
|
│ ├── stress-real-invariant.mjs # 压力:真实 DSH llm-invariant + BlockAssembler 端到端校验
|
|
@@ -300,8 +305,8 @@ default), Markdown table separator rows and horizontal rules (whitelisted by def
|
|
|
300
305
|
|
|
301
306
|
```bash
|
|
302
307
|
npm test # 功能测试(两个文件)
|
|
303
|
-
node tests/detector.test.js # 检测端到端(
|
|
304
|
-
node tests/client.test.js # 设置页组件(
|
|
308
|
+
node tests/detector.test.js # 检测端到端(82 项)
|
|
309
|
+
node tests/client.test.js # 设置页组件(29 项)
|
|
305
310
|
```
|
|
306
311
|
|
|
307
312
|
同一套用例分别驱动两个入口(`plugin/host.js` 经 `new Function` 求值、`lib/index.js` 经
|
|
@@ -353,6 +358,9 @@ when DSH is not installed. CI runs on Node 20/22/24 (matching DSH; Node 18 is no
|
|
|
353
358
|
- **代码块分档只覆盖围栏代码块**:行内代码(`` `x` ``)与缩进代码块(4 空格)仍按普通阈值判定;
|
|
354
359
|
模型忘记闭合围栏时,其后内容一律按代码块处理。倍数为 `0`(完全不检测)时,代码块内的失控复读
|
|
355
360
|
不会被截停——这是「代码再长也不误杀」的代价;默认倍数 `3` 则会在「阈值 × 3」处兜底。
|
|
361
|
+
- **片段白名单的代价**:为跨增量匹配,启用后每块最多保留(最长片段 − 1)个字符不参与检测(块结束时补投),
|
|
362
|
+
即检测最多延迟这么多字符;条目上限 64 项 × 64 字符,逐条字面量替换、不支持正则;片段表为空时走
|
|
363
|
+
零开销快路径(默认不启用,行为与未提供该功能时一致)。
|
|
356
364
|
- 检测窗口上限 1,048,576 字符:每个增量都要重写一次缓冲,成本随窗口线性增长——缓冲填满后
|
|
357
365
|
1 MiB 窗口约 0.13 ms/增量,实测 4 字符增量的平均值为 33 µs/增量(含缓冲填充期)。默认 8192 无感
|
|
358
366
|
(1.9 µs/增量,模型侧毫秒级的 token 间隔下可忽略)。
|
package/lib/client.js
CHANGED
|
@@ -26,6 +26,8 @@ window.__ModuleLoader__.load({
|
|
|
26
26
|
const React = require('react')
|
|
27
27
|
|
|
28
28
|
const NS = 'dsh-dupguard'
|
|
29
|
+
/** 构建标记:显示在设置页底部,用于判断浏览器实际加载的是哪一版(排查缓存/旧包问题)。 */
|
|
30
|
+
const BUILD_MARK = '1.7.0'
|
|
29
31
|
/** 旧版(≤ 0.1.6)register 出来的命名空间名。 */
|
|
30
32
|
const SETTINGS_NS = 'dsh-dupguard'
|
|
31
33
|
/** 新版(≥ 0.1.7)命名空间取自 loader entry id:本 bundle 插入的行 id 为 dupguard。 */
|
|
@@ -41,6 +43,15 @@ window.__ModuleLoader__.load({
|
|
|
41
43
|
empty: '白名单为空:所有字符都参与重复统计。',
|
|
42
44
|
addPlaceholder: '输入要忽略的字符(可多个)',
|
|
43
45
|
add: '添加',
|
|
46
|
+
substrings: '忽略片段(多字符,按整段匹配)',
|
|
47
|
+
substringsEmpty: '片段白名单为空:没有整段被忽略的字符串。',
|
|
48
|
+
substringAddPlaceholder: '输入要整段忽略的字符串(如 |---|)',
|
|
49
|
+
substringAdd: '添加片段',
|
|
50
|
+
substringHint: '整段字面量匹配(区分大小写、不支持正则):命中时先整段剔除,再做去空白与逐字符剔除。长片段优先。每项 ≤ {maxLen} 字符、最多 {maxCount} 项。',
|
|
51
|
+
errSubstringEmpty: '请输入非空片段。',
|
|
52
|
+
errSubstringTooLong: '片段过长:每项最多 {maxLen} 个字符。',
|
|
53
|
+
errSubstringDuplicate: '该片段已在列表中。',
|
|
54
|
+
errSubstringLimit: '片段数量已达上限({maxCount} 项)。',
|
|
44
55
|
params: '检测参数',
|
|
45
56
|
threshold: '触发阈值(连续重复次数)',
|
|
46
57
|
thresholdHint: '同一字符串连续重复达到该次数即截停(≥ 语义)。范围 {min}–{max}。',
|
|
@@ -83,6 +94,15 @@ window.__ModuleLoader__.load({
|
|
|
83
94
|
empty: 'Whitelist is empty: every character counts.',
|
|
84
95
|
addPlaceholder: 'Characters to ignore (one or more)',
|
|
85
96
|
add: 'Add',
|
|
97
|
+
substrings: 'Ignored substrings (multi-character, matched whole)',
|
|
98
|
+
substringsEmpty: 'No ignored substrings: nothing is stripped as a whole.',
|
|
99
|
+
substringAddPlaceholder: 'String to ignore as a whole (e.g. |---|)',
|
|
100
|
+
substringAdd: 'Add substring',
|
|
101
|
+
substringHint: 'Literal match of the whole substring (case-sensitive, no regex): matches are stripped first, then whitespace and per-character rules apply. Longer substrings win. Each entry ≤ {maxLen} characters, at most {maxCount} entries.',
|
|
102
|
+
errSubstringEmpty: 'Enter a non-empty substring.',
|
|
103
|
+
errSubstringTooLong: 'Substring too long: at most {maxLen} characters each.',
|
|
104
|
+
errSubstringDuplicate: 'That substring is already in the list.',
|
|
105
|
+
errSubstringLimit: 'Substring limit reached ({maxCount} entries).',
|
|
86
106
|
params: 'Detection parameters',
|
|
87
107
|
threshold: 'Threshold (consecutive repeats)',
|
|
88
108
|
thresholdHint: 'Stop once a string repeats this many times in a row (>= semantics). Range {min}–{max}.',
|
|
@@ -175,6 +195,7 @@ window.__ModuleLoader__.load({
|
|
|
175
195
|
const BOOL_FIELDS = ['stripWhitespace', 'skipCodeBlocks', 'monitorReasoning', 'monitorToolArguments']
|
|
176
196
|
const DEFAULTS = {
|
|
177
197
|
ignoredChars: ['-', '|'],
|
|
198
|
+
ignoredSubstrings: [],
|
|
178
199
|
threshold: 10,
|
|
179
200
|
codeBlockMultiplier: 3,
|
|
180
201
|
minUnitLength: 1,
|
|
@@ -185,7 +206,11 @@ window.__ModuleLoader__.load({
|
|
|
185
206
|
monitorReasoning: true,
|
|
186
207
|
monitorToolArguments: false,
|
|
187
208
|
}
|
|
188
|
-
const ALL_FIELDS = ['ignoredChars'].concat(NUMERIC_FIELDS.map((field) => field.key), BOOL_FIELDS)
|
|
209
|
+
const ALL_FIELDS = ['ignoredChars', 'ignoredSubstrings'].concat(NUMERIC_FIELDS.map((field) => field.key), BOOL_FIELDS)
|
|
210
|
+
|
|
211
|
+
/** 与宿主 lib/index.js 的常量保持一致:片段上限(每项码点数 / 条目数)。 */
|
|
212
|
+
const MAX_SUBSTRING_LENGTH = 64
|
|
213
|
+
const MAX_SUBSTRINGS = 64
|
|
189
214
|
|
|
190
215
|
/** 模板占位符替换:{name} → vars.name。 */
|
|
191
216
|
const fmt = (text, vars) =>
|
|
@@ -196,7 +221,10 @@ window.__ModuleLoader__.load({
|
|
|
196
221
|
* 保证任何异常形状到达视图时仍是 ready 且可渲染。
|
|
197
222
|
*/
|
|
198
223
|
function normalizeValue(raw) {
|
|
199
|
-
const out = {
|
|
224
|
+
const out = {
|
|
225
|
+
ignoredChars: [...DEFAULTS.ignoredChars],
|
|
226
|
+
ignoredSubstrings: [...DEFAULTS.ignoredSubstrings],
|
|
227
|
+
}
|
|
200
228
|
for (const field of NUMERIC_FIELDS) out[field.key] = DEFAULTS[field.key]
|
|
201
229
|
for (const key of BOOL_FIELDS) out[key] = DEFAULTS[key]
|
|
202
230
|
try {
|
|
@@ -204,6 +232,11 @@ window.__ModuleLoader__.load({
|
|
|
204
232
|
if (Array.isArray(raw.ignoredChars)) {
|
|
205
233
|
out.ignoredChars = raw.ignoredChars.filter((ch) => typeof ch === 'string')
|
|
206
234
|
}
|
|
235
|
+
if (Array.isArray(raw.ignoredSubstrings)) {
|
|
236
|
+
out.ignoredSubstrings = raw.ignoredSubstrings
|
|
237
|
+
.filter((item) => typeof item === 'string' && item.length > 0)
|
|
238
|
+
.slice(0, MAX_SUBSTRINGS)
|
|
239
|
+
}
|
|
207
240
|
for (const field of NUMERIC_FIELDS) {
|
|
208
241
|
const value = raw[field.key]
|
|
209
242
|
if (Number.isSafeInteger(value) && value >= field.min && value <= field.max) out[field.key] = value
|
|
@@ -219,7 +252,10 @@ window.__ModuleLoader__.load({
|
|
|
219
252
|
/** 归一化值为编辑表单:数值字段存字符串,便于输入中途的状态。 */
|
|
220
253
|
function toForm(value) {
|
|
221
254
|
const normalized = normalizeValue(value)
|
|
222
|
-
const form = {
|
|
255
|
+
const form = {
|
|
256
|
+
ignoredChars: normalized.ignoredChars,
|
|
257
|
+
ignoredSubstrings: normalized.ignoredSubstrings,
|
|
258
|
+
}
|
|
223
259
|
for (const field of NUMERIC_FIELDS) form[field.key] = String(normalized[field.key])
|
|
224
260
|
for (const key of BOOL_FIELDS) form[key] = normalized[key]
|
|
225
261
|
return form
|
|
@@ -248,6 +284,7 @@ window.__ModuleLoader__.load({
|
|
|
248
284
|
const [form, setForm] = React.useState(null)
|
|
249
285
|
const [writeState, setWriteState] = React.useState(null)
|
|
250
286
|
const [errors, setErrors] = React.useState({})
|
|
287
|
+
const [draftSub, setDraftSub] = React.useState('')
|
|
251
288
|
const dirty = React.useRef({})
|
|
252
289
|
const lastRemote = React.useRef(null)
|
|
253
290
|
|
|
@@ -344,28 +381,35 @@ window.__ModuleLoader__.load({
|
|
|
344
381
|
}
|
|
345
382
|
: null
|
|
346
383
|
|
|
347
|
-
// ----
|
|
348
|
-
//
|
|
349
|
-
//
|
|
384
|
+
// ---- 写路径:设置控制器(新版 remote.settings,旧版 settingsScope)。 ----
|
|
385
|
+
// 写成功与否**只以宿主应答为准**:宿主返回的视图是写入前的快照(配置由 loader
|
|
386
|
+
// 异步重载),拿它跟本地期望值比对必然不一致——这正是此前「明明保存成功却报
|
|
387
|
+
// 保存失败、重开后新值不见了」的原因。官方实现同样是信任 response.ok。
|
|
350
388
|
const snapshotValue = () => {
|
|
351
389
|
const current = controller.getSnapshot()
|
|
352
390
|
return current !== undefined && current.value !== undefined ? current.value : {}
|
|
353
391
|
}
|
|
354
|
-
const sameList = (a, b) => Array.isArray(a) && Array.isArray(b) &&
|
|
355
|
-
a.length === b.length && a.every((item, index) => item === b[index])
|
|
356
392
|
const currentForm = () => (form === null ? remoteForm : form)
|
|
357
|
-
const runWrite = (operation
|
|
393
|
+
const runWrite = (operation) => {
|
|
358
394
|
setWriteState('saving')
|
|
395
|
+
// 失败时带上宿主返回的原始原因(writeReason)与构建标记:
|
|
396
|
+
// 标记同时用于判断浏览器加载的是哪一版(排查旧包/缓存)。
|
|
397
|
+
const failureText = () => {
|
|
398
|
+
const reason = typeof props.writeReason === 'function' ? props.writeReason() : null
|
|
399
|
+
const suffix = reason === null || reason === undefined ? '(宿主未给出原因)' : String(reason)
|
|
400
|
+
return t('saveFailed') + ' [diag ' + BUILD_MARK + '] ' + suffix
|
|
401
|
+
}
|
|
359
402
|
Promise.resolve()
|
|
360
403
|
.then(() => operation())
|
|
361
|
-
.then(() => setWriteState(
|
|
404
|
+
.then((accepted) => setWriteState(accepted === false ? 'error:' + failureText() : 'saved'), (error) => {
|
|
362
405
|
setWriteState('error:' + String((error && error.message) || error))
|
|
363
406
|
})
|
|
364
407
|
}
|
|
365
|
-
|
|
366
|
-
|
|
367
|
-
|
|
368
|
-
|
|
408
|
+
/** 通用列表字段写入(字符白名单 / 片段白名单共用)。 */
|
|
409
|
+
const commitFieldList = (field, next) => {
|
|
410
|
+
dirty.current[field] = true
|
|
411
|
+
setForm({ ...currentForm(), [field]: next })
|
|
412
|
+
runWrite(() => controller.set(field, next))
|
|
369
413
|
}
|
|
370
414
|
const commitNumber = (field) => {
|
|
371
415
|
const message = fieldError(field)
|
|
@@ -376,7 +420,7 @@ window.__ModuleLoader__.load({
|
|
|
376
420
|
const value = parsed[field.key]
|
|
377
421
|
setErrors((prev) => ({ ...prev, [field.key]: null }))
|
|
378
422
|
// 值未变化时不写入:避免仅仅聚焦/失焦就把默认值写进用户层,
|
|
379
|
-
//
|
|
423
|
+
// 污染配置文件并让「恢复默认」失去意义。
|
|
380
424
|
if (snapshotValue()[field.key] === value) {
|
|
381
425
|
dirty.current[field.key] = false
|
|
382
426
|
setForm({ ...currentForm(), [field.key]: String(value) })
|
|
@@ -384,27 +428,20 @@ window.__ModuleLoader__.load({
|
|
|
384
428
|
}
|
|
385
429
|
dirty.current[field.key] = true
|
|
386
430
|
setForm({ ...currentForm(), [field.key]: String(value) })
|
|
387
|
-
runWrite(() => controller.set(field.key, value)
|
|
431
|
+
runWrite(() => controller.set(field.key, value))
|
|
388
432
|
}
|
|
389
433
|
const commitBool = (key) => {
|
|
390
434
|
const value = shown[key] !== true
|
|
391
435
|
dirty.current[key] = true
|
|
392
436
|
setForm({ ...currentForm(), [key]: value })
|
|
393
|
-
runWrite(() => controller.set(key, value)
|
|
437
|
+
runWrite(() => controller.set(key, value))
|
|
394
438
|
}
|
|
395
439
|
const resetAll = () => {
|
|
396
440
|
dirty.current = {}
|
|
397
441
|
setErrors({})
|
|
398
|
-
|
|
399
|
-
|
|
400
|
-
() =>
|
|
401
|
-
setForm(toForm(snapshotValue()))
|
|
402
|
-
const current = controller.getSnapshot()
|
|
403
|
-
const user = current !== undefined ? current.user : undefined
|
|
404
|
-
if (user === undefined || user === null) return true
|
|
405
|
-
return ALL_FIELDS.every((key) => user[key] === undefined)
|
|
406
|
-
},
|
|
407
|
-
)
|
|
442
|
+
// 恢复默认 = 逐字段 unset;接受与否同样只看宿主应答。
|
|
443
|
+
runWrite(() => Promise.all(ALL_FIELDS.map((key) => controller.unset(key)))
|
|
444
|
+
.then((results) => results.every((item) => item !== false)))
|
|
408
445
|
}
|
|
409
446
|
|
|
410
447
|
const add = () => {
|
|
@@ -421,9 +458,40 @@ window.__ModuleLoader__.load({
|
|
|
421
458
|
next.push(ch)
|
|
422
459
|
changed = true
|
|
423
460
|
}
|
|
424
|
-
if (changed)
|
|
461
|
+
if (changed) commitFieldList('ignoredChars', next)
|
|
425
462
|
}
|
|
426
|
-
const remove = (ch) =>
|
|
463
|
+
const remove = (ch) => commitFieldList('ignoredChars', shown.ignoredChars.filter((item) => item !== ch))
|
|
464
|
+
|
|
465
|
+
/**
|
|
466
|
+
* 片段白名单:与字符白名单不同,**不按码点拆分**——整段就是一个条目。
|
|
467
|
+
* 校验空值/超长/重复/超量,错误就地显示且不写入。
|
|
468
|
+
*/
|
|
469
|
+
const addSubstring = () => {
|
|
470
|
+
const text = draftSub.trim()
|
|
471
|
+
const fail = (message) => setErrors((prev) => ({ ...prev, substrings: message }))
|
|
472
|
+
if (text.length === 0) {
|
|
473
|
+
fail(t('errSubstringEmpty'))
|
|
474
|
+
return
|
|
475
|
+
}
|
|
476
|
+
const length = [...text].length
|
|
477
|
+
if (length > MAX_SUBSTRING_LENGTH) {
|
|
478
|
+
fail(fmt(t('errSubstringTooLong'), { maxLen: MAX_SUBSTRING_LENGTH }))
|
|
479
|
+
return
|
|
480
|
+
}
|
|
481
|
+
if (shown.ignoredSubstrings.indexOf(text) !== -1) {
|
|
482
|
+
fail(t('errSubstringDuplicate'))
|
|
483
|
+
return
|
|
484
|
+
}
|
|
485
|
+
if (shown.ignoredSubstrings.length >= MAX_SUBSTRINGS) {
|
|
486
|
+
fail(fmt(t('errSubstringLimit'), { maxCount: MAX_SUBSTRINGS }))
|
|
487
|
+
return
|
|
488
|
+
}
|
|
489
|
+
setDraftSub('')
|
|
490
|
+
setErrors((prev) => ({ ...prev, substrings: null }))
|
|
491
|
+
commitFieldList('ignoredSubstrings', [...shown.ignoredSubstrings, text])
|
|
492
|
+
}
|
|
493
|
+
const removeSubstring = (item) =>
|
|
494
|
+
commitFieldList('ignoredSubstrings', shown.ignoredSubstrings.filter((entry) => entry !== item))
|
|
427
495
|
|
|
428
496
|
const fieldRow = (field) => {
|
|
429
497
|
const message = errors[field.key] !== undefined && errors[field.key] !== null
|
|
@@ -497,6 +565,36 @@ window.__ModuleLoader__.load({
|
|
|
497
565
|
React.createElement('button', { className: 'dg-btn', type: 'button', onClick: add }, t('add')),
|
|
498
566
|
),
|
|
499
567
|
|
|
568
|
+
React.createElement('p', { className: 'dg-note' }, t('substrings')),
|
|
569
|
+
shown.ignoredSubstrings.length === 0
|
|
570
|
+
? React.createElement('p', { className: 'dg-empty' }, t('substringsEmpty'))
|
|
571
|
+
: React.createElement('div', { className: 'dg-chips' },
|
|
572
|
+
shown.ignoredSubstrings.map((item) => React.createElement('span', { className: 'dg-chip', key: item },
|
|
573
|
+
item,
|
|
574
|
+
React.createElement('button', {
|
|
575
|
+
className: 'dg-chip-remove',
|
|
576
|
+
type: 'button',
|
|
577
|
+
onClick: () => removeSubstring(item),
|
|
578
|
+
'aria-label': 'remove',
|
|
579
|
+
}, '\u00d7')))),
|
|
580
|
+
React.createElement('div', { className: 'dg-add' },
|
|
581
|
+
React.createElement('input', {
|
|
582
|
+
className: 'dg-input',
|
|
583
|
+
placeholder: t('substringAddPlaceholder'),
|
|
584
|
+
value: draftSub,
|
|
585
|
+
onChange: (event) => setDraftSub(event.target.value),
|
|
586
|
+
onKeyDown: (event) => {
|
|
587
|
+
if (event.key === 'Enter') addSubstring()
|
|
588
|
+
},
|
|
589
|
+
}),
|
|
590
|
+
React.createElement('button', { className: 'dg-btn', type: 'button', onClick: addSubstring }, t('substringAdd')),
|
|
591
|
+
),
|
|
592
|
+
React.createElement('p', { className: 'dg-field-hint' },
|
|
593
|
+
fmt(t('substringHint'), { maxLen: MAX_SUBSTRING_LENGTH, maxCount: MAX_SUBSTRINGS })),
|
|
594
|
+
errors.substrings === undefined || errors.substrings === null
|
|
595
|
+
? null
|
|
596
|
+
: React.createElement('p', { className: 'dg-field-error' }, errors.substrings),
|
|
597
|
+
|
|
500
598
|
React.createElement('p', { className: 'dg-note' }, t('params')),
|
|
501
599
|
React.createElement('div', { className: 'dg-fields' },
|
|
502
600
|
NUMERIC_FIELDS.map((field) => fieldRow(field)),
|
|
@@ -513,6 +611,11 @@ window.__ModuleLoader__.load({
|
|
|
513
611
|
: writeState === 'saved'
|
|
514
612
|
? React.createElement('p', { className: 'dg-note' }, t('saved'))
|
|
515
613
|
: React.createElement('p', { className: 'dg-note dg-error' }, writeState),
|
|
614
|
+
React.createElement('p', { className: 'dg-note' },
|
|
615
|
+
fmt(t('loadingDiag'), {
|
|
616
|
+
state: mirrorSnap ? String(mirrorSnap.status) : 'unknown',
|
|
617
|
+
diag: (mirrorSnap && mirrorSnap.diag ? String(mirrorSnap.diag) : '') + '|构建 ' + BUILD_MARK,
|
|
618
|
+
})),
|
|
516
619
|
),
|
|
517
620
|
)
|
|
518
621
|
}
|
|
@@ -657,6 +760,32 @@ window.__ModuleLoader__.load({
|
|
|
657
760
|
if (!closed && mine === generation) loaded = true
|
|
658
761
|
}
|
|
659
762
|
}
|
|
763
|
+
/** 写入失败时面板上展示的宿主原始原因(诊断用)。 */
|
|
764
|
+
let lastWriteError = null
|
|
765
|
+
|
|
766
|
+
/** 只尝试属于本插件的命名空间名:describe 报出的 ns 与其去掉组合前缀的形式。 */
|
|
767
|
+
const namespaceCandidates = () => {
|
|
768
|
+
const list = [resolvedNamespace]
|
|
769
|
+
const cut = resolvedNamespace.lastIndexOf(':')
|
|
770
|
+
if (cut !== -1 && cut + 1 < resolvedNamespace.length) {
|
|
771
|
+
const stripped = resolvedNamespace.slice(cut + 1)
|
|
772
|
+
if (list.indexOf(stripped) === -1) list.push(stripped)
|
|
773
|
+
}
|
|
774
|
+
return list
|
|
775
|
+
}
|
|
776
|
+
|
|
777
|
+
const attempt = async (namespace, owned, expected) => {
|
|
778
|
+
let response
|
|
779
|
+
try {
|
|
780
|
+
response = await settings.mutate(namespace, owned, expected)
|
|
781
|
+
} catch (error) {
|
|
782
|
+
return { ok: false, error: errorText(error) }
|
|
783
|
+
}
|
|
784
|
+
if (response === null || typeof response !== 'object') return { ok: false, error: '宿主应答无效' }
|
|
785
|
+
if (response.ok === true) return { ok: true, value: response.value }
|
|
786
|
+
return { ok: false, error: responseError(response, '宿主拒绝了该写入') }
|
|
787
|
+
}
|
|
788
|
+
|
|
660
789
|
const mutate = (ops) => {
|
|
661
790
|
const owned = ops.map((op) => (op.op === 'set'
|
|
662
791
|
? { op: 'set', path: [...op.path], value: op.value }
|
|
@@ -665,26 +794,40 @@ window.__ModuleLoader__.load({
|
|
|
665
794
|
if (closed || channel === null) return false
|
|
666
795
|
// 首次写入前先完成一次 describe:否则命名空间与 revision 都还是初始值。
|
|
667
796
|
if (!loaded) await load()
|
|
668
|
-
|
|
669
|
-
|
|
670
|
-
|
|
671
|
-
|
|
672
|
-
await
|
|
673
|
-
|
|
797
|
+
const failures = []
|
|
798
|
+
for (const namespace of namespaceCandidates()) {
|
|
799
|
+
console.info('[dupguard] mutate → ' + namespace + ' revision=' + String(revision) +
|
|
800
|
+
' paths=' + JSON.stringify(owned.map((op) => op.path)))
|
|
801
|
+
let result = await attempt(namespace, owned, revision)
|
|
802
|
+
if (!result.ok) {
|
|
803
|
+
// 常见于并发写入导致的 revision 过期:回读后按新 revision 重试一次。
|
|
804
|
+
console.warn('[dupguard] mutate 被拒(' + namespace + '):' + result.error + ',回读后重试')
|
|
805
|
+
await load()
|
|
806
|
+
result = await attempt(namespace, owned, revision)
|
|
807
|
+
}
|
|
808
|
+
if (result.ok) {
|
|
809
|
+
lastWriteError = null
|
|
810
|
+
if (!closed) acceptView(result.value)
|
|
811
|
+
publish(undefined, { status: mirror.status, error: null, diag: '已写入命名空间 ' + namespace })
|
|
812
|
+
return true
|
|
813
|
+
}
|
|
814
|
+
failures.push(namespace + ':' + result.error)
|
|
674
815
|
}
|
|
675
|
-
|
|
676
|
-
|
|
677
|
-
|
|
678
|
-
}
|
|
679
|
-
|
|
680
|
-
return true
|
|
816
|
+
lastWriteError = failures.join(';')
|
|
817
|
+
// 写失败:回读宿主真实状态,避免界面停留在错误的本地值;并把原因显示到面板。
|
|
818
|
+
await load()
|
|
819
|
+
publish(undefined, { status: mirror.status, error: '写入被拒绝 —— ' + lastWriteError, diag: 'mutate 失败' })
|
|
820
|
+
return false
|
|
681
821
|
})
|
|
682
822
|
tail = task.then(() => {}, () => {})
|
|
683
823
|
return task
|
|
684
824
|
}
|
|
825
|
+
/** 供组件展示的最近一次写入失败原因。 */
|
|
826
|
+
const writeError = () => lastWriteError
|
|
685
827
|
return {
|
|
686
828
|
load,
|
|
687
829
|
mutate,
|
|
830
|
+
writeError,
|
|
688
831
|
close: () => {
|
|
689
832
|
closed = true
|
|
690
833
|
},
|
|
@@ -763,12 +906,20 @@ window.__ModuleLoader__.load({
|
|
|
763
906
|
controller: {
|
|
764
907
|
subscribe,
|
|
765
908
|
getSnapshot: () => snapshot,
|
|
766
|
-
set: (field, value) =>
|
|
767
|
-
|
|
768
|
-
|
|
769
|
-
|
|
770
|
-
|
|
771
|
-
|
|
909
|
+
set: (field, value) => {
|
|
910
|
+
if (channel === null) {
|
|
911
|
+
lastWriteError = '设置通道未就绪(channel=null)'
|
|
912
|
+
return Promise.resolve(false)
|
|
913
|
+
}
|
|
914
|
+
return channel.mutate([{ op: 'set', path: [field], value }])
|
|
915
|
+
},
|
|
916
|
+
unset: (field) => {
|
|
917
|
+
if (channel === null) {
|
|
918
|
+
lastWriteError = '设置通道未就绪(channel=null)'
|
|
919
|
+
return Promise.resolve(false)
|
|
920
|
+
}
|
|
921
|
+
return channel.mutate([{ op: 'unset', path: [field] }])
|
|
922
|
+
},
|
|
772
923
|
},
|
|
773
924
|
mirror: {
|
|
774
925
|
subscribe,
|
|
@@ -781,6 +932,8 @@ window.__ModuleLoader__.load({
|
|
|
781
932
|
},
|
|
782
933
|
/** 更新面板上的通道诊断文本(不改变数据状态)。 */
|
|
783
934
|
note: (text) => publish(undefined, { status: mirror.status, error: mirror.error, diag: text }),
|
|
935
|
+
/** 最近一次写入失败的宿主原因(旧版通道不提供时返回 null)。 */
|
|
936
|
+
writeReason: () => (channel !== null && typeof channel.writeError === 'function' ? channel.writeError() : null),
|
|
784
937
|
disposed: () => disposed,
|
|
785
938
|
/** 是否已有可用通道(旧版接上后即可停止新版探测)。 */
|
|
786
939
|
hasChannel: () => channel !== null,
|
|
@@ -921,6 +1074,7 @@ window.__ModuleLoader__.load({
|
|
|
921
1074
|
// 传函数而非快照:远程/只读状态可能由 describe 的 writable 或
|
|
922
1075
|
// remote.$host.isLoopback 稍后确定,渲染时必须取最新值。
|
|
923
1076
|
isLoopback: () => bridge.isLoopback(),
|
|
1077
|
+
writeReason: () => bridge.writeReason(),
|
|
924
1078
|
})
|
|
925
1079
|
ctx.slots.inject('settings.section', () => ctx.slots.register({
|
|
926
1080
|
name: 'settings.section',
|
package/lib/index.js
CHANGED
|
@@ -35,6 +35,15 @@ const LIMITS = {
|
|
|
35
35
|
codeBlockMultiplier: { min: 0, max: 100 },
|
|
36
36
|
}
|
|
37
37
|
|
|
38
|
+
/**
|
|
39
|
+
* 片段白名单(ignoredSubstrings)的上限。
|
|
40
|
+
*
|
|
41
|
+
* 两者同时界定「跨增量保留的尾巴长度(≤ 最长片段 - 1 码点)」与单增量匹配成本,
|
|
42
|
+
* 因此是代码常量而非可调设置。
|
|
43
|
+
*/
|
|
44
|
+
const IGNORED_SUBSTRINGS_MAX_COUNT = 64
|
|
45
|
+
const IGNORED_SUBSTRING_MAX_LENGTH = 64
|
|
46
|
+
|
|
38
47
|
/**
|
|
39
48
|
* 把字段标记为 volatile(DSH 的「可热改」约定)。
|
|
40
49
|
*
|
|
@@ -71,6 +80,7 @@ function markVolatile(schema) {
|
|
|
71
80
|
function createSettingsSchema(z) {
|
|
72
81
|
return z.object({
|
|
73
82
|
ignoredChars: markVolatile(z.array(z.string()).default([...CONFIG.ignoredChars])),
|
|
83
|
+
ignoredSubstrings: markVolatile(z.array(z.string()).default([...CONFIG.ignoredSubstrings])),
|
|
74
84
|
threshold: markVolatile(z.number().min(LIMITS.threshold.min).max(LIMITS.threshold.max).step(1).default(CONFIG.threshold)),
|
|
75
85
|
minUnitLength: markVolatile(z.number().min(LIMITS.minUnitLength.min).max(LIMITS.minUnitLength.max).step(1).default(CONFIG.minUnitLength)),
|
|
76
86
|
maxUnitLength: markVolatile(z.number().min(LIMITS.maxUnitLength.min).max(LIMITS.maxUnitLength.max).step(1).default(CONFIG.maxUnitLength)),
|
|
@@ -114,6 +124,13 @@ const CONFIG = {
|
|
|
114
124
|
// 模型忘记闭合围栏时,其后内容都按代码块处理。
|
|
115
125
|
skipCodeBlocks: true,
|
|
116
126
|
codeBlockMultiplier: 3,
|
|
127
|
+
// 片段白名单:整段匹配的多字符串(字面量、区分大小写、不支持正则)。
|
|
128
|
+
// 命中时先整段剔除,再做去空白与逐字符剔除 —— 用于 `|---|`、`------` 这类
|
|
129
|
+
// 由多个字符组成的固定片段:逐字符白名单只能忽略单个字符,组合片段的重复
|
|
130
|
+
// 仍会被计入。长片段优先匹配,避免 `---` 抢先破坏 `-----`。
|
|
131
|
+
// 上限:每项 ≤ IGNORED_SUBSTRING_MAX_LENGTH 码点、数组 ≤ IGNORED_SUBSTRINGS_MAX_COUNT 项。
|
|
132
|
+
// 代价:为跨增量匹配,每块最多保留(最长片段 - 1)个码点不参与检测(tail 延迟)。
|
|
133
|
+
ignoredSubstrings: [],
|
|
117
134
|
// 是否同时检测思考(reasoning)文本。默认开启:思考中的复读同样消耗 token,
|
|
118
135
|
// 应立即截停。注意:正常思考中若连续重复同一字符串 10 次以上(如"等等等等"),
|
|
119
136
|
// 也会被截停,属预期行为;需要只检测可见输出时可置为 false。
|
|
@@ -213,13 +230,87 @@ function stripIgnoredChars(text, ignored) {
|
|
|
213
230
|
return out
|
|
214
231
|
}
|
|
215
232
|
|
|
216
|
-
/** 供检测使用的增量清洗:去空白 +
|
|
217
|
-
function
|
|
218
|
-
let piece = config.stripWhitespace ? stripWhitespace(
|
|
233
|
+
/** 供检测使用的增量清洗:去空白 + 移除白名单字符(片段剔除在此之前完成)。 */
|
|
234
|
+
function sanitizePiece(text, config) {
|
|
235
|
+
let piece = config.stripWhitespace ? stripWhitespace(text) : text
|
|
219
236
|
if (config.ignoredChars.length > 0) piece = stripIgnoredChars(piece, config.ignoredChars)
|
|
220
237
|
return piece
|
|
221
238
|
}
|
|
222
239
|
|
|
240
|
+
/**
|
|
241
|
+
* 增量片段剔除器(每个文本块一份)。
|
|
242
|
+
*
|
|
243
|
+
* 语义:把配置里的片段当**字面量子串**整段剔除,长片段优先(避免 `---` 抢先破坏 `-----`),
|
|
244
|
+
* 剔除后再交给 sanitizePiece(去空白 → 逐字符白名单)。
|
|
245
|
+
*
|
|
246
|
+
* 跨增量:为避免片段被增量边界切断,末尾最多保留(最长片段 - 1)个**码点**不输出,
|
|
247
|
+
* 等下一次增量拼回后再匹配;flush() 吐出尾巴(块结束/流结束时调用),保证尾部文本仍参与检测。
|
|
248
|
+
* 代价:检测最多延迟(最长片段 - 1)个码点。
|
|
249
|
+
*
|
|
250
|
+
* 片段表为空时走零开销快路径(不缓冲、原样返回),因此默认行为与未启用该功能时完全一致。
|
|
251
|
+
*
|
|
252
|
+
* @param getPatterns - 读取当前片段表(支持热更新)。
|
|
253
|
+
* @param maxLength - 片段长度上限(决定保留的尾巴长度)。
|
|
254
|
+
*/
|
|
255
|
+
function createSubstringStripper(getPatterns, maxLength) {
|
|
256
|
+
let buffer = ''
|
|
257
|
+
let cachedSource = null
|
|
258
|
+
let cachedSorted = []
|
|
259
|
+
const keepLength = () => Math.max(1, maxLength) - 1
|
|
260
|
+
|
|
261
|
+
/** 长片段优先的片段表(按数组引用缓存,热更新后自动重建)。 */
|
|
262
|
+
const sortedPatterns = () => {
|
|
263
|
+
const list = getPatterns()
|
|
264
|
+
if (list !== cachedSource || list.length !== cachedSorted.length) {
|
|
265
|
+
cachedSource = list
|
|
266
|
+
cachedSorted = [...list].sort((left, right) => [...right].length - [...left].length)
|
|
267
|
+
}
|
|
268
|
+
return cachedSorted
|
|
269
|
+
}
|
|
270
|
+
|
|
271
|
+
/** 从 text 中移除所有片段出现(字面量,非正则)。 */
|
|
272
|
+
const removePatterns = (text, patterns) => {
|
|
273
|
+
let out = text
|
|
274
|
+
for (const pattern of patterns) {
|
|
275
|
+
if (pattern.length === 0) continue
|
|
276
|
+
if (out.indexOf(pattern) !== -1) out = out.split(pattern).join('')
|
|
277
|
+
}
|
|
278
|
+
return out
|
|
279
|
+
}
|
|
280
|
+
|
|
281
|
+
return {
|
|
282
|
+
/** 消费一段增量,返回可交给后续清洗的文本(可能为空串)。 */
|
|
283
|
+
push(delta) {
|
|
284
|
+
const patterns = sortedPatterns()
|
|
285
|
+
if (patterns.length === 0) {
|
|
286
|
+
// 快路径:未配置片段时不缓冲;若此前残留尾巴(刚被清空配置),先吐出。
|
|
287
|
+
if (buffer.length === 0) return delta
|
|
288
|
+
const carried = buffer
|
|
289
|
+
buffer = ''
|
|
290
|
+
return carried + delta
|
|
291
|
+
}
|
|
292
|
+
buffer += delta
|
|
293
|
+
const cleaned = removePatterns(buffer, patterns)
|
|
294
|
+
const codePoints = [...cleaned]
|
|
295
|
+
// 保留长度按**实际最长片段**计算(而非配置上限),否则检测会被无谓地拖后。
|
|
296
|
+
const longest = [...patterns[0]].length
|
|
297
|
+
const keep = Math.min(Math.max(0, longest - 1), keepLength(), codePoints.length)
|
|
298
|
+
if (keep === 0) {
|
|
299
|
+
buffer = ''
|
|
300
|
+
return cleaned
|
|
301
|
+
}
|
|
302
|
+
buffer = codePoints.slice(codePoints.length - keep).join('')
|
|
303
|
+
return codePoints.slice(0, codePoints.length - keep).join('')
|
|
304
|
+
},
|
|
305
|
+
/** 吐出保留的尾巴(块结束/流结束)。 */
|
|
306
|
+
flush() {
|
|
307
|
+
const out = buffer
|
|
308
|
+
buffer = ''
|
|
309
|
+
return out
|
|
310
|
+
},
|
|
311
|
+
}
|
|
312
|
+
}
|
|
313
|
+
|
|
223
314
|
/**
|
|
224
315
|
* 流式围栏代码块过滤器(每个文本块一份状态)。
|
|
225
316
|
*
|
|
@@ -416,6 +507,50 @@ function normalizeIgnoredChars(raw, fallback) {
|
|
|
416
507
|
/** 上一次「无效白名单条目」告警的签名,用于去重(参数每次调用都会重算)。 */
|
|
417
508
|
let lastIgnoredCharsWarning = ''
|
|
418
509
|
|
|
510
|
+
/**
|
|
511
|
+
* 片段白名单归一化:丢弃空值/非字符串、超长(> 64 码点)与超量(> 64 项)条目,
|
|
512
|
+
* 并去重。丢弃项只告警一次(同样内容),避免每次模型调用刷屏。
|
|
513
|
+
*/
|
|
514
|
+
function normalizeIgnoredSubstrings(raw, fallback) {
|
|
515
|
+
if (!Array.isArray(raw)) return [...fallback]
|
|
516
|
+
const kept = []
|
|
517
|
+
const dropped = []
|
|
518
|
+
for (const entry of raw) {
|
|
519
|
+
if (typeof entry !== 'string' || entry.length === 0) {
|
|
520
|
+
dropped.push(String(entry))
|
|
521
|
+
continue
|
|
522
|
+
}
|
|
523
|
+
if ([...entry].length > IGNORED_SUBSTRING_MAX_LENGTH) {
|
|
524
|
+
dropped.push(entry)
|
|
525
|
+
continue
|
|
526
|
+
}
|
|
527
|
+
if (kept.indexOf(entry) !== -1) continue
|
|
528
|
+
if (kept.length >= IGNORED_SUBSTRINGS_MAX_COUNT) {
|
|
529
|
+
dropped.push(entry)
|
|
530
|
+
continue
|
|
531
|
+
}
|
|
532
|
+
kept.push(entry)
|
|
533
|
+
}
|
|
534
|
+
if (dropped.length > 0) {
|
|
535
|
+
const signature = JSON.stringify(dropped)
|
|
536
|
+
if (signature !== lastSubstringWarning) {
|
|
537
|
+
lastSubstringWarning = signature
|
|
538
|
+
console.warn(
|
|
539
|
+
'[dupguard] 片段白名单条目无效,已忽略:' + JSON.stringify(dropped) +
|
|
540
|
+
'(要求非空字符串、每项 ≤ ' + String(IGNORED_SUBSTRING_MAX_LENGTH) +
|
|
541
|
+
' 码点、最多 ' + String(IGNORED_SUBSTRINGS_MAX_COUNT) + ' 项)。'
|
|
542
|
+
)
|
|
543
|
+
}
|
|
544
|
+
}
|
|
545
|
+
return kept
|
|
546
|
+
}
|
|
547
|
+
|
|
548
|
+
/** 上一次「无效片段白名单条目」告警的签名,用于去重。 */
|
|
549
|
+
let lastSubstringWarning = ''
|
|
550
|
+
|
|
551
|
+
/** 上一次打印的生效参数摘要,用于去重(旧版设置通道每次变更都会回调)。 */
|
|
552
|
+
let lastRuntimeSummary = ''
|
|
553
|
+
|
|
419
554
|
/**
|
|
420
555
|
* 为一次 llm/stream 调用创建守卫。
|
|
421
556
|
* 每次模型调用都会新建一份状态,互不干扰。
|
|
@@ -439,6 +574,7 @@ function createStreamGuard(options, runtime) {
|
|
|
439
574
|
stripped: '', // 去空白后的滚动窗口:仅用于检测
|
|
440
575
|
fence: createFenceFilter(), // 围栏代码块状态(skipCodeBlocks 开启时使用)
|
|
441
576
|
lastRunCode: undefined, // 上一片段的代码块归属(用于跨围栏边界清空缓冲)
|
|
577
|
+
substrings: createSubstringStripper(() => runtime.ignoredSubstrings, IGNORED_SUBSTRING_MAX_LENGTH),
|
|
442
578
|
toolCallId: undefined,
|
|
443
579
|
toolCallName: undefined,
|
|
444
580
|
toolCallArguments: '',
|
|
@@ -467,7 +603,7 @@ function createStreamGuard(options, runtime) {
|
|
|
467
603
|
if (skipInsideCode && b.lastRunCode !== undefined && run.code !== b.lastRunCode) b.stripped = ''
|
|
468
604
|
b.lastRunCode = run.code
|
|
469
605
|
if (skipInsideCode && run.code) continue
|
|
470
|
-
const piece =
|
|
606
|
+
const piece = sanitizePiece(b.substrings.push(run.text), runtime)
|
|
471
607
|
if (piece.length === 0) continue
|
|
472
608
|
b.stripped = (b.stripped + piece).slice(-runtime.detectionWindow)
|
|
473
609
|
const threshold = run.code ? codeThreshold : runtime.threshold
|
|
@@ -477,6 +613,21 @@ function createStreamGuard(options, runtime) {
|
|
|
477
613
|
return null
|
|
478
614
|
}
|
|
479
615
|
|
|
616
|
+
/**
|
|
617
|
+
* 块结束时吐出片段剔除器保留的尾巴(≤ 最长片段 - 1 码点),
|
|
618
|
+
* 让它仍参与检测;阈值沿用该块最后一次片段的代码块归属。
|
|
619
|
+
*/
|
|
620
|
+
function flushSubstringTail(b) {
|
|
621
|
+
if (b === undefined || b.substrings === undefined) return null
|
|
622
|
+
const tail = sanitizePiece(b.substrings.flush(), runtime)
|
|
623
|
+
if (tail.length === 0) return null
|
|
624
|
+
b.stripped = (b.stripped + tail).slice(-runtime.detectionWindow)
|
|
625
|
+
const threshold = b.lastRunCode === true
|
|
626
|
+
? runtime.threshold * (runtime.skipCodeBlocks === true ? runtime.codeBlockMultiplier : 1)
|
|
627
|
+
: runtime.threshold
|
|
628
|
+
return findRepeatedTail(b.stripped, threshold, runtime.minUnitLength, runtime.maxUnitLength)
|
|
629
|
+
}
|
|
630
|
+
|
|
480
631
|
/** 依据 StreamChunk 协议累积状态;命中时置 stopped。 */
|
|
481
632
|
function feed(chunk) {
|
|
482
633
|
switch (chunk.type) {
|
|
@@ -507,7 +658,7 @@ function createStreamGuard(options, runtime) {
|
|
|
507
658
|
if (chunk.name !== undefined) b.toolCallName = chunk.name
|
|
508
659
|
b.toolCallArguments += chunk.argumentsDelta
|
|
509
660
|
if (runtime.monitorToolArguments) {
|
|
510
|
-
const piece =
|
|
661
|
+
const piece = sanitizePiece(b.substrings.push(chunk.argumentsDelta), runtime)
|
|
511
662
|
b.stripped = (b.stripped + piece).slice(-runtime.detectionWindow)
|
|
512
663
|
const hit = findRepeatedTail(b.stripped, runtime.threshold, runtime.minUnitLength, runtime.maxUnitLength)
|
|
513
664
|
if (hit !== null) stopped = hit
|
|
@@ -515,11 +666,21 @@ function createStreamGuard(options, runtime) {
|
|
|
515
666
|
return
|
|
516
667
|
}
|
|
517
668
|
case 'block-end': {
|
|
669
|
+
const b = blocks.get(chunk.index)
|
|
670
|
+
const tailHit = flushSubstringTail(b)
|
|
671
|
+
if (tailHit !== null) stopped = tailHit
|
|
518
672
|
blocks.delete(chunk.index)
|
|
519
673
|
return
|
|
520
674
|
}
|
|
521
675
|
case 'usage':
|
|
522
|
-
case 'finish':
|
|
676
|
+
case 'finish': {
|
|
677
|
+
// 流结束时可能还有未收到 block-end 的块:把尾巴补上,避免尾部文本漏检。
|
|
678
|
+
for (const b of blocks.values()) {
|
|
679
|
+
const tailHit = flushSubstringTail(b)
|
|
680
|
+
if (tailHit !== null) stopped = tailHit
|
|
681
|
+
}
|
|
682
|
+
return
|
|
683
|
+
}
|
|
523
684
|
default:
|
|
524
685
|
return
|
|
525
686
|
}
|
|
@@ -621,6 +782,7 @@ function installStandingMountPatch(ctx) {
|
|
|
621
782
|
function settingsBase() {
|
|
622
783
|
return {
|
|
623
784
|
ignoredChars: [...CONFIG.ignoredChars],
|
|
785
|
+
ignoredSubstrings: [...CONFIG.ignoredSubstrings],
|
|
624
786
|
threshold: CONFIG.threshold,
|
|
625
787
|
minUnitLength: CONFIG.minUnitLength,
|
|
626
788
|
maxUnitLength: CONFIG.maxUnitLength,
|
|
@@ -714,6 +876,7 @@ function applySettingsToRuntime(runtime, value) {
|
|
|
714
876
|
}
|
|
715
877
|
const pickBool = (key) => (typeof source[key] === 'boolean' ? source[key] : fallback[key])
|
|
716
878
|
runtime.ignoredChars = normalizeIgnoredChars(source.ignoredChars, fallback.ignoredChars)
|
|
879
|
+
runtime.ignoredSubstrings = normalizeIgnoredSubstrings(source.ignoredSubstrings, fallback.ignoredSubstrings)
|
|
717
880
|
runtime.threshold = pickInt('threshold')
|
|
718
881
|
runtime.minUnitLength = pickInt('minUnitLength')
|
|
719
882
|
runtime.maxUnitLength = pickInt('maxUnitLength')
|
|
@@ -748,7 +911,10 @@ function readConfigValue(config, key) {
|
|
|
748
911
|
|
|
749
912
|
/** 把插件 Config 归一化成 applySettingsToRuntime 需要的设置值对象。 */
|
|
750
913
|
function settingsValueFromConfig(config) {
|
|
751
|
-
const value = {
|
|
914
|
+
const value = {
|
|
915
|
+
ignoredChars: readConfigValue(config, 'ignoredChars'),
|
|
916
|
+
ignoredSubstrings: readConfigValue(config, 'ignoredSubstrings'),
|
|
917
|
+
}
|
|
752
918
|
for (const key of Object.keys(LIMITS)) value[key] = readConfigValue(config, key)
|
|
753
919
|
for (const key of ['stripWhitespace', 'skipCodeBlocks', 'monitorReasoning', 'monitorToolArguments']) {
|
|
754
920
|
value[key] = readConfigValue(config, key)
|
|
@@ -756,6 +922,18 @@ function settingsValueFromConfig(config) {
|
|
|
756
922
|
return value
|
|
757
923
|
}
|
|
758
924
|
|
|
925
|
+
/**
|
|
926
|
+
* 生效参数摘要(宿主日志):确认宿主加载的版本与「设置是否真的进来了」。
|
|
927
|
+
* 片段白名单是否被读到,从这里一眼可见。
|
|
928
|
+
*/
|
|
929
|
+
function summarizeRuntime(runtime) { return '阈值 ' + String(runtime.threshold) +
|
|
930
|
+
'|窗口 ' + String(runtime.detectionWindow) +
|
|
931
|
+
'|最大单元 ' + String(runtime.maxUnitLength) +
|
|
932
|
+
'|字符白名单 ' + String(runtime.ignoredChars.length) + ' 项' +
|
|
933
|
+
'|片段白名单 ' + String(runtime.ignoredSubstrings.length) + ' 项' +
|
|
934
|
+
(runtime.ignoredSubstrings.length > 0 ? '(' + runtime.ignoredSubstrings.join('、') + ')' : '')
|
|
935
|
+
}
|
|
936
|
+
|
|
759
937
|
/**
|
|
760
938
|
* 从 DSH 的 settings 服务注册检测参数(仅常驻版);设置页可动态调整全部字段。
|
|
761
939
|
* 返回注销器;settings 服务缺失(无头环境)时不做任何事。
|
|
@@ -780,6 +958,12 @@ async function installSettings(ctx, runtime) {
|
|
|
780
958
|
})
|
|
781
959
|
const sync = () => {
|
|
782
960
|
applySettingsToRuntime(runtime, scope.get())
|
|
961
|
+
// 生效参数清单(同内容只打印一次):便于在宿主日志确认设置是否真的进来了。
|
|
962
|
+
const summary = summarizeRuntime(runtime)
|
|
963
|
+
if (summary !== lastRuntimeSummary) {
|
|
964
|
+
lastRuntimeSummary = summary
|
|
965
|
+
console.log('[dupguard] 生效参数:' + summary)
|
|
966
|
+
}
|
|
783
967
|
}
|
|
784
968
|
sync()
|
|
785
969
|
const stop = scope.watch(sync)
|
|
@@ -828,9 +1012,11 @@ function apply(ctx, config) {
|
|
|
828
1012
|
applySettingsToRuntime(next, settingsValueFromConfig(config))
|
|
829
1013
|
return next
|
|
830
1014
|
}
|
|
831
|
-
fromConfig() // 启动时先算一次:日志中暴露非法配置
|
|
1015
|
+
const startup = fromConfig() // 启动时先算一次:日志中暴露非法配置
|
|
832
1016
|
runtimeSource = fromConfig
|
|
833
1017
|
console.log('[dupguard] 设置参数取自插件 Config(DSH ≥ 0.1.7 模型;命名空间为 loader entry id)')
|
|
1018
|
+
// 生效参数清单:用于确认宿主加载的版本与「设置是否真的进来了」(片段白名单是否生效靠它判断)。
|
|
1019
|
+
console.log('[dupguard] 生效参数:' + summarizeRuntime(startup))
|
|
834
1020
|
})
|
|
835
1021
|
// llm/stream:包裹每次流式模型调用的瀑布事件。
|
|
836
1022
|
// 监听器返回包装后的 AsyncIterable,即成为本次调用对消费方可见的流。
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "dsh-dupguard",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.7.0",
|
|
4
4
|
"description": "Real-time repetition guard for DeepSeek Harness (DSH): stops model generation when the same string repeats >=10 times in the streamed output. 实时检测 DSH 大模型流式输出中的重复内容,同一字符串重复十次以上立即停止生成。",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"dsh-plugin",
|
package/plugin/host.js
CHANGED
|
@@ -17,6 +17,15 @@
|
|
|
17
17
|
// 均可在设置中动态调整并持久化;两版默认值保持一致。
|
|
18
18
|
// ============================================================================
|
|
19
19
|
|
|
20
|
+
/**
|
|
21
|
+
* 片段白名单(ignoredSubstrings)的上限。
|
|
22
|
+
*
|
|
23
|
+
* 两者同时界定「跨增量保留的尾巴长度(≤ 最长片段 - 1 码点)」与单增量匹配成本,
|
|
24
|
+
* 因此是代码常量而非可调设置。
|
|
25
|
+
*/
|
|
26
|
+
const IGNORED_SUBSTRINGS_MAX_COUNT = 64
|
|
27
|
+
const IGNORED_SUBSTRING_MAX_LENGTH = 64
|
|
28
|
+
|
|
20
29
|
const CONFIG = {
|
|
21
30
|
// 触发阈值:同一字符串连续重复次数达到该值时停止输出(用户需求:重复十次以上)。
|
|
22
31
|
// 语义为「>= threshold」,即第 10 次重复出现时就触发。
|
|
@@ -35,6 +44,13 @@ const CONFIG = {
|
|
|
35
44
|
// (如 "|---|---|"),正常表格输出会大量连续出现,不应视为复读。
|
|
36
45
|
// 默认忽略连字符与竖线;需要更严格的检测时可改为空数组 []。
|
|
37
46
|
ignoredChars: ['-', '|'],
|
|
47
|
+
// 片段白名单:整段匹配的多字符串(字面量、区分大小写、不支持正则)。
|
|
48
|
+
// 命中时先整段剔除,再做去空白与逐字符剔除 —— 用于 `|---|`、`------` 这类
|
|
49
|
+
// 由多个字符组成的固定片段:逐字符白名单只能忽略单个字符,组合片段的重复
|
|
50
|
+
// 仍会被计入。长片段优先匹配,避免 `---` 抢先破坏 `-----`。
|
|
51
|
+
// 上限:每项 ≤ IGNORED_SUBSTRING_MAX_LENGTH 码点、数组 ≤ IGNORED_SUBSTRINGS_MAX_COUNT 项。
|
|
52
|
+
// 代价:为跨增量匹配,每块最多保留(最长片段 - 1)个码点不参与检测(tail 延迟)。
|
|
53
|
+
ignoredSubstrings: [],
|
|
38
54
|
// 围栏代码块(``` / ~~~)内的重复检测按倍数分三档:
|
|
39
55
|
// ≥2 → 放宽:块内改用 threshold × codeBlockMultiplier 判定(默认 3),
|
|
40
56
|
// 既能放过正常代码,又能兜住真正的失控复读;
|
|
@@ -108,13 +124,130 @@ function stripIgnoredChars(text, ignored) {
|
|
|
108
124
|
return out
|
|
109
125
|
}
|
|
110
126
|
|
|
111
|
-
/** 供检测使用的增量清洗:去空白 +
|
|
112
|
-
function
|
|
113
|
-
let piece = config.stripWhitespace ? stripWhitespace(
|
|
127
|
+
/** 供检测使用的增量清洗:去空白 + 移除白名单字符(片段剔除在此之前完成)。 */
|
|
128
|
+
function sanitizePiece(text, config) {
|
|
129
|
+
let piece = config.stripWhitespace ? stripWhitespace(text) : text
|
|
114
130
|
if (config.ignoredChars.length > 0) piece = stripIgnoredChars(piece, config.ignoredChars)
|
|
115
131
|
return piece
|
|
116
132
|
}
|
|
117
133
|
|
|
134
|
+
/**
|
|
135
|
+
* 增量片段剔除器(每个文本块一份)。
|
|
136
|
+
*
|
|
137
|
+
* 语义:把配置里的片段当**字面量子串**整段剔除,长片段优先(避免 `---` 抢先破坏 `-----`),
|
|
138
|
+
* 剔除后再交给 sanitizePiece(去空白 → 逐字符白名单)。
|
|
139
|
+
*
|
|
140
|
+
* 跨增量:为避免片段被增量边界切断,末尾最多保留(最长片段 - 1)个**码点**不输出,
|
|
141
|
+
* 等下一次增量拼回后再匹配;flush() 吐出尾巴(块结束/流结束时调用),保证尾部文本仍参与检测。
|
|
142
|
+
* 代价:检测最多延迟(最长片段 - 1)个码点。
|
|
143
|
+
*
|
|
144
|
+
* 片段表为空时走零开销快路径(不缓冲、原样返回),因此默认行为与未启用该功能时完全一致。
|
|
145
|
+
*
|
|
146
|
+
* 与 lib/index.js 中的同名实现保持一致(两个入口行为必须相同)。
|
|
147
|
+
*
|
|
148
|
+
* @param getPatterns - 读取当前片段表(支持热更新)。
|
|
149
|
+
* @param maxLength - 片段长度上限(决定保留的尾巴长度)。
|
|
150
|
+
*/
|
|
151
|
+
function createSubstringStripper(getPatterns, maxLength) {
|
|
152
|
+
let buffer = ''
|
|
153
|
+
let cachedSource = null
|
|
154
|
+
let cachedSorted = []
|
|
155
|
+
const keepLength = () => Math.max(1, maxLength) - 1
|
|
156
|
+
|
|
157
|
+
/** 长片段优先的片段表(按数组引用缓存,热更新后自动重建)。 */
|
|
158
|
+
const sortedPatterns = () => {
|
|
159
|
+
const list = getPatterns()
|
|
160
|
+
if (list !== cachedSource || list.length !== cachedSorted.length) {
|
|
161
|
+
cachedSource = list
|
|
162
|
+
cachedSorted = [...list].sort((left, right) => [...right].length - [...left].length)
|
|
163
|
+
}
|
|
164
|
+
return cachedSorted
|
|
165
|
+
}
|
|
166
|
+
|
|
167
|
+
/** 从 text 中移除所有片段出现(字面量,非正则)。 */
|
|
168
|
+
const removePatterns = (text, patterns) => {
|
|
169
|
+
let out = text
|
|
170
|
+
for (const pattern of patterns) {
|
|
171
|
+
if (pattern.length === 0) continue
|
|
172
|
+
if (out.indexOf(pattern) !== -1) out = out.split(pattern).join('')
|
|
173
|
+
}
|
|
174
|
+
return out
|
|
175
|
+
}
|
|
176
|
+
|
|
177
|
+
return {
|
|
178
|
+
/** 消费一段增量,返回可交给后续清洗的文本(可能为空串)。 */
|
|
179
|
+
push(delta) {
|
|
180
|
+
const patterns = sortedPatterns()
|
|
181
|
+
if (patterns.length === 0) {
|
|
182
|
+
// 快路径:未配置片段时不缓冲;若此前残留尾巴(刚被清空配置),先吐出。
|
|
183
|
+
if (buffer.length === 0) return delta
|
|
184
|
+
const carried = buffer
|
|
185
|
+
buffer = ''
|
|
186
|
+
return carried + delta
|
|
187
|
+
}
|
|
188
|
+
buffer += delta
|
|
189
|
+
const cleaned = removePatterns(buffer, patterns)
|
|
190
|
+
const codePoints = [...cleaned]
|
|
191
|
+
// 保留长度按**实际最长片段**计算(而非配置上限),否则检测会被无谓地拖后。
|
|
192
|
+
const longest = [...patterns[0]].length
|
|
193
|
+
const keep = Math.min(Math.max(0, longest - 1), keepLength(), codePoints.length)
|
|
194
|
+
if (keep === 0) {
|
|
195
|
+
buffer = ''
|
|
196
|
+
return cleaned
|
|
197
|
+
}
|
|
198
|
+
buffer = codePoints.slice(codePoints.length - keep).join('')
|
|
199
|
+
return codePoints.slice(0, codePoints.length - keep).join('')
|
|
200
|
+
},
|
|
201
|
+
/** 吐出保留的尾巴(块结束/流结束)。 */
|
|
202
|
+
flush() {
|
|
203
|
+
const out = buffer
|
|
204
|
+
buffer = ''
|
|
205
|
+
return out
|
|
206
|
+
},
|
|
207
|
+
}
|
|
208
|
+
}
|
|
209
|
+
|
|
210
|
+
/**
|
|
211
|
+
* 片段白名单归一化:丢弃空值/非字符串、超长(> 64 码点)与超量(> 64 项)条目,
|
|
212
|
+
* 并去重。丢弃项只告警一次(同样内容),避免每次模型调用刷屏。
|
|
213
|
+
*/
|
|
214
|
+
function normalizeIgnoredSubstrings(raw, fallback) {
|
|
215
|
+
if (!Array.isArray(raw)) return [...fallback]
|
|
216
|
+
const kept = []
|
|
217
|
+
const dropped = []
|
|
218
|
+
for (const entry of raw) {
|
|
219
|
+
if (typeof entry !== 'string' || entry.length === 0) {
|
|
220
|
+
dropped.push(String(entry))
|
|
221
|
+
continue
|
|
222
|
+
}
|
|
223
|
+
if ([...entry].length > IGNORED_SUBSTRING_MAX_LENGTH) {
|
|
224
|
+
dropped.push(entry)
|
|
225
|
+
continue
|
|
226
|
+
}
|
|
227
|
+
if (kept.indexOf(entry) !== -1) continue
|
|
228
|
+
if (kept.length >= IGNORED_SUBSTRINGS_MAX_COUNT) {
|
|
229
|
+
dropped.push(entry)
|
|
230
|
+
continue
|
|
231
|
+
}
|
|
232
|
+
kept.push(entry)
|
|
233
|
+
}
|
|
234
|
+
if (dropped.length > 0) {
|
|
235
|
+
const signature = JSON.stringify(dropped)
|
|
236
|
+
if (signature !== lastSubstringWarning) {
|
|
237
|
+
lastSubstringWarning = signature
|
|
238
|
+
console.warn(
|
|
239
|
+
'[dupguard] 片段白名单条目无效,已忽略:' + JSON.stringify(dropped) +
|
|
240
|
+
'(要求非空字符串、每项 ≤ ' + String(IGNORED_SUBSTRING_MAX_LENGTH) +
|
|
241
|
+
' 码点、最多 ' + String(IGNORED_SUBSTRINGS_MAX_COUNT) + ' 项)。'
|
|
242
|
+
)
|
|
243
|
+
}
|
|
244
|
+
}
|
|
245
|
+
return kept
|
|
246
|
+
}
|
|
247
|
+
|
|
248
|
+
/** 上一次「无效片段白名单条目」告警的签名,用于去重。 */
|
|
249
|
+
let lastSubstringWarning = ''
|
|
250
|
+
|
|
118
251
|
/**
|
|
119
252
|
* 流式围栏代码块过滤器(每个文本块一份状态)。
|
|
120
253
|
*
|
|
@@ -286,6 +419,7 @@ function createStreamGuard(options) {
|
|
|
286
419
|
stripped: '', // 去空白后的滚动窗口:仅用于检测
|
|
287
420
|
fence: createFenceFilter(), // 围栏代码块状态(skipCodeBlocks 开启时使用)
|
|
288
421
|
lastRunCode: undefined, // 上一片段的代码块归属(用于跨围栏边界清空缓冲)
|
|
422
|
+
substrings: createSubstringStripper(() => CONFIG.ignoredSubstrings, IGNORED_SUBSTRING_MAX_LENGTH),
|
|
289
423
|
toolCallId: undefined,
|
|
290
424
|
toolCallName: undefined,
|
|
291
425
|
toolCallArguments: '',
|
|
@@ -314,7 +448,7 @@ function createStreamGuard(options) {
|
|
|
314
448
|
if (skipInsideCode && b.lastRunCode !== undefined && run.code !== b.lastRunCode) b.stripped = ''
|
|
315
449
|
b.lastRunCode = run.code
|
|
316
450
|
if (skipInsideCode && run.code) continue
|
|
317
|
-
const piece =
|
|
451
|
+
const piece = sanitizePiece(b.substrings.push(run.text), CONFIG)
|
|
318
452
|
if (piece.length === 0) continue
|
|
319
453
|
b.stripped = (b.stripped + piece).slice(-CONFIG.detectionWindow)
|
|
320
454
|
const threshold = run.code ? codeThreshold : CONFIG.threshold
|
|
@@ -324,6 +458,21 @@ function createStreamGuard(options) {
|
|
|
324
458
|
return null
|
|
325
459
|
}
|
|
326
460
|
|
|
461
|
+
/**
|
|
462
|
+
* 块结束时吐出片段剔除器保留的尾巴(≤ 最长片段 - 1 码点),
|
|
463
|
+
* 让它仍参与检测;阈值沿用该块最后一次片段的代码块归属。
|
|
464
|
+
*/
|
|
465
|
+
function flushSubstringTail(b) {
|
|
466
|
+
if (b === undefined || b.substrings === undefined) return null
|
|
467
|
+
const tail = sanitizePiece(b.substrings.flush(), CONFIG)
|
|
468
|
+
if (tail.length === 0) return null
|
|
469
|
+
b.stripped = (b.stripped + tail).slice(-CONFIG.detectionWindow)
|
|
470
|
+
const threshold = b.lastRunCode === true
|
|
471
|
+
? CONFIG.threshold * (CONFIG.skipCodeBlocks === true ? CONFIG.codeBlockMultiplier : 1)
|
|
472
|
+
: CONFIG.threshold
|
|
473
|
+
return findRepeatedTail(b.stripped, threshold, CONFIG.minUnitLength, CONFIG.maxUnitLength)
|
|
474
|
+
}
|
|
475
|
+
|
|
327
476
|
/** 依据 StreamChunk 协议累积状态;命中时置 stopped。 */
|
|
328
477
|
function feed(chunk) {
|
|
329
478
|
switch (chunk.type) {
|
|
@@ -354,7 +503,7 @@ function createStreamGuard(options) {
|
|
|
354
503
|
if (chunk.name !== undefined) b.toolCallName = chunk.name
|
|
355
504
|
b.toolCallArguments += chunk.argumentsDelta
|
|
356
505
|
if (CONFIG.monitorToolArguments) {
|
|
357
|
-
const piece =
|
|
506
|
+
const piece = sanitizePiece(b.substrings.push(chunk.argumentsDelta), CONFIG)
|
|
358
507
|
b.stripped = (b.stripped + piece).slice(-CONFIG.detectionWindow)
|
|
359
508
|
const hit = findRepeatedTail(b.stripped, CONFIG.threshold, CONFIG.minUnitLength, CONFIG.maxUnitLength)
|
|
360
509
|
if (hit !== null) stopped = hit
|
|
@@ -362,11 +511,21 @@ function createStreamGuard(options) {
|
|
|
362
511
|
return
|
|
363
512
|
}
|
|
364
513
|
case 'block-end': {
|
|
514
|
+
const b = blocks.get(chunk.index)
|
|
515
|
+
const tailHit = flushSubstringTail(b)
|
|
516
|
+
if (tailHit !== null) stopped = tailHit
|
|
365
517
|
blocks.delete(chunk.index)
|
|
366
518
|
return
|
|
367
519
|
}
|
|
368
520
|
case 'usage':
|
|
369
|
-
case 'finish':
|
|
521
|
+
case 'finish': {
|
|
522
|
+
// 流结束时可能还有未收到 block-end 的块:把尾巴补上,避免尾部文本漏检。
|
|
523
|
+
for (const b of blocks.values()) {
|
|
524
|
+
const tailHit = flushSubstringTail(b)
|
|
525
|
+
if (tailHit !== null) stopped = tailHit
|
|
526
|
+
}
|
|
527
|
+
return
|
|
528
|
+
}
|
|
370
529
|
default:
|
|
371
530
|
return
|
|
372
531
|
}
|