dsh-yolo-mode 0.5.0 → 0.5.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,41 @@
2
2
 
3
3
  本项目遵循 [Semantic Versioning](https://semver.org/)。
4
4
 
5
+ ## [0.5.1] - 2026-09-22
6
+
7
+ ### 修复:裁判对推理型模型恒失败(Issue #1)
8
+
9
+ - **默认 `judge.maxTokens` 256 → 4096**(`lib/policy.js` 默认配置与 `normalizeConfig`
10
+ 合并值、`lib/judge.js` 兜底值同步)。根因:推理型裁判模型把 token 预算全部消耗在
11
+ reasoning 块上,`maxTokens` 过小时 `content` 为空(`finish_reason=length`),裁判抛
12
+ `BAD_OUTPUT` 并按 `error` 回退——`balanced` 预设下表现为「每次都转人工」,而
13
+ `judgeConfigured:true` 掩盖了失败。用户实测 4096 为可用值(256/1024 输出为 0)。
14
+ - **审计条目新增 `error` 字段**(`lib/index.js`):裁判失败时写入 `JudgeError.code`
15
+ (`BAD_OUTPUT` / `TIMEOUT` / `STREAM_ERROR` / `NO_ADAPTER` 等),使
16
+ `outcome:"delegate"` 的条目可区分「裁判失败转人工」与「裁判主动授意转人工」;
17
+ 成功路径字段形状不变(不新增 `error`)。
18
+ - **失败可见**(`lib/state.js`):新增 `stats.judgeFailures` 计数与
19
+ `judgeHealth.lastError/lastErrorTime`,经 `getStatusPayload` 以
20
+ `judgeErrors: { count, lastError?, lastErrorTime? }` 暴露;裁判异常不再仅停留于
21
+ `logger.warn`。
22
+ - **文档**:README 标注推理型模型必须调大 `judge.maxTokens`(默认已改 4096),并补充
23
+ 审计 `error` 字段与 `judgeErrors` 载荷说明。
24
+
25
+ ### 测试
26
+
27
+ - `test/policy.test.mjs`:默认 `maxTokens` 断言更新为 4096。
28
+ - `test/judge.test.mjs`:新增「未显式传 `maxTokens` → 默认 4096 传入 `llm.stream`」用例。
29
+ - `test/state.test.mjs`:新增裁判失败时 `judgeFailures` 递增、`judgeErrors` 载荷装配用例。
30
+ - `test/probe.test.mjs`:新增真实-Cordis 用例——裁判产出非 JSON(`BAD_OUTPUT`)→
31
+ 审计条目带 `error:"BAD_OUTPUT"` 且回退 delegate;成功路径断言 recent 条目**不含** `error`。
32
+
33
+ ### 兼容性
34
+
35
+ - **新增 `dsh.compatibility` 声明**(DSH STORE 上架契约):逐版本声明
36
+ `dshReleases` 兼容矩阵——`0.1.5-rc.1` / `0.1.5-rc.2` / `0.1.6-alpha.2` 均为
37
+ `compatible`(三版本已在本机真实装载运行,插件正常加载、零错误);`node` 范围
38
+ `>=20`,与 `engines.node` 一致。未实测的版本不声明(扫描时按 `unknown` 处理)。
39
+
5
40
  ## [0.5.0] - 2026-09-04
6
41
 
7
42
  ### 兼容:DSH 0.1.2-alpha.4
package/README.md CHANGED
@@ -6,10 +6,10 @@
6
6
 
7
7
  [![npm version](https://img.shields.io/npm/v/dsh-yolo-mode)](https://www.npmjs.com/package/dsh-yolo-mode)
8
8
  [![license](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
9
- [![DeepSeek Harness](https://img.shields.io/badge/DeepSeek%20Harness-0.1.2--alpha.4-blue)](https://github.com/deepseek-ai/deepseek-harness)
9
+ [![DeepSeek Harness](https://img.shields.io/badge/DeepSeek%20Harness-0.1.2--rc.1-blue)](https://github.com/deepseek-ai/deepseek-harness)
10
10
  [![Awesome DSH Plugin](https://awesome-dsh-plugin.com/badge.svg)](https://awesome-dsh-plugin.com)
11
11
 
12
- > **兼容性**:v0.5.0 支持 DSH **0.1.2-alpha.4+**(真实-Cordis 探针 125/125 通过;settings 已迁移至 alpha.4 的 `SettingsProvider` 类服务)。旧版 DSH(0.1.0-rc.6 / rc.8)请使用最后兼容的 npm 版本 **0.4.1**。
12
+ > **兼容性**:v0.5.1 支持 DSH **0.1.2-alpha.4+**,已在当前运行时 **0.1.2-rc.1** 实测(真实-Cordis 探针通过;settings 已迁移至 alpha.4 的 `SettingsProvider` 类服务)。旧版 DSH(0.1.0-rc.6 / rc.8)请使用最后兼容的 npm 版本 **0.4.1**。
13
13
 
14
14
  ---
15
15
 
@@ -77,11 +77,13 @@ dsh plugin --profile web add <项目绝对路径>
77
77
  | `judge.model` | `string` | `''` | 裁判模型;与 provider 同非空才启用裁判 |
78
78
  | `judge.systemPrompt` | `string` | `''` | 裁判 system prompt;空 = 按预设取默认 |
79
79
  | `judge.timeoutMs` | `number` | `20000` | 单次裁判超时(毫秒) |
80
- | `judge.maxTokens` | `number` | `256` | 裁判输出最大 token 数 |
80
+ | `judge.maxTokens` | `number` | `4096` | 裁判输出最大 token 数(推理型模型需调大,见下方说明) |
81
81
  | `judge.concurrency` | `number` | `2` | 并发裁判上限,溢出按错误回退 |
82
82
  | `includeSubagents` | `boolean` | `true` | 子代理会话是否同样裁决 |
83
83
  | `auditFile` | `string` | `''` | 审计日志路径;空 = `%TEMP%/dsh-yolo/judge.log` |
84
84
 
85
+ > ⚠️ **推理型模型必须调大 `judge.maxTokens`**:若裁判模型会先输出一段内部推理(reasoning / chain-of-thought)再给出结论 JSON,则 token 预算会被推理内容消耗;预算过小时 `content` 为空、`finish_reason=length`,裁判以 `BAD_OUTPUT` 失败并回退为「转人工」(默认预设 `balanced` 下表现为**每次都弹人工审批**)。默认值已从 `256` 提升至 **`4096`**;若仍遇到回退,请继续调大(如 `8192`)。审计日志中此类失败会带 `error: "BAD_OUTPUT"`。
86
+
85
87
  ### 权限层级(`levels`)
86
88
 
87
89
  ```yaml
@@ -130,7 +132,8 @@ levels:
130
132
  - **防回环**:裁判 prompt 与 agent 上下文隔离,防止模型借 Web 审批回环自批准 `danger-full-access`。
131
133
  - **不改写策略**:仅在 `ask` 策略下作为应答者,不改变 DSH 的沙箱 / 审批词汇。
132
134
  - **默认保守**:默认预设 `balanced`(不确定转人工),不默认启用 `permissive` / `yolo`。
133
- - **审计**:每次裁决落一行 JSONL,含 `{time, sessionId, origin, toolName, callId?, targetMode, currentMode, justification, decision, outcome, reason?}`。
135
+ - **审计**:每次裁决落一行 JSONL,含 `{time, sessionId, origin, toolName, callId?, targetMode, currentMode, justification, decision, outcome, reason?, error?}`。`error` 仅在**裁判失败**时出现(值为 `JudgeError` 错误码,如 `BAD_OUTPUT` / `TIMEOUT` / `STREAM_ERROR` / `NO_ADAPTER`);`outcome:"delegate"` 且有 `error`=裁判失败转人工,无 `error`=裁判主动转人工,二者可区分。
136
+ - **失败可见**:statusView 载荷携带 `judgeErrors: { count, lastError?, lastErrorTime? }`(累计裁判失败次数与最近一次错误码),裁判异常不再只停留在 `logger.warn`。
134
137
 
135
138
  ## 开发
136
139
 
package/lib/index.js CHANGED
@@ -247,13 +247,17 @@ export function apply(ctx, rawConfig) {
247
247
 
248
248
  // 7. 裁决映射(judge 走裁判,含未配置/失败/不确定回退)。
249
249
  const judge = getJudge();
250
- let judgeReason; // 裁判 reason 仅在 judge 路径产出时记录(审计 reason?)。
250
+ let judgeReason; // 裁判 reason 仅在 judge 路径产出时记录(审计 reason 字段)。
251
+ let judgeError; // 裁判失败的错误码(审计 error 字段;成功/非 judge 路径为 undefined)。
251
252
  const result = await (async () => {
252
253
  if (decision === 'allow') return { outcome: 'allowed-once' };
253
254
  if (decision === 'deny') return { outcome: 'rejected' };
254
255
  if (decision === 'delegate') return { delegate: true };
255
256
  // decision === 'judge'
256
- if (!judge) return fallback('error', cfg);
257
+ if (!judge) {
258
+ judgeError = 'NO_ADAPTER'; // provider/model 未配置:与 judge.js 的 NO_ADAPTER 同语义
259
+ return fallback('error', cfg);
260
+ }
257
261
  try {
258
262
  const r = await judge({
259
263
  toolName: req.toolName,
@@ -268,6 +272,7 @@ export function apply(ctx, rawConfig) {
268
272
  if (r.decision === 'deny') return { outcome: 'rejected' };
269
273
  return fallback('unsure', cfg); // 不确定
270
274
  } catch (err) {
275
+ judgeError = (err && typeof err === 'object' && err.code) ? String(err.code) : 'UNKNOWN';
271
276
  logger.warn('LLM 裁判失败,按预设 error 回退', errorDescriptor(err));
272
277
  return fallback('error', cfg);
273
278
  }
@@ -275,7 +280,8 @@ export function apply(ctx, rawConfig) {
275
280
 
276
281
  const outcome = result.delegate ? 'delegate' : result.outcome;
277
282
 
278
- // 8. 审计。
283
+ // 8. 审计。裁判失败时带 error 字段(区分「裁判失败」与「裁判主动转人工」);
284
+ // 成功路径字段形状不变(不新增 error)。
279
285
  audit({
280
286
  time: Date.now(),
281
287
  sessionId: (session && session.id) || (req.agent && req.agent.id),
@@ -288,6 +294,7 @@ export function apply(ctx, rawConfig) {
288
294
  decision,
289
295
  outcome,
290
296
  reason: judgeReason,
297
+ ...(judgeError !== undefined ? { error: judgeError } : {}),
291
298
  });
292
299
 
293
300
  // 9. delegate → next() 透明委托;否则返回归一化 outcome。
package/lib/judge.js CHANGED
@@ -151,7 +151,7 @@ function throwAbort(signal, upstream) {
151
151
  * @param {string} opts.model model id
152
152
  * @param {string} [opts.systemPrompt] 空/缺省 → 内置裁判 prompt(含防回环要求)
153
153
  * @param {number} [opts.timeoutMs=20000] 单次裁判调用超时(毫秒)
154
- * @param {number} [opts.maxTokens=256] 最大输出 token
154
+ * @param {number} [opts.maxTokens=4096] 最大输出 token(推理型模型需调大,见 README)
155
155
  * @param {number} [opts.concurrency=2] 信号量上限
156
156
  * @param {AbortSignal} [opts.signal] 上游取消信号(ABORTED 时中止;可空)
157
157
  * @returns {Function} async judge(input) -> {{decision:'allow'|'deny'|'unsure', reason:string}}
@@ -159,7 +159,7 @@ function throwAbort(signal, upstream) {
159
159
  export function createJudge({ llm, provider, model, systemPrompt, timeoutMs, maxTokens, concurrency, signal }) {
160
160
  const sys = typeof systemPrompt === 'string' && systemPrompt.trim() !== '' ? systemPrompt : DEFAULT_SYSTEM_PROMPT
161
161
  const ms = isPositiveInt(timeoutMs) ? timeoutMs : 20000
162
- const mt = isPositiveInt(maxTokens) ? maxTokens : 256
162
+ const mt = isPositiveInt(maxTokens) ? maxTokens : 4096
163
163
  const cap = isPositiveInt(concurrency) ? concurrency : 2
164
164
 
165
165
  // 信号量计数(活跃调用数)。进入者先同步占位,用后的 try/finally 释放。
package/lib/policy.js CHANGED
@@ -87,7 +87,7 @@ function defaultConfig() {
87
87
  preset: 'balanced',
88
88
  modes: ['workspace-write'],
89
89
  levels: {},
90
- judge: { provider: '', model: '', systemPrompt: '', timeoutMs: 20000, maxTokens: 256, concurrency: 2 },
90
+ judge: { provider: '', model: '', systemPrompt: '', timeoutMs: 20000, maxTokens: 4096, concurrency: 2 },
91
91
  includeSubagents: true,
92
92
  auditFile: '',
93
93
  }
@@ -187,7 +187,7 @@ export function normalizeConfig(raw) {
187
187
  model: jraw.model === undefined ? '' : jraw.model,
188
188
  systemPrompt: jraw.systemPrompt === undefined ? '' : jraw.systemPrompt,
189
189
  timeoutMs: jraw.timeoutMs === undefined ? 20000 : jraw.timeoutMs,
190
- maxTokens: jraw.maxTokens === undefined ? 256 : jraw.maxTokens,
190
+ maxTokens: jraw.maxTokens === undefined ? 4096 : jraw.maxTokens,
191
191
  concurrency: jraw.concurrency === undefined ? 2 : jraw.concurrency,
192
192
  }
193
193
  for (const field of ['provider', 'model', 'systemPrompt']) {
package/lib/state.js CHANGED
@@ -20,8 +20,13 @@ import { resolveAuditFile } from './audit.js'
20
20
  /** recent 环形缓冲上限(design.md §11.1/12.4,≤20)。 */
21
21
  export const RECENT_CAP = 20;
22
22
 
23
- /** 模块级统计计数(设计 §3.8/12.3;主条目内存写,桥接条目只读)。 */
24
- export const stats = { total: 0, allowed: 0, rejected: 0, delegated: 0 };
23
+ /** 模块级统计计数(设计 §3.8/12.3;主条目内存写,桥接条目只读)。
24
+ * judgeFailures 统计裁判失败次数(审计条目带 error 字段的次数),使「裁判挂了」
25
+ * 在状态面板可见、区别于「裁判主动转人工」。 */
26
+ export const stats = { total: 0, allowed: 0, rejected: 0, delegated: 0, judgeFailures: 0 };
27
+
28
+ /** 最近一次裁判失败的健康信息(code + 时间戳);无失败时 lastError 为 undefined。 */
29
+ export const judgeHealth = { lastError: undefined, lastErrorTime: undefined };
25
30
 
26
31
  /** recent 环形缓冲:新条目 unshift 到头部,超过上限截断尾部(倒序,≤RECENT_CAP)。 */
27
32
  const recent = [];
@@ -55,7 +60,8 @@ export function sessionOrigin(session) {
55
60
  * 记录一次裁决到模块级统计与 recent 环(主条目 audit() 调用)。
56
61
  * @param {object} entry 审计条目:
57
62
  * { time, sessionId, origin, toolName?, callId?, targetMode, currentMode?,
58
- * justification, decision, outcome, reason? }
63
+ * justification, decision, outcome, reason?, error? }
64
+ * error 仅在裁判失败时存在(JudgeError.code);成功路径不携带该字段。
59
65
  */
60
66
  export function recordDecision(entry) {
61
67
  recent.unshift({
@@ -65,6 +71,7 @@ export function recordDecision(entry) {
65
71
  decision: entry.decision,
66
72
  outcome: entry.outcome,
67
73
  ...(entry.reason !== undefined ? { reason: entry.reason } : {}),
74
+ ...(entry.error !== undefined ? { error: entry.error } : {}),
68
75
  });
69
76
  if (recent.length > RECENT_CAP) recent.length = RECENT_CAP;
70
77
 
@@ -72,13 +79,18 @@ export function recordDecision(entry) {
72
79
  if (entry.outcome === 'delegate') stats.delegated += 1;
73
80
  else if (entry.outcome === 'allowed-once') stats.allowed += 1;
74
81
  else if (entry.outcome === 'rejected') stats.rejected += 1;
82
+ if (entry.error !== undefined) {
83
+ stats.judgeFailures += 1;
84
+ judgeHealth.lastError = entry.error;
85
+ judgeHealth.lastErrorTime = entry.time;
86
+ }
75
87
  }
76
88
 
77
89
  /**
78
90
  * 以浅拷贝装配 statusView 载荷(design.md §12.4)。
79
91
  * @param {object} cfg 含 { preset, judge?: { provider?, model? }, auditFile? } 的配置字形
80
92
  * (主条目传 effectiveConfig();桥接条目传 settings.describe 推导的 view 字形)
81
- * @returns {{preset:string, judgeConfigured:boolean, presetDefaults:object, stats:object, recent:Array<object>, auditFile:string}}
93
+ * @returns {{preset:string, judgeConfigured:boolean, presetDefaults:object, stats:object, judgeErrors:object, recent:Array<object>, auditFile:string}}
82
94
  */
83
95
  export function getStatusPayload(cfg) {
84
96
  const c = cfg && typeof cfg === 'object' ? cfg : {};
@@ -88,6 +100,13 @@ export function getStatusPayload(cfg) {
88
100
  judgeConfigured: Boolean(j.provider && j.model),
89
101
  presetDefaults,
90
102
  stats: { ...stats },
103
+ // 裁判失败健康信息:累计失败次数 + 最近一次错误码/时间(无失败时 lastError 缺省)。
104
+ judgeErrors: {
105
+ count: stats.judgeFailures,
106
+ ...(judgeHealth.lastError !== undefined
107
+ ? { lastError: judgeHealth.lastError, lastErrorTime: judgeHealth.lastErrorTime }
108
+ : {}),
109
+ },
91
110
  recent: recent.map((r) => ({ ...r })),
92
111
  // 生效的审计日志路径(resolveAuditFile 与主条目 audit() 同一解析规则),
93
112
  // 供客户端「打开日志」按钮展示/复用。
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "dsh-yolo-mode",
3
- "version": "0.5.0",
3
+ "version": "0.5.1",
4
4
  "description": "dsh-yolo-mode —— DeepSeek Harness 双面包插件:当会话处于可写沙箱模式且审批策略为 ask 时,用大模型自动裁决沙箱升权申请,支持内置预设与自定义权限层级,并提供宿主 settings + 自发布设置桥(/yolo-mode)与 Web 客户端 UI。",
5
5
  "type": "module",
6
6
  "main": "lib/index.js",
@@ -41,14 +41,14 @@
41
41
  "peerDependencies": {
42
42
  "react": "^18.2.0",
43
43
  "@deepseek-ai/cordis": "^4.0.1",
44
- "@deepseek-ai/dsh-llm": "^0.1.2-alpha.4",
45
- "@deepseek-ai/dsh-timeout": "^0.1.2-alpha.4",
46
- "@deepseek-ai/dsh-settings": "^0.1.2-alpha.4",
44
+ "@deepseek-ai/dsh-llm": "^0.1.5-rc.1",
45
+ "@deepseek-ai/dsh-timeout": "^0.1.5-rc.1",
46
+ "@deepseek-ai/dsh-settings": "^0.1.5-rc.1",
47
47
  "@deepseek-ai/dsh-host-apiproxy": "^0.1.1-rc.2",
48
48
  "@deepseek-ai/dsh-client-runtime": "^0.1.1-rc.2",
49
- "@deepseek-ai/dsh-client-connection": "^0.1.2-alpha.4",
50
- "@deepseek-ai/dsh-client-ui-slots": "^0.1.2-alpha.4",
51
- "@deepseek-ai/dsh-client-locale": "^0.1.2-alpha.4",
49
+ "@deepseek-ai/dsh-client-connection": "^0.1.5-rc.1",
50
+ "@deepseek-ai/dsh-client-ui-slots": "^0.1.5-rc.1",
51
+ "@deepseek-ai/dsh-client-locale": "^0.1.5-rc.1",
52
52
  "@deepseek-ai/schemastery": "^3.18.1"
53
53
  },
54
54
  "devDependencies": {
@@ -76,6 +76,14 @@
76
76
  "connection",
77
77
  "remote"
78
78
  ]
79
+ },
80
+ "compatibility": {
81
+ "dshReleases": {
82
+ "0.1.5-rc.1": "compatible",
83
+ "0.1.5-rc.2": "compatible",
84
+ "0.1.6-alpha.2": "compatible"
85
+ },
86
+ "node": ">=20"
79
87
  }
80
88
  }
81
- }
89
+ }