dsh-jev-guard 0.5.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (52) hide show
  1. package/CHANGELOG.md +285 -0
  2. package/CHANGELOG.zh-CN.md +271 -0
  3. package/DEPLOY.md +202 -0
  4. package/DEPLOY.zh-CN.md +200 -0
  5. package/LICENSE +21 -0
  6. package/README.md +316 -0
  7. package/README.zh-CN.md +315 -0
  8. package/START-HERE.md +97 -0
  9. package/START-HERE.zh-CN.md +97 -0
  10. package/adapters/README.md +37 -0
  11. package/adapters/README.zh-CN.md +37 -0
  12. package/adapters/dsh/index.js +502 -0
  13. package/bin/guard.mjs +634 -0
  14. package/config.example.json +52 -0
  15. package/cordis.patch.yml +120 -0
  16. package/docs/AGENT-TASK-dsh.md +134 -0
  17. package/docs/AGENT-TASK-dsh.zh-CN.md +131 -0
  18. package/docs/ARCHITECTURE.md +118 -0
  19. package/docs/ARCHITECTURE.zh-CN.md +117 -0
  20. package/docs/DECISIONS.md +469 -0
  21. package/docs/DECISIONS.zh-CN.md +449 -0
  22. package/docs/DSH-INTEGRATION.md +178 -0
  23. package/docs/DSH-INTEGRATION.zh-CN.md +171 -0
  24. package/docs/MEASUREMENTS.md +433 -0
  25. package/docs/MEASUREMENTS.zh-CN.md +450 -0
  26. package/docs/USER-INTERVENTION.md +141 -0
  27. package/docs/USER-INTERVENTION.zh-CN.md +143 -0
  28. package/docs/VERIFICATION.md +279 -0
  29. package/docs/VERIFICATION.zh-CN.md +278 -0
  30. package/lib/audit.js +228 -0
  31. package/lib/gate.js +720 -0
  32. package/lib/i18n.js +575 -0
  33. package/lib/quota.js +389 -0
  34. package/lib/rules.js +174 -0
  35. package/lib/token.js +154 -0
  36. package/lib/verdict.js +285 -0
  37. package/package.json +82 -0
  38. package/tools/check-doc-pairs.mjs +158 -0
  39. package/tools/extract-commands.mjs +156 -0
  40. package/tools/gate-cli.mjs +240 -0
  41. package/tools/probe-prompt-lang.mjs +238 -0
  42. package/tools/probe-scripts.mjs +143 -0
  43. package/tools/report-result.mjs +146 -0
  44. package/tools/selftest-audit.mjs +93 -0
  45. package/tools/selftest-entry.mjs +177 -0
  46. package/tools/selftest-i18n.mjs +177 -0
  47. package/tools/selftest-quota.mjs +260 -0
  48. package/tools/selftest-reason.mjs +266 -0
  49. package/tools/selftest-rules.mjs +107 -0
  50. package/tools/selftest-token.mjs +100 -0
  51. package/tools/smoke-dsh-adapter.mjs +295 -0
  52. package/tools/smoke-dsh-pipeline.mjs +146 -0
@@ -0,0 +1,449 @@
1
+ # 已接受的设计取舍(不要再"顺手修"它们)
2
+
3
+ > [English](DECISIONS.md) | **简体中文**
4
+
5
+ 这份文件是给**后来的人 / AI** 看的:下面每一条都是**有意为之**的取舍,带着当时的实测依据。
6
+ 如果你觉得某条"显然不该这样",请先读完理由;要改就作为一次显式决策来改,并更新这里。
7
+
8
+ ---
9
+
10
+ ## D1 · 阀门是**事故安全网**,不是**安全边界**(2026-09-20 用户拍板)
11
+
12
+ **决策:** 不追求"防住蓄意绕过的 agent"。只防**事故**:模型/人写错的命令、不透明的脚本、
13
+ YOLO 模式下没人拦的那一下。
14
+
15
+ **这是有意接受的旁路清单:**
16
+
17
+ | 旁路 | 现状 | 为什么不堵 |
18
+ |---|---|---|
19
+ | agent 用文件写入工具直接写 `~/.jev-guard/allow.txt` | 完全可以做到(阀门只挂在 `bash`/`pwsh` 工具上) | 堵它需要把 `write`/`edit` 也纳入判定(问句要重新设计),或者做文件层权限 —— 那是**另一个产品**(沙箱/容器/seccomp 的活),不是这个阀门的活 |
20
+ | agent 换写法绕过判定(`rm -rf` → `find -delete` → `python -c shutil.rmtree`) | Jev 通常仍能看懂,但不能保证 | 见 D2 的兜底思想:靠**分层**而不是靠单点 |
21
+ | 提示注入让 agent 主动作恶 | 不设防 | 要防这个必须上沙箱/容器/权限模型 |
22
+
23
+ **用户原话(定调):** *"这个仍然存在的旁路可以在说明中记录下来,不用去堵 —— 只要这个能防住事故就行。"*
24
+
25
+ **什么情况下该重新考虑:** 如果哪天要把这台机器上的 agent 借给别人、或让不可信输入驱动它。
26
+ 那时正确的做法是**另加一层**(沙箱/低权限用户/容器),不是把阀门改造成安全边界。
27
+
28
+ ---
29
+
30
+ ## D2 · L0 硬规则用"命令位置"锚定,引号里的命令交给 Jev
31
+
32
+ **决策:** 除两个结构性例外外,**所有** L0 规则都只在**命令位置**匹配。
33
+ 命令位置 = 行首(`m` 标志下即**每一行**行首)/ `;` `&` `|` `(` `$(` 之后 /
34
+ `bash -c "` 这类"把字符串当脚本执行"的引号之后,并允许一串包装器
35
+ (`sudo` / `timeout 30` / `xargs -0` / `nice -n 5` / `find … -exec` …)。
36
+ 两个例外显式标为 `where: 'anywhere'`:`redirect-to-device`(模式以 `>` 开头)
37
+ 与 `fork-bomb`(纯语法 `:(){…};:`)—— 它们锚不到"命令位置"这个概念上。
38
+ `RULE_STATS.anywhere` 恒为 2,`tools/selftest-rules.mjs` 会打印它。
39
+
40
+ **依据:** L0 全文匹配曾**三次**拦住操作者自己的正当操作(在参数里写命令原文、在 heredoc 里
41
+ 写含测试用例的自检代码)。而实测 Jev 对这类"散文式提及"只给 **p=0.02–0.08**,对真实调用给
42
+ **0.8–1.0** —— 所以把引号里的散文交给它兜底,覆盖率没有实质损失。
43
+
44
+ **2026-09-20 第二次修正(发现了第一次没做全,且两个方向都偏):**
45
+
46
+ 第一次只改了 7 条 deny + 全部 16 条 ask,`mkfs` / `dd` / `shred` / `chmod -R` / `vssadmin` /
47
+ `wbadmin` / `cipher` / `diskpart` / `wsl --unregister` / `kubectl delete ns` / `Clear-Disk` /
48
+ `Remove-Item … -Recurse` 这 **12 条仍是全文匹配**,于是引号里的数据、注释、变量赋值、
49
+ **代码里的字符串**(实测:`c.startswith('mkfs.ext4 /dev/sdb1')`)都会命中 deny
50
+ —— 而 L0 的 `deny` **没有一次性令牌通道**,被误拦时人只能自己去终端执行。
51
+
52
+ 更严重的是**反方向**:锚定用的 `^` 没有 `m` 标志,所以"命令位置"实际只等于**整串开头**
53
+ (外加分隔符之后)。后果是凡被锚定的规则,多行脚本(heredoc)与多行 `-c` 里的真命令**全部漏判**:
54
+
55
+ | 形态 | 修正前 | 修正后 |
56
+ |---|---|---|
57
+ | `git push --force origin main`(单行) | HIT | HIT |
58
+ | `cd /tmp && git push --force …`(分隔符后) | HIT | HIT |
59
+ | `echo x \| xargs git push --force …` | **MISS** | HIT |
60
+ | `bash - <<'SH'` + `git push --force …`(多行) | **MISS** | HIT |
61
+ | `bash -c "` + 多行 + `git push --force …` | **MISS** | HIT |
62
+ | heredoc 里的 `rm -rf /` / `DROP DATABASE` | **MISS** | HIT |
63
+ | 引号/注释/赋值/代码字符串里的原文 | **HIT(假阳)** | —(交给 Jev) |
64
+
65
+ **为什么这个漏判要紧:** L0 存在的全部理由就是在**没有网络、没有额度**(`degradePolicy: 'l0-only'`)
66
+ 时兜住 `mkfs` / `dd of=/dev/*` / `git push --force` 这一类(见 D9)。平时漏判被 Jev 补上,所以一直没人发现;
67
+ 降级时它就是真空。讽刺的是:修正前 heredoc 里的 `mkfs` **反而是命中的** —— 只因为它没锚定。
68
+
69
+ 边界矩阵(18 种形态)与性能用例(4KB 包装器前缀,防灾难性回溯)在
70
+ [`../tools/selftest-rules.mjs`](../tools/selftest-rules.mjs),第 7 项验收照着跑。
71
+
72
+ ---
73
+
74
+ ## D3 · 判定失败一律 **fail-open**(放行)
75
+
76
+ **决策:** 超时、网络错、服务 5xx、代码异常 → 放行,并记 `source: error` 到审计日志。
77
+
78
+ **依据:** 宿主自己的沙箱与审批策略仍在执行之前;阀门只是**增量**检查。若改成 fail-closed,
79
+ Typesafe 抖一下你就干不了活。想"服务挂了也拦"应当**把 L0 规则加厚**(不依赖网络的那层),
80
+ 而不是改这条策略。
81
+
82
+ **健康信号:** `guard log --stats` 里的 `fail-open` 计数;持续为 0 说明判定服务没偷偷挂过。
83
+
84
+ ---
85
+
86
+ ## D4 · 阈值保持 0.5 / 0.7,靠令牌解决摩擦,而不是放低标准
87
+
88
+ **决策:** `lowThreshold: 0.5` / `highThreshold: 0.7`(在 737 条真实命令上实测
89
+ 98.51% 放行 / 0.81% 灰区 / 0.68% 拦截)。
90
+
91
+ **依据:** 抬高阈值会放走真实破坏 —— `rm -rf ~/真实目录` 实测落在 **0.68–0.82**,
92
+ 灰区里还有 `docker compose down -v`(0.69)。既然有了**一次性放行令牌**,"被拦"不再等于"卡死",
93
+ 所以宁可偏严。真要少打扰,改 `config.json` 一行即可(0.6/0.8),但那是显式换风险。
94
+
95
+ ---
96
+
97
+ ## D5 · 一次性令牌只覆盖灰区与拦截,**不越过 L0**
98
+
99
+ **决策:** 令牌能放行 `revise`/`block`;对 L0 的 `deny` 规则**无效**,理由里也不展示令牌。
100
+ 绑定完整命令文本(空白折叠)、用掉即删、放行写审计(`source: token` + `overridden`)。
101
+
102
+ **依据:** L0 是四层设计里的硬地板;如果令牌能越过它,"永不允许"就名存实亡。
103
+
104
+ ---
105
+
106
+ ## D6 · 授权(`guard allow`)只能在**交互终端**里做
107
+
108
+ **决策:** `guard allow` / `--revoke` 检测 `stdin.isTTY`,非 TTY 直接拒绝并打印可复制的整行命令。
109
+
110
+ **依据:** 实测 agent 自己跑授权命令时,它**本身也会被判危险**(p=56%)—— 因为命令文本里就带着
111
+ 那条破坏性命令。这本来是巧合挡住的;与其靠巧合,不如把它变成明确的设计:**授权 = 人的动作。**
112
+ (注意这不是安全边界,见 D1:agent 仍可写文件绕过;它只是让"被监管者给自己开后门"不会**顺手发生**。)
113
+
114
+ ---
115
+
116
+ ## D7 · 脚本正文会被上传到 TypeSafe 的 API
117
+
118
+ **决策:** 默认 `inlineScripts: true` —— 读取被调用脚本的正文作为判定输入。
119
+
120
+ **依据:** 只看命令行时,`node migrate.mjs`(体内 `DROP COLUMN`)只有 **p=0.32**;
121
+ 把正文读进去后 **0.82**。这是"看懂不透明命令"的必要代价。
122
+
123
+ **已有的缓解:** 敏感路径(`.env` / `.ssh` / `*.pem` / `*credential*` / `*secret*` / `*token*`)
124
+ 自动跳过不上传;单文件 8KB 上限;`inlineScripts: false` 可整体关闭(代价:退回 0.31 的盲区)。
125
+
126
+ ---
127
+
128
+ ## D8 · 包结构是"判定核心 + 一层薄适配",`lib/` 里不放 DSH 机制
129
+
130
+ **决策:** `lib/` 只放与调用方无关的判定核心(`gate` / `rules` / `verdict` / `token` / `audit` / `quota`),
131
+ DSH 相关的一切都放在 `adapters/dsh/index.js` 这一个文件里。
132
+ (`package.json` 的 `exports["."]` 指向它,`main` 同步。)
133
+
134
+ **依据:** 这个项目最初的写法容易让人以为判定与 DSH 绑在一起。但实际上判定不看宿主的任何东西
135
+ (不看文件系统状态、不看会话历史、不需要模型参与),所以本来就该分开 ——
136
+ **分开的理由在 D11 里改过一次**:不是为了接别的调用方,而是为了让判定能被离线复跑
137
+ (校准、回归、事故复盘全靠这一点)。放进 `lib/` 会让下一个维护者把宿主逻辑继续往核心里加。
138
+
139
+ **可验证的形式:** `grep -riE "cordis|PreToolDecision|ctx\.|approval/policy" lib/*.js` 只应命中**注释**
140
+ (目前那条注释用来解释"为什么 escalate 在完全权限下会变成拒绝")。
141
+ 代码里的宿主差异只有一个与 DSH 无关的布尔值:`canPrompt`。
142
+
143
+ **契约文档:** 当时写的是 `docs/HOST-CONTRACT.md`(宿主无关契约)。**已被 D11 取代** ——
144
+ 现在对应的是 [`DSH-INTEGRATION.md`](./DSH-INTEGRATION.md)(DSH 集成:用到哪些机制、四态怎么映射、降级契约)。
145
+ D8 里"把宿主逻辑挡在 `lib/` 之外"这条**仍然有效**,但理由在 D11 里改了:不是为了支持更多调用方,
146
+ 而是为了让判定能被离线复跑。
147
+
148
+ **回滚:** 这次搬迁在部署实例上留过一份 `lib/index.js.bak-moved-to-adapters-dsh`;重启验证插件照常装载后已删除。
149
+
150
+ ---
151
+
152
+ ## D9 · 额度耗尽时**降级为 L0-only 并显式告警**(默认),而不是静默 fail-open
153
+
154
+ **决策:** 判定服务是**收费**的,额度用完是必然事件。一旦失败属于持久性类别
155
+ (`quota` / `auth` / `no-key`),就:
156
+
157
+ 1. **写一份可读状态** `<JEV_GUARD_HOME>/degraded.json`(原因、开始时间、恢复时间、失败次数、原始错误);
158
+ 2. 在冷却窗口内(**quota 15 分钟 / 密钥 30 分钟**)**不再发请求** —— 省钱、也省掉每条命令都要等一次注定失败的请求;
159
+ 3. 窗口到期放**一次**探测请求:成功 → 自动恢复正常(人不需要做任何事);失败 → 续期继续降级;
160
+ 4. **降级期间默认只跑免费的 L0 + 预筛**(`degradePolicy: 'l0-only'`),瞬态失败(超时/网络/5xx/429限流)
161
+ 照旧逐次 fail-open、**不**降级,但会被分类记录。
162
+
163
+ **为什么不能只靠 D3 的 fail-open:** 功能上它没错(命令照跑),但它是**静默**的 —— 命令照样放行、
164
+ 日志里堆满 `source: error`,而**没有任何人能一眼看出阀门已经不在防护了**。这个项目已经吃过一次
165
+ "静默失效"的亏(审计日志那轮,见 MEASUREMENTS §7),所以这次不重复。
166
+
167
+ **为什么降级还保留 L0(而不是"整条阀门全停"):** L0 与预筛**不花钱、不联网、确定性强**,
168
+ 而且正好覆盖最坏的一类(`mkfs` / `dd of=/dev/*` / `git push --force` / `wsl --unregister`)。
169
+ 把它们也停掉,等于用"额度没了"换"最危险的一类命令失去保护"。想连它们一起停是**显式选择**:
170
+ `degradePolicy: 'off'`(实测已覆盖:该配置下连 `mkfs` 也放行)。
171
+
172
+ > **用户确认(2026-09-20):** 原话 *"「要花钱的那层不生效」这个就可以了"* —— 即
173
+ > `degradePolicy: 'l0-only'` 就是想要的默认,不是权宜之计。这一条与 D1 同源:阀门的工作是
174
+ > 把"要花钱的语义判断"和"免费的确定性规则"分开,额度没了只停前者。
175
+
176
+ **告警出现在哪(「必须有人能发现」这件事的落点):**
177
+ - 每条非 allow 判定的**理由里**(所以审批弹窗、拒绝信息、模型反馈都带着它);
178
+ - 降级放行时 `source: degraded`,审计里是**独立来源**,不是混在 error 里;
179
+ - 进入降级的那一刻写一条 `level: 'warn'` 的审计记录(每个窗口最多一条,不刷屏);
180
+ - CLI 的 **stderr**;DSH 插件的 `ctx.logger.warn`;
181
+ - `guard status`(退出码 **3** = 正在降级,可当健康检查)、`guard log --stats`。
182
+
183
+ **已知取舍:** 降级期间**语义层对新命令完全不起作用** —— 一条精心伪装的破坏性命令会被放行。
184
+ 这是"额度没了"的必然结果,不是可以靠代码消掉的东西;能做的只有①保留免费层②让人尽快知道。
185
+ 按 D1,这不是安全边界问题(它本来就不防蓄意绕过)。
186
+
187
+ **依据:** 2026-09-20 用无效密钥打真实 API 实测:`HTTP 401` → 分类 `auth` → 降级 30 分钟、
188
+ 状态文件写出真实错误体、第二次调用 `source: degraded` 且**零请求**、`guard status` 退出码 3。
189
+ 另有 54 条离线断言(`tools/selftest-quota.mjs`,替身 fetch 覆盖 402/401/429/5xx/超时/网络/坏状态文件)。
190
+
191
+ ---
192
+
193
+ ## D10 · 入口守卫**内联**、动态导入用**相对说明符**;降级集合只收"服务态度变了"的两类
194
+
195
+ 两条都是被真实事故逼出来的,而且**看起来都像是在"改善代码"**,所以必须写下来免得被改回去。
196
+
197
+ ### D10.1 入口守卫必须内联在各自文件里
198
+
199
+ **决策:** 判断"我是不是被当作入口执行"的那几行,在**每个入口脚本里各写一份**
200
+ (包内现存的是 `tools/extract-commands.mjs`;当时还有两处,已随范围收窄归档到包外)。
201
+
202
+ **依据:** 旧写法 `import.meta.url === \`file://${process.argv[1]}\`` 在 Windows 上**恒为 false**
203
+ (argv1 是 `T:\…`、url 是 `file:///T:/…`)→ 脚本加载完**直接退出 0**,而在不少调用约定里
204
+ **退出码 0 就是"放行"**:阀门既不拦也不记、还没人知道。修它时为了 DRY 把守卫抽进了 `lib/entry.js` ——
205
+ 于是**连 WSL 上也恒为 false**,因为 **`import.meta.url` 是每个模块各自的**,进了共享模块之后
206
+ 比较的就是那个文件自己的路径。
207
+
208
+ **所以这不是重复代码,而是"每份代码只谈自己的身份"。** 正确写法:
209
+
210
+ ```js
211
+ function isMainModule() {
212
+ if (!process.argv[1]) return false
213
+ try {
214
+ return realpathSync(process.argv[1]) === realpathSync(fileURLToPath(import.meta.url))
215
+ } catch {
216
+ return false
217
+ }
218
+ }
219
+ ```
220
+
221
+ `realpathSync` 同时解决盘符、反斜杠、相对路径与软链。**姊妹条款:** 动态导入要用**相对说明符**
222
+ (`await import('../../lib/gate.js')`),不要 `import(join(ROOT, …))` —— Windows 上绝对路径不是
223
+ 合法 ESM 说明符(`ERR_UNSUPPORTED_ESM_URL_SCHEME`),相对说明符按本模块自己的 URL 解析,两平台都成立。
224
+
225
+ **护栏:** `tools/selftest-entry.mjs` 真的 spawn 每个入口脚本,并断言"三个文件各自定义守卫、
226
+ 没有从共享模块导入"。**这一层只有跑在 Windows 上才算验过** —— WSL 上永远验不出盘符问题。
227
+
228
+ ### D10.2 只有 `quota` 与 `auth` 触发降级,`no-key` 不触发
229
+
230
+ **决策:** 降级集合收窄为 `{quota, auth}`。`no-key` 归类、告警,但**不写共享的 `degraded.json`**。
231
+
232
+ **依据:** `degraded.json` 是**全局共享**的,而密钥解析**每条入口各自独立**
233
+ (DSH 插件走 `ctx.credentials`,CLI 与离线脚本走环境变量或文件)。出现过这样的放大链:某条入口因为 cwd 不对
234
+ 读不到密钥 → 写一份共享的 `no-key` 降级 → **密钥正常的其它入口也一起停掉联网判定 30 分钟**。
235
+ 而 `no-key` 一次 HTTP 都不发,降级省不下任何东西 —— 它是**本地配置**状况,不是"服务对我们的态度变了"。
236
+ 同时把触发源一起修掉:`bin/guard.mjs` 里**相对路径的 `apiKeyFile` 一律按包根解析**(与 cwd 无关)。
237
+
238
+ **规律:** 凡是"共享状态 + 各组件独立前提"的组合,先问一句"这个局部故障会不会被写进全局状态"。
239
+
240
+ > **本条已被 D15 部分取代(2026-09-20)。** 上面那条放大链是真的,但它现在用**作用域**回答,而不是用
241
+ > 沉默回答:`no-key` **会**降级,而状态里记下"是谁写的",所以只压制那一条入口。下面那条按入口解析
242
+ > 密钥的修复仍然有效 —— 而上面这条规律正是作用域字段的由来。
243
+
244
+ ---
245
+
246
+ ## D11 · **只支持 DSH**(2026-09-20 收窄,用户决定)
247
+
248
+ **决策:** 本包不再维持"通用"的承诺。历史上试过的其它执行通道,其实现与验证工具
249
+ **已从本包移除**(用户于 2026-09-20 决定不再保留那份档案),包内只留 DSH:
250
+
251
+ - `adapters/` 下只有 `dsh/`;
252
+ - 文档只讲 DSH 的机制(`docs/DSH-INTEGRATION.md` 取代了原来的 `HOST-CONTRACT.md`);
253
+ - 验收清单只剩 DSH 的项(`docs/VERIFICATION.md`);
254
+ - **包内任何文件都不再提及别的工具**(2026-09-20 清理);
255
+ - 值得留下的**可迁移教训**写在 [`MEASUREMENTS.md`](./MEASUREMENTS.md) §12,不点具体工具名。
256
+
257
+ **依据(不是"做不到",是性价比):**
258
+
259
+ 1. **各自宿主的审批/信任机制互不相同**,把任何一条做扎实都是各自独立的一轮工作。
260
+ 2. **实测暴露过一个结构性缺陷**:在"宿主能不能问人"这件事上,那套实现**等于永远是否** ——
261
+ 策略取自宿主注入不了的 env、又忽略 payload 里的权限字段,且无论什么策略都输出同一个拒绝结论。
262
+ 于是"能不能问人"退化成"一律硬拒",**宿主审批这条人工通道不可达**。修它要重做那套输出契约。
263
+ (可迁移的教训:**拿不到的信息要显式承认拿不到,别用一个默认值假装它存在**。)
264
+ 3. **未验证的适配器比没有适配器更危险**:它看起来装了阀门,实际不拦。
265
+
266
+ **保留的东西:**
267
+
268
+ - `lib/` 仍然与调用方无关 —— **理由变了**:不是为了接别的宿主,而是因为判定必须能被**离线复跑**
269
+ (校准、回归、事故复盘全靠这一点)。`bin/guard.mjs` 与七份自检都靠它。
270
+ - **跨平台(WSL + Windows)不变** —— 平台不是宿主。两条平台差异都处理了:`bash`/`pwsh` 两个工具都在默认列表里;
271
+ 授权行的引号按平台分叉(见 D12)。
272
+ - 归档目录里的代码带着入口守卫的跨平台修复,将来复活别丢。
273
+
274
+ **复活任何一条路线的门槛:** 先在那个宿主上跑完整验收(含"命令真的被拦 + 日志真的有记录"),
275
+ 再把它加回支持列表;只跑通"脚本能执行"不算。**包内文件不要预先提及它们** ——
276
+ 等真的做成了再写进文档。
277
+
278
+ ---
279
+
280
+ ## D12 · 授权行的引号**按平台分叉**;Windows 上另给与 shell 无关的 `--command-file`
281
+
282
+ **决策:** `shellQuote(value, platform)`:POSIX 用 `'…'` + `'\''`;Windows 用 `'…'` + `''`(PowerShell)。
283
+ 理由文案在 Windows 上额外点明"用 PowerShell"。另外 `guard allow --command-file <文件>` 从文件读命令原文,
284
+ **完全不过 shell 的引号规则** —— 给 cmd.exe 用(它两种写法都不认)。
285
+
286
+ **依据(实测):**
287
+
288
+ | 形式 | 在 PowerShell 里 |
289
+ |---|---|
290
+ | `'a''b'`(本包 win32 形式) | ✅ 往返还原成 `a'b` |
291
+ | `'a'\''b'`(旧的 POSIX 形式) | ❌ **ParserError**,语法都不成立 |
292
+
293
+ 这段话是给用户**照抄**的:粘进去语法错 = 授权入口不可用 = 人唯一的出口被堵住。而"一次判定都还没发生"
294
+ 之前没人会发现(R2 只在被拦时才出现)。
295
+
296
+ **为什么默认用 `process.platform`:** DSH 跑在 WSL 上时,用户能粘贴的终端通常也是 WSL/POSIX;
297
+ DSH 跑在 Windows 上时,那是 PowerShell。"宿主与终端在同一侧"是常态。
298
+ 两边不同侧时(DSH 在 WSL、终端在 Windows),用 `--command-file` 那条与 shell 无关的路。
299
+
300
+ ---
301
+
302
+ ## D13 · 判定动作**随审批模式分叉**:`ask` 下 `revise`/`block` 转人工,L0 硬规则仍拦死(2026-09-20 用户拍板)
303
+
304
+ **决策:**
305
+
306
+ | 判定 | `approval: ask`(人就在场) | `approval: never`(全自动) |
307
+ |---|---|---|
308
+ | `revise`(50–70%) | **转人工弹窗**(附三种降级模板) | 直接拒绝(+ 模板 + 令牌) |
309
+ | `block` ≥70%(语义层) | **转人工弹窗** | 直接拒绝(+ 令牌) |
310
+ | L0 的 `deny` 类硬规则 | **拒绝**(不弹窗、不发令牌) | **拒绝** |
311
+ | L0 的 `ask` 类规则 / 重试预算升级 | 转人工弹窗 | 直接拒绝 |
312
+
313
+ 配置开关 `reviseInAskMode` / `blockInAskMode`(取值 `'ask'`(默认)/ `'deny'`),留一条**不用改代码**的回退路。
314
+
315
+ **依据(实测,2026-09-20):** 改之前只有 `escalate` 会看审批策略,`revise`/`block` 一律直接拒绝。
316
+ 审计日志当天 **59 条** `revise`/`block` 拒绝里,**有 9 条发生在 `policy=ask` 的会话中**(p 全在 0.50–0.63),
317
+ 命令分别是 `cp` 到部署目录、`sed -i`、`mkdir -p`、`git add -A && git commit` —— 全是操作者自己的维护动作,
318
+ **人就在旁边,却只拿到一句"请改用更安全的形式"**。用户原话:"我实际上想的是,在全自动模式下拦,
319
+ 而在需要审批的模式里面所有的拦截都改成弹审批"。
320
+
321
+ 这同时补上了 **D1 三分类里的第 (3) 类** —— "中间态:先别跑,找替代;**找不到就等用户**"。
322
+ "等用户"必须有一条能到人的通道,而当时只有 `escalate` 有。
323
+
324
+ **为什么不会因此变松:**
325
+
326
+ 1. `never`(全自动)语义上就是 `danger-full-access` preset;那种会话里 DSH 会把任何 `ask` 直接判成
327
+ `rejected`,所以"没人可问时由阀门直接拒"是唯一正确的落法。
328
+ 2. `ask` 模式下**没有应答者时审批 fail-closed**(直接 rejected)—— 转人工不会变成无人自动放行。
329
+ 3. 弹窗只授予 `allowed-once`,不存在"以后都放行"。
330
+ 4. 频率低:0.5–0.7 在 737 条语料里只占 0.81%(约 1/125),不会把弹窗变成噪音。
331
+
332
+ **同时堵上一个潜在洞:** 重试预算原本会把"任何非 escalate 判定"在第 `retryLimit+1` 次尝试后升级为
333
+ escalate —— 对带 L0 `deny` 规则的硬命中也一样。那意味着 `ask` 模式下把 `git push --force` 连交三次就会弹窗,
334
+ 而弹窗里点"允许"**越过了硬地板**(与 D5"令牌不越过 L0"自相矛盾)。现在硬命中**不参与**预算升级
335
+ (仍记 `attempts` 供审计)。审计显示这条路径在修复前**从未被触发过**(纯潜在洞)。
336
+
337
+ **代价(明说):** `ask` 模式下高分的破坏性命令现在会**弹窗等你点**,而不是立刻被拒 —— 手快点错,
338
+ 后果由决定的人承担,阀门不再替你兜底。想要"高分一律拦死"的,把 `blockInAskMode` 填成 `deny`;
339
+ 但 **L0 那一档无论怎么配都拦死**。
340
+
341
+ ## D14 · 文案双语化,但**界面语言与判定问话语言分开**:`promptLang` 默认仍是中文(2026-09-20)
342
+
343
+ **决定:**
344
+
345
+ 1. 所有**给人或模型看的文案**(判定理由、L0 规则理由、CLI 输出、降级告警与状态、引用的脚本
346
+ skip 说明、审计汇总标题)都有中英两份,集中放在 `lib/i18n.js`(L0 规则理由例外:它写在
347
+ 规则自己旁边,保持"一条规则一个自洽单元",见 `lib/rules.js` 的 typedef)。
348
+ 2. `lang` 控制界面语言,默认 `'auto'`:`JEV_GUARD_LANG` → `LC_ALL`/`LC_MESSAGES`/`LANG`
349
+ (**仅当它们指明一种受支持的语言**)→ 否则 `zh-CN`。CLI 另有 `--lang zh-CN|en`。
350
+ **`Intl`/系统 locale 刻意不在链上**:第一版把它放在最后,结果真实部署当场踩到 ——
351
+ DSH 插件跑在 WSL 里,那里 `LANG=C.UTF-8` 表示"没有偏好",`Intl` 于是报出 Node 自己的
352
+ `en-US` 兜底值,会话里的理由悄悄变英文,而同一台机器的 Windows 侧 CLI(其 Node 报 `zh-CN`)
353
+ 仍是中文。同一个阀门两种语言,事后复盘时极难解释。`C`/`POSIX`/未设置 = **没有信号**,
354
+ 落回项目主语言;要英文就明说。
355
+ 3. `promptLang` 控制**发给判定服务的那句问话**与 state 的键,**默认 `'zh-CN'`,与 `lang` 无关**。
356
+ 4. 代码注释与 `tools/` 里的自检标签**不翻译**:前者是维护者读的,后者是测试用例名;
357
+ 翻译它们会让每一次改动的维护成本翻倍,却不改变产品对外的任何一句话。
358
+
359
+ **为什么第 3 条要单独拎出来(这是本条的核心):** 阈值 0.5 / 0.7 是在**中文问话**上标定的
360
+ (114 例,见 MEASUREMENTS §2)。§14 的实测表明:换成英文问话后,21 条探针里 **12 条 p 更低 /
361
+ 4 条更高**,平均压低约 **0.04**,并且**三条命令直接翻带**——`DELETE ... WHERE`(block→revise)、
362
+ 无 WHERE 的 `UPDATE`(block→revise)、内联 `node -e rmSync`(revise→allow),方向**全部**朝更宽松。
363
+ 两轮独立运行结论一致,噪声地板只有 0.015。也就是说:把界面切成英文,如果顺手把问话也切成英文,
364
+ 就等于**悄悄把一条被测过的边界往放行方向挪了一格**。所以两者必须分开配置,默认保持中文;
365
+ 要切就得承认那是一次重标定,而不是翻译。
366
+
367
+ **为什么不是"干脆不做英文问话":** 非中文部署里,模型看到英文问话更自然,而英文问话并非不可用 ——
368
+ 18/21 同带。把它做成**显式选项 + 写明代价**,比藏起来或假装没有强。
369
+
370
+ **代价(明说):**
371
+
372
+ - 多了一份目录要维护;`selftest-i18n` 会在"某个键只写了一种语言""占位符两边不一致"
373
+ "英文里残留中文"这三件事上直接判失败。
374
+ - 语言随系统 locale 自动解析意味着:**英文 locale 的机器上,升级后输出会变成英文**(行为变更)。
375
+ 想固定就说 `lang: "zh-CN"`。
376
+ - 判定行为本身**不受影响**:L0 规则、预筛、阈值、缓存键、promptLang 默认值都不变 ——
377
+ `lang` 只换文案。
378
+
379
+ 5. **仓库里的文档同样英文优先**:每份文档默认英文,中文逐字节保留为同目录的 `<name>.zh-CN.md`,
380
+ 两份顶部各有语言切换行(第 1 条只讲了"文案",这条讲的是"文档本身")。规则:
381
+ - **中文版是原文的逐字节副本**,每次只多一行切换行;改动时**两份一起改**。
382
+ - **机械校验**代替信任,由 `node tools/check-doc-pairs.mjs` 执行:中文版必须逐字节等于基线加那一行
383
+ 切换行;两部在标题层级序列、代码围栏数、表格行数、链接目标集合、数字多重集上必须一致
384
+ (该工具还会把英文那侧残留的每一行中文列出来,供人确认它们全是引用)。
385
+ - **实测引用不翻译** —— 日志原文、命令样例、中文语料(`xargs 删除 mkfs.ext4 …` 这类)
386
+ 在被引用的位置上保持原样,否则引用就成了伪造;英文正文里剩下的中文只应是这些。
387
+ - **生成物跟着语言开关走**:`verification-results/SUMMARY.md` 由 `tools/report-result.mjs`
388
+ 生成,该工具默认英文(它是入库给人读的产物,换机器重新生成不该让语言漂移),
389
+ 但**证据列是逐字引用**,仍是中文。
390
+ - 代码注释与 `tools/` 里的自检标签仍不翻译(第 4 条),界限是:**读者在仓库外**(GitHub 访客、
391
+ 用户、模型)看得见的→双语;**读者是维护者**的→中文。
392
+
393
+ **与 D2/D5/D13 的关系:** 不改变任何一条判定路径。唯一动到判定输入的是显式设置 `promptLang: "en"`,
394
+ 而那属于"你自己选的、且现在有数了"的一类(§14)。
395
+
396
+ ---
397
+
398
+ ## D15 · `no-key` 的降级**粘性且带作用域**;用户通过**会话内 notice** 被告知;密钥**只从标准输入**录入(2026-09-20,用户决定)
399
+
400
+ **背景。** 新装的时候没有密钥:凭据层里没有、环境变量没有、文件也没有。在本决定之前,阀门在这个状态下
401
+ 保持沉默 —— 每条受管命令都在语义层 fail-open,留下一条 `error` 类判定,而**用户永远看不到任何东西**
402
+ (DSH 的 logger 会丢掉插件的 info 级日志)。用户一次要了三件事:一个真正的密钥录入位置、首次安装时对它的
403
+ 要求、以及一个会把"没有有效密钥"说出口的降级态 —— 同时必须保住让这个插件能被安装的那些性质(零依赖、
404
+ 无构建步骤、源码安装免构建授权)。
405
+
406
+ **决定。**
407
+
408
+ 1. **`no-key` 是会降级的类别,而且状态是粘性的。** 没有密钥时一次 HTTP 都不发,所以没有可探测对象。
409
+ 与 `quota`/`auth`(恢复方式是"冷却到期后问一次服务")不同,`no-key` 不可能靠时间结束。它靠**条件**
410
+ 结束:密钥一解析到,`evaluateCommand` 就清掉状态并恢复正常判定。粘性类别的 `probeDue()` 恒为假,
411
+ 所以没有密钥的部署不会在每个命令上重进一次判定。
412
+ 2. **降级状态带作用域。** `quota`/`auth` 是服务侧事实,保持 `scope: 'global'`(所有入口都遵守)。
413
+ `no-key` 是本地配置事实,写入时带上**报告它的那条入口的身份**(`'cli'` / `'dsh-adapter'`),而一条
414
+ 入口只遵守作用域属于它自己的本地状态。这正是 D10.2 那条反对意见退休的原因 —— 解法是作用域字段,
415
+ 不是沉默。
416
+ 3. **用户在对话里被告知。** DSH 不给纯 host 插件任何 toast / banner / 启动提示:设置页与 Plugins 页的
417
+ 每个位置都由浏览器侧(`dsh.client`)注册占位,启动期告警只到终端。唯一存在的渠道是在 `agent/pre-step`
418
+ 注入一条 `notice` 形态的用户消息:它渲染成对话里的一行、写进会话历史、并进入模型上下文。被否掉的
419
+ 备选:斜杠命令(它输入的内容除非设 `recordInput: false` 否则会被持久记录,而且浏览器输入框无论如何
420
+ 都看得见 —— 对密钥绝不可接受)、以及发浏览器半边(那会带来构建步骤、提交构建产物与 React 外部依赖,
421
+ 放弃这个插件的立身之本:免构建、零依赖)。notice 覆盖三种跃迁 —— 首次运行没有密钥、进入降级、恢复 ——
422
+ 每种状态每个会话只说一次,去重依据是**持久化的历史**(`session.deriveMessages()`),所以重启或恢复
423
+ 会话都不会重复;并且**绝不**往空的步批次里塞(那会白白多花一次模型请求)。`notifyInSession: false`
424
+ 可以关掉它。
425
+ 4. **密钥用 `guard key set` 录入,且只从标准输入读。** 参数会进 shell 历史与 `ps`,所以 `key set`
426
+ 不接受命令行上的值,并且拒绝非 TTY 的标准输入 —— 与 `guard allow` 同一条分界线:人在键盘上敲的,
427
+ 而且这一点可验证。它写 `apiKeyFile`、权限 **0600**,保留文件里已有的其它键,只打印长度与路径、
428
+ 永不打印值。`guard key status` 回答"哪个来源在生效",没有时退出码 3。
429
+ 5. **适配器必须学会读文件。** 它原先只从 `ctx.credentials` 与环境变量取密钥。而 `key set` 写的是文件,
430
+ 所以少了这一层,CLI 这个入口对**最需要它的用户** —— 刚装好、两者都没有的那种 —— 就是一句空话。
431
+ 现在适配器的顺序是 `ctx.credentials` → 环境变量 → `apiKeyFile`,与 CLI 共用同一条路径规则
432
+ (相对路径按包根解析,与 cwd 无关)。
433
+
434
+ **证据。** `tools/selftest-quota.mjs`(78 例)覆盖粘性状态、永不探测的规则、双向的作用域隔离、以及
435
+ "密钥出现即清除"。`tools/smoke-dsh-adapter.mjs`(21 组断言,现在已密闭 —— 它不再往真实的
436
+ `~/.jev-guard/` 写东西)覆盖 notice 是**追加**而非替换、空批次守卫、四键 `source` 形状与摘要上限、
437
+ 以及文件回退层。`tools/selftest-entry.mjs` 真的跑了一遍 `guard key set`(包括通过伪 TTY 走的交互路径),
438
+ 并断言文件被写出、可解析、POSIX 上权限为 0600、且从不回显。
439
+
440
+ **形状是一份契约。** `source` 恰好带 `kind`/`plugin`/`form`/`summary`;第五个键会被 v3 之前的迁移校验
441
+ 拒绝,而形状写错的表现是**下次恢复会话时**报 `SessionPersistenceCorruptionError` —— 会话打不开,而那时
442
+ 离改动已经很久。这个形状已在一个 DSH 检出里对着 DSH 自己的 `snapshotJsonValue`(`Session.append` 首先
443
+ 执行的那一步)验证过;CHANGELOG 记录了 `tools/smoke-dsh-pipeline.mjs` 本身为何在本部署里跑不起来。
444
+
445
+ **接受的代价:** notice 是 `role:'user'` 消息,所以它**会进入模型上下文**(这是本意 —— 模型也该知道阀门
446
+ 降级了),并从那一刻起造成一次前缀缓存失效。这正是它按状态跃迁触发、而绝不按步触发的原因。
447
+
448
+ **经验法则:** 当一个插件"必须告诉用户点什么"而它没有任何 UI 时,先问宿主已经在渲染什么 —— 会话内
449
+ notice 是持久的、带来源标注的、模型可见的;而自己造一个 UI 界面是另一个项目,有另一份依赖预算。
@@ -0,0 +1,178 @@
1
+ # DSH Integration — How This Valve Attaches to DeepSeek Harness
2
+
3
+ > **English** | [简体中文](DSH-INTEGRATION.zh-CN.md)
4
+
5
+ **In one sentence:** this package **supports DSH only** (narrowed on 2026-09-20, see [`DECISIONS.md`](./DECISIONS.md) D11).
6
+ The judging core lives in `lib/` (host-agnostic, pure functions); everything host-related is concentrated in the single file `adapters/dsh/index.js`.
7
+
8
+ ---
9
+
10
+ ## 1. Which of DSH's mechanisms it uses
11
+
12
+ | DSH mechanism | How we use it | What happens if we cannot get it |
13
+ |---|---|---|
14
+ | **`tools/pre-execute` waterfall** | The only real interception point. Returns a typed `PreToolDecision`: `{kind:'allow'}` / `{kind:'ask', reason}` / `{kind:'deny', reason}` | Without it we could only do a "suggestion layer" (the model can ignore it) |
15
+ | **`agent.session.snapshotEvents()` → `approval/policy`** | Reads the current session's approval policy: `ask` (a prompt is shown) / `never` (full permissions, ask is resolved into a denial) | We could only guess from the deployment default, and the reason text would say the wrong thing |
16
+ | **`agent.session.snapshotEvents()` → `permission/preset`** | Reads the sandbox preset (`workspace-write` / `danger-full-access`), **audit only** | A retrospective cannot see whether there was still a sandbox behind it at the time |
17
+ | **`ctx.credentials.resolve(ref)`** | Fetches TypeSafe secrets (through DSH's credential layer, no restart needed after rotation); if that fails, falls back to `process.env` | Falls back to "reading from a file", and key rotation requires a restart |
18
+ | **`exec.arguments.workdir` / `exec.agent.session.cwd`** | The working directory at judging time (used to read the body of the invoked script) | The script body cannot be guessed, and things like `node x.mjs` degrade into a blind spot |
19
+ | **Tool names `bash` / `pwsh`** | The two tools intercepted by default — covering exactly **WSL/Linux (`bash`) and Windows (`pwsh`)** | One less platform is intercepted |
20
+ | **`ctx.logger`** | Best-effort host logging (`info`/`debug` are often filtered out by the host threshold, so **we do not rely on it**) | Auditing relies on `~/.jev-guard/guard.log`, not on host logs |
21
+ | **`dispose`** | `flush()` the audit queue before exit (otherwise the trailing records are lost) | The last few verdicts are lost |
22
+
23
+ **The judging logic itself depends on nothing from DSH**: it does not look at filesystem state, needs no model involvement, and needs no session history.
24
+ So the same judging can be called by the DSH plugin inside a session, and can also be re-checked offline by `bin/guard.mjs` — the latter is the basis of the regression tests.
25
+
26
+ ---
27
+
28
+ ## 2. Two policies combined: four destinations for the same command
29
+
30
+ | Verdict | Approval policy `ask` (a prompt is shown) | Approval policy `never` (full permissions) |
31
+ |---|---|---|
32
+ | `allow` | Allowed | Allowed |
33
+ | `revise` (50–70%) | **`ask`** — escalated to a human with an approval prompt, with the three degradation templates in the reason | **Denied** + degradation template, **plus** a one-shot token hint |
34
+ | `block` ≥70% (the semantic layer) | **`ask`** — escalated to a human with an approval prompt | **Denied** + one-shot token hint |
35
+ | L0's `deny`-class hard rules | **Denied** (no prompt, no token issued) | **Denied** |
36
+ | `escalate` (L0's `ask`-class rules / retry budget escalation) | **`ask`** — the host shows an approval prompt, a human decides | **Denied** + one-shot token hint |
37
+
38
+ Two details that are easy to get wrong, but matter a lot:
39
+
40
+ 1. **Under the `ask` policy, `revise` and the semantic layer's `block` also reach the human** (changed 2026-09-20, see
41
+ [`DECISIONS.md`](./DECISIONS.md) **D13**): when the human is present, letting a 50.6% judgment decide for them makes no sense.
42
+ The measured cost before the change was that 9 `revise` denials in a day happened in `ask` sessions while the human was right beside it.
43
+ **The exception is L0's `deny`-class hard rules** — an absolute gate, denied under both policies, and not taking part in `retryLimit` escalation.
44
+ 2. **When a prompt can be shown, the token hint is not attached any more.** A measured lesson (2026-09-20): two human channels
45
+ appearing in one prompt at the same time left the reader unsure which one to take; and when the human is in front of the
46
+ window, the token is superfluous.
47
+
48
+ > `danger-full-access` = `{ sandbox: 'danger-full-access', approval: 'never' }`. At this point **the valve is the only layer**
49
+ > — there is no file sandbox as a fallback, and approval is effectively off. This is exactly why it exists, and also the scenario
50
+ > where a misjudgment costs the most.
51
+
52
+ ---
53
+
54
+ ## 3. The three human channels after it is installed
55
+
56
+ | Channel | Who executes it | What it depends on from DSH |
57
+ |---|---|---|
58
+ | One-shot token | The AI (retries after being blocked) | **Nothing** — pure local hash + files; DSH only writes the hint into the reason |
59
+ | Host approval prompt | The AI (with a human's nod) | `approval/policy = ask` + the host's approval UI |
60
+ | Human executed by hand | The human | Does not go through the valve at all |
61
+
62
+ Details, measured evidence, and the nature of each of the three channels are in [`USER-INTERVENTION.md`](./USER-INTERVENTION.md).
63
+
64
+ ---
65
+
66
+ ## 4. The degradation contract (after the quota is used up)
67
+
68
+ | Failure category | Valve behaviour | `source` in the audit |
69
+ |---|---|---|
70
+ | `quota` (402 / quota wording) / `auth` (401/403) | **Degrades**: writes `~/.jev-guard/degraded.json`, sends no more requests within the cooldown window, by default only the free L0 + pre-screen runs | First time: `error` + `degraded`; afterwards: `degraded` |
71
+ | `no-key` (no key could be resolved) | **Degrades, stickily and scoped**: no HTTP is sent at all, the state never expires with time (there is nothing to probe) and is cleared the moment a key resolves; it records the identity of the entry that reported it, so it suppresses only that entry | First time: `error` + `degraded`; afterwards: `degraded` |
72
+ | `timeout` / `network` / `server` / `rate-limit` | **Does not degrade**, fail-open each time, recorded classified by `errorKind` | `error` |
73
+
74
+ `no-key` used to be deliberately non-degrading (D10.2): it is a local configuration condition with zero HTTP cost, while
75
+ `degraded.json` is globally shared — one path being unable to read a secret should not stop the other paths too.
76
+ **D15 keeps that objection and answers it with scope instead of silence**: a local state records *who wrote it*, and an entry
77
+ obeys only its own (`'cli'` / `'dsh-adapter'`). Service-side states stay `global`.
78
+
79
+ ### 4b. The in-session notice: the only way a host-only plugin can speak to the user
80
+
81
+ DSH gives a plugin without `dsh.client` **no** toast, banner or startup notice — every settings/Plugins seat is a browser-side
82
+ registration, and boot warnings reach the terminal only. So the plugin mounts a second event:
83
+
84
+ | Mechanism | Where | What it is for | What the user sees |
85
+ |---|---|---|---|
86
+ | **`agent/pre-step` waterfall** | `adapters/dsh/index.js` | Appends one `notice`-form user message: first run with no key (the demand, carrying the exact recording command), entering a degraded state, recovering — one per state per session | A row in the conversation, collapsed to `jev-guard · <summary>` and expandable to the body |
87
+
88
+ Three properties are load-bearing, and each is asserted by `tools/smoke-dsh-adapter.mjs`:
89
+
90
+ 1. **It appends, it never replaces.** The decision's `messages` array *is* the whole batch for that step — a listener returning its
91
+ own array silently swallows the user's message. The handler always `await next()` first and returns `[...decision.messages, notice]`.
92
+ 2. **It never injects into an empty batch** (`decision.messages.length === 0 && (step === 1 || messages.length > 0)`): a non-empty
93
+ decision opens a step, so adding a message there would spend a whole extra model request just to say something.
94
+ 3. **The message shape is a contract.** `source` carries exactly `kind` / `plugin` / `form` / `summary`, `summary` is ≤120 characters
95
+ (it becomes the collapsed row's title), and `id` / `role` are set. A malformed message surfaces as
96
+ `SessionPersistenceCorruptionError` at the **next resume** — a session that will not open, far from the change that caused it. The
97
+ shape is therefore validated against DSH's own `snapshotJsonValue` (the step `Session.append` runs first) by a test that must run
98
+ inside a DSH checkout.
99
+
100
+ Deduplication runs against the **durable history** (`agent.session.deriveMessages()`), not an in-memory flag: a harness restart or a
101
+ session resume starts a fresh plugin instance, and a notice already written into the history must not be repeated. `notifyInSession:
102
+ false` turns the whole channel off.
103
+
104
+ The notice is a `role:'user'` message, so it **enters the model's context** — intended (the model should know the valve is degraded),
105
+ and the reason it fires per state transition rather than per step: each one costs a prefix-cache miss from that point on.
106
+
107
+ ### 4c. Where the key comes from (all three layers, and why the third exists)
108
+
109
+ | Order | Source | Notes |
110
+ |---|---|---|
111
+ | 1 | `ctx.credentials.resolve(ref)` | DSH's own credential store (`~/.dsh/.credentials.yaml`); a rotation needs no restart |
112
+ | 2 | the process environment | the variable named by `apiKeyEnv` |
113
+ | 3 | `apiKeyFile` (default `secrets.json` in the package root) | what `guard key set` writes; a relative path resolves against the **package root**, independent of cwd — the same rule the CLI uses |
114
+
115
+ The third layer is what makes "record your key with the CLI" true for a fresh install, which has neither a credential-layer entry
116
+ nor an environment variable. The value is never printed and never logged.
117
+
118
+ Three things to note on the DSH side:
119
+
120
+ 1. **Convey `verdict.warning` to the human** — the plugin writes it into the denial reason, the `level:'warn'` audit record and
121
+ `ctx.logger.warn`. **Do not treat `source: 'degraded'` as "judged harmless"**; it means "this one did not go through semantic judgment".
122
+ 2. **`guard status` can serve as a health check** — while degraded **the exit code is 3**.
123
+ 3. The cooldown length and "which layer is kept while degraded" are in `config.json`: `quotaCooldownMs` (15 minutes) / `authCooldownMs` (30 minutes) /
124
+ `degradePolicy` (`'l0-only'` default / `'off'` = the whole valve paused). When the window expires it automatically lets **one** probe through, and a success restores it.
125
+
126
+ ---
127
+
128
+ ## 5. Silent failure when the plugin is written wrong, and why the real deployment must verify it
129
+
130
+ The worst failure mode on the DSH side is **a plugin that silently is not attached**: no error, no log, commands run as usual.
131
+ On 2026-09-20 **three layers** of same-kind accidents really did happen (the entry guard was constantly false on Windows; after
132
+ extracting it into a shared module for DRY it failed on WSL too; a dynamic `import()` using an absolute path threw
133
+ `ERR_UNSUPPORTED_ESM_URL_SCHEME` outright on Windows). The full retrospective is in [`MEASUREMENTS.md`](./MEASUREMENTS.md) §10. Three rules:
134
+
135
+ 1. **Inline the entry guard, cross-platform:**
136
+ `realpathSync(process.argv[1]) === realpathSync(fileURLToPath(import.meta.url))`,
137
+ **do not extract it into a shared module** (`import.meta.url` follows the module, so once extracted it is constantly false).
138
+ 2. **Use relative specifiers for dynamic imports** (`await import('../../lib/gate.js')`), not absolute path strings.
139
+ 3. **Verification looks at side effects, not at "whether there was an error":** run a command that **is bound to be intercepted**
140
+ and confirm that it **really is intercepted**; then check whether `~/.jev-guard/guard.log` has that record. `tools/selftest-entry.mjs` is the automated
141
+ version of this, and it counts as verified only after **running it once on Windows and once on WSL**.
142
+
143
+ ---
144
+
145
+ ## 6. Every failure is allowed through (fail-open)
146
+
147
+ Timeout (1800 ms by default), network error, service 5xx, code exception → **allowed** and recorded as `source: error`.
148
+
149
+ The reason is not "we do not care": DSH itself still has the sandbox preset (`permission/preset`) taking effect **after** the valve
150
+ (unless it is `danger-full-access`). The valve is an **incremental** check, not the only line of defence. If it were changed to fail-closed,
151
+ one hiccup of the judging service and you cannot work — a thing that "prevents accidents" would have become a thing that
152
+ "manufactures accidents". To achieve "blocked even when the service is down", the right approach is **to thicken the L0 rules**
153
+ (that layer has no network access and depends on no service), not to change this policy. See [`DECISIONS.md`](./DECISIONS.md) D3.
154
+
155
+ ---
156
+
157
+ ## 7. Boundaries: it is an accident net, not a security boundary
158
+
159
+ **Reading this line is enough; there is no need to guess further.** What it prevents is **accidents** — commands the model or a human
160
+ wrote wrong, opaque scripts, the moment under full permissions when nobody intercepts. It does **not** prevent deliberate bypass
161
+ (changing the wording, encoding, writing the authorisation file directly).
162
+
163
+ This is not unfinished work, it is an **explicit decision**: the full rationale, the list of known bypasses, and "under what
164
+ circumstances this should be reconsidered" are in [`DECISIONS.md`](./DECISIONS.md) **D1**. To defend against malice, the right approach
165
+ is **to add another layer** (sandbox / low-privilege user / container), not to rebuild this valve into a security boundary.
166
+
167
+ ---
168
+
169
+ ## 8. Glossary (do not confuse the three "trusts")
170
+
171
+ | Term | What it refers to | Who maintains it | Does it persist |
172
+ |---|---|---|---|
173
+ | **One-shot token** | Allows **the exact text of one command** once | The valve (`~/.jev-guard/allow.txt`) | Not persistent: deleted once used, not replayable |
174
+ | **DSH approval** | A nod for **this one call** | DSH | Only "allow once", **no permanent allow** (see [`USER-INTERVENTION.md`](./USER-INTERVENTION.md) §2.1) |
175
+ | **L0 static rules** | Hard rules with no network (21 deny + 16 ask) | `lib/rules.js` | A token **cannot get past** `deny` |
176
+
177
+ Two numbering systems must also be distinguished: `L0 / L1 / L2` are **judging layers** (see [`ARCHITECTURE.md`](./ARCHITECTURE.md)),
178
+ and `allow / revise / block / escalate` are **judging results**. The former is "how it was computed", the latter is "what was computed".