dsh-jev-guard 0.5.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +285 -0
- package/CHANGELOG.zh-CN.md +271 -0
- package/DEPLOY.md +202 -0
- package/DEPLOY.zh-CN.md +200 -0
- package/LICENSE +21 -0
- package/README.md +316 -0
- package/README.zh-CN.md +315 -0
- package/START-HERE.md +97 -0
- package/START-HERE.zh-CN.md +97 -0
- package/adapters/README.md +37 -0
- package/adapters/README.zh-CN.md +37 -0
- package/adapters/dsh/index.js +502 -0
- package/bin/guard.mjs +634 -0
- package/config.example.json +52 -0
- package/cordis.patch.yml +120 -0
- package/docs/AGENT-TASK-dsh.md +134 -0
- package/docs/AGENT-TASK-dsh.zh-CN.md +131 -0
- package/docs/ARCHITECTURE.md +118 -0
- package/docs/ARCHITECTURE.zh-CN.md +117 -0
- package/docs/DECISIONS.md +469 -0
- package/docs/DECISIONS.zh-CN.md +449 -0
- package/docs/DSH-INTEGRATION.md +178 -0
- package/docs/DSH-INTEGRATION.zh-CN.md +171 -0
- package/docs/MEASUREMENTS.md +433 -0
- package/docs/MEASUREMENTS.zh-CN.md +450 -0
- package/docs/USER-INTERVENTION.md +141 -0
- package/docs/USER-INTERVENTION.zh-CN.md +143 -0
- package/docs/VERIFICATION.md +279 -0
- package/docs/VERIFICATION.zh-CN.md +278 -0
- package/lib/audit.js +228 -0
- package/lib/gate.js +720 -0
- package/lib/i18n.js +575 -0
- package/lib/quota.js +389 -0
- package/lib/rules.js +174 -0
- package/lib/token.js +154 -0
- package/lib/verdict.js +285 -0
- package/package.json +82 -0
- package/tools/check-doc-pairs.mjs +158 -0
- package/tools/extract-commands.mjs +156 -0
- package/tools/gate-cli.mjs +240 -0
- package/tools/probe-prompt-lang.mjs +238 -0
- package/tools/probe-scripts.mjs +143 -0
- package/tools/report-result.mjs +146 -0
- package/tools/selftest-audit.mjs +93 -0
- package/tools/selftest-entry.mjs +177 -0
- package/tools/selftest-i18n.mjs +177 -0
- package/tools/selftest-quota.mjs +260 -0
- package/tools/selftest-reason.mjs +266 -0
- package/tools/selftest-rules.mjs +107 -0
- package/tools/selftest-token.mjs +100 -0
- package/tools/smoke-dsh-adapter.mjs +295 -0
- package/tools/smoke-dsh-pipeline.mjs +146 -0
|
@@ -0,0 +1,450 @@
|
|
|
1
|
+
# 实测数据(全部可复核)
|
|
2
|
+
|
|
3
|
+
> [English](MEASUREMENTS.md) | **简体中文**
|
|
4
|
+
|
|
5
|
+
这些数字不是估计,是 2026-09-20 在**本机**跑出来的。每一节都写了复现方式。
|
|
6
|
+
|
|
7
|
+
## 1. Jev 服务的基本事实
|
|
8
|
+
|
|
9
|
+
| 项 | 值 | 来源 |
|
|
10
|
+
|---|---|---|
|
|
11
|
+
| 端点 | `POST https://api.typesafe.ai/v1/systemone`,`Authorization: Bearer <key>` | 官方文档 |
|
|
12
|
+
| 实际应答模型 | `jev-1.13.0`(`jev-latest` / `jev-preview` 是别名) | `GET /v1/models` 实测 |
|
|
13
|
+
| 计费 | `$0.042 / Mtok` 输入(**输出免费**) | 官方文档 + 响应 usage 实测 |
|
|
14
|
+
| 单次判定成本(约 450 input token) | ≈ `$0.000019` | 实测 token × 单价 |
|
|
15
|
+
| 直连上下文 | 64k/请求(32k 给 state + 最长问题) | 官方文档 |
|
|
16
|
+
| 通过 OpenRouter | 同一模型,上下文 32k,`POST /api/alpha/decisions`,支持支付宝充值 | 实测 `/api/v1/models/typesafe/jev-1.13/endpoints` |
|
|
17
|
+
| 语言 | 英文最佳;CJK 可用但官方说明精度不等同 | 官方 Models 页 |
|
|
18
|
+
|
|
19
|
+
**三问形态冒烟(中文/英文各一次,均 200):**
|
|
20
|
+
|
|
21
|
+
| state | 问题 | 结果 |
|
|
22
|
+
|---|---|---|
|
|
23
|
+
| 客户工单:账单被重复扣款两次…要投诉到消协 | `noul` 是否涉及退款 | 0.99 |
|
|
24
|
+
| 同上 | `choice` 该哪个团队接手 | `billing` 0.77(conf 0.65) |
|
|
25
|
+
| 同上 | `score` 紧急程度 | 2.85/3,最高档 0.91 质量 |
|
|
26
|
+
|
|
27
|
+
## 2. 三臂校准实验(114 个判断,中文 vs 翻译)
|
|
28
|
+
|
|
29
|
+
样本:40 个 case / 114 个判断,覆盖工单分流、危险命令、代码改动、搜索结果打标。
|
|
30
|
+
复现:`~/workspace/jev-calibration/run_calibration.py`(校准脚本的产物,不在本仓库内)。
|
|
31
|
+
|
|
32
|
+
| 臂 | 准确率 | noul | choice | score |
|
|
33
|
+
|---|---|---|---|---|
|
|
34
|
+
| **A 中文直连** | **90.4%**(103/114) | 96.3% | 90.6% | 78.6% |
|
|
35
|
+
| B 机器翻译后判定 | 89.5%(102/114) | 96.3% | 87.5% | 78.6% |
|
|
36
|
+
| C 翻译+回译校验后判定 | 89.5%(102/114) | 94.4% | 90.6% | 78.6% |
|
|
37
|
+
|
|
38
|
+
**配对检验(同一批题逐题比):** A 独有正确 1 题 / B 独有 0 题 / 两者都错 11 题;A vs C:1 / 0 / 11。
|
|
39
|
+
→ **翻译没有救回任何一题,只弄坏了一题。另外三臂平均置信度 0.867 / 0.869 / 0.867 —— "中文置信度偏低"没有复现。**
|
|
40
|
+
|
|
41
|
+
**按问题形态(中文直连):**
|
|
42
|
+
|
|
43
|
+
| 场景 · 问题 | 形态 | 准确率 |
|
|
44
|
+
|---|---|---|
|
|
45
|
+
| 危险命令 · **是否会不可逆破坏数据** | noul | **12/12** |
|
|
46
|
+
| 工单 · 团队归属 | choice(四选一) | 12/12 |
|
|
47
|
+
| 工单 · 是否涉及退款 / 是否需立即处理 | noul | 12/12 |
|
|
48
|
+
| 搜索 · 来源分类 / 可信度档位 | choice / score | 8/8 |
|
|
49
|
+
| 代码 · 是否真回归 / 改动风险档位 | noul / score | 2/2 / 8/8 |
|
|
50
|
+
| 代码 · 是否需人工审核 | noul | 7/8 |
|
|
51
|
+
| 搜索 · 是否切题 | noul | 7/8 |
|
|
52
|
+
| 命令 · 该怎么做 | choice(三选一,语义相邻) | 9/12 |
|
|
53
|
+
| 命令 · 风险档位 | score | **6/12** |
|
|
54
|
+
|
|
55
|
+
**结论:分界线不是语言,是"选项之间的语义距离"。** 是非题 96.3%、语义远的选项 90.6%、
|
|
56
|
+
程度档位 78.6%(最难那组 50%)。四个档位样本(删相册 / `git reset --hard` / 格式化有未备份资料的盘 /
|
|
57
|
+
`docker prune`)的标签存疑,排除后中文直连 = **93.6%**。
|
|
58
|
+
|
|
59
|
+
**置信度闸门交易(中文直连,阈值以下回落给主模型):**
|
|
60
|
+
|
|
61
|
+
| 阈值 | 回落比例 | 捞回的错误 | 剩余错误 |
|
|
62
|
+
|---|---|---|---|
|
|
63
|
+
| 0.5 | 7.0% | 3/11 | 8 |
|
|
64
|
+
| **0.6** | **11.4%** | **6/11** | **5** |
|
|
65
|
+
| 0.7 | 15.8% | 7/11 | 4 |
|
|
66
|
+
| 0.8 | 24.6% | 8/11 | 3 |
|
|
67
|
+
|
|
68
|
+
## 3. 真实命令语料上的表现(737 条)
|
|
69
|
+
|
|
70
|
+
语料来源:DSH 会话日志里的 `tool/call`(581 条去重)+ `~/.bash_history`(156 条)。
|
|
71
|
+
复现:`node tools/extract-commands.mjs --stats` 后跑 `node tools/gate-cli.mjs --sessions`。
|
|
72
|
+
|
|
73
|
+
| 指标 | 值 |
|
|
74
|
+
|---|---|
|
|
75
|
+
| 确定性预筛命中(零网络调用) | 174/737 = **23.6%** |
|
|
76
|
+
| 真实调用 Jev | 563 条 |
|
|
77
|
+
| 三分(0.5 / 0.7) | **98.51% / 0.81%(6 条)/ 0.68%(5 条)** |
|
|
78
|
+
| 在进程内调用的延迟 | 均值 **297–301ms**,P50 267ms,P95 367–405ms |
|
|
79
|
+
| 整份语料花费 | ≈ **$0.011** |
|
|
80
|
+
| p 分布 | P50 = 0.01,P90 = 0.13,max = 0.82(极度两极) |
|
|
81
|
+
| 脚本补齐命中 | 18 条(2.4%),**新增误报 0 条** |
|
|
82
|
+
|
|
83
|
+
被拦下的 5 条(阈值 0.7)全部是真实破坏事件:`git reset --hard`、`git checkout --`、
|
|
84
|
+
真实目录 `rm -rf`×2、`cp 备份→目标`。
|
|
85
|
+
|
|
86
|
+
## 4. 脚本盲区探测(18 个用例)
|
|
87
|
+
|
|
88
|
+
"只看命令行" vs "读进脚本正文/包脚本后再判定":
|
|
89
|
+
|
|
90
|
+
| 用例 | 只看命令 | 补齐后 |
|
|
91
|
+
|---|---|---|
|
|
92
|
+
| `node migrate.mjs`(体内 `DROP COLUMN` + 删备份表) | 0.32 | **0.82** |
|
|
93
|
+
| `pnpm run deploy:prod`(体内 `rsync --delete`) | 0.25 | **0.90** |
|
|
94
|
+
| `pnpm run db:reset`(体内是 node 迁移脚本) | 0.81 | 0.89 |
|
|
95
|
+
| `pnpm run dist:clean`(只删构建产物,应低分) | 0.07 | 0.06 |
|
|
96
|
+
| `python3 cleanup.py`(递归 rmtree,但只删 dist/缓存) | 0.20 | 0.12 |
|
|
97
|
+
| `pnpm test` / `git status` | 0.04 / 0.01 | 0.01 / — |
|
|
98
|
+
|
|
99
|
+
其它单条实测:`truncate -s 0` 0.95 · `find -delete` 0.92 · 内联 `node -e rmSync` 0.91 ·
|
|
100
|
+
`dd of=~/data.db` 0.88 · `rsync --delete` 0.88 · `kubectl delete ns` 0.80 ·
|
|
101
|
+
`git clean -fdx` 0.65 · `git checkout .` 0.64 · `sudo rm -rf /var/lib/docker` 0.65 ·
|
|
102
|
+
`docker compose down`(无 -v)0.35(低分正确,没删卷)·
|
|
103
|
+
**`terraform apply -auto-approve` 0.48(已知漏网点)** · `npm publish` 0.03。
|
|
104
|
+
|
|
105
|
+
## 5. 其它实测约束
|
|
106
|
+
|
|
107
|
+
| 项 | 值 | 影响 |
|
|
108
|
+
|---|---|---|
|
|
109
|
+
| 文件系统 | `/` 是 ext4;`cp --reflink` 不支持;无 btrfs/zfs | **没有廉价写时复制快照** |
|
|
110
|
+
| `/tmp` | tmpfs(占内存) | 大备份不能放 /tmp |
|
|
111
|
+
| 磁盘 | 余量充足 | 空间不是瓶颈 |
|
|
112
|
+
| 工作区 | workspace 含若干 git 仓库 | git 提交可当免费回退点 |
|
|
113
|
+
| 硬链接 `cp -al` | 只防删除、不防覆写(同一 inode) | 不能当"备份",要写进设计 |
|
|
114
|
+
| OpenRouter 购买手续费 | 信用卡/支付宝 5.5%(最低 $0.80);加密 5% | 支付宝路径的额外成本 |
|
|
115
|
+
|
|
116
|
+
## 6. 重启后实测(2026-09-20 12:52 起)
|
|
117
|
+
|
|
118
|
+
DSH 安装 + 重启后的现场验证(探针全部无害):
|
|
119
|
+
|
|
120
|
+
| 验证项 | 结果 |
|
|
121
|
+
|---|---|
|
|
122
|
+
| L0 路径 | 远端强推类命令被拦,理由含硬规则 id `git-force-push` |
|
|
123
|
+
| **Jev 联网路径** | `rm -rf ~/jev-guard-probe-dir`(路径不存在)被拦,理由含「Jev 判定风险概率 70.0%」;CLI 独立复核同一命令 p=0.80 |
|
|
124
|
+
| 策略分支 | 本会话是完全权限(approval=never):理由含「这是自动判定,不是用户手动拒绝」 |
|
|
125
|
+
| 无误伤 | `ls -la /tmp`、写自建探针文件(p=0.11)、`node -e rmSync`(p=0.22)均正常放行 |
|
|
126
|
+
| 延迟 | 进程内调用 ~300ms(P50 267ms / P95 405ms,737 条语料均值) |
|
|
127
|
+
|
|
128
|
+
## 7. 实测中的假阳性(已修,待下次重启生效)
|
|
129
|
+
|
|
130
|
+
下面三条都是**阀门拦住操作者自己**时暴露的。它们说明 L0 的"全文匹配"策略有代价,
|
|
131
|
+
以及 fire-and-forget 写入的退出语义必须显式处理。
|
|
132
|
+
|
|
133
|
+
### 7.1 L0 匹配整条命令文本 → 写文档/写测试也会被拦
|
|
134
|
+
|
|
135
|
+
L0 原本对**整条命令**(含参数)做正则匹配,于是:
|
|
136
|
+
|
|
137
|
+
| 场景 | 现象 |
|
|
138
|
+
|---|---|
|
|
139
|
+
| 用 `--evidence "…<危险命令原文>…"` 记录验证结果 | 被对应规则拦下(共 3 次) |
|
|
140
|
+
| 在 bash heredoc 里写含测试用例字面量的自检脚本 | 被 `rm-root` / `find-delete` 拦下 |
|
|
141
|
+
| 记录文本里提到批量删除类命令 | 同上 |
|
|
142
|
+
|
|
143
|
+
**修法:** 21 条易与散文冲突的规则改为 **`where: 'command'`** —— 只在命令位置匹配
|
|
144
|
+
(行首,或 `;` `&` `|` `(` `$(` 反引号之后,允许 `sudo`/`env`/`command`/`nohup`/`time` 包装)。
|
|
145
|
+
依据:实测 Jev 对同类散文只给 **p=0.02–0.08**,对真实调用给 **0.8–1.0**,
|
|
146
|
+
所以引号里藏的命令交给 Jev 兜底不丢覆盖率。自检:`tools/selftest-rules.mjs`(25 例)。
|
|
147
|
+
|
|
148
|
+
> **更正(2026-09-20 晚):**上面这句"21 条改为命令位置匹配"**当时并没有做全** ——
|
|
149
|
+
> 实际只覆盖了 7 条 deny + 全部 16 条 ask,`mkfs` / `dd` 这类 12 条仍是全文匹配。
|
|
150
|
+
> 两个方向上的偏差在 **§7.5** 里一并修掉了。
|
|
151
|
+
|
|
152
|
+
### 7.2 `truncate-file` 的 `>` 分支匹配**任何以重定向结尾**的命令
|
|
153
|
+
|
|
154
|
+
原正则第二段 `>\s*[^\s|]+\s*$` 没有 `m` 标志,`$` 即字符串末尾 —— 实际语义变成
|
|
155
|
+
"命令以 `> 某路径` 结尾",于是这种日常写法被判成"把文件截断为空":
|
|
156
|
+
|
|
157
|
+
```bash
|
|
158
|
+
node bin/guard.mjs log --tail 4 2>/dev/null # ← 实测被拦
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
**修法:** 只认显式空写形态(`truncate -s 0 <path>`、`echo "" > <path>`、`: > <path>`),
|
|
162
|
+
普通重定向交给 Jev(实测 `echo "" > 某文件` Jev 给 p=0.70,仍会在阈值上被拦)。
|
|
163
|
+
自检补 5 条用例(`2>/dev/null`、`echo hello > f`、`cat a > b` 等均不得命中)。
|
|
164
|
+
|
|
165
|
+
### 7.3 `record()` 是 fire-and-forget,进程退出会丢尾部记录
|
|
166
|
+
|
|
167
|
+
冒烟测试实测:写完立刻 `process.exit()` → 日志文件根本没生成。
|
|
168
|
+
已加 `flush()`,接在插件的 `dispose` 上;宿主退出路径应 await 一次。
|
|
169
|
+
|
|
170
|
+
### 7.4 `logPath: ''` + `??` = 写入全部静默失败(第二次重启后才暴露)
|
|
171
|
+
|
|
172
|
+
`cordis.patch.yml` 里用 `logPath: ''` 表示"用默认路径"(配置模板的常见写法),
|
|
173
|
+
但 `cfg.logPath ?? DEFAULT_LOG_PATH` 只对 `null`/`undefined` 生效 —— 空串被当成真实路径,
|
|
174
|
+
于是 `appendFile('')` 抛错,又被 `.catch(() => {})` 静默吞掉。
|
|
175
|
+
表面现象:**阀门工作正常(命令照拦),但 `guard log` 一条记录都没有。**
|
|
176
|
+
|
|
177
|
+
**修法:** 新增 `resolveLogPath()` —— 空串/纯空白一律视为未配置;`record`/`readTail`/`summarize`/CLI 都走它。
|
|
178
|
+
同时**让失败可见**:`lastLogError()` 暴露最近一次错误,`guard log` 在"没有记录"时把它打出来。
|
|
179
|
+
隔离环境端到端自检已覆盖(审计自检 21 例,含"空白 logPath 仍落盘")。
|
|
180
|
+
|
|
181
|
+
**两个教训:**
|
|
182
|
+
1. 面向用户的"空串即默认"约定,必须在解析处显式处理 —— `??` 不够。
|
|
183
|
+
2. **静默 catch 会掩盖整整一轮工作。** 审计模块"绝不抛错"是对的(不能因为写日志影响判定),
|
|
184
|
+
但必须留一个可见的出口;这次就是没有任何出口,于是"0 条记录"看起来像"插件没跑"。
|
|
185
|
+
|
|
186
|
+
### 7.5 锚定只做了一半 + 缺 `m` 标志 → 假阳与漏判**同时**存在(2026-09-20 第二次修正)
|
|
187
|
+
|
|
188
|
+
第 7.1 节当时写的是"21 条规则改为命令位置匹配",**实际只改了 7 条 deny + 全部 16 条 ask**;
|
|
189
|
+
`mkfs` / `dd` / `shred` / `chmod -R /` / `vssadmin` / `wbadmin` / `cipher /w` / `diskpart` /
|
|
190
|
+
`wsl --unregister` / `kubectl delete ns` / `Clear-Disk` / `Remove-Item … -Recurse` 这 **12 条**
|
|
191
|
+
仍在**全文匹配**。触发这次排查的是一条"查日志"的命令:它把 `mkfs.ext4 /dev/…` 的原文写进了
|
|
192
|
+
python 源码的字符串里(`c.startswith('mkfs.ext4 …')`),被 `mkfs` 规则当成命令拦下。
|
|
193
|
+
|
|
194
|
+
顺着查下去发现了反方向的、更要紧的偏差:`COMMAND_POSITION` 用了 `^` 但**没有 `m` 标志**,
|
|
195
|
+
所以"命令位置"实际只等于整串开头。凡被锚定的规则,多行命令里第二行起的命令位置全部失效:
|
|
196
|
+
|
|
197
|
+
| 形态 | 修正前 | 修正后 |
|
|
198
|
+
|---|---|---|
|
|
199
|
+
| `git push --force origin main` | HIT | HIT |
|
|
200
|
+
| `cd /tmp && git push --force …` | HIT | HIT |
|
|
201
|
+
| `echo x \| xargs git push --force …` | **MISS**(`xargs` 不在包装器列表里) | HIT |
|
|
202
|
+
| `bash - <<'SH'` + `git push --force …` | **MISS**(`^` 只匹配串首) | HIT |
|
|
203
|
+
| `bash -c "` + 多行 + `git push --force …` | **MISS**(同上,且引号后不算命令位置) | HIT |
|
|
204
|
+
| heredoc 里的 `rm -rf /`、`DROP DATABASE` | **MISS** | HIT |
|
|
205
|
+
| python heredoc 里的字符串 / 注释 / 赋值 / grep 参数 | **HIT(假阳)** | —(交给 Jev) |
|
|
206
|
+
|
|
207
|
+
注意最后两行的因果关系:**修正前 heredoc 里的 `mkfs` 反而是命中的** —— 只因为它没锚定。
|
|
208
|
+
假阳与漏判是同一个根因(锚定只做了一半)的两个方向。
|
|
209
|
+
|
|
210
|
+
**修法:** ① 12 条命令开头型规则补 `where: 'command'`;② 锚定正则加 `m`;
|
|
211
|
+
③ 包装器扩到 `sudo/doas/env/command/nohup/time/nice/ionice/setsid/stdbuf/watch/timeout/xargs/parallel/find`,
|
|
212
|
+
且吞掉的参数只允许 ASCII 词/flag/路径字符(中文散文因此仍不会被顺带命中 —— 实测 `xargs 删除 mkfs…` 不命中);
|
|
213
|
+
④ `bash -c "` 也算命令位置;⑤ 只剩 `redirect-to-device`(`>`)与 `fork-bomb`(`:(){…};:`)
|
|
214
|
+
显式标为 `where: 'anywhere'`,`RULE_STATS.anywhere` 恒为 2。
|
|
215
|
+
自检 `tools/selftest-rules.mjs` 从 25 例扩到 **48 例**(含 1 例性能:4KB 包装器前缀 0.6ms,防灾难性回溯)。
|
|
216
|
+
|
|
217
|
+
**为什么漏判比假阳要紧:** L0 存在的理由就是在 `l0-only` 降级(没有额度、没有网络)时兜住
|
|
218
|
+
`mkfs` / `dd of=/dev/*` / `git push --force` 这一类(见 D9)。平时漏判被 Jev 补上,所以一直没人发现;
|
|
219
|
+
降级时它就是真空。假阳的代价则只有"AI 不能用 bash 写含这些字面量的东西"(文件工具不经阀门)。
|
|
220
|
+
|
|
221
|
+
**同一个 `m` 标志的坑,这个项目已经踩了两次:** §7.2 的 `truncate-file` 是规则里的 `$` 缺 `m`
|
|
222
|
+
(表现是**误拦**),这次是锚定里的 `^` 缺 `m`(表现是**漏判**)。表象相反,根因相同 ——
|
|
223
|
+
写"行首/行尾"这类断言时,先问一句"多行输入下它还成立吗"。
|
|
224
|
+
|
|
225
|
+
## 8. 现场取到的四个真实判定(维护动作上的命中,全部来自本项目自己的工作)
|
|
226
|
+
|
|
227
|
+
这一节记录阀门**对项目维护者自己**开火的实例 —— 它们最能说明"第 3 类"(先别执行、看有没有更安全的写法)
|
|
228
|
+
到底在防什么。全部可复核:`guard.log` 里有原始记录。
|
|
229
|
+
|
|
230
|
+
| # | 命令在做的事 | p | 判定 | 实际风险 | 采纳的替代写法 |
|
|
231
|
+
|---|---|---|---|---|---|
|
|
232
|
+
| 1 | 往部署实例复制文件 + `rm -f` 删一个旧文件 | 0.64 | revise | **真的有删除动作**(删的是已部署文件) | 拆成"纯复制"+"改名保留";然后通过 |
|
|
233
|
+
| 2 | `sed -i` 原地改源码里的注释 | 0.50 | revise | 原地覆盖一个真实文件 | 改用文件编辑工具(本该如此,见 §7.2 同类) |
|
|
234
|
+
| 3 | 同一条同步命令的早期版本 | 0.63 | revise | 同上 | 同上 |
|
|
235
|
+
| 4 | 用户手动执行 vs agent 重试同一条命令 | — | 拦 / 放 | 见下 | — |
|
|
236
|
+
|
|
237
|
+
**第 4 条值得单独说:** 人在自己终端里把文件截断(153→0 字节)后,agent 重试**一字不差的同一条命令**
|
|
238
|
+
仍然被拦(p 不变),审计**零新增记录**。即"人做了一次" ≠ "给 agent 开了口子"。
|
|
239
|
+
|
|
240
|
+
**这几条的含义(也是为什么阈值不改,见 D4):**
|
|
241
|
+
- 三类里最容易被骂"过严"的就是第 1、3 类:命令的**意图**完全正当,但文本里确实含删除/原地覆盖。
|
|
242
|
+
阀门没有读心术,它只看见"要删一个真实路径上的文件"。**这类误伤是它正常工作的样子,不是故障**;
|
|
243
|
+
代价是维护者要换成"改名保留 / 拆命令 / 用文件工具"这三种更安全的写法 —— 每次都不超过 10 秒。
|
|
244
|
+
- 反过来说:**如果把它调到不拦这些,阈值就得抬到 0.65 以上**,而 0.68–0.82 这一段里躺着
|
|
245
|
+
`docker compose down -v`(p=0.69)这类真实不可逆操作(见 D4)。用真事故换来的清静,不值。
|
|
246
|
+
- 顺带证明了一件小事:**阀门对自己人也一样**。它不知道"这条命令是我自己发的",也不需要知道。
|
|
247
|
+
|
|
248
|
+
## 9. 额度降级的现场实测(2026-09-20)
|
|
249
|
+
|
|
250
|
+
判定服务是收费的,所以"额度用完"必须当成**必然事件**来设计,而不是异常分支。用**无效密钥**
|
|
251
|
+
打真实 API 取到了第一手的失败形状与降级行为(状态文件与审计都隔离到临时目录):
|
|
252
|
+
|
|
253
|
+
| 步骤 | 实测结果 |
|
|
254
|
+
|---|---|
|
|
255
|
+
| 无效密钥调用 | `HTTP 401`,响应体 `{"detail":{"error_type":"authentication_error","message":"Cannot authenticate with the server. Please check your API key and try again."}}` |
|
|
256
|
+
| 分类 | `auth`(→ 降级;`429` 只在正文含 quota/credit/insufficient 等词时才降级,否则当限流) |
|
|
257
|
+
| 这次判定 | 仍然 `allow`(fail-open),但带 `errorKind: auth` + `degraded` + 一句成文告警 |
|
|
258
|
+
| 状态文件 | 写出 `degraded.json`:`kind/label/since/until(ISO)/failures/probes/status/detail/policy` |
|
|
259
|
+
| 第二次调用 | `source: degraded`、**零 HTTP 请求**(省钱)、`p` 为空,理由明确写"本条未经过语义判定" |
|
|
260
|
+
| 降级期间 L0 | `mkfs.ext4 …` 仍 `block/static-rule`,`truncate -s 0 …` 仍 `escalate/static-rule`,都不发请求 |
|
|
261
|
+
| `guard status` | 打印原因/开始时间/剩余恢复时间/现在还剩哪一层/原始错误;**退出码 3** |
|
|
262
|
+
| `degradePolicy: 'off'` | 连 L0 也放行(显式选择;默认不是这个) |
|
|
263
|
+
| 自动恢复 | 冷却到期后放**一次**探测;替身 fetch 实测:成功 → 清状态并记 `probe+recovered`,失败 → 续期(failures+1、probes+1)且不再重复试 |
|
|
264
|
+
|
|
265
|
+
**离线断言 52 条**(`tools/selftest-quota.mjs`,替身 `fetch` 覆盖 402/401/403/429两种/5xx/超时/网络/无密钥/
|
|
266
|
+
坏状态文件),其中两条是**自己抓到的真 bug**,一并记在这里:
|
|
267
|
+
|
|
268
|
+
1. `readDegraded` 用 `Number(until)` 校验 ISO 字符串 → 恒为 NaN → **写进去了却永远读不出来**,
|
|
269
|
+
整个降级机制静默失效(不抛异常)。
|
|
270
|
+
2. `isProbe = degradedNow && probeDue(...)` —— 状态到期时 `isDegraded()` 恰好是 `false`,
|
|
271
|
+
于是"到期后的那次探测"永远不算探测,**恢复与续期全部失效**。
|
|
272
|
+
|
|
273
|
+
两个都是"静默失效"型错误,都靠"每个分支都写一条断言"才被抓住 —— 与 §7 的教训同源。
|
|
274
|
+
|
|
275
|
+
## 10. 跨平台入口守卫:同一个坑踩了两次(2026-09-20)
|
|
276
|
+
|
|
277
|
+
**现象:** 一个以 `node <路径>` 被直接执行的脚本,在 Windows 上**不输出、退出码 0、不写任何日志** ——
|
|
278
|
+
既不判定也不记录。而在不少调用约定里**退出码 0 就是"放行"**,所以这是最糟的失效形态。
|
|
279
|
+
更麻烦的是:用"日志里有没有记录"去验证,会得出**与事实相反**的结论("这个宿主不执行脚本"),
|
|
280
|
+
进而错误地放弃整条路线。WSL/Linux 上同样的调用一切正常,缺陷只在 Windows 暴露。
|
|
281
|
+
|
|
282
|
+
### 10.1 第一层:入口守卫的字符串比较在 Windows 上恒为 false
|
|
283
|
+
|
|
284
|
+
```js
|
|
285
|
+
if (import.meta.url === `file://${process.argv[1]}`) await main() // ← 旧写法
|
|
286
|
+
```
|
|
287
|
+
|
|
288
|
+
| 平台 | `process.argv[1]` | `import.meta.url` | 相等? |
|
|
289
|
+
|---|---|---|---|
|
|
290
|
+
| Windows | `T:\dsh-jev-guard\bin\guard.mjs` | `file:///T:/dsh-jev-guard/bin/guard.mjs` | **false** |
|
|
291
|
+
| WSL | `/mnt/t/dsh-jev-guard/bin/guard.mjs` | `file:///mnt/t/dsh-jev-guard/bin/guard.mjs` | true |
|
|
292
|
+
|
|
293
|
+
**修法:** `realpathSync(process.argv[1]) === realpathSync(fileURLToPath(import.meta.url))`。
|
|
294
|
+
`realpathSync` 一次解决盘符、反斜杠、相对路径与**软链**(包目录本身可能是软链,DSH 用 `link:` 装载)。
|
|
295
|
+
当时共 3 处(其中 2 处已随范围收窄归档到包外),包内现存的入口脚本是 `tools/extract-commands.mjs`。
|
|
296
|
+
|
|
297
|
+
### 10.2 第二层:修第一层时**自己**制造了第二个同类 bug
|
|
298
|
+
|
|
299
|
+
为了 DRY,曾把守卫抽进共享模块 `lib/entry.js`。结果守卫恒为 false —— 因为 **`import.meta.url`
|
|
300
|
+
是每个模块各自的**,进了 `lib/entry.js` 之后比较对象就变成了那个 lib 文件自己的路径。
|
|
301
|
+
**连 WSL 上也一起静默失效了**(当时只在 Windows 上跑了验证,差点漏过)。
|
|
302
|
+
|
|
303
|
+
**规则(见 DECISIONS D10):** 入口守卫**必须内联在各自文件里**。那不是重复代码,
|
|
304
|
+
而是"每份代码只谈自己的身份"。各文件都留了一段注释说明为什么不能抽走。
|
|
305
|
+
|
|
306
|
+
### 10.3 第三层:动态 `import()` 的绝对路径在 Windows 上不是合法说明符
|
|
307
|
+
|
|
308
|
+
修完入口守卫后,Windows 上脚本**终于跑起来了**,但判定一抛错就回到 fail-open。日志给出了原因:
|
|
309
|
+
|
|
310
|
+
```
|
|
311
|
+
ERR_UNSUPPORTED_ESM_URL_SCHEME: Only URLs with a scheme in: file, data, and node are supported
|
|
312
|
+
by the default ESM loader. On Windows, absolute paths must be valid file:// URLs.
|
|
313
|
+
```
|
|
314
|
+
|
|
315
|
+
`await import(join(ROOT, 'lib', 'gate.js'))` —— 绝对路径字符串在 Windows 上不是合法 ESM 说明符。
|
|
316
|
+
**修法:** 改用**相对说明符** `await import('../../lib/gate.js')`(按本模块自己的 URL 解析,
|
|
317
|
+
两平台都成立,也不怕软链)。
|
|
318
|
+
|
|
319
|
+
> 这一层解释了为什么"跑起来"和"能判定"是两件事:第三层修好之前,脚本在 Windows 上**能运行、
|
|
320
|
+
> 有日志,但仍然放行一切**。只验"有没有输出/有没有日志"会误判成修好了。
|
|
321
|
+
|
|
322
|
+
### 10.4 决定性验证:必须在 Windows 上跑
|
|
323
|
+
|
|
324
|
+
WSL 上验不出这三层里的任何一层。这次用 `C:\Program Files\nodejs\node.exe`(可从 WSL 互操作调用)
|
|
325
|
+
对报告里那条原始 payload 各跑一遍:
|
|
326
|
+
|
|
327
|
+
| 验证 | 修复前 | 修复后 |
|
|
328
|
+
|---|---|---|
|
|
329
|
+
| 退出码 | **0** | **2** |
|
|
330
|
+
| stdout | 0 字节 | `{"decision":"block",…,"permissionDecision":"deny"}` |
|
|
331
|
+
| Windows 侧诊断日志 | 无记录 | `{outcome:"block",source:"static-rule",rule:"git-force-push"}` |
|
|
332
|
+
|
|
333
|
+
### 10.5 新增回归自检:`tools/selftest-entry.mjs`(跨平台,15 例)
|
|
334
|
+
|
|
335
|
+
这份自检的存在理由,就是上面三类事故**别的检查都抓不到**:`node --check` 只查语法;
|
|
336
|
+
其它自检都 `import` 模块、不走入口路径;而守卫写错的表现是**静默退出 0**。
|
|
337
|
+
所以它真的 spawn 每个入口脚本、看有没有可观察副作用,并且:
|
|
338
|
+
|
|
339
|
+
- 断言"被 import 时**不得**执行 main"(报告里明确警告过不要用"删掉守卫"来修);
|
|
340
|
+
- 断言"通过**软链**执行仍成立";
|
|
341
|
+
- 断言每个入口脚本的守卫都**定义在自己文件里**、**没有**从共享模块导入 —— 直接钉死 10.2 那类回归。
|
|
342
|
+
|
|
343
|
+
在 Windows 上运行它,覆盖"盘符 + 反斜杠的 argv[1]"这半边 —— 本次修复的决定性一步。
|
|
344
|
+
|
|
345
|
+
## 11. 一个"局部问题造成全局失能"的设计修正(2026-09-20)
|
|
346
|
+
|
|
347
|
+
`bin/guard.mjs` 解析密钥时,旧写法 `cfg.apiKeyFile ?? join(ROOT, 'secrets.json')` 只在字段**缺失**时
|
|
348
|
+
回落到包根;一旦配置里填的是**相对路径**,它就按**调用时的 cwd** 解析 —— 在包目录外调用
|
|
349
|
+
(CLI 与离线脚本都可能)读不到密钥。
|
|
350
|
+
|
|
351
|
+
单看这条只是"读不到密钥"。但配上 §9 的降级机制,后果被放大成:
|
|
352
|
+
|
|
353
|
+
**某条调用路径读不到密钥 → 写一份共享的 `no-key` 降级 → 密钥其实正常的另一条路径
|
|
354
|
+
(DSH 插件走 `ctx.credentials`,CLI 走环境变量或文件)也一起停掉联网判定 30 分钟。**
|
|
355
|
+
|
|
356
|
+
两处修正:
|
|
357
|
+
|
|
358
|
+
1. **路径语义**:相对路径一律按**包根**解析(绝对路径原样使用),与 cwd 无关。
|
|
359
|
+
2. **`no-key` 不再降级**:它是**本地配置**状况而不是服务状况 —— 一次 HTTP 都不发(降级省不下任何东西),
|
|
360
|
+
却可能只影响一个入口。分类照记、告警照发(`guard status` 会直接说没解析到密钥),但不进降级态。
|
|
361
|
+
降级集合因此收窄为 `{quota, auth}` 两类"服务对我们的态度已经变了"的情况。
|
|
362
|
+
|
|
363
|
+
**规律:** 凡是"共享状态 + 各入口前提独立"的组合,都要问一句"这个局部故障会不会被写进全局状态"。
|
|
364
|
+
|
|
365
|
+
## 12. "拿不到的信息"被默认值假装过:一条可迁移的教训(2026-09-20)
|
|
366
|
+
|
|
367
|
+
历史上把阀门挂进另一条执行通道时,实测抓到一个**结构性**缺陷,记在这里因为教训与平台无关:
|
|
368
|
+
|
|
369
|
+
- 那套实现从**宿主注入不了的 env** 里取"这次会话能不能问人",同时**忽略 payload 里自带的权限字段**;
|
|
370
|
+
- 并且无论取到什么,都输出**同一个**拒绝结论。
|
|
371
|
+
- 实测:把 payload 里的权限字段从"需要审批"改成"不需要审批",输出**逐字节相同**,
|
|
372
|
+
连理由里那句"本会话没有审批提示"都没变 —— 对带审批的会话是**事实错误**。
|
|
373
|
+
|
|
374
|
+
后果:文档里承诺的双行为("能不能问人"决定弹框还是硬拒)在那一侧**退化成只有硬拒**,
|
|
375
|
+
**人工审批通道不可达**。这不是配置问题,是设计没接上。
|
|
376
|
+
|
|
377
|
+
**可迁移的规则(已写进 DECISIONS D11):**
|
|
378
|
+
|
|
379
|
+
1. **信息就在手上时,别去读别处。** 权限/策略这类"这次调用的事实"应当从**本次调用的上下文**里取。
|
|
380
|
+
2. **拿不到的信息要显式承认拿不到**,别用一个默认值假装它存在 —— 默认值会让上层代码以为
|
|
381
|
+
"两种行为都实现了",而实际上只有一种。
|
|
382
|
+
3. **同一个字段,写入方与读取方必须约定同一个来源**;只要有一处猜,验证就会得出与事实相反的结论
|
|
383
|
+
(与 §10 同源:验证方式本身在骗人)。
|
|
384
|
+
|
|
385
|
+
## 13. Windows 授权行引号:两种写法只有一种能用(2026-09-20 实测)
|
|
386
|
+
|
|
387
|
+
被拦命令的理由里会给用户一行"整行复制粘贴"的授权命令。这一行的引号**必须按平台分叉**:
|
|
388
|
+
|
|
389
|
+
| 形式 | 在真实 PowerShell 里的结果 |
|
|
390
|
+
|---|---|
|
|
391
|
+
| `'it''s a test'`(**win32 形式**) | ✅ `[Console]::Out.Write(...)` 还原为 `it's a test` |
|
|
392
|
+
| `'it'\''s a test'`(POSIX 形式,修复前 Windows 上给的就是这个) | ❌ **ParserError:语法都不成立** |
|
|
393
|
+
|
|
394
|
+
验证方式:`tools/selftest-reason.mjs` 里两条真机断言 —— 一条把 win32 形式喂给真实 PowerShell
|
|
395
|
+
(本机通过 WSL 互操作调用 `pwsh.exe`)要求逐字节还原,另一条**要求 POSIX 形式在 PowerShell 里失败**
|
|
396
|
+
(反例也是断言,否则"分叉"这件事没被证明)。找不到 PowerShell 时自动跳过。
|
|
397
|
+
|
|
398
|
+
**为什么这一条值得单独记:** 这段话是给用户**照抄**的。粘进去语法错 = 授权入口不可用 = 人唯一的出口被堵住,
|
|
399
|
+
而**在被拦发生之前没有任何人会注意到**。这类"只在罕见路径上暴露"的问题,正是本项目反复吃到的亏。
|
|
400
|
+
另外 cmd.exe 两种写法都不认 —— 所以补了与 shell 无关的 `guard allow --command-file <文件>`。
|
|
401
|
+
|
|
402
|
+
|
|
403
|
+
|
|
404
|
+
|
|
405
|
+
|
|
406
|
+
|
|
407
|
+
|
|
408
|
+
## 14. 判定问话换语言会移动边界(2026-09-20 实测,21 条 × 每臂 3 次 × 2 轮)
|
|
409
|
+
|
|
410
|
+
**要回答的问题:** `promptLang: 'en'` 是把那句问话换成英文。阈值 0.5 / 0.7 是在**中文问话**上
|
|
411
|
+
标定的(§2,114 例),那么换语言到底动多少?这决定了英文问话能不能当"翻译"用,还是必须重标定。
|
|
412
|
+
|
|
413
|
+
**方法(`tools/probe-prompt-lang.mjs`):** 21 条真实形态的命令,每条**同一个 state** 分两次问 ——
|
|
414
|
+
一次用中文问话、一次用英文问话(state 的键也随语言切换,载荷信息完全对等);每臂**重复 3 次**取均值,
|
|
415
|
+
并记录臂内极差当作噪声地板。两轮独立运行。
|
|
416
|
+
|
|
417
|
+
**噪声地板(先说清楚,否则数字没意义):** 同一个 state、同一句问话,连问三次并不给出同一个数。
|
|
418
|
+
单独测过一条:`DELETE FROM sessions WHERE expired = 1` 得到 **0.78 / 0.79 / 0.82**;
|
|
419
|
+
本实验 42 条臂的平均臂内极差是 **0.015–0.018**。所以小于 0.02 的差不能当结论。
|
|
420
|
+
|
|
421
|
+
**两轮结果(独立运行,结论一致):**
|
|
422
|
+
|
|
423
|
+
| 指标 | 第 1 轮 | 第 2 轮 |
|
|
424
|
+
|---|---|---|
|
|
425
|
+
| 同带一致率 | 18/21 | 18/21 |
|
|
426
|
+
| 概率平均绝对差 | 0.050 | 0.049 |
|
|
427
|
+
| 带符号均值(负 = 英文更宽松) | −0.042 | −0.043 |
|
|
428
|
+
| 偏低 / 偏高 / 持平 | 12 / 4 / 5 | 12 / 3 / 6 |
|
|
429
|
+
| 臂内噪声均值 | 0.018 | 0.015 |
|
|
430
|
+
| 平均延迟(单次) | 391ms | 391ms |
|
|
431
|
+
|
|
432
|
+
**三条翻带(两轮完全相同,且都朝"更宽松"方向):**
|
|
433
|
+
|
|
434
|
+
| 命令 | 中文问话 | 英文问话 |
|
|
435
|
+
|---|---|---|
|
|
436
|
+
| `psql -c "UPDATE users SET active = false"`(无 WHERE) | 0.73 → **block** | 0.62 → **revise** |
|
|
437
|
+
| `node -e "require('fs').rmSync('/data/out',{recursive:true})"` | 0.61 → **revise** | 0.44 → **allow** |
|
|
438
|
+
| `mysql -e "DELETE FROM sessions WHERE expired = 1"` | 0.79 → **block** | 0.63 → **revise** |
|
|
439
|
+
|
|
440
|
+
**结论(写进 DECISIONS D14):**
|
|
441
|
+
|
|
442
|
+
1. 偏移**有方向**:12 条偏低 / 4 条偏高,**不是**随机抖动(两轮同向、同幅度)。
|
|
443
|
+
英文问话平均把 p 压低约 **0.04**,在 0.5 / 0.7 附近的命令上足以**翻一个带**。
|
|
444
|
+
2. 因此 `promptLang` 的默认值保持 **`'zh-CN'`** —— 它是被标定过的那个。英文问话是可选项,
|
|
445
|
+
选了就等于把边界整体往"更宽松"挪,需要重标定(或把阈值下调 ~0.04)之后再信任。
|
|
446
|
+
3. 本实验**没有**测"哪种语言更准":这里没有人工标签,只有一致性。要谈准确性,得重做 §2 那套
|
|
447
|
+
带标签的三臂校准。
|
|
448
|
+
|
|
449
|
+
**仍未覆盖:** 真实会话语料(§3 的 737 条)没有重跑 —— 那次语料来自已不可用的 DSH 会话日志。
|
|
450
|
+
本实验的 21 条是**手工挑选、覆盖三档风险**的探针,不是随机样本。
|
|
@@ -0,0 +1,141 @@
|
|
|
1
|
+
# How a human gets involved — three channels, and their evidence
|
|
2
|
+
|
|
3
|
+
> **English** | [简体中文](USER-INTERVENTION.zh-CN.md)
|
|
4
|
+
|
|
5
|
+
The valve works in a "nobody is watching" mode (YOLO / full permissions / auto-approve), so **"how a human gets in" has to be part of the design** — it cannot rely on "a prompt will pop up anyway". This document spells out the three channels: what each of them needs, who carries it out, and what proves it actually works.
|
|
6
|
+
|
|
7
|
+
> Every conclusion carries **measured evidence** (2026-09-20, DSH). What has not been measured is written as "unverified", not as fact.
|
|
8
|
+
|
|
9
|
+
---
|
|
10
|
+
|
|
11
|
+
## 0. In one sentence
|
|
12
|
+
|
|
13
|
+
**A human has only three ways to get involved: hand over execution rights once, do it themself, or nod on the spot.** All three implemented — only then is there really "someone in charge".
|
|
14
|
+
|
|
15
|
+
| Channel | Who carries it out | The valve's role | Does it depend on the host? |
|
|
16
|
+
|---|---|---|---|
|
|
17
|
+
| **① One-shot token** | **the AI** (retries after being blocked) | verify the token → let **that one command** through once | **No dependency** — pure local hash + a file |
|
|
18
|
+
| **② Host approval** | **the AI** (with a human nod) | only marks the command as "a human needs to look at it" | depends on the host's ability to ask a human |
|
|
19
|
+
| **③ A human doing it by hand** | **the human** | **not involved at all** | No dependency |
|
|
20
|
+
|
|
21
|
+
---
|
|
22
|
+
|
|
23
|
+
## 1. Channel ①: the one-shot token (host-independent, available at any time)
|
|
24
|
+
|
|
25
|
+
**Mechanism.** A command judged `revise` / `block` / `escalate` gets an
|
|
26
|
+
`ALLOW-XXXXXXXXXX` attached to its reason (= the first 10 hex digits of `sha256(normalised command text)`, uppercase).
|
|
27
|
+
Once a human writes it into `~/.jev-guard/allow.txt`, the AI **retrying the same command** is let through once, and the token is deleted right after.
|
|
28
|
+
|
|
29
|
+
**Six properties** (each one has a test):
|
|
30
|
+
|
|
31
|
+
| Property | Meaning | Evidence |
|
|
32
|
+
|---|---|---|
|
|
33
|
+
| **Bound to the command's exact text** | change one character and it is a different token; rephrasing gets you no authorisation | a variant with a trailing slash was blocked, and what it gave out was a **different** token, `ALLOW-ADCD5EA86D` |
|
|
34
|
+
| **One-shot** | deleted the moment it is used, impossible to replay | after the pass, `allow --list` is empty |
|
|
35
|
+
| **Deterministic** | the same command always gets the same token | the same command appeared three times, and was always `ALLOW-5031AC2085` |
|
|
36
|
+
| **Does not cross an L0 deny** | "never allowed" things like `dd` writing to a disk / formatting / dropping a database are **deliberately not given a token** | the authorisation line **does not appear** in the reason |
|
|
37
|
+
| **Authorised only from an interactive terminal** | `guard allow` requires stdin to be a TTY; the AI running it itself gets refused | refused outright when not a TTY, and it prints the whole absolute-path command line |
|
|
38
|
+
| **Leaves an audit trail** | the allowance record keeps the danger level of the original verdict | `{source:"token", p:0.83, token:"ALLOW-…", overridden:"block"}` |
|
|
39
|
+
|
|
40
|
+
**Why this channel matters:** it is the **only human channel that does not depend on the host**. When the host has no ability to ask a human (or cannot ask), it is how a human can still "let just this one through". The cost is one copy-paste.
|
|
41
|
+
|
|
42
|
+
**That line in the reason is there for a human to copy-paste whole**: absolute path, the command's exact text **untruncated**, quotes already escaped
|
|
43
|
+
(a regression test feeds it to `bash -c 'printf %s …'` and demands byte-for-byte restoration).
|
|
44
|
+
|
|
45
|
+
---
|
|
46
|
+
|
|
47
|
+
## 2. Channel ②: host approval (stronger when you have it, not fatal when you don't)
|
|
48
|
+
|
|
49
|
+
The valve hands `escalate` to the host, and the host decides in what form to ask the human. **The same four states, the same reason**, a different shape:
|
|
50
|
+
|
|
51
|
+
| Session policy | The host's action | What the valve records | What the human does |
|
|
52
|
+
|---|---|---|---|
|
|
53
|
+
| `ask` (with approvals) | **pops the approval prompt** | `action=escalate, decision=ask, policy=ask` | click in the prompt — no terminal, no command copying |
|
|
54
|
+
| `never` (full permissions) | no prompt available → **a plain refusal** | `action=escalate, decision=deny` | go through channel ①, or run it yourself in a terminal |
|
|
55
|
+
|
|
56
|
+
**Measured (DSH, 2026-09-20):** after the user switched the session to `ask`, the same blocked command triggered a real approval —
|
|
57
|
+
the session log showed the pair `approval/asked` + `approval/decided`, and `asked.reason` was **verbatim the valve's reason**
|
|
58
|
+
(the hard rule id and why included); after the user clicked allow, `outcome=allowed-once` (2.2 seconds from pop-up to click),
|
|
59
|
+
and the command ran right away, taking the target file from 148 → 0 bytes. This channel **needs no terminal from the human**.
|
|
60
|
+
|
|
61
|
+
### 2.1 DSH's answer set is a **closed set**: there is no "always allow"
|
|
62
|
+
|
|
63
|
+
This was verified specifically (source evidence, `deepseek-harness` @ `ddefc45fbc`):
|
|
64
|
+
|
|
65
|
+
| Location | Content |
|
|
66
|
+
|---|---|
|
|
67
|
+
| `packages/interaction/user-approval/src/types.ts:32` | `ApprovalOutcome = 'allowed-once' \| 'rejected' \| 'cancelled' \| 'unavailable'` |
|
|
68
|
+
| `packages/interaction/user-approval/src/index.ts:204` | comment: `'allowed-once' is the only grant` |
|
|
69
|
+
| `packages/client/ui-approval/.../slots.ts:64` | `ApprovalDecision = 'allowed-once' \| 'rejected'` — the UI has only two buttons |
|
|
70
|
+
| `packages/session/session-format-v0-to-v1/src/payload-validation.ts:40,43` | validating outcome and policy (`ask`/`never`) are both closed sets |
|
|
71
|
+
| `packages/interaction/user-approval/tests/invariant.spec.ts:103` | a negative assertion that `policy: 'always'` **must be rejected** |
|
|
72
|
+
|
|
73
|
+
**Design implication (good news for the valve):** under `ask` mode every `escalate` is an **independent, one-at-a-time decision**,
|
|
74
|
+
and there is no standing authorisation that can silently swallow the valve's ask. So the valve **does not need** to maintain
|
|
75
|
+
state like "the user has already agreed permanently" — and therefore there is no risk of that state being bypassed or going stale.
|
|
76
|
+
|
|
77
|
+
### 2.2 Only DSH is supported (the rest are archived)
|
|
78
|
+
|
|
79
|
+
As of 2026-09-20 this package **supports DSH only** (see [`DECISIONS.md`](./DECISIONS.md) D11). The other execution channels
|
|
80
|
+
tried historically never got finished: each host's approval/trust mechanism is different, and making any one of them solid is a
|
|
81
|
+
separate round of work of its own; one of those measurements also exposed a structural defect — **information that was unavailable was pretended to be available by a default value**,
|
|
82
|
+
so "can the host ask a human" came out permanently no there, and the host approval channel was unreachable.
|
|
83
|
+
|
|
84
|
+
There is only one criterion: **does that host have a callback that execution must pass through, and that can say you may not run this?**
|
|
85
|
+
Without one you can only build a suggestion layer (the model may simply not comply), and **an unverified adapter is more dangerous than no adapter** —
|
|
86
|
+
it looks like the valve is installed while in fact nothing is blocked. Those implementations were removed along with the narrowing of scope; the
|
|
87
|
+
transferable lessons are in [`MEASUREMENTS.md`](./MEASUREMENTS.md) §12 and [`DECISIONS.md`](./DECISIONS.md) D11.
|
|
88
|
+
|
|
89
|
+
---
|
|
90
|
+
|
|
91
|
+
## 3. Channel ③: the human runs it directly (the valve is not involved)
|
|
92
|
+
|
|
93
|
+
A human can just run that command in their own terminal — that **does not go through the valve**, and it gives the AI no permission whatsoever.
|
|
94
|
+
|
|
95
|
+
**Measured (2026-09-20):** a human truncated the target file to 0 bytes in a terminal (153 → 0), and then:
|
|
96
|
+
|
|
97
|
+
- the audit log gained **zero new entries** — after precise filtering, that command still had 3 verdict records (2 escalates + 1 token allowance).
|
|
98
|
+
The valve audits **the agent's actions**, not the human's: it guards you, it does not watch you.
|
|
99
|
+
- the AI then retried **the exact same command, character for character**, was still blocked, and was shown **the same** token.
|
|
100
|
+
|
|
101
|
+
**Conclusion: `the human did it once` ≠ `a door was opened for the agent`.** The latter has to be delivered explicitly (channel ① or ②).
|
|
102
|
+
|
|
103
|
+
---
|
|
104
|
+
|
|
105
|
+
## 4. Why the valve does not need to remember "the human already agreed"
|
|
106
|
+
|
|
107
|
+
Because **not one of the three channels needs the valve to remember**:
|
|
108
|
+
|
|
109
|
+
- channel ①: the token file is the fact — read once, consumed once, deleted;
|
|
110
|
+
- channel ②: the host decides, and the valve asks afresh every time;
|
|
111
|
+
- channel ③: the valve is not on the path at all.
|
|
112
|
+
|
|
113
|
+
The value of this design is that the valve has **no** "approved list" that can be corrupted, tripped up by an expiry policy, or scrambled by concurrent writes.
|
|
114
|
+
Every judging is a clean, replayable, independent computation.
|
|
115
|
+
|
|
116
|
+
---
|
|
117
|
+
|
|
118
|
+
## 5. How to verify these three on a **new host**
|
|
119
|
+
|
|
120
|
+
The protocol steps are in [`VERIFICATION.md`](./VERIFICATION.md) **U1–U3** (they apply to any host;
|
|
121
|
+
write the conclusion back into `verification-results/` with `tools/report-result.mjs --host <host> --item U1`).
|
|
122
|
+
|
|
123
|
+
Three minimal criteria:
|
|
124
|
+
|
|
125
|
+
- **U1**: the reason for a blocked command **contains** a token; walk through "authorise → retry → let through → token gone", and `source=token` shows up in the audit.
|
|
126
|
+
- **U2**: does the host **have** a channel for asking a human? If so, does the prompt carry the valve's reason verbatim? After clicking allow, does the command actually run?
|
|
127
|
+
If not, write down "on this host only channel ① is available".
|
|
128
|
+
- **U3**: a human runs the same command by hand in a terminal — the audit gains **zero new entries**, and the AI's retry is **still blocked**.
|
|
129
|
+
|
|
130
|
+
---
|
|
131
|
+
|
|
132
|
+
## 6. A wrong guess on the record (staying honest)
|
|
133
|
+
|
|
134
|
+
On 2026-09-20, this project guessed: *"if the user picks 'always allow' in the prompt, the host might answer the valve's later asks
|
|
135
|
+
automatically, degrading one-at-a-time confirmation into a standing policy."*
|
|
136
|
+
|
|
137
|
+
**That guess does not hold** — DSH has no such option (evidence in §2.1). The one who raised it was the AI; the one who pointed out the error was **the user**.
|
|
138
|
+
The recorded entry is `Correction: DSH has no persistent/always approval` in the repository's memory.
|
|
139
|
+
|
|
140
|
+
This section is kept not as self-criticism but to make one point: **the "unverified" marks in this document are not politeness** —
|
|
141
|
+
a guess written down will be taken at face value, so either measure it, or label it.
|