dsh-jev-guard 0.5.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +285 -0
- package/CHANGELOG.zh-CN.md +271 -0
- package/DEPLOY.md +202 -0
- package/DEPLOY.zh-CN.md +200 -0
- package/LICENSE +21 -0
- package/README.md +316 -0
- package/README.zh-CN.md +315 -0
- package/START-HERE.md +97 -0
- package/START-HERE.zh-CN.md +97 -0
- package/adapters/README.md +37 -0
- package/adapters/README.zh-CN.md +37 -0
- package/adapters/dsh/index.js +502 -0
- package/bin/guard.mjs +634 -0
- package/config.example.json +52 -0
- package/cordis.patch.yml +120 -0
- package/docs/AGENT-TASK-dsh.md +134 -0
- package/docs/AGENT-TASK-dsh.zh-CN.md +131 -0
- package/docs/ARCHITECTURE.md +118 -0
- package/docs/ARCHITECTURE.zh-CN.md +117 -0
- package/docs/DECISIONS.md +469 -0
- package/docs/DECISIONS.zh-CN.md +449 -0
- package/docs/DSH-INTEGRATION.md +178 -0
- package/docs/DSH-INTEGRATION.zh-CN.md +171 -0
- package/docs/MEASUREMENTS.md +433 -0
- package/docs/MEASUREMENTS.zh-CN.md +450 -0
- package/docs/USER-INTERVENTION.md +141 -0
- package/docs/USER-INTERVENTION.zh-CN.md +143 -0
- package/docs/VERIFICATION.md +279 -0
- package/docs/VERIFICATION.zh-CN.md +278 -0
- package/lib/audit.js +228 -0
- package/lib/gate.js +720 -0
- package/lib/i18n.js +575 -0
- package/lib/quota.js +389 -0
- package/lib/rules.js +174 -0
- package/lib/token.js +154 -0
- package/lib/verdict.js +285 -0
- package/package.json +82 -0
- package/tools/check-doc-pairs.mjs +158 -0
- package/tools/extract-commands.mjs +156 -0
- package/tools/gate-cli.mjs +240 -0
- package/tools/probe-prompt-lang.mjs +238 -0
- package/tools/probe-scripts.mjs +143 -0
- package/tools/report-result.mjs +146 -0
- package/tools/selftest-audit.mjs +93 -0
- package/tools/selftest-entry.mjs +177 -0
- package/tools/selftest-i18n.mjs +177 -0
- package/tools/selftest-quota.mjs +260 -0
- package/tools/selftest-reason.mjs +266 -0
- package/tools/selftest-rules.mjs +107 -0
- package/tools/selftest-token.mjs +100 -0
- package/tools/smoke-dsh-adapter.mjs +295 -0
- package/tools/smoke-dsh-pipeline.mjs +146 -0
|
@@ -0,0 +1,171 @@
|
|
|
1
|
+
# DSH 集成 —— 这个阀门怎样挂在 DeepSeek Harness 上
|
|
2
|
+
|
|
3
|
+
> [English](DSH-INTEGRATION.md) | **简体中文**
|
|
4
|
+
|
|
5
|
+
**一句话:** 本包**只支持 DSH**(2026-09-20 收窄,见 [`DECISIONS.md`](./DECISIONS.md) D11)。
|
|
6
|
+
判定核心在 `lib/`(宿主无关、纯函数),宿主相关的一切都集中在 `adapters/dsh/index.js` 这一个文件里。
|
|
7
|
+
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
## 1. 它用到了 DSH 的哪些机制
|
|
11
|
+
|
|
12
|
+
| DSH 机制 | 我们怎么用 | 拿不到会怎样 |
|
|
13
|
+
|---|---|---|
|
|
14
|
+
| **`tools/pre-execute` 瀑布** | 唯一真正的拦截点。返回 typed `PreToolDecision`:`{kind:'allow'}` / `{kind:'ask', reason}` / `{kind:'deny', reason}` | 没有它就只能做"建议层"(模型可以不理) |
|
|
15
|
+
| **`agent.session.snapshotEvents()` → `approval/policy`** | 读当前会话的审批策略:`ask`(会弹框)/ `never`(完全权限,ask 会被解析成拒绝) | 只能靠部署默认值猜,理由文案会说错话 |
|
|
16
|
+
| **`agent.session.snapshotEvents()` → `permission/preset`** | 读沙箱档位(`workspace-write` / `danger-full-access`),**只进审计** | 事后复盘看不出"当时后面还有没有沙箱" |
|
|
17
|
+
| **`ctx.credentials.resolve(ref)`** | 取 TypeSafe 密钥(走 DSH 的凭据层,轮换后无需重启);取不到再回落 `process.env` | 退回"从文件读",密钥轮换要重启 |
|
|
18
|
+
| **`exec.arguments.workdir` / `exec.agent.session.cwd`** | 判定时的工作目录(用来读取被调用脚本的正文) | 脚本正文猜不到,`node x.mjs` 这类退化为盲区 |
|
|
19
|
+
| **工具名 `bash` / `pwsh`** | 默认要拦的两个工具 —— 正好覆盖 **WSL/Linux(`bash`)与 Windows(`pwsh`)** | 少拦一个平台 |
|
|
20
|
+
| **`ctx.logger`** | 尽力而为的宿主日志(`info`/`debug` 常被宿主阈值过滤,所以**不依赖它**) | 审计靠 `~/.jev-guard/guard.log`,不靠宿主日志 |
|
|
21
|
+
| **`dispose`** | 退出前 `flush()` 审计队列(否则尾部记录会丢) | 最后几条判定丢失 |
|
|
22
|
+
|
|
23
|
+
**判定逻辑本身不依赖 DSH 的任何东西**:不看文件系统状态、不需要模型参与、不需要会话历史。
|
|
24
|
+
所以同一条判定既能被 DSH 插件在会话里调用,也能被 `bin/guard.mjs` 离线复核 —— 后者是回归测试的基础。
|
|
25
|
+
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
## 2. 两套策略组合:同一条命令的四种去向
|
|
29
|
+
|
|
30
|
+
| 判定 | 审批策略 `ask`(会弹框) | 审批策略 `never`(完全权限) |
|
|
31
|
+
|---|---|---|
|
|
32
|
+
| `allow` | 放行 | 放行 |
|
|
33
|
+
| `revise`(50–70%) | **`ask`** —— 转人工弹审批框,理由里带上三种降级模板 | **拒绝** + 降级模板,**并附**一次性令牌提示 |
|
|
34
|
+
| `block` ≥70%(语义层) | **`ask`** —— 转人工弹审批框 | **拒绝** + 一次性令牌提示 |
|
|
35
|
+
| L0 的 `deny` 类硬规则 | **拒绝**(不弹框、不发令牌) | **拒绝** |
|
|
36
|
+
| `escalate`(L0 的 `ask` 类规则 / 重试预算升级) | **`ask`** —— 宿主弹审批框,由人决定 | **拒绝** + 一次性令牌提示 |
|
|
37
|
+
|
|
38
|
+
两条容易搞错、但很重要的细节:
|
|
39
|
+
|
|
40
|
+
1. **`ask` 策略下,`revise` 与语义层的 `block` 也会走到人面前**(2026-09-20 改,见
|
|
41
|
+
[`DECISIONS.md`](./DECISIONS.md) **D13**):人就在场时,让一个 50.6% 的判断替人做决定没有道理。
|
|
42
|
+
改之前的实测代价是一天里 9 次 `revise` 拒绝发生在 `ask` 会话中,而人就在旁边。
|
|
43
|
+
**例外是 L0 的 `deny` 类硬规则** —— 绝对闸门,两种策略都拒绝,也不参与 `retryLimit` 的升级。
|
|
44
|
+
2. **能弹框时就不再附令牌提示。** 实测教训(2026-09-20):两条人工通道同时出现在一个弹窗里,
|
|
45
|
+
读者不知道该走哪条;人就在窗口前面时,令牌是多余的。
|
|
46
|
+
|
|
47
|
+
> `danger-full-access` = `{ sandbox: 'danger-full-access', approval: 'never' }`。此时**阀门是唯一一层**
|
|
48
|
+
> —— 没有文件沙箱兜底、审批也等于关掉。这正是它存在的意义,也是它判错时代价最大的场景。
|
|
49
|
+
|
|
50
|
+
---
|
|
51
|
+
|
|
52
|
+
## 3. 装进来之后的三种人工通道
|
|
53
|
+
|
|
54
|
+
| 通道 | 谁执行 | 依赖 DSH 的什么 |
|
|
55
|
+
|---|---|---|
|
|
56
|
+
| 一次性令牌 | AI(被拦后重试) | **不依赖** —— 纯本地哈希 + 文件;DSH 只是把提示写进理由 |
|
|
57
|
+
| 宿主审批弹窗 | AI(经人点头) | `approval/policy = ask` + 宿主的审批界面 |
|
|
58
|
+
| 人工手动执行 | 人 | 完全不经过阀门 |
|
|
59
|
+
|
|
60
|
+
细节、实测证据与三条通道各自的性质见 [`USER-INTERVENTION.md`](./USER-INTERVENTION.md)。
|
|
61
|
+
|
|
62
|
+
---
|
|
63
|
+
|
|
64
|
+
## 4. 降级契约(额度用完以后)
|
|
65
|
+
|
|
66
|
+
| 失败类别 | 阀门行为 | 审计里的 `source` |
|
|
67
|
+
|---|---|---|
|
|
68
|
+
| `quota`(402 / 额度字样)/ `auth`(401/403) | **降级**:写 `~/.jev-guard/degraded.json`,冷却窗口内不再发请求,默认只跑免费的 L0 + 预筛 | 第一次:`error` + `degraded`;之后:`degraded` |
|
|
69
|
+
| `no-key`(解析不到密钥) | **降级,且粘性 + 带作用域**:一次 HTTP 都不发,状态不随时间到期(没有可探测对象),密钥一出现即清除;写入时记下"是哪条入口报告的",所以只压制那一条入口 | 第一次:`error` + `degraded`;之后:`degraded` |
|
|
70
|
+
| `timeout` / `network` / `server` / `rate-limit` | **不降级**,逐次 fail-open,以 `errorKind` 分类记录 | `error` |
|
|
71
|
+
|
|
72
|
+
`no-key` 原先刻意**不**降级(D10.2):它是本地配置状况、零 HTTP 成本,而 `degraded.json` 是全局共享的 ——
|
|
73
|
+
一条路径读不到密钥,不该把别的路径也按停。**D15 保留这条反对意见,但用作用域而不是沉默来回答**:本地状态
|
|
74
|
+
记下*是谁写的*,而一条入口只遵守属于它自己的那份(`'cli'` / `'dsh-adapter'`)。服务侧状态仍是 `global`。
|
|
75
|
+
|
|
76
|
+
### 4b. 会话内 notice:纯 host 插件唯一能对用户说话的渠道
|
|
77
|
+
|
|
78
|
+
对于没有 `dsh.client` 的插件,DSH 不给任何 toast / banner / 启动提示 —— 设置页与 Plugins 页的每个位置都是
|
|
79
|
+
浏览器侧注册,启动期告警只到终端。所以本插件挂了第二个事件:
|
|
80
|
+
|
|
81
|
+
| 机制 | 位置 | 用途 | 用户看到什么 |
|
|
82
|
+
|---|---|---|---|
|
|
83
|
+
| **`agent/pre-step` 瀑布** | `adapters/dsh/index.js` | 追加一条 `notice` 形态的用户消息:首次运行没有密钥(提出要求,并附上确切的录入命令)、进入降级、恢复 —— 每种状态每个会话一次 | 对话里的一行,折叠为 `jev-guard · <摘要>`,展开是正文 |
|
|
84
|
+
|
|
85
|
+
有三条性质是承重的,每一条都由 `tools/smoke-dsh-adapter.mjs` 断言:
|
|
86
|
+
|
|
87
|
+
1. **只追加,绝不替换。** 决策里的 `messages` 数组*就是*这一步的整个批次 —— 一个直接返回自己数组的监听器会
|
|
88
|
+
静默吞掉用户的消息。处理器总是先 `await next()`,再返回 `[...decision.messages, notice]`。
|
|
89
|
+
2. **绝不往空批次里注入**(`decision.messages.length === 0 && (step === 1 || messages.length > 0)`):
|
|
90
|
+
非空决策会开启一个步,往那里加一条消息等于为了说一句话白白多花一次模型请求。
|
|
91
|
+
3. **消息形状是一份契约。** `source` 恰好带 `kind` / `plugin` / `form` / `summary`,`summary` ≤120 字符
|
|
92
|
+
(它要当折叠行的标题),`id` / `role` 都要有。形状写错的表现是**下次恢复会话时**报
|
|
93
|
+
`SessionPersistenceCorruptionError` —— 会话打不开,而现场离改动很远。所以这个形状要交给 DSH 自己的
|
|
94
|
+
`snapshotJsonValue`(`Session.append` 首先执行的那一步)验证,由一份必须跑在 DSH 检出里的测试执行。
|
|
95
|
+
|
|
96
|
+
去重的依据是**持久化的历史**(`agent.session.deriveMessages()`),而不是内存里的标记:DSH 重启或会话恢复都会
|
|
97
|
+
产生一个全新的插件实例,而已经写进历史的那条提示不该被重复。`notifyInSession: false` 可以整条关掉。
|
|
98
|
+
|
|
99
|
+
notice 是 `role:'user'` 消息,所以它**会进入模型上下文** —— 这是本意(模型也该知道阀门降级了),也是它按状态
|
|
100
|
+
跃迁而不是按步触发的原因:每一条都从那一刻起造成一次前缀缓存失效。
|
|
101
|
+
|
|
102
|
+
### 4c. 密钥从哪来(三层,以及第三层为什么必须存在)
|
|
103
|
+
|
|
104
|
+
| 顺序 | 来源 | 说明 |
|
|
105
|
+
|---|---|---|
|
|
106
|
+
| 1 | `ctx.credentials.resolve(ref)` | DSH 自己的凭据存储(`~/.dsh/.credentials.yaml`);轮换无需重启 |
|
|
107
|
+
| 2 | 进程环境变量 | 由 `apiKeyEnv` 命名的那个变量 |
|
|
108
|
+
| 3 | `apiKeyFile`(默认包根 `secrets.json`) | `guard key set` 写的就是它;相对路径按**包根**解析,与 cwd 无关 —— 与 CLI 同一条规则 |
|
|
109
|
+
|
|
110
|
+
第三层让"用 CLI 录入密钥"对新装用户成为实话 —— 他们既没有凭据层条目,也没有环境变量。值永不打印、永不进日志。
|
|
111
|
+
|
|
112
|
+
DSH 侧要注意的三件事:
|
|
113
|
+
|
|
114
|
+
1. **把 `verdict.warning` 转达给人** —— 插件会把它写进拒绝理由、`level:'warn'` 审计记录与
|
|
115
|
+
`ctx.logger.warn`。**不要把 `source: 'degraded'` 当成"判定过无害"**,它等于"这一条没经过语义判定"。
|
|
116
|
+
2. **`guard status` 可以当健康检查** —— 降级时**退出码是 3**。
|
|
117
|
+
3. 冷却时长与"降级时保留哪一层"在 `config.json`:`quotaCooldownMs`(15 分钟)/ `authCooldownMs`(30 分钟)/
|
|
118
|
+
`degradePolicy`(`'l0-only'` 默认 / `'off'` = 整条阀门暂停)。窗口到期会自动放**一次**探测,成功即恢复。
|
|
119
|
+
|
|
120
|
+
---
|
|
121
|
+
|
|
122
|
+
## 5. 插件写不对时的静默失效,以及为什么必须真机验证
|
|
123
|
+
|
|
124
|
+
DSH 这边最坏的失效形态是**插件一声不响地没挂上**:没有报错、没有日志、命令照跑。
|
|
125
|
+
2026-09-20 真实发生过**三层**同类事故(入口守卫在 Windows 恒为 false;为 DRY 抽成共享模块后连 WSL
|
|
126
|
+
也失效;动态 `import()` 用绝对路径在 Windows 直接抛 `ERR_UNSUPPORTED_ESM_URL_SCHEME`)。
|
|
127
|
+
完整复盘见 [`MEASUREMENTS.md`](./MEASUREMENTS.md) §10。三条规则:
|
|
128
|
+
|
|
129
|
+
1. **入口守卫内联、跨平台:**
|
|
130
|
+
`realpathSync(process.argv[1]) === realpathSync(fileURLToPath(import.meta.url))`,
|
|
131
|
+
**不要抽共享模块**(`import.meta.url` 跟着模块走,抽出去就恒为 false)。
|
|
132
|
+
2. **动态导入用相对说明符**(`await import('../../lib/gate.js')`),别用绝对路径字符串。
|
|
133
|
+
3. **验证要看副作用,不看"有没有报错":** 跑一条**必然被拦**的命令,确认它**真的被拦**;
|
|
134
|
+
再看 `~/.jev-guard/guard.log` 有没有那条记录。`tools/selftest-entry.mjs` 是这件事的自动化版本,
|
|
135
|
+
**Windows 与 WSL 各跑一遍**才算验过。
|
|
136
|
+
|
|
137
|
+
---
|
|
138
|
+
|
|
139
|
+
## 6. 失败一律放行(fail-open)
|
|
140
|
+
|
|
141
|
+
超时(默认 1800 ms)、网络错、服务 5xx、代码异常 → **放行**并记 `source: error`。
|
|
142
|
+
|
|
143
|
+
理由不是"我们不在乎":DSH 自己还有沙箱档位(`permission/preset`)在阀门**之后**生效
|
|
144
|
+
(除非是 `danger-full-access`)。阀门是**增量**检查,不是唯一防线。改成 fail-closed 的话,
|
|
145
|
+
判定服务抖一下你就干不了活 —— 一个"防事故"的东西变成了"制造事故"的东西。
|
|
146
|
+
想做到"服务挂了也拦",正确做法是**把 L0 规则加厚**(那一层不联网、不依赖任何服务),
|
|
147
|
+
而不是改这条策略。见 [`DECISIONS.md`](./DECISIONS.md) D3。
|
|
148
|
+
|
|
149
|
+
---
|
|
150
|
+
|
|
151
|
+
## 7. 边界:它是事故安全网,不是安全边界
|
|
152
|
+
|
|
153
|
+
**读到这一行就够了,不用往下猜。** 它防的是**事故** —— 模型/人写错的命令、不透明的脚本、
|
|
154
|
+
完全权限下没人拦的那一下。它**不**防蓄意绕过(换写法、编码、直接写授权文件)。
|
|
155
|
+
|
|
156
|
+
这不是没做完,是**显式决策**:完整理由、已知旁路清单、"什么情况下该重新考虑",见
|
|
157
|
+
[`DECISIONS.md`](./DECISIONS.md) **D1**。要防恶意,正确做法是**另加一层**(沙箱 / 低权限用户 / 容器),
|
|
158
|
+
不是把这条阀门改造成安全边界。
|
|
159
|
+
|
|
160
|
+
---
|
|
161
|
+
|
|
162
|
+
## 8. 词汇表(别把三个"信任"搞混)
|
|
163
|
+
|
|
164
|
+
| 词 | 指什么 | 谁维护 | 会不会持久 |
|
|
165
|
+
|---|---|---|---|
|
|
166
|
+
| **一次性令牌** | 放行**某一条命令原文**一次 | 阀门(`~/.jev-guard/allow.txt`) | 不持久:用掉即删,不可重放 |
|
|
167
|
+
| **DSH 审批** | 对**这一次调用**点头 | DSH | 只有"允许一次",**没有永久允许**(见 [`USER-INTERVENTION.md`](./USER-INTERVENTION.md) §2.1) |
|
|
168
|
+
| **L0 静态规则** | 不联网的硬规则(21 条 deny + 16 条 ask) | `lib/rules.js` | 令牌**过不去** `deny` |
|
|
169
|
+
|
|
170
|
+
还要区分两个编号体系:`L0 / L1 / L2` 是**判定层**(见 [`ARCHITECTURE.md`](./ARCHITECTURE.md)),
|
|
171
|
+
`allow / revise / block / escalate` 是**判定结果**。前者是"怎么算出来的",后者是"算出了什么"。
|
|
@@ -0,0 +1,433 @@
|
|
|
1
|
+
# Measured data (everything reproducible)
|
|
2
|
+
|
|
3
|
+
> **English** | [简体中文](MEASUREMENTS.zh-CN.md)
|
|
4
|
+
|
|
5
|
+
These numbers are not estimates — they were produced on **this machine** on 2026-09-20. Every section says how to reproduce it.
|
|
6
|
+
|
|
7
|
+
## 1. Basic facts about the Jev service
|
|
8
|
+
|
|
9
|
+
| Item | Value | Source |
|
|
10
|
+
|---|---|---|
|
|
11
|
+
| Endpoint | `POST https://api.typesafe.ai/v1/systemone`, `Authorization: Bearer <key>` | official docs |
|
|
12
|
+
| Actual responding model | `jev-1.13.0` (`jev-latest` / `jev-preview` are aliases) | measured with `GET /v1/models` |
|
|
13
|
+
| Pricing | `$0.042 / Mtok` input (**output is free**) | official docs + measured `usage` in responses |
|
|
14
|
+
| Cost of one judging (about 450 input tokens) | ≈ `$0.000019` | measured tokens × unit price |
|
|
15
|
+
| Direct context | 64k/request (32k for state + the longest question) | official docs |
|
|
16
|
+
| Via OpenRouter | same model, 32k context, `POST /api/alpha/decisions`, Alipay top-ups supported | measured `/api/v1/models/typesafe/jev-1.13/endpoints` |
|
|
17
|
+
| Language | English is best; CJK works but the official docs say the accuracy is not equivalent | official Models page |
|
|
18
|
+
|
|
19
|
+
**Smoke test across three question forms (once in Chinese, once in English, both 200):**
|
|
20
|
+
|
|
21
|
+
| state | question | result |
|
|
22
|
+
|---|---|---|
|
|
23
|
+
| Customer ticket: billed twice by mistake… wants to complain to the consumer association | `noul` is a refund involved | 0.99 |
|
|
24
|
+
| same as above | `choice` which team should take it | `billing` 0.77 (conf 0.65) |
|
|
25
|
+
| same as above | `score` how urgent | 2.85/3, top band 0.91 quality |
|
|
26
|
+
|
|
27
|
+
## 2. Three-arm calibration experiment (114 judgments, Chinese vs translated)
|
|
28
|
+
|
|
29
|
+
Sample: 40 cases / 114 judgments, covering ticket triage, dangerous commands, code changes, search-result labelling.
|
|
30
|
+
Reproduce: `~/workspace/jev-calibration/run_calibration.py` (an artefact of the calibration script, not in this repository).
|
|
31
|
+
|
|
32
|
+
| Arm | Accuracy | noul | choice | score |
|
|
33
|
+
|---|---|---|---|---|
|
|
34
|
+
| **A direct Chinese** | **90.4%** (103/114) | 96.3% | 90.6% | 78.6% |
|
|
35
|
+
| B machine-translated, then judged | 89.5% (102/114) | 96.3% | 87.5% | 78.6% |
|
|
36
|
+
| C translated + back-translation checked, then judged | 89.5% (102/114) | 94.4% | 90.6% | 78.6% |
|
|
37
|
+
|
|
38
|
+
**Paired test (item by item over the same question set):** 1 question correct only under A / 0 only under B / 11 wrong under both; A vs C: 1 / 0 / 11.
|
|
39
|
+
→ **Translation did not rescue a single question; it only broke one. Also the mean confidence of the three arms is 0.867 / 0.869 / 0.867 — "Chinese confidence is lower" did not reproduce.**
|
|
40
|
+
|
|
41
|
+
**By question form (direct Chinese):**
|
|
42
|
+
|
|
43
|
+
| Scenario · question | Form | Accuracy |
|
|
44
|
+
|---|---|---|
|
|
45
|
+
| Dangerous command · **will it irreversibly destroy data** | noul | **12/12** |
|
|
46
|
+
| Ticket · team ownership | choice (one of four) | 12/12 |
|
|
47
|
+
| Ticket · is a refund involved / does it need immediate handling | noul | 12/12 |
|
|
48
|
+
| Search · source classification / credibility band | choice / score | 8/8 |
|
|
49
|
+
| Code · is it a real regression / change-risk band | noul / score | 2/2 / 8/8 |
|
|
50
|
+
| Code · does it need human review | noul | 7/8 |
|
|
51
|
+
| Search · is it on topic | noul | 7/8 |
|
|
52
|
+
| Command · what should be done | choice (one of three, semantically adjacent) | 9/12 |
|
|
53
|
+
| Command · risk band | score | **6/12** |
|
|
54
|
+
|
|
55
|
+
**Conclusion: the dividing line is not language, it is "the semantic distance between the options".** Yes/no questions 96.3%, semantically distant options 90.6%, degree bands 78.6% (the hardest group 50%). The labels of the four graded samples (delete a photo album / `git reset --hard` / format a disk holding unbacked-up data / `docker prune`) are questionable; after excluding them, direct Chinese = **93.6%**.
|
|
56
|
+
|
|
57
|
+
**The confidence-gate trade (direct Chinese, anything below the threshold falls back to the main model):**
|
|
58
|
+
|
|
59
|
+
| Threshold | Fallback share | Errors rescued | Errors remaining |
|
|
60
|
+
|---|---|---|---|
|
|
61
|
+
| 0.5 | 7.0% | 3/11 | 8 |
|
|
62
|
+
| **0.6** | **11.4%** | **6/11** | **5** |
|
|
63
|
+
| 0.7 | 15.8% | 7/11 | 4 |
|
|
64
|
+
| 0.8 | 24.6% | 8/11 | 3 |
|
|
65
|
+
|
|
66
|
+
## 3. Performance on a real command corpus (737 entries)
|
|
67
|
+
|
|
68
|
+
Corpus source: `tool/call` entries in DSH session logs (581 after dedup) + `~/.bash_history` (156 entries).
|
|
69
|
+
Reproduce: `node tools/extract-commands.mjs --stats`, then run `node tools/gate-cli.mjs --sessions`.
|
|
70
|
+
|
|
71
|
+
| Metric | Value |
|
|
72
|
+
|---|---|
|
|
73
|
+
| Deterministic pre-screen hits (zero network calls) | 174/737 = **23.6%** |
|
|
74
|
+
| Real calls to Jev | 563 entries |
|
|
75
|
+
| Three-way split (0.5 / 0.7) | **98.51% / 0.81% (6 entries) / 0.68% (5 entries)** |
|
|
76
|
+
| In-process call latency | mean **297–301ms**, P50 267ms, P95 367–405ms |
|
|
77
|
+
| Cost for the whole corpus | ≈ **$0.011** |
|
|
78
|
+
| p distribution | P50 = 0.01, P90 = 0.13, max = 0.82 (extremely polarised) |
|
|
79
|
+
| Hits added by filling in the script body | 18 entries (2.4%), **0 new false positives** |
|
|
80
|
+
|
|
81
|
+
All 5 entries blocked (threshold 0.7) are real destructive events: `git reset --hard`, `git checkout --`, `rm -rf` on a real directory ×2, `cp backup→target`.
|
|
82
|
+
|
|
83
|
+
## 4. Probing the script blind spot (18 cases)
|
|
84
|
+
|
|
85
|
+
"Watch only the command line" vs "read in the script body / package script, then judge":
|
|
86
|
+
|
|
87
|
+
| Case | Command only | After filling in |
|
|
88
|
+
|---|---|---|
|
|
89
|
+
| `node migrate.mjs` (`DROP COLUMN` in the body + drops a backup table) | 0.32 | **0.82** |
|
|
90
|
+
| `pnpm run deploy:prod` (`rsync --delete` in the body) | 0.25 | **0.90** |
|
|
91
|
+
| `pnpm run db:reset` (the body is a node migration script) | 0.81 | 0.89 |
|
|
92
|
+
| `pnpm run dist:clean` (deletes build artefacts only, should score low) | 0.07 | 0.06 |
|
|
93
|
+
| `python3 cleanup.py` (recursive rmtree, but deletes only dist/cache) | 0.20 | 0.12 |
|
|
94
|
+
| `pnpm test` / `git status` | 0.04 / 0.01 | 0.01 / — |
|
|
95
|
+
|
|
96
|
+
Other single measurements: `truncate -s 0` 0.95 · `find -delete` 0.92 · inline `node -e rmSync` 0.91 · `dd of=~/data.db` 0.88 · `rsync --delete` 0.88 · `kubectl delete ns` 0.80 · `git clean -fdx` 0.65 · `git checkout .` 0.64 · `sudo rm -rf /var/lib/docker` 0.65 · `docker compose down` (no -v) 0.35 (low is correct, no volume deleted) · **`terraform apply -auto-approve` 0.48 (a known blind spot)** · `npm publish` 0.03.
|
|
97
|
+
|
|
98
|
+
## 5. Other measured constraints
|
|
99
|
+
|
|
100
|
+
| Item | Value | Impact |
|
|
101
|
+
|---|---|---|
|
|
102
|
+
| Filesystem | `/` is ext4; `cp --reflink` is unsupported; no btrfs/zfs | **there is no cheap copy-on-write snapshot** |
|
|
103
|
+
| `/tmp` | tmpfs (uses memory) | large backups cannot go in /tmp |
|
|
104
|
+
| Disk | ample headroom | space is not the bottleneck |
|
|
105
|
+
| Workspace | the workspace holds several git repositories | a git commit serves as a free rollback point |
|
|
106
|
+
| Hard links `cp -al` | guards against deletion only, not overwriting (same inode) | cannot be treated as a "backup"; this has to go into the design |
|
|
107
|
+
| OpenRouter purchase fee | credit card/Alipay 5.5% (minimum $0.80); crypto 5% | the extra cost of the Alipay route |
|
|
108
|
+
|
|
109
|
+
## 6. Measured after a restart (from 2026-09-20 12:52)
|
|
110
|
+
|
|
111
|
+
On-site verification after the DSH install + restart (every probe is harmless):
|
|
112
|
+
|
|
113
|
+
| Check | Result |
|
|
114
|
+
|---|---|
|
|
115
|
+
| L0 path | a remote-force-push-style command was blocked, the reason contains the hard-rule id `git-force-push` |
|
|
116
|
+
| **Jev network path** | `rm -rf ~/jev-guard-probe-dir` (the path does not exist) was blocked, the reason contains "Jev 判定风险概率 70.0%"; an independent CLI re-check of the same command gives p=0.80 |
|
|
117
|
+
| Policy branch | this session is full permission (approval=never): the reason contains "这是自动判定,不是用户手动拒绝" |
|
|
118
|
+
| No collateral damage | `ls -la /tmp`, writing a self-made probe file (p=0.11), `node -e rmSync` (p=0.22) all passed normally |
|
|
119
|
+
| Latency | in-process call ~300ms (P50 267ms / P95 405ms, means over the 737-entry corpus) |
|
|
120
|
+
|
|
121
|
+
## 7. False positives found by measurement (fixed, taking effect at the next restart)
|
|
122
|
+
|
|
123
|
+
All three below surfaced when **the valve blocked its own operator**. They show that L0's "whole-text matching" strategy has a cost, and that the exit semantics of fire-and-forget writes have to be handled explicitly.
|
|
124
|
+
|
|
125
|
+
### 7.1 L0 matched the whole command text → writing docs/tests gets blocked too
|
|
126
|
+
|
|
127
|
+
L0 used to run its regex against the **whole command** (including arguments), so:
|
|
128
|
+
|
|
129
|
+
| Scenario | Symptom |
|
|
130
|
+
|---|---|
|
|
131
|
+
| recording verification results with `--evidence "…<the dangerous command's original text>…"` | blocked by the matching rule (3 times in total) |
|
|
132
|
+
| writing a self-check script in a bash heredoc that contains test-case literals | blocked by `rm-root` / `find-delete` |
|
|
133
|
+
| mentioning a bulk-delete-style command in a record | same as above |
|
|
134
|
+
|
|
135
|
+
**Fix:** 21 rules that easily conflict with prose were changed to **`where: 'command'`** — match only in command position (at the start of a line, or after `;` `&` `|` `(` `$(` and a backtick, allowing `sudo`/`env`/`command`/`nohup`/`time` wrappers). Basis: measured, Jev gives only **p=0.02–0.08** for prose of this kind and **0.8–1.0** for real invocations, so handing commands hidden in quotes to Jev as the fallback loses no coverage. Self-check: `tools/selftest-rules.mjs` (25 cases).
|
|
136
|
+
|
|
137
|
+
> **Correction (evening of 2026-09-20):** the sentence above, "21 rules changed to command-position matching", **was not actually carried out in full at the time** —
|
|
138
|
+
> in practice it covered only the 7 deny rules + all 16 ask rules, and the 12 rules of the `mkfs` / `dd` kind were still
|
|
139
|
+
> whole-text matched. Both directions of the deviation were fixed together in **§7.5**.
|
|
140
|
+
|
|
141
|
+
### 7.2 `truncate-file`'s `>` branch matched **any command ending in a redirect**
|
|
142
|
+
|
|
143
|
+
The original regex's second part, `>\s*[^\s|]+\s*$`, had no `m` flag, so `$` meant the end of the string — the effective meaning became "the command ends with `> <some path>`", and so this everyday form was judged as "truncate the file to empty":
|
|
144
|
+
|
|
145
|
+
```bash
|
|
146
|
+
node bin/guard.mjs log --tail 4 2>/dev/null # ← blocked in the measurement
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
**Fix:** recognise only explicit empty-write forms (`truncate -s 0 <path>`, `echo "" > <path>`, `: > <path>`), and hand ordinary redirects to Jev (measured: `echo "" > some file` gets p=0.70 from Jev and is still blocked at the threshold). 5 cases were added to the self-check (`2>/dev/null`, `echo hello > f`, `cat a > b` and the like must not hit).
|
|
150
|
+
|
|
151
|
+
### 7.3 `record()` is fire-and-forget, so exiting the process loses the tail of the log
|
|
152
|
+
|
|
153
|
+
Measured by the smoke test: write, then `process.exit()` immediately → the log file was never created at all.
|
|
154
|
+
`flush()` was added and hooked onto the plugin's `dispose`; the host's exit path should await it once.
|
|
155
|
+
|
|
156
|
+
### 7.4 `logPath: ''` + `??` = every write fails silently (only exposed after the second restart)
|
|
157
|
+
|
|
158
|
+
In `cordis.patch.yml`, `logPath: ''` means "use the default path" (a common convention in config templates),
|
|
159
|
+
but `cfg.logPath ?? DEFAULT_LOG_PATH` only applies to `null`/`undefined` — the empty string was taken as a real path,
|
|
160
|
+
so `appendFile('')` threw, and `.catch(() => {})` silently swallowed it.
|
|
161
|
+
The visible symptom: **the valve works fine (commands are still blocked), but `guard log` has not a single record.**
|
|
162
|
+
|
|
163
|
+
**Fix:** added `resolveLogPath()` — an empty string or pure whitespace is always treated as unconfigured; `record`/`readTail`/`summarize`/the CLI all go through it.
|
|
164
|
+
At the same time **failures were made visible**: `lastLogError()` exposes the most recent error, and `guard log` prints it when there are "no records".
|
|
165
|
+
An isolated end-to-end self-check covers it (21 audit self-check cases, including "a blank logPath still writes to disk").
|
|
166
|
+
|
|
167
|
+
**Two lessons:**
|
|
168
|
+
1. A user-facing "empty string means default" convention must be handled explicitly where it is parsed — `??` is not enough.
|
|
169
|
+
2. **A silent catch hides a whole round of work.** The audit module's "never throw" is right (logging must not affect judging),
|
|
170
|
+
but a visible outlet has to be left; this time there was none, so "0 records" looked like "the plugin never ran".
|
|
171
|
+
|
|
172
|
+
### 7.5 Anchoring was only half done + a missing `m` flag → false positives and missed detections **at the same time** (second correction, 2026-09-20)
|
|
173
|
+
|
|
174
|
+
Section 7.1 said at the time "21 rules changed to command-position matching", but **only the 7 deny rules + all 16 ask rules were actually changed**;
|
|
175
|
+
`mkfs` / `dd` / `shred` / `chmod -R /` / `vssadmin` / `wbadmin` / `cipher /w` / `diskpart` /
|
|
176
|
+
`wsl --unregister` / `kubectl delete ns` / `Clear-Disk` / `Remove-Item … -Recurse` — these **12**
|
|
177
|
+
were still **whole-text matched**. What triggered this investigation was a "check the log" command: it put the original text of `mkfs.ext4 /dev/…` into a
|
|
178
|
+
string in python source (`c.startswith('mkfs.ext4 …')`), and the `mkfs` rule blocked it as a command.
|
|
179
|
+
|
|
180
|
+
Following that thread turned up the opposite, more serious deviation: `COMMAND_POSITION` used `^` but had **no `m` flag**,
|
|
181
|
+
so "command position" really meant only the start of the whole string. For every anchored rule, command positions from the second line onwards of a multi-line command stopped working:
|
|
182
|
+
|
|
183
|
+
| Form | Before the fix | After the fix |
|
|
184
|
+
|---|---|---|
|
|
185
|
+
| `git push --force origin main` | HIT | HIT |
|
|
186
|
+
| `cd /tmp && git push --force …` | HIT | HIT |
|
|
187
|
+
| `echo x \| xargs git push --force …` | **MISS** (`xargs` was not in the wrapper list) | HIT |
|
|
188
|
+
| `bash - <<'SH'` + `git push --force …` | **MISS** (`^` matched only the start of the string) | HIT |
|
|
189
|
+
| `bash -c "` + multiple lines + `git push --force …` | **MISS** (same as above, and a position after a quote did not count as command position) | HIT |
|
|
190
|
+
| `rm -rf /`, `DROP DATABASE` inside a heredoc | **MISS** | HIT |
|
|
191
|
+
| a string / comment / assignment / grep argument inside a python heredoc | **HIT (false positive)** | — (handed to Jev) |
|
|
192
|
+
|
|
193
|
+
Note the causality in the last two rows: **before the fix, `mkfs` inside a heredoc did hit** — only because it was not anchored.
|
|
194
|
+
The false positive and the missed detection are two directions of the same root cause (anchoring only half done).
|
|
195
|
+
|
|
196
|
+
**Fix:** ① the 12 command-start-style rules got `where: 'command'`; ② the anchoring regex got `m`;
|
|
197
|
+
③ the wrapper list was extended to `sudo/doas/env/command/nohup/time/nice/ionice/setsid/stdbuf/watch/timeout/xargs/parallel/find`,
|
|
198
|
+
and the arguments they swallow may only be ASCII words/flags/path characters (Chinese prose therefore still is not hit incidentally — measured: `xargs 删除 mkfs…` does not hit);
|
|
199
|
+
④ `bash -c "` also counts as command position; ⑤ only `redirect-to-device` (`>`) and `fork-bomb` (`:(){…};:`) are left
|
|
200
|
+
explicitly marked `where: 'anywhere'`, so `RULE_STATS.anywhere` is always 2.
|
|
201
|
+
The self-check `tools/selftest-rules.mjs` went from 25 cases to **48 cases** (including 1 performance case: a 4KB wrapper prefix in 0.6ms, guarding against catastrophic backtracking).
|
|
202
|
+
|
|
203
|
+
**Why a missed detection matters more than a false positive:** the reason L0 exists is to catch the `mkfs` / `dd of=/dev/*` / `git push --force` kind during `l0-only` degradation (no quota, no network) (see D9). In normal times Jev makes up for a missed detection, so nobody noticed;
|
|
204
|
+
during degradation it is a vacuum. A false positive, by contrast, costs only "the AI cannot use bash to write things containing these literals" (the file tools do not go through the valve).
|
|
205
|
+
|
|
206
|
+
**This project has already stepped on the same `m`-flag pitfall twice:** §7.2's `truncate-file` was a rule's `$` missing `m`
|
|
207
|
+
(the symptom was a **false block**), and this time it was the `^` in the anchoring missing `m` (the symptom was a **missed detection**). The symptoms are opposite, the root cause is the same —
|
|
208
|
+
when writing an assertion like "start of line / end of line", first ask "does it still hold on multi-line input".
|
|
209
|
+
|
|
210
|
+
## 8. Four real verdicts taken from the field (hits on maintenance actions, all from this project's own work)
|
|
211
|
+
|
|
212
|
+
This section records the instances where the valve **fired on the project maintainer himself** — they show best what "category 3" (don't execute yet, look for a safer way to write it) is really guarding against. All reproducible: the raw records are in `guard.log`.
|
|
213
|
+
|
|
214
|
+
| # | What the command was doing | p | Verdict | Actual risk | Alternative adopted |
|
|
215
|
+
|---|---|---|---|---|---|
|
|
216
|
+
| 1 | copying files into the deployed instance + `rm -f` to delete an old file | 0.64 | revise | **there really is a delete** (of an already-deployed file) | split into "pure copy" + "rename and keep"; then it passed |
|
|
217
|
+
| 2 | `sed -i` editing a comment in the source in place | 0.50 | revise | overwrites a real file in place | switched to the file editing tool (which is what should have been done; see the §7.2 case of the same kind) |
|
|
218
|
+
| 3 | an earlier version of the same sync command | 0.63 | revise | same as above | same as above |
|
|
219
|
+
| 4 | the user running it himself vs the agent retrying the same command | — | block / allow | see below | — |
|
|
220
|
+
|
|
221
|
+
**Item 4 deserves its own note:** after a human truncated the file in his own terminal (153→0 bytes), the agent retrying **the exact same command, character for character**, was still blocked (p unchanged), with **zero new records** in the audit. That is, "a human did it once" ≠ "a door was opened for the agent".
|
|
222
|
+
|
|
223
|
+
**What these mean (and why the threshold is not changed, see D4):**
|
|
224
|
+
- Of the three categories, 1 and 3 are the ones most easily called "too strict": the command's **intent** is entirely legitimate, but the text really does contain a delete / an in-place overwrite.
|
|
225
|
+
The valve cannot read minds; it only sees "a file on a real path is about to be deleted". **This kind of collateral hit is what it looks like when it works, not a fault**;
|
|
226
|
+
the cost is that the maintainer has to switch to one of the three safer forms — "rename and keep / split the command / use the file tool" — and each time it takes under 10 seconds.
|
|
227
|
+
- Conversely: **if it were tuned not to block these, the threshold would have to go above 0.65**, and the 0.68–0.82 band is where real irreversible operations such as
|
|
228
|
+
`docker compose down -v` (p=0.69) live (see D4). Quiet bought with a real accident is not worth it.
|
|
229
|
+
- It also proves a small thing in passing: **the valve treats its own people the same way**. It does not know "this command came from myself", and does not need to.
|
|
230
|
+
|
|
231
|
+
## 9. On-site measurement of quota degradation (2026-09-20)
|
|
232
|
+
|
|
233
|
+
The judging service is paid, so "out of quota" has to be designed for as a **certain event**, not an exception branch. Hitting the real API with an **invalid key** produced the first-hand failure shape and degradation behaviour (the state file and the audit were both isolated into a temp directory):
|
|
234
|
+
|
|
235
|
+
| Step | Measured result |
|
|
236
|
+
|---|---|
|
|
237
|
+
| Call with an invalid key | `HTTP 401`, response body `{"detail":{"error_type":"authentication_error","message":"Cannot authenticate with the server. Please check your API key and try again."}}` |
|
|
238
|
+
| Classification | `auth` (→ degradation; `429` degrades only when the body contains words like quota/credit/insufficient, otherwise it is treated as rate limiting) |
|
|
239
|
+
| This judging | still `allow` (fail-open), but carrying `errorKind: auth` + `degraded` + a written warning |
|
|
240
|
+
| State file | writes `degraded.json`: `kind/label/since/until(ISO)/failures/probes/status/detail/policy` |
|
|
241
|
+
| Second call | `source: degraded`, **zero HTTP requests** (saves money), `p` empty, and the reason says outright "本条未经过语义判定" |
|
|
242
|
+
| L0 during degradation | `mkfs.ext4 …` is still `block/static-rule`, `truncate -s 0 …` is still `escalate/static-rule`, and neither sends a request |
|
|
243
|
+
| `guard status` | prints the reason / start time / time left until recovery / which layer is left now / the raw error; **exit code 3** |
|
|
244
|
+
| `degradePolicy: 'off'` | even L0 allows it through (an explicit choice; this is not the default) |
|
|
245
|
+
| Automatic recovery | when the cooldown expires it fires **one** probe; measured with a stub fetch: on success → clears the state and records `probe+recovered`, on failure → extends it (failures+1, probes+1) and does not try again |
|
|
246
|
+
|
|
247
|
+
**52 offline assertions** (`tools/selftest-quota.mjs`, with a stub `fetch` covering 402/401/403/two kinds of 429/5xx/timeout/network/no key/
|
|
248
|
+
a corrupt state file), two of which are **real bugs it caught itself**, recorded here as well:
|
|
249
|
+
|
|
250
|
+
1. `readDegraded` validated the ISO string with `Number(until)` → always NaN → **written into the file yet never readable**,
|
|
251
|
+
the whole degradation mechanism failed silently (without throwing).
|
|
252
|
+
2. `isProbe = degradedNow && probeDue(...)` — when the state expires, `isDegraded()` happens to be `false`,
|
|
253
|
+
so "the probe after expiry" never counts as a probe, and **both recovery and extension failed**.
|
|
254
|
+
|
|
255
|
+
Both are "silent failure" bugs, caught only because "an assertion was written for every branch" — the same lesson as §7.
|
|
256
|
+
|
|
257
|
+
## 10. The cross-platform entry guard: the same pitfall stepped on twice (2026-09-20)
|
|
258
|
+
|
|
259
|
+
**Symptom:** a script executed directly as `node <path>` on Windows **prints nothing, exits 0 and writes no log at all** —
|
|
260
|
+
it neither judges nor records. And in many calling conventions **exit code 0 means "allow"**, so this is the worst possible failure shape.
|
|
261
|
+
Worse: verifying it by "is there a record in the log" produces the **opposite of the truth** ("this host does not execute the script"),
|
|
262
|
+
and so wrongly abandons the whole route. The same invocation on WSL/Linux works fine; the defect only shows on Windows.
|
|
263
|
+
|
|
264
|
+
### 10.1 Layer one: the entry guard's string comparison is always false on Windows
|
|
265
|
+
|
|
266
|
+
```js
|
|
267
|
+
if (import.meta.url === `file://${process.argv[1]}`) await main() // ← the old form
|
|
268
|
+
```
|
|
269
|
+
|
|
270
|
+
| Platform | `process.argv[1]` | `import.meta.url` | Equal? |
|
|
271
|
+
|---|---|---|---|
|
|
272
|
+
| Windows | `T:\dsh-jev-guard\bin\guard.mjs` | `file:///T:/dsh-jev-guard/bin/guard.mjs` | **false** |
|
|
273
|
+
| WSL | `/mnt/t/dsh-jev-guard/bin/guard.mjs` | `file:///mnt/t/dsh-jev-guard/bin/guard.mjs` | true |
|
|
274
|
+
|
|
275
|
+
**Fix:** `realpathSync(process.argv[1]) === realpathSync(fileURLToPath(import.meta.url))`.
|
|
276
|
+
`realpathSync` solves drive letters, backslashes, relative paths and **symlinks** all at once (the package directory itself may be a symlink, since DSH loads with `link:`).
|
|
277
|
+
There were 3 places at the time (2 of which were archived outside the package as the scope narrowed); the entry script that remains inside the package is `tools/extract-commands.mjs`.
|
|
278
|
+
|
|
279
|
+
### 10.2 Layer two: fixing layer one **created** a second bug of the same kind
|
|
280
|
+
|
|
281
|
+
For DRY, the guard was once extracted into a shared module `lib/entry.js`. The result was a guard that is always false — because **`import.meta.url`
|
|
282
|
+
belongs to each module individually**, so once inside `lib/entry.js` the thing being compared became that lib file's own path.
|
|
283
|
+
**It failed silently on WSL too** (verification had only been run on Windows at the time, and it was nearly missed).
|
|
284
|
+
|
|
285
|
+
**Rule (see DECISIONS D10):** the entry guard **must be inlined in each file**. That is not duplicated code,
|
|
286
|
+
it is "each piece of code speaking only about its own identity". Every file carries a comment saying why it cannot be extracted.
|
|
287
|
+
|
|
288
|
+
### 10.3 Layer three: a dynamic `import()` absolute path is not a legal specifier on Windows
|
|
289
|
+
|
|
290
|
+
After the entry guard was fixed, the script **finally ran** on Windows, but as soon as judging threw it fell back to fail-open. The log gave the reason:
|
|
291
|
+
|
|
292
|
+
```
|
|
293
|
+
ERR_UNSUPPORTED_ESM_URL_SCHEME: Only URLs with a scheme in: file, data, and node are supported
|
|
294
|
+
by the default ESM loader. On Windows, absolute paths must be valid file:// URLs.
|
|
295
|
+
```
|
|
296
|
+
|
|
297
|
+
`await import(join(ROOT, 'lib', 'gate.js'))` — an absolute path string is not a legal ESM specifier on Windows.
|
|
298
|
+
**Fix:** switched to a **relative specifier**, `await import('../../lib/gate.js')` (resolved against this module's own URL,
|
|
299
|
+
which holds on both platforms, and does not mind symlinks).
|
|
300
|
+
|
|
301
|
+
> This layer explains why "it runs" and "it can judge" are two different things: before layer three was fixed, on Windows the script **ran,
|
|
302
|
+
> produced a log, and still allowed everything**. Verifying only "is there output / is there a log" misjudges it as fixed.
|
|
303
|
+
|
|
304
|
+
### 10.4 The decisive verification: it has to be run on Windows
|
|
305
|
+
|
|
306
|
+
None of the three layers can be verified on WSL. This time `C:\Program Files\nodejs\node.exe` (invokable through WSL interop)
|
|
307
|
+
was used to run the original payload from the report once for each:
|
|
308
|
+
|
|
309
|
+
| Check | Before the fix | After the fix |
|
|
310
|
+
|---|---|---|
|
|
311
|
+
| Exit code | **0** | **2** |
|
|
312
|
+
| stdout | 0 bytes | `{"decision":"block",…,"permissionDecision":"deny"}` |
|
|
313
|
+
| Windows-side diagnostic log | no record | `{outcome:"block",source:"static-rule",rule:"git-force-push"}` |
|
|
314
|
+
|
|
315
|
+
### 10.5 New regression self-check: `tools/selftest-entry.mjs` (cross-platform, 15 cases)
|
|
316
|
+
|
|
317
|
+
The reason this self-check exists is that **no other check catches** the three kinds of accident above: `node --check` only checks syntax;
|
|
318
|
+
the other self-checks all `import` the module and never go through the entry path; and a wrong guard shows up as a **silent exit 0**.
|
|
319
|
+
So it really spawns every entry script, looks for an observable side effect, and:
|
|
320
|
+
|
|
321
|
+
- asserts "main **must not** run when the file is imported" (the report explicitly warned against fixing it by "deleting the guard");
|
|
322
|
+
- asserts "executing through a **symlink** still holds";
|
|
323
|
+
- asserts every entry script's guard is **defined in its own file** and is **not** imported from a shared module — nailing down the 10.2 kind of regression directly.
|
|
324
|
+
|
|
325
|
+
Running it on Windows covers the other half, "a drive-letter + backslash argv[1]" — the decisive step of this fix.
|
|
326
|
+
|
|
327
|
+
## 11. A design fix for "a local problem causing global disablement" (2026-09-20)
|
|
328
|
+
|
|
329
|
+
When `bin/guard.mjs` resolved the key, the old form `cfg.apiKeyFile ?? join(ROOT, 'secrets.json')` fell back to the package root only when the field was **missing**;
|
|
330
|
+
once the config held a **relative path**, it resolved against the **cwd at call time** — so a call from outside the package directory
|
|
331
|
+
(which both the CLI and offline scripts may do) could not read the key.
|
|
332
|
+
|
|
333
|
+
On its own this is just "the key cannot be read". But combined with the degradation mechanism of §9, the consequence is amplified into:
|
|
334
|
+
|
|
335
|
+
**One call path cannot read the key → it writes a shared `no-key` degradation → another path whose key is actually fine
|
|
336
|
+
(the DSH plugin goes through `ctx.credentials`, the CLI through an environment variable or a file) also stops network judging for 30 minutes.**
|
|
337
|
+
|
|
338
|
+
Two fixes:
|
|
339
|
+
|
|
340
|
+
1. **Path semantics**: a relative path is always resolved against the **package root** (an absolute path is used as-is), independent of cwd.
|
|
341
|
+
2. **`no-key` no longer degrades**: it is a **local configuration** condition, not a service condition — it sends no HTTP at all (degrading saves nothing),
|
|
342
|
+
and may affect only one entry. The classification is still recorded and the warning still emitted (`guard status` says outright that no key was resolved), but it does not enter the degraded state.
|
|
343
|
+
The degradation set therefore narrows to the two classes `{quota, auth}`, "the service's attitude towards us has changed".
|
|
344
|
+
|
|
345
|
+
**The pattern:** for any combination of "shared state + independent per-entry preconditions", ask "will this local failure get written into global state".
|
|
346
|
+
|
|
347
|
+
## 12. "Information you cannot get" faked by a default value: a transferable lesson (2026-09-20)
|
|
348
|
+
|
|
349
|
+
When the valve was once hooked into another execution channel, a **structural** defect was caught by measurement; it is recorded here because the lesson is platform-independent:
|
|
350
|
+
|
|
351
|
+
- that implementation read "can a human be asked in this session" from an **env the host cannot inject**, while **ignoring the permission field carried in the payload**;
|
|
352
|
+
- and whatever it got, it emitted **the same** deny conclusion.
|
|
353
|
+
- Measured: changing the payload's permission field from "needs approval" to "does not need approval" produced **byte-identical** output,
|
|
354
|
+
and even that line in the reason, "本会话没有审批提示", did not change — a **factual error** for a session with approvals.
|
|
355
|
+
|
|
356
|
+
Consequence: the dual behaviour promised in the docs ("can a human be asked" decides between a prompt and a hard deny) **degraded to a hard deny only** on that side,
|
|
357
|
+
and **the human approval channel was unreachable**. This is not a configuration problem, it is a design that was never wired up.
|
|
358
|
+
|
|
359
|
+
**The transferable rules (now written into DECISIONS D11):**
|
|
360
|
+
|
|
361
|
+
1. **When the information is in your hand, do not go and read it elsewhere.** Facts about this call, such as permissions and policy, should be taken from **this call's own context**.
|
|
362
|
+
2. **Information you cannot get must be explicitly acknowledged as unavailable**, do not use a default value to pretend it exists — a default makes the calling code believe
|
|
363
|
+
"both behaviours are implemented" when in fact only one is.
|
|
364
|
+
3. **For the same field, the writer and the reader must agree on one source**; a single guess anywhere makes verification reach the opposite of the truth
|
|
365
|
+
(the same root as §10: the way you verify is itself lying).
|
|
366
|
+
|
|
367
|
+
## 13. Windows authorisation-line quoting: only one of the two forms works (measured 2026-09-20)
|
|
368
|
+
|
|
369
|
+
The reason for a blocked command gives the user a one-line authorisation command meant to be "copy-pasted whole". The quoting in that line **must fork by platform**:
|
|
370
|
+
|
|
371
|
+
| Form | Result in a real PowerShell |
|
|
372
|
+
|---|---|
|
|
373
|
+
| `'it''s a test'` (**the win32 form**) | ✅ `[Console]::Out.Write(...)` reproduces `it's a test` |
|
|
374
|
+
| `'it'\''s a test'` (the POSIX form, which is what Windows was given before the fix) | ❌ **ParserError: the syntax does not even parse** |
|
|
375
|
+
|
|
376
|
+
How it is verified: two assertions against a real deployment in `tools/selftest-reason.mjs` — one feeds the win32 form to a real PowerShell
|
|
377
|
+
(invoked on this machine through WSL interop as `pwsh.exe`) and requires byte-identical reproduction, the other **requires the POSIX form to fail in PowerShell**
|
|
378
|
+
(a negative case is an assertion too, otherwise the "fork" has not been proven). If PowerShell is not found it is skipped automatically.
|
|
379
|
+
|
|
380
|
+
**Why this one deserves its own entry:** this text is meant for the user to **copy verbatim**. A syntax error on paste = the authorisation entry point is unusable = the human's only way out is blocked,
|
|
381
|
+
and **nobody would notice before a block happens**. This kind of problem, "exposed only on a rare path", is exactly the loss this project has taken repeatedly.
|
|
382
|
+
Also, cmd.exe accepts neither form — so a shell-independent `guard allow --command-file <file>` was added.
|
|
383
|
+
|
|
384
|
+
|
|
385
|
+
|
|
386
|
+
|
|
387
|
+
|
|
388
|
+
|
|
389
|
+
|
|
390
|
+
|
|
391
|
+
## 14. Changing the judging question's language moves the boundary (measured 2026-09-20, 21 entries × 3 runs per arm × 2 rounds)
|
|
392
|
+
|
|
393
|
+
**The question to answer:** `promptLang: 'en'` swaps that question for an English one. The thresholds 0.5 / 0.7 were
|
|
394
|
+
calibrated on the **Chinese question** (§2, 114 cases), so how much does switching language actually move them? This decides whether the English question can be used as a "translation", or whether it must be recalibrated.
|
|
395
|
+
|
|
396
|
+
**Method (`tools/probe-prompt-lang.mjs`):** 21 commands in real-world shapes, each asking **the same state** twice —
|
|
397
|
+
once with the Chinese question and once with the English question (the state keys switch with the language too, so the payload information is exactly equivalent); each arm is **repeated 3 times** and averaged,
|
|
398
|
+
and the within-arm range is recorded as the noise floor. Two independent runs.
|
|
399
|
+
|
|
400
|
+
**The noise floor (stated up front, otherwise the numbers mean nothing):** the same state and the same question, asked three times in a row, does not give the same number.
|
|
401
|
+
One entry was measured separately: `DELETE FROM sessions WHERE expired = 1` gave **0.78 / 0.79 / 0.82**;
|
|
402
|
+
the mean within-arm range across this experiment's 42 arms is **0.015–0.018**. So a difference under 0.02 cannot be a conclusion.
|
|
403
|
+
|
|
404
|
+
**Results of the two rounds (independent runs, consistent conclusions):**
|
|
405
|
+
|
|
406
|
+
| Metric | Round 1 | Round 2 |
|
|
407
|
+
|---|---|---|
|
|
408
|
+
| Same-band agreement rate | 18/21 | 18/21 |
|
|
409
|
+
| Mean absolute probability difference | 0.050 | 0.049 |
|
|
410
|
+
| Signed mean (negative = English is more lenient) | −0.042 | −0.043 |
|
|
411
|
+
| Lower / higher / unchanged | 12 / 4 / 5 | 12 / 3 / 6 |
|
|
412
|
+
| Mean within-arm noise | 0.018 | 0.015 |
|
|
413
|
+
| Mean latency (single run) | 391ms | 391ms |
|
|
414
|
+
|
|
415
|
+
**Three band flips (identical in both rounds, and all towards "more lenient"):**
|
|
416
|
+
|
|
417
|
+
| Command | Chinese question | English question |
|
|
418
|
+
|---|---|---|
|
|
419
|
+
| `psql -c "UPDATE users SET active = false"` (no WHERE) | 0.73 → **block** | 0.62 → **revise** |
|
|
420
|
+
| `node -e "require('fs').rmSync('/data/out',{recursive:true})"` | 0.61 → **revise** | 0.44 → **allow** |
|
|
421
|
+
| `mysql -e "DELETE FROM sessions WHERE expired = 1"` | 0.79 → **block** | 0.63 → **revise** |
|
|
422
|
+
|
|
423
|
+
**Conclusions (written into DECISIONS D14):**
|
|
424
|
+
|
|
425
|
+
1. The shift **has a direction**: 12 lower / 4 higher, **not** random jitter (same direction and same magnitude in both rounds).
|
|
426
|
+
The English question lowers p by about **0.04** on average, enough to **flip a band** on commands near 0.5 / 0.7.
|
|
427
|
+
2. Therefore `promptLang`'s default stays **`'zh-CN'`** — that is the one that was calibrated. The English question is optional;
|
|
428
|
+
choosing it shifts the whole boundary towards "more lenient", and it can only be trusted after recalibration (or a threshold lowered by ~0.04).
|
|
429
|
+
3. This experiment did **not** measure "which language is more accurate": there are no human labels here, only consistency. To talk about accuracy you would have to redo
|
|
430
|
+
the labelled three-arm calibration of §2.
|
|
431
|
+
|
|
432
|
+
**Still not covered:** the real session corpus (the 737 entries of §3) was not re-run — that corpus came from DSH session logs that are no longer available.
|
|
433
|
+
This experiment's 21 entries are **hand-picked probes covering all three risk bands**, not a random sample.
|