dsh-jev-guard 0.5.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (52) hide show
  1. package/CHANGELOG.md +285 -0
  2. package/CHANGELOG.zh-CN.md +271 -0
  3. package/DEPLOY.md +202 -0
  4. package/DEPLOY.zh-CN.md +200 -0
  5. package/LICENSE +21 -0
  6. package/README.md +316 -0
  7. package/README.zh-CN.md +315 -0
  8. package/START-HERE.md +97 -0
  9. package/START-HERE.zh-CN.md +97 -0
  10. package/adapters/README.md +37 -0
  11. package/adapters/README.zh-CN.md +37 -0
  12. package/adapters/dsh/index.js +502 -0
  13. package/bin/guard.mjs +634 -0
  14. package/config.example.json +52 -0
  15. package/cordis.patch.yml +120 -0
  16. package/docs/AGENT-TASK-dsh.md +134 -0
  17. package/docs/AGENT-TASK-dsh.zh-CN.md +131 -0
  18. package/docs/ARCHITECTURE.md +118 -0
  19. package/docs/ARCHITECTURE.zh-CN.md +117 -0
  20. package/docs/DECISIONS.md +469 -0
  21. package/docs/DECISIONS.zh-CN.md +449 -0
  22. package/docs/DSH-INTEGRATION.md +178 -0
  23. package/docs/DSH-INTEGRATION.zh-CN.md +171 -0
  24. package/docs/MEASUREMENTS.md +433 -0
  25. package/docs/MEASUREMENTS.zh-CN.md +450 -0
  26. package/docs/USER-INTERVENTION.md +141 -0
  27. package/docs/USER-INTERVENTION.zh-CN.md +143 -0
  28. package/docs/VERIFICATION.md +279 -0
  29. package/docs/VERIFICATION.zh-CN.md +278 -0
  30. package/lib/audit.js +228 -0
  31. package/lib/gate.js +720 -0
  32. package/lib/i18n.js +575 -0
  33. package/lib/quota.js +389 -0
  34. package/lib/rules.js +174 -0
  35. package/lib/token.js +154 -0
  36. package/lib/verdict.js +285 -0
  37. package/package.json +82 -0
  38. package/tools/check-doc-pairs.mjs +158 -0
  39. package/tools/extract-commands.mjs +156 -0
  40. package/tools/gate-cli.mjs +240 -0
  41. package/tools/probe-prompt-lang.mjs +238 -0
  42. package/tools/probe-scripts.mjs +143 -0
  43. package/tools/report-result.mjs +146 -0
  44. package/tools/selftest-audit.mjs +93 -0
  45. package/tools/selftest-entry.mjs +177 -0
  46. package/tools/selftest-i18n.mjs +177 -0
  47. package/tools/selftest-quota.mjs +260 -0
  48. package/tools/selftest-reason.mjs +266 -0
  49. package/tools/selftest-rules.mjs +107 -0
  50. package/tools/selftest-token.mjs +100 -0
  51. package/tools/smoke-dsh-adapter.mjs +295 -0
  52. package/tools/smoke-dsh-pipeline.mjs +146 -0
@@ -0,0 +1,117 @@
1
+ # 架构:为什么是四层,以及为什么判定与拦截必须分开
2
+
3
+ > [English](ARCHITECTURE.md) | **简体中文**
4
+
5
+ ## 一句话
6
+
7
+ **可移植的是判断,不是权力。** 判定可以做成一个与调用方无关的函数(本包就是这么做的);
8
+ 而**强制拦截必须落在 DSH 自己的执行前拦截位上**。把这两件事分开,方案才成立 ——
9
+ 也是它能在 Windows 与 WSL 上跑同一套判定的原因。
10
+
11
+ ## 四层
12
+
13
+ ```
14
+ ┌──────────── jev-guard ────────────┐
15
+ │ L0 静态硬规则 不联网 · 不可覆盖 │
16
+ DSH ──原生插件────▶ │ L1 Jev 语义判定 300ms · 四态 · 缓存 │ ──▶ TypeSafe API
17
+ │ L2 执行前快照(规划中) │
18
+ (只支持 DSH) ─────▶│ 三张皮:CLI · 库 · 离线自检 │
19
+ └──────────────────────────────────┘
20
+ ```
21
+
22
+ ### L0 — 静态硬规则(不联网)
23
+
24
+ **职责:** "永不允许"和"必须人工确认"两类清单。`lib/rules.js`,21 条 deny + 16 条 ask。
25
+
26
+ **为什么单独一层:** 它必须不依赖网络、不依赖模型、不可被任何下游覆盖。DSH 里对应的是
27
+ `ctx.tools.guard()` —— 文档原话:may deny or abstain, **never force-allow**。
28
+ 拿语义模型当硬规则的实现,等于把安全建立在一句网络请求上。
29
+
30
+ ### L1 — Jev 语义判定
31
+
32
+ **职责:** 灰区判断。"这条命令/脚本会不会不可逆地删或覆盖真实数据?"
33
+
34
+ **为什么是这一个问句:** 114 项实测里,同一个问句在中文下 **12/12**;而同批样本里
35
+ "该怎么做"(语义相邻的三选一)只有 75%、"风险档位"只有 50%。**是非题准,程度题不准** ——
36
+ 所以档位由代码合成,不问模型。
37
+
38
+ **四态而不是三态:** "让模型再试一次"(秒级闭环)和"等人来"(可能几小时)是两条不同的控制流。
39
+ 混在一起要么把人拖进小事,要么让模型在同一个坑里换写法无限重试。`revise` 必须附三种确定性
40
+ 降级模板(只读等价 / 缩小作用域 / 先备份),因为它给不出命令 —— 它不产出文本。
41
+
42
+ ### L2 — 执行前快照(规划中)
43
+
44
+ **职责:** 判定不可能 100% 准,这一层兜住误判。**先只做 git provider**(覆盖率最高、代价最低),
45
+ 可枚举且体积在预算内才快照;枚举不出或超预算就直接拦。本机实测约束见 `docs/MEASUREMENTS.md` §5。
46
+
47
+ ### L3 — 人工出口
48
+
49
+ 不在这条流水线里,但属于设计的一部分:**一次性令牌**(放行某一条命令一次)、
50
+ **DSH 审批弹窗**(`approval: ask`)、**人手动执行**。三条通道各自的性质与实测证据见
51
+ [`USER-INTERVENTION.md`](./USER-INTERVENTION.md)。
52
+
53
+ ## 数据流(一次判定)
54
+
55
+ ```
56
+ 命令文本 + cwd
57
+ │
58
+ ├─ L0 命中? ──是──▶ block / escalate(结束,零网络)
59
+ │
60
+ ├─ 确定性预筛命中? ──是──▶ allow(结束,零网络)
61
+ │ (只读命令 / 可重建路径的 rm;复合命令逐段判断)
62
+ │
63
+ ├─ 缓存命中? ──是──▶ 返回上次结论(仍走当前策略映射)
64
+ │
65
+ ├─ 补齐 state:
66
+ │ · 命令里调用的脚本正文(`node x.mjs` → 读 x.mjs)
67
+ │ · package.json 里的脚本体(`pnpm run deploy:prod` → 读 scripts)
68
+ │ · 敏感路径跳过、8KB 上限
69
+ │
70
+ └─ Jev 一问 ──▶ p
71
+ p < low → allow
72
+ low ≤ p < high → revise(附模板)
73
+ p ≥ high → block
74
+ 超时/报错 → allow(fail-open)
75
+ ```
76
+
77
+ ## 宿主映射(同一个 verdict,不同宿主不同落地)
78
+
79
+ | verdict | DSH(审批策略 `ask`) | DSH(审批策略 `never`,即完全权限) |
80
+ |---|---|---|
81
+ | allow | 放行 | 放行 |
82
+ | revise | 拒绝 + 三种降级模板 | 拒绝 + 一次性令牌提示 |
83
+ | block | 拒绝 | 拒绝 + 一次性令牌提示 |
84
+ | escalate | **弹审批框**(由人决定) | **直接拒绝** + 一次性令牌提示 |
85
+
86
+ > 最后一行是这个项目里最容易踩的坑:**完全权限模式下 `ask` 不是"弹窗",而是"被拒绝"**。
87
+ > DSH 源码里 `approval: 'never'` 的定义就是 *never prompt anyone: every ask resolves rejected*。
88
+ > 如果在这种会话里直接返回 `ask`,模型会收到一句**错话**("the user rejected tool bash");
89
+ > 所以我们自己拒绝,并把真实原因(命中哪条硬规则、这不是人手动拒绝)写清楚。
90
+
91
+ ## 为什么"建议层"不能当防线
92
+
93
+ 任何"把判定塞给模型自己决定要不要问"的做法(提示词 / 一个可选工具)都**不是拦截**:
94
+
95
+ - **提示词** = 模型可以不遵守;
96
+ - **一个供模型调用的判定工具** = 模型**自己决定**调不调用。
97
+
98
+ 两者都无法拦下别的工具调用。本包因此只走**执行前拦截位**这条强制路线:
99
+ DSH 的 `tools/pre-execute`。判定本身仍然做成与调用方无关的函数,是为了能**离线复跑**,
100
+ 不是为了把权力交出去。
101
+
102
+ ## 已知盲区(设计里认下的,不是 bug)
103
+
104
+ | 盲区 | 现状 | 对策 |
105
+ |---|---|---|
106
+ | 基础设施工具 | `terraform apply -auto-approve` 实测 0.48,低于阈值会放行 | 给这类命令单独压低阈值,或加 L0 ask 规则 |
107
+ | 运行时才算出的路径 | 脚本按配置决定删哪个目录,只能靠源码猜 | 补 state 只能补正文;这类归入"枚举不出就拦" |
108
+ | 内联长代码 | `node -e` 短句可见(0.91),超长 heredoc / 变量拼接看不见 | 已知;必要时把 heredoc 也当"脚本正文"处理 |
109
+ | 非 shell 工具 | 文件写入工具、`run_code` 尚未接入 | 扩 `tools` 列表(DSH 侧改 config 即可) |
110
+ | 远端/云端状态 | `git push --force`、`npm publish`、删云资源 | 已进 L0;其余靠并行加问"会不会改远端状态" |
111
+ | 蓄意绕过 | 换壳/编码/绝对路径 | **不设防** —— 这是事故安全网,不是安全边界 |
112
+ | 授权文件被直接改写 | 有文件写入类工具的 agent 可以直接写 `~/.jev-guard/allow.txt`(等效于自己给自己发授权) | **有意保留,不堵** —— 见 [DECISIONS.md](./DECISIONS.md) **D1**;这不是 bug,是范围声明 |
113
+
114
+ > 表里最后两行是一件事的两种说法:这个阀门防的是**事故**(模型没看清、命令写顺手了、状态搞混了),
115
+ > 不是**对手**。用户对这个范围的定调原文见 D1。**看到这两行不要"顺手修掉"**:
116
+ > 那等于要求把文件写入类工具也纳入判定 + 做文件层权限,是另一个产品的工作量,
117
+ > 而收益只针对"蓄意绕过"这个本项目明确不覆盖的场景。
@@ -0,0 +1,469 @@
1
+ # Accepted design trade-offs (do not "helpfully fix" them)
2
+
3
+ > **English** | [简体中文](DECISIONS.zh-CN.md)
4
+
5
+ This file is for **whoever comes later — human or AI**: every entry below is a trade-off that was **made on purpose**, with the measurement evidence from the time.
6
+ If you think one of them "obviously shouldn't be this way", read the reasoning through first; if you want to change it, change it as an explicit decision and update this file.
7
+
8
+ ---
9
+
10
+ ## D1 · The valve is an **accident net**, not a **security boundary** (user's call, 2026-09-20)
11
+
12
+ **Decision:** We do not aim to "stop an agent that deliberately bypasses". We only prevent **accidents**: a mistyped command by a model/human, an opaque script,
13
+ the one shot nobody stops in YOLO mode.
14
+
15
+ **This is the bypass list we deliberately accept:**
16
+
17
+ | Bypass | Current state | Why we don't plug it |
18
+ |---|---|---|
19
+ | The agent writes `~/.jev-guard/allow.txt` directly with a file-writing tool | Entirely possible (the valve hangs only on the `bash`/`pwsh` tools) | Plugging it would mean bringing `write`/`edit` into judging too (the question would have to be redesigned), or doing file-level permissions — that is **another product** (the job of a sandbox/container/seccomp), not this valve's job |
20
+ | The agent rewrites the command to bypass judging (`rm -rf` → `find -delete` → `python -c shutil.rmtree`) | Jev usually still understands it, but there is no guarantee | See D2's fallback idea: rely on **layering**, not on a single point |
21
+ | Prompt injection makes the agent act maliciously | Not defended | Defending against this requires a sandbox/container/permission model |
22
+
23
+ **The user's own words (the tone-setter):** *"This bypass that still exists can be recorded in the documentation, no need to plug it — as long as this can prevent accidents, that's enough."*
24
+
25
+ **When to reconsider this:** if some day you want to lend the agent on this machine to someone else, or let untrusted input drive it.
26
+ The right move then is to **add another layer** (sandbox / low-privilege user / container), not to rebuild the valve into a security boundary.
27
+
28
+ ---
29
+
30
+ ## D2 · L0 hard rules are anchored on "command position"; commands inside quotes go to Jev
31
+
32
+ **Decision:** Except for two structural exceptions, **all** L0 rules match only at **command position**.
33
+ Command position = start of line (under the `m` flag, the start of **every line**) / after `;` `&` `|` `(` `$(` /
34
+ after the quotes of the `bash -c "` kind, where "a string is executed as a script", with a chain of wrappers allowed
35
+ (`sudo` / `timeout 30` / `xargs -0` / `nice -n 5` / `find … -exec` …).
36
+ The two exceptions are explicitly marked `where: 'anywhere'`: `redirect-to-device` (the pattern starts with `>`)
37
+ and `fork-bomb` (pure syntax `:(){…};:`) — they cannot be anchored to the concept of "command position".
38
+ `RULE_STATS.anywhere` is always 2, and `tools/selftest-rules.mjs` prints it.
39
+
40
+ **Evidence:** whole-text matching in L0 **three times** blocked the operator's own legitimate operations (writing the command text inside an argument,
41
+ writing self-check code containing test cases inside a heredoc). And measured, Jev gives these "prose-style mentions" only **p=0.02–0.08**, while for a real invocation it gives
42
+ **0.8–1.0** — so handing the prose inside quotes to it as a fallback costs no real coverage.
43
+
44
+ **Second correction, 2026-09-20 (the first pass turned out to be incomplete, and wrong in both directions):**
45
+
46
+ The first pass changed only the 7 deny rules + all 16 ask rules, while `mkfs` / `dd` / `shred` / `chmod -R` / `vssadmin` /
47
+ `wbadmin` / `cipher` / `diskpart` / `wsl --unregister` / `kubectl delete ns` / `Clear-Disk` /
48
+ `Remove-Item … -Recurse` — these **12 rules still matched the whole text**, so data inside quotes, comments, variable assignments and
49
+ **strings inside code** (measured: `c.startswith('mkfs.ext4 /dev/sdb1')`) all hit deny
50
+ — and L0's `deny` has **no one-shot token channel**, so when it is wrongly blocked a human can only go run it in a terminal themself.
51
+
52
+ More serious is the **opposite direction**: the `^` used for anchoring had no `m` flag, so "command position" actually equalled only **the start of the whole string**
53
+ (plus after a separator). The consequence was that for any anchored rule, real commands inside multi-line scripts (heredocs) and multi-line `-c` were **all missed**:
54
+
55
+ | Form | Before the fix | After the fix |
56
+ |---|---|---|
57
+ | `git push --force origin main` (single line) | HIT | HIT |
58
+ | `cd /tmp && git push --force …` (after a separator) | HIT | HIT |
59
+ | `echo x \| xargs git push --force …` | **MISS** | HIT |
60
+ | `bash - <<'SH'` + `git push --force …` (multi-line) | **MISS** | HIT |
61
+ | `bash -c "` + multi-line + `git push --force …` | **MISS** | HIT |
62
+ | `rm -rf /` / `DROP DATABASE` inside a heredoc | **MISS** | HIT |
63
+ | the raw text inside quotes/comments/assignments/code strings | **HIT (false positive)** | — (handed to Jev) |
64
+
65
+ **Why this missed detection matters:** the whole reason L0 exists is to hold the `mkfs` / `dd of=/dev/*` / `git push --force` class
66
+ when there is **no network and no quota** (`degradePolicy: 'l0-only'`) (see D9). In normal times Jev covers the missed detection, so nobody ever found out;
67
+ while degraded it is a vacuum. Ironically: before the fix, `mkfs` inside a heredoc **did hit** — only because it was not anchored.
68
+
69
+ The boundary matrix (18 forms) and the performance cases (a 4KB wrapper prefix, against catastrophic backtracking) are in
70
+ [`../tools/selftest-rules.mjs`](../tools/selftest-rules.mjs); acceptance item 7 is run against them.
71
+
72
+ ---
73
+
74
+ ## D3 · A failed judging is always **fail-open** (allow)
75
+
76
+ **Decision:** timeout, network error, service 5xx, code exception → allow, and record `source: error` in the audit log.
77
+
78
+ **Evidence:** the host's own sandbox and approval policy are still enforced before execution; the valve is only an **incremental** check. If it were changed to fail-closed,
79
+ one wobble at Typesafe and you can't get any work done. To "block even when the service is down" you should **thicken the L0 rules** (the layer that does not depend on the network),
80
+ not change this policy.
81
+
82
+ **Health signal:** the `fail-open` counter in `guard log --stats`; if it stays 0, the judging service has never quietly gone down.
83
+
84
+ ---
85
+
86
+ ## D4 · The thresholds stay 0.5 / 0.7 — solve the friction with tokens, not by lowering the bar
87
+
88
+ **Decision:** `lowThreshold: 0.5` / `highThreshold: 0.7` (measured on 737 real commands:
89
+ 98.51% allowed / 0.81% grey zone / 0.68% blocked).
90
+
91
+ **Evidence:** raising the threshold lets real destruction through — `rm -rf ~/真实目录` measured at **0.68–0.82**,
92
+ and the grey zone also contains `docker compose down -v` (0.69). Now that a **one-shot allow token** exists, "blocked" no longer equals "stuck",
93
+ so we would rather err strict. If you really want less interruption, one line in `config.json` changes it (0.6/0.8), but that is an explicit trade of risk.
94
+
95
+ ---
96
+
97
+ ## D5 · A one-shot token covers only the grey zone and blocks; it **does not cross L0**
98
+
99
+ **Decision:** a token can allow `revise`/`block`; for L0's `deny` rules it is **ineffective**, and no token is shown in the reason either.
100
+ It is bound to the full command text (whitespace collapsed), deleted once used, and the allow is written to the audit (`source: token` + `overridden`).
101
+
102
+ **Evidence:** L0 is the hard floor in the four-layer design; if a token could cross it, "never allowed" would exist in name only.
103
+
104
+ ---
105
+
106
+ ## D6 · Authorisation (`guard allow`) can only be done in an **interactive terminal**
107
+
108
+ **Decision:** `guard allow` / `--revoke` check `stdin.isTTY`; a non-TTY is refused outright and the whole line is printed, ready to copy.
109
+
110
+ **Evidence:** measured, when the agent runs the authorisation command itself it is **also judged dangerous** (p=56%) — because the destructive command is right there in
111
+ the command text. This used to be stopped by coincidence; rather than rely on coincidence, make it an explicit design point: **authorisation = a human action.**
112
+ (Note this is not a security boundary, see D1: the agent can still write a file to bypass it; it just means "the supervised party opening a back door for itself" will not **happen casually**.)
113
+
114
+ ---
115
+
116
+ ## D7 · The script body gets uploaded to TypeSafe's API
117
+
118
+ **Decision:** `inlineScripts: true` by default — read the body of the invoked script as judging input.
119
+
120
+ **Evidence:** looking at the command line alone, `node migrate.mjs` (with `DROP COLUMN` inside) scores only **p=0.32**;
121
+ with the body read in, **0.82**. This is the necessary price for "understanding opaque commands".
122
+
123
+ **Existing mitigations:** sensitive paths (`.env` / `.ssh` / `*.pem` / `*credential*` / `*secret*` / `*token*`)
124
+ are automatically skipped and not uploaded; an 8KB per-file cap; `inlineScripts: false` turns it off entirely (the price: falling back to a 0.31 blind spot).
125
+
126
+ ---
127
+
128
+ ## D8 · The package structure is "judging core + a thin adapter layer"; no DSH mechanics in `lib/`
129
+
130
+ **Decision:** `lib/` holds only the caller-independent judging core (`gate` / `rules` / `verdict` / `token` / `audit` / `quota`),
131
+ and everything DSH-related lives in the single file `adapters/dsh/index.js`.
132
+ (`package.json`'s `exports["."]` points at it, and `main` is in sync.)
133
+
134
+ **Evidence:** the way this project was first written makes it easy to assume judging is tied to DSH. But in fact judging looks at nothing belonging to the host
135
+ (it doesn't look at filesystem state, doesn't look at session history, doesn't need a model to take part), so it should have been separated anyway —
136
+ **the reason for separating was changed once, in D11**: not to take on other callers, but so that judging can be **replayed offline**
137
+ (calibration, regression and accident retrospectives all depend on this). Putting it in `lib/` would lead the next maintainer to keep adding host logic into the core.
138
+
139
+ **A verifiable form:** `grep -riE "cordis|PreToolDecision|ctx\.|approval/policy" lib/*.js` should hit only **comments**
140
+ (the comment there currently explains "why escalate turns into a rejection under full permissions").
141
+ The only host difference in the code is one boolean unrelated to DSH: `canPrompt`.
142
+
143
+ **Contract document:** at the time it was `docs/HOST-CONTRACT.md` (a host-independent contract). **It has been superseded by D11** —
144
+ what corresponds now is [`DSH-INTEGRATION.md`](./DSH-INTEGRATION.md) (DSH integration: which mechanisms are used, how the four states map, the degradation contract).
145
+ D8's "keep host logic out of `lib/`" **is still in force**, but the reason changed in D11: not to support more callers,
146
+ but so that judging can be replayed offline.
147
+
148
+ **Rollback:** this move left a `lib/index.js.bak-moved-to-adapters-dsh` on the deployed instance; after verifying on restart that the plugin loaded as usual, it was deleted.
149
+
150
+ ---
151
+
152
+ ## D9 · When the quota runs out, **degrade to L0-only and warn explicitly** (default), rather than a silent fail-open
153
+
154
+ **Decision:** the judging service is **paid**, so running out of quota is a certain event. Once a failure falls into a persistent category
155
+ (`quota` / `auth` / `no-key`), we:
156
+
157
+ 1. **Write a readable state file** `<JEV_GUARD_HOME>/degraded.json` (reason, start time, recovery time, failure count, the raw error);
158
+ 2. Within the cooldown window (**quota 15 minutes / key 30 minutes**) **send no more requests** — saving money, and saving every command from waiting on a request that is bound to fail;
159
+ 3. When the window expires, send **one** probe request: success → recover automatically (a human need do nothing); failure → extend and stay degraded;
160
+ 4. **During degradation, run only the free L0 + pre-screen by default** (`degradePolicy: 'l0-only'`); transient failures (timeout/network/5xx/429 rate limiting)
161
+ still fail open per occurrence and do **not** degrade, but are classified and recorded.
162
+
163
+ **Why D3's fail-open alone is not enough:** functionally it is not wrong (commands still run), but it is **silent** — commands keep being allowed,
164
+ the log fills up with `source: error`, and **nobody can see at a glance that the valve is no longer protecting anything**. This project has already been burned once by
165
+ a "silent failure" (the audit-log round, see MEASUREMENTS §7), so this time we don't repeat it.
166
+
167
+ **Why degradation still keeps L0 (rather than "stopping the whole valve"):** L0 and the pre-screen **cost nothing, need no network, and are deterministic**,
168
+ and they happen to cover the worst class (`mkfs` / `dd of=/dev/*` / `git push --force` / `wsl --unregister`).
169
+ Stopping them too amounts to trading "the quota is gone" for "the most dangerous class of commands loses its protection". Stopping them as well is an **explicit choice**:
170
+ `degradePolicy: 'off'` (measured and covered: with that config even `mkfs` is allowed).
171
+
172
+ > **User confirmation (2026-09-20):** in their own words, *"if the layer that costs money is not in effect, that's fine"* — that is,
173
+ > `degradePolicy: 'l0-only'` is the desired default, not a stopgap. This entry shares its root with D1: the valve's job is
174
+ > to separate "the semantic judgment that costs money" from "the deterministic rules that are free"; when the quota is gone, only the former stops.
175
+
176
+ **Where the warning shows up (the landing place of "someone must be able to find out"):**
177
+ - In the **reason of every non-allow verdict** (so the approval prompt, the rejection message and the model feedback all carry it);
178
+ - On a degraded allow, `source: degraded`, a **separate source** in the audit, not mixed into error;
179
+ - At the moment degradation starts, one audit entry with `level: 'warn'` (at most one per window, so it doesn't spam);
180
+ - The CLI's **stderr**; the DSH plugin's `ctx.logger.warn`;
181
+ - `guard status` (exit code **3** = currently degraded, usable as a health check), `guard log --stats`.
182
+
183
+ **Known trade-off:** during degradation **the semantic layer has no effect at all on new commands** — a carefully disguised destructive command will be allowed.
184
+ That is the inevitable consequence of "the quota is gone", not something code can remove; all we can do is ① keep the free layer ② let people find out quickly.
185
+ Per D1, this is not a security-boundary problem (it never defended against deliberate bypass anyway).
186
+
187
+ **Evidence:** on 2026-09-20 a real API was hit with an invalid key: `HTTP 401` → classified `auth` → degraded for 30 minutes,
188
+ the state file wrote out the real error body, the second call was `source: degraded` with **zero requests**, and `guard status` exited 3.
189
+ There are also 54 offline assertions (`tools/selftest-quota.mjs`, with a stand-in fetch covering 402/401/429/5xx/timeout/network/bad state file).
190
+
191
+ ---
192
+
193
+ ## D10 · The entry guard is **inlined**, dynamic imports use **relative specifiers**; the degradation set takes in only the two classes where "the service's attitude changed"
194
+
195
+ Both were forced out by real accidents, and **both look like "improving the code"**, so they must be written down to keep them from being changed back.
196
+
197
+ ### D10.1 The entry guard must be inlined in each file
198
+
199
+ **Decision:** the lines that decide "am I being executed as the entry point?" are **written separately in each entry script**
200
+ (the one left in the package is `tools/extract-commands.mjs`; there used to be two more, archived outside the package as the scope narrowed).
201
+
202
+ **Evidence:** the old form `import.meta.url === \`file://${process.argv[1]}\`` is **always false** on Windows
203
+ (argv1 is `T:\…`, the url is `file:///T:/…`) → the script **exits 0 immediately** after loading, and under many calling conventions
204
+ **exit code 0 means "allow"**: the valve neither blocks nor records, and nobody knows. When fixing it, DRY led the guard to be pulled into `lib/entry.js` —
205
+ and so it became **always false on WSL too**, because **`import.meta.url` is each module's own**; once it is in a shared module,
206
+ what gets compared is that file's own path.
207
+
208
+ **So this is not duplicated code, it is "each copy speaks only about its own identity".** The correct form:
209
+
210
+ ```js
211
+ function isMainModule() {
212
+ if (!process.argv[1]) return false
213
+ try {
214
+ return realpathSync(process.argv[1]) === realpathSync(fileURLToPath(import.meta.url))
215
+ } catch {
216
+ return false
217
+ }
218
+ }
219
+ ```
220
+
221
+ `realpathSync` also handles drive letters, backslashes, relative paths and symlinks. **Companion clause:** a dynamic import must use a **relative specifier**
222
+ (`await import('../../lib/gate.js')`), not `import(join(ROOT, …))` — on Windows an absolute path is not a valid ESM specifier
223
+ (`ERR_UNSUPPORTED_ESM_URL_SCHEME`), whereas a relative specifier resolves against this module's own URL, which holds on both platforms.
224
+
225
+ **Guard rail:** `tools/selftest-entry.mjs` really spawns each entry script and asserts that "the three files each define the guard and
226
+ do not import it from a shared module". **This layer only counts as verified when run on Windows** — on WSL the drive-letter problem can never be caught.
227
+
228
+ ### D10.2 Only `quota` and `auth` trigger degradation; `no-key` does not
229
+
230
+ **Decision:** the degradation set narrows to `{quota, auth}`. `no-key` is classified and warned about, but **does not write the shared `degraded.json`**.
231
+
232
+ **Evidence:** `degraded.json` is **globally shared**, while key resolution **is independent per entry**
233
+ (the DSH plugin uses `ctx.credentials`, the CLI and the offline scripts use an environment variable or a file). This amplification chain has occurred: one entry could not
234
+ read the key because the cwd was wrong → wrote a shared `no-key` degradation → **other entries with a working key also stopped network judging for 30 minutes**.
235
+ And `no-key` does not send a single HTTP request, so degrading saves nothing — it is a **local configuration** condition, not "the service's attitude toward us changed".
236
+ At the same time the trigger itself was fixed: in `bin/guard.mjs`, **a relative-path `apiKeyFile` is always resolved against the package root** (independent of cwd).
237
+
238
+ **The rule of thumb:** whenever you have "shared state + each component's own premises", first ask "can this local failure get written into global state?".
239
+
240
+ > **Partly superseded by D15 (2026-09-20).** The amplification chain described above is real, but it is now
241
+ > answered by **scope** rather than by staying silent: `no-key` *does* degrade, and the state records the
242
+ > identity of the entry that wrote it, so only that entry is suppressed. The per-entry key-resolution fix
243
+ > below still stands — and the rule of thumb above is exactly what produced the scope field.
244
+
245
+ ---
246
+
247
+ ## D11 · **DSH only** (narrowed 2026-09-20, user decision)
248
+
249
+ **Decision:** this package no longer keeps a "generic" promise. The other execution channels tried historically, together with their implementation and verification tooling,
250
+ **have been removed from this package** (on 2026-09-20 the user decided not to keep that archive), and only DSH is left:
251
+
252
+ - under `adapters/` there is only `dsh/`;
253
+ - the docs talk only about DSH's mechanisms (`docs/DSH-INTEGRATION.md` replaced the old `HOST-CONTRACT.md`);
254
+ - the acceptance checklist keeps only DSH's items (`docs/VERIFICATION.md`);
255
+ - **no file in the package mentions other tools any more** (cleaned up 2026-09-20);
256
+ - the **transferable lessons** worth keeping are written in [`MEASUREMENTS.md`](./MEASUREMENTS.md) §12, without naming specific tools.
257
+
258
+ **Evidence (it is not "we can't", it is cost-effectiveness):**
259
+
260
+ 1. **Each host's approval/trust mechanism differs**, and doing any one of them solidly is a separate round of work.
261
+ 2. **A structural defect was exposed by measurement**: on the question "can the host ask a human", that implementation **was permanently equivalent to no** —
262
+ the policy came from an env the host could not inject, it ignored the permission fields in the payload, and whatever the policy it emitted the same rejection verdict.
263
+ So "can it ask a human" degraded into "always hard-reject", and **the host-approval human channel was unreachable**. Fixing it would mean redoing that output contract.
264
+ (The transferable lesson: **information you cannot obtain must be acknowledged as unobtainable; don't use a default value to pretend it exists.**)
265
+ 3. **An unverified adapter is more dangerous than no adapter**: it looks like it has the valve installed, but in fact it doesn't block.
266
+
267
+ **What is kept:**
268
+
269
+ - `lib/` is still caller-independent — **the reason changed**: not to take on other hosts, but because judging must be **replayable offline**
270
+ (calibration, regression and accident retrospectives all depend on this). `bin/guard.mjs` and the seven self-checks depend on it.
271
+ - **Cross-platform (WSL + Windows) is unchanged** — a platform is not a host. Both platform differences are handled: both the `bash`/`pwsh` tools are in the default list;
272
+ the quotes in the authorisation line fork by platform (see D12).
273
+ - The code in the archive directory carries the cross-platform fix for the entry guard; don't lose it if it is ever revived.
274
+
275
+ **The bar for reviving any route:** first run the full acceptance on that host (including "the command really gets blocked + the log really has a record"),
276
+ then add it back to the supported list; merely "the script runs" does not count. **Files in the package must not mention them in advance** —
277
+ write them into the docs once they are actually done.
278
+
279
+ ---
280
+
281
+ ## D12 · The quotes in the authorisation line **fork by platform**; on Windows there is also a shell-independent `--command-file`
282
+
283
+ **Decision:** `shellQuote(value, platform)`: POSIX uses `'…'` + `'\''`; Windows uses `'…'` + `''` (PowerShell).
284
+ The reason text additionally points out "use PowerShell" on Windows. Also `guard allow --command-file <file>` reads the raw command from the file,
285
+ **completely bypassing the shell's quoting rules** — for cmd.exe (which understands neither form).
286
+
287
+ **Evidence (measured):**
288
+
289
+ | Form | In PowerShell |
290
+ |---|---|
291
+ | `'a''b'` (this package's win32 form) | ✅ round-trips back to `a'b` |
292
+ | `'a'\''b'` (the old POSIX form) | ❌ **ParserError**, the syntax isn't even valid |
293
+
294
+ This text is for the user to **copy verbatim**: paste it and get a syntax error = the authorisation entry is unusable = the human's only exit is blocked. And nobody would find out before
295
+ "not a single verdict has happened yet" (R2 only appears when something is blocked).
296
+
297
+ **Why `process.platform` by default:** when DSH runs on WSL, the terminal the user can paste into is usually WSL/POSIX too;
298
+ when DSH runs on Windows, that is PowerShell. "Host and terminal on the same side" is the norm.
299
+ When they are on different sides (DSH on WSL, the terminal on Windows), use the `--command-file` route, which is shell-independent.
300
+
301
+ ---
302
+
303
+ ## D13 · The judging action **forks with the approval mode**: under `ask`, `revise`/`block` go to a human, and L0 hard rules still hard-block (user's call, 2026-09-20)
304
+
305
+ **Decision:**
306
+
307
+ | Verdict | `approval: ask` (a human is present) | `approval: never` (fully automatic) |
308
+ |---|---|---|
309
+ | `revise` (50–70%) | **goes to a human approval prompt** (with the three downgrade templates) | rejected outright (+ template + token) |
310
+ | `block` ≥70% (semantic layer) | **goes to a human approval prompt** | rejected outright (+ token) |
311
+ | L0's `deny`-class hard rules | **rejected** (no prompt, no token issued) | **rejected** |
312
+ | L0's `ask`-class rules / retry-budget escalation | goes to a human approval prompt | rejected outright |
313
+
314
+ The config switches `reviseInAskMode` / `blockInAskMode` (values `'ask'` (default) / `'deny'`) leave a fallback path that needs **no code change**.
315
+
316
+ **Evidence (measured, 2026-09-20):** before the change only `escalate` looked at the approval policy; `revise`/`block` were always rejected outright.
317
+ Of the **59** `revise`/`block` rejections in that day's audit log, **9 happened in sessions with `policy=ask`** (all p in 0.50–0.63);
318
+ the commands were `cp` to a deployment directory, `sed -i`, `mkdir -p`, `git add -A && git commit` — all the operator's own maintenance actions,
319
+ **with the human right there, yet getting only a "please use a safer form"**. The user's own words: "what I actually had in mind was to block
320
+ in fully automatic mode, and in the modes that need approval turn all the blocks into an approval prompt".
321
+
322
+ This also closes **class (3) in D1's three-way classification** — "the middle state: don't run it yet, look for an alternative; **if you can't find one, wait for the user**".
323
+ "Wait for the user" needs a channel that can reach a human, and at the time only `escalate` had one.
324
+
325
+ **Why this does not make it looser:**
326
+
327
+ 1. `never` (fully automatic) is semantically the `danger-full-access` preset; in such a session DSH judges any `ask` directly as
328
+ `rejected`, so "when there is nobody to ask, the valve rejects directly" is the only correct landing.
329
+ 2. Under `ask`, **with no responder the approval fails closed** (rejected outright) — going to a human does not become an unattended automatic allow.
330
+ 3. The prompt grants only `allowed-once`; there is no "allow from now on".
331
+ 4. The frequency is low: 0.5–0.7 is only 0.81% of the 737-command corpus (about 1/125), so it won't turn prompts into noise.
332
+
333
+ **It also plugs a latent hole:** the retry budget used to escalate "any non-escalate verdict" to
334
+ escalate after the `retryLimit+1`-th attempt — including for a hard hit on a rule with an L0 `deny` rule. That meant under `ask` mode, submitting `git push --force` three times in a row would pop a prompt,
335
+ and clicking "allow" in it **crossed the hard floor** (contradicting D5's "a token does not cross L0"). Now a hard hit **does not take part** in budget escalation
336
+ (it still records `attempts` for the audit). The audit shows this path was **never triggered** before the fix (a purely latent hole).
337
+
338
+ **The price (stated plainly):** under `ask` mode, a high-scoring destructive command now **pops a prompt and waits for your click**, instead of being rejected at once — click too fast and get it wrong,
339
+ and the consequence is borne by whoever decided; the valve no longer covers for you. If you want "high scores always hard-blocked", set `blockInAskMode` to `deny`;
340
+ but **the L0 tier hard-blocks no matter how it is configured**.
341
+
342
+ ## D14 · Copy is bilingual, but **the interface language and the language of the judging question are separate**: `promptLang` still defaults to Chinese (2026-09-20)
343
+
344
+ **Decision:**
345
+
346
+ 1. All **copy meant for a human or a model** (verdict reasons, L0 rule reasons, CLI output, degradation warnings and status, the referenced script's
347
+ skip note, audit summary titles) exists in both Chinese and English, kept together in `lib/i18n.js` (the L0 rule reasons are the exception: they are written
348
+ next to the rule itself, keeping "one rule, one self-contained unit", see the typedef in `lib/rules.js`).
349
+ 2. `lang` controls the interface language, default `'auto'`: `JEV_GUARD_LANG` → `LC_ALL`/`LC_MESSAGES`/`LANG`
350
+ (**only when they name a supported language**) → otherwise `zh-CN`. The CLI also has `--lang zh-CN|en`.
351
+ **`Intl`/the system locale is deliberately not in the chain**: the first version put it last, and the real deployment tripped over it right away —
352
+ the DSH plugin runs in WSL, where `LANG=C.UTF-8` means "no preference", so `Intl` reported Node's own
353
+ `en-US` fallback, the reasons in the session quietly turned English, while the CLI on the Windows side of the same machine (whose Node reports `zh-CN`)
354
+ stayed Chinese. One valve, two languages; extremely hard to explain in a retrospective afterwards. `C`/`POSIX`/unset = **no signal**,
355
+ falling back to the project's primary language; if you want English, say so explicitly.
356
+ 3. `promptLang` controls **the one question sent to the judging service** and the state keys, and **defaults to `'zh-CN'`, independent of `lang`**.
357
+ 4. Code comments and the self-check labels in `tools/` are **not translated**: the former are read by maintainers, the latter are test-case names;
358
+ translating them would double the maintenance cost of every change, without changing a single sentence of the product's outward-facing copy.
359
+
360
+ **Why item 3 has to be pulled out on its own (this is the core of this entry):** the thresholds 0.5 / 0.7 were calibrated on the **Chinese question**
361
+ (114 cases, see MEASUREMENTS §2). The measurements in §14 show: after switching to an English question, of the 21 probes **12 had a lower p /
362
+ 4 a higher one**, the average dropping by about **0.04**, and **three commands flipped bands outright** — `DELETE ... WHERE` (block→revise),
363
+ an `UPDATE` without WHERE (block→revise), inline `node -e rmSync` (revise→allow), with the direction **all** toward the more permissive side.
364
+ Two independent runs agreed, and the noise floor is only 0.015. That is to say: switching the interface to English and casually switching the question to English as well
365
+ is equivalent to **quietly moving a measured boundary one notch toward allow**. So the two must be configured separately, defaulting to Chinese;
366
+ switching must be acknowledged as a recalibration, not a translation.
367
+
368
+ **Why not "just don't do an English question at all":** in a non-Chinese deployment, an English question is more natural for the model, and an English question is not unusable —
369
+ 18/21 agree. Making it an **explicit option with the price written down** beats hiding it or pretending it doesn't exist.
370
+
371
+ **The price (stated plainly):**
372
+
373
+ - One more catalogue to maintain; `selftest-i18n` fails outright on three things: "a key written in only one language",
374
+ "placeholders that differ between the two sides", "Chinese left over in the English".
375
+ - Automatically resolving the language from the system locale means: **on a machine with an English locale, the output turns English after the upgrade** (a behaviour change).
376
+ If you want it fixed, say `lang: "zh-CN"`.
377
+ - The judging behaviour itself is **unaffected**: L0 rules, the pre-screen, thresholds, cache keys and the promptLang default all stay the same —
378
+ `lang` only swaps copy.
379
+
380
+ 5. **The repository's documents are English-first too**: every document defaults to English, with the
381
+ Chinese kept byte-for-byte as `<name>.zh-CN.md` in the same directory and a language line at the top
382
+ of both (point 1 above was about *messages*; this one is about the documents themselves). The rules:
383
+ - **the Chinese file is a byte-for-byte copy of the original**, differing only by that one language
384
+ line; **change both together**;
385
+ - **mechanical comparison instead of trust**, run by `node tools/check-doc-pairs.mjs`: the Chinese
386
+ file must equal the baseline byte-for-byte plus that one language line, and the two sides must
387
+ agree on heading-level sequence, code-fence count, table row count, link-target set and numeric
388
+ multiset (the tool also prints every Chinese line left in the English file, so you can check that
389
+ they are all quotations);
390
+ - **quoted measurements are not translated** — log lines, command samples and Chinese corpus entries
391
+ (such as `xargs 删除 mkfs.ext4 …`) stay exactly as observed wherever they are quoted; a translated
392
+ quote would be a forged one. Any Chinese left in an English document should only ever be of this kind;
393
+ - **generated files follow the language switch**: `verification-results/SUMMARY.md` is produced by
394
+ `tools/report-result.mjs`, which defaults to English (it is a committed artifact read by people, and
395
+ regenerating it on another machine should not make its language drift), while its evidence column
396
+ remains a verbatim quote, i.e. Chinese;
397
+ - code comments and the labels inside `tools/` are still not translated (point 4). The line is:
398
+ **what a reader outside the repository sees** (GitHub visitors, users, models) → bilingual;
399
+ **what only a maintainer reads** → Chinese.
400
+
401
+ **Relationship to D2/D5/D13:** it changes no judging path. The only thing that touches the judging input is explicitly setting `promptLang: "en"`,
402
+ and that belongs to the category "you chose it yourself, and now you have the numbers" (§14).
403
+
404
+ ---
405
+
406
+ ## D15 · `no-key` degrades **stickily and scoped**; the user is told through an **in-session notice**; the key is recorded from **stdin only** (2026-09-20, user decision)
407
+
408
+ **Context.** A fresh install has no key: no credential-layer entry, no environment variable, no file. Before
409
+ this decision the valve stayed silent in that state — each gated command failed open at the semantic layer
410
+ with an `error`-class verdict, and nothing the user would ever see (DSH's logger drops plugin info-level
411
+ lines). The user asked for three things at once: a real place to record the key, a first-run demand for it,
412
+ and a degraded state that says "no valid key" out loud — while keeping the properties that make the plugin
413
+ installable at all (zero dependencies, no build step, source install with no build approval).
414
+
415
+ **Decisions.**
416
+
417
+ 1. **`no-key` is a degrading kind, and the state is sticky.** A missing key sends no HTTP request at all, so
418
+ there is nothing to probe. Unlike `quota`/`auth` — whose recovery is "ask the service once the cooldown
419
+ expires" — `no-key` cannot end with time. It ends with its **condition**: the moment a key resolves,
420
+ `evaluateCommand` clears the state and judges normally. `probeDue()` is false for sticky kinds for good,
421
+ so a keyless deployment never re-enters the judge on every command.
422
+ 2. **Degradation carries a scope.** `quota`/`auth` are service-side facts and stay `scope: 'global'` (every
423
+ entry obeys them). `no-key` is a local configuration fact, written with the **identity of the entry** that
424
+ reported it (`'cli'` / `'dsh-adapter'`), and an entry obeys only a local state whose scope is its own.
425
+ That is what retires the objection in D10.2 — the fix is the scope field, not silence.
426
+ 3. **The user is told inside the conversation.** DSH gives a host-only plugin no toast, no banner and no
427
+ startup notice: every settings/Plugins seat is claimed by a browser-side (`dsh.client`) registration, and
428
+ startup warnings reach the terminal only. The one channel that exists is injecting a `notice`-form user
429
+ message at `agent/pre-step`: it renders as a conversation row, is written into the session history, and
430
+ enters the model's context. Rejected alternatives: a slash command (its typed input is durably logged
431
+ unless `recordInput: false`, and the browser composer sees it regardless — never acceptable for a secret)
432
+ and shipping a browser half (that adds a build step, a committed bundle and React externals, giving up the
433
+ zero-build / zero-dependency property this plugin is built on). The notice fires on three transitions —
434
+ first run with no key, entering a degraded state, recovering — **one per state per session**, deduplicated
435
+ against the durable history (`session.deriveMessages()`) so a restart or a resume does not repeat it, and
436
+ it is **never** injected into an empty step batch (that would spend a whole extra model request).
437
+ `notifyInSession: false` turns it off.
438
+ 4. **The key is recorded with `guard key set`, from stdin only.** Arguments land in the shell history and in
439
+ `ps`, so `key set` accepts no value on the command line and refuses a non-TTY stdin — the same boundary
440
+ `guard allow` uses: a human at a keyboard, and that is verifiable. It writes the `apiKeyFile` in **0600**,
441
+ preserving other keys already in that file, and prints the length and the path, never the value.
442
+ `guard key status` answers "which source wins" and exits 3 when none does.
443
+ 5. **The adapter had to learn the file.** It resolved the key from `ctx.credentials` and the environment
444
+ only. Because `key set` writes a file, the CLI entry point would have been an empty promise for exactly
445
+ the users who need it most — a fresh install with neither. The adapter now resolves
446
+ `ctx.credentials` → environment → `apiKeyFile`, sharing the CLI's path rule (a relative path resolves
447
+ against the package root, independent of cwd).
448
+
449
+ **Evidence.** `tools/selftest-quota.mjs` (78 cases) covers the sticky state, the never-probe rule, scope
450
+ isolation in both directions, and the clear-on-key. `tools/smoke-dsh-adapter.mjs` (21 assertions, now
451
+ hermetic — it no longer writes into the real `~/.jev-guard/`) covers the notice being **appended** rather
452
+ than replacing, the empty-batch guard, the four-key `source` shape and the summary bound, and the file
453
+ fallback. `tools/selftest-entry.mjs` runs `guard key set` for real (including the interactive path through a
454
+ fake TTY) and asserts the file is written, parseable, 0600 on POSIX, and never echoed.
455
+
456
+ **The shape is a contract.** `source` carries exactly `kind`/`plugin`/`form`/`summary`; a fifth key is
457
+ rejected by the pre-v3 migration validator, and a malformed message surfaces as
458
+ `SessionPersistenceCorruptionError` at the **next resume** — a session that will not open, long after the
459
+ change. That shape was verified against DSH's own `snapshotJsonValue` (the step `Session.append` performs
460
+ first) from inside a DSH checkout; the CHANGELOG records why `tools/smoke-dsh-pipeline.mjs` itself could not
461
+ be run in this deployment.
462
+
463
+ **Trade-off accepted:** a notice is a `role:'user'` message, so it **enters the model's context** (intended —
464
+ the model should know the valve is degraded) and costs one prefix-cache miss from that point on. That is why
465
+ it fires per state transition, never per step.
466
+
467
+ **Rule of thumb:** when a plugin "must tell the user something" and ships no UI, first ask what the host
468
+ already renders — a conversation notice is durable, attributed and model-visible; inventing a UI surface is
469
+ a different project with a different dependency budget.