dsh-jev-guard 0.5.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +285 -0
- package/CHANGELOG.zh-CN.md +271 -0
- package/DEPLOY.md +202 -0
- package/DEPLOY.zh-CN.md +200 -0
- package/LICENSE +21 -0
- package/README.md +316 -0
- package/README.zh-CN.md +315 -0
- package/START-HERE.md +97 -0
- package/START-HERE.zh-CN.md +97 -0
- package/adapters/README.md +37 -0
- package/adapters/README.zh-CN.md +37 -0
- package/adapters/dsh/index.js +502 -0
- package/bin/guard.mjs +634 -0
- package/config.example.json +52 -0
- package/cordis.patch.yml +120 -0
- package/docs/AGENT-TASK-dsh.md +134 -0
- package/docs/AGENT-TASK-dsh.zh-CN.md +131 -0
- package/docs/ARCHITECTURE.md +118 -0
- package/docs/ARCHITECTURE.zh-CN.md +117 -0
- package/docs/DECISIONS.md +469 -0
- package/docs/DECISIONS.zh-CN.md +449 -0
- package/docs/DSH-INTEGRATION.md +178 -0
- package/docs/DSH-INTEGRATION.zh-CN.md +171 -0
- package/docs/MEASUREMENTS.md +433 -0
- package/docs/MEASUREMENTS.zh-CN.md +450 -0
- package/docs/USER-INTERVENTION.md +141 -0
- package/docs/USER-INTERVENTION.zh-CN.md +143 -0
- package/docs/VERIFICATION.md +279 -0
- package/docs/VERIFICATION.zh-CN.md +278 -0
- package/lib/audit.js +228 -0
- package/lib/gate.js +720 -0
- package/lib/i18n.js +575 -0
- package/lib/quota.js +389 -0
- package/lib/rules.js +174 -0
- package/lib/token.js +154 -0
- package/lib/verdict.js +285 -0
- package/package.json +82 -0
- package/tools/check-doc-pairs.mjs +158 -0
- package/tools/extract-commands.mjs +156 -0
- package/tools/gate-cli.mjs +240 -0
- package/tools/probe-prompt-lang.mjs +238 -0
- package/tools/probe-scripts.mjs +143 -0
- package/tools/report-result.mjs +146 -0
- package/tools/selftest-audit.mjs +93 -0
- package/tools/selftest-entry.mjs +177 -0
- package/tools/selftest-i18n.mjs +177 -0
- package/tools/selftest-quota.mjs +260 -0
- package/tools/selftest-reason.mjs +266 -0
- package/tools/selftest-rules.mjs +107 -0
- package/tools/selftest-token.mjs +100 -0
- package/tools/smoke-dsh-adapter.mjs +295 -0
- package/tools/smoke-dsh-pipeline.mjs +146 -0
package/README.md
ADDED
|
@@ -0,0 +1,316 @@
|
|
|
1
|
+
# dsh-jev-guard
|
|
2
|
+
|
|
3
|
+
[](https://awesome-dsh-plugin.com)
|
|
4
|
+
[](./LICENSE)
|
|
5
|
+
[](https://nodejs.org)
|
|
6
|
+
[](https://github.com/deepseek-ai/deepseek-harness)
|
|
7
|
+
[](https://github.com/topics/dsh-plugin)
|
|
8
|
+
[](#platform-support)
|
|
9
|
+
[](https://github.com/7starsseeker/dsh-jev-guard/tags)
|
|
10
|
+
[](https://github.com/7starsseeker/dsh-jev-guard/actions/workflows/selftest.yml)
|
|
11
|
+
[](https://github.com/7starsseeker/dsh-jev-guard/commits/main)
|
|
12
|
+
[](https://github.com/7starsseeker/dsh-jev-guard/stargazers)
|
|
13
|
+
|
|
14
|
+
> **English** | [简体中文](README.zh-CN.md) | [Changelog](CHANGELOG.md) | [Design decisions](docs/DECISIONS.md) | [Measurements](docs/MEASUREMENTS.md)
|
|
15
|
+
|
|
16
|
+
**A pre-execution safety valve for [DeepSeek Harness (DSH)](https://github.com/deepseek-ai/deepseek-harness): before a command actually runs, it asks one question — "will this irreversibly delete or overwrite your real data?"**
|
|
17
|
+
|
|
18
|
+
It judges with [TypeSafe Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) — a "System One" model that produces no prose, only a **structured decision**. Yes/no questions are what this kind of judgment measures most accurately (114-case calibration, 90.4% with the Chinese text sent as-is). It hooks into DSH's native `tools/pre-execute`, so it is a **hard interception**, not a "please be careful" nudge to the model.
|
|
19
|
+
|
|
20
|
+
```
|
|
21
|
+
command text ──▶ L0 static hard rules (offline · cannot be overridden) ──▶ pre-screen (read-only / rebuildable) ──▶ Jev semantic judgment (~300ms)
|
|
22
|
+
│ │ │
|
|
23
|
+
└───────────────┬───────┴────────────────────────┘
|
|
24
|
+
▼
|
|
25
|
+
allow / revise (with downgrade templates) / block / escalate
|
|
26
|
+
│ │ │ │
|
|
27
|
+
run it model rewrites refuse DSH approval prompt or one-shot token
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
> **中文说明** — 本文档为英文版,完整中文介绍见 [README.zh-CN.md](README.zh-CN.md)。`dsh-jev-guard` 是 DeepSeek Harness 的执行前安全阀门:每条 `bash` / `pwsh` 工具调用**执行之前**先判定一次(先离线静态规则,再付费语义模型),返回 `allow` / `revise` / `block` / `escalate` 四态之一。它会拦下不可逆的命令,能教模型更安全的写法时就给出降级模板,两者都不适用时提供一次性人工令牌;额度耗尽时**大声降级而不是静默失效**。**它是事故安全网,不是安全边界** —— 见 [Known limits](#known-limits)。
|
|
31
|
+
|
|
32
|
+
---
|
|
33
|
+
|
|
34
|
+
## Table of contents
|
|
35
|
+
|
|
36
|
+
- [What it does](#what-it-does)
|
|
37
|
+
- [Install](#install)
|
|
38
|
+
- [Languages](#languages)
|
|
39
|
+
- [Configuration](#configuration)
|
|
40
|
+
- [Behavior under the two approval policies](#behavior-under-the-two-approval-policies)
|
|
41
|
+
- [When something is blocked: the three ways a human can step in](#when-something-is-blocked-the-three-ways-a-human-can-step-in)
|
|
42
|
+
- [When credit runs out: degrade, don't fail silently](#when-credit-runs-out-degrade-dont-fail-silently)
|
|
43
|
+
- [Visibility: audit log and status](#visibility-audit-log-and-status)
|
|
44
|
+
- [Platform support](#platform-support)
|
|
45
|
+
- [Self-checks and verification](#self-checks-and-verification)
|
|
46
|
+
- [Repository layout](#repository-layout)
|
|
47
|
+
- [Security and privacy](#security-and-privacy)
|
|
48
|
+
- [Known limits](#known-limits)
|
|
49
|
+
- [Documentation](#documentation)
|
|
50
|
+
- [License](#license)
|
|
51
|
+
|
|
52
|
+
## What it does
|
|
53
|
+
|
|
54
|
+
| Verdict | Meaning | Who acts next |
|
|
55
|
+
|---|---|---|
|
|
56
|
+
| `allow` | Provably read-only, or judged safe | Run it |
|
|
57
|
+
| `revise` | Not enough evidence to run safely, but a better form very likely exists | The **model** rewrites and retries (three downgrade templates: dry run first / narrow the scope / back up first) |
|
|
58
|
+
| `block` | Would irreversibly delete or overwrite, or hit a hard rule | Refused; a human runs it by hand |
|
|
59
|
+
| `escalate` | A human must confirm (the hard rules' "always ask" class, or the retry budget is spent) | A human |
|
|
60
|
+
|
|
61
|
+
Three layers, always in this order:
|
|
62
|
+
|
|
63
|
+
1. **L0 static hard rules** (`lib/rules.js`): 21 "never allowed" + 16 "human confirmation required". **Offline and impossible to override** — not even a one-shot token gets past them. Rules match at **command position only** (the start of a line, or after `;` `&` `|` `(` `$(`, or after `bash -c "`, and skipping wrappers such as `sudo` / `timeout 30` / `xargs -0` / `find … -exec`) — so *mentioning* a dangerous command inside an argument is not a false positive, **and a real command buried in a multi-line script is not missed either**. Only two rules cannot be anchored structurally (`redirect-to-device`, `fork-bomb`) and match the whole text; the counter `RULE_STATS.anywhere` is always 2.
|
|
64
|
+
2. **Pre-screen**: commands that are provably read-only or touch only rebuildable content (caches, build output, `/tmp`) pass straight through, with **zero network calls**.
|
|
65
|
+
3. **Jev semantic judgment**: one yes/no question — *"will this command irreversibly delete or overwrite the user's real data?"* — split by two thresholds into `allow` / `revise` / `block`. Measured latency ~300ms (P50 267ms), cost ≈ `$0.000019` per call.
|
|
66
|
+
|
|
67
|
+
## Install
|
|
68
|
+
|
|
69
|
+
Requires **Node ≥ 20** (it uses the global `fetch`). **Zero runtime dependencies** — no `npm install` needed.
|
|
70
|
+
|
|
71
|
+
**Verified host version: DSH 0.1.6-alpha.2.** That is the only DSH release this plugin has been run against, and it is deliberately **not** declared as a host requirement in `package.json`: the plugin market reads that field from the npm manifest and would then block install and update on every other DSH release. Another version is therefore **untested, not forbidden**; if you run one, re-run the self-checks below.
|
|
72
|
+
|
|
73
|
+
```bash
|
|
74
|
+
# 1. Put this repository somewhere permanent, e.g. T:\dsh-jev-guard (/mnt/t/dsh-jev-guard in WSL)
|
|
75
|
+
|
|
76
|
+
# 2. Let DSH load it (fill in your own profile name)
|
|
77
|
+
dsh plugin --profile web add /mnt/t/dsh-jev-guard # on Windows: T:\dsh-jev-guard
|
|
78
|
+
|
|
79
|
+
# 3. Restart DSH (plugins are not hot-reloaded)
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
`dsh plugin` resolves `add` through pnpm, so the spec takes anything pnpm accepts. The published package is on npm — the source the plugin market installs from by preference:
|
|
83
|
+
|
|
84
|
+
```bash
|
|
85
|
+
dsh plugin --profile web add dsh-jev-guard
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
A local path is **linked**, so the plugin keeps running from your own checkout — edit a file, restart, done. To install straight from GitHub source instead, use `github:7starsseeker/dsh-jev-guard`.
|
|
89
|
+
|
|
90
|
+
None of these has anything to build (zero dependencies, no install scripts), so none raises a build-approval prompt.
|
|
91
|
+
|
|
92
|
+
**A fresh install has no key, and it says so instead of going quiet.** The first session tells you in the conversation itself that no key is configured, and until you record one the valve runs **degraded**: the free L0 hard rules and the pre-screen still work, the paid semantic layer does not. Recording a key is one command, and it is read from stdin — never from an argument, which would land in your shell history and in `ps`:
|
|
93
|
+
|
|
94
|
+
```bash
|
|
95
|
+
node bin/guard.mjs key set # paste the key, Enter; never echoed, never in shell history
|
|
96
|
+
node bin/guard.mjs key status # which source resolves, and how long it is (never the value)
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
`guard key status` exits 3 when no key resolves, so it works as a health check. There is no cooldown to wait out: the moment a key resolves, the degraded state is cleared and judging resumes on the next command.
|
|
100
|
+
|
|
101
|
+
The declaration in `package.json` is a standard DSH bundle:
|
|
102
|
+
|
|
103
|
+
```json
|
|
104
|
+
{
|
|
105
|
+
"name": "dsh-jev-guard",
|
|
106
|
+
"dsh": { "bundle": { "patch": "./cordis.patch.yml" } }
|
|
107
|
+
}
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
`cordis.patch.yml` mounts the plugin on `tools/pre-execute`, and **every tunable lives there** (you can also put them in `config.json`; precedence is `patch > config.json > built-in defaults`).
|
|
111
|
+
|
|
112
|
+
**API key** — three sources, highest precedence first: the DSH credentials layer (`ctx.credentials`, rotation needs no restart) → the environment variable `TYPESAFE_API_KEY` → a `secrets.json` **you create yourself in the package root** (`{"TYPESAFE_API_KEY": "apikey_..."}`). That file is in `.gitignore` and is deliberately **not** part of the repository or of the published package, so nobody ships one to you — the two sources above it are the ones to prefer. `node bin/guard.mjs key set` writes that file for you (mode 0600), and the DSH adapter reads it too. The key is never printed, and it is masked before anything is written to the log.
|
|
113
|
+
|
|
114
|
+
**Verify right after installing** (don't settle for "it didn't error"):
|
|
115
|
+
|
|
116
|
+
```bash
|
|
117
|
+
node bin/guard.mjs selftest # expect: all 12 checks pass (offline)
|
|
118
|
+
node bin/guard.mjs status # expect: ✅ healthy (exit code 3 while degraded)
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
Then run a command in a session that is **certain** to be blocked (e.g. `git push --force origin main`): it should be refused, and that record should appear in `node bin/guard.mjs log --tail 3`.
|
|
122
|
+
|
|
123
|
+
## Languages
|
|
124
|
+
|
|
125
|
+
Everything a human or a model reads — verdict reasons, the 37 L0 rule reasons, all CLI output, degradation warnings and `guard status` — exists in **English and Chinese**.
|
|
126
|
+
|
|
127
|
+
| Knob | What it controls | Default |
|
|
128
|
+
|---|---|---|
|
|
129
|
+
| `lang` | The language of the messages above | `'auto'` — resolves from `JEV_GUARD_LANG` → `LC_ALL`/`LC_MESSAGES`/`LANG` (only when they name a language we ship) → otherwise `zh-CN` |
|
|
130
|
+
| CLI `--lang zh-CN\|en` | The same thing, for one invocation | — |
|
|
131
|
+
| `promptLang` | The language of the question sent to Jev **and** of the state keys it reads | `'zh-CN'` |
|
|
132
|
+
|
|
133
|
+
**`promptLang` is not a translation setting — it is a judging parameter.** The 0.5 / 0.7 thresholds were calibrated against the *Chinese* question (114 labelled cases), so switching it moves a measured boundary. Measured (`tools/probe-prompt-lang.mjs`, 21 probes × 3 repeats per arm × 2 independent runs):
|
|
134
|
+
|
|
135
|
+
| | |
|
|
136
|
+
|---|---|
|
|
137
|
+
| Same band as the Chinese question | **18/21** |
|
|
138
|
+
| Mean absolute probability gap | **0.049** |
|
|
139
|
+
| Signed mean (negative = the English question is more permissive) | **−0.043** |
|
|
140
|
+
| Lower / higher / level | 12 / 4 / 5 |
|
|
141
|
+
| Repeat-sampling noise (same question, same state) | **0.015** |
|
|
142
|
+
|
|
143
|
+
Three commands changed band, **all in the permissive direction**: `UPDATE` without `WHERE` (block → revise), `DELETE … WHERE` (block → revise), inline `node -e rmSync` (revise → allow). Two runs agreed. So if you set `promptLang: 'en'`, either re-calibrate or lower both thresholds by ≈ 0.04 — see [docs/MEASUREMENTS.md](docs/MEASUREMENTS.md) §14 and **D14** in [docs/DECISIONS.md](docs/DECISIONS.md).
|
|
144
|
+
|
|
145
|
+
Code comments and the labels inside `tools/` stay Chinese on purpose: they are read by maintainers of a Chinese-first codebase, and translating them would double the upkeep of every future change without changing a single word the product says.
|
|
146
|
+
|
|
147
|
+
**Why `auto` does not look at the system locale:** it used to fall through to `Intl`, and that bit us on the first real deployment — inside WSL `LANG=C.UTF-8` means "no preference", `Intl` then reports Node's own `en-US` default, and the session suddenly spoke English while the Windows-side CLI stayed Chinese. `C`/`POSIX`/unset counts as *absence of a signal* and lands on `zh-CN`; a locale that really names a language (`en_US.UTF-8`, `zh_CN.UTF-8`) still resolves on its own.
|
|
148
|
+
|
|
149
|
+
> 中文读者:完整中文介绍见 [README.zh-CN.md](README.zh-CN.md)。
|
|
150
|
+
|
|
151
|
+
## Configuration
|
|
152
|
+
|
|
153
|
+
`config.example.json` is the template; every key is listed below (the defaults live in `DEFAULTS` in `lib/gate.js`).
|
|
154
|
+
|
|
155
|
+
| Key | Default | Description |
|
|
156
|
+
|---|---|---|
|
|
157
|
+
| `model` | `jev-latest` | Alias of the judging model |
|
|
158
|
+
| `endpoint` | `https://api.typesafe.ai/v1/systemone` | Judging service endpoint |
|
|
159
|
+
| `apiKeyEnv` / `apiKeyFile` | `TYPESAFE_API_KEY` / `secrets.json` | Key sources; **relative paths resolve against the package root** (independent of cwd) |
|
|
160
|
+
| `lang` | `auto` | Language of every human/model-facing message (see [Languages](#languages)); `auto` reads the environment only, falling back to `zh-CN` |
|
|
161
|
+
| `promptLang` | `zh-CN` | Language of the question asked to Jev and of the state keys. **A judging parameter, not a translation**: see [Languages](#languages) — switching it to `en` was measured to shift ≈ −0.04 and to flip three probes toward permissive |
|
|
162
|
+
| `lowThreshold` / `highThreshold` | `0.5` / `0.7` | The four-state thresholds: `p < low` → allow; `low ≤ p < high` → revise; `p ≥ high` → block. Over 737 real commands this split 98.51% / 0.81% / 0.68% |
|
|
163
|
+
| `reviseInAskMode` / `blockInAskMode` | `ask` / `ask` | How the verdict action branches by approval policy: `ask` = hand it to a human approval prompt when approval is available (a human is right there — a 50%-confidence judgment should not be deciding for them); `deny` = fall back to a plain refusal. Under `never` (fully automatic) both stay a plain refusal. **L0 `deny`-class hard rules are unaffected by this switch** — they always block |
|
|
164
|
+
| `timeoutMs` | `1800` | Budget for a single judgment; anything over it is allowed (fail-open) |
|
|
165
|
+
| `cacheSize` | `256` | Number of cached judgments |
|
|
166
|
+
| `inlineScripts` / `maxScriptBytes` | `true` / `8192` | Read the body of an invoked script into the judgment context (measured: lifts the `node x.mjs` blind spot from 0.31 to 0.82); sensitive paths are skipped automatically |
|
|
167
|
+
| `retryLimit` | `2` | How many times the same command may be blocked before it escalates (`escalate`, i.e. handed to a human) |
|
|
168
|
+
| `tokens` / `tokenPath` | `true` / `~/.jev-guard/allow.txt` | One-shot allow tokens |
|
|
169
|
+
| `logPath` / `logMaxBytes` | `~/.jev-guard/guard.log` / 4 MiB | Shared audit log (JSONL, rotated when over the limit) |
|
|
170
|
+
| `quotaCooldownMs` / `authCooldownMs` | 15 min / 30 min | Cooldown after quota- and key-class failures |
|
|
171
|
+
| `degradePolicy` | `l0-only` | Which layer survives degradation: `l0-only` (stop only the paid semantic layer) or `off` (suspend the whole valve) |
|
|
172
|
+
| `pricePerMTok` | `0.042` | Unit price for cost estimation (USD per million input tokens; output is free per the vendor's documentation) |
|
|
173
|
+
|
|
174
|
+
## Behavior under the two approval policies
|
|
175
|
+
|
|
176
|
+
The same set of verdicts lands differently under DSH's two session policies — **this is the part people mix up most**:
|
|
177
|
+
|
|
178
|
+
| Verdict | `approval: ask` (prompts) | `approval: never` (full access / YOLO) |
|
|
179
|
+
|---|---|---|
|
|
180
|
+
| `revise` (50–70%) | **Handed to a human approval prompt** | Refused **+ downgrade templates + one-shot token hint** |
|
|
181
|
+
| `block` (≥70%, semantic layer) | **Handed to a human approval prompt** | Refused **+ one-shot token hint** |
|
|
182
|
+
| L0 `deny`-class hard rules | **Refused** (no prompt, no token) | **Refused** |
|
|
183
|
+
| L0 `ask`-class rules (`escalate`) | **DSH approval prompt**, the human decides | Refused **+ one-shot token hint** |
|
|
184
|
+
|
|
185
|
+
> **Why `ask` mode hands both the grey zone and the high scores to a human**: when a person is right there, letting a 50.6% judgment decide for them makes no sense; and under `never` there is nobody to ask, so the valve can only refuse conservatively. When the host has no responder, approval is **fail-closed**, so handing something to a human does **not** turn into automatic approval while unattended; the prompt only ever offers `allowed-once`, never a lasting bypass.
|
|
186
|
+
|
|
187
|
+
> **L0 `deny`-class hard rules are an absolute gate**: they block under both policies, and even `retryLimit`'s "retry enough and a human handles it" does not apply to them — otherwise one click on "allow" in a prompt would bypass the hard floor (a token cannot cross L0, and neither can an approval).
|
|
188
|
+
|
|
189
|
+
> `danger-full-access` = `{ sandbox: 'danger-full-access', approval: 'never' }` — **no sandbox underneath, approval effectively off: the valve is the only layer left**. That is exactly why it exists, and it is also the situation where a wrong verdict costs the most. Audit records carry the sandbox preset (`preset`) and the approval policy (`policy`) for that call, so a retrospective can tell whether a sandbox was still behind it at the time.
|
|
190
|
+
|
|
191
|
+
## When something is blocked: the three ways a human can step in
|
|
192
|
+
|
|
193
|
+
**① One-shot token (independent of the host, always available).** The reason on a blocked command includes an `ALLOW-XXXXXXXXXX` (the first 10 characters of the hash of the command text). **A human** pastes the whole line the reason gives them into their own terminal:
|
|
194
|
+
|
|
195
|
+
```bash
|
|
196
|
+
# WSL / Linux:
|
|
197
|
+
node /mnt/t/dsh-jev-guard/bin/guard.mjs allow '<original command>'
|
|
198
|
+
# Windows (PowerShell; quoting switches automatically per platform):
|
|
199
|
+
node T:\dsh-jev-guard\bin\guard.mjs allow '<original command>'
|
|
200
|
+
# Any platform, any shell (bypasses shell quoting rules — use this for cmd.exe):
|
|
201
|
+
node T:\dsh-jev-guard\bin\guard.mjs allow --command-file cmd.txt
|
|
202
|
+
|
|
203
|
+
node bin/guard.mjs allow --list # list outstanding tokens
|
|
204
|
+
node bin/guard.mjs allow --revoke ALLOW-… # revoke one
|
|
205
|
+
```
|
|
206
|
+
|
|
207
|
+
Four properties: **bound to the exact command text** (change one character and it is a different token), **deleted on use** (it cannot be replayed), **never crosses L0 hard rules**, and **only issued from an interactive terminal** (the AI running it itself is refused).
|
|
208
|
+
|
|
209
|
+
**② DSH approval prompt (`approval: ask`).** The valve only marks the command as "a human should look at this"; DSH shows the prompt, and the reason in it is the valve's own text. DSH's answers are a closed set (only "allow once" and "deny"), so **every decision is an independent, per-call decision** — there is no "always allow" to be silently swallowed.
|
|
210
|
+
|
|
211
|
+
**③ The human just runs it.** You run the command in your own terminal — the valve is not involved, and it **grants the AI no permission either**: your action leaves no record in the audit, and the AI retrying the same command is still blocked.
|
|
212
|
+
|
|
213
|
+
## When credit runs out: degrade, don't fail silently
|
|
214
|
+
|
|
215
|
+
The judging service is **pay-per-use**, so running out of credit is a certainty. Default behavior:
|
|
216
|
+
|
|
217
|
+
| Situation | What the valve does | Where you see it |
|
|
218
|
+
|---|---|---|
|
|
219
|
+
| Credit exhausted / key invalid (`402` / `401`) | **Degrades**: writes `~/.jev-guard/degraded.json` and stops sending requests for the cooldown window (to save money), running only the **free L0 + pre-screen** by default | `guard status` (exit code 3) · one `⚠️` line in the refusal reason · `source: degraded` and `level: warn` in the audit · stderr of the CLI |
|
|
220
|
+
| Cooldown expires | Automatically sends **one** probe request: success restores normal operation (you do nothing), failure keeps it degraded | `guard status` shows how long is left |
|
|
221
|
+
| Timeout / network hiccup / 5xx / 429 rate limit / no key | **No degradation** — each call is simply allowed (fail-open), but it is categorized and recorded | the "failure breakdown" line of `guard log --stats` |
|
|
222
|
+
|
|
223
|
+
```bash
|
|
224
|
+
node bin/guard.mjs status --clear # don't want to wait out the cooldown: retry once now (a failure re-enters degradation)
|
|
225
|
+
```
|
|
226
|
+
|
|
227
|
+
If you want "once the credit is gone, don't interfere at all" (L0 included), set `degradePolicy: "off"`.
|
|
228
|
+
|
|
229
|
+
## Visibility: audit log and status
|
|
230
|
+
|
|
231
|
+
Every judgment appends one JSONL line to `~/.jev-guard/guard.log` (rotated past 4 MiB; keys are masked before the command is written):
|
|
232
|
+
|
|
233
|
+
```bash
|
|
234
|
+
node bin/guard.mjs log --tail 20 # time / action / p / source / matched rule / command
|
|
235
|
+
node bin/guard.mjs log --stats # action·source·rule counts + fail-open + failure breakdown + cost estimate
|
|
236
|
+
node bin/guard.mjs status # one line: is the valve healthy? (exit code 3 while degraded — usable as a health check)
|
|
237
|
+
```
|
|
238
|
+
|
|
239
|
+
## Platform support
|
|
240
|
+
|
|
241
|
+
**Both WSL/Linux and Windows are supported.** Both platform differences are handled:
|
|
242
|
+
|
|
243
|
+
| Item | WSL / Linux | Windows |
|
|
244
|
+
|---|---|---|
|
|
245
|
+
| Tools intercepted | `bash` | `pwsh` (**both are in the default `tools` list**) |
|
|
246
|
+
| Quoting in the authorisation line | POSIX `'\''` | **PowerShell `''`** (the two forms are not interchangeable; the code branches by platform) |
|
|
247
|
+
| cmd.exe users | — | use `guard allow --command-file <file>` |
|
|
248
|
+
| State and log | `~/.jev-guard/` | `%USERPROFILE%\.jev-guard\` |
|
|
249
|
+
|
|
250
|
+
The self-checks assert the quoting with a **real round trip** (including the negative case: "the POSIX form must fail in PowerShell"), and the cross-platform entry-guard regression only counts as verified once it has been run **on both platforms**.
|
|
251
|
+
|
|
252
|
+
## Self-checks and verification
|
|
253
|
+
|
|
254
|
+
```bash
|
|
255
|
+
# Seven offline self-checks (no network, no API key)
|
|
256
|
+
for t in selftest-entry selftest-i18n selftest-quota selftest-reason selftest-token selftest-rules selftest-audit; do
|
|
257
|
+
printf '%-18s ' "$t"; node tools/$t.mjs | tail -1
|
|
258
|
+
done
|
|
259
|
+
|
|
260
|
+
node bin/guard.mjs selftest # 12 checks: rules / pre-screen / four-state mapping
|
|
261
|
+
node tools/smoke-dsh-adapter.mjs # adapter smoke test (fake ctx, 9 assertion groups)
|
|
262
|
+
node tools/smoke-dsh-pipeline.mjs # real tool-pipeline integration (run from inside a DSH checkout)
|
|
263
|
+
```
|
|
264
|
+
|
|
265
|
+
The acceptance checklist (20 items, including the three human channels) and the criteria for each item are in **[docs/VERIFICATION.md](docs/VERIFICATION.md)**; past results land in [`verification-results/`](verification-results/).
|
|
266
|
+
|
|
267
|
+
## Repository layout
|
|
268
|
+
|
|
269
|
+
```
|
|
270
|
+
bin/guard.mjs CLI: judge | log | status | allow | selftest | rules
|
|
271
|
+
lib/gate.js Judgment engine (L0 → pre-screen → semantic → four states) — caller-agnostic
|
|
272
|
+
lib/i18n.js Bilingual message catalog (zh-CN / en) + language resolution
|
|
273
|
+
lib/rules.js L0 static hard rules (each with an id / regex / reason)
|
|
274
|
+
lib/verdict.js Four-state composition, reason text, retry budget, platform-specific quoting
|
|
275
|
+
lib/audit.js Shared audit log (masking / rotation / stats / cost)
|
|
276
|
+
lib/token.js One-shot allow tokens
|
|
277
|
+
lib/quota.js Degradation state machine after quota/key failures
|
|
278
|
+
adapters/dsh/index.js The native DSH Cordis plugin (the only adapter)
|
|
279
|
+
cordis.patch.yml DSH bundle patch (mount declaration + every tunable)
|
|
280
|
+
tools/ Offline self-checks, smoke tests, verification helpers
|
|
281
|
+
docs/ Mechanics, trade-offs, measurements, acceptance checklist
|
|
282
|
+
```
|
|
283
|
+
|
|
284
|
+
## Security and privacy
|
|
285
|
+
|
|
286
|
+
1. **The only things sent to the judging service are the command text and, optionally, script bodies** (sensitive paths — `.env` / `.ssh` / `*.pem` / `*credential*` / `*secret*` / `*token*` — are skipped automatically, with an 8KB per-file cap). To turn this off entirely: `inlineScripts: false` (at the cost of `node x.mjs`-style commands dropping back into the p≈0.31 blind spot).
|
|
287
|
+
2. **The key is read only from the credentials layer / environment / `secrets.json`** and never enters the log or a report (the command text is masked before writing).
|
|
288
|
+
3. **Every failure is fail-open**: when the judging service is unavailable nothing is blocked — DSH's own sandbox preset (anything except `danger-full-access`) still applies before execution. If you want "block even when the service is down", thicken the L0 rules rather than switching to fail-closed.
|
|
289
|
+
4. **It does not defend against deliberate bypass**: rewriting the command, encoding it, or writing the authorisation file directly can all get around it. Defending against malicious injection takes a sandbox / a low-privilege user / a container.
|
|
290
|
+
|
|
291
|
+
## Known limits
|
|
292
|
+
|
|
293
|
+
**It is an accident net, not a security boundary.** It guards against *accidents* — a mistyped command, an opaque script, the moment nobody stops when full access is granted; it does **not** guard against an *adversary*. This is not unfinished work but an explicit decision: the known and deliberately kept bypasses, together with "under what circumstances to reconsider", are written up in **D1** of [docs/DECISIONS.md](docs/DECISIONS.md) — **please do not "helpfully" seal them.**
|
|
294
|
+
|
|
295
|
+
Two more deliberate omissions: judging **does not simulate filesystem state** (it will not reason "that file is empty anyway"), and it does not accept "this command is harmless" arguments that would require reading runtime state — that is precisely the crack accidents come in through.
|
|
296
|
+
|
|
297
|
+
> Every document ships **in English by default, with a Chinese sibling** (`*.zh-CN.md`, linked from a language line at the top of each file). Every number in this README is reproducible from them.
|
|
298
|
+
|
|
299
|
+
## Documentation
|
|
300
|
+
|
|
301
|
+
Each document below is English by default; add `.zh-CN` before `.md` (e.g. `docs/DECISIONS.zh-CN.md`) for the Chinese version, which is kept in step with it.
|
|
302
|
+
|
|
303
|
+
| Document | Contents |
|
|
304
|
+
|---|---|
|
|
305
|
+
| [docs/DSH-INTEGRATION.md](docs/DSH-INTEGRATION.md) | Which DSH mechanisms it uses, how the four states map, the degradation contract, and why "installed" ≠ "actually blocking" |
|
|
306
|
+
| [docs/USER-INTERVENTION.md](docs/USER-INTERVENTION.md) | The three human channels + measured evidence |
|
|
307
|
+
| [docs/DECISIONS.md](docs/DECISIONS.md) | Accepted design trade-offs D1–D14 (**read before changing anything**) |
|
|
308
|
+
| [docs/MEASUREMENTS.md](docs/MEASUREMENTS.md) | Every measured number, latency/cost, incident retrospectives |
|
|
309
|
+
| [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) | The judgment layers, and why judging and intercepting must stay separate |
|
|
310
|
+
| [docs/VERIFICATION.md](docs/VERIFICATION.md) | The acceptance checklist and per-item criteria |
|
|
311
|
+
| [DEPLOY.md](DEPLOY.md) | Deployment manual (including the Windows variant and rollback) |
|
|
312
|
+
| [START-HERE.md](START-HERE.md) | Packing/configuration instructions to hand to an AI on another machine |
|
|
313
|
+
|
|
314
|
+
## License
|
|
315
|
+
|
|
316
|
+
[MIT](LICENSE)
|
package/README.zh-CN.md
ADDED
|
@@ -0,0 +1,315 @@
|
|
|
1
|
+
# dsh-jev-guard
|
|
2
|
+
|
|
3
|
+
[](https://awesome-dsh-plugin.com)
|
|
4
|
+
[](./LICENSE)
|
|
5
|
+
[](https://nodejs.org)
|
|
6
|
+
[](https://github.com/deepseek-ai/deepseek-harness)
|
|
7
|
+
[](https://github.com/topics/dsh-plugin)
|
|
8
|
+
[](#平台支持)
|
|
9
|
+
[](https://github.com/7starsseeker/dsh-jev-guard/tags)
|
|
10
|
+
[](https://github.com/7starsseeker/dsh-jev-guard/actions/workflows/selftest.yml)
|
|
11
|
+
[](https://github.com/7starsseeker/dsh-jev-guard/commits/main)
|
|
12
|
+
[](https://github.com/7starsseeker/dsh-jev-guard/stargazers)
|
|
13
|
+
|
|
14
|
+
> [English](README.md) | **简体中文** | [更新日志](CHANGELOG.md) | [设计取舍](docs/DECISIONS.md) | [实测数据](docs/MEASUREMENTS.md)
|
|
15
|
+
|
|
16
|
+
**给 [DeepSeek Harness (DSH)](https://github.com/deepseek-ai/deepseek-harness) 用的执行前安全阀门:在命令真正跑起来之前,先问一次"它会不会不可逆地删掉或覆盖你的真实数据?"**
|
|
17
|
+
|
|
18
|
+
判定用 [TypeSafe Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) —— 一个不产出文本、只产出**结构化决策**的 "System One" 模型。这类判断用是非问句实测最准(114 例校准,中文直送 90.4%)。挂载点是 DSH 原生的 `tools/pre-execute`,所以它是**强制拦截**,不是"提醒模型自觉"。
|
|
19
|
+
|
|
20
|
+
```
|
|
21
|
+
命令文本 ──▶ L0 静态硬规则(不联网·不可覆盖)──▶ 预筛(只读/可重建)──▶ Jev 语义判定(~300ms)
|
|
22
|
+
│ │ │
|
|
23
|
+
└───────────────┬───────┴────────────────────────┘
|
|
24
|
+
▼
|
|
25
|
+
allow / revise(附降级模板)/ block / escalate
|
|
26
|
+
│ │ │ │
|
|
27
|
+
直接执行 模型换写法 拒绝 DSH 审批弹窗 或 一次性令牌
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
> **英文摘要** — `dsh-jev-guard` is a pre-execution safety valve for DeepSeek Harness. It judges every `bash` / `pwsh` tool call **before** it runs — offline static rules first, then a paid semantic model — and returns one of four states (`allow` / `revise` / `block` / `escalate`). It blocks unrecoverable commands, teaches the model a safer form when it can, offers a one-shot human token when neither is right, and **degrades loudly instead of silently** when the judging API runs out of credit. **It is an accident net, not a security boundary** — see [已知边界](#已知边界).
|
|
31
|
+
|
|
32
|
+
---
|
|
33
|
+
|
|
34
|
+
## 目录
|
|
35
|
+
|
|
36
|
+
- [它做什么](#它做什么)
|
|
37
|
+
- [安装](#安装)
|
|
38
|
+
- [语言](#语言)
|
|
39
|
+
- [配置](#配置)
|
|
40
|
+
- [两种审批策略下的行为](#两种审批策略下的行为)
|
|
41
|
+
- [被拦了怎么办:人的三条介入通道](#被拦了怎么办人的三条介入通道)
|
|
42
|
+
- [额度用完会怎样:降级而不是静默失效](#额度用完会怎样降级而不是静默失效)
|
|
43
|
+
- [看得见:审计日志与状态](#看得见审计日志与状态)
|
|
44
|
+
- [平台支持](#平台支持)
|
|
45
|
+
- [自检与验收](#自检与验收)
|
|
46
|
+
- [目录结构](#目录结构)
|
|
47
|
+
- [安全与隐私](#安全与隐私)
|
|
48
|
+
- [已知边界](#已知边界)
|
|
49
|
+
- [文档](#文档)
|
|
50
|
+
- [License](#license)
|
|
51
|
+
|
|
52
|
+
## 它做什么
|
|
53
|
+
|
|
54
|
+
| 判定 | 含义 | 谁接手 |
|
|
55
|
+
|---|---|---|
|
|
56
|
+
| `allow` | 确定性只读,或判定为安全 | 直接执行 |
|
|
57
|
+
| `revise` | 证据不足以安全执行,但很可能有更好的写法 | **模型**换写法重试(附三种降级模板:先演练 / 缩小范围 / 先备份) |
|
|
58
|
+
| `block` | 会不可逆地删或覆盖,或命中硬规则 | 拒绝;由人手动执行 |
|
|
59
|
+
| `escalate` | 必须有人确认(硬规则的"必问"项,或重试预算耗尽) | 人 |
|
|
60
|
+
|
|
61
|
+
分三层,顺序固定:
|
|
62
|
+
|
|
63
|
+
1. **L0 静态硬规则**(`lib/rules.js`):21 条"永不允许" + 16 条"必须人工确认"。**不联网、不可被覆盖**,连一次性令牌也过不去。规则只在**命令位置**匹配(每一行行首,或 `;` `&` `|` `(` `$(` 之后,或 `bash -c "` 之后,并允许 `sudo`/`timeout 30`/`xargs -0`/`find … -exec` 这类包装器)—— 所以"在参数里提到危险命令"不会被误伤,**而多行脚本里的真命令也不会被漏掉**。只有两条结构上锚不了的规则(`redirect-to-device`、`fork-bomb`)是全文匹配,计数器 `RULE_STATS.anywhere` 恒为 2。
|
|
64
|
+
2. **预筛**:可证明只读或只影响可重建内容(缓存、构建产物、`/tmp`)的命令直接放行,**零网络调用**。
|
|
65
|
+
3. **Jev 语义判定**:一次是非问句 —— *"这条命令会不可逆地删除或覆盖用户的真实数据吗?"* —— 按两个阈值切成 `allow` / `revise` / `block`。实测延迟 ~300ms(P50 267ms),成本 ≈ `$0.000019`/次。
|
|
66
|
+
|
|
67
|
+
## 安装
|
|
68
|
+
|
|
69
|
+
要求 **Node ≥ 20**(用到全局 `fetch`)。**零运行时依赖**,不需要 `npm install`。
|
|
70
|
+
|
|
71
|
+
**已验证的宿主版本:DSH 0.1.6-alpha.2。** 这是本插件唯一跑过的 DSH 版本,且**刻意不在 `package.json` 里声明为宿主要求**:插件市场会从 npm manifest 读这个字段,一旦声明就会在其他所有 DSH 版本上拦住安装与更新。所以别的版本是**没验过,而不是被禁止**;换版本后请重跑下面的自检。
|
|
72
|
+
|
|
73
|
+
```bash
|
|
74
|
+
# 1. 把本仓库放到一个固定的位置,例如 T:\dsh-jev-guard(WSL 里是 /mnt/t/dsh-jev-guard)
|
|
75
|
+
|
|
76
|
+
# 2. 让 DSH 装载它(profile 名按你的实际 profile 填)
|
|
77
|
+
dsh plugin --profile web add /mnt/t/dsh-jev-guard # Windows 侧: T:\dsh-jev-guard
|
|
78
|
+
|
|
79
|
+
# 3. 重启 DSH(插件没有热加载)
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
`dsh plugin` 的 `add` 走 pnpm 解析,所以 spec 接受 pnpm 接受的一切。发布的包已上 npm —— 也就是插件市场优先采用的安装源:
|
|
83
|
+
|
|
84
|
+
```bash
|
|
85
|
+
dsh plugin --profile web add dsh-jev-guard
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
本地路径是**链接**装法,插件始终跑在你自己那份 checkout 上 —— 改文件、重启,就生效。想直接从 GitHub 源码装,用 `github:7starsseeker/dsh-jev-guard`。
|
|
89
|
+
|
|
90
|
+
以上几种都没有东西要构建(零依赖、无安装脚本),所以都不会触发构建授权。
|
|
91
|
+
|
|
92
|
+
**新装的时候没有密钥 —— 它会自己说出来,而不是装死。** 第一个会话会在对话里直接告诉你"没有配置密钥";在录入之前,阀门处于**降级**:免费的 L0 硬规则与预筛照常工作,付费的语义层不工作。录入只要一条命令,而且只从标准输入读 —— 绝不接受参数,那会进 shell 历史与 `ps`:
|
|
93
|
+
|
|
94
|
+
```bash
|
|
95
|
+
node bin/guard.mjs key set # 粘贴密钥后回车;不回显、不进 shell 历史
|
|
96
|
+
node bin/guard.mjs key status # 当前哪个来源在生效、密钥多长(永不打印值)
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
`guard key status` 在没有密钥时退出码 3,可以直接当健康检查。这里没有需要等的冷却:密钥一解析到,降级状态当场清除,下一条命令就恢复完整判定。
|
|
100
|
+
|
|
101
|
+
`package.json` 里的声明是一个标准 DSH bundle:
|
|
102
|
+
|
|
103
|
+
```json
|
|
104
|
+
{
|
|
105
|
+
"name": "dsh-jev-guard",
|
|
106
|
+
"dsh": { "bundle": { "patch": "./cordis.patch.yml" } }
|
|
107
|
+
}
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
`cordis.patch.yml` 把插件插到 `tools/pre-execute` 上,**所有可调参数都在那里**(也可以写在 `config.json` 里,优先级 `patch > config.json > 内置默认`)。
|
|
111
|
+
|
|
112
|
+
**密钥**三种来源,优先级从高到低:DSH 凭据层(`ctx.credentials`,轮换后无需重启)→ 环境变量 `TYPESAFE_API_KEY` → **你自己在包根建的** `secrets.json`(`{"TYPESAFE_API_KEY": "apikey_..."}`)。该文件写在 `.gitignore` 里,**刻意不入仓库、也不进发布包**,没人会替你带一份 —— 前两种来源才是首选。`node bin/guard.mjs key set` 会替你写这个文件(权限 0600),DSH 适配器也读它。密钥从不被打印,写入日志前会掩码。
|
|
113
|
+
|
|
114
|
+
**装好后立刻验一次**(别看"没报错"):
|
|
115
|
+
|
|
116
|
+
```bash
|
|
117
|
+
node bin/guard.mjs selftest # 期望:12 项全部通过(不联网)
|
|
118
|
+
node bin/guard.mjs status # 期望:✅ 正常(降级时退出码为 3)
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
再在会话里跑一条**必然被拦**的命令(例如 `git push --force origin main`),它应该被拒,并且 `node bin/guard.mjs log --tail 3` 里能看到那条记录。
|
|
122
|
+
|
|
123
|
+
## 语言
|
|
124
|
+
|
|
125
|
+
**给人或模型看的话**都有中英两份:判定理由、37 条 L0 规则的理由、全部 CLI 输出、降级告警与 `guard status`。
|
|
126
|
+
|
|
127
|
+
| 开关 | 管什么 | 默认 |
|
|
128
|
+
|---|---|---|
|
|
129
|
+
| `lang` | 上面那些文案的语言 | `'auto'` —— 按 `JEV_GUARD_LANG` → `LC_ALL`/`LC_MESSAGES`/`LANG`(仅当它们指明受支持的语言)→ 否则 `zh-CN` |
|
|
130
|
+
| CLI `--lang zh-CN\|en` | 同一条命令的临时指定 | — |
|
|
131
|
+
| `promptLang` | **发给 Jev 的那句问话**与它读的 state 的键 | `'zh-CN'` |
|
|
132
|
+
|
|
133
|
+
**`promptLang` 不是翻译开关,而是一个判定参数。** 阈值 0.5 / 0.7 是在**中文问话**上标定的(114 例),换语言就等于挪动这条被测过的边界。实测(`tools/probe-prompt-lang.mjs`,21 条探针 × 每臂 3 次 × 2 轮独立运行):
|
|
134
|
+
|
|
135
|
+
| 指标 | 结果 |
|
|
136
|
+
|---|---|
|
|
137
|
+
| 与中文问话同带 | **18/21** |
|
|
138
|
+
| 概率平均绝对差 | **0.049** |
|
|
139
|
+
| 带符号均值(负 = 英文问话更宽松) | **−0.043** |
|
|
140
|
+
| 偏低 / 偏高 / 持平 | 12 / 4 / 5 |
|
|
141
|
+
| 重复采样噪声(同问话同状态) | **0.015** |
|
|
142
|
+
|
|
143
|
+
三条命令翻了带,而且**方向全部朝放行**:无 `WHERE` 的 `UPDATE`(block → revise)、`DELETE … WHERE`(block → revise)、内联 `node -e rmSync`(revise → allow)。两轮结论一致。所以真要设 `promptLang: 'en'`,请先重标定,或把两个阈值下调约 0.04 —— 见 [docs/MEASUREMENTS.md](docs/MEASUREMENTS.md) §14 与 [docs/DECISIONS.md](docs/DECISIONS.md) **D14**。
|
|
144
|
+
|
|
145
|
+
代码注释与 `tools/` 里的自检标签**刻意保持中文**:它们由本仓库的维护者读,双语化只会让每次改动的维护成本翻倍,而不改变产品对外说的任何一句话。
|
|
146
|
+
|
|
147
|
+
**`auto` 为什么不看系统 locale:** 本插件最初也会落到 `Intl`,结果第一次真实部署就踩到 —— WSL 里 `LANG=C.UTF-8` 表示"没有偏好",`Intl` 于是报出 Node 自己的 `en-US` 兜底值,会话里的理由**悄悄变成英文**,而 Windows 侧 CLI 仍是中文。`C`/`POSIX`/未设置一律视作**没有信号**,落在 `zh-CN`;真正指明语言的 locale(`en_US.UTF-8`、`zh_CN.UTF-8`)照常生效。
|
|
148
|
+
|
|
149
|
+
> English readers: the default README is [README.md](README.md).
|
|
150
|
+
|
|
151
|
+
## 配置
|
|
152
|
+
|
|
153
|
+
`config.example.json` 是模板;全部键如下(默认值写在 `lib/gate.js` 的 `DEFAULTS`)。
|
|
154
|
+
|
|
155
|
+
| 键 | 默认 | 说明 |
|
|
156
|
+
|---|---|---|
|
|
157
|
+
| `model` | `jev-latest` | 判定模型别名 |
|
|
158
|
+
| `endpoint` | `https://api.typesafe.ai/v1/systemone` | 判定服务地址 |
|
|
159
|
+
| `apiKeyEnv` / `apiKeyFile` | `TYPESAFE_API_KEY` / `secrets.json` | 密钥来源;**相对路径按包根解析**(与 cwd 无关) |
|
|
160
|
+
| `lang` | `auto` | 全部给人/模型看的文案的语言(见[语言](#语言));`auto` 只读环境变量,兜底 `zh-CN` |
|
|
161
|
+
| `promptLang` | `zh-CN` | 发给 Jev 的问话与 state 键的语言。**它是判定参数而非翻译开关**:实测切 `en` 会把 p 平均压低约 0.04,并让三条探针翻向放行,见[语言](#语言) |
|
|
162
|
+
| `lowThreshold` / `highThreshold` | `0.5` / `0.7` | 四态阈值:`p < low` → allow;`low ≤ p < high` → revise;`p ≥ high` → block。737 条真实命令上得到 98.51% / 0.81% / 0.68% 三分 |
|
|
163
|
+
| `reviseInAskMode` / `blockInAskMode` | `ask` / `ask` | 判定动作怎么随审批模式分叉:`ask` = 审批可用时转人工弹窗(人就在场,不该让 50% 的判断替人做决定);`deny` = 退回直接拒绝。`never`(全自动)下两者都仍是直接拒绝。**L0 的 `deny` 类硬规则不受此开关影响** —— 它永远拦死 |
|
|
164
|
+
| `timeoutMs` | `1800` | 单次判定预算;超时一律放行(fail-open) |
|
|
165
|
+
| `cacheSize` | `256` | 判定缓存条数 |
|
|
166
|
+
| `inlineScripts` / `maxScriptBytes` | `true` / `8192` | 把被调用脚本的正文读进判定状态(实测把 `node x.mjs` 这类盲区从 0.31 提到 0.82);敏感路径自动跳过 |
|
|
167
|
+
| `retryLimit` | `2` | 同一条命令被拦多少次后升级为 `escalate`(交人处理) |
|
|
168
|
+
| `tokens` / `tokenPath` | `true` / `~/.jev-guard/allow.txt` | 一次性放行令牌 |
|
|
169
|
+
| `logPath` / `logMaxBytes` | `~/.jev-guard/guard.log` / 4 MiB | 共享审计日志(JSONL,超限轮转) |
|
|
170
|
+
| `quotaCooldownMs` / `authCooldownMs` | 15 min / 30 min | 额度、密钥类失败后的降级冷却 |
|
|
171
|
+
| `degradePolicy` | `l0-only` | 降级时保留哪一层:`l0-only`(只停要花钱的语义层)或 `off`(整条阀门暂停) |
|
|
172
|
+
| `pricePerMTok` | `0.042` | 成本估算单价(美元/百万输入 token;输出按官方说明免费) |
|
|
173
|
+
|
|
174
|
+
## 两种审批策略下的行为
|
|
175
|
+
|
|
176
|
+
同一套判定,在 DSH 的两种会话策略下落地不同 —— **这点最容易搞混**:
|
|
177
|
+
|
|
178
|
+
| 判定 | `approval: ask`(会弹框) | `approval: never`(完全权限 / YOLO) |
|
|
179
|
+
|---|---|---|
|
|
180
|
+
| `revise`(50–70%) | **转人工弹审批框** | 拒绝 **+ 降级模板 + 一次性令牌提示** |
|
|
181
|
+
| `block`(≥70%,语义层) | **转人工弹审批框** | 拒绝 **+ 一次性令牌提示** |
|
|
182
|
+
| L0 的 `deny` 类硬规则 | **拒绝**(不弹框、不发令牌) | **拒绝** |
|
|
183
|
+
| L0 的 `ask` 类规则(`escalate`) | **DSH 弹审批框**,由人决定 | 拒绝 **+ 一次性令牌提示** |
|
|
184
|
+
|
|
185
|
+
> **为什么 `ask` 模式下灰区和高分都交给人**:人就在场时,让一个 50.6% 的判断替人做决定没有道理;而 `never` 模式下没人可问,只能由阀门保守地拒。宿主没有应答者时审批是 **fail-closed**,所以转人工**不会**在无人值守时变成自动放行;弹窗只给 `allowed-once`,不留长期旁路。
|
|
186
|
+
>
|
|
187
|
+
> **L0 的 `deny` 类硬规则是绝对闸门**:两种模式都拦死,连 `retryLimit` 的"反复重试就交人"也不适用于它 —— 否则弹窗里点一次"允许"就绕过了硬地板(令牌不能越过 L0,审批同样不能)。
|
|
188
|
+
|
|
189
|
+
> `danger-full-access` = `{ sandbox: 'danger-full-access', approval: 'never' }` —— **没有沙箱兜底、审批也等于关掉,阀门是唯一一层**。这正是它存在的意义,也是它判错时代价最大的场景。审计记录里会带上当次的沙箱档位(`preset`)与审批策略(`policy`),事后复盘能看出"当时后面还有没有沙箱"。
|
|
190
|
+
|
|
191
|
+
## 被拦了怎么办:人的三条介入通道
|
|
192
|
+
|
|
193
|
+
**① 一次性令牌(与宿主无关,任何时候都在)。** 被拦命令的理由里会附一个 `ALLOW-XXXXXXXXXX`(命令文本的哈希前 10 位),**人**在自己的终端里粘贴理由给的那一整行:
|
|
194
|
+
|
|
195
|
+
```bash
|
|
196
|
+
# WSL / Linux:
|
|
197
|
+
node /mnt/t/dsh-jev-guard/bin/guard.mjs allow '<原命令>'
|
|
198
|
+
# Windows(PowerShell;引号按平台自动切换):
|
|
199
|
+
node T:\dsh-jev-guard\bin\guard.mjs allow '<原命令>'
|
|
200
|
+
# 任何平台、任何 shell(不经过 shell 引号规则 —— cmd.exe 用这个):
|
|
201
|
+
node T:\dsh-jev-guard\bin\guard.mjs allow --command-file cmd.txt
|
|
202
|
+
|
|
203
|
+
node bin/guard.mjs allow --list # 看待用令牌
|
|
204
|
+
node bin/guard.mjs allow --revoke ALLOW-… # 撤销
|
|
205
|
+
```
|
|
206
|
+
|
|
207
|
+
四条性质:**绑定命令原文**(换一个字就是另一个令牌)、**用掉即删**(无法重放)、**不越过 L0 硬规则**、**只在交互终端授权**(AI 自己跑会被拒)。
|
|
208
|
+
|
|
209
|
+
**② DSH 审批弹窗(`approval: ask`)。** 阀门只负责把命令标成"需要人看",由 DSH 弹框;弹窗里的理由就是阀门原文。DSH 的答案是闭集(只有"允许一次"和"拒绝"),所以**每次都是一次独立的逐次决定**,没有"永久允许"可以被静默吞掉。
|
|
210
|
+
|
|
211
|
+
**③ 人直接执行。** 你在自己终端里跑那条命令 —— 阀门不参与,也**不会因此给 AI 任何权限**:审计里不会有你那次动作的记录,AI 重试同一条命令仍然会被拦。
|
|
212
|
+
|
|
213
|
+
## 额度用完会怎样:降级而不是静默失效
|
|
214
|
+
|
|
215
|
+
判定服务是**按量收费**的,额度用完是必然事件。默认行为:
|
|
216
|
+
|
|
217
|
+
| 情况 | 阀门怎么办 | 你能在哪里看到 |
|
|
218
|
+
|---|---|---|
|
|
219
|
+
| 额度耗尽 / 密钥失效(`402` / `401`) | **降级**:写 `~/.jev-guard/degraded.json`,冷却窗口内不再发请求(省钱),默认只跑**免费的 L0 + 预筛** | `guard status`(退出码 3)· 拒绝理由里的一句 `⚠️` · 审计里的 `source: degraded` 与 `level: warn` · CLI 的 stderr |
|
|
220
|
+
| 冷却到期 | 自动放**一次**探测请求:成功即恢复(你不用做任何事),失败继续降级 | `guard status` 会显示还剩多久 |
|
|
221
|
+
| 超时 / 网络抖 / 5xx / 429 限流 / 无密钥 | **不降级**,只逐次放行(fail-open),但会被分类记录 | `guard log --stats` 的"失败分类"一行 |
|
|
222
|
+
|
|
223
|
+
```bash
|
|
224
|
+
node bin/guard.mjs status --clear # 不想等冷却:立刻重试一次(失败会再次进入降级)
|
|
225
|
+
```
|
|
226
|
+
|
|
227
|
+
想"额度没了就彻底别插手"(连 L0 也停)就设 `degradePolicy: "off"`。
|
|
228
|
+
|
|
229
|
+
## 看得见:审计日志与状态
|
|
230
|
+
|
|
231
|
+
每个判定都会追加一行 JSONL 到 `~/.jev-guard/guard.log`(超 4 MiB 轮转,命令写入前掩码密钥):
|
|
232
|
+
|
|
233
|
+
```bash
|
|
234
|
+
node bin/guard.mjs log --tail 20 # 时间 / 动作 / p / 来源 / 命中规则 / 命令
|
|
235
|
+
node bin/guard.mjs log --stats # 动作·来源·规则计数 + fail-open + 失败分类 + 成本估算
|
|
236
|
+
node bin/guard.mjs status # 一句话:阀门是好的吗?(降级时退出码 3,可当健康检查)
|
|
237
|
+
```
|
|
238
|
+
|
|
239
|
+
## 平台支持
|
|
240
|
+
|
|
241
|
+
**WSL/Linux 与 Windows 都支持。** 两处平台差异都已处理:
|
|
242
|
+
|
|
243
|
+
| 项 | WSL / Linux | Windows |
|
|
244
|
+
|---|---|---|
|
|
245
|
+
| 拦的工具 | `bash` | `pwsh`(**两个都在默认 `tools` 列表里**) |
|
|
246
|
+
| 授权行引号 | POSIX `'\''` | **PowerShell `''`**(两种写法不通用,已按平台分叉) |
|
|
247
|
+
| cmd.exe 用户 | — | 用 `guard allow --command-file <文件>` |
|
|
248
|
+
| 状态与日志 | `~/.jev-guard/` | `%USERPROFILE%\.jev-guard\` |
|
|
249
|
+
|
|
250
|
+
自检里对引号做了**真机往返断言**(含"POSIX 形式在 PowerShell 里必须失败"的反例);入口守卫的跨平台回归也必须在**两个平台各跑一遍**才算验过。
|
|
251
|
+
|
|
252
|
+
## 自检与验收
|
|
253
|
+
|
|
254
|
+
```bash
|
|
255
|
+
# 七份离线自检(不需要网络、不需要密钥)
|
|
256
|
+
for t in selftest-entry selftest-i18n selftest-quota selftest-reason selftest-token selftest-rules selftest-audit; do
|
|
257
|
+
printf '%-18s ' "$t"; node tools/$t.mjs | tail -1
|
|
258
|
+
done
|
|
259
|
+
|
|
260
|
+
node bin/guard.mjs selftest # 12 项:规则/预筛/四态映射
|
|
261
|
+
node tools/smoke-dsh-adapter.mjs # 适配器冒烟(假 ctx,9 组断言)
|
|
262
|
+
node tools/smoke-dsh-pipeline.mjs # 真实工具管线集成(需在 DSH 检出目录内跑)
|
|
263
|
+
```
|
|
264
|
+
|
|
265
|
+
验收清单(20 项,含三条人工通道)与逐项判据见 **[docs/VERIFICATION.md](docs/VERIFICATION.md)**;
|
|
266
|
+
历次结论落在 [`verification-results/`](verification-results/)。
|
|
267
|
+
|
|
268
|
+
## 目录结构
|
|
269
|
+
|
|
270
|
+
```
|
|
271
|
+
bin/guard.mjs CLI:judge | log | status | allow | selftest | rules
|
|
272
|
+
lib/gate.js 判定引擎(L0 → 预筛 → 语义 → 四态)—— 与调用方无关
|
|
273
|
+
lib/i18n.js 双语文案目录(zh-CN / en)与语言解析
|
|
274
|
+
lib/rules.js L0 静态硬规则(每条带 id / 正则 / 理由)
|
|
275
|
+
lib/verdict.js 四态合成、理由文案、重试预算、平台相关引号
|
|
276
|
+
lib/audit.js 共享审计日志(掩码 / 轮转 / 汇总 / 成本)
|
|
277
|
+
lib/token.js 一次性放行令牌
|
|
278
|
+
lib/quota.js 额度/密钥失败后的降级状态机
|
|
279
|
+
adapters/dsh/index.js DSH 原生 Cordis 插件(唯一的适配器)
|
|
280
|
+
cordis.patch.yml DSH bundle patch(装载声明 + 全部可调参数)
|
|
281
|
+
tools/ 离线自检、冒烟测试、验证辅助
|
|
282
|
+
docs/ 机制、取舍、实测、验收清单
|
|
283
|
+
```
|
|
284
|
+
|
|
285
|
+
## 安全与隐私
|
|
286
|
+
|
|
287
|
+
1. **会发给判定服务的只有命令文本 + 可选脚本正文**(敏感路径 `.env` / `.ssh` / `*.pem` / `*credential*` / `*secret*` / `*token*` 自动跳过,单文件 8KB 上限)。想彻底关闭:`inlineScripts: false`(代价是 `node x.mjs` 这类命令退回 p≈0.31 的盲区)。
|
|
288
|
+
2. **密钥只从凭据层 / 环境变量 / `secrets.json` 读取**,不进日志、不进报告(命令文本写入前掩码)。
|
|
289
|
+
3. **失败一律放行(fail-open)**:判定服务不可用时不拦任何东西 —— DSH 自己的沙箱档位(除 `danger-full-access` 外)仍在执行之前。想"服务挂了也拦",加厚 L0 规则,而不是改成 fail-closed。
|
|
290
|
+
4. **不防蓄意绕过**:换写法、编码、直接写授权文件都可能绕开。防恶意注入要靠沙箱 / 低权限用户 / 容器。
|
|
291
|
+
|
|
292
|
+
## 已知边界
|
|
293
|
+
|
|
294
|
+
**它是事故安全网,不是安全边界。** 它防的是*事故* —— 写错的命令、不透明的脚本、完全权限下没人拦的那一下;它**不**防*对手*。这不是没做完,是显式决策:已知且有意保留的旁路、以及"什么情况下该重新考虑",都写在 [docs/DECISIONS.md](docs/DECISIONS.md) 的 **D1**,**请不要"顺手把它堵上"**。
|
|
295
|
+
|
|
296
|
+
同样刻意的两条:判定**不模拟文件系统状态**(不会推理"反正那个文件已经是空的"),也不接受"这条命令没害处"这类需要读运行时状态的辩解 —— 那正是事故钻进来的缝。
|
|
297
|
+
|
|
298
|
+
## 文档
|
|
299
|
+
|
|
300
|
+
下面的文档**默认是英文**;把扩展名写成 `*.zh-CN.md`(例如 `docs/DECISIONS.zh-CN.md`)就是与它同步的中文版。
|
|
301
|
+
|
|
302
|
+
| 文档 | 内容 |
|
|
303
|
+
|---|---|
|
|
304
|
+
| [docs/DSH-INTEGRATION.md](docs/DSH-INTEGRATION.md) | 它用 DSH 的哪些机制、四态怎么映射、降级契约、为什么"装上了≠真的在拦" |
|
|
305
|
+
| [docs/USER-INTERVENTION.md](docs/USER-INTERVENTION.md) | 人的三条介入通道 + 实测证据 |
|
|
306
|
+
| [docs/DECISIONS.md](docs/DECISIONS.md) | 已接受的设计取舍 D1–D14(**改之前先读**) |
|
|
307
|
+
| [docs/MEASUREMENTS.md](docs/MEASUREMENTS.md) | 全部实测数字、延迟/成本、事故复盘 |
|
|
308
|
+
| [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) | 判定分层、为何判定与拦截必须分开 |
|
|
309
|
+
| [docs/VERIFICATION.md](docs/VERIFICATION.md) | 验收清单与逐项判据 |
|
|
310
|
+
| [DEPLOY.md](DEPLOY.md) | 部署手册(含 Windows 变体与回滚) |
|
|
311
|
+
| [START-HERE.md](START-HERE.md) | 交给另一台机器上的 AI 的装箱/配置说明 |
|
|
312
|
+
|
|
313
|
+
## License
|
|
314
|
+
|
|
315
|
+
[MIT](LICENSE)
|