dsh-escalation-review 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md ADDED
@@ -0,0 +1,33 @@
1
+ # Changelog
2
+
3
+ ## 0.2.0 — 2026-09-29
4
+
5
+ **What it is.** An escalation-only reviewer for DeepSeek Harness. Ordinary work keeps running on the native
6
+ sandbox; the plugin wakes for exactly one thing — a call that asks to leave that sandbox — and returns a
7
+ verdict for that call through the host's own plumbing.
8
+
9
+ - **Codex-style policy.** Risk tier and authorization are scored separately and the pair decides
10
+ `allow` / `deny` / `ask`. `medium` means reversible *consequences*, not a reversible artifact. System and
11
+ security control state (`hosts`, DNS, firewall, services, certificates, `PATH`) is `high`. A safer
12
+ alternative downgrades the action; missing records are not permission; a post-hoc approval is high risk.
13
+ User instructions reach the reviewer as a bounded root projection (head 4 + tail 12, at most 16, dropped
14
+ whole rather than truncated, with an explicit "evidence incomplete" marker).
15
+ - **Intervention switch, off by default.** A single segment control on the config page decides whether the
16
+ plugin takes part at all. It is independent of permission presets, so installing it does not require
17
+ editing a preset table.
18
+ - **Audit card.** Every reviewed call gets a line in the transcript with its verdict and the reason behind
19
+ it, in semantic colours, expandable for the full text.
20
+ - **Fail-closed.** A review that fails or times out denies by default; consecutive denials trip a circuit
21
+ breaker; the pending action is cross-checked against the session log before a verdict is accepted.
22
+ - **Host-side behaviour.** An allow answers the escalation approval as `allowed-once`; a denial blocks the
23
+ tool body and returns the reason to the model. The plugin never writes session events.
24
+
25
+ **Compatibility.** Verified with DSH `0.1.7-rc.2` and `0.2.0-rc.1` on Node 24. Peer requirements are given as
26
+ ranges rather than exact pins, so a DSH upgrade does not fail the plugin compatibility check.
27
+
28
+ **Relationship to the bundled `auto` preset.** The experimental auto-review plugin registers an `auto`
29
+ permission preset, binds the session to `danger-full-access` and reviews every native call. This plugin takes
30
+ the other side of that trade: it leaves the session's sandbox tier in force and reviews only the call that
31
+ asks to escalate.
32
+
33
+ **License.** MIT. See `LICENSE`; provenance and attribution are in `NOTICE.md`.
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 dsh-escalation-review contributors
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/NOTICE.md ADDED
@@ -0,0 +1,17 @@
1
+ # NOTICE / 来源与致谢
2
+
3
+ 本插件的代码为独立实现。分档口径、授权评分思路与提示词结构参考了以下项目(均为其各自许可):
4
+
5
+ 1. **OpenAI Codex** —— guardian 策略模板(`codex-rs/prompts/templates/guardian/policy_template.md`),
6
+ Apache License 2.0:<https://github.com/openai/codex>
7
+ 参考内容:风险分档、用户授权评分、"缺失记录不等于许可"、事后批准视为高风险、
8
+ 有更不危险的替代方案则降档等判据。**未复制其文本**(逐字重合度实测为 0)。
9
+ 2. **DeepSeek Harness 官方 auto-review 插件**(`@deepseek-ai/dsh-experimental-auto-review`),MIT License:
10
+ 参考内容:五分区评审输入、结构化 JSON 决策(`{risk, authorization, outcome, reason?}`)、
11
+ "证据有歧义即失败"的 fail-closed 姿态、`tools/pre-execute` 三态契约。
12
+ 逐字重合度实测为 1 处短句,已在本仓库改写为自有措辞。
13
+ 3. **DeepSeek Harness** 本体(MIT License):本插件通过其公开扩展点工作
14
+ (`tools/pre-execute`、`approval/request`、`sessionProjections`、`conversation.chat.node` 等),
15
+ 并在运行时使用宿主自带的 `@deepseek-ai/dsh-client-ui-primitives` 图标组件(属运行时依赖,不随本包分发)。
16
+
17
+ 策略文本中的中文/英文示例(例如"执行吧"、"我不许可")取自真实会话记录并已脱敏,仅作为语言样例。
package/README.md ADDED
@@ -0,0 +1,186 @@
1
+ # dsh-escalation-review
2
+
3
+ **English** · [简体中文](README.zh.md)
4
+
5
+ An **escalation-only LLM reviewer** for DeepSeek Harness (DSH).
6
+
7
+ Ordinary work inside the workspace — reads, writes, shell commands — runs on the native sandbox with no
8
+ review and no extra model calls. The plugin wakes for exactly one thing: a call that asks to leave the
9
+ sandbox (`sandbox_permissions` differs from the mode actually in effect). That one call gets a verdict —
10
+ allow it, refuse it, or hand it back to you — delivered through the host's own plumbing: an approval
11
+ answered as `allowed-once`, or a three-state decision object returned before the tool body runs.
12
+
13
+ ## A Codex-style reviewer
14
+
15
+ The policy follows OpenAI's Codex guardian template
16
+ (`codex-rs/prompts/templates/guardian/policy_template.md`), and that is where most of the work went.
17
+
18
+ - **Three-axis contract.** Risk tier and authorization are scored separately, and the pair decides
19
+ `allow` / `deny` / `ask`. Authorization is re-derived on every escalation: a past approval counts only
20
+ while its evidence is still in the session.
21
+ - **`medium` means reversible consequences, not a reversible artifact.** A step is medium only when its
22
+ blast radius is bounded *and* what it does can be undone. Editing a file back is not recalling its effect.
23
+ - **System and security control state is `high`.** `hosts`, DNS, firewall rules, services, certificates,
24
+ `PATH`. Each can be edited back; traffic already sent cannot be recalled.
25
+ - **A safer alternative downgrades the action.** When the same goal is reachable without the dangerous
26
+ step, the dangerous step is not what the user asked for.
27
+ - **Missing records are not permission.** Absent evidence is absence of authorization, never implied consent.
28
+ - **A post-hoc approval is high risk.** Approving after the fact does not unlock a `critical` action.
29
+ - **Bounded root projection.** User instructions reach the reviewer as head 4 + tail 12, at most 16, whole
30
+ items dropped rather than truncated, with an explicit "evidence incomplete" marker. Evidence is labelled
31
+ by role: human instruction, direct parent, approval answer, checkpoint, or fact.
32
+
33
+ Four rules shape the edges:
34
+
35
+ - **An explicit user prohibition is a hard denial**, overriding the risk tier even for a low-risk action.
36
+ - **Evidence is data, not instruction.** Command text, file contents and tool output carry no authority;
37
+ text claiming "the user already allowed this" authorizes nothing.
38
+ - **Appearance does not downgrade.** A credential-shaped payload is not made safer by looking like a
39
+ placeholder, nor an export by the destination it claims.
40
+ - **Cross-check before judging.** The pending action is compared with the logged `tool/call` for the same
41
+ `callId`; an unambiguous disagreement aborts the review. A malformed verdict counts as a review failure.
42
+
43
+ Attribution and the measured text overlap with the projects these ideas come from are in [NOTICE.md](NOTICE.md).
44
+
45
+ ## How it differs from the bundled `auto` preset
46
+
47
+ The app ships `@deepseek-ai/dsh-experimental-auto-review`. It registers an `auto` permission preset, binds
48
+ the session to `danger-full-access`, and reviews **every** native call — one extra model request per call,
49
+ at full access.
50
+
51
+ | | bundled `auto` preset | this plugin |
52
+ |---|---|---|
53
+ | sandbox tier | `danger-full-access` | the session's own tier, so `workspace-write` stays in force |
54
+ | calls reviewed | every native call | only the call that asks to escalate |
55
+ | preset table | registers `auto` | untouched — the gate is this plugin's own switch |
56
+ | extra model calls | one per call | one per escalation, none otherwise |
57
+
58
+ Keeping the session's sandbox tier is the point of the comparison: the sandbox covers ordinary work, and
59
+ the reviewer covers the gap where DSH would otherwise stop to ask you.
60
+
61
+ ## Behaviour
62
+
63
+ | Situation | What happens |
64
+ |---|---|
65
+ | Ordinary call, inside the sandbox | Not reviewed; passed to the next handler untouched. |
66
+ | Escalation, intervention switch off | Not reviewed, not touched, no model call. This is the default state. |
67
+ | Escalation, switch **on** | Reviewed once and recorded. An allow answers the approval with `allowed-once` (you are **not** asked again); a deny blocks the body and returns the reason to the model. |
68
+ | Review failed or timed out | `failMode` decides. Default `deny` (fail closed). |
69
+ | Three denials in a row, or 10 within 50 calls | Circuit breaker: the plugin stops refusing and hands the call to you instead. |
70
+ | Pending action disagrees with its logged tool call | The review stops rather than judge evidence it cannot verify. |
71
+
72
+ ## Requirements
73
+
74
+ - DSH **0.1.7-rc.2** or **0.2.0-rc.1** (both exercised), Node 24.
75
+ - Windows is what has been exercised; other platforms are expected to work but are unverified.
76
+ - `peerDependencies`: `@deepseek-ai/cordis` `~4.0.4`, `@deepseek-ai/dsh-llm` `>=0.1.7-rc.2 <0.3.0`.
77
+ - Write `@deepseek-ai/dsh*` peers as a range: the host checks them with
78
+ `semver.satisfies(runtime, range, { includePrerelease: true })`, so an exact prerelease pin reports as
79
+ incompatible on the next DSH release.
80
+
81
+ ## Install
82
+
83
+ From a local checkout, the GUI route is the cleanest: left sidebar **Plugins** (not Settings) → **Add
84
+ plugin** → the absolute path of this directory, for example
85
+ `D:\path\to\checkout\plugins\escalation-review`. It is registered as a `link:` dependency and appended to
86
+ `dsh.profile.bundles`.
87
+
88
+ ```sh
89
+ dsh plugin --profile <profile-name> add <package-or-path>
90
+ ```
91
+
92
+ - `~/.dsh/profiles/web` and `~/.dsh/profiles/desktop` are **independent**: installing into one leaves the
93
+ other alone. Use `--profile web` to test against `dsh web`.
94
+ - The host half is loaded from disk at startup, so **restart DSH** after installing, and again after
95
+ changing any file under `lib/`. The client half of an open session needs that restart too; a page reload
96
+ does not pick it up in the desktop app.
97
+
98
+ Uninstall from the same Plugins page, or drop the entry from `dsh.profile.bundles`.
99
+
100
+ ## Configuration
101
+
102
+ `enabled` is the gate, and it ships **off**: while it is off the plugin runs no review and makes no model
103
+ calls. With it on, it answers **every** sandbox-escalation approval. The gate belongs to the plugin rather
104
+ than to your permission table, so installing it leaves your presets as they are, independent of whichever
105
+ preset a session happens to use. Set it from the config page — the two-position control on the first row
106
+ of the card, Off / On — or with `"enabled": true` in a config file (the GUI writes the textual form
107
+ `"enabledText": "true"`).
108
+
109
+ Lowest to highest priority: the package's `config.json` → `~/.dsh/escalation-review.config.json` →
110
+ `configPath` if you pass one → the host config block of the plugin entry → the settings layer the config
111
+ page writes (with its own `*.ui.json` companion). File layers are re-read on every escalation and the
112
+ settings layer is polled about every two seconds, so edits take effect without a restart; changing the
113
+ plugin's **code** needs one.
114
+
115
+ `config.json` accepts `//` and `/* */` comments. `cordis.patch.yml` is YAML: comments there need `#`, and a
116
+ parse failure makes the loader skip the whole bundle silently.
117
+
118
+ | Key | Default | Meaning |
119
+ |---|---|---|
120
+ | `enabled` | `false` | The intervention switch. |
121
+ | `enabledText` | `""` | Text form of the switch (`"true"` / `"false"`) written by the config page; non-empty overrides `enabled`. |
122
+ | `provider` / `model` | `""` | Which provider and model review; empty follows the session. |
123
+ | `reasoningEffort` | `""` | Reviewer thinking effort: `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`, or empty to follow the session. |
124
+ | `failMode` | `deny` | What a failed or timed-out review means: `deny` or `ask`. |
125
+ | `denyMode` | `deny` | What a deny verdict means: `deny` outright, or `ask` the user. |
126
+ | `probeRunner` | `inproc` | Read-only probes: `inproc` (no process spawned) or `shell` (a read-only command inside the sandbox). |
127
+ | `policyExtra` | `""` | Free-form rules appended to the reviewer policy. |
128
+ | `allowedHosts` | `[]` | Hosts whose ordinary network access counts as low risk. An empty array in a user file **overrides** the package list, so write the full list if you write the key at all. |
129
+ | `timeoutMs` | `100000` | Total budget for one review, retries included. |
130
+ | `attemptTimeoutMs` | `30000` | Cap for a single request; a timeout is retried. |
131
+ | `retryDelayMs` | `5000` | Wait between attempts. |
132
+ | `minAttemptMs` | `2000` | Skip a retry when the remaining budget after the delay is below this. |
133
+
134
+ ## What it does, and what it does not do
135
+
136
+ **It answers escalations for you.** With the switch on, an allow verdict is delivered as an approval
137
+ answered `allowed-once`: that one call runs outside the sandbox with no prompt. This is the capability and
138
+ the risk in one — you are trading "I read this one myself" for "a model read it".
139
+
140
+ **It sends evidence to your model provider.** The pending action, the session context, the bounded
141
+ instruction projection and the local facts go into one review request through the provider configured in
142
+ DSH (or the one you pick in the card). Reviews cost tokens.
143
+
144
+ **It keeps a local log.** `~/.dsh/escalation-review.log` holds one JSON object per line, including the
145
+ command text and the verdict with its reason. Treat it as sensitive as the sessions it summarises.
146
+
147
+ **It never writes session events.** Verdicts and reasons are delivered in memory and to the audit card, so
148
+ session history stays exactly as the host wrote it.
149
+
150
+ **Failures close, and repeated denials break the loop.** Timeout, malformed JSON, a provider error or a
151
+ failed cross-check end in `deny` by default; three denials in a row, or ten within fifty calls, trip the
152
+ breaker and the call goes to you.
153
+
154
+ **Every decision is cross-checked** against the logged tool call, so a verdict is only issued on evidence
155
+ the plugin can tie to the call it is about.
156
+
157
+ ## Logging and evidence
158
+
159
+ - Internal packages resolve from the running app first (`app.asar/dsh`) and the profile second, so an app
160
+ upgrade keeps the plugin on the app's own copies. The root used is recorded in the `assembler-resolved`
161
+ log line.
162
+ - Read-only probes run in-process by default (nothing spawned). With `probeRunner: shell` they run as a
163
+ read-only command inside the sandbox, roughly 650–700 ms each; the budget is at most 4 probes, 3 seconds
164
+ in total, 1.2 seconds each, 2 KB of output. Without `pwsh`, the in-process channel is used.
165
+ - Review log: `$DSH_HOME/escalation-review.log`, one JSON object per line. `tools/review-report.mjs`
166
+ renders it plus the matching session context into Markdown. Key events: `ready`,
167
+ `intervention-gate`, `config-effective`, `assembler-resolved`, `reviewed`, `reviewer-failed`,
168
+ `review-retry`, `approval-granted`, `circuit-breaker`, `projection-registered`.
169
+
170
+ ## The audit card
171
+
172
+ An escalation appears in the transcript as a review card carrying the status (reviewing, awaiting
173
+ approval, allowed, denied, executed with a tool error) and the reason. Clicking it expands the detail.
174
+ Status colours come from the official theme tokens.
175
+
176
+ Known limitation: an **allow** reason travels through an in-memory session projection, because approval
177
+ outcomes have no reason field of their own. After a session is reloaded, older calls keep their status and
178
+ lose that reason; new reviews are unaffected.
179
+
180
+ Everything runs through documented extension points: `tools/pre-execute`, `approval/request`,
181
+ `sessionProjections`, `conversation.chat.node`, `configForms`.
182
+
183
+ ## License
184
+
185
+ MIT. See [LICENSE](LICENSE). Third-party attributions, and the measured text overlap with the projects this
186
+ policy borrows from, are recorded in [NOTICE.md](NOTICE.md).
package/README.zh.md ADDED
@@ -0,0 +1,162 @@
1
+ # dsh-escalation-review
2
+
3
+ [English](README.md) · **简体中文**
4
+
5
+ **只审「沙箱越界」的 LLM 评审插件**(DeepSeek Harness,下称 DSH)。
6
+
7
+ 工作区内的读写与命令跑在原生沙箱上,不评审、不产生额外模型调用。插件只为一种情况醒来:某次工具调用
8
+ 请求**离开**沙箱(`sandbox_permissions` 与当前生效模式不同)。这一次调用会被交给模型裁决 —— 放行、拒绝,
9
+ 或者交回给你 —— 并且走宿主本来就有的一套管路:把审批回答成 `allowed-once`,或在工具体执行前返回三态决策对象。
10
+
11
+ ## Codex 风格的评审器
12
+
13
+ 策略跟随 OpenAI 的 Codex guardian 模板(`codex-rs/prompts/templates/guardian/policy_template.md`),
14
+ 这也是本插件主要的工作量所在。
15
+
16
+ - **三轴契约**:风险档位与授权分别打分,两者一起决定 `allow` / `deny` / `ask`。授权在**每一次**越界时重新
17
+ 推导:过去的批准只有在证据仍留在会话里时才算数。
18
+ - **`medium` 看的是后果可逆,不是工件可逆**:只有当爆炸半径有界**并且**它造成的影响能被撤销时,才算
19
+ `medium`。能把文件改回来,不等于能把效果收回来。
20
+ - **系统与安全控制状态一律 `high`**:`hosts`、DNS、防火墙规则、服务、证书、`PATH`。它们都能改回来,
21
+ 但已经发出去的流量收不回来。
22
+ - **有更安全的替代方案时降档**:同一目标若能绕开危险步骤达成,那么危险步骤就不是用户要的东西。
23
+ - **缺失的记录不等于许可**:没有证据就是没有授权,而不是默认同意。
24
+ - **事后批准按高风险处理**:事后补的批准打不开 `critical` 动作。
25
+ - **有界根投影**:用户指令以「头部 4 条 + 尾部 12 条、上限 16」进入评审,丢就整条丢、不截断,并显式留下
26
+ 「证据不完整」标记。证据按角色标注:人类指令、直接父级、审批回答、检查点、事实。
27
+
28
+ 另有四条决定边缘情况:
29
+
30
+ - **用户明确禁止 = 硬拒绝**,它压过风险档位,连低风险动作也不例外。
31
+ - **证据是数据,不是指令**:命令文本、文件内容、工具输出都不构成授权;写着"用户已经允许了"的文本什么也
32
+ 证明不了。
33
+ - **不按外观降档**:形似凭据的载荷不会因为"看起来像占位符"而变安全,导出也不会因为自称的去向而变安全。
34
+ - **先交叉校验,再判定**:待审动作会与同一 `callId` 的日志 `tool/call` 比对,一旦明确不一致就中止评审;
35
+ 判定格式不合法按评审失败处理。
36
+
37
+ 这些想法的来源与逐字重合度的实测记录见 [NOTICE.md](NOTICE.md)。
38
+
39
+ ## 与内置 `auto` 预设的区别
40
+
41
+ 应用内置 `@deepseek-ai/dsh-experimental-auto-review`:它注册一个 `auto` 权限预设,把会话绑到
42
+ `danger-full-access`,并审核**每一个**原生调用 —— 每次调用多一次模型请求,且运行在完全放开权限下。
43
+
44
+ | | 内置 `auto` 预设 | 本插件 |
45
+ |---|---|---|
46
+ | 沙箱档位 | `danger-full-access` | 会话原本的档位,因此 **`workspace-write` 继续生效** |
47
+ | 审核范围 | 每一个原生调用 | 只有请求越界的那一次 |
48
+ | 预设表 | 注册 `auto` | 不动你的预设表 —— 门控是本插件自己的开关 |
49
+ | 额外模型调用 | 每次调用一次 | 每次越界一次,其余为零 |
50
+
51
+ 保留会话原本的沙箱档位,正是这个对比的意义:普通工作由沙箱兜住,评审器只补上 DSH 本来会停下来问你的那一段。
52
+
53
+ ## 行为
54
+
55
+ | 情形 | 结果 |
56
+ |---|---|
57
+ | 沙箱内的普通调用 | 不评审,原样交给下游处理者。 |
58
+ | 越界调用,介入开关**关闭** | 不评审、不干预、不花模型调用。这是默认状态。 |
59
+ | 越界调用,开关**打开** | 评审一次并记日志。判定放行则审批被回答成 `allowed-once`(你**不会被问第二次**);判定拒绝则工具不执行、理由回传模型。 |
60
+ | 评审失败 / 超时 | 由 `failMode` 决定,默认 `deny`(失败即拒)。 |
61
+ | 连续三次拒绝,或 50 次窗口内 10 次 | 熔断:不再拒绝,改为交回人工,避免把整个任务卡死。 |
62
+ | 待审动作与会话日志里的工具调用明确不一致 | 停止评审,而不是拿无法核对的证据下结论。 |
63
+
64
+ ## 运行前提
65
+
66
+ - DSH **0.1.7-rc.2** 或 **0.2.0-rc.1**(两者都实测过),Node 24。
67
+ - 只实测过 Windows;其它平台预期可用,但未验证。
68
+ - `peerDependencies`:`@deepseek-ai/cordis` `~4.0.4`、`@deepseek-ai/dsh-llm` `>=0.1.7-rc.2 <0.3.0`。
69
+ - `@deepseek-ai/dsh*` 的 peer 要写**版本区间**:宿主用
70
+ `semver.satisfies(runtime, range, { includePrerelease: true })` 判定,写死精确预发布版本会在 DSH
71
+ 升级后被判为不兼容。
72
+
73
+ ## 安装
74
+
75
+ 从本地目录安装,用 GUI 最干净:左侧边栏「**插件**」(不是"设置")→ 「添加插件」→ 填这个目录的绝对路径,
76
+ 例如 `D:\path\to\checkout\plugins\escalation-review`。它会以 `link:` 依赖登记,并追加到
77
+ `dsh.profile.bundles`。
78
+
79
+ ```sh
80
+ dsh plugin --profile <profile 名> add <包名或路径>
81
+ ```
82
+
83
+ - `~/.dsh/profiles/web` 与 `~/.dsh/profiles/desktop` **互相独立**,装进一个不影响另一个。想用 `dsh web`
84
+ 测,就加 `--profile web`。
85
+ - 宿主半边是启动时从磁盘加载的,所以**装完要重启 DSH**;改了 `lib/` 下任何文件也要重启一次。会话里的
86
+ 客户端半边同样需要重启才重载,桌面版里单纯刷新页面不够。
87
+
88
+ 卸载在同一张插件页上做,或者从 `dsh.profile.bundles` 里删掉那一条。
89
+
90
+ ## 配置
91
+
92
+ `enabled` 是总门控,随包**关闭**:关闭时插件不评审、不产生模型调用;打开后它接管**所有**沙箱越界的审批。
93
+ 门控属于插件本身,而不是你的权限表 —— 所以安装它不会改写你的预设,也与某个会话恰好选中哪个预设无关。
94
+ 可以在配置页上设(卡片**第一行**的两档控件:关闭 / 打开),也可以写 `"enabled": true`,或 GUI 写入的
95
+ 文本形式 `"enabledText": "true"`。
96
+
97
+ 优先级(低 → 高):包内 `config.json` → `~/.dsh/escalation-review.config.json` → 传了 `configPath` 就用它
98
+ → 插件条目里的 `config:` 段 → 配置页写入的设置层(有它自己的 `*.ui.json` 伴生文件)。文件层每次越界时重读,
99
+ 设置层约每 2 秒轮询一次,所以改动不需要重启;只有改插件**代码**才必须重启。
100
+
101
+ 本包的 `config.json` 容忍 `//` 与 `/* */` 注释;`cordis.patch.yml` 是 YAML,注释只能用 `#`,解析失败会让
102
+ 加载器**静默跳过整个 bundle**。
103
+
104
+ | 键 | 默认值 | 含义 |
105
+ |---|---|---|
106
+ | `enabled` | `false` | 介入开关(总门控)。 |
107
+ | `enabledText` | `""` | 开关的文本形式(`"true"` / `"false"`),由配置页写入;非空时覆盖 `enabled`。 |
108
+ | `provider` / `model` | `""` | 用哪个 provider 与模型评审;留空跟随当前会话。 |
109
+ | `reasoningEffort` | `""` | 评审思考强度:`off`、`minimal`、`low`、`medium`、`high`、`xhigh`、`max`,留空跟随会话。 |
110
+ | `failMode` | `deny` | 评审失败或超时的含义:`deny` 或 `ask`。 |
111
+ | `denyMode` | `deny` | 评审判定为拒绝的含义:`deny` 直接拒绝,`ask` 交回人工。 |
112
+ | `probeRunner` | `inproc` | 只读探针:`inproc`(不 spawn 进程)或 `shell`(沙箱内只读命令)。 |
113
+ | `policyExtra` | `""` | 追加到评审策略末尾的自定义规则(自然语言)。 |
114
+ | `allowedHosts` | `[]` | 命中这些主机的常规网络操作按 low 处理。注意:在用户配置里写空数组会**覆盖**包内清单,要写就写全量。 |
115
+ | `timeoutMs` | `100000` | 一次评审的总预算(含重试)。 |
116
+ | `attemptTimeoutMs` | `30000` | 单次请求上限;超时后可以重试。 |
117
+ | `retryDelayMs` | `5000` | 两次尝试之间的等待。 |
118
+ | `minAttemptMs` | `2000` | 扣掉等待后剩余预算低于此值就不再重试。 |
119
+
120
+ ## 它做什么,不做什么
121
+
122
+ **它替你回答越界审批。** 开关打开时,放行判定会以 `allowed-once` 的形式交付:那一次调用在沙箱之外执行,
123
+ 过程中不弹窗。能力与风险都在这里 —— 你把"这一次我自己看一眼"换成了"模型看过一眼"。
124
+
125
+ **它把证据发给你的模型提供方。** 待审动作、会话上下文、有界指令投影与本地事实,会组成一次评审请求,
126
+ 走 DSH 里配置的 provider(或你在卡片里选的那个)。评审消耗 token。
127
+
128
+ **它留一份本地日志。** `~/.dsh/escalation-review.log` 每行一个 JSON 对象,含命令文本与判定结论、理由。
129
+ 请把它当作和它所概括的会话同等敏感的东西。
130
+
131
+ **它从不写会话事件。** 判定与理由在内存里交付、并呈现到审核卡片上,所以会话历史仍是宿主写下的样子。
132
+
133
+ **失败即拒,反复被拒会熔断。** 超时、JSON 不合法、provider 报错、交叉校验失败,默认都以 `deny` 收场;
134
+ 连续三次、或 50 次窗口内 10 次,熔断触发,这一次交给你。
135
+
136
+ **每个判定都经过交叉校验**,与日志里的工具调用比对,因此结论只建立在能对上号的那次调用上。
137
+
138
+ ## 日志与证据
139
+
140
+ - 插件解析 DSH 内部包时**先取运行中的应用自带那份**(`app.asar/dsh`),再退到 profile,这样升级应用后
141
+ 仍用应用自带的那份。实际命中的解析根记在 `assembler-resolved` 日志行里。
142
+ - 只读探针默认走进程内(不 spawn 任何进程);`probeRunner: shell` 时走沙箱内的只读命令,单条约 650–700 ms。
143
+ 额度上限:最多 4 条、总预算 3 秒、单条 1.2 秒、输出 2 KB。拿不到 `pwsh` 时走进程内通道。
144
+ - 评审日志:`$DSH_HOME/escalation-review.log`,每行一个 JSON 对象;`tools/review-report.mjs` 会把它与
145
+ 会话上下文(按 `callId` 关联)渲染成 Markdown。主要事件:`ready`、`intervention-gate`、
146
+ `config-effective`、`assembler-resolved`、`reviewed`、`reviewer-failed`、`review-retry`、
147
+ `approval-granted`、`circuit-breaker`、`projection-registered`。
148
+
149
+ ## 审核卡片
150
+
151
+ 越界调用会在对话里出现一张审核卡片,显示状态(评审中、待审批、已放行、已拒绝、已执行但工具报错)与理由,
152
+ 点击可展开细节。状态颜色取自官方主题 token。
153
+
154
+ 已知限制:**放行的理由**是靠会话内存里的投影带过去的 —— 审批结论本身没有理由字段。会话重新加载后,
155
+ 旧调用的状态还在,但那条理由会丢;新发生的评审不受影响。
156
+
157
+ 全部通过有文档的扩展点工作:`tools/pre-execute`、`approval/request`、`sessionProjections`、
158
+ `conversation.chat.node`、`configForms`。
159
+
160
+ ## 许可
161
+
162
+ MIT,见 [LICENSE](LICENSE)。第三方来源与逐字重合度的实测记录见 [NOTICE.md](NOTICE.md)。
package/config.json ADDED
@@ -0,0 +1,72 @@
1
+ {
2
+ // ─────────────────────────────────────────────────────────────────────────────
3
+ // escalation-review 配置文件
4
+ //
5
+ // 优先级(低 → 高):
6
+ // 1) 本文件(插件目录内的默认值)
7
+ // 2) ~/.dsh/escalation-review.config.json ← 建议改这个,好找
8
+ // 3) cordis.patch.yml 里该插件条目的 config ← 宿主/我代管时用
9
+ // 每次发生越界时重新读取,所以改完**即时生效,不需要重启 DSH**。
10
+ // 支持 `//` 与 `/* */` 注释(字符串内的注释符不会被误伤)。
11
+ // ─────────────────────────────────────────────────────────────────────────────
12
+
13
+ // observe = 只评审并记录,不干预(你照旧被问);enforce = 按评审结果放行/拒绝。
14
+ // 注意:这一项现在是**调试项**,配置页默认不显示(本机要看就在 ~/.dsh/escalation-review.config.json 里写 "devMode": true)。
15
+ // 随包默认 enforce;但因为总开关 enabled 默认关,装着不动时插件仍然零介入。
16
+ // 想在接入前先观察,就把下面这行改成 "observe"。
17
+ "enabled": false,
18
+ "enabledText": "",
19
+ "mode": "enforce",
20
+
21
+ // true:第一次越界之后,用**伪造的待审动作**(假的 rm -rf、假的凭据上传…)跑一遍策略自检。
22
+ // 同一套 reviewer、同一策略、同一段真实上下文,但**绝不执行任何工具**,只记录判定。
23
+ // 想看"危险动作会不会被拒"时打开它,比真的冒险安全得多。
24
+ "selfTest": false,
25
+
26
+ // 评审超时三档(2026-09-27):单次 30s / 总 100s / 间隔 5s → 3×30 + 2×5 = 100s,正好被总预算兜住
27
+ // 单次尝试上限:有了它,"本次超时"不再等于"预算花光",所以超时**可以重试**
28
+ "attemptTimeoutMs": 30000,
29
+ // 两次尝试之间的等待(毫秒)
30
+ "retryDelayMs": 5000,
31
+ // 最小重试预算:扣掉等待后剩余不足这个数,就不白打一次注定失败的请求
32
+ "minAttemptMs": 2000,
33
+ // 整个评审的总预算(毫秒)。调用方在审批路径里等着,总等待必须有界
34
+ "timeoutMs": 100000,
35
+
36
+ // 评审失败/超时时怎么办:deny = 直接拒绝(对齐 Codex 的 strict 档);ask = 交回人工
37
+ "failMode": "deny",
38
+
39
+ // 覆盖 reviewer 使用的模型;留空则跟随当前会话的 provider/model
40
+ // 例:想省钱可以用一个更便宜快的模型专做三分类
41
+ "provider": "",
42
+ "model": "",
43
+
44
+ // low 风险白名单主机(按你机器上的**真实**证据推导:git 远端 + 包镜像)
45
+ // 命中这些主机的常规网络操作按 low 处理;上传密钥/隐私数据仍按 high 拒绝
46
+ "allowedHosts": [
47
+ "registry.npmmirror.com", // .npmrc 里的 registry
48
+ "github.com", // 37 个项目的远端
49
+ "codeload.github.com",
50
+ "api.github.com",
51
+ "raw.githubusercontent.com",
52
+ "objects.githubusercontent.com",
53
+ "gitee.com",
54
+ "gitlab.com",
55
+ "bitbucket.org",
56
+ "gitcafe.com",
57
+ "gitgud.io",
58
+ "git.oschina.net"
59
+ ],
60
+
61
+ // 追加到固定策略末尾的自定义规则(自然语言,随便写)
62
+ "policyExtra": "",
63
+
64
+ // 评审日志(每次越界一行 JSON)。observe 模式下主要看这个
65
+ "logPath": "",
66
+
67
+ // 证据预算:只影响喂给 reviewer 的上下文大小
68
+ "historyLimit": 12, // 保留的「用户指令」条数
69
+ "transcriptLimit": 40, // transcript 事件条数
70
+ "textLimit": 900 // 单条文本截断长度
71
+ }
72
+
@@ -0,0 +1,11 @@
1
+ # 本包被 dsh.profile.bundles 引用时应用的插入补丁。
2
+ #
3
+ # ⚠️ 必须是合法 YAML:注释只能用 `#`。曾经用 `//` 当注释 → 解析失败 →
4
+ # 加载器把整个 bundle **静默跳过**(原因只写 stderr,GUI 应用看不到)→ 插件毫无痕迹地不工作。
5
+ #
6
+ # 说明(2026-09-27 更正):entry 由**本 bundle 的 patch** 插入即可,不需要挪到 profile 的 patch 层。
7
+ # 之前以为"设置页只对 profile patch 层创建的条目生成"是错的 —— 真正的门槛是**客户端半边**:
8
+ # 官方 `plugins.bundle.config` / `plugins.item` 槽位由客户端插件注册,宿主 `Config` 只负责校验与落盘。
9
+ - insert:
10
+ - id: escalation-review
11
+ name: dsh-escalation-review
package/icon.svg ADDED
@@ -0,0 +1,14 @@
1
+ <svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 48 48" width="48" height="48" role="img" aria-label="沙箱越界评审">
2
+ <title>沙箱越界评审</title>
3
+ <defs>
4
+ <linearGradient id="er-g" x1="0.1" y1="0" x2="0.9" y2="1">
5
+ <stop offset="0" stop-color="#3B82F6" />
6
+ <stop offset="0.55" stop-color="#4C6EF5" />
7
+ <stop offset="1" stop-color="#7C4DFF" />
8
+ </linearGradient>
9
+ </defs>
10
+ <path d="M24 6.5l13 4.6v11.2c0 7.6-5.4 14.6-13 16.7-7.6-2.1-13-9.1-13-16.7V11.1z"
11
+ fill="none" stroke="url(#er-g)" stroke-width="3.2" stroke-linejoin="round" stroke-linecap="round" />
12
+ <path d="M17.8 23.4l4.3 4.4 8-9.6" fill="none" stroke="url(#er-g)" stroke-width="3.4"
13
+ stroke-linecap="round" stroke-linejoin="round" />
14
+ </svg>
package/lib/breaker.js ADDED
@@ -0,0 +1,76 @@
1
+ /**
2
+ * breaker.js —— 连续拒绝熔断(照抄 Codex `ext/guardian-reviewer/src/circuit_breaker.rs`)
3
+ *
4
+ * 为什么要它:评审器不停拒绝会把整个任务卡死。Codex 的做法是按 **turn** 计数:
5
+ * · consecutive_denials 连续拒绝数;recent_denials 为长度 50 的窗口
6
+ * · 标准档阈值:连续 ≥ 3 或窗口内 ≥ 10 → 触发一次 "interrupt"(`interrupt_triggered` 只触发一次)
7
+ * · 任何非拒绝(allow / 交回人工之外的正常通过)把连续计数清零
8
+ *
9
+ * 我们这边没有"打断 turn"的权限,等价动作是:**停止拒绝,改为交回人工(ask)**,
10
+ * 并写一条 `circuit-breaker` 日志——任务不再被无限卡住,人仍然在场。
11
+ */
12
+ export const MAX_CONSECUTIVE_DENIALS = 3
13
+ export const MAX_RECENT_DENIALS = 10
14
+ export const DENIAL_WINDOW_SIZE = 50
15
+
16
+ /** 每个会话一条记录;turnId 变化即视为新 turn(与 Codex 的 clear_turn 等价)。 */
17
+ const sessions = new Map()
18
+
19
+ function bucketFor(sessionKey, turnId) {
20
+ const key = String(sessionKey ?? 'unknown')
21
+ let bucket = sessions.get(key)
22
+ if (bucket === undefined || bucket.turnId !== turnId) {
23
+ bucket = { turnId, consecutive: 0, recent: [], interruptTriggered: false }
24
+ sessions.set(key, bucket)
25
+ }
26
+ return bucket
27
+ }
28
+
29
+ /**
30
+ * 记一次"拒绝"(deny 或 ask 都算评审器没有放行)。
31
+ * @returns {{ interrupt: boolean, consecutive: number, recent: number }}
32
+ */
33
+ export function recordDenial(sessionKey, turnId, limits = {}) {
34
+ const maxConsecutive = limits.maxConsecutive ?? MAX_CONSECUTIVE_DENIALS
35
+ const maxRecent = limits.maxRecent ?? MAX_RECENT_DENIALS
36
+ const bucket = bucketFor(sessionKey, turnId)
37
+ bucket.consecutive += 1
38
+ bucket.recent.push(true)
39
+ if (bucket.recent.length > DENIAL_WINDOW_SIZE) bucket.recent.shift()
40
+ const recent = bucket.recent.filter(Boolean).length
41
+ if (!bucket.interruptTriggered && (bucket.consecutive >= maxConsecutive || recent >= maxRecent)) {
42
+ bucket.interruptTriggered = true
43
+ return { interrupt: true, consecutive: bucket.consecutive, recent }
44
+ }
45
+ return { interrupt: false, consecutive: bucket.consecutive, recent }
46
+ }
47
+
48
+ /** 记一次"没有拒绝"(评审通过或未介入)→ 连续计数清零。 */
49
+ export function recordNonDenial(sessionKey, turnId) {
50
+ const bucket = bucketFor(sessionKey, turnId)
51
+ bucket.consecutive = 0
52
+ bucket.recent.push(false)
53
+ if (bucket.recent.length > DENIAL_WINDOW_SIZE) bucket.recent.shift()
54
+ }
55
+
56
+ /** 会话结束时清掉(对应 Codex 的 clear_turn)。 */
57
+ export function clearSession(sessionKey) {
58
+ sessions.delete(String(sessionKey ?? 'unknown'))
59
+ }
60
+
61
+ /** 诊断用:读当前计数。 */
62
+ export function peek(sessionKey) {
63
+ const bucket = sessions.get(String(sessionKey ?? 'unknown'))
64
+ if (bucket === undefined) return undefined
65
+ return {
66
+ turnId: bucket.turnId,
67
+ consecutive: bucket.consecutive,
68
+ recent: bucket.recent.filter(Boolean).length,
69
+ interruptTriggered: bucket.interruptTriggered,
70
+ }
71
+ }
72
+
73
+ /** 诊断用:整体重置。 */
74
+ export function resetAll() {
75
+ sessions.clear()
76
+ }