@duke-dsh-plugins/dsh-agent-approval 1.9.0 → 1.10.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +9 -10
- package/client.js +175 -171
- package/index.js +107 -224
- package/package.json +2 -2
- package/typert.host.js +13 -5
package/README.md
CHANGED
|
@@ -23,11 +23,11 @@
|
|
|
23
23
|
| 🛡 **权限菜单第四项** | `/permission` 菜单新增 **自动审批** 预设;选中即开启,切到其他预设自动关闭,跨重启保持 |
|
|
24
24
|
| 🤖 新权限模式 | 开启后:沙箱基线固定 `workspace-write`,审批策略切到 `ask`(内部接管),**不再弹人工审批** |
|
|
25
25
|
| 🤖 自动裁决(默认 LLM 直连) | 每次提权请求由**一次 LLM 直连调用**裁决(v1.8.0 起默认):与子代理同一套审批人格 / 提示词 / 结构化裁决 `{decision, riskLevel, rationale}`,但**不创建审批子会话**(零上下文污染);设置页可切回「隔离子代理」(一次性 `spawn` 子代理:独立会话、零工具、只读材料) |
|
|
26
|
-
| 🕵️ 自动审查(逐调用,实验) |
|
|
26
|
+
| 🕵️ 自动审查(逐调用,实验) | 与自动审批**完全独立**的第二种模式(需先在设置页配置 Jev,配置即启用):**Full access 基线**,每个工具调用(含 PTC 内层)执行前经 Jev 判定一次,风险调用**直接拒绝、body 不执行、不转人工**(fail-closed);规则表与会话内信任先行短路降噪;`/agent-review on\|off` 或菜单「自动审查」开启。v1.9.0 起**有效代码展开**:命令引用的解释器脚本(node / python / bash / ps1)内容会被读出并放进裁决 state,agent 临时写的脚本按**实际代码**受审而不是只看命令行(仅读 workspace 内文件,越界只标注、永不读取) |
|
|
27
27
|
| ⛔ 风险即拒绝 | 破坏性 / 不可逆 / 越界(含修改操作系统或其他应用数据)/ 理由与实际命令不符 → 直接 `reject`;仅"安全、可逆、与任务相符、理由诚实"才 `approve`——项目自身的安装/部署脚本写其文档指定路径属任务所需 |
|
|
28
28
|
| 🔒 Fail-closed | 审批 Agent 启动失败、超时(可配 30s–600s)、结果不合法 → 一律按拒绝处理,绝不静默放行 |
|
|
29
|
-
| ⚙️ 审批模型可配置 | 设置页选择 Provider + Model
|
|
30
|
-
| ⚡ TypeSafe Jev
|
|
29
|
+
| ⚙️ 审批模型可配置 | 设置页选择 Provider + Model(**只能是 LLM 路由**),不选则固定用 **Harness 默认模型**(不跟随请求会话,口径稳定);选择与超时**持久保存**,重启不丢 |
|
|
30
|
+
| ⚡ TypeSafe Jev 决策模型(仅自动审查) | Jev(System One 结构化决策模型)是**自动审查模式专属**的判定后端:设置页「自动审查」卡片独立配置(API Key / Endpoint / 模型 / 置信度阈值),直连其 API 用类型化问题(Choice/Noul)毫秒级返回带校准概率的裁决;置信度低于阈值按 fail-closed 处理,审计理由由概率合成(也可设 `TYPESAFE_API_KEY`)。**v1.10.0 起不再是自动审批的判定选项**——Jev 的风险判断弱于 LLM,而提权正是人工审批原本要守住的路径 |
|
|
31
31
|
| 📋 审计记录(随会话) | 会话窗口顶部的**「审批」标签页**(轨迹旁)查看本会话全部审批:结论 / 风险等级 / 模型 / 耗时 / 理由;悬停看完整理由与**精确工具参数**;审批 Agent 的会话 id 可回溯完整推理;已批准行可一键**「加白」**存为放行规则。记录存在**会话存储目录内的独立文件**——随会话恢复,删除会话即随之删除 |
|
|
32
32
|
| 🔁 可逆开关 | 权限菜单「自动审批」预设、`/agent-approval on\|off` 命令两条等价路径;关闭时**恢复开启前的权限旋钮** |
|
|
33
33
|
|
|
@@ -42,10 +42,9 @@
|
|
|
42
42
|
└─ 本插件 prepend 抢占(先于人工弹窗 answerer)
|
|
43
43
|
├─ (默认)LLM 直连裁决(一次 stream 调用 · 同套人格/提示词/裁决格式 · 零子会话)
|
|
44
44
|
├─ (可选)spawn 审批 Agent(独立会话 · 零工具 · 结构化裁决 · 不会递归审批)
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
└─ 超时/故障/低置信/取消 → fail-closed(按拒绝处理)
|
|
45
|
+
│ ├─ approve → allowed-once(该次放行)
|
|
46
|
+
│ ├─ reject → rejected(风险操作,最终拒绝)
|
|
47
|
+
│ └─ 超时/故障/低置信/取消 → fail-closed(按拒绝处理)
|
|
49
48
|
└─ 记入审计(会话目录内的独立文件,「审批」标签页可见)
|
|
50
49
|
```
|
|
51
50
|
|
|
@@ -81,15 +80,15 @@ dsh plugin --profile web add https://github.com/MoonlitDropOfBlood/dsh-agent-app
|
|
|
81
80
|
1. **开启**:在 `/permission` 菜单选 **自动审批**,或输入 `/agent-approval on`。
|
|
82
81
|
2. **自动裁决**:之后该会话里的提权请求(例如命令被沙箱拒绝后带 `sandbox_permissions` 的重试)不再弹窗,由审批 Agent 在后台裁决并放行/拒绝。
|
|
83
82
|
3. **审计**:会话窗口顶部的**「审批」标签页**(轨迹旁)查看本会话的审批记录;悬停"审批理由"看完整理由与工具参数;已批准行可「加白」存为放行规则。记录存在会话存储目录内的独立文件,删除会话即随之删除;v1.4 的旧全局记录用 `node scripts/migrate-records.mjs` 一次性迁移(`--dry-run` 预览)。
|
|
84
|
-
4.
|
|
85
|
-
5.
|
|
83
|
+
4. **配置自动审批**:设置 → **自动审批** 设置审批模型(Harness 默认模型或指定 Provider/Model——**只能是 LLM 路由**)、**裁决方式**(LLM 直连 / 隔离子代理)、审批超时与放行/拒绝规则。
|
|
84
|
+
5. **配置自动审查(独立)**:设置页「自动审查」卡片填 Jev 的 API Key(可调 Endpoint / 模型版本 / 置信度阈值)——**保存即启用**逐调用审查(新开会话自动进入,菜单出现 **自动审查** 预设,或执行 `/agent-review on`):该会话切到 Full access 基线,**每个工具调用**执行前经 Jev 判定,风险调用直接拒绝(不转人工);关闭恢复原旋钮。
|
|
86
85
|
6. **关闭**:菜单切回其他预设,或 `/agent-approval off`,恢复开启前的沙箱模式与审批策略。
|
|
87
86
|
|
|
88
87
|
## 目录结构
|
|
89
88
|
|
|
90
89
|
```
|
|
91
90
|
dsh-agent-approval/
|
|
92
|
-
├── index.js # Host 半:AgentApprovalService(审批瀑布抢占 + LLM
|
|
91
|
+
├── index.js # Host 半:AgentApprovalService(审批瀑布抢占 + LLM 直连/子代理裁决 + Jev 逐调用审查 + 审计)
|
|
93
92
|
├── client.js # Client 半:设置页「自动审批」+ 会话「审批」审计标签页 bundle
|
|
94
93
|
├── typert.host.js # Typert Host manifest(agentApproval 10 个 Remote 方法的描述)
|
|
95
94
|
├── cordis.patch.yml # dsh bundle patch(挂载行 + permission 预设表覆盖)
|
package/client.js
CHANGED
|
@@ -11,12 +11,16 @@
|
|
|
11
11
|
* session — restored with it after a restart, gone when the session is
|
|
12
12
|
* deleted. Rows offer the one-click「加白」rule shortcut.
|
|
13
13
|
*
|
|
14
|
-
* 2. A "自动审批" page in the Settings panel (`settings.section`)
|
|
15
|
-
*
|
|
16
|
-
*
|
|
17
|
-
*
|
|
18
|
-
*
|
|
19
|
-
*
|
|
14
|
+
* 2. A "自动审批" page in the Settings panel (`settings.section`). It
|
|
15
|
+
* presents the two modes as SEPARATE blocks since v1.10.0:
|
|
16
|
+
* - 自动审批: approval model picker (harness provider + model, or the
|
|
17
|
+
* harness default), judge invocation mode (direct LLM vs isolated
|
|
18
|
+
* subagent — an LLM either way, Jev is not selectable here), the
|
|
19
|
+
* fail-closed timeout, enabled sessions and the allow/deny rules.
|
|
20
|
+
* - 自动审查: the per-call review mode's own TypeSafe Jev settings
|
|
21
|
+
* (API key / model / endpoint / confidence gate) plus the global
|
|
22
|
+
* "new sessions enter 自动审查" switch. Configuring Jev here is what
|
|
23
|
+
* enables the mode; the approval judge is untouched by it.
|
|
20
24
|
*
|
|
21
25
|
* Session-level on/off lives in the /permission menu (the "自动审批"
|
|
22
26
|
* preset, registered by the package's cordis.patch.yml bundle patch) and the
|
|
@@ -551,8 +555,9 @@ window.__ModuleLoader__.load({
|
|
|
551
555
|
const setReviewDefault = reviewDefaultSlot[1];
|
|
552
556
|
const timeoutSlot = React.useState("");
|
|
553
557
|
const setTimeoutDraft = timeoutSlot[1];
|
|
554
|
-
// Jev backend drafts
|
|
555
|
-
//
|
|
558
|
+
// Jev backend drafts — the 自动审查 card's own configuration, shown
|
|
559
|
+
// unconditionally since v1.10.0 (it no longer hangs off a judge
|
|
560
|
+
// provider selection).
|
|
556
561
|
const jevKeySlot = React.useState("");
|
|
557
562
|
const setJevKey = jevKeySlot[1];
|
|
558
563
|
const jevEndpointSlot = React.useState("");
|
|
@@ -667,9 +672,19 @@ window.__ModuleLoader__.load({
|
|
|
667
672
|
}
|
|
668
673
|
remote
|
|
669
674
|
.setJevConfig({ apiKey: jevKey, endpoint: jevEndpoint, model: jevModel, confidence: conf })
|
|
670
|
-
.then(() => {
|
|
675
|
+
.then((res) => {
|
|
676
|
+
const v = pick(res) || {};
|
|
677
|
+
setReviewGate(v.reviewAvailable === true);
|
|
678
|
+
setReviewDefault(v.reviewDefault === true);
|
|
671
679
|
refresh();
|
|
672
|
-
|
|
680
|
+
// Configuring Jev opens the review gate; the Host turns the
|
|
681
|
+
// per-call review default on at that exact moment (v1.10.0).
|
|
682
|
+
setNote(
|
|
683
|
+
v.reviewAvailable === true
|
|
684
|
+
? "Jev 配置已保存:自动审查已就绪" +
|
|
685
|
+
(v.reviewDefault === true ? ",新开会话自动进入逐调用审查" : "")
|
|
686
|
+
: "Jev 配置已保存,但 API Key 仍为空:自动审查保持关闭",
|
|
687
|
+
);
|
|
673
688
|
})
|
|
674
689
|
.catch((e) => setNote("保存失败:" + (e && e.message ? e.message : String(e))));
|
|
675
690
|
};
|
|
@@ -735,30 +750,22 @@ window.__ModuleLoader__.load({
|
|
|
735
750
|
.catch(() => {});
|
|
736
751
|
};
|
|
737
752
|
|
|
738
|
-
//
|
|
739
|
-
//
|
|
753
|
+
// v1.10.0: the harness model directory is the WHOLE list — Jev is no
|
|
754
|
+
// longer a judge provider (it judges the per-call review mode only),
|
|
755
|
+
// so there is no synthetic "typesafe" entry to prepend.
|
|
740
756
|
const providerOptions = [{ id: "", name: "默认(Harness 默认模型)" }].concat(
|
|
741
|
-
[{ id: "typesafe", name: "TypeSafe Jev(决策模型·直连 API)" }],
|
|
742
757
|
dir ? dir.providers : [],
|
|
743
758
|
);
|
|
744
|
-
const modelOptions =
|
|
745
|
-
provider ===
|
|
746
|
-
|
|
747
|
-
{ provider: "typesafe", id: "jev-latest", name: "jev-latest(跟随最新版本)" },
|
|
748
|
-
{ provider: "typesafe", id: "jev-1.13.0", name: "jev-1.13.0(锁定版本)" },
|
|
749
|
-
]
|
|
750
|
-
: [{ provider: "", id: "", name: "默认(Harness 默认模型)" }].concat(
|
|
751
|
-
dir && dir.models ? dir.models.filter((m) => m.provider === provider) : [],
|
|
752
|
-
);
|
|
759
|
+
const modelOptions = [{ provider: "", id: "", name: "默认(Harness 默认模型)" }].concat(
|
|
760
|
+
dir && dir.models ? dir.models.filter((m) => m.provider === provider) : [],
|
|
761
|
+
);
|
|
753
762
|
const defaultHint =
|
|
754
|
-
|
|
755
|
-
? "
|
|
756
|
-
|
|
757
|
-
|
|
758
|
-
|
|
759
|
-
|
|
760
|
-
dir.defaultSelection.model
|
|
761
|
-
: "未配置时使用 Harness 默认模型路由";
|
|
763
|
+
dir && dir.defaultSelection
|
|
764
|
+
? "未配置时使用 Harness 默认模型;当前默认路由:" +
|
|
765
|
+
dir.defaultSelection.provider +
|
|
766
|
+
" / " +
|
|
767
|
+
dir.defaultSelection.model
|
|
768
|
+
: "未配置时使用 Harness 默认模型路由";
|
|
762
769
|
|
|
763
770
|
return h(
|
|
764
771
|
"div",
|
|
@@ -766,13 +773,17 @@ window.__ModuleLoader__.load({
|
|
|
766
773
|
h(
|
|
767
774
|
"div",
|
|
768
775
|
{ className: "aapr-card" },
|
|
769
|
-
h("h3", null, "
|
|
776
|
+
h("h3", null, "自动审批与自动审查"),
|
|
770
777
|
h(
|
|
771
778
|
"div",
|
|
772
779
|
{ className: "aapr-muted" },
|
|
773
|
-
"
|
|
780
|
+
"两种互相独立的权限模式。",
|
|
781
|
+
h("br", null),
|
|
782
|
+
"① 自动审批:workspace-write 基线;工具请求提权时由 LLM 审批 Agent 自动裁决(v1.10.0 起判定模型只能是 LLM)。在输入框 /permission 菜单选择「自动审批」预设,或执行 /agent-approval on|off 为会话开启。",
|
|
774
783
|
h("br", null),
|
|
775
|
-
"
|
|
784
|
+
"② 自动审查:danger-full-access 基线;每个工具调用执行前经 TypeSafe Jev 判定一次(见下方独立卡片)。在 /permission 菜单选择「自动审查」预设,或执行 /agent-review on|off。",
|
|
785
|
+
h("br", null),
|
|
786
|
+
"每个会话的审批审计记录在该会话窗口顶部的「审批」标签页(轨迹旁),随会话保存。",
|
|
776
787
|
),
|
|
777
788
|
note !== "" ? h("div", { className: "aapr-muted" }, note) : null,
|
|
778
789
|
),
|
|
@@ -794,7 +805,7 @@ window.__ModuleLoader__.load({
|
|
|
794
805
|
value: provider,
|
|
795
806
|
onChange: (e) => {
|
|
796
807
|
setProvider(e.target.value);
|
|
797
|
-
setModel(
|
|
808
|
+
setModel("");
|
|
798
809
|
},
|
|
799
810
|
},
|
|
800
811
|
providerOptions.map((p) =>
|
|
@@ -812,7 +823,7 @@ window.__ModuleLoader__.load({
|
|
|
812
823
|
className: "aapr-select",
|
|
813
824
|
value: model,
|
|
814
825
|
onChange: (e) => setModel(e.target.value),
|
|
815
|
-
disabled: dir === null
|
|
826
|
+
disabled: dir === null,
|
|
816
827
|
},
|
|
817
828
|
modelOptions.map((m) =>
|
|
818
829
|
h("option", { key: m.provider + "/" + m.id, value: m.id }, m.id === "" ? m.name : m.name + "(" + m.id + ")"),
|
|
@@ -822,144 +833,33 @@ window.__ModuleLoader__.load({
|
|
|
822
833
|
h(ui.Button, { variant: "primary", size: "sm", onClick: saveModel }, "保存"),
|
|
823
834
|
),
|
|
824
835
|
h("div", { className: "aapr-muted" }, defaultHint),
|
|
825
|
-
|
|
826
|
-
|
|
827
|
-
|
|
828
|
-
|
|
829
|
-
|
|
830
|
-
|
|
831
|
-
|
|
832
|
-
"裁决方式:",
|
|
833
|
-
h(
|
|
834
|
-
"select",
|
|
835
|
-
{
|
|
836
|
-
className: "aapr-select",
|
|
837
|
-
value: judgeMode,
|
|
838
|
-
onChange: (e) => saveJudgeMode(e.target.value),
|
|
839
|
-
},
|
|
840
|
-
h("option", { value: "llm" }, "LLM 直连(默认,不创建子会话)"),
|
|
841
|
-
h("option", { value: "subagent" }, "隔离子代理(legacy,创建审批子会话)"),
|
|
842
|
-
),
|
|
843
|
-
),
|
|
844
|
-
)
|
|
845
|
-
: null,
|
|
846
|
-
provider !== "typesafe"
|
|
847
|
-
? h(
|
|
848
|
-
"div",
|
|
849
|
-
{ className: "aapr-muted" },
|
|
850
|
-
"LLM 直连与隔离子代理是同一套审批人格、提示词与裁决格式({decision, riskLevel, rationale}),仅调用方式不同:前者一次直连模型调用完成裁决、不启动审批子代理(零上下文污染);后者每次裁决创建一个独立子会话(v1.8.0 前的唯一行为)。TypeSafe Jev 后端不受此设置影响。",
|
|
851
|
-
)
|
|
852
|
-
: null,
|
|
853
|
-
),
|
|
854
|
-
provider === "typesafe"
|
|
855
|
-
? h(
|
|
856
|
-
"div",
|
|
857
|
-
{ className: "aapr-card" },
|
|
858
|
-
h("h3", null, "TypeSafe Jev 配置"),
|
|
859
|
-
h(
|
|
860
|
-
"div",
|
|
861
|
-
{ className: "aapr-muted" },
|
|
862
|
-
"Jev 是结构化决策模型(System One):审批时直连 TypeSafe API,不创建审批子会话,毫秒级返回带校准概率的裁决。审计「理由」由概率分布合成(Jev 本身不生成文字);置信度低于阈值时按 fail-closed 处理(记 unavailable,不批准也不记拒绝)。对中文任务上下文的准确率略低于英语。API Key 明文保存在本机 config.json;留空时使用环境变量 TYPESAFE_API_KEY。",
|
|
863
|
-
),
|
|
864
|
-
h(
|
|
865
|
-
"div",
|
|
866
|
-
{ className: "aapr-row" },
|
|
867
|
-
h(
|
|
868
|
-
"label",
|
|
869
|
-
null,
|
|
870
|
-
"API Key:",
|
|
871
|
-
h("input", {
|
|
872
|
-
className: "aapr-input",
|
|
873
|
-
type: "password",
|
|
874
|
-
placeholder: "TYPESAFE_API_KEY",
|
|
875
|
-
value: jevKey,
|
|
876
|
-
onChange: (e) => setJevKey(e.target.value),
|
|
877
|
-
}),
|
|
878
|
-
),
|
|
879
|
-
h(
|
|
880
|
-
"label",
|
|
881
|
-
null,
|
|
882
|
-
"模型:",
|
|
883
|
-
h("input", {
|
|
884
|
-
className: "aapr-input",
|
|
885
|
-
placeholder: "jev-latest",
|
|
886
|
-
value: jevModel,
|
|
887
|
-
onChange: (e) => setJevModel(e.target.value),
|
|
888
|
-
}),
|
|
889
|
-
),
|
|
890
|
-
),
|
|
891
|
-
h(
|
|
892
|
-
"div",
|
|
893
|
-
{ className: "aapr-row" },
|
|
894
|
-
h(
|
|
895
|
-
"label",
|
|
896
|
-
null,
|
|
897
|
-
"Endpoint:",
|
|
898
|
-
h("input", {
|
|
899
|
-
className: "aapr-input aapr-input-wide",
|
|
900
|
-
placeholder: "https://api.typesafe.ai/v1/systemone",
|
|
901
|
-
value: jevEndpoint,
|
|
902
|
-
onChange: (e) => setJevEndpoint(e.target.value),
|
|
903
|
-
}),
|
|
904
|
-
),
|
|
905
|
-
h(
|
|
906
|
-
"label",
|
|
907
|
-
null,
|
|
908
|
-
"置信度阈值:",
|
|
909
|
-
h("input", {
|
|
910
|
-
className: "aapr-input",
|
|
911
|
-
type: "number",
|
|
912
|
-
step: "0.05",
|
|
913
|
-
min: "0.01",
|
|
914
|
-
max: "0.99",
|
|
915
|
-
value: jevConf,
|
|
916
|
-
onChange: (e) => setJevConf(e.target.value),
|
|
917
|
-
}),
|
|
918
|
-
),
|
|
919
|
-
h(ui.Button, { variant: "primary", size: "sm", onClick: saveJev }, "保存"),
|
|
920
|
-
),
|
|
921
|
-
)
|
|
922
|
-
: null,
|
|
923
|
-
provider === "typesafe"
|
|
924
|
-
? h(
|
|
925
|
-
"div",
|
|
926
|
-
{ className: "aapr-card" },
|
|
927
|
-
h("h3", null, "自动审查(逐调用审查)"),
|
|
928
|
-
h(
|
|
929
|
-
"div",
|
|
930
|
-
{ className: "aapr-row" },
|
|
931
|
-
h(
|
|
932
|
-
"label",
|
|
933
|
-
null,
|
|
934
|
-
"逐调用审查:",
|
|
935
|
-
h(
|
|
936
|
-
"select",
|
|
937
|
-
{
|
|
938
|
-
className: "aapr-select",
|
|
939
|
-
value: reviewDefault ? "on" : "off",
|
|
940
|
-
disabled: state !== null && state.reviewAvailable !== true,
|
|
941
|
-
onChange: (e) => saveReviewDefault(e.target.value === "on"),
|
|
942
|
-
},
|
|
943
|
-
h("option", { value: "on" }, "开(新会话自动审查每个工具调用)"),
|
|
944
|
-
h("option", { value: "off" }, "关(新会话按默认预设)"),
|
|
945
|
-
),
|
|
946
|
-
),
|
|
947
|
-
),
|
|
836
|
+
h(
|
|
837
|
+
"div",
|
|
838
|
+
{ className: "aapr-row" },
|
|
839
|
+
h(
|
|
840
|
+
"label",
|
|
841
|
+
null,
|
|
842
|
+
"裁决方式:",
|
|
948
843
|
h(
|
|
949
|
-
"
|
|
950
|
-
{
|
|
951
|
-
|
|
952
|
-
|
|
953
|
-
|
|
954
|
-
|
|
955
|
-
|
|
956
|
-
|
|
957
|
-
: state.reviewAvailable === true
|
|
958
|
-
? "当前状态:Jev 判定可用。"
|
|
959
|
-
: "当前状态:Jev 判定不可用(缺 API Key),开关已禁用,新会话不会自动审查。",
|
|
844
|
+
"select",
|
|
845
|
+
{
|
|
846
|
+
className: "aapr-select",
|
|
847
|
+
value: judgeMode,
|
|
848
|
+
onChange: (e) => saveJudgeMode(e.target.value),
|
|
849
|
+
},
|
|
850
|
+
h("option", { value: "llm" }, "LLM 直连(默认,不创建子会话)"),
|
|
851
|
+
h("option", { value: "subagent" }, "隔离子代理(legacy,创建审批子会话)"),
|
|
960
852
|
),
|
|
961
|
-
)
|
|
962
|
-
|
|
853
|
+
),
|
|
854
|
+
),
|
|
855
|
+
h(
|
|
856
|
+
"div",
|
|
857
|
+
{ className: "aapr-muted" },
|
|
858
|
+
"LLM 直连与隔离子代理是同一套审批人格、提示词与裁决格式({decision, riskLevel, rationale}),仅调用方式不同:前者一次直连模型调用完成裁决、不启动审批子代理(零上下文污染);后者每次裁决创建一个独立子会话(v1.8.0 前的唯一行为)。",
|
|
859
|
+
h("br", null),
|
|
860
|
+
"自动审批的判定模型只能是 LLM:TypeSafe Jev 的结构化决策在此处不可选(它的风险判断比 LLM 弱,而这里正是人工审批原本要守住的路径)。Jev 只用于下方的「自动审查」模式,两者互不影响。",
|
|
861
|
+
),
|
|
862
|
+
),
|
|
963
863
|
h(
|
|
964
864
|
"div",
|
|
965
865
|
{ className: "aapr-card" },
|
|
@@ -977,6 +877,110 @@ window.__ModuleLoader__.load({
|
|
|
977
877
|
h(ui.Button, { variant: "primary", size: "sm", onClick: saveTimeout }, "保存"),
|
|
978
878
|
),
|
|
979
879
|
),
|
|
880
|
+
h(
|
|
881
|
+
"div",
|
|
882
|
+
{ className: "aapr-card" },
|
|
883
|
+
h("h3", null, "自动审查(逐调用审查)"),
|
|
884
|
+
h(
|
|
885
|
+
"div",
|
|
886
|
+
{ className: "aapr-muted" },
|
|
887
|
+
"自动审查是独立的第二种权限模式,与上面的自动审批互不相关:它不看提权,而是以 danger-full-access 为基线,对每个工具调用(含 PTC 内层调用,外层 run_code 传输除外)在执行前用 TypeSafe Jev 判定一次——风险调用直接拒绝、body 不执行、不转人工(fail-closed,拒绝即最终结论)。",
|
|
888
|
+
h("br", null),
|
|
889
|
+
"Jev 是结构化决策模型(System One):直连 TypeSafe API,不创建审批子会话,毫秒级返回带校准概率的裁决。审计「理由」由概率分布合成(Jev 本身不生成文字);置信度低于阈值时按 fail-closed 处理(记 unavailable,不放行也不记拒绝)。对中文任务上下文的准确率略低于英语。API Key 明文保存在本机 config.json;留空时使用环境变量 TYPESAFE_API_KEY。",
|
|
890
|
+
h("br", null),
|
|
891
|
+
"保存一个可用的 API Key 即完成配置,自动审查随之启用(新开会话自动进入);下面同一张卡片可以随时关掉它。",
|
|
892
|
+
),
|
|
893
|
+
h(
|
|
894
|
+
"div",
|
|
895
|
+
{ className: "aapr-row" },
|
|
896
|
+
h(
|
|
897
|
+
"label",
|
|
898
|
+
null,
|
|
899
|
+
"逐调用审查:",
|
|
900
|
+
h(
|
|
901
|
+
"select",
|
|
902
|
+
{
|
|
903
|
+
className: "aapr-select",
|
|
904
|
+
value: reviewDefault ? "on" : "off",
|
|
905
|
+
disabled: state !== null && state.reviewAvailable !== true,
|
|
906
|
+
onChange: (e) => saveReviewDefault(e.target.value === "on"),
|
|
907
|
+
},
|
|
908
|
+
h("option", { value: "on" }, "开(新会话自动审查每个工具调用)"),
|
|
909
|
+
h("option", { value: "off" }, "关(新会话按默认预设)"),
|
|
910
|
+
),
|
|
911
|
+
),
|
|
912
|
+
),
|
|
913
|
+
h(
|
|
914
|
+
"div",
|
|
915
|
+
{ className: "aapr-row" },
|
|
916
|
+
h(
|
|
917
|
+
"label",
|
|
918
|
+
null,
|
|
919
|
+
"API Key:",
|
|
920
|
+
h("input", {
|
|
921
|
+
className: "aapr-input",
|
|
922
|
+
type: "password",
|
|
923
|
+
placeholder: "TYPESAFE_API_KEY",
|
|
924
|
+
value: jevKey,
|
|
925
|
+
onChange: (e) => setJevKey(e.target.value),
|
|
926
|
+
}),
|
|
927
|
+
),
|
|
928
|
+
h(
|
|
929
|
+
"label",
|
|
930
|
+
null,
|
|
931
|
+
"模型:",
|
|
932
|
+
h("input", {
|
|
933
|
+
className: "aapr-input",
|
|
934
|
+
placeholder: "jev-latest",
|
|
935
|
+
value: jevModel,
|
|
936
|
+
onChange: (e) => setJevModel(e.target.value),
|
|
937
|
+
}),
|
|
938
|
+
),
|
|
939
|
+
),
|
|
940
|
+
h(
|
|
941
|
+
"div",
|
|
942
|
+
{ className: "aapr-row" },
|
|
943
|
+
h(
|
|
944
|
+
"label",
|
|
945
|
+
null,
|
|
946
|
+
"Endpoint:",
|
|
947
|
+
h("input", {
|
|
948
|
+
className: "aapr-input aapr-input-wide",
|
|
949
|
+
placeholder: "https://api.typesafe.ai/v1/systemone",
|
|
950
|
+
value: jevEndpoint,
|
|
951
|
+
onChange: (e) => setJevEndpoint(e.target.value),
|
|
952
|
+
}),
|
|
953
|
+
),
|
|
954
|
+
h(
|
|
955
|
+
"label",
|
|
956
|
+
null,
|
|
957
|
+
"置信度阈值:",
|
|
958
|
+
h("input", {
|
|
959
|
+
className: "aapr-input",
|
|
960
|
+
type: "number",
|
|
961
|
+
step: "0.05",
|
|
962
|
+
min: "0.01",
|
|
963
|
+
max: "0.99",
|
|
964
|
+
value: jevConf,
|
|
965
|
+
onChange: (e) => setJevConf(e.target.value),
|
|
966
|
+
}),
|
|
967
|
+
),
|
|
968
|
+
h(ui.Button, { variant: "primary", size: "sm", onClick: saveJev }, "保存"),
|
|
969
|
+
),
|
|
970
|
+
h(
|
|
971
|
+
"div",
|
|
972
|
+
{ className: "aapr-muted" },
|
|
973
|
+
"规则表与会话内信任缓存先行短路降噪;命中拒绝规则、Jev 判拒、低置信、超时、网络故障一律拒绝该调用(工具卡片显示 AGENT_REVIEW_DENIED 详情与风险理由)。",
|
|
974
|
+
h("br", null),
|
|
975
|
+
"已存在的会话不受此开关影响,可用 /permission 菜单「自动审查」或 /agent-review on|off 单独切换。审计逐调用记录在「审批」标签页(工具列标注「逐调用」)。",
|
|
976
|
+
h("br", null),
|
|
977
|
+
state === null
|
|
978
|
+
? "状态加载中…"
|
|
979
|
+
: state.reviewAvailable === true
|
|
980
|
+
? "当前状态:Jev 判定可用,自动审查已就绪。"
|
|
981
|
+
: "当前状态:Jev 判定不可用(缺 API Key),开关已禁用,新会话不会自动审查。",
|
|
982
|
+
),
|
|
983
|
+
),
|
|
980
984
|
h(
|
|
981
985
|
"div",
|
|
982
986
|
{ className: "aapr-card" },
|
package/index.js
CHANGED
|
@@ -31,11 +31,14 @@
|
|
|
31
31
|
* `callId`) plus the asker's stated reason. A rejection must name the
|
|
32
32
|
* concrete, credible risk the operation creates (destructive /
|
|
33
33
|
* irreversible / out-of-scope / dishonest); vague unease is approved.
|
|
34
|
-
* v1.
|
|
35
|
-
*
|
|
36
|
-
*
|
|
37
|
-
*
|
|
38
|
-
*
|
|
34
|
+
* v1.10.0: the judge is ALWAYS an LLM — either one direct stream call or
|
|
35
|
+
* the spawn child. The TypeSafe Jev "System One" decision model (v1.6.0
|
|
36
|
+
* through v1.9.x could judge escalations too, as the synthetic provider
|
|
37
|
+
* id `typesafe`) is NO LONGER selectable here: its calibrated-but-shallow
|
|
38
|
+
* risk judgement is a weaker safety net than an LLM judge on the
|
|
39
|
+
* escalation path, which is exactly the path a human approval would have
|
|
40
|
+
* guarded. Jev now backs the separate per-call review mode ONLY
|
|
41
|
+
* (`_reviewWithJev`), configured independently from this judge.
|
|
39
42
|
*
|
|
40
43
|
* 3. FAIL CLOSED — any infrastructure fault, timeout, malformed verdict, or
|
|
41
44
|
* cancellation maps to the fail-closed approval outcomes
|
|
@@ -110,29 +113,35 @@ const DATA_DIR = join(process.env.DSH_HOME || join(homedir(), ".dsh"), "agent-ap
|
|
|
110
113
|
const CONFIG_FILE = join(DATA_DIR, "config.json");
|
|
111
114
|
|
|
112
115
|
/**
|
|
113
|
-
* v1.6.0
|
|
114
|
-
* model (https://api.typesafe.ai/v1/systemone)
|
|
115
|
-
* it answers typed questions (Choice / Score / Noul) over
|
|
116
|
-
* calibrated probability distributions in ~70–500ms
|
|
117
|
-
*
|
|
118
|
-
*
|
|
119
|
-
*
|
|
120
|
-
*
|
|
121
|
-
*
|
|
122
|
-
* confidence
|
|
123
|
-
* grant, and (below the gate) not a recorded rejection either.
|
|
116
|
+
* v1.6.0 (removed from the escalation judge in v1.10.0): the TypeSafe Jev
|
|
117
|
+
* "System One" decision model (https://api.typesafe.ai/v1/systemone). It does
|
|
118
|
+
* not generate text — it answers typed questions (Choice / Score / Noul) over
|
|
119
|
+
* one `state` with calibrated probability distributions in ~70–500ms, and it
|
|
120
|
+
* cannot appear in `llm.listProviders()` (not a chat route), hence the direct
|
|
121
|
+
* HTTP call. It now judges the PER-CALL REVIEW mode only (`_reviewWithJev`),
|
|
122
|
+
* configured independently of the approval judge. Fail-closed is preserved end
|
|
123
|
+
* to end there as well: any transport fault, non-200, malformed answer, or a
|
|
124
|
+
* confidence below the configured gate denies the call (or, on the low
|
|
125
|
+
* confidence branch, records `unavailable` without a verdict).
|
|
124
126
|
*/
|
|
125
|
-
const JEV_PROVIDER = "typesafe";
|
|
126
127
|
const JEV_DEFAULT_MODEL = "jev-latest";
|
|
127
128
|
const JEV_DEFAULT_ENDPOINT = "https://api.typesafe.ai/v1/systemone";
|
|
128
129
|
const JEV_DEFAULT_CONFIDENCE = 0.5;
|
|
130
|
+
/**
|
|
131
|
+
* The retired escalation-judge provider id (v1.6.0–v1.9.x). Kept ONLY to
|
|
132
|
+
* recognize and drop it from persisted configs: a config.json still carrying
|
|
133
|
+
* `model.provider: "typesafe"` migrates to the harness default route, and
|
|
134
|
+
* `setModel` treats it as "clear the override". Never used to route a judge.
|
|
135
|
+
*/
|
|
136
|
+
const JEV_LEGACY_PROVIDER = "typesafe";
|
|
129
137
|
|
|
130
138
|
/**
|
|
131
|
-
* The typed questions sent to Jev
|
|
132
|
-
* (
|
|
133
|
-
*
|
|
134
|
-
*
|
|
135
|
-
*
|
|
139
|
+
* The typed questions sent to Jev — the shared policy base for
|
|
140
|
+
* `JEV_REVIEW_QUESTIONS` (the only sender since v1.10.0). Policy lives in the
|
|
141
|
+
* criteria descriptions (Jev reads instructions literally and injects domain
|
|
142
|
+
* knowledge only through state + criteria); the wording mirrors the LLM
|
|
143
|
+
* judge's APPROVE conditions / REJECT list / 误杀治理 so both judge to the
|
|
144
|
+
* same standard. `decision` and `riskLevel` are Choices (discrete options +
|
|
136
145
|
* probabilities + confidence); `concreteRisk` is a Noul probe whose
|
|
137
146
|
* probability is folded into the synthesized audit rationale.
|
|
138
147
|
*/
|
|
@@ -164,9 +173,10 @@ const JEV_QUESTIONS = {
|
|
|
164
173
|
};
|
|
165
174
|
|
|
166
175
|
/**
|
|
167
|
-
* v1.8.0: the per-call review mode (`agent-review` preset)
|
|
168
|
-
* criteria as `JEV_QUESTIONS` — only the
|
|
169
|
-
* from "escalation request" to the
|
|
176
|
+
* v1.8.0: the per-call review mode (`agent-review` preset) — since v1.10.0
|
|
177
|
+
* the ONLY Jev caller. The SAME policy criteria as `JEV_QUESTIONS` — only the
|
|
178
|
+
* decision instruction wording adapts from "escalation request" to the
|
|
179
|
+
* pending tool call. v1.9.0 adds guidance
|
|
170
180
|
* for the `effectiveCode` state field (the actual interpreter scripts the
|
|
171
181
|
* call would run): judge the code when present; reduced visibility alone is
|
|
172
182
|
* never a rejection reason (误杀治理 holds); and a source-edit diff that
|
|
@@ -609,10 +619,13 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
609
619
|
*/
|
|
610
620
|
this._reviewDefault = false;
|
|
611
621
|
/**
|
|
612
|
-
* TypeSafe Jev direct backend settings
|
|
613
|
-
* the
|
|
614
|
-
*
|
|
615
|
-
*
|
|
622
|
+
* TypeSafe Jev direct backend settings — v1.10.0: the INDEPENDENT
|
|
623
|
+
* configuration of the 自动审查 (per-call review) mode, no longer part of
|
|
624
|
+
* the approval judge's provider selection. Configuring it (a resolvable
|
|
625
|
+
* key) is what opens the review gate; see `_jevGateOk`. The API key lives
|
|
626
|
+
* in plaintext on this machine only (config.json, same trust domain as
|
|
627
|
+
* the rest of the settings); an empty key falls back to the
|
|
628
|
+
* TYPESAFE_API_KEY env var.
|
|
616
629
|
*/
|
|
617
630
|
this._jev = {
|
|
618
631
|
apiKey: "",
|
|
@@ -876,7 +889,7 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
876
889
|
return "自动审批 is already ON for this session — switch modes through the /permission menu";
|
|
877
890
|
}
|
|
878
891
|
if (!this._jevGateOk()) {
|
|
879
|
-
return "自动审查 requires the Jev judge:
|
|
892
|
+
return "自动审查 requires the Jev judge: configure a Jev API key in Settings → 自动审批 → 「自动审查」first";
|
|
880
893
|
}
|
|
881
894
|
this._enableCore(session, agent, "review");
|
|
882
895
|
if (this._presetRegistered(REVIEW_PRESET_NAME)) {
|
|
@@ -907,13 +920,14 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
907
920
|
}
|
|
908
921
|
|
|
909
922
|
/**
|
|
910
|
-
* v1.8.0 gate for the per-call review mode
|
|
911
|
-
*
|
|
912
|
-
*
|
|
913
|
-
*
|
|
923
|
+
* v1.8.0, re-based in v1.10.0: the gate for the per-call review mode. The
|
|
924
|
+
* Jev backend is the review mode's own, independent configuration — the
|
|
925
|
+
* gate is simply "a Jev API key resolves" (config.json or
|
|
926
|
+
* TYPESAFE_API_KEY). It no longer requires switching the APPROVAL judge to
|
|
927
|
+
* Jev (that option is gone as of v1.10.0), and it never reads `_model`.
|
|
914
928
|
*/
|
|
915
929
|
_jevGateOk() {
|
|
916
|
-
return this.
|
|
930
|
+
return this._jevEffective().key !== "";
|
|
917
931
|
}
|
|
918
932
|
|
|
919
933
|
/**
|
|
@@ -938,7 +952,7 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
938
952
|
durationMs: 0,
|
|
939
953
|
childSessionId: "",
|
|
940
954
|
rationale:
|
|
941
|
-
"自动审查 requires the Jev judge (Settings →
|
|
955
|
+
"自动审查 requires the Jev judge (Settings → 自动审批 → 「自动审查」卡片: configure a Jev API key); falling back to the 自动审批 preset (fail closed)",
|
|
942
956
|
mode: "review",
|
|
943
957
|
});
|
|
944
958
|
const presets = this.ctx.get("permissionPresets");
|
|
@@ -1288,7 +1302,13 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
1288
1302
|
typeof cfg.model.provider === "string" &&
|
|
1289
1303
|
typeof cfg.model.model === "string"
|
|
1290
1304
|
) {
|
|
1291
|
-
|
|
1305
|
+
// v1.10.0 migration: a persisted `typesafe` judge route (v1.6.0–
|
|
1306
|
+
// v1.9.x) is no longer a valid judge — fall back to the harness
|
|
1307
|
+
// default route instead of keeping a dead provider on the wire.
|
|
1308
|
+
this._model =
|
|
1309
|
+
cfg.model.provider === JEV_LEGACY_PROVIDER
|
|
1310
|
+
? { provider: "", model: "" }
|
|
1311
|
+
: { provider: cfg.model.provider, model: cfg.model.model };
|
|
1292
1312
|
}
|
|
1293
1313
|
if (cfg.judgeMode === JUDGE_MODE_LLM || cfg.judgeMode === JUDGE_MODE_SUBAGENT) {
|
|
1294
1314
|
this._judgeMode = cfg.judgeMode;
|
|
@@ -1489,10 +1509,6 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
1489
1509
|
* records display — "p/m" = selected, "default(p/m)" = harness default.
|
|
1490
1510
|
*/
|
|
1491
1511
|
_judgeRoute() {
|
|
1492
|
-
if (this._model.provider === JEV_PROVIDER) {
|
|
1493
|
-
const model = this._jevEffective().model;
|
|
1494
|
-
return { provider: JEV_PROVIDER, model: model, label: "jev(" + model + ")" };
|
|
1495
|
-
}
|
|
1496
1512
|
if (this._model.provider !== "" && this._model.model !== "") {
|
|
1497
1513
|
return {
|
|
1498
1514
|
provider: this._model.provider,
|
|
@@ -1558,13 +1574,11 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
1558
1574
|
return "allowed-once";
|
|
1559
1575
|
}
|
|
1560
1576
|
|
|
1561
|
-
// 3. The judge
|
|
1562
|
-
//
|
|
1563
|
-
// direct ctx.llm.stream() call (no subagent either —
|
|
1564
|
-
// judgeMode === "subagent" spawns the judge child
|
|
1565
|
-
|
|
1566
|
-
return this._judgeWithJev(session, req, argsRaw, base, trustKey);
|
|
1567
|
-
}
|
|
1577
|
+
// 3. The judge — always an LLM (v1.10.0: Jev is no longer selectable
|
|
1578
|
+
// here; it judges the per-call review mode only). The DEFAULT "llm"
|
|
1579
|
+
// mode is one direct ctx.llm.stream() call (no subagent either —
|
|
1580
|
+
// v1.8.0); only judgeMode === "subagent" spawns the judge child
|
|
1581
|
+
// through `spawn`.
|
|
1568
1582
|
if (this._judgeMode !== JUDGE_MODE_SUBAGENT) {
|
|
1569
1583
|
return this._judgeWithLlmStream(session, req, argsRaw, base, trustKey);
|
|
1570
1584
|
}
|
|
@@ -1703,7 +1717,7 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
1703
1717
|
* exactly — same persona, same `_judgePrompt`, same VERDICT_SCHEMA verdict
|
|
1704
1718
|
* contract — only the invocation differs: no subagent session is created
|
|
1705
1719
|
* (zero judge-side context pollution; `childSessionId` stays empty).
|
|
1706
|
-
*
|
|
1720
|
+
* Fail-closed contract:
|
|
1707
1721
|
* - no concrete route / llm fault / non-'stop' finish / malformed verdict
|
|
1708
1722
|
* / timeout → `unavailable`
|
|
1709
1723
|
* - request cancelled mid-flight → `cancelled`
|
|
@@ -1969,13 +1983,12 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
1969
1983
|
return { decision: value.decision, riskLevel: value.riskLevel, rationale: value.rationale };
|
|
1970
1984
|
}
|
|
1971
1985
|
|
|
1972
|
-
// ---- the TypeSafe Jev direct backend
|
|
1986
|
+
// ---- the TypeSafe Jev direct backend (自动审查 only, since v1.10.0) ---------
|
|
1973
1987
|
|
|
1974
1988
|
/**
|
|
1975
1989
|
* Effective Jev settings with env fallback and clamping applied. The key
|
|
1976
1990
|
* may come from config.json or the TYPESAFE_API_KEY environment variable;
|
|
1977
|
-
* an absent key keeps the
|
|
1978
|
-
* `unavailable` (fail closed) until one is configured.
|
|
1991
|
+
* an absent key keeps the review gate closed, so 自动审查 cannot be armed.
|
|
1979
1992
|
*/
|
|
1980
1993
|
_jevEffective() {
|
|
1981
1994
|
const key = String(this._jev.apiKey || process.env.TYPESAFE_API_KEY || "").trim();
|
|
@@ -2023,110 +2036,12 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
2023
2036
|
}
|
|
2024
2037
|
|
|
2025
2038
|
/**
|
|
2026
|
-
*
|
|
2027
|
-
*
|
|
2028
|
-
*
|
|
2029
|
-
*
|
|
2030
|
-
*
|
|
2031
|
-
* - request cancelled mid-flight → `cancelled`
|
|
2032
|
-
* - overall timeout (the same `this._timeoutMs` budget) → `unavailable`
|
|
2033
|
-
* - confidence below the configured gate → `unavailable` (the model is
|
|
2034
|
-
* not sure enough to decide: never a grant, and not a recorded
|
|
2035
|
-
* rejection either — the v1.4.0 误杀治理 applies symmetrically)
|
|
2036
|
-
* Jev does not generate text, so the audit rationale is synthesized from
|
|
2037
|
-
* the returned distributions; the served model version (`body.model`,
|
|
2038
|
-
* which resolves aliases like jev-latest) is what the audit displays.
|
|
2039
|
+
* The single POST to the System One endpoint; resolves the parsed body.
|
|
2040
|
+
* `questions` defaults to the per-call review set (the only caller since
|
|
2041
|
+
* v1.10.0). `state` leaves this machine by design — that is the review
|
|
2042
|
+
* judge's whole point (see the workspace-confined code expansion in
|
|
2043
|
+
* `_reviewEffectiveCode`).
|
|
2039
2044
|
*/
|
|
2040
|
-
async _judgeWithJev(session, req, argsRaw, base, trustKey) {
|
|
2041
|
-
const cfg = this._jevEffective();
|
|
2042
|
-
if (cfg.key === "") {
|
|
2043
|
-
this._record(session, {
|
|
2044
|
-
...base,
|
|
2045
|
-
outcome: "unavailable",
|
|
2046
|
-
riskLevel: "-",
|
|
2047
|
-
model: "jev(" + cfg.model + ")",
|
|
2048
|
-
rationale: "Jev backend selected but no API key configured (Settings → 自动审批, or the TYPESAFE_API_KEY environment variable)",
|
|
2049
|
-
});
|
|
2050
|
-
return "unavailable";
|
|
2051
|
-
}
|
|
2052
|
-
if (typeof fetch !== "function") {
|
|
2053
|
-
this._record(session, { ...base, outcome: "unavailable", riskLevel: "-", model: "jev(" + cfg.model + ")", rationale: "fetch is unavailable in this runtime" });
|
|
2054
|
-
return "unavailable";
|
|
2055
|
-
}
|
|
2056
|
-
|
|
2057
|
-
const startedAt = Date.now();
|
|
2058
|
-
const controller = new AbortController();
|
|
2059
|
-
const signal = req.signal;
|
|
2060
|
-
const onAbort = () => controller.abort();
|
|
2061
|
-
if (signal && typeof signal.addEventListener === "function") {
|
|
2062
|
-
signal.addEventListener("abort", onAbort, { once: true });
|
|
2063
|
-
}
|
|
2064
|
-
|
|
2065
|
-
let winner;
|
|
2066
|
-
try {
|
|
2067
|
-
const state = this._jevStateOf(session, req, argsRaw);
|
|
2068
|
-
winner = await Promise.race([
|
|
2069
|
-
this._jevRequest(cfg, state, controller.signal)
|
|
2070
|
-
.then((body) => ({ kind: "result", body: body }))
|
|
2071
|
-
.catch((error) => ({
|
|
2072
|
-
kind: "fault",
|
|
2073
|
-
error: error,
|
|
2074
|
-
aborted: error && error.name === "AbortError",
|
|
2075
|
-
})),
|
|
2076
|
-
(signal
|
|
2077
|
-
? new Promise((resolve) => {
|
|
2078
|
-
if (signal.aborted) {
|
|
2079
|
-
resolve(true);
|
|
2080
|
-
return;
|
|
2081
|
-
}
|
|
2082
|
-
signal.addEventListener("abort", () => resolve(true), { once: true });
|
|
2083
|
-
})
|
|
2084
|
-
: Promise.resolve(false)
|
|
2085
|
-
).then((v) => ({ kind: "aborted", aborted: v })),
|
|
2086
|
-
this.ctx.timeout(this._timeoutMs).then(() => ({ kind: "timeout" })),
|
|
2087
|
-
]);
|
|
2088
|
-
} finally {
|
|
2089
|
-
if (signal && typeof signal.removeEventListener === "function") {
|
|
2090
|
-
signal.removeEventListener("abort", onAbort);
|
|
2091
|
-
}
|
|
2092
|
-
// Whether we lost the race to timeout/cancel or the request already
|
|
2093
|
-
// settled, closing the transport is always safe.
|
|
2094
|
-
try {
|
|
2095
|
-
controller.abort();
|
|
2096
|
-
} catch (e) {
|
|
2097
|
-
/* controller abort never blocks the outcome */
|
|
2098
|
-
}
|
|
2099
|
-
}
|
|
2100
|
-
const durationMs = Date.now() - startedAt;
|
|
2101
|
-
|
|
2102
|
-
if (winner.kind === "result") {
|
|
2103
|
-
return this._jevVerdict(session, winner.body, cfg, base, trustKey, durationMs);
|
|
2104
|
-
}
|
|
2105
|
-
if (winner.kind === "aborted") {
|
|
2106
|
-
this._record(session, { ...base, outcome: "cancelled", riskLevel: "-", model: "jev(" + cfg.model + ")", rationale: "request cancelled while Jev was judging" });
|
|
2107
|
-
return "cancelled";
|
|
2108
|
-
}
|
|
2109
|
-
if (winner.kind === "timeout") {
|
|
2110
|
-
this._record(session, {
|
|
2111
|
-
...base,
|
|
2112
|
-
outcome: "unavailable",
|
|
2113
|
-
riskLevel: "-",
|
|
2114
|
-
model: "jev(" + cfg.model + ")",
|
|
2115
|
-
rationale: "Jev request timed out after " + String(this._timeoutMs) + "ms (fail closed)",
|
|
2116
|
-
});
|
|
2117
|
-
return "unavailable";
|
|
2118
|
-
}
|
|
2119
|
-
if (winner.aborted) {
|
|
2120
|
-
this._record(session, { ...base, outcome: "cancelled", riskLevel: "-", model: "jev(" + cfg.model + ")", rationale: "request cancelled while Jev was judging" });
|
|
2121
|
-
return "cancelled";
|
|
2122
|
-
}
|
|
2123
|
-
this._record(session, { ...base, outcome: "unavailable", riskLevel: "-", model: "jev(" + cfg.model + ")", rationale: "Jev request failed: " + errText(winner.error) });
|
|
2124
|
-
return "unavailable";
|
|
2125
|
-
}
|
|
2126
|
-
|
|
2127
|
-
/** The single POST to the System One endpoint; resolves the parsed body.
|
|
2128
|
-
* `questions` defaults to the escalation set; the review path passes
|
|
2129
|
-
* `JEV_REVIEW_QUESTIONS`. */
|
|
2130
2045
|
async _jevRequest(cfg, state, abortSignal, questions) {
|
|
2131
2046
|
const response = await fetch(cfg.endpoint, {
|
|
2132
2047
|
method: "POST",
|
|
@@ -2137,7 +2052,7 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
2137
2052
|
body: JSON.stringify({
|
|
2138
2053
|
state: state,
|
|
2139
2054
|
model: cfg.model,
|
|
2140
|
-
questions: questions === undefined ?
|
|
2055
|
+
questions: questions === undefined ? JEV_REVIEW_QUESTIONS : questions,
|
|
2141
2056
|
}),
|
|
2142
2057
|
signal: abortSignal,
|
|
2143
2058
|
});
|
|
@@ -2156,9 +2071,9 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
2156
2071
|
}
|
|
2157
2072
|
|
|
2158
2073
|
/**
|
|
2159
|
-
* Parse + gate one Jev response into a normalized verdict, shared by
|
|
2160
|
-
* escalation path
|
|
2161
|
-
*
|
|
2074
|
+
* Parse + gate one Jev response into a normalized verdict, shared by every
|
|
2075
|
+
* review call (the escalation path stopped using Jev in v1.10.0) so the
|
|
2076
|
+
* review judge always holds to the same standard:
|
|
2162
2077
|
* - `{ kind: "malformed", served }` — any missing/out-of-shape answer
|
|
2163
2078
|
* - `{ kind: "low-confidence", served, choice, riskLevel, confidence, gate }`
|
|
2164
2079
|
* - `{ kind: "verdict", served, choice, riskLevel, rationale }`
|
|
@@ -2231,56 +2146,6 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
2231
2146
|
return { kind: "verdict", served: served, choice: choice, riskLevel: riskChoice, rationale: rationale };
|
|
2232
2147
|
}
|
|
2233
2148
|
|
|
2234
|
-
/**
|
|
2235
|
-
* Map a Jev response to the same outcomes the subagent path produces.
|
|
2236
|
-
* Returns the waterfall outcome string; records the audit line itself.
|
|
2237
|
-
*/
|
|
2238
|
-
_jevVerdict(session, body, cfg, base, trustKey, durationMs) {
|
|
2239
|
-
const parsed = this._jevParse(body, cfg);
|
|
2240
|
-
const label = "jev(" + parsed.served + ")";
|
|
2241
|
-
base.durationMs = durationMs;
|
|
2242
|
-
|
|
2243
|
-
if (parsed.kind === "malformed") {
|
|
2244
|
-
this._record(session, {
|
|
2245
|
-
...base,
|
|
2246
|
-
outcome: "unavailable",
|
|
2247
|
-
riskLevel: "-",
|
|
2248
|
-
model: label,
|
|
2249
|
-
rationale: "Jev returned no valid verdict shape (decision/riskLevel/concreteRisk incomplete)",
|
|
2250
|
-
});
|
|
2251
|
-
return "unavailable";
|
|
2252
|
-
}
|
|
2253
|
-
if (parsed.kind === "low-confidence") {
|
|
2254
|
-
this._record(session, {
|
|
2255
|
-
...base,
|
|
2256
|
-
outcome: "unavailable",
|
|
2257
|
-
riskLevel: parsed.riskLevel,
|
|
2258
|
-
model: label,
|
|
2259
|
-
rationale:
|
|
2260
|
-
"Jev confidence " + parsed.confidence.toFixed(2) + " is below the gate " + parsed.gate.toFixed(2) + " (decision draft: " + parsed.choice + ") — fail closed",
|
|
2261
|
-
});
|
|
2262
|
-
return "unavailable";
|
|
2263
|
-
}
|
|
2264
|
-
|
|
2265
|
-
const approved = parsed.choice === "approve";
|
|
2266
|
-
this._record(session, {
|
|
2267
|
-
...base,
|
|
2268
|
-
outcome: approved ? "allowed-once" : "rejected",
|
|
2269
|
-
riskLevel: parsed.riskLevel,
|
|
2270
|
-
model: label,
|
|
2271
|
-
rationale: trunc(parsed.rationale, 600),
|
|
2272
|
-
});
|
|
2273
|
-
if (approved && trustKey !== undefined) {
|
|
2274
|
-
let set = this._trusted.get(session.id);
|
|
2275
|
-
if (set === undefined) {
|
|
2276
|
-
set = new Set();
|
|
2277
|
-
this._trusted.set(session.id, set);
|
|
2278
|
-
}
|
|
2279
|
-
set.add(trustKey);
|
|
2280
|
-
}
|
|
2281
|
-
return approved ? "allowed-once" : "rejected";
|
|
2282
|
-
}
|
|
2283
|
-
|
|
2284
2149
|
// ---- v1.8.0 per-call review mode (agent-review) ----------------------------
|
|
2285
2150
|
|
|
2286
2151
|
/**
|
|
@@ -2573,12 +2438,12 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
2573
2438
|
}
|
|
2574
2439
|
|
|
2575
2440
|
/**
|
|
2576
|
-
* Judge one pending call through the Jev HTTP API
|
|
2577
|
-
*
|
|
2578
|
-
*
|
|
2579
|
-
* low confidence, timeouts and faults all
|
|
2580
|
-
* fallback). Returns `undefined` to allow,
|
|
2581
|
-
* decision.
|
|
2441
|
+
* Judge one pending call through the Jev HTTP API — the review mode's only
|
|
2442
|
+
* judge, and the only remaining Jev caller since v1.10.0. Fail-closed
|
|
2443
|
+
* through `_jevParse` (same validation, same confidence gate) but every
|
|
2444
|
+
* outcome is FINAL: rejections, low confidence, timeouts and faults all
|
|
2445
|
+
* deny the call (no human fallback). Returns `undefined` to allow,
|
|
2446
|
+
* otherwise a pre-execute decision.
|
|
2582
2447
|
*/
|
|
2583
2448
|
async _reviewWithJev(session, exec, argsRaw, base, trustKey) {
|
|
2584
2449
|
const cfg = this._jevEffective();
|
|
@@ -2776,15 +2641,18 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
2776
2641
|
|
|
2777
2642
|
/**
|
|
2778
2643
|
* Set the judge model override. Empty strings clear it (the judge then runs
|
|
2779
|
-
* on the harness default route, never the requester's).
|
|
2644
|
+
* on the harness default route, never the requester's). v1.10.0: Jev is no
|
|
2645
|
+
* longer a judge provider — a legacy `typesafe` selection migrates to the
|
|
2646
|
+
* harness default route. Persisted.
|
|
2780
2647
|
*/
|
|
2781
2648
|
async setModel(request) {
|
|
2782
2649
|
const provider = request && typeof request.provider === "string" ? request.provider : "";
|
|
2783
2650
|
const model = request && typeof request.model === "string" ? request.model : "";
|
|
2784
|
-
if (provider ===
|
|
2785
|
-
//
|
|
2786
|
-
//
|
|
2787
|
-
|
|
2651
|
+
if (provider === JEV_LEGACY_PROVIDER) {
|
|
2652
|
+
// v1.6.0–v1.9.x routed the escalation judge to Jev. That option is gone
|
|
2653
|
+
// (Jev judges the per-call review mode only), so an old client still
|
|
2654
|
+
// sending it clears the override instead of resurrecting a dead route.
|
|
2655
|
+
this._model = { provider: "", model: "" };
|
|
2788
2656
|
} else {
|
|
2789
2657
|
this._model =
|
|
2790
2658
|
provider !== "" && model !== "" ? { provider, model } : { provider: "", model: "" };
|
|
@@ -2828,20 +2696,35 @@ export class AgentApprovalService extends TypertRemoteService {
|
|
|
2828
2696
|
}
|
|
2829
2697
|
|
|
2830
2698
|
/**
|
|
2831
|
-
* Set the TypeSafe Jev backend settings (only provided fields change)
|
|
2832
|
-
*
|
|
2833
|
-
*
|
|
2699
|
+
* Set the TypeSafe Jev backend settings (only provided fields change) — the
|
|
2700
|
+
* 自动审查 mode's own, independent configuration (v1.10.0: Jev is not a
|
|
2701
|
+
* judge provider, so nothing here touches the approval route). Saving a
|
|
2702
|
+
* resolvable key OPENS the review gate, and opening the gate for the first
|
|
2703
|
+
* time turns 自动审查 on for new sessions (`_reviewDefault`) — "配置了 Jev
|
|
2704
|
+
* 就启用自动审查". The same card switches it back off. `confidence` is the
|
|
2705
|
+
* gate below which Jev's answer is not trusted and the call is denied;
|
|
2706
|
+
* clamped to [0.01, 0.99]. Persisted.
|
|
2834
2707
|
*/
|
|
2835
2708
|
async setJevConfig(request) {
|
|
2836
2709
|
const r = request && typeof request === "object" ? request : {};
|
|
2710
|
+
const gateBefore = this._jevGateOk();
|
|
2837
2711
|
if (typeof r.apiKey === "string") this._jev.apiKey = r.apiKey.trim();
|
|
2838
2712
|
if (typeof r.endpoint === "string") this._jev.endpoint = r.endpoint.trim();
|
|
2839
2713
|
if (typeof r.model === "string") this._jev.model = r.model.trim();
|
|
2840
2714
|
if (typeof r.confidence === "number" && Number.isFinite(r.confidence)) {
|
|
2841
2715
|
this._jev.confidence = Math.min(0.99, Math.max(0.01, r.confidence));
|
|
2842
2716
|
}
|
|
2717
|
+
const gateAfter = this._jevGateOk();
|
|
2718
|
+
if (gateAfter && !gateBefore) this._reviewDefault = true;
|
|
2843
2719
|
this._persistConfig();
|
|
2844
|
-
return {
|
|
2720
|
+
return {
|
|
2721
|
+
ok: true,
|
|
2722
|
+
value: {
|
|
2723
|
+
jev: this._jevShape(),
|
|
2724
|
+
reviewAvailable: gateAfter,
|
|
2725
|
+
reviewDefault: this._reviewDefault,
|
|
2726
|
+
},
|
|
2727
|
+
};
|
|
2845
2728
|
}
|
|
2846
2729
|
|
|
2847
2730
|
/** Set the judge timeout (clamped to [MIN, MAX] milliseconds). Persisted. */
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@duke-dsh-plugins/dsh-agent-approval",
|
|
3
|
-
"version": "1.
|
|
4
|
-
"description": "Agent-decided approvals for DeepSeek Harness: an 自动审批 permission mode where every sandbox escalation is judged automatically (default judge: one direct LLM call with no subagent session; optional isolated judge subagent
|
|
3
|
+
"version": "1.10.0",
|
|
4
|
+
"description": "Agent-decided approvals for DeepSeek Harness: an 自动审批 permission mode where every sandbox escalation is judged automatically by an LLM (default judge: one direct LLM call with no subagent session; optional isolated judge subagent) plus an independent 自动审查 per-call review mode (Full access base; every tool call reviewed before execution by the TypeSafe Jev decision model, with interpreter scripts the command references expanded into the judge state so temp scripts are judged by their actual code; risky calls rejected with no human fallback), with a configurable judge model and a per-session audit trail in the conversation window's 审批 tab.",
|
|
5
5
|
"repository": {
|
|
6
6
|
"type": "git",
|
|
7
7
|
"url": "git+https://github.com/MoonlitDropOfBlood/dsh-agent-approval.git"
|
package/typert.host.js
CHANGED
|
@@ -115,9 +115,17 @@ const setReviewDefaultValueSchema = z
|
|
|
115
115
|
})
|
|
116
116
|
.readonly();
|
|
117
117
|
|
|
118
|
+
/**
|
|
119
|
+
* v1.10.0: the TypeSafe Jev backend belongs to the 自动审查 (per-call review)
|
|
120
|
+
* mode alone — it is NOT a judge provider. `setJevConfig` therefore reports
|
|
121
|
+
* the review gate and the (possibly auto-enabled) review default so the
|
|
122
|
+
* client can reflect "配置了 Jev 就启用自动审查" without a second round trip.
|
|
123
|
+
*/
|
|
118
124
|
const setJevValueSchema = z
|
|
119
125
|
.object({
|
|
120
126
|
jev: jevConfigSchema,
|
|
127
|
+
reviewAvailable: z.boolean(),
|
|
128
|
+
reviewDefault: z.boolean(),
|
|
121
129
|
})
|
|
122
130
|
.readonly();
|
|
123
131
|
|
|
@@ -557,17 +565,17 @@ export const TYPERT = {
|
|
|
557
565
|
kind: "method",
|
|
558
566
|
name: "setModel",
|
|
559
567
|
signature: "@Remote('setModel') async setModel(request: AgentApprovalSetModelRequest): Promise<AgentApprovalSetModelResult>",
|
|
560
|
-
summary: "Set the judge model route (empty strings =
|
|
568
|
+
summary: "Set the judge model route (empty strings = use the harness default model; the escalation judge is always an LLM).",
|
|
561
569
|
jsDoc:
|
|
562
|
-
"/**\n * Set provider/model used by the approval judge; empty strings clear the override
|
|
570
|
+
"/**\n * Set provider/model used by the approval judge; empty strings clear the override (the judge then uses the harness default model). Since v1.10.0 Jev is not a judge provider — a legacy 'typesafe' selection clears the override.\n * @param request - { provider, model }.\n * @returns the stored route.\n */",
|
|
563
571
|
},
|
|
564
572
|
{
|
|
565
573
|
kind: "method",
|
|
566
574
|
name: "setJevConfig",
|
|
567
575
|
signature: "@Remote('setJevConfig') async setJevConfig(request: AgentApprovalSetJevRequest): Promise<AgentApprovalSetJevResult>",
|
|
568
|
-
summary: "Set the TypeSafe Jev backend settings (API key, endpoint, model, confidence gate; persisted).",
|
|
576
|
+
summary: "Set the TypeSafe Jev backend settings of the 自动审查 mode (API key, endpoint, model, confidence gate; persisted; a resolvable key opens the review gate and enables 自动审查 for new sessions).",
|
|
569
577
|
jsDoc:
|
|
570
|
-
"/**\n * Update the Jev backend settings
|
|
578
|
+
"/**\n * Update the Jev backend settings of the per-call review mode (independent of the approval judge). Only provided fields change; confidence (the fail-closed gate) is clamped to [0.01, 0.99]. Saving a resolvable API key opens the review gate and turns 自动审查 on for new sessions.\n * @param request - { apiKey, endpoint, model, confidence }.\n * @returns the stored Jev settings plus the review gate and the review default.\n */",
|
|
571
579
|
},
|
|
572
580
|
{
|
|
573
581
|
kind: "method",
|
|
@@ -642,7 +650,7 @@ export const TYPERT = {
|
|
|
642
650
|
{
|
|
643
651
|
name: "AgentApprovalSetJevResult",
|
|
644
652
|
declaration:
|
|
645
|
-
"export type AgentApprovalSetJevResult = { ok: true; value: { readonly jev: AgentApprovalJevConfig } } | { ok: false; error: { code: string; message?: string } };",
|
|
653
|
+
"export type AgentApprovalSetJevResult = { ok: true; value: { readonly jev: AgentApprovalJevConfig; readonly reviewAvailable: boolean; readonly reviewDefault: boolean } } | { ok: false; error: { code: string; message?: string } };",
|
|
646
654
|
},
|
|
647
655
|
{
|
|
648
656
|
name: "AgentApprovalEnabledSession",
|