pi-verdict 0.10.0 → 0.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -8,7 +8,7 @@
8
8
 
9
9
  **pi-verdict is a minimal permission gate for [pi](https://pi.dev) in the style of Claude Code's auto mode: every tool call gets checked before it runs — allow, deny, or ask you first.**
10
10
 
11
- - Minimal — just 1k+ lines of code
11
+ - Minimal — just 2k lines of code
12
12
  - Built-in danger rules and your own allow/deny rules settle the clear cases first, at zero latency
13
13
  - Everything else goes to a model classifier that sees the conversation context
14
14
  - Any uncertainty or failure fails closed; nothing ever runs silently
@@ -69,7 +69,7 @@ pi-verdict runs on both [pi](https://github.com/badlogic/pi-mono) and [oh-my-pi]
69
69
  | user rules | `~/.pi/agent/config/pi-verdict.json` | `~/.omp/agent/config/pi-verdict.json` |
70
70
  | credential file (S0 hard deny) | `~/.pi/agent/auth.json` | `~/.omp/agent/auth.json` |
71
71
 
72
- - `/automode` — show current status: on/off + shadow-cache stats for the session
72
+ - `/automode` — show current status: on/off
73
73
  - `/automode on`
74
74
  - `/automode off`
75
75
  - `ctrl+shift+a` — toggle the master switch silently (the always-on footer is the only feedback; rebind or disable via `toggleShortcut`)
@@ -97,25 +97,32 @@ pi-verdict runs on both [pi](https://github.com/badlogic/pi-mono) and [oh-my-pi]
97
97
  "~/.zshrc",
98
98
  "~/.bashrc"
99
99
  ],
100
+ "ignoreTools": [
101
+ "todo",
102
+ "ask_user_question",
103
+ "memory_write",
104
+ "memory_search"
105
+ ],
100
106
  "builtinDenyFloor": true,
101
107
  "classifierModel": null,
102
108
  "toggleShortcut": "ctrl+shift+a",
103
109
  "audit": false,
104
110
  "notifyAllows": false,
111
+ "classifierMinConfidence": null,
105
112
  "classifierFallbackModel": null,
106
- "classifierFallbackConfidence": 50,
107
- "classifierFallbackMode": "shadow"
113
+ "classifierFallbackMode": "enforce"
108
114
  }
109
115
  ```
110
116
 
111
117
  - `allow`/`deny` are JS regex arrays; **`deny` wins over `allow`**, both beat the classifier
112
- - `denyPaths` are plain paths you declare **protected** — touches trigger a terminal ask you adjudicate (non-interactive → deny); the classifier never learns the paths themselves, only that they exist. `grep`/`find`/`ls` compare their whole **search scope**: an omitted `path` (pi's default: the current directory) or a parent directory of a declared path triggers the ask as well. A fresh install pre-fills a **starter list** (`~/.ssh/`, `~/.gnupg`, `~/.mc`, shell rc/profile files), active from the first session after the initial run (any config change applies to new sessions) — a pre-filled *user declaration*, not a built-in floor: edit or empty it freely, add your own (`~/Documents/private`, …) alongside; existing configs are never rewritten
118
+ - `denyPaths` are plain paths you declare **protected** — touches trigger a terminal ask you adjudicate (non-interactive → deny); the classifier never learns the paths themselves, only that they exist. `grep`/`find`/`ls` compare their whole **search scope**: an omitted `path` (pi's default: the current directory) or a parent directory of a declared path triggers the ask as well. A fresh install pre-fills a **starter list** (`~/.ssh/`, `~/.gnupg`, `~/.mc`, shell rc/profile files)
119
+ - `ignoreTools` names uncovered tools (`todo`, `web_search`, MCP/custom tools) that skip adjudication — **allow with zero model calls**; entries naming covered tools (`bash`/`read`/`write`/`edit`/`grep`/`find`/`ls`/`powershell`) are inert: those stay governed by the deny floor and your allow/deny rules, and the self-protection layer always runs first. A fresh install pre-fills a **starter list** (`todo`, `ask_user_question`, `memory_write`, `memory_search` — observed harmless across the 1265-verdict production audit). Caveat: an exempted tool loses the classifier's `denyPaths` existence-hint vigilance (uncovered tools never hit the path extractor anyway)
113
120
  - `builtinDenyFloor: false` turns off the built-in danger/path floor (your risk; the self-protection layer below always stays on)
114
121
  - `classifierModel` pins the classifier model, e.g. `"zai/glm-5.3-flash:low"` (thinking suffix supported; default: session model with thinking off)
115
122
  - `classifierModel: "typesafe/jev-latest"` opts into the bundled **jev decisions adapter** — gray-zone verdicts via TypeSafe's jev (OpenRouter by default, or TypeSafe's official API directly with `PI_VERDICT_JEV_TRANSPORT=typesafe`); experimental, see [ADR-0003](docs/adr/0003-jev-decisions-adapter.md)
116
123
  - `audit: true` records every **gray-zone adjudication** (the full transcript sent to the classifier, its raw response, the parsed verdict) as JSONL under `~/.pi/agent/verdicts/<sessionId>.jsonl` — one file per session, the 20 most recent kept. Interactive asks also record your answer (`userAnswer` ground truth, written after the confirm resolves), and protected-path asks are recorded too (#62); rule allow/deny stays unaudited. Local-only and full-fidelity (protected-path plaintext may appear — it never leaves your machine; [ADR-0002](docs/adr/0002-deny-paths-deterministic-ask.md) boundary note); the agent can neither read nor write the directory. `/automode` shows the audit state and path while on
117
- - `notifyAllows: true` notifies on every **classifier allow** (reason + action line — e.g. jev's probability breakdown); default `false` keeps passes silent. Mechanical passes (your own allow rules, protected-path confirms) never notify; shadow-cache annotations stay debug-only; with both switches on the notification appears once
118
- - `classifierFallbackModel` (optional, [ADR-0004](docs/adr/0004-classifier-fallback-cascade.md)) adds a **second-layer classifier** consulted only when the first layer is uncertain (ask / fail-closed / jev confidence below `classifierFallbackConfidence`, default 50); `classifierFallbackMode: "shadow"` (default) observes without changing verdicts, `"enforce"` escalates strictness only (a safety ratchet — never relaxes; a failed fallback denies the triggered call). Off unless set — a natural pairing: jev first + a haiku-class fallback
124
+ - `notifyAllows: true` notifies on every **classifier allow** (reason + action line — e.g. jev's probability breakdown); default `false` keeps passes silent. Mechanical passes (your own allow rules, protected-path confirms) never notify; with both switches on the notification appears once
125
+ - `classifierMinConfidence` (optional, [ADR-0004](docs/adr/0004-classifier-fallback-cascade.md)) sets the **confidence floor**: a jev verdict below it is demoted — cascaded to `classifierFallbackModel` if set (`enforce`, the default = the second layer adjudicates, except a demoted **deny or ask** can never be auto-relaxed to an allow — a fail-closed layer emitted no verdict, so its rescue stands; `shadow` = records its opinion only and you are asked — `/automode` hints the activation switch), otherwise asked of you directly. At/above the floor the first layer is autonomous. A natural pairing: jev first + a haiku/flash-class fallback
119
126
 
120
127
  No built-in allowlist — every "always allow" claim is yours ([why](docs/configuration.md#why-no-built-in-allowlist)). Full reference: [docs/configuration.md](docs/configuration.md).
121
128
 
@@ -134,7 +141,7 @@ No built-in allowlist — every "always allow" claim is yours ([why](docs/config
134
141
  - **Hosts**: pi only. On omp the setting warns and falls back to the session model; and it must never be selected as the session model (no text generation — selecting it warns)
135
142
  - **Escape hatch**: `PI_VERDICT_JEV_URL` overrides the active transport's endpoint (OpenRouter's is an alpha API)
136
143
 
137
- jev's calibrated confidence is exactly what the fallback cascade keys on — pair it with a second layer (`"classifierFallbackModel": "anthropic/claude-haiku-4-5"`) to route its low-confidence calls to a deeper model ([ADR-0004](docs/adr/0004-classifier-fallback-cascade.md)).
144
+ jev's calibrated confidence is exactly what the confidence floor keys on — pair it with a second layer (`"classifierMinConfidence", "classifierFallbackModel"`) so its low-confidence calls go to a deeper model instead of standing ([ADR-0004](docs/adr/0004-classifier-fallback-cascade.md)).
138
145
 
139
146
  ### Self-protection (the gate guards itself — [ADR-0001](docs/adr/0001-self-protection-layer.md))
140
147
 
@@ -145,6 +152,8 @@ The gate's own files — the config and the installed extension copy — are **u
145
152
 
146
153
  Requires pi ≥ 0.84. Works in interactive and non-interactive (`-p`/json/rpc) sessions; in non-interactive modes `ask` degrades to `deny`.
147
154
 
155
+ ---
156
+
148
157
  ## How it compares
149
158
 
150
159
  | | three-state verdict | classifier sees context | fail direction | runtime deps |
@@ -173,6 +182,7 @@ tool_call
173
182
  │ ├─ your rules: user deny beats user allow
174
183
  │ ├─ denyPaths (ADR-0002): protected paths → terminal ask,
175
184
  │ │ before user allow; classifier sees an existence hint only
185
+ │ ├─ ignoreTools: your declared uncovered tools → allow, zero model calls
176
186
  │ └─ no built-in allowlist — every "always allow" claim is yours to make
177
187
  │
178
188
  ├─ 2. Gray zone → model classifier (defaults to session model — "self-reflection")
@@ -185,7 +195,6 @@ tool_call
185
195
  ├─ deny → block, reason returned to the agent
186
196
  └─ ask → human confirm; non-interactive modes degrade to deny
187
197
 
188
- [shadow cache] observe-only telemetry alongside 2/3, never changes a verdict
189
198
  ```
190
199
 
191
200
  **fail-closed**: classifier exception / timeout (25s) / contract violation → deny. Never silently allow.
@@ -194,7 +203,7 @@ tool_call
194
203
 
195
204
  Design decisions here are settled by measurement, and the lab notes ship with the repo:
196
205
 
197
- - [`research/cache-sim`](research/cache-sim/README.md) — replayed 1.2k+ real classifier verdicts to measure verdict-cache hit rate (**3.2%** → cache deferred, shadow-mode telemetry built instead)
206
+ - [`research/cache-sim`](research/cache-sim/README.md) — replayed 1.2k+ real classifier verdicts to measure verdict-cache hit rate (**3.2%** → serving cache declined; the runtime shadow telemetry built afterwards measured 3.3% and was later removed too, #73)
198
207
  - [`research/thinking-param-blackhole.md`](research/thinking-param-blackhole.md) — three-layer forensic root-cause of thinking models burning the classifier budget; why the fix is `thinkingEnabled: false`
199
208
  - [`research/rule-engine-sim`](research/rule-engine-sim/README.md) — measured a tree-sitter rule-engine port against 746 real bash calls (**absorbs 0 gray calls**) and rejected it
200
209
  - [`research/pi-permission-landscape.md`](research/pi-permission-landscape.md) — the competitive landscape this README's positioning is checked against
@@ -210,7 +219,6 @@ Design decisions here are settled by measurement, and the lab notes ship with th
210
219
  - AGENTS.md is not passed to the classifier as downweighted intent evidence (Claude Code does this)
211
220
  - parallel gray-zone calls are adjudicated serially
212
221
  - self-reflection means the session model adjudicates — point `--auto-mode-model` at a lighter model if verdict latency/cost matters (open question tracked in the issue tracker)
213
- - shadow cache is observe-only by decision; the serving switch is a one-line change once measured hit rates justify it
214
222
  - `denyPaths` bash extraction is token-level ([ADR-0002](docs/adr/0002-deny-paths-deterministic-ask.md)): command substitution, base64-embedded paths and external script contents produce no hit signal — those calls fall back to the classifier's existence-hint vigilance. MCP and custom tools bypass the extractor entirely (their gray-zone adjudication still carries the hint). Path normalization is base-tier only (ADR-0002): a nonexistent target written through a symlinked directory rebuilds no real form and produces no hit — that indirection falls to the hint vigilance too (the ancestor-rebuilding tier applies to the self-protection layer and the sensitivity floor, not denyPaths). Honest framing, same as the self-protection substring precedent: the deterministic layer is obfuscatable, which is exactly why a hit routes to *you* rather than silently deciding
215
223
  - `denyPaths` bash tokens contain no spaces: a *declared* path containing spaces cannot be spelled in a bash command in a way the extractor sees — `cat "/path with space/x"` splits into two tokens and never hits (file tools still hit, their path is not tokenized). A glob covering the final segment of a base (`cat /proj/pers*` against `denyPaths: ["/proj/personal"]`) also misses — the base's own name never appears literally. A recursive search issued from a shell misses in both spellings — no path argument (defaults to the cwd, e.g. a bare `rg foo`) or a parent-directory argument (`rg foo <parent-of-a-declared-path>`): an argument-less command contributes no token at all and bash tokens otherwise compare one-directionally, while the file tools' bidirectional subtree compare covers the same shapes issued through `grep`/`find`/`ls`. All three holes fall back to the classifier's existence hint, alongside substitution/base64 above
216
224
  - self-protection bash matching is substring regex — obfuscatable; the tamper-detection backstop catches within-session bypasses, but a cross-session baseline (hash + change confirmation at startup, incl. upgrade UX) is phase 2 per [ADR-0001](docs/adr/0001-self-protection-layer.md)
@@ -225,7 +233,7 @@ The name: the three-state **verdict** is the core concept. The UX keeps `/automo
225
233
  ```bash
226
234
  bun install
227
235
  bun run typecheck
228
- bun test # offline stub tests: self-protection, tamper detection, deny floor, user rules, denyPaths, bypass regression, classifier retry, shadow cache, commands, toggle shortcut
236
+ bun test # offline stub tests: self-protection, tamper detection, deny floor, user rules, denyPaths, bypass regression, classifier retry, commands, toggle shortcut
229
237
  ```
230
238
 
231
239
  Issue tracker and decision records live in the GitHub issues ("map" issue #1 indexes them).
package/README.zh-CN.md CHANGED
@@ -8,7 +8,7 @@
8
8
 
9
9
  **pi-verdict 是 [pi](https://pi.dev) 的 Claude Code 风格的 Auto mode 式的极简权限门禁:每次工具调用执行前先过检查——放行、拦截,或先问你。**
10
10
 
11
- - 只有1k行左右的极简代码
11
+ - 只有2k行左右的极简代码
12
12
  - 内置危险规则与你的 allow/deny 规则以零延迟先行裁决明确情形
13
13
  - 其余交给携带会话上下文的模型分类器
14
14
  - 任何不确定或失败一律 fail-closed, 绝不静默放行
@@ -70,7 +70,7 @@ pi-verdict 同时支持 [pi](https://github.com/badlogic/pi-mono) 与 [oh-my-pi]
70
70
  | 用户规则 | `~/.pi/agent/config/pi-verdict.json` | `~/.omp/agent/config/pi-verdict.json` |
71
71
  | 凭据文件(S0 硬 deny) | `~/.pi/agent/auth.json` | `~/.omp/agent/auth.json` |
72
72
 
73
- - `/automode` —— 显示当前状态:开/关 + 本会话影子缓存统计
73
+ - `/automode` —— 显示当前状态:开/关
74
74
  - `/automode on`
75
75
  - `/automode off`
76
76
  - `ctrl+shift+a` —— 静默切换主开关(footer 始终显示为唯一反馈;键位可经 `toggleShortcut` 重绑或禁用)
@@ -99,25 +99,32 @@ pi-verdict 同时支持 [pi](https://github.com/badlogic/pi-mono) 与 [oh-my-pi]
99
99
  "~/.zshrc",
100
100
  "~/.bashrc"
101
101
  ],
102
+ "ignoreTools": [
103
+ "todo",
104
+ "ask_user_question",
105
+ "memory_write",
106
+ "memory_search"
107
+ ],
102
108
  "builtinDenyFloor": true,
103
109
  "classifierModel": null,
104
110
  "toggleShortcut": "ctrl+shift+a",
105
111
  "audit": false,
106
112
  "notifyAllows": false,
113
+ "classifierMinConfidence": null,
107
114
  "classifierFallbackModel": null,
108
- "classifierFallbackConfidence": 50,
109
- "classifierFallbackMode": "shadow"
115
+ "classifierFallbackMode": "enforce"
110
116
  }
111
117
  ```
112
118
 
113
119
  - `allow`/`deny` 为 JS 正则数组;**`deny` 优先于 `allow`**,两者都优先于分类器
114
- - `denyPaths` 是你声明**受保护**的普通路径列表:触碰触发**终局 ask** 由你裁决(非交互降级 deny);分类器只被告知路径**存在**,路径明文永不出本机。`grep`/`find`/`ls` 按**整个搜索范围**比较:省略 `path`(pi 默认:当前目录)或传入位于声明路径之上的父目录,同样触发 ask。全新安装会预填一份**入门列表**(`~/.ssh/`、`~/.gnupg`、`~/.mc`、shell rc/profile 文件),自初次运行后的第一个会话起生效(一切配置变更均自新会话生效)——它是预填的*用户声明*而非内置 floor:可随意增删清空,也可与自己的路径(`~/Documents/private`、……)并列;既有配置永不被改写
120
+ - `denyPaths` 是你声明**受保护**的普通路径列表:触碰触发**终局 ask** 由你裁决(非交互降级 deny);分类器只被告知路径**存在**,路径明文永不出本机。`grep`/`find`/`ls` 按**整个搜索范围**比较:省略 `path`(pi 默认:当前目录)或传入位于声明路径之上的父目录,同样触发 ask。全新安装会预填一份**入门列表**(`~/.ssh/`、`~/.gnupg`、`~/.mc`、shell rc/profile 文件)
121
+ - `ignoreTools` 列出规则未覆盖的工具(`todo`、`web_search`、MCP/自定义工具):**直接放行、零模型调用**;列出已覆盖工具(`bash`/`read`/`write`/`edit`/`grep`/`find`/`ls`/`powershell`)的条目无效:它们仍受 deny floor 与你的 allow/deny 规则约束,自保护层也永远先行。全新安装会预填一份**入门列表**(`todo`、`ask_user_question`、`memory_write`、`memory_search`——来自项目 1265 条生产审计的观察) 注意:被豁免的工具失去分类器对 `denyPaths` 的存在性话术警戒(未覆盖工具本就不进路径提取器)
115
122
  - `builtinDenyFloor: false` 整体关闭内置危险/路径拦截(风险自担;下方自保护层永远开启)
116
123
  - `classifierModel` 指定分类器模型,如 `"zai/glm-5.3-flash:low"`(支持思考后缀;缺省 = 会话模型且显式关思考)
117
124
  - `classifierModel: "typesafe/jev-latest"` 启用随包的 **jev 决策适配器**——灰区裁决经 TypeSafe jev 完成(默认 OpenRouter,或 `PI_VERDICT_JEV_TRANSPORT=typesafe` 直连官方 API);实验性质,详见 [ADR-0003](docs/adr/0003-jev-decisions-adapter.md)
118
125
  - `audit: true` 把每次**灰区裁决**(发给分类器的完整转录、其原始响应、解析出的裁决)以 JSONL 记录到 `~/.pi/agent/verdicts/<sessionId>.jsonl`——按会话一分文件,保留最近 20 个。交互式 ask 还会记录你的应答(`userAnswer` ground truth,确认结束后落盘),protected-path ask 也入审计(#62);规则 allow/deny 仍不入。仅存本机且全保真(受保护路径明文可能出现——永不出本机;[ADR-0002](docs/adr/0002-deny-paths-deterministic-ask.md) 边界注);agent 对该目录读写双拒。开启时 `/automode` 会显示审计状态与路径
119
- - `notifyAllows: true` 对每次 **classifier 放行**发通知(reason + action 行——如 jev 的概率分解);默认 `false` 保持放行静默。机械放行(你自己的 allow 规则、protected-path 确认)永不通知;shadow 标注仍属 debug;两开关同开时通知只出现一次
120
- - `classifierFallbackModel`(可选,[ADR-0004](docs/adr/0004-classifier-fallback-cascade.md))添加**第二层分类器**,仅当第一层不确定时征询(ask / fail-closed / jev confidence 低于 `classifierFallbackConfidence`,默认 50);`classifierFallbackMode: "shadow"`(默认)只观察不改判,`"enforce"` 仅升严(安全棘轮——永不放宽;fallback 失败时该次触发调用 deny)。未设置即完全关闭——天然搭配:jev 打头 + haiku 级兜底
126
+ - `notifyAllows: true` 对每次 **classifier 放行**发通知(reason + action 行——如 jev 的概率分解);默认 `false` 保持放行静默。机械放行(你自己的 allow 规则、protected-path 确认)永不通知;两开关同开时通知只出现一次
127
+ - `classifierMinConfidence`(可选,[ADR-0004](docs/adr/0004-classifier-fallback-cascade.md))设定**置信地板**:低于它的 jev 裁决被降级——配置了 `classifierFallbackModel` 则级联(`enforce`,默认 = 第二层全权裁决,但降级的 **deny 与 ask** 永不被自动放宽为 allow——fail-closed 未产生任何裁决,其获救裁决照常生效;`shadow` = 只记录意见、由你裁决——`/automode` 会提示激活开关),否则直接问你。不低于地板时第一层自主。天然搭配:jev 打头 + haiku/flash 级兜底
121
128
 
122
129
  没有内置白名单——每一条「永远放行」声明都归你([为什么](docs/configuration.md#why-no-built-in-allowlist))。完整参考:[docs/configuration.md](docs/configuration.md)。
123
130
 
@@ -136,7 +143,7 @@ pi-verdict 同时支持 [pi](https://github.com/badlogic/pi-mono) 与 [oh-my-pi]
136
143
  - **宿主**:仅支持pi。omp 上该设置会警告并回退会话模型。也绝不能选作会话主模型(不生成文本,选中即警告)
137
144
  - **逃生口**:`PI_VERDICT_JEV_URL` 可覆盖当前 transport 的端点(OpenRouter 侧为 alpha 接口)
138
145
 
139
- jev 的校准 confidence 正是回退级联的触发依据——搭配第二层使用(`"classifierFallbackModel": "anthropic/claude-haiku-4-5"`),把低置信调用交给更深的模型([ADR-0004](docs/adr/0004-classifier-fallback-cascade.md))。
146
+ jev 的校准 confidence 正是置信地板的判定依据——搭配第二层使用(`"classifierMinConfidence", "classifierFallbackModel"`),让低置信调用交给更深的模型而非直接生效([ADR-0004](docs/adr/0004-classifier-fallback-cascade.md))。
140
147
 
141
148
  ### 自保护(门禁守护自身——[ADR-0001](docs/adr/0001-self-protection-layer.md))
142
149
 
@@ -147,6 +154,8 @@ jev 的校准 confidence 正是回退级联的触发依据——搭配第二层
147
154
 
148
155
  需要 pi ≥ 0.84。交互与非交互(`-p`/json/rpc)会话均支持;非交互模式下 `ask` 降级为 `deny`。
149
156
 
157
+ ---
158
+
150
159
  ## 与品类对比
151
160
 
152
161
  | | 三态裁决 | 分类器携带上下文 | fail 方向 | 运行时依赖 |
@@ -175,6 +184,7 @@ tool_call
175
184
  │ ├─ 用户规则:deny 优先于 allow
176
185
  │ ├─ denyPaths(ADR-0002):受保护路径 → 终局 ask,先于用户 allow;
177
186
  │ │ 分类器只见存在性话术
187
+ │ ├─ ignoreTools:用户声明的未覆盖工具 → 直接放行,零模型调用
178
188
  │ └─ 无内置白名单 —— 「永远放行」的声明由你自己做
179
189
  │
180
190
  ├─ 2. 灰区 → 模型分类器(默认继承会话模型 —— "自省")
@@ -187,7 +197,6 @@ tool_call
187
197
  ├─ deny → 拦截,理由回传 agent
188
198
  └─ ask → 人工确认;非交互模式降级为 deny
189
199
 
190
- [影子缓存] observe-only 遥测,与 2/3 并行,永不改变裁决
191
200
  ```
192
201
 
193
202
  **fail-closed**:分类器异常 / 超时(25s)/ 输出违反契约 → 拦截,绝不静默放行。
@@ -212,7 +221,6 @@ tool_call
212
221
  - AGENTS.md 未作为降权意图证据传入分类器(Claude Code 有此设计)
213
222
  - 并行灰区调用串行裁决
214
223
  - 自省意味着会话模型亲自裁决 —— 若延迟/成本敏感,用 `--auto-mode-model` 指向轻量模型(开放问题见 issue tracker)
215
- - 影子缓存按决议仅观察不生效;实测命中率达标后,生效开关是一行改动
216
224
  - `denyPaths` 的 bash 提取是 token 级([ADR-0002](docs/adr/0002-deny-paths-deterministic-ask.md)):命令替换、base64 内嵌路径、外部脚本内容不产生命中信号——这些调用回落到分类器的存在性话术警戒。MCP 与自定义工具完全绕过提取器(其灰区裁决仍带话术)。路径归一化亦为基础档(ADR-0002):经符号链接目录写入尚不存在的目标不重建真实形、不产生命中——该间接路径同样由话术警戒覆盖(祖先重建档只适用于自保护层与路径敏感度 floor,不适用 denyPaths)。诚实表述,与自保护子串正则同例:确定性层可被混淆——这正是命中交由**你**裁决而非静默决定的原因
217
225
  - `denyPaths` 的 bash token 不含空格:**声明路径本身含空格时**,bash 拼写无法被提取器识别——`cat "/path with space/x"` 被拆成两个 token 永不命中(文件类工具仍命中,其路径不经 token 化)。glob 覆盖基名末段(`denyPaths: ["/proj/personal"]` 时 `cat /proj/pers*`)同样漏过——基名自身从未字面出现。经 shell 发起的递归搜索在两种拼写下都漏过——不带路径参数(默认搜 cwd,如裸 `rg foo`)或带父目录参数(`rg foo <声明路径的父目录>`):无参命令根本不产生 token,带参时 bash token 只做单向比较;同一形状经 `grep`/`find`/`ls` 工具发起则由双向子树比较覆盖。三个洞与上述替换/base64 一样回落到分类器的存在性话术
218
226
  - 自保护 bash 匹配是子串正则——可被混淆绕过;变更检测兜底覆盖会话内绕过,跨会话基线(启动时哈希比对与变更确认,含升级 UX)按 ADR-0001 为二期
@@ -227,7 +235,7 @@ tool_call
227
235
  ```bash
228
236
  bun install
229
237
  bun run typecheck
230
- bun test # 离线桩测试:自保护 / 变更检测 / deny floor / 用户规则 / denyPaths / 绕过回归 / 分类器重试 / 影子缓存 / 命令 / toggle 快捷键
238
+ bun test # 离线桩测试:自保护 / 变更检测 / deny floor / 用户规则 / denyPaths / 绕过回归 / 分类器重试 / 命令 / toggle 快捷键
231
239
  ```
232
240
 
233
241
  Issue tracker 与决策记录在 GitHub issues(「地图」issue #1 为索引)。
@@ -39,14 +39,10 @@
39
39
  * Structure: the pipeline is adjudicate() — a zero-UI module returning a Verdict
40
40
  * value object (source: rule|protected-path|classifier|fail-closed, plus a
41
41
  * `degraded` flag for ask→deny in non-interactive sessions); the tool_call
42
- * handler maps verdicts to UI (notify/confirm/select) by source × degraded and
42
+ * handler maps verdicts to UI (notify/confirm/select) by source (the degraded
43
+ * context is implicit in the protected-path source's deny wording) and
43
44
  * runs IntegrityWatch (ADR-0001) as a pre-pipeline gate-integrity check.
44
45
  *
45
- * Shadow cache (observe-only, #7): gray-zone verdicts are replayed against a
46
- * double-key LRU(128) to measure would-be hit rate; recorded, never applied
47
- * (verdicts always come from the model), accumulating pi field data for the
48
- * "should a serving cache ship" question (#5 decision).
49
- *
50
46
  * fail-closed: classifier exception/timeout/contract violation → deny; in
51
47
  * non-interactive modes (no UI) ask → deny.
52
48
  *
@@ -60,8 +56,7 @@
60
56
  * suffix (pi-native --model syntax; default off
61
57
  * = thinking explicitly disabled)
62
58
  * PI_AUTO_MODE_MODEL env-var form of the above
63
- * --auto-mode-debug notify on every verdict (incl. allows); shadow
64
- * cache annotation on
59
+ * --auto-mode-debug notify on every verdict (incl. allows)
65
60
  * PI_AUTO_MODE_DEBUG=1 env-var form of the above (kept for compat)
66
61
  * <agentDir>/config/pi-verdict.json user rules: { allow: [regex], deny: [regex],
67
62
  * denyPaths: [path], builtinDenyFloor,
@@ -74,9 +69,9 @@
74
69
  * Known prototype simplifications (see README "Status & limitations"):
75
70
  * - no built-in bash allowlist; danger detection is regex floor (no AST parsing)
76
71
  * — unknown shapes go to the classifier
77
- * - serving verdict cache deferred (#5 decision): currently observe-only shadow
78
- * telemetry, revisit once measured; no circuit breaker (revisit signals =
79
- * deny-storm cost blowup / long non-interactive runs)
72
+ * - no serving verdict cache (#5 decision; runtime shadow telemetry removed #73 —
73
+ * re-evaluation belongs to offline replay tooling); no circuit breaker (revisit
74
+ * signals = deny-storm cost blowup / long non-interactive runs)
80
75
  * - AGENTS.md not passed to the classifier as downweighted intent evidence
81
76
  * - denyPaths bash extraction is token-level: command substitution, base64-
82
77
  * embedded paths and external script contents produce no hit signal — those
@@ -254,6 +249,14 @@ interface UserRules {
254
249
  deny: RegExp[];
255
250
  /** User-declared protected paths (ADR-0002): plain paths, tool-owned normalization; hit → ask */
256
251
  denyPaths: string[];
252
+ /** User-declared tool passthrough: names of tools OUTSIDE the command/file
253
+ * families (todo, web_search, MCP/custom tools, …) that skip adjudication —
254
+ * session-metadata/read-only tools the owner exempts, same stance as user allow
255
+ * rules: no built-in passthrough, every exemption is the user's own claim.
256
+ * Entries naming covered tools (bash/read/write/edit/grep/find/ls/powershell)
257
+ * are inert: those are governed by the deny floor and user allow/deny rules,
258
+ * which this list can never weaken. */
259
+ ignoreTools: string[];
257
260
  /** 内置 deny floor 开关(危险正则 + 路径敏感度 deny),默认 true;关闭后依赖用户规则与分类器 */
258
261
  builtinDenyFloor: boolean;
259
262
  /** 分类器模型 spec(provider/id);null = 未配置(自省继承会话模型) */
@@ -264,15 +267,19 @@ interface UserRules {
264
267
  audit: boolean;
265
268
  /** Allow visibility (#60): info notification on classifier allows; mechanical passes stay silent. Default off. */
266
269
  notifyAllows: boolean;
267
- /** #63: second-layer classifier spec (provider/id[:thinking]); null = the cascade is entirely off */
270
+ /** #67: autonomy floor for the first layer — a jev verdict with confidence strictly
271
+ * below this is demoted (cascaded to the fallback if configured, else asked of the
272
+ * user; non-interactive degrades to deny). null = floor off. */
273
+ classifierMinConfidence: number | null;
274
+ /** #63/#67: second-layer model spec (provider/id[:thinking]); consulted on demotion
275
+ * and fail-closed only. null = no second layer. */
268
276
  classifierFallbackModel: string | null;
269
- /** #63: trigger when the first layer's jev confidence is strictly below this (0–100). Default 50. */
270
- classifierFallbackConfidence: number;
271
- /** #63: "shadow" (default — observe-only, verdicts unchanged) | "enforce" (safety ratchet: the fallback may only escalate strictness, never relax) */
277
+ /** #67: does the second layer adjudicate cascaded calls ("enforce", default since
278
+ * 0.12.0) or only record its opinion while the human decides ("shadow")? */
272
279
  classifierFallbackMode: "shadow" | "enforce";
273
280
  }
274
281
 
275
- const EMPTY_RULES: UserRules = { allow: [], deny: [], denyPaths: [], builtinDenyFloor: true, classifierModel: null, toggleShortcut: DEFAULT_TOGGLE_SHORTCUT, audit: false, notifyAllows: false, classifierFallbackModel: null, classifierFallbackConfidence: 50, classifierFallbackMode: "shadow" };
282
+ const EMPTY_RULES: UserRules = { allow: [], deny: [], denyPaths: [], ignoreTools: [], builtinDenyFloor: true, classifierModel: null, toggleShortcut: DEFAULT_TOGGLE_SHORTCUT, audit: false, notifyAllows: false, classifierMinConfidence: null, classifierFallbackModel: null, classifierFallbackMode: "enforce" };
276
283
 
277
284
  /** This module's own file location (import.meta.url resolved; null = unresolvable). */
278
285
  const OWN_FILE_PATH: string | null = (() => {
@@ -321,7 +328,6 @@ function userConfigPath(): string {
321
328
  }
322
329
 
323
330
  const USER_CONFIG_TEMPLATE = `${JSON.stringify({
324
- _hint: "pi-verdict user rules — full reference: https://github.com/jesset/pi-verdict/blob/main/docs/configuration.md. deny beats allow. denyPaths: protected paths, any touch asks for your confirmation (non-interactive degrades to deny); the pre-filled starter list is your declaration, edit or empty freely. builtinDenyFloor=false disables the built-in danger floor at your own risk (the self-protection layer always stays on). classifierModel pins the classifier (provider/id, e.g. zai/glm-5.3-flash; empty = session model). classifierFallbackModel (optional) adds a second-layer classifier consulted only when the first layer is uncertain (ask / fail-closed / jev confidence below classifierFallbackConfidence, default 50); mode shadow (default) observes without changing verdicts, enforce escalates strictness only. toggleShortcut sets the master-switch toggle key (null or empty disables). This file is part of the permission gate: agent-side modification is denied — edit it manually outside pi. Changes apply to new sessions.",
325
331
  allow: ["^ls\\b"],
326
332
  deny: [],
327
333
  denyPaths: [
@@ -332,20 +338,29 @@ const USER_CONFIG_TEMPLATE = `${JSON.stringify({
332
338
  "~/.zshrc",
333
339
  "~/.bashrc",
334
340
  ],
341
+ ignoreTools: [
342
+ "todo",
343
+ "ask_user_question",
344
+ "memory_write",
345
+ "memory_search",
346
+ ],
335
347
  builtinDenyFloor: true,
336
348
  classifierModel: null,
337
349
  toggleShortcut: DEFAULT_TOGGLE_SHORTCUT,
338
350
  audit: false,
339
351
  notifyAllows: false,
352
+ classifierMinConfidence: null,
340
353
  classifierFallbackModel: null,
341
- classifierFallbackConfidence: 50,
342
- classifierFallbackMode: "shadow",
354
+ classifierFallbackMode: "enforce",
343
355
  }, null, 2)}\n`;
344
356
 
345
357
  /**
346
- * 加载用户规则。首启生成带注释模板(allow 内示例默认仅 ^ls\b 可用,其余为说明占位);
347
- * 配置缺失/损坏/字段非法一律回退空规则(安全默认,不失效),非法正则收集回报,
348
- * 非法 toggleShortcut 收集警告文案(与 skipped 同经 session_start 发出)。
358
+ * Load user rules. First run generates a template (bare keys + starter lists, no
359
+ * embedded prose — the full reference is docs/configuration.md; only ^ls\b in the
360
+ * allow example is live, the rest are placeholders). A missing/malformed/invalid
361
+ * config falls back to empty rules (safe default — the gate never disables);
362
+ * invalid regexes are collected and reported, an invalid toggleShortcut collects
363
+ * a warning text (both surface via session_start alongside `skipped`).
349
364
  */
350
365
  function loadUserRules(): { rules: UserRules; skipped: string[]; shortcutWarning: string | null } {
351
366
  try {
@@ -357,7 +372,7 @@ function loadUserRules(): { rules: UserRules; skipped: string[]; shortcutWarning
357
372
  } catch { /* 只读环境静默跳过 */ }
358
373
  return { rules: EMPTY_RULES, skipped: [], shortcutWarning: null };
359
374
  }
360
- let raw: { allow?: unknown; deny?: unknown; denyPaths?: unknown; builtinDenyFloor?: unknown; classifierModel?: unknown; toggleShortcut?: unknown; audit?: unknown; notifyAllows?: unknown; classifierFallbackModel?: unknown; classifierFallbackConfidence?: unknown; classifierFallbackMode?: unknown };
375
+ let raw: { allow?: unknown; deny?: unknown; denyPaths?: unknown; ignoreTools?: unknown; builtinDenyFloor?: unknown; classifierModel?: unknown; toggleShortcut?: unknown; audit?: unknown; notifyAllows?: unknown; classifierFallbackModel?: unknown; classifierFallbackConfidence?: unknown; classifierMinConfidence?: unknown; classifierFallbackMode?: unknown };
361
376
  try {
362
377
  raw = JSON.parse(fs.readFileSync(p, "utf8")) as typeof raw;
363
378
  } catch (err) {
@@ -385,11 +400,21 @@ function loadUserRules(): { rules: UserRules; skipped: string[]; shortcutWarning
385
400
  }
386
401
  return [x.trim()];
387
402
  });
403
+ // ignoreTools entries are plain tool names: only type-valid non-empty strings
404
+ // survive; anything else joins the same one-shot warning channel
405
+ const ignoreTools = (Array.isArray(raw.ignoreTools) ? raw.ignoreTools : []).flatMap((x) => {
406
+ if (typeof x !== "string" || !x.trim()) {
407
+ if (x !== undefined && x !== null) skipped.push(`ignoreTools: ${JSON.stringify(x)}`);
408
+ return [];
409
+ }
410
+ return [x.trim()];
411
+ });
388
412
  const shortcut = resolveToggleShortcut(raw.toggleShortcut);
389
- // #63: fallback cascade keys — invalid values skip into the one-shot warning channel and default (50 / shadow)
390
- const fbConfRaw = raw.classifierFallbackConfidence;
391
- const fbConfOk = typeof fbConfRaw === "number" && Number.isFinite(fbConfRaw) && fbConfRaw >= 0 && fbConfRaw <= 100;
392
- if (fbConfRaw !== undefined && !fbConfOk) skipped.push(`classifierFallbackConfidence: ${JSON.stringify(fbConfRaw)}`);
413
+ // #63/#67: confidence-floor keys — invalid values skip into the one-shot warning channel
414
+ if (raw.classifierFallbackConfidence !== undefined) skipped.push("classifierFallbackConfidence: renamed to classifierMinConfidence (0.11.0) — key ignored");
415
+ const minConfRaw = raw.classifierMinConfidence;
416
+ const minConfOk = typeof minConfRaw === "number" && Number.isFinite(minConfRaw) && minConfRaw >= 0 && minConfRaw <= 100;
417
+ if (minConfRaw !== undefined && minConfRaw !== null && !minConfOk) skipped.push(`classifierMinConfidence: ${JSON.stringify(minConfRaw)}`);
393
418
  const fbModeRaw = raw.classifierFallbackMode;
394
419
  if (fbModeRaw !== undefined && fbModeRaw !== "shadow" && fbModeRaw !== "enforce") skipped.push(`classifierFallbackMode: ${JSON.stringify(fbModeRaw)}`);
395
420
  return {
@@ -397,14 +422,15 @@ function loadUserRules(): { rules: UserRules; skipped: string[]; shortcutWarning
397
422
  allow: compile(raw.allow),
398
423
  deny: compile(raw.deny),
399
424
  denyPaths,
425
+ ignoreTools,
400
426
  builtinDenyFloor: raw.builtinDenyFloor !== false,
401
427
  classifierModel: typeof raw.classifierModel === "string" && raw.classifierModel.trim() ? raw.classifierModel.trim() : null,
402
428
  toggleShortcut: shortcut.key,
403
429
  audit: raw.audit === true,
404
430
  notifyAllows: raw.notifyAllows === true,
405
431
  classifierFallbackModel: typeof raw.classifierFallbackModel === "string" && raw.classifierFallbackModel.trim() ? raw.classifierFallbackModel.trim() : null,
406
- classifierFallbackConfidence: fbConfOk ? fbConfRaw : 50,
407
- classifierFallbackMode: fbModeRaw === "enforce" ? "enforce" : "shadow",
432
+ classifierMinConfidence: minConfOk ? minConfRaw : null,
433
+ classifierFallbackMode: fbModeRaw === undefined ? "enforce" : fbModeRaw === "enforce" ? "enforce" : "shadow", // invalid values land on the conservative shadow (standing invalid-config precedent); the key-less default is enforce
408
434
  },
409
435
  skipped,
410
436
  shortcutWarning: shortcut.warning,
@@ -918,7 +944,8 @@ class IntegrityWatch {
918
944
  * 2. user deny → deny (beats allow)
919
945
  * 3. denyPaths hit → terminal ask (ADR-0002: the declaring user adjudicates; before user allow)
920
946
  * 4. user allow → allow
921
- * 5. base (path tools' default allow/gray; everything else gray) → classifier
947
+ * 5. base (path tools' default allow/gray; uncovered tools: ignoreTools hit → allow
948
+ * passthrough, else gray) → classifier
922
949
  */
923
950
  function classifyByRules(toolName: string, input: Record<string, unknown>, cwd: string, user: UserRules, prot: ProtectedSet, denyPathBases: string[]): RuleResult {
924
951
  // 第 0 层:自保护层(ADR-0001)——先于一切,不可经任何配置豁免
@@ -941,7 +968,12 @@ function classifyByRules(toolName: string, input: Record<string, unknown>, cwd:
941
968
  const p = typeof input.path === "string" ? input.path : undefined;
942
969
  base = p ? classifyPath(toolName, p, cwd, false, user.builtinDenyFloor) : { verdict: "allow" };
943
970
  } else {
944
- base = { verdict: "gray", reason: `tool not covered by built-in rules: ${toolName}` };
971
+ // ignoreTools passthrough: user-declared exempt tools skip adjudication;
972
+ // reached only after self-protection and the deny floor, so neither can be
973
+ // weakened by it. Everything else uncovered stays gray for the classifier.
974
+ base = user.ignoreTools.includes(toolName)
975
+ ? { verdict: "allow", reason: `user ignoreTools passthrough: ${toolName}` }
976
+ : { verdict: "gray", reason: `tool not covered by built-in rules: ${toolName}` };
945
977
  }
946
978
  if (base.verdict === "deny") return base; // 内置 floor:deny 优先于一切用户规则
947
979
 
@@ -1027,7 +1059,6 @@ export type PipelineHost = Pick<ExtensionContext["sessionManager"], "getBranch"
1027
1059
  /**
1028
1060
  * 从会话分支收集精简转录原料:user 消息行与 assistant 工具调用行。
1029
1061
  * 丢弃 assistant 叙述/thinking 与 toolResult(注入面与 token 大头)。
1030
- * 影子缓存的 contextKey 与 buildTranscript 同源(同一批 user 行),保证键与模型输入一致。
1031
1062
  */
1032
1063
  function collectTranscriptParts(host: PipelineHost): { userLines: string[]; toolLines: string[] } {
1033
1064
  const userLines: string[] = [];
@@ -1266,161 +1297,42 @@ async function classifyWithModel(
1266
1297
  }
1267
1298
 
1268
1299
  // ============================================================================
1269
- // 影子缓存:双键命中率遥测(observe-only,#7;设计定案见 #5)
1270
- //
1271
- // 键设计(#5 定案):
1272
- // commandKey = hash(toolName + JSON.stringify(input) + cwd) —— 不做命令规范化
1273
- // contextKey = hash(最近 5 条 sanitized user 行,与 transcript 同源同窗口)
1274
- // 行为:
1275
- // 每次灰区裁决前查 would-be 命中;真实模型 allow/deny 回写(LRU 128,上下文变更覆写);
1276
- // ask 与 fail-closed 不入缓存;命中时对比缓存裁决与本次模型裁决(反事实一致性)。
1277
- // 永不生效:裁决永远来自模型,此处只记录。
1300
+ // Confidence cascade stats (#63/#67: observe-first, session-memory state)
1278
1301
  // ============================================================================
1279
1302
 
1280
- const SHADOW_LRU_MAX = 128;
1281
-
1282
- type ShadowVerdict = "allow" | "deny";
1283
- interface ShadowEntry {
1284
- ctxKey: string;
1285
- verdict: ShadowVerdict;
1286
- }
1287
-
1288
- /** FNV-1a 32 位摘要:仅会话内键用,非密码学 */
1289
- function fnv1a(s: string): string {
1290
- let h = 0x811c9dc5;
1291
- for (let i = 0; i < s.length; i++) {
1292
- h ^= s.charCodeAt(i);
1293
- h = Math.imul(h, 0x01000193);
1294
- }
1295
- return (h >>> 0).toString(16);
1296
- }
1297
-
1298
- interface ShadowStats {
1299
- gray: number; // 灰区裁决总数(含 ask/fail-closed)
1300
- hits: number; // 双键命中(would-be)
1301
- missNoEntry: number;
1302
- missCtx: number;
1303
- cmdRepeats: number; // 命令键重复(忽略 context 的上界口径)
1304
- divergeDangerous: number; // 命中且缓存 allow → 模型 deny(若缓存生效会放过本次拦截)
1305
- divergeConservative: number; // 命中且缓存 deny → 模型 allow
1306
- }
1307
-
1308
- type ShadowProbe =
1309
- | { result: "hit"; entry: ShadowEntry }
1310
- | { result: "no-entry" }
1311
- | { result: "ctx-changed"; prevVerdict: ShadowVerdict };
1312
-
1313
- class ShadowCache {
1314
- private lru = new Map<string, ShadowEntry>();
1315
- private seen = new Set<string>();
1316
- readonly stats: ShadowStats = { gray: 0, hits: 0, missNoEntry: 0, missCtx: 0, cmdRepeats: 0, divergeDangerous: 0, divergeConservative: 0 };
1317
-
1318
- /** 会话重置:清空 LRU 与统计(#5 定案:会话内存态) */
1319
- reset(): void {
1320
- this.lru.clear();
1321
- this.seen.clear();
1322
- Object.assign(this.stats, { gray: 0, hits: 0, missNoEntry: 0, missCtx: 0, cmdRepeats: 0, divergeDangerous: 0, divergeConservative: 0 });
1323
- }
1324
-
1325
- /** 灰区裁决前置查询(仅遥测,不影响裁决) */
1326
- probe(commandKey: string, ctxKey: string): ShadowProbe {
1327
- this.stats.gray++;
1328
- if (this.seen.has(commandKey)) this.stats.cmdRepeats++;
1329
- else this.seen.add(commandKey);
1330
- const entry = this.lru.get(commandKey);
1331
- if (!entry) {
1332
- this.stats.missNoEntry++;
1333
- return { result: "no-entry" };
1334
- }
1335
- if (entry.ctxKey !== ctxKey) {
1336
- this.stats.missCtx++;
1337
- return { result: "ctx-changed", prevVerdict: entry.verdict };
1338
- }
1339
- this.stats.hits++;
1340
- // LRU 位置刷新,保留原裁决(命中即重放)
1341
- this.lru.delete(commandKey);
1342
- this.lru.set(commandKey, entry);
1343
- return { result: "hit", entry };
1344
- }
1345
-
1346
- /** 真实模型 allow/deny 裁决后回写;ask 与 fail-closed 不入 */
1347
- record(commandKey: string, ctxKey: string, verdict: ShadowVerdict): void {
1348
- this.lru.delete(commandKey);
1349
- this.lru.set(commandKey, { ctxKey, verdict });
1350
- if (this.lru.size > SHADOW_LRU_MAX) {
1351
- const oldest = this.lru.keys().next().value;
1352
- if (oldest !== undefined) this.lru.delete(oldest);
1353
- }
1354
- }
1355
-
1356
- /** 命中后的反事实一致性计数(仅与可缓存裁决对比;ask/fail-closed 不可比) */
1357
- countDivergence(cached: ShadowVerdict, actual: ShadowVerdict): void {
1358
- if (cached === actual) return;
1359
- if (cached === "allow" && actual === "deny") this.stats.divergeDangerous++;
1360
- else this.stats.divergeConservative++;
1361
- }
1362
-
1363
- /** /automode 展示用摘要 */
1364
- summary(): string {
1365
- const s = this.stats;
1366
- if (s.gray === 0) return "shadow cache: no gray-zone verdicts yet this session";
1367
- const rate = ((100 * s.hits) / s.gray).toFixed(1);
1368
- return `shadow cache: gray ${s.gray} · two-key hits ${s.hits} (${rate}%) · miss no-entry ${s.missNoEntry}/ctx-changed ${s.missCtx} · cmd repeats ${s.cmdRepeats} · divergence dangerous ${s.divergeDangerous}/conservative ${s.divergeConservative}`;
1369
- }
1370
- }
1371
-
1372
- function shadowCommandKey(toolName: string, input: Record<string, unknown>, cwd: string): string {
1373
- return fnv1a(`${toolName}\u0000${JSON.stringify(input)}\u0000${cwd}`);
1374
- }
1375
-
1376
- function shadowContextKey(host: PipelineHost): string {
1377
- const { userLines } = collectTranscriptParts(host);
1378
- return fnv1a(userLines.slice(-MAX_USER_MESSAGES).join("\u0000"));
1379
- }
1380
-
1381
- function shadowTag(probe: ShadowProbe): string {
1382
- if (probe.result === "hit") return `(shadow cache: would-hit ${probe.entry.verdict})`;
1383
- if (probe.result === "ctx-changed") return `(shadow cache: miss:context-changed, previous ${probe.prevVerdict})`;
1384
- return `(shadow cache: miss:no-entry)`;
1385
- }
1386
-
1387
- // ============================================================================
1388
- // Fallback cascade stats (#63: observe-first, session-memory state; the #7 discipline)
1389
- // ============================================================================
1390
-
1391
- /** #63: ratchet strictness order — the fallback may only escalate, never relax */
1392
- const STRICTNESS_RANK: Record<"allow" | "ask" | "deny", number> = { allow: 0, ask: 1, deny: 2 };
1393
-
1394
1303
  interface FallbackStats {
1395
- triggered: number; // the gate fired (ask / fail-closed / confidence below threshold)
1396
- agreed: number; // fallback verdict no stricter than the first layer's
1397
- escalated: number; // fallback stricter than the first layer (enforce applies it; shadow observes the would-be)
1304
+ triggered: number; // the floor fired or the first layer fail-closed (with a fallback configured)
1305
+ agreed: number; // fallback verdict equals the first layer's (fail-closed defaults to deny)
1306
+ overruled: number; // fallback verdict differs (enforce applies it; shadow observes the would-be)
1398
1307
  errored: number; // fallback unresolvable or its call failed
1308
+ rescuedAllow: number; // #71: fail-closed origin the fallback ruled allow (enforce: actually allowed; shadow: would) — counted within overruled as well, surfaced distinctly for the rescue reading
1399
1309
  }
1400
1310
 
1401
1311
  class FallbackCascade {
1402
- readonly stats: FallbackStats = { triggered: 0, agreed: 0, escalated: 0, errored: 0 };
1312
+ readonly stats: FallbackStats = { triggered: 0, agreed: 0, overruled: 0, errored: 0, rescuedAllow: 0 };
1403
1313
 
1404
- /** Session reset (#7 discipline: session-memory state) */
1314
+ /** Session reset (session-memory state) */
1405
1315
  reset(): void {
1406
- Object.assign(this.stats, { triggered: 0, agreed: 0, escalated: 0, errored: 0 });
1316
+ Object.assign(this.stats, { triggered: 0, agreed: 0, overruled: 0, errored: 0, rescuedAllow: 0 });
1407
1317
  }
1408
1318
 
1409
- note(first: "allow" | "ask" | "deny", fb: "allow" | "ask" | "deny" | null): void {
1319
+ note(first: "allow" | "ask" | "deny" | null, fb: "allow" | "ask" | "deny" | null): void {
1410
1320
  this.stats.triggered++;
1411
1321
  if (fb === null) {
1412
1322
  this.stats.errored++;
1413
1323
  return;
1414
1324
  }
1415
- if (STRICTNESS_RANK[fb] > STRICTNESS_RANK[first]) this.stats.escalated++;
1325
+ // A fail-closed origin produced no first-layer verdict; its default outcome is deny
1326
+ if (first === null && fb === "allow") this.stats.rescuedAllow++;
1327
+ if ((first ?? "deny") !== fb) this.stats.overruled++;
1416
1328
  else this.stats.agreed++;
1417
1329
  }
1418
1330
 
1419
1331
  /** Summary line for /automode */
1420
1332
  summary(mode: "shadow" | "enforce"): string {
1421
1333
  const s = this.stats;
1422
- if (s.triggered === 0) return "fallback cascade: not triggered this session";
1423
- return `fallback cascade (${mode}): triggered ${s.triggered} · agreed ${s.agreed} · ${mode === "enforce" ? "escalated" : "would-escalate"} ${s.escalated} · errored ${s.errored}`;
1334
+ if (s.triggered === 0) return "confidence cascade: not triggered this session";
1335
+ return `confidence cascade (${mode}): triggered ${s.triggered} · agreed ${s.agreed} · ${mode === "enforce" ? "overruled" : "would-overrule"} ${s.overruled} · ${mode === "enforce" ? "rescued-allow" : "would-rescue-allow"} ${s.rescuedAllow} · errored ${s.errored}`;
1424
1336
  }
1425
1337
  }
1426
1338
 
@@ -1431,21 +1343,21 @@ class FallbackCascade {
1431
1343
 
1432
1344
  const AUDIT_KEEP_SESSIONS = 20;
1433
1345
 
1434
- /** #63: second-layer classifier outcome on a triggered call. The record's top-level
1435
- * fields keep first-layer semantics for corpus comparability (grill decision); the
1436
- * verdict actually applied under enforce lives in `effective` (absent in shadow). */
1346
+ /** #63/#67: second-layer outcome on a cascaded call. The record's top-level fields keep
1347
+ * first-layer semantics for corpus comparability; the verdict actually applied under
1348
+ * enforce lives in `effective` (failure rows carry the ask the human got). */
1437
1349
  export interface FallbackAudit {
1438
1350
  model: string;
1439
1351
  mode: "shadow" | "enforce";
1440
- triggeredBy: "ask" | "confidence" | "fail-closed";
1441
- /** jev confidence that fired the gate; null unless triggeredBy = "confidence" */
1352
+ triggeredBy: "confidence" | "fail-closed";
1353
+ /** jev confidence that fired the floor; null unless triggeredBy = "confidence" */
1442
1354
  confidence: number | null;
1443
1355
  /** null = the fallback call itself failed (unresolvable model, timeout, parse) */
1444
1356
  verdict: "allow" | "ask" | "deny" | null;
1445
1357
  reason: string | null;
1446
1358
  durationMs: number;
1447
1359
  error: string | null;
1448
- /** enforce mode only: the verdict applied after the ratchet */
1360
+ /** enforce mode only: the verdict applied (pre headless-degradation) */
1449
1361
  effective?: "allow" | "ask" | "deny";
1450
1362
  }
1451
1363
 
@@ -1470,7 +1382,6 @@ export interface AuditRecord {
1470
1382
  /** #62: protected-path asks are recorded too — their user answers grade the
1471
1383
  * denyPaths rules; rule allow/deny verdicts remain unaudited. */
1472
1384
  source: "model" | "fail-closed" | "protected-path";
1473
- shadow: string;
1474
1385
  degraded: boolean;
1475
1386
  /** #62 ground truth: the user's answer to an interactive ask confirm. Present only
1476
1387
  * on records whose confirm actually ran; headless/degraded asks omit it. */
@@ -1479,7 +1390,9 @@ export interface AuditRecord {
1479
1390
  answeredAt?: string;
1480
1391
  /** #62: protected-path records only — the matched path. */
1481
1392
  detail?: string;
1482
- /** #63: second-layer outcome when the uncertainty gate fired. */
1393
+ /** #67: the confidence floor fired — the first-layer verdict was demoted. */
1394
+ demoted?: true;
1395
+ /** #63/#67: second-layer outcome when the fallback was consulted. */
1483
1396
  fallback?: FallbackAudit;
1484
1397
  }
1485
1398
 
@@ -1549,7 +1462,6 @@ export class AuditLog {
1549
1462
  */
1550
1463
  export class SessionState {
1551
1464
  readonly prot: ProtectedSet;
1552
- readonly shadow = new ShadowCache();
1553
1465
  readonly fallback = new FallbackCascade();
1554
1466
  userRules: UserRules;
1555
1467
  audit: AuditLog | null;
@@ -1569,12 +1481,11 @@ export class SessionState {
1569
1481
  }
1570
1482
 
1571
1483
  /** 会话重置:重载用户规则(配置改动新会话生效)+ 按会话 cwd 重锚 denyPaths
1572
- * (ADR-0002: 每会话锚定一次)+ 清影子缓存;返回加载报告供表现层通知 */
1484
+ * (ADR-0002: 每会话锚定一次);返回加载报告供表现层通知 */
1573
1485
  reset(cwd: string): { skipped: string[]; shortcutWarning: string | null } {
1574
1486
  const loaded = loadUserRules();
1575
1487
  this.userRules = loaded.rules;
1576
1488
  this.denyPathBases = anchorDenyPaths(loaded.rules.denyPaths, cwd); // anchored to the session cwd, once (ADR-0002)
1577
- this.shadow.reset();
1578
1489
  this.fallback.reset();
1579
1490
  this.audit = this.makeAudit(loaded.rules);
1580
1491
  return { skipped: loaded.skipped, shortcutWarning: loaded.shortcutWarning };
@@ -1599,14 +1510,13 @@ export type VerdictSource = "rule" | "protected-path" | "classifier" | "fail-clo
1599
1510
 
1600
1511
  /** 判定管线的输出值对象:一次 tool_call 的完整裁决。detail 为 UI-only 明文(受保护
1601
1512
  * 路径仅入本地确认框,ADR-0002 零泄漏承诺——reason 与通知永不携带);degraded 标记
1602
- * ask 在无 UI 会话的降级产物;shadow 为影子缓存标注(仅 debug 呈现拼接用)。 */
1513
+ * ask 在无 UI 会话的降级产物。 */
1603
1514
  export interface Verdict {
1604
1515
  verdict: "allow" | "ask" | "deny";
1605
1516
  reason: string;
1606
1517
  detail?: string;
1607
1518
  source: VerdictSource;
1608
1519
  degraded: boolean;
1609
- shadow?: string;
1610
1520
  /** #62: pending audit record for an interactive ask — adjudicate defers the append so
1611
1521
  * the handler can attach the user's answer after the confirm resolves. The handler
1612
1522
  * owns the single finalize: append with userAnswer/answeredAt, or without them when
@@ -1627,74 +1537,92 @@ export interface AdjudicateEnv {
1627
1537
  getFallbackModel?: () => { model: NonNullable<ExtensionContext["model"]>; thinking: ThinkingLevel } | null;
1628
1538
  }
1629
1539
 
1630
- /** #63: should the second layer be consulted for this first-layer outcome? Precedence:
1631
- * fail-closed → ask → jev confidence strictly below the threshold. LLM reasons carry
1632
- * no numeric confidence (parseJevConfidence → null) — their gate is ask/fail-closed only. */
1633
- function fallbackTrigger(outcome: ClassifierOutcome, rules: UserRules): { triggeredBy: "ask" | "confidence" | "fail-closed"; confidence: number | null } | null {
1634
- if (!rules.classifierFallbackModel) return null;
1635
- if (outcome.source === "fail-closed") return { triggeredBy: "fail-closed", confidence: null };
1636
- if (outcome.verdict === "ask") return { triggeredBy: "ask", confidence: null };
1540
+ /** #67: the confidence floor. Below it the first layer abstains and the call cascades —
1541
+ * to the fallback if configured, else to the human (headless degrades to deny). Numeric
1542
+ * confidence exists only on jev-formatted reasons; LLM first layers never demote. */
1543
+ function confidenceDemotion(outcome: ClassifierOutcome, rules: UserRules): { confidence: number } | null {
1544
+ if (rules.classifierMinConfidence === null || outcome.source === "fail-closed") return null;
1637
1545
  const conf = parseJevConfidence(outcome.reason);
1638
- if (conf !== null && conf < rules.classifierFallbackConfidence) return { triggeredBy: "confidence", confidence: conf };
1546
+ if (conf !== null && conf < rules.classifierMinConfidence) return { confidence: conf };
1639
1547
  return null;
1640
1548
  }
1641
1549
 
1642
1550
  interface CascadeResult {
1643
- /** audit material; absent when no trigger fired */
1551
+ /** set whenever the confidence floor fired (with or without a fallback) */
1552
+ demoted?: true;
1553
+ /** audit material; present when the fallback was consulted */
1644
1554
  fb?: FallbackAudit;
1645
- /** enforce-mode override; absent = keep the first-layer verdict (shadow never overrides) */
1555
+ /** the applied outcome when the cascade changes it (pre-degradation — the caller's
1556
+ * tail applies the usual headless ask → deny rule) */
1646
1557
  effective?: { verdict: "allow" | "ask" | "deny"; reason: string; source: "classifier" | "fail-closed" };
1647
1558
  }
1648
1559
 
1649
- /** #63: run the second layer on a triggered call. Safety ratchet: the fallback may
1650
- * only escalate strictness, never relax. A failed fallback (unresolvable model or
1651
- * failed call) denies in enforce — an explicitly configured second layer must not
1652
- * silently degrade the gate to single-layer (grill decision); in shadow a failure
1653
- * is recorded and never changes the verdict. */
1654
- async function runFallbackCascade(
1560
+ /** #67: run the cascade for one triggered call. `first` is the first-layer verdict, or
1561
+ * null when the first layer never produced one (fail-closed origin). Semantics:
1562
+ * - demotion with no fallback → ask the human
1563
+ * - shadow → the fallback records its opinion; a demotion still asks the human, a
1564
+ * fail-closed deny stands
1565
+ * - enforce → the fallback adjudicates de novo, with one carve-out family (#71): a demoted
1566
+ * first-layer deny or ask may not be auto-relaxed to an allow — the human decides; a
1567
+ * fail-closed origin has no first-layer verdict, so any fallback ruling applies
1568
+ * - fallback failure/unresolvable on a cascaded call → ask the human (the tier that was
1569
+ * to adjudicate is down); headless degrades downstream */
1570
+ async function runConfidenceCascade(
1655
1571
  state: SessionState,
1656
1572
  env: AdjudicateEnv,
1657
- first: "allow" | "ask" | "deny",
1658
- trigger: { triggeredBy: "ask" | "confidence" | "fail-closed"; confidence: number | null },
1573
+ first: { verdict: "allow" | "ask" | "deny"; reason: string } | null,
1574
+ trigger: { kind: "demotion"; confidence: number } | { kind: "fail-closed" },
1659
1575
  denyPathsActive: boolean,
1660
1576
  actionLine: string,
1661
1577
  ): Promise<CascadeResult> {
1662
1578
  const rules = state.userRules;
1663
- if (!rules.classifierFallbackModel || !env.getFallbackModel) return {};
1579
+ const demotionAsk = (): CascadeResult["effective"] => ({
1580
+ verdict: "ask",
1581
+ reason: `${first!.reason} (confidence ${trigger.kind === "demotion" ? trigger.confidence : "?"}% is below your classifierMinConfidence of ${rules.classifierMinConfidence}%)`,
1582
+ source: "classifier",
1583
+ });
1584
+ const getFb = env.getFallbackModel;
1585
+ if (!rules.classifierFallbackModel || !getFb) {
1586
+ // A fail-closed without a fallback keeps its deny; a demotion asks the human
1587
+ return trigger.kind === "demotion" ? { demoted: true, effective: demotionAsk() } : {};
1588
+ }
1664
1589
  const mode = rules.classifierFallbackMode;
1665
1590
  const start = Date.now();
1666
- const base = { mode, triggeredBy: trigger.triggeredBy, confidence: trigger.confidence };
1667
- // A failed fallback (unresolvable model or failed call) records the error and, under
1668
- // enforce, denies the triggered call; `fallback.effective` carries the applied "deny"
1669
- // so failure rows read through the same sub-object as every other enforce row
1591
+ const base = { mode, triggeredBy: trigger.kind === "demotion" ? ("confidence" as const) : ("fail-closed" as const), confidence: trigger.kind === "demotion" ? trigger.confidence : null };
1592
+ const demotedMark = trigger.kind === "demotion" ? ({ demoted: true } as const) : {};
1593
+ const shadowApplied = trigger.kind === "demotion" ? { effective: demotionAsk() } : {};
1670
1594
  const failed = (model: string, error: string): CascadeResult => {
1671
- state.fallback.note(first, null);
1595
+ state.fallback.note(first?.verdict ?? null, null);
1672
1596
  const fb: FallbackAudit = { ...base, model, verdict: null, reason: null, durationMs: Date.now() - start, error };
1673
- return mode === "enforce" ? { fb: { ...fb, effective: "deny" }, effective: { verdict: "deny", reason: "fallback classifier unavailable (fail-closed)", source: "fail-closed" } } : { fb };
1597
+ if (mode === "shadow") return { ...demotedMark, fb, ...shadowApplied };
1598
+ return { ...demotedMark, fb: { ...fb, effective: "ask" }, effective: { verdict: "ask", reason: "fallback classifier unavailable (first layer abstained) — your call", source: "fail-closed" } };
1674
1599
  };
1675
- const resolved = env.getFallbackModel();
1600
+ const resolved = getFb();
1676
1601
  if (!resolved) return failed(rules.classifierFallbackModel, "fallback model unresolvable (not found or no configured auth)");
1677
1602
  const outcome = await classifyWithModel(env.host, env.signal, env.complete, resolved.model, actionLine, resolved.thinking, denyPathsActive, FALLBACK_TIMEOUT_MS);
1678
- const durationMs = Date.now() - start;
1679
1603
  if (outcome.source !== "model") return failed(resolved.model.id, outcome.reason);
1680
- state.fallback.note(first, outcome.verdict);
1681
- const fb: FallbackAudit = { ...base, model: resolved.model.id, verdict: outcome.verdict, reason: outcome.reason, durationMs, error: null };
1682
- if (mode === "enforce") {
1683
- const effective = STRICTNESS_RANK[outcome.verdict] > STRICTNESS_RANK[first] ? outcome.verdict : first;
1684
- if (effective !== first) {
1685
- return { fb: { ...fb, effective }, effective: { verdict: outcome.verdict, reason: `${outcome.reason} (second-opinion classifier escalated ${first} to ${outcome.verdict})`, source: "classifier" } };
1686
- }
1687
- return { fb: { ...fb, effective } };
1604
+ state.fallback.note(first?.verdict ?? null, outcome.verdict);
1605
+ const fb: FallbackAudit = { ...base, model: resolved.model.id, verdict: outcome.verdict, reason: outcome.reason, durationMs: Date.now() - start, error: null };
1606
+ if (mode === "shadow") return { ...demotedMark, fb, ...shadowApplied };
1607
+ // The carve-outs on second-layer authority (#71): it may not auto-relax a negative
1608
+ // first-layer verdict — a demoted deny OR ask that the fallback would allow goes to
1609
+ // the human (headless degrades to deny downstream). A fail-closed origin has no
1610
+ // first-layer verdict to relax; its fallback allow is a de novo ruling and stands.
1611
+ if (trigger.kind === "demotion" && (first?.verdict === "deny" || first?.verdict === "ask") && outcome.verdict === "allow") {
1612
+ return { demoted: true, fb: { ...fb, effective: "ask" }, effective: { verdict: "ask", reason: `${outcome.reason} (first layer said ${first!.verdict} at confidence ${trigger.confidence}%; second opinion allows — your call)`, source: "classifier" } };
1688
1613
  }
1689
- return { fb };
1614
+ return { ...demotedMark, fb: { ...fb, effective: outcome.verdict }, effective: { verdict: outcome.verdict, reason: outcome.reason, source: "classifier" } };
1690
1615
  }
1691
1616
 
1692
1617
  /**
1693
1618
  * 判定管线(CONTEXT.md「判定管线」词条的实现):自保护 → 内置 floor → 用户 deny →
1694
- * denyPaths ask → 用户 allow → 灰区分类器;ask 降级(无 UI → deny)与 fail-closed
1695
- * 内建于此,两处重复的降级实现自此唯一。零 UI:表现(notify/confirm/select)由扩展
1696
- * handler 按 source × degraded 模板呈现;变更检测(IntegrityWatch)是管线前置的
1697
- * 独立关注点,不在 adjudicate 内。导出仅为测试(内部 seam 的测试面,#35 既有模式)。
1619
+ * denyPaths ask → user allow → gray-zone classifier; ask degradation (no UI → deny)
1620
+ * and fail-closed are built in here — the two formerly duplicated degradation
1621
+ * implementations now live in one place. Zero UI: presentation (notify/confirm/
1622
+ * select) is the handler's job, keyed on source alone (the degraded context is
1623
+ * implicit in the protected-path deny wording); IntegrityWatch (tamper detection)
1624
+ * is a pre-pipeline concern, not part of adjudicate. Exported for tests only
1625
+ * (the internal-seam test surface, the standing #35 pattern).
1698
1626
  */
1699
1627
  export async function adjudicate(
1700
1628
  state: SessionState,
@@ -1712,7 +1640,7 @@ export async function adjudicate(
1712
1640
  // appends immediately. Recording stays observe-only — it never changes a verdict; write
1713
1641
  // failures stay fail-soft in the sink and surface once via drainWarning.
1714
1642
  const actionLine = toolCallLine(call.toolName, call.input);
1715
- const buildRecord = (v: Pick<AuditRecord, "verdict" | "reason" | "source" | "degraded">, raw: ClassifierOutcome["auditRaw"] | null, shadow: string): AuditRecord => ({
1643
+ const buildRecord = (v: Pick<AuditRecord, "verdict" | "reason" | "source" | "degraded">, raw: ClassifierOutcome["auditRaw"] | null): AuditRecord => ({
1716
1644
  ts: new Date().toISOString(),
1717
1645
  sessionId: env.host.getSessionId(),
1718
1646
  cwd: env.cwd,
@@ -1723,18 +1651,17 @@ export async function adjudicate(
1723
1651
  thinking: raw?.thinking ?? null,
1724
1652
  transcript: raw?.transcript ?? null,
1725
1653
  rawResponse: raw?.rawResponse ?? null,
1726
- shadow,
1727
1654
  ...v,
1728
1655
  });
1729
1656
 
1730
1657
  if (rule.verdict === "ask") {
1731
1658
  // denyPaths 命中 → ask 终局(ADR-0002):声明者本人裁决例外;无 UI 降级为 deny
1732
1659
  if (env.hasUI) {
1733
- const ppRecord: AuditRecord = { ...buildRecord({ verdict: "ask", reason: rule.reason ?? "", source: "protected-path", degraded: false }, null, "-"), detail: rule.detail };
1660
+ const ppRecord: AuditRecord = { ...buildRecord({ verdict: "ask", reason: rule.reason ?? "", source: "protected-path", degraded: false }, null), detail: rule.detail };
1734
1661
  return { verdict: "ask", reason: rule.reason ?? "", detail: rule.detail, source: "protected-path", degraded: false, ...(state.audit ? { pendingAudit: ppRecord } : {}) };
1735
1662
  }
1736
1663
  // headless: the ask degrades to deny — recorded like the gray-zone rule (the effective post-degradation verdict is what lands in the record)
1737
- state.audit?.append({ ...buildRecord({ verdict: "deny", reason: rule.reason ?? "", source: "protected-path", degraded: true }, null, "-"), detail: rule.detail });
1664
+ state.audit?.append({ ...buildRecord({ verdict: "deny", reason: rule.reason ?? "", source: "protected-path", degraded: true }, null), detail: rule.detail });
1738
1665
  return { verdict: "deny", reason: rule.reason ?? "", detail: rule.detail, source: "protected-path", degraded: true };
1739
1666
  }
1740
1667
 
@@ -1743,56 +1670,60 @@ export async function adjudicate(
1743
1670
  const resolved = env.getModel();
1744
1671
  if (!resolved) {
1745
1672
  const reason = "no classifier model available (fail-closed)";
1746
- // #63: no-model fail-closed triggers the cascade as well — the ratchet has no
1747
- // exception for first-layer absence (grill decision: enforce can never relax this
1748
- // deny; in shadow it is observability only)
1749
- const cascade = await runFallbackCascade(state, env, "deny", { triggeredBy: "fail-closed", confidence: null }, state.userRules.denyPaths.length > 0, actionLine);
1750
- const fcRecord = buildRecord({ verdict: "deny", reason, source: "fail-closed", degraded: false }, null, "-");
1673
+ // #67: a fail-closed origin cascades to the fallback if configured — under enforce
1674
+ // the fallback adjudicates de novo (superseding the 0.10.0 ratchet decision);
1675
+ // shadow records its opinion and the deny stands
1676
+ const cascade = await runConfidenceCascade(state, env, null, { kind: "fail-closed" }, state.userRules.denyPaths.length > 0, actionLine);
1677
+ const eff = cascade.effective;
1678
+ // #71: a fail-closed layer emits no negative verdict — its deny is a default, not
1679
+ // a ruling. When an enforcing fallback rescues the call, the record's top level
1680
+ // carries the applied verdict; a shadow rescue (no effective) keeps the deny.
1681
+ const effAskHeadless = eff?.verdict === "ask" && !env.hasUI;
1682
+ const fcRecord = buildRecord({ verdict: eff ? (effAskHeadless ? "deny" : eff.verdict) : "deny", reason: eff?.reason ?? reason, source: "fail-closed", degraded: effAskHeadless }, null);
1751
1683
  if (cascade.fb) fcRecord.fallback = cascade.fb;
1684
+ if (eff?.verdict === "ask" && env.hasUI) {
1685
+ return { verdict: "ask", reason: eff.reason, source: eff.source, degraded: false, ...(state.audit ? { pendingAudit: fcRecord } : {}) };
1686
+ }
1752
1687
  state.audit?.append(fcRecord);
1688
+ if (eff?.verdict === "allow") return { verdict: "allow", reason: eff.reason, source: "classifier", degraded: false };
1689
+ if (eff) return { verdict: "deny", reason: eff.reason, source: eff.source, degraded: effAskHeadless };
1753
1690
  return { verdict: "deny", reason, source: "fail-closed", degraded: false };
1754
1691
  }
1755
1692
 
1756
- // 影子缓存(observe-only):前置查询 would-be 命中,不改变任何裁决
1757
- const cmdKey = shadowCommandKey(call.toolName, call.input, env.cwd);
1758
- const ctxKey = shadowContextKey(env.host);
1759
- const probe = state.shadow.probe(cmdKey, ctxKey);
1760
-
1761
1693
  const outcome = await classifyWithModel(env.host, env.signal, env.complete, resolved.model, actionLine, resolved.thinking, state.userRules.denyPaths.length > 0);
1762
1694
 
1763
- // 影子回记:真实模型 allow/deny 入缓存;ask 与 fail-closed 不入(#5 定案);
1764
- // 命中且本次为可缓存裁决时,对比反事实一致性
1765
- if (outcome.source === "model" && outcome.verdict !== "ask") {
1766
- if (probe.result === "hit") state.shadow.countDivergence(probe.entry.verdict, outcome.verdict);
1767
- state.shadow.record(cmdKey, ctxKey, outcome.verdict);
1768
- }
1769
-
1770
- const shadow = shadowTag(probe);
1771
-
1772
- // #63 cascade: consult the second layer when the gate fires; `effective` is the
1773
- // ratchet result the returned verdict follows (shadow never overrides)
1774
- const trigger = fallbackTrigger(outcome, state.userRules);
1775
- const cascade = trigger ? await runFallbackCascade(state, env, outcome.verdict, trigger, state.userRules.denyPaths.length > 0, actionLine) : {};
1695
+ // #67 cascade: a confidence-floor demotion, or a classifier fail-closed outcome
1696
+ // (the first layer produced no verdict)
1697
+ const demotion = confidenceDemotion(outcome, state.userRules);
1698
+ const cascade = demotion || outcome.source === "fail-closed"
1699
+ ? await runConfidenceCascade(state, env, demotion ? { verdict: outcome.verdict, reason: outcome.reason } : null, demotion ? { kind: "demotion", confidence: demotion.confidence } : { kind: "fail-closed" }, state.userRules.denyPaths.length > 0, actionLine)
1700
+ : {};
1776
1701
  const effVerdict = cascade.effective?.verdict ?? outcome.verdict;
1777
1702
  const effReason = cascade.effective?.reason ?? outcome.reason;
1778
1703
  const effSource = cascade.effective?.source ?? "classifier";
1779
1704
 
1780
- // #62: record top-level keeps FIRST-layer semantics (grill decision — corpus
1781
- // comparability); the enforced outcome lives in fallback.effective and evaluators
1782
- // must read enforce rows accordingly
1783
- const firstAskDegraded = !env.hasUI && outcome.verdict === "ask";
1784
- const grayRecord = buildRecord({ verdict: firstAskDegraded ? "deny" : outcome.verdict, reason: outcome.reason, source: outcome.source, degraded: firstAskDegraded }, outcome.auditRaw ?? null, shadow);
1705
+ // #62/#67: top-level keeps first-layer semantics (corpus comparability); the applied
1706
+ // verdict lives in fallback.effective (enforce rows). Non-interactive asks of any
1707
+ // origin — native, demoted, escalated — record as their effective deny, the
1708
+ // pre-existing ask-degradation convention. #71 exception: a fail-closed first layer
1709
+ // rescued by an enforcing fallback records the applied verdict at the top level —
1710
+ // a fail-closed layer emits no negative verdict, so its default deny would distort
1711
+ // deny-rate statistics (26 observed rows, 25 actually allowed); shadow rescues keep it.
1712
+ const appliedAskHeadless = !env.hasUI && effVerdict === "ask";
1713
+ const fcRescued = cascade.effective !== undefined && outcome.source === "fail-closed";
1714
+ const grayRecord = buildRecord({ verdict: appliedAskHeadless ? "deny" : fcRescued ? effVerdict : outcome.verdict, reason: fcRescued ? effReason : outcome.reason, source: outcome.source, degraded: appliedAskHeadless }, outcome.auditRaw ?? null);
1715
+ if (cascade.demoted) grayRecord.demoted = true;
1785
1716
  if (cascade.fb) grayRecord.fallback = cascade.fb;
1786
1717
  // #62: an interactive ask defers the append to the handler finalize (ground truth);
1787
- // a headless degraded ask and every other outcome append immediately as before
1718
+ // every other outcome appends immediately as before
1788
1719
  if (effVerdict === "ask" && env.hasUI) {
1789
- return { verdict: "ask", reason: effReason, source: effSource, degraded: false, shadow, ...(state.audit ? { pendingAudit: grayRecord } : {}) };
1720
+ return { verdict: "ask", reason: effReason, source: effSource, degraded: false, ...(state.audit ? { pendingAudit: grayRecord } : {}) };
1790
1721
  }
1791
1722
  state.audit?.append(grayRecord);
1792
- if (effVerdict === "allow") return { verdict: "allow", reason: effReason, source: effSource, degraded: false, shadow };
1793
- if (effVerdict === "deny") return { verdict: "deny", reason: effReason, source: effSource, degraded: false, shadow };
1723
+ if (effVerdict === "allow") return { verdict: "allow", reason: effReason, source: effSource, degraded: false };
1724
+ if (effVerdict === "deny") return { verdict: "deny", reason: effReason, source: effSource, degraded: false };
1794
1725
  // ask:无 UI 降级为 deny(ask 降级,CONTEXT.md 词条)
1795
- return { verdict: "deny", reason: effReason, source: effSource, degraded: !env.hasUI, shadow };
1726
+ return { verdict: "deny", reason: effReason, source: effSource, degraded: true };
1796
1727
  }
1797
1728
 
1798
1729
  // ============================================================================
@@ -1815,7 +1746,7 @@ export interface AutoModeDeps {
1815
1746
  export default function autoMode(pi: ExtensionAPI, deps: AutoModeDeps = {}) {
1816
1747
  pi.registerFlag("auto-mode", { description: "Enable Auto Mode (rules + model classifier gating for tool calls)", type: "boolean", default: true });
1817
1748
  pi.registerFlag("auto-mode-model", { description: "Classifier model as provider/id[:thinking] (pi --model syntax; default: inherit session model)", type: "string" });
1818
- pi.registerFlag("auto-mode-debug", { description: "Notify every verdict incl. allows, with shadow-cache annotation", type: "boolean", default: false });
1749
+ pi.registerFlag("auto-mode-debug", { description: "Notify every verdict incl. allows", type: "boolean", default: false });
1819
1750
 
1820
1751
  let enabled = pi.getFlag("auto-mode") !== false;
1821
1752
  const debug = pi.getFlag("auto-mode-debug") === true || process.env.PI_AUTO_MODE_DEBUG === "1";
@@ -1830,19 +1761,21 @@ export default function autoMode(pi: ExtensionAPI, deps: AutoModeDeps = {}) {
1830
1761
  return { block: true, reason: blockedReason("tamper", r.reason) };
1831
1762
  }
1832
1763
 
1833
- /** Verdict → UI(本扩展唯一的裁决呈现点):按 source × degraded 查模板,文案与
1834
- * 重构前逐字节一致。受保护路径分支的通知永不携带路径明文与 action 行
1764
+ /** Verdict → UI (the extension's single presentation point): presentation keys on
1765
+ * source alone — the degraded context is implicit in the protected-path deny
1766
+ * wording. Byte-identical with the pre-refactor wording; protected-path
1767
+ * notifications never carry path plaintext or the action line
1835
1768
  * (ADR-0002 story 11:通知与 block reason 回流 agent context)。 */
1836
1769
  async function presentVerdict(v: Verdict, action: string, ctx: ExtensionContext): Promise<{ block: true; reason: string } | undefined> {
1837
1770
  if (v.verdict === "allow") {
1838
1771
  // #60 (CONTEXT.md 通知): classifier allows surface via notifyAllows OR
1839
- // debug — exactly one notification either way; the shadow suffix stays
1840
- // debug-only; mechanical passes (rule echo, protected-path confirm) stay
1841
- // debug-only — notifications carry judgment, the audit log carries completeness
1772
+ // debug — exactly one notification either way; mechanical passes
1773
+ // (rule echo, protected-path confirm) stay silent — notifications carry
1774
+ // judgment, the audit log carries completeness
1842
1775
  if (debug) {
1843
1776
  if (v.source === "rule") ctx.ui.notify(`🛡️ allow (rule): ${action}`, "info");
1844
1777
  else if (v.source === "protected-path") ctx.ui.notify("🛡️ allow (protected-path confirm)", "info");
1845
- else ctx.ui.notify(`🛡️ allow (classifier): ${v.reason}\n ${action}${v.shadow ? " " + v.shadow : ""}`, "info");
1778
+ else ctx.ui.notify(`🛡️ allow (classifier): ${v.reason}\n ${action}`, "info");
1846
1779
  } else if (state.userRules.notifyAllows && v.source === "classifier") {
1847
1780
  ctx.ui.notify(`🛡️ allow (classifier): ${v.reason}\n ${action}`, "info");
1848
1781
  }
@@ -1862,7 +1795,7 @@ export default function autoMode(pi: ExtensionAPI, deps: AutoModeDeps = {}) {
1862
1795
  ctx.ui.notify(`🛡️ Auto Mode blocked: ${v.reason}\n ${action}`, "warning");
1863
1796
  return { block: true, reason: blockedReason("rule", v.reason) };
1864
1797
  }
1865
- ctx.ui.notify(`🛡️ Auto Mode blocked: ${v.reason}\n ${action}${debug && v.shadow ? " " + v.shadow : ""}`, "warning");
1798
+ ctx.ui.notify(`🛡️ Auto Mode blocked: ${v.reason}\n ${action}`, "warning");
1866
1799
  return { block: true, reason: blockedReason("classifier", v.reason) };
1867
1800
  }
1868
1801
  // ask → 人工确认;非交互已在管线内降级,能走到这里的必有 UI
@@ -1891,7 +1824,7 @@ export default function autoMode(pi: ExtensionAPI, deps: AutoModeDeps = {}) {
1891
1824
  refreshStatus(ctx);
1892
1825
  }
1893
1826
 
1894
- // session_start:重置影子缓存(会话内存态,#5 定案)+ 重载用户规则(配置改动新会话生效)
1827
+ // session_start:重置级联计数(会话内存态)+ 重载用户规则(配置改动新会话生效)
1895
1828
  // + 重建自保护基线(ADR-0001:受保护文件的会话启动快照)
1896
1829
  pi.on("session_start", async (_event, ctx) => {
1897
1830
  const report = state.reset(ctx.cwd);
@@ -1923,16 +1856,22 @@ export default function autoMode(pi: ExtensionAPI, deps: AutoModeDeps = {}) {
1923
1856
  const denyPathsHint = () => (state.userRules.denyPaths.length > 0 ? `\ndenyPaths: ${state.userRules.denyPaths.length} active` : "");
1924
1857
  /** Status line audit hint (#54): shown only while the sink is active */
1925
1858
  const auditHint = () => (state.audit ? `\naudit: on → ${state.audit.dir}` : "");
1926
- /** Status line fallback hint (#63): shown only while the cascade is configured */
1927
- const fallbackHint = () => (state.userRules.classifierFallbackModel ? `\n${state.fallback.summary(state.userRules.classifierFallbackMode)}` : "");
1859
+ /** Status line cascade hint (#63/#67): shown while the floor or the fallback is configured */
1860
+ const fallbackHint = () => {
1861
+ if (state.userRules.classifierMinConfidence === null && !state.userRules.classifierFallbackModel) return "";
1862
+ const shadowNote = state.userRules.classifierFallbackModel && state.userRules.classifierFallbackMode === "shadow"
1863
+ ? "\nsecond layer is shadow (records only, never applies) — set classifierFallbackMode to \"enforce\" to activate it"
1864
+ : "";
1865
+ return `\n${state.fallback.summary(state.userRules.classifierFallbackMode)}${shadowNote}`;
1866
+ };
1928
1867
 
1929
1868
  pi.registerCommand("automode", {
1930
- description: "Show Auto Mode status and shadow-cache stats, or set it: /automode on|off",
1869
+ description: "Show Auto Mode status, or set it: /automode on|off",
1931
1870
  handler: async (args, ctx) => {
1932
1871
  const arg = args.trim().toLowerCase();
1933
- // 裸调用:只读状态展示,无副作用(含影子缓存统计行)
1872
+ // 裸调用:只读状态展示,无副作用
1934
1873
  if (arg === "") {
1935
- ctx.ui.notify(`${enabled ? "🛡️ Auto Mode: on" : "Auto Mode: off"}\n${state.shadow.summary()}${denyPathsHint()}${auditHint()}${fallbackHint()}\nUsage: /automode on|off${toggleHint()}`, "info");
1874
+ ctx.ui.notify(`${enabled ? "🛡️ Auto Mode: on" : "Auto Mode: off"}\n${denyPathsHint()}${auditHint()}${fallbackHint()}\nUsage: /automode on|off${toggleHint()}`, "info");
1936
1875
  return;
1937
1876
  }
1938
1877
  // 幂等设定:与现值相同不翻转,仅确认
@@ -1943,7 +1882,7 @@ export default function autoMode(pi: ExtensionAPI, deps: AutoModeDeps = {}) {
1943
1882
  const head = next
1944
1883
  ? `🛡️ Auto Mode enabled${changed ? "" : " (unchanged)"}: tool calls adjudicated by rules + classifier`
1945
1884
  : `Auto Mode disabled${changed ? "" : " (unchanged)"}: tool calls execute directly`;
1946
- ctx.ui.notify(`${head}\n${state.shadow.summary()}${fallbackHint()}`, "info");
1885
+ ctx.ui.notify(`${head}\n${fallbackHint()}`, "info");
1947
1886
  return;
1948
1887
  }
1949
1888
  // 未知参数:严格拒绝并列出用法(大小写已归一化)
@@ -2065,7 +2004,7 @@ export default function autoMode(pi: ExtensionAPI, deps: AutoModeDeps = {}) {
2065
2004
  }
2066
2005
  }
2067
2006
 
2068
- // 判定管线(零 UI)→ 呈现(source × degraded 模板)
2007
+ // 判定管线(零 UI)→ 呈现(source 模板)
2069
2008
  const verdict = await adjudicate(state, { toolName: event.toolName, input }, {
2070
2009
  cwd: ctx.cwd,
2071
2010
  hasUI: !!ctx.hasUI,
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-verdict",
3
- "version": "0.10.0",
3
+ "version": "0.12.0",
4
4
  "description": "A minimal permission gate for Pi in the style of Claude Code's auto mode",
5
5
  "author": "Jesset (https://github.com/jesset)",
6
6
  "type": "module",