@ian-pascoe/pi-guardian 0.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright 2026 Ian Pascoe
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,207 @@
1
+ # Pi Guardian
2
+
3
+ `@ian-pascoe/pi-guardian` gates a Pi agent's tool calls before they run. Routine calls run at once; risky ones go to a separate reviewer model, the Guardian, which judges the call's risk against the user's authorization. A fixed Decision Table turns that judgment into an allowed call or a binding Rejection.
4
+
5
+ Requires Node `>=22.19.0` and Pi `>=0.99.0`.
6
+
7
+ > **Guardian is a safety net, not a sandbox.** An allowed call runs with your full privileges, a reviewer model can be wrong or manipulated, and other extensions can still act on their own. Keep backups and review what agents do.
8
+
9
+ ## Install
10
+
11
+ ```bash
12
+ pi install npm:@ian-pascoe/pi-guardian
13
+ # or from this checkout
14
+ pi -e ./packages/pi-guardian/src/index.ts
15
+ ```
16
+
17
+ **Install Guardian last.** Pi runs `tool_call` handlers in extension load order, and an extension loaded after Guardian can rewrite a call's arguments after Guardian reviewed them. Guardian detects this after the call runs (see [Limitations](#limitations)).
18
+
19
+ Guardian is **enabled by default** and reviews with the session's current model unless `model` is set. Without a usable model, every call that needs review is a Review Failure.
20
+
21
+ **Choose a small, fast Guardian model with thinking off.** Every Reviewed Call waits for its review, and a large session model makes that slow and costly: in live use, Claude Opus took 6 to 17 s and about $0.27 per review, while `anthropic/claude-haiku-4-5` with `thinkingLevel: "off"` took about 1.4 s and $0.05. Set both in `/guardian` (Model, thinking level) or in settings, as in the [example](#settings). While `model` is unset, Guardian shows a one-time notice at session start, and `/guardian status` marks the model as inherited from the session.
22
+
23
+ ## How a call is judged
24
+
25
+ Every `tool_call`, including calls a tool issues itself (such as a `codemode` script's calls, which carry `parentToolCallId`), resolves to one Tool Policy:
26
+
27
+ 1. the configured `tools.<name>` setting, if any;
28
+ 2. else the built-in default:
29
+ - `allow`: `read`, `grep`, `find`, `ls`, `codemode`, `tool_search`, `todo`, `web_search`, and Context Management's `context_notes`, `context_history`, and `context_rollover`, which only touch the session's own notes, journal, and handoff;
30
+ - `edit` and `write`: `allow`, unless the target is a **Sensitive Path** or the call shares its assistant message's tool batch with a call that is not allowed without review (Pi may run them in parallel, and that call could replace the target with a link first); either is reviewed;
31
+ - `bash`: `allow` only for a **Safe Command**, otherwise reviewed;
32
+ - `terminal_start`, `terminal_send`, and `powershell`: reviewed;
33
+ 3. else the tool's `readOnlyHint: true` annotation (as reported by `pi.getAllTools()`) allows it, unless it also declares `openWorldHint: true`;
34
+ 4. else the call is reviewed.
35
+
36
+ `deny` blocks the call without a model call. `review` sends it to a Guardian Review. A nested call is judged by its own Tool Policy; the issuing call is shown to the Guardian as context only.
37
+
38
+ While Guardian's settings cannot be read (invalid `guardian` settings or an unreadable session), every call except the built-in `read`, `grep`, `find`, and `ls` is a Review Failure: the unreadable settings may have held `deny` or `review` rules, so neither they nor the built-in defaults apply. The footer shows `guardian: settings error`, and the error is reported when the session starts and in each blocked call's reason.
39
+
40
+ **Sensitive Paths** are judged after resolving `~`, `..`, `@`, `file://`, and symlinks (including dangling ones), on both the lexical and the resolved path:
41
+
42
+ - anything outside the working directory, and **every** path when the working directory is the home directory, an ancestor of it, or `/`;
43
+ - persistence and credential locations under the home directory wherever the workspace is: shell startup files (`.bashrc`, `.bash_profile`, `.profile`, `.zshrc`, `.zprofile`, `.zshenv`, and similar), `.ssh`, `.gnupg`, `.aws`, `.azure`, `.config` (including fish and systemd user units), `.local/bin`, `Library/LaunchAgents`, `.gitconfig`, `.git-credentials`, `.npmrc`, `.yarnrc`, `.pypirc`, `.netrc`, `.docker`, and `.kube`; system crontabs are outside any workspace;
44
+ - inside the workspace, any case: `.git`, `.pi`, `.agents`, `.env*` (including `.envrc`), `.husky`, `.github/workflows`, `.vscode`, and context files named `AGENTS.md`, `AGENTS.override.md`, or `CLAUDE.md` at any depth;
45
+ - Pi's agent directory (`getAgentDir()`), the session directory, and every resource Pi loaded into the session: context files, Skills (their whole directory), prompt templates, system prompt files, and extensions;
46
+ - an existing file with more than one hard link, since editing it in place changes the other paths too;
47
+ - on Windows, a path with backslashes or a drive letter, which Guardian does not judge.
48
+
49
+ Pi configuration, context files, and loaded resources are sensitive because changing them can weaken Guardian or rewrite the instructions it trusts.
50
+
51
+ A **Safe Command** is one simple command of literal words: no pipes, redirection, `;`, `&&`, `||`, `&`, subshells, grouping, command or process substitution, backticks, newlines or other control characters, variable, tilde, brace or history expansion, globbing, comments, escapes, or environment assignments, and a bare program name (no path). The built-in programs are `ls`, `pwd`, `cat`, `head`, `tail`, `wc`, `echo`, `stat`, `du`, `df`, `basename`, `dirname`, `realpath`, `which`, `whoami`, `uname`, `grep`, `rg` (without `--pre`, `--pre-glob`, or `--hostname-bin`), `find` (without `-exec`, `-execdir`, `-ok`, `-okdir`, `-delete`, `-fprint`, `-fprint0`, `-fprintf`, or `-fls`), and `git status`, `git log`, `git diff`, `git show` (without `--output`, `--ext-diff`, or `--textconv`), `git branch` (listing flags only), and `git rev-parse`, with no git options before the subcommand. When in doubt the command is reviewed. `safeCommands` adds literal prefixes: `"npm test"` allows `npm test` and `npm test -- --run`, but the rest of the command must still be literal words. An entry must itself be literal words starting with a bare program name; one such as `"./gradlew test"` could never match and is rejected as invalid settings. Avoid script runners (`npm test`, `pnpm lint`, `make`) in `safeCommands`: see [Limitations](#limitations).
52
+
53
+ ## Guardian Review
54
+
55
+ Each review is one stateless model completion without tools ([ADR-0001](docs/adr/0001-review-with-one-stateless-call.md)):
56
+
57
+ - **System prompt**: the built-in policy (evidence handling, User Authorization scoring, risk taxonomy, Pi tools), your Security Policy (`policy`), and the output contract.
58
+ - **One user message** of text blocks: context files, then evidence entries in conversation order, then the **Reviewed Call** — the tool, why it was reviewed (such as its Sensitive Path), the SHA-256 and full text of its exact arguments, the working directory, which agent issued it, and for nested calls the issuing call's tool and arguments.
59
+
60
+ The request is append-only: each review's blocks extend the previous review's, so provider prompt caches apply across reviews, even after the history outgrows the evidence budget (see below). Guardian never changes the Guarded Agent's system prompt, tools, or messages.
61
+
62
+ Pi runs a response's `tool_call` handlers one call at a time, even for parallel tool calls. Guardian therefore starts the reviews of a response's tool calls as soon as the response ends, at most four at once, and reuses each result when its call reaches `tool_call` with the same tool and arguments, so parallel calls are reviewed concurrently. It skips calls to unknown tools and calls whose arguments fail the tool's schema, which Pi never runs. A review started for a call that never arrives, or whose arguments changed before Guardian saw them, is recorded as `unused`.
63
+
64
+ Nested calls are different: a tool that issues calls concurrently, such as a `codemode` script's `Promise.all`, runs Guardian's `tool_call` handler for each at once, so their reviews run concurrently. Their **Allow once** dialogs are queued and shown one at a time.
65
+
66
+ ### Evidence and trust
67
+
68
+ **Trusted Evidence** can establish User Authorization: messages the user typed, context files, and recorded **User Overrides**. Guardian reads context files from Pi's resource loader (`getAgentsFiles()`), not from the system prompt, so a tool's prompt snippet cannot forge them. Pi loads `AGENTS.md`-style files from the working directory and its ancestors whether or not the project is trusted, so only the global file in Pi's agent directory and, in a trusted project, the others are Trusted Evidence; an untrusted project's context files are included as untrusted evidence.
69
+
70
+ Everything else is labeled **UNTRUSTED**: tool results, assistant text and reasoning, extension and summary messages, user messages an extension sent (`sendUserMessage`), the body of a Skill expanded by `/skill:` (the text the user typed after it stays trusted), and in Child Agent and Advisor sessions every user message, since it comes from another agent. A Child Agent's or Advisor's reviews instead include the **root user's** typed messages (origin `rootUser`) as Trusted Evidence, interleaved by time, while the root session runs Guardian in the same process. Each entry carries its message as JSON, so content cannot forge an evidence label.
71
+
72
+ Guardian recognizes an extension-sent message from Pi's `input` event (`source: "extension"`) and records a `pi-guardian-extension-message` session entry, so the label survives reloads. Not distinguishable, and therefore trusted like typed text: prompt templates expanded from `/name`, an extension message that another extension's `input` handler rewrote, and task messages from subagent systems other than Minimal Subagents.
73
+
74
+ The evidence budget (`evidenceBudgetTokens`) defaults to `auto`: a quarter of the Guardian model's context window, at most 32,000 tokens. Pi's chars/4 estimate undercounts Guardian's JSON-heavy requests (about 1.3× on Claude Haiku and 1.7× on Claude Opus), so Guardian scales it by a per-model factor: 1.5 until the model's provider reports usage, then the reported prompt tokens (input plus cache reads and writes) over the estimate, rounded up to a quarter and changed only by a clear margin so it stays stable.
75
+
76
+ All Trusted Evidence is always kept. The rest is a window of every entry from an anchor onward, so successive reviews share their prefix. When the evidence outgrows the budget, the anchor jumps forward past at least half a budget of the oldest untrusted entries, replaced by an omission note; the provider's cache breaks only at those jumps, at most once per half-budget of growth. The anchor is a pure function of the history, so a reload or branch switch finds the same window. Each untrusted text is capped near 2,000 tokens with a marker; Trusted Evidence is shortened, every entry with the same cap, only when it alone exceeds the budget.
77
+
78
+ The Reviewed Call is never shortened, since a cut could hide the harmful part of a call. When it does not fit the Guardian model's context window beside the policy and room for the reply, the review is a Review Failure, and the evidence budget shrinks to what is left beside it.
79
+
80
+ ### Decision Table
81
+
82
+ The model returns `{"risk_level", "user_authorization", "rationale"}` (fenced JSON is tolerated; a reply with two differing assessments is malformed). Unless `verbose` is on, the output contract asks for the rationale only at `high` or `critical` risk, the only levels the Decision Table can reject, which saves output tokens and latency; a Rejection whose assessment omits it uses the fixed reason "The Guardian gave no specific rationale." The contract is part of the system prompt, so it changes, and the cache restarts, only when `verbose` does. The Outcome is fixed:
83
+
84
+ | Risk Level | `unknown` | `low` | `medium` | `high` |
85
+ | ---------- | --------- | -------- | -------- | -------- |
86
+ | `low` | allowed | allowed | allowed | allowed |
87
+ | `medium` | allowed | allowed | allowed | allowed |
88
+ | `high` | rejected | rejected | allowed | allowed |
89
+ | `critical` | rejected | rejected | rejected | rejected |
90
+
91
+ ### Rejection
92
+
93
+ A Rejection blocks the call and tells the agent:
94
+
95
+ ```text
96
+ This action was rejected due to unacceptable risk.
97
+ Risk: high. Authorization: low.
98
+ Reason: <rationale>
99
+ Do not attempt to achieve the same outcome through a workaround, indirect execution, or variations of this call, and do not retry it. Explain the risk to the user and ask whether they want to proceed; continue only with a materially safer alternative or after the user explicitly approves this action.
100
+ ```
101
+
102
+ The user sees a warning with the tool, risk, and rationale. With `onDeny: "ask"` and an interactive UI, a dialog offers **Allow once**: a User Override. A call too long to show in the dialog (over 2,000 characters) offers **View full call**, which opens the whole call read-only in Pi's editor; **Allow once** appears only after that. Otherwise the user can authorize the action in conversation, which the next review weighs as Trusted Evidence.
103
+
104
+ **Rejection Streak**: after `maxConsecutiveRejections` (default 3; 0 disables) consecutive blocked Reviewed Calls in one request, the blocking result also asks Pi to end the turn. Blocked Review Failures count toward the streak too, so a broken Guardian cannot keep a headless agent retrying. Any allowed Reviewed Call that actually runs, including a User Override, resets it, and so does each new prompt; a call Guardian allowed but another extension then blocked does not. Pi ends the turn only when every call in the tool batch asked to stop.
105
+
106
+ ### Review Failure
107
+
108
+ No model resolved, no credentials, a provider error, a timeout (`reviewTimeoutMs`), malformed output twice in a row, a Reviewed Call too large to review in full, or unreadable Guardian settings never allow a call. With an interactive UI, a dialog offers **Allow once** (a User Override) or **Block**, after **View full call** for a long call; without one, the call is blocked with the reason and a pointer to the troubleshooting Skill. Aborting the agent's turn aborts its reviews; an aborted review blocks its call with an "aborted" reason and is not a Review Failure.
109
+
110
+ A reply without exactly one valid assessment gets one corrective retry within the same deadline: the first request plus the bad reply and a user message restating the JSON contract, so the first request's cached prefix is unchanged. Pi's provider-neutral API has no JSON mode or forced tool call, so the contract is enforced by parsing. A retried review is marked `retried` in its audit entry.
111
+
112
+ ### Audit
113
+
114
+ Every Guardian Review appends a `pi-guardian-review` session entry, which never reaches the model: tool, call ID, parent call ID, arguments (bounded) and their full SHA-256, Risk Level, User Authorization, outcome, rationale, failure, User Override, whether an allowed call actually ran, model, duration, token usage and cost summed over a retry, whether the review was retried, and argument drift. `/guardian status` derives its totals from the selected branch's entries.
115
+
116
+ The transcript shows only the reviews that need attention: Rejections, Review Failures, aborted reviews, User Overrides, and argument drift. Allowed and unused reviews are still recorded and counted, but render nothing unless `verbose` is on. Guardian shows no notice for allowed calls; while a review runs, the footer shows `guardian: reviewing <tool>`.
117
+
118
+ User Overrides return to later reviews as Trusted Evidence in structured form. The user's decision is trusted but the arguments were written by the agent, so they are a marked field, and the Guardian's rationale is left out:
119
+
120
+ ```json
121
+ {
122
+ "userOverride": {
123
+ "decision": "The user interactively allowed one call after a Rejection.",
124
+ "scope": "This authorizes only that exact call: the same tool with arguments of the same SHA-256. …",
125
+ "tool": "bash",
126
+ "argumentsSha256": "…",
127
+ "agentAuthoredArguments": "{\"command\":\"rm -rf dist\"}",
128
+ "agentAuthoredArgumentsShortened": false
129
+ }
130
+ }
131
+ ```
132
+
133
+ ## Commands
134
+
135
+ ```text
136
+ /guardian
137
+ /guardian status
138
+ /guardian on|off [--global|--project]
139
+ /guardian tool <name> <allow|review|deny|default|inherit> [--global|--project]
140
+ /guardian policy [--global|--project]
141
+ /guardian inherit [key] [--global|--project]
142
+ /guardian set <key> <JSON> [--global|--project]
143
+ ```
144
+
145
+ Without a flag, changes go to the session. In the interactive TUI, `/guardian` opens a settings menu like `/advisor`'s: a Scope row (session, trusted project, global), cycling rows for `enabled`, `thinkingLevel`, `onDeny`, and `verbose` that show the selected scope's own value or what it inherits, a model picker, a per-tool Tool Policy list, the Security Policy in Pi's editor, and typed values for the rest. Closing it records one status entry listing the changes it applied. Elsewhere `/guardian` records a status entry.
146
+
147
+ `/guardian status` records the effective settings with their sources, whether the session follows a root session, review counts (allowed, rejected, failed, aborted, overrides, argument drift), total review cost, and the last failure. The footer shows `guardian` while idle, `guardian: reviewing <tool>` during reviews, and nothing while disabled.
148
+
149
+ ## Settings
150
+
151
+ Settings live under `guardian` in Pi's global and trusted-project `settings.json`, plus session overrides. Precedence is default < global < trusted project < session; a trusted project may weaken Guardian ([ADR-0002](docs/adr/0002-trusted-projects-may-weaken-guardian.md)), while an untrusted project's settings are ignored.
152
+
153
+ | Key | Default | Meaning |
154
+ | -------------------------- | --------- | ---------------------------------------------------------------------------------------------------- |
155
+ | `enabled` | `true` | Gate tool calls. While disabled Guardian does nothing, including `deny` Tool Policies. |
156
+ | `model` | session | Guardian model as `provider/id`; absent follows the session's current model. Prefer a small one. |
157
+ | `thinkingLevel` | `"low"` | `off` … `max`, clamped to the model. |
158
+ | `tools` | `{}` | Tool Policies by tool name: `allow`, `review`, `deny`, or `null` to reset an inherited entry. |
159
+ | `safeCommands` | `[]` | Extra Safe Command prefixes; merged across scopes as a union. |
160
+ | `policy` | `""` | Security Policy: trusted destinations, forbidden actions, and other rules added to the built-in one. |
161
+ | `reviewTimeoutMs` | `60000` | Deadline for each review; a timeout is a Review Failure. |
162
+ | `evidenceBudgetTokens` | `"auto"` | Positive integer or `auto`. |
163
+ | `onDeny` | `"block"` | `block`, or `ask` to offer Allow once on a Rejection in interactive sessions. |
164
+ | `maxConsecutiveRejections` | `3` | Rejection Streak that ends the turn; `0` never ends it. |
165
+ | `verbose` | `false` | Ask for a rationale on every review and show allowed reviews in the transcript, for debugging. |
166
+
167
+ `tools` merges entry by entry across scopes: a higher scope adds or replaces entries, and `null` removes a lower scope's entry so the built-in default applies again. `set tools <JSON>` replaces that scope's whole map; `tool <name> <value>` changes one entry (`default` writes `null`, `inherit` removes the scope's entry). A configured Tool Policy overrides the built-in default, so `{"edit": "allow"}` also allows edits to Sensitive Paths.
168
+
169
+ ```json
170
+ {
171
+ "guardian": {
172
+ "model": "anthropic/claude-haiku-4-5",
173
+ "thinkingLevel": "off",
174
+ "tools": { "mcp__github__create_issue": "allow", "terminal_send": "deny" },
175
+ "safeCommands": ["tree", "file"],
176
+ "policy": "Pushing to github.com/acme/* is trusted. Never touch the production database.",
177
+ "onDeny": "ask"
178
+ }
179
+ }
180
+ ```
181
+
182
+ ## Child Agents and Advisors
183
+
184
+ Guardian loads in every session that loads it, including Minimal Subagents Child Agent sessions (print mode, no UI) and Advisor sessions. A Child Agent (detected by Minimal Subagents' `minimal-subagents.identity` entry) or an Advisor (pi-advisor's `pi-advisor-role` entry) follows its root session's effective settings live while that root runs Guardian in the same process; otherwise it falls back to its own global and project settings. Its task and other user messages come from another agent, so they are untrusted, and without UI its Review Failures and Rejections block. Settings changes from inside such a session are refused; change them in the root session.
185
+
186
+ Only these two kinds of delegated session are detected. A child session of any other subagent system is treated as a main session: its task message, sent by another agent, appears user-typed and counts as Trusted Evidence.
187
+
188
+ ## Limitations
189
+
190
+ - **Argument drift**: Guardian reviews the arguments its `tool_call` handler sees. An extension loaded after Guardian can still change them. Pi emits `tool_execution_start` before `tool_call` handlers run, so Guardian compares the reviewed arguments with the `tool_result` event's arguments instead and warns after the call has run, marking the review entry with `argumentDrift`. Install Guardian last.
191
+ - A Child Agent or Advisor sees the root user's typed messages only while the root session runs Guardian in the same process; otherwise its evidence holds only its own conversation, none of it trusted.
192
+ - Reviews cannot inspect files or run read-only checks (ADR-0001), so the policy leans conservative when evidence is missing.
193
+ - `git` read-only subcommands still honor repository configuration such as `core.fsmonitor` or `diff.external`; edits to `.git` are Sensitive Paths and therefore reviewed.
194
+ - Annotations and Safe Commands are trusted as declared; a tool that lies about `readOnlyHint` or `openWorldHint` runs without review unless you configure it.
195
+ - **Exfiltration through reads**: a read can still send data out when it reaches the network. `web_fetch` is therefore reviewed by default, since its URL can carry workspace contents to any host, and so is any tool annotated both read-only and open-world. `web_search` stays allowed: its query goes only to the configured search provider. Allow `web_fetch` with `tools.web_fetch: "allow"` only if you accept that risk.
196
+ - `safeCommands` cannot remove a lower scope's entries.
197
+ - **Script runners run unreviewed code**: a Safe Command such as `npm test`, `pnpm lint`, or `make` executes whatever the workspace's `package.json` scripts, test files, and tool configuration say, and ordinary workspace edits to those files are not reviewed. Do not add script runners to `safeCommands` unless you accept that an agent can run arbitrary code through them. Likewise, extensions that act after edits, such as pi-formatter running formatters with workspace configuration, can execute code that Guardian never reviews.
198
+ - An ordinary edit that shares a tool batch with a reviewed call is reviewed only for the assistant message's own calls; concurrent nested calls of a `codemode` script are judged one by one.
199
+ - Hard links are detected only on existing files; Guardian cannot see a link a concurrent process creates after its check.
200
+
201
+ ## Attribution
202
+
203
+ Guardian's built-in policy is adapted from the Guardian prompts of [OpenAI Codex](https://github.com/openai/codex) (`codex-rs/prompts/templates/guardian/`), Copyright 2025 OpenAI, licensed under the [Apache License, Version 2.0](http://www.apache.org/licenses/LICENSE-2.0). The adaptation (rewritten for Pi's tools, Trusted Evidence, a single stateless review, and an external Decision Table) is in [`src/guardian-prompt.ts`](src/guardian-prompt.ts). The Rejection wording follows Codex's.
204
+
205
+ ## Troubleshooting
206
+
207
+ Run `/skill:pi-guardian`, or see [`skills/pi-guardian/SKILL.md`](skills/pi-guardian/SKILL.md).
package/package.json ADDED
@@ -0,0 +1,59 @@
1
+ {
2
+ "name": "@ian-pascoe/pi-guardian",
3
+ "version": "0.0.0",
4
+ "private": false,
5
+ "description": "Model-reviewed gating of risky Pi agent tool calls before they run.",
6
+ "keywords": [
7
+ "guardian",
8
+ "pi",
9
+ "pi-extension",
10
+ "pi-package",
11
+ "security"
12
+ ],
13
+ "homepage": "https://github.com/ian-pascoe/pi-extensions/tree/main/packages/pi-guardian",
14
+ "bugs": {
15
+ "url": "https://github.com/ian-pascoe/pi-extensions/issues"
16
+ },
17
+ "license": "MIT",
18
+ "author": "Ian Pascoe <ian.g.pascoe@gmail.com>",
19
+ "repository": {
20
+ "type": "git",
21
+ "url": "git+https://github.com/ian-pascoe/pi-extensions.git",
22
+ "directory": "packages/pi-guardian"
23
+ },
24
+ "files": [
25
+ "src",
26
+ "skills",
27
+ "README.md",
28
+ "LICENSE"
29
+ ],
30
+ "type": "module",
31
+ "publishConfig": {
32
+ "access": "public",
33
+ "provenance": true
34
+ },
35
+ "scripts": {
36
+ "test": "vitest run",
37
+ "typecheck": "tsc --noEmit -p tsconfig.json"
38
+ },
39
+ "dependencies": {
40
+ "@ian-pascoe/pi-utils": "^0.3.1"
41
+ },
42
+ "peerDependencies": {
43
+ "@earendil-works/pi-ai": ">=0.99.0",
44
+ "@earendil-works/pi-coding-agent": ">=0.99.0",
45
+ "@earendil-works/pi-tui": ">=0.99.0",
46
+ "typebox": "*"
47
+ },
48
+ "engines": {
49
+ "node": ">=22.19.0"
50
+ },
51
+ "pi": {
52
+ "extensions": [
53
+ "./src/index.ts"
54
+ ],
55
+ "skills": [
56
+ "./skills"
57
+ ]
58
+ }
59
+ }
@@ -0,0 +1,24 @@
1
+ ---
2
+ name: pi-guardian
3
+ description: Diagnose Pi Guardian when tool calls are unexpectedly blocked or allowed, Guardian Reviews fail or time out, settings seem ineffective, Child Agents or Advisors behave differently, or reviews cost too much.
4
+ license: MIT
5
+ disable-model-invocation: true
6
+ ---
7
+
8
+ # Pi Guardian
9
+
10
+ Use this sequence for a live Guardian problem:
11
+
12
+ 1. Run `/guardian status`. Record the state, each setting and its source, the review counts (allowed, rejected, failed, overrides, argument drift), total cost, and the last failure. Treat unknown cost as unknown, not zero. Expand the newest `pi-guardian-review` entries in the transcript for the tool, scores, rationale, failure, model, and duration of each Guardian Review. Allowed reviews are hidden from the transcript, and low- and medium-risk reviews carry no rationale, unless `verbose` is on; turn it on (`/guardian set verbose true`) while diagnosing, then off again.
13
+ 2. If the state is `error`, read the message: invalid `guardian` settings in the global or trusted-project `settings.json`, an invalid session override, or an unreadable session. Correct it; until then every call except the built-in `read`, `grep`, `find`, and `ls` is a Review Failure, including calls the defaults would allow. A `safeCommands` entry must be literal words starting with a bare program name (not `./gradlew`).
14
+ 3. If a call was blocked without a Guardian Review, check its Tool Policy. A `deny` Tool Policy blocks without a model call. Resolution order: `tools.<name>` setting, then the built-in default (read-only tools and the Context Management tools `context_notes`, `context_history`, and `context_rollover` allowed; `edit`/`write` allowed except for Sensitive Paths and for calls sharing a tool batch with a call that needs review; `bash` allowed only for Safe Commands), then the tool's `readOnlyHint` annotation (unless it also declares `openWorldHint`), then review. `web_fetch` is reviewed by default because a URL can carry data out; allow it with `/guardian tool web_fetch allow` only if that is acceptable. Change one entry with `/guardian tool <name> <allow|review|deny|default|inherit>` at the intended scope (`--global`, `--project`, or the session by default).
15
+ 4. If a `bash` command you consider safe is reviewed, it is not a Safe Command: pipes, redirection, chaining, substitutions, variables, globs, `~`, quotes left open, or programs outside the built-in list all send it to review. Add a literal command prefix to `safeCommands` (for example `["npm test"]`); entries match leading words and the rest of the command must still be literal.
16
+ 5. If an edit is reviewed, the review entry's Reviewed Call says why. Usually its path is a Sensitive Path: outside the working directory after resolving `~`, `..`, and symlinks (everything, when the working directory is the home directory or an ancestor of it); a persistence or credential location under the home directory such as `.bashrc`, `.ssh`, `.config`, or `.local/bin`; inside `.git`, `.pi`, `.agents`, `.husky`, `.github/workflows`, `.vscode`, or a `.env*` file; a context file such as `AGENTS.md`; a resource Pi loaded; a file with more than one hard link; or under Pi's agent directory or session directory. Otherwise it shared a tool batch with a call that needed review.
17
+ 6. For Review Failures, read the failure: no model resolves (set `model` to `provider/id` or select a session model), no credentials for that model, a provider error, a timeout (`reviewTimeoutMs`), malformed output (the model did not return exactly one assessment JSON, even after one corrective retry; choose a stronger model or raise `thinkingLevel`), or a call too large to review in full (choose a model with a larger context window; Guardian never shortens the Reviewed Call). With UI, Guardian asks whether to allow once; without UI it blocks.
18
+ 7. For a Rejection you disagree with, check the rationale against Trusted Evidence: only messages the user typed, context files from the global agent directory or a trusted project, and recorded User Overrides establish User Authorization. Tool output, assistant text, Skill bodies, messages an extension sent, an untrusted project's context files, and a Child Agent's task cannot. Tell the agent in your own message that you authorize the exact action, add rules to the Security Policy (`/guardian policy`), or set `onDeny` to `ask` to allow once interactively.
19
+ 8. If the agent stopped after Rejections, the Rejection Streak reached `maxConsecutiveRejections` (default 3; 0 disables). Blocked Review Failures count too. Pi ends the turn only when every call in the tool batch asked to stop.
20
+ 9. If a Child Agent or Advisor behaves differently, check `/guardian status` in that session for `follows root`: Child Agents and Advisors follow the root session's effective settings while the root runs Guardian in the same process, and fall back to their own settings otherwise. Their user messages come from another agent and are untrusted; the root user's typed messages are Trusted Evidence for their calls while the root runs Guardian in the same process. They have no UI, so Review Failures block.
21
+ 10. If `/guardian status` reports argument drift, an extension loaded after Guardian changed the call's arguments after review. Load Guardian last.
22
+ 11. If reviews cost too much, compare the total review cost with the session cost. Allow more read-only tools or Safe Commands, lower `evidenceBudgetTokens`, choose a cheaper `model`, or lower `thinkingLevel`. With `model` unset, reviews use the session model; a small, fast model such as `anthropic/claude-haiku-4-5` with `thinkingLevel` `off` is usually far cheaper and faster. Check the review entries' token usage: if cache reads stay near the system prompt's size across reviews, report it, since evidence is meant to keep a shared prefix.
23
+
24
+ Finish when `/guardian status` shows the intended settings and the problematic call is allowed, reviewed, or blocked as its Tool Policy and Decision Table say. If a model or provider remains unavailable, report the exact failure instead of disabling Guardian.
@@ -0,0 +1,118 @@
1
+ import { isDeepStrictEqual } from "node:util";
2
+ import { Type, type Static } from "typebox";
3
+ import { Value } from "typebox/value";
4
+
5
+ export const riskLevelSchema = Type.Union([
6
+ Type.Literal("low"),
7
+ Type.Literal("medium"),
8
+ Type.Literal("high"),
9
+ Type.Literal("critical"),
10
+ ]);
11
+ export type RiskLevel = Static<typeof riskLevelSchema>;
12
+ export const userAuthorizationSchema = Type.Union([
13
+ Type.Literal("unknown"),
14
+ Type.Literal("low"),
15
+ Type.Literal("medium"),
16
+ Type.Literal("high"),
17
+ ]);
18
+ export type UserAuthorization = Static<typeof userAuthorizationSchema>;
19
+
20
+ /**
21
+ * The Guardian's required answer; extra fields are tolerated and ignored. The rationale is
22
+ * optional because, unless `verbose` is on, the Guardian gives one only for `high` or `critical`.
23
+ */
24
+ export const assessmentSchema = Type.Object({
25
+ risk_level: riskLevelSchema,
26
+ user_authorization: userAuthorizationSchema,
27
+ rationale: Type.Optional(Type.String()),
28
+ });
29
+ /** One Guardian Review's assessment. */
30
+ export interface Assessment {
31
+ risk: RiskLevel;
32
+ authorization: UserAuthorization;
33
+ rationale: string;
34
+ }
35
+
36
+ /** Outcome of a Guardian Review, derived only by the Decision Table. */
37
+ export type Outcome = "allowed" | "rejected";
38
+
39
+ /**
40
+ * Decision Table: `low` and `medium` risk are allowed; `high` risk is allowed only with at least
41
+ * `medium` User Authorization; `critical` risk is always rejected.
42
+ */
43
+ export function decide(risk: RiskLevel, authorization: UserAuthorization): Outcome {
44
+ if (risk === "low" || risk === "medium") return "allowed";
45
+ if (risk === "high" && (authorization === "medium" || authorization === "high")) return "allowed";
46
+ return "rejected";
47
+ }
48
+
49
+ /** Every top-level `{…}` span in a reply, skipping braces inside JSON strings. */
50
+ function objectSpans(text: string): string[] {
51
+ const spans: string[] = [];
52
+ let depth = 0;
53
+ let start = -1;
54
+ let inString = false;
55
+ let escaped = false;
56
+ // Every structural character is ASCII, so UTF-16 indexes are safe here.
57
+ for (let index = 0; index < text.length; index++) {
58
+ const character = text[index];
59
+ if (inString) {
60
+ if (escaped) escaped = false;
61
+ else if (character === "\\") escaped = true;
62
+ else if (character === '"') inString = false;
63
+ continue;
64
+ }
65
+ if (character === '"' && depth > 0) inString = true;
66
+ else if (character === "{") {
67
+ if (depth++ === 0) start = index;
68
+ } else if (character === "}" && depth > 0 && --depth === 0)
69
+ spans.push(text.slice(start, index + 1));
70
+ }
71
+ return spans;
72
+ }
73
+
74
+ /** Valid assessments in a reply: the whole text, or each top-level JSON object in it. */
75
+ function assessments(text: string): Assessment[] {
76
+ const found: Assessment[] = [];
77
+ for (const candidate of [text.trim(), ...objectSpans(text)]) {
78
+ let parsed: unknown;
79
+ try {
80
+ parsed = JSON.parse(candidate);
81
+ } catch {
82
+ continue;
83
+ }
84
+ if (!Value.Check(assessmentSchema, parsed)) continue;
85
+ found.push({
86
+ risk: parsed.risk_level,
87
+ authorization: parsed.user_authorization,
88
+ rationale: parsed.rationale?.trim() ?? "",
89
+ });
90
+ }
91
+ return found;
92
+ }
93
+
94
+ /**
95
+ * Parse a Guardian reply; throws a Review Failure message unless it holds exactly one valid
96
+ * assessment. Repeating the same assessment is tolerated; differing assessments are ambiguous.
97
+ */
98
+ export function parseAssessment(text: string): Assessment {
99
+ const [first, ...rest] = assessments(text);
100
+ if (!first)
101
+ throw new Error("Guardian returned malformed output (expected the assessment JSON object)");
102
+ if (rest.some((other) => !isDeepStrictEqual(other, first)))
103
+ throw new Error(
104
+ "Guardian returned malformed output (more than one differing assessment JSON object)",
105
+ );
106
+ return first;
107
+ }
108
+
109
+ /** Codex-style feedback for a Rejection, as reported to the Guarded Agent. */
110
+ export function rejectionReason(assessment: Assessment): string {
111
+ const rationale = assessment.rationale || "The Guardian gave no specific rationale.";
112
+ return [
113
+ "This action was rejected due to unacceptable risk.",
114
+ `Risk: ${assessment.risk}. Authorization: ${assessment.authorization}.`,
115
+ `Reason: ${rationale}`,
116
+ "Do not attempt to achieve the same outcome through a workaround, indirect execution, or variations of this call, and do not retry it. Explain the risk to the user and ask whether they want to proceed; continue only with a materially safer alternative or after the user explicitly approves this action.",
117
+ ].join("\n");
118
+ }
@@ -0,0 +1,169 @@
1
+ import type { SessionEntry } from "@earendil-works/pi-coding-agent";
2
+ import { Type, type Static } from "typebox";
3
+ import { Value } from "typebox/value";
4
+ import { riskLevelSchema, userAuthorizationSchema } from "./guardian-assessment.js";
5
+ import { argumentsHash, type RecordedOverride, type ToolInput } from "./guardian-evidence.js";
6
+
7
+ /** Custom session entry recording one Guardian Review; never part of model context. */
8
+ export const reviewEntryType = "pi-guardian-review";
9
+
10
+ const nullableNumber = Type.Union([Type.Number(), Type.Null()]);
11
+ const nullableString = Type.Union([Type.String(), Type.Null()]);
12
+
13
+ /**
14
+ * How a Guardian Review ended for its call: `allowed` and `rejected` come from the Decision Table,
15
+ * `failed` is a Review Failure, `aborted` followed the turn's abort, and `unused` is a review
16
+ * started ahead of a call that never reached Guardian's `tool_call` handler.
17
+ */
18
+ export const reviewOutcomeSchema = Type.Union([
19
+ Type.Literal("allowed"),
20
+ Type.Literal("rejected"),
21
+ Type.Literal("failed"),
22
+ Type.Literal("aborted"),
23
+ Type.Literal("unused"),
24
+ ]);
25
+ export type ReviewOutcome = Static<typeof reviewOutcomeSchema>;
26
+
27
+ export const reviewEntrySchema = Type.Object({
28
+ version: Type.Literal(1),
29
+ toolName: Type.String(),
30
+ toolCallId: Type.String(),
31
+ parentToolCallId: nullableString,
32
+ /** The Reviewed Call's arguments as reviewed, bounded for the journal. */
33
+ arguments: Type.String(),
34
+ /** SHA-256 of the full serialized arguments, identifying the exact call. */
35
+ argumentsSha256: Type.String(),
36
+ risk: Type.Union([riskLevelSchema, Type.Null()]),
37
+ authorization: Type.Union([userAuthorizationSchema, Type.Null()]),
38
+ outcome: reviewOutcomeSchema,
39
+ rationale: nullableString,
40
+ failure: nullableString,
41
+ /** True when the call ran because the user allowed it interactively. */
42
+ userOverride: Type.Boolean(),
43
+ /** True when the call was blocked. */
44
+ blocked: Type.Boolean(),
45
+ model: nullableString,
46
+ durationMs: Type.Number(),
47
+ usage: Type.Union([
48
+ Type.Object({
49
+ input: Type.Number(),
50
+ output: Type.Number(),
51
+ cacheRead: Type.Number(),
52
+ cacheWrite: Type.Number(),
53
+ total: Type.Number(),
54
+ }),
55
+ Type.Null(),
56
+ ]),
57
+ cost: nullableNumber,
58
+ /**
59
+ * For allowed calls: whether the call ran. Another extension's `tool_call` handler can still
60
+ * block a call Guardian allowed.
61
+ */
62
+ executed: Type.Optional(Type.Boolean()),
63
+ /** Set when the executed arguments differed from the reviewed ones. */
64
+ argumentDrift: Type.Optional(Type.Boolean()),
65
+ /** Set when a malformed reply was followed by one corrective retry. */
66
+ retried: Type.Optional(Type.Boolean()),
67
+ });
68
+ export type ReviewEntry = Static<typeof reviewEntrySchema>;
69
+
70
+ /** Totals over the selected branch's Guardian Reviews. */
71
+ export const reviewTotalsSchema = Type.Object({
72
+ reviews: Type.Number(),
73
+ allowed: Type.Number(),
74
+ rejected: Type.Number(),
75
+ failed: Type.Number(),
76
+ aborted: Type.Number(),
77
+ overrides: Type.Number(),
78
+ drift: Type.Number(),
79
+ /** Total cost in dollars; `null` when any review's cost was unknown. */
80
+ cost: nullableNumber,
81
+ lastError: nullableString,
82
+ });
83
+ export type ReviewTotals = Static<typeof reviewTotalsSchema>;
84
+
85
+ /** Valid review entries on a branch, in order, with their timestamps. */
86
+ function reviewEntries(branch: readonly SessionEntry[]) {
87
+ return branch.flatMap((entry) =>
88
+ entry.type === "custom" &&
89
+ entry.customType === reviewEntryType &&
90
+ Value.Check(reviewEntrySchema, entry.data)
91
+ ? [{ data: entry.data, timestamp: Date.parse(entry.timestamp) }]
92
+ : [],
93
+ );
94
+ }
95
+
96
+ /** Derive status totals from the branch's review entries. */
97
+ export function reviewTotals(branch: readonly SessionEntry[]): ReviewTotals {
98
+ const totals: ReviewTotals = {
99
+ reviews: 0,
100
+ allowed: 0,
101
+ rejected: 0,
102
+ failed: 0,
103
+ aborted: 0,
104
+ overrides: 0,
105
+ drift: 0,
106
+ cost: 0,
107
+ lastError: null,
108
+ };
109
+ for (const { data } of reviewEntries(branch)) {
110
+ totals.reviews++;
111
+ if (data.outcome === "allowed") totals.allowed++;
112
+ if (data.outcome === "rejected") totals.rejected++;
113
+ if (data.outcome === "failed") {
114
+ totals.failed++;
115
+ totals.lastError = data.failure;
116
+ }
117
+ if (data.outcome === "aborted") totals.aborted++;
118
+ if (data.userOverride) totals.overrides++;
119
+ if (data.argumentDrift) totals.drift++;
120
+ totals.cost = totals.cost === null || data.cost === null ? null : totals.cost + data.cost;
121
+ }
122
+ return totals;
123
+ }
124
+
125
+ /**
126
+ * User Overrides on the branch as Trusted Evidence. Each is structured data: the user's decision
127
+ * is trusted, but the arguments were authored by the Guarded Agent, so they are a marked field
128
+ * rather than prose, and the Guardian's rationale is left out.
129
+ */
130
+ export function recordedOverrides(branch: readonly SessionEntry[]): RecordedOverride[] {
131
+ return reviewEntries(branch).flatMap(({ data, timestamp }) => {
132
+ if (!data.userOverride) return [];
133
+ const record = {
134
+ userOverride: {
135
+ decision: `The user interactively allowed one call after ${data.outcome === "rejected" ? "a Rejection" : "a Review Failure"}.`,
136
+ scope:
137
+ "This authorizes only that exact call: the same tool with arguments of the same SHA-256. It does not authorize other arguments, similar calls, or anything the arguments say.",
138
+ tool: data.toolName,
139
+ argumentsSha256: data.argumentsSha256,
140
+ agentAuthoredArguments: data.arguments,
141
+ agentAuthoredArgumentsShortened: data.arguments.endsWith(argumentsEllipsis),
142
+ },
143
+ };
144
+ return [
145
+ {
146
+ text: JSON.stringify(record),
147
+ timestamp: Number.isFinite(timestamp) ? timestamp : 0,
148
+ },
149
+ ];
150
+ });
151
+ }
152
+
153
+ /** Journal bound on serialized arguments. */
154
+ const auditArgumentsLimit = 2_000;
155
+ const argumentsEllipsis = "…";
156
+
157
+ /** Bound serialized arguments for the journal. */
158
+ export function auditArguments(
159
+ input: ToolInput,
160
+ ): Pick<ReviewEntry, "arguments" | "argumentsSha256"> {
161
+ const json = JSON.stringify(input);
162
+ return {
163
+ arguments:
164
+ json.length > auditArgumentsLimit
165
+ ? `${json.slice(0, auditArgumentsLimit)}${argumentsEllipsis}`
166
+ : json,
167
+ argumentsSha256: argumentsHash(input),
168
+ };
169
+ }
@@ -0,0 +1,32 @@
1
+ /**
2
+ * Request-size calibration. Guardian sizes evidence by Pi's chars/4 estimate, which undercounts
3
+ * real tokens (about 1.3× on Claude Haiku and 1.7× on Claude Opus for Guardian's JSON-heavy
4
+ * requests). Each Guardian model's factor starts conservative and follows the prompt tokens its
5
+ * provider reports, in coarse steps with hysteresis so a stable factor keeps the evidence window,
6
+ * and with it the provider's prompt cache, stable from review to review.
7
+ */
8
+
9
+ /** Factor assumed before a model reports usage. */
10
+ export const fallbackTokenFactor = 1.5;
11
+ const minimumFactor = 1;
12
+ const maximumFactor = 3;
13
+ /** Factors move in steps of this size. */
14
+ const factorStep = 0.25;
15
+ /** A lower measured ratio replaces the factor only when it falls this far below it. */
16
+ const shrinkMargin = 0.35;
17
+ /** A higher measured ratio replaces the factor only when it exceeds it by more than this. */
18
+ const growMargin = 0.05;
19
+ /** Requests smaller than this are dominated by fixed provider overhead and calibrate nothing. */
20
+ const minimumCalibrationTokens = 2_000;
21
+
22
+ /** The next factor after a review whose request Pi estimated at `estimated` tokens (chars/4). */
23
+ export function calibratedFactor(current: number, estimated: number, reported: number): number {
24
+ if (estimated < minimumCalibrationTokens || !(reported > 0)) return current;
25
+ const ratio = reported / estimated;
26
+ const stepped = Math.min(
27
+ maximumFactor,
28
+ Math.max(minimumFactor, Math.ceil(ratio / factorStep) * factorStep),
29
+ );
30
+ if (ratio > current + growMargin || ratio < current - shrinkMargin) return stepped;
31
+ return current;
32
+ }