@agenttrail/guardrails 0.2.1 → 0.2.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,46 +1,118 @@
1
- <!-- cspell:words exfiltration kubeconfig -->
1
+ <!-- cspell:words exfiltration kubeconfig hostnames agenttrailhq OWASP slopsquatted -->
2
2
 
3
- # @agenttrail/guardrails
3
+ <div align="center">
4
4
 
5
- **A library of rules that spot dangerous commands before an AI coding agent runs them.**
5
+ <a href="https://www.agenttrail.sh">
6
+ <picture>
7
+ <source media="(prefers-color-scheme: dark)" srcset="https://www.agenttrail.sh/brand/agenttrail_logo_white.svg" />
8
+ <img src="https://www.agenttrail.sh/brand/agenttrail_logo.svg" alt="AgentTrail" width="220" />
9
+ </picture>
10
+ </a>
6
11
 
7
- 74 rules, grouped into 11 packs. Apache-2.0.
12
+ # AgentTrail Guardrails: open-source rules for AI coding agents
8
13
 
9
- ---
14
+ **A library of rules for catching dangerous commands before an AI coding agent runs them.**<br />
15
+ Plain data you can read, test, and change. Enforced in **Claude Code**, **Cursor**, and **Codex CLI**.
16
+
17
+ [![CI](https://github.com/agenttrailhq/guardrails/actions/workflows/ci.yml/badge.svg)](https://github.com/agenttrailhq/guardrails/actions/workflows/ci.yml)
18
+ [![npm version](https://img.shields.io/npm/v/@agenttrail/guardrails?color=0748FE&label=npm)](https://www.npmjs.com/package/@agenttrail/guardrails)
19
+ [![License: Apache-2.0](https://img.shields.io/badge/license-Apache--2.0-0748FE)](LICENSE)
20
+ [![Node.js 22.13+](https://img.shields.io/badge/node-%3E%3D22.13-0748FE)](https://nodejs.org/)
21
+
22
+ [Quick start](#quick-start) ·
23
+ [The packs](#the-eleven-packs) ·
24
+ [OWASP coverage](#owasp-coverage) ·
25
+ [Rule format](#a-rule-start-to-finish) ·
26
+ [Contribute a rule](#contributing-a-rule) ·
27
+ [Changelog](CHANGELOG.md)
28
+
29
+ </div>
30
+
31
+ [74 rules](https://www.agenttrail.sh/guardrails), grouped into 11 packs. Apache-2.0. Published
32
+ as `@agenttrail/guardrails` and enforced by [AgentTrail Guard](https://github.com/agenttrailhq/guard).
33
+
34
+ ![A Claude Code session: asked to split a name column and push the schema, the agent edits the Prisma schema and seed, then runs npx prisma db push --force-reset to wipe the database. AgentTrail Guard blocks it with the guardrail dd.accept-data-loss, and the agent stops](.github/assets/guard-blocks-force-reset.gif)
35
+
36
+ <sub>A real Claude Code session. The agent decides to reset the database to get a schema push
37
+ through; Guard blocks `prisma db push --force-reset` with `dd.accept-data-loss` before it runs.</sub>
10
38
 
11
39
  ## What this is, in plain terms
12
40
 
13
- When you let an AI coding agent work in your terminal, it eventually proposes something you would
14
- not have typed yourself — `git reset --hard` over a day's work, `rm -rf` on the wrong path, a
41
+ Your AI coding agent proposes commands in your terminal all day, and some are ones you would
42
+ not have typed yourself: `git reset --hard` over a day's work, `rm -rf` on the wrong path, a
15
43
  `terraform apply` against production.
16
44
 
17
45
  This package is the **list of things worth stopping**, written as data. Each entry says what to look
18
- for, how serious it is, and what should happen — allow it, ask a human first, or refuse.
46
+ for, how serious it is, and what should happen: allow it, ask a human first, or refuse.
19
47
 
20
48
  **That is all this package does.** It contains no code that watches your machine and nothing that
21
- talks to the network. It is a list. Something else has to read it and act on it — normally
49
+ talks to the network. Something else has to read the list and act on it: normally
22
50
  [`agenttrail-guard`](https://github.com/agenttrailhq/guard), which runs on your laptop and checks each
23
51
  command an agent proposes against these rules.
24
52
 
25
- Separating the two is deliberate. You can read every rule here, disagree with one, and change it,
26
- without trusting anything about the tool that enforces them.
53
+ ## Why these rules?
54
+
55
+ - **Readable.** Every rule is plain data with a title, a severity, a default action, and a
56
+ description that says what it catches **and what it misses**.
57
+ - **Proven in both directions.** Every rule ships commands it must catch and near-misses it must
58
+ leave alone, and a shared corpus of everyday commands no rule may match. A rule that fails
59
+ either side does not build.
60
+ - **Quiet on purpose.** A commit message or a `grep` that only *mentions* `rm -rf /` is not
61
+ treated as running it. See [Talking about a command is not running it](#talking-about-a-command-is-not-running-it).
62
+ - **Mapped to known risks.** The packs line up with the
63
+ [OWASP Top 10s for agentic and LLM applications](#owasp-coverage), and the gaps are stated, not
64
+ hidden.
65
+ - **Separate from the enforcer.** You can read, disagree with, and change any rule without
66
+ trusting anything about the tool that enforces it.
67
+
68
+ ## Quick start
69
+
70
+ **You need:** Node.js 22.13 or newer.
71
+
72
+ ### Enforce the rules in your agent
73
+
74
+ Most people never install this package directly. [AgentTrail Guard](https://github.com/agenttrailhq/guard)
75
+ bundles it, so installing Guard is all it takes:
76
+
77
+ ```bash
78
+ npm install -g @agenttrail/guard
79
+ agenttrail-guard init --agent claude # or: --agent cursor, --agent codex
80
+ agenttrail-guard guardrails list # every rule and its current action
81
+ agenttrail-guard guardrails show wt.reset-hard # one rule in full
82
+ agenttrail-guard guardrails set-action <id> warn # change what a rule does
83
+ agenttrail-guard guardrails allow <id> <pattern> # silence one rule for one command shape
84
+ ```
27
85
 
28
- ## Who this is for
86
+ See the [AgentTrail Guard README](https://github.com/agenttrailhq/guard#quick-start) for setup
87
+ and everything else Guard can do.
29
88
 
30
- - **You use an AI coding agent** and want a sensible default set of guardrails rather than writing
31
- your own from scratch.
32
- - **You want to see exactly what is being blocked and why.** Every rule is plain data with a written
33
- explanation, including what it *misses*.
34
- - **You want to contribute a rule** you wish had existed. See [Contributing](#contributing-a-rule).
89
+ ### Use the rules in your own code
35
90
 
36
- ## Install
91
+ Install the package if you are writing rules, or building your own tool on top of the list:
37
92
 
38
93
  ```bash
39
94
  npm install @agenttrail/guardrails
40
95
  ```
41
96
 
42
- Most people never install this directly — `agenttrail-guard` bundles it. Install it yourself if you
43
- are writing rules, or building your own tool on top of the list.
97
+ ```ts
98
+ import { RULES, getRule, rulesForPack } from "@agenttrail/guardrails/guardrails";
99
+
100
+ // Everything that refuses outright, rather than asking or warning.
101
+ const blocking = RULES.filter((rule) => rule.defaultAction === "block");
102
+ console.log(`${RULES.length} rules, ${blocking.length} block by default`);
103
+
104
+ // One pack, with each rule's severity and default action.
105
+ for (const rule of rulesForPack("secret-exposure")) {
106
+ console.log(rule.id, rule.severity, rule.defaultAction);
107
+ }
108
+
109
+ // One rule, including the limits its description admits to.
110
+ console.log(getRule("wt.reset-hard")?.description);
111
+ ```
112
+
113
+ This package has no evaluator: it tells you what each rule looks for, not whether a given
114
+ command matches. Guard does the matching. [Two ways to import it](#two-ways-to-import-it)
115
+ explains the two entry points.
44
116
 
45
117
  ## A rule, start to finish
46
118
 
@@ -53,7 +125,7 @@ Here is a complete rule. Nothing is hidden; this is the actual shape.
53
125
  "severity": "high",
54
126
  "defaultAction": "block",
55
127
  "title": "git reset --hard discards uncommitted work",
56
- "description": "Discards all uncommitted changes. Does not match `git restore` — see wt.restore-path.",
128
+ "description": "Discards all uncommitted changes. Does not match `git restore` (see wt.restore-path).",
57
129
 
58
130
  // What to look for. This one matches a Bash command against a regular expression.
59
131
  "match": {
@@ -62,7 +134,7 @@ Here is a complete rule. Nothing is hidden; this is the actual shape.
62
134
  ]
63
135
  },
64
136
 
65
- // Proof it works, in both directions — see "Every rule proves both directions" below.
137
+ // Proof it works, in both directions (see "Every rule proves both directions" below).
66
138
  "fixtures": {
67
139
  "block": ["git reset --hard"],
68
140
  "allow": ["git reset src/api.ts"]
@@ -70,12 +142,17 @@ Here is a complete rule. Nothing is hidden; this is the actual shape.
70
142
  }
71
143
  ```
72
144
 
145
+ The real rule is in [`wt-reset-hard.ts`](src/packs/working-tree/wt-reset-hard.ts); this is a
146
+ trimmed copy. With it enforced by Guard, an agent that proposes `git reset --hard` is refused
147
+ and told which rule stopped it, while `git reset src/api.ts`, which unstages one file, goes
148
+ ahead as normal.
149
+
73
150
  Reading the fields:
74
151
 
75
152
  | Field | What it means |
76
153
  |---|---|
77
154
  | `id` | A stable name. Users type it to disable or change a rule, so it never changes. |
78
- | `severity` | How bad the thing being caught is: `critical`, `high`, `medium`, `low`, `info`. **It is not a price** — this package ships no mapping from severity to money. |
155
+ | `severity` | How bad the thing being caught is: `critical`, `high`, `medium`, `low`, `info`. **It is not a price**: this package ships no mapping from severity to money. |
79
156
  | `defaultAction` | What should happen: `block` (refuse), `require_approval` (ask a human), or `warn` (allow, but say so). A user can override it. |
80
157
  | `description` | What the rule catches **and what it misses**. The honest limits are part of the rule, not a footnote. |
81
158
  | `match` | The condition. `any_of` means "any one of these is enough". |
@@ -86,26 +163,64 @@ Reading the fields:
86
163
  A rule is filed by **the harm it prevents**, never by the technique it uses to spot it.
87
164
 
88
165
  That sounds like a detail and is not. The three rules about production config, `.env` files and API
89
- endpoints all work by matching file paths — but they are *not* in `file-scope`. Someone who turned
166
+ endpoints all work by matching file paths, but they are *not* in `file-scope`. Someone who turned
90
167
  that pack off to stop path noise would otherwise silently lose their production and secret
91
168
  protection, which they never asked to turn off and would not know they had.
92
169
 
93
170
  | Pack | Rules | What it is about |
94
171
  |---|---:|---|
95
- | `working-tree` | 9 | Destroying uncommitted work or published history — `git reset --hard`, `git clean -fd`, force-push, `rm -rf`. |
96
- | `destructive-data` | 8 | Data git cannot bring back — a dropped volume, a dropped database, destructive DDL, a deleted shadow copy. |
97
- | `prod-infra` | 8 | Changing running infrastructure — Terraform, Kubernetes, Helm, cloud deletes, a deploy that names production. |
98
- | `secret-exposure` | 10 | Credentials and sensitive data leaving where they live. Mostly `warn`: reading a secret is a normal part of a normal day. |
99
- | `rce-supply-chain` | 6 | Running code nobody reviewed — pipe-to-shell, a remote runner, a redirected registry, TLS verification off. |
100
- | `safety-bypass` | 7 | Turning off a check somebody installed on purpose, or erasing the record of it — `--no-verify`, `--admin` merge, hooks disabled, host-key checking off, history and log purges, forged terminal output. |
101
- | `privilege-supply-chain` | 6 | Gaining reach or handing it out — `sudo` writes, `chmod 777`, IAM grants, persistence, publishing, new dependencies. |
102
- | `file-scope` | 4 | The agent wrote somewhere it had no business writing — its own config, the machine, git's internals, the CI definition. |
103
- | `agent-context` | 6 | The agent changing what it is or what it knows — its standing instructions, its memory, its skills and commands, its MCP servers — or starting more agents, or switching another agent's approvals off. |
104
- | `test-integrity` | 6 | The agent making its work look successful — deleting a test, weakening a runner's configuration, accepting every snapshot, switching a coverage gate off, silencing failures in bulk, or telling CI not to run. |
105
- | `exfiltration` | 4 | Moving data off the machine or opening a way in — a reverse shell, a public tunnel, a file upload, a paste service. Command channel only. |
172
+ | [`working-tree`](src/packs/working-tree/) | 9 | Destroying uncommitted work or published history: `git reset --hard`, `git clean -fd`, force-push, `rm -rf`. |
173
+ | [`destructive-data`](src/packs/destructive-data/) | 8 | Data git cannot bring back: a dropped volume, a dropped database, destructive DDL, a deleted shadow copy. |
174
+ | [`prod-infra`](src/packs/prod-infra/) | 8 | Changing running infrastructure: Terraform, Kubernetes, Helm, cloud deletes, a deploy that names production. |
175
+ | [`secret-exposure`](src/packs/secret-exposure/) | 10 | Credentials and sensitive data leaving where they live. Mostly `warn`: reading a secret is a normal part of a normal day. |
176
+ | [`rce-supply-chain`](src/packs/rce-supply-chain/) | 6 | Running code nobody reviewed: pipe-to-shell, a remote runner, a redirected registry, TLS verification off. |
177
+ | [`safety-bypass`](src/packs/safety-bypass/) | 7 | Turning off a check somebody installed on purpose, or erasing the record of it: `--no-verify`, `--admin` merge, hooks disabled, host-key checking off, history and log purges, forged terminal output. |
178
+ | [`privilege-supply-chain`](src/packs/privilege-supply-chain/) | 6 | Gaining reach or handing it out: `sudo` writes, `chmod 777`, IAM grants, persistence, publishing, new dependencies. |
179
+ | [`file-scope`](src/packs/file-scope/) | 4 | The agent wrote somewhere it had no business writing: its own config, the machine, git's internals, the CI definition. |
180
+ | [`agent-context`](src/packs/agent-context/) | 6 | The agent changing what it is or what it knows (its standing instructions, its memory, its skills and commands, its MCP servers), or starting more agents, or switching another agent's approvals off. |
181
+ | [`test-integrity`](src/packs/test-integrity/) | 6 | The agent making its work look successful: deleting a test, weakening a runner's configuration, accepting every snapshot, switching a coverage gate off, silencing failures in bulk, or telling CI not to run. |
182
+ | [`exfiltration`](src/packs/exfiltration/) | 4 | Moving data off the machine or opening a way in: a reverse shell, a public tunnel, a file upload, a paste service. Command channel only. |
106
183
 
107
184
  Pack names appear in user config files, so renaming one is a breaking change, not a tidy-up.
108
185
 
186
+ Each pack name links to its source folder: one TypeScript file per rule, so you can read any rule
187
+ end to end. Every rule, with its full description and tested examples, is also in the
188
+ [guardrail library](https://www.agenttrail.sh/guardrails) on the AgentTrail website, where you
189
+ can paste a command to see which rules match it.
190
+
191
+ ## OWASP coverage
192
+
193
+ Several packs exist because of risks named in OWASP's two lists for AI systems: the
194
+ [Top 10 for Agentic Applications](https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/)
195
+ (ASI01 to ASI10, 2026 edition) and the
196
+ [Top 10 for LLM Applications](https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/)
197
+ (LLM01 to LLM10, 2026 edition). A rule cannot see a prompt or a model's reasoning, only the
198
+ action that comes out of it, so these rules address a risk at the point where it turns into a
199
+ command or a file change.
200
+
201
+ | Pack | Agentic risks | LLM risks |
202
+ |---|---|---|
203
+ | `working-tree`, `destructive-data`, `prod-infra`, `file-scope` | ASI02 Tool Misuse and Exploitation | LLM03 Excessive Agency |
204
+ | `rce-supply-chain` | ASI05 Unexpected Code Execution, ASI04 Agentic Supply Chain | LLM04 Supply Chain |
205
+ | `privilege-supply-chain` | ASI03 Identity and Privilege Abuse, ASI04 Agentic Supply Chain | LLM04 Supply Chain |
206
+ | `secret-exposure`, `exfiltration` | ASI02 Tool Misuse and Exploitation, ASI03 Identity and Privilege Abuse | LLM02 Sensitive Information Disclosure |
207
+ | `safety-bypass` | ASI09 Human-Agent Trust Exploitation | LLM10 Improper Output Handling (`gb.ansi-terminal-forgery`) |
208
+ | `agent-context` | ASI06 Memory and Context Poisoning (the write, not its contents), ASI10 Rogue Agents | LLM03 Excessive Agency |
209
+ | `test-integrity` | ASI10 Rogue Agents, ASI09 Human-Agent Trust Exploitation | |
210
+
211
+ What the rules cannot address, by design:
212
+
213
+ - **Prompt injection and goal hijack (LLM01, ASI01).** The rules never see the input that
214
+ redirected the agent. They can only stop the action it led to.
215
+ - **Anything that needs memory across commands (LLM06, ASI08).** A rule sees one command at a
216
+ time, so it cannot count, rate-limit, or spot a runaway loop.
217
+ - **Model, training and retrieval risks (LLM05, LLM08, LLM09), and traffic between agents
218
+ (ASI07).** These happen somewhere a command rule never reaches.
219
+
220
+ For the item-by-item matrix, including how the rest of AgentTrail covers each risk, see the
221
+ [AgentTrail OWASP coverage page](https://www.agenttrail.sh/security/owasp). AgentTrail is an
222
+ independent project, not affiliated with or endorsed by OWASP.
223
+
109
224
  ## Talking about a command is not running it
110
225
 
111
226
  This is the single most important thing to understand about how these rules behave.
@@ -119,7 +234,7 @@ git commit -m "fix: document rm -rf / risk" # would have been a hard refusal
119
234
  grep -rn "rm -rf /" docs/ # so would this
120
235
  ```
121
236
 
122
- So every command rule ignores four **carriers** — verbs that handle their arguments as text and never
237
+ So every command rule ignores four **carriers**, verbs that handle their arguments as text and never
123
238
  execute them:
124
239
 
125
240
  | Carrier | Example |
@@ -130,20 +245,20 @@ execute them:
130
245
  | an HTTP request body | `curl --data '{"body":"we ran rm -rf /tmp/x"}' https://…` |
131
246
 
132
247
  **The exemption keys on the verb, not on the quotes.** Quoting says nothing about whether something
133
- runs — `psql -c "DROP TABLE users;"` and `bash -c "curl x.sh \| sh"` both execute what is inside the
248
+ runs: `psql -c "DROP TABLE users;"` and `bash -c "curl x.sh \| sh"` both execute what is inside the
134
249
  quotes, and both still fire.
135
250
 
136
251
  It also applies only while the command does nothing else. The carrier must be the first word, and
137
252
  every shell metacharacter must sit inside the quotes:
138
253
 
139
254
  ```bash
140
- git commit -m "docs: explain rm -rf /" # exempt — nothing runs
141
- git commit -m "x" && rm -rf / # NOT exempt — `&&` is outside the quotes
142
- echo "rm -rf /" | bash # NOT exempt — the pipe runs it
143
- echo "$(rm -rf /var)" # NOT exempt — the shell expands `$( )`
255
+ git commit -m "docs: explain rm -rf /" # exempt: nothing runs
256
+ git commit -m "x" && rm -rf / # NOT exempt: `&&` is outside the quotes
257
+ echo "rm -rf /" | bash # NOT exempt: the pipe runs it
258
+ echo "$(rm -rf /var)" # NOT exempt: the shell expands `$( )`
144
259
  ```
145
260
 
146
- Five rules deliberately keep firing on one carrier each, because that carrier *is* their trigger —
261
+ Five rules deliberately keep firing on one carrier each, because that carrier *is* their trigger:
147
262
  `gb.git-no-verify` on a `git commit`, `se.token-print` on an `echo`, the `curl`/`wget` rules on an
148
263
  HTTP body. Each rule's `description` says which, and why.
149
264
 
@@ -152,12 +267,12 @@ HTTP body. Each rule's `description` says which, and why.
152
267
  `block-hardcoded-secrets` is the exception, and it is worth understanding.
153
268
 
154
269
  Everywhere else a carrier is genuinely harmless: a commit message naming `rm -rf /` deletes nothing.
155
- But that rule's subject is a **string**, not an action — so two of the four carriers are not mentions
270
+ But that rule's subject is a **string**, not an action, so two of the four carriers are not mentions
156
271
  at all. They are the exposure itself:
157
272
 
158
273
  | Carrier | What happens to the key | Exempt? |
159
274
  |---|---|---|
160
- | `grep -rn AKIA .` | searched for, goes nowhere — and this is how you find a key to rotate | **yes** |
275
+ | `grep -rn AKIA .` | searched for, goes nowhere, and this is how you find a key to rotate | **yes** |
161
276
  | `echo "AKIA…"` | transient terminal output | **yes** |
162
277
  | `git commit -m "…AKIA…"` | written into history, then pushed | **no** |
163
278
  | `curl --data "…AKIA…"` | sent to a remote host | **no** |
@@ -165,26 +280,8 @@ at all. They are the exposure itself:
165
280
  The cost, stated in the other direction: documenting a real-looking key in a commit message is still
166
281
  blocked. Redact the body of the key, or use a placeholder short enough to fail the length check.
167
282
 
168
- **Known limits of the carrier logic**, in the same spirit as the rest of this file. The carrier must
169
- be the first word — a single leading `sudo` is tolerated, because it changes privilege rather than
170
- meaning, but a runner prefix is not: `pnpm exec rg …`, `npx …` and `xargs -0 grep …` still fire,
171
- since "some program eventually runs a search" is a much weaker claim than "this command is a search".
172
- At most four quoted arguments are recognised. A carrier that can be made to execute through a flag —
173
- `ack --pager='…'`, `rg --pre <cmd>` — is still treated as a mention; closing that needs information
174
- the checker does not have. And an MCP tool whose input carries the same text is not exempt either,
175
- because exempting a JSON blob would exempt a shell-running MCP server along with it.
176
-
177
- **MCP coverage.** A command rule fires on `Bash`, `PowerShell` *and* any `mcp__*` tool: the guard hands
178
- the checker an MCP call's serialized `tool_input` as the same command text every command rule reads, so
179
- a command shape run through an MCP server — `{"command":"rm -rf /"}` — is caught, not ignored. Two
180
- honest limits follow from that. First, a rule whose pattern is anchored to the start of the command
181
- (`^…` or a command-position class) may not fire inside the JSON, where the shape sits after a `"`
182
- rather than at a command boundary; the `\b`-anchored rules — most of the corpus — do fire. Second, the
183
- quoted-mention exemptions are shell-only, so an MCP payload that merely *names* a command in a text
184
- field (`{"title":"fix the rm -rf / bug"}`) is matched the same as one that runs it — a JSON blob cannot
185
- be told apart from a shell-running MCP server. File rules match by path on whichever file tool a client
186
- uses. No rule is shell-only by design; a rule that does not reach the MCP channel does so because its
187
- pattern, not its label, does not match the serialized shape.
283
+ The finer limits of the carrier logic, and how command rules apply to MCP tools, are in the
284
+ [Reference](#reference).
188
285
 
189
286
  ## What these rules deliberately do not catch
190
287
 
@@ -193,29 +290,29 @@ somebody forgot.
193
290
 
194
291
  - **Nothing about the web.** Pages an agent fetches are not checked, and there are no URL rules.
195
292
  - **Nothing inside a file.** The checker sees a file's *path*, never its contents. A secret typed into
196
- a source file, SQL built by string concatenation, a missing auth check — none of it is visible.
293
+ a source file, SQL built by string concatenation, a missing auth check: none of it is visible.
197
294
  Rules that would need it are absent rather than approximated.
198
295
  - **Nothing about where you are.** No working directory, no project root, no git branch, no cloud
199
- profile reaches the checker — it gets one command and nothing else. So "the agent wrote outside the
296
+ profile reaches the checker; it gets one command and nothing else. So "the agent wrote outside the
200
297
  project" **cannot be written as a rule**, and `file-scope` is limited to well-known absolute paths
201
298
  for good. For the same reason, a rule cannot tell a scratch database from a production one.
202
- - **Nothing hidden inside a quoted payload.** Where the danger is inside a quoted argument —
203
- `psql -c "<sql>"`, `python -c "<code>"` — a text rule can only guess. In a long script, a match says
204
- very little about what the script actually does.
299
+ - **Nothing hidden inside a quoted payload.** Where the danger is inside a quoted argument
300
+ (`psql -c "<sql>"`, `python -c "<code>"`), a text rule can only guess. In a long script, a match
301
+ says very little about what the script actually does.
205
302
  - **Nothing a wrapper hides.** `./deploy.sh` that runs `terraform apply -auto-approve` inside it is
206
303
  just a shell script from the outside.
207
304
  - **Nothing recurring.** There is no counting. "The same mistake three times this week" needs memory
208
305
  across commands, and a single command has none.
209
306
 
210
307
  Each rule's own `description` names its specific misses. Read those before trusting a rule to cover a
211
- case — they are written to be believed, not to sell.
308
+ case; they are written to be believed, not to sell.
212
309
 
213
310
  ## Every rule proves both directions
214
311
 
215
312
  Every rule ships at least one **`block`** example and at least one **`allow`** example. A rule missing
216
313
  either does not build.
217
314
 
218
- - **`block` means "this rule must match."** It does *not* mean the agent is refused — most packs
315
+ - **`block` means "this rule must match."** It does *not* mean the agent is refused; most packs
219
316
  default to asking or warning.
220
317
  - **`allow` means "this rule must NOT match."** *This is the half that matters.* Anyone can write a
221
318
  rule that catches `rm -rf /`. The hard part is not firing on `rm -rf ./node_modules` forty times a
@@ -228,13 +325,13 @@ Examples come in two kinds, and each must use the right one:
228
325
  | a command | `Bash`, `PowerShell`, a search query, an MCP tool's input |
229
326
  | a file path | `Edit`, `Write`, `Read`, `MultiEdit`, `NotebookEdit` |
230
327
 
231
- Giving a path-matching rule a command example makes it pass **without testing anything** — it matches
328
+ Giving a path-matching rule a command example makes it pass **without testing anything**: it matches
232
329
  nothing, which reads as proof of quietness and proves only that the path never reached the rule. CI
233
330
  rejects that.
234
331
 
235
332
  ## Contributing a rule
236
333
 
237
- Rules are meant to be contributed. The bar is not "clever regex" — it is **does it fire on the real
334
+ Rules are meant to be contributed. The bar is not "clever regex"; it is **does it fire on the real
238
335
  thing, and stay quiet on the near-miss**.
239
336
 
240
337
  **1. Write it**, following the shape above. Give it a `description` that says what it misses.
@@ -258,35 +355,79 @@ agenttrail-guard guardrails validate ./my-rule.json
258
355
  does **not** answer *"does it actually fire on the command I think it does?"* That needs the real
259
356
  checker, which is not part of this package.
260
357
 
261
- **So the real test runs in CI, on your pull request** — the same check, on the same machine, for
358
+ **So the real test runs in CI, on your pull request**: the same check, on the same machine, for
262
359
  everyone. You get the shape check instantly here and the real answer there, which is where it has to
263
360
  run to be trusted anyway.
264
361
 
265
- Your rule is also run against a **quiet corpus**: 328 everyday commands and paths that no rule may
266
- match at all. Your own negative example only proves your rule is quiet on the near-miss *you* thought
267
- of. The quiet corpus is what catches a Terraform rule firing on `pnpm test`.
362
+ Your rule is also run against a **quiet corpus**: hundreds of everyday commands and paths that no
363
+ rule may match at all. Your own negative example only proves your rule is quiet on the near-miss
364
+ *you* thought of. The quiet corpus is what catches a Terraform rule firing on `pnpm test`.
268
365
 
269
366
  ### Shapes that will not validate
270
367
 
271
- Three are rejected outright, each because it produces a rule that *looks* enforced and is not — the
368
+ Three are rejected outright, each because it produces a rule that *looks* enforced and is not: the
272
369
  worst failure a security tool can have.
273
370
 
274
371
  | Rejected | Why |
275
372
  |---|---|
276
- | `scope` | It compares against ids that are always UUIDs, never a vendor name — so a scoped rule matches nothing, forever, silently. |
373
+ | `scope` | It compares against ids that are always UUIDs, never a vendor name, so a scoped rule matches nothing, forever, silently. |
277
374
  | Numeric conditions | Token counts and durations are all zero *before* a command runs. "Greater than" can never fire; "less than" fires on everything. |
278
375
  | A command matcher and a file matcher in one condition | No real command carries both, so the condition can never be true. Split it into two under `any_of`. |
279
376
 
280
377
  A regular expression is also rejected if it nests unbounded repetition (`(a+)+`). The guard fails
281
- **open** under a time limit, so a pattern that backtracks does not merely run slowly — it lets the
378
+ **open** under a time limit, so a pattern that backtracks does not merely run slowly; it lets the
282
379
  command through.
283
380
 
284
381
  One more is caught in CI rather than by the shape check, because it cannot be caught earlier: **a
285
- single-item brace list in a tool name.** `"{Bash}"` is a glob pattern, and it does not match `Bash` —
382
+ single-item brace list in a tool name.** `"{Bash}"` is a glob pattern, and it does not match `Bash`,
286
383
  so the rule matches nothing, forever, with no error anywhere. Write a single tool plainly as
287
384
  `"Bash"`; braces are for real alternatives, `"{Bash,PowerShell}"`.
288
385
 
289
- ## Two ways to import it
386
+ ### Ready to open a pull request?
387
+
388
+ [CONTRIBUTING.md](CONTRIBUTING.md) has a pre-submission checklist, local build and test
389
+ commands, and where to report what. Please follow the [Code of Conduct](CODE_OF_CONDUCT.md),
390
+ and report vulnerabilities privately as described in [SECURITY.md](SECURITY.md).
391
+
392
+ ## Stability and versioning
393
+
394
+ Rule ids and pack names end up in people's config files, so they are treated as public API:
395
+
396
+ - **A rule id never changes.** Renaming or removing a rule would silently undo an override
397
+ someone set on it, so either one is a breaking change.
398
+ - **A pack name never changes** either, for the same reason: a renamed pack silently
399
+ re-enables a pack someone had turned off.
400
+ - **A new rule or pack is a feature.** Guard turns every pack on unless you turned it off, so a
401
+ rule added in a release you install starts enforcing with no step of your own.
402
+
403
+ Releases are cut automatically from [Conventional Commits](https://www.conventionalcommits.org/).
404
+ Before 1.0, a breaking change ships as a minor release and a new rule as a patch release, so
405
+ skim the [changelog](CHANGELOG.md) before you upgrade.
406
+
407
+ ## FAQ
408
+
409
+ **Can I use these rules without Guard?** Yes. They are plain data under Apache-2.0. You need your
410
+ own matcher to apply them, because this package does not contain one.
411
+
412
+ **A rule fires on something legitimate. What do I do?** In Guard, `agenttrail-guard guardrails
413
+ allow <id> <pattern>` silences that one rule for that one command shape, without turning the rule
414
+ off. Then please [open an issue](https://github.com/agenttrailhq/guardrails/issues) with a
415
+ sanitized example, so the rule can learn the near-miss.
416
+
417
+ **Why didn't a rule fire on my command?** Start with the rule's `description`: it lists what the
418
+ rule misses. Then check [what these rules deliberately do not catch](#what-these-rules-deliberately-do-not-catch).
419
+ To test a command against every rule, paste it into the [online checker](https://www.agenttrail.sh/guardrails#check).
420
+
421
+ **Why is `secret-exposure` mostly `warn`?** Reading a secret is a normal part of a normal day.
422
+ Refusing every `.env` read would be turned off within a week. Moving a secret off the machine is
423
+ held for approval instead.
424
+
425
+ **How do I turn off a whole pack?** `agenttrail-guard guardrails disable <pack>`. See
426
+ [Tune a guardrail](https://github.com/agenttrailhq/guard#tune-a-guardrail) in the Guard README.
427
+
428
+ ## Reference
429
+
430
+ ### Two ways to import it
290
431
 
291
432
  ```ts
292
433
  import { RULES, getRule } from "@agenttrail/guardrails/guardrails"; // just the rules
@@ -300,16 +441,62 @@ They are separate because the guard starts a fresh process on **every single com
300
441
  under a ten-second ceiling. It cannot afford to load a validator it never calls, or to re-check 74
301
442
  rules that were already checked before release.
302
443
 
303
- ## Where these rules came from
444
+ ### Known limits of the carrier logic
445
+
446
+ In the same spirit as the rest of this file:
447
+
448
+ - **The carrier must be the first word.** A single leading `sudo` is tolerated, because it changes
449
+ privilege rather than meaning, but a runner prefix is not: `pnpm exec rg …`, `npx …` and
450
+ `xargs -0 grep …` still fire, since "some program eventually runs a search" is a much weaker claim
451
+ than "this command is a search".
452
+ - **At most four quoted arguments are recognised.**
453
+ - **A carrier that can be made to execute through a flag** (`ack --pager='…'`, `rg --pre <cmd>`) is
454
+ still treated as a mention; closing that needs information the checker does not have.
455
+ - **An MCP tool whose input carries the same text is not exempt either,** because exempting a JSON
456
+ blob would exempt a shell-running MCP server along with it.
457
+
458
+ ### MCP coverage
459
+
460
+ A command rule fires on `Bash`, `PowerShell` *and* any `mcp__*` tool: the guard hands the checker an
461
+ MCP call's serialized `tool_input` as the same command text every command rule reads, so a command
462
+ shape run through an MCP server (`{"command":"rm -rf /"}`) is caught, not ignored. Two honest limits
463
+ follow from that:
464
+
465
+ 1. **Start-anchored patterns may miss.** A rule whose pattern is anchored to the start of the command
466
+ (`^…` or a command-position class) may not fire inside the JSON, where the shape sits after a `"`
467
+ rather than at a command boundary; the `\b`-anchored rules (most of the corpus) do fire.
468
+ 2. **Mentions are not exempt.** The quoted-mention exemptions are shell-only, so an MCP payload that
469
+ merely *names* a command in a text field (`{"title":"fix the rm -rf / bug"}`) is matched the same
470
+ as one that runs it; a JSON blob cannot be told apart from a shell-running MCP server.
471
+
472
+ File rules match by path on whichever file tool a client uses. No rule is shell-only by design; a
473
+ rule that does not reach the MCP channel does so because its pattern, not its label, does not match
474
+ the serialized shape.
475
+
476
+ ### Where these rules came from
304
477
 
305
478
  Independently authored. No block list, pattern or wording is copied from any other project.
306
479
 
307
- Where a rule's shape follows an obvious convention — an `rm -rf` pattern looks like an `rm -rf`
308
- pattern — that is two people meeting the same shell, not one copying the other.
480
+ Where a rule's shape follows an obvious convention (an `rm -rf` pattern looks like an `rm -rf`
481
+ pattern), that is two people meeting the same shell, not one copying the other.
482
+
483
+ ## Part of AgentTrail
309
484
 
310
- Four vendors' own tools inspired specific rules through their *documented failure modes*, not their
311
- code.
485
+ These rules are part of [AgentTrail](https://www.agenttrail.sh), which builds guardrails and
486
+ observability for AI coding agents:
487
+
488
+ - **[AgentTrail Guard](https://www.agenttrail.sh/agenttrail-guard)** enforces them on every
489
+ tool call that reaches its hooks, on your machine. It is free, open source, and needs no
490
+ account. Its code is in [agenttrailhq/guard](https://github.com/agenttrailhq/guard).
491
+ - **[AgentTrail OS](https://www.agenttrail.sh/agenttrail-os)** uses the same open rule library,
492
+ and adds searchable session history and team workflows. See
493
+ [AgentTrail Guard vs AgentTrail OS](https://www.agenttrail.sh/guard-vs-os).
312
494
 
313
495
  ## License
314
496
 
315
497
  Apache-2.0. See [LICENSE](./LICENSE).
498
+
499
+ Rule ideas and reports go to [Guardrails issues](https://github.com/agenttrailhq/guardrails/issues).
500
+ If you rely on these rules, a star helps other developers find them.
501
+
502
+ Made with ❤️ by [agenttrail.sh](https://www.agenttrail.sh)
@@ -3651,8 +3651,8 @@ function isPack(value) {
3651
3651
  }
3652
3652
 
3653
3653
  // src/stamp.ts
3654
- var CATALOG_VERSION = "0.2.1";
3655
- var CATALOG_PUBLISHED_AT = "2026-09-29T03:49:22Z";
3654
+ var CATALOG_VERSION = "0.2.2";
3655
+ var CATALOG_PUBLISHED_AT = "2026-10-06T13:41:00Z";
3656
3656
 
3657
3657
  // src/rules.ts
3658
3658
  var RULES_BY_PACK = {
@@ -3702,4 +3702,4 @@ export {
3702
3702
  getRule,
3703
3703
  rulesForPack
3704
3704
  };
3705
- //# sourceMappingURL=chunk-SLWHBH7H.js.map
3705
+ //# sourceMappingURL=chunk-GNN4TGRY.js.map