@agenttrail/guardrails 0.2.1 → 0.2.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +274 -87
- package/dist/{chunk-SLWHBH7H.js → chunk-GNN4TGRY.js} +3 -3
- package/dist/{chunk-SLWHBH7H.js.map → chunk-GNN4TGRY.js.map} +1 -1
- package/dist/guardrails.cjs +2 -2
- package/dist/guardrails.cjs.map +1 -1
- package/dist/guardrails.d.cts +2 -2
- package/dist/guardrails.d.ts +2 -2
- package/dist/guardrails.js +1 -1
- package/dist/index.cjs +2 -2
- package/dist/index.cjs.map +1 -1
- package/dist/index.js +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -1,46 +1,118 @@
|
|
|
1
|
-
<!-- cspell:words exfiltration kubeconfig -->
|
|
1
|
+
<!-- cspell:words exfiltration kubeconfig hostnames agenttrailhq OWASP slopsquatted -->
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
<div align="center">
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
<a href="https://www.agenttrail.sh">
|
|
6
|
+
<picture>
|
|
7
|
+
<source media="(prefers-color-scheme: dark)" srcset="https://www.agenttrail.sh/brand/agenttrail_logo_white.svg" />
|
|
8
|
+
<img src="https://www.agenttrail.sh/brand/agenttrail_logo.svg" alt="AgentTrail" width="220" />
|
|
9
|
+
</picture>
|
|
10
|
+
</a>
|
|
6
11
|
|
|
7
|
-
|
|
12
|
+
# AgentTrail Guardrails: open-source rules for AI coding agents
|
|
8
13
|
|
|
9
|
-
|
|
14
|
+
**A library of rules for catching dangerous commands before an AI coding agent runs them.**<br />
|
|
15
|
+
Plain data you can read, test, and change. Enforced in **Claude Code**, **Cursor**, and **Codex CLI**.
|
|
16
|
+
|
|
17
|
+
[](https://github.com/agenttrailhq/guardrails/actions/workflows/ci.yml)
|
|
18
|
+
[](https://www.npmjs.com/package/@agenttrail/guardrails)
|
|
19
|
+
[](LICENSE)
|
|
20
|
+
[](https://nodejs.org/)
|
|
21
|
+
|
|
22
|
+
[Quick start](#quick-start) ·
|
|
23
|
+
[The packs](#the-eleven-packs) ·
|
|
24
|
+
[OWASP coverage](#owasp-coverage) ·
|
|
25
|
+
[Rule format](#a-rule-start-to-finish) ·
|
|
26
|
+
[Contribute a rule](#contributing-a-rule) ·
|
|
27
|
+
[Changelog](CHANGELOG.md)
|
|
28
|
+
|
|
29
|
+
</div>
|
|
30
|
+
|
|
31
|
+
[74 rules](https://www.agenttrail.sh/guardrails), grouped into 11 packs. Apache-2.0. Published
|
|
32
|
+
as `@agenttrail/guardrails` and enforced by [AgentTrail Guard](https://github.com/agenttrailhq/guard).
|
|
33
|
+
|
|
34
|
+

|
|
35
|
+
|
|
36
|
+
<sub>A real Claude Code session. The agent decides to reset the database to get a schema push
|
|
37
|
+
through; Guard blocks `prisma db push --force-reset` with `dd.accept-data-loss` before it runs.</sub>
|
|
10
38
|
|
|
11
39
|
## What this is, in plain terms
|
|
12
40
|
|
|
13
|
-
|
|
14
|
-
not have typed yourself
|
|
41
|
+
Your AI coding agent proposes commands in your terminal all day, and some are ones you would
|
|
42
|
+
not have typed yourself: `git reset --hard` over a day's work, `rm -rf` on the wrong path, a
|
|
15
43
|
`terraform apply` against production.
|
|
16
44
|
|
|
17
45
|
This package is the **list of things worth stopping**, written as data. Each entry says what to look
|
|
18
|
-
for, how serious it is, and what should happen
|
|
46
|
+
for, how serious it is, and what should happen: allow it, ask a human first, or refuse.
|
|
19
47
|
|
|
20
48
|
**That is all this package does.** It contains no code that watches your machine and nothing that
|
|
21
|
-
talks to the network.
|
|
49
|
+
talks to the network. Something else has to read the list and act on it: normally
|
|
22
50
|
[`agenttrail-guard`](https://github.com/agenttrailhq/guard), which runs on your laptop and checks each
|
|
23
51
|
command an agent proposes against these rules.
|
|
24
52
|
|
|
25
|
-
|
|
26
|
-
|
|
53
|
+
## Why these rules?
|
|
54
|
+
|
|
55
|
+
- **Readable.** Every rule is plain data with a title, a severity, a default action, and a
|
|
56
|
+
description that says what it catches **and what it misses**.
|
|
57
|
+
- **Proven in both directions.** Every rule ships commands it must catch and near-misses it must
|
|
58
|
+
leave alone, and a shared corpus of everyday commands no rule may match. A rule that fails
|
|
59
|
+
either side does not build.
|
|
60
|
+
- **Quiet on purpose.** A commit message or a `grep` that only *mentions* `rm -rf /` is not
|
|
61
|
+
treated as running it. See [Talking about a command is not running it](#talking-about-a-command-is-not-running-it).
|
|
62
|
+
- **Mapped to known risks.** The packs line up with the
|
|
63
|
+
[OWASP Top 10s for agentic and LLM applications](#owasp-coverage), and the gaps are stated, not
|
|
64
|
+
hidden.
|
|
65
|
+
- **Separate from the enforcer.** You can read, disagree with, and change any rule without
|
|
66
|
+
trusting anything about the tool that enforces it.
|
|
67
|
+
|
|
68
|
+
## Quick start
|
|
69
|
+
|
|
70
|
+
**You need:** Node.js 22.13 or newer.
|
|
71
|
+
|
|
72
|
+
### Enforce the rules in your agent
|
|
73
|
+
|
|
74
|
+
Most people never install this package directly. [AgentTrail Guard](https://github.com/agenttrailhq/guard)
|
|
75
|
+
bundles it, so installing Guard is all it takes:
|
|
76
|
+
|
|
77
|
+
```bash
|
|
78
|
+
npm install -g @agenttrail/guard
|
|
79
|
+
agenttrail-guard init --agent claude # or: --agent cursor, --agent codex
|
|
80
|
+
agenttrail-guard guardrails list # every rule and its current action
|
|
81
|
+
agenttrail-guard guardrails show wt.reset-hard # one rule in full
|
|
82
|
+
agenttrail-guard guardrails set-action <id> warn # change what a rule does
|
|
83
|
+
agenttrail-guard guardrails allow <id> <pattern> # silence one rule for one command shape
|
|
84
|
+
```
|
|
27
85
|
|
|
28
|
-
|
|
86
|
+
See the [AgentTrail Guard README](https://github.com/agenttrailhq/guard#quick-start) for setup
|
|
87
|
+
and everything else Guard can do.
|
|
29
88
|
|
|
30
|
-
|
|
31
|
-
your own from scratch.
|
|
32
|
-
- **You want to see exactly what is being blocked and why.** Every rule is plain data with a written
|
|
33
|
-
explanation, including what it *misses*.
|
|
34
|
-
- **You want to contribute a rule** you wish had existed. See [Contributing](#contributing-a-rule).
|
|
89
|
+
### Use the rules in your own code
|
|
35
90
|
|
|
36
|
-
|
|
91
|
+
Install the package if you are writing rules, or building your own tool on top of the list:
|
|
37
92
|
|
|
38
93
|
```bash
|
|
39
94
|
npm install @agenttrail/guardrails
|
|
40
95
|
```
|
|
41
96
|
|
|
42
|
-
|
|
43
|
-
|
|
97
|
+
```ts
|
|
98
|
+
import { RULES, getRule, rulesForPack } from "@agenttrail/guardrails/guardrails";
|
|
99
|
+
|
|
100
|
+
// Everything that refuses outright, rather than asking or warning.
|
|
101
|
+
const blocking = RULES.filter((rule) => rule.defaultAction === "block");
|
|
102
|
+
console.log(`${RULES.length} rules, ${blocking.length} block by default`);
|
|
103
|
+
|
|
104
|
+
// One pack, with each rule's severity and default action.
|
|
105
|
+
for (const rule of rulesForPack("secret-exposure")) {
|
|
106
|
+
console.log(rule.id, rule.severity, rule.defaultAction);
|
|
107
|
+
}
|
|
108
|
+
|
|
109
|
+
// One rule, including the limits its description admits to.
|
|
110
|
+
console.log(getRule("wt.reset-hard")?.description);
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
This package has no evaluator: it tells you what each rule looks for, not whether a given
|
|
114
|
+
command matches. Guard does the matching. [Two ways to import it](#two-ways-to-import-it)
|
|
115
|
+
explains the two entry points.
|
|
44
116
|
|
|
45
117
|
## A rule, start to finish
|
|
46
118
|
|
|
@@ -53,7 +125,7 @@ Here is a complete rule. Nothing is hidden; this is the actual shape.
|
|
|
53
125
|
"severity": "high",
|
|
54
126
|
"defaultAction": "block",
|
|
55
127
|
"title": "git reset --hard discards uncommitted work",
|
|
56
|
-
"description": "Discards all uncommitted changes. Does not match `git restore`
|
|
128
|
+
"description": "Discards all uncommitted changes. Does not match `git restore` (see wt.restore-path).",
|
|
57
129
|
|
|
58
130
|
// What to look for. This one matches a Bash command against a regular expression.
|
|
59
131
|
"match": {
|
|
@@ -62,7 +134,7 @@ Here is a complete rule. Nothing is hidden; this is the actual shape.
|
|
|
62
134
|
]
|
|
63
135
|
},
|
|
64
136
|
|
|
65
|
-
// Proof it works, in both directions
|
|
137
|
+
// Proof it works, in both directions (see "Every rule proves both directions" below).
|
|
66
138
|
"fixtures": {
|
|
67
139
|
"block": ["git reset --hard"],
|
|
68
140
|
"allow": ["git reset src/api.ts"]
|
|
@@ -70,12 +142,17 @@ Here is a complete rule. Nothing is hidden; this is the actual shape.
|
|
|
70
142
|
}
|
|
71
143
|
```
|
|
72
144
|
|
|
145
|
+
The real rule is in [`wt-reset-hard.ts`](src/packs/working-tree/wt-reset-hard.ts); this is a
|
|
146
|
+
trimmed copy. With it enforced by Guard, an agent that proposes `git reset --hard` is refused
|
|
147
|
+
and told which rule stopped it, while `git reset src/api.ts`, which unstages one file, goes
|
|
148
|
+
ahead as normal.
|
|
149
|
+
|
|
73
150
|
Reading the fields:
|
|
74
151
|
|
|
75
152
|
| Field | What it means |
|
|
76
153
|
|---|---|
|
|
77
154
|
| `id` | A stable name. Users type it to disable or change a rule, so it never changes. |
|
|
78
|
-
| `severity` | How bad the thing being caught is: `critical`, `high`, `medium`, `low`, `info`. **It is not a price
|
|
155
|
+
| `severity` | How bad the thing being caught is: `critical`, `high`, `medium`, `low`, `info`. **It is not a price**: this package ships no mapping from severity to money. |
|
|
79
156
|
| `defaultAction` | What should happen: `block` (refuse), `require_approval` (ask a human), or `warn` (allow, but say so). A user can override it. |
|
|
80
157
|
| `description` | What the rule catches **and what it misses**. The honest limits are part of the rule, not a footnote. |
|
|
81
158
|
| `match` | The condition. `any_of` means "any one of these is enough". |
|
|
@@ -86,26 +163,64 @@ Reading the fields:
|
|
|
86
163
|
A rule is filed by **the harm it prevents**, never by the technique it uses to spot it.
|
|
87
164
|
|
|
88
165
|
That sounds like a detail and is not. The three rules about production config, `.env` files and API
|
|
89
|
-
endpoints all work by matching file paths
|
|
166
|
+
endpoints all work by matching file paths, but they are *not* in `file-scope`. Someone who turned
|
|
90
167
|
that pack off to stop path noise would otherwise silently lose their production and secret
|
|
91
168
|
protection, which they never asked to turn off and would not know they had.
|
|
92
169
|
|
|
93
170
|
| Pack | Rules | What it is about |
|
|
94
171
|
|---|---:|---|
|
|
95
|
-
| `working-tree` | 9 | Destroying uncommitted work or published history
|
|
96
|
-
| `destructive-data` | 8 | Data git cannot bring back
|
|
97
|
-
| `prod-infra` | 8 | Changing running infrastructure
|
|
98
|
-
| `secret-exposure` | 10 | Credentials and sensitive data leaving where they live. Mostly `warn`: reading a secret is a normal part of a normal day. |
|
|
99
|
-
| `rce-supply-chain` | 6 | Running code nobody reviewed
|
|
100
|
-
| `safety-bypass` | 7 | Turning off a check somebody installed on purpose, or erasing the record of it
|
|
101
|
-
| `privilege-supply-chain` | 6 | Gaining reach or handing it out
|
|
102
|
-
| `file-scope` | 4 | The agent wrote somewhere it had no business writing
|
|
103
|
-
| `agent-context` | 6 | The agent changing what it is or what it knows
|
|
104
|
-
| `test-integrity` | 6 | The agent making its work look successful
|
|
105
|
-
| `exfiltration` | 4 | Moving data off the machine or opening a way in
|
|
172
|
+
| [`working-tree`](src/packs/working-tree/) | 9 | Destroying uncommitted work or published history: `git reset --hard`, `git clean -fd`, force-push, `rm -rf`. |
|
|
173
|
+
| [`destructive-data`](src/packs/destructive-data/) | 8 | Data git cannot bring back: a dropped volume, a dropped database, destructive DDL, a deleted shadow copy. |
|
|
174
|
+
| [`prod-infra`](src/packs/prod-infra/) | 8 | Changing running infrastructure: Terraform, Kubernetes, Helm, cloud deletes, a deploy that names production. |
|
|
175
|
+
| [`secret-exposure`](src/packs/secret-exposure/) | 10 | Credentials and sensitive data leaving where they live. Mostly `warn`: reading a secret is a normal part of a normal day. |
|
|
176
|
+
| [`rce-supply-chain`](src/packs/rce-supply-chain/) | 6 | Running code nobody reviewed: pipe-to-shell, a remote runner, a redirected registry, TLS verification off. |
|
|
177
|
+
| [`safety-bypass`](src/packs/safety-bypass/) | 7 | Turning off a check somebody installed on purpose, or erasing the record of it: `--no-verify`, `--admin` merge, hooks disabled, host-key checking off, history and log purges, forged terminal output. |
|
|
178
|
+
| [`privilege-supply-chain`](src/packs/privilege-supply-chain/) | 6 | Gaining reach or handing it out: `sudo` writes, `chmod 777`, IAM grants, persistence, publishing, new dependencies. |
|
|
179
|
+
| [`file-scope`](src/packs/file-scope/) | 4 | The agent wrote somewhere it had no business writing: its own config, the machine, git's internals, the CI definition. |
|
|
180
|
+
| [`agent-context`](src/packs/agent-context/) | 6 | The agent changing what it is or what it knows (its standing instructions, its memory, its skills and commands, its MCP servers), or starting more agents, or switching another agent's approvals off. |
|
|
181
|
+
| [`test-integrity`](src/packs/test-integrity/) | 6 | The agent making its work look successful: deleting a test, weakening a runner's configuration, accepting every snapshot, switching a coverage gate off, silencing failures in bulk, or telling CI not to run. |
|
|
182
|
+
| [`exfiltration`](src/packs/exfiltration/) | 4 | Moving data off the machine or opening a way in: a reverse shell, a public tunnel, a file upload, a paste service. Command channel only. |
|
|
106
183
|
|
|
107
184
|
Pack names appear in user config files, so renaming one is a breaking change, not a tidy-up.
|
|
108
185
|
|
|
186
|
+
Each pack name links to its source folder: one TypeScript file per rule, so you can read any rule
|
|
187
|
+
end to end. Every rule, with its full description and tested examples, is also in the
|
|
188
|
+
[guardrail library](https://www.agenttrail.sh/guardrails) on the AgentTrail website, where you
|
|
189
|
+
can paste a command to see which rules match it.
|
|
190
|
+
|
|
191
|
+
## OWASP coverage
|
|
192
|
+
|
|
193
|
+
Several packs exist because of risks named in OWASP's two lists for AI systems: the
|
|
194
|
+
[Top 10 for Agentic Applications](https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/)
|
|
195
|
+
(ASI01 to ASI10, 2026 edition) and the
|
|
196
|
+
[Top 10 for LLM Applications](https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/)
|
|
197
|
+
(LLM01 to LLM10, 2026 edition). A rule cannot see a prompt or a model's reasoning, only the
|
|
198
|
+
action that comes out of it, so these rules address a risk at the point where it turns into a
|
|
199
|
+
command or a file change.
|
|
200
|
+
|
|
201
|
+
| Pack | Agentic risks | LLM risks |
|
|
202
|
+
|---|---|---|
|
|
203
|
+
| `working-tree`, `destructive-data`, `prod-infra`, `file-scope` | ASI02 Tool Misuse and Exploitation | LLM03 Excessive Agency |
|
|
204
|
+
| `rce-supply-chain` | ASI05 Unexpected Code Execution, ASI04 Agentic Supply Chain | LLM04 Supply Chain |
|
|
205
|
+
| `privilege-supply-chain` | ASI03 Identity and Privilege Abuse, ASI04 Agentic Supply Chain | LLM04 Supply Chain |
|
|
206
|
+
| `secret-exposure`, `exfiltration` | ASI02 Tool Misuse and Exploitation, ASI03 Identity and Privilege Abuse | LLM02 Sensitive Information Disclosure |
|
|
207
|
+
| `safety-bypass` | ASI09 Human-Agent Trust Exploitation | LLM10 Improper Output Handling (`gb.ansi-terminal-forgery`) |
|
|
208
|
+
| `agent-context` | ASI06 Memory and Context Poisoning (the write, not its contents), ASI10 Rogue Agents | LLM03 Excessive Agency |
|
|
209
|
+
| `test-integrity` | ASI10 Rogue Agents, ASI09 Human-Agent Trust Exploitation | |
|
|
210
|
+
|
|
211
|
+
What the rules cannot address, by design:
|
|
212
|
+
|
|
213
|
+
- **Prompt injection and goal hijack (LLM01, ASI01).** The rules never see the input that
|
|
214
|
+
redirected the agent. They can only stop the action it led to.
|
|
215
|
+
- **Anything that needs memory across commands (LLM06, ASI08).** A rule sees one command at a
|
|
216
|
+
time, so it cannot count, rate-limit, or spot a runaway loop.
|
|
217
|
+
- **Model, training and retrieval risks (LLM05, LLM08, LLM09), and traffic between agents
|
|
218
|
+
(ASI07).** These happen somewhere a command rule never reaches.
|
|
219
|
+
|
|
220
|
+
For the item-by-item matrix, including how the rest of AgentTrail covers each risk, see the
|
|
221
|
+
[AgentTrail OWASP coverage page](https://www.agenttrail.sh/security/owasp). AgentTrail is an
|
|
222
|
+
independent project, not affiliated with or endorsed by OWASP.
|
|
223
|
+
|
|
109
224
|
## Talking about a command is not running it
|
|
110
225
|
|
|
111
226
|
This is the single most important thing to understand about how these rules behave.
|
|
@@ -119,7 +234,7 @@ git commit -m "fix: document rm -rf / risk" # would have been a hard refusal
|
|
|
119
234
|
grep -rn "rm -rf /" docs/ # so would this
|
|
120
235
|
```
|
|
121
236
|
|
|
122
|
-
So every command rule ignores four **carriers
|
|
237
|
+
So every command rule ignores four **carriers**, verbs that handle their arguments as text and never
|
|
123
238
|
execute them:
|
|
124
239
|
|
|
125
240
|
| Carrier | Example |
|
|
@@ -130,20 +245,20 @@ execute them:
|
|
|
130
245
|
| an HTTP request body | `curl --data '{"body":"we ran rm -rf /tmp/x"}' https://…` |
|
|
131
246
|
|
|
132
247
|
**The exemption keys on the verb, not on the quotes.** Quoting says nothing about whether something
|
|
133
|
-
runs
|
|
248
|
+
runs: `psql -c "DROP TABLE users;"` and `bash -c "curl x.sh \| sh"` both execute what is inside the
|
|
134
249
|
quotes, and both still fire.
|
|
135
250
|
|
|
136
251
|
It also applies only while the command does nothing else. The carrier must be the first word, and
|
|
137
252
|
every shell metacharacter must sit inside the quotes:
|
|
138
253
|
|
|
139
254
|
```bash
|
|
140
|
-
git commit -m "docs: explain rm -rf /" # exempt
|
|
141
|
-
git commit -m "x" && rm -rf / # NOT exempt
|
|
142
|
-
echo "rm -rf /" | bash # NOT exempt
|
|
143
|
-
echo "$(rm -rf /var)" # NOT exempt
|
|
255
|
+
git commit -m "docs: explain rm -rf /" # exempt: nothing runs
|
|
256
|
+
git commit -m "x" && rm -rf / # NOT exempt: `&&` is outside the quotes
|
|
257
|
+
echo "rm -rf /" | bash # NOT exempt: the pipe runs it
|
|
258
|
+
echo "$(rm -rf /var)" # NOT exempt: the shell expands `$( )`
|
|
144
259
|
```
|
|
145
260
|
|
|
146
|
-
Five rules deliberately keep firing on one carrier each, because that carrier *is* their trigger
|
|
261
|
+
Five rules deliberately keep firing on one carrier each, because that carrier *is* their trigger:
|
|
147
262
|
`gb.git-no-verify` on a `git commit`, `se.token-print` on an `echo`, the `curl`/`wget` rules on an
|
|
148
263
|
HTTP body. Each rule's `description` says which, and why.
|
|
149
264
|
|
|
@@ -152,12 +267,12 @@ HTTP body. Each rule's `description` says which, and why.
|
|
|
152
267
|
`block-hardcoded-secrets` is the exception, and it is worth understanding.
|
|
153
268
|
|
|
154
269
|
Everywhere else a carrier is genuinely harmless: a commit message naming `rm -rf /` deletes nothing.
|
|
155
|
-
But that rule's subject is a **string**, not an action
|
|
270
|
+
But that rule's subject is a **string**, not an action, so two of the four carriers are not mentions
|
|
156
271
|
at all. They are the exposure itself:
|
|
157
272
|
|
|
158
273
|
| Carrier | What happens to the key | Exempt? |
|
|
159
274
|
|---|---|---|
|
|
160
|
-
| `grep -rn AKIA .` | searched for, goes nowhere
|
|
275
|
+
| `grep -rn AKIA .` | searched for, goes nowhere, and this is how you find a key to rotate | **yes** |
|
|
161
276
|
| `echo "AKIA…"` | transient terminal output | **yes** |
|
|
162
277
|
| `git commit -m "…AKIA…"` | written into history, then pushed | **no** |
|
|
163
278
|
| `curl --data "…AKIA…"` | sent to a remote host | **no** |
|
|
@@ -165,26 +280,8 @@ at all. They are the exposure itself:
|
|
|
165
280
|
The cost, stated in the other direction: documenting a real-looking key in a commit message is still
|
|
166
281
|
blocked. Redact the body of the key, or use a placeholder short enough to fail the length check.
|
|
167
282
|
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
meaning, but a runner prefix is not: `pnpm exec rg …`, `npx …` and `xargs -0 grep …` still fire,
|
|
171
|
-
since "some program eventually runs a search" is a much weaker claim than "this command is a search".
|
|
172
|
-
At most four quoted arguments are recognised. A carrier that can be made to execute through a flag —
|
|
173
|
-
`ack --pager='…'`, `rg --pre <cmd>` — is still treated as a mention; closing that needs information
|
|
174
|
-
the checker does not have. And an MCP tool whose input carries the same text is not exempt either,
|
|
175
|
-
because exempting a JSON blob would exempt a shell-running MCP server along with it.
|
|
176
|
-
|
|
177
|
-
**MCP coverage.** A command rule fires on `Bash`, `PowerShell` *and* any `mcp__*` tool: the guard hands
|
|
178
|
-
the checker an MCP call's serialized `tool_input` as the same command text every command rule reads, so
|
|
179
|
-
a command shape run through an MCP server — `{"command":"rm -rf /"}` — is caught, not ignored. Two
|
|
180
|
-
honest limits follow from that. First, a rule whose pattern is anchored to the start of the command
|
|
181
|
-
(`^…` or a command-position class) may not fire inside the JSON, where the shape sits after a `"`
|
|
182
|
-
rather than at a command boundary; the `\b`-anchored rules — most of the corpus — do fire. Second, the
|
|
183
|
-
quoted-mention exemptions are shell-only, so an MCP payload that merely *names* a command in a text
|
|
184
|
-
field (`{"title":"fix the rm -rf / bug"}`) is matched the same as one that runs it — a JSON blob cannot
|
|
185
|
-
be told apart from a shell-running MCP server. File rules match by path on whichever file tool a client
|
|
186
|
-
uses. No rule is shell-only by design; a rule that does not reach the MCP channel does so because its
|
|
187
|
-
pattern, not its label, does not match the serialized shape.
|
|
283
|
+
The finer limits of the carrier logic, and how command rules apply to MCP tools, are in the
|
|
284
|
+
[Reference](#reference).
|
|
188
285
|
|
|
189
286
|
## What these rules deliberately do not catch
|
|
190
287
|
|
|
@@ -193,29 +290,29 @@ somebody forgot.
|
|
|
193
290
|
|
|
194
291
|
- **Nothing about the web.** Pages an agent fetches are not checked, and there are no URL rules.
|
|
195
292
|
- **Nothing inside a file.** The checker sees a file's *path*, never its contents. A secret typed into
|
|
196
|
-
a source file, SQL built by string concatenation, a missing auth check
|
|
293
|
+
a source file, SQL built by string concatenation, a missing auth check: none of it is visible.
|
|
197
294
|
Rules that would need it are absent rather than approximated.
|
|
198
295
|
- **Nothing about where you are.** No working directory, no project root, no git branch, no cloud
|
|
199
|
-
profile reaches the checker
|
|
296
|
+
profile reaches the checker; it gets one command and nothing else. So "the agent wrote outside the
|
|
200
297
|
project" **cannot be written as a rule**, and `file-scope` is limited to well-known absolute paths
|
|
201
298
|
for good. For the same reason, a rule cannot tell a scratch database from a production one.
|
|
202
|
-
- **Nothing hidden inside a quoted payload.** Where the danger is inside a quoted argument
|
|
203
|
-
`psql -c "<sql>"`, `python -c "<code>"`
|
|
204
|
-
very little about what the script actually does.
|
|
299
|
+
- **Nothing hidden inside a quoted payload.** Where the danger is inside a quoted argument
|
|
300
|
+
(`psql -c "<sql>"`, `python -c "<code>"`), a text rule can only guess. In a long script, a match
|
|
301
|
+
says very little about what the script actually does.
|
|
205
302
|
- **Nothing a wrapper hides.** `./deploy.sh` that runs `terraform apply -auto-approve` inside it is
|
|
206
303
|
just a shell script from the outside.
|
|
207
304
|
- **Nothing recurring.** There is no counting. "The same mistake three times this week" needs memory
|
|
208
305
|
across commands, and a single command has none.
|
|
209
306
|
|
|
210
307
|
Each rule's own `description` names its specific misses. Read those before trusting a rule to cover a
|
|
211
|
-
case
|
|
308
|
+
case; they are written to be believed, not to sell.
|
|
212
309
|
|
|
213
310
|
## Every rule proves both directions
|
|
214
311
|
|
|
215
312
|
Every rule ships at least one **`block`** example and at least one **`allow`** example. A rule missing
|
|
216
313
|
either does not build.
|
|
217
314
|
|
|
218
|
-
- **`block` means "this rule must match."** It does *not* mean the agent is refused
|
|
315
|
+
- **`block` means "this rule must match."** It does *not* mean the agent is refused; most packs
|
|
219
316
|
default to asking or warning.
|
|
220
317
|
- **`allow` means "this rule must NOT match."** *This is the half that matters.* Anyone can write a
|
|
221
318
|
rule that catches `rm -rf /`. The hard part is not firing on `rm -rf ./node_modules` forty times a
|
|
@@ -228,13 +325,13 @@ Examples come in two kinds, and each must use the right one:
|
|
|
228
325
|
| a command | `Bash`, `PowerShell`, a search query, an MCP tool's input |
|
|
229
326
|
| a file path | `Edit`, `Write`, `Read`, `MultiEdit`, `NotebookEdit` |
|
|
230
327
|
|
|
231
|
-
Giving a path-matching rule a command example makes it pass **without testing anything
|
|
328
|
+
Giving a path-matching rule a command example makes it pass **without testing anything**: it matches
|
|
232
329
|
nothing, which reads as proof of quietness and proves only that the path never reached the rule. CI
|
|
233
330
|
rejects that.
|
|
234
331
|
|
|
235
332
|
## Contributing a rule
|
|
236
333
|
|
|
237
|
-
Rules are meant to be contributed. The bar is not "clever regex"
|
|
334
|
+
Rules are meant to be contributed. The bar is not "clever regex"; it is **does it fire on the real
|
|
238
335
|
thing, and stay quiet on the near-miss**.
|
|
239
336
|
|
|
240
337
|
**1. Write it**, following the shape above. Give it a `description` that says what it misses.
|
|
@@ -258,35 +355,79 @@ agenttrail-guard guardrails validate ./my-rule.json
|
|
|
258
355
|
does **not** answer *"does it actually fire on the command I think it does?"* That needs the real
|
|
259
356
|
checker, which is not part of this package.
|
|
260
357
|
|
|
261
|
-
**So the real test runs in CI, on your pull request
|
|
358
|
+
**So the real test runs in CI, on your pull request**: the same check, on the same machine, for
|
|
262
359
|
everyone. You get the shape check instantly here and the real answer there, which is where it has to
|
|
263
360
|
run to be trusted anyway.
|
|
264
361
|
|
|
265
|
-
Your rule is also run against a **quiet corpus**:
|
|
266
|
-
match at all. Your own negative example only proves your rule is quiet on the near-miss
|
|
267
|
-
of. The quiet corpus is what catches a Terraform rule firing on `pnpm test`.
|
|
362
|
+
Your rule is also run against a **quiet corpus**: hundreds of everyday commands and paths that no
|
|
363
|
+
rule may match at all. Your own negative example only proves your rule is quiet on the near-miss
|
|
364
|
+
*you* thought of. The quiet corpus is what catches a Terraform rule firing on `pnpm test`.
|
|
268
365
|
|
|
269
366
|
### Shapes that will not validate
|
|
270
367
|
|
|
271
|
-
Three are rejected outright, each because it produces a rule that *looks* enforced and is not
|
|
368
|
+
Three are rejected outright, each because it produces a rule that *looks* enforced and is not: the
|
|
272
369
|
worst failure a security tool can have.
|
|
273
370
|
|
|
274
371
|
| Rejected | Why |
|
|
275
372
|
|---|---|
|
|
276
|
-
| `scope` | It compares against ids that are always UUIDs, never a vendor name
|
|
373
|
+
| `scope` | It compares against ids that are always UUIDs, never a vendor name, so a scoped rule matches nothing, forever, silently. |
|
|
277
374
|
| Numeric conditions | Token counts and durations are all zero *before* a command runs. "Greater than" can never fire; "less than" fires on everything. |
|
|
278
375
|
| A command matcher and a file matcher in one condition | No real command carries both, so the condition can never be true. Split it into two under `any_of`. |
|
|
279
376
|
|
|
280
377
|
A regular expression is also rejected if it nests unbounded repetition (`(a+)+`). The guard fails
|
|
281
|
-
**open** under a time limit, so a pattern that backtracks does not merely run slowly
|
|
378
|
+
**open** under a time limit, so a pattern that backtracks does not merely run slowly; it lets the
|
|
282
379
|
command through.
|
|
283
380
|
|
|
284
381
|
One more is caught in CI rather than by the shape check, because it cannot be caught earlier: **a
|
|
285
|
-
single-item brace list in a tool name.** `"{Bash}"` is a glob pattern, and it does not match `Bash
|
|
382
|
+
single-item brace list in a tool name.** `"{Bash}"` is a glob pattern, and it does not match `Bash`,
|
|
286
383
|
so the rule matches nothing, forever, with no error anywhere. Write a single tool plainly as
|
|
287
384
|
`"Bash"`; braces are for real alternatives, `"{Bash,PowerShell}"`.
|
|
288
385
|
|
|
289
|
-
|
|
386
|
+
### Ready to open a pull request?
|
|
387
|
+
|
|
388
|
+
[CONTRIBUTING.md](CONTRIBUTING.md) has a pre-submission checklist, local build and test
|
|
389
|
+
commands, and where to report what. Please follow the [Code of Conduct](CODE_OF_CONDUCT.md),
|
|
390
|
+
and report vulnerabilities privately as described in [SECURITY.md](SECURITY.md).
|
|
391
|
+
|
|
392
|
+
## Stability and versioning
|
|
393
|
+
|
|
394
|
+
Rule ids and pack names end up in people's config files, so they are treated as public API:
|
|
395
|
+
|
|
396
|
+
- **A rule id never changes.** Renaming or removing a rule would silently undo an override
|
|
397
|
+
someone set on it, so either one is a breaking change.
|
|
398
|
+
- **A pack name never changes** either, for the same reason: a renamed pack silently
|
|
399
|
+
re-enables a pack someone had turned off.
|
|
400
|
+
- **A new rule or pack is a feature.** Guard turns every pack on unless you turned it off, so a
|
|
401
|
+
rule added in a release you install starts enforcing with no step of your own.
|
|
402
|
+
|
|
403
|
+
Releases are cut automatically from [Conventional Commits](https://www.conventionalcommits.org/).
|
|
404
|
+
Before 1.0, a breaking change ships as a minor release and a new rule as a patch release, so
|
|
405
|
+
skim the [changelog](CHANGELOG.md) before you upgrade.
|
|
406
|
+
|
|
407
|
+
## FAQ
|
|
408
|
+
|
|
409
|
+
**Can I use these rules without Guard?** Yes. They are plain data under Apache-2.0. You need your
|
|
410
|
+
own matcher to apply them, because this package does not contain one.
|
|
411
|
+
|
|
412
|
+
**A rule fires on something legitimate. What do I do?** In Guard, `agenttrail-guard guardrails
|
|
413
|
+
allow <id> <pattern>` silences that one rule for that one command shape, without turning the rule
|
|
414
|
+
off. Then please [open an issue](https://github.com/agenttrailhq/guardrails/issues) with a
|
|
415
|
+
sanitized example, so the rule can learn the near-miss.
|
|
416
|
+
|
|
417
|
+
**Why didn't a rule fire on my command?** Start with the rule's `description`: it lists what the
|
|
418
|
+
rule misses. Then check [what these rules deliberately do not catch](#what-these-rules-deliberately-do-not-catch).
|
|
419
|
+
To test a command against every rule, paste it into the [online checker](https://www.agenttrail.sh/guardrails#check).
|
|
420
|
+
|
|
421
|
+
**Why is `secret-exposure` mostly `warn`?** Reading a secret is a normal part of a normal day.
|
|
422
|
+
Refusing every `.env` read would be turned off within a week. Moving a secret off the machine is
|
|
423
|
+
held for approval instead.
|
|
424
|
+
|
|
425
|
+
**How do I turn off a whole pack?** `agenttrail-guard guardrails disable <pack>`. See
|
|
426
|
+
[Tune a guardrail](https://github.com/agenttrailhq/guard#tune-a-guardrail) in the Guard README.
|
|
427
|
+
|
|
428
|
+
## Reference
|
|
429
|
+
|
|
430
|
+
### Two ways to import it
|
|
290
431
|
|
|
291
432
|
```ts
|
|
292
433
|
import { RULES, getRule } from "@agenttrail/guardrails/guardrails"; // just the rules
|
|
@@ -300,16 +441,62 @@ They are separate because the guard starts a fresh process on **every single com
|
|
|
300
441
|
under a ten-second ceiling. It cannot afford to load a validator it never calls, or to re-check 74
|
|
301
442
|
rules that were already checked before release.
|
|
302
443
|
|
|
303
|
-
|
|
444
|
+
### Known limits of the carrier logic
|
|
445
|
+
|
|
446
|
+
In the same spirit as the rest of this file:
|
|
447
|
+
|
|
448
|
+
- **The carrier must be the first word.** A single leading `sudo` is tolerated, because it changes
|
|
449
|
+
privilege rather than meaning, but a runner prefix is not: `pnpm exec rg …`, `npx …` and
|
|
450
|
+
`xargs -0 grep …` still fire, since "some program eventually runs a search" is a much weaker claim
|
|
451
|
+
than "this command is a search".
|
|
452
|
+
- **At most four quoted arguments are recognised.**
|
|
453
|
+
- **A carrier that can be made to execute through a flag** (`ack --pager='…'`, `rg --pre <cmd>`) is
|
|
454
|
+
still treated as a mention; closing that needs information the checker does not have.
|
|
455
|
+
- **An MCP tool whose input carries the same text is not exempt either,** because exempting a JSON
|
|
456
|
+
blob would exempt a shell-running MCP server along with it.
|
|
457
|
+
|
|
458
|
+
### MCP coverage
|
|
459
|
+
|
|
460
|
+
A command rule fires on `Bash`, `PowerShell` *and* any `mcp__*` tool: the guard hands the checker an
|
|
461
|
+
MCP call's serialized `tool_input` as the same command text every command rule reads, so a command
|
|
462
|
+
shape run through an MCP server (`{"command":"rm -rf /"}`) is caught, not ignored. Two honest limits
|
|
463
|
+
follow from that:
|
|
464
|
+
|
|
465
|
+
1. **Start-anchored patterns may miss.** A rule whose pattern is anchored to the start of the command
|
|
466
|
+
(`^…` or a command-position class) may not fire inside the JSON, where the shape sits after a `"`
|
|
467
|
+
rather than at a command boundary; the `\b`-anchored rules (most of the corpus) do fire.
|
|
468
|
+
2. **Mentions are not exempt.** The quoted-mention exemptions are shell-only, so an MCP payload that
|
|
469
|
+
merely *names* a command in a text field (`{"title":"fix the rm -rf / bug"}`) is matched the same
|
|
470
|
+
as one that runs it; a JSON blob cannot be told apart from a shell-running MCP server.
|
|
471
|
+
|
|
472
|
+
File rules match by path on whichever file tool a client uses. No rule is shell-only by design; a
|
|
473
|
+
rule that does not reach the MCP channel does so because its pattern, not its label, does not match
|
|
474
|
+
the serialized shape.
|
|
475
|
+
|
|
476
|
+
### Where these rules came from
|
|
304
477
|
|
|
305
478
|
Independently authored. No block list, pattern or wording is copied from any other project.
|
|
306
479
|
|
|
307
|
-
Where a rule's shape follows an obvious convention
|
|
308
|
-
pattern
|
|
480
|
+
Where a rule's shape follows an obvious convention (an `rm -rf` pattern looks like an `rm -rf`
|
|
481
|
+
pattern), that is two people meeting the same shell, not one copying the other.
|
|
482
|
+
|
|
483
|
+
## Part of AgentTrail
|
|
309
484
|
|
|
310
|
-
|
|
311
|
-
|
|
485
|
+
These rules are part of [AgentTrail](https://www.agenttrail.sh), which builds guardrails and
|
|
486
|
+
observability for AI coding agents:
|
|
487
|
+
|
|
488
|
+
- **[AgentTrail Guard](https://www.agenttrail.sh/agenttrail-guard)** enforces them on every
|
|
489
|
+
tool call that reaches its hooks, on your machine. It is free, open source, and needs no
|
|
490
|
+
account. Its code is in [agenttrailhq/guard](https://github.com/agenttrailhq/guard).
|
|
491
|
+
- **[AgentTrail OS](https://www.agenttrail.sh/agenttrail-os)** uses the same open rule library,
|
|
492
|
+
and adds searchable session history and team workflows. See
|
|
493
|
+
[AgentTrail Guard vs AgentTrail OS](https://www.agenttrail.sh/guard-vs-os).
|
|
312
494
|
|
|
313
495
|
## License
|
|
314
496
|
|
|
315
497
|
Apache-2.0. See [LICENSE](./LICENSE).
|
|
498
|
+
|
|
499
|
+
Rule ideas and reports go to [Guardrails issues](https://github.com/agenttrailhq/guardrails/issues).
|
|
500
|
+
If you rely on these rules, a star helps other developers find them.
|
|
501
|
+
|
|
502
|
+
Made with ❤️ by [agenttrail.sh](https://www.agenttrail.sh)
|
|
@@ -3651,8 +3651,8 @@ function isPack(value) {
|
|
|
3651
3651
|
}
|
|
3652
3652
|
|
|
3653
3653
|
// src/stamp.ts
|
|
3654
|
-
var CATALOG_VERSION = "0.2.
|
|
3655
|
-
var CATALOG_PUBLISHED_AT = "2026-
|
|
3654
|
+
var CATALOG_VERSION = "0.2.2";
|
|
3655
|
+
var CATALOG_PUBLISHED_AT = "2026-10-06T13:41:00Z";
|
|
3656
3656
|
|
|
3657
3657
|
// src/rules.ts
|
|
3658
3658
|
var RULES_BY_PACK = {
|
|
@@ -3702,4 +3702,4 @@ export {
|
|
|
3702
3702
|
getRule,
|
|
3703
3703
|
rulesForPack
|
|
3704
3704
|
};
|
|
3705
|
-
//# sourceMappingURL=chunk-
|
|
3705
|
+
//# sourceMappingURL=chunk-GNN4TGRY.js.map
|