@alvin0/ai-agent-sdk-sandbox 0.1.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +299 -0
- package/dist/index.d.ts +731 -0
- package/dist/index.d.ts.map +1 -0
- package/dist/index.js +1310 -0
- package/dist/index.js.map +1 -0
- package/package.json +67 -0
package/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 alvin0 (chaulamdinhai) <chaulamdinhai@gmail.com>
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
package/README.md
ADDED
|
@@ -0,0 +1,299 @@
|
|
|
1
|
+
# @alvin0/ai-agent-sdk-sandbox
|
|
2
|
+
|
|
3
|
+
The sandbox contract: the file-effect vocabulary, per-call policy resolution,
|
|
4
|
+
the writable-root algebra every backend shares, and the rules that keep a broken
|
|
5
|
+
sandbox from reading as a denied command.
|
|
6
|
+
|
|
7
|
+
Runtime: **Universal**. This package imports nothing — no `node:` builtins, no
|
|
8
|
+
dependencies — so it runs anywhere the core SDK does. Enforcement lives in
|
|
9
|
+
[`@alvin0/ai-agent-sdk-sandbox-node`](../sandbox-node#readme).
|
|
10
|
+
|
|
11
|
+
```bash
|
|
12
|
+
pnpm add @alvin0/ai-agent-sdk-sandbox
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
## Modes
|
|
16
|
+
|
|
17
|
+
`SandboxMode` governs **file effects only**. Network reachability and process
|
|
18
|
+
visibility are deliberately outside this vocabulary; they belong to their own
|
|
19
|
+
seam, and a mode that claimed to cover them would be lying.
|
|
20
|
+
|
|
21
|
+
| Mode | File effects |
|
|
22
|
+
| --- | --- |
|
|
23
|
+
| `read-only` | No writes anywhere; the backend still permits the `/dev/null` sink a shell needs |
|
|
24
|
+
| `workspace-write` | Writes under the workspace root, plus any temp roots the consumer explicitly grants |
|
|
25
|
+
| `danger-full-access` | No confinement; the provider is never consulted |
|
|
26
|
+
|
|
27
|
+
## The authorization boundary
|
|
28
|
+
|
|
29
|
+
Everything a tool sends is model-authored JSON, so a policy input that widens
|
|
30
|
+
authority is one the model can grant itself. `mode` and `entries` on a request
|
|
31
|
+
are therefore the untrusted half: a requested mode is honoured only when it is
|
|
32
|
+
at least as strict as the session's own, and requested entries are intersected
|
|
33
|
+
with deployment and hardening layers. A `deny` can narrow `read`, and `read` can
|
|
34
|
+
narrow `write`, but neither can reopen a standing denial. A requested `write`
|
|
35
|
+
is refused outright.
|
|
36
|
+
|
|
37
|
+
Widening goes through a capability instead of data:
|
|
38
|
+
|
|
39
|
+
```ts
|
|
40
|
+
import { approveSandboxEscalation, resolveSandboxPolicy } from '@alvin0/ai-agent-sdk-sandbox'
|
|
41
|
+
|
|
42
|
+
// After the host has actually authorized it — a human prompt, a policy engine.
|
|
43
|
+
const approval = approveSandboxEscalation({ mode: 'workspace-write' })
|
|
44
|
+
const policy = resolveSandboxPolicy({ cwd, sessionMode, approval }, defaults)
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
Only a value this module minted is accepted. `JSON.parse('{"approved":true}')`
|
|
48
|
+
is refused, which is what a forged approval arriving in a tool payload looks
|
|
49
|
+
like. Deployment configuration goes in `defaults`, which is the trusted half and
|
|
50
|
+
may grant anything.
|
|
51
|
+
|
|
52
|
+
## Per-call policy
|
|
53
|
+
|
|
54
|
+
Policy is resolved per capability call, never fixed on the provider. Two
|
|
55
|
+
consumers can confine under different boundaries at the same instant, and an
|
|
56
|
+
approved escalation is a *new call with a wider policy* — not a mutation of
|
|
57
|
+
shared state.
|
|
58
|
+
|
|
59
|
+
```ts
|
|
60
|
+
import { confiningPolicy, resolveSandboxPolicy } from '@alvin0/ai-agent-sdk-sandbox'
|
|
61
|
+
|
|
62
|
+
const resolved = resolveSandboxPolicy(
|
|
63
|
+
{ cwd: session.cwd, sessionMode: session.mode, mode: approvedOverride },
|
|
64
|
+
{ mode: 'read-only', workspaceRoot: deploymentRoot },
|
|
65
|
+
)
|
|
66
|
+
|
|
67
|
+
const policy = confiningPolicy(resolved)
|
|
68
|
+
// `undefined` under danger-full-access: spawn the original argv, ask nothing.
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
Precedence is fixed: an approved explicit mode outranks the session's mode,
|
|
72
|
+
which outranks the deployment default. A session's cwd is its `workspace-write`
|
|
73
|
+
boundary; the configured root is the fallback for agentless calls.
|
|
74
|
+
|
|
75
|
+
## Reading a command for what it does
|
|
76
|
+
|
|
77
|
+
The file seam cannot tell `systemctl status nginx` from `systemctl restart
|
|
78
|
+
nginx`: both are argv, neither writes a file the policy governs, and one
|
|
79
|
+
observes while the other changes the machine. `classifyExec` reads the command
|
|
80
|
+
semantically so a harness can decide before any of it runs.
|
|
81
|
+
|
|
82
|
+
```ts
|
|
83
|
+
classifyExec(['systemctl', 'status', 'nginx']) // observe -> allow
|
|
84
|
+
classifyExec(['systemctl', 'restart', 'nginx']) // service-control -> ask-approval
|
|
85
|
+
classifyExec(['aws', 'configure', 'list']) // credential -> deny
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
| Capability | Default outcome |
|
|
89
|
+
| --- | --- |
|
|
90
|
+
| `observe` | `allow` |
|
|
91
|
+
| `use` | `allow-scoped` |
|
|
92
|
+
| `modify`, `service-control`, `package-install`, `privilege` | `ask-approval` |
|
|
93
|
+
| `credential`, `critical` | `deny` |
|
|
94
|
+
| `unknown` | `ask-approval` |
|
|
95
|
+
|
|
96
|
+
This is a classifier, and a classifier is a guess. Two rules keep the guess from
|
|
97
|
+
becoming a hazard: a command it does not recognise is **never allowed**, and a
|
|
98
|
+
command that hides others — a shell string, a pipeline, a chain, an argv with
|
|
99
|
+
separators in it — is decided by the riskiest thing inside it rather than by its
|
|
100
|
+
wrapper.
|
|
101
|
+
|
|
102
|
+
The script is walked one character at a time rather than split with a pattern,
|
|
103
|
+
because `grep -E 'a|b'` puts a separator inside a quoted word: a pattern either
|
|
104
|
+
splits there, inventing a command out of a regex, or refuses to split wherever a
|
|
105
|
+
quote appears. Quotes and escapes are respected, so `echo "hi; rm -rf /etc"` is
|
|
106
|
+
one command printing text while `echo hi && rm -rf /etc` is two, and the second
|
|
107
|
+
decides.
|
|
108
|
+
|
|
109
|
+
What it deliberately does not do: expand variables, resolve `$(...)`, follow a
|
|
110
|
+
script file, or know what an unrecognised binary does. `eval "$CMD"` classifies
|
|
111
|
+
as unrecognised, which asks — it does not read what `$CMD` holds.
|
|
112
|
+
|
|
113
|
+
Two of its rules exist because real model output demanded them. Asked to show
|
|
114
|
+
AWS credentials, a model proposed `aws configure list` and `aws sts
|
|
115
|
+
get-caller-identity` — neither names `~/.aws/credentials`, so a rule matching
|
|
116
|
+
credential *paths* saw nothing while the secret was read inside the tool. And
|
|
117
|
+
`if [ -f package.json ]; then npm test; fi`, read as one argv, names `[` and
|
|
118
|
+
looks like a test while what it runs is the test suite.
|
|
119
|
+
|
|
120
|
+
It decides; it does not enforce. `confine()` and `fence()` are what hold.
|
|
121
|
+
|
|
122
|
+
## Wiring it into an agent
|
|
123
|
+
|
|
124
|
+
This package does not depend on `@alvin0/ai-agent-sdk-core`, and core has no
|
|
125
|
+
sandbox slot. A session already has the two seams needed: an interceptor decides
|
|
126
|
+
and the approval broker asks.
|
|
127
|
+
|
|
128
|
+
```ts
|
|
129
|
+
const sandboxInterceptor: ToolInterceptor = {
|
|
130
|
+
name: 'sandbox:exec',
|
|
131
|
+
before: async (call, next) => {
|
|
132
|
+
const argv = argvOf(call) // undefined for a non-exec tool
|
|
133
|
+
if (argv === undefined) return await next()
|
|
134
|
+
const verdict = classifyExec(argv)
|
|
135
|
+
if (verdict.outcome === 'deny') return { kind: 'deny', reason: verdict.reason }
|
|
136
|
+
if (verdict.outcome !== 'ask-approval') return await next()
|
|
137
|
+
asked.add(call.callId)
|
|
138
|
+
return { kind: 'ask', reason: verdict.reason } // the broker asks a person
|
|
139
|
+
},
|
|
140
|
+
// Reaching `around` for a call that asked IS the answer. Mint the capability
|
|
141
|
+
// here, never from tool arguments.
|
|
142
|
+
around: async (call, next) => {
|
|
143
|
+
if (!asked.delete(call.callId)) return await next()
|
|
144
|
+
granted.set(call.callId, approveSandboxEscalation({ entries: escalationFor(call) }))
|
|
145
|
+
try { return await next() } finally { granted.delete(call.callId) }
|
|
146
|
+
},
|
|
147
|
+
}
|
|
148
|
+
|
|
149
|
+
agent.createSession({ tools: [runCommand], interceptors: [sandboxInterceptor], approvals })
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
The tool body reads the approval by `ctx.callId`, passes it to
|
|
153
|
+
`resolveSandboxPolicy`, and then confines or fences. `classifyExec` decides;
|
|
154
|
+
this package never holds anything.
|
|
155
|
+
|
|
156
|
+
## Network reach
|
|
157
|
+
|
|
158
|
+
`SandboxMode` governs file effects and says so. Reachability is a second,
|
|
159
|
+
independent axis on the same policy, because folding it into the mode would
|
|
160
|
+
make the mode claim something it does not decide — and because the two are
|
|
161
|
+
enforced by different mechanisms, so a host can provide one without the other.
|
|
162
|
+
|
|
163
|
+
| Reach | Meaning |
|
|
164
|
+
| --- | --- |
|
|
165
|
+
| `deny` | No network at all |
|
|
166
|
+
| `loopback` | Loopback only — for a proxy or language server the deployment runs |
|
|
167
|
+
| `allow-all` | Unrestricted (the default, and what this package did before the seam) |
|
|
168
|
+
|
|
169
|
+
```ts
|
|
170
|
+
resolveSandboxPolicy({ cwd, network: 'deny' }, { mode: 'read-only', workspaceRoot, network: 'deny' })
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
Reach narrows exactly like authority does: a request may tighten its own, never
|
|
174
|
+
loosen it, and only a minted approval widens it. A deployment running anything
|
|
175
|
+
untrusted should default to `deny` and widen per call.
|
|
176
|
+
|
|
177
|
+
Unix sockets are **not** governed by this axis — they are filesystem objects,
|
|
178
|
+
so a host daemon socket is closed by a `deny` entry, not by a network mode. The
|
|
179
|
+
two seams compose; neither substitutes for the other.
|
|
180
|
+
|
|
181
|
+
## Deny-list or allow-list
|
|
182
|
+
|
|
183
|
+
By default the host is readable and a policy closes paths one at a time. That
|
|
184
|
+
is a deny-list: it protects what someone remembered to name, and the path nobody
|
|
185
|
+
thought about stays open.
|
|
186
|
+
|
|
187
|
+
```ts
|
|
188
|
+
{ mode: 'read-only', workspaceRoot, baseline: 'deny',
|
|
189
|
+
entries: [{ path: '/var/log/my-api', access: 'read' }] }
|
|
190
|
+
```
|
|
191
|
+
|
|
192
|
+
`baseline: 'deny'` inverts it — nothing is readable until an entry says so —
|
|
193
|
+
which is how a request to investigate one service is scoped to that service's
|
|
194
|
+
logs rather than to every log on the host.
|
|
195
|
+
|
|
196
|
+
It is enforced by `fence(policy)`, not by the kernel profiles, and
|
|
197
|
+
`confine()` refuses it rather than pretending: inverting a mount profile means
|
|
198
|
+
binding only what a program needs, and the set a program needs to start at all
|
|
199
|
+
is specific to an OS build.
|
|
200
|
+
|
|
201
|
+
## Approvals are spent
|
|
202
|
+
|
|
203
|
+
An approval is consumed the first time a policy is resolved with it. A person
|
|
204
|
+
approving "read this file" approved one read, and a grant that survives its own
|
|
205
|
+
operation is a grant nobody is still watching.
|
|
206
|
+
|
|
207
|
+
```ts
|
|
208
|
+
approveSandboxEscalation({ entries: [{ path: '/etc/app/config.yaml', access: 'write' }] })
|
|
209
|
+
approveSandboxEscalation({ mode: 'workspace-write', scope: 'session', expiresAt })
|
|
210
|
+
```
|
|
211
|
+
|
|
212
|
+
Note what the first grant does *not* do: it never mentions a mode, so the
|
|
213
|
+
policy stays `read-only` and exactly one file becomes writable. Raising the mode
|
|
214
|
+
instead would make the whole workspace writable, and the named resource
|
|
215
|
+
decorative.
|
|
216
|
+
|
|
217
|
+
## Nested carve-outs
|
|
218
|
+
|
|
219
|
+
A single writable root is not enough. An agent that may write in a repository
|
|
220
|
+
must still be kept out of `.git`, or it can install a hook that runs arbitrary
|
|
221
|
+
code on the next `git` invocation. Entries express that as an overlapping list
|
|
222
|
+
resolved by **path specificity** — the deepest matching entry wins:
|
|
223
|
+
|
|
224
|
+
```ts
|
|
225
|
+
const entries = [
|
|
226
|
+
{ path: '/repo', access: 'write' },
|
|
227
|
+
{ path: '/repo/a', access: 'deny' },
|
|
228
|
+
{ path: '/repo/a/b', access: 'write' },
|
|
229
|
+
]
|
|
230
|
+
// /repo/x → write · /repo/a/x → deny · /repo/a/b/x → write
|
|
231
|
+
```
|
|
232
|
+
|
|
233
|
+
`grantLayers(policy)` resolves a policy into exactly that: an ordered stack of
|
|
234
|
+
layers, broadest first, each overriding the ones beneath it for its own subtree.
|
|
235
|
+
Order is the semantics, and it is why the model is a stack rather than a pair of
|
|
236
|
+
sets — a set of "granted roots" has nowhere to record a grant that lives *inside*
|
|
237
|
+
something denied, so the third line above would silently vanish.
|
|
238
|
+
|
|
239
|
+
```ts
|
|
240
|
+
grantLayers(policy)
|
|
241
|
+
// write mode /repo
|
|
242
|
+
// read protected /repo/.git ← and .ssh, .aws, .netrc, …
|
|
243
|
+
// deny entry /repo/vendor
|
|
244
|
+
// write entry /repo/vendor/cache
|
|
245
|
+
```
|
|
246
|
+
|
|
247
|
+
`PROTECTED_SUBPATHS` (`.git`, `.ssh`, `.aws`, `.netrc`, …) is layered under every
|
|
248
|
+
granted root automatically. They are `read`, not `deny`: this is a write
|
|
249
|
+
boundary, so they stay readable. An explicit entry at the same depth outranks a
|
|
250
|
+
generated one, so a deployment can deliberately reopen one.
|
|
251
|
+
|
|
252
|
+
A layer that would not change the access already in force is dropped, so the
|
|
253
|
+
result carries no rule that does nothing. `writableRoots(policy)` flattens the
|
|
254
|
+
same layers for callers that only want the two lists.
|
|
255
|
+
|
|
256
|
+
This one function is what every enforcement path reads — the kernel profiles and
|
|
257
|
+
the in-process fence. Deriving them separately is how a profile and a fence
|
|
258
|
+
silently drift into disagreeing about what is writable.
|
|
259
|
+
|
|
260
|
+
## Classification
|
|
261
|
+
|
|
262
|
+
Two failures look identical in a shell and mean opposite things:
|
|
263
|
+
|
|
264
|
+
- **denied** — confinement worked and blocked the command.
|
|
265
|
+
- **runner failure** — the sandbox itself refused or crashed; the command
|
|
266
|
+
*never ran*.
|
|
267
|
+
|
|
268
|
+
Reporting the second as the first sends a model off rewriting correct code.
|
|
269
|
+
`classifyOutcome` checks runner failure first, and never infers it from an exit
|
|
270
|
+
code alone: a rule needs a nonzero exit, its own exit-code gate, and a fatal
|
|
271
|
+
signature on a line that survives informational exclusion.
|
|
272
|
+
|
|
273
|
+
```ts
|
|
274
|
+
import { annotateStderr, classifyOutcome } from '@alvin0/ai-agent-sdk-sandbox'
|
|
275
|
+
|
|
276
|
+
const classification = classifyOutcome(
|
|
277
|
+
{ exitCode: result.status, stderr: result.stderr, signal: result.signal },
|
|
278
|
+
confined, // carries the wrapping backend's own dialect
|
|
279
|
+
)
|
|
280
|
+
const stderr = annotateStderr(result.stderr, classification, policy.mode)
|
|
281
|
+
```
|
|
282
|
+
|
|
283
|
+
Denial signatures are matched against **the backend that actually wrapped the
|
|
284
|
+
command**, never a cross-backend union — a union claims denials a given backend
|
|
285
|
+
never produces. A `SIGSYS` kill is treated as a denial without matching any
|
|
286
|
+
text, because a seccomp kill is unambiguous.
|
|
287
|
+
|
|
288
|
+
## Fail closed
|
|
289
|
+
|
|
290
|
+
`SandboxUnavailableError` (`SANDBOX_UNAVAILABLE`) is thrown when no backend can
|
|
291
|
+
enforce a confining policy. Silently running unconfined is never legal: an
|
|
292
|
+
operator who configured a boundary and sees no error believes it is enforced.
|
|
293
|
+
|
|
294
|
+
## Exports
|
|
295
|
+
|
|
296
|
+
`resolveSandboxPolicy` · `confiningPolicy` · `narrowPolicy` · `writableRoots` ·
|
|
297
|
+
`unreadablePaths` · `accessFor` · `orderEntries` · `createFsFence` ·
|
|
298
|
+
`classifyOutcome` · `annotateStderr` · `sandboxViolation` · the path algebra
|
|
299
|
+
(`normalizePath`, `containsPath`, `dedupeRoots`, …) and every contract type.
|