@alvin0/ai-agent-sdk-sandbox 0.1.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 alvin0 (chaulamdinhai) <chaulamdinhai@gmail.com>
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,299 @@
1
+ # @alvin0/ai-agent-sdk-sandbox
2
+
3
+ The sandbox contract: the file-effect vocabulary, per-call policy resolution,
4
+ the writable-root algebra every backend shares, and the rules that keep a broken
5
+ sandbox from reading as a denied command.
6
+
7
+ Runtime: **Universal**. This package imports nothing — no `node:` builtins, no
8
+ dependencies — so it runs anywhere the core SDK does. Enforcement lives in
9
+ [`@alvin0/ai-agent-sdk-sandbox-node`](../sandbox-node#readme).
10
+
11
+ ```bash
12
+ pnpm add @alvin0/ai-agent-sdk-sandbox
13
+ ```
14
+
15
+ ## Modes
16
+
17
+ `SandboxMode` governs **file effects only**. Network reachability and process
18
+ visibility are deliberately outside this vocabulary; they belong to their own
19
+ seam, and a mode that claimed to cover them would be lying.
20
+
21
+ | Mode | File effects |
22
+ | --- | --- |
23
+ | `read-only` | No writes anywhere; the backend still permits the `/dev/null` sink a shell needs |
24
+ | `workspace-write` | Writes under the workspace root, plus any temp roots the consumer explicitly grants |
25
+ | `danger-full-access` | No confinement; the provider is never consulted |
26
+
27
+ ## The authorization boundary
28
+
29
+ Everything a tool sends is model-authored JSON, so a policy input that widens
30
+ authority is one the model can grant itself. `mode` and `entries` on a request
31
+ are therefore the untrusted half: a requested mode is honoured only when it is
32
+ at least as strict as the session's own, and requested entries are intersected
33
+ with deployment and hardening layers. A `deny` can narrow `read`, and `read` can
34
+ narrow `write`, but neither can reopen a standing denial. A requested `write`
35
+ is refused outright.
36
+
37
+ Widening goes through a capability instead of data:
38
+
39
+ ```ts
40
+ import { approveSandboxEscalation, resolveSandboxPolicy } from '@alvin0/ai-agent-sdk-sandbox'
41
+
42
+ // After the host has actually authorized it — a human prompt, a policy engine.
43
+ const approval = approveSandboxEscalation({ mode: 'workspace-write' })
44
+ const policy = resolveSandboxPolicy({ cwd, sessionMode, approval }, defaults)
45
+ ```
46
+
47
+ Only a value this module minted is accepted. `JSON.parse('{"approved":true}')`
48
+ is refused, which is what a forged approval arriving in a tool payload looks
49
+ like. Deployment configuration goes in `defaults`, which is the trusted half and
50
+ may grant anything.
51
+
52
+ ## Per-call policy
53
+
54
+ Policy is resolved per capability call, never fixed on the provider. Two
55
+ consumers can confine under different boundaries at the same instant, and an
56
+ approved escalation is a *new call with a wider policy* — not a mutation of
57
+ shared state.
58
+
59
+ ```ts
60
+ import { confiningPolicy, resolveSandboxPolicy } from '@alvin0/ai-agent-sdk-sandbox'
61
+
62
+ const resolved = resolveSandboxPolicy(
63
+ { cwd: session.cwd, sessionMode: session.mode, mode: approvedOverride },
64
+ { mode: 'read-only', workspaceRoot: deploymentRoot },
65
+ )
66
+
67
+ const policy = confiningPolicy(resolved)
68
+ // `undefined` under danger-full-access: spawn the original argv, ask nothing.
69
+ ```
70
+
71
+ Precedence is fixed: an approved explicit mode outranks the session's mode,
72
+ which outranks the deployment default. A session's cwd is its `workspace-write`
73
+ boundary; the configured root is the fallback for agentless calls.
74
+
75
+ ## Reading a command for what it does
76
+
77
+ The file seam cannot tell `systemctl status nginx` from `systemctl restart
78
+ nginx`: both are argv, neither writes a file the policy governs, and one
79
+ observes while the other changes the machine. `classifyExec` reads the command
80
+ semantically so a harness can decide before any of it runs.
81
+
82
+ ```ts
83
+ classifyExec(['systemctl', 'status', 'nginx']) // observe -> allow
84
+ classifyExec(['systemctl', 'restart', 'nginx']) // service-control -> ask-approval
85
+ classifyExec(['aws', 'configure', 'list']) // credential -> deny
86
+ ```
87
+
88
+ | Capability | Default outcome |
89
+ | --- | --- |
90
+ | `observe` | `allow` |
91
+ | `use` | `allow-scoped` |
92
+ | `modify`, `service-control`, `package-install`, `privilege` | `ask-approval` |
93
+ | `credential`, `critical` | `deny` |
94
+ | `unknown` | `ask-approval` |
95
+
96
+ This is a classifier, and a classifier is a guess. Two rules keep the guess from
97
+ becoming a hazard: a command it does not recognise is **never allowed**, and a
98
+ command that hides others — a shell string, a pipeline, a chain, an argv with
99
+ separators in it — is decided by the riskiest thing inside it rather than by its
100
+ wrapper.
101
+
102
+ The script is walked one character at a time rather than split with a pattern,
103
+ because `grep -E 'a|b'` puts a separator inside a quoted word: a pattern either
104
+ splits there, inventing a command out of a regex, or refuses to split wherever a
105
+ quote appears. Quotes and escapes are respected, so `echo "hi; rm -rf /etc"` is
106
+ one command printing text while `echo hi && rm -rf /etc` is two, and the second
107
+ decides.
108
+
109
+ What it deliberately does not do: expand variables, resolve `$(...)`, follow a
110
+ script file, or know what an unrecognised binary does. `eval "$CMD"` classifies
111
+ as unrecognised, which asks — it does not read what `$CMD` holds.
112
+
113
+ Two of its rules exist because real model output demanded them. Asked to show
114
+ AWS credentials, a model proposed `aws configure list` and `aws sts
115
+ get-caller-identity` — neither names `~/.aws/credentials`, so a rule matching
116
+ credential *paths* saw nothing while the secret was read inside the tool. And
117
+ `if [ -f package.json ]; then npm test; fi`, read as one argv, names `[` and
118
+ looks like a test while what it runs is the test suite.
119
+
120
+ It decides; it does not enforce. `confine()` and `fence()` are what hold.
121
+
122
+ ## Wiring it into an agent
123
+
124
+ This package does not depend on `@alvin0/ai-agent-sdk-core`, and core has no
125
+ sandbox slot. A session already has the two seams needed: an interceptor decides
126
+ and the approval broker asks.
127
+
128
+ ```ts
129
+ const sandboxInterceptor: ToolInterceptor = {
130
+ name: 'sandbox:exec',
131
+ before: async (call, next) => {
132
+ const argv = argvOf(call) // undefined for a non-exec tool
133
+ if (argv === undefined) return await next()
134
+ const verdict = classifyExec(argv)
135
+ if (verdict.outcome === 'deny') return { kind: 'deny', reason: verdict.reason }
136
+ if (verdict.outcome !== 'ask-approval') return await next()
137
+ asked.add(call.callId)
138
+ return { kind: 'ask', reason: verdict.reason } // the broker asks a person
139
+ },
140
+ // Reaching `around` for a call that asked IS the answer. Mint the capability
141
+ // here, never from tool arguments.
142
+ around: async (call, next) => {
143
+ if (!asked.delete(call.callId)) return await next()
144
+ granted.set(call.callId, approveSandboxEscalation({ entries: escalationFor(call) }))
145
+ try { return await next() } finally { granted.delete(call.callId) }
146
+ },
147
+ }
148
+
149
+ agent.createSession({ tools: [runCommand], interceptors: [sandboxInterceptor], approvals })
150
+ ```
151
+
152
+ The tool body reads the approval by `ctx.callId`, passes it to
153
+ `resolveSandboxPolicy`, and then confines or fences. `classifyExec` decides;
154
+ this package never holds anything.
155
+
156
+ ## Network reach
157
+
158
+ `SandboxMode` governs file effects and says so. Reachability is a second,
159
+ independent axis on the same policy, because folding it into the mode would
160
+ make the mode claim something it does not decide — and because the two are
161
+ enforced by different mechanisms, so a host can provide one without the other.
162
+
163
+ | Reach | Meaning |
164
+ | --- | --- |
165
+ | `deny` | No network at all |
166
+ | `loopback` | Loopback only — for a proxy or language server the deployment runs |
167
+ | `allow-all` | Unrestricted (the default, and what this package did before the seam) |
168
+
169
+ ```ts
170
+ resolveSandboxPolicy({ cwd, network: 'deny' }, { mode: 'read-only', workspaceRoot, network: 'deny' })
171
+ ```
172
+
173
+ Reach narrows exactly like authority does: a request may tighten its own, never
174
+ loosen it, and only a minted approval widens it. A deployment running anything
175
+ untrusted should default to `deny` and widen per call.
176
+
177
+ Unix sockets are **not** governed by this axis — they are filesystem objects,
178
+ so a host daemon socket is closed by a `deny` entry, not by a network mode. The
179
+ two seams compose; neither substitutes for the other.
180
+
181
+ ## Deny-list or allow-list
182
+
183
+ By default the host is readable and a policy closes paths one at a time. That
184
+ is a deny-list: it protects what someone remembered to name, and the path nobody
185
+ thought about stays open.
186
+
187
+ ```ts
188
+ { mode: 'read-only', workspaceRoot, baseline: 'deny',
189
+ entries: [{ path: '/var/log/my-api', access: 'read' }] }
190
+ ```
191
+
192
+ `baseline: 'deny'` inverts it — nothing is readable until an entry says so —
193
+ which is how a request to investigate one service is scoped to that service's
194
+ logs rather than to every log on the host.
195
+
196
+ It is enforced by `fence(policy)`, not by the kernel profiles, and
197
+ `confine()` refuses it rather than pretending: inverting a mount profile means
198
+ binding only what a program needs, and the set a program needs to start at all
199
+ is specific to an OS build.
200
+
201
+ ## Approvals are spent
202
+
203
+ An approval is consumed the first time a policy is resolved with it. A person
204
+ approving "read this file" approved one read, and a grant that survives its own
205
+ operation is a grant nobody is still watching.
206
+
207
+ ```ts
208
+ approveSandboxEscalation({ entries: [{ path: '/etc/app/config.yaml', access: 'write' }] })
209
+ approveSandboxEscalation({ mode: 'workspace-write', scope: 'session', expiresAt })
210
+ ```
211
+
212
+ Note what the first grant does *not* do: it never mentions a mode, so the
213
+ policy stays `read-only` and exactly one file becomes writable. Raising the mode
214
+ instead would make the whole workspace writable, and the named resource
215
+ decorative.
216
+
217
+ ## Nested carve-outs
218
+
219
+ A single writable root is not enough. An agent that may write in a repository
220
+ must still be kept out of `.git`, or it can install a hook that runs arbitrary
221
+ code on the next `git` invocation. Entries express that as an overlapping list
222
+ resolved by **path specificity** — the deepest matching entry wins:
223
+
224
+ ```ts
225
+ const entries = [
226
+ { path: '/repo', access: 'write' },
227
+ { path: '/repo/a', access: 'deny' },
228
+ { path: '/repo/a/b', access: 'write' },
229
+ ]
230
+ // /repo/x → write · /repo/a/x → deny · /repo/a/b/x → write
231
+ ```
232
+
233
+ `grantLayers(policy)` resolves a policy into exactly that: an ordered stack of
234
+ layers, broadest first, each overriding the ones beneath it for its own subtree.
235
+ Order is the semantics, and it is why the model is a stack rather than a pair of
236
+ sets — a set of "granted roots" has nowhere to record a grant that lives *inside*
237
+ something denied, so the third line above would silently vanish.
238
+
239
+ ```ts
240
+ grantLayers(policy)
241
+ // write mode /repo
242
+ // read protected /repo/.git ← and .ssh, .aws, .netrc, …
243
+ // deny entry /repo/vendor
244
+ // write entry /repo/vendor/cache
245
+ ```
246
+
247
+ `PROTECTED_SUBPATHS` (`.git`, `.ssh`, `.aws`, `.netrc`, …) is layered under every
248
+ granted root automatically. They are `read`, not `deny`: this is a write
249
+ boundary, so they stay readable. An explicit entry at the same depth outranks a
250
+ generated one, so a deployment can deliberately reopen one.
251
+
252
+ A layer that would not change the access already in force is dropped, so the
253
+ result carries no rule that does nothing. `writableRoots(policy)` flattens the
254
+ same layers for callers that only want the two lists.
255
+
256
+ This one function is what every enforcement path reads — the kernel profiles and
257
+ the in-process fence. Deriving them separately is how a profile and a fence
258
+ silently drift into disagreeing about what is writable.
259
+
260
+ ## Classification
261
+
262
+ Two failures look identical in a shell and mean opposite things:
263
+
264
+ - **denied** — confinement worked and blocked the command.
265
+ - **runner failure** — the sandbox itself refused or crashed; the command
266
+ *never ran*.
267
+
268
+ Reporting the second as the first sends a model off rewriting correct code.
269
+ `classifyOutcome` checks runner failure first, and never infers it from an exit
270
+ code alone: a rule needs a nonzero exit, its own exit-code gate, and a fatal
271
+ signature on a line that survives informational exclusion.
272
+
273
+ ```ts
274
+ import { annotateStderr, classifyOutcome } from '@alvin0/ai-agent-sdk-sandbox'
275
+
276
+ const classification = classifyOutcome(
277
+ { exitCode: result.status, stderr: result.stderr, signal: result.signal },
278
+ confined, // carries the wrapping backend's own dialect
279
+ )
280
+ const stderr = annotateStderr(result.stderr, classification, policy.mode)
281
+ ```
282
+
283
+ Denial signatures are matched against **the backend that actually wrapped the
284
+ command**, never a cross-backend union — a union claims denials a given backend
285
+ never produces. A `SIGSYS` kill is treated as a denial without matching any
286
+ text, because a seccomp kill is unambiguous.
287
+
288
+ ## Fail closed
289
+
290
+ `SandboxUnavailableError` (`SANDBOX_UNAVAILABLE`) is thrown when no backend can
291
+ enforce a confining policy. Silently running unconfined is never legal: an
292
+ operator who configured a boundary and sees no error believes it is enforced.
293
+
294
+ ## Exports
295
+
296
+ `resolveSandboxPolicy` · `confiningPolicy` · `narrowPolicy` · `writableRoots` ·
297
+ `unreadablePaths` · `accessFor` · `orderEntries` · `createFsFence` ·
298
+ `classifyOutcome` · `annotateStderr` · `sandboxViolation` · the path algebra
299
+ (`normalizePath`, `containsPath`, `dedupeRoots`, …) and every contract type.