@namzu/sandbox 15.0.0 → 17.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +324 -0
- package/README.md +223 -0
- package/dist/backends/aci-standby-pool/index.d.ts +22 -4
- package/dist/backends/aci-standby-pool/index.d.ts.map +1 -1
- package/dist/backends/aci-standby-pool/index.js +31 -7
- package/dist/backends/aci-standby-pool/index.js.map +1 -1
- package/dist/backends/docker/index.d.ts +408 -26
- package/dist/backends/docker/index.d.ts.map +1 -1
- package/dist/backends/docker/index.js +1173 -168
- package/dist/backends/docker/index.js.map +1 -1
- package/dist/backends/firecracker/transport.d.ts +156 -1
- package/dist/backends/firecracker/transport.d.ts.map +1 -1
- package/dist/backends/firecracker/transport.js +223 -29
- package/dist/backends/firecracker/transport.js.map +1 -1
- package/dist/backends/http-worker-client.d.ts +64 -2
- package/dist/backends/http-worker-client.d.ts.map +1 -1
- package/dist/backends/http-worker-client.js +78 -7
- package/dist/backends/http-worker-client.js.map +1 -1
- package/dist/backends/kubernetes/egress-policy.d.ts +193 -102
- package/dist/backends/kubernetes/egress-policy.d.ts.map +1 -1
- package/dist/backends/kubernetes/egress-policy.js +321 -146
- package/dist/backends/kubernetes/egress-policy.js.map +1 -1
- package/dist/backends/kubernetes/per-sandbox-policy.d.ts +6 -6
- package/dist/backends/kubernetes/per-sandbox-policy.d.ts.map +1 -1
- package/dist/backends/kubernetes/per-sandbox-policy.js +21 -53
- package/dist/backends/kubernetes/per-sandbox-policy.js.map +1 -1
- package/dist/backends/kubernetes/transport.d.ts +7 -0
- package/dist/backends/kubernetes/transport.d.ts.map +1 -1
- package/dist/backends/kubernetes/transport.js.map +1 -1
- package/dist/egress/proxy.d.ts +47 -2
- package/dist/egress/proxy.d.ts.map +1 -1
- package/dist/egress/proxy.js +31 -7
- package/dist/egress/proxy.js.map +1 -1
- package/dist/index.d.ts +130 -6
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +55 -5
- package/dist/index.js.map +1 -1
- package/package.json +4 -4
- package/src/backends/aci-standby-pool/index.ts +37 -7
- package/src/backends/docker/index.ts +1475 -196
- package/src/backends/firecracker/transport.ts +387 -36
- package/src/backends/http-worker-client.ts +89 -5
- package/src/backends/kubernetes/egress-policy.ts +455 -187
- package/src/backends/kubernetes/per-sandbox-policy.ts +21 -66
- package/src/backends/kubernetes/transport.ts +7 -0
- package/src/egress/proxy.ts +65 -8
- package/src/index.ts +162 -5
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,329 @@
|
|
|
1
1
|
# @namzu/sandbox
|
|
2
2
|
|
|
3
|
+
## 17.0.0
|
|
4
|
+
|
|
5
|
+
### Major Changes
|
|
6
|
+
|
|
7
|
+
- 1722448: **A `container:docker` sandbox with an egress allowlist now needs an internal
|
|
8
|
+
network, a proxy image, and `container-network` reachability. Take this version
|
|
9
|
+
if you use one; otherwise nothing you install changes shape.**
|
|
10
|
+
|
|
11
|
+
The allowlist tier used to be enforced by `HTTP_PROXY` and nothing else. The
|
|
12
|
+
proxy ran in the process that created the sandbox, on the host's loopback, and
|
|
13
|
+
the sandbox kept ordinary bridge networking with full outbound reachability —
|
|
14
|
+
`--add-host namzu-egress:host-gateway` was the only thing pointing traffic at
|
|
15
|
+
it. Anything inside the container that opened a socket directly reached the
|
|
16
|
+
network with the allowlist unconsulted. It is now a sibling container on an
|
|
17
|
+
`--internal` network the sandbox is also on, which has no route off it: the
|
|
18
|
+
sandbox's traffic reaches the internet only through that container, and
|
|
19
|
+
everything else a host attaches to that network is a container on the sandbox's
|
|
20
|
+
own subnet.
|
|
21
|
+
|
|
22
|
+
**What a host using `EgressPolicy` of `static` or `resolver` must now do**, all
|
|
23
|
+
three refused at `create()` rather than downgraded:
|
|
24
|
+
|
|
25
|
+
1. `network` must be an `--internal` network: `docker network create
|
|
26
|
+
--internal <name>`. The network's `Internal` flag is read back from the
|
|
27
|
+
daemon; the name is never trusted.
|
|
28
|
+
2. `hostReachability` must be `'container-network'`. A published host port
|
|
29
|
+
needs a route out and an internal network has none — the refusal names the
|
|
30
|
+
mode to move to. **A host-side consumer (the CLI, direct dev) that reached
|
|
31
|
+
the worker on `127.0.0.1:<port>` under an allowlist policy must move to a
|
|
32
|
+
consumer on the internal network.**
|
|
33
|
+
3. `egressProxyImage` must name the proxy image:
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
pnpm --filter @namzu/sandbox build
|
|
37
|
+
docker build -f packages/sandbox/egress-proxy/Dockerfile -t <tag> packages/sandbox
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
It is a second image because the sandbox image is a string this backend
|
|
41
|
+
cannot read, and the bind-mount alternative breaks on the remote-daemon
|
|
42
|
+
deployment `container-network` exists for.
|
|
43
|
+
|
|
44
|
+
`HTTP_PROXY` and friends are still set and now direct traffic rather than
|
|
45
|
+
permit it: a tool that honours them goes through the boundary, one that ignores
|
|
46
|
+
them fails `Network unreachable` instead of bypassing the policy.
|
|
47
|
+
|
|
48
|
+
**Four behaviour changes to know about, all of them consequences of the
|
|
49
|
+
boundary existing:**
|
|
50
|
+
|
|
51
|
+
- A `resolver` policy is resolved at `create()` and at each
|
|
52
|
+
`setNetworkPolicy()`, not per request. A resolver that rotates between those
|
|
53
|
+
moments is not picked up until the next one.
|
|
54
|
+
- `setNetworkPolicy()` replaces the proxy container rather than swapping the
|
|
55
|
+
allowlist in place. The window between the two has no proxy in it and fails
|
|
56
|
+
**closed** — no request is permitted that the new policy would refuse.
|
|
57
|
+
- `brokeredCredentials` are passed to the proxy container's environment. They
|
|
58
|
+
still never enter the sandbox; they are now readable by anything with access
|
|
59
|
+
to the docker daemon, which should be treated as credential access. (On the
|
|
60
|
+
way there the value goes through the `docker` CLI child's ENVIRONMENT, not its
|
|
61
|
+
argv — an argv is world-readable in `/proc/<pid>/cmdline` on Linux.) Note also
|
|
62
|
+
that `createSandboxProvider` has never forwarded `brokeredCredentials` at all;
|
|
63
|
+
that pre-existing gap is unchanged.
|
|
64
|
+
- `egressProxyUpstreamNetwork` (default `bridge`) is where the proxy reaches
|
|
65
|
+
the internet. Anything else attached to that network can reach the proxy and
|
|
66
|
+
use it, so name a dedicated one on a shared daemon. `'none'`, and the internal
|
|
67
|
+
network itself, are refused: either leaves the proxy with no route out.
|
|
68
|
+
- Everything else on the SANDBOX's internal network is reachable from the
|
|
69
|
+
sandbox — a second sandbox, a second sandbox's proxy. That is a property of
|
|
70
|
+
the network rather than of this tier; give each sandbox its own internal
|
|
71
|
+
network when they should not see one another.
|
|
72
|
+
|
|
73
|
+
`deny-all` and `allow-all` are unaffected beyond the internal network
|
|
74
|
+
`deny-all` already required. `EgressProxy` gains two optional options
|
|
75
|
+
(`bindHost`, `selfNames`) and the container entrypoint sets them; the class is
|
|
76
|
+
otherwise unchanged.
|
|
77
|
+
|
|
78
|
+
- f1fd331: The container backend's worker now authenticates its caller, and a worker that
|
|
79
|
+
cannot refuses to start on a routable bind. This is a `major` because the wire
|
|
80
|
+
shape of a public surface changed: every route but `GET /healthz` now requires
|
|
81
|
+
`Authorization: Bearer <NAMZU_SANDBOX_TOKEN>`, a missing or wrong token is
|
|
82
|
+
`401 {"error":"unauthorized"}`, and a worker with no token will not listen on
|
|
83
|
+
anything but loopback. A token that is set but cannot be answered — empty,
|
|
84
|
+
padded, or carrying a character an HTTP header cannot hold — now refuses to
|
|
85
|
+
start rather than leaving a worker that looks authenticated and refuses its own
|
|
86
|
+
host.
|
|
87
|
+
|
|
88
|
+
Nothing in the package's TypeScript API changed — no export was removed,
|
|
89
|
+
renamed or narrowed, and `HttpWorkerClient` and `execViaHttpWorker` only gained
|
|
90
|
+
an optional token parameter. What changed is the contract between this backend
|
|
91
|
+
and the worker process it runs, and that contract is deployed separately: the
|
|
92
|
+
worker lives in `packages/sandbox/worker/server.js`, which the tarball does not
|
|
93
|
+
ship, so the image is built from whatever checkout the operator has.
|
|
94
|
+
|
|
95
|
+
**What breaks, and for whom.**
|
|
96
|
+
|
|
97
|
+
- **A worker image rebuilt from this release, paired with a host that predates
|
|
98
|
+
it.** The old host injects no token, so the new worker refuses to start on
|
|
99
|
+
the default `0.0.0.0` bind and every `create()` fails at readiness with the
|
|
100
|
+
container already exited. The refusal names the variable. Bring the host up
|
|
101
|
+
in the same step, or set the escape below.
|
|
102
|
+
- **Anything that talks to the worker without going through this backend.** The
|
|
103
|
+
route contract is now authenticated; a probe, a health-check script that
|
|
104
|
+
calls `/execute`, or a hand-rolled client must present the token.
|
|
105
|
+
- **The standby-pool backend.** A pooled worker built from this release will
|
|
106
|
+
refuse to boot on its routable default bind, because this backend has no way
|
|
107
|
+
to hand it a credential. The claim API admits exactly one property override,
|
|
108
|
+
and it is not `env` — it is a config map, which on Linux reaches the container
|
|
109
|
+
as a file mount under `/mnt/configmap/<containername>/<key>`, not as an
|
|
110
|
+
environment variable, while the worker reads its token from `process.env` at
|
|
111
|
+
startup. The platform does not validate config-map values either, and its own
|
|
112
|
+
guidance is that a value affecting application security belongs in an
|
|
113
|
+
environment variable. So the per-instance credential for this backend is NOT
|
|
114
|
+
implemented here, and the change that would close the gap — a per-claim value
|
|
115
|
+
in the config map, read by the worker — is written down in
|
|
116
|
+
`docs/sdk/container-sandbox-worker.md` rather than half-built. A token on the
|
|
117
|
+
shared profile does not fix it either: the worker would boot and then `401`
|
|
118
|
+
every call the backend makes, because this backend's client is constructed
|
|
119
|
+
without a token and it has no field to carry one.
|
|
120
|
+
|
|
121
|
+
**A host and a worker from the same release need no configuration.** The
|
|
122
|
+
container backend mints 32 random bytes per `create()`, hands the value to the
|
|
123
|
+
docker CLI in its own environment (resolved by a valueless
|
|
124
|
+
`--env NAMZU_SANDBOX_TOKEN`, so it is in no argv), and sends it as a bearer
|
|
125
|
+
header on every call it makes. Upgrading both together is the whole migration.
|
|
126
|
+
The token is per instance, never written to disk, and dies with the container.
|
|
127
|
+
It is readable where a peer in the same container or a caller who can already
|
|
128
|
+
talk to the docker daemon could read it — `docker inspect` shows it on the
|
|
129
|
+
container config for the container's life — which is why it is per-instance
|
|
130
|
+
rather than shared, and why it is a defence in depth behind network placement
|
|
131
|
+
rather than a replacement for it.
|
|
132
|
+
|
|
133
|
+
**The standby pool: place it first, decide about the credential second.** If you
|
|
134
|
+
run that backend, the control that actually carries the exposure is the network,
|
|
135
|
+
and it is the one to get right before anything else — `aci-standby-pool` already
|
|
136
|
+
refuses to claim a container group without `subnetId` (unless you set
|
|
137
|
+
`allowPublicAddress: true` and mean it), so the group sits on a private address
|
|
138
|
+
and that network is what stands in front of the worker. The refusal is unchanged
|
|
139
|
+
by this release; what has changed is that it is now the first half of the answer
|
|
140
|
+
rather than the whole of it. The second half is the escape below, which is what
|
|
141
|
+
makes a pooled worker start at all today: put
|
|
142
|
+
`NAMZU_SANDBOX_ALLOW_UNAUTHENTICATED=1` on the container group profile, with the
|
|
143
|
+
group on a private address. It is needed there only because this backend cannot
|
|
144
|
+
present a credential, and it turns a worker that cannot boot into one that
|
|
145
|
+
serves inside that address. Read the standby-pool bullet under "What breaks"
|
|
146
|
+
before taking that step, and note that a token on the shared profile is not an
|
|
147
|
+
alternative to it.
|
|
148
|
+
|
|
149
|
+
**For a deployment the container backend creates**, the escape is a migration
|
|
150
|
+
tool rather than a destination: it is what keeps you running while the host and
|
|
151
|
+
the worker image are brought up in the same step.
|
|
152
|
+
|
|
153
|
+
**To keep the old behaviour on purpose**, set
|
|
154
|
+
`NAMZU_SANDBOX_ALLOW_UNAUTHENTICATED=1` in the worker's environment — the
|
|
155
|
+
container backend's `options.env` for a sandbox it creates, or the standby
|
|
156
|
+
pool's container group profile for a pooled one. That gives the credential up
|
|
157
|
+
rather than deferring it: with it set and no token, every route but `/healthz`
|
|
158
|
+
is open to whoever can route to the container, which is exactly what the
|
|
159
|
+
default used to be. On the pool, the address comes first and this second; on
|
|
160
|
+
every other backend, both the host and the image should be upgraded in one step
|
|
161
|
+
instead. Only `1`, `true`, `yes` and `on` count, in any case and with no
|
|
162
|
+
surrounding whitespace; everything else — `= 0`, `= false`, `= " yes "` — means
|
|
163
|
+
off, so the flag cannot turn itself on.
|
|
164
|
+
|
|
165
|
+
Also unchanged, and worth repeating from the README: the transport is plain
|
|
166
|
+
HTTP, so this is a bearer token over the wire and it is defence in depth behind
|
|
167
|
+
network placement, not a replacement for it. See
|
|
168
|
+
`docs/sdk/container-sandbox-worker.md`.
|
|
169
|
+
|
|
170
|
+
### Minor Changes
|
|
171
|
+
|
|
172
|
+
- 01fbc9b: Three additive declarations, no default changed and nothing removed:
|
|
173
|
+
|
|
174
|
+
- `MicroVMBackendConfig.onExecTiming` — the owned Firecracker tier's provider
|
|
175
|
+
config gains an optional per-exec timing hook.
|
|
176
|
+
- `FirecrackerTransportTiming` — exported from the package entry point, the
|
|
177
|
+
shape that hook is called with.
|
|
178
|
+
- `VsockTransportOptions.onExecTiming` — the same hook at the transport level,
|
|
179
|
+
for a host that builds a `VsockAgentTransport` itself.
|
|
180
|
+
|
|
181
|
+
A host that sets none of them sends, receives and waits for exactly what it did
|
|
182
|
+
before: the timing accumulator is created only when the hook is set, and the
|
|
183
|
+
`undefined` checks that skip creating it skip every clock read that would have
|
|
184
|
+
filled it.
|
|
185
|
+
|
|
186
|
+
The hook reports one `exec()`'s wall clock as named phases — `dialMs`,
|
|
187
|
+
`reserveMs`, `executeMs` and `drainMs`, the four the Kubernetes tier's
|
|
188
|
+
`onTiming` already reports, plus `firstFrameMs`, `terminatorMs` and
|
|
189
|
+
`peerCloseMs` for the intervals inside the execute round trip. The three
|
|
190
|
+
sub-phases are absent, rather than zero, when their phase was never reached; the
|
|
191
|
+
four base phases are always present and use `0` for "this never happened" (a
|
|
192
|
+
call that failed at the dial reports `executeMs: 0`). The payload is durations
|
|
193
|
+
only: never the agent token, a command, its arguments, or any output.
|
|
194
|
+
|
|
195
|
+
One behaviour change worth naming, because it is the reason the hook is useful:
|
|
196
|
+
the transport's execution adapter and its `RemoteExecutionController` are now
|
|
197
|
+
built per call instead of once per transport. That controller holds no per-call
|
|
198
|
+
state, so nothing observable changes — but an adapter built once had nowhere
|
|
199
|
+
per-call to accumulate into, and a single accumulator on the transport would
|
|
200
|
+
have two concurrent `exec()` calls writing each other's phases. The Kubernetes
|
|
201
|
+
tier's own transport has been arranged this way since it was written.
|
|
202
|
+
|
|
203
|
+
`POST_RESPONSE_CLOSE_TIMEOUT_MS` (1 s) is unchanged and its behaviour is
|
|
204
|
+
unchanged; it is now documented as what it always was — a reject-only guard
|
|
205
|
+
that fails a socket whose peer never closes, never a wait a successful call
|
|
206
|
+
pays. If your host puts a relay between this process and the guest, `peerCloseMs`
|
|
207
|
+
is the number that tells you whether that relay holds the FIN: a hold **under**
|
|
208
|
+
a second is reported there on a call that resolved, and a hold **at or past**
|
|
209
|
+
a second fails the call with `vsock transport: exec peer did not close after
|
|
210
|
+
terminator`, reporting no `peerCloseMs` at all — the close in that case is the
|
|
211
|
+
one this transport causes itself when the guard fires, and its own constant is
|
|
212
|
+
not published as an elapsed time. So a resolved exec's fixed cost cannot be
|
|
213
|
+
hiding in that window: whichever way the window goes, it is either reported or
|
|
214
|
+
it is a rejection, and never a silent second.
|
|
215
|
+
|
|
216
|
+
### Patch Changes
|
|
217
|
+
|
|
218
|
+
- 656598a: Nothing a consumer installs changes. The one edited file is
|
|
219
|
+
`src/backends/docker/__tests__/leaf-permissions.smoke.test.ts`, which
|
|
220
|
+
`package.json#files` excludes from the tarball; runtime behaviour, exports, types
|
|
221
|
+
and defaults are untouched.
|
|
222
|
+
|
|
223
|
+
The `Sandbox smoke` workflow has been red on `main` since the docker hardening
|
|
224
|
+
landed (`481fb8ff`, `--read-only` on by default). That case asserts that uid 1001
|
|
225
|
+
cannot `mkdir` into the unbound `/mnt/user-data`, and it pinned the refusal to
|
|
226
|
+
one spelling — `permission denied`. With a read-only rootfs the kernel answers
|
|
227
|
+
`read-only file system` instead, because EROFS is consulted before the DAC check
|
|
228
|
+
that would have produced EACCES. The property the case exists for never changed
|
|
229
|
+
(a bound writable leaf would let the `mkdir` succeed, and the `--rc` assertion
|
|
230
|
+
still catches that); only the kernel's wording did.
|
|
231
|
+
|
|
232
|
+
The assertion now accepts either refusal and says why both are legitimate, so the
|
|
233
|
+
next change to which check fires first is a one-line read rather than a day of
|
|
234
|
+
red. `dash` and Docker are both absent from the machine this was written on, so
|
|
235
|
+
the fix is verified by the workflow that runs this file, not locally.
|
|
236
|
+
|
|
237
|
+
- Updated dependencies [4e8cf5c]
|
|
238
|
+
- Updated dependencies [19abb2c]
|
|
239
|
+
- @namzu/sdk@42.0.0
|
|
240
|
+
|
|
241
|
+
## 16.0.0
|
|
242
|
+
|
|
243
|
+
### Major Changes
|
|
244
|
+
|
|
245
|
+
- 11f7ea2: **A `.domain` allowlist entry is expanded on the config-level egress translation, a `specs` list on the applied object is refused, and a deployment that applied the old object must re-apply it before its next `create()`.**
|
|
246
|
+
|
|
247
|
+
Everything here is stated against **`@namzu/sandbox@15.0.0`, the published release this one ships on top of** — its bytes as downloaded from the registry, not this branch's base. `15.0.0` emitted the config-level `.domain` entry verbatim and said so itself: its release notes call "that a config-level `.domain` entry matches no answer at all" a pre-existing defect "left exactly as it was and deferred to its own change". This is that change, and the deferral is what the migration below is for.
|
|
248
|
+
|
|
249
|
+
`config.egress.policy = { kind: 'static', allowedHosts: ['.example.com'] }` under `engine: 'cilium'` therefore emits `toFQDNs: [{ matchName: example.com }, { matchPattern: '*.example.com' }]` where `15.0.0` emitted `toFQDNs: [{ matchName: '.example.com' }]` — the entry VERBATIM, leading dot included. No DNS answer carries a name with a leading dot, so the old object matched nothing, and nothing refused it on the way in: the shipped admission fence's `matchConditions` scope it to the sandbox host's ServiceAccount, so an operator applying the config-level object is not matched by it at all. `verifyEgressPolicyApplied` then deep-equalled what was sent, the create reported success, and the domain and every subdomain it was asked to allow were DENIED. With `ciliumNarrowing.dnsNames` on it was worse than useless, because the DNS-visibility rule carried the same unmatched names — the DNS proxy denied the lookup the `toFQDNs` half needs before that half could ever learn an address.
|
|
250
|
+
|
|
251
|
+
`SandboxNetworkPolicy.allowedHosts` has always said `.example.com` is the domain **and its subdomains** — `packages/sdk/src/types/sandbox/index.ts` — and `src/egress/allowlist.ts` implements exactly that for the docker backend's proxy. The Kubernetes PER-SANDBOX translation already honoured it. The config-level translation now reuses the same helper, so the entry becomes `matchName: example.com` plus `matchPattern: '*.example.com'`, and the DNS-visibility rule carries the names that expansion admits — `matchPattern: '*.<host>.<suffix>'` under every cluster search suffix, which is what the per-sandbox path already emitted.
|
|
252
|
+
|
|
253
|
+
**Every consumer-visible change, against `15.0.0`.** Nothing changes for an allowlist with no leading-dot entry: the expansion returns `[{ matchName: entry }]` for an exact host, the DNS rule's wildcard branch is keyed on the leading dot alone, so every other entry's bytes are identical, with and without `dnsNames`, and no entry's letter case is rewritten. The rest is this list, and each item is a create or a compile that behaves differently after the upgrade:
|
|
254
|
+
|
|
255
|
+
1. **A config-level `.domain` entry changes the object it emits**, so `verifyEgressPolicyApplied` — which compares the applied object to the translation exactly — refuses your next `createKubernetesWorkspace` or `create()` with `KubernetesEgressPolicyMismatchError` until you **re-apply the manifest this backend now translates**: the same `kubectl apply` you ran for the old one, against a cluster object that admitted nothing. That is the whole migration.
|
|
256
|
+
2. **A config-level allowlist entry that is not a hostname is refused, by name.** `.com`, `.`, `..example.com`, `*`, `*.example.com`, a URL, a port suffix and an IP literal were each passed straight into the emitted `toFQDNs` by `15.0.0`: the per-sandbox writer validated them, the config-level translation did not. Refused now, with `KubernetesNetworkPolicyHostError` and nothing written, because both halves of what they used to do are wrong — as emitted by `15.0.0` they became a `matchName` that matched nothing, and with the expansion above `.com` would instead emit `matchPattern: '*.com'` and grant every name under a public suffix, silently, on a path no admission policy covers. `.example.com`, `example.com` and `api.example.com` are all still accepted, and a `resolver` policy's returned hosts are held to the same grammar. **A create whose allowlist carries one of these works on `15.0.0` and fails after this upgrade**, until the entry is corrected.
|
|
257
|
+
3. **`KubernetesNetworkPolicyHostError` loses its third constructor argument.** `15.0.0` ships `constructor(host, reason, options?: { stateHostnameGrammar?: boolean })` on that class, which is exported from the package root; this release ships `constructor(host, reason)`. A consumer that constructed one with three arguments fails to COMPILE (an extra argument is an error in TypeScript), not merely at runtime. The option existed only to suppress one sentence in the config-level refusal, whose reason for existing — a translation that reaches the object as a literal `matchName` admitting nothing — is gone, because the entry expands on both paths now.
|
|
258
|
+
4. **The `.domain` plus `tlsServerNames` refusal reads as one sentence on both paths.** What is refused is unchanged, and the refusal is still `KubernetesNetworkPolicyHostError` with nothing written; the words changed, because `15.0.0`'s config-level message described a translation that did not expand the entry — which no longer exists. A caller matching that message's text sees new text.
|
|
259
|
+
5. **A live egress object carrying a `specs` list is refused outright.** A `CiliumNetworkPolicy` carries EITHER one `spec` or a `specs` list, and a rule in either one enforces; the named comparison reads `spec.podSelector`/`spec.endpointSelector`, `spec.policyTypes` and `spec.egress` and nothing else. `15.0.0` read `spec` and reported a match, so an object whose `spec` matched the translation and whose `specs` entry stayed inside the allowlist verified and created — and under `verify: 'named-object-only'`, where nothing else looked at it, so did one whose `specs` entry allowed more. It is a mismatch now, with `KubernetesEgressPolicyMismatchError` naming `specs`. The translation never emits a `specs` list, and the shipped admission fence refuses one for the same reason, so every affected object is a hand-edited one.
|
|
260
|
+
|
|
261
|
+
**That re-apply passes, which is not a detail.** Under the default `egress.verify: 'union'` the union check reads the namespace's policies and refuses any that allows more than the translation — and with the expansion in place that has to include the named object, or the fix would refuse the very object it tells an operator to apply: the allowance this check builds records a `toFQDNs` entry by its `matchName` alone, so the expansion's `matchPattern` would read as a widening, permanently, however many times the manifest was re-applied. The check therefore leaves exactly the document the named comparison read — the one built from the object's own `spec` — unjudged, while still reading the named object for the one thing that comparison cannot answer: whether it puts this pod in egress default-deny. So a deployment whose only applied policy is the object this backend names still creates; a second policy — anything else that selects the pod and allows more — is still refused by name; and a `specs` entry is judged like any other object's rules, which is item 5's rule seen from this side.
|
|
262
|
+
|
|
263
|
+
**Why `major`, not `patch`.** Bump intent is a claim about the consumer of the version they are on, and against the published `15.0.0` three arguments carry it. (i) A deployment whose config-level allowlist contains a `.domain` entry has a create that succeeds today and fails after this upgrade until an operator re-applies — the emission it depended on was useless, but the failure the upgrade produces is one an operator sees and has to act on. (ii) A deployment whose allowlist contains one of the entries in item 2 has a create that succeeds today and fails with a named refusal after it. (iii) A consumer that constructs the public `KubernetesNetworkPolicyHostError` with its documented third argument no longer compiles, which is a removed part of the public API whatever it did. `@namzu/sandbox`'s own precedent is the same shape: the egress-kinds release was `major` because a deployment with a second policy selecting the sandbox pods began failing `create()` where it used to succeed.
|
|
264
|
+
|
|
265
|
+
**Also in this release, and already published in `15.0.0`'s notes rather than here:** `ciliumNarrowing` and the per-sandbox `setNetworkPolicy` policies. Their changeset files were consumed by the `15.0.0` release and are deleted from this branch for that reason — the text lives in the published CHANGELOG, and a pending changeset for it would publish it a second time.
|
|
266
|
+
|
|
267
|
+
**Not proven here.** That a Cilium L7 proxy admits `example.com` and its subdomains and nothing else under the emitted pair, and that an uppercase `matchName` (`API.example.com`, emitted exactly as written, as every release has) is admitted by one. This repository has no cluster with a Cilium data plane, and the `kind` cluster its Kubernetes tests use enforces no policy at all, so the translation is pinned and the ENFORCEMENT of it is not.
|
|
268
|
+
|
|
269
|
+
- 481fb8f: `container:docker` completes its hardening baseline, and a container's root
|
|
270
|
+
filesystem is now read-only.
|
|
271
|
+
|
|
272
|
+
**Before:** the backend passed `--cap-drop=ALL` and
|
|
273
|
+
`--security-opt=no-new-privileges` and nothing else of its own, so every path
|
|
274
|
+
inside the container was writable, the container's IPC namespace was whatever
|
|
275
|
+
the host daemon's `default-ipc-mode` said about sharing it, and CPU was the one
|
|
276
|
+
resource of the three that had no way to be bounded at all.
|
|
277
|
+
|
|
278
|
+
**After:** `--ipc private` and `--read-only` are applied to every container, the
|
|
279
|
+
four paths that have to stay writable are mounted `--tmpfs` (`/tmp`, `/var/tmp`,
|
|
280
|
+
`/workspace`, `/home/namzu`, each named with its reason in
|
|
281
|
+
`src/backends/docker/index.ts`), and a new `cpuLimit` renders `--cpus`.
|
|
282
|
+
|
|
283
|
+
**What breaks.** A workload that writes inside the container outside those four
|
|
284
|
+
paths and the layout's own RW binds — `/opt`, `/srv`, `/etc`, or the `HOME` of an
|
|
285
|
+
image whose user is not `namzu` — now fails with `EROFS` instead of succeeding.
|
|
286
|
+
|
|
287
|
+
A second and much narrower break, in the same class: `--ipc private` makes this
|
|
288
|
+
container's IPC namespace un-joinable. Docker's `--ipc container:<name>` is gated
|
|
289
|
+
on the target having a shared-memory directory to enter, and a container created
|
|
290
|
+
`private` has none — so a container that reached into a sandbox's shared memory,
|
|
291
|
+
semaphores or message queues that way, which it could because these sandboxes
|
|
292
|
+
were created `shareable` on any host whose daemon was configured with
|
|
293
|
+
`default-ipc-mode: shareable`, is now refused by the daemon.
|
|
294
|
+
|
|
295
|
+
Scratch also lives in RAM now rather than on the container's writable layer. A
|
|
296
|
+
temp file larger than half the host's RAM, or larger than `--memory` when the
|
|
297
|
+
host set one, fails with `ENOSPC` or is OOM-killed, where writing it to disk used
|
|
298
|
+
to succeed. Any workload that spills more than its memory budget into `/tmp`,
|
|
299
|
+
`/var/tmp` or `/workspace` is in that group.
|
|
300
|
+
Everything the reference image does keeps working: `/tmp` stays executable
|
|
301
|
+
(docker's own `--tmpfs` default is `noexec`, which would have turned
|
|
302
|
+
`gcc -o /tmp/a.out … && /tmp/a.out` into `Permission denied`), and `HOME` stays
|
|
303
|
+
writable, which is what LibreOffice, matplotlib and npm need.
|
|
304
|
+
|
|
305
|
+
**What to do.** For an image of your own, name what it needs writable in
|
|
306
|
+
`writableRootfsPaths`; each entry becomes a `--tmpfs`, so it is scratch rather
|
|
307
|
+
than persistence. For a workload that spills more scratch than its memory
|
|
308
|
+
budget, move the spill rather than the baseline: `layout.scratch` is a bind to a
|
|
309
|
+
host directory and is still disk-backed, so give the layout one on a host
|
|
310
|
+
directory with room and point the workload at it with the per-call `env` option
|
|
311
|
+
— `TMPDIR` set to that container path — which keeps the read-only root
|
|
312
|
+
filesystem, the four paths above and the resource bounds. To give the
|
|
313
|
+
filesystem back — every path inside the container writable again — set
|
|
314
|
+
`readOnlyRootfs: false`. That turns off that one control: `--ipc private` is
|
|
315
|
+
applied whatever it says, so what it produces is the argv from before plus that
|
|
316
|
+
one flag, not the argv from before.
|
|
317
|
+
|
|
318
|
+
`cpuLimit` is opt-in and has no default, deliberately: a number chosen here
|
|
319
|
+
would throttle a run that finishes inside its timeout today, and the right value
|
|
320
|
+
is a property of the host's machine. It is new surface, and no workload that
|
|
321
|
+
works today can fail because of it.
|
|
322
|
+
|
|
323
|
+
SemVer: **major**, because two defaults changed in ways a working workload can
|
|
324
|
+
fail under. The additive parts (`cpuLimit`, `writableRootfsPaths`,
|
|
325
|
+
`readOnlyRootfs`) would be a `minor` on their own.
|
|
326
|
+
|
|
3
327
|
## 15.0.0
|
|
4
328
|
|
|
5
329
|
### Major Changes
|
package/README.md
CHANGED
|
@@ -52,6 +52,7 @@ const provider = createSandboxProvider({
|
|
|
52
52
|
image: 'namzu-sandbox:latest',
|
|
53
53
|
network: 'namzu-tasks',
|
|
54
54
|
labels: { 'example.task-id': taskId },
|
|
55
|
+
cpuLimit: 2,
|
|
55
56
|
},
|
|
56
57
|
layout: {
|
|
57
58
|
outputs: { source: { type: 'hostDir', hostPath: `/srv/tasks/${taskId}/outputs` } },
|
|
@@ -63,6 +64,117 @@ const provider = createSandboxProvider({
|
|
|
63
64
|
})
|
|
64
65
|
```
|
|
65
66
|
|
|
67
|
+
## Container hardening baseline
|
|
68
|
+
|
|
69
|
+
`container:docker` confines every container it starts, and the argv it builds is
|
|
70
|
+
pinned by a test (`src/backends/docker/__tests__/hardening.test.ts`) so a change
|
|
71
|
+
to the baseline has to get past it on purpose. Each flag carries its own
|
|
72
|
+
argument in `src/backends/docker/index.ts`; in one line each:
|
|
73
|
+
|
|
74
|
+
| Flag | Why |
|
|
75
|
+
|---|---|
|
|
76
|
+
| `--cap-drop=ALL` | No Linux capability, with no re-add list. `CAP_DAC_OVERRIDE` alone walks past the layout's read-only binds, and `NET_ADMIN` is what would put a default route back on a `deny-all` network. |
|
|
77
|
+
| `--security-opt=no-new-privileges` | A setuid binary inside the image cannot re-escalate. |
|
|
78
|
+
| `--ipc private` | This container's IPC namespace is not joinable. `shareable` (docker's other daemon default) hands every container a namespace of its own as well, but leaves it joinable by name with `--ipc container:<name>`. |
|
|
79
|
+
| `--read-only` | The image is not a place the workload writes; the layout's `outputs` and `scratch` binds are separate mounts and stay writable. |
|
|
80
|
+
| `--tmpfs <path>` | The paths inside the container that do stay writable — see below. |
|
|
81
|
+
| `--memory`, `--pids-limit`, `--cpus` | Bounds the host sets through `defaultMemoryLimitMb`, `defaultMaxProcesses` and `cpuLimit`. All three are unset by default, because the right number is a property of the host's machine and of the workload. |
|
|
82
|
+
|
|
83
|
+
**What stays writable under `--read-only`.** Four paths, all `--tmpfs`:
|
|
84
|
+
`/tmp`, `/var/tmp`, `/workspace` and `/home/namzu`. The first three are where a
|
|
85
|
+
workload's own scratch goes — `/tmp` is `TMPDIR`, where pip builds wheels and
|
|
86
|
+
where a program compiled in the sandbox is run, so these mounts are deliberately
|
|
87
|
+
executable (docker's own `--tmpfs` default is `noexec`, which would turn that
|
|
88
|
+
into `Permission denied` on a file that is plainly executable). `/home/namzu` is
|
|
89
|
+
the reference image's `HOME`: LibreOffice refuses a headless conversion without a
|
|
90
|
+
writable user profile, and matplotlib, fontconfig, npm and `pip install --user`
|
|
91
|
+
all keep caches there. A host that points `image` at its own build says what
|
|
92
|
+
that image needs with `writableRootfsPaths` — the backend cannot read an image's
|
|
93
|
+
writable set, and the alternative to asking is guessing. `readOnlyRootfs: false`
|
|
94
|
+
turns that control off — the container filesystem is writable again — and turns
|
|
95
|
+
off nothing else: `--cap-drop=ALL`, `--security-opt=no-new-privileges` and
|
|
96
|
+
`--ipc private` are applied whatever it says.
|
|
97
|
+
|
|
98
|
+
Because those four paths are tmpfs, they are RAM, not the container's writable
|
|
99
|
+
layer: scratch larger than half the host's RAM (or than `--memory`, the tighter
|
|
100
|
+
of the two when the host sets one) fails with `ENOSPC` rather than spilling onto
|
|
101
|
+
the host's disk. A run that writes temp files bigger than its memory budget does
|
|
102
|
+
not have to give up the baseline over it: `layout.scratch` is a bind to a host
|
|
103
|
+
directory and stays disk-backed, so a host with room on disk mounts one there
|
|
104
|
+
and points the workload at it — `TMPDIR` set to that container path through the
|
|
105
|
+
per-call `env` option — which keeps the read-only root filesystem and the four
|
|
106
|
+
paths above. `readOnlyRootfs: false` is the last resort rather than the first:
|
|
107
|
+
it buys the container's own writable layer back at the cost of the control.
|
|
108
|
+
|
|
109
|
+
**Three controls are deliberately absent.** A seccomp profile: docker applies
|
|
110
|
+
its built-in one to every container and nothing here asks for anything looser,
|
|
111
|
+
so what is missing is a *tighter* profile, and a hand-written one cannot be
|
|
112
|
+
verified here against the reference image's toolchain — a profile that blocks a
|
|
113
|
+
syscall chromium or LibreOffice needs breaks the sandbox, which is worse than
|
|
114
|
+
the gap it closes. Set `seccomp-profile` in the daemon's `daemon.json` if you
|
|
115
|
+
want one. `--userns-remap`: it is a daemon property (`userns-remap` in
|
|
116
|
+
`daemon.json`), not a `docker run` flag — a container only chooses between the
|
|
117
|
+
namespaces the daemon already made (`--userns=host|private`), so whether a
|
|
118
|
+
remapped mapping exists is settled before this backend's argv is read and no
|
|
119
|
+
flag here could settle it. Enable it on the host and every container in this
|
|
120
|
+
tier gets uid 0 mapped to an unprivileged uid outside.
|
|
121
|
+
`--user`: supported through `runAsUser` and unset by default, because `--user`
|
|
122
|
+
overrides the image's own choice and the reference image already ends with
|
|
123
|
+
`USER namzu`.
|
|
124
|
+
|
|
125
|
+
## Container egress boundary
|
|
126
|
+
|
|
127
|
+
An egress policy of `static` or `resolver` — a host allowlist — is enforced by
|
|
128
|
+
the egress proxy running as a **sibling container**, not by the process that
|
|
129
|
+
created the sandbox. The proxy is dual-homed: `docker run` puts it on an
|
|
130
|
+
ordinary network so it has a route to the internet, and `docker network
|
|
131
|
+
connect --alias namzu-egress` adds the `--internal` network the sandbox is on.
|
|
132
|
+
The sandbox joins that internal network alone and has no route off it — a route
|
|
133
|
+
it cannot put back, because `--cap-drop=ALL` removed `NET_ADMIN`.
|
|
134
|
+
`--add-host namzu-egress:host-gateway` and the loopback proxy that needed it are
|
|
135
|
+
gone.
|
|
136
|
+
|
|
137
|
+
"No route off it" is the accurate half of that, and it is worth being exact
|
|
138
|
+
about the other: the internal network is a subnet, so the sandbox can also
|
|
139
|
+
reach whatever else a host attaches there — a second sandbox, a second
|
|
140
|
+
sandbox's proxy. Give each sandbox its own internal network when they should
|
|
141
|
+
not see one another; see
|
|
142
|
+
[docs/sdk/sandbox-egress.md](../../docs/sdk/sandbox-egress.md).
|
|
143
|
+
|
|
144
|
+
`HTTP_PROXY`, `http_proxy`, `HTTPS_PROXY`, `https_proxy` and `NO_PROXY` are
|
|
145
|
+
still set on the sandbox, and what they are has changed. They no longer permit
|
|
146
|
+
traffic, they direct it: a tool that honours them sends its request through the
|
|
147
|
+
boundary, and a tool that ignores them — a binary that does not read proxy
|
|
148
|
+
environment, `curl --noproxy '*'`, a raw socket — has nowhere to send anything
|
|
149
|
+
and fails with `Network unreachable`. Before this, that second tool reached the
|
|
150
|
+
network with the allowlist unconsulted.
|
|
151
|
+
|
|
152
|
+
Three things a host must supply, each refused at `create()` rather than
|
|
153
|
+
downgraded: an `--internal` network in `network` (`docker network create
|
|
154
|
+
--internal <name>`), `hostReachability: 'container-network'` (a published host
|
|
155
|
+
port needs a route out that this network does not have), and
|
|
156
|
+
`egressProxyImage` — the proxy image, built from
|
|
157
|
+
`packages/sandbox/egress-proxy/Dockerfile` exactly as the sandbox image is
|
|
158
|
+
built from `packages/sandbox/worker/Dockerfile`. Nothing in this repository
|
|
159
|
+
pushes an image; both are built by hand and named by tag in the config.
|
|
160
|
+
|
|
161
|
+
The proxy container carries the same baseline the sandbox does
|
|
162
|
+
(`--cap-drop=ALL`, `--security-opt=no-new-privileges`, `--ipc private`,
|
|
163
|
+
`--read-only`), because it is the process standing between untrusted code and
|
|
164
|
+
the internet. `deny-all` and `allow-all` need none of this beyond the internal
|
|
165
|
+
network `deny-all` already required.
|
|
166
|
+
|
|
167
|
+
What the boundary does not cover — domain fronting inside a `CONNECT` tunnel, a
|
|
168
|
+
`resolver` policy resolved at `create()` and `setNetworkPolicy()` rather than
|
|
169
|
+
per request, `setNetworkPolicy()` replacing the proxy container rather than
|
|
170
|
+
swapping a list in place, brokered credentials now readable by anything with
|
|
171
|
+
daemon access, the proxy's listener being reachable by whatever else shares its
|
|
172
|
+
upstream network, and everything else the sandbox shares its own internal
|
|
173
|
+
network with — is stated in
|
|
174
|
+
[docs/sdk/sandbox-egress.md](../../docs/sdk/sandbox-egress.md), along with the
|
|
175
|
+
argv-level evidence this change is verified by and the fact that no test here
|
|
176
|
+
starts a container.
|
|
177
|
+
|
|
66
178
|
## Protocol readiness and cancellation
|
|
67
179
|
|
|
68
180
|
Pass `SandboxExecOptions.signal` to stop a command on any shipped backend. A
|
|
@@ -522,6 +634,117 @@ sees the agent's own settings — `NAMZU_SANDBOX_WORKSPACE` among them. A
|
|
|
522
634
|
workload that needs a value in its terminal passes it in `env` on the
|
|
523
635
|
`openTerminal` call, which still wins over everything else, `TERM` included.
|
|
524
636
|
|
|
637
|
+
## The container worker's control API, and its token
|
|
638
|
+
|
|
639
|
+
The container backend runs `worker/server.js` — a different process from the
|
|
640
|
+
guest agent above, with a different wire: plain HTTP on a port Docker forwards,
|
|
641
|
+
not a framed stream over a socket. It serves `GET /healthz`,
|
|
642
|
+
`POST /execute`, `POST /executions/reserve`, `POST /cancel`, `POST /read-file`
|
|
643
|
+
and `POST /write-file`, and until recently it authenticated nothing: any peer
|
|
644
|
+
that could route to the container could run a command or read and write a file
|
|
645
|
+
inside it. What made that defensible was the network the container is attached
|
|
646
|
+
to, which is a property of a deployment rather than of the worker, and is
|
|
647
|
+
absent wherever a control plane can put the container on a public address.
|
|
648
|
+
|
|
649
|
+
Every route but `GET /healthz` now requires
|
|
650
|
+
`Authorization: Bearer <NAMZU_SANDBOX_TOKEN>`. A missing, empty or different
|
|
651
|
+
token is answered `401` with `{"error":"unauthorized"}` and nothing else — no
|
|
652
|
+
expected/got, no length, no hint about which routes exist — and the request
|
|
653
|
+
never reaches a handler, so a refused `/write-file` writes nothing. The
|
|
654
|
+
comparison is a fixed-width SHA-256 digest compare, so neither the value nor
|
|
655
|
+
its length is learnable by probing, and the gate runs before every dispatch,
|
|
656
|
+
including the 404, so an unauthenticated caller cannot tell a real route from a
|
|
657
|
+
missing one, or a wrong token from a missing one. `/healthz` requires no token
|
|
658
|
+
and never echoes one, because it is what the host polls before it has any other
|
|
659
|
+
business with the worker, and it answers with liveness and the protocol version
|
|
660
|
+
only. That exemption is an exact match on the whole URL, so `POST /healthz`,
|
|
661
|
+
`GET /healthz?x=1` and `GET /healthz/` are gated like anything else. It also
|
|
662
|
+
leaves exactly one bit readable without a credential, deliberately: the
|
|
663
|
+
retiring-worker check answers before the `/healthz` dispatch, so a worker that
|
|
664
|
+
has poisoned itself answers `503 {"error":"worker_retiring"}` where a healthy
|
|
665
|
+
one answers `200`. That is the drain signal, and the readiness probe — which
|
|
666
|
+
has no credential to present — is who reads it; on every route that does
|
|
667
|
+
anything, a caller without the token gets the same `401` from a retiring worker
|
|
668
|
+
and a serving one.
|
|
669
|
+
|
|
670
|
+
**The token is minted per instance, by whoever starts the container.** The
|
|
671
|
+
container backend generates 32 random bytes at `create()` time and hands the
|
|
672
|
+
value to the docker CLI in ITS environment, resolved by the valueless
|
|
673
|
+
`--env NAMZU_SANDBOX_TOKEN`. It shares a destination with three variables that
|
|
674
|
+
already exist — the workspace path and the read/write roots land in the
|
|
675
|
+
container's environment through the same `--env` mechanism — and nothing else:
|
|
676
|
+
those three are rendered in the argv as `--env K=V`, values and all, while this
|
|
677
|
+
one is valueless, which is docker's form for "take the value from the CLI's own
|
|
678
|
+
environment". So it is in no argv, `ps` on the host does not show it, and a
|
|
679
|
+
failed `docker run` renders its argv with every `--env` value redacted, in all
|
|
680
|
+
four spellings of the flag, so the error a host logs does not carry it either.
|
|
681
|
+
What it IS visible in, said plainly: the container's own config, so
|
|
682
|
+
`docker inspect <name>` shows it for the container's life to anyone who can
|
|
683
|
+
already talk to the daemon, and the worker's `/proc` inside the container to a
|
|
684
|
+
workload sharing its uid. That is why it is per-instance and why it dies with
|
|
685
|
+
the container — an image-level secret would be shared by every container ever
|
|
686
|
+
built from it and readable by anything that can pull it, and a profile-level one
|
|
687
|
+
would be shared by every instance claimed from a pool. The variable carries the
|
|
688
|
+
`NAMZU_SANDBOX_` prefix, which is what keeps it out of every command the sandbox
|
|
689
|
+
runs: the worker strips that prefix from the environment it hands to a spawned
|
|
690
|
+
command, so a sandboxed task cannot read the credential out of its own
|
|
691
|
+
environment.
|
|
692
|
+
|
|
693
|
+
**A worker with no token refuses to start on a routable address.** The bind
|
|
694
|
+
default is `0.0.0.0` and stays that way: a published container port forwards to
|
|
695
|
+
the container's interface address rather than to its loopback, so narrowing the
|
|
696
|
+
default disables the container backend instead of hardening it. The credential
|
|
697
|
+
is what makes that default defensible, so its absence fails closed:
|
|
698
|
+
|
|
699
|
+
| What the worker was given | What it does |
|
|
700
|
+
|---|---|
|
|
701
|
+
| `NAMZU_SANDBOX_TOKEN` set | Requires it on every route but `/healthz`, whatever the bind address |
|
|
702
|
+
| No token, bound to loopback (`127.0.0.1`, `::1`, `localhost`) | Starts, unauthenticated. Nothing outside the container's own network namespace can open that socket, and refusing here would break the host-beside-docker dev case while closing nothing |
|
|
703
|
+
| No token, bound anywhere else — including the `0.0.0.0` default | Refuses to start, exit 1, naming the variable and every way out |
|
|
704
|
+
| `NAMZU_SANDBOX_TOKEN` set but **empty**, or with leading/trailing whitespace | Refuses to start in every mode. An empty value is the shape an injected secret takes when the injection resolved to nothing; a padded one is read from a trimmed header and so can never be presented, leaving a worker that looks authenticated and refuses its own host for the container's life |
|
|
705
|
+
| `NAMZU_SANDBOX_TOKEN` set to something **no HTTP header can carry** — a code point above U+00FF, or a C0 control other than HTAB | Refuses to start in every mode. A header value is one byte per character, so the client's own `fetch` throws before the request leaves the host for the first, and the HTTP parser on this side drops the connection for the second. Same "boots authenticated, refuses its host" state, caught at boot instead of at the first call |
|
|
706
|
+
| No token, routable, and `NAMZU_SANDBOX_ALLOW_UNAUTHENTICATED=1` | Starts, unauthenticated, on purpose |
|
|
707
|
+
|
|
708
|
+
**The escape, and what it gives up.**
|
|
709
|
+
`NAMZU_SANDBOX_ALLOW_UNAUTHENTICATED=1` is the explicit, named way to keep an
|
|
710
|
+
existing deployment running unauthenticated. It gives up the credential
|
|
711
|
+
entirely rather than deferring it: every route but `/healthz` is then open to
|
|
712
|
+
whoever can route to the container, which is exactly the pre-token behaviour
|
|
713
|
+
and exactly what the credential was added to remove. Only `1`, `true`, `yes` and
|
|
714
|
+
`on`, in any case and with no surrounding whitespace, turn it on: `= 0` and
|
|
715
|
+
`= false` mean off, and so does everything else, because a flag whose
|
|
716
|
+
off-spelling turns it on is a trap — and so is a security escape that accepts
|
|
717
|
+
the shape a value takes when an editor adds a space to it. A configured token
|
|
718
|
+
always wins over the escape, since the escape can only mean "serve without a
|
|
719
|
+
credential", never "ignore the one I was given".
|
|
720
|
+
|
|
721
|
+
**A worker this host did not create must be provisioned with the token by
|
|
722
|
+
whoever does.** There is no channel back: the worker is handed its credential
|
|
723
|
+
at startup and never publishes it, so a warm pool, a shared container-group
|
|
724
|
+
profile or a container someone else started cannot be authenticated against by
|
|
725
|
+
a host that has no way to learn what it holds. The standby-pool backend cannot
|
|
726
|
+
send a per-instance token: the one property override its claim API admits is a
|
|
727
|
+
config map, and a config map arrives as a file mount under `/mnt/configmap`,
|
|
728
|
+
not as an environment variable — while this worker reads its token from
|
|
729
|
+
`process.env` at startup and never looks at the filesystem for one. That is a
|
|
730
|
+
channel that exists and a credential that still does not go in it, for two
|
|
731
|
+
reasons rather than none: the worker would not read it, and Microsoft's own
|
|
732
|
+
guidance is that config map values are not validated by the runtime and are not
|
|
733
|
+
where a value affecting application security belongs. A token on the shared
|
|
734
|
+
profile would be one credential for every instance claimed from the pool, and
|
|
735
|
+
this backend has no field to present one anyway — so what runs there today is a
|
|
736
|
+
private address (`subnetId`) plus the worker-side escape. The full answer, and
|
|
737
|
+
the change that would close the gap, are in
|
|
738
|
+
`docs/sdk/container-sandbox-worker.md`.
|
|
739
|
+
|
|
740
|
+
**The transport is not confidential.** `Authorization: Bearer` over plain HTTP
|
|
741
|
+
is replayable by anything on the path, so this is defence in depth behind
|
|
742
|
+
network placement, not a replacement for it. On the container backend the
|
|
743
|
+
worker is reached over Docker's port-forward on host loopback, or by DNS name
|
|
744
|
+
on a private bridge; it does not make a public-address deployment safe, which is
|
|
745
|
+
why the standby-pool backend still refuses to claim one without `subnetId`. The
|
|
746
|
+
egress proxy is unrelated to any of this and checks nothing inbound.
|
|
747
|
+
|
|
525
748
|
## Firecracker workspace channels
|
|
526
749
|
|
|
527
750
|
The Firecracker backend exposes two optional same-sandbox channels. Call
|
|
@@ -132,10 +132,28 @@ export declare function assertEnforceable(options: SandboxBackendOptions): void;
|
|
|
132
132
|
* so nothing looks wrong. A caller who never heard of `subnetId` gets a
|
|
133
133
|
* working sandbox on the internet and no signal at all.
|
|
134
134
|
*
|
|
135
|
-
* What is on that address matters: `worker/server.js`
|
|
136
|
-
*
|
|
137
|
-
*
|
|
138
|
-
*
|
|
135
|
+
* What is on that address matters: `worker/server.js` binds every interface
|
|
136
|
+
* and, inside a private network, that was the boundary doing the work — the
|
|
137
|
+
* worker authenticated nobody. With a public address there is no boundary
|
|
138
|
+
* left, and the worker's `/execute` is reachable by anyone.
|
|
139
|
+
*
|
|
140
|
+
* The worker now requires a per-instance `Authorization: Bearer` token on
|
|
141
|
+
* every route but `/healthz`, and REFUSES TO START on a routable bind when
|
|
142
|
+
* it has none. This backend cannot supply one, and the reason this comment
|
|
143
|
+
* used to give for that was wrong: it said the claim API's single admitted
|
|
144
|
+
* override — a config map — is no channel into the container group. It is
|
|
145
|
+
* one, on Linux a file mount under `/mnt/configmap/<containername>/<key>`
|
|
146
|
+
* carrying caller-supplied per-claim values. The credential is declined
|
|
147
|
+
* anyway, for two reasons that survive their source: the worker reads its
|
|
148
|
+
* token from `process.env` at startup and never looks for a file, so a
|
|
149
|
+
* mounted value is not read; and Microsoft's own guidance is that config
|
|
150
|
+
* map values are not validated by the runtime and that values affecting
|
|
151
|
+
* application security "should be made available to the container using
|
|
152
|
+
* environment variables". The full answer, and the change that would close
|
|
153
|
+
* the gap, is in `docs/sdk/container-sandbox-worker.md`. What that means
|
|
154
|
+
* for an operator is written down there too, rather than left to be
|
|
155
|
+
* discovered at startup; this refusal is unchanged, and now stands in
|
|
156
|
+
* front of a worker that would refuse the routable case itself.
|
|
139
157
|
*
|
|
140
158
|
* Defaulting to refusal rather than to a warning, because a warning on a
|
|
141
159
|
* path that otherwise succeeds is read once and never again.
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"file":"index.d.ts","sourceRoot":"","sources":["../../../src/backends/aci-standby-pool/index.ts"],"names":[],"mappings":"AAAA;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;GAwCG;AAEH,OAAO,KAAK,EAEX,8BAA8B,EAU9B,MAAM,YAAY,CAAA;AAGnB,OAAO,KAAK,EAAE,cAAc,EAAE,qBAAqB,EAAE,MAAM,gBAAgB,CAAA;AAc3E;;;;;GAKG;AACH,MAAM,MAAM,gBAAgB,GAAG,MAAM,OAAO,CAAC,MAAM,CAAC,CAAA;AAEpD,MAAM,WAAW,mCAAmC;IACnD,QAAQ,CAAC,cAAc,EAAE,MAAM,CAAA;IAC/B,QAAQ,CAAC,aAAa,EAAE,MAAM,CAAA;IAC9B,QAAQ,CAAC,QAAQ,EAAE,MAAM,CAAA;IACzB;;;;OAIG;IACH,QAAQ,CAAC,qBAAqB,EAAE,MAAM,CAAA;IACtC;;;OAGG;IACH,QAAQ,CAAC,+BAA+B,EAAE,MAAM,CAAA;IAChD;;;OAGG;IACH,QAAQ,CAAC,6BAA6B,CAAC,EAAE,MAAM,CAAA;IAC/C;;;OAGG;IACH,QAAQ,CAAC,MAAM,EAAE,8BAA8B,CAAA;IAC/C;;OAEG;IACH,QAAQ,CAAC,WAAW,EAAE,gBAAgB,CAAA;IACtC;;;;;;OAMG;IACH,QAAQ,CAAC,QAAQ,CAAC,EAAE,MAAM,CAAA;IAC1B;;;;;;;;;;;;;OAaG;IACH,QAAQ,CAAC,kBAAkB,CAAC,EAAE,OAAO,CAAA;IACrC,QAAQ,CAAC,mBAAmB,CAAC,EAAE,MAAM,CAAA;IACrC,QAAQ,CAAC,cAAc,CAAC,EAAE,MAAM,CAAA;IAChC;;OAEG;IACH,QAAQ,CAAC,UAAU,CAAC,EAAE,MAAM,CAAA;IAC5B,QAAQ,CAAC,aAAa,CAAC,EAAE,MAAM,CAAA;IAC/B;;;;;;OAMG;IACH,QAAQ,CAAC,mBAAmB,CAAC,EAAE,MAAM,CAAA;CACrC;AASD;;;;GAIG;AACH,wBAAgB,0BAA0B,CACzC,MAAM,EAAE,mCAAmC,GACzC,cAAc,CAiBhB;AAoMD,wBAAgB,iBAAiB,CAAC,OAAO,EAAE,qBAAqB,GAAG,IAAI,CActE;AAED
|
|
1
|
+
{"version":3,"file":"index.d.ts","sourceRoot":"","sources":["../../../src/backends/aci-standby-pool/index.ts"],"names":[],"mappings":"AAAA;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;GAwCG;AAEH,OAAO,KAAK,EAEX,8BAA8B,EAU9B,MAAM,YAAY,CAAA;AAGnB,OAAO,KAAK,EAAE,cAAc,EAAE,qBAAqB,EAAE,MAAM,gBAAgB,CAAA;AAc3E;;;;;GAKG;AACH,MAAM,MAAM,gBAAgB,GAAG,MAAM,OAAO,CAAC,MAAM,CAAC,CAAA;AAEpD,MAAM,WAAW,mCAAmC;IACnD,QAAQ,CAAC,cAAc,EAAE,MAAM,CAAA;IAC/B,QAAQ,CAAC,aAAa,EAAE,MAAM,CAAA;IAC9B,QAAQ,CAAC,QAAQ,EAAE,MAAM,CAAA;IACzB;;;;OAIG;IACH,QAAQ,CAAC,qBAAqB,EAAE,MAAM,CAAA;IACtC;;;OAGG;IACH,QAAQ,CAAC,+BAA+B,EAAE,MAAM,CAAA;IAChD;;;OAGG;IACH,QAAQ,CAAC,6BAA6B,CAAC,EAAE,MAAM,CAAA;IAC/C;;;OAGG;IACH,QAAQ,CAAC,MAAM,EAAE,8BAA8B,CAAA;IAC/C;;OAEG;IACH,QAAQ,CAAC,WAAW,EAAE,gBAAgB,CAAA;IACtC;;;;;;OAMG;IACH,QAAQ,CAAC,QAAQ,CAAC,EAAE,MAAM,CAAA;IAC1B;;;;;;;;;;;;;OAaG;IACH,QAAQ,CAAC,kBAAkB,CAAC,EAAE,OAAO,CAAA;IACrC,QAAQ,CAAC,mBAAmB,CAAC,EAAE,MAAM,CAAA;IACrC,QAAQ,CAAC,cAAc,CAAC,EAAE,MAAM,CAAA;IAChC;;OAEG;IACH,QAAQ,CAAC,UAAU,CAAC,EAAE,MAAM,CAAA;IAC5B,QAAQ,CAAC,aAAa,CAAC,EAAE,MAAM,CAAA;IAC/B;;;;;;OAMG;IACH,QAAQ,CAAC,mBAAmB,CAAC,EAAE,MAAM,CAAA;CACrC;AASD;;;;GAIG;AACH,wBAAgB,0BAA0B,CACzC,MAAM,EAAE,mCAAmC,GACzC,cAAc,CAiBhB;AAoMD,wBAAgB,iBAAiB,CAAC,OAAO,EAAE,qBAAqB,GAAG,IAAI,CActE;AAED;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;GAkCG;AACH,wBAAgB,0BAA0B,CAAC,MAAM,EAAE;IAClD,QAAQ,CAAC,EAAE,MAAM,CAAA;IACjB,kBAAkB,CAAC,EAAE,OAAO,CAAA;CAC5B,GAAG,IAAI,CAOP"}
|