@namzu/sandbox 15.0.0 → 16.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +86 -0
- package/README.md +59 -0
- package/dist/backends/docker/index.d.ts +169 -6
- package/dist/backends/docker/index.d.ts.map +1 -1
- package/dist/backends/docker/index.js +480 -84
- package/dist/backends/docker/index.js.map +1 -1
- package/dist/backends/kubernetes/egress-policy.d.ts +193 -102
- package/dist/backends/kubernetes/egress-policy.d.ts.map +1 -1
- package/dist/backends/kubernetes/egress-policy.js +321 -146
- package/dist/backends/kubernetes/egress-policy.js.map +1 -1
- package/dist/backends/kubernetes/per-sandbox-policy.d.ts +6 -6
- package/dist/backends/kubernetes/per-sandbox-policy.d.ts.map +1 -1
- package/dist/backends/kubernetes/per-sandbox-policy.js +21 -53
- package/dist/backends/kubernetes/per-sandbox-policy.js.map +1 -1
- package/dist/index.d.ts +63 -5
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +33 -5
- package/dist/index.js.map +1 -1
- package/package.json +1 -1
- package/src/backends/docker/index.ts +595 -99
- package/src/backends/kubernetes/egress-policy.ts +455 -187
- package/src/backends/kubernetes/per-sandbox-policy.ts +21 -66
- package/src/index.ts +73 -5
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,91 @@
|
|
|
1
1
|
# @namzu/sandbox
|
|
2
2
|
|
|
3
|
+
## 16.0.0
|
|
4
|
+
|
|
5
|
+
### Major Changes
|
|
6
|
+
|
|
7
|
+
- 11f7ea2: **A `.domain` allowlist entry is expanded on the config-level egress translation, a `specs` list on the applied object is refused, and a deployment that applied the old object must re-apply it before its next `create()`.**
|
|
8
|
+
|
|
9
|
+
Everything here is stated against **`@namzu/sandbox@15.0.0`, the published release this one ships on top of** — its bytes as downloaded from the registry, not this branch's base. `15.0.0` emitted the config-level `.domain` entry verbatim and said so itself: its release notes call "that a config-level `.domain` entry matches no answer at all" a pre-existing defect "left exactly as it was and deferred to its own change". This is that change, and the deferral is what the migration below is for.
|
|
10
|
+
|
|
11
|
+
`config.egress.policy = { kind: 'static', allowedHosts: ['.example.com'] }` under `engine: 'cilium'` therefore emits `toFQDNs: [{ matchName: example.com }, { matchPattern: '*.example.com' }]` where `15.0.0` emitted `toFQDNs: [{ matchName: '.example.com' }]` — the entry VERBATIM, leading dot included. No DNS answer carries a name with a leading dot, so the old object matched nothing, and nothing refused it on the way in: the shipped admission fence's `matchConditions` scope it to the sandbox host's ServiceAccount, so an operator applying the config-level object is not matched by it at all. `verifyEgressPolicyApplied` then deep-equalled what was sent, the create reported success, and the domain and every subdomain it was asked to allow were DENIED. With `ciliumNarrowing.dnsNames` on it was worse than useless, because the DNS-visibility rule carried the same unmatched names — the DNS proxy denied the lookup the `toFQDNs` half needs before that half could ever learn an address.
|
|
12
|
+
|
|
13
|
+
`SandboxNetworkPolicy.allowedHosts` has always said `.example.com` is the domain **and its subdomains** — `packages/sdk/src/types/sandbox/index.ts` — and `src/egress/allowlist.ts` implements exactly that for the docker backend's proxy. The Kubernetes PER-SANDBOX translation already honoured it. The config-level translation now reuses the same helper, so the entry becomes `matchName: example.com` plus `matchPattern: '*.example.com'`, and the DNS-visibility rule carries the names that expansion admits — `matchPattern: '*.<host>.<suffix>'` under every cluster search suffix, which is what the per-sandbox path already emitted.
|
|
14
|
+
|
|
15
|
+
**Every consumer-visible change, against `15.0.0`.** Nothing changes for an allowlist with no leading-dot entry: the expansion returns `[{ matchName: entry }]` for an exact host, the DNS rule's wildcard branch is keyed on the leading dot alone, so every other entry's bytes are identical, with and without `dnsNames`, and no entry's letter case is rewritten. The rest is this list, and each item is a create or a compile that behaves differently after the upgrade:
|
|
16
|
+
|
|
17
|
+
1. **A config-level `.domain` entry changes the object it emits**, so `verifyEgressPolicyApplied` — which compares the applied object to the translation exactly — refuses your next `createKubernetesWorkspace` or `create()` with `KubernetesEgressPolicyMismatchError` until you **re-apply the manifest this backend now translates**: the same `kubectl apply` you ran for the old one, against a cluster object that admitted nothing. That is the whole migration.
|
|
18
|
+
2. **A config-level allowlist entry that is not a hostname is refused, by name.** `.com`, `.`, `..example.com`, `*`, `*.example.com`, a URL, a port suffix and an IP literal were each passed straight into the emitted `toFQDNs` by `15.0.0`: the per-sandbox writer validated them, the config-level translation did not. Refused now, with `KubernetesNetworkPolicyHostError` and nothing written, because both halves of what they used to do are wrong — as emitted by `15.0.0` they became a `matchName` that matched nothing, and with the expansion above `.com` would instead emit `matchPattern: '*.com'` and grant every name under a public suffix, silently, on a path no admission policy covers. `.example.com`, `example.com` and `api.example.com` are all still accepted, and a `resolver` policy's returned hosts are held to the same grammar. **A create whose allowlist carries one of these works on `15.0.0` and fails after this upgrade**, until the entry is corrected.
|
|
19
|
+
3. **`KubernetesNetworkPolicyHostError` loses its third constructor argument.** `15.0.0` ships `constructor(host, reason, options?: { stateHostnameGrammar?: boolean })` on that class, which is exported from the package root; this release ships `constructor(host, reason)`. A consumer that constructed one with three arguments fails to COMPILE (an extra argument is an error in TypeScript), not merely at runtime. The option existed only to suppress one sentence in the config-level refusal, whose reason for existing — a translation that reaches the object as a literal `matchName` admitting nothing — is gone, because the entry expands on both paths now.
|
|
20
|
+
4. **The `.domain` plus `tlsServerNames` refusal reads as one sentence on both paths.** What is refused is unchanged, and the refusal is still `KubernetesNetworkPolicyHostError` with nothing written; the words changed, because `15.0.0`'s config-level message described a translation that did not expand the entry — which no longer exists. A caller matching that message's text sees new text.
|
|
21
|
+
5. **A live egress object carrying a `specs` list is refused outright.** A `CiliumNetworkPolicy` carries EITHER one `spec` or a `specs` list, and a rule in either one enforces; the named comparison reads `spec.podSelector`/`spec.endpointSelector`, `spec.policyTypes` and `spec.egress` and nothing else. `15.0.0` read `spec` and reported a match, so an object whose `spec` matched the translation and whose `specs` entry stayed inside the allowlist verified and created — and under `verify: 'named-object-only'`, where nothing else looked at it, so did one whose `specs` entry allowed more. It is a mismatch now, with `KubernetesEgressPolicyMismatchError` naming `specs`. The translation never emits a `specs` list, and the shipped admission fence refuses one for the same reason, so every affected object is a hand-edited one.
|
|
22
|
+
|
|
23
|
+
**That re-apply passes, which is not a detail.** Under the default `egress.verify: 'union'` the union check reads the namespace's policies and refuses any that allows more than the translation — and with the expansion in place that has to include the named object, or the fix would refuse the very object it tells an operator to apply: the allowance this check builds records a `toFQDNs` entry by its `matchName` alone, so the expansion's `matchPattern` would read as a widening, permanently, however many times the manifest was re-applied. The check therefore leaves exactly the document the named comparison read — the one built from the object's own `spec` — unjudged, while still reading the named object for the one thing that comparison cannot answer: whether it puts this pod in egress default-deny. So a deployment whose only applied policy is the object this backend names still creates; a second policy — anything else that selects the pod and allows more — is still refused by name; and a `specs` entry is judged like any other object's rules, which is item 5's rule seen from this side.
|
|
24
|
+
|
|
25
|
+
**Why `major`, not `patch`.** Bump intent is a claim about the consumer of the version they are on, and against the published `15.0.0` three arguments carry it. (i) A deployment whose config-level allowlist contains a `.domain` entry has a create that succeeds today and fails after this upgrade until an operator re-applies — the emission it depended on was useless, but the failure the upgrade produces is one an operator sees and has to act on. (ii) A deployment whose allowlist contains one of the entries in item 2 has a create that succeeds today and fails with a named refusal after it. (iii) A consumer that constructs the public `KubernetesNetworkPolicyHostError` with its documented third argument no longer compiles, which is a removed part of the public API whatever it did. `@namzu/sandbox`'s own precedent is the same shape: the egress-kinds release was `major` because a deployment with a second policy selecting the sandbox pods began failing `create()` where it used to succeed.
|
|
26
|
+
|
|
27
|
+
**Also in this release, and already published in `15.0.0`'s notes rather than here:** `ciliumNarrowing` and the per-sandbox `setNetworkPolicy` policies. Their changeset files were consumed by the `15.0.0` release and are deleted from this branch for that reason — the text lives in the published CHANGELOG, and a pending changeset for it would publish it a second time.
|
|
28
|
+
|
|
29
|
+
**Not proven here.** That a Cilium L7 proxy admits `example.com` and its subdomains and nothing else under the emitted pair, and that an uppercase `matchName` (`API.example.com`, emitted exactly as written, as every release has) is admitted by one. This repository has no cluster with a Cilium data plane, and the `kind` cluster its Kubernetes tests use enforces no policy at all, so the translation is pinned and the ENFORCEMENT of it is not.
|
|
30
|
+
|
|
31
|
+
- 481fb8f: `container:docker` completes its hardening baseline, and a container's root
|
|
32
|
+
filesystem is now read-only.
|
|
33
|
+
|
|
34
|
+
**Before:** the backend passed `--cap-drop=ALL` and
|
|
35
|
+
`--security-opt=no-new-privileges` and nothing else of its own, so every path
|
|
36
|
+
inside the container was writable, the container's IPC namespace was whatever
|
|
37
|
+
the host daemon's `default-ipc-mode` said about sharing it, and CPU was the one
|
|
38
|
+
resource of the three that had no way to be bounded at all.
|
|
39
|
+
|
|
40
|
+
**After:** `--ipc private` and `--read-only` are applied to every container, the
|
|
41
|
+
four paths that have to stay writable are mounted `--tmpfs` (`/tmp`, `/var/tmp`,
|
|
42
|
+
`/workspace`, `/home/namzu`, each named with its reason in
|
|
43
|
+
`src/backends/docker/index.ts`), and a new `cpuLimit` renders `--cpus`.
|
|
44
|
+
|
|
45
|
+
**What breaks.** A workload that writes inside the container outside those four
|
|
46
|
+
paths and the layout's own RW binds — `/opt`, `/srv`, `/etc`, or the `HOME` of an
|
|
47
|
+
image whose user is not `namzu` — now fails with `EROFS` instead of succeeding.
|
|
48
|
+
|
|
49
|
+
A second and much narrower break, in the same class: `--ipc private` makes this
|
|
50
|
+
container's IPC namespace un-joinable. Docker's `--ipc container:<name>` is gated
|
|
51
|
+
on the target having a shared-memory directory to enter, and a container created
|
|
52
|
+
`private` has none — so a container that reached into a sandbox's shared memory,
|
|
53
|
+
semaphores or message queues that way, which it could because these sandboxes
|
|
54
|
+
were created `shareable` on any host whose daemon was configured with
|
|
55
|
+
`default-ipc-mode: shareable`, is now refused by the daemon.
|
|
56
|
+
|
|
57
|
+
Scratch also lives in RAM now rather than on the container's writable layer. A
|
|
58
|
+
temp file larger than half the host's RAM, or larger than `--memory` when the
|
|
59
|
+
host set one, fails with `ENOSPC` or is OOM-killed, where writing it to disk used
|
|
60
|
+
to succeed. Any workload that spills more than its memory budget into `/tmp`,
|
|
61
|
+
`/var/tmp` or `/workspace` is in that group.
|
|
62
|
+
Everything the reference image does keeps working: `/tmp` stays executable
|
|
63
|
+
(docker's own `--tmpfs` default is `noexec`, which would have turned
|
|
64
|
+
`gcc -o /tmp/a.out … && /tmp/a.out` into `Permission denied`), and `HOME` stays
|
|
65
|
+
writable, which is what LibreOffice, matplotlib and npm need.
|
|
66
|
+
|
|
67
|
+
**What to do.** For an image of your own, name what it needs writable in
|
|
68
|
+
`writableRootfsPaths`; each entry becomes a `--tmpfs`, so it is scratch rather
|
|
69
|
+
than persistence. For a workload that spills more scratch than its memory
|
|
70
|
+
budget, move the spill rather than the baseline: `layout.scratch` is a bind to a
|
|
71
|
+
host directory and is still disk-backed, so give the layout one on a host
|
|
72
|
+
directory with room and point the workload at it with the per-call `env` option
|
|
73
|
+
— `TMPDIR` set to that container path — which keeps the read-only root
|
|
74
|
+
filesystem, the four paths above and the resource bounds. To give the
|
|
75
|
+
filesystem back — every path inside the container writable again — set
|
|
76
|
+
`readOnlyRootfs: false`. That turns off that one control: `--ipc private` is
|
|
77
|
+
applied whatever it says, so what it produces is the argv from before plus that
|
|
78
|
+
one flag, not the argv from before.
|
|
79
|
+
|
|
80
|
+
`cpuLimit` is opt-in and has no default, deliberately: a number chosen here
|
|
81
|
+
would throttle a run that finishes inside its timeout today, and the right value
|
|
82
|
+
is a property of the host's machine. It is new surface, and no workload that
|
|
83
|
+
works today can fail because of it.
|
|
84
|
+
|
|
85
|
+
SemVer: **major**, because two defaults changed in ways a working workload can
|
|
86
|
+
fail under. The additive parts (`cpuLimit`, `writableRootfsPaths`,
|
|
87
|
+
`readOnlyRootfs`) would be a `minor` on their own.
|
|
88
|
+
|
|
3
89
|
## 15.0.0
|
|
4
90
|
|
|
5
91
|
### Major Changes
|
package/README.md
CHANGED
|
@@ -52,6 +52,7 @@ const provider = createSandboxProvider({
|
|
|
52
52
|
image: 'namzu-sandbox:latest',
|
|
53
53
|
network: 'namzu-tasks',
|
|
54
54
|
labels: { 'example.task-id': taskId },
|
|
55
|
+
cpuLimit: 2,
|
|
55
56
|
},
|
|
56
57
|
layout: {
|
|
57
58
|
outputs: { source: { type: 'hostDir', hostPath: `/srv/tasks/${taskId}/outputs` } },
|
|
@@ -63,6 +64,64 @@ const provider = createSandboxProvider({
|
|
|
63
64
|
})
|
|
64
65
|
```
|
|
65
66
|
|
|
67
|
+
## Container hardening baseline
|
|
68
|
+
|
|
69
|
+
`container:docker` confines every container it starts, and the argv it builds is
|
|
70
|
+
pinned by a test (`src/backends/docker/__tests__/hardening.test.ts`) so a change
|
|
71
|
+
to the baseline has to get past it on purpose. Each flag carries its own
|
|
72
|
+
argument in `src/backends/docker/index.ts`; in one line each:
|
|
73
|
+
|
|
74
|
+
| Flag | Why |
|
|
75
|
+
|---|---|
|
|
76
|
+
| `--cap-drop=ALL` | No Linux capability, with no re-add list. `CAP_DAC_OVERRIDE` alone walks past the layout's read-only binds, and `NET_ADMIN` is what would put a default route back on a `deny-all` network. |
|
|
77
|
+
| `--security-opt=no-new-privileges` | A setuid binary inside the image cannot re-escalate. |
|
|
78
|
+
| `--ipc private` | This container's IPC namespace is not joinable. `shareable` (docker's other daemon default) hands every container a namespace of its own as well, but leaves it joinable by name with `--ipc container:<name>`. |
|
|
79
|
+
| `--read-only` | The image is not a place the workload writes; the layout's `outputs` and `scratch` binds are separate mounts and stay writable. |
|
|
80
|
+
| `--tmpfs <path>` | The paths inside the container that do stay writable — see below. |
|
|
81
|
+
| `--memory`, `--pids-limit`, `--cpus` | Bounds the host sets through `defaultMemoryLimitMb`, `defaultMaxProcesses` and `cpuLimit`. All three are unset by default, because the right number is a property of the host's machine and of the workload. |
|
|
82
|
+
|
|
83
|
+
**What stays writable under `--read-only`.** Four paths, all `--tmpfs`:
|
|
84
|
+
`/tmp`, `/var/tmp`, `/workspace` and `/home/namzu`. The first three are where a
|
|
85
|
+
workload's own scratch goes — `/tmp` is `TMPDIR`, where pip builds wheels and
|
|
86
|
+
where a program compiled in the sandbox is run, so these mounts are deliberately
|
|
87
|
+
executable (docker's own `--tmpfs` default is `noexec`, which would turn that
|
|
88
|
+
into `Permission denied` on a file that is plainly executable). `/home/namzu` is
|
|
89
|
+
the reference image's `HOME`: LibreOffice refuses a headless conversion without a
|
|
90
|
+
writable user profile, and matplotlib, fontconfig, npm and `pip install --user`
|
|
91
|
+
all keep caches there. A host that points `image` at its own build says what
|
|
92
|
+
that image needs with `writableRootfsPaths` — the backend cannot read an image's
|
|
93
|
+
writable set, and the alternative to asking is guessing. `readOnlyRootfs: false`
|
|
94
|
+
turns that control off — the container filesystem is writable again — and turns
|
|
95
|
+
off nothing else: `--cap-drop=ALL`, `--security-opt=no-new-privileges` and
|
|
96
|
+
`--ipc private` are applied whatever it says.
|
|
97
|
+
|
|
98
|
+
Because those four paths are tmpfs, they are RAM, not the container's writable
|
|
99
|
+
layer: scratch larger than half the host's RAM (or than `--memory`, the tighter
|
|
100
|
+
of the two when the host sets one) fails with `ENOSPC` rather than spilling onto
|
|
101
|
+
the host's disk. A run that writes temp files bigger than its memory budget does
|
|
102
|
+
not have to give up the baseline over it: `layout.scratch` is a bind to a host
|
|
103
|
+
directory and stays disk-backed, so a host with room on disk mounts one there
|
|
104
|
+
and points the workload at it — `TMPDIR` set to that container path through the
|
|
105
|
+
per-call `env` option — which keeps the read-only root filesystem and the four
|
|
106
|
+
paths above. `readOnlyRootfs: false` is the last resort rather than the first:
|
|
107
|
+
it buys the container's own writable layer back at the cost of the control.
|
|
108
|
+
|
|
109
|
+
**Three controls are deliberately absent.** A seccomp profile: docker applies
|
|
110
|
+
its built-in one to every container and nothing here asks for anything looser,
|
|
111
|
+
so what is missing is a *tighter* profile, and a hand-written one cannot be
|
|
112
|
+
verified here against the reference image's toolchain — a profile that blocks a
|
|
113
|
+
syscall chromium or LibreOffice needs breaks the sandbox, which is worse than
|
|
114
|
+
the gap it closes. Set `seccomp-profile` in the daemon's `daemon.json` if you
|
|
115
|
+
want one. `--userns-remap`: it is a daemon property (`userns-remap` in
|
|
116
|
+
`daemon.json`), not a `docker run` flag — a container only chooses between the
|
|
117
|
+
namespaces the daemon already made (`--userns=host|private`), so whether a
|
|
118
|
+
remapped mapping exists is settled before this backend's argv is read and no
|
|
119
|
+
flag here could settle it. Enable it on the host and every container in this
|
|
120
|
+
tier gets uid 0 mapped to an unprivileged uid outside.
|
|
121
|
+
`--user`: supported through `runAsUser` and unset by default, because `--user`
|
|
122
|
+
overrides the image's own choice and the reference image already ends with
|
|
123
|
+
`USER namzu`.
|
|
124
|
+
|
|
66
125
|
## Protocol readiness and cancellation
|
|
67
126
|
|
|
68
127
|
Pass `SandboxExecOptions.signal` to stop a command on any shipped backend. A
|
|
@@ -22,7 +22,7 @@
|
|
|
22
22
|
*/
|
|
23
23
|
import { type ContainerSandboxLayout, type ResolvedContainerSandboxLayout } from '@namzu/sdk';
|
|
24
24
|
import type { BrokeredCredential, EgressProxyOptions } from '../../egress/index.js';
|
|
25
|
-
import { type EgressPolicy, type SandboxBackend } from '../../index.js';
|
|
25
|
+
import { type EgressPolicy, type SandboxBackend, type SandboxBackendOptions } from '../../index.js';
|
|
26
26
|
/**
|
|
27
27
|
* Backend-specific tuning. Most hosts use the defaults; advanced
|
|
28
28
|
* deployments override `image` to point at their own pre-built
|
|
@@ -47,13 +47,66 @@ export interface DockerBackendInternalConfig {
|
|
|
47
47
|
/**
|
|
48
48
|
* `--user` value for the container, e.g. `'1000:1000'` or `'nobody'`.
|
|
49
49
|
*
|
|
50
|
-
* Left unset by default because
|
|
51
|
-
*
|
|
52
|
-
*
|
|
53
|
-
* non-root
|
|
54
|
-
*
|
|
50
|
+
* Left unset by default because `--user` does not ADD a non-root user, it
|
|
51
|
+
* OVERRIDES the image's own choice of one. The reference image ends with
|
|
52
|
+
* `USER namzu` (uid 1001, its `/workspace` chowned to match), so this
|
|
53
|
+
* backend's default is already non-root for the image it ships — a
|
|
54
|
+
* hard-coded uid here would replace that with a guess, and the guess is
|
|
55
|
+
* wrong for any image whose files are owned by someone else, which
|
|
56
|
+
* surfaces as `EACCES` on a path the workload was told it could write.
|
|
57
|
+
* Set it when the image does not declare a user of its own, or when the
|
|
58
|
+
* host wants a different one than it declares.
|
|
55
59
|
*/
|
|
56
60
|
readonly runAsUser?: string;
|
|
61
|
+
/**
|
|
62
|
+
* CPU cores the container may use, rendered as `--cpus`. Unset by default.
|
|
63
|
+
*
|
|
64
|
+
* `--memory` and `--pids-limit` bound what a workload can take from the
|
|
65
|
+
* host, and CPU had no equivalent at all — no default, no knob — which
|
|
66
|
+
* reads as an oversight rather than a decision. It stays unset for the
|
|
67
|
+
* same reason neither of those two has a numeric default: the right value
|
|
68
|
+
* is a property of the host's machine and of what the workload is for, and
|
|
69
|
+
* any number this backend picked would silently throttle a run that
|
|
70
|
+
* finishes inside its timeout today. A host that wants the bound says what
|
|
71
|
+
* it is; the value is a decimal (`--cpus 1.5` is one and a half cores'
|
|
72
|
+
* worth of time, not a rounding).
|
|
73
|
+
*
|
|
74
|
+
* It lives on this config rather than beside `memoryLimitMb` on the
|
|
75
|
+
* per-call options because the documented deployment constructs one
|
|
76
|
+
* provider per task, so construction time IS per-task — and a control
|
|
77
|
+
* added to the tier-agnostic per-call shape would have to be refused by
|
|
78
|
+
* the ACI and kubernetes backends, which cannot apply a per-sandbox CPU
|
|
79
|
+
* limit any more than they can apply the memory and process ones.
|
|
80
|
+
*/
|
|
81
|
+
readonly cpuLimit?: number;
|
|
82
|
+
/**
|
|
83
|
+
* Mount the container's root filesystem read-only. Default `true`.
|
|
84
|
+
*
|
|
85
|
+
* See {@link HARDENING_ARGS} for why the default is on and
|
|
86
|
+
* {@link renderWritableRootfsArgs} for the paths that stay writable while
|
|
87
|
+
* it is. Set it to `false` to make every path inside the container
|
|
88
|
+
* writable again, which is what a host whose image writes somewhere the
|
|
89
|
+
* writable set cannot describe needs, and which is why the switch exists
|
|
90
|
+
* instead of an unwritten rule that the baseline is absolute. It turns off
|
|
91
|
+
* that one control and nothing else: `--cap-drop=ALL`,
|
|
92
|
+
* `--security-opt=no-new-privileges` and `--ipc private` are applied
|
|
93
|
+
* whatever this says. It is a config field rather than an argument so that
|
|
94
|
+
* turning it off is a line somebody wrote on purpose, and not the default
|
|
95
|
+
* anyone gets by not looking.
|
|
96
|
+
*/
|
|
97
|
+
readonly readOnlyRootfs?: boolean;
|
|
98
|
+
/**
|
|
99
|
+
* Extra paths to keep writable under `--read-only`, each mounted `--tmpfs`.
|
|
100
|
+
*
|
|
101
|
+
* The default set ({@link DEFAULT_WRITABLE_ROOTFS_PATHS}) is the reference
|
|
102
|
+
* image's needs, read off its Dockerfile; this is how a host that points
|
|
103
|
+
* `image` somewhere else says what ITS image needs, because the backend
|
|
104
|
+
* cannot read that out of an image and guessing is what these paths would
|
|
105
|
+
* otherwise be. A path the layout already mounts is refused rather than
|
|
106
|
+
* mounted twice (`Duplicate mount point`), and setting this at all beside
|
|
107
|
+
* `readOnlyRootfs: false` is refused as a contradiction.
|
|
108
|
+
*/
|
|
109
|
+
readonly writableRootfsPaths?: readonly string[];
|
|
57
110
|
/**
|
|
58
111
|
* Credentials the egress proxy stamps on, per host.
|
|
59
112
|
*
|
|
@@ -207,6 +260,116 @@ export declare function assertNetworkCarriesThePolicy(network: string, reachabil
|
|
|
207
260
|
* receives is the failure this shape exists to make testable.
|
|
208
261
|
*/
|
|
209
262
|
export declare function egressProxyOptions(config: Pick<DockerBackendInternalConfig, 'brokeredCredentials' | 'allowInwardFor'>, policy: EgressPolicy): EgressProxyOptions;
|
|
263
|
+
/** The backend config the hardening flags are rendered from. */
|
|
264
|
+
export type DockerHardeningConfig = Pick<DockerBackendInternalConfig, 'cpuLimit' | 'layout' | 'readOnlyRootfs' | 'writableRootfsPaths'>;
|
|
265
|
+
/**
|
|
266
|
+
* Refuse rootfs options that cannot both be honoured.
|
|
267
|
+
*
|
|
268
|
+
* `writableRootfsPaths` beside `readOnlyRootfs: false` is a contradiction: with
|
|
269
|
+
* a writable root filesystem every path is already writable, so the tmpfs
|
|
270
|
+
* mounts would either be dropped (a control accepted and not applied) or take a
|
|
271
|
+
* directory off the image for no reason. Refusing is the honest answer, and it
|
|
272
|
+
* is the same one the sibling backends give a per-sandbox control they cannot
|
|
273
|
+
* express.
|
|
274
|
+
*
|
|
275
|
+
* Called at construction and again where the argv is built, so a config that
|
|
276
|
+
* reaches `create()` by some path other than `buildDockerBackend` is refused
|
|
277
|
+
* too.
|
|
278
|
+
*/
|
|
279
|
+
export declare function assertRootfsOptionsAreCoherent(config: DockerHardeningConfig): void;
|
|
280
|
+
/**
|
|
281
|
+
* Refuse a `--cpus` value that cannot mean what it says.
|
|
282
|
+
*
|
|
283
|
+
* This covers non-finite and non-positive values and does NOT claim to cover
|
|
284
|
+
* every bound the daemon would refuse. The difference is worth stating, because
|
|
285
|
+
* the two classes fail in different places and only one of them is decidable
|
|
286
|
+
* here. A negative, `NaN` or `Infinity` renders into the argv as text the
|
|
287
|
+
* daemon either rejects or turns into a bound nobody asked for, and `0` is the
|
|
288
|
+
* opposite of a bound (`NanoCPUs` of zero is how a container says "no CPU
|
|
289
|
+
* limit"), so a host that wrote one of those hears about it during wiring
|
|
290
|
+
* rather than as a container that never came up.
|
|
291
|
+
*
|
|
292
|
+
* The upper bound is not ours to check. Moby's `verifyPlatformContainerResources`
|
|
293
|
+
* refuses `NanoCPUs` above the DAEMON host's CPU count (`"range of CPUs is from
|
|
294
|
+
* 0.01 to N.00, as there are only N CPUs available"`), and the same function
|
|
295
|
+
* deliberately sets no floor of its own on Linux, leaving that to the kernel.
|
|
296
|
+
* Neither number is knowable from here: the `docker` binary this backend drives
|
|
297
|
+
* can be pointed at a daemon on another machine (`DOCKER_HOST`), and even
|
|
298
|
+
* locally `os.cpus().length` is this machine's view rather than the daemon's
|
|
299
|
+
* own `runtime.NumCPU()`. Refusing on a guess at it would break a host whose
|
|
300
|
+
* daemon has more cores than the process driving it, which is a worse failure
|
|
301
|
+
* than the one it would catch — those arrive from the daemon with its own
|
|
302
|
+
* message, at spawn, where every other daemon-side refusal arrives too.
|
|
303
|
+
*/
|
|
304
|
+
export declare function assertCpuLimitIsRenderable(cpuLimit: number | undefined): void;
|
|
305
|
+
/**
|
|
306
|
+
* `--tmpfs` flags for the paths that stay writable under `--read-only`.
|
|
307
|
+
*
|
|
308
|
+
* See {@link DEFAULT_WRITABLE_ROOTFS_PATHS} for the paths themselves and why
|
|
309
|
+
* each is there. Returns nothing when the read-only root filesystem is off, and
|
|
310
|
+
* the two ways a host names paths that cannot be mounted — a contradiction with
|
|
311
|
+
* `readOnlyRootfs: false`, or a path the layout already mounts — are refusals
|
|
312
|
+
* rather than a silently shorter list.
|
|
313
|
+
*/
|
|
314
|
+
export declare function renderWritableRootfsArgs(config: DockerHardeningConfig): string[];
|
|
315
|
+
/**
|
|
316
|
+
* The confinement preamble for one container, in argv order.
|
|
317
|
+
*
|
|
318
|
+
* A function rather than a bare constant because `--read-only` is switchable
|
|
319
|
+
* and the flags that follow it describe what stays writable while it is on:
|
|
320
|
+
* `readOnlyRootfs: false` removes both the flag and the mounts. That is the only
|
|
321
|
+
* thing it removes. Everything in {@link HARDENING_ARGS} is applied
|
|
322
|
+
* unconditionally and no field can turn one of those off, so the argv for
|
|
323
|
+
* `readOnlyRootfs: false` is the argv this backend produced before any of this
|
|
324
|
+
* existed PLUS `--ipc private` — those two flags are the whole previous argv,
|
|
325
|
+
* and `--ipc private` is now unconditional. `--ipc` is not folded under this
|
|
326
|
+
* switch, because the field names the root filesystem: a host that turned the
|
|
327
|
+
* read-only rootfs off would be turning IPC isolation off as well, silently,
|
|
328
|
+
* for a reason the name of the field does not say. A switch has to mean one
|
|
329
|
+
* thing.
|
|
330
|
+
*/
|
|
331
|
+
export declare function renderHardeningArgs(config: DockerHardeningConfig): string[];
|
|
332
|
+
/**
|
|
333
|
+
* Everything {@link buildDockerRunArgs} renders, as a value.
|
|
334
|
+
*
|
|
335
|
+
* The pieces that come from the daemon or from the host are inputs rather than
|
|
336
|
+
* lookups: which network the container attaches to, and whether an egress proxy
|
|
337
|
+
* is listening and on which port. Both are already resolved by the caller, and
|
|
338
|
+
* reading them here would put a daemon call back inside the function whose
|
|
339
|
+
* whole point is that it needs none.
|
|
340
|
+
*/
|
|
341
|
+
export interface DockerRunArgvInput {
|
|
342
|
+
readonly config: DockerBackendInternalConfig;
|
|
343
|
+
readonly options: SandboxBackendOptions;
|
|
344
|
+
readonly containerName: string;
|
|
345
|
+
readonly network: string;
|
|
346
|
+
readonly hostReachability: 'host-port' | 'container-network';
|
|
347
|
+
/**
|
|
348
|
+
* Port the host-side egress proxy listens on, when one is running. Absent
|
|
349
|
+
* means no proxy, and no proxy environment is passed in — which is not the
|
|
350
|
+
* same fact as a proxy that was configured and is unreachable.
|
|
351
|
+
*/
|
|
352
|
+
readonly egressProxyPort?: number;
|
|
353
|
+
}
|
|
354
|
+
/**
|
|
355
|
+
* The complete `docker run` argv, as a value.
|
|
356
|
+
*
|
|
357
|
+
* Extracted for the same reason {@link resolveNetwork} and
|
|
358
|
+
* {@link egressProxyOptions} were: everything downstream of it needs a running
|
|
359
|
+
* Docker daemon, so a confinement flag that never reached the argv — or one
|
|
360
|
+
* that reached it in an order that cancels another — could only be caught by an
|
|
361
|
+
* operator noticing its effect missing in production. Spawning a fake `docker`
|
|
362
|
+
* and reading back what it was handed proves what the fake was told and nothing
|
|
363
|
+
* about the container the daemon would build. Here the whole baseline is one
|
|
364
|
+
* array, and an edit that drops a flag fails a test rather than a deployment.
|
|
365
|
+
*
|
|
366
|
+
* Order matters in exactly two places, and both are asserted by the test that
|
|
367
|
+
* pins this: the image is the last argument, because everything after it is a
|
|
368
|
+
* command for the container rather than a flag for docker; and every flag that
|
|
369
|
+
* takes a value is pushed as two argv entries rather than one string, so no
|
|
370
|
+
* value is ever re-split by anything downstream.
|
|
371
|
+
*/
|
|
372
|
+
export declare function buildDockerRunArgs(input: DockerRunArgvInput): string[];
|
|
210
373
|
/**
|
|
211
374
|
* Validate and resolve a {@link ContainerSandboxLayout}. Returns a
|
|
212
375
|
* {@link ResolvedContainerSandboxLayout} with every container path
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"file":"index.d.ts","sourceRoot":"","sources":["../../../src/backends/docker/index.ts"],"names":[],"mappings":"AAAA;;;;;;;;;;;;;;;;;;;;;GAqBG;AAIH,OAAO,EACN,KAAK,sBAAsB,EAE3B,KAAK,8BAA8B,EAmBnC,MAAM,YAAY,CAAA;AAEnB,OAAO,KAAK,EACX,kBAAkB,EAClB,kBAAkB,EAElB,MAAM,uBAAuB,CAAA;AAE9B,OAAO,EAEN,KAAK,YAAY,EACjB,KAAK,cAAc,
|
|
1
|
+
{"version":3,"file":"index.d.ts","sourceRoot":"","sources":["../../../src/backends/docker/index.ts"],"names":[],"mappings":"AAAA;;;;;;;;;;;;;;;;;;;;;GAqBG;AAIH,OAAO,EACN,KAAK,sBAAsB,EAE3B,KAAK,8BAA8B,EAmBnC,MAAM,YAAY,CAAA;AAEnB,OAAO,KAAK,EACX,kBAAkB,EAClB,kBAAkB,EAElB,MAAM,uBAAuB,CAAA;AAE9B,OAAO,EAEN,KAAK,YAAY,EACjB,KAAK,cAAc,EACnB,KAAK,qBAAqB,EAC1B,MAAM,gBAAgB,CAAA;AAcvB;;;;;;;;;;;GAWG;AACH,MAAM,WAAW,2BAA2B;IAC3C,QAAQ,CAAC,KAAK,EAAE,MAAM,CAAA;IACtB;;;;OAIG;IACH,QAAQ,CAAC,MAAM,EAAE,8BAA8B,CAAA;IAC/C,QAAQ,CAAC,YAAY,CAAC,EAAE,MAAM,CAAA;IAE9B;;;;;;;;;;;;OAYG;IACH,QAAQ,CAAC,SAAS,CAAC,EAAE,MAAM,CAAA;IAE3B;;;;;;;;;;;;;;;;;;;OAmBG;IACH,QAAQ,CAAC,QAAQ,CAAC,EAAE,MAAM,CAAA;IAE1B;;;;;;;;;;;;;;OAcG;IACH,QAAQ,CAAC,cAAc,CAAC,EAAE,OAAO,CAAA;IAEjC;;;;;;;;;;OAUG;IACH,QAAQ,CAAC,mBAAmB,CAAC,EAAE,SAAS,MAAM,EAAE,CAAA;IAEhD;;;;;;;;;OASG;IACH,QAAQ,CAAC,mBAAmB,CAAC,EAAE,SAAS,kBAAkB,EAAE,CAAA;IAE5D;;;;;;;;;;;;;OAaG;IACH,QAAQ,CAAC,cAAc,CAAC,EAAE,SAAS,MAAM,EAAE,CAAA;IAE3C,QAAQ,CAAC,OAAO,CAAC,EAAE,MAAM,GAAG,QAAQ,GAAG,MAAM,CAAA;IAC7C,QAAQ,CAAC,mBAAmB,CAAC,EAAE,MAAM,CAAA;IACrC,QAAQ,CAAC,cAAc,CAAC,EAAE,MAAM,CAAA;IAChC;;;;;;;;;;OAUG;IACH,QAAQ,CAAC,OAAO,CAAC,EAAE,MAAM,GAAG,OAAO,GAAG,MAAM,CAAA;IAC5C;;;;;;;;;;;;;;;;OAgBG;IACH,QAAQ,CAAC,gBAAgB,CAAC,EAAE,WAAW,GAAG,mBAAmB,CAAA;IAC7D;;;;;;;OAOG;IACH,QAAQ,CAAC,MAAM,CAAC,EAAE,QAAQ,CAAC,MAAM,CAAC,MAAM,EAAE,MAAM,CAAC,CAAC,CAAA;CAClD;AAOD;;;;GAIG;AACH,wBAAgB,kBAAkB,CAAC,MAAM,EAAE,2BAA2B,GAAG,cAAc,CAwBtF;AAED;;;;;;;;;;;;;;;;GAgBG;AACH,wBAAgB,cAAc,CAC7B,UAAU,EAAE,MAAM,EAClB,MAAM,EAAE,YAAY,GAAG,SAAS,EAChC,QAAQ,UAAQ,GACd,MAAM,CAkBR;AAED;;;;;;;GAOG;AACH,wBAAsB,mBAAmB,CAAC,MAAM,EAAE,YAAY,GAAG,OAAO,CAAC,SAAS,MAAM,EAAE,CAAC,CAM1F;AAED,0EAA0E;AAC1E,wBAAgB,gBAAgB,CAAC,MAAM,EAAE,YAAY,GAAG,SAAS,GAAG,OAAO,CAE1E;AAED;;;;;;;;GAQG;AACH,wBAAgB,iBAAiB,CAAC,qBAAqB,EAAE,MAAM,GAAG,OAAO,CAExE;AAED;;;;;;;;;;;;;;;;;;;;;;;;;;;GA2BG;AACH,wBAAgB,6BAA6B,CAC5C,OAAO,EAAE,MAAM,EACf,YAAY,EAAE,WAAW,GAAG,mBAAmB,EAC/C,MAAM,EAAE,YAAY,GAAG,SAAS,EAChC,qBAAqB,EAAE,MAAM,GAC3B,IAAI,CAcN;AAED;;;;;;;;GAQG;AACH,wBAAgB,kBAAkB,CACjC,MAAM,EAAE,IAAI,CAAC,2BAA2B,EAAE,qBAAqB,GAAG,gBAAgB,CAAC,EACnF,MAAM,EAAE,YAAY,GAClB,kBAAkB,CASpB;AA2ND,gEAAgE;AAChE,MAAM,MAAM,qBAAqB,GAAG,IAAI,CACvC,2BAA2B,EAC3B,UAAU,GAAG,QAAQ,GAAG,gBAAgB,GAAG,qBAAqB,CAChE,CAAA;AAqDD;;;;;;;;;;;;;GAaG;AACH,wBAAgB,8BAA8B,CAAC,MAAM,EAAE,qBAAqB,GAAG,IAAI,CAMlF;AAED;;;;;;;;;;;;;;;;;;;;;;;GAuBG;AACH,wBAAgB,0BAA0B,CAAC,QAAQ,EAAE,MAAM,GAAG,SAAS,GAAG,IAAI,CAO7E;AAED;;;;;;;;GAQG;AACH,wBAAgB,wBAAwB,CAAC,MAAM,EAAE,qBAAqB,GAAG,MAAM,EAAE,CA0ChF;AAED;;;;;;;;;;;;;;;GAeG;AACH,wBAAgB,mBAAmB,CAAC,MAAM,EAAE,qBAAqB,GAAG,MAAM,EAAE,CAM3E;AAED;;;;;;;;GAQG;AACH,MAAM,WAAW,kBAAkB;IAClC,QAAQ,CAAC,MAAM,EAAE,2BAA2B,CAAA;IAC5C,QAAQ,CAAC,OAAO,EAAE,qBAAqB,CAAA;IACvC,QAAQ,CAAC,aAAa,EAAE,MAAM,CAAA;IAC9B,QAAQ,CAAC,OAAO,EAAE,MAAM,CAAA;IACxB,QAAQ,CAAC,gBAAgB,EAAE,WAAW,GAAG,mBAAmB,CAAA;IAC5D;;;;OAIG;IACH,QAAQ,CAAC,eAAe,CAAC,EAAE,MAAM,CAAA;CACjC;AAED;;;;;;;;;;;;;;;;;GAiBG;AACH,wBAAgB,kBAAkB,CAAC,KAAK,EAAE,kBAAkB,GAAG,MAAM,EAAE,CA+GtE;AAskBD;;;;;;;;;;;;GAYG;AACH,wBAAgB,aAAa,CAAC,MAAM,EAAE,sBAAsB,GAAG,8BAA8B,CAmI5F;AAkCD,wBAAgB,qBAAqB,CAAC,MAAM,EAAE,8BAA8B,GAAG,MAAM,EAAE,CA+BtF;AAED,wBAAgB,wBAAwB,CAAC,MAAM,EAAE,8BAA8B,GAAG,MAAM,CAUvF;AAED;;;;;GAKG;AACH,wBAAgB,yBAAyB,CAAC,MAAM,EAAE,8BAA8B,GAAG,MAAM,CAKxF"}
|