@specific.dev/spectest 0.70.0 → 0.71.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -142,6 +142,119 @@ export interface K3sOptions {
142
142
  /** Port the in-cluster registry listens on (plain HTTP). */
143
143
  const K3S_REGISTRY_PORT = 5000;
144
144
 
145
+ /** The cluster's own containerd root, inside the k3s container. */
146
+ const CLUSTER_STORE_TARGET = "/var/lib/rancher/k3s/agent/containerd";
147
+
148
+
149
+ /**
150
+ * Written before the server starts and removed by the ready probe.
151
+ *
152
+ * It survives in the store only when a boot never reached a live API
153
+ * server, which is the one thing the sweep below cannot work out for
154
+ * itself. Blacksmith's rule, from the other end: they refuse to commit a
155
+ * cache when the integrity check *did not run*, not only when it failed.
156
+ */
157
+ const BOOT_MARKER = ".spectest-boot-incomplete";
158
+ /**
159
+ * Bring the cluster's image store back into service, or throw it away.
160
+ *
161
+ * The store is the cluster's containerd root on a directory that outlives
162
+ * this container (`NESTED_STORE_ROOT`), so a build finds coredns, traefik,
163
+ * the local-path provisioner and every image the project deployed already
164
+ * pulled **and already extracted**. Extraction is 48x the I/O of the
165
+ * download (`STICKY_DISKS.md` §7.2), which is why the zot mirror never
166
+ * helped here and why an extracted store is the whole prize.
167
+ *
168
+ * What it inherits with them is one build's crash state: the store was
169
+ * captured from a running guest, and the teardown's `docker rm -f` killed
170
+ * that cluster where it stood. **Deleting the dead files is not enough,
171
+ * and the failure is not subtle.** containerd's index still lists the
172
+ * previous build's containers, so the CRI plugin offers them to a kubelet
173
+ * whose API server has never heard of them; the kubelet spends its startup
174
+ * reconciling containers whose tasks are gone (`Failed to create existing
175
+ * container … task not found`) and does not register its own Node in time,
176
+ * and the CSI plugin — which waits on that Node with a fixed budget —
177
+ * calls `Fatalf`. The kubelet dies, k3s exits 2, and what the author sees
178
+ * is a rollout that timed out against a hostname that no longer resolves.
179
+ * Measured: 2 of 2 delta restores, against 2 of 2 clean without the store.
180
+ *
181
+ * So the sweep removes those containers through **containerd itself**,
182
+ * which is the only thing that can edit its index. The k3s image ships a
183
+ * standalone `containerd` and `ctr`, so a throwaway daemon runs against
184
+ * the store with the CRI plugin disabled, deletes the previous build's
185
+ * containers and their tasks, and stops. Images, content blobs and
186
+ * extracted snapshots — the whole cache — stay. It is the same discipline
187
+ * `spectest-image-cache-up` applies to the guest's own store one level up,
188
+ * where it sweeps the `moby` namespace before dockerd starts.
189
+ *
190
+ * Two failures cost a cache and never a build:
191
+ *
192
+ * 1. **A sweep that cannot run wipes the store.** A daemon that will not
193
+ * start leaves behind exactly the containers this cluster fatals on,
194
+ * so carrying on is the one option that is known bad.
195
+ * 2. **A store that did not work is never inherited twice.** The marker
196
+ * is written before the server starts and removed by the ready probe,
197
+ * so it survives only when a boot never reached a live API server.
198
+ * That is the backstop for the sweep being incomplete in some way we
199
+ * have not seen: the cost is one cold build's pulls, and it repairs
200
+ * itself. Blacksmith's rule from the other end — they refuse to commit
201
+ * a cache when the integrity check *did not run*, not only when it
202
+ * failed.
203
+ */
204
+ /** Plugins the sweep daemon does not need. Every plugin is startup time,
205
+ * and the sweep sits on the critical path of the cluster's boot. CRI is
206
+ * off for a second reason: this daemon exists to edit an index, and a CRI
207
+ * plugin would set about being a container runtime on a store we are
208
+ * seconds from handing to the real one. */
209
+ const SWEEP_DISABLED_PLUGINS = [
210
+ "io.containerd.grpc.v1.cri",
211
+ "io.containerd.snapshotter.v1.btrfs",
212
+ "io.containerd.snapshotter.v1.native",
213
+ "io.containerd.snapshotter.v1.aufs",
214
+ "io.containerd.snapshotter.v1.zfs",
215
+ "io.containerd.snapshotter.v1.devmapper",
216
+ "io.containerd.snapshotter.v1.stargz",
217
+ "io.containerd.snapshotter.v1.fuse-overlayfs",
218
+ ];
219
+
220
+ const STORE_SWEEP_SH =
221
+ `S=${CLUSTER_STORE_TARGET}; M="$S/${BOOT_MARKER}"; A=/run/spectest-sweep.sock; ` +
222
+ `sweep_store() { ` +
223
+ `printf '%s\\n' 'version = 2' ` +
224
+ `'disabled_plugins = [${SWEEP_DISABLED_PLUGINS.map((p) => `"${p}"`).join(", ")}]' ` +
225
+ `> /run/spectest-sweep.toml || return 1; ` +
226
+ `containerd -c /run/spectest-sweep.toml --root "$S" --state /run/spectest-sweep ` +
227
+ `--address "$A" > /run/spectest-sweep.log 2>&1 & ` +
228
+ `cd_pid=$!; i=0; ` +
229
+ `while [ ! -S "$A" ] && [ $i -lt 600 ]; do sleep 0.1; i=$((i+1)); done; ` +
230
+ `if [ ! -S "$A" ]; then kill $cd_pid 2>/dev/null; return 1; fi; ` +
231
+ // One invocation for the whole set, not two per container. Each `ctr` is
232
+ // a Go binary start plus a gRPC round trip, and a build leaves ~20
233
+ // containers behind — the per-container form spent seconds of the
234
+ // cluster's boot on process starts alone (measured: ready 3.6s -> 7.0s).
235
+ `ids=$(timeout 30 ctr -a "$A" -n k8s.io containers ls -q 2>/dev/null); ` +
236
+ `if [ -n "$ids" ]; then ` +
237
+ `timeout 60 ctr -a "$A" -n k8s.io tasks rm -f $ids >/dev/null 2>&1 || true; ` +
238
+ `timeout 60 ctr -a "$A" -n k8s.io containers rm $ids >/dev/null 2>&1 || true; fi; ` +
239
+ // containerd 2.x keeps sandboxes in a store of their own; on the 1.7 k3s
240
+ // ships they are ordinary containers and this is a no-op.
241
+ `sb=$(timeout 30 ctr -a "$A" -n k8s.io sandboxes ls -q 2>/dev/null); ` +
242
+ `[ -n "$sb" ] && timeout 60 ctr -a "$A" -n k8s.io sandboxes rm $sb >/dev/null 2>&1; ` +
243
+ `kill $cd_pid 2>/dev/null; wait $cd_pid 2>/dev/null; ` +
244
+ `rm -rf "$S/io.containerd.runtime.v2.task" "$S/tmpmounts" /run/spectest-sweep "$A" 2>/dev/null || true; ` +
245
+ `return 0; }; ` +
246
+ `if [ -e "$M" ]; then ` +
247
+ `echo 'spectest: the previous cluster on this image store never became ready; starting from an empty one' >&2; ` +
248
+ `rm -rf "$S"/* "$S"/.[!.]* 2>/dev/null || true; ` +
249
+ `elif [ -d "$S/io.containerd.metadata.v1.bolt" ]; then ` +
250
+ // Timed and printed always: the sweep is inside the cluster's ready time,
251
+ // so it is the first thing to suspect if that time moves.
252
+ `T0=$(date +%s); ` +
253
+ `if sweep_store; then echo "spectest: reused the cluster image store at $S (swept in $(($(date +%s)-T0))s)"; ` +
254
+ `else echo 'spectest: could not sweep the inherited image store; starting from an empty one' >&2; ` +
255
+ `rm -rf "$S"/* "$S"/.[!.]* 2>/dev/null || true; fi; fi; ` +
256
+ `mkdir -p "$S" && : > "$M" 2>/dev/null || true; `;
257
+
145
258
  /**
146
259
  * `extraArgs` entries the component overrides anyway. Passing one of
147
260
  * these looks like it works — the flag really is appended to the k3s
@@ -1013,53 +1126,22 @@ ${ports}${volumeMounts}${volumes}
1013
1126
  }
1014
1127
 
1015
1128
  /**
1016
- * Host-side `zot` pull-through cache layout (local Firecracker provider
1017
- * only). One zot instance per upstream registry, all bound to the
1018
- * `spectest-br0` gateway `10.42.0.1` on the ports below — **kept in sync
1019
- * with `scripts/install-zot.sh`**. We mirror the cluster's containerd
1020
- * through these so every image pull reuses the shared host cache instead
1021
- * of hitting the public registry, and we list the canonical upstream as
1022
- * a fallback endpoint so a missing/cold mirror only ever slows a pull,
1023
- * never breaks it.
1129
+ * Docker Hub through Google's public mirror, with the canonical upstream
1130
+ * as the fallback endpoint so a cold or unavailable mirror only ever
1131
+ * slows a pull, never breaks it. The same route the guest's own dockerd
1132
+ * takes (golden's daemon.json) and its in-VM BuildKit. Every other
1133
+ * registry is pulled direct. Nothing here points at a host-side service:
1134
+ * the cluster's own containerd store lives in the environment and is
1135
+ * captured with it (CONTAINER_STORE.md).
1024
1136
  */
1025
- const ZOT_MIRRORS: Array<{ registry: string; port: number; upstream: string }> = [
1026
- { registry: "docker.io", port: 5000, upstream: "https://registry-1.docker.io" },
1027
- { registry: "ghcr.io", port: 5001, upstream: "https://ghcr.io" },
1028
- { registry: "quay.io", port: 5002, upstream: "https://quay.io" },
1029
- { registry: "registry.k8s.io", port: 5003, upstream: "https://registry.k8s.io" },
1030
- { registry: "public.ecr.aws", port: 5004, upstream: "https://public.ecr.aws" },
1031
- { registry: "gcr.io", port: 5005, upstream: "https://gcr.io" },
1032
- { registry: "mcr.microsoft.com", port: 5006, upstream: "https://mcr.microsoft.com" },
1033
- ];
1034
-
1035
- /**
1036
- * Discover the host-side image cache gateway by reading the same
1037
- * `registry-mirrors` entry the in-VM dockerd already uses (baked into
1038
- * the golden rootfs's `/etc/docker/daemon.json`). Returns the
1039
- * gateway host (`"10.42.0.1"`) when present, or `null` when there's no
1040
- * host cache, in which case the cluster pulls every image direct. Runs inside the daemon (VM) at `index.ts` load time, so
1041
- * the result is stable per host and never poisons the warm-template
1042
- * cache.
1043
- */
1044
- function detectHostMirrorGateway(): string | null {
1045
- try {
1046
- const cfg = JSON.parse(
1047
- readFileSync("/etc/docker/daemon.json", "utf8"),
1048
- ) as { "registry-mirrors"?: string[] };
1049
- const first = cfg["registry-mirrors"]?.[0];
1050
- return first ? new URL(first).hostname || null : null;
1051
- } catch {
1052
- return null;
1053
- }
1054
- }
1137
+ const HUB_MIRROR = { registry: "docker.io", mirror: "https://mirror.gcr.io", upstream: "https://registry-1.docker.io" };
1055
1138
 
1056
1139
  /**
1057
1140
  * Build `/etc/rancher/k3s/registries.yaml`. k3s reads this **once, at
1058
1141
  * startup**, to configure its embedded containerd — which is why it has
1059
1142
  * to be seeded via `files` (a pre-start bind mount) rather than a
1060
1143
  * `setup` hook. Two jobs:
1061
- * 1. Mirror the cluster's image pulls through the host `zot` cache
1062
- * (omitted when there's no host cache).
1144
+ * 1. Mirror the cluster's Docker Hub pulls through `mirror.gcr.io`.
1063
1145
  * 2. Trust the in-cluster registry, addressed as `<key>.internal:5000`
1064
1146
  * (the `{{SPECTEST_SERVICE}}` token is expanded to the cluster's
1065
1147
  * service key when the file is written). Image *references* use
@@ -1067,24 +1149,15 @@ function detectHostMirrorGateway(): string | null {
1067
1149
  * endpoint `http://127.0.0.1:5000` — the hostNetwork registry pod
1068
1150
  * shares the node's netns, so this needs no in-container DNS and
1069
1151
  * can't be broken by a clobbered `hostnames`.
1070
- * Returns `null` when there's nothing to configure (no host cache and
1071
- * `registry` disabled), in which case no file is injected.
1072
1152
  */
1073
1153
  function buildRegistriesYaml(registryEnabled: boolean): string | null {
1074
- const gateway = detectHostMirrorGateway();
1075
- if (!gateway && !registryEnabled) return null;
1076
-
1077
- const lines: string[] = ["mirrors:"];
1078
- if (gateway) {
1079
- for (const { registry, port, upstream } of ZOT_MIRRORS) {
1080
- lines.push(
1081
- ` "${registry}":`,
1082
- ` endpoint:`,
1083
- ` - "http://${gateway}:${port}"`,
1084
- ` - "${upstream}"`,
1085
- );
1086
- }
1087
- }
1154
+ const lines: string[] = [
1155
+ "mirrors:",
1156
+ ` "${HUB_MIRROR.registry}":`,
1157
+ ` endpoint:`,
1158
+ ` - "${HUB_MIRROR.mirror}"`,
1159
+ ` - "${HUB_MIRROR.upstream}"`,
1160
+ ];
1088
1161
  if (registryEnabled) {
1089
1162
  const host = `{{SPECTEST_SERVICE}}.internal:${K3S_REGISTRY_PORT}`;
1090
1163
  lines.push(
@@ -1913,12 +1986,18 @@ export function k3s(opts: K3sOptions = {}) {
1913
1986
  // is affected.
1914
1987
  "mount --make-rshared / 2>/dev/null || " +
1915
1988
  "echo 'spectest: could not make / rshared; CSI node plugins may fail to publish volumes' >&2; " +
1989
+ STORE_SWEEP_SH +
1916
1990
  `exec ${serverArgs}`;
1917
1991
  // Plain /readyz probe. On a warm zot cache the cluster's images are
1918
1992
  // already local, so the first boot completes in seconds; the
1919
1993
  // first-ever boot on a cold-cache host pulls through the mirror and
1920
1994
  // can take a couple of minutes (covered by readyTimeoutSecs).
1921
- const readyCmd = "kubectl get --raw=/readyz >/dev/null 2>&1";
1995
+ //
1996
+ // Clearing the boot marker is part of the probe on purpose: "the API
1997
+ // server answers" is exactly the condition that proves this cluster came
1998
+ // up on the store it was given, and it is the only signal the sweep can
1999
+ // read on the next boot. See STORE_SWEEP_SH.
2000
+ const readyCmd = `kubectl get --raw=/readyz >/dev/null 2>&1 && rm -f ${CLUSTER_STORE_TARGET}/${BOOT_MARKER}`;
1922
2001
  const def = {
1923
2002
  image: { type: "registry", reference: `rancher/k3s:${version}` },
1924
2003
  command: cmd,
@@ -1934,15 +2013,30 @@ export function k3s(opts: K3sOptions = {}) {
1934
2013
  // the same netns and are reachable the same way, without appearing
1935
2014
  // in this list (it's documentation, not a firewall).
1936
2015
  ports: registryEnabled ? [80, 443, 6443, K3S_REGISTRY_PORT] : [80, 443, 6443],
1937
- // NOTE: do NOT mount /var/lib/rancher/k3s/agent/containerd as a cache
1938
- // volume. It was tried (to spare a recreated cluster re-pulling its
1939
- // system images on delta restores) and a fresh k3s server against the
1940
- // previous container's containerd store — killed un-cleanly by the
1941
- // teardown's `docker rm -f` — wedged the apiserver minutes in
1942
- // (rollouts never settled, pod listing started failing). The zot
1943
- // mirror already makes those re-pulls cheap; the residual win wasn't
1944
- // worth the recovery semantics of a crash-state store under a fresh
1945
- // cluster db.
2016
+ // The cluster's image store, on a volume that outlives this
2017
+ // container. Without it the store is the container's writable layer,
2018
+ // which `docker rm -f` deletes at teardown — so every build re-pulls
2019
+ // and, far more expensively, re-extracts coredns, traefik, the
2020
+ // local-path provisioner and everything the project deploys.
2021
+ // Measured on the one project on the image cache: 20.6 s per
2022
+ // build, identical on the cold and delta tiers, i.e. the one term
2023
+ // neither the delta restore nor the cache image store reached.
2024
+ //
2025
+ // `cache` names a cache disk, which is the ordinary way any project
2026
+ // asks for one — there is nothing Kubernetes-shaped in the platform
2027
+ // for this. Two clusters in one environment may name the same disk
2028
+ // and still get separate containerd roots, because a volume is rooted
2029
+ // per service on whatever disk it names; sharing a root would corrupt
2030
+ // it, since two containerd daemons cannot share a store (bolt takes
2031
+ // an exclusive lock).
2032
+ //
2033
+ // This reverses a NOTE that stood here for a year: a fresh k3s server
2034
+ // over a store killed un-cleanly by `docker rm -f` wedged the
2035
+ // apiserver minutes in, and nothing recovered. What was missing was
2036
+ // the discipline `base.rs::IMAGE_CACHE_UP_SH` applies to the guest's own
2037
+ // store — sweep the crash state, and never inherit a store that did
2038
+ // not work. Both are in STORE_SWEEP_SH.
2039
+ volumes: [{ target: CLUSTER_STORE_TARGET }],
1946
2040
  ...(registriesYaml
1947
2041
  ? {
1948
2042
  files: [