runcloud 0.1.119 → 0.1.121
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/dist/version.js
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
export const CLI_VERSION = '0.1.
|
|
1
|
+
export const CLI_VERSION = '0.1.121';
|
package/package.json
CHANGED
|
Binary file
|
|
@@ -286,20 +286,51 @@ printed, and lifecycle markers, interleaved in the order they happened.
|
|
|
286
286
|
the file rather than deleting it, so a post-mortem read works on a sandbox that
|
|
287
287
|
no longer exists.
|
|
288
288
|
|
|
289
|
-
|
|
289
|
+
Three limits that change what you can conclude from it:
|
|
290
290
|
|
|
291
|
+
- A destroyed sandbox's console gets removed after 14 days (2 weeks), so a
|
|
292
|
+
post-mortem has a deadline. A running or paused sandbox's log is never removed,
|
|
293
|
+
however long it has sat quiet.
|
|
291
294
|
- It is capped at 10 MiB and trimmed to the newest 5 MiB, so on a chatty workload
|
|
292
295
|
early output is gone, not merely paged out. An empty-looking start is a trim,
|
|
293
296
|
not proof the process printed nothing.
|
|
294
297
|
- `systemd-run` output goes to the journal, not the console, so those units never
|
|
295
|
-
appear here at all.
|
|
298
|
+
appear here at all. `journalctl -u <unit>` reaches it, but only from inside a
|
|
299
|
+
running sandbox, so that output is unrecoverable once the sandbox is destroyed
|
|
300
|
+
and needs a resume on a paused one. A workload whose output has to survive a
|
|
301
|
+
post-mortem should write to stdout rather than a unit.
|
|
302
|
+
|
|
303
|
+
Two details about reading the file: the `created` marker is appended to a shell
|
|
304
|
+
prompt line rather than starting one, so match `--- sandbox` anywhere in the line
|
|
305
|
+
rather than anchoring to the start, and boot output dwarfs everything else, so
|
|
306
|
+
reach for the tail before the head.
|
|
307
|
+
|
|
308
|
+
### What the metrics actually settle
|
|
309
|
+
|
|
310
|
+
`runcloud sandbox metrics <id> --json` carries typed counters, and they are how
|
|
311
|
+
you refute a theory instead of arguing about it. The natural wrong answer on a
|
|
312
|
+
128 MiB sandbox is always "it OOMed":
|
|
313
|
+
|
|
314
|
+
| field | what a zero proves |
|
|
315
|
+
| --- | --- |
|
|
316
|
+
| `memory_oom_kills`, `memory_oom_events` | nothing was OOM-killed, so a dead process exited on its own |
|
|
317
|
+
| `cpu_throttled_periods`, `cpu_throttled_millicores` | the CPU reservation was not starving it |
|
|
318
|
+
| lifetime `network` totals | at ~0 bytes in, no request ever arrived, so the fault is upstream of the guest |
|
|
319
|
+
|
|
320
|
+
Each has a `*_valid` companion; when that is false the counter is unavailable
|
|
321
|
+
rather than zero, and a zero you cannot trust proves nothing.
|
|
322
|
+
|
|
323
|
+
`memory_bytes` is the exception and it misleads. It is measured host-side and
|
|
324
|
+
includes page cache, so it routinely reads several times the sandbox's `mem_mb`
|
|
325
|
+
reservation. That is not a leak and not an impending OOM. Believe
|
|
326
|
+
`memory_oom_kills` over it, every time.
|
|
296
327
|
|
|
297
328
|
### Lifecycle markers
|
|
298
329
|
|
|
299
330
|
The host writes a marker into the console on every transition:
|
|
300
331
|
|
|
301
332
|
```
|
|
302
|
-
--- sandbox
|
|
333
|
+
--- sandbox paused: timeout-sweep at 2026-08-05T11:02:11Z ---
|
|
303
334
|
```
|
|
304
335
|
|
|
305
336
|
The events are `created`, `paused`, `resumed`, `stopped`, and `destroyed`. The
|
|
@@ -307,22 +338,52 @@ token after the colon is the control plane's actor, written verbatim so it can b
|
|
|
307
338
|
matched rather than parsed; the sentence around it is not a contract. Only stops
|
|
308
339
|
carry a reason, so `created` and `resumed` appear without one.
|
|
309
340
|
|
|
310
|
-
|
|
311
|
-
|
|
312
|
-
|
|
313
|
-
|
|
314
|
-
|
|
|
315
|
-
|
|
|
316
|
-
| `
|
|
317
|
-
| `
|
|
318
|
-
| `
|
|
319
|
-
| `
|
|
320
|
-
| `
|
|
341
|
+
Read the **event** before the actor. A paused sandbox still exists and its disk is
|
|
342
|
+
intact; a destroyed one is gone. Telling someone their work is lost when it is
|
|
343
|
+
warm-parked and resumable is the most expensive mistake available here.
|
|
344
|
+
|
|
345
|
+
| actor | ends up | what it means |
|
|
346
|
+
| --- | --- | --- |
|
|
347
|
+
| `timeout-sweep` | paused | hit its `--timeout` lifetime cap. Not a crash, and the work is still there |
|
|
348
|
+
| `idle-sweep` | paused | idle-paused after `--idle-pause` seconds. Not a failure at all |
|
|
349
|
+
| `retention-sweep` | destroyed | the 48h parked window expired and it was reaped for real |
|
|
350
|
+
| `pause` `resume` `stop` `archive` `restore` | as asked | a caller requested exactly this |
|
|
351
|
+
| `api` `api-force` | destroyed | a caller destroyed it. `api-force` bypassed a wedged guest |
|
|
352
|
+
| `health-sweep` | interrupted | the platform judged it unhealthy |
|
|
353
|
+
| `host-reconcile` `reap-sweep` | destroyed | the platform reclaimed it |
|
|
354
|
+
| `boot-failure` `placement-timeout` `request-timeout` `scheduler` | never ran | capacity or image, not the workload |
|
|
355
|
+
| `image-built` `image-build-failure` | varies | an async image build finished or failed |
|
|
356
|
+
| `startup-recover` `volume-lease-invalid` `replica-fork-resume-failure` | varies | platform recovery paths |
|
|
357
|
+
|
|
358
|
+
**Getting a parked sandbox back.** `timeout-sweep` and `idle-sweep` leave a full
|
|
359
|
+
snapshot, so `runcloud sandbox resume <id>` restores it warm. You have 48h before
|
|
360
|
+
`retention-sweep` destroys it for real. Resuming restarts the lifetime window from
|
|
361
|
+
now rather than clearing it, so a sandbox parked by `timeout-sweep` will park again
|
|
362
|
+
after the same interval; raise the cap first with the SDK's
|
|
363
|
+
`cloud.sandboxes.setTimeout(id, seconds)`, which has no CLI equivalent.
|
|
364
|
+
|
|
365
|
+
**The `--timeout` trap.** `--timeout` is a wall-clock lifetime cap, wholly separate
|
|
366
|
+
from `--idle-pause`. It fires on a fully busy sandbox, and `--persistent` /
|
|
367
|
+
`--idle-pause 0` do **not** hold it off, despite `--persistent` reading as "never
|
|
368
|
+
pause when idle". Default is 300s, ceiling 24h.
|
|
321
369
|
|
|
322
370
|
A sandbox that failed before it ever reached a host has no console, because the
|
|
323
371
|
file is created at boot. Those carry a `Last error` row on
|
|
324
372
|
`runcloud sandbox get` instead, which is then the only record of the cause.
|
|
325
373
|
|
|
374
|
+
### "Not reachable" is usually not a network problem
|
|
375
|
+
|
|
376
|
+
Exposure is a separate record, not a field on the sandbox, so `sandbox get` omits
|
|
377
|
+
it entirely when none exists rather than showing it as empty. Absence of a
|
|
378
|
+
`Hostname` row therefore means **no public route exists**, not that the CLI
|
|
379
|
+
declined to print one. `runcloud sandbox domain list <id>` states it outright and
|
|
380
|
+
names the fix. A sandbox created without `--expose` is unreachable from outside
|
|
381
|
+
however healthy the workload is.
|
|
382
|
+
|
|
383
|
+
When a route does exist and the port still refuses, the guest is usually bound to
|
|
384
|
+
`127.0.0.1` rather than `0.0.0.0`, which is invisible from outside and looks
|
|
385
|
+
identical to a dead app.
|
|
386
|
+
|
|
326
387
|
## Guardrails
|
|
327
388
|
|
|
328
389
|
- Destroy every sandbox created during a task unless the user explicitly asks
|