runcloud 0.1.120 → 0.1.121

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/dist/version.js CHANGED
@@ -1 +1 @@
1
- export const CLI_VERSION = '0.1.120';
1
+ export const CLI_VERSION = '0.1.121';
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "runcloud",
3
- "version": "0.1.120",
3
+ "version": "0.1.121",
4
4
  "description": "Create and control run.cloud remote mobile simulators and cloud sandboxes",
5
5
  "license": "Apache-2.0",
6
6
  "keywords": [
@@ -295,14 +295,42 @@ Three limits that change what you can conclude from it:
295
295
  early output is gone, not merely paged out. An empty-looking start is a trim,
296
296
  not proof the process printed nothing.
297
297
  - `systemd-run` output goes to the journal, not the console, so those units never
298
- appear here at all. Read `journalctl -u <unit>` inside the sandbox instead.
298
+ appear here at all. `journalctl -u <unit>` reaches it, but only from inside a
299
+ running sandbox, so that output is unrecoverable once the sandbox is destroyed
300
+ and needs a resume on a paused one. A workload whose output has to survive a
301
+ post-mortem should write to stdout rather than a unit.
302
+
303
+ Two details about reading the file: the `created` marker is appended to a shell
304
+ prompt line rather than starting one, so match `--- sandbox` anywhere in the line
305
+ rather than anchoring to the start, and boot output dwarfs everything else, so
306
+ reach for the tail before the head.
307
+
308
+ ### What the metrics actually settle
309
+
310
+ `runcloud sandbox metrics <id> --json` carries typed counters, and they are how
311
+ you refute a theory instead of arguing about it. The natural wrong answer on a
312
+ 128 MiB sandbox is always "it OOMed":
313
+
314
+ | field | what a zero proves |
315
+ | --- | --- |
316
+ | `memory_oom_kills`, `memory_oom_events` | nothing was OOM-killed, so a dead process exited on its own |
317
+ | `cpu_throttled_periods`, `cpu_throttled_millicores` | the CPU reservation was not starving it |
318
+ | lifetime `network` totals | at ~0 bytes in, no request ever arrived, so the fault is upstream of the guest |
319
+
320
+ Each has a `*_valid` companion; when that is false the counter is unavailable
321
+ rather than zero, and a zero you cannot trust proves nothing.
322
+
323
+ `memory_bytes` is the exception and it misleads. It is measured host-side and
324
+ includes page cache, so it routinely reads several times the sandbox's `mem_mb`
325
+ reservation. That is not a leak and not an impending OOM. Believe
326
+ `memory_oom_kills` over it, every time.
299
327
 
300
328
  ### Lifecycle markers
301
329
 
302
330
  The host writes a marker into the console on every transition:
303
331
 
304
332
  ```
305
- --- sandbox destroyed: timeout-sweep at 2026-08-05T11:02:11Z ---
333
+ --- sandbox paused: timeout-sweep at 2026-08-05T11:02:11Z ---
306
334
  ```
307
335
 
308
336
  The events are `created`, `paused`, `resumed`, `stopped`, and `destroyed`. The
@@ -310,22 +338,52 @@ token after the colon is the control plane's actor, written verbatim so it can b
310
338
  matched rather than parsed; the sentence around it is not a contract. Only stops
311
339
  carry a reason, so `created` and `resumed` appear without one.
312
340
 
313
- | actor | what it means |
314
- | --- | --- |
315
- | `timeout-sweep` | hit its `--timeout` lifetime cap. Not a crash |
316
- | `idle-sweep` | idle-paused after `--idle-pause` seconds. Resumable, and not a failure |
317
- | `retention-sweep` | destroyed after its 48h parked window expired |
318
- | `pause` `resume` `stop` `archive` `restore` | a caller requested exactly this |
319
- | `api` `api-force` | a caller destroyed it. `api-force` bypassed a wedged guest |
320
- | `health-sweep` `host-reconcile` `reap-sweep` | the platform reclaimed it |
321
- | `boot-failure` `placement-timeout` `request-timeout` `scheduler` | it never started. Capacity or image, not the workload |
322
- | `image-built` `image-build-failure` | an async image build finished or failed |
323
- | `startup-recover` `volume-lease-invalid` `replica-fork-resume-failure` | platform recovery paths |
341
+ Read the **event** before the actor. A paused sandbox still exists and its disk is
342
+ intact; a destroyed one is gone. Telling someone their work is lost when it is
343
+ warm-parked and resumable is the most expensive mistake available here.
344
+
345
+ | actor | ends up | what it means |
346
+ | --- | --- | --- |
347
+ | `timeout-sweep` | paused | hit its `--timeout` lifetime cap. Not a crash, and the work is still there |
348
+ | `idle-sweep` | paused | idle-paused after `--idle-pause` seconds. Not a failure at all |
349
+ | `retention-sweep` | destroyed | the 48h parked window expired and it was reaped for real |
350
+ | `pause` `resume` `stop` `archive` `restore` | as asked | a caller requested exactly this |
351
+ | `api` `api-force` | destroyed | a caller destroyed it. `api-force` bypassed a wedged guest |
352
+ | `health-sweep` | interrupted | the platform judged it unhealthy |
353
+ | `host-reconcile` `reap-sweep` | destroyed | the platform reclaimed it |
354
+ | `boot-failure` `placement-timeout` `request-timeout` `scheduler` | never ran | capacity or image, not the workload |
355
+ | `image-built` `image-build-failure` | varies | an async image build finished or failed |
356
+ | `startup-recover` `volume-lease-invalid` `replica-fork-resume-failure` | varies | platform recovery paths |
357
+
358
+ **Getting a parked sandbox back.** `timeout-sweep` and `idle-sweep` leave a full
359
+ snapshot, so `runcloud sandbox resume <id>` restores it warm. You have 48h before
360
+ `retention-sweep` destroys it for real. Resuming restarts the lifetime window from
361
+ now rather than clearing it, so a sandbox parked by `timeout-sweep` will park again
362
+ after the same interval; raise the cap first with the SDK's
363
+ `cloud.sandboxes.setTimeout(id, seconds)`, which has no CLI equivalent.
364
+
365
+ **The `--timeout` trap.** `--timeout` is a wall-clock lifetime cap, wholly separate
366
+ from `--idle-pause`. It fires on a fully busy sandbox, and `--persistent` /
367
+ `--idle-pause 0` do **not** hold it off, despite `--persistent` reading as "never
368
+ pause when idle". Default is 300s, ceiling 24h.
324
369
 
325
370
  A sandbox that failed before it ever reached a host has no console, because the
326
371
  file is created at boot. Those carry a `Last error` row on
327
372
  `runcloud sandbox get` instead, which is then the only record of the cause.
328
373
 
374
+ ### "Not reachable" is usually not a network problem
375
+
376
+ Exposure is a separate record, not a field on the sandbox, so `sandbox get` omits
377
+ it entirely when none exists rather than showing it as empty. Absence of a
378
+ `Hostname` row therefore means **no public route exists**, not that the CLI
379
+ declined to print one. `runcloud sandbox domain list <id>` states it outright and
380
+ names the fix. A sandbox created without `--expose` is unreachable from outside
381
+ however healthy the workload is.
382
+
383
+ When a route does exist and the port still refuses, the guest is usually bound to
384
+ `127.0.0.1` rather than `0.0.0.0`, which is invisible from outside and looks
385
+ identical to a dead app.
386
+
329
387
  ## Guardrails
330
388
 
331
389
  - Destroy every sandbox created during a task unless the user explicitly asks