runcloud 0.1.119 → 0.1.121

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/dist/version.js CHANGED
@@ -1 +1 @@
1
- export const CLI_VERSION = '0.1.119';
1
+ export const CLI_VERSION = '0.1.121';
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "runcloud",
3
- "version": "0.1.119",
3
+ "version": "0.1.121",
4
4
  "description": "Create and control run.cloud remote mobile simulators and cloud sandboxes",
5
5
  "license": "Apache-2.0",
6
6
  "keywords": [
@@ -286,20 +286,51 @@ printed, and lifecycle markers, interleaved in the order they happened.
286
286
  the file rather than deleting it, so a post-mortem read works on a sandbox that
287
287
  no longer exists.
288
288
 
289
- Two limits that change what you can conclude from it:
289
+ Three limits that change what you can conclude from it:
290
290
 
291
+ - A destroyed sandbox's console gets removed after 14 days (2 weeks), so a
292
+ post-mortem has a deadline. A running or paused sandbox's log is never removed,
293
+ however long it has sat quiet.
291
294
  - It is capped at 10 MiB and trimmed to the newest 5 MiB, so on a chatty workload
292
295
  early output is gone, not merely paged out. An empty-looking start is a trim,
293
296
  not proof the process printed nothing.
294
297
  - `systemd-run` output goes to the journal, not the console, so those units never
295
- appear here at all. Read `journalctl -u <unit>` inside the sandbox instead.
298
+ appear here at all. `journalctl -u <unit>` reaches it, but only from inside a
299
+ running sandbox, so that output is unrecoverable once the sandbox is destroyed
300
+ and needs a resume on a paused one. A workload whose output has to survive a
301
+ post-mortem should write to stdout rather than a unit.
302
+
303
+ Two details about reading the file: the `created` marker is appended to a shell
304
+ prompt line rather than starting one, so match `--- sandbox` anywhere in the line
305
+ rather than anchoring to the start, and boot output dwarfs everything else, so
306
+ reach for the tail before the head.
307
+
308
+ ### What the metrics actually settle
309
+
310
+ `runcloud sandbox metrics <id> --json` carries typed counters, and they are how
311
+ you refute a theory instead of arguing about it. The natural wrong answer on a
312
+ 128 MiB sandbox is always "it OOMed":
313
+
314
+ | field | what a zero proves |
315
+ | --- | --- |
316
+ | `memory_oom_kills`, `memory_oom_events` | nothing was OOM-killed, so a dead process exited on its own |
317
+ | `cpu_throttled_periods`, `cpu_throttled_millicores` | the CPU reservation was not starving it |
318
+ | lifetime `network` totals | at ~0 bytes in, no request ever arrived, so the fault is upstream of the guest |
319
+
320
+ Each has a `*_valid` companion; when that is false the counter is unavailable
321
+ rather than zero, and a zero you cannot trust proves nothing.
322
+
323
+ `memory_bytes` is the exception and it misleads. It is measured host-side and
324
+ includes page cache, so it routinely reads several times the sandbox's `mem_mb`
325
+ reservation. That is not a leak and not an impending OOM. Believe
326
+ `memory_oom_kills` over it, every time.
296
327
 
297
328
  ### Lifecycle markers
298
329
 
299
330
  The host writes a marker into the console on every transition:
300
331
 
301
332
  ```
302
- --- sandbox destroyed: timeout-sweep at 2026-08-05T11:02:11Z ---
333
+ --- sandbox paused: timeout-sweep at 2026-08-05T11:02:11Z ---
303
334
  ```
304
335
 
305
336
  The events are `created`, `paused`, `resumed`, `stopped`, and `destroyed`. The
@@ -307,22 +338,52 @@ token after the colon is the control plane's actor, written verbatim so it can b
307
338
  matched rather than parsed; the sentence around it is not a contract. Only stops
308
339
  carry a reason, so `created` and `resumed` appear without one.
309
340
 
310
- | actor | what it means |
311
- | --- | --- |
312
- | `timeout-sweep` | hit its `--timeout` lifetime cap. Not a crash |
313
- | `idle-sweep` | idle-paused after `--idle-pause` seconds. Resumable, and not a failure |
314
- | `retention-sweep` | destroyed after its 48h parked window expired |
315
- | `pause` `resume` `stop` `archive` `restore` | a caller requested exactly this |
316
- | `api` `api-force` | a caller destroyed it. `api-force` bypassed a wedged guest |
317
- | `health-sweep` `host-reconcile` `reap-sweep` | the platform reclaimed it |
318
- | `boot-failure` `placement-timeout` `request-timeout` `scheduler` | it never started. Capacity or image, not the workload |
319
- | `image-built` `image-build-failure` | an async image build finished or failed |
320
- | `startup-recover` `volume-lease-invalid` `replica-fork-resume-failure` | platform recovery paths |
341
+ Read the **event** before the actor. A paused sandbox still exists and its disk is
342
+ intact; a destroyed one is gone. Telling someone their work is lost when it is
343
+ warm-parked and resumable is the most expensive mistake available here.
344
+
345
+ | actor | ends up | what it means |
346
+ | --- | --- | --- |
347
+ | `timeout-sweep` | paused | hit its `--timeout` lifetime cap. Not a crash, and the work is still there |
348
+ | `idle-sweep` | paused | idle-paused after `--idle-pause` seconds. Not a failure at all |
349
+ | `retention-sweep` | destroyed | the 48h parked window expired and it was reaped for real |
350
+ | `pause` `resume` `stop` `archive` `restore` | as asked | a caller requested exactly this |
351
+ | `api` `api-force` | destroyed | a caller destroyed it. `api-force` bypassed a wedged guest |
352
+ | `health-sweep` | interrupted | the platform judged it unhealthy |
353
+ | `host-reconcile` `reap-sweep` | destroyed | the platform reclaimed it |
354
+ | `boot-failure` `placement-timeout` `request-timeout` `scheduler` | never ran | capacity or image, not the workload |
355
+ | `image-built` `image-build-failure` | varies | an async image build finished or failed |
356
+ | `startup-recover` `volume-lease-invalid` `replica-fork-resume-failure` | varies | platform recovery paths |
357
+
358
+ **Getting a parked sandbox back.** `timeout-sweep` and `idle-sweep` leave a full
359
+ snapshot, so `runcloud sandbox resume <id>` restores it warm. You have 48h before
360
+ `retention-sweep` destroys it for real. Resuming restarts the lifetime window from
361
+ now rather than clearing it, so a sandbox parked by `timeout-sweep` will park again
362
+ after the same interval; raise the cap first with the SDK's
363
+ `cloud.sandboxes.setTimeout(id, seconds)`, which has no CLI equivalent.
364
+
365
+ **The `--timeout` trap.** `--timeout` is a wall-clock lifetime cap, wholly separate
366
+ from `--idle-pause`. It fires on a fully busy sandbox, and `--persistent` /
367
+ `--idle-pause 0` do **not** hold it off, despite `--persistent` reading as "never
368
+ pause when idle". Default is 300s, ceiling 24h.
321
369
 
322
370
  A sandbox that failed before it ever reached a host has no console, because the
323
371
  file is created at boot. Those carry a `Last error` row on
324
372
  `runcloud sandbox get` instead, which is then the only record of the cause.
325
373
 
374
+ ### "Not reachable" is usually not a network problem
375
+
376
+ Exposure is a separate record, not a field on the sandbox, so `sandbox get` omits
377
+ it entirely when none exists rather than showing it as empty. Absence of a
378
+ `Hostname` row therefore means **no public route exists**, not that the CLI
379
+ declined to print one. `runcloud sandbox domain list <id>` states it outright and
380
+ names the fix. A sandbox created without `--expose` is unreachable from outside
381
+ however healthy the workload is.
382
+
383
+ When a route does exist and the port still refuses, the guest is usually bound to
384
+ `127.0.0.1` rather than `0.0.0.0`, which is invisible from outside and looks
385
+ identical to a dead app.
386
+
326
387
  ## Guardrails
327
388
 
328
389
  - Destroy every sandbox created during a task unless the user explicitly asks