@shardflux/sdk 0.10.2 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -3,7 +3,30 @@
3
3
  Every API the README shows is available from the version named here. Below 1.0, a minor release may break
4
4
  compatibility; breaking changes are marked **Breaking**.
5
5
 
6
- ## 0.10.2 (not yet published)
6
+ ## 0.11.0 (not yet published)
7
+
8
+ ### A resume that restarted processes says so (cold boot, contracts §12)
9
+
10
+ Additive. When the platform's VM runtime changed after a suspend and no host can restore the memory snapshot, the
11
+ cell resumes the workspace by booting its checkpoint's disk (a cold boot), automatically, also when a tool call wakes
12
+ the workspace. The workspace runs and its files are as of the suspend, but every process was restarted. The API passes
13
+ the resume's result through; this release reads it.
14
+
15
+ - `ServerTiming.memoryRestored` (`boolean | null`): `result.memory_restored`. `true` when memory and running
16
+ processes came back; `false` for `cold_boot` (and `reset_blank_layer`); `null` when the result does not say (an API
17
+ from before the field), never read as `false`.
18
+ - `ServerTiming.coldBootReason` (`string | null`): `result.cold_boot_reason` with `resumePath` `cold_boot`, today
19
+ `runtime_changed`.
20
+ - `ServerTiming.resumePath` documents `cold_boot` and `reset_blank_layer`.
21
+ - `formatTiming()` prints `resume from cold_boot: processes restarted (runtime_changed)` when memory was not restored;
22
+ other resumes print as before.
23
+ - Where it shows: `workspace.lastTiming` and the `done` progress event after `resume({ wait: true })`, `wake()` and a
24
+ tool call that woke the workspace; the finished resume operation's `result` (`memory_restored`,
25
+ `cold_boot_reason`, and `cold_boot` with operator details such as `files_as_of`, whose shape may change).
26
+ - The new `ServerTiming` fields are optional in the type (always set by the SDK), so timings built by hand still
27
+ type-check. No return type or behaviour changes: a wake that cold-booted still resolves `true`.
28
+
29
+ ## 0.10.2
7
30
 
8
31
  Fix: the SDK reports 0.10.2. 0.10.1 was published reporting 0.10.0, its previous version; the build now
9
32
  fails when `SDK_VERSION` differs from package.json.
package/README.md CHANGED
@@ -10,7 +10,7 @@ tool calls into the workspace ([tool-call capture](#tool-call-capture-070)).
10
10
  > **Early access.** Shardflux is in early access. The API is versioned (`/v1`), but this SDK is
11
11
  > below 1.0: a minor release may contain breaking changes (see [Compatibility](#compatibility)).
12
12
 
13
- > **Versions.** This README describes 0.10.2. Anything marked **(0.10.0+)** is not in 0.9.0, **(0.9.0+)** not in 0.8.x, **(0.8.0+)** not in 0.7.x,
13
+ > **Versions.** This README describes 0.11.0. Anything marked **(0.11.0+)** is not in 0.10.x, **(0.10.0+)** not in 0.9.0, **(0.9.0+)** not in 0.8.x, **(0.8.0+)** not in 0.7.x,
14
14
  > **(0.7.0+)** not in 0.6.x and **(0.6.0+)** not in 0.5.0; [CHANGELOG.md](./CHANGELOG.md) lists what each version added. Check yours with
15
15
  > `npm ls @shardflux/sdk` or the exported `SDK_VERSION`.
16
16
 
@@ -64,7 +64,8 @@ console.log(formatTiming(again.lastTiming!)); // (0.6.0+) where the r
64
64
  ```
65
65
 
66
66
  `open()` waits until the workspace is running. Opening the same key again never resets it: files,
67
- installed packages and running processes are still there. The same program is in
67
+ installed packages and running processes are still there (a resume restarts processes only in the rare cold boot,
68
+ see "A resume can restart processes" below). The same program is in
68
69
  [`examples/quickstart.ts`](./examples/quickstart.ts).
69
70
 
70
71
  With 0.5.0, wait for the suspend by its operation instead:
@@ -266,6 +267,26 @@ const cell = workspace.cell({ transitionTimeoutMs: 30_000 }); // give up waking
266
267
  await cell.exec.run(['make', 'test']); // resumes the workspace first if it is suspended
267
268
  ```
268
269
 
270
+ **A resume can restart processes (0.11.0+).** A resume normally restores memory and running processes from the
271
+ checkpoint. When the platform's VM runtime changed after the suspend and no host can restore that memory snapshot, the
272
+ resume boots the checkpoint's disk instead (a cold boot), automatically, also when a tool call wakes the workspace. The
273
+ workspace runs, its files are as of the suspend, but every process was restarted, as after a reboot: start dev
274
+ servers, databases and background jobs again. You can tell from the resume's timing:
275
+
276
+ ```ts
277
+ const t = workspace.lastTiming; // after resume({ wait: true }), wake(), or a tool call that woke the workspace
278
+ if (t?.server?.memoryRestored === false) {
279
+ // t.server.resumePath === 'cold_boot', t.server.coldBootReason === 'runtime_changed'
280
+ }
281
+ ```
282
+
283
+ - `ServerTiming.memoryRestored`: `true` when memory and processes came back, `false` when the workspace booted
284
+ (`resumePath` `cold_boot`, or `reset_blank_layer` after a reset), `null` when the API did not say (older APIs).
285
+ - `ServerTiming.coldBootReason`: why it booted (`runtime_changed`), else null. `formatTiming()` prints
286
+ `resume from cold_boot: processes restarted (runtime_changed)`.
287
+ - The finished resume operation carries the same in `result` (`memory_restored`, `cold_boot_reason`, and
288
+ `cold_boot` with details such as `files_as_of`; its shape may change, so do not depend on it).
289
+
269
290
  `suspend`, `resume`, `fork`, `snapshot` and `delete` return the lifecycle operation.
270
291
  `cloud.workspaces.waitForOperation(id)` waits for it (default timeout 5 minutes): each poll asks the API to hold
271
292
  the response until the operation changes (`Prefer: wait`, at most 20 s per request), so completion arrives within
@@ -320,7 +341,8 @@ Here the time went to waiting for a host with capacity; the network and the VM s
320
341
  issuing the first tool token (together).
321
342
  - **server** timing comes from the operation itself (one database clock): `queued` is creation until it began
322
343
  running, including any wait for capacity; `ran` is the cell's work (placement, boot or restore, guest readiness).
323
- `start` / `resume from` and `boot to ready` / `host …` are what the cell reported.
344
+ `start` / `resume from` and `boot to ready` / `host …` are what the cell reported. A resume that booted the saved
345
+ disk instead of restoring memory reads `resume from cold_boot: processes restarted (runtime_changed)` (0.11.0+).
324
346
  - **outside the server** is your total minus the operation's: network, TLS, polling latency, view and token. A large
325
347
  value with a small server total points at the connection between you and the API, not at the workspace.
326
348
  - **retries** lists transient failures the SDK retried (cause and backoff).
package/dist/http.d.ts CHANGED
@@ -1,5 +1,5 @@
1
1
  import type { RetryRecord } from './progress.js';
2
- export declare const SDK_VERSION = "0.10.2";
2
+ export declare const SDK_VERSION = "0.11.0";
3
3
  export interface RequestOptions {
4
4
  query?: Record<string, string | number | boolean | undefined | null>;
5
5
  json?: unknown;
package/dist/http.js CHANGED
@@ -8,7 +8,7 @@
8
8
  */
9
9
  import { ShardfluxApiError, ShardfluxProtocolError, apiError, isErrorBody } from "./errors.js";
10
10
  import { describeFailure } from "./progress.js";
11
- export const SDK_VERSION = '0.10.2';
11
+ export const SDK_VERSION = '0.11.0';
12
12
  export const defaultSleep = (ms) => new Promise((r) => setTimeout(r, ms));
13
13
  /**
14
14
  * The fetch the SDK uses when none is given. On runtimes whose bundled undici is 8.x (Node 26) it sends
@@ -70,8 +70,25 @@ export interface ServerTiming {
70
70
  startPath: string | null;
71
71
  /** `result.warm_fallback`: why a start that could be warm booted instead. */
72
72
  warmFallback: string | null;
73
- /** `result.resume_path`: `local_cache`, `prestaged` (copied to this host ahead of the resume, contracts §23) or `download` (the checkpoint had to be fetched first). */
73
+ /**
74
+ * `result.resume_path`: `local_cache`, `prestaged` (copied to this host ahead of the resume, contracts §23),
75
+ * `download` (the checkpoint had to be fetched first), `cold_boot` (0.11.0+: no host could restore the memory
76
+ * snapshot, so the checkpoint's disk was booted; see `memoryRestored`) or `reset_blank_layer` (the first start after
77
+ * a reset).
78
+ */
74
79
  resumePath: string | null;
80
+ /**
81
+ * `result.memory_restored` (0.11.0+): `true` when memory and running processes came back from the checkpoint;
82
+ * `false` when the workspace booted instead (`cold_boot`, `reset_blank_layer`): its files are as of the suspend, but
83
+ * every process was restarted, as after a reboot. `null` when the result does not say (an older API, or not a
84
+ * resume). Always set by `serverTiming()`; optional only so timings built by hand still type-check.
85
+ */
86
+ memoryRestored?: boolean | null;
87
+ /**
88
+ * `result.cold_boot_reason` (0.11.0+), with `resumePath` `cold_boot`: why the memory could not be restored, e.g.
89
+ * `runtime_changed` (the platform's VM runtime changed after the suspend). Null otherwise.
90
+ */
91
+ coldBootReason?: string | null;
75
92
  /** `result.boot_to_ready_ms`: VM start until the guest agent answered. */
76
93
  bootToReadyMs: number | null;
77
94
  /** `result.host_timings_ms`: the host's own steps (restore: load, after_restore, ready, …). */
@@ -175,6 +192,9 @@ export declare function traced<T>(trace: Trace, fn: () => Promise<T>): Promise<T
175
192
  * server: queued 33.40 s, ran 620 ms, total 34.02 s; start warm, boot to ready 79 ms
176
193
  * outside the server: 161 ms
177
194
  * retries: 1 (POST /v1/workspaces/open, HTTP 503 unavailable, after 200 ms)
195
+ *
196
+ * A resume that booted the workspace instead of restoring its memory (0.11.0+) says so in the server line:
197
+ * `resume from cold_boot: processes restarted (runtime_changed)`.
178
198
  */
179
199
  export declare function formatTiming(t: LifecycleTiming): string;
180
200
  export {};
package/dist/progress.js CHANGED
@@ -63,6 +63,9 @@ export function serverTiming(op) {
63
63
  startPath: str(r.start_path),
64
64
  warmFallback: str(r.warm_fallback),
65
65
  resumePath: str(r.resume_path),
66
+ // Unknown (null) unless the result says so: a result without the field is never read as "not restored".
67
+ memoryRestored: typeof r.memory_restored === 'boolean' ? r.memory_restored : null,
68
+ coldBootReason: str(r.cold_boot_reason),
66
69
  bootToReadyMs: num(r.boot_to_ready_ms),
67
70
  hostTimingsMs,
68
71
  };
@@ -222,6 +225,9 @@ const fmt = (ms) => (ms === null ? '?' : ms < 1000 ? `${Math.round(ms)} ms` : `$
222
225
  * server: queued 33.40 s, ran 620 ms, total 34.02 s; start warm, boot to ready 79 ms
223
226
  * outside the server: 161 ms
224
227
  * retries: 1 (POST /v1/workspaces/open, HTTP 503 unavailable, after 200 ms)
228
+ *
229
+ * A resume that booted the workspace instead of restoring its memory (0.11.0+) says so in the server line:
230
+ * `resume from cold_boot: processes restarted (runtime_changed)`.
225
231
  */
226
232
  export function formatTiming(t) {
227
233
  const ids = [t.workspaceId ? `workspace ${t.workspaceId}` : null, t.operationId ? `operation ${t.operationId}` : null].filter(Boolean).join(', ');
@@ -242,7 +248,7 @@ export function formatTiming(t) {
242
248
  const parts = [s.queuedMs !== null ? `queued ${fmt(s.queuedMs)}` : null, s.runMs !== null ? `ran ${fmt(s.runMs)}` : null, s.totalMs !== null ? `total ${fmt(s.totalMs)}` : `still ${s.state}`];
243
249
  const how = [
244
250
  s.startPath ? `start ${s.startPath}${s.warmFallback ? ` (warm fallback: ${s.warmFallback})` : ''}` : null,
245
- s.resumePath ? `resume from ${s.resumePath}` : null,
251
+ s.resumePath ? `resume from ${s.resumePath}${s.memoryRestored === false ? `: processes restarted${s.coldBootReason ? ` (${s.coldBootReason})` : ''}` : ''}` : null,
246
252
  s.bootToReadyMs !== null ? `boot to ready ${fmt(s.bootToReadyMs)}` : null,
247
253
  s.hostTimingsMs ? `host ${Object.entries(s.hostTimingsMs).map(([k, v]) => `${k} ${fmt(v)}`).join(', ')}` : null,
248
254
  ].filter(Boolean);
@@ -65,7 +65,9 @@ export declare class Workspace {
65
65
  /**
66
66
  * Where the time went in the last lifecycle call made through this handle: open(), wake() (also when a tool call
67
67
  * woke the workspace), or a lifecycle call with `wait`. Null for handles from get()/list() until such a call.
68
- * `formatTiming(workspace.lastTiming)` prints it.
68
+ * `formatTiming(workspace.lastTiming)` prints it. After a resume or wake, `server.memoryRestored === false` (0.11.0+)
69
+ * means the workspace booted from its saved disk instead of restoring its memory (`server.resumePath` `cold_boot`,
70
+ * reason in `server.coldBootReason`): files are as of the suspend, running processes were restarted.
69
71
  */
70
72
  get lastTiming(): LifecycleTiming | null;
71
73
  get id(): string;
@@ -172,7 +174,9 @@ export declare class Workspace {
172
174
  * Tool calls wake a suspended workspace by themselves, so this is rarely needed. With `wait` (0.9.0) it is one held
173
175
  * request (contracts §22.6): the handle takes the running view and a tool token for `agentLabel`/`tools` (defaults:
174
176
  * those given to open()), so `cell()` calls with that label and tools start at once. A workspace that is already
175
- * running is ShardfluxApiError 409 `conflict` (`already_running`).
177
+ * running is ShardfluxApiError 409 `conflict` (`already_running`). The finished operation's `result.memory_restored`
178
+ * (also `lastTiming.server.memoryRestored`, 0.11.0+) is false when the resume booted the saved disk instead
179
+ * (`resume_path` `cold_boot`): files kept, processes restarted.
176
180
  */
177
181
  resume(opts: WaitedResumeOptions): Promise<FinishedOperation>;
178
182
  resume(opts?: ResumeOptions): Promise<Operation>;
@@ -265,6 +269,10 @@ export declare class Workspace {
265
269
  * running workspace and a tool token for `agentLabel`/`tools`, which this handle keeps, so a tool call that woke the
266
270
  * workspace is retried at once (refused call, resume, call). Timing: one `request` phase with reason `held`. A server
267
271
  * that does not hold the request answers at once; the wake then waits for the operation and reads the view.
272
+ *
273
+ * A wake whose resume could not restore the workspace's memory (0.11.0+; the platform's VM runtime changed after the
274
+ * suspend) still resolves `true`: the workspace runs from its saved disk, with every process restarted. `lastTiming`
275
+ * and the `done` progress event say so (`server.memoryRestored === false`, `server.resumePath` `cold_boot`).
268
276
  */
269
277
  wake(opts?: WakeOptions): Promise<boolean>;
270
278
  /**
package/dist/workspace.js CHANGED
@@ -42,7 +42,9 @@ export class Workspace {
42
42
  /**
43
43
  * Where the time went in the last lifecycle call made through this handle: open(), wake() (also when a tool call
44
44
  * woke the workspace), or a lifecycle call with `wait`. Null for handles from get()/list() until such a call.
45
- * `formatTiming(workspace.lastTiming)` prints it.
45
+ * `formatTiming(workspace.lastTiming)` prints it. After a resume or wake, `server.memoryRestored === false` (0.11.0+)
46
+ * means the workspace booted from its saved disk instead of restoring its memory (`server.resumePath` `cold_boot`,
47
+ * reason in `server.coldBootReason`): files are as of the suspend, running processes were restarted.
46
48
  */
47
49
  get lastTiming() {
48
50
  return this.#lastTiming ?? this.#openTrace?.finished ?? null;
@@ -390,6 +392,10 @@ export class Workspace {
390
392
  * running workspace and a tool token for `agentLabel`/`tools`, which this handle keeps, so a tool call that woke the
391
393
  * workspace is retried at once (refused call, resume, call). Timing: one `request` phase with reason `held`. A server
392
394
  * that does not hold the request answers at once; the wake then waits for the operation and reads the view.
395
+ *
396
+ * A wake whose resume could not restore the workspace's memory (0.11.0+; the platform's VM runtime changed after the
397
+ * suspend) still resolves `true`: the workspace runs from its saved disk, with every process restarted. `lastTiming`
398
+ * and the `done` progress event say so (`server.memoryRestored === false`, `server.resumePath` `cold_boot`).
393
399
  */
394
400
  async wake(opts = {}) {
395
401
  // A file-first workspace runs from creation and is never suspended (contracts §29.7): nothing to wake.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@shardflux/sdk",
3
- "version": "0.10.2",
3
+ "version": "0.11.0",
4
4
  "type": "module",
5
5
  "description": "Shardflux TypeScript SDK: open persistent agent workspaces by key and give your agent workspace tools (exec, files, processes, PTY, git, browser).",
6
6
  "license": "Apache-2.0",