@shardflux/sdk 0.10.2 → 0.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +24 -1
- package/README.md +25 -3
- package/dist/http.d.ts +1 -1
- package/dist/http.js +1 -1
- package/dist/progress.d.ts +21 -1
- package/dist/progress.js +7 -1
- package/dist/workspace.d.ts +10 -2
- package/dist/workspace.js +7 -1
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -3,7 +3,30 @@
|
|
|
3
3
|
Every API the README shows is available from the version named here. Below 1.0, a minor release may break
|
|
4
4
|
compatibility; breaking changes are marked **Breaking**.
|
|
5
5
|
|
|
6
|
-
## 0.
|
|
6
|
+
## 0.11.0 (not yet published)
|
|
7
|
+
|
|
8
|
+
### A resume that restarted processes says so (cold boot, contracts §12)
|
|
9
|
+
|
|
10
|
+
Additive. When the platform's VM runtime changed after a suspend and no host can restore the memory snapshot, the
|
|
11
|
+
cell resumes the workspace by booting its checkpoint's disk (a cold boot), automatically, also when a tool call wakes
|
|
12
|
+
the workspace. The workspace runs and its files are as of the suspend, but every process was restarted. The API passes
|
|
13
|
+
the resume's result through; this release reads it.
|
|
14
|
+
|
|
15
|
+
- `ServerTiming.memoryRestored` (`boolean | null`): `result.memory_restored`. `true` when memory and running
|
|
16
|
+
processes came back; `false` for `cold_boot` (and `reset_blank_layer`); `null` when the result does not say (an API
|
|
17
|
+
from before the field), never read as `false`.
|
|
18
|
+
- `ServerTiming.coldBootReason` (`string | null`): `result.cold_boot_reason` with `resumePath` `cold_boot`, today
|
|
19
|
+
`runtime_changed`.
|
|
20
|
+
- `ServerTiming.resumePath` documents `cold_boot` and `reset_blank_layer`.
|
|
21
|
+
- `formatTiming()` prints `resume from cold_boot: processes restarted (runtime_changed)` when memory was not restored;
|
|
22
|
+
other resumes print as before.
|
|
23
|
+
- Where it shows: `workspace.lastTiming` and the `done` progress event after `resume({ wait: true })`, `wake()` and a
|
|
24
|
+
tool call that woke the workspace; the finished resume operation's `result` (`memory_restored`,
|
|
25
|
+
`cold_boot_reason`, and `cold_boot` with operator details such as `files_as_of`, whose shape may change).
|
|
26
|
+
- The new `ServerTiming` fields are optional in the type (always set by the SDK), so timings built by hand still
|
|
27
|
+
type-check. No return type or behaviour changes: a wake that cold-booted still resolves `true`.
|
|
28
|
+
|
|
29
|
+
## 0.10.2
|
|
7
30
|
|
|
8
31
|
Fix: the SDK reports 0.10.2. 0.10.1 was published reporting 0.10.0, its previous version; the build now
|
|
9
32
|
fails when `SDK_VERSION` differs from package.json.
|
package/README.md
CHANGED
|
@@ -10,7 +10,7 @@ tool calls into the workspace ([tool-call capture](#tool-call-capture-070)).
|
|
|
10
10
|
> **Early access.** Shardflux is in early access. The API is versioned (`/v1`), but this SDK is
|
|
11
11
|
> below 1.0: a minor release may contain breaking changes (see [Compatibility](#compatibility)).
|
|
12
12
|
|
|
13
|
-
> **Versions.** This README describes 0.
|
|
13
|
+
> **Versions.** This README describes 0.11.0. Anything marked **(0.11.0+)** is not in 0.10.x, **(0.10.0+)** not in 0.9.0, **(0.9.0+)** not in 0.8.x, **(0.8.0+)** not in 0.7.x,
|
|
14
14
|
> **(0.7.0+)** not in 0.6.x and **(0.6.0+)** not in 0.5.0; [CHANGELOG.md](./CHANGELOG.md) lists what each version added. Check yours with
|
|
15
15
|
> `npm ls @shardflux/sdk` or the exported `SDK_VERSION`.
|
|
16
16
|
|
|
@@ -64,7 +64,8 @@ console.log(formatTiming(again.lastTiming!)); // (0.6.0+) where the r
|
|
|
64
64
|
```
|
|
65
65
|
|
|
66
66
|
`open()` waits until the workspace is running. Opening the same key again never resets it: files,
|
|
67
|
-
installed packages and running processes are still there
|
|
67
|
+
installed packages and running processes are still there (a resume restarts processes only in the rare cold boot,
|
|
68
|
+
see "A resume can restart processes" below). The same program is in
|
|
68
69
|
[`examples/quickstart.ts`](./examples/quickstart.ts).
|
|
69
70
|
|
|
70
71
|
With 0.5.0, wait for the suspend by its operation instead:
|
|
@@ -266,6 +267,26 @@ const cell = workspace.cell({ transitionTimeoutMs: 30_000 }); // give up waking
|
|
|
266
267
|
await cell.exec.run(['make', 'test']); // resumes the workspace first if it is suspended
|
|
267
268
|
```
|
|
268
269
|
|
|
270
|
+
**A resume can restart processes (0.11.0+).** A resume normally restores memory and running processes from the
|
|
271
|
+
checkpoint. When the platform's VM runtime changed after the suspend and no host can restore that memory snapshot, the
|
|
272
|
+
resume boots the checkpoint's disk instead (a cold boot), automatically, also when a tool call wakes the workspace. The
|
|
273
|
+
workspace runs, its files are as of the suspend, but every process was restarted, as after a reboot: start dev
|
|
274
|
+
servers, databases and background jobs again. You can tell from the resume's timing:
|
|
275
|
+
|
|
276
|
+
```ts
|
|
277
|
+
const t = workspace.lastTiming; // after resume({ wait: true }), wake(), or a tool call that woke the workspace
|
|
278
|
+
if (t?.server?.memoryRestored === false) {
|
|
279
|
+
// t.server.resumePath === 'cold_boot', t.server.coldBootReason === 'runtime_changed'
|
|
280
|
+
}
|
|
281
|
+
```
|
|
282
|
+
|
|
283
|
+
- `ServerTiming.memoryRestored`: `true` when memory and processes came back, `false` when the workspace booted
|
|
284
|
+
(`resumePath` `cold_boot`, or `reset_blank_layer` after a reset), `null` when the API did not say (older APIs).
|
|
285
|
+
- `ServerTiming.coldBootReason`: why it booted (`runtime_changed`), else null. `formatTiming()` prints
|
|
286
|
+
`resume from cold_boot: processes restarted (runtime_changed)`.
|
|
287
|
+
- The finished resume operation carries the same in `result` (`memory_restored`, `cold_boot_reason`, and
|
|
288
|
+
`cold_boot` with details such as `files_as_of`; its shape may change, so do not depend on it).
|
|
289
|
+
|
|
269
290
|
`suspend`, `resume`, `fork`, `snapshot` and `delete` return the lifecycle operation.
|
|
270
291
|
`cloud.workspaces.waitForOperation(id)` waits for it (default timeout 5 minutes): each poll asks the API to hold
|
|
271
292
|
the response until the operation changes (`Prefer: wait`, at most 20 s per request), so completion arrives within
|
|
@@ -320,7 +341,8 @@ Here the time went to waiting for a host with capacity; the network and the VM s
|
|
|
320
341
|
issuing the first tool token (together).
|
|
321
342
|
- **server** timing comes from the operation itself (one database clock): `queued` is creation until it began
|
|
322
343
|
running, including any wait for capacity; `ran` is the cell's work (placement, boot or restore, guest readiness).
|
|
323
|
-
`start` / `resume from` and `boot to ready` / `host …` are what the cell reported.
|
|
344
|
+
`start` / `resume from` and `boot to ready` / `host …` are what the cell reported. A resume that booted the saved
|
|
345
|
+
disk instead of restoring memory reads `resume from cold_boot: processes restarted (runtime_changed)` (0.11.0+).
|
|
324
346
|
- **outside the server** is your total minus the operation's: network, TLS, polling latency, view and token. A large
|
|
325
347
|
value with a small server total points at the connection between you and the API, not at the workspace.
|
|
326
348
|
- **retries** lists transient failures the SDK retried (cause and backoff).
|
package/dist/http.d.ts
CHANGED
package/dist/http.js
CHANGED
|
@@ -8,7 +8,7 @@
|
|
|
8
8
|
*/
|
|
9
9
|
import { ShardfluxApiError, ShardfluxProtocolError, apiError, isErrorBody } from "./errors.js";
|
|
10
10
|
import { describeFailure } from "./progress.js";
|
|
11
|
-
export const SDK_VERSION = '0.
|
|
11
|
+
export const SDK_VERSION = '0.11.0';
|
|
12
12
|
export const defaultSleep = (ms) => new Promise((r) => setTimeout(r, ms));
|
|
13
13
|
/**
|
|
14
14
|
* The fetch the SDK uses when none is given. On runtimes whose bundled undici is 8.x (Node 26) it sends
|
package/dist/progress.d.ts
CHANGED
|
@@ -70,8 +70,25 @@ export interface ServerTiming {
|
|
|
70
70
|
startPath: string | null;
|
|
71
71
|
/** `result.warm_fallback`: why a start that could be warm booted instead. */
|
|
72
72
|
warmFallback: string | null;
|
|
73
|
-
/**
|
|
73
|
+
/**
|
|
74
|
+
* `result.resume_path`: `local_cache`, `prestaged` (copied to this host ahead of the resume, contracts §23),
|
|
75
|
+
* `download` (the checkpoint had to be fetched first), `cold_boot` (0.11.0+: no host could restore the memory
|
|
76
|
+
* snapshot, so the checkpoint's disk was booted; see `memoryRestored`) or `reset_blank_layer` (the first start after
|
|
77
|
+
* a reset).
|
|
78
|
+
*/
|
|
74
79
|
resumePath: string | null;
|
|
80
|
+
/**
|
|
81
|
+
* `result.memory_restored` (0.11.0+): `true` when memory and running processes came back from the checkpoint;
|
|
82
|
+
* `false` when the workspace booted instead (`cold_boot`, `reset_blank_layer`): its files are as of the suspend, but
|
|
83
|
+
* every process was restarted, as after a reboot. `null` when the result does not say (an older API, or not a
|
|
84
|
+
* resume). Always set by `serverTiming()`; optional only so timings built by hand still type-check.
|
|
85
|
+
*/
|
|
86
|
+
memoryRestored?: boolean | null;
|
|
87
|
+
/**
|
|
88
|
+
* `result.cold_boot_reason` (0.11.0+), with `resumePath` `cold_boot`: why the memory could not be restored, e.g.
|
|
89
|
+
* `runtime_changed` (the platform's VM runtime changed after the suspend). Null otherwise.
|
|
90
|
+
*/
|
|
91
|
+
coldBootReason?: string | null;
|
|
75
92
|
/** `result.boot_to_ready_ms`: VM start until the guest agent answered. */
|
|
76
93
|
bootToReadyMs: number | null;
|
|
77
94
|
/** `result.host_timings_ms`: the host's own steps (restore: load, after_restore, ready, …). */
|
|
@@ -175,6 +192,9 @@ export declare function traced<T>(trace: Trace, fn: () => Promise<T>): Promise<T
|
|
|
175
192
|
* server: queued 33.40 s, ran 620 ms, total 34.02 s; start warm, boot to ready 79 ms
|
|
176
193
|
* outside the server: 161 ms
|
|
177
194
|
* retries: 1 (POST /v1/workspaces/open, HTTP 503 unavailable, after 200 ms)
|
|
195
|
+
*
|
|
196
|
+
* A resume that booted the workspace instead of restoring its memory (0.11.0+) says so in the server line:
|
|
197
|
+
* `resume from cold_boot: processes restarted (runtime_changed)`.
|
|
178
198
|
*/
|
|
179
199
|
export declare function formatTiming(t: LifecycleTiming): string;
|
|
180
200
|
export {};
|
package/dist/progress.js
CHANGED
|
@@ -63,6 +63,9 @@ export function serverTiming(op) {
|
|
|
63
63
|
startPath: str(r.start_path),
|
|
64
64
|
warmFallback: str(r.warm_fallback),
|
|
65
65
|
resumePath: str(r.resume_path),
|
|
66
|
+
// Unknown (null) unless the result says so: a result without the field is never read as "not restored".
|
|
67
|
+
memoryRestored: typeof r.memory_restored === 'boolean' ? r.memory_restored : null,
|
|
68
|
+
coldBootReason: str(r.cold_boot_reason),
|
|
66
69
|
bootToReadyMs: num(r.boot_to_ready_ms),
|
|
67
70
|
hostTimingsMs,
|
|
68
71
|
};
|
|
@@ -222,6 +225,9 @@ const fmt = (ms) => (ms === null ? '?' : ms < 1000 ? `${Math.round(ms)} ms` : `$
|
|
|
222
225
|
* server: queued 33.40 s, ran 620 ms, total 34.02 s; start warm, boot to ready 79 ms
|
|
223
226
|
* outside the server: 161 ms
|
|
224
227
|
* retries: 1 (POST /v1/workspaces/open, HTTP 503 unavailable, after 200 ms)
|
|
228
|
+
*
|
|
229
|
+
* A resume that booted the workspace instead of restoring its memory (0.11.0+) says so in the server line:
|
|
230
|
+
* `resume from cold_boot: processes restarted (runtime_changed)`.
|
|
225
231
|
*/
|
|
226
232
|
export function formatTiming(t) {
|
|
227
233
|
const ids = [t.workspaceId ? `workspace ${t.workspaceId}` : null, t.operationId ? `operation ${t.operationId}` : null].filter(Boolean).join(', ');
|
|
@@ -242,7 +248,7 @@ export function formatTiming(t) {
|
|
|
242
248
|
const parts = [s.queuedMs !== null ? `queued ${fmt(s.queuedMs)}` : null, s.runMs !== null ? `ran ${fmt(s.runMs)}` : null, s.totalMs !== null ? `total ${fmt(s.totalMs)}` : `still ${s.state}`];
|
|
243
249
|
const how = [
|
|
244
250
|
s.startPath ? `start ${s.startPath}${s.warmFallback ? ` (warm fallback: ${s.warmFallback})` : ''}` : null,
|
|
245
|
-
s.resumePath ? `resume from ${s.resumePath}` : null,
|
|
251
|
+
s.resumePath ? `resume from ${s.resumePath}${s.memoryRestored === false ? `: processes restarted${s.coldBootReason ? ` (${s.coldBootReason})` : ''}` : ''}` : null,
|
|
246
252
|
s.bootToReadyMs !== null ? `boot to ready ${fmt(s.bootToReadyMs)}` : null,
|
|
247
253
|
s.hostTimingsMs ? `host ${Object.entries(s.hostTimingsMs).map(([k, v]) => `${k} ${fmt(v)}`).join(', ')}` : null,
|
|
248
254
|
].filter(Boolean);
|
package/dist/workspace.d.ts
CHANGED
|
@@ -65,7 +65,9 @@ export declare class Workspace {
|
|
|
65
65
|
/**
|
|
66
66
|
* Where the time went in the last lifecycle call made through this handle: open(), wake() (also when a tool call
|
|
67
67
|
* woke the workspace), or a lifecycle call with `wait`. Null for handles from get()/list() until such a call.
|
|
68
|
-
* `formatTiming(workspace.lastTiming)` prints it.
|
|
68
|
+
* `formatTiming(workspace.lastTiming)` prints it. After a resume or wake, `server.memoryRestored === false` (0.11.0+)
|
|
69
|
+
* means the workspace booted from its saved disk instead of restoring its memory (`server.resumePath` `cold_boot`,
|
|
70
|
+
* reason in `server.coldBootReason`): files are as of the suspend, running processes were restarted.
|
|
69
71
|
*/
|
|
70
72
|
get lastTiming(): LifecycleTiming | null;
|
|
71
73
|
get id(): string;
|
|
@@ -172,7 +174,9 @@ export declare class Workspace {
|
|
|
172
174
|
* Tool calls wake a suspended workspace by themselves, so this is rarely needed. With `wait` (0.9.0) it is one held
|
|
173
175
|
* request (contracts §22.6): the handle takes the running view and a tool token for `agentLabel`/`tools` (defaults:
|
|
174
176
|
* those given to open()), so `cell()` calls with that label and tools start at once. A workspace that is already
|
|
175
|
-
* running is ShardfluxApiError 409 `conflict` (`already_running`).
|
|
177
|
+
* running is ShardfluxApiError 409 `conflict` (`already_running`). The finished operation's `result.memory_restored`
|
|
178
|
+
* (also `lastTiming.server.memoryRestored`, 0.11.0+) is false when the resume booted the saved disk instead
|
|
179
|
+
* (`resume_path` `cold_boot`): files kept, processes restarted.
|
|
176
180
|
*/
|
|
177
181
|
resume(opts: WaitedResumeOptions): Promise<FinishedOperation>;
|
|
178
182
|
resume(opts?: ResumeOptions): Promise<Operation>;
|
|
@@ -265,6 +269,10 @@ export declare class Workspace {
|
|
|
265
269
|
* running workspace and a tool token for `agentLabel`/`tools`, which this handle keeps, so a tool call that woke the
|
|
266
270
|
* workspace is retried at once (refused call, resume, call). Timing: one `request` phase with reason `held`. A server
|
|
267
271
|
* that does not hold the request answers at once; the wake then waits for the operation and reads the view.
|
|
272
|
+
*
|
|
273
|
+
* A wake whose resume could not restore the workspace's memory (0.11.0+; the platform's VM runtime changed after the
|
|
274
|
+
* suspend) still resolves `true`: the workspace runs from its saved disk, with every process restarted. `lastTiming`
|
|
275
|
+
* and the `done` progress event say so (`server.memoryRestored === false`, `server.resumePath` `cold_boot`).
|
|
268
276
|
*/
|
|
269
277
|
wake(opts?: WakeOptions): Promise<boolean>;
|
|
270
278
|
/**
|
package/dist/workspace.js
CHANGED
|
@@ -42,7 +42,9 @@ export class Workspace {
|
|
|
42
42
|
/**
|
|
43
43
|
* Where the time went in the last lifecycle call made through this handle: open(), wake() (also when a tool call
|
|
44
44
|
* woke the workspace), or a lifecycle call with `wait`. Null for handles from get()/list() until such a call.
|
|
45
|
-
* `formatTiming(workspace.lastTiming)` prints it.
|
|
45
|
+
* `formatTiming(workspace.lastTiming)` prints it. After a resume or wake, `server.memoryRestored === false` (0.11.0+)
|
|
46
|
+
* means the workspace booted from its saved disk instead of restoring its memory (`server.resumePath` `cold_boot`,
|
|
47
|
+
* reason in `server.coldBootReason`): files are as of the suspend, running processes were restarted.
|
|
46
48
|
*/
|
|
47
49
|
get lastTiming() {
|
|
48
50
|
return this.#lastTiming ?? this.#openTrace?.finished ?? null;
|
|
@@ -390,6 +392,10 @@ export class Workspace {
|
|
|
390
392
|
* running workspace and a tool token for `agentLabel`/`tools`, which this handle keeps, so a tool call that woke the
|
|
391
393
|
* workspace is retried at once (refused call, resume, call). Timing: one `request` phase with reason `held`. A server
|
|
392
394
|
* that does not hold the request answers at once; the wake then waits for the operation and reads the view.
|
|
395
|
+
*
|
|
396
|
+
* A wake whose resume could not restore the workspace's memory (0.11.0+; the platform's VM runtime changed after the
|
|
397
|
+
* suspend) still resolves `true`: the workspace runs from its saved disk, with every process restarted. `lastTiming`
|
|
398
|
+
* and the `done` progress event say so (`server.memoryRestored === false`, `server.resumePath` `cold_boot`).
|
|
393
399
|
*/
|
|
394
400
|
async wake(opts = {}) {
|
|
395
401
|
// A file-first workspace runs from creation and is never suspended (contracts §29.7): nothing to wake.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@shardflux/sdk",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.11.0",
|
|
4
4
|
"type": "module",
|
|
5
5
|
"description": "Shardflux TypeScript SDK: open persistent agent workspaces by key and give your agent workspace tools (exec, files, processes, PTY, git, browser).",
|
|
6
6
|
"license": "Apache-2.0",
|