@shardflux/sdk 0.11.0 → 0.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,13 +1,59 @@
1
1
  # Changelog
2
2
 
3
- Every API the README shows is available from the version named here. Below 1.0, a minor release may break
4
- compatibility; breaking changes are marked **Breaking**.
3
+ Every API the README shows is available from the version named here. Breaking changes ship in minor releases and are
4
+ marked **Breaking**.
5
5
 
6
- ## 0.11.0 (not yet published)
6
+ ## 0.12.0 (release candidate)
7
7
 
8
- ### A resume that restarted processes says so (cold boot, contracts §12)
8
+ Completed exec output streams are drained before releasing their HTTP connections, with a bounded cleanup if a peer does not close.
9
9
 
10
- Additive. When the platform's VM runtime changed after a suspend and no host can restore the memory snapshot, the
10
+ Pooled HTTP/1.1 on Node 26 with pinned undici 8.10.2; custom fetch stays unchanged. Labels on open/list and setLabels, typed idle(), keepalive() and setIdlePolicy(). Protocol errors carry source; isWorkspaceGone recognizes only an explicit API workspace_deleted refusal. Exec sessions accept stdin_open and offset-addressed exec.input. Failed workspaces recover through resume/auto-wake with the same ID. Completed deletions free keys for new IDs.
11
+
12
+ Requires the QM integration backend release for labels, failed recovery, key reuse and pipe stdin. No production deployment has occurred from this branch.
13
+
14
+ ### Instant suspend: durable storage in the result
15
+
16
+ Additive. A suspend returns as soon as the workspace is sealed on its host; the copy lands in durable storage right
17
+ after. This release reads that from the result and can wait for it.
18
+
19
+ - `suspend({ durable: true })` (`SuspendOptions.durable`, also `cloud.workspaces.suspend(id, { durable: true })`):
20
+ resolves once `result.durable` is `true`. It implies `wait` and uses the same `timeoutMs`/`signal` (one budget for
21
+ the suspend and the copy). Rejects with `DurabilityLostError` (new; `durability` with the reason) when the copy
22
+ cannot be made, and `OperationTimeoutError` with `durable: true` when the time runs out (the copy continues).
23
+ `suspend({ wait: true })` still resolves when the suspend succeeds, without waiting for the copy.
24
+ - `cloud.workspaces.waitForDurable(operationOrId, waitOptions)`: the same wait for an operation you hold (a suspend, or
25
+ a fork of a running workspace).
26
+ - `isDurable(op)`, `durabilityOf(op)` (`Durability`: `state` `pending` | `durable` | `lost`, `checkpointId`,
27
+ `generationId`, `localCommitAt`, `durableBy`, `durableAt`, `localCommitToDurableMs`, `overdueAt`, `reason`) and
28
+ `lostSuspendOf(op)` (`LostSuspend`, on a resume result). `ServerTiming` gains `durable`, `suspendPath`, `durability`
29
+ and `lostSuspend`; progress phase `durable` (reason `overdue` past `durableBy`); `formatTiming()` prints them.
30
+ - `Operation.result` is typed with `durable`, `suspend_path`, `durability` and `lost_suspend` (OpenAPI).
31
+ - `OperationFailedError.workspaceActive`: a suspend-when-idle canceled with `workspace_active` (the workspace was in
32
+ use; nothing changed and it keeps running). `workspace.waitUntilReady()` resolves for it instead of throwing, as
33
+ `wake()` already did.
34
+
35
+ Requires the instant-suspend cell release for `durable: false` results; against other servers every succeeded suspend
36
+ is already durable and `durable: true` resolves at once.
37
+
38
+ ## 0.11.1
39
+
40
+ Wording: messages and JSDoc say what to do, without internals (no behaviour change).
41
+
42
+ - `OperationTimeoutError` on a queued start: `... it stays queued server side until <deadline_at> and fails with
43
+ capacity_unavailable if it has not started by then.` (the same words as the Python SDK).
44
+ - JSDoc (IDE hovers): `sendFeedback` goes to the Shardflux team; `capacity_unavailable` is retryable: the start passed
45
+ its deadline and nothing was started, send it again; `no_execution_host`: the execution cannot be placed right now;
46
+ `host_feature_unavailable`: the call is not available for this workspace, use the fallback in the hint;
47
+ `resumePath` `cold_boot`: the saved disk was booted after a platform runtime change, processes restarted; an
48
+ execution runs to completion; on Node 26 every request uses a fresh connection (`Connection: close`) unless
49
+ `SHARDFLUX_HTTP_KEEPALIVE=1`.
50
+ - The `formatTiming()` example is a production resume (413 ms).
51
+
52
+ ## 0.11.0
53
+
54
+ ### A resume that restarted processes says so (cold boot)
55
+
56
+ Additive. After a platform runtime change, the
11
57
  cell resumes the workspace by booting its checkpoint's disk (a cold boot), automatically, also when a tool call wakes
12
58
  the workspace. The workspace runs and its files are as of the suspend, but every process was restarted. The API passes
13
59
  the resume's result through; this release reads it.
@@ -80,7 +126,7 @@ from the API's OpenAPI); the account client's `setSpendPolicy()` gains the overa
80
126
  `KnownErrorReason` adds them, the spend-policy refusals (422 `overage_unavailable`, `spend_cap_required`,
81
127
  `spend_cap_below_minimum`, `spend_cap_above_plan_price`, `spend_cap_below_charges`) and 409 `version_mismatch`.
82
128
 
83
- ### Suspend when idle (contracts §20.6)
129
+ ### Suspend when idle
84
130
 
85
131
  - `workspace.suspendWhenIdle({ afterSeconds, idempotencyKey? })` and `cloud.workspaces.suspendWhenIdle(id, {
86
132
  afterSeconds })` (POST /v1/workspaces/{id}/suspend-when-idle): the workspace is suspended once it has been idle for
@@ -94,15 +140,14 @@ from the API's OpenAPI); the account client's `setSpendPolicy()` gains the overa
94
140
  count as the next turn and cancel the request.
95
141
  - Types `SuspendRequest`, `SuspendWhenIdleOptions`, `SuspendWhenIdleResult`, `SuspendWhenIdleResponse`.
96
142
 
97
- ## 0.9.0 (not yet published; npm `latest` is 0.8.0)
143
+ ## 0.9.0
98
144
 
99
145
  Elastic compute (decision 0007): file tools and wake hints for parked workspaces. Additive; older APIs and cell
100
146
  gateways keep working (the new calls answer 404 there).
101
147
 
102
148
  ### File search, patches with revisions, the wake hint
103
149
 
104
- Needs a cell gateway with the contracts §26 routes (`files/search`, `files/patch`, `wake-hint`, revisions); an older
105
- gateway does not serve them and returns no revisions.
150
+ Uses the `files/search`, `files/patch` and `wake-hint` routes and file revisions.
106
151
 
107
152
  - `cell.files.search(path, pattern, opts)`: content search under a directory (literal or RE2 with `regex`,
108
153
  `caseInsensitive`, `include`/`exclude` globs, `maxMatches`, `maxFileBytes`, `contextLines`). Returns the gateway's
@@ -137,24 +182,22 @@ gateway does not serve them and returns no revisions.
137
182
  of a sleeping workspace `offline_unavailable` and `offline_budget` (409 `workspace_not_running`: woken and retried
138
183
  like any) and `offline_changed` (503, retried; served by the running workspace).
139
184
  - Reads of a suspended workspace without a held token: the API now issues tool tokens for a suspended workspace
140
- (contracts §26.4), so `read`, `readText`, `readWithInfo`, `stat`, `list` and `search` of a suspended workspace are
185
+ , so `read`, `readText`, `readWithInfo`, `stat`, `list` and `search` of a suspended workspace are
141
186
  served from its disk without waking it also for a handle that fetches its first token after the suspend (before,
142
187
  the token was refused and the call woke the workspace). Any other call still wakes it: the cell refuses it with 409
143
188
  `workspace_not_running`. Needs that API; with an older one the token is refused and the call wakes it as before.
144
189
  - The agent-tool runner sends no wake hint for `read_file`, `list_files` and `search_files`: a sleeping workspace
145
190
  serves them from its disk, and the hint would wake a suspended one (or restore a hibernated one) for nothing.
146
- - Hosts without the features (contracts §26.7): `KnownErrorReason` adds `host_feature_unavailable` (409 `conflict`,
147
- not retryable, `details.feature` `file_search` or `file_patch`): the workspace runs on a host agent that predates
148
- the call, until it runs on an upgraded host. It is neither retried nor answered with a wake. Such a host also
149
- returns no revisions (`FileInfo.revision`, `readWithInfo().revision` are absent).
191
+ - Workspaces without the features: `KnownErrorReason` adds `host_feature_unavailable` (409 `conflict`,
192
+ not retryable, `details.feature` `file_search` or `file_patch`): the call is not available for the workspace. It is neither retried nor answered with a wake. Such a workspace also returns no revisions (`FileInfo.revision`, `readWithInfo().revision` are absent).
150
193
  - `JsonSchema.pattern`, checked by `validateArgs()`.
151
194
  - Types: `FileSearchOptions`, `FileSearchResponse`, `FileSearchRequest`, `FileSearchResult`, `FileSearchMatch`,
152
195
  `FilePatchParams`, `FilePatchEdit`, `FilePatchRequest`, `FilePatchResult`, `FileEdit`, `FileRevision`,
153
196
  `FileReadResult`, `ServedFrom`, `Residency`, `WakeHintResult`, `HintOptions`, `HintResult`.
154
197
 
155
- ### File-first workspaces (contracts §29)
198
+ ### File-first workspaces
156
199
 
157
- Needs an API with `FILE_FIRST_WORKSPACES` on and a cell that serves file-first workspaces (contracts §29.8). A
200
+ Needs an API with `FILE_FIRST_WORKSPACES` on and a cell that serves file-first workspaces. A
158
201
  file-first workspace has no VM between executions: its state is a versioned file tree under /home/user, and each
159
202
  command runs in a fresh VM whose changed files become the next tree revision. Processful workspaces are unchanged.
160
203
 
@@ -203,7 +246,7 @@ command runs in a fresh VM whose changed files become the next tree revision. Pr
203
246
  - Types: `WorkspaceMode`, `ExecutionResult`, `ExecutionResultBody`, `ExecutionChange`, `ExecutionState`,
204
247
  `ExecutionError`, `ExecutionRunOptions`, `ExecutionGetOptions`, `TreeRevisionOptions`.
205
248
 
206
- ### Waking a suspended workspace is one request (contracts §22.6)
249
+ ### Waking a suspended workspace is one request
207
250
 
208
251
  Needs an API with the held resume for the single request; against an older API every call below works as in 0.7.0.
209
252
 
@@ -222,9 +265,9 @@ Needs an API with the held resume for the single request; against an older API e
222
265
  choose the token that comes back; `cell()`'s own wake passes its label and tools. `WorkspacesApi.requestResume()`
223
266
  is the bare request (`ResumeAnswer`, `ResumeRequestOptions`, `ResumeResponse`).
224
267
  - `CellClient`: after a wake that put a new token into the client's manager, the retry uses it instead of fetching one.
225
- ### The account plane and the version check (contracts §30)
268
+ ### The account plane and the version check
226
269
 
227
- Needs an API with CLI sessions and client versions (contracts §30). Every API is additive; the one change in
270
+ Needs an API with CLI sessions and client versions. Every API is additive; the one change in
228
271
  behavior is the automatic version check (below), which makes one background request per process.
229
272
 
230
273
  ### Account plane: `ShardfluxAccount`
@@ -279,9 +322,9 @@ behavior is the automatic version check (below), which makes one background requ
279
322
  Off with `SHARDFLUX_NO_UPDATE_CHECK=1` (also `true`, `yes`, `on`) or `NO_UPDATE_NOTIFIER=1`. `fetchBillingCatalog()`
280
323
  does not check.
281
324
 
282
- ### Feedback straight to the founder (POST /v1/feedback)
325
+ ### Feedback straight to the Shardflux team (POST /v1/feedback)
283
326
 
284
- - `cloud.sendFeedback({ message, category?, context? })` sends feedback to the Shardflux founder by email and returns
327
+ - `cloud.sendFeedback({ message, category?, context? })` sends feedback to the Shardflux team by email and returns
285
328
  `{ id, receivedAt, duplicate }`. `category`: `bug`, `confusing`, `missing`, `idea`, `praise` or `other` (default).
286
329
  `context`: `agent`, `client`, `workspace`, `requestId`, `errorCode`, `command`, `page` (sent in snake_case);
287
330
  `client` defaults to `shardflux-sdk-ts/<SDK_VERSION>`.
@@ -316,7 +359,7 @@ Types only; nothing changes at run time and the API is unchanged.
316
359
 
317
360
  ## 0.7.0 (2026-09-28)
318
361
 
319
- Needs an API with the template editor (contracts §24); every new field is additive and older fields are unchanged.
362
+ Needs an API with the template editor; every new field is additive and older fields are unchanged.
320
363
 
321
364
  ### Template editor: build a template from template.yaml
322
365
 
@@ -337,7 +380,7 @@ Needs an API with the template editor (contracts §24); every new field is addit
337
380
  storage refused a PUT: `status`, `code` such as `BadDigest`; never the presigned URL). Exported: `packDirectory`,
338
381
  `readTemplateFile`, `parseTemplateText`, the tar writer (`tarHeader`, `tarPadding`, `tarEnd`).
339
382
 
340
- ### Template editor: the API surface (contracts §24.6)
383
+ ### Template editor: the API surface
341
384
 
342
385
  - `templates.uploads.request({ sha256, size, kind })`, `templates.uploads.put(bytes | Blob | stream, { kind, sha256?,
343
386
  size? })` (PUT with exactly the presigned headers, skipped when the organization has the bytes, confirmed after) and
@@ -406,8 +449,7 @@ docs/decisions/0006-tool-call-capture.md (shared with the Python SDK 0.3.0).
406
449
 
407
450
  ### Starts that wait for capacity end
408
451
 
409
- The API no longer lets a start (open, resume, restore, fork) wait in `capacity_pending` forever. One that no host
410
- could admit 15 minutes after it was created fails with `capacity_unavailable` and `retryable: true`: nothing was
452
+ The API no longer lets a start (open, resume, restore, fork) wait in `capacity_pending` forever. One still queued 15 minutes after it was created fails with `capacity_unavailable` and `retryable: true`: nothing was
411
453
  started, the concurrency slot is released, and a suspended workspace stays suspended with its state. Before, a VM could
412
454
  boot (and bill) long after every wait had given up.
413
455
 
@@ -460,8 +502,7 @@ boot (and bill) long after every wait had given up.
460
502
  receives the first tool token in the same response; it pre-connects to the workspace's cell meanwhile.
461
503
  - `waitForOperation()` and `templates.builds.waitForBuild()` use server-held polls (at most 20 s per request) and
462
504
  fall back to backoff (250 ms doubling to 5 s, ±20 % jitter) against a server without them.
463
- - On Node 26 the default `fetch` sends `Connection: close` (its bundled undici 8 can stall a request on a reused
464
- keep-alive connection for tens of seconds). `SHARDFLUX_HTTP_KEEPALIVE=1` or your own `fetch` changes that.
505
+ - On Node 26 the default `fetch` sends `Connection: close` . `SHARDFLUX_HTTP_KEEPALIVE=1` or your own `fetch` changes that.
465
506
 
466
507
  ### Suspended workspaces wake on use
467
508
 
package/README.md CHANGED
@@ -7,14 +7,14 @@ it later with its disk and memory intact, and fork it. Hand your agent framework
7
7
  tools (exec, files, processes, PTY, git, browser) that plug into any model provider, and save your harness's own
8
8
  tool calls into the workspace ([tool-call capture](#tool-call-capture-070)).
9
9
 
10
- > **Early access.** Shardflux is in early access. The API is versioned (`/v1`), but this SDK is
11
- > below 1.0: a minor release may contain breaking changes (see [Compatibility](#compatibility)).
10
+ > **Compatibility.** The API is versioned (`/v1`). Breaking changes ship only in minor releases and are marked
11
+ > **Breaking** in the changelog (see [Compatibility](#compatibility)).
12
12
 
13
13
  > **Versions.** This README describes 0.11.0. Anything marked **(0.11.0+)** is not in 0.10.x, **(0.10.0+)** not in 0.9.0, **(0.9.0+)** not in 0.8.x, **(0.8.0+)** not in 0.7.x,
14
14
  > **(0.7.0+)** not in 0.6.x and **(0.6.0+)** not in 0.5.0; [CHANGELOG.md](./CHANGELOG.md) lists what each version added. Check yours with
15
15
  > `npm ls @shardflux/sdk` or the exported `SDK_VERSION`.
16
16
 
17
- - ESM only, no runtime dependencies, Node.js 24 or later. Reading a YAML template file uses the optional peer
17
+ - ESM only, Node.js 24 or later. Reading a YAML template file uses the optional peer
18
18
  dependency `yaml` (`npm install yaml`); JSON template files need nothing.
19
19
  - Typed from the published OpenAPI documents.
20
20
  - Retries, idempotency keys, operation polling and tool-token refresh are handled for you.
@@ -64,8 +64,7 @@ console.log(formatTiming(again.lastTiming!)); // (0.6.0+) where the r
64
64
  ```
65
65
 
66
66
  `open()` waits until the workspace is running. Opening the same key again never resets it: files,
67
- installed packages and running processes are still there (a resume restarts processes only in the rare cold boot,
68
- see "A resume can restart processes" below). The same program is in
67
+ installed packages and running processes are still there. The same program is in
69
68
  [`examples/quickstart.ts`](./examples/quickstart.ts).
70
69
 
71
70
  With 0.5.0, wait for the suspend by its operation instead:
@@ -159,17 +158,17 @@ const { data, revision: current, servedFrom } = await cell.files.readWithInfo('/
159
158
  replaces the whole file instead; `expectedRevision: 'absent'` requires that the file does not exist yet. A changed
160
159
  file is 409 `conflict` with `reason` `revision_mismatch` and `details.current_revision`; an edit that does not match
161
160
  exactly once is 422 `edit_not_found` or `edit_ambiguous` with `details.index`.
162
- - A suspended workspace whose disk is still on a host is read, listed and searched there without waking it; such
161
+ - A suspended workspace is read, listed and searched from its saved disk without waking it; such
163
162
  results say `servedFrom: 'disk'` (`served_from` on search results). Everything else wakes it as usual. This also
164
163
  works for a handle without a tool token from before the suspend: the API issues tokens for suspended workspaces.
165
- - While the fleet is being upgraded, a workspace may run on a host that predates search and patches: they are 409
166
- `conflict` with `reason` `host_feature_unavailable` and `details.feature` (`file_search`, `file_patch`), not
167
- retryable (read and write the file, or run `grep` with `exec`, instead), and revisions are omitted.
164
+ - If search or patches are not available for a workspace, the call fails with 409 `conflict`, `reason`
165
+ `host_feature_unavailable` and `details.feature` (`file_search`, `file_patch`), not retryable: run `grep` with
166
+ `exec`, or read then write the file, instead. Revisions are omitted there.
168
167
 
169
168
  ### Wake hint (0.9.0+)
170
169
 
171
- An idle running workspace may be parked by its host (frozen or hibernated) and is restored by the next tool call.
172
- `workspace.hint()` tells the host a tool call is coming so the restore starts earlier: call it when your model starts
170
+ An idle running workspace is parked and restored by the next tool call. `workspace.hint()` says a tool call is
171
+ coming so the restore starts earlier: call it when your model starts
173
172
  emitting a tool call, before its arguments are complete. It is cheap and returns at once; a suspended workspace is
174
173
  resumed in the background (`result.wake`). The agent tools below send it when each call starts.
175
174
 
@@ -188,7 +187,7 @@ await workspace.resume({ wait: true }); // or simply open() the key again
188
187
 
189
188
  const { workspace: copy } = await workspace.fork({ key: 'customer-42/experiment' }, { wait: true });
190
189
 
191
- await copy.delete(); // tool access ends immediately; keys are never reused
190
+ await copy.delete(); // tool access ends immediately; the key can be reused after deletion finishes
192
191
  ```
193
192
 
194
193
  ### Requested or finished
@@ -200,25 +199,47 @@ means depends on `wait`:
200
199
  | --- | --- | --- |
201
200
  | `await workspace.suspend()` | the suspend is **requested** (usually `queued`; the workspace is still running) | the operation (`Operation`) |
202
201
  | `await workspace.suspend({ wait: true })` **(0.6.0+)** | the suspend has **finished** (`workspace.state` is then `suspended`) | the succeeded operation (`FinishedOperation`) |
202
+ | `await workspace.suspend({ durable: true })` **(0.12.0+)** | the suspend has finished **and** its copy is in durable storage (`result.durable` is `true`) | the succeeded operation (`FinishedOperation`) |
203
203
 
204
204
  `wait` also takes `WaitOptions` (`timeoutMs`, default 5 minutes; `signal`; `onProgress`). A failed operation throws
205
205
  `OperationFailedError`. Running out of time throws `OperationTimeoutError`, and the operation continues server side:
206
206
  wait again with `cloud.workspaces.waitForOperation(err.operationId)`. Without `wait`, the returned operation is the
207
207
  handle for the work in progress: pass its `id` to `waitForOperation()` when you need it finished.
208
208
 
209
- **A start waits for capacity for at most 15 minutes.** An open, resume, restore or fork that no host can admit yet
210
- waits in `capacity_pending`. Its `error.details.deadline_at` says when it gives up; the `phase` progress event carries
211
- it as `deadlineAt` **(0.6.2+)**, and so does `OperationTimeoutError` when your wait ends first. A start still pending
212
- at the deadline fails with `capacity_unavailable`: nothing was started, the concurrency slot is released, and a
213
- suspended workspace stays suspended with its state. `OperationFailedError.retryable` **(0.6.2+)** is `true` for it,
214
- so you can tell "retry later" from a definitive failure. The SDK does not retry it for you.
209
+ **Durable storage (0.12.0+).** A suspend returns as soon as the workspace is sealed on its host, typically in a few
210
+ hundred ms, and its RAM and CPU are released at that moment. `result.durable` turns `true` when the copy lands in
211
+ durable storage, typically within a second; until then it is `false` and `result.durability` shows the copy's
212
+ progress. Most code needs nothing more: a suspended workspace resumes, is read and is forked the same way either way.
213
+ When your code must know the copy is durable (before deleting a local artifact, or at the end of a job), ask for it:
214
+
215
+ ```ts
216
+ import { durabilityOf } from '@shardflux/sdk';
217
+
218
+ const op = await workspace.suspend({ durable: true }); // resolves once result.durable is true
219
+ durabilityOf(op); // { state: 'durable', checkpointId, localCommitAt, durableAt, localCommitToDurableMs, ... }
220
+ ```
221
+
222
+ - `durable: true` implies `wait` and shares its `timeoutMs` and `signal` (e.g. `{ durable: true, wait: { timeoutMs: 60_000 } }`).
223
+ If the time runs out first, `OperationTimeoutError` has `durable: true` and the copy continues server side.
224
+ - `cloud.workspaces.waitForDurable(operationOrId)` does the same for an operation you already hold, such as a
225
+ finished fork of a running workspace, whose result carries `durable` and `durability` the same way.
226
+ - `isDurable(op)`, `durabilityOf(op)` and `lastTiming.server.durable` / `.durability` read the fields; `formatTiming()`
227
+ prints `sealed on host, durable 435 ms later`. Results from before 0.12.0 servers count as durable.
228
+ - The errors reference on docs.shardflux.dev lists what `durable: true` can reject with.
229
+
230
+ **Start deadlines.** A queued open, resume, restore or fork (`capacity_pending`) has a deadline 15 minutes after it
231
+ was created. Its `error.details.deadline_at` carries the deadline; the `phase` progress event carries it as
232
+ `deadlineAt` **(0.6.2+)**, and so does `OperationTimeoutError` when your wait ends first. A start still pending at the
233
+ deadline fails with `capacity_unavailable`: nothing was started, the concurrency slot is released, and a suspended
234
+ workspace stays suspended with its state. `OperationFailedError.retryable` **(0.6.2+)** is `true` for it, so you can
235
+ tell a retryable failure from a definitive one.
215
236
 
216
237
  ```ts
217
238
  try {
218
239
  await workspace.resume({ wait: true });
219
240
  } catch (err) {
220
241
  if (err instanceof OperationFailedError && err.retryable) {
221
- // err.errorCode === 'capacity_unavailable': no host had room; nothing changed. Try again later.
242
+ // err.errorCode === 'capacity_unavailable': nothing changed; retry.
222
243
  } else throw err;
223
244
  }
224
245
  ```
@@ -229,7 +250,7 @@ A running workspace is billed while it is awake, and its idle policy waits a whi
229
250
  agent's turn ends, ask for a suspend once the workspace has been idle for a short time instead:
230
251
 
231
252
  ```ts
232
- const { suspendRequest } = await workspace.suspendWhenIdle({ afterSeconds: 60 }); // 30..3600
253
+ const { suspendRequest } = await workspace.suspendWhenIdle({ afterSeconds: 60 }); // 0..3600; 0 = as soon as it is idle
233
254
  workspace.suspendRequest; // { requested_at, after_seconds, not_before } until it applies or is cancelled
234
255
  await workspace.cancelSuspendWhenIdle(); // idempotent
235
256
  ```
@@ -242,7 +263,7 @@ await workspace.cancelSuspendWhenIdle(); // idempotent
242
263
  - It applies under every idle policy, `never` included, and never delays a suspend the policy would do sooner.
243
264
  - When a suspend is already in progress, the result's `operation` is that suspend and nothing is recorded.
244
265
  - Errors: `ShardfluxApiError` 409 with `reason` `not_running`, `operation_in_progress`, `session_lifetime` or
245
- `workspace_deleted`, and 422 `validation_failed` for `afterSeconds` outside 30..3600. A file-first workspace is never
266
+ `workspace_deleted`, and 422 `validation_failed` for `afterSeconds` outside 0..3600. A file-first workspace is never
246
267
  suspended: `NotSupportedForModeError` (409 `not_supported_for_mode`).
247
268
  - By id: `cloud.workspaces.suspendWhenIdle(id, { afterSeconds })` and `cloud.workspaces.cancelSuspendWhenIdle(id)`.
248
269
 
@@ -267,11 +288,11 @@ const cell = workspace.cell({ transitionTimeoutMs: 30_000 }); // give up waking
267
288
  await cell.exec.run(['make', 'test']); // resumes the workspace first if it is suspended
268
289
  ```
269
290
 
270
- **A resume can restart processes (0.11.0+).** A resume normally restores memory and running processes from the
271
- checkpoint. When the platform's VM runtime changed after the suspend and no host can restore that memory snapshot, the
272
- resume boots the checkpoint's disk instead (a cold boot), automatically, also when a tool call wakes the workspace. The
273
- workspace runs, its files are as of the suspend, but every process was restarted, as after a reboot: start dev
274
- servers, databases and background jobs again. You can tell from the resume's timing:
291
+ **Detecting a cold resume (0.11.0+).** A resume restores memory and running processes from the checkpoint. After a
292
+ platform runtime update, a resume can boot from the saved disk instead of restoring memory (a cold boot), also when a
293
+ tool call wakes the workspace; check `memoryRestored`. The workspace runs with its files as of the suspend, and its
294
+ processes start fresh, as after a reboot: start dev servers, databases and background jobs again. The resume's timing
295
+ says which it was:
275
296
 
276
297
  ```ts
277
298
  const t = workspace.lastTiming; // after resume({ wait: true }), wake(), or a tool call that woke the workspace
@@ -285,19 +306,18 @@ if (t?.server?.memoryRestored === false) {
285
306
  - `ServerTiming.coldBootReason`: why it booted (`runtime_changed`), else null. `formatTiming()` prints
286
307
  `resume from cold_boot: processes restarted (runtime_changed)`.
287
308
  - The finished resume operation carries the same in `result` (`memory_restored`, `cold_boot_reason`, and
288
- `cold_boot` with details such as `files_as_of`; its shape may change, so do not depend on it).
309
+ `cold_boot` with details such as `files_as_of`; it is informational and not part of the stable contract).
289
310
 
290
311
  `suspend`, `resume`, `fork`, `snapshot` and `delete` return the lifecycle operation.
291
312
  `cloud.workspaces.waitForOperation(id)` waits for it (default timeout 5 minutes): each poll asks the API to hold
292
313
  the response until the operation changes (`Prefer: wait`, at most 20 s per request), so completion arrives within
293
314
  one round trip of the commit. Against an API without bounded waits (or with `serverWait: false`) it polls with
294
315
  backoff (250 ms doubling to 5 s, ±20 % jitter). If the timeout passes, it throws `OperationTimeoutError` and the
295
- operation keeps running server side (a start waiting for capacity until its `deadlineAt`); wait for it again with the
316
+ operation keeps running server side (a queued start until its `deadlineAt`); wait for it again with the
296
317
  same call. `templates.builds.waitForBuild()` waits the same way. A waited `open()` issues the first tool token
297
318
  together with the final workspace read, so the first tool call starts at once.
298
319
 
299
- On Node 26 the default fetch sends `Connection: close`: its bundled undici 8 can stall a request on a reused
300
- keep-alive connection for tens of seconds. Pass your own `fetch`, or set `SHARDFLUX_HTTP_KEEPALIVE=1`, to change that.
320
+ From 0.12.0, Node 26 reuses TLS connections through a private HTTP/1.1 pool. Other runtimes keep native fetch. A supplied `fetch` stays unchanged. `SHARDFLUX_HTTP_KEEPALIVE=0` forces connection closure for diagnosis; `=1` opts into native pooling. No global dispatcher is installed.
301
321
 
302
322
  List and look up workspaces:
303
323
 
@@ -321,28 +341,31 @@ says where the time went, and `formatTiming()` prints it:
321
341
 
322
342
  ```ts
323
343
  const ws = await cloud.workspaces.open({ key: 'customer-42/main', template: 'python-node-browser' });
344
+ await ws.suspend({ wait: true });
345
+ await ws.resume({ wait: true });
324
346
  console.log(formatTiming(ws.lastTiming!));
325
347
  ```
326
348
 
327
- A slow open (34 s instead of the usual second) then reads, for example:
349
+ The resume then reads, for example:
328
350
 
329
351
  ```text
330
- open 34.18 s, succeeded (workspace 01a0e5a8-3edd-74ba-b489-d62b8925e342, operation 01a0e5a8-3ef0-7ecb-975e-dff2d5ca6e33)
331
- client: request 20.01 s (held) → capacity_pending 13.52 s (no_ready_host) → running 590 ms → view 42 ms ∥ token 61 ms
332
- server: queued 33.40 s, ran 620 ms, total 34.02 s; start warm, boot to ready 79 ms
333
- outside the server: 161 ms
352
+ resume 413 ms, succeeded (workspace 01a0ead6-92e5-7d19-ab7a-f565bf4676cb, operation 01a0ead6-bd61-737a-ae88-66e2d01a6e25)
353
+ client: request 218 ms → queued 195 ms
354
+ server: queued 51 ms, ran 290 ms, total 341 ms; resume from local_cache, boot to ready 211 ms, host disk 0 ms, load 6 ms, ready 62 ms, total 211 ms, after_restore 36 ms
355
+ outside the server: 72 ms
334
356
  ```
335
357
 
336
- Here the time went to waiting for a host with capacity; the network and the VM start were fast.
358
+ This resume brought the workspace back from its host's local cache, memory and running processes included, in 341 ms on the server and 413 ms end to end.
337
359
 
338
360
  - **client** phases are on your monotonic clock: the request (a held open waits on the server, `held`), then each
339
361
  operation state the SDK observed while waiting, with the server's reason (`capacity_pending` / `no_ready_host`:
340
- waiting for a host; `running` / `template_downloading`: fetching the template), then reading the workspace and
362
+ queued to start; `running` / `template_downloading`: fetching the template), then reading the workspace and
341
363
  issuing the first tool token (together).
342
364
  - **server** timing comes from the operation itself (one database clock): `queued` is creation until it began
343
- running, including any wait for capacity; `ran` is the cell's work (placement, boot or restore, guest readiness).
344
- `start` / `resume from` and `boot to ready` / `host …` are what the cell reported. A resume that booted the saved
345
- disk instead of restoring memory reads `resume from cold_boot: processes restarted (runtime_changed)` (0.11.0+).
365
+ running, including any time in `capacity_pending`; `ran` is the cell's work (placement, boot or restore, guest
366
+ readiness). `start` / `resume from` and `boot to ready` / `host …` are what the cell reported. A resume that booted
367
+ the saved disk instead of restoring memory reads `resume from cold_boot: processes restarted (runtime_changed)`
368
+ (0.11.0+).
346
369
  - **outside the server** is your total minus the operation's: network, TLS, polling latency, view and token. A large
347
370
  value with a small server total points at the connection between you and the API, not at the workspace.
348
371
  - **retries** lists transient failures the SDK retried (cause and backoff).
@@ -426,13 +449,13 @@ await cell.files.write('/home/user/app/main.py', 'print("bye")\n', { ifTreeRevis
426
449
  ```
427
450
 
428
451
  - **Executions are idempotent by id.** `executionId` defaults to a fresh `ex-<uuid>`; pass your own to make a call safe
429
- to repeat across processes. Network failures and retryable 5xx answers (503 `no_execution_host`: no host has room;
430
- the SDK waits `Retry-After`) are retried with the same id, at most `maxRetries` (5) times, and an answer that takes
431
- longer than `attemptTimeoutMs` (300 s) is awaited again with the same id. The cell runs a command once per id: a
452
+ to repeat across processes. Network failures and retryable 5xx answers (503 `no_execution_host`: the execution
453
+ cannot be placed right now; the SDK waits `Retry-After`) are retried with the same id, at most `maxRetries` (5)
454
+ times, and an answer that takes longer than `attemptTimeoutMs` (300 s) is awaited again with the same id. The cell runs a command once per id: a
432
455
  repeated call gets the recorded result (`r.replayed`). The SDK never retries with a new id: a `failed` or `lost`
433
456
  result is returned, and running it again is your decision.
434
457
  - `ws.executions.get(id, { waitMs })` reads an execution: `pending` while it runs, its result when it ended (kept 7
435
- days). Use it after a call that stopped waiting (an aborted `signal`): an execution cannot be canceled.
458
+ days). Use it after a call that stopped waiting (an aborted `signal`): an execution runs to completion.
436
459
  - While an execution runs, other executions and file writes are refused with 409 `workspace_busy`
437
460
  (`execution_in_progress`); the SDK waits them out within `transitionTimeoutMs` (120 s).
438
461
  - Only paths under /home/user exist (`outside_tree_root` otherwise). Reads, writes, `search()` and `patch()` work as on
@@ -440,7 +463,7 @@ await cell.files.write('/home/user/app/main.py', 'print("bye")\n', { ifTreeRevis
440
463
  - Calls a file-first workspace does not have fail with `NotSupportedForModeError` before any request: exec sessions
441
464
  (`cell.exec.*`), PTY, processes, git, browser, `changes()`, `suspend`, `resume`, `snapshot`, `fork`, `reset` and
442
465
  `saveAsTemplate`. On a processful workspace `executions` and `ifTreeRevision` are refused the same way.
443
- - Reopening a key with another `mode` is 409 `mode_mismatch`; a deployment without file-first workspaces answers 422
466
+ - Reopening a key with another `mode` is 409 `mode_mismatch`; an account without file-first workspaces gets 422
444
467
  `mode_not_available`; a legacy template 409 `layout_unsupported`.
445
468
 
446
469
  ## Templates: file tree, diff and dev mode
@@ -689,10 +712,9 @@ pending: past that, while the workspace takes no writes, a call is not recorded
689
712
  `onError` gets `queue_full`). A write is retried for `retryWindowMs` (120 s) and then dropped (`write_failed`).
690
713
  Observing an async tool marks its promise handled: if your code never awaits a wrapped tool's promise and it rejects,
691
714
  Node reports no `unhandledRejection` for it (the error is still in the index; Node has no way to observe a rejection
692
- without handling it). About one write per call: more than roughly 100 calls per second per
693
- workspace reaches the pending limit. `Response`, `ReadableStream`, Node streams and `Blob` results are never read
694
- (`meta.note: "stream_not_captured"`). On template v1 the files API writes as root, so the agent can read the captured
695
- files but not change them. The format is specified in `docs/decisions/0006-tool-call-capture.md`.
715
+ without handling it). About one write per call: the default limits handle about 100 calls per second per workspace;
716
+ raise the pending limits for more. `Response`, `ReadableStream`, Node streams and `Blob` results are never read
717
+ (`meta.note: "stream_not_captured"`). On template v1, captured files are read-only to the agent.
696
718
 
697
719
  ## Secrets
698
720
 
@@ -803,10 +825,9 @@ changes.
803
825
  (`details.reason`, e.g. `not_session`, `draft_not_found`, `legacy_disk_layout`; see `KnownErrorReason`).
804
826
  - `OperationFailedError`: an awaited operation ended `failed` or `canceled` (`errorCode`,
805
827
  `retryable` **(0.6.2+)**, `operation`, `timing`). `retryable` is the operation error's own flag: `true` for
806
- `capacity_unavailable` (no host could admit the start before its deadline; retry later), `false` for a definitive
807
- failure.
828
+ `capacity_unavailable` (the start passed its deadline; retry it), `false` for a definitive failure.
808
829
  - `OperationTimeoutError`: waiting gave up; the operation continues (`operationId`, `lastState`, `lastReason`,
809
- `deadlineAt` **(0.6.2+)** while it waits for capacity, `timing`).
830
+ `deadlineAt` **(0.6.2+)** while a start is queued, `timing`).
810
831
  - `ShardfluxProtocolError`: a response was not the documented shape.
811
832
  - `NotSupportedForModeError` **(0.9.0+)**, a `ShardfluxApiError` (409 `conflict`, `reason` `not_supported_for_mode`):
812
833
  the call does not exist for the workspace's `mode`; `local` is true when the SDK refused it without a request.
@@ -829,14 +850,14 @@ carries `reason` **(0.9.0+)**:
829
850
  `details.spend_cap` is `{ cap_minor, effective_cap_minor, charges_minor, currency }` (minor units, `null` when the plan
830
851
  has no overage). An older API sends no `reason`. Do not retry these in a loop.
831
852
 
832
- Retryable 429/502/503/504 refusals (for example 503 `host_capacity`, when the workspace's host has no room to restore
833
- it right now, or `wake_failed`) are retried after `Retry-After` for reads, searches and calls that carry an
853
+ Retryable 429/502/503/504 refusals (for example 503 `host_capacity` when the workspace cannot be woken right now, or
854
+ `wake_failed`) are retried after `Retry-After` for reads, searches and calls that carry an
834
855
  Idempotency-Key (writes and patches); other calls surface them with `retryable: true` and `retryAfterSeconds`.
835
856
  A read of a sleeping workspace that its disk cannot answer (409 `workspace_not_running` with `reason`
836
857
  `offline_unavailable` or `offline_budget`) wakes the workspace and is retried like any `workspace_not_running`; 503
837
858
  `offline_changed` (the disk changed during the read) is retried and served by the running workspace. 409 `conflict`
838
- `host_feature_unavailable` (the workspace's host predates the call, `details.feature`) is neither retried nor
839
- woken: it lasts until the workspace runs on an upgraded host.
859
+ `host_feature_unavailable` (search or patches are not available for the workspace, `details.feature`) is neither
860
+ retried nor woken: run `grep` with `exec`, or read then write the file, instead.
840
861
 
841
862
  ## Usage and overage
842
863
 
@@ -870,10 +891,10 @@ is charged on the next invoice until the charges reach the spend cap.
870
891
 
871
892
  ## Feedback (0.9.0+)
872
893
 
873
- `cloud.sendFeedback()` sends a message straight to the Shardflux founder, who reads every one. If you or your coding
894
+ `cloud.sendFeedback()` sends a message straight to the Shardflux team, who read every one. If you or your coding
874
895
  agent hit something while building with Shardflux, send it the moment it happens: a call that failed unexpectedly, an
875
- error or doc that was confusing, something missing or slow, a workaround you needed. Short and specific beats polished;
876
- the request id and error code let the founder find the logs.
896
+ error or doc that was confusing, something missing, a workaround you needed. Short and specific beats polished;
897
+ the request id and error code let the team find the logs.
877
898
 
878
899
  ```ts
879
900
  try {
@@ -881,7 +902,7 @@ try {
881
902
  } catch (err) {
882
903
  if (err instanceof ShardfluxApiError) {
883
904
  await cloud.sendFeedback({
884
- message: 'open failed with capacity_unavailable twice in 10 minutes; expected a start within a minute',
905
+ message: 'open rejected template "python-node" with template_not_found; expected a suggestion of the closest slug',
885
906
  category: 'bug',
886
907
  context: { requestId: err.requestId, errorCode: err.code, workspace: 'acme/demo', agent: 'claude-code' },
887
908
  });
@@ -933,7 +954,7 @@ or no longer supported, it emits one warning:
933
954
 
934
955
  - The SDK follows the API's `/v1` contract. New fields, enum values and error codes can appear in
935
956
  any release; ignore unknown fields.
936
- - While below 1.0, a breaking change bumps the minor version (0.7 to 0.8).
957
+ - Breaking changes ship only in minor releases (0.7 to 0.8) and are marked **Breaking** in the changelog.
937
958
  - `SDK_VERSION` is exported; requests send `User-Agent: shardflux-sdk-ts/<version>`.
938
959
  - Examples in this README, in `examples/` and on shardflux.dev name the version they need. The examples on the
939
960
  website and in the console are checked against the version published on npm before they ship.
@@ -941,3 +962,44 @@ or no longer supported, it emits one warning:
941
962
  ## License
942
963
 
943
964
  Apache-2.0
965
+
966
+
967
+ ## Workspace integrations (0.12.0+)
968
+
969
+ ```ts
970
+ const ws = await cloud.workspaces.open({ key: 'qm/project-42', template: 'python-node-browser',
971
+ labels: { scope: 'project-42', owner: 'qm' }, idlePolicy: 'never' });
972
+ const matches = await cloud.workspaces.list({ labels: { scope: 'project-42' } });
973
+ await ws.setLabels({ scope: 'project-42', owner: 'qm' }); // replaces all labels; {} clears
974
+ await ws.setIdlePolicy('adaptive'); // null restores the inherited policy
975
+ const idle = await ws.idle(); // read-only, does not wake or record activity
976
+ await ws.keepalive(60); // seconds; never shortens an existing keepalive
977
+ await ws.suspendWhenIdle({ afterSeconds: 2 }); // one safe server-side request; no /idle + suspend race
978
+ ```
979
+
980
+ Labels are exact string pairs (at most 50, keys 1–64 and values 0–256 characters); they are metadata, not secrets.
981
+ A failed workspace is recoverable: `await ws.wake()` and ordinary auto-waking calls restart that ID using its existing disk. They do not create a replacement workspace. A failed attempt still raises its operation error; there is no unbounded restart loop.
982
+ After `await ws.delete({ wait: true })`, opening its key creates a **new ID**. The old ID, tokens and history remain deleted. A key stays reserved while deletion is in progress.
983
+
984
+ For a background command with a pipe (no PTY):
985
+
986
+ ```ts
987
+ const cell = ws.cell();
988
+ const session = await cell.exec.start({ session_id: 'qm-worker-1', argv: ['cat'], stdin_open: true });
989
+ let ack = await cell.exec.input(session.session_id, 'hello\n', { offset: 0 });
990
+ ack = await cell.exec.input(session.session_id, '', { offset: ack.offset, close: true });
991
+ ```
992
+
993
+ Input frames are at most 64 KiB. The acknowledged offset counts bytes accepted into the pipe, not bytes consumed by the program. A blocked writer may receive a partial acknowledgement; continue from the returned offset. Retrying the identical last frame is safe. `close` sends EOF only once the entire frame is accepted. Do not combine `stdin_open` with the existing one-shot `stdin` option. The transport writes no separate input log or payload file; program output and full-state snapshots retain their normal persistence. Full-state suspend preserves the pipe. Older hosts/guests explicitly refuse this feature until upgraded (`exec_stdin` / `exec_stdin.v1`). Running sessions count as work within the documented idle-command window; long-running services should declare keepalives or choose `never`.
994
+
995
+ ```ts
996
+ import { isWorkspaceGone, ShardfluxProtocolError } from '@shardflux/sdk';
997
+ // isWorkspaceGone(error) is true only for an explicit API workspace_deleted refusal.
998
+ // A protocol error, a cell 404, a scoped API 404 or a network failure is never deletion evidence.
999
+ ```
1000
+
1001
+ `ShardfluxProtocolError.source` is `api` or `cell` for SDK responses (`unknown` only for an error constructed without a source by older caller code). It remains a protocol failure even when `status === 404`. Never delete local data based on an HTTP status alone.
1002
+
1003
+ ### Repositories with a minimum release age
1004
+
1005
+ Keep your repository's age policy. Pin an exact version that has aged enough; do not exempt the entire `@shardflux/*` scope. The first SDK, 0.5.0, was published September 26, 2026 at 21:13:52 UTC and reaches seven days on October 3 at that time. New versions, including 0.12.0, need their own seven days after publication. Before then no registry setting on our side can make them eligible. `npm view @shardflux/sdk time --json` shows publication times. A targeted exception for an inspected exact release is a repository-owner decision, not an installation requirement we bypass. Older versions do not contain this release's fixes.
package/dist/account.d.ts CHANGED
@@ -1,5 +1,5 @@
1
1
  /**
2
- * The account plane with a user session (0.9.0; contracts §30): what a person does in the web app, from code.
2
+ * The account plane with a user session (0.9.0): what a person does in the web app, from code.
3
3
  *
4
4
  * // Before a session exists (no token needed)
5
5
  * await ShardfluxAccount.register({ email, password, displayName });
@@ -102,7 +102,7 @@ export interface ShardfluxAccountOptions {
102
102
  sessionToken?: string;
103
103
  /** Default https://api.shardflux.dev. */
104
104
  baseUrl?: string;
105
- /** Default: the runtime's fetch, with `Connection: close` on Node 26 (see defaultFetch in http.ts). */
105
+ /** Default: pooled HTTP/1.1 on Node 26+, native fetch on other runtimes (see defaultFetch in http.ts). */
106
106
  fetch?: typeof fetch;
107
107
  userAgent?: string;
108
108
  /** Per-request timeout (ms), default 30 s. */
@@ -451,7 +451,7 @@ export declare class ShardfluxAccount {
451
451
  /** The signed-in user and their memberships. */
452
452
  me(): Promise<Me>;
453
453
  /**
454
- * Send feedback straight to the Shardflux founder as the signed-in user (0.9.0+; POST /v1/feedback), optionally about
454
+ * Send feedback straight to the Shardflux team as the signed-in user (0.9.0+; POST /v1/feedback), optionally about
455
455
  * one of your organizations (`organizationId`). Otherwise the same as `Shardflux.sendFeedback`: use it while you work,
456
456
  * returns `{ id, receivedAt, duplicate }`, never retried, 429 `rate_limited` per user.
457
457
  */
package/dist/account.js CHANGED
@@ -571,7 +571,7 @@ export class ShardfluxAccount {
571
571
  const b = body;
572
572
  if (b && typeof b === 'object' && typeof b.session_token === 'string') {
573
573
  if (!isSessionToken(b.session_token))
574
- throw new ShardfluxProtocolError('the API returned a session_token that is not a CLI session token (sfu_...)', 200);
574
+ throw new ShardfluxProtocolError('the API returned a session_token that is not a CLI session token (sfu_...)', 200, 'api');
575
575
  this.#token = b.session_token;
576
576
  const expiresAt = typeof b.session_expires_at === 'string' ? b.session_expires_at : '';
577
577
  await this.#onSessionToken?.({ token: b.session_token, expiresAt });
@@ -583,7 +583,7 @@ export class ShardfluxAccount {
583
583
  return this.#ctx.http.json('GET', '/v1/me', {}, this.#ctx.authorization);
584
584
  }
585
585
  /**
586
- * Send feedback straight to the Shardflux founder as the signed-in user (0.9.0+; POST /v1/feedback), optionally about
586
+ * Send feedback straight to the Shardflux team as the signed-in user (0.9.0+; POST /v1/feedback), optionally about
587
587
  * one of your organizations (`organizationId`). Otherwise the same as `Shardflux.sendFeedback`: use it while you work,
588
588
  * returns `{ id, receivedAt, duplicate }`, never retried, 429 `rate_limited` per user.
589
589
  */