@shardflux/sdk 0.10.2 → 0.11.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,9 +1,46 @@
1
1
  # Changelog
2
2
 
3
- Every API the README shows is available from the version named here. Below 1.0, a minor release may break
4
- compatibility; breaking changes are marked **Breaking**.
5
-
6
- ## 0.10.2 (not yet published)
3
+ Every API the README shows is available from the version named here. Breaking changes ship in minor releases and are
4
+ marked **Breaking**.
5
+
6
+ ## 0.11.1
7
+
8
+ Wording: messages and JSDoc say what to do, without internals (no behaviour change).
9
+
10
+ - `OperationTimeoutError` on a queued start: `... it stays queued server side until <deadline_at> and fails with
11
+ capacity_unavailable if it has not started by then.` (the same words as the Python SDK).
12
+ - JSDoc (IDE hovers): `sendFeedback` goes to the Shardflux team; `capacity_unavailable` is retryable: the start passed
13
+ its deadline and nothing was started, send it again; `no_execution_host`: the execution cannot be placed right now;
14
+ `host_feature_unavailable`: the call is not available for this workspace, use the fallback in the hint;
15
+ `resumePath` `cold_boot`: the saved disk was booted after a platform runtime change, processes restarted; an
16
+ execution runs to completion; on Node 26 every request uses a fresh connection (`Connection: close`) unless
17
+ `SHARDFLUX_HTTP_KEEPALIVE=1`.
18
+ - The `formatTiming()` example is a production resume (413 ms).
19
+
20
+ ## 0.11.0
21
+
22
+ ### A resume that restarted processes says so (cold boot)
23
+
24
+ Additive. After a platform runtime change, the
25
+ cell resumes the workspace by booting its checkpoint's disk (a cold boot), automatically, also when a tool call wakes
26
+ the workspace. The workspace runs and its files are as of the suspend, but every process was restarted. The API passes
27
+ the resume's result through; this release reads it.
28
+
29
+ - `ServerTiming.memoryRestored` (`boolean | null`): `result.memory_restored`. `true` when memory and running
30
+ processes came back; `false` for `cold_boot` (and `reset_blank_layer`); `null` when the result does not say (an API
31
+ from before the field), never read as `false`.
32
+ - `ServerTiming.coldBootReason` (`string | null`): `result.cold_boot_reason` with `resumePath` `cold_boot`, today
33
+ `runtime_changed`.
34
+ - `ServerTiming.resumePath` documents `cold_boot` and `reset_blank_layer`.
35
+ - `formatTiming()` prints `resume from cold_boot: processes restarted (runtime_changed)` when memory was not restored;
36
+ other resumes print as before.
37
+ - Where it shows: `workspace.lastTiming` and the `done` progress event after `resume({ wait: true })`, `wake()` and a
38
+ tool call that woke the workspace; the finished resume operation's `result` (`memory_restored`,
39
+ `cold_boot_reason`, and `cold_boot` with operator details such as `files_as_of`, whose shape may change).
40
+ - The new `ServerTiming` fields are optional in the type (always set by the SDK), so timings built by hand still
41
+ type-check. No return type or behaviour changes: a wake that cold-booted still resolves `true`.
42
+
43
+ ## 0.10.2
7
44
 
8
45
  Fix: the SDK reports 0.10.2. 0.10.1 was published reporting 0.10.0, its previous version; the build now
9
46
  fails when `SDK_VERSION` differs from package.json.
@@ -57,7 +94,7 @@ from the API's OpenAPI); the account client's `setSpendPolicy()` gains the overa
57
94
  `KnownErrorReason` adds them, the spend-policy refusals (422 `overage_unavailable`, `spend_cap_required`,
58
95
  `spend_cap_below_minimum`, `spend_cap_above_plan_price`, `spend_cap_below_charges`) and 409 `version_mismatch`.
59
96
 
60
- ### Suspend when idle (contracts §20.6)
97
+ ### Suspend when idle
61
98
 
62
99
  - `workspace.suspendWhenIdle({ afterSeconds, idempotencyKey? })` and `cloud.workspaces.suspendWhenIdle(id, {
63
100
  afterSeconds })` (POST /v1/workspaces/{id}/suspend-when-idle): the workspace is suspended once it has been idle for
@@ -71,15 +108,14 @@ from the API's OpenAPI); the account client's `setSpendPolicy()` gains the overa
71
108
  count as the next turn and cancel the request.
72
109
  - Types `SuspendRequest`, `SuspendWhenIdleOptions`, `SuspendWhenIdleResult`, `SuspendWhenIdleResponse`.
73
110
 
74
- ## 0.9.0 (not yet published; npm `latest` is 0.8.0)
111
+ ## 0.9.0
75
112
 
76
113
  Elastic compute (decision 0007): file tools and wake hints for parked workspaces. Additive; older APIs and cell
77
114
  gateways keep working (the new calls answer 404 there).
78
115
 
79
116
  ### File search, patches with revisions, the wake hint
80
117
 
81
- Needs a cell gateway with the contracts §26 routes (`files/search`, `files/patch`, `wake-hint`, revisions); an older
82
- gateway does not serve them and returns no revisions.
118
+ Uses the `files/search`, `files/patch` and `wake-hint` routes and file revisions.
83
119
 
84
120
  - `cell.files.search(path, pattern, opts)`: content search under a directory (literal or RE2 with `regex`,
85
121
  `caseInsensitive`, `include`/`exclude` globs, `maxMatches`, `maxFileBytes`, `contextLines`). Returns the gateway's
@@ -114,24 +150,22 @@ gateway does not serve them and returns no revisions.
114
150
  of a sleeping workspace `offline_unavailable` and `offline_budget` (409 `workspace_not_running`: woken and retried
115
151
  like any) and `offline_changed` (503, retried; served by the running workspace).
116
152
  - Reads of a suspended workspace without a held token: the API now issues tool tokens for a suspended workspace
117
- (contracts §26.4), so `read`, `readText`, `readWithInfo`, `stat`, `list` and `search` of a suspended workspace are
153
+ , so `read`, `readText`, `readWithInfo`, `stat`, `list` and `search` of a suspended workspace are
118
154
  served from its disk without waking it also for a handle that fetches its first token after the suspend (before,
119
155
  the token was refused and the call woke the workspace). Any other call still wakes it: the cell refuses it with 409
120
156
  `workspace_not_running`. Needs that API; with an older one the token is refused and the call wakes it as before.
121
157
  - The agent-tool runner sends no wake hint for `read_file`, `list_files` and `search_files`: a sleeping workspace
122
158
  serves them from its disk, and the hint would wake a suspended one (or restore a hibernated one) for nothing.
123
- - Hosts without the features (contracts §26.7): `KnownErrorReason` adds `host_feature_unavailable` (409 `conflict`,
124
- not retryable, `details.feature` `file_search` or `file_patch`): the workspace runs on a host agent that predates
125
- the call, until it runs on an upgraded host. It is neither retried nor answered with a wake. Such a host also
126
- returns no revisions (`FileInfo.revision`, `readWithInfo().revision` are absent).
159
+ - Workspaces without the features: `KnownErrorReason` adds `host_feature_unavailable` (409 `conflict`,
160
+ not retryable, `details.feature` `file_search` or `file_patch`): the call is not available for the workspace. It is neither retried nor answered with a wake. Such a workspace also returns no revisions (`FileInfo.revision`, `readWithInfo().revision` are absent).
127
161
  - `JsonSchema.pattern`, checked by `validateArgs()`.
128
162
  - Types: `FileSearchOptions`, `FileSearchResponse`, `FileSearchRequest`, `FileSearchResult`, `FileSearchMatch`,
129
163
  `FilePatchParams`, `FilePatchEdit`, `FilePatchRequest`, `FilePatchResult`, `FileEdit`, `FileRevision`,
130
164
  `FileReadResult`, `ServedFrom`, `Residency`, `WakeHintResult`, `HintOptions`, `HintResult`.
131
165
 
132
- ### File-first workspaces (contracts §29)
166
+ ### File-first workspaces
133
167
 
134
- Needs an API with `FILE_FIRST_WORKSPACES` on and a cell that serves file-first workspaces (contracts §29.8). A
168
+ Needs an API with `FILE_FIRST_WORKSPACES` on and a cell that serves file-first workspaces. A
135
169
  file-first workspace has no VM between executions: its state is a versioned file tree under /home/user, and each
136
170
  command runs in a fresh VM whose changed files become the next tree revision. Processful workspaces are unchanged.
137
171
 
@@ -180,7 +214,7 @@ command runs in a fresh VM whose changed files become the next tree revision. Pr
180
214
  - Types: `WorkspaceMode`, `ExecutionResult`, `ExecutionResultBody`, `ExecutionChange`, `ExecutionState`,
181
215
  `ExecutionError`, `ExecutionRunOptions`, `ExecutionGetOptions`, `TreeRevisionOptions`.
182
216
 
183
- ### Waking a suspended workspace is one request (contracts §22.6)
217
+ ### Waking a suspended workspace is one request
184
218
 
185
219
  Needs an API with the held resume for the single request; against an older API every call below works as in 0.7.0.
186
220
 
@@ -199,9 +233,9 @@ Needs an API with the held resume for the single request; against an older API e
199
233
  choose the token that comes back; `cell()`'s own wake passes its label and tools. `WorkspacesApi.requestResume()`
200
234
  is the bare request (`ResumeAnswer`, `ResumeRequestOptions`, `ResumeResponse`).
201
235
  - `CellClient`: after a wake that put a new token into the client's manager, the retry uses it instead of fetching one.
202
- ### The account plane and the version check (contracts §30)
236
+ ### The account plane and the version check
203
237
 
204
- Needs an API with CLI sessions and client versions (contracts §30). Every API is additive; the one change in
238
+ Needs an API with CLI sessions and client versions. Every API is additive; the one change in
205
239
  behavior is the automatic version check (below), which makes one background request per process.
206
240
 
207
241
  ### Account plane: `ShardfluxAccount`
@@ -256,9 +290,9 @@ behavior is the automatic version check (below), which makes one background requ
256
290
  Off with `SHARDFLUX_NO_UPDATE_CHECK=1` (also `true`, `yes`, `on`) or `NO_UPDATE_NOTIFIER=1`. `fetchBillingCatalog()`
257
291
  does not check.
258
292
 
259
- ### Feedback straight to the founder (POST /v1/feedback)
293
+ ### Feedback straight to the Shardflux team (POST /v1/feedback)
260
294
 
261
- - `cloud.sendFeedback({ message, category?, context? })` sends feedback to the Shardflux founder by email and returns
295
+ - `cloud.sendFeedback({ message, category?, context? })` sends feedback to the Shardflux team by email and returns
262
296
  `{ id, receivedAt, duplicate }`. `category`: `bug`, `confusing`, `missing`, `idea`, `praise` or `other` (default).
263
297
  `context`: `agent`, `client`, `workspace`, `requestId`, `errorCode`, `command`, `page` (sent in snake_case);
264
298
  `client` defaults to `shardflux-sdk-ts/<SDK_VERSION>`.
@@ -293,7 +327,7 @@ Types only; nothing changes at run time and the API is unchanged.
293
327
 
294
328
  ## 0.7.0 (2026-09-28)
295
329
 
296
- Needs an API with the template editor (contracts §24); every new field is additive and older fields are unchanged.
330
+ Needs an API with the template editor; every new field is additive and older fields are unchanged.
297
331
 
298
332
  ### Template editor: build a template from template.yaml
299
333
 
@@ -314,7 +348,7 @@ Needs an API with the template editor (contracts §24); every new field is addit
314
348
  storage refused a PUT: `status`, `code` such as `BadDigest`; never the presigned URL). Exported: `packDirectory`,
315
349
  `readTemplateFile`, `parseTemplateText`, the tar writer (`tarHeader`, `tarPadding`, `tarEnd`).
316
350
 
317
- ### Template editor: the API surface (contracts §24.6)
351
+ ### Template editor: the API surface
318
352
 
319
353
  - `templates.uploads.request({ sha256, size, kind })`, `templates.uploads.put(bytes | Blob | stream, { kind, sha256?,
320
354
  size? })` (PUT with exactly the presigned headers, skipped when the organization has the bytes, confirmed after) and
@@ -383,8 +417,7 @@ docs/decisions/0006-tool-call-capture.md (shared with the Python SDK 0.3.0).
383
417
 
384
418
  ### Starts that wait for capacity end
385
419
 
386
- The API no longer lets a start (open, resume, restore, fork) wait in `capacity_pending` forever. One that no host
387
- could admit 15 minutes after it was created fails with `capacity_unavailable` and `retryable: true`: nothing was
420
+ The API no longer lets a start (open, resume, restore, fork) wait in `capacity_pending` forever. One still queued 15 minutes after it was created fails with `capacity_unavailable` and `retryable: true`: nothing was
388
421
  started, the concurrency slot is released, and a suspended workspace stays suspended with its state. Before, a VM could
389
422
  boot (and bill) long after every wait had given up.
390
423
 
@@ -437,8 +470,7 @@ boot (and bill) long after every wait had given up.
437
470
  receives the first tool token in the same response; it pre-connects to the workspace's cell meanwhile.
438
471
  - `waitForOperation()` and `templates.builds.waitForBuild()` use server-held polls (at most 20 s per request) and
439
472
  fall back to backoff (250 ms doubling to 5 s, ±20 % jitter) against a server without them.
440
- - On Node 26 the default `fetch` sends `Connection: close` (its bundled undici 8 can stall a request on a reused
441
- keep-alive connection for tens of seconds). `SHARDFLUX_HTTP_KEEPALIVE=1` or your own `fetch` changes that.
473
+ - On Node 26 the default `fetch` sends `Connection: close` . `SHARDFLUX_HTTP_KEEPALIVE=1` or your own `fetch` changes that.
442
474
 
443
475
  ### Suspended workspaces wake on use
444
476
 
package/README.md CHANGED
@@ -7,10 +7,10 @@ it later with its disk and memory intact, and fork it. Hand your agent framework
7
7
  tools (exec, files, processes, PTY, git, browser) that plug into any model provider, and save your harness's own
8
8
  tool calls into the workspace ([tool-call capture](#tool-call-capture-070)).
9
9
 
10
- > **Early access.** Shardflux is in early access. The API is versioned (`/v1`), but this SDK is
11
- > below 1.0: a minor release may contain breaking changes (see [Compatibility](#compatibility)).
10
+ > **Compatibility.** The API is versioned (`/v1`). Breaking changes ship only in minor releases and are marked
11
+ > **Breaking** in the changelog (see [Compatibility](#compatibility)).
12
12
 
13
- > **Versions.** This README describes 0.10.2. Anything marked **(0.10.0+)** is not in 0.9.0, **(0.9.0+)** not in 0.8.x, **(0.8.0+)** not in 0.7.x,
13
+ > **Versions.** This README describes 0.11.0. Anything marked **(0.11.0+)** is not in 0.10.x, **(0.10.0+)** not in 0.9.0, **(0.9.0+)** not in 0.8.x, **(0.8.0+)** not in 0.7.x,
14
14
  > **(0.7.0+)** not in 0.6.x and **(0.6.0+)** not in 0.5.0; [CHANGELOG.md](./CHANGELOG.md) lists what each version added. Check yours with
15
15
  > `npm ls @shardflux/sdk` or the exported `SDK_VERSION`.
16
16
 
@@ -158,17 +158,17 @@ const { data, revision: current, servedFrom } = await cell.files.readWithInfo('/
158
158
  replaces the whole file instead; `expectedRevision: 'absent'` requires that the file does not exist yet. A changed
159
159
  file is 409 `conflict` with `reason` `revision_mismatch` and `details.current_revision`; an edit that does not match
160
160
  exactly once is 422 `edit_not_found` or `edit_ambiguous` with `details.index`.
161
- - A suspended workspace whose disk is still on a host is read, listed and searched there without waking it; such
161
+ - A suspended workspace is read, listed and searched from its saved disk without waking it; such
162
162
  results say `servedFrom: 'disk'` (`served_from` on search results). Everything else wakes it as usual. This also
163
163
  works for a handle without a tool token from before the suspend: the API issues tokens for suspended workspaces.
164
- - While the fleet is being upgraded, a workspace may run on a host that predates search and patches: they are 409
165
- `conflict` with `reason` `host_feature_unavailable` and `details.feature` (`file_search`, `file_patch`), not
166
- retryable (read and write the file, or run `grep` with `exec`, instead), and revisions are omitted.
164
+ - If search or patches are not available for a workspace, the call fails with 409 `conflict`, `reason`
165
+ `host_feature_unavailable` and `details.feature` (`file_search`, `file_patch`), not retryable: run `grep` with
166
+ `exec`, or read then write the file, instead. Revisions are omitted there.
167
167
 
168
168
  ### Wake hint (0.9.0+)
169
169
 
170
- An idle running workspace may be parked by its host (frozen or hibernated) and is restored by the next tool call.
171
- `workspace.hint()` tells the host a tool call is coming so the restore starts earlier: call it when your model starts
170
+ An idle running workspace is parked and restored by the next tool call. `workspace.hint()` says a tool call is
171
+ coming so the restore starts earlier: call it when your model starts
172
172
  emitting a tool call, before its arguments are complete. It is cheap and returns at once; a suspended workspace is
173
173
  resumed in the background (`result.wake`). The agent tools below send it when each call starts.
174
174
 
@@ -205,19 +205,19 @@ means depends on `wait`:
205
205
  wait again with `cloud.workspaces.waitForOperation(err.operationId)`. Without `wait`, the returned operation is the
206
206
  handle for the work in progress: pass its `id` to `waitForOperation()` when you need it finished.
207
207
 
208
- **A start waits for capacity for at most 15 minutes.** An open, resume, restore or fork that no host can admit yet
209
- waits in `capacity_pending`. Its `error.details.deadline_at` says when it gives up; the `phase` progress event carries
210
- it as `deadlineAt` **(0.6.2+)**, and so does `OperationTimeoutError` when your wait ends first. A start still pending
211
- at the deadline fails with `capacity_unavailable`: nothing was started, the concurrency slot is released, and a
212
- suspended workspace stays suspended with its state. `OperationFailedError.retryable` **(0.6.2+)** is `true` for it,
213
- so you can tell "retry later" from a definitive failure. The SDK does not retry it for you.
208
+ **Start deadlines.** A queued open, resume, restore or fork (`capacity_pending`) has a deadline 15 minutes after it
209
+ was created. Its `error.details.deadline_at` carries the deadline; the `phase` progress event carries it as
210
+ `deadlineAt` **(0.6.2+)**, and so does `OperationTimeoutError` when your wait ends first. A start still pending at the
211
+ deadline fails with `capacity_unavailable`: nothing was started, the concurrency slot is released, and a suspended
212
+ workspace stays suspended with its state. `OperationFailedError.retryable` **(0.6.2+)** is `true` for it, so you can
213
+ tell a retryable failure from a definitive one.
214
214
 
215
215
  ```ts
216
216
  try {
217
217
  await workspace.resume({ wait: true });
218
218
  } catch (err) {
219
219
  if (err instanceof OperationFailedError && err.retryable) {
220
- // err.errorCode === 'capacity_unavailable': no host had room; nothing changed. Try again later.
220
+ // err.errorCode === 'capacity_unavailable': nothing changed; retry.
221
221
  } else throw err;
222
222
  }
223
223
  ```
@@ -228,7 +228,7 @@ A running workspace is billed while it is awake, and its idle policy waits a whi
228
228
  agent's turn ends, ask for a suspend once the workspace has been idle for a short time instead:
229
229
 
230
230
  ```ts
231
- const { suspendRequest } = await workspace.suspendWhenIdle({ afterSeconds: 60 }); // 30..3600
231
+ const { suspendRequest } = await workspace.suspendWhenIdle({ afterSeconds: 60 }); // 0..3600; 0 = as soon as it is idle
232
232
  workspace.suspendRequest; // { requested_at, after_seconds, not_before } until it applies or is cancelled
233
233
  await workspace.cancelSuspendWhenIdle(); // idempotent
234
234
  ```
@@ -241,7 +241,7 @@ await workspace.cancelSuspendWhenIdle(); // idempotent
241
241
  - It applies under every idle policy, `never` included, and never delays a suspend the policy would do sooner.
242
242
  - When a suspend is already in progress, the result's `operation` is that suspend and nothing is recorded.
243
243
  - Errors: `ShardfluxApiError` 409 with `reason` `not_running`, `operation_in_progress`, `session_lifetime` or
244
- `workspace_deleted`, and 422 `validation_failed` for `afterSeconds` outside 30..3600. A file-first workspace is never
244
+ `workspace_deleted`, and 422 `validation_failed` for `afterSeconds` outside 0..3600. A file-first workspace is never
245
245
  suspended: `NotSupportedForModeError` (409 `not_supported_for_mode`).
246
246
  - By id: `cloud.workspaces.suspendWhenIdle(id, { afterSeconds })` and `cloud.workspaces.cancelSuspendWhenIdle(id)`.
247
247
 
@@ -266,17 +266,37 @@ const cell = workspace.cell({ transitionTimeoutMs: 30_000 }); // give up waking
266
266
  await cell.exec.run(['make', 'test']); // resumes the workspace first if it is suspended
267
267
  ```
268
268
 
269
+ **Detecting a cold resume (0.11.0+).** A resume restores memory and running processes from the checkpoint. After a
270
+ platform runtime update, a resume can boot from the saved disk instead of restoring memory (a cold boot), also when a
271
+ tool call wakes the workspace; check `memoryRestored`. The workspace runs with its files as of the suspend, and its
272
+ processes start fresh, as after a reboot: start dev servers, databases and background jobs again. The resume's timing
273
+ says which it was:
274
+
275
+ ```ts
276
+ const t = workspace.lastTiming; // after resume({ wait: true }), wake(), or a tool call that woke the workspace
277
+ if (t?.server?.memoryRestored === false) {
278
+ // t.server.resumePath === 'cold_boot', t.server.coldBootReason === 'runtime_changed'
279
+ }
280
+ ```
281
+
282
+ - `ServerTiming.memoryRestored`: `true` when memory and processes came back, `false` when the workspace booted
283
+ (`resumePath` `cold_boot`, or `reset_blank_layer` after a reset), `null` when the API did not say (older APIs).
284
+ - `ServerTiming.coldBootReason`: why it booted (`runtime_changed`), else null. `formatTiming()` prints
285
+ `resume from cold_boot: processes restarted (runtime_changed)`.
286
+ - The finished resume operation carries the same in `result` (`memory_restored`, `cold_boot_reason`, and
287
+ `cold_boot` with details such as `files_as_of`; it is informational and not part of the stable contract).
288
+
269
289
  `suspend`, `resume`, `fork`, `snapshot` and `delete` return the lifecycle operation.
270
290
  `cloud.workspaces.waitForOperation(id)` waits for it (default timeout 5 minutes): each poll asks the API to hold
271
291
  the response until the operation changes (`Prefer: wait`, at most 20 s per request), so completion arrives within
272
292
  one round trip of the commit. Against an API without bounded waits (or with `serverWait: false`) it polls with
273
293
  backoff (250 ms doubling to 5 s, ±20 % jitter). If the timeout passes, it throws `OperationTimeoutError` and the
274
- operation keeps running server side (a start waiting for capacity until its `deadlineAt`); wait for it again with the
294
+ operation keeps running server side (a queued start until its `deadlineAt`); wait for it again with the
275
295
  same call. `templates.builds.waitForBuild()` waits the same way. A waited `open()` issues the first tool token
276
296
  together with the final workspace read, so the first tool call starts at once.
277
297
 
278
- On Node 26 the default fetch sends `Connection: close`: its bundled undici 8 can stall a request on a reused
279
- keep-alive connection for tens of seconds. Pass your own `fetch`, or set `SHARDFLUX_HTTP_KEEPALIVE=1`, to change that.
298
+ On Node 26 the default fetch uses a fresh connection for every request (`Connection: close`);
299
+ `SHARDFLUX_HTTP_KEEPALIVE=1`, or your own `fetch`, reuses connections.
280
300
 
281
301
  List and look up workspaces:
282
302
 
@@ -300,27 +320,31 @@ says where the time went, and `formatTiming()` prints it:
300
320
 
301
321
  ```ts
302
322
  const ws = await cloud.workspaces.open({ key: 'customer-42/main', template: 'python-node-browser' });
323
+ await ws.suspend({ wait: true });
324
+ await ws.resume({ wait: true });
303
325
  console.log(formatTiming(ws.lastTiming!));
304
326
  ```
305
327
 
306
- A slow open (34 s instead of the usual second) then reads, for example:
328
+ The resume then reads, for example:
307
329
 
308
330
  ```text
309
- open 34.18 s, succeeded (workspace 01a0e5a8-3edd-74ba-b489-d62b8925e342, operation 01a0e5a8-3ef0-7ecb-975e-dff2d5ca6e33)
310
- client: request 20.01 s (held) → capacity_pending 13.52 s (no_ready_host) → running 590 ms → view 42 ms ∥ token 61 ms
311
- server: queued 33.40 s, ran 620 ms, total 34.02 s; start warm, boot to ready 79 ms
312
- outside the server: 161 ms
331
+ resume 413 ms, succeeded (workspace 01a0ead6-92e5-7d19-ab7a-f565bf4676cb, operation 01a0ead6-bd61-737a-ae88-66e2d01a6e25)
332
+ client: request 218 ms → queued 195 ms
333
+ server: queued 51 ms, ran 290 ms, total 341 ms; resume from local_cache, boot to ready 211 ms, host disk 0 ms, load 6 ms, ready 62 ms, total 211 ms, after_restore 36 ms
334
+ outside the server: 72 ms
313
335
  ```
314
336
 
315
- Here the time went to waiting for a host with capacity; the network and the VM start were fast.
337
+ This resume brought the workspace back from its host's local cache, memory and running processes included, in 341 ms on the server and 413 ms end to end.
316
338
 
317
339
  - **client** phases are on your monotonic clock: the request (a held open waits on the server, `held`), then each
318
340
  operation state the SDK observed while waiting, with the server's reason (`capacity_pending` / `no_ready_host`:
319
- waiting for a host; `running` / `template_downloading`: fetching the template), then reading the workspace and
341
+ queued to start; `running` / `template_downloading`: fetching the template), then reading the workspace and
320
342
  issuing the first tool token (together).
321
343
  - **server** timing comes from the operation itself (one database clock): `queued` is creation until it began
322
- running, including any wait for capacity; `ran` is the cell's work (placement, boot or restore, guest readiness).
323
- `start` / `resume from` and `boot to ready` / `host …` are what the cell reported.
344
+ running, including any time in `capacity_pending`; `ran` is the cell's work (placement, boot or restore, guest
345
+ readiness). `start` / `resume from` and `boot to ready` / `host …` are what the cell reported. A resume that booted
346
+ the saved disk instead of restoring memory reads `resume from cold_boot: processes restarted (runtime_changed)`
347
+ (0.11.0+).
324
348
  - **outside the server** is your total minus the operation's: network, TLS, polling latency, view and token. A large
325
349
  value with a small server total points at the connection between you and the API, not at the workspace.
326
350
  - **retries** lists transient failures the SDK retried (cause and backoff).
@@ -404,13 +428,13 @@ await cell.files.write('/home/user/app/main.py', 'print("bye")\n', { ifTreeRevis
404
428
  ```
405
429
 
406
430
  - **Executions are idempotent by id.** `executionId` defaults to a fresh `ex-<uuid>`; pass your own to make a call safe
407
- to repeat across processes. Network failures and retryable 5xx answers (503 `no_execution_host`: no host has room;
408
- the SDK waits `Retry-After`) are retried with the same id, at most `maxRetries` (5) times, and an answer that takes
409
- longer than `attemptTimeoutMs` (300 s) is awaited again with the same id. The cell runs a command once per id: a
431
+ to repeat across processes. Network failures and retryable 5xx answers (503 `no_execution_host`: the execution
432
+ cannot be placed right now; the SDK waits `Retry-After`) are retried with the same id, at most `maxRetries` (5)
433
+ times, and an answer that takes longer than `attemptTimeoutMs` (300 s) is awaited again with the same id. The cell runs a command once per id: a
410
434
  repeated call gets the recorded result (`r.replayed`). The SDK never retries with a new id: a `failed` or `lost`
411
435
  result is returned, and running it again is your decision.
412
436
  - `ws.executions.get(id, { waitMs })` reads an execution: `pending` while it runs, its result when it ended (kept 7
413
- days). Use it after a call that stopped waiting (an aborted `signal`): an execution cannot be canceled.
437
+ days). Use it after a call that stopped waiting (an aborted `signal`): an execution runs to completion.
414
438
  - While an execution runs, other executions and file writes are refused with 409 `workspace_busy`
415
439
  (`execution_in_progress`); the SDK waits them out within `transitionTimeoutMs` (120 s).
416
440
  - Only paths under /home/user exist (`outside_tree_root` otherwise). Reads, writes, `search()` and `patch()` work as on
@@ -418,7 +442,7 @@ await cell.files.write('/home/user/app/main.py', 'print("bye")\n', { ifTreeRevis
418
442
  - Calls a file-first workspace does not have fail with `NotSupportedForModeError` before any request: exec sessions
419
443
  (`cell.exec.*`), PTY, processes, git, browser, `changes()`, `suspend`, `resume`, `snapshot`, `fork`, `reset` and
420
444
  `saveAsTemplate`. On a processful workspace `executions` and `ifTreeRevision` are refused the same way.
421
- - Reopening a key with another `mode` is 409 `mode_mismatch`; a deployment without file-first workspaces answers 422
445
+ - Reopening a key with another `mode` is 409 `mode_mismatch`; an account without file-first workspaces gets 422
422
446
  `mode_not_available`; a legacy template 409 `layout_unsupported`.
423
447
 
424
448
  ## Templates: file tree, diff and dev mode
@@ -667,10 +691,9 @@ pending: past that, while the workspace takes no writes, a call is not recorded
667
691
  `onError` gets `queue_full`). A write is retried for `retryWindowMs` (120 s) and then dropped (`write_failed`).
668
692
  Observing an async tool marks its promise handled: if your code never awaits a wrapped tool's promise and it rejects,
669
693
  Node reports no `unhandledRejection` for it (the error is still in the index; Node has no way to observe a rejection
670
- without handling it). About one write per call: more than roughly 100 calls per second per
671
- workspace reaches the pending limit. `Response`, `ReadableStream`, Node streams and `Blob` results are never read
672
- (`meta.note: "stream_not_captured"`). On template v1 the files API writes as root, so the agent can read the captured
673
- files but not change them. The format is specified in `docs/decisions/0006-tool-call-capture.md`.
694
+ without handling it). About one write per call: the default limits handle about 100 calls per second per workspace;
695
+ raise the pending limits for more. `Response`, `ReadableStream`, Node streams and `Blob` results are never read
696
+ (`meta.note: "stream_not_captured"`). On template v1, captured files are read-only to the agent.
674
697
 
675
698
  ## Secrets
676
699
 
@@ -781,10 +804,9 @@ changes.
781
804
  (`details.reason`, e.g. `not_session`, `draft_not_found`, `legacy_disk_layout`; see `KnownErrorReason`).
782
805
  - `OperationFailedError`: an awaited operation ended `failed` or `canceled` (`errorCode`,
783
806
  `retryable` **(0.6.2+)**, `operation`, `timing`). `retryable` is the operation error's own flag: `true` for
784
- `capacity_unavailable` (no host could admit the start before its deadline; retry later), `false` for a definitive
785
- failure.
807
+ `capacity_unavailable` (the start passed its deadline; retry it), `false` for a definitive failure.
786
808
  - `OperationTimeoutError`: waiting gave up; the operation continues (`operationId`, `lastState`, `lastReason`,
787
- `deadlineAt` **(0.6.2+)** while it waits for capacity, `timing`).
809
+ `deadlineAt` **(0.6.2+)** while a start is queued, `timing`).
788
810
  - `ShardfluxProtocolError`: a response was not the documented shape.
789
811
  - `NotSupportedForModeError` **(0.9.0+)**, a `ShardfluxApiError` (409 `conflict`, `reason` `not_supported_for_mode`):
790
812
  the call does not exist for the workspace's `mode`; `local` is true when the SDK refused it without a request.
@@ -807,14 +829,14 @@ carries `reason` **(0.9.0+)**:
807
829
  `details.spend_cap` is `{ cap_minor, effective_cap_minor, charges_minor, currency }` (minor units, `null` when the plan
808
830
  has no overage). An older API sends no `reason`. Do not retry these in a loop.
809
831
 
810
- Retryable 429/502/503/504 refusals (for example 503 `host_capacity`, when the workspace's host has no room to restore
811
- it right now, or `wake_failed`) are retried after `Retry-After` for reads, searches and calls that carry an
832
+ Retryable 429/502/503/504 refusals (for example 503 `host_capacity` when the workspace cannot be woken right now, or
833
+ `wake_failed`) are retried after `Retry-After` for reads, searches and calls that carry an
812
834
  Idempotency-Key (writes and patches); other calls surface them with `retryable: true` and `retryAfterSeconds`.
813
835
  A read of a sleeping workspace that its disk cannot answer (409 `workspace_not_running` with `reason`
814
836
  `offline_unavailable` or `offline_budget`) wakes the workspace and is retried like any `workspace_not_running`; 503
815
837
  `offline_changed` (the disk changed during the read) is retried and served by the running workspace. 409 `conflict`
816
- `host_feature_unavailable` (the workspace's host predates the call, `details.feature`) is neither retried nor
817
- woken: it lasts until the workspace runs on an upgraded host.
838
+ `host_feature_unavailable` (search or patches are not available for the workspace, `details.feature`) is neither
839
+ retried nor woken: run `grep` with `exec`, or read then write the file, instead.
818
840
 
819
841
  ## Usage and overage
820
842
 
@@ -848,10 +870,10 @@ is charged on the next invoice until the charges reach the spend cap.
848
870
 
849
871
  ## Feedback (0.9.0+)
850
872
 
851
- `cloud.sendFeedback()` sends a message straight to the Shardflux founder, who reads every one. If you or your coding
873
+ `cloud.sendFeedback()` sends a message straight to the Shardflux team, who read every one. If you or your coding
852
874
  agent hit something while building with Shardflux, send it the moment it happens: a call that failed unexpectedly, an
853
- error or doc that was confusing, something missing or slow, a workaround you needed. Short and specific beats polished;
854
- the request id and error code let the founder find the logs.
875
+ error or doc that was confusing, something missing, a workaround you needed. Short and specific beats polished;
876
+ the request id and error code let the team find the logs.
855
877
 
856
878
  ```ts
857
879
  try {
@@ -859,7 +881,7 @@ try {
859
881
  } catch (err) {
860
882
  if (err instanceof ShardfluxApiError) {
861
883
  await cloud.sendFeedback({
862
- message: 'open failed with capacity_unavailable twice in 10 minutes; expected a start within a minute',
884
+ message: 'open rejected template "python-node" with template_not_found; expected a suggestion of the closest slug',
863
885
  category: 'bug',
864
886
  context: { requestId: err.requestId, errorCode: err.code, workspace: 'acme/demo', agent: 'claude-code' },
865
887
  });
@@ -911,7 +933,7 @@ or no longer supported, it emits one warning:
911
933
 
912
934
  - The SDK follows the API's `/v1` contract. New fields, enum values and error codes can appear in
913
935
  any release; ignore unknown fields.
914
- - While below 1.0, a breaking change bumps the minor version (0.7 to 0.8).
936
+ - Breaking changes ship only in minor releases (0.7 to 0.8) and are marked **Breaking** in the changelog.
915
937
  - `SDK_VERSION` is exported; requests send `User-Agent: shardflux-sdk-ts/<version>`.
916
938
  - Examples in this README, in `examples/` and on shardflux.dev name the version they need. The examples on the
917
939
  website and in the console are checked against the version published on npm before they ship.
package/dist/account.d.ts CHANGED
@@ -1,5 +1,5 @@
1
1
  /**
2
- * The account plane with a user session (0.9.0; contracts §30): what a person does in the web app, from code.
2
+ * The account plane with a user session (0.9.0): what a person does in the web app, from code.
3
3
  *
4
4
  * // Before a session exists (no token needed)
5
5
  * await ShardfluxAccount.register({ email, password, displayName });
@@ -451,7 +451,7 @@ export declare class ShardfluxAccount {
451
451
  /** The signed-in user and their memberships. */
452
452
  me(): Promise<Me>;
453
453
  /**
454
- * Send feedback straight to the Shardflux founder as the signed-in user (0.9.0+; POST /v1/feedback), optionally about
454
+ * Send feedback straight to the Shardflux team as the signed-in user (0.9.0+; POST /v1/feedback), optionally about
455
455
  * one of your organizations (`organizationId`). Otherwise the same as `Shardflux.sendFeedback`: use it while you work,
456
456
  * returns `{ id, receivedAt, duplicate }`, never retried, 429 `rate_limited` per user.
457
457
  */
package/dist/account.js CHANGED
@@ -583,7 +583,7 @@ export class ShardfluxAccount {
583
583
  return this.#ctx.http.json('GET', '/v1/me', {}, this.#ctx.authorization);
584
584
  }
585
585
  /**
586
- * Send feedback straight to the Shardflux founder as the signed-in user (0.9.0+; POST /v1/feedback), optionally about
586
+ * Send feedback straight to the Shardflux team as the signed-in user (0.9.0+; POST /v1/feedback), optionally about
587
587
  * one of your organizations (`organizationId`). Otherwise the same as `Shardflux.sendFeedback`: use it while you work,
588
588
  * returns `{ id, receivedAt, duplicate }`, never retried, 429 `rate_limited` per user.
589
589
  */