@shardflux/sdk 0.12.0 → 0.13.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -3,6 +3,195 @@
3
3
  Every API the README shows is available from the version named here. Breaking changes ship in minor releases and are
4
4
  marked **Breaking**.
5
5
 
6
+ ## 0.14.0 (not yet published)
7
+
8
+ ### Workspaces working at once
9
+
10
+ Additive. The plan's workspace number caps workspaces working at the same moment; stored workspaces are bounded by
11
+ Retained state.
12
+
13
+ - A tool call that finds every workspace the plan allows to work at once busy is refused with 429 `quota_exceeded`
14
+ (retryable, `Retry-After`; `details.limit` `concurrent_workspaces`, `limit_value`, `current`,
15
+ `retry_after_seconds`). The call did not run, so the SDK sends the same request again after `Retry-After` (at most
16
+ 5 s) within `maxRetries`, for every call: exec starts, stdin, PTY create, keepalive and process signals included.
17
+ A refusal that outlasts the retries is a `ShardfluxApiError` with `retryAfterSeconds`. `executions.run()` keeps
18
+ these retries and does not add its own on top; the wake hint is not retried (it never waits).
19
+ - `pty.read()` rejects with the refusal (`ShardfluxApiError`, e.g. 429 `quota_exceeded`, retryable) when an accepted
20
+ attach is closed with an error body and a 4xxx close code, instead of returning empty output.
21
+ - The API's 403 `quota_exceeded` is not retried, as before. New `details.limit`: `retained_state` (`limit_value` and
22
+ `current` in GiB): opening a new key and forking are refused while the plan's Retained state is used up; existing
23
+ workspaces keep running, waking, suspending and resuming. Usage summaries report that allowance with `enforcement`
24
+ `storage_block` and `cap_state` `storage_blocked`.
25
+ - Regenerated contract types: `ErrorCode` includes `quota_exceeded` for the workspace gateway; allowance
26
+ `enforcement` adds `storage_block`, `cap_state` adds `storage_blocked`.
27
+
28
+ ## 0.13.1 (not yet published)
29
+
30
+ Patch, additive: the transport, a documented value, a fix for tool calls across a move and self-healing workspaces.
31
+ `@shardflux/cli` 0.8.1, `@shardflux/mcp` 0.7.1 and the `shardflux` bundle 0.10.1 follow (`^0.13.1`).
32
+
33
+ ### Workspaces recover by themselves
34
+
35
+ Additive. When the machine a workspace runs on fails, the workspace is suspended, and its next use (a tool call's wake,
36
+ `resume()`, `open()`) restores it: from its own disk, files kept, or from its newest checkpoint. Nothing changes for
37
+ code that already handles a cold resume.
38
+
39
+ - `cold_boot_reason` / `ServerTiming.coldBootReason` can be `host_lost`: the resume booted the workspace's disk
40
+ (`resumePath` `cold_boot`, `memoryRestored` false, processes restarted). New type `ColdBootReason`
41
+ (`'runtime_changed' | 'host_lost'`, open to other strings; the field still accepts any string).
42
+ - `hostLostOf(op)` and `ServerTiming.hostLost`: `result.host_lost` camelCased (`HostLost`: `detectedAt`,
43
+ `restoredFrom` `'disk' | 'checkpoint'`, `restoredCheckpointId`, `stateAsOf`). A resume from the checkpoint has its
44
+ state as of `stateAsOf`; a suspend that found the machine failed succeeds with `durable: true` and only
45
+ `detectedAt` (`restoredFrom` null).
46
+ - `formatTiming()` prints `processes restarted (host_lost)` for the disk, `restored checkpoint <id> (host_lost; state as
47
+ of <time>)` for the checkpoint, and `host_lost (detected <time>)` for the suspend.
48
+ - Documented operation error codes: `resume_required` (a fork or snapshot of such a workspace before its resume; resume
49
+ it first; not retryable) and `workspace_storage_unavailable` with `details.reason` `host_lost` (not retryable).
50
+ `KnownErrorReason` gains `host_lost`.
51
+
52
+ ### Connection reuse
53
+
54
+ Additive. The call after an agent's pause between tool calls reuses its connection instead of opening a new one.
55
+
56
+ - On Node 26 the pooled transport keeps an idle connection reusable for 5 minutes (was 4 s), so a request seconds or
57
+ minutes after the last one skips the TCP and TLS handshake. Pooled sockets send TCP keepalive probes after 60 s
58
+ idle, so NATs and firewalls along the way keep the connection open. Idle connections never keep a process alive. A
59
+ supplied `fetch`, other runtimes and `SHARDFLUX_HTTP_KEEPALIVE` work as before.
60
+ - A request the SDK retries (GET, a request with an `Idempotency-Key`, a read-only POST such as files search) whose
61
+ connection closes before the response arrives is sent again at once on a new connection. That one resend has no
62
+ backoff and does not count against `maxRetries`; timing records and progress `retry` events list it with
63
+ `delayMs: 0`. A further failure follows `maxRetries` and backoff as before, and a POST without an
64
+ `Idempotency-Key` is never sent twice.
65
+
66
+ ### Faster suspend and resume
67
+
68
+ Additive. `resumePath` / `resume_path` can be `thaw`: a resume that arrives while its suspend is still being
69
+ written continues the same VM in place. Nothing else changes for your code.
70
+
71
+ ### Tool calls across a move
72
+
73
+ Fix. A tool call made while its workspace finishes a resume or a move now runs once the workspace runs, instead of
74
+ failing with `409 conflict` (`workspace_not_running`).
75
+
76
+ - When the call's tool token was refused as not running and the wake then found the workspace already running (the
77
+ resume or move committed in between; `wake` resolves `false`), the call is retried with a current token: the held
78
+ resume's, or a new one. A refused call never ran, so the retry cannot run it twice. Before, the refusal surfaced.
79
+ - A refusal that comes back while the API keeps reporting the workspace running is retried after 500 ms, within the
80
+ same 3 wakes and `transitionTimeoutMs`, then surfaces as before.
81
+ - `onProgress` sees the retry as a `tool` event of type `retry`, cause `conflict workspace_not_running (the workspace
82
+ runs; new token)`.
83
+ - The CLI, the MCP server and the `shardflux` bundle get it through their `^0.13.0` dependency.
84
+
85
+ ## 0.13.0
86
+
87
+ ### Elastic memory
88
+
89
+ Additive. A workspace may be promised more memory than it holds while idle, and the host grows it when a command needs
90
+ it (opt-in per workspace).
91
+
92
+ - `Caps.allocation_mode` (`'fixed' | 'elastic'`, type `AllocationMode`) and `Caps.memory_mib_held` on `open()`,
93
+ `workspaces.fork()` and `workspace.fork()`. Given caps replace the stored ones (caps without `allocation_mode` make
94
+ the workspace fixed); omitted caps keep the stored layout.
95
+ - `workspace.memory` (`WorkspaceMemory`: `{allocation_mode, promised_mib, held_mib, plugged_mib}`, null on an older
96
+ API) and `workspace.allocationMode` (`'fixed'` on an older API). `WorkspaceView` carries `caps.allocation_mode`,
97
+ `caps.memory_mib_held` and `memory` (regenerated contract types).
98
+ - `RunResult.memoryGrow` (`MemoryGrow | null`): the grow the exec's start waited for (`outcome` delivered, partial,
99
+ missed or failed; `from_mib`, `want_mib`, `got_mib`, `deliver_ms`), from the start's session. `ExecSession`
100
+ carries `memory_grow`. The `exec` agent tool adds `memory_grow` to its result when a grow ran.
101
+ - `KnownErrorReason` adds `allocation_mode_not_available`, `requires_elastic` and `exceeds_memory_mib` (422
102
+ `validation_failed`; elastic with file-first is the existing `not_supported_for_mode`).
103
+
104
+ ### Burst execution
105
+
106
+ Additive. One heavy, run-to-completion command can run on a larger, short-lived burst VM on the workspace's host, and
107
+ its file changes are applied back.
108
+
109
+ - `RunOptions.burst` (`'never' | 'always'`), `burstVcpus` and `burstMemoryMib` on `exec.run()`; `ExecStartRequest`
110
+ carries `burst`, `burst_vcpus` and `burst_memory_mib` (regenerated contract types).
111
+ - `RunResult.burst` (`BurstSummary | null`): host, size, method, whether the changes were applied, the change counts,
112
+ `leftover_killed`, `overhead_ms`, `excluded_paths`, `timings`, `replayed`. `ExecSession.burst` carries it; types
113
+ `BurstSummary` and `BurstError` are exported.
114
+ - `exec.run()` rejects with `ShardfluxApiError` 409 `burst_unavailable` or `burst_apply_failed` also when the burst
115
+ fails after its start (the output stream's final error event, or `burst.error` on the ended session), without
116
+ reconnecting; it does not try to cancel a burst when its `signal` aborts (a burst cannot be canceled).
117
+ - `ErrorCode` adds `burst_unavailable` and `burst_apply_failed`; `KnownErrorReason` adds `burst_mode_not_supported`,
118
+ `burst_not_supported`, `burst_size_exceeds_plan`, `not_available`, `shared_volumes`, `fence_not_drained`,
119
+ `apply_pending`, `park_failed`, `workspace_resumed`, `interrupted`, `burst_lost`, `disk_full`, `apply_failed`,
120
+ `reverted` and `revert_failed`.
121
+ - `workspaceTools(ws, { burst: true })` (opt-in, default off) gives the processful `exec` tool the inputs `burst`,
122
+ `burst_vcpus` and `burst_memory_mib`, and its result adds `burst`. The default tool schemas are unchanged.
123
+
124
+ ### Background commands in the agent tools
125
+
126
+ Additive. An agent can start a long command, keep working with the other tools, read its progress and stop it.
127
+
128
+ - The processful `exec` tool takes `background: true`: it starts the command and returns `{session_id, state}` at
129
+ once (a failed start adds `error: {code, message, reason}`, as the foreground `exec` gives it). `timeout_ms` goes up
130
+ to 86400000 (a day); with `background` it is sent only when given. The file-first `exec` is unchanged.
131
+ - New tool `exec_read` (permission `exec`): the session's `state`, `exit_code`, `term_signal`, `timed_out`,
132
+ `canceled` and output. Without offsets it returns the last `maxOutputBytes` of each stream; `stdout_offset` /
133
+ `stderr_offset` read from there, and the result's `next_stdout_offset` / `next_stderr_offset` continue where it
134
+ stopped (`truncated` when more output follows). Text is cut at character boundaries only. `wait_ms` (up to 60000)
135
+ waits for the command to exit and returns as soon as it does. `burst` and `error` when the session has them.
136
+ - New tool `exec_cancel` (permission `exec`): SIGTERM, then SIGKILL after `grace_ms` (1-60000, default 5000);
137
+ returns `{session_id, state, exit_code, canceled}`.
138
+ - With `workspaceTools(ws, { burst: true })`, `background: true` together with `burst: 'always'` is refused with
139
+ `ToolArgumentError` before any request (a burst holds the workspace until it applies its changes); `burst: 'never'`
140
+ with `background` is an ordinary background start.
141
+ - `exec_read` and `exec_cancel` follow `exec` in `workspaceTools()` on processful workspaces; a file-first workspace
142
+ does not offer them. Both send the wake hint like the other tools.
143
+
144
+ ### Held fork
145
+
146
+ Additive. A waited fork returns the running copy and its tool token in one request.
147
+
148
+ - `workspaces.fork()` and `workspace.fork()` with `wait`: when the API holds the fork (`Prefer: wait`, answered 200
149
+ with `Preference-Applied`), the copy's handle adopts the returned running view and final-epoch tool token, with no
150
+ operation poll, view refresh or token request. The request's time counts against `wait.timeoutMs`.
151
+ - `ForkOptions` and `WaitedForkOptions` (exported) add `agentLabel` and `tools`, which choose that token. They are sent
152
+ only on a server-held fork and need an API with held fork. `wait: { serverWait: false }` keeps polling.
153
+ - An API that does not apply `Prefer` answers 202 as before; the call then polls the operation and refreshes the copy.
154
+ - The fork's progress `request` phase has reason `held` when it asks the server to wait.
155
+
156
+ ### Exec start and output in one request
157
+
158
+ - `exec.run()` starts the command and follows its output in one request: the start asks for NDJSON
159
+ (`Accept: application/x-ndjson`) and a cell that supports it answers with the output stream. A dropped stream
160
+ reconnects by session and byte offsets, never starting the command again. A cell that answers the start with the
161
+ session JSON gets the output request after it, as before; a start that failed still rejects with `ExecStartError`
162
+ (or the burst's error).
163
+ - `RunResult.memoryGrow` and `burst` are unchanged: the combined stream's exit event carries the start's
164
+ `memory_grow`.
165
+
166
+ ### Immutable paths
167
+
168
+ A template may declare directories as immutable: every workspace of the template mounts them read-only from the
169
+ template's newest published version, picked up at its next cold boot or resume.
170
+
171
+ - Recipe v2 `immutable` (`TemplateRecipeV2.immutable`, regenerated): the list in `template.yaml`, sent as written by
172
+ `buildFromFile()`, `buildFromRecipe()` and `builds.create()`. The stored recipe v2 of a build
173
+ (`TemplateBuildRecipeV2`) and an exported recipe (`versions.recipe()`) carry it.
174
+ - `TemplateVersion.immutable` (`TemplateVersionImmutable`: `{paths, bytes}`, or null when the version declares none),
175
+ on every version of `templates.get()` and on `open_version`.
176
+ - `TemplateBuild.immutable_paths`: the list the produced version declares (the recipe's, else the open version's).
177
+ `TemplateBuild.failure.details` carries code-specific fields: `immutable_path_missing` names `details.path`.
178
+ - `workspace.immutableVersion` (`template.immutable_version` of the view): the template version whose immutable paths
179
+ the workspace has mounted, which can be newer than `template.version`; null when it mounts none.
180
+ - `KnownErrorReason` adds `immutable_path_removed` (422; `details.removed`, `details.open_version`),
181
+ `immutable_paths_unsupported_base` (422; `details.base`, `details.required_feature`, `details.paths`) and
182
+ `read_only_path` (409 `conflict` from the files API for a write under an immutable path).
183
+ - **Breaking:** the update policy left the API. Removed: the `UpdatePolicy` type, `TemplateSummary.update_policy`, the
184
+ regenerated `update_policy` of the workspace view, of `TemplateDefaults` and of the `defaults` inputs, and the reason
185
+ `update_policy_not_available`. It was always `pinned`; code that read it drops the read.
186
+
187
+ ### Exec cancel grace
188
+
189
+ - `exec.cancel(id, graceMs?)` takes `graceMs` 0 to 60000 (0 or omitted: 5000), as the cell API now enforces (422
190
+ `validation_failed` outside it), and waits for the answer up to the grace plus 15 s when that is longer than the
191
+ client's `timeoutMs`: a command that ignores SIGTERM is answered after the grace and the SIGKILL, not a timeout.
192
+ - The cancel that `exec.run` sends when its `signal` aborts passes `killGraceMs` capped at 60000, so a larger
193
+ `killGraceMs` still cancels the command.
194
+
6
195
  ## 0.12.0 (release candidate)
7
196
 
8
197
  Completed exec output streams are drained before releasing their HTTP connections, with a bounded cleanup if a peer does not close.
@@ -336,7 +525,6 @@ behavior is the automatic version check (below), which makes one background requ
336
525
  - New exports: `FeedbackCategory`, `FeedbackContext`, `FeedbackReceipt`, `SendFeedbackParams`, `AccountFeedbackParams`,
337
526
  `FEEDBACK_CATEGORIES`, `FEEDBACK_MESSAGE_MAX_LENGTH`.
338
527
 
339
-
340
528
  ## 0.8.0 (2026-09-28)
341
529
 
342
530
  Types only; nothing changes at run time and the API is unchanged.
package/README.md CHANGED
@@ -108,8 +108,9 @@ const listing = await cell.files.list('/home/user');
108
108
  await cell.files.remove('/home/user/data.bin');
109
109
  ```
110
110
 
111
- `exec.run()` resumes from byte offsets if the output stream drops; it never starts the command
112
- twice. Aborting its `signal` also cancels the command in the workspace. File writes are atomic
111
+ `exec.run()` starts the command and streams its output in one request **(0.13.0+)**, resumes from
112
+ byte offsets if the output stream drops, and never starts the command twice. Aborting its `signal`
113
+ also cancels the command in the workspace. File writes are atomic
113
114
  and durable (acknowledged after fsync).
114
115
 
115
116
  `cwd` is an absolute path (commands start in `/home/user` without one). The API refuses a relative
@@ -206,6 +207,12 @@ means depends on `wait`:
206
207
  wait again with `cloud.workspaces.waitForOperation(err.operationId)`. Without `wait`, the returned operation is the
207
208
  handle for the work in progress: pass its `id` to `waitForOperation()` when you need it finished.
208
209
 
210
+ **A waited fork is one request (0.13.0+).** `fork(target, { wait: true })` asks the API to hold the fork until the
211
+ copy runs, and the answer carries the running copy and a tool token for it: the copy's handle is ready with no
212
+ operation poll, view refresh or token request. `agentLabel` and `tools` pick the token it brings back. Its timing is one
213
+ `request` phase with reason `held`; `wait: { serverWait: false }` polls instead. An API without the held fork answers
214
+ at once, and the SDK then waits for the operation and refreshes the copy, as before.
215
+
209
216
  **Durable storage (0.12.0+).** A suspend returns as soon as the workspace is sealed on its host, typically in a few
210
217
  hundred ms, and its RAM and CPU are released at that moment. `result.durable` turns `true` when the copy lands in
211
218
  durable storage, typically within a second; until then it is `false` and `result.durability` shows the copy's
@@ -289,10 +296,10 @@ await cell.exec.run(['make', 'test']); // resumes the wo
289
296
  ```
290
297
 
291
298
  **Detecting a cold resume (0.11.0+).** A resume restores memory and running processes from the checkpoint. After a
292
- platform runtime update, a resume can boot from the saved disk instead of restoring memory (a cold boot), also when a
293
- tool call wakes the workspace; check `memoryRestored`. The workspace runs with its files as of the suspend, and its
294
- processes start fresh, as after a reboot: start dev servers, databases and background jobs again. The resume's timing
295
- says which it was:
299
+ platform runtime update, or after the machine the workspace ran on failed (0.13.1+), a resume can boot from the saved
300
+ disk instead of restoring memory (a cold boot), also when a tool call wakes the workspace; check `memoryRestored`. The
301
+ workspace runs with its files kept, and its processes start fresh, as after a reboot: start dev servers, databases and
302
+ background jobs again. The resume's timing says which it was:
296
303
 
297
304
  ```ts
298
305
  const t = workspace.lastTiming; // after resume({ wait: true }), wake(), or a tool call that woke the workspace
@@ -303,11 +310,20 @@ if (t?.server?.memoryRestored === false) {
303
310
 
304
311
  - `ServerTiming.memoryRestored`: `true` when memory and processes came back, `false` when the workspace booted
305
312
  (`resumePath` `cold_boot`, or `reset_blank_layer` after a reset), `null` when the API did not say (older APIs).
306
- - `ServerTiming.coldBootReason`: why it booted (`runtime_changed`), else null. `formatTiming()` prints
307
- `resume from cold_boot: processes restarted (runtime_changed)`.
313
+ - `ServerTiming.coldBootReason` (type `ColdBootReason`): why it booted, `runtime_changed` or `host_lost`
314
+ **(0.13.1+)**, else null. `formatTiming()` prints `resume from cold_boot: processes restarted (runtime_changed)`.
308
315
  - The finished resume operation carries the same in `result` (`memory_restored`, `cold_boot_reason`, and
309
316
  `cold_boot` with details such as `files_as_of`; it is informational and not part of the stable contract).
310
317
 
318
+ **Workspaces recover by themselves (0.13.1+).** When the machine a workspace runs on fails, the workspace is suspended,
319
+ and its next use (a tool call, `resume()`, `open()`) restores it; there is nothing extra to call.
320
+ `hostLostOf(op)` and `lastTiming.server.hostLost` (`HostLost`) say how:
321
+
322
+ - `restoredFrom: 'disk'`: it booted from its own disk, files kept (`coldBootReason` `host_lost`, processes restarted).
323
+ - `restoredFrom: 'checkpoint'`: it resumed its newest checkpoint (`restoredCheckpointId`); its state is as of
324
+ `stateAsOf`. `formatTiming()` adds `restored checkpoint <id> (host_lost; state as of <time>)`.
325
+ - A `suspend()` that finds the machine failed succeeds (`result.durable` true, `hostLost.detectedAt`).
326
+
311
327
  `suspend`, `resume`, `fork`, `snapshot` and `delete` return the lifecycle operation.
312
328
  `cloud.workspaces.waitForOperation(id)` waits for it (default timeout 5 minutes): each poll asks the API to hold
313
329
  the response until the operation changes (`Prefer: wait`, at most 20 s per request), so completion arrives within
@@ -317,7 +333,15 @@ operation keeps running server side (a queued start until its `deadlineAt`); wai
317
333
  same call. `templates.builds.waitForBuild()` waits the same way. A waited `open()` issues the first tool token
318
334
  together with the final workspace read, so the first tool call starts at once.
319
335
 
320
- From 0.12.0, Node 26 reuses TLS connections through a private HTTP/1.1 pool. Other runtimes keep native fetch. A supplied `fetch` stays unchanged. `SHARDFLUX_HTTP_KEEPALIVE=0` forces connection closure for diagnosis; `=1` opts into native pooling. No global dispatcher is installed.
336
+ From 0.12.0, Node 26 reuses TLS connections through a private HTTP/1.1 pool. From 0.13.0, an idle connection stays
337
+ reusable for 5 minutes, so the call after an agent's pause between tool calls skips the TCP and TLS handshake, and
338
+ pooled sockets send TCP keepalive probes after 60 s idle. Idle connections never keep a process alive. Other runtimes
339
+ keep native fetch. A supplied `fetch` stays unchanged. `SHARDFLUX_HTTP_KEEPALIVE=0` forces connection closure for
340
+ diagnosis; `=1` opts into native pooling. No global dispatcher is installed.
341
+
342
+ A request the SDK retries (safe methods, requests with an `Idempotency-Key`, read-only POSTs) whose connection closes
343
+ before the response arrives is sent again at once on a new connection **(0.13.0+)**, without backoff and outside
344
+ `maxRetries`; timing records list it with `delayMs: 0`.
321
345
 
322
346
  List and look up workspaces:
323
347
 
@@ -334,6 +358,46 @@ sessions, template drafts and test instances. `findByKey()` searches every lifet
334
358
  workspace, and a tombstone only when no live workspace has the key. Use it rather than `listAll({ keyPrefix })` for
335
359
  lookups by key.
336
360
 
361
+ ### Elastic memory (0.13.0+)
362
+
363
+ Promise a workspace a lot of memory while it holds only what it uses. An elastic workspace idles at its held floor
364
+ (`memory_mib_held`, default 1024) and grows before each command starts, so compilers, test runners and V8 size
365
+ themselves from the memory they will actually get. After 30 s without work it shrinks back. The layout takes effect at
366
+ the workspace's next start.
367
+
368
+ ```ts
369
+ const ws = await cloud.workspaces.open({
370
+ key: 'customer-42/repo-7',
371
+ template: 'node',
372
+ caps: { memory_mib: 8192, allocation_mode: 'elastic' },
373
+ });
374
+ ws.memory; // { allocation_mode: 'elastic', promised_mib: 8192, held_mib: 1024, plugged_mib: 3072 }
375
+
376
+ const r = await ws.cell().exec.run(['npm', 'test']);
377
+ r.memoryGrow; // { outcome: 'delivered', from_mib: 1024, want_mib: 4096, got_mib: 4096, deliver_ms: 175 }
378
+ ```
379
+
380
+ Given `caps` replace the stored ones, so send `allocation_mode` with every `caps` you pass; an open without `caps`
381
+ keeps the stored layout. The exec agent tool adds `memory_grow` to its result when the command's start grew the VM.
382
+
383
+ ### Burst execution (0.13.0+)
384
+
385
+ Run one heavy command, such as a cold build or a full test suite, on a larger VM without resizing the workspace. With
386
+ `burst: 'always'` the command runs on a short-lived burst VM over a copy of the workspace, output streams as usual,
387
+ and its file changes are applied back byte for byte when it exits.
388
+
389
+ ```ts
390
+ const r = await ws.cell().exec.run(['go', 'build', './...'], { burst: 'always', burstVcpus: 16 });
391
+ r.exitCode; // 0
392
+ r.burst; // { host: 'local', vcpus: 16, memory_mib: 8192, method: 'layer', applied: true,
393
+ // written_files: 3, written_dirs: 1, removed: 1, written_bytes: 40960, overhead_ms: 641, ... }
394
+ ```
395
+
396
+ A burst carries back the command's file changes; processes it started and tmpfs content stay in the burst VM, and the
397
+ workspace's own processes resume where they were. A burst is not signalled or canceled. If applying the changes
398
+ fails (`burst_apply_failed`), retry `exec.run` with the same `sessionId` to finish the apply. The agent tools offer
399
+ bursts when asked: `workspaceTools(ws, { burst: true })`.
400
+
337
401
  ### Timing and progress (0.6.0+)
338
402
 
339
403
  Every open, wake and waited lifecycle call is traced. `workspace.lastTiming` (and `err.timing` when the call fails)
@@ -365,7 +429,7 @@ This resume brought the workspace back from its host's local cache, memory and r
365
429
  running, including any time in `capacity_pending`; `ran` is the cell's work (placement, boot or restore, guest
366
430
  readiness). `start` / `resume from` and `boot to ready` / `host …` are what the cell reported. A resume that booted
367
431
  the saved disk instead of restoring memory reads `resume from cold_boot: processes restarted (runtime_changed)`
368
- (0.11.0+).
432
+ (0.11.0+), or `(host_lost)` (0.13.1+) after the machine the workspace ran on failed.
369
433
  - **outside the server** is your total minus the operation's: network, TLS, polling latency, view and token. A large
370
434
  value with a small server total points at the connection between you and the API, not at the workspace.
371
435
  - **retries** lists transient failures the SDK retried (cause and backoff).
@@ -569,6 +633,27 @@ workspace keeps running so you can inspect it, and `workspace.startup` names the
569
633
  The next open runs the failed step again. Versions report their `settings`, platform templates their `category`
570
634
  (`os` or `stack`), builds the `denied_hosts` their build network refused (add them to `build.network.extra_hosts`).
571
635
 
636
+ ### Immutable paths (0.13.0+)
637
+
638
+ Directories a template declares `immutable` follow the template: every workspace of it mounts them read-only from the
639
+ template's newest published version, so the tools, models or data you ship there reach existing workspaces with your
640
+ next version, at their next cold boot or resume. Everything else in a workspace stays its own.
641
+
642
+ ```yaml
643
+ # template.yaml
644
+ immutable: [/opt/acme] # read-only in every workspace; follows the newest version
645
+ ```
646
+
647
+ ```ts
648
+ const t = await cloud.templates.get('acme-dev');
649
+ console.log(t.open_version?.immutable); // { paths: ['/opt/acme'], bytes: 100663296 }
650
+ const ws = await cloud.workspaces.open({ key: 'customer-42/main', template: 'acme-dev' });
651
+ console.log(ws.template.version, ws.immutableVersion); // created from v1; /opt/acme shows v2
652
+ ```
653
+
654
+ A build without `immutable` keeps the list of the template's open version, and each new version keeps every path of
655
+ it: add paths, never remove them. A build reports the list its version declares in `immutable_paths`.
656
+
572
657
  ## Agent tools
573
658
 
574
659
  `workspaceTools(workspace)` returns tools with a name, a description, a JSON Schema for the
@@ -603,6 +688,20 @@ instead of being overwritten. Each call first sends `workspace.hint()` without w
603
688
  that off, e.g. when you send the hint yourself as the model starts a tool call), except `read_file`, `list_files` and
604
689
  `search_files`: a sleeping workspace answers them from its disk without waking.
605
690
 
691
+ **Background commands (0.13.0+).** On a processful workspace `exec` takes `background: true` for anything that runs
692
+ longer than a few minutes (a build, a test suite, a training run, a server): it returns `{ session_id, state }` at
693
+ once and the model keeps using the other tools. `exec_read` reads the command's state, exit code and output (the end
694
+ of each stream, or from the `next_stdout_offset` / `next_stderr_offset` of the previous read; `wait_ms` waits for the
695
+ exit and returns as soon as it happens), and `exec_cancel` stops it. `timeout_ms` (up to a day) sets the longest the
696
+ command may run; the workspace stays awake while it does. A burst (`burst: 'always'`) runs in the foreground.
697
+
698
+ ```ts
699
+ const tools = workspaceTools(workspace);
700
+ const { session_id } = (await executeToolCall(tools, { name: 'exec', input: { command: 'make -j8 test', background: true, timeout_ms: 7_200_000 } })) as { session_id: string };
701
+ // ... other tool calls ...
702
+ const progress = await executeToolCall(tools, { name: 'exec_read', input: { session_id, wait_ms: 30_000 } });
703
+ ```
704
+
606
705
  For a file-first workspace (0.9.0+) the tools are `exec` and the files tools only: `exec` runs each command as an
607
706
  execution and adds `execution_id`, `state`, `tree_revision` and `changed` to its result, and the process, terminal,
608
707
  git and browser tools are not offered. `workspaceTools(ws, { mode })` builds the definitions without touching the
@@ -859,6 +958,14 @@ A read of a sleeping workspace that its disk cannot answer (409 `workspace_not_r
859
958
  `host_feature_unavailable` (search or patches are not available for the workspace, `details.feature`) is neither
860
959
  retried nor woken: run `grep` with `exec`, or read then write the file, instead.
861
960
 
961
+ Immutable paths **(0.13.0+)**: a files write under an immutable path is 409 `conflict` with `reason` `read_only_path`
962
+ (not retried; write elsewhere, or change the template). A build is refused with 422 `validation_failed` and `reason`
963
+ `immutable_path_removed` when its recipe drops a path of the template's open version (`details.removed`; keep them),
964
+ `immutable_paths_unsupported_base` when its base cannot carry immutable paths (`details.base`; build on a newer version
965
+ of that base), or `invalid_path` (`details.field` `recipe.immutable[<i>]`). A build that fails on them has
966
+ `failure.code` `immutable_path_missing` (`failure.details.path` is not a directory in the built filesystem) or
967
+ `immutable_image_too_large`.
968
+
862
969
  ## Usage and overage
863
970
 
864
971
  `cloud.usage` reads the organization's usage (API keys see organization totals and their own project's workspaces):
@@ -990,7 +1097,7 @@ let ack = await cell.exec.input(session.session_id, 'hello\n', { offset: 0 });
990
1097
  ack = await cell.exec.input(session.session_id, '', { offset: ack.offset, close: true });
991
1098
  ```
992
1099
 
993
- Input frames are at most 64 KiB. The acknowledged offset counts bytes accepted into the pipe, not bytes consumed by the program. A blocked writer may receive a partial acknowledgement; continue from the returned offset. Retrying the identical last frame is safe. `close` sends EOF only once the entire frame is accepted. Do not combine `stdin_open` with the existing one-shot `stdin` option. The transport writes no separate input log or payload file; program output and full-state snapshots retain their normal persistence. Full-state suspend preserves the pipe. Older hosts/guests explicitly refuse this feature until upgraded (`exec_stdin` / `exec_stdin.v1`). Running sessions count as work within the documented idle-command window; long-running services should declare keepalives or choose `never`.
1100
+ Input frames are at most 64 KiB. The acknowledged offset counts bytes accepted into the pipe, not bytes consumed by the program. A blocked writer may receive a partial acknowledgement; continue from the returned offset. Retrying the identical last frame is safe. `close` sends EOF only once the entire frame is accepted. Do not combine `stdin_open` with the existing one-shot `stdin` option. The transport writes no separate input log or payload file; program output and full-state snapshots retain their normal persistence. Full-state suspend preserves the pipe. Where a workspace does not offer it, the start is refused with 409 `conflict`, `reason` `host_feature_unavailable` and `details.feature: 'exec_stdin'`, not retryable. Running sessions count as work within the documented idle-command window; long-running services should declare keepalives or choose `never`.
994
1101
 
995
1102
  ```ts
996
1103
  import { isWorkspaceGone, ShardfluxProtocolError } from '@shardflux/sdk';
package/dist/cell.d.ts CHANGED
@@ -22,6 +22,7 @@
22
22
  * have before sending them (NotSupportedForModeError).
23
23
  */
24
24
  import type { components, paths } from './generated/cell-api.js';
25
+ import { ShardfluxApiError } from './errors.js';
25
26
  import type { WorkspaceMode } from './errors.js';
26
27
  import { ExecutionResult } from './executions.js';
27
28
  import type { ExecutionGetOptions, ExecutionRunOptions } from './executions.js';
@@ -31,6 +32,24 @@ import type { ToolTokenManager } from './tokens.js';
31
32
  type S = components['schemas'];
32
33
  export type ExecStartRequest = S['ExecStartRequest'];
33
34
  export type ExecSession = S['ExecSession'];
35
+ /**
36
+ * Elastic memory (0.13.0): the host grew the workspace's memory before an exec started, and the exec
37
+ * waited for it. `outcome` delivered | partial | missed | failed; `from_mib`, `want_mib`, `got_mib` are guest memory
38
+ * before, aimed for and after; `deliver_ms` the time to deliver. On the session the start returned; absent when no
39
+ * grow ran (fixed workspaces, or already at the exec size).
40
+ */
41
+ export type MemoryGrow = S['MemoryGrow'];
42
+ /**
43
+ * Burst execution (0.13.0; needs an organization entitlement): the session of
44
+ * an exec started with `burst: 'always'` ran on a larger, short-lived burst VM on the workspace's host, and its file
45
+ * changes were applied back (`applied`). `host`, `vcpus`, `memory_mib`, `method`, the change counts
46
+ * (`written_files`, `written_dirs`, `removed`, `written_bytes`), `leftover_killed` (processes the command left behind:
47
+ * they do not survive a burst), `overhead_ms`, `excluded_paths`, `timings`, `replayed`; `error` when the burst failed
48
+ * after the start answered.
49
+ */
50
+ export type BurstSummary = S['BurstSummary'];
51
+ /** The failure recorded on a burst session (`burst.error`): code `burst_unavailable` or `burst_apply_failed`. */
52
+ export type BurstError = S['BurstError'];
34
53
  export type ExecInputResult = S['ExecInputResult'];
35
54
  export type OutputEvent = S['OutputEvent'];
36
55
  export type PtyOpenRequest = S['PtyOpenRequest'];
@@ -155,8 +174,9 @@ export interface CellClientOptions {
155
174
  sleep?: (ms: number) => Promise<void>;
156
175
  /**
157
176
  * Wakes a suspended workspace (resume, or join the active resume/open) within `timeoutMs` and resolves once it runs;
158
- * resolves `false` when there was nothing to wake (the API already reports it running), and the refusal then
159
- * surfaces. It throws when the wake fails (OperationFailedError) or outlasts `timeoutMs` (OperationTimeoutError).
177
+ * resolves `false` when there was nothing to wake (the API already reports it running): the refused call is then
178
+ * retried with a current token, since the transition that refused it ended meanwhile (0.13.1+). It throws when the
179
+ * wake fails (OperationFailedError) or outlasts `timeoutMs` (OperationTimeoutError).
160
180
  * Workspace.cell() supplies `workspace.wake()`. A call refused with `workspace_not_running` wakes the workspace and is
161
181
  * retried; a refused call was never executed, so the retry cannot duplicate it. `null` surfaces
162
182
  * the refusal instead. A wake that put a new token into this client's manager (the held resume;
@@ -223,6 +243,10 @@ export interface RunResult {
223
243
  session: ExecSession;
224
244
  /** Output-stream reconnections performed (gateway restarts, network drops). */
225
245
  reconnects: number;
246
+ /** The memory grow the start of this exec waited for (0.13.0; elastic workspaces), null when none ran. */
247
+ memoryGrow: MemoryGrow | null;
248
+ /** The burst summary (0.13.0; `burst: 'always'`), null for an ordinary exec. */
249
+ burst: BurstSummary | null;
226
250
  }
227
251
  export interface RunOptions {
228
252
  sessionId?: string;
@@ -254,7 +278,24 @@ export interface RunOptions {
254
278
  * (details.reason `secret_not_available`, details.names) and nothing runs; a name also present in `env` is 422.
255
279
  */
256
280
  secretRefs?: string[];
281
+ /**
282
+ * Burst execution (0.13.0; needs the organization entitlement `policy.burst_exec`): `'always'` runs this
283
+ * run-to-completion command on a larger, short-lived burst VM on the workspace's
284
+ * host over a copy of the workspace disk, and applies its file changes back when it exits (`result.burst`).
285
+ * Processes it starts and tmpfs content (/tmp when tmpfs, /dev/shm) do not survive; the workspace's own processes
286
+ * pause meanwhile and writes to it are refused (409). Not with `stdin` (422 burst_not_supported). Refusals: 409
287
+ * `burst_unavailable` (details.reason not_available, layout_unsupported, host_capacity, ...), 409
288
+ * `burst_apply_failed` (details.applied_entries/pending_entries; retry with the same `sessionId` to finish the apply).
289
+ * A burst cannot be canceled: `signal` stops following it, not the command.
290
+ */
291
+ burst?: 'never' | 'always';
292
+ /** Burst VM vCPUs (1-32; default the host's, 8, or the plan's ceiling when lower). */
293
+ burstVcpus?: number;
294
+ /** Burst VM memory in MiB (512-65536; default the host's, 8192, or the plan's ceiling when lower). */
295
+ burstMemoryMib?: number;
257
296
  }
297
+ /** The API error of a burst's recorded failure (`burst.error` of its session). Internal (the agent tools use it too). */
298
+ export declare function burstFailure(e: BurstError): ShardfluxApiError;
258
299
  /** Parses an NDJSON byte stream into objects (tolerates CRLF and a final unterminated line). */
259
300
  export declare function ndjson<T>(body: ReadableStream<Uint8Array>): AsyncGenerator<T>;
260
301
  export declare class CellClient {
@@ -280,8 +321,10 @@ export declare class CellClient {
280
321
  * One authorized request. Refreshes the token once on stale_epoch / 401. Lifecycle transitions,
281
322
  * bounded in total by `transitionTimeoutMs`: a call refused with `workspace_busy` is retried after `Retry-After`; one
282
323
  * refused with `workspace_not_running` (or whose token cannot be minted because the workspace is not running) wakes
283
- * the workspace through `wake`, given the time left, and is retried with a fresh token, at most 3 wakes per call.
284
- * Refused calls were never executed, so retrying is safe. When the budget is spent the refusal surfaces.
324
+ * the workspace through `wake`, given the time left, and is retried with a fresh token, at most 3 wakes per call. A
325
+ * wake that finds the workspace already running (`false`) is retried the same way: the transition that refused the
326
+ * call ended meanwhile (0.13.1+). Refused calls were never executed, so retrying is safe. When the budget is spent
327
+ * the refusal surfaces.
285
328
  */
286
329
  request(method: string, path: string, init?: RequestOptions & TransitionOptions): Promise<Response>;
287
330
  /**
@@ -308,6 +351,10 @@ export declare class CellClient {
308
351
  signal?: AbortSignal;
309
352
  }) => Promise<ExecInputResult>;
310
353
  signal: (sessionId: string, signal: Signal, onlyLeader?: boolean) => Promise<ExecSession>;
354
+ /**
355
+ * SIGTERM to the process group, SIGKILL after `graceMs` (0 to 60000; 0 or omitted is 5000). Resolves once the
356
+ * command has ended; the request waits for the grace.
357
+ */
311
358
  cancel: (sessionId: string, graceMs?: number) => Promise<ExecSession>;
312
359
  /**
313
360
  * Starts (or re-attaches to) a session and collects its output until it exits, reconnecting