@shardflux/sdk 0.11.1 → 0.13.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -3,6 +3,148 @@
3
3
  Every API the README shows is available from the version named here. Breaking changes ship in minor releases and are
4
4
  marked **Breaking**.
5
5
 
6
+ ## 0.13.0 (not yet published)
7
+
8
+ ### Elastic memory
9
+
10
+ Additive. A workspace may be promised more memory than it holds while idle, and the host grows it when a command needs
11
+ it (opt-in per workspace).
12
+
13
+ - `Caps.allocation_mode` (`'fixed' | 'elastic'`, type `AllocationMode`) and `Caps.memory_mib_held` on `open()`,
14
+ `workspaces.fork()` and `workspace.fork()`. Given caps replace the stored ones (caps without `allocation_mode` make
15
+ the workspace fixed); omitted caps keep the stored layout.
16
+ - `workspace.memory` (`WorkspaceMemory`: `{allocation_mode, promised_mib, held_mib, plugged_mib}`, null on an older
17
+ API) and `workspace.allocationMode` (`'fixed'` on an older API). `WorkspaceView` carries `caps.allocation_mode`,
18
+ `caps.memory_mib_held` and `memory` (regenerated contract types).
19
+ - `RunResult.memoryGrow` (`MemoryGrow | null`): the grow the exec's start waited for (`outcome` delivered, partial,
20
+ missed or failed; `from_mib`, `want_mib`, `got_mib`, `deliver_ms`), from the start's session. `ExecSession`
21
+ carries `memory_grow`. The `exec` agent tool adds `memory_grow` to its result when a grow ran.
22
+ - `KnownErrorReason` adds `allocation_mode_not_available`, `requires_elastic` and `exceeds_memory_mib` (422
23
+ `validation_failed`; elastic with file-first is the existing `not_supported_for_mode`).
24
+
25
+ ### Burst execution
26
+
27
+ Additive. One heavy, run-to-completion command can run on a larger, short-lived burst VM on the workspace's host, and
28
+ its file changes are applied back.
29
+
30
+ - `RunOptions.burst` (`'never' | 'always'`), `burstVcpus` and `burstMemoryMib` on `exec.run()`; `ExecStartRequest`
31
+ carries `burst`, `burst_vcpus` and `burst_memory_mib` (regenerated contract types).
32
+ - `RunResult.burst` (`BurstSummary | null`): host, size, method, whether the changes were applied, the change counts,
33
+ `leftover_killed`, `overhead_ms`, `excluded_paths`, `timings`, `replayed`. `ExecSession.burst` carries it; types
34
+ `BurstSummary` and `BurstError` are exported.
35
+ - `exec.run()` rejects with `ShardfluxApiError` 409 `burst_unavailable` or `burst_apply_failed` also when the burst
36
+ fails after its start (the output stream's final error event, or `burst.error` on the ended session), without
37
+ reconnecting; it does not try to cancel a burst when its `signal` aborts (a burst cannot be canceled).
38
+ - `ErrorCode` adds `burst_unavailable` and `burst_apply_failed`; `KnownErrorReason` adds `burst_mode_not_supported`,
39
+ `burst_not_supported`, `burst_size_exceeds_plan`, `not_available`, `shared_volumes`, `fence_not_drained`,
40
+ `apply_pending`, `park_failed`, `workspace_resumed`, `interrupted`, `burst_lost`, `disk_full`, `apply_failed`,
41
+ `reverted` and `revert_failed`.
42
+ - `workspaceTools(ws, { burst: true })` (opt-in, default off) gives the processful `exec` tool the inputs `burst`,
43
+ `burst_vcpus` and `burst_memory_mib`, and its result adds `burst`. The default tool schemas are unchanged.
44
+
45
+ ### Background commands in the agent tools
46
+
47
+ Additive. An agent can start a long command, keep working with the other tools, read its progress and stop it.
48
+
49
+ - The processful `exec` tool takes `background: true`: it starts the command and returns `{session_id, state}` at
50
+ once (a failed start adds `error: {code, message, reason}`, as the foreground `exec` gives it). `timeout_ms` goes up
51
+ to 86400000 (a day); with `background` it is sent only when given. The file-first `exec` is unchanged.
52
+ - New tool `exec_read` (permission `exec`): the session's `state`, `exit_code`, `term_signal`, `timed_out`,
53
+ `canceled` and output. Without offsets it returns the last `maxOutputBytes` of each stream; `stdout_offset` /
54
+ `stderr_offset` read from there, and the result's `next_stdout_offset` / `next_stderr_offset` continue where it
55
+ stopped (`truncated` when more output follows). Text is cut at character boundaries only. `wait_ms` (up to 60000)
56
+ waits for the command to exit and returns as soon as it does. `burst` and `error` when the session has them.
57
+ - New tool `exec_cancel` (permission `exec`): SIGTERM, then SIGKILL after `grace_ms` (1-60000, default 5000);
58
+ returns `{session_id, state, exit_code, canceled}`.
59
+ - With `workspaceTools(ws, { burst: true })`, `background: true` together with `burst: 'always'` is refused with
60
+ `ToolArgumentError` before any request (a burst holds the workspace until it applies its changes); `burst: 'never'`
61
+ with `background` is an ordinary background start.
62
+ - `exec_read` and `exec_cancel` follow `exec` in `workspaceTools()` on processful workspaces; a file-first workspace
63
+ does not offer them. Both send the wake hint like the other tools.
64
+
65
+ ### Held fork
66
+
67
+ Additive. A waited fork returns the running copy and its tool token in one request.
68
+
69
+ - `workspaces.fork()` and `workspace.fork()` with `wait`: when the API holds the fork (`Prefer: wait`, answered 200
70
+ with `Preference-Applied`), the copy's handle adopts the returned running view and final-epoch tool token, with no
71
+ operation poll, view refresh or token request. The request's time counts against `wait.timeoutMs`.
72
+ - `ForkOptions` and `WaitedForkOptions` (exported) add `agentLabel` and `tools`, which choose that token. They are sent
73
+ only on a server-held fork and need an API with held fork. `wait: { serverWait: false }` keeps polling.
74
+ - An API that does not apply `Prefer` answers 202 as before; the call then polls the operation and refreshes the copy.
75
+ - The fork's progress `request` phase has reason `held` when it asks the server to wait.
76
+
77
+ ### Exec start and output in one request
78
+
79
+ - `exec.run()` starts the command and follows its output in one request: the start asks for NDJSON
80
+ (`Accept: application/x-ndjson`) and a cell that supports it answers with the output stream. A dropped stream
81
+ reconnects by session and byte offsets, never starting the command again. A cell that answers the start with the
82
+ session JSON gets the output request after it, as before; a start that failed still rejects with `ExecStartError`
83
+ (or the burst's error).
84
+ - `RunResult.memoryGrow` and `burst` are unchanged: the combined stream's exit event carries the start's
85
+ `memory_grow`.
86
+
87
+ ### Immutable paths
88
+
89
+ A template may declare directories as immutable: every workspace of the template mounts them read-only from the
90
+ template's newest published version, picked up at its next cold boot or resume.
91
+
92
+ - Recipe v2 `immutable` (`TemplateRecipeV2.immutable`, regenerated): the list in `template.yaml`, sent as written by
93
+ `buildFromFile()`, `buildFromRecipe()` and `builds.create()`. The stored recipe v2 of a build
94
+ (`TemplateBuildRecipeV2`) and an exported recipe (`versions.recipe()`) carry it.
95
+ - `TemplateVersion.immutable` (`TemplateVersionImmutable`: `{paths, bytes}`, or null when the version declares none),
96
+ on every version of `templates.get()` and on `open_version`.
97
+ - `TemplateBuild.immutable_paths`: the list the produced version declares (the recipe's, else the open version's).
98
+ `TemplateBuild.failure.details` carries code-specific fields: `immutable_path_missing` names `details.path`.
99
+ - `workspace.immutableVersion` (`template.immutable_version` of the view): the template version whose immutable paths
100
+ the workspace has mounted, which can be newer than `template.version`; null when it mounts none.
101
+ - `KnownErrorReason` adds `immutable_path_removed` (422; `details.removed`, `details.open_version`),
102
+ `immutable_paths_unsupported_base` (422; `details.base`, `details.required_feature`, `details.paths`) and
103
+ `read_only_path` (409 `conflict` from the files API for a write under an immutable path).
104
+ - **Breaking:** the update policy left the API. Removed: the `UpdatePolicy` type, `TemplateSummary.update_policy`, the
105
+ regenerated `update_policy` of the workspace view, of `TemplateDefaults` and of the `defaults` inputs, and the reason
106
+ `update_policy_not_available`. It was always `pinned`; code that read it drops the read.
107
+
108
+ ### Exec cancel grace
109
+
110
+ - `exec.cancel(id, graceMs?)` takes `graceMs` 0 to 60000 (0 or omitted: 5000), as the cell API now enforces (422
111
+ `validation_failed` outside it), and waits for the answer up to the grace plus 15 s when that is longer than the
112
+ client's `timeoutMs`: a command that ignores SIGTERM is answered after the grace and the SIGKILL, not a timeout.
113
+ - The cancel that `exec.run` sends when its `signal` aborts passes `killGraceMs` capped at 60000, so a larger
114
+ `killGraceMs` still cancels the command.
115
+
116
+ ## 0.12.0 (release candidate)
117
+
118
+ Completed exec output streams are drained before releasing their HTTP connections, with a bounded cleanup if a peer does not close.
119
+
120
+ Pooled HTTP/1.1 on Node 26 with pinned undici 8.10.2; custom fetch stays unchanged. Labels on open/list and setLabels, typed idle(), keepalive() and setIdlePolicy(). Protocol errors carry source; isWorkspaceGone recognizes only an explicit API workspace_deleted refusal. Exec sessions accept stdin_open and offset-addressed exec.input. Failed workspaces recover through resume/auto-wake with the same ID. Completed deletions free keys for new IDs.
121
+
122
+ Requires the QM integration backend release for labels, failed recovery, key reuse and pipe stdin. No production deployment has occurred from this branch.
123
+
124
+ ### Instant suspend: durable storage in the result
125
+
126
+ Additive. A suspend returns as soon as the workspace is sealed on its host; the copy lands in durable storage right
127
+ after. This release reads that from the result and can wait for it.
128
+
129
+ - `suspend({ durable: true })` (`SuspendOptions.durable`, also `cloud.workspaces.suspend(id, { durable: true })`):
130
+ resolves once `result.durable` is `true`. It implies `wait` and uses the same `timeoutMs`/`signal` (one budget for
131
+ the suspend and the copy). Rejects with `DurabilityLostError` (new; `durability` with the reason) when the copy
132
+ cannot be made, and `OperationTimeoutError` with `durable: true` when the time runs out (the copy continues).
133
+ `suspend({ wait: true })` still resolves when the suspend succeeds, without waiting for the copy.
134
+ - `cloud.workspaces.waitForDurable(operationOrId, waitOptions)`: the same wait for an operation you hold (a suspend, or
135
+ a fork of a running workspace).
136
+ - `isDurable(op)`, `durabilityOf(op)` (`Durability`: `state` `pending` | `durable` | `lost`, `checkpointId`,
137
+ `generationId`, `localCommitAt`, `durableBy`, `durableAt`, `localCommitToDurableMs`, `overdueAt`, `reason`) and
138
+ `lostSuspendOf(op)` (`LostSuspend`, on a resume result). `ServerTiming` gains `durable`, `suspendPath`, `durability`
139
+ and `lostSuspend`; progress phase `durable` (reason `overdue` past `durableBy`); `formatTiming()` prints them.
140
+ - `Operation.result` is typed with `durable`, `suspend_path`, `durability` and `lost_suspend` (OpenAPI).
141
+ - `OperationFailedError.workspaceActive`: a suspend-when-idle canceled with `workspace_active` (the workspace was in
142
+ use; nothing changed and it keeps running). `workspace.waitUntilReady()` resolves for it instead of throwing, as
143
+ `wake()` already did.
144
+
145
+ Requires the instant-suspend cell release for `durable: false` results; against other servers every succeeded suspend
146
+ is already durable and `durable: true` resolves at once.
147
+
6
148
  ## 0.11.1
7
149
 
8
150
  Wording: messages and JSDoc say what to do, without internals (no behaviour change).
package/README.md CHANGED
@@ -14,7 +14,7 @@ tool calls into the workspace ([tool-call capture](#tool-call-capture-070)).
14
14
  > **(0.7.0+)** not in 0.6.x and **(0.6.0+)** not in 0.5.0; [CHANGELOG.md](./CHANGELOG.md) lists what each version added. Check yours with
15
15
  > `npm ls @shardflux/sdk` or the exported `SDK_VERSION`.
16
16
 
17
- - ESM only, no runtime dependencies, Node.js 24 or later. Reading a YAML template file uses the optional peer
17
+ - ESM only, Node.js 24 or later. Reading a YAML template file uses the optional peer
18
18
  dependency `yaml` (`npm install yaml`); JSON template files need nothing.
19
19
  - Typed from the published OpenAPI documents.
20
20
  - Retries, idempotency keys, operation polling and tool-token refresh are handled for you.
@@ -108,8 +108,9 @@ const listing = await cell.files.list('/home/user');
108
108
  await cell.files.remove('/home/user/data.bin');
109
109
  ```
110
110
 
111
- `exec.run()` resumes from byte offsets if the output stream drops; it never starts the command
112
- twice. Aborting its `signal` also cancels the command in the workspace. File writes are atomic
111
+ `exec.run()` starts the command and streams its output in one request **(0.13.0+)**, resumes from
112
+ byte offsets if the output stream drops, and never starts the command twice. Aborting its `signal`
113
+ also cancels the command in the workspace. File writes are atomic
113
114
  and durable (acknowledged after fsync).
114
115
 
115
116
  `cwd` is an absolute path (commands start in `/home/user` without one). The API refuses a relative
@@ -187,7 +188,7 @@ await workspace.resume({ wait: true }); // or simply open() the key again
187
188
 
188
189
  const { workspace: copy } = await workspace.fork({ key: 'customer-42/experiment' }, { wait: true });
189
190
 
190
- await copy.delete(); // tool access ends immediately; keys are never reused
191
+ await copy.delete(); // tool access ends immediately; the key can be reused after deletion finishes
191
192
  ```
192
193
 
193
194
  ### Requested or finished
@@ -199,12 +200,40 @@ means depends on `wait`:
199
200
  | --- | --- | --- |
200
201
  | `await workspace.suspend()` | the suspend is **requested** (usually `queued`; the workspace is still running) | the operation (`Operation`) |
201
202
  | `await workspace.suspend({ wait: true })` **(0.6.0+)** | the suspend has **finished** (`workspace.state` is then `suspended`) | the succeeded operation (`FinishedOperation`) |
203
+ | `await workspace.suspend({ durable: true })` **(0.12.0+)** | the suspend has finished **and** its copy is in durable storage (`result.durable` is `true`) | the succeeded operation (`FinishedOperation`) |
202
204
 
203
205
  `wait` also takes `WaitOptions` (`timeoutMs`, default 5 minutes; `signal`; `onProgress`). A failed operation throws
204
206
  `OperationFailedError`. Running out of time throws `OperationTimeoutError`, and the operation continues server side:
205
207
  wait again with `cloud.workspaces.waitForOperation(err.operationId)`. Without `wait`, the returned operation is the
206
208
  handle for the work in progress: pass its `id` to `waitForOperation()` when you need it finished.
207
209
 
210
+ **A waited fork is one request (0.13.0+).** `fork(target, { wait: true })` asks the API to hold the fork until the
211
+ copy runs, and the answer carries the running copy and a tool token for it: the copy's handle is ready with no
212
+ operation poll, view refresh or token request. `agentLabel` and `tools` pick the token it brings back. Its timing is one
213
+ `request` phase with reason `held`; `wait: { serverWait: false }` polls instead. An API without the held fork answers
214
+ at once, and the SDK then waits for the operation and refreshes the copy, as before.
215
+
216
+ **Durable storage (0.12.0+).** A suspend returns as soon as the workspace is sealed on its host, typically in a few
217
+ hundred ms, and its RAM and CPU are released at that moment. `result.durable` turns `true` when the copy lands in
218
+ durable storage, typically within a second; until then it is `false` and `result.durability` shows the copy's
219
+ progress. Most code needs nothing more: a suspended workspace resumes, is read and is forked the same way either way.
220
+ When your code must know the copy is durable (before deleting a local artifact, or at the end of a job), ask for it:
221
+
222
+ ```ts
223
+ import { durabilityOf } from '@shardflux/sdk';
224
+
225
+ const op = await workspace.suspend({ durable: true }); // resolves once result.durable is true
226
+ durabilityOf(op); // { state: 'durable', checkpointId, localCommitAt, durableAt, localCommitToDurableMs, ... }
227
+ ```
228
+
229
+ - `durable: true` implies `wait` and shares its `timeoutMs` and `signal` (e.g. `{ durable: true, wait: { timeoutMs: 60_000 } }`).
230
+ If the time runs out first, `OperationTimeoutError` has `durable: true` and the copy continues server side.
231
+ - `cloud.workspaces.waitForDurable(operationOrId)` does the same for an operation you already hold, such as a
232
+ finished fork of a running workspace, whose result carries `durable` and `durability` the same way.
233
+ - `isDurable(op)`, `durabilityOf(op)` and `lastTiming.server.durable` / `.durability` read the fields; `formatTiming()`
234
+ prints `sealed on host, durable 435 ms later`. Results from before 0.12.0 servers count as durable.
235
+ - The errors reference on docs.shardflux.dev lists what `durable: true` can reject with.
236
+
208
237
  **Start deadlines.** A queued open, resume, restore or fork (`capacity_pending`) has a deadline 15 minutes after it
209
238
  was created. Its `error.details.deadline_at` carries the deadline; the `phase` progress event carries it as
210
239
  `deadlineAt` **(0.6.2+)**, and so does `OperationTimeoutError` when your wait ends first. A start still pending at the
@@ -295,8 +324,7 @@ operation keeps running server side (a queued start until its `deadlineAt`); wai
295
324
  same call. `templates.builds.waitForBuild()` waits the same way. A waited `open()` issues the first tool token
296
325
  together with the final workspace read, so the first tool call starts at once.
297
326
 
298
- On Node 26 the default fetch uses a fresh connection for every request (`Connection: close`);
299
- `SHARDFLUX_HTTP_KEEPALIVE=1`, or your own `fetch`, reuses connections.
327
+ From 0.12.0, Node 26 reuses TLS connections through a private HTTP/1.1 pool. Other runtimes keep native fetch. A supplied `fetch` stays unchanged. `SHARDFLUX_HTTP_KEEPALIVE=0` forces connection closure for diagnosis; `=1` opts into native pooling. No global dispatcher is installed.
300
328
 
301
329
  List and look up workspaces:
302
330
 
@@ -313,6 +341,46 @@ sessions, template drafts and test instances. `findByKey()` searches every lifet
313
341
  workspace, and a tombstone only when no live workspace has the key. Use it rather than `listAll({ keyPrefix })` for
314
342
  lookups by key.
315
343
 
344
+ ### Elastic memory (0.13.0+)
345
+
346
+ Promise a workspace a lot of memory while it holds only what it uses. An elastic workspace idles at its held floor
347
+ (`memory_mib_held`, default 1024) and grows before each command starts, so compilers, test runners and V8 size
348
+ themselves from the memory they will actually get. After 30 s without work it shrinks back. The layout takes effect at
349
+ the workspace's next start.
350
+
351
+ ```ts
352
+ const ws = await cloud.workspaces.open({
353
+ key: 'customer-42/repo-7',
354
+ template: 'node',
355
+ caps: { memory_mib: 8192, allocation_mode: 'elastic' },
356
+ });
357
+ ws.memory; // { allocation_mode: 'elastic', promised_mib: 8192, held_mib: 1024, plugged_mib: 3072 }
358
+
359
+ const r = await ws.cell().exec.run(['npm', 'test']);
360
+ r.memoryGrow; // { outcome: 'delivered', from_mib: 1024, want_mib: 4096, got_mib: 4096, deliver_ms: 175 }
361
+ ```
362
+
363
+ Given `caps` replace the stored ones, so send `allocation_mode` with every `caps` you pass; an open without `caps`
364
+ keeps the stored layout. The exec agent tool adds `memory_grow` to its result when the command's start grew the VM.
365
+
366
+ ### Burst execution (0.13.0+)
367
+
368
+ Run one heavy command, such as a cold build or a full test suite, on a larger VM without resizing the workspace. With
369
+ `burst: 'always'` the command runs on a short-lived burst VM over a copy of the workspace, output streams as usual,
370
+ and its file changes are applied back byte for byte when it exits.
371
+
372
+ ```ts
373
+ const r = await ws.cell().exec.run(['go', 'build', './...'], { burst: 'always', burstVcpus: 16 });
374
+ r.exitCode; // 0
375
+ r.burst; // { host: 'local', vcpus: 16, memory_mib: 8192, method: 'layer', applied: true,
376
+ // written_files: 3, written_dirs: 1, removed: 1, written_bytes: 40960, overhead_ms: 641, ... }
377
+ ```
378
+
379
+ A burst carries back the command's file changes; processes it started and tmpfs content stay in the burst VM, and the
380
+ workspace's own processes resume where they were. A burst is not signalled or canceled. If applying the changes
381
+ fails (`burst_apply_failed`), retry `exec.run` with the same `sessionId` to finish the apply. The agent tools offer
382
+ bursts when asked: `workspaceTools(ws, { burst: true })`.
383
+
316
384
  ### Timing and progress (0.6.0+)
317
385
 
318
386
  Every open, wake and waited lifecycle call is traced. `workspace.lastTiming` (and `err.timing` when the call fails)
@@ -548,6 +616,27 @@ workspace keeps running so you can inspect it, and `workspace.startup` names the
548
616
  The next open runs the failed step again. Versions report their `settings`, platform templates their `category`
549
617
  (`os` or `stack`), builds the `denied_hosts` their build network refused (add them to `build.network.extra_hosts`).
550
618
 
619
+ ### Immutable paths (0.13.0+)
620
+
621
+ Directories a template declares `immutable` follow the template: every workspace of it mounts them read-only from the
622
+ template's newest published version, so the tools, models or data you ship there reach existing workspaces with your
623
+ next version, at their next cold boot or resume. Everything else in a workspace stays its own.
624
+
625
+ ```yaml
626
+ # template.yaml
627
+ immutable: [/opt/acme] # read-only in every workspace; follows the newest version
628
+ ```
629
+
630
+ ```ts
631
+ const t = await cloud.templates.get('acme-dev');
632
+ console.log(t.open_version?.immutable); // { paths: ['/opt/acme'], bytes: 100663296 }
633
+ const ws = await cloud.workspaces.open({ key: 'customer-42/main', template: 'acme-dev' });
634
+ console.log(ws.template.version, ws.immutableVersion); // created from v1; /opt/acme shows v2
635
+ ```
636
+
637
+ A build without `immutable` keeps the list of the template's open version, and each new version keeps every path of
638
+ it: add paths, never remove them. A build reports the list its version declares in `immutable_paths`.
639
+
551
640
  ## Agent tools
552
641
 
553
642
  `workspaceTools(workspace)` returns tools with a name, a description, a JSON Schema for the
@@ -582,6 +671,20 @@ instead of being overwritten. Each call first sends `workspace.hint()` without w
582
671
  that off, e.g. when you send the hint yourself as the model starts a tool call), except `read_file`, `list_files` and
583
672
  `search_files`: a sleeping workspace answers them from its disk without waking.
584
673
 
674
+ **Background commands (0.13.0+).** On a processful workspace `exec` takes `background: true` for anything that runs
675
+ longer than a few minutes (a build, a test suite, a training run, a server): it returns `{ session_id, state }` at
676
+ once and the model keeps using the other tools. `exec_read` reads the command's state, exit code and output (the end
677
+ of each stream, or from the `next_stdout_offset` / `next_stderr_offset` of the previous read; `wait_ms` waits for the
678
+ exit and returns as soon as it happens), and `exec_cancel` stops it. `timeout_ms` (up to a day) sets the longest the
679
+ command may run; the workspace stays awake while it does. A burst (`burst: 'always'`) runs in the foreground.
680
+
681
+ ```ts
682
+ const tools = workspaceTools(workspace);
683
+ const { session_id } = (await executeToolCall(tools, { name: 'exec', input: { command: 'make -j8 test', background: true, timeout_ms: 7_200_000 } })) as { session_id: string };
684
+ // ... other tool calls ...
685
+ const progress = await executeToolCall(tools, { name: 'exec_read', input: { session_id, wait_ms: 30_000 } });
686
+ ```
687
+
585
688
  For a file-first workspace (0.9.0+) the tools are `exec` and the files tools only: `exec` runs each command as an
586
689
  execution and adds `execution_id`, `state`, `tree_revision` and `changed` to its result, and the process, terminal,
587
690
  git and browser tools are not offered. `workspaceTools(ws, { mode })` builds the definitions without touching the
@@ -838,6 +941,14 @@ A read of a sleeping workspace that its disk cannot answer (409 `workspace_not_r
838
941
  `host_feature_unavailable` (search or patches are not available for the workspace, `details.feature`) is neither
839
942
  retried nor woken: run `grep` with `exec`, or read then write the file, instead.
840
943
 
944
+ Immutable paths **(0.13.0+)**: a files write under an immutable path is 409 `conflict` with `reason` `read_only_path`
945
+ (not retried; write elsewhere, or change the template). A build is refused with 422 `validation_failed` and `reason`
946
+ `immutable_path_removed` when its recipe drops a path of the template's open version (`details.removed`; keep them),
947
+ `immutable_paths_unsupported_base` when its base cannot carry immutable paths (`details.base`; build on a newer version
948
+ of that base), or `invalid_path` (`details.field` `recipe.immutable[<i>]`). A build that fails on them has
949
+ `failure.code` `immutable_path_missing` (`failure.details.path` is not a directory in the built filesystem) or
950
+ `immutable_image_too_large`.
951
+
841
952
  ## Usage and overage
842
953
 
843
954
  `cloud.usage` reads the organization's usage (API keys see organization totals and their own project's workspaces):
@@ -941,3 +1052,44 @@ or no longer supported, it emits one warning:
941
1052
  ## License
942
1053
 
943
1054
  Apache-2.0
1055
+
1056
+
1057
+ ## Workspace integrations (0.12.0+)
1058
+
1059
+ ```ts
1060
+ const ws = await cloud.workspaces.open({ key: 'qm/project-42', template: 'python-node-browser',
1061
+ labels: { scope: 'project-42', owner: 'qm' }, idlePolicy: 'never' });
1062
+ const matches = await cloud.workspaces.list({ labels: { scope: 'project-42' } });
1063
+ await ws.setLabels({ scope: 'project-42', owner: 'qm' }); // replaces all labels; {} clears
1064
+ await ws.setIdlePolicy('adaptive'); // null restores the inherited policy
1065
+ const idle = await ws.idle(); // read-only, does not wake or record activity
1066
+ await ws.keepalive(60); // seconds; never shortens an existing keepalive
1067
+ await ws.suspendWhenIdle({ afterSeconds: 2 }); // one safe server-side request; no /idle + suspend race
1068
+ ```
1069
+
1070
+ Labels are exact string pairs (at most 50, keys 1–64 and values 0–256 characters); they are metadata, not secrets.
1071
+ A failed workspace is recoverable: `await ws.wake()` and ordinary auto-waking calls restart that ID using its existing disk. They do not create a replacement workspace. A failed attempt still raises its operation error; there is no unbounded restart loop.
1072
+ After `await ws.delete({ wait: true })`, opening its key creates a **new ID**. The old ID, tokens and history remain deleted. A key stays reserved while deletion is in progress.
1073
+
1074
+ For a background command with a pipe (no PTY):
1075
+
1076
+ ```ts
1077
+ const cell = ws.cell();
1078
+ const session = await cell.exec.start({ session_id: 'qm-worker-1', argv: ['cat'], stdin_open: true });
1079
+ let ack = await cell.exec.input(session.session_id, 'hello\n', { offset: 0 });
1080
+ ack = await cell.exec.input(session.session_id, '', { offset: ack.offset, close: true });
1081
+ ```
1082
+
1083
+ Input frames are at most 64 KiB. The acknowledged offset counts bytes accepted into the pipe, not bytes consumed by the program. A blocked writer may receive a partial acknowledgement; continue from the returned offset. Retrying the identical last frame is safe. `close` sends EOF only once the entire frame is accepted. Do not combine `stdin_open` with the existing one-shot `stdin` option. The transport writes no separate input log or payload file; program output and full-state snapshots retain their normal persistence. Full-state suspend preserves the pipe. Where a workspace does not offer it, the start is refused with 409 `conflict`, `reason` `host_feature_unavailable` and `details.feature: 'exec_stdin'`, not retryable. Running sessions count as work within the documented idle-command window; long-running services should declare keepalives or choose `never`.
1084
+
1085
+ ```ts
1086
+ import { isWorkspaceGone, ShardfluxProtocolError } from '@shardflux/sdk';
1087
+ // isWorkspaceGone(error) is true only for an explicit API workspace_deleted refusal.
1088
+ // A protocol error, a cell 404, a scoped API 404 or a network failure is never deletion evidence.
1089
+ ```
1090
+
1091
+ `ShardfluxProtocolError.source` is `api` or `cell` for SDK responses (`unknown` only for an error constructed without a source by older caller code). It remains a protocol failure even when `status === 404`. Never delete local data based on an HTTP status alone.
1092
+
1093
+ ### Repositories with a minimum release age
1094
+
1095
+ Keep your repository's age policy. Pin an exact version that has aged enough; do not exempt the entire `@shardflux/*` scope. The first SDK, 0.5.0, was published September 26, 2026 at 21:13:52 UTC and reaches seven days on October 3 at that time. New versions, including 0.12.0, need their own seven days after publication. Before then no registry setting on our side can make them eligible. `npm view @shardflux/sdk time --json` shows publication times. A targeted exception for an inspected exact release is a repository-owner decision, not an installation requirement we bypass. Older versions do not contain this release's fixes.
package/dist/account.d.ts CHANGED
@@ -102,7 +102,7 @@ export interface ShardfluxAccountOptions {
102
102
  sessionToken?: string;
103
103
  /** Default https://api.shardflux.dev. */
104
104
  baseUrl?: string;
105
- /** Default: the runtime's fetch, with `Connection: close` on Node 26 (see defaultFetch in http.ts). */
105
+ /** Default: pooled HTTP/1.1 on Node 26+, native fetch on other runtimes (see defaultFetch in http.ts). */
106
106
  fetch?: typeof fetch;
107
107
  userAgent?: string;
108
108
  /** Per-request timeout (ms), default 30 s. */
package/dist/account.js CHANGED
@@ -571,7 +571,7 @@ export class ShardfluxAccount {
571
571
  const b = body;
572
572
  if (b && typeof b === 'object' && typeof b.session_token === 'string') {
573
573
  if (!isSessionToken(b.session_token))
574
- throw new ShardfluxProtocolError('the API returned a session_token that is not a CLI session token (sfu_...)', 200);
574
+ throw new ShardfluxProtocolError('the API returned a session_token that is not a CLI session token (sfu_...)', 200, 'api');
575
575
  this.#token = b.session_token;
576
576
  const expiresAt = typeof b.session_expires_at === 'string' ? b.session_expires_at : '';
577
577
  await this.#onSessionToken?.({ token: b.session_token, expiresAt });
package/dist/cell.d.ts CHANGED
@@ -22,6 +22,7 @@
22
22
  * have before sending them (NotSupportedForModeError).
23
23
  */
24
24
  import type { components, paths } from './generated/cell-api.js';
25
+ import { ShardfluxApiError } from './errors.js';
25
26
  import type { WorkspaceMode } from './errors.js';
26
27
  import { ExecutionResult } from './executions.js';
27
28
  import type { ExecutionGetOptions, ExecutionRunOptions } from './executions.js';
@@ -31,6 +32,25 @@ import type { ToolTokenManager } from './tokens.js';
31
32
  type S = components['schemas'];
32
33
  export type ExecStartRequest = S['ExecStartRequest'];
33
34
  export type ExecSession = S['ExecSession'];
35
+ /**
36
+ * Elastic memory (0.13.0): the host grew the workspace's memory before an exec started, and the exec
37
+ * waited for it. `outcome` delivered | partial | missed | failed; `from_mib`, `want_mib`, `got_mib` are guest memory
38
+ * before, aimed for and after; `deliver_ms` the time to deliver. On the session the start returned; absent when no
39
+ * grow ran (fixed workspaces, or already at the exec size).
40
+ */
41
+ export type MemoryGrow = S['MemoryGrow'];
42
+ /**
43
+ * Burst execution (0.13.0; needs an organization entitlement): the session of
44
+ * an exec started with `burst: 'always'` ran on a larger, short-lived burst VM on the workspace's host, and its file
45
+ * changes were applied back (`applied`). `host`, `vcpus`, `memory_mib`, `method`, the change counts
46
+ * (`written_files`, `written_dirs`, `removed`, `written_bytes`), `leftover_killed` (processes the command left behind:
47
+ * they do not survive a burst), `overhead_ms`, `excluded_paths`, `timings`, `replayed`; `error` when the burst failed
48
+ * after the start answered.
49
+ */
50
+ export type BurstSummary = S['BurstSummary'];
51
+ /** The failure recorded on a burst session (`burst.error`): code `burst_unavailable` or `burst_apply_failed`. */
52
+ export type BurstError = S['BurstError'];
53
+ export type ExecInputResult = S['ExecInputResult'];
34
54
  export type OutputEvent = S['OutputEvent'];
35
55
  export type PtyOpenRequest = S['PtyOpenRequest'];
36
56
  export type PtySession = S['PtySession'];
@@ -50,6 +70,8 @@ export type FilePatchResult = S['FilePatchResult'];
50
70
  /** A content SHA-256 (lowercase hex), or `absent` for a path that must not exist. */
51
71
  export type FileRevision = S['FileRevision'];
52
72
  export type WakeHintResult = S['WakeHintResult'];
73
+ export type IdleStatus = S['IdleStatus'];
74
+ export type KeepaliveResult = S['KeepaliveResult'];
53
75
  /** Where a running workspace's VM is: resident, frozen, hibernated or restoring. */
54
76
  export type Residency = WakeHintResult['residency'];
55
77
  /** `X-Served-From`: `disk` when a read was served from a suspended or hibernated workspace's disk. */
@@ -220,6 +242,10 @@ export interface RunResult {
220
242
  session: ExecSession;
221
243
  /** Output-stream reconnections performed (gateway restarts, network drops). */
222
244
  reconnects: number;
245
+ /** The memory grow the start of this exec waited for (0.13.0; elastic workspaces), null when none ran. */
246
+ memoryGrow: MemoryGrow | null;
247
+ /** The burst summary (0.13.0; `burst: 'always'`), null for an ordinary exec. */
248
+ burst: BurstSummary | null;
223
249
  }
224
250
  export interface RunOptions {
225
251
  sessionId?: string;
@@ -251,7 +277,24 @@ export interface RunOptions {
251
277
  * (details.reason `secret_not_available`, details.names) and nothing runs; a name also present in `env` is 422.
252
278
  */
253
279
  secretRefs?: string[];
280
+ /**
281
+ * Burst execution (0.13.0; needs the organization entitlement `policy.burst_exec`): `'always'` runs this
282
+ * run-to-completion command on a larger, short-lived burst VM on the workspace's
283
+ * host over a copy of the workspace disk, and applies its file changes back when it exits (`result.burst`).
284
+ * Processes it starts and tmpfs content (/tmp when tmpfs, /dev/shm) do not survive; the workspace's own processes
285
+ * pause meanwhile and writes to it are refused (409). Not with `stdin` (422 burst_not_supported). Refusals: 409
286
+ * `burst_unavailable` (details.reason not_available, layout_unsupported, host_capacity, ...), 409
287
+ * `burst_apply_failed` (details.applied_entries/pending_entries; retry with the same `sessionId` to finish the apply).
288
+ * A burst cannot be canceled: `signal` stops following it, not the command.
289
+ */
290
+ burst?: 'never' | 'always';
291
+ /** Burst VM vCPUs (1-32; default the host's, 8, or the plan's ceiling when lower). */
292
+ burstVcpus?: number;
293
+ /** Burst VM memory in MiB (512-65536; default the host's, 8192, or the plan's ceiling when lower). */
294
+ burstMemoryMib?: number;
254
295
  }
296
+ /** The API error of a burst's recorded failure (`burst.error` of its session). Internal (the agent tools use it too). */
297
+ export declare function burstFailure(e: BurstError): ShardfluxApiError;
255
298
  /** Parses an NDJSON byte stream into objects (tolerates CRLF and a final unterminated line). */
256
299
  export declare function ndjson<T>(body: ReadableStream<Uint8Array>): AsyncGenerator<T>;
257
300
  export declare class CellClient {
@@ -296,7 +339,19 @@ export declare class CellClient {
296
339
  follow?: boolean;
297
340
  signal?: AbortSignal;
298
341
  }) => Promise<AsyncGenerator<OutputEvent>>;
342
+ /** Write at the acknowledged offset (initially 0); a repeated identical last frame cannot duplicate input.
343
+ * At most 64 KiB per call. A partial acknowledgement requires continuing from the returned offset.
344
+ * Start with stdin_open: true. close sends EOF after this frame is fully accepted. */
345
+ input: (sessionId: string, data: string | Uint8Array, opts: {
346
+ offset: number;
347
+ close?: boolean;
348
+ signal?: AbortSignal;
349
+ }) => Promise<ExecInputResult>;
299
350
  signal: (sessionId: string, signal: Signal, onlyLeader?: boolean) => Promise<ExecSession>;
351
+ /**
352
+ * SIGTERM to the process group, SIGKILL after `graceMs` (0 to 60000; 0 or omitted is 5000). Resolves once the
353
+ * command has ended; the request waits for the grace.
354
+ */
300
355
  cancel: (sessionId: string, graceMs?: number) => Promise<ExecSession>;
301
356
  /**
302
357
  * Starts (or re-attaches to) a session and collects its output until it exits, reconnecting
@@ -442,6 +497,10 @@ export declare class CellClient {
442
497
  * workspace answers 409 `workspace_not_running` (`Workspace.hint()` then wakes it in the background).
443
498
  */
444
499
  wakeHint(signal?: AbortSignal): Promise<WakeHintResult>;
500
+ /** Read idle signals without recording activity or waking the workspace. */
501
+ idle(signal?: AbortSignal): Promise<IdleStatus>;
502
+ /** Declare work for seconds (1..the server maximum). Never shortens a previous keepalive. */
503
+ keepalive(seconds: number, signal?: AbortSignal): Promise<KeepaliveResult>;
445
504
  /**
446
505
  * One page of the workspace's changes against its template (needs the `files` tool). File-first workspaces: each
447
506
  * execution's result lists what it changed (`changed`); this route is refused (NotSupportedForModeError).