@shardflux/sdk 0.11.1 → 0.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +142 -0
- package/README.md +158 -6
- package/dist/account.d.ts +1 -1
- package/dist/account.js +1 -1
- package/dist/cell.d.ts +59 -0
- package/dist/cell.js +108 -18
- package/dist/client.d.ts +55 -10
- package/dist/client.js +105 -9
- package/dist/errors.d.ts +54 -6
- package/dist/errors.js +50 -5
- package/dist/feedback.js +1 -1
- package/dist/generated/app-api.d.ts +553 -54
- package/dist/generated/cell-api.d.ts +233 -9
- package/dist/http.d.ts +9 -4
- package/dist/http.js +39 -16
- package/dist/index.d.ts +12 -7
- package/dist/index.js +4 -2
- package/dist/lifecycle.d.ts +24 -1
- package/dist/lifecycle.js +10 -3
- package/dist/progress.d.ts +62 -1
- package/dist/progress.js +61 -0
- package/dist/templates.d.ts +25 -3
- package/dist/tools.d.ts +7 -0
- package/dist/tools.js +275 -7
- package/dist/workspace.d.ts +39 -10
- package/dist/workspace.js +46 -3
- package/package.json +4 -1
package/CHANGELOG.md
CHANGED
|
@@ -3,6 +3,148 @@
|
|
|
3
3
|
Every API the README shows is available from the version named here. Breaking changes ship in minor releases and are
|
|
4
4
|
marked **Breaking**.
|
|
5
5
|
|
|
6
|
+
## 0.13.0 (not yet published)
|
|
7
|
+
|
|
8
|
+
### Elastic memory
|
|
9
|
+
|
|
10
|
+
Additive. A workspace may be promised more memory than it holds while idle, and the host grows it when a command needs
|
|
11
|
+
it (opt-in per workspace).
|
|
12
|
+
|
|
13
|
+
- `Caps.allocation_mode` (`'fixed' | 'elastic'`, type `AllocationMode`) and `Caps.memory_mib_held` on `open()`,
|
|
14
|
+
`workspaces.fork()` and `workspace.fork()`. Given caps replace the stored ones (caps without `allocation_mode` make
|
|
15
|
+
the workspace fixed); omitted caps keep the stored layout.
|
|
16
|
+
- `workspace.memory` (`WorkspaceMemory`: `{allocation_mode, promised_mib, held_mib, plugged_mib}`, null on an older
|
|
17
|
+
API) and `workspace.allocationMode` (`'fixed'` on an older API). `WorkspaceView` carries `caps.allocation_mode`,
|
|
18
|
+
`caps.memory_mib_held` and `memory` (regenerated contract types).
|
|
19
|
+
- `RunResult.memoryGrow` (`MemoryGrow | null`): the grow the exec's start waited for (`outcome` delivered, partial,
|
|
20
|
+
missed or failed; `from_mib`, `want_mib`, `got_mib`, `deliver_ms`), from the start's session. `ExecSession`
|
|
21
|
+
carries `memory_grow`. The `exec` agent tool adds `memory_grow` to its result when a grow ran.
|
|
22
|
+
- `KnownErrorReason` adds `allocation_mode_not_available`, `requires_elastic` and `exceeds_memory_mib` (422
|
|
23
|
+
`validation_failed`; elastic with file-first is the existing `not_supported_for_mode`).
|
|
24
|
+
|
|
25
|
+
### Burst execution
|
|
26
|
+
|
|
27
|
+
Additive. One heavy, run-to-completion command can run on a larger, short-lived burst VM on the workspace's host, and
|
|
28
|
+
its file changes are applied back.
|
|
29
|
+
|
|
30
|
+
- `RunOptions.burst` (`'never' | 'always'`), `burstVcpus` and `burstMemoryMib` on `exec.run()`; `ExecStartRequest`
|
|
31
|
+
carries `burst`, `burst_vcpus` and `burst_memory_mib` (regenerated contract types).
|
|
32
|
+
- `RunResult.burst` (`BurstSummary | null`): host, size, method, whether the changes were applied, the change counts,
|
|
33
|
+
`leftover_killed`, `overhead_ms`, `excluded_paths`, `timings`, `replayed`. `ExecSession.burst` carries it; types
|
|
34
|
+
`BurstSummary` and `BurstError` are exported.
|
|
35
|
+
- `exec.run()` rejects with `ShardfluxApiError` 409 `burst_unavailable` or `burst_apply_failed` also when the burst
|
|
36
|
+
fails after its start (the output stream's final error event, or `burst.error` on the ended session), without
|
|
37
|
+
reconnecting; it does not try to cancel a burst when its `signal` aborts (a burst cannot be canceled).
|
|
38
|
+
- `ErrorCode` adds `burst_unavailable` and `burst_apply_failed`; `KnownErrorReason` adds `burst_mode_not_supported`,
|
|
39
|
+
`burst_not_supported`, `burst_size_exceeds_plan`, `not_available`, `shared_volumes`, `fence_not_drained`,
|
|
40
|
+
`apply_pending`, `park_failed`, `workspace_resumed`, `interrupted`, `burst_lost`, `disk_full`, `apply_failed`,
|
|
41
|
+
`reverted` and `revert_failed`.
|
|
42
|
+
- `workspaceTools(ws, { burst: true })` (opt-in, default off) gives the processful `exec` tool the inputs `burst`,
|
|
43
|
+
`burst_vcpus` and `burst_memory_mib`, and its result adds `burst`. The default tool schemas are unchanged.
|
|
44
|
+
|
|
45
|
+
### Background commands in the agent tools
|
|
46
|
+
|
|
47
|
+
Additive. An agent can start a long command, keep working with the other tools, read its progress and stop it.
|
|
48
|
+
|
|
49
|
+
- The processful `exec` tool takes `background: true`: it starts the command and returns `{session_id, state}` at
|
|
50
|
+
once (a failed start adds `error: {code, message, reason}`, as the foreground `exec` gives it). `timeout_ms` goes up
|
|
51
|
+
to 86400000 (a day); with `background` it is sent only when given. The file-first `exec` is unchanged.
|
|
52
|
+
- New tool `exec_read` (permission `exec`): the session's `state`, `exit_code`, `term_signal`, `timed_out`,
|
|
53
|
+
`canceled` and output. Without offsets it returns the last `maxOutputBytes` of each stream; `stdout_offset` /
|
|
54
|
+
`stderr_offset` read from there, and the result's `next_stdout_offset` / `next_stderr_offset` continue where it
|
|
55
|
+
stopped (`truncated` when more output follows). Text is cut at character boundaries only. `wait_ms` (up to 60000)
|
|
56
|
+
waits for the command to exit and returns as soon as it does. `burst` and `error` when the session has them.
|
|
57
|
+
- New tool `exec_cancel` (permission `exec`): SIGTERM, then SIGKILL after `grace_ms` (1-60000, default 5000);
|
|
58
|
+
returns `{session_id, state, exit_code, canceled}`.
|
|
59
|
+
- With `workspaceTools(ws, { burst: true })`, `background: true` together with `burst: 'always'` is refused with
|
|
60
|
+
`ToolArgumentError` before any request (a burst holds the workspace until it applies its changes); `burst: 'never'`
|
|
61
|
+
with `background` is an ordinary background start.
|
|
62
|
+
- `exec_read` and `exec_cancel` follow `exec` in `workspaceTools()` on processful workspaces; a file-first workspace
|
|
63
|
+
does not offer them. Both send the wake hint like the other tools.
|
|
64
|
+
|
|
65
|
+
### Held fork
|
|
66
|
+
|
|
67
|
+
Additive. A waited fork returns the running copy and its tool token in one request.
|
|
68
|
+
|
|
69
|
+
- `workspaces.fork()` and `workspace.fork()` with `wait`: when the API holds the fork (`Prefer: wait`, answered 200
|
|
70
|
+
with `Preference-Applied`), the copy's handle adopts the returned running view and final-epoch tool token, with no
|
|
71
|
+
operation poll, view refresh or token request. The request's time counts against `wait.timeoutMs`.
|
|
72
|
+
- `ForkOptions` and `WaitedForkOptions` (exported) add `agentLabel` and `tools`, which choose that token. They are sent
|
|
73
|
+
only on a server-held fork and need an API with held fork. `wait: { serverWait: false }` keeps polling.
|
|
74
|
+
- An API that does not apply `Prefer` answers 202 as before; the call then polls the operation and refreshes the copy.
|
|
75
|
+
- The fork's progress `request` phase has reason `held` when it asks the server to wait.
|
|
76
|
+
|
|
77
|
+
### Exec start and output in one request
|
|
78
|
+
|
|
79
|
+
- `exec.run()` starts the command and follows its output in one request: the start asks for NDJSON
|
|
80
|
+
(`Accept: application/x-ndjson`) and a cell that supports it answers with the output stream. A dropped stream
|
|
81
|
+
reconnects by session and byte offsets, never starting the command again. A cell that answers the start with the
|
|
82
|
+
session JSON gets the output request after it, as before; a start that failed still rejects with `ExecStartError`
|
|
83
|
+
(or the burst's error).
|
|
84
|
+
- `RunResult.memoryGrow` and `burst` are unchanged: the combined stream's exit event carries the start's
|
|
85
|
+
`memory_grow`.
|
|
86
|
+
|
|
87
|
+
### Immutable paths
|
|
88
|
+
|
|
89
|
+
A template may declare directories as immutable: every workspace of the template mounts them read-only from the
|
|
90
|
+
template's newest published version, picked up at its next cold boot or resume.
|
|
91
|
+
|
|
92
|
+
- Recipe v2 `immutable` (`TemplateRecipeV2.immutable`, regenerated): the list in `template.yaml`, sent as written by
|
|
93
|
+
`buildFromFile()`, `buildFromRecipe()` and `builds.create()`. The stored recipe v2 of a build
|
|
94
|
+
(`TemplateBuildRecipeV2`) and an exported recipe (`versions.recipe()`) carry it.
|
|
95
|
+
- `TemplateVersion.immutable` (`TemplateVersionImmutable`: `{paths, bytes}`, or null when the version declares none),
|
|
96
|
+
on every version of `templates.get()` and on `open_version`.
|
|
97
|
+
- `TemplateBuild.immutable_paths`: the list the produced version declares (the recipe's, else the open version's).
|
|
98
|
+
`TemplateBuild.failure.details` carries code-specific fields: `immutable_path_missing` names `details.path`.
|
|
99
|
+
- `workspace.immutableVersion` (`template.immutable_version` of the view): the template version whose immutable paths
|
|
100
|
+
the workspace has mounted, which can be newer than `template.version`; null when it mounts none.
|
|
101
|
+
- `KnownErrorReason` adds `immutable_path_removed` (422; `details.removed`, `details.open_version`),
|
|
102
|
+
`immutable_paths_unsupported_base` (422; `details.base`, `details.required_feature`, `details.paths`) and
|
|
103
|
+
`read_only_path` (409 `conflict` from the files API for a write under an immutable path).
|
|
104
|
+
- **Breaking:** the update policy left the API. Removed: the `UpdatePolicy` type, `TemplateSummary.update_policy`, the
|
|
105
|
+
regenerated `update_policy` of the workspace view, of `TemplateDefaults` and of the `defaults` inputs, and the reason
|
|
106
|
+
`update_policy_not_available`. It was always `pinned`; code that read it drops the read.
|
|
107
|
+
|
|
108
|
+
### Exec cancel grace
|
|
109
|
+
|
|
110
|
+
- `exec.cancel(id, graceMs?)` takes `graceMs` 0 to 60000 (0 or omitted: 5000), as the cell API now enforces (422
|
|
111
|
+
`validation_failed` outside it), and waits for the answer up to the grace plus 15 s when that is longer than the
|
|
112
|
+
client's `timeoutMs`: a command that ignores SIGTERM is answered after the grace and the SIGKILL, not a timeout.
|
|
113
|
+
- The cancel that `exec.run` sends when its `signal` aborts passes `killGraceMs` capped at 60000, so a larger
|
|
114
|
+
`killGraceMs` still cancels the command.
|
|
115
|
+
|
|
116
|
+
## 0.12.0 (release candidate)
|
|
117
|
+
|
|
118
|
+
Completed exec output streams are drained before releasing their HTTP connections, with a bounded cleanup if a peer does not close.
|
|
119
|
+
|
|
120
|
+
Pooled HTTP/1.1 on Node 26 with pinned undici 8.10.2; custom fetch stays unchanged. Labels on open/list and setLabels, typed idle(), keepalive() and setIdlePolicy(). Protocol errors carry source; isWorkspaceGone recognizes only an explicit API workspace_deleted refusal. Exec sessions accept stdin_open and offset-addressed exec.input. Failed workspaces recover through resume/auto-wake with the same ID. Completed deletions free keys for new IDs.
|
|
121
|
+
|
|
122
|
+
Requires the QM integration backend release for labels, failed recovery, key reuse and pipe stdin. No production deployment has occurred from this branch.
|
|
123
|
+
|
|
124
|
+
### Instant suspend: durable storage in the result
|
|
125
|
+
|
|
126
|
+
Additive. A suspend returns as soon as the workspace is sealed on its host; the copy lands in durable storage right
|
|
127
|
+
after. This release reads that from the result and can wait for it.
|
|
128
|
+
|
|
129
|
+
- `suspend({ durable: true })` (`SuspendOptions.durable`, also `cloud.workspaces.suspend(id, { durable: true })`):
|
|
130
|
+
resolves once `result.durable` is `true`. It implies `wait` and uses the same `timeoutMs`/`signal` (one budget for
|
|
131
|
+
the suspend and the copy). Rejects with `DurabilityLostError` (new; `durability` with the reason) when the copy
|
|
132
|
+
cannot be made, and `OperationTimeoutError` with `durable: true` when the time runs out (the copy continues).
|
|
133
|
+
`suspend({ wait: true })` still resolves when the suspend succeeds, without waiting for the copy.
|
|
134
|
+
- `cloud.workspaces.waitForDurable(operationOrId, waitOptions)`: the same wait for an operation you hold (a suspend, or
|
|
135
|
+
a fork of a running workspace).
|
|
136
|
+
- `isDurable(op)`, `durabilityOf(op)` (`Durability`: `state` `pending` | `durable` | `lost`, `checkpointId`,
|
|
137
|
+
`generationId`, `localCommitAt`, `durableBy`, `durableAt`, `localCommitToDurableMs`, `overdueAt`, `reason`) and
|
|
138
|
+
`lostSuspendOf(op)` (`LostSuspend`, on a resume result). `ServerTiming` gains `durable`, `suspendPath`, `durability`
|
|
139
|
+
and `lostSuspend`; progress phase `durable` (reason `overdue` past `durableBy`); `formatTiming()` prints them.
|
|
140
|
+
- `Operation.result` is typed with `durable`, `suspend_path`, `durability` and `lost_suspend` (OpenAPI).
|
|
141
|
+
- `OperationFailedError.workspaceActive`: a suspend-when-idle canceled with `workspace_active` (the workspace was in
|
|
142
|
+
use; nothing changed and it keeps running). `workspace.waitUntilReady()` resolves for it instead of throwing, as
|
|
143
|
+
`wake()` already did.
|
|
144
|
+
|
|
145
|
+
Requires the instant-suspend cell release for `durable: false` results; against other servers every succeeded suspend
|
|
146
|
+
is already durable and `durable: true` resolves at once.
|
|
147
|
+
|
|
6
148
|
## 0.11.1
|
|
7
149
|
|
|
8
150
|
Wording: messages and JSDoc say what to do, without internals (no behaviour change).
|
package/README.md
CHANGED
|
@@ -14,7 +14,7 @@ tool calls into the workspace ([tool-call capture](#tool-call-capture-070)).
|
|
|
14
14
|
> **(0.7.0+)** not in 0.6.x and **(0.6.0+)** not in 0.5.0; [CHANGELOG.md](./CHANGELOG.md) lists what each version added. Check yours with
|
|
15
15
|
> `npm ls @shardflux/sdk` or the exported `SDK_VERSION`.
|
|
16
16
|
|
|
17
|
-
- ESM only,
|
|
17
|
+
- ESM only, Node.js 24 or later. Reading a YAML template file uses the optional peer
|
|
18
18
|
dependency `yaml` (`npm install yaml`); JSON template files need nothing.
|
|
19
19
|
- Typed from the published OpenAPI documents.
|
|
20
20
|
- Retries, idempotency keys, operation polling and tool-token refresh are handled for you.
|
|
@@ -108,8 +108,9 @@ const listing = await cell.files.list('/home/user');
|
|
|
108
108
|
await cell.files.remove('/home/user/data.bin');
|
|
109
109
|
```
|
|
110
110
|
|
|
111
|
-
`exec.run()`
|
|
112
|
-
|
|
111
|
+
`exec.run()` starts the command and streams its output in one request **(0.13.0+)**, resumes from
|
|
112
|
+
byte offsets if the output stream drops, and never starts the command twice. Aborting its `signal`
|
|
113
|
+
also cancels the command in the workspace. File writes are atomic
|
|
113
114
|
and durable (acknowledged after fsync).
|
|
114
115
|
|
|
115
116
|
`cwd` is an absolute path (commands start in `/home/user` without one). The API refuses a relative
|
|
@@ -187,7 +188,7 @@ await workspace.resume({ wait: true }); // or simply open() the key again
|
|
|
187
188
|
|
|
188
189
|
const { workspace: copy } = await workspace.fork({ key: 'customer-42/experiment' }, { wait: true });
|
|
189
190
|
|
|
190
|
-
await copy.delete(); // tool access ends immediately;
|
|
191
|
+
await copy.delete(); // tool access ends immediately; the key can be reused after deletion finishes
|
|
191
192
|
```
|
|
192
193
|
|
|
193
194
|
### Requested or finished
|
|
@@ -199,12 +200,40 @@ means depends on `wait`:
|
|
|
199
200
|
| --- | --- | --- |
|
|
200
201
|
| `await workspace.suspend()` | the suspend is **requested** (usually `queued`; the workspace is still running) | the operation (`Operation`) |
|
|
201
202
|
| `await workspace.suspend({ wait: true })` **(0.6.0+)** | the suspend has **finished** (`workspace.state` is then `suspended`) | the succeeded operation (`FinishedOperation`) |
|
|
203
|
+
| `await workspace.suspend({ durable: true })` **(0.12.0+)** | the suspend has finished **and** its copy is in durable storage (`result.durable` is `true`) | the succeeded operation (`FinishedOperation`) |
|
|
202
204
|
|
|
203
205
|
`wait` also takes `WaitOptions` (`timeoutMs`, default 5 minutes; `signal`; `onProgress`). A failed operation throws
|
|
204
206
|
`OperationFailedError`. Running out of time throws `OperationTimeoutError`, and the operation continues server side:
|
|
205
207
|
wait again with `cloud.workspaces.waitForOperation(err.operationId)`. Without `wait`, the returned operation is the
|
|
206
208
|
handle for the work in progress: pass its `id` to `waitForOperation()` when you need it finished.
|
|
207
209
|
|
|
210
|
+
**A waited fork is one request (0.13.0+).** `fork(target, { wait: true })` asks the API to hold the fork until the
|
|
211
|
+
copy runs, and the answer carries the running copy and a tool token for it: the copy's handle is ready with no
|
|
212
|
+
operation poll, view refresh or token request. `agentLabel` and `tools` pick the token it brings back. Its timing is one
|
|
213
|
+
`request` phase with reason `held`; `wait: { serverWait: false }` polls instead. An API without the held fork answers
|
|
214
|
+
at once, and the SDK then waits for the operation and refreshes the copy, as before.
|
|
215
|
+
|
|
216
|
+
**Durable storage (0.12.0+).** A suspend returns as soon as the workspace is sealed on its host, typically in a few
|
|
217
|
+
hundred ms, and its RAM and CPU are released at that moment. `result.durable` turns `true` when the copy lands in
|
|
218
|
+
durable storage, typically within a second; until then it is `false` and `result.durability` shows the copy's
|
|
219
|
+
progress. Most code needs nothing more: a suspended workspace resumes, is read and is forked the same way either way.
|
|
220
|
+
When your code must know the copy is durable (before deleting a local artifact, or at the end of a job), ask for it:
|
|
221
|
+
|
|
222
|
+
```ts
|
|
223
|
+
import { durabilityOf } from '@shardflux/sdk';
|
|
224
|
+
|
|
225
|
+
const op = await workspace.suspend({ durable: true }); // resolves once result.durable is true
|
|
226
|
+
durabilityOf(op); // { state: 'durable', checkpointId, localCommitAt, durableAt, localCommitToDurableMs, ... }
|
|
227
|
+
```
|
|
228
|
+
|
|
229
|
+
- `durable: true` implies `wait` and shares its `timeoutMs` and `signal` (e.g. `{ durable: true, wait: { timeoutMs: 60_000 } }`).
|
|
230
|
+
If the time runs out first, `OperationTimeoutError` has `durable: true` and the copy continues server side.
|
|
231
|
+
- `cloud.workspaces.waitForDurable(operationOrId)` does the same for an operation you already hold, such as a
|
|
232
|
+
finished fork of a running workspace, whose result carries `durable` and `durability` the same way.
|
|
233
|
+
- `isDurable(op)`, `durabilityOf(op)` and `lastTiming.server.durable` / `.durability` read the fields; `formatTiming()`
|
|
234
|
+
prints `sealed on host, durable 435 ms later`. Results from before 0.12.0 servers count as durable.
|
|
235
|
+
- The errors reference on docs.shardflux.dev lists what `durable: true` can reject with.
|
|
236
|
+
|
|
208
237
|
**Start deadlines.** A queued open, resume, restore or fork (`capacity_pending`) has a deadline 15 minutes after it
|
|
209
238
|
was created. Its `error.details.deadline_at` carries the deadline; the `phase` progress event carries it as
|
|
210
239
|
`deadlineAt` **(0.6.2+)**, and so does `OperationTimeoutError` when your wait ends first. A start still pending at the
|
|
@@ -295,8 +324,7 @@ operation keeps running server side (a queued start until its `deadlineAt`); wai
|
|
|
295
324
|
same call. `templates.builds.waitForBuild()` waits the same way. A waited `open()` issues the first tool token
|
|
296
325
|
together with the final workspace read, so the first tool call starts at once.
|
|
297
326
|
|
|
298
|
-
|
|
299
|
-
`SHARDFLUX_HTTP_KEEPALIVE=1`, or your own `fetch`, reuses connections.
|
|
327
|
+
From 0.12.0, Node 26 reuses TLS connections through a private HTTP/1.1 pool. Other runtimes keep native fetch. A supplied `fetch` stays unchanged. `SHARDFLUX_HTTP_KEEPALIVE=0` forces connection closure for diagnosis; `=1` opts into native pooling. No global dispatcher is installed.
|
|
300
328
|
|
|
301
329
|
List and look up workspaces:
|
|
302
330
|
|
|
@@ -313,6 +341,46 @@ sessions, template drafts and test instances. `findByKey()` searches every lifet
|
|
|
313
341
|
workspace, and a tombstone only when no live workspace has the key. Use it rather than `listAll({ keyPrefix })` for
|
|
314
342
|
lookups by key.
|
|
315
343
|
|
|
344
|
+
### Elastic memory (0.13.0+)
|
|
345
|
+
|
|
346
|
+
Promise a workspace a lot of memory while it holds only what it uses. An elastic workspace idles at its held floor
|
|
347
|
+
(`memory_mib_held`, default 1024) and grows before each command starts, so compilers, test runners and V8 size
|
|
348
|
+
themselves from the memory they will actually get. After 30 s without work it shrinks back. The layout takes effect at
|
|
349
|
+
the workspace's next start.
|
|
350
|
+
|
|
351
|
+
```ts
|
|
352
|
+
const ws = await cloud.workspaces.open({
|
|
353
|
+
key: 'customer-42/repo-7',
|
|
354
|
+
template: 'node',
|
|
355
|
+
caps: { memory_mib: 8192, allocation_mode: 'elastic' },
|
|
356
|
+
});
|
|
357
|
+
ws.memory; // { allocation_mode: 'elastic', promised_mib: 8192, held_mib: 1024, plugged_mib: 3072 }
|
|
358
|
+
|
|
359
|
+
const r = await ws.cell().exec.run(['npm', 'test']);
|
|
360
|
+
r.memoryGrow; // { outcome: 'delivered', from_mib: 1024, want_mib: 4096, got_mib: 4096, deliver_ms: 175 }
|
|
361
|
+
```
|
|
362
|
+
|
|
363
|
+
Given `caps` replace the stored ones, so send `allocation_mode` with every `caps` you pass; an open without `caps`
|
|
364
|
+
keeps the stored layout. The exec agent tool adds `memory_grow` to its result when the command's start grew the VM.
|
|
365
|
+
|
|
366
|
+
### Burst execution (0.13.0+)
|
|
367
|
+
|
|
368
|
+
Run one heavy command, such as a cold build or a full test suite, on a larger VM without resizing the workspace. With
|
|
369
|
+
`burst: 'always'` the command runs on a short-lived burst VM over a copy of the workspace, output streams as usual,
|
|
370
|
+
and its file changes are applied back byte for byte when it exits.
|
|
371
|
+
|
|
372
|
+
```ts
|
|
373
|
+
const r = await ws.cell().exec.run(['go', 'build', './...'], { burst: 'always', burstVcpus: 16 });
|
|
374
|
+
r.exitCode; // 0
|
|
375
|
+
r.burst; // { host: 'local', vcpus: 16, memory_mib: 8192, method: 'layer', applied: true,
|
|
376
|
+
// written_files: 3, written_dirs: 1, removed: 1, written_bytes: 40960, overhead_ms: 641, ... }
|
|
377
|
+
```
|
|
378
|
+
|
|
379
|
+
A burst carries back the command's file changes; processes it started and tmpfs content stay in the burst VM, and the
|
|
380
|
+
workspace's own processes resume where they were. A burst is not signalled or canceled. If applying the changes
|
|
381
|
+
fails (`burst_apply_failed`), retry `exec.run` with the same `sessionId` to finish the apply. The agent tools offer
|
|
382
|
+
bursts when asked: `workspaceTools(ws, { burst: true })`.
|
|
383
|
+
|
|
316
384
|
### Timing and progress (0.6.0+)
|
|
317
385
|
|
|
318
386
|
Every open, wake and waited lifecycle call is traced. `workspace.lastTiming` (and `err.timing` when the call fails)
|
|
@@ -548,6 +616,27 @@ workspace keeps running so you can inspect it, and `workspace.startup` names the
|
|
|
548
616
|
The next open runs the failed step again. Versions report their `settings`, platform templates their `category`
|
|
549
617
|
(`os` or `stack`), builds the `denied_hosts` their build network refused (add them to `build.network.extra_hosts`).
|
|
550
618
|
|
|
619
|
+
### Immutable paths (0.13.0+)
|
|
620
|
+
|
|
621
|
+
Directories a template declares `immutable` follow the template: every workspace of it mounts them read-only from the
|
|
622
|
+
template's newest published version, so the tools, models or data you ship there reach existing workspaces with your
|
|
623
|
+
next version, at their next cold boot or resume. Everything else in a workspace stays its own.
|
|
624
|
+
|
|
625
|
+
```yaml
|
|
626
|
+
# template.yaml
|
|
627
|
+
immutable: [/opt/acme] # read-only in every workspace; follows the newest version
|
|
628
|
+
```
|
|
629
|
+
|
|
630
|
+
```ts
|
|
631
|
+
const t = await cloud.templates.get('acme-dev');
|
|
632
|
+
console.log(t.open_version?.immutable); // { paths: ['/opt/acme'], bytes: 100663296 }
|
|
633
|
+
const ws = await cloud.workspaces.open({ key: 'customer-42/main', template: 'acme-dev' });
|
|
634
|
+
console.log(ws.template.version, ws.immutableVersion); // created from v1; /opt/acme shows v2
|
|
635
|
+
```
|
|
636
|
+
|
|
637
|
+
A build without `immutable` keeps the list of the template's open version, and each new version keeps every path of
|
|
638
|
+
it: add paths, never remove them. A build reports the list its version declares in `immutable_paths`.
|
|
639
|
+
|
|
551
640
|
## Agent tools
|
|
552
641
|
|
|
553
642
|
`workspaceTools(workspace)` returns tools with a name, a description, a JSON Schema for the
|
|
@@ -582,6 +671,20 @@ instead of being overwritten. Each call first sends `workspace.hint()` without w
|
|
|
582
671
|
that off, e.g. when you send the hint yourself as the model starts a tool call), except `read_file`, `list_files` and
|
|
583
672
|
`search_files`: a sleeping workspace answers them from its disk without waking.
|
|
584
673
|
|
|
674
|
+
**Background commands (0.13.0+).** On a processful workspace `exec` takes `background: true` for anything that runs
|
|
675
|
+
longer than a few minutes (a build, a test suite, a training run, a server): it returns `{ session_id, state }` at
|
|
676
|
+
once and the model keeps using the other tools. `exec_read` reads the command's state, exit code and output (the end
|
|
677
|
+
of each stream, or from the `next_stdout_offset` / `next_stderr_offset` of the previous read; `wait_ms` waits for the
|
|
678
|
+
exit and returns as soon as it happens), and `exec_cancel` stops it. `timeout_ms` (up to a day) sets the longest the
|
|
679
|
+
command may run; the workspace stays awake while it does. A burst (`burst: 'always'`) runs in the foreground.
|
|
680
|
+
|
|
681
|
+
```ts
|
|
682
|
+
const tools = workspaceTools(workspace);
|
|
683
|
+
const { session_id } = (await executeToolCall(tools, { name: 'exec', input: { command: 'make -j8 test', background: true, timeout_ms: 7_200_000 } })) as { session_id: string };
|
|
684
|
+
// ... other tool calls ...
|
|
685
|
+
const progress = await executeToolCall(tools, { name: 'exec_read', input: { session_id, wait_ms: 30_000 } });
|
|
686
|
+
```
|
|
687
|
+
|
|
585
688
|
For a file-first workspace (0.9.0+) the tools are `exec` and the files tools only: `exec` runs each command as an
|
|
586
689
|
execution and adds `execution_id`, `state`, `tree_revision` and `changed` to its result, and the process, terminal,
|
|
587
690
|
git and browser tools are not offered. `workspaceTools(ws, { mode })` builds the definitions without touching the
|
|
@@ -838,6 +941,14 @@ A read of a sleeping workspace that its disk cannot answer (409 `workspace_not_r
|
|
|
838
941
|
`host_feature_unavailable` (search or patches are not available for the workspace, `details.feature`) is neither
|
|
839
942
|
retried nor woken: run `grep` with `exec`, or read then write the file, instead.
|
|
840
943
|
|
|
944
|
+
Immutable paths **(0.13.0+)**: a files write under an immutable path is 409 `conflict` with `reason` `read_only_path`
|
|
945
|
+
(not retried; write elsewhere, or change the template). A build is refused with 422 `validation_failed` and `reason`
|
|
946
|
+
`immutable_path_removed` when its recipe drops a path of the template's open version (`details.removed`; keep them),
|
|
947
|
+
`immutable_paths_unsupported_base` when its base cannot carry immutable paths (`details.base`; build on a newer version
|
|
948
|
+
of that base), or `invalid_path` (`details.field` `recipe.immutable[<i>]`). A build that fails on them has
|
|
949
|
+
`failure.code` `immutable_path_missing` (`failure.details.path` is not a directory in the built filesystem) or
|
|
950
|
+
`immutable_image_too_large`.
|
|
951
|
+
|
|
841
952
|
## Usage and overage
|
|
842
953
|
|
|
843
954
|
`cloud.usage` reads the organization's usage (API keys see organization totals and their own project's workspaces):
|
|
@@ -941,3 +1052,44 @@ or no longer supported, it emits one warning:
|
|
|
941
1052
|
## License
|
|
942
1053
|
|
|
943
1054
|
Apache-2.0
|
|
1055
|
+
|
|
1056
|
+
|
|
1057
|
+
## Workspace integrations (0.12.0+)
|
|
1058
|
+
|
|
1059
|
+
```ts
|
|
1060
|
+
const ws = await cloud.workspaces.open({ key: 'qm/project-42', template: 'python-node-browser',
|
|
1061
|
+
labels: { scope: 'project-42', owner: 'qm' }, idlePolicy: 'never' });
|
|
1062
|
+
const matches = await cloud.workspaces.list({ labels: { scope: 'project-42' } });
|
|
1063
|
+
await ws.setLabels({ scope: 'project-42', owner: 'qm' }); // replaces all labels; {} clears
|
|
1064
|
+
await ws.setIdlePolicy('adaptive'); // null restores the inherited policy
|
|
1065
|
+
const idle = await ws.idle(); // read-only, does not wake or record activity
|
|
1066
|
+
await ws.keepalive(60); // seconds; never shortens an existing keepalive
|
|
1067
|
+
await ws.suspendWhenIdle({ afterSeconds: 2 }); // one safe server-side request; no /idle + suspend race
|
|
1068
|
+
```
|
|
1069
|
+
|
|
1070
|
+
Labels are exact string pairs (at most 50, keys 1–64 and values 0–256 characters); they are metadata, not secrets.
|
|
1071
|
+
A failed workspace is recoverable: `await ws.wake()` and ordinary auto-waking calls restart that ID using its existing disk. They do not create a replacement workspace. A failed attempt still raises its operation error; there is no unbounded restart loop.
|
|
1072
|
+
After `await ws.delete({ wait: true })`, opening its key creates a **new ID**. The old ID, tokens and history remain deleted. A key stays reserved while deletion is in progress.
|
|
1073
|
+
|
|
1074
|
+
For a background command with a pipe (no PTY):
|
|
1075
|
+
|
|
1076
|
+
```ts
|
|
1077
|
+
const cell = ws.cell();
|
|
1078
|
+
const session = await cell.exec.start({ session_id: 'qm-worker-1', argv: ['cat'], stdin_open: true });
|
|
1079
|
+
let ack = await cell.exec.input(session.session_id, 'hello\n', { offset: 0 });
|
|
1080
|
+
ack = await cell.exec.input(session.session_id, '', { offset: ack.offset, close: true });
|
|
1081
|
+
```
|
|
1082
|
+
|
|
1083
|
+
Input frames are at most 64 KiB. The acknowledged offset counts bytes accepted into the pipe, not bytes consumed by the program. A blocked writer may receive a partial acknowledgement; continue from the returned offset. Retrying the identical last frame is safe. `close` sends EOF only once the entire frame is accepted. Do not combine `stdin_open` with the existing one-shot `stdin` option. The transport writes no separate input log or payload file; program output and full-state snapshots retain their normal persistence. Full-state suspend preserves the pipe. Where a workspace does not offer it, the start is refused with 409 `conflict`, `reason` `host_feature_unavailable` and `details.feature: 'exec_stdin'`, not retryable. Running sessions count as work within the documented idle-command window; long-running services should declare keepalives or choose `never`.
|
|
1084
|
+
|
|
1085
|
+
```ts
|
|
1086
|
+
import { isWorkspaceGone, ShardfluxProtocolError } from '@shardflux/sdk';
|
|
1087
|
+
// isWorkspaceGone(error) is true only for an explicit API workspace_deleted refusal.
|
|
1088
|
+
// A protocol error, a cell 404, a scoped API 404 or a network failure is never deletion evidence.
|
|
1089
|
+
```
|
|
1090
|
+
|
|
1091
|
+
`ShardfluxProtocolError.source` is `api` or `cell` for SDK responses (`unknown` only for an error constructed without a source by older caller code). It remains a protocol failure even when `status === 404`. Never delete local data based on an HTTP status alone.
|
|
1092
|
+
|
|
1093
|
+
### Repositories with a minimum release age
|
|
1094
|
+
|
|
1095
|
+
Keep your repository's age policy. Pin an exact version that has aged enough; do not exempt the entire `@shardflux/*` scope. The first SDK, 0.5.0, was published September 26, 2026 at 21:13:52 UTC and reaches seven days on October 3 at that time. New versions, including 0.12.0, need their own seven days after publication. Before then no registry setting on our side can make them eligible. `npm view @shardflux/sdk time --json` shows publication times. A targeted exception for an inspected exact release is a repository-owner decision, not an installation requirement we bypass. Older versions do not contain this release's fixes.
|
package/dist/account.d.ts
CHANGED
|
@@ -102,7 +102,7 @@ export interface ShardfluxAccountOptions {
|
|
|
102
102
|
sessionToken?: string;
|
|
103
103
|
/** Default https://api.shardflux.dev. */
|
|
104
104
|
baseUrl?: string;
|
|
105
|
-
/** Default:
|
|
105
|
+
/** Default: pooled HTTP/1.1 on Node 26+, native fetch on other runtimes (see defaultFetch in http.ts). */
|
|
106
106
|
fetch?: typeof fetch;
|
|
107
107
|
userAgent?: string;
|
|
108
108
|
/** Per-request timeout (ms), default 30 s. */
|
package/dist/account.js
CHANGED
|
@@ -571,7 +571,7 @@ export class ShardfluxAccount {
|
|
|
571
571
|
const b = body;
|
|
572
572
|
if (b && typeof b === 'object' && typeof b.session_token === 'string') {
|
|
573
573
|
if (!isSessionToken(b.session_token))
|
|
574
|
-
throw new ShardfluxProtocolError('the API returned a session_token that is not a CLI session token (sfu_...)', 200);
|
|
574
|
+
throw new ShardfluxProtocolError('the API returned a session_token that is not a CLI session token (sfu_...)', 200, 'api');
|
|
575
575
|
this.#token = b.session_token;
|
|
576
576
|
const expiresAt = typeof b.session_expires_at === 'string' ? b.session_expires_at : '';
|
|
577
577
|
await this.#onSessionToken?.({ token: b.session_token, expiresAt });
|
package/dist/cell.d.ts
CHANGED
|
@@ -22,6 +22,7 @@
|
|
|
22
22
|
* have before sending them (NotSupportedForModeError).
|
|
23
23
|
*/
|
|
24
24
|
import type { components, paths } from './generated/cell-api.js';
|
|
25
|
+
import { ShardfluxApiError } from './errors.js';
|
|
25
26
|
import type { WorkspaceMode } from './errors.js';
|
|
26
27
|
import { ExecutionResult } from './executions.js';
|
|
27
28
|
import type { ExecutionGetOptions, ExecutionRunOptions } from './executions.js';
|
|
@@ -31,6 +32,25 @@ import type { ToolTokenManager } from './tokens.js';
|
|
|
31
32
|
type S = components['schemas'];
|
|
32
33
|
export type ExecStartRequest = S['ExecStartRequest'];
|
|
33
34
|
export type ExecSession = S['ExecSession'];
|
|
35
|
+
/**
|
|
36
|
+
* Elastic memory (0.13.0): the host grew the workspace's memory before an exec started, and the exec
|
|
37
|
+
* waited for it. `outcome` delivered | partial | missed | failed; `from_mib`, `want_mib`, `got_mib` are guest memory
|
|
38
|
+
* before, aimed for and after; `deliver_ms` the time to deliver. On the session the start returned; absent when no
|
|
39
|
+
* grow ran (fixed workspaces, or already at the exec size).
|
|
40
|
+
*/
|
|
41
|
+
export type MemoryGrow = S['MemoryGrow'];
|
|
42
|
+
/**
|
|
43
|
+
* Burst execution (0.13.0; needs an organization entitlement): the session of
|
|
44
|
+
* an exec started with `burst: 'always'` ran on a larger, short-lived burst VM on the workspace's host, and its file
|
|
45
|
+
* changes were applied back (`applied`). `host`, `vcpus`, `memory_mib`, `method`, the change counts
|
|
46
|
+
* (`written_files`, `written_dirs`, `removed`, `written_bytes`), `leftover_killed` (processes the command left behind:
|
|
47
|
+
* they do not survive a burst), `overhead_ms`, `excluded_paths`, `timings`, `replayed`; `error` when the burst failed
|
|
48
|
+
* after the start answered.
|
|
49
|
+
*/
|
|
50
|
+
export type BurstSummary = S['BurstSummary'];
|
|
51
|
+
/** The failure recorded on a burst session (`burst.error`): code `burst_unavailable` or `burst_apply_failed`. */
|
|
52
|
+
export type BurstError = S['BurstError'];
|
|
53
|
+
export type ExecInputResult = S['ExecInputResult'];
|
|
34
54
|
export type OutputEvent = S['OutputEvent'];
|
|
35
55
|
export type PtyOpenRequest = S['PtyOpenRequest'];
|
|
36
56
|
export type PtySession = S['PtySession'];
|
|
@@ -50,6 +70,8 @@ export type FilePatchResult = S['FilePatchResult'];
|
|
|
50
70
|
/** A content SHA-256 (lowercase hex), or `absent` for a path that must not exist. */
|
|
51
71
|
export type FileRevision = S['FileRevision'];
|
|
52
72
|
export type WakeHintResult = S['WakeHintResult'];
|
|
73
|
+
export type IdleStatus = S['IdleStatus'];
|
|
74
|
+
export type KeepaliveResult = S['KeepaliveResult'];
|
|
53
75
|
/** Where a running workspace's VM is: resident, frozen, hibernated or restoring. */
|
|
54
76
|
export type Residency = WakeHintResult['residency'];
|
|
55
77
|
/** `X-Served-From`: `disk` when a read was served from a suspended or hibernated workspace's disk. */
|
|
@@ -220,6 +242,10 @@ export interface RunResult {
|
|
|
220
242
|
session: ExecSession;
|
|
221
243
|
/** Output-stream reconnections performed (gateway restarts, network drops). */
|
|
222
244
|
reconnects: number;
|
|
245
|
+
/** The memory grow the start of this exec waited for (0.13.0; elastic workspaces), null when none ran. */
|
|
246
|
+
memoryGrow: MemoryGrow | null;
|
|
247
|
+
/** The burst summary (0.13.0; `burst: 'always'`), null for an ordinary exec. */
|
|
248
|
+
burst: BurstSummary | null;
|
|
223
249
|
}
|
|
224
250
|
export interface RunOptions {
|
|
225
251
|
sessionId?: string;
|
|
@@ -251,7 +277,24 @@ export interface RunOptions {
|
|
|
251
277
|
* (details.reason `secret_not_available`, details.names) and nothing runs; a name also present in `env` is 422.
|
|
252
278
|
*/
|
|
253
279
|
secretRefs?: string[];
|
|
280
|
+
/**
|
|
281
|
+
* Burst execution (0.13.0; needs the organization entitlement `policy.burst_exec`): `'always'` runs this
|
|
282
|
+
* run-to-completion command on a larger, short-lived burst VM on the workspace's
|
|
283
|
+
* host over a copy of the workspace disk, and applies its file changes back when it exits (`result.burst`).
|
|
284
|
+
* Processes it starts and tmpfs content (/tmp when tmpfs, /dev/shm) do not survive; the workspace's own processes
|
|
285
|
+
* pause meanwhile and writes to it are refused (409). Not with `stdin` (422 burst_not_supported). Refusals: 409
|
|
286
|
+
* `burst_unavailable` (details.reason not_available, layout_unsupported, host_capacity, ...), 409
|
|
287
|
+
* `burst_apply_failed` (details.applied_entries/pending_entries; retry with the same `sessionId` to finish the apply).
|
|
288
|
+
* A burst cannot be canceled: `signal` stops following it, not the command.
|
|
289
|
+
*/
|
|
290
|
+
burst?: 'never' | 'always';
|
|
291
|
+
/** Burst VM vCPUs (1-32; default the host's, 8, or the plan's ceiling when lower). */
|
|
292
|
+
burstVcpus?: number;
|
|
293
|
+
/** Burst VM memory in MiB (512-65536; default the host's, 8192, or the plan's ceiling when lower). */
|
|
294
|
+
burstMemoryMib?: number;
|
|
254
295
|
}
|
|
296
|
+
/** The API error of a burst's recorded failure (`burst.error` of its session). Internal (the agent tools use it too). */
|
|
297
|
+
export declare function burstFailure(e: BurstError): ShardfluxApiError;
|
|
255
298
|
/** Parses an NDJSON byte stream into objects (tolerates CRLF and a final unterminated line). */
|
|
256
299
|
export declare function ndjson<T>(body: ReadableStream<Uint8Array>): AsyncGenerator<T>;
|
|
257
300
|
export declare class CellClient {
|
|
@@ -296,7 +339,19 @@ export declare class CellClient {
|
|
|
296
339
|
follow?: boolean;
|
|
297
340
|
signal?: AbortSignal;
|
|
298
341
|
}) => Promise<AsyncGenerator<OutputEvent>>;
|
|
342
|
+
/** Write at the acknowledged offset (initially 0); a repeated identical last frame cannot duplicate input.
|
|
343
|
+
* At most 64 KiB per call. A partial acknowledgement requires continuing from the returned offset.
|
|
344
|
+
* Start with stdin_open: true. close sends EOF after this frame is fully accepted. */
|
|
345
|
+
input: (sessionId: string, data: string | Uint8Array, opts: {
|
|
346
|
+
offset: number;
|
|
347
|
+
close?: boolean;
|
|
348
|
+
signal?: AbortSignal;
|
|
349
|
+
}) => Promise<ExecInputResult>;
|
|
299
350
|
signal: (sessionId: string, signal: Signal, onlyLeader?: boolean) => Promise<ExecSession>;
|
|
351
|
+
/**
|
|
352
|
+
* SIGTERM to the process group, SIGKILL after `graceMs` (0 to 60000; 0 or omitted is 5000). Resolves once the
|
|
353
|
+
* command has ended; the request waits for the grace.
|
|
354
|
+
*/
|
|
300
355
|
cancel: (sessionId: string, graceMs?: number) => Promise<ExecSession>;
|
|
301
356
|
/**
|
|
302
357
|
* Starts (or re-attaches to) a session and collects its output until it exits, reconnecting
|
|
@@ -442,6 +497,10 @@ export declare class CellClient {
|
|
|
442
497
|
* workspace answers 409 `workspace_not_running` (`Workspace.hint()` then wakes it in the background).
|
|
443
498
|
*/
|
|
444
499
|
wakeHint(signal?: AbortSignal): Promise<WakeHintResult>;
|
|
500
|
+
/** Read idle signals without recording activity or waking the workspace. */
|
|
501
|
+
idle(signal?: AbortSignal): Promise<IdleStatus>;
|
|
502
|
+
/** Declare work for seconds (1..the server maximum). Never shortens a previous keepalive. */
|
|
503
|
+
keepalive(seconds: number, signal?: AbortSignal): Promise<KeepaliveResult>;
|
|
445
504
|
/**
|
|
446
505
|
* One page of the workspace's changes against its template (needs the `files` tool). File-first workspaces: each
|
|
447
506
|
* execution's result lists what it changed (`changed`); this route is refused (NotSupportedForModeError).
|