@shardflux/sdk 0.11.1 → 0.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -3,6 +3,38 @@
3
3
  Every API the README shows is available from the version named here. Breaking changes ship in minor releases and are
4
4
  marked **Breaking**.
5
5
 
6
+ ## 0.12.0 (release candidate)
7
+
8
+ Completed exec output streams are drained before releasing their HTTP connections, with a bounded cleanup if a peer does not close.
9
+
10
+ Pooled HTTP/1.1 on Node 26 with pinned undici 8.10.2; custom fetch stays unchanged. Labels on open/list and setLabels, typed idle(), keepalive() and setIdlePolicy(). Protocol errors carry source; isWorkspaceGone recognizes only an explicit API workspace_deleted refusal. Exec sessions accept stdin_open and offset-addressed exec.input. Failed workspaces recover through resume/auto-wake with the same ID. Completed deletions free keys for new IDs.
11
+
12
+ Requires the QM integration backend release for labels, failed recovery, key reuse and pipe stdin. No production deployment has occurred from this branch.
13
+
14
+ ### Instant suspend: durable storage in the result
15
+
16
+ Additive. A suspend returns as soon as the workspace is sealed on its host; the copy lands in durable storage right
17
+ after. This release reads that from the result and can wait for it.
18
+
19
+ - `suspend({ durable: true })` (`SuspendOptions.durable`, also `cloud.workspaces.suspend(id, { durable: true })`):
20
+ resolves once `result.durable` is `true`. It implies `wait` and uses the same `timeoutMs`/`signal` (one budget for
21
+ the suspend and the copy). Rejects with `DurabilityLostError` (new; `durability` with the reason) when the copy
22
+ cannot be made, and `OperationTimeoutError` with `durable: true` when the time runs out (the copy continues).
23
+ `suspend({ wait: true })` still resolves when the suspend succeeds, without waiting for the copy.
24
+ - `cloud.workspaces.waitForDurable(operationOrId, waitOptions)`: the same wait for an operation you hold (a suspend, or
25
+ a fork of a running workspace).
26
+ - `isDurable(op)`, `durabilityOf(op)` (`Durability`: `state` `pending` | `durable` | `lost`, `checkpointId`,
27
+ `generationId`, `localCommitAt`, `durableBy`, `durableAt`, `localCommitToDurableMs`, `overdueAt`, `reason`) and
28
+ `lostSuspendOf(op)` (`LostSuspend`, on a resume result). `ServerTiming` gains `durable`, `suspendPath`, `durability`
29
+ and `lostSuspend`; progress phase `durable` (reason `overdue` past `durableBy`); `formatTiming()` prints them.
30
+ - `Operation.result` is typed with `durable`, `suspend_path`, `durability` and `lost_suspend` (OpenAPI).
31
+ - `OperationFailedError.workspaceActive`: a suspend-when-idle canceled with `workspace_active` (the workspace was in
32
+ use; nothing changed and it keeps running). `workspace.waitUntilReady()` resolves for it instead of throwing, as
33
+ `wake()` already did.
34
+
35
+ Requires the instant-suspend cell release for `durable: false` results; against other servers every succeeded suspend
36
+ is already durable and `durable: true` resolves at once.
37
+
6
38
  ## 0.11.1
7
39
 
8
40
  Wording: messages and JSDoc say what to do, without internals (no behaviour change).
package/README.md CHANGED
@@ -14,7 +14,7 @@ tool calls into the workspace ([tool-call capture](#tool-call-capture-070)).
14
14
  > **(0.7.0+)** not in 0.6.x and **(0.6.0+)** not in 0.5.0; [CHANGELOG.md](./CHANGELOG.md) lists what each version added. Check yours with
15
15
  > `npm ls @shardflux/sdk` or the exported `SDK_VERSION`.
16
16
 
17
- - ESM only, no runtime dependencies, Node.js 24 or later. Reading a YAML template file uses the optional peer
17
+ - ESM only, Node.js 24 or later. Reading a YAML template file uses the optional peer
18
18
  dependency `yaml` (`npm install yaml`); JSON template files need nothing.
19
19
  - Typed from the published OpenAPI documents.
20
20
  - Retries, idempotency keys, operation polling and tool-token refresh are handled for you.
@@ -187,7 +187,7 @@ await workspace.resume({ wait: true }); // or simply open() the key again
187
187
 
188
188
  const { workspace: copy } = await workspace.fork({ key: 'customer-42/experiment' }, { wait: true });
189
189
 
190
- await copy.delete(); // tool access ends immediately; keys are never reused
190
+ await copy.delete(); // tool access ends immediately; the key can be reused after deletion finishes
191
191
  ```
192
192
 
193
193
  ### Requested or finished
@@ -199,12 +199,34 @@ means depends on `wait`:
199
199
  | --- | --- | --- |
200
200
  | `await workspace.suspend()` | the suspend is **requested** (usually `queued`; the workspace is still running) | the operation (`Operation`) |
201
201
  | `await workspace.suspend({ wait: true })` **(0.6.0+)** | the suspend has **finished** (`workspace.state` is then `suspended`) | the succeeded operation (`FinishedOperation`) |
202
+ | `await workspace.suspend({ durable: true })` **(0.12.0+)** | the suspend has finished **and** its copy is in durable storage (`result.durable` is `true`) | the succeeded operation (`FinishedOperation`) |
202
203
 
203
204
  `wait` also takes `WaitOptions` (`timeoutMs`, default 5 minutes; `signal`; `onProgress`). A failed operation throws
204
205
  `OperationFailedError`. Running out of time throws `OperationTimeoutError`, and the operation continues server side:
205
206
  wait again with `cloud.workspaces.waitForOperation(err.operationId)`. Without `wait`, the returned operation is the
206
207
  handle for the work in progress: pass its `id` to `waitForOperation()` when you need it finished.
207
208
 
209
+ **Durable storage (0.12.0+).** A suspend returns as soon as the workspace is sealed on its host, typically in a few
210
+ hundred ms, and its RAM and CPU are released at that moment. `result.durable` turns `true` when the copy lands in
211
+ durable storage, typically within a second; until then it is `false` and `result.durability` shows the copy's
212
+ progress. Most code needs nothing more: a suspended workspace resumes, is read and is forked the same way either way.
213
+ When your code must know the copy is durable (before deleting a local artifact, or at the end of a job), ask for it:
214
+
215
+ ```ts
216
+ import { durabilityOf } from '@shardflux/sdk';
217
+
218
+ const op = await workspace.suspend({ durable: true }); // resolves once result.durable is true
219
+ durabilityOf(op); // { state: 'durable', checkpointId, localCommitAt, durableAt, localCommitToDurableMs, ... }
220
+ ```
221
+
222
+ - `durable: true` implies `wait` and shares its `timeoutMs` and `signal` (e.g. `{ durable: true, wait: { timeoutMs: 60_000 } }`).
223
+ If the time runs out first, `OperationTimeoutError` has `durable: true` and the copy continues server side.
224
+ - `cloud.workspaces.waitForDurable(operationOrId)` does the same for an operation you already hold, such as a
225
+ finished fork of a running workspace, whose result carries `durable` and `durability` the same way.
226
+ - `isDurable(op)`, `durabilityOf(op)` and `lastTiming.server.durable` / `.durability` read the fields; `formatTiming()`
227
+ prints `sealed on host, durable 435 ms later`. Results from before 0.12.0 servers count as durable.
228
+ - The errors reference on docs.shardflux.dev lists what `durable: true` can reject with.
229
+
208
230
  **Start deadlines.** A queued open, resume, restore or fork (`capacity_pending`) has a deadline 15 minutes after it
209
231
  was created. Its `error.details.deadline_at` carries the deadline; the `phase` progress event carries it as
210
232
  `deadlineAt` **(0.6.2+)**, and so does `OperationTimeoutError` when your wait ends first. A start still pending at the
@@ -295,8 +317,7 @@ operation keeps running server side (a queued start until its `deadlineAt`); wai
295
317
  same call. `templates.builds.waitForBuild()` waits the same way. A waited `open()` issues the first tool token
296
318
  together with the final workspace read, so the first tool call starts at once.
297
319
 
298
- On Node 26 the default fetch uses a fresh connection for every request (`Connection: close`);
299
- `SHARDFLUX_HTTP_KEEPALIVE=1`, or your own `fetch`, reuses connections.
320
+ From 0.12.0, Node 26 reuses TLS connections through a private HTTP/1.1 pool. Other runtimes keep native fetch. A supplied `fetch` stays unchanged. `SHARDFLUX_HTTP_KEEPALIVE=0` forces connection closure for diagnosis; `=1` opts into native pooling. No global dispatcher is installed.
300
321
 
301
322
  List and look up workspaces:
302
323
 
@@ -941,3 +962,44 @@ or no longer supported, it emits one warning:
941
962
  ## License
942
963
 
943
964
  Apache-2.0
965
+
966
+
967
+ ## Workspace integrations (0.12.0+)
968
+
969
+ ```ts
970
+ const ws = await cloud.workspaces.open({ key: 'qm/project-42', template: 'python-node-browser',
971
+ labels: { scope: 'project-42', owner: 'qm' }, idlePolicy: 'never' });
972
+ const matches = await cloud.workspaces.list({ labels: { scope: 'project-42' } });
973
+ await ws.setLabels({ scope: 'project-42', owner: 'qm' }); // replaces all labels; {} clears
974
+ await ws.setIdlePolicy('adaptive'); // null restores the inherited policy
975
+ const idle = await ws.idle(); // read-only, does not wake or record activity
976
+ await ws.keepalive(60); // seconds; never shortens an existing keepalive
977
+ await ws.suspendWhenIdle({ afterSeconds: 2 }); // one safe server-side request; no /idle + suspend race
978
+ ```
979
+
980
+ Labels are exact string pairs (at most 50, keys 1–64 and values 0–256 characters); they are metadata, not secrets.
981
+ A failed workspace is recoverable: `await ws.wake()` and ordinary auto-waking calls restart that ID using its existing disk. They do not create a replacement workspace. A failed attempt still raises its operation error; there is no unbounded restart loop.
982
+ After `await ws.delete({ wait: true })`, opening its key creates a **new ID**. The old ID, tokens and history remain deleted. A key stays reserved while deletion is in progress.
983
+
984
+ For a background command with a pipe (no PTY):
985
+
986
+ ```ts
987
+ const cell = ws.cell();
988
+ const session = await cell.exec.start({ session_id: 'qm-worker-1', argv: ['cat'], stdin_open: true });
989
+ let ack = await cell.exec.input(session.session_id, 'hello\n', { offset: 0 });
990
+ ack = await cell.exec.input(session.session_id, '', { offset: ack.offset, close: true });
991
+ ```
992
+
993
+ Input frames are at most 64 KiB. The acknowledged offset counts bytes accepted into the pipe, not bytes consumed by the program. A blocked writer may receive a partial acknowledgement; continue from the returned offset. Retrying the identical last frame is safe. `close` sends EOF only once the entire frame is accepted. Do not combine `stdin_open` with the existing one-shot `stdin` option. The transport writes no separate input log or payload file; program output and full-state snapshots retain their normal persistence. Full-state suspend preserves the pipe. Older hosts/guests explicitly refuse this feature until upgraded (`exec_stdin` / `exec_stdin.v1`). Running sessions count as work within the documented idle-command window; long-running services should declare keepalives or choose `never`.
994
+
995
+ ```ts
996
+ import { isWorkspaceGone, ShardfluxProtocolError } from '@shardflux/sdk';
997
+ // isWorkspaceGone(error) is true only for an explicit API workspace_deleted refusal.
998
+ // A protocol error, a cell 404, a scoped API 404 or a network failure is never deletion evidence.
999
+ ```
1000
+
1001
+ `ShardfluxProtocolError.source` is `api` or `cell` for SDK responses (`unknown` only for an error constructed without a source by older caller code). It remains a protocol failure even when `status === 404`. Never delete local data based on an HTTP status alone.
1002
+
1003
+ ### Repositories with a minimum release age
1004
+
1005
+ Keep your repository's age policy. Pin an exact version that has aged enough; do not exempt the entire `@shardflux/*` scope. The first SDK, 0.5.0, was published September 26, 2026 at 21:13:52 UTC and reaches seven days on October 3 at that time. New versions, including 0.12.0, need their own seven days after publication. Before then no registry setting on our side can make them eligible. `npm view @shardflux/sdk time --json` shows publication times. A targeted exception for an inspected exact release is a repository-owner decision, not an installation requirement we bypass. Older versions do not contain this release's fixes.
package/dist/account.d.ts CHANGED
@@ -102,7 +102,7 @@ export interface ShardfluxAccountOptions {
102
102
  sessionToken?: string;
103
103
  /** Default https://api.shardflux.dev. */
104
104
  baseUrl?: string;
105
- /** Default: the runtime's fetch, with `Connection: close` on Node 26 (see defaultFetch in http.ts). */
105
+ /** Default: pooled HTTP/1.1 on Node 26+, native fetch on other runtimes (see defaultFetch in http.ts). */
106
106
  fetch?: typeof fetch;
107
107
  userAgent?: string;
108
108
  /** Per-request timeout (ms), default 30 s. */
package/dist/account.js CHANGED
@@ -571,7 +571,7 @@ export class ShardfluxAccount {
571
571
  const b = body;
572
572
  if (b && typeof b === 'object' && typeof b.session_token === 'string') {
573
573
  if (!isSessionToken(b.session_token))
574
- throw new ShardfluxProtocolError('the API returned a session_token that is not a CLI session token (sfu_...)', 200);
574
+ throw new ShardfluxProtocolError('the API returned a session_token that is not a CLI session token (sfu_...)', 200, 'api');
575
575
  this.#token = b.session_token;
576
576
  const expiresAt = typeof b.session_expires_at === 'string' ? b.session_expires_at : '';
577
577
  await this.#onSessionToken?.({ token: b.session_token, expiresAt });
package/dist/cell.d.ts CHANGED
@@ -31,6 +31,7 @@ import type { ToolTokenManager } from './tokens.js';
31
31
  type S = components['schemas'];
32
32
  export type ExecStartRequest = S['ExecStartRequest'];
33
33
  export type ExecSession = S['ExecSession'];
34
+ export type ExecInputResult = S['ExecInputResult'];
34
35
  export type OutputEvent = S['OutputEvent'];
35
36
  export type PtyOpenRequest = S['PtyOpenRequest'];
36
37
  export type PtySession = S['PtySession'];
@@ -50,6 +51,8 @@ export type FilePatchResult = S['FilePatchResult'];
50
51
  /** A content SHA-256 (lowercase hex), or `absent` for a path that must not exist. */
51
52
  export type FileRevision = S['FileRevision'];
52
53
  export type WakeHintResult = S['WakeHintResult'];
54
+ export type IdleStatus = S['IdleStatus'];
55
+ export type KeepaliveResult = S['KeepaliveResult'];
53
56
  /** Where a running workspace's VM is: resident, frozen, hibernated or restoring. */
54
57
  export type Residency = WakeHintResult['residency'];
55
58
  /** `X-Served-From`: `disk` when a read was served from a suspended or hibernated workspace's disk. */
@@ -296,6 +299,14 @@ export declare class CellClient {
296
299
  follow?: boolean;
297
300
  signal?: AbortSignal;
298
301
  }) => Promise<AsyncGenerator<OutputEvent>>;
302
+ /** Write at the acknowledged offset (initially 0); a repeated identical last frame cannot duplicate input.
303
+ * At most 64 KiB per call. A partial acknowledgement requires continuing from the returned offset.
304
+ * Start with stdin_open: true. close sends EOF after this frame is fully accepted. */
305
+ input: (sessionId: string, data: string | Uint8Array, opts: {
306
+ offset: number;
307
+ close?: boolean;
308
+ signal?: AbortSignal;
309
+ }) => Promise<ExecInputResult>;
299
310
  signal: (sessionId: string, signal: Signal, onlyLeader?: boolean) => Promise<ExecSession>;
300
311
  cancel: (sessionId: string, graceMs?: number) => Promise<ExecSession>;
301
312
  /**
@@ -442,6 +453,10 @@ export declare class CellClient {
442
453
  * workspace answers 409 `workspace_not_running` (`Workspace.hint()` then wakes it in the background).
443
454
  */
444
455
  wakeHint(signal?: AbortSignal): Promise<WakeHintResult>;
456
+ /** Read idle signals without recording activity or waking the workspace. */
457
+ idle(signal?: AbortSignal): Promise<IdleStatus>;
458
+ /** Declare work for seconds (1..the server maximum). Never shortens a previous keepalive. */
459
+ keepalive(seconds: number, signal?: AbortSignal): Promise<KeepaliveResult>;
445
460
  /**
446
461
  * One page of the workspace's changes against its template (needs the `files` tool). File-first workspaces: each
447
462
  * execution's result lists what it changed (`changed`); this route is refused (NotSupportedForModeError).
package/dist/cell.js CHANGED
@@ -89,7 +89,7 @@ async function parseJson(res, what) {
89
89
  return JSON.parse(text);
90
90
  }
91
91
  catch {
92
- throw new ShardfluxProtocolError(`${what}: response is not JSON`, res.status);
92
+ throw new ShardfluxProtocolError(`${what}: response is not JSON`, res.status, 'cell');
93
93
  }
94
94
  }
95
95
  /** Delay before a retry of an execution: Retry-After (at most 30 s), else 0.5 s doubling to 8 s. */
@@ -282,7 +282,7 @@ export class CellClient {
282
282
  return (text.length === 0 ? undefined : JSON.parse(text));
283
283
  }
284
284
  catch {
285
- throw new ShardfluxProtocolError(`${method} ${path}: response is not JSON`, res.status);
285
+ throw new ShardfluxProtocolError(`${method} ${path}: response is not JSON`, res.status, 'cell');
286
286
  }
287
287
  }
288
288
  async #bytes(method, path, init = {}) {
@@ -329,9 +329,20 @@ export class CellClient {
329
329
  ...(opts.signal ? { signal: opts.signal } : {}),
330
330
  });
331
331
  if (!res.body)
332
- throw new ShardfluxProtocolError('exec output: empty body', res.status);
332
+ throw new ShardfluxProtocolError('exec output: empty body', res.status, 'cell');
333
333
  return ndjson(res.body);
334
334
  },
335
+ /** Write at the acknowledged offset (initially 0); a repeated identical last frame cannot duplicate input.
336
+ * At most 64 KiB per call. A partial acknowledgement requires continuing from the returned offset.
337
+ * Start with stdin_open: true. close sends EOF after this frame is fully accepted. */
338
+ input: (sessionId, data, opts) => {
339
+ const refusal = this.#noSessions('exec.input');
340
+ if (refusal)
341
+ return Promise.reject(refusal);
342
+ return this.#json('POST', this.#p('/v1/workspaces/{workspace_id}/exec/{session_id}/stdin', { session_id: sessionId }), {
343
+ json: { data: b64(data), offset: opts.offset, close: opts.close ?? false }, ...(opts.signal ? { signal: opts.signal } : {}),
344
+ });
345
+ },
335
346
  signal: (sessionId, signal, onlyLeader = false) => {
336
347
  const refusal = this.#noSessions('exec.signal');
337
348
  if (refusal)
@@ -404,9 +415,14 @@ export class CellClient {
404
415
  const maxReconnects = opts.maxReconnects ?? 10;
405
416
  for (;;) {
406
417
  let exited;
418
+ const streamAbort = new AbortController();
419
+ const signal = opts.signal ? AbortSignal.any([opts.signal, streamAbort.signal]) : streamAbort.signal;
420
+ let drainTimer;
407
421
  try {
408
- const events = await this.exec.output(sessionId, { stdoutOffset: so, stderrOffset: se, follow: true, ...(opts.signal ? { signal: opts.signal } : {}) });
422
+ const events = await this.exec.output(sessionId, { stdoutOffset: so, stderrOffset: se, follow: true, signal });
409
423
  for await (const ev of events) {
424
+ if (exited)
425
+ continue;
410
426
  if (ev.type === 'output' && ev.data !== undefined) {
411
427
  const bytes = unb64(ev.data);
412
428
  const start = ev.offset ?? (ev.stream === 'stderr' ? se : so);
@@ -426,7 +442,9 @@ export class CellClient {
426
442
  }
427
443
  else if (ev.type === 'exit') {
428
444
  exited = ev.session ?? (await this.exec.get(sessionId));
429
- break;
445
+ // Consume the terminal HTTP framing so the connection can be reused.
446
+ // A peer that never closes after exit must not hold the completed run forever.
447
+ drainTimer = setTimeout(() => streamAbort.abort(), 250);
430
448
  }
431
449
  else if (ev.type === 'error' && ev.error) {
432
450
  throw new ShardfluxApiError(502, ev.error, 'cell');
@@ -437,9 +455,12 @@ export class CellClient {
437
455
  if (opts.signal?.aborted || this.closed)
438
456
  throw e;
439
457
  const transient = !(e instanceof ShardfluxApiError) || e.retryable;
440
- if (!transient || reconnects >= maxReconnects)
458
+ if (!exited && (!transient || reconnects >= maxReconnects))
441
459
  throw e;
442
460
  }
461
+ finally {
462
+ clearTimeout(drainTimer);
463
+ }
443
464
  if (exited) {
444
465
  session = exited;
445
466
  break;
@@ -449,7 +470,7 @@ export class CellClient {
449
470
  throw this.#closer.signal.reason;
450
471
  reconnects += 1;
451
472
  if (reconnects > maxReconnects)
452
- throw new ShardfluxProtocolError(`exec ${sessionId}: output stream kept dropping`, 0);
473
+ throw new ShardfluxProtocolError(`exec ${sessionId}: output stream kept dropping`, 0, 'cell');
453
474
  await this.#opts.sleep(Math.min(2_000, 100 * 2 ** reconnects));
454
475
  const now = await this.exec.get(sessionId);
455
476
  if (now.state !== 'starting' && now.state !== 'running' && so >= now.stdout_size && se >= now.stderr_size) {
@@ -708,7 +729,7 @@ export class CellClient {
708
729
  finish(new ShardfluxApiError(502, m.error, 'cell'));
709
730
  }
710
731
  });
711
- ws.addEventListener('error', () => finish(new ShardfluxProtocolError('pty attach WebSocket failed', 0)));
732
+ ws.addEventListener('error', () => finish(new ShardfluxProtocolError('pty attach WebSocket failed', 0, 'cell')));
712
733
  ws.addEventListener('close', () => finish());
713
734
  if (this.closed)
714
735
  onClose();
@@ -875,6 +896,14 @@ export class CellClient {
875
896
  return Promise.resolve({ residency: 'resident' });
876
897
  return this.#json('POST', this.#p('/v1/workspaces/{workspace_id}/wake-hint'), { wake: false, busy: false, timeoutMs: 10_000, ...(signal ? { signal } : {}) });
877
898
  }
899
+ /** Read idle signals without recording activity or waking the workspace. */
900
+ idle(signal) {
901
+ return this.#json('GET', this.#p('/v1/workspaces/{workspace_id}/idle'), { wake: false, ...(signal ? { signal } : {}) });
902
+ }
903
+ /** Declare work for seconds (1..the server maximum). Never shortens a previous keepalive. */
904
+ keepalive(seconds, signal) {
905
+ return this.#json('POST', this.#p('/v1/workspaces/{workspace_id}/keepalive'), { json: { seconds }, wake: false, ...(signal ? { signal } : {}) });
906
+ }
878
907
  // ---- changes against the template (layered workspaces) ------------------
879
908
  /**
880
909
  * One page of the workspace's changes against its template (needs the `files` tool). File-first workspaces: each
package/dist/client.d.ts CHANGED
@@ -22,7 +22,7 @@ import type { SaveAsTemplateParams, SaveAsTemplateResponse } from './templates.j
22
22
  import { UsageApi } from './usage.js';
23
23
  import { VolumesApi } from './volumes.js';
24
24
  import type { VersionCheckOption } from './version-check.js';
25
- import type { FinishedOperation, InternalLifecycleOptions, LifecycleOptions, ResumeOptions, WaitedLifecycleOptions, WaitedResumeOptions } from './lifecycle.js';
25
+ import type { FinishedOperation, InternalLifecycleOptions, LifecycleOptions, ResumeOptions, SuspendOptions, WaitedLifecycleOptions, WaitedResumeOptions, WaitedSuspendOptions } from './lifecycle.js';
26
26
  import type { ProgressListener } from './progress.js';
27
27
  import { CaptureRegistry } from './capture.js';
28
28
  import type { FeedbackReceipt, SendFeedbackParams } from './feedback.js';
@@ -121,7 +121,7 @@ export interface ShardfluxOptions {
121
121
  apiKey: string;
122
122
  /** Default https://api.shardflux.dev (override with `baseUrl`). */
123
123
  baseUrl?: string;
124
- /** Default: the runtime's fetch; on Node 26 every request uses a fresh connection (`Connection: close`) unless SHARDFLUX_HTTP_KEEPALIVE=1. */
124
+ /** Default: pooled HTTP/1.1 on Node 26+, native fetch on other runtimes (see defaultFetch in http.ts). */
125
125
  fetch?: typeof fetch;
126
126
  userAgent?: string;
127
127
  /** Per-request timeout (ms), default 30 s. */
@@ -172,7 +172,11 @@ export interface WaitOptions {
172
172
  /** Progress while waiting: each observed state (queued, capacity_pending, running with its reason), retries, and `done` with the timing. */
173
173
  onProgress?: ProgressListener;
174
174
  }
175
+ export type IdlePolicy = 'adaptive' | 'never' | `fixed:${number}`;
175
176
  export interface OpenParams {
177
+ /** Searchable metadata; supplied labels replace the existing map. */
178
+ labels?: Record<string, string>;
179
+ idlePolicy?: IdlePolicy;
176
180
  key: string;
177
181
  template: string;
178
182
  caps?: Caps;
@@ -225,6 +229,8 @@ export interface OpenParams {
225
229
  onProgress?: ProgressListener;
226
230
  }
227
231
  export interface ListParams {
232
+ /** All supplied labels must match exactly. */
233
+ labels?: Record<string, string>;
228
234
  state?: WorkspaceView['observed_state'];
229
235
  desiredState?: WorkspaceView['desired_state'];
230
236
  keyPrefix?: string;
@@ -292,6 +298,15 @@ export declare class WorkspacesApi {
292
298
  * (`err.retryable` true: nothing was started; send it again); the SDK does not retry it.
293
299
  */
294
300
  waitForOperation(operationId: string, opts?: WaitOptions): Promise<Operation>;
301
+ /**
302
+ * Waits for the durable copy of a succeeded suspend or fork (0.12.0+): resolves with the operation once
303
+ * `result.durable` is true, typically within a second of the suspend. An operation already durable (or one whose result
304
+ * predates the field) resolves at once without a request. Polls GET /v1/operations/{id} (`pollIntervalMs`, default
305
+ * 250 ms, doubling up to `maxPollIntervalMs`, default 1 000 ms). Throws DurabilityLostError when the copy cannot be
306
+ * made (`durability.state` `lost`), OperationFailedError if the operation did not succeed, and OperationTimeoutError
307
+ * (`durable: true`) after `timeoutMs` (default 300 000 ms; the copy continues server side). `signal` aborts the wait.
308
+ */
309
+ waitForDurable(operation: string | Operation, opts?: WaitOptions): Promise<FinishedOperation>;
295
310
  /** One operation (GET /v1/operations/{id}); lifecycle operations stay pollable after a workspace is deleted. */
296
311
  getOperation(operationId: string, opts?: {
297
312
  signal?: AbortSignal;
@@ -300,6 +315,10 @@ export declare class WorkspacesApi {
300
315
  agentLabel?: string;
301
316
  tools?: ToolName[];
302
317
  }): Promise<Workspace>;
318
+ /** Replace labels. An empty map clears them. */
319
+ setLabels(workspaceId: string, labels: Record<string, string>): Promise<Workspace>;
320
+ /** null clears the override, restoring the template or platform policy. */
321
+ setIdlePolicy(workspaceId: string, idlePolicy: IdlePolicy | null): Promise<Workspace>;
303
322
  list(params?: ListParams): Promise<Page<Workspace>>;
304
323
  /** Iterates every page. */
305
324
  listAll(params?: Omit<ListParams, 'cursor'>): AsyncGenerator<Workspace>;
@@ -317,10 +336,12 @@ export declare class WorkspacesApi {
317
336
  delete(workspaceId: string, opts?: LifecycleOptions): Promise<Operation>;
318
337
  /**
319
338
  * Suspends the workspace (memory and processes checkpointed). Resolves when the suspend is REQUESTED: the returned
320
- * operation is usually still `queued`. Pass `{ wait: true }` to resolve once it has FINISHED (`succeeded`).
339
+ * operation is usually still `queued`. Pass `{ wait: true }` to resolve once it has FINISHED (`succeeded`): the
340
+ * workspace is sealed on its host, typically in a few hundred ms, and `result.durable` turns true when the copy lands
341
+ * in durable storage, typically within a second. `{ durable: true }` (0.12.0+) resolves only then (see SuspendOptions).
321
342
  */
322
- suspend(workspaceId: string, opts: WaitedLifecycleOptions): Promise<FinishedOperation>;
323
- suspend(workspaceId: string, opts?: LifecycleOptions): Promise<Operation>;
343
+ suspend(workspaceId: string, opts: WaitedSuspendOptions): Promise<FinishedOperation>;
344
+ suspend(workspaceId: string, opts?: SuspendOptions): Promise<Operation>;
324
345
  /**
325
346
  * Resumes a suspended workspace. Resolves when the resume is requested; with `wait`, once the workspace runs. With
326
347
  * `wait` (0.9.0) the request is held by the server until the workspace runs (one request, timing phase
package/dist/client.js CHANGED
@@ -1,4 +1,4 @@
1
- import { OperationFailedError, OperationTimeoutError, ShardfluxApiError, ShardfluxProtocolError } from "./errors.js";
1
+ import { DurabilityLostError, OperationFailedError, OperationTimeoutError, ShardfluxApiError, ShardfluxProtocolError } from "./errors.js";
2
2
  import { HttpClient, SDK_VERSION, SERVER_WAIT_MAX_S, defaultFetch, defaultSleep, pollWithWait, randomId } from "./http.js";
3
3
  import { Workspace } from "./workspace.js";
4
4
  import { AuditApi } from "./audit.js";
@@ -9,7 +9,7 @@ import { UsageApi } from "./usage.js";
9
9
  import { VolumesApi } from "./volumes.js";
10
10
  import { versionCheckHook } from "./version-check.js";
11
11
  import { AFTER_WAIT, HELD_RESUME, TRACE, runLifecycle, waitOptionsOf } from "./lifecycle.js";
12
- import { Trace, combineListeners, traced } from "./progress.js";
12
+ import { Trace, combineListeners, durabilityOf, isDurable, traced } from "./progress.js";
13
13
  import { CaptureRegistry } from "./capture.js";
14
14
  import { sendFeedback } from "./feedback.js";
15
15
  /**
@@ -96,6 +96,10 @@ export class WorkspacesApi {
96
96
  body.secrets = params.secrets;
97
97
  if (params.inputs !== undefined)
98
98
  body.inputs = params.inputs;
99
+ if (params.labels !== undefined)
100
+ body.labels = params.labels;
101
+ if (params.idlePolicy !== undefined)
102
+ body.idle_policy = params.idlePolicy;
99
103
  if (params.lifetime !== undefined)
100
104
  body.lifetime = params.lifetime;
101
105
  if (params.mode !== undefined)
@@ -246,6 +250,61 @@ export class WorkspacesApi {
246
250
  interval = Math.min(maxInterval, interval * 2);
247
251
  }
248
252
  }
253
+ /**
254
+ * Waits for the durable copy of a succeeded suspend or fork (0.12.0+): resolves with the operation once
255
+ * `result.durable` is true, typically within a second of the suspend. An operation already durable (or one whose result
256
+ * predates the field) resolves at once without a request. Polls GET /v1/operations/{id} (`pollIntervalMs`, default
257
+ * 250 ms, doubling up to `maxPollIntervalMs`, default 1 000 ms). Throws DurabilityLostError when the copy cannot be
258
+ * made (`durability.state` `lost`), OperationFailedError if the operation did not succeed, and OperationTimeoutError
259
+ * (`durable: true`) after `timeoutMs` (default 300 000 ms; the copy continues server side). `signal` aborts the wait.
260
+ */
261
+ async waitForDurable(operation, opts = {}) {
262
+ const inherited = opts[TRACE];
263
+ const id = typeof operation === 'string' ? operation : operation.id;
264
+ const trace = inherited ?? new Trace('wait', combineListeners(this.#ctx().onProgress, opts.onProgress), { operationId: id });
265
+ const run = async () => {
266
+ const { sleep, http } = this.#ctx();
267
+ const timeoutMs = opts.timeoutMs ?? 300_000;
268
+ const maxInterval = opts.maxPollIntervalMs ?? 1_000;
269
+ let interval = opts.pollIntervalMs ?? 250;
270
+ const started = Date.now();
271
+ const aborted = () => (opts.signal?.reason instanceof Error ? opts.signal.reason : new Error('aborted'));
272
+ let op = typeof operation === 'string' ? null : operation;
273
+ for (;;) {
274
+ if (opts.signal?.aborted)
275
+ throw aborted();
276
+ if (op === null) {
277
+ const { body } = await pollWithWait(http, `/v1/operations/${encodeURIComponent(id)}`, this.#ctx().authorization, 0, opts.signal, trace.onRetry);
278
+ op = body.operation;
279
+ trace.observe(op);
280
+ }
281
+ if (op.state !== 'succeeded') {
282
+ if (TERMINAL.has(op.state))
283
+ throw new OperationFailedError(op);
284
+ // Not finished yet (a fork or suspend passed by id): wait for it first, then for its copy.
285
+ op = await this.waitForOperation(id, { ...opts, timeoutMs: Math.max(1, timeoutMs - (Date.now() - started)), [TRACE]: trace });
286
+ }
287
+ if (isDurable(op) !== false)
288
+ return op;
289
+ const durability = durabilityOf(op);
290
+ if (durability?.state === 'lost')
291
+ throw new DurabilityLostError(op);
292
+ trace.phase('durable', durability?.overdueAt ? 'overdue' : null);
293
+ const waited = Date.now() - started;
294
+ if (waited >= timeoutMs)
295
+ throw new OperationTimeoutError(op, waited, true);
296
+ const jitter = interval * 0.2 * (Math.random() * 2 - 1);
297
+ const delay = Math.max(10, Math.min(interval + jitter, timeoutMs - waited));
298
+ await (opts.signal ? abortableSleep(sleep, delay, opts.signal) : sleep(delay));
299
+ interval = Math.min(maxInterval, interval * 2);
300
+ op = null;
301
+ }
302
+ };
303
+ if (inherited)
304
+ return run();
305
+ trace.phase('request');
306
+ return traced(trace, run);
307
+ }
249
308
  /** One operation (GET /v1/operations/{id}); lifecycle operations stay pollable after a workspace is deleted. */
250
309
  async getOperation(operationId, opts = {}) {
251
310
  const { operation } = await this.#http.json('GET', `/v1/operations/${encodeURIComponent(operationId)}`, opts.signal ? { signal: opts.signal } : {}, this.#auth);
@@ -257,12 +316,21 @@ export class WorkspacesApi {
257
316
  async get(workspaceId, opts = {}) {
258
317
  return this.#wrap(await this.#getView(workspaceId), opts);
259
318
  }
319
+ /** Replace labels. An empty map clears them. */
320
+ async setLabels(workspaceId, labels) {
321
+ return this.#wrap(await this.#http.json('PUT', `/v1/workspaces/${encodeURIComponent(workspaceId)}/labels`, { json: { labels } }, this.#auth));
322
+ }
323
+ /** null clears the override, restoring the template or platform policy. */
324
+ async setIdlePolicy(workspaceId, idlePolicy) {
325
+ return this.#wrap(await this.#http.json('PUT', `/v1/workspaces/${encodeURIComponent(workspaceId)}/idle-policy`, { json: { idle_policy: idlePolicy } }, this.#auth));
326
+ }
260
327
  async list(params = {}) {
261
328
  const page = await this.#http.json('GET', '/v1/workspaces', {
262
329
  query: {
263
330
  state: params.state,
264
331
  desired_state: params.desiredState,
265
332
  key_prefix: params.keyPrefix,
333
+ labels: params.labels === undefined ? undefined : JSON.stringify(params.labels),
266
334
  project_id: params.projectId,
267
335
  organization_id: params.organizationId,
268
336
  include_deleted: params.includeDeleted,
@@ -406,11 +474,11 @@ export class WorkspacesApi {
406
474
  const res = await this.#http.jsonWithHeaders('POST', `/v1/workspaces/${encodeURIComponent(workspaceId)}/resume`, init, this.#auth);
407
475
  const b = res.body;
408
476
  if (!b || typeof b !== 'object' || !b.workspace)
409
- throw new ShardfluxProtocolError('resume: response has no workspace', res.status);
477
+ throw new ShardfluxProtocolError('resume: response has no workspace', res.status, 'api');
410
478
  const pa = res.headers.get('preference-applied');
411
479
  const ready = held && res.status === 200 && pa !== null && /\bwait\s*=/i.test(pa);
412
480
  if (!ready && !b.operation)
413
- throw new ShardfluxProtocolError('resume: response has no operation', res.status);
481
+ throw new ShardfluxProtocolError('resume: response has no operation', res.status, 'api');
414
482
  const toolToken = ready ? (b.tool_token ?? null) : null;
415
483
  this.#noteToken(toolToken);
416
484
  return { ready, operation: b.operation ?? null, workspace: b.workspace, toolToken, requestId: res.headers.get('x-request-id') };
package/dist/errors.d.ts CHANGED
@@ -1,6 +1,6 @@
1
1
  import type { components as AppComponents } from './generated/app-api.js';
2
2
  import type { components as CellComponents } from './generated/cell-api.js';
3
- import type { LifecycleTiming } from './progress.js';
3
+ import type { Durability, LifecycleTiming } from './progress.js';
4
4
  export type AppErrorBody = AppComponents['schemas']['ErrorBody'];
5
5
  export type AppErrorCode = AppErrorBody['error']['code'];
6
6
  export type CellErrorCode = CellComponents['schemas']['ErrorCode'];
@@ -141,8 +141,12 @@ export declare function apiError(status: number, body: ErrorBodyLike, source: 'a
141
141
  /** The response was not the documented shape (e.g. a proxy error page). */
142
142
  export declare class ShardfluxProtocolError extends Error {
143
143
  readonly status: number;
144
- constructor(message: string, status: number);
144
+ /** The responding surface; unknown only for errors constructed by older callers. Never proof of deletion. */
145
+ readonly source: 'api' | 'cell' | 'unknown';
146
+ constructor(message: string, status: number, source?: 'api' | 'cell' | 'unknown');
145
147
  }
148
+ /** True only for an explicit API tombstone refusal. A 404, even a structured one, may be routing or scope. */
149
+ export declare function isWorkspaceGone(error: unknown): boolean;
146
150
  type Operation = AppComponents['schemas']['Operation'];
147
151
  /**
148
152
  * Waiting for an operation ran out of time. The operation keeps running server side:
@@ -163,13 +167,24 @@ export declare class OperationTimeoutError extends Error {
163
167
  readonly waitedMs: number;
164
168
  /** Where the time went: client phases, retries and the operation's own server timing so far. */
165
169
  timing: LifecycleTiming | undefined;
166
- constructor(op: Operation, waitedMs: number);
170
+ /**
171
+ * `durable` (0.12.0+): the wait was for the durable copy of a succeeded suspend or fork (`suspend({ durable: true })`,
172
+ * `waitForDurable()`); `lastState` is then `succeeded` and the copy continues server side.
173
+ */
174
+ readonly durable: boolean;
175
+ constructor(op: Operation, waitedMs: number, durable?: boolean);
167
176
  }
168
- /** The operation reached `failed` or `canceled`. */
177
+ /**
178
+ * The operation reached `failed` or `canceled`. A suspend-when-idle that found the workspace active (a tool call after
179
+ * the request, an attached stream) is `canceled` with `errorCode` `workspace_active` (`workspaceActive` true, 0.12.0+):
180
+ * nothing changed and the workspace keeps running. `waitUntilReady()` and `wake()` treat it as running, not as a failure.
181
+ */
169
182
  export declare class OperationFailedError extends Error {
170
183
  readonly operation: Operation;
171
184
  readonly operationId: string;
172
185
  readonly errorCode: string | null;
186
+ /** 0.12.0+: a suspend-when-idle canceled because the workspace was active; it keeps running, nothing changed. */
187
+ readonly workspaceActive: boolean;
173
188
  /**
174
189
  * `operation.error.retryable` (false when absent): the same request may succeed later. `capacity_unavailable` is
175
190
  * retryable: the start passed its deadline and nothing was started; send it again. The SDK never retries a failed
@@ -180,4 +195,18 @@ export declare class OperationFailedError extends Error {
180
195
  timing: LifecycleTiming | undefined;
181
196
  constructor(op: Operation);
182
197
  }
198
+ /**
199
+ * `suspend({ durable: true })` or `waitForDurable()` (0.12.0+): the suspend (or fork) succeeded, but its durable copy
200
+ * could not be made (`durability.state` `lost`). `durability` has the reason; see the lifecycle reference for what the
201
+ * next resume restores.
202
+ */
203
+ export declare class DurabilityLostError extends Error {
204
+ readonly operation: Operation;
205
+ readonly operationId: string;
206
+ readonly workspaceId: string | null;
207
+ readonly durability: Durability;
208
+ /** Where the time went. */
209
+ timing: LifecycleTiming | undefined;
210
+ constructor(op: Operation);
211
+ }
183
212
  export {};