@shardflux/sdk 0.6.0 → 0.6.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -3,7 +3,33 @@
3
3
  Every API the README shows is available from the version named here. Below 1.0, a minor release may break
4
4
  compatibility; breaking changes are marked **Breaking**.
5
5
 
6
- ## 0.6.0 (not yet published; npm `latest` is 0.5.0)
6
+ ## 0.6.2 (not yet published; npm `latest` is 0.6.1)
7
+
8
+ ### Starts that wait for capacity end
9
+
10
+ The API no longer lets a start (open, resume, restore, fork) wait in `capacity_pending` forever. One that no host
11
+ could admit 15 minutes after it was created fails with `capacity_unavailable` and `retryable: true`: nothing was
12
+ started, the concurrency slot is released, and a suspended workspace stays suspended with its state. Before, a VM could
13
+ boot (and bill) long after every wait had given up.
14
+
15
+ - `OperationFailedError.retryable`: the operation error's `retryable` flag (`false` when absent). `true` for
16
+ `capacity_unavailable` (retry later), `false` for a definitive failure. The SDK does not retry it itself.
17
+ - `OperationTimeoutError.deadlineAt`: when a wait gives up while the start is still `capacity_pending`, the time it
18
+ stops waiting for a host (`error.details.deadline_at`); the message says so. Null in other states.
19
+ - `onProgress`: a `capacity_pending` `phase` event carries `deadlineAt` when the API reports it.
20
+ - Docs: waits no longer say a pending start continues indefinitely.
21
+
22
+ ## 0.6.1 (2026-09-28)
23
+
24
+ - `formatTiming()` joined two phases that followed each other with `∥` (ran together) when their 0.1 ms times summed
25
+ with a floating-point error (1000.2 + 300.1 > 1300.3); it now prints `→`.
26
+ - `open()`: the wait's `signal` also aborts the held request (up to 20 s); before, it only ended the polls after it,
27
+ and an abort during the held request was retried as a network failure.
28
+ - `waitForOperation()` on its own records its first poll as a `request` phase; before, a held first poll's time
29
+ belonged to no phase (a 1 s wait could show `queued 0 ms`).
30
+ - README: the timing example is the formatter's real output (durations under a second print in ms).
31
+
32
+ ## 0.6.0 (2026-09-28)
7
33
 
8
34
  ### Lifecycle calls: requested or finished
9
35
 
package/README.md CHANGED
@@ -9,7 +9,7 @@ tools (exec, files, processes, PTY, git, browser) that plug into any model provi
9
9
  > **Early access.** Shardflux is in early access. The API is versioned (`/v1`), but this SDK is
10
10
  > below 1.0: a minor release may contain breaking changes (see [Compatibility](#compatibility)).
11
11
 
12
- > **Versions.** This README describes 0.6.0. Anything marked **(0.6.0+)** is not in 0.5.0;
12
+ > **Versions.** This README describes 0.6.2. Anything marked **(0.6.0+)** is not in 0.5.0;
13
13
  > [CHANGELOG.md](./CHANGELOG.md) lists what each version added. Check yours with
14
14
  > `npm ls @shardflux/sdk` or the exported `SDK_VERSION`.
15
15
 
@@ -127,6 +127,23 @@ means depends on `wait`:
127
127
  wait again with `cloud.workspaces.waitForOperation(err.operationId)`. Without `wait`, the returned operation is the
128
128
  handle for the work in progress: pass its `id` to `waitForOperation()` when you need it finished.
129
129
 
130
+ **A start waits for capacity for at most 15 minutes.** An open, resume, restore or fork that no host can admit yet
131
+ waits in `capacity_pending`. Its `error.details.deadline_at` says when it gives up; the `phase` progress event carries
132
+ it as `deadlineAt` **(0.6.2+)**, and so does `OperationTimeoutError` when your wait ends first. A start still pending
133
+ at the deadline fails with `capacity_unavailable`: nothing was started, the concurrency slot is released, and a
134
+ suspended workspace stays suspended with its state. `OperationFailedError.retryable` **(0.6.2+)** is `true` for it,
135
+ so you can tell "retry later" from a definitive failure. The SDK does not retry it for you.
136
+
137
+ ```ts
138
+ try {
139
+ await workspace.resume({ wait: true });
140
+ } catch (err) {
141
+ if (err instanceof OperationFailedError && err.retryable) {
142
+ // err.errorCode === 'capacity_unavailable': no host had room; nothing changed. Try again later.
143
+ } else throw err;
144
+ }
145
+ ```
146
+
130
147
  **Suspended workspaces wake on use.** A tool call on a suspended workspace resumes it (or joins the resume or open
131
148
  already running), then runs. A call made during a suspend or resume waits for the transition to finish. The call
132
149
  never runs twice: the cell executes nothing it refused.
@@ -148,9 +165,9 @@ await cell.exec.run(['make', 'test']); // resumes the wo
148
165
  the response until the operation changes (`Prefer: wait`, at most 20 s per request), so completion arrives within
149
166
  one round trip of the commit. Against an API without bounded waits (or with `serverWait: false`) it polls with
150
167
  backoff (250 ms doubling to 5 s, ±20 % jitter). If the timeout passes, it throws `OperationTimeoutError` and the
151
- operation keeps running server side; wait for it again with the same call. `templates.builds.waitForBuild()` waits
152
- the same way. A waited `open()` issues the first tool token together with the final workspace read, so the first
153
- tool call starts at once.
168
+ operation keeps running server side (a start waiting for capacity until its `deadlineAt`); wait for it again with the
169
+ same call. `templates.builds.waitForBuild()` waits the same way. A waited `open()` issues the first tool token
170
+ together with the final workspace read, so the first tool call starts at once.
154
171
 
155
172
  On Node 26 the default fetch sends `Connection: close`: its bundled undici 8 can stall a request on a reused
156
173
  keep-alive connection for tens of seconds. Pass your own `fetch`, or set `SHARDFLUX_HTTP_KEEPALIVE=1`, to change that.
@@ -184,9 +201,9 @@ A slow open (34 s instead of the usual second) then reads, for example:
184
201
 
185
202
  ```text
186
203
  open 34.18 s, succeeded (workspace 01a0e5a8-3edd-74ba-b489-d62b8925e342, operation 01a0e5a8-3ef0-7ecb-975e-dff2d5ca6e33)
187
- client: request 20.01 s (held) → capacity_pending 13.52 s (no_ready_host) → running 0.59 s → view 42 ms ∥ token 61 ms
188
- server: queued 33.40 s, ran 0.62 s, total 34.02 s; start warm, boot to ready 79 ms
189
- outside the server: 0.16 s
204
+ client: request 20.01 s (held) → capacity_pending 13.52 s (no_ready_host) → running 590 ms → view 42 ms ∥ token 61 ms
205
+ server: queued 33.40 s, ran 620 ms, total 34.02 s; start warm, boot to ready 79 ms
206
+ outside the server: 161 ms
190
207
  ```
191
208
 
192
209
  Here the time went to waiting for a host with capacity; the network and the VM start were fast.
@@ -338,8 +355,11 @@ await workspace.cell().exec.run(['python3', 'agent.py']); // sees $OPEN
338
355
  `message`, `requestId`, `retryable`, `details`, `operationId`, `retryAfterSeconds`, and `reason`
339
356
  (`details.reason`, e.g. `not_session`, `draft_not_found`, `legacy_disk_layout`; see `KnownErrorReason`).
340
357
  - `OperationFailedError`: an awaited operation ended `failed` or `canceled` (`errorCode`,
341
- `operation`, `timing`).
342
- - `OperationTimeoutError`: waiting gave up; the operation continues (`operationId`, `timing`).
358
+ `retryable` **(0.6.2+)**, `operation`, `timing`). `retryable` is the operation error's own flag: `true` for
359
+ `capacity_unavailable` (no host could admit the start before its deadline; retry later), `false` for a definitive
360
+ failure.
361
+ - `OperationTimeoutError`: waiting gave up; the operation continues (`operationId`, `lastState`, `lastReason`,
362
+ `deadlineAt` **(0.6.2+)** while it waits for capacity, `timing`).
343
363
  - `ShardfluxProtocolError`: a response was not the documented shape.
344
364
 
345
365
  Treat unknown error codes and reasons as generic errors: show `message`, and use `retryable`.
package/dist/client.d.ts CHANGED
@@ -92,7 +92,11 @@ export interface ForkTarget {
92
92
  lifetime?: WorkspaceLifetime;
93
93
  }
94
94
  export interface WaitOptions {
95
- /** Give up waiting after this long (default 300 000 ms); the operation continues server side. */
95
+ /**
96
+ * Give up waiting after this long (default 300 000 ms); the operation continues server side. A start waiting for
97
+ * capacity (`capacity_pending`) does so until its deadline (`error.details.deadline_at`, 15 minutes after it was
98
+ * created) and then fails with `capacity_unavailable` (retryable; nothing was started).
99
+ */
96
100
  timeoutMs?: number;
97
101
  /** First poll delay (default 250 ms); doubles up to `maxPollIntervalMs` with jitter. Used when the server does not wait. */
98
102
  pollIntervalMs?: number;
@@ -200,7 +204,9 @@ export declare class WorkspacesApi {
200
204
  * that does not wait is polled with bounded exponential backoff (+-20 % jitter). Resolves when the operation
201
205
  * succeeds; throws OperationFailedError when it fails or is canceled, OperationTimeoutError after `timeoutMs` (the
202
206
  * operation keeps running and can be awaited again). Both errors carry the wait's `timing`; `onProgress` sees each
203
- * state change (queued, capacity_pending, running, with the server's reason) as it is observed.
207
+ * state change (queued, capacity_pending, running, with the server's reason) as it is observed. A start waiting in
208
+ * capacity_pending gives up at its deadline (`deadlineAt` on the phase event) and fails with `capacity_unavailable`
209
+ * (`err.retryable` true: nothing was started, retry later); the SDK does not retry it.
204
210
  */
205
211
  waitForOperation(operationId: string, opts?: WaitOptions): Promise<Operation>;
206
212
  /** One operation (GET /v1/operations/{id}); lifecycle operations stay pollable after a workspace is deleted. */
package/dist/client.js CHANGED
@@ -81,7 +81,8 @@ export class WorkspacesApi {
81
81
  const waitOpts = params.wait === false ? null : (params.wait ?? {});
82
82
  const trace = new Trace('open', combineListeners(this.#ctx().onProgress, params.onProgress, waitOpts?.onProgress));
83
83
  return traced(trace, async () => {
84
- const init = { json: body, idempotencyKey: params.idempotencyKey ?? randomId('open-'), onRetry: trace.onRetry };
84
+ // The wait's signal also ends the held request (it can take up to 20 s), not only the polls after it.
85
+ const init = { json: body, idempotencyKey: params.idempotencyKey ?? randomId('open-'), onRetry: trace.onRetry, ...(waitOpts?.signal ? { signal: waitOpts.signal } : {}) };
85
86
  const started = Date.now();
86
87
  const timeoutMs = waitOpts?.timeoutMs ?? 300_000;
87
88
  if (waitOpts && waitOpts.serverWait !== false) {
@@ -169,13 +170,19 @@ export class WorkspacesApi {
169
170
  * that does not wait is polled with bounded exponential backoff (+-20 % jitter). Resolves when the operation
170
171
  * succeeds; throws OperationFailedError when it fails or is canceled, OperationTimeoutError after `timeoutMs` (the
171
172
  * operation keeps running and can be awaited again). Both errors carry the wait's `timing`; `onProgress` sees each
172
- * state change (queued, capacity_pending, running, with the server's reason) as it is observed.
173
+ * state change (queued, capacity_pending, running, with the server's reason) as it is observed. A start waiting in
174
+ * capacity_pending gives up at its deadline (`deadlineAt` on the phase event) and fails with `capacity_unavailable`
175
+ * (`err.retryable` true: nothing was started, retry later); the SDK does not retry it.
173
176
  */
174
177
  async waitForOperation(operationId, opts = {}) {
175
178
  const inherited = opts[TRACE];
176
179
  const trace = inherited ?? new Trace('wait', combineListeners(this.#ctx().onProgress, opts.onProgress), { operationId });
177
180
  const run = () => this.#poll(operationId, opts, trace);
178
- return inherited ? run() : traced(trace, run);
181
+ if (inherited)
182
+ return run();
183
+ // A wait on its own starts with its first poll: a held poll can take seconds before the first state is known.
184
+ trace.phase('request');
185
+ return traced(trace, run);
179
186
  }
180
187
  async #poll(operationId, opts, trace) {
181
188
  const { sleep, http, authorization } = this.#ctx();
package/dist/errors.d.ts CHANGED
@@ -53,13 +53,19 @@ type Operation = AppComponents['schemas']['Operation'];
53
53
  /**
54
54
  * Waiting for an operation ran out of time. The operation keeps running server side:
55
55
  * resume with `cloud.workspaces.waitForOperation(err.operationId)` or call open() again
56
- * (it returns the same operation while it is active).
56
+ * (it returns the same operation while it is active). A start still waiting for capacity
57
+ * (`capacity_pending`) gives up at `deadlineAt` and then fails with `capacity_unavailable`.
57
58
  */
58
59
  export declare class OperationTimeoutError extends Error {
59
60
  readonly operationId: string;
60
61
  readonly workspaceId: string | null;
61
62
  readonly lastState: Operation['state'];
62
63
  readonly lastReason: string | null;
64
+ /**
65
+ * When the operation was last `capacity_pending`: when it gives up waiting for a host (`error.details.deadline_at`,
66
+ * RFC 3339) and fails with `capacity_unavailable`. Null in other states or from an API that does not report it.
67
+ */
68
+ readonly deadlineAt: string | null;
63
69
  readonly waitedMs: number;
64
70
  /** Where the time went: client phases, retries and the operation's own server timing so far. */
65
71
  timing: LifecycleTiming | undefined;
@@ -70,6 +76,12 @@ export declare class OperationFailedError extends Error {
70
76
  readonly operation: Operation;
71
77
  readonly operationId: string;
72
78
  readonly errorCode: string | null;
79
+ /**
80
+ * `operation.error.retryable` (false when absent): the same request may succeed later. `capacity_unavailable` is
81
+ * retryable: no host could admit the start before its deadline, nothing was started, retry later. The SDK never
82
+ * retries a failed operation itself.
83
+ */
84
+ readonly retryable: boolean;
73
85
  /** Where the time went before the operation failed. */
74
86
  timing: LifecycleTiming | undefined;
75
87
  constructor(op: Operation);
package/dist/errors.js CHANGED
@@ -1,3 +1,4 @@
1
+ import { capacityDeadlineOf } from "./progress.js";
1
2
  export function isErrorBody(v) {
2
3
  if (typeof v !== 'object' || v === null || !('error' in v))
3
4
  return false;
@@ -45,23 +46,32 @@ export class ShardfluxProtocolError extends Error {
45
46
  /**
46
47
  * Waiting for an operation ran out of time. The operation keeps running server side:
47
48
  * resume with `cloud.workspaces.waitForOperation(err.operationId)` or call open() again
48
- * (it returns the same operation while it is active).
49
+ * (it returns the same operation while it is active). A start still waiting for capacity
50
+ * (`capacity_pending`) gives up at `deadlineAt` and then fails with `capacity_unavailable`.
49
51
  */
50
52
  export class OperationTimeoutError extends Error {
51
53
  operationId;
52
54
  workspaceId;
53
55
  lastState;
54
56
  lastReason;
57
+ /**
58
+ * When the operation was last `capacity_pending`: when it gives up waiting for a host (`error.details.deadline_at`,
59
+ * RFC 3339) and fails with `capacity_unavailable`. Null in other states or from an API that does not report it.
60
+ */
61
+ deadlineAt;
55
62
  waitedMs;
56
63
  /** Where the time went: client phases, retries and the operation's own server timing so far. */
57
64
  timing = undefined;
58
65
  constructor(op, waitedMs) {
59
- super(`Operation ${op.id} (${op.kind}) is still ${op.state}${op.state_reason ? ` (${op.state_reason})` : ''} after ${Math.round(waitedMs)} ms; it continues server side.`);
66
+ const deadlineAt = capacityDeadlineOf(op);
67
+ const after = deadlineAt ? `it keeps waiting for a host server side until ${deadlineAt} and fails with capacity_unavailable if none admits it by then.` : 'it continues server side.';
68
+ super(`Operation ${op.id} (${op.kind}) is still ${op.state}${op.state_reason ? ` (${op.state_reason})` : ''} after ${Math.round(waitedMs)} ms; ${after}`);
60
69
  this.name = 'OperationTimeoutError';
61
70
  this.operationId = op.id;
62
71
  this.workspaceId = op.workspace_id;
63
72
  this.lastState = op.state;
64
73
  this.lastReason = op.state_reason ?? null;
74
+ this.deadlineAt = deadlineAt;
65
75
  this.waitedMs = waitedMs;
66
76
  }
67
77
  }
@@ -70,6 +80,12 @@ export class OperationFailedError extends Error {
70
80
  operation;
71
81
  operationId;
72
82
  errorCode;
83
+ /**
84
+ * `operation.error.retryable` (false when absent): the same request may succeed later. `capacity_unavailable` is
85
+ * retryable: no host could admit the start before its deadline, nothing was started, retry later. The SDK never
86
+ * retries a failed operation itself.
87
+ */
88
+ retryable;
73
89
  /** Where the time went before the operation failed. */
74
90
  timing = undefined;
75
91
  constructor(op) {
@@ -79,5 +95,6 @@ export class OperationFailedError extends Error {
79
95
  this.operation = op;
80
96
  this.operationId = op.id;
81
97
  this.errorCode = code;
98
+ this.retryable = op.error?.retryable === true;
82
99
  }
83
100
  }