@shardflux/sdk 0.13.0 → 0.13.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -3,7 +3,86 @@
3
3
  Every API the README shows is available from the version named here. Breaking changes ship in minor releases and are
4
4
  marked **Breaking**.
5
5
 
6
- ## 0.13.0 (not yet published)
6
+ ## 0.14.0 (not yet published)
7
+
8
+ ### Workspaces working at once
9
+
10
+ Additive. The plan's workspace number caps workspaces working at the same moment; stored workspaces are bounded by
11
+ Retained state.
12
+
13
+ - A tool call that finds every workspace the plan allows to work at once busy is refused with 429 `quota_exceeded`
14
+ (retryable, `Retry-After`; `details.limit` `concurrent_workspaces`, `limit_value`, `current`,
15
+ `retry_after_seconds`). The call did not run, so the SDK sends the same request again after `Retry-After` (at most
16
+ 5 s) within `maxRetries`, for every call: exec starts, stdin, PTY create, keepalive and process signals included.
17
+ A refusal that outlasts the retries is a `ShardfluxApiError` with `retryAfterSeconds`. `executions.run()` keeps
18
+ these retries and does not add its own on top; the wake hint is not retried (it never waits).
19
+ - `pty.read()` rejects with the refusal (`ShardfluxApiError`, e.g. 429 `quota_exceeded`, retryable) when an accepted
20
+ attach is closed with an error body and a 4xxx close code, instead of returning empty output.
21
+ - The API's 403 `quota_exceeded` is not retried, as before. New `details.limit`: `retained_state` (`limit_value` and
22
+ `current` in GiB): opening a new key and forking are refused while the plan's Retained state is used up; existing
23
+ workspaces keep running, waking, suspending and resuming. Usage summaries report that allowance with `enforcement`
24
+ `storage_block` and `cap_state` `storage_blocked`.
25
+ - Regenerated contract types: `ErrorCode` includes `quota_exceeded` for the workspace gateway; allowance
26
+ `enforcement` adds `storage_block`, `cap_state` adds `storage_blocked`.
27
+
28
+ ## 0.13.1 (not yet published)
29
+
30
+ Patch, additive: the transport, a documented value, a fix for tool calls across a move and self-healing workspaces.
31
+ `@shardflux/cli` 0.8.1, `@shardflux/mcp` 0.7.1 and the `shardflux` bundle 0.10.1 follow (`^0.13.1`).
32
+
33
+ ### Workspaces recover by themselves
34
+
35
+ Additive. When the machine a workspace runs on fails, the workspace is suspended, and its next use (a tool call's wake,
36
+ `resume()`, `open()`) restores it: from its own disk, files kept, or from its newest checkpoint. Nothing changes for
37
+ code that already handles a cold resume.
38
+
39
+ - `cold_boot_reason` / `ServerTiming.coldBootReason` can be `host_lost`: the resume booted the workspace's disk
40
+ (`resumePath` `cold_boot`, `memoryRestored` false, processes restarted). New type `ColdBootReason`
41
+ (`'runtime_changed' | 'host_lost'`, open to other strings; the field still accepts any string).
42
+ - `hostLostOf(op)` and `ServerTiming.hostLost`: `result.host_lost` camelCased (`HostLost`: `detectedAt`,
43
+ `restoredFrom` `'disk' | 'checkpoint'`, `restoredCheckpointId`, `stateAsOf`). A resume from the checkpoint has its
44
+ state as of `stateAsOf`; a suspend that found the machine failed succeeds with `durable: true` and only
45
+ `detectedAt` (`restoredFrom` null).
46
+ - `formatTiming()` prints `processes restarted (host_lost)` for the disk, `restored checkpoint <id> (host_lost; state as
47
+ of <time>)` for the checkpoint, and `host_lost (detected <time>)` for the suspend.
48
+ - Documented operation error codes: `resume_required` (a fork or snapshot of such a workspace before its resume; resume
49
+ it first; not retryable) and `workspace_storage_unavailable` with `details.reason` `host_lost` (not retryable).
50
+ `KnownErrorReason` gains `host_lost`.
51
+
52
+ ### Connection reuse
53
+
54
+ Additive. The call after an agent's pause between tool calls reuses its connection instead of opening a new one.
55
+
56
+ - On Node 26 the pooled transport keeps an idle connection reusable for 5 minutes (was 4 s), so a request seconds or
57
+ minutes after the last one skips the TCP and TLS handshake. Pooled sockets send TCP keepalive probes after 60 s
58
+ idle, so NATs and firewalls along the way keep the connection open. Idle connections never keep a process alive. A
59
+ supplied `fetch`, other runtimes and `SHARDFLUX_HTTP_KEEPALIVE` work as before.
60
+ - A request the SDK retries (GET, a request with an `Idempotency-Key`, a read-only POST such as files search) whose
61
+ connection closes before the response arrives is sent again at once on a new connection. That one resend has no
62
+ backoff and does not count against `maxRetries`; timing records and progress `retry` events list it with
63
+ `delayMs: 0`. A further failure follows `maxRetries` and backoff as before, and a POST without an
64
+ `Idempotency-Key` is never sent twice.
65
+
66
+ ### Faster suspend and resume
67
+
68
+ Additive. `resumePath` / `resume_path` can be `thaw`: a resume that arrives while its suspend is still being
69
+ written continues the same VM in place. Nothing else changes for your code.
70
+
71
+ ### Tool calls across a move
72
+
73
+ Fix. A tool call made while its workspace finishes a resume or a move now runs once the workspace runs, instead of
74
+ failing with `409 conflict` (`workspace_not_running`).
75
+
76
+ - When the call's tool token was refused as not running and the wake then found the workspace already running (the
77
+ resume or move committed in between; `wake` resolves `false`), the call is retried with a current token: the held
78
+ resume's, or a new one. A refused call never ran, so the retry cannot run it twice. Before, the refusal surfaced.
79
+ - A refusal that comes back while the API keeps reporting the workspace running is retried after 500 ms, within the
80
+ same 3 wakes and `transitionTimeoutMs`, then surfaces as before.
81
+ - `onProgress` sees the retry as a `tool` event of type `retry`, cause `conflict workspace_not_running (the workspace
82
+ runs; new token)`.
83
+ - The CLI, the MCP server and the `shardflux` bundle get it through their `^0.13.0` dependency.
84
+
85
+ ## 0.13.0
7
86
 
8
87
  ### Elastic memory
9
88
 
@@ -446,7 +525,6 @@ behavior is the automatic version check (below), which makes one background requ
446
525
  - New exports: `FeedbackCategory`, `FeedbackContext`, `FeedbackReceipt`, `SendFeedbackParams`, `AccountFeedbackParams`,
447
526
  `FEEDBACK_CATEGORIES`, `FEEDBACK_MESSAGE_MAX_LENGTH`.
448
527
 
449
-
450
528
  ## 0.8.0 (2026-09-28)
451
529
 
452
530
  Types only; nothing changes at run time and the API is unchanged.
package/README.md CHANGED
@@ -296,10 +296,10 @@ await cell.exec.run(['make', 'test']); // resumes the wo
296
296
  ```
297
297
 
298
298
  **Detecting a cold resume (0.11.0+).** A resume restores memory and running processes from the checkpoint. After a
299
- platform runtime update, a resume can boot from the saved disk instead of restoring memory (a cold boot), also when a
300
- tool call wakes the workspace; check `memoryRestored`. The workspace runs with its files as of the suspend, and its
301
- processes start fresh, as after a reboot: start dev servers, databases and background jobs again. The resume's timing
302
- says which it was:
299
+ platform runtime update, or after the machine the workspace ran on failed (0.13.1+), a resume can boot from the saved
300
+ disk instead of restoring memory (a cold boot), also when a tool call wakes the workspace; check `memoryRestored`. The
301
+ workspace runs with its files kept, and its processes start fresh, as after a reboot: start dev servers, databases and
302
+ background jobs again. The resume's timing says which it was:
303
303
 
304
304
  ```ts
305
305
  const t = workspace.lastTiming; // after resume({ wait: true }), wake(), or a tool call that woke the workspace
@@ -310,11 +310,20 @@ if (t?.server?.memoryRestored === false) {
310
310
 
311
311
  - `ServerTiming.memoryRestored`: `true` when memory and processes came back, `false` when the workspace booted
312
312
  (`resumePath` `cold_boot`, or `reset_blank_layer` after a reset), `null` when the API did not say (older APIs).
313
- - `ServerTiming.coldBootReason`: why it booted (`runtime_changed`), else null. `formatTiming()` prints
314
- `resume from cold_boot: processes restarted (runtime_changed)`.
313
+ - `ServerTiming.coldBootReason` (type `ColdBootReason`): why it booted, `runtime_changed` or `host_lost`
314
+ **(0.13.1+)**, else null. `formatTiming()` prints `resume from cold_boot: processes restarted (runtime_changed)`.
315
315
  - The finished resume operation carries the same in `result` (`memory_restored`, `cold_boot_reason`, and
316
316
  `cold_boot` with details such as `files_as_of`; it is informational and not part of the stable contract).
317
317
 
318
+ **Workspaces recover by themselves (0.13.1+).** When the machine a workspace runs on fails, the workspace is suspended,
319
+ and its next use (a tool call, `resume()`, `open()`) restores it; there is nothing extra to call.
320
+ `hostLostOf(op)` and `lastTiming.server.hostLost` (`HostLost`) say how:
321
+
322
+ - `restoredFrom: 'disk'`: it booted from its own disk, files kept (`coldBootReason` `host_lost`, processes restarted).
323
+ - `restoredFrom: 'checkpoint'`: it resumed its newest checkpoint (`restoredCheckpointId`); its state is as of
324
+ `stateAsOf`. `formatTiming()` adds `restored checkpoint <id> (host_lost; state as of <time>)`.
325
+ - A `suspend()` that finds the machine failed succeeds (`result.durable` true, `hostLost.detectedAt`).
326
+
318
327
  `suspend`, `resume`, `fork`, `snapshot` and `delete` return the lifecycle operation.
319
328
  `cloud.workspaces.waitForOperation(id)` waits for it (default timeout 5 minutes): each poll asks the API to hold
320
329
  the response until the operation changes (`Prefer: wait`, at most 20 s per request), so completion arrives within
@@ -324,7 +333,15 @@ operation keeps running server side (a queued start until its `deadlineAt`); wai
324
333
  same call. `templates.builds.waitForBuild()` waits the same way. A waited `open()` issues the first tool token
325
334
  together with the final workspace read, so the first tool call starts at once.
326
335
 
327
- From 0.12.0, Node 26 reuses TLS connections through a private HTTP/1.1 pool. Other runtimes keep native fetch. A supplied `fetch` stays unchanged. `SHARDFLUX_HTTP_KEEPALIVE=0` forces connection closure for diagnosis; `=1` opts into native pooling. No global dispatcher is installed.
336
+ From 0.12.0, Node 26 reuses TLS connections through a private HTTP/1.1 pool. From 0.13.0, an idle connection stays
337
+ reusable for 5 minutes, so the call after an agent's pause between tool calls skips the TCP and TLS handshake, and
338
+ pooled sockets send TCP keepalive probes after 60 s idle. Idle connections never keep a process alive. Other runtimes
339
+ keep native fetch. A supplied `fetch` stays unchanged. `SHARDFLUX_HTTP_KEEPALIVE=0` forces connection closure for
340
+ diagnosis; `=1` opts into native pooling. No global dispatcher is installed.
341
+
342
+ A request the SDK retries (safe methods, requests with an `Idempotency-Key`, read-only POSTs) whose connection closes
343
+ before the response arrives is sent again at once on a new connection **(0.13.0+)**, without backoff and outside
344
+ `maxRetries`; timing records list it with `delayMs: 0`.
328
345
 
329
346
  List and look up workspaces:
330
347
 
@@ -412,7 +429,7 @@ This resume brought the workspace back from its host's local cache, memory and r
412
429
  running, including any time in `capacity_pending`; `ran` is the cell's work (placement, boot or restore, guest
413
430
  readiness). `start` / `resume from` and `boot to ready` / `host …` are what the cell reported. A resume that booted
414
431
  the saved disk instead of restoring memory reads `resume from cold_boot: processes restarted (runtime_changed)`
415
- (0.11.0+).
432
+ (0.11.0+), or `(host_lost)` (0.13.1+) after the machine the workspace ran on failed.
416
433
  - **outside the server** is your total minus the operation's: network, TLS, polling latency, view and token. A large
417
434
  value with a small server total points at the connection between you and the API, not at the workspace.
418
435
  - **retries** lists transient failures the SDK retried (cause and backoff).
package/dist/cell.d.ts CHANGED
@@ -174,8 +174,9 @@ export interface CellClientOptions {
174
174
  sleep?: (ms: number) => Promise<void>;
175
175
  /**
176
176
  * Wakes a suspended workspace (resume, or join the active resume/open) within `timeoutMs` and resolves once it runs;
177
- * resolves `false` when there was nothing to wake (the API already reports it running), and the refusal then
178
- * surfaces. It throws when the wake fails (OperationFailedError) or outlasts `timeoutMs` (OperationTimeoutError).
177
+ * resolves `false` when there was nothing to wake (the API already reports it running): the refused call is then
178
+ * retried with a current token, since the transition that refused it ended meanwhile (0.13.1+). It throws when the
179
+ * wake fails (OperationFailedError) or outlasts `timeoutMs` (OperationTimeoutError).
179
180
  * Workspace.cell() supplies `workspace.wake()`. A call refused with `workspace_not_running` wakes the workspace and is
180
181
  * retried; a refused call was never executed, so the retry cannot duplicate it. `null` surfaces
181
182
  * the refusal instead. A wake that put a new token into this client's manager (the held resume;
@@ -320,8 +321,10 @@ export declare class CellClient {
320
321
  * One authorized request. Refreshes the token once on stale_epoch / 401. Lifecycle transitions,
321
322
  * bounded in total by `transitionTimeoutMs`: a call refused with `workspace_busy` is retried after `Retry-After`; one
322
323
  * refused with `workspace_not_running` (or whose token cannot be minted because the workspace is not running) wakes
323
- * the workspace through `wake`, given the time left, and is retried with a fresh token, at most 3 wakes per call.
324
- * Refused calls were never executed, so retrying is safe. When the budget is spent the refusal surfaces.
324
+ * the workspace through `wake`, given the time left, and is retried with a fresh token, at most 3 wakes per call. A
325
+ * wake that finds the workspace already running (`false`) is retried the same way: the transition that refused the
326
+ * call ended meanwhile (0.13.1+). Refused calls were never executed, so retrying is safe. When the budget is spent
327
+ * the refusal surfaces.
325
328
  */
326
329
  request(method: string, path: string, init?: RequestOptions & TransitionOptions): Promise<Response>;
327
330
  /**
package/dist/cell.js CHANGED
@@ -1,4 +1,4 @@
1
- import { ExecStartError, NotSupportedForModeError, ShardfluxApiError, ShardfluxProtocolError } from "./errors.js";
1
+ import { ExecStartError, NotSupportedForModeError, ShardfluxApiError, ShardfluxProtocolError, apiError, isErrorBody, isWorkingQuotaRefusal } from "./errors.js";
2
2
  import { EXECUTION_ID, ExecutionResult, newExecutionId } from "./executions.js";
3
3
  import { HttpClient, defaultSleep, randomId, treeRevisionOf } from "./http.js";
4
4
  import { describeFailure, emitTo } from "./progress.js";
@@ -20,8 +20,13 @@ export const DEFAULT_TRANSITION_TIMEOUT_MS = 120_000;
20
20
  * capture writes recorded before them (read-your-writes). Workspace.cell() sets it; the capture's own client does not.
21
21
  */
22
22
  export const CAPTURE_BARRIER = Symbol('shardflux.captureBarrier');
23
- /** Wakes per call at most: a workspace that keeps being suspended again surfaces the refusal. */
23
+ /**
24
+ * Wakes per call at most: a workspace that keeps being suspended again, or keeps refusing the call while the API reports
25
+ * it running, surfaces the refusal.
26
+ */
24
27
  const MAX_WAKES = 3;
28
+ /** The pause before retrying a call refused again while the API reports the workspace running (from the second time). */
29
+ const RUNNING_AGAIN_PAUSE_MS = 500;
25
30
  /** exec.cancel grace: the workspace waits 5 s without one and takes at most 60 s. */
26
31
  const DEFAULT_CANCEL_GRACE_MS = 5_000;
27
32
  const MAX_CANCEL_GRACE_MS = 60_000;
@@ -63,6 +68,30 @@ export async function* ndjson(body) {
63
68
  reader.releaseLock();
64
69
  }
65
70
  }
71
+ /**
72
+ * A refused attach the gateway had to accept (cell-api.yaml "WebSocket close codes": an allowed Origin cannot read the
73
+ * HTTP answer to a failed upgrade): its only text frame is exactly the ErrorBody, then a close with 4000 + the HTTP
74
+ * status (4410 for not_found). A 4429 without a body (its close reason is `{"code","retryable",...}`) is the same
75
+ * refusal in short. Undefined for any other close.
76
+ */
77
+ function attachRefusal(code, reason, body) {
78
+ const status = code === 4410 ? 404 : code !== undefined && code >= 4000 && code < 5000 ? code - 4000 : undefined;
79
+ if (body) {
80
+ const hinted = body.error.details?.retry_after_seconds;
81
+ return apiError(status ?? 502, body, 'cell', typeof hinted === 'number' && hinted >= 0 ? hinted : undefined);
82
+ }
83
+ if (code !== 4429)
84
+ return undefined;
85
+ let short = {};
86
+ try {
87
+ short = JSON.parse(reason ?? '');
88
+ }
89
+ catch {
90
+ // Not the compact JSON: the close code alone says it.
91
+ }
92
+ const errCode = typeof short.code === 'string' ? short.code : 'rate_limited';
93
+ return apiError(429, { error: { code: errCode, message: `The attach was refused (${errCode}); retry shortly.`, request_id: typeof short.request_id === 'string' ? short.request_id : '', retryable: short.retryable !== false } }, 'cell');
94
+ }
66
95
  class ByteSink {
67
96
  #max;
68
97
  #chunks = [];
@@ -108,6 +137,10 @@ const executionRetryDelayMs = (failures, retryAfterSeconds) => retryAfterSeconds
108
137
  const attemptTimedOut = (err) => err instanceof Error && err.name === 'TimeoutError';
109
138
  /** A transport failure or a retryable 429/5xx: the execution request may be sent again with the same id. */
110
139
  function transientExecutionFailure(err) {
140
+ // The working-at-once refusal was already retried by the HTTP layer within the client's maxRetries: a second loop
141
+ // here would multiply that budget, so it surfaces.
142
+ if (isWorkingQuotaRefusal(err))
143
+ return null;
111
144
  if (err instanceof ShardfluxApiError) {
112
145
  const status = err.status === 429 || err.status === 502 || err.status === 503 || err.status === 504;
113
146
  return status && err.retryable ? { retryAfterSeconds: err.retryAfterSeconds } : null;
@@ -213,8 +246,10 @@ export class CellClient {
213
246
  * One authorized request. Refreshes the token once on stale_epoch / 401. Lifecycle transitions,
214
247
  * bounded in total by `transitionTimeoutMs`: a call refused with `workspace_busy` is retried after `Retry-After`; one
215
248
  * refused with `workspace_not_running` (or whose token cannot be minted because the workspace is not running) wakes
216
- * the workspace through `wake`, given the time left, and is retried with a fresh token, at most 3 wakes per call.
217
- * Refused calls were never executed, so retrying is safe. When the budget is spent the refusal surfaces.
249
+ * the workspace through `wake`, given the time left, and is retried with a fresh token, at most 3 wakes per call. A
250
+ * wake that finds the workspace already running (`false`) is retried the same way: the transition that refused the
251
+ * call ended meanwhile (0.13.1+). Refused calls were never executed, so retrying is safe. When the budget is spent
252
+ * the refusal surfaces.
218
253
  */
219
254
  async request(method, path, init = {}) {
220
255
  const closer = this.#closer.signal;
@@ -237,6 +272,7 @@ export class CellClient {
237
272
  const retry = (cause, delayMs, attempt) => emit?.({ type: 'retry', retry: { atMs: Math.round((performance.now() - t0) * 10) / 10, request: `${method} ${path}`, attempt, cause, delayMs } });
238
273
  let refreshed = false;
239
274
  let wakes = 0;
275
+ let runningAgain = 0;
240
276
  let attempts = 0;
241
277
  for (;;) {
242
278
  try {
@@ -270,8 +306,17 @@ export class CellClient {
270
306
  if (notRunning && wakeAllowed && this.#wake && wakes < MAX_WAKES && left > 0) {
271
307
  wakes += 1;
272
308
  const before = this.tokens.current;
273
- if ((await this.#wake(left, signal)) === false)
274
- throw err; // running per the API: nothing to wait for
309
+ if ((await this.#wake(left, signal)) === false) {
310
+ // Running per the API: the refusal was answered from a transition that ended before the wake looked (a
311
+ // resume or a move committed in between: the token request was refused while the workspace was resuming,
312
+ // and the wake found it running). The refused call never ran, so it is retried with a current token. A
313
+ // refusal that comes back while the API keeps reporting the workspace running is retried after a pause,
314
+ // within the same wakes and time budget, then surfaces.
315
+ const pause = runningAgain++ === 0 ? 0 : Math.max(0, Math.min(deadline - Date.now(), RUNNING_AGAIN_PAUSE_MS));
316
+ retry(`${err.code === 'conflict' ? 'conflict workspace_not_running' : err.code} (the workspace runs; new token)`, pause, (attempts += 1));
317
+ if (pause > 0)
318
+ await this.#opts.sleep(pause);
319
+ }
275
320
  // A held resume handed this client a token of the woken workspace: use it. Otherwise the
276
321
  // old token is of the previous epoch: fetch a new one.
277
322
  if (this.tokens.current === before)
@@ -744,6 +789,7 @@ export class CellClient {
744
789
  let next = opts.offset ?? 0;
745
790
  let session = null;
746
791
  let exited = false;
792
+ let refusal;
747
793
  await new Promise((resolve, reject) => {
748
794
  const onClose = () => finish(this.#closer.signal.reason instanceof Error ? this.#closer.signal.reason : new Error('closed'));
749
795
  const ws = new WS(url, { headers: { authorization: `Bearer ${token}` } });
@@ -760,8 +806,10 @@ export class CellClient {
760
806
  catch {
761
807
  // already closed
762
808
  }
763
- if (err)
764
- reject(err);
809
+ // A refusal body whose close did not arrive within the read: the body alone says it.
810
+ const failure = err ?? (refusal ? attachRefusal(undefined, undefined, refusal) : undefined);
811
+ if (failure)
812
+ reject(failure);
765
813
  else
766
814
  resolve();
767
815
  };
@@ -775,6 +823,12 @@ export class CellClient {
775
823
  if (typeof e.data !== 'string')
776
824
  return;
777
825
  const m = JSON.parse(e.data);
826
+ // A refused attach (e.g. 4429 quota_exceeded: every working slot is taken): the bare ErrorBody, then the close.
827
+ if (!('type' in m)) {
828
+ if (isErrorBody(m))
829
+ refusal = m;
830
+ return;
831
+ }
778
832
  if (m.type === 'output' && m.data) {
779
833
  const bytes = unb64(m.data);
780
834
  sink.push(bytes);
@@ -791,7 +845,7 @@ export class CellClient {
791
845
  }
792
846
  });
793
847
  ws.addEventListener('error', () => finish(new ShardfluxProtocolError('pty attach WebSocket failed', 0, 'cell')));
794
- ws.addEventListener('close', () => finish());
848
+ ws.addEventListener('close', (e) => finish(attachRefusal(e.code, e.reason, refusal)));
795
849
  if (this.closed)
796
850
  onClose();
797
851
  else
@@ -955,7 +1009,8 @@ export class CellClient {
955
1009
  // Nothing of a file-first workspace sleeps: the cell would answer `resident`.
956
1010
  if (this.mode === 'file_first')
957
1011
  return Promise.resolve({ residency: 'resident' });
958
- return this.#json('POST', this.#p('/v1/workspaces/{workspace_id}/wake-hint'), { wake: false, busy: false, timeoutMs: 10_000, ...(signal ? { signal } : {}) });
1012
+ // A hint never waits: a 429 quota_exceeded (every working slot is taken) surfaces at once, like its other refusals.
1013
+ return this.#json('POST', this.#p('/v1/workspaces/{workspace_id}/wake-hint'), { wake: false, busy: false, quotaRetry: false, timeoutMs: 10_000, ...(signal ? { signal } : {}) });
959
1014
  }
960
1015
  /** Read idle signals without recording activity or waking the workspace. */
961
1016
  idle(signal) {
package/dist/errors.d.ts CHANGED
@@ -70,8 +70,19 @@ export type ErrorCode = AppErrorCode | CellErrorCode;
70
70
  * 409 conflict read_only_path (a files write under an immutable path, from the cell gateway). The reason
71
71
  * `update_policy_not_available` is gone with the update policy. A build that fails on them carries `failure.code`
72
72
  * immutable_path_missing (details.path) or immutable_image_too_large.
73
+ * Host loss (0.13.1; the machine a workspace ran on failed, and the workspace restores itself on its next use): the
74
+ * operation error `workspace_storage_unavailable` carries details.reason host_lost (and details.workspace_state
75
+ * `suspended`) when neither the workspace's disk nor a checkpoint could be restored (not retryable). A fork or snapshot
76
+ * of such a workspace before its resume fails with the operation error `resume_required` (details.reason host_lost;
77
+ * not retryable): resume the workspace first.
78
+ * Workspaces working at once (0.14.0): the cell gateway answers a tool call that finds every working slot of the plan
79
+ * taken with 429 `quota_exceeded` (retryable, Retry-After; details.limit `concurrent_workspaces`, limit_value, current,
80
+ * retry_after_seconds); the call did not run and the SDK retries it like any transient refusal. The API's 403
81
+ * `quota_exceeded` (not retryable) names details.limit `concurrent_workspaces` on open, resume and fork, and
82
+ * `retained_state` (details.limit_value and details.current in GiB) when opening a new key or forking with the plan's
83
+ * Retained state used up. `details.limit` is not a reason: these are not in this union.
73
84
  */
74
- export type KnownErrorReason = 'invalid_recipe' | 'base_not_layered' | 'language_unavailable' | 'language_conflict' | 'invalid_package' | 'too_many_files' | 'platform_owned_path' | 'upload_required' | 'upload_missing' | 'upload_digest_mismatch' | 'upload_too_large' | 'extra_hosts_without_auto' | 'invalid_settings' | 'services_unsupported' | 'input_required' | 'input_unknown' | 'input_invalid' | 'egress_widening' | 'reserved_session_id' | 'env_collision' | 'reserved_template_slug' | 'package_index_unavailable' | 'package_not_found' | 'startup_failed' | 'service_not_ready' | 'secrets_unavailable' | 'workspace_not_running' | 'operation_in_progress' | 'workspace_deleted' | 'secret_not_available' | 'legacy_disk_layout' | 'not_session' | 'session_lifetime' | 'lifetime_mismatch' | 'not_resettable' | 'template_not_layered' | 'draft_exists' | 'draft_stale' | 'build_in_progress' | 'file_list_unavailable' | 'file_list_indexing' | 'guest_feature_unavailable' | 'confirm_destructive_required' | 'reserved_key_prefix' | 'invalid_defaults' | 'invalid_path' | 'too_many_acknowledged_findings' | 'template_dev_mode_role' | 'draft_not_found' | 'version_not_found' | 'path_not_found' | 'revision_mismatch' | 'edit_not_found' | 'edit_ambiguous' | 'edit_not_text' | 'patch_invalid' | 'host_capacity' | 'wake_failed' | 'workspace_fenced' | 'offline_unavailable' | 'offline_budget' | 'offline_changed' | 'host_feature_unavailable' | 'not_supported_for_mode' | 'mode_mismatch' | 'mode_not_available' | 'layout_unsupported' | 'tree_revision_mismatch' | 'outside_tree_root' | 'execution_in_progress' | 'execution_id_reused' | 'operation_id_reused' | 'no_execution_host' | 'lease_expired' | 'host_unreachable' | 'host_restarted' | 'tree_moved' | 'blob_missing' | 'blob_corrupt' | 'exec_failed_to_start' | 'invalid_cwd' | 'allowance_used' | 'overage_paused' | 'spend_cap_reached' | 'overage_unavailable' | 'spend_cap_required' | 'spend_cap_below_minimum' | 'spend_cap_above_plan_price' | 'spend_cap_below_charges' | 'version_mismatch' | 'allocation_mode_not_available' | 'requires_elastic' | 'exceeds_memory_mib' | 'burst_mode_not_supported' | 'burst_not_supported' | 'burst_size_exceeds_plan' | 'not_available' | 'shared_volumes' | 'fence_not_drained' | 'apply_pending' | 'park_failed' | 'workspace_resumed' | 'interrupted' | 'burst_lost' | 'disk_full' | 'apply_failed' | 'reverted' | 'revert_failed' | 'immutable_path_removed' | 'immutable_paths_unsupported_base' | 'read_only_path';
85
+ export type KnownErrorReason = 'invalid_recipe' | 'base_not_layered' | 'language_unavailable' | 'language_conflict' | 'invalid_package' | 'too_many_files' | 'platform_owned_path' | 'upload_required' | 'upload_missing' | 'upload_digest_mismatch' | 'upload_too_large' | 'extra_hosts_without_auto' | 'invalid_settings' | 'services_unsupported' | 'input_required' | 'input_unknown' | 'input_invalid' | 'egress_widening' | 'reserved_session_id' | 'env_collision' | 'reserved_template_slug' | 'package_index_unavailable' | 'package_not_found' | 'startup_failed' | 'service_not_ready' | 'secrets_unavailable' | 'workspace_not_running' | 'operation_in_progress' | 'workspace_deleted' | 'secret_not_available' | 'legacy_disk_layout' | 'not_session' | 'session_lifetime' | 'lifetime_mismatch' | 'not_resettable' | 'template_not_layered' | 'draft_exists' | 'draft_stale' | 'build_in_progress' | 'file_list_unavailable' | 'file_list_indexing' | 'guest_feature_unavailable' | 'confirm_destructive_required' | 'reserved_key_prefix' | 'invalid_defaults' | 'invalid_path' | 'too_many_acknowledged_findings' | 'template_dev_mode_role' | 'draft_not_found' | 'version_not_found' | 'path_not_found' | 'revision_mismatch' | 'edit_not_found' | 'edit_ambiguous' | 'edit_not_text' | 'patch_invalid' | 'host_capacity' | 'wake_failed' | 'workspace_fenced' | 'offline_unavailable' | 'offline_budget' | 'offline_changed' | 'host_feature_unavailable' | 'not_supported_for_mode' | 'mode_mismatch' | 'mode_not_available' | 'layout_unsupported' | 'tree_revision_mismatch' | 'outside_tree_root' | 'execution_in_progress' | 'execution_id_reused' | 'operation_id_reused' | 'no_execution_host' | 'lease_expired' | 'host_unreachable' | 'host_restarted' | 'tree_moved' | 'blob_missing' | 'blob_corrupt' | 'exec_failed_to_start' | 'invalid_cwd' | 'allowance_used' | 'overage_paused' | 'spend_cap_reached' | 'overage_unavailable' | 'spend_cap_required' | 'spend_cap_below_minimum' | 'spend_cap_above_plan_price' | 'spend_cap_below_charges' | 'version_mismatch' | 'allocation_mode_not_available' | 'requires_elastic' | 'exceeds_memory_mib' | 'burst_mode_not_supported' | 'burst_not_supported' | 'burst_size_exceeds_plan' | 'not_available' | 'shared_volumes' | 'fence_not_drained' | 'apply_pending' | 'park_failed' | 'workspace_resumed' | 'interrupted' | 'burst_lost' | 'disk_full' | 'apply_failed' | 'reverted' | 'revert_failed' | 'immutable_path_removed' | 'immutable_paths_unsupported_base' | 'read_only_path' | 'host_lost';
75
86
  /** A known reason, or any other string the server sends (reasons are open-ended). */
76
87
  export type ErrorReason = KnownErrorReason | (string & {});
77
88
  export interface ErrorBodyLike {
@@ -106,6 +117,13 @@ export declare class ShardfluxApiError extends Error {
106
117
  readonly treeRevision: number | undefined;
107
118
  constructor(status: number, body: ErrorBodyLike, source: 'api' | 'cell', retryAfterSeconds?: number, treeRevision?: number);
108
119
  }
120
+ /**
121
+ * The cell gateway's working-at-once refusal (0.14.0+): 429 `quota_exceeded`, retryable, with Retry-After. The gateway
122
+ * refuses the call at admission, before the workspace is touched, so nothing ran and any request (exec start, stdin,
123
+ * PTY create, keepalive included) may be sent again. The API's 403 `quota_exceeded` is a different refusal (not
124
+ * retryable) and never matches.
125
+ */
126
+ export declare function isWorkingQuotaRefusal(err: unknown): err is ShardfluxApiError;
109
127
  /** processful (the default: one VM keeps processes, memory and files) or file_first. */
110
128
  export type WorkspaceMode = AppComponents['schemas']['WorkspaceMode'];
111
129
  /**
@@ -197,6 +215,7 @@ export declare class OperationTimeoutError extends Error {
197
215
  * The operation reached `failed` or `canceled`. A suspend-when-idle that found the workspace active (a tool call after
198
216
  * the request, an attached stream) is `canceled` with `errorCode` `workspace_active` (`workspaceActive` true, 0.12.0+):
199
217
  * nothing changed and the workspace keeps running. `waitUntilReady()` and `wake()` treat it as running, not as a failure.
218
+ * `resume_required` (0.13.1+): a fork or snapshot of a workspace whose machine failed; resume it, then call again.
200
219
  */
201
220
  export declare class OperationFailedError extends Error {
202
221
  readonly operation: Operation;
package/dist/errors.js CHANGED
@@ -40,6 +40,15 @@ export class ShardfluxApiError extends Error {
40
40
  this.reason = typeof reason === 'string' ? reason : undefined;
41
41
  }
42
42
  }
43
+ /**
44
+ * The cell gateway's working-at-once refusal (0.14.0+): 429 `quota_exceeded`, retryable, with Retry-After. The gateway
45
+ * refuses the call at admission, before the workspace is touched, so nothing ran and any request (exec start, stdin,
46
+ * PTY create, keepalive included) may be sent again. The API's 403 `quota_exceeded` is a different refusal (not
47
+ * retryable) and never matches.
48
+ */
49
+ export function isWorkingQuotaRefusal(err) {
50
+ return err instanceof ShardfluxApiError && err.source === 'cell' && err.status === 429 && err.code === 'quota_exceeded' && err.retryable;
51
+ }
43
52
  /**
44
53
  * The call does not exist for the workspace's mode: 409 `conflict` with details.reason
45
54
  * `not_supported_for_mode`, `details.mode` (the workspace's mode) and `details.operation`. File-first workspaces have no
@@ -187,6 +196,7 @@ export class OperationTimeoutError extends Error {
187
196
  * The operation reached `failed` or `canceled`. A suspend-when-idle that found the workspace active (a tool call after
188
197
  * the request, an attached stream) is `canceled` with `errorCode` `workspace_active` (`workspaceActive` true, 0.12.0+):
189
198
  * nothing changed and the workspace keeps running. `waitUntilReady()` and `wake()` treat it as running, not as a failure.
199
+ * `resume_required` (0.13.1+): a fork or snapshot of a workspace whose machine failed; resume it, then call again.
190
200
  */
191
201
  export class OperationFailedError extends Error {
192
202
  operation;
@@ -5630,6 +5630,21 @@ export interface components {
5630
5630
  } & {
5631
5631
  [key: string]: unknown;
5632
5632
  };
5633
+ host_lost?: {
5634
+ /** @description RFC 3339: when the host failure was detected; the workspace was suspended then. */
5635
+ detected_at: string;
5636
+ /**
5637
+ * @description resume and open: disk when the workspace booted from its disk (files kept, processes and memory not); checkpoint when it was restored from its newest checkpoint (see state_as_of). Absent on a suspend.
5638
+ * @enum {string}
5639
+ */
5640
+ restored_from?: "disk" | "checkpoint";
5641
+ /** @description restored_from checkpoint: the checkpoint this resume restored. */
5642
+ restored_checkpoint_id?: string;
5643
+ /** @description restored_from checkpoint, RFC 3339: when that checkpoint committed; the workspace state is as of then. */
5644
+ state_as_of?: string;
5645
+ } & {
5646
+ [key: string]: unknown;
5647
+ };
5633
5648
  } & {
5634
5649
  [key: string]: unknown;
5635
5650
  };
@@ -28007,8 +28022,8 @@ export interface operations {
28007
28022
  display_name: string;
28008
28023
  unit: string;
28009
28024
  meters: ("cpu_seconds" | "memory_gib_seconds" | "storage_gib_seconds" | "egress_bytes" | "ingress_bytes" | "volume_storage_gib_seconds")[];
28010
- /** @description hard_cap: exhausting it refuses opens/resumes (402 allowance_exhausted) and running workspaces are suspended; spare_capacity: no allowance cap (runs on reclaimable spare capacity); reported: shown and notified, never enforced by refusing access; egress_block (outbound_transfer_gb, contracts §18): at 100% outbound internet traffic is blocked by an organization egress override until upgrade, purchase or the next period (workspaces keep running, starts are not refused). */
28011
- enforcement: "hard_cap" | "spare_capacity" | "reported" | "egress_block";
28025
+ /** @description hard_cap: exhausting it refuses opens/resumes (402 allowance_exhausted) and running workspaces are suspended; spare_capacity: no allowance cap (runs on reclaimable spare capacity); reported: shown and notified, never enforced by refusing access; egress_block (outbound_transfer_gb, contracts §18): at 100% outbound internet traffic is blocked by an organization egress override until upgrade, purchase or the next period (workspaces keep running, starts are not refused); storage_block (retained_state_gib, contracts §40): at 100% opening a new workspace key and forking are refused (403 quota_exceeded, details.limit retained_state) until retained state is below the allowance (existing workspaces keep running, waking, suspending and resuming; nothing is deleted or charged). */
28026
+ enforcement: "hard_cap" | "spare_capacity" | "reported" | "egress_block" | "storage_block";
28012
28027
  included: number | null;
28013
28028
  used: number;
28014
28029
  remaining: number | null;
@@ -28016,8 +28031,8 @@ export interface operations {
28016
28031
  included_meter_units: number | null;
28017
28032
  used_meter_units: number | null;
28018
28033
  remaining_meter_units: number | null;
28019
- /** @description ok: below 80 %. warning: 80 % or more of the allowance. exhausted: a hard cap is used up and starts are refused (402 allowance_exhausted; see exhausted_reason). overage: a CPU-hours or RAM GiB-hours allowance is used up while opt-in overage is on and below its spend cap, so starts are admitted and the usage past it is charged (spend_cap). over_allowance: a reported allowance is exceeded (never refused). egress_blocked: the outbound transfer allowance is used up and outbound traffic is blocked. uncapped: runs on spare capacity. not_included: the plan does not define it. */
28020
- cap_state: "ok" | "warning" | "exhausted" | "overage" | "over_allowance" | "egress_blocked" | "uncapped" | "not_included";
28034
+ /** @description ok: below 80 %. warning: 80 % or more of the allowance. exhausted: a hard cap is used up and starts are refused (402 allowance_exhausted; see exhausted_reason). overage: a CPU-hours or RAM GiB-hours allowance is used up while opt-in overage is on and below its spend cap, so starts are admitted and the usage past it is charged (spend_cap). over_allowance: a reported allowance is exceeded (never refused). egress_blocked: the outbound transfer allowance is used up and outbound traffic is blocked. storage_blocked: the retained state allowance is used up and new workspaces and forks are refused (403 quota_exceeded). uncapped: runs on spare capacity. not_included: the plan does not define it. */
28035
+ cap_state: "ok" | "warning" | "exhausted" | "overage" | "over_allowance" | "egress_blocked" | "storage_blocked" | "uncapped" | "not_included";
28021
28036
  }[];
28022
28037
  meters: {
28023
28038
  /** @enum {string} */
@@ -29062,8 +29077,8 @@ export interface operations {
29062
29077
  display_name: string;
29063
29078
  unit: string;
29064
29079
  meters: ("cpu_seconds" | "memory_gib_seconds" | "storage_gib_seconds" | "egress_bytes" | "ingress_bytes" | "volume_storage_gib_seconds")[];
29065
- /** @description hard_cap: exhausting it refuses opens/resumes (402 allowance_exhausted) and running workspaces are suspended; spare_capacity: no allowance cap (runs on reclaimable spare capacity); reported: shown and notified, never enforced by refusing access; egress_block (outbound_transfer_gb, contracts §18): at 100% outbound internet traffic is blocked by an organization egress override until upgrade, purchase or the next period (workspaces keep running, starts are not refused). */
29066
- enforcement: "hard_cap" | "spare_capacity" | "reported" | "egress_block";
29080
+ /** @description hard_cap: exhausting it refuses opens/resumes (402 allowance_exhausted) and running workspaces are suspended; spare_capacity: no allowance cap (runs on reclaimable spare capacity); reported: shown and notified, never enforced by refusing access; egress_block (outbound_transfer_gb, contracts §18): at 100% outbound internet traffic is blocked by an organization egress override until upgrade, purchase or the next period (workspaces keep running, starts are not refused); storage_block (retained_state_gib, contracts §40): at 100% opening a new workspace key and forking are refused (403 quota_exceeded, details.limit retained_state) until retained state is below the allowance (existing workspaces keep running, waking, suspending and resuming; nothing is deleted or charged). */
29081
+ enforcement: "hard_cap" | "spare_capacity" | "reported" | "egress_block" | "storage_block";
29067
29082
  included: number | null;
29068
29083
  used: number;
29069
29084
  remaining: number | null;
@@ -29071,8 +29086,8 @@ export interface operations {
29071
29086
  included_meter_units: number | null;
29072
29087
  used_meter_units: number | null;
29073
29088
  remaining_meter_units: number | null;
29074
- /** @description ok: below 80 %. warning: 80 % or more of the allowance. exhausted: a hard cap is used up and starts are refused (402 allowance_exhausted; see exhausted_reason). overage: a CPU-hours or RAM GiB-hours allowance is used up while opt-in overage is on and below its spend cap, so starts are admitted and the usage past it is charged (spend_cap). over_allowance: a reported allowance is exceeded (never refused). egress_blocked: the outbound transfer allowance is used up and outbound traffic is blocked. uncapped: runs on spare capacity. not_included: the plan does not define it. */
29075
- cap_state: "ok" | "warning" | "exhausted" | "overage" | "over_allowance" | "egress_blocked" | "uncapped" | "not_included";
29089
+ /** @description ok: below 80 %. warning: 80 % or more of the allowance. exhausted: a hard cap is used up and starts are refused (402 allowance_exhausted; see exhausted_reason). overage: a CPU-hours or RAM GiB-hours allowance is used up while opt-in overage is on and below its spend cap, so starts are admitted and the usage past it is charged (spend_cap). over_allowance: a reported allowance is exceeded (never refused). egress_blocked: the outbound transfer allowance is used up and outbound traffic is blocked. storage_blocked: the retained state allowance is used up and new workspaces and forks are refused (403 quota_exceeded). uncapped: runs on spare capacity. not_included: the plan does not define it. */
29090
+ cap_state: "ok" | "warning" | "exhausted" | "overage" | "over_allowance" | "egress_blocked" | "storage_blocked" | "uncapped" | "not_included";
29076
29091
  }[];
29077
29092
  meters: {
29078
29093
  /** @enum {string} */
@@ -191,7 +191,7 @@ export interface paths {
191
191
  * Client messages are JSON `ExecClientMessage`. Close codes: 1000 after
192
192
  * the `exit` event, 1001 on gateway shutdown (reconnect with offsets),
193
193
  * 4401 unauthenticated, 4403 forbidden, 4409 stale epoch / workspace
194
- * busy / workspace not running, 4410 workspace gone, 4429 rate limited,
194
+ * busy / workspace not running, 4410 workspace gone, 4429 rate limited / quota exceeded,
195
195
  * 4500/4503/4504 internal/unavailable/timeout (4000 + HTTP status; see
196
196
  * "WebSocket close codes" above; the close reason is compact JSON with
197
197
  * the error `code`). A refused upgrade from an allowed Origin is
@@ -265,14 +265,18 @@ export interface paths {
265
265
  get?: never;
266
266
  put?: never;
267
267
  /**
268
- * SIGTERM the process group, SIGKILL after grace
269
- * @description Sends SIGTERM to the session's process group and SIGKILL after
270
- * `grace_ms` (0 to 60000; 0 or omitted is 5000). Answers when the
271
- * command has ended: at once when SIGTERM ends it, otherwise after the
272
- * grace and the SIGKILL. A `grace_ms` outside the range is `422
273
- * validation_failed`
268
+ * Stop the command and every process it started (SIGTERM, SIGKILL after grace)
269
+ * @description Sends SIGTERM to every process the command started, background and
270
+ * setsid processes included, and SIGKILL after `grace_ms` (0 to 60000;
271
+ * 0 or omitted is 5000). Answers when they have all exited: at once when
272
+ * SIGTERM ends them, otherwise after the grace and the SIGKILL. A lost
273
+ * command is stopped the same way; a command that already exited keeps
274
+ * its result and what it left running is stopped. `503
275
+ * dependency_unavailable` (retryable) while a process cannot exit yet.
276
+ * Template versions before python-node-browser 13, ubuntu-24.04 4,
277
+ * ubuntu-22.04 2 and debian-12 2 signal the command's process group and
278
+ * return an ended session as it is. A `grace_ms` outside the range is `422 validation_failed`
274
279
  * (`details.field` `grace_ms`, `details.minimum`, `details.maximum`).
275
- * An ended session is returned as it is.
276
280
  */
277
281
  post: operations["execCancel"];
278
282
  delete?: never;
@@ -876,7 +880,7 @@ export type webhooks = Record<string, never>;
876
880
  export interface components {
877
881
  schemas: {
878
882
  /** @enum {string} */
879
- ErrorCode: "bad_request" | "unauthenticated" | "forbidden" | "not_found" | "conflict" | "stale_epoch" | "workspace_not_running" | "validation_failed" | "payload_too_large" | "range_not_satisfiable" | "rate_limited" | "capacity_pending" | "dependency_unavailable" | "timeout" | "internal_error" | "service_unavailable" | "workspace_busy" | "burst_unavailable" | "burst_apply_failed";
883
+ ErrorCode: "bad_request" | "unauthenticated" | "forbidden" | "not_found" | "conflict" | "stale_epoch" | "workspace_not_running" | "validation_failed" | "payload_too_large" | "range_not_satisfiable" | "rate_limited" | "capacity_pending" | "dependency_unavailable" | "timeout" | "internal_error" | "service_unavailable" | "workspace_busy" | "burst_unavailable" | "burst_apply_failed" | "quota_exceeded";
880
884
  ErrorBody: {
881
885
  error: {
882
886
  code: components["schemas"]["ErrorCode"];
package/dist/http.d.ts CHANGED
@@ -1,5 +1,5 @@
1
1
  import type { RetryRecord } from './progress.js';
2
- export declare const SDK_VERSION = "0.13.0";
2
+ export declare const SDK_VERSION = "0.13.1";
3
3
  export interface RequestOptions {
4
4
  query?: Record<string, string | number | boolean | undefined | null>;
5
5
  json?: unknown;
@@ -10,6 +10,11 @@ export interface RequestOptions {
10
10
  idempotencyKey?: string;
11
11
  /** The request has no effect a retry could duplicate (a read-only POST such as files/search): retried like a GET. */
12
12
  idempotent?: boolean;
13
+ /**
14
+ * false: the cell gateway's 429 `quota_exceeded` surfaces at once instead of being retried (the wake hint, which
15
+ * never waits). Default true.
16
+ */
17
+ quotaRetry?: boolean;
13
18
  signal?: AbortSignal;
14
19
  /** Override the client's default request timeout (ms); 0 disables it (streams). */
15
20
  timeoutMs?: number;
@@ -33,13 +38,45 @@ export interface HttpOptions {
33
38
  onSuccess?: (() => void) | undefined;
34
39
  }
35
40
  export declare const defaultSleep: (ms: number) => Promise<void>;
41
+ /**
42
+ * How long an idle pooled connection stays reusable (ms): 5 minutes, instead of undici's 4 s default, so the request
43
+ * after an agent's usual pause between tool calls (15 s to a few minutes) skips a new TCP and TLS handshake. Far below
44
+ * the 3600 s idle timeout of the load balancer in front of the API and the cell endpoints, so the server side never
45
+ * closes a connection for idleness while the client still treats it as reusable; below the AWS NAT gateway's 350 s idle
46
+ * timeout; and the same limit Chrome keeps for used idle sockets. Also caps a server's `Keep-Alive: timeout` hint.
47
+ */
48
+ export declare const KEEP_ALIVE_TIMEOUT_MS = 300000;
49
+ /**
50
+ * TCP keepalive probes start after a pooled socket has been idle this long (ms), so NATs and firewalls with shorter
51
+ * idle timeouts than the pool's (Azure SNAT: 4 minutes) keep the connection instead of silently dropping it. Equal to
52
+ * undici's own default, set explicitly so it is visible and tested.
53
+ */
54
+ export declare const TCP_KEEPALIVE_INITIAL_DELAY_MS = 60000;
55
+ /** The private pool's undici Agent options (exported for tests). */
56
+ export declare const POOL_OPTIONS: {
57
+ readonly allowH2: false;
58
+ readonly pipelining: 1;
59
+ readonly keepAliveTimeout: 300000;
60
+ readonly keepAliveMaxTimeout: 300000;
61
+ readonly connect: {
62
+ readonly keepAlive: true;
63
+ readonly keepAliveInitialDelay: 60000;
64
+ };
65
+ };
36
66
  /**
37
67
  * Reuses TLS connections. Node 26's bundled undici 8.9 can stall reused connections: use pinned undici with a
38
- * private HTTP/1.1-only dispatcher. Other runtimes retain native fetch. SHARDFLUX_HTTP_KEEPALIVE=0 forces close
39
- * for diagnosis; =1 retains the historical opt-in to native pooling. An explicitly supplied fetch is untouched.
68
+ * private HTTP/1.1-only dispatcher whose idle connections stay reusable for KEEP_ALIVE_TIMEOUT_MS. Other runtimes
69
+ * retain native fetch. SHARDFLUX_HTTP_KEEPALIVE=0 forces close for diagnosis; =1 retains the historical opt-in to
70
+ * native pooling. An explicitly supplied fetch is untouched. Idle pooled sockets do not keep the process alive.
40
71
  */
41
72
  export declare function defaultFetch(env?: Record<string, string | undefined> | undefined, undiciVersion?: string | undefined): typeof fetch;
42
73
  export declare function buildUrl(base: string, path: string, query?: RequestOptions['query']): string;
74
+ /**
75
+ * True when fetch failed because its connection closed or reset before a response arrived: what a request sees when
76
+ * it went out on a pooled connection the other side closed while it sat idle. Connect failures (ECONNREFUSED,
77
+ * ENOTFOUND, UND_ERR_CONNECT_TIMEOUT), timeouts and aborts are not this.
78
+ */
79
+ export declare function isStaleConnectionError(err: unknown): boolean;
43
80
  /** `X-Tree-Revision` (file-first workspaces): a non-negative integer, else null. */
44
81
  export declare function treeRevisionOf(headers: Headers): number | null;
45
82
  /** Turns a non-2xx response into a ShardfluxApiError (or a protocol error for undocumented bodies). */
package/dist/http.js CHANGED
@@ -4,18 +4,44 @@
4
4
  * retries only where a retry cannot duplicate an effect (safe methods,
5
5
  * requests carrying an Idempotency-Key, and read-only POSTs marked
6
6
  * `idempotent`). A retryable 429/502/503/504 (e.g. 503 `host_capacity`, 503
7
- * `wake_failed`) waits `Retry-After` (at most 5 s) before the retry.
7
+ * `wake_failed`) waits `Retry-After` (at most 5 s) before the retry. Such a request whose connection closed before
8
+ * any response (a pooled connection gone stale while idle) is sent again at once, once, outside `maxRetries`. The cell
9
+ * gateway's 429 `quota_exceeded` (every working slot of the plan is taken) is retried for any request (0.14.0+): it is
10
+ * refused at admission, so nothing ran.
8
11
  */
9
- import { ShardfluxApiError, ShardfluxProtocolError, apiError, isErrorBody } from "./errors.js";
12
+ import { ShardfluxApiError, ShardfluxProtocolError, apiError, isErrorBody, isWorkingQuotaRefusal } from "./errors.js";
10
13
  import { describeFailure } from "./progress.js";
11
- export const SDK_VERSION = '0.13.0';
14
+ export const SDK_VERSION = '0.13.1';
12
15
  export const defaultSleep = (ms) => new Promise((r) => setTimeout(r, ms));
16
+ /**
17
+ * How long an idle pooled connection stays reusable (ms): 5 minutes, instead of undici's 4 s default, so the request
18
+ * after an agent's usual pause between tool calls (15 s to a few minutes) skips a new TCP and TLS handshake. Far below
19
+ * the 3600 s idle timeout of the load balancer in front of the API and the cell endpoints, so the server side never
20
+ * closes a connection for idleness while the client still treats it as reusable; below the AWS NAT gateway's 350 s idle
21
+ * timeout; and the same limit Chrome keeps for used idle sockets. Also caps a server's `Keep-Alive: timeout` hint.
22
+ */
23
+ export const KEEP_ALIVE_TIMEOUT_MS = 300_000;
24
+ /**
25
+ * TCP keepalive probes start after a pooled socket has been idle this long (ms), so NATs and firewalls with shorter
26
+ * idle timeouts than the pool's (Azure SNAT: 4 minutes) keep the connection instead of silently dropping it. Equal to
27
+ * undici's own default, set explicitly so it is visible and tested.
28
+ */
29
+ export const TCP_KEEPALIVE_INITIAL_DELAY_MS = 60_000;
30
+ /** The private pool's undici Agent options (exported for tests). */
31
+ export const POOL_OPTIONS = {
32
+ allowH2: false,
33
+ pipelining: 1,
34
+ keepAliveTimeout: KEEP_ALIVE_TIMEOUT_MS,
35
+ keepAliveMaxTimeout: KEEP_ALIVE_TIMEOUT_MS,
36
+ connect: { keepAlive: true, keepAliveInitialDelay: TCP_KEEPALIVE_INITIAL_DELAY_MS },
37
+ };
13
38
  /** One private HTTP/1.1 pool on Node 26+, shared by SDK clients. No global dispatcher changes. */
14
39
  let nodeFetch;
15
40
  /**
16
41
  * Reuses TLS connections. Node 26's bundled undici 8.9 can stall reused connections: use pinned undici with a
17
- * private HTTP/1.1-only dispatcher. Other runtimes retain native fetch. SHARDFLUX_HTTP_KEEPALIVE=0 forces close
18
- * for diagnosis; =1 retains the historical opt-in to native pooling. An explicitly supplied fetch is untouched.
42
+ * private HTTP/1.1-only dispatcher whose idle connections stay reusable for KEEP_ALIVE_TIMEOUT_MS. Other runtimes
43
+ * retain native fetch. SHARDFLUX_HTTP_KEEPALIVE=0 forces close for diagnosis; =1 retains the historical opt-in to
44
+ * native pooling. An explicitly supplied fetch is untouched. Idle pooled sockets do not keep the process alive.
19
45
  */
20
46
  export function defaultFetch(env = globalThis.process?.env, undiciVersion = globalThis.process?.versions?.undici) {
21
47
  // Why a fresh connection on Node 26 (its bundled HTTP client, measured): docs/progress/startup-latency.md.
@@ -30,7 +56,7 @@ export function defaultFetch(env = globalThis.process?.env, undiciVersion = glob
30
56
  return base;
31
57
  return async (input, init) => {
32
58
  nodeFetch ??= import('undici').then(({ Agent, fetch: pooledFetch }) => {
33
- const dispatcher = new Agent({ allowH2: false, pipelining: 1 });
59
+ const dispatcher = new Agent({ ...POOL_OPTIONS, connect: { ...POOL_OPTIONS.connect } });
34
60
  return async (input, init) => {
35
61
  const response = await pooledFetch(input, { ...init, dispatcher });
36
62
  // Web-standard runtime shape; undici and DOM iterator declarations differ.
@@ -49,6 +75,25 @@ export function buildUrl(base, path, query) {
49
75
  }
50
76
  return u.toString();
51
77
  }
78
+ /** Codes (on fetch's `cause` chain) of a connection that closed or reset before the response arrived. */
79
+ const STALE_CONNECTION_CODES = new Set(['UND_ERR_SOCKET', 'ECONNRESET', 'EPIPE', 'UND_ERR_CLOSED']);
80
+ /**
81
+ * True when fetch failed because its connection closed or reset before a response arrived: what a request sees when
82
+ * it went out on a pooled connection the other side closed while it sat idle. Connect failures (ECONNREFUSED,
83
+ * ENOTFOUND, UND_ERR_CONNECT_TIMEOUT), timeouts and aborts are not this.
84
+ */
85
+ export function isStaleConnectionError(err) {
86
+ if (!(err instanceof TypeError))
87
+ return false;
88
+ const seen = new Set();
89
+ for (let c = err.cause; typeof c === 'object' && c !== null && !seen.has(c); c = c.cause) {
90
+ seen.add(c);
91
+ const code = c.code;
92
+ if (typeof code === 'string' && STALE_CONNECTION_CODES.has(code))
93
+ return true;
94
+ }
95
+ return false;
96
+ }
52
97
  function retryAfterSeconds(res) {
53
98
  const v = res.headers.get('retry-after');
54
99
  if (v === null)
@@ -89,6 +134,9 @@ export class HttpClient {
89
134
  const safe = method === 'GET' || method === 'HEAD';
90
135
  const retriable = safe || init.idempotencyKey !== undefined || init.idempotent === true;
91
136
  const sleep = this.opts.sleep ?? defaultSleep;
137
+ // Retries counted against maxRetries (they also set the backoff); `attempt` counts every send.
138
+ let counted = 0;
139
+ let staleReplayed = false;
92
140
  for (let attempt = 0;; attempt += 1) {
93
141
  const headers = {
94
142
  accept: init.accept ?? 'application/json',
@@ -122,9 +170,20 @@ export class HttpClient {
122
170
  catch (err) {
123
171
  if (init.signal?.aborted)
124
172
  throw err;
125
- if (!retriable || attempt >= this.opts.maxRetries)
173
+ if (!retriable)
174
+ throw err;
175
+ if (!staleReplayed && !signal?.aborted && isStaleConnectionError(err)) {
176
+ // The connection closed before any response, typically a pooled one the server or a NAT closed while idle:
177
+ // send again at once on a fresh connection, once per request and outside maxRetries (as Go's net/http
178
+ // replays a request that failed on a reused connection when it is idempotent or carries an Idempotency-Key).
179
+ staleReplayed = true;
180
+ init.onRetry?.({ request: `${method} ${path}`, attempt: attempt + 1, cause: describeFailure(err), delayMs: 0 });
181
+ continue;
182
+ }
183
+ if (counted >= this.opts.maxRetries)
126
184
  throw err;
127
- const delayMs = Math.min(2_000, 200 * 2 ** attempt);
185
+ const delayMs = Math.min(2_000, 200 * 2 ** counted);
186
+ counted += 1;
128
187
  init.onRetry?.({ request: `${method} ${path}`, attempt: attempt + 1, cause: describeFailure(err, headersAbort?.signal.aborted ? headersMs : timeoutMs), delayMs });
129
188
  await sleep(delayMs);
130
189
  continue;
@@ -146,10 +205,15 @@ export class HttpClient {
146
205
  const error = await errorFrom(res, this.opts.source);
147
206
  const retryableStatus = res.status === 429 || res.status === 502 || res.status === 503 || res.status === 504;
148
207
  const retryableError = error instanceof ShardfluxApiError ? error.retryable && retryableStatus : retryableStatus;
149
- if (!retriable || !retryableError || attempt >= this.opts.maxRetries)
208
+ // The working-at-once quota (cell 429 quota_exceeded): the gateway refuses the call at admission, before the
209
+ // workspace is touched, so nothing ran and even a non-idempotent request (exec start, stdin, PTY create,
210
+ // keepalive) is safe to send again. Same Retry-After wait and retry budget as every other retry.
211
+ const admissionRefusal = init.quotaRetry !== false && isWorkingQuotaRefusal(error);
212
+ if ((!retriable && !admissionRefusal) || !retryableError || counted >= this.opts.maxRetries)
150
213
  throw error;
151
214
  const hinted = error instanceof ShardfluxApiError ? error.retryAfterSeconds : undefined;
152
- const delayMs = Math.min(5_000, hinted !== undefined ? hinted * 1000 : 200 * 2 ** attempt);
215
+ const delayMs = Math.min(5_000, hinted !== undefined ? hinted * 1000 : 200 * 2 ** counted);
216
+ counted += 1;
153
217
  init.onRetry?.({ request: `${method} ${path}`, attempt: attempt + 1, cause: describeFailure(error), delayMs });
154
218
  await sleep(delayMs);
155
219
  }
package/dist/index.d.ts CHANGED
@@ -17,8 +17,8 @@ export type { components as CellComponents, paths as CellPaths } from './generat
17
17
  export { BillingApi, Shardflux, WorkspacesApi, fetchBillingCatalog, pickByKey } from './client.js';
18
18
  export type { AgentSession, AllocationMode, BillingCatalog, BillingSubscription, Caps, CheckoutSession, DiskLayout, Entitlements, FindByKeyOptions, ForkTarget, Invoice, InvoicePage, LifetimeFilter, ListParams, Me, OpenParams, OpenResponse, Operation, Page, PortalSession, PurposeFilter, ResetWorkspaceBody, ResumeAnswer, ResumeRequestOptions, ResumeResponse, ShardfluxOptions, SuspendRequest, SuspendWhenIdleOptions, SuspendWhenIdleResponse, SuspendWhenIdleResult, WaitOptions, WorkspaceInputs, WorkspaceLifetime, WorkspaceMemory, WorkspaceOrigin, WorkspacePurpose, WorkspaceView, } from './client.js';
19
19
  export type { FinishedOperation, ForkOptions, LifecycleOptions, ResumeOptions, SuspendOptions, WaitedForkOptions, WaitedLifecycleOptions, WaitedResumeOptions, WaitedSuspendOptions } from './lifecycle.js';
20
- export { durabilityOf, formatTiming, isDurable, lostSuspendOf } from './progress.js';
21
- export type { Durability, LostSuspend, LifecycleAction, LifecyclePhase, LifecycleTiming, ProgressEvent, ProgressListener, RetryRecord, ServerTiming, TimingOutcome, TimingPhase, } from './progress.js';
20
+ export { durabilityOf, formatTiming, hostLostOf, isDurable, lostSuspendOf } from './progress.js';
21
+ export type { ColdBootReason, Durability, HostLost, LostSuspend, LifecycleAction, LifecyclePhase, LifecycleTiming, ProgressEvent, ProgressListener, RetryRecord, ServerTiming, TimingOutcome, TimingPhase, } from './progress.js';
22
22
  export { FEEDBACK_CATEGORIES, FEEDBACK_MESSAGE_MAX_LENGTH } from './feedback.js';
23
23
  export type { AccountFeedbackParams, FeedbackCategory, FeedbackContext, FeedbackReceipt, SendFeedbackParams } from './feedback.js';
24
24
  export { UsageApi } from './usage.js';
package/dist/index.js CHANGED
@@ -1,5 +1,5 @@
1
1
  export { BillingApi, Shardflux, WorkspacesApi, fetchBillingCatalog, pickByKey } from "./client.js";
2
- export { durabilityOf, formatTiming, isDurable, lostSuspendOf } from "./progress.js";
2
+ export { durabilityOf, formatTiming, hostLostOf, isDurable, lostSuspendOf } from "./progress.js";
3
3
  export { FEEDBACK_CATEGORIES, FEEDBACK_MESSAGE_MAX_LENGTH } from "./feedback.js";
4
4
  export { UsageApi } from "./usage.js";
5
5
  export { TemplateBuildTimeoutError, TemplateBuildsApi, TemplateDraftApi, TemplatePackagesApi, TemplateUploadError, TemplateUploadsApi, TemplateVersionTestInstancesApi, TemplateVersionsApi, TemplatesApi, buildSettled, saveAsTemplateBody, } from "./templates.js";
@@ -74,9 +74,10 @@ export interface ServerTiming {
74
74
  warmFallback: string | null;
75
75
  /**
76
76
  * `result.resume_path`: `local_cache`, `prestaged` (copied to this host ahead of the resume),
77
- * `download` (the checkpoint had to be fetched first), `cold_boot` (0.11.0+: the saved disk was booted after a
78
- * platform runtime change; processes restarted; see `memoryRestored`) or `reset_blank_layer` (the first start after a
79
- * reset).
77
+ * `download` (the checkpoint had to be fetched first), `cold_boot` (0.11.0+: the saved disk was booted instead of
78
+ * restoring memory, see `coldBootReason`; processes restarted; see `memoryRestored`), `reset_blank_layer` (the first
79
+ * start after a reset) or `thaw` (0.13.1+: resumed while its suspend was still being written; the same VM continued
80
+ * in place).
80
81
  */
81
82
  resumePath: string | null;
82
83
  /**
@@ -87,10 +88,11 @@ export interface ServerTiming {
87
88
  */
88
89
  memoryRestored?: boolean | null;
89
90
  /**
90
- * `result.cold_boot_reason` (0.11.0+), with `resumePath` `cold_boot`: why the memory could not be restored, e.g.
91
- * `runtime_changed` (the platform's VM runtime changed after the suspend). Null otherwise.
91
+ * `result.cold_boot_reason` (0.11.0+), with `resumePath` `cold_boot`: why the memory could not be restored:
92
+ * `runtime_changed` (the platform's VM runtime changed after the suspend) or `host_lost` (0.13.1+: the machine the
93
+ * workspace ran on failed; it booted from its disk, files kept, see `hostLost`). Null otherwise.
92
94
  */
93
- coldBootReason?: string | null;
95
+ coldBootReason?: ColdBootReason | null;
94
96
  /**
95
97
  * `result.durable` (0.12.0+), suspend and fork: `true` when the capture is in durable storage, `false` while its
96
98
  * durable copy is being written (see `durability`). `null` when the result does not say (other kinds, an older API).
@@ -105,6 +107,12 @@ export interface ServerTiming {
105
107
  * checkpoint before it (`restoredCheckpointId`, state as of `stateAsOf`). Null otherwise.
106
108
  */
107
109
  lostSuspend?: LostSuspend | null;
110
+ /**
111
+ * `result.host_lost` (0.13.1+): the machine the workspace ran on failed. A resume (or open) that restored the
112
+ * workspace says from what (`restoredFrom` `disk`, or `checkpoint` with `stateAsOf`); a suspend that found it so
113
+ * succeeds with only `detectedAt`. Null otherwise.
114
+ */
115
+ hostLost?: HostLost | null;
108
116
  /** `result.boot_to_ready_ms`: VM start until the guest agent answered. */
109
117
  bootToReadyMs: number | null;
110
118
  /** `result.host_timings_ms`: the host's own steps (restore: load, after_restore, ready, …). */
@@ -143,6 +151,30 @@ export interface LostSuspend {
143
151
  /** The restored state is as of this time (RFC 3339). */
144
152
  stateAsOf: string | null;
145
153
  }
154
+ /**
155
+ * `result.cold_boot_reason` (0.11.0+): `runtime_changed` (the platform's VM runtime changed after the suspend) or
156
+ * `host_lost` (0.13.1+: the machine the workspace ran on failed). Any other string is a reason this version does not
157
+ * know.
158
+ */
159
+ export type ColdBootReason = 'runtime_changed' | 'host_lost' | (string & {});
160
+ /**
161
+ * `result.host_lost` (0.13.1+), camelCased: the machine the workspace ran on failed. The workspace was moved to
162
+ * `suspended` at that moment, and its next use (a resume, a tool call's wake, an open) restored it. See the lifecycle
163
+ * reference.
164
+ */
165
+ export interface HostLost {
166
+ /** When the failure was detected (RFC 3339). */
167
+ detectedAt: string | null;
168
+ /**
169
+ * What the resume restored: `disk` (the workspace's own disk, files kept; `coldBootReason` `host_lost`, processes
170
+ * restarted) or `checkpoint` (its newest checkpoint, see `stateAsOf`). Null on a suspend's result.
171
+ */
172
+ restoredFrom: 'disk' | 'checkpoint' | null;
173
+ /** `checkpoint`: the checkpoint the resume restored. */
174
+ restoredCheckpointId: string | null;
175
+ /** `checkpoint`: the restored state is as of this time (RFC 3339); changes after it are not in the workspace. */
176
+ stateAsOf: string | null;
177
+ }
146
178
  /**
147
179
  * The durable copy of a suspend or fork operation (0.12.0+): `result.durability` camelCased, or null when the result has
148
180
  * none (the suspend was stored durably before it completed, or another kind).
@@ -150,6 +182,11 @@ export interface LostSuspend {
150
182
  export declare function durabilityOf(op: Pick<Operation, 'result'>): Durability | null;
151
183
  /** A resume operation's `result.lost_suspend` camelCased (0.12.0+), or null. */
152
184
  export declare function lostSuspendOf(op: Pick<Operation, 'result'>): LostSuspend | null;
185
+ /**
186
+ * An operation's `result.host_lost` camelCased (0.13.1+), or null: on a resume or open that restored a workspace whose
187
+ * machine failed, and on a suspend that found it so (`detectedAt` only).
188
+ */
189
+ export declare function hostLostOf(op: Pick<Operation, 'result'>): HostLost | null;
153
190
  /**
154
191
  * Whether a suspend or fork operation's capture is in durable storage (0.12.0+): `result.durable`, else `true` for a
155
192
  * succeeded suspend or fork whose result predates the field, else null.
@@ -255,7 +292,10 @@ export declare function traced<T>(trace: Trace, fn: () => Promise<T>): Promise<T
255
292
  *
256
293
  * A call that retried a request adds `retries: <n> (<request>, <cause>, after <delay>)`.
257
294
  * A resume that booted the workspace instead of restoring its memory (0.11.0+) says so in the server line:
258
- * `resume from cold_boot: processes restarted (runtime_changed)`.
295
+ * `resume from cold_boot: processes restarted (runtime_changed)`, or `(host_lost)` (0.13.1+) when the machine the
296
+ * workspace ran on failed and it booted from its disk. A resume that restored the newest checkpoint after such a
297
+ * failure adds `restored checkpoint <id> (host_lost; state as of <time>)`; a suspend that found the machine failed
298
+ * adds `host_lost (detected <time>)`.
259
299
  */
260
300
  export declare function formatTiming(t: LifecycleTiming): string;
261
301
  export {};
package/dist/progress.js CHANGED
@@ -40,6 +40,22 @@ export function lostSuspendOf(op) {
40
40
  stateAsOf: str(r.state_as_of),
41
41
  };
42
42
  }
43
+ /**
44
+ * An operation's `result.host_lost` camelCased (0.13.1+), or null: on a resume or open that restored a workspace whose
45
+ * machine failed, and on a suspend that found it so (`detectedAt` only).
46
+ */
47
+ export function hostLostOf(op) {
48
+ const h = op.result?.host_lost;
49
+ if (typeof h !== 'object' || h === null || Array.isArray(h))
50
+ return null;
51
+ const r = h;
52
+ return {
53
+ detectedAt: str(r.detected_at),
54
+ restoredFrom: r.restored_from === 'disk' || r.restored_from === 'checkpoint' ? r.restored_from : null,
55
+ restoredCheckpointId: str(r.restored_checkpoint_id),
56
+ stateAsOf: str(r.state_as_of),
57
+ };
58
+ }
43
59
  /**
44
60
  * Whether a suspend or fork operation's capture is in durable storage (0.12.0+): `result.durable`, else `true` for a
45
61
  * succeeded suspend or fork whose result predates the field, else null.
@@ -122,6 +138,7 @@ export function serverTiming(op) {
122
138
  suspendPath: str(r.suspend_path),
123
139
  durability: durabilityOf(op),
124
140
  lostSuspend: lostSuspendOf(op),
141
+ hostLost: hostLostOf(op),
125
142
  bootToReadyMs: num(r.boot_to_ready_ms),
126
143
  hostTimingsMs,
127
144
  };
@@ -275,6 +292,19 @@ export async function traced(trace, fn) {
275
292
  }
276
293
  }
277
294
  const fmt = (ms) => (ms === null ? '?' : ms < 1000 ? `${Math.round(ms)} ms` : `${(ms / 1000).toFixed(2)} s`);
295
+ /**
296
+ * The server-line part of `result.host_lost` (0.13.1+): the checkpoint a resume restored, or a suspend's detection time.
297
+ * A resume from the disk says nothing more here: its cold boot reads `processes restarted (host_lost)`.
298
+ */
299
+ function hostLostText(h) {
300
+ if (!h)
301
+ return null;
302
+ if (h.restoredFrom === 'checkpoint')
303
+ return `restored checkpoint ${h.restoredCheckpointId ?? '?'} (host_lost; state as of ${h.stateAsOf ?? '?'})`;
304
+ if (h.restoredFrom === null)
305
+ return `host_lost${h.detectedAt ? ` (detected ${h.detectedAt})` : ''}`;
306
+ return null;
307
+ }
278
308
  /**
279
309
  * A human-readable account of a timing, for logs and bug reports. A resume in production:
280
310
  *
@@ -285,7 +315,10 @@ const fmt = (ms) => (ms === null ? '?' : ms < 1000 ? `${Math.round(ms)} ms` : `$
285
315
  *
286
316
  * A call that retried a request adds `retries: <n> (<request>, <cause>, after <delay>)`.
287
317
  * A resume that booted the workspace instead of restoring its memory (0.11.0+) says so in the server line:
288
- * `resume from cold_boot: processes restarted (runtime_changed)`.
318
+ * `resume from cold_boot: processes restarted (runtime_changed)`, or `(host_lost)` (0.13.1+) when the machine the
319
+ * workspace ran on failed and it booted from its disk. A resume that restored the newest checkpoint after such a
320
+ * failure adds `restored checkpoint <id> (host_lost; state as of <time>)`; a suspend that found the machine failed
321
+ * adds `host_lost (detected <time>)`.
289
322
  */
290
323
  export function formatTiming(t) {
291
324
  const ids = [t.workspaceId ? `workspace ${t.workspaceId}` : null, t.operationId ? `operation ${t.operationId}` : null].filter(Boolean).join(', ');
@@ -308,6 +341,7 @@ export function formatTiming(t) {
308
341
  s.startPath ? `start ${s.startPath}${s.warmFallback ? ` (warm fallback: ${s.warmFallback})` : ''}` : null,
309
342
  s.resumePath ? `resume from ${s.resumePath}${s.memoryRestored === false ? `: processes restarted${s.coldBootReason ? ` (${s.coldBootReason})` : ''}` : ''}` : null,
310
343
  s.lostSuspend ? `restored ${s.lostSuspend.restoredCheckpointId ?? 'no checkpoint'} (latest suspend ${s.lostSuspend.checkpointId} ${s.lostSuspend.reason ?? 'lost'})` : null,
344
+ hostLostText(s.hostLost),
311
345
  s.suspendPath === 'local_commit' ? 'sealed on host' : null,
312
346
  s.durability?.state === 'durable' ? `durable${s.durability.localCommitToDurableMs !== null ? ` ${fmt(s.durability.localCommitToDurableMs)} later` : ''}` : s.durability ? `durable copy ${s.durability.state}` : null,
313
347
  s.bootToReadyMs !== null ? `boot to ready ${fmt(s.bootToReadyMs)}` : null,
package/dist/usage.d.ts CHANGED
@@ -56,7 +56,9 @@ export declare class UsageApi {
56
56
  constructor(ctx: () => ClientContext);
57
57
  /**
58
58
  * Current-period usage per meter, allowances with enforcement and cap state (`overage` while opt-in overage covers
59
- * usage past a CPU-hours or RAM GiB-hours allowance), measurement freshness, `exhausted_reason` (the 402
59
+ * usage past a CPU-hours or RAM GiB-hours allowance; `storage_blocked` (0.14.0+, enforcement `storage_block`) while
60
+ * the Retained state allowance is used up: opening a new key and forking are refused with 403 `quota_exceeded`,
61
+ * details.limit `retained_state`), measurement freshness, `exhausted_reason` (the 402
60
62
  * allowance_exhausted reason while starts are refused: `allowance_used`, `overage_paused`, `spend_cap_reached`) and
61
63
  * `spend_cap` (0.10.0).
62
64
  */
package/dist/usage.js CHANGED
@@ -9,7 +9,9 @@ export class UsageApi {
9
9
  }
10
10
  /**
11
11
  * Current-period usage per meter, allowances with enforcement and cap state (`overage` while opt-in overage covers
12
- * usage past a CPU-hours or RAM GiB-hours allowance), measurement freshness, `exhausted_reason` (the 402
12
+ * usage past a CPU-hours or RAM GiB-hours allowance; `storage_blocked` (0.14.0+, enforcement `storage_block`) while
13
+ * the Retained state allowance is used up: opening a new key and forking are refused with 403 `quota_exceeded`,
14
+ * details.limit `retained_state`), measurement freshness, `exhausted_reason` (the 402
13
15
  * allowance_exhausted reason while starts are refused: `allowance_used`, `overage_paused`, `spend_cap_reached`) and
14
16
  * `spend_cap` (0.10.0).
15
17
  */
@@ -67,7 +67,9 @@ export declare class Workspace {
67
67
  * woke the workspace), or a lifecycle call with `wait`. Null for handles from get()/list() until such a call.
68
68
  * `formatTiming(workspace.lastTiming)` prints it. After a resume or wake, `server.memoryRestored === false` (0.11.0+)
69
69
  * means the workspace booted from its saved disk instead of restoring its memory (`server.resumePath` `cold_boot`,
70
- * reason in `server.coldBootReason`): files are as of the suspend, running processes were restarted.
70
+ * reason in `server.coldBootReason`): files are as of the suspend, running processes were restarted. Reason
71
+ * `host_lost` (0.13.1+): the machine the workspace ran on failed and it booted from its disk, files kept;
72
+ * `server.hostLost` says what the resume restored (`hostLostOf(op)` on the operation).
71
73
  */
72
74
  get lastTiming(): LifecycleTiming | null;
73
75
  get id(): string;
@@ -174,6 +176,8 @@ export declare class Workspace {
174
176
  * resolve once it has FINISHED, with `workspace.state` then `suspended`. That is as soon as the workspace is sealed on
175
177
  * its host, typically in a few hundred ms. `result.durable` (also `lastTiming.server.durable`, 0.12.0+) turns true
176
178
  * when the copy lands in durable storage, typically within a second; `{ durable: true }` resolves only then.
179
+ * A suspend that finds the machine the workspace ran on failed succeeds at once (0.13.1+): `result.durable` is true
180
+ * and `result.host_lost` (`hostLostOf(op)`) has `detectedAt`; the next resume restores the workspace.
177
181
  *
178
182
  * await workspace.suspend({ wait: true });
179
183
  * await workspace.suspend({ durable: true }); // 0.12.0+: also wait for the durable copy
@@ -205,11 +209,16 @@ export declare class Workspace {
205
209
  * (also `lastTiming.server.memoryRestored`, 0.11.0+) is false when the resume booted the saved disk instead
206
210
  * (`resume_path` `cold_boot`): files kept, processes restarted. `result.lost_suspend` (also
207
211
  * `lastTiming.server.lostSuspend`, `lostSuspendOf(op)`, 0.12.0+) names a suspend this resume could not restore and the
208
- * checkpoint it restored instead (see the lifecycle reference).
212
+ * checkpoint it restored instead (see the lifecycle reference). `result.host_lost` (also `lastTiming.server.hostLost`,
213
+ * `hostLostOf(op)`, 0.13.1+): the machine the workspace ran on failed and this resume restored it, from its disk
214
+ * (`cold_boot_reason` `host_lost`) or from its newest checkpoint (`state_as_of`).
209
215
  */
210
216
  resume(opts: WaitedResumeOptions): Promise<FinishedOperation>;
211
217
  resume(opts?: ResumeOptions): Promise<Operation>;
212
- /** Takes a snapshot. Resolves when it is REQUESTED; with `{ wait: true }`, once it is taken. */
218
+ /**
219
+ * Takes a snapshot. Resolves when it is REQUESTED; with `{ wait: true }`, once it is taken. After the machine the
220
+ * workspace ran on failed, resume it first: until then the snapshot fails `resume_required` (0.13.1+).
221
+ */
213
222
  snapshot(opts: WaitedLifecycleOptions & {
214
223
  label?: string;
215
224
  }): Promise<FinishedOperation>;
@@ -218,7 +227,8 @@ export declare class Workspace {
218
227
  }): Promise<Operation>;
219
228
  /**
220
229
  * Forks into a new key. Resolves when the fork is REQUESTED (the copy's handle is returned at once); with
221
- * `{ wait: true }`, once the copy exists, with its handle refreshed.
230
+ * `{ wait: true }`, once the copy exists, with its handle refreshed. After the machine the workspace ran on failed,
231
+ * resume it first: until then the fork fails `resume_required` (0.13.1+).
222
232
  */
223
233
  fork(target: ForkTarget, opts: WaitedForkOptions): Promise<{
224
234
  operation: FinishedOperation;
@@ -301,8 +311,9 @@ export declare class Workspace {
301
311
  * that does not hold the request answers at once; the wake then waits for the operation and reads the view.
302
312
  *
303
313
  * A wake whose resume could not restore the workspace's memory (0.11.0+; the platform's VM runtime changed after the
304
- * suspend) still resolves `true`: the workspace runs from its saved disk, with every process restarted. `lastTiming`
305
- * and the `done` progress event say so (`server.memoryRestored === false`, `server.resumePath` `cold_boot`).
314
+ * suspend, or, 0.13.1+, the machine the workspace ran on failed) still resolves `true`: the workspace runs from its
315
+ * saved disk, with every process restarted. `lastTiming` and the `done` progress event say so
316
+ * (`server.memoryRestored === false`, `server.resumePath` `cold_boot`, `server.coldBootReason`).
306
317
  */
307
318
  wake(opts?: WakeOptions): Promise<boolean>;
308
319
  /**
package/dist/workspace.js CHANGED
@@ -44,7 +44,9 @@ export class Workspace {
44
44
  * woke the workspace), or a lifecycle call with `wait`. Null for handles from get()/list() until such a call.
45
45
  * `formatTiming(workspace.lastTiming)` prints it. After a resume or wake, `server.memoryRestored === false` (0.11.0+)
46
46
  * means the workspace booted from its saved disk instead of restoring its memory (`server.resumePath` `cold_boot`,
47
- * reason in `server.coldBootReason`): files are as of the suspend, running processes were restarted.
47
+ * reason in `server.coldBootReason`): files are as of the suspend, running processes were restarted. Reason
48
+ * `host_lost` (0.13.1+): the machine the workspace ran on failed and it booted from its disk, files kept;
49
+ * `server.hostLost` says what the resume restored (`hostLostOf(op)` on the operation).
48
50
  */
49
51
  get lastTiming() {
50
52
  return this.#lastTiming ?? this.#openTrace?.finished ?? null;
@@ -438,8 +440,9 @@ export class Workspace {
438
440
  * that does not hold the request answers at once; the wake then waits for the operation and reads the view.
439
441
  *
440
442
  * A wake whose resume could not restore the workspace's memory (0.11.0+; the platform's VM runtime changed after the
441
- * suspend) still resolves `true`: the workspace runs from its saved disk, with every process restarted. `lastTiming`
442
- * and the `done` progress event say so (`server.memoryRestored === false`, `server.resumePath` `cold_boot`).
443
+ * suspend, or, 0.13.1+, the machine the workspace ran on failed) still resolves `true`: the workspace runs from its
444
+ * saved disk, with every process restarted. `lastTiming` and the `done` progress event say so
445
+ * (`server.memoryRestored === false`, `server.resumePath` `cold_boot`, `server.coldBootReason`).
443
446
  */
444
447
  async wake(opts = {}) {
445
448
  // A file-first workspace runs from creation and is never suspended: nothing to wake.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@shardflux/sdk",
3
- "version": "0.13.0",
3
+ "version": "0.13.1",
4
4
  "type": "module",
5
5
  "description": "Shardflux TypeScript SDK: open persistent agent workspaces by key and give your agent workspace tools (exec, files, processes, PTY, git, browser).",
6
6
  "license": "Apache-2.0",