@shardflux/sdk 0.13.0 → 0.13.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +80 -2
- package/README.md +25 -8
- package/dist/cell.d.ts +7 -4
- package/dist/cell.js +65 -10
- package/dist/errors.d.ts +20 -1
- package/dist/errors.js +10 -0
- package/dist/generated/app-api.d.ts +23 -8
- package/dist/generated/cell-api.d.ts +13 -9
- package/dist/http.d.ts +40 -3
- package/dist/http.js +74 -10
- package/dist/index.d.ts +2 -2
- package/dist/index.js +1 -1
- package/dist/progress.d.ts +47 -7
- package/dist/progress.js +35 -1
- package/dist/usage.d.ts +3 -1
- package/dist/usage.js +3 -1
- package/dist/workspace.d.ts +17 -6
- package/dist/workspace.js +6 -3
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -3,7 +3,86 @@
|
|
|
3
3
|
Every API the README shows is available from the version named here. Breaking changes ship in minor releases and are
|
|
4
4
|
marked **Breaking**.
|
|
5
5
|
|
|
6
|
-
## 0.
|
|
6
|
+
## 0.14.0 (not yet published)
|
|
7
|
+
|
|
8
|
+
### Workspaces working at once
|
|
9
|
+
|
|
10
|
+
Additive. The plan's workspace number caps workspaces working at the same moment; stored workspaces are bounded by
|
|
11
|
+
Retained state.
|
|
12
|
+
|
|
13
|
+
- A tool call that finds every workspace the plan allows to work at once busy is refused with 429 `quota_exceeded`
|
|
14
|
+
(retryable, `Retry-After`; `details.limit` `concurrent_workspaces`, `limit_value`, `current`,
|
|
15
|
+
`retry_after_seconds`). The call did not run, so the SDK sends the same request again after `Retry-After` (at most
|
|
16
|
+
5 s) within `maxRetries`, for every call: exec starts, stdin, PTY create, keepalive and process signals included.
|
|
17
|
+
A refusal that outlasts the retries is a `ShardfluxApiError` with `retryAfterSeconds`. `executions.run()` keeps
|
|
18
|
+
these retries and does not add its own on top; the wake hint is not retried (it never waits).
|
|
19
|
+
- `pty.read()` rejects with the refusal (`ShardfluxApiError`, e.g. 429 `quota_exceeded`, retryable) when an accepted
|
|
20
|
+
attach is closed with an error body and a 4xxx close code, instead of returning empty output.
|
|
21
|
+
- The API's 403 `quota_exceeded` is not retried, as before. New `details.limit`: `retained_state` (`limit_value` and
|
|
22
|
+
`current` in GiB): opening a new key and forking are refused while the plan's Retained state is used up; existing
|
|
23
|
+
workspaces keep running, waking, suspending and resuming. Usage summaries report that allowance with `enforcement`
|
|
24
|
+
`storage_block` and `cap_state` `storage_blocked`.
|
|
25
|
+
- Regenerated contract types: `ErrorCode` includes `quota_exceeded` for the workspace gateway; allowance
|
|
26
|
+
`enforcement` adds `storage_block`, `cap_state` adds `storage_blocked`.
|
|
27
|
+
|
|
28
|
+
## 0.13.1 (not yet published)
|
|
29
|
+
|
|
30
|
+
Patch, additive: the transport, a documented value, a fix for tool calls across a move and self-healing workspaces.
|
|
31
|
+
`@shardflux/cli` 0.8.1, `@shardflux/mcp` 0.7.1 and the `shardflux` bundle 0.10.1 follow (`^0.13.1`).
|
|
32
|
+
|
|
33
|
+
### Workspaces recover by themselves
|
|
34
|
+
|
|
35
|
+
Additive. When the machine a workspace runs on fails, the workspace is suspended, and its next use (a tool call's wake,
|
|
36
|
+
`resume()`, `open()`) restores it: from its own disk, files kept, or from its newest checkpoint. Nothing changes for
|
|
37
|
+
code that already handles a cold resume.
|
|
38
|
+
|
|
39
|
+
- `cold_boot_reason` / `ServerTiming.coldBootReason` can be `host_lost`: the resume booted the workspace's disk
|
|
40
|
+
(`resumePath` `cold_boot`, `memoryRestored` false, processes restarted). New type `ColdBootReason`
|
|
41
|
+
(`'runtime_changed' | 'host_lost'`, open to other strings; the field still accepts any string).
|
|
42
|
+
- `hostLostOf(op)` and `ServerTiming.hostLost`: `result.host_lost` camelCased (`HostLost`: `detectedAt`,
|
|
43
|
+
`restoredFrom` `'disk' | 'checkpoint'`, `restoredCheckpointId`, `stateAsOf`). A resume from the checkpoint has its
|
|
44
|
+
state as of `stateAsOf`; a suspend that found the machine failed succeeds with `durable: true` and only
|
|
45
|
+
`detectedAt` (`restoredFrom` null).
|
|
46
|
+
- `formatTiming()` prints `processes restarted (host_lost)` for the disk, `restored checkpoint <id> (host_lost; state as
|
|
47
|
+
of <time>)` for the checkpoint, and `host_lost (detected <time>)` for the suspend.
|
|
48
|
+
- Documented operation error codes: `resume_required` (a fork or snapshot of such a workspace before its resume; resume
|
|
49
|
+
it first; not retryable) and `workspace_storage_unavailable` with `details.reason` `host_lost` (not retryable).
|
|
50
|
+
`KnownErrorReason` gains `host_lost`.
|
|
51
|
+
|
|
52
|
+
### Connection reuse
|
|
53
|
+
|
|
54
|
+
Additive. The call after an agent's pause between tool calls reuses its connection instead of opening a new one.
|
|
55
|
+
|
|
56
|
+
- On Node 26 the pooled transport keeps an idle connection reusable for 5 minutes (was 4 s), so a request seconds or
|
|
57
|
+
minutes after the last one skips the TCP and TLS handshake. Pooled sockets send TCP keepalive probes after 60 s
|
|
58
|
+
idle, so NATs and firewalls along the way keep the connection open. Idle connections never keep a process alive. A
|
|
59
|
+
supplied `fetch`, other runtimes and `SHARDFLUX_HTTP_KEEPALIVE` work as before.
|
|
60
|
+
- A request the SDK retries (GET, a request with an `Idempotency-Key`, a read-only POST such as files search) whose
|
|
61
|
+
connection closes before the response arrives is sent again at once on a new connection. That one resend has no
|
|
62
|
+
backoff and does not count against `maxRetries`; timing records and progress `retry` events list it with
|
|
63
|
+
`delayMs: 0`. A further failure follows `maxRetries` and backoff as before, and a POST without an
|
|
64
|
+
`Idempotency-Key` is never sent twice.
|
|
65
|
+
|
|
66
|
+
### Faster suspend and resume
|
|
67
|
+
|
|
68
|
+
Additive. `resumePath` / `resume_path` can be `thaw`: a resume that arrives while its suspend is still being
|
|
69
|
+
written continues the same VM in place. Nothing else changes for your code.
|
|
70
|
+
|
|
71
|
+
### Tool calls across a move
|
|
72
|
+
|
|
73
|
+
Fix. A tool call made while its workspace finishes a resume or a move now runs once the workspace runs, instead of
|
|
74
|
+
failing with `409 conflict` (`workspace_not_running`).
|
|
75
|
+
|
|
76
|
+
- When the call's tool token was refused as not running and the wake then found the workspace already running (the
|
|
77
|
+
resume or move committed in between; `wake` resolves `false`), the call is retried with a current token: the held
|
|
78
|
+
resume's, or a new one. A refused call never ran, so the retry cannot run it twice. Before, the refusal surfaced.
|
|
79
|
+
- A refusal that comes back while the API keeps reporting the workspace running is retried after 500 ms, within the
|
|
80
|
+
same 3 wakes and `transitionTimeoutMs`, then surfaces as before.
|
|
81
|
+
- `onProgress` sees the retry as a `tool` event of type `retry`, cause `conflict workspace_not_running (the workspace
|
|
82
|
+
runs; new token)`.
|
|
83
|
+
- The CLI, the MCP server and the `shardflux` bundle get it through their `^0.13.0` dependency.
|
|
84
|
+
|
|
85
|
+
## 0.13.0
|
|
7
86
|
|
|
8
87
|
### Elastic memory
|
|
9
88
|
|
|
@@ -446,7 +525,6 @@ behavior is the automatic version check (below), which makes one background requ
|
|
|
446
525
|
- New exports: `FeedbackCategory`, `FeedbackContext`, `FeedbackReceipt`, `SendFeedbackParams`, `AccountFeedbackParams`,
|
|
447
526
|
`FEEDBACK_CATEGORIES`, `FEEDBACK_MESSAGE_MAX_LENGTH`.
|
|
448
527
|
|
|
449
|
-
|
|
450
528
|
## 0.8.0 (2026-09-28)
|
|
451
529
|
|
|
452
530
|
Types only; nothing changes at run time and the API is unchanged.
|
package/README.md
CHANGED
|
@@ -296,10 +296,10 @@ await cell.exec.run(['make', 'test']); // resumes the wo
|
|
|
296
296
|
```
|
|
297
297
|
|
|
298
298
|
**Detecting a cold resume (0.11.0+).** A resume restores memory and running processes from the checkpoint. After a
|
|
299
|
-
platform runtime update,
|
|
300
|
-
tool call wakes the workspace; check `memoryRestored`. The
|
|
301
|
-
processes start fresh, as after a reboot: start dev servers, databases and
|
|
302
|
-
says which it was:
|
|
299
|
+
platform runtime update, or after the machine the workspace ran on failed (0.13.1+), a resume can boot from the saved
|
|
300
|
+
disk instead of restoring memory (a cold boot), also when a tool call wakes the workspace; check `memoryRestored`. The
|
|
301
|
+
workspace runs with its files kept, and its processes start fresh, as after a reboot: start dev servers, databases and
|
|
302
|
+
background jobs again. The resume's timing says which it was:
|
|
303
303
|
|
|
304
304
|
```ts
|
|
305
305
|
const t = workspace.lastTiming; // after resume({ wait: true }), wake(), or a tool call that woke the workspace
|
|
@@ -310,11 +310,20 @@ if (t?.server?.memoryRestored === false) {
|
|
|
310
310
|
|
|
311
311
|
- `ServerTiming.memoryRestored`: `true` when memory and processes came back, `false` when the workspace booted
|
|
312
312
|
(`resumePath` `cold_boot`, or `reset_blank_layer` after a reset), `null` when the API did not say (older APIs).
|
|
313
|
-
- `ServerTiming.coldBootReason
|
|
314
|
-
`resume from cold_boot: processes restarted (runtime_changed)`.
|
|
313
|
+
- `ServerTiming.coldBootReason` (type `ColdBootReason`): why it booted, `runtime_changed` or `host_lost`
|
|
314
|
+
**(0.13.1+)**, else null. `formatTiming()` prints `resume from cold_boot: processes restarted (runtime_changed)`.
|
|
315
315
|
- The finished resume operation carries the same in `result` (`memory_restored`, `cold_boot_reason`, and
|
|
316
316
|
`cold_boot` with details such as `files_as_of`; it is informational and not part of the stable contract).
|
|
317
317
|
|
|
318
|
+
**Workspaces recover by themselves (0.13.1+).** When the machine a workspace runs on fails, the workspace is suspended,
|
|
319
|
+
and its next use (a tool call, `resume()`, `open()`) restores it; there is nothing extra to call.
|
|
320
|
+
`hostLostOf(op)` and `lastTiming.server.hostLost` (`HostLost`) say how:
|
|
321
|
+
|
|
322
|
+
- `restoredFrom: 'disk'`: it booted from its own disk, files kept (`coldBootReason` `host_lost`, processes restarted).
|
|
323
|
+
- `restoredFrom: 'checkpoint'`: it resumed its newest checkpoint (`restoredCheckpointId`); its state is as of
|
|
324
|
+
`stateAsOf`. `formatTiming()` adds `restored checkpoint <id> (host_lost; state as of <time>)`.
|
|
325
|
+
- A `suspend()` that finds the machine failed succeeds (`result.durable` true, `hostLost.detectedAt`).
|
|
326
|
+
|
|
318
327
|
`suspend`, `resume`, `fork`, `snapshot` and `delete` return the lifecycle operation.
|
|
319
328
|
`cloud.workspaces.waitForOperation(id)` waits for it (default timeout 5 minutes): each poll asks the API to hold
|
|
320
329
|
the response until the operation changes (`Prefer: wait`, at most 20 s per request), so completion arrives within
|
|
@@ -324,7 +333,15 @@ operation keeps running server side (a queued start until its `deadlineAt`); wai
|
|
|
324
333
|
same call. `templates.builds.waitForBuild()` waits the same way. A waited `open()` issues the first tool token
|
|
325
334
|
together with the final workspace read, so the first tool call starts at once.
|
|
326
335
|
|
|
327
|
-
From 0.12.0, Node 26 reuses TLS connections through a private HTTP/1.1 pool.
|
|
336
|
+
From 0.12.0, Node 26 reuses TLS connections through a private HTTP/1.1 pool. From 0.13.0, an idle connection stays
|
|
337
|
+
reusable for 5 minutes, so the call after an agent's pause between tool calls skips the TCP and TLS handshake, and
|
|
338
|
+
pooled sockets send TCP keepalive probes after 60 s idle. Idle connections never keep a process alive. Other runtimes
|
|
339
|
+
keep native fetch. A supplied `fetch` stays unchanged. `SHARDFLUX_HTTP_KEEPALIVE=0` forces connection closure for
|
|
340
|
+
diagnosis; `=1` opts into native pooling. No global dispatcher is installed.
|
|
341
|
+
|
|
342
|
+
A request the SDK retries (safe methods, requests with an `Idempotency-Key`, read-only POSTs) whose connection closes
|
|
343
|
+
before the response arrives is sent again at once on a new connection **(0.13.0+)**, without backoff and outside
|
|
344
|
+
`maxRetries`; timing records list it with `delayMs: 0`.
|
|
328
345
|
|
|
329
346
|
List and look up workspaces:
|
|
330
347
|
|
|
@@ -412,7 +429,7 @@ This resume brought the workspace back from its host's local cache, memory and r
|
|
|
412
429
|
running, including any time in `capacity_pending`; `ran` is the cell's work (placement, boot or restore, guest
|
|
413
430
|
readiness). `start` / `resume from` and `boot to ready` / `host …` are what the cell reported. A resume that booted
|
|
414
431
|
the saved disk instead of restoring memory reads `resume from cold_boot: processes restarted (runtime_changed)`
|
|
415
|
-
(0.11.0+).
|
|
432
|
+
(0.11.0+), or `(host_lost)` (0.13.1+) after the machine the workspace ran on failed.
|
|
416
433
|
- **outside the server** is your total minus the operation's: network, TLS, polling latency, view and token. A large
|
|
417
434
|
value with a small server total points at the connection between you and the API, not at the workspace.
|
|
418
435
|
- **retries** lists transient failures the SDK retried (cause and backoff).
|
package/dist/cell.d.ts
CHANGED
|
@@ -174,8 +174,9 @@ export interface CellClientOptions {
|
|
|
174
174
|
sleep?: (ms: number) => Promise<void>;
|
|
175
175
|
/**
|
|
176
176
|
* Wakes a suspended workspace (resume, or join the active resume/open) within `timeoutMs` and resolves once it runs;
|
|
177
|
-
* resolves `false` when there was nothing to wake (the API already reports it running)
|
|
178
|
-
*
|
|
177
|
+
* resolves `false` when there was nothing to wake (the API already reports it running): the refused call is then
|
|
178
|
+
* retried with a current token, since the transition that refused it ended meanwhile (0.13.1+). It throws when the
|
|
179
|
+
* wake fails (OperationFailedError) or outlasts `timeoutMs` (OperationTimeoutError).
|
|
179
180
|
* Workspace.cell() supplies `workspace.wake()`. A call refused with `workspace_not_running` wakes the workspace and is
|
|
180
181
|
* retried; a refused call was never executed, so the retry cannot duplicate it. `null` surfaces
|
|
181
182
|
* the refusal instead. A wake that put a new token into this client's manager (the held resume;
|
|
@@ -320,8 +321,10 @@ export declare class CellClient {
|
|
|
320
321
|
* One authorized request. Refreshes the token once on stale_epoch / 401. Lifecycle transitions,
|
|
321
322
|
* bounded in total by `transitionTimeoutMs`: a call refused with `workspace_busy` is retried after `Retry-After`; one
|
|
322
323
|
* refused with `workspace_not_running` (or whose token cannot be minted because the workspace is not running) wakes
|
|
323
|
-
* the workspace through `wake`, given the time left, and is retried with a fresh token, at most 3 wakes per call.
|
|
324
|
-
*
|
|
324
|
+
* the workspace through `wake`, given the time left, and is retried with a fresh token, at most 3 wakes per call. A
|
|
325
|
+
* wake that finds the workspace already running (`false`) is retried the same way: the transition that refused the
|
|
326
|
+
* call ended meanwhile (0.13.1+). Refused calls were never executed, so retrying is safe. When the budget is spent
|
|
327
|
+
* the refusal surfaces.
|
|
325
328
|
*/
|
|
326
329
|
request(method: string, path: string, init?: RequestOptions & TransitionOptions): Promise<Response>;
|
|
327
330
|
/**
|
package/dist/cell.js
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
import { ExecStartError, NotSupportedForModeError, ShardfluxApiError, ShardfluxProtocolError } from "./errors.js";
|
|
1
|
+
import { ExecStartError, NotSupportedForModeError, ShardfluxApiError, ShardfluxProtocolError, apiError, isErrorBody, isWorkingQuotaRefusal } from "./errors.js";
|
|
2
2
|
import { EXECUTION_ID, ExecutionResult, newExecutionId } from "./executions.js";
|
|
3
3
|
import { HttpClient, defaultSleep, randomId, treeRevisionOf } from "./http.js";
|
|
4
4
|
import { describeFailure, emitTo } from "./progress.js";
|
|
@@ -20,8 +20,13 @@ export const DEFAULT_TRANSITION_TIMEOUT_MS = 120_000;
|
|
|
20
20
|
* capture writes recorded before them (read-your-writes). Workspace.cell() sets it; the capture's own client does not.
|
|
21
21
|
*/
|
|
22
22
|
export const CAPTURE_BARRIER = Symbol('shardflux.captureBarrier');
|
|
23
|
-
/**
|
|
23
|
+
/**
|
|
24
|
+
* Wakes per call at most: a workspace that keeps being suspended again, or keeps refusing the call while the API reports
|
|
25
|
+
* it running, surfaces the refusal.
|
|
26
|
+
*/
|
|
24
27
|
const MAX_WAKES = 3;
|
|
28
|
+
/** The pause before retrying a call refused again while the API reports the workspace running (from the second time). */
|
|
29
|
+
const RUNNING_AGAIN_PAUSE_MS = 500;
|
|
25
30
|
/** exec.cancel grace: the workspace waits 5 s without one and takes at most 60 s. */
|
|
26
31
|
const DEFAULT_CANCEL_GRACE_MS = 5_000;
|
|
27
32
|
const MAX_CANCEL_GRACE_MS = 60_000;
|
|
@@ -63,6 +68,30 @@ export async function* ndjson(body) {
|
|
|
63
68
|
reader.releaseLock();
|
|
64
69
|
}
|
|
65
70
|
}
|
|
71
|
+
/**
|
|
72
|
+
* A refused attach the gateway had to accept (cell-api.yaml "WebSocket close codes": an allowed Origin cannot read the
|
|
73
|
+
* HTTP answer to a failed upgrade): its only text frame is exactly the ErrorBody, then a close with 4000 + the HTTP
|
|
74
|
+
* status (4410 for not_found). A 4429 without a body (its close reason is `{"code","retryable",...}`) is the same
|
|
75
|
+
* refusal in short. Undefined for any other close.
|
|
76
|
+
*/
|
|
77
|
+
function attachRefusal(code, reason, body) {
|
|
78
|
+
const status = code === 4410 ? 404 : code !== undefined && code >= 4000 && code < 5000 ? code - 4000 : undefined;
|
|
79
|
+
if (body) {
|
|
80
|
+
const hinted = body.error.details?.retry_after_seconds;
|
|
81
|
+
return apiError(status ?? 502, body, 'cell', typeof hinted === 'number' && hinted >= 0 ? hinted : undefined);
|
|
82
|
+
}
|
|
83
|
+
if (code !== 4429)
|
|
84
|
+
return undefined;
|
|
85
|
+
let short = {};
|
|
86
|
+
try {
|
|
87
|
+
short = JSON.parse(reason ?? '');
|
|
88
|
+
}
|
|
89
|
+
catch {
|
|
90
|
+
// Not the compact JSON: the close code alone says it.
|
|
91
|
+
}
|
|
92
|
+
const errCode = typeof short.code === 'string' ? short.code : 'rate_limited';
|
|
93
|
+
return apiError(429, { error: { code: errCode, message: `The attach was refused (${errCode}); retry shortly.`, request_id: typeof short.request_id === 'string' ? short.request_id : '', retryable: short.retryable !== false } }, 'cell');
|
|
94
|
+
}
|
|
66
95
|
class ByteSink {
|
|
67
96
|
#max;
|
|
68
97
|
#chunks = [];
|
|
@@ -108,6 +137,10 @@ const executionRetryDelayMs = (failures, retryAfterSeconds) => retryAfterSeconds
|
|
|
108
137
|
const attemptTimedOut = (err) => err instanceof Error && err.name === 'TimeoutError';
|
|
109
138
|
/** A transport failure or a retryable 429/5xx: the execution request may be sent again with the same id. */
|
|
110
139
|
function transientExecutionFailure(err) {
|
|
140
|
+
// The working-at-once refusal was already retried by the HTTP layer within the client's maxRetries: a second loop
|
|
141
|
+
// here would multiply that budget, so it surfaces.
|
|
142
|
+
if (isWorkingQuotaRefusal(err))
|
|
143
|
+
return null;
|
|
111
144
|
if (err instanceof ShardfluxApiError) {
|
|
112
145
|
const status = err.status === 429 || err.status === 502 || err.status === 503 || err.status === 504;
|
|
113
146
|
return status && err.retryable ? { retryAfterSeconds: err.retryAfterSeconds } : null;
|
|
@@ -213,8 +246,10 @@ export class CellClient {
|
|
|
213
246
|
* One authorized request. Refreshes the token once on stale_epoch / 401. Lifecycle transitions,
|
|
214
247
|
* bounded in total by `transitionTimeoutMs`: a call refused with `workspace_busy` is retried after `Retry-After`; one
|
|
215
248
|
* refused with `workspace_not_running` (or whose token cannot be minted because the workspace is not running) wakes
|
|
216
|
-
* the workspace through `wake`, given the time left, and is retried with a fresh token, at most 3 wakes per call.
|
|
217
|
-
*
|
|
249
|
+
* the workspace through `wake`, given the time left, and is retried with a fresh token, at most 3 wakes per call. A
|
|
250
|
+
* wake that finds the workspace already running (`false`) is retried the same way: the transition that refused the
|
|
251
|
+
* call ended meanwhile (0.13.1+). Refused calls were never executed, so retrying is safe. When the budget is spent
|
|
252
|
+
* the refusal surfaces.
|
|
218
253
|
*/
|
|
219
254
|
async request(method, path, init = {}) {
|
|
220
255
|
const closer = this.#closer.signal;
|
|
@@ -237,6 +272,7 @@ export class CellClient {
|
|
|
237
272
|
const retry = (cause, delayMs, attempt) => emit?.({ type: 'retry', retry: { atMs: Math.round((performance.now() - t0) * 10) / 10, request: `${method} ${path}`, attempt, cause, delayMs } });
|
|
238
273
|
let refreshed = false;
|
|
239
274
|
let wakes = 0;
|
|
275
|
+
let runningAgain = 0;
|
|
240
276
|
let attempts = 0;
|
|
241
277
|
for (;;) {
|
|
242
278
|
try {
|
|
@@ -270,8 +306,17 @@ export class CellClient {
|
|
|
270
306
|
if (notRunning && wakeAllowed && this.#wake && wakes < MAX_WAKES && left > 0) {
|
|
271
307
|
wakes += 1;
|
|
272
308
|
const before = this.tokens.current;
|
|
273
|
-
if ((await this.#wake(left, signal)) === false)
|
|
274
|
-
|
|
309
|
+
if ((await this.#wake(left, signal)) === false) {
|
|
310
|
+
// Running per the API: the refusal was answered from a transition that ended before the wake looked (a
|
|
311
|
+
// resume or a move committed in between: the token request was refused while the workspace was resuming,
|
|
312
|
+
// and the wake found it running). The refused call never ran, so it is retried with a current token. A
|
|
313
|
+
// refusal that comes back while the API keeps reporting the workspace running is retried after a pause,
|
|
314
|
+
// within the same wakes and time budget, then surfaces.
|
|
315
|
+
const pause = runningAgain++ === 0 ? 0 : Math.max(0, Math.min(deadline - Date.now(), RUNNING_AGAIN_PAUSE_MS));
|
|
316
|
+
retry(`${err.code === 'conflict' ? 'conflict workspace_not_running' : err.code} (the workspace runs; new token)`, pause, (attempts += 1));
|
|
317
|
+
if (pause > 0)
|
|
318
|
+
await this.#opts.sleep(pause);
|
|
319
|
+
}
|
|
275
320
|
// A held resume handed this client a token of the woken workspace: use it. Otherwise the
|
|
276
321
|
// old token is of the previous epoch: fetch a new one.
|
|
277
322
|
if (this.tokens.current === before)
|
|
@@ -744,6 +789,7 @@ export class CellClient {
|
|
|
744
789
|
let next = opts.offset ?? 0;
|
|
745
790
|
let session = null;
|
|
746
791
|
let exited = false;
|
|
792
|
+
let refusal;
|
|
747
793
|
await new Promise((resolve, reject) => {
|
|
748
794
|
const onClose = () => finish(this.#closer.signal.reason instanceof Error ? this.#closer.signal.reason : new Error('closed'));
|
|
749
795
|
const ws = new WS(url, { headers: { authorization: `Bearer ${token}` } });
|
|
@@ -760,8 +806,10 @@ export class CellClient {
|
|
|
760
806
|
catch {
|
|
761
807
|
// already closed
|
|
762
808
|
}
|
|
763
|
-
|
|
764
|
-
|
|
809
|
+
// A refusal body whose close did not arrive within the read: the body alone says it.
|
|
810
|
+
const failure = err ?? (refusal ? attachRefusal(undefined, undefined, refusal) : undefined);
|
|
811
|
+
if (failure)
|
|
812
|
+
reject(failure);
|
|
765
813
|
else
|
|
766
814
|
resolve();
|
|
767
815
|
};
|
|
@@ -775,6 +823,12 @@ export class CellClient {
|
|
|
775
823
|
if (typeof e.data !== 'string')
|
|
776
824
|
return;
|
|
777
825
|
const m = JSON.parse(e.data);
|
|
826
|
+
// A refused attach (e.g. 4429 quota_exceeded: every working slot is taken): the bare ErrorBody, then the close.
|
|
827
|
+
if (!('type' in m)) {
|
|
828
|
+
if (isErrorBody(m))
|
|
829
|
+
refusal = m;
|
|
830
|
+
return;
|
|
831
|
+
}
|
|
778
832
|
if (m.type === 'output' && m.data) {
|
|
779
833
|
const bytes = unb64(m.data);
|
|
780
834
|
sink.push(bytes);
|
|
@@ -791,7 +845,7 @@ export class CellClient {
|
|
|
791
845
|
}
|
|
792
846
|
});
|
|
793
847
|
ws.addEventListener('error', () => finish(new ShardfluxProtocolError('pty attach WebSocket failed', 0, 'cell')));
|
|
794
|
-
ws.addEventListener('close', () => finish());
|
|
848
|
+
ws.addEventListener('close', (e) => finish(attachRefusal(e.code, e.reason, refusal)));
|
|
795
849
|
if (this.closed)
|
|
796
850
|
onClose();
|
|
797
851
|
else
|
|
@@ -955,7 +1009,8 @@ export class CellClient {
|
|
|
955
1009
|
// Nothing of a file-first workspace sleeps: the cell would answer `resident`.
|
|
956
1010
|
if (this.mode === 'file_first')
|
|
957
1011
|
return Promise.resolve({ residency: 'resident' });
|
|
958
|
-
|
|
1012
|
+
// A hint never waits: a 429 quota_exceeded (every working slot is taken) surfaces at once, like its other refusals.
|
|
1013
|
+
return this.#json('POST', this.#p('/v1/workspaces/{workspace_id}/wake-hint'), { wake: false, busy: false, quotaRetry: false, timeoutMs: 10_000, ...(signal ? { signal } : {}) });
|
|
959
1014
|
}
|
|
960
1015
|
/** Read idle signals without recording activity or waking the workspace. */
|
|
961
1016
|
idle(signal) {
|
package/dist/errors.d.ts
CHANGED
|
@@ -70,8 +70,19 @@ export type ErrorCode = AppErrorCode | CellErrorCode;
|
|
|
70
70
|
* 409 conflict read_only_path (a files write under an immutable path, from the cell gateway). The reason
|
|
71
71
|
* `update_policy_not_available` is gone with the update policy. A build that fails on them carries `failure.code`
|
|
72
72
|
* immutable_path_missing (details.path) or immutable_image_too_large.
|
|
73
|
+
* Host loss (0.13.1; the machine a workspace ran on failed, and the workspace restores itself on its next use): the
|
|
74
|
+
* operation error `workspace_storage_unavailable` carries details.reason host_lost (and details.workspace_state
|
|
75
|
+
* `suspended`) when neither the workspace's disk nor a checkpoint could be restored (not retryable). A fork or snapshot
|
|
76
|
+
* of such a workspace before its resume fails with the operation error `resume_required` (details.reason host_lost;
|
|
77
|
+
* not retryable): resume the workspace first.
|
|
78
|
+
* Workspaces working at once (0.14.0): the cell gateway answers a tool call that finds every working slot of the plan
|
|
79
|
+
* taken with 429 `quota_exceeded` (retryable, Retry-After; details.limit `concurrent_workspaces`, limit_value, current,
|
|
80
|
+
* retry_after_seconds); the call did not run and the SDK retries it like any transient refusal. The API's 403
|
|
81
|
+
* `quota_exceeded` (not retryable) names details.limit `concurrent_workspaces` on open, resume and fork, and
|
|
82
|
+
* `retained_state` (details.limit_value and details.current in GiB) when opening a new key or forking with the plan's
|
|
83
|
+
* Retained state used up. `details.limit` is not a reason: these are not in this union.
|
|
73
84
|
*/
|
|
74
|
-
export type KnownErrorReason = 'invalid_recipe' | 'base_not_layered' | 'language_unavailable' | 'language_conflict' | 'invalid_package' | 'too_many_files' | 'platform_owned_path' | 'upload_required' | 'upload_missing' | 'upload_digest_mismatch' | 'upload_too_large' | 'extra_hosts_without_auto' | 'invalid_settings' | 'services_unsupported' | 'input_required' | 'input_unknown' | 'input_invalid' | 'egress_widening' | 'reserved_session_id' | 'env_collision' | 'reserved_template_slug' | 'package_index_unavailable' | 'package_not_found' | 'startup_failed' | 'service_not_ready' | 'secrets_unavailable' | 'workspace_not_running' | 'operation_in_progress' | 'workspace_deleted' | 'secret_not_available' | 'legacy_disk_layout' | 'not_session' | 'session_lifetime' | 'lifetime_mismatch' | 'not_resettable' | 'template_not_layered' | 'draft_exists' | 'draft_stale' | 'build_in_progress' | 'file_list_unavailable' | 'file_list_indexing' | 'guest_feature_unavailable' | 'confirm_destructive_required' | 'reserved_key_prefix' | 'invalid_defaults' | 'invalid_path' | 'too_many_acknowledged_findings' | 'template_dev_mode_role' | 'draft_not_found' | 'version_not_found' | 'path_not_found' | 'revision_mismatch' | 'edit_not_found' | 'edit_ambiguous' | 'edit_not_text' | 'patch_invalid' | 'host_capacity' | 'wake_failed' | 'workspace_fenced' | 'offline_unavailable' | 'offline_budget' | 'offline_changed' | 'host_feature_unavailable' | 'not_supported_for_mode' | 'mode_mismatch' | 'mode_not_available' | 'layout_unsupported' | 'tree_revision_mismatch' | 'outside_tree_root' | 'execution_in_progress' | 'execution_id_reused' | 'operation_id_reused' | 'no_execution_host' | 'lease_expired' | 'host_unreachable' | 'host_restarted' | 'tree_moved' | 'blob_missing' | 'blob_corrupt' | 'exec_failed_to_start' | 'invalid_cwd' | 'allowance_used' | 'overage_paused' | 'spend_cap_reached' | 'overage_unavailable' | 'spend_cap_required' | 'spend_cap_below_minimum' | 'spend_cap_above_plan_price' | 'spend_cap_below_charges' | 'version_mismatch' | 'allocation_mode_not_available' | 'requires_elastic' | 'exceeds_memory_mib' | 'burst_mode_not_supported' | 'burst_not_supported' | 'burst_size_exceeds_plan' | 'not_available' | 'shared_volumes' | 'fence_not_drained' | 'apply_pending' | 'park_failed' | 'workspace_resumed' | 'interrupted' | 'burst_lost' | 'disk_full' | 'apply_failed' | 'reverted' | 'revert_failed' | 'immutable_path_removed' | 'immutable_paths_unsupported_base' | 'read_only_path';
|
|
85
|
+
export type KnownErrorReason = 'invalid_recipe' | 'base_not_layered' | 'language_unavailable' | 'language_conflict' | 'invalid_package' | 'too_many_files' | 'platform_owned_path' | 'upload_required' | 'upload_missing' | 'upload_digest_mismatch' | 'upload_too_large' | 'extra_hosts_without_auto' | 'invalid_settings' | 'services_unsupported' | 'input_required' | 'input_unknown' | 'input_invalid' | 'egress_widening' | 'reserved_session_id' | 'env_collision' | 'reserved_template_slug' | 'package_index_unavailable' | 'package_not_found' | 'startup_failed' | 'service_not_ready' | 'secrets_unavailable' | 'workspace_not_running' | 'operation_in_progress' | 'workspace_deleted' | 'secret_not_available' | 'legacy_disk_layout' | 'not_session' | 'session_lifetime' | 'lifetime_mismatch' | 'not_resettable' | 'template_not_layered' | 'draft_exists' | 'draft_stale' | 'build_in_progress' | 'file_list_unavailable' | 'file_list_indexing' | 'guest_feature_unavailable' | 'confirm_destructive_required' | 'reserved_key_prefix' | 'invalid_defaults' | 'invalid_path' | 'too_many_acknowledged_findings' | 'template_dev_mode_role' | 'draft_not_found' | 'version_not_found' | 'path_not_found' | 'revision_mismatch' | 'edit_not_found' | 'edit_ambiguous' | 'edit_not_text' | 'patch_invalid' | 'host_capacity' | 'wake_failed' | 'workspace_fenced' | 'offline_unavailable' | 'offline_budget' | 'offline_changed' | 'host_feature_unavailable' | 'not_supported_for_mode' | 'mode_mismatch' | 'mode_not_available' | 'layout_unsupported' | 'tree_revision_mismatch' | 'outside_tree_root' | 'execution_in_progress' | 'execution_id_reused' | 'operation_id_reused' | 'no_execution_host' | 'lease_expired' | 'host_unreachable' | 'host_restarted' | 'tree_moved' | 'blob_missing' | 'blob_corrupt' | 'exec_failed_to_start' | 'invalid_cwd' | 'allowance_used' | 'overage_paused' | 'spend_cap_reached' | 'overage_unavailable' | 'spend_cap_required' | 'spend_cap_below_minimum' | 'spend_cap_above_plan_price' | 'spend_cap_below_charges' | 'version_mismatch' | 'allocation_mode_not_available' | 'requires_elastic' | 'exceeds_memory_mib' | 'burst_mode_not_supported' | 'burst_not_supported' | 'burst_size_exceeds_plan' | 'not_available' | 'shared_volumes' | 'fence_not_drained' | 'apply_pending' | 'park_failed' | 'workspace_resumed' | 'interrupted' | 'burst_lost' | 'disk_full' | 'apply_failed' | 'reverted' | 'revert_failed' | 'immutable_path_removed' | 'immutable_paths_unsupported_base' | 'read_only_path' | 'host_lost';
|
|
75
86
|
/** A known reason, or any other string the server sends (reasons are open-ended). */
|
|
76
87
|
export type ErrorReason = KnownErrorReason | (string & {});
|
|
77
88
|
export interface ErrorBodyLike {
|
|
@@ -106,6 +117,13 @@ export declare class ShardfluxApiError extends Error {
|
|
|
106
117
|
readonly treeRevision: number | undefined;
|
|
107
118
|
constructor(status: number, body: ErrorBodyLike, source: 'api' | 'cell', retryAfterSeconds?: number, treeRevision?: number);
|
|
108
119
|
}
|
|
120
|
+
/**
|
|
121
|
+
* The cell gateway's working-at-once refusal (0.14.0+): 429 `quota_exceeded`, retryable, with Retry-After. The gateway
|
|
122
|
+
* refuses the call at admission, before the workspace is touched, so nothing ran and any request (exec start, stdin,
|
|
123
|
+
* PTY create, keepalive included) may be sent again. The API's 403 `quota_exceeded` is a different refusal (not
|
|
124
|
+
* retryable) and never matches.
|
|
125
|
+
*/
|
|
126
|
+
export declare function isWorkingQuotaRefusal(err: unknown): err is ShardfluxApiError;
|
|
109
127
|
/** processful (the default: one VM keeps processes, memory and files) or file_first. */
|
|
110
128
|
export type WorkspaceMode = AppComponents['schemas']['WorkspaceMode'];
|
|
111
129
|
/**
|
|
@@ -197,6 +215,7 @@ export declare class OperationTimeoutError extends Error {
|
|
|
197
215
|
* The operation reached `failed` or `canceled`. A suspend-when-idle that found the workspace active (a tool call after
|
|
198
216
|
* the request, an attached stream) is `canceled` with `errorCode` `workspace_active` (`workspaceActive` true, 0.12.0+):
|
|
199
217
|
* nothing changed and the workspace keeps running. `waitUntilReady()` and `wake()` treat it as running, not as a failure.
|
|
218
|
+
* `resume_required` (0.13.1+): a fork or snapshot of a workspace whose machine failed; resume it, then call again.
|
|
200
219
|
*/
|
|
201
220
|
export declare class OperationFailedError extends Error {
|
|
202
221
|
readonly operation: Operation;
|
package/dist/errors.js
CHANGED
|
@@ -40,6 +40,15 @@ export class ShardfluxApiError extends Error {
|
|
|
40
40
|
this.reason = typeof reason === 'string' ? reason : undefined;
|
|
41
41
|
}
|
|
42
42
|
}
|
|
43
|
+
/**
|
|
44
|
+
* The cell gateway's working-at-once refusal (0.14.0+): 429 `quota_exceeded`, retryable, with Retry-After. The gateway
|
|
45
|
+
* refuses the call at admission, before the workspace is touched, so nothing ran and any request (exec start, stdin,
|
|
46
|
+
* PTY create, keepalive included) may be sent again. The API's 403 `quota_exceeded` is a different refusal (not
|
|
47
|
+
* retryable) and never matches.
|
|
48
|
+
*/
|
|
49
|
+
export function isWorkingQuotaRefusal(err) {
|
|
50
|
+
return err instanceof ShardfluxApiError && err.source === 'cell' && err.status === 429 && err.code === 'quota_exceeded' && err.retryable;
|
|
51
|
+
}
|
|
43
52
|
/**
|
|
44
53
|
* The call does not exist for the workspace's mode: 409 `conflict` with details.reason
|
|
45
54
|
* `not_supported_for_mode`, `details.mode` (the workspace's mode) and `details.operation`. File-first workspaces have no
|
|
@@ -187,6 +196,7 @@ export class OperationTimeoutError extends Error {
|
|
|
187
196
|
* The operation reached `failed` or `canceled`. A suspend-when-idle that found the workspace active (a tool call after
|
|
188
197
|
* the request, an attached stream) is `canceled` with `errorCode` `workspace_active` (`workspaceActive` true, 0.12.0+):
|
|
189
198
|
* nothing changed and the workspace keeps running. `waitUntilReady()` and `wake()` treat it as running, not as a failure.
|
|
199
|
+
* `resume_required` (0.13.1+): a fork or snapshot of a workspace whose machine failed; resume it, then call again.
|
|
190
200
|
*/
|
|
191
201
|
export class OperationFailedError extends Error {
|
|
192
202
|
operation;
|
|
@@ -5630,6 +5630,21 @@ export interface components {
|
|
|
5630
5630
|
} & {
|
|
5631
5631
|
[key: string]: unknown;
|
|
5632
5632
|
};
|
|
5633
|
+
host_lost?: {
|
|
5634
|
+
/** @description RFC 3339: when the host failure was detected; the workspace was suspended then. */
|
|
5635
|
+
detected_at: string;
|
|
5636
|
+
/**
|
|
5637
|
+
* @description resume and open: disk when the workspace booted from its disk (files kept, processes and memory not); checkpoint when it was restored from its newest checkpoint (see state_as_of). Absent on a suspend.
|
|
5638
|
+
* @enum {string}
|
|
5639
|
+
*/
|
|
5640
|
+
restored_from?: "disk" | "checkpoint";
|
|
5641
|
+
/** @description restored_from checkpoint: the checkpoint this resume restored. */
|
|
5642
|
+
restored_checkpoint_id?: string;
|
|
5643
|
+
/** @description restored_from checkpoint, RFC 3339: when that checkpoint committed; the workspace state is as of then. */
|
|
5644
|
+
state_as_of?: string;
|
|
5645
|
+
} & {
|
|
5646
|
+
[key: string]: unknown;
|
|
5647
|
+
};
|
|
5633
5648
|
} & {
|
|
5634
5649
|
[key: string]: unknown;
|
|
5635
5650
|
};
|
|
@@ -28007,8 +28022,8 @@ export interface operations {
|
|
|
28007
28022
|
display_name: string;
|
|
28008
28023
|
unit: string;
|
|
28009
28024
|
meters: ("cpu_seconds" | "memory_gib_seconds" | "storage_gib_seconds" | "egress_bytes" | "ingress_bytes" | "volume_storage_gib_seconds")[];
|
|
28010
|
-
/** @description hard_cap: exhausting it refuses opens/resumes (402 allowance_exhausted) and running workspaces are suspended; spare_capacity: no allowance cap (runs on reclaimable spare capacity); reported: shown and notified, never enforced by refusing access; egress_block (outbound_transfer_gb, contracts §18): at 100% outbound internet traffic is blocked by an organization egress override until upgrade, purchase or the next period (workspaces keep running, starts are not refused). */
|
|
28011
|
-
enforcement: "hard_cap" | "spare_capacity" | "reported" | "egress_block";
|
|
28025
|
+
/** @description hard_cap: exhausting it refuses opens/resumes (402 allowance_exhausted) and running workspaces are suspended; spare_capacity: no allowance cap (runs on reclaimable spare capacity); reported: shown and notified, never enforced by refusing access; egress_block (outbound_transfer_gb, contracts §18): at 100% outbound internet traffic is blocked by an organization egress override until upgrade, purchase or the next period (workspaces keep running, starts are not refused); storage_block (retained_state_gib, contracts §40): at 100% opening a new workspace key and forking are refused (403 quota_exceeded, details.limit retained_state) until retained state is below the allowance (existing workspaces keep running, waking, suspending and resuming; nothing is deleted or charged). */
|
|
28026
|
+
enforcement: "hard_cap" | "spare_capacity" | "reported" | "egress_block" | "storage_block";
|
|
28012
28027
|
included: number | null;
|
|
28013
28028
|
used: number;
|
|
28014
28029
|
remaining: number | null;
|
|
@@ -28016,8 +28031,8 @@ export interface operations {
|
|
|
28016
28031
|
included_meter_units: number | null;
|
|
28017
28032
|
used_meter_units: number | null;
|
|
28018
28033
|
remaining_meter_units: number | null;
|
|
28019
|
-
/** @description ok: below 80 %. warning: 80 % or more of the allowance. exhausted: a hard cap is used up and starts are refused (402 allowance_exhausted; see exhausted_reason). overage: a CPU-hours or RAM GiB-hours allowance is used up while opt-in overage is on and below its spend cap, so starts are admitted and the usage past it is charged (spend_cap). over_allowance: a reported allowance is exceeded (never refused). egress_blocked: the outbound transfer allowance is used up and outbound traffic is blocked. uncapped: runs on spare capacity. not_included: the plan does not define it. */
|
|
28020
|
-
cap_state: "ok" | "warning" | "exhausted" | "overage" | "over_allowance" | "egress_blocked" | "uncapped" | "not_included";
|
|
28034
|
+
/** @description ok: below 80 %. warning: 80 % or more of the allowance. exhausted: a hard cap is used up and starts are refused (402 allowance_exhausted; see exhausted_reason). overage: a CPU-hours or RAM GiB-hours allowance is used up while opt-in overage is on and below its spend cap, so starts are admitted and the usage past it is charged (spend_cap). over_allowance: a reported allowance is exceeded (never refused). egress_blocked: the outbound transfer allowance is used up and outbound traffic is blocked. storage_blocked: the retained state allowance is used up and new workspaces and forks are refused (403 quota_exceeded). uncapped: runs on spare capacity. not_included: the plan does not define it. */
|
|
28035
|
+
cap_state: "ok" | "warning" | "exhausted" | "overage" | "over_allowance" | "egress_blocked" | "storage_blocked" | "uncapped" | "not_included";
|
|
28021
28036
|
}[];
|
|
28022
28037
|
meters: {
|
|
28023
28038
|
/** @enum {string} */
|
|
@@ -29062,8 +29077,8 @@ export interface operations {
|
|
|
29062
29077
|
display_name: string;
|
|
29063
29078
|
unit: string;
|
|
29064
29079
|
meters: ("cpu_seconds" | "memory_gib_seconds" | "storage_gib_seconds" | "egress_bytes" | "ingress_bytes" | "volume_storage_gib_seconds")[];
|
|
29065
|
-
/** @description hard_cap: exhausting it refuses opens/resumes (402 allowance_exhausted) and running workspaces are suspended; spare_capacity: no allowance cap (runs on reclaimable spare capacity); reported: shown and notified, never enforced by refusing access; egress_block (outbound_transfer_gb, contracts §18): at 100% outbound internet traffic is blocked by an organization egress override until upgrade, purchase or the next period (workspaces keep running, starts are not refused). */
|
|
29066
|
-
enforcement: "hard_cap" | "spare_capacity" | "reported" | "egress_block";
|
|
29080
|
+
/** @description hard_cap: exhausting it refuses opens/resumes (402 allowance_exhausted) and running workspaces are suspended; spare_capacity: no allowance cap (runs on reclaimable spare capacity); reported: shown and notified, never enforced by refusing access; egress_block (outbound_transfer_gb, contracts §18): at 100% outbound internet traffic is blocked by an organization egress override until upgrade, purchase or the next period (workspaces keep running, starts are not refused); storage_block (retained_state_gib, contracts §40): at 100% opening a new workspace key and forking are refused (403 quota_exceeded, details.limit retained_state) until retained state is below the allowance (existing workspaces keep running, waking, suspending and resuming; nothing is deleted or charged). */
|
|
29081
|
+
enforcement: "hard_cap" | "spare_capacity" | "reported" | "egress_block" | "storage_block";
|
|
29067
29082
|
included: number | null;
|
|
29068
29083
|
used: number;
|
|
29069
29084
|
remaining: number | null;
|
|
@@ -29071,8 +29086,8 @@ export interface operations {
|
|
|
29071
29086
|
included_meter_units: number | null;
|
|
29072
29087
|
used_meter_units: number | null;
|
|
29073
29088
|
remaining_meter_units: number | null;
|
|
29074
|
-
/** @description ok: below 80 %. warning: 80 % or more of the allowance. exhausted: a hard cap is used up and starts are refused (402 allowance_exhausted; see exhausted_reason). overage: a CPU-hours or RAM GiB-hours allowance is used up while opt-in overage is on and below its spend cap, so starts are admitted and the usage past it is charged (spend_cap). over_allowance: a reported allowance is exceeded (never refused). egress_blocked: the outbound transfer allowance is used up and outbound traffic is blocked. uncapped: runs on spare capacity. not_included: the plan does not define it. */
|
|
29075
|
-
cap_state: "ok" | "warning" | "exhausted" | "overage" | "over_allowance" | "egress_blocked" | "uncapped" | "not_included";
|
|
29089
|
+
/** @description ok: below 80 %. warning: 80 % or more of the allowance. exhausted: a hard cap is used up and starts are refused (402 allowance_exhausted; see exhausted_reason). overage: a CPU-hours or RAM GiB-hours allowance is used up while opt-in overage is on and below its spend cap, so starts are admitted and the usage past it is charged (spend_cap). over_allowance: a reported allowance is exceeded (never refused). egress_blocked: the outbound transfer allowance is used up and outbound traffic is blocked. storage_blocked: the retained state allowance is used up and new workspaces and forks are refused (403 quota_exceeded). uncapped: runs on spare capacity. not_included: the plan does not define it. */
|
|
29090
|
+
cap_state: "ok" | "warning" | "exhausted" | "overage" | "over_allowance" | "egress_blocked" | "storage_blocked" | "uncapped" | "not_included";
|
|
29076
29091
|
}[];
|
|
29077
29092
|
meters: {
|
|
29078
29093
|
/** @enum {string} */
|
|
@@ -191,7 +191,7 @@ export interface paths {
|
|
|
191
191
|
* Client messages are JSON `ExecClientMessage`. Close codes: 1000 after
|
|
192
192
|
* the `exit` event, 1001 on gateway shutdown (reconnect with offsets),
|
|
193
193
|
* 4401 unauthenticated, 4403 forbidden, 4409 stale epoch / workspace
|
|
194
|
-
* busy / workspace not running, 4410 workspace gone, 4429 rate limited,
|
|
194
|
+
* busy / workspace not running, 4410 workspace gone, 4429 rate limited / quota exceeded,
|
|
195
195
|
* 4500/4503/4504 internal/unavailable/timeout (4000 + HTTP status; see
|
|
196
196
|
* "WebSocket close codes" above; the close reason is compact JSON with
|
|
197
197
|
* the error `code`). A refused upgrade from an allowed Origin is
|
|
@@ -265,14 +265,18 @@ export interface paths {
|
|
|
265
265
|
get?: never;
|
|
266
266
|
put?: never;
|
|
267
267
|
/**
|
|
268
|
-
*
|
|
269
|
-
* @description Sends SIGTERM to the
|
|
270
|
-
* `grace_ms` (0 to 60000;
|
|
271
|
-
*
|
|
272
|
-
* grace and the SIGKILL. A
|
|
273
|
-
*
|
|
268
|
+
* Stop the command and every process it started (SIGTERM, SIGKILL after grace)
|
|
269
|
+
* @description Sends SIGTERM to every process the command started, background and
|
|
270
|
+
* setsid processes included, and SIGKILL after `grace_ms` (0 to 60000;
|
|
271
|
+
* 0 or omitted is 5000). Answers when they have all exited: at once when
|
|
272
|
+
* SIGTERM ends them, otherwise after the grace and the SIGKILL. A lost
|
|
273
|
+
* command is stopped the same way; a command that already exited keeps
|
|
274
|
+
* its result and what it left running is stopped. `503
|
|
275
|
+
* dependency_unavailable` (retryable) while a process cannot exit yet.
|
|
276
|
+
* Template versions before python-node-browser 13, ubuntu-24.04 4,
|
|
277
|
+
* ubuntu-22.04 2 and debian-12 2 signal the command's process group and
|
|
278
|
+
* return an ended session as it is. A `grace_ms` outside the range is `422 validation_failed`
|
|
274
279
|
* (`details.field` `grace_ms`, `details.minimum`, `details.maximum`).
|
|
275
|
-
* An ended session is returned as it is.
|
|
276
280
|
*/
|
|
277
281
|
post: operations["execCancel"];
|
|
278
282
|
delete?: never;
|
|
@@ -876,7 +880,7 @@ export type webhooks = Record<string, never>;
|
|
|
876
880
|
export interface components {
|
|
877
881
|
schemas: {
|
|
878
882
|
/** @enum {string} */
|
|
879
|
-
ErrorCode: "bad_request" | "unauthenticated" | "forbidden" | "not_found" | "conflict" | "stale_epoch" | "workspace_not_running" | "validation_failed" | "payload_too_large" | "range_not_satisfiable" | "rate_limited" | "capacity_pending" | "dependency_unavailable" | "timeout" | "internal_error" | "service_unavailable" | "workspace_busy" | "burst_unavailable" | "burst_apply_failed";
|
|
883
|
+
ErrorCode: "bad_request" | "unauthenticated" | "forbidden" | "not_found" | "conflict" | "stale_epoch" | "workspace_not_running" | "validation_failed" | "payload_too_large" | "range_not_satisfiable" | "rate_limited" | "capacity_pending" | "dependency_unavailable" | "timeout" | "internal_error" | "service_unavailable" | "workspace_busy" | "burst_unavailable" | "burst_apply_failed" | "quota_exceeded";
|
|
880
884
|
ErrorBody: {
|
|
881
885
|
error: {
|
|
882
886
|
code: components["schemas"]["ErrorCode"];
|
package/dist/http.d.ts
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
import type { RetryRecord } from './progress.js';
|
|
2
|
-
export declare const SDK_VERSION = "0.13.
|
|
2
|
+
export declare const SDK_VERSION = "0.13.1";
|
|
3
3
|
export interface RequestOptions {
|
|
4
4
|
query?: Record<string, string | number | boolean | undefined | null>;
|
|
5
5
|
json?: unknown;
|
|
@@ -10,6 +10,11 @@ export interface RequestOptions {
|
|
|
10
10
|
idempotencyKey?: string;
|
|
11
11
|
/** The request has no effect a retry could duplicate (a read-only POST such as files/search): retried like a GET. */
|
|
12
12
|
idempotent?: boolean;
|
|
13
|
+
/**
|
|
14
|
+
* false: the cell gateway's 429 `quota_exceeded` surfaces at once instead of being retried (the wake hint, which
|
|
15
|
+
* never waits). Default true.
|
|
16
|
+
*/
|
|
17
|
+
quotaRetry?: boolean;
|
|
13
18
|
signal?: AbortSignal;
|
|
14
19
|
/** Override the client's default request timeout (ms); 0 disables it (streams). */
|
|
15
20
|
timeoutMs?: number;
|
|
@@ -33,13 +38,45 @@ export interface HttpOptions {
|
|
|
33
38
|
onSuccess?: (() => void) | undefined;
|
|
34
39
|
}
|
|
35
40
|
export declare const defaultSleep: (ms: number) => Promise<void>;
|
|
41
|
+
/**
|
|
42
|
+
* How long an idle pooled connection stays reusable (ms): 5 minutes, instead of undici's 4 s default, so the request
|
|
43
|
+
* after an agent's usual pause between tool calls (15 s to a few minutes) skips a new TCP and TLS handshake. Far below
|
|
44
|
+
* the 3600 s idle timeout of the load balancer in front of the API and the cell endpoints, so the server side never
|
|
45
|
+
* closes a connection for idleness while the client still treats it as reusable; below the AWS NAT gateway's 350 s idle
|
|
46
|
+
* timeout; and the same limit Chrome keeps for used idle sockets. Also caps a server's `Keep-Alive: timeout` hint.
|
|
47
|
+
*/
|
|
48
|
+
export declare const KEEP_ALIVE_TIMEOUT_MS = 300000;
|
|
49
|
+
/**
|
|
50
|
+
* TCP keepalive probes start after a pooled socket has been idle this long (ms), so NATs and firewalls with shorter
|
|
51
|
+
* idle timeouts than the pool's (Azure SNAT: 4 minutes) keep the connection instead of silently dropping it. Equal to
|
|
52
|
+
* undici's own default, set explicitly so it is visible and tested.
|
|
53
|
+
*/
|
|
54
|
+
export declare const TCP_KEEPALIVE_INITIAL_DELAY_MS = 60000;
|
|
55
|
+
/** The private pool's undici Agent options (exported for tests). */
|
|
56
|
+
export declare const POOL_OPTIONS: {
|
|
57
|
+
readonly allowH2: false;
|
|
58
|
+
readonly pipelining: 1;
|
|
59
|
+
readonly keepAliveTimeout: 300000;
|
|
60
|
+
readonly keepAliveMaxTimeout: 300000;
|
|
61
|
+
readonly connect: {
|
|
62
|
+
readonly keepAlive: true;
|
|
63
|
+
readonly keepAliveInitialDelay: 60000;
|
|
64
|
+
};
|
|
65
|
+
};
|
|
36
66
|
/**
|
|
37
67
|
* Reuses TLS connections. Node 26's bundled undici 8.9 can stall reused connections: use pinned undici with a
|
|
38
|
-
* private HTTP/1.1-only dispatcher
|
|
39
|
-
* for diagnosis; =1 retains the historical opt-in to
|
|
68
|
+
* private HTTP/1.1-only dispatcher whose idle connections stay reusable for KEEP_ALIVE_TIMEOUT_MS. Other runtimes
|
|
69
|
+
* retain native fetch. SHARDFLUX_HTTP_KEEPALIVE=0 forces close for diagnosis; =1 retains the historical opt-in to
|
|
70
|
+
* native pooling. An explicitly supplied fetch is untouched. Idle pooled sockets do not keep the process alive.
|
|
40
71
|
*/
|
|
41
72
|
export declare function defaultFetch(env?: Record<string, string | undefined> | undefined, undiciVersion?: string | undefined): typeof fetch;
|
|
42
73
|
export declare function buildUrl(base: string, path: string, query?: RequestOptions['query']): string;
|
|
74
|
+
/**
|
|
75
|
+
* True when fetch failed because its connection closed or reset before a response arrived: what a request sees when
|
|
76
|
+
* it went out on a pooled connection the other side closed while it sat idle. Connect failures (ECONNREFUSED,
|
|
77
|
+
* ENOTFOUND, UND_ERR_CONNECT_TIMEOUT), timeouts and aborts are not this.
|
|
78
|
+
*/
|
|
79
|
+
export declare function isStaleConnectionError(err: unknown): boolean;
|
|
43
80
|
/** `X-Tree-Revision` (file-first workspaces): a non-negative integer, else null. */
|
|
44
81
|
export declare function treeRevisionOf(headers: Headers): number | null;
|
|
45
82
|
/** Turns a non-2xx response into a ShardfluxApiError (or a protocol error for undocumented bodies). */
|
package/dist/http.js
CHANGED
|
@@ -4,18 +4,44 @@
|
|
|
4
4
|
* retries only where a retry cannot duplicate an effect (safe methods,
|
|
5
5
|
* requests carrying an Idempotency-Key, and read-only POSTs marked
|
|
6
6
|
* `idempotent`). A retryable 429/502/503/504 (e.g. 503 `host_capacity`, 503
|
|
7
|
-
* `wake_failed`) waits `Retry-After` (at most 5 s) before the retry.
|
|
7
|
+
* `wake_failed`) waits `Retry-After` (at most 5 s) before the retry. Such a request whose connection closed before
|
|
8
|
+
* any response (a pooled connection gone stale while idle) is sent again at once, once, outside `maxRetries`. The cell
|
|
9
|
+
* gateway's 429 `quota_exceeded` (every working slot of the plan is taken) is retried for any request (0.14.0+): it is
|
|
10
|
+
* refused at admission, so nothing ran.
|
|
8
11
|
*/
|
|
9
|
-
import { ShardfluxApiError, ShardfluxProtocolError, apiError, isErrorBody } from "./errors.js";
|
|
12
|
+
import { ShardfluxApiError, ShardfluxProtocolError, apiError, isErrorBody, isWorkingQuotaRefusal } from "./errors.js";
|
|
10
13
|
import { describeFailure } from "./progress.js";
|
|
11
|
-
export const SDK_VERSION = '0.13.
|
|
14
|
+
export const SDK_VERSION = '0.13.1';
|
|
12
15
|
export const defaultSleep = (ms) => new Promise((r) => setTimeout(r, ms));
|
|
16
|
+
/**
|
|
17
|
+
* How long an idle pooled connection stays reusable (ms): 5 minutes, instead of undici's 4 s default, so the request
|
|
18
|
+
* after an agent's usual pause between tool calls (15 s to a few minutes) skips a new TCP and TLS handshake. Far below
|
|
19
|
+
* the 3600 s idle timeout of the load balancer in front of the API and the cell endpoints, so the server side never
|
|
20
|
+
* closes a connection for idleness while the client still treats it as reusable; below the AWS NAT gateway's 350 s idle
|
|
21
|
+
* timeout; and the same limit Chrome keeps for used idle sockets. Also caps a server's `Keep-Alive: timeout` hint.
|
|
22
|
+
*/
|
|
23
|
+
export const KEEP_ALIVE_TIMEOUT_MS = 300_000;
|
|
24
|
+
/**
|
|
25
|
+
* TCP keepalive probes start after a pooled socket has been idle this long (ms), so NATs and firewalls with shorter
|
|
26
|
+
* idle timeouts than the pool's (Azure SNAT: 4 minutes) keep the connection instead of silently dropping it. Equal to
|
|
27
|
+
* undici's own default, set explicitly so it is visible and tested.
|
|
28
|
+
*/
|
|
29
|
+
export const TCP_KEEPALIVE_INITIAL_DELAY_MS = 60_000;
|
|
30
|
+
/** The private pool's undici Agent options (exported for tests). */
|
|
31
|
+
export const POOL_OPTIONS = {
|
|
32
|
+
allowH2: false,
|
|
33
|
+
pipelining: 1,
|
|
34
|
+
keepAliveTimeout: KEEP_ALIVE_TIMEOUT_MS,
|
|
35
|
+
keepAliveMaxTimeout: KEEP_ALIVE_TIMEOUT_MS,
|
|
36
|
+
connect: { keepAlive: true, keepAliveInitialDelay: TCP_KEEPALIVE_INITIAL_DELAY_MS },
|
|
37
|
+
};
|
|
13
38
|
/** One private HTTP/1.1 pool on Node 26+, shared by SDK clients. No global dispatcher changes. */
|
|
14
39
|
let nodeFetch;
|
|
15
40
|
/**
|
|
16
41
|
* Reuses TLS connections. Node 26's bundled undici 8.9 can stall reused connections: use pinned undici with a
|
|
17
|
-
* private HTTP/1.1-only dispatcher
|
|
18
|
-
* for diagnosis; =1 retains the historical opt-in to
|
|
42
|
+
* private HTTP/1.1-only dispatcher whose idle connections stay reusable for KEEP_ALIVE_TIMEOUT_MS. Other runtimes
|
|
43
|
+
* retain native fetch. SHARDFLUX_HTTP_KEEPALIVE=0 forces close for diagnosis; =1 retains the historical opt-in to
|
|
44
|
+
* native pooling. An explicitly supplied fetch is untouched. Idle pooled sockets do not keep the process alive.
|
|
19
45
|
*/
|
|
20
46
|
export function defaultFetch(env = globalThis.process?.env, undiciVersion = globalThis.process?.versions?.undici) {
|
|
21
47
|
// Why a fresh connection on Node 26 (its bundled HTTP client, measured): docs/progress/startup-latency.md.
|
|
@@ -30,7 +56,7 @@ export function defaultFetch(env = globalThis.process?.env, undiciVersion = glob
|
|
|
30
56
|
return base;
|
|
31
57
|
return async (input, init) => {
|
|
32
58
|
nodeFetch ??= import('undici').then(({ Agent, fetch: pooledFetch }) => {
|
|
33
|
-
const dispatcher = new Agent({
|
|
59
|
+
const dispatcher = new Agent({ ...POOL_OPTIONS, connect: { ...POOL_OPTIONS.connect } });
|
|
34
60
|
return async (input, init) => {
|
|
35
61
|
const response = await pooledFetch(input, { ...init, dispatcher });
|
|
36
62
|
// Web-standard runtime shape; undici and DOM iterator declarations differ.
|
|
@@ -49,6 +75,25 @@ export function buildUrl(base, path, query) {
|
|
|
49
75
|
}
|
|
50
76
|
return u.toString();
|
|
51
77
|
}
|
|
78
|
+
/** Codes (on fetch's `cause` chain) of a connection that closed or reset before the response arrived. */
|
|
79
|
+
const STALE_CONNECTION_CODES = new Set(['UND_ERR_SOCKET', 'ECONNRESET', 'EPIPE', 'UND_ERR_CLOSED']);
|
|
80
|
+
/**
|
|
81
|
+
* True when fetch failed because its connection closed or reset before a response arrived: what a request sees when
|
|
82
|
+
* it went out on a pooled connection the other side closed while it sat idle. Connect failures (ECONNREFUSED,
|
|
83
|
+
* ENOTFOUND, UND_ERR_CONNECT_TIMEOUT), timeouts and aborts are not this.
|
|
84
|
+
*/
|
|
85
|
+
export function isStaleConnectionError(err) {
|
|
86
|
+
if (!(err instanceof TypeError))
|
|
87
|
+
return false;
|
|
88
|
+
const seen = new Set();
|
|
89
|
+
for (let c = err.cause; typeof c === 'object' && c !== null && !seen.has(c); c = c.cause) {
|
|
90
|
+
seen.add(c);
|
|
91
|
+
const code = c.code;
|
|
92
|
+
if (typeof code === 'string' && STALE_CONNECTION_CODES.has(code))
|
|
93
|
+
return true;
|
|
94
|
+
}
|
|
95
|
+
return false;
|
|
96
|
+
}
|
|
52
97
|
function retryAfterSeconds(res) {
|
|
53
98
|
const v = res.headers.get('retry-after');
|
|
54
99
|
if (v === null)
|
|
@@ -89,6 +134,9 @@ export class HttpClient {
|
|
|
89
134
|
const safe = method === 'GET' || method === 'HEAD';
|
|
90
135
|
const retriable = safe || init.idempotencyKey !== undefined || init.idempotent === true;
|
|
91
136
|
const sleep = this.opts.sleep ?? defaultSleep;
|
|
137
|
+
// Retries counted against maxRetries (they also set the backoff); `attempt` counts every send.
|
|
138
|
+
let counted = 0;
|
|
139
|
+
let staleReplayed = false;
|
|
92
140
|
for (let attempt = 0;; attempt += 1) {
|
|
93
141
|
const headers = {
|
|
94
142
|
accept: init.accept ?? 'application/json',
|
|
@@ -122,9 +170,20 @@ export class HttpClient {
|
|
|
122
170
|
catch (err) {
|
|
123
171
|
if (init.signal?.aborted)
|
|
124
172
|
throw err;
|
|
125
|
-
if (!retriable
|
|
173
|
+
if (!retriable)
|
|
174
|
+
throw err;
|
|
175
|
+
if (!staleReplayed && !signal?.aborted && isStaleConnectionError(err)) {
|
|
176
|
+
// The connection closed before any response, typically a pooled one the server or a NAT closed while idle:
|
|
177
|
+
// send again at once on a fresh connection, once per request and outside maxRetries (as Go's net/http
|
|
178
|
+
// replays a request that failed on a reused connection when it is idempotent or carries an Idempotency-Key).
|
|
179
|
+
staleReplayed = true;
|
|
180
|
+
init.onRetry?.({ request: `${method} ${path}`, attempt: attempt + 1, cause: describeFailure(err), delayMs: 0 });
|
|
181
|
+
continue;
|
|
182
|
+
}
|
|
183
|
+
if (counted >= this.opts.maxRetries)
|
|
126
184
|
throw err;
|
|
127
|
-
const delayMs = Math.min(2_000, 200 * 2 **
|
|
185
|
+
const delayMs = Math.min(2_000, 200 * 2 ** counted);
|
|
186
|
+
counted += 1;
|
|
128
187
|
init.onRetry?.({ request: `${method} ${path}`, attempt: attempt + 1, cause: describeFailure(err, headersAbort?.signal.aborted ? headersMs : timeoutMs), delayMs });
|
|
129
188
|
await sleep(delayMs);
|
|
130
189
|
continue;
|
|
@@ -146,10 +205,15 @@ export class HttpClient {
|
|
|
146
205
|
const error = await errorFrom(res, this.opts.source);
|
|
147
206
|
const retryableStatus = res.status === 429 || res.status === 502 || res.status === 503 || res.status === 504;
|
|
148
207
|
const retryableError = error instanceof ShardfluxApiError ? error.retryable && retryableStatus : retryableStatus;
|
|
149
|
-
|
|
208
|
+
// The working-at-once quota (cell 429 quota_exceeded): the gateway refuses the call at admission, before the
|
|
209
|
+
// workspace is touched, so nothing ran and even a non-idempotent request (exec start, stdin, PTY create,
|
|
210
|
+
// keepalive) is safe to send again. Same Retry-After wait and retry budget as every other retry.
|
|
211
|
+
const admissionRefusal = init.quotaRetry !== false && isWorkingQuotaRefusal(error);
|
|
212
|
+
if ((!retriable && !admissionRefusal) || !retryableError || counted >= this.opts.maxRetries)
|
|
150
213
|
throw error;
|
|
151
214
|
const hinted = error instanceof ShardfluxApiError ? error.retryAfterSeconds : undefined;
|
|
152
|
-
const delayMs = Math.min(5_000, hinted !== undefined ? hinted * 1000 : 200 * 2 **
|
|
215
|
+
const delayMs = Math.min(5_000, hinted !== undefined ? hinted * 1000 : 200 * 2 ** counted);
|
|
216
|
+
counted += 1;
|
|
153
217
|
init.onRetry?.({ request: `${method} ${path}`, attempt: attempt + 1, cause: describeFailure(error), delayMs });
|
|
154
218
|
await sleep(delayMs);
|
|
155
219
|
}
|
package/dist/index.d.ts
CHANGED
|
@@ -17,8 +17,8 @@ export type { components as CellComponents, paths as CellPaths } from './generat
|
|
|
17
17
|
export { BillingApi, Shardflux, WorkspacesApi, fetchBillingCatalog, pickByKey } from './client.js';
|
|
18
18
|
export type { AgentSession, AllocationMode, BillingCatalog, BillingSubscription, Caps, CheckoutSession, DiskLayout, Entitlements, FindByKeyOptions, ForkTarget, Invoice, InvoicePage, LifetimeFilter, ListParams, Me, OpenParams, OpenResponse, Operation, Page, PortalSession, PurposeFilter, ResetWorkspaceBody, ResumeAnswer, ResumeRequestOptions, ResumeResponse, ShardfluxOptions, SuspendRequest, SuspendWhenIdleOptions, SuspendWhenIdleResponse, SuspendWhenIdleResult, WaitOptions, WorkspaceInputs, WorkspaceLifetime, WorkspaceMemory, WorkspaceOrigin, WorkspacePurpose, WorkspaceView, } from './client.js';
|
|
19
19
|
export type { FinishedOperation, ForkOptions, LifecycleOptions, ResumeOptions, SuspendOptions, WaitedForkOptions, WaitedLifecycleOptions, WaitedResumeOptions, WaitedSuspendOptions } from './lifecycle.js';
|
|
20
|
-
export { durabilityOf, formatTiming, isDurable, lostSuspendOf } from './progress.js';
|
|
21
|
-
export type { Durability, LostSuspend, LifecycleAction, LifecyclePhase, LifecycleTiming, ProgressEvent, ProgressListener, RetryRecord, ServerTiming, TimingOutcome, TimingPhase, } from './progress.js';
|
|
20
|
+
export { durabilityOf, formatTiming, hostLostOf, isDurable, lostSuspendOf } from './progress.js';
|
|
21
|
+
export type { ColdBootReason, Durability, HostLost, LostSuspend, LifecycleAction, LifecyclePhase, LifecycleTiming, ProgressEvent, ProgressListener, RetryRecord, ServerTiming, TimingOutcome, TimingPhase, } from './progress.js';
|
|
22
22
|
export { FEEDBACK_CATEGORIES, FEEDBACK_MESSAGE_MAX_LENGTH } from './feedback.js';
|
|
23
23
|
export type { AccountFeedbackParams, FeedbackCategory, FeedbackContext, FeedbackReceipt, SendFeedbackParams } from './feedback.js';
|
|
24
24
|
export { UsageApi } from './usage.js';
|
package/dist/index.js
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
export { BillingApi, Shardflux, WorkspacesApi, fetchBillingCatalog, pickByKey } from "./client.js";
|
|
2
|
-
export { durabilityOf, formatTiming, isDurable, lostSuspendOf } from "./progress.js";
|
|
2
|
+
export { durabilityOf, formatTiming, hostLostOf, isDurable, lostSuspendOf } from "./progress.js";
|
|
3
3
|
export { FEEDBACK_CATEGORIES, FEEDBACK_MESSAGE_MAX_LENGTH } from "./feedback.js";
|
|
4
4
|
export { UsageApi } from "./usage.js";
|
|
5
5
|
export { TemplateBuildTimeoutError, TemplateBuildsApi, TemplateDraftApi, TemplatePackagesApi, TemplateUploadError, TemplateUploadsApi, TemplateVersionTestInstancesApi, TemplateVersionsApi, TemplatesApi, buildSettled, saveAsTemplateBody, } from "./templates.js";
|
package/dist/progress.d.ts
CHANGED
|
@@ -74,9 +74,10 @@ export interface ServerTiming {
|
|
|
74
74
|
warmFallback: string | null;
|
|
75
75
|
/**
|
|
76
76
|
* `result.resume_path`: `local_cache`, `prestaged` (copied to this host ahead of the resume),
|
|
77
|
-
* `download` (the checkpoint had to be fetched first), `cold_boot` (0.11.0+: the saved disk was booted
|
|
78
|
-
*
|
|
79
|
-
* reset).
|
|
77
|
+
* `download` (the checkpoint had to be fetched first), `cold_boot` (0.11.0+: the saved disk was booted instead of
|
|
78
|
+
* restoring memory, see `coldBootReason`; processes restarted; see `memoryRestored`), `reset_blank_layer` (the first
|
|
79
|
+
* start after a reset) or `thaw` (0.13.1+: resumed while its suspend was still being written; the same VM continued
|
|
80
|
+
* in place).
|
|
80
81
|
*/
|
|
81
82
|
resumePath: string | null;
|
|
82
83
|
/**
|
|
@@ -87,10 +88,11 @@ export interface ServerTiming {
|
|
|
87
88
|
*/
|
|
88
89
|
memoryRestored?: boolean | null;
|
|
89
90
|
/**
|
|
90
|
-
* `result.cold_boot_reason` (0.11.0+), with `resumePath` `cold_boot`: why the memory could not be restored
|
|
91
|
-
* `runtime_changed` (the platform's VM runtime changed after the suspend).
|
|
91
|
+
* `result.cold_boot_reason` (0.11.0+), with `resumePath` `cold_boot`: why the memory could not be restored:
|
|
92
|
+
* `runtime_changed` (the platform's VM runtime changed after the suspend) or `host_lost` (0.13.1+: the machine the
|
|
93
|
+
* workspace ran on failed; it booted from its disk, files kept, see `hostLost`). Null otherwise.
|
|
92
94
|
*/
|
|
93
|
-
coldBootReason?:
|
|
95
|
+
coldBootReason?: ColdBootReason | null;
|
|
94
96
|
/**
|
|
95
97
|
* `result.durable` (0.12.0+), suspend and fork: `true` when the capture is in durable storage, `false` while its
|
|
96
98
|
* durable copy is being written (see `durability`). `null` when the result does not say (other kinds, an older API).
|
|
@@ -105,6 +107,12 @@ export interface ServerTiming {
|
|
|
105
107
|
* checkpoint before it (`restoredCheckpointId`, state as of `stateAsOf`). Null otherwise.
|
|
106
108
|
*/
|
|
107
109
|
lostSuspend?: LostSuspend | null;
|
|
110
|
+
/**
|
|
111
|
+
* `result.host_lost` (0.13.1+): the machine the workspace ran on failed. A resume (or open) that restored the
|
|
112
|
+
* workspace says from what (`restoredFrom` `disk`, or `checkpoint` with `stateAsOf`); a suspend that found it so
|
|
113
|
+
* succeeds with only `detectedAt`. Null otherwise.
|
|
114
|
+
*/
|
|
115
|
+
hostLost?: HostLost | null;
|
|
108
116
|
/** `result.boot_to_ready_ms`: VM start until the guest agent answered. */
|
|
109
117
|
bootToReadyMs: number | null;
|
|
110
118
|
/** `result.host_timings_ms`: the host's own steps (restore: load, after_restore, ready, …). */
|
|
@@ -143,6 +151,30 @@ export interface LostSuspend {
|
|
|
143
151
|
/** The restored state is as of this time (RFC 3339). */
|
|
144
152
|
stateAsOf: string | null;
|
|
145
153
|
}
|
|
154
|
+
/**
|
|
155
|
+
* `result.cold_boot_reason` (0.11.0+): `runtime_changed` (the platform's VM runtime changed after the suspend) or
|
|
156
|
+
* `host_lost` (0.13.1+: the machine the workspace ran on failed). Any other string is a reason this version does not
|
|
157
|
+
* know.
|
|
158
|
+
*/
|
|
159
|
+
export type ColdBootReason = 'runtime_changed' | 'host_lost' | (string & {});
|
|
160
|
+
/**
|
|
161
|
+
* `result.host_lost` (0.13.1+), camelCased: the machine the workspace ran on failed. The workspace was moved to
|
|
162
|
+
* `suspended` at that moment, and its next use (a resume, a tool call's wake, an open) restored it. See the lifecycle
|
|
163
|
+
* reference.
|
|
164
|
+
*/
|
|
165
|
+
export interface HostLost {
|
|
166
|
+
/** When the failure was detected (RFC 3339). */
|
|
167
|
+
detectedAt: string | null;
|
|
168
|
+
/**
|
|
169
|
+
* What the resume restored: `disk` (the workspace's own disk, files kept; `coldBootReason` `host_lost`, processes
|
|
170
|
+
* restarted) or `checkpoint` (its newest checkpoint, see `stateAsOf`). Null on a suspend's result.
|
|
171
|
+
*/
|
|
172
|
+
restoredFrom: 'disk' | 'checkpoint' | null;
|
|
173
|
+
/** `checkpoint`: the checkpoint the resume restored. */
|
|
174
|
+
restoredCheckpointId: string | null;
|
|
175
|
+
/** `checkpoint`: the restored state is as of this time (RFC 3339); changes after it are not in the workspace. */
|
|
176
|
+
stateAsOf: string | null;
|
|
177
|
+
}
|
|
146
178
|
/**
|
|
147
179
|
* The durable copy of a suspend or fork operation (0.12.0+): `result.durability` camelCased, or null when the result has
|
|
148
180
|
* none (the suspend was stored durably before it completed, or another kind).
|
|
@@ -150,6 +182,11 @@ export interface LostSuspend {
|
|
|
150
182
|
export declare function durabilityOf(op: Pick<Operation, 'result'>): Durability | null;
|
|
151
183
|
/** A resume operation's `result.lost_suspend` camelCased (0.12.0+), or null. */
|
|
152
184
|
export declare function lostSuspendOf(op: Pick<Operation, 'result'>): LostSuspend | null;
|
|
185
|
+
/**
|
|
186
|
+
* An operation's `result.host_lost` camelCased (0.13.1+), or null: on a resume or open that restored a workspace whose
|
|
187
|
+
* machine failed, and on a suspend that found it so (`detectedAt` only).
|
|
188
|
+
*/
|
|
189
|
+
export declare function hostLostOf(op: Pick<Operation, 'result'>): HostLost | null;
|
|
153
190
|
/**
|
|
154
191
|
* Whether a suspend or fork operation's capture is in durable storage (0.12.0+): `result.durable`, else `true` for a
|
|
155
192
|
* succeeded suspend or fork whose result predates the field, else null.
|
|
@@ -255,7 +292,10 @@ export declare function traced<T>(trace: Trace, fn: () => Promise<T>): Promise<T
|
|
|
255
292
|
*
|
|
256
293
|
* A call that retried a request adds `retries: <n> (<request>, <cause>, after <delay>)`.
|
|
257
294
|
* A resume that booted the workspace instead of restoring its memory (0.11.0+) says so in the server line:
|
|
258
|
-
* `resume from cold_boot: processes restarted (runtime_changed)
|
|
295
|
+
* `resume from cold_boot: processes restarted (runtime_changed)`, or `(host_lost)` (0.13.1+) when the machine the
|
|
296
|
+
* workspace ran on failed and it booted from its disk. A resume that restored the newest checkpoint after such a
|
|
297
|
+
* failure adds `restored checkpoint <id> (host_lost; state as of <time>)`; a suspend that found the machine failed
|
|
298
|
+
* adds `host_lost (detected <time>)`.
|
|
259
299
|
*/
|
|
260
300
|
export declare function formatTiming(t: LifecycleTiming): string;
|
|
261
301
|
export {};
|
package/dist/progress.js
CHANGED
|
@@ -40,6 +40,22 @@ export function lostSuspendOf(op) {
|
|
|
40
40
|
stateAsOf: str(r.state_as_of),
|
|
41
41
|
};
|
|
42
42
|
}
|
|
43
|
+
/**
|
|
44
|
+
* An operation's `result.host_lost` camelCased (0.13.1+), or null: on a resume or open that restored a workspace whose
|
|
45
|
+
* machine failed, and on a suspend that found it so (`detectedAt` only).
|
|
46
|
+
*/
|
|
47
|
+
export function hostLostOf(op) {
|
|
48
|
+
const h = op.result?.host_lost;
|
|
49
|
+
if (typeof h !== 'object' || h === null || Array.isArray(h))
|
|
50
|
+
return null;
|
|
51
|
+
const r = h;
|
|
52
|
+
return {
|
|
53
|
+
detectedAt: str(r.detected_at),
|
|
54
|
+
restoredFrom: r.restored_from === 'disk' || r.restored_from === 'checkpoint' ? r.restored_from : null,
|
|
55
|
+
restoredCheckpointId: str(r.restored_checkpoint_id),
|
|
56
|
+
stateAsOf: str(r.state_as_of),
|
|
57
|
+
};
|
|
58
|
+
}
|
|
43
59
|
/**
|
|
44
60
|
* Whether a suspend or fork operation's capture is in durable storage (0.12.0+): `result.durable`, else `true` for a
|
|
45
61
|
* succeeded suspend or fork whose result predates the field, else null.
|
|
@@ -122,6 +138,7 @@ export function serverTiming(op) {
|
|
|
122
138
|
suspendPath: str(r.suspend_path),
|
|
123
139
|
durability: durabilityOf(op),
|
|
124
140
|
lostSuspend: lostSuspendOf(op),
|
|
141
|
+
hostLost: hostLostOf(op),
|
|
125
142
|
bootToReadyMs: num(r.boot_to_ready_ms),
|
|
126
143
|
hostTimingsMs,
|
|
127
144
|
};
|
|
@@ -275,6 +292,19 @@ export async function traced(trace, fn) {
|
|
|
275
292
|
}
|
|
276
293
|
}
|
|
277
294
|
const fmt = (ms) => (ms === null ? '?' : ms < 1000 ? `${Math.round(ms)} ms` : `${(ms / 1000).toFixed(2)} s`);
|
|
295
|
+
/**
|
|
296
|
+
* The server-line part of `result.host_lost` (0.13.1+): the checkpoint a resume restored, or a suspend's detection time.
|
|
297
|
+
* A resume from the disk says nothing more here: its cold boot reads `processes restarted (host_lost)`.
|
|
298
|
+
*/
|
|
299
|
+
function hostLostText(h) {
|
|
300
|
+
if (!h)
|
|
301
|
+
return null;
|
|
302
|
+
if (h.restoredFrom === 'checkpoint')
|
|
303
|
+
return `restored checkpoint ${h.restoredCheckpointId ?? '?'} (host_lost; state as of ${h.stateAsOf ?? '?'})`;
|
|
304
|
+
if (h.restoredFrom === null)
|
|
305
|
+
return `host_lost${h.detectedAt ? ` (detected ${h.detectedAt})` : ''}`;
|
|
306
|
+
return null;
|
|
307
|
+
}
|
|
278
308
|
/**
|
|
279
309
|
* A human-readable account of a timing, for logs and bug reports. A resume in production:
|
|
280
310
|
*
|
|
@@ -285,7 +315,10 @@ const fmt = (ms) => (ms === null ? '?' : ms < 1000 ? `${Math.round(ms)} ms` : `$
|
|
|
285
315
|
*
|
|
286
316
|
* A call that retried a request adds `retries: <n> (<request>, <cause>, after <delay>)`.
|
|
287
317
|
* A resume that booted the workspace instead of restoring its memory (0.11.0+) says so in the server line:
|
|
288
|
-
* `resume from cold_boot: processes restarted (runtime_changed)
|
|
318
|
+
* `resume from cold_boot: processes restarted (runtime_changed)`, or `(host_lost)` (0.13.1+) when the machine the
|
|
319
|
+
* workspace ran on failed and it booted from its disk. A resume that restored the newest checkpoint after such a
|
|
320
|
+
* failure adds `restored checkpoint <id> (host_lost; state as of <time>)`; a suspend that found the machine failed
|
|
321
|
+
* adds `host_lost (detected <time>)`.
|
|
289
322
|
*/
|
|
290
323
|
export function formatTiming(t) {
|
|
291
324
|
const ids = [t.workspaceId ? `workspace ${t.workspaceId}` : null, t.operationId ? `operation ${t.operationId}` : null].filter(Boolean).join(', ');
|
|
@@ -308,6 +341,7 @@ export function formatTiming(t) {
|
|
|
308
341
|
s.startPath ? `start ${s.startPath}${s.warmFallback ? ` (warm fallback: ${s.warmFallback})` : ''}` : null,
|
|
309
342
|
s.resumePath ? `resume from ${s.resumePath}${s.memoryRestored === false ? `: processes restarted${s.coldBootReason ? ` (${s.coldBootReason})` : ''}` : ''}` : null,
|
|
310
343
|
s.lostSuspend ? `restored ${s.lostSuspend.restoredCheckpointId ?? 'no checkpoint'} (latest suspend ${s.lostSuspend.checkpointId} ${s.lostSuspend.reason ?? 'lost'})` : null,
|
|
344
|
+
hostLostText(s.hostLost),
|
|
311
345
|
s.suspendPath === 'local_commit' ? 'sealed on host' : null,
|
|
312
346
|
s.durability?.state === 'durable' ? `durable${s.durability.localCommitToDurableMs !== null ? ` ${fmt(s.durability.localCommitToDurableMs)} later` : ''}` : s.durability ? `durable copy ${s.durability.state}` : null,
|
|
313
347
|
s.bootToReadyMs !== null ? `boot to ready ${fmt(s.bootToReadyMs)}` : null,
|
package/dist/usage.d.ts
CHANGED
|
@@ -56,7 +56,9 @@ export declare class UsageApi {
|
|
|
56
56
|
constructor(ctx: () => ClientContext);
|
|
57
57
|
/**
|
|
58
58
|
* Current-period usage per meter, allowances with enforcement and cap state (`overage` while opt-in overage covers
|
|
59
|
-
* usage past a CPU-hours or RAM GiB-hours allowance
|
|
59
|
+
* usage past a CPU-hours or RAM GiB-hours allowance; `storage_blocked` (0.14.0+, enforcement `storage_block`) while
|
|
60
|
+
* the Retained state allowance is used up: opening a new key and forking are refused with 403 `quota_exceeded`,
|
|
61
|
+
* details.limit `retained_state`), measurement freshness, `exhausted_reason` (the 402
|
|
60
62
|
* allowance_exhausted reason while starts are refused: `allowance_used`, `overage_paused`, `spend_cap_reached`) and
|
|
61
63
|
* `spend_cap` (0.10.0).
|
|
62
64
|
*/
|
package/dist/usage.js
CHANGED
|
@@ -9,7 +9,9 @@ export class UsageApi {
|
|
|
9
9
|
}
|
|
10
10
|
/**
|
|
11
11
|
* Current-period usage per meter, allowances with enforcement and cap state (`overage` while opt-in overage covers
|
|
12
|
-
* usage past a CPU-hours or RAM GiB-hours allowance
|
|
12
|
+
* usage past a CPU-hours or RAM GiB-hours allowance; `storage_blocked` (0.14.0+, enforcement `storage_block`) while
|
|
13
|
+
* the Retained state allowance is used up: opening a new key and forking are refused with 403 `quota_exceeded`,
|
|
14
|
+
* details.limit `retained_state`), measurement freshness, `exhausted_reason` (the 402
|
|
13
15
|
* allowance_exhausted reason while starts are refused: `allowance_used`, `overage_paused`, `spend_cap_reached`) and
|
|
14
16
|
* `spend_cap` (0.10.0).
|
|
15
17
|
*/
|
package/dist/workspace.d.ts
CHANGED
|
@@ -67,7 +67,9 @@ export declare class Workspace {
|
|
|
67
67
|
* woke the workspace), or a lifecycle call with `wait`. Null for handles from get()/list() until such a call.
|
|
68
68
|
* `formatTiming(workspace.lastTiming)` prints it. After a resume or wake, `server.memoryRestored === false` (0.11.0+)
|
|
69
69
|
* means the workspace booted from its saved disk instead of restoring its memory (`server.resumePath` `cold_boot`,
|
|
70
|
-
* reason in `server.coldBootReason`): files are as of the suspend, running processes were restarted.
|
|
70
|
+
* reason in `server.coldBootReason`): files are as of the suspend, running processes were restarted. Reason
|
|
71
|
+
* `host_lost` (0.13.1+): the machine the workspace ran on failed and it booted from its disk, files kept;
|
|
72
|
+
* `server.hostLost` says what the resume restored (`hostLostOf(op)` on the operation).
|
|
71
73
|
*/
|
|
72
74
|
get lastTiming(): LifecycleTiming | null;
|
|
73
75
|
get id(): string;
|
|
@@ -174,6 +176,8 @@ export declare class Workspace {
|
|
|
174
176
|
* resolve once it has FINISHED, with `workspace.state` then `suspended`. That is as soon as the workspace is sealed on
|
|
175
177
|
* its host, typically in a few hundred ms. `result.durable` (also `lastTiming.server.durable`, 0.12.0+) turns true
|
|
176
178
|
* when the copy lands in durable storage, typically within a second; `{ durable: true }` resolves only then.
|
|
179
|
+
* A suspend that finds the machine the workspace ran on failed succeeds at once (0.13.1+): `result.durable` is true
|
|
180
|
+
* and `result.host_lost` (`hostLostOf(op)`) has `detectedAt`; the next resume restores the workspace.
|
|
177
181
|
*
|
|
178
182
|
* await workspace.suspend({ wait: true });
|
|
179
183
|
* await workspace.suspend({ durable: true }); // 0.12.0+: also wait for the durable copy
|
|
@@ -205,11 +209,16 @@ export declare class Workspace {
|
|
|
205
209
|
* (also `lastTiming.server.memoryRestored`, 0.11.0+) is false when the resume booted the saved disk instead
|
|
206
210
|
* (`resume_path` `cold_boot`): files kept, processes restarted. `result.lost_suspend` (also
|
|
207
211
|
* `lastTiming.server.lostSuspend`, `lostSuspendOf(op)`, 0.12.0+) names a suspend this resume could not restore and the
|
|
208
|
-
* checkpoint it restored instead (see the lifecycle reference).
|
|
212
|
+
* checkpoint it restored instead (see the lifecycle reference). `result.host_lost` (also `lastTiming.server.hostLost`,
|
|
213
|
+
* `hostLostOf(op)`, 0.13.1+): the machine the workspace ran on failed and this resume restored it, from its disk
|
|
214
|
+
* (`cold_boot_reason` `host_lost`) or from its newest checkpoint (`state_as_of`).
|
|
209
215
|
*/
|
|
210
216
|
resume(opts: WaitedResumeOptions): Promise<FinishedOperation>;
|
|
211
217
|
resume(opts?: ResumeOptions): Promise<Operation>;
|
|
212
|
-
/**
|
|
218
|
+
/**
|
|
219
|
+
* Takes a snapshot. Resolves when it is REQUESTED; with `{ wait: true }`, once it is taken. After the machine the
|
|
220
|
+
* workspace ran on failed, resume it first: until then the snapshot fails `resume_required` (0.13.1+).
|
|
221
|
+
*/
|
|
213
222
|
snapshot(opts: WaitedLifecycleOptions & {
|
|
214
223
|
label?: string;
|
|
215
224
|
}): Promise<FinishedOperation>;
|
|
@@ -218,7 +227,8 @@ export declare class Workspace {
|
|
|
218
227
|
}): Promise<Operation>;
|
|
219
228
|
/**
|
|
220
229
|
* Forks into a new key. Resolves when the fork is REQUESTED (the copy's handle is returned at once); with
|
|
221
|
-
* `{ wait: true }`, once the copy exists, with its handle refreshed.
|
|
230
|
+
* `{ wait: true }`, once the copy exists, with its handle refreshed. After the machine the workspace ran on failed,
|
|
231
|
+
* resume it first: until then the fork fails `resume_required` (0.13.1+).
|
|
222
232
|
*/
|
|
223
233
|
fork(target: ForkTarget, opts: WaitedForkOptions): Promise<{
|
|
224
234
|
operation: FinishedOperation;
|
|
@@ -301,8 +311,9 @@ export declare class Workspace {
|
|
|
301
311
|
* that does not hold the request answers at once; the wake then waits for the operation and reads the view.
|
|
302
312
|
*
|
|
303
313
|
* A wake whose resume could not restore the workspace's memory (0.11.0+; the platform's VM runtime changed after the
|
|
304
|
-
* suspend) still resolves `true`: the workspace runs from its
|
|
305
|
-
* and the `done` progress event say so
|
|
314
|
+
* suspend, or, 0.13.1+, the machine the workspace ran on failed) still resolves `true`: the workspace runs from its
|
|
315
|
+
* saved disk, with every process restarted. `lastTiming` and the `done` progress event say so
|
|
316
|
+
* (`server.memoryRestored === false`, `server.resumePath` `cold_boot`, `server.coldBootReason`).
|
|
306
317
|
*/
|
|
307
318
|
wake(opts?: WakeOptions): Promise<boolean>;
|
|
308
319
|
/**
|
package/dist/workspace.js
CHANGED
|
@@ -44,7 +44,9 @@ export class Workspace {
|
|
|
44
44
|
* woke the workspace), or a lifecycle call with `wait`. Null for handles from get()/list() until such a call.
|
|
45
45
|
* `formatTiming(workspace.lastTiming)` prints it. After a resume or wake, `server.memoryRestored === false` (0.11.0+)
|
|
46
46
|
* means the workspace booted from its saved disk instead of restoring its memory (`server.resumePath` `cold_boot`,
|
|
47
|
-
* reason in `server.coldBootReason`): files are as of the suspend, running processes were restarted.
|
|
47
|
+
* reason in `server.coldBootReason`): files are as of the suspend, running processes were restarted. Reason
|
|
48
|
+
* `host_lost` (0.13.1+): the machine the workspace ran on failed and it booted from its disk, files kept;
|
|
49
|
+
* `server.hostLost` says what the resume restored (`hostLostOf(op)` on the operation).
|
|
48
50
|
*/
|
|
49
51
|
get lastTiming() {
|
|
50
52
|
return this.#lastTiming ?? this.#openTrace?.finished ?? null;
|
|
@@ -438,8 +440,9 @@ export class Workspace {
|
|
|
438
440
|
* that does not hold the request answers at once; the wake then waits for the operation and reads the view.
|
|
439
441
|
*
|
|
440
442
|
* A wake whose resume could not restore the workspace's memory (0.11.0+; the platform's VM runtime changed after the
|
|
441
|
-
* suspend) still resolves `true`: the workspace runs from its
|
|
442
|
-
* and the `done` progress event say so
|
|
443
|
+
* suspend, or, 0.13.1+, the machine the workspace ran on failed) still resolves `true`: the workspace runs from its
|
|
444
|
+
* saved disk, with every process restarted. `lastTiming` and the `done` progress event say so
|
|
445
|
+
* (`server.memoryRestored === false`, `server.resumePath` `cold_boot`, `server.coldBootReason`).
|
|
443
446
|
*/
|
|
444
447
|
async wake(opts = {}) {
|
|
445
448
|
// A file-first workspace runs from creation and is never suspended: nothing to wake.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@shardflux/sdk",
|
|
3
|
-
"version": "0.13.
|
|
3
|
+
"version": "0.13.1",
|
|
4
4
|
"type": "module",
|
|
5
5
|
"description": "Shardflux TypeScript SDK: open persistent agent workspaces by key and give your agent workspace tools (exec, files, processes, PTY, git, browser).",
|
|
6
6
|
"license": "Apache-2.0",
|