@shardflux/sdk 0.6.1 → 0.6.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +17 -1
- package/README.md +26 -6
- package/dist/client.d.ts +8 -2
- package/dist/client.js +3 -1
- package/dist/errors.d.ts +13 -1
- package/dist/errors.js +19 -2
- package/dist/generated/app-api.d.ts +2673 -181
- package/dist/http.d.ts +1 -1
- package/dist/http.js +1 -1
- package/dist/lifecycle.d.ts +2 -1
- package/dist/progress.d.ts +14 -2
- package/dist/progress.js +18 -4
- package/dist/workspace.d.ts +3 -1
- package/dist/workspace.js +3 -1
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -3,7 +3,23 @@
|
|
|
3
3
|
Every API the README shows is available from the version named here. Below 1.0, a minor release may break
|
|
4
4
|
compatibility; breaking changes are marked **Breaking**.
|
|
5
5
|
|
|
6
|
-
## 0.6.
|
|
6
|
+
## 0.6.2 (not yet published; npm `latest` is 0.6.1)
|
|
7
|
+
|
|
8
|
+
### Starts that wait for capacity end
|
|
9
|
+
|
|
10
|
+
The API no longer lets a start (open, resume, restore, fork) wait in `capacity_pending` forever. One that no host
|
|
11
|
+
could admit 15 minutes after it was created fails with `capacity_unavailable` and `retryable: true`: nothing was
|
|
12
|
+
started, the concurrency slot is released, and a suspended workspace stays suspended with its state. Before, a VM could
|
|
13
|
+
boot (and bill) long after every wait had given up.
|
|
14
|
+
|
|
15
|
+
- `OperationFailedError.retryable`: the operation error's `retryable` flag (`false` when absent). `true` for
|
|
16
|
+
`capacity_unavailable` (retry later), `false` for a definitive failure. The SDK does not retry it itself.
|
|
17
|
+
- `OperationTimeoutError.deadlineAt`: when a wait gives up while the start is still `capacity_pending`, the time it
|
|
18
|
+
stops waiting for a host (`error.details.deadline_at`); the message says so. Null in other states.
|
|
19
|
+
- `onProgress`: a `capacity_pending` `phase` event carries `deadlineAt` when the API reports it.
|
|
20
|
+
- Docs: waits no longer say a pending start continues indefinitely.
|
|
21
|
+
|
|
22
|
+
## 0.6.1 (2026-09-28)
|
|
7
23
|
|
|
8
24
|
- `formatTiming()` joined two phases that followed each other with `∥` (ran together) when their 0.1 ms times summed
|
|
9
25
|
with a floating-point error (1000.2 + 300.1 > 1300.3); it now prints `→`.
|
package/README.md
CHANGED
|
@@ -9,7 +9,7 @@ tools (exec, files, processes, PTY, git, browser) that plug into any model provi
|
|
|
9
9
|
> **Early access.** Shardflux is in early access. The API is versioned (`/v1`), but this SDK is
|
|
10
10
|
> below 1.0: a minor release may contain breaking changes (see [Compatibility](#compatibility)).
|
|
11
11
|
|
|
12
|
-
> **Versions.** This README describes 0.6.
|
|
12
|
+
> **Versions.** This README describes 0.6.2. Anything marked **(0.6.0+)** is not in 0.5.0;
|
|
13
13
|
> [CHANGELOG.md](./CHANGELOG.md) lists what each version added. Check yours with
|
|
14
14
|
> `npm ls @shardflux/sdk` or the exported `SDK_VERSION`.
|
|
15
15
|
|
|
@@ -127,6 +127,23 @@ means depends on `wait`:
|
|
|
127
127
|
wait again with `cloud.workspaces.waitForOperation(err.operationId)`. Without `wait`, the returned operation is the
|
|
128
128
|
handle for the work in progress: pass its `id` to `waitForOperation()` when you need it finished.
|
|
129
129
|
|
|
130
|
+
**A start waits for capacity for at most 15 minutes.** An open, resume, restore or fork that no host can admit yet
|
|
131
|
+
waits in `capacity_pending`. Its `error.details.deadline_at` says when it gives up; the `phase` progress event carries
|
|
132
|
+
it as `deadlineAt` **(0.6.2+)**, and so does `OperationTimeoutError` when your wait ends first. A start still pending
|
|
133
|
+
at the deadline fails with `capacity_unavailable`: nothing was started, the concurrency slot is released, and a
|
|
134
|
+
suspended workspace stays suspended with its state. `OperationFailedError.retryable` **(0.6.2+)** is `true` for it,
|
|
135
|
+
so you can tell "retry later" from a definitive failure. The SDK does not retry it for you.
|
|
136
|
+
|
|
137
|
+
```ts
|
|
138
|
+
try {
|
|
139
|
+
await workspace.resume({ wait: true });
|
|
140
|
+
} catch (err) {
|
|
141
|
+
if (err instanceof OperationFailedError && err.retryable) {
|
|
142
|
+
// err.errorCode === 'capacity_unavailable': no host had room; nothing changed. Try again later.
|
|
143
|
+
} else throw err;
|
|
144
|
+
}
|
|
145
|
+
```
|
|
146
|
+
|
|
130
147
|
**Suspended workspaces wake on use.** A tool call on a suspended workspace resumes it (or joins the resume or open
|
|
131
148
|
already running), then runs. A call made during a suspend or resume waits for the transition to finish. The call
|
|
132
149
|
never runs twice: the cell executes nothing it refused.
|
|
@@ -148,9 +165,9 @@ await cell.exec.run(['make', 'test']); // resumes the wo
|
|
|
148
165
|
the response until the operation changes (`Prefer: wait`, at most 20 s per request), so completion arrives within
|
|
149
166
|
one round trip of the commit. Against an API without bounded waits (or with `serverWait: false`) it polls with
|
|
150
167
|
backoff (250 ms doubling to 5 s, ±20 % jitter). If the timeout passes, it throws `OperationTimeoutError` and the
|
|
151
|
-
operation keeps running server side; wait for it again with the
|
|
152
|
-
the same way. A waited `open()` issues the first tool token
|
|
153
|
-
tool call starts at once.
|
|
168
|
+
operation keeps running server side (a start waiting for capacity until its `deadlineAt`); wait for it again with the
|
|
169
|
+
same call. `templates.builds.waitForBuild()` waits the same way. A waited `open()` issues the first tool token
|
|
170
|
+
together with the final workspace read, so the first tool call starts at once.
|
|
154
171
|
|
|
155
172
|
On Node 26 the default fetch sends `Connection: close`: its bundled undici 8 can stall a request on a reused
|
|
156
173
|
keep-alive connection for tens of seconds. Pass your own `fetch`, or set `SHARDFLUX_HTTP_KEEPALIVE=1`, to change that.
|
|
@@ -338,8 +355,11 @@ await workspace.cell().exec.run(['python3', 'agent.py']); // sees $OPEN
|
|
|
338
355
|
`message`, `requestId`, `retryable`, `details`, `operationId`, `retryAfterSeconds`, and `reason`
|
|
339
356
|
(`details.reason`, e.g. `not_session`, `draft_not_found`, `legacy_disk_layout`; see `KnownErrorReason`).
|
|
340
357
|
- `OperationFailedError`: an awaited operation ended `failed` or `canceled` (`errorCode`,
|
|
341
|
-
`operation`, `timing`).
|
|
342
|
-
|
|
358
|
+
`retryable` **(0.6.2+)**, `operation`, `timing`). `retryable` is the operation error's own flag: `true` for
|
|
359
|
+
`capacity_unavailable` (no host could admit the start before its deadline; retry later), `false` for a definitive
|
|
360
|
+
failure.
|
|
361
|
+
- `OperationTimeoutError`: waiting gave up; the operation continues (`operationId`, `lastState`, `lastReason`,
|
|
362
|
+
`deadlineAt` **(0.6.2+)** while it waits for capacity, `timing`).
|
|
343
363
|
- `ShardfluxProtocolError`: a response was not the documented shape.
|
|
344
364
|
|
|
345
365
|
Treat unknown error codes and reasons as generic errors: show `message`, and use `retryable`.
|
package/dist/client.d.ts
CHANGED
|
@@ -92,7 +92,11 @@ export interface ForkTarget {
|
|
|
92
92
|
lifetime?: WorkspaceLifetime;
|
|
93
93
|
}
|
|
94
94
|
export interface WaitOptions {
|
|
95
|
-
/**
|
|
95
|
+
/**
|
|
96
|
+
* Give up waiting after this long (default 300 000 ms); the operation continues server side. A start waiting for
|
|
97
|
+
* capacity (`capacity_pending`) does so until its deadline (`error.details.deadline_at`, 15 minutes after it was
|
|
98
|
+
* created) and then fails with `capacity_unavailable` (retryable; nothing was started).
|
|
99
|
+
*/
|
|
96
100
|
timeoutMs?: number;
|
|
97
101
|
/** First poll delay (default 250 ms); doubles up to `maxPollIntervalMs` with jitter. Used when the server does not wait. */
|
|
98
102
|
pollIntervalMs?: number;
|
|
@@ -200,7 +204,9 @@ export declare class WorkspacesApi {
|
|
|
200
204
|
* that does not wait is polled with bounded exponential backoff (+-20 % jitter). Resolves when the operation
|
|
201
205
|
* succeeds; throws OperationFailedError when it fails or is canceled, OperationTimeoutError after `timeoutMs` (the
|
|
202
206
|
* operation keeps running and can be awaited again). Both errors carry the wait's `timing`; `onProgress` sees each
|
|
203
|
-
* state change (queued, capacity_pending, running, with the server's reason) as it is observed.
|
|
207
|
+
* state change (queued, capacity_pending, running, with the server's reason) as it is observed. A start waiting in
|
|
208
|
+
* capacity_pending gives up at its deadline (`deadlineAt` on the phase event) and fails with `capacity_unavailable`
|
|
209
|
+
* (`err.retryable` true: nothing was started, retry later); the SDK does not retry it.
|
|
204
210
|
*/
|
|
205
211
|
waitForOperation(operationId: string, opts?: WaitOptions): Promise<Operation>;
|
|
206
212
|
/** One operation (GET /v1/operations/{id}); lifecycle operations stay pollable after a workspace is deleted. */
|
package/dist/client.js
CHANGED
|
@@ -170,7 +170,9 @@ export class WorkspacesApi {
|
|
|
170
170
|
* that does not wait is polled with bounded exponential backoff (+-20 % jitter). Resolves when the operation
|
|
171
171
|
* succeeds; throws OperationFailedError when it fails or is canceled, OperationTimeoutError after `timeoutMs` (the
|
|
172
172
|
* operation keeps running and can be awaited again). Both errors carry the wait's `timing`; `onProgress` sees each
|
|
173
|
-
* state change (queued, capacity_pending, running, with the server's reason) as it is observed.
|
|
173
|
+
* state change (queued, capacity_pending, running, with the server's reason) as it is observed. A start waiting in
|
|
174
|
+
* capacity_pending gives up at its deadline (`deadlineAt` on the phase event) and fails with `capacity_unavailable`
|
|
175
|
+
* (`err.retryable` true: nothing was started, retry later); the SDK does not retry it.
|
|
174
176
|
*/
|
|
175
177
|
async waitForOperation(operationId, opts = {}) {
|
|
176
178
|
const inherited = opts[TRACE];
|
package/dist/errors.d.ts
CHANGED
|
@@ -53,13 +53,19 @@ type Operation = AppComponents['schemas']['Operation'];
|
|
|
53
53
|
/**
|
|
54
54
|
* Waiting for an operation ran out of time. The operation keeps running server side:
|
|
55
55
|
* resume with `cloud.workspaces.waitForOperation(err.operationId)` or call open() again
|
|
56
|
-
* (it returns the same operation while it is active).
|
|
56
|
+
* (it returns the same operation while it is active). A start still waiting for capacity
|
|
57
|
+
* (`capacity_pending`) gives up at `deadlineAt` and then fails with `capacity_unavailable`.
|
|
57
58
|
*/
|
|
58
59
|
export declare class OperationTimeoutError extends Error {
|
|
59
60
|
readonly operationId: string;
|
|
60
61
|
readonly workspaceId: string | null;
|
|
61
62
|
readonly lastState: Operation['state'];
|
|
62
63
|
readonly lastReason: string | null;
|
|
64
|
+
/**
|
|
65
|
+
* When the operation was last `capacity_pending`: when it gives up waiting for a host (`error.details.deadline_at`,
|
|
66
|
+
* RFC 3339) and fails with `capacity_unavailable`. Null in other states or from an API that does not report it.
|
|
67
|
+
*/
|
|
68
|
+
readonly deadlineAt: string | null;
|
|
63
69
|
readonly waitedMs: number;
|
|
64
70
|
/** Where the time went: client phases, retries and the operation's own server timing so far. */
|
|
65
71
|
timing: LifecycleTiming | undefined;
|
|
@@ -70,6 +76,12 @@ export declare class OperationFailedError extends Error {
|
|
|
70
76
|
readonly operation: Operation;
|
|
71
77
|
readonly operationId: string;
|
|
72
78
|
readonly errorCode: string | null;
|
|
79
|
+
/**
|
|
80
|
+
* `operation.error.retryable` (false when absent): the same request may succeed later. `capacity_unavailable` is
|
|
81
|
+
* retryable: no host could admit the start before its deadline, nothing was started, retry later. The SDK never
|
|
82
|
+
* retries a failed operation itself.
|
|
83
|
+
*/
|
|
84
|
+
readonly retryable: boolean;
|
|
73
85
|
/** Where the time went before the operation failed. */
|
|
74
86
|
timing: LifecycleTiming | undefined;
|
|
75
87
|
constructor(op: Operation);
|
package/dist/errors.js
CHANGED
|
@@ -1,3 +1,4 @@
|
|
|
1
|
+
import { capacityDeadlineOf } from "./progress.js";
|
|
1
2
|
export function isErrorBody(v) {
|
|
2
3
|
if (typeof v !== 'object' || v === null || !('error' in v))
|
|
3
4
|
return false;
|
|
@@ -45,23 +46,32 @@ export class ShardfluxProtocolError extends Error {
|
|
|
45
46
|
/**
|
|
46
47
|
* Waiting for an operation ran out of time. The operation keeps running server side:
|
|
47
48
|
* resume with `cloud.workspaces.waitForOperation(err.operationId)` or call open() again
|
|
48
|
-
* (it returns the same operation while it is active).
|
|
49
|
+
* (it returns the same operation while it is active). A start still waiting for capacity
|
|
50
|
+
* (`capacity_pending`) gives up at `deadlineAt` and then fails with `capacity_unavailable`.
|
|
49
51
|
*/
|
|
50
52
|
export class OperationTimeoutError extends Error {
|
|
51
53
|
operationId;
|
|
52
54
|
workspaceId;
|
|
53
55
|
lastState;
|
|
54
56
|
lastReason;
|
|
57
|
+
/**
|
|
58
|
+
* When the operation was last `capacity_pending`: when it gives up waiting for a host (`error.details.deadline_at`,
|
|
59
|
+
* RFC 3339) and fails with `capacity_unavailable`. Null in other states or from an API that does not report it.
|
|
60
|
+
*/
|
|
61
|
+
deadlineAt;
|
|
55
62
|
waitedMs;
|
|
56
63
|
/** Where the time went: client phases, retries and the operation's own server timing so far. */
|
|
57
64
|
timing = undefined;
|
|
58
65
|
constructor(op, waitedMs) {
|
|
59
|
-
|
|
66
|
+
const deadlineAt = capacityDeadlineOf(op);
|
|
67
|
+
const after = deadlineAt ? `it keeps waiting for a host server side until ${deadlineAt} and fails with capacity_unavailable if none admits it by then.` : 'it continues server side.';
|
|
68
|
+
super(`Operation ${op.id} (${op.kind}) is still ${op.state}${op.state_reason ? ` (${op.state_reason})` : ''} after ${Math.round(waitedMs)} ms; ${after}`);
|
|
60
69
|
this.name = 'OperationTimeoutError';
|
|
61
70
|
this.operationId = op.id;
|
|
62
71
|
this.workspaceId = op.workspace_id;
|
|
63
72
|
this.lastState = op.state;
|
|
64
73
|
this.lastReason = op.state_reason ?? null;
|
|
74
|
+
this.deadlineAt = deadlineAt;
|
|
65
75
|
this.waitedMs = waitedMs;
|
|
66
76
|
}
|
|
67
77
|
}
|
|
@@ -70,6 +80,12 @@ export class OperationFailedError extends Error {
|
|
|
70
80
|
operation;
|
|
71
81
|
operationId;
|
|
72
82
|
errorCode;
|
|
83
|
+
/**
|
|
84
|
+
* `operation.error.retryable` (false when absent): the same request may succeed later. `capacity_unavailable` is
|
|
85
|
+
* retryable: no host could admit the start before its deadline, nothing was started, retry later. The SDK never
|
|
86
|
+
* retries a failed operation itself.
|
|
87
|
+
*/
|
|
88
|
+
retryable;
|
|
73
89
|
/** Where the time went before the operation failed. */
|
|
74
90
|
timing = undefined;
|
|
75
91
|
constructor(op) {
|
|
@@ -79,5 +95,6 @@ export class OperationFailedError extends Error {
|
|
|
79
95
|
this.operation = op;
|
|
80
96
|
this.operationId = op.id;
|
|
81
97
|
this.errorCode = code;
|
|
98
|
+
this.retryable = op.error?.retryable === true;
|
|
82
99
|
}
|
|
83
100
|
}
|