@urun-sh/core 0.4.2 → 0.4.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,107 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.4.3
4
+
5
+ - **FIX: no server status word can reach a user as a benign lie — on EITHER path.**
6
+ Wave 0 (0.4.0) made the SFU WebSocket path honest: `STATUS_TO_PHASE` maps every
7
+ server word explicitly, an unmapped one goes LOUD (`console.error` via the
8
+ diagnostic mirror + an `unknown-server-status` diagnostic + phase `unknown`
9
+ carrying the literal word), and `reason` rides every phase. **It landed on the
10
+ WebSocket path only.** A brand-new user's cold start spends ALL of its time on
11
+ the HTTP admission path, and that path was documented as *"tolerant by
12
+ design"*: `admissionUpdateFrom` collapsed every 202 status word except
13
+ `queued` into `'pending'`, dropped `reason` entirely, and `reportAdmission`
14
+ rendered `provisioning`.
15
+
16
+ The consequence was that **`waiting_for_capacity` could not reach any client
17
+ at all**. The control plane answers a starved pull pool with
18
+ `202 {status:"waiting_for_capacity", reason:"no runtime worker is registered
19
+ for this pool — waiting for GPU capacity"}`; the phase exists in this SDK and
20
+ renders "Waiting for GPU capacity", but it was reachable ONLY from `_onStatus`
21
+ — i.e. only over a WebSocket that does not exist until allocation has already
22
+ succeeded, by which time capacity is no longer the question. A user whose org
23
+ has zero registered GPU workers watched "Provisioning…" for up to 600 s and
24
+ then got `never-live` "the backend is busy or wedged". The whole server-side
25
+ capacity observer was dead end-to-end.
26
+
27
+ Both entry points now share ONE `mapServerStatus(status)` and ONE loud-unknown
28
+ branch. `AdmissionUpdate.status` **widens from `'queued' | 'pending'` to
29
+ `string`** and carries the server's `reason`, plus `decisionReason` /
30
+ `refusalSource` / `refusalMessage` so a `provisioning_refused` 202 delivers
31
+ the PROVIDER's own message ("0/12 nodes are available: 12 Insufficient
32
+ nvidia.com/gpu…") instead of silence. The additive `runtime_state` detail may
33
+ still refine the two GENERIC waiting words (`queued`/`pending`) but may no
34
+ longer overrule a specific verdict — a starved answer carries
35
+ `runtime_state:'starting'` too, and letting `starting` win is exactly how the
36
+ capacity truth was lost.
37
+
38
+ - **FIX: every pause cause the server can write now has human copy, and a
39
+ conformance guard fails the build when the two lists diverge.**
40
+ `describePauseReason` knew 3 of the server's causes; the other 16 fell through
41
+ to the verbatim arm, so the canonical first-run failure printed as
42
+ *"Paused — connect_deadline. Resume any time."* — and the only thing relating
43
+ the two vocabularies was a COMMENT claiming the list was "grepped from the
44
+ server sources". The vocabulary is now VENDORED at
45
+ `src/server-pause-reasons.json`, GENERATED by
46
+ `scripts/refresh-pause-reason-fixture.mjs` from the SFU `PauseReason` union,
47
+ the `sessions.pause_reason` VOCABULARY clause (urun-infra migration
48
+ `20270906000000`) and the SFU `RuntimeNotDeliveringReason` union, each with
49
+ its path and blob sha. `src/pause-reason-vocabulary.test.ts` fails when any
50
+ vendored cause still hits the default arm. Every line keeps the invariant:
51
+ never "ended", never "start a new session", always resumable.
52
+
53
+ - **FIX: the default copy path shows an actionable hint, not developer text.**
54
+ `describeSessionPhase`'s `error` arm rendered ``Session failed — ${reason}``,
55
+ which for a create failure is the raw gateway string — an end user read
56
+ `[urun] session allocation failed at https://session-api.usw2.prod.cloud.urun.sh:
57
+ queue_full: … (HTTP 429)`. It now delegates to `describeSessionFailure`, so
58
+ message + hint are one code path. The `create-failed` hint gains arms for
59
+ **429** (names the load shed and surfaces the server's `Retry-After`) and
60
+ **5xx** (names the outage, says the retry is automatic) — both previously told
61
+ the user to "check the app/function name and auth", which is wrong about whose
62
+ fault it is. A lapsed queued REQUEST (`queue_ttl_expired` /
63
+ `stale_poll_expired`) no longer repeats the server's *"create a new session"* —
64
+ the request lapsed, the named object did not.
65
+
66
+ - **FIX: the server's load-shed pacing is no longer advisory.** A 429 sets a
67
+ jittered `Retry-After: 2-6`; the SDK read it only on the 202
68
+ legacy-immediate branch and DROPPED it on the throw path, so the outer sweep
69
+ re-dialled an overloaded control plane on its own full-jitter schedule.
70
+ `SessionAllocationError.retryAfterSeconds` now carries it, the sweep waits the
71
+ MAX of that and its own jitter (bounded by `BACKOFF_MAX_MS`), and
72
+ `SessionPhaseError.retryAfterSeconds` surfaces it to the copy.
73
+
74
+ The pacing floor is the **sweep-wide maximum**, not the last url's: every
75
+ gateway failure overwrites `lastError`, so a paced 429 from the primary
76
+ followed by an unpaced retryable failure from a regional fallback used to
77
+ erase the primary's request entirely and re-dial it on client jitter. The
78
+ same floor now applies to `resolveViewerConnect`, whose refusal ALSO carries
79
+ the server's `Retry-After` (it paced itself against a field that was never
80
+ populated). `samePhase` compares `retryAfterSeconds`, so a refusal that only
81
+ changed how long the backend is asking for re-stamps the phase instead of
82
+ leaving the UI quoting the previous number.
83
+
84
+ - **FIX: an HTTP 408 is a timeout, not a full admission queue.** The
85
+ `create-failed` hint folded 408 into the 429 arm and told the user "its
86
+ admission queue is full" — a confident FALSE cause for a status that only
87
+ establishes that a request timed out (`backoff.ts` treats it as a generic
88
+ retryable status; the durable-poll path throws `poll_http_408`). 408 now has
89
+ its own timeout copy, and still names the server's `Retry-After` when it sent
90
+ one.
91
+
92
+ - **FIX: the display sanitizer no longer has holes.** It knew exactly TWO
93
+ gateway wrappers, so token hydration (`session allocation at <url> never got
94
+ an access token`), the redirect envelope / self-redirect (`session create at
95
+ <url> …`), the poll give-up and the region-forward loop all still showed an
96
+ internal hostname in end-user copy — and a reason that is NOTHING BUT a
97
+ wrapper (the bare allocation timeout) stripped to empty and hit the
98
+ `|| reason` fallback, printing the raw string the strip existed to clean.
99
+ Every wrapper is listed, and a residual redaction is the backstop: no gateway
100
+ URL survives into display copy by any route.
101
+
102
+ - `whenLive`'s patience clocks (45 s / 600 s) are deliberately UNCHANGED here —
103
+ they are a separate item.
104
+
3
105
  ## 0.4.2
4
106
 
5
107
  - **Lockstep release** — version bump to match the `@urun-sh/openai` omp launcher cohort