@ultimat3/realtime 19.2.0 → 19.3.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CLAUDE.md CHANGED
@@ -307,6 +307,29 @@ Tier 3 package. Channels, live queries, local-first sync. One protocol for all t
307
307
  is entitled to its patches for the whole grace window, and to get them through a hub that is
308
308
  still open.
309
309
  - Exactly one `replicator` per DB, enforced by a session-level advisory lock.
310
+ - **A start and an acquisition are MEMOISED, because `running` and `#connection` are both written
311
+ after an await (`As of 2026-09-06`).** `start()` read `if (running) return true` and then awaited
312
+ `lock.tryAcquire()`, so two overlapping calls both passed the guard, both were told they held the
313
+ lock — a holder's `tryAcquire()` answers `true` — and both ran `feed.start()`: one replication
314
+ slot, two pumps, every change published twice under two `seq` generations of one producer id,
315
+ which every sync node's `SeqGapDetector` reads as a gap and repairs by re-snapshotting the fleet.
316
+ `PgAdvisoryLock.tryAcquire` had the identical hole one layer down and a worse ending: two
317
+ concurrent calls opened two SESSIONS, the second assignment to `#connection` orphaned the first,
318
+ and `release()` closes one — so the key stayed held until the process died and no standby could
319
+ ever take the slot. Both are now the shape `packages/core/src/lifecycle.ts`'s `drain()` uses: a
320
+ synchronous guard-and-register, the promise published before anything is awaited, and the memo
321
+ cleared however it settles (a `false` is "somebody else holds it right now", which the takeover
322
+ loop asks again a backoff later). `stop()` and `release()` await the in-flight one rather than
323
+ reading their own flag — that flag is false for the whole of a start, so an unguarded teardown
324
+ returns "nothing to do" and leaves behind exactly what it was called to release.
325
+ **A start that FAILS hands the lock back, and `running` is set after the feed is pumping**
326
+ (`As of 2026-09-06`). It was set before `await feed.start()`, so a feed that rejected — a slot
327
+ already `active`, a preflight refusal — left this node holding the advisory lock and claiming to
328
+ run with nothing pumping: the takeover loop's next `start()` was answered `true` by the
329
+ `if (running)` guard without re-entering `begin`, and every standby stayed a standby of a slot
330
+ whose holder was not replicating. The release is best-effort, because the feed's failure is the
331
+ one the operator acts on and a session-scoped lock a dead connection cannot release is released
332
+ by Postgres when that session ends.
310
333
  - **The replication pump has one way out, and it closes what it held.** Both exits — a decode error
311
334
  and `nextCopyData()` returning `undefined`, which is the walsender ending the copy — run `#die`:
312
335
  record `stats().failure`, stop the confirm timer, close the connection and null it. Each one left
@@ -742,9 +765,13 @@ Tier 3 package. Channels, live queries, local-first sync. One protocol for all t
742
765
  `DEFAULT_HEARTBEAT_MS`, 15s; `0` disables) sends a `hello` — byte-identical to the opening one,
743
766
  since the frame has no resume list to leave out — plus one subscribe frame per topic, which is the
744
767
  node's presence heartbeat. It is **not** how a deploy is noticed: `socket.skewed` compares the
745
- build id recorded at the upgrade against this node's, both fixed for the socket's life, so every
746
- `hello` on one socket answers the same forever and `update-available` reaches a client on the
747
- socket it opens against the *new* node. Two silent windows and the client closes with `4000` and
768
+ build the client claims (the `hello`'s `buildId`, which `sawHello` records on every one — the
769
+ latest is the record — or `?build=` on the dial) against this node's; a client says the same
770
+ build on every beat and the node's never moves while the socket is open, so every `hello` on one
771
+ socket answers the same forever and `update-available` reaches a client on the
772
+ socket it opens against the *new* node. The hello IS read — until 2026-09-07 only the dial was,
773
+ and a dial without `?build=` was recorded as this node's own id, so a client naming its build only
774
+ in the frame was current forever. Two silent windows and the client closes with `4000` and
748
775
  arms the reconnect. It is one
749
776
  self-re-arming tick on the injected `Scheduler`, not an interval: a client is either beating on a
750
777
  live socket or backing off toward a new one, never both. The 15s is the client's OWN number:
@@ -797,7 +824,24 @@ Tier 3 package. Channels, live queries, local-first sync. One protocol for all t
797
824
  *replay* bound (what a delta resume costs to fold) and adds the byte budgets as the memory one —
798
825
  `packages/cache/src/lru.ts:1-2` states why: 4,096 queries x 1,024 patches is 4.19M retained rows
799
826
  and no number of bytes at all. `forget(qid)` is called by `LiveQueryRegistry.unsubscribe` when the
800
- last subscriber of a query id goes; it had no caller, so the ring outlived the entry.
827
+ last subscriber of a query id goes; it had no caller, so the ring outlived the entry. It is
828
+ also called by `#dropIfUnheld` on the subscribe path, `As of 2026-09-06`: an entry is created
829
+ BEFORE the snapshot read that fills it, so a cold subscribe the database refused left an entry
830
+ with no subscriber and no removal path — `unsubscribe` can only reach one through a subscription
831
+ that was never attached. `maxEntries` cold failures were therefore enough to answer
832
+ `X_SUBSCRIPTION_LIMIT` to every later subscriber for the life of the process, after the database
833
+ had recovered, because `qid` derives from client-chosen input and distinct inputs mint distinct
834
+ orphans. Dropped only when the entry is still the one in the table, holds no subscriber and has
835
+ no read in flight — a read published on it belongs to a concurrent subscriber that has not
836
+ attached yet, and dropping it there would hand that subscriber a window no change reaches.
837
+ - **A channel patch id carries the NODE that minted it** (`As of 2026-09-06`). `ChannelHub`'s
838
+ `#sequence` counts within one process, and the patch id was that counter alone — so two `sync`
839
+ replicas publishing to one topic minted the same id for the same subscriber, and a channel has
840
+ no cursor and no re-snapshot, so nothing downstream can repair a collision. `nodeId` defaults to
841
+ a per-hub `uuid()` and is declarable (the pod name) when an operator should be able to say which
842
+ node published a frame. The frame's `lsn` is deliberately left as the per-process counter:
843
+ nothing reads a channel frame's lsn as an order across nodes, and `client-frames.ts` advances a
844
+ cursor only for a registered live query, never for a topic.
801
845
  - **An error never renders a value that carries a credential.** `parsePgUrl` names `DATABASE_URL`
802
846
  rather than echoing the URL it refused — an error reaches a log, `--json`, an agent transcript
803
847
  and a ticket, and the password is in the string. Same rule as `packages/mail/src/driver-smtp.ts`.
package/README.md CHANGED
@@ -306,7 +306,7 @@ new LiveClient({ signal, connect, buildId, heartbeatMs: 15_000 }); // 0 disables
306
306
  | Default | `DEFAULT_HEARTBEAT_MS`, 15s. The client's own number and the only one: `realtime.heartbeatMs` in `app.config.ts` was deleted 2026-08-19 because nothing read it |
307
307
  | One beat | a `hello` — which carries no cursors at all; `HelloFrame` has no resume list, so a beat and an opening frame are byte-identical — plus one subscribe frame per topic held |
308
308
  | Why the topics | on the node, repeating the subscribe frame **is** the presence heartbeat; presence has no frame of its own in either direction |
309
- | Not a deploy check | `update-available` answers a skew between the build id recorded at the upgrade and the node's own, and neither can change on an open socket — so every `hello` on one socket answers the same forever. A client hears about a deploy on the socket it opens against the **new** node |
309
+ | Not a deploy check | `update-available` answers a skew between the build the client claims — the `hello`'s `buildId`, or `?build=` on the dial; every hello is read and the latest one is the record — and the node's own. A client says the same build on every beat and the node's never moves while the socket is open, so every `hello` on one socket answers the same forever. A client hears about a deploy on the socket it opens against the **new** node |
310
310
  | Silence | nothing received for **two** intervals ⇒ close `4000` (a private-use code, so it is distinguishable in a log) and arm the reconnect. Judged from the last frame of any kind, since the point is that bytes still cross |
311
311
  | Not an interval | one armed tick, re-armed by itself, on the same injected `Scheduler` the reconnect uses — a client is either beating on a live socket or backing off toward a new one, never both |
312
312
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ultimat3/realtime",
3
- "version": "19.2.0",
3
+ "version": "19.3.2",
4
4
  "description": "Three-tier realtime: channels, live queries, local-first sync — one protocol, one mutator shape",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -36,8 +36,8 @@
36
36
  "test": "bun test"
37
37
  },
38
38
  "dependencies": {
39
- "@ultimat3/core": "19.2.0",
40
- "@ultimat3/query": "19.2.0",
39
+ "@ultimat3/core": "19.3.2",
40
+ "@ultimat3/query": "19.3.2",
41
41
  "nats": "2.29.3"
42
42
  }
43
43
  }
package/src/channel.ts CHANGED
@@ -4,7 +4,7 @@
4
4
  // append-only stream, so tier 1 needs no frame of its own. That is why climbing the ladder is a
5
5
  // config change: the client's frame handler is the same code at every rung.
6
6
 
7
- import { type Actor, finiteOption, logger, renderThrowable } from '@ultimat3/core';
7
+ import { type Actor, finiteOption, logger, renderThrowable, uuid } from '@ultimat3/core';
8
8
  import { formatLsn } from './changefeed';
9
9
  import {
10
10
  isPolicyDenial,
@@ -57,6 +57,17 @@ export interface ChannelHubOptions {
57
57
  * admits unbounded distinct names inside one tenant, and a per-socket cap bounds nothing.
58
58
  */
59
59
  readonly maxTopicsPerNode?: number;
60
+ /**
61
+ * This node's mark on the patch ids it mints. Defaults to a per-hub random id, which is enough
62
+ * to keep two nodes apart; declare it (the pod name, the `sync` instance id) when an operator
63
+ * reading one frame should be able to say which node published it.
64
+ *
65
+ * A blank string is read as OMITTED, never as a mark: `??` only answers for `undefined`, so
66
+ * `nodeId: ''` — which is what an unset `POD_NAME` interpolates to — stored the empty mark and
67
+ * two hubs then minted the SAME first patch id, `:0000000000000001`. That is precisely the
68
+ * collision this field exists to prevent, arriving through the field itself.
69
+ */
70
+ readonly nodeId?: string;
60
71
  }
61
72
 
62
73
  /** Distinct topics one node bridges before `X_SUBSCRIPTION_LIMIT`. */
@@ -95,6 +106,12 @@ export class ChannelHub {
95
106
  readonly #maxTopicsPerNode: number;
96
107
  #guardFailures = 0;
97
108
  #sequence = 0n;
109
+ /**
110
+ * `#sequence` counts within one PROCESS, so it cannot identify a message across nodes: two
111
+ * `sync` replicas publishing to one topic minted the same id for the same subscriber, and a
112
+ * channel has no cursor and no re-snapshot, so nothing downstream could repair the collision.
113
+ */
114
+ readonly #nodeId: string;
98
115
  /** Set by `close()`. Read by `#open`, which is the only thing that can reach a late subscription. */
99
116
  #closed = false;
100
117
 
@@ -111,6 +128,9 @@ export class ChannelHub {
111
128
  'maxTopicsPerNode',
112
129
  options.maxTopicsPerNode ?? DEFAULT_MAX_TOPICS_PER_NODE,
113
130
  );
131
+ // Trimmed before the emptiness test: `nodeId: ' '` marks a frame with a space, which reads in
132
+ // a log as no mark at all and collides with the next hub that does the same.
133
+ this.#nodeId = options.nodeId?.trim() || uuid();
114
134
  }
115
135
 
116
136
  /** Sockets this node will deliver `name` to. The metric the fanout reads. */
@@ -230,7 +250,10 @@ export class ChannelHub {
230
250
  /** Publishes to every node. Local delivery happens via the transport bridge, never directly. */
231
251
  async publish(name: Topic, message: JsonObject): Promise<void> {
232
252
  this.#sequence += 1n;
233
- const frame = channelFrame(name, formatLsn(this.#sequence), message);
253
+ const lsn = formatLsn(this.#sequence);
254
+ // The lsn stays this node's own counter — nothing reads a channel frame's lsn as an order
255
+ // across nodes — but the patch ID is what a client keys by, so it carries the node too.
256
+ const frame = channelFrame(name, lsn, message, `${this.#nodeId}:${lsn}`);
234
257
  await this.#transport.publish(`${CHANNEL_SUBJECT_PREFIX}.${name}`, encode(frame));
235
258
  }
236
259
 
@@ -360,12 +383,22 @@ function unsubscribeWhenOpen(bridge: Bridge): void {
360
383
  );
361
384
  }
362
385
 
363
- export function channelFrame(name: Topic, lsn: string, message: JsonObject): Frame {
386
+ /**
387
+ * `id` identifies the MESSAGE and defaults to the lsn, which is what every caller outside this
388
+ * file already passes as one. `ChannelHub.publish` gives it the publishing node's mark instead:
389
+ * an lsn is a per-process counter, and two nodes on one topic mint the same one.
390
+ */
391
+ export function channelFrame(
392
+ name: Topic,
393
+ lsn: string,
394
+ message: JsonObject,
395
+ id: string = lsn,
396
+ ): Frame {
364
397
  return {
365
398
  type: 'patch',
366
399
  v: PROTOCOL_VERSION,
367
400
  sid: name,
368
401
  lsn,
369
- patches: [{ op: 'insert', id: lsn, row: message, lsn }],
402
+ patches: [{ op: 'insert', id, row: message, lsn }],
370
403
  };
371
404
  }
@@ -3,6 +3,7 @@
3
3
  // every piece of the client a frame may touch, so the blast radius of a new frame kind is a
4
4
  // reviewable list rather than "whatever the router could reach through `this`".
5
5
 
6
+ import { CLOSE } from './close-codes';
6
7
  import { advance } from './cursor';
7
8
  import type { JsonObject, JsonValue } from './json';
8
9
  import type { Registration, RowWindows } from './live-rows';
@@ -14,6 +15,17 @@ import type { Frame, PresenceMember } from './sync-protocol';
14
15
  /** Declared with the window it projects; re-exported here because the router is what writes it. */
15
16
  export type { LiveState, Registration } from './live-rows';
16
17
 
18
+ /**
19
+ * The code a `reconnect` frame closes with: `CLOSE.drain`, the same number the node uses for a
20
+ * drain it closes itself, so a log reads one code for one event whichever side closed first. It
21
+ * was 1001, and a browser refuses that from script: `WebSocket.close()` throws
22
+ * `InvalidAccessError: The close code must be either 1000, or between 3000 and 4999` — measured
23
+ * in Chrome, an uncaught exception in every tab on every node drain. The reconnect still happened,
24
+ * because the node closed the socket a moment later; the exception was the only trace.
25
+ * `HEARTBEAT_TIMEOUT_CODE` in `client.ts` is the sibling, for the other close the client makes.
26
+ */
27
+ export const RECONNECT_CODE = CLOSE.drain;
28
+
17
29
  /**
18
30
  * Everything an inbound frame is allowed to reach. Narrow on purpose — a router that took the
19
31
  * client itself could touch the reconnect timer, the socket and the outbound path, none of which
@@ -151,7 +163,7 @@ export function applyFrame<T extends TableMap>(frame: Frame, target: ClientFrame
151
163
  // Order is load-bearing: arming first is what makes the close this triggers keep the delay
152
164
  // the node assigned to *this* socket instead of falling back to a local backoff.
153
165
  target.scheduleReconnect(frame.afterMs);
154
- target.closeSocket(1001, frame.reason);
166
+ target.closeSocket(RECONNECT_CODE, frame.reason);
155
167
  return;
156
168
  }
157
169
  case 'update-available': {
package/src/client.ts CHANGED
@@ -45,7 +45,11 @@ export type {
45
45
  /** The four states a live subscription renders. Declared with the window that holds them. */
46
46
  export type { LiveState } from './live-rows';
47
47
 
48
- /** Private-use close code (4000–4999), so a heartbeat timeout is distinguishable in a log. */
48
+ /**
49
+ * Private-use close code (4000–4999), so a heartbeat timeout is distinguishable in a log. A browser
50
+ * accepts only 1000 and 3000–4999 from script; `RECONNECT_CODE` (`client-frames.ts`) is the
51
+ * sibling, for the close a `reconnect` frame makes.
52
+ */
49
53
  const HEARTBEAT_TIMEOUT_CODE = 4000;
50
54
 
51
55
  /** The default reporter: `console.error`, never core's `logger` — that writes `process.stderr`. */
@@ -395,10 +399,11 @@ export class LiveClient<T extends TableMap = TableMap> {
395
399
  * heartbeat: subscribing IS being in the room, so a client that stopped repeating it is swept
396
400
  * out of every room it is still receiving from.
397
401
  *
398
- * It is NOT how a deploy is noticed. `socket.skewed` compares the build id the upgrade recorded
399
- * against this node's, both fixed for the socket's whole life, so every `hello` on one socket
400
- * gets the same answer forever; `update-available` reaches a client on the socket it opens
401
- * against the *new* node, which is a reconnect and never a beat.
402
+ * It is NOT how a deploy is noticed. `socket.skewed` compares the build id this client says —
403
+ * the `hello`'s own `buildId`, the same on every beat, or `?build=` on the dial — against the
404
+ * node's, and neither moves while the socket is open, so every `hello` on one socket gets the
405
+ * same answer forever; `update-available` reaches a client on the socket it opens against the
406
+ * *new* node, which is a reconnect and never a beat.
402
407
  */
403
408
  #beat(): void {
404
409
  this.#send(this.#hello());
@@ -0,0 +1,21 @@
1
+ // The close codes this protocol speaks, defined once below both halves. `socket.ts` (the node's
2
+ // registry) and `client-frames.ts` (browser code) each need one of these and neither may import
3
+ // the other — the client must not pull the node's registry into the tab — so the table lives here,
4
+ // a leaf with no import at all, and `socket.ts` re-exports it for the node-side files that already
5
+ // read `CLOSE` from there.
6
+ //
7
+ // 1000–1015 are the RFC 6455 codes; 4000–4999 are private use. A browser accepts only 1000 and
8
+ // 3000–4999 from script (`WebSocket.close()` throws `InvalidAccessError` on anything else), which
9
+ // is why every code the CLIENT sends is in the private range and why `goingAway` (1001) is a code
10
+ // only the node may close with.
11
+
12
+ export const CLOSE = {
13
+ normal: 1000,
14
+ goingAway: 1001,
15
+ policy: 1008,
16
+ overloaded: 1013,
17
+ versionSkew: 4000,
18
+ idle: 4001,
19
+ /** The node drains, or the client obeys a `reconnect` frame: one event, one code, either side. */
20
+ drain: 4002,
21
+ } as const;
package/src/live-query.ts CHANGED
@@ -189,6 +189,29 @@ export class LiveQueryRegistry {
189
189
 
190
190
  const qid = queryHash(args.name, args.input);
191
191
  const entry = this.#entryFor(qid, definition, args.input);
192
+ try {
193
+ return await this.#serve(entry, sid, args);
194
+ } catch (error) {
195
+ // The entry was born above, before anything could fill it. Everything below can throw — the
196
+ // snapshot, the resume, the window's own read deadline — and `unsubscribe` is the only other
197
+ // removal path, reachable only through a subscription that in this case was never attached.
198
+ this.#dropIfUnheld(qid, entry);
199
+ throw error;
200
+ }
201
+ }
202
+
203
+ /** The half of a subscribe that runs against a live entry, so its failures can be undone. */
204
+ async #serve(
205
+ entry: QueryEntry,
206
+ sid: string,
207
+ args: {
208
+ socket: SyncSocket;
209
+ name: string;
210
+ input: JsonValue;
211
+ cursor?: LiveCursor | null;
212
+ },
213
+ ): Promise<{ subscription: LiveSubscription; frame: Frame }> {
214
+ const qid = entry.qid;
192
215
  const now = this.#clock.now().getTime();
193
216
 
194
217
  if (args.cursor) {
@@ -260,6 +283,25 @@ export class LiveQueryRegistry {
260
283
  this.#options.source.forget?.(subscription.qid);
261
284
  }
262
285
 
286
+ /**
287
+ * An entry nothing is holding, dropped with the retained patches that were kept for it — the
288
+ * same three steps `unsubscribe` takes when the last subscriber leaves, for the case where a
289
+ * subscriber never arrived. Five cold failures against `maxEntries: 5` used to refuse every
290
+ * later subscribe on the node with `X_SUBSCRIPTION_LIMIT`, forever, after the database recovered.
291
+ */
292
+ #dropIfUnheld(qid: string, entry: QueryEntry): void {
293
+ // Identity, never presence: an entry this subscribe did not create may already have been
294
+ // dropped and re-created under the same qid.
295
+ if (this.#entries.get(qid) !== entry) return;
296
+ if (entry.subscribers.size !== 0) return;
297
+ // A read published on the entry belongs to a subscriber that has not attached yet — it owns
298
+ // the entry's fate, and dropping it here would leave that subscriber holding a window no
299
+ // change is ever fanned out to.
300
+ if (entry.reading !== null) return;
301
+ this.#entries.delete(qid);
302
+ this.#options.source.forget?.(qid);
303
+ }
304
+
263
305
  unsubscribeSocket(socketId: string): void {
264
306
  for (const subscription of this.#book.ofSocket(socketId)) {
265
307
  this.unsubscribe(socketId, subscription.sid);
@@ -34,6 +34,17 @@ export class PgAdvisoryLock implements AdvisoryLock {
34
34
  readonly key: string;
35
35
  readonly #options: PgAdvisoryLockOptions;
36
36
  #connection: PgConnection | null = null;
37
+ /**
38
+ * The acquisition in flight, if any. `#connection` cannot answer "am I already taking this" —
39
+ * it is written at the END of an acquisition, with a dial, a handshake and a query awaited in
40
+ * between, so two concurrent callers both read `null` and both opened a session. Both then took
41
+ * a grant on the same key (a second `pg_try_advisory_lock` on a DIFFERENT session succeeds only
42
+ * if the first has not landed yet, and on the same key from two sessions one of them answers
43
+ * `f` — either way the loser's session was left open), and the second assignment to
44
+ * `#connection` orphaned the first: `release()` closes one, and the other holds the lock until
45
+ * the process dies. A standby then never takes over a slot whose owner has already stopped.
46
+ */
47
+ #acquiring: Promise<boolean> | null = null;
37
48
 
38
49
  constructor(options: PgAdvisoryLockOptions) {
39
50
  if (!KEY_PATTERN.test(options.key)) {
@@ -49,9 +60,27 @@ export class PgAdvisoryLock implements AdvisoryLock {
49
60
  this.#options = options;
50
61
  }
51
62
 
52
- async tryAcquire(): Promise<boolean> {
63
+ /**
64
+ * Deliberately NOT `async`: the guard and the registration have to be one synchronous step, or
65
+ * the memo has the same check-then-act hole as the field it replaces. Same shape as
66
+ * `packages/core/src/lifecycle.ts`'s `drain()` and `packages/jobs/src/worker.ts`'s `round()`.
67
+ */
68
+ tryAcquire(): Promise<boolean> {
53
69
  // Already ours — see the class comment on why a second `pg_try_advisory_lock` must not run.
54
- if (this.#connection !== null) return true;
70
+ if (this.#connection !== null) return Promise.resolve(true);
71
+ const inFlight = this.#acquiring;
72
+ if (inFlight !== null) return inFlight;
73
+ // Cleared however it settles: a `false` is "somebody else holds it right now", which the
74
+ // replicator's takeover loop asks again a backoff later, and a rejection is the database
75
+ // refusing THIS attempt. A memo either of those stuck to would answer for the process's life.
76
+ const attempt = this.#acquire().finally(() => {
77
+ if (this.#acquiring === attempt) this.#acquiring = null;
78
+ });
79
+ this.#acquiring = attempt;
80
+ return attempt;
81
+ }
82
+
83
+ async #acquire(): Promise<boolean> {
55
84
  const target = parsePgUrl(this.#options.url);
56
85
  const stream = await (this.#options.stream ?? bunPgStream)(target);
57
86
  // Plain SQL, not `replication: 'database'` — this session runs one statement, never a feed.
@@ -85,6 +114,13 @@ export class PgAdvisoryLock implements AdvisoryLock {
85
114
  }
86
115
 
87
116
  async release(): Promise<void> {
117
+ // A release that arrives mid-acquisition waits it out. `#connection` is still `null` there, so
118
+ // the early return below would answer "nothing to release" and the session would open behind
119
+ // it, holding the key with no object left intending to close it — the same orphan the memo
120
+ // exists to prevent, reached from the other side. The attempt's own failure is not this
121
+ // caller's to report: `tryAcquire`'s caller already has it.
122
+ const acquiring = this.#acquiring;
123
+ if (acquiring !== null) await acquiring.catch(() => false);
88
124
  const connection = this.#connection;
89
125
  if (connection === null) return;
90
126
  this.#connection = null;
package/src/replicator.ts CHANGED
@@ -116,6 +116,8 @@ export function createReplicator(options: ReplicatorOptions): Replicator {
116
116
  // the old one.
117
117
  let producer = uuid();
118
118
  let seq = 0;
119
+ /** The start in flight, if any — see `start()` on why `running` cannot answer that question. */
120
+ let starting: Promise<boolean> | undefined;
119
121
 
120
122
  const onChange = async (raw: ChangeEvent): Promise<void> => {
121
123
  const change = normalize(raw);
@@ -141,24 +143,66 @@ export function createReplicator(options: ReplicatorOptions): Replicator {
141
143
  });
142
144
  };
143
145
 
144
- return {
145
- async start(): Promise<boolean> {
146
- if (running) return true;
147
- if (!(await options.lock.tryAcquire())) {
148
- logger.warn('replicator standby: advisory lock held elsewhere', { key: options.lock.key });
149
- return false;
150
- }
151
- running = true;
152
- producer = uuid();
153
- seq = 0;
146
+ const begin = async (): Promise<boolean> => {
147
+ if (!(await options.lock.tryAcquire())) {
148
+ logger.warn('replicator standby: advisory lock held elsewhere', { key: options.lock.key });
149
+ return false;
150
+ }
151
+ producer = uuid();
152
+ seq = 0;
153
+ try {
154
154
  await options.feed.start(
155
155
  options.from === undefined ? { onChange } : { from: options.from, onChange },
156
156
  );
157
- logger.info('replicator started', { source: options.feed.source, key: options.lock.key });
158
- return true;
157
+ } catch (thrown) {
158
+ // A start that failed holds NOTHING. `running` stays false so the takeover loop's next
159
+ // `start()` runs `begin` again instead of being answered `true` by a memo over a feed that
160
+ // never pumped, and the lock goes back so a standby can take the slot rather than waiting on
161
+ // a holder that is not replicating.
162
+ //
163
+ // The release is best-effort by design: the feed's failure is the one the caller must see,
164
+ // and a session-scoped advisory lock a dead connection cannot release is released by
165
+ // Postgres when that session ends.
166
+ await options.lock.release().catch(() => undefined);
167
+ throw thrown;
168
+ }
169
+ // AFTER the feed is pumping, never before: `running` is what `start()` answers `true` from and
170
+ // what `stop()` reads to decide there is anything to tear down.
171
+ running = true;
172
+ logger.info('replicator started', { source: options.feed.source, key: options.lock.key });
173
+ return true;
174
+ };
175
+
176
+ return {
177
+ /**
178
+ * Deliberately NOT `async`: `running` is set after an `await` on the lock, so a guard that
179
+ * awaited anything before registering has the same hole it is meant to close. Two overlapping
180
+ * starts both passed `if (running)`, both were told they held the lock — a holder's
181
+ * `tryAcquire()` answers `true` — and both ran `feed.start()`: one replication slot with two
182
+ * pumps, publishing every change twice under two `seq` generations of one producer id, which
183
+ * every sync node's `SeqGapDetector` reads as a gap and repairs by re-snapshotting the fleet.
184
+ * Same shape as `packages/core/src/lifecycle.ts`'s `drain()`.
185
+ */
186
+ start(): Promise<boolean> {
187
+ if (running) return Promise.resolve(true);
188
+ const inFlight = starting;
189
+ if (inFlight !== undefined) return inFlight;
190
+ // Cleared however it settles: a `false` is "another node holds the lock", and the takeover
191
+ // loop asks again a `retryDelayMs` later. A memo that stuck would make this node a permanent
192
+ // standby of a slot whose holder has already died.
193
+ const attempt = begin().finally(() => {
194
+ if (starting === attempt) starting = undefined;
195
+ });
196
+ starting = attempt;
197
+ return attempt;
159
198
  },
160
199
 
161
200
  async stop(): Promise<void> {
201
+ // A stop racing a start waits it out: `running` is false for the whole of `begin`, so the
202
+ // early return below would leave the feed it is about to start pumping into a replicator
203
+ // nothing intends to stop, with the advisory lock still held.
204
+ const inFlight = starting;
205
+ if (inFlight !== undefined) await inFlight.catch(() => false);
162
206
  if (!running) return;
163
207
  running = false;
164
208
  await options.feed.stop();
package/src/socket.ts CHANGED
@@ -16,18 +16,12 @@ import {
16
16
  systemClock,
17
17
  uuid,
18
18
  } from '@ultimat3/core';
19
+ import { CLOSE } from './close-codes';
19
20
  import { encode, type Frame } from './sync-protocol';
20
21
  import { AcceptBudget } from './thundering-herd';
21
22
 
22
- export const CLOSE = {
23
- normal: 1000,
24
- goingAway: 1001,
25
- policy: 1008,
26
- overloaded: 1013,
27
- versionSkew: 4000,
28
- idle: 4001,
29
- drain: 4002,
30
- } as const;
23
+ /** Defined in `close-codes.ts`, below both halves; re-exported for the node-side files. */
24
+ export { CLOSE } from './close-codes';
31
25
 
32
26
  /**
33
27
  * The slice of Bun's `ServerWebSocket` this package uses. Structural, so tests need no server.
@@ -49,7 +43,11 @@ export interface WsLike {
49
43
 
50
44
  export interface SyncSocketOptions {
51
45
  readonly ws: WsLike;
52
- /** Build id the *client* reported in `hello`. Version skew is a first-class connection state. */
46
+ /**
47
+ * Build id the *client* reported on the dial (`?build=`), or this node's own when it sent none;
48
+ * `sawHello` overwrites it with what the `hello` frame says. Version skew is a first-class
49
+ * connection state.
50
+ */
53
51
  readonly clientBuildId: string;
54
52
  readonly serverBuildId: string;
55
53
  readonly actor?: Actor | null;
@@ -113,7 +111,6 @@ export function actorIdOf(actor: Actor | null): string | null {
113
111
  */
114
112
  export class SyncSocket {
115
113
  readonly id: string;
116
- readonly clientBuildId: string;
117
114
  readonly serverBuildId: string;
118
115
  readonly openedAt: number;
119
116
  /** Channel topics (tier 1). Live-query subscriptions are keyed separately, by sid. */
@@ -149,13 +146,14 @@ export class SyncSocket {
149
146
  readonly #clock: Clock;
150
147
  readonly #maxBufferedBytes: number;
151
148
  readonly #maxDroppedFrames: number;
149
+ #clientBuildId: string;
152
150
  #closed = false;
153
151
 
154
152
  constructor(options: SyncSocketOptions) {
155
153
  this.#ws = options.ws;
156
154
  this.#clock = options.clock ?? systemClock;
157
155
  this.id = options.id ?? uuid();
158
- this.clientBuildId = options.clientBuildId;
156
+ this.#clientBuildId = options.clientBuildId;
159
157
  this.serverBuildId = options.serverBuildId;
160
158
  this.actor = options.actor ?? null;
161
159
  this.#maxBufferedBytes = options.maxBufferedBytes ?? DEFAULT_MAX_BUFFERED_BYTES;
@@ -185,9 +183,26 @@ export class SyncSocket {
185
183
  return this.#closed;
186
184
  }
187
185
 
186
+ /** The build the client last claimed: the dial's `?build=`, then whatever its `hello` said. */
187
+ get clientBuildId(): string {
188
+ return this.#clientBuildId;
189
+ }
190
+
191
+ /**
192
+ * The `hello` frame is the documented place a client names its build, and until 2026-09-07 the
193
+ * node never read it: `clientBuildId` came from the dial's `?build=` alone and defaulted to this
194
+ * node's OWN id when the query was absent — so a client that said `hello` from any build at all
195
+ * was deemed current forever, and `update-available` never came. Measured on ai-maxxing: a page
196
+ * sending `buildId: "dev"` in every hello to a node on `46db23f57d6ef969`. The last word wins,
197
+ * and the hello is the later one; `?build=` still works for a dial that carries it.
198
+ */
199
+ sawHello(buildId: string): void {
200
+ this.#clientBuildId = buildId;
201
+ }
202
+
188
203
  /** A skewed client gets an `update-available` frame; it is never silently served a new shape. */
189
204
  get skewed(): boolean {
190
- return this.clientBuildId !== this.serverBuildId;
205
+ return this.#clientBuildId !== this.serverBuildId;
191
206
  }
192
207
 
193
208
  /** `false` means the frame was dropped by backpressure — the caller must mark state stale. */
@@ -76,6 +76,9 @@ export function createFrameRouter(options: FrameRouterOptions): FrameRouter {
76
76
  async function apply(socket: SyncSocket, frame: Frame): Promise<void> {
77
77
  switch (frame.type) {
78
78
  case 'hello': {
79
+ // Before `skewed` is asked: the frame is the client's word on its build, and the upgrade
80
+ // may have recorded none (a dial without `?build=` defaults to this node's own id).
81
+ socket.sawHello(frame.buildId);
79
82
  socket.send({
80
83
  type: 'hello',
81
84
  v: PROTOCOL_VERSION,
@@ -20,6 +20,13 @@ export interface ListenOptions {
20
20
  * defeated by the listener beside it.
21
21
  */
22
22
  readonly hostname?: string;
23
+ /**
24
+ * The grace `drain()` gives the sockets it holds on SIGTERM, in ms; the node's default when
25
+ * omitted. `0` is for a process with nobody to hand its clients to — `x dev`, one node, whose
26
+ * clients reconnect to it once it is back and to nothing in the meantime: the 5s it kept their
27
+ * patches flowing was 5s of a Ctrl-C that had nothing else to wait for.
28
+ */
29
+ readonly drainGraceMs?: number;
23
30
  }
24
31
 
25
32
  export interface SyncListener {
@@ -61,7 +68,7 @@ export function listenSyncNode(node: SyncNode, options: ListenOptions = {}): Syn
61
68
  { phase: 'accept' },
62
69
  );
63
70
  const unregister = onShutdown('realtime:sync', async () => {
64
- await node.drain();
71
+ await node.drain(options.drainGraceMs === undefined ? {} : { graceMs: options.drainGraceMs });
65
72
  await node.stop();
66
73
  server.stop();
67
74
  stopListening();
package/src/sync-node.ts CHANGED
@@ -148,6 +148,8 @@ export function createSyncNode(options: SyncNodeOptions): SyncNode {
148
148
  const grants = new GrantBook();
149
149
  const gaps = new SeqGapDetector();
150
150
  let ready = false;
151
+ /** Resolved by `teardown` when the last socket leaves, for a drain that is waiting its grace. */
152
+ let lastSocketLeft: (() => void) | null = null;
151
153
  let changes: TransportSubscription | null = null;
152
154
  let sweeping: ReturnType<typeof setInterval> | null = null;
153
155
  let reauthing: ReturnType<typeof setInterval> | null = null;
@@ -188,6 +190,7 @@ export function createSyncNode(options: SyncNodeOptions): SyncNode {
188
190
  for (const name of topics) options.hub.unsubscribe(socket, name);
189
191
  sockets.remove(socket.id);
190
192
  grants.delete(socket.id);
193
+ if (sockets.count === 0) lastSocketLeft?.();
191
194
  // A closed socket is a leave, said now rather than left to TTL: everyone else would otherwise
192
195
  // keep rendering a member who is provably gone for the rest of its window. The write is on the
193
196
  // bus and the close callback is synchronous, so it cannot be awaited here.
@@ -210,6 +213,18 @@ export function createSyncNode(options: SyncNodeOptions): SyncNode {
210
213
  * because dropping the socket from the table is three of `teardown`'s five steps and the two it
211
214
  * misses are the ones another node can see.
212
215
  */
216
+ /** `graceMs`, or until `teardown` reports the table empty — whichever comes first. */
217
+ const graceOrEmpty = (graceMs: number): Promise<void> =>
218
+ new Promise<void>((resolve) => {
219
+ const timer = setTimeout(done, graceMs);
220
+ function done(): void {
221
+ clearTimeout(timer);
222
+ lastSocketLeft = null;
223
+ resolve();
224
+ }
225
+ lastSocketLeft = done;
226
+ });
227
+
213
228
  const evict = (socket: SyncSocket, code: number, reason: string): readonly Promise<unknown>[] => {
214
229
  socket.close(code, reason);
215
230
  return teardown(socket);
@@ -447,8 +462,12 @@ export function createSyncNode(options: SyncNodeOptions): SyncNode {
447
462
  if (notified < plan.length) {
448
463
  logger.warn('sync.drain_frames_dropped', { sockets: plan.length, notified });
449
464
  }
465
+ // The grace is what a socket is OWED after its reconnect frame — its patches, until its own
466
+ // delay is up — so it is waited only while there is a socket to owe it to, and ends the
467
+ // moment the last one leaves. Unconditional, it was 5.0s of a 5.1s Ctrl-C on `x dev` with
468
+ // no browser open (measured 2026-09-06): a drain of zero sockets, sleeping for nobody.
450
469
  const graceMs = drainGraceMs(drainOptions.graceMs);
451
- if (graceMs > 0) await new Promise((resolve) => setTimeout(resolve, graceMs));
470
+ if (graceMs > 0 && sockets.count > 0) await graceOrEmpty(graceMs);
452
471
  // Through `evict`, never `sockets.remove` + `grants.delete`: those are three of `teardown`'s
453
472
  // five steps, and the two they skip are the ones the rest of the fleet can see. A drained
454
473
  // socket that never left its presence set is a member every other node renders for a full
@@ -14,6 +14,12 @@ import type { AcceptBudget, Rng } from './thundering-herd';
14
14
  */
15
15
  export interface WsData {
16
16
  readonly socketId: string;
17
+ /**
18
+ * `?build=` off the dial, or this node's own id when the dial carried none. A starting value,
19
+ * not the verdict: the `hello` frame's `buildId` overwrites it (`SyncSocket.sawHello`), so a
20
+ * client that names its build only in the frame — the documented place — is not deemed current
21
+ * forever for having sent no query.
22
+ */
17
23
  readonly clientBuildId: string;
18
24
  }
19
25
 
@@ -118,6 +124,7 @@ export async function handleUpgrade(
118
124
  if (!deps.ready() || deps.socketCount() >= deps.maxConnections) return shed(deps);
119
125
  const data: WsData = {
120
126
  socketId: deps.newSocketId(),
127
+ // The node's own id is "not skewed until the hello says so", never "current forever".
121
128
  clientBuildId: url.searchParams.get('build') ?? deps.buildId,
122
129
  };
123
130
  // Before the upgrade, never after: `server.upgrade` runs `websocket.open` synchronously and does