machine-bridge-mcp 3.0.0-beta.35 → 3.0.0-beta.38

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,37 @@
1
1
  # Changelog
2
2
 
3
+ ## 3.0.0-beta.38 - 2026-08-05
4
+
5
+ ### Keep relay liveness acknowledgement off durable storage paths
6
+
7
+ - Send the Worker `pong` immediately after the authenticated socket attachment is refreshed, before any Durable Object alarm read or write. A slow alarm/storage operation can no longer delay heartbeat acknowledgement and make an otherwise healthy connection appear silent.
8
+ - Make daemon activity refresh scheduling-explicit. Heartbeats perform one alarm schedule after `pong`, and terminal tool results coalesce liveness and pending-call deadline updates into exactly one schedule instead of the previous implicit-plus-explicit pair.
9
+ - Isolate the complete event-time alarm scheduling path, not only the final alarm write. Durable deadline reads, invalid-socket cleanup, or diagnostic callback failures now become bounded observability events instead of aborting a registered call before dispatch or rejecting a WebSocket message event. The actual Durable Object `alarm()` handler remains failure-propagating so the platform can retry it.
10
+ - Add architecture regressions that require `pong` to precede alarm scheduling, forbid socket-touch helpers from acquiring hidden alarm ownership, and require one terminal-result alarm schedule.
11
+ - Record the live beta.37 incident boundary: the same launchd daemon (PID unchanged, `runs=1`) recovered a `1006 connection_interrupted` episode in 2.631 seconds on its first attempt. macOS changed the `utun5` link-quality classification from good to poor four seconds before the close while the default route remained inside the Karing system-extension tunnel. This is strong temporal correlation with OS Wi-Fi/TUN path degradation, not proof that Karing, a selected proxy node, Cloudflare, or any specific upstream component caused the close.
12
+
13
+ ## 3.0.0-beta.37 - 2026-08-04
14
+
15
+ ### Close second-order relay recovery races
16
+
17
+ - Bind asynchronous daemon-authentication proof failures to the WebSocket generation that requested them. A rejected proof from an already closed socket can no longer terminate a replacement connection that is currently connecting or ready.
18
+ - Apply explicit close-category precedence. The first specific connect, handshake, readiness, heartbeat, or Worker recovery cause survives later specific or generic close signals; only an empty or generic transport category may be upgraded.
19
+ - Separate socket cleanup from alarm ownership. Runtime-alarm invalidation no longer recursively schedules another alarm, and the final alarm deadline is recomputed after detach/rebind state changes so reconnect grace cannot inherit a stale pre-detach deadline.
20
+ - Close invalidated sockets before awaiting durable cleanup. Concurrent `error`/`close` callbacks share one cleanup Promise; successful cleanup remains terminal, while a transient failure is retried once in-event and releases its slot for a later callback without duplicating disconnected metrics or warning logs. Welcome, readiness-probe, replacement, liveness, send-failure, error, and close paths preserve the intended close reason even when storage cleanup fails.
21
+ - Mark the authenticated relay diagnostic snapshot as recovered when a probing socket becomes ready, extend outage duration through the actual readiness instant, canonicalize timestamps, and accept only stable coarse transport error classes. Failed reconnect attempts no longer erase the duration of the preceding healthy ready interval. A healthy `server_info.daemon.relay_transport` no longer reports the preceding reconnect as currently active or exposes arbitrary daemon metadata.
22
+ - Add fault-directed regressions for stale authentication promises, competing specific close causes, send-failure precedence, post-invalidation deadline recomputation, cleanup deduplication/retry, previous-ready-duration retention, ready-state diagnostic projection, and the scheduling-free cleanup architecture contract. Beta.36 was prepared but not activated; beta.37 supersedes that local candidate.
23
+
24
+ ## 3.0.0-beta.36 - 2026-08-04
25
+
26
+ ### Preserve and expose relay-disconnect evidence
27
+
28
+ - Preserve a specific connect, handshake, readiness, or heartbeat timeout classification when a later generic WebSocket error arrives before the close event. The late error can still terminate the socket, but it no longer erases the causal category used for recovery diagnosis.
29
+ - Make Worker daemon-socket cleanup idempotent. Error, close, candidate timeout, readiness timeout, liveness timeout, verified replacement, and send-failure paths converge on one expiry, pending-call detach, disconnected metric, and runtime-alarm transition, preventing duplicate Durable Object work when one socket emits both error and close. Synchronous stale-socket reclamation now retains its asynchronous cleanup with Durable Object `waitUntil` and converts storage failures into one bounded observability event instead of an unhandled rejection.
30
+ - Add a schema-versioned, privacy-bounded relay diagnostic summary to the authenticated daemon hello. The Worker sanitizes and preserves the immediately preceding reconnect episode in the daemon attachment and exposes it as authenticated `server_info.daemon.relay_transport`; endpoints, interface names, DNS data, arguments, and results remain excluded.
31
+ - Make `machine-mcp doctor` report its diagnostic scope explicitly. Doctor uses an isolated local runtime and does not inspect the running service process or its remote relay, so a green doctor result can no longer be mistaken for service WebSocket health.
32
+ - Add deterministic regressions for late-error classification, diagnostic bounding/projection, idempotent socket expiry, unified stale-candidate invalidation, retained asynchronous cleanup, authenticated server-info projection, and doctor scope.
33
+ - Pin the transitive `brace-expansion` and `undici` packages to fixed same-major releases through root overrides. The release audit discovered high-severity advisories in ESLint/Wrangler dependency paths; `npm audit fix --force` proposed an unrelated Wrangler downgrade, so beta.36 keeps the tested Wrangler/Miniflare versions while selecting `brace-expansion` 5.0.9 and `undici` 7.29.0.
34
+
3
35
  ## 3.0.0-beta.35 - 2026-08-03
4
36
 
5
37
  ### Enforce the patch-helper call contract
@@ -30,6 +30,6 @@
30
30
  "action": {
31
31
  "default_title": "Machine Bridge Browser"
32
32
  },
33
- "version_name": "3.0.0-beta.35",
33
+ "version_name": "3.0.0-beta.38",
34
34
  "key": "MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAxryYkpZhq8+VAQLHcGS9BAHQcyKX8RHGIpIwvtIVRU/rcOcE0bNdnM0aZJ/h6xWQsGDHlhvjT2+1aJaAn/9k8473BRWajzVXld961CdHYVFVHoce2hHiSJ0xydWrHMMZhAm0mN0UzjEpgZ0tMw209efcZHIvSwuxhteZMRy4kyiVjwFlOf5oXFCxRuCJnPj3AK9CmCf4XgEBuPIJ0TZmjGHOOdBvJmbCNnAWXYEo5/mf7MfCGhV4IJ1hNuhpoNQfOFKMUcw9/v/IpT62XpfXdGYTfGYCmCjC+gntK1spbkr2P4/2+sYMQtLpse71mpSNGXfcf3abU55Vpn+gncSxRQIDAQAB"
35
35
  }
@@ -267,13 +267,13 @@ Worker-name mutation is a separate identity transition. Existing state rejects a
267
267
 
268
268
  The local `RelayConnection` treats proxy selection, transport construction, WebSocket open, authentication, end-to-end readiness, and outage recovery as separate states. The shared proxy module maps WebSocket targets to standard HTTP(S) environment-proxy resolution, honors `NO_PROXY`, rejects non-HTTP(S) proxy schemes, and creates the proxy agent without exposing its URL or credentials. Invalid proxy configuration is a fatal configuration error rather than a retryable outage.
269
269
 
270
- A connection-attempt deadline terminates sockets stuck in `CONNECTING`. After open, the daemon sends `hello`; `hello_ack` establishes an authenticated relay generation and starts heartbeats, but does not resolve startup or advertise readiness. The Worker then sends a random `relay_probe`. Its result must traverse the local runtime dispatcher and `sendForSession` on that generation before `ready_ack` marks the connection usable. The daemon rejects ordinary tool calls before that state and rejects a premature readiness acknowledgement without locally recorded probe delivery. Independent handshake and readiness deadlines terminate candidates that authenticate but cannot return results. Once ready, application heartbeats require inbound activity; a silent half-open socket is terminated and reconnected. Worker error codes `daemon_transport_error` and `daemon_liveness_timeout` are connection-recovery signals, not protocol incompatibility: they terminate only the current socket, run normal disconnect cleanup, and enter bounded reconnect. The same classification is applied to the close frame if the preceding error frame is lost. Failure to send `hello`, or failure to deliver a readiness-probe result because its socket/session ended, follows the same transport-recovery path rather than being promoted to authentication or protocol failure. Unknown Worker errors, authentication rejection, malformed readiness sequencing, and identity/version mismatch remain fatal. Outage reminders run on their own exponential-backoff timer rather than depending on another transport callback.
270
+ A connection-attempt deadline terminates sockets stuck in `CONNECTING`. After open, the daemon sends `hello`; `hello_ack` establishes an authenticated relay generation and starts heartbeats, but does not resolve startup or advertise readiness. The Worker then sends a random `relay_probe`. Its result must traverse the local runtime dispatcher and `sendForSession` on that generation before `ready_ack` marks the connection usable. The daemon rejects ordinary tool calls before that state and rejects a premature readiness acknowledgement without locally recorded probe delivery. Independent handshake and readiness deadlines terminate candidates that authenticate but cannot return results. Once ready, application heartbeats require inbound activity; a silent half-open socket is terminated and reconnected. On the Worker, authenticated socket activity refresh is synchronous and `pong` is queued before Durable Object alarm inspection or mutation; heartbeat and terminal-result paths then own one explicit coalesced alarm schedule. Persistent deadline work therefore cannot sit in front of liveness acknowledgement or be scheduled twice through a hidden touch helper. Worker error codes `daemon_transport_error` and `daemon_liveness_timeout` are connection-recovery signals, not protocol incompatibility: they terminate only the current socket, run normal disconnect cleanup, and enter bounded reconnect. The same classification is applied to the close frame if the preceding error frame is lost. Failure to send `hello`, or failure to deliver a readiness-probe result because its socket/session ended, follows the same transport-recovery path rather than being promoted to authentication or protocol failure. Unknown Worker errors, authentication rejection, malformed readiness sequencing, and identity/version mismatch remain fatal. Outage reminders run on their own exponential-backoff timer rather than depending on another transport callback. Each new authenticated hello also carries a schema-versioned, bounded summary of the immediately preceding local reconnect episode. The Worker sanitizes that summary into the probing attachment; promotion to a ready daemon marks the episode recovered, canonicalizes timestamps, and retains only enumerated coarse error classes. Authenticated `server_info.daemon.relay_transport` therefore makes sub-warning-threshold interruptions diagnosable after recovery without claiming that the ready connection is still in outage or promoting every brief close to default logs.
271
271
 
272
272
  Reconnect uses bounded exponential backoff with jitter. Brief self-healing interruptions are debug-only. An unresolved outage is promoted to a rate-limited warning after a grace period, and recovery produces one summary. Raw close codes and reason strings remain debug-only.
273
273
 
274
274
  The Worker stores socket transitions in `DaemonSocketRegistry`: `candidate` before hello, `probing` after authentication, `daemon` only after the end-to-end result probe, and `expired` after terminal failure. Durable Object alarms enforce separate hello, readiness, and steady-state liveness deadlines across hibernation. A healthy incumbent remains active while a replacement is probed; a malformed, silent, incompatible, or identity-mismatched replacement is closed without displacing it. Only a verified candidate receives `ready_ack` and then replaces the old socket. Ready daemons stay live only while inbound traffic refreshes `lastSeenAt`; silent half-open or hibernation-restored sockets are reclaimed instead of advertising `daemon.connected` while tool calls time out.
275
275
 
276
- Each daemon process generates a random bounded `instance_id` at startup and includes it in every reconnect hello. Each accepted WebSocket additionally receives a random `connection_id` generation stored in its hibernation attachment. JSON-only pending calls retain an in-memory socket reference; streamed calls persist the opaque generation instead. On an unexpected socket loss, only calls owned by that generation are detached and the shared two-minute relay contract bounds recovery. During verified same-instance handover, the Worker transfers both already-detached and still-attached calls to the replacement generation before closing the incumbent. A delayed close or result from the incumbent fails the generation check and cannot mutate the rebound record. If replacement acknowledgement fails, ownership is restored to the still-open same-instance incumbent. Another process cannot inherit or resolve these calls. The local runtime mirrors that state machine by preserving active calls and completed-result envelopes until relay readiness returns. Before `ready_ack`, the Worker sends the exact IDs that still have remote waiters; the runtime cancels everything else and only then replays retained results through the verified socket. Grace expiry restores the terminal behavior: reject remote waiters, cancel local ordinary calls, terminate process trees, and discard undeliverable results. The shared execution envelope remains independent of reconnect grace. Worker operation countdown is paused while detached or transferred and resumed with its remaining budget, so handover cannot reset the normal timeout. The active stream record expiry is extended to cover each new reconnect deadline, remaining operation budget, and terminal replay window; repeated successful reconnect cycles cannot outlive and delete their own ownership record. JavaScript timers remain only the low-latency owner for JSON-only calls. Streamed operation and reconnect deadlines live in the durable call record, the earliest deadline is projected onto one Durable Object alarm, and every HTTP or WebSocket event performs a compensating overdue scan. A transient alarm-storage failure is observable but does not turn an already-dispatched operation into a false terminal failure; the next event scan remains the bounded recovery path. This does not make calls durable across daemon-process restart or machine failure; managed jobs remain the separate durable mechanism.
276
+ Each daemon process generates a random bounded `instance_id` at startup and includes it in every reconnect hello. Each accepted WebSocket additionally receives a random `connection_id` generation stored in its hibernation attachment. JSON-only pending calls retain an in-memory socket reference; streamed calls persist the opaque generation instead. On an unexpected socket loss, only calls owned by that generation are detached and the shared two-minute relay contract bounds recovery. During verified same-instance handover, the Worker transfers both already-detached and still-attached calls to the replacement generation before closing the incumbent. A delayed close or result from the incumbent fails the generation check and cannot mutate the rebound record. If replacement acknowledgement fails, ownership is restored to the still-open same-instance incumbent. Another process cannot inherit or resolve these calls. The local runtime mirrors that state machine by preserving active calls and completed-result envelopes until relay readiness returns. Before `ready_ack`, the Worker sends the exact IDs that still have remote waiters; the runtime cancels everything else and only then replays retained results through the verified socket. Grace expiry restores the terminal behavior: reject remote waiters, cancel local ordinary calls, terminate process trees, and discard undeliverable results. The shared execution envelope remains independent of reconnect grace. Worker operation countdown is paused while detached or transferred and resumed with its remaining budget, so handover cannot reset the normal timeout. The active stream record expiry is extended to cover each new reconnect deadline, remaining operation budget, and terminal replay window; repeated successful reconnect cycles cannot outlive and delete their own ownership record. JavaScript timers remain only the low-latency owner for JSON-only calls. Streamed operation and reconnect deadlines live in the durable call record, the earliest deadline is projected onto one Durable Object alarm, and every HTTP or WebSocket event performs a compensating overdue scan. Socket invalidation invoked by the alarm coordinator does not schedule recursively. After candidate/readiness/liveness transitions and call detach, the coordinator recomputes transient and durable deadlines from the resulting ownership state and performs one coalesced alarm write. The complete event-time scheduling calculation is failure-isolated: durable deadline reads, invalid-socket transitions, and alarm storage mutations become bounded observability events and do not abort a registered operation or WebSocket event. The platform-invoked alarm processor remains failure-propagating for retry, and the next ordinary event scan is the compensating recovery path. This does not make calls durable across daemon-process restart or machine failure; managed jobs remain the separate durable mechanism.
277
277
 
278
278
  ## Persistence
279
279
 
@@ -293,9 +293,9 @@ Browser-origin handling separates CORS response sharing from protocol authentica
293
293
 
294
294
  ## Observability
295
295
 
296
- Public health exposes only server identity and version. Authenticated `server_info` exposes bounded runtime status, managed-job counts, resource alias names without paths or values, relay route state without endpoint details, authenticated/probing/ready socket counts, end-to-end readiness evidence, local execution guardrails, explicit OS-enforcement gaps, and privacy-preserving capability-routing evidence. It separates the daemon capability ceiling from the authenticated account authority: `daemon.policy`/`daemon.tools` retain the pre-role ceiling, while `authorization.effective_policy`/`authorization.effective_tools` and the top-level `tools` report the role-intersected authority before any host-side filtering. It explicitly reports that the host-exposed subset is unknown to the server. The canonical MCP catalog advertises one foreground timeout contract of 1–60 seconds with tool-specific 30- or 60-second defaults, and the Worker rejects larger values before any daemon message is sent. `tool-timeout.ts` derives distinct daemon-execution and Worker-settlement deadlines, so the five-second settlement-deadline offset is not passed back to the daemon as additional execution time. `foreground-timeout.mjs` is the shared source for execution defaults and limits; `process-foreground-timeout.mjs` applies them again at the local relay execution boundary so omitted values and registered-command manifests cannot outlive the Worker response. Owner-local registered commands may still use their explicit local manifest timeout. Longer remote work uses process sessions or managed jobs rather than a synchronous foreground response. `diagnose_runtime` runs fixed local probes, explicitly reports that its own request reached the daemon, and on macOS projects the default route into a coarse VPN/TUN interception class without returning interface or endpoint data.
296
+ Public health exposes only server identity and version. Authenticated `server_info` exposes bounded runtime status, managed-job counts, resource alias names without paths or values, the current daemon connection plus its sanitized preceding-reconnect summary without endpoint details, authenticated/probing/ready socket counts, end-to-end readiness evidence, local execution guardrails, explicit OS-enforcement gaps, and privacy-preserving capability-routing evidence. It separates the daemon capability ceiling from the authenticated account authority: `daemon.policy`/`daemon.tools` retain the pre-role ceiling, while `authorization.effective_policy`/`authorization.effective_tools` and the top-level `tools` report the role-intersected authority before any host-side filtering. It explicitly reports that the host-exposed subset is unknown to the server. The canonical MCP catalog advertises one foreground timeout contract of 1–60 seconds with tool-specific 30- or 60-second defaults, and the Worker rejects larger values before any daemon message is sent. `tool-timeout.ts` derives distinct daemon-execution and Worker-settlement deadlines, so the five-second settlement-deadline offset is not passed back to the daemon as additional execution time. `foreground-timeout.mjs` is the shared source for execution defaults and limits; `process-foreground-timeout.mjs` applies them again at the local relay execution boundary so omitted values and registered-command manifests cannot outlive the Worker response. Owner-local registered commands may still use their explicit local manifest timeout. Longer remote work uses process sessions or managed jobs rather than a synchronous foreground response. `diagnose_runtime` runs fixed local probes, explicitly reports that its own request reached the daemon, and on macOS projects the default route into a coarse VPN/TUN interception class without returning interface or endpoint data.
297
297
 
298
- Foreground logging defaults to `info`; autostart uses `warn`. Authenticated readiness, persistent degradation, and recovery are user-visible state transitions. Brief relay interruptions, raw transport close details, retry timing, and all per-tool starts/successes/failures/cancellations/durations are debug-only. Unexpected local and Worker infrastructure errors are reduced to classes. Messages, strings, arrays, object depth/key counts, and serialized fields are bounded.
298
+ Foreground logging defaults to `info`; autostart uses `warn`. Authenticated readiness, persistent degradation, and recovery are user-visible state transitions. Brief relay interruptions, raw transport close details, retry timing, and all per-tool starts/successes/failures/cancellations/durations remain debug-only in logs; the bounded preceding-reconnect summary is retained in authenticated `server_info` instead. Worker socket invalidation is single-owner: `error`, `close`, heartbeat, readiness, candidate expiry, replacement, and send-failure paths converge on one idempotent expiry and detach transition. Concurrent callbacks share one per-socket cleanup Promise; successful settlement remains cached for that socket, while a failed cleanup receives one immediate retry and then releases the slot for a later lifecycle callback. Only the first claim records disconnected metrics or a WebSocket warning. Transport closure begins before durable cleanup is awaited; alarm ownership remains with the initiating event or alarm coordinator rather than the cleanup primitive. Synchronous stale-socket detection starts expiry immediately but retains the asynchronous detach/storage work with Durable Object `waitUntil`; failures become one bounded error-class event rather than an unhandled Promise rejection. Unexpected local and Worker infrastructure errors are reduced to classes. Messages, strings, arrays, object depth/key counts, and serialized fields are bounded.
299
299
 
300
300
  Cloudflare sampling is size control rather than an audit log. The project intentionally does not claim complete forensic logging. See [LOGGING.md](LOGGING.md).
301
301
 
package/docs/AUDIT.md CHANGED
@@ -1,5 +1,33 @@
1
1
  # Security and privacy audit notes
2
2
 
3
+ ## 2026-08-05 version 3.0.0-beta.38 live relay-liveness review
4
+
5
+ The activated beta.37 daemon experienced one real short transport interruption at 2026-08-05 07:38:57 local time (`2026-08-04T23:38:57.893Z`). Authenticated `server_info.daemon.relay_transport` retained close code 1006, category `connection_interrupted`, one reconnect attempt, a 2,631-millisecond outage, and a preceding ready interval of 550,090 milliseconds. The socket returned to ready state with `outage_active=false`. The launchd process remained PID 73042 with `runs=1` and `last exit code = never exited`, excluding daemon crash or service restart as the immediate cause. No stale-connection result was accepted; one `owner_missing_acknowledged` result reflected the documented at-least-once terminal replay case in which the owner had already settled and the duplicate was acknowledged to stop replay.
6
+
7
+ macOS network evidence is temporally specific but not causally complete. The default route was carried by `utun5`. At 07:38:53, four seconds before the WebSocket close, several system connection monitors changed the `utun5` link-quality classification from good to poor. The tunnel was owned by the Karing system extension (`com.nebula.karing.karingServiceSE`), and multiple tunnel-carried connections closed or reopened in the same window. Neither Karing's unified-log stream nor readable recent product logs supplied a node-change or upstream-failure record. The defensible conclusion is therefore OS Wi-Fi/TUN path degradation correlated with the close; the evidence does not prove a Karing implementation defect, a proxy-node switch, Cloudflare failure, DNS/Fake-IP behavior, or another specific upstream trigger.
8
+
9
+ The fresh source review nevertheless found a Worker-side amplification hazard. For every heartbeat, `BridgeRoom` refreshed the hibernation attachment and awaited Durable Object alarm inspection/mutation before sending `pong`. Alarm storage is normally fast, but it is an unnecessary dependency in the liveness acknowledgement path: latency or contention there could make a healthy peer appear silent. The same helper implicitly scheduled an alarm for tool results, after which matched results scheduled it a second time. Beta.38 makes attachment touch synchronous, emits `pong` first, schedules once afterward, and gives terminal results one explicit coalesced schedule. A second fault-directed pass found that the scheduling helper still allowed pre-write failures—such as a durable deadline read or invalid-socket cleanup failure—to escape. That could reject a WebSocket event after `pong`, or abort a newly registered foreground call before its daemon dispatch. Event-time scheduling now catches the complete calculation/transition/write path and reports only a bounded error class; the platform-invoked `alarm()` processor still propagates failures for retry. Architecture tests freeze heartbeat ordering and ownership, while runtime tests inject both storage-write and durable-deadline-read failures. These corrections improve resistance to Worker-side delay and storage faults; they cannot prevent a TCP/WebSocket interruption caused below the application layer.
10
+
11
+ ## 2026-08-04 version 3.0.0-beta.37 second-order relay review
12
+
13
+ A fresh review of the unactivated beta.36 candidate did not reuse its green suite as proof. It found six second-order races. An asynchronous authentication proof created for an old socket could reject after reconnect and invoke the fatal path against the replacement. A send failure and two independent specific timeout sources could overwrite an earlier, more causal close category. Moving all invalidation through the idempotent cleanup owner caused runtime-alarm callbacks to enter alarm scheduling recursively. Removing that recursion without recalculating pending state would in turn retain a deadline computed before socket detach changed operation ownership. The hello-time diagnostic snapshot still had `outage_active=true` when stored on a ready socket. Finally, the Worker bounded but did not enumerate the claimed coarse transport error class, allowing arbitrary authenticated-daemon metadata into `server_info`.
14
+
15
+ Beta.37 binds authentication completion to the originating socket, applies a specificity-preserving close-category function, and makes alarm ownership explicit. Runtime-alarm callers invalidate with rescheduling disabled; after all socket transitions, the coordinator recomputes transient and durable deadlines from the resulting state and performs one coalesced alarm write. Socket closure starts before durable detach is awaited, so a storage exception cannot leave the invalidated transport open. A per-socket cleanup registry shares concurrent callback work, retains successful settlement, retries one transient failure in-event, and allows a later callback to retry persistent failure without duplicating disconnected metrics or warnings. Promotion from `probing` to `daemon` converts the reconnect snapshot to recovered state, extends its duration through actual readiness, canonicalizes timestamps, and restricts transport error classes to the stable operational vocabulary. Failed reconnect candidates preserve the prior healthy-ready duration rather than resetting it.
16
+
17
+ New tests reproduce an old-proof rejection after a replacement socket opens, a late send failure plus Worker liveness signal after a readiness timeout, a pending deadline that becomes earlier during candidate invalidation, concurrent cleanup plus persistent-failure retry, failed reconnect attempts after a healthy connection, and the exact ready-attachment diagnostic transition. Architecture tests forbid runtime-alarm invalidation from re-entering alarm scheduling and require the cleanup owner to remain scheduling-free. The existing beta.36 tarball is stale by promotion digest and was never activated, accepted, published, tagged, or deployed; beta.37 requires a new exact candidate and live owner verification.
18
+
19
+ ## 2026-08-04 version 3.0.0-beta.36 relay-disconnect review
20
+
21
+ The reported disconnect did not coincide with a daemon restart. Launchd still owned the same process that began at 08:02:02 local time, while the Worker-facing connection identity was established again at 19:36:29. During final verification the same PID and `runs=1` service established another connection at 20:36:53 without a warning-level relay event, proving a second sub-ten-second same-process interruption. The daemon warning log contained earlier 1006 and transport-timeout outages but no record for either brief reconnect because the background service runs at `warn` and suppresses recovered interruptions shorter than ten seconds. The active socket used the operating-system network stack through a TUN route; Karing and Tailscale packet-tunnel components were both present. macOS recorded a Tailscale network-configuration change at 19:32. Around 20:34, shortly before the second reconnect, the system reported Wi-Fi link quality changing from Good to Poor, Tailscale closing and reopening several DERP paths, and repeated Karing network-path checks. This repeated correlation strengthens the system-network/TUN-path churn hypothesis, but no source proved which tunnel, upstream node, DNS/Fake-IP mapping, Worker edge, or TCP event caused either reconnect. The exact external trigger therefore remains unknown and is not represented as a code-proven Karing or Tailscale defect.
22
+
23
+ Source review found four internal defects that made this class of incident harder or more expensive to handle. First, a timeout path could set a specific close category and call `terminate()`, after which a generic WebSocket `error` event overwrote it with `relay_transport_error` before `close`. Second, Cloudflare can deliver both `webSocketError` and `webSocketClose` for one socket; both entered cleanup, repeating disconnected metrics, pending detach attempts, durable storage transactions, and alarm scheduling. Several candidate/readiness invalidation branches also expired or closed a socket outside that cleanup owner. Third, synchronous stale-socket reclamation discarded the asynchronous invalidation Promise with `void`; a Durable Object storage failure could therefore escape as an unhandled rejection and the cleanup was not retained by the event lifecycle. Fourth, the daemon already retained bounded outage state in memory, but a brief recovered interruption was unavailable from the authenticated remote `server_info`; `machine-mcp doctor` created a separate non-relay runtime and still displayed its skipped remote-relay check inside an otherwise green result.
24
+
25
+ Beta.36 keeps the first specific close category, makes attachment expiry return a one-time ownership claim, and routes every Worker invalidation through one cleanup method before closure. Stale-socket cleanup is retained through Durable Object `waitUntil`; its rejection is reduced to the coarse `daemon.socket.cleanup.failed` error class and cannot create a second rejection if observability itself fails. The next authenticated hello carries only schema 1 bounded fields: coarse application network route, outage count/active state/start/duration/attempts, close category/code, coarse transport error class, last disconnect time, and previous ready duration. The Worker validates enums, timestamps, integer ranges, metadata lengths, and schema before storing the summary in its hibernation-safe socket attachment. Only authenticated `server_info.daemon.relay_transport` returns it. Public health and default logs remain unchanged. Doctor now reports `running_service_process_inspected=false` and `remote_relay_inspected=false` and directs operators to authenticated server info.
26
+
27
+ The release dependency audit then found two newly disclosed high-severity transitive issues: `brace-expansion` 5.0.8 through ESLint/minimatch and `undici` 7.28.0 through Wrangler/Miniflare. The ordinary fix safely advanced `brace-expansion` to 5.0.9, while npm's forced remediation proposed an unrelated Wrangler downgrade even though `undici` 7.29.0 is the fixed 7.x release. Beta.36 therefore retains the already tested Wrangler 4.115.0 and Miniflare 4.20260722.1 packages but adds deterministic root overrides for `brace-expansion` 5.0.9 and `undici` 7.29.0. Complete and production dependency audits now report zero vulnerabilities, all 117 registry package signatures verify, and 38 packages carry verified attestations.
28
+
29
+ The existing 120-second same-daemon call recovery remains the continuity boundary. These changes improve causal evidence and remove duplicate Worker work; they do not guarantee that an external MCP host will keep its foreground HTTP/SSE request open during every network transition, and they cannot stabilize a third-party VPN/TUN or edge path. Long or cleanup-sensitive work still belongs in process sessions or managed jobs.
30
+
3
31
  ## 2026-08-03 version 3.0.0-beta.31 host-delivery margin review
4
32
 
5
33
  The reported “message send timed out” interruption did not coincide with a daemon crash or a current relay outage. Launchd still owned one verified beta.30 daemon process with `runs=1`; Worker and daemon versions matched; and the local security-audit chain recorded a temporally aligned `exec_command` as successfully completed after 83,514 milliseconds. During the incident the Worker showed two durable `exec_command` calls still active, the oldest at roughly 81 seconds. Both later reached terminal state, while the host ended the task. The privacy-preserving audit deliberately omits raw command text, so an exact one-to-one mapping to the UI task cannot be proven; the timestamps and active-call counts nevertheless align. This separates execution completion from message delivery: persistence can preserve a legacy result, but it cannot force a host that has abandoned the response to resume it.
package/docs/LOGGING.md CHANGED
@@ -63,9 +63,10 @@ Brief network interruptions are expected on laptop network changes, Worker deplo
63
63
  - a verified replacement is a distinct warning and permanently stops the older daemon;
64
64
  - failure to receive `hello_ack` within the handshake deadline, or `ready_ack` within the independent end-to-end readiness deadline, terminates the candidate socket and retries;
65
65
  - lack of inbound heartbeat activity terminates a half-open socket and reconnects;
66
+ - the Worker refreshes authenticated socket activity and queues `pong` before Durable Object alarm inspection or mutation, then performs one explicit coalesced schedule; storage latency is not allowed to sit ahead of heartbeat acknowledgement;
66
67
  - a late local heartbeat tick is classified as `runtime.event_loop.stall`, sends a fresh probe, and defers disconnect for a bounded recovery interval instead of being mislabeled as immediate remote failure; a macOS sleep/wake interval may legitimately produce this warning without a daemon fault.
67
68
 
68
- A WebSocket close code such as `1006` means the transport ended without a normal close handshake. It is useful for debug diagnosis but not useful as the default user message. It is not evidence that the daemon process restarted. Worker `daemon_transport_error` / `daemon_liveness_timeout` messages and their 1012 close frames are likewise retryable connection conditions, not upgrade instructions. Only an unknown/incompatible Worker error, authentication failure, or identity/version mismatch may produce the fatal protocol/configuration log and daemon exit. Default logs therefore describe the affected layer, duration, classification, and recovery behavior rather than printing raw close envelopes.
69
+ A WebSocket close code such as `1006` means the transport ended without a normal close handshake. If it recovers inside ten seconds, the warning-level service log is intentionally silent and the authenticated `daemon.relay_transport` snapshot is the post-event evidence surface. It is useful for debug diagnosis but not useful as the default user message. It is not evidence that the daemon process restarted. Worker `daemon_transport_error` / `daemon_liveness_timeout` messages and their 1012 close frames are likewise retryable connection conditions, not upgrade instructions. Only an unknown/incompatible Worker error, authentication failure, or identity/version mismatch may produce the fatal protocol/configuration log and daemon exit. Default logs therefore describe the affected layer, duration, classification, and recovery behavior rather than printing raw close envelopes.
69
70
 
70
71
  Streamed-call diagnostics are deliberately coarse. Modern request-scoped stream ownership is memory-only; legacy MCP `2025-11-25` may additionally report aggregate persistent active/detached counts, oldest age, tool-name counts, alarm mutations, unmatched-result counts, opened/coexisting/limited delivery-subscriber counts, terminal publications, live internal-subscriber sends, storage responses, and storage-race sends/failures. These counters describe legacy resumable Worker-internal storage and subscription transport; they do not prove public SSE consumption or MCP-host receipt. Worker event counters are scoped to the current isolate and say so in `metric_scope`; persisted durable calls can begin in one isolate and complete in another, so `started`, `completed`, and `failed` are not algebraically closed process-lifetime totals. The persistent pending-call snapshot is authoritative for current ownership. Logs and `server_info` must not include tool arguments, terminal results, command text, request keys, account identifiers, raw call IDs, raw connection generations, mirrored parameter values, private paths, or subscriber payloads. A stale-generation result is counted as unmatched rather than logged with its envelope.
71
72
 
@@ -126,7 +127,7 @@ This is defense in depth, not content classification. Unknown, split, transforme
126
127
 
127
128
  Security-audit enqueue/persistence failures are warning-level operational faults, but repeated failures are rate-limited per event class. The first warning is emitted immediately; duplicates within one minute are suppressed, and the next emitted warning reports the suppressed count. Warning fields contain only the tool name and a coarse error class, never the rejected audit payload or principal identifiers.
128
129
 
129
- The local execution middleware emits bounded events such as `tool.call.started`, `tool.call.completed`, `tool.call.failed`, `tool.call.slow`, and `tool.call.cancel_requested`. Stable fields include a shortened call ID, tool name, origin, duration, error code, and retryability. The Worker emits JSON events for HTTP failures and daemon socket errors. Structured values still pass through field-name and value redaction; JSON format is not permission to log arguments or results.
130
+ The local execution middleware emits bounded events such as `tool.call.started`, `tool.call.completed`, `tool.call.failed`, `tool.call.slow`, and `tool.call.cancel_requested`. Stable fields include a shortened call ID, tool name, origin, duration, error code, and retryability. The Worker emits JSON events for HTTP failures and daemon socket errors. Structured values still pass through field-name and value redaction; JSON format is not permission to log arguments or results. A retained stale-socket cleanup failure uses `daemon.socket.cleanup.failed` with only `error_class`; it contains no socket identity, endpoint, arguments, or results.
130
131
 
131
132
  `server_info` is the operational metrics surface. Local metrics include lifecycle state, active and maximum calls, oldest-call age, active-process ownership, per-tool duration buckets, and error-code counts. Worker metrics include HTTP status classes, pending internal/request-key indexes, per-tool outcomes, daemon candidate/authenticated/ready/disconnected event counts, current authenticated/probing/ready socket counts, and protocol-error counts. Metrics contain counts and bounded identifiers, not request arguments or result contents.
132
133
 
@@ -149,7 +150,7 @@ Each managed job has owner-only runner diagnostic logs. Child-step output is ret
149
150
 
150
151
  `network_route` describes only Machine Bridge's application-level proxy decision. `system-network-stack` does **not** mean a direct physical path: an operating-system VPN, TUN, packet tunnel, DNS interceptor, or endpoint-security product may still carry the connection. `network_route_scope` therefore remains `application-proxy-selection-only`.
151
152
 
152
- During an outage, remote `diagnose_runtime.runtime.relay` and local stdio `server_info.runtime.relay` expose bounded operational fields: outage count/start/duration, last close category/code, coarse transport error class, last disconnect/ready time, prior ready duration, and next retry timing. On macOS, `diagnose_runtime` may also return a coarse default-route class and `operating_system_interception` boolean. That diagnostic is returned on demand and is not promoted to default logs; interface names, IP addresses, DNS answers, proxy endpoints/credentials, Worker endpoints, tool arguments, and results remain absent. `relay.outage.active` and `relay.outage.recovered` carry the existing safe relay fields.
153
+ During an outage, remote `diagnose_runtime.runtime.relay` and local stdio `server_info.runtime.relay` expose bounded live fields: outage count/start/duration, attempts, last close category/code, coarse transport error class, last disconnect/ready time, prior ready duration, and next retry timing. After recovery, authenticated remote `server_info.daemon.relay_transport` retains the bounded preceding episode that the daemon supplied during the current connection handshake, including brief interruptions below the default warning threshold. Promotion to a ready socket sets `outage_active=false`, extends the duration through actual readiness, preserves the preceding healthy-ready duration across failed candidates, canonicalizes timestamps, and accepts only enumerated coarse operational error classes; it does not claim that the recovered connection remains in outage. On macOS, `diagnose_runtime` may also return a coarse default-route class and `operating_system_interception` boolean. That diagnostic is returned on demand and is not promoted to default logs; interface names, IP addresses, DNS answers, proxy endpoints/credentials, Worker endpoints, tool arguments, and results remain absent. `relay.outage.active` and `relay.outage.recovered` carry the existing safe relay fields.
153
154
 
154
155
  Schema 4 is strict NDJSON. Before daemon startup, both active log files are opened as owner-only regular single-link files. A schema change clears both only after validation and commits the marker only after the transition succeeds. A symlink, multiple-hard-link inode, permission error, or marker-write failure blocks startup rather than mixing formats or repeatedly erasing evidence.
155
156
 
@@ -157,7 +158,7 @@ Schema 4 is strict NDJSON. Before daemon startup, both active log files are open
157
158
 
158
159
  Canonical `full` does not remove tools based on filenames. For remote execution, the operation classifier treats credential-sensitive paths and persistence targets as hard authorization boundaries: delegated roles are denied, while owner requests remain risk-classified and audited. An MCP host, connector, model provider, desktop application, operating system, or endpoint-security layer may independently reject a request before it reaches Machine Bridge.
159
160
 
160
- Use `server_info`, `project_overview`, `machine-mcp status`, `machine-mcp doctor`, and `diagnose_runtime` to distinguish local policy from host-side enforcement. Capability-routing status is returned on demand rather than written as task logs. The in-memory observer stores a runtime-keyed task fingerprint, selected skill/match counts, recommended tool names, primary route, ambiguity class, and score gap; it never stores raw task text, static instruction content, application inventory, or route explanations. Changing the Machine Bridge profile cannot override another layer.
161
+ Use `server_info`, `project_overview`, `machine-mcp status`, `machine-mcp doctor`, and `diagnose_runtime` to distinguish local policy from host-side enforcement. `doctor` explicitly reports that its isolated runtime did not inspect the running service process or remote relay; it must not be used as a substitute for `server_info.daemon.relay_transport`. Capability-routing status is returned on demand rather than written as task logs. The in-memory observer stores a runtime-keyed task fingerprint, selected skill/match counts, recommended tool names, primary route, ambiguity class, and score gap; it never stores raw task text, static instruction content, application inventory, or route explanations. Changing the Machine Bridge profile cannot override another layer.
161
162
 
162
163
  ## Adding or changing logs
163
164
 
@@ -8,7 +8,7 @@ machine-mcp doctor
8
8
  machine-mcp service status
9
9
  ```
10
10
 
11
- `status` prints redacted profile state and verifies the deployed Worker version. Resource source paths remain redacted. `doctor` checks Node.js, the package-installed Wrangler binary, Cloudflare login, Worker health, the configured policy, the automatic-without-per-operation-prompts authorization model, and the same fixed local filesystem/process/shell/job-storage/resource probes exposed by `diagnose_runtime`. Authenticated `server_info.authorization.execution_model` reports the same contract and identifies whether the account has daemon-OS-user ambient authority. Public `/healthz` output contains only server identity and version; daemon details require an authenticated `server_info` call.
11
+ `status` prints redacted profile state and verifies the deployed Worker version. Resource source paths remain redacted. `doctor` checks Node.js, the package-installed Wrangler binary, Cloudflare login, Worker health, the configured policy, the automatic-without-per-operation-prompts authorization model, and the same fixed local filesystem/process/shell/job-storage/resource probes exposed by `diagnose_runtime`. It constructs an isolated local runtime: `diagnosticScope.running_service_process_inspected=false` and `remote_relay_inspected=false` are deliberate, so a green doctor result is not evidence that the launchd/systemd/Scheduled Task daemon retained its Worker WebSocket. Inspect authenticated `server_info.daemon.relay_transport` for the running service relay. Authenticated `server_info.authorization.execution_model` reports the authority contract and identifies whether the account has daemon-OS-user ambient authority. Public `/healthz` output contains only server identity and version; daemon details require an authenticated `server_info` call.
12
12
 
13
13
  ### Worker deployment and health convergence
14
14
 
@@ -61,7 +61,7 @@ A successful diagnostic result applies only to that probe. An MCP host can still
61
61
 
62
62
  Machine Bridge supports concurrent calls: the Worker admits 32 pending daemon calls (30 ordinary plus two reserved control calls), and the local runtime admits 16 active tool calls (14 ordinary plus two reserved control calls). The same `diagnose_runtime`/`list_roots` control set is enforced at both layers. These are capacity limits, not a single global execution queue. Modern MCP `2026-07-28` HTTP requests are independent: JSON-RPC IDs are scoped to each request/response stream, so separate clients may reuse the same numeric ID even when they share one OAuth account and token. Legacy MCP `2025-11-25` initialization still receives a signed `Mcp-Session-Id`; idempotency, explicit cancellation, and replay for that compatibility path remain session-scoped. Within the bounded two-minute recovery window, a typed request ID denotes one operation and must not be intentionally reused for new work.
63
63
 
64
- `server_info.worker.pending_calls` reports `active`, `detached`, `request_keys`, `maximum`, `oldest_ms`, `by_tool`, `transient`, and `durable_streams`. `worker.sockets_live` separately reports `authenticated`, `probing`, `ready`, and `candidates`; only `ready` sockets contribute to `daemon.connected` and `authorization.effective_tools`. A nonzero `active` count means work is in flight, not that the bridge is globally locked. `detached > 0` means the daemon WebSocket was lost and calls are inside the bounded same-daemon reconnect interval. This relay-layer state exists below both MCP eras.
64
+ `server_info.worker.pending_calls` reports `active`, `detached`, `request_keys`, `maximum`, `oldest_ms`, `by_tool`, `transient`, and `durable_streams`. `worker.sockets_live` separately reports `authenticated`, `probing`, `ready`, and `candidates`; only `ready` sockets contribute to `daemon.connected` and `authorization.effective_tools`. `daemon.relay_transport` is the bounded, daemon-supplied summary captured during the current connection handshake: it records the immediately preceding reconnect episode without endpoints, interface names, arguments, or results. A nonzero `active` count means work is in flight, not that the bridge is globally locked. `detached > 0` means the daemon WebSocket was lost and calls are inside the bounded same-daemon reconnect interval. This relay-layer state exists below both MCP eras.
65
65
 
66
66
  For modern MCP `2026-07-28`, the public response stream is the request owner: closing it cancels the transient pending call, and no request-key or replay record should remain. For legacy MCP `2025-11-25`, the signed session and typed JSON-RPC ID own bounded idempotency and explicit `notifications/cancelled`; closing a resumable public stream alone does not cancel the operation. Legacy terminal completion, explicit cancellation, timeout, or reconnect-grace expiry must eventually return active/detached/pending-call request-key counts to zero, while the separate stream-level replay identity may remain until the two-minute recovery record expires. A verified same-daemon replacement may reclaim detached relay calls after readiness, while a new daemon process cannot. Delayed results from the old socket are rejected. `detached > 0` materially beyond the two-minute grace, a modern transient call surviving response closure, or a legacy request-key count remaining after active calls reach zero is a lifecycle defect.
67
67
 
@@ -77,13 +77,13 @@ The daemon-to-Worker terminal protocol is at-least-once until `tool_result_ack`.
77
77
 
78
78
  An error naming an internal shard mapper, temporary keyspace, backfill store, connector database, or host-side cache is not automatically a Machine Bridge Worker or daemon error. Check whether the exact text appears in repository source, Worker events, daemon logs, or local process output, and whether Worker `requests.server_error` increased. If even `server_info` fails before reaching the Worker while local readiness remains healthy, preserve credentials and state; report the host/connector incident separately rather than rotating OAuth/device secrets or redeploying blindly.
79
79
 
80
- After the host path recovers, compare `server_info`, `machine-mcp doctor`, and `machine-mcp service status`: ready socket count, pending age, daemon PID/start time, relay outage fields, and local logs. A host-storage incident and a genuine stale pending call can coexist; investigate the latter independently if it exceeds its operation or reconnect deadline.
80
+ After the host path recovers, compare authenticated `server_info`, `machine-mcp doctor`, and `machine-mcp service status`: ready socket count, pending age, daemon PID/start time, `daemon.relay_transport`, and local logs. Treat doctor as local dependency/probe evidence only; it does not inspect the running service relay. A host-storage incident and a genuine stale pending call can coexist; investigate the latter independently if it exceeds its operation or reconnect deadline.
81
81
 
82
82
  ### Relay interruption messages
83
83
 
84
- A reconnect warning proves a transport interruption, not a daemon crash. Compare daemon PID and process start time with `connected_at`, `last_seen_at`, relay close category/code, outage count, and the coarse system-route diagnostic. A VPN/TUN UI may remain “connected” while its upstream route is unusable. Machine Bridge reports only coarse route/proxy classes and never logs interface names, addresses, DNS answers, proxy credentials, or Worker secrets.
84
+ A reconnect warning proves a transport interruption, not a daemon crash. Compare daemon PID and process start time with `connected_at`, `last_seen_at`, `daemon.relay_transport.last_close_category`, `last_close_code`, `outage_count`, `outage_attempts`, and the coarse network-route class. A VPN/TUN UI may remain “connected” while its upstream route is unusable. Local OS logs can be compared with the exact `last_disconnected_at` timestamp, but an interface-quality change or tunnel-process correlation is not by itself proof of which product, node, edge, or upstream failed. Machine Bridge reports only coarse route/proxy classes and never sends or logs interface names, addresses, DNS answers, proxy credentials, or Worker secrets.
85
85
 
86
- Brief retryable outages reconnect automatically. A persistent outage emits bounded summaries; identity/version mismatch, authentication rejection, and unexpected protocol messages remain permanent failures requiring version convergence or credential repair. Compare outage intervals with sleep/wake records and `diagnose_runtime.runtime.relay.heartbeat` before classifying them as active network faults; local stdio `server_info.runtime.relay.heartbeat` exposes the same state. A nonzero `event_loop_stall_count` with a large `max_event_loop_lag_ms` means the local daemon was not scheduled promptly; during recovery grace it sends a new heartbeat and deliberately postpones disconnect. That is distinct from a relay that remains silent after local scheduling has recovered. Use `--verbose` only when close codes, heartbeat deadlines, and retry delays are required.
86
+ Brief retryable outages reconnect automatically. On a ready beta.37-or-newer socket, `server_info.daemon.relay_transport.outage_active=false`; the remaining fields describe the immediately preceding reconnect episode rather than a current outage. Because the login service logs at `warn`, a recovered interruption shorter than ten seconds intentionally has no default log line; inspect the authenticated relay snapshot instead. A persistent outage emits bounded summaries; identity/version mismatch, authentication rejection, and unexpected protocol messages remain permanent failures requiring version convergence or credential repair. Compare outage intervals with sleep/wake records and `diagnose_runtime.runtime.relay.heartbeat` before classifying them as active network faults; local stdio `server_info.runtime.relay.heartbeat` exposes the same state. A nonzero `event_loop_stall_count` with a large `max_event_loop_lag_ms` means the local daemon was not scheduled promptly; during recovery grace it sends a new heartbeat and deliberately postpones disconnect. That is distinct from a relay that remains silent after local scheduling has recovered. Use `--verbose` only when close codes, heartbeat deadlines, and retry delays are required.
87
87
 
88
88
  A foreground MCP tool is not a durable job. Every advertised MCP surface accepts at most 60 seconds of daemon execution. The Worker records a settlement deadline five seconds later than the daemon execution duration for result acceptance, persistence, acknowledgement, and terminal settlement. Admission and transport latency may consume part of that interval, and it is not a guarantee that an external host will consume the final frame. Relay execution applies the same 30- or 60-second default when the field is omitted, and a registered-command manifest cannot extend a relay call past 60 seconds. An owner-local registered command may retain a longer explicit manifest timeout because it does not depend on a hosted response stream. Longer remote work belongs in `start_process` plus bounded `read_process`, or in a managed job. Keep mutation and verification in independently terminal calls when a host exposes only a foreground shell tool.
89
89
 
package/docs/TESTING.md CHANGED
@@ -78,7 +78,7 @@ The suite includes:
78
78
  - guarded state-root removal, unsafe state-root/workspace overlap rejection before creation, all-profile lock/daemon scanning, strict current-schema validation, corrupt-JSON isolation, and policy-origin persistence;
79
79
  - no filename-based sensitive-file denial under unrestricted policy;
80
80
  - shared local/Worker free-form log redaction, sensitive content under non-sensitive Worker keys, immutable local/Worker structured metadata, control-character handling, message/field bounds, suppression of both successful and failed per-tool events outside debug, service warning-level configuration, JSON-mode parity across event and direct logger methods with timestamp/stream/redaction assertions, current-schema reset, and bounded tail trimming;
81
- - deterministic relay connection lifecycle coverage for transport construction/error/deadline, failed `hello` delivery, pre-handshake `welcome` validation, separate `hello_ack` authentication and `ready_ack` end-to-end readiness, session-bound probe return and probe-delivery races, pre-ready tool rejection, premature-ready rejection, identity/version mismatch, retryable Worker hello/readiness/transport/liveness errors, retryable close-only transport/liveness delivery, fatal unknown protocol errors, autonomous outage-reminder backoff, handshake/readiness/heartbeat timeout, brief-outage suppression, sustained-outage escalation, recovery summaries, and supersession;
81
+ - deterministic relay connection lifecycle coverage for transport construction/error/deadline, failed `hello` delivery, pre-handshake `welcome` validation, separate `hello_ack` authentication and `ready_ack` end-to-end readiness, session-bound probe return and probe-delivery races, pre-ready tool rejection, premature-ready rejection, identity/version mismatch, retryable Worker hello/readiness/transport/liveness errors, retryable close-only transport/liveness delivery, preservation of a specific timeout classification when a late generic error arrives, fatal unknown protocol errors, autonomous outage-reminder backoff, handshake/readiness/heartbeat timeout, brief-outage suppression, sustained-outage escalation, recovery summaries, and supersession;
82
82
  - shared no-follow bounded-file reads for normal files, over-limit data, directories, symbolic links, and multiple-hard-link denial; typed file-mutation regressions cover create-only collisions, stale SHA-256 preconditions, missing/ambiguous edit text, malformed/stale patches, transactional rollback, Worker preservation, stdio projection, and path/content/hash non-disclosure;
83
83
  - owner-only directory enforcement rejecting final symlinks, failing closed on POSIX chmod errors, verifying `0700`, and retaining Windows portability; Worker temporary-secret lifecycle coverage for process-start-bound names, valid stale-owner reclamation, ambiguous-owner retention, `0600` mode, deletion failures, and simultaneous deployment/cleanup failures;
84
84
  - SARIF security-gate behavior for unknown findings, exact accepted rule/path matches, path mismatch rejection, rationale quality, and exception expiry;
@@ -115,7 +115,7 @@ For deterministic release validation, perform an isolated-profile smoke test wit
115
115
  `npm run control-plane-resilience:test` is the focused accident-regression gate. It exercises synchronous/asynchronous audit failures, cached-state invalidation after external alteration, count- and byte-bounded retention anchoring, POSIX/Windows process-tree fallbacks, escalation-supervisor exception isolation, mixed transient/durable Worker capacity, and the shared 30+2 / 14+2 control-plane admission contract. `npm run security-audit:test` additionally runs two independent audit workers against one owner-only state file and requires continuous sequence numbers with no lost events, covering the lock release/acquire race. Both are part of the fast plan rather than coverage-only evidence.
116
116
 
117
117
  - control-plane resilience under host pressure: local event-loop stalls versus genuine relay silence, fresh-heartbeat recovery grace, asynchronous process-group identity capture before `SIGTERM`, bounded post-signal revalidation, draining-process visibility after result settlement, two reserved diagnostic slots at both Worker and local layers under mixed transient/durable ordinary-call saturation, non-blocking audit dispatch, batched Worker persistence, queue/drop health, warning suppression, and privacy-safe audit projection;
118
- - relay outage diagnostics and recovery: application-proxy versus OS-network scope, timestamped close/outage/recovery fields, fifteen-second maximum reconnect delay, heartbeat timeout, same-instance call continuation, Worker pong/welcome send failure, and `diagnose_runtime` relay history;
118
+ - relay outage diagnostics and recovery: application-proxy versus OS-network scope, timestamped close/outage/recovery fields, bounded authenticated hello projection and recovered ready-state semantics in `server_info.daemon.relay_transport`, fifteen-second maximum reconnect delay, socket-generation-bound authentication proofs, specificity-preserving close causes, heartbeat timeout, same-instance call continuation, idempotent Worker cleanup across `error` plus `close`, scheduling-free candidate/readiness/liveness invalidation, post-detach deadline recomputation, concurrent cleanup deduplication plus failure retry, previous-ready-duration retention, retained asynchronous stale-socket cleanup with rejection containment, Worker pong-before-alarm ordering, one explicit terminal-result alarm schedule, event-time alarm write and durable-deadline-read failure isolation, pong/welcome send failure, and `diagnose_runtime` relay history;
119
119
  - recoverable managed-job terminal commits under injected result/status/delete/confirmation failures, result-only terminal reconstruction, private runtime/plan scrubbing, 24-hour staged-plan expiry, and minimal/full runner-environment inheritance;
120
120
  - cross-process security-audit serialization, continuous hash-chain sequence, uninstall blocking for audit/authorization/job transition/recovery locks, workspace recovery-envelope validation, and symlink/multiple-hard-link rejection at owner state and operational logs;
121
121
  - browser fixed identity, role-separated credentials, legacy migration, socket replacement/reconnect/timer cleanup, duplicate request rejection, response-delivery failure, and read-only broker load projection;
package/docs/UPGRADING.md CHANGED
@@ -1,5 +1,17 @@
1
1
  # Upgrading
2
2
 
3
+ ## 3.0.0-beta.38 heartbeat ordering
4
+
5
+ Beta.38 is a coordinated Worker and daemon update over the currently activated beta.37 candidate. It does not claim to eliminate TCP/WebSocket interruptions caused by Wi-Fi, VPN/TUN, proxy, edge, or upstream changes. It removes Worker-side amplification paths by sending `pong` before Durable Object alarm I/O, coalescing terminal-result deadline scheduling, and preventing event-time deadline/storage failures from aborting dispatch or WebSocket handling. Keep beta.37 active until an exact beta.38 candidate has completed local verification and owner-machine activation.
6
+
7
+ After activation, verify exact Worker/daemon/service convergence, one ready daemon, `daemon.relay_transport.outage_active=false`, zero detached calls after the two-minute grace, and successful representative file and shell operations. A forced or naturally occurring brief interruption should recover without daemon PID or launchd run-count change. Interruptions shorter than ten seconds are expected to be absent from warning-level service logs; use authenticated `server_info.daemon.relay_transport` for their close category, code, timing, and attempt count.
8
+
9
+ ## 3.0.0-beta.37 relay recovery
10
+
11
+ Beta.37 supersedes the locally prepared but never activated beta.36 candidate. Do not activate the beta.36 tarball: its promotion digest is stale after the second-order relay fixes. Beta.37 is a coordinated Worker and daemon update. Older components ignore or omit the optional relay diagnostic field, but exact convergence is required for idempotent cleanup, socket-generation-bound authentication, close-category precedence, and post-detach alarm recomputation.
12
+
13
+ After owner-authorized candidate activation, verify matching package/Worker/service versions, one verified login daemon, readiness recovery across a forced brief socket interruption, zero stale pending calls, and a bounded `server_info.daemon.relay_transport` summary whose `outage_active` field is false on the ready socket. `machine-mcp doctor` must explicitly report that it did not inspect the running service relay. Do not infer that a VPN/TUN product caused a disconnect solely from the coarse `system-network-stack` route class.
14
+
3
15
  ## Supported upgrade contract
4
16
 
5
17
  Machine Bridge supports direct upgrade from the immediately preceding published release. Obsolete transport, state, lock, browser-extension, and authorization implementations are not retained as hidden compatibility paths.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "machine-bridge-mcp",
3
- "version": "3.0.0-beta.35",
3
+ "version": "3.0.0-beta.38",
4
4
  "description": "Cross-client MCP bridge for local agent context, structured browser and application automation, files, Git, processes, resources, and durable jobs over stdio or OAuth relay.",
5
5
  "type": "module",
6
6
  "license": "MIT",
@@ -235,6 +235,8 @@
235
235
  }
236
236
  },
237
237
  "overrides": {
238
- "sharp": "0.35.3"
238
+ "sharp": "0.35.3",
239
+ "brace-expansion": "5.0.9",
240
+ "undici": "7.29.0"
239
241
  }
240
242
  }
package/src/local/cli.mjs CHANGED
@@ -26,6 +26,7 @@ import { loadServiceEnvironment } from "./service-environment.mjs";
26
26
  import { createDeviceSessionForRoot, deviceRootProviderStatus, ensurePreferredDeviceRoot } from "./device-root-provider.mjs";
27
27
  import { convergeRemoteConfiguration } from "./remote-configuration.mjs";
28
28
  import { workerHealth } from "./worker-health.mjs";
29
+ import { DOCTOR_RUNTIME_SCOPE, doctorRuntimeCheckProjection } from "./doctor-reporting.mjs";
29
30
  export { workerHealthUserReason } from "./worker-health.mjs";
30
31
  import { activeStateJobs, activeStateLocks, knownProfileStates, knownWorkerNames } from "./state-inventory.mjs";
31
32
  import {
@@ -593,16 +594,11 @@ async function doctorCommand(args) {
593
594
  } finally {
594
595
  diagnosticRuntime.stop();
595
596
  }
596
- for (const check of runtimeDiagnostics.checks) {
597
- checks.push({
598
- name: `runtime:${check.layer}`,
599
- ok: check.skipped === true || check.ok === true,
600
- detail: check.skipped ? `skipped (${check.error_class || "not applicable"})` : check.ok ? "ok" : check.error_class || "failed",
601
- });
602
- }
597
+ for (const check of runtimeDiagnostics.checks) checks.push(doctorRuntimeCheckProjection(check));
603
598
  console.log(JSON.stringify({
604
599
  ok: checks.every(check => check.ok),
605
600
  checks,
601
+ diagnosticScope: DOCTOR_RUNTIME_SCOPE,
606
602
  runtimeDiagnostics,
607
603
  state: redactState(state),
608
604
  }, null, 2));
@@ -0,0 +1,23 @@
1
+ export const DOCTOR_RUNTIME_SCOPE = Object.freeze({
2
+ running_service_process_inspected: false,
3
+ remote_relay_inspected: false,
4
+ reason: "doctor uses an isolated local runtime; inspect authenticated server_info.daemon.relay_transport for the running service relay",
5
+ });
6
+
7
+ export function doctorRuntimeCheckProjection(check = {}) {
8
+ const layer = String(check.layer || "unknown");
9
+ const skipped = check.skipped === true;
10
+ const detail = skipped && layer === "remote-relay"
11
+ ? "not inspected (doctor uses an isolated local runtime; inspect authenticated server_info.daemon.relay_transport)"
12
+ : skipped
13
+ ? `skipped (${check.error_class || "not applicable"})`
14
+ : check.ok === true
15
+ ? "ok"
16
+ : check.error_class || "failed";
17
+ return {
18
+ name: `runtime:${layer}`,
19
+ ok: skipped || check.ok === true,
20
+ applicable: !skipped,
21
+ detail,
22
+ };
23
+ }
@@ -3,7 +3,7 @@ import { classifyOperationalError } from "./log.mjs";
3
3
  import { proxyAgentForWebSocket } from "./network-proxy.mjs";
4
4
  import { RelayHeartbeatMonitor } from "./relay-heartbeat.mjs";
5
5
  import {
6
- APPLICATION_PROXY_ROUTE_SCOPE, relayOutageFields, relayRecoveryFields, relayStatusSnapshot,
6
+ APPLICATION_PROXY_ROUTE_SCOPE, preferredRelayCloseCategory, relayOutageFields, relayRecoveryFields, relayStatusSnapshot,
7
7
  } from "./relay-diagnostics.mjs";
8
8
  const DEFAULT_HEARTBEAT_INTERVAL_MS = 25_000;
9
9
  const DEFAULT_HEARTBEAT_TIMEOUT_MS = 75_000;
@@ -107,7 +107,7 @@ export class RelayConnection {
107
107
  silent_for_ms: silentForMs,
108
108
  event_loop_lag_ms: eventLoopLagMs,
109
109
  });
110
- this.pendingCloseCategory = "relay_heartbeat_timeout";
110
+ this.pendingCloseCategory = preferredRelayCloseCategory(this.pendingCloseCategory, "relay_heartbeat_timeout");
111
111
  terminateSocket(socket);
112
112
  },
113
113
  });
@@ -170,7 +170,7 @@ export class RelayConnection {
170
170
 
171
171
  interrupt(category = "relay_transport_error") {
172
172
  if (this.closed || !this.socket) return false;
173
- this.pendingCloseCategory = String(category || "relay_transport_error");
173
+ this.pendingCloseCategory = preferredRelayCloseCategory(this.pendingCloseCategory, category);
174
174
  terminateSocket(this.socket);
175
175
  return true;
176
176
  }
@@ -185,16 +185,16 @@ export class RelayConnection {
185
185
  return false;
186
186
  }
187
187
  this.logger.debug?.("remote relay welcome received");
188
- Promise.resolve(this.helloMessage(message)).then((hello) => {
188
+ Promise.resolve(this.helloMessage(message, this.status())).then((hello) => {
189
189
  if (this.socket !== socket || this.closed || this.authenticated) return;
190
190
  this.sendOnSocket(socket, hello);
191
191
  }).catch((error) => {
192
+ if (this.socket !== socket || this.closed || this.authenticated) return;
192
193
  this.logger.debug?.("could not create daemon authentication proof", { error_class: classifyOperationalError(error) });
193
194
  this.failPermanently("relay_authentication_failed");
194
195
  });
195
196
  return true;
196
197
  }
197
-
198
198
  acknowledge(message = {}) {
199
199
  const socket = this.socket;
200
200
  if (this.closed || this.authenticated || !this.isSocketOpen(socket)) return false;
@@ -216,7 +216,7 @@ export class RelayConnection {
216
216
  this.readinessTimer = this.scheduler.setTimeout(() => {
217
217
  if (this.socket !== socket || this.closed || this.ready) return;
218
218
  this.logger.debug?.("remote relay end-to-end readiness probe timed out", { timeout_ms: this.readinessTimeoutMs });
219
- this.pendingCloseCategory = "relay_readiness_timeout";
219
+ this.pendingCloseCategory = preferredRelayCloseCategory(this.pendingCloseCategory, "relay_readiness_timeout");
220
220
  terminateSocket(socket);
221
221
  }, this.readinessTimeoutMs);
222
222
  this.readinessTimer?.unref?.();
@@ -287,7 +287,7 @@ export class RelayConnection {
287
287
  if (reconnectCategory) {
288
288
  const socket = this.socket;
289
289
  if (this.closed || !socket) return true;
290
- this.pendingCloseCategory = reconnectCategory;
290
+ this.pendingCloseCategory = preferredRelayCloseCategory(this.pendingCloseCategory, reconnectCategory);
291
291
  terminateSocket(socket);
292
292
  return true;
293
293
  }
@@ -329,7 +329,7 @@ export class RelayConnection {
329
329
  this.connectTimer = this.scheduler.setTimeout(() => {
330
330
  if (this.socket !== socket || this.closed || this.isSocketOpen(socket)) return;
331
331
  this.logger.debug?.("remote relay transport connection timed out", { timeout_ms: this.connectTimeoutMs });
332
- this.pendingCloseCategory = "relay_connect_timeout";
332
+ this.pendingCloseCategory = preferredRelayCloseCategory(this.pendingCloseCategory, "relay_connect_timeout");
333
333
  terminateSocket(socket);
334
334
  }, this.connectTimeoutMs);
335
335
  this.connectTimer?.unref?.();
@@ -346,7 +346,7 @@ export class RelayConnection {
346
346
  this.handshakeTimer = this.scheduler.setTimeout(() => {
347
347
  if (this.socket !== socket || this.closed || this.ready) return;
348
348
  this.logger.debug?.("remote relay authentication acknowledgement timed out", { timeout_ms: this.handshakeTimeoutMs });
349
- this.pendingCloseCategory = "relay_handshake_timeout";
349
+ this.pendingCloseCategory = preferredRelayCloseCategory(this.pendingCloseCategory, "relay_handshake_timeout");
350
350
  terminateSocket(socket);
351
351
  }, this.handshakeTimeoutMs);
352
352
  this.handshakeTimer?.unref?.();
@@ -392,7 +392,7 @@ export class RelayConnection {
392
392
  const disconnectedAt = this.now();
393
393
  const connectedForMs = wasAuthenticated && this.connectedAt > 0 ? Math.max(0, disconnectedAt - this.connectedAt) : 0;
394
394
  this.lastDisconnectedAt = disconnectedAt;
395
- this.lastReadyDurationMs = wasReady && this.lastReadyAt > 0 ? Math.max(0, disconnectedAt - this.lastReadyAt) : 0;
395
+ if (wasReady && this.lastReadyAt > 0) this.lastReadyDurationMs = Math.max(0, disconnectedAt - this.lastReadyAt);
396
396
  this.lastCloseCode = Number(code) || 0;
397
397
  this.logger.debug?.("remote relay transport closed", {
398
398
  close_code: this.lastCloseCode,
@@ -438,7 +438,7 @@ export class RelayConnection {
438
438
  this.failPermanently("relay_authentication_failed");
439
439
  return;
440
440
  }
441
- this.pendingCloseCategory = "relay_transport_error";
441
+ this.pendingCloseCategory = preferredRelayCloseCategory(this.pendingCloseCategory, "relay_transport_error");
442
442
  terminateSocket(socket);
443
443
  });
444
444
  }
@@ -514,7 +514,7 @@ export class RelayConnection {
514
514
  return true;
515
515
  } catch (error) {
516
516
  this.lastTransportErrorClass = classifyRelayTransportError(error);
517
- this.pendingCloseCategory = "relay_transport_error";
517
+ this.pendingCloseCategory = preferredRelayCloseCategory(this.pendingCloseCategory, "relay_transport_error");
518
518
  this.logger.debug?.("remote relay send failed", { error_class: this.lastTransportErrorClass });
519
519
  terminateSocket(socket);
520
520
  return false;
@@ -1,5 +1,11 @@
1
1
  export const APPLICATION_PROXY_ROUTE_SCOPE = "application-proxy-selection-only";
2
2
 
3
+ export function preferredRelayCloseCategory(current, next) {
4
+ const existing = String(current || "");
5
+ const candidate = String(next || "relay_transport_error");
6
+ return !existing || existing === "relay_transport_error" ? candidate : existing;
7
+ }
8
+
3
9
  export function relayStatusSnapshot(state, now = Date.now()) {
4
10
  const current = Number(now) || 0;
5
11
  return {
@@ -12,6 +18,7 @@ export function relayStatusSnapshot(state, now = Date.now()) {
12
18
  reconnect_attempt: state.reconnectAttempt,
13
19
  outage_active: state.outageStartedAt > 0,
14
20
  outage_count: state.outageCount,
21
+ outage_attempts: state.outageAttempts,
15
22
  outage_started_at: isoTimestamp(state.outageStartedAt),
16
23
  outage_duration_ms: state.outageStartedAt > 0 ? Math.max(0, current - state.outageStartedAt) : 0,
17
24
  last_close_category: state.outageCount > 0 ? state.lastCloseCategory : null,
@@ -0,0 +1,24 @@
1
+ // @ts-check
2
+
3
+ import { clampInteger } from "./numbers.mjs";
4
+ import { isPlainRecord } from "./records.mjs";
5
+
6
+ export function relayHandshakeDiagnostics(value = {}) {
7
+ const status = isPlainRecord(value) ? value : {};
8
+ return {
9
+ schema_version: 1,
10
+ network_route: typeof status.network_route === "string" ? status.network_route : "unresolved",
11
+ outage_count: clampInteger(status.outage_count, 0, 0, 1_000_000_000),
12
+ outage_active: status.outage_active === true,
13
+ outage_started_at: typeof status.outage_started_at === "string" ? status.outage_started_at : null,
14
+ outage_duration_ms: clampInteger(status.outage_duration_ms, 0, 0, 31 * 24 * 60 * 60_000),
15
+ outage_attempts: clampInteger(status.outage_attempts, 0, 0, 1_000_000),
16
+ last_close_category: typeof status.last_close_category === "string" ? status.last_close_category : null,
17
+ last_close_code: Number.isSafeInteger(status.last_close_code) ? status.last_close_code : null,
18
+ last_transport_error_class: typeof status.last_transport_error_class === "string"
19
+ ? status.last_transport_error_class
20
+ : null,
21
+ last_disconnected_at: typeof status.last_disconnected_at === "string" ? status.last_disconnected_at : null,
22
+ previous_ready_duration_ms: clampInteger(status.last_ready_duration_ms, 0, 0, 365 * 24 * 60 * 60_000),
23
+ };
24
+ }
@@ -1,6 +1,7 @@
1
1
  import { Buffer } from "node:buffer";
2
2
  import relayContract from "../shared/relay-contract.json" with { type: "json" };
3
3
  import { RelayConnection } from "./relay-connection.mjs";
4
+ import { relayHandshakeDiagnostics } from "./relay-peer-diagnostics.mjs";
4
5
  import { createDaemonAuthentication, createDaemonPreflightHeaders, createDeviceSessionIdentity, validateDeviceSessionIdentity } from "./device-identity.mjs";
5
6
  import { MCP_SUPPORTED_PROTOCOL_VERSIONS, SERVER_NAME } from "./tools.mjs";
6
7
  import { normalizeAccountRole } from "./account-access.mjs";
@@ -18,17 +19,15 @@ export function createRuntimeRelayConnection(runtime, { workerUrl, deviceIdentit
18
19
  expectedServer: SERVER_NAME,
19
20
  expectedVersion: String(expectedVersion || ""),
20
21
  connectionHeaders: () => createDaemonPreflightHeaders(
21
- sessionIdentity,
22
- workerUrl,
23
- SERVER_NAME,
24
- String(expectedVersion || ""),
22
+ sessionIdentity, workerUrl, SERVER_NAME, String(expectedVersion || ""),
25
23
  ),
26
- helloMessage: async (welcome) => ({
24
+ helloMessage: async (welcome, relayStatus) => ({
27
25
  type: "hello",
28
26
  instance_id: runtime.relayInstanceId,
29
27
  tools: runtime.tools(),
30
28
  policy: runtime.policy,
31
29
  protocol_versions: MCP_SUPPORTED_PROTOCOL_VERSIONS,
30
+ relay_diagnostics: relayHandshakeDiagnostics(relayStatus),
32
31
  authentication: await createDaemonAuthentication(sessionIdentity, welcome, runtime.relayInstanceId),
33
32
  }),
34
33
  onMessage: (data, relayContext) => handleRelayData(runtime, data, relayContext),
@@ -0,0 +1,110 @@
1
+ import { sanitizeMetadataText } from "./http.ts";
2
+
3
+ const NETWORK_ROUTES = new Set([
4
+ "unresolved",
5
+ "application-http-proxy",
6
+ "system-network-stack",
7
+ "invalid-application-proxy-configuration",
8
+ ]);
9
+ const TRANSPORT_ERROR_CLASSES = new Set([
10
+ "cancelled", "timeout", "authentication_failed", "authorization_denied", "policy_denied",
11
+ "invalid_request", "not_found", "conflict", "limit_exceeded", "permission_denied",
12
+ "path_boundary", "network_error", "protocol_error", "unavailable", "integrity_error",
13
+ "execution_failed", "internal_error",
14
+ ]);
15
+ const CLOSE_CATEGORIES = new Set([
16
+ "connection_interrupted",
17
+ "relay_restarting_or_unavailable",
18
+ "relay_policy_rejected",
19
+ "relay_internal_error",
20
+ "relay_protocol_mismatch",
21
+ "relay_authentication_failed",
22
+ "relay_connect_timeout",
23
+ "relay_handshake_timeout",
24
+ "relay_readiness_timeout",
25
+ "relay_heartbeat_timeout",
26
+ "relay_transport_error",
27
+ "relay_protocol_error",
28
+ "relay_proxy_configuration",
29
+ "invalid_transport_payload",
30
+ "message_too_large",
31
+ "normal_close",
32
+ "unexpected_close",
33
+ "superseded",
34
+ ]);
35
+
36
+ export interface DaemonRelayDiagnostics {
37
+ schema_version: 1;
38
+ network_route: string;
39
+ outage_count: number;
40
+ outage_active: boolean;
41
+ outage_started_at: string | null;
42
+ outage_duration_ms: number;
43
+ outage_attempts: number;
44
+ last_close_category: string | null;
45
+ last_close_code: number | null;
46
+ last_transport_error_class: string | null;
47
+ last_disconnected_at: string | null;
48
+ previous_ready_duration_ms: number;
49
+ }
50
+
51
+ export function sanitizeDaemonRelayDiagnostics(value: unknown): DaemonRelayDiagnostics | undefined {
52
+ if (!value || typeof value !== "object" || Array.isArray(value)) return undefined;
53
+ const candidate = value as Record<string, unknown>;
54
+ if (candidate.schema_version !== 1) return undefined;
55
+ return {
56
+ schema_version: 1,
57
+ network_route: enumText(candidate.network_route, NETWORK_ROUTES, "unresolved"),
58
+ outage_count: boundedInteger(candidate.outage_count, 0, 1_000_000_000, 0),
59
+ outage_active: candidate.outage_active === true,
60
+ outage_started_at: timestamp(candidate.outage_started_at),
61
+ outage_duration_ms: boundedInteger(candidate.outage_duration_ms, 0, 31 * 24 * 60 * 60_000, 0),
62
+ outage_attempts: boundedInteger(candidate.outage_attempts, 0, 1_000_000, 0),
63
+ last_close_category: nullableEnum(candidate.last_close_category, CLOSE_CATEGORIES),
64
+ last_close_code: nullableInteger(candidate.last_close_code, 0, 4999),
65
+ last_transport_error_class: nullableEnum(candidate.last_transport_error_class, TRANSPORT_ERROR_CLASSES),
66
+ last_disconnected_at: timestamp(candidate.last_disconnected_at),
67
+ previous_ready_duration_ms: boundedInteger(candidate.previous_ready_duration_ms, 0, 365 * 24 * 60 * 60_000, 0),
68
+ };
69
+ }
70
+
71
+ function enumText(value: unknown, allowed: Set<string>, fallback: string): string {
72
+ const text = typeof value === "string" ? value : "";
73
+ return allowed.has(text) ? text : fallback;
74
+ }
75
+
76
+ function nullableEnum(value: unknown, allowed: Set<string>): string | null {
77
+ const text = typeof value === "string" ? value : "";
78
+ return allowed.has(text) ? text : null;
79
+ }
80
+
81
+ export function relayDiagnosticsAfterReady(
82
+ value: DaemonRelayDiagnostics | undefined,
83
+ readyAt?: string,
84
+ ): DaemonRelayDiagnostics | undefined {
85
+ if (!value) return undefined;
86
+ const started = Date.parse(value.outage_started_at ?? "");
87
+ const ready = Date.parse(readyAt ?? "");
88
+ const elapsed = Number.isFinite(started) && Number.isFinite(ready) ? Math.max(0, ready - started) : 0;
89
+ return {
90
+ ...value,
91
+ outage_active: false,
92
+ outage_duration_ms: Math.min(31 * 24 * 60 * 60_000, Math.max(value.outage_duration_ms, elapsed)),
93
+ };
94
+ }
95
+
96
+ function timestamp(value: unknown): string | null {
97
+ const text = sanitizeMetadataText(value, 64);
98
+ const parsed = text ? Date.parse(text) : Number.NaN;
99
+ return Number.isFinite(parsed) ? new Date(parsed).toISOString() : null;
100
+ }
101
+
102
+ function nullableInteger(value: unknown, minimum: number, maximum: number): number | null {
103
+ const number = Number(value);
104
+ return Number.isSafeInteger(number) && number >= minimum && number <= maximum ? number : null;
105
+ }
106
+
107
+ function boundedInteger(value: unknown, minimum: number, maximum: number, fallback: number): number {
108
+ const number = Number(value);
109
+ return Number.isSafeInteger(number) && number >= minimum && number <= maximum ? number : fallback;
110
+ }
@@ -1,5 +1,6 @@
1
1
  import { sanitizeDaemonChallengeAttachment } from "./daemon-auth.ts";
2
2
  import type { DaemonRole } from "./daemon-liveness.ts";
3
+ import { sanitizeDaemonRelayDiagnostics, type DaemonRelayDiagnostics } from "./daemon-relay-diagnostics.ts";
3
4
  import { sanitizeMetadataText } from "./http.ts";
4
5
  import { sanitizeDaemonPolicy, sanitizeDaemonTools, type DaemonPolicy } from "./policy.ts";
5
6
 
@@ -12,6 +13,7 @@ export interface DaemonAttachment {
12
13
  connectionId?: string;
13
14
  policy?: DaemonPolicy;
14
15
  tools?: string[];
16
+ relayDiagnostics?: DaemonRelayDiagnostics;
15
17
  authChallenge?: string;
16
18
  authIssuedAt?: number;
17
19
  authExpiresAt?: number;
@@ -35,6 +37,7 @@ export function sanitizeDaemonAttachment(value: unknown): DaemonAttachment | und
35
37
  connectionId: sanitizeConnectionId(candidate.connectionId),
36
38
  policy,
37
39
  tools: sanitizeDaemonTools(candidate.tools, policy),
40
+ relayDiagnostics: sanitizeDaemonRelayDiagnostics(candidate.relayDiagnostics),
38
41
  ...sanitizeDaemonChallengeAttachment(candidate as Record<string, unknown>),
39
42
  };
40
43
  }
@@ -1,13 +1,16 @@
1
1
  import { isLiveDaemonAttachment, withDaemonLastSeenAt, type DaemonRole } from "./daemon-liveness.ts";
2
2
  import type { DaemonPolicy } from "./policy.ts";
3
+ import { relayDiagnosticsAfterReady, type DaemonRelayDiagnostics } from "./daemon-relay-diagnostics.ts";
3
4
  import type { DaemonChallenge } from "./daemon-auth.ts";
4
5
  import { sanitizeDaemonAttachment, type DaemonAttachment } from "./daemon-socket-attachment.ts";
5
6
  interface WebSocketContext {
6
7
  getWebSockets(): WebSocket[];
7
8
  }
8
9
 
10
+ export interface DaemonSocketCleanup { task: Promise<void>; first: boolean }
9
11
  export class DaemonSocketRegistry {
10
12
  private readonly context: WebSocketContext;
13
+ private readonly cleanupTasks = new WeakMap<WebSocket, Promise<void>>();
11
14
  constructor(context: WebSocketContext) { this.context = context; }
12
15
 
13
16
  attachment(socket: WebSocket): DaemonAttachment | undefined {
@@ -54,18 +57,26 @@ export class DaemonSocketRegistry {
54
57
  } satisfies DaemonAttachment);
55
58
  }
56
59
 
57
- beginProbe(socket: WebSocket, values: { connectedAt: string; probeId: string; instanceId: string; connectionId: string; policy: DaemonPolicy; tools: string[] }): void {
60
+ beginProbe(socket: WebSocket, values: {
61
+ connectedAt: string; probeId: string; instanceId: string; connectionId: string;
62
+ policy: DaemonPolicy; tools: string[]; relayDiagnostics?: DaemonRelayDiagnostics;
63
+ }): void {
58
64
  socket.serializeAttachment({
59
65
  role: "probing", connectedAt: values.connectedAt, lastSeenAt: values.connectedAt,
60
66
  probeId: values.probeId, instanceId: values.instanceId, connectionId: values.connectionId,
61
- policy: values.policy, tools: values.tools,
67
+ policy: values.policy, tools: values.tools, relayDiagnostics: values.relayDiagnostics,
62
68
  } satisfies DaemonAttachment);
63
69
  }
64
70
 
65
71
  promote(socket: WebSocket, lastSeenAt = new Date().toISOString()): DaemonAttachment | undefined {
66
72
  const attachment = this.attachment(socket);
67
73
  if (attachment?.role !== "probing") return undefined;
68
- const ready = { ...attachment, role: "daemon" as const, lastSeenAt };
74
+ const ready = {
75
+ ...attachment,
76
+ role: "daemon" as const,
77
+ lastSeenAt,
78
+ relayDiagnostics: relayDiagnosticsAfterReady(attachment.relayDiagnostics, lastSeenAt),
79
+ };
69
80
  delete ready.probeId;
70
81
  socket.serializeAttachment(ready satisfies DaemonAttachment);
71
82
  return ready;
@@ -79,13 +90,37 @@ export class DaemonSocketRegistry {
79
90
  return touched;
80
91
  }
81
92
 
82
- expire(socket: WebSocket): void {
93
+ expire(socket: WebSocket): DaemonAttachment | undefined {
83
94
  const attachment = this.attachment(socket);
84
- if (!attachment) return;
95
+ if (!attachment || attachment.role === "expired") return undefined;
85
96
  socket.serializeAttachment({
86
97
  role: "expired", connectedAt: attachment.connectedAt, lastSeenAt: attachment.lastSeenAt,
87
98
  instanceId: attachment.instanceId, connectionId: attachment.connectionId,
99
+ relayDiagnostics: attachment.relayDiagnostics,
88
100
  } satisfies DaemonAttachment);
101
+ return attachment;
102
+ }
103
+
104
+ beginCleanup(
105
+ socket: WebSocket,
106
+ operation: (attachment: DaemonAttachment) => Promise<unknown>,
107
+ ): DaemonSocketCleanup | undefined {
108
+ const existing = this.cleanupTasks.get(socket);
109
+ if (existing) return { task: existing, first: false };
110
+ const attachment = this.attachment(socket);
111
+ if (!attachment) return undefined;
112
+ const first = attachment.role !== "expired";
113
+ if (first) this.expire(socket);
114
+ let task!: Promise<void>;
115
+ task = Promise.resolve().then(() => operation(attachment)).catch(() => operation(attachment)).then(
116
+ () => undefined,
117
+ (error) => {
118
+ if (this.cleanupTasks.get(socket) === task) this.cleanupTasks.delete(socket);
119
+ throw error;
120
+ },
121
+ );
122
+ this.cleanupTasks.set(socket, task);
123
+ return { task, first };
89
124
  }
90
125
 
91
126
  socketForConnectionId(connectionId: string): WebSocket | undefined {
@@ -11,6 +11,7 @@ import {
11
11
  isFreshDaemonCandidate,
12
12
  } from "./daemon-liveness.ts";
13
13
  import { DaemonSocketRegistry } from "./daemon-sockets.ts";
14
+ import { sanitizeDaemonRelayDiagnostics } from "./daemon-relay-diagnostics.ts";
14
15
  import { processRuntimeAlarm, scheduleRuntimeAlarm } from "./runtime-alarm.ts";
15
16
  import { consumeDaemonPreflightNonce, createDaemonChallenge, verifyDaemonAuthentication, verifyDaemonPreflight } from "./daemon-auth.ts";
16
17
  import { resolveMcpSession } from "./mcp-session.ts";
@@ -42,6 +43,7 @@ import {
42
43
  } from "./http.ts";
43
44
  import { authorizationServerMetadata } from "./worker-metadata.ts";
44
45
  import { workerBodyLimitBytes, type BridgeEnv } from "./worker-runtime-config.ts";
46
+ import { retainWorkerTask } from "./worker-task-lifetime.ts";
45
47
  import {
46
48
  MCP_DISCOVERY_TTL_MS, MCP_INSTRUCTIONS, MCP_LEGACY_PROTOCOL_VERSIONS,
47
49
  MCP_MODERN_PROTOCOL_VERSIONS, MCP_SERVER_CAPABILITIES, MCP_TOOL_LIST_TTL_MS,
@@ -58,8 +60,7 @@ import {
58
60
  closeWebSocketQuietly, daemonErrorCloseCode, isObjectRecord, rejectDaemonMessage,
59
61
  sendWebSocketQuietly, trySendWebSocket,
60
62
  } from "./websocket-protocol.ts";
61
-
62
- const SERVER_VERSION = "3.0.0-beta.35";
63
+ const SERVER_VERSION = "3.0.0-beta.38";
63
64
  const MCP_SERVER_INFO = mcpServerInfo(SERVER_VERSION);
64
65
  const MAX_DAEMON_MESSAGE_BYTES = 8 * 1024 * 1024;
65
66
  const DAEMON_RECONNECT_GRACE_MS = relayContract.reconnectGraceMs;
@@ -271,15 +272,16 @@ export class BridgeRoom extends DurableObject<BridgeEnv> {
271
272
  connectionId,
272
273
  policy: daemonPolicy,
273
274
  tools: sanitizeDaemonTools(body.tools, daemonPolicy),
275
+ relayDiagnostics: sanitizeDaemonRelayDiagnostics(body.relay_diagnostics),
274
276
  });
275
277
  this.observability.socketAuthenticated();
276
278
  try {
277
279
  ws.send(JSON.stringify({ type: "hello_ack", server: SERVER_NAME, version: SERVER_VERSION }));
278
280
  ws.send(JSON.stringify({ type: "relay_probe", id: probeId }));
279
281
  } catch {
280
- this.daemonRegistry.expire(ws);
281
- closeWebSocketQuietly(ws, 1011, "daemon readiness probe failed");
282
- await this.scheduleRuntimeAlarm();
282
+ await this.invalidateDaemonSocket(
283
+ ws, "daemon readiness probe failed", "daemon readiness probe failed", "daemon_transport_error",
284
+ );
283
285
  return;
284
286
  }
285
287
  await this.scheduleRuntimeAlarm();
@@ -292,10 +294,12 @@ export class BridgeRoom extends DurableObject<BridgeEnv> {
292
294
  }
293
295
 
294
296
  if (body.type === "heartbeat" || body.type === "ping") {
295
- await this.touchDaemonSocket(ws);
297
+ if (!this.touchDaemonSocket(ws)) return;
296
298
  if (!trySendWebSocket(ws, { type: "pong", ts: body.ts ?? Date.now() })) {
297
299
  await this.invalidateDaemonSocket(ws, "failed to acknowledge daemon heartbeat", "daemon pong failed", "daemon_transport_error");
300
+ return;
298
301
  }
302
+ await this.scheduleRuntimeAlarm();
299
303
  return;
300
304
  }
301
305
 
@@ -340,9 +344,9 @@ export class BridgeRoom extends DurableObject<BridgeEnv> {
340
344
  }
341
345
  this.observability.socketReady();
342
346
  for (const previous of previousSockets) {
343
- await this.detachDaemonSocketCalls(previous, "daemon connection replaced after verified handover");
344
- this.daemonRegistry.expire(previous);
347
+ const cleanup = this.cleanupDaemonSocket(previous, "daemon connection replaced after verified handover");
345
348
  closeWebSocketQuietly(previous, 1012, "replaced by verified daemon");
349
+ if (cleanup) await cleanup.task;
346
350
  }
347
351
  await this.scheduleRuntimeAlarm();
348
352
  return;
@@ -358,7 +362,7 @@ export class BridgeRoom extends DurableObject<BridgeEnv> {
358
362
  return;
359
363
  }
360
364
 
361
- await this.touchDaemonSocket(ws);
365
+ if (!this.touchDaemonSocket(ws)) return;
362
366
  const outcome: PendingCallOutcome = body.ok === false
363
367
  ? { ok: false, error: daemonToolError(body.error) }
364
368
  : { ok: true, value: body.result };
@@ -371,24 +375,25 @@ export class BridgeRoom extends DurableObject<BridgeEnv> {
371
375
  const decision = daemonTerminalResultDecision(transientMatched, durableSettlement);
372
376
  if (decision.acknowledge) trySendWebSocket(ws, { type: "tool_result_ack", id: body.id });
373
377
  this.observability.daemonTerminalResult(decision.disposition);
374
- if (decision.matched) await this.scheduleRuntimeAlarm();
378
+ await this.scheduleRuntimeAlarm();
375
379
  }
376
380
 
377
381
  async webSocketClose(ws: WebSocket): Promise<void> {
378
382
  if (this.streamChannel.isSubscriber(ws)) return;
379
- await this.cleanupDaemonSocket(ws, "daemon disconnected");
383
+ const cleanup = this.cleanupDaemonSocket(ws, "daemon disconnected");
384
+ if (cleanup) { await cleanup.task; await this.scheduleRuntimeAlarm(); }
380
385
  }
381
-
382
386
  async webSocketError(ws: WebSocket, error: unknown): Promise<void> {
383
387
  if (this.streamChannel.isSubscriber(ws)) return;
384
- this.observability.event("warn", "daemon.websocket.error", { error_class: workerErrorClass(error) });
385
- await this.cleanupDaemonSocket(ws, "daemon transport error");
388
+ const cleanup = this.cleanupDaemonSocket(ws, "daemon transport error");
389
+ if (!cleanup) return;
390
+ if (cleanup.first) this.observability.event("warn", "daemon.websocket.error", { error_class: workerErrorClass(error) });
391
+ await cleanup.task; await this.scheduleRuntimeAlarm();
386
392
  }
387
-
388
- private async cleanupDaemonSocket(ws: WebSocket, message: string): Promise<void> {
389
- this.observability.socketDisconnected();
390
- await this.detachDaemonSocketCalls(ws, message);
391
- await this.scheduleRuntimeAlarm();
393
+ private cleanupDaemonSocket(ws: WebSocket, message: string) {
394
+ const cleanup = this.daemonRegistry.beginCleanup(ws, (attachment) => this.detachDaemonSocketCalls(ws, message, attachment));
395
+ if (cleanup?.first) this.observability.socketDisconnected();
396
+ return cleanup;
392
397
  }
393
398
  private async handleMcp(request: Request, base: string): Promise<Response> {
394
399
  const originRejection = mcpOriginRejection(request, base, this.env.MBM_ALLOWED_ORIGINS ?? "");
@@ -707,9 +712,7 @@ export class BridgeRoom extends DurableObject<BridgeEnv> {
707
712
  },
708
713
  });
709
714
  if (!welcomed) {
710
- this.daemonRegistry.expire(server);
711
- closeWebSocketQuietly(server, 1011, "daemon welcome failed");
712
- await this.scheduleRuntimeAlarm();
715
+ await this.invalidateDaemonSocket(server, "daemon welcome failed", "daemon welcome failed", "daemon_transport_error");
713
716
  }
714
717
  return new Response(null, { status: 101, webSocket: client });
715
718
  }
@@ -742,26 +745,20 @@ export class BridgeRoom extends DurableObject<BridgeEnv> {
742
745
  return new Set(attachment.tools);
743
746
  }
744
747
 
745
- private async touchDaemonSocket(ws: WebSocket): Promise<void> {
746
- if (!this.daemonRegistry.touch(ws)) return;
747
- await this.scheduleRuntimeAlarm();
748
- }
749
-
748
+ private touchDaemonSocket(ws: WebSocket): boolean { return Boolean(this.daemonRegistry.touch(ws)); }
750
749
  private async invalidateDaemonSocket(
751
- ws: WebSocket,
752
- message: string,
753
- closeReason: string,
754
- errorCode = "daemon_liveness_timeout",
750
+ ws: WebSocket, message: string, closeReason: string,
751
+ errorCode = "daemon_liveness_timeout", scheduleAlarm = true,
755
752
  ): Promise<void> {
756
- const cleanup = this.detachDaemonSocketCalls(ws, message);
757
- this.daemonRegistry.expire(ws);
753
+ const cleanup = this.cleanupDaemonSocket(ws, message);
758
754
  sendWebSocketQuietly(ws, { type: "error", error: errorCode });
759
755
  closeWebSocketQuietly(ws, daemonErrorCloseCode(errorCode), closeReason);
760
- await cleanup;
756
+ if (cleanup) { await cleanup.task; if (scheduleAlarm) await this.scheduleRuntimeAlarm(); }
761
757
  }
762
758
 
763
- private async detachDaemonSocketCalls(ws: WebSocket, message: string): Promise<number> {
764
- const attachment = this.daemonRegistry.attachment(ws);
759
+ private async detachDaemonSocketCalls(
760
+ ws: WebSocket, message: string, attachment = this.daemonRegistry.attachment(ws),
761
+ ): Promise<number> {
765
762
  if (!attachment?.instanceId || !attachment.connectionId) {
766
763
  return await this.pending.rejectSocket(ws, () => new WorkerToolError("unavailable", message, true));
767
764
  }
@@ -778,7 +775,10 @@ export class BridgeRoom extends DurableObject<BridgeEnv> {
778
775
  for (const socket of this.daemonRegistry.readyRoleSockets()) {
779
776
  const deadline = daemonLivenessDeadlineMs(this.daemonRegistry.readyAttachment(socket));
780
777
  if (Number.isFinite(deadline) && deadline > now) continue;
781
- void this.invalidateDaemonSocket(socket, "daemon became unresponsive", "daemon liveness timeout");
778
+ retainWorkerTask(this.ctx,
779
+ this.invalidateDaemonSocket(socket, "daemon became unresponsive", "daemon liveness timeout"),
780
+ (error) => this.observability.event("error", "daemon.socket.cleanup.failed",
781
+ { error_class: workerErrorClass(error) }));
782
782
  }
783
783
  }
784
784
 
@@ -803,7 +803,7 @@ export class BridgeRoom extends DurableObject<BridgeEnv> {
803
803
  expireDurableCall: (call: import("./mcp-pending-call-store.ts").PendingStreamCallView) => this.durableCalls.expire(call),
804
804
  daemonRegistry: this.daemonRegistry,
805
805
  invalidateDaemonSocket: (socket: WebSocket, message: string, closeReason: string, errorCode?: string) =>
806
- this.invalidateDaemonSocket(socket, message, closeReason, errorCode),
806
+ this.invalidateDaemonSocket(socket, message, closeReason, errorCode, false),
807
807
  onScheduleError: (error: unknown) => this.observability.event(
808
808
  "error", "runtime.alarm.schedule.failed", { error_class: workerErrorClass(error) },
809
809
  ),
@@ -831,6 +831,7 @@ export class BridgeRoom extends DurableObject<BridgeEnv> {
831
831
  ...base,
832
832
  policy: attachment?.policy ?? null,
833
833
  tools,
834
+ relay_transport: attachment?.relayDiagnostics ?? null,
834
835
  };
835
836
  }
836
837
 
@@ -30,6 +30,6 @@ export async function writeEarliestRuntimeAlarm(input: {
30
30
  await input.storage.setAlarm(target);
31
31
  input.onMutation?.("set");
32
32
  } catch (error) {
33
- input.onError(error);
33
+ try { input.onError(error); } catch {}
34
34
  }
35
35
  }
@@ -6,7 +6,6 @@ import {
6
6
  import type { DaemonSocketRegistry } from "./daemon-sockets.ts";
7
7
  import type { McpPendingCallStore, PendingStreamCallView } from "./mcp-pending-call-store.ts";
8
8
  import type { PendingCallRegistry } from "./pending-calls.ts";
9
- import { closeWebSocketQuietly, sendWebSocketQuietly } from "./websocket-protocol.ts";
10
9
  import { writeEarliestRuntimeAlarm, type AlarmStorage } from "./runtime-alarm-storage.ts";
11
10
 
12
11
  type InvalidateDaemonSocket = (
@@ -32,15 +31,18 @@ export async function processRuntimeAlarm(context: RuntimeAlarmContext, now = Da
32
31
  if (context.durableCalls && context.expireDurableCall) {
33
32
  for (const call of await context.durableCalls.due(now)) await context.expireDurableCall(call);
34
33
  }
35
- let nextDeadline = await pendingAlarmDeadline(context, now);
34
+ let nextDeadline = Number.POSITIVE_INFINITY;
36
35
  for (const socket of context.daemonRegistry.candidateSockets()) {
37
36
  const attachment = context.daemonRegistry.attachment(socket);
38
37
  const connectedAt = Date.parse(attachment?.connectedAt ?? "");
39
38
  const deadline = connectedAt + DAEMON_HELLO_TIMEOUT_MS;
40
39
  if (!Number.isFinite(connectedAt) || deadline <= now) {
41
- context.daemonRegistry.expire(socket);
42
- sendWebSocketQuietly(socket, { type: "error", error: "daemon_hello_timeout" });
43
- closeWebSocketQuietly(socket, 1008, "daemon hello timeout");
40
+ await context.invalidateDaemonSocket(
41
+ socket,
42
+ "daemon did not complete authentication",
43
+ "daemon hello timeout",
44
+ "daemon_hello_timeout",
45
+ );
44
46
  continue;
45
47
  }
46
48
  nextDeadline = Math.min(nextDeadline, deadline);
@@ -68,6 +70,7 @@ export async function processRuntimeAlarm(context: RuntimeAlarmContext, now = Da
68
70
  }
69
71
  nextDeadline = Math.min(nextDeadline, deadline);
70
72
  }
73
+ nextDeadline = Math.min(nextDeadline, await pendingAlarmDeadline(context, now));
71
74
  await writeEarliestRuntimeAlarm({
72
75
  storage: context.storage,
73
76
  nextDeadline,
@@ -78,45 +81,50 @@ export async function processRuntimeAlarm(context: RuntimeAlarmContext, now = Da
78
81
  }
79
82
 
80
83
  export async function scheduleRuntimeAlarm(context: RuntimeAlarmContext, now = Date.now()): Promise<void> {
81
- let nextDeadline = await pendingAlarmDeadline(context, now);
82
- for (const socket of context.daemonRegistry.candidateSockets()) {
83
- const connectedAt = Date.parse(context.daemonRegistry.attachment(socket)?.connectedAt ?? "");
84
- if (!Number.isFinite(connectedAt)) {
85
- closeWebSocketQuietly(socket, 1008, "invalid daemon candidate timestamp");
86
- continue;
84
+ try {
85
+ let nextDeadline = Number.POSITIVE_INFINITY;
86
+ for (const socket of context.daemonRegistry.candidateSockets()) {
87
+ const connectedAt = Date.parse(context.daemonRegistry.attachment(socket)?.connectedAt ?? "");
88
+ if (!Number.isFinite(connectedAt)) {
89
+ await context.invalidateDaemonSocket(
90
+ socket,
91
+ "daemon candidate timestamp is invalid",
92
+ "invalid daemon candidate timestamp",
93
+ "daemon_hello_timeout",
94
+ );
95
+ continue;
96
+ }
97
+ nextDeadline = Math.min(nextDeadline, connectedAt + DAEMON_HELLO_TIMEOUT_MS);
87
98
  }
88
- nextDeadline = Math.min(nextDeadline, connectedAt + DAEMON_HELLO_TIMEOUT_MS);
89
- }
90
- for (const socket of context.daemonRegistry.probingSockets()) {
91
- const attachment = context.daemonRegistry.attachment(socket);
92
- const readyDeadline = daemonReadyDeadlineMs(attachment);
93
- const liveDeadline = daemonLivenessDeadlineMs(attachment);
94
- if (!Number.isFinite(readyDeadline) || !Number.isFinite(liveDeadline)) {
95
- await context.invalidateDaemonSocket(
96
- socket,
97
- "daemon readiness state is invalid",
98
- "daemon ready timeout",
99
- "daemon_ready_timeout",
100
- );
101
- continue;
99
+ for (const socket of context.daemonRegistry.probingSockets()) {
100
+ const attachment = context.daemonRegistry.attachment(socket);
101
+ const readyDeadline = daemonReadyDeadlineMs(attachment);
102
+ const liveDeadline = daemonLivenessDeadlineMs(attachment);
103
+ if (!Number.isFinite(readyDeadline) || !Number.isFinite(liveDeadline)) {
104
+ await context.invalidateDaemonSocket(
105
+ socket,
106
+ "daemon readiness state is invalid",
107
+ "daemon ready timeout",
108
+ "daemon_ready_timeout",
109
+ );
110
+ continue;
111
+ }
112
+ nextDeadline = Math.min(nextDeadline, readyDeadline, liveDeadline);
102
113
  }
103
- nextDeadline = Math.min(nextDeadline, readyDeadline, liveDeadline);
104
- }
105
- for (const socket of context.daemonRegistry.readyRoleSockets()) {
106
- const deadline = daemonLivenessDeadlineMs(context.daemonRegistry.readyAttachment(socket));
107
- if (!Number.isFinite(deadline)) {
108
- await context.invalidateDaemonSocket(socket, "daemon became unresponsive", "invalid daemon liveness timestamp");
109
- continue;
114
+ for (const socket of context.daemonRegistry.readyRoleSockets()) {
115
+ const deadline = daemonLivenessDeadlineMs(context.daemonRegistry.readyAttachment(socket));
116
+ if (!Number.isFinite(deadline)) {
117
+ await context.invalidateDaemonSocket(socket, "daemon became unresponsive", "invalid daemon liveness timestamp");
118
+ continue;
119
+ }
120
+ nextDeadline = Math.min(nextDeadline, deadline);
110
121
  }
111
- nextDeadline = Math.min(nextDeadline, deadline);
112
- }
113
- await writeEarliestRuntimeAlarm({
114
- storage: context.storage,
115
- nextDeadline,
116
- now,
117
- onError: context.onScheduleError,
118
- onMutation: context.onAlarmMutation,
119
- });
122
+ nextDeadline = Math.min(nextDeadline, await pendingAlarmDeadline(context, now));
123
+ await writeEarliestRuntimeAlarm({
124
+ storage: context.storage, nextDeadline, now,
125
+ onError: context.onScheduleError, onMutation: context.onAlarmMutation,
126
+ });
127
+ } catch (error) { try { context.onScheduleError(error); } catch {} }
120
128
  }
121
129
 
122
130
  async function pendingAlarmDeadline(context: RuntimeAlarmContext, now: number): Promise<number> {
@@ -0,0 +1,18 @@
1
+ interface WorkerTaskContext {
2
+ waitUntil(promise: Promise<unknown>): void;
3
+ }
4
+
5
+ export function retainWorkerTask(
6
+ context: WorkerTaskContext,
7
+ task: Promise<unknown>,
8
+ onError: (error: unknown) => void,
9
+ ): Promise<void> {
10
+ const retained = Promise.resolve(task).then(
11
+ () => undefined,
12
+ (error) => {
13
+ try { onError(error); } catch { /* diagnostic callbacks must not create a second rejection */ }
14
+ },
15
+ );
16
+ context.waitUntil(retained);
17
+ return retained;
18
+ }