machine-bridge-mcp 3.0.0-beta.144 → 3.0.0-beta.146
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +12 -0
- package/browser-extension/manifest.json +1 -1
- package/docs/ARCHITECTURE.md +1 -1
- package/docs/AUDIT.md +20 -0
- package/docs/MANAGED_JOBS.md +1 -1
- package/docs/OPERATIONS.md +3 -3
- package/docs/TESTING.md +1 -1
- package/docs/TOOL_REFERENCE.md +1 -1
- package/package.json +1 -1
- package/src/local/durable-process-initial-settlement.mjs +5 -7
- package/src/local/relay-connection.mjs +7 -3
- package/src/local/relay-diagnostics.mjs +34 -0
- package/src/local/relay-peer-diagnostics.mjs +47 -0
- package/src/local/runtime-tool-handlers.mjs +4 -1
- package/src/shared/server-metadata.json +3 -3
- package/src/shared/tool-catalog.json +1 -1
- package/src/worker/daemon-relay-diagnostics.ts +103 -0
- package/src/worker/index.ts +1 -1
- package/src/worker/server-info-tool-delivery.ts +2 -0
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,17 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 3.0.0-beta.146 - 2026-08-27
|
|
4
|
+
|
|
5
|
+
- Fix the hosted `start_job` initial-settlement result boundary exposed by the live beta.145 candidate. The shared settlement helper had explicitly copied process-carrier-only acceptance fields onto ordinary managed-job results; those properties were absent on `start_job` and therefore became JavaScript `undefined`. Machine Bridge's real tool-result normalization intentionally rejects `undefined` as non-JSON, so the job was durably accepted and could complete while the initiating hosted call returned a non-retryable `internal_error`. The helper now relies on the original accepted-object spread to preserve only fields that actually exist instead of manufacturing absent properties.
|
|
6
|
+
- Extend the regression through `normalizeToolResult()` rather than testing the bound handler result alone. This reproduces the exact boundary missed by beta.145: a short hosted `start_job` must coalesce terminal state, preserve its recovery envelope, and remain valid JSON with no unsupported value. Existing process-carrier initial settlement remains unchanged and still returns its defined `execution_mode`, `source_tool`, execution-timeout, and retry-safety fields.
|
|
7
|
+
- Reject beta.145 after live activation even though its exact candidate passed 131/131 frozen verification, install-only validation, service/Worker activation, and the activated-package OAuth canary. Two independent live hosted `start_job` probes returned `internal_error`; idempotency evidence and `read_job` proved the underlying canary job had actually executed successfully. No beta.145 candidate acceptance is recorded. Package/runtime identity advances to beta.146; hosted tool schema generation remains 15 because the public gen15 contract is unchanged and beta.146 repairs its implementation. Fresh frozen verification, candidate/install-only proof, guarded activation, activated-package canary, live short-`start_job` verification, acceptance, and exact-head provider gates are required. npm publication remains separately owner-authorized.
|
|
8
|
+
|
|
9
|
+
## 3.0.0-beta.145 - 2026-08-27
|
|
10
|
+
|
|
11
|
+
- Preserve the causal evidence for repeated awake WebSocket rebuilds instead of retaining only the latest reconnect. The local relay now keeps a newest-first `recent_outages` ring capped at eight completed reconnect episodes with bounded timestamps/durations, close/error classes, prior-ready duration/silence, coarse application-route class, and connection-stage timings. The daemon hello sanitizes that history, and Worker ready promotion synthesizes the just-completed current episode that could not yet have appeared in the pre-ready hello. Pong/application-confirmation near misses that recover without replacing WSS remain heartbeat diagnostics and are deliberately excluded from the completed-outage ring.
|
|
12
|
+
- Reduce another reproduced host-event-density amplifier without shortening durable work. Hosted `start_job` now shares the existing two-second optional managed-job initial-settlement path already used by one-step durable process carriers. A short job that reaches terminal state inside the original response returns its result with `follow_up_read_required=false`; an active job keeps the same `job_id`/recovery envelope with `follow_up_read_required=true`. Dependency waiting, thirty-minute resource admission, six-hour step ceilings, reconnect/replay safety, and local/stdio behavior remain unchanged.
|
|
13
|
+
- Supersede beta.144 because the owner-reported interruption class remains blocking and these changes alter packaged relay diagnostics, hosted `start_job` delivery semantics, tool descriptions, and shipped documentation. Package/runtime identity advances to beta.145 and hosted tool schema generation advances to 15. Fresh frozen verification, candidate/install-only proof, guarded activation, activated-package OAuth canary, live observation, acceptance, and exact-head provider gates are required before GitHub prerelease publication; any later npm publication still requires separate owner authorization, and a published beta.145 activation would restart the major-prerelease soak interval.
|
|
14
|
+
|
|
3
15
|
## 3.0.0-beta.144 - 2026-08-27
|
|
4
16
|
|
|
5
17
|
- Fix a real Windows managed-job dependency recovery defect exposed by exact-main provider CI after beta.143 acceptance. Runner-exit reconciliation already treats transient `permission_denied`/conflict/timeout/resource-unavailable failures as retryable, but a concurrently waiting downstream previously treated any secure upstream status-read error as permanent `dependency_unavailable`. Dependency polling now gives only `permission_denied`, `identity_changed`, and generic `resource_unavailable` reads a fixed 45-second monotonic recovery grace; a successful secure read clears that grace immediately, persistent unavailability still fails closed, and missing/integrity/witness-invalid evidence remains non-retryable.
|
|
@@ -30,6 +30,6 @@
|
|
|
30
30
|
"action": {
|
|
31
31
|
"default_title": "Machine Bridge Browser"
|
|
32
32
|
},
|
|
33
|
-
"version_name": "3.0.0-beta.
|
|
33
|
+
"version_name": "3.0.0-beta.146",
|
|
34
34
|
"key": "MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAxryYkpZhq8+VAQLHcGS9BAHQcyKX8RHGIpIwvtIVRU/rcOcE0bNdnM0aZJ/h6xWQsGDHlhvjT2+1aJaAn/9k8473BRWajzVXld961CdHYVFVHoce2hHiSJ0xydWrHMMZhAm0mN0UzjEpgZ0tMw209efcZHIvSwuxhteZMRy4kyiVjwFlOf5oXFCxRuCJnPj3AK9CmCf4XgEBuPIJ0TZmjGHOOdBvJmbCNnAWXYEo5/mf7MfCGhV4IJ1hNuhpoNQfOFKMUcw9/v/IpT62XpfXdGYTfGYCmCjC+gntK1spbkr2P4/2+sYMQtLpse71mpSNGXfcf3abU55Vpn+gncSxRQIDAQAB"
|
|
35
35
|
}
|
package/docs/ARCHITECTURE.md
CHANGED
|
@@ -294,7 +294,7 @@ Worker-name mutation is a separate identity transition. Existing state rejects a
|
|
|
294
294
|
|
|
295
295
|
The local `RelayConnection` treats proxy selection, transport construction, WebSocket open, authentication, end-to-end readiness, and outage recovery as separate states. The shared proxy module maps WebSocket targets to standard HTTP(S) environment-proxy resolution, honors `NO_PROXY`, rejects non-HTTP(S) proxy schemes, and creates the proxy agent without exposing its URL or credentials. Invalid proxy configuration is a fatal configuration error rather than a retryable outage.
|
|
296
296
|
|
|
297
|
-
A connection-attempt deadline terminates sockets stuck in `CONNECTING`. After open, the daemon sends `hello`; `hello_ack` establishes an authenticated relay generation and starts liveness monitoring, but does not resolve startup or advertise readiness. The Worker then sends a random `relay_probe`. Its result must traverse the local runtime dispatcher and `sendForSession` on that generation before `ready_ack` marks the connection usable. The daemon rejects ordinary tool calls before that state and rejects a premature readiness acknowledgement without locally recorded probe delivery. Independent handshake and readiness deadlines terminate candidates that authenticate but cannot return results. Once authenticated, `RelayLiveness` sends a protocol-level WebSocket Ping every five seconds. The full ten-second Pong deadline begins only after the sender write callback proves the control frame left the local queue. Ping/application-send completions are fenced to the exact current WebSocket generation, so a callback from a superseded socket cannot mutate the replacement connection; a Ping callback that never completes has its own thirty-second local dispatch bound and becomes `relay_transport_send_timeout` even when the receive direction remains active. Reaching that first-stage deadline no longer immediately kills a ready WSS: `relay-transport-confirmation.mjs` opens one bounded fifteen-second application-confirmation window and emits the existing JSON heartbeat only after the connection is fully ready. A later protocol Pong or the explicit JSON application `pong` is bidirectional transport proof and clears suspicion; unrelated application messages update receive-side liveness only and cannot prove that daemon-to-Worker writes are succeeding. `ResilientRelayConnection` concurrently prewarms signed HTTPS in standby without Worker-side takeover; confirmed WSS recovery stops that prewarm, while a real disconnect upgrades it to exact-generation takeover and preempts any stale standby request. A scheduling-responsive true black hole therefore remains bounded to one five-second probe interval plus ten-second Pong response and fifteen-second independent confirmation, while a single ten-to-fifteen-second persistent-flow stall no longer becomes an avoidable reconnect storm. A detected local event-loop stall cancels remote suspicion and follows the separate recovery-grace branch rather than being counted as network failure. Transport Ping remains active during authenticated probing, but application heartbeat/confirmation is gated on verified readiness because the Worker probing state accepts only the readiness-probe result. Same-instance reconnect also performs explicit call-ownership reconciliation: the Worker sends its still-waiting IDs; the daemon snapshots the union of active calls and unacknowledged results, completes replacement-channel readiness, then returns `resume_calls_ack.missing_ids` only for IDs absent from that ownership union. Those IDs alone may receive one same-ID transport redelivery inside the original remaining deadline; if that cannot be done safely they settle retryably with `side_effects_started=false`. Active calls and retained terminal results continue on existing ownership, and no possibly executed tool call is automatically replayed. A separate JSON application heartbeat remains at twenty-five seconds and retains a seventy-five-second application-silence timeout, so protocol-level Pong cannot mask a Worker application path that has stopped responding. On the Worker, authenticated application-heartbeat activity refresh is synchronous and its JSON `pong` is queued before Durable Object alarm inspection or mutation; heartbeat and terminal-result paths then own one explicit coalesced alarm schedule. Persistent deadline work therefore cannot sit in front of application-liveness acknowledgement or be scheduled twice through a hidden touch helper. The Worker's ninety-second daemon-liveness deadline remains an independent wider fallback across Durable Object hibernation. Worker error codes `daemon_transport_error` and `daemon_liveness_timeout` are connection-recovery signals, not protocol incompatibility: they terminate only the current socket, run normal disconnect cleanup, and enter bounded reconnect. Local protocol-watchdog expiry is classified separately as `relay_transport_timeout`; local application-silence expiry retains `relay_heartbeat_timeout`. The same Worker classification is applied to the close frame if the preceding error frame is lost. Failure to send `hello`, or failure to deliver a readiness-probe result because its socket/session ended, follows the same transport-recovery path rather than being promoted to authentication or protocol failure. Unknown Worker errors, authentication rejection, malformed readiness sequencing, and identity/version mismatch remain fatal. Outage reminders run on their own exponential-backoff timer rather than depending on another transport callback. Each new authenticated hello also carries a schema-versioned, bounded
|
|
297
|
+
A connection-attempt deadline terminates sockets stuck in `CONNECTING`. After open, the daemon sends `hello`; `hello_ack` establishes an authenticated relay generation and starts liveness monitoring, but does not resolve startup or advertise readiness. The Worker then sends a random `relay_probe`. Its result must traverse the local runtime dispatcher and `sendForSession` on that generation before `ready_ack` marks the connection usable. The daemon rejects ordinary tool calls before that state and rejects a premature readiness acknowledgement without locally recorded probe delivery. Independent handshake and readiness deadlines terminate candidates that authenticate but cannot return results. Once authenticated, `RelayLiveness` sends a protocol-level WebSocket Ping every five seconds. The full ten-second Pong deadline begins only after the sender write callback proves the control frame left the local queue. Ping/application-send completions are fenced to the exact current WebSocket generation, so a callback from a superseded socket cannot mutate the replacement connection; a Ping callback that never completes has its own thirty-second local dispatch bound and becomes `relay_transport_send_timeout` even when the receive direction remains active. Reaching that first-stage deadline no longer immediately kills a ready WSS: `relay-transport-confirmation.mjs` opens one bounded fifteen-second application-confirmation window and emits the existing JSON heartbeat only after the connection is fully ready. A later protocol Pong or the explicit JSON application `pong` is bidirectional transport proof and clears suspicion; unrelated application messages update receive-side liveness only and cannot prove that daemon-to-Worker writes are succeeding. `ResilientRelayConnection` concurrently prewarms signed HTTPS in standby without Worker-side takeover; confirmed WSS recovery stops that prewarm, while a real disconnect upgrades it to exact-generation takeover and preempts any stale standby request. A scheduling-responsive true black hole therefore remains bounded to one five-second probe interval plus ten-second Pong response and fifteen-second independent confirmation, while a single ten-to-fifteen-second persistent-flow stall no longer becomes an avoidable reconnect storm. A detected local event-loop stall cancels remote suspicion and follows the separate recovery-grace branch rather than being counted as network failure. Transport Ping remains active during authenticated probing, but application heartbeat/confirmation is gated on verified readiness because the Worker probing state accepts only the readiness-probe result. Same-instance reconnect also performs explicit call-ownership reconciliation: the Worker sends its still-waiting IDs; the daemon snapshots the union of active calls and unacknowledged results, completes replacement-channel readiness, then returns `resume_calls_ack.missing_ids` only for IDs absent from that ownership union. Those IDs alone may receive one same-ID transport redelivery inside the original remaining deadline; if that cannot be done safely they settle retryably with `side_effects_started=false`. Active calls and retained terminal results continue on existing ownership, and no possibly executed tool call is automatically replayed. A separate JSON application heartbeat remains at twenty-five seconds and retains a seventy-five-second application-silence timeout, so protocol-level Pong cannot mask a Worker application path that has stopped responding. On the Worker, authenticated application-heartbeat activity refresh is synchronous and its JSON `pong` is queued before Durable Object alarm inspection or mutation; heartbeat and terminal-result paths then own one explicit coalesced alarm schedule. Persistent deadline work therefore cannot sit in front of application-liveness acknowledgement or be scheduled twice through a hidden touch helper. The Worker's ninety-second daemon-liveness deadline remains an independent wider fallback across Durable Object hibernation. Worker error codes `daemon_transport_error` and `daemon_liveness_timeout` are connection-recovery signals, not protocol incompatibility: they terminate only the current socket, run normal disconnect cleanup, and enter bounded reconnect. Local protocol-watchdog expiry is classified separately as `relay_transport_timeout`; local application-silence expiry retains `relay_heartbeat_timeout`. The same Worker classification is applied to the close frame if the preceding error frame is lost. Failure to send `hello`, or failure to deliver a readiness-probe result because its socket/session ended, follows the same transport-recovery path rather than being promoted to authentication or protocol failure. Unknown Worker errors, authentication rejection, malformed readiness sequencing, and identity/version mismatch remain fatal. Outage reminders run on their own exponential-backoff timer rather than depending on another transport callback. Each new authenticated hello also carries a schema-versioned, bounded relay diagnostic summary, including `previous_ready_inbound_silence_ms` so the silent half-open interval before close is not collapsed into the shorter close-to-ready outage duration. The local connection additionally keeps `recent_outages`, a newest-first in-memory ring capped at eight completed reconnect episodes. The ring contains only bounded outage numbers/timestamps/durations, close/error classes, prior-ready duration/silence, coarse application-route class, and connection-stage timings; a protocol/application Pong that clears transport suspicion without rebuilding WSS is not a completed outage and stays only in heartbeat diagnostics. Because the hello is sent before the recovering WebSocket can receive its final `ready_ack`, its carried ring cannot yet contain that current episode. The Worker therefore sanitizes the prior ring into the probing attachment and, only when promotion proves end-to-end readiness, synthesizes the now-completed current episode from the already bounded scalar fields before marking `outage_active=false`. Authenticated `server_info.daemon.relay_transport` can consequently retain several sub-warning-threshold reconnects instead of losing all but the latest one, without logging them by default, persisting a tool transcript, exposing endpoints/interfaces/DNS data, or treating near-miss liveness suspicions as outages.
|
|
298
298
|
|
|
299
299
|
Reconnect uses bounded exponential backoff with jitter. Brief self-healing interruptions are debug-only. An unresolved outage is promoted to a rate-limited warning after a grace period, and recovery produces one summary. Raw close codes and reason strings remain debug-only.
|
|
300
300
|
|
package/docs/AUDIT.md
CHANGED
|
@@ -1,5 +1,25 @@
|
|
|
1
1
|
# Security and privacy audit notes
|
|
2
2
|
|
|
3
|
+
## 2026-08-27 beta.146 live start-job result-boundary follow-up
|
|
4
|
+
|
|
5
|
+
**Beta.145 is rejected by live behavior even though its underlying durable work remained healthy.** The exact candidate passed the 131-task frozen plan, candidate/install-only validation, guarded Worker/login-daemon activation, and the activated-package OAuth canary. The canary's initiating hosted `start_job` nevertheless returned public `internal_error`. Reusing its idempotency key with a deliberately different plan produced the expected conflict containing the already-bound managed-job ID; reading that exact ID proved the original canary job had completed successfully. A second harmless zero-exit hosted `start_job` independently reproduced the same public `internal_error`, while `exec_command` continued to use the shared two-second settlement path successfully. This isolates the regression to ordinary managed-job result construction rather than Worker/OAuth availability, the managed-job runner, process-carrier settlement, relay transport, or resource admission.
|
|
6
|
+
|
|
7
|
+
**The causal bug was an invalid JSON value introduced after durable acceptance.** `settleManagedJobAcceptance()` spread the accepted managed-job projection and the optional read result, then explicitly reassigned four process-carrier-only fields: `execution_mode`, `source_tool`, `execution_timeout_seconds`, and `retry_safety`. Those fields exist on `exec_command`/`run_process`/`run_local_command` durable-process acceptance but not on ordinary `start_job`; the explicit assignments therefore created own properties whose values were JavaScript `undefined`. `normalizeToolResult()` deliberately rejects `undefined`, functions, symbols, bigint, and non-finite numbers instead of silently mutating the public result. The local tool executor consequently converted an otherwise successful accepted/settled job response into `internal_error` before relay result delivery. The fix removes the redundant reassignments: the original accepted-object spread already preserves process-only fields when they genuinely exist and creates nothing when they do not.
|
|
8
|
+
|
|
9
|
+
**The previous regression test stopped one layer too early.** Beta.145 directly invoked the bound hosted `start_job` handler and asserted the expected terminal/recovery fields, but did not pass that object through the production JSON result boundary. Beta.146 extends that test through `normalizeToolResult()` and requires the recovery envelope plus `follow_up_read_required=false` to survive normalization. This would deterministically fail on beta.145's `undefined` properties. The public generation-15 contract itself does not change, so schema generation remains 15; package/runtime identity advances to beta.146 because the shipped implementation changes.
|
|
10
|
+
|
|
11
|
+
**Release consequence.** No beta.145 candidate acceptance is written. Beta.146 must repeat frozen full verification, exact candidate preparation/install-only proof, guarded activation, activated-package OAuth canary, a harmless live short-`start_job` probe that returns a JSON-valid terminal response in the initiating call, broader live relay/version observation, acceptance, and exact-head provider gates. npm publication remains the sole separate current-task owner authorization boundary. A later published beta.146 registry activation would start a new seven-day major-prerelease soak; neither beta.144 nor rejected beta.145 time carries forward.
|
|
12
|
+
|
|
13
|
+
## 2026-08-27 beta.145 repeated awake interruption observability and event-density follow-up
|
|
14
|
+
|
|
15
|
+
**Live beta.144 evidence separates a remaining physical transport class from a separate host-settlement amplifier.** The current daemon process accumulated three real WebSocket outage generations while the machine was awake. The latest retained episode was a brief close-1006 recovery after a much longer inbound-silence interval; the preceding two sub-warning-threshold reconnects had to be reconstructed indirectly because runtime state kept only cumulative `outage_count` plus the latest outage scalars. Full Wi-Fi baseline comparison falsified Wi-Fi degradation as either a necessary or sufficient condition, visible Tailscale path-change events did not consistently precede the relay silence, and the Karing system extension emitted no matching upstream timeout/reset evidence. An earlier independent event had already aligned relay 1006 with an unrelated HTTPS `unexpected EOF`, so shared system-network/VPN/TUN failure remains a supported class, but the available evidence still does not identify a specific tunnel implementation, proxy node, ISP, Cloudflare edge, or upstream hop as the initiator of every awake reset.
|
|
16
|
+
|
|
17
|
+
**The missing causal artifact was multi-episode history, not another reconnect timeout.** `RelayConnection` now retains at most eight completed reconnect episodes in memory, newest first. Each `recent_outages` entry contains only an outage sequence number, bounded disconnect/ready timestamps and duration, retry count, enumerated close/transport-error classes, previous-ready duration/silence, coarse application proxy-route class, and bounded DNS/TCP/TLS/WebSocket connection-stage timings. It contains no endpoint, interface name, address, DNS answer, certificate, close reason, account/call identity, tool argument, or result. A first-stage Pong timeout that is disproved by protocol/application confirmation without replacing the socket is not inserted, so liveness near misses remain distinguishable from actual WSS generations. The daemon hello sanitizes the existing ring before transmission. Because that hello precedes the recovering connection's final `ready_ack`, Worker promotion alone has the evidence needed to synthesize the just-completed current episode; it prepends that bounded entry and caps the attachment at eight instead of leaving authenticated `server_info` one reconnect behind.
|
|
18
|
+
|
|
19
|
+
**The same investigation exposed a narrower event-density gap rather than a defect in the existing process carriers.** The one-step `exec_command`/`run_process`/`run_local_command` path already waits up to two seconds for a newly accepted durable job to settle, eliminating many helper-plus-`read_job` pairs. The investigation still produced many explicit short multi-step `start_job` submissions followed immediately by `read_job`, which is precisely the shape the existing host-event-density guidance tries to avoid. Beta.145 factors that optional settlement routine into a managed-job helper and routes hosted `start_job` through it. A terminal job returns its result in the original response with `follow_up_read_required=false`; an active or temporarily unreadable job preserves the original acceptance/recovery envelope and continues through normal paced `read_job`. Local/stdio submission is unchanged, and no execution, dependency, admission, reconnect, replay, or six-hour step deadline is shortened or extended. Machine Bridge still cannot observe the exact external ChatGPT host condition that terminates a turn, so lower event density is a bounded mitigation for a reproduced amplifier rather than a claim to have fixed an unobservable host timeout.
|
|
20
|
+
|
|
21
|
+
**Release consequence.** The owner continues to report the interruption class during the beta.144 soak, so it is treated as blocking rather than recording soak success. These shipped runtime/catalog/documentation changes advance package identity to beta.145 and hosted tool schema generation to 15. Beta.145 must repeat frozen verification, exact candidate preparation/install-only proof, guarded Worker/service activation, activated-package OAuth canary, live relay and short-`start_job` verification, acceptance, and exact-head provider checks. npm publication remains the sole separate current-task owner authorization boundary. If beta.145 is later published and registry-verified activation completes, the major-prerelease seven-day soak begins again from that activation; elapsed beta.144 soak time is not carried forward.
|
|
22
|
+
|
|
3
23
|
## 2026-08-27 beta.144 dependency-state recovery and CI fixture follow-up
|
|
4
24
|
|
|
5
25
|
**Exact-main provider CI separated a test-harness race from a packaged Windows recovery defect.** Ubuntu/full twice timed out waiting for a Worker WebSocket `tool_call`, first with a five-second fixture bound and again after a test-only increase to ten seconds. Source tracing showed eight fixtures triggered `currentMcpCall()`/`toolCallRequest()` before calling `waitForWsMessage()`. The HTTP fetch starts immediately, while the waiter has no historical message buffer, so a fast relay dispatch could arrive before listener registration and be lost forever. Registering the waiter first removed all eight request-before-listener sites; three local Worker integration runs passed and hosted Ubuntu/full then passed. This is test evidence, not proof that production Worker dispatch was dropping messages.
|
package/docs/MANAGED_JOBS.md
CHANGED
|
@@ -15,7 +15,7 @@ Machine Bridge managed jobs reduce this failure mode by accepting the complete e
|
|
|
15
15
|
|
|
16
16
|
Recovery inventory is deliberately ordered by recoverability rather than recency alone. `list_jobs.jobs` returns at most 50 primary records, with unreadable state, active jobs, staged plans, and durable terminal history ahead of transient one-step helper history so a burst of short helpers cannot push an older long-running or pre-spawn-waiting job out of that bounded primary window. A separate `recent_process_recovery` array returns at most 16 authority-visible public handles for recent one-step terminal jobs that are still inside the thirty-minute recovery reserve but omitted from `jobs`; callers can recover the `job_id` there and continue with `read_job`. It contains no step output or internal retention class and is recovery inventory, not a polling or MCP replay/session surface. Owner/local listings also include only aggregate recent creation/churn counts; the internal `transient_process` retention class remains hidden per job.
|
|
17
17
|
|
|
18
|
-
Remote one-step `exec_command`, `run_process`, and `run_local_command` requests still become durable managed jobs before execution.
|
|
18
|
+
Remote one-step `exec_command`, `run_process`, and `run_local_command` requests still become durable managed jobs before execution. Hosted `start_job` now shares the same short two-second initial-settlement check after its explicit multi-step plan has been durably accepted. A job that reaches terminal state inside that window can therefore return its managed-job result in the original tool response with `follow_up_read_required=false`, eliminating the usual immediate second `read_job` event. A job that remains active keeps the same durable `job_id` and recovery envelope with `follow_up_read_required=true`; execution timeouts, dependency waiting, the pre-spawn resource-admission allowance, finally behavior, and later recovery semantics are unchanged. Local/stdio `start_job` does not add this hosted response wait.
|
|
19
19
|
|
|
20
20
|
This mechanism does **not** bypass host, operating-system, or endpoint-security policy. A job snapshots its execution policy/environment at acceptance; later profile changes affect new jobs, while accepted jobs continue until completion or explicit cancellation. The initial `start_job` request must still be permitted, and every local child process remains subject to local security controls.
|
|
21
21
|
|
package/docs/OPERATIONS.md
CHANGED
|
@@ -10,11 +10,11 @@ machine-mcp service status
|
|
|
10
10
|
|
|
11
11
|
Routine remote checks should use authenticated `server_info` with `detail: "summary"`; request the default/full projection only when the caller's authority permits and exact effective-tool, OAuth/account, or detailed owner observability is actually needed. Non-owner full responses intentionally retain hidden markers/counts instead of cross-principal activity, resource aliases, stable device-key identity, or daemon-only tool names. Remote `diagnose_runtime` is owner-only because its fixed probes expose machine-wide control-plane activity; narrower roles use `server_info`/`project_overview` for authority-scoped readiness and workspace state. `status` prints redacted profile state and verifies the deployed Worker version. Resource source paths remain redacted. `doctor` checks Node.js, the package-installed Wrangler binary, Cloudflare login, Worker health, the configured policy, the automatic-without-per-operation-prompts authorization model, and the same fixed local filesystem/process/shell/job-storage/resource probes exposed to the remote owner by `diagnose_runtime`. It constructs an isolated local runtime: `diagnosticScope.running_service_process_inspected=false` and `remote_relay_inspected=false` are deliberate, so a green doctor result is not evidence that the launchd/systemd/Scheduled Task daemon retained its Worker WebSocket. Inspect authenticated `server_info.daemon.relay_transport` for the running service relay. Authenticated `server_info.authorization.execution_model` reports the authority contract and identifies whether the account has daemon-OS-user ambient authority. Public `/healthz` output contains only server identity and version; daemon details require an authenticated `server_info` call.
|
|
12
12
|
|
|
13
|
-
For interruption analysis, prefer one owner `diagnose_runtime` call over a chain of inventory probes. It now includes `runtime.managed_jobs.recent_activity`, `runtime.security_audit.recent_activity`, bounded `runtime.resource_admission.waiters.diagnostics`, `runtime.system_sleep`, `runtime.event_loop_pause_analysis`, and `runtime.relay_outage_analysis`. The audit aggregate contains only counts, bounded tool names, failure totals, and calls-per-minute density derived from the existing content-free hash-chained audit log; it contains no tool arguments or results. Its `coverage=daemon_reached_relay_tool_calls_only` and `host_side_events_observable=false` fields make the evidence boundary explicit: host-only discovery/control-plane/final-delivery events are not counted. The waiter projection reports only resource-request shape and the current admission reason. On macOS with shell-capable owner diagnostics, the fixed power probe reduces `pmset` history to a small list of sleep start/end/duration/reason classes; it never returns raw power-log lines. `event_loop_pause_analysis.classification=matched_system_sleep` requires both the recorded runtime-stall end time and duration to match one of those bounded operating-system sleep intervals within a fixed tolerance. `relay_outage_analysis` separately compares the most recent completed `last_disconnected_at` -> `last_ready_at` interval with the same bounded sleep history and reports exact overlap duration/ratio. `majority_system_sleep_overlap` means at least half of that observed relay outage occurred while macOS was suspended, so a retained `connection_reset` is transport aftermath rather than sufficient evidence of a separate network root cause. `no_matching_recent_system_sleep` leaves an awake reset unassigned rather than guessing a VPN, edge, Worker, or host cause. ChatGPT host-turn termination/final-message receipt remains explicitly unobservable.
|
|
13
|
+
For interruption analysis, prefer one owner `diagnose_runtime` call over a chain of inventory probes. It now includes `runtime.managed_jobs.recent_activity`, `runtime.security_audit.recent_activity`, bounded `runtime.resource_admission.waiters.diagnostics`, `runtime.system_sleep`, `runtime.event_loop_pause_analysis`, and `runtime.relay_outage_analysis`. The audit aggregate contains only counts, bounded tool names, failure totals, and calls-per-minute density derived from the existing content-free hash-chained audit log; it contains no tool arguments or results. Its `coverage=daemon_reached_relay_tool_calls_only` and `host_side_events_observable=false` fields make the evidence boundary explicit: host-only discovery/control-plane/final-delivery events are not counted. The waiter projection reports only resource-request shape and the current admission reason. On macOS with shell-capable owner diagnostics, the fixed power probe reduces `pmset` history to a small list of sleep start/end/duration/reason classes; it never returns raw power-log lines. `event_loop_pause_analysis.classification=matched_system_sleep` requires both the recorded runtime-stall end time and duration to match one of those bounded operating-system sleep intervals within a fixed tolerance. `relay_outage_analysis` separately compares the most recent completed `last_disconnected_at` -> `last_ready_at` interval with the same bounded sleep history and reports exact overlap duration/ratio. `majority_system_sleep_overlap` means at least half of that observed relay outage occurred while macOS was suspended, so a retained `connection_reset` is transport aftermath rather than sufficient evidence of a separate network root cause. `no_matching_recent_system_sleep` leaves an awake reset unassigned rather than guessing a VPN, edge, Worker, or host cause. Full relay diagnostics additionally expose `recent_outages`, a newest-first in-memory history capped at eight completed WebSocket reconnect episodes. Each entry contains only bounded outage numbering/timestamps/durations, close/error classes, previous-ready duration/silence, coarse application-route class, and connection-stage timings. A protocol/application Pong that clears a liveness suspicion without rebuilding the WebSocket remains heartbeat evidence and is deliberately absent from `recent_outages`; the array is therefore reconnect history, not a list of every transient transport suspicion. ChatGPT host-turn termination/final-message receipt remains explicitly unobservable.
|
|
14
14
|
|
|
15
15
|
`runtime.idle_sleep_guard` also carries coarse `last_activity_started_at`, `last_activity_ended_at`, `grace_release_due_at`, `last_release_at`, `last_release_reason`, active-activity count, and `requests_system_sleep_prevention_on_ac`. These fields are diagnostic ownership evidence only. On macOS the fixed child now runs `/usr/bin/caffeinate -i -s -w <owner-pid>`: `-i` retains the existing Idle Sleep prevention, while Apple's `caffeinate` contract makes `-s` a stronger system-sleep assertion only while the Mac is on AC power. On battery, do not interpret the `-s` request as a guarantee against system sleep. Explicit sleep, lid-close policy, power loss, and operating-system behavior outside those assertion contracts remain external boundaries. A long-running managed job has its separate runner-owned assertion, while ordinary authorized relay activity retains the existing thirty-minute inactivity grace. Do not extend that grace merely because a later host failure is unobservable: a permanent or multi-hour assertion after every remote call would trade battery behavior for an unsupported host-lifecycle assumption.
|
|
16
16
|
|
|
17
|
-
Remote
|
|
17
|
+
Remote durable carriers also reduce event amplification at source. After durable acceptance, the original `exec_command`, `run_process`, or `run_local_command` response waits up to `server_info.tool_delivery.remote_process_initial_settlement_wait_ms` for a short helper to settle, while hosted `start_job` uses the same underlying two-second window exposed separately as `remote_managed_job_initial_settlement_wait_ms`. If the accepted job becomes terminal inside that window, terminal status/result are returned immediately with `follow_up_read_required=false`; otherwise the response retains the same durable recovery envelope and `follow_up_read_required=true`. This response coalescing changes only the common helper-plus-read event count; it does not reduce the 600-second one-step process execution budget, the thirty-minute pre-spawn admission allowance, dependency waiting, or the six-hour managed-job step ceiling.
|
|
18
18
|
|
|
19
19
|
Repository-controlled release and publication commands emit terminal failures as one JSON object with exactly `event` and `error` fields. The event is a fixed validated identifier; the error is bounded and passes through portable credential, URL, email, home-path, and control-character redaction before JSON serialization. Treat the object as one log record. Do not replace it with raw subprocess output or interpolate the error into a prefix string when adding a new release path.
|
|
20
20
|
|
|
@@ -93,7 +93,7 @@ After the host path recovers, compare authenticated `server_info`, `machine-mcp
|
|
|
93
93
|
|
|
94
94
|
### Relay interruption messages
|
|
95
95
|
|
|
96
|
-
A reconnect warning proves a transport interruption, not a daemon crash. Compare daemon PID and process start time with `connected_at`, `last_seen_at`, `daemon.relay_transport.last_close_category`, `last_close_code`, `outage_count`, `outage_attempts`, `previous_ready_inbound_silence_ms`, `last_connect_milestones_ms`, and the coarse network-route class. Also distinguish a planned restart from an accidental outage: a current-generation daemon sends `daemon_draining` before relay close, affected calls receive `reason=daemon_planned_drain`, and owner/full `server_info.worker.continuity_evidence` durably retains the planned-drain count/time plus bounded socket-disconnect and client-cancellation observations across Worker isolate replacement. `worker.observability.continuity` remains isolate-local and may reset; use the durable summary for post-incident correlation rather than relying on a later close-1006 inference. For `read_job`, recover with the same returned `job_id`; do not resubmit the job's underlying mutation. `last_connect_milestones_ms` contains only bounded relative timings for the most recent connection attempt phases such as DNS resolution, TCP connect, TLS establishment, HTTP rejection, and WebSocket open; `last_failed_connect_stage`, `last_failed_connect_duration_ms`, `last_failed_connect_milestones_ms`, and `last_failed_connect_http_status` retain the most recent failed attempt even after a later retry succeeds. `last_transport_error_ready` and `last_transport_error_authenticated` distinguish failure of an already-established channel from a pre-readiness connection failure. `last_transport_error_reason` is a strict privacy-safe allowlist (`connection_reset`, `connection_timeout`, `network_unreachable`, bounded DNS/TLS classes, or `unknown`) rather than the raw operating-system message. The signed HTTPS fallback retains its last error class/reason after a later successful poll while resetting the current `http_poll_failures` count, so post-recovery diagnosis can determine whether WSS and HTTPS failed through the same system-network episode. None of these fields contains a hostname, address, DNS answer, certificate, close reason, or proxy endpoint. While no daemon channel is ready, `server_info.daemon.previous_connection` retains only the last verified channel's transport, connected/last-seen/disconnected timestamps, and sanitized relay diagnostics; it excludes policy, tools, account identity, daemon instance/connection identity, call IDs, arguments, and results, and it never participates in routing or authorization. `outage_duration_ms` measures the close-to-ready recovery episode; `previous_ready_inbound_silence_ms` measures how long the preceding ready socket had stopped producing inbound transport proof before it actually closed. The second value is therefore the field that exposes a black-holed OPEN WebSocket whose visible reconnect later completes quickly. A VPN/TUN UI may remain “connected” while its upstream route is unusable. Local OS logs can be compared with the exact `last_disconnected_at` timestamp, but an interface-quality change or tunnel-process correlation is not by itself proof of which product, node, edge, or upstream failed. Correlated failure of an independent HTTPS client at the same timestamp—for example an external API `unexpected EOF` while the relay records WebSocket 1006/`connection_reset`—is stronger evidence of a shared system-network/VPN/TUN episode than of a Machine Bridge event-loop or resource-admission failure; it still does not identify the failing tunnel node or upstream provider. Machine Bridge reports only coarse route/proxy classes and never sends or logs interface names, addresses, DNS answers, proxy credentials, or Worker secrets.
|
|
96
|
+
A reconnect warning proves a transport interruption, not a daemon crash. Compare daemon PID and process start time with `connected_at`, `last_seen_at`, `daemon.relay_transport.last_close_category`, `last_close_code`, `outage_count`, `recent_outages`, `outage_attempts`, `previous_ready_inbound_silence_ms`, `last_connect_milestones_ms`, and the coarse network-route class. Also distinguish a planned restart from an accidental outage: a current-generation daemon sends `daemon_draining` before relay close, affected calls receive `reason=daemon_planned_drain`, and owner/full `server_info.worker.continuity_evidence` durably retains the planned-drain count/time plus bounded socket-disconnect and client-cancellation observations across Worker isolate replacement. `worker.observability.continuity` remains isolate-local and may reset; use the durable summary for post-incident correlation rather than relying on a later close-1006 inference. For `read_job`, recover with the same returned `job_id`; do not resubmit the job's underlying mutation. `last_connect_milestones_ms` contains only bounded relative timings for the most recent connection attempt phases such as DNS resolution, TCP connect, TLS establishment, HTTP rejection, and WebSocket open; `last_failed_connect_stage`, `last_failed_connect_duration_ms`, `last_failed_connect_milestones_ms`, and `last_failed_connect_http_status` retain the most recent failed attempt even after a later retry succeeds. `last_transport_error_ready` and `last_transport_error_authenticated` distinguish failure of an already-established channel from a pre-readiness connection failure. `last_transport_error_reason` is a strict privacy-safe allowlist (`connection_reset`, `connection_timeout`, `network_unreachable`, bounded DNS/TLS classes, or `unknown`) rather than the raw operating-system message. The signed HTTPS fallback retains its last error class/reason after a later successful poll while resetting the current `http_poll_failures` count, so post-recovery diagnosis can determine whether WSS and HTTPS failed through the same system-network episode. None of these fields contains a hostname, address, DNS answer, certificate, close reason, or proxy endpoint. While no daemon channel is ready, `server_info.daemon.previous_connection` retains only the last verified channel's transport, connected/last-seen/disconnected timestamps, and sanitized relay diagnostics; it excludes policy, tools, account identity, daemon instance/connection identity, call IDs, arguments, and results, and it never participates in routing or authorization. `outage_duration_ms` measures the close-to-ready recovery episode; `previous_ready_inbound_silence_ms` measures how long the preceding ready socket had stopped producing inbound transport proof before it actually closed. The second value is therefore the field that exposes a black-holed OPEN WebSocket whose visible reconnect later completes quickly. A VPN/TUN UI may remain “connected” while its upstream route is unusable. Local OS logs can be compared with the exact `last_disconnected_at` timestamp, but an interface-quality change or tunnel-process correlation is not by itself proof of which product, node, edge, or upstream failed. Correlated failure of an independent HTTPS client at the same timestamp—for example an external API `unexpected EOF` while the relay records WebSocket 1006/`connection_reset`—is stronger evidence of a shared system-network/VPN/TUN episode than of a Machine Bridge event-loop or resource-admission failure; it still does not identify the failing tunnel node or upstream provider. Machine Bridge reports only coarse route/proxy classes and never sends or logs interface names, addresses, DNS answers, proxy credentials, or Worker secrets.
|
|
97
97
|
|
|
98
98
|
Brief retryable outages recover automatically. On a verified current daemon channel, `server_info.daemon.relay_transport.outage_active=false`; retained fields describe the immediately preceding transport episode rather than claiming a current outage. WebSocket remains preferred and requests a protocol-level probe after five seconds. Calling `ws.ping()` only queues the control frame; it is not treated as remote-probe dispatch until the WebSocket sender's write callback confirms that the Ping actually left the local send queue. The local sender has a separate thirty-second bounded dispatch window, while a confirmed Ping retains its full ten-second Pong deadline. This deliberately prevents compression/backpressure or a slow local socket queue from spending the remote-response budget before any probe was transmitted. A protocol Pong that arrives while a Ping callback is still pending records bidirectional proof for that dispatch round; if the local write callback then completes inside the thirty-second dispatch budget, it does not arm a stale future Pong deadline. Unrelated application inbound is receive-side evidence only and cannot prove the daemon-to-Worker direction. A local queue whose Ping write callback still has not completed after thirty seconds is classified as `relay_transport_send_timeout` even if unrelated inbound traffic continues. One dispatched Ping that reaches the ten-second response deadline does not hard-kill an otherwise ready WSS: the relay runtime enters a fifteen-second `transport_confirmation_pending` window, sends the existing JSON application heartbeat as an independent path check, and prewarms signed HTTPS in standby without taking ownership away from the still-ready WSS. A protocol Pong or the explicit JSON application `pong` clears suspicion and stops standby prewarm; ordinary tool/control traffic does not, because it proves only the Worker-to-daemon receive direction. Only a second-stage confirmation window that receives no application `pong` becomes `relay_transport_timeout` and terminates the WSS. This keeps a true black hole bounded while no longer amplifying a roughly ten-to-fifteen-second persistent-flow stall into an immediate reconnect storm. `heartbeat.probe_dispatch_*`, `heartbeat.transport_confirmation_*`, and bounded sender-backlog fields distinguish local send delay, first-stage response loss, successful second-stage recovery, and confirmed two-stage failure. The separate periodic JSON application heartbeat remains twenty-five seconds with a seventy-five-second application-silence timeout, begins only after verified relay readiness, and the Worker keeps a wider ninety-second WebSocket liveness fallback. This is a detection/recovery bound, not a guarantee that a degraded network can complete another WebSocket handshake inside the same interval. WebSocket connect attempts have a thirty-second outer budget so a degraded but still valid DNS/TCP/TLS/WebSocket upgrade is not misclassified by an unrealistically narrow connection cutoff. The daemon also explicitly disables client `permessage-deflate`: the relay carries bounded control/JSON traffic, while `ws` enables compression by default on clients and compression adds sender-state/CPU overhead that can queue later frames; the stability path does not need that optional negotiation. The fallback still begins independently rather than waiting thirty seconds for WSS. On first-stage WSS liveness suspicion, the same root-certified ephemeral daemon identity prewarms signed HTTPS in standby; if WSS proves live during the second-stage confirmation, that standby poller stops. If the WSS actually disconnects, fallback switches to exact-generation takeover immediately; an in-flight standby request is aborted and replaced rather than being allowed to consume up to its own request deadline before takeover can start. That in-memory session certificate intentionally has a 24-hour maximum lifetime. It is valid for ordinary reconnects during that lifetime, but it is not silently extended: if a later WSS reconnect/authentication attempt discovers that the daemon session has expired, the daemon terminates with `relay_device_session_expired` instead of retrying forever with unusable credentials. Installed launchd/systemd/Windows service supervision restarts the failed daemon and obtains a fresh root-signed session after normal runtime cleanup; the default portable root does this without user interaction. A manually run daemon must be restarted by its operator, and a configured Secure Enclave root retains its existing user-presence requirement when the new daemon start signs the replacement session. Each fallback request has a seven-second deadline, ordinary one-second ready poll cadence, five-second standby-prewarm cadence, bounded one/two/four/five-second retry backoff after request or protocol/session failures, a 750 ms hard minimum request-start interval, and a twelve-second liveness window; a new daemon-backed call waits at most fifteen seconds for some verified daemon channel, and the measured wait is deducted from that call's original execution budget. After an established WSS disappears, the daemon explicitly marks its signed HTTP request as a takeover of the Worker-issued `connection_id` for that exact disconnected WebSocket generation. Once candidate preconditions pass, HTTPS may retire only that targeted same-instance zombie WSS that the Worker has not yet observed closing. If a newer same-instance WSS is already ready before the HTTP request arrives, the old generation no longer matches and the stale takeover remains standby instead of retiring the recovered socket. A takeover request without the exact Worker-issued WebSocket connection ID is invalid rather than being treated as an instance-only legacy takeover. Malformed, stale, wrongly targeted, or different-instance requests cannot preempt a healthy incumbent. During replacement, the daemon reconciles `resume_calls`, processes `ready_ack`, proves local readiness, and only then returns `resume_calls_ack.missing_ids`. A missing ID therefore proves both that the same daemon has no active/unacknowledged-result ownership for that call and that the replacement channel is ready. If the initiating MCP response is still open and at least one second remains in the original execution budget, the Worker may transparently retransmit exactly that same call ID, arguments, authority, and a reduced timeout. `read_job` is stricter: redelivery requires the full ten-second reconciliation headroom to remain, otherwise the Worker declines redelivery and returns retryable recovery failure rather than rewriting the call into an under-budget immediate read. If safe redelivery cannot be accepted, the call falls back to retryable `unavailable` with `side_effects_started=false`. Calls that may have executed, retained terminal results, different-daemon calls, and ambiguous mutations are never automatically replayed. Completed relay results that are still waiting for Worker acknowledgement remain bounded in daemon memory and consume the same recovery-ownership capacity as active calls: 16 total with two control-plane slots reserved for `diagnose_runtime`/`list_roots`. When ordinary recovery ownership reaches 14, another ordinary relay call is rejected before execution with retryable `limit_exceeded` and `side_effects_started=false`; the two reserved diagnostic/recovery calls remain available until total capacity reaches 16. The retained-result implementation also keeps one non-admission emergency ownership slot solely for a violated internal capacity invariant: if an already-executed result reaches retention after the normal 16-entry ceiling is unexpectedly full, that one result remains retained for acknowledgement/reconnect ownership instead of being sent unowned and later misclassified as safe to redeliver. Use of that slot emits an error-level capacity event and may make diagnostics temporarily report ownership above the normal maximum; a second such overflow is not sent. This slot is not usable admission capacity and must never be counted to raise the 16-call execution ceiling. An acknowledgement that is permanently lost cannot pin a result forever: first retention is monotonic and the result expires after the 315-second maximum Worker settlement lifetime on the next live relay heartbeat; the disconnected path still uses the shorter reconnect-grace cleanup. `diagnose_runtime.runtime.relay_result_recovery` exposes only aggregate `active_calls`, `retained_results`, active ownership, and capacity counts—never call IDs, tool arguments, or results. Compare outage intervals with sleep/wake records and `diagnose_runtime.runtime.relay.heartbeat` before classifying them as active network faults; local stdio `server_info.runtime.relay.heartbeat` exposes the same state. On macOS, an authorized, schema-valid remote tool call also enters a bounded idle-sleep guard with a fixed thirty-minute rolling inactivity grace; the service does not depend on shell-only environment overrides that launchd would not persist as configuration, and relay heartbeats do not count as user activity. Authorized relay handlers hold one shared `/usr/bin/caffeinate -i -s -w <daemon-pid>` assertion for their full execution lifetime; `-s` strengthens system-sleep prevention only on AC power while `-i` remains the baseline Idle Sleep request; concurrent handlers share the child, the thirty-minute default inactivity grace begins only after the last one settles, and a new authorized handler cancels any pending release timer so the full grace restarts after that activity settles. A remote `start_process` extends the same assertion only after resource admission succeeds and keeps it until the session child settles, so a long process session is not reduced to the handler grace window. Remote account managed-job runners independently hold `/usr/bin/caffeinate -i -s -w <runner-pid>` after their ownership claim is confirmed and persisted ownership identifies an account-backed job, then retain it through admission, steps, cleanup, and terminal persistence; local managed jobs do not acquire this remote-continuity assertion, and daemon reconnect/replacement does not own the remote runner protection. `diagnose_runtime.runtime.idle_sleep_guard` reports only daemon-side supported/enabled/active/grace/error-class state plus whether the fixed child requests the AC-only system-sleep assertion; it intentionally does not enumerate process-session or job identities. Runtime shutdown terminates process sessions before releasing the daemon guard. None of these assertions claim to prevent explicit sleep or lid-close sleep. A nonzero `event_loop_stall_count` with a large `max_event_loop_lag_ms` means the local daemon was not scheduled promptly; a short stall enters recovery grace, sends a fresh transport probe, and deliberately postpones disconnect. A large stall that aligns with `pmset` Sleep/Wake is suspension evidence. Independently, `relay_outage_analysis` can show that a close-to-ready `connection_reset` interval was itself dominated by system sleep; an awake outage with no sleep overlap remains real transport evidence without identifying which network/host layer caused it. A large `previous_ready_inbound_silence_ms` without a matching local stall remains useful pre-close half-open evidence. Use `--verbose` only when close codes, liveness deadlines, and retry delays are required.
|
|
99
99
|
|
package/docs/TESTING.md
CHANGED
|
@@ -22,7 +22,7 @@ Successful child-task stdout/stderr is intentionally suppressed so a long green
|
|
|
22
22
|
|
|
23
23
|
Remote execution carrier limits are separate from the verification plan itself. `run_process`/`run_local_command`/`exec_command` cap one detached process step at 600 seconds, while a complete `check:full` can legitimately exceed that wall-clock budget when a real install or integration dependency is slow. A remote release verification that may cross 600 seconds must therefore submit `npm run check:full` as a detached `start_job` step with a larger explicit step timeout and inspect it with `read_job`; an ordinary local terminal run is unaffected. The same rule applies below 600 seconds: a composite real-host/lifecycle/fault-injection verification with an uncertain or multi-phase upper bound must not be placed behind an arbitrary 300-second `exec_command` budget and then treat carrier SIGTERM as a test failure. Split independently terminal phases into durable job steps and size each step from the verifier's own bound plus settlement margin. Coherent short verification commands should be grouped into a repository umbrella command or one multi-step managed job rather than emitted as dozens of one-step remote carriers; this reduces host-visible event density without changing any verifier deadline or forcing a user-turn handoff. Hosted `read_job` pacing tests separately require the 40-second default, five-second terminal-detection poll, 30-second minimum nonterminal progress coalescing, hosted-only suppression of `current_step` wakeups, and preservation of local explicit-wait step progress. This does not widen any product foreground/process timeout—it prevents the verification carrier or host event stream from becoming the narrower continuity boundary than the suite being measured.
|
|
24
24
|
|
|
25
|
-
Generation-9 continuity coverage additionally requires short hosted durable-process acceptance to settle inside the original tool response when the job is already terminal, preserve the original `job_id`/recovery envelope if the optional settlement read is unavailable, and leave active jobs on the normal `read_job` path. Managed-job inventory tests must prove more than 50 newer terminal helpers cannot hide an older active job. Resource-admission tests must expose the current bounded waiter reason such as `cpu_pressure_window` without waiter identity/token/PID/path fields, and security-audit tests must aggregate recent tool density without argument/result content. macOS power-diagnostic coverage reduces representative `pmset` lines to bounded sleep intervals and requires a same-duration/same-resume-time runtime stall to classify as `matched_system_sleep` while a nonmatching pause remains unclassified. Idle-sleep-guard tests pin activity/grace/release timestamps and the coarse inactivity-expiry reason without extending the existing thirty-minute grace. These are diagnostic/recovery guarantees, not new timing verdicts for the user task.
|
|
25
|
+
Generation-9 continuity coverage additionally requires short hosted durable-process acceptance to settle inside the original tool response when the job is already terminal, preserve the original `job_id`/recovery envelope if the optional settlement read is unavailable, and leave active jobs on the normal `read_job` path. The same shared initial-settlement helper is now exercised through the actual hosted `start_job` runtime handler: a short terminal multi-step job must return `follow_up_read_required=false`, while active or unavailable optional settlement preserves its durable identity and normal `read_job` continuation. Relay coverage separately drives more than eight real reconnects and requires `recent_outages` to remain newest-first and capped at eight, while a first-stage Pong timeout cleared by application confirmation without WSS replacement must leave that completed-outage history unchanged. Local producer and Worker sanitizer fixtures bound every episode field, strip unknown/private keys, and require Worker ready promotion to synthesize the just-completed current outage that could not yet have been present in the pre-ready hello. Managed-job inventory tests must prove more than 50 newer terminal helpers cannot hide an older active job. Resource-admission tests must expose the current bounded waiter reason such as `cpu_pressure_window` without waiter identity/token/PID/path fields, and security-audit tests must aggregate recent tool density without argument/result content. macOS power-diagnostic coverage reduces representative `pmset` lines to bounded sleep intervals and requires a same-duration/same-resume-time runtime stall to classify as `matched_system_sleep` while a nonmatching pause remains unclassified. Idle-sleep-guard tests pin activity/grace/release timestamps and the coarse inactivity-expiry reason without extending the existing thirty-minute grace. These are diagnostic/recovery guarantees, not new timing verdicts for the user task.
|
|
26
26
|
|
|
27
27
|
The fast runner avoids a second npm process for package scripts that are exactly one simple `node ...` invocation and have no `pre<task>`/`post<task>` lifecycle hooks. It recreates the relevant npm lifecycle identity in the child environment; scripts containing shell syntax, lifecycle hooks, or non-Node commands still run through npm, so shell/package-manager semantics are not guessed. Parallel-safe fast tasks use a bounded worker pool that defaults to at most four workers and is further bounded by Node's `availableParallelism()`; `MBM_CHECK_CONCURRENCY` may explicitly select 1–16. Known process-heavy gates remain serial barriers. This preserves every fast-plan gate while removing repeated npm startup cost. In non-verbose plans ordinary successful child output remains suppressed, but `self-test` forwards only its sparse phase-start/phase-complete markers so a hosted hang identifies the active phase before an outer CI deadline.
|
|
28
28
|
|
package/docs/TOOL_REFERENCE.md
CHANGED
|
@@ -3091,7 +3091,7 @@ Validate and persist a durable managed-job draft without starting any process. T
|
|
|
3091
3091
|
|
|
3092
3092
|
**Start managed job**
|
|
3093
3093
|
|
|
3094
|
-
Durably accept a detached argv-based job with ordered steps, job-scoped temporary files, guaranteed-attempt finally steps, and optional depends_on links to existing managed jobs. The independent local runner continues if the MCP connection disappears. Dependencies are evaluated from durable managed-job terminal state rather than blind artifact polling: while an upstream job is active the dependent job remains queued in current_phase=dependency_wait without spawning its main child, then proceeds after all dependencies succeed or terminates with error_class=dependency_failed if an upstream later fails. A transient dependency-state read classified as permission_denied, identity_changed, or generic resource_unavailable receives a fixed 45-second recovery grace; one successful read clears that grace, while persistent unavailability fails closed as dependency_unavailable and missing/integrity/witness-invalid evidence remains immediate failure. Staged or already-failed dependencies are rejected before acceptance. On macOS, an authorized remote account job keeps a runner-owned idle-sleep assertion only after the runner ownership claim is confirmed and persisted account ownership is validated, then through dependency wait, admission, steps, cleanup, and terminal persistence; local managed jobs do not acquire this remote-continuity assertion, and it does not override explicit sleep or lid-close behavior. Use an idempotency_key to deduplicate an uncertain retry while the original job record is retained; eviction/retention expiry ends that deduplication window, so this is not permanent exactly-once execution. Use {{temp:name}}, {{resource:name}}, env_resources, or stdin_resource; registered resource contents never enter MCP arguments. Each main/finally step defaults to 600 seconds and may explicitly request up to 21600 seconds (six hours); use this managed-job surface when one continuous command legitimately needs more than the 600-second durable one-step process limit. For one coherent non-interactive workflow, prefer one multi-step managed job over a chain of one-step process carriers so host-visible event density stays bounded. Do not split a long task merely to satisfy an assumed host deadline, and do not use the user as a polling clock.
|
|
3094
|
+
Durably accept a detached argv-based job with ordered steps, job-scoped temporary files, guaranteed-attempt finally steps, and optional depends_on links to existing managed jobs. The independent local runner continues if the MCP connection disappears. Dependencies are evaluated from durable managed-job terminal state rather than blind artifact polling: while an upstream job is active the dependent job remains queued in current_phase=dependency_wait without spawning its main child, then proceeds after all dependencies succeed or terminates with error_class=dependency_failed if an upstream later fails. A transient dependency-state read classified as permission_denied, identity_changed, or generic resource_unavailable receives a fixed 45-second recovery grace; one successful read clears that grace, while persistent unavailability fails closed as dependency_unavailable and missing/integrity/witness-invalid evidence remains immediate failure. Staged or already-failed dependencies are rejected before acceptance. On macOS, an authorized remote account job keeps a runner-owned idle-sleep assertion only after the runner ownership claim is confirmed and persisted account ownership is validated, then through dependency wait, admission, steps, cleanup, and terminal persistence; local managed jobs do not acquire this remote-continuity assertion, and it does not override explicit sleep or lid-close behavior. Use an idempotency_key to deduplicate an uncertain retry while the original job record is retained; eviction/retention expiry ends that deduplication window, so this is not permanent exactly-once execution. Use {{temp:name}}, {{resource:name}}, env_resources, or stdin_resource; registered resource contents never enter MCP arguments. Each main/finally step defaults to 600 seconds and may explicitly request up to 21600 seconds (six hours); use this managed-job surface when one continuous command legitimately needs more than the 600-second durable one-step process limit. Hosted start_job keeps the original response open for the short initial-settlement window advertised by server_info.tool_delivery. If the accepted job reaches terminal state inside that window, the same response includes its terminal result with follow_up_read_required=false; otherwise it preserves the original job_id/recovery envelope with follow_up_read_required=true and normal read_job continuation. This response coalescing does not change durable ownership, dependency waiting, resource admission, or step execution deadlines. For one coherent non-interactive workflow, prefer one multi-step managed job over a chain of one-step process carriers so host-visible event density stays bounded. Do not split a long task merely to satisfy an assumed host deadline, and do not use the user as a polling clock.
|
|
3095
3095
|
|
|
3096
3096
|
| Contract field | Value |
|
|
3097
3097
|
|---|---|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "machine-bridge-mcp",
|
|
3
|
-
"version": "3.0.0-beta.
|
|
3
|
+
"version": "3.0.0-beta.146",
|
|
4
4
|
"description": "Cross-client MCP bridge for local agent context, structured browser and application automation, files, Git, processes, resources, and durable jobs over stdio or OAuth relay.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|
|
@@ -2,7 +2,7 @@ import relayContract from "../shared/relay-contract.json" with { type: "json" };
|
|
|
2
2
|
import { waitForManagedJobRead } from "./managed-job-read-wait.mjs";
|
|
3
3
|
import { ACTIVE_JOB_STATES } from "./managed-job-terminal.mjs";
|
|
4
4
|
|
|
5
|
-
export async function
|
|
5
|
+
export async function settleManagedJobAcceptance(manager, accepted, context = {}) {
|
|
6
6
|
if ((context?.origin !== "relay" && context?.authority?.origin !== "relay") || !accepted?.job_id) return accepted;
|
|
7
7
|
const hostedContext = context?.authority?.origin === "relay"
|
|
8
8
|
? context : { ...context, authority: { ...(context?.authority || {}), origin: "relay" } };
|
|
@@ -17,12 +17,6 @@ export async function settleDurableProcessAcceptance(manager, accepted, context
|
|
|
17
17
|
return {
|
|
18
18
|
...accepted,
|
|
19
19
|
...settled,
|
|
20
|
-
recovery: accepted.recovery,
|
|
21
|
-
cleanup: accepted.cleanup,
|
|
22
|
-
execution_mode: accepted.execution_mode,
|
|
23
|
-
source_tool: accepted.source_tool,
|
|
24
|
-
execution_timeout_seconds: accepted.execution_timeout_seconds,
|
|
25
|
-
retry_safety: accepted.retry_safety,
|
|
26
20
|
progress: {
|
|
27
21
|
status: settled.status,
|
|
28
22
|
current_phase: settled.current_phase ?? null,
|
|
@@ -40,3 +34,7 @@ export async function settleDurableProcessAcceptance(manager, accepted, context
|
|
|
40
34
|
};
|
|
41
35
|
}
|
|
42
36
|
}
|
|
37
|
+
|
|
38
|
+
export async function settleDurableProcessAcceptance(manager, accepted, context = {}) {
|
|
39
|
+
return settleManagedJobAcceptance(manager, accepted, context);
|
|
40
|
+
}
|
|
@@ -11,7 +11,8 @@ import {
|
|
|
11
11
|
relayHttpStatusFromError, sanitizeCloseReason, terminateSocket, tracedTlsConnection,
|
|
12
12
|
} from "./relay-connection-support.mjs";
|
|
13
13
|
import {
|
|
14
|
-
APPLICATION_PROXY_ROUTE_SCOPE, preferredRelayCloseCategory,
|
|
14
|
+
APPLICATION_PROXY_ROUTE_SCOPE, preferredRelayCloseCategory, recordRecoveredOutage, relayOutageFields,
|
|
15
|
+
relayRecoveryFields, relayStatusSnapshot,
|
|
15
16
|
} from "./relay-diagnostics.mjs";
|
|
16
17
|
import {
|
|
17
18
|
acknowledgementMismatch, isSupersededClose, readinessMismatch, reconnectDelay, relayCloseCategory, relayCloseUserCause,
|
|
@@ -89,6 +90,7 @@ export class RelayConnection {
|
|
|
89
90
|
this.outageWarningCount = 0;
|
|
90
91
|
this.lastOutageWarnAt = 0;
|
|
91
92
|
this.outageCount = 0;
|
|
93
|
+
this.recentOutages = [];
|
|
92
94
|
this.lastCloseCategory = "connection_interrupted";
|
|
93
95
|
this.lastCloseCode = 0;
|
|
94
96
|
this.transportError = new RelayTransportErrorState();
|
|
@@ -278,10 +280,11 @@ export class RelayConnection {
|
|
|
278
280
|
this.nextReconnectWallAt = 0;
|
|
279
281
|
this.lastReconnectDelayMs = 0;
|
|
280
282
|
|
|
283
|
+
const recoveredOutage = this.outageStartedAt > 0;
|
|
284
|
+
const outageMs = recoveredOutage ? Math.max(0, this.now() - this.outageStartedAt) : 0;
|
|
281
285
|
if (!this.hasConnected) {
|
|
282
286
|
this.logger.info?.("remote relay connected and end-to-end result delivery verified");
|
|
283
|
-
} else if (
|
|
284
|
-
const outageMs = Math.max(0, this.now() - this.outageStartedAt);
|
|
287
|
+
} else if (recoveredOutage) {
|
|
285
288
|
if (this.outageNoticeEmitted) {
|
|
286
289
|
const recoveryFields = relayRecoveryFields(this, outageMs);
|
|
287
290
|
this.logger.warn?.(`remote relay WebSocket restored after ${formatDuration(outageMs)} (${formatAttempts(this.outageAttempts)})`, recoveryFields);
|
|
@@ -293,6 +296,7 @@ export class RelayConnection {
|
|
|
293
296
|
});
|
|
294
297
|
}
|
|
295
298
|
}
|
|
299
|
+
if (recoveredOutage) recordRecoveredOutage(this, outageMs);
|
|
296
300
|
|
|
297
301
|
const reconnected = this.hasConnected;
|
|
298
302
|
this.hasConnected = true;
|
|
@@ -1,4 +1,5 @@
|
|
|
1
1
|
export const APPLICATION_PROXY_ROUTE_SCOPE = "application-proxy-selection-only";
|
|
2
|
+
export const RECENT_RELAY_OUTAGE_LIMIT = 8;
|
|
2
3
|
|
|
3
4
|
export function preferredRelayCloseCategory(current, next) {
|
|
4
5
|
const existing = String(current || "");
|
|
@@ -27,6 +28,7 @@ export function relayStatusSnapshot(state, now = Date.now()) {
|
|
|
27
28
|
outage_attempts: state.outageAttempts,
|
|
28
29
|
outage_started_at: isoTimestamp(state.outageStartedWallAt),
|
|
29
30
|
outage_duration_ms: state.outageStartedAt > 0 ? Math.max(0, current - state.outageStartedAt) : 0,
|
|
31
|
+
recent_outages: recentOutagesSnapshot(state.recentOutages),
|
|
30
32
|
last_close_category: state.outageCount > 0 ? state.lastCloseCategory : null,
|
|
31
33
|
last_close_code: state.outageCount > 0 ? state.lastCloseCode : null,
|
|
32
34
|
last_transport_error_class: state.outageCount > 0 ? (state.transportError?.errorClass || null) : null,
|
|
@@ -82,6 +84,29 @@ export function relayRecoveryFields(state, outageMs) {
|
|
|
82
84
|
};
|
|
83
85
|
}
|
|
84
86
|
|
|
87
|
+
export function relayRecoveredOutageSnapshot(state, outageMs) {
|
|
88
|
+
return {
|
|
89
|
+
outage_number: state.outageCount,
|
|
90
|
+
disconnected_at: isoTimestamp(state.lastDisconnectedAt),
|
|
91
|
+
ready_at: isoTimestamp(state.lastReadyWallAt),
|
|
92
|
+
duration_ms: Math.max(0, Math.round(Number(outageMs) || 0)),
|
|
93
|
+
attempts: state.outageAttempts,
|
|
94
|
+
close_category: state.lastCloseCategory,
|
|
95
|
+
close_code: state.lastCloseCode,
|
|
96
|
+
network_route: state.networkRoute,
|
|
97
|
+
last_transport_error_class: state.transportError?.errorClass || null,
|
|
98
|
+
...(state.transportError?.snapshot?.() || {}),
|
|
99
|
+
previous_ready_duration_ms: state.lastReadyDurationMs,
|
|
100
|
+
previous_ready_inbound_silence_ms: state.lastReadyInboundSilenceMs,
|
|
101
|
+
...connectTimingFields(state),
|
|
102
|
+
};
|
|
103
|
+
}
|
|
104
|
+
|
|
105
|
+
export function recordRecoveredOutage(state, outageMs) {
|
|
106
|
+
state.recentOutages.unshift(relayRecoveredOutageSnapshot(state, outageMs));
|
|
107
|
+
if (state.recentOutages.length > RECENT_RELAY_OUTAGE_LIMIT) state.recentOutages.length = RECENT_RELAY_OUTAGE_LIMIT;
|
|
108
|
+
}
|
|
109
|
+
|
|
85
110
|
function isoTimestamp(value) {
|
|
86
111
|
return Number(value) > 0 ? new Date(Number(value)).toISOString() : null;
|
|
87
112
|
}
|
|
@@ -97,3 +122,12 @@ function connectTimingFields(state) {
|
|
|
97
122
|
last_connect_milestones_ms: {}, last_connect_http_status: null,
|
|
98
123
|
};
|
|
99
124
|
}
|
|
125
|
+
|
|
126
|
+
function recentOutagesSnapshot(value) {
|
|
127
|
+
if (!Array.isArray(value)) return [];
|
|
128
|
+
return value.slice(0, RECENT_RELAY_OUTAGE_LIMIT).map((entry) => ({
|
|
129
|
+
...entry,
|
|
130
|
+
last_connect_milestones_ms: { ...(entry.last_connect_milestones_ms || {}) },
|
|
131
|
+
last_failed_connect_milestones_ms: { ...(entry.last_failed_connect_milestones_ms || {}) },
|
|
132
|
+
}));
|
|
133
|
+
}
|
|
@@ -12,6 +12,7 @@ const TRANSPORT_ERROR_REASONS = new Set([
|
|
|
12
12
|
"network_unreachable", "host_unreachable", "network_down", "local_address_unavailable", "broken_pipe",
|
|
13
13
|
"dns_not_found", "dns_temporary_failure", "tls_certificate", "tls_protocol", "multi_address_failure",
|
|
14
14
|
]);
|
|
15
|
+
const RECENT_RELAY_OUTAGE_LIMIT = 8;
|
|
15
16
|
|
|
16
17
|
export function relayHandshakeDiagnostics(value = {}) {
|
|
17
18
|
const status = isPlainRecord(value) ? value : {};
|
|
@@ -35,6 +36,7 @@ export function relayHandshakeDiagnostics(value = {}) {
|
|
|
35
36
|
outage_started_at: typeof status.outage_started_at === "string" ? status.outage_started_at : null,
|
|
36
37
|
outage_duration_ms: clampInteger(status.outage_duration_ms, 0, 0, 31 * 24 * 60 * 60_000),
|
|
37
38
|
outage_attempts: clampInteger(status.outage_attempts, 0, 0, 1_000_000),
|
|
39
|
+
recent_outages: recentOutages(status.recent_outages),
|
|
38
40
|
last_close_category: typeof status.last_close_category === "string" ? status.last_close_category : null,
|
|
39
41
|
last_close_code: Number.isSafeInteger(status.last_close_code) ? status.last_close_code : null,
|
|
40
42
|
last_transport_error_class: typeof status.last_transport_error_class === "string"
|
|
@@ -65,6 +67,51 @@ export function relayHandshakeDiagnostics(value = {}) {
|
|
|
65
67
|
};
|
|
66
68
|
}
|
|
67
69
|
|
|
70
|
+
function recentOutages(value) {
|
|
71
|
+
if (!Array.isArray(value)) return [];
|
|
72
|
+
const result = [];
|
|
73
|
+
for (const candidate of value.slice(0, RECENT_RELAY_OUTAGE_LIMIT)) {
|
|
74
|
+
if (!isPlainRecord(candidate)) continue;
|
|
75
|
+
const outageNumber = Number(candidate.outage_number);
|
|
76
|
+
if (!Number.isSafeInteger(outageNumber) || outageNumber < 1 || outageNumber > 1_000_000_000) continue;
|
|
77
|
+
result.push({
|
|
78
|
+
outage_number: outageNumber,
|
|
79
|
+
disconnected_at: boundedTimestamp(candidate.disconnected_at),
|
|
80
|
+
ready_at: boundedTimestamp(candidate.ready_at),
|
|
81
|
+
duration_ms: clampInteger(candidate.duration_ms, 0, 0, 31 * 24 * 60 * 60_000),
|
|
82
|
+
attempts: clampInteger(candidate.attempts, 0, 0, 1_000_000),
|
|
83
|
+
close_category: typeof candidate.close_category === "string" ? candidate.close_category.slice(0, 128) : null,
|
|
84
|
+
close_code: Number.isSafeInteger(candidate.close_code) && candidate.close_code >= 0 && candidate.close_code <= 4999
|
|
85
|
+
? candidate.close_code : null,
|
|
86
|
+
network_route: typeof candidate.network_route === "string" ? candidate.network_route.slice(0, 128) : "unresolved",
|
|
87
|
+
last_transport_error_class: typeof candidate.last_transport_error_class === "string"
|
|
88
|
+
? candidate.last_transport_error_class.slice(0, 128) : null,
|
|
89
|
+
last_transport_error_reason: TRANSPORT_ERROR_REASONS.has(String(candidate.last_transport_error_reason || ""))
|
|
90
|
+
? candidate.last_transport_error_reason : "unknown",
|
|
91
|
+
last_transport_error_ready: candidate.last_transport_error_ready === true,
|
|
92
|
+
last_transport_error_authenticated: candidate.last_transport_error_authenticated === true,
|
|
93
|
+
previous_ready_duration_ms: clampInteger(candidate.previous_ready_duration_ms, 0, 0, 365 * 24 * 60 * 60_000),
|
|
94
|
+
previous_ready_inbound_silence_ms: clampInteger(candidate.previous_ready_inbound_silence_ms, 0, 0, 31 * 24 * 60 * 60_000),
|
|
95
|
+
last_connect_stage: typeof candidate.last_connect_stage === "string" ? candidate.last_connect_stage.slice(0, 64) : "idle",
|
|
96
|
+
last_connect_duration_ms: clampInteger(candidate.last_connect_duration_ms, 0, 0, 10 * 60_000),
|
|
97
|
+
last_connect_milestones_ms: connectMilestones(candidate.last_connect_milestones_ms),
|
|
98
|
+
last_connect_http_status: boundedHttpStatus(candidate.last_connect_http_status),
|
|
99
|
+
last_failed_connect_stage: CONNECT_STAGES.has(String(candidate.last_failed_connect_stage || ""))
|
|
100
|
+
? candidate.last_failed_connect_stage : null,
|
|
101
|
+
last_failed_connect_duration_ms: clampInteger(candidate.last_failed_connect_duration_ms, 0, 0, 10 * 60_000),
|
|
102
|
+
last_failed_connect_milestones_ms: connectMilestones(candidate.last_failed_connect_milestones_ms),
|
|
103
|
+
last_failed_connect_http_status: boundedHttpStatus(candidate.last_failed_connect_http_status),
|
|
104
|
+
});
|
|
105
|
+
}
|
|
106
|
+
return result;
|
|
107
|
+
}
|
|
108
|
+
|
|
109
|
+
function boundedTimestamp(value) {
|
|
110
|
+
if (typeof value !== "string" || value.length > 64) return null;
|
|
111
|
+
const parsed = Date.parse(value);
|
|
112
|
+
return Number.isFinite(parsed) ? new Date(parsed).toISOString() : null;
|
|
113
|
+
}
|
|
114
|
+
|
|
68
115
|
function connectMilestones(value) {
|
|
69
116
|
const source = isPlainRecord(value) ? value : {};
|
|
70
117
|
const out = {};
|
|
@@ -1,4 +1,5 @@
|
|
|
1
1
|
import { clampInteger } from "./numbers.mjs";
|
|
2
|
+
import { settleManagedJobAcceptance } from "./durable-process-initial-settlement.mjs";
|
|
2
3
|
import { waitForManagedJobRead } from "./managed-job-read-wait.mjs";
|
|
3
4
|
|
|
4
5
|
const RUNTIME_TOOL_HANDLERS = Object.freeze({
|
|
@@ -46,7 +47,9 @@ const RUNTIME_TOOL_HANDLERS = Object.freeze({
|
|
|
46
47
|
list_local_resources: (runtime, _args, context) => runtime.managedJobManager.listResources(context),
|
|
47
48
|
generate_ssh_key_resource: (runtime, args, context) => runtime.generateSshKeyResource(args, context),
|
|
48
49
|
stage_job: (runtime, args, context) => runtime.managedJobManager.stage(args, context),
|
|
49
|
-
start_job: (runtime, args, context) =>
|
|
50
|
+
start_job: (runtime, args, context) => settleManagedJobAcceptance(
|
|
51
|
+
runtime.managedJobManager, runtime.managedJobManager.start(args, context), context,
|
|
52
|
+
),
|
|
50
53
|
list_jobs: (runtime, args, context) => runtime.managedJobManager.list(args, context),
|
|
51
54
|
read_job: (runtime, args, context) => waitForManagedJobRead({
|
|
52
55
|
args, context,
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
"supportedProtocolVersions": [
|
|
5
5
|
"2026-07-28"
|
|
6
6
|
],
|
|
7
|
-
"toolSchemaGeneration":
|
|
7
|
+
"toolSchemaGeneration": 15,
|
|
8
8
|
"instructions": [
|
|
9
9
|
"You are connected to a local workspace through machine-bridge-mcp.",
|
|
10
10
|
"Use resolve_task_capabilities only when the task needs local skill/registered-command discovery, application/browser routing, or refreshed project instructions. Straightforward file, Git, and shell work should use the already exposed tools directly instead of adding a resolver call. When capability routing is useful, pass the current user request and target path and reuse refresh.fingerprint as known_refresh_fingerprint to omit unchanged static instructions. Its execution_routing is bounded, set-level, and advisory: direct Bash remains the general escape hatch whenever the effective policy allows it. MCP transport connections are not conversation identity; refresh task-specific ranking when routing is actually needed rather than assuming prior requests established context.",
|
|
@@ -14,14 +14,14 @@
|
|
|
14
14
|
"For local application work, use structured application discovery and Accessibility actions. Arbitrary caller-supplied AppleScript or JavaScript is intentionally not accepted.",
|
|
15
15
|
"Remote mode uses a Cloudflare relay; stdio mode runs on the local machine. File and command operations execute on the user's local runtime, not in the Worker.",
|
|
16
16
|
"In remote multi-account mode, account authority is the intersection of the authenticated account role and the daemon capability ceiling. Only authorization.effective_policy and authorization.effective_tools describe that account authority; daemon.policy is not the account permission.",
|
|
17
|
-
"Use server_info with detail=summary for routine version, authority-profile, readiness, relay, pending-capacity, and timeout checks. Request the default/full projection only when exact effective tools, OAuth metadata, account identity, full isolate-local observability counters, or owner-visible durable continuity evidence are needed. worker.continuity_evidence survives Worker isolate replacement and contains only fixed-category counts/timestamps plus bounded socket-close classification; it never contains account/client identity, call IDs, tool names, arguments, results, endpoints, or close reasons. For a recovered relay interruption, compare close-to-ready outage duration with previous_ready_inbound_silence_ms: the latter captures the bounded pre-close interval in which a previously ready transport stopped proving inbound activity.",
|
|
17
|
+
"Use server_info with detail=summary for routine version, authority-profile, readiness, relay, pending-capacity, and timeout checks. Request the default/full projection only when exact effective tools, OAuth metadata, account identity, full isolate-local observability counters, or owner-visible durable continuity evidence are needed. worker.continuity_evidence survives Worker isolate replacement and contains only fixed-category counts/timestamps plus bounded socket-close classification; it never contains account/client identity, call IDs, tool names, arguments, results, endpoints, or close reasons. For a recovered relay interruption, compare close-to-ready outage duration with previous_ready_inbound_silence_ms: the latter captures the bounded pre-close interval in which a previously ready transport stopped proving inbound activity. Full relay diagnostics also retain recent_outages, a newest-first bounded history of up to eight completed WebSocket reconnect episodes; transport suspicions that recover through Pong/application confirmation without rebuilding the socket remain heartbeat evidence and are not inserted into that history.",
|
|
18
18
|
"Filename sensitivity is not classified by this server; the MCP host, connector gateway, local OS, or endpoint-security software may enforce additional independent rules.",
|
|
19
19
|
"A named full profile is canonical: it always enables writes, shell execution, unrestricted paths, the full parent environment, absolute paths, and the complete tool catalog. Any explicit narrowing is stored as custom.",
|
|
20
20
|
"The local daemon may advertise the full catalog as its capability ceiling. In remote mode, the Worker relays only the intersection allowed by the authenticated account role; a connector host may then expose an even smaller subset that Machine Bridge cannot observe or override. Conversation/surface app routing and host-cached action/tool snapshots are host-owned state, not daemon authority.",
|
|
21
21
|
"Use diagnose_runtime to distinguish a request that reached the daemon from local filesystem, process-spawn, shell, managed-job-storage, resource-admission, relay, event-density, and operating-system-suspension evidence. Owner-visible diagnostics include bounded managed-job churn, content-free recent security-audit tool-call density, the oldest resource waiters with their current pre-spawn admission reasons, and on macOS a bounded fixed sleep-history projection plus event-loop-pause and most-recent-relay-outage correlation. A matched event-loop sleep requires both pause duration and resume time to agree with the operating-system interval; relay_outage_analysis separately reports exact overlap duration/ratio for the latest recovered disconnect. majority_system_sleep_overlap means most of that outage occurred while the Mac was suspended, so a retained connection_reset is aftermath evidence rather than proof of an independent network root cause. An unmatched pause or awake relay reset is not automatically assigned another cause. The idle-sleep guard also exposes coarse activity/grace/release timestamps and whether its fixed caffeinate child requests AC-only system-sleep prevention in addition to idle-sleep prevention. Use these fields to distinguish host-visible event amplification from a job waiting for CPU/IO/memory admission or a machine that actually slept. A tool call blocked before any response cannot be diagnosed by the server; do not call that a platform disable without host-side evidence. Conversely, an execution failure that includes a child exit code or bounded child stdout/stderr proves the local process was spawned; for SSH or another nested command, treat an allowlist/forced-command refusal in that output as downstream target evidence rather than Machine Bridge policy denial. If a fresh supported conversation can invoke the same app/tool while an older conversation cannot, investigate the host conversation/surface route or cached action/tool snapshot before changing Machine Bridge policy.",
|
|
22
22
|
"Never request or return secret-file contents when a local resource alias can be used. Resources are registered through the local machine-mcp CLI; under canonical full policy, generate_ssh_key_resource can generate and register an Ed25519 key without returning private content and may be injected by path, stdin, or environment without entering MCP arguments.",
|
|
23
23
|
"stage_job persists a validated non-running draft only. It is not an approval workflow and cannot be promoted from the terminal; execution uses start_job when the effective account policy permits it, or an explicit local machine-mcp job submit PLAN.json operation.",
|
|
24
|
-
"Hosted start_job requires a caller-chosen idempotency_key before dispatch. If its acceptance response is ambiguous, retry exactly the same start_job arguments with the same key rather than creating a second logical submission or switching to another execution path. Recovery errors name the credential source as the original request and do not echo the key value. Local/stdio start_job keeps the underlying optional-key API.",
|
|
24
|
+
"Hosted start_job requires a caller-chosen idempotency_key before dispatch. If its acceptance response is ambiguous, retry exactly the same start_job arguments with the same key rather than creating a second logical submission or switching to another execution path. Recovery errors name the credential source as the original request and do not echo the key value. After durable acceptance, the original hosted start_job response waits through the short advertised managed-job initial-settlement window; a job that becomes terminal inside that window returns its terminal result with follow_up_read_required=false, while an active job preserves the same job_id/recovery envelope with follow_up_read_required=true for normal read_job continuation. This coalescing changes host-visible response count only: durable ownership, dependencies, resource admission, six-hour step ceilings, and recovery semantics are unchanged. Local/stdio start_job keeps the underlying optional-key API without the hosted settlement wait.",
|
|
25
25
|
"For durable cross-job ordering, prefer start_job.depends_on over shell loops that wait for files or markers. A dependent runner remains queued in current_phase=dependency_wait without spawning its main child until all referenced managed jobs succeed; dependency_pending_count is progress observed by hosted read_job and is coalesced by the same nonterminal progress pacing as other active-state changes. If an upstream job later fails, the dependent settles with result.error_class=dependency_failed rather than blind-waiting for an artifact that can never appear. During dependency_wait, a secure state read classified as permission_denied, identity_changed, or generic resource_unavailable receives a fixed 45-second recovery grace; one successful read clears that grace, while persistent unavailability fails closed as dependency_unavailable and missing/integrity/witness-invalid evidence remains immediate failure. Staged or already-failed dependencies are rejected before acceptance. Active dependency plans pin the referenced retained job records against capacity pruning. This is a continuity mechanism, not a reason to shorten the overall user task or force a new user turn.",
|
|
26
26
|
"Hosted read-only status and diagnostic tools must not be used as busy loops. For a known managed job, use read_job's server-side paced long-poll by default: a relay-origin active read waits inside Machine Bridge for up to the advertised managed-job read interval. Terminal settlement returns on the next bounded progress poll, but nonterminal changes are coalesced for at least the advertised thirty-second progress interval (or the caller's explicitly shorter wait), and current_step-only churn does not wake the host call by itself. This keeps multi-step long jobs observable without turning every short local step into a host-visible event. Use wait_ms=0 only for an intentional immediate checkpoint. Interactive process sessions use the separately paced read_process contract: the one-second actual output/exit blocking cap remains, while a repeated would-block request inside the fifteen-second cooldown stays inside that same MCP call until output/exit or the cooldown boundary rather than returning a rapid running checkpoint. Bounded same-response follow-up is allowed when the current task needs terminal state or additional output. Do not switch among list_jobs, server_info, diagnose_runtime, or other status surfaces merely to evade pacing. Per-call wait survival and aggregate host-response lifetime are separate constraints: do not infer that a long task can remain in one assistant response merely because each individual read succeeds. Do not infer or preempt a host/tool deadline from elapsed wall-clock time. While tool calls continue to be accepted and the task still needs the result, bounded same-response follow-up may continue; if an actual host/tool boundary ends the response, preserve the durable job/session identifier and resume that same operation later instead of resubmitting its side effect. A planned local-daemon restart may settle an in-flight hosted read_job as retryable unavailable with recovery.mode=read_same_job and the original job_id; after reconnect, read that same job_id instead of resubmitting the job's business side effect. A typed read_job not_found means the retained job record is unavailable, not that the underlying operation definitely never ran; do not blindly replay it.",
|
|
27
27
|
"For ordinary one-step remote process work whose child execution fits the 600-second process limit, exec_command, run_process, and run_local_command first commit a durable managed job, then keep the original hosted tool call open for the short advertised initial-settlement window. Acceptance transfers execution to durable ownership without forcing the current assistant response to end. If the helper settles inside that window, its terminal status/result is returned in the same tool response and no separate read_job event is required; if it remains active, the response keeps the same job_id/recovery envelope and normal read_job continuation applies. This reduces the common helper-plus-read double event without weakening durable ownership or changing the child timeout. For one coherent non-interactive workflow that needs several local commands, prefer one repository-native umbrella command or one multi-step start_job (or the smallest practical number of start_job plans) instead of emitting a chain of one-step process calls. Host-visible tool-event density is a continuity resource: batching reduces host traffic without shortening the underlying task or step lifetime. When the current task needs the result, bounded same-response read_job follow-up is allowed only when follow_up_read_required remains true; do not busy-loop or use repeated list_jobs as a substitute polling surface, and do not infer a host/tool deadline from elapsed wall-clock time. Cooperative machine-user admission is a separate pre-spawn interval and may wait up to 30 minutes; read_job.current_phase=resource_admission means no child for that step has started yet. These one-step process carriers use lower-priority terminal retention than explicit managed jobs so helper-command churn reclaims completed helper history before displacing explicit durable recovery results when such helper history is available; the managed-job store remains bounded to 512 retained states. list_jobs keeps its primary jobs window capped at 50 and prioritizes unreadable, active, staged, and durable terminal recovery state ahead of transient one-step helper history; when a recent transient terminal is retained inside the 30-minute/16-result recovery reserve but omitted from that primary window, recent_process_recovery returns up to 16 additional authority-visible public job handles without step output or internal retention metadata. This is recovery discovery after a real host/tool boundary, not a polling or MCP replay/session surface; use read_job on the recovered job_id. Owner/local capacity diagnostics distinguish durable_terminal from transient_terminal without exposing job identities. Use start_job for policy-authorized multi-step plans, job-scoped temporary_files, explicit resource injection, idempotent finally_steps, or a step that legitimately needs an execution budget above 600 seconds. Detached jobs continue after MCP disconnects and daemon replacement; return job_id/status/current_phase for later recovery only after an actual host/tool boundary is observed, external input or authorization is required, or the user explicitly requested a checkpoint.",
|
|
@@ -2681,7 +2681,7 @@
|
|
|
2681
2681
|
{
|
|
2682
2682
|
"name": "start_job",
|
|
2683
2683
|
"title": "Start managed job",
|
|
2684
|
-
"description": "Durably accept a detached argv-based job with ordered steps, job-scoped temporary files, guaranteed-attempt finally steps, and optional depends_on links to existing managed jobs. The independent local runner continues if the MCP connection disappears. Dependencies are evaluated from durable managed-job terminal state rather than blind artifact polling: while an upstream job is active the dependent job remains queued in current_phase=dependency_wait without spawning its main child, then proceeds after all dependencies succeed or terminates with error_class=dependency_failed if an upstream later fails. A transient dependency-state read classified as permission_denied, identity_changed, or generic resource_unavailable receives a fixed 45-second recovery grace; one successful read clears that grace, while persistent unavailability fails closed as dependency_unavailable and missing/integrity/witness-invalid evidence remains immediate failure. Staged or already-failed dependencies are rejected before acceptance. On macOS, an authorized remote account job keeps a runner-owned idle-sleep assertion only after the runner ownership claim is confirmed and persisted account ownership is validated, then through dependency wait, admission, steps, cleanup, and terminal persistence; local managed jobs do not acquire this remote-continuity assertion, and it does not override explicit sleep or lid-close behavior. Use an idempotency_key to deduplicate an uncertain retry while the original job record is retained; eviction/retention expiry ends that deduplication window, so this is not permanent exactly-once execution. Use {{temp:name}}, {{resource:name}}, env_resources, or stdin_resource; registered resource contents never enter MCP arguments. Each main/finally step defaults to 600 seconds and may explicitly request up to 21600 seconds (six hours); use this managed-job surface when one continuous command legitimately needs more than the 600-second durable one-step process limit. For one coherent non-interactive workflow, prefer one multi-step managed job over a chain of one-step process carriers so host-visible event density stays bounded. Do not split a long task merely to satisfy an assumed host deadline, and do not use the user as a polling clock.",
|
|
2684
|
+
"description": "Durably accept a detached argv-based job with ordered steps, job-scoped temporary files, guaranteed-attempt finally steps, and optional depends_on links to existing managed jobs. The independent local runner continues if the MCP connection disappears. Dependencies are evaluated from durable managed-job terminal state rather than blind artifact polling: while an upstream job is active the dependent job remains queued in current_phase=dependency_wait without spawning its main child, then proceeds after all dependencies succeed or terminates with error_class=dependency_failed if an upstream later fails. A transient dependency-state read classified as permission_denied, identity_changed, or generic resource_unavailable receives a fixed 45-second recovery grace; one successful read clears that grace, while persistent unavailability fails closed as dependency_unavailable and missing/integrity/witness-invalid evidence remains immediate failure. Staged or already-failed dependencies are rejected before acceptance. On macOS, an authorized remote account job keeps a runner-owned idle-sleep assertion only after the runner ownership claim is confirmed and persisted account ownership is validated, then through dependency wait, admission, steps, cleanup, and terminal persistence; local managed jobs do not acquire this remote-continuity assertion, and it does not override explicit sleep or lid-close behavior. Use an idempotency_key to deduplicate an uncertain retry while the original job record is retained; eviction/retention expiry ends that deduplication window, so this is not permanent exactly-once execution. Use {{temp:name}}, {{resource:name}}, env_resources, or stdin_resource; registered resource contents never enter MCP arguments. Each main/finally step defaults to 600 seconds and may explicitly request up to 21600 seconds (six hours); use this managed-job surface when one continuous command legitimately needs more than the 600-second durable one-step process limit. Hosted start_job keeps the original response open for the short initial-settlement window advertised by server_info.tool_delivery. If the accepted job reaches terminal state inside that window, the same response includes its terminal result with follow_up_read_required=false; otherwise it preserves the original job_id/recovery envelope with follow_up_read_required=true and normal read_job continuation. This response coalescing does not change durable ownership, dependency waiting, resource admission, or step execution deadlines. For one coherent non-interactive workflow, prefer one multi-step managed job over a chain of one-step process carriers so host-visible event density stays bounded. Do not split a long task merely to satisfy an assumed host deadline, and do not use the user as a polling clock.",
|
|
2685
2685
|
"availability": "write+direct-exec",
|
|
2686
2686
|
"annotations": {
|
|
2687
2687
|
"readOnlyHint": false,
|
|
@@ -46,6 +46,32 @@ const CLOSE_CATEGORIES = new Set([
|
|
|
46
46
|
"unexpected_close",
|
|
47
47
|
"superseded",
|
|
48
48
|
]);
|
|
49
|
+
const RECENT_RELAY_OUTAGE_LIMIT = 8;
|
|
50
|
+
|
|
51
|
+
export interface DaemonRelayOutageDiagnostics {
|
|
52
|
+
outage_number: number;
|
|
53
|
+
disconnected_at: string | null;
|
|
54
|
+
ready_at: string | null;
|
|
55
|
+
duration_ms: number;
|
|
56
|
+
attempts: number;
|
|
57
|
+
close_category: string | null;
|
|
58
|
+
close_code: number | null;
|
|
59
|
+
network_route: string;
|
|
60
|
+
last_transport_error_class: string | null;
|
|
61
|
+
last_transport_error_reason: string;
|
|
62
|
+
last_transport_error_ready: boolean;
|
|
63
|
+
last_transport_error_authenticated: boolean;
|
|
64
|
+
previous_ready_duration_ms: number;
|
|
65
|
+
previous_ready_inbound_silence_ms: number;
|
|
66
|
+
last_connect_stage: string;
|
|
67
|
+
last_connect_duration_ms: number;
|
|
68
|
+
last_connect_milestones_ms: Record<string, number>;
|
|
69
|
+
last_connect_http_status: number | null;
|
|
70
|
+
last_failed_connect_stage: string | null;
|
|
71
|
+
last_failed_connect_duration_ms: number;
|
|
72
|
+
last_failed_connect_milestones_ms: Record<string, number>;
|
|
73
|
+
last_failed_connect_http_status: number | null;
|
|
74
|
+
}
|
|
49
75
|
|
|
50
76
|
export interface DaemonRelayDiagnostics {
|
|
51
77
|
schema_version: 1;
|
|
@@ -65,6 +91,7 @@ export interface DaemonRelayDiagnostics {
|
|
|
65
91
|
outage_started_at: string | null;
|
|
66
92
|
outage_duration_ms: number;
|
|
67
93
|
outage_attempts: number;
|
|
94
|
+
recent_outages: DaemonRelayOutageDiagnostics[];
|
|
68
95
|
last_close_category: string | null;
|
|
69
96
|
last_close_code: number | null;
|
|
70
97
|
last_transport_error_class: string | null;
|
|
@@ -113,6 +140,7 @@ export function sanitizeDaemonRelayDiagnostics(value: unknown): DaemonRelayDiagn
|
|
|
113
140
|
outage_started_at: timestamp(candidate.outage_started_at),
|
|
114
141
|
outage_duration_ms: boundedInteger(candidate.outage_duration_ms, 0, 31 * 24 * 60 * 60_000, 0),
|
|
115
142
|
outage_attempts: boundedInteger(candidate.outage_attempts, 0, 1_000_000, 0),
|
|
143
|
+
recent_outages: recentOutages(candidate.recent_outages),
|
|
116
144
|
last_close_category: nullableEnum(candidate.last_close_category, CLOSE_CATEGORIES),
|
|
117
145
|
last_close_code: nullableInteger(candidate.last_close_code, 0, 4999),
|
|
118
146
|
last_transport_error_class: nullableEnum(candidate.last_transport_error_class, TRANSPORT_ERROR_CLASSES),
|
|
@@ -158,10 +186,49 @@ export function relayDiagnosticsAfterReady(
|
|
|
158
186
|
const started = Date.parse(value.outage_started_at ?? "");
|
|
159
187
|
const ready = Date.parse(readyAt ?? "");
|
|
160
188
|
const elapsed = Number.isFinite(started) && Number.isFinite(ready) ? Math.max(0, ready - started) : 0;
|
|
189
|
+
const readyTimestamp = timestamp(readyAt);
|
|
190
|
+
const recovered = value.outage_active && value.outage_count > 0 && readyTimestamp
|
|
191
|
+
? recoveredOutage(value, readyTimestamp, Math.max(value.outage_duration_ms, elapsed)) : null;
|
|
192
|
+
const recentOutages = recovered
|
|
193
|
+
? [recovered, ...value.recent_outages.filter((entry) => entry.outage_number !== recovered.outage_number)]
|
|
194
|
+
.slice(0, RECENT_RELAY_OUTAGE_LIMIT)
|
|
195
|
+
: value.recent_outages;
|
|
161
196
|
return {
|
|
162
197
|
...value,
|
|
163
198
|
outage_active: false,
|
|
164
199
|
outage_duration_ms: Math.min(31 * 24 * 60 * 60_000, Math.max(value.outage_duration_ms, elapsed)),
|
|
200
|
+
recent_outages: recentOutages,
|
|
201
|
+
};
|
|
202
|
+
}
|
|
203
|
+
|
|
204
|
+
function recoveredOutage(
|
|
205
|
+
value: DaemonRelayDiagnostics,
|
|
206
|
+
readyAt: string,
|
|
207
|
+
durationMs: number,
|
|
208
|
+
): DaemonRelayOutageDiagnostics {
|
|
209
|
+
return {
|
|
210
|
+
outage_number: value.outage_count,
|
|
211
|
+
disconnected_at: value.last_disconnected_at ?? value.outage_started_at,
|
|
212
|
+
ready_at: readyAt,
|
|
213
|
+
duration_ms: boundedInteger(durationMs, 0, 31 * 24 * 60 * 60_000, 0),
|
|
214
|
+
attempts: value.outage_attempts,
|
|
215
|
+
close_category: value.last_close_category,
|
|
216
|
+
close_code: value.last_close_code,
|
|
217
|
+
network_route: value.network_route,
|
|
218
|
+
last_transport_error_class: value.last_transport_error_class,
|
|
219
|
+
last_transport_error_reason: value.last_transport_error_reason,
|
|
220
|
+
last_transport_error_ready: value.last_transport_error_ready,
|
|
221
|
+
last_transport_error_authenticated: value.last_transport_error_authenticated,
|
|
222
|
+
previous_ready_duration_ms: value.previous_ready_duration_ms,
|
|
223
|
+
previous_ready_inbound_silence_ms: value.previous_ready_inbound_silence_ms,
|
|
224
|
+
last_connect_stage: value.last_connect_stage,
|
|
225
|
+
last_connect_duration_ms: value.last_connect_duration_ms,
|
|
226
|
+
last_connect_milestones_ms: { ...value.last_connect_milestones_ms },
|
|
227
|
+
last_connect_http_status: value.last_connect_http_status,
|
|
228
|
+
last_failed_connect_stage: value.last_failed_connect_stage,
|
|
229
|
+
last_failed_connect_duration_ms: value.last_failed_connect_duration_ms,
|
|
230
|
+
last_failed_connect_milestones_ms: { ...value.last_failed_connect_milestones_ms },
|
|
231
|
+
last_failed_connect_http_status: value.last_failed_connect_http_status,
|
|
165
232
|
};
|
|
166
233
|
}
|
|
167
234
|
|
|
@@ -191,3 +258,39 @@ function connectMilestones(value: unknown): Record<string, number> {
|
|
|
191
258
|
}
|
|
192
259
|
return result;
|
|
193
260
|
}
|
|
261
|
+
|
|
262
|
+
function recentOutages(value: unknown): DaemonRelayOutageDiagnostics[] {
|
|
263
|
+
if (!Array.isArray(value)) return [];
|
|
264
|
+
const result: DaemonRelayOutageDiagnostics[] = [];
|
|
265
|
+
for (const candidate of value.slice(0, RECENT_RELAY_OUTAGE_LIMIT)) {
|
|
266
|
+
if (!candidate || typeof candidate !== "object" || Array.isArray(candidate)) continue;
|
|
267
|
+
const entry = candidate as Record<string, unknown>;
|
|
268
|
+
const outageNumber = Number(entry.outage_number);
|
|
269
|
+
if (!Number.isSafeInteger(outageNumber) || outageNumber < 1 || outageNumber > 1_000_000_000) continue;
|
|
270
|
+
result.push({
|
|
271
|
+
outage_number: outageNumber,
|
|
272
|
+
disconnected_at: timestamp(entry.disconnected_at),
|
|
273
|
+
ready_at: timestamp(entry.ready_at),
|
|
274
|
+
duration_ms: boundedInteger(entry.duration_ms, 0, 31 * 24 * 60 * 60_000, 0),
|
|
275
|
+
attempts: boundedInteger(entry.attempts, 0, 1_000_000, 0),
|
|
276
|
+
close_category: nullableEnum(entry.close_category, CLOSE_CATEGORIES),
|
|
277
|
+
close_code: nullableInteger(entry.close_code, 0, 4999),
|
|
278
|
+
network_route: enumText(entry.network_route, NETWORK_ROUTES, "unresolved"),
|
|
279
|
+
last_transport_error_class: nullableEnum(entry.last_transport_error_class, TRANSPORT_ERROR_CLASSES),
|
|
280
|
+
last_transport_error_reason: enumText(entry.last_transport_error_reason, TRANSPORT_ERROR_REASONS, "unknown"),
|
|
281
|
+
last_transport_error_ready: entry.last_transport_error_ready === true,
|
|
282
|
+
last_transport_error_authenticated: entry.last_transport_error_authenticated === true,
|
|
283
|
+
previous_ready_duration_ms: boundedInteger(entry.previous_ready_duration_ms, 0, 365 * 24 * 60 * 60_000, 0),
|
|
284
|
+
previous_ready_inbound_silence_ms: boundedInteger(entry.previous_ready_inbound_silence_ms, 0, 31 * 24 * 60 * 60_000, 0),
|
|
285
|
+
last_connect_stage: enumText(entry.last_connect_stage, CONNECT_STAGES, "idle"),
|
|
286
|
+
last_connect_duration_ms: boundedInteger(entry.last_connect_duration_ms, 0, 10 * 60_000, 0),
|
|
287
|
+
last_connect_milestones_ms: connectMilestones(entry.last_connect_milestones_ms),
|
|
288
|
+
last_connect_http_status: nullableInteger(entry.last_connect_http_status, 100, 599),
|
|
289
|
+
last_failed_connect_stage: nullableEnum(entry.last_failed_connect_stage, CONNECT_STAGES),
|
|
290
|
+
last_failed_connect_duration_ms: boundedInteger(entry.last_failed_connect_duration_ms, 0, 10 * 60_000, 0),
|
|
291
|
+
last_failed_connect_milestones_ms: connectMilestones(entry.last_failed_connect_milestones_ms),
|
|
292
|
+
last_failed_connect_http_status: nullableInteger(entry.last_failed_connect_http_status, 100, 599),
|
|
293
|
+
});
|
|
294
|
+
}
|
|
295
|
+
return result;
|
|
296
|
+
}
|
package/src/worker/index.ts
CHANGED
|
@@ -60,7 +60,7 @@ import {
|
|
|
60
60
|
closeWebSocketQuietly, daemonErrorCloseCode, isObjectRecord, rejectDaemonMessage,
|
|
61
61
|
sendWebSocketQuietly, trySendWebSocket,
|
|
62
62
|
} from "./websocket-protocol.ts";
|
|
63
|
-
const SERVER_VERSION = "3.0.0-beta.
|
|
63
|
+
const SERVER_VERSION = "3.0.0-beta.146";
|
|
64
64
|
const MCP_SERVER_INFO = mcpServerInfo(SERVER_VERSION);
|
|
65
65
|
const MAX_DAEMON_MESSAGE_BYTES = 8 * 1024 * 1024;
|
|
66
66
|
const DAEMON_RECONNECT_GRACE_MS = relayContract.reconnectGraceMs; const NEW_CALL_RECONNECT_GRACE_MS = relayContract.newCallReconnectGraceMs;
|
|
@@ -27,6 +27,7 @@ export function remoteToolDeliveryContract(
|
|
|
27
27
|
remote_process_delivery_mode: "durable_job",
|
|
28
28
|
remote_process_acceptance_max_ms: relayContract.durableProcessAcceptanceTimeoutMs,
|
|
29
29
|
remote_process_initial_settlement_wait_ms: relayContract.durableProcessInitialSettlementWaitMs,
|
|
30
|
+
remote_managed_job_initial_settlement_wait_ms: relayContract.durableProcessInitialSettlementWaitMs,
|
|
30
31
|
remote_process_execution_timeout_max_ms: relayContract.maximumDurableProcessExecutionTimeoutMs,
|
|
31
32
|
managed_job_resource_admission_wait_max_ms: relayContract.maximumManagedJobResourceAdmissionWaitMs,
|
|
32
33
|
remote_managed_job_read_wait_default_ms: relayContract.defaultManagedJobReadWaitMs,
|
|
@@ -47,6 +48,7 @@ export function compactRemoteToolDeliveryContract(
|
|
|
47
48
|
const compact = { ...remoteToolDeliveryContract(serverVersion, subscription) };
|
|
48
49
|
delete compact.remote_managed_job_read_nonterminal_progress_minimum_ms;
|
|
49
50
|
delete compact.remote_process_initial_settlement_wait_ms;
|
|
51
|
+
delete compact.remote_managed_job_initial_settlement_wait_ms;
|
|
50
52
|
delete compact.remote_process_blocking_poll_wait_max_ms;
|
|
51
53
|
return compact;
|
|
52
54
|
}
|