machine-bridge-mcp 3.0.0-beta.102 → 3.0.0-beta.103
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +6 -0
- package/README.md +1 -1
- package/browser-extension/manifest.json +1 -1
- package/docs/ARCHITECTURE.md +1 -1
- package/docs/AUDIT.md +10 -0
- package/docs/LOGGING.md +1 -1
- package/docs/OPERATIONS.md +1 -1
- package/docs/TESTING.md +1 -1
- package/package.json +1 -1
- package/src/local/remote-activity-idle-sleep-guard.mjs +1 -1
- package/src/worker/index.ts +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,11 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 3.0.0-beta.103 - 2026-08-18
|
|
4
|
+
|
|
5
|
+
- Supersede beta.102 after owner use still reproduced a severe whole-control-plane interruption. Live beta.102 evidence separates this event from the signed HTTPS fallback: the final `PreventUserIdleSystemSleep` assertion ended, macOS entered ordinary `Idle Sleep` about five seconds later for roughly 632 seconds, and the daemon reported an aligned event-loop pause of roughly 626 seconds while launchd retained one continuously running service process. A suspended host cannot execute either WSS or HTTPS recovery. The confirmed product defect is therefore the daemon-side post-activity power lease: beta.102's five-minute grace was shorter than a normal multi-turn ChatGPT reasoning/wait interval and could release immediately before the operating system chose Idle Sleep.
|
|
6
|
+
- Extend the authorized remote-activity guard to a fixed **thirty-minute rolling inactivity lease**. Policy/account/operation authorization and argument validation still precede acquisition; concurrent relay handlers still share one daemon-bound `/usr/bin/caffeinate -i -w <daemon-pid>` child; long remote process sessions and account-backed managed-job runners retain their existing execution-lifetime ownership. After the final daemon-side activity settles, release is scheduled thirty minutes later. Any new authorized remote activity cancels the pending release and starts a fresh full thirty-minute window after the new activity settles. Relay/application heartbeats do not renew the lease, so an otherwise idle login daemon does not keep the machine awake indefinitely. Explicit sleep and lid-close behavior remain outside the guarantee.
|
|
7
|
+
- Add a red/green runtime regression for the thirty-minute default and actual release-timer wiring, update architecture contracts to reject a return to the five-minute policy, and synchronize README/architecture/operations/testing/logging guidance. These packaged changes invalidate beta.102 candidate/acceptance evidence for beta.103 release purposes. Beta.102 source/npm publication remains stopped; beta.103 requires fresh frozen verification, exact candidate/install-only preparation, new explicit owner authorization before live activation, activated-package OAuth canary, observed live verification, and new acceptance before review/publication can continue.
|
|
8
|
+
|
|
3
9
|
## 3.0.0-beta.102 - 2026-08-18
|
|
4
10
|
|
|
5
11
|
- Supersede beta.101 after a fresh live interruption disproved the assumption that faster WebSocket half-open detection was sufficient recovery. The activated beta.101 launchd daemon remained the same PID with `runs=1` and no exit or matching Sleep/Wake event, yet the daemon-to-Worker relay disappeared at `2026-08-18T04:00:07.944Z` and required about 323.8 seconds plus 13 reconnect attempts before recovery; the final retained classification was `relay_connect_timeout`, with 9,830 ms of inbound silence before the ready socket closed. During the incident Worker-local HTTP/MCP remained responsive while no authenticated/ready/candidate daemon socket existed. The host's default route was carried by `utun5`, and same-window system logs showed multi-second TLS/read stalls in other applications even while Network.framework considered the route satisfied, but the available evidence cannot identify a particular VPN/TUN/provider, Wi-Fi component, Cloudflare edge, ISP, or upstream device as the physical trigger.
|
package/README.md
CHANGED
|
@@ -195,7 +195,7 @@ For stateful GUI trajectories, owner/full callers can use the higher-level `comp
|
|
|
195
195
|
|
|
196
196
|
Remote request-owned foreground work uses the hosted reply-safe budgets described above; configurable browser/application calls may explicitly request at most 45 seconds, while remote `exec_command`, `run_process`, and `run_local_command` are durable one-step jobs with a 10-second acceptance envelope and an independent 1–600-second child execution budget after admission. The Worker retains separate settlement ownership for five additional seconds, but neither that margin nor its internal stream metrics prove that an external MCP host consumed the terminal frame. Keep mutations and validation in independently terminal calls. A timeout is a protocol result, not proof that descendant cleanup has already completed; a remote owner can inspect `diagnose_runtime.runtime.processes`, while local stdio exposes `server_info.runtime.processes`. Non-owner accounts receive authority-scoped readiness rather than machine-wide process activity. Long, cleanup-sensitive, or remotely initiated workflows should use process sessions or managed jobs; managed jobs persist ordered argv steps and `finally_steps` under owner-only local state and continue across an MCP disconnect.
|
|
197
197
|
|
|
198
|
-
On macOS, authorized remote activity uses a bounded idle-sleep assertion so ordinary system Idle Sleep does not suspend an active remote workflow. Relay handlers share the assertion for their execution lifetime plus a fixed
|
|
198
|
+
On macOS, authorized remote activity uses a bounded idle-sleep assertion so ordinary system Idle Sleep does not suspend an active remote workflow. Relay handlers share the assertion for their execution lifetime plus a fixed thirty-minute rolling inactivity grace; each new authorized remote activity cancels a pending release and restarts the full grace after the last concurrent handler settles. An admitted remote process session extends daemon-side ownership until its child settles, and an account-backed managed-job runner owns a runner-bound assertion from confirmed claim through terminal persistence. Local managed jobs do not acquire the remote-continuity assertion. These protections do not override explicit sleep or lid-close behavior.
|
|
199
199
|
|
|
200
200
|
When `run_process` or `exec_command` returns a child exit code and bounded stdout/stderr, the local process did run. For nested tools such as `ssh`, a remote forced-command usage message or command allowlist is therefore evidence from the target-side authorization layer, not evidence that Machine Bridge blocked process execution. Diagnose and change the narrowest failing layer instead of widening the `full` profile, which already removes Machine Bridge's own shell and path restrictions.
|
|
201
201
|
|
|
@@ -30,6 +30,6 @@
|
|
|
30
30
|
"action": {
|
|
31
31
|
"default_title": "Machine Bridge Browser"
|
|
32
32
|
},
|
|
33
|
-
"version_name": "3.0.0-beta.
|
|
33
|
+
"version_name": "3.0.0-beta.103",
|
|
34
34
|
"key": "MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAxryYkpZhq8+VAQLHcGS9BAHQcyKX8RHGIpIwvtIVRU/rcOcE0bNdnM0aZJ/h6xWQsGDHlhvjT2+1aJaAn/9k8473BRWajzVXld961CdHYVFVHoce2hHiSJ0xydWrHMMZhAm0mN0UzjEpgZ0tMw209efcZHIvSwuxhteZMRy4kyiVjwFlOf5oXFCxRuCJnPj3AK9CmCf4XgEBuPIJ0TZmjGHOOdBvJmbCNnAWXYEo5/mf7MfCGhV4IJ1hNuhpoNQfOFKMUcw9/v/IpT62XpfXdGYTfGYCmCjC+gntK1spbkr2P4/2+sYMQtLpse71mpSNGXfcf3abU55Vpn+gncSxRQIDAQAB"
|
|
35
35
|
}
|
package/docs/ARCHITECTURE.md
CHANGED
|
@@ -36,7 +36,7 @@ A canonical workspace receives an independent profile, Worker name, secret set,
|
|
|
36
36
|
- `runtime-capabilities.mjs` composes agent, application, browser, and effective-policy-filtered routing results, while `execution-routing.mjs` owns bounded set-level route scoring, ambiguity, fallbacks, and advisory tool projection;
|
|
37
37
|
- `runtime-tool-handlers.mjs` owns catalog-to-handler registration;
|
|
38
38
|
- `runtime-relay.mjs` owns relay construction and inbound envelope normalization, while `relay-call-recovery.mjs` owns the bounded disconnect grace, result queue, authoritative resumed-call reconciliation, same-daemon redelivery, and expiry cleanup;
|
|
39
|
-
- `macos-idle-sleep-assertion.mjs` is the single adapter boundary for the fixed macOS power primitive: `/usr/bin/caffeinate -i -w <owner-pid>` is spawned with `shell: false`, ignored stdio, and a process-lifetime owner binding; failures are best-effort and expose only a coarse error class. `remote-activity-idle-sleep-guard.mjs` composes that primitive for daemon-side relay activity. An authorized, schema-valid relay tool call begins activity only after policy/account/operation authorization and argument validation succeed; concurrent handlers share one assertion for their full execution lifetime, the fixed
|
|
39
|
+
- `macos-idle-sleep-assertion.mjs` is the single adapter boundary for the fixed macOS power primitive: `/usr/bin/caffeinate -i -w <owner-pid>` is spawned with `shell: false`, ignored stdio, and a process-lifetime owner binding; failures are best-effort and expose only a coarse error class. `remote-activity-idle-sleep-guard.mjs` composes that primitive for daemon-side relay activity. An authorized, schema-valid relay tool call begins activity only after policy/account/operation authorization and argument validation succeed; concurrent handlers share one assertion for their full execution lifetime, the fixed thirty-minute inactivity grace begins only after the last handler settles, and a new handler cancels any pending release timer so the full rolling grace restarts after the new activity settles. `process-session-remote-activity.mjs` extends the same daemon assertion beyond the `start_process` handler only after resource admission succeeds and releases it when the remote child settles, including startup failure. Remote account managed-job runners do not depend on daemon ownership: after the runner claim is confirmed and persisted ownership identifies an account-backed job, `job-runner.mjs` holds its own assertion bound to the runner PID across recovery handoff, resource admission, steps, `finally_steps`, and terminal persistence, then releases it in top-level `finally`. Local managed jobs do not acquire this remote-continuity assertion. Runtime shutdown terminates process sessions before releasing the daemon assertion. The thirty-minute relay inactivity grace is deliberately fixed rather than depending on shell environment inheritance that launchd does not persist as service configuration or treating relay heartbeats as user activity. These guards cover ordinary Idle Sleep during authorized remote work and bounded multi-turn gaps, but do not claim to defeat explicit sleep or lid-close sleep;
|
|
40
40
|
- `runtime-paths.mjs` owns runtime-directory creation, containment checks, and error-path redaction;
|
|
41
41
|
- `runtime-resource-service.mjs` owns registered-resource lookup, bounded binary/UTF-8 reads for browser/application injection, and SSH-resource registration/result projection;
|
|
42
42
|
- `security-audit-log.mjs` owns only the bounded main-thread queue and cached initializing/health projection; `security-audit-worker.mjs`, `security-audit-storage.mjs`, `security-audit-dispatch.mjs`, and `security-audit-warning.mjs` isolate startup verification, all disk/hash work, batch persistence, privacy projection, and warning suppression from result delivery;
|
package/docs/AUDIT.md
CHANGED
|
@@ -1,5 +1,15 @@
|
|
|
1
1
|
# Security and privacy audit notes
|
|
2
2
|
|
|
3
|
+
## 2026-08-18 beta.103 idle-sleep continuity correction
|
|
4
|
+
|
|
5
|
+
**Live beta.102 disproved the five-minute post-activity lease.** After beta.102 had passed candidate activation, live acceptance, cross-platform PR checks, and merge, owner use still produced a severe whole-control-plane interruption. The retained relay state showed several recovered transport episodes, but the longest contemporaneous event was not an awake network black hole: macOS power records show the final `PreventUserIdleSystemSleep` assertion disappearing and ordinary `Idle Sleep` beginning about five seconds later for roughly 632 seconds. The daemon recorded an aligned roughly 626-second event-loop pause and long inbound silence while launchd retained one continuously running service process with no termination. This proves the host was suspended rather than the daemon crashing. Because both preferred WSS and signed HTTPS fallback execute in that suspended userspace, transport redundancy cannot recover the machine while it is asleep.
|
|
6
|
+
|
|
7
|
+
**Causal defect is the inactivity-policy boundary, not the assertion primitive.** Beta.101 deliberately avoided permanently caffeinating the login daemon and chose a fixed five-minute grace after the last authorized relay activity. The reference-counting state machine itself behaved as designed: authorized/concurrent calls acquired one assertion, the last settlement armed the grace timer, and later activity would cancel that timer. The live failure demonstrates that five minutes is not long enough to model a real multi-turn remote workflow: model reasoning, user-visible waiting, and host scheduling can create a legitimate gap longer than five minutes between Machine Bridge calls. Once the timer released, the operating system was free to enter Idle Sleep immediately, making the bridge unreachable until physical wake. The earlier statement that the five-minute policy was sufficient for ordinary remote-work continuity is therefore falsified.
|
|
8
|
+
|
|
9
|
+
**Beta.103 uses a bounded thirty-minute rolling lease.** The daemon still acquires the macOS assertion only after policy/account/operation authorization and schema validation; rejected traffic and heartbeat traffic cannot keep the host awake. Concurrent authorized handlers remain reference-counted. After the last daemon-side activity settles, the assertion remains for thirty minutes; any later authorized activity cancels pending release and begins a fresh full thirty-minute inactivity window only after the new activity settles. Remote process sessions still extend ownership through child settlement, and remote account managed-job runners retain their independent runner-lifetime assertion. This is intentionally not relay-lifetime ownership: an idle but connected daemon eventually releases the assertion, avoiding permanent automatic-sleep suppression. Explicit sleep and lid-close remain unrecoverable host-power boundaries.
|
|
10
|
+
|
|
11
|
+
**Regression and release consequence.** A deterministic runtime regression first required the default constant and default constructor timer to retain the assertion for thirty minutes; it failed against beta.102's five-minute implementation, then passed after the product constant changed. The architecture release contract now requires the same thirty-minute source policy and current architecture/operations documentation. README, testing, and logging guidance describe the rolling renewal semantics and explicitly avoid heartbeat-based renewal. Because runtime source and shipped documentation change npm-package bytes, beta.102 acceptance cannot authorize beta.103. No beta.102 tag, GitHub Prerelease, npm publication, or formal published-prerelease soak had begun, so the correct recovery is a new beta.103 candidate and acceptance cycle rather than publishing or soaking the known-defective beta.102 bytes.
|
|
12
|
+
|
|
3
13
|
## 2026-08-18 beta.102 end-to-end interruption re-audit
|
|
4
14
|
|
|
5
15
|
**Live beta.101 falsified “fast detection equals recovery.”** A normal daemon-backed read reproduced the user-visible interruption while Worker-local `server_info` remained available. The Worker had no authenticated, ready, probing, or candidate daemon channel and the call exhausted the bounded recovery wait. After the daemon returned, retained evidence showed a 323,839 ms outage with 13 reconnect attempts ending in `relay_connect_timeout` and 9,830 ms of inbound silence before the previous ready WSS closed. launchd still reported the same PID, `runs=1`, and no termination signal; the matching interval contained no macOS Sleep/Wake event. A second live beta.101 interruption later recovered after about 10.6 seconds without a daemon restart. This proves the remaining product defect was not Computer Use, launchd crash, or idle sleep: beta.101 detected a broken WSS promptly but still depended on successful WSS reconnection as its only remote-daemon recovery transport. The host default route was carried by a tunnel-like interface during investigation, but available evidence does not identify a particular VPN/TUN product, node, Wi-Fi component, Cloudflare edge, ISP, or upstream device as the physical trigger.
|
package/docs/LOGGING.md
CHANGED
|
@@ -67,7 +67,7 @@ Brief network interruptions are expected on laptop network changes, Worker deplo
|
|
|
67
67
|
- failure to receive `hello_ack` within the handshake deadline, or `ready_ack` within the independent end-to-end readiness deadline, terminates the candidate socket and retries;
|
|
68
68
|
- ready transports send protocol-level WebSocket Ping every five seconds and classify ten seconds without inbound transport proof as `relay_transport_timeout`; while the daemon is scheduling-responsive, one probe interval plus that silence threshold bounds arbitrary black-hole detection to fifteen seconds, leaving reconnect/result-settlement headroom inside the ordinary twenty-second hosted execution default;
|
|
69
69
|
- a separate twenty-five-second application heartbeat refreshes Worker daemon activity and retains a seventy-five-second application-silence timeout; protocol-level Pong therefore cannot mask a Worker application path that has stopped replying. The Worker queues the heartbeat's JSON `pong` before Durable Object alarm inspection or mutation, then performs one explicit coalesced schedule, so storage latency is not allowed to sit ahead of application-liveness acknowledgement;
|
|
70
|
-
- a late local transport-watchdog tick is classified as `runtime.event_loop.stall`, sends a fresh transport probe, and defers disconnect for a bounded recovery interval instead of being mislabeled as immediate remote failure; a macOS sleep/wake interval may legitimately produce this warning without a daemon fault. An authorized, schema-valid remote tool call enters the bounded macOS idle-sleep guard only after policy/account/operation authorization and argument validation succeed; the shared assertion remains active for the full handler lifetime, remote process sessions extend it through child settlement, and the inactivity grace starts only after the last owned daemon-side activity settles. If the fixed macOS assertion child cannot be established, `runtime.idle_sleep_guard.unavailable` records only a coarse `error_class`, never argv, paths, PID, tool name, tool content, session identity, or job identity. Remote account managed-job runners use the same fixed assertion primitive but own it themselves only after runner-claim confirmation and persisted account ownership validation; local managed jobs do not acquire this remote-continuity assertion. The remote runner fallback stderr diagnostic is the fixed text `managed job idle-sleep assertion unavailable` plus a sanitized coarse `error_class`, without job name/id, workspace path, argv, environment, or captured output.
|
|
70
|
+
- a late local transport-watchdog tick is classified as `runtime.event_loop.stall`, sends a fresh transport probe, and defers disconnect for a bounded recovery interval instead of being mislabeled as immediate remote failure; a macOS sleep/wake interval may legitimately produce this warning without a daemon fault. An authorized, schema-valid remote tool call enters the bounded macOS idle-sleep guard only after policy/account/operation authorization and argument validation succeed; the shared assertion remains active for the full handler lifetime, remote process sessions extend it through child settlement, and the fixed thirty-minute rolling inactivity grace starts only after the last owned daemon-side activity settles, with each new authorized activity cancelling pending release and restarting the full grace after settlement. If the fixed macOS assertion child cannot be established, `runtime.idle_sleep_guard.unavailable` records only a coarse `error_class`, never argv, paths, PID, tool name, tool content, session identity, or job identity. Remote account managed-job runners use the same fixed assertion primitive but own it themselves only after runner-claim confirmation and persisted account ownership validation; local managed jobs do not acquire this remote-continuity assertion. The remote runner fallback stderr diagnostic is the fixed text `managed job idle-sleep assertion unavailable` plus a sanitized coarse `error_class`, without job name/id, workspace path, argv, environment, or captured output.
|
|
71
71
|
|
|
72
72
|
A WebSocket close code such as `1006` means the transport ended without a normal close handshake. If it recovers inside ten seconds, the warning-level service log is intentionally silent and the authenticated `daemon.relay_transport` snapshot is the post-event evidence surface. It is useful for debug diagnosis but not useful as the default user message. It is not evidence that the daemon process restarted. Worker `daemon_transport_error` / `daemon_liveness_timeout` messages and their 1012 close frames are likewise retryable connection conditions, not upgrade instructions. Only an unknown/incompatible Worker error, authentication failure, or identity/version mismatch may produce the fatal protocol/configuration log and daemon exit. Default logs therefore describe the affected layer, duration, classification, and recovery behavior rather than printing raw close envelopes.
|
|
73
73
|
|
package/docs/OPERATIONS.md
CHANGED
|
@@ -89,7 +89,7 @@ After the host path recovers, compare authenticated `server_info`, `machine-mcp
|
|
|
89
89
|
|
|
90
90
|
A reconnect warning proves a transport interruption, not a daemon crash. Compare daemon PID and process start time with `connected_at`, `last_seen_at`, `daemon.relay_transport.last_close_category`, `last_close_code`, `outage_count`, `outage_attempts`, `previous_ready_inbound_silence_ms`, and the coarse network-route class. While no daemon channel is ready, `server_info.daemon.previous_connection` retains only the last verified channel's transport, connected/last-seen/disconnected timestamps, and sanitized relay diagnostics; it excludes policy, tools, account identity, daemon instance/connection identity, call IDs, arguments, and results, and it never participates in routing or authorization. `outage_duration_ms` measures the close-to-ready recovery episode; `previous_ready_inbound_silence_ms` measures how long the preceding ready socket had stopped producing inbound transport proof before it actually closed. The second value is therefore the field that exposes a black-holed OPEN WebSocket whose visible reconnect later completes quickly. A VPN/TUN UI may remain “connected” while its upstream route is unusable. Local OS logs can be compared with the exact `last_disconnected_at` timestamp, but an interface-quality change or tunnel-process correlation is not by itself proof of which product, node, edge, or upstream failed. Machine Bridge reports only coarse route/proxy classes and never sends or logs interface names, addresses, DNS answers, proxy credentials, or Worker secrets.
|
|
91
91
|
|
|
92
|
-
Brief retryable outages recover automatically. On a verified current daemon channel, `server_info.daemon.relay_transport.outage_active=false`; retained fields describe the immediately preceding transport episode rather than claiming a current outage. WebSocket remains preferred and uses a five-second protocol-level probe with a ten-second inbound-silence timeout. The separate JSON application heartbeat remains twenty-five seconds with a seventy-five-second application-silence timeout, and the Worker keeps a wider ninety-second WebSocket liveness fallback. For a scheduling-responsive daemon, the WSS half-open detection horizon is at most fifteen seconds; that is a detection bound, not a guarantee that a degraded network can complete another WebSocket handshake inside the same interval. If verified WSS readiness is absent, the same root-certified ephemeral daemon identity starts the signed HTTPS fallback. Each fallback request has a seven-second deadline, ordinary one-second poll cadence, 750 ms minimum request-start interval, and twelve-second liveness window; a new daemon-backed call waits at most fifteen seconds for some verified daemon channel, and the measured wait is deducted from that call's original execution budget. After an established WSS disappears, the daemon explicitly marks its signed HTTP request as a same-instance takeover; once candidate preconditions pass, this allows HTTPS to retire a Worker-side zombie WSS that the Worker has not yet observed closing. Malformed, stale, or different-instance requests cannot preempt a healthy incumbent. During replacement, the daemon reconciles `resume_calls`, processes `ready_ack`, proves local readiness, and only then returns `resume_calls_ack.missing_ids`. A missing ID therefore proves both that the same daemon has no active/unacknowledged-result ownership for that call and that the replacement channel is ready. If the initiating MCP response is still open and at least one second remains in the original execution budget, the Worker transparently retransmits exactly that same call ID, arguments, authority, and a reduced timeout. If that safe redelivery cannot be accepted, the call falls back to retryable `unavailable` with `side_effects_started=false`. Calls that may have executed, retained terminal results, different-daemon calls, and ambiguous mutations are never automatically replayed. Compare outage intervals with sleep/wake records and `diagnose_runtime.runtime.relay.heartbeat` before classifying them as active network faults; local stdio `server_info.runtime.relay.heartbeat` exposes the same state. On macOS, an authorized, schema-valid remote tool call also enters a bounded idle-sleep guard with a fixed
|
|
92
|
+
Brief retryable outages recover automatically. On a verified current daemon channel, `server_info.daemon.relay_transport.outage_active=false`; retained fields describe the immediately preceding transport episode rather than claiming a current outage. WebSocket remains preferred and uses a five-second protocol-level probe with a ten-second inbound-silence timeout. The separate JSON application heartbeat remains twenty-five seconds with a seventy-five-second application-silence timeout, and the Worker keeps a wider ninety-second WebSocket liveness fallback. For a scheduling-responsive daemon, the WSS half-open detection horizon is at most fifteen seconds; that is a detection bound, not a guarantee that a degraded network can complete another WebSocket handshake inside the same interval. If verified WSS readiness is absent, the same root-certified ephemeral daemon identity starts the signed HTTPS fallback. Each fallback request has a seven-second deadline, ordinary one-second poll cadence, 750 ms minimum request-start interval, and twelve-second liveness window; a new daemon-backed call waits at most fifteen seconds for some verified daemon channel, and the measured wait is deducted from that call's original execution budget. After an established WSS disappears, the daemon explicitly marks its signed HTTP request as a same-instance takeover; once candidate preconditions pass, this allows HTTPS to retire a Worker-side zombie WSS that the Worker has not yet observed closing. Malformed, stale, or different-instance requests cannot preempt a healthy incumbent. During replacement, the daemon reconciles `resume_calls`, processes `ready_ack`, proves local readiness, and only then returns `resume_calls_ack.missing_ids`. A missing ID therefore proves both that the same daemon has no active/unacknowledged-result ownership for that call and that the replacement channel is ready. If the initiating MCP response is still open and at least one second remains in the original execution budget, the Worker transparently retransmits exactly that same call ID, arguments, authority, and a reduced timeout. If that safe redelivery cannot be accepted, the call falls back to retryable `unavailable` with `side_effects_started=false`. Calls that may have executed, retained terminal results, different-daemon calls, and ambiguous mutations are never automatically replayed. Compare outage intervals with sleep/wake records and `diagnose_runtime.runtime.relay.heartbeat` before classifying them as active network faults; local stdio `server_info.runtime.relay.heartbeat` exposes the same state. On macOS, an authorized, schema-valid remote tool call also enters a bounded idle-sleep guard with a fixed thirty-minute rolling inactivity grace; the service does not depend on shell-only environment overrides that launchd would not persist as configuration, and relay heartbeats do not count as user activity. Authorized relay handlers hold one shared `/usr/bin/caffeinate -i -w <daemon-pid>` assertion for their full execution lifetime; concurrent handlers share the child, the thirty-minute default inactivity grace begins only after the last one settles, and a new authorized handler cancels any pending release timer so the full grace restarts after that activity settles. A remote `start_process` extends the same assertion only after resource admission succeeds and keeps it until the session child settles, so a long process session is not reduced to the handler grace window. Remote account managed-job runners independently hold `/usr/bin/caffeinate -i -w <runner-pid>` after their ownership claim is confirmed and persisted ownership identifies an account-backed job, then retain it through admission, steps, cleanup, and terminal persistence; local managed jobs do not acquire this remote-continuity assertion, and daemon reconnect/replacement does not own the remote runner protection. `diagnose_runtime.runtime.idle_sleep_guard` reports only daemon-side supported/enabled/active/grace/error-class state; it intentionally does not enumerate process-session or job identities. Runtime shutdown terminates process sessions before releasing the daemon guard. None of these assertions claim to prevent explicit sleep or lid-close sleep. A nonzero `event_loop_stall_count` with a large `max_event_loop_lag_ms` means the local daemon was not scheduled promptly; a short stall enters recovery grace, sends a fresh transport probe, and deliberately postpones disconnect. A large stall that aligns with `pmset` Sleep/Wake is suspension evidence, while a large `previous_ready_inbound_silence_ms` without a matching local stall is stronger evidence of a half-open/network-path interruption before close. Use `--verbose` only when close codes, liveness deadlines, and retry delays are required.
|
|
93
93
|
|
|
94
94
|
A foreground MCP response is not durable delivery. Hosted synchronous calls reserve room for Worker and host settlement instead of occupying the complete interaction window: ordinary daemon-backed tools default to 20 seconds of remote execution plus a separate five-second Worker settlement margin; ordinary configurable browser/application foreground tools also default to 20 seconds, while compound `computer_observe` and `computer_act` default to 30 seconds; all configurable browser/application foreground tools retain their explicit 45-second maximum. Remote `exec_command`, `run_process`, and `run_local_command` no longer keep the child process inside that response lifetime. Each remote process request must carry a unique caller-held `idempotency_key` before dispatch; reuse that same key only when recovering an ambiguous acceptance response. The daemon commits the authorized operation as a principal-bound one-step managed job, launches it with interactive resource-admission priority, and returns a `job_id` plus `read_job` recovery metadata inside a 10-second acceptance budget; the Worker keeps a separate five-second settlement margin. If that acceptance response is lost to settlement timeout, HTTP response cancellation, or relay reconnect expiry after dispatch, the public error remains non-retryable for generic callers but carries the original key and the explicit recovery action `retry_same_tool_arguments_with_same_idempotency_key`; this reconciles against the retained job instead of authorizing a blind duplicate. The detached child may execute for up to 600 seconds after admission, but the managed runner can separately wait up to thirty minutes for cooperative machine-user resource admission before the child is spawned; the child execution deadline begins only after that admission succeeds. The shared ceiling is exposed machine-readably as `server_info.tool_delivery.managed_job_resource_admission_wait_max_ms`, because the same pre-spawn boundary applies to ordinary durable process jobs and owner `start_job` steps rather than to process tools alone. While the runner is in this pre-spawn state, `read_job.current_phase` is `resource_admission`; no command has started yet. An owner can correlate a long-running status at that phase with `diagnose_runtime.runtime.resource_admission` rather than interpreting it as a slow child process; a delegated non-owner should treat the phase itself as evidence that the child has not spawned, retain the same `job_id`, and avoid blind replay rather than attempting the owner-only machine-wide diagnostic. After admission, the phase returns to `steps`, `finally_steps`, or `recovery-cleanup` as appropriate. Completed step records preserve `duration_ms` as the total orchestration duration. Local/owner reads additionally expose `resource_admission_ms` as the pre-spawn portion so a delayed successful child can be distinguished from slow execution after the fact; delegated non-owner reads omit that machine-user scheduling timing rather than turning shared-host contention into a more precise cross-workload signal. The detached job survives MCP disconnect, relay reconnect, daemon restart, or service replacement. Non-owner process authority is unchanged: automatic durable execution still uses the delegated workspace sandbox and does not grant owner-only `start_job`. If a cached host schema omits the now-required key, the Worker rejects before daemon dispatch with a normal no-side-effect tool error and requests a `tools/list` refresh rather than surfacing a protocol-only validation failure. `start_process` remains the explicit daemon-lifetime path when interactive stdin or session-style incremental output is required, but hosted calls use a 10-second execution / 15-second settlement envelope and do not queue behind resource pressure: the first failed admission returns retryable `unavailable`; owner-local callers retain the cooperative wait. Use short `read_process` polls. A new hosted call waits at most fifteen seconds for daemon readiness, but that wait is charged against the call's existing execution budget; an in-flight disconnect likewise never pauses or extends the original absolute deadline. Owner-local stdio/CLI calls retain their synchronous local contract because they do not depend on a hosted response stream. Keep unrelated mutations and verification independently terminal, and never infer task success merely because a durable launch was accepted; read the job to a terminal state.
|
|
95
95
|
|
package/docs/TESTING.md
CHANGED
|
@@ -16,7 +16,7 @@ The explicit task lists live in `scripts/check-plan.mjs`. The fast plan retains
|
|
|
16
16
|
|
|
17
17
|
Tests are verification inputs but are not npm tarball entries under the current `package.json.files` manifest. A test-only repository change therefore does not by itself change npm package bytes or require a synthetic package version; `release-impact:check` remains authoritative and still requires a version whenever `package.json`, `package-lock.json`, or any path selected by `package.json.files` changes. A source fix that requires a regression test is versioned because of its packaged source/documentation/metadata impact, not because `tests/` is implicitly shipped.
|
|
18
18
|
|
|
19
|
-
On macOS, `scripts/run-checks.mjs` re-executes the complete fast/platform/full verification process under `/usr/bin/caffeinate -i` before it captures verification inputs. The internal `MBM_CHECK_IDLE_SLEEP_GUARD` marker prevents recursive wrapping. This establishes an idle-system-sleep assertion for the whole verification lifetime so a short child timeout cannot expire only because the machine entered Idle Sleep between dispatch and settlement. That verification-lifetime guard is distinct from production ownership: authorized, schema-valid relay handlers hold a shared `/usr/bin/caffeinate -i -w <daemon-pid>` assertion for their full execution lifetime; the default
|
|
19
|
+
On macOS, `scripts/run-checks.mjs` re-executes the complete fast/platform/full verification process under `/usr/bin/caffeinate -i` before it captures verification inputs. The internal `MBM_CHECK_IDLE_SLEEP_GUARD` marker prevents recursive wrapping. This establishes an idle-system-sleep assertion for the whole verification lifetime so a short child timeout cannot expire only because the machine entered Idle Sleep between dispatch and settlement. That verification-lifetime guard is distinct from production ownership: authorized, schema-valid relay handlers hold a shared `/usr/bin/caffeinate -i -w <daemon-pid>` assertion for their full execution lifetime; the default thirty-minute rolling inactivity grace starts only after the last concurrent handler settles, and a new authorized handler cancels pending release so the full grace restarts after it settles. A remote process session extends the same daemon assertion from successful resource admission through child settlement. A remote account managed-job runner independently holds `/usr/bin/caffeinate -i -w <runner-pid>` only after its runner claim is confirmed and persisted ownership identifies an account-backed job, then retains it through terminal persistence; local managed jobs do not acquire that remote-continuity assertion, and the remote runner protection survives daemon reconnect/replacement. Runtime shutdown terminates process sessions before releasing the daemon assertion. The production thirty-minute relay grace is fixed rather than depending on shell-only environment inheritance that launchd does not persist as service configuration or treating relay heartbeats as user activity. None of these mechanisms claims to defeat explicit sleep or lid-close sleep; those still invalidate the usefulness of a live-machine verification run.
|
|
20
20
|
|
|
21
21
|
Successful child-task stdout/stderr is intentionally suppressed so a long green plan does not overwhelm an MCP response or hide the final status behind host truncation. Progress and timing remain visible. A failed task returns bounded head/tail diagnostics for both streams. Set `MBM_CHECK_VERBOSE=1` only when an operator explicitly needs live child output; verbose mode can be large and should be run through a process session when used remotely.
|
|
22
22
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "machine-bridge-mcp",
|
|
3
|
-
"version": "3.0.0-beta.
|
|
3
|
+
"version": "3.0.0-beta.103",
|
|
4
4
|
"description": "Cross-client MCP bridge for local agent context, structured browser and application automation, files, Git, processes, resources, and durable jobs over stdio or OAuth relay.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
import { MacosIdleSleepAssertion } from "./macos-idle-sleep-assertion.mjs";
|
|
2
2
|
|
|
3
|
-
export const DEFAULT_REMOTE_ACTIVITY_IDLE_SLEEP_GRACE_MS =
|
|
3
|
+
export const DEFAULT_REMOTE_ACTIVITY_IDLE_SLEEP_GRACE_MS = 30 * 60_000;
|
|
4
4
|
|
|
5
5
|
export class RemoteActivityIdleSleepGuard {
|
|
6
6
|
constructor({ platform = process.platform, daemonPid = process.pid, graceMs = DEFAULT_REMOTE_ACTIVITY_IDLE_SLEEP_GRACE_MS,
|
package/src/worker/index.ts
CHANGED
|
@@ -57,7 +57,7 @@ import {
|
|
|
57
57
|
closeWebSocketQuietly, daemonErrorCloseCode, isObjectRecord, rejectDaemonMessage,
|
|
58
58
|
sendWebSocketQuietly, trySendWebSocket,
|
|
59
59
|
} from "./websocket-protocol.ts";
|
|
60
|
-
const SERVER_VERSION = "3.0.0-beta.
|
|
60
|
+
const SERVER_VERSION = "3.0.0-beta.103";
|
|
61
61
|
const MCP_SERVER_INFO = mcpServerInfo(SERVER_VERSION);
|
|
62
62
|
const MAX_DAEMON_MESSAGE_BYTES = 8 * 1024 * 1024;
|
|
63
63
|
const DAEMON_RECONNECT_GRACE_MS = relayContract.reconnectGraceMs; const NEW_CALL_RECONNECT_GRACE_MS = relayContract.newCallReconnectGraceMs;
|