machine-bridge-mcp 3.0.0-beta.141 → 3.0.0-beta.144
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +18 -0
- package/README.md +1 -1
- package/browser-extension/manifest.json +1 -1
- package/docs/ARCHITECTURE.md +3 -3
- package/docs/AUDIT.md +30 -0
- package/docs/LOGGING.md +2 -2
- package/docs/MANAGED_JOBS.md +4 -4
- package/docs/OPERATIONS.md +5 -5
- package/docs/TESTING.md +4 -4
- package/docs/TOOL_REFERENCE.md +3 -3
- package/package.json +1 -1
- package/src/local/macos-idle-sleep-assertion.mjs +7 -2
- package/src/local/managed-job-dependencies.mjs +15 -2
- package/src/local/managed-job-listing.mjs +3 -2
- package/src/local/managed-job-recovery-listing.mjs +18 -0
- package/src/local/remote-activity-idle-sleep-guard.mjs +2 -2
- package/src/local/runtime-diagnostic-projection.mjs +1 -0
- package/src/local/runtime-diagnostic-state.mjs +1 -0
- package/src/local/runtime-diagnostics.mjs +2 -2
- package/src/local/system-sleep-diagnostics.mjs +68 -0
- package/src/shared/server-metadata.json +4 -4
- package/src/shared/tool-catalog.json +3 -3
- package/src/worker/index.ts +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,23 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 3.0.0-beta.144 - 2026-08-27
|
|
4
|
+
|
|
5
|
+
- Fix a real Windows managed-job dependency recovery defect exposed by exact-main provider CI after beta.143 acceptance. Runner-exit reconciliation already treats transient `permission_denied`/conflict/timeout/resource-unavailable failures as retryable, but a concurrently waiting downstream previously treated any secure upstream status-read error as permanent `dependency_unavailable`. Dependency polling now gives only `permission_denied`, `identity_changed`, and generic `resource_unavailable` reads a fixed 45-second monotonic recovery grace; a successful secure read clears that grace immediately, persistent unavailability still fails closed, and missing/integrity/witness-invalid evidence remains non-retryable.
|
|
6
|
+
- Correct the Worker integration fixture race that obscured the first exact-main diagnosis. Eight tests started an MCP HTTP request before registering the WebSocket `tool_call` listener, so a fast relay dispatch could arrive before the harness was listening and be lost permanently; increasing the waiter from five to ten seconds could not fix that ordering bug. The fixture now registers the waiter before triggering the request. Three local Worker integration runs and hosted Ubuntu/full passed with the listener-first ordering.
|
|
7
|
+
- Supersede beta.143 release evidence because the dependency-recovery repair changes packaged runtime bytes and hosted orchestration semantics. The beta.143 acceptance record is removed, package/runtime identity advances to beta.144, and tool schema generation advances to 14. Fresh frozen verification, exact candidate preparation/install-only proof, guarded activation, activated-package OAuth canary, live observation, acceptance, and exact-head provider gates are required before GitHub prerelease publication. npm publication remains separately owner-authorized for the exact beta.144 version.
|
|
8
|
+
|
|
9
|
+
## 3.0.0-beta.143 - 2026-08-27
|
|
10
|
+
|
|
11
|
+
- Correct interruption diagnosis after live beta.142 evidence separated host suspension from awake transport resets. A programmatic correlation of fifteen recovered August 23 relay outages with bounded macOS `pmset` sleep intervals found thirteen sleep-dominated episodes: 32,980 of 33,390 outage seconds (98.8%) occurred while the host was suspended. `diagnose_runtime.runtime.relay_outage_analysis` now intersects the most recently completed close-to-ready relay interval with the same bounded sleep history and reports only coarse timing/overlap fields. `majority_system_sleep_overlap` therefore prevents a sleep-dominated `connection_reset` from being promoted into an unsupported independent-network root cause, while an awake reset with no sleep overlap remains real transport evidence without guessing which external layer failed.
|
|
12
|
+
- Strengthen the existing macOS remote-work power assertion on AC without extending its ownership lifetime. The fixed runtime/remote-runner primitive now launches `/usr/bin/caffeinate -i -s -w <owner-pid>` instead of `-i -w`: `-i` retains Idle Sleep prevention and Apple's `-s` assertion adds `PreventSystemSleep` only while on AC power. A live three-second probe on the affected Mac observed both assertions concurrently. The thirty-minute relay inactivity grace, process-session ownership, detached account-job ownership, fail-open error handling, and explicit/lid-close/battery limitations are unchanged; diagnostics report only whether the fixed child requests the AC-only system-sleep assertion.
|
|
13
|
+
- Keep the external-host boundary explicit. The beta.142 `UNKNOWN / TaskGroup` episode occurred during an approximately nine-second awake relay `1006/connection_reset`, while the same durable job continued and was recovered by the same `job_id`. Worker pending-call deadlines already preserve the original absolute timeout across reconnect, and the public MCP response path already emits an immediate SSE priming frame plus five-second heartbeats. This tree does not invent a host timeout or shorten the proven forty-second hosted `read_job` default without evidence. Tool schema generation advances to 13 because owner-visible `diagnose_runtime` result semantics and guidance changed. npm registry publication remains a separate explicit owner-authorization boundary.
|
|
14
|
+
|
|
15
|
+
## 3.0.0-beta.142 - 2026-08-27
|
|
16
|
+
|
|
17
|
+
- Make retained one-step process recovery discoverable after a real host/tool boundary. A saturated live 512-state store can legitimately contain 496 durable terminal jobs plus the beta.141 sixteen-result transient recovery reserve; the existing durable-first `list_jobs.jobs` window may then contain no recent process helpers even though a known helper `job_id` is still readable. `list_jobs` now keeps that primary window unchanged and adds `recent_process_recovery`, capped at 16 authority-visible public job handles for recent transient terminal results that are still inside the existing thirty-minute recovery grace but omitted from `jobs`.
|
|
18
|
+
- Preserve the request-scoped MCP boundary. The new recovery projection carries public job status only: no step output, argv, path, internal `retention_class`, conversation identity, terminal-result replay, or cross-request response session is added. It is inventory for recovering a lost `job_id`; known work still continues through `read_job`, and `list_jobs` remains a non-polling surface. The 512-state retention cap, 50-record primary inventory, 16-result/thirty-minute transient retention reserve, durable-first ordering, dependency pinning, and eviction priorities are unchanged.
|
|
19
|
+
- Advance hosted tool schema generation to 12 because `list_jobs` public result semantics and tool guidance changed. These shipped runtime, catalog, tests, and documentation changes supersede beta.141 release evidence and require a fresh prerelease verification/candidate/activation sequence before any publication. npm registry publication remains separately owner-authorized.
|
|
20
|
+
|
|
3
21
|
## 3.0.0-beta.141 - 2026-08-26
|
|
4
22
|
|
|
5
23
|
- Preserve immediate recovery evidence for remote one-step process carriers under a saturated 512-state managed-job store. The previous eviction order always discarded `transient_process` terminal results before any ordinary durable terminal history; with 510 durable terminals already retained, short `exec_command` helpers could therefore return a recoverable job ID and then become `not_found` before the next `read_job`. Beta.141 reserves at most the newest 16 otherwise-removable transient results for thirty minutes. The incoming transient counts against that bound, older/excess transient history is still evicted first, ordinary durable history is next, dependency-protected records remain pinned, and the hard 512-state cap is unchanged.
|
package/README.md
CHANGED
|
@@ -193,7 +193,7 @@ For stateful GUI trajectories, owner/full callers can use the higher-level `comp
|
|
|
193
193
|
|
|
194
194
|
## Durable work and local resources
|
|
195
195
|
|
|
196
|
-
Remote request-owned foreground work uses the hosted reply-safe budgets described above; configurable browser/application calls may explicitly request at most 45 seconds, while remote `exec_command`, `run_process`, and `run_local_command` are durable one-step jobs with a 10-second acceptance envelope and an independent 1–600-second child execution budget after admission. A continuous process that legitimately needs more than 600 seconds must use `start_job`: managed-job main/finally steps default to 600 seconds and may explicitly request up to 21,600 seconds (six hours), with resource admission occurring before that execution timer begins. The Worker retains separate settlement ownership for five additional seconds, but neither that margin nor its internal stream metrics prove that an external MCP host consumed the terminal frame. Keep unrelated mutations and validation independently terminal, but batch one coherent non-interactive command sequence into a repository umbrella command or multi-step `start_job` instead of creating one host-visible one-step job per tiny probe. A timeout is a protocol result, not proof that descendant cleanup has already completed; a remote owner can inspect `diagnose_runtime.runtime.processes`, while local stdio exposes `server_info.runtime.processes`. Non-owner accounts receive authority-scoped readiness rather than machine-wide process activity. Remote process sessions are for interactive stdin or incremental output, not a substitute for ordinary durable work: hosted `read_process` reports `status_polling_mode=paced_followup`, caps the actual output/exit blocking wait at one second, and paces a repeated would-block read inside the fifteen-second cooldown within that same MCP call until output/exit or the cooldown boundary instead of returning a rapid running checkpoint. When the current task needs more output or terminal state, the same session may be read again in the same assistant response without busy-looping. Non-interactive work should use durable `run_process`/`read_job`; multi-step, cleanup-sensitive, or daemon-restart-surviving workflows should use managed jobs, which persist ordered argv steps and `finally_steps` under owner-only local state and continue across an MCP disconnect. Durable acceptance does not force a hosted-turn handoff: active relay-origin `read_job` reports `status_polling_mode=bounded_followup` and `host_turn_handoff_recommended=false`. Its hosted default is a 40-second server-side long-poll, so an unchanged long job occupies one bounded live MCP response rather than forcing rapid host-side checkpoints. Terminal settlement returns on the next bounded five-second internal poll; nonterminal status/phase/dependency progress is coalesced for at least 30 seconds by default, and `current_step`-only churn does not wake the hosted call. `wait_ms=0` is available only when an immediate checkpoint is actually wanted, and an explicitly capable client may request up to five minutes. At the default, a synthetic 100-minute unchanged job has an anti-amplification ceiling of 150 status reads, while continuously changing nonterminal progress is separately bounded by the 30-second coalescing floor; those are density estimates rather than proof of aggregate same-response host lifetime. A known job may be followed through paced same-response `read_job` calls while those calls continue to be accepted and the task still needs the result; after an actual host/tool boundary, later recovery must continue from the same `job_id`. Completed one-step process carriers are lower-priority terminal retention than explicit managed jobs, so removable helper history is reclaimed first under the 512-state durable cap
|
|
196
|
+
Remote request-owned foreground work uses the hosted reply-safe budgets described above; configurable browser/application calls may explicitly request at most 45 seconds, while remote `exec_command`, `run_process`, and `run_local_command` are durable one-step jobs with a 10-second acceptance envelope and an independent 1–600-second child execution budget after admission. A continuous process that legitimately needs more than 600 seconds must use `start_job`: managed-job main/finally steps default to 600 seconds and may explicitly request up to 21,600 seconds (six hours), with resource admission occurring before that execution timer begins. The Worker retains separate settlement ownership for five additional seconds, but neither that margin nor its internal stream metrics prove that an external MCP host consumed the terminal frame. Keep unrelated mutations and validation independently terminal, but batch one coherent non-interactive command sequence into a repository umbrella command or multi-step `start_job` instead of creating one host-visible one-step job per tiny probe. A timeout is a protocol result, not proof that descendant cleanup has already completed; a remote owner can inspect `diagnose_runtime.runtime.processes`, while local stdio exposes `server_info.runtime.processes`. Non-owner accounts receive authority-scoped readiness rather than machine-wide process activity. Remote process sessions are for interactive stdin or incremental output, not a substitute for ordinary durable work: hosted `read_process` reports `status_polling_mode=paced_followup`, caps the actual output/exit blocking wait at one second, and paces a repeated would-block read inside the fifteen-second cooldown within that same MCP call until output/exit or the cooldown boundary instead of returning a rapid running checkpoint. When the current task needs more output or terminal state, the same session may be read again in the same assistant response without busy-looping. Non-interactive work should use durable `run_process`/`read_job`; multi-step, cleanup-sensitive, or daemon-restart-surviving workflows should use managed jobs, which persist ordered argv steps and `finally_steps` under owner-only local state and continue across an MCP disconnect. Durable acceptance does not force a hosted-turn handoff: active relay-origin `read_job` reports `status_polling_mode=bounded_followup` and `host_turn_handoff_recommended=false`. Its hosted default is a 40-second server-side long-poll, so an unchanged long job occupies one bounded live MCP response rather than forcing rapid host-side checkpoints. Terminal settlement returns on the next bounded five-second internal poll; nonterminal status/phase/dependency progress is coalesced for at least 30 seconds by default, and `current_step`-only churn does not wake the hosted call. `wait_ms=0` is available only when an immediate checkpoint is actually wanted, and an explicitly capable client may request up to five minutes. At the default, a synthetic 100-minute unchanged job has an anti-amplification ceiling of 150 status reads, while continuously changing nonterminal progress is separately bounded by the 30-second coalescing floor; those are density estimates rather than proof of aggregate same-response host lifetime. A known job may be followed through paced same-response `read_job` calls while those calls continue to be accepted and the task still needs the result; after an actual host/tool boundary, later recovery must continue from the same `job_id`. Completed one-step process carriers are lower-priority terminal retention than explicit managed jobs, so removable helper history is reclaimed first under the 512-state durable cap. `list_jobs.jobs` keeps a 50-record durable-first primary window; if a recent one-step terminal remains inside the fixed thirty-minute/16-result recovery reserve but is omitted from that window, `recent_process_recovery` returns up to 16 additional authority-visible public job handles so a later host turn can recover the `job_id` and continue with `read_job`. That secondary projection includes no step output or internal retention metadata and remains inventory rather than MCP replay/session state or a polling surface. Owner/local `capacity` diagnostics expose only coarse `durable_terminal` and `transient_terminal` counts; this improves recovery visibility without pretending that Worker acknowledgement proves an external host rendered the final assistant message. Long cross-job workflows can declare `depends_on`: the dependent job remains pre-execution `queued/dependency_wait` without spawning its main child until all upstream jobs succeed, and an upstream failure settles `dependency_failed` instead of leaving a file-poll loop waiting for an artifact that can never appear. Active/staged dependency plans pin referenced retained results until the dependency-bearing plan is terminal. A valid `job_id` that is no longer retained returns typed `not_found`; that absence is not proof that its underlying side effect never executed. `list_jobs` remains inventory rather than a substitute polling loop, and `server_info`/`diagnose_runtime` remain diagnostic surfaces rather than alternate wait channels. Elapsed minutes are not evidence that an external host deadline is near; return the durable recovery identifier for a later turn only after an actual host/tool boundary is observed, external input or authorization is required, or the user explicitly requested a checkpoint.
|
|
197
197
|
|
|
198
198
|
On macOS, authorized remote activity uses a bounded idle-sleep assertion so ordinary system Idle Sleep does not suspend an active remote workflow. Relay handlers share the assertion for their execution lifetime plus a fixed thirty-minute rolling inactivity grace; each new authorized remote activity cancels a pending release and restarts the full grace after the last concurrent handler settles. An admitted remote process session extends daemon-side ownership until its child settles, and an account-backed managed-job runner owns a runner-bound assertion from confirmed claim through terminal persistence. Local managed jobs do not acquire the remote-continuity assertion. These protections do not override explicit sleep or lid-close behavior.
|
|
199
199
|
|
|
@@ -30,6 +30,6 @@
|
|
|
30
30
|
"action": {
|
|
31
31
|
"default_title": "Machine Bridge Browser"
|
|
32
32
|
},
|
|
33
|
-
"version_name": "3.0.0-beta.
|
|
33
|
+
"version_name": "3.0.0-beta.144",
|
|
34
34
|
"key": "MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAxryYkpZhq8+VAQLHcGS9BAHQcyKX8RHGIpIwvtIVRU/rcOcE0bNdnM0aZJ/h6xWQsGDHlhvjT2+1aJaAn/9k8473BRWajzVXld961CdHYVFVHoce2hHiSJ0xydWrHMMZhAm0mN0UzjEpgZ0tMw209efcZHIvSwuxhteZMRy4kyiVjwFlOf5oXFCxRuCJnPj3AK9CmCf4XgEBuPIJ0TZmjGHOOdBvJmbCNnAWXYEo5/mf7MfCGhV4IJ1hNuhpoNQfOFKMUcw9/v/IpT62XpfXdGYTfGYCmCjC+gntK1spbkr2P4/2+sYMQtLpse71mpSNGXfcf3abU55Vpn+gncSxRQIDAQAB"
|
|
35
35
|
}
|
package/docs/ARCHITECTURE.md
CHANGED
|
@@ -32,11 +32,11 @@ A canonical workspace receives an independent profile, Worker name, secret set,
|
|
|
32
32
|
- `workspace-file-service.mjs` and `git-service.mjs` own canonical filesystem/Git operations; `file-mutation-coordinator.mjs` supplies the shared resolved-path conflict identity and serializes only conflicting paths, registers every path in a multi-file transaction before waiting, and holds those reservations until the filesystem callback settles. `workspace-file-transaction.mjs` owns flushed staging, no-overwrite target publication, whole-file compare-and-swap, patch rollback, and cleanup causality. A late patch target becomes `conflict/target_appeared`; a rollback failure after any user-file commit is a public non-retryable `patch_recovery_incomplete` settlement directing the caller to inspect affected paths, while staging-only cleanup failure remains an internal cleanup error and post-commit artifact cleanup remains a warning rather than a false rollback claim;
|
|
33
33
|
- `process-contract.mjs` owns argv shape/size validation; `process-tree-signal.mjs`, `process-tree-supervisor.mjs`, `process-tree-snapshot.mjs`, and `process-tree-ownership.mjs` separate cross-platform signaling, asynchronous escalation, bounded process-group observation, and PID/start-time ownership; `process-execution.mjs` and `process-sessions.mjs` own one-shot and interactive execution; `process-nonreplayable-settlement.mjs` maps fixed internal UI-launch/native-input helpers from definite pre-spawn failure to non-retryable unknown settlement after process start; and `process-tracker.mjs` retains active and draining process ownership until close;
|
|
34
34
|
- `shared/tool-call-capacity.mjs` defines the control-tool set and generic admission algebra; local `call-capacity.mjs` and Worker `pending-call-capacity.ts` apply it independently, while `runtime-reporting.mjs` builds privacy-aware runtime and project snapshots;
|
|
35
|
-
- `runtime-diagnostics.mjs` owns fixed local probes and their stable interpretation, while `runtime-diagnostic-state.mjs` projects privacy-safe control-plane state for remote diagnosis. On macOS, `system-sleep-diagnostics.mjs` is the only power-history parser used by that surface: it runs a fixed bounded `pmset` pipeline, returns only recent sleep interval timestamps/durations/coarse reasons,
|
|
35
|
+
- `runtime-diagnostics.mjs` owns fixed local probes and their stable interpretation, while `runtime-diagnostic-state.mjs` projects privacy-safe control-plane state for remote diagnosis. On macOS, `system-sleep-diagnostics.mjs` is the only power-history parser used by that surface: it runs a fixed bounded `pmset` pipeline, returns only recent sleep interval timestamps/durations/coarse reasons, correlates an event-loop stall only when both its resume time and duration match an operating-system sleep interval within the fixed tolerance, and independently intersects the most recently completed relay disconnect interval with those same sleep intervals. A majority-sleep relay overlap means host suspension dominated the observed outage and a retained socket reset is aftermath evidence rather than proof of a separate network root cause; an unmatched pause or awake relay reset remains unassigned instead of being guessed into another cause;
|
|
36
36
|
- `runtime-capabilities.mjs` composes agent, application, browser, and effective-policy-filtered routing results, while `execution-routing.mjs` owns bounded set-level route scoring, ambiguity, fallbacks, and advisory tool projection;
|
|
37
37
|
- `runtime-tool-handlers.mjs` owns catalog-to-handler registration;
|
|
38
38
|
- `runtime-relay.mjs` owns relay construction and inbound envelope normalization; `relay-call-recovery.mjs` owns bounded disconnect grace, authoritative resumed-call reconciliation, and same-daemon redelivery orchestration; `relay-result-retention.mjs` owns the bounded completed-but-unacknowledged result ledger plus the fail-closed automatic-redelivery proof; and `relay-recovery-admission.mjs` owns recovery-capacity accounting and its privacy-safe diagnostic projection;
|
|
39
|
-
- `macos-idle-sleep-assertion.mjs` is the single adapter boundary for the fixed macOS power primitive: `/usr/bin/caffeinate -i -w <owner-pid>` is spawned with `shell: false`, ignored stdio, and a process-lifetime owner binding; failures are best-effort and expose only a coarse error class. `remote-activity-idle-sleep-guard.mjs` composes that primitive for daemon-side relay activity. An authorized, schema-valid relay tool call begins activity only after policy/account/operation authorization and argument validation succeed; concurrent handlers share one assertion for their full execution lifetime, the fixed thirty-minute inactivity grace begins only after the last handler settles, and a new handler cancels any pending release timer so the full rolling grace restarts after the new activity settles. `process-session-remote-activity.mjs` extends the same daemon assertion beyond the `start_process` handler only after resource admission succeeds and releases it when the remote child settles, including startup failure. Remote account managed-job runners do not depend on daemon ownership: after the runner claim is confirmed and persisted ownership identifies an account-backed job, `job-runner.mjs` holds its own assertion bound to the runner PID across recovery handoff, resource admission, steps, `finally_steps`, and terminal persistence, then releases it in top-level `finally`. Local managed jobs do not acquire this remote-continuity assertion. Runtime shutdown terminates process sessions before releasing the daemon assertion. The thirty-minute relay inactivity grace is deliberately fixed rather than depending on shell environment inheritance that launchd does not persist as service configuration or treating relay heartbeats as user activity. These guards
|
|
39
|
+
- `macos-idle-sleep-assertion.mjs` is the single adapter boundary for the fixed macOS power primitive: `/usr/bin/caffeinate -i -s -w <owner-pid>` is spawned with `shell: false`, ignored stdio, and a process-lifetime owner binding; failures are best-effort and expose only a coarse error class. The `-i` assertion retains Idle Sleep prevention, while `-s` requests the stronger PreventSystemSleep assertion that macOS documents as effective only on AC power; diagnostics report only that the fixed child requests that AC-only protection, not the current power source or a guarantee against explicit/lid-close sleep. `remote-activity-idle-sleep-guard.mjs` composes that primitive for daemon-side relay activity. An authorized, schema-valid relay tool call begins activity only after policy/account/operation authorization and argument validation succeed; concurrent handlers share one assertion for their full execution lifetime, the fixed thirty-minute inactivity grace begins only after the last handler settles, and a new handler cancels any pending release timer so the full rolling grace restarts after the new activity settles. `process-session-remote-activity.mjs` extends the same daemon assertion beyond the `start_process` handler only after resource admission succeeds and releases it when the remote child settles, including startup failure. Remote account managed-job runners do not depend on daemon ownership: after the runner claim is confirmed and persisted ownership identifies an account-backed job, `job-runner.mjs` holds its own assertion bound to the runner PID across recovery handoff, resource admission, steps, `finally_steps`, and terminal persistence, then releases it in top-level `finally`. Local managed jobs do not acquire this remote-continuity assertion. Runtime shutdown terminates process sessions before releasing the daemon assertion. The thirty-minute relay inactivity grace is deliberately fixed rather than depending on shell environment inheritance that launchd does not persist as service configuration or treating relay heartbeats as user activity. These guards strengthen AC continuity without claiming to defeat explicit sleep, lid-close sleep, or battery-mode system sleep;
|
|
40
40
|
- `runtime-paths.mjs` owns runtime-directory creation, containment checks, and error-path redaction;
|
|
41
41
|
- `runtime-resource-service.mjs` owns registered-resource lookup, bounded binary/UTF-8 reads for browser/application injection, and SSH-resource registration/result projection;
|
|
42
42
|
- `security-audit-log.mjs` owns only the bounded main-thread queue and cached initializing/health projection; `security-audit-worker.mjs`, `security-audit-storage.mjs`, `security-audit-dispatch.mjs`, and `security-audit-warning.mjs` isolate startup verification, all disk/hash work, batch persistence, privacy projection, and warning suppression from result delivery;
|
|
@@ -87,7 +87,7 @@ See [Local application and browser automation](LOCAL_AUTOMATION.md).
|
|
|
87
87
|
|
|
88
88
|
### Managed job runner
|
|
89
89
|
|
|
90
|
-
`ManagedJobManager` persists bounded per-workspace job envelopes below the owner-only profile directory. Managed-job active and terminal lifecycle classifications have one shared source in `managed-job-terminal.mjs`; manager reconciliation, retention, detached-runner fatal settlement, and production full-access diagnostics consume that source instead of maintaining parallel status lists. `start_job` validates the complete plan, snapshots referenced resource metadata/hashes, writes an owner-only plan/status, and launches `job-runner.mjs` as a detached process with runner-level logs redirected to owner-only files. `stage_job` performs the same acceptance validation but writes a non-running, review-only `staged` envelope; staged records have no promotion/approval execution path, so execution requires a separate trusted `start_job` request or an explicit local `job submit` plan. `managed-job-dependency-admission.mjs` binds optional `depends_on` references to same-authority durable job identities and rejects staged or already-failed dependencies before acceptance. `managed-job-dependencies.mjs` then keeps a dependent runner in pre-execution `queued/dependency_wait` state until all upstream jobs succeed, or settles it with `dependency_failed` when one later fails; no dependent main child or process resource lease exists during that wait. `managed-job-relaunch.mjs` preserves
|
|
90
|
+
`ManagedJobManager` persists bounded per-workspace job envelopes below the owner-only profile directory. Managed-job active and terminal lifecycle classifications have one shared source in `managed-job-terminal.mjs`; manager reconciliation, retention, detached-runner fatal settlement, and production full-access diagnostics consume that source instead of maintaining parallel status lists. `start_job` validates the complete plan, snapshots referenced resource metadata/hashes, writes an owner-only plan/status, and launches `job-runner.mjs` as a detached process with runner-level logs redirected to owner-only files. `stage_job` performs the same acceptance validation but writes a non-running, review-only `staged` envelope; staged records have no promotion/approval execution path, so execution requires a separate trusted `start_job` request or an explicit local `job submit` plan. `managed-job-dependency-admission.mjs` binds optional `depends_on` references to same-authority durable job identities and rejects staged or already-failed dependencies before acceptance. `managed-job-dependencies.mjs` then keeps a dependent runner in pre-execution `queued/dependency_wait` state until all upstream jobs succeed, or settles it with `dependency_failed` when one later fails; no dependent main child or process resource lease exists during that wait. Because Windows atomic replacement/runner-exit recovery can make an otherwise valid status file transiently unreadable, dependency polling gives only `permission_denied`, `identity_changed`, and generic `resource_unavailable` reads a fixed 45-second monotonic recovery grace. One successful secure read clears that grace; persistent unavailability still fails closed as `dependency_unavailable`, while missing, corrupt, witness-mismatched, staged, or otherwise invalid evidence is never converted into a retry. `managed-job-relaunch.mjs` preserves the pre-execution distinction across a dead dependency-wait runner by restarting the original `dependency_wait` job rather than converting it to cleanup-only recovery. `managed-job-retention.mjs` owns staged expiry timing and seven-day terminal retention; `managed-job-capacity.mjs` bounds durable retained state at 512 while the public `list_jobs.jobs` primary response window remains capped at 50 records. Recent one-step terminal jobs that are still inside the fixed thirty-minute/16-result recovery reserve but omitted from that durable-first primary window are projected separately through `recent_process_recovery`, capped at 16 authority-visible public job handles and without step output or internal retention metadata. This recovery-discovery projection does not create MCP replay/session state or change request ownership. Completed one-step process carriers are marked internally as `transient_process`. Under hard capacity pressure, retention preserves at most the newest 16 otherwise-removable transient results that completed within the last thirty minutes so an accepted remote helper cannot normally lose its recovery evidence before immediate `read_job` follow-up. The incoming transient reservation counts against that 16-result window: excess or older transient history is reclaimed first, then the oldest ordinary durable terminal history, and only then a still-recent protected transient when no other removable record exists. This recovery reserve does not enlarge the 512-state cap and never overrides dependency protection. `managed-job-dependency-retention.mjs` pins terminal records referenced by active/staged dependency plans and fails closed if that protection cannot be read safely. `managed-job-directory.mjs` maps a syntactically valid but no-longer-retained job ID to fixed non-retryable `not_found`; that absence is recovery-evidence loss rather than proof that the underlying operation never executed. `managed-job-terminal-maintenance.mjs` owns post-settlement evidence validation and artifact scrubbing; `managed-job-directory-generation.mjs` binds whole-directory retirement to the exact filesystem generation that retention inspected. Retirement first revalidates the full observed generation, atomically renames that directory to an internal `retired_job_*` name carrying its device/inode identity, revalidates the moved object, and only then recursively deletes it. The retirement namespace deliberately does not match the public `MANAGED_JOB_ID` grammar, so list/read/lock scans cannot reinterpret internal cleanup state as an ordinary job. A crash after rename therefore leaves recognizable state rather than an anonymous orphan: a later maintenance pass reclaims it only when the encoded generation still matches, while a type mismatch, generation mismatch, or unreadable retired entry is projected into active-state inventory as a privacy-bounded `retired_managed_job`/`unreadable` blocker without exposing the internal filename or filesystem identity. Before destructive terminal cleanup, status/result must describe the same directory job ID, terminal state, and `finished_at` generation; the degraded `result_persisted=false` form must carry an explicit terminal-record error class. Corrupt terminal evidence is therefore retained as unreadable state and also blocks state removal instead of authorizing plan/runtime scrubbing or capacity eviction. Expiry is a real per-job state transition: it acquires `transition.lock`, re-reads the staged state, and commits through the same result-first terminal persistence path as cancellation/runner settlement. Seven-day retention is measured from terminal `finished_at`, not an older directory mtime. Admission may evict only safely removable terminal records and never active, staged, unreadable, generation-replaced, or unreclaimed retired state merely to make room; every recognized retired entry still counts toward the same hard retained-state capacity until safely removed. Cross-process create transactions serialize through an owner-identity-checked root `capacity.lock` across prune/recheck and status publication, and state inventory treats a live capacity lock as an uninstall blocker.
|
|
91
91
|
|
|
92
92
|
The runner:
|
|
93
93
|
|
package/docs/AUDIT.md
CHANGED
|
@@ -1,5 +1,35 @@
|
|
|
1
1
|
# Security and privacy audit notes
|
|
2
2
|
|
|
3
|
+
## 2026-08-27 beta.144 dependency-state recovery and CI fixture follow-up
|
|
4
|
+
|
|
5
|
+
**Exact-main provider CI separated a test-harness race from a packaged Windows recovery defect.** Ubuntu/full twice timed out waiting for a Worker WebSocket `tool_call`, first with a five-second fixture bound and again after a test-only increase to ten seconds. Source tracing showed eight fixtures triggered `currentMcpCall()`/`toolCallRequest()` before calling `waitForWsMessage()`. The HTTP fetch starts immediately, while the waiter has no historical message buffer, so a fast relay dispatch could arrive before listener registration and be lost forever. Registering the waiter first removed all eight request-before-listener sites; three local Worker integration runs passed and hosted Ubuntu/full then passed. This is test evidence, not proof that production Worker dispatch was dropping messages.
|
|
6
|
+
|
|
7
|
+
**The same provider cycle exposed a distinct production inconsistency on Windows.** A dependency-wait upstream runner was killed as part of the recovery fixture. Its runner-exit reconciliation observed a transient `permission_denied` and correctly scheduled another bounded recovery attempt, but the independently waiting downstream read the upstream status during that filesystem window. `waitForManagedJobDependencies()` converted the single read exception immediately into permanent `dependency_unavailable`, so the downstream could fail before the upstream recovery state machine converged. The two components therefore disagreed about whether the same Windows sharing/atomic-replacement condition was transient.
|
|
8
|
+
|
|
9
|
+
**Beta.144 makes dependency-state availability bounded rather than optimistic.** Only secure dependency reads classified as `permission_denied`, `identity_changed`, or generic `resource_unavailable` receive a per-dependency 45-second monotonic grace while the job remains `queued/dependency_wait`. A successful read clears the grace. Persistent failure beyond that window still becomes `dependency_unavailable`; missing state, integrity failure, witness/identity mismatch, staged state, and other invalid evidence remain immediate fail-closed conditions. Deterministic fake-clock coverage proves both one-error recovery and persistent-error expiry, while the full managed-job integration suite still proves autonomous runner recovery and dependency-failure propagation.
|
|
10
|
+
|
|
11
|
+
**Release consequence.** This repair changes packaged managed-job orchestration after beta.143 activation/acceptance, so beta.143 cannot be published from that evidence. Its acceptance record is removed, package/runtime identity advances to beta.144, and hosted tool schema generation advances to 14 because `start_job.depends_on` observable failure/recovery semantics changed. Beta.144 requires a fresh frozen full gate, candidate, install-only proof, guarded activation, activated-package OAuth canary, live observation, acceptance, and exact-head provider gates. npm registry publication remains the sole explicit current-task owner authorization boundary.
|
|
12
|
+
|
|
13
|
+
## 2026-08-27 beta.143 sleep-correlated relay continuity review
|
|
14
|
+
|
|
15
|
+
**The post-beta.142 incident record falsified the assumption that long `connection_reset` outages were primarily a VPN/TUN fault.** The same Mac retained both privacy-safe relay outage events and bounded `pmset` sleep history. A programmatic interval intersection over the fifteen recovered August 23 outages found thirteen with majority sleep overlap; aggregate outage time was 33,390 seconds and aggregate sleep overlap was 32,980 seconds (98.8%). The two no-sleep exceptions were short, roughly sixty-three and ten seconds. This establishes two distinct classes: system suspension explains the overwhelming majority of the historical long close-to-ready intervals, while short awake resets remain real transport interruptions whose upstream cause is not observable from Machine Bridge. The Karing-local-proxy live experiment was also confounded by a 224-second maintenance sleep beginning seconds after the service transition, so it does not prove that the local proxy itself is unreliable.
|
|
16
|
+
|
|
17
|
+
**The existing runtime assertion was present during a sleep it was supposed to bridge.** Power assertions showed the beta.142 `/usr/bin/caffeinate -i -w <daemon-pid>` child alive before and throughout an Idle Sleep -> maintenance-sleep sequence. The local `caffeinate(8)` contract distinguishes `-i` (prevent idle system sleep) from `-s` (prevent system sleep, valid only on AC). A bounded live probe on the same AC-powered host with `caffeinate -i -s -t 3` showed both `PreventUserIdleSystemSleep=1` and `PreventSystemSleep=1`. Beta.143 therefore adds `-s` to the one fixed assertion adapter used by authorized relay handlers, remote process sessions, and account-backed managed-job runners. It does not lengthen the existing thirty-minute inactivity grace, does not assert that `-s` is effective on battery, and does not claim to override explicit sleep or lid-close behavior.
|
|
18
|
+
|
|
19
|
+
**Diagnosis now records causality-compatible overlap instead of asking operators to correlate by eye.** `system-sleep-diagnostics.mjs` remains the only bounded `pmset` parser. In addition to the existing event-loop stall match, it intersects the most recently completed `last_disconnected_at` -> `last_ready_at` relay interval with the recent sleep intervals, merging overlaps so time is never double-counted. The owner-only result carries fixed classification plus outage duration, sleep-overlap duration/ratio, and match count; it contains no interface, endpoint, call identity, tool name, argument, result, raw power-log line, or proxy value. A majority overlap says only that system suspension dominated the observed relay outage and that a socket reset is aftermath evidence, not that every transport reset is caused by sleep.
|
|
20
|
+
|
|
21
|
+
**The user-visible `UNKNOWN / TaskGroup` remains an external-host settlement boundary, not a newly invented local timeout defect.** The observed awake interruption lasted about nine seconds. The durable release job continued without runner recovery and was later read by the same `job_id`; Worker continuity did not record a contemporaneous public request abort. Source tracing confirms pending-call reconnect preserves the original absolute deadline rather than pausing it, and the native MCP response path emits `: connected` immediately plus five-second outer keepalives. The forty-second hosted `read_job` default was previously established by live host evidence and is intentionally not shortened based on an unobservable connector error. Beta.142's recent transient recovery discovery remains the correct cross-turn fallback when a host-owned response still disappears.
|
|
22
|
+
|
|
23
|
+
**Release consequence.** Production runtime bytes, owner-visible diagnostic semantics, catalog guidance, and packaged documentation change, so the tree advances to beta.143 and hosted tool schema generation 13. Beta.142 GitHub prerelease evidence remains valid only for its exact bytes and cannot authorize beta.143. Fresh frozen verification, candidate preparation/install-only proof, guarded activation, activated-package canary, live power/relay diagnostics, acceptance, and exact-head provider gates are required before GitHub source publication. npm publication remains separately owner-authorized for the exact version.
|
|
24
|
+
|
|
25
|
+
## 2026-08-27 beta.142 recent process recovery discovery review
|
|
26
|
+
|
|
27
|
+
**The beta.141 retention reserve preserved evidence but did not guarantee that a later host turn could discover it.** Live owner inventory reproduced the boundary at the hard 512-state cap: 496 ordinary durable terminal records and the full 16 recent transient terminal reserve were retained. The bounded 50-record `list_jobs.jobs` window correctly remained durable-first, so recent one-step process helpers were absent from that primary list. Reading one of those omitted helpers by its already-known `job_id` still succeeded and returned the terminal result. The remaining defect was therefore discovery of retained recovery identity after a host/tool boundary that lost the prior tool response, not durable execution, terminal persistence, or retention.
|
|
28
|
+
|
|
29
|
+
**Beta.142 adds a second bounded projection instead of weakening the durable-first primary ordering.** `list_jobs.jobs` remains capped at 50 and keeps unreadable, active, staged, and durable terminal state ahead of transient helper history. `recent_process_recovery` independently returns at most 16 recent transient terminal jobs that passed the existing authority filter, remain inside the existing thirty-minute recovery grace, and were omitted from the primary window. The projection is built from the same validated persisted status and existing public job projection; it does not return step results or introduce another private state store. Saturated integration coverage proves that all sixteen retained recent helpers remain discoverable outside a primary window containing no transient entries, while the internal retention class remains absent from public job objects.
|
|
30
|
+
|
|
31
|
+
**The security and lifecycle boundary is intentionally unchanged.** The new array is managed-job recovery inventory, not an MCP terminal-result store, response replay identifier, conversation/session binding, or host-delivery acknowledgement. It cannot initiate a ChatGPT turn or prove that an external host rendered the previous terminal response. Authority filtering runs before either inventory projection, the retained-state cap remains 512, the primary response window remains 50, the recent transient reserve remains sixteen results for thirty minutes, and helper history still loses retention priority before ordinary durable terminal history outside that reserve. Hosted schema generation advances to 12 because the public `list_jobs` result semantics changed.
|
|
32
|
+
|
|
3
33
|
## 2026-08-26 beta.141 retained-result recovery and Windows main-CI follow-up
|
|
4
34
|
|
|
5
35
|
**The repeated user-visible interruption investigation exposed a separate recovery-evidence defect even while the underlying durable jobs remained healthy.** Owner diagnostics showed the retained managed-job store at 512 entries with 510 ordinary durable terminal records and only two transient terminal records. Two short read-only `exec_command` helpers were accepted and returned job IDs, but their follow-up `read_job` calls a few seconds later returned typed `not_found`. The cause is deterministic in the previous retention ordering: every `transient_process` terminal had lower eviction priority than every ordinary durable terminal, independent of completion age. At saturation, each new helper could therefore evict the previous helper result before the host consumed it. That does not explain or make observable the external ChatGPT turn-termination decision, but it removes recovery evidence precisely when a host/tool interruption makes that evidence most valuable.
|
package/docs/LOGGING.md
CHANGED
|
@@ -67,7 +67,7 @@ Brief network interruptions are expected on laptop network changes, Worker deplo
|
|
|
67
67
|
- failure to receive `hello_ack` within the handshake deadline, or `ready_ack` within the independent end-to-end readiness deadline, terminates the candidate socket and retries;
|
|
68
68
|
- authenticated transports request protocol-level WebSocket Ping every five seconds. A sender callback starts the full ten-second Pong deadline only after actual local dispatch. If that deadline expires on a fully ready WSS, the relay runtime records `relay.transport.suspect`, opens one fifteen-second application-confirmation window, sends a JSON heartbeat, and prewarms HTTPS in standby instead of killing the socket immediately. A later protocol Pong or explicit JSON application `pong` records `relay.transport.recovered` and preserves WSS; ordinary inbound tool/control traffic remains receive-side evidence and cannot clear transport suspicion. Only confirmation expiry records `relay.transport.confirmation_failed` and closes as `relay_transport_timeout`. A thirty-second local Ping-dispatch failure remains the distinct `relay.transport.send_timeout` / `relay_transport_send_timeout` path even if unrelated inbound traffic continues. Send-completion callbacks are exact-WebSocket-generation fenced, so a callback from a superseded socket cannot mutate current relay diagnostics or confirmation state. These anomaly events contain only bounded timing/state fields;
|
|
69
69
|
- a separate periodic twenty-five-second application heartbeat refreshes Worker daemon activity and retains a seventy-five-second application-silence timeout; it begins only after end-to-end readiness, so authenticated probing cannot send a message type that the Worker probing state does not accept. Protocol-level Pong therefore cannot mask a Worker application path that has stopped replying. The Worker queues the heartbeat's JSON `pong` before Durable Object alarm inspection or mutation, then performs one explicit coalesced schedule, so storage latency is not allowed to sit ahead of application-liveness acknowledgement;
|
|
70
|
-
- a late local transport-watchdog tick is classified as `runtime.event_loop.stall`, sends a fresh transport probe, and defers disconnect for a bounded recovery interval instead of being mislabeled as immediate remote failure; a macOS sleep/wake interval may legitimately produce this warning without a daemon fault. The heartbeat snapshot retains the lag of the last actual stall separately from the rolling maximum. Owner `diagnose_runtime` can compare that end time and duration with a bounded fixed `pmset` sleep-history projection and reports `matched_system_sleep` only when both dimensions agree within the fixed tolerance; an unmatched stall remains unclassified rather than being labeled synchronous JavaScript blockage by elimination. An authorized, schema-valid remote tool call enters the bounded macOS idle-sleep guard only after policy/account/operation authorization and argument validation succeed; the shared assertion remains active for the full handler lifetime, remote process sessions extend it through child settlement, and the fixed thirty-minute rolling inactivity grace starts only after the last owned daemon-side activity settles, with each new authorized activity cancelling pending release and restarting the full grace after settlement. The guard snapshot records coarse activity/grace/release times and release reason without tool arguments or conversation identity. If the fixed macOS assertion child cannot be established, `runtime.idle_sleep_guard.unavailable` records only a coarse `error_class`, never argv, paths, PID, tool name, tool content, session identity, or job identity. Remote account managed-job runners use the same fixed assertion primitive but own it themselves only after runner-claim confirmation and persisted account ownership validation; local managed jobs do not acquire this remote-continuity assertion. The remote runner fallback stderr diagnostic is the fixed text `managed job idle-sleep assertion unavailable` plus a sanitized coarse `error_class`, without job name/id, workspace path, argv, environment, or captured output.
|
|
70
|
+
- a late local transport-watchdog tick is classified as `runtime.event_loop.stall`, sends a fresh transport probe, and defers disconnect for a bounded recovery interval instead of being mislabeled as immediate remote failure; a macOS sleep/wake interval may legitimately produce this warning without a daemon fault. The heartbeat snapshot retains the lag of the last actual stall separately from the rolling maximum. Owner `diagnose_runtime` can compare that end time and duration with a bounded fixed `pmset` sleep-history projection and reports `matched_system_sleep` only when both dimensions agree within the fixed tolerance; an unmatched stall remains unclassified rather than being labeled synchronous JavaScript blockage by elimination. The same diagnostic also intersects the most recent completed relay disconnect interval with that bounded sleep history and reports only outage/overlap timing plus a fixed classification; `majority_system_sleep_overlap` means host suspension dominated the observed outage, so a retained `connection_reset` is aftermath evidence rather than sufficient independent-network evidence. An authorized, schema-valid remote tool call enters the bounded macOS idle-sleep guard only after policy/account/operation authorization and argument validation succeed; the shared `/usr/bin/caffeinate -i -s -w <owner-pid>` assertion remains active for the full handler lifetime, remote process sessions extend it through child settlement, and the fixed thirty-minute rolling inactivity grace starts only after the last owned daemon-side activity settles, with each new authorized activity cancelling pending release and restarting the full grace after settlement. The `-s` assertion is effective only on AC power; diagnostics report only that the fixed child requests it. The guard snapshot records coarse activity/grace/release times and release reason without tool arguments or conversation identity. If the fixed macOS assertion child cannot be established, `runtime.idle_sleep_guard.unavailable` records only a coarse `error_class`, never argv, paths, PID, tool name, tool content, session identity, or job identity. Remote account managed-job runners use the same fixed assertion primitive but own it themselves only after runner-claim confirmation and persisted account ownership validation; local managed jobs do not acquire this remote-continuity assertion. The remote runner fallback stderr diagnostic is the fixed text `managed job idle-sleep assertion unavailable` plus a sanitized coarse `error_class`, without job name/id, workspace path, argv, environment, or captured output.
|
|
71
71
|
|
|
72
72
|
A WebSocket close code such as `1006` means the transport ended without a normal close handshake, but it does not identify who initiated termination: Machine Bridge's own liveness recovery calls `terminate()` when a transport/send timeout is confirmed, and that local hard close can surface as 1006. Diagnose the cause from `last_close_category`, transport-confirmation/send-timeout evidence, and retained network milestones rather than treating 1006 itself as proof of a remote/network-initiated close. If it recovers inside ten seconds, the warning-level service log is intentionally silent and the authenticated `daemon.relay_transport` snapshot is the post-event evidence surface. It is useful for debug diagnosis but not useful as the default user message. It is not evidence that the daemon process restarted. Worker `daemon_transport_error` / `daemon_liveness_timeout` messages and their 1012 close frames are likewise retryable connection conditions, not upgrade instructions. Only an unknown/incompatible Worker error, authentication failure, or identity/version mismatch may produce the fatal protocol/configuration log and daemon exit. Default logs therefore describe the affected layer, duration, classification, and recovery behavior rather than printing raw close envelopes.
|
|
73
73
|
|
|
@@ -153,7 +153,7 @@ Each managed job has owner-only runner diagnostic logs. Child-step output is ret
|
|
|
153
153
|
|
|
154
154
|
`network_route` describes only Machine Bridge's application-level proxy decision. `system-network-stack` does **not** mean a direct physical path: an operating-system VPN, TUN, packet tunnel, DNS interceptor, or endpoint-security product may still carry the connection. `network_route_scope` therefore remains `application-proxy-selection-only`.
|
|
155
155
|
|
|
156
|
-
During an outage, remote-owner `diagnose_runtime.runtime.relay` and local stdio `server_info.runtime.relay` expose bounded live fields: outage count/start/duration, attempts, last close category/code, coarse transport error class plus strict allowlisted `last_transport_error_reason`, last disconnect/ready time, prior ready duration, prior ready inbound-silence duration, the thirty-second WSS connect budget, bounded `last_connect_milestones_ms`, transport-probe queue/dispatch/Pong state, second-stage transport-confirmation timing/recovery state, bounded sender backlog bytes, HTTPS fallback active/warming state, `https_fallback_last_takeover_ms`, and next retry timing. `https_fallback_last_takeover_ms` retains the bounded WSS-close-to-verified-HTTPS-ready interval even after WSS later reclaims primary ownership, so a longer WebSocket outage is not misreported as the same duration of whole-bridge unavailability when HTTPS recovered earlier. The fallback status separates `last_success_at` (any successful signed HTTP exchange, including standby) from `last_ready_at` (only verified ready ownership), so a successful standby poll can no longer be misread as a completed failover. Connect milestones are relative durations only and never include hostnames, addresses, DNS answers, certificates, proxy endpoints, or credentials. The current-attempt milestones and separate `last_failed_connect_*` fields are retained independently so a successful retry cannot erase the immediately preceding failed DNS/TCP/TLS/upgrade evidence; `last_transport_error_ready` and `last_transport_error_authenticated` state whether the retained transport error occurred after channel authentication/readiness. Worker-side `daemon.websocket.closed` records only a bounded close code plus `was_clean`; raw peer close reasons are deliberately omitted. An already-dispatched Worker call detached from a failed channel is retained only until the smaller of reconnect grace and that call's original remaining absolute deadline. If it settles without rebinding, the public bounded error distinguishes `original call deadline expired during reconnect` from a true full `reconnect grace expired`; neither diagnostic includes tool arguments, paths, account identity, endpoint data, or result content, and the distinction does not extend the hosted deadline. Resource-coordinator snapshot contention is likewise classified separately from execution failure: when the bounded diagnostic cannot acquire a transaction/staging lock, it reports retryable `unavailable`, `reason=coordinator_busy`, and `snapshot_available=false`. This says the diagnostic snapshot was unavailable under contention; it does not relax admission policy or claim that the underlying host pressure is Green. `outage_duration_ms` is the close-to-ready recovery interval; `last_ready_inbound_silence_ms` is the pre-close interval since the preceding ready transport last proved inbound activity. After recovery, authenticated remote `server_info.daemon.relay_transport` retains the bounded preceding episode supplied during the current connection handshake, including `previous_ready_inbound_silence_ms` and brief interruptions below the default warning threshold. Promotion to a ready socket sets `outage_active=false`, extends the outage duration through actual readiness, preserves the preceding healthy-ready duration and inbound-silence evidence across failed candidates, canonicalizes timestamps, and accepts only enumerated coarse operational error classes; it does not claim that the recovered connection remains in outage. The `local_authority_revocation_retry` category is deliberately not diagnosed as a network failure: a sustained warning directs the operator to local authority, process-session, and managed-job state while the retained Worker revocation retries on reconnection; ordinary transport categories retain network/Worker troubleshooting guidance. On macOS, `diagnose_runtime` may also return a coarse default-route class, `operating_system_interception` boolean,
|
|
156
|
+
During an outage, remote-owner `diagnose_runtime.runtime.relay` and local stdio `server_info.runtime.relay` expose bounded live fields: outage count/start/duration, attempts, last close category/code, coarse transport error class plus strict allowlisted `last_transport_error_reason`, last disconnect/ready time, prior ready duration, prior ready inbound-silence duration, the thirty-second WSS connect budget, bounded `last_connect_milestones_ms`, transport-probe queue/dispatch/Pong state, second-stage transport-confirmation timing/recovery state, bounded sender backlog bytes, HTTPS fallback active/warming state, `https_fallback_last_takeover_ms`, and next retry timing. `https_fallback_last_takeover_ms` retains the bounded WSS-close-to-verified-HTTPS-ready interval even after WSS later reclaims primary ownership, so a longer WebSocket outage is not misreported as the same duration of whole-bridge unavailability when HTTPS recovered earlier. The fallback status separates `last_success_at` (any successful signed HTTP exchange, including standby) from `last_ready_at` (only verified ready ownership), so a successful standby poll can no longer be misread as a completed failover. Connect milestones are relative durations only and never include hostnames, addresses, DNS answers, certificates, proxy endpoints, or credentials. The current-attempt milestones and separate `last_failed_connect_*` fields are retained independently so a successful retry cannot erase the immediately preceding failed DNS/TCP/TLS/upgrade evidence; `last_transport_error_ready` and `last_transport_error_authenticated` state whether the retained transport error occurred after channel authentication/readiness. Worker-side `daemon.websocket.closed` records only a bounded close code plus `was_clean`; raw peer close reasons are deliberately omitted. An already-dispatched Worker call detached from a failed channel is retained only until the smaller of reconnect grace and that call's original remaining absolute deadline. If it settles without rebinding, the public bounded error distinguishes `original call deadline expired during reconnect` from a true full `reconnect grace expired`; neither diagnostic includes tool arguments, paths, account identity, endpoint data, or result content, and the distinction does not extend the hosted deadline. Resource-coordinator snapshot contention is likewise classified separately from execution failure: when the bounded diagnostic cannot acquire a transaction/staging lock, it reports retryable `unavailable`, `reason=coordinator_busy`, and `snapshot_available=false`. This says the diagnostic snapshot was unavailable under contention; it does not relax admission policy or claim that the underlying host pressure is Green. `outage_duration_ms` is the close-to-ready recovery interval; `last_ready_inbound_silence_ms` is the pre-close interval since the preceding ready transport last proved inbound activity. After recovery, authenticated remote `server_info.daemon.relay_transport` retains the bounded preceding episode supplied during the current connection handshake, including `previous_ready_inbound_silence_ms` and brief interruptions below the default warning threshold. Promotion to a ready socket sets `outage_active=false`, extends the outage duration through actual readiness, preserves the preceding healthy-ready duration and inbound-silence evidence across failed candidates, canonicalizes timestamps, and accepts only enumerated coarse operational error classes; it does not claim that the recovered connection remains in outage. The `local_authority_revocation_retry` category is deliberately not diagnosed as a network failure: a sustained warning directs the operator to local authority, process-session, and managed-job state while the retained Worker revocation retries on reconnection; ordinary transport categories retain network/Worker troubleshooting guidance. On macOS, `diagnose_runtime` may also return a coarse default-route class, `operating_system_interception` boolean, privacy-bounded `runtime.idle_sleep_guard` state (`supported`, `enabled`, `active`, `grace_ms`, `requests_system_sleep_prevention_on_ac`, `last_error_class`), bounded `runtime.system_sleep`, and `runtime.relay_outage_analysis`. That diagnostic is returned on demand and is not promoted to default logs; interface names, IP addresses, DNS answers, proxy endpoints/credentials, Worker endpoints, tool arguments, and results remain absent. `relay.outage.active` and `relay.outage.recovered` carry the existing safe relay fields.
|
|
157
157
|
|
|
158
158
|
Schema 4 is strict NDJSON. Before daemon startup, both active log files are opened as owner-only regular single-link files. A schema change clears both only after validation and commits the marker only after the transition succeeds. A symlink, multiple-hard-link inode, permission error, or marker-write failure blocks startup rather than mixing formats or repeatedly erasing evidence.
|
|
159
159
|
|
package/docs/MANAGED_JOBS.md
CHANGED
|
@@ -13,7 +13,7 @@ can therefore stop after step 2 or 3 and leave local or remote temporary state b
|
|
|
13
13
|
|
|
14
14
|
Machine Bridge managed jobs reduce this failure mode by accepting the complete execution and cleanup plan in one call. After acceptance, an independent local runner owns the lifecycle. It does not depend on the MCP socket remaining connected.
|
|
15
15
|
|
|
16
|
-
Recovery inventory is deliberately ordered by recoverability rather than recency alone. `list_jobs` returns at most 50 records,
|
|
16
|
+
Recovery inventory is deliberately ordered by recoverability rather than recency alone. `list_jobs.jobs` returns at most 50 primary records, with unreadable state, active jobs, staged plans, and durable terminal history ahead of transient one-step helper history so a burst of short helpers cannot push an older long-running or pre-spawn-waiting job out of that bounded primary window. A separate `recent_process_recovery` array returns at most 16 authority-visible public handles for recent one-step terminal jobs that are still inside the thirty-minute recovery reserve but omitted from `jobs`; callers can recover the `job_id` there and continue with `read_job`. It contains no step output or internal retention class and is recovery inventory, not a polling or MCP replay/session surface. Owner/local listings also include only aggregate recent creation/churn counts; the internal `transient_process` retention class remains hidden per job.
|
|
17
17
|
|
|
18
18
|
Remote one-step `exec_command`, `run_process`, and `run_local_command` requests still become durable managed jobs before execution. The hosted tool call then spends only a short two-second initial-settlement window checking whether that helper already reached terminal state. A short helper can therefore return its managed-job result in the original tool response with `follow_up_read_required=false`, eliminating the usual second `read_job` event. A helper that remains active keeps the same durable `job_id` and recovery envelope with `follow_up_read_required=true`; its execution timeout, pre-spawn resource-admission allowance, and later recovery semantics are unchanged.
|
|
19
19
|
|
|
@@ -127,7 +127,7 @@ A job terminal transition is not a single best-effort write. The runner first at
|
|
|
127
127
|
|
|
128
128
|
If the runner writes a valid terminal result but exits before terminal status is committed, the manager reconstructs status from that result before considering recovery. If a dead runner instead leaves a present but non-terminal, cross-job, or otherwise invalid result object, recovery fails closed and retains the state for inspection rather than treating that corruption as an absent result and overwriting it with a cleanup retry. Lifecycle status is likewise an explicit enum, not a `not-active => terminal` inference: an unknown on-disk status is retained as an integrity failure, is not scrubbed or capacity-evicted as a completed job, blocks destructive removal through active-job inventory, and cannot be acknowledged as already finished by cancellation/revocation. While the runner is still alive in the narrow result-first/status-second settlement window, `read_job` also projects a coherent terminal status from that already-durable result in memory rather than returning an active outer status beside a terminal nested result; cleanup remains conservatively reported as pending until the persisted status catches up. This read-side projection never writes runner-owned state. These rules prevent a completed finally sequence from being replayed or observably mixed merely because the status write has not happened yet. If no valid terminal result exists, recovery may still repeat finally work, so finally steps must remain idempotent.
|
|
129
129
|
|
|
130
|
-
Unexecuted staged plans can contain stdin, environment values, and temporary scripts. They expire after 24 hours, but expiry is itself a per-job state transition: pruning must acquire the job transition lock, re-read the current status, and only then commit the non-executing terminal record through the same result-first persistence boundary. A concurrent cancel/other transition therefore wins or finishes before expiry rather than racing it. Terminal retention begins from that terminal record's `finished_at`, not from the older staged directory age, so a long-abandoned draft is not immediately deleted in the same pass that first expires it. When seven-day retention or capacity pressure eventually retires a complete job directory, removal is also generation-bound: the inspected directory is atomically quarantined under an internal `retired_job_*` generation name before recursive deletion. That internal namespace is deliberately outside the public `job_...` ID grammar, so normal list/read/lock paths cannot reinterpret cleanup state as a job. A crash after rename leaves a recognizable cleanup record. Later pruning reclaims it only when the encoded filesystem generation still matches. Verification failure never renames quarantine back onto the public job pathname; the retired evidence stays isolated for a later safe retry. Any malformed reserved `retired_job_*` name, generation mismatch, wrong type, or unreadable retired entry remains a privacy-bounded destructive-state blocker without exposing its internal filename/device/inode through ordinary job diagnostics. Every recognized retired entry still counts toward the 512-state retained-state capacity until safely removed, so namespace separation cannot become a capacity bypass. The public `job_...` namespace is reserved just as strictly: a matching name with the wrong filesystem type is retained as `unreadable`, counts toward the same capacity, blocks destructive inventory, and is shown only to owner/local diagnostics. New deterministic job admission securely inspects an existing target before any capacity eviction, so a dangling link or other invalid target cannot consume retained terminal history before the request fails. Ordinary completed-job metadata may remain under the separate seven-day retention policy. The retained-state hard cap is 512. `list_jobs` deliberately returns at most 50
|
|
130
|
+
Unexecuted staged plans can contain stdin, environment values, and temporary scripts. They expire after 24 hours, but expiry is itself a per-job state transition: pruning must acquire the job transition lock, re-read the current status, and only then commit the non-executing terminal record through the same result-first persistence boundary. A concurrent cancel/other transition therefore wins or finishes before expiry rather than racing it. Terminal retention begins from that terminal record's `finished_at`, not from the older staged directory age, so a long-abandoned draft is not immediately deleted in the same pass that first expires it. When seven-day retention or capacity pressure eventually retires a complete job directory, removal is also generation-bound: the inspected directory is atomically quarantined under an internal `retired_job_*` generation name before recursive deletion. That internal namespace is deliberately outside the public `job_...` ID grammar, so normal list/read/lock paths cannot reinterpret cleanup state as a job. A crash after rename leaves a recognizable cleanup record. Later pruning reclaims it only when the encoded filesystem generation still matches. Verification failure never renames quarantine back onto the public job pathname; the retired evidence stays isolated for a later safe retry. Any malformed reserved `retired_job_*` name, generation mismatch, wrong type, or unreadable retired entry remains a privacy-bounded destructive-state blocker without exposing its internal filename/device/inode through ordinary job diagnostics. Every recognized retired entry still counts toward the 512-state retained-state capacity until safely removed, so namespace separation cannot become a capacity bypass. The public `job_...` namespace is reserved just as strictly: a matching name with the wrong filesystem type is retained as `unreadable`, counts toward the same capacity, blocks destructive inventory, and is shown only to owner/local diagnostics. New deterministic job admission securely inspects an existing target before any capacity eviction, so a dangling link or other invalid target cannot consume retained terminal history before the request fails. Ordinary completed-job metadata may remain under the separate seven-day retention policy. The retained-state hard cap is 512. `list_jobs.jobs` deliberately returns at most 50 primary records per response so deeper recovery history does not inflate the ordinary MCP inventory window. One-step remote process carriers created by `exec_command`, `run_process`, and `run_local_command` are internally marked with the low-cardinality `transient_process` retention class; that marker is not part of the public job projection and contains no argv, path, output, or credential data. Under capacity pressure, completed transient process records are reclaimed before explicit managed-job terminal history whenever such transient records are available, except for the fixed newest-16/thirty-minute immediate recovery reserve. A retained recent process terminal that the durable-first primary `jobs` window omits may appear in `recent_process_recovery`, which is independently capped at 16 authority-visible public job handles and carries no step output or retention metadata. This prevents diagnostic/helper-command churn from preferentially destroying the recovery result of a long explicit managed job while keeping disk/privacy state and both inventory windows bounded. Active or staged plans that declare `depends_on` additionally pin those referenced retained job records against time- or capacity-based pruning until the dependency-bearing plan is terminal; if dependency protection cannot be read safely, capacity pruning fails closed rather than guessing that no dependency exists. A valid `job_id` whose directory has expired or been capacity-retired returns typed `not_found` rather than a generic execution failure. That absence means only that the retained record is unavailable; callers must not infer that the underlying operation never executed or blindly replay its side effects. A minimal-environment plan also launches its detached runner with a minimal control environment. Full parent-environment inheritance occurs only when the accepted plan explicitly captured that policy.
|
|
131
131
|
|
|
132
132
|
## Job-scoped temporary files
|
|
133
133
|
|
|
@@ -225,7 +225,7 @@ Use top-level `depends_on` when one managed job must wait for one or more earlie
|
|
|
225
225
|
}
|
|
226
226
|
```
|
|
227
227
|
|
|
228
|
-
At acceptance, each dependency is bound to its current durable job identity (`job_id`, plan hash, and creation generation). Staged dependencies and dependencies that have already failed are rejected before the new job is accepted. While any accepted dependency remains active, the dependent job stays `queued` with `current_phase=dependency_wait`; `dependency_total` and `dependency_pending_count` report progress, and the dependent job has not spawned its main child, entered process resource admission, or materialized private registered-resource/temporary-file execution copies. As upstream jobs settle, a hosted `read_job` long-poll wakes when the pending count changes. All dependencies succeeding releases the normal main-step sequence and materializes execution inputs immediately before they can be needed. An upstream job that later ends unsuccessfully makes the dependent job terminal with `result.error_class=dependency_failed` and bounded `dependency_failure` evidence instead of waiting minutes or hours for an impossible artifact. If that failure still requires declared `finally_steps`, their resource/temporary-file inputs are materialized only when the cleanup phase begins.
|
|
228
|
+
At acceptance, each dependency is bound to its current durable job identity (`job_id`, plan hash, and creation generation). Staged dependencies and dependencies that have already failed are rejected before the new job is accepted. While any accepted dependency remains active, the dependent job stays `queued` with `current_phase=dependency_wait`; `dependency_total` and `dependency_pending_count` report progress, and the dependent job has not spawned its main child, entered process resource admission, or materialized private registered-resource/temporary-file execution copies. As upstream jobs settle, a hosted `read_job` long-poll wakes when the pending count changes. A transient dependency-state read failure classified as `permission_denied`, `identity_changed`, or generic `resource_unavailable` keeps the dependent in `dependency_wait` for a fixed 45-second recovery grace instead of converting one Windows sharing/atomic-replacement race into a permanent dependency failure. A successful read clears that per-dependency grace immediately. Persistent unavailability past the grace still fails closed as `dependency_unavailable`; `not_found`, integrity failure, witness/identity mismatch, staged state, and other invalid dependency evidence remain immediate failures. All dependencies succeeding releases the normal main-step sequence and materializes execution inputs immediately before they can be needed. An upstream job that later ends unsuccessfully makes the dependent job terminal with `result.error_class=dependency_failed` and bounded `dependency_failure` evidence instead of waiting minutes or hours for an impossible artifact. If that failure still requires declared `finally_steps`, their resource/temporary-file inputs are materialized only when the cleanup phase begins.
|
|
229
229
|
|
|
230
230
|
Terminal persistence is result-first: the durable `result.json` is written before `status.json` is changed to the matching terminal state. Dependency polling therefore treats a valid terminal result for the same job as a read-only terminal projection during that narrow publication window instead of waiting for an unrelated manager read to repair the status file. It does not rewrite the upstream record. Once the dependent itself becomes terminal, `dependency_pending_count` is zero because the dependent is no longer waiting; `dependency_failure` identifies the upstream terminal cause when the job failed.
|
|
231
231
|
|
|
@@ -310,7 +310,7 @@ Never place a secret directly in `argv`, `env`, `stdin`, a temporary file's `con
|
|
|
310
310
|
|
|
311
311
|
Per-workspace jobs are stored below the owner-only profile directory. Active jobs retain an owner-only plan for crash recovery. Plan, status, result, runner identity, and lock updates use flushed atomic replacement or complete-before-visible exclusive claims. Transition/recovery locks contain ownership tokens and process start time and are removed only when their file snapshot still matches. After a terminal status is committed, the full plan is deleted, including argv, stdin, embedded temporary-file content, and resource source paths.
|
|
312
312
|
|
|
313
|
-
Retained public job data contains bounded status and redacted results. The hard capacity is 512 retained-state slots across ordinary job directories plus any recognized internal retired-cleanup entries that have not yet been safely removed; `list_jobs` still returns at most 50
|
|
313
|
+
Retained public job data contains bounded status and redacted results. The hard capacity is 512 retained-state slots across ordinary job directories plus any recognized internal retired-cleanup entries that have not yet been safely removed; `list_jobs.jobs` still returns at most 50 primary records per call, while `recent_process_recovery` may add at most 16 recent authority-visible process recovery handles that were omitted from that primary window. Terminal jobs are normally retained for up to seven days from their persisted `finished_at` settlement time, but when capacity is full the oldest safely removable terminal records are evicted to reserve a slot for a new job. Staged drafts expire after 24 hours, dependency-referenced records are protected while an active/staged dependent plan still needs them, and active, staged, unreadable, or abnormal retired state is never evicted merely to make room; if all 512 slots are occupied, new job creation returns a retryable `limit_exceeded` error whose owner-only details include coarse `retained_state`, `retired_state`, and `retired_unreadable` counts. `list_jobs.retained` remains the number of visible ordinary jobs even when only 50 are returned. That bounded inventory is recovery-first: unreadable, active, and staged state stays first, durable terminal managed-job results precede transient one-step process terminals, and only then does helper history fill the remaining response window. Owner/local responses additionally include the coarse capacity summary plus `durable_terminal` and `transient_terminal` counts so a full 512-state store can be distinguished from helper churn without exposing job identities, paths, arguments, or output. Delegated non-owner responses omit the global capacity summary. These counts improve recovery visibility only; Machine Bridge still cannot observe whether an external host consumed a terminal result or rendered a final assistant response. Private runtime copies are removed after the finally phase. Runner stdout/stderr log files contain only runner-level diagnostics; step output is not written to those operational logs.
|
|
314
314
|
|
|
315
315
|
The detached runner records a structured owner record containing PID and process start time. Recovery rejects a reused PID instead of treating an unrelated process as the active runner. Numeric-only runner records are invalid. Initial runner publication uses provisional then committed atomic generations; a claim reader is explicitly coupled to that publication protocol and may therefore retry a transient `MBM_IDENTITY_CHANGED` observation for four 1 ms attempts before failing. Each retry repeats the complete secure read and identity validation; this exception does not apply to generic or destructive file reads. Recovery-lock handoff preserves a random ownership token, and the runner removes only a lock whose PID, token, and file snapshot still match. A recovery runner does not gain authority to write terminal evidence merely by confirming its runner claim: it must first complete recovery-lock handoff. Failure or ambiguity in that bootstrap phase leaves the prior `interrupted` status and plan intact for a later safe retry instead of publishing `recovery_failed` and scrubbing recovery material. The handoff has a 30-second monotonic ownership-settlement budget. Timeout and cancellation terminate the process group/tree, retain a referenced forced-escalation timer, and clean descendants that ignore graceful termination before the runner exits.
|
|
316
316
|
|
package/docs/OPERATIONS.md
CHANGED
|
@@ -10,9 +10,9 @@ machine-mcp service status
|
|
|
10
10
|
|
|
11
11
|
Routine remote checks should use authenticated `server_info` with `detail: "summary"`; request the default/full projection only when the caller's authority permits and exact effective-tool, OAuth/account, or detailed owner observability is actually needed. Non-owner full responses intentionally retain hidden markers/counts instead of cross-principal activity, resource aliases, stable device-key identity, or daemon-only tool names. Remote `diagnose_runtime` is owner-only because its fixed probes expose machine-wide control-plane activity; narrower roles use `server_info`/`project_overview` for authority-scoped readiness and workspace state. `status` prints redacted profile state and verifies the deployed Worker version. Resource source paths remain redacted. `doctor` checks Node.js, the package-installed Wrangler binary, Cloudflare login, Worker health, the configured policy, the automatic-without-per-operation-prompts authorization model, and the same fixed local filesystem/process/shell/job-storage/resource probes exposed to the remote owner by `diagnose_runtime`. It constructs an isolated local runtime: `diagnosticScope.running_service_process_inspected=false` and `remote_relay_inspected=false` are deliberate, so a green doctor result is not evidence that the launchd/systemd/Scheduled Task daemon retained its Worker WebSocket. Inspect authenticated `server_info.daemon.relay_transport` for the running service relay. Authenticated `server_info.authorization.execution_model` reports the authority contract and identifies whether the account has daemon-OS-user ambient authority. Public `/healthz` output contains only server identity and version; daemon details require an authenticated `server_info` call.
|
|
12
12
|
|
|
13
|
-
For interruption analysis, prefer one owner `diagnose_runtime` call over a chain of inventory probes. It now includes `runtime.managed_jobs.recent_activity`, `runtime.security_audit.recent_activity`, bounded `runtime.resource_admission.waiters.diagnostics`, `runtime.system_sleep`, and `runtime.
|
|
13
|
+
For interruption analysis, prefer one owner `diagnose_runtime` call over a chain of inventory probes. It now includes `runtime.managed_jobs.recent_activity`, `runtime.security_audit.recent_activity`, bounded `runtime.resource_admission.waiters.diagnostics`, `runtime.system_sleep`, `runtime.event_loop_pause_analysis`, and `runtime.relay_outage_analysis`. The audit aggregate contains only counts, bounded tool names, failure totals, and calls-per-minute density derived from the existing content-free hash-chained audit log; it contains no tool arguments or results. Its `coverage=daemon_reached_relay_tool_calls_only` and `host_side_events_observable=false` fields make the evidence boundary explicit: host-only discovery/control-plane/final-delivery events are not counted. The waiter projection reports only resource-request shape and the current admission reason. On macOS with shell-capable owner diagnostics, the fixed power probe reduces `pmset` history to a small list of sleep start/end/duration/reason classes; it never returns raw power-log lines. `event_loop_pause_analysis.classification=matched_system_sleep` requires both the recorded runtime-stall end time and duration to match one of those bounded operating-system sleep intervals within a fixed tolerance. `relay_outage_analysis` separately compares the most recent completed `last_disconnected_at` -> `last_ready_at` interval with the same bounded sleep history and reports exact overlap duration/ratio. `majority_system_sleep_overlap` means at least half of that observed relay outage occurred while macOS was suspended, so a retained `connection_reset` is transport aftermath rather than sufficient evidence of a separate network root cause. `no_matching_recent_system_sleep` leaves an awake reset unassigned rather than guessing a VPN, edge, Worker, or host cause. ChatGPT host-turn termination/final-message receipt remains explicitly unobservable.
|
|
14
14
|
|
|
15
|
-
`runtime.idle_sleep_guard` also carries coarse `last_activity_started_at`, `last_activity_ended_at`, `grace_release_due_at`, `last_release_at`, `last_release_reason`,
|
|
15
|
+
`runtime.idle_sleep_guard` also carries coarse `last_activity_started_at`, `last_activity_ended_at`, `grace_release_due_at`, `last_release_at`, `last_release_reason`, active-activity count, and `requests_system_sleep_prevention_on_ac`. These fields are diagnostic ownership evidence only. On macOS the fixed child now runs `/usr/bin/caffeinate -i -s -w <owner-pid>`: `-i` retains the existing Idle Sleep prevention, while Apple's `caffeinate` contract makes `-s` a stronger system-sleep assertion only while the Mac is on AC power. On battery, do not interpret the `-s` request as a guarantee against system sleep. Explicit sleep, lid-close policy, power loss, and operating-system behavior outside those assertion contracts remain external boundaries. A long-running managed job has its separate runner-owned assertion, while ordinary authorized relay activity retains the existing thirty-minute inactivity grace. Do not extend that grace merely because a later host failure is unobservable: a permanent or multi-hour assertion after every remote call would trade battery behavior for an unsupported host-lifecycle assumption.
|
|
16
16
|
|
|
17
17
|
Remote one-step process carriers also reduce event amplification at source. After durable acceptance, the original `exec_command`, `run_process`, or `run_local_command` response waits up to `server_info.tool_delivery.remote_process_initial_settlement_wait_ms` for a short helper to settle. If it does, terminal status/result are returned immediately with `follow_up_read_required=false`; otherwise the response retains the durable recovery envelope and `follow_up_read_required=true`. This two-second response coalescing does not reduce the independently bounded 600-second child execution budget, the thirty-minute pre-spawn admission allowance, or the six-hour managed-job step ceiling.
|
|
18
18
|
|
|
@@ -95,7 +95,7 @@ After the host path recovers, compare authenticated `server_info`, `machine-mcp
|
|
|
95
95
|
|
|
96
96
|
A reconnect warning proves a transport interruption, not a daemon crash. Compare daemon PID and process start time with `connected_at`, `last_seen_at`, `daemon.relay_transport.last_close_category`, `last_close_code`, `outage_count`, `outage_attempts`, `previous_ready_inbound_silence_ms`, `last_connect_milestones_ms`, and the coarse network-route class. Also distinguish a planned restart from an accidental outage: a current-generation daemon sends `daemon_draining` before relay close, affected calls receive `reason=daemon_planned_drain`, and owner/full `server_info.worker.continuity_evidence` durably retains the planned-drain count/time plus bounded socket-disconnect and client-cancellation observations across Worker isolate replacement. `worker.observability.continuity` remains isolate-local and may reset; use the durable summary for post-incident correlation rather than relying on a later close-1006 inference. For `read_job`, recover with the same returned `job_id`; do not resubmit the job's underlying mutation. `last_connect_milestones_ms` contains only bounded relative timings for the most recent connection attempt phases such as DNS resolution, TCP connect, TLS establishment, HTTP rejection, and WebSocket open; `last_failed_connect_stage`, `last_failed_connect_duration_ms`, `last_failed_connect_milestones_ms`, and `last_failed_connect_http_status` retain the most recent failed attempt even after a later retry succeeds. `last_transport_error_ready` and `last_transport_error_authenticated` distinguish failure of an already-established channel from a pre-readiness connection failure. `last_transport_error_reason` is a strict privacy-safe allowlist (`connection_reset`, `connection_timeout`, `network_unreachable`, bounded DNS/TLS classes, or `unknown`) rather than the raw operating-system message. The signed HTTPS fallback retains its last error class/reason after a later successful poll while resetting the current `http_poll_failures` count, so post-recovery diagnosis can determine whether WSS and HTTPS failed through the same system-network episode. None of these fields contains a hostname, address, DNS answer, certificate, close reason, or proxy endpoint. While no daemon channel is ready, `server_info.daemon.previous_connection` retains only the last verified channel's transport, connected/last-seen/disconnected timestamps, and sanitized relay diagnostics; it excludes policy, tools, account identity, daemon instance/connection identity, call IDs, arguments, and results, and it never participates in routing or authorization. `outage_duration_ms` measures the close-to-ready recovery episode; `previous_ready_inbound_silence_ms` measures how long the preceding ready socket had stopped producing inbound transport proof before it actually closed. The second value is therefore the field that exposes a black-holed OPEN WebSocket whose visible reconnect later completes quickly. A VPN/TUN UI may remain “connected” while its upstream route is unusable. Local OS logs can be compared with the exact `last_disconnected_at` timestamp, but an interface-quality change or tunnel-process correlation is not by itself proof of which product, node, edge, or upstream failed. Correlated failure of an independent HTTPS client at the same timestamp—for example an external API `unexpected EOF` while the relay records WebSocket 1006/`connection_reset`—is stronger evidence of a shared system-network/VPN/TUN episode than of a Machine Bridge event-loop or resource-admission failure; it still does not identify the failing tunnel node or upstream provider. Machine Bridge reports only coarse route/proxy classes and never sends or logs interface names, addresses, DNS answers, proxy credentials, or Worker secrets.
|
|
97
97
|
|
|
98
|
-
Brief retryable outages recover automatically. On a verified current daemon channel, `server_info.daemon.relay_transport.outage_active=false`; retained fields describe the immediately preceding transport episode rather than claiming a current outage. WebSocket remains preferred and requests a protocol-level probe after five seconds. Calling `ws.ping()` only queues the control frame; it is not treated as remote-probe dispatch until the WebSocket sender's write callback confirms that the Ping actually left the local send queue. The local sender has a separate thirty-second bounded dispatch window, while a confirmed Ping retains its full ten-second Pong deadline. This deliberately prevents compression/backpressure or a slow local socket queue from spending the remote-response budget before any probe was transmitted. A protocol Pong that arrives while a Ping callback is still pending records bidirectional proof for that dispatch round; if the local write callback then completes inside the thirty-second dispatch budget, it does not arm a stale future Pong deadline. Unrelated application inbound is receive-side evidence only and cannot prove the daemon-to-Worker direction. A local queue whose Ping write callback still has not completed after thirty seconds is classified as `relay_transport_send_timeout` even if unrelated inbound traffic continues. One dispatched Ping that reaches the ten-second response deadline does not hard-kill an otherwise ready WSS: the relay runtime enters a fifteen-second `transport_confirmation_pending` window, sends the existing JSON application heartbeat as an independent path check, and prewarms signed HTTPS in standby without taking ownership away from the still-ready WSS. A protocol Pong or the explicit JSON application `pong` clears suspicion and stops standby prewarm; ordinary tool/control traffic does not, because it proves only the Worker-to-daemon receive direction. Only a second-stage confirmation window that receives no application `pong` becomes `relay_transport_timeout` and terminates the WSS. This keeps a true black hole bounded while no longer amplifying a roughly ten-to-fifteen-second persistent-flow stall into an immediate reconnect storm. `heartbeat.probe_dispatch_*`, `heartbeat.transport_confirmation_*`, and bounded sender-backlog fields distinguish local send delay, first-stage response loss, successful second-stage recovery, and confirmed two-stage failure. The separate periodic JSON application heartbeat remains twenty-five seconds with a seventy-five-second application-silence timeout, begins only after verified relay readiness, and the Worker keeps a wider ninety-second WebSocket liveness fallback. This is a detection/recovery bound, not a guarantee that a degraded network can complete another WebSocket handshake inside the same interval. WebSocket connect attempts have a thirty-second outer budget so a degraded but still valid DNS/TCP/TLS/WebSocket upgrade is not misclassified by an unrealistically narrow connection cutoff. The daemon also explicitly disables client `permessage-deflate`: the relay carries bounded control/JSON traffic, while `ws` enables compression by default on clients and compression adds sender-state/CPU overhead that can queue later frames; the stability path does not need that optional negotiation. The fallback still begins independently rather than waiting thirty seconds for WSS. On first-stage WSS liveness suspicion, the same root-certified ephemeral daemon identity prewarms signed HTTPS in standby; if WSS proves live during the second-stage confirmation, that standby poller stops. If the WSS actually disconnects, fallback switches to exact-generation takeover immediately; an in-flight standby request is aborted and replaced rather than being allowed to consume up to its own request deadline before takeover can start. That in-memory session certificate intentionally has a 24-hour maximum lifetime. It is valid for ordinary reconnects during that lifetime, but it is not silently extended: if a later WSS reconnect/authentication attempt discovers that the daemon session has expired, the daemon terminates with `relay_device_session_expired` instead of retrying forever with unusable credentials. Installed launchd/systemd/Windows service supervision restarts the failed daemon and obtains a fresh root-signed session after normal runtime cleanup; the default portable root does this without user interaction. A manually run daemon must be restarted by its operator, and a configured Secure Enclave root retains its existing user-presence requirement when the new daemon start signs the replacement session. Each fallback request has a seven-second deadline, ordinary one-second ready poll cadence, five-second standby-prewarm cadence, bounded one/two/four/five-second retry backoff after request or protocol/session failures, a 750 ms hard minimum request-start interval, and a twelve-second liveness window; a new daemon-backed call waits at most fifteen seconds for some verified daemon channel, and the measured wait is deducted from that call's original execution budget. After an established WSS disappears, the daemon explicitly marks its signed HTTP request as a takeover of the Worker-issued `connection_id` for that exact disconnected WebSocket generation. Once candidate preconditions pass, HTTPS may retire only that targeted same-instance zombie WSS that the Worker has not yet observed closing. If a newer same-instance WSS is already ready before the HTTP request arrives, the old generation no longer matches and the stale takeover remains standby instead of retiring the recovered socket. A takeover request without the exact Worker-issued WebSocket connection ID is invalid rather than being treated as an instance-only legacy takeover. Malformed, stale, wrongly targeted, or different-instance requests cannot preempt a healthy incumbent. During replacement, the daemon reconciles `resume_calls`, processes `ready_ack`, proves local readiness, and only then returns `resume_calls_ack.missing_ids`. A missing ID therefore proves both that the same daemon has no active/unacknowledged-result ownership for that call and that the replacement channel is ready. If the initiating MCP response is still open and at least one second remains in the original execution budget, the Worker may transparently retransmit exactly that same call ID, arguments, authority, and a reduced timeout. `read_job` is stricter: redelivery requires the full ten-second reconciliation headroom to remain, otherwise the Worker declines redelivery and returns retryable recovery failure rather than rewriting the call into an under-budget immediate read. If safe redelivery cannot be accepted, the call falls back to retryable `unavailable` with `side_effects_started=false`. Calls that may have executed, retained terminal results, different-daemon calls, and ambiguous mutations are never automatically replayed. Completed relay results that are still waiting for Worker acknowledgement remain bounded in daemon memory and consume the same recovery-ownership capacity as active calls: 16 total with two control-plane slots reserved for `diagnose_runtime`/`list_roots`. When ordinary recovery ownership reaches 14, another ordinary relay call is rejected before execution with retryable `limit_exceeded` and `side_effects_started=false`; the two reserved diagnostic/recovery calls remain available until total capacity reaches 16. The retained-result implementation also keeps one non-admission emergency ownership slot solely for a violated internal capacity invariant: if an already-executed result reaches retention after the normal 16-entry ceiling is unexpectedly full, that one result remains retained for acknowledgement/reconnect ownership instead of being sent unowned and later misclassified as safe to redeliver. Use of that slot emits an error-level capacity event and may make diagnostics temporarily report ownership above the normal maximum; a second such overflow is not sent. This slot is not usable admission capacity and must never be counted to raise the 16-call execution ceiling. An acknowledgement that is permanently lost cannot pin a result forever: first retention is monotonic and the result expires after the 315-second maximum Worker settlement lifetime on the next live relay heartbeat; the disconnected path still uses the shorter reconnect-grace cleanup. `diagnose_runtime.runtime.relay_result_recovery` exposes only aggregate `active_calls`, `retained_results`, active ownership, and capacity counts—never call IDs, tool arguments, or results. Compare outage intervals with sleep/wake records and `diagnose_runtime.runtime.relay.heartbeat` before classifying them as active network faults; local stdio `server_info.runtime.relay.heartbeat` exposes the same state. On macOS, an authorized, schema-valid remote tool call also enters a bounded idle-sleep guard with a fixed thirty-minute rolling inactivity grace; the service does not depend on shell-only environment overrides that launchd would not persist as configuration, and relay heartbeats do not count as user activity. Authorized relay handlers hold one shared `/usr/bin/caffeinate -i -w <daemon-pid>` assertion for their full execution lifetime; concurrent handlers share the child, the thirty-minute default inactivity grace begins only after the last one settles, and a new authorized handler cancels any pending release timer so the full grace restarts after that activity settles. A remote `start_process` extends the same assertion only after resource admission succeeds and keeps it until the session child settles, so a long process session is not reduced to the handler grace window. Remote account managed-job runners independently hold `/usr/bin/caffeinate -i -w <runner-pid>` after their ownership claim is confirmed and persisted ownership identifies an account-backed job, then retain it through admission, steps, cleanup, and terminal persistence; local managed jobs do not acquire this remote-continuity assertion, and daemon reconnect/replacement does not own the remote runner protection. `diagnose_runtime.runtime.idle_sleep_guard` reports only daemon-side supported/enabled/active/grace/error-class state; it intentionally does not enumerate process-session or job identities. Runtime shutdown terminates process sessions before releasing the daemon guard. None of these assertions claim to prevent explicit sleep or lid-close sleep. A nonzero `event_loop_stall_count` with a large `max_event_loop_lag_ms` means the local daemon was not scheduled promptly; a short stall enters recovery grace, sends a fresh transport probe, and deliberately postpones disconnect. A large stall that aligns with `pmset` Sleep/Wake is suspension evidence, while a large `previous_ready_inbound_silence_ms` without a matching local stall is stronger evidence of a half-open/network-path interruption before close. Use `--verbose` only when close codes, liveness deadlines, and retry delays are required.
|
|
98
|
+
Brief retryable outages recover automatically. On a verified current daemon channel, `server_info.daemon.relay_transport.outage_active=false`; retained fields describe the immediately preceding transport episode rather than claiming a current outage. WebSocket remains preferred and requests a protocol-level probe after five seconds. Calling `ws.ping()` only queues the control frame; it is not treated as remote-probe dispatch until the WebSocket sender's write callback confirms that the Ping actually left the local send queue. The local sender has a separate thirty-second bounded dispatch window, while a confirmed Ping retains its full ten-second Pong deadline. This deliberately prevents compression/backpressure or a slow local socket queue from spending the remote-response budget before any probe was transmitted. A protocol Pong that arrives while a Ping callback is still pending records bidirectional proof for that dispatch round; if the local write callback then completes inside the thirty-second dispatch budget, it does not arm a stale future Pong deadline. Unrelated application inbound is receive-side evidence only and cannot prove the daemon-to-Worker direction. A local queue whose Ping write callback still has not completed after thirty seconds is classified as `relay_transport_send_timeout` even if unrelated inbound traffic continues. One dispatched Ping that reaches the ten-second response deadline does not hard-kill an otherwise ready WSS: the relay runtime enters a fifteen-second `transport_confirmation_pending` window, sends the existing JSON application heartbeat as an independent path check, and prewarms signed HTTPS in standby without taking ownership away from the still-ready WSS. A protocol Pong or the explicit JSON application `pong` clears suspicion and stops standby prewarm; ordinary tool/control traffic does not, because it proves only the Worker-to-daemon receive direction. Only a second-stage confirmation window that receives no application `pong` becomes `relay_transport_timeout` and terminates the WSS. This keeps a true black hole bounded while no longer amplifying a roughly ten-to-fifteen-second persistent-flow stall into an immediate reconnect storm. `heartbeat.probe_dispatch_*`, `heartbeat.transport_confirmation_*`, and bounded sender-backlog fields distinguish local send delay, first-stage response loss, successful second-stage recovery, and confirmed two-stage failure. The separate periodic JSON application heartbeat remains twenty-five seconds with a seventy-five-second application-silence timeout, begins only after verified relay readiness, and the Worker keeps a wider ninety-second WebSocket liveness fallback. This is a detection/recovery bound, not a guarantee that a degraded network can complete another WebSocket handshake inside the same interval. WebSocket connect attempts have a thirty-second outer budget so a degraded but still valid DNS/TCP/TLS/WebSocket upgrade is not misclassified by an unrealistically narrow connection cutoff. The daemon also explicitly disables client `permessage-deflate`: the relay carries bounded control/JSON traffic, while `ws` enables compression by default on clients and compression adds sender-state/CPU overhead that can queue later frames; the stability path does not need that optional negotiation. The fallback still begins independently rather than waiting thirty seconds for WSS. On first-stage WSS liveness suspicion, the same root-certified ephemeral daemon identity prewarms signed HTTPS in standby; if WSS proves live during the second-stage confirmation, that standby poller stops. If the WSS actually disconnects, fallback switches to exact-generation takeover immediately; an in-flight standby request is aborted and replaced rather than being allowed to consume up to its own request deadline before takeover can start. That in-memory session certificate intentionally has a 24-hour maximum lifetime. It is valid for ordinary reconnects during that lifetime, but it is not silently extended: if a later WSS reconnect/authentication attempt discovers that the daemon session has expired, the daemon terminates with `relay_device_session_expired` instead of retrying forever with unusable credentials. Installed launchd/systemd/Windows service supervision restarts the failed daemon and obtains a fresh root-signed session after normal runtime cleanup; the default portable root does this without user interaction. A manually run daemon must be restarted by its operator, and a configured Secure Enclave root retains its existing user-presence requirement when the new daemon start signs the replacement session. Each fallback request has a seven-second deadline, ordinary one-second ready poll cadence, five-second standby-prewarm cadence, bounded one/two/four/five-second retry backoff after request or protocol/session failures, a 750 ms hard minimum request-start interval, and a twelve-second liveness window; a new daemon-backed call waits at most fifteen seconds for some verified daemon channel, and the measured wait is deducted from that call's original execution budget. After an established WSS disappears, the daemon explicitly marks its signed HTTP request as a takeover of the Worker-issued `connection_id` for that exact disconnected WebSocket generation. Once candidate preconditions pass, HTTPS may retire only that targeted same-instance zombie WSS that the Worker has not yet observed closing. If a newer same-instance WSS is already ready before the HTTP request arrives, the old generation no longer matches and the stale takeover remains standby instead of retiring the recovered socket. A takeover request without the exact Worker-issued WebSocket connection ID is invalid rather than being treated as an instance-only legacy takeover. Malformed, stale, wrongly targeted, or different-instance requests cannot preempt a healthy incumbent. During replacement, the daemon reconciles `resume_calls`, processes `ready_ack`, proves local readiness, and only then returns `resume_calls_ack.missing_ids`. A missing ID therefore proves both that the same daemon has no active/unacknowledged-result ownership for that call and that the replacement channel is ready. If the initiating MCP response is still open and at least one second remains in the original execution budget, the Worker may transparently retransmit exactly that same call ID, arguments, authority, and a reduced timeout. `read_job` is stricter: redelivery requires the full ten-second reconciliation headroom to remain, otherwise the Worker declines redelivery and returns retryable recovery failure rather than rewriting the call into an under-budget immediate read. If safe redelivery cannot be accepted, the call falls back to retryable `unavailable` with `side_effects_started=false`. Calls that may have executed, retained terminal results, different-daemon calls, and ambiguous mutations are never automatically replayed. Completed relay results that are still waiting for Worker acknowledgement remain bounded in daemon memory and consume the same recovery-ownership capacity as active calls: 16 total with two control-plane slots reserved for `diagnose_runtime`/`list_roots`. When ordinary recovery ownership reaches 14, another ordinary relay call is rejected before execution with retryable `limit_exceeded` and `side_effects_started=false`; the two reserved diagnostic/recovery calls remain available until total capacity reaches 16. The retained-result implementation also keeps one non-admission emergency ownership slot solely for a violated internal capacity invariant: if an already-executed result reaches retention after the normal 16-entry ceiling is unexpectedly full, that one result remains retained for acknowledgement/reconnect ownership instead of being sent unowned and later misclassified as safe to redeliver. Use of that slot emits an error-level capacity event and may make diagnostics temporarily report ownership above the normal maximum; a second such overflow is not sent. This slot is not usable admission capacity and must never be counted to raise the 16-call execution ceiling. An acknowledgement that is permanently lost cannot pin a result forever: first retention is monotonic and the result expires after the 315-second maximum Worker settlement lifetime on the next live relay heartbeat; the disconnected path still uses the shorter reconnect-grace cleanup. `diagnose_runtime.runtime.relay_result_recovery` exposes only aggregate `active_calls`, `retained_results`, active ownership, and capacity counts—never call IDs, tool arguments, or results. Compare outage intervals with sleep/wake records and `diagnose_runtime.runtime.relay.heartbeat` before classifying them as active network faults; local stdio `server_info.runtime.relay.heartbeat` exposes the same state. On macOS, an authorized, schema-valid remote tool call also enters a bounded idle-sleep guard with a fixed thirty-minute rolling inactivity grace; the service does not depend on shell-only environment overrides that launchd would not persist as configuration, and relay heartbeats do not count as user activity. Authorized relay handlers hold one shared `/usr/bin/caffeinate -i -s -w <daemon-pid>` assertion for their full execution lifetime; `-s` strengthens system-sleep prevention only on AC power while `-i` remains the baseline Idle Sleep request; concurrent handlers share the child, the thirty-minute default inactivity grace begins only after the last one settles, and a new authorized handler cancels any pending release timer so the full grace restarts after that activity settles. A remote `start_process` extends the same assertion only after resource admission succeeds and keeps it until the session child settles, so a long process session is not reduced to the handler grace window. Remote account managed-job runners independently hold `/usr/bin/caffeinate -i -s -w <runner-pid>` after their ownership claim is confirmed and persisted ownership identifies an account-backed job, then retain it through admission, steps, cleanup, and terminal persistence; local managed jobs do not acquire this remote-continuity assertion, and daemon reconnect/replacement does not own the remote runner protection. `diagnose_runtime.runtime.idle_sleep_guard` reports only daemon-side supported/enabled/active/grace/error-class state plus whether the fixed child requests the AC-only system-sleep assertion; it intentionally does not enumerate process-session or job identities. Runtime shutdown terminates process sessions before releasing the daemon guard. None of these assertions claim to prevent explicit sleep or lid-close sleep. A nonzero `event_loop_stall_count` with a large `max_event_loop_lag_ms` means the local daemon was not scheduled promptly; a short stall enters recovery grace, sends a fresh transport probe, and deliberately postpones disconnect. A large stall that aligns with `pmset` Sleep/Wake is suspension evidence. Independently, `relay_outage_analysis` can show that a close-to-ready `connection_reset` interval was itself dominated by system sleep; an awake outage with no sleep overlap remains real transport evidence without identifying which network/host layer caused it. A large `previous_ready_inbound_silence_ms` without a matching local stall remains useful pre-close half-open evidence. Use `--verbose` only when close codes, liveness deadlines, and retry delays are required.
|
|
99
99
|
|
|
100
100
|
A foreground MCP response is not durable delivery. Hosted synchronous calls reserve room for Worker and host settlement instead of occupying the complete interaction window: ordinary daemon-backed tools default to 20 seconds of remote execution plus a separate five-second Worker settlement margin; ordinary configurable browser/application foreground tools also default to 20 seconds, while compound `computer_observe` and `computer_act` default to 30 seconds; all configurable browser/application foreground tools retain their explicit 45-second maximum. Remote `exec_command`, `run_process`, and `run_local_command` no longer keep the child process inside that response lifetime. Each remote process request must carry a unique caller-held `idempotency_key` before dispatch; reuse that same key only when recovering an ambiguous acceptance response. The daemon commits the authorized operation as a principal-bound one-step managed job, launches it with interactive resource-admission priority, and returns a `job_id` plus `read_job` recovery metadata inside a 10-second acceptance budget; the Worker keeps a separate five-second settlement margin. If that acceptance response is lost to settlement timeout, HTTP response cancellation, or relay reconnect expiry after dispatch, the public error remains non-retryable for generic callers but carries the original key and the explicit recovery action `retry_same_tool_arguments_with_same_idempotency_key`; this reconciles against the retained job instead of authorizing a blind duplicate. The detached child may execute for up to 600 seconds after admission, but the managed runner can separately wait up to thirty minutes for cooperative machine-user resource admission before the child is spawned; the child execution deadline begins only after that admission succeeds. The shared ceiling is exposed machine-readably as `server_info.tool_delivery.managed_job_resource_admission_wait_max_ms`, because the same pre-spawn boundary applies to ordinary durable process jobs and owner `start_job` steps rather than to process tools alone. While the runner is in this pre-spawn state, `read_job.current_phase` is `resource_admission`; no command has started yet. An owner can correlate a long-running status at that phase with `diagnose_runtime.runtime.resource_admission` rather than interpreting it as a slow child process; a delegated non-owner should treat the phase itself as evidence that the child has not spawned, retain the same `job_id`, and avoid blind replay rather than attempting the owner-only machine-wide diagnostic. After admission, the phase returns to `steps`, `finally_steps`, or `recovery-cleanup` as appropriate. Completed step records preserve `duration_ms` as the total orchestration duration. Local/owner reads additionally expose `resource_admission_ms` as the pre-spawn portion so a delayed successful child can be distinguished from slow execution after the fact; delegated non-owner reads omit that machine-user scheduling timing rather than turning shared-host contention into a more precise cross-workload signal. The detached job survives MCP disconnect, relay reconnect, daemon restart, or service replacement. Non-owner process authority is unchanged: automatic durable execution still uses the delegated workspace sandbox and does not grant owner-only `start_job`. If a cached host schema omits a current required field, the Worker rejects before daemon dispatch with a normal no-side-effect tool error and requests a `tools/list` refresh rather than surfacing a protocol-only validation failure. Discovery instructions and tool descriptions both carry orchestration semantics, so `server/discover` and `tools/list` each advertise `ttlMs=0` and every host-visible tool description carries `Tool schema generation N`. `server_info.tool_delivery.tool_schema_generation`, `tool_schema_server_version`, `discovery_ttl_ms`, and `tool_list_ttl_ms` identify the live contract; `host_visible_schema_known_to_server=false` is equally important because a healthy new daemon/Worker cannot prove that an external host discarded an older cached action/tool snapshot. `host_turn_deadline_observable=false` means Machine Bridge cannot pre-compute the external assistant-turn deadline, while `managed_jobs_detached_from_mcp_response=true` records that an accepted durable job is not owned by that response lifetime. After an activation that changes hosted semantics, compare the live `server_info` generation and changed invocation behavior with the governed Workspace Action control snapshot when that product layer is applicable; automation may perform the supported refresh/review path without another conversational approval. Host-internal cache inspection is intentionally excluded from operational release verification. `start_process` remains the explicit daemon-lifetime path when interactive stdin or session-style incremental output is required, but hosted calls use a 10-second execution / 15-second settlement envelope and do not queue behind resource pressure: the first failed admission returns retryable `unavailable`; owner-local callers retain the cooperative wait. Hosted `read_process` supports paced same-response follow-up: each actual output/exit blocking wait lasts at most one second. If another would-block remote read arrives inside the fifteen-second blocking cooldown, the daemon keeps that same MCP call open until output/exit or the cooldown boundary rather than returning an immediate running checkpoint; the Worker reserves enough execution/settlement headroom for that server-side pacing. Results use `status_polling_mode=paced_followup` while the process remains live, plus `blocking_poll_throttled` and `next_blocking_poll_after_ms`; callers must not busy-loop and should respect that cooldown. A new hosted call waits at most fifteen seconds for daemon readiness, but that wait is charged against the call's existing execution budget; an in-flight disconnect likewise never pauses or extends the original absolute deadline. Pending-call reconnect retention is also bounded by the smaller of reconnect grace and that original remaining deadline, and diagnostics distinguish `original call deadline expired during reconnect` from a true full `reconnect grace expired` rather than labeling both cases as the latter. Owner-local stdio/CLI calls retain their synchronous local contract because they do not depend on a hosted response stream. Keep unrelated mutations and verification independently terminal, and never infer task success merely because a durable launch was accepted. For one coherent non-interactive sequence, prefer a repository umbrella command or multi-step `start_job` rather than creating many one-step durable process carriers. If the current task needs the result, `read_job` may follow the known durable job repeatedly in the same assistant response until terminal state while calls continue to be accepted; active relay reads report `status_polling_mode=bounded_followup` and no longer recommend forced handoff. The normal hosted read is a 40-second server-side long-poll. Terminal settlement returns on the next bounded five-second internal poll; nonterminal status/phase/dependency progress is coalesced for at least 30 seconds by default, and `current_step`-only churn does not wake the hosted call. `wait_ms=0` is the explicit immediate-checkpoint mode, while clients with independently verified longer request lifetimes may explicitly request up to five minutes. The default is intentionally below the maximum because live host evidence showed that the former five-minute default could outlive the host tool invocation even though the durable job itself remained healthy. The coalescing floor reduces host-visible event density and does not shorten the managed job, the assistant task, or the six-hour managed-step ceiling. Do not busy-loop, do not replace server-side pacing with rapid immediate reads, do not use repeated `list_jobs`, `server_info`, or `diagnose_runtime` calls as substitute polling surfaces, and do not infer or preempt a host/tool deadline from elapsed wall-clock time. Return the `job_id`, status, and current phase for later recovery only after an actual host/tool boundary is observed, external input or authorization is required, or the user explicitly requested a checkpoint; only a terminal status is task-completion evidence.
|
|
101
101
|
|
|
@@ -288,11 +288,11 @@ machine-mcp job submit plan.json
|
|
|
288
288
|
|
|
289
289
|
Registry changes apply to newly submitted jobs without restarting the daemon. Active jobs use the resource snapshot accepted with their plan.
|
|
290
290
|
|
|
291
|
-
Policy changes affect new submissions. Cancel accepted running jobs explicitly when revoking execution authority. A staged plan is a non-running draft and has no terminal promotion path. A trusted owner uses `start_job`; an explicit local operator may submit a reviewed JSON plan with `job submit`. A managed job transitions through `queued`, `running`, `cleaning`, and a terminal status such as `succeeded`, `failed`, or `cancelled`. Cleanup-specific terminal variants report a failed finally phase. If a runner PID dies after main execution has begun, the next daemon/job-CLI reconciliation marks the job interrupted, removes stale private runtime copies, and runs the finally phase in recovery mode. A runner that dies while still in pre-execution `queued/dependency_wait` is instead relaunched as the same dependency-wait job, because no dependent main side effect has started yet. Recovery never replays ordinary business steps: a terminal `recovered` status means recovery/finally work completed after the original runner was interrupted, not that the main steps succeeded. The result compatibility field `recovered=true` only records that recovery mode produced the terminal result; `recovery_failed` and `recovery_exhausted` remain failures even when that field is true, so automation must branch on terminal `status`. Automatic recovery is capped at three attempts; persistent failure becomes `recovery_exhausted`.
|
|
291
|
+
Policy changes affect new submissions. Cancel accepted running jobs explicitly when revoking execution authority. A staged plan is a non-running draft and has no terminal promotion path. A trusted owner uses `start_job`; an explicit local operator may submit a reviewed JSON plan with `job submit`. A managed job transitions through `queued`, `running`, `cleaning`, and a terminal status such as `succeeded`, `failed`, or `cancelled`. Cleanup-specific terminal variants report a failed finally phase. If a runner PID dies after main execution has begun, the next daemon/job-CLI reconciliation marks the job interrupted, removes stale private runtime copies, and runs the finally phase in recovery mode. A runner that dies while still in pre-execution `queued/dependency_wait` is instead relaunched as the same dependency-wait job, because no dependent main side effect has started yet. During that wait, secure dependency-state reads classified as `permission_denied`, `identity_changed`, or generic `resource_unavailable` receive a fixed 45-second recovery grace so a transient Windows sharing/atomic-replacement window does not permanently fail the downstream; a successful read clears the grace, while persistent unavailability and missing/integrity/witness-invalid evidence still fail closed. Recovery never replays ordinary business steps: a terminal `recovered` status means recovery/finally work completed after the original runner was interrupted, not that the main steps succeeded. The result compatibility field `recovered=true` only records that recovery mode produced the terminal result; `recovery_failed` and `recovery_exhausted` remain failures even when that field is true, so automation must branch on terminal `status`. Automatic recovery is capped at three attempts; persistent failure becomes `recovery_exhausted`.
|
|
292
292
|
|
|
293
293
|
Use job-scoped `temporary_files` for local helpers. For remote maintenance, prefer `ssh ... sh -s` with the remote script in step `stdin`; this avoids remote temporary scripts. Explicit remote cleanup belongs in idempotent `finally_steps`.
|
|
294
294
|
|
|
295
|
-
Uninstall refuses to remove local state while any managed job remains active. Active plans are needed for recovery and are owner-only. Terminal jobs delete their full plans. Up to 512 managed-job states are retained, while `list_jobs` returns at most 50
|
|
295
|
+
Uninstall refuses to remove local state while any managed job remains active. Active plans are needed for recovery and are owner-only. Terminal jobs delete their full plans. Up to 512 managed-job states are retained, while `list_jobs.jobs` returns at most 50 primary records per response. Terminal results normally remain for up to seven days but capacity pressure may retire unprotected terminal history earlier, while staged drafts expire after 24 hours and active/staged/unreadable state is never evicted merely to make room. Completed one-step `exec_command`/`run_process`/`run_local_command` carriers are internally lower-priority `transient_process` retention; when removable transient history exists, it is retired before explicit managed-job terminal history. Under saturation, at most the newest 16 such results are protected for thirty minutes. If one of those retained recent process results is omitted from the durable-first primary `jobs` window, `list_jobs.recent_process_recovery` returns up to 16 additional authority-visible public job handles so a later host turn can recover the `job_id` and continue with `read_job`; the projection contains no step output or internal retention metadata and is inventory, not a polling or MCP replay/session surface. Active/staged jobs with `depends_on` pin their referenced retained jobs so a long workflow cannot lose an upstream result that it still needs. This keeps helper-command churn from preferentially destroying durable recovery evidence while keeping both response windows bounded. A later `read_job` for a valid ID whose record is no longer retained returns typed `not_found`; that is a loss of recovery evidence, not proof that the underlying operation never executed, so do not blindly replay side effects. Step output is never copied to ordinary daemon logs.
|
|
296
296
|
|
|
297
297
|
See [MANAGED_JOBS.md](MANAGED_JOBS.md).
|
|
298
298
|
|
package/docs/TESTING.md
CHANGED
|
@@ -16,7 +16,7 @@ The explicit task lists live in `scripts/check-plan.mjs`. The fast plan retains
|
|
|
16
16
|
|
|
17
17
|
Tests are verification inputs but are not npm tarball entries under the current `package.json.files` manifest. A test-only repository change therefore does not by itself change npm package bytes or require a synthetic package version; `release-impact:check` remains authoritative and still requires a version whenever `package.json`, `package-lock.json`, or any path selected by `package.json.files` changes. A source fix that requires a regression test is versioned because of its packaged source/documentation/metadata impact, not because `tests/` is implicitly shipped.
|
|
18
18
|
|
|
19
|
-
On macOS, `scripts/run-checks.mjs` re-executes the complete fast/platform/full verification process under `/usr/bin/caffeinate -i` before it captures verification inputs. The internal `MBM_CHECK_IDLE_SLEEP_GUARD` marker prevents recursive wrapping. This establishes an idle-system-sleep assertion for the whole verification lifetime so a short child timeout cannot expire only because the machine entered Idle Sleep between dispatch and settlement. That verification-lifetime guard is distinct from production ownership: authorized, schema-valid relay handlers hold a shared `/usr/bin/caffeinate -i -w <daemon-pid>` assertion for their full execution lifetime; the default thirty-minute rolling inactivity grace starts only after the last concurrent handler settles, and a new authorized handler cancels pending release so the full grace restarts after it settles. A remote process session extends the same daemon assertion from successful resource admission through child settlement. A remote account managed-job runner independently holds `/usr/bin/caffeinate -i -w <runner-pid>` only after its runner claim is confirmed and persisted ownership identifies an account-backed job, then retains it through terminal persistence; local managed jobs do not acquire that remote-continuity assertion, and the remote runner protection survives daemon reconnect/replacement. Runtime shutdown terminates process sessions before releasing the daemon assertion. The production thirty-minute relay grace is fixed rather than depending on shell-only environment inheritance that launchd does not persist as service configuration or treating relay heartbeats as user activity. None of these mechanisms claims to defeat explicit sleep
|
|
19
|
+
On macOS, `scripts/run-checks.mjs` re-executes the complete fast/platform/full verification process under `/usr/bin/caffeinate -i` before it captures verification inputs. The internal `MBM_CHECK_IDLE_SLEEP_GUARD` marker prevents recursive wrapping. This establishes an idle-system-sleep assertion for the whole verification lifetime so a short child timeout cannot expire only because the machine entered Idle Sleep between dispatch and settlement. That verification-lifetime guard is distinct from production ownership: authorized, schema-valid relay handlers hold a shared `/usr/bin/caffeinate -i -s -w <daemon-pid>` assertion for their full execution lifetime; `-s` requests the stronger system-sleep assertion only effective on AC, while `-i` remains the baseline Idle Sleep assertion. The default thirty-minute rolling inactivity grace starts only after the last concurrent handler settles, and a new authorized handler cancels pending release so the full grace restarts after it settles. A remote process session extends the same daemon assertion from successful resource admission through child settlement. A remote account managed-job runner independently holds `/usr/bin/caffeinate -i -s -w <runner-pid>` only after its runner claim is confirmed and persisted ownership identifies an account-backed job, then retains it through terminal persistence; local managed jobs do not acquire that remote-continuity assertion, and the remote runner protection survives daemon reconnect/replacement. Runtime shutdown terminates process sessions before releasing the daemon assertion. The production thirty-minute relay grace is fixed rather than depending on shell-only environment inheritance that launchd does not persist as service configuration or treating relay heartbeats as user activity. None of these mechanisms claims to defeat explicit sleep, lid-close sleep, or battery-mode system sleep; those still invalidate the usefulness of a live-machine verification run.
|
|
20
20
|
|
|
21
21
|
Successful child-task stdout/stderr is intentionally suppressed so a long green plan does not overwhelm an MCP response or hide the final status behind host truncation. Progress and timing remain visible. A failed task returns bounded head/tail diagnostics for both streams. Set `MBM_CHECK_VERBOSE=1` only when an operator explicitly needs live child output; verbose mode can be large and should be run through a process session when used remotely.
|
|
22
22
|
|
|
@@ -71,8 +71,8 @@ The suite includes:
|
|
|
71
71
|
- managed-job non-timing approval, resource validation/redaction, output, discard, and cleanup/recovery steps with a named 120-second success budget; the aggregate-output proof uses four steps with a 600-second observer, while the managed-job tree fixture uses a 180-second timeout and 150-second descendant-readiness window; timeout/tree-kill and cancellation remain independent semantic tests;
|
|
72
72
|
- deterministic child-settlement coverage for delayed `exit`/`close` delivery, including POSIX zombie detection and recovery of the real exit code without weakening genuine timeout/tree-kill semantics;
|
|
73
73
|
- ordinary managed-job terminal observation uses 480 seconds, exceeding the longest three-phase 3×120-second fixture plan plus startup margin without changing production timeout semantics;
|
|
74
|
-
- managed-job retention regressions fill the 512-state store while the inventory
|
|
75
|
-
- managed-job dependency regressions prove that a dependent job remains `queued/dependency_wait` without spawning its main child, all-success releases execution, later upstream failure settles `dependency_failed` without a long artifact wait, already-failed/staged dependencies are rejected before acceptance, dependency progress wakes hosted long-polls, and a SIGKILLed dependency-wait runner is automatically relaunched as the same pre-execution job without user input. Dependency orchestration fixtures drive upstream pending/terminal state through persisted job evidence and use the exact runtime Node version probe for any downstream child, rather than using real host resource pressure as a timing mechanism; resource-admission behavior remains covered independently by its dedicated suite;
|
|
74
|
+
- managed-job retention regressions fill the 512-state store while the primary inventory remains capped at 50 records; they prove that saturated capacity preserves the newest 16 transient one-step results inside the thirty-minute recovery grace, that those recent helpers remain recoverable through the separate at-most-16 `recent_process_recovery` public-handle projection when the durable-first primary window omits them, that internal retention metadata is not projected, that the incoming helper counts against the bounded reserve so a seventeenth recent result reclaims the oldest transient rather than growing the reserve, that older/excess transient history remains preferred eviction material before ordinary durable history, that an active/staged `depends_on` plan pins its referenced terminal record against capacity pruning, and that a valid-but-unretained job ID returns privacy-safe non-retryable `not_found` rather than generic `execution_failed` or a path-bearing message;
|
|
75
|
+
- managed-job dependency regressions prove that a dependent job remains `queued/dependency_wait` without spawning its main child, all-success releases execution, later upstream failure settles `dependency_failed` without a long artifact wait, already-failed/staged dependencies are rejected before acceptance, dependency progress wakes hosted long-polls, and a SIGKILLed dependency-wait runner is automatically relaunched as the same pre-execution job without user input. A deterministic fake-monotonic-clock fixture injects the same wrapped Windows `permission_denied` shape produced by secure managed-state reads: one transient failure must recover on the next read, while persistent unavailability must remain bounded by the fixed 45-second dependency-state recovery grace and then fail closed as `dependency_unavailable`. Missing/integrity/witness failures are not included in that retry class. Dependency orchestration fixtures drive upstream pending/terminal state through persisted job evidence and use the exact runtime Node version probe for any downstream child, rather than using real host resource pressure as a timing mechanism; resource-admission behavior remains covered independently by its dedicated suite;
|
|
76
76
|
- managed-job runner-claim coverage injects an atomic-publication `MBM_IDENTITY_CHANGED` observation and proves the publication-coupled reader retries to the stable generation while preserving ordinary claim validation; the full local self-test separately exercises the real resource-CLI runner publication path that exposed the race;
|
|
77
77
|
- managed-job CLI list/inspect/submit/read success and rejection fixtures with a distinct 120-second subprocess budget plus structured status/signal/error diagnostics; explicit job-step timeout and cancellation contracts remain independently short;
|
|
78
78
|
- local-self managed-job CLI subprocesses explicitly clear `NODE_V8_COVERAGE`; the top-level local-self remains instrumented, while dedicated CLI-entrypoint and managed-job fixtures provide the gated module evidence without recursive profiler startup;
|
|
@@ -130,7 +130,7 @@ The suite includes:
|
|
|
130
130
|
- shared no-follow bounded-file reads plus present-path inspection where only `ENOENT` means absence for security/persistence evidence; EIO/EACCES injection covers pairing, service environment, service owner, global config, activation/soak records, and conformance checkout state. Optional project metadata separately proves BigInt path/descriptor identity mismatch rejects before reading bytes, descriptor cleanup, bounded byte-limit validation, and conservative omission of unsupported discovery data. Normal files, over-limit data, directories, symbolic links, and multiple-hard-link denial remain covered; typed file-mutation regressions cover create-only collisions, stale SHA-256 preconditions, missing/ambiguous edit text, malformed/stale patches, transactional rollback, Worker preservation, stdio projection, and path/content/hash non-disclosure;
|
|
131
131
|
- owner-only directory enforcement rejecting final symlinks, failing closed on POSIX chmod errors, verifying `0700`, and retaining Windows portability; Worker temporary-secret lifecycle coverage for process-start-bound names, valid stale-owner reclamation, ambiguous-owner retention, `0600` mode, deletion failures, and simultaneous deployment/cleanup failures;
|
|
132
132
|
- SARIF security-gate behavior for unknown findings, exact accepted rule/path matches, path mismatch rejection, rationale quality, and exception expiry;
|
|
133
|
-
- deterministic property tests over hostile browser-protocol byte strings, canonical/custom policy combinations, argv bounds/NULs, and a real direct process proving shell metacharacters remain literal argv; timeout-alignment tests prove remote process tools require a caller-held `idempotency_key`, reserve only a 10-second durable-acceptance budget plus the separate Worker settlement margin, and keep their detached one-step execution budget independently capped at 600 seconds. Cached schemas that omit the new recovery key fail as normal no-side-effect tool results with refresh guidance, while a cached schema that still permits the former five-second `read_process` maximum receives one-second-limit / `next_blocking_poll_after_ms` / durable-work guidance rather than “short polling” or forced handoff; over-limit and malformed durable timeouts still fail before dispatch. Worker pending-call regressions also reject non-finite, non-positive, non-integer, and over-contract operation/reconnect delays before timer/alarm state mutates; local relay-envelope tests reject malformed call IDs, tool-name shapes, authorization field types, missing/string/negative timeouts, and over-contract timeouts instead of coercing them. Ordinary daemon tools default to 20 seconds plus the separate Worker settlement margin, ordinary configurable browser/application tools default to 20 seconds, compound `computer_observe`/`computer_act` default to 30 seconds, all configurable browser/application tools remain explicitly capped at 45 seconds, relay-origin `start_process` has a 10-second execution budget, and `read_process` actual output/exit blocking waits remain capped at one second with a fifteen-second would-block cooldown. Regression coverage requires a repeated would-block request inside that cooldown to remain in the same MCP call until output/exit or the cooldown boundary rather than return a rapid running checkpoint, with Worker execution/settlement budget covering that pacing. A `wait_for_exit=true` cooldown regression emits frequent stdout while the child remains live and proves those ordinary output notifications do not release the exit-wait pacing stage early. Live relay reads report `status_polling_mode=paced_followup` and no forced handoff; terminal reads report `terminal`. Relay-origin active `read_job` reports `status_polling_mode=bounded_followup` with `host_turn_handoff_recommended=false`, terminal job reads report `terminal`, and relay-origin `list_jobs` remains an `inventory` surface. Managed-job wait regressions prove missing hosted `wait_ms` defaults to a host-safe 40-second server-side long-poll even for stale hosts that do not know the parameter, explicit zero remains immediate, the explicit maximum remains five minutes for clients that can carry it, progress returns early, cancellation interrupts internal waits, five-second progress probes do not perform full runner reconciliation, full reconcile is bounded to thirty-second intervals inside the advertised wait, the monotonic deadline starts before the initial full read, an unchanged timeout does not add another heavyweight post-deadline reconcile, healthy hosted runner ownership uses an async process-start probe that leaves the Node event loop runnable instead of spawning synchronous `ps`, the default 40-second wait bounds per-call interaction density but a `duration / wait interval` response count is only planning data rather than proof of aggregate host-response survival, and the dedicated `read_job` execution budget leaves ten seconds of reconciliation headroom—50/55 seconds execution/settlement at the default and 310/315 seconds at the explicit maximum. Recovery-delay regressions prove the requested wait shrinks inside the original absolute deadline, and an execution window below the ten-second reconciliation headroom is rejected before dispatch/redelivery instead of being rewritten into an under-budget immediate read; ordinary relay tools retain their 50-second ceiling. Hosted tool descriptions permit bounded same-response job/session follow-up while forbidding busy loops, rapid immediate-checkpoint substitution, and listing/diagnostic-surface polling substitutions. They also explicitly forbid inferring or preempting a host/tool deadline from elapsed wall-clock time: continuation stops only after an actual host/tool boundary is observed, external input/authorization is required, or the user requested a checkpoint. Tool-schema regressions require real `server/discover` and `tools/list` output to advertise `ttlMs=0`, current 2026-07-28 discovery to advertise `tools.listChanged=true`, legacy initialization compatibility to retain `listChanged=false`, and a `subscriptions/listen` request for `toolsListChanged` to receive the supported-subset acknowledgement plus `notifications/tools/list_changed` while the subscription stays open through the short observation window and then closes on explicit cancellation or the advertised 10-second server lease. Subscription-capacity coverage proves one account is capped at 8 active streams, four accounts can collectively fill the 32-stream Durable Object ceiling without exceeding their own limits, the next request fails before another stream is allocated, and an idempotent release returns exactly one reusable slot. Subscription-open coverage requires reservation alone not to count as a successfully opened stream, successful stream construction to mark only that authenticated account, cancellation and lease expiry each to release active capacity while retaining bounded opened-stream evidence, and opened-stream history to evict older entries beyond 64 accounts. Worker integration additionally proves that aborting the public HTTP client releases the Durable Object slot no later than the bounded lease even when the Workers runtime supplies no reliable disconnect signal. Release-documentation regressions separately bind hosted-client freshness to the client's real update model: server-opened subscription evidence cannot prove external client receipt; when ChatGPT workspace governance freezes actions, the **Action control** snapshot after automation performs the supported Refresh/review path is the product-level publication evidence, and opaque host-internal cache inspection is intentionally excluded from release acceptance. Recreation/republication is reserved for a governed Workspace snapshot that cannot be updated in place or remains stale/partial/mixed after the supported refresh path, or for separate workspace governance that explicitly requires it. Live acceptance must also cross at least one previous-generation argument or behavior boundary with a harmless invocation probe—beta.112 uses a terminal `read_job` with `wait_ms=40001` and a non-executing staged step with `timeout_seconds=3601`, while generation 6 additionally requires a valid-but-unretained `job_id` to return typed non-retryable `not_found`—because displayed definitions alone do not prove the actual invocation/daemon path. Discovery instructions must contain the current continuation contract, every host-visible tool description must carry `Tool schema generation N`, and `server_info.tool_delivery` must expose the same generation plus live server version, `discovery_ttl_ms=0`, `tool_list_ttl_ms=0`, `host_visible_schema_known_to_server=false`, `host_turn_deadline_observable=false`, and `managed_jobs_detached_from_mcp_response=true`. Relay job/process results additionally carry the current schema generation and host-deadline non-observability so live backend evidence remains diagnostic even when an external host has cached an older tool description. Detached/rebound calls cannot extend their original absolute deadline and disconnect diagnostics distinguish an original deadline exhausted during reconnect from a true reconnect-grace expiry, and owner-local registered commands retain their local manifest budget. Computer Use regressions additionally prove that application observation consumes one end-to-end screenshot/Accessibility/window-revalidation deadline and that action preflight/dispatch/verification/post-observation consume one action deadline without reclassifying a deadline exhaustion as stale/unavailable. Runtime infrastructure regressions prove the macOS remote-activity idle-sleep guard uses fixed non-shell `/usr/bin/caffeinate -i -w <daemon-pid>`, renews one bounded timer without duplicate children, releases after inactivity, and remains disabled off macOS. Runtime self-tests additionally prove the daemon independently rejects relay durable execution without the recovery credential, remote process acceptance returns a recoverable `job_id`, the job outlives the completed MCP call, the Worker's 600-second durable execution field survives the daemon's narrower local foreground-schema defense without widening local request-scoped calls, operator execution preserves the delegated workspace sandbox without gaining owner-only `start_job`, and durable results are retrievable through `read_job`. Worker integration deliberately withholds one daemon acceptance result until the settlement deadline, verifies the Worker sends cancellation for that transient call, and then proves the MCP error still carries the original `idempotency_key` plus the exact same-key replay action instead of losing recovery context; process-tree tests also assert Darwin uses a target-PGID `ps` query, preserve the global inspection budget, and repeatedly prove anti-`SIGTERM` descendants exit after foreground timeout;
|
|
133
|
+
- deterministic property tests over hostile browser-protocol byte strings, canonical/custom policy combinations, argv bounds/NULs, and a real direct process proving shell metacharacters remain literal argv; timeout-alignment tests prove remote process tools require a caller-held `idempotency_key`, reserve only a 10-second durable-acceptance budget plus the separate Worker settlement margin, and keep their detached one-step execution budget independently capped at 600 seconds. Cached schemas that omit the new recovery key fail as normal no-side-effect tool results with refresh guidance, while a cached schema that still permits the former five-second `read_process` maximum receives one-second-limit / `next_blocking_poll_after_ms` / durable-work guidance rather than “short polling” or forced handoff; over-limit and malformed durable timeouts still fail before dispatch. Worker pending-call regressions also reject non-finite, non-positive, non-integer, and over-contract operation/reconnect delays before timer/alarm state mutates; local relay-envelope tests reject malformed call IDs, tool-name shapes, authorization field types, missing/string/negative timeouts, and over-contract timeouts instead of coercing them. Ordinary daemon tools default to 20 seconds plus the separate Worker settlement margin, ordinary configurable browser/application tools default to 20 seconds, compound `computer_observe`/`computer_act` default to 30 seconds, all configurable browser/application tools remain explicitly capped at 45 seconds, relay-origin `start_process` has a 10-second execution budget, and `read_process` actual output/exit blocking waits remain capped at one second with a fifteen-second would-block cooldown. Regression coverage requires a repeated would-block request inside that cooldown to remain in the same MCP call until output/exit or the cooldown boundary rather than return a rapid running checkpoint, with Worker execution/settlement budget covering that pacing. A `wait_for_exit=true` cooldown regression emits frequent stdout while the child remains live and proves those ordinary output notifications do not release the exit-wait pacing stage early. Live relay reads report `status_polling_mode=paced_followup` and no forced handoff; terminal reads report `terminal`. Relay-origin active `read_job` reports `status_polling_mode=bounded_followup` with `host_turn_handoff_recommended=false`, terminal job reads report `terminal`, and relay-origin `list_jobs` remains an `inventory` surface. Managed-job wait regressions prove missing hosted `wait_ms` defaults to a host-safe 40-second server-side long-poll even for stale hosts that do not know the parameter, explicit zero remains immediate, the explicit maximum remains five minutes for clients that can carry it, progress returns early, cancellation interrupts internal waits, five-second progress probes do not perform full runner reconciliation, full reconcile is bounded to thirty-second intervals inside the advertised wait, the monotonic deadline starts before the initial full read, an unchanged timeout does not add another heavyweight post-deadline reconcile, healthy hosted runner ownership uses an async process-start probe that leaves the Node event loop runnable instead of spawning synchronous `ps`, the default 40-second wait bounds per-call interaction density but a `duration / wait interval` response count is only planning data rather than proof of aggregate host-response survival, and the dedicated `read_job` execution budget leaves ten seconds of reconciliation headroom—50/55 seconds execution/settlement at the default and 310/315 seconds at the explicit maximum. Recovery-delay regressions prove the requested wait shrinks inside the original absolute deadline, and an execution window below the ten-second reconciliation headroom is rejected before dispatch/redelivery instead of being rewritten into an under-budget immediate read; ordinary relay tools retain their 50-second ceiling. Hosted tool descriptions permit bounded same-response job/session follow-up while forbidding busy loops, rapid immediate-checkpoint substitution, and listing/diagnostic-surface polling substitutions. They also explicitly forbid inferring or preempting a host/tool deadline from elapsed wall-clock time: continuation stops only after an actual host/tool boundary is observed, external input/authorization is required, or the user requested a checkpoint. Tool-schema regressions require real `server/discover` and `tools/list` output to advertise `ttlMs=0`, current 2026-07-28 discovery to advertise `tools.listChanged=true`, legacy initialization compatibility to retain `listChanged=false`, and a `subscriptions/listen` request for `toolsListChanged` to receive the supported-subset acknowledgement plus `notifications/tools/list_changed` while the subscription stays open through the short observation window and then closes on explicit cancellation or the advertised 10-second server lease. Subscription-capacity coverage proves one account is capped at 8 active streams, four accounts can collectively fill the 32-stream Durable Object ceiling without exceeding their own limits, the next request fails before another stream is allocated, and an idempotent release returns exactly one reusable slot. Subscription-open coverage requires reservation alone not to count as a successfully opened stream, successful stream construction to mark only that authenticated account, cancellation and lease expiry each to release active capacity while retaining bounded opened-stream evidence, and opened-stream history to evict older entries beyond 64 accounts. Worker integration additionally proves that aborting the public HTTP client releases the Durable Object slot no later than the bounded lease even when the Workers runtime supplies no reliable disconnect signal. Release-documentation regressions separately bind hosted-client freshness to the client's real update model: server-opened subscription evidence cannot prove external client receipt; when ChatGPT workspace governance freezes actions, the **Action control** snapshot after automation performs the supported Refresh/review path is the product-level publication evidence, and opaque host-internal cache inspection is intentionally excluded from release acceptance. Recreation/republication is reserved for a governed Workspace snapshot that cannot be updated in place or remains stale/partial/mixed after the supported refresh path, or for separate workspace governance that explicitly requires it. Live acceptance must also cross at least one previous-generation argument or behavior boundary with a harmless invocation probe—beta.112 uses a terminal `read_job` with `wait_ms=40001` and a non-executing staged step with `timeout_seconds=3601`, while generation 6 additionally requires a valid-but-unretained `job_id` to return typed non-retryable `not_found`—because displayed definitions alone do not prove the actual invocation/daemon path. Discovery instructions must contain the current continuation contract, every host-visible tool description must carry `Tool schema generation N`, and `server_info.tool_delivery` must expose the same generation plus live server version, `discovery_ttl_ms=0`, `tool_list_ttl_ms=0`, `host_visible_schema_known_to_server=false`, `host_turn_deadline_observable=false`, and `managed_jobs_detached_from_mcp_response=true`. Relay job/process results additionally carry the current schema generation and host-deadline non-observability so live backend evidence remains diagnostic even when an external host has cached an older tool description. Detached/rebound calls cannot extend their original absolute deadline and disconnect diagnostics distinguish an original deadline exhausted during reconnect from a true reconnect-grace expiry, and owner-local registered commands retain their local manifest budget. Computer Use regressions additionally prove that application observation consumes one end-to-end screenshot/Accessibility/window-revalidation deadline and that action preflight/dispatch/verification/post-observation consume one action deadline without reclassifying a deadline exhaustion as stale/unavailable. Runtime infrastructure regressions prove the macOS remote-activity idle-sleep guard uses fixed non-shell `/usr/bin/caffeinate -i -s -w <daemon-pid>`, renews one bounded timer without duplicate children, releases after inactivity, and remains disabled off macOS. Runtime self-tests additionally prove the daemon independently rejects relay durable execution without the recovery credential, remote process acceptance returns a recoverable `job_id`, the job outlives the completed MCP call, the Worker's 600-second durable execution field survives the daemon's narrower local foreground-schema defense without widening local request-scoped calls, operator execution preserves the delegated workspace sandbox without gaining owner-only `start_job`, and durable results are retrievable through `read_job`. Worker integration deliberately withholds one daemon acceptance result until the settlement deadline, verifies the Worker sends cancellation for that transient call, and then proves the MCP error still carries the original `idempotency_key` plus the exact same-key replay action instead of losing recovery context; process-tree tests also assert Darwin uses a target-PGID `ps` query, preserve the global inspection budget, and repeatedly prove anti-`SIGTERM` descendants exit after foreground timeout;
|
|
134
134
|
- service/activation recovery tests prove an ambiguous launchd `bootout` against an initially active service preserves `restore_required` even when unload cannot be verified, and persistent activation retries previous-provider start/bootstrap until the exact prior service runtime is verified ready instead of abandoning recovery after one transient start failure. Candidate installation now also proves owner state remains pending through service start and exact daemon/Worker verification, commit occurs only after convergence, a recovered start commits only after recovery convergence, owner-commit failure stops the verified provider, and repeated convergence failure stops an uncommitted active candidate. The activation regression preserves the original `autostart_stop_failed` error after successful rollback and aggregates primary plus rollback failures when recovery cannot converge;
|
|
135
135
|
- prototype-shaped command, action, role, profile, form-field, keyboard, and resource names proving that inherited object properties are never interpreted as dispatch or authority; current-schema malformed OAuth roles are repaired to disabled reviewer accounts with credential revocation;
|
|
136
136
|
- canonical Worker deployment URL extraction proving unrelated `/mcp`, `/healthz`, path-bearing, and wrong-name URLs cannot be persisted as upload evidence;
|
package/docs/TOOL_REFERENCE.md
CHANGED
|
@@ -2790,7 +2790,7 @@ Terminate a live server-managed process tree with graceful or forced termination
|
|
|
2790
2790
|
|
|
2791
2791
|
**Diagnose runtime layers**
|
|
2792
2792
|
|
|
2793
|
-
Run fixed, non-user-controlled local probes and return privacy-safe control-plane state to distinguish MCP policy, local filesystem, process-spawn, shell, managed-job storage, resource-admission, relay, event-density, and operating-system-suspension evidence. Owner/local diagnostics include bounded managed-job churn, content-free recent security-audit tool-call aggregates, bounded resource-waiter summaries with current pre-spawn admission reasons, and on macOS
|
|
2793
|
+
Run fixed, non-user-controlled local probes and return privacy-safe control-plane state to distinguish MCP policy, local filesystem, process-spawn, shell, managed-job storage, resource-admission, relay, event-density, and operating-system-suspension evidence. Owner/local diagnostics include bounded managed-job churn, content-free recent security-audit tool-call aggregates, bounded resource-waiter summaries with current pre-spawn admission reasons, and on macOS bounded sleep-history correlation for both runtime stalls and the most recently recovered relay outage; they never expose tool arguments/results, waiter IDs/tokens/PIDs, private paths, or raw power logs. A successful response proves the request reached the local daemon; it cannot diagnose a host refusal that blocks the tool call itself or observe ChatGPT final-message receipt.
|
|
2794
2794
|
|
|
2795
2795
|
| Contract field | Value |
|
|
2796
2796
|
|---|---|
|
|
@@ -3091,7 +3091,7 @@ Validate and persist a durable managed-job draft without starting any process. T
|
|
|
3091
3091
|
|
|
3092
3092
|
**Start managed job**
|
|
3093
3093
|
|
|
3094
|
-
Durably accept a detached argv-based job with ordered steps, job-scoped temporary files, guaranteed-attempt finally steps, and optional depends_on links to existing managed jobs. The independent local runner continues if the MCP connection disappears. Dependencies are evaluated from durable managed-job terminal state rather than blind artifact polling: while an upstream job is active the dependent job remains queued in current_phase=dependency_wait without spawning its main child, then proceeds after all dependencies succeed or terminates with error_class=dependency_failed if an upstream later fails. Staged or already-failed dependencies are rejected before acceptance. On macOS, an authorized remote account job keeps a runner-owned idle-sleep assertion only after the runner ownership claim is confirmed and persisted account ownership is validated, then through dependency wait, admission, steps, cleanup, and terminal persistence; local managed jobs do not acquire this remote-continuity assertion, and it does not override explicit sleep or lid-close behavior. Use an idempotency_key to deduplicate an uncertain retry while the original job record is retained; eviction/retention expiry ends that deduplication window, so this is not permanent exactly-once execution. Use {{temp:name}}, {{resource:name}}, env_resources, or stdin_resource; registered resource contents never enter MCP arguments. Each main/finally step defaults to 600 seconds and may explicitly request up to 21600 seconds (six hours); use this managed-job surface when one continuous command legitimately needs more than the 600-second durable one-step process limit. For one coherent non-interactive workflow, prefer one multi-step managed job over a chain of one-step process carriers so host-visible event density stays bounded. Do not split a long task merely to satisfy an assumed host deadline, and do not use the user as a polling clock.
|
|
3094
|
+
Durably accept a detached argv-based job with ordered steps, job-scoped temporary files, guaranteed-attempt finally steps, and optional depends_on links to existing managed jobs. The independent local runner continues if the MCP connection disappears. Dependencies are evaluated from durable managed-job terminal state rather than blind artifact polling: while an upstream job is active the dependent job remains queued in current_phase=dependency_wait without spawning its main child, then proceeds after all dependencies succeed or terminates with error_class=dependency_failed if an upstream later fails. A transient dependency-state read classified as permission_denied, identity_changed, or generic resource_unavailable receives a fixed 45-second recovery grace; one successful read clears that grace, while persistent unavailability fails closed as dependency_unavailable and missing/integrity/witness-invalid evidence remains immediate failure. Staged or already-failed dependencies are rejected before acceptance. On macOS, an authorized remote account job keeps a runner-owned idle-sleep assertion only after the runner ownership claim is confirmed and persisted account ownership is validated, then through dependency wait, admission, steps, cleanup, and terminal persistence; local managed jobs do not acquire this remote-continuity assertion, and it does not override explicit sleep or lid-close behavior. Use an idempotency_key to deduplicate an uncertain retry while the original job record is retained; eviction/retention expiry ends that deduplication window, so this is not permanent exactly-once execution. Use {{temp:name}}, {{resource:name}}, env_resources, or stdin_resource; registered resource contents never enter MCP arguments. Each main/finally step defaults to 600 seconds and may explicitly request up to 21600 seconds (six hours); use this managed-job surface when one continuous command legitimately needs more than the 600-second durable one-step process limit. For one coherent non-interactive workflow, prefer one multi-step managed job over a chain of one-step process carriers so host-visible event density stays bounded. Do not split a long task merely to satisfy an assumed host deadline, and do not use the user as a polling clock.
|
|
3095
3095
|
|
|
3096
3096
|
| Contract field | Value |
|
|
3097
3097
|
|---|---|
|
|
@@ -3306,7 +3306,7 @@ Durably accept a detached argv-based job with ordered steps, job-scoped temporar
|
|
|
3306
3306
|
|
|
3307
3307
|
**List managed jobs**
|
|
3308
3308
|
|
|
3309
|
-
List
|
|
3309
|
+
List detached managed-job recovery state without returning step output. The primary jobs array remains capped at 50 and prioritizes unreadable, active, staged, and durable terminal recovery state before transient one-step helper history so helper churn cannot hide an older recoverable long-running job or its terminal result. If a recent one-step process result is still retained inside the fixed 30-minute/16-result recovery reserve but omitted from that primary window, recent_process_recovery returns up to 16 additional authority-visible public job handles so a later host turn can recover the job_id and continue with read_job. This is bounded recovery discovery only: it is not an MCP terminal-result replay/session store and must not be used for polling. The durable retained-state store is larger than these response windows. Owner/local callers also receive coarse retained-state capacity including durable-versus-transient terminal counts plus recent job-creation/churn aggregates without internal retired filenames, filesystem identities, argv, paths, or output.
|
|
3310
3310
|
|
|
3311
3311
|
| Contract field | Value |
|
|
3312
3312
|
|---|---|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "machine-bridge-mcp",
|
|
3
|
-
"version": "3.0.0-beta.
|
|
3
|
+
"version": "3.0.0-beta.144",
|
|
4
4
|
"description": "Cross-client MCP bridge for local agent context, structured browser and application automation, files, Git, processes, resources, and durable jobs over stdio or OAuth relay.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|
|
@@ -16,7 +16,7 @@ export class MacosIdleSleepAssertion {
|
|
|
16
16
|
if (this.child) return true;
|
|
17
17
|
let child = null;
|
|
18
18
|
try {
|
|
19
|
-
child = this.spawnProcess("/usr/bin/caffeinate", ["-i", "-w", String(this.processId)], {
|
|
19
|
+
child = this.spawnProcess("/usr/bin/caffeinate", ["-i", "-s", "-w", String(this.processId)], {
|
|
20
20
|
stdio: "ignore", shell: false, windowsHide: true,
|
|
21
21
|
});
|
|
22
22
|
this.child = child;
|
|
@@ -60,6 +60,11 @@ export class MacosIdleSleepAssertion {
|
|
|
60
60
|
}
|
|
61
61
|
|
|
62
62
|
snapshot() {
|
|
63
|
-
return {
|
|
63
|
+
return {
|
|
64
|
+
supported: this.supported,
|
|
65
|
+
active: Boolean(this.child),
|
|
66
|
+
requests_system_sleep_prevention_on_ac: Boolean(this.child),
|
|
67
|
+
last_error_class: this.lastErrorClass,
|
|
68
|
+
};
|
|
64
69
|
}
|
|
65
70
|
}
|
|
@@ -4,13 +4,15 @@ import { join } from "node:path";
|
|
|
4
4
|
import { createMonotonicDeadline } from "./monotonic-deadline.mjs";
|
|
5
5
|
import { managedJobDependencyCount, managedJobDependencyLabel } from "./managed-job-dependency-metadata.mjs";
|
|
6
6
|
import { resolveManagedJobDirectory } from "./managed-job-directory.mjs";
|
|
7
|
-
import { readJson, readRequiredJson } from "./managed-job-storage.mjs";
|
|
7
|
+
import { readJson, readRequiredJson, resourceErrorClass } from "./managed-job-storage.mjs";
|
|
8
8
|
import {
|
|
9
9
|
ACTIVE_JOB_STATES, isTerminalManagedJobResult, isTerminalManagedJobStatus, terminalStatusFromResult,
|
|
10
10
|
} from "./managed-job-terminal.mjs";
|
|
11
11
|
|
|
12
12
|
const DEPENDENCY_POLL_INTERVAL_MS = 1000;
|
|
13
13
|
const MAX_DEPENDENCY_WAIT_MS = 14 * 24 * 60 * 60 * 1000;
|
|
14
|
+
export const DEPENDENCY_STATE_READ_RECOVERY_GRACE_MS = 45_000;
|
|
15
|
+
const RETRYABLE_DEPENDENCY_READ_ERROR_CLASSES = new Set(["identity_changed", "permission_denied", "resource_unavailable"]);
|
|
14
16
|
|
|
15
17
|
export class ManagedJobDependencyError extends Error {
|
|
16
18
|
constructor(errorClass, message, details = {}) {
|
|
@@ -59,6 +61,7 @@ export async function waitForManagedJobDependencies({
|
|
|
59
61
|
}
|
|
60
62
|
|
|
61
63
|
const deadline = createMonotonicDeadline(MAX_DEPENDENCY_WAIT_MS, now);
|
|
64
|
+
const unavailableSinceById = new Map();
|
|
62
65
|
let previousPending = -1;
|
|
63
66
|
while (!deadline.expired()) {
|
|
64
67
|
throwIfCancelled();
|
|
@@ -69,11 +72,21 @@ export async function waitForManagedJobDependencies({
|
|
|
69
72
|
try {
|
|
70
73
|
status = await readStatus(jobId);
|
|
71
74
|
} catch (error) {
|
|
75
|
+
const errorClass = resourceErrorClass(error);
|
|
76
|
+
const elapsedMs = deadline.elapsedMs();
|
|
77
|
+
const unavailableSinceMs = unavailableSinceById.get(jobId) ?? elapsedMs;
|
|
78
|
+
if (RETRYABLE_DEPENDENCY_READ_ERROR_CLASSES.has(errorClass)
|
|
79
|
+
&& elapsedMs - unavailableSinceMs < DEPENDENCY_STATE_READ_RECOVERY_GRACE_MS) {
|
|
80
|
+
unavailableSinceById.set(jobId, unavailableSinceMs);
|
|
81
|
+
pending += 1;
|
|
82
|
+
continue;
|
|
83
|
+
}
|
|
72
84
|
throw new ManagedJobDependencyError("dependency_unavailable", "managed job dependency state is no longer available", {
|
|
73
85
|
dependency_job_id: jobId,
|
|
74
|
-
cause_class: managedJobDependencyLabel(
|
|
86
|
+
cause_class: managedJobDependencyLabel(errorClass),
|
|
75
87
|
});
|
|
76
88
|
}
|
|
89
|
+
unavailableSinceById.delete(jobId);
|
|
77
90
|
assertDependencyWitness(status, witness, jobId);
|
|
78
91
|
if (managedJobDependencySucceeded(status)) continue;
|
|
79
92
|
if (isTerminalManagedJobStatus(status?.status)) {
|
|
@@ -6,9 +6,9 @@ import { readJson, resourceErrorClass, safeReadDir } from "./managed-job-storage
|
|
|
6
6
|
import { MANAGED_JOB_ID } from "./managed-job-directory.mjs";
|
|
7
7
|
import { managedJobCapacitySnapshot, MAX_JOBS, MAX_LISTED_JOBS } from "./managed-job-capacity.mjs";
|
|
8
8
|
import { managedJobRecentActivity } from "./managed-job-activity.mjs";
|
|
9
|
+
import { recentProcessRecoveryJobs } from "./managed-job-recovery-listing.mjs";
|
|
9
10
|
import { ACTIVE_JOB_STATES, isTerminalManagedJobStatus } from "./managed-job-terminal.mjs";
|
|
10
11
|
import { clampInteger } from "./numbers.mjs";
|
|
11
|
-
|
|
12
12
|
export function listManagedJobs({ jobRoot, args, context, logger, reconcileStatus, assertKnownStatus, maximumLimit = MAX_LISTED_JOBS }) {
|
|
13
13
|
const limit = clampInteger(args.limit, 20, 1, Math.min(MAX_JOBS, maximumLimit));
|
|
14
14
|
const records = [];
|
|
@@ -36,11 +36,13 @@ export function listManagedJobs({ jobRoot, args, context, logger, reconcileStatu
|
|
|
36
36
|
|| String(right.job?.created_at || "").localeCompare(String(left.job?.created_at || ""))
|
|
37
37
|
|| String(left.job?.job_id || "").localeCompare(String(right.job?.job_id || "")));
|
|
38
38
|
const visibleJobs = records.slice(0, limit).map((record) => record.job);
|
|
39
|
+
const recentProcessRecovery = recentProcessRecoveryJobs(records, visibleJobs);
|
|
39
40
|
const capacity = managedJobCapacitySnapshot(jobRoot);
|
|
40
41
|
const durableTerminal = records.filter((record) => record.retentionClass !== "transient_process" && isTerminalManagedJobStatus(String(record.job?.status || ""))).length;
|
|
41
42
|
const transientTerminal = records.filter((record) => record.retentionClass === "transient_process" && isTerminalManagedJobStatus(String(record.job?.status || ""))).length;
|
|
42
43
|
return {
|
|
43
44
|
jobs: visibleJobs,
|
|
45
|
+
recent_process_recovery: recentProcessRecovery,
|
|
44
46
|
retained: records.length,
|
|
45
47
|
maximum: MAX_JOBS,
|
|
46
48
|
...hostedManagedJobListStatus(visibleJobs, context),
|
|
@@ -49,7 +51,6 @@ export function listManagedJobs({ jobRoot, args, context, logger, reconcileStatu
|
|
|
49
51
|
} : {}),
|
|
50
52
|
};
|
|
51
53
|
}
|
|
52
|
-
|
|
53
54
|
function recoveryPriority(job, retentionClass) {
|
|
54
55
|
if (job?.status === "unreadable") return 0;
|
|
55
56
|
if (ACTIVE_JOB_STATES.has(String(job?.status || ""))) return 1;
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
// @ts-check
|
|
2
|
+
|
|
3
|
+
import { TRANSIENT_PROCESS_RECOVERY_SLOTS, transientProcessWithinRecoveryGrace } from "./managed-job-retention-policy.mjs";
|
|
4
|
+
|
|
5
|
+
export function recentProcessRecoveryJobs(records, visibleJobs, now = Date.now()) {
|
|
6
|
+
const visibleJobIds = new Set(visibleJobs.map((job) => job.job_id));
|
|
7
|
+
return records
|
|
8
|
+
.filter((record) => !visibleJobIds.has(record.job.job_id)
|
|
9
|
+
&& transientProcessWithinRecoveryGrace({
|
|
10
|
+
retention_class: record.retentionClass,
|
|
11
|
+
finished_at: record.job.finished_at,
|
|
12
|
+
}, Number.NaN, now))
|
|
13
|
+
.sort((left, right) => String(right.job?.finished_at || "").localeCompare(String(left.job?.finished_at || ""))
|
|
14
|
+
|| String(right.job?.created_at || "").localeCompare(String(left.job?.created_at || ""))
|
|
15
|
+
|| String(left.job?.job_id || "").localeCompare(String(right.job?.job_id || "")))
|
|
16
|
+
.slice(0, TRANSIENT_PROCESS_RECOVERY_SLOTS)
|
|
17
|
+
.map((record) => record.job);
|
|
18
|
+
}
|
|
@@ -67,9 +67,9 @@ export class RemoteActivityIdleSleepGuard {
|
|
|
67
67
|
snapshot() {
|
|
68
68
|
const assertion = this.assertion.snapshot();
|
|
69
69
|
return {
|
|
70
|
-
supported: assertion.supported,
|
|
71
|
-
enabled: this.enabled,
|
|
70
|
+
supported: assertion.supported, enabled: this.enabled,
|
|
72
71
|
active: assertion.active,
|
|
72
|
+
requests_system_sleep_prevention_on_ac: assertion.requests_system_sleep_prevention_on_ac,
|
|
73
73
|
active_activities: this.activeActivities,
|
|
74
74
|
grace_ms: this.graceMs,
|
|
75
75
|
...this.timeline.snapshot(),
|
|
@@ -15,6 +15,7 @@ export function diagnosticInterpretation() {
|
|
|
15
15
|
managed_job_inventory_priority: "list_jobs prioritizes unreadable, active, and staged recovery state ahead of terminal history so short helper churn cannot hide a recoverable long-running job",
|
|
16
16
|
security_audit_recent_activity: "owner-visible security_audit.recent_activity aggregates content-free hash-chained relay tool events that reached the daemon. Host-only schema discovery, control-plane actions, and final-response delivery do not reach this audit, so its count is not the total host event count and it does not observe ChatGPT host turn termination or final response receipt",
|
|
17
17
|
system_sleep_history: "on macOS, runtime.system_sleep is a bounded fixed pmset projection of recent sleep intervals. event_loop_pause_analysis=matched_system_sleep means the runtime pause ended with a same-duration operating-system sleep interval; it is evidence of machine suspension rather than proof that JavaScript synchronously blocked the event loop",
|
|
18
|
+
relay_outage_sleep_correlation: "runtime.relay_outage_analysis compares the most recently recovered relay disconnect interval with the bounded macOS sleep history. majority_system_sleep_overlap means most of that observed relay outage occurred while the machine was suspended, so a connection reset is transport aftermath rather than sufficient evidence of an independent network root cause",
|
|
18
19
|
idle_sleep_guard_timeline: "runtime.idle_sleep_guard exposes only coarse activity/release timestamps and release reason so a later sleep can be compared with Machine Bridge power-assertion ownership without exposing tool arguments or host conversation identity",
|
|
19
20
|
};
|
|
20
21
|
}
|
|
@@ -28,6 +28,7 @@ export function diagnosticControlPlaneState(state = {}, relay = null) {
|
|
|
28
28
|
idle_sleep_guard: state.idleSleepGuard ?? null,
|
|
29
29
|
system_sleep: state.systemSleep ?? null,
|
|
30
30
|
event_loop_pause_analysis: state.eventLoopPauseAnalysis ?? null,
|
|
31
|
+
relay_outage_analysis: state.relayOutageAnalysis ?? null,
|
|
31
32
|
},
|
|
32
33
|
};
|
|
33
34
|
}
|
|
@@ -7,7 +7,7 @@ import { systemNetworkRouteCheck } from "./system-network-route.mjs";
|
|
|
7
7
|
import { diagnosticControlPlaneState } from "./runtime-diagnostic-state.mjs";
|
|
8
8
|
import { resourceAdmissionDiagnostic } from "./resource-admission-diagnostics.mjs";
|
|
9
9
|
import { diagnosticActivityProjection, diagnosticInterpretation } from "./runtime-diagnostic-projection.mjs";
|
|
10
|
-
import { correlateEventLoopStallWithSystemSleep, systemSleepDiagnostic } from "./system-sleep-diagnostics.mjs";
|
|
10
|
+
import { correlateEventLoopStallWithSystemSleep, correlateRelayOutageWithSystemSleep, systemSleepDiagnostic } from "./system-sleep-diagnostics.mjs";
|
|
11
11
|
export const RUNTIME_DIAGNOSTIC_PROCESS_TIMEOUT_MS = 30_000;
|
|
12
12
|
export async function diagnoseRuntime({
|
|
13
13
|
policy,
|
|
@@ -98,7 +98,7 @@ export async function diagnoseRuntime({
|
|
|
98
98
|
const managedJobs = typeof managedJobManager.status === "function" ? managedJobManager.status(context) : null;
|
|
99
99
|
const activity = diagnosticActivityProjection(controlPlaneState, admissionDiagnostic.snapshot, managedJobs, context);
|
|
100
100
|
activity.state.systemSleep = sleepDiagnostic.snapshot;
|
|
101
|
-
activity.state.eventLoopPauseAnalysis = correlateEventLoopStallWithSystemSleep(relay, sleepDiagnostic.snapshot);
|
|
101
|
+
activity.state.eventLoopPauseAnalysis = correlateEventLoopStallWithSystemSleep(relay, sleepDiagnostic.snapshot); activity.state.relayOutageAnalysis = correlateRelayOutageWithSystemSleep(relay, sleepDiagnostic.snapshot);
|
|
102
102
|
const resources = managedJobManager.listResources();
|
|
103
103
|
checks.push({
|
|
104
104
|
layer: "local-resource-registry",
|
|
@@ -62,6 +62,74 @@ export function correlateEventLoopStallWithSystemSleep(relay, snapshot) {
|
|
|
62
62
|
return stallProjection(match ? "matched_system_sleep" : "no_matching_recent_system_sleep", heartbeat, match || null);
|
|
63
63
|
}
|
|
64
64
|
|
|
65
|
+
export function correlateRelayOutageWithSystemSleep(relay, snapshot) {
|
|
66
|
+
const startedAt = Date.parse(String(relay?.last_disconnected_at || ""));
|
|
67
|
+
if (!(startedAt > 0)) return relayOutageProjection("no_recorded_relay_outage");
|
|
68
|
+
const endedAt = Date.parse(String(relay?.last_ready_at || ""));
|
|
69
|
+
if (relay?.outage_active === true || !(endedAt >= startedAt)) {
|
|
70
|
+
return relayOutageProjection("relay_outage_active", {
|
|
71
|
+
outageStartedAt: new Date(startedAt).toISOString(),
|
|
72
|
+
});
|
|
73
|
+
}
|
|
74
|
+
const outageDurationMs = endedAt - startedAt;
|
|
75
|
+
if (snapshot?.supported === false) {
|
|
76
|
+
return relayOutageProjection("unsupported_platform", {
|
|
77
|
+
outageStartedAt: new Date(startedAt).toISOString(), outageEndedAt: new Date(endedAt).toISOString(), outageDurationMs,
|
|
78
|
+
});
|
|
79
|
+
}
|
|
80
|
+
if (snapshot?.available !== true) {
|
|
81
|
+
return relayOutageProjection("system_sleep_history_unavailable", {
|
|
82
|
+
outageStartedAt: new Date(startedAt).toISOString(), outageEndedAt: new Date(endedAt).toISOString(), outageDurationMs,
|
|
83
|
+
});
|
|
84
|
+
}
|
|
85
|
+
const overlap = sleepOverlap(startedAt, endedAt, snapshot.recent_sleep_intervals || []);
|
|
86
|
+
const ratio = outageDurationMs > 0 ? overlap.durationMs / outageDurationMs : 0;
|
|
87
|
+
const classification = overlap.durationMs <= 0
|
|
88
|
+
? "no_matching_recent_system_sleep"
|
|
89
|
+
: ratio >= 0.5 ? "majority_system_sleep_overlap" : "partial_system_sleep_overlap";
|
|
90
|
+
return relayOutageProjection(classification, {
|
|
91
|
+
outageStartedAt: new Date(startedAt).toISOString(),
|
|
92
|
+
outageEndedAt: new Date(endedAt).toISOString(),
|
|
93
|
+
outageDurationMs,
|
|
94
|
+
sleepOverlapMs: overlap.durationMs,
|
|
95
|
+
sleepOverlapRatio: Number(ratio.toFixed(4)),
|
|
96
|
+
matchedSleepCount: overlap.count,
|
|
97
|
+
});
|
|
98
|
+
}
|
|
99
|
+
|
|
100
|
+
function sleepOverlap(startedAt, endedAt, intervals) {
|
|
101
|
+
const segments = [];
|
|
102
|
+
for (const interval of intervals) {
|
|
103
|
+
const sleepStart = Date.parse(String(interval?.started_at || ""));
|
|
104
|
+
const sleepEnd = Date.parse(String(interval?.ended_at || ""));
|
|
105
|
+
if (!(sleepStart >= 0) || !(sleepEnd > sleepStart)) continue;
|
|
106
|
+
const start = Math.max(startedAt, sleepStart);
|
|
107
|
+
const end = Math.min(endedAt, sleepEnd);
|
|
108
|
+
if (end > start) segments.push([start, end]);
|
|
109
|
+
}
|
|
110
|
+
segments.sort((left, right) => left[0] - right[0]);
|
|
111
|
+
let durationMs = 0; let count = 0; let activeStart = null; let activeEnd = null;
|
|
112
|
+
for (const [start, end] of segments) {
|
|
113
|
+
if (activeStart === null) { activeStart = start; activeEnd = end; count += 1; continue; }
|
|
114
|
+
if (start <= activeEnd) { activeEnd = Math.max(activeEnd, end); continue; }
|
|
115
|
+
durationMs += activeEnd - activeStart; activeStart = start; activeEnd = end; count += 1;
|
|
116
|
+
}
|
|
117
|
+
if (activeStart !== null) durationMs += activeEnd - activeStart;
|
|
118
|
+
return { durationMs, count };
|
|
119
|
+
}
|
|
120
|
+
|
|
121
|
+
function relayOutageProjection(classification, values = {}) {
|
|
122
|
+
return {
|
|
123
|
+
classification,
|
|
124
|
+
outage_started_at: values.outageStartedAt || null,
|
|
125
|
+
outage_ended_at: values.outageEndedAt || null,
|
|
126
|
+
outage_duration_ms: Number(values.outageDurationMs) || 0,
|
|
127
|
+
sleep_overlap_ms: Number(values.sleepOverlapMs) || 0,
|
|
128
|
+
sleep_overlap_ratio: Number(values.sleepOverlapRatio) || 0,
|
|
129
|
+
matched_sleep_count: Number(values.matchedSleepCount) || 0,
|
|
130
|
+
};
|
|
131
|
+
}
|
|
132
|
+
|
|
65
133
|
function stallProjection(classification, heartbeat, matchedSleep) {
|
|
66
134
|
return {
|
|
67
135
|
classification,
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
"supportedProtocolVersions": [
|
|
5
5
|
"2026-07-28"
|
|
6
6
|
],
|
|
7
|
-
"toolSchemaGeneration":
|
|
7
|
+
"toolSchemaGeneration": 14,
|
|
8
8
|
"instructions": [
|
|
9
9
|
"You are connected to a local workspace through machine-bridge-mcp.",
|
|
10
10
|
"Use resolve_task_capabilities only when the task needs local skill/registered-command discovery, application/browser routing, or refreshed project instructions. Straightforward file, Git, and shell work should use the already exposed tools directly instead of adding a resolver call. When capability routing is useful, pass the current user request and target path and reuse refresh.fingerprint as known_refresh_fingerprint to omit unchanged static instructions. Its execution_routing is bounded, set-level, and advisory: direct Bash remains the general escape hatch whenever the effective policy allows it. MCP transport connections are not conversation identity; refresh task-specific ranking when routing is actually needed rather than assuming prior requests established context.",
|
|
@@ -18,13 +18,13 @@
|
|
|
18
18
|
"Filename sensitivity is not classified by this server; the MCP host, connector gateway, local OS, or endpoint-security software may enforce additional independent rules.",
|
|
19
19
|
"A named full profile is canonical: it always enables writes, shell execution, unrestricted paths, the full parent environment, absolute paths, and the complete tool catalog. Any explicit narrowing is stored as custom.",
|
|
20
20
|
"The local daemon may advertise the full catalog as its capability ceiling. In remote mode, the Worker relays only the intersection allowed by the authenticated account role; a connector host may then expose an even smaller subset that Machine Bridge cannot observe or override. Conversation/surface app routing and host-cached action/tool snapshots are host-owned state, not daemon authority.",
|
|
21
|
-
"Use diagnose_runtime to distinguish a request that reached the daemon from local filesystem, process-spawn, shell, managed-job-storage, resource-admission, relay, event-density, and operating-system-suspension evidence. Owner-visible diagnostics include bounded managed-job churn, content-free recent security-audit tool-call density, the oldest resource waiters with their current pre-spawn admission reasons, and on macOS a bounded fixed sleep-history projection plus event-loop-pause correlation. A matched
|
|
21
|
+
"Use diagnose_runtime to distinguish a request that reached the daemon from local filesystem, process-spawn, shell, managed-job-storage, resource-admission, relay, event-density, and operating-system-suspension evidence. Owner-visible diagnostics include bounded managed-job churn, content-free recent security-audit tool-call density, the oldest resource waiters with their current pre-spawn admission reasons, and on macOS a bounded fixed sleep-history projection plus event-loop-pause and most-recent-relay-outage correlation. A matched event-loop sleep requires both pause duration and resume time to agree with the operating-system interval; relay_outage_analysis separately reports exact overlap duration/ratio for the latest recovered disconnect. majority_system_sleep_overlap means most of that outage occurred while the Mac was suspended, so a retained connection_reset is aftermath evidence rather than proof of an independent network root cause. An unmatched pause or awake relay reset is not automatically assigned another cause. The idle-sleep guard also exposes coarse activity/grace/release timestamps and whether its fixed caffeinate child requests AC-only system-sleep prevention in addition to idle-sleep prevention. Use these fields to distinguish host-visible event amplification from a job waiting for CPU/IO/memory admission or a machine that actually slept. A tool call blocked before any response cannot be diagnosed by the server; do not call that a platform disable without host-side evidence. Conversely, an execution failure that includes a child exit code or bounded child stdout/stderr proves the local process was spawned; for SSH or another nested command, treat an allowlist/forced-command refusal in that output as downstream target evidence rather than Machine Bridge policy denial. If a fresh supported conversation can invoke the same app/tool while an older conversation cannot, investigate the host conversation/surface route or cached action/tool snapshot before changing Machine Bridge policy.",
|
|
22
22
|
"Never request or return secret-file contents when a local resource alias can be used. Resources are registered through the local machine-mcp CLI; under canonical full policy, generate_ssh_key_resource can generate and register an Ed25519 key without returning private content and may be injected by path, stdin, or environment without entering MCP arguments.",
|
|
23
23
|
"stage_job persists a validated non-running draft only. It is not an approval workflow and cannot be promoted from the terminal; execution uses start_job when the effective account policy permits it, or an explicit local machine-mcp job submit PLAN.json operation.",
|
|
24
24
|
"Hosted start_job requires a caller-chosen idempotency_key before dispatch. If its acceptance response is ambiguous, retry exactly the same start_job arguments with the same key rather than creating a second logical submission or switching to another execution path. Recovery errors name the credential source as the original request and do not echo the key value. Local/stdio start_job keeps the underlying optional-key API.",
|
|
25
|
-
"For durable cross-job ordering, prefer start_job.depends_on over shell loops that wait for files or markers. A dependent runner remains queued in current_phase=dependency_wait without spawning its main child until all referenced managed jobs succeed; dependency_pending_count is progress observed by hosted read_job and is coalesced by the same nonterminal progress pacing as other active-state changes. If an upstream job later fails, the dependent settles with result.error_class=dependency_failed rather than blind-waiting for an artifact that can never appear. Staged or already-failed dependencies are rejected before acceptance. Active dependency plans pin the referenced retained job records against capacity pruning. This is a continuity mechanism, not a reason to shorten the overall user task or force a new user turn.",
|
|
25
|
+
"For durable cross-job ordering, prefer start_job.depends_on over shell loops that wait for files or markers. A dependent runner remains queued in current_phase=dependency_wait without spawning its main child until all referenced managed jobs succeed; dependency_pending_count is progress observed by hosted read_job and is coalesced by the same nonterminal progress pacing as other active-state changes. If an upstream job later fails, the dependent settles with result.error_class=dependency_failed rather than blind-waiting for an artifact that can never appear. During dependency_wait, a secure state read classified as permission_denied, identity_changed, or generic resource_unavailable receives a fixed 45-second recovery grace; one successful read clears that grace, while persistent unavailability fails closed as dependency_unavailable and missing/integrity/witness-invalid evidence remains immediate failure. Staged or already-failed dependencies are rejected before acceptance. Active dependency plans pin the referenced retained job records against capacity pruning. This is a continuity mechanism, not a reason to shorten the overall user task or force a new user turn.",
|
|
26
26
|
"Hosted read-only status and diagnostic tools must not be used as busy loops. For a known managed job, use read_job's server-side paced long-poll by default: a relay-origin active read waits inside Machine Bridge for up to the advertised managed-job read interval. Terminal settlement returns on the next bounded progress poll, but nonterminal changes are coalesced for at least the advertised thirty-second progress interval (or the caller's explicitly shorter wait), and current_step-only churn does not wake the host call by itself. This keeps multi-step long jobs observable without turning every short local step into a host-visible event. Use wait_ms=0 only for an intentional immediate checkpoint. Interactive process sessions use the separately paced read_process contract: the one-second actual output/exit blocking cap remains, while a repeated would-block request inside the fifteen-second cooldown stays inside that same MCP call until output/exit or the cooldown boundary rather than returning a rapid running checkpoint. Bounded same-response follow-up is allowed when the current task needs terminal state or additional output. Do not switch among list_jobs, server_info, diagnose_runtime, or other status surfaces merely to evade pacing. Per-call wait survival and aggregate host-response lifetime are separate constraints: do not infer that a long task can remain in one assistant response merely because each individual read succeeds. Do not infer or preempt a host/tool deadline from elapsed wall-clock time. While tool calls continue to be accepted and the task still needs the result, bounded same-response follow-up may continue; if an actual host/tool boundary ends the response, preserve the durable job/session identifier and resume that same operation later instead of resubmitting its side effect. A planned local-daemon restart may settle an in-flight hosted read_job as retryable unavailable with recovery.mode=read_same_job and the original job_id; after reconnect, read that same job_id instead of resubmitting the job's business side effect. A typed read_job not_found means the retained job record is unavailable, not that the underlying operation definitely never ran; do not blindly replay it.",
|
|
27
|
-
"For ordinary one-step remote process work whose child execution fits the 600-second process limit, exec_command, run_process, and run_local_command first commit a durable managed job, then keep the original hosted tool call open for the short advertised initial-settlement window. Acceptance transfers execution to durable ownership without forcing the current assistant response to end. If the helper settles inside that window, its terminal status/result is returned in the same tool response and no separate read_job event is required; if it remains active, the response keeps the same job_id/recovery envelope and normal read_job continuation applies. This reduces the common helper-plus-read double event without weakening durable ownership or changing the child timeout. For one coherent non-interactive workflow that needs several local commands, prefer one repository-native umbrella command or one multi-step start_job (or the smallest practical number of start_job plans) instead of emitting a chain of one-step process calls. Host-visible tool-event density is a continuity resource: batching reduces host traffic without shortening the underlying task or step lifetime. When the current task needs the result, bounded same-response read_job follow-up is allowed only when follow_up_read_required remains true; do not busy-loop or use repeated list_jobs as a substitute polling surface, and do not infer a host/tool deadline from elapsed wall-clock time. Cooperative machine-user admission is a separate pre-spawn interval and may wait up to 30 minutes; read_job.current_phase=resource_admission means no child for that step has started yet. These one-step process carriers use lower-priority terminal retention than explicit managed jobs so helper-command churn reclaims completed helper history before displacing explicit durable recovery results when such helper history is available; the managed-job store remains bounded to 512 retained states
|
|
27
|
+
"For ordinary one-step remote process work whose child execution fits the 600-second process limit, exec_command, run_process, and run_local_command first commit a durable managed job, then keep the original hosted tool call open for the short advertised initial-settlement window. Acceptance transfers execution to durable ownership without forcing the current assistant response to end. If the helper settles inside that window, its terminal status/result is returned in the same tool response and no separate read_job event is required; if it remains active, the response keeps the same job_id/recovery envelope and normal read_job continuation applies. This reduces the common helper-plus-read double event without weakening durable ownership or changing the child timeout. For one coherent non-interactive workflow that needs several local commands, prefer one repository-native umbrella command or one multi-step start_job (or the smallest practical number of start_job plans) instead of emitting a chain of one-step process calls. Host-visible tool-event density is a continuity resource: batching reduces host traffic without shortening the underlying task or step lifetime. When the current task needs the result, bounded same-response read_job follow-up is allowed only when follow_up_read_required remains true; do not busy-loop or use repeated list_jobs as a substitute polling surface, and do not infer a host/tool deadline from elapsed wall-clock time. Cooperative machine-user admission is a separate pre-spawn interval and may wait up to 30 minutes; read_job.current_phase=resource_admission means no child for that step has started yet. These one-step process carriers use lower-priority terminal retention than explicit managed jobs so helper-command churn reclaims completed helper history before displacing explicit durable recovery results when such helper history is available; the managed-job store remains bounded to 512 retained states. list_jobs keeps its primary jobs window capped at 50 and prioritizes unreadable, active, staged, and durable terminal recovery state ahead of transient one-step helper history; when a recent transient terminal is retained inside the 30-minute/16-result recovery reserve but omitted from that primary window, recent_process_recovery returns up to 16 additional authority-visible public job handles without step output or internal retention metadata. This is recovery discovery after a real host/tool boundary, not a polling or MCP replay/session surface; use read_job on the recovered job_id. Owner/local capacity diagnostics distinguish durable_terminal from transient_terminal without exposing job identities. Use start_job for policy-authorized multi-step plans, job-scoped temporary_files, explicit resource injection, idempotent finally_steps, or a step that legitimately needs an execution budget above 600 seconds. Detached jobs continue after MCP disconnects and daemon replacement; return job_id/status/current_phase for later recovery only after an actual host/tool boundary is observed, external input or authorization is required, or the user explicitly requested a checkpoint.",
|
|
28
28
|
"Prefer sending remote shell programs through a process stdin rather than creating remote helper files. If temporary files are necessary, use {{temp:name}} or explicit finally_steps.",
|
|
29
29
|
"Managed-job resource redaction is defense in depth, not a guarantee against transformed or partial secret output. Use capture_output=discard for steps that may echo credentials.",
|
|
30
30
|
"Do not use shell encoding, renaming, or alternate tools to bypass host or platform safety policy.",
|
|
@@ -2408,7 +2408,7 @@
|
|
|
2408
2408
|
{
|
|
2409
2409
|
"name": "diagnose_runtime",
|
|
2410
2410
|
"title": "Diagnose runtime layers",
|
|
2411
|
-
"description": "Run fixed, non-user-controlled local probes and return privacy-safe control-plane state to distinguish MCP policy, local filesystem, process-spawn, shell, managed-job storage, resource-admission, relay, event-density, and operating-system-suspension evidence. Owner/local diagnostics include bounded managed-job churn, content-free recent security-audit tool-call aggregates, bounded resource-waiter summaries with current pre-spawn admission reasons, and on macOS
|
|
2411
|
+
"description": "Run fixed, non-user-controlled local probes and return privacy-safe control-plane state to distinguish MCP policy, local filesystem, process-spawn, shell, managed-job storage, resource-admission, relay, event-density, and operating-system-suspension evidence. Owner/local diagnostics include bounded managed-job churn, content-free recent security-audit tool-call aggregates, bounded resource-waiter summaries with current pre-spawn admission reasons, and on macOS bounded sleep-history correlation for both runtime stalls and the most recently recovered relay outage; they never expose tool arguments/results, waiter IDs/tokens/PIDs, private paths, or raw power logs. A successful response proves the request reached the local daemon; it cannot diagnose a host refusal that blocks the tool call itself or observe ChatGPT final-message receipt.",
|
|
2412
2412
|
"availability": "always",
|
|
2413
2413
|
"annotations": {
|
|
2414
2414
|
"readOnlyHint": true,
|
|
@@ -2681,7 +2681,7 @@
|
|
|
2681
2681
|
{
|
|
2682
2682
|
"name": "start_job",
|
|
2683
2683
|
"title": "Start managed job",
|
|
2684
|
-
"description": "Durably accept a detached argv-based job with ordered steps, job-scoped temporary files, guaranteed-attempt finally steps, and optional depends_on links to existing managed jobs. The independent local runner continues if the MCP connection disappears. Dependencies are evaluated from durable managed-job terminal state rather than blind artifact polling: while an upstream job is active the dependent job remains queued in current_phase=dependency_wait without spawning its main child, then proceeds after all dependencies succeed or terminates with error_class=dependency_failed if an upstream later fails. Staged or already-failed dependencies are rejected before acceptance. On macOS, an authorized remote account job keeps a runner-owned idle-sleep assertion only after the runner ownership claim is confirmed and persisted account ownership is validated, then through dependency wait, admission, steps, cleanup, and terminal persistence; local managed jobs do not acquire this remote-continuity assertion, and it does not override explicit sleep or lid-close behavior. Use an idempotency_key to deduplicate an uncertain retry while the original job record is retained; eviction/retention expiry ends that deduplication window, so this is not permanent exactly-once execution. Use {{temp:name}}, {{resource:name}}, env_resources, or stdin_resource; registered resource contents never enter MCP arguments. Each main/finally step defaults to 600 seconds and may explicitly request up to 21600 seconds (six hours); use this managed-job surface when one continuous command legitimately needs more than the 600-second durable one-step process limit. For one coherent non-interactive workflow, prefer one multi-step managed job over a chain of one-step process carriers so host-visible event density stays bounded. Do not split a long task merely to satisfy an assumed host deadline, and do not use the user as a polling clock.",
|
|
2684
|
+
"description": "Durably accept a detached argv-based job with ordered steps, job-scoped temporary files, guaranteed-attempt finally steps, and optional depends_on links to existing managed jobs. The independent local runner continues if the MCP connection disappears. Dependencies are evaluated from durable managed-job terminal state rather than blind artifact polling: while an upstream job is active the dependent job remains queued in current_phase=dependency_wait without spawning its main child, then proceeds after all dependencies succeed or terminates with error_class=dependency_failed if an upstream later fails. A transient dependency-state read classified as permission_denied, identity_changed, or generic resource_unavailable receives a fixed 45-second recovery grace; one successful read clears that grace, while persistent unavailability fails closed as dependency_unavailable and missing/integrity/witness-invalid evidence remains immediate failure. Staged or already-failed dependencies are rejected before acceptance. On macOS, an authorized remote account job keeps a runner-owned idle-sleep assertion only after the runner ownership claim is confirmed and persisted account ownership is validated, then through dependency wait, admission, steps, cleanup, and terminal persistence; local managed jobs do not acquire this remote-continuity assertion, and it does not override explicit sleep or lid-close behavior. Use an idempotency_key to deduplicate an uncertain retry while the original job record is retained; eviction/retention expiry ends that deduplication window, so this is not permanent exactly-once execution. Use {{temp:name}}, {{resource:name}}, env_resources, or stdin_resource; registered resource contents never enter MCP arguments. Each main/finally step defaults to 600 seconds and may explicitly request up to 21600 seconds (six hours); use this managed-job surface when one continuous command legitimately needs more than the 600-second durable one-step process limit. For one coherent non-interactive workflow, prefer one multi-step managed job over a chain of one-step process carriers so host-visible event density stays bounded. Do not split a long task merely to satisfy an assumed host deadline, and do not use the user as a polling clock.",
|
|
2685
2685
|
"availability": "write+direct-exec",
|
|
2686
2686
|
"annotations": {
|
|
2687
2687
|
"readOnlyHint": false,
|
|
@@ -2889,7 +2889,7 @@
|
|
|
2889
2889
|
{
|
|
2890
2890
|
"name": "list_jobs",
|
|
2891
2891
|
"title": "List managed jobs",
|
|
2892
|
-
"description": "List
|
|
2892
|
+
"description": "List detached managed-job recovery state without returning step output. The primary jobs array remains capped at 50 and prioritizes unreadable, active, staged, and durable terminal recovery state before transient one-step helper history so helper churn cannot hide an older recoverable long-running job or its terminal result. If a recent one-step process result is still retained inside the fixed 30-minute/16-result recovery reserve but omitted from that primary window, recent_process_recovery returns up to 16 additional authority-visible public job handles so a later host turn can recover the job_id and continue with read_job. This is bounded recovery discovery only: it is not an MCP terminal-result replay/session store and must not be used for polling. The durable retained-state store is larger than these response windows. Owner/local callers also receive coarse retained-state capacity including durable-versus-transient terminal counts plus recent job-creation/churn aggregates without internal retired filenames, filesystem identities, argv, paths, or output.",
|
|
2893
2893
|
"availability": "always",
|
|
2894
2894
|
"annotations": {
|
|
2895
2895
|
"readOnlyHint": true,
|
package/src/worker/index.ts
CHANGED
|
@@ -60,7 +60,7 @@ import {
|
|
|
60
60
|
closeWebSocketQuietly, daemonErrorCloseCode, isObjectRecord, rejectDaemonMessage,
|
|
61
61
|
sendWebSocketQuietly, trySendWebSocket,
|
|
62
62
|
} from "./websocket-protocol.ts";
|
|
63
|
-
const SERVER_VERSION = "3.0.0-beta.
|
|
63
|
+
const SERVER_VERSION = "3.0.0-beta.144";
|
|
64
64
|
const MCP_SERVER_INFO = mcpServerInfo(SERVER_VERSION);
|
|
65
65
|
const MAX_DAEMON_MESSAGE_BYTES = 8 * 1024 * 1024;
|
|
66
66
|
const DAEMON_RECONNECT_GRACE_MS = relayContract.reconnectGraceMs; const NEW_CALL_RECONNECT_GRACE_MS = relayContract.newCallReconnectGraceMs;
|