machine-bridge-mcp 3.0.0-beta.141 → 3.0.0-beta.146
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +30 -0
- package/README.md +1 -1
- package/browser-extension/manifest.json +1 -1
- package/docs/ARCHITECTURE.md +4 -4
- package/docs/AUDIT.md +50 -0
- package/docs/LOGGING.md +2 -2
- package/docs/MANAGED_JOBS.md +5 -5
- package/docs/OPERATIONS.md +7 -7
- package/docs/TESTING.md +5 -5
- package/docs/TOOL_REFERENCE.md +3 -3
- package/package.json +1 -1
- package/src/local/durable-process-initial-settlement.mjs +5 -7
- package/src/local/macos-idle-sleep-assertion.mjs +7 -2
- package/src/local/managed-job-dependencies.mjs +15 -2
- package/src/local/managed-job-listing.mjs +3 -2
- package/src/local/managed-job-recovery-listing.mjs +18 -0
- package/src/local/relay-connection.mjs +7 -3
- package/src/local/relay-diagnostics.mjs +34 -0
- package/src/local/relay-peer-diagnostics.mjs +47 -0
- package/src/local/remote-activity-idle-sleep-guard.mjs +2 -2
- package/src/local/runtime-diagnostic-projection.mjs +1 -0
- package/src/local/runtime-diagnostic-state.mjs +1 -0
- package/src/local/runtime-diagnostics.mjs +2 -2
- package/src/local/runtime-tool-handlers.mjs +4 -1
- package/src/local/system-sleep-diagnostics.mjs +68 -0
- package/src/shared/server-metadata.json +6 -6
- package/src/shared/tool-catalog.json +3 -3
- package/src/worker/daemon-relay-diagnostics.ts +103 -0
- package/src/worker/index.ts +1 -1
- package/src/worker/server-info-tool-delivery.ts +2 -0
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,35 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 3.0.0-beta.146 - 2026-08-27
|
|
4
|
+
|
|
5
|
+
- Fix the hosted `start_job` initial-settlement result boundary exposed by the live beta.145 candidate. The shared settlement helper had explicitly copied process-carrier-only acceptance fields onto ordinary managed-job results; those properties were absent on `start_job` and therefore became JavaScript `undefined`. Machine Bridge's real tool-result normalization intentionally rejects `undefined` as non-JSON, so the job was durably accepted and could complete while the initiating hosted call returned a non-retryable `internal_error`. The helper now relies on the original accepted-object spread to preserve only fields that actually exist instead of manufacturing absent properties.
|
|
6
|
+
- Extend the regression through `normalizeToolResult()` rather than testing the bound handler result alone. This reproduces the exact boundary missed by beta.145: a short hosted `start_job` must coalesce terminal state, preserve its recovery envelope, and remain valid JSON with no unsupported value. Existing process-carrier initial settlement remains unchanged and still returns its defined `execution_mode`, `source_tool`, execution-timeout, and retry-safety fields.
|
|
7
|
+
- Reject beta.145 after live activation even though its exact candidate passed 131/131 frozen verification, install-only validation, service/Worker activation, and the activated-package OAuth canary. Two independent live hosted `start_job` probes returned `internal_error`; idempotency evidence and `read_job` proved the underlying canary job had actually executed successfully. No beta.145 candidate acceptance is recorded. Package/runtime identity advances to beta.146; hosted tool schema generation remains 15 because the public gen15 contract is unchanged and beta.146 repairs its implementation. Fresh frozen verification, candidate/install-only proof, guarded activation, activated-package canary, live short-`start_job` verification, acceptance, and exact-head provider gates are required. npm publication remains separately owner-authorized.
|
|
8
|
+
|
|
9
|
+
## 3.0.0-beta.145 - 2026-08-27
|
|
10
|
+
|
|
11
|
+
- Preserve the causal evidence for repeated awake WebSocket rebuilds instead of retaining only the latest reconnect. The local relay now keeps a newest-first `recent_outages` ring capped at eight completed reconnect episodes with bounded timestamps/durations, close/error classes, prior-ready duration/silence, coarse application-route class, and connection-stage timings. The daemon hello sanitizes that history, and Worker ready promotion synthesizes the just-completed current episode that could not yet have appeared in the pre-ready hello. Pong/application-confirmation near misses that recover without replacing WSS remain heartbeat diagnostics and are deliberately excluded from the completed-outage ring.
|
|
12
|
+
- Reduce another reproduced host-event-density amplifier without shortening durable work. Hosted `start_job` now shares the existing two-second optional managed-job initial-settlement path already used by one-step durable process carriers. A short job that reaches terminal state inside the original response returns its result with `follow_up_read_required=false`; an active job keeps the same `job_id`/recovery envelope with `follow_up_read_required=true`. Dependency waiting, thirty-minute resource admission, six-hour step ceilings, reconnect/replay safety, and local/stdio behavior remain unchanged.
|
|
13
|
+
- Supersede beta.144 because the owner-reported interruption class remains blocking and these changes alter packaged relay diagnostics, hosted `start_job` delivery semantics, tool descriptions, and shipped documentation. Package/runtime identity advances to beta.145 and hosted tool schema generation advances to 15. Fresh frozen verification, candidate/install-only proof, guarded activation, activated-package OAuth canary, live observation, acceptance, and exact-head provider gates are required before GitHub prerelease publication; any later npm publication still requires separate owner authorization, and a published beta.145 activation would restart the major-prerelease soak interval.
|
|
14
|
+
|
|
15
|
+
## 3.0.0-beta.144 - 2026-08-27
|
|
16
|
+
|
|
17
|
+
- Fix a real Windows managed-job dependency recovery defect exposed by exact-main provider CI after beta.143 acceptance. Runner-exit reconciliation already treats transient `permission_denied`/conflict/timeout/resource-unavailable failures as retryable, but a concurrently waiting downstream previously treated any secure upstream status-read error as permanent `dependency_unavailable`. Dependency polling now gives only `permission_denied`, `identity_changed`, and generic `resource_unavailable` reads a fixed 45-second monotonic recovery grace; a successful secure read clears that grace immediately, persistent unavailability still fails closed, and missing/integrity/witness-invalid evidence remains non-retryable.
|
|
18
|
+
- Correct the Worker integration fixture race that obscured the first exact-main diagnosis. Eight tests started an MCP HTTP request before registering the WebSocket `tool_call` listener, so a fast relay dispatch could arrive before the harness was listening and be lost permanently; increasing the waiter from five to ten seconds could not fix that ordering bug. The fixture now registers the waiter before triggering the request. Three local Worker integration runs and hosted Ubuntu/full passed with the listener-first ordering.
|
|
19
|
+
- Supersede beta.143 release evidence because the dependency-recovery repair changes packaged runtime bytes and hosted orchestration semantics. The beta.143 acceptance record is removed, package/runtime identity advances to beta.144, and tool schema generation advances to 14. Fresh frozen verification, exact candidate preparation/install-only proof, guarded activation, activated-package OAuth canary, live observation, acceptance, and exact-head provider gates are required before GitHub prerelease publication. npm publication remains separately owner-authorized for the exact beta.144 version.
|
|
20
|
+
|
|
21
|
+
## 3.0.0-beta.143 - 2026-08-27
|
|
22
|
+
|
|
23
|
+
- Correct interruption diagnosis after live beta.142 evidence separated host suspension from awake transport resets. A programmatic correlation of fifteen recovered August 23 relay outages with bounded macOS `pmset` sleep intervals found thirteen sleep-dominated episodes: 32,980 of 33,390 outage seconds (98.8%) occurred while the host was suspended. `diagnose_runtime.runtime.relay_outage_analysis` now intersects the most recently completed close-to-ready relay interval with the same bounded sleep history and reports only coarse timing/overlap fields. `majority_system_sleep_overlap` therefore prevents a sleep-dominated `connection_reset` from being promoted into an unsupported independent-network root cause, while an awake reset with no sleep overlap remains real transport evidence without guessing which external layer failed.
|
|
24
|
+
- Strengthen the existing macOS remote-work power assertion on AC without extending its ownership lifetime. The fixed runtime/remote-runner primitive now launches `/usr/bin/caffeinate -i -s -w <owner-pid>` instead of `-i -w`: `-i` retains Idle Sleep prevention and Apple's `-s` assertion adds `PreventSystemSleep` only while on AC power. A live three-second probe on the affected Mac observed both assertions concurrently. The thirty-minute relay inactivity grace, process-session ownership, detached account-job ownership, fail-open error handling, and explicit/lid-close/battery limitations are unchanged; diagnostics report only whether the fixed child requests the AC-only system-sleep assertion.
|
|
25
|
+
- Keep the external-host boundary explicit. The beta.142 `UNKNOWN / TaskGroup` episode occurred during an approximately nine-second awake relay `1006/connection_reset`, while the same durable job continued and was recovered by the same `job_id`. Worker pending-call deadlines already preserve the original absolute timeout across reconnect, and the public MCP response path already emits an immediate SSE priming frame plus five-second heartbeats. This tree does not invent a host timeout or shorten the proven forty-second hosted `read_job` default without evidence. Tool schema generation advances to 13 because owner-visible `diagnose_runtime` result semantics and guidance changed. npm registry publication remains a separate explicit owner-authorization boundary.
|
|
26
|
+
|
|
27
|
+
## 3.0.0-beta.142 - 2026-08-27
|
|
28
|
+
|
|
29
|
+
- Make retained one-step process recovery discoverable after a real host/tool boundary. A saturated live 512-state store can legitimately contain 496 durable terminal jobs plus the beta.141 sixteen-result transient recovery reserve; the existing durable-first `list_jobs.jobs` window may then contain no recent process helpers even though a known helper `job_id` is still readable. `list_jobs` now keeps that primary window unchanged and adds `recent_process_recovery`, capped at 16 authority-visible public job handles for recent transient terminal results that are still inside the existing thirty-minute recovery grace but omitted from `jobs`.
|
|
30
|
+
- Preserve the request-scoped MCP boundary. The new recovery projection carries public job status only: no step output, argv, path, internal `retention_class`, conversation identity, terminal-result replay, or cross-request response session is added. It is inventory for recovering a lost `job_id`; known work still continues through `read_job`, and `list_jobs` remains a non-polling surface. The 512-state retention cap, 50-record primary inventory, 16-result/thirty-minute transient retention reserve, durable-first ordering, dependency pinning, and eviction priorities are unchanged.
|
|
31
|
+
- Advance hosted tool schema generation to 12 because `list_jobs` public result semantics and tool guidance changed. These shipped runtime, catalog, tests, and documentation changes supersede beta.141 release evidence and require a fresh prerelease verification/candidate/activation sequence before any publication. npm registry publication remains separately owner-authorized.
|
|
32
|
+
|
|
3
33
|
## 3.0.0-beta.141 - 2026-08-26
|
|
4
34
|
|
|
5
35
|
- Preserve immediate recovery evidence for remote one-step process carriers under a saturated 512-state managed-job store. The previous eviction order always discarded `transient_process` terminal results before any ordinary durable terminal history; with 510 durable terminals already retained, short `exec_command` helpers could therefore return a recoverable job ID and then become `not_found` before the next `read_job`. Beta.141 reserves at most the newest 16 otherwise-removable transient results for thirty minutes. The incoming transient counts against that bound, older/excess transient history is still evicted first, ordinary durable history is next, dependency-protected records remain pinned, and the hard 512-state cap is unchanged.
|
package/README.md
CHANGED
|
@@ -193,7 +193,7 @@ For stateful GUI trajectories, owner/full callers can use the higher-level `comp
|
|
|
193
193
|
|
|
194
194
|
## Durable work and local resources
|
|
195
195
|
|
|
196
|
-
Remote request-owned foreground work uses the hosted reply-safe budgets described above; configurable browser/application calls may explicitly request at most 45 seconds, while remote `exec_command`, `run_process`, and `run_local_command` are durable one-step jobs with a 10-second acceptance envelope and an independent 1–600-second child execution budget after admission. A continuous process that legitimately needs more than 600 seconds must use `start_job`: managed-job main/finally steps default to 600 seconds and may explicitly request up to 21,600 seconds (six hours), with resource admission occurring before that execution timer begins. The Worker retains separate settlement ownership for five additional seconds, but neither that margin nor its internal stream metrics prove that an external MCP host consumed the terminal frame. Keep unrelated mutations and validation independently terminal, but batch one coherent non-interactive command sequence into a repository umbrella command or multi-step `start_job` instead of creating one host-visible one-step job per tiny probe. A timeout is a protocol result, not proof that descendant cleanup has already completed; a remote owner can inspect `diagnose_runtime.runtime.processes`, while local stdio exposes `server_info.runtime.processes`. Non-owner accounts receive authority-scoped readiness rather than machine-wide process activity. Remote process sessions are for interactive stdin or incremental output, not a substitute for ordinary durable work: hosted `read_process` reports `status_polling_mode=paced_followup`, caps the actual output/exit blocking wait at one second, and paces a repeated would-block read inside the fifteen-second cooldown within that same MCP call until output/exit or the cooldown boundary instead of returning a rapid running checkpoint. When the current task needs more output or terminal state, the same session may be read again in the same assistant response without busy-looping. Non-interactive work should use durable `run_process`/`read_job`; multi-step, cleanup-sensitive, or daemon-restart-surviving workflows should use managed jobs, which persist ordered argv steps and `finally_steps` under owner-only local state and continue across an MCP disconnect. Durable acceptance does not force a hosted-turn handoff: active relay-origin `read_job` reports `status_polling_mode=bounded_followup` and `host_turn_handoff_recommended=false`. Its hosted default is a 40-second server-side long-poll, so an unchanged long job occupies one bounded live MCP response rather than forcing rapid host-side checkpoints. Terminal settlement returns on the next bounded five-second internal poll; nonterminal status/phase/dependency progress is coalesced for at least 30 seconds by default, and `current_step`-only churn does not wake the hosted call. `wait_ms=0` is available only when an immediate checkpoint is actually wanted, and an explicitly capable client may request up to five minutes. At the default, a synthetic 100-minute unchanged job has an anti-amplification ceiling of 150 status reads, while continuously changing nonterminal progress is separately bounded by the 30-second coalescing floor; those are density estimates rather than proof of aggregate same-response host lifetime. A known job may be followed through paced same-response `read_job` calls while those calls continue to be accepted and the task still needs the result; after an actual host/tool boundary, later recovery must continue from the same `job_id`. Completed one-step process carriers are lower-priority terminal retention than explicit managed jobs, so removable helper history is reclaimed first under the 512-state durable cap
|
|
196
|
+
Remote request-owned foreground work uses the hosted reply-safe budgets described above; configurable browser/application calls may explicitly request at most 45 seconds, while remote `exec_command`, `run_process`, and `run_local_command` are durable one-step jobs with a 10-second acceptance envelope and an independent 1–600-second child execution budget after admission. A continuous process that legitimately needs more than 600 seconds must use `start_job`: managed-job main/finally steps default to 600 seconds and may explicitly request up to 21,600 seconds (six hours), with resource admission occurring before that execution timer begins. The Worker retains separate settlement ownership for five additional seconds, but neither that margin nor its internal stream metrics prove that an external MCP host consumed the terminal frame. Keep unrelated mutations and validation independently terminal, but batch one coherent non-interactive command sequence into a repository umbrella command or multi-step `start_job` instead of creating one host-visible one-step job per tiny probe. A timeout is a protocol result, not proof that descendant cleanup has already completed; a remote owner can inspect `diagnose_runtime.runtime.processes`, while local stdio exposes `server_info.runtime.processes`. Non-owner accounts receive authority-scoped readiness rather than machine-wide process activity. Remote process sessions are for interactive stdin or incremental output, not a substitute for ordinary durable work: hosted `read_process` reports `status_polling_mode=paced_followup`, caps the actual output/exit blocking wait at one second, and paces a repeated would-block read inside the fifteen-second cooldown within that same MCP call until output/exit or the cooldown boundary instead of returning a rapid running checkpoint. When the current task needs more output or terminal state, the same session may be read again in the same assistant response without busy-looping. Non-interactive work should use durable `run_process`/`read_job`; multi-step, cleanup-sensitive, or daemon-restart-surviving workflows should use managed jobs, which persist ordered argv steps and `finally_steps` under owner-only local state and continue across an MCP disconnect. Durable acceptance does not force a hosted-turn handoff: active relay-origin `read_job` reports `status_polling_mode=bounded_followup` and `host_turn_handoff_recommended=false`. Its hosted default is a 40-second server-side long-poll, so an unchanged long job occupies one bounded live MCP response rather than forcing rapid host-side checkpoints. Terminal settlement returns on the next bounded five-second internal poll; nonterminal status/phase/dependency progress is coalesced for at least 30 seconds by default, and `current_step`-only churn does not wake the hosted call. `wait_ms=0` is available only when an immediate checkpoint is actually wanted, and an explicitly capable client may request up to five minutes. At the default, a synthetic 100-minute unchanged job has an anti-amplification ceiling of 150 status reads, while continuously changing nonterminal progress is separately bounded by the 30-second coalescing floor; those are density estimates rather than proof of aggregate same-response host lifetime. A known job may be followed through paced same-response `read_job` calls while those calls continue to be accepted and the task still needs the result; after an actual host/tool boundary, later recovery must continue from the same `job_id`. Completed one-step process carriers are lower-priority terminal retention than explicit managed jobs, so removable helper history is reclaimed first under the 512-state durable cap. `list_jobs.jobs` keeps a 50-record durable-first primary window; if a recent one-step terminal remains inside the fixed thirty-minute/16-result recovery reserve but is omitted from that window, `recent_process_recovery` returns up to 16 additional authority-visible public job handles so a later host turn can recover the `job_id` and continue with `read_job`. That secondary projection includes no step output or internal retention metadata and remains inventory rather than MCP replay/session state or a polling surface. Owner/local `capacity` diagnostics expose only coarse `durable_terminal` and `transient_terminal` counts; this improves recovery visibility without pretending that Worker acknowledgement proves an external host rendered the final assistant message. Long cross-job workflows can declare `depends_on`: the dependent job remains pre-execution `queued/dependency_wait` without spawning its main child until all upstream jobs succeed, and an upstream failure settles `dependency_failed` instead of leaving a file-poll loop waiting for an artifact that can never appear. Active/staged dependency plans pin referenced retained results until the dependency-bearing plan is terminal. A valid `job_id` that is no longer retained returns typed `not_found`; that absence is not proof that its underlying side effect never executed. `list_jobs` remains inventory rather than a substitute polling loop, and `server_info`/`diagnose_runtime` remain diagnostic surfaces rather than alternate wait channels. Elapsed minutes are not evidence that an external host deadline is near; return the durable recovery identifier for a later turn only after an actual host/tool boundary is observed, external input or authorization is required, or the user explicitly requested a checkpoint.
|
|
197
197
|
|
|
198
198
|
On macOS, authorized remote activity uses a bounded idle-sleep assertion so ordinary system Idle Sleep does not suspend an active remote workflow. Relay handlers share the assertion for their execution lifetime plus a fixed thirty-minute rolling inactivity grace; each new authorized remote activity cancels a pending release and restarts the full grace after the last concurrent handler settles. An admitted remote process session extends daemon-side ownership until its child settles, and an account-backed managed-job runner owns a runner-bound assertion from confirmed claim through terminal persistence. Local managed jobs do not acquire the remote-continuity assertion. These protections do not override explicit sleep or lid-close behavior.
|
|
199
199
|
|
|
@@ -30,6 +30,6 @@
|
|
|
30
30
|
"action": {
|
|
31
31
|
"default_title": "Machine Bridge Browser"
|
|
32
32
|
},
|
|
33
|
-
"version_name": "3.0.0-beta.
|
|
33
|
+
"version_name": "3.0.0-beta.146",
|
|
34
34
|
"key": "MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAxryYkpZhq8+VAQLHcGS9BAHQcyKX8RHGIpIwvtIVRU/rcOcE0bNdnM0aZJ/h6xWQsGDHlhvjT2+1aJaAn/9k8473BRWajzVXld961CdHYVFVHoce2hHiSJ0xydWrHMMZhAm0mN0UzjEpgZ0tMw209efcZHIvSwuxhteZMRy4kyiVjwFlOf5oXFCxRuCJnPj3AK9CmCf4XgEBuPIJ0TZmjGHOOdBvJmbCNnAWXYEo5/mf7MfCGhV4IJ1hNuhpoNQfOFKMUcw9/v/IpT62XpfXdGYTfGYCmCjC+gntK1spbkr2P4/2+sYMQtLpse71mpSNGXfcf3abU55Vpn+gncSxRQIDAQAB"
|
|
35
35
|
}
|
package/docs/ARCHITECTURE.md
CHANGED
|
@@ -32,11 +32,11 @@ A canonical workspace receives an independent profile, Worker name, secret set,
|
|
|
32
32
|
- `workspace-file-service.mjs` and `git-service.mjs` own canonical filesystem/Git operations; `file-mutation-coordinator.mjs` supplies the shared resolved-path conflict identity and serializes only conflicting paths, registers every path in a multi-file transaction before waiting, and holds those reservations until the filesystem callback settles. `workspace-file-transaction.mjs` owns flushed staging, no-overwrite target publication, whole-file compare-and-swap, patch rollback, and cleanup causality. A late patch target becomes `conflict/target_appeared`; a rollback failure after any user-file commit is a public non-retryable `patch_recovery_incomplete` settlement directing the caller to inspect affected paths, while staging-only cleanup failure remains an internal cleanup error and post-commit artifact cleanup remains a warning rather than a false rollback claim;
|
|
33
33
|
- `process-contract.mjs` owns argv shape/size validation; `process-tree-signal.mjs`, `process-tree-supervisor.mjs`, `process-tree-snapshot.mjs`, and `process-tree-ownership.mjs` separate cross-platform signaling, asynchronous escalation, bounded process-group observation, and PID/start-time ownership; `process-execution.mjs` and `process-sessions.mjs` own one-shot and interactive execution; `process-nonreplayable-settlement.mjs` maps fixed internal UI-launch/native-input helpers from definite pre-spawn failure to non-retryable unknown settlement after process start; and `process-tracker.mjs` retains active and draining process ownership until close;
|
|
34
34
|
- `shared/tool-call-capacity.mjs` defines the control-tool set and generic admission algebra; local `call-capacity.mjs` and Worker `pending-call-capacity.ts` apply it independently, while `runtime-reporting.mjs` builds privacy-aware runtime and project snapshots;
|
|
35
|
-
- `runtime-diagnostics.mjs` owns fixed local probes and their stable interpretation, while `runtime-diagnostic-state.mjs` projects privacy-safe control-plane state for remote diagnosis. On macOS, `system-sleep-diagnostics.mjs` is the only power-history parser used by that surface: it runs a fixed bounded `pmset` pipeline, returns only recent sleep interval timestamps/durations/coarse reasons,
|
|
35
|
+
- `runtime-diagnostics.mjs` owns fixed local probes and their stable interpretation, while `runtime-diagnostic-state.mjs` projects privacy-safe control-plane state for remote diagnosis. On macOS, `system-sleep-diagnostics.mjs` is the only power-history parser used by that surface: it runs a fixed bounded `pmset` pipeline, returns only recent sleep interval timestamps/durations/coarse reasons, correlates an event-loop stall only when both its resume time and duration match an operating-system sleep interval within the fixed tolerance, and independently intersects the most recently completed relay disconnect interval with those same sleep intervals. A majority-sleep relay overlap means host suspension dominated the observed outage and a retained socket reset is aftermath evidence rather than proof of a separate network root cause; an unmatched pause or awake relay reset remains unassigned instead of being guessed into another cause;
|
|
36
36
|
- `runtime-capabilities.mjs` composes agent, application, browser, and effective-policy-filtered routing results, while `execution-routing.mjs` owns bounded set-level route scoring, ambiguity, fallbacks, and advisory tool projection;
|
|
37
37
|
- `runtime-tool-handlers.mjs` owns catalog-to-handler registration;
|
|
38
38
|
- `runtime-relay.mjs` owns relay construction and inbound envelope normalization; `relay-call-recovery.mjs` owns bounded disconnect grace, authoritative resumed-call reconciliation, and same-daemon redelivery orchestration; `relay-result-retention.mjs` owns the bounded completed-but-unacknowledged result ledger plus the fail-closed automatic-redelivery proof; and `relay-recovery-admission.mjs` owns recovery-capacity accounting and its privacy-safe diagnostic projection;
|
|
39
|
-
- `macos-idle-sleep-assertion.mjs` is the single adapter boundary for the fixed macOS power primitive: `/usr/bin/caffeinate -i -w <owner-pid>` is spawned with `shell: false`, ignored stdio, and a process-lifetime owner binding; failures are best-effort and expose only a coarse error class. `remote-activity-idle-sleep-guard.mjs` composes that primitive for daemon-side relay activity. An authorized, schema-valid relay tool call begins activity only after policy/account/operation authorization and argument validation succeed; concurrent handlers share one assertion for their full execution lifetime, the fixed thirty-minute inactivity grace begins only after the last handler settles, and a new handler cancels any pending release timer so the full rolling grace restarts after the new activity settles. `process-session-remote-activity.mjs` extends the same daemon assertion beyond the `start_process` handler only after resource admission succeeds and releases it when the remote child settles, including startup failure. Remote account managed-job runners do not depend on daemon ownership: after the runner claim is confirmed and persisted ownership identifies an account-backed job, `job-runner.mjs` holds its own assertion bound to the runner PID across recovery handoff, resource admission, steps, `finally_steps`, and terminal persistence, then releases it in top-level `finally`. Local managed jobs do not acquire this remote-continuity assertion. Runtime shutdown terminates process sessions before releasing the daemon assertion. The thirty-minute relay inactivity grace is deliberately fixed rather than depending on shell environment inheritance that launchd does not persist as service configuration or treating relay heartbeats as user activity. These guards
|
|
39
|
+
- `macos-idle-sleep-assertion.mjs` is the single adapter boundary for the fixed macOS power primitive: `/usr/bin/caffeinate -i -s -w <owner-pid>` is spawned with `shell: false`, ignored stdio, and a process-lifetime owner binding; failures are best-effort and expose only a coarse error class. The `-i` assertion retains Idle Sleep prevention, while `-s` requests the stronger PreventSystemSleep assertion that macOS documents as effective only on AC power; diagnostics report only that the fixed child requests that AC-only protection, not the current power source or a guarantee against explicit/lid-close sleep. `remote-activity-idle-sleep-guard.mjs` composes that primitive for daemon-side relay activity. An authorized, schema-valid relay tool call begins activity only after policy/account/operation authorization and argument validation succeed; concurrent handlers share one assertion for their full execution lifetime, the fixed thirty-minute inactivity grace begins only after the last handler settles, and a new handler cancels any pending release timer so the full rolling grace restarts after the new activity settles. `process-session-remote-activity.mjs` extends the same daemon assertion beyond the `start_process` handler only after resource admission succeeds and releases it when the remote child settles, including startup failure. Remote account managed-job runners do not depend on daemon ownership: after the runner claim is confirmed and persisted ownership identifies an account-backed job, `job-runner.mjs` holds its own assertion bound to the runner PID across recovery handoff, resource admission, steps, `finally_steps`, and terminal persistence, then releases it in top-level `finally`. Local managed jobs do not acquire this remote-continuity assertion. Runtime shutdown terminates process sessions before releasing the daemon assertion. The thirty-minute relay inactivity grace is deliberately fixed rather than depending on shell environment inheritance that launchd does not persist as service configuration or treating relay heartbeats as user activity. These guards strengthen AC continuity without claiming to defeat explicit sleep, lid-close sleep, or battery-mode system sleep;
|
|
40
40
|
- `runtime-paths.mjs` owns runtime-directory creation, containment checks, and error-path redaction;
|
|
41
41
|
- `runtime-resource-service.mjs` owns registered-resource lookup, bounded binary/UTF-8 reads for browser/application injection, and SSH-resource registration/result projection;
|
|
42
42
|
- `security-audit-log.mjs` owns only the bounded main-thread queue and cached initializing/health projection; `security-audit-worker.mjs`, `security-audit-storage.mjs`, `security-audit-dispatch.mjs`, and `security-audit-warning.mjs` isolate startup verification, all disk/hash work, batch persistence, privacy projection, and warning suppression from result delivery;
|
|
@@ -87,7 +87,7 @@ See [Local application and browser automation](LOCAL_AUTOMATION.md).
|
|
|
87
87
|
|
|
88
88
|
### Managed job runner
|
|
89
89
|
|
|
90
|
-
`ManagedJobManager` persists bounded per-workspace job envelopes below the owner-only profile directory. Managed-job active and terminal lifecycle classifications have one shared source in `managed-job-terminal.mjs`; manager reconciliation, retention, detached-runner fatal settlement, and production full-access diagnostics consume that source instead of maintaining parallel status lists. `start_job` validates the complete plan, snapshots referenced resource metadata/hashes, writes an owner-only plan/status, and launches `job-runner.mjs` as a detached process with runner-level logs redirected to owner-only files. `stage_job` performs the same acceptance validation but writes a non-running, review-only `staged` envelope; staged records have no promotion/approval execution path, so execution requires a separate trusted `start_job` request or an explicit local `job submit` plan. `managed-job-dependency-admission.mjs` binds optional `depends_on` references to same-authority durable job identities and rejects staged or already-failed dependencies before acceptance. `managed-job-dependencies.mjs` then keeps a dependent runner in pre-execution `queued/dependency_wait` state until all upstream jobs succeed, or settles it with `dependency_failed` when one later fails; no dependent main child or process resource lease exists during that wait. `managed-job-relaunch.mjs` preserves
|
|
90
|
+
`ManagedJobManager` persists bounded per-workspace job envelopes below the owner-only profile directory. Managed-job active and terminal lifecycle classifications have one shared source in `managed-job-terminal.mjs`; manager reconciliation, retention, detached-runner fatal settlement, and production full-access diagnostics consume that source instead of maintaining parallel status lists. `start_job` validates the complete plan, snapshots referenced resource metadata/hashes, writes an owner-only plan/status, and launches `job-runner.mjs` as a detached process with runner-level logs redirected to owner-only files. `stage_job` performs the same acceptance validation but writes a non-running, review-only `staged` envelope; staged records have no promotion/approval execution path, so execution requires a separate trusted `start_job` request or an explicit local `job submit` plan. `managed-job-dependency-admission.mjs` binds optional `depends_on` references to same-authority durable job identities and rejects staged or already-failed dependencies before acceptance. `managed-job-dependencies.mjs` then keeps a dependent runner in pre-execution `queued/dependency_wait` state until all upstream jobs succeed, or settles it with `dependency_failed` when one later fails; no dependent main child or process resource lease exists during that wait. Because Windows atomic replacement/runner-exit recovery can make an otherwise valid status file transiently unreadable, dependency polling gives only `permission_denied`, `identity_changed`, and generic `resource_unavailable` reads a fixed 45-second monotonic recovery grace. One successful secure read clears that grace; persistent unavailability still fails closed as `dependency_unavailable`, while missing, corrupt, witness-mismatched, staged, or otherwise invalid evidence is never converted into a retry. `managed-job-relaunch.mjs` preserves the pre-execution distinction across a dead dependency-wait runner by restarting the original `dependency_wait` job rather than converting it to cleanup-only recovery. `managed-job-retention.mjs` owns staged expiry timing and seven-day terminal retention; `managed-job-capacity.mjs` bounds durable retained state at 512 while the public `list_jobs.jobs` primary response window remains capped at 50 records. Recent one-step terminal jobs that are still inside the fixed thirty-minute/16-result recovery reserve but omitted from that durable-first primary window are projected separately through `recent_process_recovery`, capped at 16 authority-visible public job handles and without step output or internal retention metadata. This recovery-discovery projection does not create MCP replay/session state or change request ownership. Completed one-step process carriers are marked internally as `transient_process`. Under hard capacity pressure, retention preserves at most the newest 16 otherwise-removable transient results that completed within the last thirty minutes so an accepted remote helper cannot normally lose its recovery evidence before immediate `read_job` follow-up. The incoming transient reservation counts against that 16-result window: excess or older transient history is reclaimed first, then the oldest ordinary durable terminal history, and only then a still-recent protected transient when no other removable record exists. This recovery reserve does not enlarge the 512-state cap and never overrides dependency protection. `managed-job-dependency-retention.mjs` pins terminal records referenced by active/staged dependency plans and fails closed if that protection cannot be read safely. `managed-job-directory.mjs` maps a syntactically valid but no-longer-retained job ID to fixed non-retryable `not_found`; that absence is recovery-evidence loss rather than proof that the underlying operation never executed. `managed-job-terminal-maintenance.mjs` owns post-settlement evidence validation and artifact scrubbing; `managed-job-directory-generation.mjs` binds whole-directory retirement to the exact filesystem generation that retention inspected. Retirement first revalidates the full observed generation, atomically renames that directory to an internal `retired_job_*` name carrying its device/inode identity, revalidates the moved object, and only then recursively deletes it. The retirement namespace deliberately does not match the public `MANAGED_JOB_ID` grammar, so list/read/lock scans cannot reinterpret internal cleanup state as an ordinary job. A crash after rename therefore leaves recognizable state rather than an anonymous orphan: a later maintenance pass reclaims it only when the encoded generation still matches, while a type mismatch, generation mismatch, or unreadable retired entry is projected into active-state inventory as a privacy-bounded `retired_managed_job`/`unreadable` blocker without exposing the internal filename or filesystem identity. Before destructive terminal cleanup, status/result must describe the same directory job ID, terminal state, and `finished_at` generation; the degraded `result_persisted=false` form must carry an explicit terminal-record error class. Corrupt terminal evidence is therefore retained as unreadable state and also blocks state removal instead of authorizing plan/runtime scrubbing or capacity eviction. Expiry is a real per-job state transition: it acquires `transition.lock`, re-reads the staged state, and commits through the same result-first terminal persistence path as cancellation/runner settlement. Seven-day retention is measured from terminal `finished_at`, not an older directory mtime. Admission may evict only safely removable terminal records and never active, staged, unreadable, generation-replaced, or unreclaimed retired state merely to make room; every recognized retired entry still counts toward the same hard retained-state capacity until safely removed. Cross-process create transactions serialize through an owner-identity-checked root `capacity.lock` across prune/recheck and status publication, and state inventory treats a live capacity lock as an uninstall blocker.
|
|
91
91
|
|
|
92
92
|
The runner:
|
|
93
93
|
|
|
@@ -294,7 +294,7 @@ Worker-name mutation is a separate identity transition. Existing state rejects a
|
|
|
294
294
|
|
|
295
295
|
The local `RelayConnection` treats proxy selection, transport construction, WebSocket open, authentication, end-to-end readiness, and outage recovery as separate states. The shared proxy module maps WebSocket targets to standard HTTP(S) environment-proxy resolution, honors `NO_PROXY`, rejects non-HTTP(S) proxy schemes, and creates the proxy agent without exposing its URL or credentials. Invalid proxy configuration is a fatal configuration error rather than a retryable outage.
|
|
296
296
|
|
|
297
|
-
A connection-attempt deadline terminates sockets stuck in `CONNECTING`. After open, the daemon sends `hello`; `hello_ack` establishes an authenticated relay generation and starts liveness monitoring, but does not resolve startup or advertise readiness. The Worker then sends a random `relay_probe`. Its result must traverse the local runtime dispatcher and `sendForSession` on that generation before `ready_ack` marks the connection usable. The daemon rejects ordinary tool calls before that state and rejects a premature readiness acknowledgement without locally recorded probe delivery. Independent handshake and readiness deadlines terminate candidates that authenticate but cannot return results. Once authenticated, `RelayLiveness` sends a protocol-level WebSocket Ping every five seconds. The full ten-second Pong deadline begins only after the sender write callback proves the control frame left the local queue. Ping/application-send completions are fenced to the exact current WebSocket generation, so a callback from a superseded socket cannot mutate the replacement connection; a Ping callback that never completes has its own thirty-second local dispatch bound and becomes `relay_transport_send_timeout` even when the receive direction remains active. Reaching that first-stage deadline no longer immediately kills a ready WSS: `relay-transport-confirmation.mjs` opens one bounded fifteen-second application-confirmation window and emits the existing JSON heartbeat only after the connection is fully ready. A later protocol Pong or the explicit JSON application `pong` is bidirectional transport proof and clears suspicion; unrelated application messages update receive-side liveness only and cannot prove that daemon-to-Worker writes are succeeding. `ResilientRelayConnection` concurrently prewarms signed HTTPS in standby without Worker-side takeover; confirmed WSS recovery stops that prewarm, while a real disconnect upgrades it to exact-generation takeover and preempts any stale standby request. A scheduling-responsive true black hole therefore remains bounded to one five-second probe interval plus ten-second Pong response and fifteen-second independent confirmation, while a single ten-to-fifteen-second persistent-flow stall no longer becomes an avoidable reconnect storm. A detected local event-loop stall cancels remote suspicion and follows the separate recovery-grace branch rather than being counted as network failure. Transport Ping remains active during authenticated probing, but application heartbeat/confirmation is gated on verified readiness because the Worker probing state accepts only the readiness-probe result. Same-instance reconnect also performs explicit call-ownership reconciliation: the Worker sends its still-waiting IDs; the daemon snapshots the union of active calls and unacknowledged results, completes replacement-channel readiness, then returns `resume_calls_ack.missing_ids` only for IDs absent from that ownership union. Those IDs alone may receive one same-ID transport redelivery inside the original remaining deadline; if that cannot be done safely they settle retryably with `side_effects_started=false`. Active calls and retained terminal results continue on existing ownership, and no possibly executed tool call is automatically replayed. A separate JSON application heartbeat remains at twenty-five seconds and retains a seventy-five-second application-silence timeout, so protocol-level Pong cannot mask a Worker application path that has stopped responding. On the Worker, authenticated application-heartbeat activity refresh is synchronous and its JSON `pong` is queued before Durable Object alarm inspection or mutation; heartbeat and terminal-result paths then own one explicit coalesced alarm schedule. Persistent deadline work therefore cannot sit in front of application-liveness acknowledgement or be scheduled twice through a hidden touch helper. The Worker's ninety-second daemon-liveness deadline remains an independent wider fallback across Durable Object hibernation. Worker error codes `daemon_transport_error` and `daemon_liveness_timeout` are connection-recovery signals, not protocol incompatibility: they terminate only the current socket, run normal disconnect cleanup, and enter bounded reconnect. Local protocol-watchdog expiry is classified separately as `relay_transport_timeout`; local application-silence expiry retains `relay_heartbeat_timeout`. The same Worker classification is applied to the close frame if the preceding error frame is lost. Failure to send `hello`, or failure to deliver a readiness-probe result because its socket/session ended, follows the same transport-recovery path rather than being promoted to authentication or protocol failure. Unknown Worker errors, authentication rejection, malformed readiness sequencing, and identity/version mismatch remain fatal. Outage reminders run on their own exponential-backoff timer rather than depending on another transport callback. Each new authenticated hello also carries a schema-versioned, bounded
|
|
297
|
+
A connection-attempt deadline terminates sockets stuck in `CONNECTING`. After open, the daemon sends `hello`; `hello_ack` establishes an authenticated relay generation and starts liveness monitoring, but does not resolve startup or advertise readiness. The Worker then sends a random `relay_probe`. Its result must traverse the local runtime dispatcher and `sendForSession` on that generation before `ready_ack` marks the connection usable. The daemon rejects ordinary tool calls before that state and rejects a premature readiness acknowledgement without locally recorded probe delivery. Independent handshake and readiness deadlines terminate candidates that authenticate but cannot return results. Once authenticated, `RelayLiveness` sends a protocol-level WebSocket Ping every five seconds. The full ten-second Pong deadline begins only after the sender write callback proves the control frame left the local queue. Ping/application-send completions are fenced to the exact current WebSocket generation, so a callback from a superseded socket cannot mutate the replacement connection; a Ping callback that never completes has its own thirty-second local dispatch bound and becomes `relay_transport_send_timeout` even when the receive direction remains active. Reaching that first-stage deadline no longer immediately kills a ready WSS: `relay-transport-confirmation.mjs` opens one bounded fifteen-second application-confirmation window and emits the existing JSON heartbeat only after the connection is fully ready. A later protocol Pong or the explicit JSON application `pong` is bidirectional transport proof and clears suspicion; unrelated application messages update receive-side liveness only and cannot prove that daemon-to-Worker writes are succeeding. `ResilientRelayConnection` concurrently prewarms signed HTTPS in standby without Worker-side takeover; confirmed WSS recovery stops that prewarm, while a real disconnect upgrades it to exact-generation takeover and preempts any stale standby request. A scheduling-responsive true black hole therefore remains bounded to one five-second probe interval plus ten-second Pong response and fifteen-second independent confirmation, while a single ten-to-fifteen-second persistent-flow stall no longer becomes an avoidable reconnect storm. A detected local event-loop stall cancels remote suspicion and follows the separate recovery-grace branch rather than being counted as network failure. Transport Ping remains active during authenticated probing, but application heartbeat/confirmation is gated on verified readiness because the Worker probing state accepts only the readiness-probe result. Same-instance reconnect also performs explicit call-ownership reconciliation: the Worker sends its still-waiting IDs; the daemon snapshots the union of active calls and unacknowledged results, completes replacement-channel readiness, then returns `resume_calls_ack.missing_ids` only for IDs absent from that ownership union. Those IDs alone may receive one same-ID transport redelivery inside the original remaining deadline; if that cannot be done safely they settle retryably with `side_effects_started=false`. Active calls and retained terminal results continue on existing ownership, and no possibly executed tool call is automatically replayed. A separate JSON application heartbeat remains at twenty-five seconds and retains a seventy-five-second application-silence timeout, so protocol-level Pong cannot mask a Worker application path that has stopped responding. On the Worker, authenticated application-heartbeat activity refresh is synchronous and its JSON `pong` is queued before Durable Object alarm inspection or mutation; heartbeat and terminal-result paths then own one explicit coalesced alarm schedule. Persistent deadline work therefore cannot sit in front of application-liveness acknowledgement or be scheduled twice through a hidden touch helper. The Worker's ninety-second daemon-liveness deadline remains an independent wider fallback across Durable Object hibernation. Worker error codes `daemon_transport_error` and `daemon_liveness_timeout` are connection-recovery signals, not protocol incompatibility: they terminate only the current socket, run normal disconnect cleanup, and enter bounded reconnect. Local protocol-watchdog expiry is classified separately as `relay_transport_timeout`; local application-silence expiry retains `relay_heartbeat_timeout`. The same Worker classification is applied to the close frame if the preceding error frame is lost. Failure to send `hello`, or failure to deliver a readiness-probe result because its socket/session ended, follows the same transport-recovery path rather than being promoted to authentication or protocol failure. Unknown Worker errors, authentication rejection, malformed readiness sequencing, and identity/version mismatch remain fatal. Outage reminders run on their own exponential-backoff timer rather than depending on another transport callback. Each new authenticated hello also carries a schema-versioned, bounded relay diagnostic summary, including `previous_ready_inbound_silence_ms` so the silent half-open interval before close is not collapsed into the shorter close-to-ready outage duration. The local connection additionally keeps `recent_outages`, a newest-first in-memory ring capped at eight completed reconnect episodes. The ring contains only bounded outage numbers/timestamps/durations, close/error classes, prior-ready duration/silence, coarse application-route class, and connection-stage timings; a protocol/application Pong that clears transport suspicion without rebuilding WSS is not a completed outage and stays only in heartbeat diagnostics. Because the hello is sent before the recovering WebSocket can receive its final `ready_ack`, its carried ring cannot yet contain that current episode. The Worker therefore sanitizes the prior ring into the probing attachment and, only when promotion proves end-to-end readiness, synthesizes the now-completed current episode from the already bounded scalar fields before marking `outage_active=false`. Authenticated `server_info.daemon.relay_transport` can consequently retain several sub-warning-threshold reconnects instead of losing all but the latest one, without logging them by default, persisting a tool transcript, exposing endpoints/interfaces/DNS data, or treating near-miss liveness suspicions as outages.
|
|
298
298
|
|
|
299
299
|
Reconnect uses bounded exponential backoff with jitter. Brief self-healing interruptions are debug-only. An unresolved outage is promoted to a rate-limited warning after a grace period, and recovery produces one summary. Raw close codes and reason strings remain debug-only.
|
|
300
300
|
|
package/docs/AUDIT.md
CHANGED
|
@@ -1,5 +1,55 @@
|
|
|
1
1
|
# Security and privacy audit notes
|
|
2
2
|
|
|
3
|
+
## 2026-08-27 beta.146 live start-job result-boundary follow-up
|
|
4
|
+
|
|
5
|
+
**Beta.145 is rejected by live behavior even though its underlying durable work remained healthy.** The exact candidate passed the 131-task frozen plan, candidate/install-only validation, guarded Worker/login-daemon activation, and the activated-package OAuth canary. The canary's initiating hosted `start_job` nevertheless returned public `internal_error`. Reusing its idempotency key with a deliberately different plan produced the expected conflict containing the already-bound managed-job ID; reading that exact ID proved the original canary job had completed successfully. A second harmless zero-exit hosted `start_job` independently reproduced the same public `internal_error`, while `exec_command` continued to use the shared two-second settlement path successfully. This isolates the regression to ordinary managed-job result construction rather than Worker/OAuth availability, the managed-job runner, process-carrier settlement, relay transport, or resource admission.
|
|
6
|
+
|
|
7
|
+
**The causal bug was an invalid JSON value introduced after durable acceptance.** `settleManagedJobAcceptance()` spread the accepted managed-job projection and the optional read result, then explicitly reassigned four process-carrier-only fields: `execution_mode`, `source_tool`, `execution_timeout_seconds`, and `retry_safety`. Those fields exist on `exec_command`/`run_process`/`run_local_command` durable-process acceptance but not on ordinary `start_job`; the explicit assignments therefore created own properties whose values were JavaScript `undefined`. `normalizeToolResult()` deliberately rejects `undefined`, functions, symbols, bigint, and non-finite numbers instead of silently mutating the public result. The local tool executor consequently converted an otherwise successful accepted/settled job response into `internal_error` before relay result delivery. The fix removes the redundant reassignments: the original accepted-object spread already preserves process-only fields when they genuinely exist and creates nothing when they do not.
|
|
8
|
+
|
|
9
|
+
**The previous regression test stopped one layer too early.** Beta.145 directly invoked the bound hosted `start_job` handler and asserted the expected terminal/recovery fields, but did not pass that object through the production JSON result boundary. Beta.146 extends that test through `normalizeToolResult()` and requires the recovery envelope plus `follow_up_read_required=false` to survive normalization. This would deterministically fail on beta.145's `undefined` properties. The public generation-15 contract itself does not change, so schema generation remains 15; package/runtime identity advances to beta.146 because the shipped implementation changes.
|
|
10
|
+
|
|
11
|
+
**Release consequence.** No beta.145 candidate acceptance is written. Beta.146 must repeat frozen full verification, exact candidate preparation/install-only proof, guarded activation, activated-package OAuth canary, a harmless live short-`start_job` probe that returns a JSON-valid terminal response in the initiating call, broader live relay/version observation, acceptance, and exact-head provider gates. npm publication remains the sole separate current-task owner authorization boundary. A later published beta.146 registry activation would start a new seven-day major-prerelease soak; neither beta.144 nor rejected beta.145 time carries forward.
|
|
12
|
+
|
|
13
|
+
## 2026-08-27 beta.145 repeated awake interruption observability and event-density follow-up
|
|
14
|
+
|
|
15
|
+
**Live beta.144 evidence separates a remaining physical transport class from a separate host-settlement amplifier.** The current daemon process accumulated three real WebSocket outage generations while the machine was awake. The latest retained episode was a brief close-1006 recovery after a much longer inbound-silence interval; the preceding two sub-warning-threshold reconnects had to be reconstructed indirectly because runtime state kept only cumulative `outage_count` plus the latest outage scalars. Full Wi-Fi baseline comparison falsified Wi-Fi degradation as either a necessary or sufficient condition, visible Tailscale path-change events did not consistently precede the relay silence, and the Karing system extension emitted no matching upstream timeout/reset evidence. An earlier independent event had already aligned relay 1006 with an unrelated HTTPS `unexpected EOF`, so shared system-network/VPN/TUN failure remains a supported class, but the available evidence still does not identify a specific tunnel implementation, proxy node, ISP, Cloudflare edge, or upstream hop as the initiator of every awake reset.
|
|
16
|
+
|
|
17
|
+
**The missing causal artifact was multi-episode history, not another reconnect timeout.** `RelayConnection` now retains at most eight completed reconnect episodes in memory, newest first. Each `recent_outages` entry contains only an outage sequence number, bounded disconnect/ready timestamps and duration, retry count, enumerated close/transport-error classes, previous-ready duration/silence, coarse application proxy-route class, and bounded DNS/TCP/TLS/WebSocket connection-stage timings. It contains no endpoint, interface name, address, DNS answer, certificate, close reason, account/call identity, tool argument, or result. A first-stage Pong timeout that is disproved by protocol/application confirmation without replacing the socket is not inserted, so liveness near misses remain distinguishable from actual WSS generations. The daemon hello sanitizes the existing ring before transmission. Because that hello precedes the recovering connection's final `ready_ack`, Worker promotion alone has the evidence needed to synthesize the just-completed current episode; it prepends that bounded entry and caps the attachment at eight instead of leaving authenticated `server_info` one reconnect behind.
|
|
18
|
+
|
|
19
|
+
**The same investigation exposed a narrower event-density gap rather than a defect in the existing process carriers.** The one-step `exec_command`/`run_process`/`run_local_command` path already waits up to two seconds for a newly accepted durable job to settle, eliminating many helper-plus-`read_job` pairs. The investigation still produced many explicit short multi-step `start_job` submissions followed immediately by `read_job`, which is precisely the shape the existing host-event-density guidance tries to avoid. Beta.145 factors that optional settlement routine into a managed-job helper and routes hosted `start_job` through it. A terminal job returns its result in the original response with `follow_up_read_required=false`; an active or temporarily unreadable job preserves the original acceptance/recovery envelope and continues through normal paced `read_job`. Local/stdio submission is unchanged, and no execution, dependency, admission, reconnect, replay, or six-hour step deadline is shortened or extended. Machine Bridge still cannot observe the exact external ChatGPT host condition that terminates a turn, so lower event density is a bounded mitigation for a reproduced amplifier rather than a claim to have fixed an unobservable host timeout.
|
|
20
|
+
|
|
21
|
+
**Release consequence.** The owner continues to report the interruption class during the beta.144 soak, so it is treated as blocking rather than recording soak success. These shipped runtime/catalog/documentation changes advance package identity to beta.145 and hosted tool schema generation to 15. Beta.145 must repeat frozen verification, exact candidate preparation/install-only proof, guarded Worker/service activation, activated-package OAuth canary, live relay and short-`start_job` verification, acceptance, and exact-head provider checks. npm publication remains the sole separate current-task owner authorization boundary. If beta.145 is later published and registry-verified activation completes, the major-prerelease seven-day soak begins again from that activation; elapsed beta.144 soak time is not carried forward.
|
|
22
|
+
|
|
23
|
+
## 2026-08-27 beta.144 dependency-state recovery and CI fixture follow-up
|
|
24
|
+
|
|
25
|
+
**Exact-main provider CI separated a test-harness race from a packaged Windows recovery defect.** Ubuntu/full twice timed out waiting for a Worker WebSocket `tool_call`, first with a five-second fixture bound and again after a test-only increase to ten seconds. Source tracing showed eight fixtures triggered `currentMcpCall()`/`toolCallRequest()` before calling `waitForWsMessage()`. The HTTP fetch starts immediately, while the waiter has no historical message buffer, so a fast relay dispatch could arrive before listener registration and be lost forever. Registering the waiter first removed all eight request-before-listener sites; three local Worker integration runs passed and hosted Ubuntu/full then passed. This is test evidence, not proof that production Worker dispatch was dropping messages.
|
|
26
|
+
|
|
27
|
+
**The same provider cycle exposed a distinct production inconsistency on Windows.** A dependency-wait upstream runner was killed as part of the recovery fixture. Its runner-exit reconciliation observed a transient `permission_denied` and correctly scheduled another bounded recovery attempt, but the independently waiting downstream read the upstream status during that filesystem window. `waitForManagedJobDependencies()` converted the single read exception immediately into permanent `dependency_unavailable`, so the downstream could fail before the upstream recovery state machine converged. The two components therefore disagreed about whether the same Windows sharing/atomic-replacement condition was transient.
|
|
28
|
+
|
|
29
|
+
**Beta.144 makes dependency-state availability bounded rather than optimistic.** Only secure dependency reads classified as `permission_denied`, `identity_changed`, or generic `resource_unavailable` receive a per-dependency 45-second monotonic grace while the job remains `queued/dependency_wait`. A successful read clears the grace. Persistent failure beyond that window still becomes `dependency_unavailable`; missing state, integrity failure, witness/identity mismatch, staged state, and other invalid evidence remain immediate fail-closed conditions. Deterministic fake-clock coverage proves both one-error recovery and persistent-error expiry, while the full managed-job integration suite still proves autonomous runner recovery and dependency-failure propagation.
|
|
30
|
+
|
|
31
|
+
**Release consequence.** This repair changes packaged managed-job orchestration after beta.143 activation/acceptance, so beta.143 cannot be published from that evidence. Its acceptance record is removed, package/runtime identity advances to beta.144, and hosted tool schema generation advances to 14 because `start_job.depends_on` observable failure/recovery semantics changed. Beta.144 requires a fresh frozen full gate, candidate, install-only proof, guarded activation, activated-package OAuth canary, live observation, acceptance, and exact-head provider gates. npm registry publication remains the sole explicit current-task owner authorization boundary.
|
|
32
|
+
|
|
33
|
+
## 2026-08-27 beta.143 sleep-correlated relay continuity review
|
|
34
|
+
|
|
35
|
+
**The post-beta.142 incident record falsified the assumption that long `connection_reset` outages were primarily a VPN/TUN fault.** The same Mac retained both privacy-safe relay outage events and bounded `pmset` sleep history. A programmatic interval intersection over the fifteen recovered August 23 outages found thirteen with majority sleep overlap; aggregate outage time was 33,390 seconds and aggregate sleep overlap was 32,980 seconds (98.8%). The two no-sleep exceptions were short, roughly sixty-three and ten seconds. This establishes two distinct classes: system suspension explains the overwhelming majority of the historical long close-to-ready intervals, while short awake resets remain real transport interruptions whose upstream cause is not observable from Machine Bridge. The Karing-local-proxy live experiment was also confounded by a 224-second maintenance sleep beginning seconds after the service transition, so it does not prove that the local proxy itself is unreliable.
|
|
36
|
+
|
|
37
|
+
**The existing runtime assertion was present during a sleep it was supposed to bridge.** Power assertions showed the beta.142 `/usr/bin/caffeinate -i -w <daemon-pid>` child alive before and throughout an Idle Sleep -> maintenance-sleep sequence. The local `caffeinate(8)` contract distinguishes `-i` (prevent idle system sleep) from `-s` (prevent system sleep, valid only on AC). A bounded live probe on the same AC-powered host with `caffeinate -i -s -t 3` showed both `PreventUserIdleSystemSleep=1` and `PreventSystemSleep=1`. Beta.143 therefore adds `-s` to the one fixed assertion adapter used by authorized relay handlers, remote process sessions, and account-backed managed-job runners. It does not lengthen the existing thirty-minute inactivity grace, does not assert that `-s` is effective on battery, and does not claim to override explicit sleep or lid-close behavior.
|
|
38
|
+
|
|
39
|
+
**Diagnosis now records causality-compatible overlap instead of asking operators to correlate by eye.** `system-sleep-diagnostics.mjs` remains the only bounded `pmset` parser. In addition to the existing event-loop stall match, it intersects the most recently completed `last_disconnected_at` -> `last_ready_at` relay interval with the recent sleep intervals, merging overlaps so time is never double-counted. The owner-only result carries fixed classification plus outage duration, sleep-overlap duration/ratio, and match count; it contains no interface, endpoint, call identity, tool name, argument, result, raw power-log line, or proxy value. A majority overlap says only that system suspension dominated the observed relay outage and that a socket reset is aftermath evidence, not that every transport reset is caused by sleep.
|
|
40
|
+
|
|
41
|
+
**The user-visible `UNKNOWN / TaskGroup` remains an external-host settlement boundary, not a newly invented local timeout defect.** The observed awake interruption lasted about nine seconds. The durable release job continued without runner recovery and was later read by the same `job_id`; Worker continuity did not record a contemporaneous public request abort. Source tracing confirms pending-call reconnect preserves the original absolute deadline rather than pausing it, and the native MCP response path emits `: connected` immediately plus five-second outer keepalives. The forty-second hosted `read_job` default was previously established by live host evidence and is intentionally not shortened based on an unobservable connector error. Beta.142's recent transient recovery discovery remains the correct cross-turn fallback when a host-owned response still disappears.
|
|
42
|
+
|
|
43
|
+
**Release consequence.** Production runtime bytes, owner-visible diagnostic semantics, catalog guidance, and packaged documentation change, so the tree advances to beta.143 and hosted tool schema generation 13. Beta.142 GitHub prerelease evidence remains valid only for its exact bytes and cannot authorize beta.143. Fresh frozen verification, candidate preparation/install-only proof, guarded activation, activated-package canary, live power/relay diagnostics, acceptance, and exact-head provider gates are required before GitHub source publication. npm publication remains separately owner-authorized for the exact version.
|
|
44
|
+
|
|
45
|
+
## 2026-08-27 beta.142 recent process recovery discovery review
|
|
46
|
+
|
|
47
|
+
**The beta.141 retention reserve preserved evidence but did not guarantee that a later host turn could discover it.** Live owner inventory reproduced the boundary at the hard 512-state cap: 496 ordinary durable terminal records and the full 16 recent transient terminal reserve were retained. The bounded 50-record `list_jobs.jobs` window correctly remained durable-first, so recent one-step process helpers were absent from that primary list. Reading one of those omitted helpers by its already-known `job_id` still succeeded and returned the terminal result. The remaining defect was therefore discovery of retained recovery identity after a host/tool boundary that lost the prior tool response, not durable execution, terminal persistence, or retention.
|
|
48
|
+
|
|
49
|
+
**Beta.142 adds a second bounded projection instead of weakening the durable-first primary ordering.** `list_jobs.jobs` remains capped at 50 and keeps unreadable, active, staged, and durable terminal state ahead of transient helper history. `recent_process_recovery` independently returns at most 16 recent transient terminal jobs that passed the existing authority filter, remain inside the existing thirty-minute recovery grace, and were omitted from the primary window. The projection is built from the same validated persisted status and existing public job projection; it does not return step results or introduce another private state store. Saturated integration coverage proves that all sixteen retained recent helpers remain discoverable outside a primary window containing no transient entries, while the internal retention class remains absent from public job objects.
|
|
50
|
+
|
|
51
|
+
**The security and lifecycle boundary is intentionally unchanged.** The new array is managed-job recovery inventory, not an MCP terminal-result store, response replay identifier, conversation/session binding, or host-delivery acknowledgement. It cannot initiate a ChatGPT turn or prove that an external host rendered the previous terminal response. Authority filtering runs before either inventory projection, the retained-state cap remains 512, the primary response window remains 50, the recent transient reserve remains sixteen results for thirty minutes, and helper history still loses retention priority before ordinary durable terminal history outside that reserve. Hosted schema generation advances to 12 because the public `list_jobs` result semantics changed.
|
|
52
|
+
|
|
3
53
|
## 2026-08-26 beta.141 retained-result recovery and Windows main-CI follow-up
|
|
4
54
|
|
|
5
55
|
**The repeated user-visible interruption investigation exposed a separate recovery-evidence defect even while the underlying durable jobs remained healthy.** Owner diagnostics showed the retained managed-job store at 512 entries with 510 ordinary durable terminal records and only two transient terminal records. Two short read-only `exec_command` helpers were accepted and returned job IDs, but their follow-up `read_job` calls a few seconds later returned typed `not_found`. The cause is deterministic in the previous retention ordering: every `transient_process` terminal had lower eviction priority than every ordinary durable terminal, independent of completion age. At saturation, each new helper could therefore evict the previous helper result before the host consumed it. That does not explain or make observable the external ChatGPT turn-termination decision, but it removes recovery evidence precisely when a host/tool interruption makes that evidence most valuable.
|
package/docs/LOGGING.md
CHANGED
|
@@ -67,7 +67,7 @@ Brief network interruptions are expected on laptop network changes, Worker deplo
|
|
|
67
67
|
- failure to receive `hello_ack` within the handshake deadline, or `ready_ack` within the independent end-to-end readiness deadline, terminates the candidate socket and retries;
|
|
68
68
|
- authenticated transports request protocol-level WebSocket Ping every five seconds. A sender callback starts the full ten-second Pong deadline only after actual local dispatch. If that deadline expires on a fully ready WSS, the relay runtime records `relay.transport.suspect`, opens one fifteen-second application-confirmation window, sends a JSON heartbeat, and prewarms HTTPS in standby instead of killing the socket immediately. A later protocol Pong or explicit JSON application `pong` records `relay.transport.recovered` and preserves WSS; ordinary inbound tool/control traffic remains receive-side evidence and cannot clear transport suspicion. Only confirmation expiry records `relay.transport.confirmation_failed` and closes as `relay_transport_timeout`. A thirty-second local Ping-dispatch failure remains the distinct `relay.transport.send_timeout` / `relay_transport_send_timeout` path even if unrelated inbound traffic continues. Send-completion callbacks are exact-WebSocket-generation fenced, so a callback from a superseded socket cannot mutate current relay diagnostics or confirmation state. These anomaly events contain only bounded timing/state fields;
|
|
69
69
|
- a separate periodic twenty-five-second application heartbeat refreshes Worker daemon activity and retains a seventy-five-second application-silence timeout; it begins only after end-to-end readiness, so authenticated probing cannot send a message type that the Worker probing state does not accept. Protocol-level Pong therefore cannot mask a Worker application path that has stopped replying. The Worker queues the heartbeat's JSON `pong` before Durable Object alarm inspection or mutation, then performs one explicit coalesced schedule, so storage latency is not allowed to sit ahead of application-liveness acknowledgement;
|
|
70
|
-
- a late local transport-watchdog tick is classified as `runtime.event_loop.stall`, sends a fresh transport probe, and defers disconnect for a bounded recovery interval instead of being mislabeled as immediate remote failure; a macOS sleep/wake interval may legitimately produce this warning without a daemon fault. The heartbeat snapshot retains the lag of the last actual stall separately from the rolling maximum. Owner `diagnose_runtime` can compare that end time and duration with a bounded fixed `pmset` sleep-history projection and reports `matched_system_sleep` only when both dimensions agree within the fixed tolerance; an unmatched stall remains unclassified rather than being labeled synchronous JavaScript blockage by elimination. An authorized, schema-valid remote tool call enters the bounded macOS idle-sleep guard only after policy/account/operation authorization and argument validation succeed; the shared assertion remains active for the full handler lifetime, remote process sessions extend it through child settlement, and the fixed thirty-minute rolling inactivity grace starts only after the last owned daemon-side activity settles, with each new authorized activity cancelling pending release and restarting the full grace after settlement. The guard snapshot records coarse activity/grace/release times and release reason without tool arguments or conversation identity. If the fixed macOS assertion child cannot be established, `runtime.idle_sleep_guard.unavailable` records only a coarse `error_class`, never argv, paths, PID, tool name, tool content, session identity, or job identity. Remote account managed-job runners use the same fixed assertion primitive but own it themselves only after runner-claim confirmation and persisted account ownership validation; local managed jobs do not acquire this remote-continuity assertion. The remote runner fallback stderr diagnostic is the fixed text `managed job idle-sleep assertion unavailable` plus a sanitized coarse `error_class`, without job name/id, workspace path, argv, environment, or captured output.
|
|
70
|
+
- a late local transport-watchdog tick is classified as `runtime.event_loop.stall`, sends a fresh transport probe, and defers disconnect for a bounded recovery interval instead of being mislabeled as immediate remote failure; a macOS sleep/wake interval may legitimately produce this warning without a daemon fault. The heartbeat snapshot retains the lag of the last actual stall separately from the rolling maximum. Owner `diagnose_runtime` can compare that end time and duration with a bounded fixed `pmset` sleep-history projection and reports `matched_system_sleep` only when both dimensions agree within the fixed tolerance; an unmatched stall remains unclassified rather than being labeled synchronous JavaScript blockage by elimination. The same diagnostic also intersects the most recent completed relay disconnect interval with that bounded sleep history and reports only outage/overlap timing plus a fixed classification; `majority_system_sleep_overlap` means host suspension dominated the observed outage, so a retained `connection_reset` is aftermath evidence rather than sufficient independent-network evidence. An authorized, schema-valid remote tool call enters the bounded macOS idle-sleep guard only after policy/account/operation authorization and argument validation succeed; the shared `/usr/bin/caffeinate -i -s -w <owner-pid>` assertion remains active for the full handler lifetime, remote process sessions extend it through child settlement, and the fixed thirty-minute rolling inactivity grace starts only after the last owned daemon-side activity settles, with each new authorized activity cancelling pending release and restarting the full grace after settlement. The `-s` assertion is effective only on AC power; diagnostics report only that the fixed child requests it. The guard snapshot records coarse activity/grace/release times and release reason without tool arguments or conversation identity. If the fixed macOS assertion child cannot be established, `runtime.idle_sleep_guard.unavailable` records only a coarse `error_class`, never argv, paths, PID, tool name, tool content, session identity, or job identity. Remote account managed-job runners use the same fixed assertion primitive but own it themselves only after runner-claim confirmation and persisted account ownership validation; local managed jobs do not acquire this remote-continuity assertion. The remote runner fallback stderr diagnostic is the fixed text `managed job idle-sleep assertion unavailable` plus a sanitized coarse `error_class`, without job name/id, workspace path, argv, environment, or captured output.
|
|
71
71
|
|
|
72
72
|
A WebSocket close code such as `1006` means the transport ended without a normal close handshake, but it does not identify who initiated termination: Machine Bridge's own liveness recovery calls `terminate()` when a transport/send timeout is confirmed, and that local hard close can surface as 1006. Diagnose the cause from `last_close_category`, transport-confirmation/send-timeout evidence, and retained network milestones rather than treating 1006 itself as proof of a remote/network-initiated close. If it recovers inside ten seconds, the warning-level service log is intentionally silent and the authenticated `daemon.relay_transport` snapshot is the post-event evidence surface. It is useful for debug diagnosis but not useful as the default user message. It is not evidence that the daemon process restarted. Worker `daemon_transport_error` / `daemon_liveness_timeout` messages and their 1012 close frames are likewise retryable connection conditions, not upgrade instructions. Only an unknown/incompatible Worker error, authentication failure, or identity/version mismatch may produce the fatal protocol/configuration log and daemon exit. Default logs therefore describe the affected layer, duration, classification, and recovery behavior rather than printing raw close envelopes.
|
|
73
73
|
|
|
@@ -153,7 +153,7 @@ Each managed job has owner-only runner diagnostic logs. Child-step output is ret
|
|
|
153
153
|
|
|
154
154
|
`network_route` describes only Machine Bridge's application-level proxy decision. `system-network-stack` does **not** mean a direct physical path: an operating-system VPN, TUN, packet tunnel, DNS interceptor, or endpoint-security product may still carry the connection. `network_route_scope` therefore remains `application-proxy-selection-only`.
|
|
155
155
|
|
|
156
|
-
During an outage, remote-owner `diagnose_runtime.runtime.relay` and local stdio `server_info.runtime.relay` expose bounded live fields: outage count/start/duration, attempts, last close category/code, coarse transport error class plus strict allowlisted `last_transport_error_reason`, last disconnect/ready time, prior ready duration, prior ready inbound-silence duration, the thirty-second WSS connect budget, bounded `last_connect_milestones_ms`, transport-probe queue/dispatch/Pong state, second-stage transport-confirmation timing/recovery state, bounded sender backlog bytes, HTTPS fallback active/warming state, `https_fallback_last_takeover_ms`, and next retry timing. `https_fallback_last_takeover_ms` retains the bounded WSS-close-to-verified-HTTPS-ready interval even after WSS later reclaims primary ownership, so a longer WebSocket outage is not misreported as the same duration of whole-bridge unavailability when HTTPS recovered earlier. The fallback status separates `last_success_at` (any successful signed HTTP exchange, including standby) from `last_ready_at` (only verified ready ownership), so a successful standby poll can no longer be misread as a completed failover. Connect milestones are relative durations only and never include hostnames, addresses, DNS answers, certificates, proxy endpoints, or credentials. The current-attempt milestones and separate `last_failed_connect_*` fields are retained independently so a successful retry cannot erase the immediately preceding failed DNS/TCP/TLS/upgrade evidence; `last_transport_error_ready` and `last_transport_error_authenticated` state whether the retained transport error occurred after channel authentication/readiness. Worker-side `daemon.websocket.closed` records only a bounded close code plus `was_clean`; raw peer close reasons are deliberately omitted. An already-dispatched Worker call detached from a failed channel is retained only until the smaller of reconnect grace and that call's original remaining absolute deadline. If it settles without rebinding, the public bounded error distinguishes `original call deadline expired during reconnect` from a true full `reconnect grace expired`; neither diagnostic includes tool arguments, paths, account identity, endpoint data, or result content, and the distinction does not extend the hosted deadline. Resource-coordinator snapshot contention is likewise classified separately from execution failure: when the bounded diagnostic cannot acquire a transaction/staging lock, it reports retryable `unavailable`, `reason=coordinator_busy`, and `snapshot_available=false`. This says the diagnostic snapshot was unavailable under contention; it does not relax admission policy or claim that the underlying host pressure is Green. `outage_duration_ms` is the close-to-ready recovery interval; `last_ready_inbound_silence_ms` is the pre-close interval since the preceding ready transport last proved inbound activity. After recovery, authenticated remote `server_info.daemon.relay_transport` retains the bounded preceding episode supplied during the current connection handshake, including `previous_ready_inbound_silence_ms` and brief interruptions below the default warning threshold. Promotion to a ready socket sets `outage_active=false`, extends the outage duration through actual readiness, preserves the preceding healthy-ready duration and inbound-silence evidence across failed candidates, canonicalizes timestamps, and accepts only enumerated coarse operational error classes; it does not claim that the recovered connection remains in outage. The `local_authority_revocation_retry` category is deliberately not diagnosed as a network failure: a sustained warning directs the operator to local authority, process-session, and managed-job state while the retained Worker revocation retries on reconnection; ordinary transport categories retain network/Worker troubleshooting guidance. On macOS, `diagnose_runtime` may also return a coarse default-route class, `operating_system_interception` boolean,
|
|
156
|
+
During an outage, remote-owner `diagnose_runtime.runtime.relay` and local stdio `server_info.runtime.relay` expose bounded live fields: outage count/start/duration, attempts, last close category/code, coarse transport error class plus strict allowlisted `last_transport_error_reason`, last disconnect/ready time, prior ready duration, prior ready inbound-silence duration, the thirty-second WSS connect budget, bounded `last_connect_milestones_ms`, transport-probe queue/dispatch/Pong state, second-stage transport-confirmation timing/recovery state, bounded sender backlog bytes, HTTPS fallback active/warming state, `https_fallback_last_takeover_ms`, and next retry timing. `https_fallback_last_takeover_ms` retains the bounded WSS-close-to-verified-HTTPS-ready interval even after WSS later reclaims primary ownership, so a longer WebSocket outage is not misreported as the same duration of whole-bridge unavailability when HTTPS recovered earlier. The fallback status separates `last_success_at` (any successful signed HTTP exchange, including standby) from `last_ready_at` (only verified ready ownership), so a successful standby poll can no longer be misread as a completed failover. Connect milestones are relative durations only and never include hostnames, addresses, DNS answers, certificates, proxy endpoints, or credentials. The current-attempt milestones and separate `last_failed_connect_*` fields are retained independently so a successful retry cannot erase the immediately preceding failed DNS/TCP/TLS/upgrade evidence; `last_transport_error_ready` and `last_transport_error_authenticated` state whether the retained transport error occurred after channel authentication/readiness. Worker-side `daemon.websocket.closed` records only a bounded close code plus `was_clean`; raw peer close reasons are deliberately omitted. An already-dispatched Worker call detached from a failed channel is retained only until the smaller of reconnect grace and that call's original remaining absolute deadline. If it settles without rebinding, the public bounded error distinguishes `original call deadline expired during reconnect` from a true full `reconnect grace expired`; neither diagnostic includes tool arguments, paths, account identity, endpoint data, or result content, and the distinction does not extend the hosted deadline. Resource-coordinator snapshot contention is likewise classified separately from execution failure: when the bounded diagnostic cannot acquire a transaction/staging lock, it reports retryable `unavailable`, `reason=coordinator_busy`, and `snapshot_available=false`. This says the diagnostic snapshot was unavailable under contention; it does not relax admission policy or claim that the underlying host pressure is Green. `outage_duration_ms` is the close-to-ready recovery interval; `last_ready_inbound_silence_ms` is the pre-close interval since the preceding ready transport last proved inbound activity. After recovery, authenticated remote `server_info.daemon.relay_transport` retains the bounded preceding episode supplied during the current connection handshake, including `previous_ready_inbound_silence_ms` and brief interruptions below the default warning threshold. Promotion to a ready socket sets `outage_active=false`, extends the outage duration through actual readiness, preserves the preceding healthy-ready duration and inbound-silence evidence across failed candidates, canonicalizes timestamps, and accepts only enumerated coarse operational error classes; it does not claim that the recovered connection remains in outage. The `local_authority_revocation_retry` category is deliberately not diagnosed as a network failure: a sustained warning directs the operator to local authority, process-session, and managed-job state while the retained Worker revocation retries on reconnection; ordinary transport categories retain network/Worker troubleshooting guidance. On macOS, `diagnose_runtime` may also return a coarse default-route class, `operating_system_interception` boolean, privacy-bounded `runtime.idle_sleep_guard` state (`supported`, `enabled`, `active`, `grace_ms`, `requests_system_sleep_prevention_on_ac`, `last_error_class`), bounded `runtime.system_sleep`, and `runtime.relay_outage_analysis`. That diagnostic is returned on demand and is not promoted to default logs; interface names, IP addresses, DNS answers, proxy endpoints/credentials, Worker endpoints, tool arguments, and results remain absent. `relay.outage.active` and `relay.outage.recovered` carry the existing safe relay fields.
|
|
157
157
|
|
|
158
158
|
Schema 4 is strict NDJSON. Before daemon startup, both active log files are opened as owner-only regular single-link files. A schema change clears both only after validation and commits the marker only after the transition succeeds. A symlink, multiple-hard-link inode, permission error, or marker-write failure blocks startup rather than mixing formats or repeatedly erasing evidence.
|
|
159
159
|
|
package/docs/MANAGED_JOBS.md
CHANGED
|
@@ -13,9 +13,9 @@ can therefore stop after step 2 or 3 and leave local or remote temporary state b
|
|
|
13
13
|
|
|
14
14
|
Machine Bridge managed jobs reduce this failure mode by accepting the complete execution and cleanup plan in one call. After acceptance, an independent local runner owns the lifecycle. It does not depend on the MCP socket remaining connected.
|
|
15
15
|
|
|
16
|
-
Recovery inventory is deliberately ordered by recoverability rather than recency alone. `list_jobs` returns at most 50 records,
|
|
16
|
+
Recovery inventory is deliberately ordered by recoverability rather than recency alone. `list_jobs.jobs` returns at most 50 primary records, with unreadable state, active jobs, staged plans, and durable terminal history ahead of transient one-step helper history so a burst of short helpers cannot push an older long-running or pre-spawn-waiting job out of that bounded primary window. A separate `recent_process_recovery` array returns at most 16 authority-visible public handles for recent one-step terminal jobs that are still inside the thirty-minute recovery reserve but omitted from `jobs`; callers can recover the `job_id` there and continue with `read_job`. It contains no step output or internal retention class and is recovery inventory, not a polling or MCP replay/session surface. Owner/local listings also include only aggregate recent creation/churn counts; the internal `transient_process` retention class remains hidden per job.
|
|
17
17
|
|
|
18
|
-
Remote one-step `exec_command`, `run_process`, and `run_local_command` requests still become durable managed jobs before execution.
|
|
18
|
+
Remote one-step `exec_command`, `run_process`, and `run_local_command` requests still become durable managed jobs before execution. Hosted `start_job` now shares the same short two-second initial-settlement check after its explicit multi-step plan has been durably accepted. A job that reaches terminal state inside that window can therefore return its managed-job result in the original tool response with `follow_up_read_required=false`, eliminating the usual immediate second `read_job` event. A job that remains active keeps the same durable `job_id` and recovery envelope with `follow_up_read_required=true`; execution timeouts, dependency waiting, the pre-spawn resource-admission allowance, finally behavior, and later recovery semantics are unchanged. Local/stdio `start_job` does not add this hosted response wait.
|
|
19
19
|
|
|
20
20
|
This mechanism does **not** bypass host, operating-system, or endpoint-security policy. A job snapshots its execution policy/environment at acceptance; later profile changes affect new jobs, while accepted jobs continue until completion or explicit cancellation. The initial `start_job` request must still be permitted, and every local child process remains subject to local security controls.
|
|
21
21
|
|
|
@@ -127,7 +127,7 @@ A job terminal transition is not a single best-effort write. The runner first at
|
|
|
127
127
|
|
|
128
128
|
If the runner writes a valid terminal result but exits before terminal status is committed, the manager reconstructs status from that result before considering recovery. If a dead runner instead leaves a present but non-terminal, cross-job, or otherwise invalid result object, recovery fails closed and retains the state for inspection rather than treating that corruption as an absent result and overwriting it with a cleanup retry. Lifecycle status is likewise an explicit enum, not a `not-active => terminal` inference: an unknown on-disk status is retained as an integrity failure, is not scrubbed or capacity-evicted as a completed job, blocks destructive removal through active-job inventory, and cannot be acknowledged as already finished by cancellation/revocation. While the runner is still alive in the narrow result-first/status-second settlement window, `read_job` also projects a coherent terminal status from that already-durable result in memory rather than returning an active outer status beside a terminal nested result; cleanup remains conservatively reported as pending until the persisted status catches up. This read-side projection never writes runner-owned state. These rules prevent a completed finally sequence from being replayed or observably mixed merely because the status write has not happened yet. If no valid terminal result exists, recovery may still repeat finally work, so finally steps must remain idempotent.
|
|
129
129
|
|
|
130
|
-
Unexecuted staged plans can contain stdin, environment values, and temporary scripts. They expire after 24 hours, but expiry is itself a per-job state transition: pruning must acquire the job transition lock, re-read the current status, and only then commit the non-executing terminal record through the same result-first persistence boundary. A concurrent cancel/other transition therefore wins or finishes before expiry rather than racing it. Terminal retention begins from that terminal record's `finished_at`, not from the older staged directory age, so a long-abandoned draft is not immediately deleted in the same pass that first expires it. When seven-day retention or capacity pressure eventually retires a complete job directory, removal is also generation-bound: the inspected directory is atomically quarantined under an internal `retired_job_*` generation name before recursive deletion. That internal namespace is deliberately outside the public `job_...` ID grammar, so normal list/read/lock paths cannot reinterpret cleanup state as a job. A crash after rename leaves a recognizable cleanup record. Later pruning reclaims it only when the encoded filesystem generation still matches. Verification failure never renames quarantine back onto the public job pathname; the retired evidence stays isolated for a later safe retry. Any malformed reserved `retired_job_*` name, generation mismatch, wrong type, or unreadable retired entry remains a privacy-bounded destructive-state blocker without exposing its internal filename/device/inode through ordinary job diagnostics. Every recognized retired entry still counts toward the 512-state retained-state capacity until safely removed, so namespace separation cannot become a capacity bypass. The public `job_...` namespace is reserved just as strictly: a matching name with the wrong filesystem type is retained as `unreadable`, counts toward the same capacity, blocks destructive inventory, and is shown only to owner/local diagnostics. New deterministic job admission securely inspects an existing target before any capacity eviction, so a dangling link or other invalid target cannot consume retained terminal history before the request fails. Ordinary completed-job metadata may remain under the separate seven-day retention policy. The retained-state hard cap is 512. `list_jobs` deliberately returns at most 50
|
|
130
|
+
Unexecuted staged plans can contain stdin, environment values, and temporary scripts. They expire after 24 hours, but expiry is itself a per-job state transition: pruning must acquire the job transition lock, re-read the current status, and only then commit the non-executing terminal record through the same result-first persistence boundary. A concurrent cancel/other transition therefore wins or finishes before expiry rather than racing it. Terminal retention begins from that terminal record's `finished_at`, not from the older staged directory age, so a long-abandoned draft is not immediately deleted in the same pass that first expires it. When seven-day retention or capacity pressure eventually retires a complete job directory, removal is also generation-bound: the inspected directory is atomically quarantined under an internal `retired_job_*` generation name before recursive deletion. That internal namespace is deliberately outside the public `job_...` ID grammar, so normal list/read/lock paths cannot reinterpret cleanup state as a job. A crash after rename leaves a recognizable cleanup record. Later pruning reclaims it only when the encoded filesystem generation still matches. Verification failure never renames quarantine back onto the public job pathname; the retired evidence stays isolated for a later safe retry. Any malformed reserved `retired_job_*` name, generation mismatch, wrong type, or unreadable retired entry remains a privacy-bounded destructive-state blocker without exposing its internal filename/device/inode through ordinary job diagnostics. Every recognized retired entry still counts toward the 512-state retained-state capacity until safely removed, so namespace separation cannot become a capacity bypass. The public `job_...` namespace is reserved just as strictly: a matching name with the wrong filesystem type is retained as `unreadable`, counts toward the same capacity, blocks destructive inventory, and is shown only to owner/local diagnostics. New deterministic job admission securely inspects an existing target before any capacity eviction, so a dangling link or other invalid target cannot consume retained terminal history before the request fails. Ordinary completed-job metadata may remain under the separate seven-day retention policy. The retained-state hard cap is 512. `list_jobs.jobs` deliberately returns at most 50 primary records per response so deeper recovery history does not inflate the ordinary MCP inventory window. One-step remote process carriers created by `exec_command`, `run_process`, and `run_local_command` are internally marked with the low-cardinality `transient_process` retention class; that marker is not part of the public job projection and contains no argv, path, output, or credential data. Under capacity pressure, completed transient process records are reclaimed before explicit managed-job terminal history whenever such transient records are available, except for the fixed newest-16/thirty-minute immediate recovery reserve. A retained recent process terminal that the durable-first primary `jobs` window omits may appear in `recent_process_recovery`, which is independently capped at 16 authority-visible public job handles and carries no step output or retention metadata. This prevents diagnostic/helper-command churn from preferentially destroying the recovery result of a long explicit managed job while keeping disk/privacy state and both inventory windows bounded. Active or staged plans that declare `depends_on` additionally pin those referenced retained job records against time- or capacity-based pruning until the dependency-bearing plan is terminal; if dependency protection cannot be read safely, capacity pruning fails closed rather than guessing that no dependency exists. A valid `job_id` whose directory has expired or been capacity-retired returns typed `not_found` rather than a generic execution failure. That absence means only that the retained record is unavailable; callers must not infer that the underlying operation never executed or blindly replay its side effects. A minimal-environment plan also launches its detached runner with a minimal control environment. Full parent-environment inheritance occurs only when the accepted plan explicitly captured that policy.
|
|
131
131
|
|
|
132
132
|
## Job-scoped temporary files
|
|
133
133
|
|
|
@@ -225,7 +225,7 @@ Use top-level `depends_on` when one managed job must wait for one or more earlie
|
|
|
225
225
|
}
|
|
226
226
|
```
|
|
227
227
|
|
|
228
|
-
At acceptance, each dependency is bound to its current durable job identity (`job_id`, plan hash, and creation generation). Staged dependencies and dependencies that have already failed are rejected before the new job is accepted. While any accepted dependency remains active, the dependent job stays `queued` with `current_phase=dependency_wait`; `dependency_total` and `dependency_pending_count` report progress, and the dependent job has not spawned its main child, entered process resource admission, or materialized private registered-resource/temporary-file execution copies. As upstream jobs settle, a hosted `read_job` long-poll wakes when the pending count changes. All dependencies succeeding releases the normal main-step sequence and materializes execution inputs immediately before they can be needed. An upstream job that later ends unsuccessfully makes the dependent job terminal with `result.error_class=dependency_failed` and bounded `dependency_failure` evidence instead of waiting minutes or hours for an impossible artifact. If that failure still requires declared `finally_steps`, their resource/temporary-file inputs are materialized only when the cleanup phase begins.
|
|
228
|
+
At acceptance, each dependency is bound to its current durable job identity (`job_id`, plan hash, and creation generation). Staged dependencies and dependencies that have already failed are rejected before the new job is accepted. While any accepted dependency remains active, the dependent job stays `queued` with `current_phase=dependency_wait`; `dependency_total` and `dependency_pending_count` report progress, and the dependent job has not spawned its main child, entered process resource admission, or materialized private registered-resource/temporary-file execution copies. As upstream jobs settle, a hosted `read_job` long-poll wakes when the pending count changes. A transient dependency-state read failure classified as `permission_denied`, `identity_changed`, or generic `resource_unavailable` keeps the dependent in `dependency_wait` for a fixed 45-second recovery grace instead of converting one Windows sharing/atomic-replacement race into a permanent dependency failure. A successful read clears that per-dependency grace immediately. Persistent unavailability past the grace still fails closed as `dependency_unavailable`; `not_found`, integrity failure, witness/identity mismatch, staged state, and other invalid dependency evidence remain immediate failures. All dependencies succeeding releases the normal main-step sequence and materializes execution inputs immediately before they can be needed. An upstream job that later ends unsuccessfully makes the dependent job terminal with `result.error_class=dependency_failed` and bounded `dependency_failure` evidence instead of waiting minutes or hours for an impossible artifact. If that failure still requires declared `finally_steps`, their resource/temporary-file inputs are materialized only when the cleanup phase begins.
|
|
229
229
|
|
|
230
230
|
Terminal persistence is result-first: the durable `result.json` is written before `status.json` is changed to the matching terminal state. Dependency polling therefore treats a valid terminal result for the same job as a read-only terminal projection during that narrow publication window instead of waiting for an unrelated manager read to repair the status file. It does not rewrite the upstream record. Once the dependent itself becomes terminal, `dependency_pending_count` is zero because the dependent is no longer waiting; `dependency_failure` identifies the upstream terminal cause when the job failed.
|
|
231
231
|
|
|
@@ -310,7 +310,7 @@ Never place a secret directly in `argv`, `env`, `stdin`, a temporary file's `con
|
|
|
310
310
|
|
|
311
311
|
Per-workspace jobs are stored below the owner-only profile directory. Active jobs retain an owner-only plan for crash recovery. Plan, status, result, runner identity, and lock updates use flushed atomic replacement or complete-before-visible exclusive claims. Transition/recovery locks contain ownership tokens and process start time and are removed only when their file snapshot still matches. After a terminal status is committed, the full plan is deleted, including argv, stdin, embedded temporary-file content, and resource source paths.
|
|
312
312
|
|
|
313
|
-
Retained public job data contains bounded status and redacted results. The hard capacity is 512 retained-state slots across ordinary job directories plus any recognized internal retired-cleanup entries that have not yet been safely removed; `list_jobs` still returns at most 50
|
|
313
|
+
Retained public job data contains bounded status and redacted results. The hard capacity is 512 retained-state slots across ordinary job directories plus any recognized internal retired-cleanup entries that have not yet been safely removed; `list_jobs.jobs` still returns at most 50 primary records per call, while `recent_process_recovery` may add at most 16 recent authority-visible process recovery handles that were omitted from that primary window. Terminal jobs are normally retained for up to seven days from their persisted `finished_at` settlement time, but when capacity is full the oldest safely removable terminal records are evicted to reserve a slot for a new job. Staged drafts expire after 24 hours, dependency-referenced records are protected while an active/staged dependent plan still needs them, and active, staged, unreadable, or abnormal retired state is never evicted merely to make room; if all 512 slots are occupied, new job creation returns a retryable `limit_exceeded` error whose owner-only details include coarse `retained_state`, `retired_state`, and `retired_unreadable` counts. `list_jobs.retained` remains the number of visible ordinary jobs even when only 50 are returned. That bounded inventory is recovery-first: unreadable, active, and staged state stays first, durable terminal managed-job results precede transient one-step process terminals, and only then does helper history fill the remaining response window. Owner/local responses additionally include the coarse capacity summary plus `durable_terminal` and `transient_terminal` counts so a full 512-state store can be distinguished from helper churn without exposing job identities, paths, arguments, or output. Delegated non-owner responses omit the global capacity summary. These counts improve recovery visibility only; Machine Bridge still cannot observe whether an external host consumed a terminal result or rendered a final assistant response. Private runtime copies are removed after the finally phase. Runner stdout/stderr log files contain only runner-level diagnostics; step output is not written to those operational logs.
|
|
314
314
|
|
|
315
315
|
The detached runner records a structured owner record containing PID and process start time. Recovery rejects a reused PID instead of treating an unrelated process as the active runner. Numeric-only runner records are invalid. Initial runner publication uses provisional then committed atomic generations; a claim reader is explicitly coupled to that publication protocol and may therefore retry a transient `MBM_IDENTITY_CHANGED` observation for four 1 ms attempts before failing. Each retry repeats the complete secure read and identity validation; this exception does not apply to generic or destructive file reads. Recovery-lock handoff preserves a random ownership token, and the runner removes only a lock whose PID, token, and file snapshot still match. A recovery runner does not gain authority to write terminal evidence merely by confirming its runner claim: it must first complete recovery-lock handoff. Failure or ambiguity in that bootstrap phase leaves the prior `interrupted` status and plan intact for a later safe retry instead of publishing `recovery_failed` and scrubbing recovery material. The handoff has a 30-second monotonic ownership-settlement budget. Timeout and cancellation terminate the process group/tree, retain a referenced forced-escalation timer, and clean descendants that ignore graceful termination before the runner exits.
|
|
316
316
|
|